Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
A stakeholder asks for a “clean customer list” you can trust for an email campaign and a dashboard refresh. You have the spreadsheet export, but it contains blanks, mixed date formats, inconsistent country names, and duplicates. The risk is not that cleaning is hard. The risk is that a fast, AI-suggested cleaning step silently changes record counts or meanings, and you own the downstream failure when the campaign targets the wrong people or the dashboard shifts.
In this course, AI is a drafting and acceleration tool. You use it to propose transformations, formulas, Apps Script, or VBA, and to enumerate likely data quality issues. You do not let it decide what to drop, how to impute, or what “duplicate” means. Success criteria are the measurable conditions that define “clean enough” for the business task. You set them up front so cleaning is not an aesthetic exercise.
Your baseline rule is simple. Every AI-generated step must be runnable, reversible, and auditable in Excel or Google Sheets, with evidence you can show in a review.
Assume a named dataset customer_contacts_export.xlsx with 18,240 rows pulled from a CRM export covering 2025-01-01 to 2025-06-30. Each row is a contact. Key fields include email, full_name, signup_date, country, phone, and source. The segmentation dimension the marketing lead cares about is country and source, and the specific question is which countries are growing fastest by monthly signups.
Explore what the raw data issues look like.
Your deliverable is not “a cleaned file”. It is a bundle of artifacts that makes your work reproducible.
You produce a cleaned table where critical identifiers are standardized and parsing rules are consistent. You produce QA checks that prove you did not silently lose rows or create impossible values. You produce a change log that states exactly what was changed, what was dropped, and why.
An LLM is a large language model that generates text, formulas, and code based on patterns in its training and your inputs. It does not observe your workbook state unless you provide it. A prompt is the instruction and context you give the LLM to generate a candidate output. A hallucination is a generated claim or rule that sounds plausible but is not grounded in your data or constraints.
Confident-wrong happens often in cleaning because cleaning requires business definitions. The model can propose a deduplication key like email + full_name, but it cannot know whether shared emails represent families, shared inboxes, or data entry errors. If you accept its rule and delete rows, you created irrecoverable bias in every downstream count and rate.
See which tasks are safe to delegate versus must-verify.