Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
A support ops manager asks for a weekly view of customer complaints by category and reason, split by channel, for Q1 2025. You receive a complaints.csv extract with 48,219 rows spanning 2025-01-01 to 2025-03-31. It includes mixed date formats in created_at, blank reason values, duplicate ticket_ids from system retries, and inconsistent category labels like Billing, billing, and BILLING.
Your deliverable is not just a cleaned file. It is an analysis-ready table plus an audit trail that explains exactly what changed and why. If a cleaning choice is wrong, the analyst owns the downstream error in the KPI deck, the model, or the decision that follows. AI can accelerate drafting transformations, but it cannot accept responsibility for the consequences, so every suggested rule must be made explicit and then ratified by you.
Take a look at what this dataset can look like before and after cleaning.
An LLM is a large language model that generates text from patterns in its training data and the input you provide. In cleaning work, it generates candidate rules and code, but it does not observe your business process, so it can invent plausible but incorrect assumptions about what a null or duplicate means.
A prompt is the instruction you send to the LLM. In practice, your prompt is a specification. If you do not specify invariants like whether row count may change, the AI will fill in silent defaults, and you will inherit them.
A context window is the limited amount of text the LLM can use at once. If you paste only a sample of rows, the AI may miss edge cases that occur outside the sample, such as created_at values that appear only at month end.
A hallucination is generated content that is not supported by your data or instructions. In cleaning, hallucination commonly appears as invented category mappings, guessed missing value reasons, or confident claims that duplicates are safe to drop. You prevent this by requiring evidence in the form of checks you can run and outputs you can diff.
Explore how these terms connect to specific cleaning failure modes.
Start with profiling, because you cannot choose rules you have not measured. Profiling produces counts of nulls per column, distinct value lists for category, duplicate rates for ticket_id, and parsability rates for created_at. AI can summarize a profiling report, but you decide which anomalies are errors versus real business variation.
Next, decide rules before code. For example, “Standardize category to title case using an explicit mapping table” is a rule. “Drop duplicate ticket_id keeping the earliest created_at” is a rule. The code is only the implementation, and you own the rule.
Then generate transforms. AI can draft a pandas script, but you must verify behavior with deterministic checks. Verification means you assert invariants like row counts, uniqueness constraints, and value domain constraints, and you fail the run if they break.
Finally, document. An audit trail is a change log that records what was changed, what was dropped, and what assumptions were applied, in a form a future analyst can replay.
See how the workflow artifacts fit together from profiling to documentation.