Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
You are the analyst on a feature experiment for a new onboarding screen. The product lead wants a go or no-go by Friday, based on an A/B test that ran for 14 days on 240,000 eligible users, split 50/50. You have event data in a table like experiment_events with one row per user per day, keyed by user_id, date, variant, and outcomes like activated_7d.
AI can speed up your workflow here, but it cannot own your decision. You own the downstream failure if the rollout hurts retention because a metric was misdefined, a segment was excluded, or an assumption was violated. Treat AI as a generator of candidate summaries, code, and checklists that you verify against the data and your experiment design.
Explore the main parts of the process.
An LLM is a text model that generates plausible continuations given your input. In experiments work, that means it can draft analysis plans, write code, and summarize results, even when it is wrong in a way that reads confident. A prompt is the instruction and context you give the model. If you prompt vaguely, you invite the model to fill gaps with guesses, not with evidence.
When the model writes analysis code, that is code generation. You still have to run it, inspect it, and confirm it matches your design. When the model invents a detail not supported by your dataset or logs, that is a hallucination. In A/B testing, hallucinations show up as made-up sample sizes, unearned causal language, or a test choice that silently assumes normality.
Tool use means the model can call approved tools like a SQL runner or notebook kernel. It still does not validate measurement, it just executes instructions. You validate that the executed query matches the experiment population and exclusion rules.
See how an AI readout can go wrong in small but consequential ways.
AI helps most where the work is structured and the verification is crisp. Instrumentation QA is a good example. You can ask AI to generate candidate queries that check for missing variant assignment, unexpected event drops on a specific date, or imbalance in platform distribution. You verify by running the queries and confirming the counts reconcile with your assignment table.
Metric definitions are another leverage point. AI can draft a metric spec for activated_7d and list required exclusions, but you decide the definition and you confirm it matches how the business will interpret success. If the stakeholder thinks activation means completing step 3, but your metric is any session within 7 days, the experiment result will be operationally wrong even if the p-value is small.
AI can also draft an analysis narrative. You still compute the primary estimate, its confidence interval, and the decision relevant segmentation. AI writes the first draft, you own every claim in the final.
Try mapping common tasks to appropriate AI involvement levels.