Imagine you wrote a spam filter as a pile of if statements. If the subject contains FREE, mark it spam. If the sender ends in .ru, mark it spam. If the email contains wire transfer, mark it spam.
For a week, it looks fine, then the inbox starts to drift. New scams arrive that avoid your keywords, and normal emails get flagged because a coworker wrote “free lunch” in a calendar invite. Nothing in the mailbox is “broken” in a crash-the-program sense. The rules are still doing exactly what you wrote, and that is the problem.
With the rule-based filter failing, the next step is a system whose output changes after it sees what you consider correct. A program learns when training adjusts internal settings so the same input email leads to a better spam or not spam output than before training.
To keep this concrete, use a tiny inbox where each email comes with a label, which is the correct answer you want the program to produce such as spam or not_spam. That set of examples is data, meaning recorded inputs paired with their labels. The program does not take whole emails as magic. It uses features, which are noticeable properties you can measure from an email such as “contains the word invoice” or “sender domain ends with .biz”.
See how the same labeled emails can lead to different outputs before and after training.
Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
The key point is that training changes the filter’s behavior without you adding new if statements. The training process uses the labels to push the internal settings toward outputs that match the labels more often, so the same message that used to be misclassified becomes correctly classified after the adjustment.
Once the filter can change after examples, the next question is what kind of change counts as learning. If the program only memorizes the exact training emails, it can get a perfect score on those examples while still failing on the next scam that is worded differently.
Given a new email the filter has not seen before, what could a pattern-based system use to classify it without hand-coding every new trick. It can reuse features that often correlate with spam in the labeled set, like the presence of invoice, unusual sender domains, or a mismatch between the sender name and the domain.
Compare how an exact-match memory, a pattern-based learner, and fixed rules handle one unseen email.