Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
It is day three of an engagement with a mid size insurance broker, and the head of claims ops is asking for a chat tool that can answer adjusters fast. They want it to quote the latest policy rules, explain exceptions, and point to the exact page in the right handbook. In the same meeting, legal adds one more constraint. The tool must never show content from a client contract unless the adjuster is assigned to that account.
This is where customer-specific retrieval-augmented generation (RAG) shows up in forward deployed work. Customer-specific RAG is a setup where the AI system answers by first pulling facts from that customer’s own sources, then writing the response using those facts. It is customer-specific because the sources change over time and because the customer, not us, decides what counts as an acceptable answer.
In practice, we are not shipping a generic chatbot. We are shipping an answer workflow tied to real documents, real permissions, and real business outcomes. For the broker, “correct” means the answer matches the current handbook version, cites it, and respects account level access rules even when the question is ambiguous.
The core FDE move is to treat RAG as an end to end system, not a feature. If we only think about model prompts, we miss the parts that make it safe and dependable in production. The rest of this lesson stays anchored on this broker case, because every decision you make will connect back to what an adjuster sees in the moment they need an answer.
Production failures in RAG are rarely dramatic. They are usually quiet, plausible, and expensive. The adjuster gets an answer that sounds right, acts on it, and only later does someone notice the claim was handled under the wrong rule.
Consider this situation. The broker uploads a new “Water Damage Exceptions” bulletin on Monday. By Wednesday, the chat tool still answers using last month’s handbook because the content pipeline did not re-ingest the new file. Nothing errors, and the system may even cite a real page number. It is just the wrong page number for the current policy set.
Another common break is access boundaries. If the retrieval step can see a document, the AI system can end up summarizing it, even if the end user should not. This is why permissions are not a UI detail in RAG. They are part of the retrieval design, and they must be enforced before any text is handed to the model.
Silent misses are the other killer. A question like “Can we waive the inspection if the damage is under 2k?” might require a specific exception clause. If retrieval fails to find it, the model will often try to be helpful anyway, filling the gap with a generic sounding rule. Without measurement, you will not notice this drift until the business complains.
Explore the main ways these failures connect to business impact.
Assume drift, prove quality
If you are not tracking quality over time, quality is already changing and you just cannot see it yet.
These failure modes set up the real FDE question. What do we need to decide, and in what order, so this system stays correct as documents, users, and rules change?
Customer-specific RAG work is a chain of decisions where each link changes what the next link can safely do. If we skip ahead to embeddings or chunk sizes, we end up rewriting the system later when a business rule shows up that should have been a first class requirement.
Start with requirements that are specific enough to test. For the broker, we write acceptance criteria like “must cite the exact handbook section” and “must refuse when the user lacks access.” Those are not vague goals. They are pass fail checks we can turn into an evaluation set.
Next is data. We map the sources that will feed retrieval and how they change. Handbooks in SharePoint, bulletins from email, account contracts in a CRM export. We also map who is allowed to see what, because that permission model will shape indexing and query time filtering.
Then comes retrieval design. Retrieval is how the system selects which document snippets to show the model. The big decision is whether we can filter by permissions at query time, or whether we need separate indexes per segment to avoid leakage. For the broker, account contracts often push you toward strict filtering, because a single mistake is a serious incident.
Grounding is what we do to keep the model tied to retrieved text. Grounding usually includes rules like “answer only from provided context” and “include citations.” It also includes what the system should do when context is missing, like asking a clarifying question or returning “I cannot find that in current policy docs.”
Evaluation is where we make reliability measurable. An evaluation is a repeatable set of test questions with expected behaviors, run before and after changes. For this broker, we care about citation accuracy, refusal correctness for restricted docs, and coverage on high risk claim scenarios.
Deployment decisions turn all of this into an operating system. We decide how ingestion runs, how often indexes refresh, how we log requests safely, and how we roll out updates. Iteration is then a loop with guardrails. When adjusters report a miss, we decide if it was a data gap, a retrieval gap, or an instruction gap, and we fix the right layer.
See the stages as one connected system of choices.