Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
A customer opens a support chat and asks, “Can you explain why invoice INV-1042 was marked late?” The support assistant replies with a confident explanation and cites an internal note that belongs to a different company, including a line like “Per ACME renewal terms, waive fee on net-60.” The support UI never showed any ACME tickets to the agent or the customer, but the assistant still pulled that note into the answer. Nothing exotic happened here. The customer only needed to ask a normal question in the normal channel, and the application did the rest.
The system shape that produced the leak is small enough to fit in our heads. An end user sends a message to the support application, the application fetches related records from a database, the application assembles a prompt that includes the user’s message and the retrieved rows, and the application sends that prompt to a model provider. The model returns text, and the application displays the text as the assistant answer. The failure is observable as a cross-tenant citation. The application returned one tenant’s private data to another tenant’s user.
We can only fix the leak once we can point to the exact boundary where untrusted data crosses into a trusted action. A trust boundary is the seam where data from one side must not automatically gain the privileges of the other side, because the two sides have different incentives and different access. In our support assistant, the end user sits on one side and the application’s data access sits on the other, but the boundary repeats at every hop that can change what data gets read or revealed.
Across the full request path, we treat four things as untrusted because each can carry instructions or content that the application did not author. The end user message is untrusted because the user controls it. The database rows are untrusted because the database can return unexpected rows if the query or filter is wrong, and because retrieved text can contain instruction-shaped content that does not belong in an answer. The model output is untrusted because the model follows the text it received and can propose reads, writes, or disclosures the application never intended. Only the application’s own configuration, such as a static allowlist of tool names or a hard-coded list of columns to select, starts trusted, and even configuration only stays trusted if the application loads it from a controlled source. The trust-boundary diagram below names each hop and the data flowing across it. Let’s mark where each untrusted input enters the flow: