Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
A central Node proxy can burn a month of provider budget overnight when a compromised service key hits a permissive completion route and the proxy trusts caller supplied routing like provider=openai and model=$OPENAI_CHAT_MODEL in the request body.
In our running system, AcmeSupportApp calls the gateway on behalf of tenant t_acme_prod, the gateway fans out to $OPENAI_CHAT_MODEL, $ANTHROPIC_CHAT_MODEL, or $LLAMA_CHAT_MODEL, and a shared semantic cache sits near the gateway to save tokens on repeated prompts.
The failure is not a sophisticated model break. The attacker only needs a key that passes whatever shallow check exists today, then sends a high volume of requests with long prompts and large output limits, and the gateway forwards them because no enforceable per tenant budget was reserved before the provider call started.
Once the proxy also streams unfiltered provider output back to the caller, the same compromised key can drive unsafe content into downstream ticket replies because nothing enforces a policy on the generated text at the gateway boundary.
Given that sequence, which stage could have refused the very first request before any spend and before any unsafe output existed.
The gateway is the only component that sees the caller identity and the provider choice at the same time, so the gateway is where we can bind AcmeSupportApp to t_acme_prod and refuse any model or route outside its scope before a provider key is used.
Every other hop is missing a critical piece. The provider only sees the gateway’s provider credential, so the provider cannot distinguish tenants inside our platform, and a cache only sees a key derived from request content, so a cache cannot enforce who is allowed to spend or which model a caller may reach.
Untrusted data enters at multiple points even when the TLS connection is correct. Request headers can lie about identity if we accept caller supplied tenant ids, the request body can lie about routing and model selection if we honor it, and the provider response is untrusted output that can contain policy violating content or schema breaking structured fields. Here is the request path with those trust boundaries and enforcement points marked.
This diagram shows the practical constraint we build the rest of the course around. Only the gateway can authenticate the caller, evaluate a per request policy, reserve quota against a shared counter, and enforce both input and streamed output checks before the response crosses back into the caller’s system.
A gateway stays predictable under attack when every enforcement point has a single refusal code and the code is generated before the next expensive step. We use one refusal map for the whole course so logs, alerts, and client behavior stay consistent across providers.
A missing, invalid, revoked, or compromised caller credential yields 401 Unauthorized.
An authenticated caller that requests a route or model outside its allowlist yields 403 Forbidden, and the gateway refuses before contacting any provider host.
An exhausted rate limit or token and cost quota yields 429 Too Many Requests with Retry-After, because the client needs an unambiguous backoff contract and retries must not amplify spend.
If the gateway cannot safely enforce because the shared counter store is down or an open circuit breaker blocks checks, the gateway yields 503 Service Unavailable rather than failing open and spending without attribution.
If the provider call exceeds the gateway timeout, the gateway yields 504 Gateway Timeout and does not stream partial output as a successful response.
If we require structured output and the provider response fails schema validation, the gateway yields 422 Unprocessable Content so a downstream parser never runs on untrusted fields.
In the opening incident, the first refusal point has to be the gateway authentication plus policy check, before any quota reservation and before any provider call, because later stages no longer have the information needed to bind spend and permissions to t_acme_prod once the request has started flowing outward.