Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
An AI chatbot system is the software that takes a user’s message and returns a useful reply, usually by calling a large language model and wrapping that call in product logic like safety checks, formatting, and memory. It includes the parts that accept messages, decide what to send to the model, and deliver the reply back to the user.
To keep the system boundary clear, we separate three roles. The client app shows the conversation and sends messages. The backend owns the chat workflow and data. The model provider runs the model and returns generated text when called.
The simplest version is a single backend service that exposes a POST /chat API. The client sends {conversation_id, user_message}. The service loads the prior turns for that conversation, calls an LLM API with that context, and streams tokens back to the client as they arrive. Let’s anchor this in a one-region baseline flow:
In this baseline, the API service is the decision point. It reads and writes chat history in a chat store so a refresh does not lose the thread, and it calls the LLM API to produce the next assistant message. Logs and metrics capture request timing and errors so we can see when the system slows down or fails.
Even at modest traffic, two signals show up quickly. Users notice latency because time to first visible text determines whether the app feels responsive, and total generation time determines whether they abandon long answers. Cost also becomes visible because every model call consumes tokens, so each extra turn and each extra retrieved snippet increases spend per conversation.
A latency budget is the time we can spend from the user tapping send to the user seeing a usable response. In a streaming chatbot, that splits into time to first token and total time to finish the answer, and the baseline flow has both. Choosing a slower model raises both numbers, while adding more context often raises total time even if first token is similar. Next, we make that latency and cost tension concrete: