Our running workload is a governed enterprise assistant for employees. It answers questions using internal documents and it must cite sources for answers that claim facts. It must also enforce access control so a user only retrieves documents they are authorized to see.
This workload creates two distinct risks that drive platform choice. One risk is model behavior risk, including unsafe outputs, prompt injection, and hallucinated claims. The other risk is data boundary risk, including accidental disclosure of sensitive content through retrieval or logging. We will keep those risks separate because they have different controls and different evidence.
Let’s visualize the trust boundaries and the data paths that must be governed.
Generate custom courses on any topic — with hands-on practice, AI guidance, and visuals built in.
Already have an account?
This lesson sets a decision frame for delivering a governed enterprise assistant with Amazon Web Services (AWS). We will choose between retrieval-augmented generation (RAG), fine-tuning, and different inference options by starting from success metrics and the allowed data boundary. We will treat safety, latency, and cost as first-class acceptance criteria, not afterthoughts.
Amazon Bedrock is a fully managed, serverless service that provides API access to foundation models for generative AI applications and agents. Bedrock abstracts the serving infrastructure, so instance selection and GPU tuning are not customer-controlled decisions. Bedrock also includes managed features for RAG, safety controls, and agent development that live at the application layer.
Amazon SageMaker AI is a managed machine learning platform for building, training, customizing, and deploying models with more infrastructure and workflow control. SageMaker AI inference can host open, proprietary, or custom models by selecting instance types and serving frameworks. This shifts more responsibility to us for capacity planning, deployment configuration, and performance and cost tradeoffs.
A practical way to separate responsibilities is to ask what we want to own. Bedrock is optimized for consuming a model capability behind a managed API with token-based metering. SageMaker AI is optimized for owning the model deployment surface, including containers, instance families, and cost per instance-hour. This ownership choice usually dominates the later RAG versus fine-tuning discussion.
We can make this choice repeatable by following four steps. First, define the outcome and the evaluation method that will decide whether the assistant is acceptable. Second, define the governed data boundary, including what must never be sent to a model provider. Third, define evidence we can collect in production that proves the system stays within policy. Fourth, choose Bedrock, SageMaker AI, or both, based on which boundary we must control.
Let’s inspect the flow from requirements to a defensible service choice.