Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build an internal AI assistant on AWS as an identity-aware retrieval-augmented generation (RAG) application—not simply a model connected to a vector store. Amazon Bedrock Knowledge Bases can manage much of the content-ingestion and retrieval workflow, but your application still needs to enforce each employee’s document permissions, handle unsafe or insufficient context, and measure retrieval and answer quality. The architecture below is a practical synthesis of AWS guidance; it is not a single AWS reference implementation.

What the architecture needs to do

RAG retrieves relevant passages from approved enterprise content at question time and supplies them to a foundation model as context. The model can then produce an answer grounded in those passages. Amazon Bedrock Knowledge Bases support both retrieval that an application processes itself and retrieve-and-generate workflows that return a natural-language response with source context. See How Amazon Bedrock knowledge bases work and AWS Prescriptive Guidance on RAG fundamentals.

A production system also needs source validation, document preparation, identity-aware access checks, response handling, traceability, and ongoing evaluation. A useful logical flow is:

  1. Authenticate the employee. Use the organization’s identity provider through the assistant’s application front end.
  2. Resolve identity and policy. Application middleware determines the relevant user or policy attributes and carries them into the retrieval decision.
  3. Prepare approved content. Ingest validated sources with useful metadata, provenance, and permission information.
  4. Retrieve within the access boundary. Obtain relevant passages only after applying the employee’s authorization constraints.
  5. Generate with controlled context. Send the question and authorized passages to the selected foundation model, with appropriate safeguards.
  6. Return a traceable answer. Provide source references where available; when context is insufficient, say so rather than presenting an unsupported answer.
  7. Operate and improve the service. Protect operational and audit events, monitor failures, and evaluate changes against representative examples.

For AWS’s description of Knowledge Base workflows and options, see Retrieve data and generate AI responses with Amazon Bedrock Knowledge Bases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose who operates the RAG pipeline

AWS describes Managed Knowledge Bases and Customer-managed Knowledge Bases. The trade-off is primarily between delegating more infrastructure work and retaining more control over the pipeline; the options do not have identical capabilities.

Decision area Managed Knowledge Base Customer-managed Knowledge Base
Ingestion, indexing, storage, retrieval AWS manages the underlying infrastructure and workflow. Your organization manages the RAG pipeline and vector store.
Parsing and configuration control Less control over underlying ingestion and retrieval infrastructure. More control over ingestion, parsing, indexing, and storage configuration.
Connectors and document permissions Includes capabilities such as connectors and document-level permissions, with documented exceptions; for example, the Web Crawler connector is an exception. Do not assume the same connector or document-permission feature set is available.
Operational ownership Less infrastructure to operate directly, while application authorization and governance remain your responsibility. More direct responsibility for operating and maintaining the pipeline and vector infrastructure.
Best fit to assess When supported connectors and managed capabilities meet requirements and reducing pipeline operations is valuable. When custom ingestion, parsing, indexing, storage, or retrieval control is a requirement.

These distinctions and capability exceptions are described in AWS’s Knowledge Bases feature documentation. Because service features and regional support can change, verify both for the intended deployment region before committing to a design. Do not treat “managed” as proof that access behavior matches your identity model: validate how the actual source permissions and employee groups map to retrieval decisions.

Make employee authorization part of retrieval

The critical security boundary is whether an unauthorized passage can reach model context—not whether a user can see the final answer. If retrieval returns content a person cannot access, subsequent prompting or answer filtering is too late to make the retrieval decision correct.

AWS Prescriptive Guidance recommends carrying application user identity into the knowledge base as metadata so retrieval can apply metadata filtering. One AWS Architecture Blog pattern evaluates policy with Amazon Verified Permissions and translates the decision into a metadata filter for Bedrock retrieval. This is a pattern to evaluate against your own policy model, not a requirement to use that service. See the guidance on identity propagation and Bedrock integration and the multi-tenant Verified Permissions pattern.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Define the access decision. Establish which identity and policy attributes determine access to each source or document, including how changes to group membership and document permissions are reflected.
  2. Carry attributes through middleware. Resolve identity from the authenticated session and pass only the attributes needed to enforce the policy. Do not let a client-supplied filter alone define what the employee may retrieve.
  3. Apply authorization before context construction. Filter retrieval results using the verified identity or policy decision before passages are supplied to the model.
  4. Prove isolation with tests. Test users across departments, roles, and access levels, including negative cases where a matching document must not be returned. Inspect retrieved passages, not only the final prose answer.
  5. Keep an audit trail. Record appropriately protected evidence of the identity and authorization decision, retrieval outcome, and relevant application events so incidents can be investigated.

AWS guidance also calls for least-privilege IAM, private network paths where required, API activity auditing, and monitoring. AWS describes security as a shared responsibility: customer responsibilities depend on the selected services, data, requirements, and applicable laws. See its Bedrock integration security guidance.

Protect the full content-to-answer path

Authorization is necessary, but it is not the only RAG risk. A malicious instruction hidden in an indexed document can affect model behavior even when a user did not submit that instruction. AWS identifies this as indirect prompt-injection risk and recommends controls across ingestion, retrieval, and response generation. Its secure access guidance for generative AI covers these risks and related controls.

  • At ingestion: verify source ownership, provenance, file type, and permissions; validate content before indexing and inspect for malicious or irrelevant material. Maintain metadata and lineage so an answer can be traced to its source.
  • At storage and transfer: apply appropriate access boundaries and encryption. AWS documents KMS options for knowledge-base data processes; TLS for connections to third-party connectors or vector stores depends on the provider supporting TLS. Review the scope in Encryption of knowledge base resources.
  • At retrieval: enforce the identity or policy filter before returning passages, and test for cross-user and cross-department leakage.
  • At inference and response: use appropriately configured guardrails as one layer. AWS describes evaluating both user inputs and model responses, and guardrails can be used with Knowledge Bases. Contextual grounding can help reduce some risks, but it does not eliminate prompt injection.
  • In operations: monitor service and application behavior, protect logs, and define incident response for suspected leakage, malicious sources, or unexpected model behavior.

Guardrails do not replace authorization, source validation, or evaluation, and they do not guarantee privacy, correctness, regulatory compliance, or immunity to prompt injection. Amazon Bedrock documentation says, “We recommend that you continue to test and validate your guardrails to confirm that they meet your requirements.” See How Amazon Bedrock Guardrails works.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate retrieval separately from generated answers

A fluent response can conceal a retrieval failure, and good retrieval does not guarantee a correct answer. Evaluate both stages. Bedrock evaluation supports retrieve-only and retrieve-and-generate RAG evaluation jobs; AWS describes metrics for context relevance and coverage as well as evaluation of generated responses. See Evaluate the performance of Amazon Bedrock resources and Use metrics to understand RAG system performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a representative evaluation set

Version a dataset with realistic questions, expected supporting passages, and expected answers. Include ordinary use as well as difficult cases:

  • Questions whose supporting material exists in sources the user is authorized to access.
  • Permission-boundary cases where a similar passage exists but is not authorized for that user.
  • Stale or conflicting documents, to check whether retrieval and answers expose ambiguity rather than silently choosing an unsupported interpretation.
  • Unanswerable questions, to assess whether the system acknowledges insufficient context.
  • Adversarial examples, including malicious instructions in indexed content.

Track failures by stage

When an answer is wrong, determine whether retrieval missed the right evidence, returned irrelevant or unauthorized evidence, or whether generation mishandled adequate context. Keep retrieval relevance and coverage, answer quality and grounding, permission correctness, and refusal behavior visible as distinct review concerns. Do not treat one aggregate score as evidence that the system is ready for a consequential workflow.

Re-run tests after meaningful changes

Re-evaluate after changes to parsing or chunking, metadata, embeddings, retrieval settings, prompts, guardrail configuration, or model selection. Use human review for consequential workflows and inspect individual failures as well as aggregate results. AWS documents that evaluation jobs require access to supported evaluator models; retrieve-and-generate jobs also require the response generator model, and both must be available in the same Region. Verify current model support and regional availability in the evaluation documentation before implementation.

Plan for production operations

Before launch, define operational targets from the actual workload rather than assuming a generic RAG design implies a particular capacity, latency, availability, or cost. The architecture needs decisions for:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Ingestion health: freshness expectations, handling of failed or partial imports, and how source changes or revoked permissions are reflected in indexed content.
  • Lineage and audit: traceability from answer to retrieved source and protected records sufficient to investigate access and quality incidents.
  • Reliability: latency and availability targets, scaling behavior, rate-limit handling, and incident ownership across application and AWS service dependencies.
  • Failure behavior: a safe response when retrieval returns no authorized evidence or model invocation fails; do not silently substitute ungrounded generated output.
  • Cost visibility: attribution for model invocation and retrieval activity to applications, teams, or use cases where appropriate, then review against actual usage.
  • Change control: version prompts, retrieval configuration, source-processing choices, and evaluation results so a regression can be linked to a change.

No workload-specific capacity estimate, benchmark, cost model, or service-level target is implied by this architecture. Establish those from measured usage and current service terms for the regions and models you select.

Pre-launch decision checklist

  • Does the selected Knowledge Base mode support the required connectors, permissions, parsing controls, and operational ownership?
  • Can the application prove that each employee’s identity and policy attributes constrain retrieval before any passage reaches model context?
  • Have ingestion validation, document lineage, encryption, IAM boundaries, network requirements, and logging been addressed?
  • Does the assistant cite or otherwise identify evidence, and does it handle missing or conflicting context without overstating certainty?
  • Does a versioned evaluation set cover retrieval quality, answer grounding, authorization boundaries, unanswerable requests, and adversarial content?
  • Are regional model availability, evaluation requirements, failure handling, monitoring, and operational ownership confirmed for the intended deployment?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.