Build an internal AI assistant on AWS as an identity-aware retrieval-augmented generation (RAG) application—not simply a model connected to a vector store. Amazon Bedrock Knowledge Bases can manage much of the content-ingestion and retrieval workflow, but your application still needs to enforce each employee’s document permissions, handle unsafe or insufficient context, and measure retrieval and answer quality. The architecture below is a practical synthesis of AWS guidance; it is not a single AWS reference implementation.
What the architecture needs to do
RAG retrieves relevant passages from approved enterprise content at question time and supplies them to a foundation model as context. The model can then produce an answer grounded in those passages. Amazon Bedrock Knowledge Bases support both retrieval that an application processes itself and retrieve-and-generate workflows that return a natural-language response with source context. See How Amazon Bedrock knowledge bases work and AWS Prescriptive Guidance on RAG fundamentals.
A production system also needs source validation, document preparation, identity-aware access checks, response handling, traceability, and ongoing evaluation. A useful logical flow is:
- Authenticate the employee. Use the organization’s identity provider through the assistant’s application front end.
- Resolve identity and policy. Application middleware determines the relevant user or policy attributes and carries them into the retrieval decision.
- Prepare approved content. Ingest validated sources with useful metadata, provenance, and permission information.
- Retrieve within the access boundary. Obtain relevant passages only after applying the employee’s authorization constraints.
- Generate with controlled context. Send the question and authorized passages to the selected foundation model, with appropriate safeguards.
- Return a traceable answer. Provide source references where available; when context is insufficient, say so rather than presenting an unsupported answer.
- Operate and improve the service. Protect operational and audit events, monitor failures, and evaluate changes against representative examples.
For AWS’s description of Knowledge Base workflows and options, see Retrieve data and generate AI responses with Amazon Bedrock Knowledge Bases.
#1 Best Overall
Choose who operates the RAG pipeline
AWS describes Managed Knowledge Bases and Customer-managed Knowledge Bases. The trade-off is primarily between delegating more infrastructure work and retaining more control over the pipeline; the options do not have identical capabilities.
| Decision area | Managed Knowledge Base | Customer-managed Knowledge Base |
|---|---|---|
| Ingestion, indexing, storage, retrieval | AWS manages the underlying infrastructure and workflow. | Your organization manages the RAG pipeline and vector store. |
| Parsing and configuration control | Less control over underlying ingestion and retrieval infrastructure. | More control over ingestion, parsing, indexing, and storage configuration. |
| Connectors and document permissions | Includes capabilities such as connectors and document-level permissions, with documented exceptions; for example, the Web Crawler connector is an exception. | Do not assume the same connector or document-permission feature set is available. |
| Operational ownership | Less infrastructure to operate directly, while application authorization and governance remain your responsibility. | More direct responsibility for operating and maintaining the pipeline and vector infrastructure. |
| Best fit to assess | When supported connectors and managed capabilities meet requirements and reducing pipeline operations is valuable. | When custom ingestion, parsing, indexing, storage, or retrieval control is a requirement. |
These distinctions and capability exceptions are described in AWS’s Knowledge Bases feature documentation. Because service features and regional support can change, verify both for the intended deployment region before committing to a design. Do not treat “managed” as proof that access behavior matches your identity model: validate how the actual source permissions and employee groups map to retrieval decisions.
Rank #2
Make employee authorization part of retrieval
The critical security boundary is whether an unauthorized passage can reach model context—not whether a user can see the final answer. If retrieval returns content a person cannot access, subsequent prompting or answer filtering is too late to make the retrieval decision correct.
AWS Prescriptive Guidance recommends carrying application user identity into the knowledge base as metadata so retrieval can apply metadata filtering. One AWS Architecture Blog pattern evaluates policy with Amazon Verified Permissions and translates the decision into a metadata filter for Bedrock retrieval. This is a pattern to evaluate against your own policy model, not a requirement to use that service. See the guidance on identity propagation and Bedrock integration and the multi-tenant Verified Permissions pattern.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
- Define the access decision. Establish which identity and policy attributes determine access to each source or document, including how changes to group membership and document permissions are reflected.
- Carry attributes through middleware. Resolve identity from the authenticated session and pass only the attributes needed to enforce the policy. Do not let a client-supplied filter alone define what the employee may retrieve.
- Apply authorization before context construction. Filter retrieval results using the verified identity or policy decision before passages are supplied to the model.
- Prove isolation with tests. Test users across departments, roles, and access levels, including negative cases where a matching document must not be returned. Inspect retrieved passages, not only the final prose answer.
- Keep an audit trail. Record appropriately protected evidence of the identity and authorization decision, retrieval outcome, and relevant application events so incidents can be investigated.
AWS guidance also calls for least-privilege IAM, private network paths where required, API activity auditing, and monitoring. AWS describes security as a shared responsibility: customer responsibilities depend on the selected services, data, requirements, and applicable laws. See its Bedrock integration security guidance.
Protect the full content-to-answer path
Authorization is necessary, but it is not the only RAG risk. A malicious instruction hidden in an indexed document can affect model behavior even when a user did not submit that instruction. AWS identifies this as indirect prompt-injection risk and recommends controls across ingestion, retrieval, and response generation. Its secure access guidance for generative AI covers these risks and related controls.
Rank #4
- At ingestion: verify source ownership, provenance, file type, and permissions; validate content before indexing and inspect for malicious or irrelevant material. Maintain metadata and lineage so an answer can be traced to its source.
- At storage and transfer: apply appropriate access boundaries and encryption. AWS documents KMS options for knowledge-base data processes; TLS for connections to third-party connectors or vector stores depends on the provider supporting TLS. Review the scope in Encryption of knowledge base resources.
- At retrieval: enforce the identity or policy filter before returning passages, and test for cross-user and cross-department leakage.
- At inference and response: use appropriately configured guardrails as one layer. AWS describes evaluating both user inputs and model responses, and guardrails can be used with Knowledge Bases. Contextual grounding can help reduce some risks, but it does not eliminate prompt injection.
- In operations: monitor service and application behavior, protect logs, and define incident response for suspected leakage, malicious sources, or unexpected model behavior.
Guardrails do not replace authorization, source validation, or evaluation, and they do not guarantee privacy, correctness, regulatory compliance, or immunity to prompt injection. Amazon Bedrock documentation says, “We recommend that you continue to test and validate your guardrails to confirm that they meet your requirements.” See How Amazon Bedrock Guardrails works.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Evaluate retrieval separately from generated answers
A fluent response can conceal a retrieval failure, and good retrieval does not guarantee a correct answer. Evaluate both stages. Bedrock evaluation supports retrieve-only and retrieve-and-generate RAG evaluation jobs; AWS describes metrics for context relevance and coverage as well as evaluation of generated responses. See Evaluate the performance of Amazon Bedrock resources and Use metrics to understand RAG system performance.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBest Value
Build a representative evaluation set
Version a dataset with realistic questions, expected supporting passages, and expected answers. Include ordinary use as well as difficult cases:
- Questions whose supporting material exists in sources the user is authorized to access.
- Permission-boundary cases where a similar passage exists but is not authorized for that user.
- Stale or conflicting documents, to check whether retrieval and answers expose ambiguity rather than silently choosing an unsupported interpretation.
- Unanswerable questions, to assess whether the system acknowledges insufficient context.
- Adversarial examples, including malicious instructions in indexed content.
Track failures by stage
When an answer is wrong, determine whether retrieval missed the right evidence, returned irrelevant or unauthorized evidence, or whether generation mishandled adequate context. Keep retrieval relevance and coverage, answer quality and grounding, permission correctness, and refusal behavior visible as distinct review concerns. Do not treat one aggregate score as evidence that the system is ready for a consequential workflow.
Re-run tests after meaningful changes
Re-evaluate after changes to parsing or chunking, metadata, embeddings, retrieval settings, prompts, guardrail configuration, or model selection. Use human review for consequential workflows and inspect individual failures as well as aggregate results. AWS documents that evaluation jobs require access to supported evaluator models; retrieve-and-generate jobs also require the response generator model, and both must be available in the same Region. Verify current model support and regional availability in the evaluation documentation before implementation.
Plan for production operations
Before launch, define operational targets from the actual workload rather than assuming a generic RAG design implies a particular capacity, latency, availability, or cost. The architecture needs decisions for:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11- Ingestion health: freshness expectations, handling of failed or partial imports, and how source changes or revoked permissions are reflected in indexed content.
- Lineage and audit: traceability from answer to retrieved source and protected records sufficient to investigate access and quality incidents.
- Reliability: latency and availability targets, scaling behavior, rate-limit handling, and incident ownership across application and AWS service dependencies.
- Failure behavior: a safe response when retrieval returns no authorized evidence or model invocation fails; do not silently substitute ungrounded generated output.
- Cost visibility: attribution for model invocation and retrieval activity to applications, teams, or use cases where appropriate, then review against actual usage.
- Change control: version prompts, retrieval configuration, source-processing choices, and evaluation results so a regression can be linked to a change.
No workload-specific capacity estimate, benchmark, cost model, or service-level target is implied by this architecture. Establish those from measured usage and current service terms for the regions and models you select.
Quick Recap
Pre-launch decision checklist
- Does the selected Knowledge Base mode support the required connectors, permissions, parsing controls, and operational ownership?
- Can the application prove that each employee’s identity and policy attributes constrain retrieval before any passage reaches model context?
- Have ingestion validation, document lineage, encryption, IAM boundaries, network requirements, and logging been addressed?
- Does the assistant cite or otherwise identify evidence, and does it handle missing or conflicting context without overstating certainty?
- Does a versioned evaluation set cover retrieval quality, answer grounding, authorization boundaries, unanswerable requests, and adversarial content?
- Are regional model availability, evaluation requirements, failure handling, monitoring, and operational ownership confirmed for the intended deployment?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

