Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A chatbot that answers from your own documents is only as dependable as the knowledge behind it. Knowledge management for an AI chatbot is an ongoing operating practice: choose trustworthy sources, prepare them for retrieval, govern who and what the bot can reach, measure answer quality, and refresh content when facts or user needs change. Retrieval-augmented generation (RAG), the common pattern for answering from organization-specific information, does not remove this work. It moves the work to the content and the evaluation, where it has to be done continuously.

How RAG works, and where it breaks

In a RAG system, the application first retrieves content that is relevant to a user’s question, then passes that content to a language model as context and asks the model to answer from it. Microsoft’s RAG and Generative AI guidance in Azure AI Search describes this as a way to ground answers in information the model was not trained on.

That two-step design gives you two places to fail, and they need different fixes:

  • Retrieval failure: the document that holds the answer is never returned, or a plausible but wrong passage is returned instead.
  • Generation failure: the right passage is retrieved, but the model ignores it, answers only partly, or adds details that are not in it.

Most of the knowledge management work described below exists to make the first failure rarer and to make the second one visible when it happens.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I structure a knowledge base for an AI chatbot?

Start from the business task and the people who will ask questions, not from the folder of files you have. The sequence that holds up in practice is:

  1. Define the task. Write down what the bot should answer (for example, HR policy questions from employees, or product specification questions from support agents) and what it should refuse or hand off.
  2. Identify authoritative sources and permissions before ingestion. For each source, record the owner, the system of record, and who is allowed to see it. A copy on a shared drive is not an authoritative source unless someone owns it.
  3. Build a representative test set. Collect real or realistic questions, and include questions whose answers are not in the corpus at all. Without the second group you cannot tell whether the bot handles missing knowledge appropriately.
  4. Process files according to their structure. Headings, tables, FAQ pairs, and step-by-step procedures carry meaning in their layout. Extracting plain text can strip the column headers that make a table’s values interpretable.
  5. Split content into semantically useful units. A chunk should be able to answer something on its own: one policy clause, one procedure, one product specification.
  6. Enrich each chunk with metadata. See the field list below.
  7. Embed and index, then test options. Do not assume one chunk size or retrieval method suits every corpus. Compare the options against your representative questions and your actual content.

Metadata worth attaching to each chunk

Metadata lets the retrieval layer filter and rank content, and lets a reviewer trace an answer back to its origin. Add the fields that matter for your corpus:

Field What it does for the chatbot Example value
Title Identifies the document or section a chunk came from “Expense Policy, Section 4: Travel Meals”
Summary Gives the retriever a compact description to match against “Daily meal limits for domestic travel”
Keywords Captures terms users may type that the text does not use “per diem, food allowance”
Source Lets a reviewer open the original and check the answer Name of the policy system and document link
Date Shows how recent the content is Date the policy was last approved
Version Separates current content from superseded content “v3.2”
Access scope Limits which users can receive the chunk “Finance staff only”

Provenance is the non-negotiable field. If an answer cannot be traced to a specific source passage, nobody can check it, and checking is the only way to know whether the chatbot is correct.

Chunking is a test, not a default

Chunk boundaries decide what the retriever can find. A chunk that is too large can bury the one sentence that answers a question under unrelated text. A chunk that is too small can separate a condition from the rule it qualifies. Choose boundaries from the document’s own structure (sections, clauses, steps), then compare two or three approaches on the same test questions and keep the one that retrieves the correct passage most consistently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Microsoft’s guidance on RAG in Azure AI Search puts the principle this way: “RAG quality depends on how you prepare content for retrieval.” That sentence is attributed to the Microsoft Learn document, not to an individual author.

How do I keep chatbot answers up to date?

A knowledge base that was accurate at launch will drift. Prices change, policies are revised, products are discontinued, and new procedures replace old ones. Treat the corpus as maintained information with a set of recurring duties:

  • Name an owner for every source. The owner decides what is current and approves changes. If no one owns a source, it should not be indexed.
  • Track version and age. Use the date and version fields so that content past its review date can be flagged rather than silently served.
  • Review changes at the source. Connect updates to the authoritative system where possible, so a change in the policy repository reaches the index without someone remembering to copy it.
  • Remove or supersede obsolete content. After replacing a document, confirm that the old chunks are gone from the index, not just that the new ones were added. Otherwise the retriever can return both versions.
  • Rerun evaluation after significant updates. A new policy can change answers to questions that were never edited, so run the same test set again.
  • Ask content writers to review answers. Subject-matter owners can read sample answers in their area and spot errors quickly. Repeated poor answers on one topic often point to missing, ambiguous, or outdated documentation rather than a retrieval defect, and the fix is then a writing task.

How do I evaluate a RAG chatbot?

Evaluation is a repeatable loop, not a one-time acceptance test. Run it the same way each time so that results can be compared across changes:

  1. Collect representative questions. Include common questions, hard edge cases, and questions with no answer in the corpus.
  2. Inspect what was retrieved. Record which documents or chunks came back for each question.
  3. Judge retrieval. Are the retrieved passages relevant, and are they sufficient to answer the question?
  4. Judge the response. Is the answer grounded in the retrieved text, and does it cover what the question asked?
  5. Record gaps and user feedback. Note questions with no good source, and any feedback that users or support staff send.
  6. Make one targeted change. Change one thing (a chunk boundary, a metadata field, a source document, or the instructions given to the model) so that cause and effect stay clear.
  7. Rerun the same tests and aggregate the results. Compare against the previous run.

Track retrieval quality separately from response quality. If you combine them, a good answer produced from the wrong passage, or a bad answer produced from the right passage, will look the same in your scores.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluation dimensions

The Microsoft guidance Design and Develop a RAG Solution on Azure (Azure Architecture Center, updated June 30, 2026) names several dimensions that are useful for scoring:

Dimension Question it answers Scored on
Relevance Did retrieval return content related to the question? Retrieved chunks
Completeness Does the answer cover everything the question asked? Response
Groundedness Is every claim in the answer supported by the retrieved text? Response against retrieved chunks
Utilization Did the model actually use the retrieved content, rather than ignoring it? Response against retrieved chunks

Maintain a golden dataset

Running every question against the full corpus is often impractical. A curated golden dataset solves this: a fixed set of questions, each paired with the expected grounded answer and the source passage that supports it. Include the unanswerable questions too, with the expected response being a clear statement that the information is not available. Store the configuration (chunking settings, index version, model, and instructions) alongside each run, so that a drop in scores can be traced to a specific change.

How can I improve my chatbot’s answers?

Improvement starts with a diagnosis. Find the failing question, look at what was retrieved, and then decide whether the problem sits in retrieval, in the content, or in the model’s use of the context. The table below maps common symptoms to likely causes and first fixes:

Symptom Likely cause First fix to try
The correct document exists but is never returned Retrieval or chunking: the passage is buried or poorly described Adjust chunk boundaries, improve title, summary, and keyword fields, then retest options
A relevant passage is returned, but the answer is incomplete The chunk is too narrow, or the question needs content from several chunks Check sufficiency of retrieved content; widen the retrieved set or restructure the source
The answer includes details not present in the retrieved text Grounding failure in generation Tighten the model’s instructions to answer only from the retrieved context; re-score groundedness
The answer is confidently wrong and matches an old policy Stale or superseded content still indexed Confirm old chunks are removed; reindex and rerun the golden dataset
Poor answers recur on one topic with no clear source Content gap or ambiguity in the documentation Ask the content owner to write or clarify the source
The bot answers a question that the corpus cannot answer No handling for missing knowledge Add absent-answer test questions and configure a clear “not available” response or handoff

Make one change at a time and rerun the same tests. Improvements that cannot be attributed to a specific change are difficult to keep.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Governance and security

A chatbot that reads internal content is an access path to that content. Microsoft’s Cloud Adoption Framework guidance on governing and securing AI agents across the organization frames these controls at the organization level. In practice, a working baseline includes:

  • A named owner for the agent and for each knowledge source.
  • An inventory of deployed agents that records purpose, owner, platform, and access scope.
  • Least-privilege access: the agent connects only to the sources it needs.
  • Preservation of user permissions when the agent answers on a user’s behalf, so a user cannot retrieve content they could not open directly.
  • A review of each new source for content, permissions, and security risk before it is connected.
  • Written privacy, data residency, and retention rules for source data, any memory the agent keeps, and logs, with defined deletion and purging processes.
  • Ongoing monitoring, plus tests for prompt injection and data leakage before production and after significant changes.

The right settings depend on your jurisdiction, data classification, and risk tolerance. A public FAQ bot and an HR bot with salary information need different controls.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When RAG is the right tool, and when it is not

Microsoft’s Copilot Studio guidance on enhancing AI responses with retrieval-augmented generation describes RAG as best suited to factual questions and answers, summaries of policies, FAQs, and procedures, and retrieval of specific facts. The same guidance states that RAG is not intended for full-document comparison, policy compliance evaluation, or complex reasoning over long unstructured documents. Read this as a scope boundary for that pattern rather than a universal limit on every AI system. If your questions fall outside the first group, a different design is likely needed.

Choosing an approach

A conventional retrieval pipeline over a single index can be adequate for straightforward question answering. Query decomposition and multi-source reasoning add capability, and also add moving parts. Compare them on the criteria below:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Criterion Single-index retrieval Query decomposition Multi-source reasoning
Source complexity One corpus, consistent format One or more corpora with sub-questions Several sources with different owners and formats
Permission and governance needs One access scope to manage Access checked for each sub-query Access checked for each source, with combined results
Query complexity Single factual questions Questions that break into parts Questions that require combining evidence
Retrieval quality Depends on chunking and metadata Depends on the decomposition step as well Depends on each source and on how results are merged
Latency and operating cost Not stated in the reviewed Microsoft guidance; measure in your own tests Not stated in the reviewed Microsoft guidance; measure in your own tests Not stated in the reviewed Microsoft guidance; measure in your own tests
Implementation complexity Lowest of the three Moderate Highest of the three
Team’s ability to evaluate and maintain Feasible for a small content team Requires testing each decomposition path Requires testing each source and each merge step

If your team cannot yet run the evaluation loop described above, choose the simplest design that answers your questions, and add complexity only when a test shows the simple design failing.

Managed retrieval services

If you want managed retrieval infrastructure rather than building it yourself, Azure AI Search is documented by Microsoft for RAG content preparation and retrieval, as covered in the guidance linked earlier. Evaluation and observability tools are a separate category, and they are useful for the repeatable testing this article describes. Choose tools by how well they fit your sources, governance requirements, and team, and test them against your own golden dataset before relying on them.

Microsoft Engineering’s own example

Microsoft Engineering has published a case study on how it built “Ask Learn,” a RAG-based knowledge service, in How we built “Ask Learn,” the RAG-based knowledge service. It is a useful real-world reference for the structure, maintenance, and evaluation practices covered here.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.