Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Hybrid search for retrieval-augmented generation (RAG) runs two retrieval paths over the same internal corpus. A lexical path matches the words in a query, and a vector path matches the query’s meaning against embedded chunks. The two ranked lists are merged into one, and the top passages go to the language model. Issuing the query is the easy part. The production work is choosing how the lists are fused, proving on your own judged queries that the combination beats each path alone, and making sure every passage returned is one the asking user is allowed to read.

What hybrid search means in a RAG system

In a RAG system, retrieval decides what the model is allowed to read. A generation model can only ground its answer in the passages it receives, so a retrieval miss cannot be repaired further down the pipeline. Each retrieval method also tends to fail in a predictable direction. Lexical retrieval scores documents by how well they match the query’s terms, usually with a BM25-style formula that rewards rare terms appearing often in a passage. Vector retrieval embeds the query and each chunk into the same vector space and returns the chunks whose embeddings sit closest to the query’s.

Hybrid search keeps both paths. OpenSearch and Elastic document it as a single query that combines full-text and vector components. Azure AI Search describes the same pattern as one hybrid request whose two result sets are merged with reciprocal rank fusion (RRF). The product APIs differ, but the design choices underneath are the same: what each path retrieves, how the two result sets are combined, and which filters apply before a passage can reach the model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

BM25 vs vector search for RAG

Neither path wins in general. The Microsoft Azure Architecture Center’s RAG retrieval guidance and the OpenSearch hybrid-search documentation both point to the query mix and the corpus as the deciding factors. The pattern is consistent enough to plan around:

Query type Lexical (BM25-style) Vector (semantic)
Exact identifiers, such as a policy number or product code Strong: the exact token is either present in the passage or it is not Weaker: similarity measures meaning rather than exact characters, so a rare string can be underweighted
Exact names, acronyms and policy titles Strong when the analyzer keeps the token intact Variable: acronyms with several expansions are easy to confuse
Paraphrase, such as “how do I get travel costs back” Weaker: the question may share few words with the policy text Strong: the wording differs but the meaning matches
Concept queries across teams with different vocabularies Weaker: each team’s terms must match literally Stronger: the embeddings can relate the vocabularies
Questions that depend on a version, date or section Neither path decides this alone; metadata filters on version, date and section do

Hybrid retrieval covers both failure directions, but it does not guarantee better ordering. A corpus dense with identifiers may need the lexical path to carry more of the result list; a corpus of differently worded policies may need the vector path to. Which one, and by how much, is a measurement question.

Designing authorization for internal documents

Authorization changes more than any other part of an internal search design. The Azure AI Search hybrid overview lists filters among the capabilities available in the hybrid query context, but it does not prescribe how your identities and document permissions should be modelled. That mapping is yours to design. A workable pattern has five parts:

  • Copy each source system’s access rule onto every chunk at index time, as a list of group or principal identifiers. A chunk inherits the permissions of the document it came from, and it is not indexed without them.
  • At query time, resolve the asking user’s groups from your identity provider and pass them as a filter.
  • Apply that filter inside both the lexical path and the vector path. A filter on only one path lets the other path return passages the user cannot open, and fusion will merge them into the result list.
  • Check the permission again on the returned passages before they reach the model. Retrieval filtering is the first control, not the only one.
  • Log the filter values applied to each query. Auditability differs between platforms, and the log is what lets you answer why a given user was shown a given passage.

Two trade-offs follow. Permission changes reach the index with some delay, and during that delay a revoked user can still retrieve the content, so set a maximum acceptable lag and test revocation against it. Filtering can also leave fewer usable candidates than the query asked for. Count the survivors per user group during testing, and confirm whether your platform applies a filter before or after it selects candidates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I implement hybrid search for internal documents?

Work through the five stages in order. Fusion and evaluation both depend on how the index was built, so changing the schema later usually means reindexing.

1. Inventory and prepare the source corpus

Start with a register of every source system: its formats, document owners, update frequency and authorization rules. Extraction should keep the structure retrieval will later use: titles, section headings, tables where your parser supports them, a stable source identifier, timestamps and access-control metadata. The reviewed sources do not prescribe a parsing stack. Choose one, then measure how it handles your hardest formats, such as scanned PDFs or tables that split across pages.

Define how changes, deletions and permission updates reach the index. The stable identifier matters here. It lets a retrieved passage link back to its authoritative document, and it lets a deletion remove every chunk derived from that document. Treat these three behaviours as acceptance criteria for the pipeline:

  • Edit a document at its source. The new wording appears in both retrieval paths after your sync interval, and the old wording stops appearing.
  • Delete a document at its source. Its chunks and vectors are gone from the index after the same interval.
  • Revoke a user’s access at the source. That user’s queries stop returning the document after the same interval.

2. Chunk and index text with metadata

Chunks should hold enough local context to answer a question on their own, while fitting your model’s context budget and the retrieval design. The reviewed guidance does not give a universal chunk size, overlap or embedding model, and none of these should be fixed before you have measured them on your corpus. A practical test is to compare a few splitting strategies, such as splitting on section boundaries versus fixed token windows, and judge each by whether the correct passage appears in the results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An illustrative field set for each chunk is below. It is a starting point, not a prescribed schema:

  • source_id and chunk_id: link the passage back to its document.
  • title, section_path and text: indexed for lexical retrieval. The Azure guidance notes that preserved titles, content, keywords and entities give the lexical path more to match.
  • embedding: a vector field populated from the chunk text.
  • acl_groups: the permission metadata described in the authorization section.
  • updated_at and embedding_model: support freshness rules and reveal chunks still embedded with an outdated model.

In OpenSearch’s documented example, an ingest pipeline runs a text_embedding processor and writes the resulting vector into a mapped k-NN vector field, while the original text remains indexed for full-text search. Other platforms use their own index definitions for the same pattern, so check each one’s field types. The OpenSearch tutorial’s production guidance on dedicated ML nodes is worth reading before you run embedding inference on the cluster that serves queries.

The embedding model and preprocessing must match on both sides. Microsoft’s guidance is explicit that the same embedding model used for the chunks must embed the query, with the same preprocessing applied. Changing the model therefore means re-embedding the whole corpus, and the embedding_model field tells you which chunks have not been rebuilt yet.

3. Retrieve through both paths

Run the full-text query and the vector query together. The platforms expose this differently:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • OpenSearch uses a hybrid query over an index that contains both text and vector fields.
  • Azure AI Search accepts one hybrid request that carries full-text and vector query components.
  • Elastic documents a combined full-text and vector query as its hybrid pattern.

Two parameters need explicit values in every configuration: the candidate depth taken from each path before fusion, and the filters applied during retrieval. Candidate depth is a tuning parameter. If it is too shallow, the correct passage never reaches fusion. If it is too deep, you pay latency for candidates that never reach the model. Set it from your evaluation results rather than copying a value from another deployment.

4. Fuse the two result lists

Lexical and vector scores live on different scales, so adding them directly makes the final order depend on arbitrary magnitudes. Reciprocal rank fusion avoids this by using only each passage’s position in each list. The standard form scores a passage as the sum, across lists, of 1 divided by (k plus its rank), where k is a constant that damps the influence of top positions.

Consider two lists. The lexical path ranks passage A first and passage B second. The vector path ranks B first and passage C second. In this illustrative example, with k set to 60, A scores 1/61 (about 0.0164), B scores 1/62 + 1/61 (about 0.0325) and C scores 1/62 (about 0.0161). B ranks first because both paths returned it, even though only one of them placed it at the top.

RRF is the merge method Azure AI Search uses and the one Elastic recommends for hybrid search. OpenSearch offers it as a rank-based option alongside score-based normalization, which rescales each path’s scores to a common range and combines them with weights. Normalization fits when score margins carry meaning and you can keep the weights calibrated. Its cost is maintenance: normalized score distributions shift as the corpus changes, so the weights need retesting after reindexing. Use it only when tests show that rank position discards information you need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The value of k, or the weights in a normalized setup, is a tuning knob, and the reviewed sources do not give a default that suits every corpus. If you use OpenSearch, one more constraint applies. Its RRF documentation explains that shard-level BM25 statistics and the per-shard vector k can change which candidates each shard returns, and therefore change the fused ranks and scores. Run experiments with the same shard count as production, or the ordering you measure will not be the ordering users see.

5. Return a bounded, cited set of passages

Send the model a fixed number of the best passages, each with its source identity, location and an identifier or link back to the original document, so that answers can be verified. The number is a budget decision for your model and prompt, and the reviewed search documentation does not set it. The prompt, answer format and generation behaviour sit outside the search layer. Good retrieval improves the odds of a grounded answer but does not guarantee one, so the generation layer still needs a path for questions the corpus cannot answer.

Reciprocal rank fusion or a reranker?

They do different jobs, so they are not alternatives. RRF merges the lexical and vector lists into one candidate set. A reranker takes a candidate set that already exists and reorders it using a deeper calculation of query-document relevance. That adds processing and latency on every query where it runs. Microsoft’s RAG guidance advises comparing approaches against test queries and benchmarking relevance and latency before adopting a reranker in production.

Use this sequence to decide:

  1. Start with hybrid retrieval and RRF as the baseline. It avoids assumptions about score scales and gives you a fixed reference point.
  2. Run the evaluation set against the baseline and sort each failure into one of two groups: the correct passage is absent from the candidate list, or it is present but ranked too low.
  3. If the correct passage is missing, a reranker cannot help. Fix chunking, candidate depth, the analyzer or the embedding model first.
  4. If the correct passage is present but low, test a reranker over the same candidates. Compare hybrid alone with hybrid followed by the reranker, holding the queries, corpus and filters constant.
  5. Record the relevance gain and added latency for each query class. Enable the reranker globally only if the gain justifies the latency in your budget. If it helps only one query class, scope it to that class.

How do I evaluate RAG retrieval quality?

Evaluation needs a judged query set: representative questions, each paired with the documents or chunks that should answer it. Build it from real user questions where you can, and cover the classes below. The reviewed sources support comparing approaches on your own workload, but they do not establish a universal metric, threshold or top-k value. Set targets from your own baseline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Query class Example What it exposes
Exact names, acronyms, IDs, product codes and policy titles A policy number or a policy title typed as written Lexical coverage and analyzer behaviour
Natural-language questions and paraphrases “how do I get travel costs back” Vector recall and the contribution of the semantic path
Questions that require a section, date or document version The carry-over rule in the leave policy version in force last year Metadata fields and filter correctness
Questions the corpus cannot answer A question about a product the company does not sell Abstention in the generation layer, and how much irrelevant context retrieval passes along
Permission-sensitive questions The same question asked by users with different access levels Permission leakage and underfilled results

Compare lexical-only, vector-only, hybrid and hybrid with a reranker on the same queries, corpus and filters. Score retrieval separately from the generated answer. A wrong answer can come from missing evidence, which is a retrieval failure, or from the model ignoring good evidence, which is a generation failure. Only a retrieval measurement taken on its own can tell you which. Common ranking measures include recall at k, mean reciprocal rank and normalized discounted cumulative gain. Record latency and failure behaviour as well, including timeouts, empty result sets and errors under load.

Run the harness on production-equivalent infrastructure, including the shard layout discussed above, the same embedding model version and the same filter policy. A passing result on a small development index says little about the cluster users will query.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failure modes and how to diagnose them

Failure What it looks like First check
Query and index embedding mismatch Vector results look unrelated to the query while lexical results for the same query are reasonable Confirm the query path uses the same model and preprocessing as the chunks, and compare the stored embedding_model value with the query side
Exact-term misses A rare identifier returns semantically similar but wrong passages Confirm the lexical path returns the target document within its candidate depth, and check how the analyzer tokenizes the identifier
Overreliance on raw scores Fusion settings that worked on one corpus misorder results after reindexing Switch to rank-based fusion, or retest normalization after every corpus change
Shard-layout surprises (OpenSearch) Experiment rankings differ from production rankings Rerun the experiment with the production shard count
Unnecessary reranking Latency rises with little gain on the judged metrics Compare against hybrid-only on the same judged set, and remove the reranker where the gain is small
Permission leakage A user is shown a passage from a document they cannot open Confirm the permission filter is applied on both paths, that group resolution matches the source system, and that revocations reach the index within the agreed lag
Underfilled results Fewer passages reach the model than expected, especially for narrow user groups Count surviving candidates per user group, then review candidate depth and where filters apply
Stale or duplicated content A deleted or superseded passage still returns Run the delete and update acceptance tests, and check whether more than one chunk set exists for a single source_id
Vendor quickstart treated as production The demo index works, but nothing measures quality or failures Add the evaluation harness, monitoring, security controls, capacity planning and a named owner for the pipeline

Choosing a platform

The reviewed material covers three paths: OpenSearch, Azure AI Search, and Elastic (Elasticsearch). None of them is the right answer in general. The table compares what each platform’s documentation describes for hybrid retrieval. “Not stated” means the reviewed hybrid-search documentation does not describe that capability. It is not a claim that the platform lacks it.

Platform How hybrid is expressed Fusion methods documented Filters and reranking
OpenSearch A hybrid query over an index with text and vector fields Rank-based RRF and score-based normalization Not stated in the reviewed hybrid-search documentation
Azure AI Search One hybrid request carrying full-text and vector query components RRF merges the two result sets Filters are listed among the capabilities available in the hybrid query context. Reranking appears in the Azure Architecture Center RAG guidance, not in the hybrid overview
Elastic (Elasticsearch) A combined full-text and vector query RRF recommended for hybrid search; score-based option not stated Not stated in the reviewed hybrid-search documentation

The criteria that matter most usually sit outside the feature list:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Whether you want a managed service or a self-managed cluster, and what your team can operate.
  • Fit with your existing identity provider, infrastructure and data-source integrations.
  • Analyzer support for your languages and identifier formats, and the vector index options available.
  • How permission filters are implemented and how their application is logged.
  • Corpus size, update frequency, latency target and scaling approach.
  • Observability, cost model and deployment region.

The reviewed sources do not provide a consistent, current comparison of price, regional availability, service limits or feature tiers across these platforms. Confirm them with each provider before committing, because product documentation and offerings change.

The Bottom Line

Build hybrid retrieval as two paths over one index, use the same embedding model on both the chunk and query sides, and start with RRF as the fusion baseline. Treat score normalization, a reranker and tuned candidate depth as additions that your judged queries must justify. Keep the permission filter on both paths from the first day, because no amount of ranking tuning repairs a leak.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.