The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Build hybrid search by combining lexical (keyword or full-text) retrieval with vector retrieval, then evaluate every candidate configuration against the same versioned queries and relevance judgments. Keep the corpus, index, embedding model, and search settings identifiable too: a fixed query list alone cannot make a comparison reproducible if the documents or labels change. Choose the setup that meets your measured relevance and operational needs; no fusion method or weighting is best for every workload.
What hybrid search combines
Hybrid search retrieves candidates through both term-based matching and vector or semantic matching, then merges them into one ranked result list. Lexical retrieval can match exact terms and identifiers; vector retrieval can find semantically related content even when wording differs. Their value depends on the corpus and the queries users actually submit.
OpenSearch defines hybrid search as combining keyword and semantic search. Elastic describes its implementation as running full-text and vector search in one request, while Azure AI Search also combines text and vector results. These are vendor-specific implementations of the broader pattern; APIs, defaults, permissions, and capabilities differ. OpenSearch hybrid search, Elastic hybrid search, Azure AI Search hybrid search
Free tools Windows power users keep installed
One-click scans. No signup required.
How to build a repeatable evaluation
1. Define the search task and freeze representative queries
Assemble the exact query strings that reflect the application’s users and failure modes. Include exact terms, natural-language intent, rare identifiers, ambiguous requests, and known cases where the current system performs poorly. Save the literal strings with a query-set version; do not silently edit or replace queries between runs. OpenSearch’s Search Relevance Workbench supports manually defined query sets and uses “tv” and “led tv” as example strings. OpenSearch Search Relevance Workbench query sets
#1 Best Overall
2. Create and version relevance judgments
For each query, rate how relevant each judged document is to that query. OpenSearch defines a judgment as a relevance rating for one document-query pair, and groups ratings into judgment lists. Keep judgments tied to a version of the test collection: if documents or labels change, record that change rather than treating the new run as directly comparable to the old one. A stable query set without stable document judgments does not provide a controlled relevance comparison. OpenSearch judgments and judgment lists
3. Implement both retrieval paths and a fusion stage
In OpenSearch, the documented manual build flow is to create an embedding ingest pipeline, create an index with correctly typed text and vector fields whose dimensions match the embedding model, configure a search pipeline, ingest documents, and query with hybrid retrieval. An automated workflow can provision the ingest pipeline, index, and search pipeline when supplied with a model ID and suitable vector dimension. OpenSearch hybrid search setup
Record the corpus or index version, embedding model, and search configuration used for each run. This is a practical reproducibility record inferred from the need to control queries, judgments, and configurations; it is not a formal universal standard published by the cited documentation.
Rank #2
Choose a fusion approach to test, not assume
Fusion determines how lexical and vector results contribute to the final ordering. OpenSearch documents two broad options: normalize clause scores to a common scale and combine them, or use reciprocal rank fusion (RRF), which combines document positions without using raw score values.
| Approach | How it combines results | Useful evaluation question |
|---|---|---|
| Score normalization and combination | Normalizes clause scores and combines them, retaining score margins. | Do score differences meaningfully distinguish strong from weak matches in this corpus? |
| RRF | Combines rankings by document position and ignores raw score values. | Does rank-based merging improve judged relevance without unacceptable latency or other operational costs? |
Neither method is established as a universal winner. Score ranges and distributions can affect normalization; RRF avoids relying on comparable raw scales but discards the magnitude of those scores. Compare both where your platform supports them and decide from the same frozen evaluation set. OpenSearch fusion methods
Compare configurations on identical evidence
Run each candidate configuration against the identical query strings, judgment list, and test collection. OpenSearch’s Search Relevance Workbench supports experiments that compare two search configurations, evaluate a configuration against judgments, or optimize hybrid parameters. Its optimization process evaluates combinations of variants across the query set and scores results against judgments. Search Relevance Workbench experiments OpenSearch hybrid parameter optimization
Rank #3
The documented OpenSearch experiment space includes:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute- Normalization methods:
l2,min_max, andz_score; in the documented setup,z_scoreis limited toarithmetic_mean. - Score-combination methods:
arithmetic_mean,harmonic_mean, andgeometric_mean. - Lexical and neural weights in increments of 0.1 from 0.0 to 1.0.
- RRF rank constants of 1, 5, 10, 20, and 60; the documented RRF variants use equal weights among subqueries.
These are configuration values available in the documented experiment space, not benchmark results or evidence of an expected improvement. No generalizable hybrid-search uplift is established by the cited sources. OpenSearch experimental parameters
Evaluate relevance and operational behavior together
Look beyond an aggregate score
Use the same relevance metric and judgments for each candidate, but inspect outcomes by query category as well as in aggregate. A configuration can improve one kind of query while worsening another; exact identifiers, broad natural-language requests, and ambiguous terms may behave differently. The right choice depends on which misses matter for the application.
Rank #4
Measure latency, filtering, and throttling under representative load
Relevance is only part of a usable configuration. Record latency under representative load, whether filters behave as expected, and whether the system throttles. Azure’s guidance notes that large candidate sets, expensive vector settings, and semantic reranking can increase merge cost, latency, and throttling pressure. Its guidance suggests starting with a balanced hybrid pattern, tuning in small steps, and enabling semantic ranking only when it measurably improves relevance. Azure hybrid query guidance
Where recall is the priority, test a broader candidate set and measure the additional cost. Where precision or response time is the constraint, test more selective retrieval or reranking settings. These are trade-offs to measure on the target workload, not fixed prescriptions.
Check the result presentation
Return human-readable fields for users rather than exposing vector values as if they were meaningful text. Also interpret scores according to the fusion method: Azure notes that RRF scores have different magnitudes from pure vector similarity scores, so a low-looking RRF score is not directly comparable to a cosine-similarity value. Azure result and score guidance
Best Value
Make the winning configuration reproducible
Keep a run record that lets another engineer recreate the comparison. At minimum, associate each result with the query-set version, judgment version, corpus or index version, embedding model, and search configuration. Record relevant operational conditions, including filters and the load used for latency measurements. If any of these change, treat the result as a new experiment rather than attributing the difference solely to a fusion setting.
The final selection should be the configuration that performs acceptably on judged relevance across the query categories that matter and remains within the application’s operational constraints. The documentation describes available techniques and experiment controls, but does not establish a universally optimal weight, fusion method, or expected gain for every corpus.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

