Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universal winner. Start with reciprocal rank fusion (RRF) when your lexical and vector scores are hard to compare or you have few relevance labels. Test weighted score fusion when score gaps are meaningful and you can normalize and tune them. Consider learned fusion when you have representative judgments and can maintain a training and evaluation loop. Which method performs best depends on the retrievers, corpus, query mix, and the ranking cutoff that matters to your application.

What hybrid retrieval fusion does

A hybrid search system may retrieve candidates with more than one method—for example, keyword search such as BM25 and dense-vector search. Fusion combines their result lists into one ranking. It can change the order of documents that the retrievers found, but it cannot recover a relevant document that none of them retrieved.

The key difference among the three approaches is what each method uses as evidence: a document’s position in each list, its component scores, or a rule learned from relevance data.

Method What it combines Main advantage Main dependency
RRF Rank positions Does not require component scores to share a scale Ranked lists, candidate depth, and RRF parameters
Weighted score fusion Normalized component scores and weights Can preserve score-margin information Normalization, weights, and validation data
Learned fusion Scores or other ranking features through a fitted rule Can learn a more tailored or query-dependent combination Representative judgments and a maintained evaluation loop

How the three methods work

Reciprocal rank fusion (RRF)

A common RRF formula is score(d) = Σ 1 / (k + rank(d)), summing a document’s contribution across the ranked lists in which it appears. Here, rank(d) is its position in a list and k controls how quickly contributions diminish at lower positions. A document absent from a list contributes nothing from that list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

RRF uses positions, not original score magnitudes. If a document barely beats another on BM25 or wins by a large score margin, that difference is discarded as soon as the documents become ranks. This makes RRF useful when one retriever emits unbounded BM25 scores and another emits bounded vector similarities that are not calibrated to each other. It also means the candidate depth supplied by each retriever remains important: fusion cannot favor candidates it never receives.

OpenSearch documentation recommends RRF as a reasonable starting point before score distributions have been measured or calibrated. RRF is the fusion stage, not a semantic reranker: Microsoft Azure AI Search documentation describes semantic ranking as a separate operation that can rescore candidates after RRF.

Weighted score fusion

Weighted score fusion combines component scores, commonly after normalization, using a weighted sum or convex combination. Unlike RRF, it can preserve score margins: a strong score advantage may matter more than a small one. OpenSearch documents score-based normalization processors and weighted combinations as alternatives to rank-based fusion.

Normalization alone does not establish the right blend. You still need to choose a normalization method and weights, then check whether they improve ranking on representative queries. This approach is worth testing when score gaps contain useful signal and you have a stable way to put component scores on compatible scales.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Learned fusion

“Learned fusion” describes a family of approaches, not one algorithm. A system might fit weights from query-document relevance judgments, feed component scores to a learning-to-rank model, or learn a query-dependent rule. The richer versions can express patterns that one global weighted blend cannot, but they also need representative training data, held-out evaluation, and ongoing maintenance.

Consider it when your query mix may need different treatment across query types and you have enough judgments to train and validate a model. Compare it with both a tuned global weighted blend and an RRF baseline; the available evidence does not show that learned fusion universally beats either.

When to test each method

Your situation Start by testing Why
Score scales are incompatible, judgments are scarce, or you need a low-tuning baseline RRF It combines ranks rather than requiring comparable score values.
Score margins appear informative and you can normalize scores and validate weights Weighted score fusion It retains score information that rank-only fusion discards.
You have representative labels and can maintain training and evaluation Learned fusion You can fit weights or a richer ranking rule to observed relevance.
You do not know which retrieval signal helps which queries Compare all three on held-out query slices The best ranking depends on the target corpus and query mix.

These are starting points for an experiment, not algorithmic laws. A weak retriever or shallow candidate list can limit every fusion method, regardless of how well its combination rule is tuned.

What published comparisons do—and do not—show

Results differ across evaluation settings. In a comparison cited by OpenSearch documentation, RRF had an average NDCG@10 that was 3.86% lower than the score-based hybrid pipeline across six BEIR datasets; OpenSearch reports comparable latency and coordinator-node CPU utilization in that comparison. This is a result for that cited benchmark, not a forecast for another corpus or implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bruch, Gai, and Ingber’s 2022 study found that convex combination outperformed RRF in its in-domain and out-of-domain experiments. The authors also reported that, for the datasets in their study, the convex-combination parameter converged with less than 5% of the training data. Neither result establishes a sample requirement or winner for every production system. The paper’s comparison also found RRF sensitive to its parameters.

MTEB documentation provides another illustration of how results can vary by task. Its documented hybrid examples use equal weights; NDCG@10 was:

Task BM25 Dense RRF DBSF RSF
NanoSciFactRetrieval 0.710 0.725 0.754 0.538 0.767
NanoNFCorpusRetrieval 0.325 0.288 0.329 0.338 0.359
NanoSCIDOCSRetrieval 0.335 0.344 0.369 0.344 0.372

These task-level examples are not a general ranking of methods. The original 2009 RRF paper by Gordon V. Cormack, Charles L. A. Clarke, and Stefan Buettcher reported that RRF performed consistently better than individual systems and standard Condorcet Fuse in its experiments; that is not a head-to-head verdict against modern weighted or learned hybrid fusion.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare fusion methods on your system

  1. Hold retrieval inputs constant. Use the same lexical and dense retrievers, corpus, candidate depths, and judged queries for every fusion method. Otherwise, a changed retrieval stage can be mistaken for a fusion improvement.
  2. Separate tuning from evaluation. Split representative judged queries into tuning and held-out sets. Choose weights or train a learned model using only the tuning portion; use the held-out portion to compare candidates.
  3. Choose a metric and cutoff that match the product. Measure ranking quality at the depth where results are actually used. NDCG@10 is one example, not a mandatory choice; MTEB’s task-level examples show why a method’s result should be read in its evaluation context.
  4. Inspect query slices. Report results separately for intents and query forms such as exact names or identifiers, short keyword searches, and longer natural-language requests. An aggregate can hide a method that helps one kind of query while hurting another.
  5. Measure operational costs as well as relevance. Track serving latency, compute cost, score stability, and how often weights or a model need recalibration or retraining. OpenSearch’s cited BEIR comparison found comparable latency and coordinator CPU for its tested pipelines; other workloads and implementations need their own measurements.
  6. Repeat after meaningful changes. Re-evaluate when the corpus, query mix, or component retrievers change. Do not copy a published weight or parameter without validating it on your own data.

How to make the final choice

Use RRF as a defensible baseline when scores are not comparable or labels are limited. Promote weighted fusion if a normalized, validated blend makes better use of score margins. Move to learned fusion when representative judgments support a reliable held-out gain that justifies the added training and maintenance. Select by measured relevance at the useful cutoff, together with serving and stability requirements—not by the method’s name.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.