Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A QueryFusionRetriever result can be ranked correctly while the retriever’s cached scores have already been overwritten. The reported defect is not simply a bad ranking: fusion can mutate a mutable NodeWithScore wrapper owned by an input result list, so a later retrieval may expose altered scores. GitHub issue #23351 reports this for reciprocal_rerank and simple; the exact behavior depends on the code branch and wrapper objects involved.

Why can a correct result still leave the cache corrupted?

Fusion receives lists of NodeWithScore objects and uses node hashes to identify duplicate nodes across query results. Those are two different kinds of identity: equal node hashes mean the results refer to the same logical node, but they do not mean the lists contain the same wrapper object. A wrapper also carries a mutable score.

If fusion changes the score on an input wrapper, any cache that retains that same object sees the change too. A second risk exists even when the wrappers are distinct: if a fusion routine chooses one wrapper for a node hash and writes a combined score into it, the cache that owns that particular wrapper is changed. Correctly calculating and returning a fused ranking therefore does not establish that the input results remained intact.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Shared-wrapper aliasing: the same NodeWithScore instance appears in multiple result lists. Writing its score through one list changes what the other list sees.
  • Distinct-wrapper, same-hash aliasing: separate wrapper instances represent the same node hash. A routine can still overwrite one list’s cached score while deduplicating the node.

What does each fusion mode combine, and where is mutation reported?

The source snapshot examined for this report exposes four modes. The table describes the implementation behavior visible in that snapshot and issue/PR reports; it is not a claim that every released version behaves identically.

Mode What it combines Deduplication and score handling Mutation evidence and qualification
reciprocal_rerank Rank contributions; the implementation uses k = 60.0. Uses node hashes to retain a wrapper, calculates fused scores, orders hashes by those scores, then writes the fused score to the retained wrapper. Issue #23351 reports that a shared input wrapper can be overwritten after ranks are calculated. The current returned ranking may look right even as a cache-held wrapper acquires the fused score.
relative_score Normalized scores, adjusted by retriever weights and query count. Calculates per-result-set minimum and maximum values, normalizes and scales scores, then sums scores for duplicate hashes. The issue/PR history identifies this path as part of the earlier fix scope in PR #23333. However, the main-branch source snapshot described here still showed in-place assignments. Check the target branch and version rather than assuming a release includes the fix.
dist_based_score Scores normalized using distance-based bounds derived from mean and standard deviation. Calls the relative-score routine with those bounds and follows its in-place scaling and deduplication path in the examined source snapshot. Issue #23351 describes distance-based fusion as part of the earlier fix scope. The observed source snapshot still showed the mutation path; verify the specific code and release you use.
simple The incoming scores for each node. Deduplicates by hash and writes the maximum score into the first-seen wrapper. Issue #23351 reports a cache overwrite when distinct wrappers for the same hash have different query-dependent scores. A shared-wrapper case with equal scores can appear harmless because the maximum equals the score already present.

The issue says synchronous and asynchronous retrieval both call the same fusion functions, so the ownership concern applies to either entry point. That does not mean every code version or every cache arrangement will reproduce it.

How do the reported reproductions expose the two aliasing cases?

Shared wrapper: reciprocal-rank fusion

Issue #23351 reports testing with llama-index-core 0.14.25, Python 3.12, on Linux x86_64. In its reciprocal-rank example, the same wrapper is present in multiple query result lists. The original cached scores are reported as 0.9 and 0.1; after fusion, the issue reports values of approximately 0.0333 and 0.0164 in the original query’s cache as well as in the fused results.

The important diagnostic is the order of observations: the fusion output can contain the expected fused values, but those values are also visible in the cache afterward because the wrapper was shared. The issue’s reported figures are a reproduction, not an independently run test here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Distinct wrappers: simple fusion

The simple-mode reproduction uses separate wrappers for the same node hash, with query-dependent scores of 0.4 and 0.9. Because simple fusion keeps the maximum, the implementation writes 0.9 into the first-seen wrapper. Issue #23351 reports that this changes the first query’s cache, which it expects to remain at 0.4 (along with another score of 0.1 in that cache).

This case is easy to miss if a test only reuses the identical wrapper with identical scores: writing the maximum then makes no observable difference. Distinct wrappers with equal hashes and different scores are needed to expose this particular overwrite.

What should a regression test check?

Test both the returned fusion result and the state of every input list after fusion. Vary object identity separately from node identity; a test that covers only one cache shape cannot establish safety for the other.

  • Wrapper arrangement: reuse the exact same wrapper across result lists, then repeat with distinct wrappers carrying the same node hash.
  • Score relationship: test equal scores and different, query-dependent scores.
  • Retrieval path: cover synchronous and asynchronous entry points, since the issue reports that both dispatch to the fusion functions.
  • Returned values: assert the expected fused ordering and scores.
  • Input post-state: assert that every original/cache score is unchanged after fusion, including the first-seen wrapper and wrappers in other query lists.
  • Ownership: where the API is intended to produce fresh output wrappers, mutate a returned wrapper in the test and confirm that retriever-owned wrappers do not change.

For the issue’s examples, the cache assertions are as important as the ranking assertion: reciprocal-rank fusion should not replace the original 0.9 / 0.1 cache values with the reported fused values, and simple fusion should not replace the first query’s 0.4 with the other query’s 0.9.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What fix is proposed, and what is established about release status?

PR #23352 proposes avoiding simple-mode mutation by keeping node data alongside the maximum score, then constructing fresh NodeWithScore wrappers for output. Its description lists tests for distinct per-query wrappers, shared wrappers, output non-aliasing, and asynchronous behavior; it reports that three of the four new tests fail on main without the change. Those are claims in the PR description, not independently verified test results here.

The PR treats reciprocal-rank fusion separately, saying PR #21445 already rebuilds fresh wrappers as a side effect of adding retriever weights, while describing that work as blocked or stalled. Meanwhile, issue #23351 remains open in the GitHub snapshot underlying these reports. The evidence does not establish that all proposed changes have merged or shipped. Check the target repository branch and package version before relying on a fix, especially because the examined main-branch snapshot still showed in-place writes in simple, reciprocal-rank, and relative-score processing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.