Edison FloresHow extending one bug repro to all four fusion modes of QueryFusionRetriever exposed a mutation class that corrupts retriever caches — and the test pattern that catches it anywhere.
On 2026-10-03, michaelswissa filed a bug in llama_index: QueryFusionRetriever corrupts rankings when a caching retriever returns shared NodeWithScore wrappers. Relative and distance-based fusion normalize scores in place, so a wrapper shared across queries gets normalized for one result set and then read back as if it were another's.
His patch (#23333) fixes two of the four fusion modes. While reviewing it, I extended his reproduction script to all four modes — and the same mutation class was live in the other two:
reciprocal_rerank writes fused scores back into the retriever's own wrapperssimple mutates the first-seen wrapper for a node hash (the subtle one)Filed as #23351; the simple half is fixed in #23352. This post is the full anatomy of the bug class, because it generalizes far past this retriever.
Every fusion mode receives Dict[Tuple[str, int], List[NodeWithScore]] — query results the retriever produced. Any write into those wrappers is a write into someone else's state. If the retriever caches (per node, per query, whatever), your "harmless normalization" silently rewrites its cache.
The original issue caught the loud variant:
# _relative_score_fusion — normalizes in place
node_with_score.score = (node_with_score.score - min_score) / (max_score - min_score)
A shared wrapper normalized for result set A is then summed again for result set B. The relevant document ends up at -3.0 and ranks last. Loud, reproducible, fixed by #23333.
RRF computes ranks first, scores second, then writes back:
reranked_nodes.append(hash_to_node[hash])
reranked_nodes[-1].score = score # writes into the retriever's wrapper
The current call's ranking is correct — ranks were computed before the write-back. The corruption lands after: the cached 0.9 / 0.1 / 0.8 become RRF fused scores (~0.0333 / ~0.0164), and every subsequent retrieval from that cache returns garbage. Same class, but the symptom is invisible unless you assert on the cache itself:
mode=reciprocal_rerank ranking=['relevant', 'low', 'other'] # correct
cache after: orig=[0.0333, 0.0164] alt=[0.0333, 0.0164] # corrupted
simple looked clean
_simple_fusion dedups by node hash and writes the max back:
max_score = max(node_with_score.score or 0.0, all_nodes[hash].score or 0.0)
all_nodes[hash].score = max_score # writes into the retriever's wrapper
Run the shared-wrapper repro from #23332 against it and nothing happens — max(0.9, 0.9) is identity. The obvious test passes, the mode looks clean, and the bug ships.
The trigger requires the same node hash to arrive via distinct wrappers with different scores — exactly what a per-query cache produces when scoring depends on the query (BM25, query-weighted retrievers):
class PerQueryCacheRetriever(BaseRetriever):
def __init__(self):
super().__init__()
self.node = TextNode(text="relevant", id_="relevant")
self.results = {
"original": [NodeWithScore(node=self.node, score=0.4), ...],
"alternate": [NodeWithScore(node=self.node, score=0.9), ...],
}
Fusion runs; the dedup max-write puts 0.9 into the wrapper owned by "original"'s cached results. That query's cache is silently rewritten with another query's score. No exception, no ranking change in the current call — just corrupted state for every future retrieval.
The lesson that generalizes: a repro built from one cache shape (shared wrappers) is a single point of failure. When the mechanism is "write into inputs," enumerate the aliasing shapes — same object in many lists, distinct objects with one hash — and assert the post-state of every shape, not just the output.
Three steps, ~30 minutes per suspect function:
x.score =, x[...] = ... where x came from a parameter. In fusion_retriever.py this is literally four functions, two of which write into their arguments.base.results[...] after fusion — the outputs were never wrong; the caches were.result = fusion.retrieve("original")
result[0].score = 999.0
assert base.results["original"][0].score == 0.4 # ownership proven
Two shapes, both correct, different blast radius:
simple: keep (node, max_score) tuples in the dedup dict, build new NodeWithScore objects at the end. Smallest surface, merges alongside in-flight fixes without touching their lines.One wrinkle worth knowing before you patch a big repo: PR #21445 has been sitting on the RRF write-back since April (blocked, stalled). Any fix that touches those ten lines inherits its rebasing duty. "Is this hunk already claimed?" is now part of my checklist before writing a patch.
simple (michael's review catch: don't trade a false positive for a false negative, keep the fix scoped)simple-mode fix: 4 regression tests, 3 fail on main without it, full retrievers suite greenVerification was done live against llama-index-core 0.14.25 before filing anything, and every claim in this post comes from that run — including the numbers.
Disclosure: the investigation and repro work were AI-assisted (Claude/GLM), human-directed; the same disclosure is on the linked issues and PRs.