One bug class, four fusion modes: finding cache corruption in QueryFusionRetriever

# python# llamaindex# debugging# opensource
One bug class, four fusion modes: finding cache corruption in QueryFusionRetrieverEdison Flores

How extending one bug repro to all four fusion modes of QueryFusionRetriever exposed a mutation class that corrupts retriever caches — and the test pattern that catches it anywhere.

What happened

On 2026-10-03, michaelswissa filed a bug in llama_index: QueryFusionRetriever corrupts rankings when a caching retriever returns shared NodeWithScore wrappers. Relative and distance-based fusion normalize scores in place, so a wrapper shared across queries gets normalized for one result set and then read back as if it were another's.

His patch (#23333) fixes two of the four fusion modes. While reviewing it, I extended his reproduction script to all four modes — and the same mutation class was live in the other two:

  • reciprocal_rerank writes fused scores back into the retriever's own wrappers
  • simple mutates the first-seen wrapper for a node hash (the subtle one)

Filed as #23351; the simple half is fixed in #23352. This post is the full anatomy of the bug class, because it generalizes far past this retriever.

The class: fusion that mutates what it doesn't own

Every fusion mode receives Dict[Tuple[str, int], List[NodeWithScore]] — query results the retriever produced. Any write into those wrappers is a write into someone else's state. If the retriever caches (per node, per query, whatever), your "harmless normalization" silently rewrites its cache.

The original issue caught the loud variant:

# _relative_score_fusion — normalizes in place
node_with_score.score = (node_with_score.score - min_score) / (max_score - min_score)
Enter fullscreen mode Exit fullscreen mode

A shared wrapper normalized for result set A is then summed again for result set B. The relevant document ends up at -3.0 and ranks last. Loud, reproducible, fixed by #23333.

The quiet variant: reciprocal_rerank

RRF computes ranks first, scores second, then writes back:

reranked_nodes.append(hash_to_node[hash])
reranked_nodes[-1].score = score   # writes into the retriever's wrapper
Enter fullscreen mode Exit fullscreen mode

The current call's ranking is correct — ranks were computed before the write-back. The corruption lands after: the cached 0.9 / 0.1 / 0.8 become RRF fused scores (~0.0333 / ~0.0164), and every subsequent retrieval from that cache returns garbage. Same class, but the symptom is invisible unless you assert on the cache itself:

mode=reciprocal_rerank  ranking=['relevant', 'low', 'other']   # correct
    cache after: orig=[0.0333, 0.0164] alt=[0.0333, 0.0164]    # corrupted
Enter fullscreen mode Exit fullscreen mode

The trap: why simple looked clean

_simple_fusion dedups by node hash and writes the max back:

max_score = max(node_with_score.score or 0.0, all_nodes[hash].score or 0.0)
all_nodes[hash].score = max_score   # writes into the retriever's wrapper
Enter fullscreen mode Exit fullscreen mode

Run the shared-wrapper repro from #23332 against it and nothing happens — max(0.9, 0.9) is identity. The obvious test passes, the mode looks clean, and the bug ships.

The trigger requires the same node hash to arrive via distinct wrappers with different scores — exactly what a per-query cache produces when scoring depends on the query (BM25, query-weighted retrievers):

class PerQueryCacheRetriever(BaseRetriever):
    def __init__(self):
        super().__init__()
        self.node = TextNode(text="relevant", id_="relevant")
        self.results = {
            "original":   [NodeWithScore(node=self.node, score=0.4), ...],
            "alternate": [NodeWithScore(node=self.node, score=0.9), ...],
        }
Enter fullscreen mode Exit fullscreen mode

Fusion runs; the dedup max-write puts 0.9 into the wrapper owned by "original"'s cached results. That query's cache is silently rewritten with another query's score. No exception, no ranking change in the current call — just corrupted state for every future retrieval.

The lesson that generalizes: a repro built from one cache shape (shared wrappers) is a single point of failure. When the mechanism is "write into inputs," enumerate the aliasing shapes — same object in many lists, distinct objects with one hash — and assert the post-state of every shape, not just the output.

How to find this class anywhere

Three steps, ~30 minutes per suspect function:

  1. Grep for writes into inputs. x.score =, x[...] = ... where x came from a parameter. In fusion_retriever.py this is literally four functions, two of which write into their arguments.
  2. Assert on the caller's state, not the return value. Every regression test in #23352 checks base.results[...] after fusion — the outputs were never wrong; the caches were.
  3. Test ownership, not just values. The strongest assertion is aliasing: mutate the returned wrapper, assert the input didn't move.
result = fusion.retrieve("original")
result[0].score = 999.0
assert base.results["original"][0].score == 0.4   # ownership proven
Enter fullscreen mode Exit fullscreen mode

The fix family

Two shapes, both correct, different blast radius:

  • Per-occurrence copies at the dispatch sites — one invariant ("fusion works on copies") covering all modes. Copy each wrapper per occurrence per list, never deduped by identity: the same object appearing in two lists must become two copies, or relative normalization computes the wrong min/max per list. (An identity-keyed copy cache is the wrong-shaped optimization for this exact input.)
  • Fresh output wrappers per function — what #23352 does for simple: keep (node, max_score) tuples in the dedup dict, build new NodeWithScore objects at the end. Smallest surface, merges alongside in-flight fixes without touching their lines.

One wrinkle worth knowing before you patch a big repo: PR #21445 has been sitting on the RRF write-back since April (blocked, stalled). Any fix that touches those ten lines inherits its rebasing duty. "Is this hunk already claimed?" is now part of my checklist before writing a patch.

What landed

  • #23351 — the follow-up issue michael asked for, with the distinct-wrapper repro for simple (michael's review catch: don't trade a false positive for a false negative, keep the fix scoped)
  • #23352 — the simple-mode fix: 4 regression tests, 3 fail on main without it, full retrievers suite green

Verification was done live against llama-index-core 0.14.25 before filing anything, and every claim in this post comes from that run — including the numbers.

Disclosure: the investigation and repro work were AI-assisted (Claude/GLM), human-directed; the same disclosure is on the linked issues and PRs.