Preprint · 2026

When Does Query Expansion Help Agent Memory?

A Multi-Seed Study of Selective Expansion and Gate Reliability

Salomon Diei

School of Computer Science and Engineering, KOREATECH, Cheonan, Republic of Korea

An earlier version was presented at NAFSIK 2026

TL;DR Expanding only low-confidence queries halves the cost of LLM query expansion for agent memory with no loss in Recall@5, but the confidence gate picks queries no better than chance. Gating, not expansion, is the open problem.

SQE retrieval flow: dense retrieval first; queries whose top-1 score is below the threshold are expanded into trace probes and paraphrases, searched against the same index, and fused with reciprocal rank fusion.
Figure 1. SQE retrieval flow. Dense retrieval runs first; confident queries return the dense list unchanged. Queries whose top-1 score falls below τ are expanded into two hypothetical trace probes (searched with dense and BM25) and two paraphrases (searched with dense only), all against the same unchanged index. The original dense list and the five probe lists are fused with RRF.

Abstract

Long-horizon software agents store execution traces and later retrieve them with natural-language questions, so the query and the memory are written in different “languages”. LLM query expansion can bridge this gap, but it costs generation on every query. We study Selective Query-Side Expansion (SQE), which expands a query with hypothetical execution traces and paraphrases only when the top-1 dense retrieval score is low, then fuses all rankings with Reciprocal Rank Fusion.

Across 8 resampled memory stores with rebuilt indices (4,000 queries over SWE-smith agent steps), SQE reaches 69.4% Recall@5 while expanding 48% of queries; in a token-measured rerun it uses 55% fewer LLM tokens than always-on expansion. However, its gain over plain dense retrieval is small (+0.9 points, 95% CI [+0.1, +1.7]), Recall@1 drops, and it is statistically indistinguishable from a random gate at the same budget. An oracle analysis explains why: expansion helps only 5.3% of queries and hurts 4.2%, so even a perfect gate gains at most 5.3 points. We release all memory stores, per-method summaries, and verification scripts, and argue that confidence gating, not expansion itself, is the open problem.

Contributions

  1. A formulation of the trace–query mismatch in long-horizon software-agent memory, and a selective, query-side-only expansion method that leaves the memory store and indices untouched.
  2. A controlled multi-seed study with an executed random-gating baseline at matched expansion rate and provider-measured token costs.
  3. Evidence that the top-1-score gate is not better than random, with an oracle-gate analysis that bounds what any gate could achieve in this setting.
  4. A release of seeds, summaries, paired tests, and verification scripts so that every number in the paper can be recomputed.

Results

Retrieval over 8 resampled memory stores, 500 queries each.

69.4%SQE Recall@5
(Dense-Only 68.5%)
48%of queries expanded
−55%LLM tokens vs. Always-Expand
+0.5points vs. a random gate
(not significant)

Finding 1

Slightly better than dense retrieval at Recall@5, worse at Recall@1

The paired difference is +0.9 points (95% CI [+0.1, +1.7]). SQE recovers 140 targets dense retrieval missed but loses 105 it found. Recall@1 falls from 45.5% to 44.5%, and seed-to-seed spread (±2.7 points) is three times the method effect.

Finding 2

The gate does not beat a random gate

Against Random-Gated-Expansion at a similar budget, SQE gains +0.5 points (CI [−0.2, +1.1]); against Always-Expand, +0.1. Expanding about half the queries recovers essentially all of the benefit, but which half matters little.

Finding 3

Selectivity is still a large cost saving

SQE uses 405 tokens and 2.8 s per query versus 899 tokens and 5.4 s for Always-Expand, at equal Recall@5. Trace-only expansion is the worst operating point.

Explore the released results

Method comparison (8-seed mean)

Per-seed: Dense-Only vs. SQE (Recall@5)

Loaded live from results/multiseed/multiseed_report.json; no numbers are hard-coded.

Paired bootstrap: SQE minus each baseline

BaselineΔRecall@595% CIp
Dense-Only+0.9[+0.1, +1.7]0.015
Hybrid-RRF+5.9[+4.8, +7.0]<0.001
Always-Expand+0.1[−0.4, +0.7]0.315
Random-Gated-Expansion+0.5[−0.2, +1.1]0.085
Paraphrases-Only+1.1[+0.3, +1.9]0.005
HyDE-Traces-Only+5.6[+4.7, +6.6]<0.001

Pooled over 4,000 query rows. Seeds share about 30% of their memories, so these intervals are descriptive and likely too narrow.

Recall@5 per seed and 8-seed mean for each method
Figure 2. Recall@5 per seed (dots) and 8-seed mean (bar). The dashed line is the Dense-Only mean. Seed variance dominates every expansion effect.
Measured LLM tokens per query against 8-seed mean Recall@5
Figure 3. Cost against quality: measured LLM tokens per query (seed 42) against 8-seed mean Recall@5. SQE matches Always-Expand at roughly half the tokens, but so does the random gate.

Why the gate fails

The top-1 score finds hard queries, but not the ones expansion can fix.

61.7% vs. 75.9%

Low confidence marks hard queries, not fixable ones

On seed 42, expanded queries reach 61.7% Recall@5 and non-expanded ones 75.9%. Expansion does not reliably repair the hard ones.

2.2 pts

The threshold barely matters

Recall@5 stays within a 2.2-point band while the expansion rate goes from 0% to 100%. Leave-one-seed-out threshold selection gives +0.6 points over dense (CI [−0.1, +1.3]); BM25-agreement variants do not close the gap.

5.3% help · 4.2% hurt

There is little headroom for any gate

An oracle that expands exactly the queries expansion helps reaches 73.5% Recall@5 against 68.2% for dense. The remaining 90% of queries have the same outcome either way.

Recall@5 and expansion rate across gate thresholds on seed 42
Figure 4. Post-hoc threshold sweep on seed 42. Recall@5 (line) varies by about 2 points while the fraction of expanded queries (shaded) spans 0–100%.
Held-out oracle headroom compared with dense retrieval and always expanding
Figure 5. Held-out oracle headroom. Even a perfect per-query gate gains about 5 points over dense retrieval; always expanding gains about 1.

Method

SQE changes only the query at retrieval time. The memory store and both indices stay unchanged.

  1. 1RetrieveRun dense retrieval (bge-m3, FAISS) over the trace store and read the top-1 score.
  2. 2GateIf the top-1 score is at least τ = 0.65, return the dense list unchanged.
  3. 3ExpandOtherwise generate K = 2 hypothetical traces (dense + BM25) and P = 2 paraphrases (dense).
  4. 4FuseCombine the original list and the five probe lists with reciprocal rank fusion (k = 60).

Setup. Memory stores of 5,000 single agent steps from the SWE-smith tool split. For each seed (42–49) the store is resampled, both indices are rebuilt, and 500 targets are sampled. Expansion uses Qwen3.6-35B-A3B served with vLLM. Generated traces are retrieval probes, not claims that those executions occurred.

Limitations

Reproducibility

scripts/07_verify_experiment.py recomputes every metric from the result files and checks each table against its declared source.

BibTeX

@misc{diei2026sqe,
  title  = {When Does Query Expansion Help Agent Memory? A Multi-Seed Study of Selective Expansion and Gate Reliability},
  author = {Diei, Salomon},
  year   = {2026},
  note   = {Preprint},
  url    = {https://github.com/Salomondiei08/sqe-experiment}
}