TL;DR Expanding only low-confidence queries halves the cost of LLM query expansion for agent memory with no loss in Recall@5, but the confidence gate picks queries no better than chance. Gating, not expansion, is the open problem.
Figure 1. SQE retrieval flow. Dense retrieval runs first; confident queries return the dense list unchanged. Queries whose top-1 score falls below τ are expanded into two hypothetical trace probes (searched with dense and BM25) and two paraphrases (searched with dense only), all against the same unchanged index. The original dense list and the five probe lists are fused with RRF.
Abstract
Long-horizon software agents store execution traces and later retrieve them with natural-language questions, so the query and the memory are written in different “languages”. LLM query expansion can bridge this gap, but it costs generation on every query. We study Selective Query-Side Expansion (SQE), which expands a query with hypothetical execution traces and paraphrases only when the top-1 dense retrieval score is low, then fuses all rankings with Reciprocal Rank Fusion.
Across 8 resampled memory stores with rebuilt indices (4,000 queries over SWE-smith agent steps), SQE reaches 69.4% Recall@5 while expanding 48% of queries; in a token-measured rerun it uses 55% fewer LLM tokens than always-on expansion. However, its gain over plain dense retrieval is small (+0.9 points, 95% CI [+0.1, +1.7]), Recall@1 drops, and it is statistically indistinguishable from a random gate at the same budget. An oracle analysis explains why: expansion helps only 5.3% of queries and hurts 4.2%, so even a perfect gate gains at most 5.3 points. We release all memory stores, per-method summaries, and verification scripts, and argue that confidence gating, not expansion itself, is the open problem.
Contributions
A formulation of the trace–query mismatch in long-horizon software-agent memory, and a selective, query-side-only expansion method that leaves the memory store and indices untouched.
A controlled multi-seed study with an executed random-gating baseline at matched expansion rate and provider-measured token costs.
Evidence that the top-1-score gate is not better than random, with an oracle-gate analysis that bounds what any gate could achieve in this setting.
A release of seeds, summaries, paired tests, and verification scripts so that every number in the paper can be recomputed.
Results
Retrieval over 8 resampled memory stores, 500 queries each.
69.4%SQE Recall@5 (Dense-Only 68.5%)
48%of queries expanded
−55%LLM tokens vs. Always-Expand
+0.5points vs. a random gate (not significant)
Finding 1
Slightly better than dense retrieval at Recall@5, worse at Recall@1
The paired difference is +0.9 points (95% CI [+0.1, +1.7]). SQE recovers 140 targets dense retrieval missed but loses 105 it found. Recall@1 falls from 45.5% to 44.5%, and seed-to-seed spread (±2.7 points) is three times the method effect.
Finding 2
The gate does not beat a random gate
Against Random-Gated-Expansion at a similar budget, SQE gains +0.5 points (CI [−0.2, +1.1]); against Always-Expand, +0.1. Expanding about half the queries recovers essentially all of the benefit, but which half matters little.
Finding 3
Selectivity is still a large cost saving
SQE uses 405 tokens and 2.8 s per query versus 899 tokens and 5.4 s for Always-Expand, at equal Recall@5. Trace-only expansion is the worst operating point.
Pooled over 4,000 query rows. Seeds share about 30% of their memories, so these intervals are descriptive and likely too narrow.
Figure 2. Recall@5 per seed (dots) and 8-seed mean (bar). The dashed line is the Dense-Only mean. Seed variance dominates every expansion effect.Figure 3. Cost against quality: measured LLM tokens per query (seed 42) against 8-seed mean Recall@5. SQE matches Always-Expand at roughly half the tokens, but so does the random gate.
Why the gate fails
The top-1 score finds hard queries, but not the ones expansion can fix.
61.7% vs. 75.9%
Low confidence marks hard queries, not fixable ones
On seed 42, expanded queries reach 61.7% Recall@5 and non-expanded ones 75.9%. Expansion does not reliably repair the hard ones.
2.2 pts
The threshold barely matters
Recall@5 stays within a 2.2-point band while the expansion rate goes from 0% to 100%. Leave-one-seed-out threshold selection gives +0.6 points over dense (CI [−0.1, +1.3]); BM25-agreement variants do not close the gap.
5.3% help · 4.2% hurt
There is little headroom for any gate
An oracle that expands exactly the queries expansion helps reaches 73.5% Recall@5 against 68.2% for dense. The remaining 90% of queries have the same outcome either way.
Figure 4. Post-hoc threshold sweep on seed 42. Recall@5 (line) varies by about 2 points while the fraction of expanded queries (shaded) spans 0–100%.Figure 5. Held-out oracle headroom. Even a perfect per-query gate gains about 5 points over dense retrieval; always expanding gains about 1.
Method
SQE changes only the query at retrieval time. The memory store and both indices stay unchanged.
1RetrieveRun dense retrieval (bge-m3, FAISS) over the trace store and read the top-1 score.
2GateIf the top-1 score is at least τ = 0.65, return the dense list unchanged.
3ExpandOtherwise generate K = 2 hypothetical traces (dense + BM25) and P = 2 paraphrases (dense).
4FuseCombine the original list and the five probe lists with reciprocal rank fusion (k = 60).
Setup. Memory stores of 5,000 single agent steps from the SWE-smith tool split. For each seed (42–49) the store is resampled, both indices are rebuilt, and 500 targets are sampled. Expansion uses Qwen3.6-35B-A3B served with vLLM. Generated traces are retrieval probes, not claims that those executions occurred.
Limitations
Queries. Evaluation queries are written by an LLM that sees the target trace, which can leak wording. A human-audit packet is prepared but unlabeled.
Relevance labels. Only the exact source step counts as relevant, so Recall@k may understate all methods.
No downstream outcome. Recall@k is an intermediate signal; agent Pass@1 on SWE-bench-style tasks is not measured.
Statistics. Seeds share about 30% of their memories, so pooled intervals are likely too narrow; no multiplicity correction is applied.
Scope. One embedder, one generator, one trace source, and one memory size.
Reproducibility
scripts/07_verify_experiment.py recomputes every metric from the result files and checks each table against its declared source.
Experiment scriptsDataset preparation, retrieval, evaluation, verification, and audit tooling.
@misc{diei2026sqe,
title = {When Does Query Expansion Help Agent Memory? A Multi-Seed Study of Selective Expansion and Gate Reliability},
author = {Diei, Salomon},
year = {2026},
note = {Preprint},
url = {https://github.com/Salomondiei08/sqe-experiment}
}