When Does Query Expansion
Help Agent Memory?

A Multi-Seed Study of Selective Expansion and Gate Reliability

Preprint 2026 · Earlier version presented at NAFSIK 2026

Salomon Diei
School of Computer Science and Engineering, KOREATECH

The talk in one minute

Problem

Agents store execution traces but later search them with natural-language questions. The memory exists, but the query does not look like it.

Idea

Expand the query with hypothetical traces and paraphrases, but only when the first retrieval looks uncertain.

Result

Selective expansion halves the cost of always expanding with no loss in Recall@5. But the confidence gate is no better than random.

Why agent memory retrieval fails

What is stored

  • Terminal outputs and tracebacks
  • Test failures
  • Patch diffs and commands
  • Single agent steps: thought, action, observation

What is asked later

  • Natural-language questions
  • "How did I fix the off-by-one line number bug?"
  • Short reminders and issue summaries

The memory may be present, but the query and the memory are written in different “languages”.

The question

Can a cheap confidence signal decide which agent-memory queries are worth expanding?

How SQE works

SQE retrieval flow

Dense retrieval first. If the top-1 score is below τ = 0.65, generate 2 hypothetical traces (dense + BM25) and 2 paraphrases (dense), then fuse all lists with RRF (k = 60). The memory store and indices never change.

Evaluation design

5,000

agent steps per memory store

500

queries per seed

8

resampled stores, indices rebuilt

4,000

queries in total

Data and models

SWE-smith agent steps (tool split). bge-m3 + FAISS for dense, BM25 for sparse, Qwen3.6-35B-A3B for expansion.

The key control

A random gate runs the identical expansion code at a matched rate. It separates expanding some queries from choosing which ones.

Main result: a small Recall@5 gain

69.4%

SQE

68.5%

Dense-Only

69.2%

Always-Expand

68.9%

Random gate

Mean Recall@5 over 8 seeds. SQE is slightly better than dense retrieval at Recall@5 (+0.9 points), but worse at Recall@1 (44.5% vs 45.5%). It recovers 140 targets dense retrieval missed and loses 105 it found.

Seed variance dominates every effect

Recall@5 per seed and 8-seed mean
±2.7

points of seed-to-seed spread: three times larger than the method effect.

The gate does not beat a random gate

+0.9

vs Dense-Only

95% CI [+0.1, +1.7]

+0.5

vs Random gate

95% CI [−0.2, +1.1]. Not significant.

+0.1

vs Always-Expand

95% CI [−0.4, +0.7]

Expanding about half the queries recovers essentially all of the benefit of expanding every query, but which half matters little.

Selectivity is still a large cost saving

Tokens per query against Recall@5
−55%

LLM tokens vs Always-Expand: 405 vs 899 tokens and 2.8 s vs 5.4 s per query, at equal Recall@5.

But the random gate sits at the same point.

Why the gate fails

61.7%

Hard, not fixable

Expanded queries reach 61.7% Recall@5 vs 75.9% for the rest. The gate finds hard queries, but expansion does not repair them.

2.2 pts

Threshold barely matters

Recall@5 moves within a 2.2-point band while the expansion rate goes from 0% to 100%.

5.3%

Little headroom

Expansion helps 5.3% of queries and hurts 4.2%. Even an oracle gate gains only about 5 points.

The evidence behind the gate result

Threshold sweep on seed 42

Threshold sweep (seed 42)

Oracle gate headroom

Held-out oracle headroom

What the evidence supports

Supported

  • SQE is a retrieval-time method; memory writing is unchanged.
  • It matches Always-Expand at about half the tokens.
  • It gives a small Recall@5 gain over dense retrieval.
  • The top-1-score gate is not better than random.

Not claimed

  • No downstream agent Pass@1 result.
  • No human-validated query quality (audit packet prepared, unlabeled).
  • Intervals pool overlapping seeds, so they are descriptive.
  • One embedder, one generator, one trace source.

Next steps

1

Learned gates against random-budget controls

Predicting which queries expansion will fix is the open problem.

2

Downstream agent success

Measure whether retrieved memories change Pass@1 on SWE-bench-style tasks.

3

Human-audited queries

Label the prepared 100-query audit packet.

4

Confirmatory statistics

Pre-register one contrast and use seed- or task-clustered inference.

Thank you

Confidence gating, not expansion itself, is the open problem.

Salomon Diei · salomon@koreatech.ac.kr