Long-running AI agents may appear to maintain a broad external memory while repeatedly drawing from only a small, familiar core. A preprint by Xinyuan Song of Emory University and Zekun Cai of the University of Tokyo and LocationMind examines this uneven pattern and argues that rare memories deserve separate measurement because prediction errors can accumulate where retrieval is weakest. The paper is under review, so its results remain preliminary.
The study focuses on agents that keep information outside the language model in a memory stream, vector store, graph, temporal database, or paging system. These systems are commonly evaluated by task completion and token cost. Song and Cai propose a third measure: the shape of memory use—the distribution of attention across stored states, transitions, and evidence when only a limited amount can fit into each prompt.

Their audit found that finite retrieval repeatedly concentrates access around a compact core while leaving a long tail of less frequently visited states. The authors caution that the pattern does not always carry the same meaning. Random-walk agents with no semantic intent also produced heavy-tailed retrieval, largely because top-k systems keep returning to the same hubs. Those traces were compatible with a log-normal pattern and failed the paper’s stronger tests for a pure power law or critical dynamics.
Semantic language-model policies produced the clearest task-shaped core-and-tail behavior, according to the paper. In the audited traces, those policies showed concentration compatible with a truncated power law. The distinction matters because it separates a retrieval artifact from a pattern that may reflect what an agent treats as relevant while pursuing a goal. The researchers nevertheless avoid claiming that agent memory exhibits full self-organized criticality.
To act on the measurement, the authors introduce the Core–Tail World Model, or CTWM. The controller ranks memories using relevance and transition support, divides them into core and tail slices, and allocates prompt space according to a single concentration setting. It does not retrain the underlying language model or require the prompt to carry the complete interaction history. A summarized portion of the tail is retained rather than discarded.

On the paper’s Synthetic Graph World benchmark, CTWM preserved complete state and transition coverage while using 5.9 percent fewer prompt tokens than a graph-memory baseline. It also reduced prediction error for the bottom half of states by visit frequency by 13.6 percent. In paired comparisons, the approach produced consistent token savings on ALFWorld and cut tokens by 24.48 percent on LongMemEval while maintaining aggregate accuracy.
Those gains come with limits. The authors say their ALFWorld policy was not strong enough to support a claim of improved task success, so that benchmark result is about token efficiency rather than better completion rates. LongMemEval also revealed a weakness on multi-session recall, suggesting that aggressive core-tail budgeting should be relaxed when a query depends heavily on retrieving uncommon details. The topology and allocation sweeps are finite audits, not exhaustive tests across environments.
The broader implication is that external memory is not a neutral archive once context and retrieval budgets become binding. Its design determines which experiences stay prominent and which drift toward the edge of the agent’s working model. If later studies confirm the result, the distribution of memory access could become a practical control signal: developers could spend fewer tokens on repeatedly retrieved core material while deliberately preserving enough rare evidence to keep long-horizon behavior from becoming efficient but forgetful.

Comments
Loading comments…