Back to feed
arXiv cs.AI·

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

Signal
78
Hype
15
In three linesMemTrace is a benchmark evaluating long-term memory in LLM agents across three dimensions: memory age, question type (current state, earlier state, trajectory), and evidence conditions. Testing 13 configurations, the study finds that evidence use is the primary bottleneck (10× more often retrievable than missing), not retrieval itself.
Read source
Your take?
AI AgentsEvalsBenchmarks

Summary generated by Claude — human-verified