Back to feed
arXiv cs.AI·

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

Signal
72
Hype
18
In three linesWorldLines is a long-horizon embodied agent benchmark testing memory in dynamic household environments. The dataset includes temporally extended traces with dialogues, actions, and object/device state changes. ObsMem, an observer-grounded memory framework, maintains visibility-aware memories and action-native state trails for state-informed decisions.
Read source
Your take?
AI AgentsBenchmarksReasoningPapers

Summary generated by Claude — human-verified