Back to feed
arXiv cs.CL·

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Signal
78
Hype
15
In three linesAdaPlanBench is an interactive benchmark evaluating LLM agents' ability to adaptively plan and replan under progressively revealed world and user constraints. Built on 307 household tasks, it tests 10 leading models: best achieves 67.75% accuracy. Performance degrades with constraint accumulation, particularly for user constraints.
Read source
Your take?
AI AgentsReasoningBenchmarksEvals

Summary generated by Claude — human-verified