AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints
Signal
78
Hype
15
In three linesAdaPlanBench is an interactive benchmark evaluating LLM agents' ability to adaptively plan and replan under progressively revealed world and user constraints. Built on 307 household tasks, it tests 10 leading models: best achieves 67.75% accuracy. Performance degrades with constraint accumulation, particularly for user constraints.Read source
Your take?
Summary generated by Claude — human-verified