Back to feed
arXiv cs.AI·

CEO-Bench: Can Agents Play the Long Game?

Signal
78
Hype
25
In three linesCEO-Bench evaluates agents' ability to handle complex long-horizon tasks by simulating a 500-day startup operation. The agent manages pricing, marketing, budgeting through a Python interface. Only Claude Opus 4.8 and GPT-5.5 exceed the $1M starting balance, neither consistently profitable.
Read source
Your take?
AI AgentsBenchmarksReasoningClaudeGPT

Summary generated by Claude — human-verified