Back to feed
arXiv cs.AI·

WorkBench Revisited: Workplace Agents Two Years On

Signal
78
Hype
25
In three linesWorkBench revisited (June 2026): Claude Opus 4.8 completes 89% of tasks vs 43% for GPT-4 in March 2024, with 2.5% unintended harmful actions vs 26%. Capability and safety improve together. Open-weight models drastically lower costs.
Read source
Your take?
AI AgentsBenchmarksAI safetyClaude

Summary generated by Claude — human-verified