WorkBench Revisited: Workplace Agents Two Years On
Signal
78
Hype
25
In three linesWorkBench revisited (June 2026): Claude Opus 4.8 completes 89% of tasks vs 43% for GPT-4 in March 2024, with 2.5% unintended harmful actions vs 26%. Capability and safety improve together. Open-weight models drastically lower costs.Read source
Your take?
Summary generated by Claude — human-verified