Back to feed
arXiv cs.CL·

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Signal
78
Hype
15
In three linesTensorBench is a benchmark of 199 coding tasks (feature additions and refactoring) on an open-source compiler-based tensor framework extending PyTorch. Evaluation of 7 coding agents: pass rates from 64.8% (strongest) to 22.1% (weakest), with low inter-agent agreement (κ=0.05 for top two agents).
Read source
Your take?
BenchmarksCode generationAI AgentsEvals

Summary generated by Claude — human-verified