Back to feed
arXiv cs.CL·

Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents

Signal
78
Hype
22
In three linesDialogue SWE-Bench is an automatic benchmark for evaluating coding agents through user dialogue. Authors introduce a persona-grounded user simulator and a schema-guided agent improving baselines by 3-14%. Key finding: better coding models don't necessarily translate to better dialogue capabilities.
Read source
Your take?
BenchmarksCode generationAI AgentsPapers

Summary generated by Claude — human-verified