Back to feed
arXiv cs.CL·

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

Signal
78
Hype
25
In three linesMADE is a multilingual agentic diagnosing engine that decomposes post-evaluation analysis into planning, aggregate analysis, instance-level inspection, and grounded report synthesis. Tested on 33 model families, 11 benchmarks, and 26 languages (8.66M evaluation records), MADE outperforms strongest baselines by 47% in diagnosis quality and is preferred by human experts in 87.9% of comparisons.
Read source
Your take?
AI AgentsMulti-agentEvalsBenchmarks

Summary generated by Claude — human-verified