MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights
Signal
78
Hype
25
In three linesMADE is a multilingual agentic diagnosing engine that decomposes post-evaluation analysis into planning, aggregate analysis, instance-level inspection, and grounded report synthesis. Tested on 33 model families, 11 benchmarks, and 26 languages (8.66M evaluation records), MADE outperforms strongest baselines by 47% in diagnosis quality and is preferred by human experts in 87.9% of comparisons.Read source
Your take?
Summary generated by Claude — human-verified