X-MADAM-RAG: Diagnosing and Handling Chinese-English Evidence Conflict in Retrieval-Augmented Generation
Signal
72
Hype
18
In three linesX-MADAM-RAG diagnoses conflicts between Chinese and English evidence in RAG systems. On the X-RAMDocs-ZHEN benchmark (300 examples), the pipeline achieves 96.67% strict accuracy with Qwen2.5-7B-Instruct, but fails on naturalized variants (30% accuracy), revealing document-level extraction as the main bottleneck.Read source
Your take?
Summary generated by Claude — human-verified