Back to feed
arXiv cs.CL·

Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

Signal
78
Hype
15
In three linesXBCP, a controlled benchmark, evaluates deep research agents' ability to operate across languages. Four agents tested with dense and sparse retrievers across 12 languages show substantial degradation: evidence recall loss, reduced calibration, unreliable citations. Problems persist even when gold evidence is directly supplied.
Read source
Your take?
AI AgentsRAGBenchmarksEvals

Summary generated by Claude — human-verified