Back to feed
arXiv cs.CL·

EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

Signal
78
Hype
25
In three linesEvoBrowseComp is an evolving benchmark of 400 English and 400 Chinese questions to evaluate search agents (LLM + web tools). Unlike static BrowseComp, it uses live-web traversal and a three-agent framework (QA synthesis, information filtering, high-level guidance) to prevent contamination and parametric memorization. The benchmark auto-updates regularly.
Read source
Your take?
AI AgentsBenchmarksEvalsReasoning

Summary generated by Claude — human-verified