Back to feed
arXiv cs.CL·

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

Signal
78
Hype
25
In three linesLakeQA is a QA benchmark over 9.5 TB of heterogeneous data (Wikipedia + government sources) requiring search and multi-hop reasoning. GPT-4.5 achieves 18.37% exact-match. Evaluates LLM agents' ability to discover and analyze documents in massive data lakes.
Read source
Your take?
BenchmarksReasoningRAGAI Agents

Summary generated by Claude — human-verified