Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning
Signal
75
Hype
15
In three linesIRTS-ToolBench, a benchmark of 1,700 questions across 10 task types and 13 domains, evaluates how LLMs and AI agents handle irregular time series (asynchronous, informative missing values, variable sampling frequencies). Bridges gap between existing TSQA benchmarks (regular data) and real-world deployments.Read source
Your take?
Summary generated by Claude — human-verified