Back to feed
arXiv cs.AI·

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

Signal
75
Hype
20
In three linesLongWebBench is a benchmark evaluating long-horizon webpage generation by vision-language models. It contains 490 real-world pages for structural evaluation and 507 goal-oriented interaction tasks over 129 pages. Experiments show structural fidelity degrades with webpage length, and visually plausible generations often fail to support multi-step executable interactions.
Read source
Your take?
VisionBenchmarksAI AgentsCode generation

Summary generated by Claude — human-verified