Back to feed
arXiv cs.CL·

GENEB: Why Genomic Models Are Hard to Compare

Signal
78
Hype
15
In three linesGENEB is a large-scale diagnostic benchmark evaluating 40 genomic foundation models across 100 tasks in 13 functional categories under a unified probing protocol. Analysis shows aggregate leaderboards are unstable: model rankings vary sharply across task categories, scale provides modest and inconsistent gains, and architectural alignment frequently outweighs parameter count.
Read source
Your take?
BenchmarksPapersEvals

Summary generated by Claude — human-verified