Back to feed
The Decoder·

Microsoft Research's Lens proves detailed captions matter more than raw scale for training efficient image generators

Signal
78
Hype
25
In three linesMicrosoft Research presents Lens, a text-to-image model with 3.8 billion parameters that matches much larger rivals on benchmarks at a fraction of training cost. Key innovation: 800 million detailed captions generated by GPT-4.1 instead of vague web alt-text. Code and weights released under open-source license.
Read source
Your take?
Image generationBenchmarksOpen sourcePapers

Summary generated by Claude — human-verified