
Microsoft Research's Lens proves detailed captions matter more than raw scale for training efficient image generators
Quick Answer
Microsoft Research's Lens is a text-to-image model with 3.8 billion parameters that rivals larger models using 800 million detailed captions from GPT-4.1, achieving high performance at a lower training cost.
Quick Take
The model's code and weights are available as open-source, demonstrating the importance of detailed captions over sheer scale in training efficiency.
Key Points
- Lens achieves performance comparable to larger models with only 3.8 billion parameters.
- Utilizes 800 million detailed captions generated by GPT-4.1 for training.
- Significantly reduces training costs while maintaining benchmark results.
- Open-source code and weights are available for public use.
- Highlights the effectiveness of detailed captions over vague alternatives.
Source Excerpt
Microsoft Research presents Lens, a text-to-image model with just 3. 8 billion parameters that matches much larger rivals on benchmarks, at a fraction of the training cost. The secret sauce: 800 million detailed image captions generated by GPT-4. 1 instead of vague web alt-text. Code and weights are openly available under an open-source license.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

