Microsoft Research: detailed captions outperform raw scale for efficient image generators
Lens, Microsoft Research's new text-to-image model, reaches 3.8 billion parameters while matching much larger competitors on standard benchmarks
Published 9sem1 sourceNotable
Lire en français
≈ 25s
The fact
Efficiency stems from 800 million detailed image descriptions generated by GPT-4.1, significantly more precise than web alt-text, prioritizing quality over raw scale
Model code and weights released open-source, enabling community-driven iteration on parameter-to-quality optimization
Click the link to read an article on the topic: