The fact
Performance gains: ~1,000 tokens/sec on H100 GPU (4x faster than comparable autoregressive models), but with lower output quality.
Experimental positioning: Google offers it to developers as a research tool, acknowledging the speed-quality tradeoff.
Click the link to read an article on the topic: