The fact
This efficiency contrasts with industry's trend toward giant models consuming gigawatts of energy at inference time.
Teams can now deploy this model locally or on lightweight inference servers, reducing computational costs.
Click the link to read an article on the topic: