The fact
The problem amplifies with scale: in a cluster of 1,000 GPU nodes, a single slow unit degrades overall performance by 34 percent.
This finding forces rethinking fault tolerance and early detection of hardware issues.
Click the link to read an article on the topic: