FactaeThe Factual News

Why small language models fail at rare tasks

Small models forget rare tasks because frequent ones constantly overwrite learned skills during training

Published 9sem1 sourceNotable
Lire en français
20s

The fact

Researchers tested models from 4M to 4B parameters to identify this mechanism

Increasing target task frequency in training data may be sufficient—no need to scale up model size

Click the link to read an article on the topic:
Explore this topic
What if you saw the whole news differently?Factae cross-checks hundreds of sources worldwide to keep only the fact, no opinion. Explore the front page.
Auto-synthesis from 1 media source · identified on June 7, 2026
Back to home
Discover

Read more

All ia →