FactaeThe Factual News
Smaller KV caches fail to speed up transformers in long-context generation | Factae