FactaeThe Factual News
KVQuant enables running 70B LLMs on 8GB RAM with real-time cache compression | Factae