UK AI Security Institute finds benchmarks underestimate AI agent capabilities
A study by the UK's AI Security Institute (AISI) reveals that standard benchmarks underestimate AI agent capabilities due to compute budget constraints.
Published 5sem1 sourceNotable
Lire en français
≈ 24s
60 %
Newer models benefit the most, indicating actual progress is 60% faster than previously meas…
The fact
Software engineering task success rates improve by 25% when the token budget is increased tenfold.
Newer models benefit the most, indicating actual progress is 60% faster than previously measured.
Click the link to read an article on the topic: