MirrorCode benchmark: AI model costs $2,600 to recreate a program in 19 days
Epoch AI's MirrorCode benchmark tests AI models' ability to recreate programs without access to the original code, with Claude Opus 4.7 leading at a 56% solve rate.
Published 6sem1 sourceNotable
Lire en français
≈ 34s
The fact
One model ran continuously for 19 days on a complex task, costing $2,600, while all tested models still fail on the most challenging cases.
Claude Opus 4.7 rebuilt a 16,000-line toolkit in 14 hours, highlighting both advancements and current limitations of LLMs in code generation.
Click the link to read an article on the topic: