The fact
SWE-Explore benchmark is the first to separately test code search from actual repair effectiveness
Insufficient context significantly limits agent performance, even for the most advanced models
Click the link to read an article on the topic: