AI Now Solves 97% of Coding Benchmarks and Fails 77% of Real Software Work. Both Numbers Are True.
The leaderboard showed Claude Opus 5 resolving 97% of tasks on SWE-bench Verified, the industry standard test for AI coding ability. GPT and other frontier models clustered right behind it, most of them above 95%. T