Our benchmark runs on real bugs, not synthetic ones. MacroscopeBench asks whether a code reviewer would have caught each defect at the commit that introduced it, the same benchmark we use to build and tune Macroscope Code Review.
Read about the benchmark| Model | ||||||
|---|---|---|---|---|---|---|
GPT-6 Astra max effort | 78.0 | 67.8% | 91.7% | 6.26 | $7.87 | 3m 49s |
GPT-5.6 Sol high effort | 77.8 | 70.8% | 86.2% | 3.73 | $3.90 | 3m |
GLM 5.3 max effort | 77.4 | 76.4% | 78.4% | 2.24 | $3.86 | 12m 30s |
Grok 4.6 xhigh effort | 75.9 | 69.7% | 83.3% | 3.67 | $4.68 | 14m 41s |
Grok 4.6 high effort | 75.1 | 69.9% | 81.1% | 3.17 | $3.51 | 15m 30s |
Kimi K3 max effort | 72.9 | 66.7% | 80.3% | 2.18 | $2.82 | 8m 13s |
DeepSeek V4.1 Flash max effort | 72.0 | 83.1% | 63.5% | 1.27 | $0.75 | 16m 17s |
DeepSeek V4.1 Flash low effort | 70.7 | 77.5% | 65.0% | 1.37 | $0.38 | 8m 9s |
DeepSeek V4.1 Flash high effort | 70.5 | 77.5% | 64.6% | 1.33 | $0.38 | 7m 40s |
Claude Opus 5 high effort | 69.9 | 76.6% | 64.2% | 1.20 | $6.50 | 4m 18s |
Claude Opus 5 medium effort | 65.5 | 68.8% | 62.5% | 1.18 | $3.50 | 2m 12s |
Claude Opus 5 low effort | 59.4 | 57.4% | 61.5% | 1.12 | $1.40 | 56s |
Every model is run 3 independent times through the benchmark to account for nondeterminism and diverging results. Every metric, unless stated otherwise, is the arithmetic mean of the three independent runs.