Our benchmark runs on real bugs, not synthetic ones. MacroscopeBench asks whether a code reviewer would have caught each defect at the commit that introduced it, the same benchmark we use to build and tune Macroscope Code Review.
Read about the benchmark| Model | ||||||
|---|---|---|---|---|---|---|
Claude Opus 5.5 max effort | 82.3 | 80.6% | 84.0% | 3.01 | $11.58 | 10m 8s |
GPT-6 Sol max effort | 80.3 | 73.8% | 88.0% | 3.94 | $3.68 | 4m 57s |
Claude Opus 5.5 xhigh effort | 78.9 | 75.0% | 83.2% | 2.95 | $3.51 | 2m 52s |
GPT-6 Astra max effort | 78.0 | 67.8% | 91.7% | 6.26 | $7.87 | 3m 49s |
GPT-5.6 Sol high effort | 77.8 | 70.8% | 86.2% | 3.73 | $3.90 | 3m |
GLM 5.3 max effort | 77.4 | 76.4% | 78.4% | 2.24 | $3.86 | 12m 30s |
GPT-6 Sol xhigh effort | 77.2 | 69.0% | 87.6% | 3.86 | $2.00 | 3m 6s |
Grok 4.6 xhigh effort | 75.9 | 69.7% | 83.3% | 3.67 | $4.68 | 14m 41s |
Grok 4.6 high effort | 75.1 | 69.9% | 81.1% | 3.17 | $3.51 | 15m 30s |
GPT-6 Sol high effort | 74.3 | 65.5% | 85.9% | 3.62 | $1.23 | 2m 3s |
Claude Opus 5.5 high effort | 73.1 | 68.1% | 78.9% | 2.35 | $1.30 | 1m 3s |
Kimi K3 max effort | 72.9 | 66.7% | 80.3% | 2.18 | $2.82 | 8m 13s |
GPT-6 Luna max effort | 72.5 | 61.1% | 89.1% | 4.28 | $0.13 | 4m 1s |
DeepSeek V4.1 Flash max effort | 72.0 | 83.1% | 63.5% | 1.27 | $0.75 | 16m 17s |
DeepSeek V4.1 Flash low effort | 70.7 | 77.5% | 65.0% | 1.37 | $0.38 | 8m 9s |
DeepSeek V4.1 Flash high effort | 70.5 | 77.5% | 64.6% | 1.33 | $0.38 | 7m 40s |
Claude Opus 5.5 medium effort | 70.1 | 63.4% | 78.3% | 2.35 | $0.97 | 44s |
Claude Opus 5 high effort | 69.9 | 76.6% | 64.2% | 1.20 | $6.50 | 4m 18s |
GPT-6 Sol medium effort | 69.2 | 59.7% | 82.2% | 3.08 | $0.69 | 1m 31s |
GPT-6 Luna xhigh effort | 68.2 | 56.3% | 86.6% | 3.61 | $0.06 | 2m 2s |
GPT-6 Luna high effort | 65.6 | 53.0% | 86.0% | 3.50 | $0.05 | 1m 40s |
Claude Opus 5 medium effort | 65.5 | 68.8% | 62.5% | 1.18 | $3.50 | 2m 12s |
GPT-6 Luna medium effort | 62.4 | 50.9% | 80.5% | 2.60 | $0.03 | 1m 14s |
GPT-6 Sol low effort | 61.8 | 52.1% | 75.9% | 2.35 | $0.36 | 58s |
Claude Opus 5.5 low effort | 61.4 | 52.3% | 74.3% | 2.21 | $0.68 | 29s |
Claude Opus 5 low effort | 59.4 | 57.4% | 61.5% | 1.12 | $1.40 | 56s |
GPT-6 Luna low effort | 44.7 | 38.0% | 54.5% | 1.00 | $0.02 | 42s |
Every model is run 3 independent times through the benchmark to account for nondeterminism and diverging results. Every metric, unless stated otherwise, is the arithmetic mean of the three independent runs.