Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.6 Sol (xhigh) 84.1% $13.16 87 trace →
2 Claude Opus 4.8 (xhigh) 79.2% $26.15 183 trace →
3 Claude Sonnet 4.6 74.7% $22.13 450 trace →
4 GPT 5.5 (high) 71.8% $5.18 46 trace →
5 Claude Opus 4.6 70.9% $14.37 258 trace →
6 GLM-5.2 68.9% $27.00 219 trace →
7 Gemini 3.6 Flash 65.4% $4.87 145 trace →
8 Claude Opus 4.7 (xhigh) 57.6% $6.97 133 trace →
9 GPT 5.5 (xhigh) 54.7% $8.30 70 trace →
10 GPT 5.5 52.1% $1.56 22 trace →
11 Gemini 3.5 Flash 52.1% $6.70 180 trace →
12 Claude Opus 4.7 41.1% $4.69 89 trace →
13 GPT 5.6 Sol (medium) 36.5% $1.39 18 trace →
14 GPT 5.4 19.3% $0.46 12 trace →
15 Gemini 3 Flash 17.1% $0.24 89 trace →
16 Claude Haiku 4.5 14.4% $0.48 98 trace →
17 GPT 5 mini 10.4% $0.02 17 trace →
18 GPT 5.4 mini 8.6% $0.05 10 trace →
19 Gemini 3.1 Pro 0.0% $0.90 86 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run