Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.6 Sol (xhigh) 39.0% $10.13 64 trace →
2 GPT 5.5 (xhigh) 22.6% $11.21 70 trace →
3 GPT 5.5 (high) 21.5% $5.82 35 trace →
4 GLM-5.2 21.3% $12.22 113 trace →
5 GPT 5.6 Sol (medium) 18.8% $1.37 15 trace →
6 Gemini 3.6 Flash 18.4% $12.90 146 trace →
7 Claude Opus 4.8 (xhigh) 18.0% $21.15 142 trace →
8 GPT 5.5 15.6% $1.33 15 trace →
9 Claude Opus 4.6 14.2% $10.17 192 trace →
10 GPT 5.4 7.1% $0.33 10 trace →
11 Gemini 3.5 Flash 5.5% $3.29 100 trace →
12 Gemini 3.1 Pro 5.2% $1.43 99 trace →
13 Gemini 3 Flash 3.9% $0.34 97 trace →
14 Claude Opus 4.7 (xhigh) 3.9% $1.81 56 trace →
15 GPT 5.4 mini 3.9% $0.04 8 trace →
16 Claude Opus 4.7 3.7% $0.81 34 trace →
17 Claude Haiku 4.5 3.0% $0.24 67 trace →
18 GPT 5 mini 1.8% $0.02 16 trace →
19 Claude Sonnet 4.6 1.3% $26.62 457 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run