tinycc/tinycc

Unofficial mirror of mob development branch

2,843 c
1,978
Generated Behavioral Tests
72.4%
Best Score
GPT 5.5 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.5 (xhigh) 72.4% $6.30 65 trace →
2 Gemini 3.6 Flash 67.3% $3.22 150 trace →
3 GPT 5.6 Sol (medium) 64.9% $0.59 10 trace →
4 Gemini 3.5 Flash 59.7% $2.29 117 trace →
5 GPT 5.5 (high) 17.7% $3.43 36 trace →
6 GPT 5.5 16.0% $1.29 17 trace →
7 GPT 5.4 12.8% $0.43 12 trace →
8 GLM-5.2 11.1% $19.13 193 trace →
9 GPT 5.6 Sol (xhigh) 10.6% $6.68 41 trace →
10 Claude Sonnet 4.6 9.2% $100.12 879 trace →
11 Claude Opus 4.6 8.0% $12.01 189 trace →
12 Gemini 3 Flash 4.8% $0.16 50 trace →
13 Claude Opus 4.8 (xhigh) 4.3% $16.40 119 trace →
14 Gemini 3.1 Pro 3.8% $1.52 108 trace →
15 Claude Opus 4.7 (xhigh) 3.7% $0.75 22 trace →
16 Claude Opus 4.7 3.3% $0.45 22 trace →
17 GPT 5.4 mini 2.2% $0.02 7 trace →
18 GPT 5 mini 1.6% $0.02 12 trace →
19 Claude Haiku 4.5 0.0% $0.40 108 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run