direnv/direnv

unclutter your .profile

14,998 go medium
849
Generated Behavioral Tests
80.9%
Best Score
GPT 5.5 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.5 (xhigh) 80.9% $6.28 45 trace →
2 GPT 5.6 Sol (xhigh) 79.6% $7.90 50 trace →
3 Claude Opus 4.8 (xhigh) 76.5% $31.73 215 trace →
4 Gemini 3.6 Flash 75.6% $4.81 173 trace →
5 GPT 5.5 (high) 73.3% $3.77 35 trace →
6 GLM-5.2 62.9% $20.47 225 trace →
7 Claude Opus 4.6 62.0% $14.72 312 trace →
8 GPT 5.6 Sol (medium) 60.2% $1.61 13 trace →
9 Gemini 3.5 Flash 58.8% $7.43 283 trace →
10 Claude Sonnet 4.6 57.6% $20.85 417 trace →
11 GPT 5.5 48.8% $1.46 17 trace →
12 Claude Opus 4.7 (xhigh) 33.5% $7.67 169 trace →
13 GPT 5.4 32.9% $0.22 8 trace →
14 Claude Opus 4.7 32.2% $2.28 84 trace →
15 Gemini 3.1 Pro 27.1% $3.32 161 trace →
16 Claude Haiku 4.5 19.3% $1.04 133 trace →
17 GPT 5.4 mini 12.0% $0.05 31 trace →
18 GPT 5 mini 10.1% $0.03 19 trace →
19 Gemini 3 Flash 6.8% $0.73 127 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run