altdesktop/i3-style

🎨 Make your i3 config a little more stylish.

678 rs medium
539
Generated Behavioral Tests
93.3%
Best Score
GPT 5.6 Sol (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.6 Sol (xhigh) 93.3% $4.55 40 trace →
2 GPT 5.5 (xhigh) 93.3% $6.85 67 trace →
3 Claude Opus 4.7 (xhigh) 92.2% $18.42 206 trace →
4 GLM-5.2 91.5% $20.12 197 trace →
5 Claude Opus 4.8 (xhigh) 88.7% $28.39 141 trace →
6 GPT 5.5 (high) 87.0% $4.41 43 trace →
7 GPT 5.5 87.0% $2.21 28 trace →
8 Claude Opus 4.6 80.0% $12.87 205 trace →
9 Claude Sonnet 4.6 77.9% $16.07 322 trace →
10 Claude Opus 4.7 73.5% $7.30 165 trace →
11 GPT 5.6 Sol (medium) 72.2% $0.60 12 trace →
12 Gemini 3.5 Flash 72.0% $5.09 139 trace →
13 Gemini 3.6 Flash 69.9% $3.33 121 trace →
14 GPT 5.4 48.8% $0.31 8 trace →
15 Gemini 3 Flash 48.8% $0.53 107 trace →
16 Gemini 3.1 Pro 42.7% $2.02 133 trace →
17 GPT 5.4 mini 35.4% $0.02 8 trace →
18 Claude Haiku 4.5 34.9% $0.85 117 trace →
19 GPT 5 mini 30.2% $0.02 11 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run