peco/peco

Simplistic interactive filtering tool

7,881 go medium
1,215
Generated Behavioral Tests
89.6%
Best Score
GLM-5.2

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GLM-5.2 89.6% $25.62 210 trace →
2 GPT 5.6 Sol (xhigh) 86.0% $5.75 40 trace →
3 Claude Opus 4.8 (xhigh) 82.8% $41.44 230 trace →
4 Claude Opus 4.6 76.8% $14.19 391 trace →
5 Gemini 3.6 Flash 76.4% $5.02 181 trace →
6 Gemini 3.5 Flash 76.1% $3.84 152 trace →
7 Claude Sonnet 4.6 74.7% $15.96 351 trace →
8 Claude Opus 4.7 (xhigh) 72.3% $9.13 157 trace →
9 Gemini 3.1 Pro 71.9% $1.57 74 trace →
10 GPT 5.5 (high) 69.8% $4.35 46 trace →
11 GPT 5.5 (xhigh) 68.7% $8.02 84 trace →
12 Claude Opus 4.7 67.6% $2.50 84 trace →
13 GPT 5.4 66.6% $0.29 11 trace →
14 GPT 5.5 65.4% $1.13 16 trace →
15 GPT 5.6 Sol (medium) 61.7% $0.36 7 trace →
16 GPT 5 mini 56.9% $0.03 10 trace →
17 Claude Haiku 4.5 51.9% $1.14 146 trace →
18 GPT 5.4 mini 29.0% $0.03 8 trace →
19 Gemini 3 Flash 11.4% $0.20 45 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run