Byron/dua-cli

View disk space usage and delete unwanted data, fast.

5,794 rs medium
709
Generated Behavioral Tests
98.3%
Best Score
Claude Opus 4.8 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 Claude Opus 4.8 (xhigh) 98.3% $32.06 200 trace →
2 GPT 5.5 (xhigh) 93.1% $11.00 74 trace →
3 GLM-5.2 92.2% $25.55 187 trace →
4 Claude Opus 4.7 (xhigh) 90.8% $10.03 161 trace →
5 GPT 5.6 Sol (xhigh) 90.1% $4.98 29 trace →
6 Gemini 3.6 Flash 87.9% $2.94 114 trace →
7 Claude Opus 4.6 86.7% $10.56 277 trace →
8 Gemini 3.5 Flash 85.6% $3.76 149 trace →
9 GPT 5.5 (high) 84.6% $2.81 29 trace →
10 Claude Sonnet 4.6 83.9% $22.21 502 trace →
11 GPT 5.6 Sol (medium) 80.0% $0.56 10 trace →
12 GPT 5.5 78.3% $1.28 25 trace →
13 Claude Opus 4.7 70.8% $2.82 80 trace →
14 Claude Haiku 4.5 49.1% $1.27 137 trace →
15 GPT 5.4 40.5% $0.17 7 trace →
16 GPT 5 mini 38.9% $0.04 22 trace →
17 GPT 5.4 mini 26.0% $0.05 11 trace →
18 Gemini 3.1 Pro 4.2% $1.87 131 trace →
19 Gemini 3 Flash 0.0% $0.31 100 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run