chmln/sd

Intuitive find & replace CLI (sed alternative)

7,072 rs medium
810
Generated Behavioral Tests
97.4%
Best Score
GPT 5.5 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.5 (xhigh) 97.4% $7.94 89 trace →
2 Claude Opus 4.8 (xhigh) 96.4% $10.85 102 trace →
3 Claude Opus 4.7 (xhigh) 95.2% $13.81 211 trace →
4 GPT 5.6 Sol (xhigh) 94.2% $2.61 20 trace →
5 GLM-5.2 92.2% $14.17 190 trace →
6 GPT 5.6 Sol (medium) 91.0% $0.84 17 trace →
7 Claude Opus 4.7 90.9% $3.64 113 trace →
8 GPT 5.5 (high) 89.9% $3.30 37 trace →
9 Claude Opus 4.6 89.9% $7.07 270 trace →
10 Gemini 3.6 Flash 88.9% $2.77 176 trace →
11 Gemini 3.5 Flash 88.8% $1.50 118 trace →
12 GPT 5.5 86.0% $0.87 17 trace →
13 Claude Sonnet 4.6 85.1% $3.76 181 trace →
14 GPT 5.4 81.6% $0.22 8 trace →
15 Gemini 3 Flash 80.7% $0.25 111 trace →
16 Gemini 3.1 Pro 80.0% $1.33 122 trace →
17 GPT 5 mini 70.9% $0.03 22 trace →
18 Claude Haiku 4.5 68.6% $0.87 149 trace →
19 GPT 5.4 mini 45.8% $0.04 14 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run