nuta/nsh

A command-line shell like fish, but POSIX compatible.

966 rs medium
1,960
Generated Behavioral Tests
83.9%
Best Score
Gemini 3.1 Pro

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

18 runs
# Model Score Cost Calls
1 Gemini 3.1 Pro 83.9% $0.92 64 trace →
2 GLM-5.2 83.6% $38.51 225 trace →
3 GPT 5.6 Sol (xhigh) 82.0% $7.07 47 trace →
4 GPT 5.5 (xhigh) 81.1% $9.98 81 trace →
5 Claude Sonnet 4.6 80.6% $27.07 496 trace →
6 Claude Opus 4.8 (xhigh) 78.6% $27.98 158 trace →
7 Claude Opus 4.6 76.8% $13.22 304 trace →
8 GPT 5.5 (high) 76.3% $2.79 42 trace →
9 GPT 5 mini 76.2% $0.01 10 trace →
10 Claude Opus 4.7 (xhigh) 71.3% $5.86 116 trace →
11 Claude Opus 4.7 66.2% $3.70 92 trace →
12 Gemini 3.6 Flash 55.2% $7.28 204 trace →
13 GPT 5.5 53.0% $1.33 23 trace →
14 GPT 5.4 50.3% $0.19 8 trace →
15 Gemini 3.5 Flash 41.7% $7.17 192 trace →
16 Claude Haiku 4.5 36.8% $1.16 148 trace →
17 Gemini 3 Flash 31.3% $0.28 122 trace →
18 GPT 5.4 mini 0.0% $0.02 14 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run