sharkdp/hyperfine

A command-line benchmarking tool

27,960 rs
291
Generated Behavioral Tests
77.0%
Best Score
GPT 5.5 (high)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

16 runs
# Model Score Cost Calls
1 GPT 5.5 (high) 77.0% $2.48 25 trace →
2 GPT 5.5 49.1% $1.37 18 trace →
3 GPT 5.4 48.8% $0.25 9 trace →
4 Gemini 3.1 Pro 43.6% $1.29 55 trace →
5 Claude Haiku 4.5 24.1% $0.58 85 trace →
6 GPT 5.4 mini 7.2% $0.05 10 trace →
7 GPT 5 mini 4.5% $0.02 9 trace →
8 GPT 5.6 Sol (xhigh) 0.0% $3.73 23 trace →
9 GPT 5.5 (xhigh) 0.0% $13.44 107 trace →
10 Gemini 3.6 Flash 0.0% $6.36 195 trace →
11 Claude Opus 4.8 (xhigh) 0.0% $16.99 115 trace →
12 Claude Opus 4.7 0.0% $1.55 41 trace →
13 Claude Opus 4.6 0.0% $13.58 265 trace →
14 Claude Sonnet 4.6 0.0% $31.21 546 trace →
15 Gemini 3 Flash 0.0% $0.27 68 trace →
16 Claude Opus 4.7 (xhigh) n/a $13.94 180 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run