LuaJIT/LuaJIT

Mirror of the LuaJIT git repository

5,518 c medium
2,967
Generated Behavioral Tests
82.4%
Best Score
GPT 5.6 Sol (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.6 Sol (xhigh) 82.4% $6.13 47 trace →
2 GPT 5.5 (xhigh) 78.0% $8.42 68 trace →
3 Claude Opus 4.8 (xhigh) 73.9% $23.18 138 trace →
4 Claude Sonnet 4.6 71.5% $31.45 520 trace →
5 GPT 5.5 (high) 65.5% $2.54 26 trace →
6 Claude Opus 4.7 (xhigh) 58.4% $8.01 142 trace →
7 GPT 5.6 Sol (medium) 57.9% $0.74 13 trace →
8 GLM-5.2 52.3% $56.68 302 trace →
9 Gemini 3.6 Flash 44.5% $3.34 143 trace →
10 GPT 5.5 33.9% $1.10 15 trace →
11 Claude Opus 4.6 29.0% $19.88 259 trace →
12 Gemini 3.5 Flash 25.6% $4.40 129 trace →
13 Gemini 3 Flash 19.4% $0.15 68 trace →
14 GPT 5 mini 14.6% $0.01 8 trace →
15 GPT 5.4 7.7% $0.18 16 trace →
16 GPT 5.4 mini 5.3% $0.02 9 trace →
17 Gemini 3.1 Pro 3.6% $1.22 122 trace →
18 Claude Opus 4.7 2.9% $0.33 18 trace →
19 Claude Haiku 4.5 1.2% $0.98 157 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run