rust-ethereum/ethabi

Encode and decode smart contract invocations

525 rs medium
997
Generated Behavioral Tests
94.5%
Best Score
Claude Opus 4.8 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 Claude Opus 4.8 (xhigh) 94.5% $24.18 175 trace →
2 Claude Opus 4.7 (xhigh) 90.9% $14.23 225 trace →
3 GPT 5.6 Sol (xhigh) 89.9% $2.05 18 trace →
4 GLM-5.2 89.9% $14.89 163 trace →
5 Claude Opus 4.7 88.7% $13.93 246 trace →
6 Gemini 3.5 Flash 87.9% $7.09 234 trace →
7 Claude Sonnet 4.6 87.6% $12.55 407 trace →
8 GPT 5.5 (xhigh) 86.3% $6.65 73 trace →
9 Gemini 3.6 Flash 85.8% $4.08 147 trace →
10 GPT 5.5 (high) 83.9% $4.62 44 trace →
11 Claude Opus 4.6 81.6% $10.26 257 trace →
12 GPT 5.6 Sol (medium) 77.7% $0.71 12 trace →
13 GPT 5.5 72.2% $0.98 15 trace →
14 Claude Haiku 4.5 50.9% $1.25 159 trace →
15 GPT 5.4 48.2% $0.32 12 trace →
16 GPT 5 mini 29.7% $0.03 25 trace →
17 Gemini 3.1 Pro 7.8% $2.23 160 trace →
18 Gemini 3 Flash 7.8% $0.23 89 trace →
19 GPT 5.4 mini 0.0% $0.03 14 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run