eradman/entr

Run arbitrary commands when files change

5,551 c easy
586
Generated Behavioral Tests
95.6%
Best Score
GPT 5.6 Sol (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 GPT 5.6 Sol (xhigh) 95.6% $3.62 27 trace →
2 GPT 5.5 (xhigh) 91.8% $6.62 83 trace →
3 Claude Opus 4.8 (xhigh) 91.6% $10.41 126 trace →
4 GLM-5.2 90.9% $16.54 246 trace →
5 GPT 5.5 (high) 88.9% $1.64 23 trace →
6 Claude Sonnet 4.6 88.6% $12.14 383 trace →
7 Claude Opus 4.6 80.9% $7.88 202 trace →
8 Claude Opus 4.7 79.7% $1.89 77 trace →
9 GPT 5.6 Sol (medium) 79.2% $0.74 15 trace →
10 Gemini 3.5 Flash 79.0% $2.45 134 trace →
11 Gemini 3.6 Flash 78.0% $2.06 108 trace →
12 GPT 5.5 75.9% $0.86 19 trace →
13 Gemini 3.1 Pro 70.8% $1.98 152 trace →
14 Gemini 3 Flash 65.7% $0.25 77 trace →
15 GPT 5.4 58.5% $0.24 13 trace →
16 Claude Opus 4.7 (xhigh) 42.8% $33.60 454 trace →
17 Claude Haiku 4.5 35.5% $0.53 109 trace →
18 GPT 5.4 mini 13.3% $0.02 7 trace →
19 GPT 5 mini 1.2% $0.01 12 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run