lfos/calcurse

A text-based calendar and scheduling application

1,243 c medium
666
Generated Behavioral Tests
78.7%
Best Score
Claude Opus 4.8 (xhigh)

Hover a point for details · The line marks the Pareto frontier (best score per cost) · Click a point to see model details

19 runs
# Model Score Cost Calls
1 Claude Opus 4.8 (xhigh) 78.7% $25.73 175 trace →
2 GPT 5.5 (xhigh) 75.2% $10.93 86 trace →
3 GLM-5.2 70.6% $30.69 218 trace →
4 GPT 5.5 (high) 68.8% $4.46 35 trace →
5 GPT 5.6 Sol (medium) 62.8% $0.95 15 trace →
6 Gemini 3.6 Flash 56.2% $4.16 166 trace →
7 Gemini 3.5 Flash 56.2% $11.16 247 trace →
8 GPT 5.5 55.0% $1.85 27 trace →
9 Claude Opus 4.6 53.8% $12.74 248 trace →
10 GPT 5.6 Sol (xhigh) 52.7% $9.25 67 trace →
11 GPT 5.4 35.3% $0.54 12 trace →
12 Claude Haiku 4.5 35.0% $0.83 127 trace →
13 Claude Opus 4.7 (xhigh) 33.9% $3.30 84 trace →
14 Gemini 3 Flash 33.8% $0.43 113 trace →
15 Claude Opus 4.7 33.5% $2.88 80 trace →
16 GPT 5.4 mini 26.0% $0.05 11 trace →
17 Gemini 3.1 Pro 21.3% $1.69 84 trace →
18 GPT 5 mini 17.0% $0.01 8 trace →
19 Claude Sonnet 4.6 7.1% $31.57 528 trace →

Click a row to replay how that model rebuilt this program, or the model name to open its full run