Overall score
Overall score — higher is better.
API-equivalent cost
API-equivalent cost per full run — lower is better.
Wall time
Wall time per full run — lower is better.
codex · gpt-5.6-sol · low leads at 89.5, claude · haiku · low is 14x cheaper at 77.6; grok · grok-4.6 did not run.
| Agent | Overall | UX/UI | Frontend | Backend | Planning Audit | Bug Fix | Taste | Signature | Tokens in | Tokens out | Turns | Wall time | API-equiv. cost | Score / $ | Telemetry |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 01codex · gpt-5.6-sol · low | 89.5 | 100.0 | 68.3 | — | 79.0 | 100.0 | 100.0 | 92.5 | 3.9M | 74.1K | 7 | 28m 31s | $18.45 | 4.8 | partial |
| 02claude · haiku · low | 77.6 | — | 58.3 | 77.0 | 79.0 | 100.0 | 73.8 | 57.5 | 1.3K | 93.9K | 173 | 16m 55s | $1.31 | 59.1 | full |
| 03grok · grok-4.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | none |
Score vs. cost
API-equivalent cost per full run set against overall score, log scale. The shaded corner is cheap and good; the dashed line is the Pareto frontier — agents no other agent beats on both axes at once.
claudecodexPareto frontier
Changelog
- Sep 6, 2026v2026.09-smoke published — 3 agents, 7 tasks.
- Sep 6, 2026Task added: Inventory service (backend).
- Sep 6, 2026Task added: Fix date range overlap bug (bugfix).
- Sep 6, 2026Task added: Issue board filters (frontend).
- Sep 6, 2026Task added: Codebase audit, planted defects (planning-audit).
- Sep 6, 2026Task added: shipshit.dev landing hero (signature).
- Sep 6, 2026Task added: Landing page, five themes (taste).
- Sep 6, 2026Task added: Shipcut pricing page (ux-ui).