bench

v2026.09-smoke

codex · gpt-5.6-sol · low

89.5overallpartial
Post on X

Overall rank

#1 of 2

Cost efficiency rank

#2 of 2

Wall time rank

#2 of 2

Telemetry

partial

codex · gpt-5.6-sol · low leads this release at 89.5 overall. At 18.45 API-equivalent it costs 14.1x claude · haiku · low, the cheapest scored agent this release.

Tokens by class

Segment width is dollar share, not token share — where the API-equivalent cost actually went.

Input 3.9M $15.53Cache read 3.6M $1.44Output 74.1K $1.48

Tokens in

3.9M

Tokens out

74.1K

Turns

7

Wall time

28m 31s

API-equiv. cost

$18.45

Score / $

4.8

Category scores

UX/UI

100.0 (1 tasks)

Frontend

68.3 (1 tasks)

Backend

(1 tasks)

Planning Audit

79.0 (1 tasks)

Bug Fix

100.0 (1 tasks)

Taste

100.0 (1 tasks)

Signature

92.5 (1 tasks)

Per task

TaskGate passObjectiveSubjectiveScoreRuns
Issue board filters100%
55.6 (55.655.6, n=1)
87.5 (87.587.5, n=1)
68.3 (68.368.3, n=1)
shipshit.dev landing hero100%
92.5 (92.592.5, n=1)
92.5 (92.592.5, n=1)
Codebase audit, planted defects100%
65.0 (65.065.0, n=1)
100.0 (100.0100.0, n=1)
79.0 (79.079.0, n=1)
Inventory service0%
Shipcut pricing page100%
100.0 (100.0100.0, n=1)
100.0 (100.0100.0, n=1)
Landing page, five themes100%
100.0 (100.0100.0, n=1)
100.0 (100.0100.0, n=1)
Fix date range overlap bug100%
100.0 (100.0100.0, n=1)
100.0 (100.0100.0, n=1)
compare vs. grok · grok-4.6compare vs. claude · haiku · low