
We ran Thinking Machines Inkling through our private coding-agent bench. Same harness as the rest of the field. Signaldesk v1: 13 planted bugs in a ~7k-line full-stack app, network-off Docker sandboxes, hidden tests at score time. Across 8 models and 553 attempts, Inkling landed 6th. > 89% resolve (58/65) > $0.104 per successful fix > 61s average attempt > $6.01 total spend That puts it under gpt-5.6-sol (100%, $0.183/fix), grok-4.5 (99%, $0.074), fable-5 (99%, $0.411), opus-4.8 (97%, $0.144), and glm-5.2 (94%, $0.019). Above gemini-3.1-pro-preview (80%) and kimi-k3 (79%). Localized bugs and security tasks were clean. All three injection/auth tasks went 5/5. Easy and moderate tiers both hit 100%. The gap is deep investigation. Expert-tier resolve dropped to 20%. On sd-013 ranking-tiebreak it scored 1/5. On sd-003 empty-digest it scored 2/5. Peers cleared sd-013 at 5/5 (kimi 4/5). Behavior matches the miss pattern. Highest turn count in the run at 20.2 average turns, 2.4× gpt-5.6-sol’s 8.3, with the lowest average output tokens in the field. Fast wall clock (2nd after grok at 46s). Sparse thinking, lots of tool loops. Route high-stakes multi-file bugs elsewhere. Keep…

Independent data on Thinking Machines' Inkling across a 13-bug repair — a reality check on frontier coding- claims.
Checking sign-in…
Loading comments…