The real cost of finishing the job
We gave the 10 top AI models the same real coding, data, reasoning, and SQL tasks, then measured what each one finished, how good it was, and what it cost.
Full results
Reliability = share of tasks finished across repeated runs. Cost-to-Done = real $ to finish one task.
We gave the 10 top AI models the same real coding, data, reasoning, and SQL tasks, then measured what each one finished, how good it was, and what it cost. Cost-to-Done is the tokens each run actually used, priced at published per-model rates; reliability is pass^k across runs; quality is graded by an independent judge.
The most expensive model finished last. Claude Opus 4.8 cost about 101× the cheapest model and scored lowest of all 10. The cheapest models finished every task.
The receipts
Every graded run in this experiment. The actual task, exactly what each model returned, whether it passed, and what it cost. Filter by result and tap any run to read the answer. Proof, not a synthetic score.
Each model ran every task 8 times. Reliability is the share of runs that passed (pass^k); quality is graded by an independent judge model; Cost-to-Done is the tokens each run used, priced at published per-model rates. Run by the autonomous swarm on Recursiv. How it works →