RECURSIV//RESEARCH
← all experiments
Experiment 001 · June 4, 2026

The real cost of finishing the job

We gave the 10 top AI models the same real coding, data, reasoning, and SQL tasks, then measured what each one finished, how good it was, and what it cost.

VerdictBest value: Gemini 3.5 Flash — finishes 100% of tasks at <$0.0001 each.
10
models tested
80
graded runs
<$0.0001
cheapest cost-to-done
101.2x
cheapest vs priciest

Full results

#ModelReliabilityCost-to-DoneQuality
1
Gemini 3.5 Flash
Google
100%
<$0.0001
97
2
GPT-4o mini
OpenAI
100%
<$0.0001
90
3
DeepSeek V4 Pro
DeepSeek
100%
<$0.0001
97
4
MiniMax M3
MiniMax
100%
$0.0001
96
5
Kimi K2.6
Moonshot
88%
<$0.0001
86
6
Gemini 3.1 Pro
Google
100%
$0.0006
97
7
Grok 4.3
xAI
100%
$0.0007
96
8
Claude Sonnet 4.6
Anthropic
88%
$0.0008
95
9
GPT-5.5
OpenAI
88%
$0.0009
93
10
Claude Opus 4.8
Anthropic
75%
$0.0033
86
ReliabilityCost-to-DoneQualityranked by value (reliability per dollar) · longer bar = better
Best value· most reliability per dollar
Gemini 3.5 Flash
finishes 100% of tasks at <$0.0001 each
Most reliable· finishes the most tasks
Gemini 3.5 Flash
100% of tasks finished · <$0.0001/task
Cheapest that works· lowest cost above 80% reliable
Gemini 3.5 Flash
<$0.0001 per task · 100% finished

Reliability = share of tasks finished across repeated runs. Cost-to-Done = real $ to finish one task.

We gave the 10 top AI models the same real coding, data, reasoning, and SQL tasks, then measured what each one finished, how good it was, and what it cost. Cost-to-Done is the tokens each run actually used, priced at published per-model rates; reliability is pass^k across runs; quality is graded by an independent judge.

Verdict

The most expensive model finished last. Claude Opus 4.8 cost about 101× the cheapest model and scored lowest of all 10. The cheapest models finished every task.

The receipts

Every graded run in this experiment. The actual task, exactly what each model returned, whether it passed, and what it cost. Filter by result and tap any run to read the answer. Proof, not a synthetic score.

80 graded runs · 75 passed · 5 failedtap a run to read the answer
How it was measured

Each model ran every task 8 times. Reliability is the share of runs that passed (pass^k); quality is graded by an independent judge model; Cost-to-Done is the tokens each run used, priced at published per-model rates. Run by the autonomous swarm on Recursiv. How it works →

Run it yourself

Every number here came from running real agentic work on Recursiv. Point the same swarm at your own tasks.

Talk to us