The short answer
Three models survive the only test that matters here: nothing else in this group beats them on both intelligence and cost per finished task. Xiaomi's MiMo-V2.6-Flash is the cheapest at $0.06 per task with an Intelligence Index of 38. Anthropic's Claude Haiku 5.5 is the smartest model under a quarter per task, scoring 43 at $0.21. Meta's Muse Spark 1.3 is the smartest overall at 48, and the most expensive at $1.60 a task.
Everything else is dominated on these two axes. GPT-6 Luna matches MiMo's index of 38 but costs more per task. GPT-5.6 Luna is cheaper than Haiku yet weaker on every measure that counts. DeepSeek V4.1 Flash and Gemini 3.8 Flash both lose to Haiku on intelligence and on cost. The chart draws that frontier. The table under it shows the exact numbers.
| Model (maker, effort) | Intelligence Index | Cost per task | Cost per index point |
|---|---|---|---|
| MiMo-V2.6-Flash (Xiaomi, as listed) | 38 | $0.06 | 0.16¢ |
| GPT-6 Luna (OpenAI, max) | 38 | $0.07 | 0.18¢ |
| GPT-5.6 Luna (OpenAI, xhigh) | 35 | $0.09 | 0.26¢ |
| Claude Haiku 5.5 (Anthropic, max) | 43 | $0.21 | 0.49¢ |
| DeepSeek V4.1 Flash (DeepSeek, max) | 39 | $0.27 | 0.69¢ |
| Gemini 3.8 Flash (Google, high) | 41 | $1.24 | 3.02¢ |
| Muse Spark 1.3 (Meta, max) | 48 | $1.60 | 3.33¢ |
All seven figures come from the same snapshot: Artificial Analysis Intelligence Index v4.3.2, model pages read on October 8, 2026. Cost per task is Artificial Analysis' weighted average across the ten evaluations in the index. Cost per index point is our own division, explained below.
What "cost per finished task" actually measures
Artificial Analysis runs each model through ten evaluations, from agentic coding to long-context reasoning, and divides the total token bill by the number of tasks. That weighted average is the cost per Intelligence Index task. It captures what token prices hide: how chatty a model is, how much reasoning it needs, and how many input tokens its agentic runs consume.
The gaps are enormous. MiMo-V2.6-Flash finishes the average index task for $0.06. Muse Spark 1.3 needs $1.60, about 27 times more. Notice what the per-million-token prices would have told you: Gemini 3.8 Flash charges $0.75/$3.75 per million tokens, which looks mid-range, yet it lands at $1.24 per finished task because it works harder per task, generating 71,000 output tokens on average.
Three pricing footnotes change the picture if you build on these models. Haiku 5.5's $0.10/$0.50 rate only applies to prompts under 100K tokens; above that, Anthropic charges $0.50/$2.50, and Artificial Analysis itself warned its handling of that tiering was still provisional when we measured. Gemini 3.8 Flash runs on promotional pricing through December 31, 2026, then doubles to $1.50/$7.50 per million tokens. DeepSeek lists $0.30/$1.20 as its peak rate with off-peak hours at $0.15/$0.60. Confirm the current price on each provider's page before you budget.
The cheapest intelligence per dollar
Divide the cost per task by the Intelligence Index and you get a rough price for one unit of measured smarts. This is our calculation, not Artificial Analysis', and it comes with a real caveat: index points are not linear units of usefulness. A model at 48 is not "26% more useful" than one at 38. Still, as a value ranking it is stark.
MiMo-V2.6-Flash and GPT-6 Luna form a tier of their own at 0.16¢ and 0.18¢ per index point. GPT-5.6 Luna follows at 0.26¢. Then a gap: Haiku 5.5 at 0.49¢ and DeepSeek at 0.69¢ cost two to four times more per point of intelligence than the leaders, while buying meaningfully higher scores. Gemini 3.8 Flash (3.02¢) and Muse Spark 1.3 (3.33¢) are in a different sport: frontier-grade scores at frontier-grade prices.
The practical read: if your workload is thousands of similar tasks a day, the left side of this chart is where your margin lives. If one wrong answer costs you more than a thousand right ones save, the right side starts to look reasonable.
How often does it work the first time?
Here is the honest limitation of this whole comparison: nobody publishes a uniform, cross-model "satisfaction without iteration" rate. We will not invent one. What exists, measured independently on the same snapshot, are four single-attempt benchmarks: AutomationBench-AA (objective completion on real-world work tasks), AA-Briefcase v1.1 (rubric pass rate on knowledge work), Terminal-Bench 4.0 (agentic coding in a terminal), and AA-Omniscience accuracy (factual answers without hallucination). Our One-Shot Reliability Proxy is simply their mean.
Equal weights, no secret sauce. Every model below has all four components from the same v4.3.2 snapshot, so the ranking is apples to apples. Two of Haiku 5.5's components (AutomationBench-AA and Terminal-Bench 4.0) had to be read from chart bars and are approximate; the rest are exact page figures. And the label matters: this is a benchmark proxy for first-try success, not a measurement of user satisfaction. Your tasks are not these tasks.
| Model | AutomationBench-AA | AA-Briefcase rubric | Terminal-Bench 4.0 | Omniscience accuracy | Proxy |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 69% | 46% | 27% | 49% | 48% |
| Muse Spark 1.3 | 58% | 54% | 33% | 44% | 47% |
| Claude Haiku 5.5 | ≈58% | 54% | ≈33% | 44% | 47% |
| Gemini 3.8 Flash | 60% | 35% | 20% | 55% | 43% |
| MiMo-V2.6-Flash | 64% | 50% | 23% | 27% | 41% |
| GPT-6 Luna | 53% | 42% | 13% | 44% | 38% |
| GPT-5.6 Luna | 43% | 38% | 4% | 42% | 32% |
DeepSeek V4.1 Flash takes the proxy at 48%, carried by the best AutomationBench-AA score in the group (69%). Muse Spark 1.3 and Haiku 5.5 tie at 47%. The surprise is at the bottom: GPT-5.6 Luna manages 32%, dragged down by 4% on Terminal-Bench 4.0, and GPT-6 Luna only reaches 38%. Cheap per task does not mean reliable per attempt. That tradeoff is the whole story of the budget tier.
Speed is not responsiveness
Two numbers describe how a model feels, and they point in different directions here. Output speed is tokens per second once the model starts talking. Time to first token is how long you wait before it starts. Reasoning models think first, so the wait can dwarf the talking.
Haiku 5.5 at max effort is the extreme: 244.2 tokens per second, then a 369-second silence while it reasons. That is nearly six minutes before the first token. For a background agent grinding through documents, irrelevant. For anything interactive, disqualifying at this setting; lower effort settings answer far sooner. DeepSeek is the all-rounder: 223.1 tokens per second with a 1.2-second first token. MiMo is the odd one: the slowest output in the group at 57.5 tokens per second, but a first token in 4.56 seconds, so it feels snappy on short answers and sluggish on long ones.
One more wrinkle: speed and latency are measured on each provider's own API and move with load, routing, and your region. Treat them as indicative, not contractual.
The catch on each model
- MiMo-V2.6-Flash ($0.06/task, index 38): the cheapest finished task we have ever seen measured at this intelligence level, and open weights under the MIT license. The catch is throughput: 57.5 tokens per second is the slowest output here, and only four API providers carry it.
- GPT-6 Luna ($0.07/task, index 38): nearly as cheap as MiMo with OpenAI's distribution behind it. The catches are a 108.6-second time to first token and the weakest first-try reliability among the budget tier's leaders at 38% on our proxy.
- GPT-5.6 Luna ($0.09/task, index 35): the lowest sticker price from OpenAI in this group. The catch is capability: index 35 is the lowest here, and 4% on Terminal-Bench 4.0 means agentic coding usually needs a second attempt.
- Claude Haiku 5.5 ($0.21/task, index 43): the value frontier: the highest index under a quarter per task. The catches are Anthropic's tiering, which quintuples prices past 100K-token prompts, a new tokenizer that counts roughly 30% more tokens for the same text, and that 369-second first token at max effort.
- DeepSeek V4.1 Flash ($0.27/task, index 39): the best first-try proxy in the group at 48%, open weights under MIT, and the most responsive of the fast models. The catch is price structure: $0.30/$1.20 is the peak rate, and the index of 39 trails Haiku while costing more per task.
- Gemini 3.8 Flash ($1.24/task, index 41): strong factuality (55% Omniscience accuracy, best here) and the best DeepSWE coding score in its class. The catch is the calendar: promotional pricing ends December 31, 2026, and the list price doubles on January 1, 2027.
- Muse Spark 1.3 ($1.60/task, index 48): the smartest model in this comparison by a clear margin. The catch is the bill: $1.25/$4.25 per million tokens makes it 27 times pricier per task than MiMo. Note the fine print: an earlier xhigh configuration was measured at $0.55 per task under the older v4.3 index; the figures here are the current max configuration under v4.3.2, which is what keeps the comparison fair.
How we measured this
Every primary number on this page comes from Artificial Analysis' public model pages, read on October 8, 2026, all stating Intelligence Index v4.3.2. That version blends ten evaluations across agents, coding, general capability, and scientific reasoning. We used each model's high-reasoning configuration as listed by Artificial Analysis: max for Muse Spark 1.3, Haiku 5.5, GPT-6 Luna, and DeepSeek V4.1 Flash; high for Gemini 3.8 Flash; xhigh for GPT-5.6 Luna; and MiMo-V2.6-Flash as listed.
Cost per task, output tokens per task, speed, and time to first token are Artificial Analysis' measured figures, not vendor claims. The One-Shot Reliability Proxy and the cost-per-index-point ranking are our calculations, built only from those published figures, with the formula stated in the open. Two of Haiku 5.5's proxy components are approximate chart reads, marked with ≈.
Prices and benchmark versions move. We re-check the provider pages before publishing; you should re-check them before building. Nothing on this page is personalized advice: match the model to your own acceptance tests, not to our ranking.
Sources and data
- Artificial Analysis: MiMo-V2.6-Flash (max)
- Artificial Analysis: GPT-6 Luna (max)
- Artificial Analysis: GPT-5.6 Luna (xhigh)
- Artificial Analysis: Claude Haiku 5.5 (max)
- Artificial Analysis: DeepSeek V4.1 Flash (max)
- Artificial Analysis: Gemini 3.8 Flash (high)
- Artificial Analysis: Muse Spark 1.3 (max)
- Artificial Analysis: Intelligence Index methodology
- Xiaomi: MiMo-V2.6 announcement
- Artificial Analysis: Claude Haiku 5.5 launch analysis
Model pages were read on October 8, 2026. All state Intelligence Index v4.3.2. The proxy components and derived rankings were computed from those pages on the same date.
Frequently asked questions
Which AI model is cheapest per finished task?
MiMo-V2.6-Flash, at $0.06 per Artificial Analysis Intelligence Index task, with an index of 38. GPT-6 Luna is next at $0.07. Both are open-weights-friendly options; MiMo carries the MIT license.
Which model is the best value if I need real intelligence?
Claude Haiku 5.5: index 43 at $0.21 per task, the highest intelligence under a quarter per task in this group. Watch Anthropic's tiered pricing past 100K-token prompts.
What is the One-Shot Reliability Proxy?
Our own metric: the mean of four single-attempt benchmarks (AutomationBench-AA, AA-Briefcase rubric, Terminal-Bench 4.0, AA-Omniscience accuracy) from the same v4.3.2 snapshot. It estimates first-try success on benchmark tasks. It is not a user-satisfaction rate, and your tasks are not these tasks.
Why not just compare prices per million tokens?
Because token prices hide verbosity. Gemini 3.8 Flash looks mid-range at $0.75/$3.75 per million tokens but costs $1.24 per finished task: it generates far more tokens per task than cheaper-looking rivals. Cost per finished task is the number your bill actually follows.
Will these numbers still be true next month?
Maybe not. Providers change prices (Gemini's promo ends December 31, 2026), and Artificial Analysis reweights its index. Every figure here is dated October 8, 2026; confirm current pricing on the provider's page before you build.
Is MiMo-V2.6-Flash really usable, or just cheap?
Both, with limits. It is open weights under the MIT license with a 1M-token context window, and it leads this group on AutomationBench-AA at 64%. Its weaknesses are output speed (57.5 tokens/sec, slowest here) and factual accuracy (27% Omniscience accuracy, lowest here).
Apps mentioned in this guide
- Claude Haiku 5.5 review - High-volume, repetitive work: classification, summari…
Want the short version for a specific app? Start with our full directory of 22 reviewed money apps or see how we rank them.