AI Model Cost per Task 2026: Frontier vs Open-Weight Benchmarks
The short answer: in 2026 a single AI task costs anywhere from well under a cent to a few dollars, and the spread is almost entirely model choice. Published reference points (Artificial Analysis, Aug 2026) put Muse Spark 1.2 at ~$0.40 per task and Claude Opus 5 at ~$2.34 per task — a ~6x spread on the same class of work. This page benchmarks per-task cost across the frontier and open-weight models agencies actually route to, with methodology and dated sources.
Published per-task reference points (Artificial Analysis, Aug 2026)
| Model | Reference cost per task | Source |
|---|---|---|
| Meta Muse Spark 1.2 | ~$0.40 | Artificial Analysis (Aug 2026) |
| Claude Opus 5 | ~$2.34 | Artificial Analysis (Aug 2026) |
| DeepSeek V4 Flash (pre-hike) | ~$0.03 | Artificial Analysis (Aug 2026; pre-Aug 16 rates) |
These are vendor-independent benchmark figures, not our own measurements. Note the DeepSeek row is the pre-hike reference — after the Aug 16, 2026 increase, the per-task cost for DeepSeek V4 depends heavily on peak vs off-peak and cache hits.
Current per-1M rates (Aug 16, 2026) — the inputs for your own math
| Model | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|
| DeepSeek V4-Flash (off-peak) | $0.22 | $0.66 | Peak $0.44/$1.32; cache-hit input $0.007 off-peak / $0.014 peak |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | Peak $1.32/$3.96; cache-hit input $0.022 off-peak / $0.044 peak |
| Gemini 3.7 Flash (intro) | $0.75 | $3.75 | Intro through 2026-12-31, then $1.50/$7.50; 1M context |
| Meta Muse Glimmer (hosted, Together AI) | $0.35 | $1.50 | Local self-host = electricity after hardware amortized |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 500K context; AA Intelligence Index 61; cache hit $0.50 |
| Qwen 3.8 Max (open weights) | $2.00 | $6.00 | Hosted API list price; open weights live Aug 12, 2026 |
| Claude Sonnet 5 | $2.00 | $10.00 | Permanent as of Aug 10, 2026; Sept 1 increase cancelled |
| Kimi K3 (Moonshot) | $3.00 | $15.00 | Open weights; $20M/12mo commercial-use threshold |
| GPT-5.6 Sol | $5.00 | $30.00 | Per Artificial Analysis; OpenAI has not published official per-token rates |
Cost per task on a typical 10K-in / 2K-out workload
Illustrative math (input tokens ÷ 1M × input price + output tokens ÷ 1M × output price), not a benchmark claim — your real token counts will differ:
| Model | Input cost | Output cost | Total per task |
|---|---|---|---|
| DeepSeek V4-Flash (off-peak) | $0.0022 | $0.0013 | ~$0.004 |
| Gemini 3.7 Flash (intro) | $0.0075 | $0.0075 | ~$0.015 |
| Muse Glimmer (hosted) | $0.0035 | $0.0030 | ~$0.007 |
| Grok 4.6 | $0.020 | $0.012 | ~$0.032 |
| Qwen 3.8 Max | $0.020 | $0.012 | ~$0.032 |
| Claude Sonnet 5 | $0.020 | $0.020 | ~$0.040 |
| GPT-5.6 Sol | $0.050 | $0.060 | ~$0.110 |
At 10K tasks/month, that spread is $40/mo (DeepSeek off-peak) to $1,100/mo (GPT-5.6 Sol) on identical volume. This is why model routing is a margin lever, not a footnote.
What the 2026 data says about routing
- Small structured tasks belong on cheap models. Published reference (AA): Muse Spark 1.2 ≈ $0.40/task vs Claude Opus 5 ≈ $2.34 — a 6x spread for the same class of work. Our 10K/2K math shows DeepSeek off-peak, Gemini 3.7 Flash intro, and hosted Muse Glimmer under ~2 cents per task.
- Frontier costs cluster at the top. GPT-5.6 Sol at $5/$30 (AA) is the highest per-task on the board; Grok 4.6 at $2/$6 matches Qwen 3.8 Max and undercuts GPT-5.6 Sol by ~3.4x on the illustrative task while matching its AA Intelligence Index of 61.
- DeepSeek's hike changes the "cheapest stack" answer. Post-Aug 16, DeepSeek V4-Flash off-peak ($0.22/$0.66) is still cheap on cache-miss input but output is 2.4x the old rate; peak output ($1.32) is now above Gemini 3.7 Flash intro and hosted Muse Glimmer.
- Sustained volume flips the answer to self-host. Open weights (Muse Glimmer 30B local, DeepSeek V4 MIT weights, Qwen 3.8 Max, Kimi K3) run at electricity-cost marginal inference once hardware is amortized.
How to compute your agency's real cost per task
- Instrument a pilot: log input/output tokens per task type for 100–500 real tasks.
- Multiply by your model's per-1M rate (use the table above; apply cache hits and off-peak where eligible).
- Sum per task type, then add retry/failure multipliers — real agent workflows rarely run clean the first time (levelsio reported ~$0.06 per request and "a few dollars per task" baselines, with failure modes turning into $500 loops and $2,000 overnight bills).
- Re-run quarterly — 2026 pricing moves weekly (DeepSeek Aug 16, Gemini 3.7 Flash Aug 13, Grok 4.6 Aug 12, Claude Sonnet 5 Aug 10).
Model the blended cost of your exact task mix
Open the AI Agency Pricing Calculator →Setup fees, retainers, model strategy (DeepSeek, Gemini, Grok 4.6, Qwen, Sonnet 5, local), and margin — current 2026 rates.
Frequently asked questions
How much does an AI task cost in 2026?
It depends on the model and the task size. Published reference points (Artificial Analysis, Aug 2026): Muse Spark 1.2 ≈ $0.40 per task, Claude Opus 5 ≈ $2.34 per task, DeepSeek V4 Flash ≈ $0.03 per task at pre-hike rates. A 10K-token-in / 2K-token-out task on current 2026 rates ranges from ~$0.003 (DeepSeek V4-Flash off-peak cache-miss) to ~$0.11 (GPT-5.6 Sol), with frontier models like Claude Opus 5 and GPT-5.6 Sol at the high end for heavy tasks.
Which AI model is cheapest per task in 2026?
Among hosted APIs, DeepSeek V4-Flash off-peak (now $0.22/$0.66 per 1M after the Aug 16, 2026 increase) and Gemini 3.7 Flash intro pricing ($0.75/$3.75 per 1M through 2026-12-31) are the cheapest per task for small work, with Meta Muse Spark 1.2 at ~$0.40/task as the published small-task reference. For sustained volume, self-hosted open weights (Meta Muse Glimmer local, DeepSeek V4 MIT weights) undercut every hosted API once hardware is amortized.
What is the cost per task for Grok 4.6 vs GPT-5.6 Sol?
On a 10K-token-in / 2K-token-out task, Grok 4.6 ($2/$6 per 1M) costs about $0.032 vs GPT-5.6 Sol ($5/$30 per 1M, Artificial Analysis) at about $0.11 — roughly 3.4x cheaper per task on that workload. Grok 4.6 also has a 500K context window and an AA Intelligence Index of 61, matching GPT-5.6 Sol max.
How do I calculate my agency's cost per task?
Cost per task = (input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price). Measure real token counts per task type in a pilot, then multiply by the model's per-1M rate. Use cache hits and off-peak scheduling to cut input cost, and route small structured tasks to a cheap model while keeping frontier models for complex work.
Sources
- Artificial Analysis, model pages + cost-per-task estimates (Aug 2026): artificialanalysis.ai/models
- DeepSeek official pricing page (live verified Aug 16, 2026 — new peak/off-peak rates): api-docs.deepseek.com/quick_start/pricing
- Google, "Introducing Gemini 3.7 Flash" (Aug 13, 2026): blog.google
- SpaceXAI, "Grok 4.6" (Aug 12, 2026): x.ai/news/grok-4-6
- Anthropic, "Claude Sonnet 5" + pricing docs (Aug 10, 2026): anthropic.com
- Together AI, Muse Glimmer hosted pricing (verified Aug 12, 2026): together.ai/pricing
- Hacker News / levelsio cost-per-task baseline (Aug 5, 2026): news.ycombinator.com/item?id=45931825
Accuracy note: The three per-task reference points (Muse Spark $0.40, Claude Opus 5 $2.34, DeepSeek V4 Flash $0.03) are Artificial Analysis figures from Aug 2026 and are attributed as such; the DeepSeek row predates the Aug 16, 2026 increase. The 10K/2K per-task column is our illustrative arithmetic on published per-1M rates — labeled as such, not a benchmark claim. GPT-5.6 Sol per-token pricing is an AA estimate because OpenAI has not published official per-token rates. Per-1M rates are as of Aug 16, 2026 and move frequently — re-verify before quoting. No hands-on model testing was performed for this page; all figures trace to named sources.