:batch slug silently falls back to standard pricing at 2x the token cost unless you act. Three routes keep the 50% batch discount: OpenRouter's beta Batch API (plain slug, 24h window, text-only), the openai/flex tier (provider: {"only": ["openai/flex"]}, synchronous, GPT-5.6 family only — not gpt-5/4o/o3), or OpenAI's direct Batch API (still 50% off, 50K requests/batch, 24h completion). Audit every model string in your config with the endpoints check before quoting. Full routing guide: OpenRouter OpenAI Batch Endpoints Down: 3 Ways to Keep the 50% Discount →
ollama run muse-glimmer — official build 18 GB / 128K context / text+image; Apple Silicon MLX build muse-glimmer:30b-mlx, 21 GB). No official Meta hosted API price has been published; the same model is available hosted on Together AI at $0.35 / $1.50 per 1M tokens (input/output, $0.04 cached) — see the Local vs API cost scenario ↓. It is not frontier-grade on the hardest reasoning, the 17 GB quant carries ~1.0% benchmark degradation, and you own uptime/ops/security.
| Service Type | Setup Fee Range | Monthly Retainer | Avg Margin | Best For |
|---|---|---|---|---|
| 💬 Chatbot / Assistant | $1,500–$5,000 | $500–$1,500/mo | 65–75% | SMBs, e-commerce, service cos |
| 📧 Email Automation | $2,000–$6,000 | $750–$2,000/mo | 60–72% | Coaches, SaaS, agencies |
| 🎯 Lead Generation Bot | $3,000–$8,000 | $1,000–$3,000/mo | 55–70% | Real estate, insurance, finance |
| ✍️ Content Automation | $2,500–$7,500 | $800–$2,500/mo | 65–80% | Content creators, media, blogs |
| 🏢 Full Office Automation | $8,000–$35,000 | $2,500–$7,500/mo | 45–65% | Mid-market, growing teams |
| ⚙️ Custom AI Agent | $5,000–$25,000 | $1,500–$5,000/mo | 50–70% | Tech cos, SaaS, operations |
| 📱 Social Media Automation | $1,500–$4,500 | $600–$1,800/mo | 70–82% | Brands, coaches, ecommerce |
| 📢 ChatGPT Ads Management (EU) | $1,500–$5,000 | 15–25% of ad spend ($2K–$7.5K/mo floor) | 55–70% | SMB advertisers targeting the EU rollout (Aug 24, 2026 — 31 markets); agency-led buying window |
* Ranges reflect 2026 US market rates. Final pricing depends on complexity, client size, and your experience level. Model strategy affects margins more than list prices: open-weight stacks (Qwen 3.8 Max at $2/$6 per 1M tokens, Kimi K3, GLM-5.2) cut the compute line vs. paid frontier APIs — and Grok 4.6 (SpaceXAI, released Aug 12, 2026) now lists at the same $2/$6 per 1M tokens as Qwen 3.8 Max with a 500k context window and an Artificial Analysis Intelligence Index of 61 (matching GPT-5.6 Sol max) — 50% below GPT-5.6 Sol's current $4/$20 per 1M (promo window Aug 21 – Nov 21, 2026; previously $5/$30) on input tokens and 70% below on output tokens (x.ai → · AA →). New Aug 28, 2026: Tencent Hy4 preview (770B MoE, 49B active, Apache 2.0, native 1M context) lists at $0.834/$2.501 per 1M tokens on OpenRouter — ~40% below GLM-5.3 on input, ~43% below on output, and ~6x below Kimi K3 on output, though still above DeepSeek V4 Pro off-peak; a preview with over-verification tendencies and ~770GB FP8 weights (not typical self-host) — model it in the calculator's Model Strategy dropdown (OpenRouter → · HF → · Full analysis →). New Aug 28, 2026: Qwen3.8-Flash (Qwen, open weights Aug 26, 2026; 125B MoE / ~6B active, ~95% sparsity, qwen-community-1.0 license) lists at $0.15/$0.47 per 1M tokens (cache read $0.016) on QwenCloud and OpenRouter — below GLM-5.3-Flash at list on output and cache, below DeepSeek V4 Pro off-peak on both axes, and the lowest cache-read rate on this page; Artificial Analysis medians $0.30/$1.20; benchmarks vendor-reported, 'Next' preview variant — model it in the calculator's Model Strategy dropdown (OpenRouter → · QwenCloud → · Full analysis →). New Aug 13, 2026: Claude Sonnet 5 (Anthropic) — $2/$10 per 1M tokens, PERMANENT (Aug 10, 2026); the previously scheduled Sept 1, 2026 $3/$15 increase is CANCELLED, no scheduled change pending — input at open-weight parity, output well below premium frontier (Anthropic → · Pricing docs →). New Aug 13, 2026: Gemini 3.7 Flash (Google) is now the cheapest verified frontier-class API on this page at intro pricing $0.75/$3.75 per 1M tokens (input/output, valid through 2026-12-31, then $1.50/$7.50; 1M context, 64K max output) — undercutting Qwen 3.8 Max and Grok 4.6 on both axes during the intro window (Google blog → · Gemini API pricing →). Qwen 3.8 Max license terms finalized (Aug 12, 2026): separate Qwen license required before commercial use only if you run a Model-as-a-Service or AI Work Assistant business AND aggregate revenue exceeds US$50M over any consecutive 12 months (no numeric % — individually negotiated; Reuters' "up to 30%" is reporting, not license text; internal use exempt) — threshold 2.5x higher than Moonshot's Kimi K3 ($20M/12mo), but with a new AI Work Assistant trigger category (coding/office-productivity agents like Qoder, QwenWork) that Kimi K3 lacks; display clause identical (100M MAU or $20M/month revenue → prominently display the model name). ⚠ DeepSeek price increase LANDED (Aug 16, 2026, 16:00 UTC) — official peak/off-peak rates are live: off-peak V4 Pro $0.66/$1.98 per 1M tokens with $0.022 cache reads (peak $1.32/$3.96/$0.044); pre-hike V4-Flash $0.14/$0.28 and V4-Pro $0.435/$0.87 are superseded. Even after the increase, V4 Pro cache reads are ~45x below Claude Fable 5's $1.00/1M — the dominant line for long agent runs (see the cache-aware scenario below; the ~276x figure from @JulianGoldieSEO was computed on launch pricing and is stale). Verify api-docs.deepseek.com/quick_start/pricing before quoting. New (Aug 11, 2026): the cheapest per-workload option is now local self-host — Meta's Muse Glimmer (30B dense, Apache 2.0, 4-bit quant under 20 GB on a single consumer GPU) runs at electricity-cost marginal inference once hardware is amortized, undercutting every hosted API at sustained volume (Local vs API cost scenario ↓). New (Aug 11, 2026): for Content Automation work with image deliverables, Microsoft's MAI-Image-2.6 (announced Aug 10, 2026) is now #2 on the Arena text-to-image leaderboard (mai-image-2.6-preview, Elo 1336 ±11) — ahead of Google, Meta, and xAI, behind only OpenAI's GPT-Image-2 — with text rendering the standout gain (+91 Elo vs 2.5) for packaging, signage, and ad-copy image work. Availability: Arena now, MAI Playground this week, Foundry API soon; Foundry pricing for 2.6 is not yet published, so treat prior-gen MAI-Image-2.5 (~$48/1k images) as the reference until rates drop. GPT-5.6 Sol (Aug 6, 2026) now powers both Instant and deep reasoning for ChatGPT Plus/Pro — one consistent model with a reasoning-effort slider; GPT-5.6 Luna is the new default for Free/Go users (unlimited text chats rolling out this week/next week). NEW Aug 22, 2026: OpenAI published official GPT-5.6 Sol API pricing Aug 21, 2026 — $4/$20 per 1M input/output tokens (cached input $0.40 per 1M), a promotional rate valid Aug 21 – Nov 21, 2026 (down from $5/$30; API + eligible ChatGPT Work/Codex credits; ChatGPT Pro/Plus/Business subscription pricing unchanged) — verify current rates at OpenAI pricing before quoting. Agent Plugins 1.0.0 (Aug 6, 2026) adds a portability lever: one plugin package serves Codex, ChatGPT, Cursor, GitHub Copilot, Kiro, and VS Code, so setup fees carry a build-once discount and plugin packaging & distribution is priced as its own deliverable (see the Portability selector and Agent Plugins section ↓). New (Aug 2026): agent payment rails like Cloudflare Wallets (announced Aug 4, 2026) add a creator-set spending cap at the wallet layer — the same rails that make the capped vs uncapped spend comparison in this calculator worth running before you quote a retainer. New Aug 26, 2026: Claude Opus 5 (Anthropic) added as a selectable model strategy — $5 per 1M input / $25 per 1M output, exactly half of Fable 5 (launched late July 2026). Per Ramp AI Index (Aug 12, 2026), Fable 5 drew ~6% of Anthropic business tokens / 11.4% of Anthropic spend while GPT-5.6 Sol holds ~25% of OpenAI tokens / 23% of OpenAI spend at roughly half of Fable 5's price; FT-cited Ramp data (via Superpower Daily/CIO) reports Opus 5 overtook Fable 5 in enterprise spend after its late-July debut (attributed to FT-relayed Ramp data — Ramp's own report does not state the overtake; WinBuzzer cautions Opus arrived too late to explain the full July pattern). Select Claude Opus 5 in the Model Strategy dropdown to model the value-tier Anthropic stack (Ramp AI Index → · Anthropic Opus →).
Agent workflows rarely run clean the first time. On Aug 5, 2026, levelsio (Pieter Levels) reported burning $500 per Gauntlet Loop run — an AI-coding method that fans out subagents and loops until "utterly perfect" — then corrected it to $900 total with 95% of generated code removed. Measured baselines are ~$0.06 per request and "a few dollars per task"; failure modes (retry storms, subagent fan-out, silent misconfiguration) turn that into $500 loops and $2,000 overnight bills. Use this estimator to model what retries actually add to your spend.
On Aug 4, 2026 Cloudflare announced Cloudflare Wallets and cloudflare.pay — the buyer side of agentic commerce. An Account Wallet (human-owned) funds Virtual Wallets (agent-owned) with guardrails: a spending cap / allowance, an approved merchant allow-list, and a maximum transaction size. Agents that hit a limit must request a manual override from a human — they cannot approve escalations themselves, and the cap is enforced at the wallet's API layer, not inside the model's system prompt, so prompt injection can't lift it. Cloudflare's own example: give every employee a $100/week budget for AI inference by provisioning one Account Wallet and Virtual Wallets per employee with that rule. Fees are undisclosed so far — this estimator defaults to 0% and lets you set your own. For agencies, the takeaway is cap-constrained planning: with wallet-layer limits, worst-case spend per agent becomes a known number (the cap), so cost forecasts get a hard ceiling instead of a tail.
On Aug 25, 2026 OpenAI launched Premium seats for ChatGPT Business (announced August 2026) — a higher-usage seat tier for teams that outgrow Standard. Premium seats cost $125 per user per month billed monthly, or $100 per user per month billed annually (a 20% annual discount); Standard Business seats remain $25 per user/month, or $20 billed annually. Premium includes 5x more usage than Standard, is not subject to the five-hour usage limit, and gets predictable weekly usage resets. A workspace needs at least 2 paid seats (any mix of Standard + Premium) and, since Aug 24, 2026, caps at 200 paid seats per subscription; larger deployments move to ChatGPT Enterprise (sales-led). Model a client's seat line here, then add it to the retainer math above.
On September 14, 2026, Anthropic is permanently raising standard weekly Claude Code limits by 25% over the pre-promotion baseline for Pro, Max, Team, and seat-based Enterprise plans — but the temporary 50% boost expires the same day, leaving paid users with a net −16.7% (≈ −17%) cut in weekly capacity vs today (baseline 100 → today 150 → Sept 14: 125) [1][4][8]. Anthropic itself confirmed the cut: "Compared to today, this works out to a 17% reduction in weekly limits on Claude Code." [1][7]. The boost holds through September 13; new limits apply September 14 [1][8]. Model your team's workload here: enter what share of today's boosted cap your agency actually uses, and see what that becomes after Sept 14 plus what it would cost to keep today's capacity. Weekly limits only — 5-hour session limits, plan pricing, and API billing are unchanged [6][7][11]. Anthropic does not publish fixed weekly message/token counts, so this estimator is ratio-based (125/150 = 0.833) [4][6].
On Aug 10, 2026 Anthropic, Macquarie Asset Management, and Singapore's GIC announced Theseus Infrastructure — a platform to develop, operate, and lease purpose-built U.S. data centers to Anthropic under long-term agreements, with Anthropic as anchor tenant and Macquarie/GIC owning the platform and funding the majority of each project's equity. Anthropic separately pledged to pay 100% of grid-upgrade costs for interconnecting its data centers through its own electricity charges. No capital commitment, capacity, or lease terms were disclosed. Why this matters to your quotes: Gartner forecasts global data center electricity up 26% in 2026 to 565 TWh, passing 1,200 TWh by 2030 where "grid supply may be insufficient" — power, not chips, is becoming the binding constraint on AI supply. Capacity is already being locked in years ahead: up to 5GW of AWS Trainium (more than $100B over 10 years) and ~3.5GW of Google TPU via Broadcom from 2027. This scenario models the supply-side risk the rest of the calculator prices away: a power-constrained per-token outlook vs your baseline rate. All adjustments are ESTIMATES — projections, not published prices.
On July 30, 2026, AWS stopped opening Bedrock Agents Classic to new customers: the model catalog is frozen, and accounts without Bedrock Agents usage in the prior 12 months get HTTP 403 AccessDeniedException on CreateAgent / InvokeInlineAgent — with no exception process [1]. There is no end-of-life date and no migration deadline [1]. Classic itself is free as a platform; AgentCore bills ~12 consumption components — runtime $0.0895/vCPU-hr + $0.00945/GB-hr (active use only), gateway $0.005/1K invocations, memory $0.25/1K short-term events / $0.75/1K records/mo, web search $7.00/1K queries, policy $0.000025/request [3]. The one first-hand migration report showed a simple agent's bill moving ~$4.87 → ~$5.02/month (+$0.15) — tokens dominate [10]. This estimator models the platform delta AND the one-time migration labor — the number that actually moves your quote.
Google Cloud's new managed model routing (API Gateway, Public Preview since Aug 3, 2026) accepts your existing OpenAI-compatible chat requests, inspects the model name in each payload, and routes the call to a cheaper foundation model — with no client-side code changes. This estimator shows the potential token-cost savings from routing simple traffic to Gemini Flash-Lite instead of paying Flash/Pro rates for everything.
On August 3, 2026 Google Cloud added managed model routing to API Gateway (Public Preview). It accepts OpenAI-compatible chat requests, transcodes them in-flight, and dispatches them to Gemini, Anthropic Claude, or OpenAI models hosted in Vertex AI Model Garden. Google positions it as a managed replacement for self-hosted proxies like LiteLLM — no proxy server to host, scale, or maintain.
Routing is driven by the model name in each request payload. You define a router with a default model plus rules mapping client model strings to cheaper backends — unmatched traffic falls back to the default. Example: send all traffic to Flash, set the default to Flash-Lite, and route only complex/agentic requests to Flash. Google's own examples use google/gemini-3.5-flash-lite, google/gemini-2.5-pro, anthropic/claude-opus-4-7, and openai/gpt-oss-120b-maas. The new Gemini 3.7 Flash (Gemini API model ID gemini-3-7-flash, released Aug 13, 2026) is the workhorse Flash option for coding/agent traffic — at intro pricing $0.75/$3.75 per 1M tokens (valid through 2026-12-31, then $1.50/$7.50), it is already cheaper than 3.5 Flash on both axes, so routing simple traffic from it to Flash-Lite still saves but the gap is smaller.
Using Google's published list prices: a content agency sending 50M input + 10M output tokens/mo to Flash at $165/mo could route 80% to Flash-Lite and drop to ~$65/mo — ≈ $100/mo (~61%) saved. A multi-tier client setup on 2.5 Pro at $212.50/mo with 70% budget-tier traffic could drop to ~$100.85/mo — ≈ $111.65/mo (~53%) saved. A 5% fallback-traffic leak onto Flash-Lite instead of Flash saves ~$16/mo on that slice alone. Token volumes and split percentages are assumptions; substitute your own usage.
- Public Preview: text-only, name-based routing to MaaS models in Model Garden; request-side streaming, gRPC, WebSockets, Gemini Live, VPC-SC, and Private Service Connect unsupported.
- One-way mode: you cannot retrofit routing onto an existing gateway or remove it — switching requires a new API config + gateway.
- Single-host constraint: all models in one router must share the same hostname (global or one regional endpoint).
- Pricing gap: no model-routing-specific fee was found in the reviewed sources; confirm your exact model versions and region before quoting a client.
- No per-request observability yet: routing decisions aren't attributed per request in logs during preview.
- API Gateway — Overview of model routing (Google Cloud docs)
- API Gateway — Configure model routing (Google Cloud docs)
- Google Developers Blog — A unified API for AI model routing
- Introducing Gemini 3.7 Flash (Google blog, Aug 13 2026)
- Gemini API pricing (intro rate $0.75/$3.75 through 2026-12-31)
- Gemini 3.7 Flash model card (DeepMind)
- Vertex AI — Generative AI pricing (Gemini token rates)
- API Gateway pricing (per-call tiers)
- Google Cloud release notes (Aug 3, 2026)
- TLDR AI — Aug 5, 2026 issue
- API Gateway quotas and limits
2026-08-30: Added the Claude Code Usage Limit Cost Impact Estimator — a new section modeling the Sept 14, 2026 weekly limit change: permanent +25% over baseline replaces the temporary +50% boost → net −16.7% (≈ −17%) vs today for Pro, Max, Team, and seat-based Enterprise (baseline 100 → today 150 → Sept 14: 125). Inputs: plan, current weekly usage as % of today's boosted cap (default 100%), seat price ($/mo, default $100), seat count (default 5). Math (research brief t_6a2f4dd0): new cap = today's cap × 125/150 = 0.8333; baseline = today's cap ÷ 1.5; workload at 100% of today's cap = 120% of the new cap (150/125 = 1.2) → same workload needs 20% more headroom; cost to keep today's capacity = seat line × 1.2. Outputs: new cap as % of today, usage after Sept 14, capacity cut, dollar cost to hold capacity. Umami event claude_code_limits_estimated fires on submit. Updated the Claude Code weekly-limit FAQ (visible + FAQPage schema) from the Aug 31 promo framing to the Sept 14 schedule and added a new "how much will it cost my agency" FAQ item + schema entry (28→30 Q). Meta description/keywords updated. New post: Claude Code Usage Limit Cost Impact. Sources: Anthropic via @ClaudeDevs (Aug 29, 2026) + BleepingComputer + MacObserver + Notebookcheck + usingClaude + AICatchup + Lapaas + Startup Fortune + Help Center promo article (verified via research brief t_6a2f4dd0).
2026-08-29: Added the AgentCore Migration Estimator — a new section modeling the Bedrock Agents Classic → AgentCore migration: inputs are agent complexity (simple/managed harness, custom orchestration, multi-agent graph), # agents, sessions/day, active runtime per session (vCPU-hr + GB-hr), tool calls per run (→ gateway invocations), memory events per run, long-term memory records, web-search queries/day, policy auth requests, model + tokens per run, and blended agency rate. Outputs: monthly AgentCore platform bill vs Classic-equivalent bill (inference only — Classic is free as a platform [1]), delta/month, one-time migration labor range (hours–1 day / 1–3 weeks / 3–6 weeks [9]), labor cost range, and the "EOL? None — this is a strategy choice" note. Rates are AWS's published AgentCore list prices [3]: runtime $0.0895/vCPU-hr + $0.00945/GB-hr (active only), gateway $0.005/1K, memory $0.25/1K events / $0.75/1K records/mo, web search $7.00/1K, policy $0.000025/request. Umami event agentcore_estimated fires on submit. FAQ item + FAQPage schema entry (27→28 Q), meta description/keywords, changelog entry. New post: Bedrock Agents Classic → AgentCore Migration Cost. Sources: AWS maintenance-mode doc [1] + AgentCore pricing [3] + DEV.to AWS Builder case study [10] + byteiota labor table [9] (verified via research brief t_bd1e2ee4).
2026-08-29: Added the Wallet allowance-cap breach preset to the Agent Failure & Retry Cost Estimator — models Cloudflare Wallets' manual-override rule: when an agent hits its creator-set spending cap, over-limit requests are blocked at the wallet API layer and routed to a human for a manual override; the approval is billable human time. Math: over-cap requests/mo × (billable minutes per override ÷ 60) × billable rate/hr, added to the retry total with its own "Manual Override Approval" result card. Also labeled the Wallet Fee input in the Agent Wallet & Spend Cap Estimator as ESTIMATE with the flag "Cloudflare has not published wallet fees; this default is an estimate." — no fee schedule exists yet (update when Cloudflare ships fees; tracked in code comment + CHANGELOG). Sources: Cloudflare blog (Aug 4, 2026) + Help Net Security (Aug 5, 2026) — verified via research brief t_7746aecf (Cloudflare agent wallets chain, kanban t_da2044f8).
2026-08-28: Added Qwen3.8-Flash as a selectable model strategy — Qwen's new open-weight multimodal MoE (open-sourced Aug 26, 2026 on Hugging Face + ModelScope; production API on QwenCloud same day): 125B total / ~6B active per token, ~95% sparsity (OrcaRouter phrasing; card says "125B with 6B activated"), 262,144-token native context → 1M via YaRN, qwen-community-1.0 license (NOT Apache-2.0). Official pricing: $0.15 per 1M input / $0.47 per 1M output (cache read $0.016); China ¥0.8/¥2.7/¥0.1. OpenRouter live at $0.15/$0.47; Artificial Analysis medians $0.30/$1.20; OrcaRouter pass-through. Modeling: MODEL_COMPUTE_FACTOR.qwen38flash = 0.86 / MODEL_MARGIN_BONUS.qwen38flash = +7 — the cheapest open-weight API at list (below GLM-5.3-Flash on output and cache, below DeepSeek V4 Pro off-peak on both axes, lowest cache-read on the page at $0.016), tied with dsv4pro's 0.86 rather than below because benchmarks are vendor-reported with no independent replication, the model is two days old, GLM-5.3-Flash's 50% promo still undercuts on raw price through Sep 9, and the qwen-community-1.0 license is a new licensing layer. Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry (24→25 Q), meta description/keywords, Pricing Reference Table footnote, open-weight section narrative + sources, changelog entry. New post: Qwen3.8-Flash-Next Cost: Open-Weight MoE at $0.15/M. Sources: QwenCloud qwen3.8-flash + HF Qwen/Qwen3.8-Flash-Next + OpenRouter qwen/qwen3.8-flash + OrcaRouter release blog + Artificial Analysis + Unite.AI + Intelligent Living + SGLang docs + HN thread (verified via kanban t_ca417b6f / t_d363cc52).
2026-08-28: Added OpenAI's GPT-5.6 discount-experiment data to the model-cost comparison pages: during Jul 27 – Aug 14, 2026, Terra and Luna were discounted 50% via OpenRouter — daily Terra token usage rose 5.6x and Luna 13.8x, while Sol at list price rose only 1.1x (control); ~1/3 of users who tried a discounted model kept using it after expiry, and Terra/Luna went from 0.7% to 7.8% of OpenRouter tokens (Jevons-paradox elasticity, per OpenRouter's Aug 25 analysis + TLDR Aug 28). List prices updated: GPT-5.6 Terra $2/$12 (−20% Jul 30, from $2.50/$15), GPT-5.6 Luna $0.20/$1.20 (−80% Jul 30, from $1/$6) per 1M tokens — added to the per-1M rate tables on AI Model Cost per Task 2026 and AI Agent API Cost Calculator, plus the Sol pricing page. Sol's $4/$20 promo (Aug 21 – Nov 21, 2026) unchanged. New post: OpenAI's Discount Experiment: What 13.8x Usage Growth Means for Agency Pricing. Sources: OpenRouter Blog (Aug 25, 2026) · TLDR AI (Aug 28, 2026) · OpenAI API pricing (platform docs, verified Aug 28, 2026) — research brief t_5c8841ca (7 sources, 18 evidence quotes). No calculator-logic change.
2026-08-28: Added a PRICING-PATH ALERT + FAQ entry: OpenRouter removed all 35 openai/*:batch model slugs on Aug 26, 2026 (plus z-ai/glm-5.2:batch and moonshotai/kimi-k2.7-code:batch) — catalogue went 422→380 models, 37 of 42 departures carried the :batch suffix. The slugs still answer metadata requests, so nothing looks broken until a completion returns 404 "No endpoints found". Base models untouched (openai/gpt-5 still routes at $1.25/$10.00 per 1M; OpenAI vendor page: 96 models, ~2.8T tokens/wk); 24 batch variants from Anthropic (11), Google (10), Thinking Machines/NVIDIA/MiniMax survive — anthropic/claude-opus-5:batch still bills at the clean 50%. No calculator-logic change (batch is a routing tier, not a model price input), but quoting guidance updated: (1) OpenRouter beta Batch API — POST /api/beta/batches, plain slug, 50% of standard per-token pricing, 24h window, text-only (rejects image/audio/file parts), returns 202 + status "validating" (do not treat 202 as completion); (2) openai/flex — provider: {"only": ["openai/flex"]}, synchronous 50% (gpt-5.6-sol $1.00/$5.00 vs $2.00/$10.00), explicit opt-in required, GPT-5.6 family only (absent on gpt-5, gpt-4o, o3, gpt-4.1); (3) OpenAI direct Batch API — still 50% vs synchronous, separate rate-limit pool, 50K requests/batch, 200MB input, 2,000 batches/hr, completes within 24h. Check any slug: curl .../api/v1/models/<model>/endpoints — empty array = dead; run it against every model string in config (42 entries vanished). FAQ item + FAQPage schema entry (23→24 Q), visible FAQ item, model-pricing NEW banner, meta description/keywords, changelog entry. New post: OpenRouter OpenAI Batch Endpoints Down: 3 Ways to Keep the 50% Discount. Sources: MoClaw investigation (Aug 28, 2026) + OpenRouter /api/v1/models live read + OpenRouter Batch API Quickstart + provider routing docs + OpenAI Batch API docs (verified via kanban t_40c4dca7).
2026-08-28: Added Tencent Hy4 preview as a selectable model strategy — Tencent's largest open-weight release yet (open-sourced Aug 28, 2026): 770B total / 49B active MoE, native 1M-token context, 64K max output, Apache 2.0 weights on Hugging Face (tencent/Hy4-preview + FP8). OpenRouter lists $0.834 per 1M input / $2.501 per 1M output (cached input $0.042), Tencent Cloud sole provider — ~40% below GLM-5.3 on input, ~43% below on output, ~6x below Kimi K3 on output, still above DeepSeek V4 Pro off-peak ($0.66/$1.98). Tencent's internal blind eval (163 experts, 203 tasks): 2.99/4 vs GLM-5.3 2.92 / Kimi K3 2.94 — tier parity, not decisive. Modeling: MODEL_COMPUTE_FACTOR.hy4preview = 0.90 / MODEL_MARGIN_BONUS.hy4preview = +5 — cheaper per-token than the plain open stack at list, kept above glm53flash's 0.88 because GLM-5.3-Flash's promo/rate still wins on raw price and DeepSeek keeps the cache floor; flagged in output. Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry (22→23 Q), meta description/keywords, Pricing Reference Table footnote, open-weight section narrative + sources, changelog entry. New post: Tencent Hy4 Preview Pricing: 770B Open Weights at $0.834/M. Sources: Tencent release + hy.tencent.ai + GitHub Tencent-Hunyuan/Hy4-preview + Hugging Face tencent/Hy4-preview + OpenRouter tencent/hy4-preview + Bloomberg (Aug 28, 2026) + KuCoin flash + explainx.ai (verified via kanban t_6e671c22).
2026-08-26: Updated the ChatGPT Ads (EU) pricing preset with OpenAI's enterprise-first trajectory (Digiday, Aug 24, 2026): OpenAI is hiring a head of ads enterprise marketing ($374K–$415K + equity) to position ChatGPT Ads for enterprise advertisers and agencies; Ads Manager tooling now targets media teams (custom audiences need a 25K minimum, account-health indicators, budget alerts, CMO-ready reporting charts, automated headline/copy suggestions) and an SMB team is in build-out — enterprise/agency-led budgets scale first, SMB self-serve later. Pricing implication: enterprise engagements can command the premium end of the setup/management ranges now; keep SMB quotes as instrumented pilots. Updated: preset assumptions box, FAQ item + FAQPage schema entry (21→22 Q), meta description, changelog entry. Source: Digiday — OpenAI is sharpening its focus on enterprise advertisers (verified via kanban t_4bc87ea6).
2026-08-26: Added GLM-5.3-Flash (Z.ai) as a selectable model strategy — the anonymous Ox Alpha stealth model (free on OpenRouter/OpenCode from Aug 20) was confirmed by Z.ai on Aug 26, 2026 as GLM-5.3-Flash: 320B/18B MoE, natively multimodal, 1M-token context, MIT weights live on Hugging Face, served entirely on Chinese AI chips during the preview. Official API pricing: $0.15 per 1M input / $0.50 per 1M output (cached input $0.03), with a 50% launch promo to $0.075/$0.25 through Sep 9, 2026 — ~1/10th of the flagship GLM-5.3 rate ($1.40/$4.40) and below DeepSeek V4 Pro off-peak ($0.66/$1.98) on both axes. Modeling: MODEL_COMPUTE_FACTOR.glm53flash = 0.88 / MODEL_MARGIN_BONUS.glm53flash = +7 — cheaper per-token than the plain deepseek stack at list, kept above dsv4pro's 0.86 because the cache-aware profile still wins long agent runs on cache reads, benchmarks are vendor-reported (DeepSWE v1.1 63.4, TB2.1 84.3, AutomationBench 48.8, Code Bench 29.0 vs Opus 4.8 29.5), and the promo rate doubles after Sep 9. Also replaced the 2 stale "GLM 5.3 unconfirmed" passages (open-weight section + FAQ) with the confirmed reveal. Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry (20→21 Q), meta description/keywords, changelog entry. New post: GLM-5.3-Flash Pricing: Ox Alpha Revealed at $0.15/M. Sources: Z.ai release + Bloomberg/TechCrunch (Aug 26, 2026) · Kingy price check (Aug 26) · Wccftech · Hugging Face zai-org/GLM-5.3-Flash (verified via kanban t_f9df0f69).
2026-08-26: Added Claude Opus 5 (Anthropic) as a selectable model strategy and added Ramp AI Index adoption/enterprise-share stats to the Fable 5, Opus 5, and GPT-5.6 Sol entries: Opus 5 — $5 per 1M input / $25 per 1M output (exactly half of Fable 5; Anthropic official, launched late July 2026), modeled MODEL_COMPUTE_FACTOR.opus5 = 1.18 / MODEL_MARGIN_BONUS.opus5 = -5 (between Sol 1.15/-3 and Fable 5 1.25/-6); Fable 5 — ~6% of Anthropic business tokens / 11.4% of Anthropic spend (July 2026); GPT-5.6 Sol — ~25% of OpenAI tokens / 23% of OpenAI spend; FT-cited Ramp data reports Opus 5 overtook Fable 5 in enterprise spend after its late-July debut (attributed to FT-relayed Ramp data via Superpower Daily/CIO — WinBuzzer/BreezyScroll do not confirm the overtake). Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry, meta description/keywords, Pricing Reference Table footnote, assumptions footer (date → Aug 26), changelog entry. Sources: Ramp AI Index Aug 2026 · WinBuzzer · BreezyScroll · Superpower Daily (FT) · Anthropic official (verified via parent research brief t_06f1cfb2).
2026-08-25: Added a Meta Hatch consumer AI agent comparison row — new Consumer AI Agent Pricing section with an 8-field comparison table (provider, product, price, billing, availability, model, category, source). Meta Hatch: $199.99/mo premium tier — REPORTED, pending confirmation (The Information Jun 4 + Aug 24, 2026, via PYMNTS/RuntimeWire; final pricing NOT set; not shipped as of Aug 25, 2026), monthly billing, launching in coming weeks (late Aug–early Sep 2026 target), Watermelon model targeted for October 2026, consumer AI agent category — compared against ChatGPT Plus ($20/mo), ChatGPT Pro ($100–$200/mo) and Claude Max (up to $200/mo). Price-update mechanism documented in an HTML comment + site CHANGELOG; analytics task t_ecfe7f64 monitors sources for official pricing. FAQ item + FAQPage schema entry, meta description/keywords, assumptions footer note, changelog entry. Full explainer: Meta Hatch AI Agent Price: What Consumer AI Agents Cost (verified via parent research t_b6570320).
2026-08-25: Added the ChatGPT Business Seat Cost Estimator — a new section modeling Standard vs Premium seats with OpenAI's verified list prices (Premium $125/user/mo monthly / $100 annual (20% discount); Standard $25/$20), billing cadence (monthly vs annual), seat count (2-seat min / 200-seat cap), and the usage-multiplier assumptions (Premium = 5x more usage than Standard, no five-hour limit, weekly resets). Example: 10 Premium seats billed annually = $1,000/month ($12,000/yr); billed monthly = $1,250/month. Premium seats are now a calculator input — the Aug 10/12 "budget separately" note is superseded. Verified via parent research brief t_642ada39 (7 sources, 34 evidence quotes; announced August 2026, official launch Aug 25, 2026). FAQ item + FAQPage schema entry, meta tags, assumptions footer (date → Aug 25), changelog entry. New explainer: ChatGPT Business Premium Seats Pricing (2026).
2026-08-23: Added a Codex 20M users / banked reset note + FAQ entry: Codex and ChatGPT Work crossed 20 million active users the week of Aug 21, 2026 (up from 15M a week earlier), and OpenAI credited every paid user a banked reset (saved usage-limit credit that resets both 5-hour and weekly windows when redeemed). The reset rollout missed its Aug 21 deadline; OpenAI set a firmer 8pm PST same-evening deadline that also passed for many users, then announced a new reset for 2026-08-24T21:00:00Z (2pm PST) after finding usage inefficiencies. No calculator-logic change and no price change — plan list prices stand (Plus $20, Pro 5x $100, Pro 20x $200, Business per seat, API usage-based); the milestone is a capacity/billing-clarity story, not a price input. FAQ item + FAQPage schema entry, meta description/keywords, changelog entry. New explainer: Codex Pricing 2026: Plans, Limits & Banked Reset for Agencies (companion: Codex Just Hit 20 Million Users — Find AI Agency; verified via research brief t_72fe5cf1).
2026-08-22: New post GPT-5.6 Sol API Pricing: OpenAI Cuts Rates to $4/$20 (Aug 21 – Nov 21, 2026) — first Sol price cut since launch, effective Aug 21 through at least Nov 21, 2026.
2026-08-21: New explainer AI Client Communication Workflow: What It Costs to Automate Client Texts — OpenAI's Apple Messages plugin (Aug 20, 2026) lets ChatGPT read, search, summarize, draft, and send iMessage/SMS/RCS on Apple silicon Macs in ChatGPT Work/Codex. The plugin itself is plan-inclusive; the real cost line items are hardware (Apple silicon per seat), Work/Codex seats priced into retainer math, and compliance review (OpenAI has not documented exactly which message content leaves the machine or published managed-Mac admin guidance). No calculator-logic change — this is a client-communication workflow cost lens, not a token-price input. Sources: OpenAI release notes (Aug 20, 2026) + OpenAI Codex plugin docs + 9to5Mac/Engadget/MacRumors/TechCrunch/TNW/Unite.AI/Yahoo (verified via research brief t_c7d9dc7e).
2026-08-22: GPT-5.6 Sol API pricing updated to the verified promo rate — OpenAI published official per-token pricing for GPT-5.6 Sol on Aug 21, 2026: $4 per 1M input / $20 per 1M output / $0.40 per 1M cached input, a promotional rate valid Aug 21 – Nov 21, 2026, down from the $5/$30 estimate the calculator previously used (API + eligible ChatGPT Work/Codex credits; ChatGPT Pro/Plus/Business subscription pricing unchanged). All Sol pricing references updated to $4/$20 with the promo window: selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entries, meta description/keywords, Pricing Reference Table footnote, assumptions footer (added Sol promo note; Grok 4.6 comparison updated from 60%/80% cheaper to 50%/70% cheaper vs $4/$20), changelog entry. No other model prices changed. Calculator modeling unchanged (MODEL_COMPUTE_FACTOR.sol = 1.15 / MODEL_MARGIN_BONUS.sol = -3 — still premium tier). Cost-per-task verification: 10K-in/2K-out task drops from ~$0.110 to ~$0.080 (-27%); 100K-in/5K-out drops ~$0.650 to ~$0.500 (-23%); 20K-in/4K-out drops ~$0.220 to ~$0.160 (-27%) — input-heavy agent workloads drop ~20%+, consistent with the 20% input / 33% output cuts. Sources: OpenAI API pricing (published Aug 21, 2026), verified via research brief t_1d150dd3 (7 sources, 16 evidence quotes).
2026-08-22: Added a ChatGPT Ads (EU) pricing preset — new Service Type option "📢 ChatGPT Ads Management (EU)" with labeled inputs (monthly ad spend, management fee %) shown when selected, plus a sourced assumptions box. OpenAI expands ChatGPT Ads to 31 European markets Aug 24, 2026 (40 total); ads run on Free + Go tiers only (Go ≈ €8/mo); buying is agency-led first (OpenAI Ads Solutions + agency/tech partners; self-serve Ads Manager later this summer). Pricing model: setup $1,500–$5,000 + management fee 15–25% of ad spend, floored at $2,000/mo and capped at $7,500/mo (ad spend is pass-through — client pays OpenAI; the fee is the agency's revenue). Margin 62% base; portability/plugin factors intentionally skipped for this managed-service line. FAQ item + FAQPage schema entry, meta description/keywords, Pricing Reference Table row, assumptions footer (date → Aug 22), changelog entry. Sources: OpenAI (Aug 18, 2026) · Search Engine Land · Dataconomy · Euronews (verified via research brief t_ff98639f).
2026-08-20: Added a Claude Code 50% weekly usage-limit boost note + FAQ entry: Anthropic extended the promo through August 31, 2026 (11:59 PM PT) — the third extension since the May 13, 2026 launch — and for the first time said it hopes to make the boost permanent while warning capacity may be tight (announced Aug 18, 2026 via @ClaudeDevs; widely reported Aug 19). Boost applies automatically to Pro, Max (5x/20x), Team, and legacy seat-based Enterprise; excludes Free and consumption-based Enterprise seats; Claude Code only (CLI, IDE, desktop, web) — 5-hour session limits, Claude chat, Claude Cowork unchanged; weekly limits "return to their standard levels" after Aug 31. No calculator-logic change — the boost is a seat-plan weekly-quota layer, and this calculator models API/token and project pricing, not seat-based subscriptions. FAQ item + FAQPage schema entry, changelog entry. Example math: Max 5x at $100/mo delivers 1.5× weekly quota through Aug 31 ≈ $150/mo equivalent at standard limits. Sources: Anthropic Help Center — Claude Code May–August 2026 weekly limits promotion · @ClaudeDevs (Aug 18, 2026) (verified via research brief t_4e969eaa).
2026-08-20: Added the OpenAI safety-monitoring overhead stress-test toggle — a visible checkbox in the Model Strategy section (OFF by default) that applies a +20% multiplier to OpenAI-based strategies (GPT-5.6 Sol, Paid frontier). Basis (verified via research brief t_b34197c5): OpenAI's official post (Aug 18, 2026) estimates its expanded chain-of-thought safety monitoring adds roughly 20% overhead to the inference compute it monitors — required for all RL training/eval involving tools for models of Sol capability or higher, plus all Astra inference with tools after the Aug 7 Critical-cyber determination; OpenAI is NOT currently billing customers for this overhead (spokesperson via The Register, Aug 19, 2026; corroborated TNW — not stated in the official post), and Anthropic says its safeguards make a similar slowdown unnecessary (Axios, Aug 19, 2026). The toggle is a what-if cost stress, not a price change: no default outputs change (OFF by default); the output note explains when it applies vs. not; FAQ item + FAQPage schema entry, meta description/keywords, assumptions footer (date → Aug 20), changelog entry. Sources: OpenAI (Aug 18, 2026) · The Register (Aug 19, 2026) · TNW (Aug 19, 2026) · Axios (Aug 19, 2026). Updated Aug 20 (publish): new explainer OpenAI Inference Overhead: The 20% That Matters for AI Agency Pricing — the agency-facing read on why this is a cost-side signal (not a price input), what's in scope (Sol-class tool-enabled RL training/eval; Astra inference with tools), and how to stress-test frontier assumptions with the toggle.
2026-08-19: Added an AI supplier risk note + FAQ entry: on Aug 18, 2026 OpenAI announced a two-week pause in RL training on its latest deployment-bound models and held its largest planned frontier RL run on hold after internal evaluations flagged its upcoming Astra model at the "Critical" cyber-capability threshold under its Preparedness Framework (core training continued; no model cancelled; no API price changes — calculator math unchanged). Guidance: include fallback AI models and review provider tooling roadmaps when evaluating AI costs/risks. FAQ item + FAQPage schema entry, assumptions footer, changelog entry. Source: OpenAI (Aug 18, 2026) (verified via research brief t_bb45d771). Updated Aug 19 (eve): supplier-risk FAQ now also notes OpenAI's committed 2027 public listing (CFO Sarah Friar, CNBC Aug 19) de-risks vendor longevity on a known timeline; new explainer OpenAI's 2027 IPO Window: Re-Baseline Your Cost Assumptions. Updated Aug 20: Anthropic now expects to match or top SpaceX's record IPO (~$75B outset / $86.2B w/ overallotment) and could file by end of August at a valuation just under $1T (Bloomberg Aug 20) — the pre-IPO window is the Claude pricing-risk zone; new explainer Anthropic's Record IPO: Stress-Test Your Claude Cost Assumptions.
2026-08-16: Added DeepSeek V4 Pro (cache-aware) and Claude Fable 5 as selectable model strategies, plus a dedicated Cache-Aware Agent Runs cost scenario (#deepseek-cache-scenario) — verified via parent research brief t_40deadfe. DeepSeek's official peak/off-peak pricing LANDED 2026-08-16 16:00 UTC (off-peak V4 Pro: cache hit $0.022/1M, miss $0.66/1M, output $1.98/1M; peak: $0.044/$1.32/$3.96 — DeepSeek pricing page + news260813), superseding the PROVISIONAL pre-hike schedule. Claude Fable 5 (Anthropic, claude-fable-5): $10/$50 per 1M with $1.00/1M cache reads — the premium comparison. Modeling: MODEL_COMPUTE_FACTOR.dsv4pro = 0.86 / MODEL_MARGIN_BONUS.dsv4pro = +7 (cache-aware V4 Pro is the cheapest cache-read stack on this page for agent runs — 92% of input bills at $0.022/1M — but kept below the plain deepseek factor because the 92% hit rate is an assumption, flagged in output); MODEL_COMPUTE_FACTOR.fable5 = 1.25 / MODEL_MARGIN_BONUS.fable5 = -6 (most expensive per-token stack on the page). The ~276x cache-read gap from @JulianGoldieSEO's Aug 16 post is arithmetically correct on DeepSeek's LAUNCH pricing ($1.00/$0.003625 = 275.9x) but STALE after the increase — current gap ~45x off-peak / ~23x peak, and the 92% hit rate is an author assumption, not a published OpenRouter statistic (both caveats shown prominently in the scenario, FAQ, and assumptions). Selector options + helpers, model labels + assumption notes + ROI text, FAQ item + FAQPage schema entry (new cache-gap Q&A, DeepSeek Q&A rewritten from 'provisional' to 'landed'), meta description/keywords, Pricing Reference Table footnote, open-weight section banner, assumptions footer (date → Aug 16, 2026), changelog entry. Sources: DeepSeek pricing · DeepSeek release note · Anthropic pricing · Julian Goldie SEO full post (verified via research brief t_40deadfe).
2026-08-13: Added Claude Sonnet 5 (Anthropic) as a selectable model strategy — $2 per 1M input tokens / $10 per 1M output tokens, PERMANENT as of Aug 10, 2026. Anthropic made the launch pricing permanent and cancelled the previously scheduled Sept 1, 2026 increase to $3/$15 — no scheduled price change is pending (verified Aug 10, 2026 announcement + official Claude pricing docs; research brief t_735f2ef6). Cache: $0.20 per 1M cached input; cache write $2.50 (5m) / $4 (1h) per 1M. Modeling kept conservative: MODEL_COMPUTE_FACTOR.sonnet5 = 0.97 and MODEL_MARGIN_BONUS.sonnet5 = +3 — input at open-weight parity (Qwen 3.8 Max $2/$6), output above open-weight ($10 vs $6) but well below premium frontier (GPT-5.6 Sol $5/$30 estimate), hosted frontier API with no open-weight/self-host upside. Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry, meta tags, Pricing Reference Table footnote, assumptions footer, changelog entry. Sources: Anthropic announcement · Claude pricing docs (verified via research brief t_735f2ef6).
2026-08-13: Added Gemini 3.7 Flash (Google) as a selectable model strategy — released Aug 13, 2026; intro pricing $0.75/$3.75 per 1M tokens (input/output, output includes thinking tokens) valid through 2026-12-31, then $1.50/$7.50 from Jan 1, 2027 (post-intro equals 3.6 Flash's launch price; intro is exactly half). 1M-token context, 64K max output, free tier, 5,000 free search requests/mo shared across Gemini 3.x, 50% batch discount. vs Gemini 3.6 Flash: FrontierCode 1.1 Main 43.6% (vs 34.4%), DeepSWE v1.1 65.3% (vs 49.0%), WebDev Arena Elo 1588 (vs 1538). Modeling kept conservative: MODEL_COMPUTE_FACTOR.gemini37 = 0.93 (slightly below open-weight 0.95) and MODEL_MARGIN_BONUS.gemini37 = +4 — intro pricing undercuts Qwen 3.8 Max / Grok 4.6 ($2/$6) on both axes, but the rate expires 2026-12-31 and post-intro output ($7.50) is above open-weight ($6), and it is a hosted frontier API with no open-weight/self-host upside. Also added Gemini 3.7 Flash to the Gemini API Cost & Model Routing Savings estimator as a routable current model (model ID gemini-3-7-flash). Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry, meta tags, Pricing Reference Table footnote, assumptions footer (date → Aug 13, 2026), changelog entry. Sources: Google blog · Gemini API pricing · Model card (verified via research brief t_7ea3f430).
2026-08-12: Qwen 3.8 Max open weights landed on Hugging Face (Qwen/Qwen3.8-2.4T-A95B — 2.4T params / 95B active — plus Qwen3.8-2.4T-A95B-FP8, Aug 12 ~10:24 UTC) with the official Qwen3.8-Max License. Replaced the monitor pending final terms alert with the finalized terms: separate Qwen license required before commercial use ONLY if you/affiliates run a Model-as-a-Service or AI Work Assistant business AND aggregate revenue exceeds US$50M over any consecutive 12 months (2.5x higher than Kimi K3's $20M/12mo); the NEW AI Work Assistant trigger (independent AI product primarily for AI-assisted coding or office productivity — e.g. Qoder, QwenWork; excludes single-purpose tools, non-coding/office assistants, and assistants that are a feature of another product) is absent from Kimi K3; MaaS = third-party inference/fine-tuning access with meaningful control over inputs/params/training data (mere relaying to third-party-hosted models excluded); internal use not exposed to third parties is exempt; no numeric % rate in the license — the separate license is individually negotiated (Reuters "up to 30%" is reporting, not license text); display clause identical to Kimi K3 (100M MAU or $20M/month revenue → prominently display the model name). Updated: alert box, selector helper, Pricing Reference Table footnote, FAQ item + FAQPage schema entry, meta tags, open-weight section narrative (weights now downloadable), assumptions footer, changelog. No calculator-logic change — the $2/$6 per 1M list price stands; the license is a threshold-based commercial-licensing layer. Sources: HF model card · official LICENSE · FP8 weights (verified live 2026-08-12).
2026-08-12: Added Grok 4.6 (SpaceXAI) as a selectable model strategy — frontier API at open-weight prices: released Aug 12, 2026; $2/$6 per 1M tokens (input/output), fast variant 2x, cache hit $0.50 per 1M input (-75%, AA); 500k context and Intelligence Index 61 (= GPT-5.6 Sol max) per Artificial Analysis (x.ai announcement silent on context); vs GPT-5.6 Sol $5/$30 (AA), 60%+ cheaper on input, 80% cheaper on output. Modeling kept conservative: MODEL_COMPUTE_FACTOR.grok46 = 1.0 (neutral vs hybrid baseline) and MODEL_MARGIN_BONUS.grok46 = +2 — the verified list price matches the open-weight tier (Qwen 3.8 Max $2/$6) but it is a hosted frontier API with no open-weight licensing/self-host upside. Selector option + helper, model label + assumption note + ROI text, FAQ item + FAQPage schema entry, meta tags, Pricing Reference Table footnote, assumptions note, changelog entry. Availability: Cursor, Grok Build, SpaceXAI API, OpenRouter, Vercel, Cloudflare (sources: x.ai/news/grok-4-6, 9to5Mac, Artificial Analysis; verified via research brief t_a87cfcc9).
2026-08-12: Added Grok Bot (SpaceXAI) as the fourth AI coding agent billing model — subscription-bundled agent access — in the cost-transparency comparison: no standalone price announced (unknown / contact sales placeholder), bundled with SuperGrok Heavy, Cursor Ultra, Cursor Teams Premium; always-on agents run 24/7 on their own cloud computer; desktop (macOS) + iOS; enterprise waitlist; no usage caps disclosed. Agent-option metadata block, FAQ item + FAQPage schema entry, meta tags, and assumptions note updated. No calculator logic change — Grok Bot is not usage-metered, so it is not a calculator input (sources: x.ai/news/introducing-grok-bot, MacRumors, Oflight; verified via research brief t_e203f6d1).
2026-08-12: Fleshed out the Local / self-hosted (Meta Muse Glimmer) option with the official Ollama runtime and a same-model hosted benchmark: ollama run muse-glimmer (official 18 GB build / 128K context / text+image; Apple Silicon MLX build muse-glimmer:30b-mlx, 21 GB), ollama launch integrations (Claude Code, Codex, Pi, OpenCode, GitHub Copilot, OpenClaw, Hermes Agent). Meta has published no official hosted API price, so the scenario now benchmarks against Together AI hosting the same 30B at $0.35 / $1.50 per 1M tokens ($0.04 cached input) — verified on Together's pricing and model pages (2026-08-12) — making the local-vs-API break-even concrete for the same model, not just vs other vendors. Selector label, helper text, verified-specs table (new "Local runtime (Ollama)" row), cost-scenario narrative, FAQ answer, source list (+2), and assumptions date (Aug 12) all updated.
2026-08-11: Added Local / self-hosted (Meta Muse Glimmer) as a selectable model strategy — the cheapest compute factor on this page (0.85× vs open 0.95×, margin +8) — reflecting that the cheapest-model answer is no longer just "which API": Meta's Muse Glimmer (released Aug 10, 2026; 30B dense, Apache 2.0, 131,072+ context, official 4-bit K-Quant-17GB GGUF = 16,756,681,056 bytes / ~15.6 GiB / under 20 GB, single consumer GPU at >200 tok/s on RTX 5090 with DFlash) runs at electricity-cost marginal inference once a 24 GB-class card is amortized. New dedicated Local vs API cost scenario section (#muse-glimmer-local) with verified specs table, break-even framing vs hosted APIs (DeepSeek V4 PROVISIONAL $0.14/$0.28, Qwen 3.8 Max $2/$6, Kimi K3 $3/$15), honest caveats (1.0% quant degradation, 24–32 GB envelope for full context, not frontier on HLE/GPQA, no audio, you own ops/security), and source links (Meta model card, GGUF repo, HF blog, NVIDIA blog, TechStartups, AMD, explainx). Model Strategy selector helper, open-weight section narrative, Pricing Reference Table footnote, FAQ item + FAQPage schema entry, meta tags, and assumptions date (Aug 11) all updated. Per-workload framing — "cheapest per workload," not blanket "cheapest model."
2026-08-08: Added DeepSeek V4 (V4-Flash / V4-Pro) as a selectable model strategy with a PROVISIONAL pricing flag — DeepSeek announced (Aug 6, 2026, footnote 2 of its official pricing page) that it plans to raise overall API pricing "in the near future, with a significant increase expected," and has NOT disclosed new rates, a percentage, or an effective date as of 2026-08-08. Current list prices (V4-Flash $0.14/$0.28, V4-Pro $0.435/$0.87 per 1M tokens) are treated as pre-hike/provisional: the strategy is deliberately given a smaller discount than the pure open-weight stack so the calculator does not over-state "DeepSeek is cheapest," and the output carries an explicit verify-before-quoting note. New prominent red warning banner in the Open-Weight section, FAQ item + FAQPage schema entry, meta tags extended, sources linked (DeepSeek pricing page live 2026-08-08, TNW, SCMP, Dataconomy/Bloomberg, TechNode). Open-weight self-host escape hatch noted (MIT weights).
2026-08-08: Added a monitor pending final terms alert for Qwen 3.8 Max: Alibaba (Reuters, Aug 7) plans a revenue-share requirement for qualifying model-as-a-service use (rates TBD, effective date TBD) — the cost model no longer presents Qwen 3.8 Max as free/unrestricted at scale. Kimi K3 precedent verified at >$20M annual sales for a commercial agreement (up to 30% revenue share). New FAQ item + FAQPage schema entry; meta tags extended. Source: Reuters Aug 7, 2026.
2026-08-06: Added Qwen 3.8 Max to the open-weight model strategy (GA Aug 2–3, 2026; 2.4T-param MoE, ~95B active; $2/$6 per 1M tokens; 1M context; open weights promised ~Aug 10). Real-world usage section added with the Aug 6, 2026 45-project field report — self-reported, no artifacts; the "destroyed Fable 5" claim was walked back by the tester (Qwen strong on fast/multimodal builds; Fable 5 on huge long-running projects). Transparent cost-variability note added: API list prices are public, but self-hosted cost is hardware/quantization-dependent and not yet knowable until weights drop.
2026-08-06: Added GPT-5.6 Sol as a selectable model strategy (frontier, Instant + deep reasoning) with a clearly marked estimate and link to OpenAI's pricing page — OpenAI announced Sol now powers both Instant and deep reasoning for Plus/Pro, and GPT-5.6 Luna becomes the default for Free/Go users (unlimited text chats rolling out this week/next week). No official per-token API pricing published for either model; assumptions date/source now included in calculator output.
2026-08-05: Added Gemini API cost & model routing savings estimator and explainer (Google Cloud managed model routing, Public Preview Aug 3, 2026). Pricing sourced from Google's published Vertex AI / API Gateway list prices; scenario figures are illustrative (directional).
✓ Qwen 3.8 Max license terms FINALIZED — open weights live on Hugging Face (Aug 12, 2026)
Qwen 3.8 Max's official license requires a separate Qwen license before commercial use ONLY if you (or an affiliate) run a "Model as a Service" or "AI Work Assistant" business AND aggregate revenue exceeds US$50,000,000 over any consecutive 12 months. There is no numeric % rate in the license — the separate license is individually negotiated (Reuters' "up to 30%" is reporting, not license text). "AI Work Assistant" = an independent AI product primarily for AI-assisted coding or office productivity (e.g. Qoder, QwenWork); it excludes single-purpose tools, non-coding/office assistants, and assistants that are merely a feature of another product. MaaS = giving third parties access to inference/fine-tuning (API or hosted endpoint) with meaningful control over inputs, parameters, or training data; merely relaying to third-party-hosted models is excluded. Internal use not exposed to third parties is exempt. A separate display clause (identical to Kimi K3) requires prominently showing the model name on any commercial product/service UI above 100M monthly active users or $20M/month revenue. For most agencies — below the $50M/12mo bar and not running a public MaaS or AI Work Assistant product — the $2/$6 per 1M token list price and self-hosted math stand unchanged.
Sources: HF model card — Qwen3.8-2.4T-A95B · official Qwen3.8-Max License (verbatim) · FP8 weights · Reuters — Alibaba revenue-share plan (Aug 7, 2026, background)
⚠ DeepSeek price increase LANDED — official peak/off-peak rates effective 2026-08-16 16:00 UTC
DeepSeek's announced price increase is now in effect. Official rates (DeepSeek pricing page, live 2026-08-16): off-peak V4 Pro cache hit $0.022 / cache miss $0.66 / output $1.98 per 1M tokens; peak $0.044 / $1.32 / $3.96. The pre-hike schedule (V4-Flash $0.14/$0.28, V4-Pro $0.435/$0.87 per 1M tokens) is superseded. Even at the new rates, DeepSeek V4 Pro remains the lowest-cost cache-read API on this page for long agent runs — cache reads at $0.022/1M off-peak vs Claude Fable 5's $1.00/1M (~45x off-peak gap; the widely-shared "276x" figure was correct on DeepSeek's launch pricing, $1.00/$0.003625, before the increase). The calculator's DeepSeek V4 strategy now reflects the landed rates and a new DeepSeek V4 Pro (cache-aware) strategy models agent-run economics at a 92% hit-rate assumption (see the cache-aware scenario ↓). Re-verify api-docs.deepseek.com/quick_start/pricing before quoting. Self-hosting remains an escape hatch: V4 weights are MIT-licensed open weights (1.6T Pro / 284B Flash), so agencies with sustained workloads can serve them directly via vLLM-class tooling instead of paying API rates.
Sources: DeepSeek API pricing (live 2026-08-16) · DeepSeek release note — news260813 (price change 16:00 UTC Aug 16) · Anthropic — Claude Fable 5 pricing ($1.00 cache read) · @JulianGoldieSEO — the 276x cache-read gap (Aug 16, 2026) · Julian Goldie SEO — DeepSeek V4 Pro vs Fable 5 vs Grok 4.6 (full post) · TNW — DeepSeek warns of a 'significant' price rise (Aug 6, 2026, background) · SCMP (Aug 6, 2026, background)
Open-weight models are now a real alternative to paid frontier APIs. Alibaba's Qwen 3.8 Max (GA Aug 2–3, 2026; 2.4T-parameter MoE, ~95B active, 1M-token context) prices at $2 per 1M input tokens and $6 per 1M output tokens — the cheapest open frontier-class API on this page — with open weights live on Hugging Face since Aug 12, 2026 (Qwen/Qwen3.8-2.4T-A95B, plus an FP8 variant; official Qwen3.8-Max License — see the finalized-terms alert above). Moonshot's Kimi K3 — a 2.8T-parameter open-weight mixture-of-experts model (~104B active, 1M-token context, weights live on Hugging Face since July 27, 2026) — prices at $3 per 1M input tokens and $15 per 1M output tokens, a fraction of flagship paid APIs, while scoring within a few points of Claude Fable 5 and GPT-5.6 Sol on vendor-run coding benchmarks. Zhipu's GLM-5.2 (open weights, MIT license, 1M-token context) is the strongest open-source coding model on Terminal-Bench 2.1. NEW Aug 26, 2026: the anonymous Ox Alpha model that topped OpenRouter usage was confirmed as GLM-5.3-Flash — 320B/18B MoE, natively multimodal, 1M context, MIT weights live on Hugging Face, and official API pricing at $0.15/$0.50 per 1M tokens (50% promo $0.075/$0.25 through Sep 9, 2026) — roughly one-tenth of the flagship GLM-5.3 rate and below DeepSeek V4 Pro off-peak on both axes. NEW Aug 28, 2026: Tencent Hy4 preview — 770B total / 49B active MoE, native 1M-token context, Apache 2.0 weights live on Hugging Face — lists at $0.834/$2.501 per 1M tokens on OpenRouter (cached input $0.042), roughly 40% below GLM-5.3 on input and 43% on output, and ~6x below Kimi K3 on output, though still above DeepSeek V4 Pro off-peak ($0.66/$1.98); Tencent's internal blind eval scored it 2.99/4 vs GLM-5.3 2.92 / Kimi K3 2.94 (tier parity, not a decisive win), and the preview status + ~770GB FP8 weights (not self-hostable on typical infra) are the caveats. NEW Aug 28, 2026: Qwen3.8-Flash — Qwen's new open-weight MoE (125B total / ~6B active, multimodal, ~95% sparsity, 262K native context → 1M via YaRN, qwen-community-1.0 license) — lists at $0.15/$0.47 per 1M tokens (cache read $0.016) on QwenCloud and OpenRouter, below GLM-5.3-Flash at list on output and cache and below DeepSeek V4 Pro off-peak on both axes, with the lowest cache-read rate on this page; benchmarks are vendor-reported so far and it is a "Next" preview of the Qwen4 architecture. And now the cheapest option of all is not an API: Meta's Muse Glimmer (Aug 10, 2026; 30B dense, Apache 2.0, 4-bit quant under 20 GB) self-hosts on a single consumer GPU at electricity-cost marginal inference — see the Local vs API cost scenario ↓ for the break-even framing.
Real-world agency usage so far: Qwen 3.8 Max's headline marketing claim — "autonomous coding over 10+ days" — is an official claim, not yet independently replicated, and its Fable 5-beating ranking has been disputed by independent benchmark testing. The most-cited hands-on test so far (Aug 6, 2026, an agency-community builder with ~172K followers) reports building 45 real projects while ignoring benchmarks — 3D racing games, RPGs, websites, a full OS, a promo video, and autonomous workflows — with mixed results: some demos were poor, several builds looked better than Fable 5, and every project took only a few hours. The same tester's follow-up explicitly walked back the "destroyed Fable 5" framing: Qwen 3.8 Max shone on fast, multimodal, image-guided builds (including a single-file premium landing page), while Claude Fable 5 stayed stronger on huge, long-running projects that need consistency across massive contexts. Treat these as first-person, self-reported results with no linked artifacts or independent replication — useful as a delivery-speed datapoint, not as a benchmark.
Cost variability for open-weight models: published API list prices (Qwen 3.8 Max $2/$6, Kimi K3 $3/$15 per 1M tokens) are real, but total cost depends heavily on how you run the model. Qwen 3.8 Max's weights are now downloadable (live on Hugging Face since Aug 12, 2026), so self-hosting cost varies with hardware (DGX Spark-class vs cloud GPUs), quantization (MXPF4 vs full precision), context length, and utilization — plus the license layer: the Qwen3.8-Max License triggers a separate Qwen license only above $50M/12mo aggregate revenue for MaaS or AI Work Assistant businesses (see the alert above). Agencies that self-host trade a variable per-token bill for fixed hardware cost — the crossover point depends on your monthly token volume. Until you benchmark your own workloads, treat self-hosted cost as a range, not a fixed number.
What this means for agencies: model strategy is now a pricing lever. The calculator's Model Strategy selector reflects it — open-weight stacks trim the compute line (and lift margins ~5 pts), a local Muse Glimmer self-host trims it further (the cheapest factor on this page, ~0.85×), and frontier-only stacks carry a premium. Keep workflows model-portable across at least two providers, benchmark on your own workloads (vendor tables are not your client's workload), and treat AI spend as a managed line item, not a fixed cost. Read the full analysis of what open-weight models mean for agency margins →
Sources: Alibaba — Qwen 3.8 Max blog · QwenCloud — Qwen 3.8 Max pricing · HF model card — Qwen3.8-2.4T-A95B (open weights, live Aug 12, 2026) · official Qwen3.8-Max License · HF — Qwen3.8-2.4T-A95B-FP8 · 45-project field report (X, Aug 6 2026) · Moonshot — Kimi K3 blog · Kimi K3 API pricing · HF model card — moonshotai/Kimi-K3 · zai-org/GLM-5 · Reuters — Alibaba Qwen revenue-share plan (Aug 7, 2026) · Tencent — Introducing Hy4 preview (Aug 28, 2026) · HF model card — tencent/Hy4-preview (Apache 2.0) · OpenRouter — Tencent Hy4 preview pricing ($0.834/$2.501) · Tencent release announcement · Bloomberg — Tencent Hy4 coverage (Aug 28, 2026) · QwenCloud — Qwen3.8-Flash official model & pricing page (verified Aug 27, 2026 — $0.15/$0.47/$0.016 per 1M) · HF model card — Qwen/Qwen3.8-Flash-Next (125B total / 6B active, multimodal MoE, qwen-community-1.0; released Aug 26, 2026) · OpenRouter — qwen/qwen3.8-flash ($0.15/$0.47) · OrcaRouter — Qwen 3.8 Flash release (sparsity/architecture) · Artificial Analysis — Qwen3.8-Flash-Next (medians $0.30/$1.20) · Unite.AI — Qwen3.8-Flash-Next previews Qwen4 architecture (Aug 26, 2026) · Intelligent Living — Qwen3.8-Flash matches DeepSeek V4 Pro · SGLang docs — Qwen3.8-Flash-Next cookbook · HN thread — Qwen3.8-Flash-Next (Aug 26, 2026)
The cheapest-model answer is no longer just "which API." Meta released Muse Glimmer on August 10, 2026 (Meta Superintelligence Lab; distilled from the closed frontier model Muse Spark) — a 30B dense, Apache 2.0 open-weight model whose official 4-bit quant fits under 20 GB and runs on a single consumer GPU at >200 tok/s (RTX 5090: 233.4 tok/s with DFlash speculative decoding vs 74.9 baseline, a 3.1× speedup). For an agency running sustained agentic, eval, or coding volume, the marginal cost of inference collapses to electricity once the one-time GPU capex is amortized — a category that previously meant either a weak small model or expensive on-prem infra.
| Muse Glimmer — verified specs | Value |
|---|---|
| Parameters | 30B dense (~29.6B total: 28B text decoder + ~1.8B ViT-G/14 perception encoder) |
| License | Apache 2.0 — weights + all artifacts (BF16, 4-bit quants, DFlash drafter, encoder) |
| Context | 131,072+ (NVIDIA: "120K+") |
| Quantized size | Under 20 GB — official 4-bit K-Quant-17GB GGUF file = 16,756,681,056 bytes (~15.6 GiB) |
| Hardware | Single consumer GPU — K-Quant-17GB targets 24 GB VRAM; K-Quant-Dynamic targets 32 GB |
| Speed | >200 tok/s on RTX 5090 (74.9 baseline → 233.4 with DFlash); M4 Max 23.7→37.8; M5 Max 26.6→50.2; AMD AI Max+ ~24 tok/s; Radeon AI PRO R9700 ~53 tok/s |
| Local runtime (Ollama) | ollama run muse-glimmer — official 18 GB build / 128K context / text+image; Apple Silicon MLX build muse-glimmer:30b-mlx (21 GB) with DFlash + image input; ollama launch integrations include Claude Code, Codex, Pi, OpenCode, GitHub Copilot, OpenClaw, and Hermes Agent |
| Intended uses | Local agents, function calling, local coding, LLM-as-a-judge evaluation; multimodal (text+image in, text out); works with OpenClaw and Hermes Agent scaffolds |
Local self-host vs hosted APIs — the cost scenario
For an agency running sustained agentic / eval / coding volume, the one-time capex of a 24 GB-class consumer GPU amortizes against per-token API spend. Above a token-volume break-even, local marginal cost ≈ electricity — undercutting every hosted API on this page (DeepSeek V4 Pro at $0.022/1M off-peak cache reads is the lowest cache-read rate for agent runs; Qwen 3.8 Max $2/$6, Kimi K3 $3/$15, and premium frontier APIs are higher) — including the hosted version of Muse Glimmer itself: Meta has published no official API price, but Together AI hosts the same 30B at $0.35 / $1.50 per 1M tokens ($0.04 cached input), so a local box beats even the same model's API once your volume clears the hardware break-even.
Why local wins on volume: LLM-as-a-judge / evaluation is an explicitly named intended use — the highest-volume, lowest-margin-per-token workload agencies run. Offloading evals, batch coding, and background agents to a local Muse Glimmer box cuts the token bill without touching client-facing frontier calls. Privacy and data-residency become a feature: the model card and NVIDIA both cite private data processing and credential handling as design points, so security-conscious clients (HIPAA/SOC2-adjacent, NDA-bound) can get "AI that never leaves the building" at 30B capability.
Keep it honest — the caveats: K-Quant-17GB carries ~1.0% benchmark degradation vs full precision; full 131K context plus vision encoder plus DFlash drafter needs the 24–32 GB envelope, not a cheap 8 GB card; it is not frontier-grade on the hardest reasoning (HLE 22.0, GPQA 83.5 — roughly size-class parity); there is no audio modality; and local means you own uptime, ops, and security patching. The open-weight market also just got more competitive in the same week (NVIDIA's Nemotron 3.5 Lightning 30B MoE, Qwen3.6-27B, Gemma4-31B) — that is downward pressure on what "cheap" means broadly, both locally and in API pricing. So: "cheapest per workload," not "cheapest model."
Compare on the calculator: pick Local / self-hosted (Meta Muse Glimmer) in the Model Strategy selector to see the cheapest compute factor on this page (0.85× vs open-weight 0.95×, with a +8 margin lift) — then verify against your own token volume before quoting a local-only stack.
Sources: Meta official model card — Muse-Glimmer-30B · Meta GGUF repo (K-Quant-17GB file listing) · Hugging Face blog — Muse Glimmer · NVIDIA blog — Local AI open source models & agents · TechStartups — Meta launches Muse Glimmer (Aug 10, 2026) · AMD blog — Run Muse Glimmer 30B on AMD · Ollama library — muse-glimmer (official 18GB / 128K build) · Together AI pricing — Muse Glimmer 30B $0.35/$1.50
The number nobody's talking about is 276. On Aug 16, 2026, Julian Goldie pointed out that long-running AI agents re-read their instructions on every step — hundreds of times per task — so the price of a cache read dominates the bill, not the headline input price. DeepSeek's cache reads cost ~276x less than Claude Fable 5's on DeepSeek's launch pricing ($1.00 vs $0.003625 per 1M tokens). ⚠ That 276x figure is now stale: DeepSeek's official peak/off-peak price increase landed the same day (16:00 UTC Aug 16), so the current cache-read gap is ~45x off-peak ($1.00 vs $0.022) and ~23x peak ($0.044). Even at the new rates, the cache-aware gap is the biggest single cost lever in this calculator — model it below.
The scenario: a long agent run that re-reads a fixed context (system prompt + instructions + accumulated work) before every step. At a 92% cache hit rate — Julian's stated assumption for agent workloads (directionally plausible because the same context is re-read every step, but not a verified OpenRouter statistic) — 92% of input tokens bill at the cache-read rate, only 8% at the cache-miss rate. Output tokens bill at the full output rate. The default inputs below model a real heavy-agent workload; change them to match your client's.
AI coding agent pricing is not the per-seat number on the pricing page. That's the core finding of TrueFoundry's August 6, 2026 guide, and it's why this calculator treats model strategy, delivery risk, and retries as explicit cost levers rather than hidden line items. Coding agents bill through four structures — and the same team can see very different invoices under each:
| Billing model | Examples | What agencies should know |
|---|---|---|
| Flat per-seat | Cursor Pro, Windsurf Pro, Claude Pro | Fixed monthly fee per developer — predictable, but overage charges can appear when limits are exceeded, and limits often aren't published clearly. |
| Seat + credits | GitHub Copilot Pro, Pro+, Max | Lower base fee plus a monthly credit pool. A frontier model can burn credits up to 8× faster than a standard model on the same task. |
| Pay-per-token API | Claude Code (API mode), OpenAI Codex | No per-seat charge; billing follows token consumption. Under sustained high-volume agent workloads it can become the most expensive option, and teams routinely misbudget it without usage visibility. |
| Subscription-bundled agent access | Grok Bot (SpaceXAI) — SuperGrok Heavy, Cursor Ultra, Cursor Teams Premium | No per-token meter and no standalone price: always-on agents run 24/7 on their own cloud computer and are bundled into eligible subscription tiers. Beta, desktop (macOS) and iOS; enterprise waitlist. Zero surprise-bill risk by construction — the counterpoint to metered-billing blowups. Standalone economics not priced yet; authentication mechanics undisclosed. |
Grok Bot (SpaceXAI) — agent option metadata: Monthly pricing: unknown / contact sales — no standalone price has been announced; in beta and included with eligible subscriptions (SuperGrok Heavy, Cursor Ultra, Cursor Teams Premium); enterprise access via waitlist. Supported platforms: desktop (macOS app) and iOS. Always-on / cloud execution: yes — bots share a computer of their own in the cloud, so jobs keep running 24/7 even when the laptop is closed. Usage caps: not disclosed — no per-token meter and no published tier quotas as of Aug 12, 2026. Estimate assumption: because Grok Bot has no usage-metered price, it cannot be cost-modeled per-token; budget the eligible subscription cost as the line item (see AI Coding Agent Pricing in 2026 →).
Sources: SpaceXAI (x.ai) — Introducing Grok Bot (Aug 11, 2026) · MacRumors — Grok Bot Brings Always-On AI Agents to macOS and iOS (Aug 11, 2026) · Oflight — Grok Bot Explained
Six cost variables predict spend better than the headline price: usage volume (light autocomplete vs. all-day agent workflows can differ 20–50× in token consumption), model selection, billing model, team size, billing cycle (annual discounts often 15–20%), and plan tier fit. Free and entry tiers often carry quotas production teams exceed within weeks.
ROI note for agency pricing: if you resell or deliver AI-assisted work, model the raw-token line and the retry/failure line separately — "a task that costs one credit unit on a standard model might cost eight on a frontier model." This calculator's Model Strategy and Delivery Risk selectors exist exactly for that reason. Don't quote the sticker price; quote the consumption profile.
Source: TrueFoundry — AI Coding Agent Pricing: How to Choose the Right Plan (Aug 6, 2026). Full breakdown with the six-step budgeting checklist: AI Coding Agent Pricing in 2026 (findaiagency.com) →
Meta is preparing to launch Hatch, its consumer AI agent that takes multi-step actions (shopping, email, calendar, booking) across apps like DoorDash, Etsy, Reddit, Yelp and Outlook, with launch targeted for late August or early September 2026 and a new model codenamed Watermelon targeted for October. The premium tier is priced up to $199.99/month — REPORTED, not official: The Information reports Meta has considered it, final pricing is not set, and Hatch has not shipped as of Aug 25, 2026. For agencies the $200 anchor matters because it is the consumer price point clients will compare your retainer against — not the same number as your per-agent infrastructure cost. Full explainer: Meta Hatch AI Agent Price: What Consumer AI Agents Cost →
| Provider | Product | Price | Billing | Availability | Model | Category | Source |
|---|---|---|---|---|---|---|---|
| Meta | Hatch (reported) Full breakdown → |
$199.99/mo — pending confirmation | Monthly | Launching in coming weeks (late Aug – early Sep 2026 target) | Watermelon (targeted October 2026) | Consumer AI agent | The Information via RuntimeWire → |
| OpenAI | ChatGPT Plus | $20/mo | Monthly | Available now | GPT-5.6 Sol / Luna | Consumer AI assistant | openai.com → |
| OpenAI | ChatGPT Pro | $100–$200/mo (5× / 20× usage tiers) | Monthly | Available now | GPT-5.6 Sol | Consumer AI assistant | openai.com → |
| Anthropic | Claude Max | Up to $200/mo | Monthly | Available now | Claude (Fable 5-class) | Consumer AI assistant | anthropic.com → |
* Meta Hatch price is REPORTED, pending confirmation — The Information (Jun 4 + Aug 24, 2026, via PYMNTS and RuntimeWire) reports Meta has considered a premium tier up to $199.99/month for Hatch ("Hatch Plus", 5–10× daily capacity of the free tier). Final pricing is NOT set, and Hatch has not shipped as of Aug 25, 2026. ChatGPT / Claude prices are current list prices as of Aug 25, 2026.
Agency takeaway: a consumer subscription is not an agency infrastructure cost — model the workload, not the sticker price. See Meta Hatch AI Agent Price: What Consumer AI Agents Cost for the full analysis, or run the AI Agent API Cost Calculator for per-agent workload math.
On August 6, 2026, an open, vendor-neutral specification called Agent Plugins 1.0.0 was published for packaging Agent Skills and MCP servers into portable plugins — announced by OpenAI's developer account alongside AWS, Cursor, GitHub, and Vercel. A plugin is a directory: a minimal plugin.json manifest, a skills/ folder, an optional mcp.json, and reverse-domain client-extension namespaces. Six clients support the format at launch: Codex, ChatGPT, Cursor, GitHub Copilot, Kiro (AWS), and VS Code — with AWS's Agent Toolkit compatible (30+ curated skills) and Google's Agents CLI / Data Agent Kit already shipping it. The Technical Steering Committee comprises Amazon, Cursor, Microsoft, OpenAI, and Vercel; Google is joining as a Core Maintainer. The spec is v1.0.0 but labeled a Working Draft, and it deliberately leaves out installation, distribution, permissions, sandboxing, and trust/provenance — all flagged for future versions.
What changes for your pricing: the old cost model assumed a fixed per-platform adaptation cost — the same skill repackaged, re-configured, and re-maintained for every client platform. With one portable plugin, that per-client reimplementation cost collapses. The calculator's Portability / Agent Plugins selector models the shift three ways:
- Build-once amortization (setup discount). One package serves 2–3 compatible clients at ~15% off setup, or 4+ clients / resellable at ~30% off — instead of charging full adaptation per platform. The brief's suggested range is a ~60–90% discount off the adaptation slice per additional client; these factors apply the discount to the whole setup midpoint conservatively.
- Shared maintenance stream (retainer factor). One package + optional per-client extensions replaces N drifting forks, so retainers scale slightly down (0.97× for 2–3 clients, 0.94× for 4+).
- New revenue line (plugin packaging & distribution). Packaging the deliverable as a branded plugin bundle — and distributing it across the compatible-client market — is a priced deliverable, shown in its own result card ($750 for a 2–3 client bundle, $1,500 for a resellable plugin). Resellable plugin lines also lift margins ~3 pts.
One caution before you ship plugins: Agent Plugins v1.0.0 has no permission model, no sandboxing, no provenance verification, and no secrets handling. Agencies selling plugins should self-impose signing, code review, and least-privilege practices — and agencies buying plugins must vet them. That trust layer is itself a sellable compliance service, and the calculator's quote notes it whenever a plugin option is selected.
Client portability cuts both ways: lower switching costs are great for client trust, but they weaken "sticky" platform-based retainers — lock-in risk moves from the package format to the marketplace/install layer. Price the build once, but keep the retainer tied to ongoing value delivered, not to platform lock-in.
Sources: Google Developers Blog — Agent Plugins (Aug 6, 2026) · @OpenAIDevs announcement (X) · Agent Plugins Specification 1.0.0 · Compatible Clients · AWS Open Source Blog · Vercel blog · Read the full agency guide on Find AI Agency →
On August 4, 2026, Cloudflare announced Cloudflare Wallets and cloudflare.pay during Agents Week — the buyer side of agentic commerce, pairing with the Monetization Gateway (seller side, waitlist opened July 1, 2026). The service gives AI agents a human-readable wallet handle for paying APIs, MCP tools, content, and AI inference within limits set by the wallet's creator. Handle reservations opened on announcement day (Aug 4–5, 2026); full wallet access — onramping/offramping funds, issuing Virtual Wallets, live x402 purchases — arrives "in the coming months" with fees still undisclosed.
- Account Wallets — owned by humans/organizations using Cloudflare. Hold stablecoins, can add/remove funds, and delegate spending authority down to Virtual Wallets.
- Virtual Wallets — owned by AI agents, operate via API keys. An agent's maximum spend is capped by the limit set by the Account Wallet owner.
- Identity is optional: a Cloudflare account gets a unique web address / handle (e.g.
research.example.cloudflare.pay) that works as a stable ID when interacting with merchants — like DNS for agent identity, built on the agent's existing cryptographic keypair (Web Bot Auth).
- Spending cap / allowance (periodic) — the dominant cost variable. Cloudflare's worked example: "$100 per week budget for AI inference" per employee, or a $10 exploration budget for a cheap-to-try agent.
- Approved merchant allow-list — agents can only pay listed merchants.
- Maximum transaction size — per-purchase ceiling.
- Manual override required — agents that hit a limit request a human-approved escalation; they cannot approve it themselves. Planned anomaly controls notify admins of unusual spending (e.g. unexpectedly fast spending) for review.
- Prompt-injection immunity: caps are enforced at the wallet's API layer, not inside the model's system prompt — so a cap cannot be overridden by content the agent reads. This is the "outside-the-model enforcement standard" and the reason wallet rails matter for agencies running fleets of agents.
- Payments run on x402 — an open protocol that attaches payment instructions to HTTP requests (server returns HTTP 402 "Payment Required" with a price manifest; agent pays and retries with proof of payment). Originally introduced by Coinbase in 2025; Cloudflare co-founded the x402 Foundation.
- Micropayment economics: stablecoins (USDC on Coinbase's Base chain used as the primary example) settle in seconds at ~$0.0001 fees, making sub-cent per-call API billing viable vs credit-card interchange (1.5–3.5% + fixed fee).
- Fees, supported stablecoins/blockchains, custody partner, and onramp provider are all undisclosed — the calculator's wallet fee input defaults to 0% for exactly this reason. Treat any fee figure as an assumption until Cloudflare publishes pricing.
- Cost forecasting shifts from prediction to cap-constrained planning. Worst-case spend per agent becomes a known number (the cap) — the 2026 cost-blowup record (prompt-caching misses, retry loops, effort scaling) shows the risk is tail spend, not average spend, and wallet caps cut the tail.
- Client billing caps become programmable and auditable. Provision a Virtual Wallet per client project with an explicit cap, approved vendors, and transaction-size limits; over-limit requests route to a human for a documented override — an audit trail that doubles as a billing dispute shield.
- Reserve handles now. Handles are free to reserve today (first-come identity, like early DNS names); pricing comes later. Agencies reselling AI inference should reserve client-facing handles and watch the pricing announcement.
- Use the estimator above: enter agents × per-agent monthly spend, set a wallet fee if your rail has one, and set the creator-set cap to see capped vs uncapped spend — including the no-wallet-fee default and the exceeded-cap edge case.
Sources: Cloudflare Blog — The programmable wallet for the agentic Internet (Aug 4, 2026) · Cloudflare press release (Aug 4, 2026) · Help Net Security (Aug 5, 2026) · crypto.news (Aug 5, 2026) · TechTimes (Aug 4, 2026) — verified via parent research brief t_6e67321b.
provider: {"only": ["openai/flex"]} — synchronous, exactly half of standard, but GPT-5.6 family only (not gpt-5, gpt-4o, o3 or gpt-4.1); (3) OpenAI's own direct Batch API — still 50% off synchronous, separate rate-limit pool, 50K requests/batch, completes within 24h. Check any slug with curl https://openrouter.ai/api/v1/models/<model>/endpoints — an empty array means the ID is dead. Full routing guide: OpenRouter OpenAI Batch Endpoints Down: 3 Ways to Keep the 50% Discount.
ollama run muse-glimmer (18 GB build / 128K context / text+image; Apple Silicon MLX muse-glimmer:30b-mlx, 21 GB). Once the one-time GPU capex is amortized, marginal inference cost collapses to electricity — undercutting every hosted API on this page (DeepSeek V4 Pro at $0.022/1M off-peak cache reads is the lowest cache-read rate for agent runs; Qwen 3.8 Max $2/$6; Kimi K3 $3/$15) — including the hosted version of the same model: Meta has no official API price, but Together AI hosts Muse Glimmer at $0.35 / $1.50 per 1M tokens ($0.04 cached). Caveats: the 17 GB quant carries ~1.0% benchmark degradation, full 131K context + vision encoder + DFlash drafter needs the 24–32 GB envelope, it is not frontier-grade on the hardest reasoning (HLE 22.0, GPQA 83.5 — size-class parity), there is no audio modality, and you own uptime, ops, and security patching. In short: "cheapest per workload," not a blanket "cheapest model." See the Local vs API cost scenario ↓ for the full comparison and sources.