GLM-5.3-Flash Pricing: Ox Alpha Revealed at $0.15/M

Published August 26, 2026By ABD Legacy LLC
GLM 5.3 Flash pricing Ox Alpha model open source coding model 2026

What is GLM-5.3-Flash and how much does it cost?

GLM-5.3-Flash is Z.ai's newly revealed open-weight model — the anonymous "Ox Alpha" that topped OpenRouter usage for a week — and it lists at $0.15 per 1M input tokens / $0.50 per 1M output tokens (cached input $0.03), with a 50% launch discount to $0.075/$0.25 through September 9, 2026. Z.ai confirmed on August 26, 2026 that Ox Alpha was GLM-5.3-Flash: a 320-billion-parameter mixture-of-experts model with 18B active parameters, native text/image/video input, a 1M-token context window, and MIT-licensed weights on Hugging Face. Bloomberg framed the reveal as Z.ai "rivaling DeepSeek" — and at these prices, it does, on cost per token and on open-weight availability.

The mystery resolved in six days. On August 20, 2026, an anonymous model called stealth/ox-alpha appeared on OpenRouter and OpenCode — free, 1M-token context, multimodal, with tool calling. It quickly became the most-used model of the week, carrying billions of tokens from Claude Code and other agent traffic. Community forensics pointed at Z.ai (tokenizer matches, error codes, video fingerprints), but the lab stayed silent. On August 26, Z.ai confirmed it: Ox Alpha was a deliberately anonymous preview of GLM-5.3-Flash, served entirely on Chinese AI chips to gather feedback before a named release. The stealth route is retired; the model is now glm-5.3-flash on Z.ai and z-ai/glm-5.3-flash on OpenRouter, with published prices and MIT-licensed weights.

Ox Alpha was a launch test on Chinese silicon

Z.ai's release statement is direct: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips." The preview week reportedly sustained up to 100 trillion tokens per day of capacity on tens of thousands of domestically developed accelerators (commentary points to Huawei Ascend-class hardware) using Z.ai's Encode–Prefill–Decode (EPD) disaggregated serving architecture, which separates multimodal encoding, prompt prefill, and decoding into independently scheduled worker pools.

Two caveats on the 100T figure: it is a system-level capacity claim from Z.ai and its hosting partners, not a service-level guarantee or a personal quota — and it is not independently audited. What it does signal is that a Chinese lab can serve frontier-class open weights at mass scale without NVIDIA hardware. That is the supply-chain story underneath the price story.

GLM-5.3-Flash specs at a glance

SpecValue
DeveloperZ.ai (formerly Zhipu)
Preview nameOx Alpha (stealth/ox-alpha) — retired
Model size320B total parameters, 18B active (MoE)
Context window1,048,576 tokens; 131,072 max output on OpenRouter
InputsText, images, and video — first natively multimodal GLM-5-series model
WeightsHugging Face zai-org/GLM-5.3-Flash, MIT license
API routesglm-5.3-flash (Z.ai), z-ai/glm-5.3-flash (OpenRouter)
Deployment pathsSGLang, vLLM, TokenSpeed, KTransformers
Serving claimPreview week ran entirely on Chinese AI chips (system-level, not independently verified)

GLM-5.3-Flash pricing vs DeepSeek, GPT-5.6 Sol, and the open-weight class

The headline is the price. GLM-5.3-Flash is roughly one-tenth the cost of the flagship GLM-5.3 ($1.40/$4.40), and it lands below DeepSeek's post-increase rates on both axes:

ModelInput per 1MOutput per 1MNotes
GLM-5.3-Flash (Z.ai)$0.15 ($0.075 promo)$0.50 ($0.25 promo)MIT weights; cached input $0.03; promo through Sep 9, 2026
GLM-5.3 (Z.ai flagship)$1.40$4.40Same GLM-5 line, ~10x the Flash rate
DeepSeek V4 Pro (off-peak)$0.66$1.98Cache read $0.022; peak rates ~2x
DeepSeek V4 Flash (superseded)$0.14$0.28Old pre-increase rate; not current pricing
GPT-5.6 Sol (OpenAI)$4$20Promo through Nov 21, 2026; cached input $0.40
Qwen 3.8 Max (Alibaba)$2$6Open weights since Aug 12; license threshold $50M/12mo
Kimi K3 (Moonshot)$3$15Open weights; 1M context

Rates as published Aug 26, 2026. DeepSeek V4 Pro rates are DeepSeek's official off-peak rates effective Aug 16, 2026. GLM-5.3-Flash promo ends 24:00 September 9, 2026 UTC+8. Promo/cache figures from Z.ai's published pricing and Kingy's price check of Aug 26, 2026.

Run the per-task math on a typical agent workload — 10K input tokens, 2K output tokens per task:

At list price, GLM-5.3-Flash is about 4x cheaper than DeepSeek V4 Pro off-peak and 32x cheaper than GPT-5.6 Sol on this shape of workload. At the promo rate it is roughly 8x below DeepSeek V4 Pro. That is the same order of margin expansion the open-weight class delivered over the past year — compressed into a single release with a 1M-token multimodal context on top. For a 100,000-task monthly delivery, the difference between Sol and GLM-5.3-Flash at list is about $7,750/month in raw token cost ($8,000 vs $250) before retries and overhead.

Benchmarks: approaching Claude Opus 4.8 at 1/10th the cost

Z.ai's published model card reports GLM-5.3-Flash scores 63.4 on DeepSWE v1.1, 84.3 on Terminal-Bench 2.1, and 48.8 on AutomationBench, with its in-house Code Bench showing 29.0 at maximum effort versus 29.5 for Claude Opus 4.8. Z.ai positions the model as beating GLM-5.2 across its selected coding and agent benchmarks and approaching Claude Opus 4.8 overall — at roughly one-tenth the price.

These numbers are vendor-reported. Harnesses, time limits, context management, sampling settings, and judge models all affect results; Z.ai discloses several conditions (including a six-hour timeout on some agent tests and different context lengths across benchmarks) but no independent replication exists yet. The weights are now public, so third-party runs will land within days — treat the benchmarks as directional until then, exactly as we have with every model on this page. The practical agency test remains: does it complete your repository task within budget with fewer retries than the alternatives?

The Chinese-chip constraint: what it does and doesn't mean

Three practical implications for agencies building on GLM-5.3-Flash:

What this changes in the calculator

We have added GLM-5.3-Flash as a selectable model strategy in the AI Agency Pricing Calculator — modeled at list price $0.15/$0.50 with the promo window and the MIT-weights caveat flagged, positioned between the cache-aware DeepSeek V4 Pro and the plain open-weight stack. The practical takeaway for your quoting:

Price your AI work against current market benchmarks

Try the AI Agency Pricing Calculator →

Estimate setup fees, retainers, and margin in under a minute — then model per-agent workload costs with the AI Agent API Cost Calculator.

Frequently asked questions

What is the Ox Alpha model?

Ox Alpha was the anonymous "stealth" model that appeared on OpenRouter and OpenCode on August 20, 2026 with free access, a 1-million-token context window, and text/image/video input. On August 26, 2026, Z.ai confirmed Ox Alpha was its own GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with 18B active parameters, served anonymously for a week to gather feedback. The free preview has ended; the named model is now available on Z.ai and OpenRouter at published prices.

How much does GLM-5.3-Flash cost?

Z.ai lists GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output tokens (cached input $0.03), with a temporary 50% launch discount to $0.075 input / $0.015 cached / $0.25 output through September 9, 2026. That is roughly one-tenth of the flagship GLM-5.3 rate ($1.40/$4.40) and cheaper on output than DeepSeek V4 Pro's off-peak rate. The weights are free to download on Hugging Face under an MIT license, so self-hosted cost is hardware-dependent.

Is GLM-5.3-Flash open source?

Z.ai has published the GLM-5.3-Flash weights on Hugging Face under an MIT license. "Open weights" is the more precise description: you get the model weights and deployment guidance, not the full training data and pipeline. Supported deployment paths include SGLang, vLLM, TokenSpeed, and KTransformers. Note that a 320B-parameter model with 18B active is a serious deployment project — it is not a single-consumer-GPU model like Meta's 30B Muse Glimmer.

How does GLM-5.3-Flash compare to DeepSeek?

GLM-5.3-Flash is the first open-weight model this cycle to directly challenge DeepSeek's price-performance position (Bloomberg framed the reveal as "rivaling DeepSeek"). At $0.15/$0.50 list it undercuts DeepSeek V4 Pro's off-peak rate ($0.66/$1.98) on both axes, adds native text/image/video input and a 1M-token context, and ships MIT weights. Z.ai's vendor-reported benchmarks put it near Claude Opus 4.8 at roughly one-tenth the cost, but all benchmark numbers are vendor-reported until independent replication.

What does the Chinese-chip constraint mean for agencies?

Z.ai says the Ox Alpha preview week — up to 100 trillion tokens per day of capacity — ran entirely on tens of thousands of domestically developed Chinese AI accelerators using its Encode-Prefill-Decode serving architecture. For agencies, the practical effects are supply-chain independence from NVIDIA (a resilience plus), possible export-control scrutiny of the model line, and a serving reality: the Chinese-chip stack is Z.ai's, not yours. If you self-host the MIT weights, you still deploy on whatever hardware you have — the architecture's efficiency (roughly 3x less attention compute and 4.4x smaller KV cache than GLM-5.3) lowers the hardware bar but does not eliminate it.

Sources

Accuracy note: The Ox Alpha → GLM-5.3-Flash confirmation is sourced to Z.ai's own release and Bloomberg/TechCrunch reporting of Aug 26, 2026. Model specs (320B/18B, 1M context, 131K output, multimodal, MIT weights) are from Z.ai's model card and OpenRouter's named route, verified by Kingy's Aug 26 price check. All benchmark scores are Z.ai vendor-reported; no independent replication exists as of Aug 26, 2026. The 100T-tokens/day figure is a system-level capacity claim, not a verified measurement. GLM-5.3-Flash list price $0.15/$0.50 and promo $0.075/$0.25 (through Sep 9, 2026 24:00 UTC+8) per Z.ai published pricing. DeepSeek V4 Pro off-peak rates ($0.66/$1.98, cache $0.022) are DeepSeek's official rates effective Aug 16, 2026; the old V4-Flash $0.14/$0.28 rate is superseded. GLM-5.3 flagship $1.40/$4.40 per Z.ai docs. Per-task math uses 10K input / 2K output tokens; real agent workloads vary with retries, fan-out, and cache behavior — model your own volume in the calculator. Entity-list status of Z.ai's parent Zhipu (US BIS, Jan 2025) is public record; nothing here is legal advice.