OpenAI's Discount Experiment: What 13.8x Usage Growth Means for Agency Pricing
Key answers at a glance
- How much did OpenAI's discount experiment grow usage? Luna token usage jumped 13.8x and Terra rose 5.6x during the Jul 27 – Aug 14, 2026 window; Sol, kept at list price as the control, rose only 1.1x.
- What are the current GPT-5.6 prices per 1M tokens? Sol $4/$20 (promo at least through Nov 21, 2026), Terra $2/$12, Luna $0.20/$1.20.
- Did users stick after the discounts expired? ~1/3 — 32% of 100K+ customers retained some Terra/Luna usage; 18% ran at or above program pace; post-window daily volume was 1.38x the program average.
- What does this mean for AI agency pricing? Token costs are elastic and volatile — quote in deliverables, model scenarios, and re-rate quarterly instead of assuming a static per-token table.
What happened: OpenAI's 19-day discount experiment
OpenAI ran large discounts on two of its three GPT-5.6 models — Terra and Luna — from July 27 to August 14, 2026, via OpenRouter's platform (a 50% discount through OpenRouter, layered on top of OpenAI list prices). The result is the cleanest AI model pricing experiment of 2026, and it has direct consequences for how agencies price token-based work. These are the OpenAI discount experiment token usage numbers worth putting next to every rate card.
Usage followed price hard. During the window, daily Terra token usage rose 5.6x and daily Luna token usage jumped 13.8x, while Sol — the flagship, kept at list price — rose only 1.1x as a control. Terra and Luna went from 0.7% to 7.8% of all OpenRouter tokens between the pre-period and the program, a gain of 7.1 share points. OpenAI's family share across all models rose from 7.1% to 12.4%, occasionally cresting 15%.
OpenRouter measured against a pre-period baseline of July 8–26 and a post-period of August 15–20. The share gain was not cannibalization: competitors gave up 5.3 points and other OpenAI models just 1.9 points — roughly three-quarters of the gain came from outside OpenAI.
The experiment layered on top of real list-price cuts. The OpenAI API price cut July 2026 (effective Jul 30): Luna −80%, Terra −20% (Sol unchanged), which made the effective discounts through OpenRouter Luna 90% and Terra 60%. Sol's own cut followed — the OpenAI API price cut August 2026 to the $4/$20 promo — completing the family repricing. For the full GPT-5.6 Sol Terra Luna price comparison, launch, post-cut, and current rates line up below (GPT-5.6 pricing per million tokens, input / output):
| Model | Launch list price | After Jul 30 cuts | Current |
|---|---|---|---|
| GPT-5.6 Sol (flagship) | $5.00 / $30.00 | $5.00 / $30.00 (unchanged) | $4.00 / $20.00 (promo at least through Nov 21, 2026) |
| GPT-5.6 Terra (mid) | $2.50 / $15.00 | $2.00 / $12.00 (−20%) | $2.00 / $12.00 |
| GPT-5.6 Luna (high-volume) | $1.00 / $6.00 | $0.20 / $1.20 (−80%) | $0.20 / $1.20 |
For a deeper per-task comparison across GPT-5.6, Claude, Gemini, and open-weight models, see our AI Model Cost per Task 2026 benchmark and the GPT-5.6 Sol API pricing breakdown.
Jevons paradox in AI: cheaper tokens, more tokens
The experiment is being read as evidence of Jevons paradox in AI pricing: when the price of a resource falls, demand rises by more than the price drop — so total consumption (and often total spend) goes up. Why did OpenAI cut API prices? The measured answer is demand expansion: cheaper tokens grew usage multiples faster than they shrank per-token revenue. OpenAI board member Greg Brockman replied to the OpenRouter analysis calling Jevons paradox "counterintuitive and inspiring."
The retention data is the part agencies should actually memorize. Nearly a third of users who tried a discounted OpenAI model during the discount window kept using it after the discounts expired. The detail behind that headline: of the 100K+ customers with Terra/Luna usage during the program, about 32% retained some usage in the subsequent days, and 18% ran at or above their program pace. Retained accounts were larger than the median user — post-program daily Terra/Luna token volume ran at 1.38x the discount-period daily average.
The caveats are honest ones: the post-period is only 6 days versus the 19-day program, August 20 may be partial, and churned or internal accounts were excluded. But the direction is consistent with what economists have documented every time compute gets cheaper: lower marginal cost → more usage → often higher total consumption.
The practical translation for agencies: the token price you quote today is not the price your client's workload will run at six months from now — and the usage multiple that follows a price cut is your exposure.
What this means for consumption-based agency pricing
If your agency runs consumption-based AI agency pricing — passing token costs through to clients, or quoting per-task rates that assume a fixed token burn — the discount experiment changes the math in three ways. If you're still working out how to quote AI token costs to clients, these are the numbers to build the rate card from.
1. Price cuts expand usage faster than they shrink your bill. At an effective 90% discount, Luna usage grew 13.8x. A client workload that cost $100 in tokens at list price costs roughly $10 at the discounted rate — but if the 13.8x pattern holds, the workload itself grows and the bill lands closer to $138 of new work being done for ~$14. Your client gets dramatically more value per dollar; you get dramatically more delivery volume at a thinner per-token margin. That is great for retainer-based relationships and risky for fixed-fee-per-task quotes written against the old burn rate.
2. Discount windows have expiry dates baked in. The Sol promo is explicitly "at least through November 21, 2026." The Terra/Luna program ended August 14 and daily volume still ran 1.38x after it. Any quote that bakes a promotional rate in as a permanent cost basis will need a re-rate clause — or a margin haircut when the window closes.
3. Per-token pricing is now a competitive weapon, not a technical detail. The share shift (0.7% → 7.8%; ~3/4 from competitors) shows price is the single fastest lever in model adoption. When you quote "how much do AI agents cost to run," the answer is a moving target — the only defensible approach is to model scenarios, not single numbers. Use the AI Agency Pricing Calculator to stress-test your margins, and the AI Agent API Cost Calculator to estimate monthly agent workload costs under different rate assumptions. Our agency pricing models explainer covers when fixed retainers beat consumption-based billing — this data makes the case for hybrid models that re-rate quarterly.
Practical recommendations for quoting token costs
Quote in deliverables, bill against current rates. Price the outcome (a task, an agent deployment, a monthly workload) and track the token cost underneath it against live pricing. Our per-task cost benchmarks give you a defensible starting point.
Model the discount window explicitly. Every promo has an end date. When you see OpenAI API discounts, add them to your pricing sheet as a line item — "current rate, valid through X" — and keep the list-price fallback in the model. If the Sol promo reverts after November 21, 2026, the change should flip a margin report, not a surprise.
Assume usage grows when price falls. Token usage growth after price cut is no longer theoretical — it is measured. If you cut a client's per-token rate, expect token usage to grow — 1.38x post-program volume on retained accounts is your baseline assumption, and 13.8x is what happens at extreme discounts. Budget for delivery capacity, not just cost per token.
Track effective price, not list price. Prompt caching (cached input ≈10% of standard), Batch API (~50% discount), and input-length tiers change real costs by 2–10x. The list price per million tokens is a starting point, never the bill. Compare like-for-like across vendors on the Claude pricing page and Codex pricing page before you pick a default model.
Re-rate on a schedule, not on a fire drill. Consumption-based pricing only works if the rate card is current. Put a quarterly re-rate on the calendar — every model vendor in 2026 is moving prices within a single quarter, and the ones that cut prices are the ones that grow usage.
Run the scenario before you sign the SOW. The difference between list price and a discount-window rate on a 100K-task monthly delivery is thousands of dollars a month — the exact line the AI Agency Pricing Calculator models.
See the discount-experiment math in your cost model
Open the AI Agency Pricing Calculator →Model current rates (Sol $4/$20 promo, Terra $2/$12, Luna $0.20/$1.20) with a re-rate clause on your pricing sheet.
Bottom line
The OpenAI discount experiment turns Jevons paradox from an econ lecture into a pricing dataset: cut token prices ~90% and usage grows 13.8x, with a third of the new users sticking around at list price afterward. For agencies, the lesson is not "tokens are cheap now." It is that token costs are elastic, volatile, and directional — and every quote, retainer, and margin model should be built to survive the next price cut. The agency that treats AI model pricing 2026 as a static table will misprice the work; the agency that models the scenarios will win the renewals.
Frequently asked questions
What was OpenAI's GPT-5.6 discount experiment?
From July 27 to August 14, 2026, OpenAI ran 50% discounts on two of its three GPT-5.6 models — Terra and Luna — through OpenRouter, layered on top of OpenAI list prices. Daily Terra token usage rose 5.6x and Luna usage jumped 13.8x during the window, while Sol, kept at list price as the control, rose only 1.1x. Terra and Luna went from 0.7% to 7.8% of all OpenRouter tokens.
What are GPT-5.6 prices per million tokens in 2026?
After the July 30, 2026 list-price cuts, GPT-5.6 Luna costs $0.20 per 1M input / $1.20 per 1M output (−80%), Terra costs $2/$12 (−20%), and Sol is on a promotional $4/$20 rate valid at least through November 21, 2026 (down from $5/$30 at launch).
How should agencies quote token costs after the discount experiment?
Quote in deliverables and bill against current rates, model discount windows as line items with list-price fallbacks, assume usage grows when price falls, track effective price (caching, Batch API) rather than list price, and re-rate on a quarterly schedule. Nearly a third of users who tried a discounted OpenAI model kept using it after the discounts expired, so baked-in promotional rates need a re-rate clause.
Sources
- OpenRouter Blog, "GPT 5.6 Discounts & Jevons Paradox" (Aug 25, 2026): openrouter.ai
- TLDR AI newsletter (Aug 28, 2026): tldr.tech
- OpenAI API Pricing (platform docs, fetched Aug 28, 2026): platform.openai.com
- OrcaRouter, "OpenAI Cuts GPT-5.6 API Prices": orcarouter.ai
- Artificial Analysis, "GPT-5.6 has landed": artificialanalysis.ai
- CloudZero, "OpenAI API pricing in 2026": cloudzero.com
- Digg, "OpenRouter Discounts Trigger 13.8x Token Usage Surge": digg.com
Accuracy note: All experiment figures (window Jul 27–Aug 14, 2026; Terra 5.6x / Luna 13.8x / Sol 1.1x; retention ~1/3; share 0.7%→7.8%; family share 7.1%→12.4%; post-volume 1.38x; effective discounts Luna 90% / Terra 60%) are quoted from OpenRouter's primary analysis and cross-verified against TLDR AI, OrcaRouter, Artificial Analysis, CloudZero, and Digg. List prices are from OpenAI's official pricing page as of Aug 28, 2026. Model pricing changes without notice — re-verify before quoting client work.