The AI Stack Cost Report · August 2026 · Part 1: The Frontier Briefing. The meter moved again this month. This is StackSpend's monthly read on what the modern AI engineering stack actually costs. Also in this issue: the pricing map, the coding stack, and the image stack.
Use this when you need a current, factual read on where large language models stand in August 2026 — what shipped in July, which prices moved, and whether the model behind your product is still the right one.
The fast answer: July 2026 was the month the frontier got cheaper. Anthropic shipped Claude Opus 5 on July 24 at $5/$25 per 1M tokens — unchanged from Opus 4.8, but landing within 0.5% of Claude Fable 5's peak CursorBench score at half the cost per task. OpenAI took GPT-5.6 generally available on July 9, then on July 30 cut Luna by 80% (to $0.20/$1.20) and Terra by 20% (to $2/$12). Add Grok 4.5 at $2/$6, Meta's return with a paid API, and open-weight models at 2.8T parameters, and the picture flips from July's story: the constraint on frontier AI is no longer access or price. It's whether your team notices the price moved.
Last month's briefing was about a government inserting itself into the release path of frontier models. That story resolved — export controls lifted on June 30, Claude Fable 5 returned globally on July 1, and GPT-5.6 went generally available on July 9. What replaced it is a straightforward price war, and it is moving faster than most teams' model-selection reviews.
If you want the pricing side in depth, pair this with the LLM model pricing guide for August 2026 and the live LLM API Pricing Index, which tracks per-1M-token rates daily.
Quick answer
For an August 2026 snapshot:
- Claude Opus 5 (July 24) is the month's headline: near-Fable-5 capability at $5/$25, a 1M-token context window, and a five-level effort dial that lets you trade intelligence against token spend per request.
- GPT-5.6 is now GA (July 9) across Sol, Terra, and Luna — and got materially cheaper on July 30: Luna dropped 80% to $0.20/$1.20 and Terra 20% to $2/$12, while Sol held at $5/$30.
- Claude Sonnet 5 (July 1) is on introductory pricing of $2/$10 through August 31, 2026, moving to $3/$15 on September 1 — a scheduled increase worth putting in your calendar.
- The field filled in: xAI's Grok 4.5 ($2/$6, trained jointly with Cursor), Meta's Muse Spark 1.1 with its first paid developer API ($1.25/$4.25), Google's Gemini 3.6 Flash ($1.50/$7.50), and open-weight releases at frontier scale (Moonshot Kimi K3 at 2.8T parameters, Thinking Machines' Apache-2.0 Inkling).
- Access controls are no longer the binding constraint. Claude Mythos 5 remains limited-availability, but the June export-control episode is over and the frontier is broadly purchasable again.
Claude Opus 5: the same price, roughly double the model
On July 24, 2026, Anthropic released Claude Opus 5 at $5 input / $25 output per 1M tokens — identical to Opus 4.8, which it replaces at the top of the generally-available Opus line.
The results Anthropic published are unusual because they are framed in cost per task rather than benchmark points alone:
- On CursorBench 3.2 at maximum effort, Opus 5 lands within 0.5% of Claude Fable 5's peak score — at half the cost per task.
- On OSWorld 2.0, it surpasses Fable 5's best result at just over a third of the cost.
- On Frontier-Bench v0.1, it more than doubles Opus 4.8's score at a lower cost per task.
- On ARC-AGI 3, Anthropic reports a score three times the next-best model.
It ships with a 1M-token context window, 128K max output, and a five-level effort setting (low / medium / high / xhigh / max) so a single model can serve both cheap high-volume calls and expensive long-horizon agent runs. Anthropic notes one explicit gap: Opus 5 remains behind Mythos 5 on cybersecurity and exploit-development tasks.
For teams, the practical read is blunt. If you are paying $10/$50 for Fable 5 on agentic coding work, Opus 5 is the same money-for-quality trade at half the input price and half the output price — and the effort dial gives you a second lever you did not have in June. That is a model-selection decision worth making deliberately rather than drifting into.
Watch the tokenizer, not just the rate card. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text (Anthropic pricing docs). A like-for-like per-token comparison against an older model understates the newer model's cost per request. This is exactly the kind of gap that only shows up in the bill.
GPT-5.6: general availability, then an 80% cut
OpenAI's GPT-5.6 family — Sol (flagship), Terra (balanced), Luna (fast/cheap) — moved from the ~20-org limited preview described in last month's briefing to general availability on July 9, 2026.
Three weeks later, on July 30, OpenAI cut prices on two of the three tiers:
- Luna: −80% — from $1.00/$6.00 to $0.20 input / $1.20 output per 1M tokens
- Terra: −20% — from $2.50/$15.00 to $2.00 input / $12.00 output
- Sol: unchanged at $5.00/$30.00
OpenAI attributed the cuts to serving-efficiency gains made during GPT-5.6's own development — roughly a 20% reduction in end-to-end serving cost and a 15%+ improvement in token-generation efficiency (CNBC, VentureBeat). Coverage consistently frames the move as a response to cost-sensitive customers and to competition from cheaper Chinese and open-weight models — the same pressure visible everywhere else in this issue.
An 80% cut on a high-volume tier is not a rounding error. If you route classification, extraction, or summarisation traffic through Luna, your unit economics changed on July 30 whether or not anyone on the team noticed.
New and repriced models at a glance (August 2026)
| Model | Vendor | Input ($/1M) | Output ($/1M) | Status / notes |
|---|---|---|---|---|
| Claude Mythos 5 | Anthropic | $10.00 | $50.00 | Limited availability (Project Glasswing) |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | Restored globally July 1; 1M context |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | New July 24. 1M context, 5-level effort dial |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | Introductory through Aug 31; $3/$15 from Sep 1 |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | GA July 9; price unchanged |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | Cut 20% on July 30 (was $2.50/$15) |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Cut 80% on July 30 (was $1/$6) |
| Gemini 3.6 Flash | $1.50 | $7.50 | New July 21; 1M context, 64K output | |
| Grok 4.5 | xAI | $2.00 | $6.00 | New July 8; 500K context, cached input $0.50 |
| Muse Spark 1.1 | Meta | $1.25 | $4.25 | New July 9; Meta's first paid developer API |
Rates are per 1M tokens as published by each vendor and checked on 2 August 2026. Vendor pricing moves without notice — the LLM API Pricing Index re-syncs daily, and the monthly model changelog logs releases and deprecations as they land.
The rest of the field
July was one of the densest release months on record. The pattern is consistent: capability spread sideways, and price fell.
- xAI Grok 4.5 (July 8) is a 1.5T-parameter MoE trained jointly with Cursor, at $2/$6 with a 500K context window and $0.50 cached input. It reports 83.3% on Terminal-Bench 2.1 — genuinely competitive coding performance at a mid-tier price.
- Meta returned with Muse Spark 1.1 (July 9) and, more significantly, its first paid developer API at $1.25/$4.25 with a 1M-token context. Meta selling inference rather than only publishing weights is a structural change to the market.
- Google Gemini 3.6 Flash (July 21) at $1.50/$7.50 continues the high-volume workhorse line, with roughly 17% fewer output tokens than 3.5 Flash for the same work — a real cost reduction that never appears on a rate card.
- Open weights reached frontier scale. Moonshot's Kimi K3 (2.8T total / 104B active parameters, 1M context) released full weights on July 27. Thinking Machines shipped Inkling under Apache 2.0 on July 15, followed by Inkling-Small at $0.30/$1.20 on July 30. Meituan's LongCat-2.0, trained entirely on Chinese ASICs, is quoted around $0.038 per 1M tokens. See closed vs open AI models in 2026.
The spread between "cheapest good enough" and "newest best" is now roughly 150x on input (LongCat-2.0 at ~$0.04 vs Fable 5 at $10) — wider than at any point we have tracked. Which end of that range your traffic sits on is now a bigger cost decision than any infrastructure choice you will make this quarter.
What last month's access story did next
For continuity with July's briefing: the US government access episode has, for now, closed.
- The June 12 Commerce order that suspended foreign-national access to Claude Fable 5 and Mythos 5 was lifted on June 30, and Fable 5 returned globally on July 1.
- GPT-5.6, gated to ~20 partner organisations at the government's request in late June, reached general availability on July 9.
- Claude Mythos 5 remains limited-availability to approved customers under Project Glasswing — a commercial gate, not a regulatory one.
The June 2, 2026 executive order's pre-release review remains voluntary; no licensing regime was created. The durable lesson from the episode stands: model availability can change on a timescale shorter than your release cycle, so the model behind a production feature should be a configuration value, not a hard-coded assumption.
This section reports publicly documented government actions and is not legal advice; consult counsel for how export controls apply to your organization.
What this means for teams (and budgets)
A cheaper frontier is good news that creates a specific, unglamorous problem: savings do not apply themselves.
- Your rate card changed under you, twice. An OpenAI price cut on July 30 and a new Anthropic tier on July 24 both landed mid-month. Neither required any action from you — and neither will show up anywhere except your invoice unless you are watching spend by model.
- A scheduled increase is coming. Claude Sonnet 5's introductory $2/$10 ends August 31, 2026; September 1 traffic bills at $3/$15. That is a 50% increase on a tier many teams have quietly standardised on. See how to forecast AI costs in production.
- The premium tier is now optional in more places. Opus 5 at $5/$25 doing Fable-5-class work at half the price is the clearest downgrade-without-regret opportunity we have seen this year — but only if someone measures the quality delta on your workload, not on a leaderboard.
- Effort dials and tokenizers make per-token comparison unreliable. Opus 5's five effort levels and Anthropic's ~30% more verbose tokenizer both mean the rate card no longer predicts request cost. Only observed spend does.
That is the whole argument for instrumenting AI spend at the model level. StackSpend gives you LLM cost monitoring with per-model breakdown, daily model recommendations that flag when a cheaper model would do the same job with projected savings, cost attribution by model, feature, and team, and anomaly alerts the day a spike happens. Setup is read-only and takes about five minutes.
FAQ
What is Claude Opus 5 and how much does it cost?
Claude Opus 5, released July 24, 2026, is Anthropic's newest generally-available Opus-tier model, priced at $5 input / $25 output per 1M tokens — the same as Opus 4.8. It has a 1M-token context window, 128K max output, and a five-level effort setting. Anthropic reports it lands within 0.5% of Claude Fable 5's peak CursorBench 3.2 score at half the cost per task, and surpasses Fable 5 on OSWorld 2.0 at just over a third of the cost.
Did GPT-5.6 get cheaper?
Yes. On July 30, 2026 OpenAI cut GPT-5.6 Luna by 80%, from $1.00/$6.00 to $0.20 input / $1.20 output per 1M tokens, and GPT-5.6 Terra by 20%, from $2.50/$15.00 to $2.00/$12.00. GPT-5.6 Sol was unchanged at $5.00/$30.00. OpenAI attributed the cuts to serving-efficiency gains during GPT-5.6's development.
When does Claude Sonnet 5 introductory pricing end?
Claude Sonnet 5's introductory rate of $2 input / $10 output per 1M tokens runs through August 31, 2026. From September 1, 2026 it moves to standard pricing of $3 input / $15 output — a 50% increase. Batch pricing moves from $1/$5 to $1.50/$7.50 on the same date.
Which LLM should I use in August 2026?
There is no single winner — teams route by task, latency, and cost. For agentic coding and long-horizon work, Claude Opus 5 offers the best capability-per-dollar in the premium tier; GPT-5.6 Sol and Claude Fable 5 sit above it for the hardest problems. For high-volume production traffic, GPT-5.6 Luna at $0.20/$1.20 and Gemini 3.6 Flash are now very hard to beat on price. For budget or self-hosted workloads, open-weight models (Kimi K3, Inkling, GLM-5, DeepSeek V4, Qwen) remain 5–50x cheaper per token. Validate on your own eval set, then track spend by model — see how to choose an LLM for your workload.
Are frontier models still restricted by the US government?
Not currently for the models covered here. The June 12, 2026 Commerce order that suspended foreign-national access to Claude Fable 5 and Mythos 5 was lifted on June 30, and Fable 5 was restored globally on July 1. GPT-5.6 reached general availability on July 9. Claude Mythos 5 remains in limited availability to approved customers, which is a commercial restriction rather than a regulatory one. The June 2, 2026 executive order created a voluntary pre-release review, not a licensing regime.
How do I tell whether a price cut actually reduced my bill?
Compare spend per model before and after the change date, not the rate card. A price cut can be offset by higher token volume, a more verbose tokenizer, longer contexts, or a higher effort setting. StackSpend breaks AI spend down by provider and model with daily granularity, so a July 30 price change is visible as a step in the trend line rather than a surprise on the invoice — see AI cost monitoring and model cost attribution.
About The AI Stack Cost Report
The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs — frontier and open-weight LLMs, AI coding tools, image generation, and the hosting and infrastructure they run on. Each issue is compiled from primary pricing pages and release reporting, with source-check dates listed at the top of every article. The August 2026 issue has four parts: the frontier briefing, the pricing map, the coding stack, and the image stack.
Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily, the model changelog logs every release and deprecation by month, and the model glossary tracks each provider's family history.
A monthly snapshot tells you what the market charges; it can't tell you what your stack spent yesterday. StackSpend does — per-provider AI spend, anomaly alerts the day a spike happens, and a daily report in Slack. Setup is read-only, takes about five minutes, and starts with a free 14-day trial. Why we publish it: the meter came back.
Related reading
- LLM Model Pricing in August 2026: Every Major API and Open Model
- Best AI Coding Models (August 2026)
- The Latest in LLMs, July 2026 — the previous issue
- Closed vs Open AI Models in 2026: A Practical, Balanced Guide
- How to Choose an LLM for Your Workload
- LLM API Pricing Index — live per-1M-token rates, updated daily
References
- Anthropic — Introducing Claude Opus 5
- Anthropic — Pricing (Claude Platform docs)
- Anthropic — Introducing Claude Sonnet 5
- Anthropic — Redeploying Claude Fable 5
- OpenAI — Advancing the price-performance frontier with GPT-5.6
- CNBC — OpenAI cuts prices for two of its GPT-5.6 AI models
- VentureBeat — AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80%
- ThursdAI — July 2026 AI releases
- OpenRouter — Grok 4.5 pricing and benchmarks