The AI Stack Cost Report · October 2026 · Part 1: The Frontier Briefing. StackSpend's monthly read on what the modern AI engineering stack actually costs. There was no September issue, so this edition covers everything from early August to October 6, 2026. Also in this issue: the pricing map, the coding stack, the image stack, and Claude Code cost options.
Use this when you run LLMs in production and need a current, factual read on what changed in the last two months: which models shipped, which prices moved, and which retirement dates will force a migration before spring.
The fast answer: Both OpenAI and Anthropic moved a generation in September, and in both cases the biggest change was in the mid-tier: GPT-6.1 Sol and Claude Sonnet 5.5 at $2/$10 per 1M tokens now do work that needed a $5/$25 model in August. Prices went down almost everywhere. The work that remains is mostly calendar work: Claude Sonnet 4.5 retires November 30, 2026, Google's Flash introductory pricing doubles on January 1, 2027, and OpenAI has scheduled shutdowns for its TTS, Whisper/transcription and GPT-5.1-era models between January and April 2027.
August's briefing was about a price war. That continued, but the more useful story for October is how many dates now sit on the calendar of any team running AI in production: two scheduled price increases, at least a dozen retirement deadlines, and several defaults that changed without anyone filing a ticket.
For the per-token detail, pair this with the October 2026 pricing map and the live LLM API Pricing Index. For every release and deprecation by month, see the model changelog.
Quick answer
For an October 2026 snapshot:
- OpenAI moved to GPT-6. GPT-6 Astra (Sept 3) is the flagship at $10/$50. GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) followed on Sept 22, and GPT-6.1 Sol ($2/$10 with $0.10 cached input) on Sept 29. GPT-5.6 Sol was cut to $4/$20 on August 21 as a promotion that runs at least through November 21, 2026.
- Anthropic moved to 5.x point releases. Claude Fable 5.1 (Sept 1, $10/$50, $0.25 cache reads), Opus 5.5 (Sept 22, $4/$20) and Sonnet 5.5 (Sept 28, $2/$10). Sonnet 5's planned rise to $3/$15 was cancelled on August 10.
- Google shipped three Flash models in six weeks. Gemini 3.7 Flash (Aug 13) and 3.8 Flash (Sept 2) both cost $0.75/$3.75, and 3.6 Flash is now listed at the same rate. All three revert to $1.50/$7.50 on January 1, 2027.
- Meta, xAI and Mistral filled in the field. Meta's Muse Spark 1.3 (Sept 2) added a contributor tier that is about 92% cheaper in exchange for training rights on your data. xAI shipped Grok 4.7 (Sept 21, $2/$6). Mistral previewed Mistral Large 4 today, October 6, with weights promised for the end of the month.
- Open weights reached frontier size, but under custom licences. Qwen3.8-Max (2.4T parameters) and GLM-5.3 published weights under their own licences, not Apache or MIT. DeepSeek V4.1 Flash (Sept 10) is the main MIT-licensed exception.
- A deprecation wave started. Claude Sonnet 4.5 retires November 30. OpenAI's Assistants API shut down August 26. Several OpenAI audio and GPT-5.x models have shutdown dates between January 6 and April 1, 2027.
OpenAI: GPT-6 arrives, and GPT-5.6 Sol is cheaper for now
OpenAI's two-month changelog is the busiest of any vendor's. In order:
- August 21: GPT-5.6 Sol cut to $4/$20 (from $5/$30), which OpenAI describes as 20% lower input and 33% lower output. The pricing page says this promotional rate is available "at least through November 21, 2026". That wording is not a commitment to keep it after that date.
- September 3: GPT-6 Astra, "our most capable model, built for the hardest end-to-end work," at $10 input / $1.00 cached / $50 output. It does not support
nonereasoning, custom temperature/top_p or logprobs, and tool calling requires the Responses API. OpenAI added async tool calling, mid-turn steering, and the ability to change reasoning effort mid-conversation while keeping the cached prefix. - September 22: GPT-6 Sol ($2 / $0.20 cached / $10) and GPT-6 Luna ($0.10 / $0.01 cached / $0.50). Both are reasoning models with text and image input.
- September 25: OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna that had degraded image understanding, and recommended rerunning evals for image use cases.
- September 29: GPT-6.1 Sol at $2 / $0.10 cached / $2.50 cache write / $10, aimed at complex coding at lower cost than Astra. On the same day, OpenAI added an Ultrafast service tier for Astra.
Outside models, the platform changes that affect spend were the Agents API public beta (Sept 10, with managed Codex sessions and context compaction), GPT-Live 1 going GA for full-duplex voice at $0.05 per minute plus backend model usage, a Prompt Caching dashboard (Aug 20) with cache-miss diagnostics GA on Sept 8, and usage and cost reporting by API key (Aug 4). On October 6, OpenAI cut its usage tiers from five to three.
What it means for your bill: two things. First, Luna-class traffic got about 50% cheaper per token again: GPT-6 Luna is $0.10/$0.50 against GPT-5.6 Luna's $0.20/$1.20, but only if someone changes the model ID. Second, if you moved traffic to GPT-5.6 Sol because of the August cut, plan for the possibility that it returns to $5/$30 after November 21. GPT-6 Sol and 6.1 Sol at $2/$10 are the obvious places to move it. For the coding view, see the October coding stack.
Anthropic: three new models, one cancelled price rise, one retirement
Anthropic's changes since August, from its release notes:
- August 10: Sonnet 5 pricing finalized. The introductory $2/$10 is now standard, and the increase to $3/$15 scheduled for September 1 did not happen. Our August briefing flagged that increase; it was cancelled.
- September 1: Claude Fable 5.1 and Claude Mythos 5.1 at $10/$50 (same as Fable 5), with cache reads at $0.25 per 1M, 0.025x base input instead of the usual 0.1x. Both require 30-day data retention and are not available under zero data retention unless authorized. Mythos 5.1 remains limited-access.
- September 22: Claude Opus 5.5 at $4/$20 (Opus 5 was $5/$25), with a 1M-token context, 128K max output and always-on adaptive thinking.
- September 28: Claude Sonnet 5.5 at $2/$10.
- September 24: billing resumed for some refusals that arrive before any output (
stop_details.categoryofbio,frontier_llmorreasoning_extraction). These requests are now charged at normal model rates. - September 30: Claude Sonnet 4.5 deprecated, with retirement on the Claude API on November 30, 2026 and Sonnet 5.5 as the recommended replacement.
Opus 5.5 and Sonnet 5.5 are not drop-in swaps. On both, tool_choice of any or tool returns a 400 error. On Opus 5.5, thinking: {"type": "disabled"} and {"type": "enabled"} also return 400; you control depth with the effort parameter instead. The older computer_20251124 tool is rejected on the Claude API and Google Cloud in favour of computer_toolset_20260801. Sonnet 5.5's thinking blocks are bound to the account that produced them.
What it means for your bill: per-token prices fell (Opus 5.5 is 20% below Opus 5), but thinking can't be turned off on the new models and is billed as output. A migration that leaves effort at a high setting can cost more per request than the model it replaces, even at a lower rate. Sonnet 4.5 users have less than eight weeks to migrate, and the 400 errors above mean that migration needs a code change and a test run, not just a new model string. For how Claude Code's defaults moved, see Claude Code cost options in October 2026.
Google: cheap Flash now, a price cliff on January 1
Google shipped Gemini 3.7 Flash on August 13 and Gemini 3.8 Flash on September 2. Google describes 3.8 Flash as built for "long-horizon software engineering, autonomous agents, and complex enterprise workflows". As of October 6, Google's pricing page lists 3.6, 3.7 and 3.8 Flash at the same rate: $0.75 input / $3.75 output per 1M tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027.
Voice and audio got the rest of the attention:
- Gemini 3.8 Live and 3.8 Live Extended Thinking (Sept 15) for real-time voice, at $0.005/min audio input and $0.018/min audio output.
- Gemini 3.8 Flash TTS and Flash-Lite TTS (Sept 22), with 3.8 Flash TTS at $0.50 text input / $9.00 audio output per 1M tokens through December 31, also doubling on January 1, 2027.
- Gemini 3.5 Transcribe and Transcribe Live (Aug 26), covering 85+ languages with speaker diarization.
- Agentic video understanding (Sept 1) for 3.5–3.7 Flash models, which Google says uses up to 88% fewer tokens on long-form video.
The retirement side: Gemini Omni Flash Preview shut down September 30; antigravity-preview-05-2026 ended October 5; gemini-3.1-flash-image shuts down October 29, 2026; and on September 18 Google restricted Gemini 2.5 model access to prior active users.
What it means for your bill: any Gemini Flash workload you price today will double in unit cost on January 1 unless Google extends the promotion. If your 2027 budget was built on current Flash rates, it is understated by up to 2x for that line. Look for this in the January invoice rather than the rate card, since nothing about the API call changes. See how to forecast cloud and AI spend.
New and repriced models at a glance (October 2026)
| Model | Vendor | Input ($/1M) | Output ($/1M) | Status / notes |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | New Sept 3. $1.00 cached; Responses API for tools |
| GPT-6.1 Sol | OpenAI | $2.00 | $10.00 | New Sept 29. $0.10 cached input |
| GPT-6 Sol | OpenAI | $2.00 | $10.00 | New Sept 22; $0.20 cached |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | New Sept 22; $0.01 cached |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | Promo from Aug 21, at least through Nov 21 (was $5/$30) |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | New Sept 1; $0.25 cache reads; 30-day retention required |
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | New Sept 22; 1M context; thinking always on |
| Claude Sonnet 5.5 | Anthropic | $2.00 | $10.00 | New Sept 28; replacement for Sonnet 4.5 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | Intro price made permanent Aug 10 |
| Gemini 3.8 / 3.7 / 3.6 Flash | $0.75 | $3.75 | Through Dec 31; $1.50/$7.50 from Jan 1, 2027 | |
| Grok 4.7 | xAI | $2.00 | $6.00 | New Sept 21; 500K context; $4/$12 at ≥200K-token prompts |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | New Sept 2; 1M context |
| Muse Spark 1.3 Contributor | Meta | $0.10 | $0.20 | Prompts and completions used to train Meta models; 100 RPM |
| Mistral Large 4 | Mistral | $1.36 | $4.18 | Public preview Oct 6; weights promised end of October |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 / $0.30 | $0.60 / $1.20 | New Sept 10; off-peak / peak; MIT weights |
Rates are standard-tier list prices per 1M tokens, as published by each vendor and checked on October 6, 2026, without batch or caching discounts.
How to read this table: most of the rows that matter for production cost sit in the $0.10–$2 input band, and four of them carry an expiry date or a condition: GPT-5.6 Sol's promotion, Gemini Flash's introductory rate, Meta's contributor tier (your data) and DeepSeek's peak pricing (your time of day). The rate card now needs dates and conditions next to it to be useful.
Meta, xAI and Mistral
- Meta Muse Spark 1.3 (September 2) is a multimodal reasoning model with a 1M-token context, sold through the Meta Model API at $1.25/$4.25 standard. The new part is the contributor tier: $0.10 input / $0.20 output, in exchange for "permission to use your prompts and completions to train future Meta models". It is capped at 100 requests per minute, against 3,000 on standard.
- xAI Grok 4.7 (September 21) costs $2/$6 with a 500K context. Prompts of 200K tokens or more are billed at double ($4/$12). It shipped in Cursor first. For coding scores, see the coding stack.
- Mistral Large 4 (public preview, October 6) is a 1T-parameter MoE with 49B active, natively multimodal, at $1.36/$4.18 in preview. Mistral reports 61.7% on DeepSWE v1.1 and says the weights will be published at the end of October. Mistral also announced a €3B raise on September 8.
What it means for your bill: Meta's contributor tier is the first major vendor price that is explicitly paid partly in data. At roughly 92% off standard input pricing it will be tempting for prototypes. Before anyone points production traffic at it, the question for legal and security is whether your prompts can contain customer data. The 100 RPM limit makes it unsuitable for production at scale anyway.
Open weights: frontier-sized, with conditions
Open-weight models kept growing, and their licences got more complicated:
- Qwen3.8-Max: Alibaba's 2.4T-parameter MoE (95B active) published weights as Qwen3.8-2.4T-A95B under the custom qwen3.8-max licence, not Apache 2.0. Native context is 262K tokens, extendable to about 1M, and thinking can't be turned off on the text-only variant.
- GLM-5.3 (Z.ai, launched August 14) published weights under a custom GLM-5.3 licence. Third-party analysis reports that companies above a revenue threshold need a Z.ai security review before commercial use. We could not confirm the exact threshold on the model card, so check the licence text yourself. Codersera reports that the smaller GLM-5.3-Flash uses MIT.
- DeepSeek V4.1 Flash (September 10) is a 552B MoE (8B active on prefill, 16B on decode) with a 1M context, image input and MIT weights. DeepSeek reports V4.1 Flash performing ahead of V4-Pro. On the API it costs $0.15/$0.60 off-peak and double that during peak hours (01:00–04:00 and 06:00–10:00 UTC on weekdays). DeepSeek says the legacy
deepseek-v4-flashIDs now route to V4.1 Flash, and that from September 14deepseek-v4-prorequests also route to V4.1 Flash until V4.1 Pro launches. This follows DeepSeek's August price rise on V4.
What it means for your bill: "open weights" no longer means "use however you like". If self-hosting is part of your cost strategy, licence review now belongs in model selection alongside evals. For hosted DeepSeek, notice that a model ID you didn't change was routed to a different model. That is a quality and cost change that only shows up if you track spend and outputs by model. See tracking open vs closed model costs.
The deprecation calendar: dates that force a model swap
These are the retirement and shutdown dates from vendor documentation, as of October 6, 2026. After the date, requests fail.
| Date | Vendor | What goes away | Recommended replacement |
|---|---|---|---|
| Aug 5, 2026 (done) | Anthropic | claude-opus-4-1-20250805 | Claude Opus 5 / 4.8 |
| Aug 26, 2026 (done) | OpenAI | Assistants API | Responses API + Conversations API |
| Oct 1, 2026 (done) | OpenAI | gpt-5.4-cyber | Most capable cyber model available to you |
| Oct 29, 2026 | gemini-3.1-flash-image | See the image stack | |
| Nov 30, 2026 | Anthropic | claude-sonnet-4-5-20250929 | claude-sonnet-5-5 |
| Jan 6, 2027 | OpenAI | tts-1, tts-1-hd, gpt-4o-mini-tts snapshots | gpt-realtime-2.1-mini |
| Jan 20, 2027 | OpenAI | Legacy audio, realtime and transcription families (announced July 20) | Per OpenAI deprecations page |
| Feb 26, 2027 | OpenAI | whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize | gpt-live-transcribe or gpt-transcribe |
| Apr 1, 2027 | OpenAI | gpt-5.1, gpt-5.3-codex, gpt-5.4-nano | gpt-6-sol / gpt-6-luna |
Also on watch: Anthropic lists Claude Haiku 4.5 as retiring "not sooner than October 15, 2026" and Claude Opus 4.5 "not sooner than November 24, 2026". Neither has been deprecated yet, but Anthropic commits to at least 60 days' notice, so a deprecation notice could arrive at any point. Google restricted Gemini 2.5 access to existing active users on September 18.
How to read this table: the deprecations most likely to catch teams out are audio, not text. Whisper and the TTS-1 family sit in old voice and transcription pipelines that nobody has touched in a year, and both have replacements with different pricing shapes (per-minute sessions, realtime token rates). Text migrations are the easier half; check the voice stack first.
Action list for October
Concrete steps, in date order:
- Before November 21: decide what happens to any traffic on GPT-5.6 Sol if the $4/$20 promotion ends. GPT-6 Sol and 6.1 Sol at $2/$10 are the natural target. Run your eval set on both now.
- Before November 30: migrate off
claude-sonnet-4-5-20250929. Budget engineering time for the 400 errors on forcedtool_choiceand the new thinking parameters in Sonnet 5.5. Export Claude Console usage by API key to find every caller. - Before December 31: rebase your 2027 forecast for Gemini Flash at $1.50/$7.50, and for 3.8 Flash TTS at double its current rate. If those lines matter, compare GPT-6 Luna ($0.10/$0.50) and DeepSeek V4.1 Flash on your own workloads.
- This quarter: inventory voice and transcription pipelines.
tts-1(Jan 6), the legacy realtime families (Jan 20) andwhisper-1(Feb 26) all have hard dates. - Now: check effort settings on any workload you moved to Opus 5.5, Sonnet 5.5 or GPT-6. A lower rate card at a higher default effort can raise the cost per request.
- Now: write down a policy on Meta's contributor tier and similar data-for-discount offers before someone signs up with a work email.
- Ongoing: pin model IDs rather than aliases where cost matters. The
sonnetalias,chat-latestand DeepSeek's legacy IDs all changed which model answered during this period.
What this means for teams running AI in production
The two months since our last issue made the frontier cheaper and the bill harder to predict.
- Most of the savings require someone to act. GPT-6 Luna halves Luna-tier costs and Sonnet 5.5 or GPT-6.1 Sol can replace a $5/$25 model, but none of that happens until someone changes a model ID and checks the quality on your workload.
- Most of the increases happen automatically. Gemini Flash doubling on January 1, a promotion ending on GPT-5.6 Sol, peak-hour pricing on DeepSeek and higher default effort levels need no change on your side to raise the bill.
- Retirement deadlines are a forced model swap with a deadline. Every migration changes token counts and output lengths, and so changes cost. Teams that compare spend by model before and after a swap find out within a day whether the replacement is cheaper; everyone else finds out at month end.
This is the case for watching AI spend at the model level. StackSpend shows AI cost monitoring by provider and model with daily granularity, spend anomaly detection that flags the day a default change or a promotion ending moves your spend, and model recommendations that point out when a cheaper model, like GPT-6 Luna in place of GPT-5.6 Luna, would do the same job. The model changelog tracks the release and retirement dates above as they change. Setup is read-only and takes about five minutes.
FAQ
What were the biggest LLM releases between August and October 2026?
OpenAI's GPT-6 family (Astra on September 3, Sol and Luna on September 22, GPT-6.1 Sol on September 29), Anthropic's Claude Fable 5.1 (September 1), Opus 5.5 (September 22) and Sonnet 5.5 (September 28), Google's Gemini 3.7 Flash (August 13) and 3.8 Flash (September 2), xAI's Grok 4.7 (September 21) and Meta's Muse Spark 1.3 (September 2). Mistral Large 4 entered public preview on October 6. Among open-weight releases, DeepSeek V4.1 Flash (MIT, September 10), Qwen3.8-Max and GLM-5.3 were the largest.
When does Claude Sonnet 4.5 retire?
Anthropic deprecated claude-sonnet-4-5-20250929 on September 30, 2026, and will retire it on the Claude API on November 30, 2026. After that date, requests fail. The recommended replacement is Claude Sonnet 5.5 ($2/$10 per 1M tokens), which rejects forced tool_choice (any or tool) with a 400 error, so test the migration before switching.
Is Gemini 3.8 Flash going up in price?
Yes, unless Google changes its plans. Gemini 3.8, 3.7 and 3.6 Flash are priced at $0.75 input / $3.75 output per 1M tokens through December 31, 2026, and Google's pricing page lists $1.50/$7.50 from January 1, 2027. That doubles the unit cost of any Flash workload that doesn't change model.
Which OpenAI models are being shut down?
The Assistants API shut down on August 26, 2026, and gpt-5.4-cyber on October 1, 2026. Scheduled shutdowns: TTS models (tts-1, tts-1-hd, gpt-4o-mini-tts snapshots) on January 6, 2027; legacy audio and realtime families on January 20, 2027; whisper-1 and the gpt-4o-transcribe family on February 26, 2027; and gpt-5.1, gpt-5.3-codex and gpt-5.4-nano on April 1, 2027.
What is Meta's Muse Spark contributor tier?
It is a version of Muse Spark 1.3 priced at $0.10 input / $0.20 output per 1M tokens, compared with $1.25/$4.25 on the standard tier, in exchange for permission for Meta to use your prompts and completions to train future Meta models. It is limited to 100 requests per minute. The standard tier's data is not used for training.
How do I know whether these changes moved my AI bill?
Compare spend per model before and after each change date, not against the rate card. A cheaper model can still cost more per request at a higher effort setting, and a promotion ending changes cost with no change to your code. StackSpend breaks AI spend down by provider and model each day and flags anomalies the day they happen — see AI cost monitoring and spend anomaly detection.
About The AI Stack Cost Report
The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs: frontier and open-weight LLMs, AI coding tools, image generation, and the infrastructure they run on. Each issue is compiled from primary vendor pricing pages, changelogs and deprecation pages, with source-check dates in the frontmatter. This October 2026 issue follows the August 2026 frontier briefing; there was no September issue.
Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily, the model changelog logs releases and deprecations by month, and the model glossary tracks each provider's family history.
A monthly snapshot can't tell you what your stack spent yesterday. StackSpend can: per-provider and per-model AI spend, same-day anomaly alerts and a daily report in Slack. Setup is read-only.
Related reading
- LLM Model Pricing in October 2026
- The Best Coding Models in October 2026
- Image Generation Models in October 2026
- Claude Code Cost Options and Overage (October 2026)
- The Latest in LLMs, August 2026, the previous issue
- OpenAI vs Anthropic Pricing in 2026
- Anomaly Detection vs Budget Alerts
- LLM API Pricing Index: live per-1M-token rates, updated daily
References
- OpenAI — API changelog (August–October 2026), checked 2026-10-06
- OpenAI — API pricing (incl. GPT-5.6 Sol promotional note), checked 2026-10-06
- OpenAI — Deprecations, checked 2026-10-06
- OpenAI — GPT-6 Astra
- OpenAI — Introducing GPT-6.1 Sol
- Anthropic — Claude Platform release notes, checked 2026-10-06
- Anthropic — Model deprecations, checked 2026-10-06
- Anthropic — Pricing (Claude Platform docs)
- Anthropic — Introducing Claude Opus 5.5
- Anthropic — Introducing Claude Sonnet 5.5
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Google — Gemini API changelog, checked 2026-10-06
- Google — Gemini API pricing, checked 2026-10-06
- Google — Gemini API deprecations
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- Axios — Google updates Gemini with faster Flash model (Aug 13, 2026)
- xAI — Grok 4.7
- xAI — Models and pricing
- Meta — Model API pricing and rate limits
- Vercel AI Gateway — Muse Spark 1.3 model page
- Mistral — Introducing Mistral Large 4 (Oct 6, 2026)
- Mistral — Mistral raises €3B (Sept 8, 2026)
- DeepSeek — V4.1-Flash release news (Sept 10, 2026)
- DeepSeek — Models and pricing
- Hugging Face — deepseek-ai/DeepSeek-V4.1-Flash model card
- Caixin Global — DeepSeek launches V4 Pro and raises API prices (Aug 14, 2026)
- Hugging Face — Qwen/Qwen3.8-2.4T-A95B model card
- Qwen — Qwen3.8-Max: A New Bar for Coding and Cowork
- Hugging Face — zai-org/GLM-5.3 model card
- Codersera — GLM-5.3 launch guide (licence analysis)

