The AI Stack Cost Report · October 2026 · The Pricing Map. StackSpend's monthly read on what the modern AI engineering stack actually costs. There was no September issue, so this edition covers everything that changed between early August and October 6, 2026. Also in this issue: the frontier briefing, the coding stack, and the image stack.
Use this when you already know roughly what you spend on LLM APIs and want to know whether anything that shipped since August should change which model you pay for. If you only need a single vendor's current rate, the vendor's pricing page is faster; this guide is about the decision.
The fast answer: the mid-tier got much better without getting more expensive. Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol all cost $2/$10, and Claude Opus 5.5 is $4/$20, 20% below Opus 5. At the bottom, GPT-6 Luna is $0.10/$0.50, half of GPT-5.6 Luna. The frontier tier held at $10/$50 (GPT-6 Astra, Claude Fable 5.1). The planned Claude Sonnet 5 increase to $3/$15 did not happen. Two newer price risks did appear: DeepSeek now charges double at peak hours, and Google's Flash models carry introductory rates that double on January 1, 2027.
Between August and October, more than a dozen models from the vendors in this guide shipped or were repriced. Most teams will not have re-tested their model choice against all of them, so the useful question is which of these changes reduces your bill. This issue starts with that, then gives the full rate card for reference.
For rates between issues, the LLM API Pricing Index re-syncs published per-1M-token prices daily, and the model changelog logs releases and deprecations by month. To run these numbers on your own token volumes, use the LLM cost calculator.
Quick answer
As of October 6, 2026:
- Frontier: GPT-6 Astra and Claude Fable 5.1 at $10/$50 per 1M tokens. Fable 5.1 cache reads are $0.25; Astra's cached input is $1.00.
- Premium default: Claude Opus 5.5 at $4/$20, with $0.20 cache reads. GPT-5.6 Sol is temporarily $4/$20 on a promotion reported to run three months from August 21.
- Mid-tier ($2 input): Claude Sonnet 5.5 and Sonnet 5 ($2/$10), GPT-6.1 Sol ($2/$10, $0.10 cached), GPT-6 Sol ($2/$10), Grok 4.7 ($2/$6), Gemini 3.1 Pro ($2/$12).
- Volume tier: Gemini 3.8 Flash $0.75/$3.75 until December 31, Claude Haiku 4.5 $1/$5, GPT-5.6 Luna $0.20/$1.20, GPT-6 Luna $0.10/$0.50.
- Open weight: DeepSeek V4.1 Flash is $0.15/$0.60 off-peak and $0.30/$1.20 at peak on DeepSeek's own API. The largest new open models (Kimi K3, Qwen3.8-2.4T) cost $2–$3 input on managed hosts.
What changed since August
Our August issue was written around Claude Opus 5, GPT-5.6 and an expected Sonnet 5 price rise. Since then:
| Date (2026) | Change | What it means for your bill |
|---|---|---|
| Aug 13 | Gemini 3.7 Flash at $0.75/$3.75 introductory, half Gemini 3.6 Flash's $1.50/$7.50 launch price | Flash-tier work halved in price, but only until Dec 31 |
| Aug 13–29 | Open weights for Qwen3.8-2.4T-A95B (Aug 13) and GLM-5.3 (Aug 29) | More frontier-scale open options, priced at $1.40–$2.00 input on hosts |
| Aug 17 | DeepSeek introduces peak and off-peak pricing; V4 Pro rates rise by up to 1,100% on some token types | Same model, 2x price difference depending on the hour |
| Aug 21 | GPT-5.6 Sol cut from $5/$30 to $4/$20 for three months (reported) | Temporary; budget for the old rate returning around late November |
| Aug–Sep 1 | Claude Sonnet 5's $2/$10 made standard; the $3/$15 increase due September 1 was cancelled. Claude Fable 5.1 at $10/$50 with $0.25 cache reads (Sep 1) | The 50% Sonnet rise we flagged in August did not happen |
| Sep 2 | Gemini 3.8 Flash at $0.75/$3.75 introductory; Meta Muse Spark 1.3 at $1.25/$4.25, plus a $0.10/$0.20 contributor tier that lets Meta train on your prompts | Better Flash quality at the same introductory rate; a data-for-price trade at Meta |
| Sep 10–14 | DeepSeek V4.1 Flash replaces V4 Flash ($0.15/$0.60 off-peak); DeepSeek says deepseek-v4-pro requests route to V4.1 Flash from Sep 14 until V4.1 Pro launches | Model IDs you didn't change now serve a different model at a different rate |
| Sep 3 | GPT-6 Astra at $10/$50 | New OpenAI flagship at twice GPT-5.5's input price |
| Sep 21 | Grok 4.7 at $2/$6 | Cheapest output rate in the $2-input band |
| Sep 22 | Claude Opus 5.5 at $4/$20 ($0.20 cache reads); GPT-6 Sol and GPT-6 Luna | Premium tier 20% cheaper per token; bulk tier halved |
| Sep 28 | Claude Sonnet 5.5 at $2/$10 | Sonnet upgrade at no price change |
| Sep 29 | GPT-6.1 Sol at $2/$10, $0.10 cached input | OpenAI says it nearly matches Astra on coding at about one-fifth the cost |
| Oct 6 | Mistral Large 4 in public preview at $1.36/$4.18; weights promised by end of October | A new low-priced large-model option to test |
Corrections to our August issue: the Sonnet 5 increase to $3/$15 on September 1 was cancelled. We also listed Mistral Large 3 at roughly $2/$6 from market quotes; Mistral's own pricing page currently gives Mistral Large at $0.50/$1.50, and the new Mistral Large 4 preview is $1.36/$4.18.
What this means for your bill
Three patterns explain most of the October changes:
- If you pay for a premium model, you are probably overpaying by default. Opus 5.5 is 20% cheaper per token than Opus 5, and Anthropic says it costs about 40% less on typical workloads once fewer steps and tokens are counted. Teams still pinned to
claude-opus-5or GPT-5.6 Sol at $5/$30 are paying August prices for an older model. - $2/$10 now covers most work that needed $5/$25 in August. Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol share that rate, and Grok 4.7 undercuts it on output at $6. See the coding stack for where each holds up on quality.
- The cheapest rates now come with conditions. Gemini's Flash rates end December 31, GPT-5.6 Sol's cut is temporary, and DeepSeek charges double during peak hours. If you set budgets from list prices without tracking these conditions, forecasts will be wrong.
Which swaps actually cut cost
Illustrative workload. 100M input tokens and 20M output tokens a month, roughly a customer-support assistant or an internal RAG feature at moderate scale. List rates as of October 6, 2026, with no caching or batch discount unless stated. Your bill will differ with retries, reasoning tokens and tokenizer differences.
The formula is the same as every issue:
Monthly cost = (input_tokens / 1M × input_rate) + (output_tokens / 1M × output_rate)
| Swap | Before (monthly) | After (monthly) | Saving | The catch |
|---|---|---|---|---|
| Claude Opus 5 → Opus 5.5 | $500 + $500 = $1,000 | $400 + $400 = $800 | 20% (more if Anthropic's ~40% workload claim holds for you) | Thinking can't be turned off on Opus 5.5; set effort deliberately |
| Claude Opus 5.5 → Sonnet 5.5 | $800 | $200 + $200 = $400 | 50% | Opus still leads on judgement-heavy tasks; test on your own evals |
| GPT-5.6 Sol (August list, $5/$30) → GPT-6.1 Sol | $500 + $600 = $1,100 | $400 | 64% (50% against the $4/$20 promotional rate) | A model migration, not a config change; re-run evals |
| GPT-5.6 Luna → GPT-6 Luna | $20 + $24 = $44 | $10 + $10 = $20 | 55% | Small in dollars here; large at billions of tokens |
| Gemini 3.6 Flash (launch price) → Gemini 3.8 Flash | $150 + $150 = $300 | $75 + $75 = $150 | 50% until Dec 31; then $300 again | Google says 3.8 Flash may use more tokens at higher effort |
| DeepSeek V4.1 Flash, peak → off-peak (DeepSeek API) | $30 + $24 = $54 | $15 + $12 = $27 | 50% | Only for jobs you can schedule outside peak hours |
| Llama 3.3 70B, Together AI → DeepInfra | $104 + $20.80 = $124.80 | $10 + $6.40 = $16.40 | 87% | Same weights; check throughput, rate limits and serving quantization |
How to read this table: the first three rows are where most October savings sit, because premium and mid-tier models carry most production spend. A team that moves an Opus 5 workload to Sonnet 5.5 where quality allows goes from $1,000 to $400 a month on this volume, a 60% cut with no new vendor. The last two rows involve no model change at all. They come from when you run the work and which host serves it.
Swaps that look cheaper but aren't
- Kimi K3 instead of a $2/$10 closed model. At $2.70–$3.00 input and $13.50–$15.00 output on managed hosts, Kimi K3 costs $540–$600 on this workload, compared with $400 for GPT-6.1 Sol or Sonnet 5.5. Choose open frontier weights for control or self-hosting, not to save money.
- Gemini 3.8 Flash as a permanent budget line. It costs $150 on this workload until December 31 and $300 from January 1, 2027. At that rate it costs more than Claude Haiku 4.5 ($200) and far more than GPT-6 Luna ($20) or DeepSeek V4.1 Flash off-peak ($27).
- Treating GPT-5.6 Sol's $4/$20 as the new list price. The cut was reported as a three-month promotion from August 21. If nothing changes, this workload goes back from $800 to $1,100.
Caching changes the ranking
Cache-read pricing now varies more than headline input pricing. If 80% of input is a cached prefix (common for RAG system prompts and agent loops), the same 100M/20M workload costs, before cache-write charges:
- GPT-6.1 Sol: 20M × $2 + 80M × $0.10 + 20M × $10 = $248
- Claude Sonnet 5.5: 20M × $2 + 80M × $0.20 + 20M × $10 = $256
- Claude Opus 5.5: 20M × $4 + 80M × $0.20 + 20M × $20 = $496
Caching cuts Opus 5.5's input cost by about 76% but cuts the total bill by only 38%, because output costs stay the same. Output now accounts for most premium-tier spend, so the effort level you set matters as much as the rate card.
Price per task, not per token
Per-token tables hide the number that decides most model choices: what one completed task costs. Below is one illustrative agentic task, such as a scoped bug fix or a multi-step support resolution: 100K uncached input, 400K cache-read input, 30K output (reasoning included), with cache-write charges excluded.
| Model | Uncached input | Cache reads | Output | Per task | Per 1,000 tasks |
|---|---|---|---|---|---|
| GPT-6 Astra | $1.00 | $0.40 | $1.50 | $2.90 | $2,900 |
| Claude Fable 5.1 | $1.00 | $0.10 | $1.50 | $2.60 | $2,600 |
| Claude Opus 5 | $0.50 | $0.20 | $0.75 | $1.45 | $1,450 |
| Claude Opus 5.5 | $0.40 | $0.08 | $0.60 | $1.08 | $1,080 |
| Claude Sonnet 5.5 | $0.20 | $0.08 | $0.30 | $0.58 | $580 |
| Grok 4.7 | $0.20 | $0.20 | $0.18 | $0.58 | $580 |
| GPT-6.1 Sol | $0.20 | $0.04 | $0.30 | $0.54 | $540 |
| Gemini 3.8 Flash (introductory) | $0.075 | $0.03 | $0.11 | $0.22 | $218 |
| GPT-6 Luna | $0.01 | $0.004 | $0.015 | $0.03 | $29 |
The per-task view supports a recommendation you can test. If a $0.54 model needs a second attempt on more than half of your tasks, the $1.08 model that succeeds first time costs the same and takes less engineer time. Track success rate on first attempt alongside spend for two weeks before you standardise on the cheaper model. Models also differ in how many tokens they use for the same task (Anthropic's launch post says Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps), so measured per-task cost can differ widely from this same-token comparison.
How to read LLM pricing in October 2026
Rates are still quoted per million tokens, split into input and output. Five rules carry most of the cost math:
- Output costs 3–6x input, and reasoning bills as output. On Opus 5.5, Sonnet 5.5 and the Fable models thinking can't be switched off, only turned down.
- Cache-read multipliers are no longer uniform. Anthropic charges 0.1x base input on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1 and Mythos 5.1. OpenAI's GPT-6.1 Sol is 0.05x and GPT-6 Luna 0.1x.
- Long context has cliffs. GPT-6 Astra moves to $20/$75 in its long-context tier, Gemini 3.1 Pro to $4/$18 above 200K input, and Grok 4.7 to $4/$12 at 200K and above. Anthropic includes the full 1M context at standard rates on Claude 4.6 and later.
- Time and date now change price. DeepSeek's peak hours (01:00–04:00 and 06:00–10:00 UTC on weekdays) cost twice off-peak, and Google's introductory Flash rates double on January 1, 2027.
- Tokenizers differ. Claude 4.7 and later produce about 30% more tokens for the same text than earlier Claude models, per Anthropic, so compare per task, not per token, across vendors and generations.
Closed frontier APIs (October 2026)
List prices checked October 6, 2026, standard tier, short context.
| Provider | Model | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|---|
| Anthropic | Claude Mythos 5.1 | $10.00 | $50.00 | Limited availability (Project Glasswing); cache reads $0.25 |
| Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | New Sep 1. Cache reads $0.25 (Fable 5: $1.00); batch $5/$25 |
| Anthropic | Claude Opus 5.5 | $4.00 | $20.00 | New Sep 22. Cache reads $0.20; batch $2/$10; fast mode $8/$40 |
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | Cache reads $0.50; batch $2.50/$12.50 |
| Anthropic | Claude Sonnet 5.5 | $2.00 | $10.00 | New Sep 28. Cache reads $0.20; batch $1/$5 |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | Introductory rate made standard; $3/$15 increase cancelled |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | Cache reads $0.10; batch $0.50/$2.50 |
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | New Sep 3. Cached $1.00; long-context tier $20/$75 |
| OpenAI | GPT-6.1 Sol | $2.00 | $10.00 | New Sep 29. Cached $0.10; long-context tier $4/$15 |
| OpenAI | GPT-6 Sol | $2.00 | $10.00 | New Sep 22. Cached $0.20 |
| OpenAI | GPT-6 Luna | $0.10 | $0.50 | New Sep 22. Cached $0.01 |
| OpenAI | GPT-5.6 Sol | $4.00 | $20.00 | Cut from $5/$30 on Aug 21; reported as a three-month promotion |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | Unchanged since July 30 cut |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | Unchanged; GPT-6 Luna is half the price |
| OpenAI | GPT-5.5 / GPT-5.5 Pro | $5.00 / $30.00 | $30.00 / $180.00 | Prior flagship and highest-effort tier |
| Gemini 3.8 Flash | $0.75 | $3.75 | New Sep 2. Introductory to Dec 31; $1.50/$7.50 from Jan 1, 2027 | |
| Gemini 3.7 / 3.6 Flash | $0.75 | $3.75 | Same introductory terms listed on Google's pricing page | |
| Gemini 3.1 Pro (Preview) | $2.00 | $12.00 | $4/$18 above 200K input; still the only Pro on the Gemini API | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Unchanged | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Cheapest Gemini text tier | |
| xAI | Grok 4.7 | $2.00 | $6.00 | New Sep 21. Cached $0.50; $4/$12 at 200K+ tokens; 500K context |
| xAI | Grok 4.3 | $1.25 | $2.50 | 1M context; cached $0.20 |
| Meta | Muse Spark 1.3 | $1.25 | $4.25 | New Sep 2. Cached $0.15; contributor tier $0.10/$0.20 (Meta may train on your data; 100 RPM) |
| Mistral | Mistral Large 4 | $1.36 | $4.18 | Public preview Oct 6; 1T total / 49B active; weights promised by end of October |
| Mistral | Mistral Large (previous) | $0.50 | $1.50 | Rate shown on Mistral's pricing page |
Batch pricing is 50% off at Anthropic, OpenAI and Google. US-only inference on Claude 4.6 and later adds a 1.1x multiplier.
How to read the closed tier in October: the frontier band (GPT-6 Astra, Fable 5.1) is now 2.5x Opus 5.5 and 5x the $2 mid-tier, so it needs a task-level reason. The $2-input band is crowded with five strong models, and output rate ($6 to $12) is what separates them. At the bottom, GPT-6 Luna at $0.10/$0.50 is cheaper than every Gemini tier and than most hosted open models. For a vendor head-to-head, see OpenAI vs Anthropic pricing.
Open-weight models (October 2026)
Prices below are what managed hosts charge for each model, taken from each host's public pricing page on October 6. The weights are free; you pay for someone's hardware.
| Model | Input ($/1M) | Output ($/1M) | Where quoted / notes |
|---|---|---|---|
| Kimi K3 | $2.70–$3.00 | $13.50–$15.00 | Together $2.70/$13.50; DeepInfra $2.85/$14.25; Fireworks and Baseten $3/$15 |
| Qwen3.8-2.4T-A95B (Qwen3.8 Max) | $2.00 | $6.00 | Together and Fireworks. Weights published Aug 13; largest open-weight model to date |
| GLM-5.3 | $1.40 | $4.40 | Together, Fireworks, Baseten. Weights Aug 29 under a custom GLM-5.3 licence |
| DeepSeek V4 Pro | $1.30–$1.32 | $2.60–$3.96 | DeepInfra $1.30/$2.60; Together and Baseten (V4 Pro 0813) $1.32/$3.96. On DeepSeek's own API, see the note below |
| Inkling | $1.00 | $4.05 | Together AI |
| NVIDIA Nemotron 3 Ultra | $0.60 | $2.40 | Baseten |
| DeepSeek V4.1 Flash | $0.15–$0.30 | $0.60–$1.20 | DeepSeek API off-peak/peak; Fireworks, Together and Baseten $0.30/$1.20. Replaces V4 Flash on DeepSeek's API |
| GLM-5.3 Flash / Qwen3.8 Flash | $0.15 | $0.47–$0.50 | Together, Fireworks, Baseten |
| GPT-OSS 120B | $0.10–$0.15 | $0.50–$0.60 | Baseten $0.10/$0.50; Groq, Fireworks and Together $0.15/$0.60 |
| DeepSeek V4 Flash (older checkpoint) | $0.09–$0.13 | $0.18–$0.26 | DeepInfra $0.09/$0.18; Baseten V4-Flash-0731 $0.13/$0.26 |
| Llama 3.3 70B | $0.10–$1.04 | $0.32–$1.04 | DeepInfra $0.10/$0.32; Together $1.04/$1.04 |
| Llama 3.1 8B | $0.02–$0.14 | $0.04–$0.14 | DeepInfra $0.02/$0.04; Together (Lite) $0.14/$0.14 |
DeepSeek V4 Pro on DeepSeek's own API: the pricing page still lists V4 Pro at $1.32/$3.96 peak and $0.66/$1.98 off-peak, but DeepSeek's September 10 release note says deepseek-v4-pro requests route to V4.1 Flash at Flash rates from September 14 until V4.1 Pro launches. Check your invoice for which rate you are being billed. If you need V4 Pro itself, the managed hosts above still serve it.
Open weight does not always mean OSI open source. GLM-5.3 and several others ship under custom licences, so check the terms before you ship.
The October pattern: open-weight pricing now has two separate ends. At the top, Kimi K3, Qwen3.8 Max and GLM-5.3 cost $1.40–$3.00 input, close to or above the $2/$10 closed mid-tier. At the bottom, GPT-OSS, older DeepSeek Flash checkpoints and small Llama models still cost cents per million. DeepSeek's own API also stopped being the reliable price floor in August: its V4.1 Flash peak rate ($0.30/$1.20) is now higher than GPT-6 Luna ($0.10/$0.50).
Open-model hosting providers (October 2026)
The same weights cost different amounts depending on who serves them. The October examples below come from each host's pricing page.
| Host | Pricing model | October example rates | Best for |
|---|---|---|---|
| DeepInfra | Per-token on its own stack; cached-input discounts | DeepSeek V4 Pro $1.30/$2.60; Llama 3.3 70B $0.10/$0.32; Llama 3.1 8B $0.02/$0.04 | Lowest per-token price on most established open models |
| Fireworks AI | Standard, Priority and Fast serving paths; batch 50% off | Kimi K3 $3/$15 (Fast $4.50/$22.50); GLM-5.3 $1.40/$4.40; V4.1 Flash $0.30/$1.20 | Production serving with speed tiers and US-hosted variants |
| Together AI | Flat per-token serverless; also GPU rental and fine-tuning | Kimi K3 $2.70/$13.50; Qwen3.8-2.4T $2/$6; Inkling $1/$4.05 | Broadest catalog, including the newest large open models |
| Baseten | Pay-as-you-go Model APIs plus dedicated deployments | GPT-OSS 120B $0.10/$0.50; GLM-5.3 $1.40/$4.40; Kimi K3 $3/$15 | Managed APIs with a path to dedicated capacity |
| Groq | Per-token on LPU hardware | GPT-OSS 120B $0.15/$0.60 (~500 tok/s); GPT-OSS 20B $0.075/$0.30. Llama tiers now "contact sales" | Latency-bound chat and tool loops |
| OpenRouter | Pass-through provider pricing, no inference markup; 5.5% fee on card credit purchases | Underlying host rate plus credit fee | One API with routing and failover across hosts |
| Hugging Face Inference Providers | Pass-through provider rates, no HF markup; PRO includes $2/month compute credits | Underlying host rate | One token across many hosts, with per-model, per-provider usage breakdowns |
The host spread is still large: Llama 3.3 70B costs about 10x more on Together ($1.04/$1.04) than on DeepInfra ($0.10/$0.32). On newer models the spread is smaller. Kimi K3 varies by about 11% and GLM-5.3 is $1.40/$4.40 on all three hosts we checked, because fewer hosts serve them and those hosts tend to follow the lab's reference price. The cheapest host is not always the best value. A faster host can reduce total cost on agent loops that would otherwise wait on generation.
The pricing traps to watch
- Expiring promotions. Gemini 3.6/3.7/3.8 Flash go from $0.75/$3.75 to $1.50/$7.50 on January 1, 2027. GPT-5.6 Sol's $4/$20 is reported to run three months from August 21. Put both dates in your forecast now.
- Peak-hour pricing. DeepSeek's API costs twice as much between 01:00–04:00 and 06:00–10:00 UTC on weekdays. An interactive feature used in Asian business hours pays the peak rate on most of its traffic.
- Default model drift. Tools and aliases move up a tier on upgrade. Claude Code's default is now Opus 5.5 on every first-party plan (Claude Code cost options), and API aliases change which model they point to.
- Retired model IDs that still answer. DeepSeek serves
deepseek-v4-flashrequests with V4.1 Flash at the Flash price, and saysdeepseek-v4-prorequests have routed to V4.1 Flash since September 14. Your code doesn't change, but the model and the rate behind it do. - Data-for-price tiers. Meta's Muse Spark 1.3 contributor tier is 92% cheaper on input than standard because Meta may train on your prompts and completions. Check with legal before any customer data goes through it.
- Long-context tiers. GPT-6 Astra, Grok 4.7 and Gemini 3.1 Pro all price long prompts higher. RAG pipelines whose retrieved context grows over time can cross the threshold without any code change.
- Reasoning you can't switch off. On Opus 5.5, Sonnet 5.5 and the Fable models, thinking always bills as output. Effort is the control.
FAQ
What changed in LLM API pricing between August and October 2026?
The mid-tier improved without a price rise. Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol launched at $2/$10, and Claude Opus 5.5 at $4/$20, 20% below Opus 5. GPT-6 Luna halved the bulk tier to $0.10/$0.50. GPT-6 Astra and Claude Fable 5.1 hold the frontier at $10/$50. DeepSeek moved to peak and off-peak pricing on August 17, and the planned Claude Sonnet 5 increase to $3/$15 was cancelled.
What is the cheapest LLM API in October 2026?
Among closed APIs, GPT-6 Luna at $0.10 input and $0.50 output per 1M tokens. Among open-weight models on managed hosts, Llama 3.1 8B on DeepInfra ($0.02/$0.04) and older DeepSeek V4 Flash checkpoints ($0.09/$0.18 on DeepInfra) are lower still, but much less capable. For a capable general model, GPT-6 Luna and DeepSeek V4.1 Flash off-peak ($0.15/$0.60) are the practical floor.
Is Claude Opus 5.5 cheaper than Opus 5?
Yes. Opus 5.5 is $4/$20 per 1M tokens against Opus 5's $5/$25, and cache reads are $0.20 against $0.50. Anthropic says it costs about 40% less than Opus 5 on typical workloads once fewer steps and tokens are counted. On our 100M input / 20M output example, the list-price saving alone is $1,000 to $800 a month.
When does Gemini 3.8 Flash's introductory pricing end?
December 31, 2026. From January 1, 2027, Gemini 3.8 Flash costs $1.50 input and $7.50 output per 1M tokens, double the introductory $0.75/$3.75. Google's pricing page lists the same terms for Gemini 3.7 Flash and 3.6 Flash.
Is GPT-6.1 Sol or Claude Sonnet 5.5 cheaper?
Both list at $2/$10 per 1M tokens. GPT-6.1 Sol's cached input is $0.10 against Sonnet 5.5's $0.20 cache reads, so GPT-6.1 Sol is slightly cheaper on cache-heavy workloads: $248 against $256 a month in our 80%-cached example. Tokenizer differences and per-task efficiency usually matter more than that gap, so compare them on cost per completed task.
Track LLM spend across every provider
A list price tells you what a token costs; it can't tell you which model your code called yesterday, at what effort level, at what hour or through which host. Each of those changed price between August and October. A typical stack now spans one or two closed APIs, an open-model host and sometimes an aggregator, each invoicing separately.
StackSpend puts them on one timeline, with model-level breakdown, model recommendations that flag when a cheaper model such as Sonnet 5.5 is doing the same job as Opus 5.5, and same-day anomaly detection. For the overall picture, start with AI cost monitoring.
About The AI Stack Cost Report
The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs: frontier and open-weight LLMs, AI coding tools, image generation, and the hosting they run on. Each issue is compiled from primary pricing pages and release reporting, with source-check dates in the frontmatter. The October 2026 issue covers August through early October (there was no September issue) and has four parts: the frontier briefing, the pricing map, the coding stack, and the image stack.
Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily and the model changelog logs releases and deprecations by month.
Related reading
- LLM Model Pricing in August 2026, the previous issue
- The Best Coding Models in October 2026
- Claude Code Cost Options and Overage (October 2026)
- OpenAI vs Anthropic Pricing in 2026
- Cheapest AI API in 2026 for Chat, RAG, and Coding
- Bedrock vs Vertex AI Pricing: What Teams Actually Pay
- Why Your Claude API Bill Is So High
- LLM cost calculator
References
- Anthropic — Pricing (Claude Platform docs: model table, cache multipliers, Sonnet 5 footnote, batch, fast mode, data residency), checked 2026-10-06
- Anthropic — Introducing Claude Opus 5.5
- Anthropic — Introducing Claude Sonnet 5.5
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1
- TechCrunch — Anthropic releases Sonnet 5.5 (Sept 28, 2026)
- OpenAI — API pricing, checked 2026-10-06
- OpenAI Developer Community — Introducing GPT-6 Astra (Sept 3, 2026)
- OpenAI — Introducing GPT-6.1 Sol
- The Next Web — OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices
- Let's Data Science — OpenAI cuts GPT-5.6 Sol API prices (Aug 21, 2026)
- Outlook Business — OpenAI cuts GPT-5.6 Sol prices for developers
- Google — Gemini API pricing, checked 2026-10-06
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- AI Weekly — Google ships Gemini 3.7 Flash
- xAI — Models and pricing, checked 2026-10-06
- xAI — Grok 4.7
- DeepSeek — Models and pricing (peak/off-peak, V4.1 Flash, legacy names), checked 2026-10-06
- DeepSeek — V4.1-Flash release news (Sept 10, 2026; V4 Pro routing from Sept 14)
- Caixin Global — DeepSeek launches V4 Pro and raises API prices (Aug 14, 2026)
- Meta — Model API pricing and rate limits, checked 2026-10-06
- Let's Data Science — Meta releases Muse Spark 1.3 (Sept 2, 2026)
- Mistral — API pricing
- Mistral — Introducing Mistral Large 4 (Oct 6, 2026)
- Mistral Docs — Models overview
- Gigazine — GLM-5.3 released as an open model (Aug 29, 2026)
- NYU Shanghai RITS — Qwen3.8-2.4T-A95B: Alibaba open-weights its Max-tier flagship
- DeepInfra — Pricing
- Fireworks AI — Serverless pricing
- Together AI — Pricing
- Baseten — Pricing
- Groq — Supported models and pricing
- OpenRouter — FAQ (fees and pass-through pricing)
- Hugging Face — Inference Providers pricing

