Stack Spend

LLM Model Pricing in October 2026: What Changed and Which Swaps Cut Cost

GuidesOctober 6, 2026By 19 min read

The short answer

As of October 6, 2026, the mid-tier is where prices moved most. Claude Sonnet 5.5, GPT-6.1 Sol and GPT-6 Sol all cost $2/$10 per 1M tokens, Claude Opus 5.5 is $4/$20 (down from Opus 5's $5/$25), and GPT-6 Luna is $0.10/$0.50. The frontier tier is still $10/$50 (GPT-6 Astra, Claude Fable 5.1). On a workload of 100M input and 20M output tokens a month, moving from Opus 5 to Opus 5.5 saves 20% ($1,000 to $800), moving from Opus 5.5 to Sonnet 5.5 or GPT-6.1 Sol saves another 50% ($400), and moving from GPT-5.6 Luna to GPT-6 Luna cuts $44 to $20. Two dates to watch: GPT-5.6 Sol's $4/$20 promotional rate is reported to run three months from August 21, and Gemini 3.8 Flash doubles from $0.75/$3.75 to $1.50/$7.50 on January 1, 2027.

The AI Stack Cost Report · October 2026 · The Pricing Map. StackSpend's monthly read on what the modern AI engineering stack actually costs. There was no September issue, so this edition covers everything that changed between early August and October 6, 2026. Also in this issue: the frontier briefing, the coding stack, and the image stack.

Use this when you already know roughly what you spend on LLM APIs and want to know whether anything that shipped since August should change which model you pay for. If you only need a single vendor's current rate, the vendor's pricing page is faster; this guide is about the decision.

The fast answer: the mid-tier got much better without getting more expensive. Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol all cost $2/$10, and Claude Opus 5.5 is $4/$20, 20% below Opus 5. At the bottom, GPT-6 Luna is $0.10/$0.50, half of GPT-5.6 Luna. The frontier tier held at $10/$50 (GPT-6 Astra, Claude Fable 5.1). The planned Claude Sonnet 5 increase to $3/$15 did not happen. Two newer price risks did appear: DeepSeek now charges double at peak hours, and Google's Flash models carry introductory rates that double on January 1, 2027.

Between August and October, more than a dozen models from the vendors in this guide shipped or were repriced. Most teams will not have re-tested their model choice against all of them, so the useful question is which of these changes reduces your bill. This issue starts with that, then gives the full rate card for reference.

For rates between issues, the LLM API Pricing Index re-syncs published per-1M-token prices daily, and the model changelog logs releases and deprecations by month. To run these numbers on your own token volumes, use the LLM cost calculator.

Quick answer

As of October 6, 2026:

  • Frontier: GPT-6 Astra and Claude Fable 5.1 at $10/$50 per 1M tokens. Fable 5.1 cache reads are $0.25; Astra's cached input is $1.00.
  • Premium default: Claude Opus 5.5 at $4/$20, with $0.20 cache reads. GPT-5.6 Sol is temporarily $4/$20 on a promotion reported to run three months from August 21.
  • Mid-tier ($2 input): Claude Sonnet 5.5 and Sonnet 5 ($2/$10), GPT-6.1 Sol ($2/$10, $0.10 cached), GPT-6 Sol ($2/$10), Grok 4.7 ($2/$6), Gemini 3.1 Pro ($2/$12).
  • Volume tier: Gemini 3.8 Flash $0.75/$3.75 until December 31, Claude Haiku 4.5 $1/$5, GPT-5.6 Luna $0.20/$1.20, GPT-6 Luna $0.10/$0.50.
  • Open weight: DeepSeek V4.1 Flash is $0.15/$0.60 off-peak and $0.30/$1.20 at peak on DeepSeek's own API. The largest new open models (Kimi K3, Qwen3.8-2.4T) cost $2–$3 input on managed hosts.

What changed since August

Our August issue was written around Claude Opus 5, GPT-5.6 and an expected Sonnet 5 price rise. Since then:

Date (2026) Change What it means for your bill
Aug 13Gemini 3.7 Flash at $0.75/$3.75 introductory, half Gemini 3.6 Flash's $1.50/$7.50 launch priceFlash-tier work halved in price, but only until Dec 31
Aug 13–29Open weights for Qwen3.8-2.4T-A95B (Aug 13) and GLM-5.3 (Aug 29)More frontier-scale open options, priced at $1.40–$2.00 input on hosts
Aug 17DeepSeek introduces peak and off-peak pricing; V4 Pro rates rise by up to 1,100% on some token typesSame model, 2x price difference depending on the hour
Aug 21GPT-5.6 Sol cut from $5/$30 to $4/$20 for three months (reported)Temporary; budget for the old rate returning around late November
Aug–Sep 1Claude Sonnet 5's $2/$10 made standard; the $3/$15 increase due September 1 was cancelled. Claude Fable 5.1 at $10/$50 with $0.25 cache reads (Sep 1)The 50% Sonnet rise we flagged in August did not happen
Sep 2Gemini 3.8 Flash at $0.75/$3.75 introductory; Meta Muse Spark 1.3 at $1.25/$4.25, plus a $0.10/$0.20 contributor tier that lets Meta train on your promptsBetter Flash quality at the same introductory rate; a data-for-price trade at Meta
Sep 10–14DeepSeek V4.1 Flash replaces V4 Flash ($0.15/$0.60 off-peak); DeepSeek says deepseek-v4-pro requests route to V4.1 Flash from Sep 14 until V4.1 Pro launchesModel IDs you didn't change now serve a different model at a different rate
Sep 3GPT-6 Astra at $10/$50New OpenAI flagship at twice GPT-5.5's input price
Sep 21Grok 4.7 at $2/$6Cheapest output rate in the $2-input band
Sep 22Claude Opus 5.5 at $4/$20 ($0.20 cache reads); GPT-6 Sol and GPT-6 LunaPremium tier 20% cheaper per token; bulk tier halved
Sep 28Claude Sonnet 5.5 at $2/$10Sonnet upgrade at no price change
Sep 29GPT-6.1 Sol at $2/$10, $0.10 cached inputOpenAI says it nearly matches Astra on coding at about one-fifth the cost
Oct 6Mistral Large 4 in public preview at $1.36/$4.18; weights promised by end of OctoberA new low-priced large-model option to test

Corrections to our August issue: the Sonnet 5 increase to $3/$15 on September 1 was cancelled. We also listed Mistral Large 3 at roughly $2/$6 from market quotes; Mistral's own pricing page currently gives Mistral Large at $0.50/$1.50, and the new Mistral Large 4 preview is $1.36/$4.18.

What this means for your bill

Three patterns explain most of the October changes:

  1. If you pay for a premium model, you are probably overpaying by default. Opus 5.5 is 20% cheaper per token than Opus 5, and Anthropic says it costs about 40% less on typical workloads once fewer steps and tokens are counted. Teams still pinned to claude-opus-5 or GPT-5.6 Sol at $5/$30 are paying August prices for an older model.
  2. $2/$10 now covers most work that needed $5/$25 in August. Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol share that rate, and Grok 4.7 undercuts it on output at $6. See the coding stack for where each holds up on quality.
  3. The cheapest rates now come with conditions. Gemini's Flash rates end December 31, GPT-5.6 Sol's cut is temporary, and DeepSeek charges double during peak hours. If you set budgets from list prices without tracking these conditions, forecasts will be wrong.

Which swaps actually cut cost

Illustrative workload. 100M input tokens and 20M output tokens a month, roughly a customer-support assistant or an internal RAG feature at moderate scale. List rates as of October 6, 2026, with no caching or batch discount unless stated. Your bill will differ with retries, reasoning tokens and tokenizer differences.

The formula is the same as every issue:

Monthly cost = (input_tokens / 1M × input_rate) + (output_tokens / 1M × output_rate)
Swap Before (monthly) After (monthly) Saving The catch
Claude Opus 5 → Opus 5.5$500 + $500 = $1,000$400 + $400 = $80020% (more if Anthropic's ~40% workload claim holds for you)Thinking can't be turned off on Opus 5.5; set effort deliberately
Claude Opus 5.5 → Sonnet 5.5$800$200 + $200 = $40050%Opus still leads on judgement-heavy tasks; test on your own evals
GPT-5.6 Sol (August list, $5/$30) → GPT-6.1 Sol$500 + $600 = $1,100$40064% (50% against the $4/$20 promotional rate)A model migration, not a config change; re-run evals
GPT-5.6 Luna → GPT-6 Luna$20 + $24 = $44$10 + $10 = $2055%Small in dollars here; large at billions of tokens
Gemini 3.6 Flash (launch price) → Gemini 3.8 Flash$150 + $150 = $300$75 + $75 = $15050% until Dec 31; then $300 againGoogle says 3.8 Flash may use more tokens at higher effort
DeepSeek V4.1 Flash, peak → off-peak (DeepSeek API)$30 + $24 = $54$15 + $12 = $2750%Only for jobs you can schedule outside peak hours
Llama 3.3 70B, Together AI → DeepInfra$104 + $20.80 = $124.80$10 + $6.40 = $16.4087%Same weights; check throughput, rate limits and serving quantization

How to read this table: the first three rows are where most October savings sit, because premium and mid-tier models carry most production spend. A team that moves an Opus 5 workload to Sonnet 5.5 where quality allows goes from $1,000 to $400 a month on this volume, a 60% cut with no new vendor. The last two rows involve no model change at all. They come from when you run the work and which host serves it.

Swaps that look cheaper but aren't

  • Kimi K3 instead of a $2/$10 closed model. At $2.70–$3.00 input and $13.50–$15.00 output on managed hosts, Kimi K3 costs $540–$600 on this workload, compared with $400 for GPT-6.1 Sol or Sonnet 5.5. Choose open frontier weights for control or self-hosting, not to save money.
  • Gemini 3.8 Flash as a permanent budget line. It costs $150 on this workload until December 31 and $300 from January 1, 2027. At that rate it costs more than Claude Haiku 4.5 ($200) and far more than GPT-6 Luna ($20) or DeepSeek V4.1 Flash off-peak ($27).
  • Treating GPT-5.6 Sol's $4/$20 as the new list price. The cut was reported as a three-month promotion from August 21. If nothing changes, this workload goes back from $800 to $1,100.

Caching changes the ranking

Cache-read pricing now varies more than headline input pricing. If 80% of input is a cached prefix (common for RAG system prompts and agent loops), the same 100M/20M workload costs, before cache-write charges:

  • GPT-6.1 Sol: 20M × $2 + 80M × $0.10 + 20M × $10 = $248
  • Claude Sonnet 5.5: 20M × $2 + 80M × $0.20 + 20M × $10 = $256
  • Claude Opus 5.5: 20M × $4 + 80M × $0.20 + 20M × $20 = $496

Caching cuts Opus 5.5's input cost by about 76% but cuts the total bill by only 38%, because output costs stay the same. Output now accounts for most premium-tier spend, so the effort level you set matters as much as the rate card.

Price per task, not per token

Per-token tables hide the number that decides most model choices: what one completed task costs. Below is one illustrative agentic task, such as a scoped bug fix or a multi-step support resolution: 100K uncached input, 400K cache-read input, 30K output (reasoning included), with cache-write charges excluded.

Model Uncached input Cache reads Output Per task Per 1,000 tasks
GPT-6 Astra$1.00$0.40$1.50$2.90$2,900
Claude Fable 5.1$1.00$0.10$1.50$2.60$2,600
Claude Opus 5$0.50$0.20$0.75$1.45$1,450
Claude Opus 5.5$0.40$0.08$0.60$1.08$1,080
Claude Sonnet 5.5$0.20$0.08$0.30$0.58$580
Grok 4.7$0.20$0.20$0.18$0.58$580
GPT-6.1 Sol$0.20$0.04$0.30$0.54$540
Gemini 3.8 Flash (introductory)$0.075$0.03$0.11$0.22$218
GPT-6 Luna$0.01$0.004$0.015$0.03$29

The per-task view supports a recommendation you can test. If a $0.54 model needs a second attempt on more than half of your tasks, the $1.08 model that succeeds first time costs the same and takes less engineer time. Track success rate on first attempt alongside spend for two weeks before you standardise on the cheaper model. Models also differ in how many tokens they use for the same task (Anthropic's launch post says Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps), so measured per-task cost can differ widely from this same-token comparison.

How to read LLM pricing in October 2026

Rates are still quoted per million tokens, split into input and output. Five rules carry most of the cost math:

  1. Output costs 3–6x input, and reasoning bills as output. On Opus 5.5, Sonnet 5.5 and the Fable models thinking can't be switched off, only turned down.
  2. Cache-read multipliers are no longer uniform. Anthropic charges 0.1x base input on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1 and Mythos 5.1. OpenAI's GPT-6.1 Sol is 0.05x and GPT-6 Luna 0.1x.
  3. Long context has cliffs. GPT-6 Astra moves to $20/$75 in its long-context tier, Gemini 3.1 Pro to $4/$18 above 200K input, and Grok 4.7 to $4/$12 at 200K and above. Anthropic includes the full 1M context at standard rates on Claude 4.6 and later.
  4. Time and date now change price. DeepSeek's peak hours (01:00–04:00 and 06:00–10:00 UTC on weekdays) cost twice off-peak, and Google's introductory Flash rates double on January 1, 2027.
  5. Tokenizers differ. Claude 4.7 and later produce about 30% more tokens for the same text than earlier Claude models, per Anthropic, so compare per task, not per token, across vendors and generations.

Closed frontier APIs (October 2026)

List prices checked October 6, 2026, standard tier, short context.

Provider Model Input ($/1M) Output ($/1M) Notes
AnthropicClaude Mythos 5.1$10.00$50.00Limited availability (Project Glasswing); cache reads $0.25
AnthropicClaude Fable 5.1$10.00$50.00New Sep 1. Cache reads $0.25 (Fable 5: $1.00); batch $5/$25
AnthropicClaude Opus 5.5$4.00$20.00New Sep 22. Cache reads $0.20; batch $2/$10; fast mode $8/$40
AnthropicClaude Opus 5$5.00$25.00Cache reads $0.50; batch $2.50/$12.50
AnthropicClaude Sonnet 5.5$2.00$10.00New Sep 28. Cache reads $0.20; batch $1/$5
AnthropicClaude Sonnet 5$2.00$10.00Introductory rate made standard; $3/$15 increase cancelled
AnthropicClaude Haiku 4.5$1.00$5.00Cache reads $0.10; batch $0.50/$2.50
OpenAIGPT-6 Astra$10.00$50.00New Sep 3. Cached $1.00; long-context tier $20/$75
OpenAIGPT-6.1 Sol$2.00$10.00New Sep 29. Cached $0.10; long-context tier $4/$15
OpenAIGPT-6 Sol$2.00$10.00New Sep 22. Cached $0.20
OpenAIGPT-6 Luna$0.10$0.50New Sep 22. Cached $0.01
OpenAIGPT-5.6 Sol$4.00$20.00Cut from $5/$30 on Aug 21; reported as a three-month promotion
OpenAIGPT-5.6 Terra$2.00$12.00Unchanged since July 30 cut
OpenAIGPT-5.6 Luna$0.20$1.20Unchanged; GPT-6 Luna is half the price
OpenAIGPT-5.5 / GPT-5.5 Pro$5.00 / $30.00$30.00 / $180.00Prior flagship and highest-effort tier
GoogleGemini 3.8 Flash$0.75$3.75New Sep 2. Introductory to Dec 31; $1.50/$7.50 from Jan 1, 2027
GoogleGemini 3.7 / 3.6 Flash$0.75$3.75Same introductory terms listed on Google's pricing page
GoogleGemini 3.1 Pro (Preview)$2.00$12.00$4/$18 above 200K input; still the only Pro on the Gemini API
GoogleGemini 3.5 Flash-Lite$0.30$2.50Unchanged
GoogleGemini 3.1 Flash-Lite$0.25$1.50Cheapest Gemini text tier
xAIGrok 4.7$2.00$6.00New Sep 21. Cached $0.50; $4/$12 at 200K+ tokens; 500K context
xAIGrok 4.3$1.25$2.501M context; cached $0.20
MetaMuse Spark 1.3$1.25$4.25New Sep 2. Cached $0.15; contributor tier $0.10/$0.20 (Meta may train on your data; 100 RPM)
MistralMistral Large 4$1.36$4.18Public preview Oct 6; 1T total / 49B active; weights promised by end of October
MistralMistral Large (previous)$0.50$1.50Rate shown on Mistral's pricing page

Batch pricing is 50% off at Anthropic, OpenAI and Google. US-only inference on Claude 4.6 and later adds a 1.1x multiplier.

How to read the closed tier in October: the frontier band (GPT-6 Astra, Fable 5.1) is now 2.5x Opus 5.5 and 5x the $2 mid-tier, so it needs a task-level reason. The $2-input band is crowded with five strong models, and output rate ($6 to $12) is what separates them. At the bottom, GPT-6 Luna at $0.10/$0.50 is cheaper than every Gemini tier and than most hosted open models. For a vendor head-to-head, see OpenAI vs Anthropic pricing.

Open-weight models (October 2026)

Prices below are what managed hosts charge for each model, taken from each host's public pricing page on October 6. The weights are free; you pay for someone's hardware.

Model Input ($/1M) Output ($/1M) Where quoted / notes
Kimi K3$2.70–$3.00$13.50–$15.00Together $2.70/$13.50; DeepInfra $2.85/$14.25; Fireworks and Baseten $3/$15
Qwen3.8-2.4T-A95B (Qwen3.8 Max)$2.00$6.00Together and Fireworks. Weights published Aug 13; largest open-weight model to date
GLM-5.3$1.40$4.40Together, Fireworks, Baseten. Weights Aug 29 under a custom GLM-5.3 licence
DeepSeek V4 Pro$1.30–$1.32$2.60–$3.96DeepInfra $1.30/$2.60; Together and Baseten (V4 Pro 0813) $1.32/$3.96. On DeepSeek's own API, see the note below
Inkling$1.00$4.05Together AI
NVIDIA Nemotron 3 Ultra$0.60$2.40Baseten
DeepSeek V4.1 Flash$0.15–$0.30$0.60–$1.20DeepSeek API off-peak/peak; Fireworks, Together and Baseten $0.30/$1.20. Replaces V4 Flash on DeepSeek's API
GLM-5.3 Flash / Qwen3.8 Flash$0.15$0.47–$0.50Together, Fireworks, Baseten
GPT-OSS 120B$0.10–$0.15$0.50–$0.60Baseten $0.10/$0.50; Groq, Fireworks and Together $0.15/$0.60
DeepSeek V4 Flash (older checkpoint)$0.09–$0.13$0.18–$0.26DeepInfra $0.09/$0.18; Baseten V4-Flash-0731 $0.13/$0.26
Llama 3.3 70B$0.10–$1.04$0.32–$1.04DeepInfra $0.10/$0.32; Together $1.04/$1.04
Llama 3.1 8B$0.02–$0.14$0.04–$0.14DeepInfra $0.02/$0.04; Together (Lite) $0.14/$0.14

DeepSeek V4 Pro on DeepSeek's own API: the pricing page still lists V4 Pro at $1.32/$3.96 peak and $0.66/$1.98 off-peak, but DeepSeek's September 10 release note says deepseek-v4-pro requests route to V4.1 Flash at Flash rates from September 14 until V4.1 Pro launches. Check your invoice for which rate you are being billed. If you need V4 Pro itself, the managed hosts above still serve it.

Open weight does not always mean OSI open source. GLM-5.3 and several others ship under custom licences, so check the terms before you ship.

The October pattern: open-weight pricing now has two separate ends. At the top, Kimi K3, Qwen3.8 Max and GLM-5.3 cost $1.40–$3.00 input, close to or above the $2/$10 closed mid-tier. At the bottom, GPT-OSS, older DeepSeek Flash checkpoints and small Llama models still cost cents per million. DeepSeek's own API also stopped being the reliable price floor in August: its V4.1 Flash peak rate ($0.30/$1.20) is now higher than GPT-6 Luna ($0.10/$0.50).

Open-model hosting providers (October 2026)

The same weights cost different amounts depending on who serves them. The October examples below come from each host's pricing page.

Host Pricing model October example rates Best for
DeepInfraPer-token on its own stack; cached-input discountsDeepSeek V4 Pro $1.30/$2.60; Llama 3.3 70B $0.10/$0.32; Llama 3.1 8B $0.02/$0.04Lowest per-token price on most established open models
Fireworks AIStandard, Priority and Fast serving paths; batch 50% offKimi K3 $3/$15 (Fast $4.50/$22.50); GLM-5.3 $1.40/$4.40; V4.1 Flash $0.30/$1.20Production serving with speed tiers and US-hosted variants
Together AIFlat per-token serverless; also GPU rental and fine-tuningKimi K3 $2.70/$13.50; Qwen3.8-2.4T $2/$6; Inkling $1/$4.05Broadest catalog, including the newest large open models
BasetenPay-as-you-go Model APIs plus dedicated deploymentsGPT-OSS 120B $0.10/$0.50; GLM-5.3 $1.40/$4.40; Kimi K3 $3/$15Managed APIs with a path to dedicated capacity
GroqPer-token on LPU hardwareGPT-OSS 120B $0.15/$0.60 (~500 tok/s); GPT-OSS 20B $0.075/$0.30. Llama tiers now "contact sales"Latency-bound chat and tool loops
OpenRouterPass-through provider pricing, no inference markup; 5.5% fee on card credit purchasesUnderlying host rate plus credit feeOne API with routing and failover across hosts
Hugging Face Inference ProvidersPass-through provider rates, no HF markup; PRO includes $2/month compute creditsUnderlying host rateOne token across many hosts, with per-model, per-provider usage breakdowns

The host spread is still large: Llama 3.3 70B costs about 10x more on Together ($1.04/$1.04) than on DeepInfra ($0.10/$0.32). On newer models the spread is smaller. Kimi K3 varies by about 11% and GLM-5.3 is $1.40/$4.40 on all three hosts we checked, because fewer hosts serve them and those hosts tend to follow the lab's reference price. The cheapest host is not always the best value. A faster host can reduce total cost on agent loops that would otherwise wait on generation.

The pricing traps to watch

  • Expiring promotions. Gemini 3.6/3.7/3.8 Flash go from $0.75/$3.75 to $1.50/$7.50 on January 1, 2027. GPT-5.6 Sol's $4/$20 is reported to run three months from August 21. Put both dates in your forecast now.
  • Peak-hour pricing. DeepSeek's API costs twice as much between 01:00–04:00 and 06:00–10:00 UTC on weekdays. An interactive feature used in Asian business hours pays the peak rate on most of its traffic.
  • Default model drift. Tools and aliases move up a tier on upgrade. Claude Code's default is now Opus 5.5 on every first-party plan (Claude Code cost options), and API aliases change which model they point to.
  • Retired model IDs that still answer. DeepSeek serves deepseek-v4-flash requests with V4.1 Flash at the Flash price, and says deepseek-v4-pro requests have routed to V4.1 Flash since September 14. Your code doesn't change, but the model and the rate behind it do.
  • Data-for-price tiers. Meta's Muse Spark 1.3 contributor tier is 92% cheaper on input than standard because Meta may train on your prompts and completions. Check with legal before any customer data goes through it.
  • Long-context tiers. GPT-6 Astra, Grok 4.7 and Gemini 3.1 Pro all price long prompts higher. RAG pipelines whose retrieved context grows over time can cross the threshold without any code change.
  • Reasoning you can't switch off. On Opus 5.5, Sonnet 5.5 and the Fable models, thinking always bills as output. Effort is the control.

FAQ

What changed in LLM API pricing between August and October 2026?

The mid-tier improved without a price rise. Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol launched at $2/$10, and Claude Opus 5.5 at $4/$20, 20% below Opus 5. GPT-6 Luna halved the bulk tier to $0.10/$0.50. GPT-6 Astra and Claude Fable 5.1 hold the frontier at $10/$50. DeepSeek moved to peak and off-peak pricing on August 17, and the planned Claude Sonnet 5 increase to $3/$15 was cancelled.

What is the cheapest LLM API in October 2026?

Among closed APIs, GPT-6 Luna at $0.10 input and $0.50 output per 1M tokens. Among open-weight models on managed hosts, Llama 3.1 8B on DeepInfra ($0.02/$0.04) and older DeepSeek V4 Flash checkpoints ($0.09/$0.18 on DeepInfra) are lower still, but much less capable. For a capable general model, GPT-6 Luna and DeepSeek V4.1 Flash off-peak ($0.15/$0.60) are the practical floor.

Is Claude Opus 5.5 cheaper than Opus 5?

Yes. Opus 5.5 is $4/$20 per 1M tokens against Opus 5's $5/$25, and cache reads are $0.20 against $0.50. Anthropic says it costs about 40% less than Opus 5 on typical workloads once fewer steps and tokens are counted. On our 100M input / 20M output example, the list-price saving alone is $1,000 to $800 a month.

When does Gemini 3.8 Flash's introductory pricing end?

December 31, 2026. From January 1, 2027, Gemini 3.8 Flash costs $1.50 input and $7.50 output per 1M tokens, double the introductory $0.75/$3.75. Google's pricing page lists the same terms for Gemini 3.7 Flash and 3.6 Flash.

Is GPT-6.1 Sol or Claude Sonnet 5.5 cheaper?

Both list at $2/$10 per 1M tokens. GPT-6.1 Sol's cached input is $0.10 against Sonnet 5.5's $0.20 cache reads, so GPT-6.1 Sol is slightly cheaper on cache-heavy workloads: $248 against $256 a month in our 80%-cached example. Tokenizer differences and per-task efficiency usually matter more than that gap, so compare them on cost per completed task.

Track LLM spend across every provider

A list price tells you what a token costs; it can't tell you which model your code called yesterday, at what effort level, at what hour or through which host. Each of those changed price between August and October. A typical stack now spans one or two closed APIs, an open-model host and sometimes an aggregator, each invoicing separately.

StackSpend puts them on one timeline, with model-level breakdown, model recommendations that flag when a cheaper model such as Sonnet 5.5 is doing the same job as Opus 5.5, and same-day anomaly detection. For the overall picture, start with AI cost monitoring.

About The AI Stack Cost Report

The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs: frontier and open-weight LLMs, AI coding tools, image generation, and the hosting they run on. Each issue is compiled from primary pricing pages and release reporting, with source-check dates in the frontmatter. The October 2026 issue covers August through early October (there was no September issue) and has four parts: the frontier briefing, the pricing map, the coding stack, and the image stack.

Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily and the model changelog logs releases and deprecations by month.

References

Know where your cloud and AI spend stands — every day.

Connect providers in minutes. Get 90 days of visibility and start receiving daily cost updates before the invoice lands.

14-day free trial. No credit card required. Plans from $79/month.
LLM Pricing October 2026: Rates and Cheaper Swaps — StackSpend