The AI Stack Cost Report · October 2026 · The Coding Stack. StackSpend's monthly read on what the modern AI engineering stack actually costs. There was no September issue, so this edition covers everything that changed between early August and October 6, 2026.
Use this when you lead an engineering team and need a current answer to "which coding model, for which job, in which tool, on which billing model".
The fast answer: September brought Claude Fable 5.1 (Sept 1), Gemini 3.8 Flash (Sept 2), GPT-6 Astra (Sept 3), Grok 4.7 (Sept 21), Claude Opus 5.5 (Sept 22), Claude Sonnet 5.5 (Sept 28) and GPT-6.1 Sol (Sept 29). The mid-tier caught up: at $2/$10, Sonnet 5.5 and GPT-6.1 Sol now do work that needed a $5/$25 model in August. The routing system is unchanged (plan with a strong model, execute with a cheap one, review with a different vendor), but Claude Code moved its default model up a tier for Pro and Team Standard seats, so routing now means overriding defaults on purpose.
For raw scorecards, see AI coding models in 2026: strengths, weaknesses, and pricing. This article is about composing them into a team practice.
Quick answer
- Best default for hard coding work: Claude Opus 5.5 at $4/$20 per 1M tokens. 20% cheaper per token than Opus 5; Anthropic says about 40% cheaper on typical workloads at default settings.
- Best value for well-scoped work: Claude Sonnet 5.5 at $2/$10. Anthropic reports 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5's 66.4%.
- Best cross-vendor counterweight: GPT-6.1 Sol at $2/$10 with $0.10 cached input. OpenAI says it matches GPT-6 Astra on DeepSWE v1.1 at about one-fifth of the cost per task.
- Frontier tier, used sparingly: GPT-6 Astra and Claude Fable 5.1, both $10/$50. Escalation only.
- Cheap execution: Gemini 3.8 Flash ($0.75/$3.75 introductory, through Dec 31), GPT-6 Luna ($0.10/$0.50), DeepSeek V4.1 Flash ($0.15/$0.60 off-peak) and Grok 4.7 ($2/$6).
- The action item: Claude Code's
defaultis now Opus 5.5 on every first-party plan. Pin the model and effort level you want.
What changed since August
Our August issue was written around Claude Opus 5, GPT-5.6 Sol and Grok 4.5. Since then:
- Anthropic replaced most of its coding lineup. Fable 5.1 kept Fable 5's $10/$50 rate but cut cache reads to $0.25. Opus 5.5 dropped to $4/$20 with $0.20 cache reads. Sonnet 5.5 launched at $2/$10.
- Sonnet 5's price did not go up. The introductory $2/$10 rate became permanent, and the planned move to $3/$15 on September 1 did not happen.
- OpenAI moved to GPT-6. Astra is the flagship at $10/$50. GPT-6 Sol ($2/$10), GPT-6.1 Sol ($2/$10, $0.10 cached) and GPT-6 Luna ($0.10/$0.50) fill the tiers below it. GPT-5.6 Sol is still listed, now at $4/$20.
- Benchmarks moved generations. Vendors now headline Terminal-Bench 4.0, CursorBench 4.0 and DeepSWE v1.1 instead of SWE-bench Verified and Terminal-Bench 2.1, so August and October scores are not comparable.
- Claude Code defaults changed. Before v2.1.280,
defaultresolved to Sonnet 5 on Pro and Team Standard. It now resolves to Opus 5.5 on every first-party plan, at a default effort ofmedium. - Open-weight pricing got more complicated. DeepSeek raised V4 API prices in August and now charges peak and off-peak rates. V4.1 Flash costs $0.15/$0.60 off-peak and double that at peak.
The coding-model leaderboard (October 2026)
This is a point-in-time snapshot as of October 6, 2026. Prices are list API rates per 1M tokens, without batch or caching discounts.
| Model | Tier / role | Input–Output ($/1M) | Where it fits for code |
|---|---|---|---|
| Claude Opus 5.5 | Premium daily default | $4 / $20 | Judgement-heavy multi-file work, migrations, planning; 5-level effort dial; $0.20 cache reads |
| Claude Sonnet 5.5 | Workhorse, best value | $2 / $10 | Well-scoped bug fixes and features; leads Anthropic's Terminal-Bench 4.0 table |
| GPT-6.1 Sol | Workhorse / reviewer | $2 / $10 | Near-Astra agentic coding at about one-fifth the cost; strong independent reviewer of Claude output |
| GPT-6 Astra | Frontier | $10 / $50 | Hardest long-horizon tasks; 74% on DeepSWE v1.1 per OpenAI |
| Claude Fable 5.1 | Frontier | $10 / $50 | Long-running autonomous work; $0.25 cache reads soften agentic costs |
| Grok 4.7 | Mid-tier agentic | $2 / $6 | Shipped in Cursor first; 71.0% DeepSWE v1.1 per xAI; price doubles above 200K-token prompts |
| Gemini 3.8 Flash | Cheap execution | $0.75 / $3.75* | Google's "best reasoning and coding model yet" at Flash prices |
| Claude Haiku 4.5 | Fast sub-agent | $1 / $5 | File navigation, cheap fan-out inside Claude Code |
| GPT-6 Luna | Bulk execution | $0.10 / $0.50 | High-volume, fixed-plan codegen and mechanical edits |
| DeepSeek V4.1 Flash | Open-weight bulk | $0.15 / $0.60 off-peak | Cheapest capable option if your workload tolerates peak-hour pricing and data-residency trade-offs |
*Gemini 3.8 Flash's $0.75/$3.75 is introductory pricing through December 31, 2026. From January 1, 2027 it becomes $1.50/$7.50.
How to read this table: the frontier-to-bulk spread is now 100x on input ($10 vs $0.10), and most of the quality gains landed in the mid-tier. A team that standardised on a $5/$25 model in August can get equal or better coding results for $2/$10 today, but only by changing its routing on purpose.
Vendor-reported coding benchmarks
These are numbers each vendor published for its own models. Harnesses, effort levels and scaffolds differ between vendors, so compare within a row with care and across vendors with suspicion.
| Model | Terminal-Bench 4.0 | CursorBench 4.0 | DeepSWE v1.1 |
|---|---|---|---|
| Claude Sonnet 5.5 | 70.6% | 55.5% | — |
| Claude Opus 5.5 | 66.4% | 57.8% | — |
| Claude Mythos 5.1 (limited access) | 60.9% | — | — |
| GPT-6 Astra | 57.9% | — | 74% |
| Claude Fable 5.1 | 55.8% | 51.8% | — |
| Claude Opus 5 (August's leader) | 52.3% | 46.6% | — |
| Grok 4.7 | 37.6% | 46.3% | 71.0% |
| GPT-5.6 Sol | 37.3% | — | — |
Two things stand out. First, the cheaper Claude model wins the terminal benchmark: Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0, while Opus 5.5 still leads on FrontierCode (54.4% vs 46.2%) and CursorBench. "Sonnet for execution, Opus for judgement" holds, with a narrower gap. Second, OpenAI's mid-tier claim rests on DeepSWE: OpenAI says GPT-6.1 Sol matches Astra there at about one-fifth of the per-task cost. We could not find a primary-source Terminal-Bench 4.0 score for GPT-6.1 Sol, so it is left out rather than estimated.
Use different models for different jobs
Give each phase of the work to the cheapest model that does that phase well:
| Phase | What you need | Good picks (October 2026) | Why |
|---|---|---|---|
| Plan / architect | Deepest judgement, whole-system context | Opus 5.5 (high effort), GPT-6 Astra or Fable 5.1 for the hardest cases | Planning errors are the most expensive kind; this is where frontier tokens pay back |
| Execute / implement | Fast, reliable, follows a spec | Sonnet 5.5, GPT-6.1 Sol, Grok 4.7 | All three sit at $2 input and post strong agentic scores |
| High-volume subtasks | Cheap, parallel, mechanical | GPT-6 Luna, Gemini 3.8 Flash, Haiku 4.5, DeepSeek V4.1 Flash | Boilerplate, file navigation, sub-agent fan-out |
| Review / verify | Independent judgement from a different vendor | GPT-6.1 Sol reviewing Claude output, or Opus 5.5 reviewing OpenAI output | A reviewer trained like the author shares the author's blind spots |
The one real change since August is in the reviewer slot: GPT-6.1 Sol is cheap enough at $2/$10 to review every PR a Claude-based agent opens, not a sample.
A falsifiable recommendation
For teams on Claude Code, set the team default to Sonnet 5.5 and escalate to Opus 5.5 explicitly. Our prediction: on a team where most sessions are well-scoped bug fixes and small features, two weeks on a Sonnet 5.5 default will cut model cost per merged PR by roughly 40–50% without moving merged-PR rate or revert rate. If your revert rate climbs more than a couple of points, the prediction is wrong for your codebase, and Opus 5.5 should stay the default.
The trade-off that depends on context
Fable 5.1 / Astra vs Opus 5.5 is a cost-per-task question, not a per-token one. Both frontier models cost 2.5x Opus 5.5 per token. On a long autonomous migration, a frontier model that finishes in one attempt can still be cheaper overall than a mid-tier model that needs three attempts and an engineer's afternoon to untangle. Fable 5.1's $0.25 cache reads also shrink the gap on cache-heavy agent loops. On short interactive tasks the opposite is true: per-token price dominates, and frontier models are the wrong choice. The answer depends on task length, cache share and engineer time, which is why per-task spend matters more than any per-token table.
Effort settings and defaults: the hidden cost lever
Claude Opus 5.5, Sonnet 5.5 and GPT-6.1 Sol all expose effort levels (low through max). Higher effort means more reasoning tokens, so the same model at max can cost several times what it costs at low for the same prompt. Check which effort level any benchmark you rely on was run at.
Claude Code's current behaviour, per its model-configuration docs:
defaultresolves to Opus 5.5 on Pro, Max, Team, Enterprise and the Anthropic API.- Opus 5.5 and Sonnet 5.5 default to
mediumeffort. Most other models default tohigh. - Before v2.1.280, Pro and Team Standard seats defaulted to Sonnet 5.
That last point matters for budgeting: Opus 5.5 costs twice Sonnet 5's per-token rate. On API billing, a version bump doubled the rate on every default request. On subscriptions, expect plan limits to be reached sooner and more usage credits bought at API rates. For plan limits and overage, see Claude Code cost options in October 2026 and the earlier August overage guide.
Claude Code vs Codex vs Cursor vs Copilot in October
Every one of these tools now sells access to models from several vendors.
| Tool | What changed since August | Pricing shape (Oct 6) | Watch-out |
|---|---|---|---|
| Claude Code | Opus 5.5 is now the default model; Sonnet 5.5 available | Pro $20; Max from $100; Team Standard $25, Premium $125 per seat; usage credits at API rates | Default moved up a tier; pin it |
| OpenAI Codex | GPT-6 Astra and GPT-6.1 Sol added in September | Included in ChatGPT plans (Plus $20, Pro $100–$500, Business $20/user); paid credits beyond limits | Astra credits cost 100x Luna's per input token; Fast mode costs 2x |
| Cursor | Grok 4.7 launched in Cursor; Projects (coordinator agents, Sept 10); self-hosted machines; Rollouts and Security Review bots | Individual from $20; Teams $40/user; Enterprise pooled usage | Cloud agents and subagent fan-out can multiply token use per task |
| GitHub Copilot | Opus 5.5, Sonnet 5.5, GPT-6.1 Sol and Gemini 3.8 Flash in the model picker | Pro $10, Pro+ $39, Max $100, each with a flex allotment | Check how each model draws down the flex allotment |
The trend since August is consolidation of models and fragmentation of bills: Opus 5.5 runs inside Claude Code, Cursor and Copilot, each with its own billing unit. For how the pricing families compare, see per-seat or per-token AI coding tool pricing. For what Cursor's usage looks like broken out per model, see Cursor token usage by model.
Subscriptions vs API: the economics in October
| Subscriptions / seats | API (pay-per-token) | |
|---|---|---|
| Examples | Claude Pro/Max/Team, ChatGPT Plus/Pro/Business (Codex), Cursor, Copilot | Anthropic, OpenAI, Google, xAI and DeepSeek per-1M-token rates |
| Best for | Interactive, human-driven coding many times a day | CI agents, batch refactors, your own harness, product features |
| October shift | Defaults moved to pricier models, so limits are reached sooner | Mid-tier now $2/$10 with strong results; cache reads as low as $0.10–$0.20 |
| Risk | Unused seats, and silent overage via usage credits | Looping agents and peak-hour pricing with no monthly ceiling |
The rule of thumb holds: if a person drives the tool, a subscription usually wins; if code calls the model, the API usually wins. But every major subscription now offers a way to keep working past its limits (usage credits, purchased credits, flex allotments or usage-based billing), so a seat is no longer a cost ceiling. For a per-engineer planning number that includes seats, overflow and agents, see AI coding tool budget per engineer in 2026.
The part everyone underestimates: the combined bill
A typical October setup after following the routing advice above: Cursor seats, some engineers on Claude Code Max, Codex through ChatGPT Business, a Sonnet 5.5 CI agent on the Anthropic API, and GPT-6.1 Sol reviewing PRs through the OpenAI API. That is five billing surfaces in four units (seats, credits, flex allotments, tokens), and nobody can see the total or the figure per developer.
The likely failures are quiet: a tool update that moves the default up a tier, usage credits switched on, a review bot at max effort. None shows up until each invoice arrives separately.
StackSpend puts these on one timeline, broken out by provider, model, tool and person. See AI coding tool cost monitoring for the cross-tool view. Cursor cost monitoring and Claude Code cost analysis cover per-developer, per-model usage. Model recommendations flags when a cheaper model, such as Sonnet 5.5 in place of Opus 5.5, is doing the same job. For the process side, see how to monitor AI developer spend by user, team, and provider.
Practical takeaway
- Re-baseline your routing. August's choices are now too expensive at the mid-tier. Sonnet 5.5 and GPT-6.1 Sol at $2/$10 belong in your execution tier.
- Pin defaults and effort. Set the model and effort level explicitly in Claude Code and Cursor team settings rather than inheriting them.
- Make review cross-vendor and universal. At $2/$10, there is no cost reason to skip independent review on agent-written PRs.
- Reserve the $10/$50 tier for long-horizon work where cost per task, not per token, justifies it.
- Measure cost per merged PR, per developer, per model. That is the only number that makes routing decisions testable.
FAQ
What is the best coding model in October 2026?
For most engineering teams, Claude Opus 5.5 (released September 22, 2026, at $4/$20 per 1M tokens) is the strongest default for judgement-heavy coding. Claude Sonnet 5.5 ($2/$10) is the best value for well-scoped work and beats Opus 5.5 on Anthropic's Terminal-Bench 4.0 result (70.6% vs 66.4%). GPT-6.1 Sol ($2/$10) is the strongest alternative from another vendor.
Is Claude Sonnet 5.5 better than Opus 5.5 for coding?
It depends on the task. Sonnet 5.5 scores higher on Terminal-Bench 4.0, and Opus 5.5 scores higher on FrontierCode 1.1 (54.4% vs 46.2%) and CursorBench 4.0 (57.8% vs 55.5%), per Anthropic. Use Sonnet 5.5 for well-scoped execution at half the per-token price, and Opus 5.5 for planning, ambiguous multi-file changes and review.
Is GPT-6 Astra worth $10/$50 for coding?
Only for the hardest long-horizon tasks. OpenAI reports that GPT-6.1 Sol matches Astra on DeepSWE v1.1 at roughly one-fifth of the cost per task, so Astra should be an escalation path rather than a default.
What changed in Claude Code's default model?
Claude Code's default setting now resolves to Opus 5.5 on Pro, Max, Team, Enterprise and Anthropic API accounts, at medium effort. Before v2.1.280, Pro and Team Standard defaulted to Sonnet 5. Opus 5.5's per-token price is twice Sonnet 5's, so pin the model explicitly if you want to control cost.
How do I see what my team spends across Cursor, Claude Code, Codex and the APIs?
Each tool bills in its own unit and on its own invoice, so the combined figure has to be assembled. Pull provider and tool usage into one view broken out by person and model, as described in AI coding tool cost monitoring, and track cost per merged PR rather than per seat.
About The AI Stack Cost Report
The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs, compiled from primary vendor pricing and release pages with source-check dates in the frontmatter. This issue follows the August 2026 issue; there was no September issue.
Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily, and the model changelog logs releases and deprecations by month.
A monthly snapshot can't tell you what your own stack spent yesterday. StackSpend shows per-provider and per-tool AI spend with same-day anomaly alerts; setup is read-only.
Related reading
- The Best Coding Models in August 2026, the previous issue
- Claude Code Cost Options in October 2026
- AI Coding Tool Budget per Engineer in 2026
- Per-Seat or Per-Token? AI Coding Tool Pricing
- Cursor Token Usage by Model
- Claude Code Cost Options and Overage (August 2026)
- Cursor vs Windsurf vs Claude Code vs Antigravity: Which Agentic IDE Wins in 2026?
References
- Anthropic — Pricing (Claude Platform docs), checked 2026-10-06
- Anthropic — Introducing Claude Opus 5.5
- Anthropic — Introducing Claude Sonnet 5.5
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Claude Code Docs — Model configuration (defaults and effort)
- Claude — Plans and pricing
- OpenAI — API pricing
- OpenAI — GPT-6 Astra
- OpenAI Developer Community — Introducing GPT-6 Astra (Sept 3, 2026)
- OpenAI — Introducing GPT-6.1 Sol
- OpenAI — Using GPT-6 (latest model guide)
- The Next Web — OpenAI releases GPT-6.1 Sol at a fifth of GPT-6 Astra's token prices
- IT Brief UK — OpenAI launches GPT-6.1 Sol as cheaper coding model
- ChatGPT Learn — Codex pricing and plans
- ChatGPT Learn — Codex changelog
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- Google — Gemini API pricing
- xAI — Grok 4.7
- xAI — Models and pricing
- DeepSeek — Models and pricing
- Caixin Global — DeepSeek launches V4 Pro and raises API prices (Aug 14, 2026)
- Cursor — Changelog
- Cursor — Pricing
- GitHub — Copilot plans

