Auto-updated monthly

New & deprecated LLM models

A monthly log of what launched and what's being retired across every major model provider — each with live API prices and benchmark scores, updated automatically.

This month · August 2026

Permalink →

New models — August 2026

4 models with a public release date in August 2026, joined to live list prices and benchmark scores.

ModelProviderInput / 1MOutput / 1MBest scoreReleased
zai-org/GLM-5.3-FlashTogether AI$0.15$0.5090.2%20 Aug 2026
google/gemini-3.7-flashDeepInfra$0.75$3.7594.8%13 Aug 2026
grok-4.6xAI$2.00$6.0094.0%12 Aug 2026
Qwen/Qwen3.8-MaxDeepInfra$1.65$4.9592.7%2 Aug 2026

Deprecated & sunsetting — August 2026

15 models with an upstream retirement or sunset date in August 2026. Plan migrations before the date to avoid broken calls.

ModelProviderInput / 1MOutput / 1MBest scoreSunset date
claude-opus-4-1Anthropic$15.00$75.0073.3%Retired5 Aug 2026
gpt-5.2-chat-latestOpenAI$1.75$14.0073.8%Retired10 Aug 2026
gpt-5.3-chat-latestOpenAI$1.75$14.00Retired10 Aug 2026
llama-3.1-8b-instantGroq$0.05$0.0827.0%Retired16 Aug 2026
llama-3.3-70b-versatileGroq$0.59$0.7947.4%Retired16 Aug 2026
google/gemma-3n-E4B-itTogether AI$0.06$0.12Retired25 Aug 2026
meta-llama/Llama-Guard-4-12BTogether AI$0.20$0.20Retired25 Aug 2026
deepseek-ai/DeepSeek-V4-ProTogether AI$1.74$3.4877.6%Retired27 Aug 2026
moonshotai/Kimi-K2.7-CodeTogether AI$0.95$4.0087.9%Retired27 Aug 2026
nvidia/nemotron-3-ultra-550b-a55bTogether AI$0.60$3.60Retired27 Aug 2026
pearl-ai/gemma-4-31b-itTogether AI$0.28$0.8675.8%Retired27 Aug 2026
gemini-robotics-er-1.6-previewGoogle (Gemini)$1.00$5.00Sunsetting31 Aug 2026
mistral-medium-2505Mistral$0.40$2.0059.5%Sunsetting31 Aug 2026
mistral-medium-2508Mistral$0.40$2.0059.5%Sunsetting31 Aug 2026
mistral-medium-3-1-2508Mistral$0.40$2.00Sunsetting31 Aug 2026

Browse by month

Prices from the LLM API Pricing Index, synced daily. Release & benchmark data via Epoch AI (CC BY); deprecation dates from provider metadata. When a new model beats one you use, model recommendations surface it automatically at your token mix.

Track the models you actually use — not a list.

StackSpend connects read-only to your AI providers, flags deprecations before they break your calls, and surfaces cheaper equal-quality swaps from this same live dataset.