| GPT-5.6 Sol | OpenAI | 2026-06-26 | OpenAI’s newest frontier model for reasoning and long-horizon agentic work (Sol/Terra/Luna family). In limited preview as of July 2026; GA in coming weeks. | Frontier reasoning/coding; new "max" reasoning effort and multi-subagent "ultra" mode. Terra ($2.50/$15) and Luna ($1/$6) add cheaper tiers. | Preview/limited access (~20 orgs) as of July 2026; Sol pricing $5/$30 per 1M tokens. | Closed |
| Claude Fable 5 | Anthropic | 2026-06-01 | Anthropic’s new "Mythos-class" tier positioned above the Opus lane for the hardest frontier work (June 2026). | Highest Anthropic capability tier for complex reasoning and agentic work. | Newest/most premium tier; overkill and costly for routine workloads. | Closed |
| Claude Opus 4.8 | Anthropic | 2026-05-28 | Anthropic’s current premium reasoning tier (replaces Opus 4.6/4.7). 1M-token context by default, adaptive thinking, optional fast mode. | High quality ceiling for repo-scale coding and planning. $5/$25 per 1M tokens; fast mode up to ~2.5× output throughput. | Easy to overspend if used as the default tier. | Closed |
| Gemini Omni | Google | 2026-05-20 | Google’s any-input/any-output multimodal family (starting with video); Omni Flash rolling out via the Gemini API. | Unified multimodal generation and editing across modalities. | Rolling out through mid-2026; availability varies by surface. | Closed |
| Gemini 3.5 Flash | Google | 2026-05-13 | Google’s current widely-available flagship-class Gemini; strongest agentic/coding Gemini yet. (Gemini 3.5 Pro rolling out later.) | Flagship-level quality at Flash speeds/price; outperforms Gemini 3.1 Pro on key benchmarks. | Pro tier still rolling out; managed-platform complexity vs simpler API products. | Closed |
| GPT-5.5 | OpenAI | 2026-04-24 | OpenAI’s current GA flagship (with GPT-5.5 Pro). 1M context. Default for complex production reasoning until GPT-5.6 reaches GA. | Frontier reasoning/coding, mature ecosystem. $5/$30 per 1M tokens (cached input $0.50); GPT-5.5 Pro $30/$180. | Premium pricing; >272K-token prompts billed at 2× input / 1.5× output. | Closed |
| DeepSeek-V4 Pro | DeepSeek | 2026-04-24 | DeepSeek’s current open-weight flagship (V4 Pro/Flash). MIT-licensed, 1M-token context. | Leads agentic coding among open models; ties the closed frontier on SWE-Bench. MIT license. | Large model; self-hosting/serving costs remain significant. | Open-weight |
| Grok 4.3 | xAI | 2026-04-17 | xAI’s current flagship (follows multi-agent Grok 4.20). 1M-token context, native video input, document generation. | Competitive agentic reasoning; native video + document outputs; real-time X integration. | Closed ecosystem; smaller enterprise footprint than some rivals. Grok 5 signalled but not yet shipped. | Closed |
| GPT-5.4 | OpenAI | 2026-03-05 | Earlier OpenAI flagship (superseded by GPT-5.5) with 1M context, native computer-use, and integrated Codex coding. | Frontier reasoning, coding, and multimodal. Adjustable reasoning modes including low-latency. | Superseded by GPT-5.5/5.6; premium pricing; vendor lock-in. | Closed |
| Gemini 3.1 Pro | Google | 2026-02-19 | Earlier Google flagship (superseded by the Gemini 3.5 line) with 1M context and 65K output. Strong reasoning, coding, and agentic workflows. | Topped 12 of 18 benchmarks at launch. 80.6% SWE-Bench, 85.9% BrowseComp. Three-tier thinking modes. | Superseded by Gemini 3.5 Flash/Pro; platform complexity vs simpler API products. | Closed |
| Claude Sonnet 4.6 | Anthropic | 2026-02-17 | Anthropic’s newer daily-use frontier tier for coding, agents, and professional work. | Very strong default for coding and agentic tasks. | Still premium versus many open-weight options. | Closed |
| Claude Opus 4.6 | Anthropic | 2026-02-05 | Earlier Anthropic premium reasoning tier; superseded by Opus 4.8. | High quality ceiling for complex reasoning and coding. | Superseded by Opus 4.8; easy to overspend if used as the default tier. | Closed |
| GPT-5.3-Codex | OpenAI | 2026-02-05 | OpenAI's coding-specialist model for software-engineering and agent workflows. Capabilities now in GPT-5.4. | Industry-leading coding. Cloud sandbox and agent tooling. | Access and pricing vary by product path. | Closed |
| Mistral Large 3 | Mistral | 2025-12-15 | 675B total, 41B active. Open-weight flagship with 256K context. Apache 2.0. | Strong multilingual and enterprise appeal. Multimodal. Competitive pricing. | Newer; fewer enterprise deployments than closed rivals. | Open-weight |
| Grok 4.1 | xAI | 2025-11-17 | Earlier xAI closed flagship; superseded by Grok 4.20/4.3. | Competitive frontier-style reasoning and agentic positioning. | Superseded by Grok 4.3; closed ecosystem and smaller enterprise footprint than some rivals. | Closed |
| Claude Haiku 4.5 | Anthropic | 2025-10-01 | Anthropic’s fast, low-cost tier for high-volume and latency-sensitive workloads. | Strong price/latency for routing, extraction, and high-throughput tasks. | Lower reasoning ceiling than Sonnet/Opus. | Closed |
| GPT-5.2 | OpenAI | 2025-09-30 | OpenAI's widely used tier with strong personality, creativity support, and 400K context. Default for many production workloads. | Good balance of speed, cost, and quality. Mature ecosystem. | Knowledge cutoff Sep 2024. Newer GPT-5.5 (GA) and GPT-5.6 (preview) available for harder tasks. | Closed |
| GPT-5-mini | OpenAI | 2025-09-30 | OpenAI's cost-optimized model for lower latency and budget-sensitive workloads. | Fast, cheaper than flagship. Supports vision and tools. | Lower reasoning ceiling than GPT-5.4 or GPT-5.2. | Closed |
| SmolLM3 | Hugging Face | 2025-07-08 | Hugging Face’s compact multilingual long-context model in the 3B class. | Practical small-model deployment with strong long-context and multilingual positioning. | Not a direct substitute for frontier-scale reasoning models. | Open-weight |
| Gemini 2.5 Flash | Google | 2025-06-17 | Google’s faster and usually cheaper production-tier member of the Gemini 2.5 family. | Better speed and price/performance balance than the flagship tier for many workloads. | Lower absolute reasoning ceiling than Gemini 2.5 Pro. | Closed |
| DeepSeek-V3.2 | DeepSeek | 2025-06-01 | DeepSeek Sparse Attention for long-context. High-compute variant rivals GPT-5 and Gemini 3 Pro. | Gold medal IMO/IOI 2025. Improved agent and tool-use. MIT-compatible. | Large model; self-hosting costs remain significant. | Open-weight |
| Qwen3 | Qwen | 2025-04-29 | Qwen’s newer family with hybrid thinking modes and larger MoE options. | Strong open-weight flexibility with modern reasoning features. | Deployment and benchmark interpretation still require hands-on evaluation. | Open-weight |
| Llama 4 Scout | Meta | 2025-04-05 | 17B active params, 10M context. Natively multimodal. Outperforms Gemma 3 and Gemini 2.0 Flash-Lite. | Fits on single H100. Strong efficiency for multimodal tasks. | Newer than Llama 3; ecosystem still maturing. | Open-weight |
| Llama 4 Maverick | Meta | 2025-04-05 | 17B active, 128 experts. Beats GPT-4o and Gemini 2.0 Flash; comparable to DeepSeek V3 at half the params. | Strong open-weight frontier alternative. Natively multimodal. | Newer; benchmark interpretation still evolving. | Open-weight |
| Gemini 2.5 Pro | Google | 2025-03-25 | Google’s thinking-oriented Gemini flagship for advanced reasoning and coding. | Strong reasoning, coding, and multimodal reach. | Managed platform complexity is higher than simpler API products. | Closed |
| DeepSeek-V3-0324 | DeepSeek | 2025-03-24 | A stronger post-trained update to the DeepSeek-V3 line with notable coding and reasoning improvements. | Very competitive open-weight quality with stronger coding-oriented tuning than the original V3 launch. | Still a large model family that is costly to self-serve at scale. | Open-weight |
| Gemma 3 | Google | 2025-03-10 | Google’s open-weight model family built to give developers a lighter-weight alternative to Gemini. | Open-weight flexibility with strong small-to-mid-size deployment options. | Not intended to fully replace top closed-model capability. | Open-weight |
| DeepSeek-R1 | DeepSeek | 2025-01-20 | DeepSeek’s reasoning-focused open-weight family that became central to the reasoning-model conversation. | Strong reasoning, math, code, and open-weight mindshare. | Reasoning-style usage can still become expensive to serve at scale. | Open-weight |
| DeepSeek-V3 | DeepSeek | 2024-12-27 | DeepSeek’s flagship open-weight general model with strong frontier comparisons. | Very competitive open-weight quality and strong coding reputation. | Large-scale serving remains non-trivial for most teams. | Open-weight |
| Llama 3.3 70B | Meta | 2024-12-19 | A more practical high-quality Llama tier for strong open-weight deployment without 400B scale. | Good balance of quality and deployment practicality. | Still requires careful infra planning compared with closed APIs. | Open-weight |
| Qwen2.5-Coder | Qwen | 2024-11-12 | Qwen’s open-weight coding family designed specifically for software engineering workloads. | Strong coder specialization with broad size options and open deployment flexibility. | Still requires model-selection discipline and self-hosting or hosted open-model trade-offs. | Open-weight |
| Qwen2.5 | Qwen | 2024-10-16 | Alibaba’s broad open-weight family spanning many sizes and specialist variants. | Very wide size range, multilingual coverage, strong coder variants. | Family breadth can create selection and governance complexity. | Open-weight |
| Llama 3.2 | Meta | 2024-09-25 | Meta’s extension of the Llama family into lighter text tiers and multimodal variants. | Broader deployment range across edge, lighter text, and multimodal use cases. | Not the strongest Llama option for highest-end reasoning. | Open-weight |
| Mistral Large 2 | Mistral | 2024-07-24 | Mistral’s flagship commercial model line for enterprise-grade reasoning and code. | Strong multilingual and coding capability with enterprise appeal. | Not a straightforward open-weight option despite Mistral’s open-model reputation. | Restricted / commercial |
| Llama 3.1 405B | Meta | 2024-07-23 | Meta’s flagship open-weight frontier-scale model in the Llama 3.1 family. | Strong open-weight quality at very large scale. | Inference and serving costs are high in practice. | Open-weight |
| Claude 3.5 Sonnet | Anthropic | 2024-06-20 | Anthropic’s widely adopted mid-tier model that became a coding and reasoning default. | Excellent repo reasoning, strong cost-to-quality balance. | Closed access and output costs still matter in heavy usage. | Closed |
| Codestral | Mistral | 2024-05-29 | Mistral’s code-generation-focused model family built specifically for software engineering tasks. | Clear coding specialization and strong relevance for code-heavy workflows. | Licensing and access are more constrained than a simple fully open-weight release. | Restricted / commercial |
| DeepSeek-V2 | DeepSeek | 2024-05-01 | DeepSeek’s MoE family that raised expectations around efficient open-weight inference. | Efficiency story and strong coding-oriented attention from developers. | Now mainly relevant as a history milestone, not the default latest recommendation. | Open-weight |
| Cohere Command R+ | Cohere | 2024-04-04 | Cohere’s business and retrieval-oriented flagship built for enterprise RAG-style workloads. | Strong fit for enterprise retrieval and internal-data use cases. | Less general market mindshare than OpenAI, Anthropic, or Google. | Closed |
| Grok-1 (open release) | xAI | 2024-03-17 | xAI’s early Grok base model released as open weights under Apache 2.0. | Useful historical example of a major lab releasing an open-weight base model. | Not the practical default choice for most current production teams. | Open-weight |
| Gemini 1.5 Pro | Google | 2024-02-15 | Google’s long-context milestone model in the Gemini family. | Very large context and good multimodal platform integration. | Pricing and platform paths can be harder to model clearly. | Closed |
| Mixtral 8x7B | Mistral | 2023-12-11 | Mistral’s sparse MoE model that helped popularize open MoE deployment economics. | Strong efficiency and good quality for its active-parameter profile. | Older than newer flagship open reasoning families. | Open-weight |
| Zephyr 7B Beta | Hugging Face H4 | 2023-10-25 | A Hugging Face alignment-focused fine-tune that showed how much post-training could improve smaller open models. | Strong instruct behavior for its size and highly instructive as a post-training milestone. | Small compared with current flagship open-weight families. | Open-weight |
| Llama 2 | Meta | 2023-07-18 | Meta’s family that accelerated mainstream adoption of open-weight LLMs. | Major ecosystem impact and broad downstream adaptation. | Now clearly behind newer open-weight families. | Open-weight |
| BLOOM | Hugging Face / BigScience | 2022-07-12 | A major multilingual open model built through the BigScience collaboration and distributed through Hugging Face. | Important multilingual history and open collaboration milestone. | No longer competitive with newer open-weight families on raw quality. | Open-weight |