AI Model Glossary

A researched directory of notable LLMs and coding models across major providers, with release dates, strengths, weaknesses, and open-versus-closed status.

Model landscape reviewed July 2026 — newest models listed first.

This page is designed as a practical reference for teams evaluating model choices. Release dates refer to public launch dates for a model family or notable variant, not every patch release. The open vs closed column is about weight access: some models are fully closed APIs, some are open-weight, and some sit in a more restricted commercial middle ground. Entries are shown with the most recent launches first so the page works as a current-market reference, not just a history list.

Looking for live cost? See the LLM API pricing comparison — per-token prices across providers and inference hosts, with coding-benchmark scores, updated daily. Already running these models? StackSpend's model recommendations check your usage against this market daily and flag cheaper equal-or-better swaps.

Topic hubs

Use the glossary as an entry point into broader model topics

The glossary captures the market map. The topic hubs below branch into practical buying, implementation, and tooling questions so the page feeds the wider model SEO cluster instead of standing alone.

Provider history

How the major model families evolved

The current model landscape makes more sense when you see how providers moved: closed labs expanded into coding and multimodal workflows, while open-weight ecosystems like Llama, DeepSeek, Qwen, and Hugging Face-origin models kept raising the floor on what open models can do.

OpenAI

OpenAI set the pace for mainstream API-delivered frontier models, then extended into multimodal and coding-agent workflows. GPT-4o and GPT-4.1 were retired from ChatGPT in February 2026. GPT-5.5 is the current GA flagship; the GPT-5.6 family (Sol/Terra/Luna) is in limited preview as of July 2026.

2026-06-26
GPT-5.6 (Sol / Terra / Luna)

Limited preview (July 2026, ~20 orgs; GA in coming weeks). Sol = frontier reasoning/agentic ($5/$30), Terra = balanced ($2.50/$15), Luna = fastest/cheapest ($1/$6) per 1M tokens. Adds a "max" reasoning effort and multi-subagent "ultra" mode.

2026-04-24
GPT-5.5 / GPT-5.5 Pro

Current GA flagship. 1M context. GPT-5.5 $5/$30; GPT-5.5 Pro $30/$180 per 1M tokens; cached input $0.50.

2026-03-05
GPT-5.4

Earlier flagship with 1M context, native computer-use, and integrated Codex coding; superseded by GPT-5.5.

2026-02-05
GPT-5.3-Codex

Specialized coding agent; capabilities integrated into later GPT-5.x flagships.

2025-09-30
GPT-5.2

Widely adopted tier with strong personality and creative support. Knowledge cutoff Sep 2024.

2025-04-14
GPT-4.1

Retired Feb 2026. API-first coding and long-context family.

2024-05-13
GPT-4o

Retired Feb 2026. Introduced omni multimodal direction.

Anthropic

Anthropic became a default choice for repo-scale reasoning and long-context development work, especially in terminal-first workflows. The current lineup is Fable 5 (Mythos-class), Opus 4.8, Sonnet 4.6, and Haiku 4.5.

2026-06-01
Claude Fable 5

June 2026: a new "Mythos-class" tier positioned above the Opus lane for the hardest frontier work.

2026-05-28
Claude Opus 4.8

Current premium reasoning tier. 1M-token context by default, adaptive thinking, optional fast mode. $5/$25 per 1M tokens.

2026-02-17
Claude Sonnet 4.6

Current daily-use default; coding quality close to Opus at a lower-cost tier.

2026-02-05
Claude Opus 4.6

Earlier premium tier; superseded by Opus 4.7 and 4.8.

2025-10-01
Claude Haiku 4.5

Fast, low-cost tier for high-volume and latency-sensitive workloads.

2024-06-20
Claude 3.5 Sonnet

Marked a major quality and speed step-change for coding and general reasoning.

Google

Google split its model story between Gemini as the managed frontier family and Gemma as the open-weight developer path. Gemini 3.5 Flash leads the current line, with Gemini 3.5 Pro and the multimodal Gemini Omni rolling out through mid-2026.

2026-05-20
Gemini Omni

New any-input/any-output multimodal family (starting with video); Omni Flash rolling out to developers via the Gemini API.

2026-05-13
Gemini 3.5 Flash

Current widely-available flagship-class model; strongest agentic/coding Gemini yet, outperforming Gemini 3.1 Pro on key benchmarks. Gemini 3.5 Pro rolling out later.

2026-02-19
Gemini 3.1 Pro

Earlier flagship with 1M context; topped SWE-Bench and LiveCodeBench at launch. Superseded by the Gemini 3.5 line.

2025-06-17
Gemini 2.5 Flash

Extended the Gemini 2.5 line into a faster, more cost-sensitive production tier.

2025-03-25
Gemini 2.5 Pro

Brought stronger thinking and coding performance into the Gemini family.

2025-03-10
Gemma 3

Extended Google’s open-weight line with stronger multimodal and long-context capabilities.

2024-02-15
Gemini 1.5 Pro

Made very long context a practical buying criterion for production teams.

DeepSeek

DeepSeek became one of the most important open-weight challengers by combining strong efficiency claims with serious reasoning and coding performance. The current line is DeepSeek-V4 (Pro and Flash).

2026-04-24
DeepSeek-V4 (Pro / Flash)

Current open-weight flagship line: MIT-licensed, 1M-token context. V4 Pro leads agentic coding and ties the closed frontier on SWE-Bench.

2025-06-01
DeepSeek-V3.2

DSA attention for long-context. High-compute variant rivalled GPT-5 and Gemini 3 Pro at launch.

2025-03-24
DeepSeek-V3-0324

Updated the V3 line with stronger post-training and coding-oriented improvements.

2025-01-20
DeepSeek-R1

Made DeepSeek central to the open reasoning-model discussion.

2024-12-27
DeepSeek-V3

Showed that an open-weight MoE flagship could compete more directly with frontier closed models.

2024-05-01
DeepSeek-V2

Raised attention around efficient MoE design and lower-cost inference economics.

xAI

xAI positioned Grok as both a consumer-facing assistant brand and a developer/API family, with one notable open-weight release in Grok-1. Grok 4.3 is the current flagship; Grok 5 has been signalled but not yet shipped as of July 2026.

2026-04-17
Grok 4.3

Current flagship: 1M-token context, native video input, document generation, improved tool-calling. Follows the multi-agent Grok 4.20 (March 2026, 2M context).

2025-11-17
Grok 4.1

Earlier closed-model API phase with stronger agentic reasoning claims; superseded by Grok 4.20/4.3.

2024-03-17
Grok-1

Released as an Apache 2.0 open-weight base model after the initial beta period.

Mistral

Mistral mixed open-weight momentum with commercial model lines, especially for European buyers and code-focused teams.

2025-12-15
Mistral Large 3

675B total, 41B active. Open-weight, 256K context. Apache 2.0.

2024-07-24
Mistral Large 2

Strengthened Mistral’s enterprise and flagship positioning against larger US labs.

2024-05-29
Codestral

Marked Mistral’s explicit push into code-generation models.

2023-12-11
Mixtral 8x7B

Helped make sparse MoE models a core part of the open-weight conversation.

Hugging Face ecosystem

Hugging Face matters less as one single frontier family and more as the home of influential open-model releases, community fine-tunes, and distribution.

2025-07-08
SmolLM3

Made the case for compact, long-context, multilingual open models in the 3B range.

2023-10-25
Zephyr 7B Beta

Showed how alignment and post-training could make smaller open models highly competitive.

2022-07-12
BLOOM

One of the most important early multilingual open LLM efforts.

Qwen

Qwen grew into one of the most important open-weight ecosystems because it shipped broad size ranges, coder variants, and multilingual coverage quickly.

2025-04-29
Qwen3

Extended the line with hybrid thinking and larger MoE options.

2024-11-12
Qwen2.5-Coder

Strengthened Qwen’s position as a serious open-weight coding family.

2024-10-16
Qwen2.5

Expanded the family into a broad open-weight lineup spanning general, coder, and math variants.

Meta

Meta’s Llama family normalized the idea that top-tier open-weight models could shape the mainstream model market.

2025-04-05
Llama 4 Scout

17B active params, 10M context. Natively multimodal. Fits on single H100.

2025-04-05
Llama 4 Maverick

17B active, 128 experts. Beats GPT-4o and Gemini 2.0 Flash; comparable to DeepSeek V3 at half the params.

2024-12-19
Llama 3.3

Improved the practical 70B class for teams that needed stronger open-weight reasoning without 400B-scale costs.

2024-09-25
Llama 3.2

Extended the family into lighter text tiers and multimodal vision variants.

2024-07-23
Llama 3.1

Added the 405B tier and pushed open-weight models further into frontier territory.

2023-07-18
Llama 2

Established Llama as the most important mainstream open-weight family of its era.

Cohere

Cohere stayed especially relevant for enterprise retrieval and business-document workflows rather than general consumer-assistant mindshare.

2024-04-04
Command R+

Positioned Cohere as a serious enterprise retrieval and business-workflow model vendor.

Glossary

Notable AI models and coding models

This list is curated rather than exhaustive. It aims to cover the model families most useful for teams making real API, coding, open-weight, and platform decisions, including notable Hugging Face ecosystem entries, recent coding-model variants, and major open-weight challengers like DeepSeek and Qwen.

ModelProviderRelease dateDescriptionStrengthsWeaknessesOpen vs closed
GPT-5.6 SolOpenAI2026-06-26OpenAI’s newest frontier model for reasoning and long-horizon agentic work (Sol/Terra/Luna family). In limited preview as of July 2026; GA in coming weeks.Frontier reasoning/coding; new "max" reasoning effort and multi-subagent "ultra" mode. Terra ($2.50/$15) and Luna ($1/$6) add cheaper tiers.Preview/limited access (~20 orgs) as of July 2026; Sol pricing $5/$30 per 1M tokens.Closed
Claude Fable 5Anthropic2026-06-01Anthropic’s new "Mythos-class" tier positioned above the Opus lane for the hardest frontier work (June 2026).Highest Anthropic capability tier for complex reasoning and agentic work.Newest/most premium tier; overkill and costly for routine workloads.Closed
Claude Opus 4.8Anthropic2026-05-28Anthropic’s current premium reasoning tier (replaces Opus 4.6/4.7). 1M-token context by default, adaptive thinking, optional fast mode.High quality ceiling for repo-scale coding and planning. $5/$25 per 1M tokens; fast mode up to ~2.5× output throughput.Easy to overspend if used as the default tier.Closed
Gemini OmniGoogle2026-05-20Google’s any-input/any-output multimodal family (starting with video); Omni Flash rolling out via the Gemini API.Unified multimodal generation and editing across modalities.Rolling out through mid-2026; availability varies by surface.Closed
Gemini 3.5 FlashGoogle2026-05-13Google’s current widely-available flagship-class Gemini; strongest agentic/coding Gemini yet. (Gemini 3.5 Pro rolling out later.)Flagship-level quality at Flash speeds/price; outperforms Gemini 3.1 Pro on key benchmarks.Pro tier still rolling out; managed-platform complexity vs simpler API products.Closed
GPT-5.5OpenAI2026-04-24OpenAI’s current GA flagship (with GPT-5.5 Pro). 1M context. Default for complex production reasoning until GPT-5.6 reaches GA.Frontier reasoning/coding, mature ecosystem. $5/$30 per 1M tokens (cached input $0.50); GPT-5.5 Pro $30/$180.Premium pricing; >272K-token prompts billed at 2× input / 1.5× output.Closed
DeepSeek-V4 ProDeepSeek2026-04-24DeepSeek’s current open-weight flagship (V4 Pro/Flash). MIT-licensed, 1M-token context.Leads agentic coding among open models; ties the closed frontier on SWE-Bench. MIT license.Large model; self-hosting/serving costs remain significant.Open-weight
Grok 4.3xAI2026-04-17xAI’s current flagship (follows multi-agent Grok 4.20). 1M-token context, native video input, document generation.Competitive agentic reasoning; native video + document outputs; real-time X integration.Closed ecosystem; smaller enterprise footprint than some rivals. Grok 5 signalled but not yet shipped.Closed
GPT-5.4OpenAI2026-03-05Earlier OpenAI flagship (superseded by GPT-5.5) with 1M context, native computer-use, and integrated Codex coding.Frontier reasoning, coding, and multimodal. Adjustable reasoning modes including low-latency.Superseded by GPT-5.5/5.6; premium pricing; vendor lock-in.Closed
Gemini 3.1 ProGoogle2026-02-19Earlier Google flagship (superseded by the Gemini 3.5 line) with 1M context and 65K output. Strong reasoning, coding, and agentic workflows.Topped 12 of 18 benchmarks at launch. 80.6% SWE-Bench, 85.9% BrowseComp. Three-tier thinking modes.Superseded by Gemini 3.5 Flash/Pro; platform complexity vs simpler API products.Closed
Claude Sonnet 4.6Anthropic2026-02-17Anthropic’s newer daily-use frontier tier for coding, agents, and professional work.Very strong default for coding and agentic tasks.Still premium versus many open-weight options.Closed
Claude Opus 4.6Anthropic2026-02-05Earlier Anthropic premium reasoning tier; superseded by Opus 4.8.High quality ceiling for complex reasoning and coding.Superseded by Opus 4.8; easy to overspend if used as the default tier.Closed
GPT-5.3-CodexOpenAI2026-02-05OpenAI's coding-specialist model for software-engineering and agent workflows. Capabilities now in GPT-5.4.Industry-leading coding. Cloud sandbox and agent tooling.Access and pricing vary by product path.Closed
Mistral Large 3Mistral2025-12-15675B total, 41B active. Open-weight flagship with 256K context. Apache 2.0.Strong multilingual and enterprise appeal. Multimodal. Competitive pricing.Newer; fewer enterprise deployments than closed rivals.Open-weight
Grok 4.1xAI2025-11-17Earlier xAI closed flagship; superseded by Grok 4.20/4.3.Competitive frontier-style reasoning and agentic positioning.Superseded by Grok 4.3; closed ecosystem and smaller enterprise footprint than some rivals.Closed
Claude Haiku 4.5Anthropic2025-10-01Anthropic’s fast, low-cost tier for high-volume and latency-sensitive workloads.Strong price/latency for routing, extraction, and high-throughput tasks.Lower reasoning ceiling than Sonnet/Opus.Closed
GPT-5.2OpenAI2025-09-30OpenAI's widely used tier with strong personality, creativity support, and 400K context. Default for many production workloads.Good balance of speed, cost, and quality. Mature ecosystem.Knowledge cutoff Sep 2024. Newer GPT-5.5 (GA) and GPT-5.6 (preview) available for harder tasks.Closed
GPT-5-miniOpenAI2025-09-30OpenAI's cost-optimized model for lower latency and budget-sensitive workloads.Fast, cheaper than flagship. Supports vision and tools.Lower reasoning ceiling than GPT-5.4 or GPT-5.2.Closed
SmolLM3Hugging Face2025-07-08Hugging Face’s compact multilingual long-context model in the 3B class.Practical small-model deployment with strong long-context and multilingual positioning.Not a direct substitute for frontier-scale reasoning models.Open-weight
Gemini 2.5 FlashGoogle2025-06-17Google’s faster and usually cheaper production-tier member of the Gemini 2.5 family.Better speed and price/performance balance than the flagship tier for many workloads.Lower absolute reasoning ceiling than Gemini 2.5 Pro.Closed
DeepSeek-V3.2DeepSeek2025-06-01DeepSeek Sparse Attention for long-context. High-compute variant rivals GPT-5 and Gemini 3 Pro.Gold medal IMO/IOI 2025. Improved agent and tool-use. MIT-compatible.Large model; self-hosting costs remain significant.Open-weight
Qwen3Qwen2025-04-29Qwen’s newer family with hybrid thinking modes and larger MoE options.Strong open-weight flexibility with modern reasoning features.Deployment and benchmark interpretation still require hands-on evaluation.Open-weight
Llama 4 ScoutMeta2025-04-0517B active params, 10M context. Natively multimodal. Outperforms Gemma 3 and Gemini 2.0 Flash-Lite.Fits on single H100. Strong efficiency for multimodal tasks.Newer than Llama 3; ecosystem still maturing.Open-weight
Llama 4 MaverickMeta2025-04-0517B active, 128 experts. Beats GPT-4o and Gemini 2.0 Flash; comparable to DeepSeek V3 at half the params.Strong open-weight frontier alternative. Natively multimodal.Newer; benchmark interpretation still evolving.Open-weight
Gemini 2.5 ProGoogle2025-03-25Google’s thinking-oriented Gemini flagship for advanced reasoning and coding.Strong reasoning, coding, and multimodal reach.Managed platform complexity is higher than simpler API products.Closed
DeepSeek-V3-0324DeepSeek2025-03-24A stronger post-trained update to the DeepSeek-V3 line with notable coding and reasoning improvements.Very competitive open-weight quality with stronger coding-oriented tuning than the original V3 launch.Still a large model family that is costly to self-serve at scale.Open-weight
Gemma 3Google2025-03-10Google’s open-weight model family built to give developers a lighter-weight alternative to Gemini.Open-weight flexibility with strong small-to-mid-size deployment options.Not intended to fully replace top closed-model capability.Open-weight
DeepSeek-R1DeepSeek2025-01-20DeepSeek’s reasoning-focused open-weight family that became central to the reasoning-model conversation.Strong reasoning, math, code, and open-weight mindshare.Reasoning-style usage can still become expensive to serve at scale.Open-weight
DeepSeek-V3DeepSeek2024-12-27DeepSeek’s flagship open-weight general model with strong frontier comparisons.Very competitive open-weight quality and strong coding reputation.Large-scale serving remains non-trivial for most teams.Open-weight
Llama 3.3 70BMeta2024-12-19A more practical high-quality Llama tier for strong open-weight deployment without 400B scale.Good balance of quality and deployment practicality.Still requires careful infra planning compared with closed APIs.Open-weight
Qwen2.5-CoderQwen2024-11-12Qwen’s open-weight coding family designed specifically for software engineering workloads.Strong coder specialization with broad size options and open deployment flexibility.Still requires model-selection discipline and self-hosting or hosted open-model trade-offs.Open-weight
Qwen2.5Qwen2024-10-16Alibaba’s broad open-weight family spanning many sizes and specialist variants.Very wide size range, multilingual coverage, strong coder variants.Family breadth can create selection and governance complexity.Open-weight
Llama 3.2Meta2024-09-25Meta’s extension of the Llama family into lighter text tiers and multimodal variants.Broader deployment range across edge, lighter text, and multimodal use cases.Not the strongest Llama option for highest-end reasoning.Open-weight
Mistral Large 2Mistral2024-07-24Mistral’s flagship commercial model line for enterprise-grade reasoning and code.Strong multilingual and coding capability with enterprise appeal.Not a straightforward open-weight option despite Mistral’s open-model reputation.Restricted / commercial
Llama 3.1 405BMeta2024-07-23Meta’s flagship open-weight frontier-scale model in the Llama 3.1 family.Strong open-weight quality at very large scale.Inference and serving costs are high in practice.Open-weight
Claude 3.5 SonnetAnthropic2024-06-20Anthropic’s widely adopted mid-tier model that became a coding and reasoning default.Excellent repo reasoning, strong cost-to-quality balance.Closed access and output costs still matter in heavy usage.Closed
CodestralMistral2024-05-29Mistral’s code-generation-focused model family built specifically for software engineering tasks.Clear coding specialization and strong relevance for code-heavy workflows.Licensing and access are more constrained than a simple fully open-weight release.Restricted / commercial
DeepSeek-V2DeepSeek2024-05-01DeepSeek’s MoE family that raised expectations around efficient open-weight inference.Efficiency story and strong coding-oriented attention from developers.Now mainly relevant as a history milestone, not the default latest recommendation.Open-weight
Cohere Command R+Cohere2024-04-04Cohere’s business and retrieval-oriented flagship built for enterprise RAG-style workloads.Strong fit for enterprise retrieval and internal-data use cases.Less general market mindshare than OpenAI, Anthropic, or Google.Closed
Grok-1 (open release)xAI2024-03-17xAI’s early Grok base model released as open weights under Apache 2.0.Useful historical example of a major lab releasing an open-weight base model.Not the practical default choice for most current production teams.Open-weight
Gemini 1.5 ProGoogle2024-02-15Google’s long-context milestone model in the Gemini family.Very large context and good multimodal platform integration.Pricing and platform paths can be harder to model clearly.Closed
Mixtral 8x7BMistral2023-12-11Mistral’s sparse MoE model that helped popularize open MoE deployment economics.Strong efficiency and good quality for its active-parameter profile.Older than newer flagship open reasoning families.Open-weight
Zephyr 7B BetaHugging Face H42023-10-25A Hugging Face alignment-focused fine-tune that showed how much post-training could improve smaller open models.Strong instruct behavior for its size and highly instructive as a post-training milestone.Small compared with current flagship open-weight families.Open-weight
Llama 2Meta2023-07-18Meta’s family that accelerated mainstream adoption of open-weight LLMs.Major ecosystem impact and broad downstream adaptation.Now clearly behind newer open-weight families.Open-weight
BLOOMHugging Face / BigScience2022-07-12A major multilingual open model built through the BigScience collaboration and distributed through Hugging Face.Important multilingual history and open collaboration milestone.No longer competitive with newer open-weight families on raw quality.Open-weight

Know where your cloud and AI spend stands — every day.

Connect providers in minutes. Get 90 days of visibility and start receiving daily cost updates before the invoice lands.

14-day free trial. No credit card required. Plans from $29/month.
AI Model Glossary — StackSpend