Guides
August 2, 2026
By Andrew Day

The Latest in AI Image Generation, August 2026: Models, Quality, and Cost

A research-grounded August 2026 guide to AI image models: FLUX 3 goes multimodal, Microsoft's MAI-Image-2.5-Pro arrives token-priced, Nano Banana and GPT Image 2 hold the arenas — and why per-image pricing is quietly being replaced by per-token billing.

Share this post

Send it to someone managing cloud or AI spend.

LinkedInX

The AI Stack Cost Report · August 2026 · Part 4: The Image Stack. The meter moved again this month. This is StackSpend's monthly read on what the modern AI engineering stack actually costs. Also in this issue: the frontier briefing, the pricing map, and the coding stack.

Use this when you need a current, factual map of the AI image-generation field in August 2026 — which model leads for which job, what each one actually costs, and why "cost per image" is getting harder to calculate.

The fast answer: the leaderboard is broadly stable — GPT Image 2 leads prompt adherence and blind-preference arenas, Google's Nano Banana Pro leads image editing with Nano Banana 2 covering high-volume work from ~$0.045/image, Midjourney V8 owns aesthetics, and Imagen 4 spans $0.02–$0.06 across Fast/Standard/Ultra. What changed in July is structural, not competitive. Black Forest Labs' FLUX 3 (July 23) collapsed image, video, audio, and robot-action prediction into one multimodal model — but shipped video-first in gated early access with no public pricing, and the image tier still hadn't launched. And Microsoft's MAI-Image-2.5-Pro (July 23) is priced per token, not per image — $5/1M text input, $8/1M image input, $106/1M image output. Per-image price cards are quietly being replaced by token meters, and that breaks the forecasting method most teams use.

Image generation stopped being a one-model race in 2026. Different models win different jobs — photorealism, artistic style, text rendering, editing, speed — and the same output can cost 100x more on one provider than another. This guide covers the current leaders, what they cost, and the operational reality of tracking image spend at scale.

For the text-model side, see the latest in LLMs for August 2026 and the LLM model pricing guide.

Quick answer

For an August 2026 snapshot:

  • Best all-round / prompt adherence: GPT Image 2 (OpenAI) — top of blind-preference arenas, strongest on complex multi-part instructions and in-image text.
  • Best editing: Google Nano Banana Pro, which leads the image-editing arenas (~$0.134 per 1K/2K image, ~$0.24 at 4K).
  • Best high-volume general work: Nano Banana 2 — from ~$0.045/image at 512px to ~$0.151 at 4K, with the Batch API halving every tier.
  • Best aesthetics / art direction: Midjourney V8 (V8.1/V8.2 since March 2026).
  • Best photorealism in a cloud stack: Google Imagen 4 — Fast $0.02, Standard $0.04, Ultra $0.06 per image.
  • New, and worth watching rather than adopting: FLUX 3 (July 23) — multimodal image/video/audio/robotics, gated early access, no public pricing, image tier not yet shipped. MAI-Image-2.5-Pro (July 23) — Microsoft's highest-fidelity image model, public preview, token-priced.
  • Specialists: Ideogram (accurate text in images), Recraft (design systems + speed), Grok Imagine (xAI), Seedream 4.5 (ByteDance), Hunyuan Image 3.0 (Tencent). Stable Diffusion (SDXL / SD 3.5) remains the cheapest hackable base at ~$0.003.

What changed in July 2026

FLUX 3 (Black Forest Labs, July 23) is the most ambitious release of the month and the least usable today. It is a single multimodal flow model trained jointly on images, video, and audio that generates video with native audio up to 20 seconds, synthesizes and edits images, and predicts robot actions. It launched video-first in gated early access — API and private weights to selected partners on application — and as of late July, FLUX 3 Image had not launched and no pricing had been announced for any tier. If your pipeline runs on FLUX.2, nothing needs to change yet; plan for a re-benchmark rather than a migration.

MAI-Image-2.5-Pro (Microsoft, July 23) matters less for its quality claims than for its meter. Microsoft's highest-fidelity image model entered public preview in Microsoft Foundry priced at $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens — a token-based structure, not a per-image one, targeted at hero imagery, detailed editing, and precise in-image text rendering.

That is the trend worth planning around. GPT Image 2 is already token-metered, and MAI follows. Per-image price cards ($0.04, $0.055) are becoming an approximation layered over a token meter whose actual consumption depends on resolution, input images, and edit complexity. Two consequences:

  1. "Images × unit price" stops predicting your bill. A 4K edit with three reference images consumes vastly more than a 512px generation, on the same rate card.
  2. Image spend starts behaving like LLM spend — variable per call, sensitive to prompt and resolution, and only truly knowable after the fact. Which means it needs the same instrumentation.

The current leaders (August 2026)

GPT Image 2 (OpenAI) — the image model in ChatGPT and the API since April 2026. Tops blind-preference arenas and is strongest at complex, multi-constraint prompts. Token-priced; at equivalent high resolutions it is the premium option, working out around $0.10 per high-quality 1024×1024 image and $0.40+ at the top settings.

Google Nano Banana Pro / Nano Banana 2 — Nano Banana Pro leads the editing arenas at roughly $0.134 per 1K/2K image and $0.24 at 4K. Nano Banana 2 is the volume workhorse: ~$0.045 at 512px up to ~$0.151 at 4K, with the Batch API halving every tier — a floor near $0.022 per image.

Google Imagen 4 — top-tier photorealism inside the Google stack, with the clearest per-image tiering of any major provider: Fast $0.02 / Standard $0.04 / Ultra $0.06.

Midjourney V8 — still the aesthetic and art-direction leader, with ~5x faster rendering than V7, native 2K output, HD mode, and Draft mode for rapid iteration.

FLUX.2 (Black Forest Labs) — the open-weight workhorse until FLUX 3 Image ships: Pro ($0.055, camera-accurate optics), Dev/Flex ($0.025, self-hostable), Schnell (~$0.015, commercial-friendly).

Specialists and challengers — Ideogram for reliable text rendering; Recraft for design-system consistency and speed; xAI's Grok Imagine; Reve/Riverflow 2.0 Pro for conversational refinement; and Seedream 4.5 ($0.035) and Hunyuan Image 3.0 ($0.030) as strong quality-per-dollar entrants.

What image generation costs (August 2026)

Tier Examples Price Best for
Premium (token-metered) GPT Image 2; MAI-Image-2.5-Pro ($5/$8/$106 per 1M text-in / image-in / image-out tokens) ~$0.10–$0.40+ per image equivalent Top quality and prompt adherence; cost varies with resolution and inputs
Premium (per-image) Nano Banana Pro (~$0.134 at 1K/2K, ~$0.24 at 4K); Midjourney API ~$0.13–$0.24 Editing quality and art direction with predictable unit cost
Volume managed Nano Banana 2 (~$0.045–$0.151); Imagen 4 Fast/Standard/Ultra ($0.02/$0.04/$0.06) ~$0.02–$0.15 The default for most production image features; batch halves it
Open-weight (direct API) FLUX.2 Pro (~$0.055), Dev (~$0.025), Schnell (~$0.015); Ideogram, Recraft, SD 3.5 ~$0.015–$0.10 Control and predictable economics; self-host to cut further
Chinese frontier Seedream 4.5 (~$0.035), Hunyuan Image 3.0 (~$0.030) ~$0.03–$0.04 Strong quality-per-dollar challengers
Hosted aggregators FAL, Replicate, Together AI, Fireworks, Stability API ~$0.008–$0.04 Cheapest access to open-weight models; slight feature lag
Cheapest base Stability SDXL ~$0.003 High-volume, quality-tolerant workloads
Not yet priced FLUX 3 (early access, video-first) Evaluate later; no public pricing or general image tier as of late July

Rates are representative list prices checked on 2 August 2026; resolution, quality mode, and batch settings move all of them. The spread is the story: the same broad job can cost $0.003 on SDXL and $0.40+ on a premium model at high resolution — well over 100x. And as with open text models, the same open-weight image model costs different amounts depending on whether you call it directly or through an aggregator — the dynamic covered in the open-model hosting section of the pricing map.

How to choose an image model

  • Marketing/brand visuals where look matters most: Midjourney V8 for aesthetics; GPT Image 2 when the brief is complex and literal.
  • Editing and iterative refinement: Nano Banana Pro leads the editing arenas; Reve/Riverflow for conversational refinement.
  • Product/photoreal at scale: Imagen 4 (clear per-image tiers) or Nano Banana 2; FLUX.2 Pro when you need camera-accurate optics or on-prem control.
  • Text inside the image (ads, posters, UI): Ideogram remains the most reliable; MAI-Image-2.5-Pro targets this explicitly.
  • High volume, cost-sensitive, quality-tolerant: open-weight models (FLUX Dev/Schnell, SD 3.5) via an aggregator or self-hosted, or Nano Banana 2 on the Batch API.
  • Design-system consistency and speed: Recraft.
  • Video with native audio: FLUX 3 is the notable new entrant — but it is gated early access with no published pricing, so treat it as a research item.

Benchmark two or three candidates on your prompts and acceptance bar before committing — arena rankings are directional, not a substitute for your own eval.

The cost trap: cheap per image, expensive at scale

An image call looks trivially cheap — pennies. That is exactly why image spend gets away from teams:

  • Volume multiplies fast. At $0.04/image, a feature generating 10,000 images a day is $12,000/month — and image features encourage retries and variations, each a billable call.
  • Resolution and quality modes change the rate 3–8x. Nano Banana 2 spans $0.045 to $0.151 on resolution alone; Imagen 4 triples from Fast to Ultra.
  • Token metering hides the unit cost entirely. On GPT Image 2 and MAI-Image-2.5-Pro there is no per-image price — only tokens consumed, which vary per request. At $106 per 1M image-output tokens, a shift toward higher-resolution outputs is a large cost change with no rate-card change.
  • Retries and variations are silent multipliers. A pipeline producing four variants per request quietly 4x's the bill.
  • Multi-provider sprawl. Midjourney for hero images, an aggregator for bulk, a premium API for edits — three invoices, three units, no combined view.

None of this shows on a per-image price card; it shows on the invoice. The fix is to track image-generation spend by model and workload the same way you track LLM tokens. Connect your providers to StackSpend for a unified view of AI spend with anomaly detection and pace-to-forecast alerts, plus attribution by model, feature, and team, so a runaway generation job or a jump to a pricier tier is caught the day it starts. See AI cost monitoring.

FAQ

What is the best AI image generator in August 2026?

There is no single winner. GPT Image 2 leads blind-preference arenas and complex-prompt adherence; Google's Nano Banana Pro leads image editing; Midjourney V8 leads on aesthetics; Imagen 4 and Nano Banana 2 are the value picks for photoreal work at scale; FLUX.2 leads the open-weight tier. Pick by the specific job — realism, art direction, text rendering, editing, or cost — and benchmark on your own prompts.

How much does AI image generation cost per image?

In August 2026: Imagen 4 is $0.02 (Fast) / $0.04 (Standard) / $0.06 (Ultra); Nano Banana 2 runs ~$0.045 at 512px to ~$0.151 at 4K, halved on the Batch API; Nano Banana Pro is ~$0.134 at 1K/2K and ~$0.24 at 4K; FLUX.2 is ~$0.015–$0.055 by variant; hosted aggregators run ~$0.008–$0.04; and Stability SDXL is around $0.003. GPT Image 2 and MAI-Image-2.5-Pro are token-metered rather than per-image, working out roughly $0.10–$0.40+ per image depending on resolution and inputs.

What is FLUX 3 and can I use it yet?

FLUX 3, launched by Black Forest Labs on July 23, 2026, is a single multimodal flow model trained jointly on images, video, and audio — generating video with native audio up to 20 seconds, synthesizing and editing images, and predicting robot actions. It shipped video-first in gated early access via API and private weights to selected partners on application. As of late July 2026 the FLUX 3 Image tier had not launched and no public pricing had been announced for any tier, so FLUX.2 remains the production choice.

Why is image pricing moving from per-image to per-token?

Because modern image models are multimodal systems that consume text prompts, reference images, and edit instructions and emit variable amounts of image data. Microsoft's MAI-Image-2.5-Pro is priced at $5/1M text input tokens, $8/1M image input tokens, and $106/1M image output tokens; GPT Image 2 is likewise token-metered. The practical effect is that "number of images × unit price" no longer forecasts your bill — resolution, reference images, and edit complexity all move the cost per call, so you need observed spend rather than an estimate.

What is the cheapest way to generate images at scale?

Run an open-weight model (FLUX Schnell/Dev or Stable Diffusion) self-hosted or through a hosted aggregator such as FAL, Replicate, Together AI, or Fireworks, where per-image costs fall to roughly $0.008–$0.04 — or use Nano Banana 2 on the Batch API, which halves every tier to a floor near $0.022. Reserve premium models for images where quality materially matters, and monitor total spend so retries and high-resolution modes don't erase the savings.

How do I track AI image generation spend?

Treat it like token spend: attribute cost by model, feature, and provider, and watch for anomalies. This matters more in August than it did in July, because token-metered image models remove the per-image unit price you would otherwise budget against. StackSpend connects your image and LLM providers into one daily view with model-level breakdown and alerting — start with AI cost monitoring or cloud + AI cost monitoring if you also run infrastructure spend.

About The AI Stack Cost Report

The AI Stack Cost Report is StackSpend's monthly briefing on what the modern AI engineering stack costs — frontier and open-weight LLMs, AI coding tools, image generation, and the hosting and infrastructure they run on. Each issue is compiled from primary pricing pages and release reporting, with source-check dates listed at the top of every article. The August 2026 issue has four parts: the frontier briefing, the pricing map, the coding stack, and the image stack.

Between issues, the LLM API Pricing Index re-syncs published per-1M-token rates daily and the model changelog logs every release and deprecation by month.

A monthly snapshot tells you what the market charges; it can't tell you what your stack spent yesterday. StackSpend does — per-provider AI spend, anomaly alerts the day a spike happens, and a daily report in Slack. Setup is read-only, takes about five minutes, and starts with a free 14-day trial. Why we publish it: the meter came back.

Related reading

References

Share this post

Send it to someone managing cloud or AI spend.

LinkedInX

Know where your cloud and AI spend stands — every day.

Connect providers in minutes. Get 90 days of visibility and start receiving daily cost updates before the invoice lands.

14-day free trial. No credit card required. Plans from $29/month.
AI Image Generation Models (August 2026) — StackSpend Blog