GPT Image 2.5 Review 2026: OpenAI's New Image Model Is Already #1 — But Is It Actually Better?
Sunburst and Flare top the leading blind-preference leaderboards for both generation and editing. The more interesting story is what OpenAI actually changed under the hood — and what independent testing found when it checked the claims.
GPT Image 2.5 review 2026: OpenAI released ChatGPT Images 2.5 on September 8, 2026, alongside two new API models — GPT Image 2.5 Flare and GPT Image 2.5 Sunburst. Within days, both had climbed to the top of the two leaderboards most people in this space actually trust: Artificial Analysis' blind-preference Image Arena and LMArena's Text-to-Image rankings. That's a genuinely strong result. It's also not the most interesting part of the story.
The more useful question isn't whether GPT Image 2.5 tops a leaderboard — it does, by a real margin. It's what OpenAI actually changed to get there, whether that translates to real-world editing work, and where independent testing and early user reports diverge from the marketing claims. This review is built from OpenAI's official documentation and system card, Artificial Analysis and LMArena leaderboard data, and independently verified third-party testing — not a claim of hands-on testing of the model itself unless stated otherwise.
Yes, It's #1 — But Read the Fine Print
On Artificial Analysis' Text to Image leaderboard, GPT Image 2.5 Sunburst (max) currently leads with an Elo score of 1197 across more than 13,400 blind head-to-head comparisons — ahead of Flare (max) at 1190, and a real step above the previous GPT Image 2 (high) at 1171. On the Image Editing leaderboard, the pattern repeats: Sunburst at 1180, Flare at 1160, both clearly ahead of the next closest competitors. These rankings come from blind votes — users compare two images generated from the same prompt without knowing which model made which — which makes them harder to game than a self-reported benchmark.
LMArena tells a similar story with one important caveat: as of this review, Sunburst and Flare's LMArena scores are still marked preliminary, and their vote counts (in the low thousands) are dramatically smaller than GPT Image 2's established sample of roughly 78,000 votes on the same leaderboard. That doesn't mean the ranking is wrong — early scores often hold up — but it does mean the Arena result should be read as "leading so far, with less data behind it" rather than a fully settled verdict. The accurate summary: GPT Image 2.5 currently tops both major blind-preference leaderboards, with Artificial Analysis' result backed by a large sample and LMArena's still accumulating votes.
Flare vs. Sunburst: Two Models, Different Jobs
It's easy to assume "2.5" is one upgraded model, but OpenAI shipped two API models with genuinely different jobs, and mixing them up is the most common source of confusion in early coverage.
Flare is the speed-first default. OpenAI positions it for everyday high-quality generation — social content, product experiences, visual search, high-volume generation — and recommends it as the default choice for most applications. Sunburst is precision-first, built for situations where editing accuracy matters more than generation speed: creative production, campaign assets, and polished product imagery. Both accept the same inputs, support the same quality settings, and share identical token pricing — the difference is purely in what each is optimized for, not a tiered "better model costs more" structure.
What's Actually New
Reference Fidelity
OpenAI says 2.5 better preserves a person's identity, distinctive details, lighting, textures, and composition when generating from a reference image — directly relevant for product photography, character consistency, and branding work.
Precision Editing
Instead of regenerating an entire image, 2.5 is built to change one specific element — a product, a background, a line of copy — while preserving everything else, including existing brand treatment.
Multi-Turn Consistency
OpenAI says edits hold up better across long conversations, so a workflow of generate → edit → edit → polish doesn't reset progress with each new instruction the way earlier versions sometimes did.
Sketch-to-Image and Comment-Based Editing
You can draw a rough sketch directly in ChatGPT and use it as a visual reference, or place comments on specific spots in an image ("make this darker," "change this to blue") for localized edits without re-describing the whole scene.
Speed
OpenAI reports up to 50% lower latency than Images 2.0. Independent spot tests found real but inconsistent gains — one test measured Flare at roughly 12 seconds versus 36 seconds for GPT Image 2, while another measured larger absolute times for both but a similarly wide gap — a reminder that testing environment and settings affect the numbers significantly.
What Independent Testing Actually Found
This is the section most GPT Image 2.5 coverage skips, and it's the one that separates a genuine review from a rewritten press release. Independent research published in September specifically tested 2.5's editing claims on forgery and detail-preservation tasks, comparing Flare and Sunburst against a freshly re-run GPT Image 2 in the same week.
The results were mixed rather than a clean win. 2.5 preserved text surrounding edited regions better in some tests, and product codes came out more legible. But the researchers found no clear improvement in target-field correctness — the specific thing being edited didn't consistently come out more accurate — and fine-print improvements remained inconclusive because of OCR limitations in the test setup itself. Separately, early Reddit reaction after launch was genuinely split: some users praised the output quality, while others reported specific regressions in hands, a more processed or noisy look in some generations, and background complexity that occasionally looked worse than GPT Image 2.
None of this contradicts the leaderboard results — a model can win blind aesthetic preference votes while still having specific, task-dependent weaknesses that a targeted forensic test or a demanding user would notice. The honest summary: OpenAI's precision-editing claims hold up in some conditions and are genuinely unresolved in others, and the benchmark rankings moved quickly in the model's favor before independent, task-specific testing had fully caught up.
Pricing — Token-Based, Not Per-Image
| Input Type | Rate |
|---|---|
| Text input | $5 / 1M tokens |
| Image input | $8 / 1M tokens |
| Cached image input | $2 / 1M tokens |
| Image output | $30 / 1M tokens |
Flare and Sunburst share identical token pricing — the same rates as GPT Image 2, with no price increase attached to the upgrade. There's no flat "per image" price, because cost depends on resolution, quality setting, and how much reference-image data you feed in. Using OpenAI's own cost estimator for a standard 1024×1024 output, per-image cost runs from roughly $0.006 at the lowest quality setting up to about $0.211 at the maximum setting — a genuinely wide range, so budgeting off a single "$0.21 per image" headline figure (a number some coverage has repeated from a leaderboard cost calculation) would meaningfully overstate typical cost for lower-quality, higher-volume use cases.
Inside ChatGPT, image generation is available across all plans, including free, with the same usage limits OpenAI had in place before the 2.5 upgrade — the release didn't change how many images free-tier users can generate. On the API side, there's no free tier; usage starts at standard developer rates from your first request.
GPT Image 2.5 vs. the Competition
| Model | Maker | Open Weights | Positioning |
|---|---|---|---|
| GPT Image 2.5 (Sunburst/Flare) | OpenAI | No | Leads current blind-preference leaderboards for generation and editing |
| Nano Banana 2 (Gemini 3.1 Image) | No | Strong generalist competitor, deeply integrated into Google's ecosystem | |
| MAI-Image-2.6 | Microsoft | No | Consistently ranks in the top 3-5 on both leaderboards |
| Grok Imagine Image 2.0 | xAI | No | Competitive on generation, less benchmark presence on editing |
| Seedream 5.0 Pro | ByteDance | No | Strong regional and stylistic alternative |
Every serious competitor in this tier — Google, Microsoft, xAI, ByteDance — is a proprietary, closed model; open-weights alternatives currently sit well below this group on both leaderboards. That's worth knowing if self-hosting or local inference is a requirement for you, since none of the current leaderboard leaders, GPT Image 2.5 included, offer that option.
Strengths and Real Limitations
Who This Is Actually For
GPT Image 2.5 is well-suited to e-commerce and product teams doing targeted edits (swap a background, change a product color, update copy on an ad) where preserving everything else in the frame matters, creators who want an iterative generate-edit-edit-polish workflow instead of restarting from scratch each time, and developers building on the API who want a straightforward choice between speed (Flare) and precision (Sunburst) without juggling separate quality tiers.
It's a weaker fit if you specifically need open weights or local/offline inference — that's simply not on offer here — or if you're doing work where verified, forensic-level editing accuracy is non-negotiable, given that independent testing hasn't yet confirmed a clear improvement in that specific area despite the leaderboard wins. For anything hand-related or highly detail-critical, budget time for a manual review pass rather than assuming 2.5 solved that class of problem outright.
Expert Editorial Opinion
The interesting part of this launch isn't that OpenAI made prettier pictures — it's that the company is visibly trying to turn image generation into something closer to an editable, iterative design process than a slot machine you pull until something usable comes out. Sketch-to-image, comment-based localized edits, multi-turn consistency, and the explicit Flare/Sunburst speed-versus-precision split are all aimed at the same underlying problem: generation alone was never the hard part, staying in control of the result was.
The Artificial Analysis result is the one I'd weight most heavily among the available evidence, specifically because of the sample size. Over 13,000 blind comparisons is a real dataset, not an early signal that could flip with a few hundred more votes, and a clean win by that margin across both generation and editing is meaningful. The LMArena result, by contrast, deserves the "preliminary" label attached to it — a few thousand votes against an incumbent's 78,000 is early days, and treating that leaderboard position as equally settled would be overstating the evidence.
The independent forensic testing is the most valuable thing in this review's source material, and it's the part I'd want every reader to actually internalize rather than skip past. A model can win blind aesthetic preference at scale — "which image do you like better" — while showing no measurable improvement on a narrower, harder question like "did this edit change exactly the thing it was supposed to and nothing else." Those are genuinely different capabilities, and OpenAI's marketing leans on the first while implying the second. That's not dishonest, exactly — the claims are about "improved precision," which is directionally true in some conditions — but it's a gap worth knowing about before you trust 2.5 with an edit where getting only the intended change is the entire point.
The mixed early user reports (hands, processed-looking output, occasional background issues) are the kind of signal that's easy to dismiss as noise and shouldn't be. Every major image model launch generates some complaints; the question is whether they cluster around a specific, plausible failure mode, and reports converging on hands and fine detail rendering are exactly the class of problem diffusion-adjacent image models have historically struggled with. It doesn't erase the leaderboard win, but it's a reasonable flag for anyone doing detail-critical work to test their specific use case before committing to it at scale.
Is it worth using? For iterative creative and product-image workflows where the new editing tools genuinely solve a real friction point, yes, confidently — and the pricing didn't go up to get there. For anyone whose use case depends specifically on the claimed precision-editing improvement being reliably true, the honest answer right now is "probably, in many conditions, but independent testing hasn't fully confirmed it yet" — which is a meaningfully different statement than the leaderboard rankings alone would suggest.
Final Verdict
Score band: 8.0-8.9 — Strong: currently the leading model on major blind-preference leaderboards, with named caveats around Arena's still-preliminary data and independent testing that hasn't fully confirmed every precision-editing claim.
| Dimension | Weight | Score /10 | Why |
|---|---|---|---|
| Technical quality | 30% | 9.0/10 | A decisive, large-sample lead on Artificial Analysis' generation and editing leaderboards, tempered by independent testing that found no clear gain in target-field editing correctness |
| Price-to-value | 25% | 8.0/10 | No price increase over the prior generation, with a genuinely wide quality-tier range for controlling cost per image |
| Maturity & documentation | 20% | 7.5/10 | Excellent official documentation and a dedicated system card, offset by the model's newness and LMArena's still-preliminary ranking status |
| Ceiling & flexibility | 15% | 7.5/10 | Two purpose-built models, sketch input, multi-turn editing, and dual API access are genuinely flexible, capped by the lack of open weights or self-hosting |
| Honesty of positioning | 10% | 7.5/10 | OpenAI's claims are specific and mostly directionally accurate, though marketing language implies more editing precision than independent forensic testing has yet confirmed |
Weighted calculation: (9.0×0.30) + (8.0×0.25) + (7.5×0.20) + (7.5×0.15) + (7.5×0.10) = 8.10.

Comments
Post a Comment