Image editing benchmarks split by task, and no single model wins them all
- Source
- Artificial Analysis
- Date
We have updated the Artificial Analysis Image Editing Arena to expand the range of editing tasks we test for, from Enhancement & Restoration to Identity-Preserving Edits, and from Marketing & Advertising to UI/UX Design. Our updated evaluation measures model performance on complex edit tasks, and assesses not just which model is best overall, but which model is best for specific editing needs. The new leaderboard is live and voting is open. Image Editing models are advancing fast. AI now fits into different stages of image creation workflows, from generation through to post-production, and editing is no longer a side feature. We treat Image Editing and Reference to Image as separate benchmarks because they sit at different points in that workflow: Image Editing covers post-production, changing an image you already have; while our upcoming Reference to Image benchmark covers generating novel images from reference images. For Image Editing, we test a model's ability to make specific changes to an image while keeping everything else unchanged. This includes complex edit instructions that chain multiple different asks, as frontier models have largely saturated single-instruction edits. Different edit requests call for different models. Relighting a cinematic scene is a different problem from reworking the design of a marketing asset. We rank models on human preference across 7 editing actions, such as Object-Level Edit, Identity-Preserving Edit, and Enhancement & Restoration, and 10 real-world use cases, such as Marketing & Advertising, UI/UX Design, and Live-Action Film. The overall benchmark samples evenly across both. Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis Image Editing Leaderboard: ➤ MAI-Image-2.6-Preview leads the overall leaderboard and 4 of the 7 editing action boards: Scene & Style Edit, Text or Symbol Edits, Reasoning-Based Edit, and Enhancement & Restoration, where it is tied #1 with MAI-Image-2.5. It excels at restyling, relighting, and retouching images. ➤ GPT Image 2 (high) ranks #2 overall but #1 on Object-Level Edit and Composition & Framing. It is the strongest at precise local edits and spatial reframing. It is weaker at whole-image transformations that must keep the image's content intact, ranking #7 on Scene & Style Edit and #6 on Enhancement & Restoration. ➤ Seedream 5.0 Pro is the character and identity specialist, ranking #1 on Identity-Preserving Edit. ➤ MAI-Image-2.5-Flash is the value pick of the top 10, at $20 per 1,000 images against $211 for GPT Image 2 (high) and $90 for Seedream 5.0 Pro. See below for the editing action and use case breakdowns 🧵
The best model depends on the editing task. MAI-Image-2.6-Preview leads the overall Image Editing Leaderboard and the most category boards, taking #1 on 7 of the 17. GPT Image 2 is next with 5, concentrated in precise object-level and spatial edits.
Example Editing Action: Enhancement & Restoration measures image cleanup, including denoising, deblurring, flare and artifact removal, and restoring old or degraded photos without inventing detail. MAI-Image-2.5, MAI-Image-2.6-Preview, MAI-Image-2.5-Pro, and MAI-Image-2.5-Flash sweeps the top 4, with GPT Image 2 back at #6. The MAI models preserves minute details of the images and applies proper relighting while cleaning up noise, which is exactly what this editing action scores. The pattern extends to Scene & Style Edit, where MAI-Image-2.6-Preview leads, and GPT Image 2 falls to #7, its weakest editing action.



Example Editing Action: Composition & Framing measures 2D and 3D reframing, including camera angle changes, zoom-outs that extend a scene, and rearranging graphic layouts while preserving design, spacing, and style. GPT Image 2 leads MAI-Image-2.6-Preview here by 20 Elo, demonstrating better ability to preserve spatial coherency while changing camera angle, and to keep graphic elements intact while rearranging a 2D layout.



Context
Artificial Analysis expanded its Image Editing Arena to score models across 7 distinct editing actions, such as object-level edits and scene restyling, and 10 real-world use cases, such as marketing assets and UI mockups, instead of ranking models on one overall number. Microsoft's MAI-Image-2.6-Preview leads the overall leaderboard and 7 of the 17 category boards, including enhancement, restoration and text edits. OpenAI's GPT Image 2 (high) ranks #2 overall but wins precise object-level edits and camera-angle reframing, the tasks that require preserving spatial layout exactly. ByteDance Seed's Seedream 5.0 Pro leads Identity-Preserving Edit, the category that measures whether a character's face, outfit and accessories survive an edit.
Price varies widely across the leaders: MAI-Image-2.5-Flash runs $20 per 1,000 images, Seedream 5.0 Pro $90, and GPT Image 2 (high) $211, with MAI-Image-2.6-Preview's own pricing not yet announced since it remains in preview.
Checking sign-in…
Loading comments…







