The best AI for generating images
GPT Image 2 sits at the top of this architecture today, because it is the only model that wins both Arena boards at once: text-to-image and editing. If cost is your constraint, Nano Banana 2 delivers a close result; and if your requirement is deployment on your own hardware, only one row of this table is usable for you.
- six models, each with a documented rank
- every number extracted from vendor documentation or a leaderboard
- an Iran column with a stated measurement method
Last checked: This is a reference page. It is re-checked against the vendor sources and updated when a new version ships. Change log
Where each tool's score comes from
The ring below is the same set of weights printed in the table headers further down.
- Text-to-image conversion 30
- Image-editing capacity 20
- Cost per image 20
- Reference-input capacity 10
- Weight-openness status 10
- Access status from Iran 10
We chose these weights, and that is the only judgment call in the table. Weight them differently and the order changes.
Today ranking, built from published data
Both board scores are extracted directly from Arena, and prices and specifications come from each vendor own official documentation. An empty cell means nothing was published, not that the figure is zero.
An empty cell means we could not verify that number, not that the tool scored zero.
Behind each number
Every judged score in this table carries a written reason, and the Iran column says how it was checked. Those two are open here. The measurement trail behind each number, which row of which leaderboard and on how many votes, opens under the model it belongs to.
1 GPT Image 2
Weight-openness status 0/3 No weights are published. The official specification for this model lists only three API endpoints, and no download path of any kind is built into the architecture
Access status from Iran blocked / no working route The supported-countries list for the OpenAI API was extracted from the developers.openai.com infrastructure, and Iran is not recorded on it. The openai.com infrastructure itself returns a 403 status to this measurement server, so that list is the only accessible document, and it speaks strictly about the API rather than the ChatGPT app. No network test has been run from inside Iran.
Measurement trail (2)
Text-to-image conversion 1,381 Position 1 of 77 recorded models on the Arena text-to-image leaderboard, per the 10 August 2026 vote cutoff. The entry is recorded under the id gpt-image-2 (medium), the medium quality configuration, with 70,065 measurement votes
Image-editing capacity 1,463 Position 1 of 53 recorded models on the Arena image-edit leaderboard, per the 7 August 2026 vote cutoff and 184,189 measurement votes, again on the medium quality configuration
2 Nano Banana 2
Weight-openness status 0/3 No weights published, and Google provides no weight-retrieval path for any Gemini image unit
Access status from Iran blocked / no working route The Gemini API available-regions page lists the countries in which the interface is active, and Iran is not recorded on that list. This finding comes from reading a document, not from a network measurement.
Measurement trail (2)
Text-to-image conversion 1,264 Rank 7 of 77 units, Arena text-to-image board, log cutoff 10 August 2026, 29,102 votes. The row identifier carries a web-search suffix, indicating the search-grounded configuration. Confidence interval: 1258.7 to 1268.9
Image-editing capacity 1,385 Rank 10 of 53 units, Arena image-edit board, log cutoff 7 August 2026, 102,657 votes. Confidence interval 1381.0 to 1389.1, which overlaps the Nano Banana Pro interval in full
3 Nano Banana Pro
Weight-openness status 0/3 No weights published, and Google provides no weight-retrieval path for any Gemini image unit
Access status from Iran blocked / no working route The Gemini API available-regions page lists the countries in which the interface is active, and Iran is not recorded on that list. This finding comes from reading a document, not from a network measurement.
Measurement trail (2)
Text-to-image conversion 1,246 Rank 12 of 77 units, reported row gemini-3-pro-image-2k, log cutoff 10 August 2026, 139,851 votes. A second row for the same unit, identified as preview, sits at rank 14 with 1232.0
Image-editing capacity 1,389 Rank 8 of 53 units, reported row gemini-3-pro-image-2k, log cutoff 7 August 2026, 501,853 votes. The preview row is logged immediately behind it at 1385.3
4 Seedream 5.0 Pro
Measurement trail (2)
Text-to-image conversion 1,258 Rank 8 of 77 units, Arena text-to-image board, log cutoff 10 August 2026, 28,830 votes recorded on this row
Image-editing capacity 1,393 Rank 5 of 53 units, Arena image-edit board, log cutoff 7 August 2026, 71,266 votes recorded on this row
5 FLUX.2 [max]
Weight-openness status 0/3 The specification for this member ships without published model weights. Black Forest Labs makes the klein weights available but withholds the max weights, a constraint documented in the vendor's own comparison table.
Measurement trail (2)
Text-to-image conversion 1,162 Rank 24 of 77 models, Arena text-to-image board, vote cutoff 10 August 2026, 117,440 votes on this row
Image-editing capacity 1,262 Rank 27 of 53 models, Arena image-edit board, vote cutoff 7 August 2026, 353,111 votes on this row
6 FLUX.2 [klein] 4B
Weight-openness status 3/3 Model weights are published under an Apache 2.0 license on Hugging Face, alongside an undistilled Base variant that the vendor itself designates for fine-tuning and LoRA training, with an accompanying training guide published
Measurement trail (2)
Text-to-image conversion 1,030 On the Arena text-to-image ranking board this component is registered at position 65 of 77 models total; the data is reported at a vote cutoff of 10 August 2026 with a total of 144,567 votes for this row
Image-editing capacity 1,188 On the Arena image-edit ranking board this component sits at position 43 of 53 models total; the data is registered at a vote cutoff of 7 August 2026 with a total of 507,426 votes for this row
How this ranking is calculated
Every criterion below has a weight and a source. Change a weight and the whole table recomputes. There is no hand-placed position anywhere in this hub.
| Criterion | Weight | Evidence |
|---|---|---|
| Text-to-image conversion | 30 | Arena (Text to Image)Human-preference score on generating an image from a text prompt; carries the highest weight because it is the core task most users expect from these models |
| Image-editing capacity | 20 | Arena (Image Edit)Human preference on editing an existing image; scored independently because the two boards do not agree — several models rank higher on one and lower on the other |
| Cost per image | 20 | vendor stated specificationThe cost of one image at the cheapest size the vendor has published a figure for; a lower value is preferred. Every cell is tied to a specific size, and that condition is documented on the model own page |
| Reference-input capacity | 10 | vendor stated specificationThe capacity to accept reference images in a single generation run; for maintaining a consistent character or product, this figure matters more than any other criterion |
| Weight-openness status | 10 | defined scale, with a written reason per assignmentTaken from the licence the vendor has published, not from our judgment: a spectrum from API-only access to open weights under a free licence with a documented retraining path |
| Access status from Iran | 10 | our access column, with its method statedToday this column creates no distinction: all three candidates with documentation show the identical situation; the full explanation is in the body |
Why GPT Image 2 leads despite an empty cell
Fifty of the hundred points in this architecture belong to the two Arena boards, and this model captures both. On text-to-image conversion it scores 1381.1, while the runner-up records 1263.8, and on editing it stands at 1462.7 against 1392.6. These are wide gaps, and neither confidence interval reaches the next model.
The notable part is that it won with an empty column. OpenAI does not publish a ceiling for reference-image capacity anywhere in its documentation; its sample code accepts four images, but a sample is not a specification ceiling, so that cell was left empty and this model gave up ten points. Even with that ten-point loss, it finishes eight points ahead of second place.
One operational condition is registered directly in the leaderboard row name and gets ignored in most sources: the row that took first place is "gpt-image-2 (medium)", the medium-quality configuration. The high-quality tier at the same size costs $0.211, four times as much, and has no leaderboard row at all. The figure recorded in our price column, $0.053, belongs to the configuration that actually earned the rank.
The lowest-cost route to a strong result
Nano Banana 2. Alongside Nano Banana Pro, it is the only candidate that fills all six columns of this architecture, and the only one that combines full coverage with a high score. It captures the reference column outright (ten object images) and undercuts the first-place model on price.
And here is the finding that changes a purchasing decision: Nano Banana Pro, Google expensive flagship model, ranks third in this table and sits below its own cheaper model. It is genuinely behind on text-to-image conversion, confirmed by the confidence intervals, level on editing, and costs twice as much. If you are choosing between the two Google models, the cheaper one is almost always the correct answer.
The bottom row is the only one you can download
FLUX.2 klein 4B ranks sixth, and that should not be read as weakness. It captures the price column outright at $0.014, roughly a quarter of the first-place cost, and it is the only candidate that scores at all on weight openness. The other five all score zero there, because none of them has published weights.
That means if your operational requirement is deploying the model on your own infrastructure, this table does not offer six options — it offers one. It runs on a graphics card with about 13GB of memory, and its licence is Apache 2.0, meaning commercial use is permitted. Its quality trails the top of the table and that has to be accepted; in exchange it requires no account, no bank card, and no country of yours appearing on any list.
The Iran column creates no distinction today
This has to be stated plainly, because it cannot be read off the table itself. Three candidates carry data in this column: GPT Image 2 and both Google models. All three are in the identical situation: Iran is absent from their supported-country lists, and there is no official payment route. When every value in a column is equal, the scoring engine awards full marks to all of them. So this column produces no difference at all among the top three.
What this column actually does is dock ten points from the other three, not because their situation is worse but because their vendors have published no document to read at all. For FLUX we reviewed the full documentation index today: there is no terms page and no country list, and the terms URL on the main site returns a 404 status code. ByteDance has published nothing comparable for Seedream. In practice this column is measuring whether a vendor publishes a country list, which is not the same question as whether the service works from Iran. Until more documentation surfaces, that is the state of it, and we are not concealing it.
Two criteria deliberately not built into this architecture
First, resolution ceiling. Three vendors publish three incompatible specifications, and the three do not reconcile. Google labels its models "up to 4K" and nowhere publishes pixel dimensions for any tier from 0.5K to 4K. OpenAI is the most precise: 3840 pixels on the long edge and a total pixel count between 655,360 and 8,294,400. Black Forest Labs states 4 megapixels. A label with no number, an exact figure, and a third unit — no column can be built from that combination. These figures were kept as facts on each model own page, and no column was built.
Second, text inside the image. For Persian speakers this may be the single most important operational question, and it is exactly the one no vendor publishes a number for. Google describes its text rendering as advanced. OpenAI itself writes that its model can still struggle with text rendering. Black Forest Labs calls flex specialised for typography. These are three adjectives, and none of them is a measurement. A column built from adjectives is not a measurable criterion. The key for this metric exists in our data structure and stays empty until somebody publishes a number for it.
Why Midjourney has no presence in this architecture
Because no data has been extracted from it at all. midjourney.com returns a 403 status code to our server: the home page, the updates page, and its documentation alike. One difference from OpenAI is worth recording: with openai.com, robots.txt at least responds and does not forbid crawling, whereas midjourney.com/robots.txt is itself a 403, meaning even the rule we would be expected to follow cannot be read.
And unlike OpenAI, whose model holds first place on both boards, Midjourney has no presence on either Arena image board: seventy-seven models are listed on text-to-image and fifty-three on editing, and its name appears on neither. So we hold no facts and no evidence, and we will not construct a row with no numbers behind it. If you read somewhere that Midjourney is the best, ask where the number is.
A row that ranked on half the columns
Seedream 5.0 Pro finishes fourth while filling only two of the six columns in this architecture. It ranks high on both boards, fifth on editing and eighth on text-to-image, and that alone was enough to edge past FLUX.2 max by four tenths of a point, even though FLUX filled four columns. Read this carefully: Seedream position in this table means it performs well on the two things measured, not that it is superior to FLUX.
Those four empty columns are not an oversight on our part either. The Seedream 5.0 Pro page opens and is full of samples, but it carries no model ID, no price, no reference ceiling, and no date. The Seedream 5.0 page returns a 200 status code with an empty body. This row data coverage is exactly 50 of 100, right on the boundary: one more empty cell would have dropped it out of the table entirely into the "insufficient data" list below.
When our top pick is not the right fit
- If your requirement is Persian text rendering inside the image, this table does not answer it: no vendor publishes a figure on text quality, and we have not tested these six models on Persian ourselves.
- If you are looking for a monthly subscription with a graphical interface, this is the wrong page; this table instruments models, and its prices are per-image API costs.
- If your operational requirement is local deployment with no internet connection, only the bottom row of this table is usable, and the rest are not options at all.
- If your image needs to hold several consistent characters, disregard the order of this table and go straight to the reference-input column.
Known weaknesses and limits
- There is no resolution-ceiling column in this architecture, because three vendors publish three incompatible units.
- There is no column for text-rendering quality inside the image, because no vendor publishes a figure for it.
- The Iran column creates no distinction today among the three candidates that carry documentation.
- Prices are tied to a specific size, and for FLUX they are calculated per megapixel, meaning the table figure is a price floor rather than the cost of a large image.
- Seedream 5.0 Pro ranks on 50 percent data coverage, and its position has to be interpreted with that caveat.
- Arena scores are tied to specific serving configurations: the first-place row belongs to the medium-quality tier, and the Nano Banana 2 row carries a web-search suffix.
- Midjourney has no presence in this table because no page of it is reachable from our server and it holds no row on either board.
- We have not run these six models side by side on one prompt ourselves; every figure in this table is published and documented.
Technical verdict
If you must choose one model and budget is not the binding constraint, GPT Image 2 is the clearest choice in this architecture and it wins by a wide margin. If cost is the real constraint, choose Nano Banana 2, and you will likely not notice the difference in daily work. And if your requirement is full ownership of the model with no dependency on external infrastructure, FLUX.2 klein 4B is the only option, and its lower quality has to be accepted. Three different answers for three different constraints, and none of them is a "best" without conditions.
Frequently raised questions, with documented answers
What is the best AI for image generation
GPT Image 2, because it is the only model that wins both Arena leaderboards. If cost matters, Nano Banana 2 delivers a close result for less money.
What is the lowest-cost AI for image generation
FLUX.2 klein 4B, starting at $0.014 per image. Run it locally and no cost is paid to the vendor at all, because its weights are published under the Apache 2.0 licence.
Which AI renders Persian text inside an image
We do not know, and no vendor publishes a figure on it. Three of them describe their text rendering with adjectives like advanced, but none of them says anything specific about Persian. We have not tested this ourselves either, and until we do, we will not write anything about it.
Which image model can I run on my own computer
Only FLUX.2 klein 4B, whose weights are published under Apache 2.0 and which runs on a graphics card with about 13GB of memory. The other five models in this table publish no weights at all.
Can these be used from Iran
For GPT Image 2 and both Google models, Iran is not on the supported-country list and there is no official access route. For FLUX and Seedream, no such document has been published at all. We do not sell accounts and we do not propose ways to bypass the restriction.
Why is Midjourney not in this table
Because its site returns a 403 status code to our server, including its robots.txt, and it holds no row on either Arena image board. We hold no facts and no evidence, so we will not construct a row with no numbers behind it.
If the output has to be a publication-ready image rather than a raw model output, that is exactly the work we do in
our graphic design service