Compare AI Image Models for D&D: 372 Blind Votes, Ranked
See which AI image model wins for D&D art. 372 blind votes rank 22 models, and the character board winner is not the overall winner.

I once paid for six versions of a ruined observatory because I changed the model, prompt, and crop at the same time. I learnt nothing from the batch. I only knew that my Gold balance was smaller.
The better method is simple. Compare AI image models with one D&D brief, two models, and the same settings. Then decide what matters at the table before you generate a full character set.
That approach suits fantasy art because the right result depends on the job. A model that gives me a beautiful dragon may lose a character test because it forgets the chipped tusk. Another may create a plain portrait that stays readable when I crop it into a token.
The choice should follow the work. A generic winner is less useful than a model that handles your recurring NPC, campaign colour palette, or battlemap scale.

The live rankings sit on the image leaderboard. On 2 September 2026 that board held 372 counted votes across 22 models. The full September standings are further down this page.
See the Live Image StandingsWhy one good render proves very little
Most model comparisons are really prompt comparisons. One person writes a short prompt for one tool, then a long prompt for another. They change the aspect ratio. They pick the best result from four images on one side and the first result on the other.
That can be fun. It does not tell you which model fits your workflow.
The problem gets worse with fantasy art. A character prompt can fail in several different ways:
- the face is attractive but the permanent visual anchors disappear
- the armour looks good but the weapon becomes three weapons
- the scene has the right mood but the requested composition is wrong
- the full portrait looks strong but the token crop becomes muddy
- the model follows the character brief but ignores the campaign's visual style
I want to separate those failures. A recent benchmark paper on frontier text-to-image systems used 48 difficult prompts across four production systems. The exact benchmark is not a D&D test, but the lesson carries across. Results depend on the prompts and criteria you choose.
For a DM, the test set should match the assets you actually make. If your table needs portraits, test portraits. If you make VTT maps every week, test readable maps. A leaderboard built from product photography tells me very little about a gloomy shrine with six tactical exits.
How to compare AI image models fairly
I keep the comparison small. Two models are enough to make a decision without creating another selection project.
| Keep the same | Why it matters |
|---|---|
| Prompt text | The model sees the same creative brief |
| Aspect ratio | A portrait and a landscape crop need different composition |
| Resolution | Detail and generation cost can change with size |
| Reference images | Identity tests are meaningless if one model gets more help |
| Task | A character result should not be judged by battlemap standards |
| Number of attempts | Three results against one result creates a selection bias |
I also write the decision criteria before I look at the outputs. My usual order is prompt accuracy, character identity, table readability, style fit, and Gold cost. That order stops a dramatic background from winning when the face is wrong.
If both models miss the same detail, the prompt may be the problem. If one model misses it and the other follows it, the model is giving you useful evidence. If both results work, choose the cheaper or faster route for the next batch.

How CharGen's Blind Arena works
CharGen's redesigned AI Model Leaderboard is built for fantasy and tabletop RPG tasks. It covers image, video, audio, and writing models. The image board has separate filters for jobs such as Character, NPC, Monster, Environment, Battlemap, Item, Spell, and Handout.
The free route is a blind vote. Open Free Arena, choose a task, and review candidate A and candidate B. The model names stay hidden while you judge the outputs. The vote controls are direct:
A winswhen candidate A is strongerTiewhen both results meet the brief equally wellBoth failwhen neither result is usableB winswhen candidate B is strongerSkipwhen the task is not clear enough to judge
After a non-skip vote, the model names appear. A skipped battle keeps the names hidden. That is a small detail, but it matters. I am less likely to favour a familiar model when I see only the output and the task.
The ranking uses an Elo-style update. Models begin at 1,000. A win, loss, or tie changes the rating with a K-factor of 32. A skipped decision does not change a rating. Both fail stays as product feedback but does not award a win. The published methodology states the important limit clearly: these ratings are preference signals, not objective quality scores.
The paid route is Run a Test. Choose a format, select two different models, and enter one shared prompt. The model picker shows each current Gold estimate. Before the test starts, the page checks the exact quote again.
Both models receive the same prompt through the normal generation path. The comparison can remain in Generating while the two candidates finish. You can leave the page, return to Your tests, and vote when the results are ready. That fits my real prep better than waiting beside a spinner while a session plan gathers dust.
There is one privacy rule I would not ignore. A paid test sends the submitted prompt and every viable completed output into the free blind-voting pool. Other voters can see them. There is no separate opt-out in the current Arena flow. Do not submit a secret villain reveal, an unpublished setting, or a player document that your group expects you to keep private.

Standings as of September 2026
I pulled these numbers from the live image board on 2 September 2026. The overall board held 372 counted votes across 22 listed models.

Sixteen models have at least one counted battle. This is the overall board.
| Rank | Model | Rating | Counted votes |
|---|---|---|---|
| 1 | GPT Image 2 | 1264 | 77 |
| 2 | Seedream 5.0 Pro | 1150 | 81 |
| 3 | Nano Banana 2 | 1131 | 70 |
| 4 | Nano Banana 2 Lite | 1075 | 74 |
| 5 | Flux 2 Max | 1046 | 32 |
| 6 | Krea 2 Medium | 1044 | 25 |
| 7 | Grok Imagine Image 2.0 | 1016 | 1 |
| 8 | MAI Image 2.5 | 1012 | 22 |
| 9 | Wan 2.7 Image Pro | 1006 | 37 |
| 16 | Recraft V4 Pro | 935 | 23 |
| 17 | Qwen Image 3 | 913 | 54 |
| 18 | Kling 3.0 Image | 909 | 62 |
| 19 | Recraft V4 | 889 | 36 |
| 20 | Ernie Image Turbo | 884 | 31 |
| 21 | Ideogram V4 | 865 | 56 |
| 22 | Flux 2 Flash | 860 | 23 |
The rank numbers jump from 9 to 16. That gap is not a mistake. Six listed models still sit at the 1,000 starting rating with zero counted battles: Hunyuan Image 3.0 Instruct, Imagen 4 (Ultra), Midjourney, Qwen Image 3.0 Pro, Reve V2.1, and Riverflow 2.0 Pro. They hold ranks 10 to 15 by default. A 1,000 rating there means "not yet judged". It does not mean "average".
Two rows need a warning. Grok Imagine Image 2.0 sits at rank 7 on one single vote. That is one person's opinion, not a ranking. Compare it with Seedream 5.0 Pro, which holds second place across 81 votes. I trust the second number far more.
Read the counted-votes column before you read the rating column. It is the honest one.
The overall winner does not win every job
This is the part I did not expect. The image board splits by task, and the task boards disagree with the overall board.

| Task board | Votes | Leader | Rating | Second place |
|---|---|---|---|---|
| Character | 137 | Seedream 5.0 Pro | 1151 | GPT Image 2 (1125) |
| Monster | 96 | GPT Image 2 | 1198 | Nano Banana 2 (1100) |
| Battlemap | 52 | GPT Image 2 | 1100 | Seedream 5.0 Pro (1042) |
| Handout | 69 | GPT Image 2 | 1103 | Seedream 5.0 Pro (1087) |
GPT Image 2 wins the overall board and three of the four task boards. It loses the Character board to Seedream 5.0 Pro. Character is also the busiest board. It holds 137 of the 372 counted votes.
So the advice at the top of this post now has numbers behind it. A single "best AI image model for D&D" answer hides a real split. If you mostly make character portraits, the overall leader is not the current favourite for your job.
I would read the boards this way today:
- character portraits: start with Seedream 5.0 Pro, then test GPT Image 2 against it
- monsters: start with GPT Image 2, because its margin there is the widest on the board
- battlemaps: start with GPT Image 2, but that board holds only 52 votes, so test before you commit
- handouts: GPT Image 2 and Seedream 5.0 Pro sit 16 points apart, which is close enough to treat as a tie
These are shortlists, not verdicts. The sample is small on every task board except Character. Vote in the Free Arena and the numbers improve for everyone, including you.
My 15-minute test for D&D character art
I use this when I need a recurring NPC portrait, a player-character reference, or a small set of portraits that must look related.
1. Define the job before choosing the model
Write one sentence that says what the result must do.
For example: Create a chest-up portrait that still reads at 160 pixels wide, with one face, one signature item, and a plain background.
That sentence is better than “make it epic”. It gives me a crop, a composition, and a reason to reject a pretty result.
2. Write a fixed anchor prompt
My test prompt is concrete and short:
Half-orc harbour warden, chipped left tusk, two ritual scars above the right eyebrow, dented blue-grey scale armour, brass key ring at the belt, tired but alert expression, chest-up portrait, plain storm-grey background, readable silhouette.
The chipped tusk, scars, armour, and key ring are the anchors. I do not change them between models. If I add “painted”, “cinematic”, or “high detail” to only one run, I have changed the test.
3. Pick two models with a reason
I do not compare the whole catalogue. I pick one familiar model and one sensible rival. The rival might cost less, fit a different style, or have a better result on the relevant task leaderboard.
If the two prices are far apart, I write that down before voting. A more expensive result needs to solve a problem I care about. “It looks nicer” is not enough if the cheaper result keeps the face and uses half the Gold.
4. Judge the outputs in the same order
I score each candidate from 0 to 2 against five checks:
| Check | 0 points | 1 point | 2 points |
|---|---|---|---|
| Prompt accuracy | misses the job | follows most instructions | follows the brief cleanly |
| Identity anchors | key details missing | details are unclear | details are easy to spot |
| Token readability | muddy at small size | usable after a crop | clear without repair |
| Campaign fit | wrong mood or palette | close enough | belongs in this campaign |
| Cost fit | too costly for the use | acceptable | easy to repeat |
The score is not a scientific measurement. It is a reminder to judge the same things on both sides. I still make the final call as the DM.
5. Save the result as a campaign decision
I record the prompt, models, task, date, winner, and one sentence about why. For example:
Harbour warden portrait, 24 August 2026, Character task. Nano Banana 2 Lite won for face and key-ring detail at a lower Gold cost. Keep the same anchor prompt for the dock faction.
That note saves time later. I do not need to remember which model won a test from three weeks ago. I also do not need to repeat the same comparison for every minor NPC.
Test the model against the table job
The best test prompt changes with the asset. I use different checks for each job.
Recurring character portraits
Test face shape, hair, signature gear, and a simple background. Use a chest-up crop first. A model that handles the face but fills the image with scenery may be a poor choice for a token set.
NPC batches
Test three roles with the same visual direction: a dock guard, a ship broker, and a shrine keeper. Look for variety without losing the campaign's palette. If every output has the same jawline and coat, the model may be too narrow for your cast.
Battlemaps
Test one room with a clear entrance, cover, and threat area. Check whether the map can be read at the size you use in Battlemap Generator. Decorative detail is secondary to usable space. A beautiful map with no clear route is not ready for play.
Handouts and item cards
Test the object at the final display size. If text matters, write only a short label and check it carefully. For longer rules text, I keep the generated art and add the words in a separate editor. A model can make a good cursed compass without reliably producing a readable item description.
Video or audio tasks
The Arena also supports video, audio, and writing formats. The test should still name the job. For a spell effect, check motion clarity and duration. For ambience, check mood and prompt adherence. For dialogue, check character separation and whether the result leaves room for a DM edit.
How I read the leaderboard without fooling myself
I start with the task filter, not the overall board. Overall ratings answer a broad preference question. The Character or Battlemap task is closer to my real decision.
Then I read three numbers together: rating, counted battles, and typical Gold. A model with a high rating from a small sample is an interesting lead. It is not a settled winner. A model with many votes gives me a stronger signal, even if its rating is less exciting.
New models can sit at 1,000 with no counted battles. That does not mean they are average. It means nobody has supplied enough decisions yet. Likewise, a 100% win rate from one battle is not a reliable reason to spend a month of Gold.
The board is a shortlist. My own prompt is the final test.
When to fix the prompt instead of changing the model
Use the same prompt in a second pair of models before you rewrite everything. If all candidates ignore the same constraint, rewrite the constraint. Put it earlier. Remove a competing detail. Describe the visual result instead of explaining the character's entire history.
If one model ignores the anchor and another preserves it, keep the prompt and change the model. That is the useful kind of comparison because it tells you what the model is good at.
If both models are close, choose the one you can repeat. Campaign art is a series of decisions, not one poster. Consistent faces, predictable crops, and a clear cost matter more than a single spectacular outlier.
FAQ
What is the best AI image model for D&D?
There is no permanent winner. Choose by the asset. Compare two candidates for the exact job, then keep the model that follows the brief, preserves the anchors, fits the campaign, and costs enough to repeat.
How do I compare AI image models fairly?
Use one prompt, the same aspect ratio, the same resolution, the same reference images, and the same task. Compare the same number of attempts and write your decision criteria first.
Does CharGen's Blind Arena show model names before I vote?
No. The free Arena presents candidate A and candidate B without model names. A non-skip vote reveals the models. A skip records no rating change and keeps the names hidden.
Are Blind Arena rankings objective quality scores?
No. They are preference signals from counted votes. Read the task, rating, and battle count together. Use your own D&D brief before committing Gold to a large batch.
Can I use private campaign material in a paid Arena test?
No, not if it must stay private. The submitted prompt and viable completed outputs enter the free voting pool. Use a generic test prompt for unpublished campaign details.
My recommendation
Run one two-model test before you generate ten portraits. Keep the prompt fixed, judge the output at the size your players will see, and record the result. If the test exposes a prompt problem, fix the prompt. If it exposes a model difference, choose the model that you can use again next week.
Run a Blind Model Test