Skip to content
CreatePricing
Ai Image Model Comparison

Compare AI Image Models for D&D: 372 Blind Votes, Ranked

See which AI image model wins for D&D art. 372 blind votes rank 22 models, and the character board winner is not the overall winner.

Nikita VorontsovFounder & Lead Developer
14 min read
Two anonymous fantasy character portraits beside one shared D&D art brief on a Dungeon Master's desk

I once paid for six versions of a ruined observatory because I changed the model, prompt, and crop at the same time. I learnt nothing from the batch. I only knew that my Gold balance was smaller.

The better method is simple. Compare AI image models with one D&D brief, two models, and the same settings. Then decide what matters at the table before you generate a full character set.

That approach suits fantasy art because the right result depends on the job. A model that gives me a beautiful dragon may lose a character test because it forgets the chipped tusk. Another may create a plain portrait that stays readable when I crop it into a token.

The choice should follow the work. A generic winner is less useful than a model that handles your recurring NPC, campaign colour palette, or battlemap scale.

Two anonymous fantasy character portraits beside one shared D&D art brief on a Dungeon Master's desk

The live rankings sit on the image leaderboard. On 2 September 2026 that board held 372 counted votes across 22 models. The full September standings are further down this page.

See the Live Image Standings

Why one good render proves very little

Most model comparisons are really prompt comparisons. One person writes a short prompt for one tool, then a long prompt for another. They change the aspect ratio. They pick the best result from four images on one side and the first result on the other.

That can be fun. It does not tell you which model fits your workflow.

The problem gets worse with fantasy art. A character prompt can fail in several different ways:

  • the face is attractive but the permanent visual anchors disappear
  • the armour looks good but the weapon becomes three weapons
  • the scene has the right mood but the requested composition is wrong
  • the full portrait looks strong but the token crop becomes muddy
  • the model follows the character brief but ignores the campaign's visual style

I want to separate those failures. A recent benchmark paper on frontier text-to-image systems used 48 difficult prompts across four production systems. The exact benchmark is not a D&D test, but the lesson carries across. Results depend on the prompts and criteria you choose.

For a DM, the test set should match the assets you actually make. If your table needs portraits, test portraits. If you make VTT maps every week, test readable maps. A leaderboard built from product photography tells me very little about a gloomy shrine with six tactical exits.

How to compare AI image models fairly

I keep the comparison small. Two models are enough to make a decision without creating another selection project.

Keep the sameWhy it matters
Prompt textThe model sees the same creative brief
Aspect ratioA portrait and a landscape crop need different composition
ResolutionDetail and generation cost can change with size
Reference imagesIdentity tests are meaningless if one model gets more help
TaskA character result should not be judged by battlemap standards
Number of attemptsThree results against one result creates a selection bias

I also write the decision criteria before I look at the outputs. My usual order is prompt accuracy, character identity, table readability, style fit, and Gold cost. That order stops a dramatic background from winning when the face is wrong.

If both models miss the same detail, the prompt may be the problem. If one model misses it and the other follows it, the model is giving you useful evidence. If both results work, choose the cheaper or faster route for the next batch.

A fair fantasy art comparison keeps one prompt, one crop, and one task while two anonymous model outputs sit side by side

How CharGen's Blind Arena works

CharGen's redesigned AI Model Leaderboard is built for fantasy and tabletop RPG tasks. It covers image, video, audio, and writing models. The image board has separate filters for jobs such as Character, NPC, Monster, Environment, Battlemap, Item, Spell, and Handout.

The free route is a blind vote. Open Free Arena, choose a task, and review candidate A and candidate B. The model names stay hidden while you judge the outputs. The vote controls are direct:

  • A wins when candidate A is stronger
  • Tie when both results meet the brief equally well
  • Both fail when neither result is usable
  • B wins when candidate B is stronger
  • Skip when the task is not clear enough to judge

After a non-skip vote, the model names appear. A skipped battle keeps the names hidden. That is a small detail, but it matters. I am less likely to favour a familiar model when I see only the output and the task.

The ranking uses an Elo-style update. Models begin at 1,000. A win, loss, or tie changes the rating with a K-factor of 32. A skipped decision does not change a rating. Both fail stays as product feedback but does not award a win. The published methodology states the important limit clearly: these ratings are preference signals, not objective quality scores.

The paid route is Run a Test. Choose a format, select two different models, and enter one shared prompt. The model picker shows each current Gold estimate. Before the test starts, the page checks the exact quote again.

Both models receive the same prompt through the normal generation path. The comparison can remain in Generating while the two candidates finish. You can leave the page, return to Your tests, and vote when the results are ready. That fits my real prep better than waiting beside a spinner while a session plan gathers dust.

There is one privacy rule I would not ignore. A paid test sends the submitted prompt and every viable completed output into the free blind-voting pool. Other voters can see them. There is no separate opt-out in the current Arena flow. Do not submit a secret villain reveal, an unpublished setting, or a player document that your group expects you to keep private.

A Dungeon Master's model test note records the shared prompt, two anonymous candidates, the vote, and the next campaign art decision

Standings as of September 2026

I pulled these numbers from the live image board on 2 September 2026. The overall board held 372 counted votes across 22 listed models.

A Dungeon Master's ledger records the September 2026 Blind Arena standings beside two face-down candidate cards and a pile of vote tokens

Sixteen models have at least one counted battle. This is the overall board.

RankModelRatingCounted votes
1GPT Image 2126477
2Seedream 5.0 Pro115081
3Nano Banana 2113170
4Nano Banana 2 Lite107574
5Flux 2 Max104632
6Krea 2 Medium104425
7Grok Imagine Image 2.010161
8MAI Image 2.5101222
9Wan 2.7 Image Pro100637
16Recraft V4 Pro93523
17Qwen Image 391354
18Kling 3.0 Image90962
19Recraft V488936
20Ernie Image Turbo88431
21Ideogram V486556
22Flux 2 Flash86023

The rank numbers jump from 9 to 16. That gap is not a mistake. Six listed models still sit at the 1,000 starting rating with zero counted battles: Hunyuan Image 3.0 Instruct, Imagen 4 (Ultra), Midjourney, Qwen Image 3.0 Pro, Reve V2.1, and Riverflow 2.0 Pro. They hold ranks 10 to 15 by default. A 1,000 rating there means "not yet judged". It does not mean "average".

Two rows need a warning. Grok Imagine Image 2.0 sits at rank 7 on one single vote. That is one person's opinion, not a ranking. Compare it with Seedream 5.0 Pro, which holds second place across 81 votes. I trust the second number far more.

Read the counted-votes column before you read the rating column. It is the honest one.

The overall winner does not win every job

This is the part I did not expect. The image board splits by task, and the task boards disagree with the overall board.

Two painted panels labelled Character and Monster show that different AI image models win different D&D art jobs

Task boardVotesLeaderRatingSecond place
Character137Seedream 5.0 Pro1151GPT Image 2 (1125)
Monster96GPT Image 21198Nano Banana 2 (1100)
Battlemap52GPT Image 21100Seedream 5.0 Pro (1042)
Handout69GPT Image 21103Seedream 5.0 Pro (1087)

GPT Image 2 wins the overall board and three of the four task boards. It loses the Character board to Seedream 5.0 Pro. Character is also the busiest board. It holds 137 of the 372 counted votes.

So the advice at the top of this post now has numbers behind it. A single "best AI image model for D&D" answer hides a real split. If you mostly make character portraits, the overall leader is not the current favourite for your job.

I would read the boards this way today:

  • character portraits: start with Seedream 5.0 Pro, then test GPT Image 2 against it
  • monsters: start with GPT Image 2, because its margin there is the widest on the board
  • battlemaps: start with GPT Image 2, but that board holds only 52 votes, so test before you commit
  • handouts: GPT Image 2 and Seedream 5.0 Pro sit 16 points apart, which is close enough to treat as a tie

These are shortlists, not verdicts. The sample is small on every task board except Character. Vote in the Free Arena and the numbers improve for everyone, including you.

My 15-minute test for D&D character art

I use this when I need a recurring NPC portrait, a player-character reference, or a small set of portraits that must look related.

1. Define the job before choosing the model

Write one sentence that says what the result must do.

For example: Create a chest-up portrait that still reads at 160 pixels wide, with one face, one signature item, and a plain background.

That sentence is better than “make it epic”. It gives me a crop, a composition, and a reason to reject a pretty result.

2. Write a fixed anchor prompt

My test prompt is concrete and short:

Half-orc harbour warden, chipped left tusk, two ritual scars above the right eyebrow, dented blue-grey scale armour, brass key ring at the belt, tired but alert expression, chest-up portrait, plain storm-grey background, readable silhouette.

The chipped tusk, scars, armour, and key ring are the anchors. I do not change them between models. If I add “painted”, “cinematic”, or “high detail” to only one run, I have changed the test.

3. Pick two models with a reason

I do not compare the whole catalogue. I pick one familiar model and one sensible rival. The rival might cost less, fit a different style, or have a better result on the relevant task leaderboard.

If the two prices are far apart, I write that down before voting. A more expensive result needs to solve a problem I care about. “It looks nicer” is not enough if the cheaper result keeps the face and uses half the Gold.

4. Judge the outputs in the same order

I score each candidate from 0 to 2 against five checks:

Check0 points1 point2 points
Prompt accuracymisses the jobfollows most instructionsfollows the brief cleanly
Identity anchorskey details missingdetails are uncleardetails are easy to spot
Token readabilitymuddy at small sizeusable after a cropclear without repair
Campaign fitwrong mood or paletteclose enoughbelongs in this campaign
Cost fittoo costly for the useacceptableeasy to repeat

The score is not a scientific measurement. It is a reminder to judge the same things on both sides. I still make the final call as the DM.

5. Save the result as a campaign decision

I record the prompt, models, task, date, winner, and one sentence about why. For example:

Harbour warden portrait, 24 August 2026, Character task. Nano Banana 2 Lite won for face and key-ring detail at a lower Gold cost. Keep the same anchor prompt for the dock faction.

That note saves time later. I do not need to remember which model won a test from three weeks ago. I also do not need to repeat the same comparison for every minor NPC.

Test the model against the table job

The best test prompt changes with the asset. I use different checks for each job.

Recurring character portraits

Test face shape, hair, signature gear, and a simple background. Use a chest-up crop first. A model that handles the face but fills the image with scenery may be a poor choice for a token set.

NPC batches

Test three roles with the same visual direction: a dock guard, a ship broker, and a shrine keeper. Look for variety without losing the campaign's palette. If every output has the same jawline and coat, the model may be too narrow for your cast.

Battlemaps

Test one room with a clear entrance, cover, and threat area. Check whether the map can be read at the size you use in Battlemap Generator. Decorative detail is secondary to usable space. A beautiful map with no clear route is not ready for play.

Handouts and item cards

Test the object at the final display size. If text matters, write only a short label and check it carefully. For longer rules text, I keep the generated art and add the words in a separate editor. A model can make a good cursed compass without reliably producing a readable item description.

Video or audio tasks

The Arena also supports video, audio, and writing formats. The test should still name the job. For a spell effect, check motion clarity and duration. For ambience, check mood and prompt adherence. For dialogue, check character separation and whether the result leaves room for a DM edit.

How I read the leaderboard without fooling myself

I start with the task filter, not the overall board. Overall ratings answer a broad preference question. The Character or Battlemap task is closer to my real decision.

Then I read three numbers together: rating, counted battles, and typical Gold. A model with a high rating from a small sample is an interesting lead. It is not a settled winner. A model with many votes gives me a stronger signal, even if its rating is less exciting.

New models can sit at 1,000 with no counted battles. That does not mean they are average. It means nobody has supplied enough decisions yet. Likewise, a 100% win rate from one battle is not a reliable reason to spend a month of Gold.

The board is a shortlist. My own prompt is the final test.

When to fix the prompt instead of changing the model

Use the same prompt in a second pair of models before you rewrite everything. If all candidates ignore the same constraint, rewrite the constraint. Put it earlier. Remove a competing detail. Describe the visual result instead of explaining the character's entire history.

If one model ignores the anchor and another preserves it, keep the prompt and change the model. That is the useful kind of comparison because it tells you what the model is good at.

If both models are close, choose the one you can repeat. Campaign art is a series of decisions, not one poster. Consistent faces, predictable crops, and a clear cost matter more than a single spectacular outlier.

FAQ

What is the best AI image model for D&D?

There is no permanent winner. Choose by the asset. Compare two candidates for the exact job, then keep the model that follows the brief, preserves the anchors, fits the campaign, and costs enough to repeat.

How do I compare AI image models fairly?

Use one prompt, the same aspect ratio, the same resolution, the same reference images, and the same task. Compare the same number of attempts and write your decision criteria first.

Does CharGen's Blind Arena show model names before I vote?

No. The free Arena presents candidate A and candidate B without model names. A non-skip vote reveals the models. A skip records no rating change and keeps the names hidden.

Are Blind Arena rankings objective quality scores?

No. They are preference signals from counted votes. Read the task, rating, and battle count together. Use your own D&D brief before committing Gold to a large batch.

Can I use private campaign material in a paid Arena test?

No, not if it must stay private. The submitted prompt and viable completed outputs enter the free voting pool. Use a generic test prompt for unpublished campaign details.

My recommendation

Run one two-model test before you generate ten portraits. Keep the prompt fixed, judge the output at the size your players will see, and record the result. If the test exposes a prompt problem, fix the prompt. If it exposes a model difference, choose the model that you can use again next week.

Run a Blind Model Test