How to Compare AI Image Models for D&D Without Wasting Gold
Learn how to compare AI image models for D&D with one prompt, blind A/B votes, fair tests, and a practical CharGen Arena workflow.

I once paid for six versions of a ruined observatory because I changed the model, prompt, and crop at the same time. I learnt nothing from the batch. I only knew that my Gold balance was smaller.
The better method is simple. Compare AI image models with one D&D brief, two models, and the same settings. Then decide what matters at the table before you generate a full character set.
That approach suits fantasy art because the right result depends on the job. A model that gives me a beautiful dragon may lose a character test because it forgets the chipped tusk. Another may create a plain portrait that stays readable when I crop it into a token.
The choice should follow the work. A generic winner is less useful than a model that handles your recurring NPC, campaign colour palette, or battlemap scale.

Why one good render proves very little
Most model comparisons are really prompt comparisons. One person writes a short prompt for one tool, then a long prompt for another. They change the aspect ratio. They pick the best result from four images on one side and the first result on the other.
That can be fun. It does not tell you which model fits your workflow.
The problem gets worse with fantasy art. A character prompt can fail in several different ways:
- the face is attractive but the permanent visual anchors disappear
- the armour looks good but the weapon becomes three weapons
- the scene has the right mood but the requested composition is wrong
- the full portrait looks strong but the token crop becomes muddy
- the model follows the character brief but ignores the campaign's visual style
I want to separate those failures. A recent benchmark paper on frontier text-to-image systems used 48 difficult prompts across four production systems. The exact benchmark is not a D&D test, but the lesson carries across. Results depend on the prompts and criteria you choose.
For a DM, the test set should match the assets you actually make. If your table needs portraits, test portraits. If you make VTT maps every week, test readable maps. A leaderboard built from product photography tells me very little about a gloomy shrine with six tactical exits.
How to compare AI image models fairly
I keep the comparison small. Two models are enough to make a decision without creating another selection project.
| Keep the same | Why it matters |
|---|---|
| Prompt text | The model sees the same creative brief |
| Aspect ratio | A portrait and a landscape crop need different composition |
| Resolution | Detail and generation cost can change with size |
| Reference images | Identity tests are meaningless if one model gets more help |
| Task | A character result should not be judged by battlemap standards |
| Number of attempts | Three results against one result creates a selection bias |
I also write the decision criteria before I look at the outputs. My usual order is prompt accuracy, character identity, table readability, style fit, and Gold cost. That order stops a dramatic background from winning when the face is wrong.
If both models miss the same detail, the prompt may be the problem. If one model misses it and the other follows it, the model is giving you useful evidence. If both results work, choose the cheaper or faster route for the next batch.

How CharGen's Blind Arena works
CharGen's redesigned AI Model Leaderboard is built for fantasy and tabletop RPG tasks. It covers image, video, audio, and writing models. The image board has separate filters for jobs such as Character, NPC, Monster, Environment, Battlemap, Item, Spell, and Handout.
The free route is a blind vote. Open Free Arena, choose a task, and review candidate A and candidate B. The model names stay hidden while you judge the outputs. The vote controls are direct:
A winswhen candidate A is strongerTiewhen both results meet the brief equally wellBoth failwhen neither result is usableB winswhen candidate B is strongerSkipwhen the task is not clear enough to judge
After a non-skip vote, the model names appear. A skipped battle keeps the names hidden. That is a small detail, but it matters. I am less likely to favour a familiar model when I see only the output and the task.
The ranking uses an Elo-style update. Models begin at 1,000. A win, loss, or tie changes the rating with a K-factor of 32. A skipped decision does not change a rating. Both fail stays as product feedback but does not award a win. The published methodology states the important limit clearly: these ratings are preference signals, not objective quality scores.
The paid route is Run a Test. Choose a format, select two different models, and enter one shared prompt. The model picker shows each current Gold estimate. Before the test starts, the page checks the exact quote again.
Both models receive the same prompt through the normal generation path. The comparison can remain in Generating while the two candidates finish. You can leave the page, return to Your tests, and vote when the results are ready. That fits my real prep better than waiting beside a spinner while a session plan gathers dust.
There is one privacy rule I would not ignore. A paid test sends the submitted prompt and every viable completed output into the free blind-voting pool. Other voters can see them. There is no separate opt-out in the current Arena flow. Do not submit a secret villain reveal, an unpublished setting, or a player document that your group expects you to keep private.

My 15-minute test for D&D character art
I use this when I need a recurring NPC portrait, a player-character reference, or a small set of portraits that must look related.
1. Define the job before choosing the model
Write one sentence that says what the result must do.
For example: Create a chest-up portrait that still reads at 160 pixels wide, with one face, one signature item, and a plain background.
That sentence is better than “make it epic”. It gives me a crop, a composition, and a reason to reject a pretty result.
2. Write a fixed anchor prompt
My test prompt is concrete and short:
Half-orc harbour warden, chipped left tusk, two ritual scars above the right eyebrow, dented blue-grey scale armour, brass key ring at the belt, tired but alert expression, chest-up portrait, plain storm-grey background, readable silhouette.
The chipped tusk, scars, armour, and key ring are the anchors. I do not change them between models. If I add “painted”, “cinematic”, or “high detail” to only one run, I have changed the test.
3. Pick two models with a reason
I do not compare the whole catalogue. I pick one familiar model and one sensible rival. The rival might cost less, fit a different style, or have a better result on the relevant task leaderboard.
If the two prices are far apart, I write that down before voting. A more expensive result needs to solve a problem I care about. “It looks nicer” is not enough if the cheaper result keeps the face and uses half the Gold.
4. Judge the outputs in the same order
I score each candidate from 0 to 2 against five checks:
| Check | 0 points | 1 point | 2 points |
|---|---|---|---|
| Prompt accuracy | misses the job | follows most instructions | follows the brief cleanly |
| Identity anchors | key details missing | details are unclear | details are easy to spot |
| Token readability | muddy at small size | usable after a crop | clear without repair |
| Campaign fit | wrong mood or palette | close enough | belongs in this campaign |
| Cost fit | too costly for the use | acceptable | easy to repeat |
The score is not a scientific measurement. It is a reminder to judge the same things on both sides. I still make the final call as the DM.
5. Save the result as a campaign decision
I record the prompt, models, task, date, winner, and one sentence about why. For example:
Harbour warden portrait, 24 August 2026, Character task. Nano Banana 2 Lite won for face and key-ring detail at a lower Gold cost. Keep the same anchor prompt for the dock faction.
That note saves time later. I do not need to remember which model won a test from three weeks ago. I also do not need to repeat the same comparison for every minor NPC.
Test the model against the table job
The best test prompt changes with the asset. I use different checks for each job.
Recurring character portraits
Test face shape, hair, signature gear, and a simple background. Use a chest-up crop first. A model that handles the face but fills the image with scenery may be a poor choice for a token set.
NPC batches
Test three roles with the same visual direction: a dock guard, a ship broker, and a shrine keeper. Look for variety without losing the campaign's palette. If every output has the same jawline and coat, the model may be too narrow for your cast.
Battlemaps
Test one room with a clear entrance, cover, and threat area. Check whether the map can be read at the size you use in Battlemap Generator. Decorative detail is secondary to usable space. A beautiful map with no clear route is not ready for play.
Handouts and item cards
Test the object at the final display size. If text matters, write only a short label and check it carefully. For longer rules text, I keep the generated art and add the words in a separate editor. A model can make a good cursed compass without reliably producing a readable item description.
Video or audio tasks
The Arena also supports video, audio, and writing formats. The test should still name the job. For a spell effect, check motion clarity and duration. For ambience, check mood and prompt adherence. For dialogue, check character separation and whether the result leaves room for a DM edit.
How I read the leaderboard without fooling myself
I start with the task filter, not the overall board. Overall ratings answer a broad preference question. The Character or Battlemap task is closer to my real decision.
Then I read three numbers together: rating, counted battles, and typical Gold. A model with a high rating from a small sample is an interesting lead. It is not a settled winner. A model with many votes gives me a stronger signal, even if its rating is less exciting.
New models can sit at 1,000 with no counted battles. That does not mean they are average. It means nobody has supplied enough decisions yet. Likewise, a 100% win rate from one battle is not a reliable reason to spend a month of Gold.
The board is a shortlist. My own prompt is the final test.
When to fix the prompt instead of changing the model
Use the same prompt in a second pair of models before you rewrite everything. If all candidates ignore the same constraint, rewrite the constraint. Put it earlier. Remove a competing detail. Describe the visual result instead of explaining the character's entire history.
If one model ignores the anchor and another preserves it, keep the prompt and change the model. That is the useful kind of comparison because it tells you what the model is good at.
If both models are close, choose the one you can repeat. Campaign art is a series of decisions, not one poster. Consistent faces, predictable crops, and a clear cost matter more than a single spectacular outlier.
FAQ
What is the best AI image model for D&D?
There is no permanent winner. Choose by the asset. Compare two candidates for the exact job, then keep the model that follows the brief, preserves the anchors, fits the campaign, and costs enough to repeat.
How do I compare AI image models fairly?
Use one prompt, the same aspect ratio, the same resolution, the same reference images, and the same task. Compare the same number of attempts and write your decision criteria first.
Does CharGen's Blind Arena show model names before I vote?
No. The free Arena presents candidate A and candidate B without model names. A non-skip vote reveals the models. A skip records no rating change and keeps the names hidden.
Are Blind Arena rankings objective quality scores?
No. They are preference signals from counted votes. Read the task, rating, and battle count together. Use your own D&D brief before committing Gold to a large batch.
Can I use private campaign material in a paid Arena test?
No, not if it must stay private. The submitted prompt and viable completed outputs enter the free voting pool. Use a generic test prompt for unpublished campaign details.
My recommendation
Run one two-model test before you generate ten portraits. Keep the prompt fixed, judge the output at the size your players will see, and record the result. If the test exposes a prompt problem, fix the prompt. If it exposes a model difference, choose the model that you can use again next week.
Run a Blind Model Test