Compare Image Models Side by Side
Run a blind, paid test between two image models.
Use one prompt to compare exactly two image models. The Arena runs each model through CharGen's normal generation flow, charges the normal Gold price and hides the model names until you vote.
Choose one fantasy task
Open Run a Model Test, select Image and choose the fantasy task that best matches what you want to judge. Character, creature, environment and handout prompts reward different strengths.
Choose two different models. Arena tests are always A/B comparisons, so each result has a clear opponent and one vote.
Write one shared prompt
Both models receive the same prompt. Make it specific enough to judge. For example:
An astronomer in a midnight-blue coat studies a brass star chart beside a ruined observatory in a wide desert at blue hour, full-length composition, natural cinematic light, no text, no logo.
Do not include a model name or provider-specific instruction. The Arena applies the fixed comparison settings for each selected model, so you do not need to match size, seed or quality controls yourself.
Get the quote and start the test
Select Get exact quote. Check the cost shown beside each model and the total Gold cost. Select Spend Gold only when the pair, task, prompt and total are correct.
The test creates two ordinary generation requests. You can leave the page while they run. Return to Run a Model Test to resume the active comparison.
If a generation fails, CharGen's normal request refund path applies. Wait for the displayed Gold balance to refresh before retrying.
Vote without model bias
When both outputs are ready, compare Candidate A with Candidate B before you look for clues about the provider. Choose:
- A is better or B is better when one output follows the prompt better.
- Tie when the outputs are equally useful.
- Both poor when neither output meets the brief.
- Skip when you cannot make a fair decision. A skip keeps both model names hidden.
The model names appear only after a non-skip decision. The prompt and viable outputs then remain in the free Arena for later blind votes.
Use a consistent scoring rubric
| Criterion | What to inspect |
|---|---|
| Prompt adherence | Whether the astronomer, star chart, observatory, desert, blue hour, full-length framing and exclusions are present. |
| Composition | The placement, balance and readable full-length framing of the subject and setting. |
| Detail | Useful detail in the coat, brass chart, observatory and desert without inventing requirements. |
| Artifact control | Distracting anatomy, repeated objects, unintended text, logos or visual defects. |
Judge this prompt, not the model's reputation. The public leaderboard combines many blind votes, so one result is evidence about one run rather than a final claim about a model.
Troubleshooting
If something looks wrong
Troubleshooting
- Symptom
- The model pair cannot be selected.
- Likely cause
- The two selectors must contain different models from the chosen format.
- Next safe action
- Choose one model in each selector, then request a new quote.
- Symptom
- The quote changed after I edited the prompt or models.
- Likely cause
- A quote belongs to the exact comparison configuration shown when it was requested.
- Next safe action
- Review the updated pair and request the exact quote again before spending Gold.
- Symptom
- The comparison is still generating after I return.
- Likely cause
- One or both normal generation requests have not reached a terminal state.
- Next safe action
- Leave the page open or return later. The active match resumes automatically.