Skip to content
CreatePricing

Compare Image Models Side by Side

Run a blind, paid test between two image models.

Use one prompt to compare exactly two image models. The Arena runs each model through CharGen's normal generation flow, charges the normal Gold price and hides the model names until you vote.

Choose one fantasy task

Open Run a Model Test, select Image and choose the fantasy task that best matches what you want to judge. Character, creature, environment and handout prompts reward different strengths.

Choose two different models. Arena tests are always A/B comparisons, so each result has a clear opponent and one vote.

Write one shared prompt

Both models receive the same prompt. Make it specific enough to judge. For example:

An astronomer in a midnight-blue coat studies a brass star chart beside a ruined observatory in a wide desert at blue hour, full-length composition, natural cinematic light, no text, no logo.

Do not include a model name or provider-specific instruction. The Arena applies the fixed comparison settings for each selected model, so you do not need to match size, seed or quality controls yourself.

Get the quote and start the test

Select Get exact quote. Check the cost shown beside each model and the total Gold cost. Select Spend Gold only when the pair, task, prompt and total are correct.

The test creates two ordinary generation requests. You can leave the page while they run. Return to Run a Model Test to resume the active comparison.

If a generation fails, CharGen's normal request refund path applies. Wait for the displayed Gold balance to refresh before retrying.

Vote without model bias

When both outputs are ready, compare Candidate A with Candidate B before you look for clues about the provider. Choose:

  • A is better or B is better when one output follows the prompt better.
  • Tie when the outputs are equally useful.
  • Both poor when neither output meets the brief.
  • Skip when you cannot make a fair decision. A skip keeps both model names hidden.

The model names appear only after a non-skip decision. The prompt and viable outputs then remain in the free Arena for later blind votes.

Use a consistent scoring rubric

CriterionWhat to inspect
Prompt adherenceWhether the astronomer, star chart, observatory, desert, blue hour, full-length framing and exclusions are present.
CompositionThe placement, balance and readable full-length framing of the subject and setting.
DetailUseful detail in the coat, brass chart, observatory and desert without inventing requirements.
Artifact controlDistracting anatomy, repeated objects, unintended text, logos or visual defects.

Judge this prompt, not the model's reputation. The public leaderboard combines many blind votes, so one result is evidence about one run rather than a final claim about a model.

Troubleshooting

If something looks wrong

Troubleshooting

Symptom
The model pair cannot be selected.
Likely cause
The two selectors must contain different models from the chosen format.
Next safe action
Choose one model in each selector, then request a new quote.
Symptom
The quote changed after I edited the prompt or models.
Likely cause
A quote belongs to the exact comparison configuration shown when it was requested.
Next safe action
Review the updated pair and request the exact quote again before spending Gold.
Symptom
The comparison is still generating after I return.
Likely cause
One or both normal generation requests have not reached a terminal state.
Next safe action
Leave the page open or return later. The active match resumes automatically.