Skip to content
Create

Compare Image Models Side by Side

Prepare a fair two-model comparison without submitting.

Compare models only after you have made the request fair. A shared prompt and the smallest compatible control set make differences easier to review, but they do not make two providers produce the same image.

Define a fair shared baseline

Start with the smallest set of controls both selected models can use: one non-empty prompt, Count 1, the shared 1024x1024 size, no reference and no seed. Do not treat a control shown for the first selected model as proof that it is a common capability. The reviewed comparison screen checks shared size, not a full intersection of every Quality, reference or advanced control.

The reviewed prompt was:

An astronomer in a midnight-blue coat studies a brass star chart beside a ruined observatory in a wide desert at blue hour, full-length composition, natural cinematic light, no text, no logo.

Keep the prompt, Count and size fixed while comparing. A shared seed, when you choose to set one, controls one input to the experiment. It does not make outputs identical across providers or runs.

Select two compatible models

Open Model Comparison and select the reviewed pair in this order: Flux 2 Flash, then GPT Image 2. The reviewed screen showed both models, 2/4 selected and Ready to compare 2 models without a blocking message.

Desktop Model Comparison screen with Flux 2 Flash and GPT Image 2 selected, the shared astronomer prompt, an empty optional Seed field and a 9 Gold estimate.

The reviewed desktop setup. The pair, prompt and empty Seed field are ready to review, but this screen is not evidence of a submitted comparison.

Use two models for the first comparison. Add more only after you can explain what each additional model will test and have checked that the shared baseline still applies.

Check the aggregate estimate

With the reviewed pair, prompt, Count 1, 1024x1024 size, no reference and no seed, the comparison screen displayed an aggregate estimate of 9 Gold. That was an estimate for this reviewed state, not a charge or a future price promise.

Mobile Model Comparison screen showing Flux 2 Flash and GPT Image 2 selected with a 9 Gold aggregate estimate.

The mobile setup came from a genuine narrow-screen state. Recheck the current aggregate estimate immediately before any future submission.

Stop at the submission boundary

The reviewed evidence stopped before submission. After the picker entered its closed state, its visibly mounted exit animation intercepted submission. Keyboard activation also did not submit. The review did not retry the action.

This is the boundary of the evidence. It proves the route, selected pair, shared controls and 9 Gold estimate. It does not prove that a comparison ran.

Record that no result or charge exists

No comparison ID, constituent request, result, transaction, charge or My Library arrival exists for the reviewed setup. No Gold changed hands. Do not score, share or draw conclusions from a comparison that did not reach a terminal result.

Before a future attempt, record the displayed estimate and confirm the chosen pair and shared controls once more. After a successful submission, reconcile each model's completed result and request-linked transaction before treating the comparison as complete.

Score the results neutrally

Use this rubric only after a future comparison succeeds and both outputs are available. Give each model a 1 to 5 score for the same criteria, then add a short note explaining the score:

CriterionWhat to inspect
Prompt adherenceWhether the astronomer, star chart, observatory, desert, blue hour, full-length framing and exclusions are present.
CompositionThe placement, balance and readable full-length framing of the subject and setting.
DetailUseful detail in the coat, brass chart, observatory and desert without inventing requirements.
Artifact controlDistracting anatomy, repeated objects, unintended text, logos or visual defects.

Score the brief, not the model name. Keep notes specific enough that another reader can see why one output better fits this particular request.

Know what to check after a future comparison

After a future successful run, check that each selected model has one terminal result, then match each result to its request-linked transaction and actual charge. Confirm both results arrive in My Library after reload. Keep the shared prompt, Count, size, reference state and seed state beside the scores so the comparison can be reviewed later.

If one result is missing, charged differently or absent from My Library, stop the comparison record there and resolve the evidence before ranking the pair.

Troubleshooting

If something looks wrong

Troubleshooting

Symptom
A model pair shows a blocking message or cannot reach Ready to compare 2 models.
Likely cause
The pair may not support the current comparison state or one selected control may not be compatible.
Next safe action
Reduce the setup to one prompt, Count 1, a shared size, no reference and no seed, then confirm the live screen accepts the pair.
Symptom
The aggregate estimate changed after I adjusted a setting.
Likely cause
The configured comparison state changed.
Next safe action
Recheck every shared control and use the current aggregate amount before deciding whether to submit.
Symptom
The submit action does not create a comparison.
Likely cause
The current interface may be intercepting the action or the state may not be ready.
Next safe action
Do not assume a request exists. Confirm the visible state, then record a successful terminal result and its request-linked transaction before scoring anything.