Fantasy AI Model Leaderboard: September 2026 Standings
The September fantasy AI model leaderboard compares 548 blind votes across image, audio, writing and video models. See leaders and sample limits.

I checked the ai model leaderboard fantasy boards on 7 September 2026. The four overall tables contained 548 counted votes across image, audio, writing and video models. GPT Image 2 led the image board. Lyria 3 Pro led audio. Gemini 3.1 Flash Lite led writing. Video is still too small to call.
That is a useful result, but it is not one universal ranking. CharGen keeps a separate board for each medium because a model that makes a strong character portrait may not make a good battle theme or a coherent ten-second scene. The numbers below show where the current signal is strongest and where it is still thin.

September fantasy AI model rankings at a glance
The table uses the overall task for each modality. Image models are judged on character, monster, battlemap and handout work. Audio combines music, background music and sound effects. Writing covers story and lore, dialogue, adventure design and RPG knowledge. Video uses the overall video task.
| Board | Counted votes | Current leader | Rating | Votes for leader | My reading |
|---|---|---|---|---|---|
| Image | 372 | GPT Image 2 | 1,264 | 77 | The strongest sample in this report |
| Audio | 62 | Lyria 3 Pro | 1,114 | 13 | Useful early signal, with close music results |
| Writing | 105 | Gemini 3.1 Flash Lite | 1,104 | 17 | A readable lead, but still a small board |
| Video | 9 | Seven models share the first band | 1,015 to 1,016 | 1 each | Do not treat this as a settled ranking |
The rating starts at 1,000. A result above 1,000 shows that the model has won more preference signal than it has lost on that board. It does not show that one model is better at every prompt. It also does not make ratings from different boards directly comparable.
Image models: GPT Image 2 has the clearest lead
GPT Image 2 sits first at 1,264 from 77 counted votes. Seedream 5.0 Pro is second at 1,150 from 81 votes. Nano Banana 2 is third at 1,131 from 70 votes. That gives the image board a useful spread between the first three rows, with enough votes to make the order worth watching.
The vote counts add an important detail. Seedream 5.0 Pro has more votes than GPT Image 2, but it has a lower rating. I would not call that a contradiction. The board records pairwise preferences, not a simple popularity poll. The result can reflect which opponents each model faced and how voters judged the same fantasy brief.
The image board covers four jobs: characters, monsters, battlemaps and handouts. That mix matters. A model can produce a strong face and still fail a top-down map or a handout with readable text. I would use the image model leaderboard for the task filter, then test the two models against the exact asset I need.
Meta's official Muse Image and Muse Video announcement is a useful reminder that provider releases can add new options faster than a small community board can test them. A model name on a catalogue is not the same as a model with a trusted sample. When a new image row appears at 1,000, I read it as unjudged until voters build evidence.
My test for a character portrait is simple. I write the same brief for both candidates. I name the face, clothing, pose, light and crop. I judge the image at the size the group will see, not only at full screen. I record one reason for the winner. “Better” is not enough. “Kept the scar and showed the shield” gives me a decision I can repeat.
Audio models: Lyria 3 Pro leads a close top group
Lyria 3 Pro leads the audio board at 1,114 from 13 votes. Mureka V9 BGM follows at 1,108 from 10 votes. MiniMax Music 3.0 sits third at 1,082 from 13 votes. The top two are close, and the board includes several different audio jobs, so I would not use the overall rank as a final choice for every session.
The board covers music, background music and sound effects. A long battle track and a short tavern door sound do not need the same strengths. Open the audio leaderboard and switch to the task that matches your request. The overall table is a starting point, not a replacement for that check.
Google's official Lyria page describes the model family as a music system for tracks, clips and real-time work. The provider description gives useful product context. The CharGen board answers a different question: which available model did voters prefer for the same fantasy or RPG prompt? Those two facts can sit together without turning a release page into a benchmark.
The audio sample also shows why new model coverage needs time. Mirelo SFX1.6 appears in the current audio catalogue, but it has no counted votes on the overall board. Mureka V9.5 and Mureka V9.5 BGM are also waiting for a signal. I would not rank an unjudged model above a tested model just because the name is newer.
For a table, I would run separate prompts for three jobs. A music prompt might ask for a tense orchestral theme with a clear rise before the final round. A background prompt might ask for quiet campfire music with no strong melody. A sound-effect prompt might ask for a tavern door opening during a storm. The winner can change when the job changes.
Writing models: Gemini 3.1 Flash Lite keeps the lead
Gemini 3.1 Flash Lite leads the writing board at 1,104 from 17 votes. GPT-5.6 Luna is second at 1,050 from 8 votes. Claude Haiku 4.5 and Claude Opus 4.8 share 1,046, with 7 and 6 votes. The board lists 16 models, but one row has no counted votes, so I exclude it from quality claims.
The recent best AI model for fantasy writing report covers the writing board in more detail. The short version is that a rating cannot tell me if a model will preserve a faction secret or stop a scene at the right player decision. I need one fixed D&D brief and one editing check.
GPT-5.4 Nano is the useful warning in this table. It has 97 votes, far more than any other writing model, but a rating of 697. I cannot infer the reason from the table. It may have faced a different opponent mix or a different task mix. The data supports the result, not a cause.
I read the writing board in the same way I read the others. A lead is a shortlist. It is not permission to stop checking. For a campaign scene, I score factual accuracy, distinct voices, table use, read-aloud quality and edit effort. A plain answer that keeps the facts can beat a stylish answer that changes the NPC's motive.
Video models: the board needs more votes
The video board has nine counted votes. Seven models sit between 1,015 and 1,016 with one win each. HappyHorse 1.1 has two votes, with one win and one loss, and a rating of 1,001. Seedance Pro has two votes, with one win and one loss, and a rating of 999. The rest of the listed rows have no counted votes or have one loss.
That is not enough evidence for a stable winner. A model can reach the top band after one result. One later vote can change the order. I would not use this table to make a large video purchase decision yet.
The useful conclusion is simpler. The video model leaderboard needs more blind tests. Use the same clip brief, the same duration and the same reference image when the task supports one. Judge motion, character continuity, camera control and prompt compliance. Record the reason for the choice, not only the model name.
The small sample also changes how I describe new video models. “Currently unjudged” is a better label than “low ranked” when a row stays at 1,000. A model with no votes has not lost. It has not been tested on this board.
What the Arena rating measures
CharGen's Arena methodology uses one prompt for two models. Voters see candidate A and candidate B without model names. A non-skip vote reveals the models after the choice. That order helps reduce brand preference, although it cannot remove every source of bias.
Models start at 1,000. A win, loss or tie changes both ratings with a K-factor of 32. A “Both poor” result is kept as product feedback but does not award a win. A skipped battle does not change either rating. The board therefore captures counted preferences, not a lab score.
The method has a practical limit. A small sample can move quickly. It can also reflect the prompts people choose to submit. If the image board receives mostly character portraits, its overall ranking will say more about character portraits than about maps. The task filters exist for that reason.
The rating is also not a price score. A lower-cost model may be the right choice when it follows the brief well enough for a batch. A high-rated model may still need more retries for a particular job. The board helps me choose which comparison to run. It does not decide the budget for me.


How I use the September table
I use four checks before I select a model for a real campaign asset:
- Read the task. Character art, sound effects, dialogue and video motion are separate jobs.
- Read the vote count. A rating from one vote needs more caution than a rating from 77 votes.
- Keep the prompt fixed. Give two models the same facts, length, crop, duration or reference image.
- Judge the handoff. Keep the output that fits the next step, whether that is a token, a battlemap, a recap, a track or a cutscene.
The fourth check is the one I miss most often. A pretty output can be a bad campaign asset if it cannot be cropped, read aloud, looped, or placed on the table. A model comparison should include the work that follows the generation step.
For the current boards, my shortlist is direct. I would begin with GPT Image 2 for a new image comparison, Lyria 3 Pro for a music comparison, and Gemini 3.1 Flash Lite for a bounded writing test. I would not name a video winner yet. I would cast more votes there first.
The report also shows a useful baseline for October. The image board already has a broad sample. Audio needs more task-specific votes. Writing needs more balanced coverage around the top group. Video needs far more counted decisions before its ratings can carry much weight. A monthly report is only useful if it shows where the evidence is strong and where the next vote matters.
FAQ
Which model leads the September fantasy AI model leaderboard?
GPT Image 2 leads the image board at 1,264 from 77 counted votes. Lyria 3 Pro leads audio at 1,114 from 13 votes, and Gemini 3.1 Flash Lite leads writing at 1,104 from 17 votes. The video board has only nine votes, so it has no stable winner.
How many votes are in the September 2026 report?
The four overall boards contain 548 counted votes in total: 372 image votes, 62 audio votes, 105 writing votes and 9 video votes. The figures are a snapshot taken on 7 September 2026.
Are CharGen Arena ratings objective quality scores?
No. They are preference signals from blind pairwise choices. Models start at 1,000, and the rating changes after a counted win, loss or tie. Read the rating beside the vote count and task.
Why is the September video leaderboard difficult to read?
The video board has only nine counted votes. Several models have one win from one vote, which creates a narrow but unstable lead. More blind comparisons are needed before a video rank is useful as a buying or workflow decision.
How often should I check the fantasy AI model rankings?
Check the live board before a new project, and use the monthly report as a dated snapshot. New models and new votes can change a small board quickly, especially for video and audio tasks.
The practical choice for September is not one model for every job. Use the board that matches the asset, read the sample size, and run one controlled comparison before you commit to a larger batch.
View the Fantasy AI Model Leaderboard