Skip to content
CreatePricing
Ai Writing Model Comparison

Best AI Model for Fantasy Writing: 105 Blind Votes

Find the best AI model for fantasy writing from 105 blind votes across 16 models. See the leader, sample limits, and D&D use cases.

Nikita VorontsovFounder & Lead Developer
12 min read
A fantasy writing brief, a marked notebook, and a ranked AI model board on a Dungeon Master's desk

I needed a short scene for a faction envoy who had lied to the party three times. I did not need a 20,000-word campaign. I needed a voice, a secret, and a reason for the envoy to leave before the players drew steel.

That is the test I care about when I ask for the best AI model for fantasy writing. A model can write smooth prose and still forget the secret, ignore the time limit, or make every NPC sound like the same theatre student. RPG writing needs more than a pretty paragraph.

CharGen's live writing board gives me a useful starting point. On 4 September 2026, it listed 16 models across 105 counted battles. Gemini 3.1 Flash Lite held the top rating at 1,104 from 17 model battles. GPT-5.6 Luna followed at 1,050 from 8. That is an early lead, not a permanent answer.

A fantasy writing brief, a marked notebook, and a ranked AI model board on a Dungeon Master's desk

The board matters because the votes use fantasy and tabletop RPG prompts. General writing tests often ask for a short story, a sales email, or a novel outline. Those tasks do not show me how a model handles a faction relationship, a rules term, a player-facing clue, or a scene that must stop after 350 words.

A D&D writing test needs a D&D brief

I have learnt to distrust the word “best” when the task is missing. The right model for a boxed-text description may not be the right model for a villain monologue. A model that remembers a long setting document may still write dialogue with no room for player action.

The community is asking the same question in different ways. The September AI discussion thread in r/dndai includes prompts, campaign work, fiction writing and RPG tools in one place. The recurring practical issue is not a lack of generated words. It is deciding where those words help and where the DM must keep control.

Research points in the same direction. LitBench evaluates creative writing with human-labelled story comparisons. That is closer to a useful model test than asking one model to grade its own answer. A pairwise result still has limits, but it makes the judgement visible.

For a D&D test, I write the brief before I choose the models. Here is one I could use:

Write a 300-word scene in which a retired dwarf cartographer offers the party a map. The map hides one false road. The dwarf fears the harbour magistrate. Use plain fantasy prose, two speaking characters, one physical clue, and no combat. End when the party must decide whether to accept the map.

The fixed facts give the models something to preserve. The length and ending stop the response from becoming a loose setting essay. The player decision keeps the scene useful at the table.

How the CharGen writing board measures models

The writing model leaderboard covers five jobs: overall writing, story and lore, dialogue, adventure design, and RPG knowledge. The public page shows the current model name, rating, battle count, and typical Gold cost. The Arena methodology page explains the rating and blind-vote process in more detail.

The free route uses blind pairwise votes. You see two anonymous candidates before you vote. The model names appear after a counted decision. That order removes some brand bias. I can judge the response before I see whether it came from a model I already like.

The public board starts models at a rating of 1,000. A win, loss, or tie changes the rating. A model with no battles stays at 1,000, but that does not mean it writes at the average level. It means the board has no preference signal for it yet.

The paid route is called Run a Test. You choose two different models and submit one shared prompt. The page shows the estimated Gold cost, checks the exact quote before the request, and keeps the names hidden until you vote. This is useful when I have a real scene to test rather than a spare minute for public voting.

There is a privacy detail to remember. A viable result from a paid test can enter the free blind-voting pool. Other voters may see the prompt and the output. I would not submit an unreleased adventure, a player secret, or a setting document that my group expects me to keep private.

A blind AI writing model comparison uses one D&D brief, two anonymous responses, and a written decision note

The best AI model for fantasy writing in September 2026

Gemini 3.1 Flash Lite is the current overall leader. It has a 1,104 rating from 17 battles and a typical cost of 1 Gold. GPT-5.6 Luna is second at 1,050 from 8 battles and 2 Gold. Claude Haiku 4.5 and Claude Opus 4.8 both sit at 1,046, with 7 and 6 battles.

Here is the complete live board snapshot I recorded on 4 September 2026.

RankModelRatingBattlesTypical Gold
1Gemini 3.1 Flash Lite1,104171
2GPT-5.6 Luna1,05082
3Claude Haiku 4.51,04672
4Claude Opus 4.81,04669
5GPT-5.41,03066
6Claude Sonnet 51,02864
7GPT-5.51,025611
8Gemini 3 Flash Preview1,02072
9Gemini 3.5 Flash1,02064
10Gemini 3.7 Flash1,01763
11GPT-5.6 Terra1,00005
12Gemini 3.5 Flash Lite986181
13GPT-5.6 Sol985611
14GPT-5.4 Mini97782
15Gemini 3.6 Flash97163
16GPT-5.4 Nano697971

The table has two useful warnings. GPT-5.4 Nano has the largest battle count by a wide margin, but its rating is much lower than the rest. I cannot infer why from this snapshot. It may face a different mix of opponents or tasks. The board does not claim a cause.

GPT-5.6 Terra has a 1,000 starting rating and no battles. Treat it as unjudged. Do not read the rank as a quality result when the battle count is zero.

The close middle matters too. Claude Haiku 4.5, Claude Opus 4.8, GPT-5.4, Claude Sonnet 5, GPT-5.5, and several Gemini models sit within a narrow rating band. A small set of new votes can move them. I would not pay 9 Gold for Opus solely because it shares third place with Haiku at 2 Gold.

The September writing model standings show ratings, battle counts, and Gold costs beside a fantasy campaign notebook

What I would use the board for

The ranking is a shortlist generator. It is not a substitute for a campaign test. I use the numbers to decide which pair deserves my next prompt.

My situationPair I would testReason
I need a cheap first comparisonGemini 3.1 Flash Lite and GPT-5.6 LunaThey lead the board and cost 1 and 2 Gold per run
I want a low-cost controlGemini 3.1 Flash Lite and GPT-5.4 NanoBoth cost 1 Gold, with very different current ratings and battle counts
Voice matters more than priceClaude Haiku 4.5 and Claude Opus 4.8They share a 1,046 rating, while their typical Gold costs differ
I need a middle-cost alternativeGPT-5.4 and Claude Sonnet 5Their ratings are close, and their costs sit between the cheapest and highest options
I want to test an unjudged modelA 1,000-rated model and a board leaderThe comparison can add a signal while keeping one familiar reference

The reason I include a cheap control is simple. A high-cost answer must earn its place in the actual scene. If the 1 Gold model preserves the facts, writes a usable voice, and needs less editing, the price difference is not a small detail.

I also avoid calling a model “good at dialogue” from the overall table alone. The page has separate task filters, and the overall rating combines different writing jobs. I use the dialogue filter for dialogue, the adventure filter for encounters and hooks, and the RPG knowledge filter when rules language matters.

A fair AI writing model comparison for an RPG

My test has five parts. The prompt stays fixed. The scoring stays fixed. I only change the model.

I start with facts. Did the dwarf fear the magistrate? Did the map contain one false road? Did the response end with a decision? A fluent answer that misses one of those facts is not ready for my session.

Next, I check voice. The dwarf should sound like a person with a job and a fear, not a collection of fantasy nouns. I look for repeated sentence shapes, generic threats, and dialogue that explains information the characters already know.

Then I check table use. Can I read the scene aloud? Does it leave a clear opening for player action? Can I cut one paragraph without breaking the sequence? A short scene that gives players a choice is more useful than a longer scene that closes every door.

The final check is edit effort. I mark each change I would make. A result that needs four factual corrections is weaker than one that needs two small line edits, even if the first has more dramatic prose.

CriterionPass question
Brief accuracyDid the output keep every fixed fact and constraint?
VoiceCan I tell the speakers apart without labels?
RPG useDoes the scene create a choice, clue, hook, or action?
Read-aloud qualityCan I use it at the table without rewriting every line?
Edit effortHow many factual and structural changes are needed?

I record the winner with the prompt, date, models, task filter, and one reason. That note stops me from testing the same pair every week. It also tells me when the model changed or the campaign brief moved.

How I use a writing model inside a campaign

I do not ask an AI to write my whole campaign. That creates a large editing job and makes it hard to see which setting facts matter. I give it one bounded request and the context needed for that request.

In Campaign Studio, I keep the faction names, locations, player decisions, and open questions together. When I need a scene, I select only the records that affect it. The prompt might include the envoy's name, the party's last decision, the secret the envoy protects, and the exit condition.

I then use Generate for the narrow output I need. A campaign context prompt can ask for three rumour variants, two complications, or a scene that stops at a clear player decision. I keep the result as a draft until I check it against the campaign records.

The text generators and the model board solve different jobs. CharGen's text generators are free for every account. Run a Test is a paid comparison that uses the displayed Gold cost for the two selected models. I use the free generator when I need content now. I use the board when I want evidence for a model choice I will repeat.

That separation keeps the workflow practical. I do not spend 9 Gold to discover that my prompt forgot the villain's name. I fix the prompt first, then compare two models when the brief is stable.

What 105 battles can and cannot tell you

The current writing board is useful, but it is still young. There are 105 counted battles across 16 models. The model battle counts range from zero to 97. That spread means the ratings do not have equal support.

Pairwise voting also measures preference, not truth. A reader may prefer a vivid answer that contains a rules error. Another reader may favour a plain answer because it is easier to use. The board records that choice. It does not decide what your table values.

The task mix matters. Story and lore, dialogue, adventure design, and RPG knowledge are different jobs. A model can win a lore prompt and lose a rules prompt. Read the task filter before you apply the overall rank to a specific request.

The rating can move as new models enter the pool. A new model with six wins may sit above a familiar model with 97 losses, but the reason remains unknown until the prompt mix and more votes are available. I would report the number beside the rating every time.

That is why I like the board as a live decision aid. It gives me a visible starting point, then sends me back to my own brief. A small, honest test beats a confident universal answer.

FAQ

What is the best AI model for fantasy writing?

Gemini 3.1 Flash Lite currently leads CharGen's fantasy writing board with a 1,104 rating from 17 model battles. The sample is early, so test it against your own brief before you choose it.

How many models are on the CharGen writing leaderboard?

The live writing board lists 16 models and 105 counted battles. Ratings start at 1,000, and a model with no battles is not yet judged.

How do I compare AI writing models for an RPG?

Give two models the same D&D brief, facts, length, format, and criteria. Hide the model names while you judge accuracy, voice, table use, and edit effort.

Does a higher AI model rating prove better fantasy writing?

No. The rating is a preference signal from blind pairwise votes. Read it with the model's battle count, task type, and Gold cost.

Can I test writing models in CharGen?

Yes. Open the writing board to vote in the free Arena, or use Run a Test to compare two models with one shared prompt for the normal Gold cost.

The practical answer is clear enough for my next session. I would start with Gemini 3.1 Flash Lite, compare it with GPT-5.6 Luna on one fixed scene, and keep the winner only if it saves editing time.

Vote on the Writing Leaderboard