AI game master research
How to Compare AI Game Masters for Roleplay
A repeatable way to compare AI game masters on player agency, story continuity, rules, speed, and cost without mistaking one vivid scene for a benchmark.

A good comparison starts with the game turn
Which AI model makes the best game master? A single dramatic paragraph cannot answer that question. An RPG turn succeeds when the response follows the player's attempt, respects established facts, applies the rules, and leaves a meaningful next choice. A model can write beautiful prose while revealing a secret too early or deciding what the player's character does. Another can be brief yet preserve the entire scene.
This guide is a test plan, not a ranking. The festival encounter below is an original illustrative scenario; we have not run it as a controlled Playworlds model benchmark. The research behind this article shows that OpenAI, Anthropic, and Google recommend different prompting and reasoning controls. A fair comparison records those settings rather than treating every model as if it were the same interface.
Use a scene with several ways to fail
Imagine a masked festival where the player carries a broken clockwork key. The player promised a sibling not to accuse the guild steward without evidence and suspects a watchmaker altered its mechanism. The watchmaker noticed a replacement spring but does not know who ordered the work. The player asks about the spring while showing only the key's damaged half. This scene tests whether a response answers the question, respects the watchmaker's limited knowledge, remembers the promise, and avoids forcing an accusation.
Prepare several scenes, not ten copies of this one. Include an ordinary conversation, an uncertain action that needs a roll, an inventory constraint, a clue from an earlier session, and a consequential choice. Give each model the same visible situation and the same authoritative facts. Keep hidden facts out of an NPC's brief unless the application intentionally allows that NPC to know them.
Save the expected constraints before generating a response. For this scene: the watchmaker may identify a replacement spring, cannot name the customer, the sibling's promise remains active, and the player has not chosen whether to accuse anyone. These are checks a reviewer can apply even if two responses use completely different prose.
Run two different comparisons
First test prompt portability. Give each model the same compact scene brief and the same requested output. This answers the migration question: what happens if an existing game-master prompt is moved to a new model? Do not quietly add an extra hint for the model that struggles; the shared prompt is the condition being tested.
Then test an optimized workflow. Adjust each model's documented settings and instructions for the same gameplay goal. This answers a different question: what would a sensible production integration achieve? Record every change. If one model uses a structured decision step while another receives a plain prose request, disclose that the workflows differ. Both comparisons are useful, but combining their results into one unexplained winner is misleading.
Use the actual API model identifier, date, prompt version, context, tool definitions, output limit, and effort setting in your notes. A consumer chat session, coding assistant, gateway route, and first-party API request may have different controls and hidden context. Identify the interface before comparing responses.
Score what a player can notice
Review each turn against five questions: Did it respond to the attempted action? Did it keep established facts and NPC knowledge straight? Did it follow the recorded dice or rule result? Did it leave the player's feelings and next action to the player? Did it create a clear next decision? Mark failures with a short reason and the exact line that caused them. Do not collapse all five into a single impression of 'good writing.'
For a small exploratory test, run each scene several times and retain failures as well as successes. Reviewers can disagree about style, so define objective constraints first and score atmosphere separately. A third reviewer can resolve disputed cases without seeing the model name. This is a proposed method, not a claim that any sample size here proves statistical superiority.
Measure waiting time from submitted action to useful result, rather than only the time spent in one model call. If a game also retrieves state, resolves dice, validates output, or generates speech, those steps affect the player's wait. Record retries, fallbacks, and human repairs. The shortest answer is not necessarily the fastest route to a playable next turn.
Count complete turns, then count their cost
A useful economic measure is total provider spend across all attempts divided by successful completed turns. Suppose a workflow spends $2 over 100 attempts and completes 80 usable turns. Its cost per successful turn is $0.025. A second workflow spending $1.50 and completing 50 turns costs $0.03 per successful turn, despite a lower total bill. These figures are invented to demonstrate the calculation; they are not Playworlds prices or measurements.
Report the raw completed-turn count alongside the ratio. Add cache reads and writes, extra reasoning tokens, supporting calls, retries, and voice costs when they apply. A cached-input discount does not automatically translate into the same percentage reduction for the full adventure. Keep provider usage separate from any credits or price a game shows to players.
The result should be a decision for a specific job: perhaps one model handles a narrow NPC reply while another is tested for a knotty quest consequence. Do not infer that an overall leaderboard rank predicts performance on your own scenes. Re-test after meaningful model, prompt, or game-rule changes.
Try the game, then judge the next choice
A player does not need API logs to notice whether an adventure respects a clue or leaves room for a decision. Choose a world, give your character one immediate goal, and try a conversation followed by an uncertain action. Return later and see whether you can understand where the story stands. Those are useful questions for any AI RPG, regardless of the model name on the label.
Playworlds brings worlds, character sheets, dice, and AI narration into that flow. Available models can change with configuration, and this article does not claim a new comparison result or a particular model rollout. Its rubric is a way to evaluate the experience you actually receive, and a method our team can use when considering future integrations.