AI voice models
ElevenLabs v4: What the New Voice Model Could Bring to AI RPGs
Meet Eleven v4 and v4 Turbo, then explore what expressive speech could add to Playworlds narration, character voices, and memorable adventures.

The same words can change a scene
“You came back.” A relieved companion, a suspicious guard, and a wounded rival can say those words with entirely different intentions. In an AI roleplaying game, that delivery helps you understand a relationship before you decide what to do next.
ElevenLabs announced Eleven v4 and Eleven v4 Turbo on September 28, 2026. This introduction connects the release to Playworlds’ voice features. Our October 3 code review found Fish Audio powering the current voice pipeline; Eleven v4 is a candidate to explore, not an announced Playworlds integration.
Meet the two versions
ElevenLabs positions v4 for produced speech and v4 Turbo for interactive conversation. Its model catalog lists more than 90 supported languages for both, and a 10,000-character generation limit for standard v4. These are provider specifications, not measurements from a Playworlds session.
For a game, the useful distinction is the scene’s job. A chapter opening can make room for a deliberate performance. A guard answering a quick question needs to leave room for your next move. A larger request limit should not become an excuse for longer speeches.
Expression is a way to direct a performance
Audio tags let a writer suggest delivery inside the script. An original example is “[whispering] Keep the lantern covered. The bridge is watched.” The tag gives a performance direction; the spoken words carry the warning. ElevenLabs says v4 improves tag following, while its documentation acknowledges that results remain imperfect.
A useful RPG voice makes the warning understandable on the first listen. If the whisper hides the clue or every sentence sounds terrified, more expression has made the scene harder to play. Start with one clear intention per line and listen for meaning before adding flourishes.
Where this meets Playworlds today
Playworlds already has a place for a narrator and a cast: the current code supports streamed narration, GM and character voice selection, previews, and separate voice and music volume controls. Those features give a new voice model a concrete job: help players follow who is speaking and what matters.
The voice model performs text. It does not decide whether your persuasion succeeded, remember a hidden quest condition, or replace the game master. A convincing delivery should express the game’s resolved scene while leaving your character’s next action to you.
Listen for continuity as well as drama
Imagine meeting a ferryman at sunset and returning after a failed crossing. His mood may change, but he should still sound like someone you recognize. Our suggested listening test uses a greeting, a warning, and a reunion to check identity across different emotions.
The v4 documentation says Stability and Similarity are its two voice settings; Style and Speed sliders and SSML are not supported. Existing voice workflows therefore need an actual compatibility review. Turning up a setting is not a substitute for listening to the same character across several turns.
A practical first adventure
Choose a short Playworlds scene, preview the voices currently offered, and balance narration against music. Notice whether you recognize the speaker, catch the important name, and feel ready to act when the line ends. Those are useful questions with any speech model.
If you separately experiment with ElevenLabs, use an original scene and compare several takes. The examples in this series are listening exercises and design proposals, not recorded v4 gameplay results. Sources and the Playworlds implementation were reviewed on October 3, 2026.