Skip to content

Voice model comparisons

ElevenLabs v4 vs v3 and Fish Audio: Choosing a Voice for an AI RPG

Compare Eleven v4, v4 Turbo, v3, and Playworlds’ Fish Audio voice pipeline by expression, responsiveness, continuity, and a practical listening test.

By Playworlds · Elser.AI ·

A cloaked adventurer follows a stone path toward a glowing gateway among misty mountains.

Compare an adventure, not a single impressive clip

A beautiful monologue can hide an awkward game voice. The same model also needs to pronounce a quest item, answer a brief question, and remain recognizable when a character gets upset. This comparison starts with those player needs instead of declaring a universal winner.

We reviewed official ElevenLabs material and Playworlds’ implementation on October 3, 2026. We have not run a controlled audio benchmark. Playworlds currently routes speech through Fish Audio with the model identifier s2.1-pro; the ElevenLabs models below are alternatives for evaluation.

Eleven v4 versus Eleven v3

The official catalog lists v4 at 90+ languages and 10,000 characters per request, versus v3 at 70+ languages and 5,000 characters. ElevenLabs describes v4 as an improvement in voice accuracy and expression, but its own guidance also recommends comparing familiar voices before switching.

That matters for an established campaign. A more faithful rendering of a reference voice may still sound different from the companion a player has come to know. Compare a saved greeting and an emotional return scene; decide whether the new performance preserves the character you intended.

Standard v4 versus v4 Turbo

Standard v4 is aimed at quality-focused production; Turbo targets live interaction. ElevenLabs reports roughly 100 ms median inference latency and 150 ms median time to first speech for Turbo. Its launch measurement removes network latency, so those numbers are not a player’s click-to-audio time.

For a narrated discovery, audition standard v4 first. For a quick exchange with an NPC, include Turbo. Measure the complete wait after the player submits an action, including game-master text generation, speech delivery, and playback. A fast speech service can still sit inside a slow turn.

Where Fish Audio fits this comparison

Fish Audio is the relevant baseline because it is the speech provider selected in Playworlds’ current code. Existing casting, previews, streaming, and playback already form a working integration surface. Replacing that provider would involve those game features as well as the sound of a voice.

This is not a scorecard that treats every Fish release as equivalent. Fish’s public S2 material describes expressive speech, while the Playworlds adapter specifically names s2.1-pro. Keep that exact configured model in any future test and do not transfer a different model’s advertised results to it.

A fair listening test for a small party

Prepare five original lines: a greeting, a whispered warning, an unfamiliar place name, a two-speaker disagreement, and a calm recap. Keep the spoken words and intended performance consistent. Use the same reference recording where supported and permitted; otherwise document the casting differences.

Randomize the clips and ask listeners which words they understood, which character they recognized, and whether the emotion fit. Record several takes, failures, total wait, and billed usage. Compare cost per accepted scene, including retries, rather than pricing one lucky generation. No winner is reported here because that test has not been run.

Choose around the way you play

Our editorial recommendation is to audition v4 for major narration, include Turbo for rapid exchanges, and retain the current Fish integration as the baseline. These are proposed evaluation roles. They are not model options we are announcing in the Playworlds settings.

For players, start with the voices available in your adventure and listen through a complete scene. For creators, make a model earn a change by improving clarity, continuity, and pacing together. A better voice should make you want to answer the character, not merely replay the demo.