Skip to content

NPC voices & dialogue

Eleven v4 Turbo and NPC Dialogue: Making Room for Your Next Move

Explore how v4 Turbo could shape responsive NPC conversations, from speaker identity and turn timing to a practical three-turn dialogue exercise.

By Playworlds · Elser.AI ·

A cloaked adventurer follows a stone path toward a glowing gateway among misty mountains.

The best reply invites another turn

At a closed city gate, a clerk asks why you are carrying a royal seal. Your companion starts to answer, then waits for you. A good voiced exchange makes the speakers easy to follow and gives you a clear opening to bluff, explain, or walk away.

Eleven v4 Turbo is designed for interactive speech. This article proposes a dialogue experience built around Playworlds’ existing character voices; it does not announce Turbo support. The current Playworlds voice implementation uses Fish Audio, and this gate encounter is an original illustration.

Understand what the speed number measures

ElevenLabs’ launch report gives Turbo a median time to first speech of about 150 ms in its test, with network latency removed. That is a speech-service measurement. The time between your action and an audible NPC reply also depends on the game generating the answer and the browser beginning playback.

For the gate scene, measure from submitting “I show the seal” to hearing the first useful words. Then record when the whole reply finishes. A quickly started thirty-second lecture can interrupt a conversation more than a slightly later, concise answer. Responsiveness includes the space a character leaves you.

Keep the clerk and companion distinct

Give the clerk a measured delivery and the companion a brisk, familiar one. The distinction should survive a change of mood: the clerk can become alarmed and the companion can become gentle without losing their identities. Playworlds’ existing GM and character casting provides the starting point for that kind of listening test.

Separately, ElevenLabs documents multi-speaker generation with standard Eleven v4 and Eleven v3 through Text to Dialogue. That guide does not establish Turbo support for the same endpoint. A proposed integration must keep each speaker attached to the right line through playback; a speech model cannot infer which Playworlds character owns an utterance.

Do not perform a secret the player has not learned

In our example, the clerk recognizes the seal but has a reason to conceal that fact. A restrained reply can preserve uncertainty. An exaggerated gasp or a guilty laugh could reveal the answer too early, even when the spoken sentence looks neutral.

Direct the performance using what the scene is meant to reveal. “Polite, choosing each word carefully” creates room for interpretation. “Obviously lying” closes it. These are original creative directions to audition externally, not commands supported by the Playworlds player interface.

Make interruptions and recovery part of the audition

In a future Turbo prototype, test what happens when you move on before an old reply finishes, or when audio delivery fails. The listener should be able to understand which response belongs to the current scene. A new model deserves evaluation beyond its best uninterrupted demonstration.

Keep the written response available and avoid treating a replay as a new game action. These are proposed acceptance criteria for a provider change, not a claim that Playworlds currently offers voice interruption or a live microphone conversation with an NPC.

Try a three-turn conversation

First, ask the gate clerk a neutral question. Next, challenge one detail in the answer. Finally, choose a conciliatory response. Listen for a consistent character whose delivery changes with the situation. You should finish knowing what changed and what you can do next.

Try that structure with the voices currently available in Playworlds, or use your own script in a separate ElevenLabs audition. No Turbo gameplay benchmark was conducted for this article. Official sources and game code were reviewed on October 3, 2026.