Write a Scene, Cast a Voice for Every Character, Direct Every Line
Voice Scenes turns a script into a finished piece of audio. Write the scene line by line, give each line a speaker and a voice, and press Render. You get back one MP3 or WAV file plus a timeline of who speaks when, so you can drop it straight into a video, a game or a podcast.
The difference from ordinary text to speech is direction. The 30 directable voices, and any voice you design from a description, take a note on every line about how to say it: pacing, emotion, dialect and acting. Write "panicked shouting, fast" on one line and "quiet, slow, grieving" on the next, and the same character changes delivery the way an actor would. Small sounds like a sigh, a laugh, a breath or a short pause go right into the text where they happen.
Any voice on Reality Fabricator can join the cast: the directable voices, voices you designed in words, your own cloned voice and the standard voices, all in one scene. It is built for audio dramas, story read-alouds, game and visual-novel dialogue, podcast cold opens, trailers and table-top recaps, and for anyone who has wanted more than one flat narrator.
Open Voice Scenes on Reality Fabricator
What you can do
- Rows of speaker, voice, line and direction that you can add, remove and reorder
- 30 directable voices plus voices designed from a written description
- Per-line direction for pacing, emotion, accent and acting
- Inline sounds such as a sigh, a laugh, a breath or a short pause
- Mix directable, designed, cloned and standard voices in one scene
- One MP3 or WAV file back, with a clickable who-speaks-when timeline
- A cost estimate before you render, so you know what a scene will take
How it works
- Write the scene: one row per line, with who says it.
- Pick a voice for each speaker, or design a new one from a description.
- Add a direction to any line that needs a particular delivery.
- Render the scene, listen, adjust a line or two, and download the file.
Frequently asked questions
What does direction actually change?
Delivery. A direction such as whispering, out of breath, sarcastic, slow, or a Scottish accent changes how the line is performed, not what is said. The direction itself is never read aloud. Voices that cannot take a direction simply read the line plainly.
Can I use my own cloned voice in a scene?
Yes. Every voice you can pick elsewhere on Reality Fabricator works here, including your cloned voices, so you can play one character yourself and let designed voices play the rest.
How long can a scene be?
Up to 80 lines and 12,000 characters per render, which is roughly ten to fifteen minutes of dialogue. Longer pieces can be rendered in parts.
What does a scene cost?
Each line is billed on the voice that speaks it. Directable voices cost about twice a standard voice per line, standard and cloned voices cost the normal text-to-speech rate, and the free voices cost nothing. The editor shows an upper-bound estimate before you render.
Can I do this through the API?
Yes. With an API key from your settings, POST /api/audio/dialogue takes the same script as JSON (lines with a voice, text and optional direction) and returns a permanent audio URL with the timeline. POST /api/voices/design creates a voice from a description, and POST /api/audio/synthesize accepts a direction for a single line. The full reference is in the developer docs at /.well-known/llms-full.txt.