
Low Stability Is What Makes A Voice Human
The voice was the failure point and it was a settings problem before it was a script problem. 1 model, eleven_multilingual_v2, and 4 numbers: stability 0.30 to 0.45, similarity boost 0.70 to 0.80, style 0.30 to 0.50, speaker boost on. Stability is the load-bearing 1. Push it past 0.70 and the read goes flat and robotic, because the micro-variation is the human part and it is the first thing that gets averaged away. The same 4 numbers roll 2 ways - an energetic read at stability 0.30 with style 0.45, a contemplative read at stability 0.45 with style 0.25 - and both sit inside the same model. Voice choice is 6 premade options with 6 IDs written down, so a take can be reproduced in 6 months rather than re-cast by ear. There are no pause markers in this engine, so punctuation does the timing instead.
The Voice Was The Failure Point, And It Was A Setting
The first engine handed back a read that was fast, pitchy and unnatural, and the instinct was to rewrite the script. That was the wrong room to be standing in. The line was fine. The settings were doing the damage. 1 engine swap moved the whole problem into a table of 4 numbers, and the read came back human on the first take. The order of diagnosis is the lesson, and it is cheap: when a delivered voice reads wrong, read the settings table before touching the script, because a script written to survive a bad engine is a script written for the wrong delivery. The failure was reproducible, which means it was not taste. It was a value in a field, and a value in a field can be fixed in 1 line rather than 1 rewrite.
4 Numbers Carry The Whole Read
1 model, eleven_multilingual_v2, and 4 numbers underneath it. Stability 0.30 to 0.45. Similarity boost 0.70 to 0.80. Style 0.30 to 0.50. Speaker boost on. That is the entire delivery rig, and everything else in the tool is decoration on top of those 4 fields. Stability is the load-bearing 1, because it decides how much of the per-syllable variation survives the render. Similarity boost sits at 0.70 to 0.80 so the voice stays recognisably the same voice across takes rather than drifting a little further from itself on every generation. Style 0.30 to 0.50 buys expression, and the top of that band is where a read starts to perform rather than speak. Speaker boost stays on, because articulation is free and there is no reason to leave it off. 4 numbers, written down, and a delivery that can be reproduced, argued with and corrected.
Past 0.70, Stability Averages The Human Part Away
Stability at 0.70 and above returns a read that is flat and robotic, and the reason is not a bug. The micro-variation between 1 syllable and the next is the part the ear reads as a person, and it is the first thing that gets averaged away when the engine is told to hold its output steady. Every point of stability above the band buys consistency and spends humanity, and the spend is not linear, which is why the read does not degrade gently. It falls off a shelf somewhere past 0.70. The trap is that the higher value feels safer. It returns the same take twice, and that looks like quality right up until it is played next to a human read. Low stability is not a risk being taken. It is the setting that makes the take worth shipping, and the 0.30 to 0.45 band is where the read stops sounding manufactured.
The Same 4 Numbers Roll 2 Ways
The band is not 1 preset. It is 2 readings inside the same model. An energetic read runs stability 0.30 with style 0.45, which is the highest variation the band allows with the most expression stacked on it. A contemplative read runs stability 0.45 with style 0.25, which is the steadiest and the least performed of the 6 combinations that matter. Both of them sit inside 1 model, 1 voice and 1 settings table, so the difference between the 2 deliveries is 2 fields, and the difference is audible within the first line. Voice choice then narrows to 6 premade options with 6 IDs written down: a warm natural female read, a softer and slightly younger 1, a clear young energetic 1, a bright friendly 1, an animated higher-energy 1, and a warm deep confident 1. The count matters more than the names, because 6 is small enough to audition properly and large enough to cover the range. An ID written into a file is a take that can be rebuilt in 6 months. A voice re-cast by ear is a voice that changes every time the session moves to another desk.
The voice was the failure point, and it was a settings problem before it was a script problem.
There Are No Pause Markers, So Punctuation Does The Timing
This engine has no pause markers. The syntax that worked in the previous engine is not read here at all, and the failure mode is silent, because the render succeeds and only the timing is wrong. Punctuation carries the whole performance instead. 3 dots are a beat between clauses, and it is the strongest tool on the page because it holds 2 thoughts apart without ending either of them. A comma is a short breath. A full stop is a full stop. An em dash is a slightly longer dramatic pause, and it is the 1 that reads as spoken rather than written. Interjections are written as real words, never as bracketed stage directions. The engine either reads the bracket out loud or drops it, and both outcomes break the take: 1 adds a word the script never had, the other removes the beat the script was built around. A laugh typed as a bracketed instruction is a hope. A laugh typed as a word is a read.
The Energy Comes From The Cuts
The rule under the whole page is 1 line: low stability plus real punctuation plus real interjections equals a human voice, and energy comes from the cuts rather than from the speed of the voice. That last clause is where the money is saved, because speeding a read up is the cheapest way to fake excitement and it is also the most legible 1. A voice pushed 10 per cent faster to sound excited sounds 10 per cent more like a machine, and the audience is 1 generation that reads meaning faster than it reads polish. The cut does the work instead: shorter beats, harder transitions, the line landing where the picture already moved. The probability that a take reads as human above stability 0.70 is not 0, and most people are pricing it at 0 while they spend the afternoon rewriting a line that was never the problem. 4 numbers, 6 voice IDs, 4 punctuation marks and 1 rule. The rig fits on a card, which is the point. The method is the same 1 that produced the book, 393 pages published 8 April 2026 in paperback and on Kindle: name the rule, write the number down, and let the next session reuse it instead of rediscovering it.
1 model, 4 numbers, and the read is either human or a machine.
The map is dead. Nobody told you.
Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.
Beyond this book
Building the same thing somewhere else.
Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.
The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com
18-20 November. Online, Las Palmas, Bali.
Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.
ex-aisummit.com →