
The Shot List Is The Fix
A generative video model fails in the same places every time: skin, faces, hands, text, motion, physics and fine repeating patterns. Every one of those failures has a known cause, and a fault with a known cause can be designed out of the shot list before the render instead of repaired after it. The shot list is the fix, and it is written in 20 minutes.
The Failures Cluster In The Same 7 Places
The failure modes of a generative video model are not random. Across Wan, Veo, Sora, Kling and Runway the breaks land in the same 7 places: skin, faces, hands, text, motion, physics, and fine repeating patterns. That predictability is the useful part. A fault with a known cause can be removed from the shot list before anything renders, which is cheaper than repairing it in a grade at 3840 by 2160 after the cut has been approved. A model asked for a medium shot of a person crossing a room at 35mm will usually return something usable. The same model asked for a 90 degree head turn in a mirror, with signage in the background and the camera moving while the subject moves, will return a smear. Nothing about the model changed between those 2 requests. The requests changed. A 30 degree move between 2 shots of the same subject is enough to keep an audience oriented, and it costs nothing at 35mm. The whole discipline is 1 sentence long. The model fails in ways that are known in advance, so the work is not prompting harder, it is deleting the shot that produces the failure. You do not repair waxy skin afterwards. You do not shoot the shot that makes waxy skin. This is not a limitation to argue with. Every renderer in the stack has the same blind spots because they share the same design: each frame is generated from the prompt and the previous frame, with no persistent memory of the world. That one fact explains most of the list. Faces drift, hands fuse, reflections disagree with the room, and a fence at the wrong angle dissolves into moire, all because nothing in the system is holding a stable idea of the space between 1 frame and the next. So the shot list does 2 jobs at once. It decides what the audience sees, and it decides what the model is asked to attempt. Most productions only do the first job and then spend the difference in post.
Skin Fails In 2 Opposite Directions
Skin is where a synthetic image exposes itself first, and it fails in 2 opposite directions. The first direction is smooth. The model over-smooths to suppress noise, and pores and micro-texture disappear until the subject reads as a mannequin. The prompt words that cause it are the flattering ones: photorealistic, beautiful, perfect, flawless, and the blanket line about a cinematic film still with film grain attached to every shot. The second direction is the one that catches people out. Ask for hyper-detailed skin texture on a tight face and the model does not add detail it saw. It hallucinates detail it never saw, growing pores, veins and freckles across a face that is now 15 years older than the person in it. The ultra-realistic close-up is the worst offender in the list, which is counterintuitive, because it is also the prompt most people write when they want quality. Both directions have 1 fix and it is a framing decision. Keep faces at medium to wide. A face at 35mm or 50mm carries enough skin to read as real without asking the model to invent structure. The 85mm close-up that puts the entire frame on 10 square centimetres of cheek is the shot that produces both failures, sometimes inside 1 clip. At 100 per cent zoom both failures are unmistakable, which is why the working pair for a dialogue scene is 35mm and 50mm, not the 85mm that reads best on a monitor. Then control the language. Documentary wording holds skin steady: natural skin texture, visible pores, fine laugh lines, unretouched, natural light, negative fill. Drop photorealistic entirely. Drop cinematic as a blanket adjective, because it does not describe light, it describes a mood the model cannot verify. The rule that saves the most time is this. Never demand extreme skin detail. Specify restraint, not more detail. A face is not a texture to be maximised. It is a pixel budget to be respected.
A Held Face Drifts Because Nothing Remembers
A model generating video has no memory of what it made a moment ago. Any error introduced in 1 frame becomes the input for the next. On a face that error is visible. Hold a static close-up for 4 seconds and watch the jawline shift, the nose widen, the eyes go uneven. The face dissolves over time because nothing in the system is holding it still. The repair is scheduling, not prompting. Face shots stay short, 40 frames or less at 24 fps, or they keep moving so the drift reads as motion instead of decay. A slow push in survives. A lock-off does not. Over 20 shots the drift compounds in every one of them, and by the last setup the character is a relative of the person who walked into the first. Full profile turns are the second trap. A slight head turn is fine. A 90 degree rotation smears the features and deforms the eyes, and a full turn past 180 degrees destroys them. If the scene needs the character facing the other way, cut to it, or cross the 180 degree line with a move that shows the turn. Talking, chewing and eating close-ups produce the classic AI mouth, where the jaw deforms around every phoneme. Multiple faces in 1 frame is the fourth trap, because the background faces melt while the hero face holds. The rule is 1 hero face per frame, and everyone else stays out of focus. Emotional stacking is the last one. A prompt that asks for crying and smiling and grimacing in the same beat overloads the face and forces it to morph between states. Pick 1 state per shot and let the cut carry the change.
Hands Are The Tell
Hands are the tell. Extra fingers, fused fingers and wrong counts are the first thing an audience notices and the last thing the model gets right, because a hand is the most information-dense moving object in any frame. It carries 27 bones and dozens of joints, and it rotates through more distinct silhouettes per second than a face ever does. The fix has 2 halves and they work together. The first half is optical. Keep hands out of tight focus. Let them sit at the edge of the depth of field, or partly out of frame, or behind an object in the foreground. In a 50mm close-up the hand is the subject and the model runs out of reliable structure. If the shot genuinely needs hands, give them slow deliberate movement at 24 fps rather than gesture, because speed is what multiplies the error. The second half is grammatical. Describe where the hands start and where they finish. Do not describe the gesture in between. If the prompt says the hand reaches for the glass, the model invents the motion path between those words, and the invented path is where the fingers fuse. If the prompt says the hand rests on the table at the start of the shot and is closed around the glass at the end, the model has 2 poses and 1 interpolation, and the interpolation has somewhere to land. Avoid hand close-ups, sign language, hands holding signs or text, precise measured gestures, and any fast hand motion across the frame. A hand entering a 35mm frame at 24 fps is 1 of the highest failure rates in the entire list, and it is 1 of the easiest to remove from the shot list before anything is generated.
You do not fix waxy skin in post. You do not shoot the shot that produces it.
Text, Motion, Physics And Fine Patterns
Text is the simplest failure to design around. The model garbles writing, and it gets worse the longer and smaller the string. Storefront signage, book covers, phone screens, documents, hands holding signs, tattoos, clothing with words, watermarks: all of them produce corrupted lettering. If text has to appear, keep it to 2 short words, large in frame, and low in motion. Everything else is a repair queue. Motion is the second cluster. Fast or sudden movement produces blur artifacts, warping and smear. Combined camera and subject motion doubles the warp, which is why the safe rule is move 1 or the other, never both in the same beat. Slow-motion generated at native speed and retimed afterwards behaves. Slow-motion asked for directly, with too few source frames, judders and smears, and the fix is 60 fps at the source, not a higher percentage in post. Physics is the third. Mirrors, glass and water reflections are the hardest thing in the list, because the model has to hold 2 spaces consistent at once and it does not. The reflection in the glass will not match the room behind the subject, and at 24 fps an audience catches it in under 2 seconds. Fine repeating patterns are the last cluster. Teeth arrive too many and too even, a row of them reading as a shattered keyboard. Picket fences, venetian blinds, checkerboards and brick walls at the wrong angle dissolve into moire. The repair is the same in every case: pull the camera back, change the angle, or throw the pattern out of focus until it stops being a grid. A picket fence at 4096 pixels wide fails harder than the same fence at 1920, because there are more repeating edges for the model to lose count of.
The 1 Question And The Standing Negative
Before any prompt is written, 1 question runs: does this shot ask the model to do something it reliably fails at? If the answer is yes, the shot is redesigned. It is not prompted around, and it is not rescued later. The question takes 10 seconds and it saves the render, which is the other half of the arithmetic: a 45 minute generation that has to be run a second time costs more than the 20 minutes spent writing the shot list. Then attach the standing negative. There are 2 tiers and the split matters. Tier 1 goes on every render, still or video, on every renderer in the stack: glossy, airbrushed, plastic skin, waxy skin, CGI, 3D render, over-smooth, beauty filter, uncanny, doll-like, cartoon, text, watermark, logo, caption, jittery, fast motion, motion blur, warped, extra limbs, extra fingers, fused fingers, deformed hands, morphing face, identity drift, blurry, oversaturated, halo, vignette. Tier 2 is added when the body is the subject, and it locks anatomy and proportions: mismatched proportions, fused digits, asymmetric face, crossed eyes, uneven eyes, elongated neck, distorted torso, double image. Tier 1 is always on. Tier 2 is appended on top of it. The positive defaults are short and they do most of the work: documentary film, 35mm, natural skin texture, visible pores, unretouched, natural light, negative fill, slow deliberate movement, single subject. None of this is a workaround for 1 renderer. It is the arithmetic of any system that fails predictably. The probability that a render survives a review pass rises the moment the shot stops asking for the impossible, and it is the same move that built companies across twelve countries and filled 393 pages: name the limit first, then design inside it. The best AI film is the one that never asks the model to fail.
Skin fails in 2 opposite directions: too smooth, and too detailed.
The map is dead. Nobody told you.
Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.
Beyond this book
Building the same thing somewhere else.
Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.
The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com
18-20 November. Online, Las Palmas, Bali.
Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.
ex-aisummit.com →