
The Shot List Decides What The Model Fails
AI video models fail predictably, and the failure is fixed in the shot list rather than in the model. 4 danger zones cover almost all of it, in a known order of how badly they break. Hands go first. Faces run out of pixels. Skin comes back waxy by default. Text, reflections and fine patterns come back as guesses. The fix is never retouching.
The Failure Is Decided Before The First Frame
The clip came back with 6 fingers on a hand that was never meant to be in focus, and the note that followed said the renderer was weak. It is not the renderer. AI video models fail predictably, and the failure is already settled by the time the prompt is sent, because it lives in the shot list rather than in the model. That is the useful part of it. A defect that repeats on every attempt is a property of the shot rather than luck of the draw. 4 danger zones account for almost all of it, and they sit in a known order of how badly they break: hands first, then faces, then skin, then text and reflections and fine patterns. 1 rule sits above all 4 and costs nothing to apply. Never prompt your way out of a shot the model cannot do. Redesign the shot. The probability that a bad render was bad luck is not zero. The probability that the shot asked for something the model cannot do is far higher, and only the second 1 responds to anything a person can change. So the work moves to the page where the shot is described rather than to the software that renders it. That page is written before the shoot, which means the failure is preventable rather than merely survivable.
Hands Break First, And Most Visibly
Hands are the most information-dense moving object in any frame, and they break first. Extra fingers, fused fingers, a wrong count. If a hand is in the shot, assume it is the tell until proven otherwise. 3 tactics close the zone. Keep hands out of tight focus. Obscure them, or slow them down. And describe the start and the end position of the gesture rather than the path it travels, because the model invents the middle, and the middle is where fingers fuse. Avoid the whole class directly: hand close-ups, sign language, a hand holding a sign or a page, precise gestures, fast hand motion. Every 1 of them asks for a motion path the model has to guess at, and it guesses badly. The safe substitutes are almost free. The back of a hand. A wrist or a forearm crop. 1 phrase does more work than every setting in the renderer: no fingers visible. On a 24 frames per second timeline the 3.5 second cap is 84 frames, and a hand that is going to fuse fuses inside that window rather than after it. Written into the shot list as a fact rather than a hope, that 1 line removes more failed takes than any slider, and it survives the next model release because it describes the shot rather than the tool.
A Face Is A Pixel Budget
A face filling the frame forces the model to invent detail that was never in the pixels it was given, and the invention is where it shows. Extreme close-ups come back waxy or cratered. Long static holds let the jaw drift and the nose widen. A profile turn past 90 degrees smears the features. A 1080 line frame gives a close-up fewer pixels across the face than the model behaves as if it has, which is why medium framing keeps winning. Medium framing is the safe zone. Face shots stay at 3 seconds or under, or they move. 1 hero face per frame, because background faces melt. 1 emotion at a time, because stacked emotions morph into one another. A 60 second piece cut to that rule carries about 17 face shots, and then the cap forces a new angle or a new subject. Talking, chewing and eating close-ups hand the model the hardest problem on the board, which is the inside of a mouth in motion, and the result has a name of its own. The AI mouth. 5 rules follow, and not 1 of them is about quality. Medium frame. Short hold. Movement present. 1 face. 1 emotion. They are constraints on the shot list rather than corrections to the output, which is why they cost nothing and hold. 1 probability is worth keeping in view here. The probability that a waxy close-up could be rescued in post is close to zero. The probability that it renders correctly on the next attempt, with medium framing and biological cues, is high, and that is the only bet on the table.
Waxy Skin And The Anti-Beauty Fix
The model over-smooths to suppress noise, so the default skin is a beauty filter, and the opposite of beautiful language is what reverses it. 4 phrases do the work: natural skin texture, visible pores, unretouched, no makeup. They are not adjectives about style. They are instructions about biology, and they belong in the prompt rather than in the grade. The camera is rarely the problem. Phones shoot 10 bit colour and 48 kHz audio is standard on a laptop, and neither 1 puts a pore back into a smooth face. High-pass overlays, frequency separation and grain cannot add pores that were never in the pixels. If a render comes back waxy, it goes back through the prompt with the biological cues rather than through retouching, and 1 pass shows the difference plainly. Clip length carries the same discipline. Any clip with a face, a hand or a body stays at 3.5 seconds or under, and then cuts to a new angle or an object rather than holding on the same 1. Objects are the safe shots, and so are wide exterior frames with nobody in them. The inversion is useful to hold on to. The hardest things to light are often the easiest things to render, and the easiest things to render are the ones nobody argues about.
AI video models fail predictably, and the failure is fixed in the shot list rather than in the model.
Text, Reflections And Fine Patterns
AI garbles text. A storefront sign, a phone screen, a book cover, a tattoo, a licence plate: every 1 of them comes back as lettering that looks like language and is not, and a viewer reads it as an error inside about 1 second, whether it was shot at 35 mm or on a phone. Reflections fail in a different way. Mirrors, glass and water return a scene that does not match the 1 being shot, and the mismatch is visible even to somebody who could not say why. Fine repeating patterns are the third member of the group. Teeth, picket fences, brick walls and fabric weaves produce moire, a shimmering interference that appears when a pattern is sampled near its own frequency. Pull back or defocus, and it settles. 1 shared cause explains all 3. Each asks the model to reproduce high-frequency detail at a size where it has learned the texture rather than the content, and where the texture is the subject, the subject is the guess. So the rule generalises cleanly. Where the pattern is the point of the shot, the pattern is also the risk. Shoot the object, not the lettering on it. Shoot the room, not the reflection in the window. Shoot the face, not the teeth.
Post Cannot Rescue A Bad Render
The temptation at the end of a shoot is always the same 1, and it is the temptation that costs the most. The render comes back wrong, the grade is opened, and 2 hours go into a repair that cannot work. The reason is arithmetic rather than taste. A repair tool redistributes detail that exists in the frame. The waxy render has no pores to redistribute. The 6-fingered hand has no 5-fingered original underneath it. The garbled sign has no correct spelling inside it. What clears it is re-rendering with the cues, and redesigning the shot when the cues are not enough. That is the whole method, and it fits in 1 line. Never prompt your way out of a shot the model cannot do. The instinct underneath is the 1 that transfers. Every physical build this work sits on was specified before it was built. Companies built across 12 countries. A EUR 75 million industrial group restructured. 210 energy systems across Africa and Asia. 200 wood-gasification machines across the UK and Europe. 16 years of field work from rural Nigeria to Fukushima, and a book of 393 pages published on 8 April 2026. 25,000 applications arrive for 25 places in a cohort, and 210 alumni across 19 countries now run the same discipline in their own work. Not 1 of them was rescued at the end. 1 line carries it out of the room. The model is not the limit. The shot list is the limit, and the shot list is the part you write.
Never prompt your way out of a shot the model cannot do. Redesign the shot instead.
The map is dead. Nobody told you.
Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.
Beyond this book
Building the same thing somewhere else.
Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.
The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com
18-20 November. Online, Las Palmas, Bali.
Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.
ex-aisummit.com →