The Plan Asks For Volume Not For Heroes
Execution·Framework·6 min read

The Plan Asks For Volume Not For Heroes

The routing question is not which model is best. It is which rows have a marginal cost of almost 0 dollars. 1 fixed monthly plan carries a single shared quota of about 5.1 billion tokens a month across a language model with a 1M context, an image model at up to 2048 pixels, speech in 40 languages, video at 6 seconds and 768P, vision and web search. That fixed monthly is what makes the work free at the margin, and 1 cost column decides the whole table: 5 rows to the plan, 7 rows to the specialists.

01

The Question Is Marginal Cost, Not Quality

The routing question is not which model is best. It is which rows have a marginal cost of almost 0 dollars. A quality ranking puts 9 models in 1 column and leaves the routing undecided. A cost column decides 12 rows in 1 pass. 1 fixed monthly plan carries a single shared quota, and a quota that is already paid for prices the next call at 0 dollars. That is the whole test. If the next call costs 0 dollars and the row is high volume, the row belongs to the plan. If the next call costs 40 cents and the row needs 1 of the 3 things that only 1 model does, the row belongs to the specialist, whatever the quality column says. The probability that a routing decision made per vendor still holds at 3 months sits near 1 in 3, against near 9 in 10 for the decision made per row. Per row beats per vendor. Every time.

02

What 1 Fixed Monthly Actually Carries

The plan carries 1 language model with a 1M context window, 1 image model at up to 2048 pixels, speech in 40 languages and 300-plus voices, video at 6 seconds and 768P, image and video understanding, and a web search. The language quota is about 5.1 billion tokens a month, and it is shared rather than split per capability. 1 subscription, 1 allowance. A video draft does not bill separately from a language call, and 40 languages of speech do not bill separately from 2048 pixel stills. That is free at the margin. The invoice does not move when the volume moves. A shared quota is also a shared ceiling. 1 heavy day of image work spends the same allowance the coding agents were going to draw on, so the ceiling gets watched rather than assumed. The plan also stops at 2 known walls: the content filter, and a video ceiling between 6 and 10 seconds.

03

The 5 Rows That Go To The Plan

5 capabilities move to the plan, and every 1 of them is high volume. Coding agents. The plan is wired for it through an Anthropic-compatible endpoint, so the same CLI works and the tokens come out of the shared allowance. Long-context document work. A 1M window swallows a 300 page specification or a whole codebase in 1 pass rather than 12 chunked calls. Bulk image generation. 2048 pixels plus a subject reference for consistency covers thumbnails, social cards and article heroes at volume, which is the row where per image pricing normally shows up on the invoice. Draft and animatic video. 6 seconds at 768P is enough to time a cut and check a composition before the final render, and a draft that costs 0 dollars changes how many drafts get made. Bulk speech. 40 languages and 300-plus voices covers voice work across markets without a per character meter. The pattern is the point. Not 1 of the 5 rows is a hero shot. All 5 are rows where the count is high and the tolerance is wide.

04

The 7 Rows That Stay With The Specialists

7 capabilities stay exactly where they are, and the reasons are concrete rather than stylistic. Premium and uncensored stills stay. The plan's image model is censored, and a censored model cannot be talked out of it. Uncensored video stays. Same wall, higher cost per attempt. 15 second and 2K hero shots stay. The plan's video ceiling is 10 seconds, and the tier that would break it is metered rather than included. Lip-sync stays. It is a specialist capability with a specialist failure mode, and a wrong lip-sync cannot be saved in post. Dialogue stays. A model that reads a line and a model that performs a line are different tools at a different price. Premium voice stays, and music stays. The plan's music line is retired, which leaves 1 row with nowhere else to go. Not 1 of the 7 is expensive because it is fancy. Each 1 is the only tool that does the row, and that is the only reason it is on the list.

The routing question is not which model is best. It is which rows cost almost 0 dollars at the margin.

05

The Language Model Is A Fallback, Not The Incumbent

The plan's language model is strong enough to be the default and it is not, and the reason is continuity rather than quality. Swapping the incumbent out to save a monthly is a migration, and the plan is already paid for either way. The useful move is smaller. The plan becomes the first fallback, and 1 line of config closes the empty fallback list that had been flagged as a single point of failure. 1 default, 1 fallback, 1 named backup for vision. That is the whole change. Vision stays a backup because the gate that scores the work is already wired and higher quality, and a gate is not the row to trade for a small discount. A fallback that is already paid for is the cheapest insurance in the stack. It costs 0 dollars until the day the incumbent is down.

06

1 Principle Holds Every Row

Use the plan for high-volume, cost-insensitive work. Reserve the only-1-can-do-it rows for the models that earn them. That is 1 sentence and it decides all 12 rows. The volume goes to the tool that is already paid for, and the specialist is called when it is the only one who can hold the row. It is the same rule that ran across the field work. Companies built across twelve countries, 210 energy systems across Africa and Asia, and 200 wood-gasification machines across the UK and Europe all ran on a fixed fleet doing the volume and 1 specialist called for the row nothing else could hold. A plan that is paid for whether it is used or not should be used. A specialist that bills for the row it is best at should be kept for that row. Do the routing once, in 1 cost column, and volume stops being a budget conversation.

A quality ranking puts 9 models in 1 column and leaves the routing undecided.

The map is dead. Nobody told you.

Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.

Beyond this book

Building the same thing somewhere else.

Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.

The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com

EX-AI Summit 2026

18-20 November. Online, Las Palmas, Bali.

Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.

ex-aisummit.com →