The Coding Table Reads By Price Per Million
Execution·Framework·6 min read

The Coding Table Reads By Price Per Million

The table reads by price per million tokens, not by reputation. 5 models sit side by side on 4 axes, and the first column is 2 numbers: dollars in and dollars out. DeepSeek V4 Pro at $0.435 in and $0.87 out against Kimi K3 at $3.00 in and $15.00 out is a 17 times spread on output, and the reasoning inside the second model never turns off and bills as output, which is why the cheapest headline rate is not the cheapest run.

01

5 Models, 4 Axes, 1 Column That Decides

The table reads by price per million tokens, not by reputation. That order is the whole method. 5 models sit side by side on 4 axes: price, context, vision and reasoning. The first column is 2 numbers, dollars in and dollars out per 1,000,000 tokens. Everything else on the row qualifies them. DeepSeek V4 Pro reads $0.435 in and $0.87 out. Kimi K3 reads $3.00 in and $15.00 out. Both hold a 1,000,000 token context window, and both are the same shape of tool. Price is the only axis where they disagree by an order of magnitude. A cache hit runs $0.004 on the first row and $0.30 on the second, which is a 75 times spread on the same repeated read of the same prompt. 4 axes and 1 column. Read the column first. The reason the column wins is that the other 3 axes are filters rather than scores. A context window either holds the repo or it does not. Vision either reads the frame or it does not. Reasoning either stays on or it can be switched off. None of them are negotiated by preference.

02

The 17 Times Spread And The 30 Times Ladder

The obvious comparison is the 2 headline models: a 17 times spread on output and just under 7 times on input. The ladder underneath carries 8 rows. DeepSeek V4 Pro at $0.87 out, MiniMax M3 at $1.20, GLM-5.3-flash at $0.50, grok-build-0.1 at $2.00, Grok 4.3 at $2.50, GLM-5.3 at $4.40, Grok 4.6 at $6.00, and Kimi K3 at $15.00. Cheapest to priciest is $0.50 against $15.00, which is 30 times the same unit for the same token. The input side runs $0.15 to $3.00, which is 20 times. So the pair comparison understates the ladder. 17 times is the number that gets quoted. 30 times is the number that gets paid. Every row is quoted in the same unit, dollars per 1,000,000 tokens. Nothing needs converting, and that is what makes the spread usable on a Tuesday afternoon with a build broken. A spread that survives the unit check stops being a preference about models and becomes a budget decision about work. Most teams quote the pair. The invoice reads the ladder.

03

The Cheapest Headline Rate Is Not The Cheapest Run

The reasoning inside Kimi K3 never turns off. It is always on, and it bills as output. That makes $15.00 the floor of the output line rather than a ceiling on it, and the visible rate separates from the paid rate. The other rows keep a switch. DeepSeek V4 Pro and MiniMax M3 both carry a toggle, GLM-5.3-flash runs light reasoning, and GLM-5.3 has thinking forced on with a measured 9.8 seconds to the first token. Forced reasoning earns its cost on architecture work and burns it on a file rename. So 1 question decides the row before any benchmark is opened: how much of the output bill is reasoning that nobody requested? The probability that a model with reasoning forced on spends under half its output tokens on output that was actually asked for sits near 1 in 10, against near 9 in 10 for the same task on a model with the toggle left alone. Read the reasoning column before the price column. A rate that includes unrequested work is not a rate. 2 of the 8 rows bill reasoning by design. That is not a defect. It is a line item, and it belongs in the estimate.

04

Vision Is A Route, Not A Footnote

DeepSeek V4 Pro is text-only. It cannot see a screenshot. That single cell removes it from every task where a frame has to be read: a UI to rebuild, a rendering bug to hunt, a chart to transcribe, a receipt to file. 5 of the rows carry vision in some form. The 2 that do it natively are MiniMax M3 and Kimi K3, which is why a screenshot-to-fix loop routes there and nowhere else. The context column disqualifies rows the same way. Grok 4.6 holds 500,000 tokens and grok-build-0.1 holds 256,000, so both are out of a full-repo pass before price is discussed. 5 rows hold 1,000,000 tokens: MiniMax M3, DeepSeek V4 Pro, Kimi K3, GLM-5.3 and Grok 4.3. A missing capability is not a trade-off to weigh. It is a route that does not exist, and no price on that row changes the answer. Filter on the columns that can disqualify a row, then sort the survivors by price. That order is the whole routing table. Cheap and blind is still blind.

The table reads by price per million tokens, not by reputation.

05

Sunk Cost Flips The Ranking

MiniMax M3 sits on the Max plan at $50 a month, with roughly 5,100,000,000 M3 tokens a month already paid for. At our volume that makes it free, against pay-as-you-go on every other row. The $0.30 input rate and the $1.20 output rate stop mattering once the plan is settled. So the ranking inverts on sunk cost. M3 goes first for bulk work, and DeepSeek V4 Pro becomes the cheapest overflow at $0.87 out when the plan runs dry. M3 is Anthropic-compatible, which means Claude Code, Cursor and 11 other tools reach it without an adapter in between. The routing that falls out of it reads in 6 lines: bulk agentic work to M3, cheapest overflow to DeepSeek V4 Pro, hardest reasoning and refactors to GLM-5.3 at $4.40 out with a 128,000 token output, visual work to Kimi K3 or M3, the fast loop to grok-build-0.1 or DeepSeek Flash or GLM-5.3-flash, and the 1,000,000 token full-repo pass to M3, DeepSeek, Kimi K3 or GLM-5.3. 6 routes, 1 table, and no row wins everything. Grok takes neither end of any column, which makes it a fine 1,000,000 context generalist and a fast agent, and a row with no assignment on price or on quality. The middle of a table is where rows go when nothing about them is decided.

06

1 Probe Before The Route

Kimi K3 shipped on 16 July 2026, after our key was last tested in May 2026. So the instruction is 1 probe rather than 1 assumption: call kimi-k3 on the Moonshot endpoint and confirm the key answers on the new model. Until that answer exists on the screen, the row is treated as unavailable rather than merely expensive. A vendor's published price is a fact about the vendor. Your access to it is a fact about your key. 2 facts, 1 table. The table holds the first and says nothing at all about the second. The same rule that runs this table runs everything else. Companies built across 12 countries, a EUR 75 million industrial group restructured, 210 energy systems across Africa and Asia and 200 wood-gasification machines across the UK and Europe all run on the measured version rather than the assumed 1. Pick the row from the table, verify the key before the run, and let the work pay for itself. The cost of a wrong route is paid in tokens on every request after it, which is the only reason a price column is worth reading twice.

The first column is 2 numbers. Everything else on the row qualifies them.

The map is dead. Nobody told you.

Bali State of Mind is the survival guide for the collapse of everything you were taught to believe.

Beyond this book

Building the same thing somewhere else.

Julien Uhlig is available for advisory work, board seats and media appearances. Write to media@exventure.co.

The academy that trains the operators, across every company in the group, is EX Epic Academy - 25,000 applications, 25 seats per cohort, 210 alumni across 19 countries. academy.epicsolutiongroup.com

EX-AI Summit 2026

18-20 November. Online, Las Palmas, Bali.

Three days on what happens to work, capital and institutions when the map stops matching the ground. Seats are limited by cohort.

ex-aisummit.com →