The Cost of Reducing the Wrong Number
["opinion","production"]A model can be cheaper per token and more expensive per usable feature.
The useful comparison is the total cost of a usable outcome, not a higher number for its own sake.
A less capable model may need more loops to reach the same result. Those loops can add errors, correction, extra prompts, human review and another attempt. A more expensive model may be cheaper overall if it reaches usable output with fewer retries.
That is not an argument for always choosing the most expensive model. It is an argument for measuring the outcome rather than one input to the process. The useful comparison may be cost per usable feature, not cost per token.
Consider this illustrative comparison. One model costs less per token but needs two extra attempts and manual correction before the feature meets its acceptance criteria. Another costs more per token but reaches the same boundary with fewer calls and less review. The second model may have the lower total cost. You can only see that by counting retries, correction time and review time for the same accepted output.
Here, ‘usable feature’ means output that meets the agreed scope, quality, evaluation and support requirements. Those requirements differ by feature, so use a local baseline rather than a universal threshold. Track calls per accepted result, review time and correction work.
It is not limited to AI. A target becomes visible, then its reduction is treated as the result. The value that number was meant to buy disappears from the discussion.
The cheaper number can buy less
Suppose an effort estimate falls from 100 hours to 70. That is an efficiency gain only if the scope, quality and support remain at the agreed level. If testing or future support has been removed, the estimate went down because the work changed. If the same work can be delivered with better tools or execution, the lower effort may be a real gain.
The same pattern appears elsewhere:
- Delivery: a shorter schedule may defer testing, learning or hardening. The date improves; rework arrives later.
- Support: a smaller budget may reduce response capacity. More customer problems remain unresolved.
- Risk: a lower tolerance may be responsible when safety, cash or trust is at stake. It can also remove worthwhile options before anyone has examined them.
None of these makes reduction inherently wrong. It means the reduction needs a visible trade-off.
Choose a unit closer to the outcome you care about:
- cost per usable feature rather than cost per token;
- effort per outcome delivered rather than hours estimated;
- time to a useful release rather than an earlier date;
- support cost per resolved issue rather than headcount alone.
These are not universal formulas. They are prompts to ask what the number is standing in for.
A lower number is a genuine gain when the outcome, scope, quality and risk stay stable or improve. The point is not to defend higher costs. It is to stop a lower cost from being mistaken for better economics before the comparison is complete.
One name for this bias is ‘quantification bias’: favouring value that fits neatly into a number. The Drum’s report on Rory Sutherland’s 2026 Predictions session uses the term when describing his argument. It also describes his ‘doorman fallacy’: removing a visible task can remove less visible functions such as recognition, reassurance, security and status.
The same organisational version appears in Why Standardised Processes Fail Creative Technical Teams. A system can become easier to report upwards while making useful judgement harder at the front line.
Four questions before reducing a number
Before cutting a target, write down four things:
- What outcome is this number meant to buy?
- What unit gets closest to measuring that outcome?
- What scope, quality, capacity or opportunity changes when the number falls?
- What signal would tell us to stop, change course or restore the investment?
For the illustrative feature, the outcome is an accepted feature. The unit is total model and review cost per accepted result. The trade-off is extra retries and correction work. The stop signal is lower first-pass acceptance or review time above the agreed baseline.
Set the signal from a current baseline and the feature’s acceptance criteria, not a universal threshold. For AI, it might be a higher error rate or more review time. For delivery, it might be rising rework or unresolved defects.
If you cannot change the target, record the trade-off anyway. Name the missing scope, its owner and the date when the decision will be revisited. Then ask what support or quality standard must change to make the target workable. A lower estimate without a corresponding change in scope is a promise to discover the missing work later.
Cost matters. Time matters. Effort and risk matter too. The metric is not the outcome.
Put the outcome beside the number before deciding what to reduce. Then write down what the lower number removes and how you will notice.