Cost · · 5 min read · Updated
The cheapest model can make your content bill bigger
Model pricing is only one part of content cost. Route requests by risk and acceptance rate so a lower token price does not create more review and rework.
Choosing the lowest model rate feels like a simple way to reduce an AI content budget. It can make the budget larger when the cheaper output needs extra attempts, longer review, or a full rewrite.
The useful comparison is the cost of an accepted asset. Model usage is one line in that calculation.
A low request price is a good input. It is a poor definition of a low-cost result.
Start with the accepted asset#
For each content type, measure the complete path:
| Cost | What to include |
|---|---|
| Model and tool usage | Provider requests, retrieval, images, and exports |
| Attempts | Initial drafts, retries, and regenerated sections |
| Review | Editorial, brand, product, legal, or localization time |
| Repair | Human edits and another request after a return |
| Coordination | Brief changes, handoffs, and version comparison |
| Delay | Work that misses a campaign or launch window |
Then divide by accepted assets, rather than total drafts. A team that generates 100 drafts and accepts 20 has a different cost per result from a team that generates 30 drafts and accepts 25.
A small routing example#
Imagine two paths for a product email. A lower-priced model costs $0.08 per attempt. A stronger model costs $0.35.
| Path | Model cost | Attempts | Review and repair | Total per accepted email |
|---|---|---|---|---|
| Lower-priced model | $0.24 | 3 | 22 minutes at $90/hour | $33.24 |
| Stronger model | $0.35 | 1 | 8 minutes at $90/hour | $12.35 |
The amounts are illustrative. The point is the denominator. The stronger model costs more at the request level and less at the accepted-email level because it creates less repair.
Do the same calculation for product pages, customer stories, paid ads, internal announcements, and localization. A single model ranking across every type hides the work created after the response arrives.
Route by the work, not by a model leaderboard#
Use the lightest path that has a reliable acceptance rate for the task:
| Content request | First path | Escalate when |
|---|---|---|
| Format cleanup or metadata | Lower-cost model | Required fields are missing |
| A short variation from an approved source | Lower-cost model | The claim changes or the source is unclear |
| Product positioning | Stronger model with source context | The brief has competing goals |
| Legal, pricing, or security claims | Stronger model and human review | A source cannot support the sentence |
| Multi-market adaptation | Market-aware path | The glossary, policy, or local rule is missing |
| Final repair after a return | Original owner and a focused request | The same issue appears twice |
The routing rule should reference task facts a reviewer can see. “Use the best model” is not a rule. “Use the stronger path for a public claim with a source change in the last 30 days” is a rule someone can test.
Measure the tail#
The average hides the requests that consume the budget. Track at least:
- median and 90th-percentile model spend per accepted asset;
- attempts per accepted asset;
- return rate by content type and model path;
- review minutes per accepted asset;
- source or claim failures;
- percentage of requests escalated to a higher-cost path;
- assets accepted after the second or later attempt.
The 90th percentile gives an operator a warning case. If the typical landing page costs $14 to accept but the 90th percentile costs $68, a monthly budget based on the average will fail during a launch or a difficult brief.
The unit economics guide covers the shared denominator. The cost-per-asset guide shows how rejected drafts and reviewer time belong in the same calculation.
Give every escalation a reason#
An escalation should leave behind a reason that can improve routing later:
| Reason | What to learn |
|---|---|
| Source was incomplete | Improve the brief or source preparation |
| Claim needed a specialist | Create a content class with earlier review |
| First path produced a structural error | Change the template or model assignment |
| Audience or channel changed | Treat it as a new request, not a retry |
| The reviewer returned a preference | Decide whether the preference deserves a rule |
Do not route every difficult request to the most expensive path. Fix repeated failures at the brief, source, or template boundary. The cheapest request is often the one that does not need a second pass.
Set a stop condition#
Routing needs a point where another attempt requires a person to decide. For example:
- allow two attempts for a structured variation;
- return the item with a reason if both fail;
- require the owner to confirm that another attempt is worth the cost;
- record whether the item was accepted, rebriefed, or closed.
This prevents an automated process from spending through a weak request simply because each individual call looks inexpensive.
Review the decision every month#
Model rates, context sizes, and campaign mix change. Recalculate accepted cost when:
- the provider or model changes;
- the source set grows;
- a new market or channel is added;
- the reviewer or approval path changes;
- a template produces a different return pattern;
- the budget owner changes the required output.
Keep the old and new paths visible for one review period. A routing change should show whether the accepted cost, review time, and output quality moved together.
Deixic gives marketing and finance a view of spend by agent, current activity, approvals, and missing cost proof while work is still in progress. Use the product view to inspect spend and evidence, then keep review labor and campaign outcomes in the team’s normal planning systems.