← Blog

Cost · · 9 min read · Updated

Your AI content program is paying people to clean up cheap drafts

The model bill is the smallest line in most AI content budgets. Measure rejected drafts, review time, rework, coordination, and accepted output before calling the work cheaper.

Jonathan Haas

The model bill is $11.20. The campaign is $548.70.

Which number would you put in front of finance?

Most teams put the first one in a slide and pay the second one in people’s time. That is how an AI content program can look cheap while the team feels buried.

The expensive part is the eight drafts that never ship. Their model charges, source lookups, review comments, and abandoned work do not disappear. The twelve approved assets carry all of it.

The provider invoice is useful. It is also the easiest number to collect. It tells you what a billing account spent. It does not tell you what a campaign cost, which assets survived review, or why the rest died in the queue.

The useful unit is the accepted piece of work. Measure from brief to approval, then keep the model bill inside that larger number.

The five costs hiding behind a small model bill#

Keep these costs separate long enough to see where the work is going. Collapse them only after the team can explain each one.

1. Direct AI spend#

This is what a model or AI feature charges for requests attached to the work:

input tokens × input rate
+ cached input tokens × cached rate
+ output tokens × output rate
+ feature charges

The exact fields depend on the provider. OpenAI publishes current model rates in its API pricing and model catalog. Anthropic documents input, output, cache, batch, long-context, and tool-use pricing in its pricing guide. Keep the provider rate beside the usage event so a price change does not rewrite historical consumption.

Direct AI spend needs a work-item key beside it. A provider API key is a billing principal. It is not a campaign, asset, owner, or reason for the request.

2. Retrieval, tools, and retries#

Content work often calls a document store, search service, translation tool, image generator, or publishing system. Some calls have their own charges. Others add a large amount of context to the next model request.

Capture the tool, request, result status, attempt number, and charge when the call happens. A failed request still consumed time and may have consumed money. A retry that succeeds should remain a retry. If you merge both into one successful average, you lose the clue that the workflow is unstable.

3. Human review, rework, and coordination#

An AI draft that takes five minutes to generate can take an hour to approve. Review time is a cost even when it does not appear on an AI invoice.

active review cost = review minutes ÷ 60 × loaded hourly rate
rework cost = rework minutes ÷ 60 × loaded hourly rate
coordination cost = handoff and meeting minutes ÷ 60 × loaded hourly rate

Count the time spent finding the source, explaining an edit, reopening a ticket, and asking who owns the decision. Those minutes are not “process overhead” in the abstract. They are the labor required to get one publishable asset across the line.

4. Shared campaign work#

Research, campaign setup, localization planning, and approval design can serve several assets. Allocate that work once, with a stated rule. Equal allocation is easy to explain. Allocation by asset complexity may be more accurate. Either is better than assigning the whole amount to the last asset someone touched.

Keep paid distribution separate from content production. A media budget can be large while production is efficient. A small media test can still be expensive to create.

5. Discarded work#

Rejected drafts are a cost category, not an empty cell. Count the request, tool, review, rework, and coordination attached to an asset that was abandoned. Report the cost against accepted output so the denominator cannot hide failure.

If an event cannot be joined to a brief, task, or asset, call it unreconciled. Do not call it free. “Unknown” is a follow-up queue. Zero is a false conclusion.

The $11.20 trap#

Consider an illustrative campaign with twenty drafts and twelve approved assets. The team uses a loaded rate of $110 per hour:

Cost line Amount
Model requests $7.80
Retrieval and tool calls $3.40
First review: 2.25 hours at $110/hour $247.50
Rework: 0.75 hours at $110/hour $82.50
Handoffs and coordination: 0.25 hours at $110/hour $27.50
Shared campaign research and setup $180.00
Total production cost $548.70

The direct AI and tool bill is $11.20. The full production cost is $548.70. The direct bill is 2% of the cost that produced the campaign.

The cost per generated draft is $27.44. The cost per accepted asset is $45.73. The first number rewards activity. The second tells you what the team can publish.

The eight rejected drafts are not eight free experiments. Their cost is carried by the twelve assets that made it through. If the next campaign has the same model bill but only six accepted assets, cost per accepted asset nearly doubles.

Cheap per draft does not mean cheaper than before#

An AI cost report is not a savings report. To claim savings, compare the same accepted output with a credible baseline:

realized savings
= comparable baseline cost for accepted output
  − observed full production cost

The baseline should include the same review, source gathering, coordination, and shared work. A saved hour is not a saved dollar when the team fills that hour with more requests, more review, or more meetings. If payroll and agency scope did not change, call the result capacity or cost visibility. Do not call it cash savings.

The CFO question is not “How little did the model cost?” It is “What approved work did this spend produce, and what did we stop paying for?”

That question forces three answers:

  1. Which assets were approved, rejected, or abandoned?
  2. What human time did each outcome consume?
  3. Which old cost actually moved: agency hours, contractor scope, launch delay, or nothing yet?

The evidence fields that make the number survive a follow-up#

Attach these fields when the brief starts:

Field Why it matters
Campaign and task ID Joins every request to the same piece of work.
Asset ID and type Separates an email variant from a landing page or research note.
Owner and approver Gives someone responsibility for the result and the decision.
Source set and source version Shows which facts and claims informed the output.
Acceptance rule Defines what “done” means before the draft arrives.
Budget envelope Shows what the task was allowed to consume.

Attach these fields to each AI or tool event:

Field Why it matters
Provider and model Explains the direct charge and supports price reconciliation.
Input, cached input, and output usage Reproduces the usage calculation.
Tool calls and attempt number Exposes work hidden by a successful final response.
Status and timestamps Shows whether the event completed, failed, or waited.
Work-item ID Lets finance trace a charge back to a brief or asset.
Evidence state Distinguishes reconciled, partial, and missing cost proof.

At the end, add the reviewer, decision, rejection reason, accepted time, and final status. A cost number without an outcome is a usage report. A cost number attached to an accepted asset is an operating measure.

What the provider export cannot answer#

The export is the beginning of the investigation, not the answer. It can show usage. The team still needs to answer:

  • Why did this task make twelve attempts?
  • Why did eight drafts fail the same product requirement?
  • Why did a claim reviewer need to search three systems for a source?
  • Why did an approved asset sit for two days before publication?
  • Which events are missing a usable cost join?

Those questions are where cost turns into a workflow decision. If the failure is a stale source, fix the source. If it is an unclear acceptance rule, tighten the brief. If it is a repeated tool failure, stop paying for retries until the path is repaired.

Projection is a starting point, not evidence#

Writer’s Marketing AI ROI Calculator asks about company context, agency spend, use cases, publishing volume, review rounds, and the existing tool stack. That is a reasonable way to size a program and build a first business case.

It cannot answer which campaign consumed yesterday’s spend, which asset needed three revisions, or whether a provider response had usable usage data. The result is a forecast. The distinction matters when a forecast becomes the budget target.

Use two views together:

  1. Before the work: estimate volume, direct cost, review capacity, and the old cost you expect to replace.
  2. After the work: reconcile requests, tools, review, accepted output, rejected output, and missing evidence.
  3. At the next budget review: compare the forecast with observed cost per accepted asset and change the assumptions that missed.

Every assumption should eventually have an observed number beside it. If it does not, label it as an assumption instead of letting it become a fact through repetition.

The weekly review that makes the pain visible#

Marketing ops and finance should be able to answer these without rebuilding the story from email, tickets, and provider consoles:

  1. Which campaign used the most direct AI and tool spend?
  2. Which accepted asset had the highest full production cost?
  3. Which task created the most rejected work?
  4. Which review or approval step consumed the most human time?
  5. Which workflow has the longest wait before approval?
  6. Which events have missing or contradictory cost evidence?
  7. What will we pause if the next forecast crosses the budget envelope?

The last two questions are controls. Missing evidence is not zero cost. A threshold without a named action is not a control either.

Deixic gives the team the evidence layer behind these questions: spend by agent, activity tied to the work, gated approvals, connected tools, and an explicit signal when cost proof is missing. Pair that product view with the team’s review-time and agency ledger, then report the cost of accepted work instead of the model invoice alone.

A two-week measurement plan#

Choose one repeatable process: brief review, claim checking, product description updates, or campaign quality assurance. For two weeks:

  1. Attach a work-item ID to every model and tool request.
  2. Count submitted, accepted, rejected, and abandoned assets.
  3. Record active review minutes, coordination minutes, and the reason for rework.
  4. Measure approval delay separately from time spent editing.
  5. Reconcile the total against provider and tool bills.
  6. Compare the result with the old process using the same accepted output.

At the end, you should have direct AI and tool spend, human review and rework cost, full production cost, cost per accepted asset, and a list of evidence gaps. That is enough to find the first bad assumption and answer the follow-up question: “Which work produced that number?”