← Blog

Cost · · 5 min read · Updated

Cost per token is not the cost of a finished asset

A practical way to measure AI-assisted content by accepted assets, review time, rework, and launch delay instead of provider usage alone.

Jonathan Haas

Cost per token is a provider metric. Cost per accepted asset is a management metric.

That difference is why a marketing team can report lower model spend while its content budget gets larger. The model bill is visible. The labor that repairs, reviews, coordinates, and abandons the output is scattered across people and systems.

The unit that matters is the thing the business can use, not the request that produced it.

If a team paid for 96 drafts and published 24 assets, the denominator is 24. A dashboard that divides the provider invoice by 96 makes the work look four times cheaper than it was.

Choose a denominator someone can act on#

There are several useful units, but they answer different questions:

Unit Good for Bad for
Request Provider reconciliation Deciding whether a campaign worked
Draft Capacity planning Measuring finished work
Review item Staffing a quality process Comparing campaigns with different review rules
Accepted asset Budget and process decisions Measuring work that never reaches acceptance
Published asset Launch economics Work accepted but delayed by launch capacity

Use requests when you need to reconcile a provider bill. Use accepted or published assets when you need to manage the content program. Keep both values; do not force one to stand in for the other.

Build the full cost stack#

An accepted asset usually carries more than a model charge. A useful cost stack has at least these layers:

  1. Provider and tool spend. Model calls, retrieval, image generation, transcription, seats, and exports.
  2. Production time. The marketer or specialist who shapes the brief, checks the sources, and repairs the draft.
  3. Review time. Brand, legal, product, regional, and channel reviewers.
  4. Coordination time. Chasing owners, resolving comments, moving files, and reconciling versions.
  5. Discarded work. Requests and drafts that never become an accepted asset.
  6. Delay. A launch miss or a campaign window that closes while the item waits for a decision.

Do not hide every layer inside a single estimated hourly rate. Keep the raw facts visible, then choose the costing rule the finance team accepts. Direct provider spend and fully loaded production cost are both useful when their labels are honest.

A small example changes the conversation#

Imagine a month with 40 content briefs:

Cost layer Amount What it covers
Models and tools $1,100 Requests, seats, and exports
Production time $4,800 Briefing, drafting, and repairs
Review time $3,600 Brand, product, and legal review
Coordination $900 Handoffs and version cleanup
Abandoned work $300 Drafts with no accepted outcome
Total $10,700 24 accepted assets

The full cost per accepted asset is $446. The provider-only cost is $46. Both numbers are correct. Only one tells the marketing leader what an additional accepted asset currently costs.

Now suppose the team changes the brief and approval process. The next month produces 30 accepted assets:

Cost layer Amount
Models and tools $1,500
Production time $4,000
Review time $2,100
Coordination $600
Abandoned work $400
Total $8,600

The provider bill rose by 36%. Full cost fell to $287 per accepted asset, a 36% improvement, because the team produced more usable work with less repair and review. A model-only dashboard would call the first month better.

Track the fields that connect money to work#

Each cost event needs a small amount of business context:

Field Example
Brief ID spring-launch-email-07
Asset type Product email
Campaign Spring launch
Owner Lifecycle marketing
Attempt count 3
Review rounds 2
Outcome Accepted, published
Direct spend $4.82
Human minutes 74
Evidence state Reconciled

The exact field names can vary. The connection cannot. If finance has a model total and marketing has an asset total with no shared ID, the month-end report will depend on guesses.

This is where attribution by campaign, task, and owner matters. It gives each cost event a path back to the person who can explain the work or stop it.

Separate capacity from savings#

Lower cost per accepted asset can mean several things:

  • the team published more with the same staff;
  • the team reduced outside production spend;
  • reviewers spent less time on routine work;
  • the team avoided work that should never have started;
  • the team moved faster and captured a campaign window.

Those are different outcomes. Do not label all of them “savings.” A team that uses the freed capacity to support another launch may have created value without reducing the next invoice. The savings proof guide explains how to keep those claims separate.

Put a decision beside the metric#

A metric without a decision becomes another report. Pair each measure with a response:

Signal Decision
Accepted cost above the campaign envelope Pause new work and inspect the largest cost layer
Review minutes rising while provider spend is flat Repair the brief, source set, or approval rule
Attempt count above two Require an owner to classify the failure before another request
High draft volume and low acceptance Stop celebrating volume; inspect the denominator
Unattributed spend above the threshold Hold the amount as an exception until an owner resolves it

The goal is not perfect accounting. The goal is a cost model that changes what the team does next.

Use the cost-per-asset framework for the detailed production denominator, then add the accepted-asset view to the budget conversation. When the number has an owner, a unit, and a decision, AI content stops being a cheap line item and becomes manageable work.