← Blog

Cost · · 9 min read · Updated

Your AI budget is wrong before finance sees the invoice

A budget built from provider dollars misses retries, review, agency work, and unapproved tasks. Set limits by process, forecast from accepted output, and name the action behind every threshold.

Jonathan Haas

An AI content budget should answer one question before the month ends: what work can we still afford to do?

The provider invoice answers a different question: what did all of the requests cost after they happened?

Both numbers matter. The first is a control. The second is a reconciliation. Teams get surprised when they use the second number as their only budget process.

A budget ceiling without a work ID, an owner, and a stop action is a forecast wearing a control’s name. It can tell you that money went out. It cannot tell you what to pause before more money goes out.

Why the invoice arrives too late#

The invoice is usually grouped by account, key, workspace, or provider. Marketing work is grouped by campaign, asset, audience, and approval path. Those are different shapes.

By the time finance receives the bill:

  • a shared API key has mixed several campaigns together;
  • retries and failed tool calls have disappeared into a daily total;
  • a high-context request has made one task much more expensive than its peers;
  • reviewers have spent hours cleaning up outputs that looked cheap to generate;
  • agency or contractor work sits in another system;
  • a campaign has missed the date it was meant to support.

The invoice can be accurate and still be insufficient for the budget decision.

Give the budget three levels#

Start with a hierarchy that matches how work gets approved:

  1. Portfolio ceiling: the total monthly amount for AI-assisted content work.
  2. Process envelope: the amount reserved for a repeatable process such as brief review, product descriptions, localization, or campaign quality assurance.
  3. Task threshold: the amount or condition that requires a person to review before more work proceeds.

The portfolio ceiling tells finance how much can be spent. The process envelope tells marketing ops where the money is going. The task threshold catches an expensive request while someone can still ask why.

If nobody can name what happens when a threshold trips, the threshold is decoration. The action may be pause, approval, a lower-cost path, or an investigation. Pick one before the warning arrives.

A small budget example#

Imagine a team sets a monthly portfolio ceiling of $60,000:

Process envelope Monthly limit Stop or review condition
Product content updates $18,000 Pause when accepted cost exceeds the recent 90th percentile.
Campaign quality assurance $15,000 Require an owner when a task retries more than twice.
Localization $12,000 Review when source evidence is incomplete.
Brief and claim review $10,000 Escalate when approval age passes one business day.
Unallocated reserve $5,000 Finance approval before use.
Portfolio ceiling $60,000 Review forecast every week.

The exact amounts are illustrative. The important part is the connection between a limit and a decision. A row that says “$15,000” without a process owner or stop condition is not managing spend. It is labeling a pile of spend.

Build the baseline from observed work#

Use the last four to eight weeks of actual activity. For each process, collect:

  • submitted tasks;
  • accepted, rejected, and abandoned outputs;
  • direct model and tool spend;
  • retries and high-cost requests;
  • active review, rework, and coordination minutes;
  • approval delay and missed planned dates;
  • missing cost and source evidence.

Do not start with a percentage of the marketing budget. That creates a limit without a model of the work. Start with the processes the team expects to keep doing, then add the amount of capacity the next quarter requires.

The direct portion of the forecast should reference the provider’s current rate card. OpenAI’s API pricing and Anthropic’s Claude pricing documentation describe different model, cache, batch, context, and tool charges. A budget that stores only “AI cost per piece” will go stale when the process or provider changes.

The failure modes that make a budget lie#

Budget habit What it hides Better question
One shared API key Which campaign and owner used the money Can every event join to a process and task?
Cheapest model wins Review, rework, and repeated generation What is the cost per accepted asset?
Missing usage becomes zero Unexplained spend and broken joins Which events are unreconciled, and who fixes them?
A spend cap is the whole control The request can still waste the cap What action stops the next expensive attempt?
Agency savings are assumed Work may move in-house without cash savings Did an invoice, contract, or scope actually change?
More output is called value Rejected drafts and extra review Which accepted result improved?
An average hides the tail A few expensive tasks consume the reserve What are the median and 90th-percentile costs?

The most dangerous row is missing usage. Treating missing evidence as zero makes the forecast look healthy and leaves the unexplained spend for the invoice.

Forecast the rest of the month#

For a process with stable volume, use the full cost of accepted work:

month-end forecast
= spend so far
+ expected remaining accepted tasks × observed cost per accepted task

Use the median cost for the likely case and the 90th-percentile cost for the warning case:

likely case = spend so far + remaining tasks × median accepted-task cost
warning case = spend so far + remaining tasks × p90 accepted-task cost

Add committed agency, contractor, and tool costs separately. Do not pretend a provider forecast includes costs that live in another ledger.

For a new process, use a range rather than a single number:

low case = expected tasks × low observed cost
high case = expected tasks × high observed cost

The range should widen when the process has missing usage, many retries, or an approval path that has not been exercised. False precision is not a budget control.

Set thresholds around reasons, not only dollars#

A dollar threshold is useful. It is not enough on its own. Set a review condition when:

  • one task costs several times the process median;
  • a request retries after a tool or model failure;
  • a task switches to a long-context or high-cost model;
  • a source lookup returns no usable evidence;
  • spend arrives under a shared key with no work-item attribution;
  • first-pass acceptance falls below the agreed floor;
  • the month-end forecast crosses the process envelope.

Each condition should lead to a clear action: pause the task, ask for approval, switch to a lower-cost path, fix the source, or investigate the missing evidence. A warning that nobody can act on becomes another report.

Use variance to find the process problem#

When actual spend differs from the forecast, classify the variance before changing the budget:

Variance Likely question
More submitted tasks Did campaign volume increase, or did the process repeat work?
Higher cost per accepted task Did context grow, the model change, or retries increase?
More review time Did source quality fall, or did the approval class change?
Lower accepted output Are drafts failing the same requirement?
Longer approval delay Which reviewer or decision is blocking the date?
More agency or contractor work Did the internal process fail, or was demand outside its scope?
Missing cost proof Which provider, proxy, or task join is incomplete?

Variance is a diagnosis queue. Increasing the budget before answering the question can fund the failure mode for another month.

A fifteen-minute weekly budget review#

The operator reviewing the budget should see five decisions, not a raw export of every request:

  1. Spend to date: by process, agent, and owner.
  2. Forecast: likely and warning month-end totals, with assumptions visible.
  3. Largest variance: the task or process furthest from its recent range.
  4. Pending decisions: tasks waiting for an owner, approval, or evidence.
  5. Unreconciled usage: events that do not have a usable cost join.

Add one final line: What will stop today if the warning case becomes likely? If the answer is blank, the team has a report and no control.

Savings need a before-and-after test#

An AI budget can show that direct spend increased. It cannot prove that the old cost disappeared.

To claim savings, compare the same accepted result with a credible baseline:

realized savings
= comparable old-process cost
  − current full production cost

Include review, coordination, agency, contractor, tool, and launch-delay costs where they belong. A team that produces twice as many drafts with the same people may have gained capacity. It has not automatically cut spend.

“We saved 1,000 hours” is a capacity statement. “We reduced the agency scope by $X” is a savings statement. Finance needs to know which one happened.

What a good budget refuses to assume#

It refuses to assume every generated draft becomes an asset. It refuses to assume a lower token rate creates a lower content cost. It refuses to assume a process will keep the same context size as it grows. It refuses to turn a vendor ROI projection into an achieved result.

Writer’s calculator is useful for a starting projection because it asks about agency spend, publishing volume, review rounds, and the current tool stack. The projection still needs to meet actual task, approval, and invoice evidence before it becomes a budget fact.

The budget should become more accurate as the team observes work. Each accepted task gives the next forecast a better cost range. Each rejected task explains where review capacity is going. Each missing evidence item gets an owner.

The Deixic layer#

Deixic gives the team the operational view between a process and a provider invoice: spend by agent, current activity, connected tools, human approvals, activity history, and a visible gap when cost proof is missing. That makes the budget review actionable while work is still in progress.

Use Deixic’s product view for spend and evidence. Keep the team’s time, agency, contractor, and distribution costs in the finance ledger. Join them on the process or task ID. The result is a budget that can explain both the AI bill and the human work around it.

The first budget can be small#

Choose one process with a named owner. Set a monthly envelope from recent activity. Add a task threshold that requires review. Check the forecast once a week. At month end, compare direct spend, accepted output, review cost, approval delay, and missing evidence.

If the team cannot answer why the forecast moved, the next investment is measurement. If it can answer, the decision is clear: raise the envelope, fix the process, or stop funding work that does not produce an accepted result.