fbpx
Valuation

AI Gross Margin and Inference Costs

Your margin looks great right up until a buyer rebuilds it. Then costs you have been treating as overhead get moved into the cost of running the product, and the number drops.

In the example further down this post, a reported 82% margin becomes 69% once you move human review, search and retrieval, and AI support into the cost of delivering the thing. That one move is what a buyer is really getting at with every question they ask about your AI costs.

Inference cost is simply what you pay a model provider every time your software asks a model to do something. It used to be a rounding error. Two companies can now run the same feature for the same customer, and one of them pays 5 times what the other pays.

By the end of this post you will be able to build the 3 cost schedules a buyer will ask for, and run 1 test that tells you whether your AI money gets valued like software or like a services business using software pricing.

Why buyers treat inference costs differently from other AI SaaS gross margin questions

The risk being priced is not the size of the cost. It is who controls it.

Every other cost is one you set. You choose salaries. You sign the hosting contract. Inference cost is different, because the provider sets the price and changes it on a date they pick.

That is not theoretical. Google publishes Gemini 3.7 Flash at $0.75 per million input tokens through the end of 2026, and $1.50 from January 1, 2027, per the Gemini API pricing page. A token is roughly three quarters of a word. That cost doubles on a date you did not choose, while your prices stay put.

Now a subtler one. Anthropic’s published pricing documentation notes that its newer models use a tokenizer producing roughly 30% more tokens for the same text. Upgrade at an identical headline price and your bill still rises by about a third.

A cost you do not control

You can hold price, churn, and headcount and still watch gross margin fall because a vendor changed a rate card. Buyers underwrite that possibility, not your current margin.

So the number alone tells a buyer little, and the gross margin range buyers consider normal for SaaS is the wrong yardstick without adjusting for where your costs come from.

The anatomy of the inference cost schedule buyers ask for

Most founders report one blended gross margin. Buyers take it apart. Here is the shape they rebuild, using round illustrative numbers for a company with $1,000,000 of AI product revenue.

Waterfall stepAmountRunning total
AI product revenue$1,000,000$1,000,000
Model API cost-$180,000$820,000
Human review and quality checks-$60,000$760,000
Retrieval and vector storage-$40,000$720,000
Support attributable to the AI feature-$30,000$690,000
True AI gross margin69%

That founder counted only the model API bill. The three lines underneath get moved into cost of revenue during diligence. Human review is the one most often missed, and it says the most, because it means the model is not yet good enough to ship unsupervised.

Key takeaway

Report AI gross margin as a separate line. A blended number hides the mix, and a buyer who unblends it assumes the worst version.

What a quality of earnings review does to these numbers

A quality of earnings review is an independent accountant’s inspection of whether your reported profit is real and repeatable. On AI costs it does three things. It reclassifies anything needed to deliver the product into cost of revenue, including the contractor reviewing model output. It restates free credits and introductory rates at renewal prices. Then it runs sensitivity on rising unit costs.

These adjustments land after the price is agreed, and they move it. How a quality of earnings report affects your valuation is worth reading before you are inside a live process.

The Scale Direction Test

Here is the test I would run before a buyer does. Take your AI gross margin today. Model it at twice the usage, prices flat. Write down which way it moves.

That direction is the whole answer. Classic software margin improves with scale, because the big costs are built once and spread over more customers. Inference cost does not spread. Every call is billed again.

If your margin is flat or falling at twice the usage, your cost of goods grows in lockstep with revenue, and buyers value that shape closer to a services business than to software.

Buyers are not asking what your AI gross margin is. They are asking which way it moves when usage doubles.

The direction is engineered, not fated. Anthropic prices a cache read at one tenth of its standard input price, so repeated context stops being re-billed at full rate. Batch processing runs at a 50% discount on both input and output tokens across Anthropic, Google, and Amazon Bedrock alike.

Routing cheap work to cheap models compounds both: Claude Haiku 4.5 costs $1 per million input tokens against $5 for Claude Opus 5, a five times spread on identical volume.

Run the same illustration two ways, holding price and product fixed. The right column assumes a rising share of repeated context served from cache and non-urgent work moved to batch, which is why it climbs. These are illustrative, not measured.

Usage levelEverything through one frontier modelCached, batched, routed by task
Today69%69%
2x usage69%76%
4x usage69%81%

Put that in price terms, because a buyer will. One who decides your AI line behaves like services rarely announces a lower multiple. They carve the AI revenue out and value it separately, hold back price until the cost schedule proves itself, or push the risk into an earnout. All three land in the same place: less cash at close for the same reported ARR.

Both columns start in the same place. Only one of them is a software business by the time it doubles. This is the same reasoning buyers apply when they decide that not all recurring revenue is worth the same multiple, applied to the cost side of the income statement.

The three schedules to build before you sign an LOI

Build these once, update them monthly. They take a weekend and separate answering from scrambling.

1. Cost per unit of value. Pick the unit your customer buys: a ticket resolved, a document processed, a call summarized. Divide total AI cost by that count. Anthropic’s documentation publishes its own example at roughly 3,700 tokens per support conversation and about $37 per 10,000 tickets. Treat that as a vendor illustration, not your number: input and output tokens bill at different rates, so compute yours with the split shown on the rate card. Having one at all puts you ahead of most sellers.

2. Gross margin by product line. Split AI features from core product, costs assigned rather than spread evenly. Show one blended number and a buyer assumes the AI line is the weak one.

3. Sensitivity by usage and by price. Show margin at current usage, double, and quadruple, then again with unit costs 50% higher. That answers the Scale Direction Test in writing.

This is one input in how buyers are scoring AI risk in your business, and it is the part you can fix fastest, because it is engineering, not market position.

Frequently Asked Questions

What gross margin do buyers expect from an AI SaaS company?

Buyers care more about the trend than the level. A company at 65% with margin rising as usage grows tends to read better than one at 75% with margin flat or falling, because the first shape is software economics and the second is not. Report the direction alongside the number.

Do inference costs count as cost of revenue or operating expense?

Cost of revenue, if the model call is required to deliver the product the customer paid for. Model spend on internal research or experimentation stays in operating expense. A quality of earnings review will make this split whether or not you did, so make it yourself first.

How do I reduce inference costs before selling my company?

Three published levers move the number most: cache repeated context, which Anthropic prices at one tenth of standard input rates, batch work that is not time sensitive for a 50% discount, and route simple tasks to cheaper models rather than sending everything to a frontier model. Do this at least two quarters before going to market so the improved margin appears in the trailing numbers a buyer underwrites.

Will falling model prices fix my margin on their own?

Not reliably enough to plan around. Prices for a given capability have fallen sharply, but published rate cards also move up: Google’s Gemini 3.7 Flash input price is scheduled to double on January 1, 2027. Buyers underwrite the rate card you are exposed to, not the trend line.

Next Steps

If you are not sure whether your AI margin reads as software or services to a buyer, get a free value assessment and we will run the numbers with you.

Book a Free Value Assessment