AI inference: what happens and what each answer costs
AI inference is the work of producing an answer with a trained model. Learn how tokens, settings, retries and completed tasks affect the bill you actually pay.

Quick Answer: AI inference is the work a trained model performs to answer a request. Buyers pay for input and generated tokens, sometimes at different cached-input rates. The useful comparison is total spend per completed task, including retries and output length, not a single token price. Check billing categories before estimating cost.1,3
Summary in a mind map
AI inference: what happens and what each answer costs │ ├─ A request becomes an answer │ ├─ A trained model receives input │ ├─ Inference generates output │ └─ Completion needs a quality rule │ ├─ The bill has parts │ ├─ Uncached input can have its own rate │ ├─ Cached input can use another rate │ └─ Generated output adds usage │ ├─ Price is not task cost │ ├─ Retries remain in total spend │ ├─ Failed attempts are not completed tasks │ └─ Divide spend by accepted tasks │ └─ Measure the workload ├─ Use representative real requests ├─ Keep model settings with results └─ Recheck when work or tariffs change
What happens during AI inference?
A trained model receives a request and generates an output. That use of an existing model is inference. The training that produced the model is an earlier activity; it is not repeated from scratch for every customer question. For a support team, the request might contain a customer's message and instructions about the desired reply. The output might be a suggested answer for a person to check. Each new request is another use of the model, even if the same question appeared yesterday.
This distinction matters when a firm prices a service. A provider may have substantial costs to make a model available, while the customer sees a price for using it. NIST defines inference serving cost as the datacentre and compute expense of making a model available to end users. It defines token price as the amount users pay per token, then separates both from the end-to-end expense of completing a task.1 The buyer cannot infer the provider's serving expense from a posted usage price.
Inference also has an outcome. A generated response that fails a firm's quality test can still consume billed usage. The unit for a purchasing decision should therefore be a completed piece of work, with a defined standard for completion. If a suggested reply needs correction, a second call, or a human rewrite, the first model response was an inference event but not yet a completed customer service task. That observation follows directly from measuring the task rather than the isolated call.
Input and output tokens in the bill
A token is a model's basic unit of input and output; it may correspond to a character, part of a word or a whole word.1 A request has material going in and material coming out. In a typical text service, the input includes the instruction and the content supplied for that request. The output is the generated response. A provider can price these classes differently. NIST's comparison of DeepSeek V4 Pro and a reference model publishes separate prices for uncached input, cached input and output tokens.3 A single headline price consequently cannot represent every workload.
Consider two support jobs with the same short customer message. One sends only that message and asks for a brief classification. The other sends the message, a long policy extract and previous correspondence, then asks for a draft reply. The second job has more input. If the draft is long, it also has more output. Even before quality is considered, the token mix differs. Record the mix instead of assuming that one request has one stable price.
| Usage part | What to record | Why it changes the bill |
|---|---|---|
| Uncached input | Tokens newly sent with the request | It may have its own price.3 |
| Cached input | Tokens billed under a provider's cache category | NIST's example uses a distinct rate.3 |
| Generated output | Tokens returned during inference | The output rate can differ from the input rate.3 |
| Repeat calls | Every call made before completion | More calls mean more billed usage, if billed. |
A service may expose a separate reasoning category, or include it in its own output accounting. NIST reports that more reasoning generally improves performance while consuming more computational resources, including time, money or tokens.2 Use the provider's actual usage record for that class. The firm's worksheet should preserve the categories shown on the bill, rather than merge everything under an invented average token count.
Inference cost versus training cost
NIST defines training cost as compute, labour and other inputs spent by an AI company to create a new model.1 Inference is the later use of that model to produce an answer. A supplier's training expense can inform its business, yet a customer paying a hosted model normally receives a usage schedule. That customer should record the expense of the calls that complete its own work.
A company operating its own model has a different accounting problem from a company buying access. It may need to allocate the equipment and operating expense of serving requests, rather than applying an external token tariff. The categories still help: keep the cost of making the model available distinct from the workload's success rate. A server running cheaply per token would offer little business value if its answers force frequent retries. This is a measurement principle, not a claim that one deployment method is always cheaper.
The practical question is whose cost is under review. If the firm buys an API, record the actual amount charged for its requests. If it runs a model itself, record the serving resources it consumes over the measurement period. In either case, use the same task definition and quality threshold when comparing alternatives. Otherwise a low numerical cost can reflect a weaker answer rather than a more efficient service.
Why a low token rate can cost more
Price per token is only one multiplier. NIST gives a simple example: a model charging half as much per token but using four times as many tokens costs twice as much to finish the work.1 In its later comparison, DeepSeek V4 Pro cost $1.74 per million uncached input tokens, $0.0145 per million cached input tokens and $3.48 per million output tokens. GPT-5.4 mini cost $0.75, $0.075 and $4.50 respectively. Yet DeepSeek V4 Pro was cheaper on five of seven CAISI benchmarks; across those seven, its expense ranged from 53% less to 41% more.3 Compare solved tasks, not the cheapest input column.
Suppose two systems face the same sample of customer queries. One produces concise replies that staff accept; the other produces longer replies and needs further prompts. Even if the second system has a lower rate in one input column, it may consume more paid output or more calls. Conversely, caching repeated material might materially change the input mix where the provider bills cached input differently.3 Neither direction can be assumed before recording the actual use.
The denominator matters as much as the numerator. Dividing spend by all attempts makes an unreliable system look cheap if many attempts fail. Divide by tasks completed to the chosen standard. Keep the rejected attempts in the numerator, because those attempts used resources. Record human review separately if comparing the full cost of a workflow; NIST's token definitions alone do not turn staff time into a token charge.1
Settings that change the result
Sampling settings affect how a model selects successive tokens. NIST identifies temperature, top_p and top_k as common settings in an evaluation protocol.2 Reasoning effort is another setting: more reasoning generally improves performance but costs more computational resources in time, money or tokens.2 A change can affect wording, length or the chance that a response passes a task-specific test. Measure the same tasks under each proposed configuration.
Output limits and the amount of context sent with a request also deserve attention. NIST's DeepSeek V4 evaluation recorded context length, max_tokens, temperature, top_p, internal reasoning, system prompt and maximum thinking among its chosen settings.3 Shorter output is valuable only when it still performs the task. Cutting a reply below the information a customer needs might save tokens in the call while increasing follow-up work.
NIST treats model settings as part of a reproducible evaluation protocol.2 For a firm, that suggests writing down the configuration beside each result. If the team changes a setting halfway through a comparison, it should mark the change. An apparent price improvement can otherwise be a configuration change, a different input mix or a shift in task difficulty.
Estimating cost per completed task
- Define a completed task, such as a support reply accepted without a further model call.
- Run the same representative requests through each candidate configuration.
- Record every call, its uncached input, cached input, output, any billed reasoning, its result and any retry.
- Apply each token-class rate to every call, including failures, and total the spend.
- Divide that total by tasks completed to the agreed standard, not by attempted requests.
For a dated illustration using NIST's May 2026 DeepSeek V4 Pro rates, suppose a sample used 10,000 uncached input tokens, 20,000 cached input tokens and 5,000 output tokens, including unsuccessful calls. The token charges would be $0.0174 + $0.00029 + $0.0174 = $0.03509. If eight tasks met the completion rule, the token spend per completed task would be $0.00438625. Those are arithmetic illustrations using published rates, not measured support-desk results.3
A cost estimate is only as durable as its assumptions. State the work sample, the period, the configuration, the price schedule and the quality rule alongside the result. If the firm later sends larger documents or demands longer replies, the old estimate should not be silently reused. The formula remains simple; the observed inputs change.
What this estimate cannot decide
Inference cost does not tell a firm whether a model's answer is safe, correct or suitable for its customer. A token ledger can show expense and throughput, while the quality check must decide whether the output fulfilled the task. NIST's evaluation practice says to report model performance together with the costs incurred to achieve it.2 A firm should keep both visible when making a purchase decision.
A measured sample also does not guarantee future bills. Prices, workload mix, cache behaviour and output length may change. The DeepSeek V4 comparison describes specific models and benchmarks, not a permanent price ranking for every support desk.3 Treat the result as a dated estimate and repeat the measurement when the workload or tariff changes.
What to do next
A representative batch of support requests and a written acceptance rule provide a starting point. The record should include billed input categories, output, settings and every retry for each request. Calculate spend per accepted task and review any failures before scaling the test. For related decisions, see how to assess AI agent opportunities, where AI agents fit in small firms, and the choice between agents and automation platforms. An accounts payable workflow shows why the task definition matters beyond a chat response.
Frequently Asked Questions
Is inference the same thing as training a model?
Does a cheaper input token always make an AI task cheaper?
What are reasoning tokens in an inference bill?
Can sampling settings change the cost of an answer?
What should a small firm measure before estimating AI cost?
Sources
Want this run on your business?
AI Foundation Audit — a structured assessment of your AI footprint: integration risks, governance gaps, ROI opportunities. Delivered as a comprehensive report you can act on.
You receive your AI Opportunity Report and Implementation Brief — tailored to your business and delivered immediately.