AI Glossary · Definition
What is Inference?
Inference is the act of running a trained AI model to produce an output from a new input, such as generating a reply to a prompt.
By DAIDU EditorialUpdated
Inference, explained
Training builds the model; inference is using it. Every time you ask an assistant a question or an automation sends text to an AI API, that is inference. Inference has a cost, usually per token or per request, and a speed, often called latency. For high-volume business workflows, inference cost and speed can matter as much as quality, which is why teams match model size to the task.
Example
An e-commerce store estimates monthly inference costs before automating product descriptions for 10,000 items.
