Inference
inference-time, serving, generation, decodingDefinition
Running a trained model to get an answer — as opposed to training, which produced the weights in the first place; the phase you pay for per request and wait for in real time
Training happens once and costs a fortune; inference happens every time anyone asks anything and is the bill that scales with use. Two numbers describe it from the outside: time to first token, which is how responsive it feels, and tokens per second, which is how fast it finishes. A long prompt makes the first worse; a long answer makes the second matter.
The part where the model is actually used. Everything in a session that takes time and costs money is inference.
Avoid: expecting it to be repeatable. The same prompt can yield a different answer, which is why evaluation here means a dataset and a scorer rather than a snapshot test — and why anything that must be identical every time belongs in code, not in a prompt.