Training builds the model. Inference uses it. Once weights are fixed, the network takes new input and produces output, whether that is a label, a probability, or a generated sentence. No gradient updates happen here.
Inference is where the economics of AI show up. Training happens once, sometimes over weeks. Inference happens millions of times, often under latency budgets measured in milliseconds. A model too slow to respond in a chat window or on a phone is useless, no matter how accurate.
Common inference settings
- Batch inference for offline processing
- Real-time inference for interactive applications
- Edge inference on local devices
- Server inference in data centres
Optimization focuses on throughput and latency. Quantization, pruning, and specialized hardware all target the inference path. The same model can behave very differently depending on how it is served.
Comments
No comments yet. Be the first to share a thought.
Leave a comment