Serving — процесс making a trained model available to generate predictions or inferences for users, applications or downstream systems. Он может происходить через online inference, batch inference или offline cached predictions.
The process of making a trained model available to provide predictions through online inference or offline inference.