Serving is the process of making a trained model available to generate predictions or inferences for users, applications, or downstream systems. It can occur through online inference, batch inference, or offline cached predictions.
The process of making a trained model available to provide predictions through online inference or offline inference.