Batch inference is the process of generating predictions for multiple unlabeled examples grouped into batches. It is commonly used when real-time responses are not required and can improve throughput by using parallel hardware or scheduled processing.
The process of inferring predictions on multiple unlabeled examples divided into smaller subsets ("batches"). Batch inference can take advantage of the parallelization features of accelerator chips. That is, multiple accelerators can simultaneously infer predictions on different batches of unlabeled examples, dramatically increasing the number of inferences per second. See Production ML systems: Static versus dynamic inference in Machine Learning Crash Course for more information.