Also known as: Validation Data Set Β· Validation dataset Β· validation set
Validation data is data used to evaluate a model during development and tune choices such as hyperparameters, thresholds, or learning processes. It should be distinct from training and test data to help detect underfitting, overfitting, and generalization problems.
Data used for providing an evaluation of the trained AI system and for tuning its non-learnable parameters and its learning process in order, inter alia, to prevent underfitting or overfitting.
βdata used to compare the performance of different candidate modelsβ ISO/IEC 22989.
The subset of the dataset that performs initial evaluation against a trained model. Typically, you evaluate the trained model against the validation set several times before evaluating the model against the test set. Traditionally, you divide the examples in the dataset into the following three distinct subsets: - a training set - a validation set - a test set Ideally, each example in the dataset should belong to only one of the preceding subsets. For example, a single example shouldn't belong to both the training set and the validation set. See Datasets: Dividing the original dataset in Machine Learning Crash Course for more information.
A side-by-side comparison of Training Data and Validation Data. Understand how data used to fit the model differs from data used during development to tune choices and detect generalization problems.
A side-by-side comparison of Validation Data and Test Set. Understand why development tuning data must remain separate from independent evaluation data.