A golden dataset is a carefully curated dataset treated as ground truth for testing, benchmarking, or evaluating a model. It is often manually reviewed and used to detect regressions or compare model quality over time.
A set of manually curated data that captures ground truth. Teams can use one or more golden datasets to evaluate a model's quality. Some golden datasets capture different subdomains of ground truth. For example, a golden dataset for image classification might capture lighting conditions and image resolution.