Synthetic data is artificially generated data designed to resemble selected patterns, structures, or statistical properties of real data. It can support testing, training, privacy-preserving development, and simulation, but may still carry privacy, bias, or fidelity risks depending on how it is generated.
This is 'artificial' data generated by data synthesis algorithms. It replicates patterns and the statistical properties of real data (which may be personal data). It is generated from real data using a model trained to reproduce its characteristics and structure.
A side-by-side comparison of Synthetic Data and Personal Data. Understand how artificially generated data differs from information relating to an identified or identifiable person.
A side-by-side comparison of Data Augmentation and Synthetic Data. Understand how modifying or transforming existing data differs from generating artificial data that resembles real data.