Glossary

Synthetic data

Also called artificial data · generated data

Artificially generated data that mimics the statistical properties of real data without containing anything personally identifiable. The raw material — not the respondent.

Synthetic data is produced by a model rather than collected from a person. It is built to carry the statistical regularities of a real dataset — distributions, correlations, the shape of subgroups — without reproducing any identifiable individual record.

That is a status, not a use. The same material serves three quite different jobs: filling gaps in an incomplete dataset, standing in for a sensitive one so it never has to leave the building, and feeding the construction of a population you can then interview. Only the third has anything to do with research.

What it is not

It is not a respondent. It is the input, not the output. Synthetic data answers nothing — it describes. Sliding from one to the other is the most common vocabulary error in this market, and it is not harmless: it implies that a property verified on the raw material carries over as a guarantee on the answers produced from it. It does not.

What a buyer can check

Three questions are enough. Which real data does the generation rest on, and is that source public or proprietary? Which statistical properties were preserved, and which were lost? And the one that separates vendors: were those properties measured after generation, or merely aimed at? A supplier who describes a method without showing a measured gap has described an intention. Our method page sets out what we measure.

See also

  • Synthetic respondent — An instance of a model asked to answer questions as a human would, calibrated on data describing a real population. The output — the unit a sample is made of.
  • Synthetic population — The generated set taken as a whole, built to match the known statistical margins of a territory. The term comes from microsimulation, not from research.
  • Calibration — The operation that ties a generated population to observed real-world data — and the measurement of the gap that remains once it is done.

Further reading

Back to the glossary