Glossary
A Glossary of Synthetic Research
The words this market uses interchangeably, taken one at a time — and, for each, what a buyer can actually check.
“Data,” “respondent,” “persona” and “population” are not the same object. Blurring them is not a pedantic quarrel: it decides what you are entitled to expect from a result. Each entry fits on one page — what the term covers, what it does not, and how to hold a vendor to it.
The objects
- Synthetic data
- Artificially generated data that mimics the statistical properties of real data without containing anything personally identifiable. The raw material — not the respondent.
- Synthetic respondent
- An instance of a model asked to answer questions as a human would, calibrated on data describing a real population. The output — the unit a sample is made of.
- Synthetic persona
- The briefing given to the model: the traits, attitudes and characteristics that steer its answers. A specification, not a person.
- Synthetic population
- The generated set taken as a whole, built to match the known statistical margins of a territory. The term comes from microsimulation, not from research.
- Digital twin
- An attempt to replicate a specific, real individual rather than represent a segment. The strongest claim in the field — and therefore the most demanding to evidence.
The instruments
- Synthetic survey
- A conventional quantitative protocol — closed questions, segments, weighting — run against a population of synthetic respondents instead of a human panel.
- AI focus group
- Several synthetic respondents brought together with a moderator, reacting to a stimulus and to each other. The depth route, on a small group.
- Agent-based research
- A design in which respondents are not merely questioned but set in motion over time, so that one reaction can produce the next.
- Synthetic research
- Market research in which the answers come from simulated respondents, built from a real population, instead of being collected from a human panel.
The controls
- Calibration
- The operation that ties a generated population to observed real-world data — and the measurement of the gap that remains once it is done.
- Under-dispersion
- The commonest defect in simulated panels: answers cluster more tightly than real humans’. Disagreement disappears, and the signal goes with it.
The uses
- Ad pre-test
- Evaluating a creative before it runs — likeability, emotion, comprehension, intent, brand attribution. The study format where synthetic earns its keep first.
- Concept test
- Putting an idea that still lives on paper — a concept, a product, a name, a promise — in front of its audience, before anything is committed.
- Copy testing
- Showing draft advertising copy to a sample of the target audience before it runs, to measure whether it is noticed, understood, believed and attributed to the brand.
- Concept optimisation
- Improving a concept after it has been tested: finding which element holds it back, rewriting that element, and retesting the new version against the original until the gain stops.
- Name testing
- Showing candidate brand or product names to a sample of the target audience to measure fit, ease, recall, associations and appeal, before one is chosen.