
Synthetic Respondents: What a Vendor's Claim Actually Tells You
Three providers say “synthetic respondents” and mean three different products. Here is how to tell which one you are being sold — and why the most common accuracy claim answers the wrong question.
If you have sat through three vendor demos this year, you have heard the same two words describe three different products. That is not marketing sloppiness. “Synthetic respondents” genuinely covers several distinct constructions, and the gap between them is the gap between filling in a missing answer and inventing a person. This piece is a decoder, not a verdict.
Four things the words can mean
The clearest public taxonomy in the industry is not from a vendor. In June 2025 the research practice of Syntec Conseil, the French professional body, published a framing note that sorts the field into four cases. They are listed here in order of increasing distance from a real human.
- Imputation. A real person answered most of the survey and skipped some items; a model estimates what they would have said. The respondent exists. Statisticians have done this for decades.
- Extrapolation. The questionnaire is shortened and answers to unasked questions are predicted. The respondent is real; part of their record is not.
- Sample augmentation. Virtual respondents are generated from the real respondents of the same study, typically to reach a thin cell or a rare audience.
- Fully simulated respondents. No real person in the sample at all. This is what most people picture, and what most of the commercial excitement is about.
A provider can be excellent at case 1 and have nothing at all for case 4. Both get described as “synthetic respondents” in a deck.
Ask which of the four you are buying before you ask how accurate it is. The word “accurate” means something different in each one.
And three ways to build the fourth
Within fully simulated respondents there are again three constructions, and they are not equivalent: generic generation from a general-purpose model with a prompt; statistical extension of a smaller real sample; and grounded generation, where personas are built from real population data and the outputs are tuned against real human answers. We set this out at length in our reference guide to synthetic panels. Only the third has a defensible answer to “compared to whom?”
What the public claims look like
The major platforms have now said what they are doing, in public, which is progress worth acknowledging. At X4 in March 2026, Qualtrics announced synthetic panels for US consumers, built on a purpose-built model rather than an off-the-shelf one, with UK and Ireland, Canada, and Australia and New Zealand to follow. NIQ has published an educational note on the rise of synthetic respondents that frames them as bounded augmentation rather than substitution, and Ipsos has run public sessions on research with synthetic data. Building a purpose-built model instead of prompting a general one is a real engineering decision, and the right one.
The Qualtrics announcement also carries a number, and it is worth quoting exactly rather than paraphrasing:
delivering 12x better accuracy in matching human responses than general-purpose AI
Read that carefully. It is almost certainly true, and it is the wrong comparison for a buyer — not because the vendor is being evasive, but because the industry has settled on a baseline that flatters everyone who uses it. Twelve times better than an uncalibrated general-purpose model is a statement about how bad the uncalibrated baseline is.
The baseline that would actually tell you something
There is a ceiling to this problem, and it is not one hundred percent. Humans are not perfectly predictable by other humans. A pre-registered audit of thirty-seven language models put numbers on all three points at once: pairs of real people predicting each other reach 0.825; the best model reaches 0.714; and a Gaussian copula, a statistical method with no artificial intelligence in it whatsoever, reaches 0.688.
Those three numbers reframe every accuracy claim in the category. The distance between the best model and a decades-old statistical formula is thin. The distance between either and the human ceiling is not. A vendor who tells you they beat a naive prompt has told you they cleared the lowest bar in the room; a vendor who tells you where they sit relative to real humans on your population has told you something you can use.
One caution if you go to that paper yourself: its abstract and its body do not agree about whether the copula beats every model. The body, at §5.2, is the one to read.
Three questions that decode any provider
- Which of the four am I buying? Imputation on real respondents and fully simulated respondents are different products with different risks. If the answer moves during the conversation, that is the answer.
- Compared to whom, on which population, measured when? Not “versus general-purpose AI”. Versus real people, in the country and segment you are buying, with a date. A US calibration does not transfer to a French study.
- Can I have the answers at respondent level? Without them you cannot check the one thing that catches the most problems — whether people who share a profile disagree with each other as much as real people do. The procedure is in a test anyone can run, it costs nothing, and it works on any vendor including us.
The same standard, turned around
It would be cheap to publish this without applying it to ourselves, so: we measured our own engine against 1,016 real respondents and found that inside a group our simulated respondents disagreed about a third as much as real people did. We published that before we had a remedy, then published the correction and the residual figures in Under-dispersion, and what we did about it. The second defect in that family — simulated averages that sit in the wrong place — we have not solved, and we say so there.
That is the standard this category will be judged by eventually, and the reason ESOMAR wrote twenty questions rather than a certification: no demo can distinguish a provider who measures from a provider who asserts. Our own twenty answers, including the ones that were unflattering, are published in full.
The vocabulary. Each term on its own page, with what a buyer can check: synthetic respondent · synthetic persona · calibration · under-dispersion · the full glossary
Method over Magic.
Sources
Taxonomy. Philippe Guilbert, “Données synthétiques et études marketing & sondages d’opinion”, Syntec Conseil, June 2025 — the four use cases.
Vendor statements. Qualtrics, X4 2026 announcement (18 March 2026) · NIQ, “The rise of synthetic respondents” · Ipsos, research with synthetic data. Quoted as published; no private information is involved.
Audit. arXiv 2608.14606, thirty-seven models. Figures taken from the body, §5.2: human ceiling 0.825, best model 0.714, Gaussian copula 0.688.
On FlashInsight. 85% of what? · Calibration, and what a buyer can check · Under-dispersion, and a test anyone can run.