Evidence · Measurement · Limits
What we measure, and what we cannot yet prove.
A synthetic respondent should not be judged on a demo. This page brings together our use cases, published end to end, then what we have measured, against what, and what remains to be proven. It is updated with every new study and every new measurement.
Use cases
Real studies, published end to end: the question, the setup, what comes out, and what it lets you say. Enough to see how it works before you try it. Each one shows:
- the decision at stake and the question asked;
- who is interviewed, how, and how long fieldwork took;
- what comes out: themes, disagreements, verbatims, and one full interview;
- what the study lets you say, what it does not, and how to ask for the full report.
- Cas d’usage : vingt entretiens individuels avec des personas synthétiques sur leur propre place dans les études
Une étude qualitative exploratoire, publiée de bout en bout : la question posée, le panel de vingt profils, le dispositif, les thèmes, les désaccords, les verbatims et un entretien lu en entier. Puis ce que l’étude permet de dire, et ce qu’elle ne permet pas.
What this evidence lets you decide
You can use it to
- compare and rank concepts, messages or names before fieldwork, then take only the strongest ideas to fieldwork;
- spot objections and points of disagreement in a target early;
- ask any supplier, ourselves included, for the same measurements against a public survey.
Not to
- replace measurement with real people when a decision carries real weight;
- settle a minority view or a genuinely new category;
- assume answers are right on new questions: that test has not been run yet.
What we measure
The most common defect of synthetic respondents does not show in the averages: they agree with each other too much. So we measure the spread of answers against real people of the same profile. At 1.00, the disagreement is the same as among humans.
Across eight markets, our simulated respondents showed 0.20 to 0.76 of human spread before correction. After correction: 0.96 to 1.07 on the attitudinal layer, 0.85 to 1.00 on the personality layer.
- How reliable are synthetic respondents? Our own numbers, eight markets — The full note, market by market.
- Under-dispersion — The defect, defined and measurable.
- Under-dispersion, and what we did about it — The defect, published before it was fixed.
Measured against what
Always against surveys of real people, never against a model of them. Round 11 of the European Social Survey for France, the United Kingdom, Germany, Italy and Spain. The World Values Survey for the United States and Mexico. A national survey of 1,016 respondents run by Insights House for Morocco. Twin-2K-500, 2,058 American respondents, for the personality layer.
Each correction was checked on five replications of ten thousand generated profiles. We publish what we measure against and what we measure; we do not publish how the correction works.
What we cannot yet prove
It is a calibration, not a validation: the out-of-sample bench, with new questions, is pre-registered and has not been run.
The reference carries its own error: roughly ±0.02 on the American market.
One market has no independent control: Mexico rests on a single survey.
Our reliability index does not yet score dispersion, so we publish no confidence intervals derived from our variance.
What you can check yourself
A measurement nobody can repeat is worth little. We publish the method of the test, and the questions a buyer can ask any supplier, ourselves included.
- Under-dispersion, and a test anyone can run — The test, on public data, at no cost.
- Calibration, and what a buyer can check — What a supplier should be able to show.
- How synthetic research works — The eight steps, from the decision to the reading, and the checks.
- Synthetic respondent — Definition, reliability, and what to check before buying.
Our public commitments
We answer Esomar's 20 Questions for buyers of AI-based services in public, including the ones we cannot yet answer. The ICC/Esomar International Code, revised in 2025, requires telling clients when synthetic data or synthetic personas are used: we apply it. And we accept blind parallel tests, published whatever the outcome.
To propose a parallel test, discuss a measurement, or argue with us.
Write to FlashInsight