Evidence · Measurement · Limits

What we measure, and what we cannot yet prove.

A synthetic respondent should not be judged on a demo. This page brings together our use cases, published end to end, then what we have measured, against what, and what remains to be proven. It is updated with every new study and every new measurement.

Use cases

Real studies, published end to end: the question, the setup, what comes out, and what it lets you say. Enough to see how it works before you try it. Each one shows:

  • the decision at stake and the question asked;
  • who is interviewed, how, and how long fieldwork took;
  • what comes out: themes, disagreements, verbatims, and one full interview;
  • what the study lets you say, what it does not, and how to ask for the full report.

What this evidence lets you decide

You can use it to

  • compare and rank concepts, messages or names before fieldwork, then take only the strongest ideas to fieldwork;
  • spot objections and points of disagreement in a target early;
  • ask any supplier, ourselves included, for the same measurements against a public survey.

Not to

  • replace measurement with real people when a decision carries real weight;
  • settle a minority view or a genuinely new category;
  • assume answers are right on new questions: that test has not been run yet.

What we measure

The most common defect of synthetic respondents does not show in the averages: they agree with each other too much. So we measure the spread of answers against real people of the same profile. At 1.00, the disagreement is the same as among humans.

Across eight markets, our simulated respondents showed 0.20 to 0.76 of human spread before correction. After correction: 0.96 to 1.07 on the attitudinal layer, 0.85 to 1.00 on the personality layer.

Measured against what

Always against surveys of real people, never against a model of them. Round 11 of the European Social Survey for France, the United Kingdom, Germany, Italy and Spain. The World Values Survey for the United States and Mexico. A national survey of 1,016 respondents run by Insights House for Morocco. Twin-2K-500, 2,058 American respondents, for the personality layer.

Each correction was checked on five replications of ten thousand generated profiles. We publish what we measure against and what we measure; we do not publish how the correction works.

Open the animation full screen

What we cannot yet prove

It is a calibration, not a validation: the out-of-sample bench, with new questions, is pre-registered and has not been run.

The reference carries its own error: roughly ±0.02 on the American market.

One market has no independent control: Mexico rests on a single survey.

Our reliability index does not yet score dispersion, so we publish no confidence intervals derived from our variance.

Open the animation full screen

What you can check yourself

A measurement nobody can repeat is worth little. We publish the method of the test, and the questions a buyer can ask any supplier, ourselves included.

Our public commitments

We answer Esomar's 20 Questions for buyers of AI-based services in public, including the ones we cannot yet answer. The ICC/Esomar International Code, revised in 2025, requires telling clients when synthetic data or synthetic personas are used: we apply it. And we accept blind parallel tests, published whatever the outcome.

To propose a parallel test, discuss a measurement, or argue with us.

Write to FlashInsight