Glossary

Under-dispersion

Also called understated variance · collapsed spread

The commonest defect in simulated panels: answers cluster more tightly than real humans’. Disagreement disappears, and the signal goes with it.

Ask a hundred humans the same question and the answers spread out. Ask a hundred poorly built synthetic respondents and they bunch. That is under-dispersion, and it is the best-documented failure mode in the field.

Open the animation full screen

It does not show up in a mean — an under-dispersed mean can be perfectly accurate. It shows up in the standard deviation, in the shape of the distribution, in how much disagreement survives. The consequences are immediate all the same: two options the field separates cleanly come back artificially close, and two genuinely equivalent options come back artificially apart.

The test, which anyone can run

Ask the vendor for the full distribution of answers to one opinion question, not the mean. Compare its standard deviation to a public survey asking something close. If the synthetic spread is clearly narrower, the panel is under-dispersed — and the size of the gap tells you by how much.

What a buyer can demand

That dispersion be measured market by market rather than asserted in general. That it be compared to a public, citable reference. And that the number be given even when it is unflattering: a vendor who publishes no gap does not have a better panel, only an unmeasured one.

Frequently asked questions

What is under-dispersion in synthetic research?

It is when a simulated panel’s answers cluster more tightly than those of real respondents with the same profile: disagreement that exists in the population has not been reproduced. It is read as a ratio, the spread of simulated answers divided by the spread of human answers to the same question. At 1.00 the spread is human; below it, disagreement is missing.

Does under-dispersion change the averages?

No, which is why it goes unnoticed for so long. An under-dispersed panel returns the right mean and the wrong spread, so the summary table looks fine. When we corrected ours, the averages did not move. What changes is the ability to separate two options that the field separates.

What dispersion ratio is acceptable?

As close to 1.00 as possible, measured against a public survey rather than asserted in general. Our corrected markets sit between 0.96 and 1.07 on attitudes. A ratio well below 1.00 blurs comparisons between options; one well above it adds noise. The figure should be given market by market, before and after correction.

How can a buyer test for under-dispersion?

Ask the vendor for the full distribution of answers to one opinion question, not the mean. Take a public survey asking the same question of the same population, such as the European Social Survey or the World Values Survey, and compare the two standard deviations. If the synthetic spread is clearly narrower, the panel is under-dispersed, and the gap tells you by how much.

See also

  • Calibration — The operation that ties a generated population to observed real-world data — and the measurement of the gap that remains once it is done.
  • Synthetic respondent — An instance of a model asked to answer questions as a human would, calibrated on data describing a real population. The output — the unit a sample is made of.
  • Synthetic survey — A conventional quantitative protocol — closed questions, segments, weighting — run against a population of synthetic respondents instead of a human panel.

Further reading

Back to the glossary