Version 1 — 31 August 2026

ESOMAR's 20 Questions, answered.

ESOMAR's 2026 CEO Playbook found that 8% of research leaders believe synthetic data is reliable enough for business decisions, and 36% believe it never will be. Two-thirds do not use it at all. We think that scepticism is the correct starting position, and we do not intend to argue against it. We intend to answer it — question by question, with numbers, and with our limitations stated rather than omitted.

The questions below are ESOMAR’s, not ours. They are published in ESOMAR — 20 Questions to Help Buyers of AI-Based Services. We answer all twenty, in their order and in their wording.

A · Company profile

1. What experience and know-how does your company have in providing AI-based solutions for research?

FlashInsight is a synthetic respondent platform for market research, founded and led by Dr Laurent Florès, who has spent his career in market research and holds a doctorate in management sciences. The platform runs quantitative studies and qualitative interviews on simulated respondent panels built from official population statistics, across France, Italy, Spain, Germany, the United Kingdom, the United States, Mexico and Morocco.

Alongside the product we run a continuing research programme on the validity of the method itself. Its results — including the unfavourable ones — inform the answers below.

2. Where can AI-based services have a positive impact for research? What benefits does AI bring, and what problems does it address?

Our answer is narrow, and it matches what the industry reports doing. In ESOMAR's own figures the leading uses of synthetic data are testing and simulation (21%) and digital twins or personas (17%); synthetic panels come last at 5%.

The problem worth solving is not the cost of fieldwork. It is that most ideas are eliminated before any study takes place. A team with twenty concepts and budget for three eliminates seventeen by opinion. Simulated respondents let that elimination happen against a structured, population-anchored reading instead — before money is committed to fieldwork.

That is a real gain, and it is bounded. It improves the screening step. It does not replace the decision step.

3. What practical problems and issues have you encountered in the use and deployment of AI? What has worked well and how, and what has worked less well and why?

Three findings from our research programme that changed how we build. Each began as an assumption we held, and each was settled by measurement rather than by argument.

We assumed richer profiles produce better respondents. They do not. We measured what happens when real answers from a real person are added to a simulated profile. Accuracy plateaus almost immediately and then declines. More strikingly, on the dimension we measured — the strength of the relationship between political position and expressed attitudes — the simulated respondents were most caricatured when their profile contained only demographic attributes, and adding real answers reduced that caricature rather than building fidelity. We report this on the dimension tested; we have not shown it holds for stereotyping in general.

We assumed the values and personality layer improved everything. It does not. In a controlled comparison that layer clearly improves structure — how groups differ from one another, and in what direction. It degrades dispersion — how much individuals differ within a group. We state the trade-off rather than omit it, and closing it is the main item on our research agenda.

We assumed a validation study proved a product. It does not. A method can be validated in a study while a specific sample silently fails to differentiate. That is why we built a reliability index that runs on every sample in production, rather than relying on a study run once.

We report these because they are the reason the design is what it is. A supplier who has never found a result that contradicted their own assumption has either not looked or is not saying.

B · Is the AI capability explainable and fit for purpose?

4. Can you explain the role of AI in your service offer in simple, non-technical terms in a way that can be easily understood by researchers and stakeholders? What are the key functionalities?

We build a population, then we ask it questions.

Building the population comes first, and it is where the work is. Each simulated respondent is defined on two foundations: a socio-demographic profile drawn from the official statistics of the country concerned — the real proportions of age, gender, region, education and occupation — and a values and personality profile using variables validated by published research, distributed according to their measured frequencies in that same population.

Only then does the language model intervene. It receives that profile and a context, and produces how such a person would respond. It does not decide who exists in the sample, or in what proportion. The population comes from the statistics. The language comes from the model.

Results are read at the level of groups and of the total. They are never read as a prediction about an individual.

The key functionalities are five: building a population for a given country and target; running a quantitative study on it; running qualitative interviews with individual simulated respondents; scoring each sample for reliability before results are used; and producing the analysis, corrected for level and read at group and total level. A client can also supply their own data to calibrate a population — see question 5.

5. What is the AI model used? Are your company's AI solutions primarily developed internally, or do they integrate an existing AI system and/or involve a third party — and if so, which?

We use general-purpose commercial language models — open-weight or proprietary — accessed through an aggregation layer, and we are deliberately model-agnostic. We have changed the underlying model several times without changing our method.

That is the substance of our answer: the method takes precedence over the model. The correction does not live in the model's weights. It lives in the inputs — real population margins — and in the output layer, where levels are corrected and each sample is scored for reliability. In practice this means a model change does not invalidate the method, and we re-run our validity checks after one. We should be precise about the status of that: it is a property of the design, and we have not yet published a formal measurement of how stable our results are across models. Doing so is on our agenda.

Four consequences follow, and we consider them decisive for a research supplier. Portability: a better model can be adopted without redoing the method. Currency: when the world changes we update the statistics rather than retrain a model. Auditability: we can show you our margins, and you can reconstruct them from public sources — what a model learned during training is not inspectable in the same way, whether its weights are open or not. Transferability across countries: a fine-tuned model carries the distribution it was trained on, ours carries the distribution of the country being studied.

This architecture also changes what we can do with a client's own data, and the difference is worth stating precisely. A population can be calibrated on data the client supplies — their own segment structure, their own measured distributions, their own category knowledge — so that the simulated population matches the client's market rather than the national average.

Because our correction lives in the inputs, that data is not used to train or tune any model — not for this client, and not for any other. It shapes how the population is drawn. Every call we make carries a setting that explicitly refuses data collection, and that setting is covered by our automated tests; whether a provider honours it is governed by their contractual terms, which we check but do not enforce ourselves. Two conditions apply and both are absolute: the data supplied must contain no information identifying a person, and it remains the client's property throughout.

We also test domain fine-tuned models, for two reasons: to gain independence from commercial providers, and to verify this position rather than assert it. Published work is consistent that fine-tuning on real individual responses improves distributional alignment and reduces demographic bias.

Our own test does not support adopting one. We re-analyse a fine-tuned arm another team has released, at no training cost to us. Against a decision rule written down before we look at the results, the advantage is not large enough to outweigh the four properties above: it fails on two of the three comparators we declared. It would be dishonest to present that as a clean loss — on the accuracy of attitudinal means that arm ranks first of eleven, and on total dispersion sixth of eleven, where five prompting configurations beat it. It does not close the dispersion gap: it remains roughly a third below the human level. Two limits belong with the result. The arm is someone else's; we have not trained a model. And the dataset is entirely American, so the cost of carrying a trained model across countries remains unmeasured. The method, the figures and the limits are published, and the data they use is public.

Sources
Under-dispersion, and a test anyone can run · Twin-2K-500 (Toubia et al., CC BY 4.0)

6. How do the algorithms deliver the intended result? Can you summarise the underlying data and how it interacts with the model?

The underlying data is official public statistics. In practice: INSEE and its equivalents among national statistical institutes, Eurostat, the European Social Survey — an academic probability survey conducted face to face — and the US Census Bureau. These provide the socio-demographic margins of a population and the measured distribution of values and personality variables within it.

The sample is drawn so that each of those distributions is respected. Each drawn profile then becomes the context conditioning a generated response. Responses are aggregated back to segments and to the total, weighted by the real population structure rather than by anything the model produced. The statistics do the weighting. The model never weights itself.

One limitation belongs here rather than in a footnote. We currently draw on each distribution separately, which means the correlation between socio-demographic position and values is not fully preserved. Real individual-level records exist in the academic survey data we use, and moving to a draw that preserves that joint structure is active work. Until it ships, our populations match the country on each dimension taken separately, and approximate it on the relationship between them.

Sources
INSEE · Eurostat · European Social Survey · US Census Bureau

C · Is the AI capability trustworthy, ethical and transparent?

7. What processes verify and validate output accuracy, and are they documented?

Two layers.

A validity base measured against external references. Three proofs, each answering a different objection, and each using a different method.

All three results below are our own measurements on our own protocol. They are not figures taken from published papers, and they are not directly comparable to figures those papers report on different tasks.

Sources
Twin-2K-500 (Toubia et al., 2025) · European Social Survey

ProofQuestion it answersResult
Test–retest reliabilityIs it stable?Our simulated respondents reach 26% of the human self-consistency ceiling — the level at which real people reproduce their own answers two weeks apart.
External validityDoes it match reality?ρ ≈ 0.83 against Twin-2K-500, a published US benchmark of real individuals. ρ ≈ 0.97, mean absolute error 0.35, against the European Social Survey, France.
Theory consistencyIs it coherent on its own terms?18 of 18 directional predictions matched established theory, using no external reference. This is nomological consistency — whether the simulated population behaves as theory says a real one should — not internal consistency in the psychometric sense.

8. What are the limitations of your models, and how do you mitigate them?

Four, stated plainly.

We do not predict individuals. We reproduce a population structure. This is a boundary of the method, not a roadmap item.

Effect sizes are overstated. This is the field's oldest documented defect — identified in 2023 replication work and confirmed again in 2026 in a study of seventy pre-registered experiments, where simulated effects correlated with real effects at 0.85 while systematically overestimating their magnitude. What we do about it: a deterministic correction layer is applied to output levels. We should be precise about its status — it is an architectural response to a known defect, and we have not yet published a measurement showing by how much it reduces error against a human reference. What follows for use: our results are read as rankings and directions, not as absolute levels.

Dispersion is compressed. Simulated respondents within a group resemble one another more than real people do. We measure it and we state it here; it is not yet scored by our reliability index, and adding a dispersion component is in progress. Until it ships, any confidence interval derived from our variance would be too narrow, and we publish none.

Reliability degrades away from the ground truth. The further a question sits from the topics covered by the source statistics, and the more recent the phenomenon, the weaker the reading. Our population data refreshes on the publication cycle of national statistical institutes.

Sources
Ashokkumar et al., Nature 656, 115–122 (2026)

9. What considerations have you taken into account to design your service with a duty of care to humans in mind?

There are two groups to consider, and the second matters more than the first.

Research participants. We recruit none and we collect no responses from real people. There is no identifiable individual behind any simulated profile. The harms that research ethics exists to prevent — burden, deception, breach of confidentiality, re-identification — have no subject in our production work.

One clarification belongs here, because the two statements sit close together in this document. Our published validity results are measured against datasets of real respondents — the US benchmark and the European Social Survey. Those responses were collected by other organisations, under their own ethical approvals, and released in de-identified form. We compare our populations to that data; we do not collect it.

People affected by the decisions our results inform. This group exists, and a synthetic method carries a specific risk toward it: published research shows that language models misrepresent and flatten demographic groups, with the harm concentrated on minority populations. A method that reproduced that defect uncritically would produce conclusions that under-serve exactly the groups least represented in the model.

Three things follow, and we are candid that our coverage is partial. Our populations are anchored on official statistics rather than on the model's own priors. We should be precise about the status of that choice: the published literature documents the defect and shows that training on real individual responses reduces it. It does not establish that anchoring on aggregate population statistics solves it. That anchoring is our reasoned response to a documented problem, and our own validity results are the evidence we offer for it — not a result imported from someone else's paper. Our reliability gate refuses to display a score when a sample's axes fail to differentiate — one of the two observable forms of flattening. It does not yet detect the other form, compression of variance within a group, which is the limitation stated under question 8 and the reason a dispersion component is being added. And we are progressively formalising review procedures for outputs that touch sensitive segments, as the research on this matures into settled practice. We would rather describe that as under construction than claim a completed framework.

Sources
Wang et al., Nature Machine Intelligence (2025) — on identity flattening

D · How do you provide human oversight?

10. How do you ensure it is clear when AI technologies are being used?

Every report and teaser we deliver states that its content comes from simulated respondents and is generated by an artificial intelligence system. That notice is not written by the model: it is inserted deterministically by the single document shell through which every generated document passes, in the language of the document. We never present a simulated response as a human one, and we do not use language implying direct human measurement. Our published AI usage framework states explicitly that outputs are AI-generated, that they are not factual truths, and that they may contain errors or biases.

We are precise about the scope of that statement: it covers the reports and teasers we deliver, not every raw data export a client can generate for themselves from the platform. Extending the notice to those exported formats is in progress.

We make no claim regarding the machine-readable marking of generated content contemplated by Article 50(2) of the EU AI Act. It is not in place, and we do not present it as such.

For organisations to which the ICC/ESOMAR Code applies — its members, and those that have adopted it — transparency about the use of AI and synthetic data is part of the Code's requirements rather than a matter of preference. We hold ourselves to it.

Sources
ICC/ESOMAR International Code · EU AI Act, Article 50

11. Do you have ethical principles explicitly defined for your AI-driven solution, and how in practice does that help to determine the AI's behaviour?

Our published AI usage framework sets out what the service delivers, what it does not guarantee, and where responsibility for decisions sits. It now also carries five commitments specific to simulated responses, published at flashinsight.io/en/ai-usage-framework.

First: we will never present a simulated response as a human one — every deliverable states that it comes from simulated respondents. Second: we do not replace direct measurement where direct measurement is the appropriate instrument. Third: we publish our validation results, including the limitations we encounter. Fourth: we state what was simulated, and from what — which population, which statistical sources, which reference date, how many respondents. Fifth: we state what the method should not be used for, in the deliverable itself.

These are modelled on the governance standard Gallup published in May 2026. We adopt them because they are the right floor, and we say where they come from.

We distinguish deliberately between the two kinds of statement above. The first commitment is enforced in code — the notice is inserted by the document shell, not requested from a model. The others are commitments in the ordinary sense: they bind us, and they are not machine-enforced. We would rather mark the difference than let a reader assume the stronger reading throughout.

Sources
Gallup — Research on simulated responses, 11 May 2026

12. How does your AI solution integrate human oversight to ensure ethical compliance?

A researcher designs the study, reads the output and writes the analysis. The model produces responses; it does not produce conclusions. Our operating rule is that a study we deliver as an analysis is read and interpreted by an experienced researcher before it goes out, and today that review is direct and hands-on rather than delegated. We are precise about its scope: it covers the analyses we deliver, not every raw export a client can generate for themselves from the platform. We state it as a working practice rather than a certified control, for the reason given below.

We regard this as the substance of the answer rather than a formality. A synthetic method changes where the researcher's time goes; it does not remove the need for their judgement. Reading a result against category knowledge, recognising when an output is plausible but wrong, knowing which finding deserves a real study — these are functions of experience, and no measure of model quality substitutes for them.

Two limits of that answer, and what we are doing about each.

Human review is real but not yet traced. A reviewer who reads everything is a strong control and a weak audit trail: today we can tell you that review happened, not show you what was examined, what was changed, and on what grounds. We are building a research harness that records exactly that — which outputs were reviewed, by whom, what was flagged, and what was altered before delivery — so that oversight becomes demonstrable rather than asserted. We would rather describe it as under construction than present it as in place.

Review is currently manual throughout. Other suppliers in this category operate automated monitoring layers that flag outputs requiring human attention. We do not have one, and we say so. Our current position is that a smaller volume of work read completely by an experienced researcher is a better control than a larger volume filtered automatically — but that position has a ceiling, and the harness above is how we intend to raise it without giving up the reading.

E · What are your data governance protocols?

13. How do you assess whether training data is accurate, complete and relevant to research objectives?

We do not train models, so the question applies to our input data. That input is official public statistics and academic probability surveys — INSEE and its national equivalents, Eurostat, the European Social Survey, the US Census Bureau — sources whose sampling design, field method and margins of error are published by the bodies producing them.

Relevance is assessed per market. Before a population source enters production we verify that the distributions we draw match the published national margins, and we re-verify when the source is updated. Coverage limits are recorded explicitly: where an axis is absent from a country's published statistics, we say so rather than substitute a proxy from another country.

Freshness is a real constraint and we treat it as one. These sources publish on multi-year cycles. Between publications our populations describe the country as last measured, not as it is today.

Sources
INSEE · Eurostat · European Social Survey · US Census Bureau

14. Do you document the origin and processing of training or input data?

Yes, and this is the answer we would most like buyers to compare across suppliers.

Every population we use derives from named, public, probability-based sources. In full: INSEE (France), ISTAT (Italy), INE (Spain), Destatis (Germany), the Office for National Statistics (United Kingdom), the US Census Bureau, INEGI (Mexico) and the Haut-Commissariat au Plan (Morocco), together with Eurostat and the European Social Survey.

The processing is margin-fitting: we draw simulated profiles so that the resulting sample reproduces the published distributions of each source.

What a third party can verify today, and what it cannot. Anyone can obtain these sources — they are public, free and downloadable — and check that the distributions we say we reproduce are the ones those bodies publish. That is a real check, and it is more than most suppliers in this category can offer: a proprietary panel cannot be re-derived by a buyer at all.

What a third party cannot do from this document alone is reproduce our own measurements — the reliability figures and correlations quoted under question 7. Those are our measurements, made on our protocols, and we present them as such rather than as independently replicated results. We say so because the word “auditable” should not be allowed to cover both cases: our sources are open to anyone, our measurements are ours until someone else runs them.

We think this is the sharpest distinction in the current market. Much of the field is grounded in proprietary panels and client first-party data. That grounding may well be excellent — but the buyer cannot verify it, and for the most prominent claims in this category we have not found an evaluation conducted by a party without a financial interest in the result. Ours can be checked by anyone who wants to check it.

Sources
INSEE · ISTAT · INE · Destatis · Office for National Statistics · US Census Bureau · INEGI · Haut-Commissariat au Plan · Eurostat · European Social Survey

15. Please provide the link to your privacy notice.

Our privacy policy is published at flashinsight.io/en/privacy-policy, alongside our legal notice at flashinsight.io/en/legal-notice and our AI usage framework at flashinsight.io/en/ai-usage-framework.

It states, per processing purpose, the legal basis, the retention period, the categories of subprocessor with their role and location, and how transfers outside the European Union are governed. It also states plainly where our providers' own terms allow processing outside the Union — including for services whose main infrastructure sits in Europe — because we would rather say that than let a reader discover it in a subprocessor's contract.

Hosting, database and payments are named individually. The model access layer is described as a category rather than by supplier name, because it is a configuration setting we change: naming today's supplier would date the page. We disclose the precise services in use on request, before a study begins.

16. What steps do you take to comply with data protection law and protect the privacy of research participants?

We hold no data that identifies a research participant. No names, no contact details, nothing that traces back to an individual who originally provided a response. This is a property of the design rather than a policy applied on top of it: our populations are built from published aggregate distributions, so there is no source individual to identify in the first place.

We are deliberate about the boundary of that statement. Like any business, FlashInsight holds personal data about its customers — account details, contact information, billing records, access logs. That is ordinary processing, governed by the privacy policy linked under question 15, and it is a separate matter from the research data below.

Three categories of research data, treated differently, and worth separating precisely.

Simulated profiles correspond to no real person. They are generated from published distributions and contain no personal data.

Validation data. Our published benchmarks use academic microdata files released by the bodies that collected them, in de-identified form and under their own access conditions. These files carry no name, no contact information and no key back to a respondent's identity. We use them for methodological validation rather than to build production populations, and we hold them under the terms attached to their release.

Client data. Material uploaded for a study is processed for that study. Client content is not used to train or tune any model — every call we make carries a setting that refuses data collection, covered by our automated tests; whether a provider honours that refusal is governed by their terms, which we check but do not enforce.

Our application runs from Paris and our database is in Ireland. We are candid in the privacy policy that several of our providers are US companies whose own terms permit processing outside the Union in certain cases, and that we cannot guarantee no data will ever be processed outside it.

17. What steps do you follow to ensure AI systems are resilient to adversarial attacks, noise and other potential disruptions?

Resilience is a different question for us than for a panel operator. With no human respondents, the dominant integrity risks in survey research do not apply: fraudulent respondents, duplicate identities, professional survey-takers, incentive gaming, click farms. Each assumes a person with an interest in deceiving the system. There is none. We say this because it is a real difference, not because it excuses us from the rest.

Five risks do apply.

Instruction injection through uploaded stimulus material. A client uploads an image, document or text as a stimulus. If that content carries instructions aimed at the model, it can alter its behaviour. This is the most concrete risk in our architecture, because it is the only point where outside content enters the chain.

We treated it as a question to be tested rather than an assumption to be stated. We ran deliberate injection attempts through the live production path and observed what the simulated respondents actually returned.

The first result was a failure, and we report it as such. All eight attempts succeeded. Four families of hostile instruction were planted at the exact point where an uploaded document's text enters the chain: a plain override in French, the same in English, a block impersonating a system-level protocol update, and a request to disclose our internal instructions. Every one of them worked. The simulated respondent answered with the attacker's chosen word to every question. The qualitative path failed the same way.

We then deployed one safeguard: delimitation. Textual stimulus material supplied by a client is enclosed in explicit markers and labelled as material to be described, never as instructions to execute. The study instruction sits outside that block and overrides anything inside it, and the markers are stripped from client content beforehand so that an uploaded text cannot close the block and write outside it.

Re-measured on the same live path, with that safeguard alone, no behavioural-override attempt succeeded: zero out of eight on the quantitative path, and zero out of three on the qualitative path. The tests are kept in our codebase and can be re-run at any time.

What this safeguard does not do, stated plainly. It prevents a stimulus from steering the answers. It does not prevent the disclosure of our own system instructions to the client who supplied that stimulus: in our own re-test, the disclosure attempt still partially succeeds. We built an output filter to close that, measured it, and removed it — the same sentence is a leak when the model recites it and a legitimate answer when a respondent quotes it, and every calibration we tried traded one error for the other. Destroying a client's genuine verbatim is a worse outcome than disclosing an instruction to that same client, so we chose not to ship the filter. This is a documented, accepted residual risk, not an oversight. It is bounded: there is no cross-client exposure, because each study's material is isolated to that study.

Two further limits. The delimitation covers textual material; images are transmitted to the model as images, and text embedded in a visual is not enclosed in the same way. And four attack families against two models is a measured reduction on the set we tested — not a guarantee, and not a claim of invulnerability.

We describe the test, the failure, the fix and the part we deliberately left open, because a resilience claim that was never tested is not a claim.

Output integrity. A model can produce a prohibited term, non-compliant formatting or a fabricated statement, even when the instruction explicitly forbids it. We do not rely on the instruction as the only safeguard: deterministic filters run after generation and neutralise what the model should not have produced — prohibited terms, and writing language in reports. This is a lesson learned: removing a term from the instructions is not enough, because a model will eventually reinvent it. The injection work above sharpened it in both directions. An explicit prohibition written into the instruction did not hold on its own — but an output filter aimed at a distinction a machine cannot draw reliably (is this sentence our leaked instruction, or a respondent quoting it?) destroys legitimate answers, and we removed the one we had built. Output filtering works where the forbidden output is defined by a rule. It does not work where the same text is legitimate or not depending on who is speaking.

Model variability. Two identical runs do not produce exactly the same result, and a quality indicator that flickers between runs is not usable. Our reliability thresholds are relative rather than absolute, and calibrated to absorb that variation: a score moving by a few points between runs does not change the verdict shown.

Provider unavailability. An outage at a model provider interrupts production. Our independence from any single model is also a resilience property — the model access point is a configuration setting, so a switch does not change the method, and our validity checks are re-run afterwards.

Silent run failure. The risk specific to our architecture is not that a study stops, but that it finishes looking complete when part of it failed. Scheduled recovery jobs detect and restart stalled runs across each production path — quantitative runs, qualitative runs, population generation and report production.

What we do. We run security audits periodically, targeted at specific surfaces rather than at the system as a whole: we measure before, we fix, we measure after, and we keep the test scripts so that anyone can replay them. The injection work described above is the most recent, and it is the pattern — pick a named risk, try to break it, publish what happened.

What we do not have. No continuous testing programme, no audit by an independent third party, no certification. We do not claim otherwise, and we would rather say it than let a buyer assume it. Our position is that the risks above are the ones our architecture actually creates, and that naming them is worth more than pointing to a certification we do not hold.

18. Do you clearly define and communicate ownership of data, intellectual property rights and usage permissions?

Material a client uploads remains the client's, and is used for that client's study only. This includes data supplied to calibrate a population: it is used to shape that client's sample, it never enters a model, it is never applied to another client's work, and ownership does not transfer at any point. Our method, our processing of public population sources, and our reliability instrumentation remain ours. The contractual expression of this sits in our master services agreement and in individual engagement letters.

19. Do you restrict what can be done with the data?

Yes, and the restriction is methodological before it is contractual. Our results are valid as relative readings at group and total level. They are not valid as absolute levels, as individual predictions, or as a substitute for direct measurement in tracking. We state this in the deliverable itself, because a restriction that appears only in a contract is not one a researcher will recall at the moment it matters. We are candid that this is a stated boundary rather than a technically enforced one.

20. Are you clear about who owns the output?

The client owns the deliverable and the results of their study. We retain the right to use aggregate, anonymised methodological performance data — reliability scores and validation metrics — to improve and publish on the method itself. We publish no client content and no study result without permission.

Reading our numbers correctly

A third figure appears elsewhere in our published material and we reconcile it here rather than leave a reader to wonder. ρ = 0.94 is a rank correlation between our calibrated scores and a reference field panel, measured across a six-concept test, against a validation threshold we set at 0.80. It is a different protocol from the two under question 7 — a different question, a different comparison set, a different unit. The three figures are not versions of one another and should not be read as a range. Each belongs to the protocol that produced it, and we name the protocol every time we quote a number.

Two things the question 7 table does not claim. It is not individual-level accuracy. And 26% is a percentage of a ceiling, not a percentage of correctness — individual answers are only weakly predictable even from a person's own history, which is why the figure is expressed relative to that ceiling rather than against a notional perfect score.

Alongside that validity base, a reliability index runs on every sample in production. Each sample receives a score from 0 to 100 with a traffic light. Components, weights and thresholds are documented and version-stamped, so any change to the method invalidates prior scores rather than silently altering them.

One component is a gate rather than a score. If the axes defining a sample do not actually move the responses, the index is forced to red and no coherence score is displayed. This targets the failure mode the academic literature identifies as central to synthetic respondents: models flattening the groups they are asked to represent. In reviewing the published materials of the suppliers we could examine, we did not find an equivalent automatic control; we state that as the result of a search, not as a claim about the whole market.

Reliability is calibrated relative to an observed band, never against a notional perfect score. Absolute self-consistency scores are low across all current models, and the most defensible figure in the published literature is itself expressed relative to the human ceiling.

Four questions we would add to these twenty

ESOMAR's twenty are the right floor. Four more separate suppliers who measure from suppliers who assert.

  1. 1.Show me your calibration record — not the accuracy claim, but the record of how often stated confidence matched actual accuracy over time.
  2. 2.Who audited it? If every voice attesting to a method has a financial interest in it working, that is not independent evidence.
  3. 3.Show me out-of-sample performance on a population the method was not built around.
  4. 4.What is your refresh cadence, and what happens to accuracy between refreshes?

We do not yet have a complete answer to the first. We are building it, and we will publish it before we claim it.

Further reading

Our own public notes on the questions above. They are our work, not third-party evidence, and should be read as such.