FlashInsight — What is a synthetic panel? The fifth generation of market research.
Published onEN

What Is a Synthetic Panel? A Reference Guide to Synthetic Respondents and the Fifth Generation of Market Research

A reference guide to synthetic panels: what they are, how synthetic respondents are actually built, what the industry is really doing, what the science says, and where they create decision-grade value — framed as the fifth generation of market research.

Synthetic panelsMethodFifth generation

“Synthetic panel.” Two words that, in barely eighteen months, have become the most divisive phrase in market research. To some, they signal the end of the survey as we know it. To others, a dangerous mirage. Both camps, I will argue, are asking the wrong question.

I have spent more than twenty-five years in this industry — as a practitioner, as a former President of ESOMAR, and today as a professor at Université Paris-Panthéon-Assas and founder of FlashInsight. In that time I have rarely seen a topic generate so much heat and so little shared vocabulary. The question I am asked most often, by clients and students alike, is deceptively simple: what, exactly, is a synthetic panel?

This article is the reference guide I wish existed. It does four things, in order. First, it defines what a synthetic panel is and clears the terminological fog. Second — and this is where I will spend the most time — it explains how synthetic respondents are actually constructed from synthetic data. Third, it maps what the industry is really doing, from Toluna to Ipsos to Kantar. And fourth, it confronts the false debate between “it works” and “it can’t replace humans,” to bring us back to the only thing that matters: the decision you are trying to make.

What Is a Synthetic Panel?

That is the short answer. The longer answer requires a distinction that most of the debate ignores — the difference between a synthetic persona and a synthetic panel.

In its plain-language “Terms of engagement” series, Research Live (the editorial arm of the UK’s Market Research Society) puts it well: a synthetic persona is a representation of a single consumer archetype — a rich, queryable character. A synthetic panel is something larger: it generates results that simulate a sample of many virtual respondents, with quotas, distributions and segmentation, the way a real panel does.

The practical consequence matters. A persona is a tool for empathy and exploration — you interview it, you stress-test a message against it. A panel is a tool for measurement-like signals — you run a concept through it and read the spread of reactions. Confusing the two is the source of half the bad press synthetic research receives: people build a single persona, treat its output as a representative survey, and are then surprised when it does not behave like one.

First, Get the Vocabulary Right

Before we can evaluate anything, we have to stop using five different words as if they were one. “Synthetic data,” “synthetic respondents,” “synthetic personas,” “digital twins” and “generative agents” describe a progression — from statistical to individual, from static to dynamic — not a single object.

TermWhat it actually is
Synthetic dataArtificially generated data that mimics the statistical properties of real data without containing personally identifiable information. The raw material, not the respondent.
Synthetic respondentAn instance of an LLM asked to answer questions as a human would. The unit of a synthetic panel.
Synthetic personaThe “briefing” given to the model — the set of traits, attitudes and characteristics that guide its answers.
Digital twinAn attempt to replicate a specific, real individual rather than an archetype.
Generative agentThe most sophisticated architecture: memory, reflection and planning to simulate dynamic, long-term behaviour.

Hold on to one idea from this table: synthetic data is the input; the synthetic respondent is the output. A synthetic panel is what you get when you assemble many synthetic respondents into a structured sample. That is why the question “how do you build synthetic respondents from synthetic data?” is the real heart of the matter — and where we go next.

How Are Synthetic Respondents Built?

This is the section that the loudest voices in the debate tend to skip. Yet you cannot judge whether a synthetic panel is trustworthy without understanding how it was made. There is an enormous difference between a synthetic respondent improvised by a generic chatbot and one assembled inside a disciplined, calibrated architecture.

From a question to an answer: what a “response” really is

Start with a deceptively deep question: what is a survey response? For decades we implicitly treated it as the extraction of a pre-existing opinion, as if attitudes sat in mental drawers waiting to be opened. Cognitive science says otherwise. Answering a question is an act of construction — the response is assembled in the moment from memory, emotion, context and the framing of the question itself.

This reframing is what makes synthetic respondents conceptually plausible at all. LLMs are, fundamentally, contextual generators. If a human response is constructed rather than retrieved, then a model whose entire architecture is built for contextual generation is not as far from the task as critics assume. That does not make its answers “true.” It means the comparison with human answers is subtler than a simple pass/fail.

The three ways to build a synthetic panel

There is no single method. Drawing on the MRS / Research Live taxonomy and on my own work at FlashInsight, synthetic panels are built in three broad ways — in increasing order of rigour:

  1. Generic generation. Prompt a large, general-purpose LLM to role-play respondents from its training data alone. Fast and cheap, but ungrounded: this is where most of the documented failures come from.
  2. Statistical extension of a real sample. Use machine learning trained on a smaller real survey to generate additional respondents (“vertical scaling”) or to simulate how the same respondents would answer new questions (“horizontal scaling”). This is closer to classic data augmentation than to pure generative AI.
  3. Grounded and calibrated generation. Build personas from real population data, condition the model on them, and calibrate the outputs against real human responses for the targeted cohorts. This is the only approach that earns the word “panel” in any serious sense.

The lesson from the failures in the literature is almost always the same: they tested method 1 and concluded that the whole category was broken. As I have written before, fully synthetic audiences did not fail — poorly grounded ones did.

The anatomy of a grounded synthetic respondent

What does “grounded and calibrated” look like in practice? A defensible synthetic panel is built in layers, not in a single prompt:

  • A population frame. Start from a real-world distribution — census or official statistics (in France, INSEE; in the US, the US Census), an existing client panel, or a known segmentation — so the sample’s structure is anchored to reality, not invented.
  • Demographic sampling. Draw synthetic respondents so that age, gender, region, income and category behaviour match the quotas you would impose on a human panel.
  • Psychographic conditioning. Enrich each profile with attitudes, values and personality — drawing on validated, peer-reviewed frameworks of personality and human values — so respondents differ in how they react, not just in who they are.
  • Grounding and enrichment. Inject real context — client data, prior studies, category knowledge — through retrieval or calibration, so the model is reasoning about your market, not about an average of the internet.
  • Calibration against real data. Wherever real responses exist, tune the system so that its outputs reproduce known human distributions for those cohorts. This is the step that separates a measurement-grade tool from a plausible-sounding one.
  • Versioning and traceability. Treat prompts, parameters and data sources as you would treat a questionnaire: documented, versioned, auditable.

Design methods for the tool, don’t force it to replicate ours

Here is the principle that should guide the whole field: build methods that play to what the model does well, instead of forcing it to imitate instruments designed for humans. The closed rating scale is the clearest cautionary tale.

Ask an LLM to “rate this concept from 1 to 5” and you get artificial, under-dispersed distributions that regress to the mean and look nothing like human data. This is not a random glitch — it is a direct expression of the model’s positivity bias, its trained disposition to be agreeable, helpful and accommodating. Faced with a closed scoring question, that bias compresses the variance and inflates the top of the scale. The model is structurally unsuited to the closed-rating task, however cleverly we prompt it.

The right response is not to keep hammering the model with a human format until it complies. It is to change the method. Rather than demand a number, ask for a natural-language reaction — what LLMs are built to produce — and derive the score from that text, for instance by comparing it against anchored descriptions of each scale point. The distribution is then reconstructed from language instead of extracted by fiat, and it recovers the spread that a direct rating destroys. The lesson generalises far beyond scales: the problem is usually our method, not the model — and the fix is to design AI-native methods, not to replicate ours.

What the Industry Is Actually Doing

The synthetic panel is no longer a lab experiment. In 2025–2026 it moved from experimental to operational, and the major insight companies have taken visibly different routes.

  • Toluna launched HarmonAIze Personas — over one million synthetic personas across 15 markets and 9 languages — trained on its panel of 79M+ real consumers, with an “ACT Instant” ad-testing product that returns results roughly 20× faster than traditional testing. Their bet: scale and speed grounded in a vast real panel.
  • Qualtrics built Edge Audiences as a flexible platform offering human, synthetic, or hybrid samples, with fine-tuned models trained on 10,000+ studies and 7.5M responses, and human validation via PureSpectrum in the same workflow. Their bet: let the buyer choose where on the human–synthetic spectrum to sit.
  • Ipsos has partnered with Stanford University to develop rigorous validation frameworks for synthetic data — an explicitly academic, methodology-first stance.
  • Kantar and NIQ (BASES) have folded AI testing into their core offers, with NIQ’s clients framing the value operationally — “to learn early, fail fast, and optimise quickly,” in the words of Reckitt’s Chief Insights Officer.
  • Specialists and challengers complete the picture: Evidenza (B2B-focused, reporting ~88% accuracy across 100+ studies and an EY double-blind validation), Synthetic Users (UX research), and FlashInsight, among others.

The trade press and the conference circuit have caught up too. At ESOMAR’s 2025 congress, AI — and synthetic data specifically — was, in Ray Poynter’s phrase, “the connective tissue running through the programme.” Poynter, founder of NewMR and a long-standing ESOMAR voice on responsible research, has been among the most useful commentators precisely because he refuses both hype and dismissal: his guidance focuses on governance — what synthetic data is, how it is used, its benefits and risks, and how to implement it well. The ESOMAR “5 topics to help buyers of augmented and synthetic data” is in the same constructive spirit. Consumer psychologist Paul Marsden has likewise been chronicling the rise of “synthetic consumers” and digital twins, helping practitioners separate genuine capability from marketing gloss.

The signal across all of this is consistent: the question has shifted from whether the synthetic has a place to how to govern it.

The False Debate: “Does It Work?” vs “It Can’t Replace Humans”

Most public argument about synthetic panels collapses into a binary — revolution or mirage. It is a badly posed question. Let me take both sides seriously, because both are partly right.

What the science actually says

The documented limitations are real, and we must accept them. Studies such as Bisbee et al. (2024), published in Political Analysis, show that LLM-generated panels fail to reproduce population distributions on politically sensitive or low-base-rate variables. The recurring weaknesses: poor individual precision, under-dispersed variance, temporal instability, sensitivity to prompt wording, under-prediction of awareness and over-prediction of purchase intent, and difficulty with genuinely novel topics. Replacing a representative human sample with synthetic interviews to “measure reality” remains, in many contexts, an error.

The documented strengths are equally real. Park et al. (2024), “Generative Agent Simulations of 1,000 People”, showed that agents built from in-depth qualitative interviews reproduced benchmark survey responses with striking fidelity. Columbia’s Olivier Toubia reports digital-twin responses matching original human answers with ~85% accuracy — comparable to a human’s own consistency two weeks later. Earlier foundational work — Argyle et al. (2023), “Out of One, Many” and Brand, Israeli & Ngwe’s HBS working paper on using GPT for market research — pointed the same way: at the aggregate level, on well-scoped tasks, results can be solid.

Both findings are true; both are partial. The cases studied are not interchangeable, and conclusions drawn from one should not be transplanted onto the other.

The category error

Why do so many tests disappoint? Because we commit a category error. We take questionnaires designed for humans — Likert scales, closed questions, standardised protocols — and feed them, unchanged, to language models. It is like the first motorists hitching horses to the front of their cars. We are testing an engine’s ability to pull a cart, then blaming the engine. The skeptical studies almost always used classic elicitation applied as-is to LLMs. AI-native methods tell a very different story.

The elephant in the room

There is also an uncomfortable asymmetry in the debate. While the profession scrutinises AI’s flaws, a very real problem is quietly eroding our craft: the collapse of online panel quality. The GRIT Business & Innovation Report records a sharp rise in data-quality concerns — fraudulent respondents, bots, disengaged participants, compromised samples. Criticising synthetic respondents in the name of data quality while ignoring the degradation of human panels is comparing a theoretical ideal to a degraded reality. The honest question is not “perfect human vs imperfect synthetic.” It is how to combine imperfect sources to make better, faster, more robust decisions.

The MRS Delphi report on synthetic respondents (2024) offers the cleanest reframing: synthetic outputs are fit for many operational, short-term decisions — on condition that they sit inside a rigorous, transparent protocol.

From Panel to Decision: Where Synthetic Panels Create Value

The productive question is not “can they replace humans?” but: for what kinds of decisions are the signals from a synthetic panel already good enough? Two functions sit side by side, and conflating them is the original sin.

  • Measurement — statistical rigour, representativeness, longitudinality — remains irreplaceable for decisions that commit significant capital. This is the home of the human (and increasingly the calibrated hybrid) panel.
  • Simulation — long practised through conjoint analysis and agent-based models — has been economically transformed. What took weeks now takes hours, which lets us explore before we commit.

That is why I keep arguing that simulation, not prediction, is what synthetic research is really for. Used this way, synthetic panels already earn their place in:

  1. Ideation and exploration — generating hypotheses, objections and formulations, where you want diversity and plausibility, not precise prevalence.
  2. Pre-tests and concept screening — filtering many concepts cheaply before investing in human quant.
  3. Scenario stress-tests — mapping how contrasting segments react to crises, claims or extreme stimuli.
  4. Augmented interpretation — using LLMs after the quant to enrich verbatims, surface weak signals and reframe insight.
  5. Hard-to-reach populations — where well-constructed synthetic agents can open doors that fieldwork cannot.

In every case the decisive test is the same: not “is this true to the decimal?” but “is this reliable enough for the decision I must make now?”

A Short History: The Five Generations of Market Research

None of this is a rupture without lineage. It helps to see the synthetic panel as the latest chapter in a long story — the evolution of how our profession listens to people.

GenerationEraDefining method
FirstMid-20th centuryFace-to-face and postal surveys; the birth of representative sampling.
Second1970s–1980sTelephone interviewing (CATI); speed and national reach.
Third1980s–1990sComputer-assisted interviewing and standing access panels.
Fourth~2000sOnline research and digital panels (Giannelloni & Vernette, 2000; Toubia & Florès, 2007).
Fifth2024 onwardAI-native research: synthetic respondents, simulation and calibrated hybrids.

Each transition triggered the same reflex. When the internet arrived, institutes simply digitised existing focus groups and put questionnaires online — they treated a new paradigm as a new channel. We are at risk of repeating that mistake with AI: bolting language models onto methods designed for a different era.

The fifth generation will not be defined by who wins the synthetic-respondent debate. It will be defined by those who design methods native to AI — exploiting what it does best, framing what it does worst, and refusing the false choice between purist nostalgia and uncritical adoption.

The role of the expert changes accordingly. We become simulation architects and data guardians — designing relevant experiments, validating outputs against the real world, and translating statistical signals into strategic decisions.

How to Use Synthetic Panels Well: Three Operating Principles

If you take nothing else from this guide, take these three rules. They are the difference between a synthetic panel that informs decisions and one that quietly misleads them.

  1. Separate exploration from validation. Use AI to expand hypothesis spaces and stress-test concepts. Reserve human (or calibrated hybrid) data for decisions that commit budgets.
  2. Work on relative deltas, not absolute levels. Which concept beats which is reliable. The absolute score is not. Calibrate against real data wherever you have it, and version your protocols as you version questionnaires.
  3. Trace prompts and provenance. Trust scales with visibility, not with claimed performance. Document the training data, the generation parameters and the known biases. Then, before any high-stakes decision, impose a cheap reality check — a handful of real interviews, a small quant top-up, a benchmark against an existing survey.

Conclusion: Method Over Magic

The debate about synthetic panels, as it is usually framed, pits an idealised past against a fantasised future. The quest of those who defend representativeness is legitimate — our profession exists to give decision-makers as undistorted a reflection of reality as possible. But the illusion of “perfectly measured reality” has itself become a mirage in an age of collapsing panel quality.

A synthetic panel is not a magic button that replaces the consumer. It is a powerful new instrument — most valuable for exploration, simulation and speed, most dangerous when sold as a substitute for measurement. Built with grounding, calibration and traceability, it belongs firmly in the modern researcher’s toolkit. Built carelessly, it produces synthetic certainty, which is worse than no certainty at all.

The history of our discipline shows that every technological rupture has enriched it — provided we had the courage to change our methods, not just our tools.

There is no magic. There is method. Done right, that is enough.

Frequently Asked Questions

What is a synthetic panel?

A synthetic panel is a sample of AI-generated respondents — virtual individuals whose answers are produced by large language models — built to simulate, test, forecast or augment what a human survey sample would say. Each respondent is conditioned on a profile, and the panel reproduces quotas and distributions like a real one.

What is the difference between a synthetic panel and a synthetic persona?

A synthetic persona represents a single consumer archetype you can interview or stress-test. A synthetic panel is a structured sample of many synthetic respondents, designed to produce survey-like, aggregate signals.

How are synthetic respondents created?

Three main ways: generic generation from a general-purpose LLM; statistical extension of a smaller real sample; and grounded, calibrated generation in which personas are built from real population data and the model’s outputs are tuned against real human responses. The third is the most reliable.

Are synthetic panels accurate?

It depends entirely on construction and use. They are weak at individual precision and absolute levels, and strong at aggregate, relative comparisons on well-scoped tasks — provided the method is AI-native rather than a closed format borrowed from human surveys.

Can synthetic panels replace human panels?

Not for high-stakes measurement. They complement human data: use them for exploration, screening and simulation, and reserve human or hybrid validation for decisions that commit significant resources.

When should you use a synthetic panel?

For ideation, concept screening, scenario stress-tests, augmented interpretation, and reaching hard-to-reach audiences — whenever the decision needs a fast, directional signal rather than a precise measurement.

References and Further Reading

  • Argyle, L. P., et al. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis. doi.org/10.1017/pan.2023.2
  • Bisbee, J., et al. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis. doi.org/10.1017/pan.2024.3
  • Brand, J., Israeli, A., & Ngwe, D. (2023). Using GPT for Market Research. Harvard Business School Working Paper.
  • ESOMAR (2025). 5 Topics of Discussion to Help Buyers of Augmented and Synthetic Data. esomar.org
  • Florès, L. (2018). How open data may be leveraged to help develop a new “research paradigm”. International Journal of Market Research, 60(4). doi.org/10.1177/1470785318771449
  • Greenbook (2025). GRIT Business & Innovation Report. greenbook.org/grit
  • Ipsos & Stanford University (2025). Pioneering the future of market research with synthetic data. ipsos.com
  • Market Research Society (2024). Using Synthetic Respondents for Market Research. Delphi Report. mrs.org.uk
  • Park, J. S., et al. (2024). Generative Agent Simulations of 1,000 People. arXiv:2411.10109. arxiv.org/abs/2411.10109
  • Research Live / MRS. Terms of Engagement: Synthetic Personas & Synthetic Panels. research-live.com
  • Toluna (2025). One Million Synthetic Personas. tolunacorporate.com
  • Toubia, O., & Florès, L. (2007). Adaptive Idea Screening Using Consumers. Marketing Science, 26(3).

This article was originally published in French on L’Atelier IA.

Laurent Florès is a professor at Paris-Panthéon-Assas University, former ESOMAR President, and founder of L’Atelier IA and FlashInsight. He writes about AI, research, and the human factor.