Customer Research

Limits of Synthetic Data in Consumer Research

Synthetic data works for concept testing but fails on pricing and behavioral predictions.

Columnist · · 8 min read
Cover illustration for “Limits of Synthetic Data in Consumer Research”
Synthetic Consumer Panels · September 5, 2026 · 8 min read · 1,872 words

Synthetic data works in consumer research, but only for a narrow slice of the job. Most of the industry is currently pretending that slice is bigger than it is, and knowing exactly where the boundary sits is the whole skill right now. Qualtrics surveyed 3,200 market research professionals in November 2024 and found 73% had already used synthetic responses at least once. About seven in ten expect synthetic data to make up more than half of all consumer research within three years. That expectation is ahead of what the evidence actually supports, and the gap between the two is where the risk lives.

What synthetic data actually is in a consumer research context

Synthetic consumer panels are AI-generated respondents built from real-world material: old surveys, behavioral logs, customer reviews, public opinion trends. A model trained on that stack learns the patterns in it, then generates new answers that follow those patterns.

Hold onto that distinction, because it explains everything else in this piece: synthetic panels model what's already been observed, and they don't go out and collect anything new about the world. Two versions show up in practice. Persona-based simulation asks a model to answer as a defined type, say a budget-conscious family of four or a compliance officer at a mid-size bank. Behavioral simulation goes further and tries to predict how a whole segment would react to a price change, a new concept, or a tweaked feature.

The appeal isn't complicated. Traditional panel recruitment eats up roughly 60% of a research project's timeline, and the average connected consumer now gets 15 or more survey invitations a month, with cold-email response rates under 2% in most categories. Synthetic panels skip that grind entirely. They also reach people recruitment can't touch: enterprise CTOs, niche technical buyers, reactions to product concepts that don't exist yet. That upside is real, which is exactly why the limits deserve a precise answer instead of a shrug in either direction.

Where synthetic data performs reliably and where the evidence is strong

A 2025 review of 14 key studies found 86% reported at least partial success getting synthetic responses to mimic human ones, with half showing strong similarity. That's a real signal the technology moved past novelty.

But read the other half of that number too. Partial success across 86% of studies still means a meaningful chunk didn't replicate human behavior at all, and the headline figure hides which task types failed. The pattern underneath it: synthetic data holds up on pattern-matching, what kind of person tends to buy this, and falls apart on situational behavior, what will this one person do standing in this one store on a Tuesday.

Inside that safe zone, a handful of applications actually earn their keep. Concept screening is one; synthetic panels show less positivity bias than incentivized human panels, so they separate a strong concept from a mediocre one with a wider, more useful spread. Purchase intent modeling at the category level is another. A 2025 study titled "LLMs Reproduce Human Purchase Intent" found close alignment with real human data, provided the prompts and mapping methods behind them are built correctly. Message and positioning work benefits too, especially in B2B, where a simulated CTO persona can pressure-test a pitch before it reaches a real buyer. And for early hypothesis generation, synthetic panels are cheap and fast enough to help a team walk into human testing with sharper questions already in hand.

The bias problem synthetic data inherits from its training sources

Synthetic output only carries the quality of what trained it. Research has found that biases baked into training data don't just show up in synthetic responses, they get amplified on the way out.

Part of the mechanism sits inside how these models work. Asked to simulate a population, a language model drifts toward the modal, most-likely answer instead of reproducing the messy spread a real population actually has. Research has found that a substantial share of the correlations between variables in synthetic panels end up differing from what real human panels show. That's a structural feature of how the outputs get produced, not noise that averages out.

Geography makes it worse. Most large language models train predominantly on English-language data from North America and Western Europe. Ask one to simulate a consumer in Lagos, Manila, or São Paulo, and what comes back often reads like a Western perspective wearing a local costume, missing the regional texture a real respondent would give you. Catching that requires checking against real respondents from that market, which defeats half the point of going synthetic in the first place.

The ceiling shows up in demographic work too. Research has found that GPT-based models can't meaningfully reflect different preferences across demographic groups. Studies have found brand preference can be nudged toward demographic alignment only when a prompt explicitly names individual attributes, and even then, brand perceptions stay flat across groups that should differ in real life. Any research question that depends on genuine variance between subgroups is walking into a high failure risk here.

Why willingness-to-pay and pricing research are particularly unreliable from synthetic sources

Pricing research is probably the single riskiest place to lean on a synthetic panel. Research has found that off-the-shelf GPT estimates of willingness-to-pay are often wrong, and not wrong by a small margin either: the direction flips, meaning the model can say customers will pay more for a feature when the true answer runs the other way.

Fine-tuning against category-specific human conjoint data can recover the right sign and a rough sense of scale. But that fix needs real human data to calibrate against in the first place, which quietly cancels out the reason someone reached for synthetic data to begin with.

Layer onto that the fact that human willingness-to-pay answers are already noisy on their own. Respondents routinely understate their ceiling because they're negotiating, whether they realize it or not. Stack synthetic distortion on top of that existing noise, and the error compounds instead of canceling out. Worse, synthetic outputs cluster tightly around a mean, so a pricing study can hand back a confident-looking price band that no real population actually supports. The precision is cosmetic; the accuracy behind it isn't there. Pricing decisions, conjoint work, revenue modeling: none of it should run primarily on synthetic evidence. Use it to shape a hypothesis, never to set a number finance signs off on.

What synthetic data cannot capture in UX and usability research

Figma's 2025 AI report found 24% of designers and 40% of developers already use AI somewhere in testing. So adoption is real here too.

But usability testing runs on signals a synthetic respondent can't produce. Hesitation and task abandonment, the moment a user stalls, backtracks, or gives up on a flow, is exactly the friction a redesign needs to find, and no model trained on past data can predict where that moment lands on a UI it has never actually touched. Emotional reaction works the same way. A synthetic panel can output the word "confused." A real usability session shows confusion arriving on its own, usually somewhere the research team never expected. People do strange, specific things with software, and what a user actually does routinely diverges from whatever a model predicts as "expected" behavior.

The failure mode here is quiet, and it shows up late. A team validates a design against synthetic responses, ships it, then finds out real users are abandoning the flow at a step nobody flagged, weeks after launch, once the support tickets start piling up.

AI still has a job in UX research, just a different one than standing in for the participant. AI-moderated interview platforms can run real people through structured tasks at scale, drawing on participant networks in the tens of millions to fill studies fast. AI is also useful for finding patterns across a big pile of real session recordings, cutting the analysis workload without touching the underlying evidence. And it speeds up screening and recruiting, getting actual users into a study faster. AI as infrastructure around the research serves one purpose; synthetic data as a stand-in for the humans inside it serves a different one. Treating those as the same tool is where teams get burned.

The contamination problem that makes real-panel quality harder to trust too

Real panels were never automatically clean, so this isn't a new problem dressed up as one. Commercial panels have always carried a share of respondents speeding through surveys purely for the incentive payout, and through 2025, AI-generated respondents slipping past basic attention checks became a fast-growing addition to that same pile.

Panel quality audits have consistently found meaningful fraud rates in unmanaged commercial panels. That's a serious quality problem, and it has nothing to do with synthetic data. It predates the current AI wave by years.

Here's the irony worth sitting with. One argument for synthetic panels is that they dodge the contamination problem plaguing human panels, and there's some truth in that; a synthetic panel's composition is at least known and fully controlled by whoever built it. But synthetic data trades that risk for a different one: inherited bias and hallucination, where a model hands back a confident, coherent-sounding answer with no basis in any real consumer's actual behavior. Quality control isn't optional on either side of this. The checks just look different, and neither source certifies itself.

How to map research tasks to the right data source

Diagram: Where Synthetic Data Holds Up vs. Falls Apart. Visualizes: Visualize a two-zone decision map showing which consumer research tasks synthetic data can handle reliably versus where real respondents are non-negotiable.

Treat synthetic AI as a junior partner in the process, not a co-lead, and save it for the lower-stakes moments rather than the ones carrying real financial weight.

Synthetic data earns its place in a few spots: early concept screening and hypothesis filtering before a team commits budget to a full human study; rapid iteration on messaging and positioning, especially inside demographic and cultural contexts the training data actually covers well; stress-testing assumptions across a wide range of simulated answers as a first pass, never the final word; and early exploration of hard-to-reach B2B segments, where a rough read on a niche technical buyer beats waiting weeks for recruitment to fill a panel.

Real respondents stay non-negotiable everywhere else. Pricing, conjoint analysis, and any willingness-to-pay work feeding an actual financial decision. UX and usability testing, where the evidence is hesitation, confusion, and abandonment, none of which a model can convincingly fake. Research in non-Western markets, where training-data bias skews the output in ways hard to catch without a local respondent to check it against. Any study built around differences between demographic subgroups, since synthetic panels flatten exactly that variance. And any validation study feeding a go or no-go call on a launch or a new market.

The smarter approach here is sequencing. Run synthetic panels first to sharpen the questions and shrink the sample size the human study actually needs, then move faster into real fieldwork instead of skipping it. Build a synthetic model once, calibrated properly against real respondents, and it turns into something a team keeps running against new decisions instead of rebuilding from scratch each time.

Most teams that get burned make the same mistake: they treat a synthetic output as the final answer, instead of a fast, structured input that still needs a real human check before anything gets decided.

Sources

  1. customerexperiencedive.com
  2. quirks.com
  3. medium.com

More in Synthetic Consumer Panels