top of page

"AI Can Simulate Your Customers. It Can't Live Their Experience."

Curious People
Sep 8
7 min read

Updated: Sep 9


A question has been circulating in research and innovation circles lately, usually asked with some excitement: why not just create an AI persona of your target market and run your qualitative and quantitative research on that instead? It is faster, cheaper, and infinitely patient. So why keep bothering real people at all?


We have spent some time with this question. Our answer is not "the technology isn't good enough yet." Our answer is that the question rests on a category error — and once you see it, no amount of future technology fixes it.



Why the Temptation Is Real


It is worth being honest about why this idea is spreading so fast, because the pull is genuine and the numbers behind it are not exaggerated.


A traditional research study — proper recruitment, incentives, moderation, transcription, analysis — typically runs $5,000 to $65,000 and takes four to twelve weeks to deliver a result. A synthetic panel covering the same ground can be generated in an afternoon for somewhere between one and a few hundred dollars, sometimes less than the cost of a coffee for a full ten-thousand-persona pass. Hard-to-reach populations — niche B2B buyers, people in regulated industries, audiences scattered across time zones — that would take weeks of recruiting effort can be modelled in minutes. For a founder validating an early idea, or a team wanting directional signal before committing real budget, that arithmetic is genuinely seductive.


So it is worth asking directly: does the effort of building these personas actually outweigh what they deliver, if there is any real benefit at all? The honest answer is more interesting than a flat yes or no. On raw cost alone, the answer is no — generating a synthetic panel is cheap, often astonishingly so, and cheaper than almost anyone assumes. The resource imbalance is not in the building. It is in what building cheaply actually buys you, and what you still have to spend afterward to trust it.


Here is the part the cost comparisons leave out. A synthetic result still has to be validated against a real human benchmark before anyone can responsibly act on it — and that validation requires exactly the real research the synthetic panel was meant to replace. Skip that step, and you are not saving effort; you are deferring a cost and adding risk on top of it, because a wrong number in a strategy deck is far more expensive than a slow one. The studies that found synthetic samples matching real averages while a third of the underlying relationships pointed in the wrong direction did not fail because building the persona was hard. They failed after the persona was built, cheaply, and trusted too quickly. The resource that actually gets spent disproportionately is not compute — it is the diligence to check the output, and that diligence is precisely the step the promise of speed tempts teams to skip.


So the real cost-benefit question is not "is it expensive to make," but "what does it cost you when it's wrong and nobody checked." That is where the effort-to-benefit ratio actually breaks down — not at the point of creation, but at the point of decision.



The Numbers Tell an Interesting Story, But Not the Whole One



Start with what the evidence actually shows on accuracy, because it is more nuanced than either the hype or the backlash suggests. Calibrated synthetic respondents can hit real accuracy on well-trodden quantitative patterns. One academic study out of the University of Mannheim and ETH Zürich, tested against 57 real personal-care product surveys and 9,300 human responses, found a technique called Semantic Similarity Rating achieved roughly 90 percent of human test-retest reliability on purchase intent. Interestingly, the method only worked when the model was allowed to answer in free text, which was then mapped to a rating scale by meaning — forcing it to just pick a number on a five-point scale collapsed the answers toward the middle and produced far weaker results. The signal lived in the narrative, not the number.


But the same body of research contains a warning that matters more than the headline figure. Researchers at Prolific simulated nearly a thousand individuals using demographic personas — the most common commercial approach — and found it roughly tripled distributional error compared to simply asking the model for one honest aggregate guess. The personas collapsed onto stereotyped, modal answers. Feeding in real interview transcripts helped, but nothing beat the accuracy of just running a second batch of real humans. Other studies found something worse hiding beneath aggregate averages that looked fine: nearly a third of the underlying statistical relationships in synthetic data pointed in the wrong direction entirely, even while the averages matched. You could build a strategy on a number that looked right and was actually backwards, and never know until it was too late.


These are real, fixable-sounding engineering problems. Better prompting, better calibration, more transcripts, next year's model. Which is exactly why they are not the real argument.



Two Kinds of Tool, and a Question Neither Was Built to Answer



Here is where the conversation usually gets interesting, and where we think the entire debate has been asking the wrong question.


A weather model does not need to get wet to predict rain. A flight simulator does not need to leave the ground to model aerodynamics. Both are excellent tools precisely because prediction without instantiation is genuinely possible — a system can output accurate behaviour about a phenomenon without having lived through it.


But notice what neither tool was ever asked to do. A weather model was never commissioned to report what it felt like to be rained on. A flight simulator was never asked what it is like to fly. Their job was always behavioural output — will it rain, does the aircraft respond correctly to this input — and that is a job a sufficiently good model can, in principle, do well.


A synthetic consumer persona is asked to do something categorically different. Its entire commission is to stand in for what it was like for a real person to encounter a product, a workplace, a moment of change. That is the one thing the weather-model logic was never built to answer, because it was never the target variable in the first place. The comparison that seemed to defend synthetic personas — "prediction without instantiation happens all the time in engineering" — actually proves the opposite once you notice what's being predicted. Lived experience was never on the menu for a weather model. For a synthetic consumer, it's the only thing ordered.



An Entire Industry Has Already Ruled on This


The more we sat with this, the more we noticed that serious institutions have already faced this exact question and settled it — not through philosophy seminars, but through hard, codified rules.


Take aviation. Simulator time is not merely discouraged as a substitute for real flying hours — it is capped by federal regulation. An Airline Transport Pilot certificate requires 1,500 hours of total flight time, and a simulator, however advanced, can be credited toward only a bounded slice of that: a maximum of 25 hours toward one requirement, 50 toward another, and no more than 100 hours total across the entire 1,500-hour rule, regardless of how many extra hours a pilot logs in the box. Some categories of training device are barred from counting at all in places where higher-fidelity simulators are permitted in limited amounts. This is not a temporary rule waiting for better simulator technology. It is a permanent, written acknowledgment that rehearsal and the real thing are different in kind, not degree — and that no simulator fidelity closes that gap, because the gap was never about fidelity.


Or take music. A violinist can practise alone at home every day for years. That practice does not substitute for rehearsing with an orchestra. Orchestral rehearsal, however extensive, does not substitute for the stage. And even a full dress rehearsal on that exact stage does not substitute for the actual concert — because an audience is present, the performance is unrepeatable, and something genuine is at stake that rehearsal, by definition, cannot replicate. Each stage moves closer to the real thing. The gap never closes. It only becomes more expensive to keep pretending it isn't there.



What This Means for How Organisations Listen


We think market research is facing the same question aviation and music already answered, and it deserves the same clarity of rule rather than an accuracy debate that quietly resets every time a new model ships.


The honest distinction is not "synthetic personas are inaccurate." Some are becoming remarkably accurate at specific, bounded tasks — pre-testing a questionnaire, stress-testing a hypothesis before fieldwork, screening a large number of early concepts. Those are legitimate uses, the practice-room uses, and no one credible argues otherwise. The Market Research Society's own guidance is precise on this point: synthetic participants should be used "primarily as supplements to human judgment rather than replacements," never as the source of the actual finding.


The distinction that matters is what the output is being asked to stand in for. The moment a persona is presented as evidence of what a real customer, employee, or leader actually felt, believed, or lived through, it has quietly claimed a seat it was never eligible to hold — not because the technology is too immature, but because lived experience was never a variable it had access to in the first place. It was never in the training data as someone's experience. It was only ever a pattern of what such experiences have looked like when other people described them, averaged into something plausible.




Why This Is How HEARyou Was Built



This is not an abstract distinction for us. It is close to the founding premise of HEARyou.


Every study we run begins with the same question, asked to a real person: what is it like to be you in this experience? That question cannot be answered by a system rehearsing plausible responses to a demographic label, no matter how many transcripts it has absorbed or how convincingly it writes. It can only be answered by someone who was actually there — in the meeting that went wrong, the product that disappointed them, the moment that changed how they felt about their work.


AI has a real and valuable role in this — but it is the role of reaching more of the real thing, not replacing it. HEARyou uses AI to hold structured, consistent, interviewee-led conversations across hundreds of real people instead of a dozen, so that one research team can genuinely hear more voices without flattening any of them into an average. That is categorically different from generating a persona and asking it to speak on those people's behalf. One reaches the concert hall. The other rehearses alone at home and calls it the performance.


The practice room has real value — for questions, for hypotheses, for stress-testing an instrument before you ever use it on a real person. But no one has ever staged a concert in it, and no research finding worth acting on should be staged there either.



This piece draws on the ongoing research and internal discussions behind HEARyou and The Curious People Solutions' approach to human-centred listening.

Comments


bottom of page