Engineered, not prompted.
LLM roleplay isn't your customer. Here's how we build synthetic humans you can trust: grounded in your data, cited, and tested before they speak.
Ask an LLM to be your customer and it will improvise: fluently, confidently, and from stereotypes it never checked against your world. A synthetic human is the opposite of improvisation. It is engineered.
Everyone has met the fake customer. You ask an AI to roleplay your user and it answers in seconds: agreeable, articulate, and completely untethered from the people you actually serve. It's a costume, not a model. We wrote about why that fails in our first paper.
The wider field has landed on the same tension. Synthetic users are a genuine advance: teams use them to reach audiences at a scale and speed traditional recruitment can't match,1 to compress the cost and timeline of early research,2 and to pressure-test ideas before a line of code is written.
But the literature is unsentimental about how they fail. The Interaction Design Foundation questions whether an AI persona can stand in for a researched one at all.3 Peer-reviewed work documents a recurring habit of flattening people into stereotypes and asserting confident claims with nothing underneath.4,9 And the costume rarely fits: asked to speak for a demographic group, a model reflects that group's real opinions only loosely, and prompting it to adopt the persona barely closes the gap.8
Practitioners feel it too. Fluency was never the hard part. Trust is.
This paper is about closing that gap: the discipline we insist on at every step, so a synthetic is something you can trust, audit and rebuild. It comes down to one idea. A synthetic human is not generated in a single step. It is built, in two stages, the way you would engineer anything you intend to rely on.
Built, not bluffed.
Two stages, four phases. The first makes sense of your data; the second turns it into people. Here is what happens inside each.
Modelling the audience
Before a single persona exists, we build a grounded model of the audience hiding inside your data.
We start by reading your documents and extracting the population groups that actually live in them, not personas we'd like to find, but the segments your own evidence supports. The how-to literature shares one backbone: define the audience, then gather and model quality data;7 the discipline is in never letting a step run ungrounded. The payoff is measurable: in controlled studies, agents grounded in people's own words predict their attitudes and behaviour far more accurately than demographics alone, and they shrink the very group-to-group gaps that stereotype-driven models widen.10
Every extraction keeps a pointer back to where it came from. That grounding is what lets an audience be rebuilt or updated later: change the source, and the model can move with it.
From there we research public data (reviews, forums, news and published studies) on what your brand is known for and how your target audience actually behaves: socially, not just demographically. Each finding is grounded so it can be checked against your own content before it is used. Synthetic data is only as trustworthy as the quality assurance behind it,5 so nothing enters the model unchecked.
What comes out is not an impression of your audience. It is a structured model, and every layer of it is accountable. Generative persona work is faulted, again and again, for confident invention and cultural blind spots;4 grounding each layer in a named source is the structural answer to both.
The measurable surface of the audience: the things they worry about, how comfortable they are with technology, where they hesitate, how they prefer to be spoken to, and the factors that actually tip their decisions.
The heart of the model: multiple dimensions, each stated as a single finding with its evidence and a confidence level graded by how strong that evidence is. Each names where it came from (your documents and public data); where the evidence isn't there, we abstain rather than assert.
Research is always about a situation. We capture the pain points, the signals of comfort and the boundaries that only matter in the context you're actually putting to the test, so the audience is sharp where it counts.
The part most systems skip. Where the evidence is missing, we mark the gap explicitly rather than let the model paper over it. An honest "we don't know" is worth more than a confident invention,4 and it is exactly the move the research says most tools avoid.
Where the data is silent, so are we.
Generating the persona
With the audience modelled, we bring individuals to life: layer by layer, so coherence never breaks.
The first move is curation, not writing. Personas are built in layers so coherence holds, starting from defined behaviour (their role, their situation, the attitudes that define them and the triggers that move them) before any story is added.
Then we add the substance: history and lived detail, the texture that makes a persona read like a person rather than a profile.
Crucially, the finished persona sheet never floats free. It carries its grounding links all the way back to the archetype and audience it was drawn from, so any claim it makes can be traced to its origin.
Then, before any persona is allowed to speak, it has to pass.
A light first pass checks that the persona is internally consistent (that attitudes, history and behaviour belong to the same person) before any deeper work is spent on it.
Each persona is validated for alignment with the modelled audience through interactive testing across several turns of conversation. Only a persona that stays in character, and in evidence, is cleared to ship. This is the check the sceptics say is missing: a synthetic that is never tested against the audience it claims to represent is the one they warn about.3 And the bar is higher than a matching average: a synthetic can hit a group's mean and still miss its spread, so alignment is judged on the shape of the responses, not the headline number.12
The rigour behind our results, back-tested against human interviews → Nº 01: 9/10 teams make the same mistakes. How we keep them defensible → Trust.
Every answer, traced back to your data.
Where we agree with the sceptics
Discipline is not a claim of omniscience. On the limits, the research is right, and we designed for them rather than around them.
A synthetic human, however carefully engineered, is a complement to real human research, not a wholesale replacement for it. Lived emotion, cultural nuance and the genuinely unexpected still come from people.
What engineering buys you is a faster, cheaper, fully traceable first pass: a model good enough to sharpen the questions you take to real users, and honest enough to tell you where it is guessing.5 Done with that discipline, the results travel: independent work finds synthetic respondents reproducing human purchase-intent surveys at close to their test-retest reliability, each rating explained in the respondent's own words.11 Time is the constraint every research team feels, and a grounded synthetic is how you spend it on the right questions.
That is why abstentions are a feature, not an apology, and why the strongest programmes run synthetic and human research side by side.6 The literature's call for responsible, transparent use is one we engineered for from the first stage, not bolted on at the end.4
A sharper first pass, not the last word.
None of this is prompting. Prompting asks an LLM to imagine your customer in a single, unaccountable step. Engineering builds a model: grounded in your data, honest about its limits, validated before it speaks, and traceable after. That is the standard your research deserves.
It's time to meet your
synthetic humans.
See how we'd model your customers.
Book a call and we will run it on your data. See an audience modelled, cited and confidence-rated, then watch a synthetic human pass both validation gates before it answers a single question.
- Bain & Company. How Synthetic Customers Bring Companies Closer to the Real Ones.
- Making Science. Enhancing UX/UI Research with Synthetic Users.
- Interaction Design Foundation. Are AI-Generated Synthetic Users Replacing Personas?
- arXiv (2025). Creating and Evaluating Personas Using Generative AI.
- Snowflake. What Is Synthetic Data? Examples and Use Cases.
- AI EDAM, Cambridge University Press. Synthetic users: insights from designers' interactions with persona-based chatbots.
- MJV Innovation. How are AI models used to create synthetic users for research?
- arXiv (2023), Santurkar et al. Whose Opinions Do Language Models Reflect?
- arXiv (2023), Cheng, Piccardi & Yang. CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations.
- arXiv (2024), Park et al. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.
- arXiv (2025), Maier et al. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings.
- arXiv (2026), Moon et al. Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level.
Every figure and claim above is linked to its primary source so you can read the original. The methodology is our own.