A small product team needs to test a new appointment-booking service. Recruiting participants will take days; asking a language model to “act as” an older patient, a busy parent or a person with limited vision takes seconds. Within minutes, the team has plausible frustrations, quotations and feature requests. The output looks remarkably like research.
That resemblance creates the danger. An AI persona can be useful for widening a team’s questions, rehearsing a study and exposing obvious assumptions. It cannot, by itself, provide evidence about what a particular group experiences. The right role is not synthetic participant but hypothesis generator: a fast instrument whose predictions must be compared with situated human observation before they influence consequential design choices.

A persona is compressed evidence
Design personas existed long before generative AI. At their best, they compress observations from interviews, fieldwork, service data and support contacts into a memorable account of goals, constraints and behaviour. Their authority comes from the evidence beneath the compression. An invented name and portrait are not the research; they are an interface to research that has already happened.
A language model reverses that sequence. It can generate the polished interface first: a coherent biography, likely concerns and convincing phrases. It does so by predicting language from patterns in its training data and the prompt, not by observing the proposed service, waiting for a screen reader to announce a badly labelled button or discovering that a family shares one phone. The model produces what a situation sounds like from the outside.
This does not make the output random. It may contain valuable regularities learned from many descriptions of people and interfaces. But plausibility is not provenance. A detail can sound socially perceptive while being a stereotype, an average that fits nobody, or a projection of the designer’s own prompt.
Believability can conceal the bias
A 2024 CHI study examined 450 AI-generated persona descriptions with internal evaluators and subject-matter experts. The personas were often judged believable, informative and relatable, yet the researchers also found biases in age, occupation and stated pain points, together with a strong orientation towards the United States. Experts saw somewhat more stereotyping and less relatability than internal evaluators did. The important finding is the combination: biased output need not look crude or obviously false. It can arrive in a polished, sympathetic form (Salminen and colleagues, CHI 2024).
A 2026 ethical audit presented at AAAI compared 1,512 personas generated by three model families with human-authored responses. It found that the models disproportionately foregrounded racial markers and produced identities that were linguistically elaborate but narratively reductive. The authors describe a form of “algorithmic othering”: minority identity becomes highly visible while the person becomes less particular (AAAI 2026). Adding demographic detail to a prompt can therefore intensify the very simplification it was meant to correct.
The strongest case for simulation
The counterargument deserves more than a warning label. Human research costs time and money. Some teams have no researcher, tiny budgets and a deadline before participants can be recruited. Early ideas are numerous, and many contain predictable usability failures. If an AI agent can uncover a large share of those problems cheaply, teams can repair them before asking people to spend time on an immature prototype.
There is encouraging evidence. A paired study published on 19 March 2026 tested ten students and matched AI “twins” on the same educational web application. The researchers reported substantial overlap in qualitative themes and similarities in navigation behaviour. Yet the agents consistently underestimated cognitive load and emotional frustration. The sample and application were narrow, so the result is not a general validation of synthetic users. It shows something more useful: simulation may reproduce visible task patterns while missing the intensity and meaning of the human experience (Jaiswal and colleagues, 2026).
This suggests a division of labour. Use synthetic users to explore the possibility space: generate edge cases, challenge instructions, rehearse interview questions and predict where a task might fail. Use real people to establish whether those predictions occur, what they cost, which workarounds people invent and what the team failed to imagine.
The simulation must not erase participation
Replacing participants is not only an accuracy problem. It changes who has standing in the design process. A person invited into research can reject the team’s categories, introduce a need nobody anticipated and explain why an apparently minor obstacle is consequential. A generated persona cannot consent, contest its representation or demand that the project change. It has no stake in the outcome.
The distributional effect matters. The groups described as “hard to recruit” are often those whose circumstances are least well represented in standard datasets and most easily flattened into demographic shorthand: people with disabilities, uncommon language needs, precarious work, shared devices or distrust grounded in earlier institutional harm. Simulation is most tempting precisely where substitution is least defensible.
There are also privacy limits on the apparent alternative. A team should not construct a synthetic twin by feeding a model identifiable interview transcripts, health details or customer records without a valid purpose, authority and protection. De-identification can be fragile when narratives contain distinctive combinations of events. The need for safer research operations is not permission to turn personal histories into prompts.
Run a synthetic-to-situated gap test
A team can test the value of synthetic users without pretending they are people. Before conducting a small real usability study, give the AI the public product description, the task and a deliberately limited set of contextual assumptions. Ask it to predict five specific failure points, the observable behaviour associated with each and the evidence that would disconfirm the prediction. Do not ask it to invent quotations or inner feelings.
Then observe three to five consenting participants completing the same bounded task, using an accessible method appropriate to the service. Record task failures, hesitation, workarounds, requests for help and participant explanations. Keep the synthetic predictions separate until the observations are coded, following the same principle Alkemata proposed for preserving independent views in human–AI disagreement.
Compare the two sets in a simple ledger: predicted and observed; predicted but not observed; observed but not predicted; and impossible to judge. Add two columns that matter more than raw overlap: severity for the person and design action. A missed obstacle that blocks one participant may matter more than four correctly predicted cosmetic irritations. An invented issue can also waste effort if a team mistakes fluent simulation for demand.
The experiment does not validate the model for every population or future product. It calibrates one use in one context. If predictions repeatedly help the team prepare better questions, keep the tool in the exploratory stage. If they narrow attention, reproduce stereotypes or miss consequential barriers, change the prompts, restrict the role or stop using it.
The meaningful next decision is therefore not whether synthetic users are “good” or “bad”. It is whether a team can name what evidence the simulation is allowed to provide. Let AI generate possibilities. Let people supply experience, contradiction and stakes. Only the second can turn a plausible persona into accountable design knowledge.