Ask an ungrounded language model to "act like a mid-market operations manager evaluating your pricing page" and you will get something fluent, confident, and largely useless. Not because the model is bad at pretending, but because it is pretending from the average of the internet rather than from anything true about your users.
The single biggest lever on simulation quality is not the model, the prompt, or the panel size. It is grounding: how much of your own evidence each synthetic respondent is built from. This post is a practical guide to doing that well, and to knowing when you have not done it well enough to trust the output.
Why ungrounded personas fail
An ungrounded persona fails in two characteristic ways.
First, it regresses to the mean. Without specific evidence, the model produces the most statistically typical answer for the demographic sketch you gave it. Every "operations manager" sounds like the same operations manager. You lose exactly the thing research exists to find: the variance, the objections, the segment that behaves differently from the others.
Second, it agrees with you. Language models are trained to be helpful, and helpfulness leaks into role-play as acquiescence: synthetic respondents that like your design, understand your copy, and would definitely upgrade. If your synthetic panel loves everything you show it, you do not have a panel. You have a mirror.
Both failures shrink as grounding grows. This is not just our observation; it is one of the better-replicated findings in the field.
What the research says
The strongest evaluations in this field share a design: build an agent for a real person from rich qualitative evidence, then test whether the agent answers the way its human counterpart does on established survey instruments taken later.
The number to remember is not the raw accuracy but the ceiling it is measured against. Humans do not perfectly agree with themselves on a retest weeks later, so the honest benchmark is the human test-retest rate. Against that ceiling, interview-grounded agents consistently close most of the distance, while agents built from demographic attributes alone lag well behind.
That gap, interview-grounded versus demographics-only, is the entire game. It is why we treat grounding as the core discipline of synthetic research rather than an optional enrichment step, and it is why no responsible vendor should quote a fidelity number without telling you what the panel was grounded in.
The evidence hierarchy
Not all evidence moves fidelity equally. In practice we see a fairly stable ordering:
- Interview transcripts. The richest single source. Real people reasoning out loud about their goals, constraints, and frustrations, in their own words. Even five to ten transcripts materially change how a panel reasons, because they carry the tone and the objections that demographic sketches cannot.
- Support and qualitative streams. Tickets, Intercom or Zendesk conversations, Dovetail highlights, sales call notes. Less structured than interviews but abundant, current, and skewed toward friction, which is precisely what acquiescent agents underweight.
- Behavioral event data. Product analytics from tools like Amplitude, Mixpanel, or PostHog, or your own SDK events. Behavior grounds what people do rather than what they say, and the gap between the two is often the finding.
- Surveys and past experiments. CSV survey exports and historical A/B results give you distributions to anchor against, and, crucially, held-out outcomes to validate against later.
- Product context. Figma frames, live URLs, docs, and repos. This grounds the stimulus rather than the respondent: agents that can actually see your onboarding flow give you friction points, not generalities.
The hierarchy is a prioritization aid, not a shopping list. A panel grounded in ten transcripts and a support export will usually beat one grounded in a beautifully formatted persona document, because persona documents are already someone's interpretation. Feed the panel the raw material instead.
How grounding works in Sentia
When you upload research artifacts or connect a source, we do not paste documents into a prompt. Artifacts are parsed and broken into attributed evidence units, each linked back to its source, and each agent draws on the evidence relevant to the decision it is currently reasoning about.
Two properties of this design matter for trust:
- Provenance is explicit. Every dimension of a panel is labeled by where it came from: your evidence, the recruiter's inference, or an assumption. An absent field shows up as a surfaced gap, never a silently invented default. When a simulation report cites a rationale, you can trace it to the transcript or ticket behind it.
- Gaps stay visible. If nothing in your evidence speaks to, say, price sensitivity in a segment, the honest output is "we do not have grounding for this," not a confident guess. A panel that admits ignorance in the right places is worth far more than one that answers everything.
A practical grounding playbook
If you are setting up your first grounded panel, do it in this order:
- Start with transcripts, even a few. Five to ten interviews or sales calls covering your two or three most important segments. Rough transcripts are fine; verbatim tone is the value.
- Connect one qualitative stream. Your support tool or Dovetail workspace. This keeps the panel current without manual uploads and biases it toward real friction.
- Connect analytics. Session-level event trails give agents behavioral anchors: where users actually drop, rage-click, and return.
- Upload the quantitative history you want to predict against. Past surveys and A/B outcomes. Hold some of this back; you will want it in step 5.
- Backtest before you trust. Re-run a decision you already shipped and compare the simulated distribution to what really happened. Not just the average: check whether the spread and the segment differences look like reality. A simulation that matches the mean while every synthetic respondent gives the same answer has collapsed into an average persona, and its next prediction should not be trusted.
Design scenarios that let agents disagree
Grounding reduces sycophancy; scenario design finishes the job. Three rules of thumb:
- Never ask "do you like this?" Ask the panel to choose between real alternatives with real costs: this plan or that one, upgrade now or churn, variant A or variant B against your current design.
- Give every scenario a way to say no. Include the status quo as an option. The most informative simulated outcome is often "most of this segment would do nothing."
- Read the rationales, not just the tallies. The distribution tells you what; the reasoning tells you whether the panel understood the question. Rationales that quote your evidence are a good sign. Rationales that could apply to any product are a grounding gap.
Where this method stops
Grounded simulation is a directional instrument, and some decisions should not rest on it alone:
- Populations you have no evidence for. If you are entering a market where you have zero transcripts, tickets, or behavioral data, a synthetic panel for that market is speculation with good formatting. Talk to real people first, then ground a panel in what they said.
- High-stakes, irreversible calls. Pricing migrations, brand changes, anything contractual. Use simulation to narrow the options, then validate the finalist with real users.
- Emotionally charged or sensitive domains. Health, money under stress, safety. Simulated affect is the least validated part of this field; treat emotional signals as hypotheses only.
- Genuinely novel interactions. When no one has ever used anything like the thing you are testing, there is no behavioral prior to ground in. That is what prototypes and humans are for.
The workflow we recommend, and use ourselves, is hybrid: synthetic panels for the fast directional 80 percent of decisions, real research concentrated on the 20 percent where the stakes or the novelty demand it. Simulation does not replace that research. It makes sure that when you do spend a six-week study, you spend it on the question that deserves it.
If you are wiring up your first panel and want a second pair of eyes on your grounding setup, talk to us. We would rather help you build a panel you can distrust correctly than have you trust one you should not.