Skip to content
How-to

How to evaluate a synthetic user research platform

Ten questions that separate a grounded simulation platform from a prompt with a dashboard, why each one matters, and what a good answer sounds like. Includes the questions we would want asked of us.

6 min readSentia Labs

Every vendor in this category will tell you their agents are accurate. The claim is unfalsifiable as stated, so it carries no information, and buyers end up choosing on demo polish.

These are the ten questions we would ask on an evaluation call. They are ordered by how much they separate serious implementations from thin ones. We have included what a good answer sounds like, and where we think we answer well and where we do not, because a buyer's guide written by a vendor is worthless if it only asks questions that vendor happens to win.

1. What grounds the agents, and can I see the citation?

This is the question. Everything else is secondary.

Agents built from demographics alone reach roughly 74 percent of the human test-retest ceiling on core survey instruments. Agents built from interview evidence reach roughly 83 to 86 percent (Park et al., 2024). Whether a platform can ingest your transcripts, tickets, survey responses, and event trails is therefore not a feature comparison, it is the main variable in output quality.

A good answer names the specific sources it ingests and shows you a single agent attribute traced back to the document it came from.

A weak answer talks about model quality, parameter counts, or which foundation model is under the hood.

2. What happens when there is no evidence for a field?

Ask them to show you an under-specified population and watch what the system does.

The dangerous behaviour is silent imputation. If a platform fills an unknown attribute with a plausible-looking default and does not tell you, every downstream number inherits a fabrication you cannot see. A loud gap is a working product. A quiet default is a lie that reaches a decision.

A good answer shows you a visible gap, a provenance label, or a refusal.

3. How do you know it is right, and how would I know?

There is a difference between a vendor validating their model and your workspace having a track record.

Published benchmarks are worth something. In Nature, GPT-4 simulations predicted 70 preregistered experiments at r = 0.85, and r = 0.90 on unpublished studies, and we wrote about the limits of that result. It measures models against survey experiments, not against your customers.

So ask the second half: after I ship something, does the platform score the prediction that preceded it?

A good answer describes a closed loop, usually events fired from your product when a change is exposed to a user and when they act, and a hit rate you can inspect per workspace.

A weak answer is a single accuracy percentage with no denominator.

4. Can I backtest it on a decision I already know the outcome of?

The cheapest way to calibrate your own trust is to run a study on something you shipped last quarter, without telling the system what happened.

If a vendor resists this, ask why. It is the one test that cannot be gamed by demo selection.

5. Does it return a distribution or a verdict?

A single number is easy to present and easy to over-trust. What you want is the spread, the segments that disagree, and the reasoning behind the split.

The most valuable output of a synthetic study is frequently not the tally but one objection you had not considered, surfaced by a segment you were not focused on. A platform that compresses everything to a recommendation has thrown that away. We walk through what a full report should contain in anatomy of a simulation report.

6. What can it not do, and will they tell me unprompted?

Ask directly: what questions should I not use this for?

A vendor who cannot answer has either not thought about it or will not say. The real boundaries in this category are reasonably well understood. Genuinely novel interaction patterns have no behavioural prior to calibrate against. Emotionally loaded decisions involving health or money under stress have the least validated affective fidelity. Irreversible commitments should be narrowed by simulation and closed by humans.

Any vendor claiming their product replaces user research entirely is telling you something useful about the vendor.

7. How does the population get built, and from what data?

"Representative" is doing a lot of unexamined work in most demos.

Ask which population source shapes the distribution. Published census microdata and equivalents produce a population with the real shape of a market. A model asked to imagine a representative sample produces the internet's stereotype of one, and it will be visibly wrong in ways that are hard to catch until a segment behaves implausibly.

Also ask what happens for a market outside the vendor's home country. This is where thin implementations break first.

8. Can an agent drive it, or does it need a human in a dashboard?

This one is newer, and it is rapidly becoming the difference between a tool that gets used and one that gets forgotten.

If your team works through coding agents, a research platform reachable only through a web UI is a context switch that will quietly stop happening. Ask whether there is a REST API, a CLI, and an MCP server, and specifically whether they expose the same capabilities as the app or a reduced subset. A read-only API on top of a UI-first product is not the same thing.

9. Who owns the evidence, and what happens when I leave?

You are handing over transcripts and behavioural data about real people. Ask where it lives, whether it trains anything shared across customers, how deletion propagates across every store, and what you can export.

The answer "your data never informs another customer's simulations" should be a stated guarantee, not an inference you make from a marketing page.

10. What does it cost when it works?

The trap is per-seat pricing on a tool whose value comes from asking more questions. If every extra study costs a negotiation, teams will ration the instrument and never reach the volume that makes it worth having.

Ask what happens at ten times your expected usage.

Where we land on our own questions

We would rather you ask us these than take our word for it, so here is our honest scorecard.

Where we think we answer well. Grounding is the centre of Sentia: agents are built from your transcripts, tickets, survey responses, and event trails, every attribute carries provenance, and a field with no evidence behind it is surfaced as a gap rather than defaulted. Populations are sampled from census microdata and WorldPop rather than imagined. Predictions are scored against shipped outcomes through exposure and outcome events, so accuracy is measured per workspace rather than asserted. And the whole platform is agent-native by construction, with the same capability plane behind the app, the REST API, the sentia CLI, and an MCP server.

Where we are still earning it. A scored track record is only as good as its length, and a hit rate means little in your first month. Backtesting an old decision is the faster way to calibrate trust, which is why we point new teams at it in your first week with a synthetic panel. We also hold the position that this does not replace talking to users, which means we will sometimes tell you the answer is to go run a real study.

The question we would most want asked. Number four. Take a decision you shipped, one where you know the outcome and we do not, and run it. That single test will tell you more than any demo, including ours.

For the definitional groundwork, start with what is synthetic user research. For where this sits against your existing methods, see synthetic users vs usability testing vs user interviews.