Skip to content
Perspectives

Synthetic users vs usability testing vs user interviews

The three ways a product team gets an answer about a design, what each one actually measures, what each costs in calendar time, and a rule for picking between them on the decision in front of you.

6 min readSentia Labs

You have a prototype and a question about it. There are three practical ways to get an answer.

Sit down with users and watch them use it. Push it to an unmoderated testing platform and collect sessions. Or simulate the population and read how the simulation behaves.

They are not ranked, and the most expensive mistake is not picking the wrong one. It is assuming they measure the same thing.

Moderated interviewsUnmoderated testingSynthetic users
What it measuresObserved behaviour, plus the whyObserved behaviour, at volumeModelled behaviour, from prior evidence
Calendar timeTwo to six weeksTwo to five daysSame day
Typical n5 to 1220 to 100Hundreds to thousands
Cost of one more questionA whole new roundA new test and new incentivesNear zero
Needs a working buildA clickable prototypeA clickable prototypeA prototype, a URL, or a static frame
Reaches niche segmentsSlowly and expensivelyOnly if the panel has themYes, if you hold evidence about them
Catches the thing nobody predictedBest in classSometimesOnly via rationales, if you read them
Wrong in a way you will noticeIn the room, immediatelyIn the session replayOnly if you score it

What each one is genuinely best at

Moderated interviews are for understanding why. Nothing substitutes for watching a person fail to complete a task you were certain was obvious, then asking what they expected. This is where you learn that your mental model and theirs disagree, which is a class of finding no other method reliably produces. The constraint is arithmetic: at two to six weeks a study, most teams run a handful a year and decide everything else on opinion.

Unmoderated platforms are for volume on a settled question. When you know exactly what to measure, task completion, time on task, where people click first, an unmoderated tool gets you a real sample of real humans within the week. They are weaker when the question is exploratory, because you cannot follow up, and weaker again when you need a niche segment the panel does not carry.

Synthetic users are for screening. They are at their best when you have more options than you can afford to build. Six onboarding flows, four pricing frames, three empty states: put all of them in front of a grounded population, kill the obvious losers, and spend your scarce human research on the two survivors. That is the hybrid workflow, and we think it is the correct shape for a modern research practice.

The trap in each method

Every method has a characteristic way of producing a confident wrong answer, and knowing yours is more useful than knowing its strengths.

Interviews over-fit to five people. A vivid session from an articulate participant will outweigh a boring pattern from the other four, because stories are stickier than counts. Small n plus high salience is how teams end up redesigning around one person's edge case.

Unmoderated tests measure the task you wrote. Participants complete the task they were assigned, in a context where they have no stakes and no history with your product. High completion rates on an artificial task tell you less than they appear to.

Synthetic studies inherit their grounding. An agent built from your transcripts, tickets, and event trails models your users. An agent built from a prompt models the internet's idea of a person, and it will be fluent, agreeable, and unmoored. This is the trap that matters most, because the output looks identical either way. It is why we treat grounding as the whole game, and why we wrote a guide to it.

A rule for picking

Two questions resolve most cases.

Is the decision reversible? If you can ship it, watch it, and roll it back next week, synthetic evidence is usually enough to choose. If you cannot, and a pricing migration, a navigation overhaul, or anything with a migration path usually cannot, narrow the field with simulation and close it with people.

Do you hold evidence about this population? If you have transcripts, tickets, and behavioural data for the segment, a synthetic study will be grounded and worth acting on. If you do not, you do not have a synthetic panel, you have a prompt with a costume on. Go collect evidence, or recruit humans.

Reversible and well-evidenced: simulate and move. Irreversible or thinly evidenced: simulate to narrow, then confirm with people.

There is a third question worth asking when the interaction pattern is genuinely new. If nothing like it exists in the world, there is no behavioural prior for a simulation to calibrate against, and the honest instrument is a prototype in front of a person. Simulation is strongest on decisions that rhyme with decisions users have made before.

What actually changes when screening gets cheap

The usual pitch for this category is cost, and the framing is normally too aggressive. It is true that the marginal cost of one more question falls close to zero, and that genuinely changes what a team can afford to ask. We wrote about the consequences in the new economics of design research.

What is not true is that the savings are free. A synthetic study displaces cost rather than deleting it. You spend less on recruiting, scheduling, and incentives, and more on grounding, scenario design, and checking the result. Teams who treat it as a cheaper unmoderated test get a cheaper unmoderated test.

The failure mode to watch for is evidence theatre. When a study costs an hour, it becomes tempting to run one after the decision is made and use it as a rubber stamp. That is not fast research. It is an expensive habit with a research-shaped receipt.

Where we sit

We built Sentia for the screening role specifically, and three design choices follow.

Agents walk the actual artifact. A Figma prototype, a live URL, or a set of frames, with each agent moving through the flow rather than being asked to imagine it, and every run returns System Usability Scale scoring and accessibility findings alongside the verdict. Screening should produce the same shape of output as the study it is standing in for.

Grounding is not optional. A population is assembled from evidence you already own, each attribute traced to a citation, and gaps surfaced rather than filled with a plausible default. Census microdata and WorldPop shape the distribution so a panel reflects a real market rather than one flattened persona.

Results get scored. Our SDK emits exposure and outcome events, so when a change reaches real users the result grades the simulation that predicted it. Over time you get a measured hit rate for your own workspace instead of a vendor accuracy claim.

We are not trying to replace your research practice. The end state we want is a team running ten times as many studies, most of them synthetic, spending its human research budget on the questions that deserve a person in the room.

For the definitional groundwork, start with what is synthetic user research. If you are comparing vendors, we wrote the questions we would ask in how to evaluate a synthetic user research platform.