Skip to content
How-to

The hybrid research workflow: synthetic for direction, humans for confirmation

A working split for teams adopting simulated research. Which decisions can move on synthetic evidence alone, which need real users, and how to run the loop between them.

5 min readSentia Labs

The least useful question in this field is "synthetic or real?" Teams that frame it that way end up in one of two failure modes: they distrust simulation entirely and stay stuck at the old research tempo, or they trust it entirely and stop talking to the humans their business depends on.

The productive question is an allocation question. You have two instruments with different cost curves and different error profiles. Which decisions go to which instrument, and how do the two feed each other? This post is our working answer, the one we use ourselves and recommend to every team we onboard.

What each instrument is actually good at

A grounded synthetic panel is strong at breadth and speed: screening thirty options down to three, rank-ordering alternatives, surfacing friction points and objections you had not considered, and explaining its reasoning at a per-respondent level you could never afford to collect at scale. It is weak exactly where you would expect a model to be weak: absolute point estimates ("precisely 12.4 percent will convert"), domains where your evidence is thin, and emotionally loaded decisions where simulated affect is the least validated part of the field.

Research with real humans has the inverse profile. It is slow and expensive per decision, which means it cannot cover your decision volume. But it is the only ground truth there is. Nothing about a simulation is trustworthy except insofar as it has been, or can be, checked against real behavior.

The hybrid workflow is just the disciplined combination: synthetic evidence for the fast directional 80 percent of decisions, recruited research concentrated on the 20 percent where stakes or novelty demand it, and a feedback loop that makes the synthetic side more trustworthy every time the real world weighs in.

A decision matrix

Two questions place almost any product decision in the right quadrant. First, how expensive is being wrong? A reversible change behind a flag is cheap to be wrong about; a pricing migration or a public brand change is not. Second, how well does your evidence cover this population and this kind of question? A panel grounded in hundreds of relevant transcripts and event trails covers a checkout redesign well; the same panel covers a brand-new market segment badly.

Evidence-richEvidence-thin
Cheap to be wrongSimulate and ship. Watch the metrics you already have.Simulate, ship behind a flag, and treat the launch itself as the study.
Expensive to be wrongSimulate to narrow the field, then validate the finalist with real users.Real research first. Ground a panel in what you learn, then simulate.

The bottom-right quadrant deserves emphasis because it is where overconfident teams get hurt. If the decision is high-stakes and your evidence is thin, simulation is not a shortcut. It is speculation with good formatting, and the honest move is to go collect evidence from humans first. We wrote about what "thin" means, and how to fix it, in our guide to grounding.

The loop in practice

For a decision that warrants the full workflow, it looks like this:

  1. Frame the decision as a choice, not a rating. "Which of these three onboarding flows gets a new workspace to its first simulation fastest, and where does each one lose people?" beats "do users like our onboarding?" Forced choices with real alternatives are also your main defense against agreeable-by-default respondents.
  2. Screen wide with the panel. Run every variant you have, including the ones you suspect are bad. Cheap screening is the whole point; do not pre-filter with your own taste, because your taste is exactly what you are trying to check.
  3. Read the rationales, not just the tallies. The distribution tells you what the panel chose. The reasoning tells you whether it understood the question, and it is where the surprises live: the objection nobody on the team raised, the segment that reads your pricing page completely differently.
  4. Narrow to one or two finalists. At this point you have spent hours, not weeks, and you have a ranked field with explanations.
  5. Validate the finalist with real users when the quadrant calls for it. Five recruited sessions on one finalist is a fundamentally different spend than five sessions across thirty candidates. This is where the saved budget goes, not into a drawer.
  6. Ship, and tie the outcome back. Record what actually happened and compare it to what the simulation predicted. This is the step most teams skip, and it is the one that makes step 2 more trustworthy next quarter. In Sentia this loop is instrumented: shipped outcomes flow back against the simulations that predicted them, so accuracy is a measured number, not a vibe.

What "directional" actually means

When we say synthetic evidence is directional, we mean something specific about which of its outputs to trust, in what order:

  • Rank order is the most robust output. "Variant B beats variant A for this segment" survives model imperfection far better than any absolute number.
  • Distribution shape and segment differences come next. "Enterprise admins are split, individual contributors are enthusiastic" is a real finding worth acting on.
  • Absolute point estimates are the least trustworthy output and should be treated as hypotheses until your own backtests show they calibrate. A predicted 34 percent conversion is not a forecast you bank; it is a prior you check.

A simulation report that changes which option you ship, or which objection you fix before launch, has done its job even if every number in it is off by a factor.

Where this method stops

The hybrid split assumes your synthetic instrument has been honestly assessed. Three situations void that assumption:

  • You have never backtested. If you have not once compared a simulated distribution to a realized outcome, you do not yet know what your panel's directional accuracy is. Run one retrospective study on a decision you already shipped before you lean on prospective ones.
  • The panel has drifted from the product. Evidence ages. A panel grounded in last year's users can be confidently wrong about this year's. Recurring syncs from live sources exist precisely to shrink this gap.
  • The decision is really a values question. Some choices, in accessibility, pricing fairness, or data handling, are not empirical questions about what users will do. No research instrument, synthetic or human, answers what you should stand for.

Used inside those limits, the hybrid workflow is not a compromise between speed and rigor. It is how you get both: simulation makes research cheap enough to apply everywhere, and real research makes simulation honest enough to trust. If you are working out where the split should sit for your team, talk to us.