Skip to content
Research notes

Synthetic respondents just passed a real test

In Nature, GPT-4 simulations predicted the results of 70 preregistered experiments at r = 0.85, and r = 0.90 on unpublished studies. What this validates, what it does not, and what to do with it.

4 min readSentia Labs

Skepticism about synthetic research is healthy. The idea that a language model can stand in for your users, even directionally, is the kind of claim that deserves a hard test rather than a demo. Over the past year, the field got two of the hardest tests it has ever had, both published in Nature. This post is our reading of what they establish, what they explicitly do not, and what a product team should change because of them.

The study: 476 real experimental effects

The first result is Hewitt, Ashokkumar, Ghezae, and Willer, "Large language models can predict the results of social science experiments". The design is unusually clean. The team assembled an archive of 70 preregistered, nationally representative survey experiments run in the United States: 476 separate treatment effects, measured on more than 100,000 real participants. Then they asked GPT-4 to role-play representative samples of Americans responding to the same stimuli, and compared the simulated effects to the real ones.

The simulated predictions correlated with the actual experimental results at r = 0.85. For context, that equaled or exceeded the accuracy of expert human forecasters asked to predict the same outcomes. The result also replicated on nine large megastudies covering 346 additional treatment effects.

The control that matters most

The obvious objection to any result like this is memorization: maybe the model saw these published studies during training and is reciting, not simulating. The authors anticipated it. On the subset of studies that were unpublished at the time, and therefore could not appear in any training data, accuracy did not degrade. It rose, to r = 0.90.

That control is why this paper moves the debate. It is direct evidence that a language model prompted with a population and a stimulus is doing something predictive about human responses, not something retrospective about its training corpus.

What the paper does not validate

The honest reading has three sharp edges, and they matter as much as the headline.

  • Magnitudes were wrong even when directions were right. As coverage of the paper emphasized, GPT-4 systematically overestimated effect sizes, by roughly a factor of two. The model reliably told you which intervention would work better; it was not a trustworthy source for how much better.
  • Prediction is not understanding. The authors and commentators are explicit that a system that forecasts responses well is not thereby a system that understands people, and that treating fluent simulated respondents as if they were real opinion carries real risks, including to minority perspectives that models are known to underrepresent.
  • The domain was survey experiments. Nationally representative message and treatment experiments are close cousins of copy tests and concept tests, but they are not the whole of product research. Extrapolating to every research task would be exactly the overclaiming this field needs less of.

The second result: grounding in behavior

The other landmark is Binz and colleagues' Centaur, published in Nature in 2025. The team fine-tuned a language model on Psych-101, a dataset of trial-by-trial decisions from more than 60,000 participants making over 10 million choices across 160 classic behavioral experiments. The resulting model predicted the behavior of held-out participants better than the specialized cognitive models psychologists have refined for decades, and it generalized to modified tasks and entirely new domains.

Read together, the two papers make complementary points. Hewitt et al. show that simulation of human responses is predictive out of the box, at the level of directional effects. Centaur shows that the ceiling rises when the model is grounded in real behavioral data rather than left to its priors. That second point is the academic version of something we see constantly in practice: grounding is the lever. A model tuned on what people actually did beats a model guessing from a description of who they are.

What a product team should do with this

If you take both papers seriously, three practices follow directly:

  1. Use simulation for direction, and trust it in this order: rank order first, distribution shape and segment differences second, absolute point estimates last. The r = 0.85 result and the 2x magnitude error are two facts about the same system. A simulated 40 percent lift means "this variant likely wins," not "expect 40 percent."
  2. Ground the simulation in your own evidence. Centaur's gain came from 10 million real choices. Your equivalent is transcripts, support conversations, and behavioral event trails. An ungrounded panel is running on priors; a grounded one is running on your users.
  3. Recalibrate magnitudes against realized outcomes. A consistent 2x overestimate is not a dead end; it is a correction waiting to be learned. Any serious deployment of synthetic research should compare its predictions to what actually shipped and happened, and adjust the next forecast accordingly. This is the discipline that separates an instrument from a party trick.

Our position, stated plainly

Neither paper tested Sentia, or any commercial platform. They validate a method class, not a product, and anyone in this market who blurs that line should lose your trust. What the papers do establish is that the method class is real: simulated respondents predict directional outcomes at accuracy comparable to expert forecasters, memorization does not explain it, and behavioral grounding raises the ceiling.

Our platform is built on the assumption that the caveats in these papers are design requirements. Directional trust ordering is why our reports lead with distributions and rationales rather than a single number. Grounding is the first thing we ask about any panel. And magnitude recalibration against shipped outcomes is the flywheel the whole system turns on.

The evidence now supports a specific, bounded claim: simulation is a legitimate directional instrument for decisions about human responses, and it gets better when it is grounded and calibrated. That is not "you never need real users again." It is something more useful: a reason to put evidence behind the eighty percent of decisions that were never going to get a recruited study, and to spend your real research where it counts.