A foundation model can role-play a user. That capability is now cheap, widely available, and improving for everyone at the same time. It is also not a research method.
The method starts where the training set ends. Your users' objections, their actual drop-off points, the way your best researcher frames a tradeoff, the outcome of the pricing change you shipped last quarter: none of that is sitting in a public corpus waiting to be retrieved. It lives in your transcripts, your tickets, your event trails, and your shipped history. A panel that cannot see those things is guessing from the average of the internet. A panel that can, and that is later scored against what really happened, is a different kind of object: a standing model of a specific market.
That is the asset. The model is how you query it this quarter.
What commoditizes, and what does not
Generic synthetic respondents commoditize for the same reason generic copy and generic code do. Anyone with an API key can prompt a model to "act like a mid-market operations manager." As the models improve, that prompt gets more fluent for every team at once. Fluency is not a moat. It is a shared input, the way cloud compute became a shared input for software.
What does not commoditize is a model of your users. Two properties make that asset hard to reconstruct from the outside.
First, the evidence is yours. Interview transcripts, support threads, session-level event trails, past A/B outcomes, and the product context those people actually see: these are not public training data. They are the residue of running a company. Interview-grounded agents reach roughly 83 to 86 percent of the human test-retest ceiling on core survey instruments; agents built from demographic attributes alone reach roughly 74 percent (Park et al., 2024). That gap is not a model-quality gap. It is an evidence gap, and the practical version of closing it is in our grounding guide.
Second, the score is yours. A simulation that is never checked against a shipped outcome is a story. A simulation tied back to what users actually did, through the same exposure and outcome events your product already emits, becomes a forecast with a track record. We score those forecasts with proper rules (Brier, log-loss) rather than a matched average, which can read healthy while the distribution underneath it has collapsed. That scoring history cannot be scraped, licensed, or prompted into existence. It is earned one shipped decision at a time.
A competitor can rent the same foundation model you rent. They cannot rent the last two years of your users talking, clicking, and converting.
The loop that produces both is short enough to state in a sentence: ground a panel in attributed evidence, run a real decision through it, read the distributions and the rationales rather than a point estimate, then bind the shipped outcome back so the next forecast starts from a slightly less wrong place. The foundation model sits inside one of those four steps. It will be swapped, versioned, and occasionally failed over. Teams that treat the model as the product will spend the next decade chasing whoever shipped the latest weights. Teams that treat the loop as the product will spend it accumulating something that gets more specific to their market with every use.
Why "just use a better model" is the wrong upgrade
When a simulation looks generic, the instinct is to upgrade the model. Sometimes that helps, the way a sharper lens helps a camera. It does not help if you are pointed at the wrong scene.
An ungrounded panel on a stronger model is still an ungrounded panel. It will be more articulate about the average internet user, more confident in its acquiescence, and no closer to the segment that actually churns on your billing page. The Nature results on LLM social-science prediction are real and useful (we wrote about what they do and do not establish), but they measure models against nationally representative survey experiments, not against your customers deciding whether to upgrade. Those are different instruments. Mixing them up is how teams end up wallpapering decisions with fluent, well-formatted, wrong answers.
The upgrades that actually move fidelity sit upstream and downstream of the model. More of the right evidence in. Scenarios that force a real choice instead of a compliment. A backtest against a decision whose outcome you already know. A held-out outcome you refuse to peek at until the simulation has spoken. None of that requires waiting for the next model release, and all of it compounds independently of which lab is currently ahead.
The loop tests options. It does not invent them.
There is a second thing the training set does not contain, and it is easy to miss while admiring the first.
AI is now very good at finding a cheaper path to a goal you have already named. Give it a conversion target, a latency budget, a "make this onboarding shorter" brief, and it will generate variants, rank them, and keep going after a human team would have stopped. That is optimization, and teams should take all of it they can get. It is not judgment. Optimization searches a space; judgment changes the space. The pricing frame nobody has tried, the empty state that refuses to apologize, the decision to kill a feature the dashboard still likes: there is nothing there to imitate, because the market has not seen the move. We argued in our thesis that abundance of generation relocates scarcity from building to knowing what is worth building. This is where that lands operationally.
A grounded panel is a testing instrument, and that is the point of it. It can tell you, directionally, how a specific population is likely to respond to a choice you put in front of it. It cannot tell you which choice to invent. Three consequences follow, and they are the easiest things to skip when the instrument is fast.
Scenario design is a human act. "Do you like this?" is not a decision. "This plan or that one, upgrade now or stay, variant A or the status quo" is. The quality of a simulation is usually settled before the first agent runs, in the tradeoffs the team was willing to put on the table. More studies fail from a flattering prompt than from a weak model.
The surprising option has to be in the set. Cheap screening is wasted if you pre-filter with your own taste. Include the variant you suspect is bad. Include the one that makes a stakeholder nervous. A panel cannot prefer a move it was never shown.
Rationales are where invention gets checked. Tallies tell you what the panel chose. Reasoning tells you whether the question was understood, and whether an objection you had not considered is load-bearing. That objection is often the seed of the next invention. Treating the report as a scoreboard throws the seed away.
Two ways a cheap loop goes wrong
The risk of a fast instrument is that teams reach for it even on the questions it cannot answer, because answering something is more comfortable than sitting with a question. Beyond the evidence theater we described in the new economics of design research, two failure modes show up specifically when optimization is mistaken for judgment.
Local maxima with better charts. The team only ever tests incremental variants of the current design. The panel, asked to choose among near-copies, will pick a winner. That winner can still be the wrong product. Optimization ran; judgment never started.
Dashboard autopsies dressed up as learning. Shipping, watching metrics, and calling the aggregate "research" is how teams did this before simulation existed. Simulation's job is to move some of that learning to before the commit. Used only to decorate a decision that has already been made, it is an expensive rubber stamp.
The corrective is a sequencing rule rather than a tool setting. Invent first. Put the invention in a forced choice with a real cost and a way to say no. Simulate. Read the rationales. If the stakes are high or the evidence is thin, confirm with humans, per the hybrid workflow. Bind the shipped outcome back. Then invent again, with a slightly sharper picture of the market.
Institutional knowledge, not tribal knowledge
Most companies already have a version of user understanding. It lives in the heads of the people who ran the last study, in a persona deck from two years ago, and in the Slack thread where someone said "our users would never do that." That knowledge is real and often excellent. It is also fragile. It walks out with the researcher, it goes stale the day the product changes, it cannot be queried by a new PM at 11pm, and it cannot be scored.
Static personas fail this test by construction: they are caches of answers to questions someone thought to ask, once. A grounded reasoning agent is a view over evidence you still hold, so new questions hit the evidence rather than the cache, and new transcripts update the view without a rewrite project. Because outcomes feed back, the institution can also tell when its picture of the user is drifting, instead of discovering the drift in a quarterly autopsy. Simulation on its own does not produce any of that. Grounding, provenance, and calibration do. Without them you have a faster way to generate the same slogans.
Where this argument stops
A closed loop is not a substitute for talking to humans, and it is not a substitute for having something worth testing.
If you have no transcripts, no tickets, and no behavioral data for a population, you do not have a loop. You have a prompt. Go collect evidence first. If the decision is irreversible, a pricing migration or a contractual change or a brand bet you cannot unwind, simulation can narrow the field but should not close it. Emotionally charged decisions, the ones involving health, money under stress, or safety, have the least validated affective fidelity in this field. And a genuinely new interaction pattern has no behavioral prior to calibrate against; the honest instrument there is a prototype in front of a person.
None of that protects intuition from evidence. It only says that evidence cannot originate the option it is asked to evaluate.
Used in its proper place, the split is simple. People invent. The panel tests. Reality scores the test. The company keeps the record. The models will keep getting better at the middle step, and they should. The first step and the last are still yours.