1 June 2026·9 min read

The Surprising Boundary Between Psychology and Culture

In June 2026 we back-tested synthetic respondents against World Values Survey populations in the US and GB. The variance came out narrower than in real populations, and one finding redrew how we separate psychology from culture.

What a back-test against published population data taught us about synthetic respondents.

For the past year we've been building synthetic respondents.

Not synthetic survey answers. Not AI summaries. Not personas.

Synthetic people — with different motivations, different priorities, different ways of interpreting the world, and different psychographic structures.

And the question we hear more than any other is exactly the right one:

How do you know they're not just averages produced by a language model?

It's a fair challenge. In fact, it's probably the most important challenge facing synthetic research today.

So we decided to test it. Not with a demo. Not with a client case study. Not with a marketing benchmark. With public data.

And the results taught us something we weren't expecting.

The test

We used the World Values Survey (WVS), Wave 7 — one of the most respected datasets in social science, as the reference.

We generated synthetic populations for the United States and Great Britain, using demographic quotas aligned with national population structures. We then compared their responses against the real WVS distributions across five non-sensitive attitudes:

  • Interpersonal trust
  • Life satisfaction
  • Happiness
  • Attitudes toward competition
  • Importance of work

Our goal was not simply to reproduce averages. A model can land an average by accident. What matters is whether it reproduces variation.

Human populations disagree. They contain minorities, contradictions, and extremes. If synthetic populations collapse toward the center, they may look realistic while missing what actually makes people different.

So we evaluated four things: distributional convergence, mean alignment, variance preservation, and the differences between the two countries.

The numbers

Here are the actual results. No cherry-picking. No selective reporting. No hidden failures.

United States

ItemSyntheticHumanConvergenceVariance ratio
Trust1.721.6390.5%0.87
Work importance1.571.8883.2%0.37
Happiness2.221.8471.6%0.48
Competition4.073.3150.0%0.15
Life satisfaction7.027.2838.1%0.06

Great Britain

ItemSyntheticHumanConvergenceVariance ratio
Trust1.721.5481.6%0.80
Happiness2.181.7867.4%0.38
Competition4.253.7862.9%0.23
Work importance1.852.0559.1%0.18
Life satisfaction6.957.3447.1%0.13

How to read this

Convergence is the overlap between the synthetic and human answer distributions, reported as 1 minus total variation distance (higher is better). Variance ratio is synthetic variance divided by human variance (1.0 = same spread as real humans; below 1.0 = narrower).

Scale directions differ by item: for trust and happiness, lower means more trusting or happier; for life satisfaction (0–10), higher means more satisfied; competition runs 1 (good) to 10 (harmful); work importance runs 1 (very important) to 4 (not at all).

Average distributional convergence was approximately 65%.

The two findings below turned out to be the most useful part of the exercise.

Finding 1: The variance collapsed

Look at the variance ratios.

Interpersonal trust preserved variance reasonably well (0.87 and 0.80). But most other variables did not. Life satisfaction is the clearest case: in the United States, the variance ratio was only 0.06. Real humans spread across the full scale; synthetic respondents clustered tightly around the center.

The synthetic populations disagreed with each other less than real humans do.

This is a known tendency of language models — they regress toward the most probable answer, producing narrower distributions, fewer extremes, fewer outliers, less disagreement. Knowing the mechanism doesn't excuse the result.

It doesn't invalidate the benchmark. But it reveals a genuine technical challenge: if synthetic respondents are going to become a serious research methodology, preserving human variance may be one of the most important problems to solve. It's one we're now working on directly.

Finding 2: Close on trust levels, flat on trust differences

The second finding surprised us more.

Interpersonal trust is one of the classic examples of cross-cultural variation: in the WVS data, Great Britain reports higher trust than the United States, and the gap is well documented.

Within each country, the synthetic trust distribution overlapped the real one by 90.5% (US) and 81.6% (GB). The gap between the two countries, however, did not appear. The synthetic US and GB populations returned almost identical trust scores (1.72 in both). The real populations did not.

At first glance, that looks like a failure. We think it reveals something more interesting.

The boundary between psychology and culture

The benchmark forced a more fundamental question: what is our system actually modelling?

Differences between whole societies did not emerge on their own from individual profiles.

Our architecture is designed to create variation between individuals — different motivations, priorities, psychographic structures, and ways of making sense of the world. What it models less strongly are the institutional and historical forces that shape entire societies.

Trust isn't purely a psychological variable. It's also cultural and institutional — it emerges from social norms, civic institutions, historical experience, collective expectations, shared narratives. Those forces exist above the individual level, and our benchmark suggests they should be treated as a separate modelling layer rather than assumed to emerge automatically from individual psychology.

This isn't something we discovered by accident. It's a boundary condition. And understanding boundary conditions is part of taking a methodology seriously.

Why this matters more than it seems

If our goal had been to prove perfect population realism, this benchmark would be disappointing. Fortunately, that wasn't the most important thing we learned.

Most commercial decisions are not made between countries. They're made between customer types.

Different motivations, anxieties, priorities, and decision styles. A product team rarely needs the average trust score of a nation. It needs to understand why one segment embraces a proposition while another rejects it.

That is a question about segments, not national averages, and it is the question qualitative pre-research is built around: why each respondent answers the way it does, from its own situation.

We ran that test next: a pre-registered study of 2,600 synthetic respondents across eight markets, sealed before any data and published in full, including the prediction the data refuted.

The pre-registered study, its controls and its result.

Read the pre-registered study →

What this benchmark does not prove

This benchmark does not demonstrate predictive validity. It does not prove synthetic respondents can predict purchase intent, willingness to pay, brand choice, or product adoption. Nor does it demonstrate commercial usefulness. Those questions require prospective studies, parallel human-synthetic comparisons, and real-world decision outcomes.

It addresses a more basic question: how close do synthetic answer distributions come to real ones, item by item? The tables above are the answer for this run, with the caveats above.

A better validation framework

This project also changed how we think about validation. Population realism shouldn't be the finish line. It should be the starting point. A more useful hierarchy:

  • Level 1 — Population realism. Do synthetic populations resemble real ones?
  • Level 2 — Psychographic realism. Do different psychological profiles diverge coherently?
  • Level 3 — Reproducibility. Do repeated runs produce stable results?
  • Level 4 — Human convergence. Do findings align with parallel human research?
  • Level 5 — Predictive utility. Do the resulting insights improve decisions?

Only the final two establish commercial value. But the first three establish credibility — and credibility is the resource synthetic research needs most right now.

A note on transparency

World Values Survey Wave 7 predates today's frontier models and may be partially represented within their training corpora. For that reason we do not treat this as proof of predictive validity — a model can score well by memory rather than fidelity. We treat it as a back-test to learn from, not as validation.

We're publishing the methodology, and we invite anyone building synthetic respondents to run the same test and post their own numbers — variance ratios included. The point isn't to win an argument. It's to give the category a shared, reproducible way to ask how close is close enough.

What we learned

We started this benchmark trying to answer a simple question: are synthetic respondents just averages?

The back-test did not settle that on its own. What it gave us was sharper: synthetic answers came out narrower than human ones, and differences between whole societies did not emerge on their own from individual profiles. Both are the kind of result you only get by running the test in public.

That is useful. It tells you where to look first when you test any synthetic panel.

Synthetic research won't become credible because it gets everything right. It will become credible when we understand precisely where it gets things wrong.

References and further reading

  • Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3), 337–351.
  • Park, J. S., Zou, C. Q., Shaw, A., et al. (2024). Generative Agent Simulations of 1,000 People. Stanford University.
  • Larooij, M., & Törnberg, P. (2025). Validation is the Central Challenge for Generative Social Simulation. AI Review (Springer).
  • Morocho, E. E. T., Cima, L., Cresci, S., et al. (2026). Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents. WWW 2026 Companion.
  • Bisbee, J., Kennedy, R., & Goel, S. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis.
  • Haerpfer, C., et al. (eds.). World Values Survey: Round Seven. WVS Association.

Want to run the same back-test on your own category? Let's talk.

QualiSynth