We pre-registered a prediction in our own favor. The data said no. We published it.
A pre-registered study of 2,600 synthetic respondents across eight markets, sealed by hash before we saw a single data point. It confirmed, with sealed evidence, what we actually sell — that fixing a persona's profile changes its behavior controllably and traceably — and it falsified what we would have loved to claim: that the same layer is what makes markets differ. We publish both.
Before we saw a single data point, we sealed a prediction by cryptographic hash: that the psychometric layer our engine adds is what makes eight markets answer the same dilemma differently. It was the product-favorable prediction. We registered it, along with the instrument and the full analysis plan, and committed in writing to publish the result whatever it turned out to be.
The answer was no. But the study came back with two headlines, not one — and the one that confirms what we built is as strong as the one that falsified us. This is the writeup we committed to.
This is the pre-registered successor to our earlier exploratory study (n=5 per market). We built it to close that study's declared limitations — and to re-test, with controls, our own earlier claim about where between-market differences come from.
Read the v1 on SSRNWhat the study is
The object is not the psychology of eight countries. It is a construct-validity question for the whole field of synthetic respondents: when a synthetic population differs by country, which layer of the generation procedure produced that difference? Almost nobody isolates it, so the answer is usually a guess. We built the controls to stop guessing.
What we did prove (and it is what defines a segment)
The strongest result in the study is also the one that matters commercially. Fix a synthetic persona's value profile to the opposite pole and its behavior changes in the predicted direction, in all eight markets — the largest effect in the whole study. Concrete example: the behavior "in the end the decision is mine" appears in 76% of the growth-oriented profile versus 42% of the conservative one.
This is not a technical detail: it is exactly what the product promises. Change the persona's vector and its behavior changes — coherently, controllably, traceably. Now we have the pre-registered evidence that the mechanism does what we say it does. That is what a defined segment is: not a random sample, but a profile you can fix and defend.
The full result: two axes with different sources
The differences between markets are also real (effect sizes of 0.17 to 0.38, reproducible across three independent generation seeds). But they do not come from our psychometric layer. A control cell with the layer removed shows the same between-market structure; a flat-prompt version on the same base model shows no degradation either. Two independent controls, built to catch different things, point the same way.
What emerges is a clean separation between two axes that practice tends to conflate:
| Axis | What produces it | In this study |
|---|---|---|
| Profile divergence WITHIN a market | The psychometric vector (value profile) | Confirmed in 8/8 markets, directional predictions held |
| Differences BETWEEN markets | Country + language, as rendered by the base model | Real and reproducible, but the psychometric layer does NOT produce it |
The part most companies cannot publish: we corrected ourselves
Our first reading of the data suggested something stronger — that the cell without the psychometric layer had more between-market structure, an "inversion." It would have been a striking result. We did not trust it. We tested whether it was an artifact of our own coder being noisier on one cell than another, and it was: about 79% of the apparent inversion was measurement noise. So we withdrew the strong claim and kept the firm, smaller one: the layer adds no between-market structure. We caught our own over-claim before publishing, not after.
What we DO claim, and what we do NOT
In two columns, because the distinction is the product:
We do claim
- That fixing the psychometric profile produces a distinct, coherent, stable segment within a market — confirmed in all eight, with the predicted directional predictions holding. It is the axis the layer governs, and it is what we sell.
- That the between-market structure is real and reproducible across seeds.
- That all of this is auditable: pre-registered by hash, with 2,600 labeled responses and fixed-seed code, all open.
We do NOT claim
- That our psychometric layer captures the culture of each country. It does not — that is the prediction the data falsified. Between-market structure is rendered by the base model; a second model, given identical country and language, did not differentiate the markets at all. Market labels denote the generation context, not a country.
- That this is "validated." The blind human coding has now run — see the update at the end. Two of the six behaviours cleared the bar we had pre-registered; the gate as a whole did not, so the word still does not apply. We set that bar ourselves, before there was any data, and we report against it either way: that is what separates an auditable study from a self-reported number.
To be exact about what the psychometric layer governs, because it is easy to summarize wrong: it governs profile divergence within a market — what defines a segment, confirmed 8/8 — and not the structure between countries. Two distinct axes with distinct sources; conflating them is exactly the error this study exists to correct.
Why we published it
The result has an uncomfortable face for us: the layer our product adds does not do one of the things its market value would invite you to assume. We publish it anyway because a study you can audit — pre-registered by hash, data and code open, results against our own interest included — is worth more than a validation number nobody can check. A skeptic can re-run the whole analysis and see that we did not move the goalposts. That is the whole point.
And the face that is not uncomfortable is the one that actually matters to a buyer: we proved, with sealed evidence, that the mechanism we sell — profiles that produce distinct, traceable segments — does exactly what we say. We would rather be the people who publish the study that goes against part of their own product, and prove the part that holds, than the people who can publish neither.
Update, 6 September 2026 — the human coding is done
We said human validation was the next phase. It has now happened, and this is what it found. We recruited 27 people on Prolific and had them read 120 of the answers blind: translated into English, stripped of any cue to their market, and never told that the texts were synthetic or that the task was called coding. Before the real items, each of them answered three questions with known answers; seven failed that calibration and were excluded from the analysis, though not from payment, and the blocks they had covered were relaunched with a replacement. After those exclusions every one of the 120 texts is left with exactly two independent readings.
Two of the six behaviours came through cleanly: agreement between our sealed classifier and the human consensus reached 0.93 and 0.96. But we had written the gate as the weakest behaviour, before we had any data, and by that rule it does not pass: 0.62 against the 0.70 we had committed to. H2 and H3 therefore remain unconfirmed, exactly as the pre-registered failure clause requires.
The interesting part is why. It is not that the classifier reads the texts badly. It is that the readers disagree with each other: on four of the six behaviours, two people reading the same answer do not agree on whether the behaviour is present. Where there is no human agreement there is no criterion to measure a machine against. The problem is our taxonomy, not our coder.
That has a consequence worth stating for the whole field, and it is the most useful thing this phase produced. Our two language-model judges agreed with each other around 90% of the time on the very constructs where two people do not reach a Krippendorff alpha of 0.5. Agreement between language models is easy to obtain and easy to mistake for validity: it measures the stability of a procedure, not the existence of the category. Only blind human coding tells the two apart, which is exactly why we paid for it.
What happens next is the ordinary work of a research programme. We rewrite the definitions of the four behaviours the readers could not apply consistently, and re-code the same 120 texts — the data is already collected and sealed, so that costs a redraft and not a new study. Separately, the human respondent arm, where real people answer the same dilemma instead of coding it, is prepared and sealed; that is the step that decides whether synthetic answers converge with human ones at all. The paper is now at v2.1 with the full Phase 2 result inside it, data and code included.
The pre-registration, the 2,600 responses and the analysis code — the full trail behind this study lives on our evidence page.
See the evidenceTry the engine — no signup
QualiSynthIf you do qualitative research, we would genuinely like to know where this would mislead you. That gap is what we are after.