We pre-registered a prediction in our own favor. The data said no. We published it.
A pre-registered study of 2,600 synthetic respondents across eight markets, sealed by hash before we saw a single data point. It confirmed, with sealed evidence, what we actually sell — that fixing a persona's profile changes its behavior controllably and traceably — and it falsified what we would have loved to claim: that the same layer is what makes markets differ. We publish both.
Before we saw a single data point, we sealed a prediction by cryptographic hash: that the psychometric layer our engine adds is what makes eight markets answer the same dilemma differently. It was the product-favorable prediction. We registered it, along with the instrument and the full analysis plan, and committed in writing to publish the result whatever it turned out to be.
The answer was no. But the study came back with two headlines, not one — and the one that confirms what we built is as strong as the one that falsified us. This is the writeup we committed to.
This is the pre-registered successor to our earlier exploratory study (n=5 per market). We built it to close that study's declared limitations — and to re-test, with controls, our own earlier claim about where between-market differences come from.
Read the v1 on SSRNWhat the study is
The object is not the psychology of eight countries. It is a construct-validity question for the whole field of synthetic respondents: when a synthetic population differs by country, which layer of the generation procedure produced that difference? Almost nobody isolates it, so the answer is usually a guess. We built the controls to stop guessing.
What we did prove (and it is what defines a segment)
The strongest result in the study is also the one that matters commercially. Fix a synthetic persona's value profile to the opposite pole and its behavior changes in the predicted direction, in all eight markets — the largest effect in the whole study. Concrete example: the behavior "in the end the decision is mine" appears in 76% of the growth-oriented profile versus 42% of the conservative one.
This is not a technical detail: it is exactly what the product promises. Change the persona's vector and its behavior changes — coherently, controllably, traceably. Now we have the pre-registered evidence that the mechanism does what we say it does. That is what a defined segment is: not a random sample, but a profile you can fix and defend.
The full result: two axes with different sources
The differences between markets are also real (effect sizes of 0.17 to 0.38, reproducible across three independent generation seeds). But they do not come from our psychometric layer. A control cell with the layer removed shows the same between-market structure; a flat-prompt version on the same base model shows no degradation either. Two independent controls, built to catch different things, point the same way.
What emerges is a clean separation between two axes that practice tends to conflate:
| Axis | What produces it | In this study |
|---|---|---|
| Profile divergence WITHIN a market | The psychometric vector (value profile) | Confirmed in 8/8 markets, directional predictions held |
| Differences BETWEEN markets | Country + language, as rendered by the base model | Real and reproducible, but the psychometric layer does NOT produce it |
The part most companies cannot publish: we corrected ourselves
Our first reading of the data suggested something stronger — that the cell without the psychometric layer had more between-market structure, an "inversion." It would have been a striking result. We did not trust it. We tested whether it was an artifact of our own coder being noisier on one cell than another, and it was: about 79% of the apparent inversion was measurement noise. So we withdrew the strong claim and kept the firm, smaller one: the layer adds no between-market structure. We caught our own over-claim before publishing, not after.
What we DO claim, and what we do NOT
In two columns, because the distinction is the product:
We do claim
- That fixing the psychometric profile produces a distinct, coherent, stable segment within a market — confirmed in all eight, with the predicted directional predictions holding. It is the axis the layer governs, and it is what we sell.
- That the between-market structure is real and reproducible across seeds.
- That all of this is auditable: pre-registered by hash, with 2,600 labeled responses and fixed-seed code, all open.
We do NOT claim
- That our psychometric layer captures the culture of each country. It does not — that is the prediction the data falsified. Between-market structure is rendered by the base model; a second model, given identical country and language, did not differentiate the markets at all. Market labels denote the generation context, not a country.
- That this is "validated." No judge in this phase is human — interim reliability is an independent language-model coder. Human validation is the next phase, and only then does that word apply. This is a provisional Phase 1 result, and we say so with our head up: saying it plainly is what separates an auditable study from a self-reported number.
To be exact about what the psychometric layer governs, because it is easy to summarize wrong: it governs profile divergence within a market — what defines a segment, confirmed 8/8 — and not the structure between countries. Two distinct axes with distinct sources; conflating them is exactly the error this study exists to correct.
Why we published it
The result has an uncomfortable face for us: the layer our product adds does not do one of the things its market value would invite you to assume. We publish it anyway because a study you can audit — pre-registered by hash, data and code open, results against our own interest included — is worth more than a validation number nobody can check. A skeptic can re-run the whole analysis and see that we did not move the goalposts. That is the whole point.
And the face that is not uncomfortable is the one that actually matters to a buyer: we proved, with sealed evidence, that the mechanism we sell — profiles that produce distinct, traceable segments — does exactly what we say. We would rather be the people who publish the study that goes against part of their own product, and prove the part that holds, than the people who can publish neither.
The pre-registration, the 2,600 responses and the analysis code — the full trail behind this study lives on our evidence page.
See the evidenceTry the engine — no signup
QualiSynthIf you do qualitative research, we would genuinely like to know where this would mislead you. That gap is what we are after.