How do you justify synthetic respondents to your client?
The question a research director asks is not "do they sound real". It is "how do you know they are who you said they were?" Here is our answer, the number behind it, and the case where we got it wrong first.
Sounding real is table stakes. Any large language model clears that bar, and your client knows it. The question that decides the meeting is narrower and much harder: if you asked for people who already own the product, how do you know the people who were interviewed own it?
A segment is a promise about who is in the sample. If nobody checks the promise, the transcripts come out fluent and the study is worthless. We got that wrong before we got it right, so let us start there, and then look at what had to be built so that it cannot happen again.
What went wrong in our own study?
In August 2026 we ran an automotive study with three segments defined by past behaviour: people who already drive an electric car, people who considered one and bought petrol or diesel, and people who never considered one.
Then we read the profiles by hand. In the segment labelled "already drives an electric car", 22 of 23 respondents said they had ruled one out. In the segment labelled "never considered it", 18 of 23 said they had considered it. The three segments were the same people wearing three labels.
Every between-segment reading from that run was void, and none of it was visible in the transcripts, which were perfectly coherent. Coherent and wrong is the dangerous combination, because it survives review. The cause was mundane: a segment described in prose does not reach the respondent unless something carries it there, and past behaviour had no field to travel in.
How do we check a segment now?
Before any interview runs, every generated respondent is judged against the written segment definition, and the judge returns one of three verdicts:
- complies - the biography states what the segment requires
- contradicts - the biography states the opposite
- not stated - the biography is silent on it
Only contradictions are discarded, and every discarded respondent is regenerated until the sample is full. Two verdicts would not work: with a binary judgment you have to choose between admitting contradictions and emptying the segment, because most biographies are simply silent on any given condition. The third verdict is what makes the check safe to run.
You get the number back. Not a promise that the segment is clean, but the percentage of respondents in each segment that complied, per study.
This is the step a general-purpose chatbot does not have. Ask one to play a consumer who owns an electric car and it will oblige, instantly and convincingly, because there is nothing between the request and the answer: no sample to inspect, nobody to discard, nobody to replace, no number to report back. The hard part was never producing the conversation. It is everything that has to happen before the conversation is allowed to start.
How accurate is the judge?
180 judgments, zero errors. The measurement is 18 real stored biographies, each judged six times, across all three verdict classes. Repeating the same question over the same text is the only way to separate criterion from noise: a judge that is right once may be guessing.
On the same automotive study, re-run with the check in place, the three segments came back at 21 of 23, 23 of 23, and 23 of 23.
The hardest case we have measured is a condition stated in the negative - "does not travel" - which a generator satisfies on its own only 2 times out of 5. With the check in place, asking for five people returns five of five. The cost is a second generation round.
Can a segment be defined by behaviour alone, with no age and no occupation?
Yes. We tested four transport segments - train, bus, bicycle, and "does not travel" - with no age range and no occupation given, five people each. All four came back at 100%, with ages ranging from 29 to 83 inside each segment.
One caveat that matters when you write the brief: the check judges what your segment declares, not what the other segments declare. If your cyclists must exclude public transport users, that exclusion has to be written into the condition. Written, it is enforced: we judged train and bus respondents against "cycles and does not use public transport" and got five contradicts, zero complies. Left unwritten, nobody checks it.
Does the psychological profile change what they choose?
Start with the part that does work, because it is measurable and you can audit it yourself: the profile we ask for is the profile that arrives. Each of the four profiles scored 0.86 to 0.97 on its own target dimension against 0.40 to 0.45 on the others. The vector is set, it reaches the respondent, and it shows up in how they talk and how they rate themselves.
What it did not do was change the choice.
We ran a dedicated study: 100 respondents, 25 per psychometric profile, one shared stimulus. The four profiles chose almost identically, at chi-squared 0.17 and p = 0.982, and the top-ranked decision criterion was the same across all four.
What we can claim from this is narrow - one stimulus, one category. It does not license the broader claim that the engine flattens people in general, because without a human arm we cannot separate "the model does not differentiate" from "this decision does not depend on values". Plenty of real decisions do not.
We report it because it is the kind of finding only an instrumented system can produce. To establish that a psychological profile does not move a particular decision, you first have to set the profile, prove it was set, hold everything else constant and count. Something that cannot do those four things cannot find this out. It can only assume, in whichever direction sells better.
What does move the decision, then?
Behaviour. In the same automotive study, given a petrol-versus-electric choice:
- segment "already drives an electric car": 9 of 20 chose the electric option
- the two segments that do not own one: 0 of 41
That is the largest between-segment difference our test battery has produced, and the first one resting on segments verified respondent by respondent. Read it for what it is: these are synthetic respondents, not a market sample. The finding is that the segment definition changes the response, not that some share of real drivers buy electric. And the design does not establish causation.
When does this not work?
Every limit below is one we can state because we measure it. A tool that never inspects its own sample does not avoid these problems - it just has no way of seeing them.
- Purity is not comparable across studies. It depends on how observable your written condition is. A condition that leaves no trace in a biography - a purchase channel, for instance - makes the check switch itself off and deliver the segment unverified.
- Segments are not guaranteed mutually exclusive. They are guaranteed to match their own definition. Exclusivity is something you write.
- A malformed judgment falls back to the innocuous verdict. If the response is unparseable twice in a row, that respondent enters. It is logged, but it enters.
What can you tell your client?
That the sample was checked against the brief before anyone was interviewed, that respondents contradicting the brief were discarded and replaced, and that the compliance rate per segment is in the deliverable.
And, when they ask the harder question, that we know one thing the category tends to oversell: setting a psychological profile did not change what our respondents chose. What changed it was defining the segment by what people have actually done. We would rather tell you that now than have you find it in a debrief.
See how the check works on a segment of your own
QualiSynth