What it means
Human test-retest consistency is a reliability benchmark from research: Ask the same person the same question 2 weeks apart, and they are about 85% consistent with themselves. People are not perfectly stable, so 85% is the realistic ceiling for how repeatable a human answer is. You use this number as a yardstick to judge whether an AI or persona output is accurate enough to trust. If a system matches what a person says about as often as that person matches themselves, it is performing at human-level reliability.
Why it matters
Without a benchmark, "85% accurate" sounds like a failing grade, and teams dismiss AI outputs that are actually strong. Human test-retest consistency reframes that number. When Stanford and Google DeepMind trained synthetic personas on interview transcripts and tested them against real follow-up surveys, the personas hit 85% accuracy, which is comparable to a human answering themselves. That is the bar. It tells you when an AI persona is good enough to act on and when chasing higher numbers is chasing noise.
Say your synthetic persona answers a set of survey questions and matches your real customers 84% of the time. Measured against human test-retest consistency, that is not a flaw to fix. It is the persona behaving as reliably as the people it models, so you treat its output as a credible filter rather than holding out for an unrealistic 95%.
How to use this knowledge
Set 85% as the realistic ceiling. Judge AI and persona outputs against human consistency, not against perfection. Demanding more than people manage themselves wastes effort.
Validate before you trust. Compare persona answers to known business truths and real customer responses, the way the Stanford study checked personas against follow-up surveys.
Use it to settle "is this good enough" debates. When stakeholders question an AI output's reliability, anchor the conversation to the human test-retest baseline instead of an arbitrary target.
Keep humans on final decisions. Human-level consistency makes a persona a strong filter, not a replacement for real validation before you ship.
Growth Memo guidance
For context, that's comparable to human test-retest consistency. If you ask the same person the same question 2 weeks apart, they're about 85% consistent with themselves.
Result: 85% accuracy. The synthetic personas replicated what the actual study participants said.
Synthetic personas — AI-generated profiles whose accuracy is judged against this human consistency benchmark.
Validation benchmarks — Reality checks that pair with test-retest consistency to catch persona hallucinations.
Prompt tracking — Repeated AI runs whose variance you accept once you know human answers vary too.
Confidence score — A per-field reliability rating that complements an overall consistency benchmark.
Sycophancy bias — A failure mode where AI personas over-agree, lowering the realism that test-retest consistency assumes.

