How Sunzu simulates human behaviour.
How composable AI, customer evidence and model checks support simulations for product, design and marketing research.
Sunzu simulates human behaviour with composable AI: 20+ layers, including algorithms and small models, working together to support customer research.
Our first paper explains where simulations can help a team make better-informed decisions. This paper describes the evidence that grounds those simulations, the audience and persona models, and how to interpret our internal benchmarks.
A useful simulation must do more than produce a plausible conversation. Research identifies risks including stereotypes, caricatures and weak validation.4,9 Asking an AI to speak for a demographic group does not establish that it reflects the group’s opinions.8
Sunzu combines customer evidence, audience modelling and response checks within its composable AI system. The layer count describes the system’s composition. Its usefulness for a research task is assessed through the findings it produces and their relationship to human evidence.
Your customer insights
Data enrichment
Enrich → Synthesise
- Audience models are initialised
Behavioural modelling
Model → Tune
Your audiences
Modelling the audience
Customer evidence defines the audience’s context. The simulation explores how people in that context might respond.
We identify customer groups supported by the supplied research and documents. Interviews, surveys, support themes and other customer evidence help establish their circumstances, motivations and constraints.
Research by Park and colleagues found that agents grounded in interviews or surveys predicted responses more accurately than agents based on demographics alone.10 This supports the approach of grounding simulations in customer evidence. It does not independently validate Sunzu’s implementation.
Findings extracted for the audience model retain source links so reviewers can check them and update the model when evidence changes. Relevant public reviews, forums, news and published studies provide additional context. Findings are checked against customer evidence and assigned confidence ratings.
Sunzu’s behavioural corpus contains 1M+ profiles, 34M+ attributes and 57M+ data points, according to our internal research. It supplements the context available from customer evidence. These figures describe the underlying evidence base, not recruited participants or the sample size of a benchmark. The underlying proprietary data is not publicly disclosed.
Broader evidence can suggest what to investigate when customer data is sparse. Its relevance depends on the audience and question. Source links and confidence ratings support review; they do not, by themselves, establish that a simulated answer is correct.
The audience model has four components. These describe its contents, rather than the 20+ layers of the AI system:
The group’s documented background and circumstances, using characteristics supported by the source material.
Motivations, attitudes and preferences, with evidence and confidence ratings for each finding.
The needs, pain points and constraints relevant to the question being tested.
Gaps in the evidence, marked explicitly so missing information is not presented as a finding.
Generating the persona
A persona represents an individual within the audience model. Its responses are AI-generated simulations.
Each persona starts with a defined role, situation, attitudes and likely decision factors. Contextual detail supports responses across the situations being tested.
The persona retains links to the audience model and its evidence. This lets reviewers trace its basis while distinguishing a source finding from a modelled response.
Persona responses should be read as simulated reactions. A realistic portrait or fluent answer does not add human observations to the study.
Each persona must pass coherence and alignment checks before it is used in a study.
We check that the persona’s attitudes, background and behaviour are internally consistent.
We test responses across several turns of conversation against the audience model and its evidence. Alignment includes variation between responses: matching a group average can still hide differences in the distribution.12
Coherence and alignment assess consistency with the audience evidence. Predictive validity asks how closely findings match what people report or do.
Back-testing informs model refinement. For your own evaluation, keep the human findings used to judge results separate from the material used to build or adjust the audience.
Benchmarks against human findings
The following results come from Sunzu’s internal research on proprietary data. They measure reference findings recovered in specific back-tests. Concept, UX and ad testing use research with 500+ humans as the reference; depth interviews use human insights retested within two weeks. Each human reference is normalised to 100%. Results vary by audience and task.
-
Discover
Depth interviews
Understand problems, priorities and the language your audience uses.
- Reported sample equivalence
- 10+
- humans
Themes covered Sunzu in minutes 76%Human expert with AI in days 54%Human insights, retested in months 100% -
Frame
Concept testing
Compare ideas and identify the strongest direction for your audience.
- Reported sample equivalence
- 30+
- humans
Insights found Sunzu in minutes 85%Human expert with AI in days 58%Research, 500+ humans in months 100% -
Shape
UX testing
Find confusing steps and usability problems before release.
- Reported sample equivalence
- 20+
- humans
Defects detected Sunzu in minutes 75%Human expert with AI in days 34%Research, 500+ humans in months 100% -
Launch
Ad testing
Compare messages and creative before spending on a campaign.
- Reported sample equivalence
- 20+
- humans
Insights found Sunzu in minutes 69%Human expert with AI in days 55%Research, 500+ humans in months 100%
Read the scope below, or see the benchmark comparison table.
What the benchmarks establish
Each result measures a specific comparison with human findings. Use the benchmark that matches your research task.
Coverage of reference findings. The percentages describe themes covered, insights found or defects detected. The human reference is set to 100% to express relative coverage. That baseline does not imply that the reference research captured every possible finding.
Coverage and correctness. Recovering a reference finding and avoiding an unsupported finding are different measures. Coverage alone does not state how many additional findings were incorrect, how serious a missed defect was or whether a proposed decision improved a business outcome.
Sample equivalence. The figures shown are Sunzu’s reported estimates of result equivalence within these comparisons. They are not counts of recruited participants, a statistical margin of error or a guarantee that another study will produce the same result.
Research time. Minutes, days and months describe the compared workflows. Assess total elapsed time for your own process, including preparation, setup and review. The two-week interval belongs to the depth-interview retest reference, rather than the duration of a Sunzu study.
Proprietary research. The underlying datasets and detailed evaluation protocols are proprietary and are not published here. The results are reported by Sunzu from its own research. The external publications cited in this paper provide context for the approach; they are not independent audits of these results.
Application to your audience. A benchmark is a starting point for evaluation. Compare findings with your own human evidence, review misses and unsupported additions, and assess whether the simulation helps your team reach a useful decision.
When to work with people
Use modelled audiences to sharpen early decisions and focus the research you take to customers.
Work directly with people when a decision depends on lived experience, accessibility, cultural context or behaviour that the available evidence does not capture. Unknowns need to remain visible when the evidence cannot support a conclusion.
Some studies show promising results on specific tasks. Maier and colleagues reported purchase-intent predictions close to human test-retest reliability for personal-care product surveys.11 That is evidence for their method and dataset, not a guarantee for every audience or research question.
Research on designers’ use of persona-based chatbots also examines how these tools fit into design practice.6 We use audience models alongside human research, with their sources and limits visible to the team.
Start with a defined audience, a decision and human evidence suitable for evaluating the result. Review the simulation’s sources and limits, then measure the time and effort required to reach a useful conclusion. Our practical guide explains how to connect that evaluation to business value.
Get user insights in minutes.
Shape products people love.
Test a decision with your audience.
- Bain & Company. How Synthetic Customers Bring Companies Closer to the Real Ones.
- Making Science. Enhancing UX/UI Research with Synthetic Users.
- Interaction Design Foundation. Are AI-Generated Synthetic Users Replacing Personas?
- arXiv (2025). Creating and Evaluating Personas Using Generative AI.
- Snowflake. What Is Synthetic Data? Examples and Use Cases.
- AI EDAM, Cambridge University Press. Synthetic users: insights from designers' interactions with persona-based chatbots.
- MJV Innovation. How are AI models used to create synthetic users for research?
- arXiv (2023), Santurkar et al. Whose Opinions Do Language Models Reflect?
- arXiv (2023), Cheng, Piccardi & Yang. CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations.
- arXiv (2024), Park et al. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.
- arXiv (2025), Maier et al. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings.
- arXiv (2026), Moon et al. Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level.