White Paper · On Creating Audiences Nº 02 Book a call
Composable AI × Research methodology

How Sunzu simulates human behaviour.

How composable AI, customer evidence and model checks support simulations for product, design and marketing research.

Sunzu simulates human behaviour with composable AI: 20+ layers, including algorithms and small models, working together to support customer research.

Our first paper explains where simulations can help a team make better-informed decisions. This paper describes the evidence that grounds those simulations, the audience and persona models, and how to interpret our internal benchmarks.

A useful simulation must do more than produce a plausible conversation. Research identifies risks including stereotypes, caricatures and weak validation.4,9 Asking an AI to speak for a demographic group does not establish that it reflects the group’s opinions.8

Sunzu combines customer evidence, audience modelling and response checks within its composable AI system. The layer count describes the system’s composition. Its usefulness for a research task is assessed through the findings it produces and their relationship to human evidence.

How an audience is created A two-stage, double-diamond process. Stage one, data enrichment: customer insights are enriched and organised. Stage two, behavioural modelling: audience models are initialised, then tuned, producing your audiences. Data enrichment Behavioural modelling Enrich Synthesise Model Tune Audience models are initialised Your customer insights Your audience models

Your customer insights

  1. Data enrichment

    Enrich → Synthesise

  2. Audience models are initialised

    Behavioural modelling

    Model → Tune

Your audiences

Two stages: organise the evidence, then build and validate personas.
01
Stage one · Enrich → Synthesise

Modelling the audience

Customer evidence defines the audience’s context. The simulation explores how people in that context might respond.

We identify customer groups supported by the supplied research and documents. Interviews, surveys, support themes and other customer evidence help establish their circumstances, motivations and constraints.

Research by Park and colleagues found that agents grounded in interviews or surveys predicted responses more accurately than agents based on demographics alone.10 This supports the approach of grounding simulations in customer evidence. It does not independently validate Sunzu’s implementation.

Findings extracted for the audience model retain source links so reviewers can check them and update the model when evidence changes. Relevant public reviews, forums, news and published studies provide additional context. Findings are checked against customer evidence and assigned confidence ratings.

Sunzu’s behavioural corpus contains 1M+ profiles, 34M+ attributes and 57M+ data points, according to our internal research. It supplements the context available from customer evidence. These figures describe the underlying evidence base, not recruited participants or the sample size of a benchmark. The underlying proprietary data is not publicly disclosed.

Broader evidence can suggest what to investigate when customer data is sparse. Its relevance depends on the audience and question. Source links and confidence ratings support review; they do not, by themselves, establish that a simulated answer is correct.

The audience model has four components. These describe its contents, rather than the 20+ layers of the AI system:

01 Demographic profile Who they are

The group’s documented background and circumstances, using characteristics supported by the source material.

02 Psychographic insights How they think

Motivations, attitudes and preferences, with evidence and confidence ratings for each finding.

03 Scenario specifics What is at stake

The needs, pain points and constraints relevant to the question being tested.

04 Abstentions What we do not know

Gaps in the evidence, marked explicitly so missing information is not presented as a finding.

The four parts of the audience model.
02
Stage two · Model → Tune

Generating the persona

A persona represents an individual within the audience model. Its responses are AI-generated simulations.

Each persona starts with a defined role, situation, attitudes and likely decision factors. Contextual detail supports responses across the situations being tested.

The persona retains links to the audience model and its evidence. This lets reviewers trace its basis while distinguishing a source finding from a modelled response.

Persona responses should be read as simulated reactions. A realistic portrait or fluent answer does not add human observations to the study.

An audience portrait
Illustrative portrait of a modelled persona.

Each persona must pass coherence and alignment checks before it is used in a study.

Gate 01 / Coherence
Does it hold together?

We check that the persona’s attitudes, background and behaviour are internally consistent.

Gate 02 / Alignment
Does it fit the audience?

We test responses across several turns of conversation against the audience model and its evidence. Alignment includes variation between responses: matching a group average can still hide differences in the distribution.12

Coherence and alignment assess consistency with the audience evidence. Predictive validity asks how closely findings match what people report or do.

Back-testing informs model refinement. For your own evaluation, keep the human findings used to judge results separate from the material used to build or adjust the audience.

Benchmarks against human findings

The following results come from Sunzu’s internal research on proprietary data. They measure reference findings recovered in specific back-tests. Concept, UX and ad testing use research with 500+ humans as the reference; depth interviews use human insights retested within two weeks. Each human reference is normalised to 100%. Results vary by audience and task.

  1. Discover

    Depth interviews

    Understand problems, priorities and the language your audience uses.

    Reported sample equivalence
    10+
    humans
    Themes covered
    Sunzu in minutes 76%
    Human expert with AI in days 54%
    Human insights, retested in months 100%
  2. Frame

    Concept testing

    Compare ideas and identify the strongest direction for your audience.

    Reported sample equivalence
    30+
    humans
    Insights found
    Sunzu in minutes 85%
    Human expert with AI in days 58%
    Research, 500+ humans in months 100%
  3. Shape

    UX testing

    Find confusing steps and usability problems before release.

    Reported sample equivalence
    20+
    humans
    Defects detected
    Sunzu in minutes 75%
    Human expert with AI in days 34%
    Research, 500+ humans in months 100%
  4. Launch

    Ad testing

    Compare messages and creative before spending on a campaign.

    Reported sample equivalence
    20+
    humans
    Insights found
    Sunzu in minutes 69%
    Human expert with AI in days 55%
    Research, 500+ humans in months 100%

Read the scope below, or see the benchmark comparison table.

03
Evidence and interpretation

What the benchmarks establish

Each result measures a specific comparison with human findings. Use the benchmark that matches your research task.

Coverage of reference findings. The percentages describe themes covered, insights found or defects detected. The human reference is set to 100% to express relative coverage. That baseline does not imply that the reference research captured every possible finding.

Coverage and correctness. Recovering a reference finding and avoiding an unsupported finding are different measures. Coverage alone does not state how many additional findings were incorrect, how serious a missed defect was or whether a proposed decision improved a business outcome.

Sample equivalence. The figures shown are Sunzu’s reported estimates of result equivalence within these comparisons. They are not counts of recruited participants, a statistical margin of error or a guarantee that another study will produce the same result.

Research time. Minutes, days and months describe the compared workflows. Assess total elapsed time for your own process, including preparation, setup and review. The two-week interval belongs to the depth-interview retest reference, rather than the duration of a Sunzu study.

Proprietary research. The underlying datasets and detailed evaluation protocols are proprietary and are not published here. The results are reported by Sunzu from its own research. The external publications cited in this paper provide context for the approach; they are not independent audits of these results.

Application to your audience. A benchmark is a starting point for evaluation. Compare findings with your own human evidence, review misses and unsupported additions, and assess whether the simulation helps your team reach a useful decision.

04
Limits

When to work with people

Use modelled audiences to sharpen early decisions and focus the research you take to customers.

Work directly with people when a decision depends on lived experience, accessibility, cultural context or behaviour that the available evidence does not capture. Unknowns need to remain visible when the evidence cannot support a conclusion.

Some studies show promising results on specific tasks. Maier and colleagues reported purchase-intent predictions close to human test-retest reliability for personal-care product surveys.11 That is evidence for their method and dataset, not a guarantee for every audience or research question.

Research on designers’ use of persona-based chatbots also examines how these tools fit into design practice.6 We use audience models alongside human research, with their sources and limits visible to the team.

Start with a defined audience, a decision and human evidence suitable for evaluating the result. Review the simulation’s sources and limits, then measure the time and effort required to reach a useful conclusion. Our practical guide explains how to connect that evaluation to business value.

Get user insights in minutes.
Shape products people love.

Get started

Test a decision with your audience.

Sources & further reading
  1. Bain & Company. How Synthetic Customers Bring Companies Closer to the Real Ones.
  2. Making Science. Enhancing UX/UI Research with Synthetic Users.
  3. Interaction Design Foundation. Are AI-Generated Synthetic Users Replacing Personas?
  4. arXiv (2025). Creating and Evaluating Personas Using Generative AI.
  5. Snowflake. What Is Synthetic Data? Examples and Use Cases.
  6. AI EDAM, Cambridge University Press. Synthetic users: insights from designers' interactions with persona-based chatbots.
  7. MJV Innovation. How are AI models used to create synthetic users for research?
  8. arXiv (2023), Santurkar et al. Whose Opinions Do Language Models Reflect?
  9. arXiv (2023), Cheng, Piccardi & Yang. CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations.
  10. arXiv (2024), Park et al. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.
  11. arXiv (2025), Maier et al. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings.
  12. arXiv (2026), Moon et al. Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level.

Your PDF is ready

Your download should start automatically. If it doesn’t, use the button below.

Download PDF