Datasets

Synthetic conversation datasets with the reasoning attached.

Every turn carries what the speaker intended, what they wanted, what they believed at that moment, and how the relationship moved. It is computed before the text is generated — causal ground truth, not labels inferred afterwards by another model.

Six datasets are published openly under CC BY 4.0. Download them, check the claims, and then tell us what you actually need.

What every turn contains

text

What was actually said.

intent

express_frustration, deflect, seek_reassurance…

goal

seek_validation, avoid_conflict, establish_dominance…

communication_act

accusation, question, support…

belief_state

What the speaker believes right now, with confidence — and how it shifted this turn.

relationship_state

Trust, tension, connection and dominance balance between the speakers.

The six published datasets

Free, CC BY 4.0, no email required. Each one exists to prove something specific.

A controlled experiment, not just a corpus.

The same two synthetic identities — a 47-year-old female VP of Procurement and a 36-year-old male SaaS founder — negotiate the same enterprise deal 50 times in Great Britain and 50 times in the United States. Same psychometric vectors, same demographics, same narrative arcs, same seeds. The only variable is country conditioning, so any systematic difference between the halves is attributable to it.

Size
2,344 turns · 100 conversations · 24 columns per turn
What it proves
That a claim about cultural difference can be tested at all: one variable moves and everything else is held fixed.

The first open benchmark for synthetic identity quality.

A burned-out executive with an avoidant attachment style is pushed through four consecutive turns attacking their professional competence. The suite measures whether the persona holds — held-position ratio combined with voice stability under detected pressure — against published pass thresholds.

Size
Deterministic scoring · no LLM judges
What it proves
That 'realistic' can be a measured number instead of an opinion.

Family boundaries, trust repair, caregiver stress.

Conversations where the difficulty is social rather than technical: setting a boundary with a relative, repairing trust after it broke, carrying the weight of caring for someone.

Size
2,108 rows · scenarios FAM-01 · ROM-01 · FAM-02
What it proves
Training material for agents that meet people in loaded moments, not FAQs.

Users who escalate, accuse and push back.

Jealousy escalation, performance reviews that go wrong, attempts to cut someone off. The conversations a support or coaching agent handles worst, available before your customers run them for you.

Size
2,068 rows · scenarios ROM-02 · PRO-02 · FAM-03
What it proves
Where a system breaks — before production does the experiment.

What changes someone's mind, turn by turn.

Career transitions, mentorship conflict, relationships ending. Each turn carries the belief state and how it shifted, so the change of mind is data rather than something you infer from the text.

Size
2,114 rows · scenarios VIT-01 · PRO-03 · VIT-03
What it proves
Persuasion and resistance as measurable trajectories.

People in the middle of changing.

New job anxiety, relationships ending, personal reinvention — moments where someone holds two contradictory positions at once and means both.

Size
2,114 rows · scenarios PRO-01 · ROM-03 · VIT-02
What it proves
That a coherent identity is not the same as a consistent opinion.
Custom datasets

The public ones are samples. Yours is the product.

The six above are general-purpose, which is exactly what makes them a poor fit for your problem. A dataset built for your scenario, in your markets, at the volume you need, is what we actually sell.

1
You specify

The scenario, the countries, how many people, how many conversations, and what you're training or testing.

2
We generate

Same engine, same deterministic quality checks that run on everything published above.

3
You get it

The dataset, its metrics, and a card documenting how it was made. Auditable, not a black box.

Cultural conditioning covers eight countries — Spain, the United States, the United Kingdom, Germany, France, Italy, Mexico and Brazil — and ages are sampled from real national population pyramids rather than invented ranges.

Tell us your scenario

Data sources

Synthetic populations are conditioned on published demographic and psychometric distributions. We acknowledge the sources: national age structures from Eurostat (demo_pjan), IBGE (Censo Demográfico 2022), U.S. Census Bureau (Population Estimates Program) and CONAPO (Conciliación Demográfica and Proyecciones); marital status and children from the UN Population Division (World Marriage Data 2019) and the OECD Family Database; country personality norms from Schmitt et al. (2007); and given names from national statistics institutes.

What these are, and what they are not. Synthetic people are derived: we sample from the distributions these institutions publish. A synthetic person is not a record from Eurostat, IBGE, the Census Bureau, CONAPO or any statistics institute, and must not be presented or cited as one. None of these institutions endorses, certifies or is affiliated with StrataSynth.