Social Reasoning
Family boundaries, trust repair, caregiver stress.
What is in this dataset?
Conversations where the difficulty is social rather than technical: setting a boundary with a relative, repairing trust after it broke, carrying the weight of caring for someone.
Every turn carries what the speaker intended, what they wanted, the communication act, what they believed at that moment and how the relationship moved — computed before the text was generated, not inferred from it afterwards.
Why does it exist?
Most open dialogue data is transactional: a question, an answer, a resolution. Systems trained only on that behave badly the first time someone brings a problem that has no clean answer.
How was it built?
- Three scenarios: family boundaries, romantic trust repair, caregiver stress.
- Each speaker holds a psychological profile that persists across the whole conversation.
- Intent, goal and communication act are resolved before the sentence is rendered.
What does it prove — and what does it not?
Training material for agents that meet people in loaded moments, not FAQs.
These are synthetic conversations, not transcripts of real people. They are useful as training and test material; they are not evidence about how any real population behaves.
What can you use it for?
- Fine-tuning assistants that have to handle emotionally loaded turns.
- Testing whether a model keeps its footing when there is no clean answer.
- Studying intent and belief annotations that were computed, not inferred.
Data sources
Synthetic populations are conditioned on published demographic and psychometric distributions. We acknowledge the sources: national age structures from Eurostat (demo_pjan), IBGE (Censo Demográfico 2022), U.S. Census Bureau (Population Estimates Program) and CONAPO (Conciliación Demográfica and Proyecciones); marital status and children from the UN Population Division (World Marriage Data 2019) and the OECD Family Database; country personality norms from Schmitt et al. (2007); and given names from national statistics institutes.
What these are, and what they are not. Synthetic people are derived: we sample from the distributions these institutions publish. A synthetic person is not a record from Eurostat, IBGE, the Census Bureau, CONAPO or any statistics institute, and must not be presented or cited as one. None of these institutions endorses, certifies or is affiliated with StrataSynth.