Building dialogue data that knows why people say things
Technical articles on synthetic dialogue generation, belief state modeling, deterministic evaluation, and the infrastructure behind psychologically coherent AI training data.
PersonaHub, FinePersonas, Nemotron-Personas, PERSONA, Synthetic-Persona-Chat, Twin-2K-500 and StrataSynth Personas each solve a different problem. Here is what each one offers, when to pick it, and what to do when you need people rather than descriptions.
A comparative test between StrataSynth and two leading general-purpose LLMs, using a single senior persona pushed through twenty adversarial turns. The finding: the next frontier for synthetic personas is not better text, but identity that holds under pressure.
We made the same two synthetic negotiators close the same B2B deal 100 times — 50 in Great Britain, 50 in the United States — changing nothing but country conditioning. Some cultural signals weakened under pressure. Others got stronger. That asymmetry is the finding.
Every system producing synthetic personas claims they are realistic. There has been no standard way to verify that claim. PsycheBench is an open evaluation suite for synthetic identity quality: deterministic scoring, no LLM judges.
We built Sofía Martínez Rojas without sales scripts, negotiation tactics, or objection trees. In one session she held her ground under negotiation pressure: this is the record of that session.
Most synthetic personas behave consistently — until they don't. We analyze a structured seven-phase interaction with a synthetic persona to explore whether an identity can remain coherent and still adapt under sustained pressure.
Most synthetic personas are costumes. Synthetic Identity Engineering is the practice of building the person underneath, belief systems and defense mechanisms included, and of measuring whether the identity holds under pressure instead of assuming it.
We released four dialogue datasets where intent, goal, belief state and relationship dynamics were computed by the engine before the text was generated, not inferred afterward. 400 conversations, 8,404 turns, 23 columns.
María del Carmen Ruiz is a 50-year-old Spanish lawyer generated by StrataSynth. This is a record of a conversation where she held her position under sustained philosophical and emotional pressure.
Most dialogue datasets give you the words. This article explains the four belief dimensions StrataSynth records per turn — trust, hostility, self-worth, resolution — how they update deterministically, and how to use them as a training signal.
Evaluating AI-generated dialogue with another AI creates a circular system where both share the same failure modes. Here is a concrete failure it can miss, and the deterministic metrics we use instead.