Synthetic persona datasets compared: which one fits your project?
PersonaHub, FinePersonas, Nemotron-Personas, PERSONA, Synthetic-Persona-Chat, Twin-2K-500 and StrataSynth Personas each solve a different problem. Here is what each one offers, when to pick it, and what to do when you need people rather than descriptions.
Synthetic personas have become a standard ingredient for building and testing AI: seeding training data, simulating users, checking an assistant for bias before it meets real customers. There are now several excellent open datasets, and they are not interchangeable. Each was built for a different job.
This is a practical guide to the main ones, including ours, so you can pick the right tool for yours.
At a glance
| Size | Anchored to a real population | Languages / countries | What each persona contains | Personality as numbers | License | |
|---|---|---|---|---|---|---|
| PersonaHub | very large | no, built for diversity | English | one descriptive sentence or paragraph | no | CC BY-NC-SA 4.0 (non-commercial) |
| FinePersonas | 42 million | no, built for diversity | English | one paragraph, with topic labels and embeddings | no | Llama 3 license |
| Nemotron-Personas | 1 million per country | yes | US, France, Brazil, Japan, India, Singapore, Korea | demographics, location, several themed biographies, skills, hobbies, career goals | not in the open version | CC BY 4.0 |
| PERSONA | 1,586 | yes, US | English, US | demographic and idiosyncratic attributes, with preference feedback | - | CC BY-NC-SA 4.0 (non-commercial) |
| Synthetic-Persona-Chat | about 11,000 | no | English | a few persona sentences and a conversation between two personas | no | CC BY 4.0 |
| Twin-2K-500 | 2,058 | real people | English, US | about 500 survey answers from real participants | real measured responses | CC BY 4.0 |
| StrataSynth Personas | 10,000 | yes | Spain, Mexico, US, UK, Germany, France, Italy, Brazil, each in its own language | demographics, occupation, income, household, daily routine, speaking style, hobbies, skills, goals | Big Five and adult attachment, 0 to 1 | CC BY 4.0 |
When each one shines
PersonaHub and FinePersonas: breadth. When you need millions of different points of view to diversify synthetic instructions, math problems or domain text, nothing beats their scale. A persona is a short description, which is exactly what you want when it is a seed for a prompt.
Nemotron-Personas: scale with real demographics. A million people per country, anchored to official statistics, with rich themed biographies. If your users are in the United States, Japan or India, it is a superb starting point.
PERSONA: pluralistic alignment. Built to test whether a model can serve people with different values and preferences, with hundreds of thousands of preference judgements. A research testbed first.
Synthetic-Persona-Chat: persona-grounded dialogue. Short persona facts plus the conversation they produce. Ideal for training a model to stay consistent with a persona in chat.
Twin-2K-500: ground truth. Real people answering hundreds of questions. It is not synthetic, and that is its value: it is what you compare a simulation against.
Where StrataSynth fits
First, how ours is made. Every StrataSynth persona is generated by our own engine, the StrataSynth Humans Engine. We did not use any of the datasets above to build ours: no records, no text, no seeds. We read them only to compare, which is what this post is.
StrataSynth Personas is the open dataset layer of that larger system. The same engine sits underneath our products and the people we generate:
- StrataSynth Personas: an open dataset of population-grounded synthetic people, ready to use in AI training, evaluation and user simulation.
- StrataSynth Humans Engine: the engine that generates those people and keeps them consistent across conversations and scenarios.
- QualiSynth: qualitative research built on the engine, where synthetic consumers can be interviewed and compared before real fieldwork.
- ArenaSynth: practice for the conversations that matter, where professionals rehearse negotiations, objections and difficult decisions against synthetic counterparts.
The dataset itself was built for one job: testing and training agents that will talk to real customers in Europe and the Americas, in their own language.
- Eight countries, each in its own language. Spanish for Spain and Mexico, German, French, Italian, Portuguese for Brazil and English for the US and the UK, not translated from English. Spain, Mexico, the UK, Germany and Italy are not covered by the other population-anchored datasets above.
- Real populations, down to the household. Age, region, education, household income, employment, living arrangements and family structure follow each country’s real adult population, and in Spain, France, the US, Mexico and Brazil so do occupations, named the way people say them there. Names follow real frequencies by generation, and Spanish, Mexican and Brazilian people carry two surnames.
- Personality as numbers, and visible in the text. Every person has Big Five and adult attachment scores, and the biography, routine and hobbies are written from them. You can filter by personality and get people who read differently.
- Ready to simulate a user. Each person comes with a speaking style: tone, length, vocabulary, the phrases they reach for. Put a biography and a speaking style in a prompt and you have a customer who talks like a specific person from a specific place.
- Free for commercial use under CC BY 4.0.
From persona data to people you can talk to
A static persona is useful on its own. Many projects need something more: a person who responds to a situation, keeps a position, remembers who they are from one conversation to the next, and whose reasons you can inspect.
That is what the Humans Engine does. It models the person underneath the words: identity, personality, beliefs, goals and the state of the relationship are kept explicitly, outside the text, and the engine sets what the person intends and wants before any language is generated. The same person can be reused across interactions instead of being reinvented with every prompt.
In practice that opens several uses:
- Research. Interview synthetic consumers before committing to fieldwork.
- Agent testing. Put a conversational system in front of users who disagree, escalate or change their minds.
- Training. Rehearse negotiations, objections and difficult conversations against a consistent counterpart.
- Synthetic data. Generate conversations with the intent, goal and belief state recorded alongside the text.
Choosing the right layer
There is no single best synthetic persona dataset. The right choice depends on what you need to do.
- Millions of short persona descriptions? PersonaHub or FinePersonas.
- Population-grounded people at very large scale? Nemotron-Personas is designed for that.
- Persona-grounded conversational training data? Synthetic-Persona-Chat is a natural fit.
- Real human responses to compare against? Twin-2K-500 gives you something no synthetic dataset can: answers from real participants.
- Population-grounded people in European and American markets, in their own languages, with explicit personality? StrataSynth Personas is designed for that.
- People you can interview, challenge or simulate repeatedly? That is where the StrataSynth Humans Engine comes in.
These datasets also work well together. Use a broad one to diversify prompts, a population-anchored one to make your test users look like your market, real human data to check that the simulation holds up, and an interactive engine when you need behaviour rather than descriptions.
Go beyond the dataset
Want to see how these people behave, rather than only read about them? Describe who you want to understand. Meet them. Challenge them.
- Talk to a synthetic person now, in the browser, no account needed.
- QualiSynth: pressure-test your qualitative research before you fund the real fieldwork.
- ArenaSynth: rehearse negotiations, objections and difficult conversations against synthetic counterparts.
- Ask for a dataset built for your population or scenario: tell us what you are trying to build.
The open dataset: StrataSynth/stratasynth-personas-8-countries.
The datasets in this comparison
- PersonaHub: proj-persona/PersonaHub
- FinePersonas: argilla/FinePersonas-v0.1
- Nemotron-Personas: nvidia/Nemotron-Personas-USA and its regional versions
- PERSONA: SynthLabsAI/PERSONA
- Synthetic-Persona-Chat: google/Synthetic-Persona-Chat
- Twin-2K-500: LLM-Digital-Twin/Twin-2K-500
Details as published on each dataset’s page in September 2026.