Contents
Figure 1: A million conversations, zero annotators — if you filter hard enough
The synthetic data fundamentals article listed LLMs as the dominant technique for text-based synthetic data. Conversational data is where that technique has gone furthest. Baize used ChatGPT's self-chat to produce 653,000 multi-turn dialogues. UltraChat used two ChatGPT instances, one playing the user and one the assistant, to build a corpus on the order of a million conversations or more. Neither project hired annotators to write a single dialogue.
The reason this works is also why it needs care. One survey of this literature puts it plainly: post-training effectiveness depends more on data quality than on model scale. A pipeline that generates a million mediocre conversations loses to one that generates fifty thousand good ones and filters hard. This article covers generating conversational datasets, filtering them, and evaluating whether they were worth generating, using the framework from the last several articles in this series.
By the end, you'll know the main generation patterns (self-chat, dual-agent, persona-driven, multi-agent), a working pipeline with a filtering stage, and the failure modes that matter most for dialogue. The Generator-Critic loop from Synthetic-Persona-Chat is the most reusable idea in this whole literature.
What Makes Conversational Data Different
The text augmentation article's finding was that task structure, not generation quality, decided whether augmentation helped. Dialogue adds failure modes that single-sentence text doesn't have:
- Consistency across turns. A persona who says they're a vegetarian in turn two shouldn't order steak in turn nine. Faithfulness to the assigned persona is a named evaluation criterion in the persona-chat literature for exactly this reason.
- Role discipline. In self-chat, the "user" model tends to drift into sounding like an assistant. Real users are terse, vague, and sometimes wrong.
- Realism of the human side. Assistant turns are easy for an LLM to produce. User turns, with their disfluencies, topic jumps, and underspecified requests, are the hard part.
Pattern One: Self-Chat and Dual-Agent Generation
Baize explored two paradigms for self-play dialogue: generating a complete multi-turn dialogue in a single pass, or building it turn by turn by alternating user and assistant roles. UltraChat uses a dual-agent setup, with one ChatGPT instance generating user queries and another producing responses, so no human is involved at any step:
def generate_dialogue(llm, seed_topic, num_turns=6):
history = []
for turn in range(num_turns):
user_msg = llm.generate(
system="You are a user. Write ONE short, natural message. "
"Be informal. Sometimes be vague.",
context=history + [("topic", seed_topic)],
)
history.append(("user", user_msg))
assistant_msg = llm.generate(
system="You are a helpful support assistant.",
context=history,
)
history.append(("assistant", assistant_msg))
return history
The tradeoff is practical. A single pass is cheaper, because it takes one call per dialogue, but the model controls both sides and tends to produce suspiciously cooperative users. Turn-by-turn costs more calls but lets you give the "user" its own instructions, which is where realism is won or lost.
Pattern Two: Persona-Driven Generation
Persona conditioning is the main lever for diversity and specificity. Work on LLM role-play found that persona-based conditioning improves engagement and specificity, and several current frameworks build on it.
DiaSynth splits the process into LLM-proposed subtopic generation, role-conditioned persona generation with semantic filtering, and dialogue generation via chain-of-thought prompts that reason about speaker attributes such as familiarity, emotion, and formality. PSYDIAL generates personality-based dialogues, then routes failures back through regeneration — dialogues that miss a requirement get re-prompted and re-filtered until they pass. Multi-agent group chats (built with frameworks like AutoGen) have an LLM create personas, each with a name, qualities, and speech style, plus a conversation starter, then spawn agents that talk to each other.
The Faker article's distinction applies: Faker fills structured fields, and a persona generator fills behavior. Both help, for different reasons. Faker is still useful here for task-oriented dialogue, where you need plausible names, addresses, and order numbers to slot into templates.
Pattern Three: Seeding From Real Queries
One of the most effective realism tricks is conditioning generation on real queries instead of letting the model invent topics. One pipeline conditioned GPT-4o on queries extracted from natural datasets: assistant commands from an in-house dataset for action-triggering dialogues, and the Natural Questions dataset for information-seeking ones. The result stays closer to the scenarios you care about.
This is the fundamentals article's rule to anchor on real data, applied to prompts: your real query logs, support tickets, or search queries are a better seed than anything the model dreams up unprompted.
Pattern Four: The Generator-Critic Loop
Synthetic-Persona-Chat (20K conversations, roughly 11.8 turns per user) is the clearest example of closing the loop on quality. An LLM generator produces candidate dialogues, and a panel of LLM-based critics scores them on depth, coherence, persona-faithfulness, and toxicity. The generator is then fine-tuned on the high-quality outputs selected by multi-expert voting, and the cycle repeats.
The same logic the diffusion article found "critical" on small datasets applies: generation is cheap, so the quality of your pipeline is mostly the quality of what you throw away.
The Filtering Stage
Filtering is where most of the real work sits. A typical pipeline runs rule-based, statistical, and learned filters over generated dialogues, checking factuality, coherence, diversity, and toxicity:
from sentence_transformers import SentenceTransformer
import numpy as np
embedder = SentenceTransformer("all-MiniLM-L6-v2")
def filter_dialogues(dialogues, min_turns=4, max_turns=20, dedup_threshold=0.92):
# 1. Structural filter
kept = [d for d in dialogues if min_turns <= len(d) <= max_turns]
# 2. Embedding-based deduplication
texts = [" ".join(t[1] for t in d) for d in kept]
vecs = embedder.encode(texts, normalize_embeddings=True)
unique = []
for i, v in enumerate(vecs):
if all(np.dot(v, vecs[j]) < dedup_threshold for j in unique):
unique.append(i)
# 3. Safety and quality screening (toxicity classifier / LLM-as-judge) goes here
return [kept[i] for i in unique]
Install the embedder with pip install sentence-transformers — the all-MiniLM-L6-v2 model is the same one from the sentence transformers article. Embedding deduplication catches the quiet failure of LLM generation: a thousand "different" dialogues that are all the same conversation in different words. Turn-length filtering, toxicity screening, JSON standardization, and LLM-as-a-judge filters for unfaithful or off-topic samples round out the usual stack.
Preference Data for Alignment
Dialogue datasets aren't only for supervised fine-tuning. Preference pairs (a better and a worse response to the same prompt) drive RLHF and DPO, and they can be synthesized too. SynPO is one closed-loop example: generate tasks, produce candidate responses of varying quality, and self-compare to derive preferences. A preference dataset is a reward signal in data form, and the reward design warning applies: a flawed preference generator teaches the model to optimize the wrong thing.
Where This Connects to Distillation, RAG, and Privacy
- Distillation. SODA is literally titled "dialogue distillation." Generating training data from a stronger LLM is black-box distillation, as covered in the knowledge distillation article. Check the terms of service of any model whose outputs you train on, since providers differ on what they permit.
- RAG versus fine-tuning. If your chatbot mainly needs current facts, retrieval often beats fine-tuning on synthetic dialogues. Fine-tuning shapes behavior and tone, and retrieval supplies knowledge.
- Privacy. Synthetic patient-physician dialogues generated from clinical notes are an active application. Generating from real notes risks reproducing them, so the memorization checks and DP considerations from the privacy article apply here too.
- Class imbalance. Rare intents are the dialogue version of the minority class in the SMOTE article. Targeted generation for underrepresented intents is a legitimate use.
Evaluating the Dataset You Generated
For chatbots, the standard three pillars become:
- Utility (TSTR). Fine-tune on synthetic dialogues, then evaluate on a held-out set of real conversations. This is the number that matters.
- Fidelity. Compare turn-length distributions, vocabulary diversity, and topic coverage against real logs. Real corpora like WildChat (about a million real ChatGPT-user interactions) give you a reference distribution.
- Privacy. Check for verbatim leakage of any real source text.
Common Mistakes People Make
- Letting the same model play both sides with no role differentiation. The user turns come out polite and over-informed. Give the user role its own instructions, and persona conditioning where you can.
- Skipping deduplication. Near-duplicate dialogues inflate the dataset without adding information.
- Training on synthetic dialogues alone. Mix in real conversations and keep a real held-out test set — the model collapse discussion in the fundamentals article applies directly.
- Evaluating only on synthetic test data. A model trained and tested on the same generator's style looks great and tells you nothing about real users. Use TSTR.
- Ignoring persona faithfulness. If your pipeline never checks whether a speaker stays consistent across turns, the dataset can teach the model to contradict itself.
A Practical Starting Sequence
- Collect real seed queries from logs, tickets, or public datasets.
- Generate personas and pair them with seeds for diversity.
- Generate dialogues turn by turn, with separate instructions for the user role.
- Filter hard: structure, embedding deduplication, toxicity, and an LLM judge for faithfulness.
- Fine-tune and evaluate with TSTR on real held-out conversations.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Spoken Dialogue Systems | the dialogue-state foundation: why turns, slots, and context tracking shape what a good conversational dataset has to contain. | View on Amazon |
![]() |
Natural Language Processing with Transformers | the model side this data feeds: fine-tuning, preference methods, and how conversational quality gets measured after training. | View on Amazon |
![]() |
Python Text Generation with Transformers | hands-on generation and filtering pipelines, matching the code patterns in this article from self-chat loops to judge-based screening. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
How do you generate multi-turn dialogue data without human annotators?
With LLMs playing both sides. Baize used ChatGPT's self-chat to produce 653,000 multi-turn dialogues, and UltraChat used two ChatGPT instances — one playing the user, one the assistant — to build a corpus on the order of a million conversations, with no annotator writing a single dialogue. The practical tradeoff: a single-pass generation is cheap but produces suspiciously cooperative users, while turn-by-turn generation costs more calls but lets you give the user role its own instructions, which is where realism is won or lost.
What makes synthetic dialogue data different from other synthetic text?
Three failure modes that single-sentence text doesn't have. Consistency across turns — a persona who says they're a vegetarian in turn two shouldn't order steak in turn nine. Role discipline — in self-chat the user model tends to drift into sounding like an assistant, while real users are terse, vague, and sometimes wrong. And realism of the human side: assistant turns are easy for an LLM to produce, but user turns with their disfluencies, topic jumps, and underspecified requests are the hard part.
How do you filter generated conversations?
A typical pipeline runs rule-based, statistical, and learned filters: structural checks (turn counts, JSON standardization), embedding-based deduplication to catch a thousand rewordings of the same conversation, toxicity screening, and an LLM-as-a-judge filter for faithfulness and off-topic samples. Generation is cheap, so the quality of your pipeline is mostly the quality of what you throw away.
How do you evaluate a synthetic dialogue dataset?
Three checks. Utility (TSTR): fine-tune on synthetic dialogues, then evaluate on a held-out set of real conversations — this is the number that matters. Fidelity: compare turn-length distributions, vocabulary diversity, and topic coverage against real logs such as WildChat. Privacy: check for verbatim leakage of any real source text, especially if you generated from real notes or tickets.
Wrapping Up
Conversational synthetic data works because LLMs can play both sides of a dialogue, and it fails when the pipeline treats generation as the hard part. Self-chat and dual-agent setups give you volume, persona conditioning and real-query seeding give you diversity and realism, and the Generator-Critic loop plus a strict filtering stage give you quality. Evaluation through TSTR on real conversations tells you whether any of it helped.
Remember that quality beats quantity, the user side is where realism is hardest, and real data should stay in the mix as both training signal and test set. This article closes the loop on the text side of this series' synthetic data arc: the fundamentals article introduced LLM-based generation in the abstract, the text augmentation article warned about task-dependence, and the NLP synthetic data piece covers the surrounding generation techniques — this is the full pipeline for the case where the data is a conversation.
Now take a sample of real queries from any project you're working on, such as support tickets or search logs, and generate ten persona-driven dialogues from them. Then read them aloud and mark every turn where the user sounds more helpful than a real user would. That count tells you how much work your user-side prompts still need.


