Overview
UltraChat is a large-scale multi-turn dialogue dataset created by Tsinghua University researchers containing 1.5 million diverse conversations generated systematically using ChatGPT for training conversational AI assistants.
What’s In It
The dataset contains 1.5 million multi-turn dialogues covering questions about the world, creative writing tasks, and assistance with existing materials like summarization. Each conversation involves multiple turns with diverse topics drawn from a curated seed set.
How It’s Used
UltraChat became popular for training open instruction-tuned models. Zephyr-7B was trained on UltraChat-200K, a filtered subset. The dataset is valued for its dialogue diversity and multi-turn coherence.
Controversies
As a synthetically generated dataset, UltraChat inherits ChatGPT’s biases. Using GPT-generated data to train competing models raises questions about model collapse and whether synthetic data adds novel capabilities beyond the teacher model.