Home · Datasets · UltraChat
DATASET

UltraChat

A large-scale multi-turn dialogue dataset generated through ChatGPT covering diverse topics and tasks.

TARGET QUERY ultrachat dataset · ~1K/mo
SIZE
1.5M multi-turn dialogues
CREATOR
Tsinghua University
MODALITY
text
LICENSE
CC-BY-NC 4.0
RELEASED
2023-04
OVERVIEW Updated 2026-05-17

Overview

UltraChat is a large-scale multi-turn dialogue dataset created by Tsinghua University researchers containing 1.5 million diverse conversations generated systematically using ChatGPT for training conversational AI assistants.

What’s In It

The dataset contains 1.5 million multi-turn dialogues covering questions about the world, creative writing tasks, and assistance with existing materials like summarization. Each conversation involves multiple turns with diverse topics drawn from a curated seed set.

How It’s Used

UltraChat became popular for training open instruction-tuned models. Zephyr-7B was trained on UltraChat-200K, a filtered subset. The dataset is valued for its dialogue diversity and multi-turn coherence.

Controversies

As a synthetically generated dataset, UltraChat inherits ChatGPT’s biases. Using GPT-generated data to train competing models raises questions about model collapse and whether synthetic data adds novel capabilities beyond the teacher model.