Home · Datasets · Open Assistant Conversations
DATASET

Open Assistant Conversations

Human-generated and ranked conversation dataset for training open-source chat assistants.

TARGET QUERY open assistant dataset · ~3K/mo
SIZE
161K messages across 66K trees
CREATOR
LAION
MODALITY
text
LICENSE
Apache 2.0
RELEASED
2023-04
OVERVIEW Updated 2026-05-17

Overview

Open Assistant Conversations is a crowdsourced dataset of human-generated and ranked conversation trees created by LAION’s Open Assistant project. It represents one of the first large-scale efforts to produce human-preference data without relying on synthetic generation.

What’s In It

The dataset contains 161,443 messages in 66,497 conversation trees across 35 languages. Each message was written and ranked by human volunteers for quality, toxicity, and helpfulness. The tree structure captures multiple response options at each turn.

How It’s Used

Open Assistant data is used for supervised fine-tuning and reward model training. Its human rankings enable DPO and RLHF training without expensive API calls. The multilingual coverage makes it valuable for non-English chat model development.

Controversies

Crowdsourced data quality varies significantly. The volunteer base skews toward technical users, potentially missing everyday user interaction patterns. Some conversations were flagged for harmful content that slipped through moderation.