Home · Datasets · ShareGPT
DATASET

ShareGPT

Community-shared ChatGPT conversations used to train instruction-following open models.

TARGET QUERY sharegpt dataset · ~3K/mo
SIZE
90K conversations
CREATOR
Community
MODALITY
text
LICENSE
Various
RELEASED
2023-03
OVERVIEW Updated 2026-05-17

Overview

ShareGPT is a community-collected dataset of real ChatGPT conversations shared by users through browser extensions. It provided the first large-scale collection of how people actually interact with chat models in practice.

What’s In It

The dataset contains approximately 90,000 multi-turn conversations between users and ChatGPT, shared voluntarily. Conversations cover a wide range of topics from coding to creative writing. Length and complexity vary significantly.

How It’s Used

ShareGPT data was central to training Vicuna, one of the first open models to closely match ChatGPT quality. The real user conversations provide more naturalistic training signal than synthetic instruction datasets.

Controversies

Users who shared conversations may not have fully consented to training use. Some conversations contain personal information. Using ChatGPT outputs to train competing models potentially violates OpenAI’s terms of service.