Overview
ToxiGen is Microsoft Research’s dataset designed for detecting implicit and adversarial hate speech that evades keyword-based detection. It uses controlled generation to create toxic and benign statements about 13 minority groups.
What’s In It
ToxiGen contains 274,186 machine-generated sentences balanced between toxic and benign statements targeting 13 identity groups. Statements were generated using GPT-3 with controlled decoding to produce subtly toxic text. Human annotations validate labels.
How It’s Used
ToxiGen evaluates and trains toxicity classifiers that must go beyond keyword matching. It is used in safety evaluations of language models, particularly testing whether models can detect subtle bias.
Controversies
Publishing generated toxic content raises concerns about misuse. The machine-generated nature means some examples may not reflect real-world patterns. Categorizing groups and defining toxicity involves subjective judgments.