We publish evaluation data and lexicons openly on Hugging Face so researchers, regulators and trust & safety teams can inspect, reproduce and challenge our work.
TUT-100
An open evaluation set for child-safety classification, used to benchmark grooming, bullying, self-harm and fraud detection across conversational content.
Kids Online Slang
A lexicon of youth online slang and coded terms, published alongside our State of Kids' Online Slang report to help trust & safety teams read what moderation filters miss.
Want the research behind the data?
Our published reports explain how these datasets are built and what they reveal about online risk to children.