#ToxicChat
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation #MachineLearning #AI
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
dlvr.it
March 5, 2024 at 6:40 AM
August 20, 2025 at 11:27 PM
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, Jingbo Shang
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation. (arXiv:2310.17389v1 [cs.CL])
http://arxiv.org/abs/2310.17389
October 27, 2023 at 2:09 AM
activating steering primarily for harmful inputs. Experiments using safety benchmarks like ToxicChat & In-The-Wild Jailbreak Prompts demonstrate that our weighted steering controller significantly increases refusal rates compared to the base LLM, [5/7 of https://arxiv.org/abs/2505.20309v1]
May 28, 2025 at 5:54 AM
Lots of people evaluated gpt-oss-safeguard as well, here's the code and research notes from David Gros trying to replicate its performance on the ToxicChat dataset

github.com/roostorg/mod...
model-community/hackathon_demos/open-safeguard-dec-2025/gpt_oss_utils_dactile at main · roostorg/model-community
Share evaluation outcomes and implementation tips for using open safety models in Trust & Safety workflows - roostorg/model-community
github.com
December 22, 2025 at 3:53 PM