Tech Meridian ← LIVE FEED
RU

NEWS · RESEARCH · #142

Audit of 777K conversations shows how user mistreatment of conversational AI occurs and varies

The paper audits 777K English conversations from LMSYS-Chat-1M using two detectors — an eight-category lexicon targeting hostility at the assistant and the dataset's moderation signal — and finds they capture different, weakly overlapping phenomena. The lexicon flags insults, threats, and jailbreak coercion aimed at models while moderation labels more often mark toxic-content solicitation; together they flag about 5% of user turns, with a precision-adjusted estimate of model-directed mistreatment at about 0.90% (noted as an evaluation-arena rate). Hostility varies roughly 13-fold across models driven mainly by which users models attract rather than model behaviour, first-turn hostility spreads farther than post-response hostility, assistant apologies are associated with higher odds of next-turn hostility within conversations (but more apologetic models receive less hostility overall), and hostility shows temporal patterns (coercive openings vs. affective accumulation). The authors release the lexicon, cross-validation pipeline, and derived tables.

KEY POINTS

  1. The paper audits 777K English conversations from LMSYS-Chat-1M using two detectors — an eight-category lexicon targeting hostility at the assistant and the dataset's moderation signal — and finds they capture different, weakly overlapping phenomena.
  2. The lexicon flags insults, threats, and jailbreak coercion aimed at models while moderation labels more often mark toxic-content solicitation; together they flag about 5% of user turns, with a precision-adjusted estimate of model-directed mistreatment at about 0.90% (noted as an evaluation-arena rate).
  3. Hostility varies roughly 13-fold across models driven mainly by which users models attract rather than model behaviour, first-turn hostility spreads farther than post-response hostility, assistant apologies are associated with higher odds of next-turn hostility within conversations (but more apologetic models receive less hostility overall), and hostility shows temporal patterns (coercive openings vs.

WHY IT MATTERS

This quantifies how user-directed hostility and coercion appear in large evaluation chats, which affects interpretation of model behaviour, alignment assessments, and deployment/moderation strategies.

SOURCES & TIMELINE

1