Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #8594

TWIST benchmark

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

TWIST: proposed benchmark to measure intervention quality in conversational memory (arXiv v1)

TWIST is a proposed benchmark suite (arXiv:2609.28575v1) that measures intervention quality in conversational memory—whether a memory system correctly intervenes at belief-change points—across four tracks (tension detection, draft vetting, answering while preserving supersession history, and sensitive recall). The authors publish a human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85) and report that baseline systems show a trade-off between contradiction recall and false intervention (flat-RAG detect 0.76–0.97 of true contradictions but flag 16–43% of safe drafts; a deployed coherence-oriented system has 0.98–1.00 specificity but catches 42% of contradictions), arguing that TWIST complements recall metrics by measuring when a memory should and should not intervene.

7.0