NEWS · RESEARCH · #350
Query-Aware Source-Risk Triage for Retrieval-Augmented Generation (arXiv:2609.16564v1)
The paper proposes a pre-generation triage layer for retrieval-augmented generation that treats source-query relationships as query-dependent; it routes canonical query families and labels retrieved pages to pass, contextualize, exclude, or review. The method combines a four-dimension page score, rank-discounted family aggregation, intent-preserving query mutations, and a family-held-out router, calibrated on a 200-URL pilot and analyzed on a synthetic 20,000-row scenario with an oracle page gate to define risk-coverage targets, while noting limits on annotation reliability and real retrieval dynamics.
KEY POINTS
- The paper proposes a pre-generation triage layer for retrieval-augmented generation that treats source-query relationships as query-dependent; it routes canonical query families and labels retrieved pages to pass, contextualize, exclude, or review.
- The method combines a four-dimension page score, rank-discounted family aggregation, intent-preserving query mutations, and a family-held-out router, calibrated on a 200-URL pilot and analyzed on a synthetic 20,000-row scenario with an oracle page gate to define risk-coverage targets, while noting limits on annotation reliability and real retrieval dynamics.
- It offers an auditable, query-dependent method to prioritize which retrieved sources need review in RAG pipelines, informing safer and more precise review workflows.
WHY IT MATTERS
It offers an auditable, query-dependent method to prioritize which retrieved sources need review in RAG pipelines, informing safer and more precise review workflows.