Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #605

Closed-World Resolution and Hallucinated-Tools Benchmark (HTB) for LLM tool hallucination

New arXiv paper (arXiv:2609.19425v1) measures and benchmarks 'tool hallucination' in tool-augmented LLM agents, proposing a five-class taxonomy (H1–H5) and a training-free closed-world resolver called the Resolution Rung (registry membership plus signature check). The study reports 322 hallucinations across ten hosted models and two invocation surfaces, finds fabricated-tool calls concentrate on an unconstrained raw-JSON surface, shows model scale does not eliminate the problem, documents additional collision/shadowing risks when merging namespaces via the Model Context Protocol (M1–M5) with 154 measured hallucinations, and releases the versioned Hallucinated-Tools Benchmark (HTB).

KEY POINTS

  1. New arXiv paper (arXiv:2609.19425v1) measures and benchmarks 'tool hallucination' in tool-augmented LLM agents, proposing a five-class taxonomy (H1–H5) and a training-free closed-world resolver called the Resolution Rung (registry membership plus signature check).
  2. The study reports 322 hallucinations across ten hosted models and two invocation surfaces, finds fabricated-tool calls concentrate on an unconstrained raw-JSON surface, shows model scale does not eliminate the problem, documents additional collision/shadowing risks when merging namespaces via the Model Context Protocol (M1–M5) with 154 measured hallucinations, and releases the versioned Hallucinated-Tools Benchmark (HTB).
  3. Shows tool-call hallucinations are a structural blind spot that must be detected by a closed-world resolver before gating, and provides a benchmark (HTB) to compare defenses.

WHY IT MATTERS

Shows tool-call hallucinations are a structural blind spot that must be detected by a closed-world resolver before gating, and provides a benchmark (HTB) to compare defenses.

SOURCES & TIMELINE

1