Audit finds 91 "silent failures" in agent-tool interactions within ToolUniverse
Researchers present arXiv:2609.26836v1, an audit mechanism that examined 15 scientific tools integrated in the ToolUniverse environment and identified 91 "silent failures" where tool invocations returned incomplete or missing information without notifying the agent or user. Most failures occurred at the API (51) and wrapper (25) layers—common types were missing fields and inconsistencies in search/filtering/ranking—and the paper proposes a "contextual reliability" concept plus testing, disclosure, and monitoring measures.