Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #608

Characterizing web search behavior of conversational LLM agents across four platforms

arXiv:2609.19244v1 reports the first study of agentic Web search across four conversational platforms (ChatGPT, Claude, Grok, DeepSeek), combining real-world user interactions (in vivo) with controlled API experiments (in vitro). The paper analyzes when agents choose to invoke Web search, their query strategies, domain preferences in returned results, and how they transform results into grounded responses, finding substantial variability across platforms, platform-specific result biases, and some reliance on uncited search results.

KEY POINTS

  1. arXiv:2609.19244v1 reports the first study of agentic Web search across four conversational platforms (ChatGPT, Claude, Grok, DeepSeek), combining real-world user interactions (in vivo) with controlled API experiments (in vitro).
  2. The paper analyzes when agents choose to invoke Web search, their query strategies, domain preferences in returned results, and how they transform results into grounded responses, finding substantial variability across platforms, platform-specific result biases, and some reliance on uncited search results.
  3. The findings affect design and evaluation of conversational agents and search tools by revealing platform-dependent search behavior, attribution gaps, and trade-offs between search frequency and response quality.

WHY IT MATTERS

The findings affect design and evaluation of conversational agents and search tools by revealing platform-dependent search behavior, attribution gaps, and trade-offs between search frequency and response quality.

SOURCES & TIMELINE

1