Characterizing web search behavior of conversational LLM agents across four platforms
arXiv:2609.19244v1 reports the first study of agentic Web search across four conversational platforms (ChatGPT, Claude, Grok, DeepSeek), combining real-world user interactions (in vivo) with controlled API experiments (in vitro). The paper analyzes when agents choose to invoke Web search, their query strategies, domain preferences in returned results, and how they transform results into grounded responses, finding substantial variability across platforms, platform-specific result biases, and some reliance on uncited search results.