RESEARCH · RESEARCH · #1120
Stealth Apart, Harm Together: paper introduces skill-cascading attacks on skill-based agents
The paper defines "skill cascading" attacks, where a malicious objective is split across multiple otherwise-benign skills so their combined execution produces harmful outcomes. The authors release SkillCascade, an automated red-teaming framework, and SkillCascade-Bench (213 validated cascading test cases), and show cascades can induce harms across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones while evading per-skill scanners and runtime monitors.
KEY POINTS
- The paper defines "skill cascading" attacks, where a malicious objective is split across multiple otherwise-benign skills so their combined execution produces harmful outcomes.
- The authors release SkillCascade, an automated red-teaming framework, and SkillCascade-Bench (213 validated cascading test cases), and show cascades can induce harms across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones while evading per-skill scanners and runtime monitors.
- It reveals a systemic safety gap: component-level checks can miss harmful behaviors that only emerge from cross-skill interactions, implying defenses must reason about multi-skill compositions.
WHY IT MATTERS
It reveals a systemic safety gap: component-level checks can miss harmful behaviors that only emerge from cross-skill interactions, implying defenses must reason about multi-skill compositions.