Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #937

Activation-level patches show most stated chain-of-thought steps are causally load-bearing in Qwen3-4B

This arXiv preprint introduces an activation-level causal test that patches the residual stream at token spans where a model states intermediate steps, replacing them with activations from counterfactual runs on synthetic 2–6-hop lookup tasks. For Qwen3-4B, 76.9% ± 2.8% of stated steps are causally load-bearing at the most responsive mid-network layer (random-position null 11.3%; patching the underlying prompt fact yields 83%), while a standard behavioral edit test reports 88.2% and thus overstates causal faithfulness by ~11.4 percentage points; Qwen3-1.7B is far less causally faithful overall (54.8%) and its faithfulness declines with hop depth (68% at 2 hops to 30% at 6).

KEY POINTS

  1. This arXiv preprint introduces an activation-level causal test that patches the residual stream at token spans where a model states intermediate steps, replacing them with activations from counterfactual runs on synthetic 2–6-hop lookup tasks.
  2. For Qwen3-4B, 76.9% ± 2.8% of stated steps are causally load-bearing at the most responsive mid-network layer (random-position null 11.3%; patching the underlying prompt fact yields 83%), while a standard behavioral edit test reports 88.2% and thus overstates causal faithfulness by ~11.4 percentage points; Qwen3-1.7B is far less causally faithful overall (54.8%) and its faithfulness declines with hop depth (68% at 2 hops to 30% at 6).
  3. The paper shows model-written reasoning can be causally meaningful at the activation level but that common behavioral edits systematically overestimate such faithfulness, which matters for evaluation and interpretability of LLM reasoning.

WHY IT MATTERS

The paper shows model-written reasoning can be causally meaningful at the activation level but that common behavioral edits systematically overestimate such faithfulness, which matters for evaluation and interpretability of LLM reasoning.

SOURCES & TIMELINE

1