Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #5504

chain-of-thought (CoT)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

Activation-level patches show most stated chain-of-thought steps are causally load-bearing in Qwen3-4B

This arXiv preprint introduces an activation-level causal test that patches the residual stream at token spans where a model states intermediate steps, replacing them with activations from counterfactual runs on synthetic 2–6-hop lookup tasks. For Qwen3-4B, 76.9% ± 2.8% of stated steps are causally load-bearing at the most responsive mid-network layer (random-position null 11.3%; patching the underlying prompt fact yields 83%), while a standard behavioral edit test reports 88.2% and thus overstates causal faithfulness by ~11.4 percentage points; Qwen3-1.7B is far less causally faithful overall (54.8%) and its faithfulness declines with hop depth (68% at 2 hops to 30% at 6).

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

LogicTrack: auditing LLM reasoning with formal logic solvers

LogicTrack is a neuro-symbolic framework that auto-formalizes each Chain-of-Thought step and verifies them with automated theorem provers, introducing a Solver-Based Backtracking Reward (SBR) to score step-wise logical soundness and guide backtracking tree search at inference time. The authors also use backtracking traces to build supervised fine-tuning data and report improved reasoning-chain verifiability and final-answer pass rates across 8 benchmarks and 7 LLMs (arXiv:2609.21492v1).

6.0