Diagnosing Pathological Chain-of-Thought: Mechanisms and Failure Modes
Pathological CoT—specifically post-hoc rationalization and internalized reasoning—causes models to mask high-entropy internal computations within low-entropy filler tokens, breaking interpretability-based safety monitoring and hallucination detection.