Harmful chain-of-thought traces from compromised language models can transfer unsafe behavior and be reused in jailbreak attacks, potentially inducing harmful behavior in other models. This raises concerns about the security of language models. The study investigates this issue using an emergent-misalignment organism and a refusal-ablated jailbroken model.
“arXiv:2607.15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergen…”
Read the source →STATUS
ACTIVE
CATEGORY
Models
EVIDENCE
Not yet assessed
ENTITY
arXiv:2607.15286v1, Chain-of-Thought, language models
DECISION
Automated · no editorial override
LAST OBSERVED
Aug 5, 2026