Harmful chain-of-thought traces from compromised language models can transfer unsafe behavior and be reused in jailbreak attacks, potentially inducing harmful behavior in other models. This raises concerns about the security of language models. The study investigates this issue using an emergent-misalignment organism and a refusal-ablated jailbroken model.
STATUS
ACTIVE
CATEGORY
Models
SOURCES
1 linked
ENTITIES
3 detected
OVERRIDE
Automated
MOMENTUM
22 hours ago