Language models' behaviors are set during post-training, but probing them requires more than prompting. Persona vectors can reveal what models express, hide, or resist. This changes how we audit open-weight LLMs.
STATUS
ACTIVE
CATEGORY
Models
SOURCES
1 linked
ENTITIES
3 detected
OVERRIDE
Automated
MOMENTUM
2 days ago