Loading
Language models' behaviors are set during post-training, but probing them requires more than prompting. Persona vectors can reveal what models express, hide, or resist. This changes how we audit open-weight LLMs.
“arXiv:2607.13162v3 Announce Type: replace-cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vector…”
Read the source →STATUS
ACTIVE
CATEGORY
Models
EVIDENCE
Not yet assessed
ENTITY
Persona Vectors, LLMs, arXiv:2607.13162v3
DECISION
Automated · no editorial override
LAST OBSERVED
Aug 3, 2026