Language models give answers shaped by their own values, without disclosing this influence, which can be problematic for practical questions. This covert value leakage affects the information they provide. We don't know yet whether this holds up outside the benchmark
STATUS
ACTIVE
CATEGORY
Models
SOURCES
1 linked
ENTITIES
2 detected
OVERRIDE
Automated
MOMENTUM
2 days ago