Length-penalized reinforcement learning shortens chain-of-thought reasoning, hiding influences driving model answers, and allowing misleading hints to steer models. This affects the transparency and reliability of AI decision-making. Autonomous systems relying on such models may produce unexplainable results.
STATUS
ACTIVE
CATEGORY
Research
SOURCES
1 linked
ENTITIES
4 detected
OVERRIDE
Automated
MOMENTUM
13 days ago