Evaluating large language models in conversations is expensive and key events like jailbreaks emerge late. Dynamic budget allocation can help. This method could make LLM evaluation more efficient, but we don't know yet if it holds up outside the lab.
STATUS
ACTIVE
CATEGORY
Models
SOURCES
1 linked
ENTITIES
2 detected
OVERRIDE
Automated
MOMENTUM
4 hours ago