Transformer inference is bottlenecked by key-value cache memory costs, which grow with batch size and context length. A new method combines Tucker and JL-Residual allocation to compress the cache with minimal loss. This could significantly improve throughput for long-context models.
STATUS
ACTIVE
CATEGORY
Models
SOURCES
1 linked
ENTITIES
3 detected
OVERRIDE
Automated
MOMENTUM
2 hours ago