MikeTrendsTrends right now

search

self-pruning transformer

Trends

  1. 1
    New Self-Pruning Transformer Targets Extreme KV-Cache Compression●A Self-Pruning Transformer: Extreme KV-Cache Compression w/Universal AttentionYhnWorldUS Politics61 h ago

    A new arXiv paper describes a self-pruning transformer architecture that achieves extreme KV-cache compression using what its authors call universal attention. The approach would cut memory needed to store key-value caches during inference, a major cost in running large language models. Early discussion among developers focuses on whether the pruning method preserves model quality at high compression rates.