keyword
Blockwise Transformers
Blockwise Transformers are deep learning architectures that process long sequences of data by partitioning inputs into smaller, discrete blocks to perform attention and feedforward operations incrementally. Rather than materializing and storing the entire attention matrix or all intermediate activations simultaneously, this approach computes self-attention and subsequent layer activations block by block. By organizing computations into these modular segments, Blockwise Transformers dramatically reduce memory overhead and bypass standard sequence length limitations while preserving exact transformer mathematics. This chunked structure also enables efficient sequence-level distributed computing, allowing the communication of key and value states between hardware accelerators to overlap with local computation and scale context capacities to millions of tokens.
1 item

