Flowformer: Linearizing Transformers with Conservation Flows
Haixu WuJialong WuJiehui XuJianmin WangMingsheng Long
Proposes Flowformer, a linear-complexity Transformer architecture based on flow network conservation theory that prevents attention degeneration without imposing task-specific inductive biases across vision, language, time series, and reinforcement learning domains.
Modern transformer models drive major breakthroughs in artificial intelligence, but their core attention mechanism scales quadratically with input length. This computational bottleneck makes processing long sequences costly and restricts model scaling. While prior linear-time alternatives reduce computing requirements, they typically suffer from degenerated, near-uniform attention or rely on narrow domain assumptions (such as spatial or temporal locality) that undermine general model performance.
The article introduces Flowformer, an efficient linear-time transformer architecture based on flow network theory. The primary objective is to demonstrate that enforcing flow conservation principles across information sources and sinks enables linear computational scaling without sacrificing model accuracy or relying on restrictive domain-specific assumptions.
To evaluate this approach, the authors reformulated attention as a flow network where incoming flow conservation induces source competition (highlighting critical input tokens) and outgoing flow conservation manages sink allocation (filtering aggregated information). The model was tested across five diverse, standard benchmarks covering long sequences (Long-Range Arena), language modeling (WikiText-103), computer vision (ImageNet-1K), temporal classification (UEA archive), and offline reinforcement learning (D4RL).
Empirical evaluations show four major findings. First, Flowformer achieved the highest overall score (56.48% average accuracy) on the Long-Range Arena benchmark, outperforming canonical quadratic transformers (54.39%) and existing linear variants while maintaining high training and inference throughput. Second, in image classification on ImageNet-1K, Flowformer achieved 80.6% Top-1 accuracy, matching or exceeding full-attention baselines and significantly outperforming prior linear models that relied on temporal locality assumptions (which achieved only 68.3%). Third, in language modeling, Flowformer achieved a superior perplexity of 30.8 compared to standard transformers (33.0) and other linear mechanisms. Finally, on reinforcement learning control tasks, Flowformer maintained stable performance with an average reward of 73.5, whereas other efficient architectures suffered significant degradation (dropping to roughly 63.8–67.8).
These findings indicate that linear transformers can match or exceed full-attention architectures across varied domains without adding model parameters or custom task-specific modifications. By eliminating quadratic computational and memory scaling, Flowformer reduces computing hardware costs, decreases latency, and enables models to process much longer sequences in practical operational deployments.
Organizations developing large-scale transformer applications should evaluate flow-based linear attention as a drop-in replacement for standard attention mechanisms. The evidence supports adopting Flowformer particularly in settings processing long sequences, such as long-document analysis, high-resolution vision, and continuous control. Next steps include scaling Flowformer to large general-purpose pre-trained foundation models across broader operational environments.
While the empirical results demonstrate robust performance across established benchmarks, evaluations were conducted on moderate model sizes and standard experimental datasets. Confidence in the underlying theoretical framework and reported gains is high, though teams should conduct pilot tests before large-scale production deployment to verify performance on custom downstream tasks and ultra-large model scales.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Read the original Transformer paper first to understand the self-attention mechanism and quadratic-cost problem that Flowformer redesigns.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). This linear-attention approach uses feature maps and matrix associativity—the main family of methods Flowformer contrasts with its flow-conservation formulation.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). Performers show how random feature maps approximate softmax attention in linear time, providing a key point of comparison for Flowformer's alternative linearization.
- Paper: Linformer: Self-Attention with Linear Complexity, Sinong Wang et al. (2020). Linformer introduces low-rank attention as another influential route to linear complexity, clarifying the design trade-offs Flowformer positions itself against.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). This survey organizes efficient-Transformer methods, including linear attention, and supplies context for Flowformer's claims about earlier approaches and their limitations.
No sufficiently relevant recommendations were found.
