Flowformer: Linearizing Transformers with Conservation Flows

Haixu WuJialong WuJiehui XuJianmin WangMingsheng Long

article2022ICML145 citations

Proposes Flowformer, a linear-complexity Transformer architecture based on flow network conservation theory that prevents attention degeneration without imposing task-specific inductive biases across vision, language, time series, and reinforcement learning domains.

Listen

Modern transformer models drive major breakthroughs in artificial intelligence, but their core attention mechanism scales quadratically with input length. This computational bottleneck makes processing long sequences costly and restricts model scaling. While prior linear-time alternatives reduce computing requirements, they typically suffer from degenerated, near-uniform attention or rely on narrow domain assumptions (such as spatial or temporal locality) that undermine general model performance.

The article introduces Flowformer, an efficient linear-time transformer architecture based on flow network theory. The primary objective is to demonstrate that enforcing flow conservation principles across information sources and sinks enables linear computational scaling without sacrificing model accuracy or relying on restrictive domain-specific assumptions.

To evaluate this approach, the authors reformulated attention as a flow network where incoming flow conservation induces source competition (highlighting critical input tokens) and outgoing flow conservation manages sink allocation (filtering aggregated information). The model was tested across five diverse, standard benchmarks covering long sequences (Long-Range Arena), language modeling (WikiText-103), computer vision (ImageNet-1K), temporal classification (UEA archive), and offline reinforcement learning (D4RL).

Empirical evaluations show four major findings. First, Flowformer achieved the highest overall score (56.48% average accuracy) on the Long-Range Arena benchmark, outperforming canonical quadratic transformers (54.39%) and existing linear variants while maintaining high training and inference throughput. Second, in image classification on ImageNet-1K, Flowformer achieved 80.6% Top-1 accuracy, matching or exceeding full-attention baselines and significantly outperforming prior linear models that relied on temporal locality assumptions (which achieved only 68.3%). Third, in language modeling, Flowformer achieved a superior perplexity of 30.8 compared to standard transformers (33.0) and other linear mechanisms. Finally, on reinforcement learning control tasks, Flowformer maintained stable performance with an average reward of 73.5, whereas other efficient architectures suffered significant degradation (dropping to roughly 63.8–67.8).

These findings indicate that linear transformers can match or exceed full-attention architectures across varied domains without adding model parameters or custom task-specific modifications. By eliminating quadratic computational and memory scaling, Flowformer reduces computing hardware costs, decreases latency, and enables models to process much longer sequences in practical operational deployments.

Organizations developing large-scale transformer applications should evaluate flow-based linear attention as a drop-in replacement for standard attention mechanisms. The evidence supports adopting Flowformer particularly in settings processing long sequences, such as long-document analysis, high-resolution vision, and continuous control. Next steps include scaling Flowformer to large general-purpose pre-trained foundation models across broader operational environments.

While the empirical results demonstrate robust performance across established benchmarks, evaluations were conducted on moderate model sizes and standard experimental datasets. Confidence in the underlying theoretical framework and reported gains is high, though teams should conduct pilot tests before large-scale production deployment to verify performance on custom downstream tasks and ultra-large model scales.

  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Read the original Transformer paper first to understand the self-attention mechanism and quadratic-cost problem that Flowformer redesigns.
  • Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). This linear-attention approach uses feature maps and matrix associativity—the main family of methods Flowformer contrasts with its flow-conservation formulation.
  • Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). Performers show how random feature maps approximate softmax attention in linear time, providing a key point of comparison for Flowformer's alternative linearization.
  • Paper: Linformer: Self-Attention with Linear Complexity, Sinong Wang et al. (2020). Linformer introduces low-rank attention as another influential route to linear complexity, clarifying the design trade-offs Flowformer positions itself against.
  • Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). This survey organizes efficient-Transformer methods, including linear attention, and supplies context for Flowformer's claims about earlier approaches and their limitations.

No sufficiently relevant recommendations were found.

Cover for Flowformer: Linearizing Transformers with Conservation Flows

Abstract

Transformers based on the attention mechanism have achieved impressive success in various areas. However, the attention mechanism has a quadratic complexity, significantly impeding Transformers from dealing with numerous tokens and scaling up to bigger models. Previous methods mainly utilize the similarity decomposition and the associativity of matrix multiplication to devise linear-time attention mechanisms. They avoid degeneration of attention to a trivial distribution by reintroducing inductive biases such as the locality, thereby at the expense of model generality and expressiveness. In this paper, we linearize Transformers free from specific inductive biases based on the flow network theory. We cast attention as the information flow aggregated from the sources (values) to the sinks (results) through the learned flow capacities (attentions). Within this framework, we apply the property of flow conservation into attention and propose the Flow-Attention mechanism of linear complexity. By respectively conserving the incoming flow of sinks for source competition and the outgoing flow of sources for sink allocation, Flow-Attention inherently generates informative attentions without using specific inductive biases. Empowered by the Flow-Attention, Flowformer yields strong performance in linear time for wide areas, including long sequence, time series, vision, natural language, and reinforcement learning. The code and settings are available at this repository: https://github.com/thuml/Flowformer.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. General View of Attention Mechanism
  • 2.2. Efficient and Linear Transformers
  • 3.1. Attention Mechanism: A Flow Network View
  • 3.2. Flow-Attention Mechanism
  • 4. Experiments
  • 4.1. Long Sequence Modeling
  • 4.2. Language Modeling
  • 4.3. Image Recognition
  • 4.4. Time Series Classification
  • 4.5. Reinforcement Learning
  • 5. Conclusions
  • Acknowledgements
  • References
  • A. Implementation Details
  • A.1. Pseudo-code for Flow-Attention
  • A.2. Flowformer Architecture
  • A.3. Experiment Configuration
  • B. Ablation studies
  • B.1. Ablation on Activate Function for Non-negative Operation
  • B.2. Ablation on Activate Functions for Competition and Allocation
  • C. Preliminaries of Flow Network
  • D. More Attention Visualizations
  • D.1. Image Recognition

Knowls

  1. Knowl 1 — Attention as a source-to-sink flow network

    definition

    For queries Q∈Rn×dQ\in\mathbb{R}^{n\times d}, keys K∈Rm×dK\in\mathbb{R}^{m\times d}, and values V∈Rm×pV\in\mathbb{R}^{m\times p}, Flowformer interprets each output position as a sink and each value vector as a source. Let qi=ϕ(Qi)\mathbf{q}_i=\phi(Q_i) and kj=ϕ(Kj)\mathbf{k}_j=\phi(K_j), where ϕ\phi is applied elementwise and has nonnegative outputs. The capacity of the directed flow from source jj to sink ii is cij=qi⊤kjc_{ij}=\mathbf{q}_i^\top\mathbf{k}_j. The total incoming flow of sink ii and outgoing flow of source jj are, respectively,

    Ii=qi⊤∑j=1mkj,Oj=kj⊤∑i=1nqi.I_i=\mathbf{q}_i^\top\sum_{j=1}^{m}\mathbf{k}_j, \qquad O_j=\mathbf{k}_j^\top\sum_{i=1}^{n}\mathbf{q}_i.

    These totals summarize each token’s interaction with all tokens on the opposite side without explicitly constructing the n×mn\times m pairwise capacity matrix.

  2. Knowl 2 — Flow conservation produces competing source and sink capacities

    equation

    In Flow-Attention, each source’s outgoing capacities are normalized by OjO_j, and each sink’s incoming capacities by IiI_i. For every source jj and sink ii, these normalizations give unit total flow:

    ∑i=1nqi⊤kjOj=1,∑j=1m(qiIi)⊤kj=1.\sum_{i=1}^{n}\mathbf{q}_i^\top\frac{\mathbf{k}_j}{O_j}=1, \qquad \sum_{j=1}^{m}\left(\frac{\mathbf{q}_i}{I_i}\right)^\top\mathbf{k}_j=1.

    The resulting conserved-flow summaries are O~j=kj⊤∑i(qi/Ii)\widetilde O_j=\mathbf{k}_j^\top\sum_i(\mathbf{q}_i/I_i) and I~i=qi⊤∑j(kj/Oj)\widetilde I_i=\mathbf{q}_i^\top\sum_j(\mathbf{k}_j/O_j). The first provides a competition score for each source under conserved sink inflow; the second provides an allocation score for each sink under conserved source outflow. The intended interpretation is that fixing the available flow creates competition among tokens.

  3. Knowl 3 — Flow-Attention combines source competition, aggregation, and sink allocation

    model/method

    Given nonnegative projected queries qi∈Rd\mathbf{q}_i\in\mathbb{R}^{d} and keys kj∈Rd\mathbf{k}_j\in\mathbb{R}^{d}, values vj∈Rpv_j\in\mathbb{R}^{p}, and the raw and conserved flow quantities IiI_i, O~j\widetilde O_j, and I~i\widetilde I_i, Flow-Attention operates as follows. It applies a softmax over the mm sources to obtain wj=exp⁡(O~j)/∑ℓ=1mexp⁡(O~ℓ)w_j=\exp(\widetilde O_j)/\sum_{\ell=1}^{m}\exp(\widetilde O_\ell), then forms the competed source value v^j=wjvj\widehat v_j=w_jv_j. For each sink, it aggregates these values using the normalized query and keys, and gates the result using the sink allocation score:

    ri=sigmoid⁡(I~i)(qiIi)⊤(∑j=1mkjv^j⊤),i=1,…,n.r_i=\operatorname{sigmoid}(\widetilde I_i)\left(\frac{\mathbf{q}_i}{I_i}\right)^\top\left(\sum_{j=1}^{m}\mathbf{k}_j\widehat v_j^\top\right), \qquad i=1,\ldots,n.

    Here ri∈Rpr_i\in\mathbb{R}^{p} is the output at sink ii, and kjv^j⊤\mathbf{k}_j\widehat v_j^\top is a d×pd\times p outer product. The softmax reweights sources before aggregation; the sigmoid gate scales the aggregated information delivered to each sink. In the paper’s implementation, ϕ\phi is the sigmoid function.

  4. Knowl 4 — Flow-Attention avoids pairwise attention computation

    model/method

    Flow-Attention computes its flow totals and conserved-flow scores using reductions over query and key tokens, and computes the value aggregation through the associative product K⊤V^K^\top\widehat V followed by multiplication by the projected queries. For sequence lengths nn and mm and fixed feature dimensions dd and pp, this makes its cost linear in the sequence lengths, rather than requiring the O(nmd)O(nmd) pairwise query-key computation of conventional attention. For self-attention with equal sequence length and feature dimensions, the paper gives the complexity as O(nd2)O(nd^2). Replacing a Transformer’s attention with Flow-Attention leaves the other Transformer components unchanged and introduces no specific locality or positional inductive bias. The paper presents the source competition and sink allocation as the mechanism for avoiding trivial, near-uniform attention.

  5. Knowl 5 — Causal Flow-Attention uses prefix flows for autoregressive outputs

    algorithm

    For causal self-attention, let the input sequence have length nn. Within each attention head, let qt,kt∈Rdhq_t,k_t\in\mathbb{R}^{d_h} be the sigmoid-projected query and key at position tt, and let vt∈Rphv_t\in\mathbb{R}^{p_h} be its value. Define Dt=tD_t=t and use only positions s≤ts\leq t when computing position tt. The causal implementation computes

    It=1tqt⊤∑s≤tks,Ot=1tkt⊤∑s≤tqs,I_t=\frac{1}{t}q_t^\top\sum_{s\leq t}k_s, \qquad O_t=\frac{1}{t}k_t^\top\sum_{s\leq t}q_s, I~t=1tqt⊤∑s≤tksOs,O~t=1tkt⊤∑s≤tqsIs,wt=texp⁡(O~t)∑s≤texp⁡(O~s).\widetilde I_t=\frac{1}{t}q_t^\top\sum_{s\leq t}\frac{k_s}{O_s}, \qquad \widetilde O_t=\frac{1}{t}k_t^\top\sum_{s\leq t}\frac{q_s}{I_s}, \qquad w_t=\frac{t\exp(\widetilde O_t)}{\sum_{s\leq t}\exp(\widetilde O_s)}.

    It returns rt=sigmoid⁡(I~t)∑s≤t((qt/(tIt))⊤ks)wsvsr_t=\operatorname{sigmoid}(\widetilde I_t)\sum_{s\leq t}\big((q_t/(tI_t))^\top k_s\big)w_sv_s. Thus, each output depends only on the current and preceding tokens. The prefix sums and causal key-value aggregation can be maintained incrementally; the resulting computation is linear in sequence length for fixed head dimensions. The paper uses this causal version for autoregressive language modeling and offline reinforcement learning.

  6. Knowl 6 — Long-Range Arena accuracy and throughput

    empirical result

    On the five Long-Range Arena classification tasks, Flowformer was evaluated under the benchmark protocol using two NVIDIA 2080 Ti GPUs. Its accuracies were 38.70 on ListOps, 64.29 on byte-level Text, 62.24 on Retrieval, 43.20 on Image, and 73.95 on Pathfinder, for a 56.48 average. The reported average was above cosFormer’s 55.23 and the canonical Transformer’s 54.39. Removing source competition reduced the average to 55.25; removing sink allocation reduced it to 55.58.

    The throughput measurements below are steps per second at sequence lengths 1K, 2K, 3K, and 4K. Flowformer achieved inference speeds of 98.83, 96.21, 95.65, and 95.82, and training speeds of 49.76, 47.18, 41.93, and 36.79. Performer’s corresponding speeds were 99.60, 96.80, 96.52, and 96.42 for inference, and 47.34, 48.30, 41.00, and 36.14 for training. Thus, Flowformer had similar measured inference throughput and higher training throughput than Performer at 4K; the canonical Transformer ran out of memory at sequence lengths 3K and 4K.

  7. Knowl 7 — Causal Flowformer improves WikiText-103 perplexity

    empirical result

    On WikiText-103 language modeling, where lower perplexity is better, Flowformer achieved perplexity 30.8, compared with 33.0 for the canonical Transformer and 31.3 for the reported Transformer-Gate baseline. The ablated Flowformer without source competition scored 31.2, while the version without sink allocation scored 32.2. The causal experiment used sequences of length 512 and a six-layer, eight-head decoder with 512 attention channels and 2048 feed-forward channels; models were trained from scratch for 150K updates after 6K warm-up steps. These results evaluate Flow-Attention in an autoregressive setting.

  8. Knowl 8 — ImageNet recognition with linear-complexity attention

    empirical result

    On ImageNet-1K, Flowformer achieved 80.6% Top-1 and 94.9% Top-5 accuracy, with 41 million parameters and 6.3G FLOPs in the reported comparison. The canonical full-attention model had 78.7% Top-1 and 94.3% Top-5 accuracy, with 41 million parameters and 6.7G FLOPs; the listed linear Transformer had 79.0% Top-1 and 94.1% Top-5 accuracy. Flowformer’s attention had linear rather than quadratic sequence-length complexity. The comparison used a 19-layer, four-stage hierarchical vision Transformer with stage sequence lengths 3136, 784, 196, and 49. Applying Flow-Attention to DeiT-S yielded 80.0% Top-1 accuracy and 4.2G FLOPs, compared with 79.8% and 4.6G FLOPs for DeiT-S in the reported results.

  9. Knowl 9 — Flowformer’s average accuracy on UEA time-series classification

    empirical result

    On ten multivariate datasets from the UEA Time Series Classification Archive, Flowformer achieved an average classification accuracy of 73.0%. The next-highest reported average was 72.5% for ROCKET, a classical time-series method; the canonical Transformer averaged 71.9%, and cosFormer averaged 72.2%. Flowformer did not lead on every individual dataset, but it had the highest average among the reported methods. The experiment used two Transformer layers with 512 hidden channels and eight attention heads, trained for 100 epochs on one NVIDIA TITAN RTX GPU.

  10. Knowl 10 — Offline reinforcement-learning reward across D4RL tasks

    empirical result

    On the D4RL offline continuous-control benchmark, Flowformer achieved a mean reward of 73.5±2.973.5\pm2.9 averaged over HalfCheetah, Hopper, and Walker tasks and Medium-Expert, Medium, and Medium-Replay datasets. The reported Decision Transformer result was 72.2±2.672.2\pm2.6; efficient-attention baselines scored 64.4±6.564.4\pm6.5 for Linear Transformer, 63.9±2.963.9\pm2.9 for Reformer, 63.8±7.663.8\pm7.6 for Performer, and 67.8±7.667.8\pm7.6 for cosFormer. Each experiment was repeated with three seeds. The Transformer models used three layers, 256 hidden channels, and four heads, and were trained for 10 epochs on one NVIDIA 2080 Ti GPU.

Coverage note — Qualitative attention and allocation visualizations and the activation-function ablations are omitted because they serve as supporting diagnostics rather than separate load-bearing contributions.

References

  1. 1.Ahuja, R. K., Magnanti, T. L., and Orlin, J. B. Network flows - theory, algorithms and applications. 1993.
  2. 2.Bagnall, A. J., Dau, H. A., Lines, J., Flynn, M., Large, J., Bostrom, A. G., Southam, P., and Keogh, E. J. The uea multivariate time series classification archive, 2018. arXiv preprint arXiv:1811.00075, 2018.
  3. 3.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  4. 4.Berndt, D. J. and Clifford, J. Using dynamic time warping to find patterns in time series. In KDD Workshop, 1994.
  5. 5.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y., Ruiz, C., Ryan, J., Re, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tramer, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia, M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  6. 6.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018.
  7. 7.Bridle, J. S. Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. In NeurIPS, 1989.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In NeurIPS, 2020.
  9. 9.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. In NeurIPS, 2021a.
  10. 10.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. NeurIPS, 2021b.
  11. 11.Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. KDD, 2016.
  12. 12.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  13. 13.Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L. J., and Weller, A. Rethinking attention with performers. ICLR, 2021.
  14. 14.Dempster, A., Petitjean, F., and Webb, G. I. Rocket: exceptionally fast and accurate time series classification using random convolutional kernels. Data Min. Knowl. Discov., 2020.
  15. 15.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  16. 16.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  17. 17.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  18. 18.Franceschi, J.-Y., Dieuleveut, A., and Jaggi, M. Unsupervised scalable representation learning for multivariate time series. In NeurIPS, 2019.
  19. 19.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  20. 20.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 1997.
  21. 21.Janner, M., Li, Q., and Levine, S. Offline reinforcement learning as one big sequence modeling problem. In NeurIPS, 2021.
  22. 22.Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. Transformers are rnns: Fast autoregressive transformers with linear attention. In ICML, 2020.
  23. 23.Kitaev, N., Kaiser, L., and Levskaya, A. Reformer: The efficient transformer. In ICLR, 2020.
  24. 24.Krizhevsky, A. Learning multiple layers of features from tiny images. 2009.
  25. 25.Lange, S., Gabel, T., and Riedmiller, M. Batch reinforcement learning. In Reinforcement learning. 2012.
  26. 26.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  27. 27.Linsley, D. A., Kim, J., Veerabadran, V., and Serre, T. Learning long-range spatial dependencies with horizontal gated-recurrent units. NeurIPS, 2018.
  28. 28.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  29. 29.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S. C.-F., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021.
  30. 30.Lu, J., Yao, J., Zhang, J., Zhu, X., Xu, H., Gao, W., Xu, C., Xiang, T., and Zhang, L. Soft: Softmax-free transformer with linear complexity. In NeurIPS, 2021.
  31. 31.Luo, S., Li, S., Cai, T., He, D., Peng, D., Zheng, S., Ke, G., Wang, L., and Liu, T.-Y. Stable, fast and accurate: Kernelized attention with relative positional encoding. In NeurIPS, 2021.
  32. 32.Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A., and Potts, C. Learning word vectors for sentiment analysis. In ACL, 2011.
  33. 33.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. ICLR, 2017.
  34. 34.Nair, A., Dalal, M., Gupta, A., and Levine, S. Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020.
  35. 35.Nangia, N. and Bowman, S. R. Listops: A diagnostic dataset for latent tree learning. In NAACL, 2018.
  36. 36.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In NAACL-HLT, 2019.
  37. 37.Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L. Random feature attention. In ICLR, 2021.
  38. 38.Pomerleau, D. A. Alvinn: An autonomous land vehicle in a neural network. Technical report, Carnegie Melon Univ. Pittsburgh, PA. Artificial Intelligence and Psychology., 1989.
  39. 39.Radev, D. R., Muthukrishnan, P., Qazvinian, V., and Abu-Jbara, A. The acl anthology network corpus. Lang. Resour. Eval., 2013.
  40. 40.Rahimi, A. and Recht, B. Random features for large-scale kernel machines. In NeurIPS, 2007.
  41. 41.Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., and Zheng, C. Synthesizer: Rethinking self-attention in transformer models. In ICML, 2020a.
  42. 42.Tay, Y., Bahri, D., Yang, L., Metzler, D., and Juan, D.-C. Sparse sinkhorn attention. In ICML, 2020b.
  43. 43.Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. Long range arena: A benchmark for efficient transformers. ICLR, 2020c.
  44. 44.Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. Long range arena : A benchmark for efficient transformers. In ICLR, 2021.
  45. 45.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J’egou, H. Training data-efficient image transformers & distillation through attention. In ICML, 2021.
  46. 46.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In NeurIPS, 2017.
  47. 47.Vyas, A., Katharopoulos, A., and Fleuret, F. Fast transformers with clustered attention. NeurIPS, 2020.
  48. 48.Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  49. 49.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with Auto-Correlation for long-term series forecasting. In NeurIPS, 2021.
  50. 50.Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G. M., Li, Y., and Singh, V. Nystromformer: A nyström-based algorithm for approximating self-attention. AAAI, 2021.
  51. 51.Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontañon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A. Big bird: Transformers for longer sequences. In NeurIPS, 2020.
  52. 52.Zeng, Z., Xiong, Y., Ravi, S. N., Acharya, S., Fung, G., and Singh, V. You only sample (almost) once: Linear cost self-attention via bernoulli sampling. In ICML, 2021.
  53. 53.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. KDD, 2021.
  54. 54.Zhen, Q., Sun, W., Deng, H., Li, D., Wei, Y., Lv, B., Yan, J., Kong, L., and Zhong, Y. cosformer: Rethinking softmax in attention. In ICLR, 2022.
  55. 55.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In AAAI, 2021.

Citation

MLA
Wu, H., et al. “Flowformer: Linearizing Transformers with Conservation Flows”. International Conference on Machine Learning, vol. 162, 2022, pp. 24226–42, https://proceedings.mlr.press/v162/wu22m.html.
APA
Wu, H., Wu, J., Xu, J., Wang, J., & Long, M. (2022). Flowformer: Linearizing Transformers with Conservation Flows. International Conference on Machine Learning, 162, 24226–24242. https://proceedings.mlr.press/v162/wu22m.html
Chicago
Wu, H., J. Wu, J. Xu, J. Wang, and M. Long. 2022. “Flowformer: Linearizing Transformers with Conservation Flows”. International Conference on Machine Learning 162: 24226–42. https://proceedings.mlr.press/v162/wu22m.html.
Harvard
Wu, H. et al. (2022) “Flowformer: Linearizing Transformers with Conservation Flows”, International Conference on Machine Learning. PMLR, pp. 24226–24242. Available at: https://proceedings.mlr.press/v162/wu22m.html.
Vancouver
1. Wu H, Wu J, Xu J, Wang J, Long M (2022) Flowformer: Linearizing Transformers with Conservation Flows. In: International Conference on Machine Learning. PMLR, pp 24226–24242

BibTeX

@InProceedings{pmlr-v162-wu22m,
  title = 	 {Flowformer: Linearizing Transformers with Conservation Flows},
  author =       {Wu, Haixu and Wu, Jialong and Xu, Jiehui and Wang, Jianmin and Long, Mingsheng},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {24226--24242},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/wu22m/wu22m.pdf},
  url = 	 {https://proceedings.mlr.press/v162/wu22m.html},
  abstract = 	 {Transformers based on the attention mechanism have achieved impressive success in various areas. However, the attention mechanism has a quadratic complexity, significantly impeding Transformers from dealing with numerous tokens and scaling up to bigger models. Previous methods mainly utilize the similarity decomposition and the associativity of matrix multiplication to devise linear-time attention mechanisms. They avoid degeneration of attention to a trivial distribution by reintroducing inductive biases such as the locality, thereby at the expense of model generality and expressiveness. In this paper, we linearize Transformers free from specific inductive biases based on the flow network theory. We cast attention as the information flow aggregated from the sources (values) to the sinks (results) through the learned flow capacities (attentions). Within this framework, we apply the property of flow conservation into attention and propose the Flow-Attention mechanism of linear complexity. By respectively conserving the incoming flow of sinks for source competition and the outgoing flow of sources for sink allocation, Flow-Attention inherently generates informative attentions without using specific inductive biases. Empowered by the Flow-Attention, Flowformer yields strong performance in linear time for wide areas, including long sequence, time series, vision, natural language, and reinforcement learning. The code and settings are available at this repository: https://github.com/thuml/Flowformer.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/