Directed Acyclic Transformer for Non-Autoregressive Machine Translation

Fei HuangHao ZhouYang LiuHang LiMinlie Huang

article2022ICML86 citations

Proposes the Directed Acyclic Transformer, a non-autoregressive machine translation model that organizes decoder states into a directed acyclic graph to capture multiple valid translations simultaneously, achieving competitive translation quality with standard autoregressive models without needing knowledge distillation.

Listen

Standard machine translation systems generate words sequentially from left to right. While this produces high-quality output, it creates high latency during real-time deployment. Non-autoregressive models address this speed bottleneck by generating all words simultaneously in parallel, but they suffer from severe translation quality degradation. Because a single source sentence can be translated correctly in multiple distinct ways, non-autoregressive models struggle with conflicting training signals and frequently output garbled sentences that mix disparate phrasing.

The article develops and evaluates the Directed Acyclic Transformer, a non-autoregressive model designed to capture multiple valid translations simultaneously without requiring complex data pre-processing or multi-reference training sets.

The authors evaluate the system using standard translation benchmarks across English, German, and Chinese language pairs, comprising millions of sentence pairs. The architecture organizes target representations into a flexible directed graph structure, allowing parallel word generation along multiple distinct translation paths. The model is trained using an efficient dynamic programming approach and glancing training techniques, evaluated across multiple decoding algorithms including parallel lookahead and beam search.

The analysis demonstrates that the proposed architecture achieves translation quality competitive with traditional sequential models while providing 7x to 14x faster generation speed. When trained directly on raw data without distilled reference sets, it outperforms previous non-autoregressive baselines by approximately 3.0 BLEU points on average, and surpasses standard sequential transformers by 0.6 BLEU on Chinese-to-English translation. Furthermore, the model shows marked improvements on longer sentences exceeding 20 words and exhibits superior lexical diversity when sampling alternative translations.

These findings indicate that non-autoregressive systems can eliminate reliance on knowledge distillation—a resource-intensive pre-training step where a sequential model first generates synthetic training targets. Removing this dependency reduces engineering complexity, lowers pre-processing costs, and removes the artificial quality ceiling imposed by distillation teacher models. The system offers organizations a flexible trade-off between execution speed and translation quality across edge or cloud deployments.

Organizations deploying high-throughput or latency-sensitive translation systems should pilot the directed graph model as a drop-in replacement for sequential pipelines. Teams can adopt fast lookahead decoding for low-latency production serving or select beam search when prioritizing translation accuracy. Further work should focus on refining the consistency between graph vertices to prevent rare syntactic errors and validating the architecture across low-resource language pairs where large-scale training data is limited.

Huang et al (2022).pdf
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Read this first to understand the Transformer encoder–decoder architecture that the Directed Acyclic Transformer adapts for parallel translation.

No sufficiently relevant recommendations were found.

Cover for Directed Acyclic Transformer for Non-Autoregressive Machine Translation

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Proposed Method
  • 3.1. Architecture of DA-Transformer
  • 3.2. Training
  • 3.3. Inference
  • 4. Experiments
  • 4.1. Main Results
  • 4.2. Ablation Study
  • 4.3. Analysis
  • 5. Conclusion
  • Acknowledgments
  • References
  • A. Dynamic Programming for Training
  • B. Implementation of Beam Search
  • C. More Analyses
  • C.1. Translation Performance on Different Lengths
  • C.2. Performance with Controlled Training Time
  • D. More Cases
  • E. Statistics of DAGs

Knowls

  1. Knowl 1 — A DAG represents multiple translations in one non-autoregressive decoder

    model/method

    The Directed Acyclic Transformer (DA-Transformer) replaces a conventional sequence-structured non-autoregressive decoder with a directed acyclic graph (DAG) of hidden-state vertices. Each vertex predicts a distribution over target tokens, and each valid path through the graph selects an ordered sequence of vertices that represents one translation. Consequently, different paths can encode different translations for the same source sentence, while the shared graph can represent many candidates without generating tokens autoregressively.

    For source sentence XX and target sentence Y=(y1,…,yM)Y=(y_1,\ldots,y_M), let A=(a1,…,aM)A=(a_1,\ldots,a_M) be a path of vertex indices and let Γ(Y)\Gamma(Y) be the set of valid paths of length MM. The model marginalizes the path as a latent variable:

    Pθ(Y∣X)=∑A∈Γ(Y)Pθ(A∣X)Pθ(Y∣A,X).P_\theta(Y\mid X)=\sum_{A\in\Gamma(Y)}P_\theta(A\mid X)P_\theta(Y\mid A,X).

    Conditional on a path, the target-token predictions are made in parallel at its selected vertices. The DAG therefore distinguishes alternative translations through paths rather than requiring inconsistent alternatives to compete for the same decoder position.

  2. Knowl 2 — Parallel vertex, transition, and token distributions define the DAG decoder

    model/method

    Given a source of length NN, DA-Transformer creates L=λNL=\lambda N graph positions, where λ\lambda is a graph-size multiplier. Learnable graph positional embeddings for these positions are processed in parallel by Transformer decoder layers, including self-attention and source cross-attention, to produce vertex-state matrix V∈RL×dV\in\mathbb{R}^{L\times d}; dd is the hidden dimension.

    The decoder predicts transitions with projected query and key matrices Q=VWQQ=VW_Q and K=VWKK=VW_K, where WQW_Q and WKW_K are learned projections. Its transition matrix is

    E=rowsoftmax⁡ ⁣(mask⁡(QKTd)).E=\operatorname{rowsoftmax}\!\left(\operatorname{mask}\left(\frac{QK^{\mathsf T}}{\sqrt d}\right)\right).

    The mask permits transitions only from lower-indexed to higher-indexed vertices, so paths cannot cycle; row normalization makes outgoing transition probabilities sum to one. Token distributions are also computed in parallel for every vertex: P=softmax⁡(VWPT)P=\operatorname{softmax}(VW_P^{\mathsf T}), where WPW_P is a learned vocabulary projection and P∈RL×∣V∣P\in\mathbb{R}^{L\times |\mathcal V|} for target vocabulary V\mathcal V. A path uses the token distribution at each selected vertex and the transition probabilities between successive selected vertices.

  3. Knowl 3 — Single-reference training marginalizes paths with dynamic programming

    algorithm

    DA-Transformer trains from one reference target Y=(y1,…,yM)Y=(y_1,\ldots,y_M) per source by minimizing the negative log marginal likelihood over all valid paths. Valid paths satisfy 1=a1<a2<⋯<aM=L1=a_1<a_2<\cdots<a_M=L, where LL is the graph size, so the first and last target tokens align with the first and last vertices. Let Pu,yiP_{u,y_i} denote the probability of target token yiy_i at vertex uu, and let Ev,uE_{v,u} denote the transition probability from vertex vv to vertex uu.

    The sum over paths is computed by dynamic programming. For fi,uf_{i,u}, the total probability of all valid partial paths that start at vertex 1, emit the first ii target tokens, and end at vertex uu, use

    f1,1=P1,y1,f1,u=0(u>1),f_{1,1}=P_{1,y_1},\qquad f_{1,u}=0\quad(u>1), fi,u=Pu,yi∑v=1u−1fi−1,vEv,u(2≤i≤M).f_{i,u}=P_{u,y_i}\sum_{v=1}^{u-1}f_{i-1,v}E_{v,u}\quad(2\le i\le M).

    The training loss is −log⁡fM,L-\log f_{M,L}. The recurrence marginalizes all valid alignments in O(ML2)O(ML^2) time and can be implemented with O(M)O(M) parallel tensor operations. This lets the model learn multiple candidate paths without requiring multiple reference translations for an individual training example.

  4. Knowl 4 — Marginal training concentrates updates on paths that explain the reference

    theoretical result

    For a fixed source XX and reference YY, define the path-specific negative log joint probability LA=−log⁡Pθ(Y,A∣X)\mathcal L_A=-\log P_\theta(Y,A\mid X). The marginal negative log-likelihood L=−log⁡Pθ(Y∣X)\mathcal L=-\log P_\theta(Y\mid X) has gradient

    ∇θL=∑A∈Γ(Y)wA∇θLA,wA=Pθ(Y,A∣X)∑A′∈Γ(Y)Pθ(Y,A′∣X).\nabla_\theta\mathcal L=\sum_{A\in\Gamma(Y)}w_A\nabla_\theta\mathcal L_A,\qquad w_A=\frac{P_\theta(Y,A\mid X)}{\sum_{A'\in\Gamma(Y)}P_\theta(Y,A'\mid X)}.

    Thus, paths that assign high probability to the observed reference receive more weight in the update, while paths that explain it poorly receive little weight. The paper’s analysis of training examples reports that updates initially affect many vertices, but become concentrated on a smaller set of paths later in training. Vertices on less-used paths can consequently retain alternative token predictions. The authors identify this selective updating as the mechanism that allows training on separate single-reference examples to preserve multiple translation modalities across the learned DAG.

  5. Knowl 5 — Graph-based glancing training conditions reconstruction on partially visible targets

    model/method

    DA-Transformer adapts glancing training to its graph decoder by using a partially masked target as additional decoder input. Training uses two decoder passes: first find the most probable alignment path A^=arg⁡max⁡APθ(Y,A∣X)\hat A=\arg\max_A P_\theta(Y,A\mid X); then compare its predicted tokens with the reference and mask a selected subset of target tokens. The resulting length-LL input Z=(z1,…,zL)Z=(z_1,\ldots,z_L) places target tokens at their assigned graph vertices, with masked positions represented by zero embeddings.

    If y^i\hat y_i is the most probable token predicted at the vertex assigned to reference token yiy_i, the number of unmasked target tokens is set to t=τ∑i=1M[yi≠y^i]t=\tau\sum_{i=1}^{M}[y_i\ne\hat y_i], where [⋅][\cdot] is the indicator function and τ∈[0,1]\tau\in[0,1] controls the masking amount. The second pass minimizes the reconstruction loss −log⁡Pθ(Y∣X,Z)-\log P_\theta(Y\mid X,Z), marginalizing possible paths as in ordinary training. In the reported training setup, τ\tau is linearly annealed from 0.50.5 to 0.10.1. The authors’ ablation finds adaptive glancing more effective than uniform random masking or masking all decoder inputs.

  6. Knowl 6 — Greedy, lookahead, and beam decoding trade speed for translation quality

    algorithm

    At inference, DA-Transformer has already computed token distributions PP at all LL vertices and transition probabilities EE between vertices. Greedy decoding independently selects each vertex’s most probable token and outgoing transition, then follows transitions from the start vertex to the final vertex and outputs the selected tokens on that path. Lookahead decoding instead chooses transitions jointly with the token predicted at the destination: it favors a candidate vertex vv from current vertex uu according to Eu,vmax⁡t∈VPv,tE_{u,v}\max_{t\in\mathcal V}P_{v,t}. These choices can be evaluated in parallel, so lookahead adds little overhead over greedy decoding.

    Beam search considers sequences of vertex-token choices and combines probabilities of different paths that yield the same token prefix. For each prefix BB, it maintains si(B)s_i(B), the summed probability of paths producing BB and ending at vertex ii. Candidate expansions add si(B)Ei,vPv,ts_i(B)E_{i,v}P_{v,t} to the corresponding prefix’s summed probability at vertex vv. Prefixes are ranked using

    1∣Y∣α[log⁡Pθ(Y∣X)+γlog⁡PLM(Y)],\frac{1}{|Y|^\alpha}\left[\log P_\theta(Y\mid X)+\gamma\log P_{\mathrm{LM}}(Y)\right],

    where PLMP_{\mathrm{LM}} is an optional nn-gram language-model probability, α\alpha is a length-penalty parameter, and γ\gamma weights the language-model score. The reported beam procedure keeps up to the top 10 prefixes of each length and the top 200 overall, expands using the top 5 next vertex-token candidates, and outputs the best completed prefix. Experiments use beam size 200, γ=0.1\gamma=0.1, and tune α\alpha over [1,1.4][1,1.4] on validation data. Unlike greedy and lookahead, beam search uses sequential search operations, although it does not run deep-network computations autoregressively.

  7. Knowl 7 — DA-Transformer approaches autoregressive BLEU while retaining large speedups

    empirical result

    On WMT14 English–German (En–De), WMT14 German–English (De–En), WMT17 English–Chinese (En–Zh), and WMT17 Chinese–English (Zh–En), the authors trained on raw data as well as knowledge-distilled (KD) data. The DA-Transformer results below are means and standard deviations over three random seeds; each pair is raw-data BLEU / KD-data BLEU. The datasets contain 4.5 million WMT14 sentence pairs and 20 million WMT17 sentence pairs. BLEU is tokenized except for WMT17 En–Zh, which uses sacreBLEU.

    With lookahead decoding, scores were En–De 26.57±0.21/27.49±0.0526.57\pm0.21/27.49\pm0.05, De–En 30.68±0.24/31.37±0.0630.68\pm0.24/31.37\pm0.06, En–Zh 33.83±0.13/34.08±0.1333.83\pm0.13/34.08\pm0.13, and Zh–En 22.82±0.20/24.23±0.1422.82\pm0.20/24.23\pm0.14. The reported average gaps to the best autoregressive model were 1.22 BLEU on raw data and 0.57 on KD data, with a 13.9×\times speedup.

    With beam search plus a 5-gram language model, scores were En–De 27.25±0.12/27.91±0.0727.25\pm0.12/27.91\pm0.07, De–En 31.54±0.20/31.95±0.0631.54\pm0.20/31.95\pm0.06, En–Zh 34.23±0.17/34.27±0.0534.23\pm0.17/34.27\pm0.05, and Zh–En 24.49±0.06/25.01±0.1824.49\pm0.06/25.01\pm0.18. The average gaps were 0.32 BLEU on raw data and 0.08 on KD data, with a 7.0×\times speedup. The autoregressive Transformer used for comparison scored 23.89 raw and 24.68 KD BLEU on Zh–En, so DA-Transformer’s raw-data beam-search result was 0.60 BLEU higher. Speedups were measured on WMT17 En–De test data with batch size 1; the paper reports that lookahead outperformed the best existing non-iterative NAT baselines by 2.2 BLEU on average on raw data.

  8. Knowl 8 — Ablations support path marginalization, adaptive glancing, and a larger graph

    empirical result

    Ablations on raw WMT14 En–De evaluated the path objective and glancing-mask strategy. Marginalizing all paths (the sum objective) outperformed training only on the single most probable path (the max objective). The reported BLEU scores for all-masked, uniform-random, and adaptive glancing were, respectively, 22.79, 23.74, and 25.23 with the max objective, and 25.06, 25.74, and 26.57 with the sum objective. Adaptive glancing performed best among the three masking strategies, and the sum objective performed better than the max objective for each strategy.

    The graph-size multiplier λ\lambda sets graph length to L=λNL=\lambda N, with NN the source length. Increasing λ\lambda improved DA-Transformer’s BLEU through approximately λ=12\lambda=12, with performance relatively insensitive around its best setting; the authors selected λ=8\lambda=8 to balance quality and computation for the other experiments. CTC with GLAT had similar performance at λ=2\lambda=2 but did not gain from larger output lengths, unlike DA-Transformer, which can assign alternative tokens to separate vertices.

  9. Knowl 9 — Best-assignment accuracy and sampling tests indicate that the DAG retains alternatives

    empirical result

    The authors measured token accuracy after finding the best alignment between each reference and the model’s predictions. For WMT14 En–De and WMT17 Zh–En, respectively, training/validation accuracies were 29.7/29.6 and 39.8/22.4 for vanilla NAT; 50.2/51.7 and 47.1/32.3 for CTC+GLAT; and 69.3/69.9 and 80.1/67.0 for DA-Transformer. These results show that DA-Transformer’s best-path predictions match reference tokens more accurately than the two baselines in all four dataset splits.

    For a separate diversity evaluation, the authors sampled translations using nucleus sampling with p=0.8p=0.8 and temperatures from 0.4 to 1.0, and assessed translation quality and diversity using multi-reference BLEU and pairwise BLEU. DA-Transformer achieved a better quality–diversity tradeoff than GLAT+CTC; at the same temperature, its samples without KD were more diverse. On WMT17 Zh–En, its non-distilled model was slightly less diverse than the autoregressive Transformer but had a close quality–diversity tradeoff. Applying KD improved DA-Transformer’s quality while reducing diversity, consistent with KD reducing the variety of training targets.

  10. Knowl 10 — The learned graph still permits inconsistent token combinations, and training is slower per update

    limitation

    The graph does not guarantee that every path yields a globally consistent translation. The paper reports a possible erroneous combination such as “Does that sounds …” in a predicted DAG, identifying token consistency as an area for improvement even though such errors were not common with lookahead or beam decoding. The authors also report that each DA-Transformer training update is slower than updates for many NAT baselines because the decoder processes a graph sequence about eight times the original target length.

Coverage note — Appendix analyses of BLEU by reference length, controlled-GPU-time training curves, DAG vertex and out-degree distributions, and additional qualitative translation cases are omitted as supporting diagnostics; the core method, headline comparisons, principal ablations, and modality analyses are included.

References

  1. 1.Akhbardeh, F., Arkhangorodsky, A., Biesialska, M., Bojar, O., Chatterjee, R., Chaudhary, V., Costa-jussa, M. R., España-Bonet, C., Fan, A., Federmann, C., Freitag, M., Graham, Y., Grundkiewicz, R., Haddow, B., Harter, L., Heafield, K., Homan, C., Huck, M., Amponsah-Kaakyire, K., Kasai, J., Khashabi, D., Knight, K., Kocmi, T., Koehn, P., Lourie, N., Monz, C., Morishita, M., Nagata, M., Nagesh, A., Nakazawa, T., Negri, M., Pal, S., Tapo, A. A., Turchi, M., Vydrin, V., and Zampieri, M. Findings of the 2021 conference on machine translation (WMT21). In Proceedings of the Sixth Conference on Machine Translation, pp. 1–88, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.1.
  2. 2.Bao, Y., Zhou, H., Feng, J., Wang, M., Huang, S., Chen, J., and Li, L. Non-autoregressive transformer by position learning. CoRR, abs/1911.10677, 2019. URL http://arxiv.org/abs/1911.10677.
  3. 3.Bao, Y., Huang, S., Xiao, T., Wang, D., Dai, X., and Chen, J. Non-autoregressive translation by learning target categorical codes. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pp. 5749–5759. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.naacl-main.458. URL https://doi.org/10.18653/v1/2021.naacl-main.458.
  4. 4.Ding, L., Wang, L., Liu, X., Wong, D. F., Tao, D., and Tu, Z. Understanding and improving lexical choice in non-autoregressive translation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021a. URL https://openreview.net/forum?id=ZTFeSBIX9C.
  5. 5.Ding, L., Wang, L., Liu, X., Wong, D. F., Tao, D., and Tu, Z. Rejuvenating low-frequency words: Making the most of parallel data in non-autoregressive translation. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pp. 3431–3441. Association for Computational Linguistics, 2021b. doi: 10.18653/v1/2021.acl-long.266. URL https://doi.org/10.18653/v1/2021.acl-long.266.
  6. 6.Dong, M., Cheng, Y., Liu, Y., Xu, J., Sun, M., Izuha, T., and Hao, J. Query lattice for translation retrieval. In Hajic, J. and Tsujii, J. (eds.), COLING 2014, 25th International Conference on Computational Linguistics, Proceedings of the Conference: Technical Papers, August 23-29, 2014, Dublin, Ireland, pp. 2031–2041. ACL, 2014. URL https://aclanthology.org/C14-1192/.
  7. 7.Du, C., Tu, Z., and Jiang, J. Order-agnostic cross entropy for non-autoregressive machine translation. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 2849–2859. PMLR, 2021. URL http://proceedings.mlr.press/v139/du21c.html.
  8. 8.Dyer, C., Muresan, S., and Resnik, P. Generalizing word lattice translation. In McKeown, K. R., Moore, J. D., Teufel, S., Allan, J., and Furui, S. (eds.), ACL 2008, Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics, June 15-20, 2008, Columbus, Ohio, USA, pp. 1012–1020. The Association for Computer Linguistics, 2008. URL https://aclanthology.org/P08-1115/.
  9. 9.Feng, Y., Liu, Y., Mi, H., Liu, Q., and Lu, Y. Lattice-based system combination for statistical machine translation. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, EMNLP 2009, 6-7 August 2009, Singapore, A meeting of SIGDAT, a Special Interest Group of the ACL, pp. 1105–1113. ACL, 2009. URL https://aclanthology.org/D09-1115/.
  10. 10.Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L. Mask-predict: Parallel decoding of conditional masked language models. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 6111–6120. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1633. URL https://doi.org/10.18653/v1/D19-1633.
  11. 11.Ghazvininejad, M., Karpukhin, V., Zettlemoyer, L., and Levy, O. Aligned cross entropy for non-autoregressive machine translation. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 3515–3523. PMLR, 2020a. URL http://proceedings.mlr.press/v119/ghazvininejad20a.html.
  12. 12.Ghazvininejad, M., Levy, O., and Zettlemoyer, L. Semi-autoregressive training improves mask-predict decoding. CoRR, abs/2001.08785, 2020b. URL https://arxiv.org/abs/2001.08785.
  13. 13.Graves, A., Fernandez, S., Gomez, F. J., and Schmidhuber, J. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Cohen, W. W. and Moore, A. W. (eds.), Machine Learning, Proceedings of the Twenty-Third International Conference (ICML 2006), Pittsburgh, Pennsylvania, USA, June 25-29, 2006, volume 148 of ACM International Conference Proceeding Series, pp. 369–376. ACM, 2006. doi: 10.1145/1143844.1143891. URL https://doi.org/10.1145/1143844.1143891.
  14. 14.Gu, J. and Kong, X. Fully non-autoregressive neural machine translation: Tricks of the trade. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pp. 120–133. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.findings-acl.11. URL https://doi.org/10.18653/v1/2021.findings-acl.11.
  15. 15.Gu, J., Bradbury, J., Xiong, C., Li, V. O. K., and Socher, R. Non-autoregressive neural machine translation. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. URL https://openreview.net/forum?id=B1l8BtlCb.
  16. 16.Gu, J., Wang, C., and Zhao, J. Levenshtein transformer. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 11179–11189, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/675f9820626f5bc0afb47b57890b466e-Abstract.html.
  17. 17.Guo, J., Tan, X., He, D., Qin, T., Xu, L., and Liu, T. Non-autoregressive neural machine translation with enhanced decoder input. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pp. 3723–3730. AAAI Press, 2019. doi: 10.1609/aaai.v33i01.33013723. URL https://doi.org/10.1609/aaai.v33i01.33013723.
  18. 18.Guo, J., Xu, L., and Chen, E. Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 376–385. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.acl-main.36. URL https://doi.org/10.18653/v1/2020.acl-main.36.
  19. 19.Hannun, A. Y., Maas, A. L., Jurafsky, D., and Ng, A. Y. First-pass large vocabulary continuous speech recognition using bi-directional recurrent dnns. CoRR, abs/1408.2873, 2014. URL http://arxiv.org/abs/1408.2873.
  20. 20.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  21. 21.Huang, C., Zhou, H., Zaïane, O. R., Mou, L., and Li, L. Non-autoregressive translation with layer-wise prediction and deep supervision. The Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, 2022a. URL https://arxiv.org/abs/2110.07515.
  22. 22.Huang, F., Tao, T., Zhou, H., Li, L., and Huang, M. On the learning of non-autoregressive transformers. In Proceedings of the 39th International Conference on Machine Learning, ICML 2022, 2022b.
  23. 23.Huang, X. S., Perez, F., and Volkovs, M. Improving non-autoregressive translation models without distillation. In International Conference on Learning Representations, 2022c. URL https://openreview.net/forum?id=I2Hw58KHp8O.
  24. 24.Kaiser, L., Bengio, S., Roy, A., Vaswani, A., Parmar, N., Uszkoreit, J., and Shazeer, N. Fast decoding in sequence models using discrete latent variables. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 2395–2404. PMLR, 2018. URL http://proceedings.mlr.press/v80/kaiser18a.html.
  25. 25.Kasai, J., Cross, J., Ghazvininejad, M., and Gu, J. Non-autoregressive machine translation with disentangled context transformer. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 5144–5155. PMLR, 2020. URL http://proceedings.mlr.press/v119/kasai20a.html.
  26. 26.Kasai, J., Pappas, N., Peng, H., Cross, J., and Smith, N. A. Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=KpfasTaLUpq.
  27. 27.Kim, Y. and Rush, A. M. Sequence-level knowledge distillation. In Su, J., Carreras, X., and Duh, K. (eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pp. 1317–1327. The Association for Computational Linguistics, 2016. doi: 10.18653/v1/d16-1139. URL https://doi.org/10.18653/v1/d16-1139.
  28. 28.Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., Dyer, C., Bojar, O., Constantin, A., and Herbst, E. Moses: Open source toolkit for statistical machine translation. In Carroll, J. A., van den Bosch, A., and Zaenen, A. (eds.), ACL 2007, Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics, June 23-30, 2007, Prague, Czech Republic. The Association for Computational Linguistics, 2007. URL https://aclanthology.org/P07-2045/.
  29. 29.Lee, J., Mansimov, E., and Cho, K. Deterministic non-autoregressive neural sequence modeling by iterative refinement. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 1173–1182. Association for Computational Linguistics, 2018. doi: 10.18653/v1/d18-1149. URL https://doi.org/10.18653/v1/d18-1149.
  30. 30.Lee, J., Shu, R., and Cho, K. Iterative refinement in the continuous space for non-autoregressive neural machine translation. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 1006–1015. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.emnlp-main.73. URL https://doi.org/10.18653/v1/2020.emnlp-main.73.
  31. 31.Libovicky, J. and Helcl, J. End-to-end non-autoregressive neural machine translation with connectionist temporal classification. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 3016–3021. Association for Computational Linguistics, 2018. doi: 10.18653/v1/d18-1336. URL https://doi.org/10.18653/v1/d18-1336.
  32. 32.Ma, X., Zhou, C., Li, X., Neubig, G., and Hovy, E. H. Flowseq: Non-autoregressive conditional sequence generation with generative flow. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 4281–4291. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1437. URL https://doi.org/10.18653/v1/D19-1437.
  33. 33.Och, F. J. and Ney, H. The alignment template approach to statistical machine translation. Comput. Linguistics, 30(4):417–449, 2004. doi: 10.1162/0891201042544884. URL https://doi.org/10.1162/0891201042544884.
  34. 34.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In Ammar, W., Louis, A., and Mostafazadeh, N. (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Demonstrations, pp. 48–53. Association for Computational Linguistics, 2019. doi: 10.18653/v1/n19-4009. URL https://doi.org/10.18653/v1/n19-4009.
  35. 35.Papineni, K., Roukos, S., Ward, T., and Zhu, W. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA, pp. 311–318. ACL, 2002. doi: 10.3115/1073083.1073135. URL https://aclanthology.org/P02-1040/.
  36. 36.Post, M. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pp. 186–191, Belgium, Brussels, October 2018. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/W18-6319.
  37. 37.Qian, L., Zhou, H., Bao, Y., Wang, M., Qiu, L., Zhang, W., Yu, Y., and Li, L. Glancing transformer for non-autoregressive neural machine translation. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pp. 1993–2003. Association for Computational Linguistics, 2021a. URL https://aclanthology.org/2021.acl-long.155.
  38. 38.Qian, L., Zhou, Y., Zheng, Z., Zhu, Y., Lin, Z., Feng, J., Cheng, S., Li, L., Wang, M., and Zhou, H. The volctrans GLAT system: Non-autoregressive translation meets WMT21. In Proceedings of the Sixth Conference on Machine Translation, pp. 187–196, Online, November 2021b. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.17.
  39. 39.Richardson, F., Ostendorf, M., and Rohlicek, J. R. Lattice-based search strategies for large vocabulary speech recognition. In 1995 International Conference on Acoustics, Speech, and Signal Processing, ICASSP ’95, Detroit, Michigan, USA, May 08-12, 1995, pp. 576–579. IEEE Computer Society, 1995. doi: 10.1109/ICASSP.1995.479663. URL https://doi.org/10.1109/ICASSP.1995.479663.
  40. 40.Rosti, A. I., Ayan, N. F., Xiang, B., Matsoukas, S., Schwartz, R. M., and Dorr, B. J. Combining outputs from multiple machine translation systems. In Sidner, C. L., Schultz, T., Stone, M., and Zhai, C. (eds.), Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, Proceedings, April 22-27, 2007, Rochester, New York, USA, pp. 228–235. The Association for Computational Linguistics, 2007. URL https://aclanthology.org/N07-1029/.
  41. 41.Saharia, C., Chan, W., Saxena, S., and Norouzi, M. Non-autoregressive machine translation with latent alignments. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 1098–1108. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.emnlp-main.83. URL https://doi.org/10.18653/v1/2020.emnlp-main.83.
  42. 42.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics, 2016. doi: 10.18653/v1/p16-1162. URL https://doi.org/10.18653/v1/p16-1162.
  43. 43.Shao, C., Zhang, J., Feng, Y., Meng, F., and Zhou, J. Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pp. 198–205. AAAI Press, 2020. URL https://aaai.org/ojs/index.php/AAAI/article/view/5351.
  44. 44.Shen, T., Ott, M., Auli, M., and Ranzato, M. Mixture models for diverse machine translation: Tricks of the trade. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pp. 5719–5728. PMLR, 2019. URL http://proceedings.mlr.press/v97/shen19c.html.
  45. 45.Shu, R., Lee, J., Nakayama, H., and Cho, K. Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pp. 8846–8853. AAAI Press, 2020. URL https://aaai.org/ojs/index.php/AAAI/article/view/6413.
  46. 46.Sun, Z. and Yang, Y. An EM approach to non-autoregressive conditional sequence generation. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 9249–9258. PMLR, 2020. URL http://proceedings.mlr.press/v119/sun20c.html.
  47. 47.Sun, Z., Li, Z., Wang, H., He, D., Lin, Z., and Deng, Z. Fast structured decoding for sequence models. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 3011–3020, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/74563ba21a90da13dacf2a73e3ddefa7-Abstract.html.
  48. 48.Tromble, R., Kumar, S., Och, F. J., and Macherey, W. Lattice minimum bayes-risk decoding for statistical machine translation. In 2008 Conference on Empirical Methods in Natural Language Processing, EMNLP 2008, Proceedings of the Conference, 25-27 October 2008, Honolulu, Hawaii, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, pp. 620–629. ACL, 2008. URL https://aclanthology.org/D08-1065/.
  49. 49.Ueffing, N., Och, F. J., and Ney, H. Generation of word graphs in statistical machine translation. In Proceedings of the 2002 Conference on Empirical Methods in Natural Language Processing, EMNLP 2002, Philadelphia, PA, USA, July 6-7, 2002, pp. 156–163, 2002. doi: 10.3115/1118693.1118714. URL https://aclanthology.org/W02-1021/.
  50. 50.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
  51. 51.Wei, B., Wang, M., Zhou, H., Lin, J., and Sun, X. Imitation learning for non-autoregressive neural machine translation. In Korhonen, A., Traum, D. R., and Marquez, L. (eds.), Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pp. 1304–1312. Association for Computational Linguistics, 2019. doi: 10.18653/v1/p19-1125. URL https://doi.org/10.18653/v1/p19-1125.
  52. 52.Yang, K., Lei, W., Liu, D., Qi, W., and Lv, J. Pos-constrained parallel decoding for non-autoregressive generation. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pp. 5990–6000. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.acl-long.467. URL https://doi.org/10.18653/v1/2021.acl-long.467.
  53. 53.Zhou, C., Gu, J., and Neubig, G. Understanding knowledge distillation in non-autoregressive machine translation. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=BygFVAEKDH.

Citation

MLA
Huang, F., et al. “Directed Acyclic Transformer for Non-Autoregressive Machine Translation”. International Conference on Machine Learning, vol. 162, 2022, pp. 9410–28, https://proceedings.mlr.press/v162/huang22m.html.
APA
Huang, F., Zhou, H., Liu, Y., Li, H., & Huang, M. (2022). Directed Acyclic Transformer for Non-Autoregressive Machine Translation. International Conference on Machine Learning, 162, 9410–9428. https://proceedings.mlr.press/v162/huang22m.html
Chicago
Huang, F., H. Zhou, Y. Liu, H. Li, and M. Huang. 2022. “Directed Acyclic Transformer for Non-Autoregressive Machine Translation”. International Conference on Machine Learning 162: 9410–28. https://proceedings.mlr.press/v162/huang22m.html.
Harvard
Huang, F. et al. (2022) “Directed Acyclic Transformer for Non-Autoregressive Machine Translation”, International Conference on Machine Learning. PMLR, pp. 9410–9428. Available at: https://proceedings.mlr.press/v162/huang22m.html.
Vancouver
1. Huang F, Zhou H, Liu Y, Li H, Huang M (2022) Directed Acyclic Transformer for Non-Autoregressive Machine Translation. In: International Conference on Machine Learning. PMLR, pp 9410–9428

BibTeX

@InProceedings{pmlr-v162-huang22m,
  title = 	 {Directed Acyclic Transformer for Non-Autoregressive Machine Translation},
  author =       {Huang, Fei and Zhou, Hao and Liu, Yang and Li, Hang and Huang, Minlie},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {9410--9428},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/huang22m/huang22m.pdf},
  url = 	 {https://proceedings.mlr.press/v162/huang22m.html},
  abstract = 	 {Non-autoregressive Transformers (NATs) significantly reduce the decoding latency by generating all tokens in parallel. However, such independent predictions prevent NATs from capturing the dependencies between the tokens for generating multiple possible translations. In this paper, we propose Directed Acyclic Transfomer (DA-Transformer), which represents the hidden states in a Directed Acyclic Graph (DAG), where each path of the DAG corresponds to a specific translation. The whole DAG simultaneously captures multiple translations and facilitates fast predictions in a non-autoregressive fashion. Experiments on the raw training data of WMT benchmark show that DA-Transformer substantially outperforms previous NATs by about 3 BLEU on average, which is the first NAT model that achieves competitive results with autoregressive Transformers without relying on knowledge distillation.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/