Directed Acyclic Transformer for Non-Autoregressive Machine Translation
Fei HuangHao ZhouYang LiuHang LiMinlie Huang
Proposes the Directed Acyclic Transformer, a non-autoregressive machine translation model that organizes decoder states into a directed acyclic graph to capture multiple valid translations simultaneously, achieving competitive translation quality with standard autoregressive models without needing knowledge distillation.
Standard machine translation systems generate words sequentially from left to right. While this produces high-quality output, it creates high latency during real-time deployment. Non-autoregressive models address this speed bottleneck by generating all words simultaneously in parallel, but they suffer from severe translation quality degradation. Because a single source sentence can be translated correctly in multiple distinct ways, non-autoregressive models struggle with conflicting training signals and frequently output garbled sentences that mix disparate phrasing.
The article develops and evaluates the Directed Acyclic Transformer, a non-autoregressive model designed to capture multiple valid translations simultaneously without requiring complex data pre-processing or multi-reference training sets.
The authors evaluate the system using standard translation benchmarks across English, German, and Chinese language pairs, comprising millions of sentence pairs. The architecture organizes target representations into a flexible directed graph structure, allowing parallel word generation along multiple distinct translation paths. The model is trained using an efficient dynamic programming approach and glancing training techniques, evaluated across multiple decoding algorithms including parallel lookahead and beam search.
The analysis demonstrates that the proposed architecture achieves translation quality competitive with traditional sequential models while providing 7x to 14x faster generation speed. When trained directly on raw data without distilled reference sets, it outperforms previous non-autoregressive baselines by approximately 3.0 BLEU points on average, and surpasses standard sequential transformers by 0.6 BLEU on Chinese-to-English translation. Furthermore, the model shows marked improvements on longer sentences exceeding 20 words and exhibits superior lexical diversity when sampling alternative translations.
These findings indicate that non-autoregressive systems can eliminate reliance on knowledge distillation—a resource-intensive pre-training step where a sequential model first generates synthetic training targets. Removing this dependency reduces engineering complexity, lowers pre-processing costs, and removes the artificial quality ceiling imposed by distillation teacher models. The system offers organizations a flexible trade-off between execution speed and translation quality across edge or cloud deployments.
Organizations deploying high-throughput or latency-sensitive translation systems should pilot the directed graph model as a drop-in replacement for sequential pipelines. Teams can adopt fast lookahead decoding for low-latency production serving or select beam search when prioritizing translation accuracy. Further work should focus on refining the consistency between graph vertices to prevent rare syntactic errors and validating the architecture across low-resource language pairs where large-scale training data is limited.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Read this first to understand the Transformer encoder–decoder architecture that the Directed Acyclic Transformer adapts for parallel translation.
No sufficiently relevant recommendations were found.
