Exploiting Similarities among Languages for Machine Translation

Tomas MikolovQuoc V. LeIlya Sutskever

article2013arXiv1,659 citations

Demonstrates that learning a simple linear mapping between monolingual vector spaces can automatically translate missing words and phrases across language pairs with high accuracy using minimal bilingual data.

Listen

Building and maintaining statistical machine translation systems requires extensive bilingual dictionaries and phrase tables, which are labor-intensive, expensive, and difficult to scale across diverse language pairs. The article evaluates an automated method to generate and expand bilingual dictionaries by learning a linear mapping between geometric vector spaces of different languages using large monolingual datasets and a small starting bilingual dictionary.

The approach constructs continuous vector representations of words and short phrases independently for each language using billions of words of monolingual text. Because languages share underlying semantic structures grounded in real-world concepts, their vector spaces exhibit similar geometric arrangements. The authors learn a linear transformation matrix using a small initial bilingual seed set of the 5,000 most frequent words and test the projection on unseen words. The method was evaluated across multiple language pairs—including English paired with Spanish, Czech, and Vietnamese—using both standard benchmark data and large-scale news corpora.

The findings show that this linear mapping significantly outperforms traditional count-based co-occurrence baselines. For English-to-Spanish translation using large corpora, setting confidence thresholds yielded top-5 translation accuracy of roughly 90% across more than half of the tested vocabulary. When combined with word-spelling similarities for related languages, performance improved further, achieving 75% top-1 accuracy at a 55% coverage rate. The approach scaled effectively with corpus size and proved capable of translating infrequent words—maintaining roughly 60% top-5 accuracy for words ranked 15,000 to 19,000 in frequency—while successfully bridging structurally distant language pairs like English and Vietnamese without relying on shared alphabets.

These results demonstrate that organizations can substantially reduce the cost, time, and manual effort required to build translation resources, particularly for low-resource language pairs where parallel bilingual text is scarce. Furthermore, by calculating translation confidence scores based on vector distances, systems can filter errors, flag ambiguous translations, and automatically clean existing bilingual dictionaries, where the model identified flawed entries roughly 15% of the time in sample evaluations.

Stakeholders should adopt this linear projection technique to augment phrase tables and systematically audit existing translation databases. Next steps should focus on piloting the method in production machine translation pipelines and extending it to low-resource domains. Decision-makers should note that model accuracy depends heavily on access to large monolingual corpora—smaller datasets dropped accuracy on infrequent words from 60% to 25%—and exact-match scoring understates true performance by treating valid synonyms as translation errors.

arXiv: 1309.4168tmikolov/word2vec
  • Paper: Word Translation Without Parallel Data, Alexis Conneau et al. (2017). Generalizes linear cross-lingual embedding mapping to a completely unsupervised setting using adversarial training and Procrustes refinement without requiring parallel bilingual data.
  • Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Advances cross-lingual representation learning from linear word-level alignments to contextualized multilingual language model pretraining.
  • Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Builds a comprehensive cross-lingual evaluation benchmark (XNLI) to assess how well multilingual representations transfer semantic knowledge across languages.
  • Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Investigates the cross-lingual geometric properties and zero-shot transfer capabilities emerging within shared multilingual embedding representations.
  • Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). Enhances monolingual vector spaces with character n-grams to handle rare and morphologically rich words, directly improving the embeddings used in cross-lingual projection.
  • Paper: Neural Word Embedding as Implicit Matrix Factorization, Omer Levy et al. (2014). Provides a theoretical analysis showing that the neural skip-gram word embeddings mapped in this paper correspond mathematically to shifted PMI matrix factorization.
Cover for Exploiting Similarities among Languages for Machine Translation

Abstract

Dictionaries and phrase tables are the basis of modern statistical machine translation systems. This paper develops a method that can automate the process of generating and extending dictionaries and phrase tables. Our method can translate missing word and phrase entries by learning language structures based on large monolingual data and mapping between languages from small bilingual data. It uses distributed representation of words and learns a linear mapping between vector spaces of languages. Despite its simplicity, our method is surprisingly effective: we can achieve almost 90% precision@5 for translation of words between English and Spanish. This method makes little assumption about the languages, so it can be used to extend and refine dictionaries and translation tables for any language pairs.

Table of Contents

  • 1 Introduction
  • 2 The Skip-gram and Continuous Bag-of-Words Models
  • 3 Linear Relationships Between Languages
  • 4 Translation Matrix
  • 5 Experiments on WMT11 Datasets
  • 5.1 Setup Description
  • 5.2 Baseline Techniques
  • 5.3 Results with WMT11 Data
  • 6 Large Scale Experiments
  • 6.1 Using Distances as Confidence Measure
  • 7 Examples
  • 7.1 Spanish to English Example Translations
  • 7.2 High Confidence Translations
  • 7.3 Detection of Dictionary Errors
  • 7.4 Translation between distant language pairs: English and Vietnamese
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Linear Mapping for Cross-Lingual Word Translation

    model/method

    Given continuous word vector representations trained independently on monolingual corpora for a source language and a target language, and a seed dictionary of nn translated word pairs {xi,zi}i=1n\{x_i, z_i\}_{i=1}^n, where xi∈Rd1x_i \in \mathbb{R}^{d_1} is the distributed representation of source word ii and zi∈Rd2z_i \in \mathbb{R}^{d_2} is the distributed representation of its target translation, the translation system learns a linear transformation matrix W∈Rd2×d1W \in \mathbb{R}^{d_2 \times d_1}.

    The parameter matrix WW is optimized by minimizing the sum of squared Euclidean errors: min⁡W∑i=1n∥Wxi−zi∥2\min_W \sum_{i=1}^n \|W x_i - z_i\|^2 which is solved using stochastic gradient descent.

    At test time, given the vector representation x∈Rd1x \in \mathbb{R}^{d_1} of any source word present in the monolingual corpus, the model maps it into the target space via z^=Wx\hat{z} = W x. The translation is predicted as the target vocabulary word yy whose vector representation zy∈Rd2z_y \in \mathbb{R}^{d_2} has the highest cosine similarity to z^\hat{z}: y^=arg⁡max⁡y∈Vtarget(Wx)⊤zy∥Wx∥2∥zy∥2\hat{y} = \arg\max_{y \in V_{\text{target}}} \frac{(W x)^\top z_y}{\|W x\|_2 \|z_y\|_2} where VtargetV_{\text{target}} denotes the vocabulary of the target language.

  2. Knowl 2 — Geometric Isomorphism of Cross-Lingual Word Vector Spaces

    assumption

    Continuous distributed word representations (such as Skip-gram and Continuous Bag-of-Words embeddings) trained independently on large monolingual corpora across different natural languages exhibit isomorphic geometric arrangements for shared real-world concepts (such as semantic categories, numbers, and animals). Because cross-linguistic semantic relationships share topological structure, the global relationship between two independently learned continuous language spaces can be accurately captured by an affine linear mapping consisting of rotation and scaling.

  3. Knowl 3 — Confidence Scoring and Bilingual Dictionary Error Detection

    model/method

    Given a learned cross-lingual projection matrix W∈Rd2×d1W \in \mathbb{R}^{d_2 \times d_1} mapping source word vectors x∈Rd1x \in \mathbb{R}^{d_1} to target word vectors z∈Rd2z \in \mathbb{R}^{d_2}, a translation confidence score for a source word xx is defined as the maximum cosine similarity between its mapped representation WxW x and any word vector in the target vocabulary VtargetV_{\text{target}}: confidence(x)=max⁡i∈Vtarget(Wx)⊤zi∥Wx∥2∥zi∥2\text{confidence}(x) = \max_{i \in V_{\text{target}}} \frac{(W x)^\top z_i}{\|W x\|_2 \|z_i\|_2} If this confidence value falls below a chosen threshold, the translation candidate is rejected, allowing translation precision to be tuned against vocabulary coverage.

    Conversely, for existing bilingual dictionary entries (xi,zi)(x_i, z_i), measuring the distance or angular separation between the projected source embedding WxiW x_i and the dictionary translation embedding ziz_i provides an automated filter to identify erroneous, ambiguous, or rare dictionary translations.

  4. Knowl 4 — Cross-Lingual Translation Benchmarks on WMT11

    data/table

    Word and phrase translation accuracy was evaluated on the WMT11 datasets across English (En), Spanish (Sp), and Czech (Cz). Monolingual Continuous Bag-of-Words (CBOW) models were trained on 575M tokens (127K vocabulary) for English, 84M tokens (107K vocabulary) for Spanish, and 155M tokens (505K vocabulary) for Czech. The linear Translation Matrix (TM) was trained on the 5,000 most frequent source words and evaluated on the subsequent 1,000 words against morphological Edit Distance (ED), count-based Word Co-occurrence, and a combined model (ED + TM). Evaluation is measured by exact-match Precision@1 (P@1) and Precision@5 (P@5).

    Translation Edit Distance Word Co-occurrence Translation Matrix ED + TM Coverage
    P@1 P@5 P@1 P@5 P@1 P@5 P@1 P@5
    En →\rightarrow Sp 13% 24% 19% 30% 33% 51% 43% 60% 92.9%
    Sp →\rightarrow En 18% 27% 20% 30% 35% 52% 44% 62% 92.9%
    En →\rightarrow Cz 5% 9% 9% 17% 27% 47% 29% 50% 90.5%
    Cz →\rightarrow En 7% 11% 11% 20% 23% 42% 25% 45% 90.5%

    The linear Translation Matrix substantially outperforms both morphological and co-occurrence baselines across all pairs. Combining the Translation Matrix with Edit Distance produces large gains for morphologically related language pairs (English and Spanish), but gives negligible benefit for distant pairs (English and Czech).

  5. Knowl 5 — Source-Target Dimensionality Asymmetry in Cross-Lingual Projection

    empirical result

    When optimizing a linear translation matrix W∈Rd2×d1W \in \mathbb{R}^{d_2 \times d_1} between continuous word representations, translation accuracy is highest when the source language representation dimensionality d1d_1 is substantially larger (approximately 2×2\times to 4×4\times) than the target language dimensionality d2d_2.

    In English-to-Spanish translation experiments, optimal performance was obtained with 800-dimensional English source vectors and 200-dimensional Spanish target vectors. In Spanish-to-English translation, the highest accuracy was achieved with 800-dimensional Spanish source vectors and 300-dimensional English target vectors.

  6. Knowl 6 — Scaling with Monolingual Data Size and Vocabulary Frequency

    empirical result

    Cross-lingual linear translation accuracy scales monotonically with the size of monolingual training corpora. Scaling monolingual English and Spanish training sets from 10710^7 words up to billions of words (Google News data) increases Precision@1 from under 10% to approximately 53%, and Precision@5 from under 20% to approximately 75% on the top 5K–6K frequency band.

    When evaluated on disjoint 2,000-word test bins of decreasing frequency (from ranks 5K–7K down to 17K–19K) using a projection matrix trained only on the top 5K words:

    • Embeddings trained on large corpora retain Precision@5 of approximately 60% for rare words at ranks 15K–19K.
    • Embeddings trained on the smaller WMT11 corpus drop to 25% Precision@5 on the same 15K–19K band, demonstrating that large monolingual corpora are required to correctly position infrequent word vectors for cross-lingual alignment.
  7. Knowl 7 — Precision-Coverage Tradeoff via Confidence Thresholding

    data/table

    Applying a minimum confidence threshold to projected word translations allows trading off vocabulary coverage for higher translation precision. Evaluated on English-to-Spanish translation for words ranked 5K–6K using models trained on large Google News corpora:

    Translation Matrix Only Translation Matrix + Edit Distance
    Threshold Coverage P@1 P@5 Threshold Coverage P@1 P@5
    0.0 92.5% 53% 75% 0.0 92.5% 58% 77%
    0.5 78.4% 59% 82% 0.4 77.6% 66% 84%
    0.6 54.0% 71% 90% 0.5 55.0% 75% 91%
    0.7 17.0% 78% 91% 0.6 25.3% 85% 93%

    At a threshold of 0.6, the standalone Translation Matrix translates 54.0% of test words with 71% P@1 and 90% P@5. When combined with morphological edit distance, a threshold of 0.6 achieves 85% P@1 and 93% P@5 on 25.3% of the test vocabulary.

  8. Knowl 8 — Cross-Lingual Projection on Distant Language Pairs with Phrase Extraction

    data/table

    To translate between language pairs lacking one-to-one word boundaries, frequent multi-word collocations are merged into single tokens using statistical bigram scoring prior to Skip-gram training. On English and Vietnamese using Google News corpora (yielding 1.3B training tokens/phrases for Vietnamese), a linear translation matrix trained on 5K seed pairs achieves the following performance:

    Translation Direction Coverage Precision@1 Precision@5
    English →\rightarrow Vietnamese 87.8% 10% 30%
    Vietnamese →\rightarrow English 87.8% 24% 40%

    Morphological edit distance provides no benefit for this distant language pair. Exact-match Precision@1 for English to Vietnamese is degraded by the existence of multiple valid synonyms in the target space that are counted as incorrect under strict string matching.

  9. Knowl 9 — Normalized Count-Based Word Co-Occurrence Baseline

    model/method

    The count-based co-occurrence baseline builds continuous representations using a bilingual seed dictionary DD of size ∣D∣|D|:

    1. For each test word in the source language and candidate word in the target language, a vector of dimension ∣D∣|D| records the count of co-occurring in-dictionary words within a sliding context window of up to 10 words.
    2. To remove bias caused by differing monolingual corpus sizes, source language co-occurrence counts are divided by the corpus size ratio: csource,j′=csource,j∣Ssource∣/∣Starget∣c_{\text{source}, j}' = \frac{c_{\text{source}, j}}{|S_{\text{source}}| / |S_{\text{target}}|} where ∣Ssource∣|S_{\text{source}}| and ∣Starget∣|S_{\text{target}}| are the token counts of the source and target corpora.
    3. Counts are transformed logarithmically: vj=log⁡(1+cj′)v_j = \log(1 + c_j').
    4. Each vector is normalized to unit Euclidean length (∥v∥2=1\|v\|_2 = 1).
    5. Source vectors are directly compared against target vectors indexed by corresponding dictionary entries via cosine similarity.

Coverage note — Specific qualitative translation samples from Tables 5, 6, and 7 were omitted in favor of the quantitative benchmarks and formal error-detection methodology. The Skip-gram and CBOW loss formulations were omitted as they are prior work cited from Mikolov et al. (2013a).

References

  1. 1.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin. 2003. A neural probabilistic language model. In Journal of Machine Learning Research, pages 1137–1155.
  2. 2.Ronan Collobert and Jason Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of the 25th international conference on Machine learning, pages 160–167. ACM.
  3. 3.Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. The Journal of Machine Learning Research, 12:2493–2537.
  4. 4.Jeff Elman. 1990. Finding structure in time. In Cognitive Science, pages 179–211.
  5. 5.Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, and Dan Klein. 2008. Learning bilingual lexicons from monolingual corpora. In ACL, volume 2008, pages 771–779.
  6. 6.Eric Huang, Richard Socher, Christopher Manning, and Andrew Y Ng. 2012. Improving word representations via global context and multiple word prototypes. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1, pages 873–882. Association for Computational Linguistics.
  7. 7.Philipp Koehn and Kevin Knight. 2000. Estimating word translation probabilities from unrelated monolingual corpora using the em algorithm. In AAAI/IAAI, pages 711–715.
  8. 8.Philipp Koehn and Kevin Knight. 2002. Learning a translation lexicon from monolingual corpora. In Proceedings of the ACL-02 workshop on Unsupervised lexical acquisition-Volume 9, pages 9–16. Association for Computational Linguistics.
  9. 9.Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernocký, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In INTERSPEECH, pages 1045–1048.
  10. 10.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  11. 11.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013b. Distributed representations of phrases and their compositionality. In NIPS.
  12. 12.Tomas Mikolov, Scott Wen-tau Yih, and Geoffrey Zweig. 2013c. Linguistic regularities in continuous space word representations. In NAACL HLT.
  13. 13.Tomas Mikolov. 2012. Statistical Language Models based on Neural Networks. Ph.D. thesis, Brno University of Technology.
  14. 14.Andriy Mnih and Geoffrey E Hinton. 2008. A scalable hierarchical distributed language model. In Advances in neural information processing systems, pages 1081–1088.
  15. 15.Frederic Morin and Yoshua Bengio. 2005. Hierarchical probabilistic neural network language model. In Proceedings of the international workshop on artificial intelligence and statistics, pages 246–252.
  16. 16.David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986. Learning representations by back-propagating errors. Nature, 323(6088):533–536.
  17. 17.Richard Socher, Cliff C Lin, Andrew Ng, and Chris Manning. 2011. Parsing natural scenes and natural language with recursive neural networks. In Proceedings of the 28th International Conference on Machine Learning (ICML-11), pages 129–136.
  18. 18.Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng. 2013. Parsing with compositional vector grammars. In ACL.
  19. 19.Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010. Word representations: a simple and general method for semi-supervised learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 384–394. Association for Computational Linguistics.

Citation

MLA
Mikolov, T., et al. “Exploiting Similarities Among Languages for Machine Translation”. arXiv, 2013, https://doi.org/10.48550/arxiv.1309.4168.
APA
Mikolov, T., Le, Q. V., & Sutskever, I. (2013). Exploiting Similarities among Languages for Machine Translation. arXiv. https://doi.org/10.48550/arxiv.1309.4168
Chicago
Mikolov, T., Q. V. Le, and I. Sutskever. 2013. “Exploiting Similarities Among Languages for Machine Translation”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1309.4168.
Harvard
Mikolov, T., Le, Q.V. and Sutskever, I. (2013) “Exploiting Similarities among Languages for Machine Translation”. arXiv. Available at: https://doi.org/10.48550/arxiv.1309.4168.
Vancouver
1. Mikolov T, Le QV, Sutskever I (2013) Exploiting Similarities among Languages for Machine Translation. https://doi.org/10.48550/arxiv.1309.4168

BibTeX

@misc{https://doi.org/10.48550/arxiv.1309.4168,
  doi = {10.48550/ARXIV.1309.4168},
  url = {https://arxiv.org/abs/1309.4168},
  author = {Mikolov, Tomas and Le, Quoc V. and Sutskever, Ilya},
  keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Exploiting Similarities among Languages for Machine Translation},
  publisher = {arXiv},
  year = {2013},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission