Word Translation Without Parallel Data

Alexis ConneauGuillaume LampleMarc'Aurelio RanzatoLudovic DenoyerHervé Jégou

article2017ICLR1,815 citations

Presents an unsupervised method for aligning monolingual word embeddings that builds accurate bilingual dictionaries without parallel data or shared alphabets, matching or outperforming supervised approaches across distant language pairs.

Listen

Natural language translation and cross-lingual artificial intelligence systems traditionally depend on extensive bilingual dictionaries or parallel text corpora, in which sentences are manually translated across languages. Generating and maintaining these bilingual resources is expensive, labor-intensive, and often infeasible for low-resource or non-European languages. The article evaluates whether high-quality cross-lingual word mappings and bilingual dictionaries can be built entirely without parallel or paired data, relying strictly on independent monolingual text.

The researchers developed an unsupervised framework that aligns independently trained monolingual word embedding spaces into a shared coordinate system. The approach uses a two-step alignment process: an adversarial training phase where a neural network discriminator learns to distinguish between source and target language word vectors while a transformation matrix is trained to fool it, followed by a refinement step using an exact mathematical Procrustes alignment based on mutually confident anchor words. To evaluate matchings accurately, the team introduced a distance metric called Cross-Domain Similarity Local Scaling (CSLS) to reduce the hubness problem—a distortion where certain words appear as default nearest neighbors to many unrelated queries. They also devised an internal unsupervised validation metric to tune hyper-parameters and determine stopping criteria without requiring translated validation sets. Credibility was established by evaluating large vocabularies across diverse language pairs (including English paired with French, Spanish, German, Russian, Chinese, and Esperanto) against established supervised baselines.

The evaluation yielded several key findings. First, the unsupervised framework matched or outperformed supervised state-of-the-art baselines across multiple cross-lingual tasks; on standard English-Italian translation retrieval, the unsupervised approach reached 66.2% precision compared to 63.7% for the leading supervised baseline. Second, introducing CSLS provided large performance gains across both supervised and unsupervised systems, increasing retrieval precision by up to 7 to 8 percentage points over standard nearest-neighbor matching. Third, the system demonstrated strong cross-lingual generalization across distant languages that do not share character sets or alphabets, such as English-Russian and English-Chinese, where previous character-based heuristics failed. Fourth, on low-resource pairs lacking large parallel corpora, such as English-Esperanto, the method produced functional word-by-word sentence translation, achieving an automated translation score of 14.3 BLEU in Esperanto-to-English translation.

These findings indicate that cross-lingual systems do not strictly require costly parallel data or bilingual dictionaries to achieve high performance. Organizations can substantially lower data acquisition costs and reduce implementation timelines when deploying language models for lower-resource language pairs or niche vocabularies. The results challenge the assumption that supervision or shared alphabets are mandatory for cross-lingual vector space alignment.

Teams and practitioners should consider adopting unsupervised cross-lingual vector alignment and CSLS matching metrics when developing multilingual natural language applications, particularly where parallel data is sparse or unavailable. For end-to-end sentence translation workflows, next steps should include pairing this unsupervised lexicon induction with language models to correct syntax and word-order errors that arise in basic word-by-word translations.

A key limitation is that alignment quality relies on the structural comparability of the underlying monolingual corpora; divergence in domain topics or co-occurrence statistics across texts reduces alignment accuracy on rarer words. In addition, simpler word-by-word translations fail to capture complex polysemy and grammatical structure without downstream decoding models. Within these boundary conditions, confidence in the cross-lingual word mapping capability is high across multiple distinct language families.

Cover for Word Translation Without Parallel Data

Abstract

State-of-the-art methods for learning cross-lingual word embeddings have relied on bilingual dictionaries or parallel corpora. Recent studies showed that the need for parallel data supervision can be alleviated with character-level information. While these methods showed encouraging results, they are not on par with their supervised counterparts and are limited to pairs of languages sharing a common alphabet. In this work, we show that we can build a bilingual dictionary between two languages without using any parallel corpora, by aligning monolingual word embedding spaces in an unsupervised way. Without using any character information, our model even outperforms existing supervised methods on cross-lingual tasks for some language pairs. Our experiments demonstrate that our method works very well also for distant language pairs, like English-Russian or English-Chinese. We finally describe experiments on the English-Esperanto low-resource language pair, on which there only exists a limited amount of parallel data, to show the potential impact of our method in fully unsupervised machine translation. Our code, embeddings and dictionaries are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Model
  • 2.1 Domain-adversarial setting
  • 2.2 Refinement procedure
  • 2.3 Cross-domain similarity local scaling (CSLS)
  • 3 Training and architectural choices
  • 3.1 Architecture
  • 3.2 Discriminator inputs
  • 3.3 Orthogonality
  • 3.4 Dictionary generation
  • 3.5 Validation criterion for unsupervised model selection
  • 4 Experiments
  • 4.1 Evaluation tasks
  • 4.2 Results and discussion
  • 5 Related work
  • 6 Conclusion
  • References
  • 7 Appendix

Knowls

  1. Knowl 1 — Unsupervised Cross-Lingual Word Embedding Alignment Framework

    model/method

    The framework aligns two independently trained monolingual continuous word embedding spaces—a source space X={x1,…,xn}⊂Rd\mathcal{X} = \{x_1, \dots, x_n\} \subset \mathbb{R}^d and a target space Y={y1,…,ym}⊂Rd\mathcal{Y} = \{y_1, \dots, y_m\} \subset \mathbb{R}^d—without using any parallel text, bilingual dictionaries, or character-level heuristics. The alignment pipeline operates in two sequential stages:

    1. Adversarial Initialization: A discriminator network is trained to differentiate between mapped source word embeddings WxsW x_s and target word embeddings yty_t, while a linear transformation matrix W∈Rd×dW \in \mathbb{R}^{d \times d} is simultaneously optimized as a generator to project source embeddings into the target space to fool the discriminator.
    2. Procrustes Refinement: Using the initial mapping WW, a high-precision synthetic bilingual dictionary is constructed by identifying mutual nearest neighbors among frequent words using Cross-Domain Similarity Local Scaling (CSLS). The mapping WW is then refined by computing the closed-form orthogonal Procrustes solution on this synthetic dictionary. This step can optionally be applied iteratively.
  2. Knowl 2 — Adversarial Objective and Training Procedure for Linear Embedding Mapping

    model/method

    Given source word embeddings X={x1,…,xn}\mathcal{X} = \{x_1, \dots, x_n\} and target word embeddings Y={y1,…,ym}\mathcal{Y} = \{y_1, \dots, y_m\}, a linear mapping W∈Rd×dW \in \mathbb{R}^{d \times d} and a discriminator parameterized by θD\theta_D engage in a two-player minimax game. The discriminator outputs the probability PθD(source=1∣z)P_{\theta_D}(\text{source} = 1 \mid z) that an embedding vector z∈Rdz \in \mathbb{R}^d originates from the mapped source distribution rather than the target distribution.

    The discriminator is trained to minimize the loss: LD(θD∣W)=−1n∑i=1nlog⁡PθD(source=1∣Wxi)−1m∑i=1mlog⁡PθD(source=0∣yi)\mathcal{L}_D(\theta_D \mid W) = -\frac{1}{n} \sum_{i=1}^n \log P_{\theta_D}(\text{source} = 1 \mid W x_i) - \frac{1}{m} \sum_{i=1}^m \log P_{\theta_D}(\text{source} = 0 \mid y_i)

    The mapping WW is trained to fool the discriminator by minimizing: LW(W∣θD)=−1n∑i=1nlog⁡PθD(source=0∣Wxi)−1m∑i=1mlog⁡PθD(source=1∣yi)\mathcal{L}_W(W \mid \theta_D) = -\frac{1}{n} \sum_{i=1}^n \log P_{\theta_D}(\text{source} = 0 \mid W x_i) - \frac{1}{m} \sum_{i=1}^m \log P_{\theta_D}(\text{source} = 1 \mid y_i)

    Key implementation details:

    • Inputs: The discriminator is trained exclusively on the 50,000 most frequent words in each language, sampled uniformly at each step, to prevent degraded representations of rare words from harming alignment.
    • Discriminator Architecture: A multilayer perceptron with two hidden layers of 2048 units each, Leaky-ReLU activations, input dropout with rate 0.1, and prediction label smoothing s=0.2s = 0.2.
    • Optimization: Stochastic gradient descent with batch size 32, learning rate 0.1, and decay rate 0.95 applied to both the discriminator and WW. The learning rate is halved whenever the unsupervised validation metric decreases.
  3. Knowl 3 — Multiplicative Orthogonality Update Rule for Embedding Mapping

    model/method

    To preserve the monolingual geometry (such as ℓ2\ell_2 norms and dot products) of the source embeddings and stabilize adversarial optimization, the linear mapping matrix W∈Rd×dW \in \mathbb{R}^{d \times d} is constrained to be approximately orthogonal (W∈Od(R)W \in \mathcal{O}_d(\mathbb{R})).

    During adversarial training, standard gradient descent updates on WW are alternated with the following regularized update step: W←(1+β)W−β(WWT)WW \leftarrow (1 + \beta) W - \beta (W W^T) W where β=0.01\beta = 0.01. This step projects WW toward the manifold of orthogonal matrices after each update, ensuring that the eigenvalues of WW remain close to unit modulus.

  4. Knowl 4 — Cross-Domain Similarity Local Scaling (CSLS)

    definition

    Cross-Domain Similarity Local Scaling (CSLS) is a similarity metric between mapped source embeddings WxsW x_s and target embeddings yty_t designed to mitigate the hubness problem in high dimensions, where certain vectors become nearest neighbors to an abnormally high number of query vectors.

    CSLS constructs a bipartite neighborhood graph where each source vector is linked to its KK nearest target neighbors NT(Wxs)\mathcal{N}_T(W x_s) and each target vector is linked to its KK nearest source neighbors NS(yt)\mathcal{N}_S(y_t) according to cosine similarity cos⁡(u,v)=u⋅v∥u∥2∥v∥2\cos(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2}.

    The mean cosine similarity of a mapped source vector WxsW x_s to its target neighborhood is: rT(Wxs)=1K∑y∈NT(Wxs)cos⁡(Wxs,y)r_T(W x_s) = \frac{1}{K} \sum_{y \in \mathcal{N}_T(W x_s)} \cos(W x_s, y)

    The mean cosine similarity of a target vector yty_t to its mapped source neighborhood is: rS(yt)=1K∑Wx∈NS(yt)cos⁡(Wx,yt)r_S(y_t) = \frac{1}{K} \sum_{W x \in \mathcal{N}_S(y_t)} \cos(W x, y_t)

    The CSLS similarity score is defined as: CSLS(Wxs,yt)=2cos⁡(Wxs,yt)−rT(Wxs)−rS(yt)\text{CSLS}(W x_s, y_t) = 2 \cos(W x_s, y_t) - r_T(W x_s) - r_S(y_t)

    CSLS increases the similarity score for isolated word vectors and decreases it for vectors in dense regions (hubs). Using K=10K = 10 provides strong, robust performance without requiring cross-validation.

  5. Knowl 5 — Iterative Procrustes Refinement via Synthetic Mutual Nearest-Neighbor Dictionaries

    algorithm

    Given an initial linear mapping W∈Rd×dW \in \mathbb{R}^{d \times d} obtained via adversarial training and monolingual embedding sets X⊂Rd\mathcal{X} \subset \mathbb{R}^d and Y⊂Rd\mathcal{Y} \subset \mathbb{R}^d, the mapping is refined using an orthogonal Procrustes step on a synthetic dictionary:

    Input: Initial mapping WW, source embeddings X\mathcal{X}, target embeddings Y\mathcal{Y}, vocabulary limit VV, neighborhood size KK
    Output: Refined orthogonal mapping matrix W∗W^*
    Select the VV most frequent words from X\mathcal{X} and Y\mathcal{Y}
    Compute pairwise CSLS similarity scores CSLS(Wxs,yt)\text{CSLS}(W x_s, y_t) using neighborhood size KK
    Construct synthetic dictionary D={(xs,yt)}\mathcal{D} = \{(x_s, y_t)\} containing only mutual nearest neighbors:
        yt=arg⁡max⁡yCSLS(Wxs,y)y_t = \arg\max_y \text{CSLS}(W x_s, y) and xs=arg⁡max⁡xCSLS(Wx,yt)x_s = \arg\max_x \text{CSLS}(W x, y_t)
    Form matrices XD∈Rd×∣D∣X_{\mathcal{D}} \in \mathbb{R}^{d \times |\mathcal{D}|} and YD∈Rd×∣D∣Y_{\mathcal{D}} \in \mathbb{R}^{d \times |\mathcal{D}|} with the aligned word embeddings
    Compute singular value decomposition: UΣVT=SVD(YDXDT)U \Sigma V^T = \text{SVD}(Y_{\mathcal{D}} X_{\mathcal{D}}^T)
    Update mapping: W∗=UVTW^* = U V^T
    return W∗W^*

    The resulting matrix W∗W^* solves the Orthogonal Procrustes problem: W∗=arg⁡min⁡W∈Od(R)∥WXD−YD∥FW^* = \arg\min_{W \in \mathcal{O}_d(\mathbb{R})} \|W X_{\mathcal{D}} - Y_{\mathcal{D}}\|_F Because the synthetic dictionary constructed from adversarial initialization is already high quality, performing a single iteration is usually sufficient; additional iterations yield marginal improvements of less than 1%.

  6. Knowl 6 — Unsupervised Validation Metric for Cross-Lingual Model Selection

    model/method

    Because cross-lingual parallel validation data is unavailable in the unsupervised setting, model selection, early stopping, and hyperparameter tuning are conducted using an unsupervised validation metric that quantifies the geometric closeness between mapped source and target spaces:

    1. Select the 10,000 most frequent source words.
    2. For each selected source word xix_i, determine its translation candidate y(i)y_{(i)} in the target space via nearest-neighbor retrieval under Cross-Domain Similarity Local Scaling (CSLS): y(i)=arg⁡max⁡yCSLS(Wxi,y)y_{(i)} = \arg\max_y \text{CSLS}(W x_i, y).
    3. Compute the average cosine similarity across these 10,000 deemed translation pairs: Validation Score=110000∑i=110000cos⁡(Wxi,y(i))\text{Validation Score} = \frac{1}{10000} \sum_{i=1}^{10000} \cos(W x_i, y_{(i)})

    This unsupervised criterion correlates strongly with actual bilingual word translation precision and is used both to trigger learning rate decay during training and to select optimal hyperparameters across different language pairs.

  7. Knowl 7 — Bilingual Word Translation Retrieval Performance Across Language Pairs

    data/table

    Word translation retrieval Precision@1 (P@1, in %) evaluated on 1,500 source test queries against a 200,000-word target vocabulary using 300-dimensional fastText embeddings trained on Wikipedia.

    Method en-es es-en en-fr fr-en en-de de-en
    Supervised with fastText
    Procrustes - NN 77.4 77.3 74.9 76.1 68.4 67.7
    Procrustes - ISF 81.1 82.6 81.1 81.3 71.1 71.5
    Procrustes - CSLS 81.4 82.9 81.1 82.4 73.5 72.4
    Unsupervised with fastText
    Adv - NN 69.8 71.3 70.4 61.9 63.1 59.6
    Adv - CSLS 75.7 79.7 77.8 71.2 70.1 66.4
    Adv - Refine - NN 79.1 78.1 78.1 78.2 71.3 69.6
    Adv - Refine - CSLS 81.7 83.3 82.3 82.1 74.0 72.2
    Method en-ru ru-en en-zh zh-en en-eo eo-en
    Supervised with fastText
    Procrustes - NN 47.0 58.2 40.6 30.2 22.1 20.4
    Procrustes - ISF 49.5 63.8 35.7 37.5 29.0 27.9
    Procrustes - CSLS 51.7 63.7 42.7 36.7 29.3 25.3
    Unsupervised with fastText
    Adv - NN 29.1 41.5 18.5 22.3 13.5 12.1
    Adv - CSLS 37.2 48.1 23.4 28.3 18.6 16.6
    Adv - Refine - NN 37.3 54.3 30.9 21.9 20.7 20.6
    Adv - Refine - CSLS 44.0 59.1 32.5 31.4 28.2 25.6

    The unsupervised method combining adversarial initialization, Procrustes refinement, and CSLS retrieval (Adv - Refine - CSLS) matches or slightly exceeds the supervised Procrustes baseline across European languages (e.g., 81.7% vs 81.4% on en-es; 82.3% vs 81.1% on en-fr) and effectively aligns distant languages with non-Latin alphabets (Russian, Chinese) as well as lower-resource languages (Esperanto).

  8. Knowl 8 — Comparison of Unsupervised and Supervised Methods on English-Italian Benchmark

    data/table

    Word translation average precisions (P@1, P@5, P@10 in %) evaluated on the standard English-Italian benchmark using 1,500 source queries and 200k target words across WaCky and Wikipedia embeddings.

    English to Italian Italian to English
    Method P@1 P@5 P@10 P@1 P@5 P@10
    Supervised (WaCky corpora)
    Mikolov et al. (2013b) 33.8 48.3 53.9 24.9 41.0 47.4
    Dinu et al. (2015) 38.5 56.4 63.9 24.6 45.4 54.1
    CCA 36.1 52.7 58.1 31.0 49.9 57.0
    Artetxe et al. (2017) 39.7 54.7 60.5 33.8 52.4 59.1
    Smith et al. (2017) 43.1 60.7 66.4 38.0 58.5 63.6
    Procrustes - CSLS 44.9 61.8 66.6 38.5 57.2 63.0
    Unsupervised (WaCky corpora)
    Adv - Refine - CSLS 45.1 60.7 65.1 38.3 57.8 62.8
    Supervised (Wikipedia fastText)
    Procrustes - CSLS 63.7 78.6 81.1 56.3 76.2 80.6
    Unsupervised (Wikipedia fastText)
    Adv - Refine - CSLS 66.2 80.4 83.4 58.7 76.5 80.9

    On identical WaCky embeddings, unsupervised Adv - Refine - CSLS achieves 45.1% P@1 on English-to-Italian, outperforming all previous supervised models. Using fastText Wikipedia embeddings, it reaches 66.2% P@1, surpassing the supervised Procrustes baseline (63.7% P@1).

  9. Knowl 9 — Cross-Lingual Sentence Translation Retrieval and Word Similarity Performance

    empirical result

    The cross-lingual embedding representations were validated on sentence-level and semantic similarity tasks:

    1. Sentence Translation Retrieval (Europarl): Using IDF-weighted bag-of-words sentence embeddings for 2,000 source sentence queries against 200,000 target sentences, unsupervised Adv - Refine - CSLS achieves 65.9% P@1 (79.7% P@5, 83.1% P@10) on English-to-Italian and 69.0% P@1 (79.7% P@5, 83.1% P@10) on Italian-to-English. Supervised Procrustes - CSLS achieves 66.1% and 69.5% P@1 respectively, whereas prior supervised approaches reached at most 54.6% (en-it) and 48.9% (it-en).
    2. Cross-Lingual Word Similarity (SemEval 2017): Pearson correlations with human similarity scores:
      • English-Spanish (en-es): Supervised baseline 0.72, Adv 0.69, Adv - Refine 0.71 (official SemEval NASARI baseline: 0.64).
      • English-German (en-de): Supervised baseline 0.72, Adv 0.70, Adv - Refine 0.71 (SemEval NASARI baseline: 0.60).
      • English-Italian (en-it): Supervised baseline 0.71, Adv 0.67, Adv - Refine 0.71 (SemEval NASARI baseline: 0.65).

    Across both sentence retrieval and cross-lingual semantic similarity, the unsupervised refinement model matches supervised performance.

  10. Knowl 10 — Unsupervised Word-by-Word Machine Translation on English-Esperanto

    empirical result

    The unsupervised dictionary generation method was evaluated on machine translation for the low-resource language pair English-Esperanto using Wikipedia fastText embeddings and ~60,000 vocabulary-filtered sentence pairs from the Tatoeba corpus. Sentences were translated word-by-word without syntactic reordering or a target language model:

    Dictionary Retrieval Metric English-to-Esperanto BLEU Esperanto-to-English BLEU
    Nearest Neighbors (NN) 6.1 11.9
    CSLS 11.1 14.3

    Constructing the dictionary via CSLS improves translation by 5.0 BLEU points on English-to-Esperanto and 2.4 BLEU points on Esperanto-to-English compared to standard nearest-neighbor retrieval. On bilingual word translation retrieval, the unsupervised approach obtains 28.2% P@1 (46.5% P@5) on English-to-Esperanto and 25.6% P@1 (43.9% P@5) on Esperanto-to-English.

  11. Knowl 11 — Sensitivity of Embedding Alignment to Training Corpora and Architectures

    empirical result

    An empirical analysis of aligning two sets of English monolingual word embeddings reveals how corpus consistency and embedding models affect alignment quality across word frequency tiers (from 5k–7k to 150k–152k):

    • Identical corpus, identical model (Skip-Gram on Wikipedia with different seeds): Achieves near-perfect retrieval accuracy across all frequency bands (99.7% to 100.0%).
    • Identical corpus, different models (Skip-Gram to CBOW on Wikipedia): Maintains high accuracy across all tiers when using CSLS (96.3% on rare words at rank 150k–152k, versus 87.3% with nearest neighbors).
    • Different corpora (Wikipedia to Gigaword with fastText): Accuracy drops substantially on rare words (69.3% with nearest neighbors, 86.7% with CSLS at rank 150k–152k) due to divergence in co-occurrence statistics between corpora.
    • Different corpora and different models (Skip-Gram on Wikipedia to fastText on Gigaword): Performance further degrades to 48.0% with nearest neighbors and 67.3% with CSLS at rank 150k–152k.

    These results demonstrate that similarity in co-occurrence statistics of the underlying training corpora is the primary driver of embedding space alignability.

Coverage note — None was omitted; all primary contributed methods, algorithms, loss functions, validation criteria, and experimental evaluations across word translation, sentence retrieval, semantic similarity, machine translation, and corpus compatibility analyses are covered.

References

  1. 1.Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A Smith. Massively multilingual word embeddings. arXiv preprint arXiv:1602.01925, 2016.
  2. 2.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. Proceedings of EMNLP, 2016.
  3. 3.Mikel Artetxe, Gorka Labaka, and Eneko Agirre. Learning bilingual word embeddings with (almost) no bilingual data. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 451–462. Association for Computational Linguistics, 2017.
  4. 4.Marco Baroni, Silvia Bernardini, Adriano Ferraresi, and Eros Zanchetta. The wacky wide web: a collection of very large linguistically processed web-crawled corpora. Language resources and evaluation, 43(3):209–226, 2009.
  5. 5.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5: 135–146, 2017.
  6. 6.Jose Camacho-Collados, Mohammad Taher Pilehvar, and Roberto Navigli. Nasari: Integrating ex- plicit knowledge and corpus statistics for a multilingual representation of concepts and entities. Artificial Intelligence, 240:36–64, 2016.
  7. 7.Jose Camacho-Collados, Mohammad Taher Pilehvar, Nigel Collier, and Roberto Navigli. Semeval2017 task 2: Multilingual and cross-lingual semantic word similarity. Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval 2017), 2017.
  8. 8.Hailong Cao, Tiejun Zhao, Shu Zhang, and Yao Meng. A distribution-based model to learn bilingual word embeddings. Proceedings of COLING, 2016.
  9. 9.Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. Parseval networks: Improving robustness to adversarial examples. International Conference on Machine Learning, pp. 854–863, 2017.
  10. 10.Georgiana Dinu, Angeliki Lazaridou, and Marco Baroni. Improving zero-shot learning by mitigating the hubness problem. International Conference on Learning Representations, Workshop Track, 2015.
  11. 11.Qing Dou, Ashish Vaswani, Kevin Knight, and Chris Dyer. Unifying bayesian inference and vector space models for improved decipherment. 2015.
  12. 12.Long Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird, and Trevor Cohn. Learning crosslingual word embeddings without bilingual corpora. Proceedings of EMNLP, 2016.
  13. 13.Manaal Faruqui and Chris Dyer. Improving vector space word representations using multilingual correlation. Proceedings of EACL, 2014.
  14. 14.Pascale Fung. Compiling bilingual lexicon entries from a non-parallel english-chinese corpus. In Proceedings of the Third Workshop on Very Large Corpora, pp. 173–183, 1995.
  15. 15.Pascale Fung and Lo Yuen Yee. An ir approach for translating new words from nonparallel, comparable texts. In Proceedings of the 17th International Conference on Computational Linguistics - Volume 1, COLING ’98, pp. 414–420. Association for Computational Linguistics, 1998.
  16. 16.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Francois Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks. Journal of Machine Learning Research, 17(59):1–35, 2016.
  17. 17.Ian Goodfellow. Nips 2016 tutorial: Generative adversarial networks. arXiv preprint arXiv:1701.00160, 2016.
  18. 18.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advancesin neural information processing systems, pp. 2672–2680, 2014.
  19. 19.Stephan Gouws, Yoshua Bengio, and Greg Corrado. Bilbowa: Fast bilingual distributed representations without word alignments. In Proceedings of the 32nd International Conference on Machine Learning (ICML-15), pp. 748–756, 2015.
  20. 20.Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, and Dan Klein. Learning bilingual lexicons from monolingual corpora. In Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics, 2008.
  21. 21.Zellig S Harris. Distributional structure. Word, 10(2-3):146–162, 1954.
  22. 22.Ann Irvine and Chris Callison-Burch. Supervised bilingual lexicon induction with multiple monolingual signals. In HLT-NAACL, 2013.
  23. 23.Herve Jegou, Cordelia Schmid, Hedi Harzallah, and Jakob Verbeek. Accurate image search using the contextual dissimilarity measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(1):2–11, 2010.
  24. 24.Jeff Johnson, Matthijs Douze, and Herve J egou. Billion-scale similarity search with gpus.  arXiv preprint arXiv:1702.08734, 2017.
  25. 25.Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. Inducing crosslingual distributed representations of words. Proceedings of COLING, pp. 1459–1474, 2012.
  26. 26.Philipp Koehn and Kevin Knight. Learning a translation lexicon from monolingual corpora. In Proceedings of the ACL-02 workshop on Unsupervised lexical acquisition-Volume 9, pp. 9–16. Association for Computational Linguistics, 2002.
  27. 27.Grzegorz Kondrak, Bradley Hauer, and Garrett Nicolai. Bootstrapping unsupervised bilingual lexicon induction. In EACL, 2017.
  28. 28.Angeliki Lazaridou, Georgiana Dinu, and Marco Baroni. Hubness and pollution: Delving into crossspace mapping for zero-shot learning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, 2015.
  29. 29.Omer Levy and Yoav Goldberg. Neural word embedding as implicit matrix factorization. Advances in neural information processing systems, pp. 2177–2185, 2014.
  30. 30.Thang Luong, Richard Socher, and Christopher D Manning. Better word representations with recursive neural networks for morphology. CoNLL, pp. 104–113, 2013.
  31. 31.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. Proceedings of Workshop at ICLR, 2013a.
  32. 32.Tomas Mikolov, Quoc V Le, and Ilya Sutskever. Exploiting similarities among languages for machine translation. arXiv preprint arXiv:1309.4168, 2013b.
  33. 33.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, pp. 3111–3119, 2013c.
  34. 34.Robert Parker, David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. English gigaword. Linguistic Data Consortium, 2011.
  35. 35.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. Proceedings of EMNLP, 14:1532–1543, 2014.
  36. 36.N. Pourdamghani and K. Knight. Deciphering related languages. In EMNLP, 2017.
  37. 37.Miloš Radovanovi ȇ c, Alexandros Nanopoulos, and Mirjana Ivanovi  c. Hubs in space: Popular nearest  neighbors in high-dimensional data. Journal of Machine Learning Research, 11(Sep):2487–2531, 2010.
  38. 38.Reinhard Rapp. Identifying word translations in non-parallel texts. In Proceedings of the 33rd Annual Meeting on Association for Computational Linguistics, ACL ’95, pp. 320–322. Association for Computational Linguistics, 1995.
  39. 39.Reinhard Rapp. Automatic identification of word translations from unrelated english and german corpora. In Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics, ACL ’99. Association for Computational Linguistics, 1999.
  40. 40.S. Ravi and K. Knight. Deciphering foreign language. In ACL, 2011.
  41. 41.Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas. The earth mover’s distance as a metric for image retrieval. International journal of computer vision, 40(2):99–121, 2000.
  42. 42.Charles Schafer and David Yarowsky. Inducing translation lexicons via diverse similarity measures and bridge languages. In Proceedings of the 6th Conference on Natural Language Learning - Volume 20, COLING-02. Association for Computational Linguistics, 2002.
  43. 43.Peter H Schonemann. A generalized solution of the orthogonal procrustes problem.  Psychometrika, 31(1):1–10, 1966.
  44. 44.Samuel L Smith, David HP Turban, Steven Hamblin, and Nils Y Hammerla. Offline bilingual word vectors, orthogonal transformations and the inverted softmax. International Conference on Learning Representations, 2017.
  45. 45.Jrg Tiedemann. Parallel data, tools and interfaces in opus. In Nicoletta Calzolari (Conference Chair), Khalid Choukri, Thierry Declerck, Mehmet Uur Doan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey, may 2012. European Language Resources Association (ELRA). ISBN 978-2-9517408-7-7.
  46. 46.Shinji Umeyama. An eigendecomposition approach to weighted graph matching problems. IEEE transactions on pattern analysis and machine intelligence, 10(5):695–703, 1988.
  47. 47.Ivan Vulic and Marie-Francine Moens. Bilingual word embeddings from non-parallel documentaligned data applied to bilingual lexicon induction. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics (ACL 2015), pp. 719–725, 2015.
  48. 48.Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. Normalized word embedding and orthogonal transform for bilingual word translation. Proceedings of NAACL, 2015.
  49. 49.Lihi Zelnik-manor and Pietro Perona. Self-tuning spectral clustering. In L. K. Saul, Y. Weiss, and L. Bottou (eds.), Advances in Neural Information Processing Systems 17, pp. 1601–1608. MIT Press, 2005.
  50. 50.Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. Earth mover’s distance minimization for unsupervised bilingual lexicon induction. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 1924–1935. Association for Computational Linguistics, 2017a.
  51. 51.Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. Adversarial training for unsupervised bilingual lexicon induction. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, 2017b.
  52. 52.Will Y Zou, Richard Socher, Daniel M Cer, and Christopher D Manning. Bilingual word embeddings for phrase-based machine translation. Proceedings of EMNLP, 2013.

Citation

MLA
Conneau, A., et al. “Word Translation Without Parallel Data”. arXiv, 2017, http://arxiv.org/abs/1710.04087v3.
APA
Conneau, A., Lample, G., Ranzato, M., Denoyer, L., & Jégou, H. (2017). Word Translation Without Parallel Data. arXiv. http://arxiv.org/abs/1710.04087v3
Chicago
Conneau, A., G. Lample, M. Ranzato, L. Denoyer, and H. Jégou. 2017. “Word Translation Without Parallel Data”. arXiv. http://arxiv.org/abs/1710.04087v3.
Harvard
Conneau, A. et al. (2017) “Word Translation Without Parallel Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1710.04087v3.
Vancouver
1. Conneau A, Lample G, Ranzato M, Denoyer L, Jégou H (2017) Word Translation Without Parallel Data. arXiv

BibTeX

@article{conneau2017word,
  title = {Word Translation Without Parallel Data},
  author = {Conneau, Alexis and Lample, Guillaume and Ranzato, Marc'Aurelio and Denoyer, Ludovic and Jégou, Hervé},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1710.04087v3},
  eprint = {1710.04087}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors