Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation
Wei JinHaitao MaoZheng LiHaoming JiangChen LuoHongzhi WenHaoyu HanHanqing LuZhengyang WangRuirui Li
Presents Amazon-M2, a large-scale shopping session dataset spanning millions of interactions across six languages and locales with rich text attributes, establishing new benchmarks for multilingual session-based recommendation, cross-market domain transfer, and product title generation.
Online retailers increasingly depend on session-based recommender systems to anticipate customer intent during active browsing windows, especially when shoppers browse anonymously without historical profiles. Existing benchmark datasets for this task present major constraints: they lack rich textual attributes, omit diverse geographic and linguistic contexts, and cover limited product catalogs. To overcome these limitations, the article introduces Amazon-M2, the first large-scale, multilingual, multi-locale shopping session benchmark designed to advance personalized recommendation and text generation algorithms.
Amazon-M2 comprises real-world customer sessions spanning six global locales—the United Kingdom, Germany, Japan, Spain, France, and Italy—covering six major languages and over 1.4 million unique products across more than 3.6 million training sessions. The dataset provides extensive metadata, including product titles, descriptions, brands, and prices. The article formulates three distinct evaluation tasks: next-product recommendation within major regions, cross-locale next-product recommendation with domain shifts from data-rich to underrepresented markets, and a novel task for generating the title of the next unseen product in a session.
Evaluation of leading deep learning models alongside simple heuristics revealed several key findings. First, a basic popularity baseline frequently outperformed sophisticated neural networks on ranking metrics (Mean Reciprocal Rank), demonstrating that product popularity remains a powerful bias in large catalogs. Second, in cross-domain transfer tasks, pre-training models on major markets before fine-tuning on data-sparse regions significantly improved performance across both recall and ranking metrics. Third, although deep models retrieved relevant items effectively—achieving high recall scores of 65% to 75%—they struggled to rank them near the top of recommendation lists. Finally, in product title generation, a simple heuristic that copies the previous item's title outperformed fine-tuned multilingual text models, highlighting the critical importance of the immediately preceding interaction.
These findings indicate that existing recommendation algorithms and general-purpose language models are not fully equipped to navigate complex e-commerce catalogs or exploit unstructured textual metadata out of the box. Simply incorporating standard off-the-shelf text embeddings can degrade recommendation quality if the pre-training domain does not match product catalog structures. Transfer learning offers a cost-effective pathway to enhance recommendation quality in low-resource emerging markets without requiring massive local data collection, but models require better ranking refinement to turn high retrieval rates into immediate conversions.
Organizations developing e-commerce recommender systems should adopt pre-training and fine-tuning frameworks to bootstrap algorithms across smaller international locales while focusing engineering efforts on multi-stage ranking mechanisms. Further research is necessary to design domain-specific language models that better understand product text features and to address data imbalances. Because the data originates solely from Amazon's platform over a three-week window, stakeholders should exercise caution when generalizing these findings to non-retail domains or vastly different user behavioral settings.
- Paper: Session-based Recommendations with Recurrent Neural Networks, Balázs Hidasi et al. (2016). This seminal paper introduced session-based recommendation without user profiles using neural sequence modeling, establishing the fundamental problem formulation benchmarked by Amazon-M2.
- Paper: Self-Attentive Sequential Recommendation, Wang-Cheng Kang et al. (2018). It introduces SASRec, the standard self-attentive sequential recommendation architecture evaluated as a primary neural baseline in session datasets like Amazon-M2.
- Paper: BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer, Fei Sun et al. (2019). It provides the BERT4Rec architecture for sequential recommendation using masked transformer modeling, representing one of the key sequence models tested on session interaction data.
- Paper: Session-based Recommendation with Graph Neural Networks, Shu Wu et al. (2018). It presents graph neural network architectures for session-based next-item prediction, which form a major class of baseline models evaluated against simpler heuristics in shopping sessions.
- Paper: Are we really making much progress? A worrying analysis of recent neural recommendation approaches, Maurizio Ferrari Dacrema et al. (2019). This critique showed that simple heuristics frequently outperform complex neural recommendation algorithms, directly anticipating the empirical findings observed on the Amazon-M2 benchmark.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). It details multilingual denoising sequence-to-sequence pre-training (mBART), providing the foundation for the multilingual text generation tasks and cross-locale transfer explored in the benchmark.
- Paper: Amazon.com recommendations: item-to-item collaborative filtering, Greg Linden et al. (2003). It outlines the foundational item-to-item co-occurrence principles in Amazon e-commerce systems that underpin the strong competitive performance of session co-visitation heuristics.
- Paper: An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation, Haoyu Han et al. (2026). This work directly follows and builds upon observations from datasets like Amazon-M2 by proving that simple local transition graph heuristics can match or exceed complex generative sequential models.
- Paper: GenRec: An LLM-Backed Recommendation Ranker at Netflix, Ying Li et al. (2026). It extends the practical challenge of integrating large language models into real-world catalog ranking pipelines by deploying an LLM-backed recommendation ranker at scale.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). It scales multilingual text embedding evaluation across hundreds of languages and tasks, helping to evaluate text encoders suited for multilingual e-commerce metadata.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). It introduces versatile multilingual embeddings capable of unified dense, sparse, and multi-vector retrieval, addressing text representation challenges identified in multi-locale item catalogs.
