Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation

Wei JinHaitao MaoZheng LiHaoming JiangChen LuoHongzhi WenHaoyu HanHanqing LuZhengyang WangRuirui Li

article2023NeurIPS63 citations

Presents Amazon-M2, a large-scale shopping session dataset spanning millions of interactions across six languages and locales with rich text attributes, establishing new benchmarks for multilingual session-based recommendation, cross-market domain transfer, and product title generation.

Listen

Online retailers increasingly depend on session-based recommender systems to anticipate customer intent during active browsing windows, especially when shoppers browse anonymously without historical profiles. Existing benchmark datasets for this task present major constraints: they lack rich textual attributes, omit diverse geographic and linguistic contexts, and cover limited product catalogs. To overcome these limitations, the article introduces Amazon-M2, the first large-scale, multilingual, multi-locale shopping session benchmark designed to advance personalized recommendation and text generation algorithms.

Amazon-M2 comprises real-world customer sessions spanning six global locales—the United Kingdom, Germany, Japan, Spain, France, and Italy—covering six major languages and over 1.4 million unique products across more than 3.6 million training sessions. The dataset provides extensive metadata, including product titles, descriptions, brands, and prices. The article formulates three distinct evaluation tasks: next-product recommendation within major regions, cross-locale next-product recommendation with domain shifts from data-rich to underrepresented markets, and a novel task for generating the title of the next unseen product in a session.

Evaluation of leading deep learning models alongside simple heuristics revealed several key findings. First, a basic popularity baseline frequently outperformed sophisticated neural networks on ranking metrics (Mean Reciprocal Rank), demonstrating that product popularity remains a powerful bias in large catalogs. Second, in cross-domain transfer tasks, pre-training models on major markets before fine-tuning on data-sparse regions significantly improved performance across both recall and ranking metrics. Third, although deep models retrieved relevant items effectively—achieving high recall scores of 65% to 75%—they struggled to rank them near the top of recommendation lists. Finally, in product title generation, a simple heuristic that copies the previous item's title outperformed fine-tuned multilingual text models, highlighting the critical importance of the immediately preceding interaction.

These findings indicate that existing recommendation algorithms and general-purpose language models are not fully equipped to navigate complex e-commerce catalogs or exploit unstructured textual metadata out of the box. Simply incorporating standard off-the-shelf text embeddings can degrade recommendation quality if the pre-training domain does not match product catalog structures. Transfer learning offers a cost-effective pathway to enhance recommendation quality in low-resource emerging markets without requiring massive local data collection, but models require better ranking refinement to turn high retrieval rates into immediate conversions.

Organizations developing e-commerce recommender systems should adopt pre-training and fine-tuning frameworks to bootstrap algorithms across smaller international locales while focusing engineering efforts on multi-stage ranking mechanisms. Further research is necessary to design domain-specific language models that better understand product text features and to address data imbalances. Because the data originates solely from Amazon's platform over a three-week window, stakeholders should exercise caution when generalizing these findings to non-retail domains or vastly different user behavioral settings.

arXiv: 2307.09688
Cover for Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation

Abstract

Modeling customer shopping intentions is a crucial task for e-commerce, as it directly impacts user experience and engagement. Thus, accurately understanding customer preferences is essential for providing personalized recommendations. Session-based recommendation, which utilizes customer session data to predict their next interaction, has become increasingly popular. However, existing session datasets have limitations in terms of item attributes, user diversity, and dataset scale. As a result, they cannot comprehensively capture the spectrum of user behaviors and preferences. To bridge this gap, we present the Amazon Multilingual Multi-locale Shopping Session Dataset, namely Amazon-M2. It is the first multilingual dataset consisting of millions of user sessions from six different locales, where the major languages of products are English, German, Japanese, French, Italian, and Spanish. Remarkably, the dataset can help us enhance personalization and understanding of user preferences, which can benefit various existing tasks as well as enable new tasks. To test the potential of the dataset, we introduce three tasks in this work: (1) next-product recommendation, (2) next-product recommendation with domain shifts, and (3) next-product title generation. With the above tasks, we benchmark a range of algorithms on our proposed dataset, drawing new insights for further research and practice. In addition, based on the proposed dataset and tasks, we hosted a competition in the KDD CUP 2023² and have attracted thousands of users and submissions. The winning solutions and the associated workshop can be accessed at our website https://kddcup23.github.io/.

Table of Contents

  • 1 Introduction
  • 2 Dataset & Task Description
  • 2.1 Dataset Description
  • 2.2 Task Description
  • 3 Dataset Analysis
  • 4 Benchmark on the Proposed Three Tasks
  • 4.1 Task 1. Next-product Recommendation
  • 4.2 Task 2. Next-product Recommendation with Domain Shifts
  • 4.3 Task 3. Next-product Title Generation
  • 5 Discussion
  • 6 Conclusion
  • 7 Acknowledgement
  • References
  • A Experimental Setup
  • B More Dataset Details
  • B.1 Dataset Collection
  • B.2 Additional data analysis
  • B.3 License
  • B.4 Extended Discussion
  • C More Experimental Results
  • C.1 Additional results on NDCG@100
  • C.2 Task 1. Next-product Recommendation
  • C.3 Task 2. Next-product Recommendation with Domain Shifts
  • C.4 Task 3. Next-product Title Prediction
  • D Limitation & Broader Impact

Knowls

  1. Knowl 1 — Amazon-M2 Dataset Specification and Attribute Schema

    definition

    The Amazon Multilingual Multi-locale Shopping Session Dataset (Amazon-M2) is a large-scale dataset of anonymized customer shopping sessions paired with detailed product metadata collected from Amazon across six locales: the United Kingdom (UK), Germany (DE), Japan (JP), Spain (ES), Italy (IT), and France (FR).

    The dataset comprises two core components:

    1. User Sessions: A set of user sessions S={s1,s2,…,sn}\mathcal{S} = \{s_1, s_2, \dots, s_n\}, where each session s={e1,e2,…,et}s = \{e_1, e_2, \dots, e_t\} is a chronologically ordered sequence of product identifiers interacted with by an anonymous user within a 30-minute active window.
    2. Product Attribute Table: A product dictionary V={v1,v2,…,vm}\mathcal{V} = \{v_1, v_2, \dots, v_m\} mapping each unique Amazon Standard Identification Number (ASIN) to categorical, numerical, and multilingual textual attributes across six primary languages (English, German, Japanese, Spanish, Italian, and French).

    The attributes recorded for each product include:

    • ASIN (id): Unique Amazon Standard Identification Number.
    • Locale: Regional market code.
    • Title: Descriptive textual name of the item.
    • Brand: Manufacturer or commercial brand name.
    • Size: Dimensions or size specification.
    • Model: Specific product model or variation identifier.
    • Material Type: Composition material (e.g., metal, plastic, wood).
    • Color Text: Textual description of the color variant.
    • Author: Author name (for book products).
    • Bullet Description (desc): Bulleted textual summary of key selling points and product features.
  2. Knowl 2 — Amazon-M2 Dataset Summary Statistics Across Locales

    data/table

    The Amazon-M2 dataset contains 1,410,675 unique products, 3,606,249 training sessions (15,306,183 total train interactions), and 361,659 test sessions (1,483,822 total test interactions) spanning a three-week collection window (first two weeks for training, final week divided into two test phases). Major locales (UK, DE, JP) contain roughly 10 times more sessions and items than underrepresented locales (ES, FR, IT).

    Train Test
    Locale #Products #Sessions #Interactions Avg. Length #Sessions #Interactions Avg. Length
    UK 500,180 1,182,181 4,872,506 4.1 115,936 466,265 4.0
    DE 518,327 1,111,416 4,836,983 4.4 104,568 450,090 4.3
    JP 395,009 979,119 4,388,790 4.5 96,467 434,777 4.5
    ES 42,503 89,047 326,256 3.7 8,176 31,133 3.8
    FR 44,577 117,561 416,797 3.5 12,520 48,143 3.9
    IT 50,461 126,925 464,851 3.7 13,992 53,414 3.8
    Total 1,410,675 3,606,249 15,306,183 4.2 361,659 1,483,822 4.2
  3. Knowl 3 — Amazon-M2 Benchmark Task Formulations

    definition

    The Amazon-M2 benchmark defines three distinct downstream evaluation tasks:

    1. Task 1: Next-Product Recommendation: Given an observed sequence of interacted product IDs s={e1,e2,…,et}s = \{e_1, e_2, \dots, e_t\}, predict the product ID et+1e_{t+1} that the user interacts with next. The training and test sessions are drawn from the same high-resource locale distributions (UK, DE, JP).
    2. Task 2: Next-Product Recommendation with Domain Shifts: Models are pre-trained on sessions from the high-resource locales (UK, DE, JP) and subsequently fine-tuned and evaluated on underrepresented locales (ES, FR, IT). This tests cross-lingual and cross-market knowledge transfer under distribution shift and data scarcity.
    3. Task 3: Next-Product Title Generation: Given the textual titles of the most recent KK products {et−K+1,…,et}\{e_{t-K+1}, \dots, e_t\} in a session, generate the textual title of the next product et+1e_{t+1}. The target products in the test set do not appear in the training set, evaluating cold-start generative recommendation.
  4. Knowl 4 — Session-KNN Similarity and Scoring Formulation

    equation

    In session-based collaborative filtering via Session-KNN (SKNN), the similarity between an active session ss and a candidate historical session sj∈Ss_j \in \mathcal{S} is computed using session item overlap cosine similarity: sim(s,sj)=∣s∩sj∣∣s∣⋅∣sj∣\text{sim}(s, s_j) = \frac{|s \cap s_j|}{\sqrt{|s| \cdot |s_j|}}

    Given the set N(s)⊆S\mathcal{N}(s) \subseteq \mathcal{S} of the most similar sessions to ss, the relevance score of a candidate item ee is: score(e,s)=∑n∈N(s)sim(s,n)In(e)\text{score}(e, s) = \sum_{n \in \mathcal{N}(s)} \text{sim}(s, n) I_n(e) where In(e)∈{0,1}I_n(e) \in \{0, 1\} is an indicator function taking the value 11 if session nn contains product ee, and 00 otherwise.

  5. Knowl 5 — Empirical Characteristics of Multilingual Shopping Sessions

    empirical result

    Analysis of the Amazon-M2 session corpus reveals four key statistical and behavioral phenomena:

    • Long-tail distributions: Product interaction frequencies, session lengths, and repeated item counts all exhibit steep power-law tails across all locales. While average session lengths range from 3.5 to 4.5, extreme sessions reach lengths >100>100.
    • Cross-locale item overlap: The overlap ratio between locale aa and locale bb, defined as ∣Na∩Nb∣∣Na∣\frac{|N_a \cap N_b|}{|N_a|} where Na,NbN_a, N_b are product sets, is minimal among major locales (UK-DE is 13%, JP vs. all others ≤1%\le 1\%, UK/DE vs. JP ≤3%\le 3\%). Conversely, underrepresented locales share a high fraction of their items with Germany (DE contains 40% of ES items, 45% of FR items, and 37% of IT items), establishing structural feasibility for cross-market transfer.
    • Repeat consumption patterns: Approximately 35% of all sessions across all six locales contain repeated interactions with the same item within the active window.
    • Collaborative filtering density: In SKNN retrieval analysis, the distribution of the 10th-highest similarity score across sessions concentrates near 1.0, indicating sufficient inter-session overlap to retrieve at least 10 relevant candidate items per session.
  6. Knowl 6 — Ranking and Recommendation Evaluation Metrics

    equation

    Session recommendation quality is assessed using Mean Reciprocal Rank at depth KK (MRR@K\text{MRR}@K) and Hit Rate / Recall at depth KK (Recall@K\text{Recall}@K or Hit@K\text{Hit}@K):

    MRR@K=1N∑t∈T1Rank(t)\text{MRR}@K = \frac{1}{N} \sum_{t \in T} \frac{1}{\text{Rank}(t)} where N=∣T∣N = |T| is the total number of test sessions, and Rank(t)∈{1,…,K}\text{Rank}(t) \in \{1, \dots, K\} is the rank of the true target product in the top-KK recommendation list for session tt. If the ground-truth product does not appear in the top KK, 1Rank(t)=0\frac{1}{\text{Rank}(t)} = 0.

    Recall@K=Hit@K=nhitN\text{Recall}@K = \text{Hit}@K = \frac{n_{\text{hit}}}{N} where nhitn_{\text{hit}} is the count of test sessions in which the ground-truth product appears anywhere within the top-KK recommendations.

  7. Knowl 7 — Task 1 Next-Product Recommendation Benchmark

    empirical result

    On Task 1 (in-domain next-product recommendation on large locales UK, DE, and JP), evaluating top-KK ranking at K=100K=100 shows that a simple item Popularity heuristic outperforms deep sequential and graph models in MRR@100 across all locales, while CORE achieves the highest Recall@100.

    MRR@100 Recall@100
    Model UK DE JP Overall UK DE JP Overall
    Popularity 0.2302 0.2259 0.2766 0.2426 0.4047 0.4065 0.4683 0.4243
    GRU4Rec++ 0.1656 0.1578 0.2080 0.1757 0.3665 0.3627 0.4185 0.3808
    NARM 0.1801 0.1716 0.2255 0.1908 0.4021 0.4016 0.4559 0.4180
    STAMP 0.2050 0.1955 0.2521 0.2159 0.3370 0.3320 0.3900 0.3512
    SRGNN 0.1841 0.1767 0.2282 0.1948 0.3882 0.3862 0.4403 0.4032
    CORE 0.1510 0.1510 0.1846 0.1609 0.5591 0.5525 0.5898 0.5591
    MGS 0.1668 0.1739 0.2376 0.1907 0.5641 0.5479 0.4677 0.5194

    The strong performance of Popularity indicates severe popularity bias in large-scale shopping catalogs and underscores the difficulty deep ID-based session architectures face when ranking over 1.4 million items.

  8. Knowl 8 — Task 2 Cross-Locale Transfer Learning Benchmark

    empirical result

    On Task 2 (next-product recommendation under domain shift from high-resource locales JP, UK, DE to underrepresented locales ES, FR, IT), pre-training on high-resource data followed by fine-tuning consistently outperforms training solely on target-locale data across sequential and graph baselines.

    MRR@100 Recall@100
    Paradigm Method ES FR IT Overall ES FR IT Overall
    Heuristic Popularity 0.2854 0.2940 0.2706 0.2829 0.5118 0.5166 0.4914 0.5058
    Supervised GRU4Rec++ 0.2538 0.2734 0.2420 0.2564 0.5817 0.5995 0.5753 0.5856
    NARM 0.2598 0.2805 0.2493 0.2632 0.5894 0.6061 0.5803 0.5920
    STAMP 0.2555 0.2752 0.2430 0.2578 0.5162 0.5365 0.5116 0.5217
    SRGNN 0.2627 0.2797 0.2500 0.2640 0.5426 0.5543 0.5322 0.5429
    CORE 0.1978 0.2204 0.1941 0.2045 0.6931 0.7215 0.6937 0.7034
    MGS 0.2491 0.2775 0.2411 0.2560 0.5829 0.5811 0.5689 0.5766
    Pretrain GRU4Rec++ 0.2601 0.2856 0.2480 0.2646 0.5989 0.6236 0.5898 0.6042
    Finetune NARM 0.2733 0.2917 0.2564 0.2734 0.6094 0.6315 0.5974 0.6127
    STAMP 0.2701 0.2878 0.2584 0.2719 0.4792 0.4988 0.4637 0.4803
    SRGNN 0.2747 0.2980 0.2612 0.2779 0.5835 0.6101 0.5714 0.5883
    CORE 0.1685 0.1839 0.1632 0.1720 0.6946 0.7156 0.6851 0.6985
    MGS 0.2612 0.2870 0.2693 0.2722 0.5747 0.6133 0.5812 0.5913

    While Popularity achieves the highest MRR@100 (0.2829), fine-tuned neural models significantly outperform it in Recall@100 (e.g., CORE at 0.6985, NARM at 0.6127), indicating strong candidate retrieval capability but weaker precision ranking at the top ranks.

  9. Knowl 9 — Ablation of Multilingual Sentence BERT Initialization

    empirical result

    Initializing item representations using pre-trained Multilingual Sentence BERT embeddings derived from product titles degrades session recommendation performance relative to standard random ID initialization on underrepresented locales (ES, FR, IT).

    MRR@100 Recall@100
    Model ES FR IT Overall ES FR IT Overall
    SRGNN (Random ID) 0.2627 0.2797 0.2500 0.2640 0.5426 0.5543 0.5322 0.5429
    SRGNNF (Text Init) 0.2107 0.2546 0.2239 0.2303 0.4541 0.5132 0.4788 0.4976
    GRU4Rec++ (Random ID) 0.2538 0.2734 0.2420 0.2564 0.5817 0.5995 0.5753 0.5856
    GRU4RecF (Text Init) 0.2303 0.2651 0.2303 0.2427 0.4976 0.5631 0.5353 0.5454

    Overall MRR@100 drops from 0.2640 to 0.2303 for SRGNN and from 0.2564 to 0.2427 for GRU4Rec++. This performance penalty stems from domain misalignment between the generic pre-training objective of multilingual sentence encoders and the specialized vocabulary and structure of e-commerce product titles.

  10. Knowl 10 — Task 3 Generative Next-Product Title Prediction Benchmark

    empirical result

    On Task 3 (predicting the text title of the next unseen product in a session), generative language modeling baselines (mT5-small and mT5-base) conditioned on the last K∈{1,2,3}K \in \{1, 2, 3\} item titles fail to outperform the simple heuristic of copying the title of the immediately preceding product (Last Product Title).

    Method Validation BLEU Phase-1 Test BLEU Phase-2 Test BLEU
    mT5-small, K=1K = 1 0.2499 0.2265 0.2245
    mT5-small, K=2K = 2 0.2401 0.2176 0.2166
    mT5-small, K=3K = 3 0.2366 0.2142 0.2098
    mT5-base, K=1K = 1 0.2477 0.2251 0.2190
    Last Product Title 0.2500 0.2677 0.2655

    Key observations:

    1. Last Product Title achieves the highest BLEU across test sets (0.2677 phase 1, 0.2655 phase 2).
    2. Increasing the input history length KK from 1 to 3 degrades mT5 BLEU score monotonically (from 0.2265 to 0.2142 on Phase-1 test).
    3. Scaling model capacity from mT5-small to mT5-base does not yield performance gains (0.2265 vs. 0.2251 Phase-1 test BLEU).

Coverage note — None was omitted; all key contributions including dataset design, summary statistics, task formulations, algorithmic benchmarks across Tasks 1-3, and feature ablation are fully captured.

References

  1. 1.Folasade Olubusola Isinkaye, Yetunde O Folajimi, and Bolande Adefowoke Ojokoh. Recommendation systems: Principles, methods and evaluation. Egyptian informatics journal, 16(3):261–273, 2015.
  2. 2.Paul Resnick and Hal R Varian. Recommender systems. Communications of the ACM, 40(3):56–58, 1997.
  3. 3.Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, Yingqian Min, Zhichao Feng, Xinyan Fan, Xu Chen, Pengfei Wang, Wendi Ji, Yaliang Li, Xiaoling Wang, and Ji-Rong Wen. Recbole: Towards a unified, comprehensive and efficient framework for recommendation algorithms. In CIKM, pages 4653–4664. ACM, 2021.
  4. 4.Wenqi Fan, Xiaorui Liu, Wei Jin, Xiangyu Zhao, Jiliang Tang, and Qing Li. Graph trend filtering networks for recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 112–121, 2022.
  5. 5.Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. Deep learning based recommender system: A survey and new perspectives. ACM computing surveys (CSUR), 52(1):1–38, 2019.
  6. 6.Yiqiao Jin, Yunsheng Bai, Yanqiao Zhu, Yizhou Sun, and Wei Wang. Code recommendation for open source software developers. In Proceedings of the ACM Web Conference 2023, 2022.
  7. 7.Paul Covington, Jay Adams, and Emre Sargin. Deep neural networks for youtube recommendations. In Proceedings of the 10th ACM conference on recommender systems, pages 191–198, 2016.
  8. 8.Wayne Xin Zhao, Sui Li, Yulan He, Liwei Wang, Ji-Rong Wen, and Xiaoming Li. Exploring demographic information in social media for product recommendation. Knowledge and Information Systems, 49:61–89, 2016.
  9. 9.Xin Wayne Zhao, Yanwei Guo, Yulan He, Han Jiang, Yuexin Wu, and Xiaoming Li. We know what you want to buy: a demographic-based system for product recommendation on microblogs. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 1935–1944, 2014.
  10. 10.Shoujin Wang, Longbing Cao, Yan Wang, Quan Z Sheng, Mehmet A Orgun, and Defu Lian. A survey on session-based recommender systems. ACM Computing Surveys (CSUR), 54(7):1–38, 2021.
  11. 11.Malte Ludewig and Dietmar Jannach. Evaluation of session-based recommendation algorithms. User Modeling and User-Adapted Interaction, 28:331–390, 2018.
  12. 12.Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. Session-based recommendations with recurrent neural networks. arXiv preprint arXiv:1511.06939, 2015.
  13. 13.Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan. Session-based recommendation with graph neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 346–353, 2019.
  14. 14.Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization. In Proceedings of the 29th ACM international conference on information & knowledge management, pages 1893–1902, 2020.
  15. 15.Jing Li, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma. Neural attentive session-based recommendation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, pages 1419–1428, 2017.
  16. 16.Qiao Liu, Yifu Zeng, Refuoe Mokhosi, and Haibin Zhang. Stamp: short-term attention/memory priority model for session-based recommendation. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1831–1839, 2018.
  17. 17.Feng Yu, Yanqiao Zhu, Qiang Liu, Shu Wu, Liang Wang, and Tieniu Tan. Tagnn: Target attentive graph neural networks for session-based recommendation. In Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, pages 1921–1924, 2020.
  18. 18.David Ben-Shimon, Alexander Tsikinovsky, Michael Friedmann, Bracha Shapira, Lior Rokach, and Johannes Hoerle. Recsys challenge 2015 and the yoochoose dataset. In Proceedings of the 9th ACM Conference on Recommender Systems, pages 357–358, 2015.
  19. 19.Dressipi dataset (http://www.recsyschallenge.com/2022), 2022.
  20. 20.Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. Graph neural networks in recommender systems: a survey. ACM Computing Surveys, 55(5):1–37, 2022.
  21. 21.Tianchi. Ijcai-15 repeat buyers prediction dataset, 2018.
  22. 22.Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. Harnessing the power of llms in practice: A survey on chatgpt and beyond. arXiv preprint arXiv:2304.13712, 2023.
  23. 23.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
  24. 24.Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al. A comprehensive survey on pretrained foundation models: A history from bert to chatgpt. arXiv preprint arXiv:2302.09419, 2023.
  25. 25.Diginetica dataset (https://competitions.codalab.org/competitions/11161), 2016.
  26. 26.Hongzhi Yin, Bin Cui, Jing Li, Junjie Yao, and Chen Chen. Challenging the long tail recommendation. arXiv preprint arXiv:1205.6700, 2012.
  27. 27.Siyi Liu and Yujia Zheng. Long-tail session-based recommendation. In Proceedings of the 14th ACM Conference on Recommender Systems, pages 509–514, 2020.
  28. 28.Chen Gao, Xiangning Chen, Fuli Feng, Kai Zhao, Xiangnan He, Yong Li, and Depeng Jin. Cross-domain recommendation without sharing user-relevant data. In The world wide web conference, pages 491–502, 2019.
  29. 29.Antoine Nzeyimana and Andre Niyongabo Rubungo. Kinyabert: a morphology-aware kinyarwanda language model. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5347–5363, 2022.
  30. 30.Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, 2021.
  31. 31.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, et al. Choosing transfer languages for cross-lingual learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3125–3135, 2019.
  32. 32.Ashton Anderson, Ravi Kumar, Andrew Tomkins, and Sergei Vassilvitskii. The dynamics of repeat consumption. In Proceedings of the 23rd international conference on World wide web, pages 419–430, 2014.
  33. 33.Jun Chen, Chaokun Wang, and Jianmin Wang. Will you" reconsume" the near past? fast prediction on short-term reconsumption behaviors. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 29, 2015.
  34. 34.Austin R Benson, Ravi Kumar, and Andrew Tomkins. Modeling user consumption sequences. In Proceedings of the 25th International Conference on World Wide Web, pages 519–529, 2016.
  35. 35.Pengjie Ren, Zhumin Chen, Jing Li, Zhaochun Ren, Jun Ma, and Maarten De Rijke. Repeatnet: A repeat aware neural recommendation machine for session-based recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 4806–4813, 2019.
  36. 36.Zhiqiang Pan, Fei Cai, Wanyu Chen, and Honghui Chen. Graph co-attentive session-based recommendation. ACM Transactions on Information Systems (TOIS), 40(4):1–31, 2021.
  37. 37.Yujia Zheng, Siyi Liu, and Zailei Zhou. Balancing multi-level interactions for session-based recommendation. arXiv preprint arXiv:1910.13527, 2019.
  38. 38.Yujia Zheng, Siyi Liu, Zekun Li, and Shu Wu. Cold-start sequential recommendation via meta learner. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 4706–4713, 2021.
  39. 39.Meirui Wang, Pengjie Ren, Lei Mei, Zhumin Chen, Jun Ma, and Maarten De Rijke. A collaborative session-based recommendation approach with parallel memory modules. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, pages 345–354, 2019.
  40. 40.Balázs Hidasi and Alexandros Karatzoglou. Recurrent neural networks with top-k gains for session-based recommendations. In Proceedings of the 27th ACM international conference on information and knowledge management, pages 843–852, 2018.
  41. 41.Yupeng Hou, Binbin Hu, Zhiqiang Zhang, and Wayne Xin Zhao. Core: simple and effective session-based recommendation within consistent representation space. In Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval, pages 1796–1801, 2022.
  42. 42.Siqi Lai, Erli Meng, Fan Zhang, Chenliang Li, Bin Wang, and Aixin Sun. An attribute-driven mirror graph network for session-based recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1674–1683, 2022.
  43. 43.Nils Reimers and Iryna Gurevych. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2020.
  44. 44.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  45. 45.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  46. 46.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  47. 47.Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. Towards universal sequence representation learning for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 585–593, 2022.
  48. 48.Chenglin Li, Mingjun Zhao, Huanming Zhang, Chenyun Yu, Lei Cheng, Guoqiang Shu, Beibei Kong, and Di Niu. Recguru: Adversarial learning of generalized user representations for cross-domain recommendation. In Proceedings of the fifteenth ACM international conference on web search and data mining, pages 571–581, 2022.
  49. 49.Nouhaila Idrissi and Ahmed Zellou. A systematic literature review of sparsity issues in recommender systems. Social Network Analysis and Mining, 10:1–23, 2020.
  50. 50.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  51. 51.OpenAI. Gpt-4 technical report. ArXiv, abs/2303.08774, 2023.
  52. 52.Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. Recommender systems in the era of large language models (llms). arXiv preprint arXiv:2307.02046, 2023.
  53. 53.Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. Large language models are zero-shot rankers for recommender systems. arXiv preprint arXiv:2305.08845, 2023.
  54. 54.Xiaolei Wang, Xinyu Tang, Wayne Xin Zhao, Jingyuan Wang, and Ji-Rong Wen. Rethinking the evaluation for conversational recommendation in the era of large language models. arXiv preprint arXiv:2305.13112, 2023.
  55. 55.Junling Liu, Chao Liu, Renjie Lv, Kang Zhou, and Yan Zhang. Is chatgpt a good recommender? a preliminary study. arXiv preprint arXiv:2304.10149, 2023.
  56. 56.Wenjie Wang, Xinyu Lin, Fuli Feng, Xiangnan He, and Tat-Seng Chua. Generative recommendation: Towards next-generation recommender paradigm. arXiv preprint arXiv:2304.03516, 2023.
  57. 57.Jinming Li, Wentao Zhang, Tian Wang, Guanglei Xiong, Alan Lu, and Gerard Medioni. Gpt4rec: A generative framework for personalized recommendation and user interests interpretation. arXiv preprint arXiv:2304.03879, 2023.
  58. 58.Junjie Zhang, Ruobing Xie, Yupeng Hou, Wayne Xin Zhao, Leyu Lin, and Ji-Rong Wen. Recommendation as instruction following: A large language model empowered recommendation approach. arXiv preprint arXiv:2305.07001, 2023.
  59. 59.Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. arXiv preprint arXiv:2305.00447, 2023.
  60. 60.Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. A survey on large language models for recommendation. arXiv preprint arXiv:2305.19860, 2023.
  61. 61.Touseef Iqbal and Shaima Qureshi. The survey: Text generation models in deep learning. Journal of King Saud University-Computer and Information Sciences, 34(6):2515–2528, 2022.
  62. 62.Luciano Floridi and Massimo Chiriatti. Gpt-3: Its nature, scope, limits, and consequences. Minds and Machines, 30:681–694, 2020.
  63. 63.Kaisheng Zeng, Chengjiang Li, Lei Hou, Juanzi Li, and Ling Feng. A comprehensive survey of entity alignment for knowledge graphs. AI Open, 2:1–13, 2021.
  64. 64.Xiaojuan Zhao, Yan Jia, Aiping Li, Rong Jiang, and Yichen Song. Multi-source knowledge fusion: a survey. World Wide Web, 23:2567–2592, 2020.
  65. 65.Hsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, and Xu Sun. Aligning cross-lingual entities with multi-aspect information. arXiv preprint arXiv:1910.06575, 2019.
  66. 66.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  67. 67.Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S Yu. A comprehensive survey on graph neural networks. arXiv preprint arXiv:1901.00596, 2019.
  68. 68.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  69. 69.Hongzhi Wen, Jiayuan Ding, Wei Jin, Yiqi Wang, Yuying Xie, and Jiliang Tang. Graph neural networks for multimodal single-cell data integration. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4153–4163, 2022.
  70. 70.Jiayuan Ding, Hongzhi Wen, Wenzhuo Tang, Renming Liu, Zhaoheng Li, Julian Venegas, Runze Su, Dylan Molho, Wei Jin, Wangyang Zuo, et al. Dance: A deep learning library and benchmark for single-cell analysis. bioRxiv, pages 2022–10, 2022.
  71. 71.David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams. Convolutional networks on graphs for learning molecular fingerprints. Advances in neural information processing systems, 28, 2015.
  72. 72.Zhikai Chen, Haitao Mao, Hang Li, Wei Jin, Hongzhi Wen, Xiaochi Wei, Shuaiqiang Wang, Dawei Yin, Wenqi Fan, Hui Liu, et al. Exploring the potential of large language models (llms) in learning on graphs. arXiv preprint arXiv:2307.03393, 2023.
  73. 73.Xiaoxin He, Xavier Bresson, Thomas Laurent, and Bryan Hooi. Explanations as features: Llm-based features for text-attributed graphs. arXiv preprint arXiv:2305.19523, 2023.
  74. 74.Jiayan Guo, Lun Du, and Hengyu Liu. Gpt4graph: Can large language models understand graph structured data? an empirical evaluation and benchmarking. arXiv preprint arXiv:2305.15066, 2023.
  75. 75.Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, and Yulia Tsvetkov. Can language models solve graph problems in natural language? arXiv preprint arXiv:2305.10037, 2023.
  76. 76.Jiawei Zhang. Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt. arXiv preprint arXiv:2304.11116, 2023.
  77. 77.Hoyeop Lee, Jinbae Im, Seongwon Jang, Hyunsouk Cho, and Sehee Chung. Melu: Meta-learned user preference estimator for cold-start recommendation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1073–1082, 2019.
  78. 78.Stef Van Buuren. Flexible imputation of missing data. CRC press, 2018.
  79. 79.Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, et al. Recbole: Towards a unified, comprehensive and efficient framework for recommendation algorithms. In proceedings of the 30th acm international conference on information & knowledge management, pages 4653–4664, 2021.

Citation

MLA
Jin, W., et al. “Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation”. arXiv, 2023, http://arxiv.org/abs/2307.09688v2.
APA
Jin, W., Mao, H., Li, Z., Jiang, H., Luo, C., Wen, H., Han, H., Lu, H., Wang, Z., Li, R., Li, Z., Cheng, M. X., Goutam, R., Zhang, H., Subbian, K., Wang, S., Sun, Y., Tang, J., Yin, B., & Tang, X. (2023). Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation. arXiv. http://arxiv.org/abs/2307.09688v2
Chicago
Jin, W., H. Mao, Z. Li, et al. 2023. “Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation”. arXiv. http://arxiv.org/abs/2307.09688v2.
Harvard
Jin, W. et al. (2023) “Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2307.09688v2.
Vancouver
1. Jin W, Mao H, Li Z, et al (2023) Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation. arXiv

BibTeX

@article{jin2023amazon,
  title = {Amazon-M2: A Multilingual Multi-locale Shopping Session Dataset for Recommendation and Text Generation},
  author = {Jin, Wei and Mao, Haitao and Li, Zheng and Jiang, Haoming and Luo, Chen and Wen, Hongzhi and Han, Haoyu and Lu, Hanqing and Wang, Zhengyang and Li, Ruirui and Li, Zhen and Cheng, Monica Xiao and Goutam, Rahul and Zhang, Haiyang and Subbian, Karthik and Wang, Suhang and Sun, Yizhou and Tang, Jiliang and Yin, Bing and Tang, Xianfeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2307.09688v2},
  eprint = {2307.09688}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors