Large-scale recommendation refers to the automated process and system architecture designed to identify, rank, and deliver personalized item suggestions to users from massive candidate pools containing millions or billions of items. Because evaluating complex scoring models over an entire corpus in real time is computationally prohibitive, these systems typically operate through a multi-stage pipeline comprising candidate retrieval, ranking, and re-ranking. The initial retrieval stage rapidly filters the vast item corpus down to a manageable subset of promising candidates using efficient indexing and vector similarity search methods, after which downstream stages apply more computationally intensive models to accurately score, personalize, and order the final recommendations under strict latency and throughput constraints.