Argo: Efficient Importance Labeling for Enterprise Email Systems

Siddhant RayGanesh AnanthanarayananKevin ChianYan GuoCristina St HillJack W. StokesVictor WangJunchen Jiang

article2026arXiv0 citations

Presents Argo, an enterprise email labeling framework that matches frontier language model quality while reducing inference costs by up to 167x through cost-performance profiling and dynamic resource provisioning.

Listen

Modern organizations face overwhelming daily email volumes, making automated importance labeling critical for workplace productivity and workflow prioritization. While advanced large language models offer strong semantic reasoning and context understanding, deploying them universally across enterprise-scale workloads incurs prohibitive financial and computational costs, estimated in the billions of dollars monthly for major providers. Traditional heuristic and feature-engineered solutions are computationally inexpensive but fail to generalize accurately.

The main objective of the article is to design, implement, and evaluate Argo, an enterprise-level email labeling framework that delivers near-frontier model labeling accuracy while dramatically reducing operational inference and profiling expenses.

The evaluated approach uses a hybrid architecture that dynamically splits labeling workloads across two complementary methods based on label characteristics. For simple binary labels (such as whether an action is needed), Argo routes tasks to a lightweight, central processing unit-based text embedding classifier. For complex, non-binary labels (such as discrete priority rankings), it routes requests through an ordered cascade of small language models of increasing parameter size, escalating only when smaller models fall below confidence thresholds. An offline profiling engine navigates configuration choices using a calibration set, while an on-demand greedy resource allocation algorithm manages real-time capacity bottlenecks and cloud service penalties during peak email traffic.

Evaluation across three open-source corporate and governmental email corpora demonstrates that Argo achieves a 148-fold to 167-fold reduction in inference costs compared to baseline frontier models with negligible loss in label quality. The framework's profiler discovers optimal operating configurations at 20-fold to 640,000-fold lower search costs than standard exhaustive sweeps. Additionally, Argo's load-aware provisioning algorithm reduces peak-load scaling cost spikes by 2.2-fold to 3.8-fold compared to baseline provisioning strategies, while the underlying architecture successfully extends to email summarization tasks with 36-fold to 48-fold cost savings.

These findings demonstrate that enterprises do not need to rely uniformly on massive foundation models to achieve high-grade natural language understanding in communications. By decoupling complex from binary labeling tasks and pairing small specialized models with CPU-based classifiers, organizations can make intelligent email triage economically viable at scale. The results also show that intelligent offline profiling avoids the massive compute expenditures typically associated with searching large parameter configuration spaces.

Organizations planning large-scale natural language processing deployments should consider adopting hybrid cascades over uniform large-model architectures. Next steps include establishing small offline calibration pipelines to tune confidence thresholds and integrating enterprise policy rules—such as designating latency delays or fallback tiers during peak load periods for non-critical user tiers. Further pilot testing on live, user-specific enterprise streams and fine-tuning domain-specific small models can deliver additional accuracy gains.

The study's primary limitations stem from evaluation on public historical datasets rather than live enterprise production environments with varied real-time latency and user interaction patterns. Additionally, periodic shifts in organizational communication require ongoing re-profiling to preserve calibration accuracy. Nevertheless, the consistent performance across diverse datasets provides high confidence that Argo's structural cost reductions and quality guarantees generalize well to standard enterprise messaging systems.

  • Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). It investigates using large language models as automated data annotators to generate training data for smaller, cheaper models, establishing the cost-quality paradigm that Argo builds upon for email labeling.
  • Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). It provides foundational evidence that LLMs can accurately and cost-effectively substitute for human annotators in text-labeling tasks, motivating Argo's search for efficient LLM-based labeling schemes.
  • Paper: Snorkel: Rapid Training Data Creation with Weak Supervision, Alexander J. Ratner et al. (2017). It presents the Snorkel framework for programmatic and weak supervision, offering key conceptual groundwork for replacing expensive manual labeling with scalable, cost-effective labeling alternatives.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). It introduces few-shot in-context learning with large-scale language models, demonstrating the deep contextual labeling capability whose inference costs Argo seeks to mitigate.
Cover for Argo: Efficient Importance Labeling for Enterprise Email Systems

Abstract

Email importance labeling has long been a critical yet challenging problem for businesses and individuals. Traditional approaches; such as keyword matching, user-defined rules, and sender-based heuristics; demand extensive manual feature engineering and fail to scale effectively or generalize. Recent advances in large language models (LLMs) demonstrate strong potential and a natural fit for this task, offering deep contextual understanding and superior labeling quality. However, using LLM models like GPT-4.1 at enterprise email volumes incurs prohibitive computational costs and hinders real-world deployment. We explore the trade-off space of using alternative labeling schemes as opposed to GPT4.1 scale LLMs, with the goal of achieving near GPT level labeling quality with significantly lower cost. We develop Argo, an enterprise email labeling framework, where we construct a profiler to efficiently search the cost quality trade-off space of labeling and identify cost-efficient alternatives to labeling emails. Additionally, we design an on-demand provisioning scheme to intelligently scale Argo with real time load, to minimize cost increases during peak load inference. Over 3 open-source email datasets, Argo achieves 148-167X inference cost reduction with negligible quality degradation and 20-640000X lower profiling costs, making large-scale, context-aware email labeling practical for enterprises.

Table of Contents

  • 1 Introduction
  • 2 Background: Email Importance Labeling
  • 3 Towards obtaining better cost-quality tradeoffs for email labeling
  • 3.1 How are different email labels distributed?
  • 3.2 What if we only use an SLM cascade?
  • 3.3 What if we only use an embedding classifier?
  • 3.4 Obtaining optimal quality-cost tradeoffs need a smart system for labeling
  • 4 Argo: A System for Cost-Efficient Enterprise Email Labeling
  • 4.1 How and what do we profile?
  • 4.2 Selecting knob values with Argo’s profiler
  • 4.3 How do we manage the cost of profiling?
  • 4.4 On-demand provisioning to minimize email labeling costs under increasing load
  • 5 Enhancements to Argo
  • 6 Implementation
  • 7 Evaluation
  • 7.1 Setup
  • 7.2 End-to-End Results
  • 7.3 Breakdown and Sensitivity Results
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Argo Hybrid Labeling Architecture for Enterprise Email

    model/method

    Argo is an enterprise-scale email importance labeling framework that dynamically routes label assignments between an embedding-based neural classifier and a cascade of Small Language Models (SLMs) based on label characteristics determined during an offline profiling stage.

    For binary classification tasks (such as determining whether an email requires a reply, needs scheduling, represents an urgent matter, or requires explicit administrative action), Argo executes a 3-layer Multi-Layer Perceptron (MLP) trained on frozen text embeddings (such as text-embedding-3-large). The embedding vectors are computed once per email and reused across all binary labels, and both training and inference are offloaded to CPU instances, completing in under 1 ms1\text{ ms} per email.

    For non-binary or multi-class labels (such as overall email priority rated on an ordinal scale from 1 to 5), embedding classifiers degrade in F1F_1-score by more than 15%15\% due to class imbalance and text heterogeneity. For these labels, Argo routes emails sequentially through an SLM cascade ordered from smallest (lowest inference cost) to largest model. An SLM at stage ii produces a label and an associated confidence score derived by mapping output token log-probabilities to a linear scale [0,100][0, 100]. If the confidence score meets or exceeds a calibrated threshold TiT_i, the label is accepted; otherwise, the request cascades to model i+1i+1. If the end of the cascade is reached without exceeding the final threshold, the prediction of the largest, most expressive SLM is retained.

  2. Knowl 2 — Greedy On-Demand Instance Provisioning Algorithm for SLM Cascades

    algorithm

    Argo manages model-as-a-service instance allocations for an SLM cascade during traffic spikes using a greedy marginal-cost minimization algorithm. When email arrival exceeds the steady-state capacity of an active model instance, Argo evaluates whether it is cheaper to provision a new instance of the current model (incurring a multiplicative penalty factor p≥1p \ge 1 applied to the base instance cost cic_i) or to spill over the excess load to the next, more expensive model in the cascade with incremental per-request running cost zi+1−ziz_{i+1} - z_i.

    Input: Total demand DD, number of requests rr, instance capacity CC,
           instance base costs c1,…,cmc_1, \dots, c_m, penalty p≥1p \ge 1,
           running costs per request z1<z2<⋯<zmz_1 < z_2 < \dots < z_m,
           initial steady-state requests kik_i, initial instances nin_i,
           number of models mm in the cascade
    Output: Updated instances nin_i, request counts kik_i, and TotalCost
    Dreq←D/rD_{\text{req}} \leftarrow D / r
    for each new request do
        for i←1i \leftarrow 1 to mm do
            if kiDreq<niCk_i D_{\text{req}} < n_i C then
                ki←ki+1k_i \leftarrow k_i + 1
                break
            else
                ΔInstCost←cipni\Delta InstCost \leftarrow c_i p^{n_i}
                ΔRunCost←zi+1−zi\Delta RunCost \leftarrow z_{i+1} - z_i
                if ΔInstCost<ΔRunCost\Delta InstCost < \Delta RunCost then
                    ni←ni+1n_i \leftarrow n_i + 1
                    ki←ki+1k_i \leftarrow k_i + 1
                    break
                else
                    continue
    InstanceCost←∑i=1mcipni−1p−1InstanceCost \leftarrow \sum_{i=1}^m c_i \frac{p^{n_i} - 1}{p - 1}
    RunCost←∑i=1mkiziRunCost \leftarrow \sum_{i=1}^m k_i z_i
    TotalCost←InstanceCost+RunCostTotalCost \leftarrow InstanceCost + RunCost
    return nin_i, kik_i, TotalCostTotalCost

    Because per-request costs strictly increase up the cascade (z1<z2<⋯<zmz_1 < z_2 < \dots < z_m) and cost decisions are additive, evaluating only the immediate next tier in the cascade guarantees minimal marginal cost increase at every step.

  3. Knowl 3 — Search Space Pruning Strategies for SLM Cascade Profiling

    model/method

    Configuring an SLM cascade for email labeling across multiple knobs (model selection, cascade order, confidence thresholds, and calibration set size) spans an exhaustive search space exceeding 10810^8 configurations. Argo reduces this search space using four structural properties:

    1. Incremental Calibration Sizing: Argo starts with a minimal calibration subset of emails and profiles candidate configurations against a separate holdout validation set, expanding the calibration set iteratively until candidate configurations reach the empirical Pareto frontier of quality versus cost on the validation set, bounding unnecessary profiling on redundant emails.
    2. SLM Quality-Cost Independence: The relative quality-cost Pareto optimality of an individual SLM is invariant to the other models present in the cascade. Argo selects constituent models independently based on individual Pareto efficiency prior to cascade assembly.
    3. Monotonic Size Ordering: Because the runtime system cannot predict a priori which model will satisfy confidence requirements on an unseen email, cascading in strictly ascending order of model parameter count and token cost (z1<z2<⋯<zmz_1 < z_2 < \dots < z_m) is guaranteed to minimize cost by exhausting cheaper SLMs first.
    4. Confidence Floor Pruning: Profiling reveals that individual SLMs exhibit catastrophic accuracy drops below specific confidence floors (e.g., below 70%70\% linear confidence for Phi-4-mini and below 80%80\% for Llama-3.1-8B). Argo prunes all candidate threshold vectors containing values below these known individual failure boundaries from the grid search.
  4. Knowl 4 — Standardized Wasserstein-1 Distance for SLM Cascade Re-Profiling

    equation

    To detect distribution shifts in incoming enterprise emails that require recalibrating SLM cascade thresholds and classifier parameters, Argo measures the divergence between the baseline token log-probability confidence distribution XX and the arriving stream distribution YY using the Standardized Wasserstein-1 Distance (SWD):

    SWD=W1(FX,FY)σX\text{SWD} = \frac{W_1(F_X, F_Y)}{\sigma_X}

    where FXF_X and FYF_Y are the cumulative distribution functions of the baseline and incoming confidence scores, σX=std(X)\sigma_X = \text{std}(X) is the standard deviation of the baseline distribution, and W1(FX,FY)W_1(F_X, F_Y) is the first Wasserstein distance defined on the inverse distribution functions (quantile functions) FX−1F_X^{-1} and FY−1F_Y^{-1}:

    W1(FX,FY)=∫01∣FX−1(u)−FY−1(u)∣duW_1(F_X, F_Y) = \int_{0}^{1} \left| F_X^{-1}(u) - F_Y^{-1}(u) \right| du

    Argo triggers re-profiling (generating new golden calibration labels using a frontier LLM baseline) when SWD>1.0\text{SWD} > 1.0 or when 24 hours have elapsed since the last calibration run, whichever occurs first.

  5. Knowl 5 — End-to-End Inference Cost Reduction and Labeling Accuracy

    empirical result

    Across three enterprise email evaluation datasets (1,000 emails from the Enron dataset, 500 emails from the Fauci dataset, and 700 emails from the Hillary Clinton dataset), Argo's default Balanced operating policy achieves a 148×148\times to 167×167\times reduction in inference API cost relative to a full GPT-4.1 baseline while preserving comparable average F1F_1-score across all labels (overall priority, needs reply, is urgent, needs action, and needs scheduling).

    Compared to deploying any single standalone SLM (such as Phi-4-mini, Llama-3.1-8B, Gemma-3-27B, Qwen2.5-32B, or Llama-3.3-70B), Argo achieves 5%5\% to 15%15\% higher average F1F_1-score while lowering inference cost by 1.67×1.67\times to 17×17\times.

    Argo also provides parameterized operating points on the Pareto frontier:

    • Cost-Focus Policy: Achieves 186×186\times to 200×200\times API cost reduction relative to GPT-4.1 with a slight quality degradation.
    • Quality-Focus Policy: Achieves 120×120\times to 136×136\times API cost reduction relative to GPT-4.1, operating within 1–2%1\text{--}2\% of the oracle upper-bound cascade.
  6. Knowl 6 — Profiling Cost Optimization Across Search Space Baselines

    empirical result

    Argo's profiler reduces the computational cost of finding Pareto-optimal labeling configurations by 20×20\times to 641,160×641,160\times compared to alternative search strategies without sacrificing labeling accuracy:

    • Exhaustive Profiler: Sweeps all configuration combinations across the entire calibration set (1.0×1.0\times baseline cost).
    • Knob Independence Decomposition: Reduces profiling cost by 13×13\times by evaluating individual SLM tradeoffs independently.
    • Size-Ordered Pruning: Reduces profiling cost by 1,560×1,560\times by constraining cascades to ascending parameter order.
    • Full Argo Profiler: Combines incremental calibration sizing, model independence, monotonic ordering, and confidence floor truncation, yielding a cumulative 641,160×641,160\times reduction in profiling cost while achieving identical F1F_1 quality to the exhaustive search.
    • 1% and 10% Random Calibration Sampling Baselines: While cheaper than exhaustive search (31,000×31,000\times and 4,850×4,850\times cost reduction respectively), random sampling results in severe quality degradations of 33%33\% and 21%21\% lower F1F_1-score on downstream evaluation.
    • Reduced Cascade Baseline (profiling only the top 3 SLMs): Reduces profiling cost by 24–29×24\text{--}29\times but incurs a 12%12\% drop in F1F_1-score.
    • Reduced Thresholds Baseline (profiling thresholds >50%>50\%): Retains quality but only achieves a 12–15×12\text{--}15\times cost reduction.
  7. Knowl 7 — Provisioning Cost Reduction Under Simulated Enterprise Traffic Spikes

    empirical result

    Under simulated peak email arrival workloads on the Enron dataset where queue capacities for individual SLM instances saturate, Argo's greedy provisioning algorithm (evaluating marginal instance penalty versus cascade spillover) achieves 2.2×2.2\times to 3.8×3.8\times lower total cost increases compared to standard baseline strategies.

    When evaluating bottleneck capacities across single-SLM saturation (1 SLM at capacity) and multi-SLM saturation (3 SLMs at capacity) under penalty multipliers p=2×p = 2\times and p=5×p = 5\times:

    • The Always Increase Instances baseline (which immediately provisions new instances of overloaded SLMs despite penalty pp) suffers rapid cost inflation as pnip^{n_i} grows.
    • The Always Move Up Cascade baseline (which routes all overflow requests directly to higher-tier models without considering instance addition) incurs severe per-request running cost premiums.
    • Argo's hybrid greedy allocation tracks the lower bound across all load and penalty regimes.
  8. Knowl 8 — Model Inventory and Relative API Cost Multipliers in Argo

    data/table

    The following table lists the candidate SLMs and text embedding models evaluated in Argo's labeling pipeline. API cost reduction factors represent blended input/output token pricing (based on a 3:1 input-to-output token ratio on Azure AI Foundry) relative to the frontier GPT-4.1 baseline running identical email inputs:

    Model Name Type API Cost Reduction
    microsoft/phi-4-mini-instruct SLM 105×105\times cheaper
    meta-llama/Llama-3.1-8B-Instruct SLM 100×100\times cheaper
    google/gemma-3-27b-it SLM 90×90\times cheaper
    Qwen/Qwen2.5-32B-Instruct SLM 20×20\times cheaper
    meta-llama/Llama-3.3-70B-Instruct SLM 10×10\times cheaper
    text-embedding-3-large Embedding 100×100\times cheaper

    All SLMs support context windows of up to 128,000 tokens.

  9. Knowl 9 — Ablation of Isolated Cascading vs Isolated Embedding Classification

    empirical result

    Ablation experiments isolating Argo's two labeling components demonstrate that combining both methods is necessary to balance cost and accuracy:

    1. Cascade-Only Configuration: Forcing the SLM cascade to predict both binary and non-binary labels preserves full labeling accuracy (F1F_1-score) but reduces cost savings by 2×2\times (e.g., dropping from ∼160×\sim 160\times cost reduction down to ∼80×\sim 80\times). This occurs because generating multiple binary labels jointly within a single SLM prompt overwhelms model reasoning capabilities and degrades F1F_1-score by 15–20%15\text{--}20\%, requiring independent cascade runs for each label.
    2. Embedding-Only Configuration: Forcing a trained multi-head embedding classifier (MLP on text-embedding-3-large) to predict all labels maximizes cost reduction but causes a 13%13\% to 15%15\% degradation in F1F_1-score across non-binary labels (e.g., discrete Priority levels 1 through 5), which exceeds allowable enterprise quality tolerances.
  10. Knowl 10 — Enterprise Operator Policies, Constraints, and Summarization Extension

    model/method

    Argo incorporates modular extensions and policy overrides for enterprise deployment constraints:

    • Operator-Enforced Model Constraints: When operators enforce policy or privacy constraints (such as capping cascade depth to a maximum of 4 SLMs and excluding the Meta Llama family), Argo profiles substitute open-weight architectures (such as mistralai/Mistral-7B-Instruct-v0.2 and DeepSeek Distilled). Under these constrained settings, Argo still discovers a balanced configuration achieving 122×122\times to 148×148\times cost reduction relative to GPT-4.1.
    • Tier-Based Peak Load Policies: During severe traffic bursts, Argo can apply organizational hierarchy signals from enterprise contact cards (e.g., C-suite vs. non-priority employee tiers):
      1. Quality-Based Downgrade: Non-critical user tiers are downgraded to smaller, cheaper SLMs during capacity saturation, bypassing profiler thresholds to avoid instance expansion penalties.
      2. Delay-Based Request Buffering: Non-time-sensitive labeling requests are queued and deferred across the burst duration, staggering arrivals to remain below provisioned instance thresholds CC.
    • Email Summarization Extension: Argo adapts its cascade framework from discrete classification to abstractive summarization by replacing token F1F_1 metrics with LLM-as-a-judge scoring and evaluating SLM selection via token-frequency-weighted confidence scores. This configuration achieves a 36×36\times to 48×48\times inference cost reduction compared to GPT-4.1.

Coverage note — None was omitted; all contributed models, algorithms, search space reduction strategies, mathematical drift equations, empirical benchmarks, ablations, and system extensions have been captured.

References

  1. 1.Alibaba/Qwen Team. Qwen2.5 32b instruct. https://huggingface.co/Qwen/Qwen2.5-32B-Instruct, 2025. SLM, 20× cheaper.
  2. 2.Sakhar Alkhereyf and Owen Rambow. Work hard, play hard: Email classification on the avocado and enron corpora. In Proceedings of TextGraphs-11: the Workshop on Graph-based Methods for Natural Language Processing, pages 57–65. Association for Computational Linguistics, 2017.
  3. 3.Apple. Show emails from vip senders in mail on mac. https://support.apple.com/guide/mail/show-emails-from-vip-senders-mail40589/mac. Accessed: 2025-12-08.
  4. 4.Apple Support. Summarize notifications and reduce interruptions with apple intelligence on iphone. https://support.apple.com/is-is/guide/iphone/iph1fbe7d2b9/ios, 2025. Accessed December 8, 2025.
  5. 5.Yogesh Balaji, Rama Chellappa, and Soheil Feizi. Normalized wasserstein distance for mixture distributions with applications in adversarial learning and domain adaptation, 2019.
  6. 6.Ron Bekkerman. Automatic categorization of email into folders: Benchmark experiments on enron and sri corpora.
  7. 7.Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. Supervised learning of universal sentence representations from natural language inference data. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 670–680, Copenhagen, Denmark, September 2017. Association for Computational Linguistics.
  8. 8.DeepSeek AI. Deepseek distilled models. https://huggingface.co/deepseek-ai, 2024. Distilled variants including DeepSeek-LLM and R1-Distill.
  9. 9.Sheng Deng, Wei Wang, and Jian Sun. Hierarchical attention networks for email classification. In AAAI Conference on Artificial Intelligence, 2018.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL, 2019.
  11. 11.Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. Hybrid llm: Cost-efficient and quality-aware query routing, 2024.
  12. 12.Nicolas Ducheneaut and Victoria Bellotti. E-mail as habitat: an exploration of embedded personal information management. Interactions, 8(5):30–38, September 2001.
  13. 13.OpenAI et al. Gpt-4 technical report, 2024.
  14. 14.William Falcon and the PyTorch Lightning Team. Pytorch lightning. https://github.com/Lightning-AI/pytorch-lightning, 2025. Version 2.5.4 (accessed 2025-12-09).
  15. 15.Federal Energy Regulatory Commission. Enron email dataset. https://www.cs.cmu.edu/~enron/, 2004. Accessed 2025-12-09.
  16. 16.Google DeepMind. Gemma 3 27b it. https://huggingface.co/google/gemma-3-27b-it, 2025. SLM, 90× cheaper.
  17. 17.David Graus, David van Dijk, Manos Tsagkias, Wouter Weerkamp, and Maarten de Rijke. Recipient recommendation in enterprises using communication graphs and email content. In Proceedings of the 37th International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’14, page 1079–1082, New York, NY, USA, 2014. Association for Computing Machinery.
  18. 18.Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, and Sanjiv Kumar. Language model cascades: Token-level uncertainty and beyond. In B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, editors, International Conference on Representation Learning, volume 2024, pages 4147–4180, 2024.
  19. 19.Haibo He and Edwardo A. Garcia. Learning from imbalanced data. IEEE Trans. on Knowl. and Data Eng., 21(9):1263–1284, September 2009.
  20. 20.Yifei He, Haoxiang Wang, Bo Li, and Han Zhao. Gradual domain adaptation: Theory and algorithms. Journal of Machine Learning Research, 25(361):1–40, 2024.
  21. 21.Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. Focus: Querying large video datasets with low latency and low cost. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 269–286, Carlsbad, CA, October 2018. USENIX Association.
  22. 22.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In Proceedings of the 10th International Conference on Learning Representations (ICLR), 2022.
  23. 23.Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. Routerbench: A benchmark for multi-llm routing system, 2024.
  24. 24.Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodik, Siddhartha Sen, and Ion Stoica. Chameleon: scalable adaptation of video analytics. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication, SIGCOMM ’18, page 253–266, New York, NY, USA, 2018. Association for Computing Machinery.
  25. 25.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations (ICLR), 2015.
  26. 26.Svetlana Kiritchenko and Stan Matwin. Email classification with co-training. In Proceedings of the 2011 Conference of the Center for Advanced Studies on Collaborative Research, CASCON ’11, page 301–312, USA, 2011. IBM Corp.
  27. 27.Bryan Klimt and Yiming Yang. The enron corpus: A new dataset for email classification research. In European Conference on Machine Learning (ECML), 2004.
  28. 28.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP), 2023.
  29. 29.Andrew Lampert, Robert Dale, and Cecile Paris. Detecting emails containing requests for action. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 984–992, 2010.
  30. 30.Xunzhuo Liu, Huamin Chen, Samzong Lu, Yossi Ovadia, Guohong Wen, Hao Wu, Zhengda Tan, Jintao Zhang, Senan Zedan, Yehudit Kerido, Liav Weiss, Haichen Zhang, Bishen Yu, Asaad Balum, Noa Limoy, Abdallah Samara, Baofa Fan, Brent Salisbury, Ryan Cook, Zhijie Wang, Qiping Pan, Rehan Khan, Avishek Goswami, Houston H. Zhang, Shuyi Wang, Ziang Tang, Fang Han, Zohaib Hassan, Jianqiao Zheng, and Avinash Changrani. vllm semantic router: Signal driven decision routing for mixture-of-modality models, 2026.
  31. 31.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  32. 32.Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Wei Liu, Jian Luan, Xiwen Zhang, Nicholas D. Lane, and Mengwei Xu. Demystifying small language models for edge deployment. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14747–14764, Vienna, Austria, July 2025. Association for Computational Linguistics.
  33. 33.Wendy E. Mackay. More than just a communication system: Diversity in the use of electronic mail. In Proceedings of the ACM Conference on Computer-Supported Cooperative Work (CSCW), pages 344–353. ACM, 1988.
  34. 34.Andrew McCallum, Xuerui Wang, and Andrés Corrada-Emmanuel. Topic and role discovery in social networks with experiments on enron and academic email. J. Artif. Int. Res., 30(1):249–272, October 2007.
  35. 35.Meta AI. Llama 3.1 8b instruct. https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct, 2025. SLM, 100× cheaper.
  36. 36.Meta AI. Llama 3.3 70b instruct. https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct, 2025. SLM, 10× cheaper.
  37. 37.Microsoft. What is focused inbox? https://support.microsoft.com/en-us/office/what-is-focused-inbox-16b24373-dfa9-4139-ab19-08aa753a6055. Accessed: 2025-12-08.
  38. 38.Microsoft. Phi-4-mini-instruct. https://huggingface.co/microsoft/phi-4-mini-instruct, 2025. SLM, 105× cheaper.
  39. 39.Microsoft Learn. What is azure ai foundry? https://learn.microsoft.com/en-us/azure/ai-foundry/what-is-azure-ai-foundry, 2025. Accessed December 9, 2025.
  40. 40.Mistral AI. Mistral 7b instruct. https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2, 2024. Open-weight 7B instruction-tuned model.
  41. 41.Farnaz Moradi, Tomas Olovsson, and Philippas Tsigas. Towards modeling legitimate and unsolicited email traffic using social network properties. In Proceedings of the Fifth Workshop on Social Network Systems, SNS ’12, New York, NY, USA, 2012. Association for Computing Machinery.
  42. 42.Deepak Narayanan et al. Efficient large-scale language model training on gpu clusters. In USENIX OSDI, 2021.
  43. 43.Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica. Routellm: Learning to route llms with preference data. arXiv preprint arXiv:2406.11635, 2024.
  44. 44.ONNX Community. Onnx: Open neural network exchange. https://onnx.ai, 2025. Accessed 2025-12-09.
  45. 45.OpenAI. Openai api. https://openai.com/blog/openai-api/, 2020. Accessed December 9, 2025.
  46. 46.OpenAI. Openai api pricing — text token costs. https://platform.openai.com/docs/pricing, 2025. Accessed December 9, 2025.
  47. 47.OpenAI. text-embedding-3-large. https://platform.openai.com/docs/models, 2025. Embedding model, 100× cheaper.
  48. 48.Hani Osman. Fauci emails dataset, 2021.
  49. 49.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 2019.
  50. 50.Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China, November 2019. Association for Computational Linguistics.
  51. 51.Jigar Shetty and Jafar Adibi. The enron email dataset: Database schema and brief statistical report. In Information Retrieval Research, 2004.
  52. 52.Kai Shu, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Hassan Awadallah, Milad Shokouhi, and Susan Dumais. Learning with weak supervision for email intent detection. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’20, page 1051–1060, New York, NY, USA, 2020. Association for Computing Machinery.
  53. 53.Leslie N Smith. Cyclical learning rates for training neural networks. 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 464–472, 2017.
  54. 54.Duncan Soiffer, Steven Kolawole, and Virginia Smith. Semantic agreement enables efficient open-ended LLM cascades. In Saloni Potdar, Lina Rojas-Barahona, and Sebastien Montella, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 2499–2537, Suzhou (China), November 2025. Association for Computational Linguistics.
  55. 55.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958, 2014.
  56. 56.Computerworld Staff. Is your gmail inbox setup slowing you down? https://www.computerworld.com/article/3511582/is-your-gmail-inbox-setup-slowing-you-down.html. Accessed: 2025-12-08.
  57. 57.Jesicca Stockett. Gmail categories and inbox tabs. https://swatkb.atlassian.net/wiki/spaces/GA/pages/19661188/Gmail+Categories+and+Inbox+Tabs. Accessed: 2025-12-08.
  58. 58.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. In ACL, 2019.
  59. 59.Jian Tang, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, and Qiaozhu Mei. Line: Large-scale information network embedding. In Proceedings of the 24th International Conference on World Wide Web, WWW ’15, page 1067–1077, Republic and Canton of Geneva, CHE, 2015. International World Wide Web Conferences Steering Committee.
  60. 60.U.S. Department of State. Hillary clinton email archive. https://wikileaks.org/clinton-emails/, 2016. FOIA Release; Accessed 2025-12-09.
  61. 61.Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Improving text embeddings with large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11897–11916, Bangkok, Thailand, August 2024. Association for Computational Linguistics.
  62. 62.Wei Wang, Saghar Hosseini, Ahmed Awadallah, Paul Bennett, and Chris Quirk. Context-aware intent identification in email conversations. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1209–1218. ACM, 2019.
  63. 63.Steve Whittaker and Candace Sidner. Email overload: exploring personal information management of email. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’96, page 276–283, New York, NY, USA, 1996. Association for Computing Machinery.
  64. 64.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, October 2020. Association for Computational Linguistics.
  65. 65.Mike Wong, Murali Ramanujam, Guha Balakrishnan, and Ravi Netravali. MadEye: Boosting live video analytics accuracy with adaptive camera configurations. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 549–568, Santa Clara, CA, April 2024. USENIX Association.
  66. 66.Liu Yang, Susan Dumais, Paul Bennett, and Ahmed Awadallah. Characterizing and predicting enterprise email reply behavior. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 505–514. ACM, 2017.
  67. 67.Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. Hierarchical attention networks for document classification. In NAACL, 2016.
  68. 68.Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J. Freedman. Live video analytics at scale with approximation and Delay-Tolerance. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17), pages 377–392, Boston, MA, March 2017. USENIX Association.
  69. 69.Rui Zhang and Chen Li. Email reply and importance prediction using pre-trained language models. Information Processing and Management, 2021.

Citation

MLA
Ray, S., et al. “Argo: Efficient Importance Labeling for Enterprise Email Systems”. arXiv, 2026, http://arxiv.org/abs/2605.21604v1.
APA
Ray, S., Ananthanarayanan, G., Chian, K., Guo, Y., Hill, C. S., Stokes, J. W., Wang, V., & Jiang, J. (2026). Argo: Efficient Importance Labeling for Enterprise Email Systems. arXiv. http://arxiv.org/abs/2605.21604v1
Chicago
Ray, S., G. Ananthanarayanan, K. Chian, et al. 2026. “Argo: Efficient Importance Labeling for Enterprise Email Systems”. arXiv. http://arxiv.org/abs/2605.21604v1.
Harvard
Ray, S. et al. (2026) “Argo: Efficient Importance Labeling for Enterprise Email Systems”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.21604v1.
Vancouver
1. Ray S, Ananthanarayanan G, Chian K, Guo Y, Hill CS, Stokes JW, Wang V, Jiang J (2026) Argo: Efficient Importance Labeling for Enterprise Email Systems. arXiv

BibTeX

@article{ray2026argo,
  title = {Argo: Efficient Importance Labeling for Enterprise Email Systems},
  author = {Ray, Siddhant and Ananthanarayanan, Ganesh and Chian, Kevin and Guo, Yan and Hill, Cristina St and Stokes, Jack W. and Wang, Victor and Jiang, Junchen},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.21604v1},
  eprint = {2605.21604}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/