Two-Stage Learning to Defer with Multiple Experts

Anqi MaoChristopher MohriMehryar MohriYutao Zhong

article2023NeurIPS95 citations

Proposes a theoretically grounded framework for two-stage learning to defer with multiple experts by designing surrogate losses with rigorous H\mathcal{H}-consistency guarantees that enable efficient post-hoc deferral without retraining the underlying predictor.

Listen

Modern machine learning models, including large language models, often encounter critical real-world challenges such as costly inference and factual inaccuracies or hallucinations. A practical way to mitigate these issues is to defer challenging or uncertain inputs to external specialized experts or larger models. However, standard learning-to-defer methods require training the base predictor and the routing mechanism simultaneously from scratch. This single-stage approach is prohibitively expensive and impractical for organizations that already possess established, pre-trained base models.

The article addresses this bottleneck by developing and analyzing a principled two-stage framework for learning to defer with multiple experts. Its primary objective is to demonstrate that an organization can take an existing, fixed predictor trained with standard classification methods and train an effective deferral mechanism in a second stage while retaining strong theoretical performance guarantees.

To evaluate this framework, the authors formulated a new family of surrogate loss functions across two common deferral designs: score-based systems (which output extra routing scores) and predictor-rejector systems (which use dedicated routing functions). They derived theoretical bounds connecting the training objective directly to true decision accuracy and validated the approach empirically on benchmark image datasets (CIFAR-10 and SVHN) using residual neural networks paired with up to three expert models of varying capacity and computational cost.

The findings confirm that the proposed two-stage framework is both theoretically sound and practically effective. First, the surrogate loss functions resolve an open theoretical challenge in deferral literature by providing formal non-asymptotic consistency guarantees (H-consistency) and realizability guarantees under constant inference costs. Second, empirical evaluations demonstrate that routing inputs across multiple experts consistently improves overall accuracy. On the CIFAR-10 benchmark, the base model accuracy of 70.56% increased to 77.68% when deferring across three higher-capacity experts under zero base cost, and to 72.42% when accounting for expert inference costs. On SVHN, accuracy improved from 91.12% to 93.30% and 92.19% under the respective cost scenarios, with performance scaling smoothly as additional experts were made available.

These results have significant operational implications for high-stakes and resource-constrained environments. By decoupling the deferral stage from base model training, organizations can substantially reduce compute costs, deployment timelines, and engineering risks associated with retraining massive models. Furthermore, incorporating expert-specific inference costs allows decision-makers to explicitly balance operational expense against classification accuracy.

Based on these findings, teams deploying pre-trained models should implement two-stage deferral when upgrading system accuracy with expert models or when seeking to triage computational load. Before wide-scale adoption, engineering teams should conduct domain-specific pilot testing to tune the cost parameters that govern when to defer versus predict directly.

A key limitation identified in the article is that the cost hyperparameters assigned to experts are not automated and currently rely on cross-validation tuning. Nonetheless, there is high confidence in the foundational methods due to rigorous mathematical proofs and consistent empirical verification across multiple datasets and expert configurations.

  • Paper: Learning From Crowds, V. Raykar et al. (2010). It provides foundational principles for modeling heterogeneous annotator and expert quality without ground truth, which motivates the expert assignment mechanisms in learning to defer.
  • Paper: The foundations of cost-sensitive learning, Charles Elkan (2001). It establishes the core decision-theoretic framework for cost-sensitive classification and threshold adjustments that underpins surrogate loss design and constant cost deferral.
  • Paper: Robust Classification for Imprecise Environments, F. Provost et al. (2000). It introduces hybrid decision rules and operating-condition trade-offs between predictive models that inform predictor-rejector and deferral systems.

No sufficiently relevant recommendations were found.

Cover for Two-Stage Learning to Defer with Multiple Experts

Abstract

We study a two-stage scenario for learning to defer with multiple experts, which is crucial in practice for many applications. In this scenario, a predictor is derived in a first stage by training with a common loss function such as cross-entropy. In the second stage, a deferral function is learned to assign the most suitable expert to each input. We design a new family of surrogate loss functions for this scenario both in the score-based and the predictor-rejector settings and prove that they are supported by H-consistency bounds, which implies their Bayes-consistency. Moreover, we show that, for a constant cost function, our two-stage surrogate losses are realizable H-consistent. While the main focus of this work is a theoretical analysis, we also report the results of several experiments on CIFAR-10 and SVHN datasets.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Two-stage H-consistent surrogate loss
  • 3.1 General surrogate losses
  • 3.2 H-consistency bounds for two-stage surrogate losses
  • 3.3 H-consistency bounds for standard surrogate loss functions
  • 3.4 Realizable H-consistency
  • 4 Predictor-rejector setting
  • 5 Experiments
  • 6 Conclusion
  • References
  • A Related work
  • B Examples of two-stage score-based surrogate losses
  • C Examples of two-stage predictor-rejector surrogate losses
  • D Proof of H-consistency bounds for score-based two-stage surrogate losses (Theorem 1)
  • E Proof of H-consistency bounds for standard surrogate loss functions (Theorem 3)
  • E.1 Multinomial logistic loss
  • E.2 Sum exponential loss
  • E.3 Generalized cross-entropy loss
  • E.4 Mean absolute error loss
  • F Proof of realizable consistency for score-based two-stage surrogate losses (Theorem 5)
  • G Proof of (H, R)-consistency bounds for predictor-rejector two-stage surrogate losses (Theorem 6)
  • H Proof of realizable consistency for predictor-rejector two-stage surrogate losses (Theorem 7)

Knowls

  1. Knowl 1 — Learning-to-defer problem with multiple experts

    definition

    Let X\mathcal X be an input space, let Y=[n]={1,…,n}\mathcal Y=[n]=\{1,\ldots,n\} be the n≥2n\ge 2 prediction labels, and let h1,…,hneh_1,\ldots,h_{n_e} be fixed experts. The learner may either predict a label in Y\mathcal Y or defer to expert hjh_j, represented by the augmented label n+jn+j.

    In the score-based formulation, a hypothesis h:X×[n+ne]→Rh:\mathcal X\times[ n+n_e]\to\mathbb R selects

    h(x)=arg⁡max⁡k∈[n+ne]h(x,k),h(x)=\arg\max_{k\in[n+n_e]}h(x,k),

    with deterministic tie breaking. If h(x)∈Yh(x)\in\mathcal Y, its loss is the zero-one loss; if h(x)=n+jh(x)=n+j, its loss is the expert-specific cost cj(x,y)∈[0,1]c_j(x,y)\in[0,1]:

    Ldef(h,x,y)=1{h(x)≠y}1{h(x)∈Y}+∑j=1necj(x,y)1{h(x)=n+j}.L_{\mathrm{def}}(h,x,y)=\mathbf 1\{h(x)\ne y\}\mathbf 1\{h(x)\in\mathcal Y\}+\sum_{j=1}^{n_e}c_j(x,y)\mathbf 1\{h(x)=n+j\}.

    The complementary cost is cˉj(x,y)=1−cj(x,y)\bar c_j(x,y)=1-c_j(x,y). The paper allows input- and label-dependent costs, including cj(x,y)=αj1{hj(x)≠y}+βjc_j(x,y)=\alpha_j\mathbf 1\{h_j(x)\ne y\}+\beta_j, where αj\alpha_j measures expert error and βj\beta_j is an inference cost. The two-stage setting first learns the predictor scores for the original nn classes and then learns only the deferral scores while keeping the predictor fixed.

  2. Knowl 2 — Score-based two-stage surrogate family

    model/method

    A score hypothesis class is decomposed as H=Hp×Hd\mathcal H=\mathcal H_p\times\mathcal H_d, where hp∈Hph_p\in\mathcal H_p supplies the first nn prediction scores and hd∈Hdh_d\in\mathcal H_d supplies the nen_e deferral scores. The first stage minimizes a standard nn-class surrogate loss ℓ1\ell_1 over hph_p.

    For the second stage, define an augmented (ne+1)(n_e+1)-class score function hˉd\bar h_d by treating the predictor as class 00:

    hˉd(x,0)=max⁡y∈[n]hp(x,y),hˉd(x,j)=hd(x,j),j∈[ne].\bar h_d(x,0)=\max_{y\in[n]}h_p(x,y),\qquad \bar h_d(x,j)=h_d(x,j),\quad j\in[n_e].

    The score for class 00 is fixed during the second stage. Given any standard (ne+1)(n_e+1)-class surrogate loss ℓ2\ell_2, the proposed second-stage loss is

    Lhp(hd,x,y)=1{hp(x)=y}ℓ2(hˉd,x,0)+∑j=1necˉj(x,y)ℓ2(hˉd,x,j).L_{h_p}(h_d,x,y)=\mathbf 1\{h_p(x)=y\}\ell_2(\bar h_d,x,0)+\sum_{j=1}^{n_e}\bar c_j(x,y)\ell_2(\bar h_d,x,j).

    Thus, the first-stage loss trains the available predictor, while the second-stage loss trains the deferral scores to choose among the predictor (class 00) and the nen_e experts. The construction supports arbitrary bounded, instance-dependent costs and multiple deferral choices. The paper instantiates ℓ2\ell_2 with sum-exponential, multinomial logistic, generalized cross-entropy, and mean-absolute-error losses.

  3. Knowl 3 — H-consistency bound for score-based two-stage learning

    theoretical result

    Let Eℓ(f)=E[ℓ(f,X,Y)]\mathcal E_\ell(f)=\mathbb E[\ell(f,X,Y)] and Eℓ∗(F)=inf⁡f∈FEℓ(f)\mathcal E_\ell^*(\mathcal F)=\inf_{f\in\mathcal F}\mathcal E_\ell(f). For a loss ℓ\ell and class F\mathcal F, define the minimizability gap by

    Mℓ(F)=Eℓ∗(F)−EX[inf⁡f∈FEY∣X[ℓ(f,X,Y)]].M_\ell(\mathcal F)=\mathcal E_\ell^*(\mathcal F)-\mathbb E_X\left[\inf_{f\in\mathcal F}\mathbb E_{Y\mid X}[\ell(f,X,Y)]\right].

    Assume that the first-stage loss ℓ1\ell_1 has an Hp\mathcal H_p-consistency bound with respect to nn-class zero-one loss, and that the second-stage loss ℓ2\ell_2 has an Hd\mathcal H_d-consistency bound with respect to (ne+1)(n_e+1)-class zero-one loss. Specifically, assume there are non-decreasing concave functions Γ1,Γ2\Gamma_1,\Gamma_2 such that the corresponding zero-one excess risks plus minimizability gaps are bounded by Γ1\Gamma_1 and Γ2\Gamma_2 applied to the respective surrogate excess risks.

    Let bjb_j be finite upper bounds satisfying cˉj(x,y)≤bj\bar c_j(x,y)\le b_j and let B=∑j=1nebjB=\sum_{j=1}^{n_e}b_j. Then every score-based two-stage hypothesis h=(hp,hd)h=(h_p,h_d) satisfies

    ELdef(h)−ELdef∗(H)+MLdef(H)≤Γ1 ⁣(Eℓ1(hp)−Eℓ1∗(Hp)+Mℓ1(Hp))+(1+B)Γ2 ⁣(ELhp(hd)−ELhp∗(Hd)+MLhp(Hd)B).\begin{aligned} \mathcal E_{L_{\mathrm{def}}}(h)-\mathcal E_{L_{\mathrm{def}}}^*(\mathcal H)+M_{L_{\mathrm{def}}}(\mathcal H) \le{}&\Gamma_1\!\left(\mathcal E_{\ell_1}(h_p)-\mathcal E_{\ell_1}^*(\mathcal H_p)+M_{\ell_1}(\mathcal H_p)\right)\\ &+(1+B)\Gamma_2\!\left(\frac{\mathcal E_{L_{h_p}}(h_d)-\mathcal E_{L_{h_p}}^*(\mathcal H_d)+M_{L_{h_p}}(\mathcal H_d)}{B}\right). \end{aligned}

    When Γ2\Gamma_2 is linear, the factors 1+B1+B and 1/B1/B can be removed. Consequently, reducing the estimation errors of both stage losses quantitatively controls the excess deferral loss. When the relevant best-in-class errors equal their Bayes errors, the minimizability gaps vanish; taking the hypothesis classes to be all measurable functions then gives Bayes-consistency of the two-stage surrogate.

  4. Knowl 4 — Fixed-score multiclass consistency bounds

    theoretical result

    The paper proves the second-stage condition needed by the two-stage construction. Consider a standard KK-class classification problem with score functions for K−1K-1 classes drawn from a symmetric and complete hypothesis class, together with one additional class whose score is an arbitrary fixed function λ:X→R\lambda:\mathcal X\to\mathbb R. Let Hˉλ\bar{\mathcal H}_\lambda denote this augmented class. Symmetry means that the class-wise score functions can be chosen independently from a common family, and completeness means every individual score can attain every real value.

    For any f∈Hˉλf\in\bar{\mathcal H}_\lambda, the excess KK-class zero-one risk is bounded by

    E0-1(f)−E0-1∗(Hˉλ)≤Γ ⁣(Eℓ(f)−Eℓ∗(Hˉλ)+Mℓ(Hˉλ))−M0-1(Hˉλ).\mathcal E_{0\text{-}1}(f)-\mathcal E_{0\text{-}1}^*(\bar{\mathcal H}_\lambda) \le \Gamma\!\left(\mathcal E_\ell(f)-\mathcal E_\ell^*(\bar{\mathcal H}_\lambda)+M_\ell(\bar{\mathcal H}_\lambda)\right)-M_{0\text{-}1}(\bar{\mathcal H}_\lambda).

    The bound remains valid even though the score of one class is fixed, which is the technically important case for the second stage. For the following standard losses, with qk(x)q_k(x) the score for class kk and pq(k∣x)=eqk(x)/∑i=1Keqi(x)p_q(k\mid x)=e^{q_k(x)}/\sum_{i=1}^K e^{q_i(x)}, the corresponding functions are:

    ℓexp⁡(q,x,y)=∑k≠yeqk(x)−qy(x),Γ(t)=2t;\ell_{\exp}(q,x,y)=\sum_{k\ne y}e^{q_k(x)-q_y(x)},\qquad \Gamma(t)=\sqrt{2t}; ℓlog⁡(q,x,y)=log⁡(∑k=1Keqk(x)−qy(x)),Γ(t)=2t;\ell_{\log}(q,x,y)=\log\left(\sum_{k=1}^K e^{q_k(x)-q_y(x)}\right),\qquad \Gamma(t)=\sqrt{2t}; ℓgce(q,x,y)=1α[1−pq(y∣x)α],α∈(0,1),Γ(t)=2Kαt;\ell_{\mathrm{gce}}(q,x,y)=\frac{1}{\alpha}\left[1-p_q(y\mid x)^\alpha\right],\quad \alpha\in(0,1),\qquad \Gamma(t)=\sqrt{2K^\alpha t}; ℓmae(q,x,y)=1−pq(y∣x),Γ(t)=Kt.\ell_{\mathrm{mae}}(q,x,y)=1-p_q(y\mid x),\qquad \Gamma(t)=Kt.

    Therefore, all four standard losses can be used in the second stage despite its fixed predictor score. In particular, logistic loss in both stages gives square-root H-consistency bounds and Bayes-consistency.

  5. Knowl 5 — Realizable consistency of the score-based surrogate

    theoretical result

    A score-based surrogate is realizable H\mathcal H-consistent if, whenever a distribution admits some h∗∈Hh^*\in\mathcal H with ELdef(h∗)=0\mathcal E_{L_{\mathrm{def}}}(h^*)=0, any sequence whose surrogate excess risk converges to zero also has deferral excess risk converging to zero.

    The proposed score-based two-stage surrogate has a stronger exact result under the following conditions: H\mathcal H is closed under scaling, meaning h∈Hh\in\mathcal H implies τh∈H\tau h\in\mathcal H for every real τ\tau; the first- and second-stage losses are both multinomial logistic loss; and every expert cost is constant, cj(x,y)=βjc_j(x,y)=\beta_j.

    If h^p\hat h_p is a global minimizer of the first-stage expected logistic loss and h^d\hat h_d is a global minimizer of ELh^p\mathcal E_{L_{\hat h_p}} over the second-stage class, then every score-realizable distribution satisfies

    ELdef(h^p,h^d)=0.\mathcal E_{L_{\mathrm{def}}}(\hat h_p,\hat h_d)=0.

    Thus, under constant costs, the same two-stage losses are both Bayes-consistent through the H-consistency bounds and realizable H-consistent under restricted, scaling-closed hypothesis classes. The asymptotic version follows when both stagewise surrogate excess risks converge to zero.

  6. Knowl 6 — Predictor-rejector two-stage surrogate

    model/method

    In the predictor-rejector formulation, the predictor is a score function h:X×[n]→Rh:\mathcal X\times[n]\to\mathbb R and the deferral mechanism is a separate vector-valued function r=(r1,…,rne)r=(r_1,\ldots,r_{n_e}) from a class R\mathcal R. The predictor outputs h(x)=arg⁡max⁡y∈[n]h(x,y)h(x)=\arg\max_{y\in[n]}h(x,y). Deferral to expert jj occurs when rj(x)≤0r_j(x)\le 0 and rj(x)<ri(x)r_j(x)<r_i(x) for every i≠ji\ne j; otherwise the predictor is used. Equivalently, define r0(x)=0r_0(x)=0 and select the smallest score among r0(x),r1(x),…,rne(x)r_0(x),r_1(x),\ldots,r_{n_e}(x), where index 00 denotes the predictor.

    The target loss is

    Ldef(h,r,x,y)=1{h(x)≠y}1{r(x)=0}+∑j=1necj(x,y)1{r(x)=j}.L_{\mathrm{def}}(h,r,x,y)=\mathbf 1\{h(x)\ne y\}\mathbf 1\{r(x)=0\}+\sum_{j=1}^{n_e}c_j(x,y)\mathbf 1\{r(x)=j\}.

    For the second stage, represent the rejector as an (ne+1)(n_e+1)-class score function rˉ\bar r with rˉ(x,0)=0\bar r(x,0)=0 and rˉ(x,j)=−rj(x)\bar r(x,j)=-r_j(x). Given any standard (ne+1)(n_e+1)-class loss ℓ2\ell_2, the proposed surrogate is

    Lh(r,x,y)=1{h(x)=y}ℓ2(rˉ,x,0)+∑j=1necˉj(x,y)ℓ2(rˉ,x,j).L_h(r,x,y)=\mathbf 1\{h(x)=y\}\ell_2(\bar r,x,0)+\sum_{j=1}^{n_e}\bar c_j(x,y)\ell_2(\bar r,x,j).

    The first stage minimizes a standard classification loss for hh, and the second stage minimizes LhL_h for rr with hh fixed. The negative sign in rˉ(x,j)=−rj(x)\bar r(x,j)=-r_j(x) aligns standard classification scores, which favor large values, with the rejector rule, which selects small values of rj(x)r_j(x).

  7. Knowl 7 — H-consistency bound for predictor-rejector learning

    theoretical result

    Let H\mathcal H be the predictor class and R\mathcal R the rejector class. Assume that the first-stage loss ℓ1\ell_1 has an H\mathcal H-consistency bound with respect to standard nn-class zero-one loss, and that the second-stage loss ℓ2\ell_2 has an Rˉ\bar{\mathcal R}-consistency bound with respect to (ne+1)(n_e+1)-class zero-one loss for the augmented rejector scores rˉ(x,0)=0\bar r(x,0)=0 and rˉ(x,j)=−rj(x)\bar r(x,j)=-r_j(x). Let the corresponding non-decreasing concave functions be Γ1\Gamma_1 and Γ2\Gamma_2.

    Using B=∑j=1nebjB=\sum_{j=1}^{n_e}b_j, where cˉj(x,y)≤bj\bar c_j(x,y)\le b_j, every predictor-rejector pair satisfies

    ELdef(h,r)−ELdef∗(H,R)+MLdef(H,R)≤Γ1 ⁣(Eℓ1(h)−Eℓ1∗(H)+Mℓ1(H))+(1+B)Γ2 ⁣(ELh(r)−ELh∗(R)+MLh(R)B).\begin{aligned} \mathcal E_{L_{\mathrm{def}}}(h,r)-\mathcal E_{L_{\mathrm{def}}}^*(\mathcal H,\mathcal R)+M_{L_{\mathrm{def}}}(\mathcal H,\mathcal R) \le{}&\Gamma_1\!\left(\mathcal E_{\ell_1}(h)-\mathcal E_{\ell_1}^*(\mathcal H)+M_{\ell_1}(\mathcal H)\right)\\ &+(1+B)\Gamma_2\!\left(\frac{\mathcal E_{L_h}(r)-\mathcal E_{L_h}^*(\mathcal R)+M_{L_h}(\mathcal R)}{B}\right). \end{aligned}

    If Γ2\Gamma_2 is linear, the multiplicative factor 1+B1+B and the normalization by BB can be removed. When the two stagewise best-in-class errors equal their Bayes errors, the minimizability gaps vanish, and the bound gives the corresponding Bayes-consistency guarantee. The result establishes that the predictor-rejector formulation enjoys the same quantitative two-stage guarantee as the score-based formulation.

  8. Knowl 8 — Realizable consistency of the predictor-rejector surrogate

    theoretical result

    Assume that the predictor class H\mathcal H and rejector class R\mathcal R are both closed under scaling, every expert cost is constant, cj(x,y)=βjc_j(x,y)=\beta_j, and both stages use multinomial logistic loss. Suppose a distribution is (H,R)(\mathcal H,\mathcal R)-realizable, meaning that some pair (h∗,r∗)∈H×R(h^*,r^*)\in\mathcal H\times\mathcal R has zero expected deferral loss.

    Let h^\hat h minimize the first-stage expected logistic loss and let r^\hat r minimize the second-stage loss ELh^\mathcal E_{L_{\hat h}} over R\mathcal R. Then

    ELdef(h^,r^)=0.\mathcal E_{L_{\mathrm{def}}}(\hat h,\hat r)=0.

    Consequently, if the first-stage and second-stage surrogate excess risks converge to zero rather than being minimized exactly, the deferral excess risk also converges to zero. Hence the predictor-rejector construction is both Bayes-consistent through its H-consistency bound and realizable (H,R)(\mathcal H,\mathcal R)-consistent for constant costs.

  9. Knowl 9 — CIFAR-10 and SVHN evaluation of multiple experts

    empirical result

    The experiments evaluate the score-based two-stage method with a pre-trained predictor and a separately trained deferral model. The predictor and deferral model are ResNet-4 networks; the three available experts are ResNet-10, ResNet-16, and ResNet-28, ordered by increasing capacity. Both stages use logistic loss. Adam is trained with batch size 128128, weight decay 10−410^{-4}, the default learning rate, and no data augmentation; training lasts 15 epochs on SVHN and 50 epochs on CIFAR-10. Results are means plus standard deviations over three runs, and all entries are accuracies in percent.

    Without a base cost, the cost is cj(x,y)=1{hj(x)≠y}c_j(x,y)=\mathbf 1\{h_j(x)\ne y\}. With a base cost, the costs use cj(x,y)=1{hj(x)≠y}+βjc_j(x,y)=\mathbf 1\{h_j(x)\ne y\}+\beta_j in the experimental implementation, with βj=(0.1,0.12,0.14)\beta_j=(0.1,0.12,0.14) for SVHN and (0.3,0.32,0.34)(0.3,0.32,0.34) for CIFAR-10 as expert capacity increases. The reported overall accuracies are:

    • SVHN, no base cost: base model 91.1291.12; one expert 91.85±0.01%91.85\pm0.01\%; two experts 92.77±0.02%92.77\pm0.02\%; three experts 93.30±0.02%93.30\pm0.02\%.
    • CIFAR-10, no base cost: base model 70.5670.56; one expert 72.63±0.20%72.63\pm0.20\%; two experts 75.84±0.35%75.84\pm0.35\%; three experts 77.68±0.07%77.68\pm0.07\%.
    • SVHN, with base cost: base model 91.1291.12; one expert 91.66±0.01%91.66\pm0.01\%; two experts 92.05±0.10%92.05\pm0.10\%; three experts 92.19±0.03%92.19\pm0.03\%.
    • CIFAR-10, with base cost: base model 70.5670.56; one expert 71.73±0.06%71.73\pm0.06\%; two experts 72.31±0.31%72.31\pm0.31\%; three experts 72.42±0.12%72.42\pm0.12\%.

    Accuracy increases monotonically as more experts become available in both cost settings. Without a base cost, overall accuracy equals one minus the expected deferral loss, so the results directly demonstrate improved deferral performance; with base costs, the results show the tradeoff between accuracy improvement and inference cost.

  10. Knowl 10 — Cost selection remains an unresolved practical issue

    limitation

    The deferral loss depends on the expert cost functions cj(x,y)c_j(x,y) and their base-cost parameters, but the paper does not provide a principled procedure for selecting these costs. In practice, the costs are typically chosen by cross-validation. The experiments use manually specified neighboring base-cost values and observe similar results for nearby choices, but identifying theoretically justified costs across tasks is left for future work.

Coverage note — Proof derivations, related work, and the fully expanded algebraic forms for the non-logistic surrogate instantiations were omitted because the substantive theorem statements, general surrogate constructions, logistic implementation setting, experiments, and limitation are retained.

References

  1. 1.D. A. E. Acar, A. Gangrade, and V. Saligrama. Budget learning via bracketing. In International Conference on Artificial Intelligence and Statistics, pages 4109–4119, 2020.
  2. 2.P. Awasthi, N. Frank, A. Mao, M. Mohri, and Y. Zhong. Calibration and consistency of adversarial surrogate losses. In Advances in Neural Information Processing Systems, 2021a.
  3. 3.P. Awasthi, A. Mao, M. Mohri, and Y. Zhong. A finer calibration analysis for adversarial robustness. arXiv preprint arXiv:2105.01550, 2021b.
  4. 4.P. Awasthi, A. Mao, M. Mohri, and Y. Zhong. Multi-class H-consistency bounds. In Advances in neural information processing systems, 2022a.
  5. 5.P. Awasthi, A. Mao, M. Mohri, and Y. Zhong. H-consistency bounds for surrogate loss minimizers. In International Conference on Machine Learning, 2022b.
  6. 6.P. Awasthi, A. Mao, M. Mohri, and Y. Zhong. Theoretically grounded loss functions and algorithms for adversarial robustness. In International Conference on Artificial Intelligence and Statistics, pages 10077–10094, 2023a.
  7. 7.P. Awasthi, A. Mao, M. Mohri, and Y. Zhong. DC-programming for neural network optimizations. Journal of Global Optimization, 2023b.
  8. 8.G. Bansal, B. Nushi, E. Kamar, E. Horvitz, and D. S. Weld. Is the most accurate ai the best teammate? optimizing ai for teamwork. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11405–11414, 2021.
  9. 9.P. L. Bartlett and M. H. Wegkamp. Classification with a reject option using a hinge loss. Journal of Machine Learning Research, 9(8), 2008.
  10. 10.P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101(473):138–156, 2006.
  11. 11.N. L. C. Benz and M. G. Rodriguez. Counterfactual inference of second opinions. In Uncertainty in Artificial Intelligence, pages 453–463. PMLR, 2022.
  12. 12.S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  13. 13.Y. Cao, T. Cai, L. Feng, L. Gu, J. Gu, B. An, G. Niu, and M. Sugiyama. Generalizing consistent multi-class classification with rejection to be compatible with arbitrary losses. In Advances in neural information processing systems, 2022.
  14. 14.N. Charoenphakdee, Z. Cui, Y. Zhang, and M. Sugiyama. Classification with rejection based on cost-sensitive classification. In International Conference on Machine Learning, pages 1507–1517, 2021.
  15. 15.M.-A. Charusaie, H. Mozannar, D. Sontag, and S. Samadi. Sample efficient learning of predictors that complement humans. In International Conference on Machine Learning, pages 2972–3005, 2022.
  16. 16.C. Chow. An optimum character recognition system using decision function. IEEE T. C., 1957.
  17. 17.C. Chow. On optimum recognition error and reject tradeoff. IEEE Transactions on information theory, 16(1):41–46, 1970.
  18. 18.C. Cortes, G. DeSalvo, and M. Mohri. Learning with rejection. In International Conference on Algorithmic Learning Theory, pages 67–82, 2016a.
  19. 19.C. Cortes, G. DeSalvo, and M. Mohri. Boosting with abstention. In Advances in Neural Information Processing Systems, pages 1660–1668, 2016b.
  20. 20.C. Cortes, G. DeSalvo, and M. Mohri. Theory and algorithms for learning with rejection in binary classification. Annals of Mathematics and Artificial Intelligence, to appear, 2023.
  21. 21.A. De, P. Koley, N. Ganguly, and M. Gomez-Rodriguez. Regression under human assistance. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2611–2620, 2020.
  22. 22.R. El-Yaniv et al. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(5), 2010.
  23. 23.A. Gangrade, A. Kag, and V. Saligrama. Selective classification via one-sided prediction. In International Conference on Artificial Intelligence and Statistics, pages 2179–2187, 2021.
  24. 24.R. Gao, M. Saar-Tsechansky, M. De-Arteaga, L. Han, M. K. Lee, and M. Lease. Human-ai collaboration with bandit feedback. arXiv preprint arXiv:2105.10614, 2021.
  25. 25.Y. Geifman and R. El-Yaniv. Selective classification for deep neural networks. In Advances in neural information processing systems, 2017.
  26. 26.Y. Geifman and R. El-Yaniv. Selectivenet: A deep neural network with an integrated reject option. In International conference on machine learning, pages 2151–2159, 2019.
  27. 27.Y. Grandvalet, A. Rakotomamonjy, J. Keshet, and S. Canu. Support vector machines with a reject option. In Advances in neural information processing systems, 2008.
  28. 28.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  29. 29.P. Hemmer, S. Schellhammer, M. Vössing, J. Jakubik, and G. Satzger. Forming effective human-ai teams: Building machine learning models that complement the capabilities of multiple experts. arXiv preprint arXiv:2206.07948, 2022.
  30. 30.P. Hemmer, L. Thede, M. Vössing, J. Jakubik, and N. Kühl. Learning to defer with limited expert predictions. arXiv preprint arXiv:2304.07306, 2023.
  31. 31.R. Herbei and M. Wegkamp. Classification with reject option. Can. J. Stat., 2005.
  32. 32.S. Joshi, S. Parbhoo, and F. Doshi-Velez. Pre-emptive learning-to-defer for sequential medical decision-making under uncertainty. arXiv preprint arXiv:2109.06312, 2021.
  33. 33.A. T. Kalai, V. Kanade, and Y. Mansour. Reliable agnostic learning. Journal of Computer and System Sciences, 78(5):1481–1495, 2012.
  34. 34.E. Kamar, S. Hacker, and E. Horvitz. Combining human and machine intelligence in large-scale crowdsourcing. In AAMAS, pages 467–474, 2012.
  35. 35.G. Kerrigan, P. Smyth, and M. Steyvers. Combining human predictions with model probabilities via confusion matrices and calibration. Advances in Neural Information Processing Systems, 34:4421–4434, 2021.
  36. 36.V. Keswani, M. Lease, and K. Kenthapadi. Towards unbiased and accurate deferral to multiple experts. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 154–165, 2021.
  37. 37.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  38. 38.J. Kleinberg, H. Lakkaraju, J. Leskovec, J. Ludwig, and S. Mullainathan. Human decisions and machine predictions. The quarterly journal of economics, 133(1):237–293, 2018.
  39. 39.A. Krizhevsky. Learning multiple layers of features from tiny images. Technical report, Toronto University, 2009.
  40. 40.V. Kuznetsov, M. Mohri, and U. Syed. Multi-class deep boosting. In Advances in Neural Information Processing Systems, pages 2501–2509, 2014.
  41. 41.J. Liu, B. Gallego, and S. Barbieri. Incorporating uncertainty in learning to defer algorithms for safe computer-aided diagnosis. Scientific reports, 12(1):1762, 2022.
  42. 42.P. Long and R. Servedio. Consistency versus realizable H-consistency for multiclass classification. In International Conference on Machine Learning, pages 801–809, 2013.
  43. 43.D. Madras, E. Creager, T. Pitassi, and R. Zemel. Learning adversarially fair and transferable representations. arXiv preprint arXiv:1802.06309, 2018.
  44. 44.A. Mao, M. Mohri, and Y. Zhong. H-consistency bounds: Characterization and extensions. In Advances in Neural Information Processing Systems, 2023a.
  45. 45.A. Mao, M. Mohri, and Y. Zhong. Principled approaches for learning to defer with multiple experts. arXiv preprint arXiv:2310.14774, 2023b.
  46. 46.A. Mao, M. Mohri, and Y. Zhong. Predictor-rejector multi-class abstention: Theoretical analysis and algorithms. arXiv preprint arXiv:2310.14772, 2023c.
  47. 47.A. Mao, M. Mohri, and Y. Zhong. H-consistency bounds for pairwise misranking loss surrogates. In International conference on Machine learning, 2023d.
  48. 48.A. Mao, M. Mohri, and Y. Zhong. Ranking with abstention. In ICML 2023 Workshop The Many Facets of Preference-Based Learning, 2023e.
  49. 49.A. Mao, M. Mohri, and Y. Zhong. Theoretically grounded loss functions and algorithms for score-based multi-class abstention. arXiv preprint arXiv:2310.14770, 2023f.
  50. 50.A. Mao, M. Mohri, and Y. Zhong. Structured prediction with stronger consistency guarantees. In Advances in Neural Information Processing Systems, 2023g.
  51. 51.A. Mao, M. Mohri, and Y. Zhong. Cross-entropy loss functions: Theoretical analysis and applications. In International Conference on Machine Learning, 2023h.
  52. 52.C. Mohri, D. Andor, E. Choi, M. Collins, A. Mao, and Y. Zhong. Learning to reject with a fixed predictor: Application to decontextualization. arXiv preprint arXiv:2301.09044, 2023.
  53. 53.M. Mohri, A. Rostamizadeh, and A. Talwalkar. Foundations of Machine Learning. MIT Press, second edition, 2018.
  54. 54.H. Mozannar and D. Sontag. Consistent estimators for learning to defer to an expert. In International Conference on Machine Learning, pages 7076–7087, 2020.
  55. 55.H. Mozannar, A. Satyanarayan, and D. Sontag. Teaching humans when to defer to a classifier via exemplars. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 5323–5331, 2022.
  56. 56.H. Mozannar, H. Lang, D. Wei, P. Sattigeri, S. Das, and D. Sontag. Who should predict? exact algorithms for learning to defer to humans. In International Conference on Artificial Intelligence and Statistics, pages 10520–10545, 2023.
  57. 57.H. Narasimhan, W. Jitkrittum, A. K. Menon, A. S. Rawat, and S. Kumar. Post-hoc estimators for learning to defer to an expert. In Advances in Neural Information Processing Systems, 2022.
  58. 58.H. Narasimhan, A. K. Menon, W. Jitkrittum, and S. Kumar. Learning to reject meets ood detection: Are all abstentions created equal? arXiv preprint arXiv:2301.12386, 2023.
  59. 59.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. In Advances in Neural Information Processing Systems, 2011.
  60. 60.C. Ni, N. Charoenphakdee, J. Honda, and M. Sugiyama. On the calibration of multiclass classification with rejection. In Advances in Neural Information Processing Systems, pages 2582–2592, 2019.
  61. 61.N. Okati, A. De, and M. Rodriguez. Differentiable learning under triage. Advances in Neural Information Processing Systems, 34:9140–9151, 2021.
  62. 62.M. F. Pradier, J. Zazo, S. Parbhoo, R. H. Perlis, M. Zazzi, and F. Doshi-Velez. Preferential mixture-of-experts: Interpretable models that rely on human expertise as much as possible. AMIA Summits on Translational Science Proceedings, 2021:525, 2021.
  63. 63.M. Raghu, K. Blumer, G. Corrado, J. Kleinberg, Z. Obermeyer, and S. Mullainathan. The algorithmic automation problem: Prediction, triage, and human effort. arXiv preprint arXiv:1903.12220, 2019.
  64. 64.N. Raman and M. Yee. Improving learning-to-defer algorithms through fine-tuning. arXiv preprint arXiv:2112.10768, 2021.
  65. 65.H. G. Ramaswamy, A. Tewari, and S. Agarwal. Consistent algorithms for multiclass classification with an abstain option. Electronic Journal of Statistics, 12(1):530–554, 2018.
  66. 66.I. Steinwart. How to compare different loss functions and their risks. Constructive Approximation, 26(2):225–287, 2007.
  67. 67.E. Straitouri, A. Singla, V. B. Meresht, and M. Gomez-Rodriguez. Reinforcement learning under algorithmic triage. arXiv preprint arXiv:2109.11328, 2021.
  68. 68.E. Straitouri, L. Wang, N. Okati, and M. G. Rodriguez. Provably improving expert predictions with conformal prediction. arXiv preprint arXiv:2201.12006, 2022.
  69. 69.S. Tan, J. Adebayo, K. Inkpen, and E. Kamar. Investigating human+ machine complementarity for recidivism predictions. arXiv preprint arXiv:1808.09123, 2018.
  70. 70.R. Verma and E. Nalisnick. Calibrated learning to defer with one-vs-all classifiers. In International Conference on Machine Learning, pages 22184–22202, 2022.
  71. 71.R. Verma, D. Barrejón, and E. Nalisnick. Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles. In International Conference on Artificial Intelligence and Statistics, pages 11415–11434, 2023.
  72. 72.J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus. Emergent abilities of large language models. CoRR, abs/2206.07682, 2022.
  73. 73.Y. Wiener and R. El-Yaniv. Agnostic selective classification. In Advances in neural information processing systems, 2011.
  74. 74.B. Wilder, E. Horvitz, and E. Kamar. Learning to complement humans. In International Joint Conferences on Artificial Intelligence, pages 1526–1533, 2021.
  75. 75.M. Yuan and M. Wegkamp. Classification methods with reject option based on convex risk minimization. Journal of Machine Learning Research, 11(1), 2010.
  76. 76.M. Yuan and M. Wegkamp. SVMs with a reject option. In Bernoulli, 2011.
  77. 77.M. Zhang and S. Agarwal. Bayes consistency vs. H-consistency: The interplay between surrogate loss functions and the scoring function class. In Advances in Neural Information Processing Systems, 2020.
  78. 78.T. Zhang. Statistical behavior and consistency of classification methods based on convex risk minimization. The Annals of Statistics, 32(1):56–85, 2004.
  79. 79.J. Zhao, M. Agrawal, P. Razavi, and D. Sontag. Directing human attention in event localization for clinical timeline creation. In Machine Learning for Healthcare Conference, pages 80–102, 2021.
  80. 80.C. Zheng, G. Wu, F. Bao, Y. Cao, C. Li, and J. Zhu. Revisiting discriminative vs. generative classifiers: Theory and implications. arXiv preprint arXiv:2302.02334, 2023.
  81. 81.L. Ziyin, Z. Wang, P. P. Liang, R. Salakhutdinov, L.-P. Morency, and M. Ueda. Deep gamblers: Learning to abstain with portfolio theory. arXiv preprint arXiv:1907.00208, 2019.

Citation

MLA
Mao, A., et al. “Two-Stage Learning to Defer with Multiple Experts”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 3578–606, https://proceedings.neurips.cc/paper_files/paper/2023/file/0b17d256cf1fe1cc084922a8c6b565b7-Paper-Conference.pdf.
APA
Mao, A., Mohri, C., Mohri, M., & Zhong, Y. (2023). Two-Stage Learning to Defer with Multiple Experts. Advances in Neural Information Processing Systems, 36, 3578–3606. https://proceedings.neurips.cc/paper_files/paper/2023/file/0b17d256cf1fe1cc084922a8c6b565b7-Paper-Conference.pdf
Chicago
Mao, A., C. Mohri, M. Mohri, and Y. Zhong. 2023. “Two-Stage Learning to Defer with Multiple Experts”. Advances in Neural Information Processing Systems 36: 3578–3606. https://proceedings.neurips.cc/paper_files/paper/2023/file/0b17d256cf1fe1cc084922a8c6b565b7-Paper-Conference.pdf.
Harvard
Mao, A. et al. (2023) “Two-Stage Learning to Defer with Multiple Experts”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 3578–3606. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/0b17d256cf1fe1cc084922a8c6b565b7-Paper-Conference.pdf.
Vancouver
1. Mao A, Mohri C, Mohri M, Zhong Y (2023) Two-Stage Learning to Defer with Multiple Experts. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 3578–3606

BibTeX

@inproceedings{mao2023two,
  title = {Two-Stage Learning to Defer with Multiple Experts},
  author = {Mao, Anqi and Mohri, Christopher and Mohri, Mehryar and Zhong, Yutao},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {3578-3606},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/0b17d256cf1fe1cc084922a8c6b565b7-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors