Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Eyke HüllermeierWillem Waegeman

article2019Machine-mediated learning2,248 citations

Clarifies the critical distinction between irreducible data randomness and reducible model ignorance, providing a comprehensive framework for quantifying both aleatoric and epistemic uncertainty to build safer, more reliable machine learning systems.

Listen

As machine learning models are increasingly deployed in safety-critical domains such as medical diagnosis and autonomous driving, trusting their predictions becomes vital. Standard machine learning methods typically output single probability scores that can give a false sense of certainty, failing to reflect when an algorithm is operating outside its expertise or lacking data. When models fail without warning, the risks to human safety, operational performance, and institutional trust are substantial.

The article provides a conceptual and methodological overview of uncertainty handling in supervised machine learning, demonstrating why and how to separate uncertainty into two primary forms: aleatoric uncertainty, which stems from inherent randomness in the data-generating process and cannot be reduced by collecting more data, and epistemic uncertainty, which stems from a lack of knowledge or limited training data and can be reduced with additional information.

Through a comprehensive literature review and comparative analysis, the article evaluates classical frequentist approaches, Bayesian techniques, ensemble methods, and set-based frameworks like imprecise probabilities and conformal prediction. The authors examine how these various paradigms represent uncertainty at the level of models, parameters, and individual case-by-case predictions.

The article establishes several key findings. First, conventional single-distribution probabilistic modeling conflates aleatoric and epistemic factors because Bayesian model averaging washes out the learner's state of ignorance. Second, representing uncertainty using sets or sets of distributions, rather than single probability values, offers a more natural way to model severe data deficits without making arbitrary prior assumptions. Third, deep neural networks and complex models often exhibit high confidence even when making severe errors on unfamiliar or adversarial inputs, but capturing parameter variance via Bayesian neural networks or ensemble disagreement effectively surfaces reducible epistemic uncertainty. Fourth, set-valued predictions and classification with reject options provide practical mechanisms to manage risk by allowing a model to abstain or output multiple candidate answers when either aleatoric conflict or epistemic ignorance is too high.

These findings have immediate implications for system safety, compliance, and cost. Recognizing epistemic uncertainty enables automated systems to identify novel, out-of-distribution scenarios and delegate decisions to human operators, directly mitigating the risk of silent, catastrophic failures. Distinguishing between the two forms of uncertainty also prevents wasted expenditures on gathering more data when uncertainty is inherently irreducible, directing data collection budgets only toward areas where epistemic uncertainty can genuinely be reduced.

For future development, organizations should transition from purely point-based or single-probability outputs to uncertainty-aware architectures, such as deep ensembles, Bayesian extensions, or conformal predictors with statistically guaranteed error bounds. When uncertain, systems should adopt set-based predictions or selective abstention policies based on utility maximization. Further methodological research is needed to establish a mathematically rigorous, axiomatic foundation for decomposing total uncertainty and to standardize empirical evaluation protocols, such as accuracy-rejection benchmarks.

Readers should note that the field remains highly active and unsettled, with many proposed uncertainty decomposition methods relying on heuristic approximations. In addition, current approaches generally assume a fixed, correctly specified model space, which limits their ability to capture broader model misspecification or sudden real-world environment shifts. Practical adoption must therefore proceed with careful validation tailored to specific deployment settings.

arXiv: 1910.09457
Cover for Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Abstract

The notion of uncertainty is of major importance in machine learning and constitutes a key element of machine learning methodology. In line with the statistical tradition, uncertainty has long been perceived as almost synonymous with standard probability and probabilistic predictions. Yet, due to the steadily increasing relevance of machine learning for practical applications and related issues such as safety requirements, new problems and challenges have recently been identified by machine learning scholars, and these problems may call for new methodological developments. In particular, this includes the importance of distinguishing between (at least) two different types of uncertainty, often referred to as aleatoric and epistemic. In this paper, we provide an introduction to the topic of uncertainty in machine learning as well as an overview of attempts so far at handling uncertainty in general and formalizing this distinction in particular.

Table of Contents

  • 1 Introduction
  • 2 Sources of uncertainty in supervised learning
  • 2.1 Supervised learning and predictive uncertainty
  • 2.2 Sources of uncertainty
  • 2.3 Reducible versus irreducible uncertainty
  • 2.4 Approximation and model uncertainty
  • 3 Modeling approximation uncertainty: Set-based versus distributional representations
  • 3.1 Version space learning
  • 3.2 Bayesian inference
  • 3.3 Representing a lack of knowledge
  • 4 Machine learning methods for representing uncertainty
  • 4.1 Probability estimation via scoring, calibration, and ensembling
  • 4.2 Maximum likelihood estimation and Fisher information
  • 4.3 Generative models
  • 4.4 Gaussian processes
  • 4.5 Deep neural networks
  • 4.6 Credal sets and classifiers
  • 4.6.1 Uncertainty measures for credal sets
  • 4.6.2 Set-valued prediction
  • 4.7 Reliable classification
  • 4.7.1 Modeling the plausibility of predictions
  • 4.7.2 From plausibility to aleatoric and epistemic uncertainty
  • 4.8 Conformal prediction
  • 4.9 Set-valued prediction based on utility maximization
  • 5 Discussion and conclusion
  • A Background on uncertainty modeling
  • A.1 Sets versus distributions
  • A.2 Representation of ignorance
  • A.3 Sets of distributions
  • A.4 Distributions of sets
  • B Max-min versus sum-product aggregation
  • References

Knowls

  1. Knowl 1 — Supervised Learning Decomposition of Uncertainty: Aleatoric, Approximation, and Model Uncertainty

    definition

    In supervised learning, predictive uncertainty regarding a query instance xq∈Xx_q \in \mathcal{X} decomposes into three distinct sources:

    1. Aleatoric (statistical/irreducible) uncertainty: The variability in outcome y∈Yy \in \mathcal{Y} arising from the non-deterministic nature of the data-generating dependency, characterized by the conditional distribution p(y∣xq)=p(xq,y)/p(xq)p(y | x_q) = p(x_q, y)/p(x_q). Even with complete knowledge of the underlying distribution, point predictions generated by the pointwise Bayes predictor: f∗(x)=arg⁡min⁡y^∈Y∫Yℓ(y,y^)dP(y∣x)f^*(x) = \arg\min_{\hat{y} \in \mathcal{Y}} \int_{\mathcal{Y}} \ell(y, \hat{y}) dP(y | x) retain irreducible uncertainty.

    2. Model uncertainty: Epistemic discrepancy between the best possible hypothesis within a chosen hypothesis space H\mathcal{H}, defined as h∗=arg⁡min⁡h∈HR(h)=arg⁡min⁡h∈H∫X×Yℓ(h(x),y)dP(x,y)h^* = \arg\min_{h \in \mathcal{H}} R(h) = \arg\min_{h \in \mathcal{H}} \int_{\mathcal{X} \times \mathcal{Y}} \ell(h(x), y) dP(x, y), and the true pointwise Bayes predictor f∗f^*. It reflects potential misspecification of H\mathcal{H}.

    3. Approximation uncertainty: Epistemic discrepancy between the empirical risk minimizer h^=arg⁡min⁡h∈H1N∑i=1Nℓ(h(xi),yi)\hat{h} = \arg\min_{h \in \mathcal{H}} \frac{1}{N} \sum_{i=1}^N \ell(h(x_i), y_i) learned on a finite training set D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N and the optimal hypothesis h∗h^*. This reducible uncertainty decreases toward zero as sample size N→∞N \to \infty for a consistent learner.

  2. Knowl 2 — Information-Theoretic Uncertainty Decomposition in Bayesian Neural Networks

    model/method

    In Bayesian neural networks where network weights ww follow a posterior distribution p(w∣D)p(w | \mathcal{D}) given data D\mathcal{D}, the total predictive uncertainty for a discrete outcome y∈Yy \in \mathcal{Y} given query xx is quantified by the Shannon entropy of the predictive posterior distribution p(y∣x)=∫p(y∣w,x)p(w∣D)dwp(y | x) = \int p(y | w, x) p(w | \mathcal{D}) dw:

    H[p(y∣x)]=−∑y∈Yp(y∣x)log⁡2p(y∣x)H[p(y | x)] = -\sum_{y \in \mathcal{Y}} p(y | x) \log_2 p(y | x)

    Conditioning on fixed weights ww eliminates epistemic parameter uncertainty. The expected aleatoric uncertainty is the posterior expectation over weight configurations:

    Ep(w∣D)[H[p(y∣w,x)]]=−∫p(w∣D)(∑y∈Yp(y∣w,x)log⁡2p(y∣w,x))dw\mathbb{E}_{p(w | \mathcal{D})}[H[p(y | w, x)]] = -\int p(w | \mathcal{D}) \left( \sum_{y \in \mathcal{Y}} p(y | w, x) \log_2 p(y | w, x) \right) dw

    Epistemic uncertainty ue(x)u_e(x) is defined as the difference between total uncertainty and expected aleatoric uncertainty, which equals the mutual information I(y,w)I(y, w) between the outcome yy and the parameters ww:

    ue(x)=H[p(y∣x)]−Ep(w∣D)[H[p(y∣w,x)]]=I(y,w)u_e(x) = H[p(y | x)] - \mathbb{E}_{p(w | \mathcal{D})}[H[p(y | w, x)]] = I(y, w)

    This quantity measures the information that observing the true outcome yy would provide about the parameters ww.

  3. Knowl 3 — Ensemble-Based Approximation of Aleatoric and Epistemic Uncertainty

    model/method

    Given an ensemble of MM probabilistic predictors {h1,…,hM}\{h_1, \dots, h_M\} approximating a posterior distribution p(h∣D)p(h | \mathcal{D}) over hypothesis space H\mathcal{H}, predictive uncertainty for a query xx in a classification problem with outcome space Y\mathcal{Y} is decomposed into aleatoric and epistemic components using ensemble statistics:

    • Aleatoric uncertainty ua(x)u_a(x) is the average entropy of the individual ensemble predictions: ua(x)=−1M∑i=1M∑y∈Yp(y∣hi,x)log⁡2p(y∣hi,x)u_a(x) = -\frac{1}{M} \sum_{i=1}^M \sum_{y \in \mathcal{Y}} p(y | h_i, x) \log_2 p(y | h_i, x)

    • Total uncertainty ut(x)u_t(x) is the entropy of the averaged ensemble prediction: ut(x)=−∑y∈Y(1M∑i=1Mp(y∣hi,x))log⁡2(1M∑i=1Mp(y∣hi,x))u_t(x) = -\sum_{y \in \mathcal{Y}} \left( \frac{1}{M} \sum_{i=1}^M p(y | h_i, x) \right) \log_2 \left( \frac{1}{M} \sum_{i=1}^M p(y | h_i, x) \right)

    • Epistemic uncertainty ue(x)u_e(x) is computed as the difference: ue(x)=ut(x)−ua(x)u_e(x) = u_t(x) - u_a(x)

    This epistemic uncertainty ue(x)u_e(x) is mathematically equivalent to the Jensen-Shannon divergence among the ensemble members' predicted distributions {p(⋅∣hi,x)}i=1M\{p(\cdot | h_i, x)\}_{i=1}^M.

  4. Knowl 4 — Plausibility-Based Reliable Binary Classification Framework

    model/method

    In binary classification with classes Y={−1,+1}\mathcal{Y} = \{-1, +1\}, epistemic plausibility over hypotheses h∈Hh \in \mathcal{H} is modeled by the normalized likelihood:

    πH(h)=L(h)sup⁡h′∈HL(h′)=L(h)L(hml)\pi_{\mathcal{H}}(h) = \frac{L(h)}{\sup_{h' \in \mathcal{H}} L(h')} = \frac{L(h)}{L(h_{ml})}

    where L(h)L(h) is the likelihood of hypothesis hh on data D\mathcal{D} and hmlh_{ml} is the maximum likelihood estimator.

    For a query instance xqx_q, a probabilistic classifier hh predicting probability h(xq)=p(+1∣xq)∈[0,1]h(x_q) = p(+1 | x_q) \in [0, 1] yields class support degrees:

    π(+1∣h,xq)=max⁡(2h(xq)−1,0)\pi(+1 | h, x_q) = \max(2h(x_q) - 1, 0) π(−1∣h,xq)=max⁡(1−2h(xq),0)\pi(-1 | h, x_q) = \max(1 - 2h(x_q), 0)

    The total plausibility of each candidate class is obtained via max-min aggregation (corresponding to a Sugeno integral with respect to the possibility measure ΠH\Pi_{\mathcal{H}} on H\mathcal{H}):

    π(+1∣xq)=sup⁡h∈Hmin⁡(πH(h),π(+1∣h,xq))\pi(+1 | x_q) = \sup_{h \in \mathcal{H}} \min\left(\pi_{\mathcal{H}}(h), \pi(+1 | h, x_q)\right) π(−1∣xq)=sup⁡h∈Hmin⁡(πH(h),π(−1∣h,xq))\pi(-1 | x_q) = \sup_{h \in \mathcal{H}} \min\left(\pi_{\mathcal{H}}(h), \pi(-1 | h, x_q)\right)

  5. Knowl 5 — Quantification of Epistemic and Aleatoric Uncertainty from Plausibility Scores

    equation

    Given positive and negative class plausibilities π(+1)=π(+1∣xq)\pi(+1) = \pi(+1 | x_q) and π(−1)=π(−1∣xq)\pi(-1) = \pi(-1 | x_q) for a query xqx_q in binary classification, epistemic uncertainty ueu_e and aleatoric uncertainty uau_a are defined as:

    ue=min⁡(π(+1),π(−1))u_e = \min\left(\pi(+1), \pi(-1)\right) ua=1−max⁡(π(+1),π(−1))u_a = 1 - \max\left(\pi(+1), \pi(-1)\right)

    These measures satisfy the property ua+ue≤1u_a + u_e \le 1. Four boundary cases characterize this representation:

    • Full epistemic uncertainty (ue=1,ua=0u_e = 1, u_a = 0): π(+1)=π(−1)=1\pi(+1) = \pi(-1) = 1, which occurs when there exist at least two maximally plausible hypotheses (L(h)=L(hml)L(h) = L(h_{ml})) that completely support +1+1 and −1-1, respectively.
    • No epistemic uncertainty (ue=0u_e = 0): either π(+1)=0\pi(+1) = 0 or π(−1)=0\pi(-1) = 0. When all plausible hypotheses agree on p(+1∣xq)=αp(+1 | x_q) = \alpha and β=max⁡(α,1−α)\beta = \max(\alpha, 1 - \alpha), aleatoric uncertainty is ua=2(1−β)u_a = 2(1 - \beta).
    • Full aleatoric uncertainty (ua=1,ue=0u_a = 1, u_e = 0): occurs when β=1/2\beta = 1/2 (all plausible hypotheses assign probability 1/21/2 to both classes).
    • No uncertainty (ua=0,ue=0u_a = 0, u_e = 0): occurs when β=1\beta = 1 (all plausible hypotheses assign probability 11 or 00 to the positive class).
  6. Knowl 6 — Credal Uncertainty Decomposition via Upper Entropy and Generalized Hartley Measure

    theoretical result

    Let QQ be a credal set (a convex set of probability distributions) on a finite outcome space Y\mathcal{Y}. Total uncertainty U(Q)U(Q) decomposes additively into aleatoric uncertainty AU(Q)AU(Q) (conflict) and epistemic uncertainty EU(Q)EU(Q) (non-specificity):

    U(Q)=AU(Q)+EU(Q)U(Q) = AU(Q) + EU(Q)

    Epistemic uncertainty is quantified by the generalized Hartley measure:

    GH(Q)=∑A⊆YmQ(A)log⁡(∣A∣)GH(Q) = \sum_{A \subseteq \mathcal{Y}} m_Q(A) \log(|A|)

    where mQ:2Y→[0,1]m_Q: 2^{\mathcal{Y}} \to [0, 1] is the Möbius inverse of the lower probability capacity function νQ(A)=inf⁡q∈Qq(A)\nu_Q(A) = \inf_{q \in Q} q(A), given by mQ(A)=∑B⊆A(−1)∣A∖B∣νQ(B)m_Q(A) = \sum_{B \subseteq A} (-1)^{|A \setminus B|} \nu_Q(B).

    Total uncertainty is evaluated by the upper Shannon entropy H∗(Q)=max⁡q∈QH(q)H^*(Q) = \max_{q \in Q} H(q), yielding the disaggregation:

    H∗(Q)=(H∗(Q)−GH(Q))+GH(Q)H^*(Q) = \left( H^*(Q) - GH(Q) \right) + GH(Q)

    Alternatively, using lower entropy H∗(Q)=min⁡q∈QH(q)H_*(Q) = \min_{q \in Q} H(q):

    H∗(Q)=H∗(Q)+(H∗(Q)−H∗(Q))H^*(Q) = H_*(Q) + \left( H^*(Q) - H_*(Q) \right)

    where H∗(Q)H_*(Q) represents conflict (aleatoric uncertainty) and H∗(Q)−H∗(Q)H^*(Q) - H_*(Q) represents non-specificity (epistemic uncertainty).

  7. Knowl 7 — Dominance and Indeterminate Predictions in Credal Classification

    model/method

    In credal classification, knowledge about hypotheses is represented by a credal set C⊆HC \subseteq \mathcal{H}. For a query xqx_q, candidate outcomes are assessed by the concept of dominance: outcome y∈Yy \in \mathcal{Y} dominates y′∈Yy' \in \mathcal{Y} if yy is strictly more probable than y′y' under all distributions in the credal set:

    γ(y,y′,xq)=inf⁡h∈Cp(y∣xq,h)p(y′∣xq,h)>1\gamma(y, y', x_q) = \inf_{h \in C} \frac{p(y | x_q, h)}{p(y' | x_q, h)} > 1

    The set-valued prediction consists of all non-dominated classes Y^⊆Y\hat{Y} \subseteq \mathcal{Y}. In binary classification Y={−1,+1}\mathcal{Y} = \{-1, +1\}, interval-valued probabilities [p‾(y∣xq),pˉ(y∣xq)][\underline{p}(y | x_q), \bar{p}(y | x_q)] are computed with p‾(y∣xq)=inf⁡h∈Cp(y∣xq,h)\underline{p}(y | x_q) = \inf_{h \in C} p(y | x_q, h) and pˉ(y∣xq)=sup⁡h∈Cp(y∣xq,h)\bar{p}(y | x_q) = \sup_{h \in C} p(y | x_q, h).

    An uncertainty sampling score s(xq)s(x_q) based on pairwise dominance is defined as:

    s(xq)=−max⁡(γ(+1,−1,xq),γ(−1,+1,xq))s(x_q) = -\max\left( \gamma(+1, -1, x_q), \gamma(-1, +1, x_q) \right)

    This score increases with both the width of the probability interval pˉ(y∣xq)−p‾(y∣xq)\bar{p}(y | x_q) - \underline{p}(y | x_q) (epistemic uncertainty) and the closeness of the interval midpoint to 1/21/2 (aleatoric uncertainty).

  8. Knowl 8 — Set-Valued Prediction via Expected Utility Maximization

    model/method

    In multiclass classification with KK classes Y={y1,…,yK}\mathcal{Y} = \{y_1, \dots, y_K\}, set-valued prediction selects a non-empty prediction set Y^⊆Y\hat{Y} \subseteq \mathcal{Y} by maximizing the posterior expected set-based utility:

    Y^u∗(xq)=arg⁡max⁡Y^∈2Y∖{∅}∑y∈Yu(y,Y^)p(y∣xq)\hat{Y}^*_u(x_q) = \arg\max_{\hat{Y} \in 2^{\mathcal{Y}} \setminus \{\emptyset\}} \sum_{y \in \mathcal{Y}} u(y, \hat{Y}) p(y | x_q)

    A standard family of utility functions u:Y×(2Y∖{∅})→[0,1]u: \mathcal{Y} \times (2^{\mathcal{Y}} \setminus \{\emptyset\}) \to [0, 1] is defined by a sequence (g(1),…,g(K))∈[0,1]K(g(1), \dots, g(K)) \in [0, 1]^K:

    u(y,Y^)={0if y∉Y^g(∣Y^∣)if y∈Y^u(y, \hat{Y}) = \begin{cases} 0 & \text{if } y \notin \hat{Y} \\ g(|\hat{Y}|) & \text{if } y \in \hat{Y} \end{cases}

    To ensure proper decision behavior under uncertainty, gg must satisfy:

    1. g(1)=1g(1) = 1 (maximal utility for a correct singleton prediction).
    2. g(s)g(s) is non-increasing with set size s=∣Y^∣s = |\hat{Y}|.
    3. g(s)≥1/sg(s) \ge 1/s (risk-aversion property ensuring abstaining into a set of size ss yields higher utility than randomly guessing among the ss classes).

    Examples include discounted accuracy gδ,γ(s)=δs−γs2g_{\delta, \gamma}(s) = \frac{\delta}{s} - \frac{\gamma}{s^2}, exponential utility gexp⁡(s)=1−exp⁡(−δ/s)g_{\exp}(s) = 1 - \exp(-\delta / s), and logarithmic utility glog⁡(s)=log⁡(1+1/s)g_{\log}(s) = \log(1 + 1/s).

  9. Knowl 9 — Conformal Prediction and Confidence-Credibility Metrics

    model/method

    Given training examples (x1,y1),…,(xN,yN)(x_1, y_1), \dots, (x_N, y_N) and a query instance xN+1x_{N+1} under the assumption of exchangeability, conformal prediction performs transductive inference by testing candidate outcomes y∈Yy \in \mathcal{Y}. A nonconformity function f:X×Y→Rf: \mathcal{X} \times \mathcal{Y} \to \mathbb{R} produces scores αi=f(xi,yi)\alpha_i = f(x_i, y_i) for i∈{1,…,N}i \in \{1, \dots, N\} and αN+1=f(xN+1,y)\alpha_{N+1} = f(x_{N+1}, y). The pp-value associated with yy is:

    p(y)=#{i∈{1,…,N+1}∣αi≥αN+1}N+1p(y) = \frac{\#\{i \in \{1, \dots, N+1\} \mid \alpha_i \ge \alpha_{N+1}\}}{N + 1}

    For a significance level ϵ∈(0,1)\epsilon \in (0, 1), candidate yy is rejected if p(y)<ϵp(y) < \epsilon. The prediction region Yϵ={y∈Y∣p(y)≥ϵ}Y^\epsilon = \{y \in \mathcal{Y} \mid p(y) \ge \epsilon\} is guaranteed to satisfy marginal coverage P(yN+1∈Yϵ)≥1−ϵP(y_{N+1} \in Y^\epsilon) \ge 1 - \epsilon.

    From the sorted candidate pp-values p(1)≥p(2)≥⋯≥p(K)p_{(1)} \ge p_{(2)} \ge \dots \ge p_{(K)} for class outcomes Y={y1,…,yK}\mathcal{Y} = \{y_1, \dots, y_K\}, point prediction y^=arg⁡max⁡yp(y)\hat{y} = \arg\max_{y} p(y) is evaluated by:

    • Credibility: p(1)p_{(1)}, representing the largest significance level ϵ\epsilon for which YϵY^\epsilon is non-empty.
    • Confidence: 1−p(2)1 - p_{(2)}, representing the maximum confidence 1−ϵ1 - \epsilon for which YϵY^\epsilon remains a singleton set {y^}\{\hat{y}\}.
  10. Knowl 10 — Context-Dependence and Relativity of Aleatoric and Epistemic Uncertainty

    theoretical result

    Aleatoric and epistemic uncertainty are not absolute properties of a physical system, but are context-dependent and defined relative to the learning context (X,Y,H,P)(\mathcal{X}, \mathcal{Y}, \mathcal{H}, P):

    • Holding instance space X\mathcal{X}, output space Y\mathcal{Y}, hypothesis space H\mathcal{H}, and data distribution PP fixed, additional training instances decrease epistemic (approximation) uncertainty without altering aleatoric uncertainty.
    • If the learner is allowed to augment the instance space X\mathcal{X} with additional informative features to form a higher-dimensional space X′\mathcal{X}', overlapping class-conditional distributions in X\mathcal{X} can become completely separable in X′\mathcal{X}'.

    This augmentation converts previously irreducible aleatoric uncertainty into epistemic uncertainty: the class overlap is eliminated, but fitting the model in a higher-dimensional space requires more training data and increases approximation difficulty.

  11. Knowl 11 — Maxitive Possibility Measures vs Additive Probability Measures for Ignorance Representation

    theoretical result

    Probability measures P:2Ω→[0,1]P: 2^\Omega \to [0, 1] on a reference frame Ω\Omega are additive on disjoint sets A,B⊆ΩA, B \subseteq \Omega (P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B)) and normalized by P(Ω)=1P(\Omega) = 1. Consequently, probability distributions cannot represent complete ignorance without invoking the principle of indifference (e.g., p(ω)=1/∣Ω∣p(\omega) = 1/|\Omega|), which conflates epistemic ignorance with precise stochastic equiprobability and is not invariant under reparametrization.

    In contrast, possibility measures Π:2Ω→[0,1]\Pi: 2^\Omega \to [0, 1] generated from a possibility distribution π:Ω→[0,1]\pi: \Omega \to [0, 1] via Π(A)=sup⁡ω∈Aπ(ω)\Pi(A) = \sup_{\omega \in A} \pi(\omega) obey maxitivity:

    Π(A∪B)=max⁡(Π(A),Π(B))\Pi(A \cup B) = \max(\Pi(A), \Pi(B))

    Together with the dual necessity measure N(A)=1−Π(Aˉ)N(A) = 1 - \Pi(\bar{A}), complete ignorance is naturally represented by Π(A)=1\Pi(A) = 1 and N(A)=0N(A) = 0 for all non-empty subsets A⊂ΩA \subset \Omega, allowing an event AA to be declared fully plausible without forcing its complement Aˉ\bar{A} to be implausible.

Coverage note — Standard textbook reviews of parameter estimation via maximum likelihood/Fisher information in classical statistics, standard Gaussian process regression derivations, and general post-hoc calibration methods (e.g., Platt scaling, isotonic regression) were deliberately omitted as they represent established background rather than specific conceptual contributions of this paper.

References

  1. 1.Abellan, J., & Moral, S. (2000). A non-specificity measure for convex sets of probability distributions. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 8, 357–367.
  2. 2.Abellan, J., Klir, J., & Moral, S. (2006). Disaggregated total uncertainty measure for credal sets. International Journal of General Systems, 35(1), 29–44.
  3. 3.Aggarwal, C., Kong, X., Gu, Q., Han, J., & Yu, P. (2014). Active learning: A survey. In Data Classification: Algorithms and Applications, 571–606.
  4. 4.Antonucci, A., Corani, G., & Gabaglio, S. (2012). Active learning by the naive credal classifier. In Proceedings of the sixth European workshop on probabilistic graphical models (PGM), pp. 3–10.
  5. 5.Balasubramanian, V., Ho, S., & Vovk, V. (Eds.). (2014). Conformal prediction for reliable machine learning: Theory, adaptations and applications. Burlington: Morgan Kaufmann.
  6. 6.Barber, R., Candes, E., Ramdas, A., & Tibshirani, R. (2020). The limits of distribution-free conditional predictive inference. CoRR. (arXiv:abs/1903.04684v2).
  7. 7.Bazargami, M., & Mac-Namee, B. (2019). The elliptical basis function data descriptor network: A one-class classification approach for anomaly detection. In European conference on machine learning and knowledge discovery in databases.
  8. 8.Bernardo, J. (1979). Reference posterior distributions for Bayesian inference. Journal of the Royal Statistical Society, Series B (Methodological), 41(2), 113–147.
  9. 9.Bernardo, J. (2005). An introduction to the imprecise Dirichlet model for multinomial data. International Journal of Approximate Reasoning, 39(2–3), 123–150.
  10. 10.Bi, W., & Kwok, J. (2015). Bayes-optimal hierarchical multilabel classification. IEEE Transactions on Knowledge and Data Engineering, 27, 1–1. https://doi.org/10.1109/TKDE.2015.2441707.
  11. 11.Blum, M., & Riedmiller, M. (2013). Optimization of Gaussian process hyperparameters using Rprop. In Proceedings of ESANN, 21st European symposium on artificial neural networks, Bruges, Belgium.
  12. 12.Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.
  13. 13.Cattaneo, M. (2005). Likelihood-based statistical decisions. In Proceedings of 4th international symposium on imprecise probabilities and their applications, pp. 107–116.
  14. 14.Chow, C. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory IT, 16, 41–46.
  15. 15.Corani, G., & Zaffalon, M. (2008). Learning reliable classifiers from small or incomplete data sets: The naive credal classifier 2. Journal of Machine Learning Research, 9, 581–621.
  16. 16.Corani, G., & Zaffalon, M. (2009). Lazy naive credal classifier. In Proceedings of the 1st ACM SIGKDD workshop on knowledge discovery from uncertain data, ACM, New York, NY, USA, U ’09, pp. 30–37, https://doi.org/10.1145/1610555.1610560.
  17. 17.Cozman, F. (2000). Credal networks. Artificial Intelligence, 120(2), 199–233.
  18. 18.Csiszár, I. (2008). Axiomatic characterizations of information measures. Entropy, 10, 261–273.
  19. 19.Del Coz, J., Díez, J., & Bahamonde, A. (2009). Learning nondeterministic classifiers. The Journal of Machine Learning Research, 10, 2273–2293.
  20. 20.Deng, Y. (2014). Generalized evidence theory. CoRRarXiv:abs/404.4801.v1.
  21. 21.Denker, J., & LeCun, Y. (1991). Transforming neural-net output levels to probability distributions. In Proceedings of NIPS, advances in neural information processing systems.
  22. 22.Denoeux, T. (2014). Likelihood-based belief function: Justification and some extensions to low-quality data. International Journal of Approximate Reasoning, 55(7), 1535–1547.
  23. 23.Depeweg, S., Hernandez-Lobato, J., Doshi-Velez, F., & Udluft, S. (2018). Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning. In Proceedings of ICML, 35th international conference on machine learning, Stockholm, Sweden.
  24. 24.Der Kiureghian, A., & Ditlevsen, O. (2009). Aleatory or epistemic? Does it matter? Structural Safety, 31, 105–112.
  25. 25.Destercke, S., Dubois, D., & Chojnacki, E. (2008). Unifying practical uncertainty representations: I. Generalized p-boxes. International Journal of Approximate Reasoning, 49, 649–663.
  26. 26.DeVries, T., & Taylor, G. (2018). Learning confidence for out-of-distribution detection in neural networks. CoRRarXiv:abs/1802.04865.
  27. 27.Dubois, D. (2006). Possibility theory and statistical reasoning. Computational Statistics and Data Analysis, 51(1), 47–69.
  28. 28.Dubois, D., & Hüllermeier, E. (2007). Comparing probability measures using possibility theory: A notion of relative peakedness. International Journal of Approximate Reasoning, 45(2), 364–385.
  29. 29.Dubois, D., & Prade, H. (1988). Possibility theory. Berlin: Plenum Press.
  30. 30.Dubois, D., Prade, H., & Smets, P. (1996). Representing partial ignorance. IEEE Transactions on Systems, Man and Cybernetics, Series A, 26(3), 361–377.
  31. 31.Dubois, D., Moral, S., & Prade, H. (1997). A semantics for possibility theory based on likelihoods. Journal of Mathematical Analysis and Applications, 205(2), 359–380.
  32. 32.Endres, D., & Schindelin, J. (2003). A new metric for probability distributions. IEEE Transactions on Information Theory, 49(7), 1858–1860.
  33. 33.Flach, P. (2017). Classifier calibration. In Encyclopedia of machine learning and data mining. Berlin: Springer, pp. 210–217.
  34. 34.Freitas, A. A. (2007). A tutorial on hierarchical classification with applications in bioinformatics. In Research and trends in data mining technologies and applications, pp. 175–208.
  35. 35.Frieden, B. (2004). Science from Fisher information: A unification. Cambridge: Cambridge University Press.
  36. 36.Gal, Y., & Ghahramani, Z. (2016). Bayesian convolutional neural networks with Bernoulli approximate variational inference. In Proceedings of the ICLR workshop track.
  37. 37.Gama, J. (2012). A survey on learning from data streams: Current and future trends. Progress in Artificial Intelligence, 1(1), 45–55.
  38. 38.Gammerman, A., & Vovk, V. (2002). Prediction algorithms and confidence measures based on algorithmic randomness theory. Theoretical Computer Science, 287, 209–217.
  39. 39.Gneiting, T., & Raftery, A. (2005). Strictly proper scoring rules, prediction, and estimation. Tech. Rep. 463R, Department of Statistics, University of Washington.
  40. 40.Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. Cambridge, MA: The MIT Press.
  41. 41.Graves, A. (2011). Practical variational inference for neural networks. In Proceedings of NIPS, advances in neural information processing systems, pp. 2348–2356.
  42. 42.Hartley, R. (1928). Transmission of information. Bell Labs Technical Journal, 7(3), 535–563.
  43. 43.Hechtlinger, Y., Poczos, B., & Wasserman, L. (2019). Cautious deep learning. CoRRarXiv:abs/1805.09460 .v2.
  44. 44.Hellman, M. (1970). The nearest neighbor classification rule with a reject option. IEEE Transactions on Systems, Man and Cybernetics SMC, 6, 179–185.
  45. 45.Hendrycks, D., & Gimpel, K. (2017). A baseline for detecting misclassified and out-of-distribution examples in neural networks. In Proceedings of ICLR, international conference on learning representations.
  46. 46.Herbei, R., & Wegkamp, M. (2006). Classification with reject option. Canadian Journal of Statistics, 34(4), 709–721.
  47. 47.Hora, S. (1996). Aleatory and epistemic uncertainty in probability elicitation with an example from hazardous waste management. Reliability Engineering and System Safety, 54(2–3), 217–223.
  48. 48.Hühn, J., & Hüllermeier, E. (2009). FR3: A fuzzy rule learner for inducing reliable classifiers. IEEE Transactions on Fuzzy Systems, 17(1), 138–149.
  49. 49.Hüllermeier, E., & Brinker, K. (2008). Learning valued preference structures for solving classification problems. Fuzzy Sets and Systems, 159(18), 2337–2352.
  50. 50.Jeffreys, H. (1946). An invariant form for the prior probability in estimation problems. Proceedings of the Royal Society A, 186, 453–461.
  51. 51.Johansson, U., Löfström, T., Sundell, H., Linusson, H., Gidenstam, A., & Boström, H. (2018). Venn predictors for well-calibrated probability estimation trees. In Proceedings of COPA, 7th symposium on conformal and probabilistic prediction and applications, Maastricht, The Netherlands, pp. 3–14.
  52. 52.Jordan, M., Ghahramani, Z., Jaakkola, T., & Saul, L. (1999). An introduction to variational methods for graphical models. Machine Learning, 37(2), 183–233.
  53. 53.Kay, D. M. (1992). A practical Bayesian framework for backpropagation networks. NeuralComputation, 4(3), 448–472.
  54. 54.Kendall, A., & Gal, Y. (2017). What uncertainties do we need in Bayesian deep learning for computer vision? In Proceedings of NIPS, advances in neural information processing systems, pp. 5574–5584.
  55. 55.Khan, S. S., & Madden, M. G. (2014). One-class classification: Taxonomy of study and review of techniques. The Knowledge Engineering Review, 29(3), 345–374. https://doi.org/10.1017/S026988891300043X.
  56. 56.Klement, E., Mesiar, R., & Pap, E. (2002). Triangular norms. London: Kluwer Academic Publishers.
  57. 57.Klir, G. (1994). Measures of uncertainty in the Dempster–Shafer theory of evidence. In R. Yager, M. Fedrizzi, & J. Kacprzyk (Eds.), Advances in the Dempster–Shafer theory of evidence (pp. 35–49). New York: Wiley.
  58. 58.Klir, G., & Mariano, M. (1987). On the uniqueness of possibilistic measure of uncertainty and information. Fuzzy Sets and Systems, 24(2), 197–219.
  59. 59.Kolmogorov, A. (1965). Three approaches to the quantitative definition of information. Problems of Information Transmission, 1(1), 1–7.
  60. 60.Kruppa, J., Liu, Y., Biau, G., Kohler, M., König, I., Malley, J., et al. (2014). Probability estimation with machine learning methods for dichotomous and multi-category outcome: Theory. Biometrical Journal, 56(4), 534–563.
  61. 61.Kruse, R., Schwecke, E., & Heinsohn, J. (1991). Uncertainty and vagueness in knowledge based systems. Berlin: Springer.
  62. 62.Kull, M., & Flach, P. (2014). Reliability maps: A tool to enhance probability estimates and improve classification accuracy. In Proceedings of ECML/PKDD, European conference on machine learning and principles and practice of knowledge discovery in databases, Nancy, France, pp. 18–33.
  63. 63.Kull, M., de Menezes, T., Filho, S., & Flach, P. (2017). Beta calibration: A well-founded and easily implemented improvement on logistic calibration for binary classifiers. In Proceedings of AISTATS, 20th international conference on artificial intelligence and statistics, Fort Lauderdale, FL, USA, pp. 623–631.
  64. 64.Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of NeurIPS, 31st conference on neural information processing systems, Long Beach, California, USA.
  65. 65.Lambrou, A., Papadopoulos, H., & Gammerman, A. (2011). Reliable confidence measures for medical diagnosis with evolutionary algorithms. IEEE Transactions on Information Technology in Biomedicine, 15(1), 93–99.
  66. 66.Lassiter, D. (2020). Representing credal imprecision: from sets of measures to hierarchical Bayesian models. Philosophical Studies (Forthcoming).
  67. 67.Lee J, Bahri Y, Novak R, Schoenholz, S., Pennington, J., & Sohl-Dickstein, J. (2018a). Deep neural networks as Gaussian processes. In Proceedings of ICLR, international conference on learning representations.
  68. 68.Lee, K., Lee, K., Lee, H., & Shin, J. (2018b). A simple unified framework for detecting out-of-distribution samples and adversarial attacks. CoRR. (arXiv:abs/1807.03888.v2).
  69. 69.Liang, S., Li, Y., & Srikant, R. (2018). Enhancing the reliability of out-of-distribution image detection in neural networks. In Proceedings of ICLR, international conference on learning representations.
  70. 70.Linusson, H., Johansson, U., Boström, H., & Löfström, T. (2016). Reliable confidence predictions using conformal prediction. In Proceedings of PAKDD, 20th Pacific-Asia conference on knowledge discovery and data mining, Auckland, New Zealand.
  71. 71.Linusson, H., Johansson, U., Boström, H., & Löfström, T. (2018). Classification with reject option using conformal prediction. In Proceedings of PAKDD, 22nd Pacific-Asia conference on knowledge discovery and data mining, Melbourne, VIC, Australia.
  72. 72.Liu, F. T., Ting, K. M., & Hua Zhou, Z. (2009). Isolation forest. In Proceedings of ICDM 2008, Eighth IEEE international conference on data mining, IEEE Computer Society, pp. 413–422.
  73. 73.Maau, D. D., Cozman, F., Conaty, D., & de Campos, C. P. (2017). Credal sum-product networks. In PMLR: Proceedings of machine learning Research (ISIPTA 2017), vol 62, pp. 205–216.
  74. 74.Malinin, A., & Gales, M. (2018). Predictive uncertainty estimation via prior networks. In Proceedings of NeurIPS, 32nd conference on neural information processing systems, Montreal, Canada.
  75. 75.Matheron, G. (1975). Random sets and integral geometry. New York: Wiley.
  76. 76.Mitchell, T. (1977). Version spaces: A candidate elimination approach to rule learning. In Proceedings IJCAI-77, pp. 305–310.
  77. 77.Mitchell, T. (1980). The need for biases in learning generalizations. Tech. Rep. TR CBM–TR–117, Rutgers University.
  78. 78.Mobiny, A., Nguyen, H., Moulik, S., Garg, N., & Wu, C. (2017). DropConnect is effective in modeling uncertainty of Bayesian networks. CoRR. arXiv:abs/1906.04569.
  79. 79.Neal, R. (2012). Bayesian learning for neural networks (p. 118). Berlin: Springer.
  80. 80.Nguyen, H. (1978). On random sets and belief functions. Journal of Mathematical Analysis and Applications, 65, 531–542.
  81. 81.Nguyen, V., Destercke, S., Masson, M., & Hüllermeier, E. (2018). Reliable multi-class classification based on pairwise epistemic and aleatoric uncertainty. In Proceedings IJCAI 2018, 27th international joint conference on artificial intelligence (pp. 5089–5095). Sweden: Stockholm.
  82. 82.Nguyen, V., Destercke, S., Hüllermeier, E., & (2019). Epistemic uncertainty sampling. In Proceedings of DS 2019, 22nd international conference on discovery science, Split, Croatia.
  83. 83.Oh, S. (2017). Top-k hierarchical classification. In AAAI. AAAI Press, pp. 2450–2456.
  84. 84.Owhadi, H., Sullivan, T., McKerns, M., Ortiz, M., & Scovel, C. (2012). Optimal uncertainty quantification. CoRR. (arXiv:abs/1009.0679.v3).
  85. 85.Papadopoulos, H. (2008). Inductive conformal prediction: Theory and application to neural networks. Tools in Artificial Intelligence, 18(2), 315–330.
  86. 86.Papernot, N., & McDaniel, P. (2018). Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning. CoRRarXiv:abs/1803.04765v1.
  87. 87.Perello-Nieto, M., Filho, T. S., Kull, M., & Flach, P. (2016). Background check: A general technique to build more reliable and versatile classifiers. In Proceedings of ICDM, international conference on data mining.
  88. 88.Platt, J. (1999). Probabilistic outputs for support vector machines and comparison to regularized likelihood methods. In A. Smola, P. Bartlett, B. Schoelkopf, & D. Schuurmans (Eds.), Advances in large margin classifiers (pp. 61–74). Cambridge, MA: MIT Press.
  89. 89.Pukelsheim, F. (2006). Optimal design of experiments. Philadelphia: SIAM.
  90. 90.Ramaswamy, H. G., Tewari, A., & Agarwal, S. (2015). Consistent algorithms for multiclass classification with a reject option. CoRR. (arXiv:abs/1505.04137).
  91. 91.Rangwala, H., & Naik, A. (2017). Large scale hierarchical classification: Foundations, algorithms and applications. In The European conference on machine learning and principles and practice of knowledge discovery in databases.
  92. 92.Rényi, A. (1970). Probability theory. Amsterdam: North-Holland.
  93. 93.Sato, M., Suzuki, J., Shindo, H., & Matsumoto, Y. (2018). Interpretable adversarial perturbation in input embedding space for text. In Proceedings IJCAI 2018, 27th international joint conference on artificial intelligence (pp. 4323–4330). Sweden: Stockholm.
  94. 94.Seeger, M. (2004). Gaussian processes for machine learning. International Journal of Neural Systems, 14(2), 69–104.
  95. 95.Senge, R., Bösner, S., Dembczynski, K., Haasenritter, J., Hirsch, O., Donner-Banzhoff, N., et al. (2014). Reliable classification: Learning classifiers that distinguish aleatoric and epistemic uncertainty. Information Sciences, 255, 16–29.
  96. 96.Sensoy, M., Kaplan, L., & Kandemir, M. (2018). Evidential deep learning to quantify classification uncertainty. In Proceedings of NeurIPS, 32nd conference on neural information processing systems, Montreal, Canada.
  97. 97.Shafer, G. (1976). A mathematical theory of evidence. Princeton: Princeton University Press.
  98. 98.Shafer, G., & Vovk, V. (2008). A tutorial on conformal prediction. Journal of Machine Learning Research, 9, 371–421.
  99. 99.Shaker, M., & Hüllermeier, E. (2020). Aleatoric and epistemic uncertainty with random forests. In Proceedings of IDA 2020, 18th international symposium on intelligent data analysis, Springer, Konstanz, Germany, LNCS, vol 12080, pp. 444–456, https://doi.org/10.1007/978-3-030-44584-3_35.
  100. 100.Shilkret, N. (1971). Maxitive measure and integration. Nederl Akad Wetensch Proc Ser A 74 = Indag Math, 33, 109–116.
  101. 101.Smets, P., & Kennes, R. (1994). The transferable belief model. Artificial Intelligence, 66, 191–234.
  102. 102.Sondhi, J. F. A., Perry, J., & Simon, N. (2019). Selective prediction-set models with coverage guarantees. CoRRarXiv:abs/1906.05473.v1.
  103. 103.Sourati, J., Akcakaya, M., Erdogmus, D., Leen, T., & Dy, J. (2018). A probabilistic active learning algorithm based on Fisher information ratio. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(8), 2023–2029.
  104. 104.Sugeno, M. (1974). Theory of fuzzy integrals and its application. PhD thesis, Tokyo Institute of Technology.
  105. 105.Tan, M., & Le, Q. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of ICML, 36th international conference on machine learning, Long Beach, California.
  106. 106.Tax, D. M., & Duin, R. P. (2004). Support vector data description. Machine Learning, 54(1), 45–66. https://doi.org/10.1023/B:MACH.0000008084.60811.49.
  107. 107.Vapnik, V. (1998). Statistical learning theory. New York: Wiley.
  108. 108.Varshney, K. (2016). Engineering safety in machine learning. In Proceedings of information theory appllication workshop, La Jolla, CA.
  109. 109.Varshney, K., & Alemzadeh, H. (2016). On the safety of machine learning: Cyber-physical systems, decision sciences, and data products. CoRR. arXiv:abs/1610.01256.
  110. 110.Vovk, V., Gammerman, A., & Shafer, G. (2003). Algorithmic learning in a random world. Berlin: Springer.
  111. 111.Walley, P. (1991). Statistical reasoning with imprecise probabilities. London: Chapman and Hall.
  112. 112.Wasserman, L. (1990). Belief functions and statistical evidence. The Canadian Journal of Statistics, 18(3), 183–196.
  113. 113.Wolpert, D. (1996). The lack of a priori distinctions between learning algorithms. Neural Computation, 8(7), 1341–1390.
  114. 114.Yager, R. (1983). Entropy and specificity in a mathematical theory of evidence. International Journal of General Systems, 9, 249–260.
  115. 115.Yang, F., Wanga, H. Z., Mi, H., de Lin, C., & Cai, W. W. (2009). Using random forest for reliable classification and cost-sensitive learning for medical diagnosis. BMC Bioinformatics, 10, 118–131.
  116. 116.Yang, G., Destercke, S., & Masson, M. (2017a). Cautious classification with nested dichotomies and imprecise probabilities. Soft Computing, 21, 7447–7462.
  117. 117.Yang, G., Destercke, S., & Masson, M. (2017b). The costs of indeterminacy: How to determine them? IEEE Transactions on Cybernetics, 47, 4316–4327.
  118. 118.Zadrozny, B., & Elkan, C. (2001). Obtaining calibrated probability estimates from decision trees and Naive Bayesian classifiers. In Proceedings of ICML, international conference on machine learning, pp. 609–616.
  119. 119.Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of KDD–02, 8th international conference on knowledge discovery and data mining, Edmonton, Alberta, Canada, pp. 694–699.
  120. 120.Zaffalon, M. (2002). The naive credal classifier. Journal of Statistical Planning and Inference, 105(1), 5–21.
  121. 121.Zaffalon, M., Giorgio, C., & Deratani Mauá, D. (2012). Evaluating credal classifiers by utility-discounted predictive accuracy. The International Journal of Approximate Reasoning, 53, 1282–1301.
  122. 122.Ziyin, L., Wang, Z., Liang, P. P., Salakhutdinov, R., Morency, L. P., & Ueda, M. (2019). Deep gamblers: Learning to abstain with portfolio theory. arXiv:1907.00208.

Citation

MLA
Hüllermeier, E., and W. Waegeman. “Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods”. Machine Learning, vol. 110, no. 3, 2021, pp. 457–506, https://doi.org/10.1007/s10994-021-05946-3.
APA
Hüllermeier, E., & Waegeman, W. (2021). Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods. Machine Learning, 110(3), 457–506. https://doi.org/10.1007/s10994-021-05946-3
Chicago
Hüllermeier, E., and W. Waegeman. 2021. “Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods”. Machine Learning 110 (3): 457–506. https://doi.org/10.1007/s10994-021-05946-3.
Harvard
Hüllermeier, E. and Waegeman, W. (2021) “Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods”, Machine Learning, 110(3), pp. 457–506. Available at: https://doi.org/10.1007/s10994-021-05946-3.
Vancouver
1. Hüllermeier E, Waegeman W (2021) Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods. Machine Learning 110:457–506

BibTeX

@article{H_llermeier_2021, title={Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods}, volume={110}, ISSN={1573-0565}, url={http://dx.doi.org/10.1007/s10994-021-05946-3}, DOI={10.1007/s10994-021-05946-3}, number={3}, journal={Machine Learning}, publisher={Springer Science and Business Media LLC}, author={Hüllermeier, Eyke and Waegeman, Willem}, year={2021}, month=Mar, pages={457–506} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF