Mean Absolute Percentage Error for regression models

Arnaud de MyttenaereBoris GoldenBénédicte Le GrandFabrice Rossi

article2016Neurocomputing1,351 citations

Establishes theoretical foundations for Mean Absolute Percentage Error regression by proving the universal consistency of empirical risk minimization and demonstrating that training optimal MAPE models is equivalent to weighted Mean Absolute Error regression.

Listen

In forecasting and financial modeling, practitioners often evaluate predictive performance using the Mean Absolute Percentage Error (MAPE) rather than standard metrics like the Mean Squared Error (MSE) or Mean Absolute Error (MAE), due to its intuitive interpretation in relative percentage terms. However, while traditional regression methods rest on well-established theoretical foundations, the theoretical properties and optimization mechanics of directly training models to minimize MAPE have historically remained unexplored. This creates a disconnect between the metric used to train predictive algorithms and the business metric used to evaluate their success.

The article establishes the theoretical legitimacy and practical viability of using MAPE as a primary objective function in machine learning and regression. It demonstrates that optimal MAPE-based models exist, proves that standard empirical risk minimization reliably converges to the optimal solution, and formulates a practical method to train non-linear kernel regression models directly under this objective.

To conduct this evaluation, the analysis derived mathematical proofs linking MAPE model complexity, covering numbers, and statistical learning bounds to existing formulations for median and absolute error regression. It then adapted kernel quantile regression into a dual quadratic optimization problem incorporating instance weights inversely proportional to the target values. Finally, the researchers tested the approach on simulated benchmark datasets comprising 1,000 training points and 1,000 test points under varying degrees of vertical offset and noise to observe empirical behavior across different target ranges.

The findings show that finding the best model under MAPE is equivalent to conducting weighted median regression, where each observation is weighted inversely by its true value. Theoretically, the article proves the existence of a globally optimal MAPE regression function and establishes universal strong consistency for empirical risk minimization, provided the target variable remains bounded away from zero. Computationally, in dual formulations of kernel regression, minimizing MAPE automatically widens the optimization constraints for smaller target values, forcing the algorithm to fit low-magnitude targets much more accurately. In experimental simulations, models directly trained on MAPE outperformed standard median regression models, achieving substantial error reductions (such as decreasing test MAPE from roughly 188% down to 100% when target values hovered near zero) and naturally biasing predictions lower than the conditional median.

These results provide senior decision-makers and technical leaders with the rigorous justification needed to deploy MAPE-optimized algorithms in production environments, such as energy load forecasting, pricing expensive assets, or financial gain-loss projections. Relying on models trained directly for percentage accuracy eliminates performance loss caused by metric misalignment and protects against severe relative forecasting errors on low-value transactions. However, because MAPE penalizes relative over-predictions more heavily than under-predictions, managers should anticipate that optimal MAPE models will systematically produce conservative, downward-shifted forecasts.

Organizations should actively adopt weighted quantile regression or the presented kernel formulation whenever MAPE serves as the primary operational key performance indicator, provided the predicted values are strictly non-zero. For standard linear models, teams can readily implement this via standard weighted median regression solvers. Further research and pilot testing are recommended to extend theoretical convergence guarantees to regularized kernel estimators and to examine methods for handling datasets containing values arbitrarily close to zero.

arXiv: 1605.02541
  • Paper: Another look at measures of forecast accuracy, Rob J. Hyndman et al. (2006). It provides a foundational critique of percentage-based forecast accuracy metrics like MAPE, establishing the practical and mathematical failure modes that motivate formal theoretical analysis.
  • Paper: Stability and Generalization, Olivier Bousquet et al. (2002). It establishes the theoretical principles of algorithmic stability and generalization error bounds used in analyzing empirical risk minimization.
  • Paper: Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Peter L. Bartlett et al. (2002). It introduces Rademacher complexity and risk bound frameworks foundational for proving the universal consistency of empirical risk minimization estimators.
  • Paper: Support Vector Regression Machines, Harris Drucker et al. (1996). It lays the groundwork for non-parametric loss minimization and regularized kernel regression models.
Cover for Mean Absolute Percentage Error for regression models

Abstract

We study in this paper the consequences of using the Mean Absolute Percentage Error (MAPE) as a measure of quality for regression models. We prove the existence of an optimal MAPE model and we show the universal consistency of Empirical Risk Minimization based on the MAPE. We also show that finding the best model under the MAPE is equivalent to doing weighted Mean Absolute Error (MAE) regression, and we apply this weighting strategy to kernel regression. The behavior of the MAPE kernel regression is illustrated on simulated data.

Table of Contents

  • 1 Introduction
  • 2 General setting and notations
  • 3 Existence of the MAPE-regression function
  • 3.1 Finite values for the point-wise problem
  • 3.2 Existence of a solution for the point-wise problem
  • 3.3 Choosing the minimum
  • 4 Effects of the MAPE on complexity control
  • 4.1 Classes of functions
  • 4.2 Covering numbers
  • 4.2.1 Notations and definitions
  • 4.2.2 Supremum covering numbers
  • 4.2.3 LpL_{p} covering numbers
  • 4.3 VC-dimension
  • 4.4 Examples of Uniform Laws of Large Numbers
  • 5 Consistency and the MAPE
  • 6 MAPE kernel regression
  • 6.1 From quantile regression to MAPE regression
  • 6.1.1 Quantile regression
  • 6.2 MAPE primal problem
  • 6.2.1 MAPE dual problem
  • 6.2.2 Comparaison to the quantile regression
  • 6.3 A simulation study
  • 6.3.1 Generation of observations
  • 6.3.2 Results
  • 6.3.3 Graphical illustration
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Universal Consistency of Empirical Risk Minimization under MAPE

    theoretical result

    Let Z=(X,Y)Z = (X, Y) be a random pair taking values in Rd×R\mathbb{R}^d \times \mathbb{R} such that ∣Y∣≥YL>0|Y| \ge Y_L > 0 almost surely for a fixed constant YLY_L. Let Dn=(Xi,Yi)i=1nD_n = (X_i, Y_i)_{i=1}^n be nn independent and identically distributed copies of ZZ. Let (Gn)n≥1(\mathcal{G}_n)_{n \ge 1} be a sequence of classes of measurable functions from Rd\mathbb{R}^d to R\mathbb{R} satisfying:

    1. Gn⊂Gn+1\mathcal{G}_n \subset \mathcal{G}_{n+1} for all nn;
    2. ⋃n≥1Gn\bigcup_{n \ge 1} \mathcal{G}_n is dense in L1(μ)L^1(\mu) for any probability measure μ\mu on Rd\mathbb{R}^d;
    3. For all nn, the Vapnik-Chervonenkis dimension Vn=VCdim(H+(Gn,lMAPE))<∞V_n = \text{VCdim}(\mathcal{H}^+(\mathcal{G}_n, l_{MAPE})) < \infty, where H+(Gn,lMAPE)={(x,y,t)↦It≤∣g(x)−y∣∣y∣∣g∈Gn}\mathcal{H}^+(\mathcal{G}_n, l_{MAPE}) = \{ (x, y, t) \mapsto \mathbb{I}_{t \le \frac{|g(x) - y|}{|y|}} \mid g \in \mathcal{G}_n \};
    4. For all nn, ∥Gn∥∞=sup⁡g∈Gnsup⁡x∈Rd∣g(x)∣<∞\|\mathcal{G}_n\|_\infty = \sup_{g \in \mathcal{G}_n} \sup_{x \in \mathbb{R}^d} |g(x)| < \infty.

    If the class complexity satisfies the growth conditions:

    lim⁡n→∞Vn∥Gn∥∞2log⁡∥Gn∥∞n=0\lim_{n \to \infty} \frac{V_n \|\mathcal{G}_n\|_\infty^2 \log \|\mathcal{G}_n\|_\infty}{n} = 0

    and there exists δ>0\delta > 0 such that:

    lim⁡n→∞n1−δ∥Gn∥∞2=∞,\lim_{n \to \infty} \frac{n^{1-\delta}}{\|\mathcal{G}_n\|_\infty^2} = \infty,

    then the Empirical Risk Minimizer g^lMAPE,Gn,Dn=arg⁡min⁡g∈Gn1n∑i=1n∣g(Xi)−Yi∣∣Yi∣\hat{g}_{l_{MAPE}, \mathcal{G}_n, D_n} = \arg\min_{g \in \mathcal{G}_n} \frac{1}{n} \sum_{i=1}^n \frac{|g(X_i) - Y_i|}{|Y_i|} is universally strongly consistent, in the sense that the MAPE risk converges almost surely to the optimal risk:

    lim⁡n→∞LMAPE(g^lMAPE,Gn,Dn)=LMAPE∗=inf⁡g∈M(Rd,R)E[∣g(X)−Y∣∣Y∣](a.s.)\lim_{n \to \infty} L_{MAPE}(\hat{g}_{l_{MAPE}, \mathcal{G}_n, D_n}) = L^*_{MAPE} = \inf_{g \in \mathcal{M}(\mathbb{R}^d, \mathbb{R})} \mathbb{E}\left[\frac{|g(X) - Y|}{|Y|}\right] \quad (a.s.)

  2. Knowl 2 — Necessary and Sufficient Conditions for Finite Pointwise MAPE Loss

    theoretical result

    Let TT be a real-valued random variable. For m∈Rm \in \mathbb{R}, define the expected percentage error loss as J(m)=E[∣m−T∣∣T∣]J(m) = \mathbb{E}\left[\frac{|m - T|}{|T|}\right], adopting the conventions that a0=∞\frac{a}{0} = \infty for any a≠0a \ne 0 and 00=1\frac{0}{0} = 1 (which implies J(0)=1J(0) = 1).

    The expected loss J(m)<∞J(m) < \infty for all m∈Rm \in \mathbb{R} if and only if both of the following conditions hold:

    1. P(T=0)=0\mathbb{P}(T = 0) = 0
    2. The tail probability sums near zero converge: ∑k=1∞k P(T∈(1k+1,1k])<∞and∑k=1∞k P(T∈[−1k,−1k+1))<∞.\sum_{k=1}^\infty k \, \mathbb{P}\left(T \in \left(\frac{1}{k+1}, \frac{1}{k}\right]\right) < \infty \quad \text{and} \quad \sum_{k=1}^\infty k \, \mathbb{P}\left(T \in \left[-\frac{1}{k}, -\frac{1}{k+1}\right)\right) < \infty.

    If either condition is violated, then J(m)=∞J(m) = \infty for every m≠0m \ne 0.

  3. Knowl 3 — Existence and Global Optimality of the MAPE Regression Function

    theoretical result

    Let (X,Y)∈Rd×R(X, Y) \in \mathbb{R}^d \times \mathbb{R} be a random pair. Suppose that for every x∈Rdx \in \mathbb{R}^d, the conditional distribution of YY given X=xX = x satisfies P(Y=0∣X=x)=0\mathbb{P}(Y = 0 \mid X = x) = 0 as well as the discrete convergence criteria ∑k=1∞k P(Y∈(1k+1,1k]∣X=x)<∞\sum_{k=1}^\infty k \, \mathbb{P}\left(Y \in \left(\frac{1}{k+1}, \frac{1}{k}\right] \mid X = x\right) < \infty and ∑k=1∞k P(Y∈[−1k,−1k+1)∣X=x)<∞\sum_{k=1}^\infty k \, \mathbb{P}\left(Y \in \left[-\frac{1}{k}, -\frac{1}{k+1}\right) \mid X = x\right) < \infty.

    Then for every xx, the conditional expected loss function m↦E[∣m−Y∣∣Y∣∣X=x]m \mapsto \mathbb{E}\left[\frac{|m - Y|}{|Y|} \mid X = x\right] is convex, coercive (lim∣m∣→∞E[∣m−Y∣∣Y∣∣X=x]=∞\\lim_{|m| \to \infty} \mathbb{E}\left[\frac{|m - Y|}{|Y|} \mid X = x\right] = \infty), and possesses a nonempty, bounded, closed interval of global minima in R\mathbb{R}. Defining mMAPE(x)m_{MAPE}(x) as the midpoint of this interval yields a well-defined measurable function mMAPE:Rd→Rm_{MAPE}: \mathbb{R}^d \to \mathbb{R} that minimizes the global MAPE risk over all measurable functions M(Rd,R)\mathcal{M}(\mathbb{R}^d, \mathbb{R}):

    LMAPE(mMAPE)=inf⁡g∈M(Rd,R)E[∣g(X)−Y∣∣Y∣].L_{MAPE}(m_{MAPE}) = \inf_{g \in \mathcal{M}(\mathbb{R}^d, \mathbb{R})} \mathbb{E}\left[\frac{|g(X) - Y|}{|Y|}\right].

  4. Knowl 4 — Dual Formulation of MAPE Kernel Quantile Regression

    model/method

    Let H\mathcal{H} be a Reproducing Kernel Hilbert Space (RKHS) with feature map ϕ:Rd→H\phi: \mathbb{R}^d \to \mathcal{H} and positive-definite kernel k(x,x′)=⟨ϕ(x),ϕ(x′)⟩Hk(x, x') = \langle \phi(x), \phi(x') \rangle_\mathcal{H}. Given a training dataset (xi,yi)i=1n(x_i, y_i)_{i=1}^n with yi≠0y_i \ne 0, quantile parameter τ∈[0,1]\tau \in [0, 1], and regularization parameter λ>0\lambda > 0, the primal regularized MAPE quantile regression problem is:

    min⁡w∈H,b∈R,ξ,ξ∗1nλ∑i=1nτξi+(1−τ)ξi∗∣yi∣+12∥w∥H2\min_{w \in \mathcal{H}, b \in \mathbb{R}, \xi, \xi^*} \frac{1}{n\lambda} \sum_{i=1}^n \frac{\tau \xi_i + (1 - \tau)\xi_i^*}{|y_i|} + \frac{1}{2} \|w\|_\mathcal{H}^2

    subject to: yi−⟨ϕ(xi),w⟩H−b≤∣yi∣ξi,∀i=1,…,ny_i - \langle \phi(x_i), w \rangle_\mathcal{H} - b \le |y_i| \xi_i, \quad \forall i = 1, \dots, n ⟨ϕ(xi),w⟩H+b−yi≤∣yi∣ξi∗,∀i=1,…,n\langle \phi(x_i), w \rangle_\mathcal{H} + b - y_i \le |y_i| \xi_i^*, \quad \forall i = 1, \dots, n ξi≥0,ξi∗≥0,∀i=1,…,n.\xi_i \ge 0, \quad \xi_i^* \ge 0, \quad \forall i = 1, \dots, n.

    Letting C=1nλC = \frac{1}{n\lambda} and K∈Rn×nK \in \mathbb{R}^{n \times n} be the Gram matrix where Kij=k(xi,xj)K_{ij} = k(x_i, x_j), the corresponding Wolfe dual optimization problem is:

    max⁡α∈RnαTy−12αTKα\max_{\alpha \in \mathbb{R}^n} \alpha^T y - \frac{1}{2} \alpha^T K \alpha

    subject to: ∑i=1nαi=0andC(τ−1)∣yi∣2≤αi≤Cτ∣yi∣2,∀i=1,…,n.\sum_{i=1}^n \alpha_i = 0 \quad \text{and} \quad \frac{C(\tau - 1)}{|y_i|^2} \le \alpha_i \le \frac{C\tau}{|y_i|^2}, \quad \forall i = 1, \dots, n.

    The reconstructed regression function is f(x)=∑i=1nαik(xi,x)+bf(x) = \sum_{i=1}^n \alpha_i k(x_i, x) + b.

  5. Knowl 5 — VC-Dimension Preservation from MAE to MAPE Loss Classes

    theoretical result

    Let G\mathcal{G} be an arbitrary class of regression functions from Rd\mathbb{R}^d to R\mathbb{R}. For a loss function l:R×R→R+l: \mathbb{R} \times \mathbb{R} \to \mathbb{R}_+, define the subgraph indicator function class:

    H+(G,l)={(x,y,t)↦It≤l(g(x),y)∣g∈G}.\mathcal{H}^+(\mathcal{G}, l) = \left\{ (x, y, t) \mapsto \mathbb{I}_{t \le l(g(x), y)} \mid g \in \mathcal{G} \right\}.

    Replacing the Mean Absolute Error loss lMAE(p,y)=∣p−y∣l_{MAE}(p, y) = |p - y| by the Mean Absolute Percentage Error loss lMAPE(p,y)=∣p−y∣∣y∣l_{MAPE}(p, y) = \frac{|p - y|}{|y|} does not increase the Vapnik-Chervonenkis dimension:

    VCdim(H+(G,lMAPE))≤VCdim(H+(G,lMAE)).\text{VCdim}(\mathcal{H}^+(\mathcal{G}, l_{MAPE})) \le \text{VCdim}(\mathcal{H}^+(\mathcal{G}, l_{MAE})).

  6. Knowl 6 — Covering Number Bounds of MAPE Function Classes via MAE

    theoretical result

    Let G\mathcal{G} be a class of regression functions from Rd\mathbb{R}^d to R\mathbb{R}, and define the loss function class H(G,l)={(x,y)↦l(g(x),y)∣g∈G}\mathcal{H}(\mathcal{G}, l) = \{ (x, y) \mapsto l(g(x), y) \mid g \in \mathcal{G} \}.

    1. Supremum Norm Bound: If ∥G∥∞=sup⁡g∈Gsup⁡x∣g(x)∣<∞\|\mathcal{G}\|_\infty = \sup_{g \in \mathcal{G}} \sup_{x} |g(x)| < \infty and the target variable is bounded away from zero by YL>0Y_L > 0 such that ∣y∣≥YL|y| \ge Y_L, then for any ϵ>0\epsilon > 0: N(ϵ,H(G,lMAPE),∥⋅∥YL∞)≤N(ϵYL,H(G,lMAE),∥⋅∥∞),\mathcal{N}(\epsilon, \mathcal{H}(\mathcal{G}, l_{MAPE}), \|\cdot\|_{Y_L \infty}) \le \mathcal{N}(\epsilon Y_L, \mathcal{H}(\mathcal{G}, l_{MAE}), \|\cdot\|_\infty), where ∥h∥YL∞=sup⁡x∈Rd,∣y∣≥YL∣h(x,y)∣\|h\|_{Y_L \infty} = \sup_{x \in \mathbb{R}^d, |y| \ge Y_L} |h(x, y)|.

    2. Empirical LpL_p Norm Bound: For any training dataset Dn=(xi,yi)i=1nD_n = (x_i, y_i)_{i=1}^n with yi≠0y_i \ne 0 for all ii, and for any p≥1p \ge 1: N(ϵ,H(G,lMAPE),∥⋅∥p,Dn)≤N(ϵmin⁡1≤i≤n∣yi∣,H(G,lMAE),∥⋅∥p,Dn),\mathcal{N}(\epsilon, \mathcal{H}(\mathcal{G}, l_{MAPE}), \|\cdot\|_{p, D_n}) \le \mathcal{N}\left(\epsilon \min_{1 \le i \le n} |y_i|, \mathcal{H}(\mathcal{G}, l_{MAE}), \|\cdot\|_{p, D_n}\right), where ∥h1−h2∥p,Dn=(1n∑i=1n∣h1(xi,yi)−h2(xi,yi)∣p)1/p\|h_1 - h_2\|_{p, D_n} = \left( \frac{1}{n} \sum_{i=1}^n |h_1(x_i, y_i) - h_2(x_i, y_i)|^p \right)^{1/p}.

  7. Knowl 7 — Instance-Weighted Dual Constraint Dynamics in Kernel MAPE Regression

    theoretical result

    In comparison to standard kernel quantile regression (which has box constraints C(τ−1)≤αi≤CτC(\tau - 1) \le \alpha_i \le C\tau), kernel MAPE regression imposes instance-dependent box constraints in the dual optimization problem:

    C(τ−1)∣yi∣2≤αi≤Cτ∣yi∣2,∀i=1,…,n.\frac{C(\tau - 1)}{|y_i|^2} \le \alpha_i \le \frac{C\tau}{|y_i|^2}, \quad \forall i = 1, \dots, n.

    This leads to two properties:

    1. Small ∣yi∣|y_i| emphasis: For sample points with ∣yi∣<1|y_i| < 1, the feasible range for αi\alpha_i expands relative to the MAE case, assigning greater optimization capacity and fitting power to observations near zero where relative percentage error is most sensitive.
    2. Vanishing regularization equivalence: When the regularization penalty λ→0\lambda \to 0 (meaning C=1nλ→∞C = \frac{1}{n\lambda} \to \infty), the box constraint intervals for αi\alpha_i expand to (−∞,∞)(-\infty, \infty) for all ii under both MAE and MAPE formulations. Consequently, in the unregularized interpolation limit (f(xi)≈yif(x_i) \approx y_i), the dual solutions of MAE and MAPE kernel regressions coincide.
  8. Knowl 8 — Reduction of MAPE Empirical Risk Minimization to Weighted MAE Regression

    model/method

    Minimizing the empirical MAPE over a hypothesis class G\mathcal{G} on training data (xi,yi)i=1n(x_i, y_i)_{i=1}^n with yi≠0y_i \ne 0:

    g^=arg⁡min⁡g∈G1n∑i=1n∣g(xi)−yi∣∣yi∣\hat{g} = \arg\min_{g \in \mathcal{G}} \frac{1}{n} \sum_{i=1}^n \frac{|g(x_i) - y_i|}{|y_i|}

    is mathematically equivalent to solving a weighted Mean Absolute Error (median) regression problem with fixed instance weights wi=1∣yi∣w_i = \frac{1}{|y_i|}:

    g^=arg⁡min⁡g∈G1n∑i=1nwi∣g(xi)−yi∣.\hat{g} = \arg\min_{g \in \mathcal{G}} \frac{1}{n} \sum_{i=1}^n w_i |g(x_i) - y_i|.

    When G\mathcal{G} is a class of linear models g(x)=wTx+bg(x) = w^T x + b, this optimization reduces to a linear program solvable via interior-point or simplex algorithms. For general parametric and nonparametric models, any quantile regression algorithm supporting instance weights at quantile level τ=0.5\tau = 0.5 directly optimizes the empirical MAPE.

  9. Knowl 9 — Empirical Comparison of Kernel MAPE vs. Kernel Median Regression across Translation Shifts

    data/table

    Synthetic data were generated according to Y=sinc(X,a)+ϵ(X)Y = \text{sinc}(X, a) + \epsilon(X), where sinc(X,a)=a+sin⁡(2πX)2πX\text{sinc}(X, a) = a + \frac{\sin(2\pi X)}{2\pi X}, X∼U([−1,1])X \sim \mathcal{U}([-1, 1]), and ϵ(X)∼N(0,(0.1⋅exp⁡(1−X))2)\epsilon(X) \sim \mathcal{N}\left(0, (0.1 \cdot \exp(1 - X))^2\right), using 1000 training points and 1000 test points with a Gaussian kernel and 5-fold cross-validation to select the regularization hyperparameter CC.

    aa MAPE(y,f^MAE,a)\text{MAPE}(y, \hat{f}_{MAE, a}) (%) MAPE(y,f^MAPE,a)\text{MAPE}(y, \hat{f}_{MAPE, a}) (%) CMAEC_{MAE} CMAPEC_{MAPE}
    0.00 128.62 94.09 0.01 0.10
    0.10 187.78 100.10 0.05 0.01
    0.50 72.27 57.47 5.00 10.00
    1.00 51.39 39.53 10000.00 1.00
    2.50 10.58 10.98 5.00 1.00
    5.00 4.80 4.89 5.00 10.00
    10.00 2.39 2.40 5.00 100.00
    25.00 0.96 0.96 5.00 100000.00
    50.00 0.48 0.48 5.00 1000.00
    100.00 0.24 0.24 5.00 10000.00

    The table illustrates that when the target variable is close to zero (a≤1.00a \le 1.00), kernel MAPE regression achieves substantially lower test MAPE than kernel median (MAE) regression (e.g., 94.09% vs. 128.62% at a=0.00a = 0.00, and 100.10% vs. 187.78% at a=0.10a = 0.10). For large vertical shifts (a≥2.50a \ge 2.50), both models perform nearly identically because relative variations diminish and absolute error becomes proportional to relative error.

  10. Knowl 10 — Asymmetry and Zero-Shrinkage Characteristics of MAPE Regression

    empirical result

    Compared to Mean Absolute Error (median) regression, the Mean Absolute Percentage Error regression model exhibits three qualitative properties:

    1. Downward bias toward zero: For any positive random variable YY, the optimal point prediction under the MAPE loss is strictly below the conditional median of YY.
    2. Zero-clamping under sign ambiguity: In regions where the conditional distribution of YY given X=xX=x spans both positive and negative values around zero, the fitted MAPE estimator f^MAPE(x)\hat{f}_{MAPE}(x) collapses directly to or very near zero. Predicting y^=0\hat{y} = 0 guarantees a worst-case error bounded at exactly 100% (∣0−y∣∣y∣=1\frac{|0 - y|}{|y|} = 1), avoiding the catastrophic percentage error spikes incurred by the median predictor when it predicts a non-zero value of the incorrect sign near y=0y = 0.
    3. Non-invariance to translations: While MAE regression is translation-equivariant (shifting YY by a constant aa translates f^MAE\hat{f}_{MAE} by aa), MAPE regression alters its functional shape depending on aa because the weighting term 1∣y∣\frac{1}{|y|} changes the relative penalty of errors across the domain. As a→∞a \to \infty, f^MAPE\hat{f}_{MAPE} converges asymptotically toward f^MAE\hat{f}_{MAE}.

Coverage note — None was omitted; the theoretical bounds, consistency proofs, RKHS dual derivation, and simulation analyses have all been fully captured.

References

  1. 1.P. J. Huber, Robust estimation of a location parameter, The Annals of Mathematical Statistics 35 (1) (1964) 73–101.
  2. 2.J. S. Armstrong, F. Collopy, Error measures for generalizing about forecasting methods: Empirical comparisons, International Journal of Forecasting 8 (1) (1992) 69 – 80.
  3. 3.L. Györfi, M. Kohler, A. Krzyżak, H. Walk, A Distribution-Free Theory of Nonparametric Regression, Springer, New York, 2002.
  4. 4.M. Anthony, P. L. Bartlett, Neural Network Learning: Theoretical Foundations, Cambridge University Press, 1999.
  5. 5.R. Koenker, quantreg: Quantile regression. r package version 5.05 (2013).
  6. 6.S. Boyd, L. Vandenberghe, Convex optimization, Cambridge university press, 2004.
  7. 7.R. Koenker, G. Bassett Jr, Regression quantiles, Econometrica: journal of the Econometric Society (1978) 33–50.
  8. 8.I. Takeuchi, Q. V. Le, T. D. Sears, A. J. Smola, Nonparametric quantile estimation, The Journal of Machine Learning Research 7 (2006) 1231–1264.
  9. 9.Y. Li, Y. Liu, J. Zhu, Quantile regression in reproducing kernel hilbert spaces, Journal of the American Statistical Association 102 (477) (2007) 255–268.

Citation

MLA
de Myttenaere, A., et al. “Mean Absolute Percentage Error for Regression Models”. Neurocomputing, vol. 192, 2016, pp. 38–48, https://doi.org/10.1016/j.neucom.2015.12.114.
APA
de Myttenaere, A., Golden, B., Le Grand, B., & Rossi, F. (2016). Mean Absolute Percentage Error for regression models. Neurocomputing, 192, 38–48. https://doi.org/10.1016/j.neucom.2015.12.114
Chicago
de Myttenaere, A., B. Golden, B. Le Grand, and F. Rossi. 2016. “Mean Absolute Percentage Error for Regression Models”. Neurocomputing 192: 38–48. https://doi.org/10.1016/j.neucom.2015.12.114.
Harvard
de Myttenaere, A. et al. (2016) “Mean Absolute Percentage Error for regression models”, Neurocomputing, 192, pp. 38–48. Available at: https://doi.org/10.1016/j.neucom.2015.12.114.
Vancouver
1. de Myttenaere A, Golden B, Le Grand B, Rossi F (2016) Mean Absolute Percentage Error for regression models. Neurocomputing 192:38–48

BibTeX

@article{de_Myttenaere_2016, title={Mean Absolute Percentage Error for regression models}, volume={192}, ISSN={0925-2312}, url={http://dx.doi.org/10.1016/j.neucom.2015.12.114}, DOI={10.1016/j.neucom.2015.12.114}, journal={Neurocomputing}, publisher={Elsevier BV}, author={de Myttenaere, Arnaud and Golden, Boris and Le Grand, Bénédicte and Rossi, Fabrice}, year={2016}, month=June, pages={38–48} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF