Mean Absolute Percentage Error for regression models
Arnaud de MyttenaereBoris GoldenBénédicte Le GrandFabrice Rossi
Establishes theoretical foundations for Mean Absolute Percentage Error regression by proving the universal consistency of empirical risk minimization and demonstrating that training optimal MAPE models is equivalent to weighted Mean Absolute Error regression.
In forecasting and financial modeling, practitioners often evaluate predictive performance using the Mean Absolute Percentage Error (MAPE) rather than standard metrics like the Mean Squared Error (MSE) or Mean Absolute Error (MAE), due to its intuitive interpretation in relative percentage terms. However, while traditional regression methods rest on well-established theoretical foundations, the theoretical properties and optimization mechanics of directly training models to minimize MAPE have historically remained unexplored. This creates a disconnect between the metric used to train predictive algorithms and the business metric used to evaluate their success.
The article establishes the theoretical legitimacy and practical viability of using MAPE as a primary objective function in machine learning and regression. It demonstrates that optimal MAPE-based models exist, proves that standard empirical risk minimization reliably converges to the optimal solution, and formulates a practical method to train non-linear kernel regression models directly under this objective.
To conduct this evaluation, the analysis derived mathematical proofs linking MAPE model complexity, covering numbers, and statistical learning bounds to existing formulations for median and absolute error regression. It then adapted kernel quantile regression into a dual quadratic optimization problem incorporating instance weights inversely proportional to the target values. Finally, the researchers tested the approach on simulated benchmark datasets comprising 1,000 training points and 1,000 test points under varying degrees of vertical offset and noise to observe empirical behavior across different target ranges.
The findings show that finding the best model under MAPE is equivalent to conducting weighted median regression, where each observation is weighted inversely by its true value. Theoretically, the article proves the existence of a globally optimal MAPE regression function and establishes universal strong consistency for empirical risk minimization, provided the target variable remains bounded away from zero. Computationally, in dual formulations of kernel regression, minimizing MAPE automatically widens the optimization constraints for smaller target values, forcing the algorithm to fit low-magnitude targets much more accurately. In experimental simulations, models directly trained on MAPE outperformed standard median regression models, achieving substantial error reductions (such as decreasing test MAPE from roughly 188% down to 100% when target values hovered near zero) and naturally biasing predictions lower than the conditional median.
These results provide senior decision-makers and technical leaders with the rigorous justification needed to deploy MAPE-optimized algorithms in production environments, such as energy load forecasting, pricing expensive assets, or financial gain-loss projections. Relying on models trained directly for percentage accuracy eliminates performance loss caused by metric misalignment and protects against severe relative forecasting errors on low-value transactions. However, because MAPE penalizes relative over-predictions more heavily than under-predictions, managers should anticipate that optimal MAPE models will systematically produce conservative, downward-shifted forecasts.
Organizations should actively adopt weighted quantile regression or the presented kernel formulation whenever MAPE serves as the primary operational key performance indicator, provided the predicted values are strictly non-zero. For standard linear models, teams can readily implement this via standard weighted median regression solvers. Further research and pilot testing are recommended to extend theoretical convergence guarantees to regularized kernel estimators and to examine methods for handling datasets containing values arbitrarily close to zero.
- Paper: Another look at measures of forecast accuracy, Rob J. Hyndman et al. (2006). It provides a foundational critique of percentage-based forecast accuracy metrics like MAPE, establishing the practical and mathematical failure modes that motivate formal theoretical analysis.
- Paper: Stability and Generalization, Olivier Bousquet et al. (2002). It establishes the theoretical principles of algorithmic stability and generalization error bounds used in analyzing empirical risk minimization.
- Paper: Rademacher and Gaussian Complexities: Risk Bounds and Structural Results, Peter L. Bartlett et al. (2002). It introduces Rademacher complexity and risk bound frameworks foundational for proving the universal consistency of empirical risk minimization estimators.
- Paper: Support Vector Regression Machines, Harris Drucker et al. (1996). It lays the groundwork for non-parametric loss minimization and regularized kernel regression models.
- Paper: The M4 Competition: 100,000 time series and 61 forecasting methods, Spyros Makridakis et al. (2020). It extends empirical forecasting evaluations across massive benchmark datasets where percentage-based and scaled error metrics are tested under practical competitive conditions.
- Paper: M5 accuracy competition: Results, findings, and conclusions, Spyros Makridakis et al. (2022). It evaluates modern machine-learning regression and forecasting methods on large-scale hierarchical data using specialized loss weighting strategies.
