Faith-Shap: The Faithful Shapley Interaction Index
Che-Ping TsaiChih-Kuan YehPradeep Ravikumar
Develops Faith-Shap, a uniquely determined feature interaction index that extends standard Shapley axioms to higher-order interactions by framing explanation as optimal polynomial approximation of the underlying game value function.
Modern machine learning increasingly relies on complex black-box models whose decisions require clear interpretation. While standard game-theoretic attribution techniques assign importance values to individual features, they fail to explain complex feature interactions, such as relationships between words in text or groups of pixels in images. Previous attempts to quantify these interactions either introduced unnatural mathematical restrictions or sacrificed critical properties like efficiency, where attributions must sum to the model's total predicted outcome.
The article develops a principled framework to quantify feature interactions up to an arbitrary order while preserving standard game-theoretic axioms. The authors frame interaction attribution as the most faithful polynomial approximation to a model's underlying set function, leading to a unique formulation named the Faithful Shapley Interaction index, or Faith-Shap.
The research evaluates this framework across theoretical game formulations and empirical benchmarks. The approach demonstrates mathematical uniqueness under core axiomatic properties and assesses computational performance on tabular marketing data and sentiment classification using deep language models. The estimation procedure formulates attributions as a regularized weighted least-squares regression rather than relying solely on combinatorial sampling.
The findings establish that Faith-Shap uniquely satisfies linearity, symmetry, dummy player, and efficiency constraints in the interaction setting. Across synthetic and practical datasets, the weighted least-squares estimation required significantly fewer model evaluations to reach target accuracy compared to existing interaction indices, often reducing sample needs by several fold. Furthermore, empirical tests on language data confirmed that Faith-Shap accurately disentangles complementary and non-complementary word interactions without distorting individual feature values.
These results provide a mathematically sound and computationally viable method for auditing high-stakes artificial intelligence systems where feature interactions drive critical decisions. Unlike earlier methods that distorted lower-order or higher-order contributions, Faith-Shap offers a balanced, faithful representation of complex model behavior.
Organizations deploying complex models in regulated or high-impact environments should adopt polynomial approximation frameworks like Faith-Shap when individual feature attributions provide insufficient explanations. Future work should expand formal theoretical guarantees for sample-based approximations and explore alternative cooperative game concepts for interaction settings.
- Paper: A Unified Approach to Interpreting Model Predictions, Scott M. Lundberg et al. (2017). This foundational paper establishes the game-theoretic axiomatic framework for feature attribution and Kernel SHAP, which Faith-Shap directly generalizes to higher-order feature interactions.
- Paper: From Local Explanations to Global Understanding with Explainable AI for Trees, Scott M. Lundberg et al. (2020). This work introduces early formulations of SHAP interaction values in tree models, providing the direct baseline and practical motivation for Faith-Shap's efficient interaction index.
- Paper: Consistent Individualized Feature Attribution for Tree Ensembles, Scott M. Lundberg et al. (2018). This study analyzes mathematical consistency and local accuracy in tree-based game-theoretic attributions, foundational properties that Faith-Shap formalizes for arbitrary-order interactions.
- Paper: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Marco Tulio Ribeiro et al. (2016). This paper introduces local surrogate approximations via weighted least-squares regression, which forms the core estimation mechanism extended by Faith-Shap's polynomial formulation.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). This paper details the axiomatic approach to feature attribution, defining key invariance and completeness criteria that Faith-Shap preserves in the interaction setting.
- Paper: Learning to Estimate Shapley Values with Vision Transformers, Ian Connick Covert et al. (2023). This paper applies scalable estimation techniques to compute game-theoretic Shapley values in modern Vision Transformers, building on efficient attribution paradigms like those in Faith-Shap.
- Paper: Data Shapley in One Training Run, Jiachen T. Wang et al. (2025). This work extends cooperative game-theoretic attribution concepts beyond feature spaces into efficient data-valuation frameworks for large foundation models.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This text provides a critical theoretical critique of identifiability and statistical fragility across feature attribution and interpretability methods, offering an important perspective following Faith-Shap.
