SHAP
Overview
SHAP computes parameter importance using Shapley values from cooperative game theory (Lundberg & Lee, 2017). Each parameter receives a contribution value that represents its average marginal contribution across all possible feature coalitions.
The Importance Chart uses TreeSHAP for efficient exact Shapley value computation on Random Forest trees. Global importance is the mean across the training split, normalized to sum to 1.
Formula
Shapley value for parameter at sample :
Global SHAP importance:
Shapley values satisfy four axioms: efficiency, symmetry, linearity, dummy (a feature with no effect gets zero).
Comparison with Other Methods
| Aspect | SHAP | MDI | RF-ANOVA |
|---|---|---|---|
| Theoretical basis | Shapley axioms | Impurity reduction | Variance decomposition over leaf boxes (fANOVA) |
| High-cardinality bias | None | High | Low |
| Local interpretability | Yes | No | No |
| Cost | Medium–High | Low | Medium |
Hyperparameters
| Parameter | Value |
|---|---|
| Trees | 64 |
| Max depth | 10 |
| Max rows | 1,000 (downsampled) |
| Seed | 42 |
R² Interpretation
| R² | Meaning |
|---|---|
| ≥ 0.8 | Good fit. Scores are reliable. |
| 0.5–0.8 | Moderate. Use with caution. |
| < 0.5 | Poor fit. |
Notes
- SHAP shows global importance (mean ). Per-sample local values are not displayed.
- When features are strongly correlated, path-dependent TreeSHAP can be unstable.
When to Use
- When explainability and theoretical consistency are priorities.
- For reports requiring rigorous attribution.
- After RF-ANOVA / Permutation screening has identified top parameters.
References
- Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS 30. https://proceedings.neurips.cc/paper_files/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html
- Lundberg, S. M. et al. (2020). From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1), 56–67. https://doi.org/10.1038/s42256-019-0138-9