Ridge
Overview
Ridge regression fits a linear model with L2 regularization and uses the absolute coefficient values |β_j| as parameter importance scores.
- Sign of β_j: whether the parameter increases or decreases the objective.
- |β_j|: magnitude of linear influence (sensitivity).
Like Spearman, it is lightweight and useful for a quick global overview before running heavier methods.
Formula
Given standardized input matrix and centered objective :
Coefficient of determination:
Clipped at 0 for numerical stability.
Importance score for parameter j:
Preprocessing and holdout evaluation
Ridge is unified with the same convention as RF-ANOVA / MDI / SHAP / PFI (previously Ridge reported an in-sample R² computed on the training data itself, which overstated fit quality when its R² badge was shown next to the other methods' holdout R²).
flowchart TD
data["Trial data<br/>(X ∈ R^{N×P}, y ∈ R^N)"] --> step1["Step 1: Drop rows with any non-finite<br/>value in x or y<br/>(before fitting or scoring)"]
step1 --> step2["Step 2: 80/20 holdout split<br/>(Fisher-Yates shuffle, seeds 42/43)"]
step1 -- "Fewer than 2 valid rows" --> zero["Return zero importances,<br/>R² = 0.0"]
step2 -- "≥ 4 valid rows" --> split1["First 80% train,<br/>remaining 20% eval"]
step2 -- "< 4 valid rows" --> split2["Fall back to using all data<br/>for both train and eval"]
split1 --> step3["Step 3: Standardize/center using TRAIN<br/>statistics only<br/>(EVAL data is transformed with the<br/>same TRAIN-derived mean/std before scoring)"]
split2 --> step3
step3 --> step4["Step 4: Solve<br/>(X_train^T X_train + αI)β = X_train^T y_{c,train}"]
step4 --> step5["Step 5: R² is computed on EVAL only<br/>(data the fit never saw)"]
Because the closed-form solve is cheap (O(n·p²)), Ridge does not apply the row cap (e.g. 2,000 rows) used by the tree-ensemble methods.
Interpreting R²
| R² | Interpretation |
|---|---|
| ≥ 0.8 | Good linear fit. Scores are reliable. |
| 0.5–0.8 | Moderate. Use as reference only. |
| < 0.5 | Poor fit. Switch to RF-ANOVA / SHAP / Sobol. |
Strengths and Limitations
Strengths:
- Very fast.
- Handles mild multicollinearity (L2 regularization).
- Easy to interpret.
Limitations:
- Cannot capture nonlinearity, thresholds, or strong interactions.
- Coefficient allocation can be unstable when features are highly correlated.
When to Use
- Quick initial screening.
- Approximately linear objective function.
- When direction and magnitude of effect are needed alongside Spearman.
Typical workflow: Ridge + Spearman for screening → RF-ANOVA / SHAP / Sobol if nonlinearity is suspected.
References
- Hoerl, A. E., & Kennard, R. W. (1970). Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics, 12(1), 55–67. https://doi.org/10.1080/00401706.1970.10488634