Overview
Traditional docking approaches (AutoDock Vina, GOLD) rely on rigid-body approximations and are computationally expensive. SEARCH-ML employs a Stacked Ensemble Meta-Learning architecture to predict binding affinities with high statistical rigor using only CPU resources.
Pipeline Logic
1. Data Sanitization
Dynamic regex-based filtering cleans feature names and isolates high-variance numerical descriptors from metadata.
2. Normalization
Z-score Standardization centers data at zero mean and unit variance for stable gradient descent.
3. Feature Engineering
Calculates descriptors for Ligand (LogP, TPSA), Protein (Isoelectric point), and Pocket geometry (Volume, SASA).
4. Stratification
5% of data is isolated as a hold-out test set. The remaining 95% is used for 10-Fold Cross-Validation.
Stacked Ensemble Architecture
Figure 1: Information flow in the SEARCH-ML Ensemble.
Hyperparameter Optimization
We utilized Optuna's Tree-structured Parzen Estimator (TPE) to fine-tune each base learner. This Bayesian approach focuses on high-probability parameter regions to minimize RMSE.
Validation Strategy
The tool was validated using a high-resolution 10-fold cross-validation approach with an external hold-out set to ensure no data leakage.
| Dataset Split | Allocation | Technical Purpose |
|---|---|---|
| Training Set | 95% | Used for 10-Fold CV loops and Meta-Model training. |
| Hold-out Test | 5% | Completely isolated. Used ONLY for final metric verification. |
| Cross-Validation | 10 Folds | Generates Out-of-Fold (OOF) predictions to train the Blender without leakage. |
Comparative Benchmark
| Tool | Method | Acc. (R²) | Time (1M) |
|---|---|---|---|
| SEARCH-ML | ML Ensemble | 0.72 | ~10 sec |
| GraphDTA | GNN (GPU) | 0.67 | ~45 min |
| AutoDock Vina | Docking | 0.40 | ~300 min |
| Glide (XP) | Docking | 0.45 | ~600 min |
*SEARCH-ML uses CPU only. GraphDTA requires GPU.
Interpretability & Feature Importance
The XGBoost feature importance plot highlights that Hydrophobicity (MolLogP) and Polarity (TPSA) are the dominant drivers of binding affinity.
XGBoost Importance
CatBoost Importance
Random Forest Importance
LightGBM Importance
Performance Correlation
Scatter plots showing the correlation between Experimental (Actual) vs Predicted Affinity values on the hold-out test set.
XGBoost
Train/Test
CatBoost
Train/Test
LightGBM
Train/Test
Random Forest
Train/Test
Final Stacked Ensemble
Best Performance
Final Metrics
Unified Stacked Ensemble
The meta-model successfully aggregates weak learners to minimize variance. By stacking CatBoost, XGBoost, Random Forest, and LightGBM, we achieved a 6.2% reduction in RMSE compared to the single best estimator.