Loading...
Overview Pipeline Architecture Validation Benchmarks
Technical Report

Methodology & Validation

SEARCH-ML is a CPU-optimized ensemble framework trained on 11,689 experimentally validated complexes. It bridges the gap between speed and accuracy, outperforming traditional docking tools.

Overview

Traditional docking approaches (AutoDock Vina, GOLD) rely on rigid-body approximations and are computationally expensive. SEARCH-ML employs a Stacked Ensemble Meta-Learning architecture to predict binding affinities with high statistical rigor using only CPU resources.

0
R² Score (Test)
0
RMSE (kcal/mol)
0
Mins / 100M Mols

Pipeline Logic

google:search

1. Data Sanitization

Dynamic regex-based filtering cleans feature names and isolates high-variance numerical descriptors from metadata.

2. Normalization

Z-score Standardization centers data at zero mean and unit variance for stable gradient descent.

3. Feature Engineering

Calculates descriptors for Ligand (LogP, TPSA), Protein (Isoelectric point), and Pocket geometry (Volume, SASA).

4. Stratification

5% of data is isolated as a hold-out test set. The remaining 95% is used for 10-Fold Cross-Validation.

Stacked Ensemble Architecture

Figure 1: Information flow in the SEARCH-ML Ensemble.

Hyperparameter Optimization

We utilized Optuna's Tree-structured Parzen Estimator (TPE) to fine-tune each base learner. This Bayesian approach focuses on high-probability parameter regions to minimize RMSE.

500
Trials per Fold
SQLite
Persistent Storage
Early Stop
Automated Pruning

Validation Strategy

The tool was validated using a high-resolution 10-fold cross-validation approach with an external hold-out set to ensure no data leakage.

google:search
Dataset Split Allocation Technical Purpose
Training Set 95% Used for 10-Fold CV loops and Meta-Model training.
Hold-out Test 5% Completely isolated. Used ONLY for final metric verification.
Cross-Validation 10 Folds Generates Out-of-Fold (OOF) predictions to train the Blender without leakage.

Comparative Benchmark

Tool Method Acc. (R²) Time (1M)
SEARCH-ML ML Ensemble 0.72 ~10 sec
GraphDTA GNN (GPU) 0.67 ~45 min
AutoDock Vina Docking 0.40 ~300 min
Glide (XP) Docking 0.45 ~600 min

*SEARCH-ML uses CPU only. GraphDTA requires GPU.

Interpretability & Feature Importance

The XGBoost feature importance plot highlights that Hydrophobicity (MolLogP) and Polarity (TPSA) are the dominant drivers of binding affinity.

XGBoost Importance

XGBoost Feature Importance Plot

CatBoost Importance

CatBoost Feature Importance Plot

Random Forest Importance

Random Forest Feature Importance Plot

LightGBM Importance

LightGBM Feature Importance Plot

Performance Correlation

Scatter plots showing the correlation between Experimental (Actual) vs Predicted Affinity values on the hold-out test set.

XGBoost

Train/Test
XGBoost Correlation Plot

CatBoost

Train/Test
CatBoost Correlation Plot

LightGBM

Train/Test
LightGBM Correlation Plot

Random Forest

Train/Test
Random Forest Correlation Plot

Final Stacked Ensemble

Best Performance
Stacked Model Correlation Plot

Final Metrics

Production Ready

Unified Stacked Ensemble

The meta-model successfully aggregates weak learners to minimize variance. By stacking CatBoost, XGBoost, Random Forest, and LightGBM, we achieved a 6.2% reduction in RMSE compared to the single best estimator.

1.0118
Test RMSE
0.7188
Test R² Score