Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

43 Commits

Folders and files

Repository files navigation

MDS-Ensemble QSAR

Conformational-ensemble descriptors for flexible molecules — a molecular-dynamics approach to QSAR.

Motivation

Classical QSAR represents a molecule by descriptors computed from a single, static structure. For small, rigid drug-like molecules that is a reasonable approximation. But molecules larger than the classic "rule of 5" chemical space — macrocycles, cyclic peptides, long flexible chains, many natural products — are not static: in solution they wobble between many conformations, much like an intrinsically disordered protein (IDP) samples an ensemble rather than a fixed fold. For properties that depend on 3D shape and solvent exposure (permeability, solubility, and ultimately bioavailability), collapsing that ensemble to one conformer throws away exactly the information that matters.

Hypothesis. Describing a flexible molecule by a conformational ensemble generated with molecular-dynamics simulation (MDS) in explicit solvent yields QSAR features that are more predictive than single-conformer features — and the gain grows with molecular flexibility.

Approach

 public property data ──▶ curate ──▶ flexibility descriptors ──▶ representative set
                                                                        │
                                                                        ▼
                                        MD in solvent (water, set pH) ──▶ conformer
                                                                        ensemble
                                                                        │
                                                                        ▼
                          ensemble-averaged descriptors ──▶ QSAR ──▶ compare vs static

We prove the concept on two data-rich public benchmarks — one per property the project targets:

Dataset Property (endpoint) Source n (curated)
caco2_wang Caco-2 permeability Wang et al. 2016 / TDC ~896
esol Aqueous solubility Delaney 2004 / MoleculeNet ESOL ~1117

These two datasets are also a natural contrast: caco2_wang is rich in flexible, beyond-rule-of-5 molecules (≈261 bRo5, ≈81 macrocycles), where the ensemble hypothesis should bite; esol is dominated by small rigid molecules, a baseline where a single conformer should already suffice.

Stage 1 — data extraction (implemented)

scripts/01_build_poc_dataset.py runs the full stage-1 pipeline:

  1. Acquire — download each benchmark from a pinned, checksum-verified mirror (ensemble_qsar/data/sources.py). Cached under data/raw/.
  2. Curate — RDKit standardization: largest organic fragment, neutralize, canonical SMILES + InChIKey, drop non-organic/unparseable, and collapse duplicate structures (label aggregated by median). Every drop is counted in an auditable CurationReport (ensemble_qsar/data/curate.py).
  3. Describe — physicochemical + flexibility descriptors: rotatable bonds, Kier flexibility index, macrocycle flag, Fsp3, TPSA, cLogP, HBD/HBA, rule-of-5 violations, and a coarse flex_class (ensemble_qsar/data/descriptors.py).
  4. Select — a flexibility-stratified, chemically-diverse representative set for the MD stage: MaxMin diversity picks within rotatable-bond strata, deliberately over-sampling the flexible tail while keeping a rigid control group (ensemble_qsar/data/select.py).

Outputs land in data/processed/<dataset>/. The small stage-1 deliverables — representative_set.csv and summary.json — are committed as a snapshot; the full curated.{csv,parquet} tables are regenerable build artefacts (rerun the script) and are gitignored.

Run it

pip install -r requirements.txt
python scripts/01_build_poc_dataset.py                 # both datasets
python scripts/01_build_poc_dataset.py --datasets caco2_wang --n-select 80
python tests/test_pipeline.py                          # offline sanity checks

Stage 2a — CPU MD prep (implemented)

Turns each selected molecule into a parameterized, GPU-ready MD input using only CPU — conformer generation, pH protonation, AM1-BCC/GAFF2 (or ff19SB for peptides) parameterization — and hands off a dry system plus a leap_solvate.in recipe for the GPU stage. Class-aware routing (peptide vs small molecule) with review flags. See docs/PREP.md.

bash scripts/setup_conda_env.sh && conda activate mdsprep
python scripts/02_prep_molecule.py --from-validation T1_smoke
python scripts/03_prep_batch.py --input data/validation_subset.csv

Stage 2b — GPU MD + ensemble analysis (implemented, Colab)

notebooks/stage2b_md_colab.ipynb (thin driver over ensemble_qsar/md/) takes the prep hand-off and runs solvate → minimize → equilibrate → production → analysis on a Colab GPU, yielding each molecule's conformational ensemble with 3D-PSA/RMSD, PCA and cluster representatives. Checkpoint/resume-safe across Colab disconnects; fail-soft batch. Logic validated end-to-end on CPU OpenMM.

Stage 5 — one-command tool: SMILES → MD-ensemble report (implemented)

run.py orchestrates the whole pipeline (prep → MD → analysis) end-to-end and emits a self-contained explore_<mol_id>.html per molecule. It reuses the Stage 2/3 code as-is — each stage's skip logic makes re-runs resume. Everything for one molecule lands in a single merged directory results/<mol_id>/.

conda activate mdsprep                       # local GPU for MD; CPU for prep/analysis
python run.py "O=C(O)c1ccccc1" --name aspirin
python run.py --input molecules.csv --outdir results          # batch, fail-soft
python run.py "CCO" --production-ns 50 --n-frames 500 --force

Batch runs also write summary.csv and regenerate the index.html hub. See docs/CLI.md.

One-shot demo notebook. notebooks/demo_pipeline_colab.ipynb is the interactive counterpart for a single molecule on a Colab GPU: type a SMILES in a form, Run all, and it goes prep → MD → report with a live production progress bar and renders the self-contained explorer inline. Setup uses micromamba for the AmberTools binaries (no kernel restart). Regenerate with python scripts/build_demo_notebook.py.

Stage 3 — Ensemble descriptors + explorer (implemented)

ensemble_qsar/features/ aggregates each MD ensemble into a per-molecule QSAR feature row (3D-PSA, Rg, SASA, intramolecular H-bonds — distribution stats + population-weighted means) → features.csv. ensemble_qsar/viz/ builds a self-contained HTML explorer per molecule (scripts/20_export_viz.py): a data exporter writes viz_data.json, a molecule-agnostic viewer renders a t-SNE/PCA scatter (click a conformer), time slider, descriptor time series and cluster-representative 3D images — one file, no server or CDN.

python scripts/20_export_viz.py --md-dir data/md/<mol_id>   # one molecule
python scripts/21_build_all.py  --md-root data/md           # all + index.html + features.csv

Roadmap

See docs/ROADMAP.md. Stage 1 (data), Stage 2a (CPU prep), Stage 2b (GPU MD + ensemble analysis) and Stage 3 (ensemble descriptors + explorer) are done; next is training QSAR models on static vs ensemble features (the ensemble-vs-static comparison).

Repository layout

ensemble_qsar/data/   sources · curate · descriptors · select
scripts/              01_build_poc_dataset.py  (stage-1 orchestrator)
data/processed/       curated datasets + representative sets (committed)
data/raw/             cached mirror downloads (gitignored)
tests/                offline pipeline sanity tests
docs/ROADMAP.md       staged plan

Data provenance & licensing

Datasets are redistributed from their public mirrors solely as a reproducible convenience; original terms apply. Caco-2: Wang et al., J. Chem. Inf. Model. 2016. ESOL: Delaney, J. Chem. Inf. Comput. Sci. 2004 (MoleculeNet).

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages