Project 03 · Generative Molecular Design
A structure-conditioned molecular generation and prioritization workflow for the MDM2 binding pocket. I used a pretrained DiffSBDD model to generate candidate 3D molecules, then evaluated the resulting chemical space through validity analysis, molecular properties, applicability-domain analysis, ADMET filtering, docking, and protein–ligand interaction analysis.
Rather than beginning with a reference ligand and optimizing molecular strings, this experiment explored structure-conditioned generation. DiffSBDD was used to sample molecules directly in the three-dimensional environment of the MDM2 binding pocket.
Generation was treated as the beginning of the workflow rather than the final result. Generated structures passed through successive computational filters before structural evaluation.
Each stage reduced the generated chemical space using a different criterion, moving from basic molecular validity toward developability and structure-based prioritization.
The generated molecules were structurally distinct from the ligand-based training chemistry. This provided useful chemical novelty, but also exposed an important limitation when applying a QSAR model to diffusion-generated structures.
Generated molecules showed low structural similarity to the Nutlin-3a reference, with Tanimoto similarities spanning approximately 0.013 to 0.239.
Among 2,863 evaluated molecules, the mean nearest-training similarity was approximately 0.187 and the maximum was 0.372, below the applicability-domain threshold of 0.646.
This result changed how the QSAR predictions were interpreted. Predicted potency could still be inspected as exploratory model output, but it was not treated as validated evidence for ranking these structurally novel generated molecules.
Because the generated chemistry fell outside the ligand-based QSAR applicability domain, prioritization emphasized molecular quality, predicted developability, diversity, and structural evaluation.
Of the 2,863 QSAR-evaluable molecules, 2,154 passed the synthetic-accessibility criterion of SA ≤ 5, corresponding to approximately 75.2% of the evaluated set.
ADMET predictions were calculated for the 2,154 molecules passing the SA filter. A total of 1,025 passed the core ADMET criteria, and 786 remained after applying the additional solubility criterion.
The 786 retained molecules remained highly diverse. Butina clustering at a Tanimoto threshold of 0.60 produced 782 clusters, indicating that the filtered candidate set covered a broad range of structural solutions rather than collapsing onto a small number of closely related chemotypes.
The top 20 prioritized molecules were docked into the MDM2 binding site. Three candidates were then selected for detailed protein–ligand interaction analysis.
Best docking score among the three highlighted candidates. PLIP identified seven hydrophobic interactions under the applied interaction criteria.
Advanced through the prioritization workflow, although no qualifying protein–ligand interactions were identified by PLIP under the applied criteria.
PLIP identified two hydrophobic interactions and one hydrogen bond, providing complementary structural evidence beyond the docking score alone.
DiffSBDD successfully explored chemistry far from the reference ligand and ligand-based QSAR training space. The same novelty that makes generative modeling valuable, however, also limits confidence in predictions from models trained on narrower chemical domains.
The project demonstrates why generated molecules should not be ranked by a single predicted activity value. Applicability domain, molecular properties, ADMET, diversity, docking, and interaction analysis provide complementary evidence for deciding which generated structures are worth advancing.
This project complements the reinforcement-learning workflow by exploring a different generative paradigm: direct molecular generation conditioned on the three-dimensional protein environment.