← Back to Projects

Project 03 · Generative Molecular Design

Structure-Based Molecular Generation with DiffSBDD

A structure-conditioned molecular generation and prioritization workflow for the MDM2 binding pocket. I used a pretrained DiffSBDD model to generate candidate 3D molecules, then evaluated the resulting chemical space through validity analysis, molecular properties, applicability-domain analysis, ADMET filtering, docking, and protein–ligand interaction analysis.

DiffSBDD Diffusion Models RDKit ADMET AutoDock Vina PLIP MDM2
01 · Objective

Generating molecules from the binding pocket

Rather than beginning with a reference ligand and optimizing molecular strings, this experiment explored structure-conditioned generation. DiffSBDD was used to sample molecules directly in the three-dimensional environment of the MDM2 binding pocket.

5ZXF MDM2 protein structure used as the structural context for generation.
3,000 Molecular structures requested from the pretrained DiffSBDD generation workflow.
2,910 Unique RDKit-valid generated molecules retained after structure parsing and duplicate removal.
02 · Computational Workflow

Generation to structural prioritization

Generation was treated as the beginning of the workflow rather than the final result. Generated structures passed through successive computational filters before structural evaluation.

MDM2 Pocket PDB 5ZXF
DiffSBDD 3D molecular generation
Validity RDKit parsing + uniqueness
Chemical Space Similarity + model support
SA / ADMET Developability filtering
Diversity Clustering + prioritization
Docking AutoDock Vina
Interactions PLIP analysis
03 · Generation Funnel

From thousands of structures to a focused set

Each stage reduced the generated chemical space using a different criterion, moving from basic molecular validity toward developability and structure-based prioritization.

3,000 Structures requested from DiffSBDD
2,957 SDF molecular records generated
2,915 RDKit-valid molecules
2,910 Unique valid molecules
2,863 Molecules successfully evaluated by the QSAR pipeline
2,154 Molecules retained after SA ≤ 5 filtering
1,025 Passed core ADMET criteria
786 Retained after solubility-aware filtering
20 Top candidates advanced to molecular docking
3 Final candidates selected for detailed structural analysis
04 · Chemical Space

Novel chemistry creates a model-support challenge

The generated molecules were structurally distinct from the ligand-based training chemistry. This provided useful chemical novelty, but also exposed an important limitation when applying a QSAR model to diffusion-generated structures.

Similarity to Nutlin-3a

Generated molecules showed low structural similarity to the Nutlin-3a reference, with Tanimoto similarities spanning approximately 0.013 to 0.239.

QSAR Applicability Domain

Among 2,863 evaluated molecules, the mean nearest-training similarity was approximately 0.187 and the maximum was 0.372, below the applicability-domain threshold of 0.646.

0 / 2,863 molecules were inside the defined QSAR applicability domain

This result changed how the QSAR predictions were interpreted. Predicted potency could still be inspected as exploratory model output, but it was not treated as validated evidence for ranking these structurally novel generated molecules.

05 · Developability Filtering

Prioritizing beyond predicted potency

Because the generated chemistry fell outside the ligand-based QSAR applicability domain, prioritization emphasized molecular quality, predicted developability, diversity, and structural evaluation.

Synthetic Accessibility

Of the 2,863 QSAR-evaluable molecules, 2,154 passed the synthetic-accessibility criterion of SA ≤ 5, corresponding to approximately 75.2% of the evaluated set.

ADMET Filtering

ADMET predictions were calculated for the 2,154 molecules passing the SA filter. A total of 1,025 passed the core ADMET criteria, and 786 remained after applying the additional solubility criterion.

06 · Diversity

Preserving chemical diversity

The 786 retained molecules remained highly diverse. Butina clustering at a Tanimoto threshold of 0.60 produced 782 clusters, indicating that the filtered candidate set covered a broad range of structural solutions rather than collapsing onto a small number of closely related chemotypes.

786 Filtered molecules entering diversity analysis
782 Butina clusters at Tanimoto 0.60
20 Prioritized structures advanced to docking
07 · Structural Evaluation

Docking and protein–ligand interactions

The top 20 prioritized molecules were docked into the MDM2 binding site. Three candidates were then selected for detailed protein–ligand interaction analysis.

Candidate 0504

−6.135 kcal/mol

Best docking score among the three highlighted candidates. PLIP identified seven hydrophobic interactions under the applied interaction criteria.

Candidate 2032

−5.494 kcal/mol

Advanced through the prioritization workflow, although no qualifying protein–ligand interactions were identified by PLIP under the applied criteria.

Candidate 2654

−5.091 kcal/mol

PLIP identified two hydrophobic interactions and one hydrogen bond, providing complementary structural evidence beyond the docking score alone.

08 · Key Takeaway

Generation is only the beginning

DiffSBDD successfully explored chemistry far from the reference ligand and ligand-based QSAR training space. The same novelty that makes generative modeling valuable, however, also limits confidence in predictions from models trained on narrower chemical domains.

Integrated prioritization matters

The project demonstrates why generated molecules should not be ranked by a single predicted activity value. Applicability domain, molecular properties, ADMET, diversity, docking, and interaction analysis provide complementary evidence for deciding which generated structures are worth advancing.

Structure-based generative molecular design

This project complements the reinforcement-learning workflow by exploring a different generative paradigm: direct molecular generation conditioned on the three-dimensional protein environment.