Potency-driven exploration
Initial reinforcement-learning generation prioritized predicted MDM2 potency while maintaining molecular similarity and physicochemical constraints.
Computational Scientist | AI/ML for Drug Discovery
Interactive structural comparison of Nutlin-3a and generated MDM2 inhibitor candidates. The MDM2 receptor and Nutlin-3a reference remain fixed while selected candidates are overlaid in the binding pocket.
Iterative reward redesign progressively shifted the generated molecular population toward improved predicted potency, drug-like properties, and multi-objective constraint satisfaction.
Initial reinforcement-learning generation prioritized predicted MDM2 potency while maintaining molecular similarity and physicochemical constraints.
The reward function was redesigned to incorporate predicted bioavailability and major ADMET liabilities alongside potency and molecular quality.
Further reward refinement produced a population with higher predicted potency, improved QED and cLogP distributions, and greater satisfaction of the shared design criteria.
Across three iterative generations, reward redesign shifted the generated population toward substantially greater simultaneous satisfaction of the predefined multi-objective design criteria.
Summary statistics across unique generated molecules.
| Metric | Generation 1 | Generation 2 | Generation 3 |
|---|---|---|---|
| Unique molecules | 870 | 3,371 | 1,493 |
| Median predicted pIC50 | 6.496 | 6.805 | 7.053 |
| Shared-criteria pass rate | 0.57% | 0.86% | 36.77% |
| Median QED | 0.372 | 0.385 | 0.437 |
| Median cLogP | 6.074 | 5.904 | 5.447 |
| Median Nutlin similarity | 0.437 | 0.260 | 0.295 |
| Applicability-domain support | 0.443 | 0.297 | 0.323 |
The progression from Generation 1 to Generation 3 shows increasing median predicted potency and QED, decreasing median cLogP, and a marked increase in the fraction of molecules meeting the shared design criteria.
Structural evaluation is considered separately because population-level property improvement does not necessarily imply improved binding behavior for every selected candidate.
Selected molecules were evaluated beyond population-level property optimization using molecular docking, MM/GBSA calculations, and short molecular-dynamics simulations.
AutoDock Vina was used to evaluate candidate poses within the MDM2 binding site.
Selected complexes were compared using MM/GBSA estimates derived from simulation snapshots.
Ligand RMSD and contact persistence were examined during 1 ns implicit-solvent trajectories.
AutoDock Vina · PDB 5ZXF
| Molecule | Generation | Docking score |
|---|---|---|
| Nutlin-3a Reference | — | −8.220 kcal/mol |
| Gen2 Candidate 1 | Generation 2 | −7.219 kcal/mol |
| Gen2 Candidate 2 | Generation 2 | −7.641 kcal/mol |
| Gen2 Candidate 3 Selected | Generation 2 | −8.227 kcal/mol |
| Gen3 Candidate 1 | Generation 3 | −7.179 kcal/mol |
| Gen3 Candidate 2 | Generation 3 | −6.976 kcal/mol |
| Gen3 Candidate 3 | Generation 3 | −7.051 kcal/mol |
Gen2 Candidate 3 produced a Vina score of −8.227 kcal/mol, close to the −8.220 kcal/mol Nutlin-3a reference. Generation 3 improved several population-level design properties but did not improve docking scores relative to the strongest Generation 2 candidate.
More negative values indicate more favorable computational estimates.
MM/GBSA values are comparative computational estimates, not experimental binding free energies. Generation 3 MM/GBSA calculations were performed as a post-submission extension of the original analysis.
1 ns implicit-solvent trajectories
Generation 3 Candidate 3 produced the most negative MM/GBSA estimate, but also showed the largest ligand RMSD and weak contact persistence during the short trajectory. Candidate 2 retained more ligand–pocket contacts while maintaining a substantially more favorable MM/GBSA estimate than the Nutlin reference.
These short implicit-solvent simulations are used for computational prioritization rather than as evidence of experimental binding stability.
Evaluates whether reward redesign shifts the overall generated chemical population toward desired properties.
Evaluates plausible binding poses and docking scores within the MDM2 binding site.
Provides comparative energetic estimates for selected protein–ligand complexes.
Examines short-timescale structural behavior, ligand displacement, and contact persistence.
An integrated computational pipeline connects experimental activity data, machine-learning potency prediction, reinforcement-learning molecular generation, ADMET-aware optimization, and structure-based evaluation.
Curated MDM2 bioactivity data for target CHEMBL5023 used to construct the potency-prediction dataset.
ExtraTrees regression with molecular descriptors and fingerprints for predicted MDM2 inhibitory potency.
Reinforcement-learning molecular generation with iterative multi-objective reward redesign.
Multi-objective scoring incorporated predicted bioavailability and key ADMET liabilities alongside potency and molecular quality.
Generated molecules were prioritized using predicted properties, chemical constraints, and generation-level ranking.
AutoDock Vina evaluation of generated molecules within the MDM2 binding pocket.
Comparative energetic estimates were calculated for selected protein–ligand complexes.
Short trajectories examined ligand RMSD and protein–ligand contact persistence.
The workflow progressively narrows chemical space: experimental bioactivity data support potency modeling, the model guides generative optimization, ADMET-aware objectives reshape the molecular population, and structure-based analyses provide complementary evidence for candidate prioritization.