Generative models can produce thousands of chemically valid molecules. But in drug discovery, generating molecules is rarely the difficult part.
A molecule with high predicted potency may have poor solubility. A molecule with attractive physicochemical properties may move far outside the chemical space where the predictive model is reliable. Improving similarity to a known ligand may preserve useful chemistry but limit exploration.
This makes molecular generation fundamentally a multi-objective optimization problem.
In my MDM2–p53 inhibitor design project, I explored this problem through three generations of reinforcement-learning-driven molecular design. Rather than treating the reward function as fixed, I used the behavior of each generated population to inform the next reward design.
Starting with a Potency-Driven Objective
The first generation emphasized predicted MDM2 inhibitory activity while also considering solubility, drug-likeness, and similarity to a known reference compound.
Conceptually, the reward can be represented as:
where:
- P(x) represents predicted potency,
- S(x) represents solubility,
- Q(x) represents drug-likeness,
- T(x) represents molecular similarity,
- and the w terms determine the relative importance of each objective.
This formulation seems straightforward, but the weights encode an important scientific decision: what kind of chemistry are we asking the model to explore?
A reward dominated by potency can push generation toward molecules that exploit the predictive model while performing poorly on other properties. Conversely, placing too much weight on similarity can constrain exploration around known chemical space.
The challenge is therefore not simply maximizing reward. It is constructing a reward that represents a useful molecular-design objective.
Expanding the Objective Beyond Potency
The second generation expanded the reward to include a broader set of predicted ADMET properties.
The objective incorporated predicted potency together with drug-likeness, structural similarity, solubility, bioavailability, hERG liability, DILI risk, P-glycoprotein behavior, and CYP3A4-related properties.
This changed the optimization problem substantially.
This distinction matters because drug discovery rarely has a single optimum. A candidate may improve one property while sacrificing another. Multi-objective molecular design therefore involves navigating trade-offs rather than maximizing one isolated score.
A New Problem: Leaving the Model’s Applicability Domain
Generative optimization introduced another important issue.
A QSAR model can assign a prediction to a generated molecule even when that molecule is structurally different from the compounds used to train the model.
I therefore examined the generated molecules relative to the applicability domain (AD) of the potency model.
This revealed an important tension between generative exploration and predictive confidence. Novel chemistry is desirable—but moving too far from the training distribution can make the optimization increasingly dependent on extrapolated predictions.
This observation motivated another redesign of the reward.
Making Applicability Domain Part of the Design Loop
For the third generation, structural support from the predictive model was incorporated more explicitly into the optimization strategy.
The goal was not to force generated molecules to reproduce the training set. That would defeat much of the purpose of generative design.
Instead, the objective was to balance exploration of new chemical space with sufficient structural support for meaningful model-based prioritization.
The resulting population showed greater simultaneous satisfaction of the predefined multi-objective criteria while partially recovering QSAR structural support that had been lost during the previous generation.
This turns model validation into part of the optimization loop.
From Static Generation to Adaptive Molecular Design
The three-generation experiment changed how I think about reinforcement learning for molecular discovery.
A common representation of generative molecular design is:
But a more useful workflow can be:
In this framework, the reward function is not assumed to be perfect from the beginning.
Generated populations provide information about the behavior of the optimization system. That information can expose undesirable shifts in chemical space, conflicts among objectives, or regions where predictive models have limited structural support.
Those observations can then guide the next optimization cycle. The process becomes an adaptive computational design loop.
What I Learned
The central lesson was that molecular generation should not be evaluated simply by the number of valid molecules produced—or even by the highest predicted potency achieved.
More useful questions include:
- What objective is the model actually learning to optimize?
- What trade-offs emerge between potency, physicochemical properties, and predicted liabilities?
- How far does generated chemistry move from the predictive model’s training domain?
- Are improvements concentrated in one metric, or distributed across multiple design objectives?
- What does the generated population tell us about how the next optimization cycle should change?
For computational drug discovery, the value of generative AI may therefore lie less in producing increasingly large molecular libraries and more in creating iterative design systems in which prediction, generation, evaluation, and redesign continuously inform one another.
That is the direction I am exploring: moving from one-shot molecular generation toward adaptive, multi-objective computational design.