Generative molecular models are designed to explore chemical space. Predictive models, however, learn from a finite set of previously observed molecules.

Combining the two creates an important challenge.

A generated molecule can receive an attractive predicted potency while being structurally distant from the chemistry that supported the predictive model.

The prediction still has a numerical value. But the model may now be extrapolating rather than operating within a region of chemical space well represented by its training data.

This is where the concept of an applicability domain becomes particularly important for generative drug design.

Prediction Is Not the Same as Predictive Support

Machine-learning models can often generate predictions for inputs far beyond the examples used during training.

That computational ability should not be confused with evidence that the prediction is equally reliable everywhere in chemical space.

In molecular property modeling, the training set defines the chemistry from which the model learns relationships between molecular representation and biological or physicochemical properties.

When a new molecule resembles chemistry represented in that dataset, the model is operating in a more familiar region. As structural distance increases, predictions may depend more heavily on extrapolation.

Applicability-domain analysis asks whether a prediction is being made in a region of chemical space that is sufficiently supported by the model’s training data.

Why Generative Models Make This Problem More Important

In conventional virtual screening, a predictive model is often applied to a predefined molecular library.

Generative models change this relationship.

Instead of only evaluating existing molecules, the generative system can actively search for new structures that maximize a scoring function.

If predicted potency contributes strongly to the reward, the generator is effectively searching for molecular structures that the predictive model scores highly.

This creates a subtle risk: optimization can move toward regions where the model produces favorable predictions even though those regions have progressively weaker structural support from the training chemistry.

Predictive Model → Reward → Molecular Generation → Optimization → Increasing Chemical Exploration

The generator is doing what it was asked to do. The issue is whether the predictive model remains an equally trustworthy guide throughout that exploration.

Novelty and Predictive Reliability Pull in Different Directions

This produces a fundamental tension in generative molecular design.

We want generated molecules to explore beyond known compounds. If every generated molecule closely resembles the training data, the system may provide little meaningful chemical novelty.

But unrestricted exploration can move the generated population into regions where predictive models have limited support.

The objective is not to eliminate novelty. The objective is to understand when novelty begins to exceed the evidence supporting the predictive model.

Applicability-domain analysis therefore should not necessarily be interpreted as a binary rule that divides molecules into “good” and “bad.”

It can instead provide context for interpreting predictions and understanding the trade-off between exploration and model support.

What I Observed in Iterative MDM2 Molecular Design

I encountered this issue while developing an iterative reinforcement-learning workflow for MDM2–p53 inhibitor design.

The workflow combined a potency model with molecular generation and additional objectives related to molecular properties and predicted developability.

As the reward became more multi-objective, the generated chemistry changed.

Evaluating only predicted potency or overall reward would not fully describe what was happening. I therefore examined the generated populations relative to the structural support of the QSAR model.

This analysis showed that optimization could produce attractive predicted profiles while simultaneously moving portions of the generated population away from chemistry well represented by the potency model.

Generate → Score → Examine Chemical Space → Evaluate Model Support → Diagnose

That observation became useful feedback for the next design iteration.

From Diagnostic Metric to Design Feedback

Applicability domain is often considered after model development: build a model, make predictions, and then assess whether new compounds fall within an appropriate domain.

In an iterative generative workflow, it can play a broader role.

If a generated population systematically moves away from the predictive model’s structural support, that information can influence how the next optimization objective is designed.

Generate → Evaluate → Diagnose AD → Redesign Objective → Generate Again

In my workflow, applicability-domain analysis therefore became part of the feedback used to reason about subsequent optimization.

The aim was not to force generation back onto known compounds. Instead, the goal was to maintain useful exploration while accounting for the structural evidence supporting model-based prioritization.

Applicability Domain Is Not Uncertainty Quantification

It is also important to distinguish applicability-domain analysis from a complete estimate of predictive uncertainty.

Structural similarity or distance can indicate whether a molecule resembles the chemistry represented in the training set, but it does not capture every source of model error.

Experimental noise, dataset bias, representation choices, model architecture, endpoint heterogeneity, and other factors can all affect predictive reliability.

Applicability domain is therefore best interpreted as one layer of evidence rather than a guarantee that an individual prediction is correct.

What This Means for Generative Drug Discovery

Generative AI and predictive modeling are naturally complementary, but their interaction requires careful interpretation.

Generative models are valuable precisely because they can explore chemistry beyond what has already been observed. Predictive models, meanwhile, derive their evidence from existing data.

Effective computational design therefore requires managing the boundary between these two objectives: exploration and predictive support.

Several questions become especially useful:

  • How structurally different are generated molecules from the compounds supporting the predictive model?
  • Are the highest predicted scores concentrated in poorly supported regions of chemical space?
  • How does model support change as optimization progresses?
  • Are we rewarding genuine multi-objective improvement or increasingly relying on extrapolated predictions?
  • Can information from the generated population improve the next design iteration?
The important question is not only “What does the model predict?” but also “How much evidence does the model have for making that prediction here?”

For me, this is one of the most important lessons from connecting predictive machine learning with generative molecular design.

Applicability-domain analysis becomes more than a model validation step. It becomes part of the reasoning loop that connects molecular generation, predictive modeling, chemical exploration, and iterative design.

Maryam Taherzadeh Computational Scientist · AI/ML for Drug Discovery