Target discovery and drug design in the era of artificial intelligence

Payton Fleming, Andrey A. IvanovView original
ComprehensiveExpertmarcus voice
Ten to the power of sixty. That's a common estimate for the number of drug-like molecules that could theoretically exist. Every drug discovery program in history has sampled a fraction so small that it barely registers. The question Fleming and Ivanov ask in their review is not whether artificial intelligence can help navigate that space — it's precisely how, at each stage of discovery, from the moment you're trying to figure out what to drug all the way to the moment you're trying to make it. The answer they give is specific and methodological. Not a promise, but a map. Start with target discovery, because that's where most programs fail before chemistry even begins. Fleming and Ivanov frame the problem as one of dynamic networks rather than isolated molecular lesions. Complex diseases — cancer, neurodegeneration, and metabolic disorders — arise from dysregulated, interconnected systems that adapt to genetic, environmental, and therapeutic perturbations. A differential expression hit or a biochemical annotation can point you toward a gene, but it tells you almost nothing about whether modulating that gene will meaningfully shift a disease phenotype in a specific cellular context. The biology is distributed. The interventions need to be distributed as well. Artificial intelligence recalibrates this along three axes. The first is network-enabled inference. Graph neural networks, or GNNs, integrate interaction topology with node attributes, such as expression levels, mutation status, and perturbation responses, to produce context-aware embeddings used for link prediction and functional module detection. A method called SPIDER uses graph attention networks to merge global interactomes with condition-specific molecular features, yielding cell-type-specific protein-protein interaction networks. Discriminative network embedding applies contrastive self-supervision on interaction topology to recover functional modules, even when node annotations are sparse. New sequence-based interaction predictors like diffPALM and InteracTor extend this further, enabling systematic expansion of protein-protein interaction maps beyond what experimental bias captures. The second axis is multimodal fusion. Autoencoder architectures compress genomics, transcriptomics, proteomics, phosphoproteomics, and chemical perturbation data into unified representations. VEGA, a variational autoencoder constrained to represent pathways and regulatory modules, recovered STAT3 and OLIG2 activity in glioblastoma that standard differential expression missed. The Connectivity Map's L1000 compendium and platforms like PandaOmics link transcriptomic signatures to compounds, connecting target inference directly to chemical starting points. The third axis is perturbation-aware modeling. PDGrapher identifies candidate targets whose modulation optimally shifts diseased transcriptional states. It prioritized KDR, the gene encoding VEGFR2, in non-small-cell lung cancer based on transcriptional reversal logic. Geneformer uses large-scale single-cell pretraining and in silico gene deletion to recover dosage-sensitive regulators, identifying the TEAD4 gene as a cardiomyopathy target. PINNACLE integrates single-cell transcriptomics with protein-protein interaction topology and nominated tumor necrosis factor and interleukin 6 receptor as context-specific targets in inflammatory disease, based on cellular specificity and network centrality. These approaches also push into historically intractable target classes. Non-enzymatic proteins, transcriptional regulators, and protein-protein interfaces exert influence through connectivity rather than catalytic activity, and AI-enabled network and sequence models expose them. Fleming and Ivanov provide a concrete example: the histone methyltransferase isoform NSD3S lacks catalytic activity but uses a defined fifteen-amino-acid region to bind MYC and block its degradation via FBXW7. A noncanonical MKK3-MYC interaction stabilizes and activates MYC independent of kinase activity. These are not conventionally druggable sites by classical criteria, but they are structurally and functionally defined interaction motifs that AI can surface. Once you have a target, you need structure. The availability of structure has been transformed. AlphaFold2 achieves near-experimental accuracy for many proteins from sequence alone, and proteome-scale deployment has extended reliable structural coverage to most human proteins, enabling early pocket analysis before a single crystal is grown. AlphaFold 3 extends this to biomolecular complexes — protein-protein, protein-nucleic acid, and protein-ligand — with pose success evaluated using pocket-aligned ligand root mean square deviation thresholds below two angstroms and DockQ metrics. Example per-complex local distance difference test values from the review include eighty-two point eight for a bacterial CRP/FNR regulator bound to DNA and cyclic GMP, and eighty-three point zero for a four thousand six hundred sixty-five residue coronavirus spike complex bound by antibodies. These numbers represent a genuine expansion in what can be structurally characterized at the start of a program. That expansion feeds virtual screening at unprecedented scale. PubChem currently contains over seventy-five million unique compounds, but fewer than ten percent have biological test results. Make-on-demand libraries like Enamine REAL and WuXi GalaXi push synthetically accessible space beyond twenty billion molecules. Machine learning accelerates the traversal of that space. Deep Docking trains on a subset of docked compounds and reduces exhaustive docking requirements by up to two orders of magnitude while preserving enrichment. GNINA integrates three-dimensional convolutional neural networks into docking pipelines and shows substantial gains in enrichment over empirical scoring functions. RoseTTAFold All-Atom jointly predicts backbone, side-chain conformations, and ligand poses, improving accuracy when receptor structures are unknown or partially resolved. But structure prediction has real limits that practitioners need to internalize. AlphaFold family models output static, dominant conformations. They do not approximate solution ensembles. E3 ubiquitin ligases — natively open in their apoprotein state, observed closed only when ligand-bound — were predicted exclusively in the closed state by AlphaFold 3. Accuracy degrades substantially when median sequence alignment depth falls below roughly thirty sequences. One benchmark reported a chirality violation rate of four point four percent in PoseBusters. Co-folding models struggle to capture larger ligand-induced conformational changes, particularly loop rearrangements, and their ability to distinguish true binders from false positives in large virtual screening lists was target-dependent and often inferior to physics-based docking. The most reliable workflows remain hybrid: AI for rapid hypothesis generation, physics for mechanistic refinement, and experiment for validation. Affinity prediction is moving in the right direction, but slowly. Foundation models like Boltz-2 attempt to unify structure prediction and affinity estimation. In a prospective benchmark on five hundred fifty-seven previously unpublished SARS-CoV-2 Mac1 macrodomain ligand-protein crystal structures, diffusion-based co-folding models including AlphaFold 3, Chai-1, and Boltz-2 reproduced a substantial fraction of ligand poses with near-atomic accuracy. Boltz-2 showed the strongest correlation between predicted affinity and experimental potency among co-folding methods, but the review is clear that its affinity estimates remain only weakly correlated with docking scores and complement rather than replace physics-based screening. Shifting from structure to ligands, the transformation is equally significant. The modern era of ligand-based drug design replaces fixed fingerprints with learned, task-adaptive molecular representations. Graph neural networks treat molecules as atom-bond graphs and produce learned embeddings that outperform random forest baselines under scaffold-based splits — a stricter test of generalization across structural classes. Chemical language models operate on SMILES strings: masked-token pretraining approaches like MG-BERT and chemically informed tokenizations like FG-BERT, which masks functional groups, both show consistent performance gains across physicochemical properties, absorption distribution metabolism excretion toxicity, and phenotypic tasks after pretraining on large unlabeled corpora. The practical change is the two-stage paradigm: self-supervised pretraining on massive unlabeled chemical libraries followed by fine-tuning on smaller labeled assay datasets. This dramatically reduces the labeled-data burden for new endpoints and allows models to transfer general structural knowledge to specific biological questions. Multitask learning amplifies this further, scaling naturally to multi-property objectives. Uncertainty estimation is not optional in this framework — it's operational. Deep models can be overconfident outside their training distribution, and without calibrated confidence measures, active learning loops degrade into noise. Random forests and Gaussian processes remain widely used precisely because they provide principled uncertainty quantification in data-limited regimes. The review's prospective successes illustrate what happens when these components align: a model trained on approximately seven thousand five hundred screened compounds helped prioritize molecules leading to abaucin, a narrow-spectrum antimicrobial; deep-learning-guided phenotypic screening produced halicin. Explainability tools like integrated gradients and SHAP connect model outputs to substructures, converting predictions into medicinal chemistry hypotheses rather than black-box rankings. Generative molecular design sits downstream of all this — it's where the inference from target, structure, and ligand data gets converted into new chemistry. Reinforcement learning is the most widely adopted framework: a policy network proposes structural edits and is rewarded for predicted activity while penalized for liabilities. Fragment-based and multiparameter reinforcement learning methods optimize several properties simultaneously, mirroring lead optimization practice. Variational autoencoders establish continuous latent chemical spaces that permit interpolation between chemotypes; diffusion approaches generate chemically and structurally coherent three-dimensional outputs, particularly valuable in structure-aware settings. Transformer-based chemical language models fine-tuned on receptor-specific ligand sets have yielded experimentally validated nanomolar A2A adenosine receptor ligands. In one study, twelve prioritized binders were synthesized and seven were confirmed as dual modulators. REINVENT 4 integrates transformer and sequence-based generators with reinforcement learning and curriculum learning to support R-group replacement, linker design, scaffold hopping, and library design in a single framework. These are not demonstrations on paper — GENTRL generated DDR1 kinase inhibitors with an inhibitory concentration of ten nanomolar that were synthesized and confirmed for biochemical and cellular activity; hierarchical graph models produced DYRK1A inhibitors with an inhibitory concentration of forty-one nanomolar. The expansion into non-classical modalities is particularly striking. For PROTACs — proteolysis-targeting chimeras — transformer-based generative frameworks coupled with reinforcement learning and physics-based filtering have produced low-nanomolar degraders. Chemistry42-derived PKMYT1 inhibitors incorporated into CRBN-recruiting PROTACs achieved a DC50 of approximately one nanomolar with robust antitumor efficacy in xenografts, compared to the established BRD4 degrader AT7 with a DC50 of approximately twenty point eight nanomolar. Molecular glue discovery is becoming systematic: GlueFinder mined over two thousand six hundred human protein dimers and identified interface-adjacent pockets that support glue-mediated complex formation, recapitulating known glues like thalidomide at CRBN-SALL4. A separate integrated AI pipeline identified bufalin as a glue stabilizing the estrogen receptor alpha and STUB1 complex, with estrogen receptor alpha binding affinity of five point four micromolar and glue-induced complex stabilization yielding potent degradation with in vivo efficacy, including reversal of tamoxifen resistance in xenograft and patient-derived organoid models. For peptides and mini-proteins, diffusion frameworks like RFdiffusion generate binders conditioned on target geometry and hotspot complementarity, achieving nanomolar affinities with near-atomic agreement between predicted and measured structures by cryo-electron microscopy and surface plasmon resonance. The review's critical point about generative design is worth sitting with. These models optimize predefined scoring functions. If those scores are wrong — and they often are — the outputs are chemically unstable, synthetically inaccessible, or biologically irrelevant. The successes happened when generative objectives were tightly coupled to synthetic constraints, potency predictors, and iterative experimental feedback. That coupling brings us to retrosynthesis. Retrosynthesis and forward reaction prediction are the layers most often underestimated, and the review treats them as explicit bridges between computational design and the bench. Multi-step retrosynthetic planners like ASKCOS and AiZynthFinder, along with commercial environments like IBM RXN for Chemistry and Synthia, combine transformer-based reaction prediction with curated reaction knowledge to propose laboratory-suitable routes. RAscore provides a machine-learning synthetic accessibility metric trained on retrosynthesis planner outcomes that rapidly estimates tractability without full route enumeration. The Molecular Transformer and Chemformer achieve high accuracy in predicting reaction outcomes across diverse reaction classes, supporting feasibility checks and data-driven reaction optimization. The benchmarking problem here is acute. Training sets are biased toward well-studied chemotypes. Models that perform well on canonical reactions struggle with novel chemical space — precisely where generative design most needs them. Dataset leakage and inconsistent benchmarking practices compound the problem. This distribution mismatch is not peripheral; it's the central obstacle to end-to-end AI-driven discovery. Which brings everything together in the limitations. Fleming and Ivanov are direct: data quality and coverage constrain every layer. Biological interaction maps and perturbation datasets are unevenly distributed across genes and disease contexts. High-quality negative data is scarce. Label noise and assay variability propagate uncertainty into predictions. Models trained on cell lines do not always generalize to organoids, in vivo models, or patient samples. Genetic perturbations do not always recapitulate pharmacological modulation. Deep models are frequently overconfident outside their domain of applicability, which is why calibrated uncertainty is not a nice-to-have but an operational prerequisite. Mechanistic interpretability — linking molecular features to biological pathways and network effects — is emerging as a defining requirement, not just for scientific credibility but for regulatory confidence and experimental follow-through. The integrated loop the review advocates is not a vague future aspiration. It already exists in partial form in platforms like Chemistry42 and REINVENT 4, and in case studies where AI-generated molecules reached biochemical and in vivo validation. The architecture is: multimodal target prioritization feeds structure-based and ligand-based design, which feeds generative optimization constrained by absorption distribution metabolism excretion toxicity and synthetic feasibility assessment, which feeds experimental validation, which feeds back into models. Human domain knowledge defines meaningful objectives and experimental tests. AI accelerates prioritization and exploration within that loop. The gaps are real and documented, but so is the progress. The PKMYT1 degrader at sub-nanomolar DC50, the abaucin and halicin discoveries, the RFdiffusion binders with near-atomic structural confirmation — these are not demonstrations in held-out benchmarks. They are compounds that were made and tested. That's the standard Fleming and Ivanov hold the field to, and it's the right one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Ten to the power of sixty. That's a common estimate for the number of drug-like molecules that could theoretically exist. Every drug discovery program in history has sampled a fraction so small that it barely registers. The question Fleming and Ivanov ask in their review is not whether artificial intelligence can help navigate that space — it's precisely how, at each stage of discovery, from the moment you're trying to figure out what to drug all the way to the moment you're trying to make it. The answer they give is specific and methodological. Not a promise, but a map. Start with target discovery, because that's where most programs fail before chemistry even begins. Fleming and Ivanov frame the problem as one of dynamic networks rather than isolated molecular lesions. Complex diseases — cancer, neurodegeneration, and metabolic disorders — arise from dysregulated, interconnected systems that adapt to genetic, environmental, and therapeutic perturbations. A differential expression hit or a biochemical annotation can point you toward a gene, but it tells you almost nothing about whether modulating that gene will meaningfully shift a disease phenotype in a specific cellular context. The biology is distributed. The interventions need to be distributed as well.

Artificial intelligence recalibrates this along three axes. The first is network-enabled inference. Graph neural networks, or GNNs, integrate interaction topology with node attributes, such as expression levels, mutation status, and perturbation responses, to produce context-aware embeddings used for link prediction and functional module detection. A method called SPIDER uses graph attention networks to merge global interactomes with condition-specific molecular features, yielding cell-type-specific protein-protein interaction networks. Discriminative network embedding applies contrastive self-supervision on interaction topology to recover functional modules, even when node annotations are sparse. New sequence-based interaction predictors like diffPALM and InteracTor extend this further, enabling systematic expansion of protein-protein interaction maps beyond what experimental bias captures. The second axis is multimodal fusion. Autoencoder architectures compress genomics, transcriptomics, proteomics, phosphoproteomics, and chemical perturbation data into unified representations. VEGA, a variational autoencoder constrained to represent pathways and regulatory modules, recovered STAT3 and OLIG2 activity in glioblastoma that standard differential expression missed. The Connectivity Map's L1000 compendium and platforms like PandaOmics link transcriptomic signatures to compounds, connecting target inference directly to chemical starting points.

The third axis is perturbation-aware modeling. PDGrapher identifies candidate targets whose modulation optimally shifts diseased transcriptional states. It prioritized KDR, the gene encoding VEGFR2, in non-small-cell lung cancer based on transcriptional reversal logic. Geneformer uses large-scale single-cell pretraining and in silico gene deletion to recover dosage-sensitive regulators, identifying the TEAD4 gene as a cardiomyopathy target. PINNACLE integrates single-cell transcriptomics with protein-protein interaction topology and nominated tumor necrosis factor and interleukin 6 receptor as context-specific targets in inflammatory disease, based on cellular specificity and network centrality. These approaches also push into historically intractable target classes. Non-enzymatic proteins, transcriptional regulators, and protein-protein interfaces exert influence through connectivity rather than catalytic activity, and AI-enabled network and sequence models expose them. Fleming and Ivanov provide a concrete example: the histone methyltransferase isoform NSD3S lacks catalytic activity but uses a defined fifteen-amino-acid region to bind MYC and block its degradation via FBXW7. A noncanonical MKK3-MYC interaction stabilizes and activates MYC independent of kinase activity. These are not conventionally druggable sites by classical criteria, but they are structurally and functionally defined interaction motifs that AI can surface.

Once you have a target, you need structure. The availability of structure has been transformed. AlphaFold2 achieves near-experimental accuracy for many proteins from sequence alone, and proteome-scale deployment has extended reliable structural coverage to most human proteins, enabling early pocket analysis before a single crystal is grown. AlphaFold 3 extends this to biomolecular complexes — protein-protein, protein-nucleic acid, and protein-ligand — with pose success evaluated using pocket-aligned ligand root mean square deviation thresholds below two angstroms and DockQ metrics. Example per-complex local distance difference test values from the review include eighty-two point eight for a bacterial CRP/FNR regulator bound to DNA and cyclic GMP, and eighty-three point zero for a four thousand six hundred sixty-five residue coronavirus spike complex bound by antibodies. These numbers represent a genuine expansion in what can be structurally characterized at the start of a program. That expansion feeds virtual screening at unprecedented scale. PubChem currently contains over seventy-five million unique compounds, but fewer than ten percent have biological test results. Make-on-demand libraries like Enamine REAL and WuXi GalaXi push synthetically accessible space beyond twenty billion molecules.

Machine learning accelerates the traversal of that space. Deep Docking trains on a subset of docked compounds and reduces exhaustive docking requirements by up to two orders of magnitude while preserving enrichment. GNINA integrates three-dimensional convolutional neural networks into docking pipelines and shows substantial gains in enrichment over empirical scoring functions. RoseTTAFold All-Atom jointly predicts backbone, side-chain conformations, and ligand poses, improving accuracy when receptor structures are unknown or partially resolved. But structure prediction has real limits that practitioners need to internalize. AlphaFold family models output static, dominant conformations. They do not approximate solution ensembles. E3 ubiquitin ligases — natively open in their apoprotein state, observed closed only when ligand-bound — were predicted exclusively in the closed state by AlphaFold 3. Accuracy degrades substantially when median sequence alignment depth falls below roughly thirty sequences. One benchmark reported a chirality violation rate of four point four percent in PoseBusters.

Co-folding models struggle to capture larger ligand-induced conformational changes, particularly loop rearrangements, and their ability to distinguish true binders from false positives in large virtual screening lists was target-dependent and often inferior to physics-based docking. The most reliable workflows remain hybrid: AI for rapid hypothesis generation, physics for mechanistic refinement, and experiment for validation. Affinity prediction is moving in the right direction, but slowly. Foundation models like Boltz-2 attempt to unify structure prediction and affinity estimation. In a prospective benchmark on five hundred fifty-seven previously unpublished SARS-CoV-2 Mac1 macrodomain ligand-protein crystal structures, diffusion-based co-folding models including AlphaFold 3, Chai-1, and Boltz-2 reproduced a substantial fraction of ligand poses with near-atomic accuracy. Boltz-2 showed the strongest correlation between predicted affinity and experimental potency among co-folding methods, but the review is clear that its affinity estimates remain only weakly correlated with docking scores and complement rather than replace physics-based screening.

Shifting from structure to ligands, the transformation is equally significant. The modern era of ligand-based drug design replaces fixed fingerprints with learned, task-adaptive molecular representations. Graph neural networks treat molecules as atom-bond graphs and produce learned embeddings that outperform random forest baselines under scaffold-based splits — a stricter test of generalization across structural classes. Chemical language models operate on SMILES strings: masked-token pretraining approaches like MG-BERT and chemically informed tokenizations like FG-BERT, which masks functional groups, both show consistent performance gains across physicochemical properties, absorption distribution metabolism excretion toxicity, and phenotypic tasks after pretraining on large unlabeled corpora. The practical change is the two-stage paradigm: self-supervised pretraining on massive unlabeled chemical libraries followed by fine-tuning on smaller labeled assay datasets. This dramatically reduces the labeled-data burden for new endpoints and allows models to transfer general structural knowledge to specific biological questions. Multitask learning amplifies this further, scaling naturally to multi-property objectives.

Uncertainty estimation is not optional in this framework — it's operational. Deep models can be overconfident outside their training distribution, and without calibrated confidence measures, active learning loops degrade into noise. Random forests and Gaussian processes remain widely used precisely because they provide principled uncertainty quantification in data-limited regimes. The review's prospective successes illustrate what happens when these components align: a model trained on approximately seven thousand five hundred screened compounds helped prioritize molecules leading to abaucin, a narrow-spectrum antimicrobial; deep-learning-guided phenotypic screening produced halicin. Explainability tools like integrated gradients and SHAP connect model outputs to substructures, converting predictions into medicinal chemistry hypotheses rather than black-box rankings. Generative molecular design sits downstream of all this — it's where the inference from target, structure, and ligand data gets converted into new chemistry. Reinforcement learning is the most widely adopted framework: a policy network proposes structural edits and is rewarded for predicted activity while penalized for liabilities. Fragment-based and multiparameter reinforcement learning methods optimize several properties simultaneously, mirroring lead optimization practice.

Variational autoencoders establish continuous latent chemical spaces that permit interpolation between chemotypes; diffusion approaches generate chemically and structurally coherent three-dimensional outputs, particularly valuable in structure-aware settings. Transformer-based chemical language models fine-tuned on receptor-specific ligand sets have yielded experimentally validated nanomolar A2A adenosine receptor ligands. In one study, twelve prioritized binders were synthesized and seven were confirmed as dual modulators. REINVENT 4 integrates transformer and sequence-based generators with reinforcement learning and curriculum learning to support R-group replacement, linker design, scaffold hopping, and library design in a single framework. These are not demonstrations on paper — GENTRL generated DDR1 kinase inhibitors with an inhibitory concentration of ten nanomolar that were synthesized and confirmed for biochemical and cellular activity; hierarchical graph models produced DYRK1A inhibitors with an inhibitory concentration of forty-one nanomolar.

The expansion into non-classical modalities is particularly striking. For PROTACs — proteolysis-targeting chimeras — transformer-based generative frameworks coupled with reinforcement learning and physics-based filtering have produced low-nanomolar degraders. Chemistry42-derived PKMYT1 inhibitors incorporated into CRBN-recruiting PROTACs achieved a DC50 of approximately one nanomolar with robust antitumor efficacy in xenografts, compared to the established BRD4 degrader AT7 with a DC50 of approximately twenty point eight nanomolar. Molecular glue discovery is becoming systematic: GlueFinder mined over two thousand six hundred human protein dimers and identified interface-adjacent pockets that support glue-mediated complex formation, recapitulating known glues like thalidomide at CRBN-SALL4. A separate integrated AI pipeline identified bufalin as a glue stabilizing the estrogen receptor alpha and STUB1 complex, with estrogen receptor alpha binding affinity of five point four micromolar and glue-induced complex stabilization yielding potent degradation with in vivo efficacy, including reversal of tamoxifen resistance in xenograft and patient-derived organoid models. For peptides and mini-proteins, diffusion frameworks like RFdiffusion generate binders conditioned on target geometry and hotspot complementarity, achieving nanomolar affinities with near-atomic agreement between predicted and measured structures by cryo-electron microscopy and surface plasmon resonance.

The review's critical point about generative design is worth sitting with. These models optimize predefined scoring functions. If those scores are wrong — and they often are — the outputs are chemically unstable, synthetically inaccessible, or biologically irrelevant. The successes happened when generative objectives were tightly coupled to synthetic constraints, potency predictors, and iterative experimental feedback. That coupling brings us to retrosynthesis. Retrosynthesis and forward reaction prediction are the layers most often underestimated, and the review treats them as explicit bridges between computational design and the bench. Multi-step retrosynthetic planners like ASKCOS and AiZynthFinder, along with commercial environments like IBM RXN for Chemistry and Synthia, combine transformer-based reaction prediction with curated reaction knowledge to propose laboratory-suitable routes. RAscore provides a machine-learning synthetic accessibility metric trained on retrosynthesis planner outcomes that rapidly estimates tractability without full route enumeration. The Molecular Transformer and Chemformer achieve high accuracy in predicting reaction outcomes across diverse reaction classes, supporting feasibility checks and data-driven reaction optimization.

The benchmarking problem here is acute. Training sets are biased toward well-studied chemotypes. Models that perform well on canonical reactions struggle with novel chemical space — precisely where generative design most needs them. Dataset leakage and inconsistent benchmarking practices compound the problem. This distribution mismatch is not peripheral; it's the central obstacle to end-to-end AI-driven discovery. Which brings everything together in the limitations. Fleming and Ivanov are direct: data quality and coverage constrain every layer. Biological interaction maps and perturbation datasets are unevenly distributed across genes and disease contexts. High-quality negative data is scarce. Label noise and assay variability propagate uncertainty into predictions. Models trained on cell lines do not always generalize to organoids, in vivo models, or patient samples. Genetic perturbations do not always recapitulate pharmacological modulation. Deep models are frequently overconfident outside their domain of applicability, which is why calibrated uncertainty is not a nice-to-have but an operational prerequisite. Mechanistic interpretability — linking molecular features to biological pathways and network effects — is emerging as a defining requirement, not just for scientific credibility but for regulatory confidence and experimental follow-through.

The integrated loop the review advocates is not a vague future aspiration. It already exists in partial form in platforms like Chemistry42 and REINVENT 4, and in case studies where AI-generated molecules reached biochemical and in vivo validation. The architecture is: multimodal target prioritization feeds structure-based and ligand-based design, which feeds generative optimization constrained by absorption distribution metabolism excretion toxicity and synthetic feasibility assessment, which feeds experimental validation, which feeds back into models. Human domain knowledge defines meaningful objectives and experimental tests. AI accelerates prioritization and exploration within that loop. The gaps are real and documented, but so is the progress. The PKMYT1 degrader at sub-nanomolar DC50, the abaucin and halicin discoveries, the RFdiffusion binders with near-atomic structural confirmation — these are not demonstrations in held-out benchmarks. They are compounds that were made and tested. That's the standard Fleming and Ivanov hold the field to, and it's the right one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Computer Science