We have an exciting new drug target, we’ve validated the biology, we have a structure, now it’s time to design a compound. We don’t just need this compound to bind the target and to have activity (e.g. inhibit a kinase); ultimately, we need it to be effective, and safe, in humans. And ideally, this drug would be available in a pill that a person wouldn’t have to take more than once or twice a day (i.e. the effective human dose). Achieving this involves a complex interplay of drug properties and human physiology that affect the absorption, distribution, metabolism, and elimination of the compound in a way that allows for an effective concentration of the drug at the target for an extended period of time.
Now, what if we could predict whether a compound could reach an effective human dose very early in the development of a drug, or even incorporate that information into the design of the compound itself? That’s not as far-fetched as it may seem. In fact, several companies, including Roche1 and J&J2, are already doing just that. However, there are multiple approaches to human dose prediction, and each has its pros and cons.
Over the past decade, two major computational paradigms have emerged to address this challenge: purely data-driven, “AI” (really just advanced machine learning (ML) approaches and mechanistic physiologically based pharmacokinetic (PBPK) modeling. Increasingly, the industry is converging on hybrid approaches that combine the strengths of both.
Many ADMET properties are now routinely predicted using ML-based quantitative structure–property relationship (QSPR) models, often with excellent performance for endpoints such as solubility, permeability, plasma protein binding, or metabolic stability. Recent advances in deep learning and access to larger pharmaceutical datasets have further accelerated this trend. However, the prediction of integrated in vivo pharmacokinetic (PK) behavior, and ultimately human dose, has proven substantially more difficult. As Bassani et al. recently observed, despite the rapid increase in ML-based PK studies and larger datasets, “we did not observe a steady and substantial improvement in model predictivity.”3
The limitations are not surprising. Human dose is not a single intrinsic molecular property. Instead, it emerges from a complex interplay between absorption, distribution, metabolism, excretion, potency, formulation, route of administration, dosing frequency, and human physiology. Pure ML approaches attempt to infer these relationships directly from historical data, typically using molecular descriptors, fingerprints, or graph neural network representations as inputs. Several groups have demonstrated promising results using this strategy. Wang et al. developed ML models to predict intravenous PK parameters such as clearance and volume of distribution,4 while Lombardo et al. applied extensive datasets to improve prediction of human volume of distribution.5 Kosugi and Hosea compared ML-based clearance prediction against traditional bottom-up methods and demonstrated that ML methods can achieve competitive performance under certain conditions.6
These approaches offer clear advantages. ML models can be extremely fast, scalable, and well-suited for high-throughput screening environments where thousands of compounds may need to be evaluated rapidly. They are especially useful when the training domain is closely aligned with the chemistry space being explored. For early discovery teams, this enables rapid triaging of compounds before experimental data are available.
However, purely statistical approaches also have important limitations. Most notably, they often lack mechanistic interpretability. Predictions may be accurate within the training domain but difficult to extrapolate beyond it. ML models also struggle to explicitly incorporate dose-dependent phenomena, formulation effects, nonlinear PK, transporter-mediated disposition, or species translation. In many cases, they operate as “black boxes,” obscuring the physiological rationale behind predictions and making it challenging for scientists to understand why a compound is predicted to succeed or fail.
Mechanistic PBPK modeling provides a fundamentally different framework. Rather than directly learning dose from historical examples, PBPK models integrate compound-specific ADME parameters into mathematical representations of human physiology. Drug concentration-time profiles are simulated across tissues and organs using differential equations grounded in biology and anatomy. Importantly, these models can incorporate experimentally measured data, predicted properties, or both.
Historically, PBPK modeling was considered too computationally intensive for high-throughput discovery applications. Traditional PBPK workflows often required extensive manual parameterization and expert modelers. That perception has changed dramatically in recent years. Advances in automation, cloud computing, and streamlined PBPK architectures have enabled the emergence of high-throughput PBPK (HTPBPK) platforms capable of evaluating large virtual libraries rapidly and reproducibly.7
One of the most prominent examples is the High-Throughput Pharmacokinetics (HTPK) platform from Simulations Plus. HTPK combines ML-predicted or experimentally measured ADME properties with mechanistic PBPK simulation engines to rapidly estimate human PK and projected dose. Rather than replacing mechanistic modeling with AI, HTPK uses AI and predictive models to parameterize the mechanistic system. This distinction is critical. The PBPK framework provides physiological consistency, while ML models supply scalable predictions for the underlying inputs.
In practice, this hybrid approach offers several advantages over standalone ML PK prediction. First, it allows explicit incorporation of dose and route of administration, which are difficult for many QSPR models to capture directly. Second, the mechanistic structure improves interpretability: if predicted exposure is poor, scientists can determine whether the issue arises from clearance, permeability, solubility, first-pass metabolism, or another physiological process. Third, PBPK models naturally support species translation and scenario analysis, enabling simulations across preclinical species and humans using consistent mechanistic assumptions.
The performance of HTPK-style workflows has been encouraging. Naga et al. evaluated high-throughput PBPK predictions in discovery settings and demonstrated that mechanistic approaches could successfully inform early-stage compound prioritization and human PK estimation.1 Importantly, these models were able to provide actionable guidance while maintaining the flexibility required for medicinal chemistry optimization.
The pharmaceutical industry has increasingly embraced these hybrid strategies. Roche’s SwiftPK platform is a notable example.1 SwiftPK leverages Simulations Plus HTPK technology within Roche’s internal discovery workflows to enable rapid mechanistic PK predictions at scale. By integrating predictive ADME models with automated PBPK simulations, SwiftPK allows discovery scientists to estimate human exposure and projected dose liabilities much earlier in the design cycle. This type of platform exemplifies the broader industry movement toward combining AI-driven prediction with mechanistic modeling rather than viewing the two paradigms as competing alternatives.
Indeed, the future of predictive human dose estimation likely lies in the integration of ML and mechanistic approaches rather than exclusive reliance on either. ML models excel at rapidly predicting individual parameters from molecular structure and identifying complex nonlinear relationships in large datasets. PBPK models excel at integrating those parameters into physiologically meaningful simulations that can account for dose, route, formulation, and interspecies translation. Together, they form a complementary framework capable of supporting more informed and translational decision-making.
This convergence is particularly important as drug discovery increasingly prioritizes “developability” alongside potency. Medicinal chemists are no longer optimizing compounds solely for target affinity; they must also consider whether a molecule can realistically achieve therapeutic exposure at an acceptable human dose. Predictive dose estimation therefore becomes not just a DMPK exercise, but a central component of multi-parameter optimization.
Ultimately, the goal is not simply to predict PK endpoints more accurately, but to better understand how molecular properties translate into clinically viable medicines. Hybrid ML-mechanistic platforms such as HTPK and SwiftPK represent an important step toward that vision, enabling faster, more transparent, and more physiologically grounded predictions of human dose earlier than ever before.
(1) Bassani, D.; Andrews-Morger, A.; Zhang, J.; Docci, L.; Cecere, G.; Pähler, A.; Belubbi, T.; Laye, P.; Shih, I.; Parrott, N. J. High-Throughput Physiologically Based Pharmacokinetic Model for Rodent Pharmacokinetics Prediction Using Machine Learning-Predicted Inputs and a Large In Vivo Pharmacokinetics Data Set. Mol. Pharmaceutics 2026, 23 (3), 1606–1617. https://doi.org/10.1021/acs.molpharmaceut.5c01317.
(2) Van Rompaey, D.; Ray Chaudhuri, S.; Ahmad, M.; Cisar, J.; Van Den Bergh, A.; Ash, J.; Wu, Z.; Bryan, M. C.; Edwards, J. P.; DesJarlais, R.; Wegner, J. K.; Ceulemans, H.; Mitra, K.; Polidori, D. Toward Dose Prediction at Point of Design. J. Med. Chem. 2024, 67 (24), 22282–22290. https://doi.org/10.1021/acs.jmedchem.4c02385.
(3) Bassani, D.; Parrott, N. J.; Manevski, N.; Zhang, J. D. Another String to Your Bow: Machine Learning Prediction of the Pharmacokinetic Properties of Small Molecules. Expert Opin Drug Discov 2024, 19 (6), 683–698. https://doi.org/10.1080/17460441.2024.2348157.
(4) Wang, Y.; Liu, H.; Fan, Y.; Chen, X.; Yang, Y.; Zhu, L.; Zhao, J.; Chen, Y.; Zhang, Y. In Silico Prediction of Human Intravenous Pharmacokinetic Parameters with Improved Accuracy. J. Chem. Inf. Model. 2019, 59 (9), 3968–3980. https://doi.org/10.1021/acs.jcim.9b00300.
(5) Lombardo, F.; Bentzien, J.; Berellini, G.; Muegge, I. In Silico Models of Human PK Parameters. Prediction of Volume of Distribution Using an Extensive Data Set and a Reduced Number of Parameters. JPharmSci 2021, 110 (1), 500–509. https://doi.org/10.1016/j.xphs.2020.08.023.
(6) Kosugi, Y.; Hosea, N. Direct Comparison of Total Clearance Prediction: Computational Machine Learning Model versus Bottom-Up Approach Using In Vitro Assay. Mol. Pharmaceutics 2020, 17 (7), 2299–2309. https://doi.org/10.1021/acs.molpharmaceut.9b01294.
(7) Naga, D.; Parrott, N.; Ecker, G. F.; Olivares-Morales, A. Evaluation of the Success of High-Throughput Physiologically Based Pharmacokinetic (HT-PBPK) Modeling Predictions to Inform Early Drug Discovery. Mol. Pharmaceutics 2022, 19 (7), 2203–2216. https://doi.org/10.1021/acs.molpharmaceut.2c00040.