FDA is expanding non-animal methods while ARPA-H funds human drug-safety models. Public materials omit runtime, hardware, and deployment cost.

On July 14, 2026, researchers from Johns Hopkins and collaborating institutions published a virtual-patient model for liver cancer in PNAS. The model combines whole-body quantitative systems pharmacology with an agent-based simulation of cells inside a tumor. Its virtual patients reproduced response rates reported in clinical trials of cabozantinib and nivolumab, including their use in combination.
The result is narrower than the phrase "virtual patient" may suggest. It models one cancer, a defined set of therapies, and selected features of the tumor microenvironment. The researchers say it needs further validation before clinical use. Still, it shows what an in silico trial can do when mathematical physiology, machine learning, and patient data meet in one workflow.
It also leaves out a number that will matter if these systems move from research projects into routine drug development: the amount of compute required to run them.
The paper describes its calibration method, reports simulations of 500 virtual patients, and compares the results with independent clinical and spatial-molecular data. It does not report elapsed runtime, processor or accelerator type, or the resources required to generate the virtual cohort. That omission does not weaken the reported biology. It does make the system's operating cost and portability hard to judge.
Those questions are becoming more important because the US government is now paying to develop this class of model.
The FDA published its roadmap for reducing animal testing in April 2025. It began with monoclonal antibodies and encouraged developers to submit evidence from New Approach Methodologies, including organoids, organ-on-chip systems, in vitro assays, and computational toxicity models.
The agency's language matters. The roadmap creates a path to reduce or replace some animal-testing requirements where an alternative can provide equivalent or better evidence. It does not declare animal studies obsolete, and it does not pre-qualify the models that might replace them. On April 20, 2026, the FDA reported completing its first-year roadmap goals, including new guidance and infrastructure for evaluating non-animal methods.
ARPA-H is supplying the research money. Its Computational ADME-Tox and Physiology Analysis for Safer Therapeutics program, known as CATALYST, offers up to $125 million over four and a half years. The program funds teams building models of how drugs move through the body, how they affect organs, and where toxicity may appear before a candidate reaches a human trial.
These are milestone-based contract ceilings, not lump-sum checks. Deep Origin leads a consortium with an award of up to $31.7 million. The Charles Stark Draper Laboratory leads a separate human-data-stack project with an award of up to $56.2 million. Both awards began on September 30, 2025. Other prime awardees are working on cardiac safety, organ toxicity, antibody pharmacokinetics, and the data practices needed to make AI-supported evidence acceptable in regulated development.
CATALYST is funding the construction and validation of these systems. The public award records describe their intended capabilities but do not specify expected runtime, hardware, or cost per simulated drug candidate. At this stage, ARPA-H is buying research rather than a finished production service. Even so, the missing operational numbers make it difficult to compare one approach with another or estimate what adoption would require.
"Virtual human" is convenient shorthand for several kinds of models operating at different biological scales. They do not all use the same methods, data, or machines.
Some drug-discovery workflows begin with molecular simulation. Hybrid quantum mechanics and molecular mechanics, usually shortened to QM/MM, can model an enzyme's reactive region with quantum methods while treating the surrounding protein and solvent classically. A recent MiMiC project ran this kind of calculation on the JUWELS CPU supercomputer to study a mutant enzyme associated with glioma. That is established supercomputing work, but it is one tool used for certain molecular problems rather than the universal first layer of every virtual patient.
Physiologically based pharmacokinetic models work at a different scale. They use mathematical representations of organs and blood flow to estimate where a substance travels and how its concentration changes over time. Empa researchers combined a PBPK model with Bayesian fitting and multivariate regression to predict nanoparticle distribution in mice. Their ACS Nano paper drew from 10 published studies containing 18 biodistribution experiments.
The model's limits are useful. Its generated concentration curves achieved an overall adjusted R-squared of 0.43, and 65% of predicted points fell within threefold of the observed values. Performance varied sharply between nanoparticle cases. The researchers shared their code and called for larger datasets, nonlinear models, and more experimental validation. They did not present the system as a general replacement for mouse studies.
The Johns Hopkins liver-cancer model moves up another scale. The study calibrated a spatial quantitative systems pharmacology model using clinical and molecular data. Its simulated response rate was 17.1% for nivolumab, compared with 15% in the referenced trial, and 21.6% for nivolumab plus cabozantinib, compared with 18.1% in the trial cohort. The model also reproduced a spatial feature found in patient tissue: fibroblasts separating immune cells from cancer cells in non-responding tumors.
That is credible validation for a defined use. It is not evidence that a full human physiology model has arrived. The distinction is easy to lose when "virtual patient," "digital twin," and "virtual human" appear in the same funding announcement.
At the clinical-trial level, Unlearn's PROCOVA method uses a prognostic score derived from historical patient data as a covariate in a randomized trial. The European Medicines Agency qualified PROCOVA in September 2022 for a defined context involving Phase 2 and Phase 3 trials with continuous outcomes. The opinion says the method can increase statistical power or support a smaller sample size when its assumptions hold.
The EMA also spells out the constraints. The historical data must resemble the future trial population, the prognostic model must be validated, and sample-size calculations must account for uncertainty. According to Unlearn's account of subsequent FDA feedback, CDER viewed PROCOVA as a special case of standard covariate adjustment and said future use should be discussed with the appropriate review division for the specific product and trial.
Regulators have therefore accepted a scoped statistical procedure. They have not qualified a connected chain of molecular, organ, disease, and trial simulations as a virtual human.
Hardware disclosure alone does not make a model reproducible. Reproduction depends on access to code and data, software versions, solver settings, numerical tolerances, random seeds, and a documented execution environment. Hardware can still affect numerical behavior, and it determines whether another lab can run the work at a practical speed.
Runtime and resource consumption answer a different set of questions. How many candidate compounds can a system evaluate in a week? Can a university lab run it, or does it require a national facility or hyperscale cloud account? Does generating 500 virtual patients take minutes, days, or months? How much does each run cost, and how does the model scale when the cohort grows?
Biomedical simulation papers can publish those numbers. A 2023 whole-heart digital-twin study reported about 11.3 hours to simulate one heartbeat on eight Nvidia A100 GPUs, producing roughly 8 terabytes of data. A newer cardiac electrophysiology solver reported 23 minutes for a coarse patient-specific mesh and 302 minutes for a fine mesh. The same study ran 512 simulations concurrently across 128 compute nodes while calibrating a patient-specific model.
Those disclosures do not reveal patient data or proprietary model weights. They tell other researchers what the work costs to reproduce and whether the approach can support a larger trial.
The same gap appears in other simulation-heavy fields. At ICRA 2026, Nvidia reported robotics policies trained across millions of simulated trajectories but did not publish the GPU-hours or cluster sizes behind those runs, as Supercomputing News reported in May. Trajectory counts describe the dataset. They do not describe the infrastructure that produced it.
Drug-development models present a more complicated case because their workloads differ so widely. A regression-based PBPK model may run on a workstation. A QM/MM calculation can occupy thousands of CPU cores. A spatial model with hundreds of virtual patients may sit somewhere between them. Treating all three as generic "AI compute" hides the engineering differences that determine cost.
This is where the convergence of simulation and AI becomes a useful supercomputing story. Scientific models and learned models increasingly share infrastructure, software libraries, and accelerator queues even when their mathematics remains distinct. Funding programs and system vendors have responded by presenting simulation capacity as AI infrastructure, a shift examined in Why Simulation Supercomputers Are Being Pitched as AI Infrastructure.
The public record supports a measured conclusion. The FDA is creating pathways for non-animal evidence. ARPA-H is funding model development before those models are ready for broad regulatory use. Researchers have produced credible results in nanoparticle pharmacokinetics, liver-cancer response, and statistical trial design, each within a limited context.
The missing compute data does not invalidate those results. It leaves an important part of the deployment case unanswered. A model can be scientifically sound and still be too slow, expensive, or specialized for routine use. Conversely, a cheap model has little value if it cannot reproduce biology well enough to support a decision.
The next useful disclosures from CATALYST teams will include more than accuracy metrics. Runtime by workload, hardware configuration, cost per simulation, scaling behavior, and reproducible software environments would show whether these models can move beyond well-funded research programs. They would also let drug developers compare computational evidence with the time and expense of the studies it is meant to reduce.
For now, the phrase "virtual human" describes an ambition. The validated systems underneath it remain separate, narrow, and unevenly documented. Their science is becoming easier to evaluate. Their operating requirements are not.