How Smaller, Explainable AI Outperforms Large Models in Clinical Trials

How Explainable AI Is Transforming Clinical Trials How Explainable AI Is Transforming Clinical Trials

You’ve built a company on the premise that smaller, specialized architectures outperform large models in clinical settings. What does the data show that convinced you to go against the prevailing direction of the field?

I’ve been fortunate to learn from some extraordinary computational scientists across both medicine and artificial intelligence. After completing my doctorate at the University of Southern California, I returned to the University of Toronto, where I had the opportunity to interact with Geoffrey Hinton and members of his research group. That experience gave me a close view of how modern AI was evolving and of the ideas that would ultimately transform the field.

One of the dominant lessons from deep learning has been that performance often improves with scale: more data, more computation, and larger models. That approach has been extraordinarily successful in many areas of AI. But clinical drug development presents a very different mathematical problem. More data is often a luxury you simply do not have.

Clinical trials are inherently small, noisy, and highly heterogeneous. They are a classic example of the problem: relatively few patients, sometimes only dozens or hundreds, measured across hundreds or even thousands of clinical, genomic, imaging, and biological variables. In that setting, simply making the model larger or asking for more data is not enough. You need algorithms specifically designed to extract stable, meaningful structure from limited and complex human data.

When you throw traditional, scale-driven classifiers like random forests, gradient boosting, or deep neural networks at these raw clinical datasets, the data shows they perform the same or worse than a coin flip. Most recently, we benchmarked these standard ML classifiers on a Phase 2 clinical dataset. Without prior stratification, their predictive AUC-ROC scores hovered between an abysmal 0.27 and 0.62. They were overfitting to noise and completely averaging out the therapeutic signal. Honestly, these other methods may work on small data sets but it is the nature of patient populations that my experience has made clear is the actual problem. 

As evidence, when we retrained those same standard classifiers using the compact, trial-native features learned by NetraAI, we saw a performance leap of over 30% in AUC-ROC on psychiatric scale data, and achieved up to 100% accuracy and specificity on MRI-based neuroanatomical models. The data convinced me that the therapeutic signal is there but  it’s diluted by population averages. To find it, you need a sharper lens that respects the unique geometry of the trial itself.

Thirty to fifty patients is a sample size that most data scientists would consider too small to build on. What do these architectures do that prevents them from fitting to noise rather than signal?

If you try to perform brute-force combinatorial searches or run traditional neural nets on 30 to 50 patients, you will overfit 100% of the time, generating false models that look great in a sandbox but completely disintegrate in the real world. I also do not want to send the message that these ultrasmall data sets are ideal. They are not and the results need to be validated repeatedly. The message, however, is that there are signals in these ultrasmall data sets that can be used to augment your understanding of a patient population, and this can have a real economic impact. Ideally, when patient populations start to rise above 80, then atypical methods like those used at NetraMark provide a real advantage for the pharmaceutical companies that utilize them. 

NetraAI bypasses overfitting by completely rethinking the learning substrate. First, we don’t prune variables using crude univariate methods up front, which destroys the synergistic, multivariate interactions that define patient biology. Instead, we map the patients’ baseline characteristics into a new kind of effective high-dimensional patient geometric space where the data is allowed to organize itself dynamically. 

Second, we incorporate a long-range memory mechanism. In our algorithms, every variable has a sustained, logarithmically decaying influence on other variables across generations of the learning loop. This allows us to identify complex, hard-to-find interactions with linear complexity.

Finally, NetraAI is designed to segment the dataset into explainable and unexplainable subsets, which is our biggest mathematical defense against overfitting. We do not force the model to explain every single patient. By allowing the architecture to designate highly heterogeneous or ambiguous patients as “unknown” or “unexplainable,” we drastically reduce the noise floor, allowing us to extract stable, reproducible, and clinically coherent personas from incredibly small cohorts. This makes intuitive sense if you consider it for a moment: in a small data set there will be patients that do not look like other patients, and if you can somehow organize this and see that there are bundles of variables that explain some subgroups better than others, than the explainable ones can be used to improve your next trials. 

LLMs have demonstrated remarkable capability across a wide range of domains. Where does that capability break down when those kinds of models are applied to clinical trial data? Is it a fundamental limitation or an engineering problem that will eventually be solved? 

LLMs are extraordinarily powerful because they learn from enormous amounts of human-generated information. They are very good at recognizing patterns that have appeared repeatedly across language, literature, code and other large datasets. But clinical trial data presents almost the opposite problem. You may have only a few dozen or a few hundred patients, each described by hundreds or thousands of variables, and the clinically important signal may exist only in a small subset of those patients.

That creates a very different mathematical challenge. In a clinical trial, the average behaviour of the population can look remarkably stable while the therapeutic signal is concentrated in a small, highly structured region of the data. A drug may work because of a particular combination of three or four biological or clinical characteristics that occurs in only a fraction of the trial population. If a model is optimized primarily to learn the dominant patterns in the data, those small but important structures can look like noise and disappear into the average.

There is another issue I think is especially important. LLMs inherit an enormous amount of existing human knowledge, including our diagnostic categories, scientific literature and assumptions about disease. That is tremendously useful when you want to reason from what humanity already knows. But in clinical development, sometimes the most valuable discovery is precisely the thing we do not already know: a previously unrecognized patient population defined by an unexpected interaction among variables. A model strongly guided by the labels and concepts we already use can have difficulty discovering structures that sit outside those assumptions.

So I do not think the answer is simply that LLMs need to get bigger. Some limitations are engineering problems and will improve as these systems become better at structured data, numerical information and specialized analytical tools. But there is also a deeper architectural issue. A system built to learn broad statistical regularities from massive datasets is solving a different problem from one designed to discover sparse, explainable structure in small, heterogeneous patient populations. I therefore see the future as complementary: LLMs can interpret literature, generate hypotheses and help scientists reason about results, while specialized AI architectures tackle the core mathematical problem of discovering meaningful patient subpopulations. The most powerful systems will combine the two.

SHAP and LIME are the industry’s answer to explainability. What’s the practical difference between the kind of explanation they provide  and actual causal explanation?

The practical difference is between a post-hoc approximation of a machine’s behavior and the native, transparent logic of the biology itself.

SHAP and LIME are post-hoc explainability methods. They take an existing, highly complex black box model, which has already trained on thousands of variables and is highly prone to overfitting, and perform local linear perturbations around a prediction to approximate how the model reached its conclusion. In high stakes clinical and regulatory settings, these post-hoc approximations are dangerous because they are neither faithful nor stable. They can fluctuate wildly based on how you perturb the space, and they are essentially explaining the model, not the patient.

NetraAI, on the other hand, is built around ex-ante interpretability. The explainability is baked into the architecture itself. Our proprietary Attractor AI utilizes a look-back memory mechanism that gives the algorithms the power of causal inference. Instead of a black box score, our output is a MDS or persona defined by a compact combination of two to four original clinical variables with explicit, readable value ranges.

A biostatistician or a clinician can look at a NetraAI persona (for instance, a specific range of posterior cingulate white matter volume paired with emotional reactivity on a depression scale) and immediately evaluate its clinical plausibility. It represents a native, structural reasoning of the data that you can carry directly into the design of your next protocol, rather than an unverified statistical proxy.

The “no call” option – the decision to output nothing when the signal isn’t strong enough is unusual in a commercial AI product. What was the thinking behind building that into the architecture, and how do sponsors respond when the system returns no result?

The thinking was very simple: in a small clinical trial, forcing an answer can be more dangerous than admitting uncertainty. These datasets are often noisy, heterogeneous and underpowered, so not every patient or every region of the data contains information that can support a reliable conclusion. If an AI system is required to classify everyone, it will eventually start learning noise. We wanted NetraAI to have the scientific discipline to say, “there is not enough evidence here.”

That becomes especially important because clinical trial populations are small. If the system can identify the portions of the population where the signal is genuinely stable and informative, those patients can teach us something very valuable about where the drug may have its clearest advantage. The point is not to claim that we have solved the entire disease from one trial. It is to identify what that trial actually taught us about response, nonresponse and heterogeneity without forcing conclusions where the data do not support them.

As more trials and more patient data accumulate, those reliable pieces can begin to form a mosaic of the true biological heterogeneity of the disease. Importantly, that mosaic can emerge from the data rather than being imposed in advance by existing diagnostic labels or assumptions. For any one trial, however, the immediate goal is more practical: understand as precisely as possible what the trial did right, what it did wrong, and which patient characteristics carried the most meaningful treatment information so that this wisdom can be transported into the next study.

Sponsors sometimes find a no call uncomfortable at first because commercial AI has trained people to expect an answer every time. But once they understand the alternative, the value becomes clear. A confident-looking false signal can send the next trial in the wrong direction and cost millions of dollars. A no call tells you where the evidence stops. In clinical development, knowing what you do not know can be just as important as knowing what you do.

The FDA and EMA are skeptical of AI predictions that can’t be explained. From your direct experience working with regulators, how far has that skepticism advanced, and what are they asking for that most AI tools can’t provide?

The regulatory skepticism is highly advanced, and frankly, it is entirely justified. Global regulators are converging on very strict, explicit mandates for transparency and auditability. Look at the EU AI Act’s traceability requirements, or the FDA’s draft guidelines on AI/ML in drug development.

Regulators aren’t as skeptical of AI’s capability as they are of its defensibility. If an AI tool outputs a black box risk prediction score based on millions of parameters, a regulator can’t audit the decision logic. They can’t evaluate whether the prediction is driven by true biological mechanisms or by historical clinical biases and confounding variables hidden in some massive, poorly curated external EHR dataset.

What regulators are asking for, and what most AI tools can’t provide, is a clear, auditable pathway from the raw data to the subpopulation finding. They are asking for trial-native explanations. Because our personas are defined by a few transparent clinical variables and ranges, sponsors can present them to the FDA as clear, prospective, and statistically sound enrichment strategies.

Pre-specifying inclusion criteria before Phase 3 based on AI-identified responder subgroups is a significant regulatory commitment. What’s the conversation with the FDA when you’re asking them to accept that a subgroup identified by an algorithm should define who gets into a trial?

The conversation with the FDA is actually highly constructive, provided you speak their language of prespecification, multiplicity control, and transportability. If you walk into a regulatory meeting with a post-hoc subgroup that you dredged from your Phase 2 data using repeated, unconstrained statistical tests, the FDA will show you the door. They know that if you search a dataset long enough, you will always find a subgroup that looks statistically significant by sheer chance.

But the conversation changes completely when you show them that the subgroup was identified through a locked, stability-tested algorithm in Phase 2, and that you are transporting that exact, unchanged rule prospectively into the Phase 3 Statistical Analysis Plan (SAP) as a covariate or an alpha-controlled primary population. 

We show them our validation data, like the esmethadone REL-1017 program, where our Phase 2-derived compact rules (such as somatic GI symptoms paired with systolic blood pressure between 108 and 131 mmHg) were applied completely unchanged to two independent Phase 3 datasets. It demonstrated a massive, statistically robust separation that replicated perfectly. When the FDA sees that the algorithm is merely an upstream discovery tool, and that the Phase 3 trial itself is the rigorous, confirmatory test under strict alpha control, they accept it as a highly sophisticated, de-risked design.

We do not want to tell sponsors who should or should not be in their trials. Instead, we can give them a statistically gated way to test where a drug may work best while preserving the opportunity for a broader indication.

Think of statistical alpha as a budget. For example, a sponsor might have an overall alpha of 0.05 and pre-allocate 0.02 to test a NetraAI-defined patient subgroup and 0.03 to the full trial population. If the subgroup succeeds, a prespecified gatekeeping strategy can allow that 0.02 to be recycled back to the overall population test, giving the broader analysis access to the full 0.05.

The result is not simply a narrower trial. It is a way to give the strongest biological signal a chance to establish itself first, while preserving a statistically valid path toward demonstrating that the drug works more broadly.

From a commercial standpoint, sponsors have already written off failed Phase 2 assets. What does it take to reopen that conversation, and who in the organization is the one willing to make that call?

This is actually one of the ideas NetraMark was founded around. When I started the company in 2017, I believed that drug repositioning and resurrection could become a real business model. My view was that some failed clinical assets were not necessarily failed drugs. In some cases, the trial may simply have been unable to identify the patients in whom the drug was working. At the time, the industry was not particularly ready for that idea. Today, with the economics of drug development becoming increasingly difficult, the conversation is very different.

To reopen a failed asset, you first have to challenge the assumption that a negative population-level result tells you everything there is to know about the drug. Clinical trials average together biologically different patients, and heterogeneity, placebo response and differences in disease trajectory can dilute a meaningful treatment effect. The question we ask is: was there a coherent, explainable patient population in which the drug actually demonstrated an advantage?

That can often be investigated relatively quickly using the existing Phase 2 database. In a retrospective analysis, we look for compact, reproducible patient populations where treatment effect was concentrated and then test whether those findings survive rigorous validation. Sometimes the difference between a dead program and a plausible new development strategy can come down to only a few patient characteristics that could be translated into eligibility criteria or prospective stratification in the next trial. The important point is that we are not manufacturing a rescue story after the fact. We are asking whether the original trial contains sufficiently strong evidence to justify another experiment.

The decision to reopen an asset usually has to come from senior clinical and portfolio leadership: the CMO, Head of Clinical Development, or increasingly someone responsible for portfolio strategy and capital allocation. Biostatistics and regulatory teams are essential because any resurrection strategy has to withstand serious methodological scrutiny. But the business decision ultimately belongs to leaders willing to ask whether an asset that has been assigned a value of zero might still contain substantial unrealized value.

That was the opportunity I saw when I founded NetraMark in 2017. Drug development produces enormous amounts of expensive information, even when a trial fails. I believed then, and still believe, that better mathematics could sometimes convert that information into a second, scientifically defensible opportunity for the drug.

This is frontier science. The evidence base is early, the mechanisms aren’t fully understood, and the clinical applications are years away. How do you hold that kind of scientific conviction alongside the commercial realities of building a company in this space?

I am a mathematician and medical scientist by training, and I have spent much of my career thinking about complex systems. But my conviction does not come only from theory. For almost two decades now, I have had the unusual experience of looking at patient populations through the eyes of machine intelligence.

Over thousands of computational experiments, you begin to develop an intuition for these populations. You see again and again that diseases we describe with a single diagnostic label are often composed of very different kinds of patients. You see treatment effects appear in one part of a population and disappear when everything is averaged together. You see the same kinds of interacting structures emerge across different diseases, datasets and therapeutic areas. At some point, those observations stop feeling like isolated computational results and begin to look like a deeper property of human biology.

The mathematics gives me a language for thinking about why. Dynamical systems, interacting variables, attractor-like behaviour and complex state spaces appear at many different scales in biology. Some of the research questions I am interested in are highly exploratory, including questions at the intersection of quantum physics, neuroscience and biological organization. Others are immediately practical, such as understanding why a drug worked for one group of patients and not another in a Phase 2 trial. The scientific horizons may be very different, but my interest in both comes from the same underlying question: how does meaningful structure emerge from complexity?

Building a company requires discipline about where evidence ends and speculation begins. We do not sell frontier hypotheses to pharmaceutical companies. We give sponsors practical, explainable analyses of their clinical trial data that can be validated, challenged by their scientists and used to make better development decisions. The more speculative science belongs in our research program until the evidence earns it a place somewhere else.

So I do not see scientific conviction and commercial discipline as opposites. The conviction tells me where to keep looking. The discipline tells me what we have earned the right to claim today. After nearly twenty years and thousands of experiments, I have developed a very strong intuition about the hidden structure of patient populations. Now my job is to demonstrate that rigorously enough, often enough, that the pharmaceutical industry comes to see what I see.

Bio:

Dr. Joseph Geraci, Co-Founder, Chief Scientific and Technical Officer at NetraMark

Dr. Geraci founded NetraMark, where he led the creation of NetraAI, a mathematically augmented AI platform designed to identify clinically meaningful patient subpopulations within small, complex datasets. His work is grounded in the belief that many clinical trial failures stem not from ineffective therapies, but from poor patient stratification—a challenge his technology was built to address. Over more than a decade of development, his research and leadership helped transform that idea into a platform used to support pharmaceutical and biotechnology companies in designing smarter, more precise clinical trials.

He is affiliated with leading academic and research institutions, including Queen’s University, where he is an Associate Professor in Molecular Medicine; the Center for Biotechnology and Genomics Medicine at the Medical College of Georgia; the University of California, San Diego, where he has served as a Visiting Scientist in quantum computation and neuroscience; and the Centre for Addiction and Mental Health in Ontario, Canada. 

Written by

TechEdge AI

Techedge AI is a niche publication dedicated to keeping its audience at the forefront of the rapidly evolving AI technology landscape. With a sharp focus on emerging trends, groundbreaking innovations, and expert insights, we cover everything from C-suite interviews and industry news to in-depth articles, podcasts, press releases, and guest posts. Join us as we explore the AI technologies shaping tomorrow's world.

View all posts by TechEdge AI →

Grow Your
Brand Visibility

Looking to publish a press release, guest article, interview or podcast? Connect with us.

GET FEATURED
Subscribe

Sign up today for exclusive insights and updates.

Newsletter Signup