Two peer-reviewed journals, three weeks apart. Opposite answers on whether general or clinical AI performs better. Both contained flawed benchmarks…and there’s no referee.
This editorial by Protege co-founder and Chief Scientific Engy Ziedan was originally published in Healthcare Business Today.
Today, more than 200 million people use ChatGPT every week for healthcare questions. Almost every physician in the United States has referred to AI for clinical guidance this year.
At the start of 2024, none of these tools existed. Just two years later, tens of thousands of doctors ask AI millions of questions a week. Already, AI has forever changed how medicine is practiced.
But who grades the AI your doctor uses? How do we judge which AI is the most reliable in the doctor’s office?
This question is no longer abstract, and the answer has consequences. On June 12, a paper in Nature Medicine dropped a bombshell: general-purpose models like ChatGPT, Claude, and Gemini appeared to give better clinical guidance than healthcare-specific AI models like OpenEvidence, which is used by 65% of US physicians.
AI commentators jumped on this announcement, noting how unprecedented this was, with some indicating that this was a sign that general-purpose models were going to make vertical-specific models obsolete. A whole industry segment of well-capitalized vertical AI builders became rattled at the potential market response if general-purpose models truly beat the application-specific models.
But three weeks later, a new paper appeared in the Journal of the American Medical Association (JAMA) that concluded the exact opposite: 149 blinded medical experts spanning 30+ specialties consistently graded OpenEvidence’s healthcare model more highly than Claude, GPT, and Gemini across the board on real clinical questions: on accuracy, utility, source quality, and verifiability.
But neither study reported whether each AI model’s “settings” impacted the results. The second study noted that longer answers scored better than shorter ones. Shorter answers often reflect the default settings of models, but these settings can be changed to allow for longer reasoning or thinking.
So did OpenEvidence’s specialized model beat OpenAI’s GPT-5.5 on default settings? Even if the authors set every model to their most capable settings, are those settings cost-effective enough for actual patient care use? We don’t know the answer to either. The authors are not the AI model builders, so it’s understandable why this was not a focus. But this matters for actual care center deployment!
These two contradictory papers — both published in well-respected scientific journals — point towards a simple yet vexing problem: Who can fairly and independently assess the use of AI models in healthcare?
In some contexts, AI resembles a new treatment or medical device. It influences diagnosis, communication, triage, documentation, prioritization, referrals, and ultimately decisions that affect human health and well-being. But unlike a physician, AI is not educated through accredited institutions, examined through standardized tests, licensed by professional boards, or vetted through traditional hiring processes.
These factors point to the same conclusion: healthcare AI requires careful independent evaluation that the market is currently missing.
History tells us that we have been here before. Every single healthcare technology of the past century has resulted in regulation when it became popular.
The first option is classic government regulation. Take drugs and the history of the FDA. In 1938, Congress passed the Federal Food, Drug, and Cosmetic Act, which came after the Elixir Sulfanilamide disaster, where a physician tried to dilute a sulfanilamide drug using a toxic and untested solvent. One hundred and seven people died, many of them children. The public demanded oversight.
Today, of the 1,500 or more drugs that are submitted to the FDA every year, only 40-50 become new active ingredients that are ultimately launched into the market.
Take another example: Medical schools do not accredit themselves or determine who is fit to practice. The American Medical Association (AMA), along with the Liaison Committee on Medical Education (LCME) accredits doctors in the United States and Canada. These organizations determine educational curricula and the number of residency seats available to medical school graduates.
Or a third: the Relative Value Scale Update Committee (RUC). The Centers for Medicare and Medicaid Services relies on an independent committee of physicians, who decide what resources are required for any given surgery and what the government payment for each surgery should be. As a result, a government-recognized committee process determines the price for a knee replacement versus a cardiac bypass, rather than the market.
All these are independent or government-led interventions to provide oversight to the market, with varying degrees of effectiveness. The problem here is that the pace of innovation in AI is much, much faster than anything we have seen, to the point where the government may not be able to move quickly enough to oversee it. The irony, and reflecting the central tension of this moment, is that the FDA itself is soliciting feedback on how AI-enabled technologies might improve the efficiency, speed, and quality of decision-making in early-phase clinical trials.
Academic oversight is the other backstop for independent evaluation. However, academia is notoriously slow with peer review, taking months or longer to review submissions. We’ve also seen what happens when we try to speed it up. During the COVID-19 pandemic, incentives to speed up submission processing saw a decrease in submission quality even as the volume of submissions rose.
As an institution, academia is also rarely truly an arm’s length removed. Research examining 12,682 COVID-19 papers indexed on PubMed found that eight percent were accepted for publication on the very same day they were submitted. Even more concerning, they found serious editorial conflicts of interest in 43 percent of these papers. In many cases, authors submitting manuscripts were themselves editors at the very journals to which they submitted.
If AI model builders cannot judge themselves, the government cannot solve the oversight problem quickly, and academia cannot solve it fairly, then who can step in?
The problem is not necessarily malicious intent. The problem is that medical science is filled with choices. Which patients are included in the study? Which questions are asked in the benchmark? Which grading rubric is selected? Small changes in these decisions can materially alter performance estimates, which influence downstream decisions in the care room.
Take HealthBench, for example, which is regarded as the state-of-the-art benchmark today in healthcare AI. Of the 5,000 cases in the benchmark, 322 involve mental health, and fewer than 100 focus specifically on prenatal depression. These are benchmark design choices some may agree or disagree with — and models are evaluated and deployed based on these choices.
If healthcare AI scales without independent scrutiny by those who understand all aspects of AI model development and deployment, eventually mistakes will be made. There will be a highly visible failure: a harmful clinical recommendation, a misdiagnosis, or some other scandal salient enough to erode public trust.
AI is not slowing down. Unlike drugs, AI evolves on timescales that our existing institutions were never designed to handle. Clinical trials can take years. Peer review can take months. Regulatory guidance can take even longer. In AI, three months is like a lifetime.
We need an independent arbiter of truth for evaluating AI models for healthcare that has deep expertise in all angles of healthcare AI deployment: The AI model training and settings. The health data exposure and statistical methodology applied to healthcare. The actual in-office medical practice and nuances of patient care.
The floodgates are open. AI is already in the exam room, the ICU, and the oncology ward. We know models hallucinate. We know benchmarks can be gamed. We know the people grading the tests are the same people selling them. The only question left is whether we build an independent referee before the first high-profile crisis — or after.
P.S. – If you’re building in this space and care deeply about how healthcare AI is going to be safely deployed, we’d love to hear from you. Drop us a line at https://withprotege.ai/contact.