At three in the morning, during an inexplicably busy shift, it’s not unusual to catch a doctor opening an extra tab on the computer — not to look up the ICD-11 online, but to interact with an artificial intelligence model.
Organizing medical records, summarizing a patient’s history or even checking a diagnostic hypothesis: this kind of AI support, which until recently seemed far off, now occupies a discreet but growing space behind the scenes of medical practice.
But how far can these tools go beyond operational support and assess whether a clinical answer is reliable enough to be used directly in patient care?
A multicenter study recently published in the scientific journal npj Digital Medicine set out to answer exactly that question by placing physicians and AI agents side by side to analyze the same cases.
What did the comparison between physicians and AI agents find?
More than 400 physicians from seven different specialties took part, with varying levels of experience and practice settings. In front of them were real, anonymized clinical cases — and the task of evaluating free-text answers generated by language models.
In parallel, the researchers put another group to the same challenge: AI agents configured to simulate the profiles of these physicians. The question behind the experiment seemed simple but was hard to answer: to what extent can an automated evaluator — in practice, a language-model AI — complement or even replace human judgment?
According to the study, the evaluators’ years of experience and practice setting changed their perception of an answer’s quality — to the point of altering the ranking of the AI models depending on the profile of the group of physicians.
More consistent with one another, the AI agents delivered evaluations that followed the same line. But that consistency had a limit: they lacked something no physician missed — the ability to read the context and the particularities of each case. In other words, this is precisely what keeps them from replacing the physician’s assessment.
Does it make sense to compare physicians and algorithms directly?
For economist Alexandre Chiavegatto Filho, associate professor (livre-docente) at the Faculdade de Saúde Pública da USP and director of the Laboratory of Big Data and Predictive Analysis in Health (LABDAPS), this kind of comparison starts from a flawed premise.
In an interview with Prime Health Report, he says that “testing the performance of algorithms against the performance of physicians is, at bottom, a useless comparison.” And he concludes: “algorithms will never make life-or-death decisions without a responsible human being behind them.”
Editor-in-chief of USP’s Revista de Saúde Pública, Chiavegatto notes that this year the CFM (Brazil’s Federal Council of Medicine) published a resolution reinforcing exactly this principle: the final clinical decision remains the physician’s, because someone has to be accountable for it.
“If the algorithm makes a mistake — because sometimes it does — there has to be a human being who is responsible. If there’s no physician there, it would be the person who developed the algorithm, someone who never even saw the patient, who doesn’t even know the case,” he reflects.
Asked why so many studies insist on this kind of comparison — algorithm versus physician — the researcher says “it’s to show that the algorithm works,” or “that the algorithm has learned something relevant.” After all, if the decision were easy, no one would need an algorithm for it.
Where does AI already genuinely help in an emergency department?
While the academic debate moves forward, AI has already entered the routine of those working on the front lines of emergency care. An emergency physician and professor at the Faculdade de Medicina de Bauru da USP, Júlio Marchini supervised the Emergency Medicine Residency Program at the Faculdade de Medicina da USP (FMUSP) from 2017 to 2023.
He describes to PHR the exact moment the tool comes into the patient encounter. “AI can be used after the medical history and the initial physical exam, as support for clinical reasoning,” he explains. “It helps broaden the differential diagnosis in unusual or nonspecific cases, and it can flag warning signs of severity and patient safety issues.”
According to the physician, the tool’s greatest value shows up in rare cases. “Perhaps, in cases of methanol poisoning, AI could have suggested the suspicion more quickly,” he notes. But he is blunt: “The final decision is always the physician’s. AI doesn’t examine the patient, doesn’t assess clinical nuances directly, and it can hallucinate.”
As for the extra time this verification requires during a shift, Marchini says: “Using AI takes time, but when it suggests a relevant hypothesis, that time is actually gained.” In his view, the reliability of the answers has been improving: “As the models advance, the suggestions are increasingly realistic.”
Should AI be incorporated into the triage routine?
For those still hesitant to bring AI into the triage routine, Marchini’s advice is straightforward: “AI is a tool. It doesn’t replace the medical history, the physical exam, compassionate care or medical judgment. Used with critical thinking, it can help broaden hypotheses and review points that could compromise patient safety. The final decision remains the physician’s.”
It is striking that the study’s findings point, in a way, in the same direction as Marchini’s experience — he had the opportunity to witness the exact moment AI entered the emergency department at ICHC-FMUSP.
Neither the study nor the firsthand account points to any possibility, however remote, that AI will one day replace the physician. The study showed that AI agents deliver efficient evaluations that are aligned in general direction, but without capturing the nuances of human clinical reasoning. Marchini can describe this every day, case by case, in the “heat” of a shift.
A full professor of AI at FSP-USP, Chiavegatto explains why this gap won’t close, even with all the technological progress: it isn’t a bug that gets fixed with an update. The problem isn’t technical but structural. What’s at stake is the very architecture of medical accountability.
For Chiavegatto, this doesn’t mean algorithms have no value — it means that value takes a specific form. “Algorithms in health care will always be an aid to decision-making by the health professional,” he says.
Only from there, he says, does research move to the stage that really matters: measuring “how much this professional gains compared with deciding to act alone, without the help of this algorithm.” For the expert, it’s no longer about studies comparing “physician versus AI,” but “physician alone versus physician with AI assistance.”