new publication

Dresden researchers develop on-premise medical AI agent for reliable clinical decision support

AI agents could reliably support diagnoses and clinical decision-making in the future – provided that sensitive health data are protected and clinicians can assess how reliable individual AI-generated results are. Researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TU Dresden and Dresden University Hospital have developed an on-premise medical AI system that addresses both challenges. Their findings have been published in the journal Nature Medicine.

Li Zhang, Georg Wölflein, Dyke Ferber, Junhao Liang, Zunamys I. Carrero, Xuewei Wu, Julien Vibert, Jan Clusmann, Lino Möhrmann, Elena E. Möhrmann, Catharina Wichmann, Fabian Wolf, Tim Lenz, Jakob Nikolas Kather: On-Premise Medical AI Agents for Reliable Clinical Decision-Making, Nature Medicine, 2026.
Graphic showing a hospital, computer rack and laptop on the left, on the right side two doctors are checking the output from this ai agent data from the laptop

Large language models are increasingly capable of handling complex medical tasks. However, their use in clinical practice faces two fundamental challenges: sensitive patient data need to remain under the control of the respective institution rather than being shared with external services, and clinicians need to be able to assess reliability of AI-generated diagnoses. A research team led by Prof. Jakob N. Kather in Dresden developed a fully on-premise diagnostic AI system and evaluated it in a simulation environment in which two AI agents interact with each other: one takes on the role of the physician, the other that of the patient. Using this setup, the researchers investigated how reliably medical AI systems can make diagnostic decisions. The system is built upon the MIRA AI agent introduced in June 2026.

High diagnostic accuracy in standardized tests

The AI agent was tested using standardized clinical cases covering a range of conditions, including appendicitis, cholecystitis, pneumonia, pulmonary embolism and urinary tract infections. The cases were based on anonymized electronic patient data, including diagnoses, laboratory values, medications and physical examination findings. The physician AI agent could ask follow-up questions and request clinical findings and laboratory results. Based on this information, it generated a diagnosis and an accompanying rationale. The best locally operated AI model reached the correct diagnosis in approximately 90 percent of cases in one benchmark and 84 percent in the other. For 181 randomly selected cases, physicians additionally reviewed the diagnoses. The automated evaluation and the physicians’ consensus agreed in more than 90 percent of cases.

Consistent answers are the strongest indicator of a correct diagnosis

One of the study’s key questions was how to determine whether an AI-generated decision can be trusted. To address this, the researchers investigated several indicators of diagnostic correctness. The most informative was whether the AI arrived at the same diagnosis when processing the same case repeatedly. The more stable the diagnosis remained across multiple runs, the more likely it was to be correct. The model’s internal likelihood scores were less useful for predicting whether a diagnosis was correct. This was also evident in a stress test: when the researchers removed reliable information from the system, diagnostic accuracy dropped substantially. Although the AI diagnoses became more frequently incorrect, the model’s probability-based signals did not reflect this.

Medical AI agents to support clinical practice

Based on their findings, the researchers propose a possible framework for collaboration between clinicians and AI. Rather than either fully trusting an AI system or requiring every decision to be verified, cases with stronger reliability signals could be distinguished from more uncertain ones. Uncertain cases would always be deferred to medical professionals for review.

Prof. Jakob N. Kather, senior author of the study and Professor of Clinical Artificial Intelligence at TU Dresden, explains: “Our goal is an AI agent with selective autonomy. These systems should support clinicians in decision-making, but never take over completely. It is therefore essential that AI-generated results are understandable to clinicians and that the systems clearly indicate when their outputs are uncertain. Responsibility for diagnosis and treatment will always remain with humans.”

Local deployment enables institutional control

A second focus of the study is technical control over the AI system. The AI models and all data processing ran entirely within locally operated infrastructure. This gives medical institutions greater control over where data are processed and which AI models are used. On-premise infrastructure provides the technical basis for managing data protection, model versions, access rights and monitoring processes within the respective institution. Kather also sees this as a prerequisite for advancing medical AI responsibly in Europe: “Calls to slow down AI development are, in my view, heading in the wrong direction. In medicine, we are still at the very beginning: AI systems such as our agents are only just starting to show what is possible, and patients have so far seen very little benefit. We need to move faster, not slower. What matters is that we develop these systems in a way that is safe, transparent and preserves data sovereignty – for example, by running them entirely within the hospital. But for this to work, Europe also needs to play a role in the foundational technology itself, including the language models, rather than simply building on technologies developed elsewhere.”

Next steps towards clinical use

The study also highlights questions that need to be addressed before any potential clinical deployment. The researchers found indications of differences in diagnostic accuracy between patient groups. It remains unclear why simulated cases involving older patients in particular showed poorer results, and this will require further investigation. In addition, the findings are based on retrospective simulations using existing clinical data rather than on testing during routine clinical care.

“Our results show that we have taken an important step towards more reliable medical AI agents. Next, we want to investigate how well this approach performs under realistic conditions, with clinicians involved in the process and across a broader range of diseases and clinical data,” says Li Zhang, first author of the publication and a researcher in Prof. Kather’s team.

Another focus is efficiency. Because estimating reliability currently requires repeated runs for each case, the team is working to reduce the computational cost without compromising the quality of the results. In addition to the Dresden researchers, scientists from the National Center for Tumor Diseases (NCT) Heidelberg at Heidelberg University Hospital also contributed to the study.

Share this Post

More News

Open Doors at Saxon Ministries
Skip to content