AI in Medicine: Accurate, But Not for Everyone
The quality of artificial intelligence (AI) depends on the data used to train the underlying models. The same holds true in medicine. Yet many datasets contain significant gaps, for example when it comes to women. Dr. Dana Mahr and Nora Weinberger from Karlsruhe Institute of Technology (KIT) explain why this could lead to a structural flaw in the system, affecting far more people than previously assumed.
lookKIT: You argue that AI models in medicine exhibit bias against women. In what way?
Nora Weinberger: AI systems learn from existing data. However, clinical trials have long excluded women because the male body was considered the universal norm. Among other reasons, the female hormonal cycle was considered a confounding factor that could complicate research results. When scientists now train an algorithm based on such data, the resulting models reproduce this imbalance across millions of diagnoses and treatment recommendations. Through scaling, AI systems can even amplify existing biases.
Dr. Dana Mahr: In addition, AI models give biased datasets an appearance of technical neutrality. As a result, they seem objective and legitimate. The social and historical assumptions embedded in the data disappear from view, and the bias within the system becomes invisible. At the same time, there is no built-in mechanism capable of correcting these distortions.
Why does this matter?
Weinberger: Because AI in medicine is not a future scenario. It is already a reality. In radiology, algorithms are already scanning X-rays for signs of tumors, and in dermatology, apps are evaluating skin lesions for potential indications of cancer. This development will continue to advance. One particularly promising approach is the creation of digital patient twins.

Issue 2/2026 of the research magazine lookKIT is dedicated to the theme of the Science Year 2026.
To the magazineWhat is a digital twin?
Mahr: It is a digital model of an individual based on general study data and their own medical data: DNA, blood test results, MRI or CT images, and medications. With the help of such models, physicians can simulate how a patient might respond to a particular drug, what side effects may occur, or what could happen if Treatment A is chosen instead of Treatment B. The goal is to support medical decision-making.
Weinberger: The digital patient twin represents an attempt to personalize medicine. It holds the promise of fewer treatments based on guesswork, fewer unnecessary side effects, and greater accuracy. However, it is crucial to recognize that these models can only be as good as the data they are trained on.
So, there are blind spots in the available data?
Weinberger: In the case of a heart attack, an algorithm would reliably recognize the symptoms in men. Female symptoms, however, often present differently, including nausea and fatigue, and may therefore be overlooked by AI systems. The same applies to the most common form of heart failure among women, which remains significantly underrepresented in clinical studies. These diagnostic errors are not marginal phenomena; they strike at the very heart of healthcare.
Mahr: Moreover, this does not only concern the female body. The data also reflects social inequalities. Women in low- and middle-income countries are less likely to have access to digital infrastructure, which means their health data is systematically underrepresented in AI models. In addition, AI-assisted diagnostic systems are less reliable at recognizing symptoms in women and in people of color. So the bias is also social and structural.
Weinberger: Sometimes the bias is not in the data itself, but in the measuring device. Pulse oximeters, for example, which measure blood oxygen levels, systematically show values that are too high for Black people. If such measurements serve as input for a digital twin, the model is built on a flawed foundation. We therefore need to examine the entire data chain, from the measuring device to the AI model.
Are researchers and development teams aware of this problem?
Mahr: Awareness of gender bias remains very limited. In our scoping review of digital patient twins, over a third of the studies did not mention gender even once. Only two out of 31 studies incorporated gender into their design as a socio-biomedical concept. Most of those who develop or use these models do not actively consider whose body actually served as the model for the simulation. This is not malice; it is precisely the blind spot we are talking about.
Weinberger: The problem is real, and we must take it seriously. And it is precisely this discussion that can enable us to systematically highlight these structural inequalities for the first time – before the technology is widely adopted. This is a great opportunity. At the same time, we should be aware that the alternative to AI in medicine is not a healthcare system free of errors.
How can we improve the situation?
Mahr: We need better and more diverse data sets, greater transparency, and a systematic assessment of whether the models work equally well for all groups. Bias testing should not be a voluntary add-on. It should be an integral part of regulatory approval and clinical implementation.
Weinberger: At the same time, we need to involve patients in the development process from the very beginning, rather than just bringing them in as test subjects at the end. They can help identify which experiences, symptoms, everyday realities, and gaps in care are missing from the data sets. In other words, we need to fill gaps in our knowledge that would otherwise remain invisible without this perspective. The uncomfortable truth is that no one can stay out of this. Development teams need to adopt an interdisciplinary approach from the outset, and ministries and research agencies need to establish diversity and fairness as requirements for funding eligibility.
On the other hand, one could argue that healthcare today is better than ever. Why should we dwell on such things?
Weinberger: Because otherwise existing inequalities in healthcare will not simply persist; they will be systematically reinforced. That is the real risk. No one expects perfection, but there is a difference between a system that occasionally makes mistakes and a system that consistently performs worse for certain groups.
Mahr: Just as we expect medications to be effective and not cause harm, we should also expect an algorithm to perform its tasks equally well for everyone. For me, that would be the minimum requirement.
Isabelle Hartmann, October 1, 2026
Translation: Dipl.-Übers. Veronika Zsófia Lázár
Bias
Bias describes a systematic distortion or error in perception or judgment. There are many types of bias in AI systems.
Representation bias occurs when parts of subpopulations are missing or underrepresented in the training data, resulting in unequal treatment. However, bias in AI can also arise from the design, training, evaluation, or use of the models.
