Understanding bias in medical AI systems
AI is revolutionising health care, but bias in medical AI risks reinforcing existing health inequity.
AI medical tools are contributing to medical innovations and clinical efficiencies globally, and experts recognise their potential to revolutionise clinical medicine.
Meanwhile, there are growing concerns about the bias in medical AI technologies can exhibit and their impact on health care and treatment.
Associate Professor Sanjay Warrier, a Sydney-based breast oncology and oncoplastic surgeon, says AI is steadily entering the screening, imaging and diagnosis he relies on. “BreastScreen NSW now uses machine-reading technology called Lunit Insight MMG to support radiologists reading mammograms".
But Professor Warrier adds, “AI medical tools can perform unevenly when they've been trained on data that doesn't represent every patient. Breast density, age and ethnic background all matter here.”
As a Harvard Data Science Review paper reports, “A biased world produces biased data.” These biases often reflect social and healthcare inequalities, with experts often referring to the human element involved, as AI relies on data generated, collected, recorded, and labelled by humans.
The latest research recommends that clinicians try to identify, assess, and mitigate these biases when deploying AI-enabled tools in their practice to ensure patient safety and to prevent perpetuating inequities in health care and outcomes, especially for those already marginalised by old biases.
What is bias in medical AI?
A 2026 paper in the Journal of the American College of Emergency Physicians (JACEP) defined bias as a systematic flaw in a decision-making process that results in unfair or unintended outcomes that can be inadvertently embedded in AI algorithms or training data. The JACEP paper reports that, perhaps the most concerning, is automation bias: the tendency for clinicians to over-rely on or excessively trust AI-generated recommendations, even when its outputs are questionable or contradict their own clinical judgment.
In just one example of bias the paper refers to, a US landmark study demonstrated how an algorithm using health care costs as a proxy for health needs systematically disadvantaged African American patients. The algorithm allocated fewer resources to African American patients based on their healthcare spending, even though spending reflected unequal access to care rather than true need.
As the CSIRO reports, biased algorithms can delay diagnoses, overlook symptoms, and reinforce longstanding inequalities.
An Australian Journal of General Practice paper reports that in Australia, evidence shows that implicit bias, the unconscious attitudes health care professionals may hold, tends to underestimate Aboriginal and Torres Strait Islander people’s experience of pain, leading to less comprehensive assessment and subsequently delays in treatment because of mismanagement.
Associate Professor Sonika Tyagi leads the Digital Health & Bioinformatics research lab in the School of Computing Technologies at RMIT University. She says that beyond data-level issues, other bias types include measurement bias (when a proxy feature fails to capture the intended real-world construct) and interpretation bias (when model outputs are judged or applied inconsistently across contexts).
How bias in medical AI occurs
According to the JACEP paper, bias can manifest at multiple stages of the AI lifecycle, from data collection and algorithm design to clinical implementation and human interaction.
Professor Tiyagi says one of the major causes is the training data itself.
“If the training data over- or under-represents certain groups, languages, or viewpoints, the model absorbs those imbalances as "normal." Human feedback used to fine-tune the model may also introduce annotator bias.”
“LLMs learn by identifying statistical patterns in massive amount of training datasets, then use those patterns to predict likely responses to a query. The model optimises for statistically frequent patterns rather than factual correctness. In this way, it can amplify existing biases rather than simply reproducing them.”
“Since the model has no independent way to verify truth, it treats frequent patterns as reliable, making bias hard to detect without deliberate auditing.”
AI bias mitigation
Scientia Professor of Information Systems and Technology Management, Manju Ahuja, from the UNSW Business School defines AI bias as a social-technical issue. This means human bias needs to be considered and mitigated in the design of AI systems and throughout the AI product lifecycle.
Fairness metrics built into the design stage
Professor Tiyagi says awareness of historical and representation biases should inform decisions at the design stage. Her recommendation is in line with findings in a review paper in the Journal American Medical Informatics Association Open, which also recommends ongoing monitoring throughout the model’s lifecycle for effectiveness.
In 2025, the Australian Commission on Safety and Quality in Health Care (the Commission) released its Pragmatic AI Guidance for Clinicians, which highlights the importance of AI systems being trained on data that reflects the diversity of the patient population in which they will be used.
Professor Liz Marles, Commission Clinical Director and GP, says, “This will ensure Aboriginal and Torres Strait Islander communities and other population groups are appropriately represented.”
Prompting AI medical tools for greater transparency
Professor Ian Scott, Clinical Consultant in AI, at the Digital Health and Informatics Directorate, Metro South Hospital and Health Service, says methods to mitigate biases include context-aware prompting, in which users request the model to display its reasoning steps, which helps make biased associations more transparent.
“Another method is using retrieval-augmented generation (RAG) whereby the model is required to retrieve and cite original and authoritative source documents underpinning its reasoning (such as journal publications or clinical guidelines).”
Clinical judgement remains most important safeguard
Professor Scott says that these strategies may help reduce bias risk but are unlikely to eliminate bias.
“Clinicians must continue to apply their own judgement in interpreting model outputs for individual patient scenarios and seek verification when outputs look and sound implausible or counter-intuitive.”
As per AHPRA guidelines, practitioners must apply human judgment to any output of AI.
Professor Marles says, “It is best to view AI as a prompt or a tool, rather than an authority. If the outputs from an AI tool such as an AI scribe seems unexpected or inaccurate, further investigation is warranted.
“Generative AI has the potential to add new information to our medical notes and letters, so we always need to review and evaluate these notes for bias and accuracy.”
Warrier says screening is a good example of getting this right. “The technology supports the radiologist, but a specialist still reads every set of images and makes the call. I treat what the AI gives me as one input, not the final answer."
Clinical judgement at the centre of design
Professer Tiyagi says, “Equally important is who is in the room when these design decisions are made, since a narrow set of perspectives can leave blind spots unaddressed regardless of technical safeguards.”
Dr Warrier tells InSight+ that his practice is building its own AI tool, using Claude, to mirror the way a multidisciplinary team works through a clinical decision together. “It's still early, but it's a good example of putting clinical judgement at the centre of the design from the start.”
Key takeaways
- AI systems learn from datasets that come from human input and decisions. This data may reflect historical inequalities in care, underrepresentation of certain groups, and biased clinical decision-making.
- AI systems absorb these biases, with the risk of perpetuating these further in medical care for those already marginalised.
- AI systems should be trained on data that most accurately reflects the demographic diversity of the populations for which they will serve.
- Continuous monitoring, auditing and updating of AI systems throughout the AI lifecycle is important to identify and mitigate biases.
- Professor Marles says, “Clinicians should remain alert to and investigate outputs that are inconsistent with the clinical context or their professional judgement.”
- As Harvard University suggests, involving a range of stakeholders, such as patient advocates, in the design of AI systems can help ensure a broader perspective and they are ethically sound and culturally sensitive.
Tracey is a freelance health and medical writer, with a Bachelor of Health Promotion, who enjoys writing to help build a bridge of knowledge, ideas, issues and news between the community and health professionals. She hopes to contribute to improved health literacy, and equity in healthcare. Her work has appeared in publications including Retail Pharmacy Magazine, Retail Pharmacy Assistants Magazine, SBS Food, The Age and The Guardian.
Subscribe to the free InSight+ weekly newsletter here. It is available to all readers, not just registered medical practitioners.
More from this week
You may also like
VIEW MORENewsletters
Subscribe to the InSight+ newsletter
Immediate and free access to the latest articles
No spam, you can unsubscribe anytime you want.
By providing your information, you agree to our Access Terms and our Privacy Policy. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.