An objective framework for evaluating unrecognized bias in medical AI models predicting COVID-19 outcomes

Hossein Estiri; Zachary H Strasser; Sina Rashidian; Jeffrey G Klann; Kavishwar B Wagholikar; Thomas H McCoy; Shawn N Murphy

doi:10.1093/jamia/ocac070

An objective framework for evaluating unrecognized bias in medical AI models predicting COVID-19 outcomes

J Am Med Inform Assoc. 2022 Jul 12;29(8):1334-1341. doi: 10.1093/jamia/ocac070.

Authors

Hossein Estiri^{1

2}, Zachary H Strasser^{1

2}, Sina Rashidian^{3

4}, Jeffrey G Klann^{1

2

5}, Kavishwar B Wagholikar^{1

2}, Thomas H McCoy⁶, Shawn N Murphy^{1

5

7

8}

Affiliations

¹ Laboratory of Computer Science, Massachusetts General Hospital, Boston, Massachusetts, USA.
² Department of Medicine, Massachusetts General Hospital, Boston, Massachusetts, USA.
³ Verily Life Sciences, Boston, Massachusetts, USA.
⁴ Massachusetts General Hospital, Boston, MA 02114, USA.
⁵ Research Information Science and Computing, Mass General Brigham, Somerville, Massachusetts, USA.
⁶ Center for Quantitative Health, Massachusetts General Hospital, Boston, Massachusetts, USA.
⁷ Department of Biomedical Informatics, Harvard Medical School, Boston, Massachusetts, USA.
⁸ Department of Neurology, Massachusetts General Hospital, Boston, Massachusetts, USA.

Abstract

Objective: The increasing translation of artificial intelligence (AI)/machine learning (ML) models into clinical practice brings an increased risk of direct harm from modeling bias; however, bias remains incompletely measured in many medical AI applications. This article aims to provide a framework for objective evaluation of medical AI from multiple aspects, focusing on binary classification models.

Materials and methods: Using data from over 56 000 Mass General Brigham (MGB) patients with confirmed severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), we evaluate unrecognized bias in 4 AI models developed during the early months of the pandemic in Boston, Massachusetts that predict risks of hospital admission, ICU admission, mechanical ventilation, and death after a SARS-CoV-2 infection purely based on their pre-infection longitudinal medical records. Models were evaluated both retrospectively and prospectively using model-level metrics of discrimination, accuracy, and reliability, and a novel individual-level metric for error.

Results: We found inconsistent instances of model-level bias in the prediction models. From an individual-level aspect, however, we found most all models performing with slightly higher error rates for older patients.

Discussion: While a model can be biased against certain protected groups (ie, perform worse) in certain tasks, it can be at the same time biased towards another protected group (ie, perform better). As such, current bias evaluation studies may lack a full depiction of the variable effects of a model on its subpopulations.

Conclusion: Only a holistic evaluation, a diligent search for unrecognized bias, can provide enough information for an unbiased judgment of AI bias that can invigorate follow-up investigations on identifying the underlying roots of bias and ultimately make a change.

Keywords: COVID-19; bias; electronic health records; medical AI; predictive model.

MeSH terms

Artificial Intelligence
COVID-19*
Humans
Reproducibility of Results
Retrospective Studies
SARS-CoV-2