Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Evaluation of crowdsourced mortality prediction models as a framework for assessing AI in medicine

Timothy Bergquist, Thomas Schaffter, Yao Yan, Thomas Yu, Justin Prosser, Jifan Gao, Guanhua Chen, Łukasz Charzewski, Zofia Nawalany, Ivan Brugere, Renata Retkute, Alidivinas Prusokas, Augustinas Prusokas, Yonghwa Choi, Sanghoon Lee, Junseok Choe, Inggeol Lee, Sunkyu Kim, Jaewoo Kang, Patient Mortality Prediction DREAM Challenge Consortium, Sean D. Mooney, View ORCID ProfileJustin Guinney
doi: https://doi.org/10.1101/2021.01.18.21250072
Timothy Bergquist
1Sage Bionetworks
2Department of Biomedical Informatics and Medical Education, University of Washington
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Thomas Schaffter
1Sage Bionetworks
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yao Yan
1Sage Bionetworks
3Molecular Engineering and Sciences Institute, University of Washington
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Thomas Yu
1Sage Bionetworks
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Justin Prosser
4Institute of Translational Health Sciences, University of Washington
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Jifan Gao
5Department of Biostatistics and Medical Informatics, University of Wisconsin-Madison
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Guanhua Chen
5Department of Biostatistics and Medical Informatics, University of Wisconsin-Madison
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Łukasz Charzewski
6Proacta
7Division of Biophysics, Faculty of Physics, University of Warsaw
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Zofia Nawalany
6Proacta
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Ivan Brugere
8Department of Computer Science, University of Illinois at Chicago
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Renata Retkute
8Department of Computer Science, University of Illinois at Chicago
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Alidivinas Prusokas
9Department of Plant Sciences, University of Cambridge
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Augustinas Prusokas
10Plant and Molecular Sciences, School of Natural and Environmental Sciences, Newcastle University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Yonghwa Choi
12Department of Computer Science and Engineering, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Sanghoon Lee
12Department of Computer Science and Engineering, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Junseok Choe
12Department of Computer Science and Engineering, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Inggeol Lee
13Department of Interdisciplinary Program in Bioinformatics, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Sunkyu Kim
12Department of Computer Science and Engineering, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Jaewoo Kang
12Department of Computer Science and Engineering, College of Informatics, Korea University
13Department of Interdisciplinary Program in Bioinformatics, College of Informatics, Korea University
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Sean D. Mooney
2Department of Biomedical Informatics and Medical Education, University of Washington
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Justin Guinney
1Sage Bionetworks
2Department of Biomedical Informatics and Medical Education, University of Washington
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Justin Guinney
  • For correspondence: justin.guinney{at}sagebase.org
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Supplementary material
  • Data/Code
  • Preview PDF
Loading

Abstract

Applications of machine learning in healthcare are of high interest and have the potential to significantly improve patient care. Yet, the real-world accuracy and performance of these models on different patient subpopulations remains unclear. To address these important questions, we hosted a community challenge to evaluate different methods that predict healthcare outcomes. To overcome patient privacy concerns, we employed a Model-to-Data approach, allowing citizen scientists and researchers to train and evaluate machine learning models on private health data without direct access to that data. We focused on the prediction of all-cause mortality as the community challenge question. In total, we had 345 registered participants, coalescing into 25 independent teams, spread over 3 continents and 10 countries. The top performing team achieved a final area under the receiver operator curve of 0.947 (95% CI 0.942, 0.951) and an area under the precision-recall curve of 0.487 (95% CI 0.458, 0.499) on patients prospectively collected over a one year observation of a large health system. Post-hoc analysis after the challenge revealed that models differ in accuracy on subpopulations, delineated by race or gender, even when they are trained on the same data and have similar accuracy on the population. This is the largest community challenge focused on the evaluation of state-of-the-art machine learning methods in a healthcare system performed to date, revealing both opportunities and pitfalls of clinical AI.

Competing Interest Statement

The authors have declared no competing interest.

Funding Statement

This work was supported by the Clinical and Translational Science Awards Program National Center for Data to Health funding by the National Center for Advancing Translational Sciences at the National Institutes of Health (grant numbers U24TR002306 and UL1 TR002319). Any opinions expressed in this document are those of the Center for Data to Health community and the Institute for Translational Health Sciences and do not necessarily reflect the views of the National Center for Advancing Translational Sciences, team members, or affiliated organizations and institutions. TB, YY, SM, JG, TY, TS, and JP were supported by grant number U24TR002306. TB, JP, and SM were supported by grant number UL1 TR002319.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

The details of the IRB/oversight body that provided approval or exemption for the research described are given below:

We received an institutional review board (IRB) nonhuman subjects research designation from the University of Washington Human Subjects Research Division to construct a dataset derived from all patient records from the University of Washington Enterprise Data Warehouse.

All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.

Yes

Data Availability

Due to privacy concerns, the clinical data used in this study can not be made available.

Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license.
Back to top
PreviousNext
Posted January 20, 2021.
Download PDF

Supplementary Material

Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Evaluation of crowdsourced mortality prediction models as a framework for assessing AI in medicine
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Evaluation of crowdsourced mortality prediction models as a framework for assessing AI in medicine
Timothy Bergquist, Thomas Schaffter, Yao Yan, Thomas Yu, Justin Prosser, Jifan Gao, Guanhua Chen, Łukasz Charzewski, Zofia Nawalany, Ivan Brugere, Renata Retkute, Alidivinas Prusokas, Augustinas Prusokas, Yonghwa Choi, Sanghoon Lee, Junseok Choe, Inggeol Lee, Sunkyu Kim, Jaewoo Kang, Patient Mortality Prediction DREAM Challenge Consortium, Sean D. Mooney, Justin Guinney
medRxiv 2021.01.18.21250072; doi: https://doi.org/10.1101/2021.01.18.21250072
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Evaluation of crowdsourced mortality prediction models as a framework for assessing AI in medicine
Timothy Bergquist, Thomas Schaffter, Yao Yan, Thomas Yu, Justin Prosser, Jifan Gao, Guanhua Chen, Łukasz Charzewski, Zofia Nawalany, Ivan Brugere, Renata Retkute, Alidivinas Prusokas, Augustinas Prusokas, Yonghwa Choi, Sanghoon Lee, Junseok Choe, Inggeol Lee, Sunkyu Kim, Jaewoo Kang, Patient Mortality Prediction DREAM Challenge Consortium, Sean D. Mooney, Justin Guinney
medRxiv 2021.01.18.21250072; doi: https://doi.org/10.1101/2021.01.18.21250072

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Health Informatics
Subject Areas
All Articles
  • Addiction Medicine (349)
  • Allergy and Immunology (668)
  • Allergy and Immunology (668)
  • Anesthesia (181)
  • Cardiovascular Medicine (2648)
  • Dentistry and Oral Medicine (316)
  • Dermatology (223)
  • Emergency Medicine (399)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (942)
  • Epidemiology (12228)
  • Forensic Medicine (10)
  • Gastroenterology (759)
  • Genetic and Genomic Medicine (4103)
  • Geriatric Medicine (387)
  • Health Economics (680)
  • Health Informatics (2657)
  • Health Policy (1005)
  • Health Systems and Quality Improvement (985)
  • Hematology (363)
  • HIV/AIDS (851)
  • Infectious Diseases (except HIV/AIDS) (13695)
  • Intensive Care and Critical Care Medicine (797)
  • Medical Education (399)
  • Medical Ethics (109)
  • Nephrology (436)
  • Neurology (3882)
  • Nursing (209)
  • Nutrition (577)
  • Obstetrics and Gynecology (739)
  • Occupational and Environmental Health (695)
  • Oncology (2030)
  • Ophthalmology (585)
  • Orthopedics (240)
  • Otolaryngology (306)
  • Pain Medicine (250)
  • Palliative Medicine (75)
  • Pathology (473)
  • Pediatrics (1115)
  • Pharmacology and Therapeutics (466)
  • Primary Care Research (452)
  • Psychiatry and Clinical Psychology (3432)
  • Public and Global Health (6527)
  • Radiology and Imaging (1403)
  • Rehabilitation Medicine and Physical Therapy (814)
  • Respiratory Medicine (871)
  • Rheumatology (409)
  • Sexual and Reproductive Health (410)
  • Sports Medicine (342)
  • Surgery (448)
  • Toxicology (53)
  • Transplantation (185)
  • Urology (165)