Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Identifying Incarceration Status in the Electronic Health Record Using Natural Language Processing in Emergency Department Settings

View ORCID ProfileThomas Huang, View ORCID ProfileVimig Socrates, View ORCID ProfileAidan Gilson, View ORCID ProfileConrad Safranek, View ORCID ProfileLing Chi, Emily A. Wang, View ORCID ProfileLisa B. Puglisi, View ORCID ProfileCynthia Brandt, View ORCID ProfileR. Andrew Taylor, View ORCID ProfileKaren Wang
doi: https://doi.org/10.1101/2023.10.11.23296772
Thomas Huang
1Department of Emergency Medicine, Yale University School of Medicine
BS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Thomas Huang
Vimig Socrates
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
3Program of Computational Biology and Bioinformatics, Yale University
MS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Vimig Socrates
Aidan Gilson
1Department of Emergency Medicine, Yale University School of Medicine
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
BS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Aidan Gilson
Conrad Safranek
1Department of Emergency Medicine, Yale University School of Medicine
BS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Conrad Safranek
Ling Chi
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
BS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Ling Chi
Emily A. Wang
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
4SEICHE Center for Health and Justice, Yale School of Medicine
5Department of Medicine, Yale University School of Medicine
MD MAS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Lisa B. Puglisi
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
4SEICHE Center for Health and Justice, Yale School of Medicine
5Department of Medicine, Yale University School of Medicine
MD
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Lisa B. Puglisi
Cynthia Brandt
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
MD MPH
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Cynthia Brandt
R. Andrew Taylor
1Department of Emergency Medicine, Yale University School of Medicine
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
MD MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for R. Andrew Taylor
  • For correspondence: richard.taylor{at}yale.edu
Karen Wang
2Section for Biomedical Informatics and Data Science, Yale University School of Medicine
4SEICHE Center for Health and Justice, Yale School of Medicine
5Department of Medicine, Yale University School of Medicine
6Equity Research and Innovation Center, Yale University School of Medicine
MD, MHS
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Karen Wang
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Data/Code
  • Preview PDF
Loading

ABSTRACT

Background Incarceration is a highly prevalent social determinant of health associated with high rates of morbidity and mortality and racialized health inequities. Despite this, incarceration status is largely invisible to health services research due to poor electronic health record capture within clinical settings. Our primary objective is to develop and assess natural language processing (NLP) techniques for identifying incarceration status from clinical notes to improve clinical sciences and delivery of care for millions of individuals impacted by incarceration.

Methods We annotated 1,000 unstructured clinical notes randomly selected from the emergency department for incarceration history. Of these annotated notes, 80% were used to train the Longformer-based and RoBERTa NLP models. The remaining 20% served as the test set. Model performance was evaluated using accuracy, sensitivity, specificity, precision, F1 score and Shapley values.

Results Of annotated notes, 55.9% contained evidence for incarceration history by manual annotation. ICD-10 code identification demonstrated accuracy of 46.1%, sensitivity of 4.8%, specificity of 99.1%, precision of 87.1%, and F1 score of 0.09. RoBERTa NLP demonstrated an accuracy of 77.0%, sensitivity of 78.6%, specificity of 73.3%, precision of 80.0%, and F1 score of 0.79. Longformer NLP demonstrated an accuracy of 91.5%, sensitivity of 94.6%, specificity of 87.5%, precision of 90.6%, and F1 score of 0.93.

Conclusion The Longformer-based NLP model was effective in identifying patients’ exposure to incarceration and has potential to help address health disparities by enabling use of electronic health records to study quality of care for this patient population and identify potential areas for improvement.

Competing Interest Statement

The authors have declared no competing interest.

Funding Statement

YCCI Doris Duke Charitable Foundation: Fund to Retain Clinical Scientists Fellowship for Medical Student Research from the Yale School of Medicine Office of Student Research.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

The details of the IRB/oversight body that provided approval or exemption for the research described are given below:

This study was approved by the Yale University Institutional Review Board, which waived the need for informed consent. HIC#: 1602017249

I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.

Yes

Data Availability

All data produced in the present study are available upon reasonable request to the authors.

Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-NC-ND 4.0 International license.
Back to top
PreviousNext
Posted October 12, 2023.
Download PDF
Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Identifying Incarceration Status in the Electronic Health Record Using Natural Language Processing in Emergency Department Settings
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Identifying Incarceration Status in the Electronic Health Record Using Natural Language Processing in Emergency Department Settings
Thomas Huang, Vimig Socrates, Aidan Gilson, Conrad Safranek, Ling Chi, Emily A. Wang, Lisa B. Puglisi, Cynthia Brandt, R. Andrew Taylor, Karen Wang
medRxiv 2023.10.11.23296772; doi: https://doi.org/10.1101/2023.10.11.23296772
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Identifying Incarceration Status in the Electronic Health Record Using Natural Language Processing in Emergency Department Settings
Thomas Huang, Vimig Socrates, Aidan Gilson, Conrad Safranek, Ling Chi, Emily A. Wang, Lisa B. Puglisi, Cynthia Brandt, R. Andrew Taylor, Karen Wang
medRxiv 2023.10.11.23296772; doi: https://doi.org/10.1101/2023.10.11.23296772

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Emergency Medicine
Subject Areas
All Articles
  • Addiction Medicine (349)
  • Allergy and Immunology (668)
  • Allergy and Immunology (668)
  • Anesthesia (181)
  • Cardiovascular Medicine (2648)
  • Dentistry and Oral Medicine (316)
  • Dermatology (223)
  • Emergency Medicine (399)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (942)
  • Epidemiology (12228)
  • Forensic Medicine (10)
  • Gastroenterology (759)
  • Genetic and Genomic Medicine (4103)
  • Geriatric Medicine (387)
  • Health Economics (680)
  • Health Informatics (2657)
  • Health Policy (1005)
  • Health Systems and Quality Improvement (985)
  • Hematology (363)
  • HIV/AIDS (851)
  • Infectious Diseases (except HIV/AIDS) (13695)
  • Intensive Care and Critical Care Medicine (797)
  • Medical Education (399)
  • Medical Ethics (109)
  • Nephrology (436)
  • Neurology (3882)
  • Nursing (209)
  • Nutrition (577)
  • Obstetrics and Gynecology (739)
  • Occupational and Environmental Health (695)
  • Oncology (2030)
  • Ophthalmology (585)
  • Orthopedics (240)
  • Otolaryngology (306)
  • Pain Medicine (250)
  • Palliative Medicine (75)
  • Pathology (473)
  • Pediatrics (1115)
  • Pharmacology and Therapeutics (466)
  • Primary Care Research (452)
  • Psychiatry and Clinical Psychology (3432)
  • Public and Global Health (6527)
  • Radiology and Imaging (1403)
  • Rehabilitation Medicine and Physical Therapy (814)
  • Respiratory Medicine (871)
  • Rheumatology (409)
  • Sexual and Reproductive Health (410)
  • Sports Medicine (342)
  • Surgery (448)
  • Toxicology (53)
  • Transplantation (185)
  • Urology (165)