Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Cohort profile: St. Michael’s Hospital Tuberculosis Database (SMH-TB), a retrospective cohort of electronic health record data and variables extracted using natural language processing

View ORCID ProfileDavid Landsman, Ahmed Abdelbasit, View ORCID ProfileChristine Wang, Michael Guerzhoy, Ujash Joshi, Shaun Mathew, Chloe Pou-Prom, David Dai, Victoria Pequegnat, Joshua Murray, Kamalprit Chokar, Michaelia Banning, Muhammad Mamdani, View ORCID ProfileSharmistha Mishra, View ORCID ProfileJane Batt
doi: https://doi.org/10.1101/2020.09.11.20192419
David Landsman
1MAP Centre for Urban Health Solutions, Li Ka Shing Knowledge Institute, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for David Landsman
Ahmed Abdelbasit
2Department of Medicine, University of Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Christine Wang
2Department of Medicine, University of Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Christine Wang
Michael Guerzhoy
3Princeton University, Princeton, New Jersey, United States
4University of Toronto, Toronto, Ontario, Canada
5Li Ka Shing Knowledge Institute, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Ujash Joshi
4University of Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Shaun Mathew
6Department of Computer Science, Ryerson University, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Chloe Pou-Prom
7Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
David Dai
7Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Victoria Pequegnat
8Decision Support Services, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Joshua Murray
7Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Kamalprit Chokar
9Division of Respirology, Department of Medicine, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Michaelia Banning
7Unity Health Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Muhammad Mamdani
5Li Ka Shing Knowledge Institute, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
7Unity Health Toronto, Toronto, Ontario, Canada
10Leslie Dan Faculty of Pharmacy, University of Toronto, Canada, Toronto, Ontario, Canada
11Institute of Health Policy, Management, and Evaluation, University of Toronto, Toronto, Ontario, Canada
12Vector Institute, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Sharmistha Mishra
1MAP Centre for Urban Health Solutions, Li Ka Shing Knowledge Institute, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
2Department of Medicine, University of Toronto, Toronto, Ontario, Canada
11Institute of Health Policy, Management, and Evaluation, University of Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Sharmistha Mishra
Jane Batt
13Keenan Research Center for Biomedical Science, St. Michael’s Hospital, Unity Health Toronto, Toronto, Ontario, Canada
14Institute of Medical Science, University of Toronto, Toronto, Ontario, Canada
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for Jane Batt
  • For correspondence: Jane.batt{at}utoronto.ca
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Data/Code
  • Preview PDF
Loading

Abstract

Background Tuberculosis (TB) is a major cause of death worldwide. TB research draws heavily on clinical cohorts which can be generated using electronic health records (EHR), but granular information extracted from unstructured EHR data is limited. The St. Michael’s Hospital TB database (SMHTB) was established to address gaps in EHR-derived TB clinical cohorts and provide researchers and clinicians with detailed, granular data related to TB management and treatment.

Methods We collected and validated multiple layers of EHR data from the TB outpatient clinic at St. Michael’s Hospital, Toronto, Ontario, Canada to generate the SMH-TB database. SMH-TB contains structured data directly from the EHR, and variables generated using natural language processing (NLP) by extracting relevant information from free-text within clinic, radiology, and other notes. NLP performance was assessed using recall, precision and F1 score averaged across variable labels. We present characteristics of the cohort population using binomial proportions and 95% confidence intervals (CI), with and without adjusting for NLP misclassification errors.

Results SMH-TB currently contains retrospective patient data spanning 2011 to 2018, for a total of 3298 patients (N=3237 with at least 1 associated dictation). Performance of TB diagnosis and medication NLP rulesets surpasses 93% in recall, precision and F1 metrics, indicating good generalizability. We estimated 20% (95% CI: 18.4-21.2%) were diagnosed with active TB and 46% (95% CI: 43.8-47.2%) were diagnosed with latent TB. After adjusting for potential misclassification, the proportion of patients diagnosed with active and latent TB was 18% (95% CI: 16.8-19.7%) and 40% (95% CI: 37.8-41.6%) respectively

Conclusion SMH-TB is a unique database that includes a breadth of structured data derived from structured and unstructured EHR data. The data are available for a variety of research applications, such as clinical epidemiology, quality improvement and mathematical modelling studies.

Competing Interest Statement

The authors have declared no competing interest.

Funding Statement

Supported by the Ontario Early Researcher Award Number ER17-13-043 (to SM). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

The details of the IRB/oversight body that provided approval or exemption for the research described are given below:

Unity Health Toronto Research Ethics Board (REB 19-080)

All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.

Yes

Data Availability

The validated NLP rulesets are publicly available for use from: https://github.com/mishra-lab/tb-nlp-rulesets. Data collected in SMH-TB contains sensitive patient information and as such, researchers interested in conducting TB-related research using the data are welcome to contact the corresponding author and submit a request. The study team welcomes collaboration and use of the database, and all external requests will be screened to ensure adequate data exists to enable a collaboration. The project will then undergo the approval process of the Research and Ethics Board (REB) of Unity Health Toronto. Data provided to researchers can either be the de-identified version of the SMH-TB database, or the full identifiable version, based on their research needs and REB approval.

https://github.com/mishra-lab/tb-nlp-rulesets

Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY 4.0 International license.
Back to top
PreviousNext
Posted September 13, 2020.
Download PDF
Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Cohort profile: St. Michael’s Hospital Tuberculosis Database (SMH-TB), a retrospective cohort of electronic health record data and variables extracted using natural language processing
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Cohort profile: St. Michael’s Hospital Tuberculosis Database (SMH-TB), a retrospective cohort of electronic health record data and variables extracted using natural language processing
David Landsman, Ahmed Abdelbasit, Christine Wang, Michael Guerzhoy, Ujash Joshi, Shaun Mathew, Chloe Pou-Prom, David Dai, Victoria Pequegnat, Joshua Murray, Kamalprit Chokar, Michaelia Banning, Muhammad Mamdani, Sharmistha Mishra, Jane Batt
medRxiv 2020.09.11.20192419; doi: https://doi.org/10.1101/2020.09.11.20192419
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Cohort profile: St. Michael’s Hospital Tuberculosis Database (SMH-TB), a retrospective cohort of electronic health record data and variables extracted using natural language processing
David Landsman, Ahmed Abdelbasit, Christine Wang, Michael Guerzhoy, Ujash Joshi, Shaun Mathew, Chloe Pou-Prom, David Dai, Victoria Pequegnat, Joshua Murray, Kamalprit Chokar, Michaelia Banning, Muhammad Mamdani, Sharmistha Mishra, Jane Batt
medRxiv 2020.09.11.20192419; doi: https://doi.org/10.1101/2020.09.11.20192419

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Epidemiology
Subject Areas
All Articles
  • Addiction Medicine (349)
  • Allergy and Immunology (668)
  • Allergy and Immunology (668)
  • Anesthesia (181)
  • Cardiovascular Medicine (2648)
  • Dentistry and Oral Medicine (316)
  • Dermatology (223)
  • Emergency Medicine (399)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (942)
  • Epidemiology (12228)
  • Forensic Medicine (10)
  • Gastroenterology (759)
  • Genetic and Genomic Medicine (4103)
  • Geriatric Medicine (387)
  • Health Economics (680)
  • Health Informatics (2657)
  • Health Policy (1005)
  • Health Systems and Quality Improvement (985)
  • Hematology (363)
  • HIV/AIDS (851)
  • Infectious Diseases (except HIV/AIDS) (13695)
  • Intensive Care and Critical Care Medicine (797)
  • Medical Education (399)
  • Medical Ethics (109)
  • Nephrology (436)
  • Neurology (3882)
  • Nursing (209)
  • Nutrition (577)
  • Obstetrics and Gynecology (739)
  • Occupational and Environmental Health (695)
  • Oncology (2030)
  • Ophthalmology (585)
  • Orthopedics (240)
  • Otolaryngology (306)
  • Pain Medicine (250)
  • Palliative Medicine (75)
  • Pathology (473)
  • Pediatrics (1115)
  • Pharmacology and Therapeutics (466)
  • Primary Care Research (452)
  • Psychiatry and Clinical Psychology (3432)
  • Public and Global Health (6527)
  • Radiology and Imaging (1403)
  • Rehabilitation Medicine and Physical Therapy (814)
  • Respiratory Medicine (871)
  • Rheumatology (409)
  • Sexual and Reproductive Health (410)
  • Sports Medicine (342)
  • Surgery (448)
  • Toxicology (53)
  • Transplantation (185)
  • Urology (165)