Skip to main content
medRxiv
  • Home
  • About
  • Submit
  • ALERTS / RSS
Advanced Search

Challenges and best practices for digital unstructured data enrichment in health research: a systematic narrative review

Jana Sedlakova, Paola Daniore, Andrea Horn Wintsch, Markus Wolf, Mina Stanikic, Christina Haag, Chloé Sieber, Gerold Schneider, Kaspar Staub, Dominik Alois Ettlin, Oliver Grübner, Fabio Rinaldi, Viktor von Wyl, University of Zurich Digital Society Initiative (UZH-DSI) Health Community
doi: https://doi.org/10.1101/2022.07.28.22278137
Jana Sedlakova
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
3Institute of Biomedical Ethics and History of Medicine, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Paola Daniore
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Andrea Horn Wintsch
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
5Center for Gerontology, University of Zurich, Zurich, Switzerland
6CoupleSense: Health and Interpersonal Emotion Regulation Group, University Research Priority Program (URPP) Dynamics of Healthy Aging, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Markus Wolf
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
7Department of Psychology, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Mina Stanikic
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
4Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Christina Haag
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
4Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Chloé Sieber
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
4Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Gerold Schneider
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
8Department of Computational Linguistics, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Kaspar Staub
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
9Institute of Evolutionary Medicine, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Dominik Alois Ettlin
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
10Center of Dental Medicine, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Oliver Grübner
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
11Department of Geography, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Fabio Rinaldi
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
12Dalle Molle Institute for Artificial Intelligence (IDSIA), Switzerland
13Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland
14Fondazione Bruno Kessler, Trento, Italy
15Swiss Institute of Bioinformatics, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
Viktor von Wyl
1Digital Society Initiative, University of Zurich, Zurich, Switzerland
2Institute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland
4Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • For correspondence: viktor.vonwyl{at}uzh.ch
  • Abstract
  • Full Text
  • Info/History
  • Metrics
  • Supplementary material
  • Data/Code
  • Preview PDF
Loading

Abstract

Digital data play an increasingly important role in advancing medical research and care. However, most digital data in healthcare are in an unstructured and often not readily accessible format for research. Specifically, unstructured data are available in a non-standardized format and require substantial preprocessing and feature extraction to translate them to meaningful insights. This might hinder their potential to advance health research, prevention, and patient care delivery, as these processes are resource intensive and connected with unresolved challenges. These challenges might prevent enrichment of structured evidence bases with relevant unstructured data, which we refer to as digital unstructured data enrichment. While prevalent challenges associated with unstructured data in health research are widely reported across literature, a comprehensive interdisciplinary summary of such challenges and possible solutions to facilitate their use in combination with existing data sources is missing.

In this study, we report findings from a systematic narrative review on the seven most prevalent challenge areas connected with the digital unstructured data enrichment in the fields of cardiology, neurology and mental health along with possible solutions to address these challenges. Building on these findings, we compiled a checklist following the standard data flow in a research study to contribute to the limited available systematic guidance on digital unstructured data enrichment. This proposed checklist offers support in early planning and feasibility assessments for health research combining unstructured data with existing data sources. Finally, the sparsity and heterogeneity of unstructured data enrichment methods in our review call for a more systematic reporting of such methods to achieve greater reproducibility.

Competing Interest Statement

The authors have declared no competing interest.

Funding Statement

This study was partially funded by the Digital Society Initiative of the University of Zurich.

Author Declarations

I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.

Yes

I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.

Yes

I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).

Yes

I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.

Yes

Footnotes

  • Funding: This study was partially funded by the Digital Society Initiative (DSI).

Data Availability

All data produced in the present study are available upon reasonable request to the authors

Copyright 
The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. All rights reserved. No reuse allowed without permission.
Back to top
PreviousNext
Posted July 29, 2022.
Download PDF

Supplementary Material

Data/Code
Email

Thank you for your interest in spreading the word about medRxiv.

NOTE: Your email address is requested solely to identify you as the sender of this article.

Enter multiple addresses on separate lines or separate them with commas.
Challenges and best practices for digital unstructured data enrichment in health research: a systematic narrative review
(Your Name) has forwarded a page to you from medRxiv
(Your Name) thought you would like to see this page from the medRxiv website.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Share
Challenges and best practices for digital unstructured data enrichment in health research: a systematic narrative review
Jana Sedlakova, Paola Daniore, Andrea Horn Wintsch, Markus Wolf, Mina Stanikic, Christina Haag, Chloé Sieber, Gerold Schneider, Kaspar Staub, Dominik Alois Ettlin, Oliver Grübner, Fabio Rinaldi, Viktor von Wyl, University of Zurich Digital Society Initiative (UZH-DSI) Health Community
medRxiv 2022.07.28.22278137; doi: https://doi.org/10.1101/2022.07.28.22278137
Twitter logo Facebook logo LinkedIn logo Mendeley logo
Citation Tools
Challenges and best practices for digital unstructured data enrichment in health research: a systematic narrative review
Jana Sedlakova, Paola Daniore, Andrea Horn Wintsch, Markus Wolf, Mina Stanikic, Christina Haag, Chloé Sieber, Gerold Schneider, Kaspar Staub, Dominik Alois Ettlin, Oliver Grübner, Fabio Rinaldi, Viktor von Wyl, University of Zurich Digital Society Initiative (UZH-DSI) Health Community
medRxiv 2022.07.28.22278137; doi: https://doi.org/10.1101/2022.07.28.22278137

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero
  • Tweet Widget
  • Facebook Like
  • Google Plus One

Subject Area

  • Health Informatics
Subject Areas
All Articles
  • Addiction Medicine (349)
  • Allergy and Immunology (668)
  • Allergy and Immunology (668)
  • Anesthesia (181)
  • Cardiovascular Medicine (2648)
  • Dentistry and Oral Medicine (316)
  • Dermatology (223)
  • Emergency Medicine (399)
  • Endocrinology (including Diabetes Mellitus and Metabolic Disease) (942)
  • Epidemiology (12228)
  • Forensic Medicine (10)
  • Gastroenterology (759)
  • Genetic and Genomic Medicine (4103)
  • Geriatric Medicine (387)
  • Health Economics (680)
  • Health Informatics (2657)
  • Health Policy (1005)
  • Health Systems and Quality Improvement (985)
  • Hematology (363)
  • HIV/AIDS (851)
  • Infectious Diseases (except HIV/AIDS) (13695)
  • Intensive Care and Critical Care Medicine (797)
  • Medical Education (399)
  • Medical Ethics (109)
  • Nephrology (436)
  • Neurology (3882)
  • Nursing (209)
  • Nutrition (577)
  • Obstetrics and Gynecology (739)
  • Occupational and Environmental Health (695)
  • Oncology (2030)
  • Ophthalmology (585)
  • Orthopedics (240)
  • Otolaryngology (306)
  • Pain Medicine (250)
  • Palliative Medicine (75)
  • Pathology (473)
  • Pediatrics (1115)
  • Pharmacology and Therapeutics (466)
  • Primary Care Research (452)
  • Psychiatry and Clinical Psychology (3432)
  • Public and Global Health (6527)
  • Radiology and Imaging (1403)
  • Rehabilitation Medicine and Physical Therapy (814)
  • Respiratory Medicine (871)
  • Rheumatology (409)
  • Sexual and Reproductive Health (410)
  • Sports Medicine (342)
  • Surgery (448)
  • Toxicology (53)
  • Transplantation (185)
  • Urology (165)