De-identification of Clinical Documents Using Natural Language Processing – UROP Spring Symposium 2022

De-identification of Clinical Documents Using Natural Language Processing

photo of presenter

Jiaye Tan

Pronouns: he/him/his

Research Mentor(s): VG Vinod Vydiswaran
Co-Presenter:
Research Mentor School/College/Department: University of Michigan / Medicine
Presentation Date: April 20
Presentation Type: Oral5
Session: Session 6 – 4:40pm – 5:30 pm
Room: Breakout room 5
Authors: Vinod Vydiswaran, Jiaye Tan
Presenter: 3

Abstract

Clinical documents are replete with information of research interest. They detail the nuances of each patient’s disease, the treatments they received, and the outcomes of such treatments. The Health Insurance Portability and Accountability Act (HIPAA) requires clinical notes to be ridden of personally identifiable information, called protected health information (PHI), prior to research use. Many institutions, including the UC system, have already developed their own de-identification tools. However, the University of Michigan still lacks an effective de-identification model, while the demand for de-identified clinical data continues to rise within the school community. To satisfy such needs, we propose a hybrid model combing machine learning and rule-based approaches. To tailor it to Michigan Medicine data, our model was trained with a rich set of annotated clinical notes derived exclusively from Michigan Medicine. Our goal is to have our model achieve SOTA results on the testing set, outcompeting pre-existing models including HyDeXT, Philter, and NLM-Scrubber. If successful, our project will help de-identify an abundance of clinical documents for research use without legal concerns.

Presentation link

Biomedical Sciences

lsa logoum logo