Noah Weissman
Pronouns: he/him/his
Research Mentor(s): VG Vinod Vydiswaran
Co-Presenter: Walker, Broadbent
Research Mentor School/College/Department: University of Michigan / Medicine
Presentation Date: April 20
Presentation Type: Poster
Session: Session 6 – 4:40pm – 5:30 pm
Room: League Ballroom
Authors: Vinod Vydiswaran, Noah Weissman, Walker Broadbent
Presenter: 70
Abstract
While other Universities and Hospital systems possess tools to de-identify personal information in clinical health data, The University of Michigan does not currently possess a tool specific to the structure of Michigan Medicine data. Three student annators annotated 3 sets of clinical data each, with every clinical data set containing 50 radiology report notes and each note being annotated by at least two annotators. Annotators will mark all occurrences of Name, Telecom, Address, Ethnicity/Nationality, Date, Age, Numeric & Alphanumerical IDs, Organization, and Occupation. An algorithm compared annotations on the same sets made by different annotators and found that all annotations had consistent agreement. The results for the annotations are as follows: Set A(Precision: 0.95, Recall: 0.91, F1: 0.93), Set B (Precision: 0.88, Recall: 0.97, F1: 0.92), and Set C (Precision: 0.86, Recall: 0.88, F1: 0.87). This high agreement certified the use of these annotations to train Natural Language Processing (NLP) models for the Precision Health project. In conclusion, the radiology report notes supported consistent annotations and will be used to train NLP models for the precision health task. Once the NLP models have been trained, an algorithm will be developed to systematically obscure personally identifiable information while simultaneously preserving readability and accuracy. This project has broad implications for the entirety of research being performed at Michigan because it will give researchers access to a much larger range of data from Michigan medicine, ideally allowing for more comprehensive studies.
Biomedical Sciences



