Human Sweat Analysis (Scent II) for Disease Diagnosis: Data Alignment – UROP Spring Symposium 2024

Human Sweat Analysis (Scent II) for Disease Diagnosis: Data Alignment

Rahul Thalla

Pronouns: He/Him

Research Mentor(s): Sardar Ansari
Research Mentor School/College/Department: Weil Institute for Critical Care Research and Innovation / Medicine
Program:
Authors: Rahul Thalla, Loc Cao, Sardar Ansari
Session: Session 2: 10:00 am – 10:50 am
Poster: 37

Abstract

Human Sweat Analysis for Disease Diagnosis project aims to use gas chromatograms obtained from patients’ sweat to diagnose different diseases. Gas chromatography separates a gas mixture by passing it through a medium in which the components move at different rates. A portable gas chromatography device was designed to extract the odor from the body and break it down into carbon-based molecules called volatile organic compounds (VOCs). Since VOCs are byproducts of the metabolic processes of cells, chromatograms can give insight into patients’ physiological conditions or diseases. The device breaks body odor down into different components and generates a chromatogram signal corresponding to their concentration. Each component appears as a peak in the signal that can be matched with a VOC. This matching process can become difficult depending on the retention time drift: peaks corresponding to the same VOC can occur at different times. Varying elements of the experimental devices cause drift. Since we expect sweat chromatograms to have 40-50 peaks per signal, we chose to manually align a small subset of these signals to train a deep-learning peak alignment model, and then manually validate its performance to further fine-tune the algorithm. Both tasks will be done using a MATLAB Graphical User Interface (GUI) designed to facilitate chromatogram alignment. We are interested in quantifying the efficiency of the GUI in aligning chromatograms and in reducing the discrepancies between the ground truth (another set of annotations) and the user’s annotations. These annotations were compared over 18 signals with a total of 779 peaks and found to have a matching rate to the ground truth of 87.91%, where the two annotations identified the same peak with a tolerance of 1 second. 52 peaks (7.3%) in the ground truth annotations were detected but could not be matched to any peaks in the reference. These false positives can arise from either noise being picked up by peak detection or the VOC not being present in the reference chromatogram. The user correctly identified 53.4% of these false positives but also identified 22 more peaks. Many of these misidentified peaks between the user and the ground truth were due to peaks in close proximity to each other for which one would be aligned to the reference and the other would be labeled as a false positive. Real patient data will be soon collected to reliably diagnose up to 20 diseases.

lsa logoum logo