GEO to SEA CDM: An Automated Metadata Integration Pipeline – UROP Symposium

GEO to SEA CDM: An Automated Metadata Integration Pipeline

Adharv Prerepa

Research Mentor: Yongqun He
Mentor Department: Not Available, Medicine
Author(s): Adharv Prerepa, Anthony Huffman, Oliver He
Session: Session 5 (2:00 PM – 2:50 PM)
Presentation Type: Poster 72

Abstract

The Study-Experiment-Assay (SEA) Common Data Model (CDM) was developed to standardize the representation of biomedical experimental data, particularly in vaccine and immunology research, and currently integrates data from multiple sources such as ImmPort, VIGET, and CELLxGENE, enabling structured representation of vaccine studies across diverse experimental contexts. However, large-scale gene expression datasets from the Gene Expression Omnibus (GEO), one of the most widely used public repositories for transcriptomic data, have not yet been integrated into the SEA CDM framework, limiting the model’s coverage of public genomics data. In this study, we address this gap by integrating GEO metadata into SEA CDM through structured mapping of GEO series and sample-level information to the existing SEA CDM classes and headings. By integrating GEO metadata into SEA CDM, this work expands the scope of SEA CDM to include public transcriptomic datasets and enhances its ability to represent diverse biomedical studies within a unified data structure. The integration of GEO metadata into SEA CDM strengthens the model’s role as a framework for standardizing heterogeneous experimental data across various biomedical domains. Furthermore, it supports broader integrative analyses of vaccine-related transcriptomic studies and contributes to the development of a more comprehensive, interoperable biomedical data system.

lsa logoum logo