Yuhuan Liu
Research Mentor: Keith Feldman
Mentor Department: Learning Health Sciences, Medicine
Author(s): Yuhuan Liu, Keith Feldman
Session: Session 4 (1:00 PM – 1:50 PM)
Presentation Type: Poster 45
Abstract
Purpose: Research contributing to improving individual and population health has grown rapidly, leading to a large and diverse foundation for translational research. However, the heterogeneity in study cohort composition and study design has limited efforts to identify relevant population-specific evidence from literature. This project explores a knowledge graph embedding-based approach that uses population characteristics defined by trial inclusion criteria to retrieve and rank existing clinical trials most relevant to a given population. A secondary objective is to identify inclusion criteria that are sparsely represented in existing clinical trials to identify potentially understudied cohorts. Methods: This study used TransE embedding trained on PlaNet, a publicly available clinical trial knowledge graph linking trials to standardized disease, intervention, and eligibility-criteria concepts. For each study population, inclusion criteria embeddings are aggregated to form a generalized representation and then shifted using the learned trial-inclusion relation vector to align with trial embeddings. Relevance is measured through pairwise distances between the generated entity and known study embeddings. The resulting distance distributions are summarized using kernel density estimation (KDE), with mean distance used as a complementary measure, to evaluate how well specific population characteristics are represented across existing clinical trials. Results: Preliminary analyses indicate that similarity between entities in the embedding space is robust, with semantically related inclusion criteria located closely together. Early experiments have found that due to the specificity of each trial’s inclusion criteria, translating multiple generic inclusion criteria to identify relevant studies is somewhat noisy. However, replication of the approach on diseases for trials has been shown to be more robust. Ongoing work is seeking to refine approaches to generalizing inclusion criteria to better allow for comparisons across heterogeneous existing clinical trials. Early KDE estimates also appear promising in identifying areas of low density in the embedding space, representing fewer similar nodes in a given area. Conclusion: This work suggests that knowledge graph embedding-based methods can help identify and rank studies relevant to specific population characteristics, offering a potentially robust and scalable approach for translational research synthesis and hypothesis generation. Ongoing work focuses on improving embeddings to better capture heterogeneity in real-world clinical trials.



