Xinyi Li
Research Mentor: Yongqun He
Mentor Department: Not Available, Medicine
Author(s): Xinyi Li, Michael Xiang, Yongqun Oliver He
Session: Session 4 (1:00 PM – 1:50 PM)
Presentation Type: Poster 31
Abstract
Adhesins are surface-exposed proteins that mediate host–pathogen interactions and play a central role in bacterial colonization, making them important targets for vaccine design. In vaccine informatics, accurately identifying adhesins remains a major challenge. The Vaxign2 pipeline relies on SPAAN, a machine-learning model developed in 2005, which was trained on limited protein data and does not take advantage of recent advances in large-scale protein databases or modern representation learning methods. As the UniProt database has grown significantly and new protein language models have emerged, we believe it is time to re-evaluate and improve adhesin prediction algorithms. To address this gap, we use attention-based representations from the pretrained ESM-2 protein language model to improve feature extraction for adhesin prediction. We trained and compared five supervised learning models: a Support Vector Machine (SVM) with a linear kernel, an SVM with a radial basis function (RBF) kernel, a Random Forest (RF), a Multilayer Perceptron (MLP), and an Extreme Gradient Boosting (XGBoost) model. Then, we use standard classification metrics like MCC and accuracy rate to evaluate model performance. Our findings indicate that the SVM with an RBF kernel consistently outperforms the other models, supporting our hypothesis that modern protein language model embeddings with conventional machine-learning classifiers can significantly enhance adhesin prediction. This study demonstrates how advances in protein language modeling can be directly translated into improved tools for vaccine informatics.


