Indel Fitness Prediction from Diverse Deep Mutational Scanning Data with DDGindel – UROP Symposium

Indel Fitness Prediction from Diverse Deep Mutational Scanning Data with DDGindel

Aidan Horne

Research Mentor: Yang Li
Mentor Department: Ecology and Evolutionary Biology, LSA
Author(s): Aidan Horne, Yang Li
Session: Session 3 (11:00 AM – 11:50 AM)
Presentation Type: Poster 29

Abstract

Protein engineering is a scientific discipline that seeks to design proteins that treat diseases, counteract venoms, and act in biofuels. Research in this field seeks to find proteins or protein mutations that are more stable, active, or structurally specialized than those that are currently known or exist. With the use of highly accurate machine learning models, the change in thermodynamic fold stability, or the ??G (pronounced delta delta G), of a mutation can be predicted, reducing resource consumption in the creation and testing phases of protein design. A good protein folding stability is crucial, as a protein with poor folding stability will either misfold or fail to fold at all, resulting in a nonfunctional protein. Here, most machine learning models focus on predicting the ??G of substitution mutations, leaving predictive methods for insertion and deletion mutations (indels) sorely deficient. This project identifies how different mutation types affect the stability of proteins differently and aims to increase the number of methods to accurately predict the ??G of indels by showcasing DDGindel: a sequence-based ??G-prediction model that adapts the model architecture of DDGemb to facilitate ??G-prediction of indels through the implementation of pairwise sequence alignments. DDGemb, a substitution-only ??G-prediction model, outperformed 24 other ??G-prediction models on a set of benchmark datasets. DDGindel adapts the architecture of DDGemb to incorporate the ability to predict the ??G of indels and attempts to maintain predictive performance on substitution mutations. DDGindel is trained using 1,547,092 forward and reverse single-substitution, double-substitution, single-insertion, and single-deletion mutations on 479 proteins and domains (331 natural and 148 designed) and is tested on benchmark datasets to allow comparison with other ??G-prediction models.

lsa logoum logo