Diya Kannappan
Research Mentor(s): Brian Athey
Mentor Department: Computational Medicine & Bioinformatics
Authors: Diya Kannappan, Monica Holmes, Greg Farnum, Brian Athey
Session: Session 1 (9:00am – 9:50am)
Presentation Type: Poster 98
Abstract
Although AI-driven basecalling has propelled genome sequencing forward, the reliance on imperfect training data often introduces noise and systematic biases. While widely used, these methods compromise data fidelity, and prevent the integration of genome sequencing into clinical application, such as personalized medicine, which demand the highest accuracy. In contrast, our research champions a paradigm shift toward single-molecule sequencing (SMS) accuracy, grounded in precisely engineered plasmid-based DNA standards and a robust computational framework. Central to this approach is the creation of high-purity plasmid-based DNA standards containing single molecular species with fully characterized sequences. Verified through Sanger sequencing, restriction enzyme digestion, and capillary electrophoresis, these plasmids serve as ground truths for accurately assessing basecalling performance. We further developed an end-to-end bioinformatics pipeline that segments raw SMS reads by length, aligns each subset to its corresponding reference sequence, and calculates four key metrics: per-read quality, per-read edit distance, per-base quality, and modification accuracy. We then compared these empirically determined metrics against those predicted by the basecaller. By systematically analyzing single molecular species, this pipeline precisely quantifies sequencing errors and helps detect systematic biases. In measuring and comparing the derived metrics, we provide a more comprehensive view of SMS fidelity without relying on potentially flawed training data. This strategy lays the groundwork for standardizing single-molecule sequencing error assessment, with the ultimate goal of realizing clinically robust genome analyses. By integrating precise plasmid standards and detailed error profiling, we move closer to achieving the single molecule sequencing accuracy required for widespread adoption of genome sequencing in healthcare.



