Mingze Sun
Research Mentor: Brian Athey
Mentor Department: Computational Medicine & Bioinformatics, Medicine
Author(s): Mingze Sun, Pranjal Srivastava, Haoran Li, Gregory Farnum, Brian Athey
Session: Session 6 (3:00 PM – 3:50 PM)
Presentation Type: Poster 51
Abstract
High-throughput DNA sequencing technology is the foundation of modern genomics, with applications ranging from population-scale studies to clinical diagnostics. Over the past decade, Illumina’s short-read-long sequencing platforms and long-read-long single-molecule sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), have achieved significant improvements in sequencing accuracy. However, there are still key limitations in how sequencing accuracy is defined and measured, and there is room for further improvement. The current sequencing workflow reports predicted accuracy via a Phred quality score (Q-value) assigned by the base recognizer, which reflects the model’s estimate of the probability of base recognition errors. To quantify actual accuracy, predicted reads can be compared to a known reference sequence and observed editing errors (substitutions, insertions, deletions) can be summarized as an empirical error rate. In practice, discrepancies are often efficiently calculated with the help of comparison tools such as minimap2 to obtain observed error data by read segment or base or sequence, and benchmarking studies on Illumina, PacBio, and ONT platforms often present both Q-value distributions and comparative empirical error rates. As a result, there is no generalized framework for directly measuring empirical or sequence-resolved error rates at the single-molecule level, nor is it possible to determine whether the observed accuracy exceeds experimentally achievable limits. Of particular importance is the inability of existing methods to provide accuracy bounds limited by sample purity or signal indistinguishability. The aim of this study is to quantify and define the empirical error rate in single-molecule sequencing by redefining sequencing as a finite multi-class classification problem, using plasmid DNA standards of known sequence and purity. The framework enables direct estimation of classification error probabilities and comparison of predicted and empirical error rates. It also provides theoretical bounds on achievable accuracy.


