John Colyer
Research Mentor: Brian Athey
Mentor Department: Computational Medicine & Bioinformatics, Medicine
Author(s): John Colyer, Isaac Farnum, Gregory Farnum, Brian Athey
Session: Session 2 (10:00 AM – 10:50 AM)
Presentation Type: Poster 15
Abstract
Long-read sequencing has the potential to resolve complex genomic regions that are inaccessible to short-read sequencing, but current methods struggle with accuracy compared to the short-read methods. Highly homologous pharmacogenes, such as CYP2D6 and CYP2D7, are particularly difficult to sequence due to repetitive sequences with slight variations that are difficult to solve unambiguously. Developing well-defined synthetic reference sequences is critical for benchmarking sequencing accuracy and identifying sources of error in long-read sequencing techniques. In this project, synthetic 10,000 basepair constructs of CYP2D6 and CYP2D7 were designed and assembled to serve as known physical reference sequences for nanopore sequencing evaluation. Gene design was performed using SnapGene, where each target sequence was divided into modular fragments containing 4 base overhangs for Golden Gate assembly. Fragments were modified with primer strands and amplified with PCR, cloned into plasmid vectors using type IIS restriction enzymes, and validated through purification and Sanger sequencing. Verified fragments were subsequently assembled into full-length gene constructs through Golden Gate assembly and prepared for nanopore sequencing. This allowed for comparison between the known physical reference sequence and the experimentally obtained reads from the nanopore sequencer. This workflow allows for systematic assessment of sequencing accuracy, including base-calling errors and structural variant detection in highly homologous genomic regions. By using synthetic genes with fully defined sequences, sequencing and base-calling errors could be differentiated from biological variability. Additionally, accounting for major variations in the synthetic gene by assembling other frequently observed variants of CYP2D6 and CYP2D7 allows further control of sequencing validation. Overall, this work supports improved long-read sequencing accuracy and more reliable interpretation of pharmacogenetic data, with potential implications for drug response prediction and precision medicine.


