Efficient tool for sampling empirical reference distributions of SNPs for hypothesis testing – UROP Spring Symposium 2022

Efficient tool for sampling empirical reference distributions of SNPs for hypothesis testing

photo of presenter

Luke Miga

Pronouns: He/Him

Research Mentor(s): Matthew Patrick
Co-Presenter:
Research Mentor School/College/Department: Department of Dermatology / Medicine
Presentation Date: April 20
Presentation Type: Poster
Session: Session 1 – 10am – 10:50am
Room: League Ballroom
Authors: Luke Miga, Matthew Patrick, Lam Tsoi
Presenter: 106

Abstract

Genome-wide association (GWA) studies identify genotypes that are more likely to occur with particular phenotypes. They achieve this by comparing frequencies of single nucleotide polymorphisms (SNPs) between populations with and without the target phenotype. Often we are interested in knowing whether certain regulatory features (for example, chromatin accessibility) are enriched among the GWA signals. However, other factors such as minor allele frequency (MAF) and linkage disequilibrium (LD) also affect SNP frequencies. ldLookup is a Linux command-line tool that produces empirical reference distributions of SNPs for hypothesis testing. It was designed to 1) produce statistically valid distributions by controlling for MAF and LD, 2) produce distributions efficiently, and 3) be highly configurable and so applicable to different use cases. We present the program architecture: SNPs are stratified by MAF and number of LD surrogates, then stored in disk-based hash tables. To generate distributions, the user provides a list of input markers. Each input marker is stratified, and a random marker with similar MAF and number of LD surrogates is selected from the program database. By sampling multiple sets of markers and assessing their enrichment for regulatory features, we can form a null distribution for statistical testing. We also assess ldLookup’s performance through runtime, CPU utilization, and memory consumption benchmarks, which show that ldLookup can produce a few hundred to a few thousand reference distributions per second. Benchmarking also implicates I/O operations as a bottleneck and highlights the number of strata as a performance-sensitive parameter. ldLookup represents a performant sampling technology for GWA studies. It is a basis for more complicated statistical tests. Since it is a standard, configurable tool, it could help to avoid duplicated effort between particular GWA studies. To facilitate this, ldLookup is publicly available on Github.

Presentation link

Biomedical Sciences, Interdisciplinary

lsa logoum logo