Skip to Main Content

Evan Koch

Assistant Professor
DownloadHi-Res Photo

About

Titles

Assistant Professor

Biography

Dr. Evan Koch is a population geneticist and computational biologist with a research focus on the genetic and evolutionary basis of human traits and diseases. He earned his PhD in Ecology and Evolution at the University of Chicago in the lab of John Novembre, developing theoretical and empirical approaches to describe how mutation, selection, and drift affect genetic variation and how mutation and population history shape the distribution of complex trait variation. As a postdoctoral research fellow with Shamil Sunyaev at Harvard Medical School, he worked on a fine-scale mutation rate map for the human genome and created population genetics models and statistical methods using GWAS data to study selection. At Yale School of Medicine, his group develops quantitative methods that bridge classical population genetics with contemporary genomic and phenotypic data.

Last Updated on November 05, 2025.

Appointments

Education & Training

Research Associate
Harvard Medical School (2025)
Postdoctoral Associate
Harvard Medical School (2023)
PhD
University of Chicago, Ecology and Evolution (2018)
BS
University of Texas at Austin, Biology (2012)

Research

Overview

The research in our lab revolves around how evolutionary forces such as mutation, natural selection, and population history shape genetic variation, primarily at the population level. We are particularly interested in the impacts on human traits and diseases and develop population-genetic theory, statistical methods, and computational tools to address these questions. Our preferred strategy begins with explicit models of evolutionary processes and develops methods and techniques that make those models useful for inference from modern genomic data. Human genetics now includes sequencing studies and biobanks containing hundreds of thousands to millions of individuals, increasingly complete maps of molecular effects from high-throughput experiments, and computational predictions for nearly every possible variant. Understanding the implications of these data requires models that connect what mutations do, how they are filtered by selection, and what is ultimately observed in human populations and disease studies.

Mutation rates and rare genetic variation

New mutations are the source of all genetic variation within populations and ultimately all evolutionary change. Different sites in the genome differ by orders of magnitude in their germline mutation rates, and these differences dominate the distribution of rare variation observed in sequencing studies. These data enable fine-scale mutation-rate estimation, but also necessitate models where rare variant frequencies depend on mutation rate. We develop theory and methods to understand mutation rates and to better utilize rare variation in genetic analyses.

In previous work this has included germline mutation rate maps, along with methods for comparison and estimation of residual variation in mutation rates. We previously worked on Roulette, which provides basepair-resolution predictions across the human genome, and current efforts are focused on improving accuracy and providing a complete description of hypermutable elements in the human genome. We also developed population-genetic models for how recurrent mutations in large samples and at high mutation-rate sites affect allele frequencies, and applied these models to improve the identification of mutational processes, estimate mutation-rate distributions in hypermutable regions, and fit recent demographic history.

As sequencing samples continue to grow, identical mutations will recur independently many times in the genetic history of the sampled individuals, and the affected sites will comprise an increasing fraction of the human genome. We will develop models that use recurrence to improve fine-scale mutation-rate estimates, identify mutational processes, and model rare variation. Doing so first requires accurate detection of recurrent mutations, and we develop machine-learning methods to identify and characterize recurrent mutations in biobank-scale whole-genome sequencing data from cohorts like All of Us and the UK Biobank. We will apply these methods to generate population-wide maps of repeated mutational origin.

In addition to refining mutation-rate estimates, we will also use these data to distinguish the effects of mutability from negative selection. While an overabundance of variation relative to model predictions can indicate enhanced mutability, depletions of genetic variation indicate selective constraint and possible disease contributions. Highly mutable sites that remain invariant in enormous sequencing cohorts provide strong evidence for selection. We develop models that use this logic to identify strongly selected coding and noncoding sites, and determine which classes of strongly constrained regulatory and noncoding RNA elements, including hypermutable tRNAs and RNUs, make reproducible contributions to rare disease. Related work develops methods for rare variation in large samples with complex ancestral histories, using allele frequencies without requiring explicit models of every aspect of population history.

Natural Selection and the Genetic Architecture of Complex Traits

Complex-trait variation arises from the interplay of mutation, genetic drift, and natural selection. Biobank GWAS enable analyses of genetic architecture across large collections of traits and diseases. To explain the patterns uncovered by these studies, we modeled the evolutionary origins of genetic architecture and used these insights to infer selection from GWAS data. Application of one method across human traits and diseases found broad evidence for stabilizing selection, implying long-term selection against both risk and protective alleles for common diseases.

The study of natural selection using GWAS largely relies on single-locus models, but many properties of phenotypic selection, mutational architecture, and genetic interactions are not identifiable from single-locus models. We work on the development of two-locus models as a flexible and tractable way forward. Variant pairs provide a starting point for learning how variant effects compose on haplotypes. The expansion of variant-effect resources and GWAS has created an opportunity to learn pairwise composition by integrating effect estimates and predictions with linkage disequilibrium and haplotype structure. We develop analytical and simulation approaches that incorporate realistic features of human genomes and populations and produce predictions that can be compared directly with GWAS. These models connect patterns of signed effects among linked variants to features of trait architecture including the genomic distribution of causal mutations, pleiotropy, polygenicity, linkage disequilibrium, and recombination.

One application is to loci with multiple causal variants, an abundance of which have been uncovered by fine-mapping across GWAS and molecular QTLs. This allelic heterogeneity is a measurable footprint of evolution and mutational architecture, but existing summaries largely describe heterogeneity without considering these generative processes. We are combining fine-mapping results with models of mutation and selection that account for ascertainment. This will enable inference of how often discordant effects persist on shared haplotypes through shared selection and how multi-variant architectures vary across traits and genomic contexts. One goal here is to unravel the complexity of polygenic traits and diseases by first understanding genetic interactions within the same gene or regulatory element.

Mapping Functional Effects to Selection

Selective constraint, the propensity of the genome to remain unchanged across evolutionary time and within populations, is a precise indicator of broad functional importance. However, its application is limited by a lack of knowledge of where, biologically, constraint arises. Expanded phylogenetic and human sequencing data, together with functional genomic measurements at base-pair resolution and in individual cell types, provide an opportunity to characterize the biology mediating constraint. We develop methods that combine constraint with high-throughput assays of chromatin state, gene regulation, and molecular perturbation.

High-throughput experimental approaches and computational predictors now allow variant effects to be characterized at scale. Under negative selection, human polymorphism provides a readout of fitness effects. Our lab is developing population-genetic inference techniques that efficiently and flexibly map predictor scores to selection coefficients, allowing measurements and predictions from different molecular contexts to be compared on a common fitness scale. A major goal of this effort is to identify when and how functional assays, sequence-based predictors, and population-genetic evidence should be combined into a joint framework. This work uses mutation-aware models of negative selection and is being extended to variant pairs in order to use haplotype frequencies to test whether the combined fitness effects of nearby variants agree with expectations based on their individual molecular effects.

Natural selection is a primary force shaping genetic variation in human traits and diseases, and we examine how selective pressures emerge from phenotypic effects and variant interactions, while investigating how functional effects produce selective effects. By characterizing these processes across human populations, we aim to bridge population, functional, and statistical genomics to illuminate the mechanisms of selection in humans. These mechanisms inform disease biology and will improve the interpretation of genetic association studies.

Medical Research Interests

Biological Evolution; Computational Biology; Data Science; Genetics, Population; Genomics; Human Genetics; Models, Statistical

Research at a Glance

Publications Timeline

A big-picture view of Evan Koch's research output by year.

Publications

Featured Publications

2025

2023

2021

2019

Get In Touch

Contacts

Academic Office Number

Locations

Events