Anupama Jha, PhD
Assistant Professor of GeneticsCards
Featured Publication
Jha Lab
Education
University of Pennsylvania (2020)
Technical University of Munich (2014)
Featured Publication
Jha Lab
Education
University of Pennsylvania (2020)
Technical University of Munich (2014)
Featured Publication
Jha Lab
Education
University of Pennsylvania (2020)
Technical University of Munich (2014)
About
Copy Link
Titles
Assistant Professor of Genetics
Biography
Dr. Anupama Jha's research is focused on developing predictive machine-learning methods to understand three-dimensional genome architecture and downstream gene regulation in healthy tissues and cancers. By building computational models that connect DNA sequence, chromatin organization, and functional readouts, her work aims to uncover the principles governing genome regulation across biological contexts and to translate large-scale genomic data into interpretable models of cellular function.
Dr. Jha received a B.Tech. in Information Technology from GGSIPU, an M.S. in Informatics from the Technical University of Munich, and a Ph.D. in Computer and Information Science from the University of Pennsylvania, where she worked in the BioCiphers Lab under the mentorship of Dr. Yoseph Barash. During her doctoral training, she developed interpretable deep learning methods to study tissue-specific alternative splicing and RNA-binding protein regulatory networks, including the development of enhanced integrated gradients as a framework for improving the interpretability of deep learning models in genomics. She then completed postdoctoral training at the University of Washington in the Noble Lab, under the supervision of Dr. William Stafford Noble, prior to joining Yale in 2026.
Among her notable contributions, Dr. Jha developed TwinC, a sequence-to-function model for predicting and functionally interpreting inter-chromosomal genome architecture from DNA sequence, published in Nature Communications. She also co-developed Fibertools, a tool for DNA m6A calling and integrated long-read epigenetic and genetic analysis, published in Genome Research. Her earlier work on applying deep learning to identify common transcriptome signatures of cancer, published in Genome Biology, demonstrated the power of interpretable neural network models for uncovering disease-relevant regulatory programs. She is a collaborating member of the ENCODE, DNA Zoo, and 4D Nucleome Consortia, and is a co-author on a generalizable Hi-C foundation model for chromatin architecture published in Nature Methods.
Dr. Jha is the recipient of an NHGRI K99/R00 Pathway to Independence Award, an NVIDIA Academic Grant, and a UW Data Science Fellowship from the eScience Institute at the University of Washington. She has received travel fellowships from ISMB/ECCB and multiple best poster awards, including at the RNA Biology & Cancer Symposium. She serves on the Joint Steering Committee of the Yale-BI Biomedical Data Science Fellowship and is an active reviewer for journals including Genome Biology, Nature Communications, and PLOS Computational Biology, as well as conferences including RECOMB and ISMB.
Appointments
Genetics
Assistant ProfessorPrimary
Other Departments & Organizations
Education & Training
- Postdoctoral Scholar
- University of Washington (2025)
- PhD
- University of Pennsylvania (2020)
- MSc
- Technical University of Munich (2014)
Research
Copy Link
Overview
Jha laboratory focuses on developing large-scale computational and machine learning methods to reveal the sequence basis of 3D genome architecture and its impact on downstream gene regulation across human tissues and evolution. Leveraging large-scale and heterogeneous multi-omics data across human tissues and mammalian species, combined with novel, integrative, and interpretable machine learning methods, we study how sequence variations influence nuclear organization in humans and across mammalian evolution, how variations in nuclear organization impact transcriptional and post-transcriptional gene regulation across human tissues, and how the regulatory infrastructure is misregulated in cancer. We are creating the first comprehensive map of tissue-specific regulatory elements implicated in tissue-agnostic and tissue-specific regulation through cis- and trans-chromosomal contacts.
Public Health Interests
ORCID
0000-0003-3029-2086Jha Lab
We study 3D genome architecture and gene regulation in healthy tissues and cancer using machine learning.
View Lab Website
Research at a Glance
Publications Timeline
Publications
Featured Publications
DNA-m6A calling and integrated long-read epigenetic and genetic analysis with fibertools
Jha A, Bohaczuk S, Mao Y, Ranchalis J, Mallory B, Min A, Hamm M, Swanson E, Dubocanin D, Finkbeiner C, Li T, Whittington D, Noble W, Stergachis A, Vollger M. DNA-m6A calling and integrated long-read epigenetic and genetic analysis with fibertools. Genome Research 2024, 34: 1976-1986. PMID: 38849157, PMCID: PMC11610455, DOI: 10.1101/gr.279095.124.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and ConceptsConceptsOxford Nanopore TechnologiesLong readsEpigenetic architectureEpigenetic dataEpigenetic analysisLong-read DNA sequencingLong-read dataLong-read sequencingSingle-nucleotide resolutionVariable genomic regionsSingle-molecule sequencingSingle-moleculeSingle-molecule resolutionLong DNA moleculesPacific BiosciencesState-of-the-art toolkitsCytosine methylationGenomic regionsSequencing platformsNanopore TechnologiesDNA sequencesGenetic analysisAccurate identificationEpigenetic studiesDNA moleculesEnhanced Integrated Gradients: improving interpretability of deep learning models using splicing codes as a case study
Jha A, K. Aicher J, R. Gazzara M, Singh D, Barash Y. Enhanced Integrated Gradients: improving interpretability of deep learning models using splicing codes as a case study. Genome Biology 2020, 21: 149. PMID: 32560708, PMCID: PMC7305616, DOI: 10.1186/s13059-020-02055-7.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and ConceptsIdentifying common transcriptome signatures of cancer by interpreting deep learning models
Jha A, Quesnel-Vallières M, Wang D, Thomas-Tikhonenko A, Lynch K, Barash Y. Identifying common transcriptome signatures of cancer by interpreting deep learning models. Genome Biology 2022, 23: 117. PMID: 35581644, PMCID: PMC9112525, DOI: 10.1186/s13059-022-02681-3.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and ConceptsConceptsCore cancer pathwaysTranscriptomic signaturesTranscriptomic featuresProtein-coding gene expressionRNA-seq samplesRNA processing genesCancer pathwaysSignatures of cancerNormal tissue typesRNA-seqAberrant splicingSplice variationCancer typesGene expressionGenomic alterationsGenesSplicingCancer biologyTranscriptomeTumor typesCell proliferationConclusionsOur resultsSolid tumor typesGene signatureTissue typesPrediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC
Jha A, Hristov B, Wang X, Wang S, Greenleaf W, Kundaje A, Aiden E, Bertero A, Noble W. Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC. Nature Communications 2026, 17: 5427. PMID: 42009674, PMCID: PMC13280383, DOI: 10.1038/s41467-026-72031-5.Peer-Reviewed Original ResearchCitationsAltmetricConceptsTrans contactsGenome architectureIn situ Hi-C dataChromatin conformation assaysGM12878 cell lineHi-C dataTranscription factor bindingFunctional interpretationGenome foldingInter-chromosomalChromatin accessibilityFactor bindingGene regulationDNA sequencesConformation assaysLigation-dependentDNA architectureDNACell typesCell linesG-quadruplexFoldingGM12878AssayChromatin
2026
A generalizable Hi-C foundation model for chromatin architecture, single-cell and multiomics analysis across species
Wang X, Zhang Y, Ray S, Jha A, Fang T, Hang S, Doulatov S, Noble W, Wang S. A generalizable Hi-C foundation model for chromatin architecture, single-cell and multiomics analysis across species. Nature Methods 2026, 1-15. PMID: 42298067, DOI: 10.1038/s41592-026-03097-8.Peer-Reviewed Original ResearchAltmetricConceptsHi-C dataSingle-cell Hi-C dataIntegrated analysisRegulatory functionsAnalysis of 3D structuresEpigenomic activityChromatin architectureGenomic analysisNuclear DNACellular processesEpigenomic regulationMultiomics analysisAnalytical pipelineSingle-cellFunctional roleCell typesChromatinSpeciesGenomeDNAMultiomicsRegulationCells
2024
Machine learning-optimized targeted detection of alternative splicing
Yang K, Islas N, Jewell S, Wu D, Jha A, Radens C, Pleiss J, Lynch K, Barash Y, Choi P. Machine learning-optimized targeted detection of alternative splicing. Nucleic Acids Research 2024, 53: gkae1260. PMID: 39727154, PMCID: PMC11797022, DOI: 10.1093/nar/gkae1260.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and ConceptsConceptsRNA-seq dataRNA-seqQuantification of alternative splicingDetection of alternative splicingGTEx RNA-seq dataRNA-seq methodJunction-spanning readsMultiplex reverse transcriptionPool of primersSequencing depthSplicing eventsPrimer sequencesAlternative splicingTranscriptome analysisRNA sequencingPrimersVariation sequencingReverse transcriptionSequenceSplicingHigh-throughputComprehensive detectionTranscriptomeTranscriptionRNAEnhancing Hi-C contact matrices for loop detection with Capricorn: a multiview diffusion model
Fang T, Liu Y, Woicik A, Lu M, Jha A, Wang X, Li G, Hristov B, Liu Z, Xu H, Noble W, Wang S. Enhancing Hi-C contact matrices for loop detection with Capricorn: a multiview diffusion model. Bioinformatics 2024, 40: i471-i480. PMID: 38940142, PMCID: PMC11211821, DOI: 10.1093/bioinformatics/btae211.Peer-Reviewed Original ResearchCitationsMeSH Keywords and ConceptsConceptsHigh-coverage dataChromatin featuresHi-CResolution enhancement methodExperimental Hi-C dataHi-C dataChromatin structure analysisContact matricesF1 scoreEnhancement methodDNA sequencesMachine learning modelsMean square errorNatural imagesChromatinSource codeLoop detectionLearning modelsModel backboneResolution enhancementThree-dimensional architectureGenomeComputational methodsStructural analysisCapricorn
2023
RNA splicing analysis using heterogeneous and large RNA-seq datasets
Vaquero-Garcia J, Aicher J, Jewell S, Gazzara M, Radens C, Jha A, Norton S, Lahens N, Grant G, Barash Y. RNA splicing analysis using heterogeneous and large RNA-seq datasets. Nature Communications 2023, 14: 1230. PMID: 36869033, PMCID: PMC9984406, DOI: 10.1038/s41467-023-36585-y.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and ConceptsConceptsRNA-seqRNA splicing analysisRNA-seq datasetsRNA-seq dataTranscriptome complexityGTEx v8Splicing regulationSplicing analysisRNA splicingBiological replicatesSplice variationMAJIQSplice variantsSplicingRNASuite of algorithmsTranscriptomeBenchmark datasetsBrain subregionsSynthetic dataDatasetVariantsReplicationRegulation
2021
RNA-Binding Proteins PCBP1 and PCBP2 Are Critical Determinants of Murine Erythropoiesis
Ji X, Jha A, Humenik J, Ghanem L, Kromer A, Duncan-Lewis C, Traxler E, Weiss M, Barash Y, Liebhaber S. RNA-Binding Proteins PCBP1 and PCBP2 Are Critical Determinants of Murine Erythropoiesis. Molecular And Cellular Biology 2021, 41: e00668-20. PMID: 34180713, PMCID: PMC8384066, DOI: 10.1128/mcb.00668-20.Peer-Reviewed Original ResearchCitationsMeSH Keywords and ConceptsConceptsRNA-binding proteinsRNA-binding protein PCBP1Erythroid lineageBlood formationExon splicingComplex phenotypesPrimary erythroid progenitorsPCBP1Mouse developmentGene expressionPCBP2Fetal demisePeri-implantitisErythroid progenitorsMurine erythropoiesisLociMouse embryosErythropoietic differentiationLineagesProteinEmbryosMidgestationDifferentiationInactivationMiceMOCCASIN: a method for correcting for known and unknown confounders in RNA splicing analysis
Slaff B, Radens C, Jewell P, Jha A, Lahens N, Grant G, Thomas-Tikhonenko A, Lynch K, Barash Y. MOCCASIN: a method for correcting for known and unknown confounders in RNA splicing analysis. Nature Communications 2021, 12: 3353. PMID: 34099673, PMCID: PMC8184769, DOI: 10.1038/s41467-021-23608-9.Peer-Reviewed Original ResearchCitationsAltmetricMeSH Keywords and Concepts
Academic Achievements & Community Involvement
Copy Link
Activities
activity Encyclopedia of DNA Elements consortium (ENCODE)
2021 - PresentProfessional OrganizationsMemberactivity DNA Zoo consortium
2021 - PresentProfessional OrganizationsMemberactivity 4D Nucleome consortium (4DN)
2021 - PresentProfessional OrganizationsMemberactivity International Society for Computational Biology (ISCB)
2014 - PresentProfessional OrganizationsMemberactivity Integrative models of nuclear DNA organization
03/01/2025 - PresentLectureDepartment of Biological SciencesDetailsWest Lafayette, IN, United StatesSponsored by Purdue University
Honors
honor UW Data Science Fellow
01/01/2025, 01/01/2024, 01/01/2023, 01/01/2022, 01/01/2021National AwardeScience Institute, University of WashingtonDetailsUnited Stateshonor Travel Fellowship
01/01/2023International AwardInternational Society for Computational Biology (ICSB) European Conference on Computational Biology (ECCB)
Get In Touch
Copy Link
Contacts
Administrative Support
Locations
Sterling Hall of Medicine
Lab
333 Cedar Street
New Haven, CT 06510