Whole Genome Sequencing Data for ADRD Research
Harmonized whole genome sequencing data for Alzheimer's research is available through the Whole Genome Sequence Harmonization Study (WGS Harmonization), which combines data from three Accelerating Medicines Partnership – Alzheimer's Disease (AMP-AD) cohort studies. The study jointly calls variants across all three cohorts to produce a uniform dataset from 1,872 donors. The research was funded by the National Institute on Aging.
Genomic data for Alzheimer's disease research can often be generated on different sequencing platforms and at different depths across cohorts, which makes direct comparison difficult. This dataset processed all three cohorts through a shared pipeline, so that allele counts and quality metrics are consistent regardless of which study a sample originated from. For researchers working on Alzheimer's disease and related dementias (ADRD), this means a genomic finding in one cohort can be checked against the other two without re-deriving the underlying data from scratch.
What the Dataset Contains
The dataset contains:
- 1,872 donors with harmonized whole genome sequencing data across three AMP-AD cohorts
- 1,501 donors with matched, harmonized RNA sequencing data through the companion RNAseq Harmonization Study
- Funding: National Institute on Aging
- Data formats: a subsetted VCF reflecting study-specific batches, or a combined VCF spanning all three cohorts, plus raw FASTQ files for all samples
What Cohorts Contribute Genome Sequencing Data to This Study?
The dataset draws on three of the primary AMP-AD cohort studies, each contributing both whole-genome sequencing and, for most donors, matched RNA sequencing data:
- ROSMAP (Religious Orders Study and Rush Memory and Aging Project): 1,179 donors with harmonized WGS; 911 with both harmonized WGS and RNAseq, drawn from the dorsolateral prefrontal cortex, frontal cortex, head of caudate nucleus, posterior cingulate cortex, and temporal cortex.
- MSBB (Mount Sinai Brain Bank): 344 donors with harmonized WGS; 297 with both harmonized WGS and RNAseq, drawn from the frontal pole, inferior frontal gyrus, parahippocampal gyrus, prefrontal cortex, and superior temporal gyrus.
- MayoRNAseq: 349 donors with harmonized WGS; 293 with both harmonized.
What Researchers Can Do With This Dataset
Researchers use this dataset for cross-cohort replication, testing whether a genetic association found in one cohort holds in the other two without separately reprocessing each dataset. The 1,501 donors with both harmonized WGS and RNAseq data support multi-omic integration, pairing genetic variants with gene expression to study expression quantitative trait loci across multiple brain regions. The combined sample size, spanning nearly 1,900 donors, also increases statistical power for detecting rare and structural variants, and RNAseq spanning regions such as the dorsolateral prefrontal cortex, temporal cortex, and cerebellum supports brain-region-specific comparisons of genetic and expression signals across areas differentially affected by Alzheimer's disease pathology.
How to Access The Dataset
The WGS Harmonization Study dataset is discoverable through the AD Discovery Portal and accessible within individually permissioned AD Workbench workspaces at no cost to qualified researchers. AD Workbench provides free compute, virtual machines, and multimodal analysis tools and allows researchers to bring their own code, models, and approved external datasets into their workspace.
About the Alzheimer’s Disease Data Initiative
The Alzheimer’s Disease Data Initiative (AD Data Initiative) is a global coalition of partners working together to promote and power data sharing and research collaboration to accelerate breakthroughs in Alzheimer’s and related dementias (ADRD) research. We're on a mission to transform the field by offering secure data sharing and analytics tools and resources, all available to users at no cost.
