GNPC v1 Harmonized Dataset: Proteomics Data for Neurodegenerative Disease Research
Fragmented proteomics data and inconsistent harmonization standards have historically limited large-scale biomarker discovery for neurodegenerative disease. The Global Neurodegeneration Proteomics Consortium (GNPC) v1 Harmonized Dataset (HDS) was built to close that gap. It's the largest neurodegeneration proteomics dataset to date, with over 300 million protein measurements from 40,000 biosamples. These samples span multiple neurodegenerative conditions including ALS, FTD, Alzheimer’s, and Parkinson’s Disease alongside healthy aging cohorts, and are sourced from 23 cohorts.
The GNPC v1 HDS was co-funded by Gates Ventures and Johnson & Johnson to accelerate drug discovery and improve patient outcomes for the 57 million people worldwide suffering from neurodegenerative diseases.
It's now broadly accessible to qualified researchers worldwide through the AD Data Initiative's AD Workbench platform.
What the Dataset Contains
The dataset contains:
- 300 million+ unique protein measurements
- 40,000 biofluid samples, including plasma, serum, and cerebrospinal fluid, approximately 40% of which are longitudinal
- 20+ contributing international clinical cohorts
- 50+ harmonized clinical features per record, including diagnosis, disease stage, neuropsychological assessments, and demographics
- Conditions covered: AD, PD, ALS, FTD and healthy aging cohorts
Protein data was generated across multiple platforms, including SomaScan, Olink, and mass spectrometry, enabling cross-modality proteomic analysis.. Clinical metadata was harmonized across cohorts by a professional vendor, anonymized prior to sharing, and quality-controlled at the individual record level.
A successor dataset, the vMultiplatform HDS, is currently in development and will expand the resource further.
What Researchers Can Do with This Dataset
The GNPC v1 HDS is suited for large-scale biomarker discovery across neurodegenerative disease, cross-cohort comparative analyses, and longitudinal studies of disease progression. Its harmonized clinical metadata enables researchers to investigate relationships between protein expression, diagnosis, and clinical outcomes across diverse populations at a scale not previously available in a single resource. Active research workstreams include longitudinal profiling, cross-sectional profiling, proteogenomic analysis, and multivariate prediction, with studies targeting inflammation, genetic risk factors, and combinatorial gene-protein signatures associated with disease type and progression.
How to Access the Dataset
The GNPC v1 HDS became broadly available to the research community on July 15, 2025. It is discoverable through the AD DISCOVERY PORTAL ↗ and accessible within individually permissioned AD Workbench workspaces at no cost to qualified researchers. Findings from the dataset have been published, including the GNPC summary paper in Nature Medicine.
AD Workbench provides compute, virtual machines, and multimodal analysis tools and allows researchers to bring their own code, models, and approved external datasets into their workspace at no cost to qualified researchers.
