300 Million and Counting:
How the AD Workbench Powers the World’s Largest Protein Biomarker Discovery Effort
The Global Neurodegeneration Proteomics Consortium (GNPC) uses AD Workbench to power the world's largest protein biomarker discovery effort for neurodegenerative disease. GNPC brought together dementia researchers from leading institutions to build this dataset: 300 million proteins from 40,000 biofluid samples.
Gates Ventures and Johnson & Johnson co-funded GNPC v1 to help accelerate drug discovery for the 57 million people worldwide living with neurodegenerative disease. AD Workbench, the AD Data Initiative's research environment, was used to create, curate, and share the GNPC v1 Harmonized Dataset (HDS) with qualified researchers around the world.
Why a Proteomics Dataset?
New biomarkers could dramatically speed up diagnostics and treatment for neurodegenerative diseases including Alzheimer's disease, Parkinson's disease, ALS, and frontotemporal dementia. Historically, biomarker discovery has been slow. Data has not been harmonized across studies and data governance rules make large-scale analysis hard.
GNPC is a consortium of two dozen research institutions built to solve this. It brings together data and bio-samples from many cohort studies into one large, harmonized proteomics database, so researchers can find actionable insights faster.
Data Collection, Curation, and Sharing
Data Collection: Primary data comes from the SomaScan 7k platform, which measures thousands of proteins in plasma, serum, and cerebrospinal fluid samples. GNPC partners could choose to prioritize cross-sectional or longitudinal samples, based on flexible funding support. Each contributing cohort also provided clinical metadata for every protein record diagnosis, disease stage, neuropsychological assessments, and demographics.
Data Curation: Each sample was analyzed, quality-checked, and matched to a clinical profile from the time it was collected. A professional vendor harmonized and anonymized the clinical data across cohorts before sharing it with consortium members. In total, the v1 HDS dataset available through AD Workbench includes over 50 clinical features from 35,000 biofluid samples making it the largest biomarker discovery effort for neurodegenerative disease to date. The AD Data Initiative also created a standard Data Contributor Agreement (DCA), signed before any data transfer, that meets international regulatory guidance and ensured all harmonization was done at no cost to data providers.
Data Sharing: The GNPC HDS v1 is curated, harmonized, and now accessible on AD Workbench, the AD Data Initiative's secure, cloud-based environment for neurodegenerative research. Inside AD Workbench, researchers get multimodal analysis tools, free compute and virtual machines, and can bring their own code, models, and approved external datasets. A one-year embargo period gave GNPC members exclusive, secure access before wider release. As of summer 2025, the HDS is available on the AD Discovery Portal, a public data catalog for AD Workbench and its partner platforms.
Figure 1
THE GNPC v1 HARMONIZED DATASET IS AVAILABLE ON THE AD DISCOVERY PORTAL, ENABLING PROTEOMIC DATA SHARING AT SCALE.
The three-stage flow shows (1) multi-cohort data harmonization in the harmonization vendor workspace, (2) intra-consortium sharing during the 12-month embargo period, and (3) researcher access to the HDS via individually permissioned AD Workbench workspaces.
What's Next: GNPC v1, vMP, and Beyond
When GNPC released the HDS v1 publicly in July 2025, it also published several papers, including in Nature Medicine on the dataset and its first-year findings including a map of how plasma proteins differ across neurodegenerative diseases, a clear APOEε4 protein signature that holds up across disease types, and new proteomic "aging clock" measurements across organ systems. Since release, about 600 users have applied for access through AD Workbench in the dataset's first 9 months.
Next up: the GNPC vMultiplatform (vMP) expansion, going to consortium members in Summer 2026 and to the public in Summer 2027. vMP will add multiple proteomic platforms (increasing depth to ~17,000 protein measurements per sample), greater geographic diversity, interventional cohort data, and new AI/ML analysis capabilities within AD Workbench.


