How AD Curation Studio Powers a Novel Digital Speech Biomarker Study
AD Curation Studio is helping power SpeechDx, a three-year study using speech patterns to diagnose Alzheimer's disease and related dementias (ADRD) earlier in disease progression, when interventions are more effective.
Convened by the Alzheimer's Drug Discovery Foundation's Diagnostics Accelerator, SpeechDx follows more than 2,000 participants across nine global sites to build a gold-standard dataset of prognostic speech biomarkers.
To manage this collaboration, SpeechDx uses AD Curation Studio, part of the AD Data Initiative's product suite, to harmonize its digital speech and clinical data and make it available to permissioned researchers worldwide.
The AD Data Initiative’s partnership with SpeechDx has also provided an opportunity to expand the product suite to include digital voice ingestion tools through the Global Research and Imaging Platform (GRIP).
Why Speech as a Biomarker?
Diagnosing ADRD early enough for effective treatment remains one of the field's hardest problems. Subtle changes in speech like pacing, word-finding, and structure can appear before cognitive decline is otherwise noticeable, making speech a promising early biomarker.
But building a reliable, prognostic speech dataset means collecting, harmonizing, and protecting two very different data types at scale: raw voice recordings and clinical trial data. SpeechDx was designed to solve exactly this problem, generating a gold-standard dataset the broader research community can build on.
Data Collection, Curation, and Sharing
Data Collection: SpeechDx collects two types of data. Digital voice data comes from the SpeechDx app, pre-installed on tablets at each participant's location of choice. Built by the Global Research and Imaging Platform (GRIP), the app uses open-source speech tasks including picture description and storytelling. These tasks are designed to elicit natural speech. Clinical data, including cerebrospinal fluid and plasma samples, MRI and PET scans, and neuropsychological testing is also collected during periodic clinical site visits.
Data Curation: Voice recordings are uploaded to a secure, Azure-hosted server, where clinical managers track participant progress on a dashboard. Each recording is quality-checked, stripped of identifying information, and transcribed before moving to AD Curation Studio. Clinical data is anonymized at each trial site before upload. Inside AD Curation Studio, the voice and clinical data are cleaned, described, and organized. A defined set of harmonized clinical variables is then matched to each participant's voice data to build the combined SpeechDx dataset. From the moment data enters AD Curation Studio, it's encrypted both at rest and in transit.
Data Sharing: This analysis-ready data is shared with permissioned research groups in secure private workspaces. Data stays read-only, so research groups can't download it without permission, though they can bring their own code and models, and their code, models, and results remain private.
If you are interested in discussing whether the AD Data Initiative product suite might help strengthen your research initiatives, please contact us at info@alzheimersdata.org.

