A microphone close up in a studio setting
Accessibility

The Speech Accessibility Project is Teaching AI to Understand Impaired Speech

8 min read|Updated September 2026
Share

At a glance

The Speech Accessibility Project at the University of Illinois Urbana-Champaign has built the world's largest dataset of impaired speech, with more than 1,500 hours of recordings from roughly a thousand people with Parkinson's disease, ALS, stroke, Down syndrome, and cerebral palsy. Companies including Microsoft, Google, Amazon, Apple, and Meta use this data to retrain their speech recognition systems, and Microsoft has reported accuracy gains between 18 and 60 percent as a result.

  • Roughly 250,000 to 500,000 people in the United States live with Parkinson's disease, and millions more worldwide have conditions that change how they speak
  • Standard speech recognition systems historically failed on impaired speech, with word error rates far higher than for typical speech
  • About 2,000 participants had submitted recordings to the project as of the end of September 2025
  • Microsoft reported accuracy gains of 18 to 60 percent after training on project data
  • The dataset is available to researchers and companies who want to build more accessible voice technology

Voice assistants, dictation software, and voice activated devices have become everyday tools, but they only work if the system can understand you. For millions of people whose speech differs from the norm because of Parkinson's disease, amyotrophic lateral sclerosis, stroke, cerebral palsy, or Down syndrome, these systems simply fail. The Speech Accessibility Project, led by the University of Illinois Urbana-Champaign, is fixing that by building the datasets AI companies need to understand everyone.

The project, which launched in 2022 with backing from Amazon, Apple, Google, Meta, and Microsoft, pays people with impaired speech to record phrases and sentences. Those recordings form a private, de-identified dataset that participating companies use to retrain their speech recognition models. As of the end of September 2025, about 2,000 participants had submitted recordings, and the current data package includes 1,500 hours of speech from roughly a thousand participants. The results are already measurable. Microsoft has reported accuracy gains between 18 and 60 percent on impaired speech after training with project data.

Why Standard Speech Recognition Fails

Speech recognition systems learn from enormous datasets of recorded speech. If the training data contains mostly clear, typical voices, the resulting model will struggle with anything else. Accents have long been a known weakness. Impaired speech, sometimes called dysarthric speech, is a far bigger gap. Word error rates for dysarthric speech remained stubbornly high even as recognition of typical speech approached human parity.

The consequences go beyond convenience. Voice control is often the most accessible way to interact with technology for people with limited mobility. When a phone, smart speaker, or dictation tool cannot understand the user, it can shut them out of communication, independence, and work. The problem is not that the AI cannot be trained. The problem is that nobody had collected enough impaired speech data to train it well.

Paying Participants to Build the Dataset

The Speech Accessibility Project recruits adults in the United States and Canada with Parkinson's disease, ALS, Down syndrome, cerebral palsy, and speech effects of stroke. Participants record prompted phrases using a simple web interface from home, and they are paid for their time. This design matters. Collecting impaired speech is slow and effortful for the speaker, and compensating participants respects that effort while producing a dataset with genuine diversity of conditions, severity levels, and demographics.

By June 2024 the project had shared 235,000 speech samples with the companies funding it. The research community received its first partial dataset release in April 2024, containing more than 400 hours of speech from over 500 individuals with disabilities. That release powered the Interspeech 2025 Speech Accessibility Project Challenge, a public competition in which research teams from around the world competed to build the most accurate recognition systems for impaired speech.

Measurable Gains in the Real World

The corporate impact is concrete. Microsoft reported accuracy improvements between 18 and 60 percent after incorporating project data into its speech models. Google's related efforts grew out of Project Euphonia, a research initiative that collected over a million speech samples from people with impaired speech and led to Project Relate, an Android app that transcribes impaired speech in real time and restates it in a clear synthesized voice.

The project has also opened its data to nonprofits and companies beyond the original founders through a proposal process. Any organization building speech recognition tools can apply to use the dataset, which means the benefits can spread to augmentative communication devices, accessibility software, and niche products that serve small populations the biggest companies would never prioritize on their own.

A Model for Inclusive AI

What makes this project unusual is its structure. Competing companies, including Amazon, Apple, Google, Meta, and Microsoft, fund a shared dataset that any of them can use, alongside academic researchers. The participants who contribute the data are paid rather than mined. The dataset is designed to be private and de-identified from the start. It is a working template for how the AI industry can close representation gaps it cannot solve alone.

There is a lesson here that reaches beyond speech. Every AI system reflects its training data, and groups missing from that data get worse results. Fixing that does not require a breakthrough. It requires the unglamorous work of recruiting participants, recording sentences, compensating people fairly, and sharing the results. A thousand people recording phrases from their living rooms have already made voice technology measurably better for millions.

Common Questions

What is the Speech Accessibility Project?

The Speech Accessibility Project is a research initiative led by the University of Illinois Urbana-Champaign that collects recordings from people with impaired speech to build a private, de-identified dataset. Companies including Amazon, Apple, Google, Meta, and Microsoft use the data to improve speech recognition for people with conditions such as Parkinson's disease, ALS, stroke, cerebral palsy, and Down syndrome.

How large is the dataset?

As of the end of September 2025, about 2,000 participants had submitted recordings to the project. The current data package includes 1,500 hours of recorded speech from roughly 999 participants, and a partial release of more than 400 hours from over 500 individuals was made available to researchers in April 2024.

What accuracy improvements has the project produced?

Microsoft reported accuracy gains between 18 and 60 percent on impaired speech after training with project data. The April 2024 partial dataset release also powered the Interspeech 2025 Speech Accessibility Project Challenge, a public research competition for improving recognition of impaired speech.

Who can participate in the Speech Accessibility Project?

The project recruits residents of the United States and Canada, aged 18 and older, with Parkinson's disease, ALS, Down syndrome, cerebral palsy, or speech effects of stroke. Participants record prompted phrases from home through a web interface and are paid for their contributions.

How does this relate to Google Project Relate and Project Euphonia?

Project Euphonia is a Google research initiative that collected more than a million speech samples from people with impaired speech. It led to Project Relate, an Android app that transcribes impaired speech in real time, restates it in a clear synthesized voice, and connects to Google Assistant. The Speech Accessibility Project extends this kind of work as a cross-industry, academically led dataset shared across competing companies.

Sources: Speech Accessibility Project, University of Illinois Urbana-Champaign (Beckman Institute); Interspeech 2025 Speech Accessibility Project Challenge paper, arXiv; IPM Newsroom reporting on the project; Google Research Project Euphonia and Project Relate pages; The Verge coverage of Project Relate.