Soroosh Tayebi Arasteh

Tackling Data Scarcity in Automatic Pathological Speech Analysis

Speech has become a valuable tool in digital healthcare. It’s cheap to collect, non-invasive, and useful for detecting conditions like Parkinson’s and Alzheimer’s disease or supporting speech therapy. But speech is also biometric data, meaning it can reveal a lot about a person’s identity. Combined with strict privacy regulations around patient data, this severely limits how much pathological speech data researchers can actually access and share, which in turn limits how well deep learning models can be trained on it.

This thesis works to expand the availability of pathological speech data for research while protecting patient privacy, addressing this from three angles. First, it investigates whether pathological speech is inherently easier to re-identify than healthy speech using automatic speaker verification. The answer turned out to be disorder-specific: adults with Dysphonia faced a higher risk of re-identification, while Dysarthria carried risks similar to healthy speech, and speech intelligibility itself didn’t affect verification accuracy. For children with Cleft Lip and Palate, the recording environment mattered more than the condition itself. Combining data across different disorders actually improved verification performance, pointing to the value of pathological diversity in datasets.

Second, the thesis anonymizes real pathological speech from over 2,700 speakers across several German institutions, comparing deep-learning-based and signal-processing-based anonymization methods. Across most disorders, including Dysarthria, Dysphonia, and Cleft Lip and Palate, anonymization achieved strong privacy gains with little to no loss in diagnostic usefulness, while Dysglossia even saw slight improvements. The effect was fairly consistent across different demographic groups, though the right approach still depends on the disorder.

Third, the thesis turns to federated learning as a way to train diagnostic models across institutions without sharing raw patient data. Using speech data from three separate real-world corpora in German, Spanish, and Czech, a federated model for detecting Parkinson’s disease outperformed every institution’s local model and matched the performance of a model trained on all the data pooled together, without any of that data ever leaving its home institution.

Together, these contributions show that privacy-preserving techniques don’t have to come at the cost of diagnostic accuracy, and that disorder-specific, collaborative approaches can open up more pathological speech data for research while keeping patients’ privacy intact.