Reference standard
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
Four principles guide every validation:
We compare model outputs with a clinically meaningful endpoint or validated assessment appropriate to the intended use.
We evaluate data reflecting the intended population, languages, devices and realistic recording conditions.
We test on held-out, external or prospective data that did not influence model development.
We document sensitivity, specificity, AUROC or calibration as appropriate, together with uncertainty, subgroup, language, device and noise checks, and known limitations.
The table reports sensitivity and specificity for eight intended use cases, alongside their reference measures, cohort sizes and evaluated languages. Cohorts were balanced by class (healthy control versus pathological), age and sex, plus weight and smoking habits where applicable.
| Condition | Reference measure | Sensitivity | Specificity | Participants | Languages |
|---|---|---|---|---|---|
| ConditionStress | Reference measureSalivary cortisol | Sensitivity98% | Specificity92% | Participants400 | LanguagesFrench, Italian, English, Spanish, German and Portuguese |
| ConditionAnxiety | Reference measureGAD-7 | Sensitivity61% | Specificity79% | Participants1,132 | LanguagesFrench, Italian, English, Spanish, German, Portuguese and Chinese |
| ConditionDepression | Reference measurePHQ-9 | Sensitivity77% | Specificity83% | Participants1,933 | LanguagesFrench, Italian, Chinese, English, Spanish, German and Portuguese |
| ConditionParkinson’s disease | Reference measureUPDRS | Sensitivity98% | Specificity96% | Participants2,745 | LanguagesFrench, Slovak, English, Italian, Spanish and German |
| ConditionAlzheimer’s disease | Reference measureMMSE | Sensitivity97% | Specificity92% | Participants2,957 | LanguagesEnglish, Greek, French, Slovak, Chinese, Spanish, German and Italian |
| ConditionMild cognitive impairment | Reference measureMoCA | Sensitivity81% | Specificity76% | Participants2,389 | LanguagesSlovak, Mandarin Chinese, Spanish, English, German and Italian |
| ConditionRespiratory function | Reference measureCOPD, COVID-19 or asthma diagnosis and RQoL | Sensitivity72% | Specificity82% | Participants832 | LanguagesSpanish, English and French |
| ConditionType 2 diabetes | Reference measureHbA1c levels | Sensitivity73% | Specificity70% | Participants1,440 | LanguagesSpanish, English and French |
Supported voice biomarkers can be integrated through the Virtuosis API.
Gervaise L et al. Voice biomarkers for multi-condition health screening using a short natural speech sample: a real-world validation. International Vocal Biomarkers Conference (IVBC). 2026.
Gervaise L et al. Speech-based biomarkers for scalable detection of Alzheimer’s disease, mild cognitive impairment, and Parkinson’s disease. European Journal of Neurology. 2026.
De Kalbermatten C, Gervaise L, Qorraj M. Family voice in rare diseases (FAVOUR-RD). European Conference on Rare Diseases (ECRD). 2026.
Gervaise L. Evaluation of voice-based screening for depression across controlled and real-world speech datasets. Voice AI Symposium. Frontiers. 2026.
Mekki M, Gervaise L. Voice-based depression detection: a bio-inspired and language-invariant model. Applied Machine Learning Days (AMLD). 2026.