Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review.
Overview
- Bellini College of Artificial Intelligence, Cybersecurity and Computing, University of South Florida, 4202 East Fowler Avenue, Tampa, FL, 33620, United States, 1 8135850780
- Department of Otolaryngology Head and Neck Surgery, USF Health Voice Center, University of South Florida, Tampa, FL, United States
- Medical Engineering, College of Engineering, University of South Florida, Tampa, FL, United States
Abstract
Background: Driven by recent advances in artificial intelligence (AI), particularly in medicine, audio-based voice and speech biomarkers are increasingly investigated for various medical applications as a complementary or even alternative modality to traditional medical devices. The adoption of deep learning techniques in recent literature is motivated by their superior performance compared to classical machine learning methods. However, ethical and regulatory concerns regarding the black-box nature of these models have limited their integration into clinical workflows. Consequently, explainable artificial intelligence (XAI) has recently been used to address this issue by generating explanations for opaque model outputs. Ideally, medical XAI systems aim to provide human-understandable, clinically grounded explanations essential for enhanced AI trustworthiness and, thereby, facilitate adoption into real-world clinical settings.
Objective: We conduct a systematic literature review of XAI methods applied for explaining deep learning techniques in audio-based voice and speech clinical applications. We aim to identify what XAI methods have been used to explain the decisions of deep learning voice and speech AI systems in health care, as well as XAI-informed insights. Additionally, we aim to contextualize these findings with respect to clinical applicability and stakeholder relevance. Lastly, we identify opportunities and recommendations for future clinical audio XAI design.
Methods: We used PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses). Six electronic databases (IEEE Xplore, ACM Digital Library, Scopus, PubMed, Web of Science, and Nature) were searched for papers published between January 2015 and February 2025. Eligible studies applied explainability or interpretability methods to deep learning models for voice or speech audio in health care contexts. Risk of bias was assessed using PROBAST+AI (Prediction Model Risk of Bias Assessment Tool). The results were thematically synthesized across explainability categories, input representations, clinical domains, validation strategies, and stakeholder considerations.
Results: A total of 30 studies met the inclusion criteria. These studies used a range of explainability approaches, including gradient-based methods, perturbation-based techniques, surrogate model–based methods, model-internal representation analyses, concept-based detectors, and attention-based explanations. Applications spanned diverse clinical domains, including voice disorders, neurodegenerative diseases, psychiatric conditions, and traumatic brain injury. Overall, results indicate that most studies relied primarily on qualitative interpretation of explainability outputs, with limited quantitative validation of explanation consistency across external datasets. Furthermore, none of the included studies explicitly conducted human-in-the-loop evaluations with relevant stakeholders, highlighting a substantial gap in stakeholder alignment.
Conclusions: Current XAI practices in clinical voice and speech analysis are limited by insufficient validation, lack of domain-specific design, and misalignment with clinical stakeholder needs. This review highlights opportunities for developing validated, audio-aware, and stakeholder-centered XAI approaches to support trustworthy clinical deployment. Interpretation of these findings should consider limitations related to single-reviewer study selection, potential high-risk of bias, and the repeated use of benchmark datasets.
Reproduced under the paper's license (CC BY), from the paper cited above.
Code
The paper links to its data, not to its authors' code: see the Data section.
Tracing map
A tracing map links a paper to the code its authors published: this paper has none, so it has no map.
Data
Datasets cited
- doi:10.13026/
gzjs-0535 , at the source; found in the references - zenodo:16874898, at Zenodo; found in the references
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Funding: added National Institutes of Health: Bridge2AI; Common Fund
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 5 authors, 8 keywords, 5 MeSH terms, 119 references.
Cite
This paper
Ebraheem, M., Toghranegar, J., Bridge2AI-Voice Consortium, Bensoussan, Y., & Templeton, J. M. (2026). Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review. Journal of medical Internet research, 28, e83790. https://
BibTeX
@article{ebraheem2026exp
author = {Ebraheem, Mohamed and Toghranegar, Jamie and {Bridge2AI-Voice Consortium} and Bensoussan, Yael and Templeton, John Michael},
title = {{Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review}},
journal = {Journal of medical Internet research},
year = {2026},
month = jun,
volume = {28},
pages = {e83790},
publisher = {JMIR Publications Inc.},
issn = {1439-4456},
doi = {10.2196/
url = {https://
pmid = {42341346},
pmcid = {PMC13293602}
}
RIS
TY - JOUR
AU - Ebraheem, Mohamed
AU - Toghranegar, Jamie
AU - Bridge2AI-Voice Consortium
AU - Bensoussan, Yael
AU - Templeton, John Michael
TI - Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review
T2 - Journal of medical Internet research
J2 - J Med Internet Res
PY - 2026
DA - 2026/
VL - 28
SP - e83790
SN - 1439-4456
PB - JMIR Publications Inc.
DO - 10.2196/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.2196/
"type": "article-journal",
"title": "Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review",
"container-title": "Journal of medical Internet research",
"author": [
{
"family": "Ebraheem",
"given": "Mohamed"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"literal": "Bridge2AI-Voice Consortium"
},
{
"family": "Bensoussan",
"given": "Yael"
},
{
"family": "Templeton",
"given": "John Michael"
}
],
"container-title-short":
"volume": "28",
"page": "e83790",
"DOI": "10.2196/
"PMID": "42341346",
"PMCID": "PMC13293602",
"ISSN": "1439-4456",
"publisher": "JMIR Publications Inc.",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1002/lio2.70519
- Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects.Journal: Laryngoscope investigative otolaryngologyIn common: 1 reference, author Yael Bensoussan
- [2] doi:10.1038/s41598-026-47769-z [code]
- A multimodal explainable artificial intelligence framework for interpretable Parkinson's disease prediction.Journal: Scientific reportsIn common: 2 references
- [3] doi:10.3389/fninf.2026.1799307
- FUSION-AD: interpretable AI framework for risk assessment and subgroup discovery in Alzheimer's disease.Journal: Frontiers in neuroinformaticsIn common: 2 references
- [4] doi:10.3390/s26165121 [code]
- Electrophysiological Signatures of Sarcopenia: A Systematic Review of sEMG Features, Fatigue Indices and AI-Based Classifiers.Journal: Sensors (Basel, Switzerland)In common: 2 references
- [5] doi:10.1016/j.isci.2026.116135 [code]
- Predicting ICU in-hospital mortality from text-encoded structured EHR data using adaptive transformer layer fusion.Journal: iScienceIn common: 2 references
- [6] doi:10.1007/s40120-026-00924-0
- Artificial Intelligence and Machine Learning in Pediatric Epilepsy: A Systematic Review.Journal: Neurology and therapyIn common: 2 references
- [7] doi:10.1038/s41467-026-71555-0 [code]
- A deep representation learning model to predict response to vagus nerve stimulation.Journal: Nature communicationsIn common: 2 references
- [8] doi:10.1038/s41598-026-52658-6
- A multi-level attention CNN-transformer based framework for the detection of brain tumor using regional dual-score explainability.Journal: Scientific reportsIn common: 2 references
- [9] doi:10.1093/braincomms/fcag253 [code]
- Disease detection and classification in temporal lobe epilepsy: step-wise versus simultaneous AI decision models in a multisite neuroimaging study.Journal: Brain communicationsIn common: 2 references
- [10] doi:10.1371/journal.pone.0346575 [code]
- Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.Journal: PloS oneIn common: 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
