OSCR

Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review.

Overview

Authors: Mohamed Ebraheem1, Jamie Toghranegar2, Bridge2AI-Voice Consortium3, Yael Bensoussan2, John Michael Templeton1,3
  1. Bellini College of Artificial Intelligence, Cybersecurity and Computing, University of South Florida, 4202 East Fowler Avenue, Tampa, FL, 33620, United States, 1 8135850780
  2. Department of Otolaryngology Head and Neck Surgery, USF Health Voice Center, University of South Florida, Tampa, FL, United States
  3. Medical Engineering, College of Engineering, University of South Florida, Tampa, FL, United States
Institutions: University of South Florida (United States)
Journal: Journal of medical Internet research, volume 28, article e83790
Dates: received 9 September 2025; accepted 20 April 2026; published online 24 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.2196/83790 · PMID 42341346 · PMCID PMC13293602 · OpenAlex W4414100850
Open access: gold, a free copy (OpenAlex)
Status: data only
Categories: human (organism), cognitive (subfield)
Keywords: explainable artificial intelligence, clinical voice analysis, speech biomarkers, deep learning, interpretability, medical decision support, trustworthy AI, artificial intelligence
MeSH: Artificial Intelligence*, Speech*, Voice*, Deep Learning, Humans (* major topic)
Journal subjects: Digital Health Reviews, Clinical Information and Decision Making, Decision Support for Health Professionals, Artificial Intelligence, Responsible Health AI, Voice Phenotyping and Vocal Biomarkers
Topic: Artificial Intelligence in Healthcare and Education (Health Informatics, Medicine), according to OpenAlex
Citations: not cited yet (Europe PMC); 126 references in the paper

Abstract

Background: Driven by recent advances in artificial intelligence (AI), particularly in medicine, audio-based voice and speech biomarkers are increasingly investigated for various medical applications as a complementary or even alternative modality to traditional medical devices. The adoption of deep learning techniques in recent literature is motivated by their superior performance compared to classical machine learning methods. However, ethical and regulatory concerns regarding the black-box nature of these models have limited their integration into clinical workflows. Consequently, explainable artificial intelligence (XAI) has recently been used to address this issue by generating explanations for opaque model outputs. Ideally, medical XAI systems aim to provide human-understandable, clinically grounded explanations essential for enhanced AI trustworthiness and, thereby, facilitate adoption into real-world clinical settings.

Objective: We conduct a systematic literature review of XAI methods applied for explaining deep learning techniques in audio-based voice and speech clinical applications. We aim to identify what XAI methods have been used to explain the decisions of deep learning voice and speech AI systems in health care, as well as XAI-informed insights. Additionally, we aim to contextualize these findings with respect to clinical applicability and stakeholder relevance. Lastly, we identify opportunities and recommendations for future clinical audio XAI design.

Methods: We used PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses). Six electronic databases (IEEE Xplore, ACM Digital Library, Scopus, PubMed, Web of Science, and Nature) were searched for papers published between January 2015 and February 2025. Eligible studies applied explainability or interpretability methods to deep learning models for voice or speech audio in health care contexts. Risk of bias was assessed using PROBAST+AI (Prediction Model Risk of Bias Assessment Tool). The results were thematically synthesized across explainability categories, input representations, clinical domains, validation strategies, and stakeholder considerations.

Results: A total of 30 studies met the inclusion criteria. These studies used a range of explainability approaches, including gradient-based methods, perturbation-based techniques, surrogate model–based methods, model-internal representation analyses, concept-based detectors, and attention-based explanations. Applications spanned diverse clinical domains, including voice disorders, neurodegenerative diseases, psychiatric conditions, and traumatic brain injury. Overall, results indicate that most studies relied primarily on qualitative interpretation of explainability outputs, with limited quantitative validation of explanation consistency across external datasets. Furthermore, none of the included studies explicitly conducted human-in-the-loop evaluations with relevant stakeholders, highlighting a substantial gap in stakeholder alignment.

Conclusions: Current XAI practices in clinical voice and speech analysis are limited by insufficient validation, lack of domain-specific design, and misalignment with clinical stakeholder needs. This review highlights opportunities for developing validated, audio-aware, and stakeholder-centered XAI approaches to support trustworthy clinical deployment. Interpretation of these findings should consider limitations related to single-reviewer study selection, potential high-risk of bias, and the repeated use of benchmark datasets.

Reproduced under the paper's license (CC BY), from the paper cited above.

Code

The paper links to its data, not to its authors' code: see the Data section.

Tracing map

A tracing map links a paper to the code its authors published: this paper has none, so it has no map.

Data

Datasets cited

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Funding: added National Institutes of Health: Bridge2AI; Common Fund

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 5 authors, 8 keywords, 5 MeSH terms, 119 references.

Cite

This paper

Ebraheem, M., Toghranegar, J., Bridge2AI-Voice Consortium, Bensoussan, Y., & Templeton, J. M. (2026). Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review. Journal of medical Internet research, 28, e83790. https://doi.org/10.2196/83790

BibTeX

@article{ebraheem2026explainable,
author = {Ebraheem, Mohamed and Toghranegar, Jamie and {Bridge2AI-Voice Consortium} and Bensoussan, Yael and Templeton, John Michael},
title = {{Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review}},
journal = {Journal of medical Internet research},
year = {2026},
month = jun,
volume = {28},
pages = {e83790},
publisher = {JMIR Publications Inc.},
issn = {1439-4456},
doi = {10.2196/83790},
url = {https://doi.org/10.2196/83790},
pmid = {42341346},
pmcid = {PMC13293602}
}

RIS

TY - JOUR
AU - Ebraheem, Mohamed
AU - Toghranegar, Jamie
AU - Bridge2AI-Voice Consortium
AU - Bensoussan, Yael
AU - Templeton, John Michael
TI - Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review
T2 - Journal of medical Internet research
J2 - J Med Internet Res
PY - 2026
DA - 2026/06/24
VL - 28
SP - e83790
SN - 1439-4456
PB - JMIR Publications Inc.
DO - 10.2196/83790
UR - https://doi.org/10.2196/83790
LA - en
ER -

CSL-JSON

{
"id": "10.2196/83790",
"type": "article-journal",
"title": "Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review",
"container-title": "Journal of medical Internet research",
"author": [
{
"family": "Ebraheem",
"given": "Mohamed"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"literal": "Bridge2AI-Voice Consortium"
},
{
"family": "Bensoussan",
"given": "Yael"
},
{
"family": "Templeton",
"given": "John Michael"
}
],
"container-title-short": "J Med Internet Res",
"volume": "28",
"page": "e83790",
"DOI": "10.2196/83790",
"PMID": "42341346",
"PMCID": "PMC13293602",
"ISSN": "1439-4456",
"publisher": "JMIR Publications Inc.",
"URL": "https://doi.org/10.2196/83790",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1002/lio2.70519
Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects.
Journal: Laryngoscope investigative otolaryngology
In common: 1 reference, author Yael Bensoussan
[2] doi:10.1038/s41598-026-47769-z [code]
A multimodal explainable artificial intelligence framework for interpretable Parkinson's disease prediction.
Journal: Scientific reports
In common: 2 references
[3] doi:10.3389/fninf.2026.1799307
FUSION-AD: interpretable AI framework for risk assessment and subgroup discovery in Alzheimer's disease.
Journal: Frontiers in neuroinformatics
In common: 2 references
[4] doi:10.3390/s26165121 [code]
Electrophysiological Signatures of Sarcopenia: A Systematic Review of sEMG Features, Fatigue Indices and AI-Based Classifiers.
Journal: Sensors (Basel, Switzerland)
In common: 2 references
[5] doi:10.1016/j.isci.2026.116135 [code]
Predicting ICU in-hospital mortality from text-encoded structured EHR data using adaptive transformer layer fusion.
Journal: iScience
In common: 2 references
[6] doi:10.1007/s40120-026-00924-0
Artificial Intelligence and Machine Learning in Pediatric Epilepsy: A Systematic Review.
Journal: Neurology and therapy
In common: 2 references
[7] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: 2 references
[8] doi:10.1038/s41598-026-52658-6
A multi-level attention CNN-transformer based framework for the detection of brain tumor using regional dual-score explainability.
Journal: Scientific reports
In common: 2 references
[9] doi:10.1093/braincomms/fcag253 [code]
Disease detection and classification in temporal lobe epilepsy: step-wise versus simultaneous AI decision models in a multisite neuroimaging study.
Journal: Brain communications
In common: 2 references
[10] doi:10.1371/journal.pone.0346575 [code]
Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.
Journal: PloS one
In common: 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.