OSCR

Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects.

Overview

Authors: James Anibal1,2, Geetha Krishna Chaitanya Nama3, Shrramana Ganesh4, Yosef Nafii4, Samantha Salvi Cruz5, Veronica Daoud4, Madeleine Zanin6, Tram Le4, Parsa Khorrami3, Bradford J Wood2, David Clifton1, Jamie Toghranegar4, Yael Bensoussan4, Yael E Bensoussan, Olivier Elemento, Jean‐Christophe Bélisle‐Pipon, David Dorr, Satrajit Ghosh, Alistair Johnson, Phillip Payne
and 43 other authorsMaria E Powell, Anaïs Rameau, Vardit Ravitsky, Alexandros Sigaras, Shaheen Awan, Ruth Bahr, Donald Bolser, Kathy J Jenkins, Frank Rudzicz, Jennifer Siu, Stephanie Watts, Yassmeen Abdel‐Aty, Toufeeq Ahmed Syed, James Anibal, Steven Bedrick, Isaac Bevers, Micah Boyer, Rahul Brito, Selina A Casalino, John Costello, Enrique Diaz‐Ocampo, Mahmoud Elmahdy, Kenneth Fletcher, Alexander Gelbard, Karim Hanna, Bill Hersh, Lochana Jayachandran, Kaley Jenney, Andrea Krussel, Chloe Loewith, Tempestt Neal, Claire Premi‐Bortolotto, Sarah Rohde, Samantha Salvi Cruz, Elizabeth Silberholz, Duncan Sutherland, Venkata Swarna Mukhi Talluri, Jamie Toghranegar, Kimberly Vinson, Claire Wilson, Madeleine Zanin, Theresa Zesiewicz, Robin Zhao
  1. Computational Health Informatics Lab, Institute of Biomedical Engineering, University of Oxford, Oxford, UK
  2. Center for Interventional Oncology, NIH Clinical Center, National Institutes of Health, Bethesda, Maryland, USA
  3. Department of Computer Science and Engineering, University of South Florida, Tampa, Florida, USA
  4. USF Health Voice Center, Department of Otolaryngology‐Head & Neck Surgery, University of South Florida, Tampa, Florida, USA
  5. Vanderbilt University Medical Center, Nashville, Tennessee, USA
  6. Mount Sinai Hospital, Toronto, Ontario, Canada
Journal: Laryngoscope investigative otolaryngology, volume 11, issue 4, article e70519
Dates: received 23 February 2026; accepted 19 July 2026; published online 24 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/lio2.70519 · PMID 42639494 · PMCID PMC13502059 · OpenAlex W7204103813
Open access: gold, a free copy (OpenAlex)
Status: data only
Categories: human (organism)
Methods: Machine learning, Connectivity, Statistics, Physiology & signal measures
Keywords: artificial intelligence, bias mitigation, dataset optimization, large language models
Topic: Artificial Intelligence in Healthcare and Education (Health Informatics, Medicine), according to OpenAlex
Funding: NIH HHS (OT2 OD032720)
Citations: not cited yet (Europe PMC); 94 references in the paper
Research resources: RRID:SCR_007345

Abstract

Objectives: Past studies have shown that many clinical machine learning models have performance limitations due to imbalances in the training data. For voice data generation projects, the origin of the problem may lie in the recruiting methods used during data collection efforts. This study introduces a generative AI pipeline for “dataset decision support”, recommending recruitment decisions based on high‐dimensional insights.

Methods: The publicly available GOSSIS‐1‐eICU dataset was filtered to create patient populations that were relevant to voice data generation projects. Lab results and vital signs from the electronic health record were also used to train a neural network for prediction of disease type. Prediction uncertainty estimates were included in the dataset as approximate indicators of health complexity. To select the best recruitment choice for addressing imbalances, an open‐source large language model (LLM) was then instructed to assess dataset statistics and the characteristics of possible participants. Simulations were run in which the system constructed datasets of 250 patients.

Results: In over 90% of cases, the proposed system reduced categorical imbalances and widened continuous distributions when compared to randomly sampled counterfactual datasets (q‐value < 0.05). Variables included race, age, BMI, sex, disease type, oxygenation status, co‐morbidities, post‐operative status, Glasgow Coma Scale verbal response score, the Acute Physiology Score III, prediction uncertainty, vital signs, and lab results.

Conclusion: LLMs may provide useful, explainable recommendations when presented with dataset distribution statistics and candidate profiles. In the future, this simulated scenario may be extended to align with conditions in emergency departments or other high‐volume settings.

Level of Evidence: 3.

Reproduced under the paper's license (CC BY), from the paper cited above.

Code

The paper links to its data, not to its authors' code: see the Data section.

Tracing map

A tracing map links a paper to the code its authors published: this paper has none, so it has no map.

Data

Datasets cited

Data Availability Statement

The authors have nothing to report.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 63 authors, 4 keywords, 1 funder, 59 references, 1 RRID.

Cite

This paper

Anibal, J., Nama, G. K. C., Ganesh, S., Nafii, Y., Cruz, S. S., Daoud, V., Zanin, M., Le, T., Khorrami, P., Wood, B. J., Clifton, D., Toghranegar, J., Bensoussan, Y., Bensoussan, Y. E., Elemento, O., Bélisle‐Pipon, J., Dorr, D., Ghosh, S., Johnson, A., . . . Zhao, R. (2026). Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects. Laryngoscope investigative otolaryngology, 11(4), e70519. https://doi.org/10.1002/lio2.70519

BibTeX

@article{anibal2026large,
author = {Anibal, James and Nama, Geetha Krishna Chaitanya and Ganesh, Shrramana and Nafii, Yosef and Cruz, Samantha Salvi and Daoud, Veronica and Zanin, Madeleine and Le, Tram and Khorrami, Parsa and Wood, Bradford J and Clifton, David and Toghranegar, Jamie and Bensoussan, Yael and Bensoussan, Yael E and Elemento, Olivier and Bélisle‐Pipon, Jean‐Christophe and Dorr, David and Ghosh, Satrajit and Johnson, Alistair and Payne, Phillip and Powell, Maria E and Rameau, Anaïs and Ravitsky, Vardit and Sigaras, Alexandros and Awan, Shaheen and Bahr, Ruth and Bolser, Donald and Jenkins, Kathy J and Rudzicz, Frank and Siu, Jennifer and Watts, Stephanie and Abdel‐Aty, Yassmeen and Syed, Toufeeq Ahmed and Anibal, James and Bedrick, Steven and Bevers, Isaac and Boyer, Micah and Brito, Rahul and Casalino, Selina A and Costello, John and Diaz‐Ocampo, Enrique and Elmahdy, Mahmoud and Fletcher, Kenneth and Gelbard, Alexander and Hanna, Karim and Hersh, Bill and Jayachandran, Lochana and Jenney, Kaley and Krussel, Andrea and Loewith, Chloe and Neal, Tempestt and Premi‐Bortolotto, Claire and Rohde, Sarah and Cruz, Samantha Salvi and Silberholz, Elizabeth and Sutherland, Duncan and Talluri, Venkata Swarna Mukhi and Toghranegar, Jamie and Vinson, Kimberly and Wilson, Claire and Zanin, Madeleine and Zesiewicz, Theresa and Zhao, Robin},
title = {{Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects}},
journal = {Laryngoscope investigative otolaryngology},
year = {2026},
month = aug,
volume = {11},
number = {4},
pages = {e70519},
publisher = {Wiley},
issn = {2378-8038},
doi = {10.1002/lio2.70519},
url = {https://doi.org/10.1002/lio2.70519},
pmid = {42639494},
pmcid = {PMC13502059}
}

RIS

TY - JOUR
AU - Anibal, James
AU - Nama, Geetha Krishna Chaitanya
AU - Ganesh, Shrramana
AU - Nafii, Yosef
AU - Cruz, Samantha Salvi
AU - Daoud, Veronica
AU - Zanin, Madeleine
AU - Le, Tram
AU - Khorrami, Parsa
AU - Wood, Bradford J
AU - Clifton, David
AU - Toghranegar, Jamie
AU - Bensoussan, Yael
AU - Bensoussan, Yael E
AU - Elemento, Olivier
AU - Bélisle‐Pipon, Jean‐Christophe
AU - Dorr, David
AU - Ghosh, Satrajit
AU - Johnson, Alistair
AU - Payne, Phillip
AU - Powell, Maria E
AU - Rameau, Anaïs
AU - Ravitsky, Vardit
AU - Sigaras, Alexandros
AU - Awan, Shaheen
AU - Bahr, Ruth
AU - Bolser, Donald
AU - Jenkins, Kathy J
AU - Rudzicz, Frank
AU - Siu, Jennifer
AU - Watts, Stephanie
AU - Abdel‐Aty, Yassmeen
AU - Syed, Toufeeq Ahmed
AU - Anibal, James
AU - Bedrick, Steven
AU - Bevers, Isaac
AU - Boyer, Micah
AU - Brito, Rahul
AU - Casalino, Selina A
AU - Costello, John
AU - Diaz‐Ocampo, Enrique
AU - Elmahdy, Mahmoud
AU - Fletcher, Kenneth
AU - Gelbard, Alexander
AU - Hanna, Karim
AU - Hersh, Bill
AU - Jayachandran, Lochana
AU - Jenney, Kaley
AU - Krussel, Andrea
AU - Loewith, Chloe
AU - Neal, Tempestt
AU - Premi‐Bortolotto, Claire
AU - Rohde, Sarah
AU - Cruz, Samantha Salvi
AU - Silberholz, Elizabeth
AU - Sutherland, Duncan
AU - Talluri, Venkata Swarna Mukhi
AU - Toghranegar, Jamie
AU - Vinson, Kimberly
AU - Wilson, Claire
AU - Zanin, Madeleine
AU - Zesiewicz, Theresa
AU - Zhao, Robin
TI - Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects
T2 - Laryngoscope investigative otolaryngology
J2 - Laryngoscope Investig Otolaryngol
PY - 2026
DA - 2026/08/24
VL - 11
IS - 4
SP - e70519
SN - 2378-8038
PB - Wiley
DO - 10.1002/lio2.70519
UR - https://doi.org/10.1002/lio2.70519
LA - en
ER -

CSL-JSON

{
"id": "10.1002/lio2.70519",
"type": "article-journal",
"title": "Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects",
"container-title": "Laryngoscope investigative otolaryngology",
"author": [
{
"family": "Anibal",
"given": "James"
},
{
"family": "Nama",
"given": "Geetha Krishna Chaitanya"
},
{
"family": "Ganesh",
"given": "Shrramana"
},
{
"family": "Nafii",
"given": "Yosef"
},
{
"family": "Cruz",
"given": "Samantha Salvi"
},
{
"family": "Daoud",
"given": "Veronica"
},
{
"family": "Zanin",
"given": "Madeleine"
},
{
"family": "Le",
"given": "Tram"
},
{
"family": "Khorrami",
"given": "Parsa"
},
{
"family": "Wood",
"given": "Bradford J"
},
{
"family": "Clifton",
"given": "David"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"family": "Bensoussan",
"given": "Yael"
},
{
"family": "Bensoussan",
"given": "Yael E"
},
{
"family": "Elemento",
"given": "Olivier"
},
{
"family": "Bélisle‐Pipon",
"given": "Jean‐Christophe"
},
{
"family": "Dorr",
"given": "David"
},
{
"family": "Ghosh",
"given": "Satrajit"
},
{
"family": "Johnson",
"given": "Alistair"
},
{
"family": "Payne",
"given": "Phillip"
},
{
"family": "Powell",
"given": "Maria E"
},
{
"family": "Rameau",
"given": "Anaïs"
},
{
"family": "Ravitsky",
"given": "Vardit"
},
{
"family": "Sigaras",
"given": "Alexandros"
},
{
"family": "Awan",
"given": "Shaheen"
},
{
"family": "Bahr",
"given": "Ruth"
},
{
"family": "Bolser",
"given": "Donald"
},
{
"family": "Jenkins",
"given": "Kathy J"
},
{
"family": "Rudzicz",
"given": "Frank"
},
{
"family": "Siu",
"given": "Jennifer"
},
{
"family": "Watts",
"given": "Stephanie"
},
{
"family": "Abdel‐Aty",
"given": "Yassmeen"
},
{
"family": "Syed",
"given": "Toufeeq Ahmed"
},
{
"family": "Anibal",
"given": "James"
},
{
"family": "Bedrick",
"given": "Steven"
},
{
"family": "Bevers",
"given": "Isaac"
},
{
"family": "Boyer",
"given": "Micah"
},
{
"family": "Brito",
"given": "Rahul"
},
{
"family": "Casalino",
"given": "Selina A"
},
{
"family": "Costello",
"given": "John"
},
{
"family": "Diaz‐Ocampo",
"given": "Enrique"
},
{
"family": "Elmahdy",
"given": "Mahmoud"
},
{
"family": "Fletcher",
"given": "Kenneth"
},
{
"family": "Gelbard",
"given": "Alexander"
},
{
"family": "Hanna",
"given": "Karim"
},
{
"family": "Hersh",
"given": "Bill"
},
{
"family": "Jayachandran",
"given": "Lochana"
},
{
"family": "Jenney",
"given": "Kaley"
},
{
"family": "Krussel",
"given": "Andrea"
},
{
"family": "Loewith",
"given": "Chloe"
},
{
"family": "Neal",
"given": "Tempestt"
},
{
"family": "Premi‐Bortolotto",
"given": "Claire"
},
{
"family": "Rohde",
"given": "Sarah"
},
{
"family": "Cruz",
"given": "Samantha Salvi"
},
{
"family": "Silberholz",
"given": "Elizabeth"
},
{
"family": "Sutherland",
"given": "Duncan"
},
{
"family": "Talluri",
"given": "Venkata Swarna Mukhi"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"family": "Vinson",
"given": "Kimberly"
},
{
"family": "Wilson",
"given": "Claire"
},
{
"family": "Zanin",
"given": "Madeleine"
},
{
"family": "Zesiewicz",
"given": "Theresa"
},
{
"family": "Zhao",
"given": "Robin"
}
],
"container-title-short": "Laryngoscope Investig Otolaryngol",
"volume": "11",
"issue": "4",
"page": "e70519",
"DOI": "10.1002/lio2.70519",
"PMID": "42639494",
"PMCID": "PMC13502059",
"ISSN": "2378-8038",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/lio2.70519",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
24
]
]
}
}

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.2196/83790
Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review.
Journal: Journal of medical Internet research
In common: 1 reference, author Yael Bensoussan
[2] doi:10.1186/s12951-026-04551-7
The role of AI-assisted drug repurposing in neurological disorders: a systematic review of validation strategies, challenges and opportunities.
Journal: Journal of nanobiotechnology
In common: 2 references
[3] doi:10.1038/s44401-026-00111-1 [code]
Beyond fingerprint and voice print: cough sequence sound for identity verification.
Journal: npj health systems
In common: 2 references
[4] doi:10.1038/s43856-026-01817-x [code]
Visual prompt engineering for multimodal and irregularly sampled medical data.
Journal: Communications medicine
In common: 2 references
[5] doi:10.1371/journal.pcbi.1014302 [code]
Trial-level sequence modeling reveals hidden dynamics of dual-task interference.
Journal: PLoS computational biology
In common: 1 reference
[6] doi:10.1038/s42003-026-09938-8 [code]
Representation Transfer via Invariant Input-driven Continuous Attractors for Fast Domain Adaptation.
Journal: Communications biology
In common: 1 reference
[7] doi:10.3390/brainsci16080873
Evaluating Validation Strategies in Motor Imagery EEG: A Full-Cohort GAF-PLV Analysis and Matched Sensitivity Study.
Journal: Brain sciences
In common: 1 reference
[8] doi:10.1186/s40708-026-00320-2
Effiformer: a unified data-efficient vision transformer-CNN framework for interpretable epileptic seizure detection.
Journal: Brain informatics
In common: 1 reference
[9] doi:10.3390/bios16070394
Hybrid Edge-Cloud Asymmetric Analytics for Portable Multimodal BCI Biosensors.
Journal: Biosensors
In common: 1 reference
[10] doi:10.1371/journal.pone.0353930 [code]
MMFNet: A multi-branch multi-scale framework with adaptive sparse self-attention and cross-modal fusion for sleep stage assessment.
Journal: PloS one
In common: 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.