Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects.
Overview
and 43 other authors
Maria E Powell, Anaïs Rameau, Vardit Ravitsky, Alexandros Sigaras, Shaheen Awan, Ruth Bahr, Donald Bolser, Kathy J Jenkins, Frank Rudzicz, Jennifer Siu, Stephanie Watts, Yassmeen Abdel‐Aty, Toufeeq Ahmed Syed, James Anibal, Steven Bedrick, Isaac Bevers, Micah Boyer, Rahul Brito, Selina A Casalino, John Costello, Enrique Diaz‐Ocampo, Mahmoud Elmahdy, Kenneth Fletcher, Alexander Gelbard, Karim Hanna, Bill Hersh, Lochana Jayachandran, Kaley Jenney, Andrea Krussel, Chloe Loewith, Tempestt Neal, Claire Premi‐Bortolotto, Sarah Rohde, Samantha Salvi Cruz, Elizabeth Silberholz, Duncan Sutherland, Venkata Swarna Mukhi Talluri, Jamie Toghranegar, Kimberly Vinson, Claire Wilson, Madeleine Zanin, Theresa Zesiewicz, Robin Zhao- Computational Health Informatics Lab, Institute of Biomedical Engineering, University of Oxford, Oxford, UK
- Center for Interventional Oncology, NIH Clinical Center, National Institutes of Health, Bethesda, Maryland, USA
- Department of Computer Science and Engineering, University of South Florida, Tampa, Florida, USA
- USF Health Voice Center, Department of Otolaryngology‐Head & Neck Surgery, University of South Florida, Tampa, Florida, USA
- Vanderbilt University Medical Center, Nashville, Tennessee, USA
- Mount Sinai Hospital, Toronto, Ontario, Canada
Abstract
Objectives: Past studies have shown that many clinical machine learning models have performance limitations due to imbalances in the training data. For voice data generation projects, the origin of the problem may lie in the recruiting methods used during data collection efforts. This study introduces a generative AI pipeline for “dataset decision support”, recommending recruitment decisions based on high‐dimensional insights.
Methods: The publicly available GOSSIS‐1‐eICU dataset was filtered to create patient populations that were relevant to voice data generation projects. Lab results and vital signs from the electronic health record were also used to train a neural network for prediction of disease type. Prediction uncertainty estimates were included in the dataset as approximate indicators of health complexity. To select the best recruitment choice for addressing imbalances, an open‐source large language model (LLM) was then instructed to assess dataset statistics and the characteristics of possible participants. Simulations were run in which the system constructed datasets of 250 patients.
Results: In over 90% of cases, the proposed system reduced categorical imbalances and widened continuous distributions when compared to randomly sampled counterfactual datasets (q‐value < 0.05). Variables included race, age, BMI, sex, disease type, oxygenation status, co‐morbidities, post‐operative status, Glasgow Coma Scale verbal response score, the Acute Physiology Score III, prediction uncertainty, vital signs, and lab results.
Conclusion: LLMs may provide useful, explainable recommendations when presented with dataset distribution statistics and candidate profiles. In the future, this simulated scenario may be extended to align with conditions in emergency departments or other high‐volume settings.
Level of Evidence: 3.
Reproduced under the paper's license (CC BY), from the paper cited above.
Code
The paper links to its data, not to its authors' code: see the Data section.
Tracing map
A tracing map links a paper to the code its authors published: this paper has none, so it has no map.
Data
Datasets cited
- doi:10.13026/
gbmg-a531 , at the source; found in the references
Data Availability Statement
The authors have nothing to report.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 63 authors, 4 keywords, 1 funder, 59 references, 1 RRID.
Cite
This paper
Anibal, J., Nama, G. K. C., Ganesh, S., Nafii, Y., Cruz, S. S., Daoud, V., Zanin, M., Le, T., Khorrami, P., Wood, B. J., Clifton, D., Toghranegar, J., Bensoussan, Y., Bensoussan, Y. E., Elemento, O., Bélisle‐Pipon, J., Dorr, D., Ghosh, S., Johnson, A., . . . Zhao, R. (2026). Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects. Laryngoscope investigative otolaryngology, 11(4), e70519. https://
BibTeX
@article{anibal2026large
author = {Anibal, James and Nama, Geetha Krishna Chaitanya and Ganesh, Shrramana and Nafii, Yosef and Cruz, Samantha Salvi and Daoud, Veronica and Zanin, Madeleine and Le, Tram and Khorrami, Parsa and Wood, Bradford J and Clifton, David and Toghranegar, Jamie and Bensoussan, Yael and Bensoussan, Yael E and Elemento, Olivier and Bélisle‐Pipon, Jean‐Christophe and Dorr, David and Ghosh, Satrajit and Johnson, Alistair and Payne, Phillip and Powell, Maria E and Rameau, Anaïs and Ravitsky, Vardit and Sigaras, Alexandros and Awan, Shaheen and Bahr, Ruth and Bolser, Donald and Jenkins, Kathy J and Rudzicz, Frank and Siu, Jennifer and Watts, Stephanie and Abdel‐Aty, Yassmeen and Syed, Toufeeq Ahmed and Anibal, James and Bedrick, Steven and Bevers, Isaac and Boyer, Micah and Brito, Rahul and Casalino, Selina A and Costello, John and Diaz‐Ocampo, Enrique and Elmahdy, Mahmoud and Fletcher, Kenneth and Gelbard, Alexander and Hanna, Karim and Hersh, Bill and Jayachandran, Lochana and Jenney, Kaley and Krussel, Andrea and Loewith, Chloe and Neal, Tempestt and Premi‐Bortolotto, Claire and Rohde, Sarah and Cruz, Samantha Salvi and Silberholz, Elizabeth and Sutherland, Duncan and Talluri, Venkata Swarna Mukhi and Toghranegar, Jamie and Vinson, Kimberly and Wilson, Claire and Zanin, Madeleine and Zesiewicz, Theresa and Zhao, Robin},
title = {{Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects}},
journal = {Laryngoscope investigative otolaryngology},
year = {2026},
month = aug,
volume = {11},
number = {4},
pages = {e70519},
publisher = {Wiley},
issn = {2378-8038},
doi = {10.1002/
url = {https://
pmid = {42639494},
pmcid = {PMC13502059}
}
RIS
TY - JOUR
AU - Anibal, James
AU - Nama, Geetha Krishna Chaitanya
AU - Ganesh, Shrramana
AU - Nafii, Yosef
AU - Cruz, Samantha Salvi
AU - Daoud, Veronica
AU - Zanin, Madeleine
AU - Le, Tram
AU - Khorrami, Parsa
AU - Wood, Bradford J
AU - Clifton, David
AU - Toghranegar, Jamie
AU - Bensoussan, Yael
AU - Bensoussan, Yael E
AU - Elemento, Olivier
AU - Bélisle‐Pipon, Jean‐Christophe
AU - Dorr, David
AU - Ghosh, Satrajit
AU - Johnson, Alistair
AU - Payne, Phillip
AU - Powell, Maria E
AU - Rameau, Anaïs
AU - Ravitsky, Vardit
AU - Sigaras, Alexandros
AU - Awan, Shaheen
AU - Bahr, Ruth
AU - Bolser, Donald
AU - Jenkins, Kathy J
AU - Rudzicz, Frank
AU - Siu, Jennifer
AU - Watts, Stephanie
AU - Abdel‐Aty, Yassmeen
AU - Syed, Toufeeq Ahmed
AU - Anibal, James
AU - Bedrick, Steven
AU - Bevers, Isaac
AU - Boyer, Micah
AU - Brito, Rahul
AU - Casalino, Selina A
AU - Costello, John
AU - Diaz‐Ocampo, Enrique
AU - Elmahdy, Mahmoud
AU - Fletcher, Kenneth
AU - Gelbard, Alexander
AU - Hanna, Karim
AU - Hersh, Bill
AU - Jayachandran, Lochana
AU - Jenney, Kaley
AU - Krussel, Andrea
AU - Loewith, Chloe
AU - Neal, Tempestt
AU - Premi‐Bortolotto, Claire
AU - Rohde, Sarah
AU - Cruz, Samantha Salvi
AU - Silberholz, Elizabeth
AU - Sutherland, Duncan
AU - Talluri, Venkata Swarna Mukhi
AU - Toghranegar, Jamie
AU - Vinson, Kimberly
AU - Wilson, Claire
AU - Zanin, Madeleine
AU - Zesiewicz, Theresa
AU - Zhao, Robin
TI - Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects
T2 - Laryngoscope investigative otolaryngology
J2 - Laryngoscope Investig Otolaryngol
PY - 2026
DA - 2026/
VL - 11
IS - 4
SP - e70519
SN - 2378-8038
PB - Wiley
DO - 10.1002/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1002/
"type": "article-journal",
"title": "Large Language Models for Optimizing Patient Recruitment Decisions in Voice Data Generation Projects",
"container-title": "Laryngoscope investigative otolaryngology",
"author": [
{
"family": "Anibal",
"given": "James"
},
{
"family": "Nama",
"given": "Geetha Krishna Chaitanya"
},
{
"family": "Ganesh",
"given": "Shrramana"
},
{
"family": "Nafii",
"given": "Yosef"
},
{
"family": "Cruz",
"given": "Samantha Salvi"
},
{
"family": "Daoud",
"given": "Veronica"
},
{
"family": "Zanin",
"given": "Madeleine"
},
{
"family": "Le",
"given": "Tram"
},
{
"family": "Khorrami",
"given": "Parsa"
},
{
"family": "Wood",
"given": "Bradford J"
},
{
"family": "Clifton",
"given": "David"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"family": "Bensoussan",
"given": "Yael"
},
{
"family": "Bensoussan",
"given": "Yael E"
},
{
"family": "Elemento",
"given": "Olivier"
},
{
"family": "Bélisle‐Pipon",
"given": "Jean‐Christophe"
},
{
"family": "Dorr",
"given": "David"
},
{
"family": "Ghosh",
"given": "Satrajit"
},
{
"family": "Johnson",
"given": "Alistair"
},
{
"family": "Payne",
"given": "Phillip"
},
{
"family": "Powell",
"given": "Maria E"
},
{
"family": "Rameau",
"given": "Anaïs"
},
{
"family": "Ravitsky",
"given": "Vardit"
},
{
"family": "Sigaras",
"given": "Alexandros"
},
{
"family": "Awan",
"given": "Shaheen"
},
{
"family": "Bahr",
"given": "Ruth"
},
{
"family": "Bolser",
"given": "Donald"
},
{
"family": "Jenkins",
"given": "Kathy J"
},
{
"family": "Rudzicz",
"given": "Frank"
},
{
"family": "Siu",
"given": "Jennifer"
},
{
"family": "Watts",
"given": "Stephanie"
},
{
"family": "Abdel‐Aty",
"given": "Yassmeen"
},
{
"family": "Syed",
"given": "Toufeeq Ahmed"
},
{
"family": "Anibal",
"given": "James"
},
{
"family": "Bedrick",
"given": "Steven"
},
{
"family": "Bevers",
"given": "Isaac"
},
{
"family": "Boyer",
"given": "Micah"
},
{
"family": "Brito",
"given": "Rahul"
},
{
"family": "Casalino",
"given": "Selina A"
},
{
"family": "Costello",
"given": "John"
},
{
"family": "Diaz‐Ocampo",
"given": "Enrique"
},
{
"family": "Elmahdy",
"given": "Mahmoud"
},
{
"family": "Fletcher",
"given": "Kenneth"
},
{
"family": "Gelbard",
"given": "Alexander"
},
{
"family": "Hanna",
"given": "Karim"
},
{
"family": "Hersh",
"given": "Bill"
},
{
"family": "Jayachandran",
"given": "Lochana"
},
{
"family": "Jenney",
"given": "Kaley"
},
{
"family": "Krussel",
"given": "Andrea"
},
{
"family": "Loewith",
"given": "Chloe"
},
{
"family": "Neal",
"given": "Tempestt"
},
{
"family": "Premi‐Bortolotto",
"given": "Claire"
},
{
"family": "Rohde",
"given": "Sarah"
},
{
"family": "Cruz",
"given": "Samantha Salvi"
},
{
"family": "Silberholz",
"given": "Elizabeth"
},
{
"family": "Sutherland",
"given": "Duncan"
},
{
"family": "Talluri",
"given": "Venkata Swarna Mukhi"
},
{
"family": "Toghranegar",
"given": "Jamie"
},
{
"family": "Vinson",
"given": "Kimberly"
},
{
"family": "Wilson",
"given": "Claire"
},
{
"family": "Zanin",
"given": "Madeleine"
},
{
"family": "Zesiewicz",
"given": "Theresa"
},
{
"family": "Zhao",
"given": "Robin"
}
],
"container-title-short":
"volume": "11",
"issue": "4",
"page": "e70519",
"DOI": "10.1002/
"PMID": "42639494",
"PMCID": "PMC13502059",
"ISSN": "2378-8038",
"publisher": "Wiley",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
24
]
]
}
}
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.2196/83790
- Explainable and Interpretable AI for Voice and Speech Analysis in Clinical Care: Systematic Review.Journal: Journal of medical Internet researchIn common: 1 reference, author Yael Bensoussan
- [2] doi:10.1186/s12951-026-04551-7
- The role of AI-assisted drug repurposing in neurological disorders: a systematic review of validation strategies, challenges and opportunities.Journal: Journal of nanobiotechnologyIn common: 2 references
- [3] doi:10.1038/s44401-026-00111-1 [code]
- Beyond fingerprint and voice print: cough sequence sound for identity verification.Journal: npj health systemsIn common: 2 references
- [4] doi:10.1038/s43856-026-01817-x [code]
- Visual prompt engineering for multimodal and irregularly sampled medical data.Journal: Communications medicineIn common: 2 references
- [5] doi:10.1371/journal.pcbi.1014302 [code]
- Trial-level sequence modeling reveals hidden dynamics of dual-task interference.Journal: PLoS computational biologyIn common: 1 reference
- [6] doi:10.1038/s42003-026-09938-8 [code]
- Representation Transfer via Invariant Input-driven Continuous Attractors for Fast Domain Adaptation.Journal: Communications biologyIn common: 1 reference
- [7] doi:10.3390/brainsci16080873
- Evaluating Validation Strategies in Motor Imagery EEG: A Full-Cohort GAF-PLV Analysis and Matched Sensitivity Study.Journal: Brain sciencesIn common: 1 reference
- [8] doi:10.1186/s40708-026-00320-2
- Effiformer: a unified data-efficient vision transformer-CNN framework for interpretable epileptic seizure detection.Journal: Brain informaticsIn common: 1 reference
- [9] doi:10.3390/bios16070394
- Hybrid Edge-Cloud Asymmetric Analytics for Portable Multimodal BCI Biosensors.Journal: BiosensorsIn common: 1 reference
- [10] doi:10.1371/journal.pone.0353930 [code]
- MMFNet: A multi-branch multi-scale framework with adaptive sparse self-attention and cross-modal fusion for sleep stage assessment.Journal: PloS oneIn common: 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
