OSCR

Optimized feature gains explain and predict successes and failures of human selective listening.

Code ↔ Paper

17 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 17 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Experiment 1: effect of distractors on attentional selection in monaural conditions › Stimuli ↔ corpus/saddler_word_rec.py, the whole file · a weak match · score 0.88 · IEEE AASP CASA, Spoken Wikipedia, Common Voice, Saddler, babble, music
  2. [2] § Methods › Cochlear model stage ↔ src/audio_attention_transforms.py, lines 292–346 · score 0.77 · gammatone filter bank, half wave, convolved, domain, compressive, downsampled
  3. [3] § Methods › Cochlear model stage ↔ src/audio_transforms.py, lines 309–360 · score 0.77 · gammatone filter bank, half wave, convolved, domain, compressive, downsampled
  4. [4] § Methods › Cochlear model stage ↔ src/time_domain_cochleagram.py, lines 15–82 · score 0.76 · compressive nonlinearity, impulse response, filter bank, downsampled, cochlea, channels
  5. [5] § Methods › Experiment 2: effect of harmonicity on attentional selection in monaural conditions › Stimuli ↔ src/get_swc_popham_expmt_stim_2024.py, lines 100–237 · score 0.72 · jitter pattern, Whispered speech, STRAIGHT, sinusoidal, cut, Inharmonic
  6. [6] § Methods › Experiment 2: effect of harmonicity on attentional selection in monaural conditions › Stimuli ↔ src/get_swc_popham_expmt_stim_2024_for_model.py, lines 81–225 · score 0.71 · jitter pattern, Whispered speech, STRAIGHT, sinusoidal, Inharmonic, harmonics
  7. [7] § Methods › Cochlear model stage ↔ src/audio_attention_transforms.py, lines 292–346 · score 0.71 · gammatone filter bank, half wave, compressive, downsampled, waveform, cochlea
  8. [8] § Results › The model replicates effects of speech harmonicity ↔ src/get_swc_popham_expmt_stim_2024_for_model.py, lines 81–225 · score 0.67 · speech harmonicity, whispered speech, cue signals, human word, inharmonic, talker
  9. [9] § Methods › Training data generation › Signal augmentations ↔ src/get_swc_popham_expmt_stim_2024.py, lines 100–237 · score 0.65 · target signals, RMS normalized, cue signal, onset, cut, duration
  10. [10] § Methods › Training data generation › Signal augmentations ↔ src/get_swc_unfamiliar_distractor_stim_2024.py, lines 50–135 · score 0.64 · target signals, RMS normalized, cue signal, onset, cut, duration
  11. [11] § Methods › Training data generation › Speech corpora ↔ corpus/saddler_word_rec.py, the whole file · a weak match · score 0.60 · Common Voice, word classes, word recognition, corpus, excerpts, train
  12. [12] § Methods › Experiment 1: effect of distractors on attentional selection in monaural conditions › Stimuli ↔ src/get_swc_mono_expmt_stim_2024.py, lines 14–144 · score 0.56 · talker Mandarin, distractor signals, babble, music, English, scenes
  13. [13] § Methods › Experiment 5: the precedence effect with concurrent speech signals › Model experiment ↔ src/eval_precedence.py, lines 107–174 · score 0.53 · anechoic room, lead channel, precedence, binaural, location, SNRs
  14. [14] § Methods › Artificial neural network constituent operations › Weighted-average pooling ↔ src/custom_modules.py, lines 23–75 · score 0.52 · channel dimension, Hanning, windows, stride, Weighted, convolution
  15. [15] § Methods › Experiment 6: effect of spatial separation in azimuth and elevation › Procedure ↔ src/get_swc_unfamiliar_distractor_stim_2024.py, lines 50–135 · score 0.51 · combined signal, cue signal, SPL, dB, excerpt, stimuli
  16. [16] § Methods › Analysis of model locus of attention ↔ src/eval_symmetric_distractors.py, lines 99–175 · score 0.51 · RMS normalized, single distractor, composed, excerpts, signals, stimuli
  17. [17] § Methods › Experiment 4: spatial tuning for masked speech › Model experiment ↔ notebooks/Paper_Figs_Raw_Analysis/Byrne_et_al_model_simulation_all_arch.ipynb, lines 368–466 · score 0.50 · model simulation, Model thresholds, fitting, 18 dB, sex, azimuth

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 92 lines · 3.7 KB · MIT · 2 matches

  1. # write torch dataloader for test set
  2. import torch
  3. import pandas as pd
  4. import librosa
  5. import pickle
  6. from pathlib import Path
  7. class SaddlerSWCWordRecTest(torch.utils.data.Dataset):
  8. """
  9. Dataset for the Saddler word recognition experiment using
  10. foreground excerpts from Spoken Wikipedia. Backgrounds are
  11. either:
  12. - music from musdb
  13. - 8 talkerbabble from common voice
  14. - spectrally matched noise (SSN; match each foreground)
  15. - Festen and Plomp style modulated maskers
  16. - audioset
  17. - natural scenes from IEEE AASP CASA challenge
  18. - clean (no background)
  19. """
  20. def __init__(self, manifest_path, bg_stim_path, condition, label_type="WSN", sr=20_000):
  21. """
  22. Args:
  23. manifest_path (str): path to pandas manifest with fg and cue excerpts
  24. bg_stim_path (str): path to directory with background stimuli
  25. condition (str): background condition. Either "music", "babble", "stationary", "modulated", "audioset", "ieee_scenes", or "clean"
  26. label_type (str): Set of word class labels to use. Either "WSN" (JSIN) or "CV" common voice
  27. sr (int): sampling rate to load audio at
  28. """
  29. self.manifest = pd.read_pickle(manifest_path)
  30. self.condition = condition
  31. self.label_type = label_type
  32. self.sr = sr
  33. self.condition_dict = {'music':"background_musdb18hq",
  34. "babble":"background_cv08talkerbabble",
  35. "stationary": "background_issnstationary",
  36. "modulated": "background_issnfestenplomp",
  37. "audioset": "background_audioset",
  38. "ieee_scenes": "background_ieeeaaspcasa",
  39. }
  40. if condition == "clean":
  41. self.bg_stim = None
  42. self.test_cond_dir = None
  43. else:
  44. self.test_cond_dir = self.condition_dict[self.condition]
  45. self.bg_stim = list((bg_stim_path / self.test_cond_dir).glob("*.wav"))
  46. self.class_map, self.word_2_class = self.class_map()
  47. self.dataset_len = len(self.manifest)
  48. def class_map(self):
  49. """
  50. Loads the mapping between the word IDX and human readable word map.
  51. """
  52. if self.label_type == "WSN":
  53. ## load WSN vocab mapping
  54. word_and_speaker_encodings = pickle.load( open( "/om2/user/imgriff/projects/Auditory-Attention/word_and_speaker_encodings_jsinv3.pckl", "rb" ))
  55. class_map = word_and_speaker_encodings['word_idx_to_word']
  56. elif self.label_type == "CV":
  57. class_map = pickle.load( open("/om2/user/imgriff/datasets/commonvoice_9/en/cv_800_word_label_to_int_dict.pkl", "rb" ))
  58. word_2_class = {v:k for k,v in class_map.items()}
  59. return class_map, word_2_class
  60. def __getitem__(self, index):
  61. """
  62. Gets components of the hdf5 file that are used for training
  63. Args:
  64. index (int): index into the hdf5 file
  65. Returns:
  66. [signal, target] : the training audio (signal) containing the preprocessing
  67. which may combine the foreground and background speech, and the target idx
  68. specified by target_keys.
  69. """
  70. foreground, _ = librosa.load(self.manifest['src_fn'][index], sr=self.sr)
  71. cue, _ = librosa.load(self.manifest['cue_src_fn'][index], sr=self.sr)
  72. if self.condition == "clean":
  73. background = None
  74. else:
  75. background, _ = librosa.load(self.bg_stim[index], sr=self.sr)
  76. word = self.manifest['word'][index]
  77. word_label = self.word_2_class[word]
  78. return cue, foreground, background, word_label
  79. def __len__(self):
  80. return self.dataset_len

saddler_word_rec.py at commit bb61769, under MIT · at the source

Overview

  1. Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology,Cambridge, MA USA
  2. McGovern Institute for Brain Research, Massachusetts Institute of Technology,Cambridge, MA USA
  3. Program in Speech and Hearing Biosciences and Technology, Harvard University,Cambridge, MA USA
  4. Center for Brains, Minds, and Machines, Massachusetts Institute of Technology,Cambridge, MA USA
Institutions: Massachusetts Institute of Technology (United States); Harvard University (United States)
Journal: Nature human behaviour, volume 10, issue 5, pages 937-959
Dates: received 3 July 2025; accepted 20 January 2026; published online 13 March 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41562-026-02414-7 · PMID 41826525 · PMCID PMC13192276 · OpenAlex W7135192550
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cognitive (subfield)
Methods: Spectral & time-frequency, Connectivity, Statistics, Graphs, Machine learning
Keywords: Human behaviour, Auditory system, Computational neuroscience
MeSH: Attention*, Auditory Perception*, Neural Networks, Computer*, Speech Perception*, Acoustic Stimulation, Adult, Cues, Female, Humans, Male (* major topic)
Topic: Neuroscience and Music Perception (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: NIH (R01DC017970)
Citations: cited by 5 papers (Europe PMC); 94 references in the paper

Abstract

Attention facilitates communication by enabling selective listening to sound sources of interest. However, little is known about why attentional selection succeeds in some conditions but fails in others. While neurophysiology implicates multiplicative feature gains in selective attention, it is unclear whether such gains can explain real-world attention-driven behaviour. Here we optimized an artificial neural network with stimulus-computable feature gains to recognize a cued talker’s speech from binaural audio in ‘cocktail party’ scenarios. Though not trained to mimic humans, the model produced human-like performance across diverse real-world conditions, exhibiting selection based both on voice qualities and on spatial location as well as selection failures in conditions where humans tended to fail. It also predicted novel attentional effects that we confirmed in human experiments, and exhibited signatures of ‘late selection’ like those seen in human auditory cortex. The results suggest that human-like attentional strategies naturally arise from the optimization of feature gains for selective listening.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.

mcdermottLab/auditory_attention

License: MIT
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: bb617690c0e2f7713721650e8c1e164122874b39, 31 March 2026
Languages: Python (61), Jupyter (49), Shell (32)
Size: 166 files, 142 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (requirements.txt), 49 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (99 files), pandas (91 files), Matplotlib (64 files), seaborn (63 files), SciPy (39 files), PyTorch (28 files), h5py (21 files), Pingouin (15 files), statsmodels (8 files), scikit-learn (7 files), PyTorch Lightning (6 files), TensorFlow (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
144 files

OSF wjzvu

License: none: the authors keep all their rights
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 3 files
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
At the source: osf.io/wjzvu

Code availability

The code used for modelling and data analysis in this study is available via GitHub at https://github.com/mcdermottLab/auditory_attention and is linked from the project OSF repository93.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 142 scripts, each with its path and the digest of its content;
  • 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The processed, anonymized human data and simulation data from this study are available via OSF93. There are no restrictions on data availability, and all relevant files are provided in CSV format. No data with mandated deposition are included in this study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 3 keywords, 10 MeSH terms, 1 funder, 78 references.

Cite

This paper

Griffith, I. M., Hess, R. P., & McDermott, J. H. (2026). Optimized feature gains explain and predict successes and failures of human selective listening. Nature human behaviour, 10(5), 937-959. https://doi.org/10.1038/s41562-026-02414-7

BibTeX

@article{griffith2026optimized,
author = {Griffith, Ian M. and Hess, R. Preston and McDermott, Josh H.},
title = {{Optimized feature gains explain and predict successes and failures of human selective listening}},
journal = {Nature human behaviour},
year = {2026},
month = mar,
volume = {10},
number = {5},
pages = {937--959},
publisher = {Nature Portfolio},
issn = {2397-3374},
doi = {10.1038/s41562-026-02414-7},
url = {https://doi.org/10.1038/s41562-026-02414-7},
pmid = {41826525},
pmcid = {PMC13192276}
}

RIS

TY - JOUR
AU - Griffith, Ian M.
AU - Hess, R. Preston
AU - McDermott, Josh H.
TI - Optimized feature gains explain and predict successes and failures of human selective listening
T2 - Nature human behaviour
J2 - Nat Hum Behav
PY - 2026
DA - 2026/03/13
VL - 10
IS - 5
SP - 937
EP - 959
SN - 2397-3374
PB - Nature Portfolio
DO - 10.1038/s41562-026-02414-7
UR - https://doi.org/10.1038/s41562-026-02414-7
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41562-026-02414-7",
"type": "article-journal",
"title": "Optimized feature gains explain and predict successes and failures of human selective listening",
"container-title": "Nature human behaviour",
"author": [
{
"family": "Griffith",
"given": "Ian M."
},
{
"family": "Hess",
"given": "R. Preston"
},
{
"family": "McDermott",
"given": "Josh H."
}
],
"container-title-short": "Nat Hum Behav",
"volume": "10",
"issue": "5",
"page": "937-959",
"DOI": "10.1038/s41562-026-02414-7",
"PMID": "41826525",
"PMCID": "PMC13192276",
"ISSN": "2397-3374",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41562-026-02414-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
13
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-72146-9 [code]
Modeling attention and binding in the brain through bidirectional recurrent gating.
Journal: Nature communications
In common: PyTorch, seaborn, SciPy, 2 other tools, 11 references
[2] doi:10.7554/elife.100056 [code]
Multi-talker speech comprehension at different temporal scales in listeners with normal and impaired hearing.
Journal: eLife
In common: PyTorch, scikit-learn, pandas, 2 other tools, cognitive, 7 references
[3] doi:10.1371/journal.pbio.3003707 [code]
Behavioral engagement facilitates auditory neuron responses beyond their receptive fields.
Journal: PLoS biology
In common: 9 references
[4] doi:10.1038/s41467-026-75455-1 [code]
Shared latent representations of speech production for cross-patient speech decoding.
Journal: Nature communications
In common: PyTorch Lightning, TensorFlow, h5py, 8 other tools, cognitive
[5] doi:10.1371/journal.pbio.3003876 [code]
Competing speech streams are simultaneously represented in the human cortex during attention switching.
Journal: PLoS biology
In common: cognitive, 8 references
[6] doi:10.1016/j.isci.2026.117375 [code]
Motor priming is associated with widespread recruitment into neural ensembles and more rapid ensemble transitions.
Journal: iScience
In common: Pingouin, h5py, statsmodels, 7 other tools, 2 references
[7] doi:10.7554/elife.105953 [code]
Top-down feedback in deep neural networks leads to functional differences during audiovisual integration.
Journal: eLife
In common: PyTorch Lightning, h5py, PyTorch, 6 other tools, 2 references
[8] doi:10.1016/j.isci.2026.117180 [code]
Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.
Journal: iScience
In common: h5py, statsmodels, PyTorch, 6 other tools, 2 references
[9] doi:10.3389/fnins.2026.1756386 [code]
Attention to speech modulates distortion product otoacoustic emissions evoked by speech-derived stimuli in humans.
Journal: Frontiers in neuroscience
In common: statsmodels, seaborn, pandas, 3 other tools, cognitive, 4 references
[10] doi:10.1371/journal.pcbi.1014571 [code]
SynAPSeg: A novel dataset and image analysis framework for deep learning-based synapse detection and quantification.
Journal: PLoS computational biology
In common: Pingouin, TensorFlow, statsmodels, 7 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.