OSCR

Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § Methods and Materials › Ridge regression mapping from word embeddings to neural responses ↔ RidgeRegression.py, lines 41–130 · score 0.69 · ridge regression, cross validation, training, lagged, model, electrode
  2. [2] § Methods and Materials › Ridge regression mapping from word embeddings to neural responses ↔ ProcessECoGEmbeddings.py, lines 108–184 · score 0.66 · layer embeddings, word onset, neural responses, windows, lags, electrode
  3. [3] § Methods and Materials › Large language models to generate word representations ↔ WordEmbeddingGenerator.py, lines 148–233 · score 0.50 · attention mask, truncated, transcripts, tokens, Mistral, words

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 130 lines · 4.5 KB · no license · 1 match

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. """
  4. Ridge regression leave-one-out cross-validation for neural data alignment.
  5. This script:
  6. 1. Iterates over context lengths, subjects, and layers.
  7. 2. Performs leave-one-out cross-validation across trials.
  8. 3. Runs ridge regression with bootstrap-based model selection.
  9. 4. Saves raw trial-by-trial correlation results per subject.
  10. Dependencies:
  11. - numpy
  12. - matplotlib
  13. - scipy
  14. - ridge_utils (custom package with ridge + bootstrap_ridge)
  15. """
  16. import os
  17. import numpy as np
  18. from ridge_utils.ridge import bootstrap_ridge
  19. import matplotlib.pyplot as plt
  20. # ------------------------------------------------------------
  21. # CONFIGURATION
  22. # ------------------------------------------------------------
  23. CONTEXT_LENGTHS = ["full"]
  24. SAVE_BASE = "results_regression"
  25. DATA_BASE = "data_formatted_regression"
  26. N_SUBJECTS = 3
  27. N_LAYERS = 33
  28. N_TRIALS = 25
  29. # Ridge regression parameters
  30. ALPHAS = np.logspace(-2, 4, 8) # Regularization values (10^-2 ... 10^4)
  31. NBOOTS = 10 # Bootstrap resamples
  32. CHUNKLEN = 100 # Chunk length for bootstrap
  33. # ------------------------------------------------------------
  34. # MAIN SCRIPT
  35. # ------------------------------------------------------------
  36. print("Python script started...", flush=True)
  37. for context_len in CONTEXT_LENGTHS:
  38. save_dir = os.path.join(SAVE_BASE)
  39. if not os.path.exists(save_dir):
  40. print(f"ERROR: Saving path does not exist -> {save_dir}")
  41. break
  42. # Iterate over subjects
  43. for subject in range(1, N_SUBJECTS + 1):
  44. all_corrs_all_layers = []
  45. # Iterate over layers
  46. for layer_number in range(N_LAYERS):
  47. all_corrs = []
  48. # Leave-one-out cross-validation across trials
  49. for leave_out in range(N_TRIALS):
  50. acc_stim, acc_resp = [], []
  51. # Build training data (exclude one trial)
  52. for i in range(N_TRIALS):
  53. if i == leave_out:
  54. continue
  55. stim_path = os.path.join(
  56. DATA_BASE,
  57. f"Context_{context_len}_trimmed_4.0_unattended/Subject_{subject}/Layer_{layer_number}/Trial_{i}_embedding.npy"
  58. )
  59. resp_path = os.path.join(
  60. DATA_BASE,
  61. f"Context_{context_len}_trimmed_4.0_unattended/Subject_{subject}/Trial_{i}_response.npy"
  62. )
  63. acc_stim.append(np.load(stim_path))
  64. arr = np.load(resp_path)
  65. acc_resp.append(arr.reshape(arr.shape[0], -1))
  66. # Stack training data
  67. Rstim = np.vstack(acc_stim)
  68. Rresp = np.vstack(acc_resp)
  69. # Load held-out trial
  70. Pstim = np.load(
  71. os.path.join(
  72. DATA_BASE,
  73. f"Context_{context_len}_trimmed_4.0_unattended/Subject_{subject}/Layer_{layer_number}/Trial_{leave_out}_embedding.npy"
  74. )
  75. )
  76. arr = np.load(
  77. os.path.join(
  78. DATA_BASE,
  79. f"Context_{context_len}_trimmed_4.0_unattended/Subject_{subject}/Layer_{layer_number}/Trial_{leave_out}_response.npy"
  80. )
  81. )
  82. Presp = arr.reshape(arr.shape[0], -1)
  83. n_electrodes = arr.shape[1]
  84. # Define number of bootstrap chunks
  85. nchunks = int(len(Rresp) * 0.25 / CHUNKLEN)
  86. # Debug shapes
  87. print(Rstim.shape, Pstim.shape, Rresp.shape, Presp.shape, flush=True)
  88. # Run ridge regression with bootstrap-based model selection
  89. _, corr, _, _, _ = bootstrap_ridge(
  90. Rstim, Rresp, Pstim, Presp,
  91. alphas=ALPHAS, nboots=NBOOTS,
  92. chunklen=CHUNKLEN, nchunks=nchunks,
  93. use_corr=True, single_alpha=False
  94. )
  95. # Reshape correlations to (electrodes × lags)
  96. all_corrs.append(np.reshape(corr, (n_electrodes, 7)))
  97. # Collect results for this layer
  98. all_corrs_all_layers.append(all_corrs)
  99. print(f"Finished Layer {layer_number}, Subject {subject}, Context {context_len}", flush=True)
  100. # Save results for this subject
  101. save_path = os.path.join(save_dir, f"Subject_{subject}_all_layers.npy")
  102. np.save(save_path, all_corrs_all_layers)
  103. print(f"✅ Saved results: {save_path}")
  104. print("Script completed successfully.")

RidgeRegression.py at commit 30c0e1a, no license · at the source

Overview

  1. KU Leuven, Department of Neurosciences, ExpORL, Leuven, Belgium
  2. KU Leuven, Department of Electrical Engineering (ESAT), PSI, Leuven, Belgium
  3. Columbia University, Department of Electrical Engineering, New York, NY, United States
  4. Hofstra Northwell School of Medicine, Uniondale, NY, United States
  5. The Feinstein Institutes for Medical Research, Manhasset, NY, United States
  6. Columbia University, Department of Neurology, New York, NY, United States
  7. Columbia University, Department of Neurological Surgery, Vagelos College of Physicians and Surgeons, New York, NY, United States
Institutions: KU Leuven (Belgium); Columbia University (United States); Feinstein Institute for Medical Research (United States); Hofstra University (United States)
Journal: Imaging neuroscience (Cambridge, Mass.), volume 4, article IMAG.a.1227
Dates: received 9 September 2025; accepted 10 April 2026; published online 7 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1162/imag.a.1227 · PMID 42112161 · PMCID PMC13155409 · OpenAlex W4409824180
Open access: diamond, a free copy (OpenAlex)
Status: code verified
Categories: intracranial EEG (iEEG / ECoG / SEEG) (modality), cognitive (subfield)
Methods: Preprocessing, Connectivity, Machine learning, Statistics, Spectral & time-frequency
Keywords: auditory attention, large language models, context encoding
Topic: Speech Recognition and Synthesis (Artificial Intelligence, Computer Science), according to OpenAlex
Funding: Research Foundation Flanders (1S49823N, 1290821N); National Institute of Health; National Institute on Deafness and Other Communication Disorders; Marie-Josee and Henry R. Kravis Foundation
Citations: not cited yet (Europe PMC); 35 references in the paper

Abstract

Large language models (LLMs) capture long-range contextual structure in natural language and have recently been shown to align with the human brain’s contextualized linguistic encoding. This makes them a promising computational probe for studying how context-dependent linguistic information is represented during natural speech perception. Speech perception often occurs in multi-talker environments, where attention must dynamically select among competing streams, yet how contextual information from attended and unattended speech is neurally encoded remains underexplored. Here, we investigate how auditory attention modulates neural tracking of context-dependent linguistic representations using electrocorticography (ECoG) and stereoelectroencephalography (sEEG) recordings from three epilepsy patients engaged in a two-conversation “cocktail party” paradigm. To model neural responses to attended and unattended speech streams, we used contextual word embeddings generated by large language models. We find that LLM-derived features reliably predict brain activity for the attended stream and that contextual information from the unattended stream also contributes to neural prediction. Importantly, these contributions extend beyond low-level acoustic features and shallow syntactic information, and depend on the surrounding linguistic context. Moreover, neural tracking of the unattended stream reflects shorter-range contextual integration than that of the attended stream. Together, these findings indicate that neural responses to speech reflect context-dependent linguistic representations from multiple concurrent speech streams, with attention modulating the depth and timescale of contextual integration. Our results highlight the utility of LLMs for probing higher-level linguistic representations in complex, naturalistic listening environments.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

corentinpuffay/LLM_ECoG

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 30c0e1ac05883eeca8779a1dbe31f3a4e4ce5453, 25 August 2025
Languages: Python (5)
Size: 14 files, 5 scripts
Software Heritage: not archived
Found in: “Data and Code Availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (5 files), Matplotlib (3 files), PyTorch (3 files), SciPy (3 files), pandas (2 files), seaborn (2 files), Hugging Face Transformers (2 files), h5py (1 file), MNE-Python (1 file), scikit-learn (1 file), SymPy (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
6 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data and Code Availability

The data that support the findings of this study are available on request from the corresponding author (N.M.). The data are not publicly available due to privacy or ethical restrictions. The code base to run our analyses is available on GitHub: https://github.com/corentinpuffay/LLM_ECoG.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Authors: added Gavin Mischler (0000-0003-4776-3518); Vishal Choudhari (0009-0000-5486-5913); Jonas Vanthornhout (0000-0002-1503-599X); Ashesh D. Mehta (0000-0001-7293-1101); Catherine Schevon (0000-0002-4485-7933); Guy M. McKhann (0000-0002-9695-3564); Tom Francart (0000-0001-9734-4261); Nima Mesgarani (0000-0002-2987-759X); removed Gavin Mischler; Vishal Choudhari; Jonas Vanthornhout; Ashesh D. Mehta; Catherine Schevon; Guy M. McKhann; Tom Francart; Nima Mesgarani

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 11 authors, 3 keywords, 4 funders, 32 references.

Cite

This paper

Puffay, C., Mischler, G., Choudhari, V., Vanthornhout, J., Bickel, S., Mehta, A. D., Schevon, C., McKhann, G. M., Van hamme, H., Francart, T., & Mesgarani, N. (2026). Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech. Imaging neuroscience (Cambridge, Mass.), 4, IMAG.a.1227. https://doi.org/10.1162/imag.a.1227

BibTeX

@article{puffay2026large,
author = {Puffay, Corentin and Mischler, Gavin and Choudhari, Vishal and Vanthornhout, Jonas and Bickel, Stephan and Mehta, Ashesh D. and Schevon, Catherine and McKhann, Guy M. and Van hamme, Hugo and Francart, Tom and Mesgarani, Nima},
title = {{Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech}},
journal = {Imaging neuroscience (Cambridge, Mass.)},
year = {2026},
month = may,
volume = {4},
pages = {IMAG.a.1227},
publisher = {MIT Press},
issn = {2837-6056},
doi = {10.1162/imag.a.1227},
url = {https://doi.org/10.1162/imag.a.1227},
pmid = {42112161},
pmcid = {PMC13155409}
}

RIS

TY - JOUR
AU - Puffay, Corentin
AU - Mischler, Gavin
AU - Choudhari, Vishal
AU - Vanthornhout, Jonas
AU - Bickel, Stephan
AU - Mehta, Ashesh D.
AU - Schevon, Catherine
AU - McKhann, Guy M.
AU - Van hamme, Hugo
AU - Francart, Tom
AU - Mesgarani, Nima
TI - Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech
T2 - Imaging neuroscience (Cambridge, Mass.)
J2 - Imaging Neurosci (Camb)
PY - 2026
DA - 2026/05/07
VL - 4
SP - IMAG.a.1227
SN - 2837-6056
PB - MIT Press
DO - 10.1162/imag.a.1227
UR - https://doi.org/10.1162/imag.a.1227
LA - en
ER -

CSL-JSON

{
"id": "10.1162/imag.a.1227",
"type": "article-journal",
"title": "Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech",
"container-title": "Imaging neuroscience (Cambridge, Mass.)",
"author": [
{
"family": "Puffay",
"given": "Corentin"
},
{
"family": "Mischler",
"given": "Gavin"
},
{
"family": "Choudhari",
"given": "Vishal"
},
{
"family": "Vanthornhout",
"given": "Jonas"
},
{
"family": "Bickel",
"given": "Stephan"
},
{
"family": "Mehta",
"given": "Ashesh D."
},
{
"family": "Schevon",
"given": "Catherine"
},
{
"family": "McKhann",
"given": "Guy M."
},
{
"family": "Van hamme",
"given": "Hugo"
},
{
"family": "Francart",
"given": "Tom"
},
{
"family": "Mesgarani",
"given": "Nima"
}
],
"container-title-short": "Imaging Neurosci (Camb)",
"volume": "4",
"page": "IMAG.a.1227",
"DOI": "10.1162/imag.a.1227",
"PMID": "42112161",
"PMCID": "PMC13155409",
"ISSN": "2837-6056",
"publisher": "MIT Press",
"URL": "https://doi.org/10.1162/imag.a.1227",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
7
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.7554/elife.106543 [code]
Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.
Journal: eLife
In common: Hugging Face Transformers, MNE-Python, h5py, 7 other tools, cognitive, 6 references
[2] doi:10.1038/s41467-026-72253-7 [code]
Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds.
Journal: Nature communications
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 5 references
[3] doi:10.7554/elife.101204 [code]
Larger language models better align with neural representations of natural language.
Journal: eLife
In common: Hugging Face Transformers, PyTorch, scikit-learn, 3 other tools, intracranial EEG (iEEG / ECoG / SEEG), 5 references
[4] doi:10.1371/journal.pbio.3003876 [code]
Competing speech streams are simultaneously represented in the human cortex during attention switching.
Journal: PLoS biology
In common: cognitive, 9 references
[5] doi:10.7554/elife.100056 [code]
Multi-talker speech comprehension at different temporal scales in listeners with normal and impaired hearing.
Journal: eLife
In common: MNE-Python, PyTorch, scikit-learn, 3 other tools, cognitive, 5 references
[6] doi:10.1038/s41467-026-75240-0 [code]
Dynamic acoustic-to-categorical representations of phonemes and prosody along ventral and dorsal speech streams.
Journal: Nature communications
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, cognitive, 4 references
[7] doi:10.1038/s41597-025-06397-4 [code]
MEG-SCANS - A comprehensive magnetoencephalography speech dataset with Stories, Chirps and Noisy Sentences
Journal: n/a
In common: MNE-Python, pandas, SciPy, 2 other tools, 5 references
[8] doi:10.1016/j.isci.2026.117180 [code]
Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.
Journal: iScience
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 2 references
[9] doi:10.1162/nol.a.244 [code]
A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: Hugging Face Transformers, MNE-Python, PyTorch, 4 other tools, 3 references
[10] doi:10.1038/s42003-026-10169-0 [code]
Shared representations in brains and models reveal a two-route cortical organization during scene perception.
Journal: Communications biology
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, cognitive, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.