Quantifying the Influence of Lexical Surprisal on Acoustic Speech Encoding While Controlling for Within-Speaker Variability.
The 8 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Methods › Data Acquisition and Preprocessing ↔ Code/gmdlDataset_preprocess.m, lines 1–68 · score 0.88 · passband attenuation, stopband attenuation, Noisy channels, deviant trials, kurtosis, downsampled
- [2] § Results › EEG Signatures of Acoustic Speech Processing Are Stronger for Less Predictable Words, Beyond Variations in Speaker Articulation ↔ Code/gmdlDataset_fw_lme_4D_array_calc.m, lines 1–36 · score 0.69 · 100–200 ms, 0–100 ms, stage regression, NP surprisal, window, LME
- [3] § Methods › Modeling the Relationship Between Speech Features and EEG Responses ↔ Code/gmdlDataset_fw_lme_4D_array_calc.m, lines 1–36 · score 0.66 · 100–300 ms, stage regression, surprisal models, forward models, trained, 100 ms
- [4] § Results › EEG Signatures of Acoustic Speech Processing Are Stronger for Less Predictable Words, Beyond Variations in Speaker Articulation ↔ Code/gmdlDataset_fw_lme_prep.m, lines 1–34 · score 0.63 · 100–200 ms, 0–100 ms, LME model, NP surprisal, window, envelope
- [5] § Results › EEG Signatures of Acoustic Speech Processing Are Stronger for Less Predictable Words, Beyond Variations in Speaker Articulation ↔ Code/singleWordProsody.m, the whole file · a weak match · score 0.61 · single word, 0–100 ms, resolvability, rows, onset, envelope
- [6] § Methods › Data Acquisition and Preprocessing ↔ Code/gmdlDataset_preprocess.m, lines 71–181 · score 0.56 · EEG space, PCA, denoise, MCCA, Preprocessing, neural
- [7] § Methods › Assessing the Influence of Context‐Based Predictions on Acoustic Encoding ↔ Code/gmdlDataset_fw_mod.m, lines 1–46 · score 0.54 · word onset, 100–300 ms, speech envelope, lags, 100 ms, surprisal
- [8] § Methods › Modeling the Relationship Between Speech Features and EEG Responses ↔ Code/gmdlDataset_fw_mod.m, lines 1–46 · score 0.53 · 100–700 ms, 100–300 ms, lags, trained, TRF, envelope
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
MATLAB · 181 lines · 6.6 KB · no license · 2 matches
- %% Shyanthony Synigal Fall 2023
- % Preprocessing G.M. di Liberto's 2018 vocoding experiment data (his raw data)
- % Applying MCCA using standard trials only
- % Uses kurtosis, spectra, and probability to identify noisy channels
- % reref to the average of all channels
- % Dependencies: EEGLAB, Fieldtrip (for filtfilthd), NoiseTools (Alain de Cheveigne's toolbox)
- % This script loads the CND data
- % My data wasn't originally in CND format when I analyzed it and saved it
- % for the first time, so I tried to write it in at the end
- study_dir = pwd;
- eegRaw_dir = [study_dir '\rawNeural\'];
- eeg_dir = [study_dir '\dataCND\'];
- mcca_dir = [eeg_dir '\MCCA\'];
- dev_dir = [study_dir '\Data\'];
- if ~exist(eeg_dir,'dir'), mkdir(eeg_dir); end
- if ~exist(mcca_dir,'dir'), mkdir(mcca_dir); end
- % Bandpass filter
- fs = 512; % Sampling frequency (Hz)
- fs_new = 128; % what we'll downsample to
- fstop1 = 0.5; % Lower stopband frequency (Hz)
- fpass1 = 1; % Lower passband frequency (Hz)
- astop1 = 60; % Stopband attenuation (dB)
- apass = 1; % Passband attenuation (dB)
- fpass2 = 8; % Upper passband frequency (Hz)
- fstop2 = 8.5; % Upper stopband frequency (Hz)
- astop2 = 80; % Stopband attenuation (dB)
- h = fdesign.highpass(fstop1,fpass1,astop1,apass,fs);
- hpf = design(h,'cheby2','MatchExactly','stopband'); clear h
- h = fdesign.lowpass(fpass2,fstop2,apass,astop2,fs);
- lpf = design(h,'cheby2','MatchExactly','stopband'); clear h
- clear fpass2 fstop2 apass astop2 fstop1 fpass1 astop1 h
- % Other info: subjects, conditions, etc.
- subs = 1:14;
- nsubs = length(subs);
- conditions = {'np','c','p'}; % the presentation order
- ncond = numel(conditions);
- nchans = 128;
- badchans_all = cell(1,nsubs);
- tsec = 10; % 10 sec trials
- tlen = fs_new*tsec; %% of samp for 10 sec trials
- nPCs = 40; % number of PCs
- nMCCs = 110; % numbr of CCs
- % load([dev_dir 'chanlocs.mat']) % also saved in CND file
- load([dev_dir 'Deviants.mat']);
- standards = 1:120;
- standards(ismember(standards,deviants))=[]; %removing the deviant trials
- ntrials = numel(standards); % should be 93;
- %% Preprocess
- for cond = 1:ncond
- disp(['** Processing condition' int2str(cond) ' **']);
- for s = 1:nsubs
- allEEG = zeros(tlen,nchans,ntrials); %needs to be time x chann x trials
- load([eegRaw_dir, 'dataSub', int2str(s), '_', conditions{cond}, '.mat'],'neural')
- row = 1;
- for trial = standards
- % using standard trials only. even for clean condition.
- disp(['subject ' int2str(s) ' trial ' int2str(trial)])
- load([eegRaw_dir 'dataSub' int2str(s) '_' conditions{cond} '.mat'],'neural');
- eeg = neural.data{trial}; % [neural.data{trial} neural.extChan{1}.data{trial}]; % time x chan
- %'** Filtering **'
- EEG = filtfilthd(hpf,eeg');
- EEG = filtfilthd(lpf,EEG);
- EEG = EEG(:,1:nchans);
- % What Giovanni used to find trial beginning and endings
- startSample = neural.trialStart(1,trial);
- endSample = neural.trialEnd(1,trial);
- EEG = EEG(startSample:endSample,:); %still tim x chan
- % mastoids = mastoids(startSample:endSample,:);
- clear startSample endSample
- % '** Interpolating bad channels **' % Spline interpolate bad channels (time x chans) w/ eeglab
- EEGstruct = create_eegstruct(EEG,fs,neural.chanlocs); % input data: time x chann
- elecind1 = []; elecind2 = []; elecind3 = []; badchans = []; elecind4 = []; % for bad channels
- [~,elecind1,~,~] = pop_rejchan(EEGstruct, 'elec',1:nchans,'threshold',10,'norm','on','measure','kurt');
- [~,elecind2,~,~] = pop_rejchan(EEGstruct, 'elec',1:nchans,'threshold',5,'norm','on','measure','spec');
- [~,elecind3,~,~] = pop_rejchan(EEGstruct, 'elec',1:nchans,'threshold',10,'norm','on','measure','prob');
- elecind4 = findBadChansMean(EEGstruct,3);
- badchans = unique([elecind1 elecind2 elecind3 elecind4]); % indecies of bad channels
- badchans_all{row,s} = badchans;
- EEGstruct.badchans = badchans; % keep info about bad channels
- EEGstruct = eeg_interp(EEGstruct,badchans); % do and save interp data
- data = double(EEGstruct.data'); clear EEG eeg % time x chan x trials
- % '** Downsampling **'
- eegData = downsample(data,(fs/fs_new)); %still time x chan
- % '** Rereferening to global avg **'
- % input data = chan x time, so transpose
- eegData = reref(eegData',[],'keepref','off');
- allEEG(1:length(eegData),:,row) = eegData'; %needs to be time x chann x trials
- row = row+1;
- clear badchans data eegData %mastoids
- end
- %
- disp('** Run PCA and save **')
- % save each subject's data here and run MCCA on all subjects at once
- % xAll = sub x time x PC x trial; topcs = sub x chan x chan
- [xAll(s,:,:,:), topcs(s,:,:)] = alainPCA(allEEG,nPCs);
- clear allEEG %startIdx EEGdata
- end
- save([mcca_dir 'pca_and_badchans_standards_' conditions{cond} '.mat'],'xAll','topcs','badchans_all'); %
- disp('** Denoise **')
- filename = ([mcca_dir 'MCCA_standards_' conditions{cond} '_subject_']); %fn save name,subject#,mat
- alainMCCAdenoise(xAll,nPCs,nMCCs,filename); % xOut (denoised) = sub x chann x pc x trial
- % Data is expressed in CC space. map back to EEG space for each subject and each trial and save
- for s = 1:nsubs
- load([mcca_dir 'MCCA_standards_' conditions{cond} '_subject_' int2str(s) 'mat.mat']);
- tt = size(xx,2); % xx = time x trial x PC
- eeg_clean = NaN(size(xx,1),nchans,tt); % time, chan, trial
- for trial = 1:tt
- eeg_clean(:,:,trial) = squeeze(xx(:,trial,:))*squeeze(topcs(s,:,1:nPCs))';
- end
- disp('** Save **')
- eeg.trialPosition = 1:size(eeg_clean,3);
- eeg.data = squeeze(num2cell(eeg_clean,[1 2]))';
- eeg.chanlocs = chanlocs;
- save([eeg_dir, 'dataSub', int2str(s), '_', conditions{cond}, '.mat'],'eeg', '-v7.3')
- end
- clear xAll topcs badchans_all
- end
gmdlDataset_preprocess.m, no license · at the source
Overview
- Department of Neuroscience, University of Rochester, Rochester, New York, USA
- Del Monte Institute for Neuroscience, University of Rochester, Rochester, New York, USA
- School of Engineering, Trinity Centre for Bioengineering and Trinity College Institute of Neuroscience, Trinity College Dublin, Dublin, Ireland
- Department of Biomedical Engineering, University of Rochester, Rochester, New York, USA
- Center for Visual Science, University of Rochester, Rochester, New York, USA
Abstract
There is substantial support for the idea that the listening brain makes predictions about upcoming speech and that these predictions are integrated with sensory input to influence perception. For example, the early auditory encoding of words appears to vary based on how those words semantically relate to their preceding context, suggesting that top‐down information might feed back to affect acoustic speech processing. However, the way in which speakers enunciate words can vary based on how well those words fit with their preceding context. This presents a potential confound to the interpretation of top‐down prediction in the listener. In this study, we address this possibility by assessing the influence of probability‐based predictions (word surprisal) on electroencephalographic (EEG) indices of acoustic speech processing while controlling for variations in speaker dynamics. We analyzed EEG from 14 adults who undertook a perceptual pop‐out task in which prior information enhanced the comprehensibility of degraded speech while acoustic information was held constant. Behavioral results confirmed the manipulation's effectiveness and were mirrored in the neural indices of word surprisal processing. Importantly, a positive relationship between word surprisal and EEG tracking of word acoustics emerged for degraded speech when prior information rendered it intelligible, but was absent when it was unintelligible, despite identical acoustic input across conditions. The difference in neural effects between conditions also correlated with the corresponding difference in behavioral pop‐out. These findings support the claim that top‐down word predictability influences the acoustic encoding of natural speech, independent of variations in speaker enunciation.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.
OSF 756xn
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
13 files
- Code/
create_eegstruct.m , MATLAB, 70 lines - Code/
extractForwardAccuracies , MATLAB, 88 lines_cond_pcorr.m - Code/
findBadChansMean.m , MATLAB, 48 lines - Code/
forwardsModel_dataPrep_g , MATLAB, 36 linesmdl.m - Code/
gmdlDataset_fw_lme_4D_ar , MATLAB, 127 lines, 2 matchesray_calc.m - Code/
gmdlDataset_fw_lme_prep. , MATLAB, 163 lines, 1 matchm - Code/
gmdlDataset_fw_mod.m , MATLAB, 192 lines, 2 matches - Code/
gmdlDataset_fw_mod_P_wit , MATLAB, 156 linesh_NPsurp.m - Code/
gmdlDataset_lme_models.R , R, 194 lines - Code/
gmdlDataset_partial_corr , MATLAB, 97 lines.m - Code/
gmdlDataset_partial_corr , MATLAB, 105 lines_NPCsurp.m - Code/
gmdlDataset_preprocess.m , MATLAB, 181 lines, 2 matches - Code/
singleWordProsody.m , MATLAB, 59 lines, 1 match
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 13 scripts, each with its path and the digest of its content;
- 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
The data and scripts associated with this study are available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Wiley
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 5 keywords, 9 MeSH terms, 3 funders, 62 references.
Cite
This paper
Synigal, S. R., Broderick, M. P., & Lalor, E. C. (2026). Quantifying the Influence of Lexical Surprisal on Acoustic Speech Encoding While Controlling for Within-Speaker Variability. The European journal of neuroscience, 64(1), e70569. https://
BibTeX
@article{synigal2026quan
author = {Synigal, Shyanthony R and Broderick, Michael P and Lalor, Edmund C},
title = {{Quantifying the Influence of Lexical Surprisal on Acoustic Speech Encoding While Controlling for Within-Speaker Variability}},
journal = {The European journal of neuroscience},
year = {2026},
month = jul,
volume = {64},
number = {1},
pages = {e70569},
publisher = {Wiley},
issn = {0953-816X},
doi = {10.1111/
url = {https://
pmid = {42390025},
pmcid = {PMC13325525}
}
RIS
TY - JOUR
AU - Synigal, Shyanthony R
AU - Broderick, Michael P
AU - Lalor, Edmund C
TI - Quantifying the Influence of Lexical Surprisal on Acoustic Speech Encoding While Controlling for Within-Speaker Variability
T2 - The European journal of neuroscience
J2 - Eur J Neurosci
PY - 2026
DA - 2026/
VL - 64
IS - 1
SP - e70569
SN - 0953-816X
PB - Wiley
DO - 10.1111/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1111/
"type": "article-journal",
"title": "Quantifying the Influence of Lexical Surprisal on Acoustic Speech Encoding While Controlling for Within-Speaker Variability",
"container-title": "The European journal of neuroscience",
"author": [
{
"family": "Synigal",
"given": "Shyanthony R"
},
{
"family": "Broderick",
"given": "Michael P"
},
{
"family": "Lalor",
"given": "Edmund C"
}
],
"container-title-short":
"volume": "64",
"issue": "1",
"page": "e70569",
"DOI": "10.1111/
"PMID": "42390025",
"PMCID": "PMC13325525",
"ISSN": "0953-816X",
"publisher": "Wiley",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1523/eneuro.0069-26.2026 [code]
- Electrophysiological Indices of Hierarchical Speech Processing Differentially Reflect the Comprehension of Speech in Noise.Journal: eNeuroIn common: psych, EEGLAB, emmeans, 2 other tools, EEG, cognitive, 16 references, author Edmund C. Lalor
- [2] doi:10.1093/braincomms/fcag261 [code]
- Cortical speech envelope tracking reflects lesion-symptom profiles in post-stroke aphasia.Journal: Brain communicationsIn common: emmeans, lmerTest, cognitive, 6 references
- [3] doi:10.1038/s42003-026-10265-1 [code]
- The shape of attention reflects flexible filtering of natural speech modulations.Journal: Communications biologyIn common: EEG, cognitive, 7 references
- [4] doi:10.1016/j.neuroimage.2026.122115 [code]
- Midfrontal theta power relates to response speeding following frustrative nonreward.Journal: NeuroImageIn common: psych, EEGLAB, emmeans, 3 other tools, EEG, cognitive, 1 reference
- [5] doi:10.1371/journal.pbio.3003979 [code]
- Impaired midfrontal‑motor theta phase synchronization characterizes maladaptive motivational behavior in people with obsessive‑compulsive disorder.Journal: PLoS biologyIn common: psych, EEGLAB, emmeans, 3 other tools, EEG, 1 reference
- [6] doi:10.7554/elife.107088 [code]
- Development of auditory and spontaneous movement responses to music over the first postnatal year.Journal: eLifeIn common: EEGLAB, emmeans, Statistics and Machine Learning Toolbox, 1 other tool, EEG, 4 references
- [7] doi:10.1371/journal.pone.0353990 [code]
- Positive mood enhances accessibility of unrelated concepts in the first language but not in the foreign language.Journal: PloS oneIn common: psych, EEGLAB, emmeans, 2 other tools, EEG, cognitive, 1 reference
- [8] doi:10.1523/jneurosci.0154-26.2026 [code]
- Faster but less precise: expectation enhances response speed while reducing sensory fidelity.Journal: The Journal of neuroscience : the official journal of the Society for NeuroscienceIn common: EEGLAB, Statistics and Machine Learning Toolbox, EEG, cognitive, 5 references
- [9] doi:10.1038/s41537-026-00761-y [code]
- The role of fear learning in the development of psychosis: an EEG study utilizing a differential fear conditioning paradigm in people with psychotic vulnerability.Journal: Schizophrenia (Heidelberg, Germany)In common: psych, EEGLAB, emmeans, 2 other tools, EEG, 1 reference
- [10] doi:10.1162/imag.a.105 [code]
- Right posterior theta reflects human parahippocampal phase resetting by salient cues during goal-directed navigationJournal: n/aIn common: psych, EEGLAB, lmerTest, 2 other tools, EEG, cognitive, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 13 scripts, and 8 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:94b83b959bf1ed4f…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
