Prediction, syntax and semantic grounding in the brain and large language models.
The 6 matches
- [1] § Material and methods › Data preparation ↔ EEG_Preprocessing_OpenAccess.py, lines 118–196 · score 0.89 · 1–20 Hz, highest variance, ECG channels, Independent Component, MNE, ICs
- [2] § Material and methods › Data preparation ↔ MEG_Preprocessing_OpenAccess.py, lines 132–157 · score 0.73 · heartbeat artifact, MNE, flat, noisy, eye, filter
- [3] § Material and methods › Data preparation ↔ get_word_classes_epochs_OpenAccess.py, lines 80–163 · score 0.69 · audio signal, stimulus channel, word onset, epochs, segment, interval
- [4] § Material and methods › Analysis of LLM LLaMa 3.2 ↔ syntactic_predictability_OpenAccess.py, lines 233–308 · score 0.61 · confusion matrix, Syntactic predictability, Proper Nouns, Accuracy, classifiers, hidden
- [5] § Material and methods › Analysis of LLM LLaMa 3.2 ↔ syntactic_predictability_OpenAccess.py, lines 130–141 · score 0.57 · ReLU, syntactic prediction, linear, classifier, hidden
- [6] § Material and methods › Data preparation ↔ MEG_Preprocessing_OpenAccess.py, lines 8–43 · score 0.53 · trigger channel, stimulus channel, onset, MEG, EEG
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 194 lines · 7.4 KB · MIT · 2 matches
- import mne
- from matplotlib import pyplot as plt
- from scipy.signal import correlate
- import numpy as np
- from mne.preprocessing import ICA
- def load_meg_data(path, i):
- # load eeg data with stimuli channel
- proband = mne.io.read_raw_bti(path, rename_channels=False, preload=True)
- # find trigger in "RESPONSE"-channel, to align EEG and MEG data
- events = mne.find_events(proband, stim_channel=["RESPONSE"], output="onset")
- start_ = events[2][0]
- stop_ = events[6][0]
- if i == 24: # for subject 24 there were some setup problems, which is why we have other triggers here
- start_ = events[2][0]
- stop_ = events[22][0]
- trigger_channel = proband.copy().pick_channels(['RESPONSE']).load_data()
- meg_trigger = trigger_channel.to_data_frame()
- start = meg_trigger.loc[int(start_)]["time"]
- stop = meg_trigger.loc[int(stop_)]["time"]
- proband.crop(tmin=start, tmax=stop)
- data_meg, _ = proband.get_data(return_times=True)
- # reset the start-time and end-time to the trigger before and at the end of the measurement
- proband = mne.io.RawArray(data_meg, proband.info, first_samp=0)
- del data_meg
- # stimuli channel with the audiobook as input in "X2"
- stimuli_channel = proband.copy().pick_channels(['X2']).load_data()
- # select MEG sensors
- channel_names = ["A" + str(i) for i in range(1, 249)]
- proband = proband.pick(picks=channel_names)
- # interpolate bad sensors
- proband.info["bads"] = ["A32", "A60", "A242", "A30", "A33"]
- proband = proband.interpolate_bads(reset_bads=True)
- return proband, stimuli_channel, trigger_channel
- def correlateIC_ekg(sources, num, thresh):
- # read EEG file with EOG and ECG channels
- eog_ekg = mne.io.read_raw_fif("EEG_Preprocessed/Prob" + str(num) + "_EEGdata_raw.fif")
- eog_ekg = eog_ekg.pick(picks=["EOG", "EKG"])
- # resample to the same sampling frequency as the MEG data
- eog_ekg.resample(200)
- eog_ekg_df = eog_ekg.to_data_frame()
- channels = sources.info.ch_names
- sources = sources.to_data_frame()
- # correlate each independent component with the ECG signal
- corr = []
- for j in channels:
- corr_ = correlate(eog_ekg_df["EKG"], sources[j], "valid")
- corr.append(corr_)
- corr = [np.abs(c[0]) for c in corr]
- # find the max correlation value
- corrmax = np.max(corr)
- corr = [c/corrmax for c in corr]
- # append the index of all the channels with a higher correlation value than thresh
- idx = []
- for i in range(len(corr)):
- if corr[i] > thresh:
- idx.append(i)
- return idx
- def correlateIC_eog(sources, num, thresh):
- # read EEG file with EOG and ECG channels
- eog_ekg = mne.io.read_raw_fif("W:/hno/science/koelblna/Vakuum_MEEG/EEG_Preprocessed/Prob" + str(num) + "_EEGdata_raw.fif")
- # the EOG channel of subject 29 had a misfunction --> use of frontal channel "Fp1" instead
- if num == 29:
- eog_ekg = eog_ekg.pick(picks=["Fp1"])
- else:
- eog_ekg = eog_ekg.pick(picks=["EOG", "EKG"])
- # resample to the same sampling frequency as the MEG data
- eog_ekg.resample(200)
- eog_ekg_df = eog_ekg.to_data_frame()
- channels = sources.info.ch_names
- sources = sources.to_data_frame()
- # correlate each independent component with the EOG signal
- corr = []
- for j in channels:
- if num == 29:
- corr_ = correlate(eog_ekg_df["Fp1"], sources[j], "valid")
- else:
- corr_ = correlate(eog_ekg_df["EOG"], sources[j], "valid")
- corr.append(corr_)
- corr = [np.abs(c[0]) for c in corr]
- # find the max correlation value
- corrmax = np.max(corr)
- corr = [c/corrmax for c in corr]
- # append the index of all the channels with a higher correlation value than thresh
- idx = []
- for i in range(len(corr)):
- if corr[i] > thresh:
- idx.append(i)
- return idx
- def compute_ica_fixed(raw, num):
- # downsample data for computational efficiency
- raw.resample(200)
- # calculate ICA
- ica = ICA(n_components=50, max_iter='auto', random_state=97, method="fastica")
- ica.fit(raw)
- # get sources (Independent Components) from data
- sources = ica.get_sources(raw)
- # get all independents components that have eye or heartbeat signal in them
- idx_ekg = correlateIC_ekg(sources, num, 0.8)
- idx_eog = correlateIC_eog(sources, num, 0.8)
- sources = sources.to_data_frame()
- # concatenate the first two components (high variance) and the ECG and EOG correlated components
- exc = np.concatenate(([0, 1], np.array(idx_ekg), np.array(idx_eog)))
- exc = list(exc)
- exc = list(dict.fromkeys(exc))
- # exclude the components and reconstruct the data
- ica.exclude = exc
- ica.apply(raw)
- return ica, raw
- def preprocessing(raw, num):
- # find bad meg sensors
- auto_noisy_chs, auto_flat_chs = mne.preprocessing.find_bad_channels_maxwell(
- raw, duration=120,
- verbose=True,
- )
- # concatenate sensors with flat signal or highly noisy signal
- exc = np.concatenate((np.array(auto_noisy_chs), np.array(auto_flat_chs)))
- exc = list(exc)
- exc = list(dict.fromkeys(exc))
- meg_ch = raw.pick(picks="meg")
- meg_ch = meg_ch.info.ch_names
- channel_exc = []
- for i in exc:
- if i in meg_ch:
- channel_exc.append(i)
- # mark them as bad
- raw.info["bads"] = channel_exc
- # interpolate them
- raw = raw.interpolate_bads(reset_bads=True)
- # filter data
- raw_lh = raw.copy().filter(l_freq=1, h_freq=20)
- # use ICA to extract eye and heartbeat artifacts
- ica, raw_ica = compute_ica_fixed(raw_lh, num)
- return raw_ica
- acqtime = ["0", "20.10.23_1059", "26.10.23_1521", "27.10.23_1505", "30.10.23_1452", "02.11.23_1459",
- "03.11.23_1541", "07.11.23_1501", "09.11.23_1350", "10.11.23_1502", "13.11.23_1520", "15.11.23_1517",
- "21.11.23_1505", "24.11.23_1456", "29.11.23_1514", "30.11.23_1549", "06.12.23_1450", "12.12.23_1459",
- "13.12.23_1507", "15.12.23_1459", "18.12.23_1519", "19.12.23_1502", "20.12.23_1558", "15.01.24_1447",
- "16.01.24_1454", "22.01.24_1454", "23.01.24_1449", "24.01.24_1444", "26.01.24_1015", "26.01.24_1438",
- "29.01.24_1455", "30.01.24_1030", "31.01.24_1151", "31.01.24_1551"]
- for i in range(1, 33):
- if i == 7 or i == 13 or i == 14 or i == 18: # subjects with setup problems (e.g. earphone fell out)
- continue
- # load MEG file
- if i > 9:
- path = "nc_lin_" + str(i) + "/ling_10/" + acqtime[i] + "/1/c,rfhp1.0Hz"
- else:
- path = "nc_lin_0" + str(i) + "/ling_10/" + acqtime[i] + "/1/c,rfhp1.0Hz"
- # load meg channels, stimuli channel and trigger channel
- data, stimuli_channel, trigger_channel = load_meg_data(path, i)
- # preprocess MEG data and save it
- prep_data = preprocessing(data, i)
- fname = "Bad_Channels/Prob" + str(i) + "_fixed_ICA_prep_raw.fif"
- prep_data.save(fname, overwrite=True)
- # resample and save stimuli channel
- stimuli_channel = stimuli_channel.resample(sfreq=200)
- st_name = "Bad_Channels/Prob" + str(i) + "_stimuli_channel_raw.fif"
- stimuli_channel.save(st_name, overwrite=True)
- # resample and save trigger channel
- trigger_channel = trigger_channel.resample(sfreq=200)
- tr_name = "Bad_Channels/Prob" + str(i) + "_trigger_channel_raw.fif"
- trigger_channel.save(tr_name, overwrite=True)
MEG_Preprocessing_OpenAccess.py at commit ce17546, under MIT · at the source
Overview
- Neuroscience Lab, University Hospital Erlangen, Erlangen, Germany
- CCN Group, Pattern Recognition Lab, FAU Erlangen-Nürnberg, Erlangen, Germany
- Department of Neurosurgery, University Hospital Erlangen, Erlangen, Germany
- Department of Neurosurgery, University Hospital Halle (Saale), Halle (Saale), Germany
- Department of Neuroradiology, University Hospital Erlangen, Erlangen, Germany
- Pattern Recognition Lab, FAU Erlangen-Nürnberg, Erlangen, Germany
- Mannheim Center for Neuromodulation and Neuroprosthetics (MCNN), University Hospital Mannheim, University Heidelberg, Heidelberg, Germany
- BGU Ludwigshafen, Ludwigshafen, Germany
- ZHAW Zürich, Zürich, Switzerland
Abstract
Language comprehension involves continuous anticipation of upcoming linguistic input, requiring the rapid integration of syntactic structure and semantic information. To capture the spatio-temporal dynamics of such anticipatory processes during naturalistic language comprehension, we combined electroencephalography (EEG) and magnetoencephalography (MEG), leveraging their complementary sensitivities and high temporal resolution. Using this combined EEG-MEG approach, we investigated word-class-specific neural responses during continuous speech perception and related these findings to word class-level predictability and representational structure in a large language model. Twenty-nine healthy participants listened to a German audio book while their neural responses were recorded. Event-related fields and event-related potentials for different word classes showed highly reproducible, characteristic spatio-temporal signatures, including significant pre-onset activity for nouns, suggesting enhanced anticipatory processing of this word class. Source-space analyses revealed activity patterns extending beyond temporal regions into areas compatible with sensorimotor cortices, suggesting a deeper semantic grounding of nouns in e.g. sensory experiences than verbs. By analyzing word class-specific predictability and representational structure in the transformer-based language model Llama, we provide a computational reference frame that complements the neural findings at the level of word classes. These findings highlight the power of simultaneous MEG-EEG recordings in unraveling the predictive, syntactic, and semantic mechanisms that underlie language comprehension.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.
nikolakoe/Prediction_Syntax_Semantic_Grounding
ce17546208c78df4eda2dec6968063b3e90a491d, 13 July 2026Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
7 files
- EEG_Preprocessing_OpenAc
cess.py , Python, 196 lines, 1 match - MEG_Preprocessing_OpenAc
cess.py , Python, 194 lines, 2 matches - get_word_classes_epochs_
OpenAccess.py , Python, 163 lines, 1 match - semantic_predictability_
OpenAccess.py , Python, 124 lines - syntactic_predictability
_OpenAccess.py , Python, 436 lines, 2 matches - LICENSE, License, 21 lines
- README.md, Text, 16 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 5 scripts, each with its path and the digest of its content;
- 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:15744486, at Zenodo; found in “Data availability”
Data availability
The data is published on the public repository zenodo: (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 9 keywords, 15 MeSH terms, 2 funders, 68 references.
Cite
This paper
Kölbl, N., Rampp, S., Kaltenhäuser, M., Tziridis, K., Maier, A., Kinfe, T., Chavarriaga, R., Krauss, P., & Schilling, A. (2026). Prediction, syntax and semantic grounding in the brain and large language models. Scientific reports, 16(1), 8728. https://
BibTeX
@article{kolbl2026predic
author = {Kölbl, Nikola and Rampp, Stefan and Kaltenhäuser, Martin and Tziridis, Konstantin and Maier, Andreas and Kinfe, Thomas and Chavarriaga, Ricardo and Krauss, Patrick and Schilling, Achim},
title = {{Prediction, syntax and semantic grounding in the brain and large language models}},
journal = {Scientific reports},
year = {2026},
month = mar,
volume = {16},
number = {1},
pages = {8728},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/
url = {https://
pmid = {41807493},
pmcid = {PMC12979642}
}
RIS
TY - JOUR
AU - Kölbl, Nikola
AU - Rampp, Stefan
AU - Kaltenhäuser, Martin
AU - Tziridis, Konstantin
AU - Maier, Andreas
AU - Kinfe, Thomas
AU - Chavarriaga, Ricardo
AU - Krauss, Patrick
AU - Schilling, Achim
TI - Prediction, syntax and semantic grounding in the brain and large language models
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/
VL - 16
IS - 1
SP - 8728
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Prediction, syntax and semantic grounding in the brain and large language models",
"container-title": "Scientific reports",
"author": [
{
"family": "Kölbl",
"given": "Nikola"
},
{
"family": "Rampp",
"given": "Stefan"
},
{
"family": "Kaltenhäuser",
"given": "Martin"
},
{
"family": "Tziridis",
"given": "Konstantin"
},
{
"family": "Maier",
"given": "Andreas"
},
{
"family": "Kinfe",
"given": "Thomas"
},
{
"family": "Chavarriaga",
"given": "Ricardo"
},
{
"family": "Krauss",
"given": "Patrick"
},
{
"family": "Schilling",
"given": "Achim"
}
],
"container-title-short":
"volume": "16",
"issue": "1",
"page": "8728",
"DOI": "10.1038/
"PMID": "41807493",
"PMCID": "PMC12979642",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10108-z [code]
- Time-resolved EEG decoding reveals altered neural dynamics of affective semantic evaluation in depression and suicidality.Journal: Communications biologyIn common: MNE-Python, seaborn, scikit-learn, 4 other tools, EEG, cognitive, 5 references
- [2] doi:10.1162/nol.a.247 [code]
- Same Sentences, Different Grammars, Different Brain Responses?
: An MEG Study on Case and Agreement Encoding in Hindi and Nepali Split-Ergative Structures. Journal: Neurobiology of language (Cambridge, Mass.)In common: MNE-Python, seaborn, pandas, 3 other tools, MEG, 5 references - [3] doi:10.7554/elife.106543 [code]
- Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.Journal: eLifeIn common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
- [4] doi:10.1093/nc/niag029 [code]
- A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.Journal: Neuroscience of consciousnessIn common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, cognitive, 4 references
- [5] doi:10.1016/j.neuroimage.2026.122051 [code]
- Determining hemispheric language dominance from MEG beta-power modulations: Concordance with fMRI.Journal: NeuroImageIn common: MNE-Python, scikit-learn, pandas, 3 other tools, MEG, cognitive, 4 references
- [6] doi:10.1162/imag.a.1227 [code]
- Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.Journal: Imaging neuroscience (Cambridge, Mass.)In common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
- [7] doi:10.1038/s41593-026-02285-1 [code]
- Fixation duration on natural scenes is explained by memory encoding not processing demand.Journal: Nature neuroscienceIn common: MNE-Python, PyTorch, seaborn, 5 other tools, MEG, cognitive, 2 references
- [8] doi:10.1162/nol.a.244 [code]
- A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.Journal: Neurobiology of language (Cambridge, Mass.)In common: Hugging Face Transformers, MNE-Python, PyTorch, 4 other tools, 2 references
- [9] doi:10.1038/s41597-026-07350-9 [code]
- An open multi-center MEG-EEG dataset for studying conscious visual perception.Journal: Scientific dataIn common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, EEG, 2 references
- [10] doi:10.1038/s41467-026-75240-0 [code]
- Dynamic acoustic-to-categorical representations of phonemes and prosody along ventral and dorsal speech streams.Journal: Nature communicationsIn common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, cognitive, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 5 scripts, and 6 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:26673380e87da393…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
