OSCR

Prediction, syntax and semantic grounding in the brain and large language models.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Material and methods › Data preparation ↔ EEG_Preprocessing_OpenAccess.py, lines 118–196 · score 0.89 · 1–20 Hz, highest variance, ECG channels, Independent Component, MNE, ICs
  2. [2] § Material and methods › Data preparation ↔ MEG_Preprocessing_OpenAccess.py, lines 132–157 · score 0.73 · heartbeat artifact, MNE, flat, noisy, eye, filter
  3. [3] § Material and methods › Data preparation ↔ get_word_classes_epochs_OpenAccess.py, lines 80–163 · score 0.69 · audio signal, stimulus channel, word onset, epochs, segment, interval
  4. [4] § Material and methods › Analysis of LLM LLaMa 3.2 ↔ syntactic_predictability_OpenAccess.py, lines 233–308 · score 0.61 · confusion matrix, Syntactic predictability, Proper Nouns, Accuracy, classifiers, hidden
  5. [5] § Material and methods › Analysis of LLM LLaMa 3.2 ↔ syntactic_predictability_OpenAccess.py, lines 130–141 · score 0.57 · ReLU, syntactic prediction, linear, classifier, hidden
  6. [6] § Material and methods › Data preparation ↔ MEG_Preprocessing_OpenAccess.py, lines 8–43 · score 0.53 · trigger channel, stimulus channel, onset, MEG, EEG

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 194 lines · 7.4 KB · MIT · 2 matches

  1. import mne
  2. from matplotlib import pyplot as plt
  3. from scipy.signal import correlate
  4. import numpy as np
  5. from mne.preprocessing import ICA
  6. def load_meg_data(path, i):
  7. # load eeg data with stimuli channel
  8. proband = mne.io.read_raw_bti(path, rename_channels=False, preload=True)
  9. # find trigger in "RESPONSE"-channel, to align EEG and MEG data
  10. events = mne.find_events(proband, stim_channel=["RESPONSE"], output="onset")
  11. start_ = events[2][0]
  12. stop_ = events[6][0]
  13. if i == 24: # for subject 24 there were some setup problems, which is why we have other triggers here
  14. start_ = events[2][0]
  15. stop_ = events[22][0]
  16. trigger_channel = proband.copy().pick_channels(['RESPONSE']).load_data()
  17. meg_trigger = trigger_channel.to_data_frame()
  18. start = meg_trigger.loc[int(start_)]["time"]
  19. stop = meg_trigger.loc[int(stop_)]["time"]
  20. proband.crop(tmin=start, tmax=stop)
  21. data_meg, _ = proband.get_data(return_times=True)
  22. # reset the start-time and end-time to the trigger before and at the end of the measurement
  23. proband = mne.io.RawArray(data_meg, proband.info, first_samp=0)
  24. del data_meg
  25. # stimuli channel with the audiobook as input in "X2"
  26. stimuli_channel = proband.copy().pick_channels(['X2']).load_data()
  27. # select MEG sensors
  28. channel_names = ["A" + str(i) for i in range(1, 249)]
  29. proband = proband.pick(picks=channel_names)
  30. # interpolate bad sensors
  31. proband.info["bads"] = ["A32", "A60", "A242", "A30", "A33"]
  32. proband = proband.interpolate_bads(reset_bads=True)
  33. return proband, stimuli_channel, trigger_channel
  34. def correlateIC_ekg(sources, num, thresh):
  35. # read EEG file with EOG and ECG channels
  36. eog_ekg = mne.io.read_raw_fif("EEG_Preprocessed/Prob" + str(num) + "_EEGdata_raw.fif")
  37. eog_ekg = eog_ekg.pick(picks=["EOG", "EKG"])
  38. # resample to the same sampling frequency as the MEG data
  39. eog_ekg.resample(200)
  40. eog_ekg_df = eog_ekg.to_data_frame()
  41. channels = sources.info.ch_names
  42. sources = sources.to_data_frame()
  43. # correlate each independent component with the ECG signal
  44. corr = []
  45. for j in channels:
  46. corr_ = correlate(eog_ekg_df["EKG"], sources[j], "valid")
  47. corr.append(corr_)
  48. corr = [np.abs(c[0]) for c in corr]
  49. # find the max correlation value
  50. corrmax = np.max(corr)
  51. corr = [c/corrmax for c in corr]
  52. # append the index of all the channels with a higher correlation value than thresh
  53. idx = []
  54. for i in range(len(corr)):
  55. if corr[i] > thresh:
  56. idx.append(i)
  57. return idx
  58. def correlateIC_eog(sources, num, thresh):
  59. # read EEG file with EOG and ECG channels
  60. eog_ekg = mne.io.read_raw_fif("W:/hno/science/koelblna/Vakuum_MEEG/EEG_Preprocessed/Prob" + str(num) + "_EEGdata_raw.fif")
  61. # the EOG channel of subject 29 had a misfunction --> use of frontal channel "Fp1" instead
  62. if num == 29:
  63. eog_ekg = eog_ekg.pick(picks=["Fp1"])
  64. else:
  65. eog_ekg = eog_ekg.pick(picks=["EOG", "EKG"])
  66. # resample to the same sampling frequency as the MEG data
  67. eog_ekg.resample(200)
  68. eog_ekg_df = eog_ekg.to_data_frame()
  69. channels = sources.info.ch_names
  70. sources = sources.to_data_frame()
  71. # correlate each independent component with the EOG signal
  72. corr = []
  73. for j in channels:
  74. if num == 29:
  75. corr_ = correlate(eog_ekg_df["Fp1"], sources[j], "valid")
  76. else:
  77. corr_ = correlate(eog_ekg_df["EOG"], sources[j], "valid")
  78. corr.append(corr_)
  79. corr = [np.abs(c[0]) for c in corr]
  80. # find the max correlation value
  81. corrmax = np.max(corr)
  82. corr = [c/corrmax for c in corr]
  83. # append the index of all the channels with a higher correlation value than thresh
  84. idx = []
  85. for i in range(len(corr)):
  86. if corr[i] > thresh:
  87. idx.append(i)
  88. return idx
  89. def compute_ica_fixed(raw, num):
  90. # downsample data for computational efficiency
  91. raw.resample(200)
  92. # calculate ICA
  93. ica = ICA(n_components=50, max_iter='auto', random_state=97, method="fastica")
  94. ica.fit(raw)
  95. # get sources (Independent Components) from data
  96. sources = ica.get_sources(raw)
  97. # get all independents components that have eye or heartbeat signal in them
  98. idx_ekg = correlateIC_ekg(sources, num, 0.8)
  99. idx_eog = correlateIC_eog(sources, num, 0.8)
  100. sources = sources.to_data_frame()
  101. # concatenate the first two components (high variance) and the ECG and EOG correlated components
  102. exc = np.concatenate(([0, 1], np.array(idx_ekg), np.array(idx_eog)))
  103. exc = list(exc)
  104. exc = list(dict.fromkeys(exc))
  105. # exclude the components and reconstruct the data
  106. ica.exclude = exc
  107. ica.apply(raw)
  108. return ica, raw
  109. def preprocessing(raw, num):
  110. # find bad meg sensors
  111. auto_noisy_chs, auto_flat_chs = mne.preprocessing.find_bad_channels_maxwell(
  112. raw, duration=120,
  113. verbose=True,
  114. )
  115. # concatenate sensors with flat signal or highly noisy signal
  116. exc = np.concatenate((np.array(auto_noisy_chs), np.array(auto_flat_chs)))
  117. exc = list(exc)
  118. exc = list(dict.fromkeys(exc))
  119. meg_ch = raw.pick(picks="meg")
  120. meg_ch = meg_ch.info.ch_names
  121. channel_exc = []
  122. for i in exc:
  123. if i in meg_ch:
  124. channel_exc.append(i)
  125. # mark them as bad
  126. raw.info["bads"] = channel_exc
  127. # interpolate them
  128. raw = raw.interpolate_bads(reset_bads=True)
  129. # filter data
  130. raw_lh = raw.copy().filter(l_freq=1, h_freq=20)
  131. # use ICA to extract eye and heartbeat artifacts
  132. ica, raw_ica = compute_ica_fixed(raw_lh, num)
  133. return raw_ica
  134. acqtime = ["0", "20.10.23_1059", "26.10.23_1521", "27.10.23_1505", "30.10.23_1452", "02.11.23_1459",
  135. "03.11.23_1541", "07.11.23_1501", "09.11.23_1350", "10.11.23_1502", "13.11.23_1520", "15.11.23_1517",
  136. "21.11.23_1505", "24.11.23_1456", "29.11.23_1514", "30.11.23_1549", "06.12.23_1450", "12.12.23_1459",
  137. "13.12.23_1507", "15.12.23_1459", "18.12.23_1519", "19.12.23_1502", "20.12.23_1558", "15.01.24_1447",
  138. "16.01.24_1454", "22.01.24_1454", "23.01.24_1449", "24.01.24_1444", "26.01.24_1015", "26.01.24_1438",
  139. "29.01.24_1455", "30.01.24_1030", "31.01.24_1151", "31.01.24_1551"]
  140. for i in range(1, 33):
  141. if i == 7 or i == 13 or i == 14 or i == 18: # subjects with setup problems (e.g. earphone fell out)
  142. continue
  143. # load MEG file
  144. if i > 9:
  145. path = "nc_lin_" + str(i) + "/ling_10/" + acqtime[i] + "/1/c,rfhp1.0Hz"
  146. else:
  147. path = "nc_lin_0" + str(i) + "/ling_10/" + acqtime[i] + "/1/c,rfhp1.0Hz"
  148. # load meg channels, stimuli channel and trigger channel
  149. data, stimuli_channel, trigger_channel = load_meg_data(path, i)
  150. # preprocess MEG data and save it
  151. prep_data = preprocessing(data, i)
  152. fname = "Bad_Channels/Prob" + str(i) + "_fixed_ICA_prep_raw.fif"
  153. prep_data.save(fname, overwrite=True)
  154. # resample and save stimuli channel
  155. stimuli_channel = stimuli_channel.resample(sfreq=200)
  156. st_name = "Bad_Channels/Prob" + str(i) + "_stimuli_channel_raw.fif"
  157. stimuli_channel.save(st_name, overwrite=True)
  158. # resample and save trigger channel
  159. trigger_channel = trigger_channel.resample(sfreq=200)
  160. tr_name = "Bad_Channels/Prob" + str(i) + "_trigger_channel_raw.fif"
  161. trigger_channel.save(tr_name, overwrite=True)

MEG_Preprocessing_OpenAccess.py at commit ce17546, under MIT · at the source

Overview

Authors: Nikola Kölbl1,2, Stefan Rampp3,4,5, Martin Kaltenhäuser3, Konstantin Tziridis1, Andreas Maier6, Thomas Kinfe7,8, Ricardo Chavarriaga9, Patrick Krauss1,2,7,8, Achim Schilling1,2,7,8
  1. Neuroscience Lab, University Hospital Erlangen, Erlangen, Germany
  2. CCN Group, Pattern Recognition Lab, FAU Erlangen-Nürnberg, Erlangen, Germany
  3. Department of Neurosurgery, University Hospital Erlangen, Erlangen, Germany
  4. Department of Neurosurgery, University Hospital Halle (Saale), Halle (Saale), Germany
  5. Department of Neuroradiology, University Hospital Erlangen, Erlangen, Germany
  6. Pattern Recognition Lab, FAU Erlangen-Nürnberg, Erlangen, Germany
  7. Mannheim Center for Neuromodulation and Neuroprosthetics (MCNN), University Hospital Mannheim, University Heidelberg, Heidelberg, Germany
  8. BGU Ludwigshafen, Ludwigshafen, Germany
  9. ZHAW Zürich, Zürich, Switzerland
Journal: Scientific reports, volume 16, issue 1, article 8728
Dates: received 13 November 2025; accepted 20 February 2026; published online 10 March 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41598-026-41532-0 · PMID 41807493 · PMCID PMC12979642 · OpenAlex W7134923930
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: EEG (modality), MEG (modality), human (organism), cognitive (subfield)
Methods: Spectral & time-frequency, Statistics, Smoothing, state filtering, decompositions, Preprocessing, Evoked potentials, Machine learning
Keywords: Anticipatory activity, Semantic grounding, Natural language, Continuous speech, MEG, EEG, Embodied cognition, Neuroscience, Psychology
MeSH: Brain*, Comprehension*, Language*, Semantics*, Speech Perception*, Adult, Brain Mapping, Electroencephalography, Evoked Potentials, Female, Humans, Large Language Models, Magnetoencephalography, Male, Young Adult (* major topic)
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: European Research Council (810316); Universitätsklinikum Erlangen
Citations: cited by 1 paper (Europe PMC); 110 references in the paper

Abstract

Language comprehension involves continuous anticipation of upcoming linguistic input, requiring the rapid integration of syntactic structure and semantic information. To capture the spatio-temporal dynamics of such anticipatory processes during naturalistic language comprehension, we combined electroencephalography (EEG) and magnetoencephalography (MEG), leveraging their complementary sensitivities and high temporal resolution. Using this combined EEG-MEG approach, we investigated word-class-specific neural responses during continuous speech perception and related these findings to word class-level predictability and representational structure in a large language model. Twenty-nine healthy participants listened to a German audio book while their neural responses were recorded. Event-related fields and event-related potentials for different word classes showed highly reproducible, characteristic spatio-temporal signatures, including significant pre-onset activity for nouns, suggesting enhanced anticipatory processing of this word class. Source-space analyses revealed activity patterns extending beyond temporal regions into areas compatible with sensorimotor cortices, suggesting a deeper semantic grounding of nouns in e.g. sensory experiences than verbs. By analyzing word class-specific predictability and representational structure in the transformer-based language model Llama, we provide a computational reference frame that complements the neural findings at the level of word classes. These findings highlight the power of simultaneous MEG-EEG recordings in unraveling the predictive, syntactic, and semantic mechanisms that underlie language comprehension.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

nikolakoe/Prediction_Syntax_Semantic_Grounding

License: MIT
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: ce17546208c78df4eda2dec6968063b3e90a491d, 13 July 2026
Languages: Python (5)
Size: 79 files, 5 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (5 files), NumPy (5 files), pandas (4 files), MNE-Python (3 files), SciPy (3 files), PyTorch (2 files), Hugging Face Transformers (2 files), scikit-learn (1 file), seaborn (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
7 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The data is published on the public repository zenodo: (https://zenodo.org/records/15744486). All evaluation code is shared via GitHub: (https://github.com/nikolakoe/Prediction_Syntax_Semantic_Grounding).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 9 keywords, 15 MeSH terms, 2 funders, 68 references.

Cite

This paper

Kölbl, N., Rampp, S., Kaltenhäuser, M., Tziridis, K., Maier, A., Kinfe, T., Chavarriaga, R., Krauss, P., & Schilling, A. (2026). Prediction, syntax and semantic grounding in the brain and large language models. Scientific reports, 16(1), 8728. https://doi.org/10.1038/s41598-026-41532-0

BibTeX

@article{kolbl2026prediction,
author = {Kölbl, Nikola and Rampp, Stefan and Kaltenhäuser, Martin and Tziridis, Konstantin and Maier, Andreas and Kinfe, Thomas and Chavarriaga, Ricardo and Krauss, Patrick and Schilling, Achim},
title = {{Prediction, syntax and semantic grounding in the brain and large language models}},
journal = {Scientific reports},
year = {2026},
month = mar,
volume = {16},
number = {1},
pages = {8728},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-41532-0},
url = {https://doi.org/10.1038/s41598-026-41532-0},
pmid = {41807493},
pmcid = {PMC12979642}
}

RIS

TY - JOUR
AU - Kölbl, Nikola
AU - Rampp, Stefan
AU - Kaltenhäuser, Martin
AU - Tziridis, Konstantin
AU - Maier, Andreas
AU - Kinfe, Thomas
AU - Chavarriaga, Ricardo
AU - Krauss, Patrick
AU - Schilling, Achim
TI - Prediction, syntax and semantic grounding in the brain and large language models
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/03/10
VL - 16
IS - 1
SP - 8728
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-41532-0
UR - https://doi.org/10.1038/s41598-026-41532-0
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-41532-0",
"type": "article-journal",
"title": "Prediction, syntax and semantic grounding in the brain and large language models",
"container-title": "Scientific reports",
"author": [
{
"family": "Kölbl",
"given": "Nikola"
},
{
"family": "Rampp",
"given": "Stefan"
},
{
"family": "Kaltenhäuser",
"given": "Martin"
},
{
"family": "Tziridis",
"given": "Konstantin"
},
{
"family": "Maier",
"given": "Andreas"
},
{
"family": "Kinfe",
"given": "Thomas"
},
{
"family": "Chavarriaga",
"given": "Ricardo"
},
{
"family": "Krauss",
"given": "Patrick"
},
{
"family": "Schilling",
"given": "Achim"
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "8728",
"DOI": "10.1038/s41598-026-41532-0",
"PMID": "41807493",
"PMCID": "PMC12979642",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-41532-0",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10108-z [code]
Time-resolved EEG decoding reveals altered neural dynamics of affective semantic evaluation in depression and suicidality.
Journal: Communications biology
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, EEG, cognitive, 5 references
[2] doi:10.1162/nol.a.247 [code]
Same Sentences, Different Grammars, Different Brain Responses?: An MEG Study on Case and Agreement Encoding in Hindi and Nepali Split-Ergative Structures.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: MNE-Python, seaborn, pandas, 3 other tools, MEG, 5 references
[3] doi:10.7554/elife.106543 [code]
Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.
Journal: eLife
In common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
[4] doi:10.1093/nc/niag029 [code]
A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.
Journal: Neuroscience of consciousness
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, cognitive, 4 references
[5] doi:10.1016/j.neuroimage.2026.122051 [code]
Determining hemispheric language dominance from MEG beta-power modulations: Concordance with fMRI.
Journal: NeuroImage
In common: MNE-Python, scikit-learn, pandas, 3 other tools, MEG, cognitive, 4 references
[6] doi:10.1162/imag.a.1227 [code]
Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
[7] doi:10.1038/s41593-026-02285-1 [code]
Fixation duration on natural scenes is explained by memory encoding not processing demand.
Journal: Nature neuroscience
In common: MNE-Python, PyTorch, seaborn, 5 other tools, MEG, cognitive, 2 references
[8] doi:10.1162/nol.a.244 [code]
A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: Hugging Face Transformers, MNE-Python, PyTorch, 4 other tools, 2 references
[9] doi:10.1038/s41597-026-07350-9 [code]
An open multi-center MEG-EEG dataset for studying conscious visual perception.
Journal: Scientific data
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, EEG, 2 references
[10] doi:10.1038/s41467-026-75240-0 [code]
Dynamic acoustic-to-categorical representations of phonemes and prosody along ventral and dorsal speech streams.
Journal: Nature communications
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, MEG, cognitive, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.