OSCR

Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § Methods › Data ↔ lingpred_new/io.py, lines 79–223 · score 0.66 · 0.1–40 Hz, baseline correction, filtered, word onset, Sherlock, window
  2. [2] § Methods › Control system one: self-predictability analysis ↔ lingpred_new/plotting.py, lines 744–892 · score 0.60 · incorrect predictions, unpredicted words, GloVe, split, vectors, model
  3. [3] § Methods › MEG encoding modelling › Source selection ↔ lingpred_new/io.py, lines 79–223 · score 0.56 · post onset encoding, post word onset, window, activation, 500 ms, pre

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 699 lines · 25 KB · MIT · 2 matches

  1. import glob
  2. import shutil
  3. import os
  4. import sys
  5. from pathlib import Path
  6. import numpy as np
  7. import pickle
  8. from typing import Sequence,Union,Optional
  9. import imp
  10. import pandas as pd
  11. import tgt
  12. import mne
  13. import mne_bids
  14. import h5py
  15. from itertools import compress
  16. from scipy.io import loadmat
  17. PROJ_ROOT = '/project/3018059.03'
  18. # -----------------------------------------------------------#
  19. # Routine to load text #
  20. # -----------------------------------------------------------#
  21. def get_text_per_session(dataset: str, session: int, subject: int):
  22. """
  23. Loads text for a given dataset, session, and subject.
  24. Parameters
  25. ----------
  26. dataset : str
  27. Name of the dataset. Options: "Gwilliams" or "Armani"
  28. session : int
  29. Session for which the text is supposed to be loaded.
  30. - `range(11)` for Armani
  31. - `range(4)` for Gwilliams
  32. subject :
  33. Subject for whom the text is supposed to be loaded.
  34. Returns
  35. -------
  36. str
  37. The text for all runs as a single string.
  38. """
  39. if dataset =='Armani':
  40. runs = _runs_in_session(sess_i=session,sub_i=subject)
  41. text_all_runs = ''
  42. for run in runs:
  43. this_run_text = _load_full_text(sess_i = session, run_i = run ,without_breaks=True)
  44. text_all_runs = text_all_runs + ' ' + this_run_text
  45. return text_all_runs
  46. if dataset == 'Gwilliams':
  47. if str(session) =='0':
  48. textname = 'lw1.txt'
  49. if str(session) == '1':
  50. textname = 'cable_spool_fort.txt'
  51. if str(session) =='2':
  52. textname = 'easy_money.txt'
  53. if str(session) == '3':
  54. textname = 'the_black_willow.txt'
  55. fname = '/project/3018059.03/Lingpred/data/Gwilliams/stimuli/text/' + textname
  56. with open(fname, 'r') as file:
  57. text = file.read()
  58. return text
  59. # -----------------------------------------------------------#
  60. # NEW ROUTINES TO LOAD GWILLIAMS OR ARMANI MEG DATA #
  61. # -----------------------------------------------------------#
  62. def get_neural_data(dataset: str, sessions: list, subject: int, task = 'compr', datatype='source', channels=None, window_size=100, band=(0.1, 40), baseline=None):
  63. """
  64. Computes the average activation at each channel and lag relative to word onset.
  65. Parameters
  66. ----------
  67. dataset : str
  68. Either "Gwilliams", "Sherlock", or "Armani".
  69. subject : str or int
  70. Subject identifier. For Armani dataset: e.g., "001". For Gwilliams dataset: e.g., "01".
  71. task : str or int
  72. Task identifier. For Armani dataset: "compr". For Gwilliams dataset: 0, 1, 2, or 3, corresponding to:
  73. - 0 = lw1
  74. - 1 = cable spool fort
  75. - 2 = easy money
  76. - 3 = black willow
  77. channels : str, optional
  78. Can be 'None', 'language', 'signal', 'pre-onset', or 'post-onset'.
  79. - `None`: Returns all channels.
  80. - `'language'`: Returns only language-related channels.
  81. - `'signal'` (subject-specific): Returns only channels with relevant encoding (max value > 25% of max overall value).
  82. - `'pre-onset'` (subject-specific): Returns only channels with predominantly pre-word-onset encoding.
  83. - `'post-onset'` (subject-specific): Returns only channels with predominantly post-word-onset encoding.
  84. window_size : int, optional
  85. Window size for averaging the neural data with lag = 25 ms. If `None`, no averaging is performed.
  86. band : tuple of float
  87. Bandpass filter range to be used. Can be (0.5, 8) or (1, 40).
  88. baseline : bool or None
  89. Whether to perform baseline correction. If `None`, no correction is performed.
  90. Returns
  91. -------
  92. numpy.ndarray
  93. An array containing the average activation at each channel and lag relative to word onset.
  94. The returned array has shape (nr_channels, nr_epochs, nr_lags).
  95. pandas.DataFrame
  96. A dataframe containing the word of each epoch, the onset time, and additional information.
  97. list
  98. A list of dropped epochs.
  99. """
  100. if dataset == "Gwilliams":
  101. task = task
  102. sessions = sessions
  103. #get_sessions(dataset, subject) # These methods are still missing
  104. if channels == 'language':
  105. sources = get_language_channels(dataset)
  106. if channels == 'signal':
  107. sources = get_signal_channels(dataset, subject)
  108. if channels == 'pre-onset':
  109. sources = get_signal_channels(dataset, subject, pre=True)
  110. if channels == 'post-onset':
  111. sources = get_signal_channels(dataset, subject, post=True)
  112. else: sources = None # this includes all channels
  113. for session in sessions:
  114. bad_epochs_all_runs = [] # list with indices of dropped epochs
  115. length_all_epochs = 0 # constant to be added to the indices for each run
  116. runs = get_runs(dataset, subject, session) # This now only returns [1] for Gwilliams & Armani
  117. for run in runs:
  118. # load raw data and get a dataframe for annotations incl. word ID and onset/offset
  119. raw_data = load_raw_data(dataset, subject, session, run, task, datatype=datatype, band=band)
  120. annotations_df = get_words_onsets_offsets(raw_data, dataset, subject, session, run)
  121. onset_words = annotations_df.onset
  122. # create events
  123. events = np.c_[onset_words * raw_data.info['sfreq'],
  124. np.zeros((len(onset_words), 1)),
  125. np.ones((len(onset_words), 1))].astype(int)
  126. # now down-sample the data to 200Hz for computational reasons:
  127. # by passing the events array we make sure our timing is not jittered:
  128. raw_data, events = raw_data.resample(sfreq=200, events=events)
  129. # create epochs
  130. epochs = mne.Epochs(raw_data,
  131. events,
  132. tmin=-2,
  133. tmax=2,
  134. baseline=baseline, # (-2, 0) would be the default baseline
  135. picks=sources,
  136. metadata = annotations_df,
  137. event_repeated='drop',
  138. preload=True)
  139. # get indices of bad epochs:
  140. bad_epochs_this_run = [index + length_all_epochs for index, dl in enumerate(epochs.drop_log) if len(dl)]
  141. print(bad_epochs_this_run)
  142. length_all_epochs = length_all_epochs + len(epochs) + len(bad_epochs_this_run)
  143. bad_epochs_all_runs = bad_epochs_all_runs + bad_epochs_this_run
  144. # get average per lag
  145. if window_size:
  146. data_per_lag = get_mean_data_lag(epochs, window_size) # array of shape (nr_epochs, nr_channels, nr_lags)
  147. else:
  148. data_per_lag = epochs.get_data()
  149. # stack on top for all runs in one session & append annotations dataframe
  150. if run == runs[0] : # first run in session
  151. data_per_lag_all_runs = data_per_lag
  152. comb_annotation_df = epochs.metadata
  153. else:
  154. data_per_lag_all_runs = np.vstack((data_per_lag_all_runs, # array of shape:
  155. data_per_lag)) # (nr_all_epochs, nr_channels, nr_lags)
  156. comb_annotation_df = comb_annotation_df.append(epochs.metadata)
  157. print(bad_epochs_all_runs)
  158. # stack on top for all sessions & append annotations dataframe
  159. if session == 1 or len(sessions)==1:
  160. data_per_lag_all_sess = data_per_lag_all_runs
  161. annotation_df_all_sess = comb_annotation_df
  162. else:
  163. data_per_lag_all_sess = np.vstack((data_per_lag_all_sess, # array of shape:
  164. data_per_lag_all_runs)) # (nr_all_epochs, nr_channels, nr_lags)
  165. annotation_df_all_sess = annotation_df_all_sess.append(comb_annotation_df)
  166. # for subject 3, session 8 words 3841-4170 are scrambled: and need to be dropped from the df and the neural data:
  167. if session==8 and subject==3:
  168. annotation_df_all_sess.drop(index=annotation_df_all_sess[3841:4170].index, inplace=True)
  169. np.delete(data_per_lag_all_sess, np.arange(3841,4170))
  170. # change shape of the array such that one can easily loop over the channels: (nr_channels, nr_all_epochs, nr_lags)
  171. data_per_lag_all_sess = np.swapaxes(data_per_lag_all_sess,0,1)
  172. # return array with neural data and annotation data frame:
  173. return data_per_lag_all_sess, annotation_df_all_sess, bad_epochs_all_runs
  174. # ------------------------------------------------------------#
  175. # Associated New Auxiliary Functions: #
  176. # ------------------------------------------------------------#
  177. def drop_nans(y, X):
  178. '''
  179. Removes words containing NaN values from the neural data and GPT layer activations.
  180. Parameters
  181. ----------
  182. y : numpy.ndarray
  183. Neural data array of shape (channels, words, lags), containing the averaged MEG data.
  184. X : numpy.ndarray
  185. GPT layer activations, an array of shape (words, dimensions + 1).
  186. nan_ids : list of int
  187. List containing the indices of words that contain NaN values.
  188. Returns
  189. -------
  190. numpy.ndarray
  191. The `y` array with words containing NaNs removed.
  192. numpy.ndarray
  193. The `X` array with words containing NaNs removed.
  194. '''
  195. # get NaN values in neural data:
  196. nan_ids = get_nans(y)
  197. y = np.swapaxes(y,0,1) # swap axes such that rows == words
  198. # drop rows containing NaNs in y and X
  199. y = np.delete(y, nan_ids, axis=0)
  200. X = np.delete(X, nan_ids, axis=0)
  201. y = np.swapaxes(y,0,1) # swap axes such that rows == channels
  202. # print shapes
  203. print(X.shape, y.shape)
  204. return y, X
  205. def get_nans(neural_data):
  206. """
  207. Identifies words containing NaN values in the neural data.
  208. Parameters
  209. ----------
  210. neural_data : numpy.ndarray
  211. Array of shape (channels, words, lags), containing the averaged MEG data.
  212. Returns
  213. -------
  214. list of int
  215. List containing the indices of words that contain NaN values (to be dropped).
  216. """
  217. # initialise counter and index list
  218. count = 0
  219. nan_ids = []
  220. # loop over words (rows) to find rows containing NaNs
  221. for i, word in enumerate(neural_data[0]): # just do this for one channel
  222. if np.any(np.any(np.isnan(word))):
  223. count+=1
  224. nan_ids.append(i)
  225. print('There are {} words with NaNs which will be dropped.'.format(count))
  226. return nan_ids
  227. def get_language_channels(dataset: str, get_areas=False):
  228. '''
  229. Retrieves the list of channel names related to the language system for a given dataset.
  230. Parameters
  231. ----------
  232. dataset : str
  233. The dataset name. Can be either 'sherlock', 'Armani', or 'Gwilliams'.
  234. Returns
  235. -------
  236. list of str
  237. A list of channel names related to the language system.
  238. '''
  239. if dataset=='sherlock' or dataset=='Armani':
  240. source_info = ASH_load_source_info()
  241. ch_names = source_info['lbls_language']
  242. if get_areas:
  243. ch_names = list(compress(source_info['areas'], source_info['language_mask']))
  244. else: raise VallueError('get_language_channels is only implemented for the Armani dataset')
  245. return ch_names
  246. def get_lags(window_size, lag=0.025, start=-2, end=2):
  247. '''
  248. Returns an list of shape (nr_lags, 3) each row containing (lag_index, start_window, end_window) in sec
  249. '''
  250. starts = np.round(np.arange(start, end-window_size+lag, lag), 3)
  251. ends = np.round([s + window_size for s in starts], 3)
  252. arr = [[nr, start, end] for nr, (start, end) in enumerate(zip(starts, ends))]
  253. return arr
  254. def get_mean_data_lag(epochs, window_size):
  255. '''
  256. Param: MNE epochs object
  257. Returns average data for each window: array of shape (nr_epochs, nr_channels, nr_lags)
  258. '''
  259. # transform window_size from ms to s (round to 2 decimals to ensure precision, i.e. 0.05, 0.10, 0.15, 0.2)
  260. window_size = round(window_size/1000, 2)
  261. lags = get_lags(window_size=window_size)
  262. nr_lags = len(lags)
  263. for i in range(nr_lags):
  264. nr, start, end = lags[i]
  265. data = epochs.get_data(tmin=start, tmax=end) # ndarray of shape (nr_ep, nr_ch, nr_data)
  266. mean_data = np.mean(data, axis=2) # ndarray of shape (nr_epochs, nr_channels)
  267. del data
  268. if i==0:
  269. com_data = mean_data
  270. else:
  271. com_data = np.dstack([com_data, mean_data])
  272. return com_data # array of shape (nr_epochs, nr_channels, nr_lags)
  273. def load_raw_data(dataset: str, subject: int, session: int, run: int, task: str, datatype='raw', band=(1, 40)):
  274. """
  275. Loads raw data object using the MNE.
  276. Parameters
  277. ----------
  278. dataset : str
  279. The dataset name. Can be "Gwilliams", "sherlock", or "Armani".
  280. subject : str or int
  281. The subject identifier for which the data should be loaded.
  282. session : str or int
  283. The session for which the data should be loaded.
  284. task : str
  285. The task for which the data should be loaded. '0' for Gwilliams and 'compr' for Armani.
  286. band : tuple of float
  287. A tuple specifying the frequency band of the filtered Sherlock data.
  288. datatype : str
  289. The type of data to load:
  290. - `'filtered'` if the dataset is Gwilliams (filtered sensor data).
  291. - `'source'` if the dataset is Armani (filtered source data).
  292. Returns
  293. -------
  294. mne.io.Raw
  295. An MNE raw object with the filtered data for Gwilliams and source localized data for Armani.
  296. """
  297. # set root to path:
  298. if dataset=="Gwilliams":
  299. root='/project/3018059.03/Lingpred/data/Gwilliams/derived/'
  300. sess = str(session)
  301. if subject < 10:
  302. sub = '0' + str(subject)
  303. else:
  304. sub = str(subject)
  305. elif dataset=="Armani":
  306. root='/project/3018059.03/Lingpred/data/Armani/'
  307. sub = '00' + str(subject)
  308. if session < 10:
  309. sess = '00' + str(session)
  310. else:
  311. sess = '0' + str(session)
  312. # set path to raw data:
  313. if datatype == 'raw':
  314. datatype = 'meg'
  315. bids_path = mne_bids.BIDSPath(subject=sub,
  316. session=sess,
  317. task=task,
  318. datatype=datatype,
  319. root=root)
  320. # load data:
  321. if dataset=="Armani":
  322. # if we want the raw source localised data we need to load it from the source foulder:
  323. if datatype=='source':
  324. # file name, e.g.: 1-1_lcmv-data_0.1-40raw.fif
  325. fname = str(subject) + '-' + str(session) + '_lcmv-data_' + str(band[0]) + '-' + str(band[1]) + 'raw.fif'
  326. # handle naming of session 10
  327. if session < 10:
  328. session = '0'+ str(session)
  329. # full file path
  330. fif_path = root + 'sub-00' + str(subject) + '/ses-0' + str(session) + '/source/' + fname
  331. #read data
  332. raw = mne.io.read_raw_fif(fif_path)
  333. # else if we want raw data:
  334. else:
  335. raw = mne.io.read_raw_ctf(bids_path)
  336. if dataset=="Gwilliams":
  337. # if we want the filtered data:
  338. if datatype=='filtered':
  339. band=(0.1, 40)
  340. # file name, e.g.: 1-1_lcmv-data_0.1-40raw.fif
  341. fname = str(subject)+'-'+sess+'_filtered_data_'+str(band[0])+'-'+str(band[1])+'_task-'+task+'_'+'raw.fif'
  342. # full file path
  343. folder = '/project/3018059.03/data/Gwilliams/'
  344. fif_path = folder + 'sub-' + sub + '/ses-' + sess + '/filtered/' + fname
  345. # return False if the file doesn't exist
  346. if not Path(fif_path).is_file():
  347. return False
  348. #read data
  349. raw = mne.io.read_raw_fif(fif_path)
  350. # if we want the raw data:
  351. else:
  352. raw = mne_bids.read_raw_bids(bids_path)
  353. if dataset not in ['Gwilliams', 'Armani']:
  354. raise ValueError('Dataset variable must be either "Gwilliams", or "Armani"')
  355. # load raw data
  356. raw.load_data()
  357. return raw
  358. def get_words_onsets_offsets(raw_data, dataset:str, subject:int, session:int, run:int):
  359. """
  360. Loads word onset offsets for a given dataset, subject, session, and run.
  361. Parameters
  362. ----------
  363. raw_data : mne.io.Raw
  364. The raw data object.
  365. dataset : str
  366. The dataset name. Can be "Gwilliams", "Armani", or "Sherlock".
  367. subject : str or int
  368. The subject identifier for which the offsets should be loaded.
  369. session : str or int
  370. The session for which the offsets should be loaded.
  371. run : str or int
  372. The run for which the offsets should be loaded.
  373. Returns
  374. -------
  375. pandas.DataFrame
  376. A DataFrame containing at least two columns: 'word' and 'onset'.
  377. """
  378. if dataset == 'Gwilliams':
  379. df = raw_data.annotations.to_data_frame()
  380. df = pd.DataFrame(df.description.apply(eval).to_list())
  381. onsets = raw_data.annotations.onset
  382. # make a column for the onsets:
  383. df['onset'] = onsets
  384. # keep only word onset data
  385. df_words = df[df['kind']=='word']
  386. # keep only words which are part of the story (no pseudowords or wordlists)
  387. df_words = df_words[df_words['condition']=='sentence']
  388. # deal with shift:
  389. word_onsets = df.loc[df_words.index + 1].onset.values
  390. df_words['onset'] = word_onsets
  391. elif dataset == 'Armani':
  392. # handle naming of session 10:
  393. if session < 10:
  394. sess = '00' + str(session)
  395. else:
  396. sess = '0' + str(session)
  397. # get path to events file:
  398. dir_path = '/project/3018059.03/Lingpred/data/Armani/'
  399. filepath = 'sub-00' + str(subject) +'/' + 'ses-' + sess +'/'+ 'meg/'
  400. filename = 'sub-00' + str(subject) + '_ses-' + sess + '_task-compr_events.tsv'
  401. # read pandas DataFrame
  402. annotations = pd.read_csv(dir_path+filepath+filename, sep='\t')
  403. # type of event is separated by runs: there are at max 7 runs in each session:
  404. onset_names_list = ['word_onset_01', 'word_onset_02', 'word_onset_03', 'word_onset_04',
  405. 'word_onset_05', 'word_onset_06', 'word_onset_07', 'word_onset_08']
  406. # get only words, rename columns (identically to Gwilliams) and clean data frame
  407. df_words = annotations[annotations.type.isin(onset_names_list)]
  408. df_words = df_words[df_words.value != 'sp']
  409. df_words.rename(columns={'value':'word'}, inplace=True)
  410. df_words = clean_events(df_words, subject, session, dataset)
  411. else: raise ValueError('Dataset variable must be either "Gwilliams", or "Armani"')
  412. return df_words
  413. def get_phonemes_onsets_offsets(dataset:str, subject:int, session:int, run:int,
  414. only_word_inital_phonemes=True, correct_wav_onset=True):
  415. """
  416. Loads phoneme onset offsets for a given dataset, subject, session, and run.
  417. Parameters
  418. ----------
  419. raw_data : mne.io.Raw
  420. The raw data object.
  421. dataset : str
  422. The dataset name. Can be "Gwilliams", "Armani", or "Sherlock".
  423. subject : str or int
  424. The subject identifier for which the offsets should be loaded.
  425. session : str or int
  426. The session for which the offsets should be loaded.
  427. run : str or int
  428. The run for which the offsets should be loaded.
  429. only_word_initial_phonemes : bool, optional
  430. Whether to retrieve only word-initial phonemes. Defaults to `False`.
  431. Returns
  432. -------
  433. pandas.DataFrame
  434. A DataFrame containing at least two columns: 'phoneme' and 'onset'.
  435. """
  436. if dataset == 'Armani':
  437. # handle naming of session 10:
  438. if session < 10:
  439. sess = '00' + str(session)
  440. else:
  441. sess = '0' + str(session)
  442. # get path to events file:
  443. dir_path = '/project/3018059.03/Lingpred/data/Armani/'
  444. filepath = 'sub-00' + str(subject) +'/' + 'ses-' + sess +'/'+ 'meg/'
  445. filename = 'sub-00' + str(subject) + '_ses-' + sess + '_task-compr_events.tsv'
  446. # read pandas DataFrame for the entire session:
  447. annotations = pd.read_csv(dir_path+filepath+filename, sep='\t')
  448. # get list with word_onsets for this run:
  449. word_onset_name = ['word_onset_0{}'.format(run)]
  450. df_words = annotations[annotations.type.isin(word_onset_name)]
  451. df_words = df_words[df_words.value != 'sp']
  452. word_onsets = df_words.onset
  453. if correct_wav_onset:
  454. # now look for the timing of the audio onset for this run:
  455. index_first_word = df_words.index[0] # index for the first word in this run
  456. for i in np.arange(index_first_word, -1, -1): # interate from there backwards
  457. if annotations.iloc[i].type == 'wav_onset': # to the most recent wave onset
  458. audio_onset = annotations.iloc[i].onset # and get it's onset time
  459. break # break out of for loop
  460. # get list with phoneme_onsets for this run::
  461. onset_name = ['phoneme_onset_0{}'.format(run)]
  462. # get only phonemes and clean data frame
  463. df_phonemes = annotations[annotations.type.isin(onset_name)]
  464. if only_word_inital_phonemes:
  465. df_phonemes = df_phonemes[df_phonemes.onset.isin(word_onsets)]
  466. else: df_phonemes = df_phonemes[df_phonemes.value != 'sp']
  467. if correct_wav_onset:
  468. #convert times to audio onset times:
  469. df_phonemes.onset = df_phonemes.onset - audio_onset
  470. # add a column with the offsets:
  471. offsets = df_phonemes.onset + df_phonemes.duration
  472. df_phonemes['offset'] = offsets
  473. else: raise ValueError('Dataset variable must be "Armani"')
  474. return df_phonemes
  475. # !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
  476. # DUMMY FUNCTIONS: TO BE IMPLEMENTED ONCE WE HAVE SOURCE LEVEL DATA FOR ALL DATASETS:
  477. # !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
  478. # this function may be necessary to remove events which are not in the GPT-2 word embeddings
  479. def clean_events(annotations_words, subject, session, dataset):
  480. return annotations_words
  481. def get_sessions(dataset, subject):
  482. if dataset =='Armani':
  483. return np.arange(1,11) # Armani has 10 sessions: from 1-10
  484. else:
  485. raise ValueError("Only defined for the Armeni dataset at the moment.")
  486. # get's runs when using Micha'methods, otherwise returns [1] as a list
  487. def get_runs(dataset, subject, session):
  488. if dataset == "Armani":
  489. return _runs_in_session(session,sub_i=subject)
  490. else:
  491. return [1]
  492. def _load_full_text(sess_i,run_i,without_breaks=True):
  493. ASH_sess2runs=lambda x : {1:7,2:7,4:8,6:7,8:7}.get(x,6)
  494. """Load the full text as a single string for each run."""
  495. sess_str='0'+str(sess_i) if sess_i<10 else str(sess_i)
  496. fname=PROJ_ROOT/'Lingpred'/'data' / 'Armani' /'stimuli'/ "{}_{}.txt".format(sess_str,run_i)
  497. with open(fname, 'r') as file:
  498. full_text = file.read().replace('\n', ' ') if without_breaks else file.read()
  499. full_text=full_text.replace(' ',' ')
  500. return(full_text)
  501. def _runs_in_session(sess_i:int,sub_i:Union[None,int])->list:
  502. """from session number (and, optionally, subject number), get list of runs"""
  503. def _sess2nruns(sess_i,sub_i=None):
  504. """
  505. (subject number has to be given to account for aborted runs in pilot sub, and sub-003).
  506. (in subj3-sess8, run3 (run 7 !?) is missing)
  507. """
  508. if sub_i is None: sub_i=1
  509. if sub_i<0:
  510. sess2runs=lambda x : {1:3,2:7,4:7,6:7,8:7}.get(x,6)
  511. elif sub_i==3:
  512. sess2runs=lambda x : {1:7,2:7,4:8,6:7}.get(x,6)
  513. else:
  514. sess2runs=lambda x : {1:7,2:7,4:8,6:7,8:7}.get(x,6)
  515. return(sess2runs(sess_i))
  516. runs=list(range(1, 1 + _sess2nruns(sess_i,sub_i=sub_i)))
  517. return(runs)

io.py at commit 95c6375, under MIT · at the source

Overview

  1. Donders Institute for Brain Cognition and Behaviour, Nijmegen, Netherlands
  2. Institute of Psychology, Jagiellonian University, Kraków, Poland
  3. Amsterdam Brain and Cognition, University of Amsterdam, Amsterdam, Netherlands
Journal: eLife, volume 14, article RP106543
Dates: published online 10 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.7554/elife.106543 · PMID 41960890 · PMCID PMC13068430 · OpenAlex W4409527720
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cognitive (subfield)
Methods: Statistics, Machine learning, Preprocessing, Spectral & time-frequency, fMRI & imaging, Connectivity
Keywords: Human
MeSH: Brain*, Speech Perception*, Humans, Language (* major topic)
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: POLONEZ BIS (project No. 2022/47/P/HS6/02294); European Research Council (101000942, 945339, Skłodowska-Curie grant agreement No. 945339, No. 101000942 SURPRISE); Dutch Research Council (NWO) (VI.C.231.043)
Citations: cited by 10 papers (Europe PMC); 42 references in the paper

Abstract

The human brain is thought to constantly predict future words during language processing. Recently, a new approach emerged that aims to capture neural prediction directly by using vector representations of words (embeddings) to predict brain activity prior to word onset. Two findings have been proposed as hallmarks of neural next-word prediction: (i) significant encoding prior to word onset and (ii) its modulation by word predictability. However, natural language is rife with temporal correlations, where upcoming words share statistical information with preceding ones. This raises a critical question: Do these hallmarks emerge from the brain actively predicting future content, or might they be equally well explained by the regression model exploiting these inherent stimulus dependencies? To distinguish between these alternatives, we applied the same encoding analysis to passive control systems, i.e., representational systems that encode the stimulus but cannot predict upcoming words. We show that both hallmarks emerge in two such control systems, namely in word embeddings themselves and in speech acoustics. We further show that proposed methods to correct for these dependencies are insufficient, as the effects persist even after such corrections. Together, these results suggest that pre-onset prediction of brain activity might reflect dependencies in natural language rather than predictive computations. This questions the extent to which this new encoding-based method can be used to study prediction in the brain.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

InesSchoenmann/Lingpred

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 95c6375f8b014233febd116eddd6fecb03baccd3, 9 July 2026
Languages: Python (12), Jupyter (6)
Size: 122 files, 18 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, 6 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (12 files), pandas (7 files), Matplotlib (4 files), SciPy (4 files), h5py (3 files), MNE-Python (3 files), seaborn (3 files), MNE-BIDS (2 files), scikit-learn (2 files), PyTorch (1 file), Hugging Face Transformers (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
20 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 18 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The main dataset used here, Armeni et al., 2019’s few-subject MEG dataset, was made available with the original publication at https://doi.org/10.1038/s41597-022-01382-7. The additional multi-subject dataset by Gwilliams et al., 2022 is available at https://doi.org/10.17605/OSF.IO/AG3KJ. The stimuli and model features used in Goldstein et al., 2022b are available at https://openneuro.org/datasets/ds005574/versions/1.0.2 and the audio is available at https://www.thisamericanlife.org/631/so-a-monkey-and-a-horse-walk-into-a-bar/act-one-0. The code used for modelling analyses and plotting is available at https://github.com/InesSchoenmann/Lingpred (copy archived at Schoenmann, 2026).

The following previously published datasets were used:

Armeni K, Güçlü U, van Gerven M, Schoffelen J-M. 2022. A 10-hour within-participant magnetoencephalography narrative dataset to test models of naturalistic language comprehension. Donders Data Repository.

Zada Z, Nastase SA, Aubrey B, Jalon I, Goldstein A, Michelmann S, Wang H, Hasenfratz L, Doyle W, Friedman D, Dugan P, Melloni L, Devore S, Devinsky O, Flinker A, Hasson U. 2025. The "Podcast" ECoG dataset. OpenNeuro.

Gwilliams L, Flick G, Marantz A, Pylkkänen L, Poeppel D, King JR. 2022. MASC-MEG. Open Science Framework.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 4 authors, 1 keyword, 4 MeSH terms, 3 funders, 40 references.

Cite

This paper

Schönmann, I., Szewczyk, J., de Lange, F. P., & Heilbron, M. (2026). Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs. eLife, 14, RP106543. https://doi.org/10.7554/elife.106543

BibTeX

@article{schonmann2026stimulus,
author = {Schönmann, Inés and Szewczyk, Jakub and de Lange, Floris P and Heilbron, Micha},
title = {{Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs}},
journal = {eLife},
year = {2026},
month = apr,
volume = {14},
pages = {RP106543},
publisher = {eLife Sciences Publications, Ltd},
issn = {2050-084X},
doi = {10.7554/elife.106543},
url = {https://doi.org/10.7554/elife.106543},
pmid = {41960890},
pmcid = {PMC13068430}
}

RIS

TY - JOUR
AU - Schönmann, Inés
AU - Szewczyk, Jakub
AU - de Lange, Floris P
AU - Heilbron, Micha
TI - Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs
T2 - eLife
J2 - Elife
PY - 2026
DA - 2026/04/10
VL - 14
SP - RP106543
SN - 2050-084X
PB - eLife Sciences Publications, Ltd
DO - 10.7554/elife.106543
UR - https://doi.org/10.7554/elife.106543
LA - en
ER -

CSL-JSON

{
"id": "10.7554/elife.106543",
"type": "article-journal",
"title": "Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs",
"container-title": "eLife",
"author": [
{
"family": "Schönmann",
"given": "Inés"
},
{
"family": "Szewczyk",
"given": "Jakub"
},
{
"family": "de Lange",
"given": "Floris P"
},
{
"family": "Heilbron",
"given": "Micha"
}
],
"container-title-short": "Elife",
"volume": "14",
"page": "RP106543",
"DOI": "10.7554/elife.106543",
"PMID": "41960890",
"PMCID": "PMC13068430",
"ISSN": "2050-084X",
"publisher": "eLife Sciences Publications, Ltd",
"URL": "https://doi.org/10.7554/elife.106543",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1162/imag.a.1227 [code]
Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Hugging Face Transformers, MNE-Python, h5py, 7 other tools, cognitive, 6 references
[2] doi:10.7554/elife.101204 [code]
Larger language models better align with neural representations of natural language.
Journal: eLife
In common: Hugging Face Transformers, PyTorch, scikit-learn, 3 other tools, 7 references
[3] doi:10.1038/s41467-026-72253-7 [code]
Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds.
Journal: Nature communications
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 5 references
[4] doi:10.1073/pnas.2422097122
Hierarchical dynamic coding coordinates speech comprehension in the human brain
Journal: n/a
In common: OSF ag3kj, cognitive, 5 references
[5] doi:10.1038/s41467-026-75662-w [code]
Distinct Roles of Deep and Superficial Cortical Layers in Tone Prediction, Comparison, and Adaptation in Human Auditory Cortices.
Journal: Nature communications
In common: MNE-Python, seaborn, scikit-learn, 4 other tools, cognitive, author Floris P de Lange
[6] doi:10.1167/jov.26.5.7 [code]
Representations in vision and language converge in a shared, multidimensional space of perceived similarities.
Journal: Journal of vision
In common: h5py, PyTorch, seaborn, 5 other tools, cognitive, 3 references
[7] doi:10.1162/nol.a.244 [code]
A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.
Journal: Neurobiology of language (Cambridge, Mass.)
In common: Hugging Face Transformers, MNE-Python, PyTorch, 4 other tools, 3 references
[8] doi:10.1016/j.isci.2026.117180 [code]
Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.
Journal: iScience
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 2 references
[9] doi:10.1038/s41598-026-41532-0 [code]
Prediction, syntax and semantic grounding in the brain and large language models.
Journal: Scientific reports
In common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
[10] doi:10.1038/s41597-025-05174-7 [code]
A large-scale MEG and EEG dataset for object recognition in naturalistic scenes
Journal: n/a
In common: MNE-BIDS, MNE-Python, h5py, 7 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.