Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.
The 3 matches
- [1] § Methods › Data ↔ lingpred_new/io.py, lines 79–223 · score 0.66 · 0.1–40 Hz, baseline correction, filtered, word onset, Sherlock, window
- [2] § Methods › Control system one: self-predictability analysis ↔ lingpred_new/plotting.py, lines 744–892 · score 0.60 · incorrect predictions, unpredicted words, GloVe, split, vectors, model
- [3] § Methods › MEG encoding modelling › Source selection ↔ lingpred_new/io.py, lines 79–223 · score 0.56 · post onset encoding, post word onset, window, activation, 500 ms, pre
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 699 lines · 25 KB · MIT · 2 matches
- import glob
- import shutil
- import os
- import sys
- from pathlib import Path
- import numpy as np
- import pickle
- from typing import Sequence,Union,Optional
- import imp
- import pandas as pd
- import tgt
- import mne
- import mne_bids
- import h5py
- from itertools import compress
- from scipy.io import loadmat
- PROJ_ROOT = '/project/3018059.03'
- # -----------------------------------------------------------#
- # Routine to load text #
- # -----------------------------------------------------------#
- def get_text_per_session(dataset: str, session: int, subject: int):
- """
- Loads text for a given dataset, session, and subject.
- Parameters
- ----------
- dataset : str
- Name of the dataset. Options: "Gwilliams" or "Armani"
- session : int
- Session for which the text is supposed to be loaded.
- - `range(11)` for Armani
- - `range(4)` for Gwilliams
- subject :
- Subject for whom the text is supposed to be loaded.
- Returns
- -------
- str
- The text for all runs as a single string.
- """
- if dataset =='Armani':
- runs = _runs_in_session(sess_i=session,sub_i=subject)
- text_all_runs = ''
- for run in runs:
- this_run_text = _load_full_text(sess_i = session, run_i = run ,without_breaks=True)
- text_all_runs = text_all_runs + ' ' + this_run_text
- return text_all_runs
- if dataset == 'Gwilliams':
- if str(session) =='0':
- textname = 'lw1.txt'
- if str(session) == '1':
- textname = 'cable_spool_fort.txt'
- if str(session) =='2':
- textname = 'easy_money.txt'
- if str(session) == '3':
- textname = 'the_black_willow.txt'
- fname = '/project/3018059.03/Lingpred/data/Gwilliams/stimuli/text/' + textname
- with open(fname, 'r') as file:
- text = file.read()
- return text
- # -----------------------------------------------------------#
- # NEW ROUTINES TO LOAD GWILLIAMS OR ARMANI MEG DATA #
- # -----------------------------------------------------------#
- def get_neural_data(dataset: str, sessions: list, subject: int, task = 'compr', datatype='source', channels=None, window_size=100, band=(0.1, 40), baseline=None):
- """
- Computes the average activation at each channel and lag relative to word onset.
- Parameters
- ----------
- dataset : str
- Either "Gwilliams", "Sherlock", or "Armani".
- subject : str or int
- Subject identifier. For Armani dataset: e.g., "001". For Gwilliams dataset: e.g., "01".
- task : str or int
- Task identifier. For Armani dataset: "compr". For Gwilliams dataset: 0, 1, 2, or 3, corresponding to:
- - 0 = lw1
- - 1 = cable spool fort
- - 2 = easy money
- - 3 = black willow
- channels : str, optional
- Can be 'None', 'language', 'signal', 'pre-onset', or 'post-onset'.
- - `None`: Returns all channels.
- - `'language'`: Returns only language-related channels.
- - `'signal'` (subject-specific): Returns only channels with relevant encoding (max value > 25% of max overall value).
- - `'pre-onset'` (subject-specific): Returns only channels with predominantly pre-word-onset encoding.
- - `'post-onset'` (subject-specific): Returns only channels with predominantly post-word-onset encoding.
- window_size : int, optional
- Window size for averaging the neural data with lag = 25 ms. If `None`, no averaging is performed.
- band : tuple of float
- Bandpass filter range to be used. Can be (0.5, 8) or (1, 40).
- baseline : bool or None
- Whether to perform baseline correction. If `None`, no correction is performed.
- Returns
- -------
- numpy.ndarray
- An array containing the average activation at each channel and lag relative to word onset.
- The returned array has shape (nr_channels, nr_epochs, nr_lags).
- pandas.DataFrame
- A dataframe containing the word of each epoch, the onset time, and additional information.
- list
- A list of dropped epochs.
- """
- if dataset == "Gwilliams":
- task = task
- sessions = sessions
- #get_sessions(dataset, subject) # These methods are still missing
- if channels == 'language':
- sources = get_language_channels(dataset)
- if channels == 'signal':
- sources = get_signal_channels(dataset, subject)
- if channels == 'pre-onset':
- sources = get_signal_channels(dataset, subject, pre=True)
- if channels == 'post-onset':
- sources = get_signal_channels(dataset, subject, post=True)
- else: sources = None # this includes all channels
- for session in sessions:
- bad_epochs_all_runs = [] # list with indices of dropped epochs
- length_all_epochs = 0 # constant to be added to the indices for each run
- runs = get_runs(dataset, subject, session) # This now only returns [1] for Gwilliams & Armani
- for run in runs:
- # load raw data and get a dataframe for annotations incl. word ID and onset/offset
- raw_data = load_raw_data(dataset, subject, session, run, task, datatype=datatype, band=band)
- annotations_df = get_words_onsets_offsets(raw_data, dataset, subject, session, run)
- onset_words = annotations_df.onset
- # create events
- events = np.c_[onset_words * raw_data.info['sfreq'],
- np.zeros((len(onset_words), 1)),
- np.ones((len(onset_words), 1))].astype(int)
- # now down-sample the data to 200Hz for computational reasons:
- # by passing the events array we make sure our timing is not jittered:
- raw_data, events = raw_data.resample(sfreq=200, events=events)
- # create epochs
- epochs = mne.Epochs(raw_data,
- events,
- tmin=-2,
- tmax=2,
- baseline=baseline, # (-2, 0) would be the default baseline
- picks=sources,
- metadata = annotations_df,
- event_repeated='drop',
- preload=True)
- # get indices of bad epochs:
- bad_epochs_this_run = [index + length_all_epochs for index, dl in enumerate(epochs.drop_log) if len(dl)]
- print(bad_epochs_this_run)
- length_all_epochs = length_all_epochs + len(epochs) + len(bad_epochs_this_run)
- bad_epochs_all_runs = bad_epochs_all_runs + bad_epochs_this_run
- # get average per lag
- if window_size:
- data_per_lag = get_mean_data_lag(epochs, window_size) # array of shape (nr_epochs, nr_channels, nr_lags)
- else:
- data_per_lag = epochs.get_data()
- # stack on top for all runs in one session & append annotations dataframe
- if run == runs[0] : # first run in session
- data_per_lag_all_runs = data_per_lag
- comb_annotation_df = epochs.metadata
- else:
- data_per_lag_all_runs = np.vstack((data_per_lag_all_runs, # array of shape:
- data_per_lag)) # (nr_all_epochs, nr_channels, nr_lags)
- comb_annotation_df = comb_annotation_df.append(epochs.metadata)
- print(bad_epochs_all_runs)
- # stack on top for all sessions & append annotations dataframe
- if session == 1 or len(sessions)==1:
- data_per_lag_all_sess = data_per_lag_all_runs
- annotation_df_all_sess = comb_annotation_df
- else:
- data_per_lag_all_sess = np.vstack((data_per_lag_all_sess, # array of shape:
- data_per_lag_all_runs)) # (nr_all_epochs, nr_channels, nr_lags)
- annotation_df_all_sess = annotation_df_all_sess.append(comb_annotation_df)
- # for subject 3, session 8 words 3841-4170 are scrambled: and need to be dropped from the df and the neural data:
- if session==8 and subject==3:
- annotation_df_all_sess.drop(index=annotation_df_all_sess[3841:4170].index, inplace=True)
- np.delete(data_per_lag_all_sess, np.arange(3841,4170))
- # change shape of the array such that one can easily loop over the channels: (nr_channels, nr_all_epochs, nr_lags)
- data_per_lag_all_sess = np.swapaxes(data_per_lag_all_sess,0,1)
- # return array with neural data and annotation data frame:
- return data_per_lag_all_sess, annotation_df_all_sess, bad_epochs_all_runs
- # ------------------------------------------------------------#
- # Associated New Auxiliary Functions: #
- # ------------------------------------------------------------#
- def drop_nans(y, X):
- '''
- Removes words containing NaN values from the neural data and GPT layer activations.
- Parameters
- ----------
- y : numpy.ndarray
- Neural data array of shape (channels, words, lags), containing the averaged MEG data.
- X : numpy.ndarray
- GPT layer activations, an array of shape (words, dimensions + 1).
- nan_ids : list of int
- List containing the indices of words that contain NaN values.
- Returns
- -------
- numpy.ndarray
- The `y` array with words containing NaNs removed.
- numpy.ndarray
- The `X` array with words containing NaNs removed.
- '''
- # get NaN values in neural data:
- nan_ids = get_nans(y)
- y = np.swapaxes(y,0,1) # swap axes such that rows == words
- # drop rows containing NaNs in y and X
- y = np.delete(y, nan_ids, axis=0)
- X = np.delete(X, nan_ids, axis=0)
- y = np.swapaxes(y,0,1) # swap axes such that rows == channels
- # print shapes
- print(X.shape, y.shape)
- return y, X
- def get_nans(neural_data):
- """
- Identifies words containing NaN values in the neural data.
- Parameters
- ----------
- neural_data : numpy.ndarray
- Array of shape (channels, words, lags), containing the averaged MEG data.
- Returns
- -------
- list of int
- List containing the indices of words that contain NaN values (to be dropped).
- """
- # initialise counter and index list
- count = 0
- nan_ids = []
- # loop over words (rows) to find rows containing NaNs
- for i, word in enumerate(neural_data[0]): # just do this for one channel
- if np.any(np.any(np.isnan(word))):
- count+=1
- nan_ids.append(i)
- print('There are {} words with NaNs which will be dropped.'.format(count))
- return nan_ids
- def get_language_channels(dataset: str, get_areas=False):
- '''
- Retrieves the list of channel names related to the language system for a given dataset.
- Parameters
- ----------
- dataset : str
- The dataset name. Can be either 'sherlock', 'Armani', or 'Gwilliams'.
- Returns
- -------
- list of str
- A list of channel names related to the language system.
- '''
- if dataset=='sherlock' or dataset=='Armani':
- source_info = ASH_load_source_info()
- ch_names = source_info['lbls_language']
- if get_areas:
- ch_names = list(compress(source_info['areas'], source_info['language_mask']))
- else: raise VallueError('get_language_channels is only implemented for the Armani dataset')
- return ch_names
- def get_lags(window_size, lag=0.025, start=-2, end=2):
- '''
- Returns an list of shape (nr_lags, 3) each row containing (lag_index, start_window, end_window) in sec
- '''
- starts = np.round(np.arange(start, end-window_size+lag, lag), 3)
- ends = np.round([s + window_size for s in starts], 3)
- arr = [[nr, start, end] for nr, (start, end) in enumerate(zip(starts, ends))]
- return arr
- def get_mean_data_lag(epochs, window_size):
- '''
- Param: MNE epochs object
- Returns average data for each window: array of shape (nr_epochs, nr_channels, nr_lags)
- '''
- # transform window_size from ms to s (round to 2 decimals to ensure precision, i.e. 0.05, 0.10, 0.15, 0.2)
- window_size = round(window_size/1000, 2)
- lags = get_lags(window_size=window_size)
- nr_lags = len(lags)
- for i in range(nr_lags):
- nr, start, end = lags[i]
- data = epochs.get_data(tmin=start, tmax=end) # ndarray of shape (nr_ep, nr_ch, nr_data)
- mean_data = np.mean(data, axis=2) # ndarray of shape (nr_epochs, nr_channels)
- del data
- if i==0:
- com_data = mean_data
- else:
- com_data = np.dstack([com_data, mean_data])
- return com_data # array of shape (nr_epochs, nr_channels, nr_lags)
- def load_raw_data(dataset: str, subject: int, session: int, run: int, task: str, datatype='raw', band=(1, 40)):
- """
- Loads raw data object using the MNE.
- Parameters
- ----------
- dataset : str
- The dataset name. Can be "Gwilliams", "sherlock", or "Armani".
- subject : str or int
- The subject identifier for which the data should be loaded.
- session : str or int
- The session for which the data should be loaded.
- task : str
- The task for which the data should be loaded. '0' for Gwilliams and 'compr' for Armani.
- band : tuple of float
- A tuple specifying the frequency band of the filtered Sherlock data.
- datatype : str
- The type of data to load:
- - `'filtered'` if the dataset is Gwilliams (filtered sensor data).
- - `'source'` if the dataset is Armani (filtered source data).
- Returns
- -------
- mne.io.Raw
- An MNE raw object with the filtered data for Gwilliams and source localized data for Armani.
- """
- # set root to path:
- if dataset=="Gwilliams":
- root='/project/3018059.03/Lingpred/data/Gwilliams/derived/'
- sess = str(session)
- if subject < 10:
- sub = '0' + str(subject)
- else:
- sub = str(subject)
- elif dataset=="Armani":
- root='/project/3018059.03/Lingpred/data/Armani/'
- sub = '00' + str(subject)
- if session < 10:
- sess = '00' + str(session)
- else:
- sess = '0' + str(session)
- # set path to raw data:
- if datatype == 'raw':
- datatype = 'meg'
- bids_path = mne_bids.BIDSPath(subject=sub,
- session=sess,
- task=task,
- datatype=datatype,
- root=root)
- # load data:
- if dataset=="Armani":
- # if we want the raw source localised data we need to load it from the source foulder:
- if datatype=='source':
- # file name, e.g.: 1-1_lcmv-data_0.1-40raw.fif
- fname = str(subject) + '-' + str(session) + '_lcmv-data_' + str(band[0]) + '-' + str(band[1]) + 'raw.fif'
- # handle naming of session 10
- if session < 10:
- session = '0'+ str(session)
- # full file path
- fif_path = root + 'sub-00' + str(subject) + '/ses-0' + str(session) + '/source/' + fname
- #read data
- raw = mne.io.read_raw_fif(fif_path)
- # else if we want raw data:
- else:
- raw = mne.io.read_raw_ctf(bids_path)
- if dataset=="Gwilliams":
- # if we want the filtered data:
- if datatype=='filtered':
- band=(0.1, 40)
- # file name, e.g.: 1-1_lcmv-data_0.1-40raw.fif
- fname = str(subject)+'-'+sess+'_filtered_data_'+str(band[0])+'-'+str(band[1])+'_task-'+task+'_'+'raw.fif'
- # full file path
- folder = '/project/3018059.03/data/Gwilliams/'
- fif_path = folder + 'sub-' + sub + '/ses-' + sess + '/filtered/' + fname
- # return False if the file doesn't exist
- if not Path(fif_path).is_file():
- return False
- #read data
- raw = mne.io.read_raw_fif(fif_path)
- # if we want the raw data:
- else:
- raw = mne_bids.read_raw_bids(bids_path)
- if dataset not in ['Gwilliams', 'Armani']:
- raise ValueError('Dataset variable must be either "Gwilliams", or "Armani"')
- # load raw data
- raw.load_data()
- return raw
- def get_words_onsets_offsets(raw_data, dataset:str, subject:int, session:int, run:int):
- """
- Loads word onset offsets for a given dataset, subject, session, and run.
- Parameters
- ----------
- raw_data : mne.io.Raw
- The raw data object.
- dataset : str
- The dataset name. Can be "Gwilliams", "Armani", or "Sherlock".
- subject : str or int
- The subject identifier for which the offsets should be loaded.
- session : str or int
- The session for which the offsets should be loaded.
- run : str or int
- The run for which the offsets should be loaded.
- Returns
- -------
- pandas.DataFrame
- A DataFrame containing at least two columns: 'word' and 'onset'.
- """
- if dataset == 'Gwilliams':
- df = raw_data.annotations.to_data_frame()
- df = pd.DataFrame(df.description.apply(eval).to_list())
- onsets = raw_data.annotations.onset
- # make a column for the onsets:
- df['onset'] = onsets
- # keep only word onset data
- df_words = df[df['kind']=='word']
- # keep only words which are part of the story (no pseudowords or wordlists)
- df_words = df_words[df_words['condition']=='sentence']
- # deal with shift:
- word_onsets = df.loc[df_words.index + 1].onset.values
- df_words['onset'] = word_onsets
- elif dataset == 'Armani':
- # handle naming of session 10:
- if session < 10:
- sess = '00' + str(session)
- else:
- sess = '0' + str(session)
- # get path to events file:
- dir_path = '/project/3018059.03/Lingpred/data/Armani/'
- filepath = 'sub-00' + str(subject) +'/' + 'ses-' + sess +'/'+ 'meg/'
- filename = 'sub-00' + str(subject) + '_ses-' + sess + '_task-compr_events.tsv'
- # read pandas DataFrame
- annotations = pd.read_csv(dir_path+filepath+filename, sep='\t')
- # type of event is separated by runs: there are at max 7 runs in each session:
- onset_names_list = ['word_onset_01', 'word_onset_02', 'word_onset_03', 'word_onset_04',
- 'word_onset_05', 'word_onset_06', 'word_onset_07', 'word_onset_08']
- # get only words, rename columns (identically to Gwilliams) and clean data frame
- df_words = annotations[annotations.type.isin(onset_names_list)]
- df_words = df_words[df_words.value != 'sp']
- df_words.rename(columns={'value':'word'}, inplace=True)
- df_words = clean_events(df_words, subject, session, dataset)
- else: raise ValueError('Dataset variable must be either "Gwilliams", or "Armani"')
- return df_words
- def get_phonemes_onsets_offsets(dataset:str, subject:int, session:int, run:int,
- only_word_inital_phonemes=True, correct_wav_onset=True):
- """
- Loads phoneme onset offsets for a given dataset, subject, session, and run.
- Parameters
- ----------
- raw_data : mne.io.Raw
- The raw data object.
- dataset : str
- The dataset name. Can be "Gwilliams", "Armani", or "Sherlock".
- subject : str or int
- The subject identifier for which the offsets should be loaded.
- session : str or int
- The session for which the offsets should be loaded.
- run : str or int
- The run for which the offsets should be loaded.
- only_word_initial_phonemes : bool, optional
- Whether to retrieve only word-initial phonemes. Defaults to `False`.
- Returns
- -------
- pandas.DataFrame
- A DataFrame containing at least two columns: 'phoneme' and 'onset'.
- """
- if dataset == 'Armani':
- # handle naming of session 10:
- if session < 10:
- sess = '00' + str(session)
- else:
- sess = '0' + str(session)
- # get path to events file:
- dir_path = '/project/3018059.03/Lingpred/data/Armani/'
- filepath = 'sub-00' + str(subject) +'/' + 'ses-' + sess +'/'+ 'meg/'
- filename = 'sub-00' + str(subject) + '_ses-' + sess + '_task-compr_events.tsv'
- # read pandas DataFrame for the entire session:
- annotations = pd.read_csv(dir_path+filepath+filename, sep='\t')
- # get list with word_onsets for this run:
- word_onset_name = ['word_onset_0{}'.format(run)]
- df_words = annotations[annotations.type.isin(word_onset_name)]
- df_words = df_words[df_words.value != 'sp']
- word_onsets = df_words.onset
- if correct_wav_onset:
- # now look for the timing of the audio onset for this run:
- index_first_word = df_words.index[0] # index for the first word in this run
- for i in np.arange(index_first_word, -1, -1): # interate from there backwards
- if annotations.iloc[i].type == 'wav_onset': # to the most recent wave onset
- audio_onset = annotations.iloc[i].onset # and get it's onset time
- break # break out of for loop
- # get list with phoneme_onsets for this run::
- onset_name = ['phoneme_onset_0{}'.format(run)]
- # get only phonemes and clean data frame
- df_phonemes = annotations[annotations.type.isin(onset_name)]
- if only_word_inital_phonemes:
- df_phonemes = df_phonemes[df_phonemes.onset.isin(word_onsets)]
- else: df_phonemes = df_phonemes[df_phonemes.value != 'sp']
- if correct_wav_onset:
- #convert times to audio onset times:
- df_phonemes.onset = df_phonemes.onset - audio_onset
- # add a column with the offsets:
- offsets = df_phonemes.onset + df_phonemes.duration
- df_phonemes['offset'] = offsets
- else: raise ValueError('Dataset variable must be "Armani"')
- return df_phonemes
- # !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
- # DUMMY FUNCTIONS: TO BE IMPLEMENTED ONCE WE HAVE SOURCE LEVEL DATA FOR ALL DATASETS:
- # !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
- # this function may be necessary to remove events which are not in the GPT-2 word embeddings
- def clean_events(annotations_words, subject, session, dataset):
- return annotations_words
- def get_sessions(dataset, subject):
- if dataset =='Armani':
- return np.arange(1,11) # Armani has 10 sessions: from 1-10
- else:
- raise ValueError("Only defined for the Armeni dataset at the moment.")
- # get's runs when using Micha'methods, otherwise returns [1] as a list
- def get_runs(dataset, subject, session):
- if dataset == "Armani":
- return _runs_in_session(session,sub_i=subject)
- else:
- return [1]
- def _load_full_text(sess_i,run_i,without_breaks=True):
- ASH_sess2runs=lambda x : {1:7,2:7,4:8,6:7,8:7}.get(x,6)
- """Load the full text as a single string for each run."""
- sess_str='0'+str(sess_i) if sess_i<10 else str(sess_i)
- fname=PROJ_ROOT/'Lingpred'/'data' / 'Armani' /'stimuli'/ "{}_{}.txt".format(sess_str,run_i)
- with open(fname, 'r') as file:
- full_text = file.read().replace('\n', ' ') if without_breaks else file.read()
- full_text=full_text.replace(' ',' ')
- return(full_text)
- def _runs_in_session(sess_i:int,sub_i:Union[None,int])->list:
- """from session number (and, optionally, subject number), get list of runs"""
- def _sess2nruns(sess_i,sub_i=None):
- """
- (subject number has to be given to account for aborted runs in pilot sub, and sub-003).
- (in subj3-sess8, run3 (run 7 !?) is missing)
- """
- if sub_i is None: sub_i=1
- if sub_i<0:
- sess2runs=lambda x : {1:3,2:7,4:7,6:7,8:7}.get(x,6)
- elif sub_i==3:
- sess2runs=lambda x : {1:7,2:7,4:8,6:7}.get(x,6)
- else:
- sess2runs=lambda x : {1:7,2:7,4:8,6:7,8:7}.get(x,6)
- return(sess2runs(sess_i))
- runs=list(range(1, 1 + _sess2nruns(sess_i,sub_i=sub_i)))
- return(runs)
io.py at commit 95c6375, under MIT · at the source
Overview
- Donders Institute for Brain Cognition and Behaviour, Nijmegen, Netherlands
- Institute of Psychology, Jagiellonian University, Kraków, Poland
- Amsterdam Brain and Cognition, University of Amsterdam, Amsterdam, Netherlands
Abstract
The human brain is thought to constantly predict future words during language processing. Recently, a new approach emerged that aims to capture neural prediction directly by using vector representations of words (embeddings) to predict brain activity prior to word onset. Two findings have been proposed as hallmarks of neural next-word prediction: (i) significant encoding prior to word onset and (ii) its modulation by word predictability. However, natural language is rife with temporal correlations, where upcoming words share statistical information with preceding ones. This raises a critical question: Do these hallmarks emerge from the brain actively predicting future content, or might they be equally well explained by the regression model exploiting these inherent stimulus dependencies? To distinguish between these alternatives, we applied the same encoding analysis to passive control systems, i.e., representational systems that encode the stimulus but cannot predict upcoming words. We show that both hallmarks emerge in two such control systems, namely in word embeddings themselves and in speech acoustics. We further show that proposed methods to correct for these dependencies are insufficient, as the effects persist even after such corrections. Together, these results suggest that pre-onset prediction of brain activity might reflect dependencies in natural language rather than predictive computations. This questions the extent to which this new encoding-based method can be used to study prediction in the brain.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.
InesSchoenmann/Lingpred
95c6375f8b014233febd116eddd6fecb03baccd3, 9 July 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
20 files
- __init__.py, Python, 1 line
- gpt2/
__init__.py , Python, 1 line - gpt2/
model.py , Python, 734 lines - lingpred_audio/
__init__.py , Python, 5 lines - lingpred_audio/
audio.py , Python, 517 lines - lingpred_new/
__init__.py , Python, 5 lines - lingpred_new/
encoding_analysis.py , Python, 752 lines - lingpred_new/
io.py , Python, 699 lines, 2 matches - lingpred_new/
plotting.py , Python, 993 lines, 1 match - lingpred_new/
preprocessing.py , Python, 276 lines - lingpred_new/
utils.py , Python, 683 lines - notebooks/
Audio_Encoding.ipynb , Jupyter, 1,108 lines - notebooks/
Compute_Brainscore.ipynb , Jupyter, 41 lines - notebooks/
Compute_Selfpredictabili , Jupyter, 532 linesty.ipynb - notebooks/
Figure.ipynb , Jupyter, 141 lines - notebooks/
Make_Audio_X_Matrix.ipyn , Jupyter, 100 linesb - notebooks/
__init__.py , Python, 1 line - notebooks/
adjusting_goldstein_audi , Jupyter, 25 lineso.ipynb - LICENSE, License, 21 lines
- README.md, Text, 37 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 18 scripts, each with its path and the digest of its content;
- 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- openneuro:ds005574, at OpenNeuro; found in “Data availability”
- osf:ag3kj, at OSF; found in “Data availability”
Data availability
The main dataset used here, Armeni et al., 2019’s few-subject MEG dataset, was made available with the original publication at https://
The following previously published datasets were used:
Armeni K, Güçlü U, van Gerven M, Schoffelen J-M. 2022. A 10-hour within-participant magnetoencephalography narrative dataset to test models of naturalistic language comprehension. Donders Data Repository.
Zada Z, Nastase SA, Aubrey B, Jalon I, Goldstein A, Michelmann S, Wang H, Hasenfratz L, Doyle W, Friedman D, Dugan P, Melloni L, Devore S, Devinsky O, Flinker A, Hasson U. 2025. The "Podcast" ECoG dataset. OpenNeuro.
Gwilliams L, Flick G, Marantz A, Pylkkänen L, Poeppel D, King JR. 2022. MASC-MEG. Open Science Framework.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 4 authors, 1 keyword, 4 MeSH terms, 3 funders, 40 references.
Cite
This paper
Schönmann, I., Szewczyk, J., de Lange, F. P., & Heilbron, M. (2026). Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs. eLife, 14, RP106543. https://
BibTeX
@article{schonmann2026st
author = {Schönmann, Inés and Szewczyk, Jakub and de Lange, Floris P and Heilbron, Micha},
title = {{Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs}},
journal = {eLife},
year = {2026},
month = apr,
volume = {14},
pages = {RP106543},
publisher = {eLife Sciences Publications, Ltd},
issn = {2050-084X},
doi = {10.7554/
url = {https://
pmid = {41960890},
pmcid = {PMC13068430}
}
RIS
TY - JOUR
AU - Schönmann, Inés
AU - Szewczyk, Jakub
AU - de Lange, Floris P
AU - Heilbron, Micha
TI - Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs
T2 - eLife
J2 - Elife
PY - 2026
DA - 2026/
VL - 14
SP - RP106543
SN - 2050-084X
PB - eLife Sciences Publications, Ltd
DO - 10.7554/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.7554/
"type": "article-journal",
"title": "Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs",
"container-title": "eLife",
"author": [
{
"family": "Schönmann",
"given": "Inés"
},
{
"family": "Szewczyk",
"given": "Jakub"
},
{
"family": "de Lange",
"given": "Floris P"
},
{
"family": "Heilbron",
"given": "Micha"
}
],
"container-title-short":
"volume": "14",
"page": "RP106543",
"DOI": "10.7554/
"PMID": "41960890",
"PMCID": "PMC13068430",
"ISSN": "2050-084X",
"publisher": "eLife Sciences Publications, Ltd",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1162/imag.a.1227 [code]
- Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.Journal: Imaging neuroscience (Cambridge, Mass.)In common: Hugging Face Transformers, MNE-Python, h5py, 7 other tools, cognitive, 6 references
- [2] doi:10.7554/elife.101204 [code]
- Larger language models better align with neural representations of natural language.Journal: eLifeIn common: Hugging Face Transformers, PyTorch, scikit-learn, 3 other tools, 7 references
- [3] doi:10.1038/s41467-026-72253-7 [code]
- Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds.Journal: Nature communicationsIn common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 5 references
- [4] doi:10.1073/pnas.2422097122
- Hierarchical dynamic coding coordinates speech comprehension in the human brainJournal: n/aIn common: OSF ag3kj, cognitive, 5 references
- [5] doi:10.1038/s41467-026-75662-w [code]
- Distinct Roles of Deep and Superficial Cortical Layers in Tone Prediction, Comparison, and Adaptation in Human Auditory Cortices.Journal: Nature communicationsIn common: MNE-Python, seaborn, scikit-learn, 4 other tools, cognitive, author Floris P de Lange
- [6] doi:10.1167/jov.26.5.7 [code]
- Representations in vision and language converge in a shared, multidimensional space of perceived similarities.Journal: Journal of visionIn common: h5py, PyTorch, seaborn, 5 other tools, cognitive, 3 references
- [7] doi:10.1162/nol.a.244 [code]
- A Novel Approach to Map the Causal Impact of Brain Stimulation on Semantic Processing With Language Models.Journal: Neurobiology of language (Cambridge, Mass.)In common: Hugging Face Transformers, MNE-Python, PyTorch, 4 other tools, 3 references
- [8] doi:10.1016/j.isci.2026.117180 [code]
- Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.Journal: iScienceIn common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 2 references
- [9] doi:10.1038/s41598-026-41532-0 [code]
- Prediction, syntax and semantic grounding in the brain and large language models.Journal: Scientific reportsIn common: Hugging Face Transformers, MNE-Python, PyTorch, 6 other tools, cognitive, 1 reference
- [10] doi:10.1038/s41597-025-05174-7 [code]
- A large-scale MEG and EEG dataset for object recognition in naturalistic scenesJournal: n/aIn common: MNE-BIDS, MNE-Python, h5py, 7 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 18 scripts, and 3 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:0ba0e0d6c5c034c1…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
