OSCR

Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds.

Code ↔ Paper

9 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 9 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Results › Key model comparison results do not replicate with contiguous splits ↔ misc_code/word_count.ipynb, lines 33–135 · score 0.99 · bidirectional transformer models, predicted neural responses, unidirectional transformer models, findings relied, GPT2 family, introduced biases
  2. [2] § Results › Across-layer neural predictivity patterns between shuffled and contiguous splits are anti-correlated ↔ misc_code/word_count.ipynb, lines 33–135 · score 0.99 · highly anti correlated, later intermediate layers, early intermediate layers, intermediate layers achieved, late layers, layer trends
  3. [3] § Results › Positional information and word rate fully account for the neural predictivity of Untrained GPT2XL ↔ misc_code/word_count.ipynb, lines 20–31 · score 0.88 · surprisingly high neural, biases computations, transformer architecture, neural variance explained, untrained transformers, behavioral
  4. [4] § Methods › Language models ↔ generate_activations/LLM_v2.py, lines 46–125 · score 0.84 · XLM CLM EnFr, Transfo XL WT103, CTRL, RWKV, Llama, instruction
  5. [5] § Methods › Regression methods ↔ run_reg_scripts/helper_funcs.py, lines 676–782 · score 0.75 · ridge regression, linear regression, feature spaces, expanded, GPU, himalaya
  6. [6] § Methods › Selection of best layer ↔ generate_activations/clean_untrained.py, the whole file · a weak match · score 0.71 · intermediate layer, random seeds, static layer, best layer, untrained, LLMs
  7. [7] § Results › Positional information, word rate, and static embeddings account for the majority of the neural predictivity of trained LLMs ↔ misc_code/word_count.ipynb, lines 20–31 · score 0.63 · neural variance explained, simple confounds, neural predictivity, mapping, majority, brains
  8. [8] § Methods › Statistical testing ↔ analyze_results/figures_code/figure2.py, lines 614–688 · score 0.54 · fROI, squared error, FDR, GPT2XL, electrode, stacks
  9. [9] § Methods › Language models ↔ analyze_results/figures_code/trained_figure/trained_LLMs_varpar.py, lines 61–119 · score 0.52 · RoBERTa Large, Pile, gpt2xl, RWKV, Instruct, Llama

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 138 lines · 43 KB · MIT · 4 matches

  1. # %%
  2. import re
  3. def count_words_skip_citations_and_punctuations(text):
  4. # Remove \citet and \citep citations from the text
  5. text_without_citations = re.sub(r'\\cite[t|p]*\{[^}]*\}', '', text)
  6. # Remove all punctuation using regex
  7. text_without_punctuation = re.sub(r'[^\w\s]', '', text_without_citations)
  8. # Split the remaining text into words and count them
  9. words = text_without_punctuation.split()
  10. print(len(words))
  11. return len(words)
  12. # Example usage
  13. text = "This is a sentence with a citation \citet{author2025} and another \citep{author2025}."
  14. word_count = count_words_skip_citations_and_punctuations(text)
  15. print(word_count)
  16. # %%
  17. introduction = '''Vision neuroscience experienced a transformative shift with the introduction of deep learning models capable of human-level object recognition. Beginning with the influential work of \citet{Yamins2014-fz}, several studies demonstrated that the internal activations of such models were highly predictive of brain responses to visual stimuli \citet{Guclu2015-nz, Seeliger2018-uh, Eickenberg2017-dm}. This breakthrough, coupled with the open availability of model weights and neural datasets, gave rise to a comparative analysis framework known as "Brain-Score" \citet{Schrimpf2018-ah}. Brain-Score sought to rank deep learning models based on their neural and behavioral predictivity to identify “brain-like” models, with the goal of uncovering which model characteristics best correlated with neural and behavioral predictivity to provide insights into brain function. While this paradigm initially generated significant excitement, a growing body of research has since highlighted critical concerns, including behavioral divergences between models and brains \citet{Bowers2022-dl}, confounds inherent in the stimuli, and the surprising fragility of model performance rankings to the choice of similarity metric \citet{Soni2024-yz}.
  18. More recently, a revolution driven by deep learning has taken hold in language neuroscience. While early studies demonstrated that recurrent neural network (RNN) language models (LMs) could effectively predict brain responses \citet{wehbe-etal-2014-aligning, Jain2018-rg}, the advent of transformer large language models (LLMs) with remarkable behavioral capacities \citet{Radford2019-gr} sparked a surge in interest and research activity. This surge was fueled by pioneering studies showing that the internal activations of transformer LMs were highly effective at predicting brain responses, with their neural predictivity attributed to various factors, including their contextual representations, training objectives, and architectural properties \citet{Toneva2019-xy, Schrimpf2021-pg, Goldstein2022-ey, Caucheteux2022-dt}.
  19. The most cited among these studies to date, \citet{Schrimpf2021-pg}, exemplified the adoption of the "Brain-Score" framework in language neuroscience. The authors evaluated the neural predictivity of 43 models---including LLMs, smaller RNN LMs, and static word embeddings---across three neural datasets. Based on this comprehensive analyses, \citet{Schrimpf2021-pg} reported three main results. First, LLMs trained to predict the next word, specifically \textit{GPT2XL}, provided the best neural predictivity across all $43$ models. Second, the neural predictivity of models was positively correlated with one property of the models in particular, their next word prediction ability. These two findings were interpreted as "computationally explicit evidence that predictive processing fundamentally shapes the language comprehension mechanisms in the human brain" \citet{Schrimpf2021-pg}. Third, untrained (i.e. randomly initialized) LLMs demonstrated surprisingly high neural predictivity relative to their trained counterparts, which was interpreted as evidence that the transformer architecture may play role in biasing computations to be more brain-like. Taken together, these findings painted a picture in which the artifical intelligence community was "rapidly converging on architectures that might capture key aspects of language processing in the human mind and brain".
  20. \citet{Schrimpf2021-pg} laid the groundwork for several follow up studies which viewed LLMs not only as predictive tools, but also as candidate explanatory models of biological language processing. For example, \citet{Hosseini2024-kg} demonstrated that LLMs achieve high neural predictivity even when the scale of natural language data they are trained on mirrors the developmental language exposure of humans. \citet{Aw2024-Cl} highlighted the effects of instruction tuning on LLMs, revealing that this process enhances both their alignment with neural data and their integration of world knowledge. \citet{AlKhamissi2024-qr} further investigated the neural predictivity of untrained transformer-based LLMs, attributing it to tokenization strategy and multi-headed attention. The authors then constructed a model based on these two components that displayed high neural and behavioral predictivity, leading them to "conceptualize language processing in the human brain as an untrained feature encoder providing representations to a downstream trainable decoder that produces language output".
  21. Following the critique of deep learning models in vision, we present important evidence questioning results from \citet{Schrimpf2021-pg} and several of its follow up studies. Specifically, we find that LLM-to-brain mappings in this study were evaluated with incorrect or biased methodological choices, and simple confounds were not properly accounted for. First, we show that when using shuffled test splits, as done in \citet{Schrimpf2021-pg} as well as several other LLM-to-brain mapping studies \citet{Oota2022-bc, Aw2024-Cl, Hosseini2024-kg, Kauf2024-rh, Hosseini2024-zd, Mischler2024}, a trivial model that encodes temporal autocorrelation outperforms \textit{GPT2XL} on these three neural datasets. When switching to contiguous test splits, we find that the main results in \citet{Schrimpf2021-pg} are not robust. Specifically, their results on trained models either do not replicate or are dependent on sub-par methods to extract activations from models that are biased towards certain classes of models (specifically \textit{GPT2}). Furthermore, simple confounds, namely positional information and word rate, match or outperform the neural predictivity of \textit{GPT2XL} on two of the three datasets, and match or outperform the neural predictivity of untrained \textit{GPT2XL} on all three neural datasets. Finally, a combination of simple confounds and non-contextual word embeddings accounted for the majority of the neural variance explained by trained \textit{GPT2XL} on all three datasets, and the simple confounds alone accounted for all of the neural variance explained by untrained \textit{GPT2XL} across all three datasets. Together, question prior conclusions made on these datasets that LLMs capture key aspects of biological language processing.'''
  22. intro_wc =count_words_skip_citations_and_punctuations(introduction)
  23. # %%
  24. result_sec_1 = '''We evaluated the neural predictivity of several models on three neural datasets: (1) \textit{Pereira2018} (fMRI - passages) \cite{Pereira2018-ry}, (2) \textit{Fedorenko2016} \cite{Fedorenko2016-eg} (ECoG - sentences), and (3) \textit{Blank2014}, \cite{Blank2014-iz} (fMRI - stories). In \textit{Pereira2018}, $n=10$ participants read short passages presented one sentence at a time, and a single fMRI volume (TR) was acquired after presentation of each sentence. In \textit{Fedorenko2016}, ECoG recordings were made while $n=5$ participants read $52$ sentences. In \textit{Blank2014}, fMRI signals were acquired while $n=5$ participants listened to $8$ stories. Following \citet{Schrimpf2021-pg}, we place the most weight on \textit{Pereira2018}, and focus all analyses on language-selective voxels (\textit{Pereira2018}) / electrodes (\textit{Fedorenko2016}) / fROIs (\textit{Blank2014}). For additional details on these three neural datasets, see \ref{sec:neural_data}.
  25. We focused the majority of our results on \textit{GPT2XL} because \textit{GPT2XL} achieved the best neural predictivity out of $43$ models in \citet{Schrimpf2021-pg}, it was the main model used in \citet{Kauf2024-rh}, and \citet{Hosseini2024-kg} also focused their analyses on \textit{GPT2} style models. Furthermore, many language encoding studies which use other neural datasets also focus their analyses on \textit{GPT2} style models \citep{Caucheteux2021-nv, Caucheteux2023-yf, Goldstein2022-ey}. When evaluating the neural predictivty of a model, we first passed as input to the model the same natural language stimuli that the participants received. We then extracted model activations (see \ref{sec:llm feature pooling} for more details on the extraction procedure), and trained linear regressions to predict brain responses from each layer of the model.
  26. To assess a model's neural predictivity, we used nested cross-validation. In each outer fold, the subset of brain responses and model activations used for fitting the regression is termed the train set, and the held-out subset of brain responses and model activations used to evaluate the regression is termed the test set. For models with multiple layers, we focus our analyses on the layer which achieves the highest neural predictivity on the test set (averaged across outer folds). (\ref{sec:select_best_layer}).'''
  27. r1_wc = count_words_skip_citations_and_punctuations(result_sec_1)
  28. result_sec_2 = '''predictivity than \textit{GPT2XL} when using shuffled test splits}
  29. The majority of studies evaluating the neural predictivity of language models have employed contiguous test splits \citep{Huth2016-tp, Antonello2023-ab, Antonello2024-co, Caucheteux2021-nv, Caucheteux2022-dt, Caucheteux2023-yf, goldstein2024the, Goldstein2022-ey} [MAYBE ADD MORE CITES HERE], where a temporally contiguous chunk of brain responses (and the associated model activations) are held-out for testing. By contrast, several studies using these neural datasets \citep{AlKhamissi2024-qr, Aw2024-Cl, Kauf2024-rh, Oota2022-bc, Hosseini2024-kg, Hosseini2024-zd}, most notably \citet{Schrimpf2021-pg}, used shuffled test splits, where brain responses were arbitrarily placed into the test set. To provide an example with the \textit{Pereira2018} dataset, contiguous test splits mean that brain responses for entire passages are held out for testing, whereas shuffled test splits mean that it is possible for some brain responses from a given passage to be in the train set, and other brain responses \textit{from the same passage} to be in the test set.
  30. Shuffled test splits are known to be problematic in the field of language neural encoding because brain responses are temporally autocorrelated for reasons that are not entirely stimulus-driven \citep{Zada2023-rx}. Due to temporal autocorrelation, a model can achieve high neural predictivity when using shuffled test splits simply because it represents nearby time points similarly, rather than because it extracts features from the natural language stimuli that are present (or correlated) with those in brain responses. To demonstrate this, we constructed a trivial model which operates only on the principle that stimuli nearby in time should be assigned similar representations. We term this model the Orthogonal Autocorrelated Sequences Model (\textit{OASM}) because the activations of \textit{OASM} are orthogonal for distinct passages/sentences/stories, and autocorrelated for stimuli within a passage/sentence/story (see \ref{sec:OASM} for additional details on its construction).
  31. We first compared the neural predictivity of \textit{OASM} to that of \textit{GPT2XL} when using largely the same methodological choices implemented in \citet{Schrimpf2021-pg}. These methodological choices were fitting the mapping between model activations and brain responses with ordinary least squares (OLS) regression (\ref{sec:regression}) and extracting activations from \textit{GPT2XL} using the last token (word) of the stimuli, which we term \textit{GPT2XL-LT} (\ref{sec:llm feature pooling}). Under these settings, \textit{OASM} predicts brain responses significantly better than \textit{GPT2XL-LT} on \textit{Pereira2018} and \textit{Blank2014} across participants (paired t-test, $\alpha=0.05$), and is on-par with \textit{GPT2XL-LT} on \textit{Fedorenko2016} (Figure \ref{fig:schrimpfcomp}a, top left corner). To evaluate the robustness of this finding, we compared the performance of \textit{OASM} to \textit{GPT2XL} when using two other activation extraction methods for \textit{GPT2XL}: taking the mean across tokens (mean pooling, referred to as \textit{GPT2XL-MP}), and taking the sum across tokens (sum pooling, referred to as \textit{GPT2XL-SP}). We included mean pooling because it was used in \citet{Kauf2024-rh}, and sum pooling because it is somewhat analogous to the process of convolving the hemodynamic response function (HRF) with model activations (\ref{sec:llm feature pooling}). \textit{OASM} achieved significantly better neural predictivity than \textit{GPT2XL-MP} and \textit{GPT2XL-SP} on \textit{Pereira2018} and \textit{Blank2014}, and performed on par with these two models on \textit{Fedorenko2016}.
  32. We hypothesized that the gap between \textit{GPT2XL} and \textit{OASM} was partially due to the use of OLS regression, which does not impose a regularization term. This is because regularization tends to help higher dimensional predictors, and a given layer of \textit{GPT2XL} is higher dimensional than \textit{OASM}. We found that while using L2-regularized regression reduced the gap between \textit{OASM} and \textit{GPT2XL} on average, \textit{OASM} still significantly outperformed all variants of \textit{GPT2XL} on \textit{Pereira2018} and \textit{Blank2014}, and performed on par with all variants of \textit{GPT2XL} on \textit{Fedorenko2016} (Figure \ref{fig:schrimpfcomp}a, top right corner).
  33. We find that the neural predictivity of \textit{OASM} is at or very close to $0$ across all methodological choices with contiguous splits, which should be the case given that \textit{OASM} represents distinct passages/sentences/stories orthogonally. We also find that L2-regularized regression leads to higher performance for all \textit{GPT2XL} activation extraction variants when using contiguous test splits as well. For this reason, and because L2-regularized regression is the standard in the field, \citep{Huth2016-tp, Antonello2023-ab, Antonello2024-co, Caucheteux2021-nv, Caucheteux2022-dt, Caucheteux2023-yf, goldstein2024the, Goldstein2022-ey}, we used L2-regularized regression for the remainder of our analyses. '''
  34. r2_wc = count_words_skip_citations_and_punctuations(result_sec_2)
  35. result_sec_3 = '''While \textit{OASM} explains similar or more neural variance than \textit{GPT2XL} when using shuffled test splits, it remains unclear how much \textit{OASM} accounts for the neural variance \textit{GPT2XL} explains. In other words, to what extent is the neural predictivity of \textit{GPT2XL} on shuffled test splits attributable to the fact that it simply represents nearby timepoints similarly?
  36. We addressed this question in three ways. First, we examined whether the pattern in neural predictivity of \textit{OASM} and \textit{GPT2XL} was positively correlated across voxels/electrodes/fROIs. Second, we more directly addressed this question by performing a variance partitioning style analysis at the participant-level using $R^2$. In brief, this variance partitioning analysis involved quantifying the neural predictivity of a model which combines activations from both \textit{OASM} and \textit{GPT2XL} (\textit{OASM+GPT2XL}). Given the neural predictivity of \textit{OASM+GPT2XL}, we computed the percentage of neural variance that \textit{GPT2XL} explains that is also explained by \textit{OASM}, termed $\Omega\textsubscript{\textit{GPT2XL}}(OASM)$, using the formula outlined in \ref{sec:frac_var_llm}. Finally, because $\Omega$ is defined at the participant level, we quantified the percentage of voxels/electrodes/fROIs where \textit{OASM+GPT2XL} explained significantly more neural variance than \textit{OASM} alone \citet{Benjamini1995-kn} (see \ref{sec:stats} for details on statistical testing). When quantifying this percentage, we only considered the subset of voxels/electrodes/fROIs in which \textit{GPT2XL} predicted brain responses significantly better than chance.
  37. Across all three datasets, neural predictivity across voxels/electrodes/fROIs was correlated between \textit{OASM} and \textit{GPT2XL} (Figure \ref{fig:schrimpfcomp}c, d). Furthermore, \textit{OASM} accounted for over $80\%$ of the neural variance that \textit{GPT2XL} explains in \textit{Pereira2018}, over $50\%$ in \textit{Fedorenko2016}, and nearly $100\%$ in \textit{Blank2014} (Figure \ref{fig:schrimpfcomp}e). When considering only the subset of voxels/electrodes/fROIs which \textit{GPT2XL} predicted better than chance, \textit{GPT2XL} explained signficant neural variance over \textit{OASM} in around $20\%$ of voxels, $50\%$ of electrodes, and $0$ fROIs. In all, these results show that the majority of the neural predictivty of \textit{GPT2XL} is confounded with a model that simply represents nearby time points similarly/
  38. '''
  39. r3_wc = count_words_skip_citations_and_punctuations(result_sec_3)
  40. result_sec_4 = '''In the previous analyses, we presented evidence showing that shuffled test splits are not a reliable method to assess neural predictivity. It is therefore important to assess to what extent can switching from shuffled test splits to contiguous test splits impact LLM to brain mapping findings. To begin answering this question, we examined the pattern in neural predictivity across the layers of \textit{GPT2XL} when using shuffled and contiguous test splits. The pattern in across layer predictivity is important because studies which employed shuffled test splits using these datasets exclusively focus their analyses on either the LLM layer which achieves the highest neural predictivity, \citet{Schrimpf2021-pg, Hosseini2024-kg, Kauf2024-rh, Aw2024-Cl}, or the last LLM layer \citet{Oota2022-bc}.
  41. We find that on \textit{Pereira2018}, the across layer performance is highly anti-correlated between shuffled and contiguous splits across all activation extraction variants of \textit{GPT2XL} (Figure \ref{fig:schrimpfcomp}f).
  42. When using shuffled splits, the early and late layers of \textit{GPT2XL} achieve the highest neural predictivity when evaluating. These across layer values are consistent with the code accompanying \citet{Schrimpf2021-pg} and with \citet{Kauf2024-rh}; they are not consistent the across layer pattern displayed in Figure $2$c of \citet{Schrimpf2021-pg} for reasons we are unaware of. When using contiguous test splits intermediate layers achieve the best neural predictivity. Given that we selected the best layer of \textit{GPT2XL} for each of the two experiments in \textit{Pereira2018} separately, we also show that these same trends hold when evaluating across layer patterns within each experiment separately (Supp). On \textit{Fedorenko2016}, the across layer trends are more similar across shuffled and contiguous test splits, with later intermediate layers generally performing the best on shuffled splits and early intermediate layers performing the best on contiguous splits ((Figure \ref{fig:schrimpfcomp}e). Finally, on \textit{Blank2014}, the across layer performance is also highly anti-correlated between shuffled and contiguous splits; later intermediate layers performed the best when using shuffled splits, whereas the early layers performed the best when using contiguous splits ((Figure \ref{fig:schrimpfcomp}e). We thus show that on $2$ out of the $3$ neural datasets, the pattern in across layer neural predictivity flips between shuffled and contiguous test splits, providing a clear example of how shuffled test splits can impact findings on LLM to brain mapping studies. '''
  43. r4_wc = count_words_skip_citations_and_punctuations(result_sec_4)
  44. result_sec_5 = '''Building on our observation that shuffled test splits can alter across-layer predictivity patterns, we sought to replicate the key model comparison analyses of \citet{Schrimpf2021-pg} using contiguous test splits and L2-regularized regression.
  45. We evaluated 28 of the 29 bidirectional transformer models, 6 of the 9 unidirectional transformer models, and 2 of the 3 word embedding models from \citet{Schrimpf2021-pg}, alongside 4 recent unidirectional transformer models (Llama-3.2 family) and 4 unidirectional RNNs (RWKV-4 family). These models were compared using the best activation extraction method for each model. Previously, \citet{Schrimpf2021-pg} reported that auto-regressive transformer models (e.g. GPT2 family predicted neural responses uniquely well, that transformers outperformed recurrent models, and that contextual models far surpassed static word embeddings. However, these findings relied on shuffled splits and last-token activation extraction, which may have introduced biases.
  46. Our results challenge these conclusions. Using contiguous splits and optimal activation extraction methods, we find no evidence that auto-regressive transformer models (e.g., GPT-2, Llama) outperform bidirectional transformers in predicting neural responses. Similarly, recurrent models achieve comparable performance to transformers (Fig. 3a). The only consistent trend across datasets is that static word embedding models (e.g., GloVe, word2vec) perform markedly worse than contextual models.
  47. We further introduce a simple baseline, the Position and Word Rate (PWR) model, which encodes only positional and word rate information. While PWR matches static models in the \textit{Pereira2018} dataset, it rivals contextual models in the \textit{Fedorenko2016} dataset and outperforms them in the \textit{Blank2014} dataset, revealing the potential impact of simple confounds in these datasets.
  48. To examine how activation extraction methods influence model comparisons, we tested multiple methods and found that trends were generally robust. However, the last-token method biased results in the \textit{Pereira2018} dataset, significantly disadvantaging bidirectional models (Welch’s t-test, $p=4.66 \times 10^{-7}$), though this was also the lowest-performing activation extraction method for every model.
  49. When reintroducing shuffled splits, unidirectional models (RWKV-4, GPT-2, Llama-3.2) consistently outperformed bidirectional models across datasets and activation extraction methods. This suggests that the use of shuffled train-test splits likely drove the apparent superiority of unidirectional models in \citet{Schrimpf2021-pg}.
  50. We next examined the consistency of neural predictivity trends across datasets by calculating Pearson correlations between model performances (Pearson $r$). Using contiguous splits and the best activation extraction method for each model, significant correlations emerged when all models were included, but vanished when static word embedding models were excluded.
  51. This pattern persisted when using individual activation extraction methods. Significant correlations were always found when including all models, but when excluding static models, only the mean-pooling method between \textit{Pereira2018} and \textit{Blank2014} yielded a significant result.
  52. Under shuffled splits, significant correlations appeared universally across datasets and extraction methods, regardless of whether static models were included. These findings indicate that shuffled splits in \citet{Schrimpf2021-pg} likely exaggerated the perceived consistency of model comparisons across datasets.
  53. Finally, we reassessed the most impactful result from \citet{Schrimpf2021-pg}: the correlation between next-word prediction (perplexity) and neural predictivity. Using contiguous splits and optimal activation extraction methods, we found significant correlations across datasets only when static word embedding models were included. When restricted to contextual models, these correlations disappeared (Fig. 5a).
  54. This result holds across most activation extraction methods, except in the \textit{Pereira2018} dataset with the last-token method, where a significant correlation persisted. Under shuffled splits, correlations were more widespread but only consistent across datasets and extraction methods for the \textit{Fedorenko2016} dataset and the mean-pooling method. Overall, our findings suggest that the correlation between neural predictivity and next-word prediction is less robust than previously reported, being sensitive to methodological choices.
  55. Given the influential role that this correlation between next word prediction and neural predictivity has taken in the literature, we further investigate what methodological deviations from \citet{Schrimpf2021-pg} might be most responsible for our null findings. First, we note that the choice to exclude 7 of the 43 models would not have affected the finding of significant correlation had we used the same neural predictivity values reported by \citet{Schrimpf2021-pg}. In fact, when using the same neural predictivity values reported by \citet{Schrimpf2021-pg} the model exclusions in this work actually increase the reported Pearson correlations compared to what was reported by \citet{Schrimpf2021-pg} (SUPP FIG X). Hence, it is not our exclusion of certain models that is responsible for this difference in results.
  56. Our analyses from the model comparison section reveal that non-contextual models and \textit{PWR} perform surprisingly well relative to contextual models. Motivated by these findings, we sought to answer two key questions. First, would combining the activations of a non-contextual model and \textit{PWR} close the gap in neural predictivity with contextual models? Second, to what extent do non-contextual models and/or \textit{PWR} account for the same neural variance explained by contextual models? To address these questions, we focus on \textit{GloVe} and \textit{GPT2XL} as representative examples of non-contextual and contextual models, respectively. We replicate our analyses with \textit{RoBERTa-large} in Extended Data Figure X.
  57. On \textit{Pereira2018}, the neural predictivity of \textit{PWR+GloVe} was higher than that of \textit{PWR} or \textit{GloVe} alone (Figure\ref{fig:trained_var_par}a). Combining \textit{PWR} with \textit{GloVe} did not yield significant increases in neural predictivity over either model alone in \textit{Fedorenko2016} and \textit{Blank2014}, and so we elected to only use \textit{PWR} to address the second question in these two datasets.
  58. We next sought to examine how much of the neural variance of \textit{GPT2XL} that \textit{PWR+GloVe} accounts for in \textit{Pereira2018}, and \textit{PWR} alone explains in \textit{Fedorenko2016} and \textit{Blank2014}. The neural predictivity between \textit{PWR+GloVe / PWR} and \textit{GPT2XL} across voxels/electrodes/fROIs was strongly correlated (Figure \ref{fig:trained_var_par}b, c). Furthermore, \textit{PWR+GloVe} accounted for over $85\%$ of the neural variance explained by \textit{GPT2XL} in \textit{Pereira2018}, and \textit{PWR} accounted for over $80\%$ and nearly $100\%$ of the neural variance \textit{GPT2XL} explained in \textit{Fedorenko2016} and \textit{Blank2014}, respectively. Finally, when considering only the subset of voxels/electrodes/fROIs which \textit{GPT2XL} significantly explained, \textit{GPT2XL} explained significant neural variance over \textit{PWR+GloVe} in around $10\%$ of voxels, and over \textit{PWR} in around $5\%$ of electrodes and $3\%$ of fROIs.
  59. Given that the \textit{PWR+GloVe} model explained some neural variance over \textit{PWR} alone on \textit{Pereria2018}, we also explained two other models for this dataset: sense-specific word embeddings (\textit{SENSE}) and contextual syntactic representations (\textit{SYNT)}. These models are more complex than \textit{GloVe}, but are less complex than \textit{GPT2XL}. We found that \textit{SENSE} and \textit{SYNTAX} did not explain significant neural variance over \textit{PWR+GloVe} (Extended Data Figure X).
  60. In summary, we find that a combination of positional information, word rate, and non-contextual embeddings predicts almost as much or more neural variance than \textit{GPT2XL}. Furthermore, these models account for upwards of $80\%$ of the neural variance that \textit{GPT2XL} explains across all three datasets. This suggests that even when using contiguous test splits, there are relatively simple explanations underlying the high neural predictivity of contextual models on these three datasets.
  61. Having accounted for the majority of the neural variance for trained \textit{GPT2XL}, we finally turned to untrained \textit{GPT2XL} (\textit{GPT2XLU)}. The neural predictivity of untrained transformers was used to argue that the transformer architecture biases computations to be more brain-like \citet{Schrimpf2021-pg}, and inspired the development of a novel architecture that reportedly achieved state-of-the-art neural and behavioral alignment across several datasets \citet{AlKhamissi2024-qr}. We hypothesized that when evaluated on contiguous test splits, the neural predictivity of \textit{GPT2XLU} could be fully accounted for by \textit{PWR}.
  62. We found that \textit{PWR} performed on par or significantly better than all three variants of \textit{GPT2XLU} across all three neural datasets (Figure \ref{fig:untrained_var_par}a). Furthermore, the neural predictivity of \textit{PWR} and \textit{GPT2XLU} was strongly correlated across voxels/electrodes/fROIs (Figure \ref{fig:untrained_var_par}b,c). \textit{PWR} accounted for essentially all ($> 98\%$) of the neural variance of \textit{GPT2XLU} across all three datasets, and there were no voxels/electrodes/fROIs where \textit{GPT2XLU} explained significant neural variance over \textit{PWR}. We therefore find that a combination of positional and word rate information not only explains equal or more neural variance than \textit{GPT2XLU}, but these simple features also explain nearly all of the neural variance that \textit{GPT2XLU} explains.'''
  63. r5_wc = count_words_skip_citations_and_punctuations(result_sec_5)
  64. discussion = '''
  65. Beyond \citet{Schrimpf2021-pg}, our study questions results on several prior studies which performed LLM-to-brain mappings on these datasets using shuffled tests splits (\citet{AlKhamissi2024-qr, Kauf2024-rh, Oota2022-bc, Hosseini2024-kg, Hosseini2024-zd, Aw2024-Cl}). Among the studies that have use shuffled test splits, some have note that autoregressive transformers can achieve high neural predictivity when trained on little or even no training data \citet{Hosseini2024-kg, AlKhamissi2024-qr}. \citet{Hosseini2024-kg} responded to the common criticism that LLMs are not viable models of human language processing because they are trained on massive corpora of text by showing that LLMs, specifically \textit{GPT2} style models, can achieve high neural predictivity on \textit{Pereira2018} when trained on developmentally realistic amounts of data. Importantly, when using shuffled test splits, we find that a model that is trained on no text and which only knows passage/sentence/story boundaries achieves comparable or higher neural predictivity than fully trained \textit{GPT2XL}, suggesting that such results are not indicative of a model's developmental plausibility. \citet{AlKhamissi2024-qr} attributed the neural predictivity of untrained, auto-regressive LLMs to two key components: multi-headed attention and byte-pair encoding. By contrast, we show that when using contiguous test splits position and word rate fully account for the neural predictivity of \textit{GPT2XLU}.
  66. Another major question that previous studies have attempted to address using \textit{Pereira2018} with shuffled test splits is whether semantic or syntactic information is the primary driver of the LLM-to-brain mapping \citet{Kauf2024-rh, Oota2022-bc}. \citet{Kauf2024-rh} showed through a series of linguistic perturbations that lexical-semantic information, not syntactic structure, accounted for the majority of the mapping between \textit{GPT2XL} and brain responses. \citet{Oota2022-bc} showed that fine-tuning \textit{BERT} on syntax-related natural language processing tasks (NLP) led to greater improvements in neural predictivity relative to other NLP tasks. We evaluated this question with contiguous test splits by using a lexico-semantic model, \textit{GloVe}, and a syntactic model based on \citet{Caucheteux2021-nv}. We found that combining \textit{GloVe} with position and word rate was important for accounting for the neural variance that \textit{GPT2XL} explains. By contrast, the syntactic model did not explain neural variance over position and word rate. Therefore, our analyses suggest that on this neural dataset, lexico-semantic information can explain LLM-to-brain mappings beyond simple confounds, whereas the role of syntax cannot be decoupled from simpler explanations.
  67. Other previous studies using shuffled test splits have observed correlations between neural predictivity and certain model performance metrics. For instance, \citet{Aw2024-Cl} reported that instruction-tuning large language models (LLMs) enhances their neural predictivity on \textit{Pereira2018} and \textit{Blank2014}, suggesting that this improvement stems from instruction-tuning increasing the models’ world knowledge. \citet{Mischler2024} reported that as LLMs achieve better performance on semantic benchmark tasks, their neural predictivity increases on an intracranial EEG dataset. Crucially, given that we demonstrated that the relationship between neural predictivity and other model performance metrics (e.g. next-word prediction) can change substantially when shifting from shuffled to contiguous test splits and with different activation extraction methods, we recommend evaluating whether these relationships are robust.
  68. \citet{Hosseini2024-zd} noted that although \textit{GPT2XL} performed particularly well on \textit{Pereira2018}, most contextual models also performed fairly well relative to a "noise ceiling” estimated via inter-subject predictivity. This led them to hypothesize that “universal representations” — those consistent across models — underlie the mapping between large language models (LLMs) and brain activity. To test this hypothesis, the authors constructed a new neural dataset specifically comprising "low agreement" sentences, whose representations varied significantly across models. They found that the ratio between LLM neural predictivity and the noise ceiling was much lower on this dataset than on \textit{Pereira2018}, which they interpreted in support of their hypothesis. However, this discrepancy between the two datasets can be attributed to a simpler explanation. In \textit{Pereira2018}, the noise ceiling is inflated in the same manner as LLM neural predictivity due to the “OASM-like” structure of brain responses, where brain responses are more similar within a passage than across distinct passages. The “low agreement” dataset, on the other hand, consisted of isolated, unrelated sentences presented one at a time in a consistent order across participants in several runs. Because previous sentences were not provided as context and temporally adjacent sentences were not similar, model activations no longer exhibited "OASM-like" properties like they did in \textit{Pereira2018}, reducing the artificial inflation of neural predictivity scores. However, since brain responses remained temporally autocorrelated due to the fixed presentation order, the noise ceiling remained inflated when using shuffled test splits. Therefore, the observed differences in neural predictivity relative to the noise ceiling between \textit{Pereira2018} and the “low agreement” dataset are a consequence of shuffled test splits and dataset design choices. Our results with contiguous test splits do bolster the perspective from Hosseini that the neural predictivity of contextual models are similar on these datasets, and our variance partitioning analyses on a representative unidirectional model (\textit{GPT2XL}) and a representative bidirectional model (\textit{RoBERTA-large}) indicate that these models both are largely predicting the same neural variance.
  69. %Of the aforementioned studies that use shuffled train-test splits, the only one that explicitly states this is \citet{Kauf2024-rh}. The other seven studies do not describe how their train-test splits were constructed beyond stating how many folds were used for cross-validation, and the use of shuffled test splits was only evident through their associated codebases. determining whether shuffling had occurred required us to check their code
  70. Although the majority of studies in the field use contiguous splits, the potential impact of shuffled test splits on prior influential findings has not been extensively examined. Two factors may contribute to this. First, the use of shuffled train-test splits is often not explicitly reported. Among the eight studies we identified as employing this approach, only one—Kauf et al. 2024—clearly stated this choice; the others generally described the number of cross-validation folds without detailing their construction. Second, the methodological implications of shuffled train-test splits have been underexplored. For instance, Kauf et al. 2024 acknowledges that shuffled splits can “inflate” neural predictivity compared to contiguous splits but utilizes shuffled test splits exclusively for their main analyses. Our results highlight that this inflation varies across models, resulting in markedly different patterns of across-layer and across-model performance.
  71. Kauf et al. 2024 further discusses both merits and limitations of shuffled test splits, citing increased semantic coverage in the training set as a benefit and acknowledging temporal autocorrelation as a drawback. However, their exclusive use of shuffled test splits for their main analyses, without providing complementary analyses using contiguous splits, may inadvertently signal to the field that the advantages outweigh the drawbacks. Our findings suggest that the issue of temporal autocorrelation is substantial, as the majority of neural variance explained by GPT2-XL is confounded with temporal structure. Additionally, the purported benefit of increased semantic coverage appears limited in the Pereira2018 dataset, which includes multiple passages even for specific topics like beekeeping. By designing our test splits to leverage this feature, we demonstrate that contiguous splits can effectively balance semantic diversity and methodological rigor.
  72. In addition to avoiding shuffled train-test splits, our findings underscore the importance of systematically investigating multiple activation extraction methods in fMRI datasets. In Pereira2018, we observe that the most commonly used approach—last-token extraction—generally performs the worst, particularly disadvantaging bidirectional transformers. Beyond model comparison analyses, Jain et al. (2022) demonstrated that the standard activation extraction method for narrative comprehension datasets (analogous to sum-pooling) can lead to artifacts where model dimensions with the longest timescales are mapped to brain regions with the shortest timescales, such as primary auditory cortex. This occurs because the regression effectively repurposes these long-timescale representations to predict word-rate-related responses. More broadly, this highlights a fundamental concern when interpreting LLM-to-brain mappings: the associations between model representations and brain regions are heavily influenced by the assumptions inherent in the chosen activation extraction method.
  73. Although comparisons of model performances without including simple confounds in the regression are common, several other studies have accounted for simple confounds. Caucheteux et al. 2021 control for word rate as well as phonological features before examining how much of the encoding performance of LLMs can be explained by syntax and semantics. Reddy et al. 2021 iteratively accounts for simpler features, such as punctuation, before analyzing the neural predictivity of more complex syntactic models and LLM models. LeBel et al. 2021 and DeHeer et al. 2017 also iteratively account for the neural predictivity of simpler acoustic and phonetic features before introducing more complex, semantic models. We encourage the continued use of these low-level confounding predictors, especially when authors seek to draw scientific inferences about higher-level representations from neural predictivity differences between models. 
  74. The consideration of positional confounds specifically has been relatively limited. Antonello et al. (2023) was the first LLM encoding study to highlight this issue, identifying positional signals in the initial 100 seconds of story stimuli. To mitigate the influence of these signals, they excluded the first 100 seconds of the held-out story used for testing. Similarly, we observe significant positional effects at the beginnings of stories in the Blank2014 dataset, and the neural variance predicted by GPT2XL is largely confounded with these signals. This suggests that positional confounds may be pervasive across narrative comprehension datasets. Such confounds pose particular challenges for analyses examining how the length of prior context provided to an LLM affects its neural predictivity, as these analyses are often interpreted as reflecting the timescale of language integration in the brain. For example, in Pereira2018, we find that positional confounds substantially account for the difference in neural predictivity between contextualized and decontextualized LLM representations, indicating that much of this variance stems from positional signals rather than passage-level contextual integration. 
  75. We find that the correlation between neural predictivity and next-word prediction can depend on choice of datasets, models, and feature extraction methods. Results in the literature are similarly mixed. For instance, Caucheteux et al. (2022) trained 36 transformer models from scratch and found that although trained models exhibited higher neural predictivity than untrained models, neural predictivity ceased scaling with next-word prediction relatively early in training. Similarly, Pasquiou et al. (2022) reported a correlation between perplexity and neural predictivity within, but not across, model classes. On the other hand, Antonello et al. 2023, Hong et al. 2024 and Bonasse-Gahot et al. 2024 did report significant correlations between neural predictivity and next word prediction. 
  76. In light of these complexities, large-scale analyses like ours, which incorporate diverse datasets, models, and methodologies, are essential for evaluating whether any specific model attribute is reliably associated with neural predictivity. Such analyses have proven similarly insightful in the visual domain, where they have likewise highlighted methodological fragility (Soni et al., 2024) and revealed a striking similarity in the neural predictivity of deep learning models regardless of their training objective or architecture (Conwell et al., 2024).
  77. '''
  78. d_wc = count_words_skip_citations_and_punctuations(discussion)
  79. r1_wc + r2_wc + r3_wc + r4_wc + r5_wc + d_wc
  80. # %%

word_count.ipynb at commit 046135d, under MIT · at the source

Overview

Authors: Nima Hadidi1,2, Ebrahim Feghhi1,2, Bryan H Song3, Idan A Blank2,4,5, Jonathan C Kao1,2,3,6
  1. Department of Electrical and Computer Engineering, University of California, Los Angeles, Los Angeles, CA USA
  2. Neuroscience Interdepartmental Program, University of California, Los Angeles, Los Angeles, CA USA
  3. Department of Computer Science, University of California, Los Angeles, Los Angeles, CA USA
  4. Department of Linguistics, University of California, Los Angeles, Los Angeles, CA USA
  5. Department of Psychology, University of California, Los Angeles, Los Angeles, CA USA
  6. Department of Neurobiology, University of California, Los Angeles, Los Angeles, CA USA
Institutions: University of California, Los Angeles (United States)
Journal: Nature communications, volume 17, issue 1, article 5769
Dates: received 14 April 2025; accepted 7 April 2026; published online 27 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-72253-7 · PMID 42045182 · PMCID PMC13324000 · OpenAlex W7156034438
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), methods / tools (subfield)
Methods: Preprocessing, Connectivity, Statistics, Machine learning, Spectral & time-frequency, fMRI & imaging
Keywords: Neural encoding, Language
MeSH: Brain*, Large Language Models*, Humans (* major topic)
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: National Science Foundation (1943467); Foundation for the National Institutes of Health (R01NS121097)
Citations: cited by 5 papers (Europe PMC); 47 references in the paper

Abstract

Emerging research seeks to draw neuroscientific insights from the neural predictivity of large language models (LLMs). However, as results rapidly proliferate, there is a growing need for large-scale assessments of their robustness. Here, we analyze a wide range of models and methodological approaches across three widely used neural datasets. We find that the use of shuffled train-test splits has contributed to findings that are influential but spurious. Furthermore, how activations are extracted from LLMs can bias results in favor of specific model classes. Lastly, we find that confounding variables, particularly positional signals and word rate, perform competitively with trained LLMs and fully account for the neural predictivity of untrained LLMs on these neural datasets. Although many studies in the field avoid these pitfalls, our results indicate that some apparent alignment between LLMs and brains has emerged from non-robust methods and overlooked confounds.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.

ebrahimfeghhi/beyond-brainscore

License: MIT
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 046135d414aa19285b5aa838a8730700fae582ea, 2 March 2026
Languages: Jupyter (46), Python (41), Shell (6)
Size: 955 files, 93 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (environment.yaml, requirements.txt, requirements_updated.txt), 46 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (78 files), Matplotlib (44 files), pandas (34 files), seaborn (28 files), SciPy (20 files), scikit-learn (15 files), Nilearn (14 files), PyTorch (7 files), NiBabel (6 files), Hugging Face Transformers (5 files), Plotly (2 files), h5py (1 file), xarray (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
95 files

Code availability

Code for this manuscript is available at the following link: Code (https://github.com/ebrahimfeghhi/beyond-brainscore/).

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 93 scripts, each with its path and the digest of its content;
  • 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

The neural data, as well as associated texts, for the three datasets have been deposited at the Figshare database at the following link: Figshare (https://figshare.com/s/42e9f147aa0471e12eea?file=58227502). Source data for the manuscript is provided at the following link: Source Data (https://github.com/ebrahimfeghhi/beyond-brainscore/tree/main/source_data).

Code for this manuscript is available at the following link: Code (https://github.com/ebrahimfeghhi/beyond-brainscore/).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 2 keywords, 3 MeSH terms, 2 funders, 25 references.

Cite

This paper

Hadidi, N., Feghhi, E., Song, B. H., Blank, I. A., & Kao, J. C. (2026). Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds. Nature communications, 17(1), 5769. https://doi.org/10.1038/s41467-026-72253-7

BibTeX

@article{hadidi2026spurious,
author = {Hadidi, Nima and Feghhi, Ebrahim and Song, Bryan H and Blank, Idan A and Kao, Jonathan C},
title = {{Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds}},
journal = {Nature communications},
year = {2026},
month = apr,
volume = {17},
number = {1},
pages = {5769},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-72253-7},
url = {https://doi.org/10.1038/s41467-026-72253-7},
pmid = {42045182},
pmcid = {PMC13324000}
}

RIS

TY - JOUR
AU - Hadidi, Nima
AU - Feghhi, Ebrahim
AU - Song, Bryan H
AU - Blank, Idan A
AU - Kao, Jonathan C
TI - Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/04/27
VL - 17
IS - 1
SP - 5769
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-72253-7
UR - https://doi.org/10.1038/s41467-026-72253-7
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-72253-7",
"type": "article-journal",
"title": "Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds",
"container-title": "Nature communications",
"author": [
{
"family": "Hadidi",
"given": "Nima"
},
{
"family": "Feghhi",
"given": "Ebrahim"
},
{
"family": "Song",
"given": "Bryan H"
},
{
"family": "Blank",
"given": "Idan A"
},
{
"family": "Kao",
"given": "Jonathan C"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "5769",
"DOI": "10.1038/s41467-026-72253-7",
"PMID": "42045182",
"PMCID": "PMC13324000",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-72253-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
27
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.7554/elife.106543 [code]
Stimulus dependencies-rather than next-word prediction-can explain pre-onset brain encoding in naturalistic listening designs.
Journal: eLife
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 5 references
[2] doi:10.1162/imag.a.1227 [code]
Large language models reveal the neural tracking of linguistic context in attended and unattended multi-talker speech.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Hugging Face Transformers, h5py, PyTorch, 6 other tools, 5 references
[3] doi:10.1038/s42003-026-10169-0 [code]
Shared representations in brains and models reveal a two-route cortical organization during scene perception.
Journal: Communications biology
In common: Hugging Face Transformers, Plotly, h5py, 8 other tools, 3 references
[4] doi:10.7554/elife.107933 [code]
Modality-agnostic decoding of vision and language from fMRI.
Journal: eLife
In common: Hugging Face Transformers, Nilearn, h5py, 8 other tools, 2 references
[5] doi:10.1016/j.isci.2026.117180 [code]
Developmental changes in similarity between neural representations of mental arithmetic and artificial neural networks.
Journal: iScience
In common: Hugging Face Transformers, Nilearn, h5py, 8 other tools, 2 references
[6] doi:10.1038/s41467-026-76098-y [code]
A single computational objective can produce specialization of streams in visual cortex.
Journal: Nature communications
In common: xarray, Hugging Face Transformers, h5py, 8 other tools, 1 reference
[7] doi:10.1038/s41597-026-07248-6 [code]
A large-scale fMRI dataset for vision-language semantic association.
Journal: Scientific data
In common: Nilearn, h5py, NiBabel, 6 other tools, methods / tools, 3 references
[8] doi:10.1038/s41586-026-10691-5 [code]
Mapping the neuronal building blocks of human language with language models.
Journal: Nature
In common: scikit-learn, pandas, SciPy, 2 other tools, 7 references
[9] doi:10.1371/journal.pone.0346575 [code]
Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.
Journal: PloS one
In common: xarray, Hugging Face Transformers, Plotly, 8 other tools, methods / tools
[10] doi:10.1162/imag.a.1283 [code]
The language network responds robustly to sentences across tasks.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Nilearn, h5py, seaborn, 5 other tools, 4 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.