OSCR

Mapping the neuronal building blocks of human language with language models.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches
  1. [1] § Combinatorial linguistic encoding ↔ analysis.zip/computation/figure/Figure3.ipynb, lines 1–25 · score 0.73 · network layers, word utterance, Swap sentence, peak predictivities, normalized predictivity, population predictivities
  2. [2] § Methods › Predictivities of neurons across language models › Neural predictivity controls ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm_swap_sentence.py, lines 1–43 · score 0.71 · length matched, swap sentence, neuronal population, predictivities, embeddings, models
  3. [3] § Methods › Predictivities of neurons across language models › Neural predictivity dynamics and mapping ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm_by_time_shuffle.py, lines 1–41 · score 0.69 · temporal window, neural activity, neuronal firing rates, shifted, post, dynamics
  4. [4] § Methods › Predictivities of neurons across language models › Language model predictivity ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_randtok.py, lines 1–38 · score 0.67 · neuronal firing rates, aligned neuronal, population predictivities, neuronal predictivities, model
  5. [5] § Linguistic representations by neurons ↔ analysis.zip/computation/code/B1_probe_decoding_single_word.py, lines 1–25 · score 0.64 · Decoding linguistic features, single word, linguistic representations, training, neural, activity
  6. [6] § Diversity and distribution of neurons ↔ analysis.zip/computation/figure/Figure4.ipynb, lines 1–32 · score 0.64 · anatomical areas, posterior frontal, brain areas, posterior temporal, selective neurons, brightness
  7. [7] § Methods › Predictivities of neurons across language models › Language model predictivity ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm.py, lines 1–35 · score 0.62 · neuronal firing rates, aligned neuronal, population predictivities, temporal, Language, model
  8. [8] § Methods › Single-neuronal and local field responses to linguistic features › Comparing neuronal activities across hemispheres and cortical areas ↔ analysis.zip/computation/figure/Figure4.ipynb, lines 152–204 · score 0.60 · posterior frontal, anterior temporal, posterior temporal
  9. [9] § Methods › Predictivities of neurons across language models › Language model predictivity ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm_by_time_shuffle.py, lines 1–41 · score 0.59 · randomly permuting, shuffled control, baseline, predictivity, permutation, activities
  10. [10] § Combinatorial linguistic encoding ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm_swap_sentence.py, lines 1–43 · score 0.58 · Swap sentence, population predictivities, Neuronal predictivity, LLaMA, semantic, embeddings
  11. [11] § Local cortical organization ↔ analysis.zip/computation/figure/Figure5.ipynb, lines 1–25 · score 0.57 · coincidence matrix, LFP response, LFP activities, microelectrode, single neuronal, channels
  12. [12] § Linguistic representations by neurons ↔ analysis.zip/computation/code/function_code.py, lines 414–440 · score 0.57 · logistic regression, L2, multinomial, classifier, training, Decoding
  13. [13] § Methods › Single-neuronal and local field responses to linguistic features › Single-neuronal analysis › Modulation of neuronal responses ↔ analysis.zip/computation/code/A2_probe_selectivity_zscore.py, lines 1–26 · score 0.57 · neuronal modulation, neuronal firing rates, magnitude, quantifies, linguistic feature, scores
  14. [14] § Methods › Predictivities of neurons across language models › Neural predictivity dynamics and mapping ↔ analysis.zip/computation/code/C1_neuron_predictivity_mean_various_llm.py, lines 1–35 · score 0.56 · temporal window, neuronal firing rates, dynamics, mapping, predictivity, population
  15. [15] § Methods › Neuronal decoding analysis › Decoding linguistic features from neural activity › Decoding model ↔ analysis.zip/computation/code/function_code.py, lines 414–440 · score 0.55 · logistic regression, L2, multinomial, classifier, decode
  16. [16] § Combinatorial linguistic encoding ↔ analysis.zip/computation/figure/Figure3.ipynb, lines 1–25 · score 0.54 · neurons peak predictivities, exponential fit, population predictivity, sided permutation, utterance, regression

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 436 lines · 17 KB · CC-BY-4.0 · 2 matches

  1. # %% [markdown]
  2. # # Figure 3: Organization and Temporal Dynamics of Linguistic Information Encoding
  3. #
  4. # This section generates the analysis and visualizations for Figure 3, which explores how various models (syntactic, semantic, and contextual) predict neuronal population activity, as well as the temporal dynamics of this predictivity.
  5. #
  6. # ### Figure Panels Breakdown:
  7. # * **Panel b:** Population predictivities. Box plots showing distribution (Q1–Q3), median, and mean, evaluated against chance distributions (1-sided permutation tests).
  8. # * **Panel c:** Contextual integration analysis. Exponential fit demonstrating the relationship between context length and predictive performance.
  9. # * **Panel d:** Validation controls. Performance comparison against swap-sentence and pseudo-sentence null models.
  10. # * **Panel e:** Temporal and layered dynamics. Normalized predictivity heatmaps across time (aligned to word utterance) and network layers.
  11. # * **Panel f:** Peak predictivity mapping. Relationship between single-neuron peak predictivity, peak timing, DNN layer, and specific linguistic feature selectivity.
  12. #
  13. # ### Dependencies & Data Requirements:
  14. # * Requires Ridge regression outputs (`result_llm/`).
  15. # * Requires pre-computed predictivity scores (C1/C2/C3/D1).
  16. # %%
  17. import numpy as np
  18. import pandas as pd
  19. import matplotlib.pyplot as plt
  20. import copy
  21. import matplotlib as mpl
  22. from matplotlib.pyplot import rc
  23. plt.rcParams['font.size'] = '10'
  24. mpl.rcParams['pdf.fonttype'] = 42
  25. # %%
  26. onehot = np.load('../result_llm/combined/C3_explained_r_syn.npy')
  27. onehot_shuffle = np.load('../result_llm/combined/C3_explained_r_syn_shuffle.npy')
  28. onehot_w2v = np.load('../result_llm/combined/C3_explained_r_sem_syn_exclude_oov.npy')
  29. onehot_w2v_shuffle = np.load('../result_llm/combined/C3_explained_r_sem_syn_exclude_oov_shuffle.npy')
  30. exp_w2v = np.load('../result_llm/combined/C3_explained_r_word2vec_exclude_oov.npy')
  31. exp_w2v_shuffle = np.load('../result_llm/combined/C3_explained_r_word2vec_exclude_oov_shuffle.npy')
  32. exp_r3 = np.load('../result_llm/combined/C1_explained_r_combined_normalize_mean_vicuna7b.npy') # dim = session x time (-2 to 2 second, 0.1 s / bin) x layer
  33. exp_r3_shuffle = np.load('../result_llm/combined/C1_explained_r_combined_normalize_mean_vicuna7b_by_time_shuffle.npy') # dim = session x time (-2 to 2 second, 0.1 s / bin) x layer
  34. # %%
  35. # %% [markdown]
  36. # # Figure 3b: Comparison of Neuronal Predictivity: Syntactic, Word2Vec, and LLM Models
  37. # %%
  38. np.random.seed(10)
  39. ss = 0
  40. ee = 3
  41. plt.figure(figsize=(5,6))
  42. values = onehot[:,ss:ee,0].mean(1)
  43. plt.plot(np.random.normal(-4,0.1,14),values,'C0.',alpha=1)
  44. bp1 = plt.boxplot(values, positions = [-4],sym='', widths=[0.3])
  45. for median in bp1['medians']:
  46. median.set(color='C3', linewidth=2)
  47. # print(values.mean(), '\t',values.std(), '\t',np.median(values))
  48. plt.plot([-4-0.1,-3.9],[values.mean(), values.mean()],'r-')
  49. values = onehot_shuffle[:,ss:ee,0].mean(1)
  50. plt.plot(np.random.normal(-3.5,0.1,14),values,'C0.',alpha=1)
  51. bp1 = plt.boxplot(values, positions = [-3.5],sym='', widths=[0.3])
  52. for median in bp1['medians']:
  53. median.set(color='gray', linewidth=2)
  54. plt.plot([-3.5-0.1,-3.4],[values.mean(), values.mean()],'r-')
  55. values = exp_w2v[:,ss:ee,0].mean(1)
  56. plt.plot(np.random.normal(-2,0.1,14),values,'C0.',alpha=1)
  57. bp1 = plt.boxplot(values, positions = [-2],sym='', widths=[0.3])
  58. for median in bp1['medians']:
  59. median.set(color='C2', linewidth=2)
  60. # print(values.mean(), '\t',values.std(), '\t',np.median(values))
  61. plt.plot([-2-0.1,-1.9],[values.mean(), values.mean()],'r-')
  62. values = exp_w2v_shuffle[:,ss:ee,0].mean(1)
  63. plt.plot(np.random.normal(-1.5,0.1,14),values,'C0.',alpha=1)
  64. bp1 = plt.boxplot(values, positions = [-1.5],sym='', widths=[0.3])
  65. for median in bp1['medians']:
  66. median.set(color='gray', linewidth=2)
  67. plt.plot([-1.5-0.1,-1.4],[values.mean(), values.mean()],'r-')
  68. values = onehot_w2v[:,ss:ee,0].mean(1)
  69. plt.plot(np.random.normal(0,0.1,14),values,'C0.',alpha=1)
  70. bp1 = plt.boxplot(values, positions = [0],sym='', widths=[0.3])
  71. for median in bp1['medians']:
  72. median.set(color='C1', linewidth=2)
  73. # print(values.mean(), '\t',values.std(), '\t',np.median(values))
  74. plt.plot([-0.1,0.1],[values.mean(), values.mean()],'r-')
  75. values = onehot_w2v_shuffle[:,ss:ee,0].mean(1)
  76. plt.plot(np.random.normal(0.5,0.1,14),values,'C0.',alpha=1)
  77. bp1 = plt.boxplot(values, positions = [0.5],sym='', widths=[0.3])
  78. for median in bp1['medians']:
  79. median.set(color='gray', linewidth=2)
  80. plt.plot([0.5-0.1,0.6],[values.mean(), values.mean()],'r-')
  81. values = exp_r3[:,(ss+17):(ee+17),5].mean(1)
  82. plt.plot(np.random.normal(2,0.1,14),values,'C0.',alpha=1)
  83. bp1 = plt.boxplot(values, positions = [2],sym='', widths=[0.3])
  84. for median in bp1['medians']:
  85. median.set(color='C0', linewidth=2)
  86. # print(values.mean(), '\t',values.std(), '\t',np.median(values))
  87. plt.plot([2-0.1,2.1],[values.mean(), values.mean()],'r-')
  88. values = exp_r3_shuffle[:,(ss+17):(ee+17),5].mean(1)
  89. plt.plot(np.random.normal(2.5,0.1,14),values,'C0.',alpha=1)
  90. bp1 = plt.boxplot(values, positions = [2.5],sym='', widths=[0.3])
  91. for median in bp1['medians']:
  92. median.set(color='gray', linewidth=2)
  93. plt.plot([2.5-0.1,2.6],[values.mean(), values.mean()],'r-')
  94. plt.yticks(np.arange(-0.1,0.26,0.1))
  95. plt.xticks(np.arange(-4,3,2),['syn','word2vec','syn_w2v','LLM'])
  96. plt.ylabel('predictivity')
  97. plt.ylim([-0.3,0.3])
  98. plt.legend(['orange: median; red: mean'])
  99. plt.savefig('figure3/predictivity_by_syn_w2v_llm.pdf')
  100. # %%
  101. # %% [markdown]
  102. # ### Statistical Significance Analysis: Pairwise Model Comparison
  103. # %%
  104. from scipy import stats
  105. def perm_test_boot(vals1_, vals2_):
  106. pooled = np.hstack([vals1_,vals2_])
  107. np.random.shuffle(pooled)
  108. star_in = pooled[:len(vals1_)]
  109. star_out = pooled[-len(vals2_):]
  110. return star_in.mean()-star_out.mean()
  111. def permutation_test_1side(vals1, vals2, nsamp=1000, seed=None):
  112. if seed is not None:
  113. np.random.seed(seed)
  114. vals1 = np.array(vals1)
  115. vals2 = np.array(vals2)
  116. delta = vals1.mean()-vals2.mean()
  117. est = np.array(list(map(lambda x: perm_test_boot(vals1, vals2),range(nsamp))))
  118. diff_count = len(np.where(est>=delta)[0])
  119. p2 = float(diff_count)/float(nsamp)
  120. return p2
  121. def permutation_test_2side(vals1, vals2, nsamp=1000, seed=None):
  122. if seed is not None:
  123. np.random.seed(seed)
  124. vals1 = np.array(vals1)
  125. vals2 = np.array(vals2)
  126. delta = vals1.mean()-vals2.mean()
  127. est = np.array(list(map(lambda x: perm_test_boot(vals1, vals2),range(nsamp))))
  128. diff_count = len(np.where(np.abs(est)>=np.abs(delta))[0])
  129. p = float(diff_count)/float(nsamp)
  130. return p
  131. def check_significance(data_array, point):
  132. # Filter out NaNs to get clean stats
  133. clean_data = data_array[~np.isnan(data_array)]
  134. mu = np.mean(clean_data)
  135. sigma = np.std(clean_data, ddof=1) # Use ddof=1 for sample standard deviation
  136. # Calculate Z-score: (x - mu) / sigma
  137. z_score = (point - mu) / sigma
  138. # Get the two-tailed p-value
  139. # stats.norm.sf is the survival function (1 - CDF)
  140. p_value = stats.norm.sf(abs(z_score)) * 2
  141. return z_score, p_value
  142. # %%
  143. # Performs 1-sided permutation testing (n=100,000) to evaluate whether the predictive performance differences between the Contextual (LLM) and Syntactic/Semantic models are statistically significant.
  144. ss = 0
  145. ee = 3
  146. values0 = exp_r3[:,(ss+17):(ee+17),5] # time offsets start from -2,000 s for this dataset.
  147. values1 = onehot[:,ss:ee,0]
  148. values2 = exp_w2v[:,ss:ee,0]
  149. values3 = onehot_w2v[:,ss:ee,0]
  150. print('contextual v.s. syn; \tp = ',permutation_test_1side(values0.reshape(-1), values1.reshape(-1), nsamp=100000, seed=46))
  151. print('contextual v.s. sem; \tp = ',permutation_test_1side(values0.reshape(-1), values2.reshape(-1), nsamp=100000, seed=46))
  152. print('contextual v.s. syn+sem; \tp = ',permutation_test_1side(values0.reshape(-1), values3.reshape(-1), nsamp=100000, seed=46))
  153. print('syn v.s. sem; \tp = ',permutation_test_1side(values1.reshape(-1), values2.reshape(-1), nsamp=100000, seed=46))
  154. print('syn v.s. syn+sem; \tp = ',permutation_test_1side(values1.reshape(-1), values3.reshape(-1), nsamp=100000, seed=46))
  155. print('sem v.s. syn+sem; \tp = ',permutation_test_1side(values2.reshape(-1), values3.reshape(-1), nsamp=100000, seed=46))
  156. # %%
  157. # Statistical Validation: Baseline vs. Shuffle Controls
  158. # Performs 1-sided permutation testing (n=100,000) to verify that the predictivity scores for each model architecture (syntactic, semantic, and combined) significantly outperform their respective shuffled null distributions.
  159. values1s = onehot_shuffle[:,ss:ee,0]
  160. values2s = exp_w2v_shuffle[:,ss:ee,0]
  161. values3s = onehot_w2v_shuffle[:,ss:ee,0]
  162. print('syn v.s. shuffle; \tp = ',permutation_test_1side(values1.reshape(-1), values1s.reshape(-1), nsamp=100000, seed=20))
  163. print('sem v.s. shuffle; \tp = ',permutation_test_1side(values2.reshape(-1), values2s.reshape(-1), nsamp=100000, seed=20))
  164. print('syn+sem v.s. shuffle; \tp = ',permutation_test_1side(values3.reshape(-1), values3s.reshape(-1), nsamp=100000, seed=20))
  165. # %%
  166. # %% [markdown]
  167. # # Figure 3c: Predictivity by preceding words
  168. # %%
  169. gram_n = np.array([1,3,5,7,10])
  170. exp_r3 = np.zeros([14,33,len(gram_n)])
  171. for ind, i_n in enumerate(gram_n):
  172. exp_r3[:,:,ind] = np.load('../result_llm/combined/C2_explained_r_combined_vicuna7b_'+str(i_n)+'gram.npy')
  173. exp_r3.shape
  174. # %%
  175. from scipy.optimize import curve_fit
  176. def exponential_func(x, a, b, c):
  177. """
  178. Exponential model function for curve fitting.
  179. y = a * exp(b * x) + c
  180. """
  181. return a * np.exp(b * x) + c
  182. # %%
  183. plt.figure(figsize=(3,5))
  184. maxind = exp_r3.mean(0).reshape(-1,5).argmax(0)
  185. # plt.bar(gram_n[iinds],exp_r3.mean(0)[25:45,:].max(0).max(0)[iinds],alpha=0.5,color='gray')
  186. output = []
  187. for ii in range(exp_r3.shape[2]):
  188. values = exp_r3[:,:,ii].reshape(14,-1)[:,maxind[ii]]
  189. mmean = values.mean()
  190. sstd = values.std()/np.sqrt(values.shape[0])
  191. output.append(mmean)
  192. plt.bar(gram_n[ii], mmean, color='C0',alpha=0.5)
  193. plt.errorbar(gram_n[ii], mmean, sstd, color='gray')
  194. # for jj in range(exp_r3.shape[0])
  195. plt.plot(np.random.normal(gram_n[ii], 0.1, 14), values,'C0.')
  196. plt.xlim([0,11])
  197. plt.xlabel('# of preceding words')
  198. plt.ylabel('predictivity')
  199. popt, pcov = curve_fit(exponential_func, gram_n, output, p0=(-8,-0.08,0.08))
  200. a_fit, b_fit, c_fit = popt
  201. # Create a smooth x-range for plotting the fitted curve
  202. x_fit = np.linspace(min(gram_n), max(gram_n), 100)
  203. y_fit = exponential_func(x_fit, a_fit, b_fit, c_fit)
  204. plt.plot(x_fit, y_fit, label=f'Fitted Curve ($y={a_fit:.2f}e^{{{b_fit:.2f}x}}+{c_fit:.2f}$)', color='red', linewidth=2)
  205. plt.legend(fontsize=10)
  206. plt.savefig('figure3/predictivity_by_preceding_words_errorbars_scatter.pdf')
  207. # %%
  208. # %%
  209. # %% [markdown]
  210. # # Figure 3d: Sentence-context specificity of neuronal activity
  211. # %%
  212. randtok = np.load('../result_llm/combined/C1_explained_r_combined_vicuna7b_random_token_randtok.npy')
  213. swap = np.load('../result_llm/combined/C1_explained_r_combined_normalize_mean_swap_sentence_vicuna7b.npy')
  214. # %%
  215. plt.figure(figsize=(16,10))
  216. plt.subplot(2,2,1)
  217. plt.bar(np.arange(randtok.shape[1]),randtok.mean(axis=0), width=0.5, alpha=0.8,color='C0')
  218. mmean = randtok.mean(0)
  219. sstd = randtok.std(0)/np.sqrt(randtok.shape[0])
  220. plt.errorbar(np.arange(randtok.shape[1]), mmean, sstd, color='gray',fmt='none')
  221. for ii in range(randtok.shape[1]):
  222. plt.plot(np.random.normal(ii, 0.1, 14), randtok[:,ii],'.',color='gray', alpha=0.5)
  223. plt.ylim([-0.3,0.3])
  224. plt.yticks(np.arange(-30,31,10)/100)
  225. # plt.xticks(np.arange(5,11,5))
  226. plt.xlabel('layers')
  227. plt.ylabel('predictivity')
  228. plt.title('Pseudo-sentence input control')
  229. plt.subplot(2,2,2)
  230. plt.bar(np.arange(swap.shape[1]),swap.mean(axis=0), width=0.5, alpha=0.8,color='C0')
  231. mmean = swap.mean(0)
  232. sstd = swap.std(0)/np.sqrt(swap.shape[0])
  233. plt.errorbar(np.arange(swap.shape[1]), mmean, sstd, color='gray',fmt='none')
  234. for ii in range(swap.shape[1]):
  235. plt.plot(np.random.normal(ii, 0.1, 14), swap[:,ii],'.',color='gray', alpha=0.5)
  236. plt.ylim([-0.3,0.3])
  237. plt.yticks(np.arange(-30,31,10)/100)
  238. plt.xlabel('layers')
  239. plt.ylabel('predictivity')
  240. plt.title('swap')
  241. plt.savefig('figure3/vicuna7b_pseudoSentence_Swapping_controls_bars_scatter.pdf')
  242. # %%
  243. # %% [markdown]
  244. # # Figure 3e: Temporal structure of neuronal activity
  245. # %%
  246. folder_all = ['p1_1', 'p1_2', 'p2_1', 'p2_2', 'p3_1', 'p3_2', 'p4_1', 'p5_1', 'p5_2', 'p6_1', 'p7_1', 'p7_2', 'p8_1', 'p8_2']
  247. nn = 0
  248. ww = 0
  249. neuron_nums = []
  250. for folder in folder_all:
  251. neuron_data= np.load('../../../data_processed/'+folder+'/self_FR_aligned2word.npy')
  252. # print(folder, neuron_data.shape[:2])
  253. nn += neuron_data.shape[0]
  254. ww += neuron_data.shape[1]
  255. neuron_nums.append(neuron_data.shape[0])
  256. neuron_nums = np.array(neuron_nums)
  257. # %%
  258. exp_r3 = np.load('../result_llm/combined/C1_explained_r_combined_normalize_mean_vicuna7b.npy')
  259. weights_normalized = neuron_nums / np.sum(neuron_nums)
  260. weights_reshaped = weights_normalized[:, np.newaxis, np.newaxis]
  261. weighted_exp_r3 = exp_r3 * weights_reshaped
  262. weighted_average = np.sum(weighted_exp_r3, axis=0)
  263. # %%
  264. plt.imshow(weighted_average.transpose()/(weighted_average.max()),vmin=0, vmax=1, cmap='magma')
  265. plt.colorbar(orientation='horizontal')
  266. plt.xticks(np.arange(0,41,10),(np.arange(-20,21,10)/10).astype(int))
  267. # plt.xlim([0,80])
  268. plt.xlabel('time (s)')
  269. plt.ylabel('layer')
  270. plt.savefig('figure3/population_predictivity_weight_mean.pdf')
  271. # %%
  272. # This provides an additional non-weighted version for comparison.
  273. plt.imshow(exp_r3.mean(0).transpose()/(exp_r3.mean(0).max()),vmin=0, vmax=1, cmap='magma')
  274. plt.colorbar(orientation='horizontal')
  275. plt.xticks(np.arange(0,41,10),(np.arange(-20,21,10)/10).astype(int))
  276. # plt.xlim([0,80])
  277. plt.xlabel('time (s)')
  278. plt.ylabel('layer')
  279. # %%
  280. exp_r3_indiv = np.load('../result_llm/combined/D1_explained_r_combined_indiv_vicuna7b.npy')
  281. plt.imshow(np.nanmean(exp_r3_indiv,0).T/np.nanmean(exp_r3_indiv,0).max(),vmin=0,vmax=1,cmap='magma')
  282. plt.colorbar(orientation='horizontal')
  283. plt.xticks(np.arange(0,41,10),(np.arange(-20,21,10)/10).astype(int))
  284. # plt.xlim([-1,39])
  285. plt.xlabel('time (s)')
  286. plt.ylabel('layer')
  287. plt.savefig('figure3/explained_r_image_vicuna_indiv_magma.pdf')
  288. # %%
  289. # %% [markdown]
  290. # # Figure 3f: Population organization of linguistic representations
  291. # %%
  292. probe_all = ['Constituency_level','Closing_depth','Constituents','Parts_of_speech','Dependency_depth','Dependencies','Word_position','Word_pitch']
  293. neuron_n = 579
  294. all_probes = np.zeros([0,neuron_n])
  295. for i in range(len(probe_all)):
  296. temp_method = 'ranksum'
  297. temp = np.load('../result_feature/combined/A1_'+probe_all[i]+'_self_fdr_p_'+temp_method+'.npy')
  298. all_probes = np.vstack((all_probes, temp[:neuron_n].reshape([1,-1])))
  299. trans_ind = np.load('trans_ind.npy')
  300. all_probes = all_probes[:,trans_ind]
  301. # %%
  302. layer = exp_r3_indiv.shape[2]
  303. mmap = np.zeros([10, layer, 40]) # feature x layer x time
  304. for dt3 in range(40):
  305. exp_floor = copy.deepcopy(exp_r3_indiv)
  306. exp_floor[exp_floor<0] = 0
  307. for i in range(8):
  308. sel_neu = (all_probes[i,:]<0.05)
  309. mmap[i,:,dt3] = np.nanmean(exp_floor[sel_neu, dt3, :],axis=0)
  310. i += 1
  311. sel_neu = (all_probes.min(axis=0)>0.05)
  312. mmap[i,:, dt3] = np.nanmean(exp_floor[sel_neu, dt3, :], axis=0)
  313. # %%
  314. from matplotlib.patches import Ellipse
  315. # reorder = np.array([2, 3,1,4,5, 0,6,7,8])
  316. reorder = np.array([1,2,4,3,0,5,6,7,8])
  317. scale_factor = 25
  318. xy_scale_factor = 2.25
  319. fig, ax = plt.subplots(1)
  320. for i in range(len(reorder)-3):
  321. argm = np.argmax(mmap[reorder[i]][:,5:20])
  322. argx, argy = np.unravel_index(argm, mmap[reorder[i]][:,5:20].shape)
  323. # ax.plot(argy, argx, 'o', alpha=0.5, markersize = mmap[reorder[i]][:,5:20][argx, argy]*500)
  324. width = (0 + mmap[reorder[i]][:,5:20][argx, argy])*scale_factor
  325. height = (0 + mmap[reorder[i]][:,5:20][argx, argy])*scale_factor*xy_scale_factor
  326. ellipse = Ellipse((argy, argx), width, height, angle=0,
  327. edgecolor='none', facecolor='C'+str(i), lw=2, alpha=0.5)
  328. ax.add_patch(ellipse)
  329. width = (mmap[reorder[i]][:,5:20].std(0)[argy] + mmap[reorder[i]][:,5:20][argx, argy]*2)*scale_factor
  330. height = (mmap[reorder[i]][:,5:20].std(1)[argx] + mmap[reorder[i]][:,5:20][argx, argy]*2)*scale_factor*xy_scale_factor
  331. ellipse = Ellipse((argy, argx), width, height, angle=0,
  332. edgecolor='C'+str(i), facecolor='C'+str(i), lw=1, alpha=0.1,label='_nolegend_')
  333. ax.add_patch(ellipse)
  334. ww = 0.05*scale_factor
  335. hh = ww*xy_scale_factor
  336. stand1 = Ellipse((15,10), ww, hh, angle=0, edgecolor='none', facecolor='gray', lw=2, alpha=0.5)
  337. ax.add_patch(stand1)
  338. ww = 0.1*scale_factor
  339. hh = ww*xy_scale_factor
  340. stand1 = Ellipse((15,10), ww, hh, angle=0, edgecolor='gray', facecolor='gray', lw=1, alpha=0.1)
  341. ax.add_patch(stand1)
  342. ax.set_ylim([0,35])
  343. ax.set_xlim([0,20])
  344. ax.set_xticks(np.arange(0,16,5),(np.arange(-15,1,5)/10))
  345. ax.set_yticks(np.arange(0,35,5))
  346. plt.text(15,5,'scale: 0.05, 0.1')
  347. plt.legend(np.array(probe_all)[reorder[:6]])
  348. plt.xlabel('time (s)')
  349. plt.ylabel('layer')
  350. plt.savefig('figure3/max_predictivity_by_layer_time_probe.pdf')
  351. # %%

Figure3.ipynb, under CC-BY-4.0 · at the source

Overview

Authors: Jing Cai1, Yoav Kfir1, Mohsen Jamali1, Hesen Huang1, Young Joon Kim1, Sydney S Cash2,3,4,5, Ziv M Williams1,3,4,5
  1. Department of Neurosurgery, Massachusetts General Hospital, Harvard Medical School, Boston, MA USA
  2. Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA USA
  3. Center for Neurotechnology and Neurorecovery, Department of Neurology, Massachusetts General Hospital, Boston, MA USA
  4. Harvard-MIT Division of Health Sciences and Technology, Boston, MA USA
  5. Harvard Medical School, Program in Neuroscience, Boston, MA USA
Journal: Nature, volume 656, issue 8127, pages 425-433
Dates: received 19 September 2024; accepted 21 May 2026; published online 17 June 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41586-026-10691-5 · PMID 42310453 · PMCID PMC13468166 · OpenAlex W7165039482
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cognitive (subfield)
Methods: Spectral & time-frequency, Preprocessing, Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Single-unit activity, calcium imaging
Keywords: Language, Neural decoding
MeSH: Brain Mapping*, Language*, Models, Neurological*, Natural Language Processing*, Neurons*, Adult, Female, Humans, Male, Semantics, Speech, Temporal Lobe, Young Adult (* major topic)
Topic: Neurobiology of Language and Bilingualism (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: cited by 3 papers (Europe PMC); 126 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

Zenodo 20128462

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 6 files
Software Heritage: not checked
Found in: “Code availability”
Holds: README, environment (requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: Matplotlib (22 files), NumPy (22 files), pandas (22 files), scikit-learn (13 files), statsmodels (5 files), SciPy (4 files), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
23 files
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it points to the authors' code: Zenodo 20128462

Read it in the paper: doi.org/10.1038/s41586-026-10691-5.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 22 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s41586-026-10691-5.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 2 keywords, 13 MeSH terms, 108 references.

Cite

This paper

Cai, J., Kfir, Y., Jamali, M., Huang, H., Kim, Y. J., Cash, S. S., & Williams, Z. M. (2026). Mapping the neuronal building blocks of human language with language models. Nature, 656(8127), 425-433. https://doi.org/10.1038/s41586-026-10691-5

BibTeX

@article{cai2026mapping,
author = {Cai, Jing and Kfir, Yoav and Jamali, Mohsen and Huang, Hesen and Kim, Young Joon and Cash, Sydney S and Williams, Ziv M},
title = {{Mapping the neuronal building blocks of human language with language models}},
journal = {Nature},
year = {2026},
month = jun,
volume = {656},
number = {8127},
pages = {425--433},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/s41586-026-10691-5},
url = {https://doi.org/10.1038/s41586-026-10691-5},
pmid = {42310453},
pmcid = {PMC13468166}
}

RIS

TY - JOUR
AU - Cai, Jing
AU - Kfir, Yoav
AU - Jamali, Mohsen
AU - Huang, Hesen
AU - Kim, Young Joon
AU - Cash, Sydney S
AU - Williams, Ziv M
TI - Mapping the neuronal building blocks of human language with language models
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/06/17
VL - 656
IS - 8127
SP - 425
EP - 433
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/s41586-026-10691-5
UR - https://doi.org/10.1038/s41586-026-10691-5
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41586-026-10691-5",
"type": "article-journal",
"title": "Mapping the neuronal building blocks of human language with language models",
"container-title": "Nature",
"author": [
{
"family": "Cai",
"given": "Jing"
},
{
"family": "Kfir",
"given": "Yoav"
},
{
"family": "Jamali",
"given": "Mohsen"
},
{
"family": "Huang",
"given": "Hesen"
},
{
"family": "Kim",
"given": "Young Joon"
},
{
"family": "Cash",
"given": "Sydney S"
},
{
"family": "Williams",
"given": "Ziv M"
}
],
"container-title-short": "Nature",
"volume": "656",
"issue": "8127",
"page": "425-433",
"DOI": "10.1038/s41586-026-10691-5",
"PMID": "42310453",
"PMCID": "PMC13468166",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41586-026-10691-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
17
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-75745-8 [code]
A language network in the individualized functional connectomes of 1199 human brains doing arbitrary tasks.
Journal: Nature communications
In common: scikit-learn, pandas, SciPy, 2 other tools, cognitive, 12 references
[2] doi:10.1038/s41586-026-10448-0 [code]
Plasticity and language in the anaesthetized human hippocampus.
Journal: Nature
In common: SciPy, Matplotlib, NumPy, cognitive, 3 references, 2 authors
[3] doi:10.1038/s41586-026-10653-x [code]
A mosaic of whole-body representations on the human precentral gyrus.
Journal: Nature
In common: scikit-learn, SciPy, Matplotlib, 1 other tool, 3 references, 2 authors
[4] doi:10.1038/s41467-026-75455-1 [code]
Shared latent representations of speech production for cross-patient speech decoding.
Journal: Nature communications
In common: statsmodels, scikit-learn, pandas, 3 other tools, cognitive, 7 references
[5] doi:10.1126/sciadv.aec0518 [code]
Hybrid spatial organization and evidence for magnitude-independent neural coding of linguistic information during sentence production.
Journal: Science advances
In common: 10 references
[6] doi:10.1162/imag.a.1283 [code]
The language network responds robustly to sentences across tasks.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: statsmodels, scikit-learn, pandas, 3 other tools, 7 references
[7] doi:10.1038/s41467-026-72253-7 [code]
Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds.
Journal: Nature communications
In common: scikit-learn, pandas, SciPy, 2 other tools, 7 references
[8] doi:10.7554/elife.109400 [code]
High-fidelity neural speech reconstruction through an efficient acoustic-linguistic dual-pathway framework.
Journal: eLife
In common: SciPy, Matplotlib, NumPy, cognitive, 8 references
[9] doi:10.1038/s41467-026-76598-x
Preserved topography, lateralization, selectivity, and functional connectivity of the language network in older brains.
Journal: Nature communications
In common: cognitive, 8 references
[10] doi:10.1038/s41593-026-02258-4 [code]
Laminar organization of cellular microcircuits modulating human interictal epileptiform discharges.
Journal: Nature neuroscience
In common: statsmodels, scikit-learn, pandas, 3 other tools, 5 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.