OSCR

Specific expansion of motor cortical projections in a singing mouse.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Methods › MAPseq data analysis › Heatmaps and cell-type labelling ↔ MAPseq/ED_fig5_UMI_thresholds_sexes.ipynb, lines 76–132 · score 0.64 · CT cells, PT cells, HB, striatum, thalamus, HY
  2. [2] § Methods › MAPseq data analysis › Heatmaps and cell-type labelling ↔ MAPseq/ED_fig4_infectivity_downsampling.ipynb, lines 100–156 · score 0.63 · CT cells, PT cells, striatum, thalamus, infectivity, HY
  3. [3] § Methods › MAPseq data analysis › Degree and motif analyses ↔ MAPseq/fig4_motifs.ipynb, lines 354–406 · score 0.56 · unobserved neurons, Motif proportions, population, probabilities, cells, brain
  4. [4] § Methods › MAPseq data analysis › Degree and motif analyses ↔ MAPseq/ED_fig4_infectivity_downsampling.ipynb, lines 513–565 · score 0.56 · unobserved neurons, Motif proportions, population, probabilities, cells, brain
  5. [5] § Testing models of brain-wide cortical projections ↔ MAPseq/ED_fig4_infectivity_downsampling.ipynb, lines 100–156 · score 0.54 · pyramidal tract, intratelencephalic, striatum, thalamus, cortical, PT
  6. [6] § Testing models of brain-wide cortical projections ↔ MAPseq/ED_fig5_UMI_thresholds_sexes.ipynb, lines 76–132 · score 0.54 · pyramidal tract, intratelencephalic, striatum, thalamus, cortical, PT
  7. [7] § Methods › MAPseq tissue processing ↔ MAPseq/ED_fig4_infectivity_downsampling.ipynb, lines 59–96 · score 0.53 · projecting neurons, Injection sites, Allen, infections, dissection

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 998 lines · 35 KB · no license · 4 matches

  1. # %%
  2. ###### load packages
  3. import pandas as pd
  4. import numpy as np
  5. import matplotlib.pyplot as plt
  6. import seaborn as sns
  7. import math
  8. from itertools import combinations, chain
  9. from scipy import stats
  10. from upsetplot import from_memberships
  11. from scipy.stats import chi2, binomtest
  12. from statsmodels.distributions.empirical_distribution import ECDF # for generating empirical cdfs
  13. import matplotlib.lines as mlines # needed for custom legend
  14. import os
  15. # needed for editable text in svg format
  16. plt.rcParams['font.family'] = ['sans-serif']
  17. plt.rcParams['font.sans-serif'] = ['Arial']
  18. plt.rcParams['text.usetex'] = False
  19. plt.rcParams['svg.fonttype'] = 'none'
  20. # load metadata
  21. metadata = pd.read_csv("metadata.csv")
  22. # import custom colormaps
  23. from colormaps import blue_cmp, orange_cmp
  24. # specify blue/orange colors for qualitative intervals/data
  25. blue_qual = [blue_cmp.colors[50], blue_cmp.colors[100], blue_cmp.colors[150], blue_cmp.colors[200], blue_cmp.colors[250]]
  26. orange_qual = [orange_cmp.colors[36], orange_cmp.colors[72], orange_cmp.colors[108], orange_cmp.colors[144], orange_cmp.colors[180], orange_cmp.colors[216], orange_cmp.colors[252]]
  27. %matplotlib inline
  28. # %% [markdown]
  29. # # Load data
  30. # %%
  31. # set-up out paths
  32. out_path = '/your/path/here/'
  33. # %%
  34. # load data
  35. in_path = 'your/path/here'
  36. # load binarized data
  37. omc_bin = []
  38. for idx, row in metadata.iterrows():
  39. df = pd.read_csv(os.path.join(in_path, row["mice"]+"_"+row["dataset"]+"_bin.csv"), index_col=0)
  40. omc_bin.append(df)
  41. omc_countN = []
  42. for idx, row in metadata.iterrows():
  43. df = pd.read_csv(os.path.join(in_path, row["mice"]+"_"+row["dataset"]+"_countN.csv"), index_col=0)
  44. omc_countN.append(df)
  45. # %% [markdown]
  46. # # Functions
  47. # %% [markdown]
  48. # ## Processing functions
  49. # %%
  50. def clean_up_data(df_dirty, to_drop = ['OB', 'ACAi', 'ACAc', 'HIP'], inj_site="OMCi"):
  51. """Clean up datasets so all matrices are in the same format. Function
  52. (1) drops unwanted columns, e.g. negative controls or dissections of other injection sites.
  53. (2) renames RN (allen acronym) as BS (my brainstem acronym)
  54. (3) drops neurons w/ 0 projections after dropped columns
  55. Args:
  56. df_dirty (DataFrame): Pandas dataframe (neurons x area) that needs to be processed
  57. to_drop (list, optional): columns to drop. Defaults to ['OB', 'ACAi', 'ACAc', 'HIP'].
  58. inj_site (str, optional): Injection site. Defaults to "OMCi".
  59. Returns:
  60. DataFrame: Cleaned up data in dataframe format.
  61. """
  62. # 1. drop unused areas
  63. dropped = df_dirty.drop(to_drop, axis=1)
  64. # 2. change RN to bs
  65. replaced = dropped.rename(columns={'RN':'BS'})
  66. # 3. drop neurons w/ 0 projections after removing negative regions
  67. if type(inj_site)==str:
  68. nodes = replaced.drop([inj_site], axis=1).sum(axis=1)
  69. else:
  70. nodes = replaced.drop(inj_site, axis=1).sum(axis=1)
  71. n_idx = nodes > 0 # non-zero projecting neurons
  72. clean = replaced[n_idx]
  73. return clean
  74. # %%
  75. def sort_by_celltype(proj, it_areas=["OMCc", "AUD", "STR"], ct_areas=["TH"],
  76. pt_areas=["AMY","HY","SNr","SCm","PG","PAG","BS"],
  77. sort=True):
  78. """
  79. Function takes in projection matrix and outputs matrix sorted by the 3 major celltypes:
  80. - IT = intratelencephalic (projects to cortical and/or Striatum), type = 10
  81. - CT = corticalthalamic (projects to thalamus w/o projection to brainstem), type = 100
  82. - PT = pyramidal tract (projects to brainstem += other areas), type = 1000
  83. Returns single dataframe with cells sorted and labelled by 3 cell types (IT/CT/PT)
  84. Args:
  85. proj (DataFrame): pd.DataFrame of BC x area. Entries can be normalized BC or binary.
  86. it_areas (list, optional): Areas to determine IT cells. Defaults to ["OMCc", "AUD", "STR"].
  87. ct_areas (list, optional): Areas to determine CT cells. Defaults to ["TH"]. Don't actually use this...
  88. pt_areas (list, optional): Areas to determine PT cells. Defaults to ["AMY","HY","SNr","SCm","PG","PAG","BS"].
  89. sort (bool, optional): Whether to sort by cell type or return w/ original index. Defaults to True.
  90. Returns:
  91. df_out (DataFrame): Returns dataframe with extra column (type) labelled and sorted by cell type
  92. """
  93. ds=proj.copy()
  94. # Isolate PT cells
  95. pt_counts = ds[pt_areas].sum(axis=1)
  96. pt_idx = ds[pt_counts>0].index
  97. ds_pt = ds.loc[pt_idx,:]
  98. ds_pt['type'] = "PT"
  99. # Isolate remaining non-PT cells
  100. ds_npt = ds.drop(pt_idx)
  101. # Identify CT cells by thalamus projection
  102. th_idx = ds_npt['TH'] > 0
  103. ds_th = ds_npt[th_idx]
  104. if sort:
  105. ds_th = ds_th.sort_values('TH', ascending=False)
  106. ds_th['type'] = "CT"
  107. # Identify IT cells by the remaining cells (non-PT, non-CT)
  108. ds_nth = ds_npt[~th_idx]
  109. if sort:
  110. ds_nth = ds_nth.sort_values(it_areas,ascending=False)
  111. ds_nth['type'] = "IT"
  112. # combine IT and CT cells
  113. ds_npt = pd.concat([ds_nth, ds_th])
  114. # combine IT/CT and PT cells
  115. if sort:
  116. sorted = pd.concat([ds_npt,ds_pt],ignore_index=True)
  117. df_out=sorted.reset_index(drop=True)
  118. else:
  119. df_out = pd.concat([ds_npt,ds_pt]).sort_index()
  120. return(df_out)
  121. # %%
  122. def dfs_preprocess_counts(df_list, drop=["type"],
  123. norm_by="inj_median", inj_site="OMCi", medians=None):
  124. """Take dataframe and process it for downstream analysis. Can be binarized or
  125. non-binarized data.
  126. Args:
  127. df_list (list): List of dataframes,
  128. drop (list, optional): List of columns to drop before returning. Defaults to ["OMCi", "type"].
  129. norm_by (str, optional): How to normalize counts. Can be "inj_median" or 'all_median' or "given". Defaults to "inj_median".
  130. inj_site (str, optional): What column to use for injection norm. Defaults to "OMCi".
  131. medians (list, optional): if norm_by=="given", uses list of medians (same length as df_list) to normalize dataframes. Defaults to None.
  132. Returns:
  133. out_list (list): List of dataframes same order as input list.
  134. """
  135. out_list = []
  136. norms_out = []
  137. for i in range(len(df_list)):
  138. df = df_list[i].drop(drop, axis=1)
  139. if norm_by=="inj_row":
  140. # normalize on neuron by neuron basis to injection site value
  141. out_df = df.div(df[inj_site], axis=0)
  142. elif norm_by=="given":
  143. out_df = df/medians[i]
  144. elif norm_by=="inj_mean":
  145. vals = df[inj_site].values.flatten()
  146. mean = np.mean(vals)
  147. norms_out.append(mean)
  148. out_df = df/mean
  149. else:
  150. # normalize by non-zero median of injection site
  151. if norm_by=="inj_median":
  152. vals = df[inj_site].values.flatten() # added injection site median
  153. # normalize by median of all non-zero values median across whole dataset
  154. elif norm_by=="all_median":
  155. vals = df.values.flatten()
  156. idx = vals.nonzero() # only use non-zero ncounts for determining median
  157. plot = vals[idx]
  158. median = np.median(plot)
  159. norms_out.append(median)
  160. out_df = df/median
  161. for j in range(len(drop)):
  162. out_df[drop[j]] = df_list[i][drop[j]] # add dropped columns back in, can preserve labels
  163. out_list.append(out_df)
  164. if len(norms_out)>0:
  165. return(norms_out, out_list)
  166. else:
  167. return(out_list)
  168. # %%
  169. def dfs_to_cdf(df_list, plot_areas, resolution=1000, metadata=metadata):
  170. """Takes in list of DFs of count(N) data and returns dataframe w/ cdf data that can be plotted.
  171. Returned Dataframe includes metadata
  172. Args:
  173. df_list (list): list of DataFrames of count(N) data.
  174. plot_areas (list): List of strings of areas to calculate cdfs.
  175. resolution (int, optional): Used to determine resolution of cdf line. Defaults to 1000.
  176. medatadata (df, optional): Metadata where row corresponds to df_list indices. Defaults to metadata.
  177. Returns:
  178. cdf_df (DataFrame): Dataframe of cdf summary values
  179. all_ecdfs (list): Empirical cdf
  180. """
  181. # combine all DFs into one df labelled w/ metadata
  182. all_bc = pd.DataFrame(columns=list(df_list[0].columns)+["mice", "species", "dataset"])
  183. for i in range(metadata.shape[0]):
  184. df = df_list[i].copy(deep=True)
  185. df['mice'] = metadata.loc[i, 'mice']
  186. df['species'] = metadata.loc[i, "species"]
  187. df['dataset'] = metadata.loc[i, "dataset"]
  188. all_bc = pd.concat([all_bc, df])
  189. all_bc = all_bc.reset_index(drop=True)
  190. cdf_df = pd.DataFrame(columns=["x", "cdf", "mice", "species", "dataset", "area"])
  191. all_ecdfs = {}
  192. # calculate cdf by area, then add by mouse
  193. for area in plot_areas:
  194. # just use nonzero BC
  195. area_idx = all_bc[area] > 0
  196. area_bc = all_bc.loc[area_idx, [area, "mice", "species", "dataset"]]
  197. # get min/max for each area to set cdf bounds
  198. area_min = area_bc[area].min()
  199. area_max = area_bc[area].max()
  200. for i in range(metadata.shape[0]):
  201. micei = metadata.loc[i, 'mice']
  202. mice_bc = area_bc[area_bc['mice']==micei]
  203. if mice_bc[area].sum()==0:
  204. print("NO BARCODES, cannot compute ECDF for", area, metadata.loc[i,'mice'])
  205. else:
  206. # print(area, metadata.loc[i,"mice"])
  207. ecdf = ECDF(mice_bc[area])
  208. x = np.logspace(np.log10(area_min), np.log10(area_max), num=resolution)
  209. y = ecdf(x)
  210. int = pd.DataFrame({"x":x, "cdf":y, "mice":metadata.loc[i,"mice"], "species":metadata.loc[i,"species"],
  211. "dataset":metadata.loc[i,"dataset"], "area":area})
  212. cdf_df = pd.concat([cdf_df, int])
  213. all_ecdfs[micei+"_"+area] = ecdf
  214. return(cdf_df, all_ecdfs)
  215. # %%
  216. def dfs_to_proportions(df_list, drop=["OMCi", "type"], keep=None, cell_type=None, meta=metadata,
  217. inj_site="OMCi", aud_rename=False):
  218. """Output dataframe of proportions in format that can be plotted with seaborn
  219. Args:
  220. df_list (list):
  221. - List of dataframes of neurons/BC by areas
  222. drop (list, optional):
  223. - Defaults to ["OMCi", "type"]
  224. - list of areas/columns to drop before calculating proportions
  225. cell_type (string, optional):
  226. - Specify cell types in df, either IT, CT or PT
  227. - Defaults to None
  228. Returns:
  229. plot_df (pandas_dataframe):
  230. - returns dataframe in format for seaborn plotting
  231. - columns = areas, and other metadata
  232. """
  233. plot_df = pd.DataFrame(columns=["area", "proportion", "mice", "species", "dataset"])
  234. if cell_type == "IT":
  235. drop.extend([inj_site, 'TH', 'HY', 'AMY', 'SNr', 'SCm', 'PG',
  236. 'PAG', 'BS'])
  237. elif cell_type == "PT":
  238. if aud_rename:
  239. drop.extend([inj_site,inj_site[:-1]+"c", 'AUD/TEa', "STR"])
  240. else:
  241. drop.extend([inj_site,inj_site[:-1]+"c", 'AUD', "STR"])
  242. if keep:
  243. drop = []
  244. mice = meta["mice"]
  245. species = meta["species"]
  246. dataset = meta["dataset"]
  247. for i in range(len(df_list)):
  248. df = df_list[i].drop(drop, axis=1)
  249. if keep:
  250. df = df.loc[:, keep] # just subset keep columns
  251. bc_sum = df.sum()
  252. proportion = bc_sum/df.shape[0]
  253. df_add = pd.DataFrame({"area":proportion.index.values, "proportion":proportion.values,
  254. "mice":meta.loc[i,"mice"], "species":meta.loc[i,"species"], "dataset":meta.loc[i,"dataset"]})
  255. # add cell_type if specified
  256. if cell_type:
  257. df_add["type"] = cell_type
  258. plot_df = pd.concat([plot_df, df_add])
  259. return plot_df
  260. # %%
  261. def downsample_neurons(data, meta=metadata, random_state=10, species="MMus", sample_ns=None):
  262. """Given list of dataframe, sample from combined neurons/cells (with replacement), in equivalent numbers
  263. to singing mouse (or sample_ns) cells/brain. Returns list where each element is dataframe of neurons w/ numbers equivalent to ns.
  264. Args:
  265. data (list): List of pandas dataframes where each row is a different cell/neuron
  266. metadata (DataFrame, optional): Metadata used to determine which df in data is lab/singing mouse.
  267. Defaults to metadata.
  268. random_state (int, optional): Set random state to use for repeatable sampling.
  269. Defaults to 10.
  270. species (str, optional): Species to sample. Defaults to "MMus".
  271. sample_ns (list, optional): List of int use as sample size per 'brain', if none defaults to singing mouse brain size.
  272. Defaults to None.
  273. """
  274. all = [data[i] for i in range(metadata.shape[0]) if metadata.loc[i,'species']==species]
  275. all = pd.concat(all).reset_index(drop=True)
  276. # print("mm_all.shape[0]", mm_all.shape[0])
  277. pool = all.copy()
  278. samp = []
  279. ns = []
  280. if sample_ns:
  281. ns = sample_ns
  282. else:
  283. # get steg sample neuron sizes
  284. for i in range(metadata.shape[0]):
  285. if metadata.loc[i,'species']=="STeg":
  286. ns.append(data[i].shape[0])
  287. for i in range(len(ns)):
  288. n = ns[i]
  289. int = pool.sample(n, random_state=random_state+i) # can't have same random_state for every round or will sample the same neurons
  290. samp.append(int.reset_index(drop=True))
  291. return(samp)
  292. # %%
  293. def statistic_testing(df, sp1="MMus", sp2="STeg", to_plot='proportion',
  294. groupby="area", test="ttest"):
  295. """output dataframe based on comparison of species individuals
  296. output dataframe can be used for making volcano plot
  297. Args:
  298. df (DataFrame): DataFrame of repeated values across species and inviduals.
  299. sp1 (str, optional): Group1 to use as comparison. Defaults to "MMus".
  300. sp2 (str, optional): Group1 to use as comparison. Defaults to "STeg".
  301. to_plot (str, optional): Column to calculate comparison. Defaults to 'proportion'.
  302. groupby (str, optional): Categories within to do comparison. Defaults to "area".
  303. test (str, optional): What test to use ["ttest", "mannwhitneyu"]. Defaults to "ttest".
  304. Returns:
  305. plot (DataFrame): pd.DataFrame w/ pvalues, means across species, and log. Can be used for volcano plot.
  306. """
  307. groups = sorted(df[groupby].unique())
  308. # sp1
  309. sp1_df = df[df["species"]==sp1]
  310. sp1_array = sp1_df.pivot_table(columns='mice', values=to_plot, index=groupby).values
  311. sp2_df = df[df["species"]==sp2]
  312. sp2_array = sp2_df.pivot_table(columns='mice', values=to_plot, index=groupby).values
  313. if test=="ttest":
  314. results = stats.ttest_ind(sp1_array, sp2_array, axis=1)
  315. elif test=="mannwhitneyu":
  316. results = stats.mannwhitneyu(sp1_array, sp2_array, axis=1)
  317. p_vals = results[1]
  318. plot = pd.DataFrame({groupby:groups, "p-value":p_vals})
  319. plot[sp1+"_mean"] = sp1_array.mean(axis=1)
  320. plot[sp1+"_sem"] = stats.sem(sp1_array, axis=1)
  321. plot[sp2+"_mean"] = sp2_array.mean(axis=1)
  322. plot[sp2+"_sem"] = stats.sem(sp2_array, axis=1)
  323. # plot["effect_size"] = (plot["st_mean"]-plot["mm_mean"]) / (plot["st_mean"] + plot["mm_mean"]) # modulation index
  324. plot["fold_change"] = plot[sp2+"_mean"]/(plot[sp1+"_mean"])
  325. plot["log2_fc"] = np.log2(plot["fold_change"])
  326. plot["nlog10_p"] = -np.log10(plot["p-value"])
  327. plot["p<0.05"] = plot["p-value"]<0.05
  328. return(plot)
  329. # %%
  330. def estimate_n_total(df, plot_areas):
  331. """Calculate estimated N_total based on observed number of neuron per areas.
  332. Works with IT and PT areas with as many areas as wanted. See Han et al., 2017
  333. for more detail/formula.
  334. Args:
  335. df (DataFrame): Binary matrix of BC x areas
  336. areas (list): List of areas using to calculat motifs. Should be list of strings.
  337. Returns:
  338. n_total (int): Estimated n total for input data.
  339. """
  340. n_obs = df.shape[0]
  341. n_areas = df.sum()
  342. all_terms = []
  343. for k in range(1, len(plot_areas)+1):
  344. combos = list(combinations(plot_areas, k))
  345. term = 0
  346. for i in range(len(combos)):
  347. product = 1
  348. for j in range(len(combos[i])):
  349. n_area = n_areas[combos[i][j]]
  350. product = product*n_area
  351. term = term + product
  352. all_terms.append(term)
  353. # need to subtract first term from n_obs
  354. all_terms[0] = n_obs - all_terms[0]
  355. # multiply every other by -1 starting w/ 3rd term
  356. for l in range(len(all_terms)):
  357. if l>=2 and l%2==0:
  358. all_terms[l] = -1*all_terms[l]
  359. # find roots of polynomial
  360. roots = np.roots(all_terms)
  361. # convert to real numbers if numbers complex
  362. if isinstance(roots[0],complex):
  363. reals = []
  364. for num in roots:
  365. if num.imag==0:
  366. reals.append(num.real)
  367. else:
  368. reals = roots.copy()
  369. # pick root that is more than n_obs
  370. # if can't find root more than n_obs, return n_obs as n_total
  371. n_total = n_obs
  372. for num in reals:
  373. if num > n_obs:
  374. n_total = round(num)
  375. return(n_total)
  376. # %%
  377. def df_to_motif_proportion(df, areas, proportion=True, subset=None):
  378. """Make series to feed into upset plot based on given data and area list
  379. Args:
  380. df (pd.DataFrame): df containing bc x area
  381. areas (list): List of areas (str) to be plotted
  382. proportion (bool, optional): Whether to output counts or proportions. Defaults to True.
  383. subset (str, optional): Whether to subset on certain brain area (e.g. 'PAG'). Defaults to None.
  384. Returns:
  385. plot_s (DataFrame): dataframe of proprotion/count of each motif for input data.
  386. """
  387. # generate all combinations of areas in true/false list
  388. area_comb = []
  389. for i in range(len(areas)):
  390. n = i+1
  391. area_comb.append(list(combinations(areas, n)))
  392. area_comb_list = list(chain.from_iterable(area_comb)) # flatten list
  393. memberships = from_memberships(area_comb_list) # generate true/false
  394. area_comb_TF = memberships.index.values # extract array w/ true/false values
  395. area_comb_names = memberships.index.names # get order of areas
  396. # calculate number of neurons of each motif
  397. comb_count = []
  398. for tf in area_comb_TF:
  399. neurons = df
  400. for i in range(len(area_comb_names)):
  401. # subset dataset on presence/absence of proj
  402. neurons = neurons[neurons[area_comb_names[i]]==tf[i]]
  403. if proportion: # return proportion
  404. comb_count.append(neurons.shape[0]/df.shape[0])
  405. else: # return count
  406. comb_count.append(neurons.shape[0])
  407. plot_s = from_memberships(area_comb_list, data=comb_count)
  408. if subset:
  409. motif_areas = plot_s.index.names
  410. subset_idx = motif_areas.index(subset)
  411. idx = [i for i, x in enumerate(plot_s.index) if x[subset_idx]]
  412. plot_s = plot_s[idx]
  413. return(plot_s)
  414. # %%
  415. def df_to_motif_estimated_proportions(data, combinations, adjust_total=False):
  416. """Given dataframe of cells and index of combinations (generated from df_to_motif_prportions),
  417. return series similar to output for df_to_motif_proportions. Proportions are estimated
  418. by multiplying the bulk/population proporitons/probabilities.
  419. Args:
  420. data (DataFrame): DataFrame of BC x areas, binary
  421. combinations (MultiIndex): Index from output of df-to_motif_proportions
  422. adjust_total (bool): whether to adjust the matrix w/ unobserved neurons. Defaults to False.
  423. Returns:
  424. """
  425. # adjust n_total (add in 0 projectors), if not done
  426. if adjust_total:
  427. # calculate n_total
  428. n_obs = data.shape[0]
  429. n_total = estimate_n_total(data, combinations.names)
  430. n_unobs = np.array(n_total )- np.array(n_obs)
  431. unobs_df = pd.DataFrame(0, index=np.arange(n_unobs), columns=data.columns)
  432. df = pd.concat([data, unobs_df]).reset_index(drop=True)
  433. else:
  434. df = data.copy()
  435. # get brain areas specified in combinations
  436. names = list(combinations.names)
  437. # subset df to just columns used in motifs
  438. df_subset = df.loc[:,names]
  439. # get bulk proportions across those areas
  440. bulk_prop = df_subset.sum(axis=0)/df.shape[0]
  441. # calculate expected motif proportion based on product of bulk proportions
  442. expected_proportions = []
  443. names = combinations.names
  444. for i in range(combinations.shape[0]):
  445. motif = combinations[i]
  446. product = 1
  447. for j in range(len(names)):
  448. if motif[j]: # if area in motif, multiply by bulk probability
  449. product = product * bulk_prop[names[j]]
  450. else: # if area not in motif, multiply by 1-bulk probability (chance of not projecting to area)
  451. product = product * (1-bulk_prop[names[j]])
  452. expected_proportions.append(product)
  453. motif_expected_prop = pd.Series(expected_proportions, index=combinations)
  454. return(motif_expected_prop)
  455. # %%
  456. def TF_to_motifs(index):
  457. """Generate motif strings from True/False MultiIndex genreated by df_to_motif_proportion()
  458. Args:
  459. index (MultiIndex): True/False for area combinations
  460. Returns:
  461. array of strings for area combinations
  462. """
  463. motifs_strings = []
  464. for r in index:
  465. motif = ""
  466. for i in range(len(index.names)):
  467. if r[i]:
  468. motif = motif+index.names[i]+"_"
  469. motifs_strings.append(motif)
  470. return(motifs_strings)
  471. # %% [markdown]
  472. # ## Plotting functions
  473. # %%
  474. def dot_plot(data, subset=None, title=None, err="se", add_legend=False,
  475. to_plot="proportion", ylim=(0), fig_size=(3.5,3.5),
  476. jitter=False, alpha=0.3):
  477. """Plot closed circle of value per area for data.
  478. Args:
  479. data (DataFrame): Dataframe of proportions, included downsampled data.
  480. subset (list, optional): List of what to subset on to plot. Defaults to None.
  481. title (str, optional): Title for plot. Defaults to None.
  482. err (str, optional): Error to plot, can be "ci", "pi", "se", or "sd". Defaults to "se".
  483. add_legend (bool, optional): Whether to add legend labeling mean/err. Defaults to False.
  484. to_plot (str, optional): Column to plot. Defaults to "proportion".
  485. ylim (tuple, optional): lower bound for yaxis. Defaults to (0).
  486. fig_size (tuple, optional): Dimensions (in inches) of plot. Defaults to (3.5,3.5).
  487. jitter (bool, optional): Whether to have overlapping points. Defaults to False.
  488. alpha (float, optional): How much transparency for points. Defaults to 0.3.
  489. """
  490. if subset:
  491. df = data[data[subset[0]]==subset[1]]
  492. else:
  493. df = data.copy()
  494. fig, ax = plt.subplots()
  495. sns.stripplot(x='species', y=to_plot, hue="species",
  496. dodge=False, data=df, jitter=jitter,
  497. size=10, alpha=alpha)
  498. sns.pointplot(df, x="species", y=to_plot, hue="species",
  499. dodge=False, marker="+",
  500. ls="", zorder=10, errorbar="se")
  501. ax.set_xlabel("")
  502. plt.title(title, size=20)
  503. plt.ylim(ylim) # make sure y axis starts at 0
  504. ax.spines['right'].set_visible(False)
  505. ax.spines['top'].set_visible(False)
  506. # set figure size
  507. fig = plt.gcf()
  508. fig.set_size_inches(fig_size[0],fig_size[1])
  509. if add_legend:
  510. legend = mlines.Line2D([], [], color="black", marker="[]", linewidth=0, label="mean, "+err)
  511. plt.legend(handles=[legend], loc="lower right")
  512. else:
  513. plt.legend([],[], frameon=False)
  514. return(fig)
  515. # %%
  516. def plot_cdf(data, plot_areas, log=True, title="", color_by="species", colors=[blue_cmp.colors[255], orange_cmp.colors[255]],
  517. individual=True, meta=metadata, legend=True, fig_size=(3,3), calc_cdf=True):
  518. """Takes in countN data and returns cdf plots
  519. Args:
  520. data (list): list of DataFrames of count(N), where each element in animal
  521. plot_areas (list): list of strings of areas to include in final output
  522. log (bool, optional): Whether to use log on axis scale. Defaults to True.
  523. title (str, optional): figure title. Defaults to "".
  524. color_by (str, optional): Can be "mice", "species", or "dataset", what to label as metadata. Defaults to "species".
  525. colors (list, optional): colors used to label cdfs. Defaults to [blue_cmp.colors[255], orange_cmp.colors[255]].
  526. individual (bool, optional): _description_. Defaults to True.
  527. meta (_type_, optional): _description_. Defaults to metadata.
  528. calc_cdf (bool, optional): Whether to calcualte cdf or not. Defaults to True.
  529. """
  530. # calculate ecdf per animal and put into dataframe
  531. if calc_cdf:
  532. cdf_df, foo = dfs_to_cdf(data, plot_areas=plot_areas, metadata=meta)
  533. else:
  534. cdf_df = data.copy()
  535. # calculate number of axes needed
  536. n = math.ceil(len(plot_areas)/5) # round up divide by 4 = axs rows
  537. if len(plot_areas)==1:
  538. fig, ax = plt.subplots(1,1, figsize=(5,5))
  539. ax_list = [ax]
  540. else:
  541. fig, axs = plt.subplots(n, 5, figsize=(20, 5*n))
  542. ax_list = axs.flat
  543. i = 0
  544. for ax in ax_list:
  545. if i < len(plot_areas):
  546. area = plot_areas[i]
  547. plot = cdf_df[cdf_df['area']==area]
  548. groups = plot[color_by].unique()
  549. plot_1 = plot[plot[color_by] ==groups[0]]
  550. plot_2 = plot[plot[color_by] ==groups[1]]
  551. if individual:
  552. sns.lineplot(plot_1, x="x", y="cdf", estimator=None, units="mice", color=colors[0], ax=ax) # plots individual mice
  553. sns.lineplot(plot_2, x="x", y="cdf", estimator=None, units="mice", color=colors[1], ax=ax) # plots individual mice
  554. else:
  555. sns.lineplot(plot_1, x="x", y="cdf", ax=ax) # plots mean ci95
  556. sns.lineplot(plot_2, x="x", y="cdf", ax=ax) # plots mean ci95
  557. # get rid of top and right axis
  558. ax.spines['right'].set_visible(False)
  559. ax.spines['top'].set_visible(False)
  560. if log:
  561. ax.set_xscale("log")
  562. ax.set_xlabel("Normalized Counts")
  563. ax.set_title(area)
  564. i+=1
  565. else:
  566. ax.axis('off')
  567. # create cutom legend
  568. if legend:
  569. colors = [colors[0], colors[1]]
  570. lines = [mlines.Line2D([0], [0], color=c, linewidth=3) for c in colors]
  571. labels = [groups[0], groups[1]]
  572. fig.legend(lines,labels, bbox_to_anchor=(0.75, 0.935))
  573. # increase text size
  574. plt.rcParams.update({'font.size': 12})
  575. if title!="":
  576. plt.suptitle(title, y=0.93, size=20)
  577. elif title=="":
  578. plt.suptitle("By "+color_by, y=0.93, size=20)
  579. # erase minor ticks
  580. plt.minorticks_off()
  581. # set figure size
  582. fig = plt.gcf()
  583. fig.set_size_inches(fig_size[0],fig_size[1])
  584. return(fig)
  585. # %% [markdown]
  586. # # Data pre-processing
  587. # %%
  588. # initial processing
  589. # binary processing
  590. omc_clean = [clean_up_data(df) for df in omc_bin]
  591. omc_type = [sort_by_celltype(df) for df in omc_clean]
  592. # normalized count (countN) processing
  593. omc_cleanN = [clean_up_data(df) for df in omc_countN]
  594. omc_typeN = [sort_by_celltype(df) for df in omc_cleanN]
  595. means, omc_preprocessN = dfs_preprocess_counts(omc_typeN, norm_by="inj_mean") # normalize by injection mean
  596. medians, omc_preprocessN_med = dfs_preprocess_counts(omc_typeN, norm_by="inj_median")
  597. # %% [markdown]
  598. # # Infectivity
  599. # - Plots for Extended Data fig. 4
  600. # %%
  601. # Plot unique barcodes by species
  602. infect_df = pd.DataFrame(columns=["Unique Barcodes", "mice", "species", "dataset"])
  603. for i in range(metadata.shape[0]):
  604. infect_df.loc[i,"Unique Barcodes"] = omc_type[i].shape[0]
  605. infect_df.loc[i, "mice"] = metadata.loc[i,"mice"]
  606. infect_df.loc[i,"species"] = metadata.loc[i, "species"]
  607. infect_df.loc[i, "dataset"] = metadata.loc[i, "dataset"]
  608. print("Mean")
  609. display(infect_df.groupby("species")["Unique Barcodes"].mean())
  610. print("SEM")
  611. infect_df.groupby("species")["Unique Barcodes"].sem()
  612. # %%
  613. dot_plot(infect_df, to_plot="Unique Barcodes", fig_size=(3,3))
  614. plt.yscale("log")
  615. plt.ylim(100, 30000)
  616. plt.minorticks_off()
  617. # plt.savefig(out_path+"infectivity_uBC_dotplot.svg", dpi=300, bbox_inches="tight")
  618. plt.show()
  619. # %%
  620. # plot barcode count per neuron for injection site (OMCi)
  621. # NO WITHIN SAMPLE NORM - i.e. only normalized to spike-ins
  622. plot_cdf(omc_typeN, plot_areas=["OMCi"], color_by="species",
  623. title=None, legend=False, individual=False, fig_size=(2,2))
  624. plt.title("OMCi")
  625. # plt.savefig(out_path+"omci_rnaspike_norm_cdf.svg", dpi=300, bbox_inches="tight")
  626. plt.show()
  627. # %%
  628. # take medians for each animal
  629. omc_typeN[0]["OMCi"].median()
  630. medians_omci = []
  631. for i in range(metadata.shape[0]):
  632. medians_omci.append(omc_typeN[i]["OMCi"].median())
  633. omci_med = metadata.copy()
  634. omci_med["inj_med"] = medians_omci
  635. omci_med["area"] = "OMCi"
  636. # calcualte mean, sem for each species
  637. print("Mean")
  638. print(omci_med.groupby("species")["inj_med"].mean())
  639. print("\nSEM")
  640. print(omci_med.groupby("species")["inj_med"].sem())
  641. # Mann whitney u test
  642. mm_med = omci_med[omci_med["species"]=="MMus"]
  643. st_med = omci_med[omci_med["species"]=="STeg"]
  644. print(stats.mannwhitneyu(mm_med["inj_med"], st_med["inj_med"]))
  645. dot_plot(omci_med, to_plot="inj_med")
  646. plt.show()
  647. # %% [markdown]
  648. # # Downsampling
  649. # %%
  650. # seperate it cells
  651. omc_it = [df[df['type']=="IT"] for df in omc_type]
  652. # %% [markdown]
  653. # # Downsampling proportions
  654. # %%
  655. # seperate binarized counts into celltypes
  656. it_bin = [df[df["type"]=="IT"].reset_index(drop=True) for df in omc_type]
  657. pt_bin = [df[df["type"]=="PT"].reset_index(drop=True) for df in omc_type]
  658. it_areas = ["OMCc", "AUD", "STR"]
  659. pt_areas = ["TH", "HY", "AMY", "PAG", "SNr", "SCm", "PG", "BS"]
  660. # calculate area proprotions
  661. it_prop = dfs_to_proportions(it_bin, cell_type="IT", drop=["OMCi", "type"])
  662. pt_prop = dfs_to_proportions(pt_bin, cell_type="PT", drop=["OMCi", "type"])
  663. # %%
  664. # downsample mmus neurons
  665. mmus_down_meta = pd.DataFrame({"mice":["mm_down1", "mm_down2", "mm_down3", "mm_down4",
  666. "mm_down5", "mm_down6", "mm_down7"],
  667. "species":["MMus\ndownsample"]*7,
  668. "dataset":["aggregate"]*7})
  669. it_mmus_down = downsample_neurons(it_bin)
  670. pt_mmus_down = downsample_neurons(pt_bin)
  671. # calculate proportions
  672. it_mm_down_prop = dfs_to_proportions(it_mmus_down, cell_type="IT", drop=["OMCi", "type"],
  673. meta=mmus_down_meta)
  674. pt_mm_down_prop = dfs_to_proportions(pt_mmus_down, cell_type="PT", drop=["OMCi", "type"],
  675. meta=mmus_down_meta)
  676. # add to proportion dataframes
  677. it_prop_all = pd.concat([it_prop, it_mm_down_prop])
  678. pt_prop_all = pd.concat([pt_prop, pt_mm_down_prop])
  679. # %%
  680. # plot downsampled dot plots
  681. # plot it areas
  682. for area in it_areas:
  683. dot_plot(it_prop_all, subset=["area",area], title=area, add_legend=False,
  684. jitter=True)
  685. plt.ylabel("Propotion")
  686. # plt.savefig(out_path+"OMC_"+area+"_proportion_dot_downsample.svg", dpi=300, bbox_inches="tight")
  687. plt.show()
  688. # plot pt areas
  689. for area in pt_areas:
  690. dot_plot(pt_prop_all, subset=["area",area], title=area, add_legend=False,
  691. jitter=True)
  692. plt.ylabel("Propotion")
  693. # plt.savefig(out_path+"OMC_"+area+"_proportion_dot_downsample.svg", dpi=300, bbox_inches="tight")
  694. plt.show()
  695. # %%
  696. # significance testing - IT
  697. print("\nMANNWHITNEYU - mmus v. steg")
  698. mannwhitney = statistic_testing(it_prop_all, test="mannwhitneyu", sp1="MMus", sp2="STeg")
  699. display(mannwhitney)
  700. print("\nMANNWHITNEYU - mmus downsample v. steg")
  701. mannwhitney = statistic_testing(it_prop_all, test="mannwhitneyu", sp1="MMus\ndownsample", sp2="STeg")
  702. display(mannwhitney)
  703. # %%
  704. # significance testing - PT
  705. print("\nMANNWHITNEYU - mmus v. steg")
  706. mannwhitney = statistic_testing(pt_prop_all, test="mannwhitneyu", sp1="MMus", sp2="STeg")
  707. display(mannwhitney)
  708. print("\nMANNWHITNEYU - mmus downsample v. steg")
  709. mannwhitney = statistic_testing(pt_prop_all, test="mannwhitneyu", sp1="MMus\ndownsample", sp2="STeg")
  710. display(mannwhitney)
  711. # %% [markdown]
  712. # # Downsampling motifs
  713. # %%
  714. # calculate estimates for motifs
  715. plot_areas = ["OMCc", "AUD", "STR"]
  716. # Estimate n-totals
  717. n_totals = [estimate_n_total(omc_it[i], plot_areas) for i in range(len(omc_it))]
  718. # Count obs motifs
  719. n_obs_motifs = [df_to_motif_proportion(df, areas=plot_areas, proportion=False) for df in omc_it]
  720. motifs = n_obs_motifs[0].index
  721. # convert to proportions using adjusted n_totals
  722. p_obs_motifs = [n_obs_motifs[i]/n_totals[i] for i in range(len(n_totals))]
  723. # calculate expected proportions based on independent bulk probabilities adjusted for n_total
  724. p_expected_motifs = [df_to_motif_estimated_proportions(df, motifs, adjust_total=True) for df in omc_it]
  725. motif_strings = TF_to_motifs(motifs)
  726. # put into dataframe - note p_obs is adjusted for n_total
  727. it_motifs_df = []
  728. for i in range(len(n_totals)):
  729. df = pd.DataFrame({"motifs":motif_strings, "n_shape":omc_it[i].shape[0], "n_total":n_totals[i],
  730. "n_obs":n_obs_motifs[i], "p_obs":p_obs_motifs[i], "p_exp":p_expected_motifs[i],
  731. "mice":metadata.loc[i,"mice"], "species":metadata.loc[i,"species"]})
  732. it_motifs_df.append(df)
  733. it_motifs_df = pd.concat(it_motifs_df).reset_index(drop=True)
  734. it_motifs_df
  735. # %%
  736. # downsample MMus and calculate motif proportions
  737. # downsample
  738. mmus_down_meta = pd.DataFrame({"mice":["mm_down1", "mm_down2", "mm_down3", "mm_down4",
  739. "mm_down5", "mm_down6", "mm_down7"],
  740. "species":["MMus\ndownsample"]*7,
  741. "dataset":["aggregate"]*7})
  742. it_mmus_down = downsample_neurons(omc_it)
  743. # %%
  744. # calculate estimates for motifs
  745. plot_areas = ["OMCc", "AUD", "STR"]
  746. # Estimate n-totals
  747. n_totals = [estimate_n_total(it_mmus_down[i], plot_areas) for i in range(len(it_mmus_down))]
  748. # Count obs motifs
  749. n_obs_motifs = [df_to_motif_proportion(df, areas=plot_areas, proportion=False) for df in it_mmus_down]
  750. motifs = n_obs_motifs[0].index
  751. # convert to proportions using adjusted n_totals
  752. p_obs_motifs = [n_obs_motifs[i]/n_totals[i] for i in range(len(n_totals))]
  753. # calculate expected proportions based on independent bulk probabilities adjusted for n_total
  754. p_expected_motifs = [df_to_motif_estimated_proportions(df, motifs, adjust_total=True) for df in it_mmus_down]
  755. motif_strings = TF_to_motifs(motifs)
  756. # put into dataframe - note p_obs is adjusted for n_total
  757. it_motifs_df_mmd = []
  758. for i in range(len(n_totals)):
  759. df = pd.DataFrame({"motifs":motif_strings, "n_shape":it_mmus_down[i].shape[0], "n_total":n_totals[i],
  760. "n_obs":n_obs_motifs[i], "p_obs":p_obs_motifs[i], "p_exp":p_expected_motifs[i],
  761. "mice":mmus_down_meta.loc[i,"mice"], "species":mmus_down_meta.loc[i,"species"]})
  762. it_motifs_df_mmd.append(df)
  763. it_motifs_df_mmd = pd.concat(it_motifs_df_mmd).reset_index(drop=True)
  764. # it_motifs_df_mmd
  765. # %%
  766. # plot
  767. plot_df = pd.concat([it_motifs_df, it_motifs_df_mmd])
  768. for motif in motif_strings:
  769. dot_plot(plot_df, subset=["motifs",motif], to_plot="p_obs", title=motif, add_legend=False,
  770. jitter=True)
  771. plt.ylabel("Propotion")
  772. # plt.savefig(out_path+"OMC_"+motif+"_proportion_dot_downsample.svg", dpi=300, bbox_inches="tight")
  773. plt.show()
  774. # %%
  775. # Significance testing
  776. print("\nMANNWHITNEYU - mmus v. steg")
  777. mannwhitney = statistic_testing(plot_df, to_plot="p_obs", groupby="motifs", test="mannwhitneyu", sp1="MMus", sp2="STeg")
  778. display(mannwhitney)
  779. print("\nMANNWHITNEYU - mmus downsample v. steg")
  780. mannwhitney = statistic_testing(plot_df, to_plot="p_obs", groupby="motifs", test="mannwhitneyu", sp1="MMus\ndownsample", sp2="STeg")
  781. display(mannwhitney)

ED_fig4_infectivity_downsampling.ipynb at commit 9b5fa18, no license · at the source

Overview

  1. Cold Spring Harbor Laboratory,Cold Spring Harbor, NY USA
  2. School of Biological Sciences, Cold Spring Harbor Laboratory,Cold Spring Harbor, NY USA
Institutions: Cold Spring Harbor Laboratory (United States)
Journal: Nature, volume 655, issue 8122, pages 438-446
Dates: received 27 November 2024; accepted 26 March 2026; published online 6 May 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41586-026-10458-y · PMID 42092144 · PMCID PMC13345907 · OpenAlex W7160437351
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism), systems (subfield)
Methods: Connectivity, Statistics, Spectral & time-frequency, Evoked potentials, Graphs
Keywords: Neuroscience, Speciation
MeSH: Motor Cortex*, Rodentia*, Vocalization, Animal*, Animals, Auditory Cortex, Female, Male, Mice, Neural Pathways, Species Specificity (* major topic)
Topic: Animal Vocal Communication and Behavior (Developmental Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NINDS NIH HHS (RF1 NS132046)
Citations: cited by 2 papers (Europe PMC); 81 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

singingmicelab/Isko-2026-MAPseq-code

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 9b5fa186b8cb4fee3e0f806829ff0eacb97c62ff, 5 February 2026
Languages: Jupyter (4), Python (1)
Size: 10 files, 5 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, 4 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (5 files), NumPy (5 files), pandas (4 files), SciPy (4 files), seaborn (4 files), statsmodels (3 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
6 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41586-026-10458-y.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41586-026-10458-y.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 2 keywords, 10 MeSH terms, 1 funder, 75 references.

Cite

This paper

Isko, E. C., Harpole, C. E., Zheng, X. M., Zhan, H., Davis, M. B., Zador, A. M., & Banerjee, A. (2026). Specific expansion of motor cortical projections in a singing mouse. Nature, 655(8122), 438-446. https://doi.org/10.1038/s41586-026-10458-y

BibTeX

@article{isko2026specific,
author = {Isko, Emily C. and Harpole, Clifford E. and Zheng, Xiaoyue Mike and Zhan, Huiqing and Davis, Martin B. and Zador, Anthony M. and Banerjee, Arkarup},
title = {{Specific expansion of motor cortical projections in a singing mouse}},
journal = {Nature},
year = {2026},
month = may,
volume = {655},
number = {8122},
pages = {438--446},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/s41586-026-10458-y},
url = {https://doi.org/10.1038/s41586-026-10458-y},
pmid = {42092144},
pmcid = {PMC13345907}
}

RIS

TY - JOUR
AU - Isko, Emily C.
AU - Harpole, Clifford E.
AU - Zheng, Xiaoyue Mike
AU - Zhan, Huiqing
AU - Davis, Martin B.
AU - Zador, Anthony M.
AU - Banerjee, Arkarup
TI - Specific expansion of motor cortical projections in a singing mouse
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/05/06
VL - 655
IS - 8122
SP - 438
EP - 446
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/s41586-026-10458-y
UR - https://doi.org/10.1038/s41586-026-10458-y
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41586-026-10458-y",
"type": "article-journal",
"title": "Specific expansion of motor cortical projections in a singing mouse",
"container-title": "Nature",
"author": [
{
"family": "Isko",
"given": "Emily C."
},
{
"family": "Harpole",
"given": "Clifford E."
},
{
"family": "Zheng",
"given": "Xiaoyue Mike"
},
{
"family": "Zhan",
"given": "Huiqing"
},
{
"family": "Davis",
"given": "Martin B."
},
{
"family": "Zador",
"given": "Anthony M."
},
{
"family": "Banerjee",
"given": "Arkarup"
}
],
"container-title-short": "Nature",
"volume": "655",
"issue": "8122",
"page": "438-446",
"DOI": "10.1038/s41586-026-10458-y",
"PMID": "42092144",
"PMCID": "PMC13345907",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41586-026-10458-y",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
6
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s12859-026-06490-4 [code]
Tissueformer: extending single-cell foundation models to predict population-level phenotypes.
Journal: BMC bioinformatics
In common: seaborn, pandas, SciPy, 2 other tools, mouse, 1 reference, author Anthony Zador
[2] doi:10.1038/s41467-026-72057-9 [code]
Sex-specific behavioral feedback modulates sensorimotor processing and drives flexible social behavior.
Journal: Nature communications
In common: statsmodels, seaborn, pandas, 3 other tools, systems, 2 references
[3] doi:10.1016/j.isci.2026.117136 [code]
Complex pectoral-fin-driven locomotion underlies refined visuomotor behavior in the cichlid &lt;i&gt;Astatotilapia burtoni&lt;/i&gt;.
Journal: iScience
In common: pandas, SciPy, Matplotlib, 1 other tool, systems, 3 references
[4] doi:10.1038/s42003-026-10276-y [code]
The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.
Journal: Communications biology
In common: statsmodels, seaborn, pandas, 3 other tools, 2 references
[5] doi:10.1038/s41593-026-02376-z [code]
A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions.
Journal: Nature neuroscience
In common: statsmodels, seaborn, pandas, 3 other tools, mouse, 2 references
[6] doi:10.1126/sciadv.adv3770 [code]
Evolution of a central dopamine circuit underlies adaptation of a light-evoked sensorimotor response in the blind cavefish.
Journal: Science advances
In common: seaborn, pandas, SciPy, 2 other tools, systems, 2 references
[7] doi:10.1371/journal.pcbi.1014024 [code]
Social familiarity strengthens neural and vocal responses to conspecific calls in zebra finches.
Journal: PLoS computational biology
In common: seaborn, pandas, SciPy, 2 other tools, 2 references
[8] doi:10.1093/pnasnexus/pgag055 [code]
Comparative transcriptomics reveals differences in cortical cell type organization between metatherian and eutherian mammals.
Journal: PNAS nexus
In common: seaborn, pandas, SciPy, 2 other tools, mouse, 2 references
[9] doi:10.1038/s41593-026-02388-9 [code]
Hippocampal CA3 connectomics reveals a gradient of mossy fiber inputs and selective feedforward inhibition onto pyramidal cells.
Journal: Nature neuroscience
In common: seaborn, pandas, SciPy, 2 other tools, systems, mouse, 2 references
[10] doi:10.1038/s41593-026-02362-5 [code]
Replay of procedural memory is independent of the hippocampus.
Journal: Nature neuroscience
In common: statsmodels, seaborn, pandas, 3 other tools, mouse, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.