OSCR

A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Methods › Quantification of feature separability ↔ morphology/lobe_separability.ipynb, lines 358–450 · score 0.98 · random forest classifier, balanced accuracy score, class_weight, max_depth, max_features, min_samples_split
  2. [2] § Methods › Classification of layer 2/3 neurons ↔ morphology/lobe_separability.ipynb, lines 358–450 · score 0.86 · StandardScaler, cross validation, balanced accuracy, stratified, fold, train
  3. [3] § Methods › Morphology type clustering ↔ preprocessing/mtype_clustering_pyr_23.py, lines 440–519 · score 0.83 · Calinski Harabasz, Davies Bouldin, Silhouette coefficient, clusters, aggregation, metrics
  4. [4] § Results › Cellular RNA sequencing validates key corresponding mouse−human regions in brain lobes ↔ transcriptomics/gene_corresponding_dotplot.ipynb, lines 159–296 · score 0.79 · superior parietal, transverse temporal, cuneus, precentral, MOp, parahippocampal
  5. [5] § Results › A framework for joint mouse−human cortical analysis of single neurons in corresponding mouse and human brain regions ↔ transcriptomics/gene_corresponding_dotplot.ipynb, lines 159–296 · score 0.63 · temporal lobe, human brains, mouse brains, frontal, parietal, FL
  6. [6] § Methods › Multimodal lobe separability › Molecular ↔ transcriptomics/transcriptome_analysis_merfish.ipynb, lines 508–620 · score 0.60 · LAMP5 LHX6, MERFISH, pairwise, FL, PL, TL
  7. [7] § Results › Cellular RNA sequencing validates key corresponding mouse−human regions in brain lobes ↔ transcriptomics/transcriptome_analysis_merfish.ipynb, lines 796–938 · score 0.55 · Homologous gene, mouse regions, NP, transcriptomic, metrics, ranks

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 1,080 lines · 46 KB · MIT · 2 matches

  1. # %%
  2. import pandas as pd
  3. import numpy as np
  4. from sklearn.decomposition import PCA
  5. import os
  6. import seaborn as sns
  7. import matplotlib
  8. from matplotlib import pyplot as plt
  9. import sys
  10. sys.path.append('..')
  11. from basicfunc import MouseAnatomyTree, bootstrap_stat_test
  12. import scipy
  13. from sklearn.metrics import normalized_mutual_info_score, adjusted_mutual_info_score, mutual_info_score
  14. import pymrmr
  15. from sklearn.model_selection import cross_validate, LeaveOneOut, StratifiedKFold
  16. from sklearn.ensemble import RandomForestClassifier
  17. from sklearn.neighbors import KNeighborsClassifier
  18. from sklearn.linear_model import LogisticRegression
  19. from sklearn.preprocessing import StandardScaler, LabelEncoder
  20. from sklearn.metrics import confusion_matrix
  21. from sklearn import metrics
  22. from sklearn.discriminant_analysis import QuadraticDiscriminantAnalysis
  23. from sklearn.tree import DecisionTreeClassifier
  24. from sklearn.svm import SVC
  25. from sklearn.model_selection import GridSearchCV
  26. from scipy.stats import wilcoxon
  27. import warnings
  28. warnings.filterwarnings("ignore")
  29. # %%
  30. used_morpho_features = ['Center Shift', 'Average Contraction',
  31. 'Average Bifurcation Angle Remote',
  32. 'Max Branch Order', 'Number of Bifurcations', 'Total Length',
  33. 'Max Path Distance',
  34. 'Average Euclidean Distance', 'Average Path Distance', '3D Density',
  35. 'Volume',]
  36. mapped_cols = ['Center Shift','Avg. Straightness','Avg. Remote Bifur Angle',
  37. 'Max Branch Order','Bifur Num','Total Length','Max Path Dist','Avg. Euclidean Dist',
  38. 'Avg. Path Dist','3D Density','Volume']
  39. feat_mapping_dict = dict(zip(used_morpho_features,mapped_cols))
  40. homo_color_lut = {
  41. 'FL':np.array([255,206,72])/255.0,
  42. 'PL':np.array([153,217,234])/255.0,
  43. 'TL':np.array([118,137,211])/255.0}
  44. # mouse
  45. color1 = np.array([29/255,140/255,67/255])
  46. # human
  47. color2 = np.array([255,102,102])/255.0
  48. # reference dataset
  49. color3=np.array([191,200,215])/255.0
  50. # %%
  51. # Whether to only consider pyramidal cells in layer 2/3
  52. pyr23_only = False
  53. # %%
  54. mouse_anatomy_tree = MouseAnatomyTree('../Data/external/tree.json')
  55. # cell type
  56. df_ct_mouse = pd.read_csv(r'..\Data\metadata\mouse_celltype.csv',index_col=0)
  57. if pyr23_only:
  58. df_ct_mouse = df_ct_mouse[df_ct_mouse['layer'].isin(['2/3'])]
  59. df_ct_mouse = df_ct_mouse[df_ct_mouse['is_pyramidal']==1]
  60. else:
  61. df_ct_mouse = df_ct_mouse[df_ct_mouse['layer'].isin(['2/3', '5'])]
  62. df_ct_mouse = df_ct_mouse[df_ct_mouse['is_pyramidal']==1]
  63. pass
  64. df_ct_human = pd.read_csv(r'..\Data\metadata\human_celltype.csv',index_col=0)
  65. if pyr23_only:
  66. df_ct_human = df_ct_human[df_ct_human['layer'].isin(['L2/3'])]
  67. df_ct_human = df_ct_human[df_ct_human['is_pyramidal']==1]
  68. else:
  69. df_ct_human = df_ct_human[df_ct_human['is_pyramidal']==1]
  70. pass
  71. # morphology features of cropped swc
  72. crop_thres_list = [50,]
  73. restem = 8
  74. path_morpho = r'..\Data\Morphology'
  75. path_nbranch = r'..\Data\Morphology'
  76. dict_df_morpho_mouse = {}
  77. for crop_thres in crop_thres_list:
  78. tmpdf = pd.read_csv(os.path.join(path_morpho, f"mouse_apical_morphology_restem{restem}_crop{crop_thres}.csv"), index_col=0)[used_morpho_features]
  79. # tmpdf['3D Density'] = (tmpdf['3D Density']+0).apply(np.log10)
  80. tmpdf['3D Density'] = (tmpdf['Total Length']/tmpdf['Volume']).apply(np.log10)
  81. tmpdf['Volume'] = (tmpdf['Volume']+1).apply(np.log10)
  82. tmpdf.columns = [feat_mapping_dict[x] for x in tmpdf.columns]
  83. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"mouse_apical_n_branch_num_restem{restem}_crop{crop_thres}.csv"), index_col=0)
  84. for i in range(6,11): del tmptmpdf[f'{i}']
  85. del tmptmpdf['1']
  86. tmptmpdf.columns = [f"L{int(x)} Branch Num" for x in tmptmpdf.columns]
  87. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  88. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  89. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"mouse_apical_n_branch_length_restem{restem}_crop{crop_thres}.csv"), index_col=0)
  90. for i in range(6,11): del tmptmpdf[f'{i}']
  91. tmptmpdf.columns = [f"L{int(x)} Branch Length" for x in tmptmpdf.columns]
  92. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  93. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  94. tmpdf.columns = ["A_" + x for x in tmpdf.columns]
  95. tmptmpdf = pd.read_csv(os.path.join(path_morpho, f"mouse_basal_morphology_restem{restem}_crop{crop_thres}.csv"), index_col=0)[used_morpho_features]
  96. # tmpdf['3D Density'] = (tmpdf['3D Density']+0).apply(np.log10)
  97. tmptmpdf['3D Density'] = (tmptmpdf['Total Length']/tmptmpdf['Volume']).apply(np.log10)
  98. tmptmpdf['Volume'] = (tmptmpdf['Volume']+1).apply(np.log10)
  99. tmptmpdf.columns = [feat_mapping_dict[x] for x in tmptmpdf.columns]
  100. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  101. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"mouse_basal_n_branch_num_restem{restem}_crop{crop_thres}.csv"), index_col=0)
  102. for i in range(6,11): del tmptmpdf[f'{i}']
  103. tmptmpdf.columns = [f"L{int(x)} Branch Num" for x in tmptmpdf.columns]
  104. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  105. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  106. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"mouse_basal_n_branch_length_restem{restem}_crop{crop_thres}.csv"), index_col=0)
  107. for i in range(6,11): del tmptmpdf[f'{i}']
  108. tmptmpdf.columns = [f"L{int(x)} Branch Length" for x in tmptmpdf.columns]
  109. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  110. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  111. tmpdf.columns = ["B_" + x if not x.startswith('A_') else x for x in tmpdf.columns]
  112. dict_df_morpho_mouse[crop_thres] = tmpdf
  113. dict_df_morpho_human = {}
  114. for crop_thres in crop_thres_list:
  115. tmpdf = pd.read_csv(os.path.join(path_morpho, f"human_apical_morphology_crop{crop_thres}.csv"), index_col=0)[used_morpho_features]
  116. # tmpdf['3D Density'] = (tmpdf['3D Density']+0).apply(np.log10)
  117. tmpdf['3D Density'] = (tmpdf['Total Length']/tmpdf['Volume']).apply(np.log10)
  118. tmpdf['Volume'] = (tmpdf['Volume']+1).apply(np.log10)
  119. tmpdf.columns = [feat_mapping_dict[x] for x in tmpdf.columns]
  120. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"human_apical_n_branch_num_crop{crop_thres}.csv"), index_col=0)
  121. for i in range(6,11): del tmptmpdf[f'{i}']
  122. del tmptmpdf['1']
  123. tmptmpdf.columns = [f"L{int(x)} Branch Num" for x in tmptmpdf.columns]
  124. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  125. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  126. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"human_apical_n_branch_length_crop{crop_thres}.csv"), index_col=0)
  127. for i in range(6,11): del tmptmpdf[f'{i}']
  128. tmptmpdf.columns = [f"L{int(x)} Branch Length" for x in tmptmpdf.columns]
  129. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  130. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  131. tmpdf.columns = ["A_" + x for x in tmpdf.columns]
  132. tmptmpdf = pd.read_csv(os.path.join(path_morpho, f"human_basal_morphology_crop{crop_thres}.csv"), index_col=0)[used_morpho_features]
  133. # tmpdf['3D Density'] = (tmpdf['3D Density']+0).apply(np.log10)
  134. tmptmpdf['3D Density'] = (tmptmpdf['Total Length']/tmptmpdf['Volume']).apply(np.log10)
  135. tmptmpdf['Volume'] = (tmptmpdf['Volume']+1).apply(np.log10)
  136. tmptmpdf.columns = [feat_mapping_dict[x] for x in tmptmpdf.columns]
  137. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  138. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"human_basal_n_branch_num_crop{crop_thres}.csv"), index_col=0)
  139. for i in range(6,11): del tmptmpdf[f'{i}']
  140. tmptmpdf.columns = [f"L{int(x)} Branch Num" for x in tmptmpdf.columns]
  141. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  142. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  143. tmptmpdf = pd.read_csv(os.path.join(path_nbranch, f"human_basal_n_branch_length_crop{crop_thres}.csv"), index_col=0)
  144. for i in range(6,11): del tmptmpdf[f'{i}']
  145. tmptmpdf.columns = [f"L{int(x)} Branch Length" for x in tmptmpdf.columns]
  146. # tmpdf = pd.concat([tmpdf, tmptmpdf.iloc[:,tmptmpdf.sum(axis=0).values>0]], axis=1)
  147. tmpdf = pd.concat([tmpdf, tmptmpdf], axis=1)
  148. tmpdf.columns = ["B_" + x if not x.startswith('A_') else x for x in tmpdf.columns]
  149. dict_df_morpho_human[crop_thres] = tmpdf
  150. # %%
  151. def compute_histograms(arr1,arr2,outlier: str='minmax'):
  152. arr1 = np.array(arr1)
  153. arr2 = np.array(arr2)
  154. if outlier not in ['minmax','boxplot']:
  155. raise ValueError(f'invalid outlier value: {outlier}')
  156. lowb, highb = 0, 0
  157. if outlier=='boxplot':
  158. Q11,Q12,Q13=np.percentile(arr1, [25,50,75])
  159. IQR1 = Q13-Q11
  160. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  161. Q21,Q22,Q23=np.percentile(arr2, [25,50,75])
  162. IQR2 = Q23-Q21
  163. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  164. lowb = min(lowb1, lowb2)
  165. highb = max(highb1, highb2)
  166. elif outlier=='minmax':
  167. lowb = min(np.min(arr1),np.min(arr2))
  168. highb = max(np.max(arr1),np.max(arr2))
  169. arr1_filter = arr1[(arr1>=lowb)&(arr1<=highb)]
  170. arr2_filter = arr2[(arr2>=lowb)&(arr2<=highb)]
  171. arr_concat = np.hstack([arr1_filter,arr2_filter])
  172. ct_arr_concat = [0]*len(arr1_filter)+[1]*len(arr2_filter)
  173. _, bin_edges = np.histogram(arr_concat,bins='auto')
  174. bin_num = len(bin_edges) - 1
  175. interval = bin_edges[1]-bin_edges[0]
  176. bin_lowb = np.min(bin_edges)
  177. bin_highb = np.max(bin_edges)
  178. hist_fv=[]
  179. for i,v in enumerate(arr_concat):
  180. if v<bin_lowb:
  181. # hist_fv.append(0)
  182. # print(v,bin_lowb)
  183. continue
  184. elif v>bin_highb:
  185. # hist_fv.append(bin_num-1)
  186. # print(v,bin_highb)
  187. continue
  188. elif v==bin_highb:
  189. hist_fv.append((v-bin_lowb)//interval-1)
  190. else:
  191. hist_fv.append((v-bin_lowb)//interval)
  192. return hist_fv[:len(arr1_filter)], hist_fv[len(arr1_filter):]
  193. def DIFF_feat_ct(arr1,arr2,method='JSD',stat_test=True,group1_meta=None,group2_meta=None):
  194. assert method in ['NMI','JSD','MI','AMI','CORR'], f'unsupported method: {method}'
  195. def JSD_score(p,q):
  196. M=(p+q)/2
  197. output = 0.5*scipy.stats.entropy(p,M,base=2)+0.5*scipy.stats.entropy(q, M,base=2)
  198. if np.isnan(output):
  199. return 0.0
  200. else:
  201. return output
  202. Q11,Q12,Q13=np.percentile(arr1, [25,50,75])
  203. IQR1 = Q13-Q11
  204. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  205. Q21,Q22,Q23=np.percentile(arr2, [25,50,75])
  206. IQR2 = Q23-Q21
  207. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  208. lowb = min(lowb1, lowb2)
  209. highb = max(highb1, highb2)
  210. # arr_concat = np.hstack((arr1[(arr1>=lowb)&(arr1<highb)],arr2[(arr2>=lowb)&(arr2<highb)]))
  211. # ct_arr = [0]*len(arr1[(arr1>=lowb)&(arr1<highb)])+[1]*len(arr2[(arr2>=lowb)&(arr2<highb)])
  212. arr_concat = np.hstack([arr1,arr2])
  213. ct_arr = [0]*len(arr1)+[1]*len(arr2)
  214. # _, bin_edges = np.histogram(arr_concat[(arr_concat>=lowb)&(arr_concat<=highb)],bins='auto')
  215. _, bin_edges = np.histogram(arr_concat,bins='auto')
  216. bin_num = len(bin_edges)-1
  217. interval = bin_edges[1]-bin_edges[0]
  218. tmparr1 = arr_concat[:len(arr1)]
  219. tmparr2 = arr_concat[len(arr1):]
  220. if stat_test:
  221. pv,sig,cohens_d,pv_mlm = bootstrap_stat_test(arr1,arr2,'mwu',group1_meta=group1_meta,group2_meta=group2_meta)
  222. else:
  223. pv,sig,cohens_d,pv_mlm = None, None,None, None
  224. bin_lowb = np.min(bin_edges)
  225. bin_highb = np.max(bin_edges)
  226. if method in ['NMI','MI','AMI','CORR']:
  227. hist_fv=[]
  228. for i,v in enumerate(arr_concat):
  229. if v<bin_lowb:
  230. # hist_fv.append(0)
  231. # print(v,bin_lowb)
  232. continue
  233. elif v>bin_highb:
  234. # hist_fv.append(bin_num-1)
  235. # print(v,bin_highb)
  236. continue
  237. elif v==bin_highb:
  238. hist_fv.append((v-bin_lowb)//interval-1)
  239. else:
  240. hist_fv.append((v-bin_lowb)//interval)
  241. if method=='NMI':
  242. score = normalized_mutual_info_score(hist_fv,ct_arr)
  243. elif method=='MI':
  244. score = mutual_info_score(hist_fv,ct_arr)
  245. elif method=='AMI':
  246. score = adjusted_mutual_info_score(hist_fv,ct_arr)
  247. elif method=='CORR':
  248. score = abs(scipy.stats.pointbiserialr(hist_fv,ct_arr).correlation)
  249. elif method == 'JSD':
  250. hist_fv1 = np.zeros(bin_num)
  251. hist_fv2 = np.zeros(bin_num)
  252. for i,v in enumerate(tmparr1):
  253. if v<bin_lowb:
  254. # hist_fv1.append(0)
  255. continue
  256. elif v>bin_highb:
  257. # hist_fv1.append(bin_num-1)
  258. continue
  259. elif v==bin_highb:
  260. hist_fv1[int((v-bin_lowb)//interval-1)]+=1
  261. else:
  262. hist_fv1[int((v-bin_lowb)//interval)]+=1
  263. for i,v in enumerate(tmparr2):
  264. if v<bin_lowb:
  265. # hist_fv2.append(0)
  266. continue
  267. elif v>bin_highb:
  268. # hist_fv2.append(bin_num-1)
  269. continue
  270. elif v==bin_highb:
  271. hist_fv2[int((v-bin_lowb)//interval-1)]+=1
  272. else:
  273. hist_fv2[int((v-bin_lowb)//interval)]+=1
  274. hist_fv1 = np.asarray(hist_fv1)/(np.sum(hist_fv1)+1e-10)
  275. hist_fv2 = np.asarray(hist_fv2)/(np.sum(hist_fv2)+1e-10)
  276. score = JSD_score(hist_fv1,hist_fv2)
  277. return score,pv,sig,cohens_d,pv_mlm
  278. # %%
  279. '''Comparison of the difference between regions'''
  280. select_cropthres = [50]
  281. df_morpho_mouse = pd.DataFrame()
  282. df_morpho_human = pd.DataFrame()
  283. for crop_thres in select_cropthres:
  284. tmpdf = dict_df_morpho_mouse[crop_thres].copy()
  285. tmpdf.columns = [f"R{crop_thres}_"+x for x in tmpdf.columns]
  286. df_morpho_mouse = pd.concat([df_morpho_mouse, tmpdf], axis=1)
  287. tmpdf = dict_df_morpho_human[crop_thres].copy()
  288. tmpdf.columns = [f"R{crop_thres}_"+x for x in tmpdf.columns]
  289. df_morpho_human = pd.concat([df_morpho_human, tmpdf], axis=1)
  290. df_morpho_mouse.dropna(how='any', axis=0, inplace=True)
  291. df_morpho_human.dropna(how='any', axis=0, inplace=True)
  292. print(df_morpho_mouse.shape)
  293. # %%
  294. SCORE_LIST = ["balanced_accuracy","precision","recall","f1_macro","f1_weighted","roc_auc",]
  295. def classification(X, y, algorithm: str,cv=5, n_components=0.95):
  296. n_jobs = -1
  297. X=np.asarray(X)
  298. y=np.asarray(y)
  299. # standardization
  300. scaler = StandardScaler().fit(X)
  301. X_standard = scaler.transform(X)
  302. if n_components==None:
  303. X_standard_pca = X_standard.copy()
  304. else:
  305. pca = PCA(n_components)
  306. X_standard_pca = pca.fit_transform(X_standard)
  307. enc = LabelEncoder().fit(np.unique(y))
  308. y_numeric = enc.transform(y)
  309. # shuffle
  310. training_data = np.hstack((X_standard_pca, y_numeric.reshape(-1,1)))
  311. np.random.seed(821)
  312. np.random.shuffle(training_data)
  313. X = training_data[:, :-1]
  314. y = training_data[:, -1]
  315. if algorithm=='RF':
  316. classifier=RandomForestClassifier(n_jobs=n_jobs,class_weight=None,random_state=821)
  317. elif algorithm=='DT':
  318. classifier=DecisionTreeClassifier(class_weight='balanced',random_state=821)
  319. elif algorithm=='LR':
  320. classifier=LogisticRegression(n_jobs=n_jobs,class_weight='balanced',random_state=821)
  321. elif algorithm=='QDA':
  322. classifier=QuadraticDiscriminantAnalysis()
  323. elif algorithm=='KNN':
  324. classifier=KNeighborsClassifier(n_neighbors=5, weights='uniform',
  325. algorithm="auto", leaf_size=30, n_jobs=n_jobs)
  326. elif algorithm=='SVM':
  327. classifier=SVC(random_state=821)
  328. else:
  329. raise ValueError(f'no model: {algorithm}')
  330. if algorithm=='KNN':
  331. loo = LeaveOneOut()
  332. count=0
  333. y_true=[]
  334. y_pred=[]
  335. for train,test in loo.split(X):
  336. count+=1
  337. classifier.fit(X[train],y[train])
  338. result = classifier.predict(X[test])
  339. y_true.append(y[test][0])
  340. y_pred.append(result[0])
  341. cm = confusion_matrix(y_true, y_pred)
  342. cm_normalized = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]
  343. # print("accuracy:",metrics.accuracy_score(y_true,y_pred))
  344. print("balanced_accuracy:",metrics.balanced_accuracy_score(y_true,y_pred))
  345. # print("F1-score:",metrics.f1_score(y_true,y_pred,average ='macro'))
  346. else:
  347. kf = StratifiedKFold(n_splits=cv,shuffle=True,random_state=821)
  348. params = {
  349. 'n_estimators': [100, 200, 300, 400, 500],
  350. 'min_samples_split' :[2,3,4],
  351. 'max_depth': [4,5],
  352. 'max_features': ['sqrt'],
  353. 'class_weight': ['balanced'],
  354. 'random_state': [821]
  355. }
  356. grid_rf = GridSearchCV(classifier, param_grid=params, cv=kf, n_jobs=n_jobs, refit=False,
  357. scoring='balanced_accuracy', return_train_score=True).fit(X, y)
  358. print('Best parameters:', grid_rf.best_params_)
  359. mean_train_score = grid_rf.cv_results_['mean_train_score'][grid_rf.best_index_]
  360. std_train_score = grid_rf.cv_results_['std_train_score'][grid_rf.best_index_]
  361. mean_test_score = grid_rf.cv_results_['mean_test_score'][grid_rf.best_index_]
  362. std_test_score = grid_rf.cv_results_['std_test_score'][grid_rf.best_index_]
  363. print(f"train:{mean_train_score:.3f}±{std_train_score:.3f}",f"test:{mean_test_score:.3f}±{std_test_score:.3f}")
  364. classifier = type(classifier)(**grid_rf.best_params_)
  365. # classifier = grid_rf.best_estimator_
  366. cv = cross_validate(classifier, X, y, cv=kf, n_jobs=n_jobs, scoring=SCORE_LIST, return_train_score=True)
  367. # # print(f"{cv['test_score'].mean():.3f} +/- {cv['test_score'].std():.3f}")
  368. return [cv[f'test_{x}'].mean() for x in SCORE_LIST],[cv[f'test_{x}'].std() for x in SCORE_LIST], [cv[f'train_{x}'].mean() for x in SCORE_LIST],[cv[f'train_{x}'].std() for x in SCORE_LIST]
  369. return None
  370. # %%
  371. homo_ct_list = ['FL','PL']
  372. from sklearn.preprocessing import LabelEncoder
  373. # homo_ct_list = list(human_ct_lut_tmp.keys())
  374. # preparation for data
  375. tmpindice = df_morpho_mouse.index
  376. tmpindice = np.intersect1d(tmpindice, df_ct_mouse[df_ct_mouse['homologous'].isin(homo_ct_list)].index)
  377. df_morpho_mouse_homo = df_morpho_mouse.copy().loc[tmpindice]
  378. df_morpho_mouse_homo['homologous'] = df_ct_mouse.loc[tmpindice]['homologous']
  379. tmpindice = df_morpho_human.index
  380. tmpindice = np.intersect1d(tmpindice, df_ct_human[df_ct_human['homologous'].isin(homo_ct_list)].index)
  381. df_morpho_human_homo = df_morpho_human.copy().loc[tmpindice]
  382. df_morpho_human_homo['homologous'] = df_ct_human.loc[tmpindice]['homologous']
  383. # standardization
  384. arr_morpho_mouse = df_morpho_mouse_homo.iloc[:,:-1].values.copy()
  385. stand_mean_mouse = arr_morpho_mouse.mean(axis=0)
  386. stand_std_mouse = arr_morpho_mouse.std(axis=0)
  387. arr_morpho_stand_mouse = (arr_morpho_mouse - stand_mean_mouse) / stand_std_mouse
  388. arr_morpho_human = df_morpho_human_homo.iloc[:,:-1].values.copy()
  389. stand_mean_human = arr_morpho_human.mean(axis=0)
  390. stand_std_human = arr_morpho_human.std(axis=0)
  391. arr_morpho_stand_human = (arr_morpho_human - stand_mean_human) / stand_std_human
  392. mouse_drop_col = df_morpho_mouse_homo.columns[:-1][np.isnan(arr_morpho_stand_mouse).all(axis=0)]
  393. human_drop_col = df_morpho_human_homo.columns[:-1][np.isnan(arr_morpho_stand_human).all(axis=0)]
  394. df_morpho_mouse_homo = df_morpho_mouse_homo.drop(columns=mouse_drop_col)
  395. df_morpho_human_homo = df_morpho_human_homo.drop(columns=human_drop_col)
  396. arr_morpho_stand_mouse = arr_morpho_stand_mouse[:,~np.isnan(arr_morpho_stand_mouse).all(axis=0)]
  397. arr_morpho_stand_human = arr_morpho_stand_human[:,~np.isnan(arr_morpho_stand_human).all(axis=0)]
  398. print(homo_ct_list, 'human shape:', arr_morpho_stand_human.shape, 'mouse shape:', arr_morpho_stand_mouse.shape)
  399. print("human", df_morpho_human_homo['homologous'].value_counts())
  400. print("mouse", df_morpho_mouse_homo['homologous'].value_counts())
  401. # %%
  402. def js_divergence(arr1, arr2):
  403. def kl_divergence(mu1, cov1, mu2, cov2):
  404. # Ensure the covariance matrices are invertible
  405. cov2_inv = np.linalg.inv(cov2)
  406. # Dimensionality of the data
  407. k = mu1.shape[0]
  408. # Compute the KL divergence
  409. kl_div = 0.5 * (np.log(np.linalg.det(cov2) / np.linalg.det(cov1))
  410. - k
  411. + np.trace(cov2_inv @ cov1)
  412. + (mu2 - mu1).T @ cov2_inv @ (mu2 - mu1))
  413. return kl_div
  414. # Compute means and covariances for both distributions
  415. mu1 = np.mean(arr1, axis=0)
  416. mu2 = np.mean(arr2, axis=0)
  417. cov1 = np.cov(arr1.T)
  418. cov2 = np.cov(arr2.T)
  419. # Compute the midpoint distribution's parameters
  420. mu_m = 0.5 * (mu1 + mu2)
  421. cov_m = 0.5 * (cov1 + cov2)
  422. # Calculate the JS divergence
  423. js_div = 0.5 * kl_divergence(mu1, cov1, mu_m, cov_m) + 0.5 * kl_divergence(mu2, cov2, mu_m, cov_m)
  424. return js_div
  425. print('Human')
  426. tmppca = PCA(0.99)
  427. tmpdf = df_morpho_human_homo[df_morpho_human_homo['homologous'].isin(homo_ct_list)]
  428. tmparr = tmpdf.iloc[:,:-1]
  429. tmparr = (tmparr-tmparr.mean())/tmparr.std()
  430. tmparr.dropna(axis=1,inplace=True)
  431. print(js_divergence(tmparr[tmpdf['homologous']==homo_ct_list[0]],
  432. tmparr[tmpdf['homologous']==homo_ct_list[1]]))
  433. tmparr_pca = tmppca.fit_transform(tmparr)
  434. print(js_divergence(tmparr_pca[tmpdf['homologous']==homo_ct_list[0]], tmparr_pca[tmpdf['homologous']==homo_ct_list[1]]))
  435. print('Mouse')
  436. tmppca = PCA(0.99)
  437. tmpdf = df_morpho_mouse_homo[df_morpho_mouse_homo['homologous'].isin(homo_ct_list)]
  438. tmparr = tmpdf.iloc[:,:-1]
  439. tmparr = (tmparr-tmparr.mean())/tmparr.std()
  440. tmparr.dropna(axis=1,inplace=True)
  441. print(js_divergence(tmparr[tmpdf['homologous']==homo_ct_list[0]],
  442. tmparr[tmpdf['homologous']==homo_ct_list[1]]))
  443. tmparr_pca = tmppca.fit_transform(tmparr)
  444. print(js_divergence(tmparr_pca[tmpdf['homologous']==homo_ct_list[0]], tmparr_pca[tmpdf['homologous']==homo_ct_list[1]]))
  445. # %%
  446. '''feature selection'''
  447. from mrmr import mrmr_classif
  448. df_morpho_human_homo_stand = df_morpho_human_homo.copy()
  449. df_morpho_human_homo_stand.iloc[:,:-1] = arr_morpho_stand_human
  450. df_human_mrmrdata = df_morpho_human_homo_stand.copy()
  451. df_human_mrmrdata = df_human_mrmrdata[list(df_human_mrmrdata.columns[-1:])+list(df_human_mrmrdata.columns[:-1])]
  452. for i,ct in enumerate(homo_ct_list):
  453. df_human_mrmrdata.loc[df_human_mrmrdata[df_human_mrmrdata['homologous']==ct].index,'homologous']=i+1
  454. for col in df_human_mrmrdata.columns[1:]:
  455. tmparr1=df_human_mrmrdata.loc[df_human_mrmrdata[df_human_mrmrdata['homologous']==1].index,col]
  456. tmparr2=df_human_mrmrdata.loc[df_human_mrmrdata[df_human_mrmrdata['homologous']==2].index,col]
  457. tmparr1,tmparr2 = compute_histograms(tmparr1,tmparr2,outlier='minmax')
  458. df_human_mrmrdata.loc[df_human_mrmrdata[df_human_mrmrdata['homologous']==1].index,col] = tmparr1
  459. df_human_mrmrdata.loc[df_human_mrmrdata[df_human_mrmrdata['homologous']==2].index,col] = tmparr2
  460. selected_feat_human = pymrmr.mRMR(df_human_mrmrdata.astype(int),'MIQ',df_human_mrmrdata.shape[1]-1)
  461. # selected_feat_human = mrmr_classif(X=df_human_mrmrdata.iloc[:,1:].astype(int), y=df_human_mrmrdata.homologous.astype(int), K=40,relevance='f')
  462. selected_feat_human[:5]
  463. # %%
  464. '''feature selection'''
  465. df_morpho_mouse_homo_stand = df_morpho_mouse_homo.copy()
  466. df_morpho_mouse_homo_stand.iloc[:,:-1] = arr_morpho_stand_mouse
  467. df_mouse_mrmrdata = df_morpho_mouse_homo_stand.copy()
  468. df_mouse_mrmrdata = df_mouse_mrmrdata[list(df_mouse_mrmrdata.columns[-1:])+list(df_mouse_mrmrdata.columns[:-1])]
  469. for i,ct in enumerate(homo_ct_list):
  470. df_mouse_mrmrdata.loc[df_mouse_mrmrdata[df_mouse_mrmrdata['homologous']==ct].index,'homologous']=i+1
  471. for col in df_mouse_mrmrdata.columns[1:]:
  472. tmparr1=df_mouse_mrmrdata.loc[df_mouse_mrmrdata[df_mouse_mrmrdata['homologous']==1].index,col]
  473. tmparr2=df_mouse_mrmrdata.loc[df_mouse_mrmrdata[df_mouse_mrmrdata['homologous']==2].index,col]
  474. tmparr1,tmparr2 = compute_histograms(tmparr1,tmparr2,outlier='minmax')
  475. df_mouse_mrmrdata.loc[df_mouse_mrmrdata[df_mouse_mrmrdata['homologous']==1].index,col] = tmparr1
  476. df_mouse_mrmrdata.loc[df_mouse_mrmrdata[df_mouse_mrmrdata['homologous']==2].index,col] = tmparr2
  477. selected_feat_mouse = pymrmr.mRMR(df_mouse_mrmrdata.astype(int),'MIQ',df_mouse_mrmrdata.shape[1]-1)
  478. # selected_feat_mouse = mrmr_classif(X=df_mouse_mrmrdata.iloc[:,1:], y=df_mouse_mrmrdata.homologous, K=40)
  479. selected_feat_mouse[:5]
  480. # %%
  481. ct1,ct2=homo_ct_list
  482. JSD_list_m = []
  483. JSD_list_h = []
  484. for tmpi,feat in enumerate(df_morpho_mouse_homo_stand.columns[:-1]):
  485. tmparr1 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==1]
  486. tmparr2 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==2]
  487. nmi_score,_,_,_,_ = DIFF_feat_ct(tmparr1,tmparr2,method='JSD',stat_test=False)
  488. JSD_list_m.append(nmi_score)
  489. for tmpi,feat in enumerate(df_morpho_human_homo_stand.columns[:-1]):
  490. tmparr1 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==1]
  491. tmparr2 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==2]
  492. nmi_score,_,_,_,_ = DIFF_feat_ct(tmparr1,tmparr2,method='JSD',stat_test=False)
  493. JSD_list_h.append(nmi_score)
  494. JSD_list_m = np.array(JSD_list_m)
  495. JSD_list_h = np.array(JSD_list_h)
  496. xy_max = np.max([np.max(JSD_list_h),np.max(JSD_list_m)])
  497. JSD_list_m = JSD_list_m/xy_max
  498. JSD_list_h = JSD_list_h/xy_max
  499. fig,ax = plt.subplots(1,1,figsize=(3.3,3.3))
  500. ax.set_aspect('equal', adjustable='box')
  501. sns.scatterplot(x=JSD_list_h,y=JSD_list_m,color='black',zorder=11,edgecolor='gray')
  502. ax.plot((0,1),(0,1),color='gray',zorder=12)
  503. ax.grid(ls='--')
  504. ax.set_xlabel('Human-normalized JSD')
  505. ax.set_ylabel('Mouse-normalized JSD')
  506. ax.set_title('All features')
  507. plt.show()
  508. # %%
  509. # df_sourcedata = pd.DataFrame(np.array([JSD_list_h, JSD_list_m]).T, columns=['human','mouse'])
  510. # df_sourcedata['feature'] = df_morpho_human_homo_stand.columns[:-1]
  511. # df_sourcedata.to_csv(rf"..\Tables\source_data\Fig_6_JSD_{homo_ct_list[0]}_{homo_ct_list[1]}.csv", index=False)
  512. # %%
  513. # %%
  514. fig,axs = plt.subplots(1,2,figsize=(3,3),sharey=True)
  515. bias=0.0
  516. alpha=0.5
  517. barlw=2.0
  518. axs[1].invert_yaxis()
  519. selecttop = 5
  520. eps=1e-10
  521. np.random.seed(821)
  522. rand_tmparr1 = np.random.randn(np.sum(df_human_mrmrdata.homologous==1))
  523. np.random.seed(821)
  524. rand_tmparr2 = np.random.randn(np.sum(df_human_mrmrdata.homologous==2))
  525. rand_nmi_score,_1,_2,_,_ = DIFF_feat_ct(rand_tmparr1,rand_tmparr2,method='JSD',stat_test=False)
  526. nmi_score_list_hh = []
  527. nmi_score_list_hm = []
  528. for tmpi,feat in enumerate(selected_feat_human[:selecttop]):
  529. tmparr1 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==1]
  530. tmparr2 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==2]
  531. nmi_score_h,_1,_2,_,_ = DIFF_feat_ct(tmparr1,tmparr2, method='JSD',stat_test=False)
  532. nmi_score_list_hh.append(nmi_score_h)
  533. # print(nmi_score,len(tmparr1),len(tmparr2),_1,_2)
  534. tmparr1 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==1]
  535. tmparr2 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==2]
  536. nmi_score_m,_1,_2,_,_ = DIFF_feat_ct(tmparr1,tmparr2, method='JSD',stat_test=False)
  537. nmi_score_list_hm.append(nmi_score_m)
  538. # print(nmi_score,len(tmparr1),len(tmparr2),_1,_2)
  539. axs[1].barh(tmpi+bias,nmi_score_h/xy_max,zorder=11,color=color2,edgecolor=color2**2,alpha=alpha,lw=barlw,height=0.4)
  540. print(selected_feat_human[:selecttop][tmpi])
  541. # axs[1].text(0,tmpi,selected_feat_human[:selecttop][tmpi],verticalalignment='center',horizontalalignment='left',zorder=13)
  542. # print(nmi_score)
  543. axs[1].axvline(rand_nmi_score/xy_max,ls='--',color='gray',lw=barlw,zorder=12)
  544. axs[1].grid(axis='y',which='both',linestyle='--')
  545. axs[1].spines['top'].set_visible(False)
  546. axs[1].spines['right'].set_visible(False)
  547. # axs[1].set_yticks([])
  548. axs[1].set_ylim(selecttop, -1)
  549. axs[1].set_yticks(np.arange(0,selecttop)+0.425)
  550. axs[1].set_yticklabels([])
  551. axs[1].set_xlim(right=1*1.05)
  552. arrx1list = np.array([])
  553. arrx2list = np.array([])
  554. arry1list = np.array([])
  555. arry2list = np.array([])
  556. ct1,ct2=homo_ct_list
  557. for tmpi, feat in enumerate(selected_feat_human[:selecttop]):
  558. tmparr1 = df_morpho_human_homo_stand[feat].values[df_morpho_human_homo_stand['homologous']==ct1].copy()
  559. tmparr2 = df_morpho_human_homo_stand[feat].values[df_morpho_human_homo_stand['homologous']==ct2].copy()
  560. Q11,Q12,Q13=np.percentile(tmparr1, [25,50,75])
  561. IQR1 = Q13-Q11
  562. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  563. Q21,Q22,Q23=np.percentile(tmparr2, [25,50,75])
  564. IQR2 = Q23-Q21
  565. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  566. lowb = min(lowb1, lowb2)
  567. highb = max(highb1, highb2)
  568. # tmparr1 = tmparr1[(tmparr1>=lowb)&(tmparr1<=highb)]
  569. # tmparr2 = tmparr2[(tmparr2>=lowb)&(tmparr2<=highb)]
  570. tmparr_min = min(np.min(tmparr1),np.min(tmparr2))
  571. tmparr_max = max(np.max(tmparr1),np.max(tmparr2))
  572. tmparr1 = (tmparr1-tmparr_min) / (tmparr_max-tmparr_min+eps)
  573. tmparr2 = (tmparr2-tmparr_min) / (tmparr_max-tmparr_min+eps)
  574. arrx1list = np.hstack([arrx1list,tmparr1])
  575. arrx2list = np.hstack([arrx2list,tmparr2])
  576. arry1list = np.hstack([arry1list,np.array([tmpi]*len(tmparr1))])
  577. arry2list = np.hstack([arry2list,np.array([tmpi]*len(tmparr2))])
  578. g=sns.violinplot(x=np.hstack([arrx1list,arrx2list]),y=np.hstack([arry1list,arry2list]).astype(int),
  579. hue=np.array([ct1]*len(arrx1list)+[ct2]*len(arrx2list)),
  580. split=True,inner='quartile',orient='h',ax=axs[0],cut=0,bw_adjust=1.5,width=0.8,density_norm='width')
  581. ct1_c = homo_color_lut.get(ct1)
  582. ct2_c = homo_color_lut.get(ct2)
  583. for ind, violin in enumerate(g.findobj(matplotlib.collections.PolyCollection)):
  584. if ind % 2 != 0:
  585. violin.set_facecolor(ct1_c)
  586. violin.set_edgecolor(ct1_c**2)
  587. else:
  588. violin.set_facecolor(ct2_c)
  589. violin.set_edgecolor(ct2_c**2)
  590. g.legend_.set_visible(False)
  591. axs[0].spines['top'].set_visible(False)
  592. axs[0].spines['right'].set_visible(False)
  593. axs[0].grid(axis='y',which='both',linestyle='--')
  594. axs[0].set_yticks([])
  595. axs[0].set_xlabel('normalized\ndistribution')
  596. axs[1].set_xlabel('normalized JSD')
  597. max_JSD_value = axs[1].get_xlim()[1]
  598. # max_JSD_value = 1
  599. fig.tight_layout(pad=1.08)
  600. plt.show()
  601. # %%
  602. fig,axs = plt.subplots(1,2,figsize=(3,3),sharey=True)
  603. bias=0.0
  604. alpha=0.5
  605. barlw=2.0
  606. axs[1].invert_yaxis()
  607. selecttop = 5
  608. eps=1e-10
  609. np.random.seed(821)
  610. rand_tmparr1 = np.random.randn(np.sum(df_mouse_mrmrdata.homologous==1))
  611. np.random.seed(821)
  612. rand_tmparr2 = np.random.randn(np.sum(df_mouse_mrmrdata.homologous==2))
  613. rand_nmi_score,_,_,_,_ = DIFF_feat_ct(rand_tmparr1,rand_tmparr2,method='JSD',stat_test=False)
  614. nmi_score_list_mh = []
  615. nmi_score_list_mm = []
  616. for tmpi,feat in enumerate(selected_feat_mouse[:selecttop]):
  617. tmparr1 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==1]
  618. tmparr2 = df_morpho_mouse_homo_stand[feat].values[df_mouse_mrmrdata.homologous==2]
  619. nmi_score_m,_,_,_,_ = DIFF_feat_ct(tmparr1,tmparr2,method='JSD',stat_test=False)
  620. nmi_score_list_mm.append(nmi_score_m)
  621. # print(nmi_score,len(tmparr1),len(tmparr2),_1,_2)
  622. tmparr1 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==1]
  623. tmparr2 = df_morpho_human_homo_stand[feat].values[df_human_mrmrdata.homologous==2]
  624. nmi_score_h,_,_,_,_ = DIFF_feat_ct(tmparr1,tmparr2,method='JSD',stat_test=False)
  625. nmi_score_list_mh.append(nmi_score_h)
  626. # print(nmi_score,len(tmparr1),len(tmparr2),_1,_2)
  627. axs[1].barh(tmpi+bias,nmi_score_m/xy_max,zorder=11,color=color1,edgecolor=color1**2,alpha=alpha,lw=barlw,height=0.4)
  628. print(selected_feat_mouse[:selecttop][tmpi])
  629. # axs[1].text(0,tmpi,selected_feat_mouse[:selecttop][tmpi],verticalalignment='center',horizontalalignment='left',zorder=13)
  630. # print(nmi_score)
  631. axs[1].axvline(rand_nmi_score/xy_max,ls='--',color='gray',lw=barlw,zorder=12)
  632. axs[1].grid(axis='y',which='both',linestyle='--')
  633. axs[1].spines['top'].set_visible(False)
  634. axs[1].spines['right'].set_visible(False)
  635. # axs[1].set_yticks([])
  636. axs[1].set_ylim(selecttop, -1)
  637. axs[1].set_yticks(np.arange(0,selecttop)+0.425)
  638. axs[1].set_yticklabels([])
  639. axs[1].set_xlim(right=max_JSD_value)
  640. arrx1list = np.array([])
  641. arrx2list = np.array([])
  642. arry1list = np.array([])
  643. arry2list = np.array([])
  644. ct1,ct2=homo_ct_list
  645. for tmpi, feat in enumerate(selected_feat_mouse[:selecttop]):
  646. tmparr1 = df_morpho_mouse_homo_stand[feat].values[df_morpho_mouse_homo_stand['homologous']==ct1].copy()
  647. tmparr2 = df_morpho_mouse_homo_stand[feat].values[df_morpho_mouse_homo_stand['homologous']==ct2].copy()
  648. Q11,Q12,Q13=np.percentile(tmparr1, [25,50,75])
  649. IQR1 = Q13-Q11
  650. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  651. Q21,Q22,Q23=np.percentile(tmparr2, [25,50,75])
  652. IQR2 = Q23-Q21
  653. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  654. lowb = min(lowb1, lowb2)
  655. highb = max(highb1, highb2)
  656. # tmparr1 = tmparr1[(tmparr1>=lowb)&(tmparr1<=highb)]
  657. # tmparr2 = tmparr2[(tmparr2>=lowb)&(tmparr2<=highb)]
  658. tmparr_min = min(np.min(tmparr1),np.min(tmparr2))
  659. tmparr_max = max(np.max(tmparr1),np.max(tmparr2))
  660. tmparr1 = (tmparr1-tmparr_min) / (tmparr_max-tmparr_min+eps)
  661. tmparr2 = (tmparr2-tmparr_min) / (tmparr_max-tmparr_min+eps)
  662. arrx1list = np.hstack([arrx1list,tmparr1])
  663. arrx2list = np.hstack([arrx2list,tmparr2])
  664. arry1list = np.hstack([arry1list,np.array([tmpi]*len(tmparr1))])
  665. arry2list = np.hstack([arry2list,np.array([tmpi]*len(tmparr2))])
  666. g=sns.violinplot(x=np.hstack([arrx1list,arrx2list]),y=np.hstack([arry1list,arry2list]).astype(int),
  667. hue=np.array([ct1]*len(arrx1list)+[ct2]*len(arrx2list)),
  668. split=True,inner='quartile',orient='h',ax=axs[0],cut=0,width=0.8,bw_adjust=1.5,density_norm='width')
  669. ct1_c = homo_color_lut.get(ct1)
  670. ct2_c = homo_color_lut.get(ct2)
  671. for ind, violin in enumerate(g.findobj(matplotlib.collections.PolyCollection)):
  672. if ind % 2 != 0:
  673. violin.set_facecolor(ct1_c)
  674. violin.set_edgecolor(ct1_c**2)
  675. else:
  676. violin.set_facecolor(ct2_c)
  677. violin.set_edgecolor(ct2_c**2)
  678. g.legend_.set_visible(False)
  679. axs[0].spines['top'].set_visible(False)
  680. axs[0].spines['right'].set_visible(False)
  681. axs[0].grid(axis='y',which='both',linestyle='--')
  682. axs[0].set_yticks([])
  683. axs[0].set_xlabel('normalized\ndistribution')
  684. axs[1].set_xlabel('normalized JSD')
  685. fig.tight_layout(pad=1.08)
  686. plt.show()
  687. # %%
  688. df_sourcedata = pd.DataFrame()
  689. for tmpi, feat in enumerate(selected_feat_human[:selecttop]):
  690. tmparr1 = df_morpho_human_homo_stand[feat].values[df_morpho_human_homo_stand['homologous']==ct1].copy()
  691. tmparr2 = df_morpho_human_homo_stand[feat].values[df_morpho_human_homo_stand['homologous']==ct2].copy()
  692. Q11,Q12,Q13=np.percentile(tmparr1, [25,50,75])
  693. IQR1 = Q13-Q11
  694. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  695. Q21,Q22,Q23=np.percentile(tmparr2, [25,50,75])
  696. IQR2 = Q23-Q21
  697. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  698. lowb = min(lowb1, lowb2)
  699. highb = max(highb1, highb2)
  700. tmparr_min = min(np.min(tmparr1),np.min(tmparr2))
  701. tmparr_max = max(np.max(tmparr1),np.max(tmparr2))
  702. tmparr1 = (tmparr1-tmparr_min) / (tmparr_max-tmparr_min+eps)
  703. tmparr2 = (tmparr2-tmparr_min) / (tmparr_max-tmparr_min+eps)
  704. df_sourcedata = pd.concat([
  705. df_sourcedata,
  706. pd.DataFrame({
  707. 'feature': [feat]*len(tmparr1),
  708. 'value': tmparr1,
  709. 'homologous': [ct1]*len(tmparr1),
  710. 'species': ['human']*len(tmparr1)
  711. }),
  712. pd.DataFrame({
  713. 'feature': [feat]*len(tmparr2),
  714. 'value': tmparr2,
  715. 'homologous': [ct2]*len(tmparr2),
  716. 'species': ['human']*len(tmparr2)
  717. })
  718. ], ignore_index=True)
  719. for tmpi, feat in enumerate(selected_feat_mouse[:selecttop]):
  720. tmparr1 = df_morpho_mouse_homo_stand[feat].values[df_morpho_mouse_homo_stand['homologous']==ct1].copy()
  721. tmparr2 = df_morpho_mouse_homo_stand[feat].values[df_morpho_mouse_homo_stand['homologous']==ct2].copy()
  722. Q11,Q12,Q13=np.percentile(tmparr1, [25,50,75])
  723. IQR1 = Q13-Q11
  724. lowb1,highb1 = Q11-1.5*IQR1, Q13+1.5*IQR1
  725. Q21,Q22,Q23=np.percentile(tmparr2, [25,50,75])
  726. IQR2 = Q23-Q21
  727. lowb2,highb2 = Q21-1.5*IQR2, Q23+1.5*IQR2
  728. lowb = min(lowb1, lowb2)
  729. highb = max(highb1, highb2)
  730. tmparr_min = min(np.min(tmparr1),np.min(tmparr2))
  731. tmparr_max = max(np.max(tmparr1),np.max(tmparr2))
  732. tmparr1 = (tmparr1-tmparr_min) / (tmparr_max-tmparr_min+eps)
  733. tmparr2 = (tmparr2-tmparr_min) / (tmparr_max-tmparr_min+eps)
  734. df_sourcedata = pd.concat([
  735. df_sourcedata,
  736. pd.DataFrame({
  737. 'feature': [feat]*len(tmparr1),
  738. 'value': tmparr1,
  739. 'homologous': [ct1]*len(tmparr1),
  740. 'species': ['mouse']*len(tmparr1)
  741. }),
  742. pd.DataFrame({
  743. 'feature': [feat]*len(tmparr2),
  744. 'value': tmparr2,
  745. 'homologous': [ct2]*len(tmparr2),
  746. 'species': ['mouse']*len(tmparr2)
  747. })
  748. ], ignore_index=True)
  749. # df_sourcedata.to_csv(rf"E:\ZhixiYun\Projects\fMOST_atlas\Tables\source_data\Fig_6_distribution_{homo_ct_list[0]}_{homo_ct_list[1]}.csv", index=False)
  750. df_sourcedata
  751. # %%
  752. # xy_max = np.max([np.max(nmi_score_list_hh),np.max(nmi_score_list_hm),np.max(nmi_score_list_mh),np.max(nmi_score_list_mm)])
  753. fig,ax = plt.subplots(1,1,figsize=(3.3,3.3))
  754. ax.set_aspect('equal', adjustable='box')
  755. sns.scatterplot(x=np.array(nmi_score_list_hh)/xy_max,y=np.array(nmi_score_list_hm)/xy_max,color=color2,zorder=11,edgecolor='gray')
  756. ax.plot((0,1),(0,1),color='gray',zorder=12)
  757. ax.grid(ls='--')
  758. ax.set_xlabel('Human-normalized JSD')
  759. ax.set_ylabel('Mouse-normalized JSD')
  760. ax.set_title('Top-5 features')
  761. # fig.savefig(rf'E:\ZhixiYun\Projects\fMOST_atlas\Figures\Homologous\analysis\diff_region_diff\diff_distribution\top5_mRMR\h_top5_features_{ct1}_{ct2}_JSD.svg',
  762. # transparent=True, bbox_inches='tight', dpi=500)
  763. plt.show()
  764. fig,ax = plt.subplots(1,1,figsize=(3.3,3.3))
  765. ax.set_aspect('equal', adjustable='box')
  766. sns.scatterplot(x=np.array(nmi_score_list_mh)/xy_max,y=np.array(nmi_score_list_mm)/xy_max,color=color1,zorder=11,edgecolor='gray')
  767. ax.plot((0,1),(0,1),color='gray',zorder=12)
  768. ax.grid(ls='--')
  769. ax.set_xlabel('Human-normalized JSD')
  770. ax.set_ylabel('Mouse-normalized JSD')
  771. ax.set_title('Top-5 features')
  772. plt.show()
  773. # %%
  774. selected_model = ['RF']
  775. mean_mat_human = np.zeros((len(selected_feat_human),len(selected_model),len(SCORE_LIST)))
  776. std_mat_human = np.zeros((len(selected_feat_human),len(selected_model),len(SCORE_LIST)))
  777. train_mean_mat_human = np.zeros((len(selected_feat_human),len(selected_model),len(SCORE_LIST)))
  778. train_std_mat_human = np.zeros((len(selected_feat_human),len(selected_model),len(SCORE_LIST)))
  779. for i in range(len(selected_feat_human)):
  780. feat = selected_feat_human[i]
  781. for j in range(len(selected_model)):
  782. model = selected_model[j]
  783. # mRMR features
  784. test_mean_,test_std_,train_mean_,train_std_=classification(df_morpho_human_homo[selected_feat_human[:i+1]],df_morpho_human_homo['homologous'],model,n_components=None)
  785. # # PCA features
  786. # test_mean_,test_std_,train_mean_,train_std_=classification(df_morpho_human_homo[selected_feat_human],df_morpho_human_homo['homologous'],model,n_components=i+1)
  787. mean_mat_human[i,j] = test_mean_
  788. std_mat_human[i,j] = test_std_
  789. train_mean_mat_human[i,j] = train_mean_
  790. train_std_mat_human[i,j] = train_std_
  791. # %%
  792. df_sourcedata = pd.concat([
  793. pd.concat([
  794. pd.Series(['human']*train_mean_mat_human.shape[0], name="species"),
  795. pd.Series(['train']*train_mean_mat_human.shape[0], name="stage"),
  796. pd.Series(np.arange(train_mean_mat_human.shape[0])+1, name="feature_num"),
  797. pd.DataFrame(train_mean_mat_human.squeeze(axis=1), columns=SCORE_LIST)
  798. ], axis=1),
  799. pd.concat([
  800. pd.Series(['human']*mean_mat_human.shape[0], name="species"),
  801. pd.Series(['test']*mean_mat_human.shape[0], name="stage"),
  802. pd.Series(np.arange(mean_mat_human.shape[0])+1, name="feature_num"),
  803. pd.DataFrame(mean_mat_human.squeeze(axis=1), columns=SCORE_LIST)
  804. ], axis=1),
  805. ], ignore_index=True)
  806. # df_sourcedata.to_csv(rf"E:\ZhixiYun\Projects\fMOST_atlas\Tables\source_data\Fig_6A_EDF_classification_scores_human_{homo_ct_list[0]}_{homo_ct_list[1]}.csv", index=False)
  807. # %%
  808. selected_model = ['RF']
  809. mean_mat_mouse = np.zeros((len(selected_feat_mouse),len(selected_model),len(SCORE_LIST)))
  810. std_mat_mouse = np.zeros((len(selected_feat_mouse),len(selected_model),len(SCORE_LIST)))
  811. train_mean_mat_mouse = np.zeros((len(selected_feat_mouse),len(selected_model),len(SCORE_LIST)))
  812. train_std_mat_mouse = np.zeros((len(selected_feat_mouse),len(selected_model),len(SCORE_LIST)))
  813. for i in range(len(selected_feat_mouse)):
  814. feat = selected_feat_mouse[i]
  815. for j in range(len(selected_model)):
  816. model = selected_model[j]
  817. # mRMR features
  818. test_mean_,test_std_,train_mean_,train_std_=classification(df_morpho_mouse_homo[selected_feat_mouse[:i+1]],df_morpho_mouse_homo['homologous'],model, n_components=None)
  819. # # PCA features
  820. # test_mean_,test_std_,train_mean_,train_std_=classification(df_morpho_mouse_homo[selected_feat_mouse],df_morpho_mouse_homo['homologous'],model,n_components=i+1)
  821. mean_mat_mouse[i,j] = test_mean_
  822. std_mat_mouse[i,j] = test_std_
  823. train_mean_mat_mouse[i,j] = train_mean_
  824. train_std_mat_mouse[i,j] = train_std_
  825. # %%
  826. df_sourcedata = pd.concat([
  827. pd.concat([
  828. pd.Series(['mouse']*train_mean_mat_mouse.shape[0], name="species"),
  829. pd.Series(['train']*train_mean_mat_mouse.shape[0], name="stage"),
  830. pd.Series(np.arange(train_mean_mat_mouse.shape[0])+1, name="feature_num"),
  831. pd.DataFrame(train_mean_mat_mouse.squeeze(axis=1), columns=SCORE_LIST)
  832. ], axis=1),
  833. pd.concat([
  834. pd.Series(['mouse']*mean_mat_mouse.shape[0], name="species"),
  835. pd.Series(['test']*mean_mat_mouse.shape[0], name="stage"),
  836. pd.Series(np.arange(mean_mat_mouse.shape[0])+1, name="feature_num"),
  837. pd.DataFrame(mean_mat_mouse.squeeze(axis=1), columns=SCORE_LIST)
  838. ], axis=1),
  839. ], ignore_index=True)
  840. # df_sourcedata.to_csv(rf"E:\ZhixiYun\Projects\fMOST_atlas\Tables\source_data\Fig_6A_EDF_classification_scores_mouse_{homo_ct_list[0]}_{homo_ct_list[1]}.csv", index=False)
  841. # %%
  842. t_stat, p_value = wilcoxon(mean_mat_human[::2,0,0].flatten(),mean_mat_mouse[::2,0,0].flatten())
  843. print(t_stat, p_value)
  844. def cohens_d(a, b):
  845. n1,n2 = len(a), len(b)
  846. pooled_std = np.sqrt(((n1-1)*np.std(a, ddof=1)**2 + (n2-1)*np.std(b, ddof=1)**2) / (n1+n2-2))
  847. cohens_d = (np.mean(a) - np.mean(b)) / pooled_std
  848. return cohens_d
  849. cohens_d(mean_mat_human[::2,0,0].flatten(),mean_mat_mouse[::2,0,0].flatten())
  850. # %%
  851. for model_idx,model in enumerate(selected_model):
  852. for score_idx,score in enumerate(SCORE_LIST):
  853. fig, ax = plt.subplots(1,1,figsize=(5,2))
  854. # fig.set_facecolor(np.array([243,255,162])/255.0)
  855. # ax.set_facecolor(np.array([231,251,191])/255.0)
  856. bias=0.0
  857. ax.fill_between(x=np.arange(1,len(mean_mat_mouse)+1,2)-bias, y1=mean_mat_mouse[::2,model_idx,score_idx]-std_mat_mouse[::2,model_idx,score_idx],
  858. y2=mean_mat_mouse[::2,model_idx,score_idx]+std_mat_mouse[::2,model_idx,score_idx],color=color1**0.5,alpha=0.5,lw=0,zorder=10)
  859. ax.plot(np.arange(1,len(mean_mat_mouse)+1,2)-bias, mean_mat_mouse[::2,model_idx,score_idx],color=color1,lw=1,label='mouse',zorder=11)
  860. ax.scatter(np.arange(1,len(mean_mat_mouse)+1,2)-bias, mean_mat_mouse[::2,model_idx,score_idx],color=color1,lw=0,zorder=12,s=25)
  861. ax.fill_between(x=np.arange(1,len(mean_mat_human)+1,2)-bias, y1=mean_mat_human[::2,model_idx,score_idx]-std_mat_human[::2,model_idx,score_idx],
  862. y2=mean_mat_human[::2,model_idx,score_idx]+std_mat_human[::2,model_idx,score_idx],color=color2**0.5,alpha=0.5,lw=0,zorder=10)
  863. ax.plot(np.arange(1,len(mean_mat_human)+1,2)-bias, mean_mat_human[::2,model_idx,score_idx],color=color2,lw=1,label='human',zorder=11)
  864. ax.scatter(np.arange(1,len(mean_mat_human)+1,2)-bias, mean_mat_human[::2,model_idx,score_idx],color=color2,lw=0,zorder=12,s=25)
  865. # # # with train score
  866. # ax.fill_between(x=np.arange(1,len(train_mean_mat_mouse)+1,1)-bias, y1=train_mean_mat_mouse[:,model_idx,score_idx]-train_std_mat_mouse[:,model_idx,score_idx],
  867. # y2=train_mean_mat_mouse[:,model_idx,score_idx]+train_std_mat_mouse[:,model_idx,score_idx],color=color1**0.5,alpha=0.5,lw=0,zorder=10)
  868. # ax.plot(np.arange(1,len(train_mean_mat_mouse)+1,1)-bias, train_mean_mat_mouse[:,model_idx,score_idx],color=color1**1,lw=1,label='mouse_train',zorder=11,ls='--')
  869. # ax.scatter(np.arange(1,len(train_mean_mat_mouse)+1,1)-bias, train_mean_mat_mouse[:,model_idx,score_idx],color=color1**1,lw=0,zorder=12,s=30,marker='^')
  870. # ax.fill_between(x=np.arange(1,len(train_mean_mat_human)+1,1)-bias, y1=train_mean_mat_human[:,model_idx,score_idx]-train_std_mat_human[:,model_idx,score_idx],
  871. # y2=train_mean_mat_human[:,model_idx,score_idx]+train_std_mat_human[:,model_idx,score_idx],color=color2**0.5,alpha=0.5,lw=0,zorder=10)
  872. # ax.plot(np.arange(1,len(train_mean_mat_human)+1,1)-bias, train_mean_mat_human[:,model_idx,score_idx],color=color2**1,lw=1,label='human_train',zorder=11,ls='--')
  873. # ax.scatter(np.arange(1,len(train_mean_mat_human)+1,1)-bias, train_mean_mat_human[:,model_idx,score_idx],color=color2**1,lw=0,zorder=12,s=30,marker='^')
  874. ax.set_title(f"classification of {homo_ct_list[0]}&{homo_ct_list[1]}")
  875. ax.set_ylabel(score)
  876. ax.set_xlabel('feature number')
  877. ax.set_xticks([1]+np.arange(20,len(mean_mat_mouse)+1,20).tolist())
  878. # ax.legend(loc=0).set_zorder(13)
  879. ax.spines['top'].set_visible(False)
  880. ax.spines['right'].set_visible(False)
  881. ax.grid(axis='y',ls='--')
  882. plt.show()
  883. # break
  884. # %%

lobe_separability.ipynb at commit a0cc151, under MIT · at the source

Overview

Authors: Zhixi Yun1,2,3, Wen Ye1,4, Nan Ji5,6,7, Yujin Wang5,8, Jing Rong1, Tian Li5,6,7, Yingxin Li1,3,4, Qi Shen1,2, Ruilong Wang1, Hairuo Zhang1, Hao Xu9, Nan Mo1, Binghua Song1, Xin Chen1, Yi Wang10, Yiying Mai5,6,7, Ming Ding11, Song-Lin Ding12, Lijuan Liu1,4, Liwei Zhang5,6,7,8,13, Hanchuan Peng3,14
14 affiliations
  1. New Cornerstone Science Laboratory, Institute for Brain and Intelligence, Southeast University,Nanjing, China
  2. School of Computer Science and Engineering, Southeast University,Nanjing, China
  3. New Cornerstone Science Laboratory, Institute for Brain and Intelligence, Fudan University,Shanghai, China
  4. School of Biological Science & Medical Engineering, Southeast University,Nanjing, China
  5. Department of Neurosurgery, Beijing Tiantan Hospital, Capital Medical University,Beijing, China
  6. China National Clinical Research Center for Neurological Diseases,Beijing, China
  7. Beijing Key Laboratory of Brain Tumor,Beijing, China
  8. Chinese Institutes for Medical Research, Capital Medical University,Beijing, China
  9. Division of Life Sciences and Medicine, University of Science and Technology of China,Hefei, China
  10. Pediatric Epilepsy Center, Peking University First Hospital,Beijing, China
  11. School of Instrumentation and Optoelectronic Engineering, Beihang University,Beijing, China
  12. Allen Institute for Brain Science,Seattle, WA USA
  13. Beijing Neurosurgical Institute,Beijing, China
  14. Shanghai Academy of Natural Sciences, Fudan University,Shanghai, China
Journal: Nature neuroscience, volume 29, issue 9, pages 2296-2310
Dates: received 10 October 2025; accepted 12 June 2026; published online 14 July 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41593-026-02376-z · PMID 42449132 · PMCID PMC13533831 · OpenAlex W7168298252
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), mouse (organism)
Methods: Preprocessing, Connectivity, Spectral & time-frequency, Statistics, Smoothing, state filtering, decompositions, Machine learning, Physiology & signal measures
Keywords: Computational biology and bioinformatics, Neuroscience
MeSH: Cerebral Cortex*, Dendrites*, Neurons*, Animals, Female, Humans, Male, Mice, Species Specificity (* major topic)
Journal subjects: Resource
Topic: Cell Image Analysis Techniques (Biophysics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: cited by 1 paper (Europe PMC); 85 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

vaa3d/vaa3d_tools

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 475e4ca92d4e51de10f1c05d80cef6615432c087, 25 April 2024
Languages: C/C++ (14643), C++ (3300), C (1444), MATLAB (401), Python (302), Shell (249), JavaScript (227), Perl (28), Java (24), Jupyter (18), Fortran (17), R (12), CUDA (2), SAS (1)
Size: 70,755 files, 20,668 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, tests, documentation
Not found: CITATION.cff, environment file, continuous integration
Tools: NumPy (45 files), Image Processing Toolbox (10 files), Matplotlib (8 files), Keras (7 files), OpenCV (7 files), Statistics and Machine Learning Toolbox (5 files), h5py (2 files), scikit-learn (2 files), SciPy (2 files), Curve Fitting Toolbox (1 file), Signal Processing Toolbox (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
2,000 files

YunZhixi98/Mouse_Human_Comparison

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: a0cc151a00384543a21c95c61479e65112ba6b5d, 14 July 2026
Languages: Jupyter (7), Python (6), R (1)
Size: 17 files, 14 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, 7 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (13 files), pandas (12 files), SciPy (9 files), Matplotlib (8 files), seaborn (6 files), scikit-learn (5 files), anndata (2 files), Scanpy (2 files), ggplot2 (1 file), Seurat (1 file), statsmodels (1 file), tidyverse (1 file), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
16 files

Vaa3D

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: the link answers
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source: github.com/Vaa3D

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41593-026-02376-z.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 2,012 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41593-026-02376-z.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 21 authors, 2 keywords, 9 MeSH terms, 81 references.

Cite

This paper

Yun, Z., Ye, W., Ji, N., Wang, Y., Rong, J., Li, T., Li, Y., Shen, Q., Wang, R., Zhang, H., Xu, H., Mo, N., Song, B., Chen, X., Wang, Y., Mai, Y., Ding, M., Ding, S.-L., Liu, L., . . . Peng, H. (2026). A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions. Nature neuroscience, 29(9), 2296-2310. https://doi.org/10.1038/s41593-026-02376-z

BibTeX

@article{yun2026framework,
author = {Yun, Zhixi and Ye, Wen and Ji, Nan and Wang, Yujin and Rong, Jing and Li, Tian and Li, Yingxin and Shen, Qi and Wang, Ruilong and Zhang, Hairuo and Xu, Hao and Mo, Nan and Song, Binghua and Chen, Xin and Wang, Yi and Mai, Yiying and Ding, Ming and Ding, Song-Lin and Liu, Lijuan and Zhang, Liwei and Peng, Hanchuan},
title = {{A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions}},
journal = {Nature neuroscience},
year = {2026},
month = jul,
volume = {29},
number = {9},
pages = {2296--2310},
publisher = {Nature Portfolio},
issn = {1097-6256},
doi = {10.1038/s41593-026-02376-z},
url = {https://doi.org/10.1038/s41593-026-02376-z},
pmid = {42449132},
pmcid = {PMC13533831}
}

RIS

TY - JOUR
AU - Yun, Zhixi
AU - Ye, Wen
AU - Ji, Nan
AU - Wang, Yujin
AU - Rong, Jing
AU - Li, Tian
AU - Li, Yingxin
AU - Shen, Qi
AU - Wang, Ruilong
AU - Zhang, Hairuo
AU - Xu, Hao
AU - Mo, Nan
AU - Song, Binghua
AU - Chen, Xin
AU - Wang, Yi
AU - Mai, Yiying
AU - Ding, Ming
AU - Ding, Song-Lin
AU - Liu, Lijuan
AU - Zhang, Liwei
AU - Peng, Hanchuan
TI - A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions
T2 - Nature neuroscience
J2 - Nat Neurosci
PY - 2026
DA - 2026/07/14
VL - 29
IS - 9
SP - 2296
EP - 2310
SN - 1097-6256
PB - Nature Portfolio
DO - 10.1038/s41593-026-02376-z
UR - https://doi.org/10.1038/s41593-026-02376-z
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41593-026-02376-z",
"type": "article-journal",
"title": "A framework for comparative analysis of human and mouse cortical neuron dendrites in corresponding brain regions",
"container-title": "Nature neuroscience",
"author": [
{
"family": "Yun",
"given": "Zhixi"
},
{
"family": "Ye",
"given": "Wen"
},
{
"family": "Ji",
"given": "Nan"
},
{
"family": "Wang",
"given": "Yujin"
},
{
"family": "Rong",
"given": "Jing"
},
{
"family": "Li",
"given": "Tian"
},
{
"family": "Li",
"given": "Yingxin"
},
{
"family": "Shen",
"given": "Qi"
},
{
"family": "Wang",
"given": "Ruilong"
},
{
"family": "Zhang",
"given": "Hairuo"
},
{
"family": "Xu",
"given": "Hao"
},
{
"family": "Mo",
"given": "Nan"
},
{
"family": "Song",
"given": "Binghua"
},
{
"family": "Chen",
"given": "Xin"
},
{
"family": "Wang",
"given": "Yi"
},
{
"family": "Mai",
"given": "Yiying"
},
{
"family": "Ding",
"given": "Ming"
},
{
"family": "Ding",
"given": "Song-Lin"
},
{
"family": "Liu",
"given": "Lijuan"
},
{
"family": "Zhang",
"given": "Liwei"
},
{
"family": "Peng",
"given": "Hanchuan"
}
],
"container-title-short": "Nat Neurosci",
"volume": "29",
"issue": "9",
"page": "2296-2310",
"DOI": "10.1038/s41593-026-02376-z",
"PMID": "42449132",
"PMCID": "PMC13533831",
"ISSN": "1097-6256",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41593-026-02376-z",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-72935-2 [code]
Spindle neurons in human cortex possess distinctive firing properties and transcriptomic signatures.
Journal: Nature communications
In common: Curve Fitting Toolbox, UMAP, Scanpy, 12 other tools, 5 references
[2] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: Keras, UMAP, anndata, 12 other tools, 4 references
[3] doi:10.1073/pnas.2533168123 [code]
Dendritic morphology and synaptic nonlinearities enhance functional complexity in human cortical neurons.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: h5py, scikit-learn, pandas, 3 other tools, 10 references
[4] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: UMAP, anndata, Scanpy, 10 other tools, mouse, 4 references
[5] doi:10.1186/s13073-026-01704-z [code]
Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.
Journal: Genome medicine
In common: UMAP, anndata, Scanpy, 9 other tools, mouse, 5 references
[6] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: UMAP, anndata, Scanpy, 10 other tools, mouse, 4 references
[7] doi:10.1038/s41556-026-02009-4 [code]
An atlas of primate insular cortex reveals a signal-processing strategy in von Economo neurons.
Journal: Nature cell biology
In common: Scanpy, seaborn, pandas, 3 other tools, 10 references
[8] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: Keras, UMAP, anndata, 11 other tools, 1 reference
[9] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: anndata, Scanpy, Seurat, 9 other tools, 4 references
[10] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Keras, UMAP, Seurat, 12 other tools, mouse

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.