OSCR

Whole-genome duplication shaped cell-type evolution in the vertebrate brain.

Code ↔ Paper

26 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 26 matches · 5 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Clustering and annotation ↔ 0.atlas/5.iterative_clustering/Run.iterative_clustering.R, the whole file · a weak match · score 1.00 · de.score.th, de_param, iter_clust, max.dim, merge_cl, min.genes
  2. [2] § Methods › Identification of cell-type-specific TFs and conserved sets for cell-type families ↔ 3.markers/3.nsforest/find_minimum_combinations_of_TFs.py, lines 39–75 · score 0.96 · BinaryFirst_high, gene_selection, n_binary_genes, n_top_genes, n_trees, nsforesting.NSForest
  3. [3] § Methods › Identification of cell-type-specific TFs and conserved sets for cell-type families ↔ 3.markers/3.nsforest/find_minimum_combinations_of_markers.py, lines 35–71 · score 0.96 · BinaryFirst_high, gene_selection, n_binary_genes, n_top_genes, n_trees, nsforesting.NSForest
  4. [4] § Methods › Cell-type nonspecific dominant expression ↔ 5.paralog_expression_evolution/Dominant expression/s2.check_dominant_within_species_paralogs.R, lines 55–105 · score 0.75 · friedman_test, wilcox_test, paralogue families, Bonferroni, pairwise, dominant
  5. [5] § Methods › Identification of marker genes at the cell-type family and cluster level ↔ 0.atlas/7.final_Bflo/step2.harmony.R, lines 113–183 · score 0.73 · logfc.threshold, min.pct, FindAllMarkers, pos, clusters, genes
  6. [6] § Methods › Atlas integration and cross-species mapping ↔ 0.atlas/6.SAMap/SAMap_run_vertebrate_neurons.py, lines 1–42 · score 0.72 · crossK, hom_edge_mode, Pearson, SAMap, pairwise, Atlas
  7. [7] § Methods › Atlas integration and cross-species mapping ↔ 0.atlas/6.SAMap/SAMap_run_vertebrate_nonneurons.py, lines 1–42 · score 0.72 · crossK, hom_edge_mode, Pearson, SAMap, pairwise, Atlas
  8. [8] § Methods › Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb, lines 84–153 · score 0.68 · vertebrate ancestral states, gene families, loss, binarized, amniote, homologous
  9. [9] § Methods › Clustering and annotation ↔ cytograph/clustering/polished_louvain.py, lines 128–241 · score 0.64 · Louvain clustering, min cells, resolution, iterative, component, max
  10. [10] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/s1.get_matrix.R, the whole file · a weak match · score 0.63 · GABAergic, Glutamatergic neurons, regional genes, De, location, AST
  11. [11] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/s1.get_matrix.R, the whole file · a weak match · score 0.63 · GABAergic, glutamatergic neurons, Regional genes, telencephalon, sum, astrocytes
  12. [12] § Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb, lines 84–153 · score 0.63 · Ancestral state, neo, sub, gene family, Binarized, amniotes
  13. [13] § Methods › Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s3.calculate_divergence_across_species.ipynb, lines 14–30 · score 0.61 · retained orthogroups, OrthoFinder, Gene relationships, copy, SSDs, paralogue
  14. [14] § Methods › Gene regulatory network analysis ↔ 8.GRNs_detection/1.mouse/pyscenic_run.py, lines 28–50 · score 0.60 · motif annotation, target gene, pySCENIC, ranking, mouse, TFs
  15. [15] § Methods › Gene regulatory network analysis ↔ 8.GRNs_detection/1.human/pyscenic_run.py, lines 28–51 · score 0.60 · motif annotation, target gene, pySCENIC, ranking, TFs, Human
  16. [16] § Methods › RNA velocity and multipotency analyses in amphioxus ↔ cytograph/preprocessing/doublet_finder.py, lines 46–173 · score 0.58 · n_neighbors, Dimensionality reduction, variance, log, transformation, preprocessed
  17. [17] § Methods › GO annotations and enrichment analyses ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/plot_and_check_regionalisation.v2.ipynb, lines 324–366 · score 0.57 · enrichGO, db, cutoff, hs, mm, enriched
  18. [18] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/plot_and_check_regionalisation.v2.ipynb, lines 45–99 · score 0.57 · Comparison matrix, regional genes, AST, gradient, log10, Variance
  19. [19] § Methods › Cell-type nonspecific dominant expression ↔ 5.paralog_expression_evolution/Dominant expression/s3.plot_dominate_expression_frequency_for_individual_species.by_pairwise_wilcox.ipynb, lines 110–132 · score 0.55 · dominant copy, wilcox, pairwise, pct, exp, SSD
  20. [20] § Methods › Identifying gene relationships for orthologues, paralogues, ohnologues and SSD paralogues ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s3.calculate_divergence_across_species.ipynb, lines 14–30 · score 0.55 · OrthoFinder, gene relationships, matched, duplicated, orthogroups, WGD
  21. [21] § Methods › GO annotations and enrichment analyses ↔ cytograph/species/species.py, the whole file · a weak match · score 0.54 · Danio rerio, musculus, sapiens, mouse, human, species
  22. [22] § Lasting effect of WGD on cell types ↔ cytograph/species/mouse.py, lines 961–1020 · score 0.54 · Nr2f1, Nr2f2, mice, species
  23. [23] § Methods › Identifying gene relationships for orthologues, paralogues, ohnologues and SSD paralogues ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s4.paralog_switching_calculation.updated.ipynb, lines 12–46 · score 0.54 · OrthoFinder, gene relationships, orthology, duplicated, WGD, paralogues
  24. [24] § Methods › Identification of marker genes at the cell-type family and cluster level ↔ 3.markers/DEGs_Seurat/FindMarkers.R, the whole file · a weak match · score 0.52 · FindAllMarkers, ROC, Seurat, wilcox, genes, cell
  25. [25] § Dosage selection across cell types ↔ 5.paralog_expression_evolution/Dominant expression/s3.plot_dominate_expression_frequency_for_individual_species.by_pairwise_wilcox.ipynb, lines 110–132 · score 0.51 · dominant copy, individual species, insign, pct, exp, SSD
  26. [26] § Linking WGD to cell-type evolution ↔ 10.case_studies/3.amphi_glia_evo_devo/3.subset_glia/subset_glia.ipynb, lines 10–14 · score 0.51 · N4 stage, neural tube, RNA, glia, brain, adult

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 340 lines · 16 KB · GPL-3.0 · 2 matches

  1. # %%
  2. suppressPackageStartupMessages({
  3. require(igraph)
  4. library(dplyr)
  5. library(ggvenn)
  6. library(tidyr)
  7. library(ggpubr)
  8. library(rstatix)
  9. library(stringr)
  10. })
  11. # %%
  12. markers <- readRDS('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/01.data/03.markers/allmarkers.wilcox.FDR001_FC1.5.rds')
  13. # %%
  14. # expressed genes for each species at cell type family level
  15. #markers <- readRDS('exp_domain/combined.exp_tri_score.rds')
  16. # %%
  17. Amniote_homologous_celltypes <- read.delim('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/01.data/02.atlas_final/3.plotting/8.plot_dotplot_cross_species/celltype_ordering.v2.txt', header = F)$V1
  18. Vertebrate_homologous_celltypes <- Amniote_homologous_celltypes[!Amniote_homologous_celltypes %in%
  19. c('Oligodendrocyte precursor cells','Oligodendrocytes')]
  20. # %%
  21. duplicate_pairs <- readRDS("Combined.SSD_WGD.rds")
  22. # %%
  23. orthogroups <- read.delim('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/02.gene_relationships/run4/results/Ortho_pipeline/OrthoFinder/Orthogroups/Orthogroups.tsv')
  24. # at least one copy for 4 species
  25. orthogroups <- orthogroups %>% select(c('Orthogroup', 'Pmar', 'Pvit', 'Mmus', 'Hsap')) %>%
  26. filter(Pmar != '' & Pvit != '' & Mmus != '' & Hsap != '')
  27. # %%
  28. # calculate number of genes by calcualting commas in it
  29. count_commas <- function(x) {
  30. sapply(gregexpr(",", x), function(match) ifelse(match[1] == -1, 0, length(match)))
  31. }
  32. number_genes <- data.frame(apply(orthogroups, c(1,2), count_commas))
  33. # retain orthogroups with only 5 copies max for each four species
  34. orthogroups <- orthogroups[which(number_genes$Pmar <= 5 & number_genes$Pvit <= 5 & number_genes$Mmus <= 5 & number_genes$Hsap <= 5), ]
  35. # %%
  36. dim(orthogroups)
  37. # %%
  38. WGD_genes <- unique(c(duplicate_pairs[duplicate_pairs$type == 'WGD', 'dup1'], duplicate_pairs[duplicate_pairs$type == 'WGD', 'dup2']))
  39. SSD_genes <- unique(c(duplicate_pairs[duplicate_pairs$type == 'SSD', 'dup1'], duplicate_pairs[duplicate_pairs$type == 'SSD', 'dup2']))
  40. paralog_genes <- unique(c(duplicate_pairs[,'dup1'], duplicate_pairs[,'dup2']))
  41. # pay attention! at least a pair (since it's in the same gene family:orthogroup)
  42. # a pair of ohnologues or SSD paralogues
  43. WGD_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
  44. sum(sapply(row, function(cell) {
  45. sum(WGD_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
  46. })) >= 3
  47. }), ]
  48. SSD_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
  49. sum(sapply(row, function(cell) {
  50. sum(SSD_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
  51. })) >= 3
  52. }), ]
  53. Paralog_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
  54. sum(sapply(row, function(cell) {
  55. sum(paralog_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
  56. })) >= 3
  57. }), ]
  58. # %%
  59. #check the number of orthogroups being classified as both WGD_orthogroups and SSD_orthogroups
  60. shared <- semi_join(WGD_orthogroups, SSD_orthogroups, by = colnames(SSD_orthogroups))
  61. dim(Paralog_orthogroups)
  62. dim(WGD_orthogroups)
  63. dim(SSD_orthogroups)
  64. dim(shared)
  65. # %%
  66. get_gene_ID <- function(row, species){
  67. return(paste(species, unlist(strsplit(as.character(row[[species]]), ", ")), sep = '_'))
  68. }
  69. # %%
  70. # for only amniotes
  71. get_results_amniotes <- function(orthogroup){
  72. results <- Reduce(rbind, apply(orthogroup, 1, FUN = function(row){
  73. # get gene IDs for each species
  74. Hsap = get_gene_ID(row, 'Hsap')
  75. Mmus = get_gene_ID(row, 'Mmus')
  76. Pvit = get_gene_ID(row, 'Pvit')
  77. # construct exp matrix
  78. mtx1 = data.frame(matrix(0, nrow = length(c(Hsap, Mmus, Pvit)), ncol = length(Amniote_homologous_celltypes)))
  79. rownames(mtx1) <- c(Hsap, Mmus, Pvit)
  80. colnames(mtx1) <- Amniote_homologous_celltypes
  81. # use markers to binarize mtx
  82. for (g in rownames(mtx1)){
  83. c <- markers[match(g, markers$species_gene), "cluster"]
  84. if (!is.na(c)){
  85. mtx1[g, colnames(mtx1) %in% c] <- 1
  86. }
  87. }
  88. Hsap_tmp <- colSums(mtx1[grepl('Hsap', rownames(mtx1)),])
  89. Mmus_tmp <- colSums(mtx1[grepl('Mmus', rownames(mtx1)),])
  90. Pvit_tmp <- colSums(mtx1[grepl('Pvit', rownames(mtx1)),])
  91. Ancestral_states1 <- ifelse(
  92. (Hsap_tmp > 0 & Mmus_tmp > 0) | (Hsap_tmp > 0 & Pvit_tmp > 0) | (Mmus_tmp > 0 & Pvit_tmp > 0),1,0)
  93. if (sum(Ancestral_states1[Ancestral_states1 == 1]) > 0){
  94. # calculate sub- and neo- for species Human compared to ancestral state of amniotes and lamprey
  95. info_mtx <- sweep(mtx1, 2, Ancestral_states1, FUN = "-")
  96. # calculate sub- number for human (I also considered loss-of-function for all paralogs in that family, in that case, wouldn't be count as sub)
  97. Hsap_sub <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == -1)
  98. Hsap_loss <- (length(Hsap)*sum(colSums(info_mtx[grepl('Hsap', rownames(mtx1)), ]) == length(Hsap)))
  99. Hsap_sub <- Hsap_sub - Hsap_loss
  100. # calculate neo- number for human
  101. Hsap_neo <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == 1)
  102. # calculate sub- number for mouse
  103. Mmus_sub <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == -1)
  104. Mmus_loss <- (length(Mmus)*sum(colSums(info_mtx[grepl('Mmus', rownames(mtx1)), ]) == length(Mmus)))
  105. Mmus_sub <- Mmus_sub - Mmus_loss
  106. # calculate neo- number for mouse
  107. Mmus_neo <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == 1)
  108. # calculate sub- number for lizard
  109. Pvit_sub <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == -1)
  110. Pvit_loss <- (length(Pvit)*sum(colSums(info_mtx[grepl('Pvit', rownames(mtx1)), ]) == length(Pvit)))
  111. Pvit_sub <- Pvit_sub - Pvit_loss
  112. # calculate neo- number for lizard
  113. Pvit_neo <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == 1)
  114. info <- rbind(c(row[['Orthogroup']], Hsap_sub, Hsap_sub/length(grepl('Hsap', rownames(mtx1))), "sub", "Hsap"),
  115. c(row[['Orthogroup']], Hsap_neo, Hsap_neo/length(grepl('Hsap', rownames(mtx1))), "neo", "Hsap"),
  116. c(row[['Orthogroup']], Hsap_loss, Hsap_loss/length(grepl('Hsap', rownames(mtx1))), "loss", "Hsap"),
  117. c(row[['Orthogroup']], Mmus_sub, Mmus_sub/length(grepl('Mmus', rownames(mtx1))), "sub", "Mmus"),
  118. c(row[['Orthogroup']], Mmus_neo, Mmus_neo/length(grepl('Mmus', rownames(mtx1))), "neo", "Mmus"),
  119. c(row[['Orthogroup']], Mmus_loss, Mmus_loss/length(grepl('Mmus', rownames(mtx1))), "loss", "Mmus"),
  120. c(row[['Orthogroup']], Pvit_sub, Pvit_sub/length(grepl('Pvit', rownames(mtx1))), "sub", "Pvit"),
  121. c(row[['Orthogroup']], Pvit_neo, Pvit_neo/length(grepl('Pvit', rownames(mtx1))), "neo", "Pvit"),
  122. c(row[['Orthogroup']], Pvit_loss, Pvit_loss/length(grepl('Pvit', rownames(mtx1))), "loss", "Pvit"))
  123. } else {
  124. info <- NULL
  125. #info_mtx <- NULL
  126. }
  127. return(info)
  128. }))
  129. results <- data.frame(results)
  130. colnames(results) <- c("orthogroups", "number", "ratio", "fun", "species")
  131. results$ratio <- as.double(results$ratio)
  132. results$number <- as.numeric(results$number)
  133. results$species <- factor(results$species, levels = c('Hsap', 'Mmus', 'Pvit'))
  134. return(results)
  135. }
  136. # %%
  137. # for all vertebrates, if 3/4 used as marker then 1
  138. get_results_vertebrates <- function(orthogroup){
  139. results <- Reduce(rbind, apply(orthogroup, 1, FUN = function(row){
  140. # get gene IDs for each species
  141. Hsap = get_gene_ID(row, 'Hsap')
  142. Mmus = get_gene_ID(row, 'Mmus')
  143. Pvit = get_gene_ID(row, 'Pvit')
  144. Pmar = get_gene_ID(row, 'Pmar')
  145. # construct exp matrix
  146. mtx1 = data.frame(matrix(0, nrow = length(c(Hsap, Mmus, Pvit, Pmar)), ncol = length(Vertebrate_homologous_celltypes)))
  147. rownames(mtx1) <- c(Hsap, Mmus, Pvit, Pmar)
  148. colnames(mtx1) <- Vertebrate_homologous_celltypes
  149. # use markers to binarize mtx
  150. for (g in rownames(mtx1)){
  151. c <- markers[match(g, markers$species_gene), "cluster"]
  152. if (!is.na(c)){
  153. mtx1[g, colnames(mtx1) %in% c] <- 1
  154. }
  155. }
  156. Hsap_tmp <- colSums(mtx1[grepl('Hsap', rownames(mtx1)),])
  157. Mmus_tmp <- colSums(mtx1[grepl('Mmus', rownames(mtx1)),])
  158. Pvit_tmp <- colSums(mtx1[grepl('Pvit', rownames(mtx1)),])
  159. Pmar_tmp <- colSums(mtx1[grepl('Pmar', rownames(mtx1)),])
  160. Ancestral_states1 <- ifelse(
  161. (Hsap_tmp > 0 & Mmus_tmp > 0 & Pvit_tmp > 0) | (Hsap_tmp > 0 & Mmus_tmp > 0 & Pmar_tmp > 0) |
  162. (Hsap_tmp > 0 & Pmar_tmp > 0 & Pvit_tmp > 0) | (Mmus_tmp > 0 & Pvit_tmp > 0 & Pmar_tmp > 0),1,0)
  163. if (sum(Ancestral_states1[Ancestral_states1 == 1]) > 0){
  164. # calculate sub- and neo- for species Human compared to ancestral state of amniotes and lamprey
  165. info_mtx <- sweep(mtx1, 2, Ancestral_states1, FUN = "-")
  166. # calculate sub- number for human
  167. Hsap_sub <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == -1)
  168. Hsap_loss <- (length(Hsap)*sum(colSums(info_mtx[grepl('Hsap', rownames(mtx1)), ]) == length(Hsap)))
  169. Hsap_sub <- Hsap_sub - Hsap_loss
  170. # calculate neo- number for human
  171. Hsap_neo <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == 1)
  172. # calculate sub- number for mouse
  173. Mmus_sub <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == -1)
  174. Mmus_loss <- (length(Mmus)*sum(colSums(info_mtx[grepl('Mmus', rownames(mtx1)), ]) == length(Mmus)))
  175. Mmus_sub <- Mmus_sub - Mmus_loss
  176. # calculate neo- number for mouse
  177. Mmus_neo <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == 1)
  178. # calculate sub- number for lizard
  179. Pvit_sub <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == -1)
  180. Pvit_loss <- (length(Pvit)*sum(colSums(info_mtx[grepl('Pvit', rownames(mtx1)), ]) == length(Pvit)))
  181. Pvit_sub <- Pvit_sub - Pvit_loss
  182. # calculate neo- number for lizard
  183. Pvit_neo <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == 1)
  184. # calculate sub- number for lamprey
  185. Pmar_sub <- sum(info_mtx[grepl('Pmar', rownames(mtx1)), ] == -1)
  186. Pmar_loss <- (length(Pmar)*sum(colSums(info_mtx[grepl('Pmar', rownames(mtx1)), ]) == length(Pmar)))
  187. Pmar_sub <- Pmar_sub - Pmar_loss
  188. # calculate neo- number for lamprey
  189. Pmar_neo <- sum(info_mtx[grepl('Pmar', rownames(mtx1)), ] == 1)
  190. info <- rbind(c(row[['Orthogroup']], Hsap_sub, Hsap_sub/length(Hsap), "sub", "Hsap"),
  191. c(row[['Orthogroup']], Hsap_neo, Hsap_neo/length(Hsap), "neo", "Hsap"),
  192. c(row[['Orthogroup']], Hsap_loss, Hsap_loss/length(Hsap), "loss", "Hsap"),
  193. c(row[['Orthogroup']], Mmus_sub, Mmus_sub/length(Mmus), "sub", "Mmus"),
  194. c(row[['Orthogroup']], Mmus_neo, Mmus_neo/length(Mmus), "neo", "Mmus"),
  195. c(row[['Orthogroup']], Mmus_loss, Mmus_loss/length(Mmus), "loss", "Mmus"),
  196. c(row[['Orthogroup']], Pvit_sub, Pvit_sub/length(Pvit), "sub", "Pvit"),
  197. c(row[['Orthogroup']], Pvit_neo, Pvit_neo/length(Pvit), "neo", "Pvit"),
  198. c(row[['Orthogroup']], Pvit_loss, Pvit_loss/length(Pvit), "loss", "Pvit"),
  199. c(row[['Orthogroup']], Pmar_sub, Pmar_sub/length(Pmar), "sub", "Pmar"),
  200. c(row[['Orthogroup']], Pmar_neo, Pmar_neo/length(Pmar), "neo", "Pmar"),
  201. c(row[['Orthogroup']], Pmar_loss, Pmar_loss/length(Pmar), "loss", "Pmar"))
  202. } else {
  203. info <- NULL
  204. #info_mtx <- NULL
  205. }
  206. return(info)
  207. }))
  208. results <- data.frame(results)
  209. colnames(results) <- c("orthogroups", "number", "ratio", "fun", "species")
  210. results$ratio <- as.double(results$ratio)
  211. results$number <- as.numeric(results$number)
  212. results$species <- factor(results$species, levels = c('Hsap', 'Mmus', 'Pvit', 'Pmar'))
  213. return(results)
  214. }
  215. # %%
  216. WGD_res_vertebrates <- get_results_vertebrates(WGD_orthogroups)
  217. SSD_res_vertebrates <- get_results_vertebrates(SSD_orthogroups)
  218. Paralog_res_vertebrates <- get_results_vertebrates(Paralog_orthogroups)
  219. # %%
  220. WGD_res_vertebrates$type <- 'WGD'
  221. SSD_res_vertebrates$type <- 'SSD'
  222. results <- rbind(WGD_res_vertebrates, SSD_res_vertebrates)
  223. results$type <- factor(results$type, levels = c('WGD', 'SSD'))
  224. # %%
  225. head(results)
  226. # %%
  227. # print the total number of changes explained by each way
  228. a1 = sum(results[results$fun == "sub" & results$type == 'WGD', "number"])
  229. a2 = sum(results[results$fun == "neo" & results$type == 'WGD', "number"])
  230. a3 = sum(results[results$fun == "loss" & results$type == 'WGD', "number"])
  231. b1 = sum(results[results$fun == "sub" & results$type == 'SSD', "number"])
  232. b2 = sum(results[results$fun == "neo" & results$type == 'SSD', "number"])
  233. b3 = sum(results[results$fun == "loss" & results$type == 'SSD', "number"])
  234. c1 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "sub", "number"])
  235. c2 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "neo", "number"])
  236. c3 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "loss", "number"])
  237. # %%
  238. a1/(a1+a2+a3)
  239. a2/(a1+a2+a3)
  240. a3/(a1+a2+a3)
  241. b1/(b1+b2+b3)
  242. b2/(b1+b2+b3)
  243. b3/(b1+b2+b3)
  244. c1/(c1+c2+c3)
  245. c2/(c1+c2+c3)
  246. c3/(c1+c2+c3)
  247. # %%
  248. results %>% filter(fun == 'neo' & species == 'Hsap') %>% arrange(orthogroups, fun) %>% filter(orthogroups %in% shared$Orthogroup) %>% filter(number > 3)
  249. # %%
  250. sum(table(ops$orthogroups) == 2)
  251. # %%
  252. my_comparisons <- list(c("sub", "neo"))
  253. p1 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "ratio", color = "fun", palette = "npg") +
  254. stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
  255. facet_grid(vars(type), vars(species))
  256. p2 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "number", color = "fun", palette = "npg") +
  257. stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
  258. facet_grid(vars(type), vars(species))
  259. # %%
  260. ggsave("vertebrates.marker.ratio.sub_neo.pdf", p1, width = 5, height = 6)
  261. ggsave("vertebrates.marker.number.sub_neo.pdf", p2, width = 5, height = 6)
  262. #ggsave("vertebrates.exp_Triscore.ratio.sub_neo.pdf", p1, width = 5, height = 6)
  263. #ggsave("vertebrates.exp_Triscore.number.sub_neo.pdf", p2, width = 5, height = 6)
  264. # %%
  265. WGD_res_amniotes <- get_results_amniotes(WGD_orthogroups)
  266. SSD_res_amniotes <- get_results_amniotes(SSD_orthogroups)
  267. WGD_res_amniotes$type <- 'WGD'
  268. SSD_res_amniotes$type <- 'SSD'
  269. results <- rbind(WGD_res_amniotes, SSD_res_amniotes)
  270. results$type <- factor(results$type, levels = c('WGD', 'SSD'))
  271. # %%
  272. my_comparisons <- list(c("sub", "neo"))
  273. p3 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "ratio", color = "fun", palette = "npg") +
  274. stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
  275. facet_grid(vars(type), vars(species))
  276. p4 <- results %>% filter(fun != 'loss')%>% ggboxplot(x = "fun", y = "number", color = "fun", palette = "npg") +
  277. stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
  278. facet_grid(vars(type), vars(species))
  279. ggsave("amniotes.marker.ratio.sub_neo.pdf", p3, width = 5, height = 6)
  280. ggsave("amniotes.marker.number.sub_neo.pdf", p4, width = 5, height = 6)
  281. #ggsave("amniotes.exp_Triscore.ratio.sub_neo.latest.pdf", p3, width = 5, height = 6)
  282. #ggsave("amniotes.exp_Triscore.number.sub_neo.latest.pdf", p4, width = 5, height = 6)
  283. # %%
  284. # %%
  285. # %%
  286. # not included in paper, pls ignore
  287. my_comparisons <- list(c("WGD", "SSD"))
  288. results %>% ggboxplot(x = "type", y = "ratio", color = "type", palette = "npg") +
  289. stat_compare_means(comparisons = my_comparisons, method = "wilcox") +
  290. facet_grid(vars(fun), vars(species))
  291. # %%

s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb at commit 808269e, under GPL-3.0 · at the source

Overview

Authors: Yuanzhen Zhu1, Shuai Zhang2, Jiankai Wei1,3,4, Diego Dolgetta-Garcia1, Katia Jindrich1, Huimin Liu2, Chenggang Shi2,5,6, Rongrong Pan2,7, Yuwei Chen2, Yan Xu2, Qiye Li8,9, Günter P Wagner10,11, Peter W H Holland1, Guang Li2, Sebastian M Shimeld1
  1. Department of Biology, University of Oxford, Oxford, UK
  2. State Key Laboratory of Cellular Stress Biology, School of Life Sciences, Xiamen University, Xiamen, China
  3. Fang Zongxi Center for Marine EvoDevo, MoE Key Laboratory of Marine Genetics and Breeding, College of Marine Life Sciences, Ocean University of China, Qingdao, China
  4. MoE Key Laboratory of Evolution and Marine Biodiversity, Institute of Evolution and Marine Biodiversity, Ocean University of China, Qingdao, China
  5. Mountain Ecological Restoration and Biodiversity Conservation Key Laboratory of Sichuan Province, Chengdu Institute of Biology, Chinese Academy of Sciences, Chengdu, China
  6. China–Croatia Belt and Road Joint Laboratory on Biodiversity and Ecosystem Services, Chengdu Institute of Biology, Chinese Academy of Sciences, Chengdu, China
  7. Department of Basic Medical Sciences, Qiannan Medical College for Nationalities, Duyun, China
  8. State Key Laboratory of Genome and Multi-Omics Technologies and Shenzhen Key Laboratory of Forensics, BGI Research, Shenzhen, China
  9. BGI Research, Wuhan, China
  10. Department of Ecology and Evolutionary Biology, Yale University, New Haven, CT USA
  11. Department of Evolutionary Biology, University of Vienna, Vienna, Austria
Journal: Nature, volume 656, issue 8128, pages 659-669
Dates: received 24 June 2025; accepted 6 May 2026; published online 10 June 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41586-026-10629-x · PMID 42271056 · PMCID PMC13489949 · OpenAlex W7164165509
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), mouse (organism), other (organism), cellular / molecular (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning
Keywords: Molecular evolution, Genomics, Evolutionary developmental biology, Cellular neuroscience
MeSH: Brain*, Evolution, Molecular*, Gene Duplication*, Genome*, Lancelets*, Vertebrates*, Animals, Humans, Lampreys, Lizards, Mice, Phylogeny, Single-Cell Analysis, Transcription Factors, Transcriptome (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Biotechnology and Biological Sciences Research Council (BB/X015203/1)
Citations: not cited yet (Europe PMC); 108 references in the paper

Abstract

The complex brains of vertebrates have more cell types than those of their closest relatives. Whole-genome duplications (WGDs) occurred during early vertebrate evolution1, but it is unclear whether the duplicated genes (ohnologues) facilitated cell-type evolution. Here using brain single-cell transcriptomes from five chordates—human2, mouse3, lizard4, lamprey5 and amphioxus—we report that many cell-type families with conserved core transcription factors in vertebrates do not show one-to-one homology with amphioxus. Moreover, ohnologues, particularly those from the first WGD, were more important than small-scale duplication paralogues for vertebrate cell-type evolution. To explore whether ohnologues are mechanistically important for this process, we predicted ancestral cell-type states and compared them to amphioxus and experimentally investigated macroglia. The findings indicate that ohnologues had a role in early vertebrate cell-type diversification. Moreover, by examining paralogue expression across cell types and species, we show that expression changes were mainly driven by dosage selection and subfunctionalization. We also link ohnologues to cellular diversity at different anatomical and cell-type scales. Our findings demonstrate the importance of WGDs for the evolution of early vertebrate brain complexity and highlight that the resultant ohnologues continued to capacitate cell-type evolution long after they were formed.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 26 matches between paragraphs and lines of code.

DiracZhu1998/WGD2celltype_evolution

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 808269e1c0c13331cf19a21284904ccbb26bf54d, 30 March 2026
Languages: Jupyter (47), R (22), Python (18), Shell (1)
Size: 388 files, 88 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, 47 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (41 files), Seurat (38 files), pandas (23 files), ggpubr (18 files), Scanpy (18 files), ggplot2 (17 files), igraph (17 files), NumPy (13 files), anndata (12 files), Matplotlib (8 files), rstatix (6 files), circlize (4 files), cowplot (4 files), reshape2 (4 files), reticulate (3 files), SciPy (3 files), DESeq2 (2 files), seaborn (2 files), clusterProfiler (1 file), Harmony (1 file), lme4 (1 file), patchwork (1 file), scVelo (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
90 files

fmarletaz/hagfish

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 68842e98e6515a12869c951fc1b82026fa0df103, 14 November 2023
Languages: R (8), Python (8), Jupyter (3), Shell (1), Julia (1)
Size: 22,165 files, 21 scripts
Software Heritage: archived
Found in: the text, “Identifying gene relationships for orthologues, ”
Holds: README, 3 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (8 files), NumPy (4 files), pandas (3 files), ggplot2 (2 files), ggpubr (2 files), pheatmap (2 files), SciPy (2 files), BEDTools (1 file), DataFrames.jl (1 file), Distributions.jl (1 file), Matplotlib (1 file), Pingouin (1 file), Plots.jl (1 file), seaborn (1 file), statsmodels (1 file), Turing.jl (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
22 files

linnarsson-lab/adult-human-brain

License: BSD-2-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 2b5aa12dffb8cc1bcb3979c409cb6ac907546b2e, 5 June 2026
Languages: Python (96), Jupyter (17)
Size: 130 files, 113 scripts
Software Heritage: archived
Found in: “Data availability”
Holds: README, license file, environment (setup.cfg, setup.py), 17 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (74 files), SciPy (35 files), Matplotlib (34 files), scikit-learn (26 files), pandas (12 files), NetworkX (8 files), seaborn (7 files), Numba (6 files), igraph (3 files), statsmodels (1 file), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
97 files

figshare 29327111

License: GPL 3.0+
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 6 files
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

Code availability

All code and scripts associated with this analysis are publicly available from the GitHub repository (https://github.com/DiracZhu1998/WGD2celltype_evolution).

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 204 scripts, each with its path and the digest of its content;
  • 26 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

The reference genome, gene models, functional annotations of protein-coding genes, full marker gene list of each cell cluster and important intermediate files are deposited in a Figshare repository (10.6084/m9.figshare.29327111)108. All amphioxus bulk RNA-seq, scRNA-seq and snRNA-seq data, as well as lamprey embryonic neural tube scRNA-seq data, have been deposited in the CNCB database under BioProject accession number PRJCA059549 (https://ngdc.cncb.ac.cn/bioproject/browse/PRJCA059549). Data analysed from publicly available sources are described in the paper, available from GitHub and Figshare repositories and other online resources, and are listed as follows: (1) human brain atlas (human_adult_GRCh38-3.0.0.h5ad) from https://github.com/linnarsson-lab/adult-human-brain; (2) mouse brain atlas (l5_all.loom) from http://mousebrain.org/adolescent/downloads.html; (3) lizard brain atlas (Pogona_vitticeps_cells_Science_2022.rds) from https://datashare.mpcdf.mpg.de/s/WBX59YhJKKebizb#editor; (4) lamprey brain atlases (adult_diencephalon.rds, adult_rhombencephalon.rds, adult_mesencephalon.rds, adult_telencephalon.rds and lamprey_adult_whole_brain.rds) from https://downloads.kaessmannlab.org/lamprey/; (5) human eye and lung scRNA, preprocessed by the Human Cell Atlas, from https://datasets.cellxgene.cziscience.com/64175889-d600-4b58-97ea-e74be80206e5.rds and https://datasets.cellxgene.cziscience.com/b351804c-293e-4aeb-9c4c-043db67f4540.rds; (6) three species of CN datasets (GSM4873765_mouse_data.RData.gz, GSM4873766_human_data.RData.gz and GSM4873767_chicken_data.RData.gz) from https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE160471; and (7) Ohnologs (v.2) database (http://ohnologs.curie.fr).

All code and scripts associated with this analysis are publicly available from the GitHub repository (https://github.com/DiracZhu1998/WGD2celltype_evolution).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 15 authors, 4 keywords, 15 MeSH terms, 1 funder, 100 references.

Cite

This paper

Zhu, Y., Zhang, S., Wei, J., Dolgetta-Garcia, D., Jindrich, K., Liu, H., Shi, C., Pan, R., Chen, Y., Xu, Y., Li, Q., Wagner, G. P., Holland, P. W. H., Li, G., & Shimeld, S. M. (2026). Whole-genome duplication shaped cell-type evolution in the vertebrate brain. Nature, 656(8128), 659-669. https://doi.org/10.1038/s41586-026-10629-x

BibTeX

@article{zhu2026whole,
author = {Zhu, Yuanzhen and Zhang, Shuai and Wei, Jiankai and Dolgetta-Garcia, Diego and Jindrich, Katia and Liu, Huimin and Shi, Chenggang and Pan, Rongrong and Chen, Yuwei and Xu, Yan and Li, Qiye and Wagner, Günter P and Holland, Peter W H and Li, Guang and Shimeld, Sebastian M},
title = {{Whole-genome duplication shaped cell-type evolution in the vertebrate brain}},
journal = {Nature},
year = {2026},
month = jun,
volume = {656},
number = {8128},
pages = {659--669},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/s41586-026-10629-x},
url = {https://doi.org/10.1038/s41586-026-10629-x},
pmid = {42271056},
pmcid = {PMC13489949}
}

RIS

TY - JOUR
AU - Zhu, Yuanzhen
AU - Zhang, Shuai
AU - Wei, Jiankai
AU - Dolgetta-Garcia, Diego
AU - Jindrich, Katia
AU - Liu, Huimin
AU - Shi, Chenggang
AU - Pan, Rongrong
AU - Chen, Yuwei
AU - Xu, Yan
AU - Li, Qiye
AU - Wagner, Günter P
AU - Holland, Peter W H
AU - Li, Guang
AU - Shimeld, Sebastian M
TI - Whole-genome duplication shaped cell-type evolution in the vertebrate brain
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/06/10
VL - 656
IS - 8128
SP - 659
EP - 669
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/s41586-026-10629-x
UR - https://doi.org/10.1038/s41586-026-10629-x
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41586-026-10629-x",
"type": "article-journal",
"title": "Whole-genome duplication shaped cell-type evolution in the vertebrate brain",
"container-title": "Nature",
"author": [
{
"family": "Zhu",
"given": "Yuanzhen"
},
{
"family": "Zhang",
"given": "Shuai"
},
{
"family": "Wei",
"given": "Jiankai"
},
{
"family": "Dolgetta-Garcia",
"given": "Diego"
},
{
"family": "Jindrich",
"given": "Katia"
},
{
"family": "Liu",
"given": "Huimin"
},
{
"family": "Shi",
"given": "Chenggang"
},
{
"family": "Pan",
"given": "Rongrong"
},
{
"family": "Chen",
"given": "Yuwei"
},
{
"family": "Xu",
"given": "Yan"
},
{
"family": "Li",
"given": "Qiye"
},
{
"family": "Wagner",
"given": "Günter P"
},
{
"family": "Holland",
"given": "Peter W H"
},
{
"family": "Li",
"given": "Guang"
},
{
"family": "Shimeld",
"given": "Sebastian M"
}
],
"container-title-short": "Nature",
"volume": "656",
"issue": "8128",
"page": "659-669",
"DOI": "10.1038/s41586-026-10629-x",
"PMID": "42271056",
"PMCID": "PMC13489949",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41586-026-10629-x",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: scVelo, Harmony, UMAP, 23 other tools, genetics / omics, cellular / molecular, 6 references
[2] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Turing.jl, DataFrames.jl, Plots.jl, 24 other tools, mouse, cellular / molecular
[3] doi:10.1126/sciadv.aed2952 [code]
Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.
Journal: Science advances
In common: scVelo, BEDTools, UMAP, 16 other tools, cellular / molecular, 7 references
[4] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: reticulate, rstatix, UMAP, 19 other tools, genetics / omics, 4 references
[5] doi:10.1016/j.celrep.2026.117073 [code]
Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.
Journal: Cell reports
In common: BEDTools, reticulate, Pingouin, 19 other tools, genetics / omics, mouse, cellular / molecular, 2 references
[6] doi:10.1126/sciadv.aeg3223 [code]
The extreme diversity of retinal amacrine cells has deep evolutionary roots.
Journal: Science advances
In common: reticulate, rstatix, anndata, 17 other tools, genetics / omics, cellular / molecular, 3 references
[7] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: BEDTools, reticulate, rstatix, 19 other tools, genetics / omics, 2 references
[8] doi:10.1038/s41586-026-10214-2 [code]
Multidimensional profiling of heterogeneity in supratentorial ependymomas.
Journal: Nature
In common: Harmony, reticulate, anndata, 19 other tools, genetics / omics, mouse, 1 reference
[9] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: Harmony, reticulate, UMAP, 19 other tools, genetics / omics, cellular / molecular
[10] doi:10.1038/s42003-026-10034-0 [code]
Region- and cell type-specific changes in gene expression in the cerebellum after classical fear conditioning.
Journal: Communications biology
In common: reticulate, rstatix, anndata, 16 other tools, genetics / omics, mouse, cellular / molecular, 4 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.