Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
The 26 matches · 5 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
- [1] § Methods › Clustering and annotation ↔ 0.atlas/5.iterative_clustering/Run.iterative_clustering.R, the whole file · a weak match · score 1.00 · de.score.th, de_param, iter_clust, max.dim, merge_cl, min.genes
- [2] § Methods › Identification of cell-type-specific TFs and conserved sets for cell-type families ↔ 3.markers/3.nsforest/find_minimum_combinations_of_TFs.py, lines 39–75 · score 0.96 · BinaryFirst_high, gene_selection, n_binary_genes, n_top_genes, n_trees, nsforesting.NSForest
- [3] § Methods › Identification of cell-type-specific TFs and conserved sets for cell-type families ↔ 3.markers/3.nsforest/find_minimum_combinations_of_markers.py, lines 35–71 · score 0.96 · BinaryFirst_high, gene_selection, n_binary_genes, n_top_genes, n_trees, nsforesting.NSForest
- [4] § Methods › Cell-type nonspecific dominant expression ↔ 5.paralog_expression_evolution/Dominant expression/s2.check_dominant_within_species_paralogs.R, lines 55–105 · score 0.75 · friedman_test, wilcox_test, paralogue families, Bonferroni, pairwise, dominant
- [5] § Methods › Identification of marker genes at the cell-type family and cluster level ↔ 0.atlas/7.final_Bflo/step2.harmony.R, lines 113–183 · score 0.73 · logfc.threshold, min.pct, FindAllMarkers, pos, clusters, genes
- [6] § Methods › Atlas integration and cross-species mapping ↔ 0.atlas/6.SAMap/SAMap_run_vertebrate_neurons.py, lines 1–42 · score 0.72 · crossK, hom_edge_mode, Pearson, SAMap, pairwise, Atlas
- [7] § Methods › Atlas integration and cross-species mapping ↔ 0.atlas/6.SAMap/SAMap_run_vertebrate_nonneurons.py, lines 1–42 · score 0.72 · crossK, hom_edge_mode, Pearson, SAMap, pairwise, Atlas
- [8] § Methods › Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb, lines 84–153 · score 0.68 · vertebrate ancestral states, gene families, loss, binarized, amniote, homologous
- [9] § Methods › Clustering and annotation ↔ cytograph/clustering/polished_louvain.py, lines 128–241 · score 0.64 · Louvain clustering, min cells, resolution, iterative, component, max
- [10] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/s1.get_matrix.R, the whole file · a weak match · score 0.63 · GABAergic, Glutamatergic neurons, regional genes, De, location, AST
- [11] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/s1.get_matrix.R, the whole file · a weak match · score 0.63 · GABAergic, glutamatergic neurons, Regional genes, telencephalon, sum, astrocytes
- [12] § Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb, lines 84–153 · score 0.63 · Ancestral state, neo, sub, gene family, Binarized, amniotes
- [13] § Methods › Subfunctionalization and neofunctionalization ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s3.calculate_divergence_across_species.ipynb, lines 14–30 · score 0.61 · retained orthogroups, OrthoFinder, Gene relationships, copy, SSDs, paralogue
- [14] § Methods › Gene regulatory network analysis ↔ 8.GRNs_detection/1.mouse/pyscenic_run.py, lines 28–50 · score 0.60 · motif annotation, target gene, pySCENIC, ranking, mouse, TFs
- [15] § Methods › Gene regulatory network analysis ↔ 8.GRNs_detection/1.human/pyscenic_run.py, lines 28–51 · score 0.60 · motif annotation, target gene, pySCENIC, ranking, TFs, Human
- [16] § Methods › RNA velocity and multipotency analyses in amphioxus ↔ cytograph/preprocessing/doublet_finder.py, lines 46–173 · score 0.58 · n_neighbors, Dimensionality reduction, variance, log, transformation, preprocessed
- [17] § Methods › GO annotations and enrichment analyses ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/plot_and_check_regionalisation.v2.ipynb, lines 324–366 · score 0.57 · enrichGO, db, cutoff, hs, mm, enriched
- [18] § Genome duplication and regional identity ↔ 10.case_studies/1.macroglia/AST/5.regionalisation_genes/plot_and_check_regionalisation.v2.ipynb, lines 45–99 · score 0.57 · Comparison matrix, regional genes, AST, gradient, log10, Variance
- [19] § Methods › Cell-type nonspecific dominant expression ↔ 5.paralog_expression_evolution/Dominant expression/s3.plot_dominate_expression_frequency_for_individual_species.by_pairwise_wilcox.ipynb, lines 110–132 · score 0.55 · dominant copy, wilcox, pairwise, pct, exp, SSD
- [20] § Methods › Identifying gene relationships for orthologues, paralogues, ohnologues and SSD paralogues ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s3.calculate_divergence_across_species.ipynb, lines 14–30 · score 0.55 · OrthoFinder, gene relationships, matched, duplicated, orthogroups, WGD
- [21] § Methods › GO annotations and enrichment analyses ↔ cytograph/species/species.py, the whole file · a weak match · score 0.54 · Danio rerio, musculus, sapiens, mouse, human, species
- [22] § Lasting effect of WGD on cell types ↔ cytograph/species/mouse.py, lines 961–1020 · score 0.54 · Nr2f1, Nr2f2, mice, species
- [23] § Methods › Identifying gene relationships for orthologues, paralogues, ohnologues and SSD paralogues ↔ 5.paralog_expression_evolution/Paralog shift sub_vs_neo/s4.paralog_switching_calculation.updated.ipynb, lines 12–46 · score 0.54 · OrthoFinder, gene relationships, orthology, duplicated, WGD, paralogues
- [24] § Methods › Identification of marker genes at the cell-type family and cluster level ↔ 3.markers/DEGs_Seurat/FindMarkers.R, the whole file · a weak match · score 0.52 · FindAllMarkers, ROC, Seurat, wilcox, genes, cell
- [25] § Dosage selection across cell types ↔ 5.paralog_expression_evolution/Dominant expression/s3.plot_dominate_expression_frequency_for_individual_species.by_pairwise_wilcox.ipynb, lines 110–132 · score 0.51 · dominant copy, individual species, insign, pct, exp, SSD
- [26] § Linking WGD to cell-type evolution ↔ 10.case_studies/3.amphi_glia_evo_devo/3.subset_glia/subset_glia.ipynb, lines 10–14 · score 0.51 · N4 stage, neural tube, RNA, glia, brain, adult
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 340 lines · 16 KB · GPL-3.0 · 2 matches
- # %%
- suppressPackageStartupMessages({
- require(igraph)
- library(dplyr)
- library(ggvenn)
- library(tidyr)
- library(ggpubr)
- library(rstatix)
- library(stringr)
- })
- # %%
- markers <- readRDS('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/01.data/03.markers/allmarkers.wilcox.FDR001_FC1.5.rds')
- # %%
- # expressed genes for each species at cell type family level
- #markers <- readRDS('exp_domain/combined.exp_tri_score.rds')
- # %%
- Amniote_homologous_celltypes <- read.delim('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/01.data/02.atlas_final/3.plotting/8.plot_dotplot_cross_species/celltype_ordering.v2.txt', header = F)$V1
- Vertebrate_homologous_celltypes <- Amniote_homologous_celltypes[!Amniote_homologous_celltypes %in%
- c('Oligodendrocyte precursor cells','Oligodendrocytes')]
- # %%
- duplicate_pairs <- readRDS("Combined.SSD_WGD.rds")
- # %%
- orthogroups <- read.delim('/mnt/data01/yuanzhen/01.Vertebrate_cell_evo/02.gene_relationships/run4/results/Ortho_pipeline/OrthoFinder/Orthogroups/Orthogroups.tsv')
- # at least one copy for 4 species
- orthogroups <- orthogroups %>% select(c('Orthogroup', 'Pmar', 'Pvit', 'Mmus', 'Hsap')) %>%
- filter(Pmar != '' & Pvit != '' & Mmus != '' & Hsap != '')
- # %%
- # calculate number of genes by calcualting commas in it
- count_commas <- function(x) {
- sapply(gregexpr(",", x), function(match) ifelse(match[1] == -1, 0, length(match)))
- }
- number_genes <- data.frame(apply(orthogroups, c(1,2), count_commas))
- # retain orthogroups with only 5 copies max for each four species
- orthogroups <- orthogroups[which(number_genes$Pmar <= 5 & number_genes$Pvit <= 5 & number_genes$Mmus <= 5 & number_genes$Hsap <= 5), ]
- # %%
- dim(orthogroups)
- # %%
- WGD_genes <- unique(c(duplicate_pairs[duplicate_pairs$type == 'WGD', 'dup1'], duplicate_pairs[duplicate_pairs$type == 'WGD', 'dup2']))
- SSD_genes <- unique(c(duplicate_pairs[duplicate_pairs$type == 'SSD', 'dup1'], duplicate_pairs[duplicate_pairs$type == 'SSD', 'dup2']))
- paralog_genes <- unique(c(duplicate_pairs[,'dup1'], duplicate_pairs[,'dup2']))
- # pay attention! at least a pair (since it's in the same gene family:orthogroup)
- # a pair of ohnologues or SSD paralogues
- WGD_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
- sum(sapply(row, function(cell) {
- sum(WGD_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
- })) >= 3
- }), ]
- SSD_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
- sum(sapply(row, function(cell) {
- sum(SSD_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
- })) >= 3
- }), ]
- Paralog_orthogroups <- orthogroups[apply(orthogroups, 1, function(row) {
- sum(sapply(row, function(cell) {
- sum(paralog_genes %in% unlist(strsplit(as.character(cell), ", "))) >= 2
- })) >= 3
- }), ]
- # %%
- #check the number of orthogroups being classified as both WGD_orthogroups and SSD_orthogroups
- shared <- semi_join(WGD_orthogroups, SSD_orthogroups, by = colnames(SSD_orthogroups))
- dim(Paralog_orthogroups)
- dim(WGD_orthogroups)
- dim(SSD_orthogroups)
- dim(shared)
- # %%
- get_gene_ID <- function(row, species){
- return(paste(species, unlist(strsplit(as.character(row[[species]]), ", ")), sep = '_'))
- }
- # %%
- # for only amniotes
- get_results_amniotes <- function(orthogroup){
- results <- Reduce(rbind, apply(orthogroup, 1, FUN = function(row){
- # get gene IDs for each species
- Hsap = get_gene_ID(row, 'Hsap')
- Mmus = get_gene_ID(row, 'Mmus')
- Pvit = get_gene_ID(row, 'Pvit')
- # construct exp matrix
- mtx1 = data.frame(matrix(0, nrow = length(c(Hsap, Mmus, Pvit)), ncol = length(Amniote_homologous_celltypes)))
- rownames(mtx1) <- c(Hsap, Mmus, Pvit)
- colnames(mtx1) <- Amniote_homologous_celltypes
- # use markers to binarize mtx
- for (g in rownames(mtx1)){
- c <- markers[match(g, markers$species_gene), "cluster"]
- if (!is.na(c)){
- mtx1[g, colnames(mtx1) %in% c] <- 1
- }
- }
- Hsap_tmp <- colSums(mtx1[grepl('Hsap', rownames(mtx1)),])
- Mmus_tmp <- colSums(mtx1[grepl('Mmus', rownames(mtx1)),])
- Pvit_tmp <- colSums(mtx1[grepl('Pvit', rownames(mtx1)),])
- Ancestral_states1 <- ifelse(
- (Hsap_tmp > 0 & Mmus_tmp > 0) | (Hsap_tmp > 0 & Pvit_tmp > 0) | (Mmus_tmp > 0 & Pvit_tmp > 0),1,0)
- if (sum(Ancestral_states1[Ancestral_states1 == 1]) > 0){
- # calculate sub- and neo- for species Human compared to ancestral state of amniotes and lamprey
- info_mtx <- sweep(mtx1, 2, Ancestral_states1, FUN = "-")
- # calculate sub- number for human (I also considered loss-of-function for all paralogs in that family, in that case, wouldn't be count as sub)
- Hsap_sub <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == -1)
- Hsap_loss <- (length(Hsap)*sum(colSums(info_mtx[grepl('Hsap', rownames(mtx1)), ]) == length(Hsap)))
- Hsap_sub <- Hsap_sub - Hsap_loss
- # calculate neo- number for human
- Hsap_neo <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == 1)
- # calculate sub- number for mouse
- Mmus_sub <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == -1)
- Mmus_loss <- (length(Mmus)*sum(colSums(info_mtx[grepl('Mmus', rownames(mtx1)), ]) == length(Mmus)))
- Mmus_sub <- Mmus_sub - Mmus_loss
- # calculate neo- number for mouse
- Mmus_neo <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == 1)
- # calculate sub- number for lizard
- Pvit_sub <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == -1)
- Pvit_loss <- (length(Pvit)*sum(colSums(info_mtx[grepl('Pvit', rownames(mtx1)), ]) == length(Pvit)))
- Pvit_sub <- Pvit_sub - Pvit_loss
- # calculate neo- number for lizard
- Pvit_neo <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == 1)
- info <- rbind(c(row[['Orthogroup']], Hsap_sub, Hsap_sub/length(grepl('Hsap', rownames(mtx1))), "sub", "Hsap"),
- c(row[['Orthogroup']], Hsap_neo, Hsap_neo/length(grepl('Hsap', rownames(mtx1))), "neo", "Hsap"),
- c(row[['Orthogroup']], Hsap_loss, Hsap_loss/length(grepl('Hsap', rownames(mtx1))), "loss", "Hsap"),
- c(row[['Orthogroup']], Mmus_sub, Mmus_sub/length(grepl('Mmus', rownames(mtx1))), "sub", "Mmus"),
- c(row[['Orthogroup']], Mmus_neo, Mmus_neo/length(grepl('Mmus', rownames(mtx1))), "neo", "Mmus"),
- c(row[['Orthogroup']], Mmus_loss, Mmus_loss/length(grepl('Mmus', rownames(mtx1))), "loss", "Mmus"),
- c(row[['Orthogroup']], Pvit_sub, Pvit_sub/length(grepl('Pvit', rownames(mtx1))), "sub", "Pvit"),
- c(row[['Orthogroup']], Pvit_neo, Pvit_neo/length(grepl('Pvit', rownames(mtx1))), "neo", "Pvit"),
- c(row[['Orthogroup']], Pvit_loss, Pvit_loss/length(grepl('Pvit', rownames(mtx1))), "loss", "Pvit"))
- } else {
- info <- NULL
- #info_mtx <- NULL
- }
- return(info)
- }))
- results <- data.frame(results)
- colnames(results) <- c("orthogroups", "number", "ratio", "fun", "species")
- results$ratio <- as.double(results$ratio)
- results$number <- as.numeric(results$number)
- results$species <- factor(results$species, levels = c('Hsap', 'Mmus', 'Pvit'))
- return(results)
- }
- # %%
- # for all vertebrates, if 3/4 used as marker then 1
- get_results_vertebrates <- function(orthogroup){
- results <- Reduce(rbind, apply(orthogroup, 1, FUN = function(row){
- # get gene IDs for each species
- Hsap = get_gene_ID(row, 'Hsap')
- Mmus = get_gene_ID(row, 'Mmus')
- Pvit = get_gene_ID(row, 'Pvit')
- Pmar = get_gene_ID(row, 'Pmar')
- # construct exp matrix
- mtx1 = data.frame(matrix(0, nrow = length(c(Hsap, Mmus, Pvit, Pmar)), ncol = length(Vertebrate_homologous_celltypes)))
- rownames(mtx1) <- c(Hsap, Mmus, Pvit, Pmar)
- colnames(mtx1) <- Vertebrate_homologous_celltypes
- # use markers to binarize mtx
- for (g in rownames(mtx1)){
- c <- markers[match(g, markers$species_gene), "cluster"]
- if (!is.na(c)){
- mtx1[g, colnames(mtx1) %in% c] <- 1
- }
- }
- Hsap_tmp <- colSums(mtx1[grepl('Hsap', rownames(mtx1)),])
- Mmus_tmp <- colSums(mtx1[grepl('Mmus', rownames(mtx1)),])
- Pvit_tmp <- colSums(mtx1[grepl('Pvit', rownames(mtx1)),])
- Pmar_tmp <- colSums(mtx1[grepl('Pmar', rownames(mtx1)),])
- Ancestral_states1 <- ifelse(
- (Hsap_tmp > 0 & Mmus_tmp > 0 & Pvit_tmp > 0) | (Hsap_tmp > 0 & Mmus_tmp > 0 & Pmar_tmp > 0) |
- (Hsap_tmp > 0 & Pmar_tmp > 0 & Pvit_tmp > 0) | (Mmus_tmp > 0 & Pvit_tmp > 0 & Pmar_tmp > 0),1,0)
- if (sum(Ancestral_states1[Ancestral_states1 == 1]) > 0){
- # calculate sub- and neo- for species Human compared to ancestral state of amniotes and lamprey
- info_mtx <- sweep(mtx1, 2, Ancestral_states1, FUN = "-")
- # calculate sub- number for human
- Hsap_sub <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == -1)
- Hsap_loss <- (length(Hsap)*sum(colSums(info_mtx[grepl('Hsap', rownames(mtx1)), ]) == length(Hsap)))
- Hsap_sub <- Hsap_sub - Hsap_loss
- # calculate neo- number for human
- Hsap_neo <- sum(info_mtx[grepl('Hsap', rownames(mtx1)), ] == 1)
- # calculate sub- number for mouse
- Mmus_sub <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == -1)
- Mmus_loss <- (length(Mmus)*sum(colSums(info_mtx[grepl('Mmus', rownames(mtx1)), ]) == length(Mmus)))
- Mmus_sub <- Mmus_sub - Mmus_loss
- # calculate neo- number for mouse
- Mmus_neo <- sum(info_mtx[grepl('Mmus', rownames(mtx1)), ] == 1)
- # calculate sub- number for lizard
- Pvit_sub <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == -1)
- Pvit_loss <- (length(Pvit)*sum(colSums(info_mtx[grepl('Pvit', rownames(mtx1)), ]) == length(Pvit)))
- Pvit_sub <- Pvit_sub - Pvit_loss
- # calculate neo- number for lizard
- Pvit_neo <- sum(info_mtx[grepl('Pvit', rownames(mtx1)), ] == 1)
- # calculate sub- number for lamprey
- Pmar_sub <- sum(info_mtx[grepl('Pmar', rownames(mtx1)), ] == -1)
- Pmar_loss <- (length(Pmar)*sum(colSums(info_mtx[grepl('Pmar', rownames(mtx1)), ]) == length(Pmar)))
- Pmar_sub <- Pmar_sub - Pmar_loss
- # calculate neo- number for lamprey
- Pmar_neo <- sum(info_mtx[grepl('Pmar', rownames(mtx1)), ] == 1)
- info <- rbind(c(row[['Orthogroup']], Hsap_sub, Hsap_sub/length(Hsap), "sub", "Hsap"),
- c(row[['Orthogroup']], Hsap_neo, Hsap_neo/length(Hsap), "neo", "Hsap"),
- c(row[['Orthogroup']], Hsap_loss, Hsap_loss/length(Hsap), "loss", "Hsap"),
- c(row[['Orthogroup']], Mmus_sub, Mmus_sub/length(Mmus), "sub", "Mmus"),
- c(row[['Orthogroup']], Mmus_neo, Mmus_neo/length(Mmus), "neo", "Mmus"),
- c(row[['Orthogroup']], Mmus_loss, Mmus_loss/length(Mmus), "loss", "Mmus"),
- c(row[['Orthogroup']], Pvit_sub, Pvit_sub/length(Pvit), "sub", "Pvit"),
- c(row[['Orthogroup']], Pvit_neo, Pvit_neo/length(Pvit), "neo", "Pvit"),
- c(row[['Orthogroup']], Pvit_loss, Pvit_loss/length(Pvit), "loss", "Pvit"),
- c(row[['Orthogroup']], Pmar_sub, Pmar_sub/length(Pmar), "sub", "Pmar"),
- c(row[['Orthogroup']], Pmar_neo, Pmar_neo/length(Pmar), "neo", "Pmar"),
- c(row[['Orthogroup']], Pmar_loss, Pmar_loss/length(Pmar), "loss", "Pmar"))
- } else {
- info <- NULL
- #info_mtx <- NULL
- }
- return(info)
- }))
- results <- data.frame(results)
- colnames(results) <- c("orthogroups", "number", "ratio", "fun", "species")
- results$ratio <- as.double(results$ratio)
- results$number <- as.numeric(results$number)
- results$species <- factor(results$species, levels = c('Hsap', 'Mmus', 'Pvit', 'Pmar'))
- return(results)
- }
- # %%
- WGD_res_vertebrates <- get_results_vertebrates(WGD_orthogroups)
- SSD_res_vertebrates <- get_results_vertebrates(SSD_orthogroups)
- Paralog_res_vertebrates <- get_results_vertebrates(Paralog_orthogroups)
- # %%
- WGD_res_vertebrates$type <- 'WGD'
- SSD_res_vertebrates$type <- 'SSD'
- results <- rbind(WGD_res_vertebrates, SSD_res_vertebrates)
- results$type <- factor(results$type, levels = c('WGD', 'SSD'))
- # %%
- head(results)
- # %%
- # print the total number of changes explained by each way
- a1 = sum(results[results$fun == "sub" & results$type == 'WGD', "number"])
- a2 = sum(results[results$fun == "neo" & results$type == 'WGD', "number"])
- a3 = sum(results[results$fun == "loss" & results$type == 'WGD', "number"])
- b1 = sum(results[results$fun == "sub" & results$type == 'SSD', "number"])
- b2 = sum(results[results$fun == "neo" & results$type == 'SSD', "number"])
- b3 = sum(results[results$fun == "loss" & results$type == 'SSD', "number"])
- c1 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "sub", "number"])
- c2 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "neo", "number"])
- c3 = sum(Paralog_res_vertebrates[Paralog_res_vertebrates$fun == "loss", "number"])
- # %%
- a1/(a1+a2+a3)
- a2/(a1+a2+a3)
- a3/(a1+a2+a3)
- b1/(b1+b2+b3)
- b2/(b1+b2+b3)
- b3/(b1+b2+b3)
- c1/(c1+c2+c3)
- c2/(c1+c2+c3)
- c3/(c1+c2+c3)
- # %%
- results %>% filter(fun == 'neo' & species == 'Hsap') %>% arrange(orthogroups, fun) %>% filter(orthogroups %in% shared$Orthogroup) %>% filter(number > 3)
- # %%
- sum(table(ops$orthogroups) == 2)
- # %%
- my_comparisons <- list(c("sub", "neo"))
- p1 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "ratio", color = "fun", palette = "npg") +
- stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
- facet_grid(vars(type), vars(species))
- p2 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "number", color = "fun", palette = "npg") +
- stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
- facet_grid(vars(type), vars(species))
- # %%
- ggsave("vertebrates.marker.ratio.sub_neo.pdf", p1, width = 5, height = 6)
- ggsave("vertebrates.marker.number.sub_neo.pdf", p2, width = 5, height = 6)
- #ggsave("vertebrates.exp_Triscore.ratio.sub_neo.pdf", p1, width = 5, height = 6)
- #ggsave("vertebrates.exp_Triscore.number.sub_neo.pdf", p2, width = 5, height = 6)
- # %%
- WGD_res_amniotes <- get_results_amniotes(WGD_orthogroups)
- SSD_res_amniotes <- get_results_amniotes(SSD_orthogroups)
- WGD_res_amniotes$type <- 'WGD'
- SSD_res_amniotes$type <- 'SSD'
- results <- rbind(WGD_res_amniotes, SSD_res_amniotes)
- results$type <- factor(results$type, levels = c('WGD', 'SSD'))
- # %%
- my_comparisons <- list(c("sub", "neo"))
- p3 <- results %>% filter(fun != 'loss') %>% ggboxplot(x = "fun", y = "ratio", color = "fun", palette = "npg") +
- stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
- facet_grid(vars(type), vars(species))
- p4 <- results %>% filter(fun != 'loss')%>% ggboxplot(x = "fun", y = "number", color = "fun", palette = "npg") +
- stat_compare_means(comparisons = my_comparisons, method = "wilcox", paired = TRUE) +
- facet_grid(vars(type), vars(species))
- ggsave("amniotes.marker.ratio.sub_neo.pdf", p3, width = 5, height = 6)
- ggsave("amniotes.marker.number.sub_neo.pdf", p4, width = 5, height = 6)
- #ggsave("amniotes.exp_Triscore.ratio.sub_neo.latest.pdf", p3, width = 5, height = 6)
- #ggsave("amniotes.exp_Triscore.number.sub_neo.latest.pdf", p4, width = 5, height = 6)
- # %%
- # %%
- # %%
- # not included in paper, pls ignore
- my_comparisons <- list(c("WGD", "SSD"))
- results %>% ggboxplot(x = "type", y = "ratio", color = "type", palette = "npg") +
- stat_compare_means(comparisons = my_comparisons, method = "wilcox") +
- facet_grid(vars(fun), vars(species))
- # %%
s2.get_vertebrate_ancestral_states_and_check_sub_neo_loss.latest_gene_family_reconsider.ipynb at commit 808269e, under GPL-3.0 · at the source
Overview
- Department of Biology, University of Oxford, Oxford, UK
- State Key Laboratory of Cellular Stress Biology, School of Life Sciences, Xiamen University, Xiamen, China
- Fang Zongxi Center for Marine EvoDevo, MoE Key Laboratory of Marine Genetics and Breeding, College of Marine Life Sciences, Ocean University of China, Qingdao, China
- MoE Key Laboratory of Evolution and Marine Biodiversity, Institute of Evolution and Marine Biodiversity, Ocean University of China, Qingdao, China
- Mountain Ecological Restoration and Biodiversity Conservation Key Laboratory of Sichuan Province, Chengdu Institute of Biology, Chinese Academy of Sciences, Chengdu, China
- China–Croatia Belt and Road Joint Laboratory on Biodiversity and Ecosystem Services, Chengdu Institute of Biology, Chinese Academy of Sciences, Chengdu, China
- Department of Basic Medical Sciences, Qiannan Medical College for Nationalities, Duyun, China
- State Key Laboratory of Genome and Multi-Omics Technologies and Shenzhen Key Laboratory of Forensics, BGI Research, Shenzhen, China
- BGI Research, Wuhan, China
- Department of Ecology and Evolutionary Biology, Yale University, New Haven, CT USA
- Department of Evolutionary Biology, University of Vienna, Vienna, Austria
Abstract
The complex brains of vertebrates have more cell types than those of their closest relatives. Whole-genome duplications (WGDs) occurred during early vertebrate evolution1, but it is unclear whether the duplicated genes (ohnologues) facilitated cell-type evolution. Here using brain single-cell transcriptomes from five chordates—human2, mouse3, lizard4, lamprey5 and amphioxus—we report that many cell-type families with conserved core transcription factors in vertebrates do not show one-to-one homology with amphioxus. Moreover, ohnologues, particularly those from the first WGD, were more important than small-scale duplication paralogues for vertebrate cell-type evolution. To explore whether ohnologues are mechanistically important for this process, we predicted ancestral cell-type states and compared them to amphioxus and experimentally investigated macroglia. The findings indicate that ohnologues had a role in early vertebrate cell-type diversification. Moreover, by examining paralogue expression across cell types and species, we show that expression changes were mainly driven by dosage selection and subfunctionalization. We also link ohnologues to cellular diversity at different anatomical and cell-type scales. Our findings demonstrate the importance of WGDs for the evolution of early vertebrate brain complexity and highlight that the resultant ohnologues continued to capacitate cell-type evolution long after they were formed.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 26 matches between paragraphs and lines of code.
DiracZhu1998/WGD2celltype_evolution
808269e1c0c13331cf19a21284904ccbb26bf54d, 30 March 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
90 files
- 0.atlas/
1.preprocessing/ , Python, 62 linesHsap/ step1.preprocess_data.py - 0.atlas/
1.preprocessing/ , Python, 74 linesMmus/ step1.loom_to_h5ad_subse t_to_Seurat.py - 0.atlas/
1.preprocessing/ , Jupyter, 58 linesPmar/ merge_lamprey_allRDS.ipy nb - 0.atlas/
1.preprocessing/ , Python, 18 linesPmar/ replace_hyphens_undersco res.py - 0.atlas/
1.preprocessing/ , Jupyter, 102 linesPvit/ add_annotations_info_and _convert2h5ad.ipynb - 0.atlas/
1.preprocessing/ , Jupyter, 48 lineseye_lung/ mod.ipynb - 0.atlas/
2.anndata/ , Jupyter, 73 linesTune_celltype_label_and_ subset_neurons_vs_non_ne urons.ipynb - 0.atlas/
4.Seurat/ , R, 19 linesconvert_scanpy_to_seurat .R - 0.atlas/
5.iterative_clustering/ , R, 236 linesClusterPlotFunctions.R - 0.atlas/
5.iterative_clustering/ , R, 73 lines, 1 matchRun.iterative_clustering .R - 0.atlas/
5.iterative_clustering/ , R, 33 linesget_top_reference_for_at las_iterative_clustering .R - 0.atlas/
6.SAMap/ , Python, 49 lines, 1 matchSAMap_run_vertebrate_neu rons.py - 0.atlas/
6.SAMap/ , Python, 49 lines, 1 matchSAMap_run_vertebrate_non neurons.py - 0.atlas/
6.SAMap/ , Python, 120 linesextract_important_inform ations.v2.py - 0.atlas/
7.final/ , Jupyter, 120 liness1.transfer_annotated_in formation_to_individual_ atlas.ipynb - 0.atlas/
7.final/ , R, 19 liness2.convert_scanpy_to_seu rat.R - 0.atlas/
7.final/ , Jupyter, 96 liness3.generate_orthologs_se urat_atlas.ipynb - 0.atlas/
7.final/ , Jupyter, 105 liness4.generate_orthogroup_s eurat_atlas.ipynb - 0.atlas/
7.final_Bflo/ , R, 155 linesstep1.cellQC.R - 0.atlas/
7.final_Bflo/ , R, 184 lines, 1 matchstep2.harmony.R - 10.case_studies/
1.macroglia/ , Jupyter, 28 linesAST/ 1.get_subtype/ subset.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 165 linesAST/ 3.subtype_identification / plot_UMAP.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 187 linesAST/ 3.subtype_identification / subtype_dotplot.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 498 lines, 2 matchesAST/ 5.regionalisation_genes/ plot_and_check_regionali sation.v2.ipynb - 10.case_studies/
1.macroglia/ , R, 58 lines, 2 matchesAST/ 5.regionalisation_genes/ s1.get_matrix.R - 10.case_studies/
1.macroglia/ , R, 75 linesAST/ 5.regionalisation_genes/ s2.LMM.R - 10.case_studies/
1.macroglia/ , R, 27 linesAST/ 5.regionalisation_genes/ s3.marker.R - 10.case_studies/
1.macroglia/ , Jupyter, 40 linesEpen/ 1.get_subtype/ subset.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 130 linesEpen/ 3.subtype_identification / plot_UMAP.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 25 linesOligo_OPC/ 1.get_subtype/ subset.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 139 linesOligo_OPC/ 3.subtype_identification / plot_UMAP.ipynb - 10.case_studies/
1.macroglia/ , Jupyter, 89 linesOligo_OPC/ 3.subtype_identification / subtype_dotplot.ipynb - 10.case_studies/
2.cerellar_nuclei/ , Jupyter, 110 lines1.update_atlas/ update_info_rds.v2.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 103 lines1.evo_cell_tree/ 0.prep_data/ generate_orthogroup_sum_ exp_seurat_for_glia_tree .ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 94 lines1.evo_cell_tree/ 1.build_tree_with_specif icity_indices/ 1.chordate_glia_tree_on_ metagene.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 157 lines2.SAMap_glia_relationshi ps/ chord_for_neuron_atlas.i pynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 150 lines2.SAMap_glia_relationshi ps/ chord_for_non_neuron_atl as.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 46 lines, 1 match3.subset_glia/ subset_glia.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 77 lines4.CytoTrace2/ cytotrace.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 281 lines5.keyTFs_macroglia/ find_cases_AST_Epen_usin g_paralogous_genes.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 265 lines5.keyTFs_macroglia/ find_cases_AST_OPC_using _paralogous_genes.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 280 lines5.keyTFs_macroglia/ find_cases_AST_Oligo_usi ng_paralogous_genes.ipyn b - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 67 lines6.scvelo/ 1.convert2h5ad.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 52 lines6.scvelo/ 2.merge_velo_info_into_h 5ad.ipynb - 10.case_studies/
3.amphi_glia_evo_devo/ , Jupyter, 61 lines6.scvelo/ 3.scvelo.ipynb - 10.case_studies/
4.hypothalamus/ , Jupyter, 81 lines0.atlas/ Ciona_CNS_atlas_prep.ipy nb - 10.case_studies/
4.hypothalamus/ , Jupyter, 58 lines0.atlas/ Mouse_hypo.ipynb - 10.case_studies/
4.hypothalamus/ , Jupyter, 25 lines0.atlas/ Mouse_hypo_subset_gene.i pynb - 10.case_studies/
4.hypothalamus/ , Jupyter, 109 lines2.plotting/ chord_for_neuron_hypotha lamus.ipynb - 10.case_studies/
4.hypothalamus/ , Jupyter, 104 lines2.plotting/ chord_for_non_neuron_hyp othalamus.ipynb - 2.gene_relationships/
paralogue_inference/ , Jupyter, 55 linesinferring_paralogs_v2.ip ynb - 3.markers/
3.nsforest/ , Python, 77 lines, 1 matchfind_minimum_combination s_of_TFs.py - 3.markers/
3.nsforest/ , Python, 73 lines, 1 matchfind_minimum_combination s_of_markers.py - 3.markers/
DEGs_Seurat/ , R, 28 lines, 1 matchFindMarkers.R - 4.ohno_para_significance
/ , R, 147 lines0bin/ Ohnologs_stats.v2.R - 4.ohno_para_significance
/ , R, 144 lines0bin/ Paralogs_stats.v2.R - 4.ohno_para_significance
/ , Jupyter, 85 lines0bin/ get_background_genes_fro m_atlas.ipynb - 4.ohno_para_significance
/ , R, 154 lines0bin/ non_Ohnolog_Paralogs_sta ts.v2.R - 4.ohno_para_significance
/ , Jupyter, 93 lines3.plot_EDFig4de/ summary_plot.ipynb - 4.ohno_para_significance
/ , Shell, 12 lines3.plot_EDFig4de/ work.sh - 4.ohno_para_significance
/ , Jupyter, 189 lines3.plot_Fig2/ summary_plot.ipynb - 4.ohno_para_significance
/ , Jupyter, 181 lines3.plot_Fig2/ summary_plot.subtype.ipy nb - 5.paralog_expression_evo
lution/ , Jupyter, 93 linesDominant expression/ s1.get_matrix_expression _data_and_pct_data.ipynb - 5.paralog_expression_evo
lution/ , R, 110 lines, 1 matchDominant expression/ s2.check_dominant_within _species_paralogs.R - 5.paralog_expression_evo
lution/ , Jupyter, 70 linesDominant expression/ s3.plot_dominate_express ion_frequency_for_indivi dual_species.by_Friedman .ipynb - 5.paralog_expression_evo
lution/ , Jupyter, 140 lines, 2 matchesDominant expression/ s3.plot_dominate_express ion_frequency_for_indivi dual_species.by_pairwise _wilcox.ipynb - 5.paralog_expression_evo
lution/ , R, 295 linesDominant expression/ s4.plot_dominant_express ion_across_species.R - 5.paralog_expression_evo
lution/ , Jupyter, 358 linesDominant expression/ s4.plot_dominant_express ion_across_species.ipynb - 5.paralog_expression_evo
lution/ , R, 41 linesParalog shift sub_vs_neo/ Trinarization_score.R - 5.paralog_expression_evo
lution/ , Jupyter, 56 linesParalog shift sub_vs_neo/ s1.get_combined_SSD_WGD_ pairs.ipynb - 5.paralog_expression_evo
lution/ , R, 35 linesParalog shift sub_vs_neo/ s1.get_expression_domain _Triscore.R - 5.paralog_expression_evo
lution/ , Jupyter, 340 lines, 2 matchesParalog shift sub_vs_neo/ s2.get_vertebrate_ancest ral_states_and_check_sub _neo_loss.latest_gene_fa mily_reconsider.ipynb - 5.paralog_expression_evo
lution/ , Jupyter, 169 lines, 2 matchesParalog shift sub_vs_neo/ s3.calculate_divergence_ across_species.ipynb - 5.paralog_expression_evo
lution/ , Jupyter, 361 lines, 1 matchParalog shift sub_vs_neo/ s4.paralog_switching_cal culation.updated.ipynb - 8.GRNs_detection/
1.human/ , Python, 20 lines0.bin/ s0.modify_info_file.py - 8.GRNs_detection/
1.human/ , R, 76 lines0.bin/ s1.convert_feather2conse nsusEnsembl.R - 8.GRNs_detection/
1.human/ , R, 25 lines0.bin/ s1b.subset_exp_info.R - 8.GRNs_detection/
1.human/ , Python, 36 lines0.bin/ s2.changeTFname.py - 8.GRNs_detection/
1.human/ , Python, 44 lines0.bin/ s2.change_motif_name.py - 8.GRNs_detection/
1.human/ , Python, 142 lines, 1 matchpyscenic_run.py - 8.GRNs_detection/
1.human/ , Python, 25 linessave_logos.py - 8.GRNs_detection/
1.mouse/ , Python, 20 lines0.bin/ s0.modify_info_file.py - 8.GRNs_detection/
1.mouse/ , R, 71 lines0.bin/ s1.convert_feather2conse nsusEnsembl.R - 8.GRNs_detection/
1.mouse/ , R, 25 lines0.bin/ s1b.subset_exp_info.R - 8.GRNs_detection/
1.mouse/ , Python, 36 lines0.bin/ s2.changeTFname.py - 8.GRNs_detection/
1.mouse/ , Python, 44 lines0.bin/ s2.change_motif_name.py - 8.GRNs_detection/
1.mouse/ , Python, 142 lines, 1 matchpyscenic_run.py - 8.GRNs_detection/
1.mouse/ , Python, 25 linessave_logos.py - LICENSE, License, 674 lines
- README.md, Text, 11 lines
fmarletaz/hagfish
68842e98e6515a12869c951fc1b82026fa0df103, 14 November 2023Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
22 files
- Functional/
ext_go.r , R, 345 lines - Gene families/
gene_families.ipynb , Jupyter, 858 lines - Paralogons/
Tetraploidy_cl.ipynb , Jupyter, 1,699 lines - Paralogons/
molecular_dating/ , Jupyter, 62 linesUntitled.ipynb - Paralogons/
molecular_dating/ , R, 145 linesplot_chrono.r - Paralogons/
relaxed/ , R, 69 linespgon-draw.r - Paralogons/
sel-pgon2.py , Python, 76 lines - Paralogons/
strict/ , R, 69 linespgon-draw.r - Synteny/
ideogram.r , R, 196 lines - Synteny/
plot_clg_hagfish.r , R, 305 lines - Transcriptomics/
subfunc.r , R, 184 lines - Transcriptomics/
tau_re.r , R, 310 lines - hox/
concat_ali.py , Python, 97 lines - rediploidization/
scripts/ , Python, 128 linesadd_outgr.py - rediploidization/
scripts/ , Python, 249 linesget_clgb_cyclo_fams.py - rediploidization/
scripts/ , Python, 354 linesget_informative_fams.py - rediploidization/
scripts/ , Python, 43 linesget_subali.py - rediploidization/
scripts/ , Python, 144 linesparse_au.py - rediploidization/
scripts/ , Python, 88 linesplot_au.py - rediploidization/
scripts/ , Shell, 52 linesprototype_au_test_many.s h - wgd_modelling/
whale_brspecDLWGD.jl , Julia, 56 lines - README.md, Text, 19 lines
linnarsson-lab/adult-human-brain
2b5aa12dffb8cc1bcb3979c409cb6ac907546b2e, 5 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
97 files
- cytograph/
__init__.py , Python, 15 lines - cytograph/
_version.py , Python, 1 line - cytograph/
annotation/ , Python, 4 lines__init__.py - cytograph/
annotation/ , Python, 142 linesauto_annotator.py - cytograph/
annotation/ , Python, 104 linesauto_auto_annotator.py - cytograph/
annotation/ , Python, 35 linescell_cycle_annotator.py - cytograph/
clustering/ , Python, 4 lines__init__.py - cytograph/
clustering/ , Python, 64 linescluster_validator.py - cytograph/
clustering/ , Python, 241 lines, 1 matchpolished_louvain.py - cytograph/
clustering/ , Python, 59 linespolished_surprise.py - cytograph/
clustering/ , Python, 89 linesunpolished_louvain.py - cytograph/
decomposition/ , Python, 252 linesHPF.py - cytograph/
decomposition/ , Python, 303 linesHPF_accel.py - cytograph/
decomposition/ , Python, 3 lines__init__.py - cytograph/
decomposition/ , Python, 59 linesincremental_residuals_pc a.py - cytograph/
decomposition/ , Python, 99 linespca.py - cytograph/
embedding/ , Python, 1 line__init__.py - cytograph/
embedding/ , Python, 115 linesart_of_tsne.py - cytograph/
embedding/ , Python, 34 linescustom_init.py - cytograph/
enrichment/ , Python, 8 lines__init__.py - cytograph/
enrichment/ , Python, 86 linesbinary_differential_expr ession.py - cytograph/
enrichment/ , Python, 65 linesenrichment.py - cytograph/
enrichment/ , Python, 69 linesfeature_selection_by_dev iance.py - cytograph/
enrichment/ , Python, 144 linesfeature_selection_by_enr ichment.py - cytograph/
enrichment/ , Python, 185 linesfeature_selection_by_mul tilevel_enrichment.py - cytograph/
enrichment/ , Python, 56 linesfeature_selection_by_var iance.py - cytograph/
enrichment/ , Python, 75 linesgsea.py - cytograph/
enrichment/ , Python, 49 linesneighborhood_enrichment. py - cytograph/
enrichment/ , Python, 58 linestrinarizer.py - cytograph/
manifold/ , Python, 3 lines__init__.py - cytograph/
manifold/ , Python, 366 linesbalanced_knn.py - cytograph/
manifold/ , Python, 57 linesgraph_skeletonizer.py - cytograph/
manifold/ , Python, 103 linespoisson_pooling.py - cytograph/
metrics/ , Python, 2 lines__init__.py - cytograph/
metrics/ , Python, 47 lineslaplacian_score.py - cytograph/
metrics/ , Python, 68 linespoisson_proximity.py - cytograph/
metrics/ , Python, 142 linesspecial_metrics.py - cytograph/
pipeline/ , Python, 6 lines__init__.py - cytograph/
pipeline/ , Python, 128 linesaggregator.py - cytograph/
pipeline/ , Python, 516 linescommands.py - cytograph/
pipeline/ , Python, 109 linesconfig.py - cytograph/
pipeline/ , Python, 211 linescytograph.py - cytograph/
pipeline/ , Python, 245 linesengine.py - cytograph/
pipeline/ , Python, 165 linespunchcards.py - cytograph/
pipeline/ , Python, 20 linesutils.py - cytograph/
pipeline/ , Python, 519 linesworkflow.py - cytograph/
plotting/ , Python, 29 linesTF_heatmap.py - cytograph/
plotting/ , Python, 20 lines__init__.py - cytograph/
plotting/ , Python, 88 linesbatch_covariates.py - cytograph/
plotting/ , Python, 47 linesbuckets.py - cytograph/
plotting/ , Python, 26 linescell_cycle.py - cytograph/
plotting/ , Python, 72 linescolors.py - cytograph/
plotting/ , Python, 31 linesdecision_boundary.py - cytograph/
plotting/ , Python, 51 linesdendrogram.py - cytograph/
plotting/ , Python, 135 linesdoublets_plots.py - cytograph/
plotting/ , Python, 46 linesembedded_velocity.py - cytograph/
plotting/ , Python, 30 linesfactors.py - cytograph/
plotting/ , Python, 42 linesgene_velocity.py - cytograph/
plotting/ , Python, 166 linesheatmap.py - cytograph/
plotting/ , Python, 95 linesmanifold.py - cytograph/
plotting/ , Python, 26 linesmarkerheatmap.py - cytograph/
plotting/ , Python, 47 linesmetromap.py - cytograph/
plotting/ , Python, 12 linesmidpoint_normalize.py - cytograph/
plotting/ , Python, 91 linespunchcard_selection.py - cytograph/
plotting/ , Python, 120 linesqc_plots.py - cytograph/
plotting/ , Python, 85 linesradius_characteristics.p y - cytograph/
plotting/ , Python, 56 linesscatter.py - cytograph/
plotting/ , Python, 35 linesumi_genes.py - cytograph/
postprocessing/ , Python, 2 lines__init__.py - cytograph/
postprocessing/ , Python, 77 linesmerge_subset.py - cytograph/
postprocessing/ , Python, 186 linessplit_subset.py - cytograph/
preprocessing/ , Python, 4 lines__init__.py - cytograph/
preprocessing/ , Python, 173 lines, 1 matchdoublet_finder.py - cytograph/
preprocessing/ , Python, 104 linesgenotyping.py - cytograph/
preprocessing/ , Python, 68 linesnormalizer.py - cytograph/
preprocessing/ , Python, 24 linesqc_functions.py - cytograph/
preprocessing/ , Python, 26 linesutils.py - cytograph/
species/ , Python, 1 line__init__.py - cytograph/
species/ , Python, 1,801 lineshuman.py - cytograph/
species/ , Python, 1,508 lines, 1 matchmouse.py - cytograph/
species/ , Python, 154 lines, 1 matchspecies.py - cytograph/
utils.py , Python, 104 lines - cytograph/
visualization/ , Python, 2 lines__init__.py - cytograph/
visualization/ , Python, 281 linescolors.py - cytograph/
visualization/ , Python, 214 linesplot_overview.py - cytograph/
visualization/ , Python, 134 linesscatter.py - notebooks/
Preprint/ , Jupyter, 1,638 linesFigure1.ipynb - notebooks/
Preprint/ , Jupyter, 504 linesFigure2.ipynb - notebooks/
Preprint/ , Jupyter, 1,726 linesFigure3.ipynb - notebooks/
Preprint/ , Jupyter, 1,116 linesFigure5.ipynb - notebooks/
Preprint/ , Jupyter, 288 linesFigure5_Diffxpy.ipynb - notebooks/
Preprint/ , Jupyter, 605 linesFigure6.ipynb - notebooks/
Preprint/ , Jupyter, 513 linesFigureS1.ipynb - notebooks/
Preprint/ , Jupyter, 456 linesFigureS10.ipynb - notebooks/
Preprint/ , Jupyter, 482 linesFigureS7.ipynb - repository limit reached (2,000 files or 30 MB): the rest is at the source (18 files)
- LICENSE, License, 25 lines
- README.md, Text, 75 lines
figshare 29327111
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
Code availability
All code and scripts associated with this analysis are publicly available from the GitHub repository (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 204 scripts, each with its path and the digest of its content;
- 26 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE160471, at NCBI GEO; found in “Data availability”
Data Availability Statement
The reference genome, gene models, functional annotations of protein-coding genes, full marker gene list of each cell cluster and important intermediate files are deposited in a Figshare repository (10.6084/
All code and scripts associated with this analysis are publicly available from the GitHub repository (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Nature Portfolio
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 15 authors, 4 keywords, 15 MeSH terms, 1 funder, 100 references.
Cite
This paper
Zhu, Y., Zhang, S., Wei, J., Dolgetta-Garcia, D., Jindrich, K., Liu, H., Shi, C., Pan, R., Chen, Y., Xu, Y., Li, Q., Wagner, G. P., Holland, P. W. H., Li, G., & Shimeld, S. M. (2026). Whole-genome duplication shaped cell-type evolution in the vertebrate brain. Nature, 656(8128), 659-669. https://
BibTeX
@article{zhu2026whole,
author = {Zhu, Yuanzhen and Zhang, Shuai and Wei, Jiankai and Dolgetta-Garcia, Diego and Jindrich, Katia and Liu, Huimin and Shi, Chenggang and Pan, Rongrong and Chen, Yuwei and Xu, Yan and Li, Qiye and Wagner, Günter P and Holland, Peter W H and Li, Guang and Shimeld, Sebastian M},
title = {{Whole-genome duplication shaped cell-type evolution in the vertebrate brain}},
journal = {Nature},
year = {2026},
month = jun,
volume = {656},
number = {8128},
pages = {659--669},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/
url = {https://
pmid = {42271056},
pmcid = {PMC13489949}
}
RIS
TY - JOUR
AU - Zhu, Yuanzhen
AU - Zhang, Shuai
AU - Wei, Jiankai
AU - Dolgetta-Garcia, Diego
AU - Jindrich, Katia
AU - Liu, Huimin
AU - Shi, Chenggang
AU - Pan, Rongrong
AU - Chen, Yuwei
AU - Xu, Yan
AU - Li, Qiye
AU - Wagner, Günter P
AU - Holland, Peter W H
AU - Li, Guang
AU - Shimeld, Sebastian M
TI - Whole-genome duplication shaped cell-type evolution in the vertebrate brain
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/
VL - 656
IS - 8128
SP - 659
EP - 669
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Whole-genome duplication shaped cell-type evolution in the vertebrate brain",
"container-title": "Nature",
"author": [
{
"family": "Zhu",
"given": "Yuanzhen"
},
{
"family": "Zhang",
"given": "Shuai"
},
{
"family": "Wei",
"given": "Jiankai"
},
{
"family": "Dolgetta-Garcia",
"given": "Diego"
},
{
"family": "Jindrich",
"given": "Katia"
},
{
"family": "Liu",
"given": "Huimin"
},
{
"family": "Shi",
"given": "Chenggang"
},
{
"family": "Pan",
"given": "Rongrong"
},
{
"family": "Chen",
"given": "Yuwei"
},
{
"family": "Xu",
"given": "Yan"
},
{
"family": "Li",
"given": "Qiye"
},
{
"family": "Wagner",
"given": "Günter P"
},
{
"family": "Holland",
"given": "Peter W H"
},
{
"family": "Li",
"given": "Guang"
},
{
"family": "Shimeld",
"given": "Sebastian M"
}
],
"container-title-short":
"volume": "656",
"issue": "8128",
"page": "659-669",
"DOI": "10.1038/
"PMID": "42271056",
"PMCID": "PMC13489949",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: scVelo, Harmony, UMAP, 23 other tools, genetics / omics, cellular / molecular, 6 references
- [2] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: Turing.jl, DataFrames.jl, Plots.jl, 24 other tools, mouse, cellular / molecular
- [3] doi:10.1126/sciadv.aed2952 [code]
- Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.Journal: Science advancesIn common: scVelo, BEDTools, UMAP, 16 other tools, cellular / molecular, 7 references
- [4] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: reticulate, rstatix, UMAP, 19 other tools, genetics / omics, 4 references
- [5] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: BEDTools, reticulate, Pingouin, 19 other tools, genetics / omics, mouse, cellular / molecular, 2 references
- [6] doi:10.1126/sciadv.aeg3223 [code]
- The extreme diversity of retinal amacrine cells has deep evolutionary roots.Journal: Science advancesIn common: reticulate, rstatix, anndata, 17 other tools, genetics / omics, cellular / molecular, 3 references
- [7] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: BEDTools, reticulate, rstatix, 19 other tools, genetics / omics, 2 references
- [8] doi:10.1038/s41586-026-10214-2 [code]
- Multidimensional profiling of heterogeneity in supratentorial ependymomas.Journal: NatureIn common: Harmony, reticulate, anndata, 19 other tools, genetics / omics, mouse, 1 reference
- [9] doi:10.1016/j.xcrm.2026.102651 [code]
- Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.Journal: Cell reports. MedicineIn common: Harmony, reticulate, UMAP, 19 other tools, genetics / omics, cellular / molecular
- [10] doi:10.1038/s42003-026-10034-0 [code]
- Region- and cell type-specific changes in gene expression in the cerebellum after classical fear conditioning.Journal: Communications biologyIn common: reticulate, rstatix, anndata, 16 other tools, genetics / omics, mouse, cellular / molecular, 4 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 4 repositories of the authors' code, each at its verified commit and with its license, 204 scripts, and 26 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:db029566fa4d5784…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
