Comparative analysis of the cellular landscape in mammalian striatum.
The 23 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Methods › Single-nucleus RNA sequencing, cell type annotation, and cellular compositional abundance analysis ↔ 05_DGE_analysis/00_prep_pseudobulk_counts.R, lines 484–535 · score 0.98 · k.weight, LogNormalize, FindVariableFeatures, IntegrateData, FindIntegrationAnchors, SelectIntegrationFeatures
- [2] § Methods › Single-nucleus RNA sequencing dataset alignment and quality control ↔ 03_SPNs/00_get_orthologs.sh, the whole file · a weak match · score 0.97 · Mustela putorius furo, Callithrix jacchus, Homo Sapiens, Macaca mulatta, Mus musculus, Pan troglodytes
- [3] § Methods › Bat interneuron subtype analyses ↔ 02_Bat/04_WGCNA.R, lines 128–166 · score 0.97 · soft threshold power, blockwiseModules, deepSplit, maxBlockSize, minCoreKME, minKMEtoStay
- [4] § Methods › Bat interneuron subtype analyses ↔ 02_Human/06_WGCNA.R, lines 133–172 · score 0.96 · soft threshold power, blockwiseModules, deepSplit, maxBlockSize, minCoreKME, minKMEtoStay
- [5] § Results › Identification of striatal cell populations of each species ↔ QC_plots.R, lines 353–401 · score 0.85 · Il1rapl2, Ppp1r1b, Csf1r, Apbb1ip, CHRM2, GFAP
- [6] § Results › Non-primates have lower eSPN to SPN proportions ↔ 03_SPNs/03_eSPN_D1D2_hybrid_correlation.R, lines 43–82 · score 0.78 · SEMA5B, SV2B, eSPNs, CRYM, KREMEN1, SGK1
- [7] § Results › Non-primates have lower eSPN to SPN proportions ↔ 03_SPNs/03_eSPN_D1D2_hybrid_correlation.R, lines 43–82 · score 0.78 · SEMA5B, SV2B, eSPNs, CRYM, KREMEN1, SGK1
- [8] § Methods › Human-specific differentially expressed gene (HS DEG) analysis ↔ 02_Human/06_WGCNA.R, lines 235–277 · score 0.70 · HS DEGs, gene enrichment, GO term, cutoffs, WGCNA, modules
- [9] § Results › Differential gene expression analysis reveals stronger human specific change in SPNs than glia ↔ 02_Human/06_WGCNA.R, lines 235–277 · score 0.69 · red module genes, blue module genes, GO term, WGCNA, upregulated, enrichment
- [10] § Methods › Human-specific differentially expressed gene (HS DEG) analysis ↔ 05_DGE_analysis/01_primates_pseudobulk_DESeq2.R, lines 134–176 · score 0.67 · humanized age, DESeq2, covariates, pseudobulk, sex, log10
- [11] § Methods › Single-nucleus RNA sequencing, cell type annotation, and cellular compositional abundance analysis ↔ 05_DGE_analysis/00_prep_pseudobulk_counts.R, lines 94–137 · score 0.65 · FindIntegrationAnchors, SelectIntegrationFeatures, orthologous genes, Seurat, tool, NCBI
- [12] § Methods › Single-nucleus RNA sequencing dataset alignment and quality control ↔ 01_Generate_NHP_gtfs/MKREF_LIFTOFF_CHIMP.sh, lines 1–39 · score 0.63 · panTro5, mkref, GTF, liftoff, hg38, genomes
- [13] § Methods › Differentially expressed gene (DEG) and correlation analysis ↔ 03_SPNs/03_eSPN_D1D2_hybrid_correlation.R, lines 124–186 · score 0.63 · D1 D2 hybrid, gene expression, heatmap, correlation, SPN, macaque
- [14] § Methods › Differentially expressed gene (DEG) and correlation analysis ↔ 03_SPNs/03_eSPN_D1D2_hybrid_correlation.R, lines 124–180 · score 0.63 · D1 D2 hybrid, gene expression, heatmap, correlation, SPN, macaque
- [15] § Methods › Single-nucleus RNA sequencing, cell type annotation, and cellular compositional abundance analysis ↔ 03_SPNs/01_INTEGRATION_AND_CLUSTERING.R, lines 210–253 · score 0.61 · FindIntegrationAnchors, SelectIntegrationFeatures, SRR13808461, SRR13808459, NCBI, orthologous
- [16] § Results › Identification of interneurons found primarily in the bat ↔ 04_Interneurons/01_Integration&Clustering.R, lines 1076–1143 · score 0.61 · PDGFD PTHLH PVALB, FOXP2 EYA2, CCK VIP, FOXP2 TSHZ2, cluster, SST
- [17] § Results › Identification of striatal cell populations of each species ↔ 05_DGE_analysis/00_prep_pseudobulk_counts.R, lines 436–482 · score 0.59 · PPP1R1B, CSF1R, ARHGAP15, PDGFRA, AQP4, PENK
- [18] § Results › Identification of interneurons found primarily in the bat ↔ 04_Interneurons/04_Correlation_withinBatPutamen.R, lines 42–104 · score 0.57 · correlation matrix, FOXP2 EYA2, FOXP2 TSHZ2, gene expression, LMO3, interneurons
- [19] § Methods › Differentially expressed gene (DEG) and correlation analysis ↔ 04_Interneurons/01_Integration&Clustering.R, lines 1800–1842 · score 0.57 · BuildClusterTree, PlotClusterTree, distances, matrix, clustering
- [20] § Results › Identification of interneurons found primarily in the bat ↔ 04_Interneurons/03_DEG_Analysis.R, lines 311–350 · score 0.55 · GO term enrichment, upregulated genes, FOXP2 TSHZ2, enriched, interneurons, bat
- [21] § Results › Non-primates have lower eSPN to SPN proportions ↔ 03_SPNs/02_eSPN_markers.R, lines 395–444 · score 0.55 · RUNX1T1, DCLK1, eSPN, TSHZ1, CASZ1, FOXP2
- [22] § Methods › Differentially expressed gene (DEG) and correlation analysis ↔ 05_DGE_analysis/02_DEG_figures.R, lines 317–360 · score 0.54 · log2 fold change, downregulated, upregulated, DEGs, putamen, genes
- [23] § Methods › Differentially expressed gene (DEG) and correlation analysis ↔ 05_DGE_analysis/01_primates_pseudobulk_DESeq2.R, lines 134–176 · score 0.52 · log2 fold change, volcano, DEGs, sum, genes
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 1,667 lines · 69 KB · MIT · 3 matches
- # load packages
- library(patchwork)
- library(Seurat)
- library(rhdf5)
- library(DropletUtils)
- library(DropletQC)
- library(dplyr)
- library(Matrix)
- library(ggplot2)
- library(plyr)
- library(tidyverse)
- library(tidyr)
- library(ggpubr)
- library(reshape2)
- library(rio)
- library(data.table)
- library(harmony)
- library(ggrepel)
- library(edgeR)
- library(variancePartition)
- library(Matrix.utils)
- library(SingleCellExperiment)
- library(RColorBrewer)
- set.seed(1234)
- source("/project/Neuroinformatics_Core/Konopka_lab/s422071/SCRIPTS_pr/SCRIPTS/utility_functions.R")
- library(curl)
- #conda install bioconda::r-wgcna
- #BiocManager::install('WGCNA')
- library(WGCNA)
- library(flashClust)
- #remotes::install_url("https://cran.r-project.org/src/contrib/Archive/matrixStats/matrixStats_1.1.0.tar.gz")
- library(matrixStats)
- library(RColorBrewer)
- ####
- ## PREPARE DATASETS
- ####
- # read the data
- human_str <- readRDS(paste0(file_dir,"human_integrated_caudate_putamen_ANNOTATED.RDS"))
- chimp_str = readRDS(paste0(file_dir,"chimp_integrated_caudate_putamen_ANNOTATED.RDS"))
- macaque_str = readRDS(paste0(file_dir,"macaque_integrated_caudate_putamen_ANNOTATED.RDS"))
- marmoset_str = readRDS(paste0(file_dir,"marmoset_integrated_caudate_putamen_ANNOTATED.RDS"))
- mouse_str = readRDS(paste0(file_dir,"Mouse_Caudate_Annotated_FINAL.RDS"))
- mouse_str$Tissue = rep("Caudoputamen", nrow(mouse_str[[]]))
- mouse_str$Species = rep("Mouse", nrow(mouse_str[[]]))
- mouse_str$id = paste0(mouse_str$orig.ident, mouse_str$Tissue)
- bat_str = readRDS(paste0(file_dir,"bat_integrated_caudate_putamen_ANNOTATED.RDS"))
- ferret_str = readRDS(paste0(file_dir,"Ferret_Caudate_Krienen_ANNOTATED.RDS"))
- ferret_str$newannot_2 = ferret_str$newannot
- ferret_str$newannot = ferret_str$broad_annot
- # Reshape and combine metadata
- human_meta = [email hidden]
- human_meta$Species = 'Human'
- chimp_meta = [email hidden]
- chimp_meta$Species = 'Chimp'
- macaque_meta = [email hidden]
- macaque_meta$Species = 'Macaque'
- marmoset_meta = [email hidden]
- marmoset_meta$Species = 'Marmoset'
- bat_meta = [email hidden]
- bat_meta$Species = 'Bat'
- sub_obj_new_list <- list(
- Human = human_str,
- Chimp = chimp_str,
- Macaque = macaque_str,
- Marmoset = marmoset_str,
- Mouse = mouse_str,
- Bat = bat_str,
- Ferret = ferret_str
- )
- saveRDS(sub_obj_new_list, file = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/sub_obj_new_list_11_12_2025.RDS")
- # make RNA as active assay
- a <- lapply(sub_obj_new_list, function(x) {
- DefaultAssay(x) <- "RNA"
- x@assays$SCT = NULL
- x[["RNA"]] <- as(object = x[["RNA"]], Class = "Assay")
- x
- })
- b_new = Reduce(merge, a)
- ### microglia cleanup
- seurM = subset(b_new, subset = newannot_2 == "Microglia")
- # Found ortholog genes from human pr coding genes and extracted the genes which are ortholog in all 7 species using ncbi datasets tool
- ortho_genes <- read.table("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/metadata/orthologs_in_7_species_human_pr_codingOnly_2.csv", header = TRUE)$symbol #15055 pr coding ortho genes
- ### subset the genes and the Microglia from each Species
- # subset cells and metadata
- # Extract the count matrix from the Seurat object (default is "RNA" assay)
- mat = seurM@assays$RNA$counts
- new_mat <- mat[rownames(mat) %in% ortho_genes,]
- # generate new Seurat object
- glia_merged <- CreateSeuratObject(count = new_mat, meta.data = seurM[[]])
- ####
- ## SET VARIABLES
- ####
- marksToPlot <- c('NRGN', 'SYT1', 'SNAP25', 'GAD1', 'GAD2', 'DRD1', 'DRD2', 'FOXP2', 'TAC1', 'PENK', 'SST', 'NPY', 'MOG', 'PLP1', 'MBP', 'MOBP', 'SLC1A2', 'SLC1A3', 'APOE', 'GPC5', 'GLI3', 'AQP4', 'CSF1R', 'ARHGAP15', 'PLDL1', 'CASZ1', 'PTPRZ1', 'PCDH15', 'SOX6', 'FYN', 'BCAS1', 'ENPP6', 'GPR17', 'FLT1', 'DUSP1', 'COBLL1', 'CFTR', 'CHAT', 'ADARB2', 'LHX6', 'MEIS2', 'TAC3', 'VIP', 'TH', 'CCK', 'TTC34', 'CROCC', 'FHAD1', 'SOX4', 'PDGFRB', 'PDGFRA', 'SATB2', 'SLC17A7', 'EYA2', 'PDGFD', 'PTHLH', 'PVALB', 'SPARCL1', 'RSPO2', 'RMST', 'LMO3', 'TSHZ2', "CALB1", "CALB2", "NOS1", "TSHZ1", "OPRM1", "CHST9","GRM8", "PPP1R1B", 'GRIK3', 'CXCL14')
- pref = 'Microglia_AllSpecies_AllTissues_human_pr_coding_orthologs'
- ####
- ## INTEGRATE ACROSS SPECIES
- ####
- # Split data
- seurM = glia_merged
- seurM[["RNA"]] = as(object = seurM[["RNA"]], Class = "Assay")
- seurM$id = paste0(seurM$orig.ident, '_', seurM$Tissue)
- seurML = SplitObject(seurM, split.by = "id")
- # normalize and identify variable features for each dataset independently
- seurML = lapply(X = seurML, FUN = function(x) {
- x = NormalizeData(x)
- x = FindVariableFeatures(x, selection.method = "vst", nfeatures = 2000)
- })
- # select features that are repeatedly variable across datasets for integration run PCA on each
- # dataset using these features
- features = SelectIntegrationFeatures(object.list = seurML)
- print(paste0('Number of genes to use for integration: ', length(features)))
- seurML = lapply(X = seurML, FUN = function(x) {
- x = ScaleData(x, features = features, verbose = FALSE)
- x = RunPCA(x, features = features, verbose = FALSE, npcs = 30)
- })
- anchors = FindIntegrationAnchors(object.list = seurML, anchor.features = features, normalization.method = "LogNormalize", reduction = "rpca")
- # save memory prior integration
- gc()
- # this command creates an 'integrated' data assay
- options(future.globals.maxSize = 16000 * 1024^2)
- allseur_integrated = IntegrateData(anchorset = anchors, k.weight = 30, normalization.method = 'LogNormalize')
- saveRDS(allseur_integrated, '/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/Integrated_rpca_Microglia_AllSpecies_ALLTissues_human_pr_coding_orthologs.RDS')
- # Clustering
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_1")
- pdf(paste0(pref, "_INTEGRATED_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- DefaultAssay(allseur_integrated) <- "integrated"
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_INTEGRATED_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_PRIMATE_ANNOT.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T, group.by = 'newannot') + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_INTEGRATED_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_INTEGRATED_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- saveRDS(allseur_integrated, paste0('/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_1/', pref, '_integrated_CLUSTERING1.RDS'))
- allseur_integrated = readRDS( paste0('/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_1/', pref, '_integrated_CLUSTERING1.RDS'))
- ################# CLUSTERING_2 ##########################
- ##Cluster#,CellType
- #12*,MOL+micro
- #19*,Neu+Micro
- #21*,Endo+micro
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_2/")
- toremove = c(12,19,21)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_2.RDS"))
- allseur_integrated = readRDS(paste0(pref,"_integrated_CLUSTERING_2.RDS"))
- ##################### CLUSTERING_3 #################
- #14*, MOL+micro
- #19*, MOL+micro
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_3/")
- toremove = c(14,19)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_3.RDS"))
- allseur_integrated = readRDS(paste0(pref,"_integrated_CLUSTERING_3.RDS"))
- ################### CLUSTERING_4 ###################
- #15*,Endo
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/CLUSTERING_4/")
- toremove = c(15)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_4.RDS"))
- ## annotate
- (mapnames <-setNames(rep("Microglia", 27),c(0:26)))
- # Create the broadannot
- allseur_integrated$Micro_annot = unname(mapnames[allseur_integrated[["seurat_clusters"]][,1]]) # to extract the annotations and not the cluster #s
- table(allseur_integrated$Micro_annot)
- #Microglia
- # 32961
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MICROGLIA/ANNOTATED")
- # QC PLOTS
- pdf(paste0(pref, "_Clusters.pdf"))
- DimPlot(allseur_integrated, group.by = 'Micro_annot', label = T, raster = T, label.size = 7) +
- NoLegend() +
- ggtitle('')
- dev.off()
- pdf(paste0(pref, "_Depth_Gene.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Micro_annot', y = 'nFeature_RNA') +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_Depth_UMI.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Micro_annot', y = 'nCount_RNA', ylim = c(0,20000)) +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- pdf(paste0(pref, "_NuclearFraction.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Micro_annot', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # MARKER GENE PLOTS
- DefaultAssay(allseur_integrated) = 'RNA'
- seurM_corr = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_MarkersDotPlot.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, group.by = "Micro_annot", features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'Micro_annot', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'Micro_annot', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_ANNOTATED.RDS"))
- ## add these annotations to the original seu obj
- # extract the newest annotations
- b$detailed_annot <- as.character(b$newannot_2)
- # Match cell names between objects
- common_cells <- intersect(rownames([email hidden]), rownames([email hidden]))
- # Update annotations for those cells
- b$detailed_annot[common_cells] <- allseur_integrated$Micro_annot[common_cells]
- b$is_micro = ifelse(rownames([email hidden]) %in% rownames([email hidden]), "yes", "no")
- # Identify microglia cells in b that are not in the integrated object
- b_keep = subset(b, subset = newannot_2 != "Microglia" | is_micro == "yes")
- ############################################
- ####### MOL + OPC
- ################################################
- ### subset the genes and the MOL+OPC from each Species
- # subset cells and metadata
- subset_b = subset(b_keep, subset = newannot_2 %in% c("MOL", "OPC")) # already orthologs only
- # Extract the count matrix from the Seurat object (default is "RNA" assay)
- mat = subset_b@assays$RNA$counts
- new_mat <- mat[rownames(mat) %in% ortho_genes,]
- # generate new Seurat object
- glia_merged <- CreateSeuratObject(count = new_mat, meta.data = subset_b[[]])
- ####
- ## SET VARIABLES
- ####
- # PPP1R1B IS AN SPN MARKER!
- # TSHZ1 IS A PATCH MARKER
- marksToPlot <- c('NRGN', 'SYT1', 'SNAP25', 'GAD1', 'GAD2', 'DRD1', 'DRD2', 'FOXP2', 'TAC1', 'PENK', 'SST', 'NPY', 'MOG', 'PLP1', 'MBP', 'MOBP', 'SLC1A2', 'SLC1A3', 'AQP4', 'CSF1R', 'ARHGAP15', 'PLD1','CASZ1', 'PTPRZ1', 'PCDH15', 'SOX6', 'FYN', 'BCAS1', 'ENPP6', 'GPR17', 'FLT1', 'DUSP1', 'COBLL1', 'CFTR', 'CHAT', 'ADARB2', 'LHX6', 'MEIS2', 'TAC3', 'VIP', 'TH', 'CCK', 'TTC34', 'CROCC', 'FHAD1', 'SOX4', 'PDGFRB', 'PDGFRA', 'SATB2', 'SLC17A7', 'EYA2', 'PDGFD', 'PTHLH', 'PVALB', 'SPARCL1', 'RSPO2', 'RMST', 'LMO3', 'TSHZ2', "CALB1", "CALB2", "NOS1", "TSHZ1", "OPRM1", "CHST9","GRM8", "PPP1R1B", 'GRIK3', 'CXCL14')
- pref = 'MOL_OPC_AllSpecies_AllTissues_human_pr_coding_orthologs'
- ####
- ## INTEGRATE ACROSS SPECIES
- ####
- #### use old seurat version
- # Split data
- seurM = glia_merged
- seurM[["RNA"]] = as(object = seurM[["RNA"]], Class = "Assay")
- seurM$id = paste0(seurM$orig.ident, '_', seurM$Tissue)
- seurML = SplitObject(seurM, split.by = "id")
- # normalize and identify variable features for each dataset independently
- seurML = lapply(X = seurML, FUN = function(x) {
- x = NormalizeData(x)
- x = FindVariableFeatures(x, selection.method = "vst", nfeatures = 2000)
- })
- # select features that are repeatedly variable across datasets for integration run PCA on each
- # dataset using these features
- features = SelectIntegrationFeatures(object.list = seurML)
- print(paste0('Number of genes to use for integration: ', length(features)))
- seurML = lapply(X = seurML, FUN = function(x) {
- x = ScaleData(x, features = features, verbose = FALSE)
- x = RunPCA(x, features = features, verbose = FALSE, npcs = 30)
- })
- anchors = FindIntegrationAnchors(object.list = seurML, anchor.features = features, normalization.method = "LogNormalize", reduction = "rpca")
- # save memory prior integration
- gc()
- # this command creates an 'integrated' data assay
- #(kweg = floor(min(table(seurM$orig.ident))/10)*10) #90
- options(future.globals.maxSize = 16000 * 1024^2)
- allseur_integrated = IntegrateData(anchorset = anchors, k.weight = 30, normalization.method = 'LogNormalize')
- saveRDS(allseur_integrated, '/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/Integrated_rpca_MOL_OPC_AllSpecies_ALLTissues_human_pr_coding_orthologs.RDS')
- # Clustering
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_1")
- pdf(paste0(pref, "_INTEGRATED_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- DefaultAssay(allseur_integrated) <- "integrated"
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_INTEGRATED_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_PRIMATE_ANNOT.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T, group.by = 'newannot') + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_INTEGRATED_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_INTEGRATED_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- saveRDS(allseur_integrated, paste0('/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_1/', pref, '_integrated_CLUSTERING1.RDS'))
- ################# CLUSTERING_2 ##########################
- ##Cluster#,CellType
- #17*,Neu+MOL
- #19*,MOL+Ast
- #25*,Neu+MOL
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_2/")
- toremove = c(17,19,25)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_2.RDS"))
- allseur_integrated= readRDS(paste0(pref,"_integrated_CLUSTERING_2.RDS"))
- ##################### CLUSTERING_3 #################
- #21*, MOL+micro
- #24*, neu+MOL
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_3/")
- toremove = c(21,24)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_3.RDS"))
- ############### CLUSTERING_4 #################
- #21*,MOL+OPC
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_4/")
- toremove = c(21)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_4.RDS"))
- ######################### ANNOTATE #################
- ###Custer#,cell_type_annotation
- #0,MOL
- #1,MOL
- #2,MOL
- #3,MOL
- #4,MOL
- #5,MOL
- #6,MOL
- #7,MOL
- #8,MOL
- #9,MOL
- #10,OPC
- #11,MOL
- #12,MOL
- #13,MOL
- #14,OPC
- #15,MOL
- #16,MOL
- #17,OPC
- #18,OPC
- #19,MOL
- #20,MOL
- #21,COP
- #22,OPC
- #23,OPC
- ## annotate
- (mapnames <-setNames(c("MOL","MOL","MOL","MOL","MOL","MOL","MOL","MOL","MOL",
- "MOL","OPC","MOL","MOL","MOL","OPC","MOL","MOL","OPC","OPC","MOL",
- "MOL","COP","OPC","OPC"),c(0:23)))
- # Create the broadannot
- allseur_integrated$MOL_OPC_annot = unname(mapnames[allseur_integrated[["seurat_clusters"]][,1]]) # to extract the annotations and not the cluster #s
- table(allseur_integrated$MOL_OPC_annot)
- # COP MOL OPC
- # 447 263193 28394
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/ANNOTATED")
- # QC PLOTS
- pdf(paste0(pref, "_Clusters.pdf"))
- DimPlot(allseur_integrated, group.by = 'MOL_OPC_annot', label = T, raster = T, label.size = 7) +
- NoLegend() +
- ggtitle('')
- dev.off()
- pdf(paste0(pref, "_Depth_Gene.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'MOL_OPC_annot', y = 'nFeature_RNA') +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_Depth_UMI.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'MOL_OPC_annot', y = 'nCount_RNA', ylim = c(0,20000)) +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- pdf(paste0(pref, "_NuclearFraction.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'MOL_OPC_annot', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # MARKER GENE PLOTS
- DefaultAssay(allseur_integrated) = 'RNA'
- seurM_corr = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_MarkersDotPlot.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, group.by = "MOL_OPC_annot", features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'MOL_OPC_annot', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'MOL_OPC_annot', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_ANNOTATED.RDS"))
- ### put these annots to the original seurat obj
- b = readRDS("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- # Match cell names between objects
- common_cells <- intersect(rownames([email hidden]), rownames([email hidden]))
- # Update annotations for those cells
- b$detailed_annot = b$newannot_2
- b$detailed_annot[common_cells] <- allseur_integrated$MOL_OPC_annot[common_cells]
- b$is_mol = ifelse(rownames([email hidden]) %in% common_cells, "yes", "no")
- # Identify cells in b that are not in the integrated object
- b_keep <- subset(
- b,
- subset = !(newannot_2 %in% c("MOL", "OPC", "COP")) | is_mol == "yes")
- )
- # check
- Idents(b_keep) = b_keep$detailed_annot
- DotPlot(b_keep, features = marksToPlot) + rotate_x_text(45)
- # save
- saveRDS(b_keep,"/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- ############################################
- ####### Astrocyte
- ################################################
- ### subset the genes and the MOL+OPC from each Species
- # subset cells and metadata
- # Subset Seurat object to only include 'Microglia' cells and exclude the specified 'id' values bcs they have <10 microglial cells
- seurM = subset(b_keep, subset = newannot_2 %in% c("Astrocyte")) # already orthologs only
- # Extract the count matrix from the Seurat object (default is "RNA" assay)
- mat = seurM@assays$RNA$counts
- new_mat <- mat[rownames(mat) %in% ortho_genes,]
- # generate new Seurat object
- glia_merged <- CreateSeuratObject(count = new_mat, meta.data = subset_b[[]])
- ####
- ## SET VARIABLES
- ####
- marksToPlot <- c('NRGN', 'SYT1', 'SNAP25', 'GAD1', 'GAD2', 'DRD1', 'DRD2', 'FOXP2', 'TAC1', 'PENK', 'SST', 'NPY', 'MOG', 'PLP1', 'MBP', 'MOBP', 'SLC1A2', 'SLC1A3', 'AQP4', 'CSF1R', 'ARHGAP15', 'PLD1','CASZ1', 'PTPRZ1', 'PCDH15', 'SOX6', 'FYN', 'BCAS1', 'ENPP6', 'GPR17', 'FLT1', 'DUSP1', 'COBLL1', 'CFTR', 'CHAT', 'ADARB2', 'LHX6', 'MEIS2', 'TAC3', 'VIP', 'TH', 'CCK', 'TTC34', 'CROCC', 'FHAD1', 'SOX4', 'PDGFRB', 'PDGFRA', 'SATB2', 'SLC17A7', 'EYA2', 'PDGFD', 'PTHLH', 'PVALB', 'SPARCL1', 'RSPO2', 'RMST', 'LMO3', 'TSHZ2', "CALB1", "CALB2", "NOS1", "TSHZ1", "OPRM1", "CHST9","GRM8", "PPP1R1B", 'GRIK3', 'CXCL14')
- pref = 'Astrocyte_AllSpecies_AllTissues_human_pr_coding_orthologs'
- ####
- ## INTEGRATE ACROSS SPECIES
- ####
- #### use old seurat version
- # Split data
- seurM = glia_merged
- seurM[["RNA"]] = as(object = seurM[["RNA"]], Class = "Assay")
- seurM$id = paste0(seurM$orig.ident, '_', seurM$Tissue)
- seurML = SplitObject(seurM, split.by = "id")
- # normalize and identify variable features for each dataset independently
- seurML = lapply(X = seurML, FUN = function(x) {
- x = NormalizeData(x)
- x = FindVariableFeatures(x, selection.method = "vst", nfeatures = 2000)
- })
- # select features that are repeatedly variable across datasets for integration run PCA on each
- # dataset using these features
- features = SelectIntegrationFeatures(object.list = seurML)
- print(paste0('Number of genes to use for integration: ', length(features)))
- seurML = lapply(X = seurML, FUN = function(x) {
- x = ScaleData(x, features = features, verbose = FALSE)
- x = RunPCA(x, features = features, verbose = FALSE, npcs = 30)
- })
- anchors = FindIntegrationAnchors(object.list = seurML, anchor.features = features, normalization.method = "LogNormalize", reduction = "rpca")
- # save memory prior integration
- gc()
- # this command creates an 'integrated' data assay
- #(kweg = floor(min(table(seurM$orig.ident))/10)*10) #90
- options(future.globals.maxSize = 16000 * 1024^2)
- allseur_integrated = IntegrateData(anchorset = anchors, k.weight = 30, normalization.method = 'LogNormalize')
- saveRDS(allseur_integrated, '/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/Integrated_rpca_Astrocyte_AllSpecies_ALLTissues_human_pr_coding_orthologs.RDS')
- # Clustering
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_ASTROCYTE/CLUSTERING_1")
- pdf(paste0(pref, "_INTEGRATED_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- DefaultAssay(allseur_integrated) <- "integrated"
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_INTEGRATED_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_PRIMATE_ANNOT.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T, group.by = 'newannot') + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_INTEGRATED_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_INTEGRATED_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- saveRDS(allseur_integrated, paste0('/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_MOL_OPC/CLUSTERING_1/', pref, '_integrated_CLUSTERING1.RDS'))
- ################# CLUSTERING_2 ##########################
- ##Cluster#,CellType
- #10*,MOL+Ast
- #12*,MOL+Ast
- #13*,Neu+Ast
- #15*,MOL+Ast
- #21*,MOL+Ast
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_ASTROCYTE/CLUSTERING_2/")
- toremove = c(10,12,13,15,21)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_2.RDS"))
- ################# CLUSTERING_2 ##########################
- ##Cluster#,CellType
- #14*,Neu+Ast
- #15*,MOL+Ast
- #17*,Endo+Ast
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_ASTROCYTE/CLUSTERING_3/")
- toremove = c(14,15,17)
- allseur_integrated = subset(allseur_integrated, subset = seurat_clusters %in% toremove, invert = T)
- # Clustering
- DefaultAssay(allseur_integrated) = 'integrated'
- allseur_integrated = ScaleData(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunPCA(allseur_integrated, verbose = FALSE)
- allseur_integrated = RunUMAP(allseur_integrated, dims = 1:20, reduction = 'pca')
- # Basic Plots
- pdf(paste0(pref, "_UMAP.pdf"))
- DimPlot(allseur_integrated, group.by = 'Species', raster = T)
- dev.off()
- allseur_integrated = FindNeighbors(allseur_integrated, dims = 1:20, reduction = 'pca')
- allseur_integrated = FindClusters(allseur_integrated, resolution = 1)
- pdf(paste0(pref, "_CLUSTERS.pdf"))
- DimPlot(allseur_integrated, label = T, raster = T) + NoLegend()
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'seurat_clusters', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # Plot previously identified markers
- DefaultAssay(allseur_integrated) = 'RNA'
- allseur_integrated = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "INTEGRATED_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_DEPTH.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'nFeature_RNA', color = 'black') + NoLegend() +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "INTEGRATED_NUCLEAR_FRACTION.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'seurat_clusters', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_CLUSTERING_3.RDS"))
- ############ ANNOTATE #############
- (mapnames <-setNames(rep("Astrocyte", 22),c(0:21)))
- # Create the broadannot
- allseur_integrated$Ast_annot = unname(mapnames[allseur_integrated[["seurat_clusters"]][,1]]) # to extract the annotations and not the cluster #s
- table(allseur_integrated$Ast_annot)
- #Astrocyte
- # 64660
- setwd("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/01_ASTROCYTE/ANNOTATED")
- # QC PLOTS
- pdf(paste0(pref, "_Clusters.pdf"))
- DimPlot(allseur_integrated, group.by = 'Ast_annot', label = T, raster = T, label.size = 7) +
- NoLegend() +
- ggtitle('')
- dev.off()
- pdf(paste0(pref, "_Depth_Gene.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Ast_annot', y = 'nFeature_RNA') +
- rotate_x_text(90)
- dev.off()
- pdf(paste0(pref, "_Depth_UMI.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Ast_annot', y = 'nCount_RNA', ylim = c(0,20000)) +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- pdf(paste0(pref, "_NuclearFraction.pdf"))
- ggboxplot(allseur_integrated[[]], x = 'Ast_annot', y = 'intronRat') +
- rotate_x_text(90) +
- theme(text=element_text(size=20, face = 'bold'))
- dev.off()
- # MARKER GENE PLOTS
- DefaultAssay(allseur_integrated) = 'RNA'
- seurM_corr = NormalizeData(allseur_integrated)
- pdf(paste0(pref, "_MarkersDotPlot.pdf"), width = 20, height = 10)
- DotPlot(allseur_integrated, group.by = "Ast_annot", features = marksToPlot) +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- stackedbarplot(allseur_integrated[[]], groupx = 'Ast_annot', groupfill = 'orig.ident', paste0(pref, '_STACK_SAMPLES'), wd = 20)
- stackedbarplot(allseur_integrated[[]], groupx = 'Ast_annot', groupfill = 'Species', paste0(pref, '_STACK_SPECIES'))
- # save
- saveRDS(allseur_integrated, paste0(pref,"_integrated_ANNOTATED.RDS"))
- ## add these annotations to the original seu obj
- # extract the newest annotations
- # Match cell names between objects
- b_keep$is_ast = ifelse(rownames([email hidden]) %in% rownames([email hidden]), "yes", "no")
- # Identify cells in b that are not in the integrated object
- b_keep_2 = subset(b_keep, subset = newannot_2 != "Astrocyte" | is_ast == "yes")
- # check
- Idents(b_keep_2) = b_keep_2$detailed_annot
- DotPlot(b_keep_2, features = marksToPlot) + rotate_x_text(45)
- # save
- saveRDS(b_keep_2,"/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- ####################### POST GLIA CLEANUP ###########################
- file_outdir = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/"
- species_list <- c("human", "chimp", "macaque", "marmoset", "mouse", "bat", "ferret")
- primates_list <- c("human", "chimp", "macaque", "marmoset")
- #### new obj AFTER GLIA CLEANUP
- b = readRDS("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- b$CellType = b$detailed_annot
- b$CellType = gsub("COP", "OPC", b$CellType)
- # put the interneuron subtype annotations to this seu obj
- interneurons_seu = readRDS("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/Interneurons_AllSpecies_AllTissues_ANNOTATED.RDS")
- ## add these annotations to the original seu obj
- # extract the newest annotations
- # Match cell names between objects
- b$is_interneu = ifelse(rownames([email hidden]) %in% rownames([email hidden]), "yes", "no")
- # Identify cells in b that are not in the integrated object
- b_keep = subset(b, subset = detailed_annot != "Non_SPN" | is_interneu == "yes")
- # order the rownames of smaller seu obj according to the bigger one
- ordered_meta_interneurons_seu = [email hidden][match(rownames([email hidden]), rownames([email hidden])),]
- b_keep$is_interneu = ordered_meta_interneurons_seu$newannot
- b_keep$a = ifelse(is.na(b_keep$is_interneu), as.character(b_keep$detailed_annot), as.character(b_keep$is_interneu))
- # check
- table(b_keep$detailed_annot, b_keep$a)
- # re-assign
- b_keep$broadannot_2 = b_keep$detailed_annot
- b_keep$detailed_annot = b_keep$a
- # orgnize metadata
- b_keep$is_micro = NULL
- b_keep$is_mol = NULL
- b_keep$keep_cell = NULL
- b_keep$is_ast = NULL
- b_keep$is_interneu = NULL
- b_keep$a = NULL
- b_keep$broad_annot = NULL
- b_keep$tissue = NULL
- b_keep$id = paste0(b_keep$orig.ident, "_", b_keep$Tissue)
- ### remove human caudate sample
- b_keep_2 = subset(b_keep, subset = id %in% "Sample_242999_Caudate", invert = T)
- #### find the sex and age for all
- orig_meta <- [email hidden] # safe copy
- ### add info to the metadata
- a = orig_meta
- a[grep("v324", a$orig.ident),]$Sex = "Male"
- a[grep("v321", a$orig.ident),]$Sex = "Female"
- a$Sex = sub("female", "Female", a$Sex)
- a$Sex = sub("male", "Male", a$Sex)
- a[grep("Bat", a$Species),]$Sex = "Male"
- a[grep("Bat", a$Species),]$Age = "3y"
- a[grep("marm027", a$orig.ident),]$Sex = "Male"
- a[grep("marm028", a$orig.ident),]$Sex = "Female"
- a[grep("marm029", a$orig.ident),]$Sex = "Male"
- a[grep("marm027", a$orig.ident),]$Age = "2y4m"
- a[grep("marm028", a$orig.ident),]$Age = "3y2m"
- a[grep("marm029", a$orig.ident),]$Age = "2y6m"
- a[grep("v324", a$orig.ident),]$Age = "2035d"
- a[grep("v321", a$orig.ident),]$Age = "2051d"
- #a[grep("SRR13808459", a$orig.ident),]$Sex = "Female"
- a[grep("SRR13808462", a$orig.ident),]$Sex = "Female"
- a[grep("SRR13808463", a$orig.ident),]$Sex = "Female"
- #a[grep("SRR13808466", a$orig.ident),]$Sex = "Female"
- #a[grep("SRR13808467", a$orig.ident),]$Sex = "Female"
- a[grep("SRR11921037", a$orig.ident),]$Sex = "Female"
- a[grep("SRR11921038", a$orig.ident),]$Sex = "Female"
- a[grep("SRR11921037", a$orig.ident),]$Age = "42d"
- a[grep("SRR11921038", a$orig.ident),]$Age = "42d"
- patterns <- c(paste0("SRR1192100", 5:9), paste0("SRR119210", 10:12))
- pattern_all <- paste(patterns, collapse = "|")
- idx <- grep(pattern_all, a$orig.ident)
- a[idx, ]$Age <- "70d"
- a[idx, ]$Sex = "Male"
- a$sex = NULL
- a$age = NULL
- a$cellbarc = str_split_i(rownames(a), "-", 1)
- a$cellID = paste0(a$cellbarc, "_", a$id)
- a$Sex = sub("FeMale", "Female", a$Sex)
- ### add humanized ages
- # read species traits taken from https://genomics.senescence.info/species/entry.php?species=Mus_musculus, https://genomics.senescence.info/species/entry.php?species=Mustela_nigripes, https://genomics.senescence.info/species/entry.php?species=Phyllostomus_hastatus, https://animaldiversity.org/accounts/Phyllostomus_hastatus/, https://genomics.senescence.info/species/entry.php?species=Phyllostomus_discolor
- outdir = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/revision/"
- species_traits = read.csv(paste0(outdir, "primate_lifeTraits.csv"))
- sample_metadata = read.csv(paste0(outdir, "comparative_striatum_sample_metadata.csv"), sep = "\t")
- sample_metadata$id = as.character(paste0(sample_metadata$Barcode.ID, "_", sample_metadata$Brain.Area))
- a$id = as.character(a$id)
- # add the ages in years to the single cell-level metadata
- meta_df <- merge(a, sample_metadata, by = "id", all.x = TRUE)
- meta_df$Age <- meta_df$Age.x
- meta_df$Sex <- meta_df$Sex.x
- meta_df$Species <- meta_df$Species.x
- meta_df <- meta_df[, !grepl("\\.x$|\\.y$", names(meta_df))]
- meta_df$Brain.Area = NULL
- # Add orig.ident only if rownames are not already "cellbarc-orig.ident"
- rownames(meta_df) <- ifelse(
- grepl("[-_]", meta_df$cellbarc), # check for presence of "-" or "_"
- meta_df$cellbarc, # if present, keep as is
- paste0(meta_df$cellbarc, "-", meta_df$orig.ident) # else, append orig.ident
- )
- # restore original cell order (this prevents scrambling)
- meta_df <- meta_df[rownames(a), ]
- # Check if same set but different order
- setequal(rownames(meta_df), rownames(a))
- humanize_all_ages <- function(meta_df, species_traits) {
- # Initialize vector for humanized ages (same length and order)
- humanized_ages <- numeric(nrow(meta_df))
- # Get unique species in the metadata
- species_list <- unique(meta_df$Species)
- for (species in species_list) {
- # Convert ages to numerical
- species_ages <- meta_df$Age_in_years[meta_df$Species == species]
- species_ages_num <- as.numeric(species_ages)
- if (species %in% c("Human")) {
- # For humans, humanized age = original age
- humanized_ages[meta_df$Species == species] <- species_ages_num
- } else {
- # Check if species exists in lifetraits
- if (species %in% colnames(species_traits)) {
- # Build linear model: species ~ Human (lifetraits)
- model <- lm(species_traits[[species]] ~ species_traits$Human)
- # Calculate humanized ages using the model coefficients
- humanized <- (species_ages_num - model$coefficients[1]) / model$coefficients[2]
- # Assign back to correct rows in humanized_ages vector
- humanized_ages[meta_df$Species == species] <- humanized
- } else {
- warning(paste("Species", species, "not found in lifetraits. Assigning NA"))
- humanized_ages[meta_df$Species == species] <- NA
- }
- }
- }
- return(humanized_ages)
- }
- meta_df$Humanized_age = humanize_all_ages(meta_df, species_traits)
- # update the metadata of the seu obj
- # Join by cell ID (ensure cell names match exactly)
- [email hidden] = meta_df
- # check
- marksToPlot <- c('NRGN', 'SYT1', 'SNAP25', 'GAD1', 'GAD2', 'DRD1', 'DRD2', 'FOXP2', 'TAC1', 'PENK', 'SST', 'NPY', 'MOG', 'PLP1', 'MBP', 'MOBP', 'SLC1A2', 'SLC1A3', 'APOE', 'GPC5', 'GLI3', 'AQP4', 'CSF1R', 'ARHGAP15', 'PLDL1', 'CASZ1', 'PTPRZ1', 'PCDH15', 'SOX6', 'FYN', 'BCAS1', 'ENPP6', 'GPR17', 'FLT1', 'DUSP1', 'COBLL1', 'CFTR', 'CHAT', 'ADARB2', 'LHX6', 'MEIS2', 'TAC3', 'VIP', 'TH', 'CCK', 'TTC34', 'CROCC', 'FHAD1', 'SOX4', 'PDGFRB', 'PDGFRA', 'SATB2', 'SLC17A7', 'EYA2', 'PDGFD', 'PTHLH', 'PVALB', 'SPARCL1', 'RSPO2', 'RMST', 'LMO3', 'TSHZ2', "CALB1", "CALB2", "NOS1", "TSHZ1", "OPRM1", "CHST9","GRM8", "PPP1R1B", 'GRIK3', 'CXCL14')
- # Plot previously identified markers
- DefaultAssay(b_keep_2) = 'RNA'
- b_keep_2 = NormalizeData(b_keep_2)
- pdf(paste0("merged_clean_AllCellTypes_AllTissues_seu_detailedAnnot_DOTPLOT.pdf"), width = 20, height = 10)
- DotPlot(b_keep_2, features = marksToPlot, group.by = "detailed_annot") +
- xlab('') +
- ylab('') +
- theme(axis.text.x=element_text(size=20, face = 'bold'),
- axis.text.y=element_text(size=20, face = 'bold'),
- legend.text = element_text(size=20, face = 'bold'),
- legend.title = element_text(size=20, face = 'bold')) +
- rotate_x_text(45)
- dev.off()
- # save
- saveRDS(b_keep_2,"/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- ##### generate pseudobulk counts
- file_outdir = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/"
- species_list <- c("human", "chimp", "macaque", "marmoset", "mouse", "bat", "ferret")
- primates_list <- c("human", "chimp", "macaque", "marmoset")
- #### new obj AFTER GLIA CLEANUP
- b = readRDS("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/merged_clean_AllCellTypes_AllTissues_seu.RDS")
- sub_obj_new_list = SplitObject(b, split.by = "Species")
- names(sub_obj_new_list) = tolower(names(sub_obj_new_list))
- meta_df = [email hidden]
- ### run function to generate pseudobulk count matrices for each species and celltype. Keep cellTypes which are present in at least 3 of the samples for a given Species-Tissue as well as genes with at least 1 copy of at least 3 of the samples for a given Species-Tissue. Generate the pseudobulk count matrix for all species-cellTypes except ferret which has only 2 samples using utilities function: pseudobulk_species = function(seurObj, ctype, features = c())
- pseudobulk_species
- function(seurObj, ctype, n_copy, sample_number){
- require(plyr)
- require(dplyr)
- require(tidyverse)
- require(tidyr)
- require(Seurat)
- require(ggpubr)
- require(reshape2)
- require(data.table)
- require(rio)
- require(scran)
- require(scater)
- require(SingleCellExperiment)
- require(edgeR)
- require(DESeq2)
- if(!('Sample' %in% colnames(seurObj[[]]))){stop('Please put your donors into Sample column')}
- if(!('CellType' %in% colnames(seurObj[[]]))){stop('Please put your cell types into CellType column')}
- # Load RNA
- meta = [email hidden]
- metakeep = c('CellType', 'Sample', 'Species')
- meta = meta[,metakeep]
- # Subset the given cell type
- subSeur = subset(seurObj, subset = CellType %in% ctype)
- # Create SCE assay
- DefaultAssay(subSeur) = 'RNA'
- subSCE = as.SingleCellExperiment(subSeur)
- # Pseudobulk SCE assay
- subGroups = colData(subSCE)[, c("Species", "Sample", "CellType")]
- subPseudoSCE = sumCountsAcrossCells(subSCE, subGroups)
- subGroups = colData(subPseudoSCE)[, c("Species", "Sample", "CellType")]
- subGroups$Sample = factor(subGroups$Sample)
- subAggMat = subPseudoSCE@assays@data$sum
- colnames(subAggMat) = subGroups$Sample
- # binarize the matrix
- binary_subAggMat = as.matrix((subAggMat >= n_copy) + 0) # genes with at least "n_copy" copy(ies) will be 1
- exp_list = list()
- # Keep genes expressed in at least "sample_number" samples of any Tissue
- for(tissue in unique(subSeur$Tissue)){
- tissue_columns = grepl(tissue, colnames(binary_subAggMat), ignore.case = TRUE)
- exp_list[[tissue]] = which(rowSums(binary_subAggMat[, tissue_columns]) >= sample_number) %>% names
- }
- expgns = Reduce(intersect, exp_list)
- cat("Length of genes", unique(subSeur$Species), ctype, "=", length(expgns), "\n")
- subAggMat = subAggMat[expgns,]
- return(subAggMat)
- }
- pb_count_list = list()
- for (sp in species_list) {
- sub_obj_new = sub_obj_new_list[[sp]]
- celltypes = unique(sub_obj_new$CellType)
- for (ctype in celltypes) {
- if (sp != "ferret") {
- pb_count_list[[sp]][[ctype]] <- pseudobulk_species(
- seurObj = sub_obj_new,
- ctype = ctype,
- n_copy = 1,
- sample_number = 3
- )
- } else {
- pb_count_list[["ferret"]][[ctype]] <- NULL
- }
- }
- }
- #save
- saveRDS(pb_count_list, file = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/pseudobulk_count_matrices_AllSpecies_AllCellTypes_broadannot_PostgliaCleanup.RDS")
- ### Genes were filtered to prep the pseudobulk counts (> 1 copy in >= 3 samples of that Species-Tissue) To continue with downstream analysis, get only the common genes across primates
- common_genes_per_celltype = list()
- primate_pb_count_list = pb_count_list[primates_list]
- for (ctype in names(primate_pb_count_list$human)) {
- # Extract the gene vectors for this cell type across all species
- genes_across_species <- lapply(primate_pb_count_list, function(species_count) {
- rownames(species_count[[ctype]])
- })
- # Remove any NULLs (if some species lack this cell type)
- genes_across_species <- genes_across_species[!sapply(genes_across_species, is.null)]
- # Take intersection across species
- common_genes_per_celltype[[ctype]] <- Reduce(intersect, genes_across_species)
- }
- # Inspect results
- str(common_genes_per_celltype)
- # save
- saveRDS(common_genes_per_celltype, file = "/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/pseudobulk_common_genes_withinPrimates_per_celltype_broadannot_postgliaCleanup.RDS")
- common_genes_per_celltype = readRDS("/project/Neuroinformatics_Core/Konopka_lab/s422071/workdir_pr/s422071/projects/comparative_striatum/seu_objs/pseudobulk_common_genes_withinPrimates_per_celltype_broadannot_postgliaCleanup.RDS")
- ##### subset each count matrix with these genes
- sub_seu_pb_count_list = list()
- pb_meta_list = list()
- meta_df_list = list()
- ## Generate singleCellExperiment object
- for (sp in primates_list){
- for (t in unique(sub_obj_new_list[[sp]]$Tissue)){
- # change COP to OPC
- sub_obj_new_list[[sp]]$CellType <- gsub("COP", "OPC", sub_obj_new_list[[sp]]$CellType)
- for (ctype in unique(sub_obj_new_list[[sp]]$CellType)){
- # get the seu obj
- seu = sub_obj_new_list[[sp]]
- # subset for the given cell type and tissue
- sub_seu = subset(seu, subset = CellType %in% ctype & Tissue %in% t)
- # subset pseudobulk count mat with common genes
- sub_seu_pb_count = primate_pb_count_list[[sp]][[ctype]][, grepl(t, colnames(primate_pb_count_list[[sp]][[ctype]]))]
- # subset the pseudocount matrix with common genes for each celltype
- sub_seu_pb_count_list[[t]][[ctype]][[sp]] = sub_seu_pb_count[common_genes_per_celltype[[ctype]],]
- # generate pseudobulk metadata
- sub_seu_meta <- [email hidden] %>%
- as.data.frame() %>%
- dplyr::select(Species, Sample, Tissue, Sex, Age, Age_in_years, Humanized_age, CellType, id, nCount_RNA)
- ## calculate the sum of the total UMI at the log scale for each sample
- tibble_meta <- as_tibble(sub_seu_meta) %>%
- dplyr::group_by(Sample) %>%
- dplyr::summarise(
- log10_sum_ncount_RNA = log10(sum(nCount_RNA, na.rm = TRUE)),
- n = dplyr::n(),
- .groups = "drop"
- )
- tibble_meta = as.data.frame(tibble_meta)
- # aggregate based on sample
- a <- sub_seu_meta %>% distinct(Sample, .keep_all = TRUE)
- rownames(a) = a$Sample
- # merge two dataframes
- meta_df_sample = merge(a,tibble_meta ,by = "Sample", all.x = TRUE)
- rownames(meta_df_sample) = meta_df_sample$Sample
- # Check matching of matrix columns and metadata rows
- all(colnames(sub_seu_pb_count) == rownames(meta_df_sample))
- # if it says FALSE due to order not being the same. Order:
- # order the metadata rownames based on the col names of the count matrix
- #rownames(meta_df_sample) <- rownames(meta_df_sample)[order(match(rownames(meta_df_sample), colnames(sub_seu_pb_count)))]
- # Check matching of matrix columns and metadata rows
- #all(colnames(sub_seu_pb_count) == rownames(meta_df_sample))
- # save in a list
- pb_meta_list[[t]][[ctype]][[sp]] = meta_df_sample
- meta_df_list[[t]][[ctype]][[sp]] = [email hidden]
- #rownames(meta_df_list[[t]][[ctype]][[sp]]) <- meta_df_list[[t]][[ctype]][[sp]]$Sample
- # get sub_seu_pb_count_list for the next step
- }}}
- saveRDS(pb_meta_list, file = paste0(file_outdir,"primates_pb_meta_list_PostgliaCleanup.RDS"))
- saveRDS(meta_df_list, file = paste0(file_outdir,"primates_meta_df_list_PostgliaCleanup.RDS"))
- saveRDS(sub_seu_pb_count_list, file = paste0(file_outdir,"primates_sub_seu_pb_count_list_PostgliaCleanup.RDS"))
- ### generate a list of metadata and pseudobulk counts for the same tissue and celltype
- merged_pb_meta <- list()
- merged_sub_seu_pb_count <- list()
- merged_meta_df <- list()
- for (t in names(pb_meta_list)){
- for (ctype in names(pb_meta_list[[t]])){
- # Combine all species' metadata rows
- merged_pb_meta[[t]][[ctype]] <- do.call(rbind, pb_meta_list[[t]][[ctype]])
- merged_sub_seu_pb_count[[t]][[ctype]] <- do.call(cbind, sub_seu_pb_count_list[[t]][[ctype]])
- merged_meta_df[[t]][[ctype]] <- do.call(rbind, meta_df_list[[t]][[ctype]])
- }}
- # save
- saveRDS(merged_pb_meta, file = paste0(file_outdir,"primates_merged_pb_meta_PostgliaCleanup.RDS"))
- saveRDS(merged_sub_seu_pb_count, file = paste0(file_outdir,"primates_merged_sub_seu_pb_count_PostgliaCleanup.RDS"))
- saveRDS(merged_meta_df, file = paste0(file_outdir,"primates_merged_meta_df_PostgliaCleanup.RDS"))
- ############################################################
- sessionInfo()
- R version 4.2.3 (2023-03-15)
- Platform: x86_64-conda-linux-gnu (64-bit)
- Running under: Red Hat Enterprise Linux Server 7.9 (Maipo)
- Matrix products: default
- BLAS/LAPACK: /cm/shared/apps/rstudio-desktop/2022.12.0/lib/libopenblasp-r0.3.21.so
- locale:
- [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
- [3] LC_TIME=en_US.UTF-8 LC_COLLATE=en_US.UTF-8
- [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
- [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
- [9] LC_ADDRESS=C LC_TELEPHONE=C
- [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
- attached base packages:
- [1] stats4 stats graphics grDevices utils datasets methods
- [8] base
- other attached packages:
- [1] fastDummies_1.7.3 flashClust_1.01-2
- [3] WGCNA_1.73 fastcluster_1.3.0
- [5] dynamicTreeCut_1.63-1 curl_5.2.2
- [7] Matrix.utils_0.9.8 variancePartition_1.28.9
- [9] BiocParallel_1.32.5 edgeR_3.40.2
- [11] limma_3.54.0 ggrepel_0.9.3
- [13] DESeq2_1.38.3 harmony_1.1.0
- [15] Rcpp_1.0.10 data.table_1.14.8
- [17] rio_1.0.1 reshape2_1.4.4
- [19] ggpubr_0.6.0 lubridate_1.9.3
- [21] forcats_1.0.0 stringr_1.5.0
- [23] purrr_1.0.1 readr_2.1.4
- [25] tidyr_1.3.0 tibble_3.2.1
- [27] tidyverse_2.0.0 plyr_1.8.9
- [29] ggplot2_3.4.4 Matrix_1.6-5
- [31] dplyr_1.1.4 DropletQC_0.0.0.9000
- [33] DropletUtils_1.18.1 SingleCellExperiment_1.20.0
- [35] SummarizedExperiment_1.28.0 Biobase_2.58.0
- [37] GenomicRanges_1.50.0 GenomeInfoDb_1.34.9
- [39] IRanges_2.32.0 S4Vectors_0.36.0
- [41] BiocGenerics_0.44.0 MatrixGenerics_1.10.0
- [43] matrixStats_0.63.0 rhdf5_2.42.1
- [45] BPCells_0.3.1 Seurat_5.3.0
- [47] SeuratObject_5.1.0 sp_1.6-0
- [49] patchwork_1.3.0.9000
- loaded via a namespace (and not attached):
- [1] scattermore_1.2 R.methodsS3_1.8.2
- [3] knitr_1.48 bit64_4.0.5
- [5] irlba_2.3.5.1 DelayedArray_0.24.0
- [7] R.utils_2.12.2 rpart_4.1.19
- [9] KEGGREST_1.38.0 RCurl_1.98-1.12
- [11] doParallel_1.0.17 generics_0.1.3
- [13] preprocessCore_1.60.2 RhpcBLASctl_0.23-42
- [15] cowplot_1.1.1 RSQLite_2.3.1
- [17] RANN_2.6.1 future_1.33.1
- [19] bit_4.0.5 tzdb_0.3.0
- [21] spatstat.data_3.0-0 httpuv_1.6.9
- [23] xfun_0.47 hms_1.1.2
- [25] evaluate_0.20 promises_1.2.0.1
- [27] fansi_1.0.6 progress_1.2.2
- [29] caTools_1.18.2 igraph_1.4.2
- [31] DBI_1.1.3 geneplotter_1.76.0
- [33] htmlwidgets_1.6.2 spatstat.geom_3.0-6
- [35] ellipsis_0.3.2 RSpectra_0.16-1
- [37] backports_1.4.1 annotate_1.76.0
- [39] aod_1.3.3 deldir_1.0-6
- [41] sparseMatrixStats_1.10.0 vctrs_0.6.5
- [43] ROCR_1.0-11 abind_1.4-5
- [45] cachem_1.0.7 withr_2.5.2
- [47] grr_0.9.5 progressr_0.13.0
- [49] checkmate_2.3.0 sctransform_0.4.1
- [51] prettyunits_1.1.1 mclust_6.0.0
- [53] goftest_1.2-3 cluster_2.1.4
- [55] dotCall64_1.1-0 lazyeval_0.2.2
- [57] crayon_1.5.2 spatstat.explore_3.0-6
- [59] pkgconfig_2.0.3 nlme_3.1-162
- [61] nnet_7.3-18 rlang_1.1.4
- [63] globals_0.16.2 lifecycle_1.0.3
- [65] miniUI_0.1.1.1 polyclip_1.10-4
- [67] RcppHNSW_0.4.1 lmtest_0.9-40
- [69] carData_3.0-5 Rhdf5lib_1.20.0
- [71] boot_1.3-28.1 zoo_1.8-11
- [73] base64enc_0.1-3 ggridges_0.5.4
- [75] png_0.1-8 viridisLite_0.4.1
- [77] bitops_1.0-7 R.oo_1.25.0
- [79] KernSmooth_2.23-20 spam_2.10-0
- [81] rhdf5filters_1.10.1 Biostrings_2.66.0
- [83] blob_1.2.3 DelayedMatrixStats_1.20.0
- [85] parallelly_1.35.0 spatstat.random_3.1-3
- [87] remaCor_0.0.16 rstatix_0.7.2
- [89] ggsignif_0.6.4 beachmat_2.14.0
- [91] scales_1.3.0 memoise_2.0.1
- [93] magrittr_2.0.3 ica_1.0-3
- [95] gplots_3.1.3 zlibbioc_1.44.0
- [97] compiler_4.2.3 dqrng_0.3.0
- [99] RColorBrewer_1.1-3 lme4_1.1-35.1
- [101] fitdistrplus_1.1-8 cli_3.6.2
- [103] XVector_0.38.0 listenv_0.9.0
- [105] pbapply_1.7-0 htmlTable_2.4.2
- [107] Formula_1.2-5 MASS_7.3-58.3
- [109] tidyselect_1.2.0 stringi_1.7.12
- [111] locfit_1.5-9.8 grid_4.2.3
- [113] tools_4.2.3 timechange_0.2.0
- [115] future.apply_1.10.0 parallel_4.2.3
- [117] rstudioapi_0.14 foreign_0.8-85
- [119] foreach_1.5.2 gridExtra_2.3
- [121] EnvStats_2.8.1 farver_2.1.1
- [123] Rtsne_0.16 digest_0.6.31
- [125] shiny_1.7.4 car_3.1-2
- [127] broom_1.0.3 scuttle_1.8.0
- [129] later_1.3.0 RcppAnnoy_0.0.20
- [131] httr_1.4.5 AnnotationDbi_1.60.2
- [133] Rdpack_2.6 colorspace_2.1-0
- [135] XML_3.99-0.14 tensor_1.5
- [137] reticulate_1.42.0 splines_4.2.3
- [139] uwot_0.1.14 spatstat.utils_3.1-4
- [141] plotly_4.10.1 xtable_1.8-4
- [143] jsonlite_1.8.4 nloptr_2.0.3
- [145] R6_2.5.1 Hmisc_5.1-1
- [147] pillar_1.9.0 htmltools_0.5.8.1
- [149] mime_0.12 glue_1.6.2
- [151] fastmap_1.1.1 minqa_1.2.6
- [153] codetools_0.2-19 mvtnorm_1.2-3
- [155] utf8_1.2.4 lattice_0.21-8
- [157] spatstat.sparse_3.0-0 pbkrtest_0.5.2
- [159] gtools_3.9.4 GO.db_3.16.0
- [161] survival_3.5-3 rmarkdown_2.28
- [163] munsell_0.5.0 GenomeInfoDbData_1.2.9
- [165] iterators_1.0.14 impute_1.72.3
- [167] HDF5Array_1.26.0 gtable_0.3.3
- [169] rbibutils_2.2.16
00_prep_pseudobulk_counts.R at commit b8e3704, under MIT · at the source
Overview
- Department of Neurobiology, UCLA David Geffen School of Medicine, Los Angeles, CA USA
- Department of Neuroscience, Peter O’Donnell Jr. Brain Institute, UT Southwestern Medical Center, Dallas, TX USA
- Present Address: Division of Genetics and Genomics, Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School, Boston, MA USA
- School of Biology, University of St Andrews, St Andrews, UK
- Keeling Center for Comparative Medicine and Research, The University of Texas, MD Anderson Cancer Center, Bastrop, TX USA
- Department of Anthropology and Center for the Advanced Study of Human Paleobiology, The George Washington University, Washington, DC USA
Abstract
The dorsal striatum is important for highly specialized functions including movement, learning, and habit formation. However, it is not known if species-specialized behaviors are associated with cellular specializations in the striatum. Here, we compared single-nucleus RNA sequencing (snRNA-seq) data from human, chimpanzee, rhesus macaque, common marmoset, and pale spear-nosed bat caudate (CN) and putamen (Pu) separately as well as mouse caudoputamen (C-Pu), which represents divergence among species spanning approximately 94 million years of evolution. We observed a lower neuron-to-glia ratio in primate striata compared to non-primates, reflecting the allometric scaling of neuron density and relative glia density invariance in larger brains. Among neurons, eccentric spiny projection neurons (eSPNs) - an SPN of unknown function - showed significantly lower proportions in non-primate striata for both CN and Pu. Focusing on the heterogeneity within interneurons, we identified two bat striatal interneuron cell types that are nearly absent in other species: which express LMO3, and co-express FOXP2 and TSHZ2. Other striatal interneurons also exhibited differential abundance between primates and non-primates. In summary, we provide a comprehensive snRNA-seq dataset of dorsal striatum, identify two distinct, previously uncharacterized populations of bat interneurons, and uncover fundamental cellular composition differences between primate and non-primate striata.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 23 matches between paragraphs and lines of code.
konopkalab/Comparative_striatum
b8e37042199a12756ee1ae145790fc9862f6cfc6, 15 April 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
86 files
- 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_CHIMP.s , Shell, 22 linesh - 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_MACAQUE , Shell, 30 lines.sh - 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_MARMOSE , Shell, 22 linesT.sh - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_CHIMP.sh , Shell, 48 lines, 1 match - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_MACAQUE.sh , Shell, 48 lines - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_MARMOSET.s , Shell, 48 linesh - 02_Bat/
00_DEMULTIPLEX.sh , Shell, 145 lines - 02_Bat/
01_GENERATE_COUNT_MATRIX , Shell, 49 lines.sh - 02_Bat/
02_CELLBENDER.sh , Shell, 13 lines - 02_Bat/
03_ANNOTATION.R , R, 950 lines - 02_Bat/
04_WGCNA.R , R, 547 lines, 1 match - 02_Chimp/
00_mkfastq_1.sh , Shell, 54 lines - 02_Chimp/
00_mkfastq_2.sh , Shell, 54 lines - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 51 linesP.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 54 linesP_2.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 53 linesP_3.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 44 linesP_4.sh - 02_Chimp/
02_CELLBENDER.sh , Shell, 13 lines - 02_Chimp/
03_ANNOTATE_CHIMP_CAUDAT , R, 396 linesE.R - 02_Chimp/
03_ANNOTATE_CHIMP_PUTAME , R, 391 linesN.R - 02_Chimp/
04_INTERNEURON_CLUSTER.R , R, 292 lines - 02_Chimp/
05_NEURON_CLUSTER.R , R, 636 lines - 02_Ferret/
00_CONVERT_BAM_TO_10X_FA , Shell, 72 linesSTQ.sh - 02_Ferret/
01_GENERATE_COUNT_MATRIX , Shell, 48 lines.sh - 02_Ferret/
02_PREP_FOR_CELLBENDER.R , R, 118 lines - 02_Ferret/
03_CELLBENDER.sh , Shell, 43 lines - 02_Ferret/
04_ANNOTATION.R , R, 325 lines - 02_Human/
00_mkfastq.sh , Shell, 40 lines - 02_Human/
01_COUNT_CELLRANGER_1.sh , Shell, 58 lines - 02_Human/
01_COUNT_CELLRANGER_2.sh , Shell, 61 lines - 02_Human/
01_COUNT_CELLRANGER_3.sh , Shell, 59 lines - 02_Human/
01_COUNT_CELLRANGER_4.sh , Shell, 54 lines - 02_Human/
01_COUNT_CELLRANGER_5.sh , Shell, 54 lines - 02_Human/
02_CELLBENDER.sh , Shell, 19 lines - 02_Human/
03_CAUDATEPUTAMEN_CLUSTE , R, 356 linesR.R - 02_Human/
04_NEURON_CLUSTER.R , R, 481 lines - 02_Human/
05_INTERNEURON_CLUSTER.R , R, 226 lines - 02_Human/
06_WGCNA.R , R, 510 lines, 3 matches - 02_Macaque/
01_CELLRANGER_COUNT_1.sh , Shell, 66 lines - 02_Macaque/
02_CELLBENDER.sh , Shell, 13 lines - 02_Macaque/
03_ANNOTATION_MACAQUE_CA , R, 525 linesUDATE.R - 02_Macaque/
03_ANNOTATION_MACAQUE_PU , R, 306 linesTAMEN.R - 02_Macaque/
04_INTEGRATION_CAUDATE.R , R, 212 lines - 02_Macaque/
04_INTEGRATION_PUTAMEN.R , R, 219 lines - 02_Macaque/
05_CAUDATE_PUTAMEN_INTEG , R, 408 linesRATE_ANNOTATE.R - 02_Macaque/
He_et_al_Macaque/ , Shell, 74 lines00_fasterq_dump.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_1.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_2.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_3.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 13 lines02_CELLBENDER.sh - 02_Macaque/
He_et_al_Macaque/ , R, 361 lines03_ANNOTATION_CAUDATE.R - 02_Macaque/
He_et_al_Macaque/ , R, 445 lines03_ANNOTATION_PUTAMEN.R - 02_Marmoset/
03_INTEGRATE_ANNOTATE_DA , R, 538 linesTASETS.R - 02_Marmoset/
Krienen_et_al/ , Shell, 34 lines00_DOWNLOAD_FASTQ_FROM_U RL.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 81 lines01_ARRANGE_FASTQs.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 69 lines02_GENERATE_COUNT_MATRIC ES.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 11 lines03_CELLBENDER.sh - 02_Marmoset/
Krienen_et_al/ , R, 793 lines04_ANNOTATION.R - 02_Marmoset/
Lin_et_al/ , Shell, 77 lines00_DOWNLOAD_FASTQs.sh - 02_Marmoset/
Lin_et_al/ , Shell, 45 lines01_GENERATE_COUNT_MATRIX .sh - 02_Marmoset/
Lin_et_al/ , Shell, 11 lines02_CELLBENDER.sh - 02_Marmoset/
Lin_et_al/ , R, 538 lines03_ANNOTATE.R - 02_Mouse/
01_GENERATE_COUNT_MATRIX , Shell, 51 lines_NEXTFLOW.sh - 02_Mouse/
02_PREPARE_FOR_CELLBENDE , R, 49 linesR.R - 02_Mouse/
03_CELLBENDER.sh , Shell, 15 lines - 02_Mouse/
04_CAUDOPUTAMEN_ANNOTATE , R, 279 lines.R - 03_SPNs/
00_get_orthologs.sh , Shell, 90 lines, 1 match - 03_SPNs/
01_INTEGRATION_AND_CLUST , R, 2,281 lines, 1 matchERING.R - 03_SPNs/
02_eSPN_markers.R , R, 444 lines, 1 match - 03_SPNs/
03_eSPN_D1D2_hybrid_corr , R, 186 lines, 2 matcheselation.R - 03_SPNs/
04_SPN_seurat_obj_conver , R, 32 linest_seurat_obj_to_AnnData. R - 03_SPNs/
05_Cell_type_proportion_ , Jupyter, 460 linesanalysis_scCODA.ipynb - 04_Interneurons/
01_Integration& , R, 3,782 lines, 2 matchesClustering.R - 04_Interneurons/
02_Interneurons_scCODA.i , Jupyter, 586 linespynb - 04_Interneurons/
03_DEG_Analysis.R , R, 467 lines, 1 match - 04_Interneurons/
04_Correlation_withinBat , R, 104 lines, 1 matchPutamen.R - 04_Interneurons/
05_Fig5A_scCODA.py , Python, 77 lines - 05_DGE_analysis/
00_prep_pseudobulk_count , R, 1,667 lines, 3 matchess.R - 05_DGE_analysis/
01_primates_pseudobulk_D , R, 353 lines, 2 matchesESeq2.R - 05_DGE_analysis/
02_DEG_figures.R , R, 827 lines, 1 match - GLIA_TO_NEURON.R, R, 517 lines
- QC_plots.R, R, 1,047 lines, 1 match
- python_functions.py, Python, 394 lines
- scCODA_Glia_to_Neuron.py
, Python, 86 lines - LICENSE, License, 21 lines
- README.md, Text, 15 lines
Zenodo 19266844
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
- 28 September 2026: the link answers (HTTP 200)
86 files
- 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_CHIMP.s , Shell, 22 linesh - 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_MACAQUE , Shell, 30 lines.sh - 01_Generate_NHP_gtfs/
LIFTOFF_HUMAN_TO_MARMOSE , Shell, 22 linesT.sh - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_CHIMP.sh , Shell, 48 lines - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_MACAQUE.sh , Shell, 48 lines - 01_Generate_NHP_gtfs/
MKREF_LIFTOFF_MARMOSET.s , Shell, 48 linesh - 02_Bat/
00_DEMULTIPLEX.sh , Shell, 145 lines - 02_Bat/
01_GENERATE_COUNT_MATRIX , Shell, 49 lines.sh - 02_Bat/
02_CELLBENDER.sh , Shell, 13 lines - 02_Bat/
03_ANNOTATION.R , R, 950 lines - 02_Bat/
04_WGCNA.R , R, 547 lines - 02_Chimp/
00_mkfastq_1.sh , Shell, 54 lines - 02_Chimp/
00_mkfastq_2.sh , Shell, 54 lines - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 51 linesP.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 54 linesP_2.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 53 linesP_3.sh - 02_Chimp/
01_CELLRANGER_COUNT_CHIM , Shell, 44 linesP_4.sh - 02_Chimp/
02_CELLBENDER.sh , Shell, 13 lines - 02_Chimp/
03_ANNOTATE_CHIMP_CAUDAT , R, 396 linesE.R - 02_Chimp/
03_ANNOTATE_CHIMP_PUTAME , R, 391 linesN.R - 02_Chimp/
04_INTERNEURON_CLUSTER.R , R, 292 lines - 02_Chimp/
05_NEURON_CLUSTER.R , R, 636 lines - 02_Ferret/
00_CONVERT_BAM_TO_10X_FA , Shell, 72 linesSTQ.sh - 02_Ferret/
01_GENERATE_COUNT_MATRIX , Shell, 48 lines.sh - 02_Ferret/
02_PREP_FOR_CELLBENDER.R , R, 118 lines - 02_Ferret/
03_CELLBENDER.sh , Shell, 43 lines - 02_Ferret/
04_ANNOTATION.R , R, 325 lines - 02_Human/
00_mkfastq.sh , Shell, 40 lines - 02_Human/
01_COUNT_CELLRANGER_1.sh , Shell, 58 lines - 02_Human/
01_COUNT_CELLRANGER_2.sh , Shell, 61 lines - 02_Human/
01_COUNT_CELLRANGER_3.sh , Shell, 59 lines - 02_Human/
01_COUNT_CELLRANGER_4.sh , Shell, 54 lines - 02_Human/
01_COUNT_CELLRANGER_5.sh , Shell, 54 lines - 02_Human/
02_CELLBENDER.sh , Shell, 19 lines - 02_Human/
03_CAUDATEPUTAMEN_CLUSTE , R, 356 linesR.R - 02_Human/
04_NEURON_CLUSTER.R , R, 481 lines - 02_Human/
05_INTERNEURON_CLUSTER.R , R, 226 lines - 02_Human/
06_WGCNA.R , R, 510 lines - 02_Macaque/
01_CELLRANGER_COUNT_1.sh , Shell, 66 lines - 02_Macaque/
02_CELLBENDER.sh , Shell, 13 lines - 02_Macaque/
03_ANNOTATION_MACAQUE_CA , R, 525 linesUDATE.R - 02_Macaque/
03_ANNOTATION_MACAQUE_PU , R, 306 linesTAMEN.R - 02_Macaque/
04_INTEGRATION_CAUDATE.R , R, 212 lines - 02_Macaque/
04_INTEGRATION_PUTAMEN.R , R, 219 lines - 02_Macaque/
05_CAUDATE_PUTAMEN_INTEG , R, 408 linesRATE_ANNOTATE.R - 02_Macaque/
He_et_al_Macaque/ , Shell, 74 lines00_fasterq_dump.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_1.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_2.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 51 lines01_CELLRANGER_COUNT_3.sh - 02_Macaque/
He_et_al_Macaque/ , Shell, 13 lines02_CELLBENDER.sh - 02_Macaque/
He_et_al_Macaque/ , R, 361 lines03_ANNOTATION_CAUDATE.R - 02_Macaque/
He_et_al_Macaque/ , R, 445 lines03_ANNOTATION_PUTAMEN.R - 02_Marmoset/
03_INTEGRATE_ANNOTATE_DA , R, 538 linesTASETS.R - 02_Marmoset/
Krienen_et_al/ , Shell, 34 lines00_DOWNLOAD_FASTQ_FROM_U RL.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 81 lines01_ARRANGE_FASTQs.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 69 lines02_GENERATE_COUNT_MATRIC ES.sh - 02_Marmoset/
Krienen_et_al/ , Shell, 11 lines03_CELLBENDER.sh - 02_Marmoset/
Krienen_et_al/ , R, 793 lines04_ANNOTATION.R - 02_Marmoset/
Lin_et_al/ , Shell, 77 lines00_DOWNLOAD_FASTQs.sh - 02_Marmoset/
Lin_et_al/ , Shell, 45 lines01_GENERATE_COUNT_MATRIX .sh - 02_Marmoset/
Lin_et_al/ , Shell, 11 lines02_CELLBENDER.sh - 02_Marmoset/
Lin_et_al/ , R, 538 lines03_ANNOTATE.R - 02_Mouse/
01_GENERATE_COUNT_MATRIX , Shell, 51 lines_NEXTFLOW.sh - 02_Mouse/
02_PREPARE_FOR_CELLBENDE , R, 49 linesR.R - 02_Mouse/
03_CELLBENDER.sh , Shell, 15 lines - 02_Mouse/
04_CAUDOPUTAMEN_ANNOTATE , R, 279 lines.R - 03_SPNs/
00_get_orthologs.sh , Shell, 90 lines - 03_SPNs/
01_INTEGRATION_AND_CLUST , R, 2,281 linesERING.R - 03_SPNs/
02_eSPN_markers.R , R, 444 lines - 03_SPNs/
03_eSPN_D1D2_hybrid_corr , R, 180 lines, 2 matcheselation.R - 03_SPNs/
04_SPN_seurat_obj_conver , R, 32 linest_seurat_obj_to_AnnData. R - 03_SPNs/
05_Cell_type_proportion_ , Jupyter, 460 linesanalysis_scCODA.ipynb - 04_Interneurons/
01_Integration& , R, 3,782 linesClustering.R - 04_Interneurons/
02_Interneurons_scCODA.i , Jupyter, 586 linespynb - 04_Interneurons/
03_DEG_Analysis.R , R, 467 lines - 04_Interneurons/
04_Correlation_withinBat , R, 104 linesPutamen.R - 04_Interneurons/
05_Fig5A_scCODA.py , Python, 77 lines - 05_DGE_analysis/
00_prep_pseudobulk_count , R, 1,667 liness.R - 05_DGE_analysis/
01_primates_pseudobulk_D , R, 353 linesESeq2.R - 05_DGE_analysis/
02_DEG_figures.R , R, 827 lines - GLIA_TO_NEURON.R, R, 517 lines
- QC_plots.R, R, 1,047 lines
- python_functions.py, Python, 394 lines
- scCODA_Glia_to_Neuron.py
, Python, 86 lines - LICENSE, License, 21 lines
- README.md, Text, 15 lines
Code availability
All the scripts used in the study were deposited to https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 168 scripts, each with its path and the digest of its content;
- 23 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE293075, at NCBI GEO; found in “Data availability”
Data Availability Statement
The human, chimpanzee, and pale spear-nosed bat dorsal striatum snRNA-seq data generated in this study have been deposited in the GEO database under accession code GSE293075 (https://
All the scripts used in the study were deposited to https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 28 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 12 authors, 2 keywords, 13 MeSH terms, 9 funders, 99 references.
Cite
This paper
Buyukkahraman, G., Caglayan, E., Hörpel, S. G., Zhang, Y., van Tussenbroek, I. A., Orozco, C. G., Oh, E., Hopkins, W. D., Sherwood, C. C., Roberts, T. F., Vernes, S. C., & Konopka, G. (2026). Comparative analysis of the cellular landscape in mammalian striatum. Nature communications, 17(1), 6793. https://
BibTeX
@article{buyukkahraman20
author = {Buyukkahraman, Gozde and Caglayan, Emre and Hörpel, Stephen G and Zhang, Yaqiang and van Tussenbroek, Ine A and Orozco, Carlos G and Oh, Emily and Hopkins, William D and Sherwood, Chet C and Roberts, Todd F and Vernes, Sonja C and Konopka, Genevieve},
title = {{Comparative analysis of the cellular landscape in mammalian striatum}},
journal = {Nature communications},
year = {2026},
month = may,
volume = {17},
number = {1},
pages = {6793},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42185272},
pmcid = {PMC13385833}
}
RIS
TY - JOUR
AU - Buyukkahraman, Gozde
AU - Caglayan, Emre
AU - Hörpel, Stephen G
AU - Zhang, Yaqiang
AU - van Tussenbroek, Ine A
AU - Orozco, Carlos G
AU - Oh, Emily
AU - Hopkins, William D
AU - Sherwood, Chet C
AU - Roberts, Todd F
AU - Vernes, Sonja C
AU - Konopka, Genevieve
TI - Comparative analysis of the cellular landscape in mammalian striatum
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 6793
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Comparative analysis of the cellular landscape in mammalian striatum",
"container-title": "Nature communications",
"author": [
{
"family": "Buyukkahraman",
"given": "Gozde"
},
{
"family": "Caglayan",
"given": "Emre"
},
{
"family": "Hörpel",
"given": "Stephen G"
},
{
"family": "Zhang",
"given": "Yaqiang"
},
{
"family": "van Tussenbroek",
"given": "Ine A"
},
{
"family": "Orozco",
"given": "Carlos G"
},
{
"family": "Oh",
"given": "Emily"
},
{
"family": "Hopkins",
"given": "William D"
},
{
"family": "Sherwood",
"given": "Chet C"
},
{
"family": "Roberts",
"given": "Todd F"
},
{
"family": "Vernes",
"given": "Sonja C"
},
{
"family": "Konopka",
"given": "Genevieve"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "6793",
"DOI": "10.1038/
"PMID": "42185272",
"PMCID": "PMC13385833",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
25
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: rpy2, WGCNA, Harmony, 19 other tools, genetics / omics, cellular / molecular, 2 references
- [2] doi:10.1038/s41586-026-10214-2 [code]
- Multidimensional profiling of heterogeneity in supratentorial ependymomas.Journal: NatureIn common: Cell Ranger, Harmony, SingleCellExperiment, 18 other tools, genetics / omics, mouse
- [3] doi:10.1038/s41593-026-02300-5 [code]
- Integrated single-cell and spatial transcriptomic profiling in ALS uncovers peripheral-to-central immune infiltration and reprogramming.Journal: Nature neuroscienceIn common: Cell Ranger, Harmony, SAMtools, 17 other tools, genetics / omics, cellular / molecular
- [4] doi:10.1002/imt2.70163 [code]
- Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.Journal: iMetaIn common: Harmony, SingleCellExperiment, edgeR, 17 other tools, genetics / omics, mouse, cellular / molecular
- [5] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: rpy2, SAMtools, edgeR, 16 other tools, genetics / omics, mouse, cellular / molecular
- [6] doi:10.1038/s41398-026-04200-5 [code]
- Postmortem brain single-nucleus and bulk gene expression analyses identify shared and distinct abnormalities in bipolar disorder and major depressive disorder.Journal: Translational psychiatryIn common: WGCNA, Harmony, SingleCellExperiment, 14 other tools, genetics / omics, cellular / molecular, 2 references
- [7] doi:10.1016/j.cpblue.2026.100007 [code]
- An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.Journal: Cell press blueIn common: rpy2, SingleCellExperiment, SAMtools, 16 other tools
- [8] doi:10.1016/j.celrep.2026.117500 [code]
- Spatio-molecular gene expression reflects dorsal anterior cingulate cortex structure and function in the human brain.Journal: Cell reportsIn common: Cell Ranger, Harmony, SingleCellExperiment, 13 other tools, genetics / omics, cellular / molecular, 1 reference
- [9] doi:10.7554/elife.93640 [code]
- Sibling chimerism among microglia in marmosets.Journal: eLifeIn common: Harmony, SingleCellExperiment, anndata, 12 other tools, non-human primate, genetics / omics, cellular / molecular, 3 references
- [10] doi:10.1016/j.cell.2026.05.026 [code]
- The critical role of the endogenous immune compartment after CAR T cell therapy in recurrent GBM.Journal: CellIn common: Harmony, SingleCellExperiment, edgeR, 13 other tools, genetics / omics, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 168 scripts, and 23 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:2647197ef334a38b…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
