OSCR

Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification.

Code ↔ Paper

23 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 23 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Methods › Data integration and annotation (module 3) ↔ R_code/Module_3_part_1.R, lines 135–214 · score 0.98 · mitochondrial gene expression, IntegrateData, anchor.features, k.filter, k.score, FindIntegrationAnchors
  2. [2] § Methods › Sample processing and filtering of Sst+ cells (module 1) ↔ R_code/Module_3_part_1.R, lines 135–214 · score 0.96 · mitochondrial gene expression, RunUMAP, principal component, FindClusters, FindNeighbors, FindVariableFeatures
  3. [3] § Methods › Sample processing and filtering of Sst+ cells (module 1) ↔ R_code/Module_1.R, lines 42–81 · score 0.95 · Cell cycle phase, CellCycleScoring, mitochondrial gene expression, FindVariableFeatures, NormalizeData, ScaleData
  4. [4] § Methods › Trajectory and pseudotime analysis ↔ R_code/Trajectory_analysis.R, lines 97–141 · score 0.94 · diffusion components, root cell, diffusion map, early sample cells, gene modules, DPT
  5. [5] § Methods › WOT analysis ↔ R_code/WaddingtonOT_analysis.R, lines 195–254 · score 0.90 · median fate probabilities, growth_iters, optimal_transport, possible cluster identities, growth rates, command
  6. [6] § Methods › Cell annotation to adult counterparts ↔ R_code/MapMyCell.R, lines 1–39 · score 0.78 · reference taxonomy, hierarchical mapping, mouse brain, MapMyCells, CCN20230722, algorithm
  7. [7] § Methods › Saturation analysis ↔ R_code/Saturation_analysis.R, lines 174–246 · score 0.77 · MetaNeighborUS, cluster centroids, inflection point, AUROC, saturation, 95 %
  8. [8] § Methods › Integrability testing of datasets (module 2) › Neighborhood composition › k-NN computation ↔ R_code/Module_2.R, lines 117–168 · score 0.76 · evaluate neighborhood composition, global fraction, reference atlas, NN, graph, cells
  9. [9] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_mergeClusters.R, lines 80–130 · score 0.76 · cophenetic distance, DEscore, cluster pairs, logfc.threshold, dendrogram, genes
  10. [10] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_mergeClusters.R, lines 80–130 · score 0.75 · FindMarkers, DEscore, merging clusters, logfc.threshold, log10, cutoff
  11. [11] § Methods › Integrability testing of datasets (module 2) › Anchor points ↔ R_code/Module_2_number_of_anchors_threshold_definition.R, lines 1–51 · score 0.74 · acceptance criterion, FindIntegrationAnchors, Atlas v0, cca, Seurat, subset
  12. [12] § Methods › Random forest classification and atlas comparison ↔ R_code/Cluster_validation.R, lines 557–636 · score 0.72 · diagonal accuracy, v2 v1, classification accuracy, Cross talk, predicted, matrices
  13. [13] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Iterative_Clustering_pipeline_functions.R, lines 206–228 · score 0.70 · FindMarkers, DEscore, logfc.threshold, pos, log10, adj
  14. [14] § Methods › Cell annotation to adult counterparts ↔ Figures/MapMyCell/mapmycell_Base_atlas.rmd, lines 65–89 · score 0.68 · hierarchical mapping, reference taxonomy, mouse brain, MapMyCells, CCN20230722, algorithm
  15. [15] § Methods › Cortical and striatal integration ↔ R_code/Module_3_part_1.R, lines 251–316 · score 0.66 · FindTransferAnchors, TransferData, major clusters, UMAP, v0, module
  16. [16] § Results › Dev-SST-v2 reveals 26 clusters of cortical SST+ inhibitory neurons ↔ R_code/Cluster_validation.R, lines 557–636 · score 0.63 · diagonal accuracy, v2 v1, cross talk, bars, Martinotti, Seurat
  17. [17] § Methods › Cell annotation validation ↔ R_code/Cluster_validation.R, lines 268–334 · score 0.62 · min.pct, logfc.threshold, ranked, validate, cross, predict
  18. [18] § Methods › Integrability testing of datasets (module 2) › Neighborhood composition ↔ R_code/Module_2.R, lines 53–114 · score 0.60 · IntegrateData, RunPCA, ScaleData, PCs, anchors, SST
  19. [19] § Results › LRP1 and LRP2.1 are distinct cell types with divergent trajectories ↔ R_code/Trajectory_analysis.R, lines 25–95 · score 0.59 · synapse assembly, LRP2.1, axon, ANTLER, trajectory, LRP1
  20. [20] § Results › Generalized computational pipeline to assess integrability of different datasets ↔ R_code/Module_2.R, lines 117–168 · score 0.58 · nearest neighbors, rejection rate, BET, NN, fraction, batch
  21. [21] § Methods › Cortical and striatal integration ↔ R_code/Merfish_annotation.R, lines 45–94 · score 0.53 · FindTransferAnchors, TransferData, UMAP, Seurat, v2, genes
  22. [22] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_IterativeClustering.R, the whole file · a weak match · score 0.51 · Clustering stability evaluation, iteration, Jaccard, resolutions, subsamplings, v2
  23. [23] § Results › Dev-SST-v2 reveals 26 clusters of cortical SST+ inhibitory neurons ↔ R_code/Saturation_analysis.R, lines 142–171 · score 0.50 · robust Hausdorff distance, Inflection points, Saturation, subsampled, Dev, v2

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 317 lines · 19 KB · no license · 3 matches

  1. ## Installation needed
  2. if (!requireNamespace("pacman", quietly = TRUE)) install.packages("pacman")
  3. pacman::p_load(
  4. Seurat, dittoSeq, ggplot2, gtools, dplyr, R.utils, colorspace, tidyr,
  5. purrr, Matrix, patchwork, BiocManager
  6. )
  7. devtools::install_github("crazyhottommy/scclusteval")
  8. BiocManager::install("scDblFinder")
  9. BiocManager::install("MAST")
  10. BiocManager::install("dittoSeq")
  11. ########################
  12. # Load required packages
  13. ########################
  14. library(Seurat)
  15. library(dittoSeq)
  16. library(ggplot2)
  17. library(gtools)
  18. library(dplyr)
  19. library(dittoSeq)
  20. library(scDblFinder)
  21. library(R.utils)
  22. library(colorspace)
  23. library(scclusteval)
  24. library(tidyr)
  25. library(purrr)
  26. library(Matrix)
  27. library(patchwork)
  28. library(MAST)
  29. source("Seurat_Utils.R")
  30. colors_ditto<-dittoColors()
  31. names(colors_ditto)<-as.character(c(0:(length(colors_ditto)-1)))
  32. ###########
  33. # Load data
  34. ###########
  35. # Load previously Sst+ filtered cells from samples passing integrability tests
  36. base_atlas <- readRDS("integrated_all_Sst_filtered.rds")
  37. base_atlas$batch <- paste0("BaseAtlas_",base_atlas$orig.ident)
  38. E16_Lim4 <- readRDS("lim4_E16_Sst_filtered.rds")
  39. E16_Lim4$batch <- "E16_Lim4"
  40. P1_Lim3 <- readRDS('lim3_P1_Sst_filtered.rds')
  41. P1_Lim3$batch <- "P1_Lim3"
  42. P1_Lim5 <- readRDS("lim5_P1_Sst_filtered.rds")
  43. P1_Lim5$batch <- "P1_Lim5"
  44. test_dataset_filtered_sst<-readRDS("P5_WT1_WT23_Sst_filtered.rds")
  45. P5_WT1_Lim1 <- subset(test_dataset_filtered_sst, subset= orig.ident %in% "lim_P5_sorted_sept23_WT1")
  46. P5_WT1_Lim1$batch <- "P5_WT1_Lim1"
  47. P5_WT23_Lim2 <- subset(test_dataset_filtered_sst, subset= orig.ident %in% "lim_P5_sorted_sept23_WT23")
  48. P5_WT23_Lim2$batch <- "P5_WT23_Lim2"
  49. E18_Lippi<-readRDS("Lippi_lab_e18.5_Sst_filtered.rds")
  50. E18_Lippi$batch <- "E18_Lippi"
  51. P1_EMI014_Lim<-readRDS("lim_P1_EMI014_Sst_filtered.rds")
  52. P1_EMI014_Lim$batch <- "P1_EMI014_Lim"
  53. Wu_data <- readRDS("P2_P7_GSE272706_Sst_filtered.RDS")
  54. list_Wu <-SplitObject(Wu_data, split.by = "orig.ident")
  55. Wu_p2 <- list_Wu[["GSM8409508_P2MGE_Fezf2HET_BaxcHET"]]
  56. Wu_p2$batch <- "P2_Wu"
  57. E16_EMI018 <-readRDS("lim_E16.5_EMI018_Sst_filtered.rds")
  58. E16_EMI018$batch <- 'E16_EMI018'
  59. # Set the default assay to RNA for analysis
  60. DefaultAssay(base_atlas) <- "RNA"
  61. DefaultAssay(E16_Lim4) <- "RNA"
  62. DefaultAssay(P1_Lim3) <- "RNA"
  63. DefaultAssay(P1_Lim5) <- "RNA"
  64. DefaultAssay(P5_WT1_Lim1) <- "RNA"
  65. DefaultAssay(P5_WT23_Lim2) <- "RNA"
  66. DefaultAssay(E18_Lippi) <- "RNA"
  67. DefaultAssay(P1_EMI014_Lim) <- "RNA"
  68. DefaultAssay(Wu_p2)<-"RNA"
  69. DefaultAssay(E16_EMI018)<-"RNA"
  70. # Split the base atlas into separate objects based on the sampleID
  71. integrated_all_list <- SplitObject(base_atlas, split.by = "batch")
  72. # Add the cleaned dataset to the list of samples
  73. integrated_all_list$E16_Lim4 <- E16_Lim4
  74. integrated_all_list$P1_Lim3 <- P1_Lim3
  75. integrated_all_list$P1_Lim5 <- P1_Lim5
  76. integrated_all_list$P5_WT1_Lim1 <- P5_WT1_Lim1
  77. integrated_all_list$P5_WT23_Lim2 <- P5_WT23_Lim2
  78. integrated_all_list$E18_Lippi <- E18_Lippi
  79. integrated_all_list$P1_EMI014_Lim <- P1_EMI014_Lim
  80. integrated_all_list$Wu_p2 <- Wu_p2
  81. integrated_all_list$E16_EMI018 <- E16_EMI018
  82. #########################
  83. # Doublets identification
  84. #########################
  85. # First, we annotate cells according to Atlas-v0 clusters using label transfer.
  86. # The predicted clusters will be used to identify potential doublets.
  87. # Normalize the data, identify variable features, and scale the Atlas-v0 dataset.
  88. base_atlas <- NormalizeData(base_atlas)
  89. base_atlas <- FindVariableFeatures(base_atlas)
  90. base_atlas <- ScaleData(base_atlas, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  91. base_atlas <- RunPCA(base_atlas,npcs = 50)
  92. # Initialize empty vector to store cluster annotation
  93. minor_label <- c()
  94. # Perform label transfer
  95. for (sample in c("E16_Lim4","P1_Lim3","P1_Lim5","P5_WT1_Lim1","P5_WT23_Lim2","E18_Lippi","P1_EMI014_Lim","Wu_p2","E16_EMI018")) {
  96. integrated_all_list[[sample]] <- NormalizeData(integrated_all_list[[sample]])
  97. integrated_all_list[[sample]] <- FindVariableFeatures(integrated_all_list[[sample]])
  98. integrated_all_list[[sample]] <- ScaleData(integrated_all_list[[sample]], vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  99. integrated_all_list[[sample]] <- RunPCA(integrated_all_list[[sample]],npcs = 50)
  100. # Find transfer anchors between the Atlas-v0 (reference) and the samples
  101. transfer.anchors <- FindTransferAnchors(reference = base_atlas, query = integrated_all_list[[sample]], dims = 1:40, reference.reduction = "pca", features=intersect(rownames(base_atlas), rownames(integrated_all_list[[sample]])))
  102. # Transfer labels based on the reference
  103. predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$cluster_label_trained_with_all, dims = 1:40)
  104. # Add the predicted cluster annotation to the sample object
  105. integrated_all_list[[sample]]$minor_label_transferAnchors_BaseAtlas <- predictions[colnames(integrated_all_list[[sample]]),'predicted.id']
  106. # Append the predicted cluster annotation to the overall list
  107. minor_pred <- predictions$predicted.id
  108. names(minor_pred) <- rownames(predictions)
  109. minor_label <- c(minor_label,minor_pred)
  110. }
  111. for (sample in c("BaseAtlas_E16","BaseAtlas_lim_P5_fixed_sorted","BaseAtlas_P1","BaseAtlas_P5") ) {
  112. integrated_all_list[[sample]]$minor_label_transferAnchors_BaseAtlas <- integrated_all_list[[sample]]$cluster_label_trained_with_all
  113. }
  114. # Doublets Prediction Using predicted clusters
  115. # Initialize an empty vector to store doublet annotations
  116. doublets_sceDblF <- c()
  117. # Loop through all datasets in the integrated list and predict doublets
  118. for (i in 1:length(integrated_all_list)){
  119. sceDblF <- scDblFinder(integrated_all_list[[i]]@assays$RNA@counts,dbr =0.07, clusters=integrated_all_list[[i]]$minor_label_transferAnchors_BaseAtlas)
  120. # Extract doublet annotations from the results
  121. doublets_anno <- as.vector(sceDblF@colData$scDblFinder.class)
  122. names(doublets_anno) <- row.names(sceDblF@colData)
  123. # Append the doublet annotations to the list
  124. doublets_sceDblF <- c(doublets_sceDblF,doublets_anno)
  125. # Add the doublet classification to the Seurat object metadata
  126. integrated_all_list[[i]]$doublets <- doublets_anno[colnames(integrated_all_list[[i]])]
  127. # Subset the Seurat object to retain only the singlets
  128. integrated_all_list[[i]] <- subset(integrated_all_list[[i]], subset=doublets=="singlet")
  129. }
  130. ################################
  131. # CCA Integration and clustering
  132. ################################
  133. # Normalize data and find variable features
  134. for(sample in names(integrated_all_list)){
  135. obj <- integrated_all_list[[sample]]
  136. obj <- NormalizeData(obj)
  137. obj <- FindVariableFeatures(obj)
  138. integrated_all_list[[sample]] <- obj
  139. }
  140. # Find integration anchors between datasets
  141. i.anchors <- FindIntegrationAnchors(object.list = integrated_all_list, dims = 1:30, reduction = 'cca', scale = T, k.anchor = 5, k.filter = 100, k.score = 15, anchor.features = 3000)
  142. # Integrate the data using the anchors
  143. integrated_v2 <- IntegrateData(anchorset = i.anchors, dims = 1:30, normalization.method ='LogNormalize', k.weight=100)
  144. DefaultAssay(integrated_v2) <- 'integrated'
  145. integrated_v2 <- ScaleData(integrated_v2, verbose = FALSE, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  146. integrated_v2 <- RunPCA(integrated_v2, npcs = 60)
  147. ElbowPlot(integrated_v2, ndims = 60)
  148. # Determine the optimal number of principal components (PCs) for downstream analysis
  149. data.use.integrated <- PrepDR(object = integrated_v2, genes.use = VariableFeatures(object = integrated_v2), use.imputed = F, assay.type = "integrated")
  150. path_data <- getwd()
  151. nPCs.data.use5 <- PCA_estimate_nPC(data.use.integrated,
  152. whereto = paste0(path_data, "/optimal_nPCs_5_integrated.RDS"),
  153. k = 5, by.nPC = 5, from.nPC = 30, to.nPC = 50) #check ElbowPlot and adjust range acoordingly
  154. nPCs.data.use <- PCA_estimate_nPC(data.use.integrated,
  155. whereto = paste0(path_data, "/optimal_nPCs_integrated.RDS"),
  156. k = 5, by.nPC = 1, from.nPC = nPCs.data.use5 - 5, to.nPC = nPCs.data.use5 + 5)
  157. integrated_v2 <- RunUMAP(integrated_v2, dims=1:nPCs.data.use)
  158. integrated_v2 <- FindNeighbors(integrated_v2, dims=1:nPCs.data.use)
  159. integrated_v2 <- FindClusters(integrated_v2, resolution = c(0.4))
  160. ######################################
  161. # PERFORM QC AND REMOVE STRESSED CELLS
  162. ######################################
  163. DefaultAssay(integrated_v2) <- 'RNA'
  164. integrated_v2 <- ScaleData(integrated_v2)
  165. # Extract UMAP embeddings (2D coordinates for visualization)
  166. umap <- Embeddings(integrated_v2,reduction = "umap")
  167. # Loop through a list of stress-related genes to visualize their expression on UMAP and export in pdf
  168. for (gene in c("Hsp90b1","Hspa5","Mapk8")){
  169. # Get the gene expression data
  170. data <- integrated_v2@assays$RNA@data[gene,]
  171. # Combine the UMAP coordinates with gene expression values
  172. umap_gene <- cbind(umap[colnames((integrated_v2@assays$RNA@data)),1:2], data)
  173. # Order the data by expression values (ascending order)
  174. umap_gene <- umap_gene[order(umap_gene[,3], decreasing = FALSE),]
  175. colnames(umap_gene) = c("umap_1","umap_2","Expression")
  176. umap_gene <- as.data.frame(umap_gene)
  177. pdf(paste("GeneExpression_",gene,"_umap.pdf",sep=""))
  178. #print(ggplot(umap_gene, aes(umap_1, umap_2)) + geom_point(aes(colour = Expression), size=1) + scale_color_continuous_sequential(palette='Purple_Yellow') + theme(panel.background = element_rect(fill='white', colour='black')) + theme(legend.position="none"))
  179. print(ggplot(umap_gene, aes(umap_1, umap_2)) + geom_point(aes(colour = Expression), size=1) + scale_color_continuous_sequential(palette='Purple_Yellow') + theme(panel.background = element_rect(fill='white', colour='black'))+ ggtitle( paste0(gene)) )
  180. dev.off()
  181. }
  182. # Create UMAP plots based on feature count (nFeature) and mitochondrial percentage (percent.mt)
  183. dittoDimPlot(integrated_v2, "nFeature_RNA", reduction.use = "umap", min.color = "lightgrey", max.color = "blue")
  184. dittoDimPlot(integrated_v2, "percent.mt", reduction.use = "umap", min.color = "lightgrey", max.color = "blue")
  185. # Create UMAP plots for different clustering resolutions and export in pdf
  186. pdf('umap_plot_atlas_v2_clusters.pdf')
  187. Idents(integrated_v2) <- "integrated_snn_res.0.4"
  188. DimPlot(integrated_v2, reduction = "umap",raster=FALSE, label=TRUE) + ggtitle("Resolution: 0.4")
  189. dev.off()
  190. # Set the clustering identity to the specific resolution (0.4) and generate bar plots of samples per cluster
  191. Idents(integrated_v2)<-'integrated_snn_res.0.4'
  192. dittoBarPlot(integrated_v2, var = "batch", group.by = "integrated_snn_res.0.4") + labs(title = NULL)
  193. # Generate violin plots for expression of various genes by cluster (based on the integrated SNN resolution 0.4)
  194. VlnPlot(integrated_v2, "Hsp90b1", pt.size = 0, group.by = "integrated_snn_res.0.4")
  195. VlnPlot(integrated_v2, "Hspa5", pt.size = 0, group.by = "integrated_snn_res.0.4")
  196. VlnPlot(integrated_v2, "Mapk8", pt.size = 0, group.by = "integrated_snn_res.0.4")
  197. # Add a new feature for ribosomal RNA percentage (ribo genes)
  198. integrated_v2[["percent.ribo"]] <- PercentageFeatureSet(integrated_v2, pattern = "^Rp[Sl]")
  199. # Violin plot for ribosomal RNA percentage by cluster
  200. VlnPlot(integrated_v2, "percent.ribo", pt.size = 0, group.by = "integrated_snn_res.0.4")
  201. # Violin plot for mitochondrial RNA percentage by cluster
  202. VlnPlot(integrated_v2, "percent.mt", pt.size = 0, group.by = "integrated_snn_res.0.4")
  203. # Violin plots for total RNA count and feature count by cluster
  204. VlnPlot(integrated_v2, "nCount_RNA", pt.size = 0, group.by = "integrated_snn_res.0.4") + ylim(0,30000)
  205. VlnPlot(integrated_v2, "nFeature_RNA", pt.size = 0, group.by = "integrated_snn_res.0.4")
  206. # Filter cells - remove clusters with high expression of stress markers or mitochondrial genes, as well as clusters with low nFeatures or unbalanced sample composition.
  207. # Specify clusters you want to remove
  208. integrated_v2_filtered<-subset(integrated_v2, subset=integrated_snn_res.0.4 %in% c(8,14,15), invert = TRUE)
  209. #################################
  210. # CCA Integration after filtering
  211. #################################
  212. DefaultAssay(integrated_v2_filtered) <- 'RNA'
  213. integrated_all_list <- SplitObject(integrated_v2_filtered, split.by = "batch")
  214. # Loop over each sample and normalize the data, then find variable features
  215. for(sample in names(integrated_all_list)){
  216. # Normalize the data for each sample
  217. obj <- integrated_all_list[[sample]]
  218. obj <- NormalizeData(obj)
  219. obj <- FindVariableFeatures(obj)
  220. # Save the updated sample object back to the list
  221. integrated_all_list[[sample]] <- obj
  222. }
  223. # Find integration anchors between the different samples (using CCA for dimensionality reduction)
  224. i.anchors <- FindIntegrationAnchors(object.list = integrated_all_list, dims = 1:30, reduction='cca',scale=T, k.anchor=5, k.filter=100, k.score=15,anchor.features=3000)
  225. integrated_v2_filtered <- IntegrateData(anchorset = i.anchors, dims = 1:30, normalization.method = 'LogNormalize', k.weight=80)
  226. DefaultAssay(integrated_v2_filtered) <- 'integrated'
  227. integrated_v2_filtered <- ScaleData(integrated_v2_filtered, verbose = FALSE, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  228. integrated_v2_filtered <-RunPCA(integrated_v2_filtered, npcs = 60)
  229. ElbowPlot(integrated_v2, ndims = 60)
  230. # Determine the optimal number of principal components (PCs) for downstream analysis
  231. data.use.integrated<- PrepDR(object = integrated_v2_filtered, genes.use = VariableFeatures(object = integrated_v2_filtered), use.imputed = F, assay.type = "integrated")
  232. path_data <- getwd()
  233. nPCs.data.use5 <- PCA_estimate_nPC(data.use.integrated,
  234. whereto = paste0(path_data, "/optimal_nPCs_5_integrated_filtered.RDS"),
  235. k = 5, by.nPC = 5, from.nPC = 40, to.nPC = 60) # Check Elbow plot and set the range for optimal PC
  236. nPCs.data.use <- PCA_estimate_nPC(data.use.integrated,
  237. whereto = paste0(path_data, "/optimal_nPCs_integrated_filtered.RDS"),
  238. k = 2, by.nPC = 1, from.nPC = nPCs.data.use5 - 5, to.nPC = nPCs.data.use5 + 5)
  239. integrated_v2_filtered <- RunUMAP(integrated_v2_filtered, dims=1:nPCs.data.use)
  240. integrated_v2_filtered <- FindNeighbors(integrated_v2_filtered, dims=1:nPCs.data.use)
  241. integrated_v2_filtered <- FindClusters(integrated_v2_filtered, resolution = c(0.4,0.5,0.6,0.7))
  242. # Here, another round of quality checking is performed
  243. ##################
  244. # Cells annotation
  245. ##################
  246. # We annotate cells to major classes using Atlas-v0 as reference
  247. DefaultAssay(integrated_v2_filtered) <- 'RNA'
  248. integrated_all_list <- SplitObject(integrated_v2_filtered, split.by = "batch")
  249. base_atlas <- subset(integrated_v2_filtered, subset= batch %in% c("BaseAtlas_E16","BaseAtlas_lim_P5_fixed_sorted","BaseAtlas_P1","BaseAtlas_P5"))
  250. DefaultAssay(base_atlas) <- 'RNA'
  251. base_atlas <- NormalizeData(base_atlas)
  252. base_atlas <- FindVariableFeatures(base_atlas)
  253. base_atlas <- ScaleData(base_atlas, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  254. base_atlas<-RunPCA(base_atlas,npcs = 60)
  255. # Initialize empty vectors to store predicted labels and scores
  256. minor_label <- c()
  257. major_label <- c()
  258. pred_score_minor <- c()
  259. pred_score_major <- c()
  260. # Loop through each sample, normalize, identify variable features, scale data, and run PCA
  261. for (sample in c("E16_Lim4","P1_Lim3","P1_Lim5","P5_WT1_Lim1","P5_WT23_Lim2","E18_Lippi","P1_EMI014_Lim","Wu_p2","E16_EMI018")) {
  262. # Normalize, find variable features, and scale the data for each sample
  263. integrated_all_list[[sample]] <- NormalizeData(integrated_all_list[[sample]])
  264. integrated_all_list[[sample]] <- FindVariableFeatures(integrated_all_list[[sample]])
  265. integrated_all_list[[sample]] <- ScaleData(integrated_all_list[[sample]], vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
  266. integrated_all_list[[sample]] <- RunPCA(integrated_all_list[[sample]],npcs = 50)
  267. # Perform anchor finding for label transfer based on PCA reduction and intersected features
  268. transfer.anchors <- FindTransferAnchors(reference = base_atlas, query = integrated_all_list[[sample]], dims = 1:40,reference.reduction = "pca", features=intersect(rownames(base_atlas), rownames(integrated_all_list[[sample]])))
  269. # Transfer the predicted minor labels (clusters)
  270. predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$cluster_label_trained_with_all, dims = 1:40)
  271. minor_pred <- predictions$predicted.id
  272. names(minor_pred) <- rownames(predictions)
  273. minor_label <- c(minor_label,minor_pred)
  274. # Capture the prediction score for each minor label prediction
  275. pred_score_tmp <- predictions$prediction.score.max
  276. names(pred_score_tmp) <- rownames(predictions)
  277. pred_score_minor <- c(pred_score_minor,pred_score_tmp)
  278. # Transfer the predicted major labels (major subsets)
  279. predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$major_cluster_label_trained_with_all, dims = 1:40)
  280. major_pred <- predictions$predicted.id
  281. names(major_pred) <- rownames(predictions)
  282. major_label <- c(major_label,major_pred)
  283. # Capture the prediction score for each major label prediction
  284. pred_score_tmp <- predictions$prediction.score.max
  285. names(pred_score_tmp) <- rownames(predictions)
  286. pred_score_major <- c(pred_score_major,pred_score_tmp)
  287. }
  288. saveRDS(minor_label,"Minor_label_labelTranfer.RDS")
  289. saveRDS(major_label,"Major_label_labelTranfer.RDS")
  290. # Assign minor and major labels to the meta-data of the integrated filtered dataset
  291. integrated_v2_filtered$minor_label_transferAnchors_BaseAtlas <- integrated_v2_filtered$cluster_label_trained_with_all
  292. [email hidden][names(minor_label),'minor_label_transferAnchors_BaseAtlas'] <- minor_label
  293. integrated_v2_filtered$major_label_transferAnchors_BaseAtlas <- integrated_v2_filtered$major_cluster_label_trained_with_all
  294. [email hidden][names(major_label),'major_label_transferAnchors_BaseAtlas'] <- major_label
  295. saveRDS(integrated_v2_filtered, "Integrated_atlas_V2_filtered_annotated.RDS")
  296. # Generate UMAP plots for major labels
  297. ClusterCol <- c('#35C1D5', '#4A2884', '#E05F36')
  298. DimPlot(integrated_v2_filtered, group.by="major_label_transferAnchors_BaseAtlas", reduction = 'umap', cols=ClusterCol)
  299. #######################################
  300. # SAVE objects for clusters computation
  301. #######################################
  302. integrated_v2_filtered_LRP <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "LRP")
  303. saveRDS(integrated_v2_filtered_LRP,'Integrated_atlas_v2_LRP.RDS')
  304. integrated_v2_filtered_Martinotti <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "Martinotti")
  305. saveRDS(integrated_v2_filtered_Martinotti,'Integrated_atlas_v2_Martinotti.RDS')
  306. integrated_v2_filtered_NonMartinotti <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "Non-Martinotti")
  307. saveRDS(integrated_v2_filtered_NonMartinotti,'Integrated_atlas_v2_NonMartinotti.RDS')

Module_3_part_1.R at commit cca2307, no license · at the source

Overview

Authors: Minhui Liu1,2, Facundo Ferrero Restelli1,2, Elia Micoli1,2, Giulia Barbiera3, Rani Moors1,2, Evelien Nouboers1,2, Malou Reverendo1, Jessica Xinyun Du4, Hannah Bertels2, Dimitris Konstantopoulos3, Keimpe Wierda1, Aya Takeoka5, Giordano Lippi4, Lynette Lim1,2
  1. VIB-KU Leuven Center for Neuroscience, Leuven, Belgium
  2. Department of Neurosciences, Katholieke Universiteit (KU) Leuven, Leuven, Belgium
  3. Genevia Technologies Oy, Tampere, Finland
  4. Department of Neuroscience, Scripps Research Institute, La Jolla, CA USA
  5. RIKEN Center for Brain Science, Saitama, Japan
Journal: Nature neuroscience, volume 29, issue 9, pages 2283-2295
Dates: received 11 February 2025; accepted 24 June 2026; published online 3 August 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41593-026-02387-w · PMID 42547805 · PMCID PMC13533839 · OpenAlex W7172317050
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), mouse (organism), cellular / molecular (subfield)
Methods: Smoothing, state filtering, decompositions, Machine learning, Evoked potentials, Connectivity, Statistics, fMRI & imaging
Keywords: Cell type diversity, Cell fate and cell lineage, Differentiation, Neuronal development, Molecular neuroscience
MeSH: Cerebral Cortex*, Neural Inhibition*, Neurons*, Transcriptome*, Animals, Mice, Neurodevelopment, Single-Cell Analysis, Single-Cell Gene Expression Analysis, Somatostatin (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Vlaams Instituut voor Biotechnologie (Flanders Institute for Biotechnology) (Lim lab); NINDS NIH HHS (F31 NS118982); NIMH NIH HHS (RF1 MH126719); U.S. Department of Health & Human Services | NIH | Center for Scientific Review (NIH Center for Scientific Review) (F31NS118982, RF1MH126719); International Foundation for Research in Paraplegia (Internationale Stiftung für Forschung in Paraplegie) (P188)
Citations: not cited yet (Europe PMC); 73 references in the paper
Research resources: rabbit anti-Dach1 RRID:AB_2230330, rabbit antinitric oxide synthase 1—NOS1 RRID:AB_572256, WT C57BL/6J RRID:IMSR_JAX:000664, RC:FLTG RRID:IMSR_JAX:026932, Nkx2.1flp RRID:IMSR_JAX:028577, RCE RRID:MMRRC_032037-JAX, RRID:SCR_024672

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 23 matches between paragraphs and lines of code.

Limlab-VIBCBD/Dev-SST-atlas

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: cca2307aea99a332c5f0899c68af487142c797ac, 14 September 2026
Languages: R (16)
Size: 54 files, 16 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (Nextflow_pipeline/docker/Dockerfile), 1 notebook
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: Seurat (15 files), tidyverse (15 files), ggplot2 (6 files), circlize (2 files), ComplexHeatmap (2 files), patchwork (2 files), reshape2 (2 files), anndata (1 file), cowplot (1 file), mgcv (1 file), pandas (1 file), randomForest (1 file), Scanpy (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
17 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41593-026-02387-w.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 16 scripts, each with its path and the digest of its content;
  • 23 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41593-026-02387-w.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 14 authors, 5 keywords, 10 MeSH terms, 5 funders, 70 references, 7 RRIDs.

Cite

This paper

Liu, M., Restelli, F. F., Micoli, E., Barbiera, G., Moors, R., Nouboers, E., Reverendo, M., Du, J. X., Bertels, H., Konstantopoulos, D., Wierda, K., Takeoka, A., Lippi, G., & Lim, L. (2026). Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification. Nature neuroscience, 29(9), 2283-2295. https://doi.org/10.1038/s41593-026-02387-w

BibTeX

@article{liu2026developing,
author = {Liu, Minhui and Restelli, Facundo Ferrero and Micoli, Elia and Barbiera, Giulia and Moors, Rani and Nouboers, Evelien and Reverendo, Malou and Du, Jessica Xinyun and Bertels, Hannah and Konstantopoulos, Dimitris and Wierda, Keimpe and Takeoka, Aya and Lippi, Giordano and Lim, Lynette},
title = {{Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification}},
journal = {Nature neuroscience},
year = {2026},
month = aug,
volume = {29},
number = {9},
pages = {2283--2295},
publisher = {Nature Portfolio},
issn = {1097-6256},
doi = {10.1038/s41593-026-02387-w},
url = {https://doi.org/10.1038/s41593-026-02387-w},
pmid = {42547805},
pmcid = {PMC13533839}
}

RIS

TY - JOUR
AU - Liu, Minhui
AU - Restelli, Facundo Ferrero
AU - Micoli, Elia
AU - Barbiera, Giulia
AU - Moors, Rani
AU - Nouboers, Evelien
AU - Reverendo, Malou
AU - Du, Jessica Xinyun
AU - Bertels, Hannah
AU - Konstantopoulos, Dimitris
AU - Wierda, Keimpe
AU - Takeoka, Aya
AU - Lippi, Giordano
AU - Lim, Lynette
TI - Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification
T2 - Nature neuroscience
J2 - Nat Neurosci
PY - 2026
DA - 2026/08/03
VL - 29
IS - 9
SP - 2283
EP - 2295
SN - 1097-6256
PB - Nature Portfolio
DO - 10.1038/s41593-026-02387-w
UR - https://doi.org/10.1038/s41593-026-02387-w
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41593-026-02387-w",
"type": "article-journal",
"title": "Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification",
"container-title": "Nature neuroscience",
"author": [
{
"family": "Liu",
"given": "Minhui"
},
{
"family": "Restelli",
"given": "Facundo Ferrero"
},
{
"family": "Micoli",
"given": "Elia"
},
{
"family": "Barbiera",
"given": "Giulia"
},
{
"family": "Moors",
"given": "Rani"
},
{
"family": "Nouboers",
"given": "Evelien"
},
{
"family": "Reverendo",
"given": "Malou"
},
{
"family": "Du",
"given": "Jessica Xinyun"
},
{
"family": "Bertels",
"given": "Hannah"
},
{
"family": "Konstantopoulos",
"given": "Dimitris"
},
{
"family": "Wierda",
"given": "Keimpe"
},
{
"family": "Takeoka",
"given": "Aya"
},
{
"family": "Lippi",
"given": "Giordano"
},
{
"family": "Lim",
"given": "Lynette"
}
],
"container-title-short": "Nat Neurosci",
"volume": "29",
"issue": "9",
"page": "2283-2295",
"DOI": "10.1038/s41593-026-02387-w",
"PMID": "42547805",
"PMCID": "PMC13533839",
"ISSN": "1097-6256",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41593-026-02387-w",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
3
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-71595-6 [code]
A single-cell and spatial atlas of early human olfactory development.
Journal: Nature communications
In common: mgcv, anndata, circlize, 9 other tools, genetics / omics, 3 references
[2] doi:10.1371/journal.pbio.3003757 [code]
Cell type-agnostic transcriptomic signatures enable uniform comparisons of neural maturation.
Journal: PLoS biology
In common: anndata, circlize, Scanpy, 5 other tools, genetics / omics, mouse, 5 references
[3] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 2 references
[4] doi:10.1038/s41593-026-02316-x [code]
Single-cell multi-omic atlas and morphogen screening informs midbrain and hindbrain organoid engineering.
Journal: Nature neuroscience
In common: anndata, Scanpy, Seurat, 6 other tools, genetics / omics, mouse, 4 references
[5] doi:10.1038/s41467-026-73305-8 [code]
Comparative analysis of the cellular landscape in mammalian striatum.
Journal: Nature communications
In common: anndata, Scanpy, ComplexHeatmap, 7 other tools, genetics / omics, mouse, cellular / molecular, 2 references
[6] doi:10.1038/s41593-026-02300-5 [code]
Integrated single-cell and spatial transcriptomic profiling in ALS uncovers peripheral-to-central immune infiltration and reprogramming.
Journal: Nature neuroscience
In common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 1 reference
[7] doi:10.1038/s41467-026-74038-4 [code]
Semaglutide attenuates neuroinflammation in male mice.
Journal: Nature communications
In common: mgcv, anndata, Scanpy, 7 other tools, mouse, cellular / molecular, 2 references
[8] doi:10.1038/s41467-026-74171-0 [code]
Cluster replicability in single-cell and single-nucleus atlases of the mouse brain.
Journal: Nature communications
In common: anndata, circlize, Scanpy, 6 other tools, genetics / omics, mouse, cellular / molecular, 3 references
[9] doi:10.1186/s13073-026-01704-z [code]
Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.
Journal: Genome medicine
In common: anndata, Scanpy, ComplexHeatmap, 7 other tools, genetics / omics, mouse, cellular / molecular, 2 references
[10] doi:10.1126/sciadv.aeg3223 [code]
The extreme diversity of retinal amacrine cells has deep evolutionary roots.
Journal: Science advances
In common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.