Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification.
The 23 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Methods › Data integration and annotation (module 3) ↔ R_code/Module_3_part_1.R, lines 135–214 · score 0.98 · mitochondrial gene expression, IntegrateData, anchor.features, k.filter, k.score, FindIntegrationAnchors
- [2] § Methods › Sample processing and filtering of Sst+ cells (module 1) ↔ R_code/Module_3_part_1.R, lines 135–214 · score 0.96 · mitochondrial gene expression, RunUMAP, principal component, FindClusters, FindNeighbors, FindVariableFeatures
- [3] § Methods › Sample processing and filtering of Sst+ cells (module 1) ↔ R_code/Module_1.R, lines 42–81 · score 0.95 · Cell cycle phase, CellCycleScoring, mitochondrial gene expression, FindVariableFeatures, NormalizeData, ScaleData
- [4] § Methods › Trajectory and pseudotime analysis ↔ R_code/Trajectory_analysis.R, lines 97–141 · score 0.94 · diffusion components, root cell, diffusion map, early sample cells, gene modules, DPT
- [5] § Methods › WOT analysis ↔ R_code/WaddingtonOT_analysis.R, lines 195–254 · score 0.90 · median fate probabilities, growth_iters, optimal_transport, possible cluster identities, growth rates, command
- [6] § Methods › Cell annotation to adult counterparts ↔ R_code/MapMyCell.R, lines 1–39 · score 0.78 · reference taxonomy, hierarchical mapping, mouse brain, MapMyCells, CCN20230722, algorithm
- [7] § Methods › Saturation analysis ↔ R_code/Saturation_analysis.R, lines 174–246 · score 0.77 · MetaNeighborUS, cluster centroids, inflection point, AUROC, saturation, 95 %
- [8] § Methods › Integrability testing of datasets (module 2) › Neighborhood composition › k-NN computation ↔ R_code/Module_2.R, lines 117–168 · score 0.76 · evaluate neighborhood composition, global fraction, reference atlas, NN, graph, cells
- [9] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_mergeClusters.R, lines 80–130 · score 0.76 · cophenetic distance, DEscore, cluster pairs, logfc.threshold, dendrogram, genes
- [10] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_mergeClusters.R, lines 80–130 · score 0.75 · FindMarkers, DEscore, merging clusters, logfc.threshold, log10, cutoff
- [11] § Methods › Integrability testing of datasets (module 2) › Anchor points ↔ R_code/Module_2_number_of_anchors_threshold_definition.R, lines 1–51 · score 0.74 · acceptance criterion, FindIntegrationAnchors, Atlas v0, cca, Seurat, subset
- [12] § Methods › Random forest classification and atlas comparison ↔ R_code/Cluster_validation.R, lines 557–636 · score 0.72 · diagonal accuracy, v2 v1, classification accuracy, Cross talk, predicted, matrices
- [13] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Iterative_Clustering_pipeline_functions.R, lines 206–228 · score 0.70 · FindMarkers, DEscore, logfc.threshold, pos, log10, adj
- [14] § Methods › Cell annotation to adult counterparts ↔ Figures/MapMyCell/mapmycell_Base_atlas.rmd, lines 65–89 · score 0.68 · hierarchical mapping, reference taxonomy, mouse brain, MapMyCells, CCN20230722, algorithm
- [15] § Methods › Cortical and striatal integration ↔ R_code/Module_3_part_1.R, lines 251–316 · score 0.66 · FindTransferAnchors, TransferData, major clusters, UMAP, v0, module
- [16] § Results › Dev-SST-v2 reveals 26 clusters of cortical SST+ inhibitory neurons ↔ R_code/Cluster_validation.R, lines 557–636 · score 0.63 · diagonal accuracy, v2 v1, cross talk, bars, Martinotti, Seurat
- [17] § Methods › Cell annotation validation ↔ R_code/Cluster_validation.R, lines 268–334 · score 0.62 · min.pct, logfc.threshold, ranked, validate, cross, predict
- [18] § Methods › Integrability testing of datasets (module 2) › Neighborhood composition ↔ R_code/Module_2.R, lines 53–114 · score 0.60 · IntegrateData, RunPCA, ScaleData, PCs, anchors, SST
- [19] § Results › LRP1 and LRP2.1 are distinct cell types with divergent trajectories ↔ R_code/Trajectory_analysis.R, lines 25–95 · score 0.59 · synapse assembly, LRP2.1, axon, ANTLER, trajectory, LRP1
- [20] § Results › Generalized computational pipeline to assess integrability of different datasets ↔ R_code/Module_2.R, lines 117–168 · score 0.58 · nearest neighbors, rejection rate, BET, NN, fraction, batch
- [21] § Methods › Cortical and striatal integration ↔ R_code/Merfish_annotation.R, lines 45–94 · score 0.53 · FindTransferAnchors, TransferData, UMAP, Seurat, v2, genes
- [22] § Methods › Iterative clustering of Dev-SST-v0, Dev-SST-v1, Dev-SST-v2 ↔ R_code/Module_3_part_2_IterativeClustering.R, the whole file · a weak match · score 0.51 · Clustering stability evaluation, iteration, Jaccard, resolutions, subsamplings, v2
- [23] § Results › Dev-SST-v2 reveals 26 clusters of cortical SST+ inhibitory neurons ↔ R_code/Saturation_analysis.R, lines 142–171 · score 0.50 · robust Hausdorff distance, Inflection points, Saturation, subsampled, Dev, v2
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 317 lines · 19 KB · no license · 3 matches
- ## Installation needed
- if (!requireNamespace("pacman", quietly = TRUE)) install.packages("pacman")
- pacman::p_load(
- Seurat, dittoSeq, ggplot2, gtools, dplyr, R.utils, colorspace, tidyr,
- purrr, Matrix, patchwork, BiocManager
- )
- devtools::install_github("crazyhottommy/scclusteval")
- BiocManager::install("scDblFinder")
- BiocManager::install("MAST")
- BiocManager::install("dittoSeq")
- ########################
- # Load required packages
- ########################
- library(Seurat)
- library(dittoSeq)
- library(ggplot2)
- library(gtools)
- library(dplyr)
- library(dittoSeq)
- library(scDblFinder)
- library(R.utils)
- library(colorspace)
- library(scclusteval)
- library(tidyr)
- library(purrr)
- library(Matrix)
- library(patchwork)
- library(MAST)
- source("Seurat_Utils.R")
- colors_ditto<-dittoColors()
- names(colors_ditto)<-as.character(c(0:(length(colors_ditto)-1)))
- ###########
- # Load data
- ###########
- # Load previously Sst+ filtered cells from samples passing integrability tests
- base_atlas <- readRDS("integrated_all_Sst_filtered.rds")
- base_atlas$batch <- paste0("BaseAtlas_",base_atlas$orig.ident)
- E16_Lim4 <- readRDS("lim4_E16_Sst_filtered.rds")
- E16_Lim4$batch <- "E16_Lim4"
- P1_Lim3 <- readRDS('lim3_P1_Sst_filtered.rds')
- P1_Lim3$batch <- "P1_Lim3"
- P1_Lim5 <- readRDS("lim5_P1_Sst_filtered.rds")
- P1_Lim5$batch <- "P1_Lim5"
- test_dataset_filtered_sst<-readRDS("P5_WT1_WT23_Sst_filtered.rds")
- P5_WT1_Lim1 <- subset(test_dataset_filtered_sst, subset= orig.ident %in% "lim_P5_sorted_sept23_WT1")
- P5_WT1_Lim1$batch <- "P5_WT1_Lim1"
- P5_WT23_Lim2 <- subset(test_dataset_filtered_sst, subset= orig.ident %in% "lim_P5_sorted_sept23_WT23")
- P5_WT23_Lim2$batch <- "P5_WT23_Lim2"
- E18_Lippi<-readRDS("Lippi_lab_e18.5_Sst_filtered.rds")
- E18_Lippi$batch <- "E18_Lippi"
- P1_EMI014_Lim<-readRDS("lim_P1_EMI014_Sst_filtered.rds")
- P1_EMI014_Lim$batch <- "P1_EMI014_Lim"
- Wu_data <- readRDS("P2_P7_GSE272706_Sst_filtered.RDS")
- list_Wu <-SplitObject(Wu_data, split.by = "orig.ident")
- Wu_p2 <- list_Wu[["GSM8409508_P2MGE_Fezf2HET_BaxcHET"]]
- Wu_p2$batch <- "P2_Wu"
- E16_EMI018 <-readRDS("lim_E16.5_EMI018_Sst_filtered.rds")
- E16_EMI018$batch <- 'E16_EMI018'
- # Set the default assay to RNA for analysis
- DefaultAssay(base_atlas) <- "RNA"
- DefaultAssay(E16_Lim4) <- "RNA"
- DefaultAssay(P1_Lim3) <- "RNA"
- DefaultAssay(P1_Lim5) <- "RNA"
- DefaultAssay(P5_WT1_Lim1) <- "RNA"
- DefaultAssay(P5_WT23_Lim2) <- "RNA"
- DefaultAssay(E18_Lippi) <- "RNA"
- DefaultAssay(P1_EMI014_Lim) <- "RNA"
- DefaultAssay(Wu_p2)<-"RNA"
- DefaultAssay(E16_EMI018)<-"RNA"
- # Split the base atlas into separate objects based on the sampleID
- integrated_all_list <- SplitObject(base_atlas, split.by = "batch")
- # Add the cleaned dataset to the list of samples
- integrated_all_list$E16_Lim4 <- E16_Lim4
- integrated_all_list$P1_Lim3 <- P1_Lim3
- integrated_all_list$P1_Lim5 <- P1_Lim5
- integrated_all_list$P5_WT1_Lim1 <- P5_WT1_Lim1
- integrated_all_list$P5_WT23_Lim2 <- P5_WT23_Lim2
- integrated_all_list$E18_Lippi <- E18_Lippi
- integrated_all_list$P1_EMI014_Lim <- P1_EMI014_Lim
- integrated_all_list$Wu_p2 <- Wu_p2
- integrated_all_list$E16_EMI018 <- E16_EMI018
- #########################
- # Doublets identification
- #########################
- # First, we annotate cells according to Atlas-v0 clusters using label transfer.
- # The predicted clusters will be used to identify potential doublets.
- # Normalize the data, identify variable features, and scale the Atlas-v0 dataset.
- base_atlas <- NormalizeData(base_atlas)
- base_atlas <- FindVariableFeatures(base_atlas)
- base_atlas <- ScaleData(base_atlas, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- base_atlas <- RunPCA(base_atlas,npcs = 50)
- # Initialize empty vector to store cluster annotation
- minor_label <- c()
- # Perform label transfer
- for (sample in c("E16_Lim4","P1_Lim3","P1_Lim5","P5_WT1_Lim1","P5_WT23_Lim2","E18_Lippi","P1_EMI014_Lim","Wu_p2","E16_EMI018")) {
- integrated_all_list[[sample]] <- NormalizeData(integrated_all_list[[sample]])
- integrated_all_list[[sample]] <- FindVariableFeatures(integrated_all_list[[sample]])
- integrated_all_list[[sample]] <- ScaleData(integrated_all_list[[sample]], vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- integrated_all_list[[sample]] <- RunPCA(integrated_all_list[[sample]],npcs = 50)
- # Find transfer anchors between the Atlas-v0 (reference) and the samples
- transfer.anchors <- FindTransferAnchors(reference = base_atlas, query = integrated_all_list[[sample]], dims = 1:40, reference.reduction = "pca", features=intersect(rownames(base_atlas), rownames(integrated_all_list[[sample]])))
- # Transfer labels based on the reference
- predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$cluster_label_trained_with_all, dims = 1:40)
- # Add the predicted cluster annotation to the sample object
- integrated_all_list[[sample]]$minor_label_transferAnchors_BaseAtlas <- predictions[colnames(integrated_all_list[[sample]]),'predicted.id']
- # Append the predicted cluster annotation to the overall list
- minor_pred <- predictions$predicted.id
- names(minor_pred) <- rownames(predictions)
- minor_label <- c(minor_label,minor_pred)
- }
- for (sample in c("BaseAtlas_E16","BaseAtlas_lim_P5_fixed_sorted","BaseAtlas_P1","BaseAtlas_P5") ) {
- integrated_all_list[[sample]]$minor_label_transferAnchors_BaseAtlas <- integrated_all_list[[sample]]$cluster_label_trained_with_all
- }
- # Doublets Prediction Using predicted clusters
- # Initialize an empty vector to store doublet annotations
- doublets_sceDblF <- c()
- # Loop through all datasets in the integrated list and predict doublets
- for (i in 1:length(integrated_all_list)){
- sceDblF <- scDblFinder(integrated_all_list[[i]]@assays$RNA@counts,dbr =0.07, clusters=integrated_all_list[[i]]$minor_label_transferAnchors_BaseAtlas)
- # Extract doublet annotations from the results
- doublets_anno <- as.vector(sceDblF@colData$scDblFinder.class)
- names(doublets_anno) <- row.names(sceDblF@colData)
- # Append the doublet annotations to the list
- doublets_sceDblF <- c(doublets_sceDblF,doublets_anno)
- # Add the doublet classification to the Seurat object metadata
- integrated_all_list[[i]]$doublets <- doublets_anno[colnames(integrated_all_list[[i]])]
- # Subset the Seurat object to retain only the singlets
- integrated_all_list[[i]] <- subset(integrated_all_list[[i]], subset=doublets=="singlet")
- }
- ################################
- # CCA Integration and clustering
- ################################
- # Normalize data and find variable features
- for(sample in names(integrated_all_list)){
- obj <- integrated_all_list[[sample]]
- obj <- NormalizeData(obj)
- obj <- FindVariableFeatures(obj)
- integrated_all_list[[sample]] <- obj
- }
- # Find integration anchors between datasets
- i.anchors <- FindIntegrationAnchors(object.list = integrated_all_list, dims = 1:30, reduction = 'cca', scale = T, k.anchor = 5, k.filter = 100, k.score = 15, anchor.features = 3000)
- # Integrate the data using the anchors
- integrated_v2 <- IntegrateData(anchorset = i.anchors, dims = 1:30, normalization.method ='LogNormalize', k.weight=100)
- DefaultAssay(integrated_v2) <- 'integrated'
- integrated_v2 <- ScaleData(integrated_v2, verbose = FALSE, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- integrated_v2 <- RunPCA(integrated_v2, npcs = 60)
- ElbowPlot(integrated_v2, ndims = 60)
- # Determine the optimal number of principal components (PCs) for downstream analysis
- data.use.integrated <- PrepDR(object = integrated_v2, genes.use = VariableFeatures(object = integrated_v2), use.imputed = F, assay.type = "integrated")
- path_data <- getwd()
- nPCs.data.use5 <- PCA_estimate_nPC(data.use.integrated,
- whereto = paste0(path_data, "/optimal_nPCs_5_integrated.RDS"),
- k = 5, by.nPC = 5, from.nPC = 30, to.nPC = 50) #check ElbowPlot and adjust range acoordingly
- nPCs.data.use <- PCA_estimate_nPC(data.use.integrated,
- whereto = paste0(path_data, "/optimal_nPCs_integrated.RDS"),
- k = 5, by.nPC = 1, from.nPC = nPCs.data.use5 - 5, to.nPC = nPCs.data.use5 + 5)
- integrated_v2 <- RunUMAP(integrated_v2, dims=1:nPCs.data.use)
- integrated_v2 <- FindNeighbors(integrated_v2, dims=1:nPCs.data.use)
- integrated_v2 <- FindClusters(integrated_v2, resolution = c(0.4))
- ######################################
- # PERFORM QC AND REMOVE STRESSED CELLS
- ######################################
- DefaultAssay(integrated_v2) <- 'RNA'
- integrated_v2 <- ScaleData(integrated_v2)
- # Extract UMAP embeddings (2D coordinates for visualization)
- umap <- Embeddings(integrated_v2,reduction = "umap")
- # Loop through a list of stress-related genes to visualize their expression on UMAP and export in pdf
- for (gene in c("Hsp90b1","Hspa5","Mapk8")){
- # Get the gene expression data
- data <- integrated_v2@assays$RNA@data[gene,]
- # Combine the UMAP coordinates with gene expression values
- umap_gene <- cbind(umap[colnames((integrated_v2@assays$RNA@data)),1:2], data)
- # Order the data by expression values (ascending order)
- umap_gene <- umap_gene[order(umap_gene[,3], decreasing = FALSE),]
- colnames(umap_gene) = c("umap_1","umap_2","Expression")
- umap_gene <- as.data.frame(umap_gene)
- pdf(paste("GeneExpression_",gene,"_umap.pdf",sep=""))
- #print(ggplot(umap_gene, aes(umap_1, umap_2)) + geom_point(aes(colour = Expression), size=1) + scale_color_continuous_sequential(palette='Purple_Yellow') + theme(panel.background = element_rect(fill='white', colour='black')) + theme(legend.position="none"))
- print(ggplot(umap_gene, aes(umap_1, umap_2)) + geom_point(aes(colour = Expression), size=1) + scale_color_continuous_sequential(palette='Purple_Yellow') + theme(panel.background = element_rect(fill='white', colour='black'))+ ggtitle( paste0(gene)) )
- dev.off()
- }
- # Create UMAP plots based on feature count (nFeature) and mitochondrial percentage (percent.mt)
- dittoDimPlot(integrated_v2, "nFeature_RNA", reduction.use = "umap", min.color = "lightgrey", max.color = "blue")
- dittoDimPlot(integrated_v2, "percent.mt", reduction.use = "umap", min.color = "lightgrey", max.color = "blue")
- # Create UMAP plots for different clustering resolutions and export in pdf
- pdf('umap_plot_atlas_v2_clusters.pdf')
- Idents(integrated_v2) <- "integrated_snn_res.0.4"
- DimPlot(integrated_v2, reduction = "umap",raster=FALSE, label=TRUE) + ggtitle("Resolution: 0.4")
- dev.off()
- # Set the clustering identity to the specific resolution (0.4) and generate bar plots of samples per cluster
- Idents(integrated_v2)<-'integrated_snn_res.0.4'
- dittoBarPlot(integrated_v2, var = "batch", group.by = "integrated_snn_res.0.4") + labs(title = NULL)
- # Generate violin plots for expression of various genes by cluster (based on the integrated SNN resolution 0.4)
- VlnPlot(integrated_v2, "Hsp90b1", pt.size = 0, group.by = "integrated_snn_res.0.4")
- VlnPlot(integrated_v2, "Hspa5", pt.size = 0, group.by = "integrated_snn_res.0.4")
- VlnPlot(integrated_v2, "Mapk8", pt.size = 0, group.by = "integrated_snn_res.0.4")
- # Add a new feature for ribosomal RNA percentage (ribo genes)
- integrated_v2[["percent.ribo"]] <- PercentageFeatureSet(integrated_v2, pattern = "^Rp[Sl]")
- # Violin plot for ribosomal RNA percentage by cluster
- VlnPlot(integrated_v2, "percent.ribo", pt.size = 0, group.by = "integrated_snn_res.0.4")
- # Violin plot for mitochondrial RNA percentage by cluster
- VlnPlot(integrated_v2, "percent.mt", pt.size = 0, group.by = "integrated_snn_res.0.4")
- # Violin plots for total RNA count and feature count by cluster
- VlnPlot(integrated_v2, "nCount_RNA", pt.size = 0, group.by = "integrated_snn_res.0.4") + ylim(0,30000)
- VlnPlot(integrated_v2, "nFeature_RNA", pt.size = 0, group.by = "integrated_snn_res.0.4")
- # Filter cells - remove clusters with high expression of stress markers or mitochondrial genes, as well as clusters with low nFeatures or unbalanced sample composition.
- # Specify clusters you want to remove
- integrated_v2_filtered<-subset(integrated_v2, subset=integrated_snn_res.0.4 %in% c(8,14,15), invert = TRUE)
- #################################
- # CCA Integration after filtering
- #################################
- DefaultAssay(integrated_v2_filtered) <- 'RNA'
- integrated_all_list <- SplitObject(integrated_v2_filtered, split.by = "batch")
- # Loop over each sample and normalize the data, then find variable features
- for(sample in names(integrated_all_list)){
- # Normalize the data for each sample
- obj <- integrated_all_list[[sample]]
- obj <- NormalizeData(obj)
- obj <- FindVariableFeatures(obj)
- # Save the updated sample object back to the list
- integrated_all_list[[sample]] <- obj
- }
- # Find integration anchors between the different samples (using CCA for dimensionality reduction)
- i.anchors <- FindIntegrationAnchors(object.list = integrated_all_list, dims = 1:30, reduction='cca',scale=T, k.anchor=5, k.filter=100, k.score=15,anchor.features=3000)
- integrated_v2_filtered <- IntegrateData(anchorset = i.anchors, dims = 1:30, normalization.method = 'LogNormalize', k.weight=80)
- DefaultAssay(integrated_v2_filtered) <- 'integrated'
- integrated_v2_filtered <- ScaleData(integrated_v2_filtered, verbose = FALSE, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- integrated_v2_filtered <-RunPCA(integrated_v2_filtered, npcs = 60)
- ElbowPlot(integrated_v2, ndims = 60)
- # Determine the optimal number of principal components (PCs) for downstream analysis
- data.use.integrated<- PrepDR(object = integrated_v2_filtered, genes.use = VariableFeatures(object = integrated_v2_filtered), use.imputed = F, assay.type = "integrated")
- path_data <- getwd()
- nPCs.data.use5 <- PCA_estimate_nPC(data.use.integrated,
- whereto = paste0(path_data, "/optimal_nPCs_5_integrated_filtered.RDS"),
- k = 5, by.nPC = 5, from.nPC = 40, to.nPC = 60) # Check Elbow plot and set the range for optimal PC
- nPCs.data.use <- PCA_estimate_nPC(data.use.integrated,
- whereto = paste0(path_data, "/optimal_nPCs_integrated_filtered.RDS"),
- k = 2, by.nPC = 1, from.nPC = nPCs.data.use5 - 5, to.nPC = nPCs.data.use5 + 5)
- integrated_v2_filtered <- RunUMAP(integrated_v2_filtered, dims=1:nPCs.data.use)
- integrated_v2_filtered <- FindNeighbors(integrated_v2_filtered, dims=1:nPCs.data.use)
- integrated_v2_filtered <- FindClusters(integrated_v2_filtered, resolution = c(0.4,0.5,0.6,0.7))
- # Here, another round of quality checking is performed
- ##################
- # Cells annotation
- ##################
- # We annotate cells to major classes using Atlas-v0 as reference
- DefaultAssay(integrated_v2_filtered) <- 'RNA'
- integrated_all_list <- SplitObject(integrated_v2_filtered, split.by = "batch")
- base_atlas <- subset(integrated_v2_filtered, subset= batch %in% c("BaseAtlas_E16","BaseAtlas_lim_P5_fixed_sorted","BaseAtlas_P1","BaseAtlas_P5"))
- DefaultAssay(base_atlas) <- 'RNA'
- base_atlas <- NormalizeData(base_atlas)
- base_atlas <- FindVariableFeatures(base_atlas)
- base_atlas <- ScaleData(base_atlas, vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- base_atlas<-RunPCA(base_atlas,npcs = 60)
- # Initialize empty vectors to store predicted labels and scores
- minor_label <- c()
- major_label <- c()
- pred_score_minor <- c()
- pred_score_major <- c()
- # Loop through each sample, normalize, identify variable features, scale data, and run PCA
- for (sample in c("E16_Lim4","P1_Lim3","P1_Lim5","P5_WT1_Lim1","P5_WT23_Lim2","E18_Lippi","P1_EMI014_Lim","Wu_p2","E16_EMI018")) {
- # Normalize, find variable features, and scale the data for each sample
- integrated_all_list[[sample]] <- NormalizeData(integrated_all_list[[sample]])
- integrated_all_list[[sample]] <- FindVariableFeatures(integrated_all_list[[sample]])
- integrated_all_list[[sample]] <- ScaleData(integrated_all_list[[sample]], vars.to.regress = c("nFeature_RNA",'percent.mt','ccDiff'))
- integrated_all_list[[sample]] <- RunPCA(integrated_all_list[[sample]],npcs = 50)
- # Perform anchor finding for label transfer based on PCA reduction and intersected features
- transfer.anchors <- FindTransferAnchors(reference = base_atlas, query = integrated_all_list[[sample]], dims = 1:40,reference.reduction = "pca", features=intersect(rownames(base_atlas), rownames(integrated_all_list[[sample]])))
- # Transfer the predicted minor labels (clusters)
- predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$cluster_label_trained_with_all, dims = 1:40)
- minor_pred <- predictions$predicted.id
- names(minor_pred) <- rownames(predictions)
- minor_label <- c(minor_label,minor_pred)
- # Capture the prediction score for each minor label prediction
- pred_score_tmp <- predictions$prediction.score.max
- names(pred_score_tmp) <- rownames(predictions)
- pred_score_minor <- c(pred_score_minor,pred_score_tmp)
- # Transfer the predicted major labels (major subsets)
- predictions <- TransferData(anchorset = transfer.anchors, refdata = base_atlas$major_cluster_label_trained_with_all, dims = 1:40)
- major_pred <- predictions$predicted.id
- names(major_pred) <- rownames(predictions)
- major_label <- c(major_label,major_pred)
- # Capture the prediction score for each major label prediction
- pred_score_tmp <- predictions$prediction.score.max
- names(pred_score_tmp) <- rownames(predictions)
- pred_score_major <- c(pred_score_major,pred_score_tmp)
- }
- saveRDS(minor_label,"Minor_label_labelTranfer.RDS")
- saveRDS(major_label,"Major_label_labelTranfer.RDS")
- # Assign minor and major labels to the meta-data of the integrated filtered dataset
- integrated_v2_filtered$minor_label_transferAnchors_BaseAtlas <- integrated_v2_filtered$cluster_label_trained_with_all
- [email hidden][names(minor_label),'minor_label_transferAnchors_BaseAtlas'] <- minor_label
- integrated_v2_filtered$major_label_transferAnchors_BaseAtlas <- integrated_v2_filtered$major_cluster_label_trained_with_all
- [email hidden][names(major_label),'major_label_transferAnchors_BaseAtlas'] <- major_label
- saveRDS(integrated_v2_filtered, "Integrated_atlas_V2_filtered_annotated.RDS")
- # Generate UMAP plots for major labels
- ClusterCol <- c('#35C1D5', '#4A2884', '#E05F36')
- DimPlot(integrated_v2_filtered, group.by="major_label_transferAnchors_BaseAtlas", reduction = 'umap', cols=ClusterCol)
- #######################################
- # SAVE objects for clusters computation
- #######################################
- integrated_v2_filtered_LRP <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "LRP")
- saveRDS(integrated_v2_filtered_LRP,'Integrated_atlas_v2_LRP.RDS')
- integrated_v2_filtered_Martinotti <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "Martinotti")
- saveRDS(integrated_v2_filtered_Martinotti,'Integrated_atlas_v2_Martinotti.RDS')
- integrated_v2_filtered_NonMartinotti <- subset(integrated_v2_filtered, subset=major_label_transferAnchors_BaseAtlas == "Non-Martinotti")
- saveRDS(integrated_v2_filtered_NonMartinotti,'Integrated_atlas_v2_NonMartinotti.RDS')
Module_3_part_1.R at commit cca2307, no license · at the source
Overview
- VIB-KU Leuven Center for Neuroscience, Leuven, Belgium
- Department of Neurosciences, Katholieke Universiteit (KU) Leuven, Leuven, Belgium
- Genevia Technologies Oy, Tampere, Finland
- Department of Neuroscience, Scripps Research Institute, La Jolla, CA USA
- RIKEN Center for Brain Science, Saitama, Japan
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 23 matches between paragraphs and lines of code.
Limlab-VIBCBD/Dev-SST-atlas
cca2307aea99a332c5f0899c68af487142c797ac, 14 September 2026Availability: 1 check, the latest on 26 September 2026: the link answers
- 26 September 2026: the link answers
17 files
- Figures/
MapMyCell/ , R, 174 lines, 1 matchmapmycell_Base_atlas.rmd - Nextflow_pipeline/
bin/ , R, 72 linesprocessing_functions.R - R_code/
Cluster_validation.R , R, 636 lines, 3 matches - R_code/
Iterative_Clustering_pip , R, 313 lines, 1 matcheline_functions.R - R_code/
MapMyCell.R , R, 200 lines, 1 match - R_code/
Merfish_annotation.R , R, 94 lines, 1 match - R_code/
Module_1.R , R, 188 lines, 1 match - R_code/
Module_2.R , R, 169 lines, 3 matches - R_code/
Module_2_number_of_ancho , R, 144 lines, 1 matchrs_threshold_definition. R - R_code/
Module_3_part_1.R , R, 317 lines, 3 matches - R_code/
Module_3_part_2_Iterativ , R, 44 lines, 1 matcheClustering.R - R_code/
Module_3_part_2_mergeClu , R, 131 lines, 2 matchessters.R - R_code/
Saturation_analysis.R , R, 289 lines, 2 matches - R_code/
Seurat_Utils.R , R, 66 lines - R_code/
Trajectory_analysis.R , R, 262 lines, 2 matches - R_code/
WaddingtonOT_analysis.R , R, 517 lines, 1 match - README.md, Text, 16 lines
Code availability statement
The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: Limlab-VIBCBD/
Dev-SST-atlas - it says that the code is available on request
Read it in the paper: doi.org/10.1038/s41593-026-02387-w.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 16 scripts, each with its path and the digest of its content;
- 23 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- biostudies:S-BSST3005, at BioStudies; found in “Data availability”
- geo:GSE280655, at NCBI GEO; found in “Data availability”
- zenodo:20139187, at Zenodo; found in “Data availability”
Data availability statement
The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to 3 datasets: BioStudies S-BSST3005, NCBI GEO GSE280655, Zenodo 20139187
- it says that the data are available on request
Read it in the paper: doi.org/10.1038/s41593-026-02387-w.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 3, 28 September 2026
- Publisher: n/a → Nature Portfolio
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 14 authors, 5 keywords, 10 MeSH terms, 5 funders, 70 references, 7 RRIDs.
Cite
This paper
Liu, M., Restelli, F. F., Micoli, E., Barbiera, G., Moors, R., Nouboers, E., Reverendo, M., Du, J. X., Bertels, H., Konstantopoulos, D., Wierda, K., Takeoka, A., Lippi, G., & Lim, L. (2026). Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification. Nature neuroscience, 29(9), 2283-2295. https://
BibTeX
@article{liu2026developi
author = {Liu, Minhui and Restelli, Facundo Ferrero and Micoli, Elia and Barbiera, Giulia and Moors, Rani and Nouboers, Evelien and Reverendo, Malou and Du, Jessica Xinyun and Bertels, Hannah and Konstantopoulos, Dimitris and Wierda, Keimpe and Takeoka, Aya and Lippi, Giordano and Lim, Lynette},
title = {{Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification}},
journal = {Nature neuroscience},
year = {2026},
month = aug,
volume = {29},
number = {9},
pages = {2283--2295},
publisher = {Nature Portfolio},
issn = {1097-6256},
doi = {10.1038/
url = {https://
pmid = {42547805},
pmcid = {PMC13533839}
}
RIS
TY - JOUR
AU - Liu, Minhui
AU - Restelli, Facundo Ferrero
AU - Micoli, Elia
AU - Barbiera, Giulia
AU - Moors, Rani
AU - Nouboers, Evelien
AU - Reverendo, Malou
AU - Du, Jessica Xinyun
AU - Bertels, Hannah
AU - Konstantopoulos, Dimitris
AU - Wierda, Keimpe
AU - Takeoka, Aya
AU - Lippi, Giordano
AU - Lim, Lynette
TI - Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification
T2 - Nature neuroscience
J2 - Nat Neurosci
PY - 2026
DA - 2026/
VL - 29
IS - 9
SP - 2283
EP - 2295
SN - 1097-6256
PB - Nature Portfolio
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification",
"container-title": "Nature neuroscience",
"author": [
{
"family": "Liu",
"given": "Minhui"
},
{
"family": "Restelli",
"given": "Facundo Ferrero"
},
{
"family": "Micoli",
"given": "Elia"
},
{
"family": "Barbiera",
"given": "Giulia"
},
{
"family": "Moors",
"given": "Rani"
},
{
"family": "Nouboers",
"given": "Evelien"
},
{
"family": "Reverendo",
"given": "Malou"
},
{
"family": "Du",
"given": "Jessica Xinyun"
},
{
"family": "Bertels",
"given": "Hannah"
},
{
"family": "Konstantopoulos",
"given": "Dimitris"
},
{
"family": "Wierda",
"given": "Keimpe"
},
{
"family": "Takeoka",
"given": "Aya"
},
{
"family": "Lippi",
"given": "Giordano"
},
{
"family": "Lim",
"given": "Lynette"
}
],
"container-title-short":
"volume": "29",
"issue": "9",
"page": "2283-2295",
"DOI": "10.1038/
"PMID": "42547805",
"PMCID": "PMC13533839",
"ISSN": "1097-6256",
"publisher": "Nature Portfolio",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
3
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-71595-6 [code]
- A single-cell and spatial atlas of early human olfactory development.Journal: Nature communicationsIn common: mgcv, anndata, circlize, 9 other tools, genetics / omics, 3 references
- [2] doi:10.1371/journal.pbio.3003757 [code]
- Cell type-agnostic transcriptomic signatures enable uniform comparisons of neural maturation.Journal: PLoS biologyIn common: anndata, circlize, Scanpy, 5 other tools, genetics / omics, mouse, 5 references
- [3] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 2 references
- [4] doi:10.1038/s41593-026-02316-x [code]
- Single-cell multi-omic atlas and morphogen screening informs midbrain and hindbrain organoid engineering.Journal: Nature neuroscienceIn common: anndata, Scanpy, Seurat, 6 other tools, genetics / omics, mouse, 4 references
- [5] doi:10.1038/s41467-026-73305-8 [code]
- Comparative analysis of the cellular landscape in mammalian striatum.Journal: Nature communicationsIn common: anndata, Scanpy, ComplexHeatmap, 7 other tools, genetics / omics, mouse, cellular / molecular, 2 references
- [6] doi:10.1038/s41593-026-02300-5 [code]
- Integrated single-cell and spatial transcriptomic profiling in ALS uncovers peripheral-to-central immune infiltration and reprogramming.Journal: Nature neuroscienceIn common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 1 reference
- [7] doi:10.1038/s41467-026-74038-4 [code]
- Semaglutide attenuates neuroinflammation in male mice.Journal: Nature communicationsIn common: mgcv, anndata, Scanpy, 7 other tools, mouse, cellular / molecular, 2 references
- [8] doi:10.1038/s41467-026-74171-0 [code]
- Cluster replicability in single-cell and single-nucleus atlases of the mouse brain.Journal: Nature communicationsIn common: anndata, circlize, Scanpy, 6 other tools, genetics / omics, mouse, cellular / molecular, 3 references
- [9] doi:10.1186/s13073-026-01704-z [code]
- Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.Journal: Genome medicineIn common: anndata, Scanpy, ComplexHeatmap, 7 other tools, genetics / omics, mouse, cellular / molecular, 2 references
- [10] doi:10.1126/sciadv.aeg3223 [code]
- The extreme diversity of retinal amacrine cells has deep evolutionary roots.Journal: Science advancesIn common: anndata, circlize, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 16 scripts, and 23 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:b37a710abde9e7c1…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
