OSCR

A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Technical Validation › Anchor-based label transfer and validation ↔ Code.R, lines 361–448 · score 0.89 · FindTransferAnchors, MapQuery, reference.reduction, wt_pink1ko, LogNormalize, pcaproject
  2. [2] § Methods › Single-nucleus RNA-seq data processing and analysis ↔ Code.R, lines 1–63 · score 0.84 · FindNeighbors, IntegrateLayers, variable features, Seurat, QC, resolution
  3. [3] § Background & Summary ↔ Code.R, lines 201–258 · score 0.67 · intestinal epithelial cells, goblet cells, stem cells, colonocytes, clustering, genes
  4. [4] § Technical Validation › Anchor-based label transfer and validation ↔ Code.R, lines 260–310 · score 0.59 · Col1a1, Acta2, Lyz2, Pdgfra, Ptprc, fibroblasts
  5. [5] § Technical Validation › Anchor-based label transfer and validation ↔ Code.R, lines 739–816 · score 0.57 · decrease Gini, decrease accuracy, MDA, MDG, class, predicted
  6. [6] § Technical Validation › Anchor-based label transfer and validation ↔ Code.R, lines 739–816 · score 0.56 · Decrease Gini, Decrease Accuracy, Random forest, MDA, MDG, predicted cell
  7. [7] § Background & Summary ↔ Code.R, lines 201–258 · score 0.55 · intestinal epithelium, stem cells, DEGs, colonocytes, immune, gene

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 816 lines · 28 KB · CC-BY-4.0 · 7 matches

  1. library(Seurat)
  2. library(SeuratObject)
  3. library(dplyr)
  4. library(Matrix)
  5. library(ggplot2)
  6. library(patchwork)
  7. # Loading datasets
  8. WT_data <- Read10X(data.dir = "WT")
  9. WT <- CreateSeuratObject(counts = WT_data, project = "WT")
  10. PINK1_data <- Read10X(data.dir = "PINK1")
  11. PINK1 <- CreateSeuratObject(counts = PINK1_data, project = "PINK1")
  12. # Assigning the dataset metadata
  13. WT$dataset <- "WT"
  14. PINK1$dataset <- "PINK1"
  15. WT <- JoinLayers(WT)
  16. PINK1 <- JoinLayers(PINK1)
  17. # Merging datasets
  18. merged <- merge(WT, y = PINK1, add.cell.ids = c("WT", "PINK1"), project = "Gut_tissue")
  19. Idents(object = merged) <- "Gut_tissue"
  20. merged[["percent.mt"]] <- PercentageFeatureSet(merged, pattern = "^mt-")
  21. # QC metrics
  22. VlnPlot(merged, features = c("nFeature_RNA", "nCount_RNA", "percent.mt"), ncol = 3)
  23. merged <- subset(merged, subset = nFeature_RNA > 200 & nFeature_RNA < 3000 &
  24. nCount_RNA > 500 & nCount_RNA < 10000 &
  25. percent.mt < 10)
  26. FeatureScatter(merged, feature1 = "nCount_RNA", feature2 = "nFeature_RNA", raster = FALSE)
  27. dim(WT)
  28. dim(merged)
  29. head(merged)
  30. # preprocessing
  31. merged <- NormalizeData(merged)
  32. merged <- FindVariableFeatures(merged)
  33. merged <- ScaleData(merged)
  34. merged <- RunPCA(merged)
  35. DimPlot(merged, reduction = "pca", group.by = "orig.ident") + ggtitle("PCA After Integration")
  36. ElbowPlot(merged)
  37. merged <- FindNeighbors(merged, dims = 1:20, reduction = "pca")
  38. merged <- FindClusters(merged, resolution = 2, cluster.name = "unintegrated_clusters")
  39. merged <- RunUMAP(merged, dims = 1:30, reduction = "pca", reduction.name = "umap.unintegrated")
  40. merged <- RunTSNE(merged, dims = 1:30, reduction = "pca", reduction.name = "tsne.unintegrated")
  41. DimPlot(merged, reduction = "umap.unintegrated", group.by = "orig.ident") + ggtitle("umap Before Integration")
  42. DimPlot(merged, reduction = "tsne.unintegrated", group.by = "orig.ident") +
  43. ggtitle("t-SNE Before Integration")
  44. merged <- IntegrateLayers(
  45. object = merged, method = CCAIntegration,
  46. orig.reduction = "pca", new.reduction = "integrated.cca",
  47. verbose = FALSE
  48. )
  49. merged <- FindNeighbors(merged, reduction = "integrated.cca", dims = 1:20)
  50. merged <- FindClusters(merged, resolution = 0.2, cluster.name = "cca_clusters")
  51. merged <- RunUMAP(merged, reduction = "integrated.cca", dims = 1:20, reduction.name = "umap.cca")
  52. DimPlot(merged, reduction = "umap.cca", group.by = "orig.ident") + ggtitle("umap After Integration")
  53. merged <- RunTSNE(merged, reduction = "integrated.cca", dims = 1:20, reduction.name = "tsne.cca")
  54. DimPlot(merged, reduction = "tsne.cca", group.by = "orig.ident") + ggtitle("tsne After Integration")
  55. # Create a vector of different resolutions
  56. #resolutions <- c(0.1, 0.15, 0.2, 0.25, 0.3, 0.5)
  57. # Loop through resolutions and find clusters at each level
  58. #for (res in resolutions) {
  59. #merged <- FindClusters(merged, resolution = res, cluster.name = paste0("res_", res))
  60. #}
  61. #library(patchwork) # For combining plots
  62. # Generate UMAP plots for different resolutions
  63. #plots <- lapply(resolutions, function(res) {
  64. #DimPlot(merged, reduction = "umap.cca", group.by = paste0("res_", res)) + ggtitle(paste("Resolution:", res))
  65. #})
  66. # Combine all plots
  67. #wrap_plots(plots, ncol = 2)
  68. #table(merged$seurat_clusters) # Default clustering (last applied resolution)
  69. #for (res in resolutions) {
  70. #print(table([email hidden][[paste0("res_", res)]]))
  71. #}
  72. DimPlot(merged, reduction = "umap.cca", label = TRUE)
  73. DimPlot(merged, reduction = "tsne.cca", label = TRUE)
  74. merged<-JoinLayers(merged)
  75. [email hidden]
  76. Idents(merged) <- "cca_clusters"
  77. # colors selection
  78. cluster_colors <- c(
  79. "#1f78b4", "#33a02c", "#e31a1c", "#ff7f00", "#6a3d9a",
  80. "#b15928", "#b2df8a", "#fb9a99", "#fdbf6f", "#cab2d6",
  81. "#a6cee3", "#ffff99", "#8c510a", "#d73027", "#4575b4",
  82. "#f46d43", "#74add1", "#d53e4f"
  83. )
  84. # UMAP for WT (Keeping Cluster Colors)
  85. p1 <- DimPlot(
  86. merged,
  87. reduction = "tsne.cca",
  88. group.by = "cca_clusters",
  89. cells = WhichCells(merged, expression = orig.ident == "WT"),
  90. cols = cluster_colors,
  91. pt.size = 0.1
  92. ) +
  93. ggtitle("WT") +
  94. theme_minimal() +
  95. theme(
  96. plot.title = element_text(size = 18, face = "bold"),
  97. legend.position = "right",
  98. axis.text = element_text(size = 14),
  99. axis.title = element_text(size = 16)
  100. )
  101. # UMAP for PINK1 KO (Keeping Cluster Colors)
  102. p2 <- DimPlot(
  103. merged,
  104. reduction = "tsne.cca",
  105. group.by = "cca_clusters",
  106. cells = WhichCells(merged, expression = orig.ident == "PINK1"),
  107. cols = cluster_colors,
  108. pt.size = 0.1
  109. ) +
  110. ggtitle("PINK1 KO") +
  111. theme_minimal() +
  112. theme(
  113. plot.title = element_text(size = 18, face = "bold"),
  114. legend.position = "right",
  115. axis.text = element_text(size = 14),
  116. axis.title = element_text(size = 16)
  117. )
  118. # Combining both plots side by side
  119. final_plot <- p1 + p2
  120. table(merged$cca_clusters, merged$orig.ident)
  121. DEGs <- FindAllMarkers(merged, only.pos = TRUE, min.pct = 0.25, logfc.threshold = 0.25)
  122. View(DEGs)
  123. # Count total cells per condition (WT & PINK1 KO)
  124. total_cells <- table(merged$orig.ident)
  125. table(merged$orig.ident)
  126. # Count the number of cells in each cluster per condition
  127. cluster_counts <- as.data.frame(table(merged$cca_clusters, merged$orig.ident))
  128. # Rename columns
  129. colnames(cluster_counts) <- c("cca_clusters", "dataset", "cluster_counts")
  130. # Calculate relative abundance (% of total cells in each condition)
  131. cluster_counts <- cluster_counts %>%
  132. group_by(dataset) %>%
  133. mutate(Relative_Abundance = (cluster_counts / sum(cluster_counts)) * 100)
  134. head(cluster_counts)
  135. # Plotting the relative abundance of each cluster in WT vs. PINK1 KO
  136. ggplot(cluster_counts, aes(x = cca_clusters, y = Relative_Abundance, fill = dataset)) +
  137. geom_bar(stat = "identity", position = "dodge", width = 0.7) +
  138. theme_minimal() +
  139. labs(title = "",
  140. x = "Cluster",
  141. y = "Relative Abundance (%)") +
  142. theme(
  143. axis.text.x = element_text(angle = 45, hjust = 1, size = 14, face = "bold"),
  144. axis.text.y = element_text(size = 12, face = "bold"),
  145. axis.title = element_text(size = 14, face = "bold"),
  146. plot.title = element_text(size = 14, face = "bold", hjust = 0.5),
  147. legend.title = element_text(size = 14, face = "bold"),
  148. legend.text = element_text(size = 12)
  149. ) +
  150. scale_fill_manual(values = c("WT" = "#1f78b4", "PINK1" = "#e31a1c"))
  151. VlnPlot(merged, features = c("Pink1", "Map2", "Syt1"), raster = FALSE, pt.size = 0, group.by = "cca_clusters")
  152. top_DEGs <- DEGs %>%
  153. group_by(cluster) %>%
  154. top_n(n = 3, wt = avg_log2FC) %>%
  155. pull(gene)
  156. print(top_DEGs)
  157. View(DEGs)
  158. # color palette
  159. custom_colors <- c("#66c2a5", "#fc8d62", "#8da0cb", "#e78ac3", "#a6d854",
  160. "#ffd92f", "#e5c494", "#b3b3b3", "#1f78b4", "#33a02c")
  161. # dot plot
  162. dot_plot <- DotPlot(merged, features = top_DEGs, cols = custom_colors, dot.scale = 6) +
  163. theme_minimal() +
  164. labs(title = "Top DEGs Across Clusters", x = "Clusters", y = "Genes") +
  165. theme(
  166. axis.text.x = element_text(angle = 45, hjust = 1, size = 9, face = "bold"),
  167. axis.text.y = element_text(size = 9, face = "bold"),
  168. plot.title = element_text(size = 14, face = "bold"),
  169. legend.position = "right"
  170. )
  171. print(dot_plot)
  172. ############ cell type assingnment
  173. merged$cell_type <- as.character(Idents(merged))
  174. merged$cell_type[merged$seurat_clusters == 0] <- "Stem cell"
  175. merged$cell_type[merged$seurat_clusters == 1] <- "Stem cell"
  176. merged$cell_type[merged$seurat_clusters == 2] <- "Goblet cells"
  177. merged$cell_type[merged$seurat_clusters == 3] <- "Goblet cells"
  178. merged$cell_type[merged$seurat_clusters == 4] <- "Colonocytes"
  179. merged$cell_type[merged$seurat_clusters == 5] <- "Myofibroblasts"
  180. merged$cell_type[merged$seurat_clusters == 6] <- "Pericytes"
  181. merged$cell_type[merged$seurat_clusters == 7] <- "Stem cell"
  182. merged$cell_type[merged$seurat_clusters == 8] <- "Intestinal epithelial cells"
  183. merged$cell_type[merged$seurat_clusters == 9] <- "Immune cells"
  184. merged$cell_type[merged$seurat_clusters == 10] <- "Immune cells"
  185. merged$cell_type[merged$seurat_clusters == 11] <- "Colonocytes"
  186. merged$cell_type[merged$seurat_clusters == 12] <- "Enteroendocrine cells"
  187. merged$cell_type[merged$seurat_clusters == 13] <- "Endothelial cells"
  188. merged$cell_type[merged$seurat_clusters == 14] <- "Endothelial cells"
  189. merged$cell_type[merged$seurat_clusters == 15] <- "Schwann cells"
  190. merged$cell_type[merged$seurat_clusters == 16] <- "Stem cell"
  191. merged$cell_type[merged$seurat_clusters == 17] <- "Schwann cells"
  192. Idents(merged) <- merged$cell_type
  193. my_colors <- c(
  194. "#E64B35", "#4DBBD5", "#00A087", "#3C5488", "#F39B7F", "#8491B4", "#91D1C2",
  195. "#DC0000", "#7E6148", "#B09C85", "#FFDC91", "#00A6D6", "#FF61CC", "#1B9E77",
  196. "#D95F02", "#7570B3", "#E7298A", "#66A61E", "#E6AB02", "#A6761D", "#A6CEE3",
  197. "#FB9A99", "#FDBF6F", "#CAB2D6", "#B3DE69", "#BC80BD", "#9E9AC8",
  198. "#8DD3C7", "#FCCDE5", "#D9D9D9", "#FFB3AB", "#A1D99B", "#B3CDE3", "#C49C94",
  199. "#F2B447", "#C7CEEA", "#FF6F91", "#88CCEE", "#DDCC77", "#A594F9" # NEW FINAL COLOR
  200. )
  201. umap_plot <- DimPlot(merged, reduction = "tsne.cca", label = FALSE, cols = cluster_colors) +
  202. theme_classic(base_size = 14) +
  203. theme(
  204. axis.text = element_blank(),
  205. axis.ticks = element_blank(),
  206. axis.title = element_blank(),
  207. legend.title = element_text(size = 12, face = "bold"),
  208. legend.text = element_text(size = 10)
  209. ) +
  210. ggtitle("")
  211. ggsave("wt_vs_pink1.png", plot = umap_plot, width = 5.1, height = 3.1, units = "in", dpi = 300, bg = "transparent")
  212. # Define Ordered Marker Genes
  213. features_ordered <- c(
  214. "Fcgbp", "Zg16", "Muc2", # Goblet Cells
  215. "Ptprc", "Lyz2", # Myeloid Immune Cells
  216. "Cd19", "Ms4a1", # Lymphoid Immune Cells
  217. "Plp1", "Ngfr", "Sox10", # Schwann Cells
  218. "Pecam1", "Lyve1", "Flt4", # Endothelial Cells
  219. "Chga", "Chgb", "Tph1", "Neurod1", # Enteroendocrine Cells
  220. "Col1a1", "Pdgfra", "Acta2", "Vim", # Fibroblasts
  221. "Mki67", "Pcna", # Proliferating Cells
  222. "Prom1", # Adult Stem Cell
  223. "Tfrc", "Clock", "Ccnd2", "Fut9", # Stem Cells
  224. "mt-Nd4", "Cox6a1", "Rplp2", # Mitochondria-Rich Cells
  225. "Tex9", "Cluh", # Migrating Cells
  226. "Slc26a3", "Slc15a1", "Epcam" # Colonocytes/Enterocytes
  227. )
  228. library(Seurat)
  229. library(ggplot2)
  230. # Define ordered marker genes
  231. features_ordered <- c(
  232. "Chga", "Neurod1", # Enteroendocrine
  233. "Lgr5", "Axin2", # Stem Cells
  234. "Prom1", "Car1", # Colonocytes
  235. "Slc15a1", "Epcam", # Enterocytes
  236. "Col1a1", "Pdgfra", # Fibroblasts
  237. "Zg16", "Muc2", # Goblet Cells
  238. "Pecam1", "Lyve1", # Endothelial
  239. "Ptprc", "Lyz2", # Immune Cells
  240. "Acta2", "Vim", # Myofibroblasts
  241. "Plp1", "Ngfr" # Schwann Cells
  242. )
  243. # Create DotPlot
  244. dot_plot <- DotPlot(merged, features = features_ordered) +
  245. RotatedAxis() +
  246. ggtitle("Marker Genes Across Gut Cell Types") +
  247. scale_color_gradientn(
  248. colors = c("#e0ecf4", "#9ecae1", "#fdae61", "#f46d43", "#a50026"), # Blue to red custom palette
  249. name = "Expression"
  250. ) +
  251. scale_size(range = c(1, 6), name = "% Cells") + # Adjust dot sizes
  252. theme_classic(base_size = 14) +
  253. theme(
  254. axis.text.x = element_text(angle = 45, hjust = 1, size = 10, face = "bold"),
  255. axis.text.y = element_text(size = 10, face = "bold"),
  256. plot.title = element_text(hjust = 0.5, size = 16, face = "bold"),
  257. panel.border = element_rect(color = "black", fill = NA, size = 1) # Add box line
  258. )
  259. # UMAP for WT
  260. p1 <- DimPlot(
  261. merged,
  262. reduction = "tsne.cca",
  263. cells = WhichCells(merged, expression = orig.ident == "WT"),
  264. cols = cluster_colors,
  265. pt.size = 0.1
  266. ) +
  267. ggtitle("WT") +
  268. theme_classic(base_size = 14) + # Clean base theme
  269. theme(
  270. plot.title = element_text(size = 18, face = "bold", hjust = 0.5),
  271. axis.text = element_text(size = 12),
  272. axis.title = element_text(size = 14),
  273. legend.position = "none", # Hide legend
  274. panel.border = element_rect(color = "black", fill = NA, size = 1), # Add box line
  275. panel.grid = element_blank() # Remove grid lines
  276. )
  277. # UMAP for PINK1 KO
  278. p2 <- DimPlot(
  279. merged,
  280. reduction = "tsne.cca",
  281. cells = WhichCells(merged, expression = orig.ident == "PINK1"),
  282. cols = cluster_colors,
  283. pt.size = 0.1
  284. ) +
  285. ggtitle("PINK1 KO") +
  286. theme_classic(base_size = 14) +
  287. theme(
  288. plot.title = element_text(size = 18, face = "bold", hjust = 0.5),
  289. axis.text = element_text(size = 12),
  290. axis.title = element_text(size = 14),
  291. legend.position = "none",
  292. panel.border = element_rect(color = "black", fill = NA, size = 1),
  293. panel.grid = element_blank()
  294. )
  295. # Combine plots
  296. library(patchwork)
  297. final_plot <- p1 + p2
  298. final_plot
  299. ggsave("wt_vs_pink1_diff.png", plot = final_plot, width = 5.1, height = 2.8, units = "in", dpi = 300, bg = "transparent")
  300. library(dplyr)
  301. library(ggplot2)
  302. # Creating a data frame of counts
  303. metadata_df <- as.data.frame([email hidden])
  304. summary_df <- metadata_df %>%
  305. group_by(orig.ident, cell_type) %>%
  306. summarize(count = n(), .groups = "drop")
  307. # Calculating proportions
  308. summary_df <- summary_df %>%
  309. group_by(orig.ident) %>%
  310. mutate(proportion = count / sum(count))
  311. # colors for cell types
  312. my_colors <- cluster_colors
  313. # Stacked bar plot
  314. ggplot(summary_df, aes(x = orig.ident, y = proportion, fill = cell_type)) +
  315. geom_col(color = "black", size = 0.3) + # Thin black outline
  316. scale_y_continuous(labels = scales::percent_format()) +
  317. scale_fill_manual(values = my_colors) +
  318. labs(
  319. title = "",
  320. x = "Condition",
  321. y = "Cell Type Proportion",
  322. fill = "Cell Type"
  323. ) +
  324. theme_classic(base_size = 14) +
  325. theme(
  326. axis.text.x = element_text(angle = 45, hjust = 1, size = 12, face = "bold"),
  327. axis.text.y = element_text(size = 12),
  328. plot.title = element_text(size = 16, face = "bold", hjust = 0.5),
  329. panel.border = element_rect(color = "black", fill = NA, size = 1) # Add box line
  330. )
  331. #####################################label transfer validation######################
  332. reference <- mouse_colon_male
  333. query <- wt_pink1ko
  334. reference <- NormalizeData(reference)
  335. reference <- FindVariableFeatures(reference)
  336. query <- NormalizeData(query)
  337. query <- FindVariableFeatures(query)
  338. anchors <- FindTransferAnchors(
  339. reference = reference,
  340. query = query,
  341. normalization.method = "LogNormalize",
  342. reduction = "pcaproject",
  343. reference.reduction = "pca"
  344. )
  345. reference <- RunUMAP(reference, reduction = "pca", dims = 1:20, return.model = T)
  346. reference[["umap.new"]] <- CreateDimReducObject(
  347. embeddings = reference[["umap"]]@cell.embeddings,
  348. key = "UMAPnew_", # Set the key with an underscore
  349. assay = "RNA"
  350. )
  351. Validation <- MapQuery(
  352. anchorset = anchors,
  353. query = query,
  354. reference = reference,
  355. refdata = list(celltype = reference$Annotation),
  356. reference.reduction = "pca",
  357. reduction.model = "umap"
  358. )
  359. colors <- c(
  360. "#e6194B", "#3cb44b", "#ffe119", "#4363d8", "#f58231", "#911eb4",
  361. "#46f0f0", "#f032e6", "#bcf60c", "#fabebe", "#008080", "#e6beff",
  362. "#9A6324", "#fffac8", "#800000", "#aaffc3", "#808000", "#ffd8b1",
  363. "#0000FF", "#FF1493", "#00FF7F", "#FF4500", "#1E90FF", "#FF00FF",
  364. "#ADFF2F", "#FF6347", "#20B2AA"
  365. )
  366. table(Validation$predicted.celltype) # See label distribution
  367. p1 <- DimPlot(Validation, group.by = "cell_type", reduction = "tsne.cca", label = FALSE, cols = colors) +
  368. ggtitle("Actual cell type")
  369. p2 <- DimPlot(Validation, group.by = "predicted.celltype", reduction = "tsne.cca", label = FALSE, cols = colors) +
  370. ggtitle("Predicted cell type")
  371. p1 + p2
  372. ########### validation dataet to show predictic score distrtibution in density plot
  373. library(ggplot2)
  374. library(viridis)
  375. # Create a data frame of all predicted cell type scores
  376. scores <- [email hidden]$predicted.celltype.score
  377. cell_types <- [email hidden]$predicted.celltype
  378. data <- data.frame(cell_type = cell_types, score = scores)
  379. # Set the y-axis scale limits to be between 0 and the maximum density value
  380. max_density <- max(density(data$score)$y)
  381. ylim_max <- max_density + max_density*0.2 # Set a little extra space at the top
  382. ggplot(data, aes(x = score, fill = cell_type)) +
  383. geom_density(alpha = 0.5) +
  384. xlab("Predicted cell type score") +
  385. ylab("Density") +
  386. ggtitle("Distribution of predicted cell type scores") +
  387. ylim(0, ylim_max)
  388. # Generate a slightly darker palette
  389. num_types <- length(unique(data$cell_type))
  390. palette_colors <- viridis(num_types, option = "C")
  391. # Create a histogram of all predicted cell type scores
  392. validation_score <- ggplot(data, aes(x = score, fill = cell_type)) +
  393. geom_histogram(alpha = 0.5, position = "identity", bins = 30) +
  394. xlab("Predicted cell type score") +
  395. ylab("Count") +
  396. scale_fill_manual(values = palette_colors) +
  397. labs(
  398. title = "Distribution of Predicted Cell Type Scores",
  399. x = "Predicted Cell Type Score",
  400. y = "Cell Count",
  401. fill = "Predicted Cell Type"
  402. ) +
  403. theme_classic(base_size = 14) + # No background, no grid
  404. theme(
  405. plot.title = element_text(face = "bold", size = 14, hjust = 0.5),
  406. axis.title = element_text(face = "bold"),
  407. axis.text = element_text(color = "black"),
  408. legend.title = element_text(face = "bold"),
  409. legend.key.size = unit(0.4, "cm"),
  410. legend.text = element_text(size = 10),
  411. legend.position = "right",
  412. panel.border = element_rect(colour = "black", fill = NA, linewidth = 0.7) # Add outer box
  413. )
  414. ggsave("validation_score.png", plot = validation_score, width = 6.19, height = 3.42, units = "in", dpi = 300, bg = "transparent")
  415. ggplot(data, aes(x = score, fill = cell_type)) +
  416. geom_histogram(bins = 30, color = "black", fill = "#FFB3BA") +
  417. facet_wrap(~ cell_type, scales = "free_y", ncol = 5) +
  418. theme_minimal() +
  419. labs(
  420. title = "Predicted Cell Type Score Distribution per Cell Type",
  421. x = "Predicted Score", y = "Cell Count"
  422. ) +
  423. theme(strip.text = element_text(face = "bold"))
  424. table(Validation$predicted.celltype)
  425. table(Validation_high$predicted.celltype)
  426. Validation_high <- subset(Validation, subset = predicted.celltype.score > 0.5)
  427. table(Validation$predicted.celltype.score > 0.5)
  428. DimPlot(Validation_high, group.by = "predicted.celltype", reduction = "tsne.cca", label = FALSE, cols = colors) +
  429. ggtitle("Predicted cell type")
  430. p1 <- DimPlot(Validation_high, group.by = "cell_type", reduction = "tsne.cca", label = FALSE, cols = colors) +
  431. ggtitle("Actual cell type")
  432. p2 <- DimPlot(Validation_high, group.by = "predicted.celltype", reduction = "tsne.cca", label = TRUE, cols = colors) +
  433. ggtitle("Predicted cell type")
  434. p1 + p2
  435. table(Validation_high$predicted.celltype)
  436. table(Validation$predicted.celltype)
  437. library(randomForest)
  438. library(caret)
  439. library(pheatmap)
  440. library(ggplot2)
  441. library(multiROC)
  442. #1. data
  443. expr_mat <- GetAssayData(Validation_high, slot = "data") %>% as.matrix() %>% t()
  444. labels <- as.factor([email hidden]$predicted.celltype) # Use predicted label
  445. num_features <- 2000
  446. var_genes <- FindVariableFeatures(Validation_high, nfeatures = num_features) %>% VariableFeatures()
  447. expr_mat_reduced <- expr_mat[, var_genes]
  448. # 2. Build ML dataframe
  449. ml_data <- as.data.frame(expr_mat_reduced)
  450. ml_data$cell_type <- labels
  451. colnames(ml_data) <- make.names(colnames(ml_data))
  452. # Train/Test Split
  453. set.seed(123)
  454. trainIndex <- createDataPartition(ml_data$cell_type, p = 0.7, list = FALSE)
  455. train_data <- ml_data[trainIndex, ]
  456. test_data <- ml_data[-trainIndex, ]
  457. # Train RF Classifier
  458. set.seed(123)
  459. rf_model <- randomForest(cell_type ~ ., data = train_data, ntree = 500, importance = TRUE)
  460. # Predict on test set
  461. test_predictions <- predict(rf_model, newdata = test_data)
  462. test_probs <- predict(rf_model, newdata = test_data, type = "prob")
  463. # ----------------------------
  464. # Confusion Matrix & Enhanced Visualization
  465. # ----------------------------
  466. conf_mat <- confusionMatrix(test_predictions, test_data$cell_type)
  467. print(conf_mat)
  468. # Convert to matrix for heatmap
  469. cm <- conf_mat$table
  470. cm_df <- as.data.frame(cm)
  471. colnames(cm_df) <- c("Prediction", "Reference", "Freq")
  472. cm_df$Freq_log <- log1p(cm_df$Freq)
  473. # ggplot2 heatmap (log scale for low counts)
  474. ggplot(cm_df, aes(x = Reference, y = Prediction, fill = Freq_log)) +
  475. geom_tile(color = "grey90", linewidth = 0.3) +
  476. geom_text(aes(label = Freq), size = 3.5, color = "black") +
  477. scale_fill_gradientn(colors = c("white", "lightyellow", "orange", "red", "darkred"),
  478. name = "log(Freq+1)",
  479. guide = guide_colorbar(barwidth = 10, barheight = 1)) +
  480. labs(title = "Random Forest Confusion Matrix",
  481. x = "Actual Cell Type", y = "Predicted Cell Type") +
  482. theme_minimal(base_size = 13) +
  483. theme(
  484. axis.text.x = element_text(angle = 45, hjust = 1),
  485. plot.title = element_text(hjust = 0.5, face = "bold"),
  486. legend.position = "top"
  487. )
  488. # ----------------------------
  489. # Filter Confusion Matrix: remove all-zero rows/columns
  490. # ----------------------------
  491. # Extract raw confusion matrix
  492. cm <- conf_mat$table
  493. # Keep only classes that appear in both predicted and actual
  494. keep_classes <- intersect(rownames(cm)[rowSums(cm) > 0],
  495. colnames(cm)[colSums(cm) > 0])
  496. cm_filtered <- cm[keep_classes, keep_classes]
  497. # Convert to dataframe for ggplot
  498. cm_df <- as.data.frame(cm_filtered)
  499. colnames(cm_df) <- c("Prediction", "Reference", "Freq")
  500. cm_df$Freq_log <- log1p(cm_df$Freq)
  501. ggplot(cm_df, aes(x = Reference, y = Prediction, fill = Freq_log)) +
  502. geom_tile(color = "grey90", linewidth = 0.3) +
  503. geom_text(aes(label = Freq), size = 3.5, color = "black") +
  504. scale_fill_gradientn(colors = c("white", "grey90", "lightyellow","red", "darkred"),
  505. name = "log(Freq+1)",
  506. guide = guide_colorbar(barwidth = 10, barheight = 1)) +
  507. labs(title = "Random Forest Confusion Matrix",
  508. x = "Actual Cell Type", y = "Predicted Cell Type") +
  509. theme_minimal(base_size = 13) +
  510. theme(
  511. axis.text.x = element_text(angle = 45, hjust = 1),
  512. plot.title = element_text(hjust = 0.5, face = "bold"),
  513. legend.position = "top"
  514. )
  515. # Filtered cm_df is already prepared
  516. cm_df$Freq_log <- log1p(cm_df$Freq)
  517. random <- ggplot(cm_df, aes(x = Reference, y = Prediction, fill = Freq_log)) +
  518. # Draw tiles only where Freq > 0
  519. geom_tile(data = subset(cm_df, Freq > 0), color = "grey90", linewidth = 0.3) +
  520. # Add text labels only for nonzero cells
  521. geom_text(data = subset(cm_df, Freq > 0),
  522. aes(label = Freq), size = 3.5, color = "black") +
  523. scale_fill_gradientn(colors = c("white", "grey90", "lightyellow","red", "darkred"),
  524. name = "log(Freq+1)",
  525. guide = guide_colorbar(barwidth = 10, barheight = 1),
  526. na.value = "white") + # Make zero cells blank
  527. labs(title = "Random Forest Confusion Matrix",
  528. x = "Actual Cell Type", y = "Predicted Cell Type") +
  529. theme_minimal(base_size = 13) +
  530. theme(
  531. axis.text.x = element_text(angle = 45, hjust = 1),
  532. plot.title = element_text(hjust = 0.5, face = "bold"),
  533. legend.position = "top"
  534. )
  535. # assumes: cm_df with columns Reference, Prediction, Freq
  536. cm_df$Freq_log <- log1p(cm_df$Freq)
  537. random <- ggplot(subset(cm_df, Freq > 0),
  538. aes(x = Reference, y = Prediction, fill = Freq_log)) +
  539. geom_tile(color = "grey90", linewidth = 0.3) +
  540. scale_fill_gradientn(
  541. colors = c("white", "grey90", "lightyellow","red", "darkred"),
  542. name = "log(Freq+1)",
  543. guide = guide_colorbar(barwidth = 10, barheight = 1),
  544. na.value = "white"
  545. ) +
  546. coord_fixed(expand = FALSE) +
  547. labs(
  548. title = "Random Forest Confusion Matrix",
  549. x = "Actual Cell Type", y = "Predicted Cell Type"
  550. ) +
  551. theme_classic(base_size = 13) +
  552. theme(
  553. legend.position = "top",
  554. plot.title = element_text(hjust = 0.5, face = "bold"),
  555. axis.text.x = element_text(angle = 45, hjust = 1),
  556. panel.border = element_rect(color = "black", fill = NA, linewidth = 0.7) # boxed frame
  557. # no panel.grid in theme_classic(), so gridlines are gone
  558. )
  559. random
  560. ggsave("rf.png", plot = random, width = 6.00, height = 6.13, units = "in", dpi = 300, bg = "transparent")
  561. getwd()
  562. # Extract confusion matrix
  563. cm <- conf_mat$table
  564. # Calculate per-class accuracy (Recall)
  565. class_accuracy <- diag(cm) / rowSums(cm)
  566. # Remove NA (where row sum = 0, i.e., no true cells of that type)
  567. class_accuracy <- class_accuracy[!is.na(class_accuracy)]
  568. # Make dataframe
  569. acc_df <- data.frame(
  570. CellType = names(class_accuracy),
  571. Accuracy = class_accuracy
  572. )
  573. # Plot as horizontal barplot
  574. library(ggplot2)
  575. ggplot(acc_df, aes(x = reorder(CellType, Accuracy), y = Accuracy)) +
  576. geom_col(fill = "#9ABB50") +
  577. geom_text(aes(label = sprintf("%.2f", Accuracy)),
  578. hjust = -0.2, size = 3.5, color = "black") +
  579. scale_y_continuous(limits = c(0, 1.05), expand = c(0,0)) +
  580. coord_flip() +
  581. labs(title = "Per-Class prediction accuracy (random forest)",
  582. x = "Cell Type", y = "Accuracy (recall)") +
  583. theme_minimal(base_size = 14) +
  584. theme(
  585. axis.text.x = element_text(size = 12),
  586. axis.text.y = element_text(size = 12),
  587. plot.title = element_text(hjust = 0.5, face = "bold")
  588. )
  589. library(ggplot2)
  590. library(ggplot2)
  591. ggplot(acc_df, aes(x = reorder(CellType, Accuracy), y = Accuracy)) +
  592. geom_col(fill = "#9ABB50", width = 0.6) + # Sleek bar width
  593. scale_y_continuous(limits = c(0, 1.05), expand = c(0,0)) +
  594. coord_flip() +
  595. labs(
  596. title = "Per-Class Prediction Accuracy",
  597. x = NULL,
  598. y = "Accuracy (Recall)"
  599. ) +
  600. theme_classic(base_size = 10) +
  601. theme(
  602. axis.text.x = element_text(size = 10, color = "black"),
  603. axis.text.y = element_text(size = 10, color = "black"),
  604. axis.ticks.y = element_blank(),
  605. plot.title = element_text(hjust = 0.5, face = "bold", size = 10),
  606. panel.grid.major.y = element_blank(),
  607. panel.grid.minor = element_blank()
  608. )
  609. library(dplyr)
  610. library(ggplot2)
  611. # Extract importance from trained RF
  612. imp_mat <- randomForest::importance(rf_model)
  613. imp_df <- data.frame(Gene = rownames(imp_mat), imp_mat, row.names = NULL)
  614. ## ---- Top 20 by MeanDecreaseAccuracy (MDA) ----
  615. top20_mda <- imp_df %>%
  616. filter(!is.na(MeanDecreaseAccuracy)) %>%
  617. arrange(desc(MeanDecreaseAccuracy)) %>%
  618. slice_head(n = 20)
  619. p_mda <- ggplot(top20_mda,
  620. aes(x = MeanDecreaseAccuracy,
  621. y = reorder(Gene, MeanDecreaseAccuracy))) +
  622. geom_point(size = 4, color = "#E69F00") + # orange dots
  623. labs(title = "Top 20 Features by Mean Decrease Accuracy",
  624. x = "Mean Decrease Accuracy", y = NULL) +
  625. theme_classic(base_size = 12) +
  626. theme(
  627. axis.text.y = element_text(size = 9),
  628. plot.title = element_text(hjust = 0.5, face = "bold"),
  629. panel.border = element_rect(color = "black", fill = NA, linewidth = 0.8)
  630. )
  631. ggsave("rf_top20_MDA_dotplot.png", p_mda, width = 7, height = 6, dpi = 300, bg = "white")
  632. ## ---- Top 20 by MeanDecreaseGini (MDG) ----
  633. top20_mdg <- imp_df %>%
  634. filter(!is.na(MeanDecreaseGini)) %>%
  635. arrange(desc(MeanDecreaseGini)) %>%
  636. slice_head(n = 20)
  637. p_mdg <- ggplot(top20_mdg,
  638. aes(x = MeanDecreaseGini,
  639. y = reorder(Gene, MeanDecreaseGini))) +
  640. geom_point(size = 4, color = "#56B4E9") + # blue dots
  641. labs(title = "Top 20 Features by Mean Decrease Gini",
  642. x = "Mean Decrease Gini", y = NULL) +
  643. theme_classic(base_size = 12) +
  644. theme(
  645. axis.text.y = element_text(size = 9),
  646. plot.title = element_text(hjust = 0.5, face = "bold"),
  647. panel.border = element_rect(color = "black", fill = NA, linewidth = 0.8)
  648. )
  649. ggsave("rf_top20_MDG_dotplot.png", p_mdg, width = 7, height = 6, dpi = 300, bg = "white")
  650. ########################
  651. # build a tiny data frame of diagonal positions (all class levels)
  652. lev <- levels(factor(cm_df$Reference, levels = unique(cm_df$Reference)))
  653. diag_df <- data.frame(
  654. Reference = factor(lev, levels = lev),
  655. Prediction = factor(lev, levels = lev)
  656. )
  657. diag_color <- "#1f1f1f" # <- change this to any accent (e.g., "#9C27B0", "#FF5722")
  658. random <- ggplot(subset(cm_df, Freq > 0),
  659. aes(x = Reference, y = Prediction, fill = log1p(Freq))) +
  660. geom_tile(color = "grey90", linewidth = 0.3) +
  661. # make the diagonal pop: thin white halo + colored stroke
  662. geom_tile(data = diag_df, fill = NA, color = "white", linewidth = 2.0) +
  663. geom_tile(data = diag_df, fill = NA, color = diag_color, linewidth = 1.1) +
  664. scale_fill_gradientn(
  665. colors = c("#FFFFFF", "#EEEEEE", "#F4F1D0", "#E57373", "#B71C1C"),
  666. name = "log(Freq+1)",
  667. guide = guide_colorbar(barwidth = 10, barheight = 1),
  668. na.value = "white"
  669. ) +
  670. coord_fixed(expand = FALSE) +
  671. labs(title = "Random Forest Confusion Matrix",
  672. x = "Actual Cell Type", y = "Predicted Cell Type") +
  673. theme_classic(base_size = 13) +
  674. theme(
  675. legend.position = "top",
  676. plot.title = element_text(hjust = 0.5, face = "bold"),
  677. axis.text.x = element_text(angle = 45, hjust = 1),
  678. panel.border = element_rect(color = "black", fill = NA, linewidth = 0.7)
  679. )

Code.R, under CC-BY-4.0 · at the source

Overview

Authors: Muhammad Junaid1,2,3, Soo Jung Park4, Yiseul Bae2,3,4, Eun Jeong Lee2,3,4, Su Bin Lim1,2,3
  1. Department of Biochemistry & Molecular Biology, Ajou University School of Medicine,Suwon, 16499 South Korea
  2. Department of Biomedical Science, Graduate School of Ajou University,Suwon, 16499 South Korea
  3. BK21 R&E initiative for Advanced Precision Medicine, Ajou University School of Medicine,Suwon, 16499 South Korea
  4. Department of Brain Science, Ajou University School of Medicine,Suwon, 16499 South Korea
Institutions: Ajou University (South Korea)
Journal: Scientific data, volume 13, issue 1, article 819
Dates: received 8 December 2025; accepted 31 March 2026; published online 6 April 2026
Type: Data paper · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41597-026-07193-4 · PMID 41942478 · PMCID PMC13230531 · OpenAlex W7150806569
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), mouse (organism), Parkinson's (population), methods / tools (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning
MeSH: Colon*, Protein Kinases*, Animals, Cell Nucleus, Mice, Mice, Knockout, PTEN-Induced Putative Kinase, RNA-Seq, Sequence Analysis, RNA (* major topic)
Journal subjects: Data Descriptor
Topic: Genetic factors in colorectal cancer (Pathology and Forensic Medicine, Medicine), according to OpenAlex
Funding: National Research Foundation of Korea (2022R1C1C1005741, RS-2023-00217595, RS-2025-02217836, RS-2020-NR049588, RS-2020-NR046270, 2022R1C1C1004756); Korea Health Industry Development Institute (HR22C1734, RS-2025-02310331)
Citations: not cited yet (Europe PMC); 22 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

figshare 30702425

License: CC-BY-4.0
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Languages: R (1)
Size: 3 files, 1 script
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: caret (1 file), ggplot2 (1 file), patchwork (1 file), pheatmap (1 file), randomForest (1 file), Seurat (1 file), tidyverse (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
  • 29 September 2026: the link answers (HTTP 200)
1 file
  • Code.R, R, 816 lines, 7 matches
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it points to the authors' code: figshare 30702425

Read it in the paper: doi.org/10.1038/s41597-026-07193-4.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s41597-026-07193-4.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 9 MeSH terms, 2 funders, 21 references.

Cite

This paper

Junaid, M., Park, S. J., Bae, Y., Lee, E. J., & Bin Lim, S. (2026). A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice. Scientific data, 13(1), 819. https://doi.org/10.1038/s41597-026-07193-4

BibTeX

@article{junaid2026single,
author = {Junaid, Muhammad and Park, Soo Jung and Bae, Yiseul and Lee, Eun Jeong and Bin Lim, Su},
title = {{A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice}},
journal = {Scientific data},
year = {2026},
month = apr,
volume = {13},
number = {1},
pages = {819},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/s41597-026-07193-4},
url = {https://doi.org/10.1038/s41597-026-07193-4},
pmid = {41942478},
pmcid = {PMC13230531}
}

RIS

TY - JOUR
AU - Junaid, Muhammad
AU - Park, Soo Jung
AU - Bae, Yiseul
AU - Lee, Eun Jeong
AU - Bin Lim, Su
TI - A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/04/06
VL - 13
IS - 1
SP - 819
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/s41597-026-07193-4
UR - https://doi.org/10.1038/s41597-026-07193-4
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41597-026-07193-4",
"type": "article-journal",
"title": "A single-nucleus RNA-seq dataset of the colon in Pink1-deficient and wild-type mice",
"container-title": "Scientific data",
"author": [
{
"family": "Junaid",
"given": "Muhammad"
},
{
"family": "Park",
"given": "Soo Jung"
},
{
"family": "Bae",
"given": "Yiseul"
},
{
"family": "Lee",
"given": "Eun Jeong"
},
{
"family": "Bin Lim",
"given": "Su"
}
],
"container-title-short": "Sci Data",
"volume": "13",
"issue": "1",
"page": "819",
"DOI": "10.1038/s41597-026-07193-4",
"PMID": "41942478",
"PMCID": "PMC13230531",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41597-026-07193-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
6
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.7717/peerj.21426 [code]
Integrated transcriptomic identification and validation reveal key autophagy-associated biomarkers in sleep deprivation.
Journal: PeerJ
In common: randomForest, caret, pheatmap, 4 other tools, genetics / omics, mouse
[2] doi:10.3390/ijms27156925 [code]
XGBoost-SHAP Interpretable Modeling Identifies and Validates an Eight-Gene Biomarker for Hepatic Encephalopathy Risk Prediction in Cirrhosis.
Journal: International journal of molecular sciences
In common: randomForest, caret, pheatmap, 4 other tools, genetics / omics
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: randomForest, caret, pheatmap, 4 other tools, mouse
[4] doi:10.1038/s41467-026-75723-0 [code]
Spatial transcriptomics reveals distinct cell type dynamics following opioid dependence in female mice with the common human μ-opioid receptor variant Oprm1 A118G.
Journal: Nature communications
In common: randomForest, pheatmap, Seurat, 3 other tools, genetics / omics, mouse
[5] doi:10.1002/advs.76205 [code]
CHCHD10 Mitigates Alzheimer's Disease-Related Phenotypes in Association With Epigenetic Remodeling in Directly Reprogrammed Neurons.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: randomForest, pheatmap, Seurat, 3 other tools, genetics / omics
[6] doi:10.3390/ijms27104466 [code]
Uncovering the Key Circuit FOSL2/FOS/EGR3/EGR1, Contributing to the Hyperexcitability of Excitatory Neurons in the Epileptic Temporal Cortex and Hippocampus.
Journal: International journal of molecular sciences
In common: caret, pheatmap, Seurat, 3 other tools, genetics / omics
[7] doi:10.1038/s41467-026-77170-3 [code]
DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.
Journal: Nature communications
In common: randomForest, caret, pheatmap, 2 other tools, genetics / omics
[8] doi:10.1016/j.celrep.2026.117298 [code]
Midbrain endocannabinoids actuate dopamine-based action selection.
Journal: Cell reports
In common: randomForest, caret, patchwork, 2 other tools, mouse
[9] doi:10.1038/s41593-026-02387-w [code]
Developing mouse inhibitory neuron single-cell transcriptomes reveal distinct modes of cell-type diversification.
Journal: Nature neuroscience
In common: randomForest, Seurat, patchwork, 2 other tools, genetics / omics, mouse
[10] doi:10.1016/j.nbd.2026.107379 [code]
DYRK1A and Parkinson's disease, facts and hypotheses.
Journal: Neurobiology of disease
In common: ggplot2, tidyverse, Parkinson's, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.