OSCR

A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Materials and methods › Datasets ↔ rakd7/Figure2_Markdown.Rmd, lines 200–326 · score 0.93 · L4 Rorb, L5 Deptor, L6B Ctgf, L6A Foxp2, L1 Ndnf, cortical area
  2. [2] § Materials and methods › Analysis of cell morphologies › Data preparation ↔ rakd7/Figure3_Markdown.Rmd, lines 116–245 · score 0.87 · log transformed, Inertia Tensor, Redundant features, Hu Moment features, morphometric features, duplicated
  3. [3] § Materials and methods › Analysis of cell morphologies › Selecting the number of clusters ↔ rakd7/Figure3_Markdown.Rmd, lines 252–350 · score 0.86 · factoextra package, find_curve_elbow, fviz_nbclust, pathviewr package, cluster sum, squares
  4. [4] § Materials and methods › Analysis of cell morphologies › Data preparation ↔ rakd7/revised_Figure8_Markdown.Rmd, lines 101–147 · score 0.85 · minor axis lengths, aspect ratio, Inertia Tensor, Hu Moment, cell
  5. [5] § Materials and methods › Analysis of cell morphologies › Laminar distribution of size and shape categories ↔ rakd7/revised_Figure7_Markdown.Rmd, lines 336–432 · score 0.82 · Monte Carlo resampling, Laminar enrichment, layer combination, Empirical, binomial, fraction
  6. [6] § Materials and methods › Analysis of cell morphologies › Selecting the number of clusters ↔ rakd7/revised_Figure6_Heatmap_Markdown.Rmd, lines 100–185 · score 0.82 · factoextra package, find_curve_elbow, fviz_nbclust, pathviewr package, squares, sum
  7. [7] § Materials and methods › Analysis of cell morphologies › Cluster visualization ↔ rakd7/Figure3_Markdown.Rmd, lines 652–772 · score 0.81 · global structure, denSNE algorithm, density preserving, random seed, perplexity, arrangement
  8. [8] § Materials and methods › Image processing and cell morphology measurement › Identifying cortical layers ↔ rakd7/Figure2_Markdown.Rmd, lines 200–326 · score 0.71 · standardized cortical depth, cortical layer boundaries, laminar density, profiles, genes, cells
  9. [9] § Materials and methods › Analysis of cell morphologies › Cluster visualization ↔ rakd7/Figures4-5_Markdown.Rmd, lines 250–323 · score 0.68 · denSNE, random seed, scored morphometric, density preserving, perplexity, algorithm
  10. [10] § Materials and methods › Analysis of cell morphologies › Cluster phenotyping to generate size and shape categories ↔ rakd7/Figures4-5_Markdown.Rmd, lines 536–610 · score 0.64 · ward.D2, hierarchical clustering, phenotype matrix, linkage
  11. [11] § Results › High-dimensional analysis reveals large variation in PV+ soma morphology ↔ rakd7/Figure3_Markdown.Rmd, lines 652–772 · score 0.63 · Cluster centroids, density preserving, denSNE, RSKC cluster, median, weights
  12. [12] § Results › Laminar organization of PV+ soma morphologies reveals a shared structural logic with area-dependent modulation ↔ rakd7/revised_Figure7_Markdown.Rmd, lines 549–619 · score 0.58 · bubble charts, Laminar enrichment, cortical areas, Cohen, cortical layers, enriched
  13. [13] § Materials and methods › Analysis of cell morphologies › Robust sparse K-means clustering ↔ rakd7/revised_Figure8_Markdown.Rmd, lines 245–303 · score 0.57 · Robust Sparse, feature weighting, iterative, L1, RSKC, clusters
  14. [14] § Materials and methods › Analysis of cell morphologies › Robust sparse K-means clustering ↔ rakd7/Figure3_Markdown.Rmd, lines 384–476 · score 0.56 · Robust Sparse, feature weighting, iterative, L1, RSKC, clusters
  15. [15] § Results › Phenotyping reveals distinct size and shape classes of PV+ soma morphology ↔ rakd7/revised_Figure6_Heatmap_Markdown.Rmd, lines 250–305 · score 0.50 · contour complexity, asymmetric, polarized, compact, heatmap, module

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R Markdown · 773 lines · 29 KB · no license · 5 matches

  1. ---
  2. title: "Atlas of Parvalbumin-Positive Interneurons in the Mouse Visual and Somatosensory Cortex: A Morphology Resource"
  3. subtitle: |
  4. Figure 3
  5. *Markdown Correspondence: [email hidden]*
  6. author: "Maheshwar Panday, Leanne Monteiro, Ahad Daudi, Kathryn M. Murphy*"
  7. date: "`r Sys.Date()`"
  8. output:
  9. pdf_document:
  10. latex_engine: xelatex
  11. ---
  12. ```{r setup, include=FALSE}
  13. knitr::opts_chunk$set(echo = TRUE)
  14. ```
  15. # Loading Packages
  16. ```{r load.packages, message = FALSE, warning = FALSE}
  17. # ==================================================
  18. # Core Data Science & Data Handling
  19. # ==================================================
  20. library(tidyverse) # data wrangling, transformation, and visualization (
  21. library(here) # filepath and directory management
  22. # ==================================================
  23. # Clustering & Dimensionality Reduction
  24. # ==================================================
  25. library(RSKC) # robust and sparse k-means clustering
  26. library(Rtsne) # t-SNE dimensionality reduction for cluster visualization
  27. library(densvis) # density-preserving t-SNE embedding
  28. # ==================================================
  29. # Elbow Plot & Cluster Evaluation
  30. # ==================================================
  31. library(factoextra) # clustering diagnostics and elbow plots
  32. library(pathviewr) # elbow/cluster visualization utilities
  33. library(elbow) # elbow method for optimal cluster selection
  34. # ==================================================
  35. # Visualization & Plot Enhancement
  36. # ==================================================
  37. library(ggforce) # advanced ggplot extensions (facet_wrap_paginate, etc.)
  38. library(ggpubr) # publication-ready plots and plot arrangement
  39. library(ggExtra) # marginal density/histogram plots (ggMarginal)
  40. library(see) # additional geoms and themes (geom_halfviolin)
  41. library(ggrepel) # non-overlapping text labels in plots
  42. # ==================================================
  43. # Color Palettes & Fonts
  44. # ==================================================
  45. library(RColorBrewer) # qualitative and sequential color palettes
  46. library(viridis) # perceptually uniform color palettes
  47. library(showtext) # load and use custom fonts in plots
  48. # ==================================================
  49. # Heatmaps & Matrix Visualization
  50. # ==================================================
  51. library(pheatmap) # heatmap generation for clustering and expression matrices
  52. library(gplots) # heatmap.2 and additional plotting utilities
  53. # ==================================================
  54. # Plot Layout & Figure Composition
  55. # ==================================================
  56. library(gridExtra) # arrange multiple plots into grids
  57. library(patchwork) # compose multi-panel ggplot figures
  58. # ==================================================
  59. # Graphics Devices & Export
  60. # ==================================================
  61. library(R.devices) # devEval and advanced graphics device handling
  62. ```
  63. ```{r export.path, include = FALSE }
  64. export.path <- here::here("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Exploration/Jan2024_Pvalb_Variability_Study/30Jul2024_PvalbV1S1_RSKC")
  65. if(!dir.exists(export.path)) {
  66. dir.create(export.path, recursive = TRUE)
  67. }
  68. ```
  69. # Load Dataframes
  70. This chunk loads the dataframes needed to prepare the morphology data matrix for Robust Sparse K-Means Clustering
  71. Objects Created:
  72. | Object | Description |
  73. | -------------------------- | ---------------------------------------- |
  74. | `CellProfiler.Dataframe` | full cell morphology dataset loaded |
  75. | `column.sorting.dataframe` | column inclusion and exclusion reference |
  76. ```{r dataframes, message = FALSE, warning = FALSE }
  77. ## --------------------------------------------- ##
  78. ##### Read in the full CellProfiler Dataframe #####
  79. ## --------------------------------------------- ##
  80. CellProfiler.Dataframe <- read.csv("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Data/Jan2024_Pvalb_Variability/Analysis_Data/Pvalb_V1_S1_5Specimen_CellMorphologyNorm.csv")
  81. # remove any cells that lie below the bottom of the cortex :
  82. CellProfiler.Dataframe <- CellProfiler.Dataframe %>% filter (norm_y<= 1)
  83. # secondary csv - what columns to include/exclude
  84. column.sorting.dataframe <- read.csv("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Data/Jan2024_Pvalb_Variability/Analysis_Data/CellProfiler_ColumnSort_RKSC.csv")
  85. ```
  86. # Prepare Dataframes for RSKC
  87. This chunk prepares the CellProfiler dataset for RSKC by separating clustering features (97) from the descriptive metadata using the column-sorting dataframe. It creates a working.exp feature matrix and working.ids descriptor dataframe. It adds. unique object number for cell tracking and identifies/removes redundant, duplicated morphometric features. The seven Hu image moments are log-transformed to normalise their distributions and redundant columns are removed. Finally, the feature matrix is z-scored to place all features into a common numerical space for RSKC.
  88. Objects Created :
  89. | Object | Description |
  90. | ----------------------- | --------------------------------------- |
  91. | `RSKC.features` | columns included in clustering |
  92. | `RSKC.descriptors` | columns excluded from clustering |
  93. | `working.exp.raw` | raw clustering feature dataframe |
  94. | `working.ids` | descriptor dataframe with cell IDs |
  95. | `duplicate_columns` | list of identical feature columns |
  96. | `redundant.columns` | predefined redundant feature names |
  97. | `redundant.features.df` | dataframe of redundant features |
  98. | `columns.to.drop` | redundant columns removed from features |
  99. | `huMoment_columns` | Hu Moment feature column names |
  100. | `working.exp` | z-scored clustering feature matrix |
  101. ```{r dataframe.preparation, message = FALSE, warning = FALSE }
  102. #==================================#
  103. ###=== Dataframe Preparation ===###
  104. #==================================#
  105. ## -------------------------------------------------------------- ##
  106. ##### use the sorting dataframe to prepare the RSKC dataframes #####
  107. ## -------------------------------------------------------------- ##
  108. # features - these are the morphometric parameters used
  109. RSKC.features <-
  110. column.sorting.dataframe$column_names[
  111. column.sorting.dataframe$include_in_RSKC == "yes"
  112. ]
  113. # descriptors - these are the metdata columns
  114. RSKC.descriptors <-
  115. column.sorting.dataframe$column_names[
  116. column.sorting.dataframe$include_in_RSKC == "no"
  117. ]
  118. # subset the CellProfiler.Dataframe into features and descriptors
  119. # (using var names that match RSKC code )
  120. working.exp.raw <- CellProfiler.Dataframe[, RSKC.features]
  121. # inspection - are the right columns included in RSKC?
  122. # print (colnames (working.exp))
  123. # view(working.exp)
  124. working.ids <- CellProfiler.Dataframe[, !names(CellProfiler.Dataframe)%in%
  125. RSKC.features]
  126. # working.ids <- CellProfiler.Dataframe[, RSKC.descriptors]
  127. # inspection - are the right columns being set as descriptors
  128. # print (colnames (working.ids))
  129. # add a unique object number to identify each cell
  130. working.ids <- cbind(UniqueObjectNumber = 1:nrow(working.ids), working.ids)
  131. # inspection step - was this column added correctly?
  132. # view(working.ids)
  133. ## ----------------------------------------------------------- ##
  134. ##### removing the redundant - perfectly duplicated columns #####
  135. ## ----------------------------------------------------------- ##
  136. # Create an empty list to store sets of duplicate columns
  137. duplicate_columns <- list()
  138. # Loop over each column
  139. for (i in 1:(ncol(working.exp.raw) - 1)) {
  140. for (j in (i + 1):ncol(working.exp.raw)) {
  141. # Check if the columns are identical
  142. if (all(working.exp.raw[, i] == working.exp.raw[, j])) {
  143. # Add the pair of duplicate columns to the list
  144. duplicate_columns <- c(duplicate_columns,
  145. list(c(names(working.exp.raw)[i],
  146. names(working.exp.raw)[j])))
  147. }
  148. }
  149. }
  150. # Print the sets of duplicate columns
  151. print(paste("duplicated columns are :",duplicate_columns))
  152. # identify redundant columns
  153. redundant.columns <- c("AreaShape_Area",
  154. "AreaShape_CentralMoment_0_0",
  155. "AreaShape_SpatialMoment_0_0",
  156. "AreaShape_InertiaTensor_0_1",
  157. "AreaShape_InertiaTensor_1_0")
  158. # put the redundant columns in a dataframe of their own (inspect them)
  159. redundant.features.df <- working.exp.raw[redundant.columns]
  160. # view(redundant.features.df) # inspection step
  161. # drop the columns that are redundant from the working.exp
  162. columns.to.drop <- c("AreaShape_CentralMoment_0_0",
  163. "AreaShape_SpatialMoment_0_0",
  164. "AreaShape_InertiaTensor_0_1")
  165. working.exp.raw <- working.exp.raw[, !(names(working.exp.raw) %in%
  166. columns.to.drop)]
  167. ## ---------------------------------------------- ##
  168. ##### log transformation of Hu Moment Features #####
  169. ## ---------------------------------------------- ##
  170. # Extract columns with "HuMoment" in their names
  171. huMoment_columns <- grep("HuMoment", names(working.exp.raw), value = TRUE)
  172. # Apply the transformation to the extracted columns
  173. # This is to make the magnitudes of the Hu Moments
  174. # comparable to all the other morphometrics
  175. working.exp.raw[huMoment_columns] <- lapply(working.exp.raw[huMoment_columns],
  176. function(x) {
  177. -1 * sign(x) * log10(abs(x) + 1e-10)# Adding a small constant to avoid log(0)
  178. })
  179. ## ------------------------------------------------- ##
  180. ##### z-score the working.exp RKSC feature values #####
  181. ## ------------------------------------------------- ##
  182. # Z-score to make all features numerically comparable for analysis
  183. working.exp <- scale(working.exp.raw)
  184. # inspection step - view the features as they are passed to RSKC
  185. #head (working.exp)
  186. ```
  187. # Elbow plot - use this to determine an optimal number of clusters
  188. This code calculates the optimal number of clusters for the prepared dataset using three methods: fviz_nbclust() (factoextra) to create a standard elbow plot, find_curve_elbow() (pathviewr) to detect the elbow point numerically, and elbow() (ahasverus/elbow) as an alternative method. The within-cluster sum of squares (wss) values are extracted and combined with cluster sizes into a dataframe for analysis. The resulting elbow plots are generated and stored as R plot objects for inspection. Finally, the selected optimal number of clusters (k.clusters) and the visual plot (elbow.point) are ready for export or further analysis.
  189. This code was run on a MacBookPro M3Pro Apple Silicon Chip, 18GB Unified Memory
  190. runtime for this code was ~ 2.45 minutes
  191. Objects Created :
  192. | Object | Description |
  193. | --------------- | ------------------------------------ |
  194. | `df.elbow` | fviz_nbclust elbow plot object |
  195. | `wss_values` | within-cluster sum of squares values |
  196. | `cluster_sizes` | numeric vector of cluster indices |
  197. | `df` | dataframe of cluster sizes and WSS |
  198. | `k.clusters` | optimal number of clusters (numeric) |
  199. | `elbow.point` | recorded elbow plot object |
  200. ```{r elbow.plot, fig.align = "center", fig.height = 4, fig.width = 5, warning = FALSE, message = FALSE}
  201. start.time <- Sys.time()
  202. ##========================##
  203. ###==== Elbow Plots =====###
  204. ##========================##
  205. ## ---------------------------------------------------------- ##
  206. ##### 1. Using the fviz_nbclust() from factoextra package. #####
  207. ## We will also plot the Elbow plot ##
  208. ## ---------------------------------------------------------- ##
  209. df.elbow <- fviz_nbclust(working.exp, kmeans, method = c("wss"),
  210. k.max=70)+
  211. theme(axis.text.x = element_text(size=10,angle = 90, vjust = 0.5))
  212. #Plot the Elbow Plot
  213. df.elbow
  214. ## ------------------------------------------------------ ##
  215. ##### 2. Using find_curve_elbow from pathviewr package #####
  216. ## ------------------------------------------------------ ##
  217. # This is done by drawing an (imaginary) line between the first observation and
  218. # the final observation. Then, the distance between that line and each
  219. # observation is calculated. The "elbow" of the curve is
  220. # the observation that maximizes this distance.
  221. # Extract the within sums of squares values calculated using fviz_nbclust
  222. wss_values<- df.elbow$data$y
  223. # Adding the cluster sizes
  224. cluster_sizes <- 1:length(wss_values)
  225. # Combining the wss and the cluster sizes into a dataframe
  226. df <- cbind (cluster_sizes,wss_values) %>% as.data.frame()
  227. find_curve_elbow(df, plot_curve = TRUE)
  228. k.clusters <- find_curve_elbow(df, plot_curve = F) # store optimal k
  229. ## -------------------------------------------------------------------------- ##
  230. ##### 3. Using elbow from devtools::install_github("ahasverus/elbow") #####
  231. # More information: https://nicolascasajus.fr/elbow/articles/introduction.html #
  232. ## -------------------------------------------------------------------------- ##
  233. elbow(data = df) # output k= 13
  234. elbow.point <- recordPlot()
  235. plot.new() ## clean up device
  236. elbow.point # redraw
  237. # # Save plot
  238. # devEval(type = "pdf",
  239. # path = paste(export.path),
  240. # name = "morphometric elbowplot 12 oligo_neuron",
  241. # height = 7, width = 7,
  242. # expr = {
  243. # print(elbow.point)
  244. # }
  245. # )
  246. ## -------------------------------- ##
  247. #### 4. Exporting the Elbow Plot #####
  248. ## -------------------------------- ##
  249. # # Set the file path and name for the PDF file
  250. #pdf(file = paste0(export.path, "/ElbowPlot_clustering_k.pdf"),
  251. # height = 7, width = 10)
  252. #
  253. # # Plot the graph
  254. print(elbow.point)
  255. #
  256. # # Close the PDF device
  257. #dev.off()
  258. end.time <- Sys.time()
  259. elapsed.time <- end.time-start.time
  260. print (elapsed.time)
  261. ```
  262. # cleaning up the elbow plot
  263. make it visually presentable
  264. ```{r elbow.plot.cleanup , fig.align = "center", fig.height = 4, fig.width = 5}
  265. # Create the elbow plot
  266. cleaned.up.elbowplot <- ggplot(df, aes(x = cluster_sizes, y = wss_values)) +
  267. geom_point() + # Show dots for each data point
  268. geom_line() + # Connect the dots with lines
  269. labs(x ="",
  270. y = "") +
  271. scale_x_continuous(breaks = seq(0, 70, by = 10)) +
  272. # red vertical line at the optimal k
  273. geom_vline(xintercept = k.clusters, color = "red", linetype = 1) +
  274. theme_classic() + # Use a classic theme
  275. theme(
  276. axis.text.x = element_text(hjust = 1, vjust = 0.5,
  277. family = "Helvetica", size = 8),
  278. axis.text.y = element_text(family = "Helvetica", size = 8))
  279. # Print the plot
  280. print(cleaned.up.elbowplot)
  281. ggsave(paste0(export.path, "/CleanedUp_ElbowPlot.pdf"),
  282. plot = cleaned.up.elbowplot,
  283. height = 4, width = 3.5)
  284. ```
  285. # Clustering morphology data with RSKC
  286. This code performs Robust & Sparse K-means Clustering (RSKC) on the prepared morphometric feature matrix (working.exp) using the optimal cluster number (k.clusters). It stores the cluster assignments and feature weights for each run, along with relevant metadata for each cell. After clustering, the cluster labels are combined with cell coordinates and specimen/region metadata for visualization. Finally, the RSKC labels are exported as a CSV file for downstream analysis or plotting. The for loop architecutre supports multiple runs of RSKC
  287. This code was run on a MacBookPro M3Pro Apple Silicon Chip, 18GB Unified Memory
  288. Approximate Runtime : ~1.02 hrs
  289. Objects Created :
  290. | Object | Description |
  291. | -------------- | --------------------------------------------- |
  292. | `clust.num` | number of clusters for RSKC |
  293. | `runs` | number of RSKC iterations |
  294. | `rskc.results` | list storing RSKC output objects |
  295. | `rskc.labels` | dataframe of cluster assignments and metadata |
  296. | `rskc.weights` | dataframe of feature weights per run |
  297. | `i` | loop index for runs |
  298. ```{r performing.RSKC, message = FALSE, warning = FALSE }
  299. # RSKC: Perform Robust & Sparse K-means Clustering ####
  300. # Select number of clusters
  301. clust.num <- k.clusters
  302. # Select number of iterations
  303. runs <- 1
  304. # Prepare objects to store RSKC results
  305. rskc.results <- list()
  306. rskc.labels <- working.ids %>%
  307. dplyr::select(UniqueObjectNumber, Metadata_SampleID,
  308. Metadata_SpecimenID, Metadata_Region,
  309. norm_x, norm_y)
  310. rskc.weights <- data.frame("Morphometric" = colnames(working.exp))
  311. # seed is initalised for reproducibility
  312. set.seed(12345)
  313. for (i in 1:runs) {
  314. # Perform RSKC
  315. rskc.results[[i]] <- RSKC(d = working.exp,
  316. alpha = 0.1,
  317. nstart = 1000,
  318. ncl = clust.num,
  319. scaling = F,
  320. L1 = sqrt(ncol(working.exp)))
  321. # Store cluster assignments for run i
  322. rskc.labels[i+4] <- rskc.results[[i]]$labels
  323. colnames(rskc.labels)[i+4] <- paste("Run_", i, sep = "")
  324. # Store feature weights for run i
  325. rskc.weights[i+1] <- rskc.results[[i]]$weights
  326. colnames(rskc.weights)[i+1] <- paste("Run_", i, sep = "")
  327. # Uncomment to track progress
  328. print(paste(i,"/", runs, sep = ""))
  329. }
  330. ## ---------------------------------------- ##
  331. ##### Melt the RSKC Dataframe for ggplot #####
  332. ## ---------------------------------------- ##
  333. # add the relevant metadata columns to the rskc.labels dataframe
  334. rskc.labels$norm_x<- working.ids$norm_x
  335. rskc.labels$norm_y <- working.ids$norm_y
  336. rskc.labels$Metadata_SpecimenID <- working.ids$Metadata_SpecimenID
  337. rskc.labels$Metadata_Region <- working.ids$Metadata_Region
  338. # view(rskc.labels) # uncomment to inspect the labels
  339. # export the RSKC labels as a csv
  340. # write.csv(rskc.labels,
  341. # file = file.path(export.path,
  342. # paste0(clust.num, "Pvalb_V1_S1_RSKC_labels.csv")),
  343. # row.names = FALSE)
  344. ```
  345. this chunk prepares the summary dataframe of RSKC feature weights.
  346. ```{r prepare.summary.weights}
  347. # Calculate the mean and standard deviation for each feature
  348. rskc.weights.summary <- rskc.weights %>%
  349. gather(run, weight, -Morphometric) %>%
  350. group_by(Morphometric) %>%
  351. summarise(
  352. mean = mean(weight, na.rm = TRUE),
  353. sd = sd(weight, na.rm = TRUE),
  354. se = sd / sqrt(n())
  355. ) %>%
  356. arrange(desc(mean))
  357. ```
  358. ```{r include = FALSE }
  359. # this chunk produces the first version of the RSKC weights plot -features are ordered by decreasing weight
  360. # a correctly formatted version of the plot is included in a separate markdown for phenotyping and denSNE plotting .
  361. # using the feature weights outpur for each trial, take the average weight for each feature and plot that in a histogram.
  362. # add error bars to the plot - standard error
  363. # annotate the bars with the average feature weight rounded to five decimal places (legibility and ease of reference)
  364. ## ----------------------------------------- ##
  365. ##### calculate average weights for each feature ####
  366. ## ----------------------------------------- ##
  367. # Calculate the mean and standard deviation for each feature
  368. rskc.weights.summary <- rskc.weights %>%
  369. gather(run, weight, -Morphometric) %>%
  370. group_by(Morphometric) %>%
  371. summarise(
  372. mean = mean(weight, na.rm = TRUE),
  373. sd = sd(weight, na.rm = TRUE),
  374. se = sd / sqrt(n())
  375. ) %>%
  376. arrange(desc(mean)) # Order by mean in decreasing order
  377. # remove the "AreaShape_" prefix from all the morphometric names in the dataframe
  378. rskc.weights.summary$Morphometric <- gsub("AreaShape_", "", rskc.weights.summary$Morphometric)
  379. # Change the order of Morphometric factor levels to match the order of the means
  380. rskc.weights.summary$Morphometric <- factor(rskc.weights.summary$Morphometric, levels = rskc.weights.summary$Morphometric)
  381. ## -------------------------- ##
  382. ##### plotting the weights #####
  383. ## -------------------------- ##
  384. # Plot the weights with error bars
  385. weights.plot <- ggplot(rskc.weights.summary, aes(x = Morphometric, y = mean)) +
  386. geom_bar(stat = "identity", fill = "lightsteelblue2", colour = "black") +
  387. geom_errorbar(aes(ymin = mean - se, ymax = mean + se), width = 0.2) +
  388. geom_text(aes(y = mean/2, label = round(mean, 5)),
  389. angle = 90, family = "Helvetica", size = 2)+
  390. theme_classic() +
  391. theme(axis.text.x = element_text(angle = 90, hjust = 1,
  392. vjust = 0.5, family = "Helvetica", size = 7),
  393. axis.text.y = element_text(family = "Helvetica", size = 7),
  394. axis.title.x = element_text(family = "Helvetica", size = 16),
  395. axis.title.y = element_text(family = "Helvetica", size = 16)) +
  396. #labs (x = "", y = "")
  397. labs(x = "Morphometric", y = "RSKC Feature Weight")
  398. # Print the plot
  399. print(weights.plot)
  400. ## -------------------------- ##
  401. ##### export the data and plot #####
  402. ## -------------------------- ##
  403. # # Export the weights for 100 runs
  404. # write.csv(rskc.weights, file = paste0(export.path, "/RSKC_Weights.csv"), row.names = FALSE)
  405. #
  406. # # Export the summary dataframe
  407. #write.csv(rskc.weights.summary, file = paste0(export.path, "/Summary_RSKC_Weights.csv"), row.names = FALSE)
  408. # Export the barplot with error bars
  409. #ggsave(filename = file.path(export.path, "Sep23_RSKC_WeightsPlot.pdf"), plot = weights.plot, height = 7, width = 10)
  410. ```
  411. # Multiplying feature weights to the z-scored features
  412. prepare morphology data for cluster visualisation as described in Balsor et al., 2021.
  413. This code prepares a weighted and ordered feature matrix for downstream t-SNE visualization using the RSKC feature weights. It reorders the z-scored morphometric features based on their average weights and multiplies each column by its corresponding mean weight to emphasize important features. Cluster assignments from RSKC are then added to the weighted dataframe, creating a combined dataset with both metadata and weighted morphometrics. Finally, both the weighted and unweighted z-scored dataframes with cluster labels are stored for analysis and visualization.
  414. | Object | Description |
  415. | ---------------------------------- | ------------------------------------------- |
  416. | `working.exp` | dataframe of z-scored morphometric features |
  417. | `working.exp.ordered` | features reordered by importance |
  418. | `weighted.working.exp.ordered` | weighted features for t-SNE |
  419. | `order` | morphometric feature ordering vector |
  420. | `CellProfiler.weighted.scored.df` | combined metadata and weighted features |
  421. | `CellProfiler.scored.clustered.df` | combined metadata and unweighted features |
  422. | `rownames(rskc.weights.summary)` | morphometric names assigned to rows |
  423. | `working.ids$cluster` | cluster assignments added to metadata |
  424. ```{r weighted.working.exp, message = FALSE, warning = FALSE }
  425. # remove the AreaShape_ prefix from the working.exp column names
  426. colnames(working.exp) <- sub("^AreaShape_", "", colnames(working.exp))
  427. # Remove "AreaShape_" prefix from Morphometric column
  428. rskc.weights.summary$Morphometric <-
  429. sub("^AreaShape_", "", rskc.weights.summary$Morphometric)
  430. ## ----------------------------------------- ##
  431. ##### weighted z-scored dataframe for tsne ####
  432. ## ----------------------------------------- ##
  433. # the column containing row names is the morphometric column
  434. rownames(rskc.weights.summary) <- rskc.weights.summary$Morphometric
  435. # convert to dataframe
  436. working.exp <- as.data.frame(working.exp)
  437. # obtain the order of rows - decreasing weight order of features
  438. order <- rskc.weights.summary$Morphometric
  439. # Reorder the columns of working.exp
  440. working.exp.ordered <- working.exp %>% select(order)
  441. # view(working.exp.ordered) # inspection step
  442. # sweep function to multiply each column of working.exp
  443. # by the corresponding average weight
  444. # MARGIN = 2 indicates the operation is applied to the columns
  445. weighted.working.exp.ordered <- sweep(working.exp.ordered,
  446. MARGIN = 2,
  447. STATS = rskc.weights.summary$mean,
  448. `*` )
  449. # view(weighted.working.exp.ordered) # inspection step
  450. # merge the cluster assignments to the weighted.exp.labels
  451. weighted.working.exp.ordered <- as.data.frame(weighted.working.exp.ordered)
  452. working.ids$cluster <-rskc.labels$Run_1
  453. CellProfiler.weighted.scored.df <- cbind(working.ids,
  454. weighted.working.exp.ordered )
  455. # inspection steps - uncomment to view the dataframes :
  456. # view(CellProfiler.weighted.scored.df)
  457. # print (head(CellProfiler.weighted.scored.df))
  458. # append the cluster labels to the z-scored dataframe
  459. CellProfiler.scored.clustered.df <- cbind(working.ids,working.exp.ordered)
  460. ## -- Inspection Steps : view the dataframes -- ##
  461. # view(CellProfiler.scored.clustered.df)
  462. # view(CellProfiler.weighted.scored.df)
  463. # print (head(CellProfiler.scored.clustered.df))
  464. # print (head(CellProfiler.weighted.scored.df))
  465. # print (colnames(CellProfiler.scored.clustered.df))
  466. ## -- Export the clustered z-scored dataframe -- ##
  467. # write.csv(CellProfiler.scored.clustered.df,
  468. # file= paste0(export.path,
  469. # "/Scored_Clustered_Pvalb_30Jul2024.csv"))
  470. ```
  471. # running densne
  472. perplexity term : 10
  473. random seed term :13579
  474. visualisation of clusters assigned from RSKC - to be used as a visualisation - the weights from RSKC smooth out the spacing and arrangement of clusters when they are mapped down from High D to 2D as a result of the densne algorithm.
  475. cluster assignment is currently using the clusters assigned from the first run of RSKC but average weights over the 100 trials are used in multiplying weights to feature z-scores
  476. This code was run on a MacBookPro M3 Pro Apple Silicon Chip, 18GB Unified Memory
  477. Approximate Runtime : ~ 49.3 seconds
  478. ```{r rskc.cluster.densne, fig.align="center", fig.height=5, fig.width=5, out.width="100%", message = FALSE}
  479. ## ----------------------------------------------- ##
  480. ##### Densne for a specific seed and perplexity #####
  481. ## ----------------------------------------------- ##
  482. # set the colour palette
  483. color.palette.vector.dense = c("darkgoldenrod2", "lightcoral", "cadetblue1",
  484. "olivedrab1", "pink1", "maroon3",
  485. "chartreuse3", "turquoise3", "navyblue",
  486. "slateblue3","plum", "deepskyblue", "khaki2")
  487. # Set your specific seed and perplexity
  488. # seed for initialisation
  489. # perplexity - influences local/global structure preservation
  490. seed <- 13579
  491. perplexity <-10
  492. # Set the seed for reproducibility
  493. set.seed(seed)
  494. # Run the density-preserving t-SNE algorithm
  495. dense_tsne <- densne(weighted.working.exp.ordered,
  496. perplexity = perplexity,
  497. check_duplicates = FALSE)
  498. # denSNE coordiantes are saved as a dataframe
  499. dense_tsne.df <- as.data.frame(dense_tsne)
  500. # Combine with the z-scored dataframe
  501. dense_tsne.df <- cbind(CellProfiler.weighted.scored.df, dense_tsne.df)
  502. # set the cluster column to factor - discrete labels not continuous numbers
  503. dense_tsne.df$RSKC_clusters <-
  504. factor(CellProfiler.weighted.scored.df$cluster)
  505. dense_tsne.df$Metadata_SpecimenID <-
  506. factor(CellProfiler.weighted.scored.df$Metadata_SpecimenID)
  507. # Add the seed and perplexity as columns
  508. dense_tsne.df$seed <- seed
  509. dense_tsne.df$perplexity <- perplexity
  510. # add the cluster label mapping for median sizes to the dataframe
  511. # 1. Calculate the median size for each cluster
  512. cluster_medians <- dense_tsne.df %>%
  513. group_by(RSKC_clusters) %>%
  514. summarise(median_size = median(Area, na.rm = TRUE)) %>%
  515. arrange(median_size) # ascending order: smallest to largest
  516. # 2. Create a mapping vector: cluster ID -> size rank
  517. median.size.order.mapping <- setNames(
  518. seq_len(nrow(cluster_medians)), # ranks 1, 2, ..., n
  519. cluster_medians$RSKC_clusters # names are cluster IDs
  520. )
  521. # 3. Apply this mapping to your dataframe
  522. dense_tsne.df$RSKC_size_clusters <- factor(
  523. median.size.order.mapping[as.character(dense_tsne.df$RSKC_clusters)],
  524. levels = seq_len(nrow(cluster_medians)) # ensures order is 1..n
  525. )
  526. # 4. Optional: inspect the mapping
  527. median.size.order.mapping
  528. # set the RSKC_clusters column as a factor
  529. dense_tsne.df$RSKC_clusters <- factor(dense_tsne.df$RSKC_clusters)
  530. # find the cluster centroids in the Dense space
  531. centroids <- aggregate(cbind(V1, V2) ~ RSKC_clusters, dense_tsne.df, mean)
  532. # Plotting Dense t-SNE plot
  533. p <- ggplot(dense_tsne.df, aes(x = V1, y = V2), size = 1.25) +
  534. geom_point(aes(color = as.factor(RSKC_size_clusters))) +
  535. scale_color_manual(values = color.palette.vector.dense) +
  536. theme_void() +
  537. theme(legend.position = "right")
  538. #theme (
  539. #legend.position = "none",
  540. #axis.title = element_text(family = "Helvetica", size = 100),
  541. #axis.text = element_text (family = "Helvetica", size = 80)
  542. #)
  543. # plot the centroids
  544. p <- p + geom_point(data = centroids,
  545. aes(x = V1, y = V2),
  546. color = 'black',
  547. size = 5,
  548. shape = 4)
  549. print(p)
  550. ## -- export the plot as a pdf -- ##
  551. # ggsave(paste0(export.path,
  552. # "/2OrderedSizeCluster_perp10_seed13579_denSNE_PvalbV1S1.pdf"),
  553. # plot = p, width = 10, height = 10)
  554. ## -- write the dense t-sne dataframe to csv -- ##
  555. # write.csv(dense_tsne.df, paste0(export.path,
  556. # "/denSNE_weighted_zscored_size_clusters_PvalbV1S1.csv"), row.names = FALSE)
  557. ```

Figure3_Markdown.Rmd, no license · at the source

Overview

Authors: Maheshwar Panday1, Leanne Monteiro1, Ahad Daudi2, Kathryn M. Murphy1,2
  1. McMaster Neuroscience Graduate Program, McMaster University, Hamilton, ON, Canada
  2. Department of Psychology, Neuroscience and Behaviour, McMaster University, Hamilton, ON, Canada
Institutions: McMaster University (Canada)
Journal: Frontiers in neuroscience, volume 20, article 1848222
Dates: received 5 April 2026; accepted 12 May 2026; published online 10 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fnins.2026.1848222 · PMID 42359344 · PMCID PMC13290952 · OpenAlex W7164160799
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism), systems (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing
Keywords: cortical circuits, experience-dependent plasticity, laminar organization, parvalbumin interneurons, soma morphology, visual cortex (V1)
Topic: Neural dynamics and brain function (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 59 references in the paper

Abstract

Parvalbumin-positive (PV+) inhibitory interneurons are central components of experience-dependent plasticity in the visual cortex (V1). Anatomically, they are distributed across cortical layers and are traditionally classified as basket, chandelier, or bipolar cells based on dendritic and axonal arborization patterns. In parallel, physiological, transcriptomic, connectomic, and multimodal studies have revealed substantial diversity within PV+ populations, raising the question of how this diversity is organized within cortical circuits. In contrast, light microscopy studies based on soma labeling have primarily quantified PV+ cells using size and density, and the potential of soma morphology to capture this diverse organization remains unclear. To address this, we developed a high-throughput, data-driven approach to quantify PV+ soma morphology in > 14,000 cells from mouse V1 and somatosensory cortex (S1). Using 97 morphological features combined with clustering, phenotyping, and laminar mapping, we identified structured diversity in PV+ somas. PV+ cells were organized into 13 morphological clusters along partially independent gradients of size and shape. Phenotyping identified four size and five shape categories that describe PV+ cell diversity across cortical areas. Mapping these categories onto cortical layers revealed a structured organization in which specific morphologies are enriched within distinct laminar compartments. This organization aligns with cortical architecture and suggests that PV+ interneuron morphology is systematically related to circuit structure. These findings demonstrate that substantial morphological information can be extracted from standard PV+ labeling approaches using quantitative analysis. Together, this work provides a high-dimensional atlas of PV+ interneuron soma morphology in mouse V1 and S1 and establishes a framework for linking cellular anatomy to circuit organization and experience-dependent plasticity.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

OSF 8f2nr

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Languages: R (6)
Size: 14 files, 6 scripts
Software Heritage: not checked
Found in: “Data availability statement”
Holds: 6 notebooks
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (6 files), patchwork (5 files), cowplot (4 files), ggplot2 (4 files), pheatmap (3 files), easystats (2 files), circlize (1 file), ComplexHeatmap (1 file), ggpubr (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
6 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 6 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The datasets and code presented in this study can be found in the online repository 10.17605/OSF.IO/8F2NR (https://doi.org/10.17605/OSF.IO/8F2NR). The names of the repository/repositories and accession number(s) can be found in the article/Supplementary material.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 4 authors, 6 keywords, 1 funder, 53 references.

Cite

This paper

Panday, M., Monteiro, L., Daudi, A., & Murphy, K. M. (2026). A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex. Frontiers in neuroscience, 20, 1848222. https://doi.org/10.3389/fnins.2026.1848222

BibTeX

@article{panday2026high,
author = {Panday, Maheshwar and Monteiro, Leanne and Daudi, Ahad and Murphy, Kathryn M.},
title = {{A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex}},
journal = {Frontiers in neuroscience},
year = {2026},
month = jun,
volume = {20},
pages = {1848222},
publisher = {Frontiers Media SA},
issn = {1662-4548},
doi = {10.3389/fnins.2026.1848222},
url = {https://doi.org/10.3389/fnins.2026.1848222},
pmid = {42359344},
pmcid = {PMC13290952}
}

RIS

TY - JOUR
AU - Panday, Maheshwar
AU - Monteiro, Leanne
AU - Daudi, Ahad
AU - Murphy, Kathryn M.
TI - A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex
T2 - Frontiers in neuroscience
J2 - Front Neurosci
PY - 2026
DA - 2026/06/10
VL - 20
SP - 1848222
SN - 1662-4548
PB - Frontiers Media SA
DO - 10.3389/fnins.2026.1848222
UR - https://doi.org/10.3389/fnins.2026.1848222
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fnins.2026.1848222",
"type": "article-journal",
"title": "A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex",
"container-title": "Frontiers in neuroscience",
"author": [
{
"family": "Panday",
"given": "Maheshwar"
},
{
"family": "Monteiro",
"given": "Leanne"
},
{
"family": "Daudi",
"given": "Ahad"
},
{
"family": "Murphy",
"given": "Kathryn M."
}
],
"container-title-short": "Front Neurosci",
"volume": "20",
"page": "1848222",
"DOI": "10.3389/fnins.2026.1848222",
"PMID": "42359344",
"PMCID": "PMC13290952",
"ISSN": "1662-4548",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fnins.2026.1848222",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41586-026-10512-9 [code]
Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.
Journal: Nature
In common: circlize, ComplexHeatmap, pheatmap, 4 other tools, mouse, 4 references
[2] doi:10.1038/s41467-026-74753-y [code]
A human-specific microRNA controls the timing of excitatory synaptogenesis.
Journal: Nature communications
In common: easystats, circlize, ComplexHeatmap, 6 other tools
[3] doi:10.1016/j.celrep.2026.117500 [code]
Spatio-molecular gene expression reflects dorsal anterior cingulate cortex structure and function in the human brain.
Journal: Cell reports
In common: easystats, circlize, ComplexHeatmap, 6 other tools
[4] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: circlize, ComplexHeatmap, pheatmap, 4 other tools, mouse, 2 references
[5] doi:10.1038/s44318-026-00806-z [code]
Interspecific diversity in the neuronal composition of the mammalian cortex arises from heterochrony in neurogenesis.
Journal: The EMBO journal
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, 1 reference
[6] doi:10.1038/s41586-026-10214-2 [code]
Multidimensional profiling of heterogeneity in supratentorial ependymomas.
Journal: Nature
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse, 1 reference
[7] doi:10.1038/s41593-026-02367-0 [code]
A reproducible three-dimensional model of human brain tissue to investigate physiological and disease-associated microglia phenotypes.
Journal: Nature neuroscience
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, 1 reference
[8] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse
[9] doi:10.1038/s41467-026-76341-6 [code]
Neonatal inflammation disrupts a temporally restricted postnatal Numb-enriched microglial state in mice.
Journal: Nature communications
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse
[10] doi:10.7717/peerj.21426 [code]
Integrated transcriptomic identification and validation reveal key autophagy-associated biomarkers in sleep deprivation.
Journal: PeerJ
In common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.