A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex.
The 15 matches
- [1] § Materials and methods › Datasets ↔ rakd7/Figure2_Markdown.Rmd, lines 200–326 · score 0.93 · L4 Rorb, L5 Deptor, L6B Ctgf, L6A Foxp2, L1 Ndnf, cortical area
- [2] § Materials and methods › Analysis of cell morphologies › Data preparation ↔ rakd7/Figure3_Markdown.Rmd, lines 116–245 · score 0.87 · log transformed, Inertia Tensor, Redundant features, Hu Moment features, morphometric features, duplicated
- [3] § Materials and methods › Analysis of cell morphologies › Selecting the number of clusters ↔ rakd7/Figure3_Markdown.Rmd, lines 252–350 · score 0.86 · factoextra package, find_curve_elbow, fviz_nbclust, pathviewr package, cluster sum, squares
- [4] § Materials and methods › Analysis of cell morphologies › Data preparation ↔ rakd7/revised_Figure8_Markdown.Rmd, lines 101–147 · score 0.85 · minor axis lengths, aspect ratio, Inertia Tensor, Hu Moment, cell
- [5] § Materials and methods › Analysis of cell morphologies › Laminar distribution of size and shape categories ↔ rakd7/revised_Figure7_Markdown.Rmd, lines 336–432 · score 0.82 · Monte Carlo resampling, Laminar enrichment, layer combination, Empirical, binomial, fraction
- [6] § Materials and methods › Analysis of cell morphologies › Selecting the number of clusters ↔ rakd7/revised_Figure6_Heatmap_Markdown.Rmd, lines 100–185 · score 0.82 · factoextra package, find_curve_elbow, fviz_nbclust, pathviewr package, squares, sum
- [7] § Materials and methods › Analysis of cell morphologies › Cluster visualization ↔ rakd7/Figure3_Markdown.Rmd, lines 652–772 · score 0.81 · global structure, denSNE algorithm, density preserving, random seed, perplexity, arrangement
- [8] § Materials and methods › Image processing and cell morphology measurement › Identifying cortical layers ↔ rakd7/Figure2_Markdown.Rmd, lines 200–326 · score 0.71 · standardized cortical depth, cortical layer boundaries, laminar density, profiles, genes, cells
- [9] § Materials and methods › Analysis of cell morphologies › Cluster visualization ↔ rakd7/Figures4-5_Markdown.Rmd, lines 250–323 · score 0.68 · denSNE, random seed, scored morphometric, density preserving, perplexity, algorithm
- [10] § Materials and methods › Analysis of cell morphologies › Cluster phenotyping to generate size and shape categories ↔ rakd7/Figures4-5_Markdown.Rmd, lines 536–610 · score 0.64 · ward.D2, hierarchical clustering, phenotype matrix, linkage
- [11] § Results › High-dimensional analysis reveals large variation in PV+ soma morphology ↔ rakd7/Figure3_Markdown.Rmd, lines 652–772 · score 0.63 · Cluster centroids, density preserving, denSNE, RSKC cluster, median, weights
- [12] § Results › Laminar organization of PV+ soma morphologies reveals a shared structural logic with area-dependent modulation ↔ rakd7/revised_Figure7_Markdown.Rmd, lines 549–619 · score 0.58 · bubble charts, Laminar enrichment, cortical areas, Cohen, cortical layers, enriched
- [13] § Materials and methods › Analysis of cell morphologies › Robust sparse K-means clustering ↔ rakd7/revised_Figure8_Markdown.Rmd, lines 245–303 · score 0.57 · Robust Sparse, feature weighting, iterative, L1, RSKC, clusters
- [14] § Materials and methods › Analysis of cell morphologies › Robust sparse K-means clustering ↔ rakd7/Figure3_Markdown.Rmd, lines 384–476 · score 0.56 · Robust Sparse, feature weighting, iterative, L1, RSKC, clusters
- [15] § Results › Phenotyping reveals distinct size and shape classes of PV+ soma morphology ↔ rakd7/revised_Figure6_Heatmap_Markdown.Rmd, lines 250–305 · score 0.50 · contour complexity, asymmetric, polarized, compact, heatmap, module
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 773 lines · 29 KB · no license · 5 matches
- ---
- title: "Atlas of Parvalbumin-Positive Interneurons in the Mouse Visual and Somatosensory Cortex: A Morphology Resource"
- subtitle: |
- Figure 3
- *Markdown Correspondence: [email hidden]*
- author: "Maheshwar Panday, Leanne Monteiro, Ahad Daudi, Kathryn M. Murphy*"
- date: "`r Sys.Date()`"
- output:
- pdf_document:
- latex_engine: xelatex
- ---
- ```{r setup, include=FALSE}
- knitr::opts_chunk$set(echo = TRUE)
- ```
- # Loading Packages
- ```{r load.packages, message = FALSE, warning = FALSE}
- # ==================================================
- # Core Data Science & Data Handling
- # ==================================================
- library(tidyverse) # data wrangling, transformation, and visualization (
- library(here) # filepath and directory management
- # ==================================================
- # Clustering & Dimensionality Reduction
- # ==================================================
- library(RSKC) # robust and sparse k-means clustering
- library(Rtsne) # t-SNE dimensionality reduction for cluster visualization
- library(densvis) # density-preserving t-SNE embedding
- # ==================================================
- # Elbow Plot & Cluster Evaluation
- # ==================================================
- library(factoextra) # clustering diagnostics and elbow plots
- library(pathviewr) # elbow/cluster visualization utilities
- library(elbow) # elbow method for optimal cluster selection
- # ==================================================
- # Visualization & Plot Enhancement
- # ==================================================
- library(ggforce) # advanced ggplot extensions (facet_wrap_paginate, etc.)
- library(ggpubr) # publication-ready plots and plot arrangement
- library(ggExtra) # marginal density/histogram plots (ggMarginal)
- library(see) # additional geoms and themes (geom_halfviolin)
- library(ggrepel) # non-overlapping text labels in plots
- # ==================================================
- # Color Palettes & Fonts
- # ==================================================
- library(RColorBrewer) # qualitative and sequential color palettes
- library(viridis) # perceptually uniform color palettes
- library(showtext) # load and use custom fonts in plots
- # ==================================================
- # Heatmaps & Matrix Visualization
- # ==================================================
- library(pheatmap) # heatmap generation for clustering and expression matrices
- library(gplots) # heatmap.2 and additional plotting utilities
- # ==================================================
- # Plot Layout & Figure Composition
- # ==================================================
- library(gridExtra) # arrange multiple plots into grids
- library(patchwork) # compose multi-panel ggplot figures
- # ==================================================
- # Graphics Devices & Export
- # ==================================================
- library(R.devices) # devEval and advanced graphics device handling
- ```
- ```{r export.path, include = FALSE }
- export.path <- here::here("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Exploration/Jan2024_Pvalb_Variability_Study/30Jul2024_PvalbV1S1_RSKC")
- if(!dir.exists(export.path)) {
- dir.create(export.path, recursive = TRUE)
- }
- ```
- # Load Dataframes
- This chunk loads the dataframes needed to prepare the morphology data matrix for Robust Sparse K-Means Clustering
- Objects Created:
- | Object | Description |
- | -------------------------- | ---------------------------------------- |
- | `CellProfiler.Dataframe` | full cell morphology dataset loaded |
- | `column.sorting.dataframe` | column inclusion and exclusion reference |
- ```{r dataframes, message = FALSE, warning = FALSE }
- ## --------------------------------------------- ##
- ##### Read in the full CellProfiler Dataframe #####
- ## --------------------------------------------- ##
- CellProfiler.Dataframe <- read.csv("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Data/Jan2024_Pvalb_Variability/Analysis_Data/Pvalb_V1_S1_5Specimen_CellMorphologyNorm.csv")
- # remove any cells that lie below the bottom of the cortex :
- CellProfiler.Dataframe <- CellProfiler.Dataframe %>% filter (norm_y<= 1)
- # secondary csv - what columns to include/exclude
- column.sorting.dataframe <- read.csv("/Users/maheshwar/Kathy Murphy Dropbox/Maheshwar Panday/MP-Adult-V1-ISH/MP-Adult-V1-ISH Matlab-R/MP-Adult-V1-ISH R Data/Jan2024_Pvalb_Variability/Analysis_Data/CellProfiler_ColumnSort_RKSC.csv")
- ```
- # Prepare Dataframes for RSKC
- This chunk prepares the CellProfiler dataset for RSKC by separating clustering features (97) from the descriptive metadata using the column-sorting dataframe. It creates a working.exp feature matrix and working.ids descriptor dataframe. It adds. unique object number for cell tracking and identifies/removes redundant, duplicated morphometric features. The seven Hu image moments are log-transformed to normalise their distributions and redundant columns are removed. Finally, the feature matrix is z-scored to place all features into a common numerical space for RSKC.
- Objects Created :
- | Object | Description |
- | ----------------------- | --------------------------------------- |
- | `RSKC.features` | columns included in clustering |
- | `RSKC.descriptors` | columns excluded from clustering |
- | `working.exp.raw` | raw clustering feature dataframe |
- | `working.ids` | descriptor dataframe with cell IDs |
- | `duplicate_columns` | list of identical feature columns |
- | `redundant.columns` | predefined redundant feature names |
- | `redundant.features.df` | dataframe of redundant features |
- | `columns.to.drop` | redundant columns removed from features |
- | `huMoment_columns` | Hu Moment feature column names |
- | `working.exp` | z-scored clustering feature matrix |
- ```{r dataframe.preparation, message = FALSE, warning = FALSE }
- #==================================#
- ###=== Dataframe Preparation ===###
- #==================================#
- ## -------------------------------------------------------------- ##
- ##### use the sorting dataframe to prepare the RSKC dataframes #####
- ## -------------------------------------------------------------- ##
- # features - these are the morphometric parameters used
- RSKC.features <-
- column.sorting.dataframe$column_names[
- column.sorting.dataframe$include_in_RSKC == "yes"
- ]
- # descriptors - these are the metdata columns
- RSKC.descriptors <-
- column.sorting.dataframe$column_names[
- column.sorting.dataframe$include_in_RSKC == "no"
- ]
- # subset the CellProfiler.Dataframe into features and descriptors
- # (using var names that match RSKC code )
- working.exp.raw <- CellProfiler.Dataframe[, RSKC.features]
- # inspection - are the right columns included in RSKC?
- # print (colnames (working.exp))
- # view(working.exp)
- working.ids <- CellProfiler.Dataframe[, !names(CellProfiler.Dataframe)%in%
- RSKC.features]
- # working.ids <- CellProfiler.Dataframe[, RSKC.descriptors]
- # inspection - are the right columns being set as descriptors
- # print (colnames (working.ids))
- # add a unique object number to identify each cell
- working.ids <- cbind(UniqueObjectNumber = 1:nrow(working.ids), working.ids)
- # inspection step - was this column added correctly?
- # view(working.ids)
- ## ----------------------------------------------------------- ##
- ##### removing the redundant - perfectly duplicated columns #####
- ## ----------------------------------------------------------- ##
- # Create an empty list to store sets of duplicate columns
- duplicate_columns <- list()
- # Loop over each column
- for (i in 1:(ncol(working.exp.raw) - 1)) {
- for (j in (i + 1):ncol(working.exp.raw)) {
- # Check if the columns are identical
- if (all(working.exp.raw[, i] == working.exp.raw[, j])) {
- # Add the pair of duplicate columns to the list
- duplicate_columns <- c(duplicate_columns,
- list(c(names(working.exp.raw)[i],
- names(working.exp.raw)[j])))
- }
- }
- }
- # Print the sets of duplicate columns
- print(paste("duplicated columns are :",duplicate_columns))
- # identify redundant columns
- redundant.columns <- c("AreaShape_Area",
- "AreaShape_CentralMoment_0_0",
- "AreaShape_SpatialMoment_0_0",
- "AreaShape_InertiaTensor_0_1",
- "AreaShape_InertiaTensor_1_0")
- # put the redundant columns in a dataframe of their own (inspect them)
- redundant.features.df <- working.exp.raw[redundant.columns]
- # view(redundant.features.df) # inspection step
- # drop the columns that are redundant from the working.exp
- columns.to.drop <- c("AreaShape_CentralMoment_0_0",
- "AreaShape_SpatialMoment_0_0",
- "AreaShape_InertiaTensor_0_1")
- working.exp.raw <- working.exp.raw[, !(names(working.exp.raw) %in%
- columns.to.drop)]
- ## ---------------------------------------------- ##
- ##### log transformation of Hu Moment Features #####
- ## ---------------------------------------------- ##
- # Extract columns with "HuMoment" in their names
- huMoment_columns <- grep("HuMoment", names(working.exp.raw), value = TRUE)
- # Apply the transformation to the extracted columns
- # This is to make the magnitudes of the Hu Moments
- # comparable to all the other morphometrics
- working.exp.raw[huMoment_columns] <- lapply(working.exp.raw[huMoment_columns],
- function(x) {
- -1 * sign(x) * log10(abs(x) + 1e-10)# Adding a small constant to avoid log(0)
- })
- ## ------------------------------------------------- ##
- ##### z-score the working.exp RKSC feature values #####
- ## ------------------------------------------------- ##
- # Z-score to make all features numerically comparable for analysis
- working.exp <- scale(working.exp.raw)
- # inspection step - view the features as they are passed to RSKC
- #head (working.exp)
- ```
- # Elbow plot - use this to determine an optimal number of clusters
- This code calculates the optimal number of clusters for the prepared dataset using three methods: fviz_nbclust() (factoextra) to create a standard elbow plot, find_curve_elbow() (pathviewr) to detect the elbow point numerically, and elbow() (ahasverus/elbow) as an alternative method. The within-cluster sum of squares (wss) values are extracted and combined with cluster sizes into a dataframe for analysis. The resulting elbow plots are generated and stored as R plot objects for inspection. Finally, the selected optimal number of clusters (k.clusters) and the visual plot (elbow.point) are ready for export or further analysis.
- This code was run on a MacBookPro M3Pro Apple Silicon Chip, 18GB Unified Memory
- runtime for this code was ~ 2.45 minutes
- Objects Created :
- | Object | Description |
- | --------------- | ------------------------------------ |
- | `df.elbow` | fviz_nbclust elbow plot object |
- | `wss_values` | within-cluster sum of squares values |
- | `cluster_sizes` | numeric vector of cluster indices |
- | `df` | dataframe of cluster sizes and WSS |
- | `k.clusters` | optimal number of clusters (numeric) |
- | `elbow.point` | recorded elbow plot object |
- ```{r elbow.plot, fig.align = "center", fig.height = 4, fig.width = 5, warning = FALSE, message = FALSE}
- start.time <- Sys.time()
- ##========================##
- ###==== Elbow Plots =====###
- ##========================##
- ## ---------------------------------------------------------- ##
- ##### 1. Using the fviz_nbclust() from factoextra package. #####
- ## We will also plot the Elbow plot ##
- ## ---------------------------------------------------------- ##
- df.elbow <- fviz_nbclust(working.exp, kmeans, method = c("wss"),
- k.max=70)+
- theme(axis.text.x = element_text(size=10,angle = 90, vjust = 0.5))
- #Plot the Elbow Plot
- df.elbow
- ## ------------------------------------------------------ ##
- ##### 2. Using find_curve_elbow from pathviewr package #####
- ## ------------------------------------------------------ ##
- # This is done by drawing an (imaginary) line between the first observation and
- # the final observation. Then, the distance between that line and each
- # observation is calculated. The "elbow" of the curve is
- # the observation that maximizes this distance.
- # Extract the within sums of squares values calculated using fviz_nbclust
- wss_values<- df.elbow$data$y
- # Adding the cluster sizes
- cluster_sizes <- 1:length(wss_values)
- # Combining the wss and the cluster sizes into a dataframe
- df <- cbind (cluster_sizes,wss_values) %>% as.data.frame()
- find_curve_elbow(df, plot_curve = TRUE)
- k.clusters <- find_curve_elbow(df, plot_curve = F) # store optimal k
- ## -------------------------------------------------------------------------- ##
- ##### 3. Using elbow from devtools::install_github("ahasverus/elbow") #####
- # More information: https://nicolascasajus.fr/elbow/articles/introduction.html #
- ## -------------------------------------------------------------------------- ##
- elbow(data = df) # output k= 13
- elbow.point <- recordPlot()
- plot.new() ## clean up device
- elbow.point # redraw
- # # Save plot
- # devEval(type = "pdf",
- # path = paste(export.path),
- # name = "morphometric elbowplot 12 oligo_neuron",
- # height = 7, width = 7,
- # expr = {
- # print(elbow.point)
- # }
- # )
- ## -------------------------------- ##
- #### 4. Exporting the Elbow Plot #####
- ## -------------------------------- ##
- # # Set the file path and name for the PDF file
- #pdf(file = paste0(export.path, "/ElbowPlot_clustering_k.pdf"),
- # height = 7, width = 10)
- #
- # # Plot the graph
- print(elbow.point)
- #
- # # Close the PDF device
- #dev.off()
- end.time <- Sys.time()
- elapsed.time <- end.time-start.time
- print (elapsed.time)
- ```
- # cleaning up the elbow plot
- make it visually presentable
- ```{r elbow.plot.cleanup , fig.align = "center", fig.height = 4, fig.width = 5}
- # Create the elbow plot
- cleaned.up.elbowplot <- ggplot(df, aes(x = cluster_sizes, y = wss_values)) +
- geom_point() + # Show dots for each data point
- geom_line() + # Connect the dots with lines
- labs(x ="",
- y = "") +
- scale_x_continuous(breaks = seq(0, 70, by = 10)) +
- # red vertical line at the optimal k
- geom_vline(xintercept = k.clusters, color = "red", linetype = 1) +
- theme_classic() + # Use a classic theme
- theme(
- axis.text.x = element_text(hjust = 1, vjust = 0.5,
- family = "Helvetica", size = 8),
- axis.text.y = element_text(family = "Helvetica", size = 8))
- # Print the plot
- print(cleaned.up.elbowplot)
- ggsave(paste0(export.path, "/CleanedUp_ElbowPlot.pdf"),
- plot = cleaned.up.elbowplot,
- height = 4, width = 3.5)
- ```
- # Clustering morphology data with RSKC
- This code performs Robust & Sparse K-means Clustering (RSKC) on the prepared morphometric feature matrix (working.exp) using the optimal cluster number (k.clusters). It stores the cluster assignments and feature weights for each run, along with relevant metadata for each cell. After clustering, the cluster labels are combined with cell coordinates and specimen/region metadata for visualization. Finally, the RSKC labels are exported as a CSV file for downstream analysis or plotting. The for loop architecutre supports multiple runs of RSKC
- This code was run on a MacBookPro M3Pro Apple Silicon Chip, 18GB Unified Memory
- Approximate Runtime : ~1.02 hrs
- Objects Created :
- | Object | Description |
- | -------------- | --------------------------------------------- |
- | `clust.num` | number of clusters for RSKC |
- | `runs` | number of RSKC iterations |
- | `rskc.results` | list storing RSKC output objects |
- | `rskc.labels` | dataframe of cluster assignments and metadata |
- | `rskc.weights` | dataframe of feature weights per run |
- | `i` | loop index for runs |
- ```{r performing.RSKC, message = FALSE, warning = FALSE }
- # RSKC: Perform Robust & Sparse K-means Clustering ####
- # Select number of clusters
- clust.num <- k.clusters
- # Select number of iterations
- runs <- 1
- # Prepare objects to store RSKC results
- rskc.results <- list()
- rskc.labels <- working.ids %>%
- dplyr::select(UniqueObjectNumber, Metadata_SampleID,
- Metadata_SpecimenID, Metadata_Region,
- norm_x, norm_y)
- rskc.weights <- data.frame("Morphometric" = colnames(working.exp))
- # seed is initalised for reproducibility
- set.seed(12345)
- for (i in 1:runs) {
- # Perform RSKC
- rskc.results[[i]] <- RSKC(d = working.exp,
- alpha = 0.1,
- nstart = 1000,
- ncl = clust.num,
- scaling = F,
- L1 = sqrt(ncol(working.exp)))
- # Store cluster assignments for run i
- rskc.labels[i+4] <- rskc.results[[i]]$labels
- colnames(rskc.labels)[i+4] <- paste("Run_", i, sep = "")
- # Store feature weights for run i
- rskc.weights[i+1] <- rskc.results[[i]]$weights
- colnames(rskc.weights)[i+1] <- paste("Run_", i, sep = "")
- # Uncomment to track progress
- print(paste(i,"/", runs, sep = ""))
- }
- ## ---------------------------------------- ##
- ##### Melt the RSKC Dataframe for ggplot #####
- ## ---------------------------------------- ##
- # add the relevant metadata columns to the rskc.labels dataframe
- rskc.labels$norm_x<- working.ids$norm_x
- rskc.labels$norm_y <- working.ids$norm_y
- rskc.labels$Metadata_SpecimenID <- working.ids$Metadata_SpecimenID
- rskc.labels$Metadata_Region <- working.ids$Metadata_Region
- # view(rskc.labels) # uncomment to inspect the labels
- # export the RSKC labels as a csv
- # write.csv(rskc.labels,
- # file = file.path(export.path,
- # paste0(clust.num, "Pvalb_V1_S1_RSKC_labels.csv")),
- # row.names = FALSE)
- ```
- this chunk prepares the summary dataframe of RSKC feature weights.
- ```{r prepare.summary.weights}
- # Calculate the mean and standard deviation for each feature
- rskc.weights.summary <- rskc.weights %>%
- gather(run, weight, -Morphometric) %>%
- group_by(Morphometric) %>%
- summarise(
- mean = mean(weight, na.rm = TRUE),
- sd = sd(weight, na.rm = TRUE),
- se = sd / sqrt(n())
- ) %>%
- arrange(desc(mean))
- ```
- ```{r include = FALSE }
- # this chunk produces the first version of the RSKC weights plot -features are ordered by decreasing weight
- # a correctly formatted version of the plot is included in a separate markdown for phenotyping and denSNE plotting .
- # using the feature weights outpur for each trial, take the average weight for each feature and plot that in a histogram.
- # add error bars to the plot - standard error
- # annotate the bars with the average feature weight rounded to five decimal places (legibility and ease of reference)
- ## ----------------------------------------- ##
- ##### calculate average weights for each feature ####
- ## ----------------------------------------- ##
- # Calculate the mean and standard deviation for each feature
- rskc.weights.summary <- rskc.weights %>%
- gather(run, weight, -Morphometric) %>%
- group_by(Morphometric) %>%
- summarise(
- mean = mean(weight, na.rm = TRUE),
- sd = sd(weight, na.rm = TRUE),
- se = sd / sqrt(n())
- ) %>%
- arrange(desc(mean)) # Order by mean in decreasing order
- # remove the "AreaShape_" prefix from all the morphometric names in the dataframe
- rskc.weights.summary$Morphometric <- gsub("AreaShape_", "", rskc.weights.summary$Morphometric)
- # Change the order of Morphometric factor levels to match the order of the means
- rskc.weights.summary$Morphometric <- factor(rskc.weights.summary$Morphometric, levels = rskc.weights.summary$Morphometric)
- ## -------------------------- ##
- ##### plotting the weights #####
- ## -------------------------- ##
- # Plot the weights with error bars
- weights.plot <- ggplot(rskc.weights.summary, aes(x = Morphometric, y = mean)) +
- geom_bar(stat = "identity", fill = "lightsteelblue2", colour = "black") +
- geom_errorbar(aes(ymin = mean - se, ymax = mean + se), width = 0.2) +
- geom_text(aes(y = mean/2, label = round(mean, 5)),
- angle = 90, family = "Helvetica", size = 2)+
- theme_classic() +
- theme(axis.text.x = element_text(angle = 90, hjust = 1,
- vjust = 0.5, family = "Helvetica", size = 7),
- axis.text.y = element_text(family = "Helvetica", size = 7),
- axis.title.x = element_text(family = "Helvetica", size = 16),
- axis.title.y = element_text(family = "Helvetica", size = 16)) +
- #labs (x = "", y = "")
- labs(x = "Morphometric", y = "RSKC Feature Weight")
- # Print the plot
- print(weights.plot)
- ## -------------------------- ##
- ##### export the data and plot #####
- ## -------------------------- ##
- # # Export the weights for 100 runs
- # write.csv(rskc.weights, file = paste0(export.path, "/RSKC_Weights.csv"), row.names = FALSE)
- #
- # # Export the summary dataframe
- #write.csv(rskc.weights.summary, file = paste0(export.path, "/Summary_RSKC_Weights.csv"), row.names = FALSE)
- # Export the barplot with error bars
- #ggsave(filename = file.path(export.path, "Sep23_RSKC_WeightsPlot.pdf"), plot = weights.plot, height = 7, width = 10)
- ```
- # Multiplying feature weights to the z-scored features
- prepare morphology data for cluster visualisation as described in Balsor et al., 2021.
- This code prepares a weighted and ordered feature matrix for downstream t-SNE visualization using the RSKC feature weights. It reorders the z-scored morphometric features based on their average weights and multiplies each column by its corresponding mean weight to emphasize important features. Cluster assignments from RSKC are then added to the weighted dataframe, creating a combined dataset with both metadata and weighted morphometrics. Finally, both the weighted and unweighted z-scored dataframes with cluster labels are stored for analysis and visualization.
- | Object | Description |
- | ---------------------------------- | ------------------------------------------- |
- | `working.exp` | dataframe of z-scored morphometric features |
- | `working.exp.ordered` | features reordered by importance |
- | `weighted.working.exp.ordered` | weighted features for t-SNE |
- | `order` | morphometric feature ordering vector |
- | `CellProfiler.weighted.scored.df` | combined metadata and weighted features |
- | `CellProfiler.scored.clustered.df` | combined metadata and unweighted features |
- | `rownames(rskc.weights.summary)` | morphometric names assigned to rows |
- | `working.ids$cluster` | cluster assignments added to metadata |
- ```{r weighted.working.exp, message = FALSE, warning = FALSE }
- # remove the AreaShape_ prefix from the working.exp column names
- colnames(working.exp) <- sub("^AreaShape_", "", colnames(working.exp))
- # Remove "AreaShape_" prefix from Morphometric column
- rskc.weights.summary$Morphometric <-
- sub("^AreaShape_", "", rskc.weights.summary$Morphometric)
- ## ----------------------------------------- ##
- ##### weighted z-scored dataframe for tsne ####
- ## ----------------------------------------- ##
- # the column containing row names is the morphometric column
- rownames(rskc.weights.summary) <- rskc.weights.summary$Morphometric
- # convert to dataframe
- working.exp <- as.data.frame(working.exp)
- # obtain the order of rows - decreasing weight order of features
- order <- rskc.weights.summary$Morphometric
- # Reorder the columns of working.exp
- working.exp.ordered <- working.exp %>% select(order)
- # view(working.exp.ordered) # inspection step
- # sweep function to multiply each column of working.exp
- # by the corresponding average weight
- # MARGIN = 2 indicates the operation is applied to the columns
- weighted.working.exp.ordered <- sweep(working.exp.ordered,
- MARGIN = 2,
- STATS = rskc.weights.summary$mean,
- `*` )
- # view(weighted.working.exp.ordered) # inspection step
- # merge the cluster assignments to the weighted.exp.labels
- weighted.working.exp.ordered <- as.data.frame(weighted.working.exp.ordered)
- working.ids$cluster <-rskc.labels$Run_1
- CellProfiler.weighted.scored.df <- cbind(working.ids,
- weighted.working.exp.ordered )
- # inspection steps - uncomment to view the dataframes :
- # view(CellProfiler.weighted.scored.df)
- # print (head(CellProfiler.weighted.scored.df))
- # append the cluster labels to the z-scored dataframe
- CellProfiler.scored.clustered.df <- cbind(working.ids,working.exp.ordered)
- ## -- Inspection Steps : view the dataframes -- ##
- # view(CellProfiler.scored.clustered.df)
- # view(CellProfiler.weighted.scored.df)
- # print (head(CellProfiler.scored.clustered.df))
- # print (head(CellProfiler.weighted.scored.df))
- # print (colnames(CellProfiler.scored.clustered.df))
- ## -- Export the clustered z-scored dataframe -- ##
- # write.csv(CellProfiler.scored.clustered.df,
- # file= paste0(export.path,
- # "/Scored_Clustered_Pvalb_30Jul2024.csv"))
- ```
- # running densne
- perplexity term : 10
- random seed term :13579
- visualisation of clusters assigned from RSKC - to be used as a visualisation - the weights from RSKC smooth out the spacing and arrangement of clusters when they are mapped down from High D to 2D as a result of the densne algorithm.
- cluster assignment is currently using the clusters assigned from the first run of RSKC but average weights over the 100 trials are used in multiplying weights to feature z-scores
- This code was run on a MacBookPro M3 Pro Apple Silicon Chip, 18GB Unified Memory
- Approximate Runtime : ~ 49.3 seconds
- ```{r rskc.cluster.densne, fig.align="center", fig.height=5, fig.width=5, out.width="100%", message = FALSE}
- ## ----------------------------------------------- ##
- ##### Densne for a specific seed and perplexity #####
- ## ----------------------------------------------- ##
- # set the colour palette
- color.palette.vector.dense = c("darkgoldenrod2", "lightcoral", "cadetblue1",
- "olivedrab1", "pink1", "maroon3",
- "chartreuse3", "turquoise3", "navyblue",
- "slateblue3","plum", "deepskyblue", "khaki2")
- # Set your specific seed and perplexity
- # seed for initialisation
- # perplexity - influences local/global structure preservation
- seed <- 13579
- perplexity <-10
- # Set the seed for reproducibility
- set.seed(seed)
- # Run the density-preserving t-SNE algorithm
- dense_tsne <- densne(weighted.working.exp.ordered,
- perplexity = perplexity,
- check_duplicates = FALSE)
- # denSNE coordiantes are saved as a dataframe
- dense_tsne.df <- as.data.frame(dense_tsne)
- # Combine with the z-scored dataframe
- dense_tsne.df <- cbind(CellProfiler.weighted.scored.df, dense_tsne.df)
- # set the cluster column to factor - discrete labels not continuous numbers
- dense_tsne.df$RSKC_clusters <-
- factor(CellProfiler.weighted.scored.df$cluster)
- dense_tsne.df$Metadata_SpecimenID <-
- factor(CellProfiler.weighted.scored.df$Metadata_SpecimenID)
- # Add the seed and perplexity as columns
- dense_tsne.df$seed <- seed
- dense_tsne.df$perplexity <- perplexity
- # add the cluster label mapping for median sizes to the dataframe
- # 1. Calculate the median size for each cluster
- cluster_medians <- dense_tsne.df %>%
- group_by(RSKC_clusters) %>%
- summarise(median_size = median(Area, na.rm = TRUE)) %>%
- arrange(median_size) # ascending order: smallest to largest
- # 2. Create a mapping vector: cluster ID -> size rank
- median.size.order.mapping <- setNames(
- seq_len(nrow(cluster_medians)), # ranks 1, 2, ..., n
- cluster_medians$RSKC_clusters # names are cluster IDs
- )
- # 3. Apply this mapping to your dataframe
- dense_tsne.df$RSKC_size_clusters <- factor(
- median.size.order.mapping[as.character(dense_tsne.df$RSKC_clusters)],
- levels = seq_len(nrow(cluster_medians)) # ensures order is 1..n
- )
- # 4. Optional: inspect the mapping
- median.size.order.mapping
- # set the RSKC_clusters column as a factor
- dense_tsne.df$RSKC_clusters <- factor(dense_tsne.df$RSKC_clusters)
- # find the cluster centroids in the Dense space
- centroids <- aggregate(cbind(V1, V2) ~ RSKC_clusters, dense_tsne.df, mean)
- # Plotting Dense t-SNE plot
- p <- ggplot(dense_tsne.df, aes(x = V1, y = V2), size = 1.25) +
- geom_point(aes(color = as.factor(RSKC_size_clusters))) +
- scale_color_manual(values = color.palette.vector.dense) +
- theme_void() +
- theme(legend.position = "right")
- #theme (
- #legend.position = "none",
- #axis.title = element_text(family = "Helvetica", size = 100),
- #axis.text = element_text (family = "Helvetica", size = 80)
- #)
- # plot the centroids
- p <- p + geom_point(data = centroids,
- aes(x = V1, y = V2),
- color = 'black',
- size = 5,
- shape = 4)
- print(p)
- ## -- export the plot as a pdf -- ##
- # ggsave(paste0(export.path,
- # "/2OrderedSizeCluster_perp10_seed13579_denSNE_PvalbV1S1.pdf"),
- # plot = p, width = 10, height = 10)
- ## -- write the dense t-sne dataframe to csv -- ##
- # write.csv(dense_tsne.df, paste0(export.path,
- # "/denSNE_weighted_zscored_size_clusters_PvalbV1S1.csv"), row.names = FALSE)
- ```
Figure3_Markdown.Rmd, no license · at the source
Overview
- McMaster Neuroscience Graduate Program, McMaster University, Hamilton, ON, Canada
- Department of Psychology, Neuroscience and Behaviour, McMaster University, Hamilton, ON, Canada
Abstract
Parvalbumin-positive (PV+) inhibitory interneurons are central components of experience-dependent plasticity in the visual cortex (V1). Anatomically, they are distributed across cortical layers and are traditionally classified as basket, chandelier, or bipolar cells based on dendritic and axonal arborization patterns. In parallel, physiological, transcriptomic, connectomic, and multimodal studies have revealed substantial diversity within PV+ populations, raising the question of how this diversity is organized within cortical circuits. In contrast, light microscopy studies based on soma labeling have primarily quantified PV+ cells using size and density, and the potential of soma morphology to capture this diverse organization remains unclear. To address this, we developed a high-throughput, data-driven approach to quantify PV+ soma morphology in > 14,000 cells from mouse V1 and somatosensory cortex (S1). Using 97 morphological features combined with clustering, phenotyping, and laminar mapping, we identified structured diversity in PV+ somas. PV+ cells were organized into 13 morphological clusters along partially independent gradients of size and shape. Phenotyping identified four size and five shape categories that describe PV+ cell diversity across cortical areas. Mapping these categories onto cortical layers revealed a structured organization in which specific morphologies are enriched within distinct laminar compartments. This organization aligns with cortical architecture and suggests that PV+ interneuron morphology is systematically related to circuit structure. These findings demonstrate that substantial morphological information can be extracted from standard PV+ labeling approaches using quantitative analysis. Together, this work provides a high-dimensional atlas of PV+ interneuron soma morphology in mouse V1 and S1 and establishes a framework for linking cellular anatomy to circuit organization and experience-dependent plasticity.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.
OSF 8f2nr
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
6 files
- rakd7/
Figure2_Markdown.Rmd , R, 440 lines, 2 matches - rakd7/
Figure3_Markdown.Rmd , R, 773 lines, 5 matches - rakd7/
Figures4-5_Markdown.Rmd , R, 1,336 lines, 2 matches - rakd7/
revised_Figure6_Heatmap_ , R, 917 lines, 2 matchesMarkdown.Rmd - rakd7/
revised_Figure7_Markdown , R, 620 lines, 2 matches.Rmd - rakd7/
revised_Figure8_Markdown , R, 700 lines, 2 matches.Rmd
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 6 scripts, each with its path and the digest of its content;
- 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability statement
The datasets and code presented in this study can be found in the online repository 10.17605/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 4 authors, 6 keywords, 1 funder, 53 references.
Cite
This paper
Panday, M., Monteiro, L., Daudi, A., & Murphy, K. M. (2026). A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex. Frontiers in neuroscience, 20, 1848222. https://
BibTeX
@article{panday2026high,
author = {Panday, Maheshwar and Monteiro, Leanne and Daudi, Ahad and Murphy, Kathryn M.},
title = {{A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex}},
journal = {Frontiers in neuroscience},
year = {2026},
month = jun,
volume = {20},
pages = {1848222},
publisher = {Frontiers Media SA},
issn = {1662-4548},
doi = {10.3389/
url = {https://
pmid = {42359344},
pmcid = {PMC13290952}
}
RIS
TY - JOUR
AU - Panday, Maheshwar
AU - Monteiro, Leanne
AU - Daudi, Ahad
AU - Murphy, Kathryn M.
TI - A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex
T2 - Frontiers in neuroscience
J2 - Front Neurosci
PY - 2026
DA - 2026/
VL - 20
SP - 1848222
SN - 1662-4548
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "A high-dimensional atlas of parvalbumin interneuron soma morphology in mouse visual and somatosensory cortex",
"container-title": "Frontiers in neuroscience",
"author": [
{
"family": "Panday",
"given": "Maheshwar"
},
{
"family": "Monteiro",
"given": "Leanne"
},
{
"family": "Daudi",
"given": "Ahad"
},
{
"family": "Murphy",
"given": "Kathryn M."
}
],
"container-title-short":
"volume": "20",
"page": "1848222",
"DOI": "10.3389/
"PMID": "42359344",
"PMCID": "PMC13290952",
"ISSN": "1662-4548",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: circlize, ComplexHeatmap, pheatmap, 4 other tools, mouse, 4 references
- [2] doi:10.1038/s41467-026-74753-y [code]
- A human-specific microRNA controls the timing of excitatory synaptogenesis.Journal: Nature communicationsIn common: easystats, circlize, ComplexHeatmap, 6 other tools
- [3] doi:10.1016/j.celrep.2026.117500 [code]
- Spatio-molecular gene expression reflects dorsal anterior cingulate cortex structure and function in the human brain.Journal: Cell reportsIn common: easystats, circlize, ComplexHeatmap, 6 other tools
- [4] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: circlize, ComplexHeatmap, pheatmap, 4 other tools, mouse, 2 references
- [5] doi:10.1038/s44318-026-00806-z [code]
- Interspecific diversity in the neuronal composition of the mammalian cortex arises from heterochrony in neurogenesis.Journal: The EMBO journalIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, 1 reference
- [6] doi:10.1038/s41586-026-10214-2 [code]
- Multidimensional profiling of heterogeneity in supratentorial ependymomas.Journal: NatureIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse, 1 reference
- [7] doi:10.1038/s41593-026-02367-0 [code]
- A reproducible three-dimensional model of human brain tissue to investigate physiological and disease-associated microglia phenotypes.Journal: Nature neuroscienceIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, 1 reference
- [8] doi:10.1002/imt2.70163 [code]
- Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.Journal: iMetaIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse
- [9] doi:10.1038/s41467-026-76341-6 [code]
- Neonatal inflammation disrupts a temporally restricted postnatal Numb-enriched microglial state in mice.Journal: Nature communicationsIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse
- [10] doi:10.7717/peerj.21426 [code]
- Integrated transcriptomic identification and validation reveal key autophagy-associated biomarkers in sleep deprivation.Journal: PeerJIn common: circlize, ComplexHeatmap, pheatmap, 5 other tools, mouse
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 6 scripts, and 15 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:c50e18e956a99e4d…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
