A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
The 20 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
- [1] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ Experiments/_Composite_score_generation/Save_Composite_Score_Brain.ipynb, lines 18–34 · score 0.95 · CelloScope, SpatialDecon, STDeconvolve, Cell DART, SpiceMix, DestVI
- [2] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ Experiments/_Deconvolution_Metrics_Calculation/Evaluation_Brain_Category.ipynb, lines 29–43 · score 0.95 · CelloScope, SpatialDecon, STDeconvolve, Cell DART, SpiceMix, DestVI
- [3] § Results › Benchmarking of spatial domain detection methods ↔ Experiments/_Domain_Detection_Metrics_Calculation/all_datasets_metric_computation.ipynb, lines 185–194 · score 0.93 · DR SC, DeepST, SpaSRL, SpatialPCA, BayesCafe, BayesSpace
- [4] § Results › Benchmarking of spatial domain detection methods ↔ Experiments/_Domain_Detection_Metrics_Calculation/Metrics.py, lines 351–421 · score 0.93 · DR SC, DeepST, SpaSRL, SpatialPCA, BayesCafe, BayesSpace
- [5] § Results › Benchmarking of spatial domain detection methods ↔ Experiments/_Domain_Detection_Metrics_Calculation/all_datasets_metric_computation.ipynb, lines 289–313 · score 0.91 · mouse breast cancer, osmFISH, chicken heart, prostate cancer, liver cancer, kidney cancer
- [6] § Results › Benchmarking of spatial domain detection methods ↔ Experiments/_Domain_Detection_Metrics_Calculation/Metrics.py, lines 944–1089 · score 0.86 · SpaSRL, SpatialPCA, BayesCafe, BayesSpace, SpaceFlow, GraphST
- [7] § Methods › Overview of SynthST ↔ SynthST/Simulator_CTP/SynthST/model.py, lines 56–99 · score 0.86 · cross entropy, weight decay, adjacency matrix, proportion matrix, L2, encoder
- [8] § Results › Benchmarking of spatial domain detection methods ↔ Experiments/_Domain_Detection_Metrics_Calculation/Metrics.py, lines 351–421 · score 0.84 · DeepST, BayesSpace, SpaceFlow, prostate cancer, GraphST, PRECAST
- [9] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ Experiments/_Domain_Detection_Metrics_Calculation/Metrics.py, lines 107–248 · score 0.78 · chicken heart, prostate cancer, liver cancer, kidney cancer, breast cancer, ground truth
- [10] § Methods › Evaluation metrics for domain detection ↔ Experiments/_Domain_Detection_Metrics_Calculation/Metrics.py, lines 423–497 · score 0.78 · Normalized Mutual, Adjusted Rand, Completeness Score, Homogeneity, ground truth, ARI
- [11] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ Experiments/_Composite_score_generation/plotting_files.py, lines 1–39 · score 0.72 · spatial Pearson, Spatial Metrics, composite score, Rare Cell, Cosine, Lee
- [12] § Methods › Overview of SynthST ↔ SynthST/Simulator_CTP/SynthST/SynthST.py, the whole file · a weak match · score 0.68 · weight decay, SynthST, GAT, optimization, inferred, loss
- [13] § Results › Cell2location, RCTD and SONAR achieve the highest accuracy in cell-type deconvolution ↔ Experiments/_Composite_score_generation/Save_Composite_Score_Brain.ipynb, lines 18–34 · score 0.67 · SpatialDWLS, DestVI, Cell2location, Redeconve, Composite Score, SSIM
- [14] § Results › Cell2location, RCTD and SONAR achieve the highest accuracy in cell-type deconvolution ↔ Experiments/_Deconvolution_Metrics_Calculation/Evaluation_Brain_Category.ipynb, lines 29–43 · score 0.66 · SpatialDWLS, DestVI, Cell2location, Redeconve, SSIM, Geary
- [15] § Extended Data ↔ SynthST/Simulator_CTP/SynthST/model.py, lines 56–99 · score 0.64 · cross entropy, adjacency matrix, proportion matrix, reconstructs, decoder, loss
- [16] § Methods › Overview of SynthST ↔ SynthST/Simulator_gene_expression/Simulate_gene_expression_notebook.ipynb, lines 58–86 · score 0.62 · SynthST, gene expression, Cell2location, signature, abundance, simulated
- [17] § Methods › Composite score metrics for cell-type deconvolution ↔ Experiments/_Composite_score_generation/plotting_files.py, lines 58–89 · score 0.58 · Spatial Score, Composite Rare, Composite Score, SSIM, Correlation, Moran
- [18] § Methods › Composite score metrics for cell-type deconvolution ↔ Experiments/_Deconvolution_Metrics_Calculation/metrics.py, lines 261–279 · score 0.55 · Recall curve, Precision, threshold, ground truth, AUPR, metric
- [19] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ SynthST/Simulator_gene_expression/Simulate_gene_expression_notebook.ipynb, lines 58–86 · score 0.52 · simulated gene expression, SynthST, signature, single cell, matrix, spots
- [20] § Results › Overview of spatial cell-type deconvolution benchmarking ↔ benchmarking_domain_detection/CCST/run_CCST.py, the whole file · a weak match · score 0.52 · single cell resolution, deep graph, gene expression, MERFISH, training, ground truth
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 1,193 lines · 47 KB · MIT · 5 matches
- # ----------------------- Imports -----------------------
- import os
- import re
- import numpy as np
- import pandas as pd
- import scanpy as sc
- import squidpy as sq
- import anndata as an
- import seaborn as sns
- import matplotlib.pyplot as plt
- from PIL import Image
- from matplotlib import font_manager as fm
- from scipy.spatial import distance_matrix
- from scipy.spatial.distance import squareform, pdist
- from sklearn.metrics import (
- adjusted_rand_score,
- normalized_mutual_info_score,
- silhouette_score,
- homogeneity_score,
- completeness_score
- )
- from sklearn.preprocessing import StandardScaler
- # ----------------------- Utility -----------------------
- def fx_1NN(i,location_in):
- location_in = np.array(location_in)
- dist_array = distance_matrix(location_in[i,:][None,:],location_in)[0,:]
- dist_array[i] = np.inf
- return np.min(dist_array)
- def fx_kNN(i,location_in,k,cluster_in):
- location_in = np.array(location_in)
- cluster_in = np.array(cluster_in)
- dist_array = distance_matrix(location_in[i,:][None,:],location_in)[0,:]
- dist_array[i] = np.inf
- ind = np.argsort(dist_array)[:k]
- cluster_use = np.array(cluster_in)
- if np.sum(cluster_use[ind]!=cluster_in[i])>(k/2):
- return 0
- else:
- return 1
- def _compute_CHAOS(clusterlabel, location):
- clusterlabel = np.array(clusterlabel)
- location = np.array(location)
- matched_location = StandardScaler().fit_transform(location)
- clusterlabel_unique = np.unique(clusterlabel)
- dist_val = np.zeros(len(clusterlabel_unique))
- count = 0
- for k in clusterlabel_unique:
- location_cluster = matched_location[clusterlabel==k,:]
- if len(location_cluster)<=2:
- continue
- n_location_cluster = len(location_cluster)
- results = [fx_1NN(i,location_cluster) for i in range(n_location_cluster)]
- dist_val[count] = np.sum(results)
- count = count + 1
- chaos = np.sum(dist_val)/len(clusterlabel)
- return np.exp(-0.5*chaos)
- def _compute_PAS(clusterlabel,location):
- clusterlabel = np.array(clusterlabel)
- location = np.array(location)
- matched_location = location
- results = [fx_kNN(i,matched_location,k=10,cluster_in=clusterlabel) for i in range(matched_location.shape[0])]
- return np.sum(results)/len(clusterlabel)
- def compute_ASW(pred,spatial_coords):
- distance_matrix = squareform(pdist(spatial_coords))
- sil = silhouette_score(X=distance_matrix, labels=pred, metric='precomputed')
- return (sil + 1)/2
- def LISI(coords, meta, label, perplexity=30, nn_eps=0):
- import rpy2.robjects as robjects
- from rpy2.robjects import pandas2ri
- pandas2ri.activate()
- from rpy2.robjects.packages import importr
- importr("lisi")
- if not isinstance(coords, pd.DataFrame):
- coords = pd.DataFrame(coords)
- if not isinstance(meta, pd.DataFrame):
- meta = pd.DataFrame(meta)
- meta = meta.loc[:, [label]]
- meta[label] = meta[label].astype(str)
- coords = robjects.conversion.py2rpy(coords)
- meta = robjects.conversion.py2rpy(meta)
- as_matrix = robjects.r["as.matrix"]
- lisi = robjects.r["compute_lisi"](as_matrix(coords), meta, label, perplexity, nn_eps)
- if isinstance(lisi, pd.DataFrame):
- lisi = lisi.values
- elif isinstance(lisi, np.recarray):
- lisi = [item[0] for item in lisi]
- return lisi
- # ----------------------- Import functions -----------------------
- def import_dataset(dataset_name, paths, mode = "evaluation"):
- """
- Load dataset given its name and a dictionary of file paths.
- Parameters
- ----------
- dataset_name : str
- Name of the dataset (used for branching logic).
- paths : dict
- Dictionary of file paths needed for that dataset.
- Must contain at least 'st_path'.
- Can contain 'gnd_path' if ground truth is from a CSV/TSV file.
- Returns
- -------
- gnd : pd.Series or pd.DataFrame
- Ground truth clusters.
- locs : pd.DataFrame
- Spatial coordinates.
- """
- # Load main AnnData object
- ST = an.read_h5ad(paths["st_path"])
- if dataset_name.startswith("DLPFC"):
- gnd = pd.read_csv(paths["gnd_path"], sep="\t")['layer_guess_reordered']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.array_row,
- "array_col": ST.obs.array_col
- })
- elif dataset_name.startswith("embryo"):
- gnd = pd.DataFrame(ST.obs['annotation']).rename(columns={"annotation": "Cluster"})
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obsm["spatial"][:, 0],
- "array_col": ST.obsm["spatial"][:, 1]
- }, index=ST.obs_names)
- elif dataset_name == "mouse_breast_cancer":
- gnd = pd.DataFrame(ST.obs['ground_truth']).rename(columns={"annotation": "Cluster"})
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.array_row,
- "array_col": ST.obs.array_col
- })
- elif dataset_name.startswith("MERFISH_brain"):
- gnd = pd.DataFrame(ST.obs['ground_truth']).rename(columns={"annotation": "Cluster"})
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obsm["spatial"][:, 0],
- "array_col": ST.obsm["spatial"][:, 1]
- }, index=ST.obs_names)
- elif dataset_name == "simulated_breast_atlas":
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obsm["spatial"][:, 0],
- "array_col": ST.obsm["spatial"][:, 1]
- }, index=ST.obs_names)
- elif dataset_name == "osmFISH":
- gnd = pd.DataFrame(ST.obs['ground_truth']).rename(columns={"annotation": "Cluster"})
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obsm["spatial"][:, 0],
- "array_col": ST.obsm["spatial"][:, 1]
- }, index=ST.obs_names)
- elif dataset_name.startswith("simulated_kidney_cancer"):
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.new_x,
- "array_col": ST.obs.new_y
- })
- elif dataset_name.startswith("simulated_breast_cancer"):
- ST.obs_names_make_unique()
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.new_x,
- "array_col": ST.obs.new_y
- })
- elif dataset_name.startswith("simulated_liver_cancer"):
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.new_x,
- "array_col": ST.obs.new_y
- })
- elif dataset_name.startswith("simulated_intestine"):
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.new_x,
- "array_col": ST.obs.new_y
- })
- elif dataset_name == "simulated_chicken_heart":
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.array_row,
- "array_col": ST.obs.array_col
- })
- elif dataset_name == "simulated_prostate_cancer":
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names)['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.x_new,
- "array_col": ST.obs.y_new
- })
- elif dataset_name == "simulated_cerebellum":
- gnd = pd.read_csv(paths["gnd_path"]).set_index(ST.obs_names.astype(int))['Ground Truth']
- gnd.index.name = "Index"
- locs = pd.DataFrame({
- "array_row": ST.obs.xcoord,
- "array_col": ST.obs.ycoord
- }).set_index(ST.obs_names.astype(int))
- else:
- raise ValueError(f"Incorrect input: {dataset_name}")
- if mode == "plot":
- locs = pd.DataFrame({
- "array_row": ST.obsm["spatial"][:, 0],
- "array_col": ST.obsm["spatial"][:, 1]
- }).set_index(ST.obs_names)
- if dataset_name == "simulated_cerebellum":
- locs.index = locs.index.astype(int)
- locs.index.name = "Index"
- return gnd, locs
- def load_prediction(dataset_name, method_name, path):
- """
- Load output file given its name and file path.
- Parameters
- ----------
- dataset_name : str
- Name of the dataset (used for branching logic).
- method_name : str
- Name of the method (used for branching logic).
- path : str
- Path to output file
- Returns
- -------
- pred : predicted clusters
- """
- if os.path.exists(path):
- if dataset_name.startswith("simulated_breast_cancer") and method_name in ["BASS","PRECAST", "BayesSpace" , "DR_SC"]:
- pred = pd.read_csv(path)
- pred = pred.set_index(pred.columns[0])
- if method_name == "PRECAST":
- pred = pred['cluster']
- pred.index = pred.index.str.replace(r'_\d+$', '', regex=True)
- elif(method_name == "banksy"):
- pred = pd.read_csv(path)
- pred = pred.set_index(pred.columns[0])
- columns_to_grab = [col for col in pred.columns if col.startswith("clust_M1_lam0.")]
- pred = pred[columns_to_grab]
- elif method_name == "giotto":
- pred = pd.read_csv(path)
- if "leiden_clus" in pred.columns:
- columns_to_grab = ["leiden_clus"]
- pred = pred.set_index(pred.columns[1])
- else:
- columns_to_grab = ["cluster"]
- pred = pred.set_index(pred.columns[0])
- pred = pred[columns_to_grab]
- elif method_name == "PRECAST" or method_name == "BayesCafe":
- pred = pd.read_csv(path)
- pred = pred.set_index(pred.columns[0])
- columns_to_grab = ["cluster"]
- pred = pred[columns_to_grab]
- elif method_name == "SpaceFlow":
- pred = pd.read_csv(path)
- pred = pred.set_index(pred.columns[1])
- columns_to_grab = ["Predicted_cell_label"]
- pred = pred[columns_to_grab]
- else:
- pred = pd.read_csv(path)
- pred = pred.set_index(pred.columns[0])
- else:
- print(f"path doesn't exist : {path}")
- print(f"{method_name} output missing for {dataset_name}")
- pred = None
- return pred
- # ----------------------- Metric Computationn -----------------------
- def compute_metrics(dataset_names, dataset_paths, method_names, pred_paths, error_log_file, output_dir, metric_names="all"):
- """
- Compute metric values given dataset and method names and paths
- Parameters
- ----------
- dataset_names : list
- List of datasets to compute on
- Options for dataset names present at bottom of this file
- dataset_paths: dictionary of dictionaries
- Dictionaries mapping dataset names to dictionaries with st_path and gnd_path needed for importing dataset
- method_names: list
- List of methods to compute on
- Options for method names present at bottom of this file
- pred_paths : dictionary of dictionaries
- Dictionary mapping dataset names and method names to path of output file
- metric_names : list or str
- List of metrics to compute or "all"
- error_log_file : str
- Path to error log file
- output_dir : str
- Path to save final metric compute files to
- Returns
- -------
- pred : predicted clusters
- """
- # Available metrics
- available_metrics = ["ARI", "NMI", "CHAOS", "PAS", "ASW", "HOM", "COM"]
- if metric_names == "all":
- metric_names = available_metrics
- # Initialize pivoted DataFrame for each metric
- method_names_dict = {
- "SCANIT": "SCANIT",
- "CCST": "CCST",
- "DeepST": "DeepST",
- "GraphST": "GraphST",
- "PROST": "PROST",
- "SpaSRL": "SpaSRL",
- "STAGATE": "STAGATE",
- "SpatialPCA": "SpatialPCA",
- "banksy": "Banksy",
- "giotto": "Giotto",
- "DR_SC": "DR.SC",
- "ISC_MEB": "ISC.MEB",
- "BayesSpace": "BayesSpace",
- "PRECAST": "PRECAST",
- "BayesCafe": "BayesCafe",
- "BASS": "BASS",
- "SpaceFlow" : "SpaceFlow",
- "IRIS" : "IRIS"
- }
- dataset_names_dict = {
- "DLPFC151507": "DLPFC 151507",
- "DLPFC151508": "DLPFC 151508",
- "DLPFC151509": "DLPFC 151509",
- "DLPFC151510": "DLPFC 151510",
- "DLPFC151669": "DLPFC 151669",
- "DLPFC151670": "DLPFC 151670",
- "DLPFC151671": "DLPFC 151671",
- "DLPFC151672": "DLPFC 151672",
- "DLPFC151673": "DLPFC 151673",
- "DLPFC151674": "DLPFC 151674",
- "DLPFC151675": "DLPFC 151675",
- "DLPFC151676": "DLPFC 151676",
- "embryo9.5": "Embryo 9.5",
- "embryo14.5": "Embryo 14.5",
- "mouse_breast_cancer": "Mouse Breast Cancer",
- "MERFISH_brain0.04": "MERFISH Brain 0.04",
- "MERFISH_brain0.09": "MERFISH Brain 0.09",
- "MERFISH_brain0.14": "MERFISH Brain 0.14",
- "MERFISH_brain0.19": "MERFISH Brain 0.19",
- "MERFISH_brain0.24": "MERFISH Brain 0.24",
- "osmFISH": "osmFISH",
- "simulated_kidney_cancer410": "Kidney Cancer 410",
- "simulated_kidney_cancer411": "Kidney Cancer 411",
- "simulated_kidney_cancer506": "Kidney Cancer 506",
- "simulated_breast_cancerER+_CID4290": "Breast Cancer ER+ CID 4290",
- "simulated_breast_cancerTNBC_CID44971": "Breast Cancer TNBC CID 44971",
- "simulated_liver_cancerHCC-1L": "Liver Cancer HCC-1L",
- "simulated_liver_cancerHCC-2L": "Liver Cancer HCC-2L",
- "simulated_liver_cancerHCC-3L": "Liver Cancer HCC-3L",
- "simulated_liver_cancerHCC-4L": "Liver Cancer HCC-4L",
- "simulated_breast_atlas": "Breast Atlas",
- "simulated_intestineA1": "Intestine A1",
- "simulated_intestineA2": "Intestine A2",
- "simulated_chicken_heart": "Chicken Heart",
- "simulated_prostate_cancer" : "Prostate Cancer",
- "simulated_cerebellum" : "Cerebellum"
- }
- metric_results = {
- metric: pd.DataFrame(
- index=list(method_names_dict.values()), columns=list(dataset_names_dict.values()), dtype=float
- ) for metric in metric_names
- }
- os.makedirs(output_dir, exist_ok=True)
- print(f"saving to {output_dir}")
- # Process each dataset
- for dataset_name in dataset_names:
- try:
- # Load dataset (ground truth and spatial coordinates)
- gnd, locs = import_dataset(dataset_name , dataset_paths[dataset_name])
- gnd = gnd.dropna()
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Dataset loading error: {dataset_name} - {e}\n")
- continue
- for method_name in method_names:
- try:
- # Load predictions
- print(f"Running on {dataset_name}_{method_name}")
- pred = load_prediction(dataset_name, method_name, pred_paths[dataset_name][method_name])
- # If predictions are missing, skip to the next method
- if pred is None:
- continue
- pred = pred[~pred.index.duplicated(keep='first')]
- # Ensure first columns are indices
- intersect_idx = gnd.index.intersection(pred.index)
- # Filter gnd and pred based on the intersection
- gnd_filtered = gnd.loc[intersect_idx]
- pred_filtered = pred.loc[intersect_idx]
- # Replace pred and gnd for further computation
- gnd_values = gnd_filtered.values.flatten()
- pred_values = pred_filtered.iloc[:, 0].values
- # Placeholder for spatial coordinates
- spatial_coords = locs.loc[intersect_idx].values
- # Compute metrics and populate the pivoted DataFrame
- for metric in metric_names:
- try:
- if metric == "ARI":
- value = adjusted_rand_score(gnd_values, pred_values)
- elif metric == "NMI":
- value = normalized_mutual_info_score(gnd_values, pred_values)
- elif metric == "CHAOS":
- value = _compute_CHAOS(pred_values, spatial_coords)
- elif metric == "PAS":
- value = _compute_PAS(pred_values, spatial_coords)
- elif metric == "ASW":
- value = compute_ASW(pred_values, spatial_coords)
- elif metric == "HOM":
- value = homogeneity_score(gnd_values, pred_values)
- elif metric == "COM":
- value = completeness_score(gnd_values, pred_values)
- else:
- continue
- # Populate the DataFrame
- metric_results[metric].loc[method_names_dict[method_name], dataset_names_dict[dataset_name]] = value
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Metric error: Dataset={dataset_name}, Method={method_name}, Metric={metric} - {e}\n")
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Prediction loading error: Dataset={dataset_name}, Method={method_name} - {e}\n")
- # Save each metric result as a CSV
- for metric, df in metric_results.items():
- try:
- output_file = os.path.join(output_dir, f"{metric}_results.csv")
- df.to_csv(output_file)
- print(f"Saved {metric} results to {output_file}")
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Metric saving error: Metric={metric} - {e}\n")
- def update_metrics(dataset_names, dataset_paths, method_names, pred_paths, error_log_file , output_dir,metric_names="all"):
- """
- Modify existing output files with new metric values given dataset and method names and paths
- Parameters
- ----------
- dataset_names : list
- List of datasets to compute on
- Options for dataset names present at bottom of this file
- dataset_paths: dictionary of dictionaries
- Dictionaries mapping dataset names to dictionaries with st_path and gnd_path needed for importing dataset
- method_names: list
- List of methods to compute on
- Options for method names present at bottom of this file
- pred_paths : dictionary of dictionaries
- Dictionary mapping dataset names and method names to path of output file
- metric_names : list or str
- List of metrics to compute or "all"
- error_log_file : str
- Path to error log file
- output_dir : str
- Path to save final metric compute files to
- Returns
- -------
- pred : predicted clusters
- """
- # Available metrics
- available_metrics = ["ARI", "NMI", "CHAOS", "PAS", "ASW", "HOM", "COM"]
- if metric_names == "all":
- metric_names = available_metrics
- method_names_dict = {
- "SCANIT": "SCANIT",
- "CCST": "CCST",
- "DeepST": "DeepST",
- "GraphST": "GraphST",
- "PROST": "PROST",
- "SpaSRL": "SpaSRL",
- "STAGATE": "STAGATE",
- "SpatialPCA": "SpatialPCA",
- "banksy": "Banksy",
- "giotto": "Giotto",
- "DR_SC": "DR.SC",
- "ISC_MEB": "ISC.MEB",
- "BayesSpace": "BayesSpace",
- "PRECAST": "PRECAST",
- "BayesCafe": "BayesCafe",
- "BASS": "BASS",
- "SpaceFlow":"SpaceFlow",
- "IRIS":"IRIS"
- }
- dataset_names_dict = {
- "DLPFC151507": "DLPFC 151507",
- "DLPFC151508": "DLPFC 151508",
- "DLPFC151509": "DLPFC 151509",
- "DLPFC151510": "DLPFC 151510",
- "DLPFC151669": "DLPFC 151669",
- "DLPFC151670": "DLPFC 151670",
- "DLPFC151671": "DLPFC 151671",
- "DLPFC151672": "DLPFC 151672",
- "DLPFC151673": "DLPFC 151673",
- "DLPFC151674": "DLPFC 151674",
- "DLPFC151675": "DLPFC 151675",
- "DLPFC151676": "DLPFC 151676",
- "embryo9.5": "Embryo 9.5",
- "embryo14.5": "Embryo 14.5",
- "mouse_breast_cancer": "Mouse Breast Cancer",
- "MERFISH_brain0.04": "MERFISH Brain 0.04",
- "MERFISH_brain0.09": "MERFISH Brain 0.09",
- "MERFISH_brain0.14": "MERFISH Brain 0.14",
- "MERFISH_brain0.19": "MERFISH Brain 0.19",
- "MERFISH_brain0.24": "MERFISH Brain 0.24",
- "osmFISH": "osmFISH",
- "simulated_kidney_cancer410": "Kidney Cancer 410",
- "simulated_kidney_cancer411": "Kidney Cancer 411",
- "simulated_kidney_cancer506": "Kidney Cancer 506",
- "simulated_breast_cancerER+_CID4290": "Breast Cancer ER+ CID 4290",
- "simulated_breast_cancerTNBC_CID44971": "Breast Cancer TNBC CID 44971",
- "simulated_liver_cancerHCC-1L": "Liver Cancer HCC-1L",
- "simulated_liver_cancerHCC-2L": "Liver Cancer HCC-2L",
- "simulated_liver_cancerHCC-3L": "Liver Cancer HCC-3L",
- "simulated_liver_cancerHCC-4L": "Liver Cancer HCC-4L",
- "simulated_breast_atlas": "Breast Atlas",
- "simulated_intestineA1": "Intestine A1",
- "simulated_intestineA2": "Intestine A2",
- "simulated_chicken_heart": "Chicken Heart",
- "simulated_prostate_cancer" : "Prostate Cancer",
- "simulated_cerebellum" : "Cerebellum"
- }
- # Initialize pivoted DataFrame for each metric
- metric_results = {}
- for metric in metric_names:
- path = f"{output_dir}/{metric}_results.csv"
- try:
- df = pd.read_csv(path,skip_blank_lines=True)
- df.set_index(df.columns[0], inplace=True, drop=True) # Set the first column as index
- df.index.name = None # Optional: remove index name
- metric_results[metric] = df
- except FileNotFoundError:
- # Create a new DataFrame if not already present
- metric_results[metric] = pd.DataFrame(index=method_names_dict.values(), columns=dataset_names_dict.values())
- print(f"Created new DataFrame for metric {metric}")
- os.makedirs(output_dir, exist_ok=True)
- # Process each dataset
- for dataset_name in dataset_names:
- try:
- # Load dataset (ground truth and spatial coordinates)
- gnd, locs = import_dataset(dataset_name , dataset_paths[dataset_name])
- gnd = gnd.dropna()
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Dataset loading error: {dataset_name} - {e}\n")
- continue
- for method_name in method_names:
- try:
- method_display_name = method_names_dict[method_name]
- # Load predictions
- print(f"Running on {dataset_name}_{method_name}")
- pred = load_prediction(dataset_name, method_name , pred_paths[dataset_name][method_name])
- # If predictions are missing, skip to the next method
- if pred is None:
- continue
- pred = pred[~pred.index.duplicated(keep='first')]
- # Ensure first columns are indices
- intersect_idx = gnd.index.intersection(pred.index)
- # Filter gnd and pred based on the intersection
- gnd_filtered = gnd.loc[intersect_idx]
- pred_filtered = pred.loc[intersect_idx]
- # Replace pred and gnd for further computation
- gnd_values = gnd_filtered.values.flatten()
- pred_values = pred_filtered.iloc[:,0].values
- # Placeholder for spatial coordinates
- spatial_coords = locs.loc[intersect_idx].values
- # Compute metrics and populate the pivoted DataFrame
- for metric in metric_names:
- try:
- if metric == "ARI":
- value = adjusted_rand_score(gnd_values, pred_values)
- elif metric == "NMI":
- value = normalized_mutual_info_score(gnd_values, pred_values)
- elif metric == "CHAOS":
- value = _compute_CHAOS(pred_values, spatial_coords)
- elif metric == "PAS":
- value = _compute_PAS(pred_values, spatial_coords)
- elif metric == "ASW":
- value = compute_ASW(pred_values, spatial_coords)
- elif metric == "HOM":
- value = homogeneity_score(gnd_values, pred_values)
- elif metric == "COM":
- value = completeness_score(gnd_values, pred_values)
- else:
- continue
- # Populate the DataFrame
- if method_display_name not in metric_results[metric].index:
- metric_results[metric].loc[method_display_name] = np.nan
- metric_results[metric].loc[method_names_dict[method_name], dataset_names_dict[dataset_name]] = value
- print(f"set value of {metric} to {value}")
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Metric error: Dataset={dataset_name}, Method={method_name}, Metric={metric} - {e}\n")
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Prediction loading error: Dataset={dataset_name}, Method={method_name} - {e}\n")
- # Save each metric result as a CSV
- for metric, df in metric_results.items():
- try:
- output_file = os.path.join(output_dir, f"{metric}_results.csv")
- df = df.iloc[:len(method_names_dict)] # Dynamic cutoff (good)
- df.to_csv(output_file)
- print(f"Saved {metric} results to {output_file}")
- except Exception as e:
- with open(error_log_file, "a") as log:
- log.write(f"Metric saving error: Metric={metric} - {e}\n")
- # ----------------------- Visualization Functions -----------------------
- def plot_stitched_heatmaps(metric_files, dataset_names, dataset_type, output_file):
- """
- Produce stitched heatmaps for all metrics
- Parameters
- ----------
- metric_files : dictionary
- Mapping of metric names to output files
- dataset_names: list
- Names of datasets to include in heatmap
- dataset_type: str
- Represents name of selected group of datasets (for example : simulated/real/DLPFC)
- Used only for naming
- output_file : str
- Path to save output image to
- """
- num_metrics = len(metric_files) + 1
- individual_plot_height = 8 # Height of each individual heatmap
- total_height = individual_plot_height * num_metrics # Total height for all metrics
- fig, axes = plt.subplots(4, 2, figsize=(10, total_height)) # Adjust figure size dynamically
- axes = axes.flatten()
- for ax, (metric, file_path) in zip(axes, metric_files.items()):
- try:
- data = pd.read_csv(file_path, index_col=0)
- data_subset = data[dataset_names]
- sns.heatmap(
- data_subset,
- cmap="coolwarm",
- cbar_kws={'label': metric},
- annot=True, # Add annotations
- fmt=".2f", # Format for annotations
- annot_kws={"size": 6 ,"weight": "bold" }, # Small font size for annotations
- ax=ax
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - {metric}" , fontweight = "bold")
- ax.set_xlabel("Datasets" , fontweight = "bold")
- ax.set_ylabel("Methods" , fontweight = "bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- except Exception as e:
- ax.set_visible(False) # Hide axes if there's an error
- print(f"Error processing metric {metric}: {e}")
- try:
- metrics_to_average = ["NMI", "ARI", "ASW", "HOM", "COM"]
- metric_dfs = []
- for metric in metrics_to_average:
- if metric in metric_files:
- data = pd.read_csv(metric_files[metric], index_col=0)
- data_subset = data[dataset_names]
- metric_dfs.append(data_subset)
- # Compute composite score
- composite_data = sum(metric_dfs) / len(metrics_to_average)
- sns.heatmap(
- composite_data,
- cmap="coolwarm",
- cbar_kws={'label': "composite"},
- annot=True, # Add annotations
- fmt=".2f", # Format for annotations
- annot_kws={"size": 6 ,"weight": "bold" }, # Small font size for annotations
- ax=axes[-1]
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - Composite" , fontweight = "bold")
- ax.set_xlabel("Datasets" , fontweight = "bold")
- ax.set_ylabel("Methods" , fontweight = "bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- except Exception as e:
- axes[-1].set_visible(False) # Hide composite score plot if there's an error
- print(f"Error computing composite score: {e}")
- plt.tight_layout() # Ensure spacing between plots
- plt.savefig(output_file, bbox_inches="tight", dpi = 200)
- plt.close()
- def plot_individual_heatmaps(metric_files, dataset_names, dataset_type, output_dir):
- """
- Produce individual heatmaps for metrics of choice
- Parameters
- ----------
- metric_files : dictionary
- Mapping of metric names to output files
- dataset_names: list
- Names of datasets to include in heatmap
- dataset_type: str
- Represents name of selected group of datasets (for example : simulated/real/DLPFC)
- Used only for naming
- output_file : str
- Path to save output image to
- """
- os.makedirs(output_dir, exist_ok=True)
- metrics = ["ARI", "NMI", "CHAOS", "PAS", "ASW", "HOM", "COM"]
- metric_dfs = []
- for metric, file_path in metric_files.items():
- try:
- data = pd.read_csv(file_path, index_col=0)
- data_subset = data[dataset_names]
- plt.figure(figsize=(8, 6))
- ax = sns.heatmap(
- data_subset,
- cmap="coolwarm",
- cbar_kws={'label': metric},
- annot=True,
- fmt=".2f",
- annot_kws={"size": 6, "weight": "bold"}
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - {metric}", fontweight="bold")
- ax.set_xlabel("Datasets", fontweight="bold")
- ax.set_ylabel("Methods", fontweight="bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- output_path = os.path.join(output_dir, f"{dataset_type}_Heatmap_{metric}.png")
- plt.tight_layout()
- plt.savefig(output_path, bbox_inches="tight")
- plt.close()
- if metric in metrics:
- metric_dfs.append(data_subset)
- except Exception as e:
- print(f"Error processing metric {metric}: {e}")
- try:
- if metric_dfs:
- composite_data_accuracy = sum(metric_dfs[0:4]) / 4
- composite_data_accuracy.to_csv(os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite_accuracy.csv"))
- plt.figure(figsize=(8, 6))
- ax = sns.heatmap(
- composite_data_accuracy,
- cmap="coolwarm",
- cbar_kws={'label': "composite - accuracy"},
- annot=True,
- fmt=".2f",
- annot_kws={"size": 6, "weight": "bold"}
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - Composite (accuracy)", fontweight="bold")
- ax.set_xlabel("Datasets", fontweight="bold")
- ax.set_ylabel("Methods", fontweight="bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- output_path = os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite_accuracy.png")
- plt.tight_layout()
- plt.savefig(output_path, bbox_inches="tight", dpi=200)
- plt.close()
- except Exception as e:
- print(f"Error computing composite accuracy score: {e}")
- try:
- if metric_dfs:
- composite_data_consistency = sum(metric_dfs[4:7]) / 3
- composite_data_consistency.to_csv(os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite_consistency.csv"))
- plt.figure(figsize=(8, 6))
- ax = sns.heatmap(
- composite_data_consistency,
- cmap="coolwarm",
- cbar_kws={'label': "composite - consistency"},
- annot=True,
- fmt=".2f",
- annot_kws={"size": 6, "weight": "bold"}
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - Composite (consistency)", fontweight="bold")
- ax.set_xlabel("Datasets", fontweight="bold")
- ax.set_ylabel("Methods", fontweight="bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- output_path = os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite_consistency.png")
- plt.tight_layout()
- plt.savefig(output_path, bbox_inches="tight", dpi=200)
- plt.close()
- except Exception as e:
- print(f"Error computing composite consistency score: {e}")
- try:
- if metric_dfs:
- composite_data = (composite_data_accuracy + composite_data_consistency)/2
- composite_data.to_csv(os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite.csv"))
- plt.figure(figsize=(8, 6))
- ax = sns.heatmap(
- composite_data,
- cmap="coolwarm",
- cbar_kws={'label': "composite"},
- annot=True,
- fmt=".2f",
- annot_kws={"size": 6, "weight": "bold"}
- )
- ax.set_title(f"{dataset_type} Dataset Heatmap - Composite", fontweight="bold")
- ax.set_xlabel("Datasets", fontweight="bold")
- ax.set_ylabel("Methods", fontweight="bold")
- labels = ax.get_xticklabels()
- bold_font = fm.FontProperties(weight='bold')
- ax.set_xticklabels(labels, rotation=45, ha="right", fontproperties=bold_font)
- y_labels = ax.get_yticklabels()
- ax.set_yticklabels(y_labels, fontproperties=bold_font)
- output_path = os.path.join(output_dir, f"{dataset_type}_Heatmap_Composite.png")
- plt.tight_layout()
- plt.savefig(output_path, bbox_inches="tight", dpi=200)
- plt.close()
- except Exception as e:
- print(f"Error computing composite score: {e}")
- def make_embedding_plots(dataset_name,dataset_paths,pred_paths,save_dir, method_names = ["ground_truth","SCANIT", "CCST" , "DeepST" , "GraphST" , "PROST" , "SpaSRL" , "STAGATE","SpatialPCA" , "banksy" , "giotto" , "DR_SC" , "ISC_MEB" , "BayesSpace" , "PRECAST" , "BayesCafe" , "BASS"],stitch=True):
- """
- Produce embedding plots for select datasets and method outputs
- Parameters
- ----------
- dataset_name : str
- Name of dataset to create plot for
- dataset_paths : dict
- Dictionary of file paths needed for that dataset.
- Must contain at least 'st_path'.
- Can contain 'gnd_path' if ground truth is from a CSV/TSV file.
- pred_paths: dict of dicts
- Dictionary of output file paths for those datasets and names
- saved as a dataset x method dictionary
- save_dir : str
- Path to save output images to
- """
- method_names_dict = {
- "SCANIT": "SCANIT",
- "CCST": "CCST",
- "DeepST": "DeepST",
- "GraphST": "GraphST",
- "PROST": "PROST",
- "SpaSRL": "SpaSRL",
- "STAGATE": "STAGATE",
- "SpatialPCA": "SpatialPCA",
- "banksy": "Banksy",
- "giotto": "Giotto",
- "DR_SC": "DR.SC",
- "ISC_MEB": "ISC.MEB",
- "BayesSpace": "BayesSpace",
- "PRECAST": "PRECAST",
- "BayesCafe": "BayesCafe",
- "BASS": "BASS",
- "SpaceFlow":"SpaceFlow",
- "IRIS" : "IRIS"
- }
- dataset_names_dict = {
- "DLPFC151507": "DLPFC 151507",
- "DLPFC151508": "DLPFC 151508",
- "DLPFC151509": "DLPFC 151509",
- "DLPFC151510": "DLPFC 151510",
- "DLPFC151669": "DLPFC 151669",
- "DLPFC151670": "DLPFC 151670",
- "DLPFC151671": "DLPFC 151671",
- "DLPFC151672": "DLPFC 151672",
- "DLPFC151673": "DLPFC 151673",
- "DLPFC151674": "DLPFC 151674",
- "DLPFC151675": "DLPFC 151675",
- "DLPFC151676": "DLPFC 151676",
- "embryo9.5": "Embryo 9.5",
- "embryo14.5": "Embryo 14.5",
- "mouse_breast_cancer": "Mouse Breast Cancer",
- "MERFISH_brain0.04": "MERFISH Brain 0.04",
- "MERFISH_brain0.09": "MERFISH Brain 0.09",
- "MERFISH_brain0.14": "MERFISH Brain 0.14",
- "MERFISH_brain0.19": "MERFISH Brain 0.19",
- "MERFISH_brain0.24": "MERFISH Brain 0.24",
- "osmFISH": "osmFISH",
- "simulated_kidney_cancer410": "Kidney Cancer 410",
- "simulated_kidney_cancer411": "Kidney Cancer 411",
- "simulated_kidney_cancer506": "Kidney Cancer 506",
- "simulated_breast_cancerER+_CID4290": "Breast Cancer ER+ CID 4290",
- "simulated_breast_cancerTNBC_CID44971": "Breast Cancer TNBC CID 44971",
- "simulated_liver_cancerHCC-1L": "Liver Cancer HCC-1L",
- "simulated_liver_cancerHCC-2L": "Liver Cancer HCC-2L",
- "simulated_liver_cancerHCC-3L": "Liver Cancer HCC-3L",
- "simulated_liver_cancerHCC-4L": "Liver Cancer HCC-4L",
- "simulated_breast_atlas": "Breast Atlas",
- "simulated_intestineA1": "Intestine A1",
- "simulated_intestineA2": "Intestine A2",
- "simulated_chicken_heart": "Chicken Heart",
- "simulated_prostate_cancer" : "Prostate Cancer",
- "simulated_cerebellum" : "Cerebellum"
- }
- gnd, locs = import_dataset(dataset_name, dataset_paths[dataset_name], mode = "plot")
- for method_name in method_names:
- if method_name == "ground_truth":
- gnd = gnd.squeeze()
- adata = an.AnnData(obs=pd.DataFrame({"cluster": gnd.astype(str)}))
- adata.obsm["spatial"] = np.array(locs)
- else:
- pred = load_prediction(dataset_name,method_name, pred_paths[dataset_name][method_name])
- if pred is None:
- print(f"output for {dataset_name} {method_name} not present")
- continue
- pred = pred[~pred.index.duplicated(keep='first')]
- intersect_idx = gnd.index.intersection(pred.index)
- pred_filtered = pred.loc[intersect_idx]
- # Handle both DataFrame and Series cases safely
- if isinstance(pred_filtered, pd.DataFrame):
- if pred_filtered.shape[1] > 1:
- # If multiple columns (e.g. banksy), take the first or warn
- print(f"Warning: {method_name} prediction for {dataset_name} has multiple columns; using the first one.")
- pred_final = pred_filtered.iloc[:, 0]
- else:
- pred_final = pred_filtered
- pred_final = pred_final.squeeze() # Ensure Series, not DataFrame
- pred_final.index = intersect_idx # Ensure correct indexing
- # pred_values = pred_filtered.iloc[:, 0].values
- spatial_coords = locs.loc[intersect_idx].values
- adata = an.AnnData(obs=pd.DataFrame({"cluster": pred_final.astype(str)}))
- adata.obsm['spatial'] = np.array(spatial_coords)
- # Ensure the directory exists
- if not os.path.exists(save_dir):
- os.makedirs(save_dir)
- print(f"Created directory: {save_dir}")
- # Generate and save the plot
- plt.rcParams['font.weight'] = 'bold' # Bold for all text
- plt.rcParams['axes.titleweight'] = 'bold' # Bold for titles
- plt.rcParams['axes.labelweight'] = 'bold'
- fig = sc.pl.embedding(
- adata,
- basis="spatial", # Use the 'spatial' embedding
- color="cluster", # Color points by the 'cluster' column
- title=f"{dataset_names_dict[dataset_name]} - {"Ground Truth" if method_name == "ground_truth" else method_names_dict[method_name]}",
- size=100,
- alpha=1,
- cmap="coolwarm",
- show = True,
- return_fig=True
- # save=f"{dataset_name}_{method_name}.png" # File will be saved in the default directory by Scanpy
- )
- # Move the file to the desired directory
- current_dir = os.getcwd()
- plot_path = os.path.join(current_dir, f"figures/spatial{dataset_name}.png")
- if not os.path.exists(os.path.join(save_dir, f"{dataset_name}")):
- os.makedirs(os.path.join(save_dir, f"{dataset_name}"))
- new_path = os.path.join(save_dir, f"{dataset_name}/{method_name}.png")
- fig.savefig(new_path, dpi=200, bbox_inches="tight", facecolor="white")
- print(f"figure saved to {new_path}")
- plt.close(fig)
- # if os.path.exists(plot_path):
- # os.rename(plot_path, new_path)
- # print(f"Moved plot to: {new_path}")
- # else:
- # print(f"Plot file not found at: {plot_path}")
- # Helper functions to stitch embedding plots together
- def get_all_files(folder_path):
- file_names = [f for f in os.listdir(folder_path) if os.path.isfile(os.path.join(folder_path, f))]
- png_files = [f for f in file_names if f.lower().endswith('.png')]
- png_files.sort()
- ground_truth = 'ground_truth.png' # Replace with the actual name or pattern for the ground truth file
- if ground_truth in png_files:
- png_files.remove(ground_truth)
- png_files.insert(0, ground_truth)
- return png_files
- def save_grid_image(grid_rows, grid_cols, folder_path, dataset_name, save_path, width, height):
- png_files = get_all_files(folder_path)
- image_paths = [os.path.join(folder_path, p) for p in png_files]
- images = [Image.open(img) for img in image_paths]
- fig, axes = plt.subplots(grid_rows, grid_cols, figsize=(width, height), facecolor="white")
- fig.subplots_adjust(wspace=0.0, hspace=0.0) # Reduced spacing between rows and columns
- axes = axes.flatten()
- for ax, img in zip(axes, images + [None] * (grid_rows * grid_cols - len(images))):
- if img is not None:
- ax.imshow(img)
- ax.axis("off") # Remove axes
- output_filename = dataset_name + "_grid_image.png"
- plt.savefig(os.path.join(save_path, output_filename), dpi=200, bbox_inches="tight", facecolor="white")
- plt.close()
- print(f"Grid saved as {output_filename} with 200 DPI.")
- # ----------------------- Composite Score Computation (over data subsets) -----------------------
- def return_custom_score(metric_files, dataset_names):
- metric_dfs = []
- for metric, file_path in metric_files.items():
- try:
- data = pd.read_csv(file_path, index_col=0)
- data_subset = data[dataset_names]
- metric_dfs.append(data_subset)
- except Exception as e:
- print(f"Error processing metric {metric}: {e}")
- try:
- if metric_dfs:
- composite_data_accuracy = sum(metric_dfs[0:4]) / 4
- except Exception as e:
- print(f"Error computing composite accuracy score: {e}")
- try:
- if metric_dfs:
- composite_data_consistency = sum(metric_dfs[4:7]) / 3
- except Exception as e:
- print(f"Error computing composite consistency score: {e}")
- try:
- if metric_dfs:
- composite_data = (composite_data_accuracy + composite_data_consistency)/2
- except Exception as e:
- print(f"Error computing composite score: {e}")
- composite_data = np.sum(composite_data, axis = 1)/len(dataset_names)
- top_5_rows = (
- composite_data # Convert to Series with MultiIndex (row, column)
- .sort_values(ascending=False) # Sort in descending order
- .head(5) # Get top 5
- .index.get_level_values(0) # Extract row names
- .tolist() # Convert to list
- )
- return(composite_data, top_5_rows)
- # ----------------------- Useful Variables (comment out to use) -----------------------
- # dataset_names = ["DLPFC151507" , "DLPFC151508" , "DLPFC151509","DLPFC151510","DLPFC151669","DLPFC151670","DLPFC151671","DLPFC151672","DLPFC151673","DLPFC151674","DLPFC151675","DLPFC151676","embryo9.5" , "embryo14.5" , "mouse_breast_cancer", "MERFISH_brain0.04", "MERFISH_brain0.09", "MERFISH_brain0.14", "MERFISH_brain0.19", "MERFISH_brain0.24","osmFISH" , "simulated_kidney_cancer410" , "simulated_kidney_cancer411" , "simulated_kidney_cancer506" , "simulated_breast_cancerER+_CID4290" , "simulated_breast_cancerTNBC_CID44971" , "simulated_liver_cancerHCC-1L", "simulated_liver_cancerHCC-2L", "simulated_liver_cancerHCC-3L", "simulated_liver_cancerHCC-4L" , "simulated_breast_atlas" , "simulated_intestineA1", "simulated_intestineA2","simulated_chicken_heart", "simulated_prostate_cancer", "simulated_cerebellum"]
- # DLPFC_dataset_names = ["DLPFC151507" , "DLPFC151508" , "DLPFC151509","DLPFC151510","DLPFC151669","DLPFC151670","DLPFC151671","DLPFC151672","DLPFC151673","DLPFC151674","DLPFC151675","DLPFC151676"]
- # real_dataset_names = ["embryo9.5" , "embryo14.5" , "mouse_breast_cancer", "MERFISH_brain0.04", "MERFISH_brain0.09", "MERFISH_brain0.14", "MERFISH_brain0.19", "MERFISH_brain0.24","osmFISH"]
- # simulated_dataset_names = ["simulated_kidney_cancer410" , "simulated_kidney_cancer411" , "simulated_kidney_cancer506" , "simulated_breast_cancerER+_CID4290" , "simulated_breast_cancerTNBC_CID44971" , "simulated_liver_cancerHCC-1L", "simulated_liver_cancerHCC-2L", "simulated_liver_cancerHCC-3L", "simulated_liver_cancerHCC-4L" , "simulated_breast_atlas" , "simulated_intestineA1", "simulated_intestineA2","simulated_chicken_heart", "simulated_prostate_cancer", "simulated_cerebellum"]
- # method_names = ["SCANIT", "CCST" , "DeepST" , "GraphST" , "PROST" , "SpaSRL" , "STAGATE","SpatialPCA" , "banksy" , "giotto" , "DR_SC" , "ISC_MEB" , "BayesSpace" , "PRECAST" , "BayesCafe" , "BASS","SpaceFlow", "IRIS"]
Metrics.py at commit d3b2bec, under MIT · at the source
Overview
- Department of Computer Science and Engineering, Indian Institute of Technology Kanpur
- Department of Mathematics and Statistics, Indian Institute of Technology Kanpur
- Department of Biological Sciences and Bioengineering, Indian Institute of Technology Kanpur
- Mehta Family Centre for Engineering in Medicine, Indian Institute of Technology Kanpur
Abstract
Spatial transcriptomic technologies enable high-resolution characterization of gene expression patterns and reconstruction of cellular architecture within tissue contexts. Two key computational problems have emerged for analyzing these datasets: spatial deconvolution, for disentangling celltype compositions at spatial locations, and spatial domain detection, for identifying spatially coherent regions within a tissue section. Although numerous methods have been developed for each task, a comprehensive and unified benchmarking study spanning diverse tissue types, spatial resolutions, and technological platforms remains lacking, hindering informed method selection by end users and impeding future methodological advancements. Here, we present spDDB (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 20 matches between paragraphs and lines of code.
Zafar-Lab/spDDB
d3b2bec027726acb47fe5a57b0ea8bb1dffea380, 7 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
114 files
- Environments/
Deconvolution/ , R, 43 linesinstallation_Rmethods.R - Environments/
Domain/ , R, 43 linesinstallation_Rmethods.R - Experiments/
_Binning_code_for_Gold_s , Jupyter, 116 linestandard_datasets/ .ipynb_checkpoints/ Binning_in_python_brain- checkpoint.ipynb - Experiments/
_Binning_code_for_Gold_s , Jupyter, 116 linestandard_datasets/ Binning_in_python_brain. ipynb - Experiments/
_Binning_code_for_Gold_s , Jupyter, 105 linestandard_datasets/ Binning_in_python_breast _cancer.ipynb - Experiments/
_Binning_code_for_Gold_s , Jupyter, 110 linestandard_datasets/ Binning_in_python_ileum_ 100um.ipynb - Experiments/
_Binning_code_for_Gold_s , Jupyter, 106 linestandard_datasets/ Binning_in_python_lung_c ancer.ipynb - Experiments/
_Binning_code_for_Gold_s , Python, 110 linestandard_datasets/ binning.py - Experiments/
_CellCharter_shape_chara , Jupyter, 101 linescterization/ .ipynb_checkpoints/ CellCharter_datasets_Rep o-checkpoint.ipynb - Experiments/
_CellCharter_shape_chara , Jupyter, 179 linescterization/ .ipynb_checkpoints/ Top_and_bottom_25_percen tile_Organ_wise-checkpoi nt.ipynb - Experiments/
_CellCharter_shape_chara , Jupyter, 101 linescterization/ CellCharter_datasets_Rep o.ipynb - Experiments/
_CellCharter_shape_chara , Jupyter, 179 linescterization/ Top_and_bottom_25_percen tile_Organ_wise.ipynb - Experiments/
_Composite_score_generat , Jupyter, 130 linesion/ .ipynb_checkpoints/ Save_Composite_Score_Bra in-checkpoint.ipynb - Experiments/
_Composite_score_generat , Jupyter, 130 lines, 2 matchesion/ Save_Composite_Score_Bra in.ipynb - Experiments/
_Composite_score_generat , Python, 110 lines, 2 matchesion/ plotting_files.py - Experiments/
_Composite_score_generat , Python, 324 linesion/ utils.py - Experiments/
_Deconvolution_Metrics_C , Jupyter, 101 linesalculation/ .ipynb_checkpoints/ Evaluation_Brain_Categor y-checkpoint.ipynb - Experiments/
_Deconvolution_Metrics_C , Python, 396 linesalculation/ .ipynb_checkpoints/ create_update_metrics-ch eckpoint.py - Experiments/
_Deconvolution_Metrics_C , Jupyter, 101 lines, 2 matchesalculation/ Evaluation_Brain_Categor y.ipynb - Experiments/
_Deconvolution_Metrics_C , Python, 396 linesalculation/ create_update_metrics.py - Experiments/
_Deconvolution_Metrics_C , Python, 308 lines, 1 matchalculation/ metrics.py - Experiments/
_Domain_Detection_Metric , Python, 1,193 lines, 5 matchess_Calculation/ Metrics.py - Experiments/
_Domain_Detection_Metric , Jupyter, 323 lines, 2 matchess_Calculation/ all_datasets_metric_comp utation.ipynb - Experiments/
_Rare_cell_type_computat , Jupyter, 245 linesion/ .ipynb_checkpoints/ Find_Rare_celltype_July_ Defination1-checkpoint.i pynb - Experiments/
_Rare_cell_type_computat , Jupyter, 206 linesion/ .ipynb_checkpoints/ Find_Regional_Rare_cellt ypes_Brain-checkpoint.ip ynb - Experiments/
_Rare_cell_type_computat , Jupyter, 245 linesion/ Find_Rare_celltype_July_ Defination1.ipynb - Experiments/
_Rare_cell_type_computat , Jupyter, 206 linesion/ Find_Regional_Rare_cellt ypes_Brain.ipynb - SynthST/
Simulator_CTP/ , Jupyter, 162 lines.ipynb_checkpoints/ Run_SynthST-checkpoint.i pynb - SynthST/
Simulator_CTP/ , Jupyter, 162 linesRun_SynthST.ipynb - SynthST/
Simulator_CTP/ , Jupyter, 110 linesSynthST/ .ipynb_checkpoints/ STAGATE_tutorial-checkpo int.ipynb - SynthST/
Simulator_CTP/ , Python, 155 linesSynthST/ .ipynb_checkpoints/ SynthST-checkpoint.py - SynthST/
Simulator_CTP/ , Python, 110 linesSynthST/ .ipynb_checkpoints/ Train_SynthST-checkpoint .py - SynthST/
Simulator_CTP/ , Python, 9 linesSynthST/ .ipynb_checkpoints/ __init__-checkpoint.py - SynthST/
Simulator_CTP/ , Python, 192 linesSynthST/ .ipynb_checkpoints/ model-checkpoint.py - SynthST/
Simulator_CTP/ , Python, 176 linesSynthST/ .ipynb_checkpoints/ utils-checkpoint.py - SynthST/
Simulator_CTP/ , Python, 155 lines, 1 matchSynthST/ SynthST.py - SynthST/
Simulator_CTP/ , Python, 111 linesSynthST/ Train_SynthST.py - SynthST/
Simulator_CTP/ , Python, 9 linesSynthST/ __init__.py - SynthST/
Simulator_CTP/ , Python, 192 lines, 2 matchesSynthST/ model.py - SynthST/
Simulator_CTP/ , Python, 176 linesSynthST/ utils.py - SynthST/
Simulator_gene_expressio , Jupyter, 122 linesn/ .ipynb_checkpoints/ Simulate_gene_expression _notebook-checkpoint.ipy nb - SynthST/
Simulator_gene_expressio , Python, 207 linesn/ .ipynb_checkpoints/ simulate_gene_expression -checkpoint.py - SynthST/
Simulator_gene_expressio , Jupyter, 122 lines, 2 matchesn/ Simulate_gene_expression _notebook.ipynb - SynthST/
Simulator_gene_expressio , Python, 207 linesn/ simulate_gene_expression .py - benchmarking_deconvoluti
on/ , Jupyter, 113 lines.ipynb_checkpoints/ Pipeline_Autogenes-check point.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 130 lines.ipynb_checkpoints/ Pipeline_Cell2Location-c heckpoint.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 144 lines.ipynb_checkpoints/ Pipeline_CellDART-checkp oint.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 123 lines.ipynb_checkpoints/ Pipeline_DestVI-checkpoi nt.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 83 lines.ipynb_checkpoints/ Pipeline_GraphST-checkpo int.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 18 lines.ipynb_checkpoints/ Pipeline_STRIDE-checkpoi nt.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 70 lines.ipynb_checkpoints/ Pipeline_Tangram-checkpo int.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 113 linesPipeline_Autogenes.ipynb - benchmarking_deconvoluti
on/ , R, 45 linesPipeline_CARD.R - benchmarking_deconvoluti
on/ , Jupyter, 130 linesPipeline_Cell2Location.i pynb - benchmarking_deconvoluti
on/ , Jupyter, 144 linesPipeline_CellDART.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 123 linesPipeline_DestVI.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 83 linesPipeline_GraphST.ipynb - benchmarking_deconvoluti
on/ , Jupyter, 50 linesPipeline_POLARIS/ .ipynb_checkpoints/ Terminal-checkpoint.ipyn b - benchmarking_deconvoluti
on/ , R, 54 linesPipeline_POLARIS/ BayesSpace_to_infer_numb er_of_domains.R - benchmarking_deconvoluti
on/ , Jupyter, 21 linesPipeline_POLARIS/ Terminal.ipynb - benchmarking_deconvoluti
on/ , Python, 542 linesPipeline_POLARIS/ datasets.py - benchmarking_deconvoluti
on/ , Python, 747 linesPipeline_POLARIS/ fit_st_layer.py - benchmarking_deconvoluti
on/ , Python, 286 linesPipeline_POLARIS/ model_st_layer.py - benchmarking_deconvoluti
on/ , Python, 250 linesPipeline_POLARIS/ models_mae.py - benchmarking_deconvoluti
on/ , Python, 366 linesPipeline_POLARIS/ train.py - benchmarking_deconvoluti
on/ , Python, 368 linesPipeline_POLARIS/ utils.py - benchmarking_deconvoluti
on/ , R, 75 linesPipeline_RCTD.R - benchmarking_deconvoluti
on/ , R, 38 linesPipeline_Redeconve.R - benchmarking_deconvoluti
on/ , R, 162 linesPipeline_SD.R - benchmarking_deconvoluti
on/ , R, 85 linesPipeline_SONAR.R - benchmarking_deconvoluti
on/ , R, 69 linesPipeline_SPOTlight.R - benchmarking_deconvoluti
on/ , R, 47 linesPipeline_STDeconvolve.R - benchmarking_deconvoluti
on/ , Jupyter, 18 linesPipeline_STRIDE.ipynb - benchmarking_deconvoluti
on/ , R, 47 linesPipeline_Seurat.R - benchmarking_deconvoluti
on/ , R, 82 linesPipeline_SpatialDWLS.R - benchmarking_deconvoluti
on/ , R, 49 linesPipeline_SpatialDecon.R - benchmarking_deconvoluti
on/ , Jupyter, 265 linesPipeline_Spicemix.ipynb - benchmarking_deconvoluti
on/ , Python, 65 linesPipeline_Stereoscope.py - benchmarking_deconvoluti
on/ , Jupyter, 70 linesPipeline_Tangram.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 88 lines.ipynb_checkpoints/ Pipeline_DeepST-checkpoi nt.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 80 lines.ipynb_checkpoints/ Pipeline_GraphST-checkpo int.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 86 lines.ipynb_checkpoints/ Pipeline_PROST-checkpoin t.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 86 lines.ipynb_checkpoints/ Pipeline_SCANIT-checkpoi nt.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 52 lines.ipynb_checkpoints/ Pipeline_STAGATE-checkpo int.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 71 lines.ipynb_checkpoints/ Pipeline_spaSRL-checkpoi nt.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 60 linesCCST/ .ipynb_checkpoints/ Pipeline_CCST-checkpoint .ipynb - benchmarking_domain_dete
ction/ , Python, 221 linesCCST/ CCST.py - benchmarking_domain_dete
ction/ , Python, 284 linesCCST/ CCST_ST_utils.py - benchmarking_domain_dete
ction/ , Jupyter, 60 linesCCST/ Pipeline_CCST.ipynb - benchmarking_domain_dete
ction/ , Python, 189 linesCCST/ data_generation.py - benchmarking_domain_dete
ction/ , Python, 80 lines, 1 matchCCST/ run_CCST.py - benchmarking_domain_dete
ction/ , R, 49 linesPipeline_BASS.R - benchmarking_domain_dete
ction/ , R, 47 linesPipeline_Banksy.R - benchmarking_domain_dete
ction/ , R, 45 linesPipeline_BayesCafe.R - benchmarking_domain_dete
ction/ , R, 32 linesPipeline_BayesSpace.R - benchmarking_domain_dete
ction/ , R, 25 linesPipeline_DRSC.R - benchmarking_domain_dete
ction/ , Jupyter, 88 linesPipeline_DeepST.ipynb - benchmarking_domain_dete
ction/ , R, 63 linesPipeline_Giotto.R - benchmarking_domain_dete
ction/ , Jupyter, 80 linesPipeline_GraphST.ipynb - benchmarking_domain_dete
ction/ , R, 45 linesPipeline_ISC_MEB.R - benchmarking_domain_dete
ction/ , R, 45 linesPipeline_PRECAST.R - benchmarking_domain_dete
ction/ , Jupyter, 86 linesPipeline_PROST.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 86 linesPipeline_SCANIT.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 52 linesPipeline_STAGATE.ipynb - benchmarking_domain_dete
ction/ , Jupyter, 71 linesPipeline_spaSRL.ipynb - docs/
source/ , Python, 36 linesconf.py - docs/
source/ , Jupyter, 178 linestutorials/ 1_SynthST_generate_synth etic_cell_type_proportio ns.ipynb - docs/
source/ , Jupyter, 138 linestutorials/ 2_SynthST_generate_synth etic_spatial_gene_expres sion.ipynb - docs/
source/ , Jupyter, 94 linestutorials/ 3_SynthST_Simulation_str ategy2_Lung_Cancer.ipynb - docs/
source/ , Jupyter, 184 linestutorials/ 4_spDDB_deconvolution_ev aluation.ipynb - docs/
source/ , Jupyter, 133 linestutorials/ 5_spDDB_rare_celltype.ip ynb - docs/
source/ , Jupyter, 189 linestutorials/ 6_spDDB_shape_characteri zation.ipynb - LICENSE, License, 21 lines
- README.md, Text, 51 lines
Code availability
Our benchmarking pipeline for both the tasks is provided as reproducible pipeline at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 112 scripts, each with its path and the digest of its content;
- 20 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zafar-lab.github.io/
spddb_datasets.github.io , at zafar-lab.github.io; found in “Data availability”
Data availability
All datasets used in the study are publicly available. The simulated spatial gene expression datasets generated for spatial deconvolution and domain detection benchmarking have been deposited on figshare and are accessible through the project website: https://
For the spatial cell-type deconvolution task, the spatial and single cell gene expression datasets used to generate simulated spatial gene expression datasets are as follows. The spatial and scRNA dataset for DLPFC are available at ref (46) and (47) respectively, Mouse Brain at ref (39), Hippocampus at ref (24), Cerebellum at ref (24), Visium HD Mouse brain at ref (45) (https://
For the domain detection task, Tthe spatial and single cell gene expression datasets for DLPFC is available at ref(46) (spatial:https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, journal, dates, 4 authors, 4 keywords, 92 references.
Cite
This paper
Shree, A., Aditya, V., Kumar, T., & Zafar, H. (2026). A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies. Research Square (preprint). https://
BibTeX
@article{shree2026compre
author = {Shree, Ajita and Aditya, V and Kumar, Tanush and Zafar, Hamim},
title = {{A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies}},
journal = {Research Square (preprint)},
year = {2026},
month = jun,
publisher = {Research Square},
issn = {2693-5015},
doi = {10.21203/
url = {https://
}
RIS
TY - JOUR
AU - Shree, Ajita
AU - Aditya, V
AU - Kumar, Tanush
AU - Zafar, Hamim
TI - A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
T2 - Research Square (preprint)
J2 - Res Sq
PY - 2026
DA - 2026/
SN - 2693-5015
PB - Research Square
DO - 10.21203/
UR - https://
ER -
CSL-JSON
{
"id": "10.21203/
"type": "article",
"title": "A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies",
"container-title": "Research Square (preprint)",
"author": [
{
"family": "Shree",
"given": "Ajita"
},
{
"family": "Aditya",
"given": "V"
},
{
"family": "Kumar",
"given": "Tanush"
},
{
"family": "Zafar",
"given": "Hamim"
}
],
"container-title-short":
"DOI": "10.21203/
"ISSN": "2693-5015",
"publisher": "Research Square",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41592-026-03194-8 [code]
- Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.Journal: Nature methodsIn common: Squidpy, rpy2, PyTorch Geometric, 17 other tools, methods / tools, genetics / omics, 14 references
- [2] doi:10.1002/advs.77003 [code]
- SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: rpy2, PyTorch Geometric, anndata, 10 other tools, methods / tools, genetics / omics, 10 references
- [3] doi:10.1016/j.isci.2026.117206 [code]
- ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.Journal: iScienceIn common: anndata, Scanpy, Pillow, 7 other tools, genetics / omics, 14 references
- [4] doi:10.1093/bioinformatics/btag540 [code]
- Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.Journal: Bioinformatics (Oxford, England)In common: Squidpy, rpy2, PyTorch Geometric, 11 other tools, methods / tools, genetics / omics, 6 references
- [5] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: rpy2, SingleCellExperiment, UMAP, 18 other tools, genetics / omics
- [6] doi:10.1093/nar/gkag706 [code]
- scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.Journal: Nucleic acids researchIn common: Squidpy, PyTorch Geometric, UMAP, 14 other tools, 3 references
- [7] doi:10.1186/s13073-026-01704-z [code]
- Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.Journal: Genome medicineIn common: SingleCellExperiment, UMAP, anndata, 14 other tools, genetics / omics, 4 references
- [8] doi:10.1093/bib/bbag404 [code]
- Navigating cell maps by deep learning integration of single-cell and spatially resolved transcriptomics.Journal: Briefings in bioinformaticsIn common: PyTorch Geometric, anndata, Scanpy, 10 other tools, genetics / omics, 7 references
- [9] doi:10.1093/bib/bbag298 [code]
- Empowering multifaceted analysis of spatial transcriptomics data with RGAST.Journal: Briefings in bioinformaticsIn common: rpy2, PyTorch Geometric, anndata, 8 other tools, methods / tools, genetics / omics, 8 references
- [10] doi:10.1093/bioinformatics/btag578 [code]
- NicheDeSig: niche-aware deconvolution and adaptive signature analysis for spatial transcriptomics.Journal: Bioinformatics (Oxford, England)In common: anndata, Scanpy, PyTorch, 4 other tools, methods / tools, genetics / omics, 12 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 112 scripts, and 20 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:856dae4e8f00cb56…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
