OSCR

SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.

Code ↔ Paper

9 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 9 matches
  1. [1] § Results › Overview of SemanticST ↔ semanticst/SemanticST_main.py, lines 28–110 · score 0.77 · min cut loss, reconstruction loss, learning embedding, learned semantic graphs, adjacency matrix, optimization
  2. [2] § Method › SemanticST › Semantic Graph Learning ↔ semanticst/model.py, lines 145–210 · score 0.70 · GCN layers, activation function, feature matrix, semantic graph, nodes, encoder
  3. [3] § Method › Mini‐Batch Training ↔ semanticst/loading_batches.py, lines 126–182 · score 0.68 · DataLoader, PyTorch, mini batch, shuffling, training, spots
  4. [4] § Method › Mini‐Batch Training ↔ semanticst/SemanticST_main.py, lines 28–110 · score 0.68 · DataLoader, spatial graph, adjacency matrix, iteration, reconstructed, PyTorch
  5. [5] § Method › SemanticST's Objective Function › Community‐Based Detection Loss Function ↔ semanticst/SemanticST_main.py, lines 264–334 · score 0.62 · min cut loss, learned embedding, optimize, semantic graphs, reconstruction, encoders
  6. [6] § Method › Data Preprocessing ↔ semanticst/preprocess.py, lines 131–147 · score 0.61 · variable genes, Scanpy, filtered, Preprocessing, Xenium, cells
  7. [7] § Method › SemanticST's Objective Function › Community‐Based Detection Loss Function ↔ semanticst/model.py, lines 75–122 · score 0.60 · Gumbel Softmax, mincut loss, Detection, embedding, matrix, graph
  8. [8] § Method › SemanticST ↔ semanticst/Semantic_graphs.py, lines 85–203 · score 0.56 · graph neural network, Semantic graph learning, GNN, loss
  9. [9] § Method › SemanticST › Fusion of the Latent Representations for Each Semantic Graph ↔ semanticst/model.py, lines 145–210 · score 0.55 · node feature matrix, encoding, fused, decoder, Fusion, encoder

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 451 lines · 18 KB · MIT · 3 matches

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. """
  4. Created on Wed Mar 27 11:31:26 2024
  5. @author: Roxana
  6. """
  7. import numpy as np
  8. import scipy.sparse as sp
  9. import torch
  10. import torch.nn.functional as F
  11. from tqdm import tqdm
  12. from .preprocess import (
  13. preprocess,
  14. construct_interaction,
  15. construct_interaction_KNN,
  16. construct_interaction_KNN_edge_index,
  17. construct_sparse_graph,
  18. sparse_mx_to_torch_sparse_tensor,
  19. fix_seed,
  20. )
  21. from .model import Encoder, DeepMinCutModel
  22. from .Semantic_graphs import SemanticGraphLearning
  23. class Semantic_batches():
  24. """
  25. A class for learning semantic representations from spatial transcriptomics (ST) data
  26. using graph-based methods.
  27. This class constructs spatial graphs, applies deep learning-based feature extraction,
  28. and integrates reconstruction loss and MinCut loss to enhance the learned embeddings.
  29. Parameters
  30. ----------
  31. adata : anndata.AnnData
  32. AnnData object containing spatial transcriptomics data.
  33. train_loader : torch.utils.data.DataLoader
  34. DataLoader for training batches of spatial transcriptomics data.
  35. test_loader : torch.utils.data.DataLoader
  36. DataLoader for testing batches of spatial transcriptomics data.
  37. num_iter : int
  38. Number of iterations (batches) per epoch.
  39. config : object
  40. Configuration object containing device, data type, and other hyperparameters.
  41. learning_rate : float, optional, default=0.001
  42. Learning rate for optimizing the model.
  43. weight_decay : float, optional, default=0.00
  44. Regularization parameter controlling weight decay in the optimizer.
  45. epochs : int, optional, default=1
  46. Number of training epochs.
  47. dim_output : int, optional, default=64
  48. Dimensionality of the output embeddings.
  49. alpha : float, optional, default=1
  50. Weight factor controlling the influence of the reconstruction loss in representation learning.
  51. beta : float, optional, default=0.1
  52. Weight factor controlling the influence of the **MinCut loss** in representation learning.
  53. Integration : bool, optional, default=True
  54. Whether to integrate additional data sources in representation learning.
  55. Attributes
  56. ----------
  57. dim_input : int
  58. Dimensionality of the input features.
  59. device : str
  60. Computation device (`'cpu'` or `'cuda'`), retrieved from config.
  61. adj : torch.Tensor
  62. Adjacency matrix representing the spatial interactions between spots.
  63. edge_index : torch.Tensor
  64. Edge list representing graph connectivity.
  65. G : list
  66. List of semantic graphs constructed during training.
  67. emb_rec : np.ndarray
  68. Learned representations from the decoder.
  69. emb_h : np.ndarray
  70. Learned representations from the encoder.
  71. Methods
  72. -------
  73. train():
  74. Trains the semantic graph learning model on ST data using reconstruction loss and MinCut loss.
  75. test():
  76. Evaluates the trained model and extracts representations.
  77. Returns
  78. -------
  79. anndata.AnnData
  80. Updated AnnData object with learned embeddings stored in `adata.obsm['emb_decoder']`
  81. and `adata.obsm['emb_encoder']`.
  82. """
  83. def __init__(
  84. self,
  85. adata,
  86. train_loader,
  87. test_loader,
  88. num_iter,
  89. config,
  90. learning_rate=0.001,
  91. weight_decay=0.00,
  92. epochs=1000,
  93. dim_output=64,
  94. alpha=1,
  95. beta=0.1,
  96. Integration=True
  97. ):
  98. self.adata = adata
  99. self.train_loader = train_loader
  100. self.test_loader = test_loader
  101. self.config = config
  102. self.learning_rate = learning_rate
  103. self.weight_decay = weight_decay
  104. self.epochs = epochs
  105. self.datatype = self.config.dtype
  106. self.alpha = alpha
  107. self.beta = beta
  108. self.Integration = Integration
  109. self.device = self.config.device
  110. fix_seed(self.config.seed)
  111. batch, _ids = next(iter(train_loader))
  112. self.num_iter = num_iter
  113. self.dim_input = batch[:, :-2].shape[1]
  114. self.dim_output = dim_output
  115. print("\n🚀 Welcome to SemanticST! 🚀\n")
  116. print("📢 Recommendation: If your dataset contains more than 40000 spots or cells, we suggest using **mini-batch training** for efficiency.")
  117. def train(self):
  118. print("\n✅ Using Mini-Batch Training for better efficiency! 🏋️‍♂️")
  119. self.min_model = DeepMinCutModel()
  120. print('Begin to train ST data...')
  121. self.loss = 0
  122. total_loss = 0
  123. batch_idx = 0
  124. self.model = Encoder(self.dim_input, self.dim_output, num_graphs=4).to(self.device)
  125. self.optimizer = torch.optim.Adam(self.model.parameters(), self.learning_rate,
  126. weight_decay=self.weight_decay)
  127. for batch_idx, (data, ids) in enumerate(tqdm(self.train_loader, total=self.num_iter, desc="Training Progress")):
  128. spot_data = data.float()
  129. num_nodes = data.shape[0]
  130. coor = spot_data[:, -2:]
  131. coor = coor.to(torch.int)
  132. W = []
  133. self_loops = torch.arange(num_nodes, device=self.device)
  134. self_loops = torch.stack([self_loops, self_loops], dim=0)
  135. if self.datatype in ['Stereo', 'Slide']:
  136. self.adj = construct_sparse_graph(coor) # scipy sparse, stays sparse
  137. gn = self.adj + sp.eye(self.adj.shape[0])
  138. self.graph_neigh = sparse_mx_to_torch_sparse_tensor(gn).to(self.device) # binary mask
  139. self.edge_index = construct_interaction_KNN_edge_index(coor).to(self.device)
  140. else:
  141. self.adj, self.graph_neigh = construct_interaction(coor)
  142. self.graph_neigh = torch.FloatTensor(self.graph_neigh + np.eye(self.adj.shape[0])).to(self.device)
  143. self.adj = torch.FloatTensor(self.adj).to(self.device)
  144. self.edge_index = self.adj.nonzero().t()
  145. self.edge_index = torch.cat((self.edge_index, self_loops), dim=1)
  146. self.loop_weight = torch.full((num_nodes,), 0.1).to(self.device)
  147. self.sg_learning = SemanticGraphLearning(spot_data[:, :-2], self.adj, self.edge_index, self.device, self.config.use_mini_batch, self.datatype)
  148. self.G = self.sg_learning.train()
  149. for i in range(len(self.G)):
  150. if self.datatype in ['Stereo', 'Slide']:
  151. self.new_edge_weight = self.G[i]
  152. else:
  153. self.new_edge_weight = torch.cat((self.G[i], self.loop_weight))
  154. W.append(self.new_edge_weight)
  155. spot_data = torch.FloatTensor(spot_data[:, :-2])
  156. feature = spot_data.to(self.device)
  157. self.adj_weighted = self.edge_weights_to_sparse(num_nodes)
  158. for epoch in range(self.epochs):
  159. self.model.train()
  160. self.hiden_feat, self.emb, self.g, self.hidden_embeddings, self.attn_weights = self.model(feature, self.edge_index, self.graph_neigh, W)
  161. self.loss_feat = F.mse_loss(feature, self.emb)
  162. self.loss_deep_mincut = self.min_model.deep_mincut_loss(self.hiden_feat, self.adj_weighted)
  163. loss = self.alpha * self.loss_feat + self.beta * self.loss_deep_mincut
  164. self.optimizer.zero_grad()
  165. loss.backward()
  166. self.optimizer.step()
  167. total_loss += loss.data.item()
  168. batch_idx += 1
  169. self.loss = total_loss / (batch_idx + 1)
  170. del self.edge_index, self.adj, spot_data, self.graph_neigh, self.G
  171. del self.loop_weight, self.adj_weighted
  172. print("Optimization finished for ST data!")
  173. with torch.no_grad():
  174. self.model.eval()
  175. Emb_e = []
  176. Emb_d = []
  177. for batch_idx, (data, ids) in enumerate(tqdm(self.test_loader, desc="Testing Progress", unit="batch")):
  178. spot_data = data.float()
  179. num_nodes = data.shape[0]
  180. coor = spot_data[:, -2:]
  181. W = []
  182. self_loops = torch.arange(num_nodes, device=self.device)
  183. self_loops = torch.stack([self_loops, self_loops], dim=0)
  184. if self.datatype in ['Stereo', 'Slide']:
  185. self.adj, self.graph_neigh = construct_interaction_KNN(coor)
  186. self.graph_neigh = torch.FloatTensor(self.graph_neigh + np.eye(self.adj.shape[0])).to(self.device)
  187. self.edge_index = construct_interaction_KNN_edge_index(coor).to(self.device)
  188. self.adj = torch.FloatTensor(self.adj).to(self.device)
  189. else:
  190. self.adj, self.graph_neigh = construct_interaction(coor)
  191. self.graph_neigh = torch.FloatTensor(self.graph_neigh + np.eye(self.adj.shape[0])).to(self.device)
  192. self.adj = torch.FloatTensor(self.adj).to(self.device)
  193. self.edge_index = self.adj.nonzero().t()
  194. self.edge_index = torch.cat((self.edge_index, self_loops), dim=1)
  195. self.loop_weight = torch.full((num_nodes,), 0.1).to(self.device)
  196. self.sg_learning = SemanticGraphLearning(spot_data[:, :-2], self.adj, self.edge_index, self.device, self.config.use_mini_batch, self.datatype)
  197. self.G = self.sg_learning.evaluate(spot_data[:, :-2])
  198. for i in range(len(self.G)):
  199. if self.datatype in ['Stereo', 'Slide']:
  200. self.new_edge_weight = self.G[i]
  201. else:
  202. self.new_edge_weight = torch.cat((self.G[i], self.loop_weight))
  203. W.append(self.new_edge_weight)
  204. spot_data = torch.FloatTensor(spot_data[:, :-2])
  205. feature = spot_data.to(self.device)
  206. self.emb_rec = self.model(feature, self.edge_index, self.graph_neigh, W)[1].detach()
  207. self.emb_h = self.model(feature, self.edge_index, self.graph_neigh, W)[0].detach()
  208. Emb_d.append(self.emb_rec.cpu().numpy())
  209. Emb_e.append(self.emb_h.cpu().numpy())
  210. Emb_d = np.concatenate(Emb_d, axis=0)
  211. Emb_e = np.concatenate(Emb_e, axis=0)
  212. self.adata.obsm['emb_decoder'] = Emb_d
  213. self.adata.obsm['emb_encoder'] = Emb_e
  214. return self.adata
  215. def edge_weights_to_sparse(self, num_nodes):
  216. A = torch.sparse_coo_tensor(
  217. self.edge_index, self.new_edge_weight,
  218. (num_nodes, num_nodes), device=self.device,
  219. )
  220. return (A + A.t()).coalesce()
  221. class Semantic():
  222. """
  223. A class for training a semantic graph learning model on full spatial transcriptomics (ST) data
  224. without mini-batch training.
  225. This class constructs spatial graphs, applies deep learning-based feature extraction,
  226. and integrates reconstruction loss and MinCut loss to enhance the learned embeddings.
  227. Parameters
  228. ----------
  229. adata : anndata.AnnData
  230. AnnData object containing spatial transcriptomics data.
  231. config : object
  232. Configuration object containing device, data type, and other hyperparameters.
  233. learning_rate : float, optional, default=0.001
  234. Learning rate for optimizing the model.
  235. weight_decay : float, optional, default=0.00
  236. Regularization parameter controlling weight decay in the optimizer.
  237. epochs : int, optional, default=1000
  238. Number of training epochs.
  239. dim_input : int, optional, default=3000
  240. Dimensionality of the input features.
  241. dim_output : int, optional, default=64
  242. Dimensionality of the output embeddings.
  243. alpha : float, optional, default=10
  244. Weight factor controlling the influence of the reconstruction loss in representation learning.
  245. beta : float, optional, default=0.1
  246. Weight factor controlling the influence of the **MinCut loss** in representation learning.
  247. Integration : bool, optional, default=True
  248. Whether to integrate additional data sources in representation learning.
  249. Attributes
  250. ----------
  251. device : str
  252. Computation device (`'cpu'` or `'cuda'`), retrieved from config.
  253. adj : torch.Tensor
  254. Adjacency matrix representing the spatial interactions between spots.
  255. edge_index : torch.Tensor
  256. Edge list representing graph connectivity.
  257. G : list
  258. List of semantic graphs constructed during training.
  259. emb_rec : np.ndarray
  260. Learned representations from the decoder.
  261. emb_h : np.ndarray
  262. Learned representations from the encoder.
  263. Methods
  264. -------
  265. train():
  266. Trains the semantic graph learning model on the full dataset without mini-batching.
  267. Returns
  268. -------
  269. anndata.AnnData
  270. Updated AnnData object with learned embeddings stored in `adata.obsm['emb_decoder']`
  271. and `adata.obsm['emb_encoder']`.
  272. """
  273. def __init__(self,
  274. adata, config,
  275. learning_rate=0.001,
  276. weight_decay=0.00,
  277. epochs=1000,
  278. dim_input=3000,
  279. dim_output=64,
  280. alpha=10,
  281. beta=0.1,
  282. lambda_factor=0.1,
  283. Integration=True
  284. ):
  285. self.adata = adata.copy()
  286. self.config = config
  287. self.learning_rate = learning_rate
  288. self.weight_decay = weight_decay
  289. self.epochs = epochs
  290. self.random_seed = self.config.seed
  291. self.alpha = alpha
  292. self.beta = beta
  293. self.Integration = Integration
  294. self.datatype = self.config.dtype
  295. self.dim_input = dim_input
  296. self.dim_output = dim_output
  297. self.device = self.config.device
  298. self.lambda_factor = lambda_factor
  299. fix_seed(self.config.seed)
  300. print("\n🚀 Welcome to SemanticST! 🚀\n")
  301. print("📢 Recommendation: If your dataset contains more than 40000 spots or cells, we suggest using **mini-batch training** for efficiency.")
  302. def train(self):
  303. print("\n✅ Using Full Dataset Training (No Mini-Batching). 🔥")
  304. preprocess(self.adata, self.datatype)
  305. self.min_model = DeepMinCutModel()
  306. print('Begin to train ST data...')
  307. data = self.adata.obsm['feat'].copy()
  308. self.dim_input = data.shape[1]
  309. num_nodes = data.shape[0]
  310. coor = self.adata.obsm['spatial']
  311. coor = torch.from_numpy(coor).to(device=self.device, dtype=torch.int)
  312. W = []
  313. self_loops = torch.arange(num_nodes, device=self.device)
  314. self_loops = torch.stack([self_loops, self_loops], dim=0)
  315. if self.datatype in ['Stereo', 'Slide']:
  316. self.adj = construct_sparse_graph(coor) # scipy sparse, stays sparse
  317. gn = self.adj + sp.eye(self.adj.shape[0])
  318. self.graph_neigh = sparse_mx_to_torch_sparse_tensor(gn).to(self.device) # binary mask
  319. self.edge_index = construct_interaction_KNN_edge_index(coor).to(self.device)
  320. else:
  321. self.adj, self.graph_neigh = construct_interaction(coor)
  322. self.graph_neigh = torch.FloatTensor(self.graph_neigh + np.eye(self.adj.shape[0])).to(self.device)
  323. self.adj = torch.FloatTensor(self.adj).to(self.device)
  324. self.edge_index = self.adj.nonzero().t()
  325. self.edge_index = torch.cat((self.edge_index, self_loops), dim=1)
  326. self.loop_weight = torch.full((num_nodes,), 0.1).to(self.device)
  327. self.sg_learning = SemanticGraphLearning(data, self.adj, self.edge_index, self.device, self.config.use_mini_batch, self.datatype, lambda_factor=self.lambda_factor)
  328. self.sg_learning.train()
  329. self.G = self.sg_learning.evaluate(data)
  330. print("Semantic Graph Learning Completed")
  331. num_graphs = len(self.G[1])
  332. self.model = Encoder(self.dim_input, self.dim_output, num_graphs=num_graphs).to(self.device)
  333. self.optimizer = torch.optim.Adam(self.model.parameters(), self.learning_rate,
  334. weight_decay=self.weight_decay)
  335. for i in range(len(self.G)):
  336. if self.datatype in ['Stereo', 'Slide']:
  337. self.new_edge_weight = self.G[i]
  338. else:
  339. self.new_edge_weight = torch.cat((self.G[i], self.loop_weight))
  340. W.append(self.new_edge_weight)
  341. data = torch.FloatTensor(data)
  342. feature = data.to(self.device)
  343. self.adj_weighted = self.edge_weights_to_sparse(num_nodes)
  344. for epoch in tqdm(range(self.epochs), desc="Feature Learning Epochs"):
  345. self.model.train()
  346. self.hiden_feat, self.emb, self.g, self.hidden_embeddings, self.attn_weights = self.model(feature, self.edge_index, self.graph_neigh, W)
  347. self.loss_feat = F.mse_loss(feature, self.emb)
  348. self.loss_deep_mincut = self.min_model.deep_mincut_loss(self.hiden_feat, self.adj_weighted)
  349. loss = self.alpha * self.loss_feat + self.beta * self.loss_deep_mincut
  350. self.optimizer.zero_grad()
  351. loss.backward()
  352. self.optimizer.step()
  353. with torch.no_grad():
  354. self.model.eval()
  355. data = torch.FloatTensor(data)
  356. feature = data.to(self.device)
  357. self.emb_rec = self.model(feature, self.edge_index, self.graph_neigh, W)[1].detach()
  358. self.emb_h = self.model(feature, self.edge_index, self.graph_neigh, W)[0].detach()
  359. self.adata.obsm['emb_decoder'] = self.emb_rec.cpu().numpy()
  360. self.adata.obsm['emb_encoder'] = self.emb_h.cpu().numpy()
  361. return self.adata
  362. def edge_weights_to_sparse(self, num_nodes):
  363. A = torch.sparse_coo_tensor(
  364. self.edge_index, self.new_edge_weight,
  365. (num_nodes, num_nodes), device=self.device,
  366. )
  367. return (A + A.t()).coalesce()
  368. def cosine_similarity(self, pred_sp, emb_sp): # pred_sp: spot x gene; emb_sp: spot x gene
  369. """
  370. Calculate cosine similarity based on predicted and reconstructed gene expression matrix.
  371. """
  372. M = torch.matmul(pred_sp, emb_sp.T)
  373. Norm_c = torch.norm(pred_sp, p=2, dim=1)
  374. Norm_s = torch.norm(emb_sp, p=2, dim=1)
  375. Norm = torch.matmul(Norm_c.reshape((pred_sp.shape[0], 1)), Norm_s.reshape((emb_sp.shape[0], 1)).T) + -5e-12
  376. M = torch.div(M, Norm)
  377. if torch.any(torch.isnan(M)):
  378. M = torch.where(torch.isnan(M), torch.full_like(M, 0.4868), M)
  379. return M

SemanticST_main.py at commit ad3f204, under MIT · at the source

Overview

  1. UNSW BioMedical Machine Learning Lab (BML), School of Biomedical Engineering, UNSW Sydney, Sydney, Australia
  2. School of Biomedical Engineering, UNSW Sydney, Sydney, Australia
  3. Tyree Institute of Health Engineering (IHealthE), UNSW Sydney, Sydney, Australia
  4. Center of Excellence in Precision Medicine and Digital Health, Department of Physiology, Faculty of Dentistry, Chulalongkorn University, Bangkok, Thailand
  5. Clinic of General, Special Care and Geriatric Dentistry, Center for Dental Medicine, University of Zurich, Zurich, Switzerland
  6. Center for Immune‐Related Diseases, Shanghai Institute of Immunology, Shanghai Jiao Tong University School of Medicine, Shanghai, China
  7. Visiting Scholar (Collaborative Projects), Center of Excellence in Precision Medicine and Digital Health, Chulalongkorn University, Bangkok, Thailand
Dates: received 14 January 2026; accepted 27 July 2026; published online 31 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/advs.77003 · PMID 42671391 · PMCID PMC13528690 · OpenAlex W7204790580
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), methods / tools (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Connectivity, Machine learning, fMRI & imaging
Keywords: deep learning, graph neural network, semantic graph, spatial transcriptomics
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Ratchadaphiseksomphot Endowment Fund, Chulalongkorn University (CTG168027); Health Systems Research Institute (69-145, 69-143); Thailand Science Research and Innovation Fund, Chulalongkorn University (HEA_FF_69_036_3200_003); UNSW (PS348543); UNSW Scientia Program Fellowship; Australian Research Council Discovery Early Career Researcher Award (DE220101210); NSW Cardiovascular Research Network Career Advancement (RG254353)
Citations: not cited yet (Europe PMC); 70 references in the paper

Abstract

Spatial transcriptomics (ST) analysis is often hindered by technical limitations and methodological biases that let dominant signals overshadow subtle but crucial biological patterns, such as rare cell types and fine‐grained heterogeneity. This is especially true for high‐complexity datasets from platforms like Xenium. We present SemanticST, a graph neural network (GNN) framework that addresses this challenge through a fundamentally different design. SemanticST is the first GNN method to implement mini‐batch training, enabling scalable analysis of massive ST datasets (validated on Xenium). Crucially, it employs a multi‐semantic graph fusion strategy that learns disentangled biological representations across tissue, using a min‐cut loss that requires neither graph corruption nor contrastive sampling. Benchmarking across diverse tissues (e.g., brain, embryo, tumor) confirms consistent superiority. It achieves up to 20% higher ARI/NMI on the gold‐standard brain cortex and uniquely delineates all mouse olfactory bulb layers and hippocampal sub‐regions. In high‐resolution breast cancer data, SemanticST identifies computationally plausible spatial domains, including a candidate rare triple receptor‐positive region and a FOXC2‐enriched EMT‐associated domain, from Xenium alone. Furthermore, SemanticST provides superior, robust multi‐sample integration on established benchmarks. SemanticST offers an essential, scalable framework for translating spatial complexity into biologically informative and testable hypotheses.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.

roxana9/SemanticST

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: ad3f2041f38927709f00df70bcda35ca19f08cf7, 7 September 2026
Languages: Python (23)
Size: 38 files, 23 scripts
Software Heritage: not archived
Found in: “Code Availability”
Holds: README, license file, environment (requirements.txt, setup.py)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: PyTorch (16 files), NumPy (12 files), Scanpy (11 files), Matplotlib (9 files), scikit-learn (9 files), pandas (8 files), PyTorch Geometric (7 files), rpy2 (6 files), SciPy (5 files), anndata (1 file), h5py (1 file), Numba (1 file), scikit-image (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
25 files

Code Availability

The SemanticST algorithm is an open‐source Python implementation available at https://github.com/roxana9/SemanticST.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 23 scripts, each with its path and the digest of its content;
  • 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

All data used in this study were obtained from previously published sources. A comprehensive list of these sources is provided in Table S3. For convenience, we have also made the compiled data available via Zenodo: https://zenodo.org/records/15339700.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, pages, dates, 8 authors, 4 keywords, 7 funders, 61 references.

Cite

This paper

Zahedi, R., Argha, A., Farbehi, N., Bakhshayeshi, I., Porntaveetus, T., Ye, Y., Lovell, N. H., & Alinejad‐Rokny, H. (2026). SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics. Advanced science (Weinheim, Baden-Wurttemberg, Germany), e77003. https://doi.org/10.1002/advs.77003

BibTeX

@article{zahedi2026semanticst,
author = {Zahedi, Roxana and Argha, Ahmadreza and Farbehi, Nona and Bakhshayeshi, Ivan and Porntaveetus, Thantrira and Ye, Youqiong and Lovell, Nigel H and Alinejad‐Rokny, Hamid},
title = {{SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics}},
journal = {Advanced science (Weinheim, Baden-Wurttemberg, Germany)},
year = {2026},
month = aug,
pages = {e77003},
publisher = {Wiley},
issn = {2198-3844},
doi = {10.1002/advs.77003},
url = {https://doi.org/10.1002/advs.77003},
pmid = {42671391},
pmcid = {PMC13528690}
}

RIS

TY - JOUR
AU - Zahedi, Roxana
AU - Argha, Ahmadreza
AU - Farbehi, Nona
AU - Bakhshayeshi, Ivan
AU - Porntaveetus, Thantrira
AU - Ye, Youqiong
AU - Lovell, Nigel H
AU - Alinejad‐Rokny, Hamid
TI - SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics
T2 - Advanced science (Weinheim, Baden-Wurttemberg, Germany)
J2 - Adv Sci (Weinh)
PY - 2026
DA - 2026/08/31
SP - e77003
SN - 2198-3844
PB - Wiley
DO - 10.1002/advs.77003
UR - https://doi.org/10.1002/advs.77003
LA - en
ER -

CSL-JSON

{
"id": "10.1002/advs.77003",
"type": "article-journal",
"title": "SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics",
"container-title": "Advanced science (Weinheim, Baden-Wurttemberg, Germany)",
"author": [
{
"family": "Zahedi",
"given": "Roxana"
},
{
"family": "Argha",
"given": "Ahmadreza"
},
{
"family": "Farbehi",
"given": "Nona"
},
{
"family": "Bakhshayeshi",
"given": "Ivan"
},
{
"family": "Porntaveetus",
"given": "Thantrira"
},
{
"family": "Ye",
"given": "Youqiong"
},
{
"family": "Lovell",
"given": "Nigel H"
},
{
"family": "Alinejad‐Rokny",
"given": "Hamid"
}
],
"container-title-short": "Adv Sci (Weinh)",
"page": "e77003",
"DOI": "10.1002/advs.77003",
"PMID": "42671391",
"PMCID": "PMC13528690",
"ISSN": "2198-3844",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/advs.77003",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
31
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/bioinformatics/btag540 [code]
Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.
Journal: Bioinformatics (Oxford, England)
In common: rpy2, PyTorch Geometric, anndata, 10 other tools, methods / tools, genetics / omics, 15 references
[2] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: rpy2, PyTorch Geometric, anndata, 10 other tools, methods / tools, genetics / omics, 10 references
[3] doi:10.1093/bib/bbag298 [code]
Empowering multifaceted analysis of spatial transcriptomics data with RGAST.
Journal: Briefings in bioinformatics
In common: rpy2, PyTorch Geometric, anndata, 9 other tools, methods / tools, genetics / omics, 11 references
[4] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: rpy2, PyTorch Geometric, anndata, 8 other tools, methods / tools, genetics / omics, 12 references
[5] doi:10.1038/s41467-026-71759-4 [code]
CellNiche represents cellular microenvironments in atlas-scale spatial omics data with contrastive learning.
Journal: Nature communications
In common: PyTorch Geometric, anndata, Scanpy, 7 other tools, 10 references
[6] doi:10.1002/advs.75969 [code]
Accurately Deciphering Tissue Heterogeneity From Spatial Multi-Modal and Multi-Omics With STransformer.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: PyTorch Geometric, anndata, Scanpy, 6 other tools, genetics / omics, 8 references
[7] doi:10.1371/journal.pcbi.1014346 [code]
StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.
Journal: PLoS computational biology
In common: rpy2, anndata, Scanpy, 7 other tools, genetics / omics, 7 references
[8] doi:10.1093/bioinformatics/btag430 [code]
SPIDER: spatially integrated denoising via embedding regularization with single cell supervision.
Journal: Bioinformatics (Oxford, England)
In common: rpy2, PyTorch Geometric, anndata, 7 other tools, genetics / omics, 6 references
[9] doi:10.1016/j.isci.2026.117206 [code]
ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.
Journal: iScience
In common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 8 references
[10] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: PyTorch Geometric, anndata, Numba, 10 other tools, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.