StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.
The 7 matches
- [1] § 5. Methods › 5.5. Cell density-aware adaptive weighting mechanism and low-confidence region protection ↔ TI/single_time/transition.py, lines 130–233 · score 0.84 · low confidence region, density adaptive, Low density, spatial weights, hole, connectivity
- [2] § 2. Results › 2.1. Overview ↔ TI/single_time/transition.py, lines 130–233 · score 0.80 · fused cost, adaptive transition matrix, cost matrix, high density, low density, spatial distance
- [3] § 2. Results › 2.1. Overview ↔ TI/single_time/lap.py, lines 381–439 · score 0.61 · terminal cells, vector field, transition matrix, adjacency, optimal, embedding
- [4] § 5. Methods › 5.4. Multi-task training objectives ↔ example.ipynb, lines 204–219 · score 0.60 · Kendall correlation, Spearman correlation, ground truth
- [5] § 5. Methods › 5.8. Selection of starting cells ↔ TI/single_time/velocity.py, lines 21–99 · score 0.55 · transition probability matrix, optimal transport, sequence, cells
- [6] § 5. Methods › 5.10. Cell velocity calculation and trajectory construction methods ↔ TI/single_time/velocity.py, lines 650–717 · score 0.53 · displacement vector, cell velocity, weighted
- [7] § 5. Methods › 5.10. Cell velocity calculation and trajectory construction methods ↔ TI/single_time/velocity.py, lines 212–310 · score 0.52 · numerical stability, transition probabilities, velocity, cell
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 834 lines · 25 KB · no license · 3 matches
- import ot
- import sys
- import scanpy as sc
- import pandas as pd
- import numpy as np
- from anndata import AnnData
- from sklearn.metrics.pairwise import euclidean_distances
- from sklearn.neighbors import NearestNeighbors
- from scipy.sparse import csr_matrix
- from scipy.stats import norm
- from numpy.random import RandomState
- import plotly.graph_objs as go
- import plotly.offline as py
- from ipywidgets import VBox
- import seaborn as sns
- from typing import Optional, Union, Literal, Tuple, List
- from .utils import nearest_neighbors, kmeans_centers
- def get_ot_matrix(
- adata: AnnData,
- data_type: str,
- alpha1: int = 1,
- alpha2: int = 1,
- random_state: Union[None, int, RandomState] = 0,
- ) -> np.ndarray:
- """
- Calculate transfer probabilities between cells.
- Using optimal transport theory based on gene expression and/or spatial location information.
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- data_type
- The type of sequencing data.
- - ``'spatial'``: for the spatial transcriptome data.
- - ``'single-cell'``: for the single-cell sequencing data.
- alpha1
- The proportion of spatial location information.
- (Default: 1)
- alpha2
- The proportion of gene expression information.
- (Default: 1)
- random_state
- Different initial states for the pca.
- (Default: 0)
- Returns
- -------
- :class:`~numpy.ndarray`
- Cell transition probability matrix.
- """
- if "X_pca" not in adata.obsm:
- print("X_pca is not in adata.obsm, automatically do PCA first.")
- sc.tl.pca(adata, svd_solver="arpack", random_state=random_state)
- newdata = adata.obsm["X_pca"]
- newdata2 = newdata.copy()
- if data_type == "spatial":
- newcoor = adata.obsm["X_spatial"]
- newcoor2 = newcoor.copy()
- # calculate physical distance
- ed_coor = euclidean_distances(newcoor, newcoor2, squared=True)
- m1 = ed_coor / sum(sum(ed_coor))
- # calculate gene expression PCA space distance
- ed_gene = euclidean_distances(newdata, newdata2, squared=True)
- m2 = ed_gene / sum(sum(ed_gene))
- M = alpha1 * m1 + alpha2 * m2
- M /= M.max()
- row, col = np.diag_indices_from(M)
- M[row, col] = M.max() * 1000000
- elif data_type == "single-cell":
- ed = euclidean_distances(newdata, newdata2, squared=True)
- M = ed / sum(sum(ed))
- M /= M.max()
- row, col = np.diag_indices_from(M)
- M[row, col] = M.max() * 1000000
- else:
- sys.exit(
- "Please give the right data type, choose from 'spatial' or 'single-cell'."
- )
- a, b = (
- np.ones((adata.n_obs,)) / adata.n_obs,
- np.ones((adata.n_obs,)) / adata.n_obs,
- )
- lambd = 1e-1
- Gs = np.array(ot.sinkhorn(a, b, M, lambd))
- return Gs
- def set_start_cells(
- adata: AnnData,
- select_way: Literal["coordinates", "cell_type"],
- cell_type: Optional[str] = None,
- start_point: Optional[Tuple[int, int]] = None,
- basis: str = "spatial",
- split: bool = False,
- n_clusters: int = 2,
- n_neigh: int = 5,
- ) -> list:
- """
- Use coordinates or cell type to manually select starting cells.
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- select_way
- Ways to select starting cells.
- (1) ``'cell_type'``: select by cell type.
- (2) ``'coordinates'``: select by coordinates.
- cell_type
- Restrict the cell type of starting cells.
- (Deafult: None)
- start_point
- The coordinates of the start point in 'coordinates' mode.
- basis
- The basis in `adata.obsm` to store position information.
- split
- Whether to split the specific type of cells into several small clusters according to cell density.
- n_clsuters
- The number of cluster centers after splitting.
- n_neigh
- The number of neighbors next to the start point/cluster center selected as the starting cell.
- Returns
- -------
- list
- The index number of selected starting cells.
- """
- if select_way == "coordinates":
- if start_point is None:
- raise ValueError(
- f"`start_point` must be specified in the 'coordinates' mode."
- )
- start_cells = nearest_neighbors(start_point, adata.obsm["X_" + basis], n_neigh)[
- 0
- ]
- if cell_type is not None:
- type_cells = np.where(adata.obs["cluster"] == cell_type)[0]
- start_cells = set(start_cells).intersection(set(type_cells))
- elif select_way == "cell_type":
- if cell_type is None:
- raise ValueError("in 'cell_type' mode, `cell_type` cannot be None.")
- start_cells = np.where(adata.obs["cluster"] == cell_type)[0]
- if split == True:
- mask = adata.obs["cluster"] == cell_type
- cell_coords = adata.obsm["X_" + basis][mask]
- cluster_centers = kmeans_centers(cell_coords, n_clusters=n_clusters)
- select_cluster_coords = adata.obsm["X_" + basis].copy()
- select_cluster_coords[np.logical_not(mask)] = 1e10
- start_cells = nearest_neighbors(
- cluster_centers, select_cluster_coords, n_neigh
- ).flatten()
- else:
- raise ValueError(f"`select_way` must choose from 'coordinates' or 'cell_type'.")
- return list(start_cells)
- def get_ptime(adata: AnnData, start_cells: list):
- """
- Get the cell pseudotime based on transition probabilities from initial cells.
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- start_cells
- List of index numbers of starting cells.
- Returns
- -------
- :class:`~numpy.ndarray`
- Ptime correspongding to cells.
- """
- select_trans = adata.obsp["trans"][start_cells]
- cell_tran = np.sum(select_trans, axis=0)
- adata.obs["tran"] = cell_tran
- cell_tran_sort = list(np.argsort(cell_tran))
- cell_tran_sort = cell_tran_sort[::-1]
- ptime = pd.Series(dtype="float32", index=adata.obs.index)
- for i in range(adata.n_obs):
- ptime[cell_tran_sort[i]] = i / (adata.n_obs - 1)
- return ptime.values
- from scipy.sparse import issparse, csr_matrix
- import numpy as np
- from sklearn.neighbors import NearestNeighbors
- def get_neigh_trans(
- adata: AnnData, basis: str, n_neigh_pos: int = 10, n_neigh_embed: int = 0
- ):
- """
- Get transport neighbors using both spatial positions and SEDR embeddings
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- basis
- The basis used in visualizing the cell position.
- n_neigh_pos
- Number of neighbors based on spatial positions.
- n_neigh_embed
- Number of neighbors based on SEDR embeddings.
- Returns
- -------
- :class:`~scipy.sparse._csr.csr_matrix`
- Sparse matrix of transition probabilities for selected neighbors.
- """
- if n_neigh_pos == 0 and n_neigh_embed == 0:
- raise ValueError(
- "Number of position and embedding neighbors cannot both be zero."
- )
- # Ensure numerical stability of transition matrix
- trans_mat = adata.obsp["trans"]
- if np.any(trans_mat < 0):
- trans_mat = np.maximum(trans_mat, 0)
- row_sums = trans_mat.sum(axis=1)
- if np.any(row_sums == 0):
- trans_mat += 1e-10
- row_sums = trans_mat.sum(axis=1)
- trans_mat = trans_mat / row_sums[:, np.newaxis]
- adata.obsp["trans"] = trans_mat
- # Spatial position neighbors
- if n_neigh_pos:
- nn_pos = NearestNeighbors(n_neighbors=n_neigh_pos, n_jobs=-1)
- nn_pos.fit(adata.obsm["X_" + basis])
- _, neigh_pos = nn_pos.kneighbors(adata.obsm["X_" + basis])
- neigh_pos = neigh_pos[:, 1:] # exclude self
- # SEDR embedding neighbors
- if n_neigh_embed:
- if "StPedf" not in adata.obsm:
- raise ValueError("StPedf embedding not found in adata.obsm")
- nn_embed = NearestNeighbors(n_neighbors=n_neigh_embed, n_jobs=-1)
- nn_embed.fit(adata.obsm["StPedf"])
- _, neigh_embed = nn_embed.kneighbors(adata.obsm["StPedf"])
- neigh_embed = neigh_embed[:, 1:] # exclude self
- # Build neighbor list
- neigh_list = []
- for i in range(adata.n_obs):
- neighbors = []
- if n_neigh_pos:
- neighbors.extend(neigh_pos[i])
- if n_neigh_embed:
- neighbors.extend(neigh_embed[i])
- # Deduplicate and exclude self
- neighbors = np.unique(neighbors)
- neighbors = neighbors[neighbors != i]
- neigh_list.append(neighbors)
- # Build sparse transition probability matrix
- indptr = [0]
- indices = []
- csr_data = []
- for i in range(adata.n_obs):
- neighbors = neigh_list[i]
- if len(neighbors) == 0:
- continue
- # Get transition probabilities and normalize
- probs = adata.obsp["trans"][i, neighbors]
- if issparse(probs):
- probs = probs.toarray().flatten()
- prob_sum = probs.sum()
- if prob_sum > 0:
- probs = probs / prob_sum
- else:
- probs = np.ones_like(probs) / len(probs)
- indices.extend(neighbors)
- csr_data.extend(probs)
- indptr.append(len(indices))
- trans_neigh_csr = csr_matrix(
- (csr_data, indices, indptr), shape=(adata.n_obs, adata.n_obs)
- )
- return trans_neigh_csr
- '''def get_neigh_trans(
- adata: AnnData, basis: str, n_neigh_pos: int = 10, n_neigh_gene: int = 0
- ):
- """
- Get the transport neighbors from two ways, position and/or gene expression
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- basis
- The basis used in visualizing the cell position.
- n_neigh_pos
- Number of neighbors based on cell positions such as spatial or umap coordinates.
- (Default: 10)
- n_neigh_gene
- Number of neighbors based on gene expression (PCA).
- (Default: 0)
- Returns
- -------
- :class:`~scipy.sparse._csr.csr_matrix`
- A sparse matrix composed of transition probabilities of selected neighbor cells.
- """
- if n_neigh_pos == 0 and n_neigh_gene == 0:
- raise ValueError(
- "the number of position neighbors and gene neighbors cannot be zero at the same time."
- )
- if n_neigh_pos:
- nn = NearestNeighbors(n_neighbors=n_neigh_pos, n_jobs=-1)
- nn.fit(adata.obsm["X_" + basis])
- dist_pos, neigh_pos = nn.kneighbors(adata.obsm["X_" + basis])
- dist_pos = dist_pos[:, 1:]
- neigh_pos = neigh_pos[:, 1:]
- neigh_pos_list = []
- for i in range(adata.n_obs):
- idx = neigh_pos[i] # embedding上的邻居
- idx2 = neigh_pos[idx] # embedding上邻居的邻居
- idx2 = np.setdiff1d(idx2, i)
- neigh_pos_list.append(np.unique(np.concatenate([idx, idx2])))
- # neigh_pos_list.append(idx)
- if n_neigh_gene:
- if "X_pca" not in adata.obsm:
- print("X_pca is not in adata.obsm, automatically do PCA first.")
- sc.tl.pca(adata)
- sc.pp.neighbors(
- adata, use_rep="X_pca", key_added="X_pca", n_neighbors=n_neigh_gene
- )
- neigh_gene = adata.obsp["distances"].indices.reshape(
- -1, adata.uns["neighbors"]["params"]["n_neighbors"] - 1
- )
- indptr = [0]
- indices = []
- csr_data = []
- count = 0
- for i in range(adata.n_obs):
- if n_neigh_pos == 0:
- n_all = neigh_gene[i]
- elif n_neigh_gene == 0:
- n_all = neigh_pos_list[i]
- else:
- n_all = np.unique(np.concatenate([neigh_pos_list[i], neigh_gene[i]]))
- count += len(n_all)
- indptr.append(count)
- indices.extend(n_all)
- csr_data.extend(
- adata.obsp["trans"][i][n_all]
- / (adata.obsp["trans"][i][n_all].sum()) # normalize
- )
- trans_neigh_csr = csr_matrix(
- (csr_data, indices, indptr), shape=(adata.n_obs, adata.n_obs)
- )
- return trans_neigh_csr'''
- '''def get_velocity(
- adata: AnnData,
- basis: str,
- n_neigh_pos: int = 10,
- n_neigh_gene: int = 0,
- grid_num=50,
- smooth=0.5,
- density=1.0,
- ) -> tuple:
- """
- Get the velocity of each cell.
- The speed can be determined in terms of the cell location and/or gene expression.
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- basis
- The label of cell coordinates, for example, `umap` or `spatial`.
- n_neigh_pos
- Number of neighbors based on cell positions such as spatial or umap coordinates.
- (Default: 10)
- n_neigh_gene
- Number of neighbors based on gene expression.
- (Default: 0)
- Returns
- -------
- tuple
- The grid coordinates and cell velocities on each grid to draw the streamplot figure.
- """
- adata.obsp["trans_neigh_csr"] = get_neigh_trans(
- adata, basis, n_neigh_pos, n_neigh_gene
- )
- position = adata.obsm["X_" + basis]
- V = np.zeros(position.shape) # 速度为2维
- for cell in range(adata.n_obs): # 循环每个细胞
- cell_u = 0.0 # 初始化细胞速度
- cell_v = 0.0
- x1 = position[cell][0] # 初始化细胞坐标
- y1 = position[cell][1]
- for neigh in adata.obsp["trans_neigh_csr"][cell].indices: # 针对每个邻居
- p = adata.obsp["trans_neigh_csr"][cell, neigh]
- if (
- adata.obs["ptime"][neigh] < adata.obs["ptime"][cell]
- ): # 若邻居的ptime小于当前的,则概率反向
- p = -p
- x2 = position[neigh][0]
- y2 = position[neigh][1]
- # 正交向量确定速度方向,乘上概率确定速度大小
- sub_u = p * (x2 - x1) / (np.sqrt((x2 - x1) ** 2 + (y2 - y1) ** 2))
- sub_v = p * (y2 - y1) / (np.sqrt((x2 - x1) ** 2 + (y2 - y1) ** 2))
- cell_u += sub_u
- cell_v += sub_v
- V[cell][0] = cell_u / adata.obsp["trans_neigh_csr"][cell].indptr[1]
- V[cell][1] = cell_v / adata.obsp["trans_neigh_csr"][cell].indptr[1]
- adata.obsm["velocity_" + basis] = V
- print(f"The velocity of cells store in 'velocity_{basis}'.")
- P_grid, V_grid = get_velocity_grid(
- adata,
- P=position,
- V=adata.obsm["velocity_" + basis],
- grid_num=grid_num,
- smooth=smooth,
- density=density,
- )
- return P_grid, V_grid'''
- def get_velocity(
- adata: AnnData,
- basis: str,
- n_neigh_pos: int = 10,
- n_neigh_embed: int = 5,
- grid_num=50,
- smooth=0.5,
- density=1.0,
- ) -> tuple:
- """
- Calculate cell velocity using both spatial positions and SEDR embeddings
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- basis
- Coordinate system label (e.g., 'spatial').
- n_neigh_pos
- Number of spatial position neighbors.
- n_neigh_embed
- Number of SEDR embedding neighbors.
- grid_num
- Resolution for grid-based visualization.
- smooth
- Smoothing factor for visualization.
- density
- Density factor for visualization.
- Returns
- -------
- tuple
- Grid positions and velocities for visualization.
- """
- # Get neighbor matrix (based on spatial positions and SEDR embeddings)
- adata.obsp["trans_neigh_csr"] = get_neigh_trans(
- adata, basis, n_neigh_pos, n_neigh_embed
- )
- position = adata.obsm["X_" + basis]
- V = np.zeros(position.shape)
- # Add small value to prevent division by zero
- epsilon = 1e-10
- for cell in range(adata.n_obs):
- cell_u = 0.0
- cell_v = 0.0
- x1, y1 = position[cell]
- neighbors = adata.obsp["trans_neigh_csr"][cell].indices
- n_neighbors = len(neighbors)
- if n_neighbors == 0:
- V[cell] = [0.0, 0.0]
- continue
- for neigh in neighbors:
- p = adata.obsp["trans_neigh_csr"][cell, neigh]
- # Pseudotime direction correction
- if adata.obs["ptime"][neigh] < adata.obs["ptime"][cell]:
- p = -p
- x2, y2 = position[neigh]
- # Calculate distance and direction
- dx = x2 - x1
- dy = y2 - y1
- dist = np.sqrt(dx**2 + dy**2) + epsilon
- # Calculate velocity components
- sub_u = p * dx / dist
- sub_v = p * dy / dist
- # Numerical stability check
- if not np.isfinite(sub_u):
- sub_u = 0.0
- if not np.isfinite(sub_v):
- sub_v = 0.0
- cell_u += sub_u
- cell_v += sub_v
- # Average velocity
- V[cell][0] = cell_u / n_neighbors
- V[cell][1] = cell_v / n_neighbors
- # Final check
- if not np.isfinite(V[cell][0]):
- V[cell][0] = 0.0
- if not np.isfinite(V[cell][1]):
- V[cell][1] = 0.0
- # Add small noise to prevent all zeros
- noise = np.random.normal(0, 1e-6, V.shape)
- V += noise
- adata.obsm["velocity_" + basis] = V
- print(f"Cell velocities stored in 'velocity_{basis}'")
- # Grid processing for visualization
- P_grid, V_grid = get_velocity_grid(
- adata,
- P=position,
- V=V,
- grid_num=grid_num,
- smooth=smooth,
- density=density,
- )
- return P_grid, V_grid
- def get_2_biggest_clusters(adata):
- """Automatic selection of starting cells to determine the direction of cell trajectories
- Args:
- adata (anndata):
- Returns:
- tuple: Contains 2 cluster with maximum sum of transition probabilities
- """
- clusters = np.unique(adata.obs["cluster"])
- cluster_trans = pd.DataFrame(index=clusters, columns=clusters)
- for start_cluster in clusters:
- for end_cluster in clusters:
- if start_cluster == end_cluster:
- cluster_trans.loc[start_cluster, end_cluster] = 0
- continue
- starts = adata.obs["cluster"] == start_cluster
- ends = adata.obs["cluster"] == end_cluster
- cluster_trans.loc[start_cluster, end_cluster] = (
- np.sum(adata.obsp["trans"][starts][:, ends])
- / np.sum(starts)
- / np.sum(ends)
- )
- highest_2_clusters = cluster_trans.stack().astype(float).idxmax()
- return highest_2_clusters
- def auto_get_start_cluster(adata, clusters: Optional[list] = None):
- """
- Select the start cluster with the largest sum of transfer probability
- Parameters
- ----------
- adata
- Anndata
- clusters, list
- Give clusters to find, by default None, each cluster will be traversed and calculated
- Returns
- -------
- str
- One cluster with maximum sum of transition probabilities
- """
- if clusters == None:
- clusters = np.unique(adata.obs["cluster"])
- cluster_prob_sum = {}
- for cluster in clusters:
- start_cells = set_start_cells(adata, select_way="cell_type", cell_type=cluster)
- adata.obs["ptime"] = get_ptime(adata, start_cells)
- cell_time_sort = adata.obs["ptime"].values.argsort()
- prob_sum = 0
- for i in range(len(cell_time_sort) - 1):
- pre = cell_time_sort[i]
- next = cell_time_sort[i + 1]
- prob = adata.obsp["trans"][pre, next]
- prob_sum += prob
- cluster_prob_sum[cluster] = prob_sum
- highest_cluster = max(cluster_prob_sum, key=cluster_prob_sum.get)
- print(
- "The auto selecting cluster is: '"
- + highest_cluster
- + "'. If there is a large discrepancy with the known biological knowledge, please manually select the starting cluster."
- )
- return highest_cluster
- def get_velocity_grid(
- adata,
- P: np.ndarray,
- V: np.ndarray,
- grid_num: int = 50,
- smooth: float = 0.5,
- density: float = 1.0,
- ) -> tuple:
- """
- Convert cell velocity to grid velocity for streamline display
- The visualization of vector field borrows idea from scTour: https://github.com/LiQian-XC/sctour/blob/main/sctour.
- Parameters
- ----------
- P
- The position of cells.
- V
- The velocity of cells.
- smooth
- The factor for scale in Gaussian pdf.
- (Default: 0.5)
- density
- grid density
- (Default: 1.0)
- Returns
- ----------
- tuple
- The embedding and unitary displacement vectors in grid level.
- """
- grids = []
- for dim in range(P.shape[1]):
- m, M = np.min(P[:, dim]), np.max(P[:, dim])
- m = m - 0.01 * np.abs(M - m)
- M = M + 0.01 * np.abs(M - m)
- gr = np.linspace(m, M, int(grid_num * density))
- grids.append(gr)
- meshes = np.meshgrid(*grids)
- P_grid = np.vstack([i.flat for i in meshes]).T
- n_neighbors = int(P.shape[0] / grid_num)
- nn = NearestNeighbors(n_neighbors=n_neighbors, n_jobs=-1)
- nn.fit(P)
- dists, neighs = nn.kneighbors(P_grid)
- scale = np.mean([grid[1] - grid[0] for grid in grids]) * smooth
- weight = norm.pdf(x=dists, scale=scale)
- p_mass = weight.sum(1)
- V_grid = (V[neighs] * weight[:, :, None]).sum(1)
- V_grid /= np.maximum(1, p_mass)[:, None]
- P_grid = np.stack(grids)
- ns = P_grid.shape[1]
- V_grid = V_grid.T.reshape(2, ns, ns)
- mass = np.sqrt((V_grid * V_grid).sum(0))
- min_mass = 1e-5
- min_mass = np.clip(min_mass, None, np.percentile(mass, 99) * 0.01)
- cutoff = mass < min_mass
- V_grid[0][cutoff] = np.nan
- adata.uns["P_grid"] = P_grid
- adata.uns["V_grid"] = V_grid
- return P_grid, V_grid
- class Lasso:
- """
- Lasso an region of interest (ROI) based on spatial cluster.
- Parameters
- ----------
- adata
- An :class:`~anndata.AnnData` object.
- """
- __sub_index = []
- sub_cells = []
- def __init__(self, adata):
- self.adata = adata
- def vi_plot(
- self,
- basis: str = "spatial",
- cell_type: Optional[str] = None,
- ):
- """
- Plot figures.
- Parameters
- ----------
- basis
- The basis in `adata.obsm` to store position information.
- (Deafult: 'spatial')
- cell_type
- Restrict the cell type of starting cells.
- (Deafult: None)
- Returns
- -------
- The container of cell scatter plot and table.
- """
- cell_types = self.adata.obs["cluster"].unique()
- colors = sns.color_palette(n_colors=len(cell_types)).as_hex()
- cluster_color = dict(zip(cell_types, colors))
- self.adata.uns["cluster_color"] = cluster_color
- df = pd.DataFrame()
- df["group_ID"] = self.adata.obs_names
- df["labels"] = self.adata.obs["cluster"].values
- df["spatial_0"] = self.adata.obsm["X_" + basis][:, 0]
- df["spatial_1"] = self.adata.obsm["X_" + basis][:, 1]
- df["color"] = df.labels.map(self.adata.uns["cluster_color"])
- py.init_notebook_mode()
- f = go.FigureWidget(
- [
- go.Scatter(
- x=df["spatial_0"],
- y=df["spatial_1"],
- mode="markers",
- marker_color=df["color"],
- )
- ]
- )
- scatter = f.data[0]
- f.layout.plot_bgcolor = "rgb(255,255,255)"
- f.layout.autosize = False
- axis_dict = dict(
- showticklabels=True,
- autorange=True,
- )
- f.layout.yaxis = axis_dict
- f.layout.xaxis = axis_dict
- f.layout.width = 600
- f.layout.height = 600
- # Create a table FigureWidget that updates on selection from points in the scatter plot of f
- t = go.FigureWidget(
- [
- go.Table(
- header=dict(
- values=["group_ID", "labels", "spatial_0", "spatial_1"],
- fill=dict(color="#C2D4FF"),
- align=["left"] * 5,
- ),
- cells=dict(
- values=[
- df[col]
- for col in ["group_ID", "labels", "spatial_0", "spatial_1"]
- ],
- fill=dict(color="#F5F8FF"),
- align=["left"] * 5,
- ),
- )
- ]
- )
- def selection_fn(trace, points, selector):
- t.data[0].cells.values = [
- df.loc[points.point_inds][col]
- for col in ["group_ID", "labels", "spatial_0", "spatial_1"]
- ]
- Lasso.__sub_index = t.data[0].cells.values[0]
- Lasso.sub_cells = np.where(self.adata.obs.index.isin(Lasso.__sub_index))[0]
- if cell_type is not None:
- type_cells = np.where(self.adata.obs["cluster"] == cell_type)[0]
- Lasso.sub_cells = sorted(
- set(Lasso.sub_cells).intersection(set(type_cells))
- )
- scatter.on_selection(selection_fn)
- # Put everything together
- return VBox((f, t))
velocity.py at commit 57c3873, no license · at the source
Overview
Abstract
Spatial transcriptomics is transforming our multidimensional understanding of cellular spatial organization and its functional mechanisms in processes such as development and disease by systematically resolving the spatial heterogeneity of gene expression within tissues. To delve deeper into the dynamic processes underlying spatial expression patterns, spatial trajectory inference integrates genetic and spatial information to reconstruct the spatial developmental trajectories of cells within tissues. This approach reveals the patterns of differentiation and dynamic changes as cellular states evolve continuously along spatial axes. However, existing methods often struggle to uniformly model the complex, nonlinear interactions between high-dimensional gene expression and spatial coordinates. Here, we introduce StPedf, whose core lies in employing a neural network with a masking mechanism to capture complex nonlinear interactions between high-dimensional genes and spatial positions. It further leverages spatial proximity information as a guiding cue, dynamically and adaptively adjusting the embedding of gene and spatial information and the weighting of spatial proximity information based on spatial density. This enables trajectory inference guided by spatial information. This enables optimal transport to derive intercellular transition matrices, reconstruct cellular differentiation trajectories, and construct pseudo-spatiotemporal maps. StPedf demonstrates superior performance over existing methods on five structurally distinct simulated datasets. Using StPedf, we successfully mapped distinct lineages in the spatial trajectories of telencephalon regeneration in the Ambystoma mexicanum, multiple malignant lineages expanding within primary tumors, and developmental spatial trajectories and pseudo-spatiotemporal maps in human dorsolateral prefrontal cortex (DLPFC). StPedf significantly enhances the accuracy and interpretability of spatial trajectory inference, providing critical technical support for revealing the dynamic patterns of cellular fate transitions within tissue microenvironments.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.
YuanZ0316/StPedf
57c3873d95126c76f8bd28d5f141f7416c749464, 8 April 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
20 files
- StPedf/
StPedf_model.py , Python, 374 lines - StPedf/
StPedf_module.py , Python, 326 lines - StPedf/
__init__.py , Python, 15 lines - StPedf/
clustering_func.py , Python, 65 lines - StPedf/
graph_func.py , Python, 258 lines - StPedf/
utils_func.py , Python, 57 lines - TI/
__init__.py , Python, 3 lines - TI/
single_time/ , Python, 611 linesPgene.py - TI/
single_time/ , Python, 37 lines__init__.py - TI/
single_time/ , Python, 136 linesanalysis.py - TI/
single_time/ , Python, 672 linesgene_regulation.py - TI/
single_time/ , Python, 665 lines, 1 matchlap.py - TI/
single_time/ , Python, 395 lines, 2 matchestransition.py - TI/
single_time/ , Python, 35 linesutils.py - TI/
single_time/ , Python, 672 linesvectorfield.py - TI/
single_time/ , Python, 834 lines, 3 matchesvelocity.py - Tutorial/
1.axolotl_brain_injury_D , Jupyter, 263 lines15_trajectory_analysis.i pynb.ipynb - Tutorial/
2.DLPFC151673.ipynb , Jupyter, 176 lines - example.ipynb, Jupyter, 219 lines, 1 match
- README.md, Text, 113 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 19 scripts, each with its path and the digest of its content;
- 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- db.cngb.org/
data_resources/ , at db.cngb.org; found in “Data Availability”project - research.libd.org/
spatiallibd , at research.libd.org; found in “Data Availability”
Data Availability
Reproducible codes: The StPedf method is implemented as a Python package, publicly available at GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Funding: added Innovative Research Group Project of the National Natural Science Foundation of China: 11831015; National Natural Science Foundation of China: 11831015, 12271216, 92370131
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 9 MeSH terms, 38 references.
Cite
This paper
Zhang, Y., Sun, Z., Shi, Z., Nan, M., Fu, Y., Ren, Q., & Gao, J. (2026). StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion. PLoS computational biology, 22(6), e1014346. https://
BibTeX
@article{zhang2026stpedf
author = {Zhang, Yuan and Sun, Ziyan and Shi, Zhixin and Nan, Mengdi and Fu, Yuhan and Ren, Qing and Gao, Jie},
title = {{StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion}},
journal = {PLoS computational biology},
year = {2026},
month = jun,
volume = {22},
number = {6},
pages = {e1014346},
publisher = {PLOS},
issn = {1553-734X},
doi = {10.1371/
url = {https://
pmid = {42247447},
pmcid = {PMC13240877}
}
RIS
TY - JOUR
AU - Zhang, Yuan
AU - Sun, Ziyan
AU - Shi, Zhixin
AU - Nan, Mengdi
AU - Fu, Yuhan
AU - Ren, Qing
AU - Gao, Jie
TI - StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion
T2 - PLoS computational biology
J2 - PLoS Comput Biol
PY - 2026
DA - 2026/
VL - 22
IS - 6
SP - e1014346
SN - 1553-734X
PB - PLOS
DO - 10.1371/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1371/
"type": "article-journal",
"title": "StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion",
"container-title": "PLoS computational biology",
"author": [
{
"family": "Zhang",
"given": "Yuan"
},
{
"family": "Sun",
"given": "Ziyan"
},
{
"family": "Shi",
"given": "Zhixin"
},
{
"family": "Nan",
"given": "Mengdi"
},
{
"family": "Fu",
"given": "Yuhan"
},
{
"family": "Ren",
"given": "Qing"
},
{
"family": "Gao",
"given": "Jie"
}
],
"container-title-short":
"volume": "22",
"issue": "6",
"page": "e1014346",
"DOI": "10.1371/
"PMID": "42247447",
"PMCID": "PMC13240877",
"ISSN": "1553-734X",
"publisher": "PLOS",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
5
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1002/advs.77003 [code]
- SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: rpy2, anndata, Scanpy, 7 other tools, genetics / omics, 7 references
- [2] doi:10.1093/bib/bbag298 [code]
- Empowering multifaceted analysis of spatial transcriptomics data with RGAST.Journal: Briefings in bioinformaticsIn common: rpy2, anndata, Scanpy, 8 other tools, genetics / omics, 5 references
- [3] doi:10.21203/rs.3.rs-9676637/v1 [code]
- A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic TechnologiesJournal: Research Square (preprint)In common: rpy2, anndata, Scanpy, 7 other tools, genetics / omics, 5 references
- [4] doi:10.1038/s41592-026-03194-8 [code]
- Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.Journal: Nature methodsIn common: rpy2, anndata, Scanpy, 7 other tools, genetics / omics, 5 references
- [5] doi:10.1093/bioinformatics/btag540 [code]
- Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.Journal: Bioinformatics (Oxford, England)In common: rpy2, anndata, Scanpy, 7 other tools, genetics / omics, 5 references
- [6] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: rpy2, anndata, Scanpy, 10 other tools, genetics / omics, 1 reference
- [7] doi:10.1093/bioinformatics/btag430 [code]
- SPIDER: spatially integrated denoising via embedding regularization with single cell supervision.Journal: Bioinformatics (Oxford, England)In common: rpy2, anndata, Scanpy, 6 other tools, genetics / omics, 5 references
- [8] doi:10.1186/s13073-026-01704-z [code]
- Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.Journal: Genome medicineIn common: anndata, Scanpy, PyTorch, 6 other tools, db.cngb.org/data_resources/project, genetics / omics, 2 references
- [9] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: rpy2, anndata, Scanpy, 8 other tools, genetics / omics, 2 references
- [10] doi:10.1016/j.stemcr.2026.103015 [code]
- Brain injury reactivates a developmental program driving genesis and integration of transient LGE-class interneurons.Journal: Stem cell reportsIn common: anndata, Scanpy, NetworkX, 8 other tools, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 19 scripts, and 7 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:b489abb0ec185fe7…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
