OSCR

mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.

Code ↔ Paper

21 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 21 matches
  1. [1] § 3 Results › 3.5 Posterior velocity variability reveals modality-specific uncertainty patterns ↔ mmVelo_tutorial_v2/src/mmvelo_multi/velocity_off_manifold_mod.py, lines 1–37 · score 0.74 · low dimensional embedding, local tangent space, manifold instability, velocity samples, fluctuation, metric
  2. [2] § 2 Methods › 2.13 Preprocessing of data › 2.13.1 10× embryonic E18 mouse brain ↔ mmVelo_tutorial_v2/src/fig_E18_mose_brain/08_benchmarking_latent.py, lines 78–137 · score 0.71 · E18 mouse, highly variable genes, Scanpy, matrix, filtered, brain
  3. [3] § 3 Results › 3.7 mmVelo estimates velocity in missing modalities ↔ mmVelo_tutorial_v2/src/fig_human_brain/01_clustering_pseudotime.py, lines 17–43 · score 0.70 · mGPC, GluN, nIPC, cyc, prog, SP
  4. [4] § 2 Methods › 2.12 Velocity inference on missing modality ↔ mmVelo_tutorial_v2/tutorial_2_human_brain_missing_modality.ipynb, lines 1–59 · score 0.69 · cross modal, human cortical development, scRNA, scATAC, missing modality, profiles
  5. [5] § 3 Results › 3.4 mmVelo reveals the dynamics of TF binding motifs in mouse hair follicle development ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/03_clustering_motif.py, lines 305–341 · score 0.65 · hair shaft cuticle, root sheath, hair follicle, medulla, cortex, motifs
  6. [6] § 3 Results › 3.4 mmVelo reveals the dynamics of TF binding motifs in mouse hair follicle development ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/03_clustering_motif_vjp.py, lines 404–427 · score 0.65 · hair shaft cuticle, root sheath, hair follicle, medulla, cortex, motifs
  7. [7] § 3 Results › 3.7 mmVelo estimates velocity in missing modalities ↔ mmVelo_tutorial_v2/src/mmvelo_tutorial/viz_human_brain.py, lines 94–124 · score 0.64 · excitatory neuron lineage, GluN, nIPC, pseudotime, mmVelo, cells
  8. [8] § 2 Methods › 2.13 Preprocessing of data › 2.13.2 SHARE-seq mouse skin (hair follicle) data ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/02_clustering_heatmap.py, lines 34–65 · score 0.61 · hair shaft cuticle, Hair follicle, TAC, medulla, cortex, mouse
  9. [9] § 2 Methods › 2.13 Preprocessing of data › 2.13.2 SHARE-seq mouse skin (hair follicle) data ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/03_clustering_motif.py, lines 305–341 · score 0.61 · hair shaft cuticle, Hair follicle, TAC, medulla, cortex, mouse
  10. [10] § 2 Methods › 2.6 Training procedure ↔ mmVelo_tutorial_v2/src/mmvelo_tutorial/train_human_brain.py, lines 257–344 · score 0.60 · fine tuning, smoothed profiles, validation, trained, ELBO, reconstruct
  11. [11] § 2 Methods › 2.6 Training procedure ↔ mmVelo_tutorial_v2/src/mmvelo_tutorial/train_mouse_brain.py, lines 139–218 · score 0.59 · fine tuning, smoothed profiles, validation, trained, ELBO, reconstruct
  12. [12] § 2 Methods › 2.6 Training procedure ↔ mmVelo_tutorial/src/streamlineplot.py, lines 1026–1114 · score 0.57 · steady state model, mRNA, RNA velocity, ratio, unspliced, dynamics
  13. [13] § 2 Methods › 2.6 Training procedure ↔ mmVelo_tutorial_v2/src/mmvelo_multi/streamlineplot.py, lines 1028–1116 · score 0.57 · steady state model, mRNA, RNA velocity, ratio, unspliced, dynamics
  14. [14] § 2 Methods › 2.3 Fine-tuning decoders for smoothed profile reconstruction ↔ mmVelo_tutorial_v2/src/mmvelo_tutorial/train_human_brain.py, lines 257–344 · score 0.56 · fine tuned, smoothed profiles, predict, neighborhood, reconstruction, mmVelo
  15. [15] § 2 Methods › 2.3 Fine-tuning decoders for smoothed profile reconstruction ↔ mmVelo_tutorial_v2/src/mmvelo_tutorial/train_mouse_brain.py, lines 139–218 · score 0.55 · fine tuned, smoothed profiles, predict, reconstruction, mmVelo, cell
  16. [16] § 3 Results › 3.7 mmVelo estimates velocity in missing modalities ↔ mmVelo_tutorial_v2/tutorial_2_human_brain_missing_modality.ipynb, lines 1–59 · score 0.54 · human cortical development, scRNA, scATAC, missing modalities, chromatin velocity, mmVelo
  17. [17] § 2 Methods › 2.11 Inferring TF-peak regulatory relationships using chromatin velocity ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/05_GRN_inference_pycistopic_res.py, lines 56–98 · score 0.54 · putative regulator TFs, peak cluster, Leiden, motif, mmVelo
  18. [18] § 2 Methods › 2.11 Inferring TF-peak regulatory relationships using chromatin velocity ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/05_GRN_inference_1st.py, lines 143–200 · score 0.53 · importance scores, mRNA, GRNBoost2, regulate, TFs, Leiden
  19. [19] § 2 Methods › 2.10 Quantification of velocity uncertainty ↔ mmVelo_tutorial_v2/src/mmvelo_multi/velocity_off_manifold.py, lines 1–29 · score 0.53 · manifold instability, latent dynamics, fluctuation, posterior, uncertainty, Quantification
  20. [20] § 2 Methods › 2.10 Quantification of velocity uncertainty ↔ mmVelo_tutorial_v2/src/mmvelo_multi/velocity_off_manifold_mod.py, lines 1–37 · score 0.52 · manifold components, velocity samples, instability, space, Quantification, uncertainty
  21. [21] § 2 Methods › 2.8 Benchmarking against velocity estimation methods ↔ mmVelo_tutorial_v2/src/fig_mouse_hair_follicle/03_clustering_motif_vjp.py, lines 404–427 · score 0.51 · hair shaft cuticle, mouse hair follicle, cortex, score, Pseudotime, cells

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 490 lines · 22 KB · no license · 2 matches

  1. """
  2. velocity_off_manifold_mod.py
  3. We quantify on- and off-manifold uncertainty of latent dynamics
  4. in mmVelo directly within each modality-specific space.
  5. Unlike the latent-space decomposition implemented in velocity_off_manifold.py,
  6. this approach constructs local tangent spaces
  7. based on denoised modality-specific representations (RNA: PCA; ATAC: TruncatedSVD/LSI).
  8. Modality-specific velocity samples are then decomposed into on- and off-manifold components
  9. within these spaces.
  10. Processing flow:
  11. Phase 1 : Collect z_hat and denoised x_RNA, x_ATAC for all cells
  12. Phase 2 : Compute low-dimensional embeddings for each modality using PCA / TruncatedSVD,
  13. and construct kNN graphs for each modality
  14. Phase 3 : Sample d for each batch and compute velocity,
  15. then project into the local tangent spaces of each modality to evaluate on/off-manifold variance
  16. Defined metrics (m ∈ {RNA, ATAC}):
  17. U_dyn^m : dynamics fluctuation uncertainty (all components)
  18. U_dyn_par^m : on-manifold dynamics fluctuation
  19. U_dyn_perp^m : off-manifold instability
  20. R_n^m : off-manifold energy ratio, defined as
  21. mean_s[ ||v_perp^m,(s)||^2 / ||v^m,(s)||^2 ]
  22. """
  23. import os
  24. from typing import Optional, Tuple
  25. import numpy as np
  26. import torch
  27. import torch.nn.functional as F
  28. import torch.distributions as dist
  29. from collections import defaultdict
  30. from sklearn.neighbors import NearestNeighbors
  31. from sklearn.decomposition import PCA, TruncatedSVD
  32. from torch.utils.data import DataLoader
  33. # ──────────────────────────────────────────────────────────────
  34. # utlities
  35. # ──────────────────────────────────────────────────────────────
  36. def _cosine_sim_variance(samples: torch.Tensor) -> torch.Tensor:
  37. """
  38. samples : (S, N, dim)
  39. returns : (N,) 各セルの cosine 類似度の分散 (Bessel 補正あり)
  40. """
  41. v_bar = samples.mean(0) # (N, dim)
  42. cos = (F.normalize(samples, dim=-1, eps=1e-8)
  43. * F.normalize(v_bar.unsqueeze(0), dim=-1, eps=1e-8)
  44. ).sum(-1) # (S, N)
  45. return cos.var(dim=0, unbiased=True) # (N,)
  46. def _get_z_hat(model, s, u, a) -> torch.Tensor:
  47. """posterior mean z_hat = (mu_r + mu_a) / 2"""
  48. qz_mu_r, _ = model.vaes[0].enc_z(s, u)
  49. qz_mu_a, _ = model.vaes[1].enc_z(a)
  50. return (qz_mu_r + qz_mu_a) / 2 # (N, z_dim)
  51. def _local_tangent_projector_batched(
  52. embed_batch: torch.Tensor,
  53. neigh_embed_batch: torch.Tensor,
  54. manifold_dim: int,
  55. ) -> torch.Tensor:
  56. """
  57. Compute the projection matrix onto the local tangent space
  58. in the low-dimensional embedding of each modality.
  59. Parameters
  60. ----------
  61. embed_batch : (n_cells, d_emb) 各細胞の埋め込み
  62. neigh_embed_batch : (n_cells, k, d_emb) 近傍細胞の埋め込み
  63. manifold_dim : r top-r PCs の数
  64. Returns
  65. -------
  66. P_n : (n_cells, d_emb, d_emb) 射影行列 P = U_n U_n^T
  67. """
  68. # 近傍を中心化
  69. centered = neigh_embed_batch - embed_batch.unsqueeze(1) # (n_cells, k, d_emb)
  70. # batched SVD: Vh[:, :r, :] が上位 r 右特異ベクトル (各行 = 主方向)
  71. _, _, Vh = torch.linalg.svd(centered, full_matrices=False) # Vh: (n_cells, min(k,d_emb), d_emb)
  72. r = min(manifold_dim, Vh.shape[1])
  73. U_n = Vh[:, :r, :].transpose(-1, -2) # (n_cells, d_emb, r)
  74. P_n = U_n @ U_n.transpose(-1, -2) # (n_cells, d_emb, d_emb)
  75. return P_n
  76. # ──────────────────────────────────────────────────────────────
  77. # Phase 1: Compute z_hat and denoised modality representations
  78. # ──────────────────────────────────────────────────────────────
  79. @torch.no_grad()
  80. def _collect_representations(
  81. model, dataloader: DataLoader, device: str
  82. ) -> Tuple[np.ndarray, np.ndarray, np.ndarray]:
  83. """
  84. Compute z_hat, x_RNA = f^s(z_hat)*C_s, x_ATAC = f^a(z_hat)*C_a
  85. Returns
  86. -------
  87. all_z_hat : (N, z_dim)
  88. all_x_rna : (N, rna_dim)
  89. all_x_atac : (N, atac_dim)
  90. """
  91. model.eval()
  92. z_hat_list = []
  93. x_rna_list = []
  94. x_atac_list = []
  95. for batch in dataloader:
  96. s, u, a = batch[0].to(device), batch[1].to(device), batch[2].to(device)
  97. z_hat = _get_z_hat(model, s, u, a)
  98. x_rna = model.vaes[0].dec_su(z_hat)[0][0] * model.norm_mat_s
  99. x_atac = model.vaes[1].dec_ald(z_hat)[0] * model.norm_mat_a
  100. z_hat_list.append(z_hat.cpu().numpy())
  101. x_rna_list.append(x_rna.cpu().numpy())
  102. x_atac_list.append(x_atac.cpu().numpy())
  103. return (
  104. np.concatenate(z_hat_list, axis=0),
  105. np.concatenate(x_rna_list, axis=0),
  106. np.concatenate(x_atac_list, axis=0),
  107. )
  108. # ──────────────────────────────────────────────────────────────
  109. # Phase 2: compute low dimensional embeddings and construct a kNN graph
  110. # ──────────────────────────────────────────────────────────────
  111. def _build_modality_embedding_and_knn(
  112. x_rna: np.ndarray,
  113. x_atac: np.ndarray,
  114. n_pca_rna: int,
  115. n_lsi_atac: int,
  116. n_neighbors: int,
  117. ) -> Tuple[np.ndarray, np.ndarray, np.ndarray, np.ndarray, np.ndarray, np.ndarray]:
  118. """
  119. For RNA and ATAC:
  120. - compute low-dimensional embeddings (PCA / TruncatedSVD)
  121. - obtain loading matrix (to project velocity onto the low-dimensional space)
  122. - construct a kNNgraph
  123. Returns
  124. -------
  125. embed_rna : (N, n_pca_rna) RNA 低次元埋め込み
  126. W_rna : (rna_dim, n_pca_rna) PCA loading 行列
  127. nn_idx_rna : (N, k) RNA kNN インデックス
  128. embed_atac : (N, n_lsi_atac) ATAC 低次元埋め込み
  129. W_atac : (atac_dim, n_lsi_atac) TruncatedSVD loading 行列
  130. nn_idx_atac : (N, k) ATAC kNN インデックス
  131. """
  132. N = x_rna.shape[0]
  133. k = min(n_neighbors, N - 1)
  134. # ── RNA: PCA ──────────────────────────────────────────────
  135. n_comp_rna = min(n_pca_rna, x_rna.shape[1], N - 1)
  136. pca_rna = PCA(n_components=n_comp_rna, whiten=False)
  137. embed_rna = pca_rna.fit_transform(x_rna).astype(np.float32) # (N, n_pca_rna)
  138. W_rna = pca_rna.components_.T.astype(np.float32) # (rna_dim, n_pca_rna)
  139. nn_rna = NearestNeighbors(n_neighbors=k, metric="euclidean", algorithm="auto")
  140. nn_rna.fit(embed_rna)
  141. nn_idx_rna = nn_rna.kneighbors(embed_rna, return_distance=False) # (N, k)
  142. # ── ATAC: TruncatedSVD (LSI) ──────────────────────────────
  143. n_comp_atac = min(n_lsi_atac, x_atac.shape[1], N - 1)
  144. svd_atac = TruncatedSVD(n_components=n_comp_atac, algorithm="randomized")
  145. embed_atac = svd_atac.fit_transform(x_atac).astype(np.float32) # (N, n_lsi_atac)
  146. W_atac = svd_atac.components_.T.astype(np.float32) # (atac_dim, n_lsi_atac)
  147. nn_atac = NearestNeighbors(n_neighbors=k, metric="euclidean", algorithm="auto")
  148. nn_atac.fit(embed_atac)
  149. nn_idx_atac = nn_atac.kneighbors(embed_atac, return_distance=False) # (N, k)
  150. return embed_rna, W_rna, nn_idx_rna, embed_atac, W_atac, nn_idx_atac
  151. # ──────────────────────────────────────────────────────────────
  152. # Phase 3: compute on/off-manifold decomposition
  153. # ──────────────────────────────────────────────────────────────
  154. @torch.no_grad()
  155. def _batch_off_manifold_mod(
  156. model,
  157. batch,
  158. # RNA modality
  159. embed_rna_batch: torch.Tensor, # (n_cells, n_pca_rna)
  160. neigh_embed_rna_batch: torch.Tensor, # (n_cells, k, n_pca_rna)
  161. W_rna: torch.Tensor, # (rna_dim, n_pca_rna)
  162. manifold_dim_rna: int,
  163. # ATAC modality
  164. embed_atac_batch: torch.Tensor, # (n_cells, n_lsi_atac)
  165. neigh_embed_atac_batch: torch.Tensor,# (n_cells, k, n_lsi_atac)
  166. W_atac: torch.Tensor, # (atac_dim, n_lsi_atac)
  167. manifold_dim_atac: int,
  168. # sampling
  169. n_samples: int,
  170. device: str,
  171. ) -> dict:
  172. """
  173. Compute the modality-specific on/off-manifold uncertainty decomposition for one batch.
  174. The velocity is projected onto the low-dimensional embeddings of each modality
  175. before decomposing it into on- and off-manifold components.
  176. Returns
  177. -------
  178. dict:
  179. z_hat (n_cells, z_dim)
  180. vel_mean_rna (n_cells, rna_dim) RNA velocity 平均
  181. vel_mean_atac (n_cells, atac_dim) ATAC velocity 平均
  182. u_dyn_rna (n_cells,)
  183. u_dyn_atac (n_cells,)
  184. u_dyn_par_rna (n_cells,)
  185. u_dyn_par_atac (n_cells,)
  186. u_dyn_perp_rna (n_cells,)
  187. u_dyn_perp_atac (n_cells,)
  188. off_ratio_rna (n_cells,)
  189. off_ratio_atac (n_cells,)
  190. """
  191. s, u, a = batch[0].to(device), batch[1].to(device), batch[2].to(device)
  192. z_hat = _get_z_hat(model, s, u, a) # (n_cells, z_dim)
  193. # ── modality-specific 局所接線空間の射影行列 ──────────────
  194. # RNA: (n_cells, n_pca_rna, n_pca_rna)
  195. P_rna = _local_tangent_projector_batched(embed_rna_batch, neigh_embed_rna_batch, manifold_dim_rna)
  196. I_rna = torch.eye(P_rna.shape[-1], device=device).unsqueeze(0).expand(z_hat.shape[0], -1, -1)
  197. # ATAC: (n_cells, n_lsi_atac, n_lsi_atac)
  198. P_atac = _local_tangent_projector_batched(embed_atac_batch, neigh_embed_atac_batch, manifold_dim_atac)
  199. I_atac = torch.eye(P_atac.shape[-1], device=device).unsqueeze(0).expand(z_hat.shape[0], -1, -1)
  200. # base decode
  201. s_base = model.vaes[0].dec_su(z_hat)[0][0] * model.norm_mat_s # (n_cells, rna_dim)
  202. a_base = model.vaes[1].dec_ald(z_hat)[0] * model.norm_mat_a # (n_cells, atac_dim)
  203. # q(d | z_hat)
  204. qd = model.vaes[0].enc_dyn(z_hat)
  205. # (high-dim velocity — for mean and total U_dyn)
  206. vel_rna_full_list = []
  207. vel_atac_full_list = []
  208. # low-dim projected velocity (for on/off decomposition)
  209. vel_rna_pca_list = []
  210. vel_atac_lsi_list = []
  211. off_ratio_rna_list = []
  212. off_ratio_atac_list = []
  213. for _ in range(n_samples):
  214. d_s = qd.rsample() # (n_cells, z_dim)
  215. # velocity in high-dim space
  216. z_dt = z_hat + model.d_coeff * d_s
  217. v_rna = model.vaes[0].dec_su(z_dt)[0][0] * model.norm_mat_s - s_base # (n_cells, rna_dim)
  218. v_atac = model.vaes[1].dec_ald(z_dt)[0] * model.norm_mat_a - a_base # (n_cells, atac_dim)
  219. vel_rna_full_list.append(v_rna)
  220. vel_atac_full_list.append(v_atac)
  221. # velocity を低次元に射影
  222. v_rna_pca = v_rna @ W_rna # (n_cells, n_pca_rna)
  223. v_atac_lsi = v_atac @ W_atac # (n_cells, n_lsi_atac)
  224. vel_rna_pca_list.append(v_rna_pca)
  225. vel_atac_lsi_list.append(v_atac_lsi)
  226. # off-manifold energy ratio (低次元空間で計算)
  227. v_rna_perp_pca = ((I_rna - P_rna) @ v_rna_pca.unsqueeze(-1)).squeeze(-1)
  228. v_atac_perp_lsi = ((I_atac - P_atac) @ v_atac_lsi.unsqueeze(-1)).squeeze(-1)
  229. norm_rna_sq = (v_rna_pca ** 2).sum(-1).clamp(min=1e-8)
  230. norm_atac_sq = (v_atac_lsi ** 2).sum(-1).clamp(min=1e-8)
  231. off_ratio_rna_list.append( (v_rna_perp_pca ** 2).sum(-1) / norm_rna_sq )
  232. off_ratio_atac_list.append((v_atac_perp_lsi ** 2).sum(-1) / norm_atac_sq )
  233. # (S, n_cells, dim) に集約
  234. rna_full_stack = torch.stack(vel_rna_full_list, dim=0) # (S, n_cells, rna_dim)
  235. atac_full_stack = torch.stack(vel_atac_full_list, dim=0)
  236. rna_pca_stack = torch.stack(vel_rna_pca_list, dim=0) # (S, n_cells, n_pca_rna)
  237. atac_lsi_stack = torch.stack(vel_atac_lsi_list, dim=0)
  238. # on-manifold / off-manifold 成分 (低次元空間)
  239. def _proj(stack, P, I):
  240. # stack: (S, n_cells, d_emb), P: (n_cells, d_emb, d_emb)
  241. v_col = stack.unsqueeze(-1) # (S, n_cells, d_emb, 1)
  242. par = (P.unsqueeze(0) @ v_col).squeeze(-1) # (S, n_cells, d_emb)
  243. perp = ((I - P).unsqueeze(0) @ v_col).squeeze(-1)
  244. return par, perp
  245. rna_par_stack, rna_perp_stack = _proj(rna_pca_stack, P_rna, I_rna)
  246. atac_par_stack, atac_perp_stack = _proj(atac_lsi_stack, P_atac, I_atac)
  247. def _var(stack):
  248. return _cosine_sim_variance(stack).cpu().numpy()
  249. return {
  250. "z_hat": z_hat.cpu().numpy(),
  251. "vel_mean_rna": rna_full_stack.mean(0).cpu().numpy(),
  252. "vel_mean_atac": atac_full_stack.mean(0).cpu().numpy(),
  253. # total (high-dim velocity)
  254. "u_dyn_rna": _var(rna_full_stack),
  255. "u_dyn_atac": _var(atac_full_stack),
  256. # on-manifold (low-dim projected)
  257. "u_dyn_par_rna": _var(rna_par_stack),
  258. "u_dyn_par_atac": _var(atac_par_stack),
  259. # off-manifold (low-dim projected)
  260. "u_dyn_perp_rna": _var(rna_perp_stack),
  261. "u_dyn_perp_atac": _var(atac_perp_stack),
  262. # off-manifold エネルギー比
  263. "off_ratio_rna": torch.stack(off_ratio_rna_list, dim=0).mean(0).cpu().numpy(),
  264. "off_ratio_atac": torch.stack(off_ratio_atac_list, dim=0).mean(0).cpu().numpy(),
  265. }
  266. # ──────────────────────────────────────────────────────────────
  267. # main function
  268. # ──────────────────────────────────────────────────────────────
  269. def compute_off_manifold_uncertainty_mod(
  270. model,
  271. dataloader: DataLoader,
  272. n_samples: int = 200,
  273. n_neighbors: int = 50,
  274. n_pca_rna: int = 50,
  275. n_lsi_atac: int = 50,
  276. manifold_dim_rna: int = 10,
  277. manifold_dim_atac: int = 10,
  278. device: str = "cuda",
  279. ) -> dict:
  280. """
  281. Compute modality-specific on/off-manifold uncertainty decomposition for all cells.
  282. Phase 1: Compute z_hat and denoised modality representations (x_RNA, x_ATAC).
  283. Phase 2: Compute low-dimension embeddings for each modality (RNA: PCA, ATAC: TruncatedSVD)
  284. and construct kNN graphs for each modality.
  285. Phase 3: Sample d for each batch and compute velocity, then estimate local tangent spaces in the low-dimensional modality space and decompose the uncertainty.
  286. Parameters
  287. ----------
  288. model : trained model (DREG_DYN)
  289. dataloader : DataLoader
  290. n_samples : number of sampling iterations S
  291. n_neighbors : number of kNN neighbors k
  292. n_pca_rna : dimension of RNA low-dimensional embedding
  293. n_lsi_atac : dimension of ATAC low-dimensional embedding
  294. manifold_dim_rna : dimension of RNA local tangent space r
  295. manifold_dim_atac: dimension of ATAC local tangent space r
  296. device : Device to run the computation on (e.g., "cuda" or "cpu")
  297. Returns
  298. -------
  299. dict:
  300. z_hat (N, z_dim)
  301. vel_mean_rna (N, rna_dim)
  302. vel_mean_atac (N, atac_dim)
  303. u_dyn_rna (N,) total dynamics fluctuation uncertainty (RNA)
  304. u_dyn_atac (N,) total dynamics fluctuation uncertainty (ATAC)
  305. u_dyn_par_rna (N,) on-manifold dynamics fluctuation (RNA)
  306. u_dyn_par_atac (N,) on-manifold dynamics fluctuation (ATAC)
  307. u_dyn_perp_rna (N,) off-manifold instability (RNA)
  308. u_dyn_perp_atac (N,) off-manifold instability (ATAC)
  309. off_ratio_rna (N,) off-manifold energy ratio R_n^RNA
  310. off_ratio_atac (N,) off-manifold energy ratio R_n^ATAC
  311. """
  312. model.to(device)
  313. model.eval()
  314. # ── Phase 1 ─────────────────────────────
  315. print("[Phase 1] Collecting z_hat, x_RNA, x_ATAC for all cells...")
  316. all_z_hat, all_x_rna, all_x_atac = _collect_representations(model, dataloader, device)
  317. N = len(all_z_hat)
  318. print(f" N={N}, rna_dim={all_x_rna.shape[1]}, atac_dim={all_x_atac.shape[1]}")
  319. # ── Phase 2────────────────────────
  320. print(f"[Phase 2] Building modality-specific embeddings and kNN graphs "
  321. f"(RNA: PCA-{n_pca_rna}, ATAC: SVD-{n_lsi_atac}, k={n_neighbors})...")
  322. (embed_rna, W_rna,
  323. nn_idx_rna,
  324. embed_atac, W_atac,
  325. nn_idx_atac) = _build_modality_embedding_and_knn(
  326. all_x_rna, all_x_atac,
  327. n_pca_rna=n_pca_rna,
  328. n_lsi_atac=n_lsi_atac,
  329. n_neighbors=n_neighbors,
  330. )
  331. print(f" embed_rna={embed_rna.shape}, embed_atac={embed_atac.shape}")
  332. # テンソル化
  333. embed_rna_t = torch.tensor(embed_rna, dtype=torch.float32, device=device) # (N, n_pca_rna)
  334. embed_atac_t = torch.tensor(embed_atac, dtype=torch.float32, device=device) # (N, n_lsi_atac)
  335. W_rna_t = torch.tensor(W_rna, dtype=torch.float32, device=device) # (rna_dim, n_pca_rna)
  336. W_atac_t = torch.tensor(W_atac, dtype=torch.float32, device=device) # (atac_dim, n_lsi_atac)
  337. nn_rna_t = torch.tensor(nn_idx_rna, dtype=torch.long, device=device) # (N, k)
  338. nn_atac_t = torch.tensor(nn_idx_atac, dtype=torch.long, device=device) # (N, k)
  339. # 近傍埋め込み配列
  340. neigh_embed_rna = embed_rna_t[nn_rna_t] # (N, k, n_pca_rna)
  341. neigh_embed_atac = embed_atac_t[nn_atac_t] # (N, k, n_lsi_atac)
  342. # ── Phase 3───────────
  343. print(f"[Phase 3] Computing modality-specific on/off-manifold decomposition "
  344. f"(n_samples={n_samples}, manifold_dim_rna={manifold_dim_rna}, "
  345. f"manifold_dim_atac={manifold_dim_atac})...")
  346. accum = defaultdict(list)
  347. cell_offset = 0
  348. for batch in dataloader:
  349. n_cells = batch[0].shape[0]
  350. sl = slice(cell_offset, cell_offset + n_cells)
  351. out = _batch_off_manifold_mod(
  352. model, batch,
  353. embed_rna_batch = embed_rna_t[sl],
  354. neigh_embed_rna_batch = neigh_embed_rna[sl],
  355. W_rna = W_rna_t,
  356. manifold_dim_rna = manifold_dim_rna,
  357. embed_atac_batch = embed_atac_t[sl],
  358. neigh_embed_atac_batch = neigh_embed_atac[sl],
  359. W_atac = W_atac_t,
  360. manifold_dim_atac = manifold_dim_atac,
  361. n_samples = n_samples,
  362. device = device,
  363. )
  364. for key, val in out.items():
  365. accum[key].append(val)
  366. cell_offset += n_cells
  367. return {key: np.concatenate(vals, axis=0) for key, vals in accum.items()}
  368. # ──────────────────────────────────────────────────────────────
  369. # save
  370. # ──────────────────────────────────────────────────────────────
  371. def save_off_manifold_uncertainty_mod(
  372. dm,
  373. results: dict,
  374. save_dir: str,
  375. filename_rna: str = "adata_rna_offmanifold_mod.loom",
  376. filename_atac: str = "adata_atac_offmanifold_mod.loom",
  377. ) -> None:
  378. """
  379. save the results of modality-specific off-manifold decomposition
  380. to AnnData and save as .loom files.
  381. RNA AnnData:
  382. layers["vel_mean"] : RNA velocity 平均 (N, rna_dim)
  383. obsm["z_hat"] : posterior mean z (N, z_dim)
  384. obs["u_dyn"] : total dynamics fluctuation uncertainty
  385. obs["u_dyn_par"] : on-manifold dynamics fluctuation
  386. obs["u_dyn_perp"] : off-manifold instability
  387. obs["off_ratio"] : off-manifold energy ratio R_n^RNA
  388. ATAC AnnData:
  389. layers["vel_mean"] : ATAC velocity 平均 (N, atac_dim)
  390. obs["u_dyn"] : total dynamics fluctuation uncertainty
  391. obs["u_dyn_par"] : on-manifold dynamics fluctuation
  392. obs["u_dyn_perp"] : off-manifold instability
  393. obs["off_ratio"] : off-manifold energy ratio R_n^ATAC
  394. """
  395. os.makedirs(save_dir, exist_ok=True)
  396. dm.adata_r.layers["vel_mean"] = results["vel_mean_rna"]
  397. dm.adata_r.obsm["z_hat"] = results["z_hat"]
  398. dm.adata_r.obs["u_dyn"] = results["u_dyn_rna"]
  399. dm.adata_r.obs["u_dyn_par"] = results["u_dyn_par_rna"]
  400. dm.adata_r.obs["u_dyn_perp"] = results["u_dyn_perp_rna"]
  401. dm.adata_r.obs["off_ratio"] = results["off_ratio_rna"]
  402. dm.adata_a.layers["vel_mean"] = results["vel_mean_atac"]
  403. dm.adata_a.obs["u_dyn"] = results["u_dyn_atac"]
  404. dm.adata_a.obs["u_dyn_par"] = results["u_dyn_par_atac"]
  405. dm.adata_a.obs["u_dyn_perp"] = results["u_dyn_perp_atac"]
  406. dm.adata_a.obs["off_ratio"] = results["off_ratio_atac"]
  407. rna_path = os.path.join(save_dir, filename_rna)
  408. atac_path = os.path.join(save_dir, filename_atac)
  409. dm.adata_r.write_loom(rna_path, write_obsm_varm=True)
  410. dm.adata_a.write_loom(atac_path, write_obsm_varm=True)
  411. print(f"Saved RNA AnnData → {rna_path}")
  412. print(f"Saved ATAC AnnData → {atac_path}")

velocity_off_manifold_mod.py at commit 3dbe5d1, no license · at the source

Overview

Authors: Satoshi Nomura1, Yasuhiro Kojima2, Kodai Minoura1, Shuto Hayashi3, Ko Abe3, Haruka Hirose3, Teppei Shimamura1,3
  1. Japanese Red Cross Aichi Medical Center, Nagoya Daiichi Hospital, Nagoya, Japan
  2. Laboratory of Computational Life Science, National Cancer Center Research Institute, Tokyo, Japan
  3. Department of Computational and Systems Biology, Division of Biological Data Science, Medical Research Laboratory, Institute for Integrated Research, Institute of Science Tokyo, Tokyo, Japan
Journal: Bioinformatics (Oxford, England), volume 42, issue 9, article btag652
Dates: received 19 November 2025; accepted 19 August 2026; published online 31 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1093/bioinformatics/btag652 · PMID 42675615 · PMCID PMC13585172 · OpenAlex W4405659442
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), mouse (organism)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning
MeSH: Single-Cell Analysis*, Software*, Animals, Brain, Chromatin, Humans, Mice, Multiomics, Transcriptome (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: JSPS KAKENHI (JP23H04938, JP26K03026, JP22H04925); Japan Agency for Medical Research and Development (AMED) (JP26gm2010002, JP26zf0127012, JP26wm0325068, JP256f0137003, JP26wm0625519, JP26nk0101112, JP26tm0424226); Japan Society for the Promotion of Science (JSPS); Medical Research Center Initiative for High Depth Omics at Institute of Science Tokyo; Multilayered Stress Diseases (JPMXP1323015483)
Citations: not cited yet (Europe PMC); 101 references in the paper

Abstract

Motivation: Single-cell multiomics reveals regulatory relationships across biological layers but captures only static snapshots, obscuring the dynamics coordinated across modalities. RNA velocity predicts transcriptome dynamics, yet cannot be extended to other layers such as the regulome, leaving chromatin accessibility dynamics unresolved.

Results: We developed mmVelo (multimodal velocity of single cells), a deep generative model that infers cell state dynamics from spliced and unspliced mRNA and projects them onto other modalities, yielding chromatin velocity at single-peak resolution. In developing mouse brain, mmVelo accurately recovered accessibility dynamics; in mouse skin, it identified transcription factors regulating accessibility. Decomposing posterior velocity variability into manifold-aligned and off-manifold components revealed modality-specific uncertainty structure, with chromatin fluctuation elevated near lineage branching. Using multiomics data as a bridge, mmVelo inferred the dynamics of missing modalities from single-modal human brain data.

Availability and implementation: Source code is freely available under the MIT license at https://github.com/nomuhyooon/mmVelo; the version and test data used here are archived at https://doi.org/10.5281/zenodo.20103609.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 21 matches between paragraphs and lines of code.

nomuhyooon/mmVelo

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 3dbe5d18cad3e2f020df8e8b74f3068f669b3b7f, 10 June 2026
Languages: Python (81), Jupyter (3)
Size: 104 files, 84 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (mmVelo_tutorial/pyproject.toml, mmVelo_tutorial/setup.cfg, mmVelo_tutorial_v2/pyproject.toml, mmVelo_tutorial_v2/setup.cfg), 3 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (76 files), Scanpy (65 files), Matplotlib (59 files), SciPy (55 files), pandas (53 files), anndata (45 files), scVelo (40 files), seaborn (40 files), PyTorch (35 files), UMAP (26 files), PyTorch Lightning (21 files), statsmodels (20 files), scikit-learn (19 files), Pyro (2 files), pysam (2 files), BEDTools (1 file), Pillow (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
85 files

Zenodo 20103609

License: CC-BY-4.0
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (76 files), Scanpy (65 files), Matplotlib (59 files), SciPy (55 files), pandas (53 files), anndata (45 files), scVelo (40 files), seaborn (40 files), PyTorch (35 files), UMAP (26 files), PyTorch Lightning (21 files), statsmodels (20 files), scikit-learn (19 files), Pyro (2 files), pysam (2 files), BEDTools (1 file), Pillow (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers (HTTP 200)
  • 26 September 2026: the link answers (HTTP 200)
85 files

Availability and implementation

Source code is freely available under the MIT license at https://github.com/nomuhyooon/mmVelo; the version and test data used here are archived at https://doi.org/10.5281/zenodo.20103609.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 168 scripts, each with its path and the digest of its content;
  • 21 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The 10× embryonic mouse brain dataset was downloaded from the 10× website at https://www.10xgenomics.com/datasets/fresh-embryonic-e-18-mouse-brain-5-k-1-standard-2-0-0. The SHARE-seq mouse skin dataset (Ma et al. 2020c) was downloaded from GEO (GSE140203 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE140203)). The 10× human cortical development dataset was downloaded from GEO (GSE162170 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE162170)). The full source code is publicly available at https://github.com/nomuhyooon/mmVelo, and an archived version corresponding to this study is available at https://doi.org/10.5281/zenodo.20103609. Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 9 MeSH terms, 5 funders, 98 references.

Cite

This paper

Nomura, S., Kojima, Y., Minoura, K., Hayashi, S., Abe, K., Hirose, H., & Shimamura, T. (2026). mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities. Bioinformatics (Oxford, England), 42(9), btag652. https://doi.org/10.1093/bioinformatics/btag652

BibTeX

@article{nomura2026mmvelo,
author = {Nomura, Satoshi and Kojima, Yasuhiro and Minoura, Kodai and Hayashi, Shuto and Abe, Ko and Hirose, Haruka and Shimamura, Teppei},
title = {{mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities}},
journal = {Bioinformatics (Oxford, England)},
year = {2026},
month = aug,
volume = {42},
number = {9},
pages = {btag652},
publisher = {Oxford University Press},
issn = {1367-4803},
doi = {10.1093/bioinformatics/btag652},
url = {https://doi.org/10.1093/bioinformatics/btag652},
pmid = {42675615},
pmcid = {PMC13585172}
}

RIS

TY - JOUR
AU - Nomura, Satoshi
AU - Kojima, Yasuhiro
AU - Minoura, Kodai
AU - Hayashi, Shuto
AU - Abe, Ko
AU - Hirose, Haruka
AU - Shimamura, Teppei
TI - mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities
T2 - Bioinformatics (Oxford, England)
J2 - Bioinformatics
PY - 2026
DA - 2026/08/01
VL - 42
IS - 9
SP - btag652
SN - 1367-4803
PB - Oxford University Press
DO - 10.1093/bioinformatics/btag652
UR - https://doi.org/10.1093/bioinformatics/btag652
LA - en
ER -

CSL-JSON

{
"id": "10.1093/bioinformatics/btag652",
"type": "article-journal",
"title": "mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities",
"container-title": "Bioinformatics (Oxford, England)",
"author": [
{
"family": "Nomura",
"given": "Satoshi"
},
{
"family": "Kojima",
"given": "Yasuhiro"
},
{
"family": "Minoura",
"given": "Kodai"
},
{
"family": "Hayashi",
"given": "Shuto"
},
{
"family": "Abe",
"given": "Ko"
},
{
"family": "Hirose",
"given": "Haruka"
},
{
"family": "Shimamura",
"given": "Teppei"
}
],
"container-title-short": "Bioinformatics",
"volume": "42",
"issue": "9",
"page": "btag652",
"DOI": "10.1093/bioinformatics/btag652",
"PMID": "42675615",
"PMCID": "PMC13585172",
"ISSN": "1367-4803",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/bioinformatics/btag652",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.crmeth.2026.101342 [code]
Interpretable learning of temporal cellular dynamics from single-cell data.
Journal: Cell reports methods
In common: scVelo, UMAP, anndata, 9 other tools, genetics / omics, 9 references
[2] doi:10.7554/elife.108950 [code]
Comprehensive RNA velocity by modeling the cascade of gene regulation, transcription, and splicing from single-cell RNA sequencing data with TSvelo.
Journal: eLife
In common: scVelo, anndata, Scanpy, 6 other tools, genetics / omics, mouse, 10 references
[3] doi:10.1038/s44320-026-00208-7 [code]
Interpretable deep generative ensemble learning for single-cell omics with Hydra.
Journal: Molecular systems biology
In common: pysam, UMAP, anndata, 8 other tools, 9 references
[4] doi:10.1038/s41467-026-74000-4 [code]
ArchVelo: archetypal velocity modeling for single-cell multi-omic trajectories.
Journal: Nature communications
In common: scVelo, anndata, Scanpy, 7 other tools, genetics / omics, mouse, 8 references
[5] doi:10.1002/advs.77986 [code]
DUET-seq: An Open-Source Droplet Platform for High-Fidelity Joint Chromatin and Transcriptome Profiling Reveals Temporal Regulatory Decoupling in Single Cells.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: scVelo, Scanpy, seaborn, 4 other tools, genetics / omics, 10 references
[6] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: scVelo, pysam, UMAP, 9 other tools, genetics / omics, 6 references
[7] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: pysam, BEDTools, UMAP, 11 other tools, genetics / omics, mouse, 3 references
[8] doi:10.1038/s41467-026-74694-6 [code]
Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Journal: Nature communications
In common: Pyro, anndata, Scanpy, 8 other tools, genetics / omics, 5 references
[9] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: pysam, BEDTools, UMAP, 9 other tools, genetics / omics, mouse, 3 references
[10] doi:10.1038/s42003-026-10462-y [code]
SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.
Journal: Communications biology
In common: pysam, BEDTools, anndata, 7 other tools, genetics / omics, mouse, 5 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.