OSCR

Similarities between <i>Ciona</i> Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?

Code ↔ Paper

1 match between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 1 match
  1. [1] § Materials and Methods › Single-cell RNA sequencing analysis ↔ src/scvi/module/_vae.py, lines 750–893 · score 0.69 · linearly decoded variational, auto encoder, LDVAE, gene expression, space, mapped

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 893 lines · 34 KB · BSD-3-Clause · 1 match

  1. from __future__ import annotations
  2. import logging
  3. import warnings
  4. from typing import TYPE_CHECKING
  5. import numpy as np
  6. import torch
  7. from torch.nn.functional import one_hot
  8. from scvi import REGISTRY_KEYS, settings
  9. from scvi.data._constants import ADATA_MINIFY_TYPE
  10. from scvi.distributions._utils import _needs_cpu_detour
  11. from scvi.module._constants import MODULE_KEYS
  12. from scvi.module.base import (
  13. BaseMinifiedModeModuleClass,
  14. EmbeddingModuleMixin,
  15. LossOutput,
  16. auto_move_data,
  17. )
  18. from scvi.utils import unsupported_if_adata_minified
  19. if TYPE_CHECKING:
  20. from collections.abc import Callable
  21. from typing import Literal
  22. from torch.distributions import Distribution
  23. logger = logging.getLogger(__name__)
  24. class VAE(EmbeddingModuleMixin, BaseMinifiedModeModuleClass):
  25. """Variational auto-encoder :cite:p:`Lopez18`.
  26. Parameters
  27. ----------
  28. n_input
  29. Number of input features.
  30. n_batch
  31. Number of batches. If ``0``, no batch correction is performed.
  32. n_labels
  33. Number of labels.
  34. n_hidden
  35. Number of nodes per hidden layer. Passed into :class:`~scvi.nn.Encoder` and
  36. :class:`~scvi.nn.DecoderSCVI`.
  37. n_latent
  38. Dimensionality of the latent space.
  39. n_layers
  40. Number of hidden layers. Passed into :class:`~scvi.nn.Encoder` and
  41. :class:`~scvi.nn.DecoderSCVI`.
  42. n_continuous_cov
  43. Number of continuous covariates.
  44. n_cats_per_cov
  45. A list of integers containing the number of categories for each categorical covariate.
  46. dropout_rate
  47. Dropout rate. Passed into :class:`~scvi.nn.Encoder` but not :class:`~scvi.nn.DecoderSCVI`.
  48. dispersion
  49. Flexibility of the dispersion parameter when ``gene_likelihood`` is either ``"nb"`` or
  50. ``"zinb"``. One of the following:
  51. * ``"gene"``: parameter is constant per gene across cells.
  52. * ``"gene-batch"``: parameter is constant per gene per batch.
  53. * ``"gene-label"``: parameter is constant per gene per label.
  54. * ``"gene-cell"``: parameter is constant per gene per cell.
  55. log_variational
  56. If ``True``, use :func:`~torch.log1p` on input data before encoding for numerical stability
  57. (not normalization).
  58. gene_likelihood
  59. Distribution to use for reconstruction in the generative process. One of the following:
  60. * ``"nb"``: :class:`~scvi.distributions.NegativeBinomial`.
  61. * ``"zinb"``: :class:`~scvi.distributions.ZeroInflatedNegativeBinomial`.
  62. * ``"poisson"``: :class:`~scvi.distributions.Poisson`.
  63. * ``"normal"``: :class:`~torch.distributions.normal.Normal`.
  64. latent_distribution
  65. Distribution to use for the latent space. One of the following:
  66. * ``"normal"``: isotropic normal.
  67. * ``"ln"``: logistic normal with normal params N(0, 1).
  68. encode_covariates
  69. If ``True``, covariates are concatenated to gene expression prior to passing through
  70. the encoder(s). Else, only the gene expression is used.
  71. deeply_inject_covariates
  72. If ``True`` and ``n_layers > 1``, covariates are concatenated to the outputs of hidden
  73. layers in the encoder(s) (if ``encoder_covariates`` is ``True``) and the decoder prior to
  74. passing through the next layer.
  75. batch_representation
  76. Method for encoding batch information. One of the following:
  77. * ``"one-hot"``: represent batches with one-hot encodings.
  78. * ``"embedding"``: represent batches with continuously-valued embeddings using
  79. :class:`~scvi.nn.Embedding`.
  80. Note that batch representations are only passed into the encoder(s) if
  81. ``encode_covariates`` is ``True``.
  82. use_batch_norm
  83. Specifies where to use :class:`~torch.nn.BatchNorm1d` in the model. One of the following:
  84. * ``"none"``: don't use batch norm in either encoder(s) or decoder.
  85. * ``"encoder"``: use batch norm only in the encoder(s).
  86. * ``"decoder"``: use batch norm only in the decoder.
  87. * ``"both"``: use batch norm in both encoder(s) and decoder.
  88. Note: if ``use_layer_norm`` is also specified, both will be applied (first
  89. :class:`~torch.nn.BatchNorm1d`, then :class:`~torch.nn.LayerNorm`).
  90. use_layer_norm
  91. Specifies where to use :class:`~torch.nn.LayerNorm` in the model. One of the following:
  92. * ``"none"``: don't use layer norm in either encoder(s) or decoder.
  93. * ``"encoder"``: use layer norm only in the encoder(s).
  94. * ``"decoder"``: use layer norm only in the decoder.
  95. * ``"both"``: use layer norm in both encoder(s) and decoder.
  96. Note: if ``use_batch_norm`` is also specified, both will be applied (first
  97. :class:`~torch.nn.BatchNorm1d`, then :class:`~torch.nn.LayerNorm`).
  98. use_size_factor_key
  99. If ``True``, use the :attr:`~anndata.AnnData.obs` column as defined by the
  100. ``size_factor_key`` parameter in the model's ``setup_anndata`` method as the scaling
  101. factor in the mean of the conditional distribution. Takes priority over
  102. ``use_observed_lib_size``.
  103. use_observed_lib_size
  104. If ``True``, use the observed library size for RNA as the scaling factor in the mean of the
  105. conditional distribution.
  106. extra_payload_autotune
  107. If ``True``, will return extra matrices in the loss output to be used during autotune
  108. library_log_means
  109. :class:`~numpy.ndarray` of shape ``(1, n_batch)`` of means of the log library sizes that
  110. parameterize the prior on library size if ``use_size_factor_key`` is ``False`` and
  111. ``use_observed_lib_size`` is ``False``.
  112. library_log_vars
  113. :class:`~numpy.ndarray` of shape ``(1, n_batch)`` of variances of the log library sizes
  114. that parameterize the prior on library size if ``use_size_factor_key`` is ``False`` and
  115. ``use_observed_lib_size`` is ``False``.
  116. var_activation
  117. Callable used to ensure positivity of the variance of the variational distribution. Passed
  118. into :class:`~scvi.nn.Encoder`. Defaults to :func:`~torch.exp`.
  119. extra_encoder_kwargs
  120. Additional keyword arguments passed into :class:`~scvi.nn.Encoder`.
  121. extra_decoder_kwargs
  122. Additional keyword arguments passed into :class:`~scvi.nn.DecoderSCVI`.
  123. batch_embedding_kwargs
  124. Keyword arguments passed into :class:`~scvi.nn.Embedding` if ``batch_representation`` is
  125. set to ``"embedding"``.
  126. """
  127. def __init__(
  128. self,
  129. n_input: int,
  130. n_batch: int = 0,
  131. n_labels: int = 0,
  132. n_hidden: int = 128,
  133. n_latent: int = 10,
  134. n_layers: int = 1,
  135. n_continuous_cov: int = 0,
  136. n_cats_per_cov: list[int] | None = None,
  137. dropout_rate: float = 0.1,
  138. dispersion: Literal["gene", "gene-batch", "gene-label", "gene-cell"] = "gene",
  139. log_variational: bool = True,
  140. gene_likelihood: Literal["zinb", "nb", "poisson", "normal"] = "zinb",
  141. latent_distribution: Literal["normal", "ln"] = "normal",
  142. encode_covariates: bool = False,
  143. deeply_inject_covariates: bool = True,
  144. batch_representation: Literal["one-hot", "embedding"] = "one-hot",
  145. use_batch_norm: Literal["encoder", "decoder", "none", "both"] = "both",
  146. use_layer_norm: Literal["encoder", "decoder", "none", "both"] = "none",
  147. use_size_factor_key: bool = False,
  148. use_observed_lib_size: bool = True,
  149. extra_payload_autotune: bool = False,
  150. library_log_means: np.ndarray | None = None,
  151. library_log_vars: np.ndarray | None = None,
  152. var_activation: Callable[[torch.Tensor], torch.Tensor] = None,
  153. extra_encoder_kwargs: dict | None = None,
  154. extra_decoder_kwargs: dict | None = None,
  155. batch_embedding_kwargs: dict | None = None,
  156. ):
  157. from scvi.nn import DecoderSCVI, Encoder
  158. super().__init__()
  159. self.dispersion = dispersion
  160. self.n_latent = n_latent
  161. self.log_variational = log_variational
  162. self.gene_likelihood = gene_likelihood
  163. self.n_batch = n_batch
  164. self.n_input = n_input
  165. self.n_labels = n_labels
  166. self.n_hidden = n_hidden
  167. self.n_layers = n_layers
  168. self.latent_distribution = latent_distribution
  169. self.encode_covariates = encode_covariates
  170. self.use_size_factor_key = use_size_factor_key
  171. self.use_observed_lib_size = use_size_factor_key or use_observed_lib_size
  172. self.extra_payload_autotune = extra_payload_autotune
  173. if not self.use_observed_lib_size:
  174. if library_log_means is None or library_log_vars is None:
  175. raise ValueError(
  176. "If not using observed_lib_size, "
  177. "must provide library_log_means and library_log_vars."
  178. )
  179. self.register_buffer("library_log_means", torch.from_numpy(library_log_means).float())
  180. self.register_buffer("library_log_vars", torch.from_numpy(library_log_vars).float())
  181. if self.dispersion == "gene":
  182. self.px_r = torch.nn.Parameter(torch.randn(n_input))
  183. elif self.dispersion == "gene-batch":
  184. self.px_r = torch.nn.Parameter(torch.randn(n_input, n_batch))
  185. elif self.dispersion == "gene-label":
  186. self.px_r = torch.nn.Parameter(torch.randn(n_input, n_labels))
  187. elif self.dispersion == "gene-cell":
  188. pass
  189. else:
  190. raise ValueError(
  191. "`dispersion` must be one of 'gene', 'gene-batch', 'gene-label', 'gene-cell'."
  192. )
  193. self.batch_representation = batch_representation
  194. if self.batch_representation == "embedding":
  195. self.init_embedding(REGISTRY_KEYS.BATCH_KEY, n_batch, **(batch_embedding_kwargs or {}))
  196. batch_dim = self.get_embedding(REGISTRY_KEYS.BATCH_KEY).embedding_dim
  197. elif self.batch_representation != "one-hot":
  198. raise ValueError("`batch_representation` must be one of 'one-hot', 'embedding'.")
  199. use_batch_norm_encoder = use_batch_norm == "encoder" or use_batch_norm == "both"
  200. use_batch_norm_decoder = use_batch_norm == "decoder" or use_batch_norm == "both"
  201. use_layer_norm_encoder = use_layer_norm == "encoder" or use_layer_norm == "both"
  202. use_layer_norm_decoder = use_layer_norm == "decoder" or use_layer_norm == "both"
  203. n_input_encoder = n_input + n_continuous_cov * encode_covariates
  204. if self.batch_representation == "embedding":
  205. n_input_encoder += batch_dim * encode_covariates
  206. cat_list = list([] if n_cats_per_cov is None else n_cats_per_cov)
  207. else:
  208. cat_list = [n_batch] + list([] if n_cats_per_cov is None else n_cats_per_cov)
  209. encoder_cat_list = cat_list if encode_covariates else None
  210. _extra_encoder_kwargs = extra_encoder_kwargs or {}
  211. self.z_encoder = Encoder(
  212. n_input_encoder,
  213. n_latent,
  214. n_cat_list=encoder_cat_list,
  215. n_layers=n_layers,
  216. n_hidden=n_hidden,
  217. dropout_rate=dropout_rate,
  218. distribution=latent_distribution,
  219. inject_covariates=deeply_inject_covariates,
  220. use_batch_norm=use_batch_norm_encoder,
  221. use_layer_norm=use_layer_norm_encoder,
  222. var_activation=var_activation,
  223. return_dist=True,
  224. **_extra_encoder_kwargs,
  225. )
  226. # l encoder goes from n_input-dimensional data to 1-d library size
  227. self.l_encoder = Encoder(
  228. n_input_encoder,
  229. 1,
  230. n_layers=1,
  231. n_cat_list=encoder_cat_list,
  232. n_hidden=n_hidden,
  233. dropout_rate=dropout_rate,
  234. inject_covariates=deeply_inject_covariates,
  235. use_batch_norm=use_batch_norm_encoder,
  236. use_layer_norm=use_layer_norm_encoder,
  237. var_activation=var_activation,
  238. return_dist=True,
  239. **_extra_encoder_kwargs,
  240. )
  241. n_input_decoder = n_latent + n_continuous_cov
  242. if self.batch_representation == "embedding":
  243. n_input_decoder += batch_dim
  244. _extra_decoder_kwargs = extra_decoder_kwargs or {}
  245. self.decoder = DecoderSCVI(
  246. n_input_decoder,
  247. n_input,
  248. n_cat_list=cat_list,
  249. n_layers=n_layers,
  250. n_hidden=n_hidden,
  251. inject_covariates=deeply_inject_covariates,
  252. use_batch_norm=use_batch_norm_decoder,
  253. use_layer_norm=use_layer_norm_decoder,
  254. scale_activation="softplus" if use_size_factor_key else "softmax",
  255. **_extra_decoder_kwargs,
  256. )
  257. def _get_inference_input(
  258. self,
  259. tensors: dict[str, torch.Tensor | None],
  260. full_forward_pass: bool = False,
  261. ) -> dict[str, torch.Tensor | None]:
  262. """Get input tensors for the inference process."""
  263. if full_forward_pass or self.minified_data_type is None:
  264. loader = "full_data"
  265. elif self.minified_data_type in [
  266. ADATA_MINIFY_TYPE.LATENT_POSTERIOR,
  267. ADATA_MINIFY_TYPE.LATENT_POSTERIOR_WITH_COUNTS,
  268. ]:
  269. loader = "minified_data"
  270. else:
  271. raise NotImplementedError(f"Unknown minified-data type: {self.minified_data_type}")
  272. if loader == "full_data":
  273. return {
  274. MODULE_KEYS.X_KEY: tensors[REGISTRY_KEYS.X_KEY],
  275. MODULE_KEYS.BATCH_INDEX_KEY: tensors[REGISTRY_KEYS.BATCH_KEY],
  276. MODULE_KEYS.CONT_COVS_KEY: tensors.get(REGISTRY_KEYS.CONT_COVS_KEY, None),
  277. MODULE_KEYS.CAT_COVS_KEY: tensors.get(REGISTRY_KEYS.CAT_COVS_KEY, None),
  278. }
  279. else:
  280. return {
  281. MODULE_KEYS.QZM_KEY: tensors[REGISTRY_KEYS.LATENT_QZM_KEY],
  282. MODULE_KEYS.QZV_KEY: tensors[REGISTRY_KEYS.LATENT_QZV_KEY],
  283. REGISTRY_KEYS.OBSERVED_LIB_SIZE: tensors[REGISTRY_KEYS.OBSERVED_LIB_SIZE],
  284. }
  285. def _get_generative_input(
  286. self,
  287. tensors: dict[str, torch.Tensor],
  288. inference_outputs: dict[str, torch.Tensor | Distribution | None],
  289. ) -> dict[str, torch.Tensor | None]:
  290. """Get input tensors for the generative process."""
  291. size_factor = tensors.get(REGISTRY_KEYS.SIZE_FACTOR_KEY, None)
  292. if size_factor is not None:
  293. size_factor = torch.log(size_factor)
  294. return {
  295. MODULE_KEYS.Z_KEY: inference_outputs[MODULE_KEYS.Z_KEY],
  296. MODULE_KEYS.LIBRARY_KEY: inference_outputs[MODULE_KEYS.LIBRARY_KEY],
  297. MODULE_KEYS.BATCH_INDEX_KEY: tensors[REGISTRY_KEYS.BATCH_KEY],
  298. MODULE_KEYS.Y_KEY: tensors[REGISTRY_KEYS.LABELS_KEY],
  299. MODULE_KEYS.CONT_COVS_KEY: tensors.get(REGISTRY_KEYS.CONT_COVS_KEY, None),
  300. MODULE_KEYS.CAT_COVS_KEY: tensors.get(REGISTRY_KEYS.CAT_COVS_KEY, None),
  301. MODULE_KEYS.SIZE_FACTOR_KEY: size_factor,
  302. }
  303. def _compute_local_library_params(
  304. self,
  305. batch_index: torch.Tensor,
  306. ) -> tuple[torch.Tensor, torch.Tensor]:
  307. """Computes local library parameters.
  308. Compute two tensors of shape (batch_index.shape[0], 1) where each
  309. element corresponds to the mean and variances, respectively, of the
  310. log library sizes in the batch the cell corresponds to.
  311. """
  312. from torch.nn.functional import linear
  313. n_batch = self.library_log_means.shape[1]
  314. local_library_log_means = linear(
  315. one_hot(batch_index.squeeze(-1), n_batch).float(), self.library_log_means
  316. )
  317. local_library_log_vars = linear(
  318. one_hot(batch_index.squeeze(-1), n_batch).float(), self.library_log_vars
  319. )
  320. return local_library_log_means, local_library_log_vars
  321. @auto_move_data
  322. def _regular_inference(
  323. self,
  324. x: torch.Tensor,
  325. batch_index: torch.Tensor,
  326. cont_covs: torch.Tensor | None = None,
  327. cat_covs: torch.Tensor | None = None,
  328. n_samples: int = 1,
  329. ) -> dict[str, torch.Tensor | Distribution | None]:
  330. """Run the regular inference process."""
  331. x_ = x
  332. if self.use_observed_lib_size:
  333. library = torch.log(x.sum(1)).unsqueeze(1)
  334. if self.log_variational:
  335. x_ = torch.log1p(x_)
  336. if cont_covs is not None and self.encode_covariates:
  337. encoder_input = torch.cat((x_, cont_covs), dim=-1)
  338. else:
  339. encoder_input = x_
  340. if cat_covs is not None and self.encode_covariates:
  341. categorical_input = torch.split(cat_covs, 1, dim=1)
  342. else:
  343. categorical_input = ()
  344. if self.batch_representation == "embedding" and self.encode_covariates:
  345. batch_rep = self.compute_embedding(REGISTRY_KEYS.BATCH_KEY, batch_index)
  346. encoder_input = torch.cat([encoder_input, batch_rep], dim=-1)
  347. qz, z = self.z_encoder(encoder_input, *categorical_input)
  348. else:
  349. qz, z = self.z_encoder(encoder_input, batch_index, *categorical_input)
  350. ql = None
  351. if not self.use_observed_lib_size:
  352. if self.batch_representation == "embedding":
  353. ql, library_encoded = self.l_encoder(encoder_input, *categorical_input)
  354. else:
  355. ql, library_encoded = self.l_encoder(
  356. encoder_input, batch_index, *categorical_input
  357. )
  358. library = library_encoded
  359. if n_samples > 1:
  360. untran_z = qz.sample((n_samples,))
  361. z = self.z_encoder.z_transformation(untran_z)
  362. if self.use_observed_lib_size:
  363. library = library.unsqueeze(0).expand(
  364. (n_samples, library.size(0), library.size(1))
  365. )
  366. else:
  367. library = ql.sample((n_samples,))
  368. return {
  369. MODULE_KEYS.Z_KEY: z,
  370. MODULE_KEYS.QZ_KEY: qz,
  371. MODULE_KEYS.QL_KEY: ql,
  372. MODULE_KEYS.LIBRARY_KEY: library,
  373. }
  374. @auto_move_data
  375. def _cached_inference(
  376. self,
  377. qzm: torch.Tensor,
  378. qzv: torch.Tensor,
  379. observed_lib_size: torch.Tensor,
  380. n_samples: int = 1,
  381. ) -> dict[str, torch.Tensor | None]:
  382. """Run the cached inference process."""
  383. from torch.distributions import Normal
  384. qz = Normal(qzm, qzv.sqrt())
  385. # use dist.sample() rather than rsample because we aren't optimizing the z here
  386. untran_z = qz.sample() if n_samples == 1 else qz.sample((n_samples,))
  387. z = self.z_encoder.z_transformation(untran_z)
  388. library = torch.log(observed_lib_size)
  389. if n_samples > 1:
  390. library = library.unsqueeze(0).expand((n_samples, library.size(0), library.size(1)))
  391. return {
  392. MODULE_KEYS.Z_KEY: z,
  393. MODULE_KEYS.QZ_KEY: qz,
  394. MODULE_KEYS.QL_KEY: None,
  395. MODULE_KEYS.LIBRARY_KEY: library,
  396. }
  397. @auto_move_data
  398. def generative(
  399. self,
  400. z: torch.Tensor,
  401. library: torch.Tensor,
  402. batch_index: torch.Tensor,
  403. cont_covs: torch.Tensor | None = None,
  404. cat_covs: torch.Tensor | None = None,
  405. size_factor: torch.Tensor | None = None,
  406. y: torch.Tensor | None = None,
  407. transform_batch: torch.Tensor | None = None,
  408. ) -> dict[str, Distribution | None]:
  409. """Run the generative process."""
  410. from torch.nn.functional import linear
  411. from scvi.distributions import (
  412. NegativeBinomial,
  413. Normal,
  414. Poisson,
  415. ZeroInflatedNegativeBinomial,
  416. )
  417. # TODO: refactor forward function to not rely on y
  418. # Likelihood distribution
  419. if cont_covs is None:
  420. decoder_input = z
  421. elif z.dim() != cont_covs.dim():
  422. decoder_input = torch.cat(
  423. [z, cont_covs.unsqueeze(0).expand(z.size(0), -1, -1)], dim=-1
  424. )
  425. else:
  426. decoder_input = torch.cat([z, cont_covs], dim=-1)
  427. if cat_covs is not None:
  428. categorical_input = torch.split(cat_covs, 1, dim=1)
  429. else:
  430. categorical_input = ()
  431. if transform_batch is not None:
  432. batch_index = torch.ones_like(batch_index) * transform_batch
  433. if not self.use_size_factor_key:
  434. size_factor = library
  435. if self.batch_representation == "embedding":
  436. batch_rep = self.compute_embedding(REGISTRY_KEYS.BATCH_KEY, batch_index)
  437. decoder_input = torch.cat([decoder_input, batch_rep], dim=-1)
  438. px_scale, px_r, px_rate, px_dropout = self.decoder(
  439. self.dispersion,
  440. decoder_input,
  441. size_factor,
  442. *categorical_input,
  443. y,
  444. )
  445. else:
  446. px_scale, px_r, px_rate, px_dropout = self.decoder(
  447. self.dispersion,
  448. decoder_input,
  449. size_factor,
  450. batch_index,
  451. *categorical_input,
  452. y,
  453. )
  454. if self.dispersion == "gene-label":
  455. px_r = linear(
  456. one_hot(y.squeeze(-1), self.n_labels).float(), self.px_r
  457. ) # px_r gets transposed - last dimension is nb genes
  458. elif self.dispersion == "gene-batch":
  459. px_r = linear(one_hot(batch_index.squeeze(-1), self.n_batch).float(), self.px_r)
  460. elif self.dispersion == "gene":
  461. px_r = self.px_r
  462. px_r = torch.exp(px_r)
  463. if self.gene_likelihood == "zinb":
  464. px = ZeroInflatedNegativeBinomial(
  465. mu=px_rate,
  466. theta=px_r,
  467. zi_logits=px_dropout,
  468. scale=px_scale,
  469. )
  470. elif self.gene_likelihood == "nb":
  471. px = NegativeBinomial(mu=px_rate, theta=px_r, scale=px_scale)
  472. elif self.gene_likelihood == "poisson":
  473. px = Poisson(rate=px_rate, scale=px_scale)
  474. elif self.gene_likelihood == "normal":
  475. px = Normal(px_rate, px_r, normal_mu=px_scale)
  476. # Priors
  477. if self.use_observed_lib_size:
  478. pl = None
  479. else:
  480. (
  481. local_library_log_means,
  482. local_library_log_vars,
  483. ) = self._compute_local_library_params(batch_index)
  484. pl = Normal(local_library_log_means, local_library_log_vars.sqrt())
  485. pz = Normal(torch.zeros_like(z), torch.ones_like(z))
  486. return {
  487. MODULE_KEYS.PX_KEY: px,
  488. MODULE_KEYS.PL_KEY: pl,
  489. MODULE_KEYS.PZ_KEY: pz,
  490. }
  491. @unsupported_if_adata_minified
  492. def loss(
  493. self,
  494. tensors: dict[str, torch.Tensor],
  495. inference_outputs: dict[str, torch.Tensor | Distribution | None],
  496. generative_outputs: dict[str, Distribution | None],
  497. kl_weight: torch.Tensor | float = 1.0,
  498. ) -> LossOutput:
  499. """Compute the loss."""
  500. from torch.distributions import kl_divergence
  501. x = tensors[REGISTRY_KEYS.X_KEY]
  502. kl_divergence_z = kl_divergence(
  503. inference_outputs[MODULE_KEYS.QZ_KEY], generative_outputs[MODULE_KEYS.PZ_KEY]
  504. ).sum(dim=-1)
  505. if not self.use_observed_lib_size:
  506. kl_divergence_l = kl_divergence(
  507. inference_outputs[MODULE_KEYS.QL_KEY], generative_outputs[MODULE_KEYS.PL_KEY]
  508. ).sum(dim=1)
  509. else:
  510. kl_divergence_l = torch.zeros_like(kl_divergence_z)
  511. reconst_loss = -generative_outputs[MODULE_KEYS.PX_KEY].log_prob(x).sum(-1)
  512. kl_local_for_warmup = kl_divergence_z
  513. kl_local_no_warmup = kl_divergence_l
  514. weighted_kl_local = kl_weight * kl_local_for_warmup + kl_local_no_warmup
  515. loss = torch.mean(reconst_loss + weighted_kl_local)
  516. # a payload to be used during autotune
  517. if self.extra_payload_autotune:
  518. extra_metrics_payload = {
  519. "z": inference_outputs["z"],
  520. "batch": tensors[REGISTRY_KEYS.BATCH_KEY],
  521. "labels": tensors[REGISTRY_KEYS.LABELS_KEY],
  522. }
  523. else:
  524. extra_metrics_payload = {}
  525. return LossOutput(
  526. loss=loss,
  527. reconstruction_loss=reconst_loss,
  528. kl_local={
  529. MODULE_KEYS.KL_L_KEY: kl_divergence_l,
  530. MODULE_KEYS.KL_Z_KEY: kl_divergence_z,
  531. },
  532. extra_metrics=extra_metrics_payload,
  533. )
  534. @torch.inference_mode()
  535. def sample(
  536. self,
  537. tensors: dict[str, torch.Tensor],
  538. n_samples: int = 1,
  539. max_poisson_rate: float = 1e8,
  540. generative_kwargs: dict | None = None,
  541. ) -> torch.Tensor:
  542. r"""Generate predictive samples from the posterior predictive distribution.
  543. The posterior predictive distribution is denoted as :math:`p(\hat{x} \mid x)`, where
  544. :math:`x` is the input data and :math:`\hat{x}` is the sampled data.
  545. We sample from this distribution by first sampling ``n_samples`` times from the posterior
  546. distribution :math:`q(z \mid x)` for a given observation, and then sampling from the
  547. likelihood :math:`p(\hat{x} \mid z)` for each of these.
  548. Parameters
  549. ----------
  550. tensors
  551. Dictionary of tensors passed into ``VAE.forward``.
  552. n_samples
  553. Number of Monte Carlo samples to draw from the distribution for each observation.
  554. max_poisson_rate
  555. The maximum value to which to clip the ``rate`` parameter of
  556. :class:`~scvi.distributions.Poisson`. Avoids numerical sampling issues when the
  557. parameter is very large due to the variance of the distribution.
  558. generative_kwargs
  559. Keyword args for ``generative()`` in fwd pass
  560. Returns
  561. -------
  562. Tensor on CPU with shape ``(n_obs, n_vars)`` if ``n_samples == 1``, else
  563. ``(n_obs, n_vars,)``.
  564. """
  565. from scvi.distributions import Poisson
  566. inference_kwargs = {"n_samples": n_samples}
  567. _, generative_outputs = self.forward(
  568. tensors,
  569. inference_kwargs=inference_kwargs,
  570. generative_kwargs=generative_kwargs,
  571. compute_loss=False,
  572. )
  573. dist = generative_outputs[MODULE_KEYS.PX_KEY]
  574. if self.gene_likelihood == "poisson":
  575. on_mps = self.device.type == "mps"
  576. rate = torch.clamp(dist.rate, max=max_poisson_rate)
  577. if _needs_cpu_detour(on_mps, torch.poisson):
  578. rate = rate.to("cpu")
  579. dist = Poisson(rate)
  580. # (n_obs, n_vars) if n_samples == 1, else (n_samples, n_obs, n_vars)
  581. samples = dist.sample()
  582. # (n_samples, n_obs, n_vars) -> (n_obs, n_vars, n_samples)
  583. samples = torch.permute(samples, (1, 2, 0)) if n_samples > 1 else samples
  584. return samples.cpu()
  585. @torch.inference_mode()
  586. @auto_move_data
  587. def marginal_ll(
  588. self,
  589. tensors: dict[str, torch.Tensor],
  590. n_mc_samples: int,
  591. return_mean: bool = False,
  592. n_mc_samples_per_pass: int = 1,
  593. ):
  594. """Compute the marginal log-likelihood of the data under the model.
  595. Parameters
  596. ----------
  597. tensors
  598. Dictionary of tensors passed into ``VAE.forward``.
  599. n_mc_samples
  600. Number of Monte Carlo samples to use for the estimation of the marginal log-likelihood.
  601. return_mean
  602. Whether to return the mean of marginal likelihoods over cells.
  603. n_mc_samples_per_pass
  604. Number of Monte Carlo samples to use per pass. This is useful to avoid memory issues.
  605. """
  606. from torch import logsumexp
  607. from torch.distributions import Normal
  608. batch_index = tensors[REGISTRY_KEYS.BATCH_KEY]
  609. to_sum = []
  610. if n_mc_samples_per_pass > n_mc_samples:
  611. warnings.warn(
  612. "Number of chunks is larger than the total number of samples, setting it to the "
  613. "number of samples",
  614. RuntimeWarning,
  615. stacklevel=settings.warnings_stacklevel,
  616. )
  617. n_mc_samples_per_pass = n_mc_samples
  618. n_passes = int(np.ceil(n_mc_samples / n_mc_samples_per_pass))
  619. for _ in range(n_passes):
  620. # Distribution parameters and sampled variables
  621. inference_outputs, _, losses = self.forward(
  622. tensors,
  623. inference_kwargs={"n_samples": n_mc_samples_per_pass},
  624. get_inference_input_kwargs={"full_forward_pass": True},
  625. )
  626. qz = inference_outputs[MODULE_KEYS.QZ_KEY]
  627. ql = inference_outputs[MODULE_KEYS.QL_KEY]
  628. z = inference_outputs[MODULE_KEYS.Z_KEY]
  629. library = inference_outputs[MODULE_KEYS.LIBRARY_KEY]
  630. # Reconstruction Loss
  631. reconst_loss = losses.dict_sum(losses.reconstruction_loss)
  632. # Log-probabilities
  633. p_z = (
  634. Normal(torch.zeros_like(qz.loc), torch.ones_like(qz.scale)).log_prob(z).sum(dim=-1)
  635. )
  636. p_x_zl = -reconst_loss
  637. q_z_x = qz.log_prob(z).sum(dim=-1)
  638. log_prob_sum = p_z + p_x_zl - q_z_x
  639. if not self.use_observed_lib_size:
  640. (
  641. local_library_log_means,
  642. local_library_log_vars,
  643. ) = self._compute_local_library_params(batch_index)
  644. p_l = (
  645. Normal(local_library_log_means, local_library_log_vars.sqrt())
  646. .log_prob(library)
  647. .sum(dim=-1)
  648. )
  649. q_l_x = ql.log_prob(library).sum(dim=-1)
  650. log_prob_sum += p_l - q_l_x
  651. if n_mc_samples_per_pass == 1:
  652. log_prob_sum = log_prob_sum.unsqueeze(0)
  653. to_sum.append(log_prob_sum)
  654. to_sum = torch.cat(to_sum, dim=0)
  655. batch_log_lkl = logsumexp(to_sum, dim=0) - np.log(n_mc_samples)
  656. if return_mean:
  657. batch_log_lkl = torch.mean(batch_log_lkl).item()
  658. else:
  659. batch_log_lkl = batch_log_lkl.cpu()
  660. return batch_log_lkl
  661. class LDVAE(VAE):
  662. """Linear-decoded Variational auto-encoder model.
  663. Implementation of :cite:p:`Svensson20`.
  664. This model uses a linear decoder, directly mapping the latent representation
  665. to gene expression levels. It still uses a deep neural network to encode
  666. the latent representation.
  667. Compared to standard VAE, this model is less powerful but can be used to
  668. inspect which genes contribute to variation in the dataset. It may also be used
  669. for all scVI tasks, like differential expression, batch correction, imputation, etc.
  670. However, batch correction may be less powerful as it assumes a linear model.
  671. Parameters
  672. ----------
  673. n_input
  674. Number of input genes
  675. n_batch
  676. Number of batches
  677. n_labels
  678. Number of labels
  679. n_hidden
  680. Number of nodes per hidden layer (for encoder)
  681. n_latent
  682. Dimensionality of the latent space
  683. n_layers_encoder
  684. Number of hidden layers used for encoder NNs
  685. dropout_rate
  686. Dropout rate for neural networks
  687. dispersion
  688. One of the following
  689. * ``'gene'`` - dispersion parameter of NB is constant per gene across cells
  690. * ``'gene-batch'`` - dispersion can differ between different batches
  691. * ``'gene-label'`` - dispersion can differ between different labels
  692. * ``'gene-cell'`` - dispersion can differ for every gene in every cell
  693. log_variational
  694. Log(data+1) prior to encoding for numerical stability. Not normalization.
  695. gene_likelihood
  696. One of
  697. * ``'nb'`` - Negative binomial distribution
  698. * ``'zinb'`` - Zero-inflated negative binomial distribution
  699. * ``'poisson'`` - Poisson distribution
  700. use_batch_norm
  701. Bool whether to use batch norm in decoder
  702. bias
  703. Bool whether to have bias term in linear decoder
  704. latent_distribution
  705. One of
  706. * ``'normal'`` - Isotropic normal
  707. * ``'ln'`` - Logistic normal with normal params N(0, 1)
  708. use_observed_lib_size
  709. Use observed library size for RNA as scaling factor in mean of conditional distribution.
  710. **kwargs
  711. """
  712. def __init__(
  713. self,
  714. n_input: int,
  715. n_batch: int = 0,
  716. n_labels: int = 0,
  717. n_hidden: int = 128,
  718. n_latent: int = 10,
  719. n_layers_encoder: int = 1,
  720. dropout_rate: float = 0.1,
  721. dispersion: str = "gene",
  722. log_variational: bool = True,
  723. gene_likelihood: str = "nb",
  724. use_batch_norm: bool = True,
  725. bias: bool = False,
  726. latent_distribution: str = "normal",
  727. use_observed_lib_size: bool = False,
  728. **kwargs,
  729. ):
  730. from scvi.nn import Encoder, LinearDecoderSCVI
  731. super().__init__(
  732. n_input=n_input,
  733. n_batch=n_batch,
  734. n_labels=n_labels,
  735. n_hidden=n_hidden,
  736. n_latent=n_latent,
  737. n_layers=n_layers_encoder,
  738. dropout_rate=dropout_rate,
  739. dispersion=dispersion,
  740. log_variational=log_variational,
  741. gene_likelihood=gene_likelihood,
  742. latent_distribution=latent_distribution,
  743. use_observed_lib_size=use_observed_lib_size,
  744. **kwargs,
  745. )
  746. self.use_batch_norm = use_batch_norm
  747. self.z_encoder = Encoder(
  748. n_input,
  749. n_latent,
  750. n_layers=n_layers_encoder,
  751. n_hidden=n_hidden,
  752. dropout_rate=dropout_rate,
  753. distribution=latent_distribution,
  754. use_batch_norm=True,
  755. use_layer_norm=False,
  756. return_dist=True,
  757. )
  758. self.l_encoder = Encoder(
  759. n_input,
  760. 1,
  761. n_layers=1,
  762. n_hidden=n_hidden,
  763. dropout_rate=dropout_rate,
  764. use_batch_norm=True,
  765. use_layer_norm=False,
  766. return_dist=True,
  767. )
  768. self.decoder = LinearDecoderSCVI(
  769. n_latent,
  770. n_input,
  771. n_cat_list=[n_batch],
  772. use_batch_norm=use_batch_norm,
  773. use_layer_norm=False,
  774. bias=bias,
  775. )
  776. @torch.inference_mode()
  777. def get_loadings(self) -> np.ndarray:
  778. """Extract per-gene weights in the linear decoder."""
  779. # This is BW, where B is diag(b) batch norm, W is the weight matrix
  780. if self.use_batch_norm is True:
  781. w = self.decoder.factor_regressor.fc_layers[0][0].weight
  782. bn = self.decoder.factor_regressor.fc_layers[0][1]
  783. sigma = torch.sqrt(bn.running_var + bn.eps)
  784. gamma = bn.weight
  785. b = gamma / sigma
  786. b_identity = torch.diag(b)
  787. loadings = torch.matmul(b_identity, w)
  788. else:
  789. loadings = self.decoder.factor_regressor.fc_layers[0][0].weight
  790. loadings = loadings.detach().cpu().numpy()
  791. if self.n_batch > 1:
  792. loadings = loadings[:, : -self.n_batch]
  793. return loadings

_vae.py at commit fa9470f, under BSD-3-Clause · at the source

Overview

Authors: Matthew J. Kourakis1, Yishen Miao2, Erin D. Newman-Smith1, Kerrianne Ryan3, William C. Smith1,2
ORCID iDs: Yishen Miao
  1. Neuroscience Research Institute, University of California Santa Barbara, Santa Barbara, California 93106
  2. Department of Molecular, Cellular and Developmental Biology, University of California, Santa Barbara, California 93106
  3. Life Sciences Centre, Dalhousie University, Halifax, Nova Scotia B3H 1A5, Canada
Institutions: University of California, Santa Barbara (United States); Dalhousie University (Canada)
Journal: eNeuro, volume 13, issue 8, pages ENEURO.0362-25.2026
Dates: received 30 September 2025; accepted 18 July 2026; published online 19 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1523/eneuro.0362-25.2026 · PMID 42637546 · PMCID PMC13509005 · OpenAlex W7204162582
Open access: gold, a free copy (OpenAlex)
Status: code verified
Methods: Evoked potentials, Statistics, Single-unit activity, calcium imaging
Keywords: cerebellum, chordate, Ciona, evolution, hindbrain
MeSH: Biological Evolution*, Cerebellum*, Ciona*, Ganglia, Spinal*, Rhombencephalon*, Animals, Ciona intestinalis (* major topic)
Journal subjects: Theory/New Concepts, Development
Topic: Developmental Biology and Gene Regulation (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Instutes of Health (NS127106); U.S. Department of Energy (DOE) (DE-SC0021978)
Citations: not cited yet (Europe PMC); 107 references in the paper

Abstract

The cerebellum, a major structure of the vertebrate central nervous system, forms the dorsal region of the hindbrain. While evidence from connectomics, neuron classification, and gene expression analyses suggests that the major CNS divisions of forebrain, midbrain, hindbrain, and spinal cord predate the vertebrate/tunicate divergence, it is unknown if the subdivision of the hindbrain into a cerebellum, or cerebellum precursor structure, does as well. In fact, conflicting views exist over whether even the most basal vertebrates, the cyclostomes, have a cerebellum homolog. We report here that the motor ganglion (MG), a structure that has been homologized to the vertebrate hindbrain in the chordate Ciona (a tunicate), has functionally distinct dorsal and ventral domains, similar to the generalized hindbrain of vertebrates. In Ciona, the dorsal MG functions as an ascending relay and processing center for peripheral sensory neurons, while the ventral domain processes and relays descending sensorimotor commands. These differences are reflected in distinct neuron classes present in the two domains, as well as in gene expression. While the dorsal MG of Ciona and HB of vertebrates differ in function and complexity, they share a number of similarities. Thus, while it is likely that the cerebellum proper evolved within vertebrates, our findings support the hypothesis that a simpler, dorsally positioned antecedent was present before the tunicate–vertebrate divergence.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.

scverse/scvi-tools

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: fa9470f9b59f47a1df1677eda44161ac14845db3, 24 September 2026
Languages: Python (342), JavaScript (1)
Size: 559 files, 343 scripts
Software Heritage: archived
Found in: the Zenodo archive record
Holds: README, license file, environment (Dockerfile, pyproject.toml), tests, continuous integration, documentation
Not found: CITATION.cff
Tools: PyTorch (165 files), NumPy (158 files), anndata (107 files), pandas (72 files), SciPy (38 files), PyTorch Lightning (22 files), scikit-learn (21 files), Pyro (18 files), Scanpy (8 files), xarray (6 files), PyTorch Geometric (5 files), h5py (4 files), Matplotlib (4 files), Biopython (1 file), Numba (1 file), seaborn (1 file), SHAP (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
345 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 343 scripts, each with its path and the digest of its content;
  • 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data and materials availability

The integrated C. robusta larva-stage single-cell RNA sequencing dataset and the complete differential expression data for all clusters shown in Figure 4C are deposited at Zenodo (https://doi.org/10.5281/zenodo.15320250).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 5 keywords, 7 MeSH terms, 2 funders, 104 references.

Cite

This paper

Kourakis, M. J., Miao, Y., Newman-Smith, E. D., Ryan, K., & Smith, W. C. (2026). Similarities between <i>Ciona</i> Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor? eNeuro, 13(8), ENEURO.0362-25.2026. https://doi.org/10.1523/eneuro.0362-25.2026

BibTeX

@article{kourakis2026similarities,
author = {Kourakis, Matthew J. and Miao, Yishen and Newman-Smith, Erin D. and Ryan, Kerrianne and Smith, William C.},
title = {{Similarities between \<i\>Ciona\</i\> Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?}},
journal = {eNeuro},
year = {2026},
month = aug,
volume = {13},
number = {8},
pages = {ENEURO.0362--25.2026},
publisher = {Society for Neuroscience},
issn = {2373-2822},
doi = {10.1523/eneuro.0362-25.2026},
url = {https://doi.org/10.1523/eneuro.0362-25.2026},
pmid = {42637546},
pmcid = {PMC13509005}
}

RIS

TY - JOUR
AU - Kourakis, Matthew J.
AU - Miao, Yishen
AU - Newman-Smith, Erin D.
AU - Ryan, Kerrianne
AU - Smith, William C.
TI - Similarities between <i>Ciona</i> Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
T2 - eNeuro
J2 - eNeuro
PY - 2026
DA - 2026/08/24
VL - 13
IS - 8
SP - ENEURO.0362
EP - 25.2026
SN - 2373-2822
PB - Society for Neuroscience
DO - 10.1523/eneuro.0362-25.2026
UR - https://doi.org/10.1523/eneuro.0362-25.2026
LA - en
ER -

CSL-JSON

{
"id": "10.1523/eneuro.0362-25.2026",
"type": "article-journal",
"title": "Similarities between <i>Ciona</i> Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?",
"container-title": "eNeuro",
"author": [
{
"family": "Kourakis",
"given": "Matthew J."
},
{
"family": "Miao",
"given": "Yishen"
},
{
"family": "Newman-Smith",
"given": "Erin D."
},
{
"family": "Ryan",
"given": "Kerrianne"
},
{
"family": "Smith",
"given": "William C."
}
],
"container-title-short": "eNeuro",
"volume": "13",
"issue": "8",
"page": "ENEURO.0362-25.2026",
"DOI": "10.1523/eneuro.0362-25.2026",
"PMID": "42637546",
"PMCID": "PMC13509005",
"ISSN": "2373-2822",
"publisher": "Society for Neuroscience",
"URL": "https://doi.org/10.1523/eneuro.0362-25.2026",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
24
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: PyTorch Geometric, anndata, Numba, 10 other tools, 1 reference
[2] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: Biopython, anndata, Numba, 10 other tools, 1 reference
[3] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: PyTorch Geometric, anndata, Numba, 10 other tools, 1 reference
[4] doi:10.1093/bioinformatics/btag652 [code]
mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.
Journal: Bioinformatics (Oxford, England)
In common: Pyro, PyTorch Lightning, anndata, 9 other tools, 1 reference
[5] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: anndata, Numba, Scanpy, 7 other tools, 3 references
[6] doi:10.1002/advs.77003 [code]
SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: PyTorch Geometric, anndata, Numba, 9 other tools, 1 reference
[7] doi:10.1093/bioinformatics/btag540 [code]
Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.
Journal: Bioinformatics (Oxford, England)
In common: PyTorch Geometric, anndata, Numba, 9 other tools, 1 reference
[8] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: PyTorch Lightning, Biopython, SHAP, 9 other tools
[9] doi:10.64898/2026.03.30.714220 [code]
An integrated single cell and spatial omics atlas of human prenatal development
Journal: bioRxiv (preprint)
In common: Biopython, anndata, Scanpy, 9 other tools, 1 reference
[10] doi:10.1038/s41467-026-74694-6 [code]
Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Journal: Nature communications
In common: Pyro, anndata, Scanpy, 8 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.