Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
The 3 matches
- [1] § Methods › Data processing › TCGA pan-gynecologic atlas ↔ sofa/utils/utils.py, lines 18–105 · score 0.60 · highly variable features, log transformed, variance
- [2] § Methods › Data processing › Pan-cancer DepMap ↔ sofa/utils/utils.py, lines 18–105 · score 0.59 · highly variable genes, log transformed
- [3] § Methods › Downstream analysis › Gene set overrepresentation analysis (ORA) ↔ sofa/plots/plots.py, lines 348–407 · score 0.52 · enrichr API, overrepresentation, background, gene
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 504 lines · 16 KB · MIT · 2 matches
- #!/usr/bin/env python3
- import pyro
- import torch
- import numpy as np
- import muon as mu
- from muon import MuData
- from sklearn.preprocessing import LabelEncoder
- from anndata import AnnData
- from ..models.SOFA import SOFA
- import pandas as pd
- import scanpy as sc
- from typing import Union
- import numpy as np
- from sklearn.preprocessing import LabelEncoder, StandardScaler
- import gseapy as gp
- from sklearn.metrics import log_loss, accuracy_score
- def get_ad(data: pd.DataFrame,
- llh: str="gaussian",
- select_hvg: bool=False,
- log: bool=False,
- scale: bool=False,
- scaling_factor: float=0.1
- ) -> AnnData:
- """
- Convert a numpy array to an AnnData object.
- Parameters:
- -----------
- data : pandas DataFrame
- The input data to be converted to AnnData object.
- name : str
- The name of the variable.
- llh : str, optional
- The likelihood of the data. It should be "gaussian", "bernoulli" or
- "categorical". Default is "gaussian".
- select_hvg: bool, optional
- whether to select highly variable features.
- log: bool, optional
- whether to log transform the data.
- scale: bool, optional
- whether to center and scale the data.
- scaling_factor: float, optional
- The scaling factor to scale the likelihood for this view.
- It is a crucial hyperparameter. Too high scaling parameters can lead to overfitting
- (Factors don't explain variance of Xmdata, but perfectly predict Ymdata).
- Too low values can lead to the model ignoring the covariate guidance. In practice
- a value of 0.1 is a good starting point. Default is 0.1.
- Returns:
- --------
- adata : AnnData
- The converted AnnData object.
- """
- data =data.loc[:,~data.columns.duplicated()]
- data_ = data.loc[~np.all(pd.isnull(data), axis=1),:]
- data = data.loc[:,~np.any(pd.isnull(data_), axis=0)]
- mask = ~np.any(pd.isnull(data), axis=1)
- mask.index = mask.index.astype(str)
- if llh == "multinomial" or llh == "bernoulli":
- if type(data) != int:
- label_encoder = LabelEncoder()
- # Apply LabelEncoder to each column
- encoded_data = data.values.flatten()
- encoded_data = label_encoder.fit_transform(encoded_data)
- encoded_data = encoded_data.reshape(data.shape)
- adata = AnnData(encoded_data, dtype=np.float32)
- label_mapping = {label: encoded_label for label, encoded_label in zip(label_encoder.classes_, label_encoder.transform(label_encoder.classes_))}
- else:
- encoded_data = data.values
- if data.shape[1] == 1:
- adata = AnnData(encoded_data.reshape(-1,1), dtype=np.float32)
- adata.var_names = data.columns
- else:
- adata = AnnData(encoded_data, dtype=np.float32)
- adata.var_names = data.columns
- data.index = data.index.astype(str)
- adata.obs_names = data.index.tolist()
- if log:
- sc.pp.log1p(adata)
- adata.obsm["mask"] = mask.values
- if select_hvg:
- adata_filtered = adata[~np.all(data.isna(),axis=1),:]
- adata = adata[:,~np.any(np.isnan(adata_filtered.X),axis=0)]
- adata_filtered = adata_filtered[:,~np.any(np.isnan(adata_filtered.X),axis=0)]
- sc.pp.highly_variable_genes(adata_filtered, n_top_genes=2000)
- adata = adata[:,adata_filtered.var["highly_variable"]]
- adata.var["highly_variable"] = adata_filtered.var["highly_variable"]
- if scale:
- scaler = StandardScaler()
- adata.X = scaler.fit_transform(adata.X)
- adata.X[adata.obsm["mask"] == False] = 0
- adata.uns["llh"] = llh
- adata.uns["scaling_factor"] = scaling_factor
- return adata
- def calc_var_explained(X_pred, X):
- """
- Calculate R2 for X and X_pred.
- Parameters
- ----------
- X_pred : numpy.array
- Predicted X.
- X : numpy.array
- Input X.
- Returns
- -------
- float
- R2 value for X and X_pred.
- """
- num = np.sum(np.square(X-X_pred))
- denom = np.sum(np.square(X))
- vexp= 1 - num/denom
- if vexp < 0:
- vexp = 0
- return vexp
- def get_var_explained_per_view_factor(model: SOFA):
- """
- Calculate the fraction of variance of each view
- that is explained by each factor.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- Returns
- -------
- numpy.array
- Array containing the fraction of variance of each view
- that is explained by each factor.
- """
- X = [i.cpu().numpy() for i in model.X]
- vexp = []
- if not hasattr(model, "Z"):
- model.Z = model.predict("Z", num_split=10000)
- if not hasattr(model, f"W"):
- model.W = [model.predict(f"W_{i}", num_split=10000) for i in range(len(X))]
- for i in range(len(X)):
- mask = model.Xmask[i].cpu().numpy()
- vexp_factor = []
- for j in range(model.num_factors):
- X_pred_factor = model.Z[mask,j, np.newaxis] @ model.W[i][np.newaxis,j,:]
- vexp_factor.append(calc_var_explained(X_pred_factor, X[i][mask,:]))
- vexp.append(np.stack(vexp_factor).reshape(model.num_factors,1))
- vexp = np.hstack(vexp)
- return vexp
- def calc_var_explained_(X_pred, X):
- """
- Calculate the fraction of variance of each view
- that is explained by each factor.
- Parameters
- ----------
- X_pred : numpy.array
- Predicted X.
- X : numpy.array
- Input X.
- Returns
- -------
- numpy.array
- Array containing the fraction of variance of each view
- that is explained by each factor.
- """
- vexp = []
- for i in range(len(X_pred)):
- num = np.sum(np.square(X-X_pred[i]))
- denom = np.sum(np.square(X))
- vexp.append(1 - num/denom)
- vexp = np.stack(vexp)
- vexp[vexp < 0] = 0
- return vexp
- def get_loadings(model: SOFA,
- view: str
- )-> pd.DataFrame:
- """
- Get the loadings of the model for a specific view.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- view : str
- Name of the view to get the loadings for.
- Returns
- -------
- pd.DataFrame
- DataFrame containing the loadings of the model for the specified view.
- """
- ind_labels = np.array([f"Factor_{i+1}" for i in range(model.num_factors)], dtype=object)
- if model.Ymdata is not None:
- guided_factors = list(model.Ymdata.mod.keys())
- for i in range(len(guided_factors)):
- s = " (" + guided_factors[i] + ")"
- ind_labels[model.design.cpu().numpy()[i,:]==1] = ind_labels[model.design.cpu().numpy()[i,:]==1] + s
- if hasattr(model, f"W"):
- W = pd.DataFrame(model.W[model.views.index(view)], index = ind_labels, columns = model.Xmdata.mod[view].var_names)
- else:
- model.W = [model.predict(f"W_{i}", num_split=10000) for i in range(len(model.X))]
- W = pd.DataFrame(model.W[model.views.index(view)], index = ind_labels, columns = model.Xmdata.mod[view].var_names)
- return W
- def get_factors(model: SOFA,
- )-> pd.DataFrame:
- """
- Get the loadings of the model for a specific view.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- Returns
- -------
- pd.DataFrame
- DataFrame containing the loadings of the model for the specified view.
- """
- col_labels = np.array([f"Factor_{i+1}" for i in range(model.num_factors)], dtype=object)
- if model.Ymdata is not None:
- guided_factors = list(model.Ymdata.mod.keys())
- for i in range(len(guided_factors)):
- s = " (" + guided_factors[i] + ")"
- col_labels[model.design.cpu().numpy()[i,:]==1] = col_labels[model.design.cpu().numpy()[i,:]==1] + s
- if hasattr(model, f"Z"):
- Z = pd.DataFrame(model.Z, index = model.Xmdata.obs.index, columns = col_labels)
- else:
- model.Z = model.predict("Z")
- Z = pd.DataFrame(model.Z, index = model.Xmdata.obs.index, columns = col_labels)
- return Z
- def get_top_loadings(model,view, factor, sign="+", top_n=100):
- """
- Get the top_n loadings of the model for a specific view.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- view : str
- Name of the view to get the loadings for.
- factor : int
- Index of the factor to get the top loadings for.
- Should be between 1 and the total number of factors.
- sign : str
- Sign of the loadings to get. Default is "+".
- top_n : int
- Number of top loadings to get. Default is 100.
- Returns
- -------
- pandas.DataFrame
- DataFrame containing the top_n loadings of the model for the specified view.
- """
- assert(sign=="+" or sign=="-")
- # correct for pythonic indexing
- factor = factor-1
- W = get_loadings(model, view)
- W = W.iloc[factor,:]
- if sign == "+":
- idx = np.argpartition(W, -top_n)[-top_n:]
- topW = W.index[idx]
- elif sign=="-":
- idx = np.argpartition(W*-1, -top_n)[-top_n:]
- topW = W.index[idx]
- return topW.tolist()
- def get_gsea_enrichment(gene_list, db, background):
- """
- Get gene set enrichment analysis results based on a gene_list using gseapy.
- Parameters
- ----------
- gene_list : list
- List of strings containing gene names.
- db : list
- List of strings containing database names to be used for enrichment analysis.
- background : list
- List of strings containing gene names to be used as background.
- Returns
- -------
- Enrichr object
- Enrichr object containing the results of the enrichment analysis.
- """
- enr = gp.enrichr(gene_list=gene_list, # or "./tests/data/gene_list.txt",
- gene_sets=[db],
- organism='human', # don't forget to set organism to the one you desired! e.g. Yeast
- outdir=None,# don't write to disk
- background=background
- )
- return enr
- def calc_rmse(X, X_pred):
- """
- Calculate the root mean squared error between X and X_pred.
- Parameters
- ----------
- X_pred : numpy.array
- Predicted X.
- X : numpy.array
- Input X.
- Returns
- -------
- float
- The root mean squared error of X of the model.
- """
- rmse = np.sqrt(np.sum(np.square(X-X_pred))/(X.shape[0]*X.shape[1]))
- return rmse
- def get_rmse(model):
- """
- Calculate the root mean squared error of the model.
- Parameters
- ----------
- model : SOFA
- THe trained SOFA model.
- Returns
- -------
- dict
- The root mean squared error of X of the model for each view.
- """
- if not hasattr(model, f"X_pred"):
- model.X_pred = [model.predict(f"X_{i}", num_split=10000) for i in range(len(model.X))]
- rmse = {}
- for i in range(len(model.X)):
- rmse[model.views[i]] = calc_rmse(model.X[i].cpu().numpy(), model.X_pred[i])
- return rmse
- def sigmoid(x):
- return 1 / (1 + np.exp(-x))
- def softmax(x):
- return np.exp(x) / np.sum(np.exp(x))
- def get_guide_error(model):
- """
- Calculate the root mean squared error for continuous, binary crossentropy for binary or
- categorical cross entropy for categorical Y of the model.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- Returns
- -------
- dict
- Containing the root mean squared error for continuous, binary crossentropy for binary or
- categorical cross entropy for categorical Y of the model.
- """
- if model.Ymdata is None:
- raise ValueError("Model does not have guide variables!")
- if not hasattr(model, "Y_pred"):
- model.Y_pred = []
- for i in range(len(model.Y)):
- model.Y_pred.append(model.predict(f"Y_{i}"))
- error = {}
- for ix, (i,j) in enumerate(zip(model.Y, model.Y_pred)):
- i = i.cpu().numpy()
- if model.guide_llh[ix] == "gaussian":
- error[model.guide_views[ix]]= calc_rmse(i,j)
- elif model.guide_llh[ix] == "bernoulli":
- error[model.guide_views[ix]] = log_loss(i,sigmoid(j))
- elif model.guide_llh[ix] == "multinomial":
- error[model.guide_views[ix]] = log_loss(i,softmax(j))
- return error
- def save_model(model, file_prefix):
- """
- Saves a model as h5mu and save files to disk.
- Model hyperparameters, input data and predictions are saved in the h5mu file
- and the model parameters are saved in the save file.
- Both files are needed to load a model and continue training.
- Parameters
- ----------
- model : SOFA
- The trained SOFA model.
- file_prefix : str
- Filename prefix to save the model as h5mu and save files.
- Returns
- -------
- tuple(str,str)
- Filenames of the saved h5mu and save files.
- """
- model_mdata = model.save_as_mudata()
- try:
- model_mdata.write(file_prefix+".h5mu")
- except RuntimeError:
- print("Metadata probably too long to save. Please save the metadata separately. \nSaving model without metadata!")
- # remove metadata if saving does not work
- model_mdata.obs = pd.DataFrame(index=model_mdata.obs.index)
- # save model without metadata
- model_mdata.write(file_prefix+".h5mu")
- except TypeError:
- print("Mixed type columns are currently not supported when saving metadata of a model. Please save the metadata separately.\nSaving model without metadata!")
- # remove metadata if saving does not work
- model_mdata.obs = pd.DataFrame(index=model_mdata.obs.index)
- # save model without metadata
- model_mdata.write(file_prefix+".h5mu")
- except Exception:
- print("Unexpected error! \nSaving model without metadata!")
- # remove metadata if saving does not work
- model_mdata.obs = pd.DataFrame(index=model_mdata.obs.index)
- # save model without metadata
- model_mdata.write(file_prefix+".h5mu")
- dict_ = pyro.get_param_store()
- dict_.save(file_prefix +".save")
- return file_prefix +".h5mu", file_prefix +".save"
- def load_model(file_prefix):
- """
- Load a saved model from disk.
- The function requires an h5mu and a save file to load model.
- Parameters
- ----------
- file_prefix : str
- Filename prefix to save the model as h5mu and save files.
- Returns
- -------
- SOFA
- The loaded SOFA model.
- """
- mdata = mu.read(file_prefix +".h5mu")
- if "guide_mod" in list(mdata.uns.keys()):
- Ymdata = MuData({i:mdata.mod[i] for i in mdata.uns["guide_mod"]})
- Xmdata = MuData({i:mdata.mod[i] for i in mdata.mod if i not in mdata.uns["guide_mod"]})
- design = mdata.uns["input_design"]
- else:
- Xmdata = MuData({i:mdata.mod[i] for i in mdata.mod})
- Ymdata = None
- design = np.array(0)
- num_factors = mdata.uns["input_num_factors"]
- # TODO find way to save and load mixed column metadata
- # check if entries in obs (metadata)
- if mdata.obs.shape[1] !=0:
- metadata = mdata.obs
- else:
- metadata = None
- horseshoe = mdata.uns["horseshoe"]
- # apparently mu.read does not read None type in uns so
- # need to check if seed is there, if not set to None
- if "seed" not in mdata.uns:
- seed = None
- else:
- seed = mdata.uns["seed"]
- pyro.set_rng_seed(seed)
- device = mdata.uns["device"]
- model = SOFA(Xmdata,
- num_factors=num_factors,
- Ymdata = Ymdata,
- design = torch.tensor(design),
- device=device,
- horseshoe=horseshoe,
- subsample=0,
- metadata = metadata,
- seed=seed)
- model.Z = mdata.uns["Z"]
- W = [mdata.uns[f"W_{i}"] for i in model.views]
- model.W = W
- model.X_pred =[mdata.uns[f"X_{i}"] for i in model.views]
- if Ymdata is not None:
- model.Y_pred = [mdata.uns[f"Y_pred_{i}"] for i in Ymdata.mod]
- model.history = mdata.uns["history"]
- # load pyro paramstore
- dict_ = pyro.get_param_store()
- dict_.load(file_prefix+".save")
- return model
utils.py at commit 2e56def, under MIT · at the source
Overview
- Collaboration for joint PhD degree between EMBL and Heidelberg University, Faculty of Biosciences,Heidelberg, Germany
- Genome Biology Unit, EMBL,Heidelberg, Germany
- Molecular Medicine Partnership Unit (MMPU), Heidelberg, Germany
- Department of Medicine V, Hematology, Oncology and Rheumatology, University Hospital Heidelberg,Heidelberg, Germany
- Department of Hematology and Oncology, University Hospital Düsseldorf,Düsseldorf, Germany
- Heidelberg University, Faculty of Medicine, and Heidelberg University Hospital, Institute for Computational Biomedicine,Heidelberg, Germany
- European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI),Hinxton, UK
Abstract
A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different ‘omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyze such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications, a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty.
Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFA’s effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.
Zenodo 14761127
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
tcapraz/SOFA
2e56defd2ef8ced64bf642c121297712b5f8b8bc, 15 October 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
17 files
- docs/
conf.py , Python, 59 lines - docs/
notebooks/ , Jupyter, 382 linesdepmap_example.ipynb - docs/
notebooks/ , Jupyter, 15 linesinstallation.ipynb - docs/
notebooks/ , Jupyter, 440 linessc_multiome_example.ipyn b - docs/
notebooks/ , Jupyter, 571 linestcga_gyn_analysis.ipynb - sofa/
__init__.py , Python, 16 lines - sofa/
models/ , Python, 701 linesSOFA.py - sofa/
plots/ , Python, 3 lines__init__.py - sofa/
plots/ , Python, 407 lines, 1 matchplots.py - sofa/
utils/ , Python, 8 lines__init__.py - sofa/
utils/ , Python, 504 lines, 2 matchesutils.py - tests/
conftest.py , Python, 59 lines - tests/
test_plots.py , Python, 57 lines - tests/
test_sofa.py , Python, 148 lines - tests/
test_utils.py , Python, 213 lines - LICENSE, License, 21 lines
- README.md, Text, 42 lines
Zenodo 19594991
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
Code availability
SOFA is implemented in the Python package biosofa and is available on github (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 15 scripts, each with its path and the digest of its content;
- 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
Processed data to reproduce the analysis and figures are available on Zenodo 10.5281/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 3 keywords, 7 MeSH terms, 1 funder, 82 references.
Cite
This paper
Capraz, T., Vöhringer, H., Kruger Serrano, K. S. A., Ramirez Flores, R. O., Saez-Rodriguez, J., & Huber, W. (2026). Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data. Nature communications, 17(1), 8725. https://
BibTeX
@article{capraz2026semi,
author = {Capraz, Tümay and Vöhringer, Harald and Kruger Serrano, Klaus Sebastian Augusto and Ramirez Flores, Ricardo Omar and Saez-Rodriguez, Julio and Huber, Wolfgang},
title = {{Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data}},
journal = {Nature communications},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {8725},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42624833},
pmcid = {PMC13493367}
}
RIS
TY - JOUR
AU - Capraz, Tümay
AU - Vöhringer, Harald
AU - Kruger Serrano, Klaus Sebastian Augusto
AU - Ramirez Flores, Ricardo Omar
AU - Saez-Rodriguez, Julio
AU - Huber, Wolfgang
TI - Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 8725
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data",
"container-title": "Nature communications",
"author": [
{
"family": "Capraz",
"given": "Tümay"
},
{
"family": "Vöhringer",
"given": "Harald"
},
{
"family": "Kruger Serrano",
"given": "Klaus Sebastian Augusto"
},
{
"family": "Ramirez Flores",
"given": "Ricardo Omar"
},
{
"family": "Saez-Rodriguez",
"given": "Julio"
},
{
"family": "Huber",
"given": "Wolfgang"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "8725",
"DOI": "10.1038/
"PMID": "42624833",
"PMCID": "PMC13493367",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
20
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1093/bioinformatics/btag652 [code]
- mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.Journal: Bioinformatics (Oxford, England)In common: Pyro, anndata, Scanpy, 8 other tools, genetics / omics, 5 references
- [2] doi:10.1038/s44320-026-00208-7 [code]
- Interpretable deep generative ensemble learning for single-cell omics with Hydra.Journal: Molecular systems biologyIn common: anndata, Scanpy, PyTorch, 6 other tools, 5 references
- [3] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: anndata, Scanpy, statsmodels, 6 other tools, genetics / omics, 3 references
- [4] doi:10.1523/eneuro.0362-25.2026 [code]
- Similarities between &
lt;i& gt;Ciona& lt;/ i& gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/ V Subdivision within a Hindbrain Precursor? Journal: eNeuroIn common: Pyro, anndata, Scanpy, 8 other tools, 1 reference - [5] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: anndata, Scanpy, statsmodels, 7 other tools, methods / tools, genetics / omics, 2 references
- [6] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: anndata, Scanpy, statsmodels, 7 other tools, genetics / omics, 2 references
- [7] doi:10.1016/j.xgen.2026.101217 [code]
- ProtoCloud: A prototypical self-explaining model for single-cell analysis.Journal: Cell genomicsIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 3 references
- [8] doi:10.64898/2026.03.30.714220 [code]
- An integrated single cell and spatial omics atlas of human prenatal developmentJournal: bioRxiv (preprint)In common: anndata, Scanpy, statsmodels, 7 other tools, 2 references
- [9] doi:10.1093/bib/bbag298 [code]
- Empowering multifaceted analysis of spatial transcriptomics data with RGAST.Journal: Briefings in bioinformaticsIn common: anndata, Scanpy, PyTorch, 6 other tools, methods / tools, genetics / omics, 2 references
- [10] doi:10.1002/advs.77003 [code]
- SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: anndata, Scanpy, PyTorch, 6 other tools, methods / tools, genetics / omics, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 15 scripts, and 3 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:af6529d890ee1350…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
