scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
The 12 matches
- [1] § Materials and methods › Standardizing pre-training corpus for accurate transcriptomic profiling ↔ scDifformer_data_process/preprocess/preprocess.py, lines 44–116 · score 0.79 · Ensembl annotated protein, standard deviation, detection, doublets, loom, mitochondrial
- [2] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.ipynb, lines 205–284 · score 0.70 · fully connected neural, neural network, pseudo spots, dimensions, deconvolution, module
- [3] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.py, lines 209–287 · score 0.70 · fully connected neural, neural network, pseudo spots, dimensions, deconvolution, module
- [4] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.ipynb, lines 1–76 · score 0.68 · logistic regression, identified cell, expression matrix, spots, transcriptome, genes
- [5] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.py, lines 6–80 · score 0.68 · logistic regression, identified cell, expression matrix, spots, transcriptome, genes
- [6] § Results › Cross-modal application of scDifformer to spatial spot deconvolution ↔ Tutorial.ipynb, lines 205–284 · score 0.67 · neural network, STdGCN, deep learning, spatial transcriptomics, deconvolution, spots
- [7] § Results › Cross-modal application of scDifformer to spatial spot deconvolution ↔ Tutorial.py, lines 209–287 · score 0.67 · neural network, STdGCN, deep learning, spatial transcriptomics, deconvolution, spots
- [8] § Materials and methods › scDifformer architecture and pre-training ↔ DSTG/models.py, lines 86–136 · score 0.57 · Adam optimizer, weight decay, L2, layer, trained, model
- [9] § Materials and methods › scDifformer architecture and pre-training ↔ pretrain.py, lines 174–261 · score 0.56 · warmup steps, learning rate scheduling, Adam, optimization, validated, trained
- [10] § Materials and methods › scDifformer gene embeddings, cell embeddings, and attention weights ↔ scgpt_spatial/model/model.py, lines 1077–1185 · score 0.53 · gene embedding, cell embeddings, vector, space, hidden, architecture
- [11] § Results › scDifformer achieves state-of-the-art performance in cell-type annotation across diverse datasets ↔ lr_baseline_crossorgan.py, lines 45–75 · score 0.52 · macro F1 score, cross validation, human, accuracy
- [12] § Materials and methods › scDifformer gene embeddings, cell embeddings, and attention weights ↔ scgpt/model/model.py, lines 915–1004 · score 0.50 · gene embedding, cell embeddings, vector, hidden, architecture, dimensional
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 287 lines · 15 KB · no license · 3 matches
- # %%
- import os
- import sys
- import warnings
- warnings.filterwarnings("ignore")
- sys.path.append(os.getcwd())
- from STdGCN.STdGCN import run_STdGCN
- '''
- This module is used to provide the path of the loading data and saving data.
- Parameters:
- sc_path: The path for loading single cell reference data.
- ST_path: The path for loading spatial transcriptomics data.
- output_path: The path for saving output files.
- The relevant file name and data format for loading:
- sc_data.tsv: The expression matrix of the single cell reference data with cells as rows and genes as columns. This file should be saved in "sc_path".
- sc_label.tsv: The cell-type annotation of sincle cell data. The table should have two columns: The cell barcode/name and the cell-type annotation information.
- This file should be saved in "sc_path".
- ST_data.tsv: The expression matrix of the spatial transcriptomics data with spots as rows and genes as columns. This file should be saved in "ST_path".
- coordinates.csv: The coordinates of the spatial transcriptomics data. The table should have three columns: Spot barcode/name, X axis (column name 'x'), and Y axis (column name 'y').
- This file should be saved in "ST_path".
- marker_genes.tsv [optional]: The gene list used to run STdGCN. Each row is a gene and no table header is permitted. This file should be saved in "sc_path".
- ST_ground_truth.tsv [optional]: The ground truth of ST data. The data should be transformed into the cell type proportions. This file should be saved in "ST_path".
- '''
- paths = {
- 'sc_path': './data/sc_data',
- 'ST_path': './data/ST_data',
- 'output_path': './output',
- }
- '''
- This module is used to preprocess the input data and identify marker genes [optional].
- Parameters:
- 'preprocess': [bool]. Select whether the input expression data needs to be preprocessed. This step includes normalization, logarithmization, selecting highly variable genes,
- regressing out mitochondrial genes, and scaling data.
- 'normalize': [bool]. When 'preprocess'=True, select whether you need to normalize each cell/spot by total counts = 10,000, so that every cell/spot has the same total
- count after normalization.
- 'log': [bool]. When 'preprocess'=True, select whether you need to logarithmize (X=log(X+1)) the expression matrix.
- 'highly_variable_genes': [bool]. When 'preprocess'=True, select whether you need to filter the highly variable genes.
- 'highly_variable_gene_num': [int or None]. When 'preprocess'=True and 'highly_variable_genes'=True, select the number of highly-variable genes to keep.
- 'regress_out': [bool]. When 'preprocess'=True, select whether you need to regress out mitochondrial genes.
- 'scale': [bool]. When 'preprocess'=True, select whether you need to scale each gene to unit variance and zero mean.
- 'PCA_components': [int]. Number of principal components to compute for principal component analysis (PCA).
- 'marker_gene_method': ['logreg', 'wilcoxon']. We used "scanpy.tl.rank_genes_groups" (https://scanpy.readthedocs.io/en/stable/generated/scanpy.tl.rank_genes_groups.html)
- to identify cell type marker genes. For marker gene selection, STdGCN provides two methods, 'wilcoxon' (Wilcoxon rank-sum) and 'logreg' (uses
- logistic regression).
- 'top_gene_per_type': [int]. The number of genes for each cell type that can be used to train STdGCN.
- 'filter_wilcoxon_marker_genes': [bool]. When 'marker_gene_method'='wilcoxon', select whether you need additional steps for gene filtering.
- 'pvals_adj_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes with corrected p-values < 'pvals_adj_threshold' were kept.
- 'log_fold_change_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes with log fold change > 'log_fold_change_threshold' were kept.
- 'min_within_group_fraction_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes expressed with fraction at least
- 'min_within_group_fraction_threshold' in the cell type were kept.
- 'max_between_group_fraction_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes expressed with fraction at most
- 'max_between_group_fraction_threshold' in the union of the rest of cell types were kept.
- '''
- find_marker_genes_paras = {
- 'preprocess': True,
- 'normalize': True,
- 'log': True,
- 'highly_variable_genes': False,
- 'highly_variable_gene_num': None,
- 'regress_out': False,
- 'PCA_components': 30,
- 'marker_gene_method': 'logreg',
- 'top_gene_per_type': 100,
- 'filter_wilcoxon_marker_genes': True,
- 'pvals_adj_threshold': 0.10,
- 'log_fold_change_threshold': 1,
- 'min_within_group_fraction_threshold': None,
- 'max_between_group_fraction_threshold': None,
- }
- '''
- This module is used to simulate pseudo-spots.
- Parameters:
- 'spot_num': [int]. The number of pseudo-spots.
- 'min_cell_num_in_spot': [int]. The minimum number of cells in a pseudo-spot.
- 'max_cell_num_in_spot': [int]. The maximum number of cells in a pseudo-spot.
- 'generation_method': ['cell' or 'celltype']. STdGCN provides two pseudo-spot simulation methods. When 'generation_method'='cell', each cell is equally selected. When
- 'generation_method'='celltype', each cell type is equally selected. See manuscript for more details.
- 'max_cell_types_in_spot': [int]. When 'generation_method'='celltype', choose the maximum number of cell types in a pseudo-spot.
- '''
- pseudo_spot_simulation_paras = {
- 'spot_num': 30000,
- 'min_cell_num_in_spot': 8,
- 'max_cell_num_in_spot': 12,
- 'generation_method': 'celltype',
- 'max_cell_types_in_spot': 4,
- }
- '''
- This module is used for real- and pseudo- spots normalization.
- Parameters:
- 'normalize': [bool]. Select whether you need to normalize each cell/spot by total counts = 10,000, so that every cell/spot has the same total count after normalization.
- 'log': [bool]. Select whether you need to logarithmize (X=log(X+1)) the expression matrix.
- 'scale': [bool]. Select whether you need to scale each gene to unit variance and zero mean.
- '''
- data_normalization_paras = {
- 'normalize': True,
- 'log': True,
- 'scale': False,
- }
- '''
- This module is used to integrate the normalized real- and pseudo- spots together to construct the real-to-pseudo-spot link graph.
- Parameters:
- 'batch_removal_method': ['mnn', 'scanorama', 'combat', None]. Considering batch effects, STdGCN provides four integration methods: mnn (mnnpy, DOI:10.1038/nbt.4091),
- scanorama (Scanorama, DOI: 10.1038/s41587-019-0113-3), combat (Combat, DOI: 10.1093/biostatistics/kxj037), None (concatenation with no batch removal).
- 'dimensionality_reduction_method': ['PCA', 'autoencoder', 'nmf', None]. When 'batch_removal_method' is not 'scanorama', select whether the data needs dimensionality reduction, and which
- dimensionality reduction method is applied.
- 'dim': [int]. When 'batch_removal_method'='scanorama', select the dimension for this method. When 'batch_removal_method' is not 'scanorama' and 'dimensionality_reduction_method' is
- not None, select the dimension of the dimensionality reduction.
- 'scale': [bool]. When 'batch_removal_method' is not 'scanorama', select whether you need to scale each gene to unit variance and zero mean.
- '''
- integration_for_adj_paras = {
- 'batch_removal_method': None,
- 'dim': 30,
- 'dimensionality_reduction_method': 'PCA',
- 'scale': True,
- }
- '''
- The module is used to construct the adjacency matrix of the expression graph, which contains three subgraphs: a real-to-pseudo-spot graph, a pseudo-spots internal graph,
- and a real-spots internal graph.
- Parameters:
- 'find_neighbor_method' ['MNN', 'KNN']. STdGCN provides two methods for link graph construction, KNN (K-nearest neighbors) and MNN (mutual nearest neighbors, DOI: 10.1038/nbt.4091).
- 'dist_method': ['euclidean', 'cosine']. The metrics used for computing paired distances between spots.
- 'corr_dist_neighbors': [int]. The number of nearest neighbors.
- 'PCA_dimensionality_reduction': [bool]. For pseudo-spots internal graph and real-spots internal graph construction, select if the data needs to use PCA dimensionality reduction before
- computing paired distances between spots.
- 'dim': [int]. When 'PCA_dimensionality_reduction'=True, select the dimension of the PCA.
- '''
- inter_exp_adj_paras = {
- 'find_neighbor_method': 'MNN',
- 'dist_method': 'cosine',
- 'corr_dist_neighbors': 20,
- }
- real_intra_exp_adj_paras = {
- 'find_neighbor_method': 'MNN',
- 'dist_method': 'cosine',
- 'corr_dist_neighbors': 10,
- 'PCA_dimensionality_reduction': False,
- 'dim': 50,
- }
- pseudo_intra_exp_adj_paras = {
- 'find_neighbor_method': 'MNN',
- 'dist_method': 'cosine',
- 'corr_dist_neighbors': 20,
- 'PCA_dimensionality_reduction': False,
- 'dim': 50,
- }
- '''
- The module is used to construct the adjacency matrix of the spatial graph.
- Parameters:
- 'space_dist_threshold': [float or None]. Only the distance between two spots smaller than 'space_dist_threshold' can be linked.
- 'link_method' ['soft', 'hard']. If spot i and j linked, A(i,j)=1 if 'link_method'='hard', while A(i,j)=1/distance(i,j) if 'link_method'='soft'. See manuscript for more details.
- '''
- spatial_adj_paras = {
- 'link_method': 'soft',
- 'space_dist_threshold': 2,
- }
- '''
- This module is used to integrate the normalized real- and pseudo- spots as the input feature for STdGCN.
- Parameters:
- 'batch_removal_method': ['mnn', 'scanorama', 'combat', None]. Considering batch effects, STdGCN provides four integration methods: mnn (mnnpy, DOI:10.1038/nbt.4091),
- scanorama (Scanorama, DOI: 10.1038/s41587-019-0113-3), combat (Combat, DOI: 10.1093/biostatistics/kxj037), None (concatenation with no batch removal).
- 'dimensionality_reduction_method': ['PCA', 'autoencoder', 'nmf', None]. When 'batch_removal_method' is not 'scanorama', select whether the data needs dimensionality reduction, and which
- dimensionality reduction method is applied.
- 'dim': [int]. When 'batch_removal_method'='scanorama', select the dimension for this method. When 'batch_removal_method' is not 'scanorama' and 'dimensionality_reduction_method' is
- not None, select the dimension of the dimensionality reduction.
- 'scale': [bool]. When 'batch_removal_method' is not 'scanorama', select whether you need to scale each gene to unit variance and zero mean.
- '''
- integration_for_feature_paras = {
- 'batch_removal_method': None,
- 'dimensionality_reduction_method': None,
- 'dim': 80,
- 'scale': True,
- }
- '''
- This module is used for setting the deep learning parameters for STdGCN.
- Parameters:
- 'epoch_n': [int]. The maximum number of epochs.
- 'dim': [int]. The dimension of the hidden layers.
- 'common_hid_layers_num': [int]. The number of GCN layers = 'common_hid_layers_num'+1.
- 'fcnn_hid_layers_num': [int]. The number of fully connected neural network layers = 'fcnn_hid_layers_num'+2.
- 'dropout': [float]. The probability of an element to be zeroed.
- 'learning_rate_SGD': [float]. Initial learning rate.
- 'weight_decay_SGD': [float]. L2 penalty.
- 'momentum': [float]. Momentum factor.
- 'dampening': [float]. Dampening for momentum.
- 'nesterov': [bool]. Enables Nesterov momentum.
- 'early_stopping_patience': [int]. Early stopping epochs.
- 'clip_grad_max_norm': [float]. Clips gradient norm of an iterable of parameters.
- #'LambdaLR_scheduler_coefficient': [float]. The coefficent of the LambdaLR scheduler fucntion: lr(epoch) = [LambdaLR_scheduler_coefficient] ^ epoch_n × learning_rate_SGD.
- 'print_loss_epoch_step': [int]. Print the loss value at every 'print_epoch_step' epoch.
- '''
- GCN_paras = {
- 'epoch_n': 3000,
- 'dim': 80,
- 'common_hid_layers_num': 1,
- 'fcnn_hid_layers_num': 1,
- 'dropout': 0,
- 'learning_rate_SGD': 2e-1,
- 'weight_decay_SGD': 3e-4,
- 'momentum': 0.9,
- 'dampening': 0,
- 'nesterov': True,
- 'early_stopping_patience': 20,
- 'clip_grad_max_norm': 1,
- #'LambdaLR_scheduler_coefficient': 0.997,
- 'print_loss_epoch_step': 20,
- }
- '''
- ## run STdGCN
- Parameters
- 'load_test_groundtruth': [bool]. Select whether you need to upload the ground truth file (ST_ground_truth.tsv) of the spatial transcriptomics data to track the performance of STdGCN.
- 'use_marker_genes': [bool]. Select whether you need the gene selection process before running STdGCN. Otherwise use common genes from single cell and spatial transcriptomics data.
- 'external_genes': [bool]. When "use_marker_genes"=True, you can upload your specified gene list (marker_genes.tsv) to run STdGCN.
- 'generate_new_pseudo_spots': [bool]. STdGCN will save the simulated pseudo-spots to "pseudo_ST.pkl". If you want to run multiple deconvolutions with the same single cell reference data,
- you don't need to simulate new pseudo-spots and set 'generate_new_pseudo_spots'=False. When 'generate_new_pseudo_spots'=False, you need to pre-move the "pseudo_ST.pkl"
- to the 'output_path' so that STdGCN can directly load the pre-simulated pseudo-spots.
- 'fraction_pie_plot': [bool]. Select whether you need to draw the pie plot of the predicted results. Based on our experience, we do not recommend to draw the pie plot when the predicted
- spot number is very large. For 1,000 spots, the plotting time is less than 2 minutes; for 2,000 spots, the plotting time is about 10 minutes; for 3,000 spots, it takes
- about 30 minutes.
- 'cell_type_distribution_plot': [bool]. Select whether you need to draw the scatter plot of the predicted results for each cell type.
- 'n_jobs': [int]. Set the number of threads used for intraop parallelism on CPU. 'n_jobs=-1' represents using all CPUs.
- 'GCN_device': ['GPU', 'CPU']. Select the device used to run GCN networks.
- '''
- results = run_STdGCN(paths,
- load_test_groundtruth = False,
- use_marker_genes = True,
- external_genes = False,
- find_marker_genes_paras = find_marker_genes_paras,
- generate_new_pseudo_spots = True,
- pseudo_spot_simulation_paras = pseudo_spot_simulation_paras,
- data_normalization_paras = data_normalization_paras,
- integration_for_adj_paras = integration_for_adj_paras,
- inter_exp_adj_paras = inter_exp_adj_paras,
- spatial_adj_paras = spatial_adj_paras,
- real_intra_exp_adj_paras = real_intra_exp_adj_paras,
- pseudo_intra_exp_adj_paras = pseudo_intra_exp_adj_paras,
- integration_for_feature_paras = integration_for_feature_paras,
- GCN_paras = GCN_paras,
- fraction_pie_plot = True,
- cell_type_distribution_plot = True,
- n_jobs = -1,
- GCN_device = 'GPU'
- )
- results.write_h5ad(paths['output_path']+'/results.h5ad')
- # %%
Tutorial.ipynb at commit d02d390, no license · at the source
Overview
- Research Center for Life Sciences Computing, Zhejiang Lab, Hangzhou, Zhejiang 311121, China
- Zhejiang Provincial Key Laboratory of Pancreatic Disease, The First Affiliated Hospital, and Institute of Translational Medicine, Zhejiang University School of Medicine, Hangzhou 311121, China
- Department of Laboratory Medicine of The First Affiliated Hospital & Liangzhu Laboratory, Zhejiang University School of Medicine, Hangzhou 311121, China
- Key Laboratory of Spatial Omics of Zhejiang Province, State Key Laboratory of Genome and Multi-omics Technologies, BGI Research, Hangzhou 310030, China
- Nebraska Food for Health Center, Department of Food Science and Technology, University of Nebraska, Lincoln, NE 68588, United States
- Department of Chemistry, The University of Manchester, Manchester M13 9PL, United Kingdom
- School of Life Science and Technology, ShanghaiTech University, Shanghai 200092, China
- Chongqing Institute of Intelligent Medicine, 799 Jingwei Avenue, Yuzhong District, Chongqing 400042, China
Abstract
Virtual cells represent a promising paradigm to understand cellular mechanisms, behavior, and dynamics. The realization of virtual cells relies on the accurate modeling of cellular dynamics from large-scale, multi-modal single-cell data. However, experiment-specific technical noise and intrinsic biological heterogeneity pose major challenges for virtual cell modeling. To address this gap, we present scDifformer, a context-aware transformer model augmented with a denoising diffusion module and a dedicated post-training phase. This three-phase design, comprising masked language model pre-training, diffusion-driven post-training, and downstream fine-tuning, directly enhances scDifformer’s ability to denoise sparse, noisy data and generalize across studies. Benchmarking across seven tissues and multiple independent studies shows that the diffusion module consistently improves cross-dataset performance, particularly in settings with strong batch effects. By combining the strengths of transformer and diffusion models, scDifformer achieves state-of-the-art performance in cell type annotation across diverse datasets. It further demonstrates robust capability in resolving immune cell identities across multiple tissues, accurately recovering key marker genes, functional pathways, and cross-tissue differentiation trajectories. Finally, by integrating scDifformer with a graph neural network, we extend its utility to spatial transcriptomics, significantly enhancing spot-level deconvolution accuracy. Altogether, scDifformer provides a scalable and biologically grounded framework for modeling heterogeneous single-cell data, offering a powerful foundation for the development of high-fidelity, multi-modal virtual cell models.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 12 matches between paragraphs and lines of code.
huggingface.co/ctheodoris/geneformer
1f7fbae4e469a5f4f1af8c111a529cfe1b3829f5, 12 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
TencentAILabHealthcare/scBERT
262fd4b91f3f1c21a6e595d03d4ef423e16ffc99, 13 December 2023Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
12 files
- attn_sum_save.py, Python, 91 lines
- finetune.py, Python, 271 lines
- lr_baseline_crossorgan.p
y , Python, 76 lines, 1 match - performer_pytorch/
__init__.py , Python, 1 line - performer_pytorch/
performer_pytorch.py , Python, 639 lines - performer_pytorch/
reversible.py , Python, 167 lines - predict.py, Python, 127 lines
- preprocess.py, Python, 27 lines
- pretrain.py, Python, 261 lines, 1 match
- utils.py, Python, 376 lines
- LICENSE, License, 673 lines
- README.md, Text, 94 lines
biomap-research/scFoundation
397631c495eddf9ad6644fc00c6ea8139e651245, 23 November 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
73 files
- DeepCDR/
plot.ipynb , Jupyter, 409 lines - DeepCDR/
prog/ , Python, 1 linelayers/ __init__.py - DeepCDR/
prog/ , Python, 178 lineslayers/ graph.py - DeepCDR/
prog/ , Python, 112 linesmodel.py - DeepCDR/
prog/ , Python, 29 linesprocess_drug.py - DeepCDR/
prog/ , Shell, 16 linesrun.sh - DeepCDR/
prog/ , Python, 320 linesrun_DeepCDR.py - DeepCDR/
prog/ , Python, 320 linesrun_DeepCDR_leave_drug.p y - DeepCDR/
prog/ , Python, 108 linesrun_pytorch_embedding.py - GEARS/
Plot.ipynb , Jupyter, 703 lines - GEARS/
gears/ , Python, 2 lines__init__.py - GEARS/
gears/ , Python, 410 linesdata_utils.py - GEARS/
gears/ , Python, 527 linesgears.py - GEARS/
gears/ , Python, 885 linesinference.py - GEARS/
gears/ , Python, 236 linesmodel.py - GEARS/
gears/ , Python, 373 linespertdata.py - GEARS/
gears/ , Python, 361 linesutils.py - GEARS/
gears/ , Python, 22 linesversion.py - GEARS/
modules/ , Python, 1 line__init__.py - GEARS/
modules/ , Python, 267 linesattention.py - GEARS/
modules/ , Python, 280 linesencoders.py - GEARS/
modules/ , Python, 160 linesmae_autobin.py - GEARS/
modules/ , Python, 708 linesperformer_module.py - GEARS/
modules/ , Python, 199 linesreversible.py - GEARS/
modules/ , Python, 231 linestransformer.py - GEARS/
run_sh/ , Shell, 37 linesrun_singlecell-adamson.s h - GEARS/
run_sh/ , Shell, 37 linesrun_singlecell-dixit.sh - GEARS/
run_sh/ , Shell, 56 linesrun_singlecell_maeautobi n-0.1B-res0-adamson.sh - GEARS/
run_sh/ , Shell, 56 linesrun_singlecell_maeautobi n-0.1B-res0-dixit.sh - GEARS/
run_sh/ , Shell, 54 linesrun_singlecell_maeautobi n-0.1B-res0-norman.sh - GEARS/
run_sh/ , Shell, 50 linesrun_singlecell_maeautobi n-demo-baseline.sh - GEARS/
run_sh/ , Shell, 54 linesrun_singlecell_maeautobi n-demo-emb-train.sh - GEARS/
run_sh/ , Shell, 55 linesrun_singlecell_maeautobi n-demo-emb.sh - GEARS/
run_sh/ , Shell, 36 linesrun_singlecell_norman.sh - GEARS/
train.py , Python, 78 lines - SCAD/
data/ , Python, 51 linesprocessing/ preprocessing_GDSC_all_d rug_logIC50_binary_inter sect.py - SCAD/
data/ , Python, 247 linessplit_norm/ split_data_SCAD_5fold_no rm.py - SCAD/
model/ , Python, 629 linesSCAD_train_binarized_5fo lds-pub.py - SCAD/
model/ , Python, 169 linesSCADmodules.py - SCAD/
plot-publish.ipynb , Jupyter, 673 lines - SCAD/
run.sh , Shell, 28 lines - SCAD/
run_embedding_bulk.py , Python, 120 lines - SCAD/
run_embedding_sc.py , Python, 108 lines - ablation/
ablation-00.ipynb , Jupyter, 255 lines - ablation/
ablation-01.ipynb , Jupyter, 73 lines - ablation/
ablation-02.ipynb , Jupyter, 114 lines - annotation/
celltype-plot.ipynb , Jupyter, 190 lines - apiexample/
client.py , Python, 171 lines - apiexample/
example.sh , Shell, 21 lines - enhancement/
Baron_evaluation.ipynb , Jupyter, 303 lines - enhancement/
PBMC68k_evaluation.ipynb , Jupyter, 222 lines - enhancement/
run.sh , Shell, 7 lines - enhancement/
run_embedding_sc.py , Python, 105 lines - genemodule/
plot_geneemb.ipynb , Jupyter, 283 lines - mapping/
mapping-publish.ipynb , Jupyter, 199 lines - model/
check_consistency.ipynb , Jupyter, 139 lines - model/
demo.sh , Shell, 43 lines - model/
finetune_model.py , Python, 75 lines - model/
get_embedding.py , Python, 263 lines - model/
load.py , Python, 198 lines - model/
pretrainmodels/ , Python, 3 lines__init__.py - model/
pretrainmodels/ , Python, 162 linesmae_autobin.py - model/
pretrainmodels/ , Python, 638 linesperformer.py - model/
pretrainmodels/ , Python, 80 linespytorchTransformer.py - model/
pretrainmodels/ , Python, 200 linesreversible.py - model/
pretrainmodels/ , Python, 61 linesselect_model.py - model/
pretrainmodels/ , Python, 43 linestransformer.py - preprocessing/
demo.ipynb , Jupyter, 51 lines - preprocessing/
demo.sh , Shell, 5 lines - preprocessing/
down.sh , Shell, 42 lines - preprocessing/
scRNA_workflow.py , Python, 127 lines - LICENSE, License, 201 lines
- README.md, Text, 100 lines
biomed-AI/CellFM
72c9f4a9580a3716058c184900ed14a65151ed8f, 7 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
85 files
- .ipynb_checkpoints/
1node_train-checkpoint.s , Shell, 40 linesh - .ipynb_checkpoints/
attention-checkpoint.py , Python, 93 lines - .ipynb_checkpoints/
data_process-checkpoint. , Python, 203 linespy - .ipynb_checkpoints/
earlystop-checkpoint.py , Python, 161 lines - .ipynb_checkpoints/
lora-checkpoint.py , Python, 30 lines - .ipynb_checkpoints/
loss_function-checkpoint , Python, 114 lines.py - .ipynb_checkpoints/
metrics-checkpoint.py , Python, 160 lines - .ipynb_checkpoints/
model-checkpoint.py , Python, 268 lines - .ipynb_checkpoints/
retention-checkpoint.py , Python, 247 lines - .ipynb_checkpoints/
scheduler_sc100m-checkpo , Shell, 47 linesint.sh - .ipynb_checkpoints/
train-checkpoint.py , Python, 142 lines - .ipynb_checkpoints/
utils-checkpoint.py , Python, 222 lines - .ipynb_checkpoints/
worker_sc100m-checkpoint , Shell, 42 lines.sh - 1node_train.sh, Shell, 40 lines
- attention.py, Python, 93 lines
- config.py, Python, 57 lines
- data_process.py, Python, 203 lines
- earlystop.py, Python, 161 lines
- get_gene_emb.py, Python, 361 lines
- lora.py, Python, 30 lines
- loss_function.py, Python, 114 lines
- metrics.py, Python, 194 lines
- model.py, Python, 270 lines
- retention.py, Python, 247 lines
- scheduler_sc100m.sh, Shell, 47 lines
- train.py, Python, 143 lines
- tutorials/
BatchIntegration/ , Jupyter, 622 linesBatchIntegration.ipynb - tutorials/
BatchIntegration/ , Python, 257 linesfinetune_intergration_mo del.py - tutorials/
BinaryclassGeneFunction. , Jupyter, 224 linesipynb - tutorials/
CellAnnotation/ , Jupyter, 257 linesCellAnnotation_finetune. ipynb - tutorials/
CellAnnotation/ , Jupyter, 252 linesCellAnnotation_zeroshot. ipynb - tutorials/
CellAnnotation/ , Python, 113 linesannotation_model.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ data/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 402 linescellot_model/ data/ cell.py - tutorials/
ChemicalPerturbation/ , Python, 60 linescellot_model/ data/ utils.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ losses/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 24 linescellot_model/ losses/ mmd.py - tutorials/
ChemicalPerturbation/ , Python, 3 linescellot_model/ models/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 247 linescellot_model/ models/ ae.py - tutorials/
ChemicalPerturbation/ , Python, 231 linescellot_model/ models/ cellfm.py - tutorials/
ChemicalPerturbation/ , Python, 166 linescellot_model/ models/ cellot.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ networks/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 135 linescellot_model/ networks/ icnns.py - tutorials/
ChemicalPerturbation/ , Python, 322 linescellot_model/ networks/ retention.py - tutorials/
ChemicalPerturbation/ , Python, 129 linescellot_model/ networks/ scret.py - tutorials/
ChemicalPerturbation/ , Python, 105 linescellot_model/ preprocess.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ train/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 123 linescellot_model/ train/ experiment.py - tutorials/
ChemicalPerturbation/ , Python, 105 linescellot_model/ train/ summary.py - tutorials/
ChemicalPerturbation/ , Python, 410 linescellot_model/ train/ train.py - tutorials/
ChemicalPerturbation/ , Python, 38 linescellot_model/ train/ utils.py - tutorials/
ChemicalPerturbation/ , Python, 68 linescellot_model/ transport.py - tutorials/
ChemicalPerturbation/ , Python, 1 linecellot_model/ utils/ __init__.py - tutorials/
ChemicalPerturbation/ , Python, 357 linescellot_model/ utils/ evaluate.py - tutorials/
ChemicalPerturbation/ , Python, 109 linescellot_model/ utils/ flags.py - tutorials/
ChemicalPerturbation/ , Python, 173 linescellot_model/ utils/ helpers.py - tutorials/
ChemicalPerturbation/ , Python, 48 linescellot_model/ utils/ loaders.py - tutorials/
ChemicalPerturbation/ , Python, 339 linescellot_model/ utils/ viz.py - tutorials/
ChemicalPerturbation/ , Python, 184 linesevaluate.py - tutorials/
ChemicalPerturbation/ , Python, 482 linesfinetune_sciplex.py - tutorials/
ChemicalPerturbation/ , Python, 17 linessetup.py - tutorials/
ChemicalPerturbation/ , Python, 80 linestrain.py - tutorials/
IdentifyingCelltypelncRN , Jupyter, 466 linesAs.ipynb - tutorials/
MulticlassGeneFunction.i , Jupyter, 287 linespynb - tutorials/
Perturbation/ , Jupyter, 50 linesGenePerturbation.ipynb - tutorials/
Perturbation/ , Python, 2 linesgears/ __init__.py - tutorials/
Perturbation/ , Python, 410 linesgears/ data_utils.py - tutorials/
Perturbation/ , Python, 548 linesgears/ gears.py - tutorials/
Perturbation/ , Python, 894 linesgears/ inference.py - tutorials/
Perturbation/ , Python, 264 linesgears/ model.py - tutorials/
Perturbation/ , Python, 395 linesgears/ pertdata.py - tutorials/
Perturbation/ , Python, 367 linesgears/ utils.py - tutorials/
Perturbation/ , Python, 22 linesgears/ version.py - tutorials/
Perturbation/ , Python, 1 linemodules/ __init__.py - tutorials/
Perturbation/ , Python, 267 linesmodules/ attention.py - tutorials/
Perturbation/ , Python, 281 linesmodules/ encoders.py - tutorials/
Perturbation/ , Python, 160 linesmodules/ mae_autobin.py - tutorials/
Perturbation/ , Python, 708 linesmodules/ performer_module.py - tutorials/
Perturbation/ , Python, 199 linesmodules/ reversible.py - tutorials/
Perturbation/ , Python, 231 linesmodules/ transformer.py - tutorials/
process.ipynb , Jupyter, 164 lines - utils.py, Python, 266 lines
- worker_sc100m.sh, Shell, 42 lines
- readme.md, Text, 111 lines
bowang-lab/scgpt
cebd6fae655b9c585a4807daa3ac31bb764f06b4, 27 April 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
59 files
- data/
cellxgene/ , Shell, 31 linesarray_build_scb.sh - data/
cellxgene/ , Shell, 21 linesarray_download_partition .sh - data/
cellxgene/ , Shell, 53 linesarray_process_allcounts. sh - data/
cellxgene/ , Python, 224 linesbuild_large_scale_data.p y - data/
cellxgene/ , Python, 71 linesbuild_soma_idx.py - data/
cellxgene/ , Shell, 10 linesbuild_soma_idx.sh - data/
cellxgene/ , Python, 33 linesdata_config.py - data/
cellxgene/ , Python, 103 linesdownload_partition.py - data/
cellxgene/ , Shell, 22 linesdownload_partition.sh - data/
cellxgene/ , Python, 41 linesexpand_gene_list.py - data/
cellxgene/ , Python, 315 linesprocess_allcounts.py - docs/
conf.py , Python, 45 lines - examples/
finetune_integration.py , Python, 798 lines - scgpt/
__init__.py , Python, 29 lines - scgpt/
data_collator.py , Python, 202 lines - scgpt/
data_sampler.py , Python, 95 lines - scgpt/
loss.py , Python, 36 lines - scgpt/
model/ , Python, 11 lines__init__.py - scgpt/
model/ , Python, 82 linesdsbn.py - scgpt/
model/ , Python, 491 linesflash_attn_compat.py - scgpt/
model/ , Python, 546 linesgeneration_model.py - scgpt/
model/ , Python, 17 linesgrad_reverse.py - scgpt/
model/ , Python, 1,039 lines, 1 matchmodel.py - scgpt/
model/ , Python, 1,085 linesmultiomic_model.py - scgpt/
preprocess.py , Python, 303 lines - scgpt/
scbank/ , Python, 19 lines__init__.py - scgpt/
scbank/ , Python, 135 linesdata.py - scgpt/
scbank/ , Python, 802 linesdatabank.py - scgpt/
scbank/ , Python, 2 linesmonitor.py - scgpt/
scbank/ , Python, 31 linessetting.py - scgpt/
tasks/ , Python, 2 lines__init__.py - scgpt/
tasks/ , Python, 280 linescell_emb.py - scgpt/
tasks/ , Python, 297 linesgrn.py - scgpt/
tokenizer/ , Python, 1 line__init__.py - scgpt/
tokenizer/ , Python, 516 linesgene_tokenizer.py - scgpt/
tokenizer/ , Python, 253 linesvocab_compat.py - scgpt/
trainer.py , Python, 685 lines - scgpt/
utils/ , Python, 1 line__init__.py - scgpt/
utils/ , Python, 642 linesutil.py - scripts/
diagnose_flash_attn.py , Python, 483 lines - tests/
__init__.py , Python, 1 line - tests/
test_flash_attn.py , Python, 404 lines - tests/
test_sampler.py , Python, 104 lines - tests/
test_scbank.py , Python, 232 lines - tests/
test_scformer.py , Python, 8 lines - tests/
test_tokenizer.py , Python, 147 lines - tutorials/
Tutorial_Annotation.ipyn , Jupyter, 1,092 linesb - tutorials/
Tutorial_Attention_GRN.i , Jupyter, 504 linespynb - tutorials/
Tutorial_GRN.ipynb , Jupyter, 312 lines - tutorials/
Tutorial_Integration.ipy , Jupyter, 809 linesnb - tutorials/
Tutorial_Multiomics.ipyn , Jupyter, 563 linesb - tutorials/
Tutorial_Perturbation.ip , Jupyter, 591 linesynb - tutorials/
Tutorial_Reference_Mappi , Jupyter, 221 linesng.ipynb - tutorials/
build_atlas_index_faiss. , Python, 378 linespy - tutorials/
zero-shot/ , Jupyter, 415 linesTutorial_ZeroShot_Integr ation.ipynb - tutorials/
zero-shot/ , Jupyter, 286 linesTutorial_ZeroShot_Integr ation_Continual_Pretrain ing.ipynb - tutorials/
zero-shot/ , Jupyter, 349 linesTutorial_ZeroShot_Refere nce_Mapping.ipynb - LICENSE, License, 21 lines
- README.md, Text, 113 lines
broadinstitute/Tangram
4c68995a418f41dc8caef567598c4d9b47781a13, 1 July 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
22 files
- cell_selection/
__init__.py , Python, 1 line - cell_selection/
cell_sampling.py , Python, 44 lines - docs/
source/ , Python, 70 linesconf.py - gene_selection/
__init__.py , Python, 4 lines - gene_selection/
celltype_specific_genes. , Python, 12 linespy - gene_selection/
highly_variable_genes.py , Python, 9 lines - gene_selection/
spapros_genes.py , Python, 13 lines - gene_selection/
spatially_variable_genes , Python, 13 lines.py - setup.py, Python, 37 lines
- tangram/
__init__.py , Python, 5 lines - tangram/
_version.py , Python, 3 lines - tangram/
mapping_optimizer.py , Python, 639 lines - tangram/
mapping_parameter_tuning , Python, 272 lines.py - tangram/
mapping_utils.py , Python, 428 lines - tangram/
plot_utils.py , Python, 724 lines - tangram/
spatial_weights.py , Python, 30 lines - tangram/
utils.py , Python, 842 lines - tests/
tangram_test.py , Python, 217 lines - tutorial_tangram_with_sq
uidpy.ipynb , Jupyter, 456 lines - tutorial_tangram_without
_squidpy.ipynb , Jupyter, 362 lines - LICENSE.md, License, 29 lines
- README.md, Text, 170 lines
luoyuanlab/stdgcn
d02d39085bd297ba26212a07170821e03b644549, 6 May 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
11 files
- STdGCN/
GCN.py , Python, 264 lines - STdGCN/
STdGCN.py , Python, 328 lines - STdGCN/
__init__.py , Python, 7 lines - STdGCN/
_version.py , Python, 3 lines - STdGCN/
adjacency_matrix.py , Python, 218 lines - STdGCN/
autoencoder.py , Python, 70 lines - STdGCN/
utils.py , Python, 358 lines - STdGCN/
visualization.py , Python, 101 lines - Tutorial.ipynb, Jupyter, 287 lines, 3 matches
- Tutorial.py, Python, 287 lines, 3 matches
- README.md, Text, 49 lines
Su-informatics-lab/DSTG
4a2f958a87ba3137c3ffa75188563af3e0245a5a, 11 August 2022Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
15 files
- DSTG/
R_utils.R , R, 230 lines - DSTG/
__init__.py , Python, 2 lines - DSTG/
convert_data.R , R, 29 lines - DSTG/
data.py , Python, 72 lines - DSTG/
evaluation.R , R, 13 lines - DSTG/
graph.py , Python, 101 lines - DSTG/
gutils.py , Python, 177 lines - DSTG/
layers.py , Python, 144 lines - DSTG/
metrics.py , Python, 175 lines - DSTG/
models.py , Python, 136 lines, 1 match - DSTG/
train.py , Python, 139 lines - DSTG/
utils.py , Python, 102 lines - setup.py, Python, 23 lines
- LICENCE, License, 21 lines
- README.md, Text, 51 lines
bowang-lab/scGPT-spatial
f9442d777ba47cd3dcac8253833a233ff5bf7262, 12 February 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
21 files
- scgpt_spatial/
__init__.py , Python, 9 lines - scgpt_spatial/
data_collator.py , Python, 459 lines - scgpt_spatial/
data_sampler.py , Python, 95 lines - scgpt_spatial/
loss.py , Python, 40 lines - scgpt_spatial/
model/ , Python, 53 linesMoE.py - scgpt_spatial/
model/ , Python, 8 lines__init__.py - scgpt_spatial/
model/ , Python, 426 linesflash_layers.py - scgpt_spatial/
model/ , Python, 17 linesgrad_reverse.py - scgpt_spatial/
model/ , Python, 165 lineslayers.py - scgpt_spatial/
model/ , Python, 1,220 lines, 1 matchmodel.py - scgpt_spatial/
preprocess.py , Python, 324 lines - scgpt_spatial/
tasks/ , Python, 1 line__init__.py - scgpt_spatial/
tasks/ , Python, 350 linescell_emb.py - scgpt_spatial/
tokenizer/ , Python, 1 line__init__.py - scgpt_spatial/
tokenizer/ , Python, 551 linesgene_tokenizer.py - scgpt_spatial/
utils/ , Python, 2 lines__init__.py - scgpt_spatial/
utils/ , Python, 44 linesspa_util.py - scgpt_spatial/
utils/ , Python, 391 linesutil.py - tutorials/
Multi_Modal_Integration_ , Jupyter, 114 linesDemo.ipynb - LICENSE, License, 21 lines
- README.md, Text, 50 lines
huggingface.co/allenxiao/scdifformer
54d1f5c9756909f9b4c83b7917844ef159d8cdca, 14 May 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
14 files
- scDifformer_data_process
/ , Python, 1 line__init__.py - scDifformer_data_process
/ , Python, 44 linesdict_token.py - scDifformer_data_process
/ , Python, 50 linesloom_batch_100.py - scDifformer_data_process
/ , Python, 71 linesnonzero_median.py - scDifformer_data_process
/ , Python, 69 linesnonzero_median_combine.p y - scDifformer_data_process
/ , Python, 1 linepreflight/ __init__.py - scDifformer_data_process
/ , Python, 65 linespreflight/ convert_data.py - scDifformer_data_process
/ , Python, 51 linespreflight/ dataset_var_x_info.py - scDifformer_data_process
/ , Python, 1 linepreprocess/ __init__.py - scDifformer_data_process
/ , Python, 146 lines, 1 matchpreprocess/ preprocess.py - scDifformer_data_process
/ , Python, 58 linesprint_data_info.py - scDifformer_data_process
/ , Python, 38 linestoken_data_combine.py - scDifformer_data_process
/ , Python, 268 linestokenizer.py - README.md, Text, 57 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 10 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 297 scripts, each with its path and the digest of its content;
- 12 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE115469, at NCBI GEO; found in the text, “Intra-dataset classification”
Data availability
The model weights and a minimun data of scDifformer are available in Huggingface project (https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 16 authors, 6 MeSH terms, 2 funders, 82 references.
Cite
This paper
Xiao, Z., Wang, W., Long, X., Zhang, W., Zhang, W., Chen, D., Xu, S., Yu, Q., Zhang, X., Huang, S., Zhang, N., Yin, Y., Huang, X., Ye, J., Zheng, J., & Guo, L. (2026). scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data. Nucleic acids research, 54(13), gkag706. https://
BibTeX
@article{xiao2026scdiffo
author = {Xiao, Zhan and Wang, Wuke and Long, Xin and Zhang, Wenbo and Zhang, Weiqiang and Chen, Duoyuan and Xu, Sujie and Yu, Qian and Zhang, Xinpeng and Huang, Shichen and Zhang, Ning and Yin, Yanbin and Huang, Xingxu and Ye, Jieping and Zheng, Jinfang and Guo, Ling},
title = {{scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data}},
journal = {Nucleic acids research},
year = {2026},
month = jul,
volume = {54},
number = {13},
pages = {gkag706},
publisher = {Oxford University Press},
issn = {0305-1048},
doi = {10.1093/
url = {https://
pmid = {42444608},
pmcid = {PMC13366050}
}
RIS
TY - JOUR
AU - Xiao, Zhan
AU - Wang, Wuke
AU - Long, Xin
AU - Zhang, Wenbo
AU - Zhang, Weiqiang
AU - Chen, Duoyuan
AU - Xu, Sujie
AU - Yu, Qian
AU - Zhang, Xinpeng
AU - Huang, Shichen
AU - Zhang, Ning
AU - Yin, Yanbin
AU - Huang, Xingxu
AU - Ye, Jieping
AU - Zheng, Jinfang
AU - Guo, Ling
TI - scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data
T2 - Nucleic acids research
J2 - Nucleic Acids Res
PY - 2026
DA - 2026/
VL - 54
IS - 13
SP - gkag706
SN - 0305-1048
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data",
"container-title": "Nucleic acids research",
"author": [
{
"family": "Xiao",
"given": "Zhan"
},
{
"family": "Wang",
"given": "Wuke"
},
{
"family": "Long",
"given": "Xin"
},
{
"family": "Zhang",
"given": "Wenbo"
},
{
"family": "Zhang",
"given": "Weiqiang"
},
{
"family": "Chen",
"given": "Duoyuan"
},
{
"family": "Xu",
"given": "Sujie"
},
{
"family": "Yu",
"given": "Qian"
},
{
"family": "Zhang",
"given": "Xinpeng"
},
{
"family": "Huang",
"given": "Shichen"
},
{
"family": "Zhang",
"given": "Ning"
},
{
"family": "Yin",
"given": "Yanbin"
},
{
"family": "Huang",
"given": "Xingxu"
},
{
"family": "Ye",
"given": "Jieping"
},
{
"family": "Zheng",
"given": "Jinfang"
},
{
"family": "Guo",
"given": "Ling"
}
],
"container-title-short":
"volume": "54",
"issue": "13",
"page": "gkag706",
"DOI": "10.1093/
"PMID": "42444608",
"PMCID": "PMC13366050",
"ISSN": "0305-1048",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s44320-026-00208-7 [code]
- Interpretable deep generative ensemble learning for single-cell omics with Hydra.Journal: Molecular systems biologyIn common: Keras, UMAP, anndata, 13 other tools, 7 references
- [2] doi:10.21203/rs.3.rs-9676637/v1 [code]
- A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic TechnologiesJournal: Research Square (preprint)In common: Squidpy, PyTorch Geometric, UMAP, 14 other tools, 3 references
- [3] doi:10.1016/j.xgen.2026.101217 [code]
- ProtoCloud: A prototypical self-explaining model for single-cell analysis.Journal: Cell genomicsIn common: UMAP, anndata, Scanpy, 8 other tools, 9 references
- [4] doi:10.1038/s41467-026-71759-4 [code]
- CellNiche represents cellular microenvironments in atlas-scale spatial omics data with contrastive learning.Journal: Nature communicationsIn common: Squidpy, PyTorch Geometric, anndata, 9 other tools, 6 references
- [5] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: Keras, UMAP, anndata, 14 other tools, 2 references
- [6] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: Squidpy, PyTorch Geometric, UMAP, 14 other tools
- [7] doi:10.1038/s41467-026-68596-w [code]
- Spatial cartography of human thymus enables the geopositioning of lineage transcription factors in rare mimetic thymic epithelial cells.Journal: Nature communicationsIn common: Squidpy, anndata, Scanpy, 12 other tools, 3 references
- [8] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: RDKit, Keras, UMAP, 14 other tools
- [9] doi:10.1093/bib/bbag404 [code]
- Navigating cell maps by deep learning integration of single-cell and spatially resolved transcriptomics.Journal: Briefings in bioinformaticsIn common: PyTorch Geometric, anndata, Scanpy, 9 other tools, 5 references
- [10] doi:10.1186/s12864-026-12965-8 [code]
- Systematic evaluation of single-cell foundation model interpretability: attention-derived edge scores add no incremental value over gene-level features for perturbation-target prediction.Journal: BMC genomicsIn common: anndata, Scanpy, NetworkX, 9 other tools, 5 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 10 repositories of the authors' code, each at its verified commit and with its license, 297 scripts, and 12 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:cef185b39aec9a37…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
- TencentAILabHealthcare/scBERT: README.md on GitHub
- biomap-research/scFoundation: README.md on GitHub
- biomed-AI/CellFM: readme.md on GitHub
- bowang-lab/scgpt: README.md on GitHub
- broadinstitute/Tangram: README.md on GitHub
- luoyuanlab/stdgcn: README.md on GitHub
- Su-informatics-lab/DSTG: README.md on GitHub
- bowang-lab/scGPT-spatial: README.md on GitHub
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
