OSCR

scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.

Code ↔ Paper

12 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 12 matches
  1. [1] § Materials and methods › Standardizing pre-training corpus for accurate transcriptomic profiling ↔ scDifformer_data_process/preprocess/preprocess.py, lines 44–116 · score 0.79 · Ensembl annotated protein, standard deviation, detection, doublets, loom, mitochondrial
  2. [2] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.ipynb, lines 205–284 · score 0.70 · fully connected neural, neural network, pseudo spots, dimensions, deconvolution, module
  3. [3] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.py, lines 209–287 · score 0.70 · fully connected neural, neural network, pseudo spots, dimensions, deconvolution, module
  4. [4] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.ipynb, lines 1–76 · score 0.68 · logistic regression, identified cell, expression matrix, spots, transcriptome, genes
  5. [5] § Materials and methods › Application of model in spatial transcriptome deconvolution ↔ Tutorial.py, lines 6–80 · score 0.68 · logistic regression, identified cell, expression matrix, spots, transcriptome, genes
  6. [6] § Results › Cross-modal application of scDifformer to spatial spot deconvolution ↔ Tutorial.ipynb, lines 205–284 · score 0.67 · neural network, STdGCN, deep learning, spatial transcriptomics, deconvolution, spots
  7. [7] § Results › Cross-modal application of scDifformer to spatial spot deconvolution ↔ Tutorial.py, lines 209–287 · score 0.67 · neural network, STdGCN, deep learning, spatial transcriptomics, deconvolution, spots
  8. [8] § Materials and methods › scDifformer architecture and pre-training ↔ DSTG/models.py, lines 86–136 · score 0.57 · Adam optimizer, weight decay, L2, layer, trained, model
  9. [9] § Materials and methods › scDifformer architecture and pre-training ↔ pretrain.py, lines 174–261 · score 0.56 · warmup steps, learning rate scheduling, Adam, optimization, validated, trained
  10. [10] § Materials and methods › scDifformer gene embeddings, cell embeddings, and attention weights ↔ scgpt_spatial/model/model.py, lines 1077–1185 · score 0.53 · gene embedding, cell embeddings, vector, space, hidden, architecture
  11. [11] § Results › scDifformer achieves state-of-the-art performance in cell-type annotation across diverse datasets ↔ lr_baseline_crossorgan.py, lines 45–75 · score 0.52 · macro F1 score, cross validation, human, accuracy
  12. [12] § Materials and methods › scDifformer gene embeddings, cell embeddings, and attention weights ↔ scgpt/model/model.py, lines 915–1004 · score 0.50 · gene embedding, cell embeddings, vector, hidden, architecture, dimensional

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 287 lines · 15 KB · no license · 3 matches

  1. # %%
  2. import os
  3. import sys
  4. import warnings
  5. warnings.filterwarnings("ignore")
  6. sys.path.append(os.getcwd())
  7. from STdGCN.STdGCN import run_STdGCN
  8. '''
  9. This module is used to provide the path of the loading data and saving data.
  10. Parameters:
  11. sc_path: The path for loading single cell reference data.
  12. ST_path: The path for loading spatial transcriptomics data.
  13. output_path: The path for saving output files.
  14. The relevant file name and data format for loading:
  15. sc_data.tsv: The expression matrix of the single cell reference data with cells as rows and genes as columns. This file should be saved in "sc_path".
  16. sc_label.tsv: The cell-type annotation of sincle cell data. The table should have two columns: The cell barcode/name and the cell-type annotation information.
  17. This file should be saved in "sc_path".
  18. ST_data.tsv: The expression matrix of the spatial transcriptomics data with spots as rows and genes as columns. This file should be saved in "ST_path".
  19. coordinates.csv: The coordinates of the spatial transcriptomics data. The table should have three columns: Spot barcode/name, X axis (column name 'x'), and Y axis (column name 'y').
  20. This file should be saved in "ST_path".
  21. marker_genes.tsv [optional]: The gene list used to run STdGCN. Each row is a gene and no table header is permitted. This file should be saved in "sc_path".
  22. ST_ground_truth.tsv [optional]: The ground truth of ST data. The data should be transformed into the cell type proportions. This file should be saved in "ST_path".
  23. '''
  24. paths = {
  25. 'sc_path': './data/sc_data',
  26. 'ST_path': './data/ST_data',
  27. 'output_path': './output',
  28. }
  29. '''
  30. This module is used to preprocess the input data and identify marker genes [optional].
  31. Parameters:
  32. 'preprocess': [bool]. Select whether the input expression data needs to be preprocessed. This step includes normalization, logarithmization, selecting highly variable genes,
  33. regressing out mitochondrial genes, and scaling data.
  34. 'normalize': [bool]. When 'preprocess'=True, select whether you need to normalize each cell/spot by total counts = 10,000, so that every cell/spot has the same total
  35. count after normalization.
  36. 'log': [bool]. When 'preprocess'=True, select whether you need to logarithmize (X=log(X+1)) the expression matrix.
  37. 'highly_variable_genes': [bool]. When 'preprocess'=True, select whether you need to filter the highly variable genes.
  38. 'highly_variable_gene_num': [int or None]. When 'preprocess'=True and 'highly_variable_genes'=True, select the number of highly-variable genes to keep.
  39. 'regress_out': [bool]. When 'preprocess'=True, select whether you need to regress out mitochondrial genes.
  40. 'scale': [bool]. When 'preprocess'=True, select whether you need to scale each gene to unit variance and zero mean.
  41. 'PCA_components': [int]. Number of principal components to compute for principal component analysis (PCA).
  42. 'marker_gene_method': ['logreg', 'wilcoxon']. We used "scanpy.tl.rank_genes_groups" (https://scanpy.readthedocs.io/en/stable/generated/scanpy.tl.rank_genes_groups.html)
  43. to identify cell type marker genes. For marker gene selection, STdGCN provides two methods, 'wilcoxon' (Wilcoxon rank-sum) and 'logreg' (uses
  44. logistic regression).
  45. 'top_gene_per_type': [int]. The number of genes for each cell type that can be used to train STdGCN.
  46. 'filter_wilcoxon_marker_genes': [bool]. When 'marker_gene_method'='wilcoxon', select whether you need additional steps for gene filtering.
  47. 'pvals_adj_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes with corrected p-values < 'pvals_adj_threshold' were kept.
  48. 'log_fold_change_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes with log fold change > 'log_fold_change_threshold' were kept.
  49. 'min_within_group_fraction_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes expressed with fraction at least
  50. 'min_within_group_fraction_threshold' in the cell type were kept.
  51. 'max_between_group_fraction_threshold': [float or None]. When 'marker_gene_method'='wilcoxon' and 'rank_gene_filter'=True, only genes expressed with fraction at most
  52. 'max_between_group_fraction_threshold' in the union of the rest of cell types were kept.
  53. '''
  54. find_marker_genes_paras = {
  55. 'preprocess': True,
  56. 'normalize': True,
  57. 'log': True,
  58. 'highly_variable_genes': False,
  59. 'highly_variable_gene_num': None,
  60. 'regress_out': False,
  61. 'PCA_components': 30,
  62. 'marker_gene_method': 'logreg',
  63. 'top_gene_per_type': 100,
  64. 'filter_wilcoxon_marker_genes': True,
  65. 'pvals_adj_threshold': 0.10,
  66. 'log_fold_change_threshold': 1,
  67. 'min_within_group_fraction_threshold': None,
  68. 'max_between_group_fraction_threshold': None,
  69. }
  70. '''
  71. This module is used to simulate pseudo-spots.
  72. Parameters:
  73. 'spot_num': [int]. The number of pseudo-spots.
  74. 'min_cell_num_in_spot': [int]. The minimum number of cells in a pseudo-spot.
  75. 'max_cell_num_in_spot': [int]. The maximum number of cells in a pseudo-spot.
  76. 'generation_method': ['cell' or 'celltype']. STdGCN provides two pseudo-spot simulation methods. When 'generation_method'='cell', each cell is equally selected. When
  77. 'generation_method'='celltype', each cell type is equally selected. See manuscript for more details.
  78. 'max_cell_types_in_spot': [int]. When 'generation_method'='celltype', choose the maximum number of cell types in a pseudo-spot.
  79. '''
  80. pseudo_spot_simulation_paras = {
  81. 'spot_num': 30000,
  82. 'min_cell_num_in_spot': 8,
  83. 'max_cell_num_in_spot': 12,
  84. 'generation_method': 'celltype',
  85. 'max_cell_types_in_spot': 4,
  86. }
  87. '''
  88. This module is used for real- and pseudo- spots normalization.
  89. Parameters:
  90. 'normalize': [bool]. Select whether you need to normalize each cell/spot by total counts = 10,000, so that every cell/spot has the same total count after normalization.
  91. 'log': [bool]. Select whether you need to logarithmize (X=log(X+1)) the expression matrix.
  92. 'scale': [bool]. Select whether you need to scale each gene to unit variance and zero mean.
  93. '''
  94. data_normalization_paras = {
  95. 'normalize': True,
  96. 'log': True,
  97. 'scale': False,
  98. }
  99. '''
  100. This module is used to integrate the normalized real- and pseudo- spots together to construct the real-to-pseudo-spot link graph.
  101. Parameters:
  102. 'batch_removal_method': ['mnn', 'scanorama', 'combat', None]. Considering batch effects, STdGCN provides four integration methods: mnn (mnnpy, DOI:10.1038/nbt.4091),
  103. scanorama (Scanorama, DOI: 10.1038/s41587-019-0113-3), combat (Combat, DOI: 10.1093/biostatistics/kxj037), None (concatenation with no batch removal).
  104. 'dimensionality_reduction_method': ['PCA', 'autoencoder', 'nmf', None]. When 'batch_removal_method' is not 'scanorama', select whether the data needs dimensionality reduction, and which
  105. dimensionality reduction method is applied.
  106. 'dim': [int]. When 'batch_removal_method'='scanorama', select the dimension for this method. When 'batch_removal_method' is not 'scanorama' and 'dimensionality_reduction_method' is
  107. not None, select the dimension of the dimensionality reduction.
  108. 'scale': [bool]. When 'batch_removal_method' is not 'scanorama', select whether you need to scale each gene to unit variance and zero mean.
  109. '''
  110. integration_for_adj_paras = {
  111. 'batch_removal_method': None,
  112. 'dim': 30,
  113. 'dimensionality_reduction_method': 'PCA',
  114. 'scale': True,
  115. }
  116. '''
  117. The module is used to construct the adjacency matrix of the expression graph, which contains three subgraphs: a real-to-pseudo-spot graph, a pseudo-spots internal graph,
  118. and a real-spots internal graph.
  119. Parameters:
  120. 'find_neighbor_method' ['MNN', 'KNN']. STdGCN provides two methods for link graph construction, KNN (K-nearest neighbors) and MNN (mutual nearest neighbors, DOI: 10.1038/nbt.4091).
  121. 'dist_method': ['euclidean', 'cosine']. The metrics used for computing paired distances between spots.
  122. 'corr_dist_neighbors': [int]. The number of nearest neighbors.
  123. 'PCA_dimensionality_reduction': [bool]. For pseudo-spots internal graph and real-spots internal graph construction, select if the data needs to use PCA dimensionality reduction before
  124. computing paired distances between spots.
  125. 'dim': [int]. When 'PCA_dimensionality_reduction'=True, select the dimension of the PCA.
  126. '''
  127. inter_exp_adj_paras = {
  128. 'find_neighbor_method': 'MNN',
  129. 'dist_method': 'cosine',
  130. 'corr_dist_neighbors': 20,
  131. }
  132. real_intra_exp_adj_paras = {
  133. 'find_neighbor_method': 'MNN',
  134. 'dist_method': 'cosine',
  135. 'corr_dist_neighbors': 10,
  136. 'PCA_dimensionality_reduction': False,
  137. 'dim': 50,
  138. }
  139. pseudo_intra_exp_adj_paras = {
  140. 'find_neighbor_method': 'MNN',
  141. 'dist_method': 'cosine',
  142. 'corr_dist_neighbors': 20,
  143. 'PCA_dimensionality_reduction': False,
  144. 'dim': 50,
  145. }
  146. '''
  147. The module is used to construct the adjacency matrix of the spatial graph.
  148. Parameters:
  149. 'space_dist_threshold': [float or None]. Only the distance between two spots smaller than 'space_dist_threshold' can be linked.
  150. 'link_method' ['soft', 'hard']. If spot i and j linked, A(i,j)=1 if 'link_method'='hard', while A(i,j)=1/distance(i,j) if 'link_method'='soft'. See manuscript for more details.
  151. '''
  152. spatial_adj_paras = {
  153. 'link_method': 'soft',
  154. 'space_dist_threshold': 2,
  155. }
  156. '''
  157. This module is used to integrate the normalized real- and pseudo- spots as the input feature for STdGCN.
  158. Parameters:
  159. 'batch_removal_method': ['mnn', 'scanorama', 'combat', None]. Considering batch effects, STdGCN provides four integration methods: mnn (mnnpy, DOI:10.1038/nbt.4091),
  160. scanorama (Scanorama, DOI: 10.1038/s41587-019-0113-3), combat (Combat, DOI: 10.1093/biostatistics/kxj037), None (concatenation with no batch removal).
  161. 'dimensionality_reduction_method': ['PCA', 'autoencoder', 'nmf', None]. When 'batch_removal_method' is not 'scanorama', select whether the data needs dimensionality reduction, and which
  162. dimensionality reduction method is applied.
  163. 'dim': [int]. When 'batch_removal_method'='scanorama', select the dimension for this method. When 'batch_removal_method' is not 'scanorama' and 'dimensionality_reduction_method' is
  164. not None, select the dimension of the dimensionality reduction.
  165. 'scale': [bool]. When 'batch_removal_method' is not 'scanorama', select whether you need to scale each gene to unit variance and zero mean.
  166. '''
  167. integration_for_feature_paras = {
  168. 'batch_removal_method': None,
  169. 'dimensionality_reduction_method': None,
  170. 'dim': 80,
  171. 'scale': True,
  172. }
  173. '''
  174. This module is used for setting the deep learning parameters for STdGCN.
  175. Parameters:
  176. 'epoch_n': [int]. The maximum number of epochs.
  177. 'dim': [int]. The dimension of the hidden layers.
  178. 'common_hid_layers_num': [int]. The number of GCN layers = 'common_hid_layers_num'+1.
  179. 'fcnn_hid_layers_num': [int]. The number of fully connected neural network layers = 'fcnn_hid_layers_num'+2.
  180. 'dropout': [float]. The probability of an element to be zeroed.
  181. 'learning_rate_SGD': [float]. Initial learning rate.
  182. 'weight_decay_SGD': [float]. L2 penalty.
  183. 'momentum': [float]. Momentum factor.
  184. 'dampening': [float]. Dampening for momentum.
  185. 'nesterov': [bool]. Enables Nesterov momentum.
  186. 'early_stopping_patience': [int]. Early stopping epochs.
  187. 'clip_grad_max_norm': [float]. Clips gradient norm of an iterable of parameters.
  188. #'LambdaLR_scheduler_coefficient': [float]. The coefficent of the LambdaLR scheduler fucntion: lr(epoch) = [LambdaLR_scheduler_coefficient] ^ epoch_n × learning_rate_SGD.
  189. 'print_loss_epoch_step': [int]. Print the loss value at every 'print_epoch_step' epoch.
  190. '''
  191. GCN_paras = {
  192. 'epoch_n': 3000,
  193. 'dim': 80,
  194. 'common_hid_layers_num': 1,
  195. 'fcnn_hid_layers_num': 1,
  196. 'dropout': 0,
  197. 'learning_rate_SGD': 2e-1,
  198. 'weight_decay_SGD': 3e-4,
  199. 'momentum': 0.9,
  200. 'dampening': 0,
  201. 'nesterov': True,
  202. 'early_stopping_patience': 20,
  203. 'clip_grad_max_norm': 1,
  204. #'LambdaLR_scheduler_coefficient': 0.997,
  205. 'print_loss_epoch_step': 20,
  206. }
  207. '''
  208. ## run STdGCN
  209. Parameters
  210. 'load_test_groundtruth': [bool]. Select whether you need to upload the ground truth file (ST_ground_truth.tsv) of the spatial transcriptomics data to track the performance of STdGCN.
  211. 'use_marker_genes': [bool]. Select whether you need the gene selection process before running STdGCN. Otherwise use common genes from single cell and spatial transcriptomics data.
  212. 'external_genes': [bool]. When "use_marker_genes"=True, you can upload your specified gene list (marker_genes.tsv) to run STdGCN.
  213. 'generate_new_pseudo_spots': [bool]. STdGCN will save the simulated pseudo-spots to "pseudo_ST.pkl". If you want to run multiple deconvolutions with the same single cell reference data,
  214. you don't need to simulate new pseudo-spots and set 'generate_new_pseudo_spots'=False. When 'generate_new_pseudo_spots'=False, you need to pre-move the "pseudo_ST.pkl"
  215. to the 'output_path' so that STdGCN can directly load the pre-simulated pseudo-spots.
  216. 'fraction_pie_plot': [bool]. Select whether you need to draw the pie plot of the predicted results. Based on our experience, we do not recommend to draw the pie plot when the predicted
  217. spot number is very large. For 1,000 spots, the plotting time is less than 2 minutes; for 2,000 spots, the plotting time is about 10 minutes; for 3,000 spots, it takes
  218. about 30 minutes.
  219. 'cell_type_distribution_plot': [bool]. Select whether you need to draw the scatter plot of the predicted results for each cell type.
  220. 'n_jobs': [int]. Set the number of threads used for intraop parallelism on CPU. 'n_jobs=-1' represents using all CPUs.
  221. 'GCN_device': ['GPU', 'CPU']. Select the device used to run GCN networks.
  222. '''
  223. results = run_STdGCN(paths,
  224. load_test_groundtruth = False,
  225. use_marker_genes = True,
  226. external_genes = False,
  227. find_marker_genes_paras = find_marker_genes_paras,
  228. generate_new_pseudo_spots = True,
  229. pseudo_spot_simulation_paras = pseudo_spot_simulation_paras,
  230. data_normalization_paras = data_normalization_paras,
  231. integration_for_adj_paras = integration_for_adj_paras,
  232. inter_exp_adj_paras = inter_exp_adj_paras,
  233. spatial_adj_paras = spatial_adj_paras,
  234. real_intra_exp_adj_paras = real_intra_exp_adj_paras,
  235. pseudo_intra_exp_adj_paras = pseudo_intra_exp_adj_paras,
  236. integration_for_feature_paras = integration_for_feature_paras,
  237. GCN_paras = GCN_paras,
  238. fraction_pie_plot = True,
  239. cell_type_distribution_plot = True,
  240. n_jobs = -1,
  241. GCN_device = 'GPU'
  242. )
  243. results.write_h5ad(paths['output_path']+'/results.h5ad')
  244. # %%

Tutorial.ipynb at commit d02d390, no license · at the source

Overview

Authors: Zhan Xiao1, Wuke Wang2, Xin Long1,3, Wenbo Zhang1, Weiqiang Zhang1, Duoyuan Chen4, Sujie Xu1, Qian Yu1, Xinpeng Zhang5, Shichen Huang6, Ning Zhang1, Yanbin Yin5, Xingxu Huang2,7, Jieping Ye1, Jinfang Zheng1,8, Ling Guo1
  1. Research Center for Life Sciences Computing, Zhejiang Lab, Hangzhou, Zhejiang 311121, China
  2. Zhejiang Provincial Key Laboratory of Pancreatic Disease, The First Affiliated Hospital, and Institute of Translational Medicine, Zhejiang University School of Medicine, Hangzhou 311121, China
  3. Department of Laboratory Medicine of The First Affiliated Hospital & Liangzhu Laboratory, Zhejiang University School of Medicine, Hangzhou 311121, China
  4. Key Laboratory of Spatial Omics of Zhejiang Province, State Key Laboratory of Genome and Multi-omics Technologies, BGI Research, Hangzhou 310030, China
  5. Nebraska Food for Health Center, Department of Food Science and Technology, University of Nebraska, Lincoln, NE 68588, United States
  6. Department of Chemistry, The University of Manchester, Manchester M13 9PL, United Kingdom
  7. School of Life Science and Technology, ShanghaiTech University, Shanghai 200092, China
  8. Chongqing Institute of Intelligent Medicine, 799 Jingwei Avenue, Yuzhong District, Chongqing 400042, China
Journal: Nucleic acids research, volume 54, issue 13, article gkag706
Dates: received 13 October 2025; accepted 27 June 2026; published online 14 July 2026; in print July 2026
Type: Research article · Language: English
License: CC BY-NC
Identifiers: DOI 10.1093/nar/gkag706 · PMID 42444608 · PMCID PMC13366050 · OpenAlex W7168291358
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: Smoothing, state filtering, decompositions, Machine learning, Connectivity, fMRI & imaging
MeSH: Models, Biological*, Single-Cell Analysis*, Software*, Algorithms, Animals, Humans (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Provincial Fiscal Subsidy Funds for Major Science and Technology Innovation Platforms (2024SSYS0007); Pioneer, Leading Goose + X (2026SSYS0002)
Citations: not cited yet (Europe PMC); 92 references in the paper

Abstract

Virtual cells represent a promising paradigm to understand cellular mechanisms, behavior, and dynamics. The realization of virtual cells relies on the accurate modeling of cellular dynamics from large-scale, multi-modal single-cell data. However, experiment-specific technical noise and intrinsic biological heterogeneity pose major challenges for virtual cell modeling. To address this gap, we present scDifformer, a context-aware transformer model augmented with a denoising diffusion module and a dedicated post-training phase. This three-phase design, comprising masked language model pre-training, diffusion-driven post-training, and downstream fine-tuning, directly enhances scDifformer’s ability to denoise sparse, noisy data and generalize across studies. Benchmarking across seven tissues and multiple independent studies shows that the diffusion module consistently improves cross-dataset performance, particularly in settings with strong batch effects. By combining the strengths of transformer and diffusion models, scDifformer achieves state-of-the-art performance in cell type annotation across diverse datasets. It further demonstrates robust capability in resolving immune cell identities across multiple tissues, accurately recovering key marker genes, functional pathways, and cross-tissue differentiation trajectories. Finally, by integrating scDifformer with a graph neural network, we extend its utility to spatial transcriptomics, significantly enhancing spot-level deconvolution accuracy. Altogether, scDifformer provides a scalable and biologically grounded framework for modeling heterogeneous single-cell data, offering a powerful foundation for the development of high-fidelity, multi-modal virtual cell models.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 12 matches between paragraphs and lines of code.

huggingface.co/ctheodoris/geneformer

License: apache
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 1f7fbae4e469a5f4f1af8c111a529cfe1b3829f5, 12 September 2026
Languages: Python (22), Jupyter (8)
Size: 87 files, 30 scripts
Software Heritage: not archived
Found in: the text, “scRNA-seq cell annotation methods”
Holds: README, environment (requirements.txt, setup.py, docs/requirements.txt), documentation, 8 notebooks
Not found: license file, CITATION.cff, tests, continuous integration
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers

TencentAILabHealthcare/scBERT

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 262fd4b91f3f1c21a6e595d03d4ef423e16ffc99, 13 December 2023
Languages: Python (10)
Size: 13 files, 10 scripts
Software Heritage: not archived
Found in: the text, “scRNA-seq cell annotation methods”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (8 files), PyTorch (8 files), pandas (6 files), Scanpy (6 files), SciPy (6 files), anndata (5 files), scikit-learn (5 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
12 files

biomap-research/scFoundation

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 397631c495eddf9ad6644fc00c6ea8139e651245, 23 November 2025
Languages: Python (42), Shell (16), Jupyter (13)
Size: 364 files, 71 scripts
Software Heritage: archived
Found in: the text, “scRNA-seq cell annotation methods”
Holds: README, license file, 13 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (28 files), pandas (27 files), PyTorch (26 files), Scanpy (19 files), scikit-learn (12 files), SciPy (12 files), seaborn (8 files), Keras (4 files), PyTorch Geometric (4 files), Matplotlib (3 files), NetworkX (3 files), h5py (2 files), RDKit (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
73 files

biomed-AI/CellFM

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 72c9f4a9580a3716058c184900ed14a65151ed8f, 7 September 2026
Languages: Python (70), Jupyter (8), Shell (6)
Size: 169 files, 84 scripts
Software Heritage: not archived
Found in: the text, “scRNA-seq cell annotation methods”
Holds: README, license file, environment (tutorials/ChemicalPerturbation/requirements.txt, tutorials/ChemicalPerturbation/setup.py), 8 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (47 files), pandas (28 files), PyTorch (21 files), SciPy (20 files), Scanpy (19 files), scikit-learn (9 files), anndata (8 files), Matplotlib (4 files), PyTorch Geometric (4 files), NetworkX (2 files), seaborn (2 files), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
85 files

bowang-lab/scgpt

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: cebd6fae655b9c585a4807daa3ac31bb764f06b4, 27 April 2026
Languages: Python (42), Jupyter (10), Shell (5)
Size: 98 files, 57 scripts
Software Heritage: not archived
Found in: the text, “scRNA-seq cell annotation methods”
Holds: README, license file, environment (poetry.lock, pyproject.toml, docs/environment.yml, docs/requirements.txt), tests, continuous integration, documentation, 10 notebooks
Not found: CITATION.cff
Tools: NumPy (27 files), PyTorch (27 files), Scanpy (17 files), anndata (13 files), SciPy (12 files), Matplotlib (11 files), pandas (10 files), scikit-learn (8 files), seaborn (6 files), NetworkX (3 files), UMAP (3 files), h5py (1 file), Numba (1 file), PyTorch Geometric (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
59 files

broadinstitute/Tangram

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 4c68995a418f41dc8caef567598c4d9b47781a13, 1 July 2025
Languages: Python (18), Jupyter (2)
Size: 108 files, 20 scripts
Software Heritage: not archived
Found in: the text, “Spatial transcriptomics deconvolution methods”
Holds: README, license file, environment (environment.yml, setup.py, docs/requirements.txt), tests, continuous integration, documentation, 2 notebooks
Not found: CITATION.cff
Tools: NumPy (12 files), pandas (9 files), Scanpy (9 files), PyTorch (5 files), Matplotlib (4 files), SciPy (4 files), seaborn (4 files), scikit-learn (3 files), Squidpy (2 files), anndata (1 file), scikit-image (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
22 files

luoyuanlab/stdgcn

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: d02d39085bd297ba26212a07170821e03b644549, 6 May 2025
Languages: Python (9), Jupyter (1)
Size: 29 files, 10 scripts
Software Heritage: not archived
Found in: the text, “Spatial transcriptomics deconvolution methods”
Holds: README, 1 notebook
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (5 files), PyTorch (5 files), Scanpy (4 files), pandas (3 files), Matplotlib (2 files), scikit-learn (2 files), anndata (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
11 files

Su-informatics-lab/DSTG

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 4a2f958a87ba3137c3ffa75188563af3e0245a5a, 11 August 2022
Languages: Python (10), R (3)
Size: 18 files, 13 scripts
Software Heritage: not archived
Found in: the text, “Spatial transcriptomics deconvolution methods”
Holds: README, license file, environment (setup.py)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (5 files), pandas (5 files), scikit-learn (3 files), SciPy (3 files), TensorFlow (3 files), NetworkX (2 files), Seurat (1 file), tidyverse (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
15 files

bowang-lab/scGPT-spatial

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: f9442d777ba47cd3dcac8253833a233ff5bf7262, 12 February 2025
Languages: Python (18), Jupyter (1)
Size: 22 files, 19 scripts
Software Heritage: not archived
Found in: the text, “Spatial transcriptomics deconvolution methods”
Holds: README, license file, 1 notebook
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: PyTorch (13 files), NumPy (9 files), pandas (6 files), anndata (3 files), Scanpy (3 files), Matplotlib (2 files), SciPy (2 files), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
21 files

huggingface.co/allenxiao/scdifformer

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 54d1f5c9756909f9b4c83b7917844ef159d8cdca, 14 May 2026
Languages: Python (13)
Size: 18 files, 13 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (scDifformer_data_process/requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: pandas (4 files), NumPy (3 files), Scanpy (3 files), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
14 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 10 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 297 scripts, each with its path and the digest of its content;
  • 12 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The model weights and a minimun data of scDifformer are available in Huggingface project (https://huggingface.co/allenxiao/scDIFFormer). Source data are provided in this paper. The data processing code of scDifformer, together with training code for fine-tune models, is available at the Huggingface repository (https://huggingface.co/allenxiao/scDIFFormer).

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 16 authors, 6 MeSH terms, 2 funders, 82 references.

Cite

This paper

Xiao, Z., Wang, W., Long, X., Zhang, W., Zhang, W., Chen, D., Xu, S., Yu, Q., Zhang, X., Huang, S., Zhang, N., Yin, Y., Huang, X., Ye, J., Zheng, J., & Guo, L. (2026). scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data. Nucleic acids research, 54(13), gkag706. https://doi.org/10.1093/nar/gkag706

BibTeX

@article{xiao2026scdifformer,
author = {Xiao, Zhan and Wang, Wuke and Long, Xin and Zhang, Wenbo and Zhang, Weiqiang and Chen, Duoyuan and Xu, Sujie and Yu, Qian and Zhang, Xinpeng and Huang, Shichen and Zhang, Ning and Yin, Yanbin and Huang, Xingxu and Ye, Jieping and Zheng, Jinfang and Guo, Ling},
title = {{scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data}},
journal = {Nucleic acids research},
year = {2026},
month = jul,
volume = {54},
number = {13},
pages = {gkag706},
publisher = {Oxford University Press},
issn = {0305-1048},
doi = {10.1093/nar/gkag706},
url = {https://doi.org/10.1093/nar/gkag706},
pmid = {42444608},
pmcid = {PMC13366050}
}

RIS

TY - JOUR
AU - Xiao, Zhan
AU - Wang, Wuke
AU - Long, Xin
AU - Zhang, Wenbo
AU - Zhang, Weiqiang
AU - Chen, Duoyuan
AU - Xu, Sujie
AU - Yu, Qian
AU - Zhang, Xinpeng
AU - Huang, Shichen
AU - Zhang, Ning
AU - Yin, Yanbin
AU - Huang, Xingxu
AU - Ye, Jieping
AU - Zheng, Jinfang
AU - Guo, Ling
TI - scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data
T2 - Nucleic acids research
J2 - Nucleic Acids Res
PY - 2026
DA - 2026/07/01
VL - 54
IS - 13
SP - gkag706
SN - 0305-1048
PB - Oxford University Press
DO - 10.1093/nar/gkag706
UR - https://doi.org/10.1093/nar/gkag706
LA - en
ER -

CSL-JSON

{
"id": "10.1093/nar/gkag706",
"type": "article-journal",
"title": "scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data",
"container-title": "Nucleic acids research",
"author": [
{
"family": "Xiao",
"given": "Zhan"
},
{
"family": "Wang",
"given": "Wuke"
},
{
"family": "Long",
"given": "Xin"
},
{
"family": "Zhang",
"given": "Wenbo"
},
{
"family": "Zhang",
"given": "Weiqiang"
},
{
"family": "Chen",
"given": "Duoyuan"
},
{
"family": "Xu",
"given": "Sujie"
},
{
"family": "Yu",
"given": "Qian"
},
{
"family": "Zhang",
"given": "Xinpeng"
},
{
"family": "Huang",
"given": "Shichen"
},
{
"family": "Zhang",
"given": "Ning"
},
{
"family": "Yin",
"given": "Yanbin"
},
{
"family": "Huang",
"given": "Xingxu"
},
{
"family": "Ye",
"given": "Jieping"
},
{
"family": "Zheng",
"given": "Jinfang"
},
{
"family": "Guo",
"given": "Ling"
}
],
"container-title-short": "Nucleic Acids Res",
"volume": "54",
"issue": "13",
"page": "gkag706",
"DOI": "10.1093/nar/gkag706",
"PMID": "42444608",
"PMCID": "PMC13366050",
"ISSN": "0305-1048",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/nar/gkag706",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s44320-026-00208-7 [code]
Interpretable deep generative ensemble learning for single-cell omics with Hydra.
Journal: Molecular systems biology
In common: Keras, UMAP, anndata, 13 other tools, 7 references
[2] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: Squidpy, PyTorch Geometric, UMAP, 14 other tools, 3 references
[3] doi:10.1016/j.xgen.2026.101217 [code]
ProtoCloud: A prototypical self-explaining model for single-cell analysis.
Journal: Cell genomics
In common: UMAP, anndata, Scanpy, 8 other tools, 9 references
[4] doi:10.1038/s41467-026-71759-4 [code]
CellNiche represents cellular microenvironments in atlas-scale spatial omics data with contrastive learning.
Journal: Nature communications
In common: Squidpy, PyTorch Geometric, anndata, 9 other tools, 6 references
[5] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: Keras, UMAP, anndata, 14 other tools, 2 references
[6] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: Squidpy, PyTorch Geometric, UMAP, 14 other tools
[7] doi:10.1038/s41467-026-68596-w [code]
Spatial cartography of human thymus enables the geopositioning of lineage transcription factors in rare mimetic thymic epithelial cells.
Journal: Nature communications
In common: Squidpy, anndata, Scanpy, 12 other tools, 3 references
[8] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: RDKit, Keras, UMAP, 14 other tools
[9] doi:10.1093/bib/bbag404 [code]
Navigating cell maps by deep learning integration of single-cell and spatially resolved transcriptomics.
Journal: Briefings in bioinformatics
In common: PyTorch Geometric, anndata, Scanpy, 9 other tools, 5 references
[10] doi:10.1186/s12864-026-12965-8 [code]
Systematic evaluation of single-cell foundation model interpretability: attention-derived edge scores add no incremental value over gene-level features for perturbation-target prediction.
Journal: BMC genomics
In common: anndata, Scanpy, NetworkX, 9 other tools, 5 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.