OSCR

SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.

Code ↔ Paper

12 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 12 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Results and discussion › SpaDC accurately detects brain structures and reveals spatial domain specific regulatory networks ↔ Tutorials/Tutorial_P22.ipynb, lines 128–157 · score 0.81 · 23687043–23687913, 25317719–25318626, 70998671–70999544, 83126754–83127607, chr14, chr18
  2. [2] § Results and discussion › SpaDC accurately captures spatial structures of the mouse embryonic brain ↔ Tutorials/Tutorial_MISAR.ipynb, lines 86–115 · score 0.81 · Cerebellar Vermis, DPallm, DPallv, Diencephalon, Hindbrain, Midbrain
  3. [3] § Methods › Graph regularized convolutional neural network ↔ SpaDC/model.py, lines 18–88 · score 0.79 · ReLU, sigmoid, flattened, kernel, tower, Linear
  4. [4] § Methods › GRN inference ↔ GRN_R/get_peak_motif_perturbation.R, lines 1–57 · score 0.72 · matched motif, DNA sequence, perturbed, positions, bp, genome
  5. [5] § Results and discussion › Overview of SpaDC ↔ SpaDC/train_SpaDC_bc.py, lines 114–179 · score 0.67 · binary cross entropy, triplet loss, BCE, anchors, trained, nearest
  6. [6] § Methods › Benchmarking metrics › Batch entropy mixing score ↔ SpaDC/utils.py, lines 328–389 · score 0.63 · batch entropy mixing, Iteratively, score, cell
  7. [7] § Results and discussion › Overview of SpaDC ↔ SpaDC/train_SpaDC.py, the whole file · a weak match · score 0.60 · binary cross entropy, BCE, shuffled, trained, loss, graph
  8. [8] § Methods › GRN inference ↔ Tutorials/Inferring_GRN_on_P22.ipynb, lines 134–237 · score 0.60 · XGBoost, functional CREs, regression, Shapley, gene, denoised
  9. [9] § Methods › Multiple spatial epigenomics data alignment ↔ SpaDC/train_SpaDC_bc.py, lines 114–179 · score 0.58 · Triplet loss, cell embeddings, margin, anchor, batch
  10. [10] § Results and discussion › SpaDC accurately captures spatial structures of the mouse embryonic brain ↔ Tutorials/Tutorial_MISAR.ipynb, lines 86–115 · score 0.55 · DPallm, DPallv, Mesenchyme, Subpallium, SpaDC
  11. [11] § Methods › Graph regularized convolutional neural network ↔ SpaDC/train_SpaDC.py, the whole file · a weak match · score 0.52 · BCE loss, optimized, weight, graph, neighbor, matrix
  12. [12] § Methods › Benchmarking metrics ↔ SpaDC/utils.py, lines 328–389 · score 0.50 · batch entropy mixing, score, clustering

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 213 lines · 6.5 KB · MIT · 2 matches

  1. # %% [markdown]
  2. # # Run SpaDC on spatial ATAC-seq data of E15.5 mouse embryonic brain
  3. # %% [markdown]
  4. # ### The data is available at: https://doi.org/10.6084/m9.figshare.30739373
  5. # %%
  6. import scanpy as sc
  7. import pandas as pd
  8. import SpaDC
  9. import torch
  10. import os
  11. # %% [markdown]
  12. # # Preparation of data
  13. # %% [markdown]
  14. # ##### The mm10.fa.gz reference genome file is not included in this repository due to its large file size. Please download it from the UCSC Genome Browser:
  15. #
  16. # https://hgdownload.soe.ucsc.edu/goldenPath/mm10/bigZips/mm10.fa.gz
  17. #
  18. # After downloading, decompress the file and place mm10.fa under the ./SpaDC/ directory before running the pipeline.
  19. # %% [markdown]
  20. # # Importing the data
  21. # %%
  22. # Load adata
  23. DATA_DIR = "./SpaDC/MISAR_seq"
  24. adata_path = os.path.join(DATA_DIR, "MISAR_seq_mouse_E15_brain_ATAC_data.h5ad") ### The adata is processed with EpiScanpy, just remove the low quality peaks and spots with min_features=10 and min_cells=10
  25. adata = sc.read_h5ad(adata_path)
  26. # %% [markdown]
  27. # ##### Note: The required information in adata are: spatial coordinates, peak-by-spot data, peaks
  28. # %% [markdown]
  29. # # Converting chromosome names to UCSC format
  30. # %% [markdown]
  31. # ##### If you would like to use the method quickly, you may skip the following section.
  32. # %% [markdown]
  33. # ##### You can directly load the preprocessed sequence data from our dataset using the "Loading standard sequencing" part.
  34. # %%
  35. index = adata.var_names
  36. index = pd.DataFrame(x.split('-') for x in index)
  37. seq, _ = SpaDC.make_bed_seqs_from_df(index, './SpaDC/mm10.fa', 1344)
  38. file = open(os.path.join(DATA_DIR, "seqs.txt"),'a')
  39. file.write('seq\n')
  40. for i in range(len(seq)):
  41. s = seq[i] + '\n'
  42. file.write(s)
  43. file.close()
  44. # %% [markdown]
  45. # #### Loading standard sequencing
  46. # %%
  47. seq_path = os.path.join(DATA_DIR, "seqs.txt")
  48. seq = pd.read_csv(seq_path, sep='\t')
  49. # %% [markdown]
  50. # # Running SpaDC
  51. # %% [markdown]
  52. # ##### After running SpaDC, the learned spot embeddings will be stored in adata.obsm['SpaDC']
  53. # %% [markdown]
  54. # ##### The trained model parameters will be saved as model.pt in the current working directory.
  55. # %%
  56. adata = SpaDC.train_SpaDC(adata, seq, n_epochs=200, batch_size=256, save_model=True, out_dir='', device=torch.device('cuda:8'))
  57. # %% [markdown]
  58. # # Clustering on SpaDC's embedding
  59. # %%
  60. sc.pp.neighbors(adata, use_rep='SpaDC')
  61. res, _ = SpaDC.getNClusters(adata, 14)
  62. sc.tl.leiden(adata, key_added='SpaDC', resolution=res)
  63. # %%
  64. # Visualization
  65. import matplotlib.pyplot as plt
  66. plt.rcParams['font.sans-serif'] = 'Arial'
  67. plt.rcParams['font.family'] = 'sans-serif'
  68. # Map SpaDC clusters to annotated spot types
  69. cluster_annotation = {
  70. '0':'1-Hindbrain',
  71. '1':'2-Skull',
  72. '2':'3-Cartilage',
  73. '3':'4-Diencephalon',
  74. '4':'5-Mesenchyme',
  75. '5':'6-DPallm',
  76. '6':'7-Midbrain',
  77. '7':'8-Cartilage',
  78. '8':'9-Subpallium',
  79. '9':'10-Cerebellar vermis',
  80. '10':'11-Thalamus',
  81. '11':'12-DPallv',
  82. '12':'13-Muscle',
  83. '13':'14-Hindbrain',
  84. }
  85. adata.obs['SpaDC'] = adata.obs['SpaDC'].map(cluster_annotation).astype('category')
  86. f, ax = plt.subplots(figsize=(6, 4))
  87. sc.pl.embedding(adata, basis='spatial1', color='SpaDC', size=100, ax=ax, show=False)
  88. plt.tight_layout()
  89. plt.show()
  90. # %% [markdown]
  91. # # Visualizing the denosing performance of SpaDC
  92. # %% [markdown]
  93. # ### before denoising
  94. # %%
  95. import episcanpy as epi
  96. epi.pp.binarize(adata)
  97. epi.pp.normalize_total(adata)
  98. epi.pp.log1p(adata)
  99. f, ax = plt.subplots(1, 4, figsize=(14, 3))
  100. sc.pl.embedding(adata, basis='spatial1', color='chr1-120602010-120602510', vmin=0, size=60, vmax=1.5, ax=ax[0], show=False, colorbar_loc=None)
  101. sc.pl.embedding(adata, basis='spatial1', color='chr12-6638859-6639359', vmin=0, size=60, vmax=1.5, ax=ax[1], show=False, colorbar_loc=None)
  102. sc.pl.embedding(adata, basis='spatial1', color='chr13-59971685-59972185', vmin=0, size=60, vmax=1.5, ax=ax[2], show=False, colorbar_loc=None)
  103. sc.pl.embedding(adata, basis='spatial1', color='chr5-131596826-131597326', vmin=0, size=60, vmax=1.5, ax=ax[3], show=False, colorbar_loc=None)
  104. ax[0].set_xlabel('')
  105. ax[1].set_xlabel('')
  106. ax[2].set_xlabel('')
  107. ax[3].set_xlabel('')
  108. ax[0].set_ylabel('Raw', fontsize=12)
  109. ax[1].set_ylabel('')
  110. ax[2].set_ylabel('')
  111. ax[3].set_ylabel('')
  112. plt.show()
  113. # %% [markdown]
  114. # ### after denoising
  115. # %%
  116. import squidpy as sq
  117. adata = sc.read_h5ad(adata_path)
  118. model_state_dict = './model.pt'
  119. # get denoise adata
  120. adata_denoise = SpaDC.get_denoise_adata(adata, seq, model_state_dict)
  121. adata_denoise.obsm['spatial'] = adata.obsm['spatial1']
  122. sc.pp.normalize_total(adata, target_sum=1e4)
  123. sc.pp.log1p(adata)
  124. sc.tl.pca(adata)
  125. sc.pp.neighbors(adata, n_neighbors=15, n_pcs=30)
  126. sq.gr.spatial_neighbors(adata, coord_type='grid', spatial_key='spatial1')
  127. adata.obsp['connectivities'] = (0.5 * adata.obsp['connectivities'] + 0.5 * adata.obsp['spatial_connectivities'])
  128. res1, _ = SpaDC.getNClusters(adata, 14)
  129. sc.tl.leiden(adata, key_added='Raw_clusters', resolution=res1)
  130. sc.pp.normalize_total(adata_denoise, target_sum=1e4)
  131. sc.pp.log1p(adata_denoise)
  132. sc.tl.pca(adata_denoise)
  133. sc.pp.neighbors(adata_denoise, n_neighbors=15, n_pcs=30)
  134. sq.gr.spatial_neighbors(adata_denoise, coord_type='grid', spatial_key='spatial')
  135. adata_denoise.obsp['connectivities'] = (0.5 * adata_denoise.obsp['connectivities'] + 0.5 * adata_denoise.obsp['spatial_connectivities'])
  136. res2, _ = SpaDC.getNClusters(adata_denoise, 14)
  137. sc.tl.leiden(adata_denoise, key_added='Denoise_clusters', resolution=res2)
  138. f, ax = plt.subplots(1, 2, figsize=(9, 4))
  139. sc.pl.embedding(adata, basis='spatial1', color='Raw_clusters', size=100, ax=ax[0], show=False)
  140. sc.pl.embedding(adata_denoise, basis='spatial', color='Denoise_clusters', size=100, ax=ax[1], show=False)
  141. ax[0].set_xlabel('')
  142. ax[1].set_xlabel('')
  143. ax[0].set_ylabel('')
  144. ax[1].set_ylabel('')
  145. plt.tight_layout()
  146. plt.show()
  147. # %%
  148. adata.X = adata_denoise.X
  149. f, ax = plt.subplots(1, 4, figsize=(14, 3))
  150. sc.pl.embedding(adata, basis='spatial1', color='chr1-120602010-120602510', vmin=0, size=60, ax=ax[0], show=False, colorbar_loc=None)
  151. sc.pl.embedding(adata, basis='spatial1', color='chr12-6638859-6639359', vmin=0, size=60, ax=ax[1], show=False, colorbar_loc=None)
  152. sc.pl.embedding(adata, basis='spatial1', color='chr13-59971685-59972185', vmin=0, size=60, ax=ax[2], show=False, colorbar_loc=None)
  153. sc.pl.embedding(adata, basis='spatial1', color='chr5-131596826-131597326', vmin=0, size=60, ax=ax[3], show=False, colorbar_loc=None)
  154. ax[0].set_xlabel('')
  155. ax[1].set_xlabel('')
  156. ax[2].set_xlabel('')
  157. ax[3].set_xlabel('')
  158. ax[0].set_ylabel('Denoise', fontsize=12)
  159. ax[1].set_ylabel('')
  160. ax[2].set_ylabel('')
  161. ax[3].set_ylabel('')
  162. plt.show()

Tutorial_MISAR.ipynb at commit bcf4ad8, under MIT · at the source

Overview

Authors: Chuanlong Ma1, Chenghui Yang2, Caiwei Zhen3, Zhentao He3, Yong Luo3, Lihua Zhang2,3
  1. School of Cyber Science and Engineering, Wuhan University,Wuhan, China
  2. School of Artificial Intelligence, Wuhan University,Wuhan, China
  3. School of Computer Science, Wuhan University,Wuhan, China
Institutions: Wuhan University (China)
Journal: Communications biology, volume 9, issue 1, article 1196
Dates: received 8 December 2025; accepted 2 June 2026; published online 6 June 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s42003-026-10462-y · PMID 42251185 · PMCID PMC13575232 · OpenAlex W7163750264
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), mouse (organism)
Methods: Preprocessing, Machine learning
Keywords: Computational models, Data mining
MeSH: Chromatin*, Chromatin Immunoprecipitation Sequencing*, Computational Biology*, Gene Regulatory Networks*, Animals, Brain, Mice (* major topic)
Topic: Genomics and Chromatin Dynamics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: cited by 1 paper (Europe PMC); 60 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 12 matches between paragraphs and lines of code.

mcllllllll/SpaDC

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bcf4ad8c408282a1954d10222c7a93fb16a56039, 10 May 2026
Languages: Python (6), Jupyter (4), R (3)
Size: 18 files, 13 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (setup.py), 4 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: PyTorch (7 files), pandas (5 files), Scanpy (5 files), Matplotlib (4 files), NumPy (3 files), SciPy (3 files), BEDTools (2 files), scikit-learn (2 files), anndata (1 file), Biopython (1 file), NetworkX (1 file), pysam (1 file), SHAP (1 file), Squidpy (1 file), tidyverse (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
15 files

Zenodo 20300757

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: PyTorch (7 files), pandas (5 files), Scanpy (5 files), Matplotlib (4 files), NumPy (3 files), SciPy (3 files), BEDTools (2 files), scikit-learn (2 files), anndata (1 file), Biopython (1 file), NetworkX (1 file), pysam (1 file), SHAP (1 file), Squidpy (1 file), tidyverse (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
15 files
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s42003-026-10462-y.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 26 scripts, each with its path and the digest of its content;
  • 12 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s42003-026-10462-y.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 2 keywords, 7 MeSH terms, 1 funder, 57 references.

Cite

This paper

Ma, C., Yang, C., Zhen, C., He, Z., Luo, Y., & Zhang, L. (2026). SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data. Communications biology, 9(1), 1196. https://doi.org/10.1038/s42003-026-10462-y

BibTeX

@article{ma2026spadc,
author = {Ma, Chuanlong and Yang, Chenghui and Zhen, Caiwei and He, Zhentao and Luo, Yong and Zhang, Lihua},
title = {{SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data}},
journal = {Communications biology},
year = {2026},
month = jun,
volume = {9},
number = {1},
pages = {1196},
publisher = {Nature Publishing Group},
issn = {2399-3642},
doi = {10.1038/s42003-026-10462-y},
url = {https://doi.org/10.1038/s42003-026-10462-y},
pmid = {42251185},
pmcid = {PMC13575232}
}

RIS

TY - JOUR
AU - Ma, Chuanlong
AU - Yang, Chenghui
AU - Zhen, Caiwei
AU - He, Zhentao
AU - Luo, Yong
AU - Zhang, Lihua
TI - SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data
T2 - Communications biology
J2 - Commun Biol
PY - 2026
DA - 2026/06/06
VL - 9
IS - 1
SP - 1196
SN - 2399-3642
PB - Nature Publishing Group
DO - 10.1038/s42003-026-10462-y
UR - https://doi.org/10.1038/s42003-026-10462-y
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s42003-026-10462-y",
"type": "article-journal",
"title": "SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data",
"container-title": "Communications biology",
"author": [
{
"family": "Ma",
"given": "Chuanlong"
},
{
"family": "Yang",
"given": "Chenghui"
},
{
"family": "Zhen",
"given": "Caiwei"
},
{
"family": "He",
"given": "Zhentao"
},
{
"family": "Luo",
"given": "Yong"
},
{
"family": "Zhang",
"given": "Lihua"
}
],
"container-title-short": "Commun Biol",
"volume": "9",
"issue": "1",
"page": "1196",
"DOI": "10.1038/s42003-026-10462-y",
"PMID": "42251185",
"PMCID": "PMC13575232",
"ISSN": "2399-3642",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s42003-026-10462-y",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
6
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: pysam, Biopython, BEDTools, 8 other tools, genetics / omics, mouse, 12 references
[2] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: Squidpy, pysam, BEDTools, 8 other tools, genetics / omics, mouse, 7 references
[3] doi:10.1038/s41467-026-73171-4 [code]
Dissecting epigenetic heterogeneity in single-cell DNA methylomes with a unified framework.
Journal: Nature communications
In common: pysam, BEDTools, anndata, 7 other tools, genetics / omics, 7 references
[4] doi:10.1038/s41467-026-71803-3 [code]
Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.
Journal: Nature communications
In common: Squidpy, BEDTools, anndata, 8 other tools, genetics / omics, mouse, 5 references
[5] doi:10.1093/bioinformatics/btag652 [code]
mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.
Journal: Bioinformatics (Oxford, England)
In common: pysam, BEDTools, anndata, 7 other tools, genetics / omics, mouse, 5 references
[6] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: Squidpy, anndata, Scanpy, 7 other tools, genetics / omics, 5 references
[7] doi:10.1016/j.celrep.2026.117073 [code]
Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.
Journal: Cell reports
In common: pysam, BEDTools, anndata, 7 other tools, genetics / omics, mouse, 4 references
[8] doi:10.1038/s41467-026-71759-4 [code]
CellNiche represents cellular microenvironments in atlas-scale spatial omics data with contrastive learning.
Journal: Nature communications
In common: Squidpy, anndata, Scanpy, 7 other tools, mouse, 5 references
[9] doi:10.1038/s41592-026-03211-w [code]
Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.
Journal: Nature methods
In common: pysam, Biopython, BEDTools, 8 other tools, genetics / omics, mouse, 2 references
[10] doi:10.64898/2026.03.30.714220 [code]
An integrated single cell and spatial omics atlas of human prenatal development
Journal: bioRxiv (preprint)
In common: Squidpy, Biopython, anndata, 8 other tools, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.