OSCR

Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Methods › Bulk ATAC-seq data processing ↔ hicompass/preprocess/atac.py, lines 299–375 · score 0.73 · bedGraph, BigWig, genomecov, bedtools, pipeline, BAM
  2. [2] § Methods › Model architecture ↔ hicompass/predicting/blocks.py, lines 95–129 · score 0.72 · DNAEncoder, ReLU, feature weighting, channel, batch, modules
  3. [3] § Methods › Meta-cell strategy and meta-cell Hi-C prediction ↔ hicompass/preprocess/atac.py, lines 299–375 · score 0.72 · bedGraph, BigWig, genomecov, bedtools, bw, bam
  4. [4] § Methods › Model architecture ↔ hicompass/predicting/blocks.py, lines 34–92 · score 0.68 · ATACDepthEncoder, depth feature, depth range, embedding, module
  5. [5] § Methods › Model training strategy ↔ hicompass/train/HicompassTrain.py, lines 206–283 · score 0.67 · cross entropy losses, MSE loss, SSIM, batch, training, model
  6. [6] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/predicting/PredictDataset.py, lines 64–205 · score 0.64 · DNA sequences, sequencing depth, genomic features, generalized CTCF, ATAC seq, error
  7. [7] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/predicting/PredictDataset.py, lines 64–205 · score 0.62 · DNA sequence, sequencing depth, genomic features, generalized CTCF, ATAC seq, genome
  8. [8] § Methods › Performance comparison with previous methods ↔ hicompass/train/HicompassModel.py, lines 47–114 · score 0.61 · AdaptiveAvgPool2d, insulation score, PyTorch, diagonal, models, training
  9. [9] § Methods › Model architecture ↔ hicompass/train/HicompassModel.py, lines 47–114 · score 0.60 · ResNet, adaptive, PyTorch, decoder, discriminator, modeling
  10. [10] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/train/HicompassTrain.py, lines 206–283 · score 0.60 · cross entropy, information weighted, discriminator, MSE, SSIM, loss
  11. [11] § Methods › Hi-C data processing ↔ hicompass/preprocess/hic_norm.py, lines 147–227 · score 0.58 · kb resolution, contrast stretching, mouse, human, hg38, optimize
  12. [12] § Results › Systematic evaluation of Hi-Compass performance ↔ hicompass/preprocess/hic_norm.py, lines 450–557 · score 0.56 · upper triangle, lower triangle, contrast stretched, bands, diagonals, Window
  13. [13] § Methods › Hi-C prediction and integration ↔ hicompass/predicting/utils.py, lines 113–178 · score 0.53 · create_cooler, prediction window, matrix, resolution, chromosomes, Hi
  14. [14] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/train/HicompassDataset.py, lines 535–581 · score 0.53 · DNA sequence, genomic features, generalized CTCF, ATAC seq, genome, depth
  15. [15] § Methods › CTCF ChIP-seq data processing ↔ hicompass/preprocess/atac.py, lines 272–297 · score 0.50 · bedGraph

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 545 lines · 20 KB · MIT · 3 matches

  1. #!/usr/bin/env python3
  2. """
  3. ATAC-seq preprocessing for Hi-Compass training data preparation.
  4. This module handles:
  5. 1. Stratified subsampling from bulk ATAC-seq BAM files
  6. 2. Conversion to BigWig format with proper chromosome filtering
  7. 3. Hierarchical directory organization for training data
  8. """
  9. import os
  10. import subprocess
  11. from pathlib import Path
  12. from typing import List, Optional, Union
  13. import logging
  14. logging.basicConfig(
  15. level=logging.INFO,
  16. format='%(asctime)s - %(levelname)s - %(message)s'
  17. )
  18. logger = logging.getLogger(__name__)
  19. class ATACPreprocessor:
  20. """Preprocesses ATAC-seq data for Hi-Compass training."""
  21. def __init__(
  22. self,
  23. input_bam: Union[str, Path],
  24. cell_type: str,
  25. output_dir: Union[str, Path],
  26. chrom_sizes: Union[str, Path],
  27. genome: str = 'hg38',
  28. depth_mode: str = 'range',
  29. depth_list: Optional[List[Union[int, float]]] = None,
  30. min_depth: float = 2e5,
  31. max_depth: float = 2e7,
  32. step: float = 2e4,
  33. seed: int = 114514,
  34. keep_intermediate: bool = False
  35. ):
  36. """
  37. Initialize ATAC preprocessor.
  38. Args:
  39. input_bam: Path to input BAM/SAM file (bulk ATAC-seq)
  40. cell_type: Cell line name (e.g., 'GM12878', 'K562')
  41. output_dir: Output directory for processed files
  42. chrom_sizes: Chromosome sizes file (e.g., hg38.chrom.sizes)
  43. genome: Genome build ('hg38', 'mm10', etc.) for chromosome filtering
  44. depth_mode: 'range' or 'list' for depth specification
  45. depth_list: Custom depths (for 'list' mode)
  46. min_depth: Minimum depth (for 'range' mode, default: 2e5)
  47. max_depth: Maximum depth (for 'range' mode, default: 2e7)
  48. step: Step size (for 'range' mode, default: 2e4)
  49. seed: Random seed for reproducibility
  50. keep_intermediate: Keep intermediate files (BAM, bedGraph)
  51. """
  52. self.input_bam = Path(input_bam)
  53. self.cell_type = self._sanitize_cell_type(cell_type)
  54. self.output_dir = Path(output_dir)
  55. self.chrom_sizes = Path(chrom_sizes)
  56. self.genome = genome
  57. self.depth_mode = depth_mode
  58. self.depth_list = [int(d) for d in depth_list] if depth_list else None
  59. self.min_depth = int(min_depth)
  60. self.max_depth = int(max_depth)
  61. self.step = int(step)
  62. self.seed = seed
  63. self.keep_intermediate = keep_intermediate
  64. # Define valid chromosomes based on genome
  65. self._set_valid_chromosomes()
  66. self.output_dir.mkdir(parents=True, exist_ok=True)
  67. self._validate_inputs()
  68. self._total_reads = None
  69. def _set_valid_chromosomes(self):
  70. """Set valid chromosome names based on genome build."""
  71. if self.genome == 'hg38':
  72. # Human: chr1-22, chrX, chrY
  73. self.valid_chroms = set([f'chr{i}' for i in range(1, 23)] + ['chrX', 'chrY'])
  74. elif self.genome == 'mm10':
  75. # Mouse: chr1-19, chrX, chrY
  76. self.valid_chroms = set([f'chr{i}' for i in range(1, 20)] + ['chrX', 'chrY'])
  77. else:
  78. # For custom genomes, read from chrom_sizes file
  79. logger.warning(f"Custom genome '{self.genome}' - will use all chromosomes from chrom_sizes")
  80. self.valid_chroms = None # Will be set from chrom_sizes file
  81. @staticmethod
  82. def _sanitize_cell_type(cell_type: str) -> str:
  83. """Sanitize cell type to valid filename characters."""
  84. import re
  85. return re.sub(r'[^a-zA-Z0-9_-]', '_', cell_type)
  86. def _validate_inputs(self):
  87. """Validate inputs and check required tools."""
  88. if not self.input_bam.exists():
  89. raise FileNotFoundError(f"Input BAM not found: {self.input_bam}")
  90. if not self.chrom_sizes.exists():
  91. raise FileNotFoundError(f"Chrom sizes not found: {self.chrom_sizes}")
  92. # Load valid chromosomes from chrom_sizes if not set
  93. if self.valid_chroms is None:
  94. self.valid_chroms = set()
  95. with open(self.chrom_sizes, 'r') as f:
  96. for line in f:
  97. chrom = line.strip().split('\t')[0]
  98. self.valid_chroms.add(chrom)
  99. logger.info(f"Loaded {len(self.valid_chroms)} chromosomes from {self.chrom_sizes.name}")
  100. # Validate depth configuration
  101. if self.depth_mode == 'list':
  102. if not self.depth_list:
  103. raise ValueError("depth_list required for mode 'list'")
  104. if any(d <= 0 for d in self.depth_list):
  105. raise ValueError("All depths must be positive")
  106. elif self.depth_mode == 'range':
  107. if self.min_depth >= self.max_depth:
  108. raise ValueError("min_depth must be < max_depth")
  109. else:
  110. raise ValueError(f"Invalid depth_mode: {self.depth_mode}")
  111. # Check required tools
  112. required_tools = ['samtools', 'bedtools', 'bedGraphToBigWig']
  113. missing_tools = []
  114. for tool in required_tools:
  115. if subprocess.run(['which', tool], capture_output=True).returncode != 0:
  116. missing_tools.append(tool)
  117. if missing_tools:
  118. raise EnvironmentError(
  119. f"Required tools not found: {', '.join(missing_tools)}\n"
  120. f"Please install: samtools, bedtools, bedGraphToBigWig (UCSC tools)"
  121. )
  122. def _get_total_reads(self) -> int:
  123. """Count total reads in BAM file."""
  124. if self._total_reads is None:
  125. logger.info("Counting reads in input BAM...")
  126. result = subprocess.run(
  127. ['samtools', 'view', '-c', str(self.input_bam)],
  128. capture_output=True, text=True, check=True
  129. )
  130. self._total_reads = int(result.stdout.strip())
  131. logger.info(f"Total reads in {self.input_bam.name}: {self._total_reads:,}")
  132. return self._total_reads
  133. @staticmethod
  134. def _format_depth(depth: int) -> str:
  135. """
  136. Format depth as scientific notation (e.g., 1e5, 2e6, 2.1e5).
  137. Examples:
  138. 100000 -> '1e5'
  139. 210000 -> '2.1e5'
  140. 2000000 -> '2e6'
  141. """
  142. exp = 0
  143. value = float(depth)
  144. while value >= 10:
  145. value /= 10
  146. exp += 1
  147. if value == int(value):
  148. return f"{int(value)}e{exp}"
  149. return f"{value:.1f}e{exp}".rstrip('0').rstrip('.')
  150. def _generate_filename(self, depth: Union[int, str], extension: str) -> str:
  151. """
  152. Generate standardized filename: {CellType}~ATAC~{Depth}.{ext}
  153. Examples:
  154. GM12878, 210000, 'bw' -> 'GM12878~ATAC~2.1e5.bw'
  155. K562, 'bulk', 'bam' -> 'K562~ATAC~bulk.bam'
  156. """
  157. if isinstance(depth, int):
  158. depth_str = self._format_depth(depth)
  159. else:
  160. depth_str = depth
  161. return f"{self.cell_type}~ATAC~{depth_str}.{extension}"
  162. def _get_depth_dir(self, depth: Union[int, str]) -> Path:
  163. """
  164. Get subdirectory for a specific depth.
  165. Creates hierarchical structure:
  166. output_dir/
  167. └── CellType~ATAC~Depth/
  168. ├── CellType~ATAC~Depth.bam
  169. ├── CellType~ATAC~Depth.sorted.bam
  170. ├── CellType~ATAC~Depth.bedgraph
  171. ├── CellType~ATAC~Depth.sorted.bedgraph
  172. └── CellType~ATAC~Depth.bw
  173. """
  174. if isinstance(depth, int):
  175. depth_str = self._format_depth(depth)
  176. else:
  177. depth_str = depth
  178. depth_dir = self.output_dir / f"{self.cell_type}~ATAC~{depth_str}"
  179. depth_dir.mkdir(parents=True, exist_ok=True)
  180. return depth_dir
  181. def get_target_depths(self) -> List[int]:
  182. """Get list of target depths based on configuration."""
  183. if self.depth_mode == 'list':
  184. depths = sorted(self.depth_list)
  185. else:
  186. depths = list(range(self.min_depth, self.max_depth + 1, self.step))
  187. logger.info(f"Target depths: {len(depths)} levels")
  188. return depths
  189. def subsample_bam(self, target_depth: int) -> Path:
  190. """
  191. Subsample BAM to target depth using samtools.
  192. Args:
  193. target_depth: Target number of reads
  194. Returns:
  195. Path to subsampled BAM in depth-specific subdirectory
  196. """
  197. total_reads = self._get_total_reads()
  198. if target_depth >= total_reads:
  199. logger.warning(
  200. f"Target {target_depth:,} >= total {total_reads:,}, using full BAM"
  201. )
  202. return self.input_bam
  203. fraction = target_depth / total_reads
  204. # Get depth-specific directory
  205. depth_dir = self._get_depth_dir(target_depth)
  206. output_bam = depth_dir / self._generate_filename(target_depth, 'bam')
  207. logger.info(f"Subsampling to {target_depth:,} reads (fraction={fraction:.6f})")
  208. # Format for samtools -s: seed.fraction
  209. fraction_str = f"{fraction:.10f}"
  210. if fraction_str.startswith('0.'):
  211. fraction_decimal = fraction_str[2:] # Remove "0."
  212. else:
  213. fraction_decimal = fraction_str
  214. fraction_decimal = fraction_decimal.rstrip('0') or '0'
  215. samtools_seed = f"{self.seed}.{fraction_decimal}"
  216. try:
  217. subprocess.run([
  218. 'samtools', 'view',
  219. '-s', samtools_seed,
  220. '-b', '-o', str(output_bam),
  221. str(self.input_bam)
  222. ], check=True, capture_output=True, text=True)
  223. logger.info(f"✓ Created: {output_bam.relative_to(self.output_dir)}")
  224. return output_bam
  225. except subprocess.CalledProcessError as e:
  226. logger.error(
  227. f"samtools failed: -s {samtools_seed}\n"
  228. f"stderr: {e.stderr}"
  229. )
  230. raise
  231. def _filter_bedgraph(self, input_bedgraph: Path, output_bedgraph: Path):
  232. """
  233. Filter bedGraph to keep only valid chromosomes.
  234. This is critical to prevent bedGraphToBigWig errors from non-standard
  235. chromosomes (e.g., chrM, random contigs, alt haplotypes).
  236. Args:
  237. input_bedgraph: Input bedGraph file (sorted)
  238. output_bedgraph: Output filtered bedGraph file
  239. """
  240. logger.info(f"Filtering chromosomes (keeping {len(self.valid_chroms)} standard chroms)...")
  241. filtered_lines = 0
  242. total_lines = 0
  243. with open(input_bedgraph, 'r') as infile, open(output_bedgraph, 'w') as outfile:
  244. for line in infile:
  245. total_lines += 1
  246. chrom = line.split('\t')[0]
  247. if chrom in self.valid_chroms:
  248. outfile.write(line)
  249. filtered_lines += 1
  250. kept_pct = (filtered_lines / total_lines * 100) if total_lines > 0 else 0
  251. logger.info(f" Kept {filtered_lines:,}/{total_lines:,} lines ({kept_pct:.1f}%)")
  252. def bam_to_bigwig(self, bam_file: Path, depth: Union[int, str]) -> Path:
  253. """
  254. Convert BAM to BigWig with chromosome filtering.
  255. Pipeline:
  256. 1. BAM -> sorted BAM (coordinate sorted)
  257. 2. sorted BAM -> bedGraph (coverage)
  258. 3. bedGraph -> sorted bedGraph (lexicographic sort)
  259. 4. sorted bedGraph -> filtered bedGraph (standard chroms only)
  260. 5. filtered bedGraph -> BigWig
  261. Args:
  262. bam_file: Input BAM file
  263. depth: Depth label (for directory structure)
  264. Returns:
  265. Path to output BigWig
  266. """
  267. depth_dir = self._get_depth_dir(depth)
  268. base = bam_file.stem
  269. sorted_bam = depth_dir / f"{base}.sorted.bam"
  270. bedgraph = depth_dir / f"{base}.bedgraph"
  271. sorted_bedgraph_temp = depth_dir / f"{base}.sorted.temp.bedgraph"
  272. sorted_bedgraph = depth_dir / f"{base}.sorted.bedgraph"
  273. bigwig = depth_dir / f"{base}.bw"
  274. try:
  275. # Step 1: Sort BAM by coordinates
  276. logger.info(f"[1/5] Sorting BAM...")
  277. subprocess.run([
  278. 'samtools', 'sort', '-o', str(sorted_bam), str(bam_file)
  279. ], check=True, capture_output=True)
  280. # Step 2: BAM to bedGraph (coverage)
  281. logger.info("[2/5] Converting to bedGraph...")
  282. with open(bedgraph, 'w') as f:
  283. subprocess.run([
  284. 'bedtools', 'genomecov', '-ibam', str(sorted_bam), '-bg'
  285. ], stdout=f, check=True)
  286. # Step 3: Sort bedGraph
  287. logger.info("[3/5] Sorting bedGraph...")
  288. with open(sorted_bedgraph_temp, 'w') as f:
  289. subprocess.run([
  290. 'sort', '-k1,1', '-k2,2n', str(bedgraph)
  291. ], stdout=f, check=True)
  292. # Step 4: Filter to standard chromosomes
  293. logger.info("[4/5] Filtering standard chromosomes...")
  294. self._filter_bedgraph(sorted_bedgraph_temp, sorted_bedgraph)
  295. # Step 5: bedGraph to BigWig
  296. logger.info("[5/5] Converting to BigWig...")
  297. subprocess.run([
  298. 'bedGraphToBigWig',
  299. str(sorted_bedgraph),
  300. str(self.chrom_sizes),
  301. str(bigwig)
  302. ], check=True, capture_output=True)
  303. logger.info(f"✓ Created: {bigwig.relative_to(self.output_dir)}")
  304. # Cleanup intermediate files
  305. if not self.keep_intermediate:
  306. for f in [sorted_bam, bedgraph, sorted_bedgraph_temp, sorted_bedgraph]:
  307. f.unlink(missing_ok=True)
  308. if bam_file != self.input_bam and bam_file.parent == depth_dir:
  309. bam_file.unlink(missing_ok=True)
  310. return bigwig
  311. except subprocess.CalledProcessError as e:
  312. logger.error(f"Conversion failed: {e}")
  313. if hasattr(e, 'stderr') and e.stderr:
  314. logger.error(f"stderr: {e.stderr}")
  315. raise
  316. def process(self, include_bulk: bool = True) -> List[Path]:
  317. """
  318. Process all target depths and optionally generate bulk BigWig.
  319. Directory structure created:
  320. output_dir/
  321. ├── CellType~ATAC~2.1e5/
  322. │ └── CellType~ATAC~2.1e5.bw
  323. ├── CellType~ATAC~4.2e5/
  324. │ └── CellType~ATAC~4.2e5.bw
  325. └── CellType~ATAC~bulk/
  326. └── CellType~ATAC~1.7e8.bw (actual depth in filename)
  327. Args:
  328. include_bulk: Also generate bulk BigWig from full BAM
  329. Returns:
  330. List of generated BigWig files
  331. """
  332. depths = self.get_target_depths()
  333. logger.info(f"\n{'='*70}")
  334. logger.info(f"Starting ATAC-seq preprocessing")
  335. logger.info(f"{'='*70}")
  336. logger.info(f"Cell type: {self.cell_type}")
  337. logger.info(f"Genome: {self.genome}")
  338. logger.info(f"Target depths: {len(depths)}")
  339. logger.info(f"Output: {self.output_dir}")
  340. logger.info(f"Structure: {self.cell_type}~ATAC~<depth>/")
  341. logger.info(f"{'='*70}\n")
  342. bigwig_files = []
  343. # Process subsampled depths
  344. for i, depth in enumerate(depths, 1):
  345. try:
  346. logger.info(f"\n[{i}/{len(depths)}] Processing depth: {depth:,}")
  347. logger.info(f"{'='*50}")
  348. bam = self.subsample_bam(depth)
  349. bw = self.bam_to_bigwig(bam, depth)
  350. bigwig_files.append(bw)
  351. except Exception as e:
  352. logger.error(f"Failed at depth {depth}: {e}")
  353. continue
  354. # Generate bulk BigWig
  355. if include_bulk:
  356. try:
  357. logger.info(f"\n[Bulk] Generating bulk BigWig...")
  358. logger.info(f"{'='*50}")
  359. # Get actual total reads
  360. bulk_depth = self._get_total_reads()
  361. logger.info(f"Bulk depth: {bulk_depth:,} reads")
  362. # Create bulk directory with 'bulk' label
  363. bulk_dir = self._get_depth_dir('bulk')
  364. # Generate filename with actual depth
  365. bulk_filename = self._generate_filename(bulk_depth, 'bam')
  366. bulk_bam = bulk_dir / bulk_filename
  367. # Copy original BAM to bulk directory
  368. logger.info("Copying input BAM to bulk directory...")
  369. subprocess.run(['cp', str(self.input_bam), str(bulk_bam)], check=True)
  370. logger.info(f"✓ Copied to: {bulk_bam.relative_to(self.output_dir)}")
  371. # Convert to BigWig (use actual depth for final filename)
  372. bulk_bw = self.bam_to_bigwig(bulk_bam, 'bulk')
  373. # Rename BigWig to include actual depth in filename
  374. # The bam_to_bigwig uses the BAM filename stem, so we need to rename
  375. final_bw_name = self._generate_filename(bulk_depth, 'bw')
  376. final_bw_path = bulk_dir / final_bw_name
  377. if bulk_bw != final_bw_path:
  378. bulk_bw.rename(final_bw_path)
  379. logger.info(f"✓ Renamed to: {final_bw_path.name}")
  380. bulk_bw = final_bw_path
  381. bigwig_files.append(bulk_bw)
  382. except Exception as e:
  383. logger.error(f"Failed to generate bulk: {e}")
  384. # Summary
  385. logger.info(f"\n{'='*70}")
  386. logger.info(f"✓ Processing Complete")
  387. logger.info(f"{'='*70}")
  388. logger.info(f"Generated {len(bigwig_files)} BigWig files:")
  389. for bw in bigwig_files:
  390. logger.info(f" • {bw.relative_to(self.output_dir)}")
  391. logger.info(f"\nOutput directory: {self.output_dir}")
  392. logger.info(f"{'='*70}\n")
  393. return bigwig_files
  394. def preprocess_atac(
  395. input_bam: str,
  396. cell_type: str,
  397. output_dir: str,
  398. chrom_sizes: str,
  399. genome: str = 'hg38',
  400. depths: Optional[List[Union[int, float]]] = None,
  401. min_depth: float = 2e5,
  402. max_depth: float = 2e7,
  403. step: float = 2e4,
  404. include_bulk: bool = True,
  405. **kwargs
  406. ) -> List[Path]:
  407. """
  408. Convenience function for ATAC-seq preprocessing.
  409. Example usage:
  410. # Using depth range
  411. bigwigs = preprocess_atac(
  412. input_bam='GM12878.bam',
  413. cell_type='GM12878',
  414. output_dir='data/ATAC/hg38',
  415. chrom_sizes='hg38.chrom.sizes',
  416. genome='hg38',
  417. min_depth=2e5,
  418. max_depth=2e7,
  419. step=2e4
  420. )
  421. # Using custom depth list
  422. bigwigs = preprocess_atac(
  423. input_bam='K562.bam',
  424. cell_type='K562',
  425. output_dir='data/ATAC/hg38',
  426. chrom_sizes='hg38.chrom.sizes',
  427. genome='hg38',
  428. depths=[1e5, 2.1e5, 5e5, 1e6]
  429. )
  430. Args:
  431. input_bam: Input BAM file path
  432. cell_type: Cell line name (e.g., 'GM12878', 'K562')
  433. output_dir: Output directory
  434. chrom_sizes: Chromosome sizes file
  435. genome: Genome build ('hg38', 'mm10', etc.)
  436. depths: Custom depth list (if provided, overrides range parameters)
  437. min_depth: Minimum depth for range mode (default: 2e5)
  438. max_depth: Maximum depth for range mode (default: 2e7)
  439. step: Step size for range mode (default: 2e4)
  440. include_bulk: Generate bulk BigWig (default: True)
  441. **kwargs: Additional arguments for ATACPreprocessor
  442. Returns:
  443. List of generated BigWig file paths
  444. """
  445. preprocessor = ATACPreprocessor(
  446. input_bam=input_bam,
  447. cell_type=cell_type,
  448. output_dir=output_dir,
  449. chrom_sizes=chrom_sizes,
  450. genome=genome,
  451. depth_mode='list' if depths else 'range',
  452. depth_list=depths,
  453. min_depth=min_depth,
  454. max_depth=max_depth,
  455. step=step,
  456. **kwargs
  457. )
  458. return preprocessor.process(include_bulk=include_bulk)

atac.py at commit 8ea1710, under MIT · at the source

Overview

Authors: Yuan-Chen Sun1, Wen-Jie Jiang2, Kang-Wen Cai1, Na-Na Wei3, Fu-Ting Lai4, Hao-Jie Wang1, Rui-Xiang Gao1, Ze-Yu Kuang4, Jia-Lu Zhou5, An Liu1, Han-Wen Zhu4, Yu-Juan Wang1, Ming Xu2,6, Hua-Jun Wu1,4,7
  1. Key laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Peking University Cancer Hospital & Institute, Beijing, China
  2. Department of Cardiology and Institute of Vascular Medicine, Peking University Third Hospital, State Key Laboratory of Vascular Homeostasis and Remodeling, NHC Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides, Beijing Key Laboratory of Cardiovascular Receptors Research, Peking University, Beijing, China
  3. Key laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Department of Lymphoma, Peking University Cancer Hospital & Institute, Beijing, China
  4. Department of Biomedical Informatics, School of Basic Medical Sciences, Peking University Health Science Center, Beijing, China
  5. Department of Gynecology and Obstetrics, Chinese PLA General Hospital, Beijing, China
  6. Research Unit of Medical Science Research Management/Basic and Clinical Research of Metabolic Cardiovascular Diseases, Chinese Academy of Medical Sciences, Beijing, China
  7. Beijing Advanced Center of Cellular Homeostasis and Aging-Related Diseases, Center for Precision Medicine Multi-Omics Research, Institute of Advanced Clinical Medicine, Peking University, Beijing, China
Journal: Nature communications, volume 17, issue 1, article 5172
Dates: received 23 July 2025; accepted 2 April 2026; published online 14 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-71877-z · PMID 41980945 · PMCID PMC13250166 · OpenAlex W7154258860
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), mouse (organism)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning
Keywords: Computational models, Machine learning, Epigenomics, Chromatin remodelling, Chromatin analysis
MeSH: Chromatin*, Deep Learning*, Genome*, Single-Cell Analysis*, Animals, Hippocampus, Humans, Mice (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Natural Science Foundation of China (National Science Foundation of China) (32470662, 32270683); Natural Science Foundation of Beijing Municipality (Beijing Natural Science Foundation) (5242006)
Citations: cited by 1 paper (Europe PMC); 71 references in the paper

Abstract

Three-dimensional genome organization controls cell-type-specific gene expression through chromatin interactions, yet systematic analysis across diverse cellular contexts remains limited by experimental constraints. Here we present Hi-Compass, a depth-aware deep learning framework that predicts cell-type-specific chromatin organization using only chromatin accessibility data as cell-type-specific input. By dynamically accommodating variability in sequencing depth, Hi-Compass enables robust predictions across the full spectrum of data scales, from sparse single-cell to high-coverage bulk profiles. Benchmarking shows that Hi-Compass achieves superior concordance with experimental Hi-C data compared to existing methods, with particularly strong recovery of high-confidence chromatin loops. Applied to peripheral blood and embryonic heart datasets, Hi-Compass resolves cell-type-specific chromatin interactions and systematically links disease-associated variants to putative target genes. The framework further enables spatially resolved chromatin interaction prediction in hippocampal tissue and demonstrates cross-species applicability through fine-tuning to mouse systems. Hi-Compass expands the capacity to study three-dimensional genome regulation across biological scales and species.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

EndeavourSyc/Hi-Compass

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 8ea17100dbc8576a66248f47587cb987e3b7f102, 22 July 2026
Languages: Python (27)
Size: 33 files, 27 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (setup.py), documentation
Not found: CITATION.cff, tests, continuous integration
Tools: NumPy (13 files), PyTorch (10 files), pandas (8 files), scikit-image (5 files), Matplotlib (2 files), BEDTools (1 file), PyTorch Lightning (1 file), pysam (1 file), SAMtools (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
29 files

Zenodo 19283061

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (13 files), PyTorch (10 files), pandas (8 files), scikit-image (5 files), Matplotlib (2 files), BEDTools (1 file), PyTorch Lightning (1 file), pysam (1 file), SAMtools (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
  • 29 September 2026: the link answers (HTTP 200)
29 files
At the source:

Code availability

The Hi-Compass framework was implemented in the ‘hicompass’ Python package, which is available at https://github.com/EndeavourSyc/Hi-Compass. The specific version of the code associated with this publication is archived in Zenodo (10.5281/zenodo.19283061).

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 54 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Data links

Data availability

The Hi-C, CTCF ChIP–seq and ATAC–seq datasets used in the study were all public data from the ENCODE (https://www.encodeproject.org/), 4DN (https://data.4dnucleome.org/) and GEO (https://www.ncbi.nlm.nih.gov/geo/) database, with the detailed information listed in the Supplementary Data 1-3. Cohesin HiChIP loop calls in GM12878 were obtained from Mumbach et al.43 and RAD21 ChIA-PET loops from Grubert et al.44 The single-cell ATAC-seq, Multiome and spatial ATAC-seq data were obtained from 10X Genomics (https://www.10xgenomics.com/datasets/), ENCODE and GEO database, with the detailed information listed in the Supplementary Data 4. GWAS SNPs data was obtained from GWAS Catalog70 (https://www.ebi.ac.uk/gwas/), eQTLs data from GTEx71 v10 (https://www.gtexportal.org/).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 14 authors, 5 keywords, 8 MeSH terms, 2 funders, 68 references.

Cite

This paper

Sun, Y.-C., Jiang, W.-J., Cai, K.-W., Wei, N.-N., Lai, F.-T., Wang, H.-J., Gao, R.-X., Kuang, Z.-Y., Zhou, J.-L., Liu, A., Zhu, H.-W., Wang, Y.-J., Xu, M., & Wu, H.-J. (2026). Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution. Nature communications, 17(1), 5172. https://doi.org/10.1038/s41467-026-71877-z

BibTeX

@article{sun2026hi,
author = {Sun, Yuan-Chen and Jiang, Wen-Jie and Cai, Kang-Wen and Wei, Na-Na and Lai, Fu-Ting and Wang, Hao-Jie and Gao, Rui-Xiang and Kuang, Ze-Yu and Zhou, Jia-Lu and Liu, An and Zhu, Han-Wen and Wang, Yu-Juan and Xu, Ming and Wu, Hua-Jun},
title = {{Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution}},
journal = {Nature communications},
year = {2026},
month = apr,
volume = {17},
number = {1},
pages = {5172},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-71877-z},
url = {https://doi.org/10.1038/s41467-026-71877-z},
pmid = {41980945},
pmcid = {PMC13250166}
}

RIS

TY - JOUR
AU - Sun, Yuan-Chen
AU - Jiang, Wen-Jie
AU - Cai, Kang-Wen
AU - Wei, Na-Na
AU - Lai, Fu-Ting
AU - Wang, Hao-Jie
AU - Gao, Rui-Xiang
AU - Kuang, Ze-Yu
AU - Zhou, Jia-Lu
AU - Liu, An
AU - Zhu, Han-Wen
AU - Wang, Yu-Juan
AU - Xu, Ming
AU - Wu, Hua-Jun
TI - Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/04/14
VL - 17
IS - 1
SP - 5172
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-71877-z
UR - https://doi.org/10.1038/s41467-026-71877-z
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-71877-z",
"type": "article-journal",
"title": "Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution",
"container-title": "Nature communications",
"author": [
{
"family": "Sun",
"given": "Yuan-Chen"
},
{
"family": "Jiang",
"given": "Wen-Jie"
},
{
"family": "Cai",
"given": "Kang-Wen"
},
{
"family": "Wei",
"given": "Na-Na"
},
{
"family": "Lai",
"given": "Fu-Ting"
},
{
"family": "Wang",
"given": "Hao-Jie"
},
{
"family": "Gao",
"given": "Rui-Xiang"
},
{
"family": "Kuang",
"given": "Ze-Yu"
},
{
"family": "Zhou",
"given": "Jia-Lu"
},
{
"family": "Liu",
"given": "An"
},
{
"family": "Zhu",
"given": "Han-Wen"
},
{
"family": "Wang",
"given": "Yu-Juan"
},
{
"family": "Xu",
"given": "Ming"
},
{
"family": "Wu",
"given": "Hua-Jun"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "5172",
"DOI": "10.1038/s41467-026-71877-z",
"PMID": "41980945",
"PMCID": "PMC13250166",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-71877-z",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pgen.1012081
ADNP regulates chromatin architecture and lineage fidelity during neural differentiation.
Journal: PLoS genetics
In common: mouse, 9 references
[2] doi:10.21203/rs.3.rs-9927928/v1 [code]
Genome-wide and allele-resolved maps of the radial architecture of the mouse genome
Journal: Research Square (preprint)
In common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 4 references
[3] doi:10.1101/gr.280394.124 [code]
De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.
Journal: Genome research
In common: pysam, BEDTools, scikit-image, 3 other tools, 4 references
[4] doi:10.1038/s42003-026-10462-y [code]
SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.
Journal: Communications biology
In common: pysam, BEDTools, PyTorch, 3 other tools, genetics / omics, mouse, 3 references
[5] doi:10.1093/bioinformatics/btag652 [code]
mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.
Journal: Bioinformatics (Oxford, England)
In common: pysam, PyTorch Lightning, BEDTools, 4 other tools, genetics / omics, mouse, 1 reference
[6] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: pysam, BEDTools, PyTorch, 3 other tools, genetics / omics, mouse, 2 references
[7] doi:10.1016/j.celrep.2026.117110 [code]
Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.
Journal: Cell reports
In common: pysam, BEDTools, SAMtools, 4 other tools, genetics / omics, 1 reference
[8] doi:10.1038/s41586-026-10512-9 [code]
Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.
Journal: Nature
In common: pysam, BEDTools, SAMtools, 3 other tools, mouse, 2 references
[9] doi:10.1038/s41592-026-03211-w [code]
Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.
Journal: Nature methods
In common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 1 reference
[10] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.