Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution.
The 15 matches
- [1] § Methods › Bulk ATAC-seq data processing ↔ hicompass/preprocess/atac.py, lines 299–375 · score 0.73 · bedGraph, BigWig, genomecov, bedtools, pipeline, BAM
- [2] § Methods › Model architecture ↔ hicompass/predicting/blocks.py, lines 95–129 · score 0.72 · DNAEncoder, ReLU, feature weighting, channel, batch, modules
- [3] § Methods › Meta-cell strategy and meta-cell Hi-C prediction ↔ hicompass/preprocess/atac.py, lines 299–375 · score 0.72 · bedGraph, BigWig, genomecov, bedtools, bw, bam
- [4] § Methods › Model architecture ↔ hicompass/predicting/blocks.py, lines 34–92 · score 0.68 · ATACDepthEncoder, depth feature, depth range, embedding, module
- [5] § Methods › Model training strategy ↔ hicompass/train/HicompassTrain.py, lines 206–283 · score 0.67 · cross entropy losses, MSE loss, SSIM, batch, training, model
- [6] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/predicting/PredictDataset.py, lines 64–205 · score 0.64 · DNA sequences, sequencing depth, genomic features, generalized CTCF, ATAC seq, error
- [7] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/predicting/PredictDataset.py, lines 64–205 · score 0.62 · DNA sequence, sequencing depth, genomic features, generalized CTCF, ATAC seq, genome
- [8] § Methods › Performance comparison with previous methods ↔ hicompass/train/HicompassModel.py, lines 47–114 · score 0.61 · AdaptiveAvgPool2d, insulation score, PyTorch, diagonal, models, training
- [9] § Methods › Model architecture ↔ hicompass/train/HicompassModel.py, lines 47–114 · score 0.60 · ResNet, adaptive, PyTorch, decoder, discriminator, modeling
- [10] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/train/HicompassTrain.py, lines 206–283 · score 0.60 · cross entropy, information weighted, discriminator, MSE, SSIM, loss
- [11] § Methods › Hi-C data processing ↔ hicompass/preprocess/hic_norm.py, lines 147–227 · score 0.58 · kb resolution, contrast stretching, mouse, human, hg38, optimize
- [12] § Results › Systematic evaluation of Hi-Compass performance ↔ hicompass/preprocess/hic_norm.py, lines 450–557 · score 0.56 · upper triangle, lower triangle, contrast stretched, bands, diagonals, Window
- [13] § Methods › Hi-C prediction and integration ↔ hicompass/predicting/utils.py, lines 113–178 · score 0.53 · create_cooler, prediction window, matrix, resolution, chromosomes, Hi
- [14] § Results › Hi-Compass enables Hi-C prediction from multi-depth ATAC-seq data ↔ hicompass/train/HicompassDataset.py, lines 535–581 · score 0.53 · DNA sequence, genomic features, generalized CTCF, ATAC seq, genome, depth
- [15] § Methods › CTCF ChIP-seq data processing ↔ hicompass/preprocess/atac.py, lines 272–297 · score 0.50 · bedGraph
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 545 lines · 20 KB · MIT · 3 matches
- #!/usr/bin/env python3
- """
- ATAC-seq preprocessing for Hi-Compass training data preparation.
- This module handles:
- 1. Stratified subsampling from bulk ATAC-seq BAM files
- 2. Conversion to BigWig format with proper chromosome filtering
- 3. Hierarchical directory organization for training data
- """
- import os
- import subprocess
- from pathlib import Path
- from typing import List, Optional, Union
- import logging
- logging.basicConfig(
- level=logging.INFO,
- format='%(asctime)s - %(levelname)s - %(message)s'
- )
- logger = logging.getLogger(__name__)
- class ATACPreprocessor:
- """Preprocesses ATAC-seq data for Hi-Compass training."""
- def __init__(
- self,
- input_bam: Union[str, Path],
- cell_type: str,
- output_dir: Union[str, Path],
- chrom_sizes: Union[str, Path],
- genome: str = 'hg38',
- depth_mode: str = 'range',
- depth_list: Optional[List[Union[int, float]]] = None,
- min_depth: float = 2e5,
- max_depth: float = 2e7,
- step: float = 2e4,
- seed: int = 114514,
- keep_intermediate: bool = False
- ):
- """
- Initialize ATAC preprocessor.
- Args:
- input_bam: Path to input BAM/SAM file (bulk ATAC-seq)
- cell_type: Cell line name (e.g., 'GM12878', 'K562')
- output_dir: Output directory for processed files
- chrom_sizes: Chromosome sizes file (e.g., hg38.chrom.sizes)
- genome: Genome build ('hg38', 'mm10', etc.) for chromosome filtering
- depth_mode: 'range' or 'list' for depth specification
- depth_list: Custom depths (for 'list' mode)
- min_depth: Minimum depth (for 'range' mode, default: 2e5)
- max_depth: Maximum depth (for 'range' mode, default: 2e7)
- step: Step size (for 'range' mode, default: 2e4)
- seed: Random seed for reproducibility
- keep_intermediate: Keep intermediate files (BAM, bedGraph)
- """
- self.input_bam = Path(input_bam)
- self.cell_type = self._sanitize_cell_type(cell_type)
- self.output_dir = Path(output_dir)
- self.chrom_sizes = Path(chrom_sizes)
- self.genome = genome
- self.depth_mode = depth_mode
- self.depth_list = [int(d) for d in depth_list] if depth_list else None
- self.min_depth = int(min_depth)
- self.max_depth = int(max_depth)
- self.step = int(step)
- self.seed = seed
- self.keep_intermediate = keep_intermediate
- # Define valid chromosomes based on genome
- self._set_valid_chromosomes()
- self.output_dir.mkdir(parents=True, exist_ok=True)
- self._validate_inputs()
- self._total_reads = None
- def _set_valid_chromosomes(self):
- """Set valid chromosome names based on genome build."""
- if self.genome == 'hg38':
- # Human: chr1-22, chrX, chrY
- self.valid_chroms = set([f'chr{i}' for i in range(1, 23)] + ['chrX', 'chrY'])
- elif self.genome == 'mm10':
- # Mouse: chr1-19, chrX, chrY
- self.valid_chroms = set([f'chr{i}' for i in range(1, 20)] + ['chrX', 'chrY'])
- else:
- # For custom genomes, read from chrom_sizes file
- logger.warning(f"Custom genome '{self.genome}' - will use all chromosomes from chrom_sizes")
- self.valid_chroms = None # Will be set from chrom_sizes file
- @staticmethod
- def _sanitize_cell_type(cell_type: str) -> str:
- """Sanitize cell type to valid filename characters."""
- import re
- return re.sub(r'[^a-zA-Z0-9_-]', '_', cell_type)
- def _validate_inputs(self):
- """Validate inputs and check required tools."""
- if not self.input_bam.exists():
- raise FileNotFoundError(f"Input BAM not found: {self.input_bam}")
- if not self.chrom_sizes.exists():
- raise FileNotFoundError(f"Chrom sizes not found: {self.chrom_sizes}")
- # Load valid chromosomes from chrom_sizes if not set
- if self.valid_chroms is None:
- self.valid_chroms = set()
- with open(self.chrom_sizes, 'r') as f:
- for line in f:
- chrom = line.strip().split('\t')[0]
- self.valid_chroms.add(chrom)
- logger.info(f"Loaded {len(self.valid_chroms)} chromosomes from {self.chrom_sizes.name}")
- # Validate depth configuration
- if self.depth_mode == 'list':
- if not self.depth_list:
- raise ValueError("depth_list required for mode 'list'")
- if any(d <= 0 for d in self.depth_list):
- raise ValueError("All depths must be positive")
- elif self.depth_mode == 'range':
- if self.min_depth >= self.max_depth:
- raise ValueError("min_depth must be < max_depth")
- else:
- raise ValueError(f"Invalid depth_mode: {self.depth_mode}")
- # Check required tools
- required_tools = ['samtools', 'bedtools', 'bedGraphToBigWig']
- missing_tools = []
- for tool in required_tools:
- if subprocess.run(['which', tool], capture_output=True).returncode != 0:
- missing_tools.append(tool)
- if missing_tools:
- raise EnvironmentError(
- f"Required tools not found: {', '.join(missing_tools)}\n"
- f"Please install: samtools, bedtools, bedGraphToBigWig (UCSC tools)"
- )
- def _get_total_reads(self) -> int:
- """Count total reads in BAM file."""
- if self._total_reads is None:
- logger.info("Counting reads in input BAM...")
- result = subprocess.run(
- ['samtools', 'view', '-c', str(self.input_bam)],
- capture_output=True, text=True, check=True
- )
- self._total_reads = int(result.stdout.strip())
- logger.info(f"Total reads in {self.input_bam.name}: {self._total_reads:,}")
- return self._total_reads
- @staticmethod
- def _format_depth(depth: int) -> str:
- """
- Format depth as scientific notation (e.g., 1e5, 2e6, 2.1e5).
- Examples:
- 100000 -> '1e5'
- 210000 -> '2.1e5'
- 2000000 -> '2e6'
- """
- exp = 0
- value = float(depth)
- while value >= 10:
- value /= 10
- exp += 1
- if value == int(value):
- return f"{int(value)}e{exp}"
- return f"{value:.1f}e{exp}".rstrip('0').rstrip('.')
- def _generate_filename(self, depth: Union[int, str], extension: str) -> str:
- """
- Generate standardized filename: {CellType}~ATAC~{Depth}.{ext}
- Examples:
- GM12878, 210000, 'bw' -> 'GM12878~ATAC~2.1e5.bw'
- K562, 'bulk', 'bam' -> 'K562~ATAC~bulk.bam'
- """
- if isinstance(depth, int):
- depth_str = self._format_depth(depth)
- else:
- depth_str = depth
- return f"{self.cell_type}~ATAC~{depth_str}.{extension}"
- def _get_depth_dir(self, depth: Union[int, str]) -> Path:
- """
- Get subdirectory for a specific depth.
- Creates hierarchical structure:
- output_dir/
- └── CellType~ATAC~Depth/
- ├── CellType~ATAC~Depth.bam
- ├── CellType~ATAC~Depth.sorted.bam
- ├── CellType~ATAC~Depth.bedgraph
- ├── CellType~ATAC~Depth.sorted.bedgraph
- └── CellType~ATAC~Depth.bw
- """
- if isinstance(depth, int):
- depth_str = self._format_depth(depth)
- else:
- depth_str = depth
- depth_dir = self.output_dir / f"{self.cell_type}~ATAC~{depth_str}"
- depth_dir.mkdir(parents=True, exist_ok=True)
- return depth_dir
- def get_target_depths(self) -> List[int]:
- """Get list of target depths based on configuration."""
- if self.depth_mode == 'list':
- depths = sorted(self.depth_list)
- else:
- depths = list(range(self.min_depth, self.max_depth + 1, self.step))
- logger.info(f"Target depths: {len(depths)} levels")
- return depths
- def subsample_bam(self, target_depth: int) -> Path:
- """
- Subsample BAM to target depth using samtools.
- Args:
- target_depth: Target number of reads
- Returns:
- Path to subsampled BAM in depth-specific subdirectory
- """
- total_reads = self._get_total_reads()
- if target_depth >= total_reads:
- logger.warning(
- f"Target {target_depth:,} >= total {total_reads:,}, using full BAM"
- )
- return self.input_bam
- fraction = target_depth / total_reads
- # Get depth-specific directory
- depth_dir = self._get_depth_dir(target_depth)
- output_bam = depth_dir / self._generate_filename(target_depth, 'bam')
- logger.info(f"Subsampling to {target_depth:,} reads (fraction={fraction:.6f})")
- # Format for samtools -s: seed.fraction
- fraction_str = f"{fraction:.10f}"
- if fraction_str.startswith('0.'):
- fraction_decimal = fraction_str[2:] # Remove "0."
- else:
- fraction_decimal = fraction_str
- fraction_decimal = fraction_decimal.rstrip('0') or '0'
- samtools_seed = f"{self.seed}.{fraction_decimal}"
- try:
- subprocess.run([
- 'samtools', 'view',
- '-s', samtools_seed,
- '-b', '-o', str(output_bam),
- str(self.input_bam)
- ], check=True, capture_output=True, text=True)
- logger.info(f"✓ Created: {output_bam.relative_to(self.output_dir)}")
- return output_bam
- except subprocess.CalledProcessError as e:
- logger.error(
- f"samtools failed: -s {samtools_seed}\n"
- f"stderr: {e.stderr}"
- )
- raise
- def _filter_bedgraph(self, input_bedgraph: Path, output_bedgraph: Path):
- """
- Filter bedGraph to keep only valid chromosomes.
- This is critical to prevent bedGraphToBigWig errors from non-standard
- chromosomes (e.g., chrM, random contigs, alt haplotypes).
- Args:
- input_bedgraph: Input bedGraph file (sorted)
- output_bedgraph: Output filtered bedGraph file
- """
- logger.info(f"Filtering chromosomes (keeping {len(self.valid_chroms)} standard chroms)...")
- filtered_lines = 0
- total_lines = 0
- with open(input_bedgraph, 'r') as infile, open(output_bedgraph, 'w') as outfile:
- for line in infile:
- total_lines += 1
- chrom = line.split('\t')[0]
- if chrom in self.valid_chroms:
- outfile.write(line)
- filtered_lines += 1
- kept_pct = (filtered_lines / total_lines * 100) if total_lines > 0 else 0
- logger.info(f" Kept {filtered_lines:,}/{total_lines:,} lines ({kept_pct:.1f}%)")
- def bam_to_bigwig(self, bam_file: Path, depth: Union[int, str]) -> Path:
- """
- Convert BAM to BigWig with chromosome filtering.
- Pipeline:
- 1. BAM -> sorted BAM (coordinate sorted)
- 2. sorted BAM -> bedGraph (coverage)
- 3. bedGraph -> sorted bedGraph (lexicographic sort)
- 4. sorted bedGraph -> filtered bedGraph (standard chroms only)
- 5. filtered bedGraph -> BigWig
- Args:
- bam_file: Input BAM file
- depth: Depth label (for directory structure)
- Returns:
- Path to output BigWig
- """
- depth_dir = self._get_depth_dir(depth)
- base = bam_file.stem
- sorted_bam = depth_dir / f"{base}.sorted.bam"
- bedgraph = depth_dir / f"{base}.bedgraph"
- sorted_bedgraph_temp = depth_dir / f"{base}.sorted.temp.bedgraph"
- sorted_bedgraph = depth_dir / f"{base}.sorted.bedgraph"
- bigwig = depth_dir / f"{base}.bw"
- try:
- # Step 1: Sort BAM by coordinates
- logger.info(f"[1/5] Sorting BAM...")
- subprocess.run([
- 'samtools', 'sort', '-o', str(sorted_bam), str(bam_file)
- ], check=True, capture_output=True)
- # Step 2: BAM to bedGraph (coverage)
- logger.info("[2/5] Converting to bedGraph...")
- with open(bedgraph, 'w') as f:
- subprocess.run([
- 'bedtools', 'genomecov', '-ibam', str(sorted_bam), '-bg'
- ], stdout=f, check=True)
- # Step 3: Sort bedGraph
- logger.info("[3/5] Sorting bedGraph...")
- with open(sorted_bedgraph_temp, 'w') as f:
- subprocess.run([
- 'sort', '-k1,1', '-k2,2n', str(bedgraph)
- ], stdout=f, check=True)
- # Step 4: Filter to standard chromosomes
- logger.info("[4/5] Filtering standard chromosomes...")
- self._filter_bedgraph(sorted_bedgraph_temp, sorted_bedgraph)
- # Step 5: bedGraph to BigWig
- logger.info("[5/5] Converting to BigWig...")
- subprocess.run([
- 'bedGraphToBigWig',
- str(sorted_bedgraph),
- str(self.chrom_sizes),
- str(bigwig)
- ], check=True, capture_output=True)
- logger.info(f"✓ Created: {bigwig.relative_to(self.output_dir)}")
- # Cleanup intermediate files
- if not self.keep_intermediate:
- for f in [sorted_bam, bedgraph, sorted_bedgraph_temp, sorted_bedgraph]:
- f.unlink(missing_ok=True)
- if bam_file != self.input_bam and bam_file.parent == depth_dir:
- bam_file.unlink(missing_ok=True)
- return bigwig
- except subprocess.CalledProcessError as e:
- logger.error(f"Conversion failed: {e}")
- if hasattr(e, 'stderr') and e.stderr:
- logger.error(f"stderr: {e.stderr}")
- raise
- def process(self, include_bulk: bool = True) -> List[Path]:
- """
- Process all target depths and optionally generate bulk BigWig.
- Directory structure created:
- output_dir/
- ├── CellType~ATAC~2.1e5/
- │ └── CellType~ATAC~2.1e5.bw
- ├── CellType~ATAC~4.2e5/
- │ └── CellType~ATAC~4.2e5.bw
- └── CellType~ATAC~bulk/
- └── CellType~ATAC~1.7e8.bw (actual depth in filename)
- Args:
- include_bulk: Also generate bulk BigWig from full BAM
- Returns:
- List of generated BigWig files
- """
- depths = self.get_target_depths()
- logger.info(f"\n{'='*70}")
- logger.info(f"Starting ATAC-seq preprocessing")
- logger.info(f"{'='*70}")
- logger.info(f"Cell type: {self.cell_type}")
- logger.info(f"Genome: {self.genome}")
- logger.info(f"Target depths: {len(depths)}")
- logger.info(f"Output: {self.output_dir}")
- logger.info(f"Structure: {self.cell_type}~ATAC~<depth>/")
- logger.info(f"{'='*70}\n")
- bigwig_files = []
- # Process subsampled depths
- for i, depth in enumerate(depths, 1):
- try:
- logger.info(f"\n[{i}/{len(depths)}] Processing depth: {depth:,}")
- logger.info(f"{'='*50}")
- bam = self.subsample_bam(depth)
- bw = self.bam_to_bigwig(bam, depth)
- bigwig_files.append(bw)
- except Exception as e:
- logger.error(f"Failed at depth {depth}: {e}")
- continue
- # Generate bulk BigWig
- if include_bulk:
- try:
- logger.info(f"\n[Bulk] Generating bulk BigWig...")
- logger.info(f"{'='*50}")
- # Get actual total reads
- bulk_depth = self._get_total_reads()
- logger.info(f"Bulk depth: {bulk_depth:,} reads")
- # Create bulk directory with 'bulk' label
- bulk_dir = self._get_depth_dir('bulk')
- # Generate filename with actual depth
- bulk_filename = self._generate_filename(bulk_depth, 'bam')
- bulk_bam = bulk_dir / bulk_filename
- # Copy original BAM to bulk directory
- logger.info("Copying input BAM to bulk directory...")
- subprocess.run(['cp', str(self.input_bam), str(bulk_bam)], check=True)
- logger.info(f"✓ Copied to: {bulk_bam.relative_to(self.output_dir)}")
- # Convert to BigWig (use actual depth for final filename)
- bulk_bw = self.bam_to_bigwig(bulk_bam, 'bulk')
- # Rename BigWig to include actual depth in filename
- # The bam_to_bigwig uses the BAM filename stem, so we need to rename
- final_bw_name = self._generate_filename(bulk_depth, 'bw')
- final_bw_path = bulk_dir / final_bw_name
- if bulk_bw != final_bw_path:
- bulk_bw.rename(final_bw_path)
- logger.info(f"✓ Renamed to: {final_bw_path.name}")
- bulk_bw = final_bw_path
- bigwig_files.append(bulk_bw)
- except Exception as e:
- logger.error(f"Failed to generate bulk: {e}")
- # Summary
- logger.info(f"\n{'='*70}")
- logger.info(f"✓ Processing Complete")
- logger.info(f"{'='*70}")
- logger.info(f"Generated {len(bigwig_files)} BigWig files:")
- for bw in bigwig_files:
- logger.info(f" • {bw.relative_to(self.output_dir)}")
- logger.info(f"\nOutput directory: {self.output_dir}")
- logger.info(f"{'='*70}\n")
- return bigwig_files
- def preprocess_atac(
- input_bam: str,
- cell_type: str,
- output_dir: str,
- chrom_sizes: str,
- genome: str = 'hg38',
- depths: Optional[List[Union[int, float]]] = None,
- min_depth: float = 2e5,
- max_depth: float = 2e7,
- step: float = 2e4,
- include_bulk: bool = True,
- **kwargs
- ) -> List[Path]:
- """
- Convenience function for ATAC-seq preprocessing.
- Example usage:
- # Using depth range
- bigwigs = preprocess_atac(
- input_bam='GM12878.bam',
- cell_type='GM12878',
- output_dir='data/ATAC/hg38',
- chrom_sizes='hg38.chrom.sizes',
- genome='hg38',
- min_depth=2e5,
- max_depth=2e7,
- step=2e4
- )
- # Using custom depth list
- bigwigs = preprocess_atac(
- input_bam='K562.bam',
- cell_type='K562',
- output_dir='data/ATAC/hg38',
- chrom_sizes='hg38.chrom.sizes',
- genome='hg38',
- depths=[1e5, 2.1e5, 5e5, 1e6]
- )
- Args:
- input_bam: Input BAM file path
- cell_type: Cell line name (e.g., 'GM12878', 'K562')
- output_dir: Output directory
- chrom_sizes: Chromosome sizes file
- genome: Genome build ('hg38', 'mm10', etc.)
- depths: Custom depth list (if provided, overrides range parameters)
- min_depth: Minimum depth for range mode (default: 2e5)
- max_depth: Maximum depth for range mode (default: 2e7)
- step: Step size for range mode (default: 2e4)
- include_bulk: Generate bulk BigWig (default: True)
- **kwargs: Additional arguments for ATACPreprocessor
- Returns:
- List of generated BigWig file paths
- """
- preprocessor = ATACPreprocessor(
- input_bam=input_bam,
- cell_type=cell_type,
- output_dir=output_dir,
- chrom_sizes=chrom_sizes,
- genome=genome,
- depth_mode='list' if depths else 'range',
- depth_list=depths,
- min_depth=min_depth,
- max_depth=max_depth,
- step=step,
- **kwargs
- )
- return preprocessor.process(include_bulk=include_bulk)
atac.py at commit 8ea1710, under MIT · at the source
Overview
- Key laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Peking University Cancer Hospital & Institute, Beijing, China
- Department of Cardiology and Institute of Vascular Medicine, Peking University Third Hospital, State Key Laboratory of Vascular Homeostasis and Remodeling, NHC Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides, Beijing Key Laboratory of Cardiovascular Receptors Research, Peking University, Beijing, China
- Key laboratory of Carcinogenesis and Translational Research (Ministry of Education/Beijing), Department of Lymphoma, Peking University Cancer Hospital & Institute, Beijing, China
- Department of Biomedical Informatics, School of Basic Medical Sciences, Peking University Health Science Center, Beijing, China
- Department of Gynecology and Obstetrics, Chinese PLA General Hospital, Beijing, China
- Research Unit of Medical Science Research Management/Basic and Clinical Research of Metabolic Cardiovascular Diseases, Chinese Academy of Medical Sciences, Beijing, China
- Beijing Advanced Center of Cellular Homeostasis and Aging-Related Diseases, Center for Precision Medicine Multi-Omics Research, Institute of Advanced Clinical Medicine, Peking University, Beijing, China
Abstract
Three-dimensional genome organization controls cell-type-specific gene expression through chromatin interactions, yet systematic analysis across diverse cellular contexts remains limited by experimental constraints. Here we present Hi-Compass, a depth-aware deep learning framework that predicts cell-type-specific chromatin organization using only chromatin accessibility data as cell-type-specific input. By dynamically accommodating variability in sequencing depth, Hi-Compass enables robust predictions across the full spectrum of data scales, from sparse single-cell to high-coverage bulk profiles. Benchmarking shows that Hi-Compass achieves superior concordance with experimental Hi-C data compared to existing methods, with particularly strong recovery of high-confidence chromatin loops. Applied to peripheral blood and embryonic heart datasets, Hi-Compass resolves cell-type-specific chromatin interactions and systematically links disease-associated variants to putative target genes. The framework further enables spatially resolved chromatin interaction prediction in hippocampal tissue and demonstrates cross-species applicability through fine-tuning to mouse systems. Hi-Compass expands the capacity to study three-dimensional genome regulation across biological scales and species.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.
EndeavourSyc/Hi-Compass
8ea17100dbc8576a66248f47587cb987e3b7f102, 22 July 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
29 files
- hicompass/
__init__.py , Python, 3 lines - hicompass/
cli.py , Python, 74 lines - hicompass/
commands/ , Python, 14 lines__init__.py - hicompass/
commands/ , Python, 283 linespredicting.py - hicompass/
commands/ , Python, 140 linespreprocess_atac.py - hicompass/
commands/ , Python, 1 linepreprocess_dna.py - hicompass/
commands/ , Python, 145 linespreprocess_hic_norm.py - hicompass/
commands/ , Python, 80 linespreprocess_hic_to_npz.py - hicompass/
commands/ , Python, 453 linestraining.py - hicompass/
predicting/ , Python, 205 lines, 2 matchesPredictDataset.py - hicompass/
predicting/ , Python, 51 linesPredictModel.py - hicompass/
predicting/ , Python, 14 lines__init__.py - hicompass/
predicting/ , Python, 457 lines, 2 matchesblocks.py - hicompass/
predicting/ , Python, 161 lineschromosome_dataset_predi ct.py - hicompass/
predicting/ , Python, 203 linesdata_feature.py - hicompass/
predicting/ , Python, 205 lines, 1 matchutils.py - hicompass/
preprocess/ , Python, 15 lines__init__.py - hicompass/
preprocess/ , Python, 545 lines, 3 matchesatac.py - hicompass/
preprocess/ , Python, 636 lines, 1 matchhic_norm.py - hicompass/
preprocess/ , Python, 304 lineshic_to_npz.py - hicompass/
train/ , Python, 916 lines, 1 matchHicompassDataset.py - hicompass/
train/ , Python, 114 lines, 2 matchesHicompassModel.py - hicompass/
train/ , Python, 547 lines, 2 matchesHicompassTrain.py - hicompass/
train/ , Python, 17 lines__init__.py - hicompass/
train/ , Python, 457 linesblocks.py - hicompass/
train/ , Python, 203 linesdata_feature.py - setup.py, Python, 59 lines
- LICENSE, License, 21 lines
- README.md, Text, 508 lines
Zenodo 19283061
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
- 29 September 2026: the link answers (HTTP 200)
29 files
- hicompass/
__init__.py , Python, 3 lines - hicompass/
cli.py , Python, 74 lines - hicompass/
commands/ , Python, 14 lines__init__.py - hicompass/
commands/ , Python, 283 linespredicting.py - hicompass/
commands/ , Python, 140 linespreprocess_atac.py - hicompass/
commands/ , Python, 1 linepreprocess_dna.py - hicompass/
commands/ , Python, 123 linespreprocess_hic_norm.py - hicompass/
commands/ , Python, 80 linespreprocess_hic_to_npz.py - hicompass/
commands/ , Python, 453 linestraining.py - hicompass/
predicting/ , Python, 205 linesPredictDataset.py - hicompass/
predicting/ , Python, 51 linesPredictModel.py - hicompass/
predicting/ , Python, 14 lines__init__.py - hicompass/
predicting/ , Python, 457 linesblocks.py - hicompass/
predicting/ , Python, 161 lineschromosome_dataset_predi ct.py - hicompass/
predicting/ , Python, 203 linesdata_feature.py - hicompass/
predicting/ , Python, 205 linesutils.py - hicompass/
preprocess/ , Python, 15 lines__init__.py - hicompass/
preprocess/ , Python, 545 linesatac.py - hicompass/
preprocess/ , Python, 519 lines, 1 matchhic_norm.py - hicompass/
preprocess/ , Python, 304 lineshic_to_npz.py - hicompass/
train/ , Python, 916 linesHicompassDataset.py - hicompass/
train/ , Python, 114 linesHicompassModel.py - hicompass/
train/ , Python, 547 linesHicompassTrain.py - hicompass/
train/ , Python, 17 lines__init__.py - hicompass/
train/ , Python, 457 linesblocks.py - hicompass/
train/ , Python, 203 linesdata_feature.py - setup.py, Python, 59 lines
- LICENSE, License, 21 lines
- README.md, Text, 508 lines
Code availability
The Hi-Compass framework was implemented in the ‘hicompass’ Python package, which is available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 54 scripts, each with its path and the digest of its content;
- 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Data links
- ncbi.nlm.nih.gov/
geo , NCBI; found in “Data availability”
Data availability
The Hi-C, CTCF ChIP–seq and ATAC–seq datasets used in the study were all public data from the ENCODE (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 14 authors, 5 keywords, 8 MeSH terms, 2 funders, 68 references.
Cite
This paper
Sun, Y.-C., Jiang, W.-J., Cai, K.-W., Wei, N.-N., Lai, F.-T., Wang, H.-J., Gao, R.-X., Kuang, Z.-Y., Zhou, J.-L., Liu, A., Zhu, H.-W., Wang, Y.-J., Xu, M., & Wu, H.-J. (2026). Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution. Nature communications, 17(1), 5172. https://
BibTeX
@article{sun2026hi,
author = {Sun, Yuan-Chen and Jiang, Wen-Jie and Cai, Kang-Wen and Wei, Na-Na and Lai, Fu-Ting and Wang, Hao-Jie and Gao, Rui-Xiang and Kuang, Ze-Yu and Zhou, Jia-Lu and Liu, An and Zhu, Han-Wen and Wang, Yu-Juan and Xu, Ming and Wu, Hua-Jun},
title = {{Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution}},
journal = {Nature communications},
year = {2026},
month = apr,
volume = {17},
number = {1},
pages = {5172},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {41980945},
pmcid = {PMC13250166}
}
RIS
TY - JOUR
AU - Sun, Yuan-Chen
AU - Jiang, Wen-Jie
AU - Cai, Kang-Wen
AU - Wei, Na-Na
AU - Lai, Fu-Ting
AU - Wang, Hao-Jie
AU - Gao, Rui-Xiang
AU - Kuang, Ze-Yu
AU - Zhou, Jia-Lu
AU - Liu, An
AU - Zhu, Han-Wen
AU - Wang, Yu-Juan
AU - Xu, Ming
AU - Wu, Hua-Jun
TI - Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 5172
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution",
"container-title": "Nature communications",
"author": [
{
"family": "Sun",
"given": "Yuan-Chen"
},
{
"family": "Jiang",
"given": "Wen-Jie"
},
{
"family": "Cai",
"given": "Kang-Wen"
},
{
"family": "Wei",
"given": "Na-Na"
},
{
"family": "Lai",
"given": "Fu-Ting"
},
{
"family": "Wang",
"given": "Hao-Jie"
},
{
"family": "Gao",
"given": "Rui-Xiang"
},
{
"family": "Kuang",
"given": "Ze-Yu"
},
{
"family": "Zhou",
"given": "Jia-Lu"
},
{
"family": "Liu",
"given": "An"
},
{
"family": "Zhu",
"given": "Han-Wen"
},
{
"family": "Wang",
"given": "Yu-Juan"
},
{
"family": "Xu",
"given": "Ming"
},
{
"family": "Wu",
"given": "Hua-Jun"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "5172",
"DOI": "10.1038/
"PMID": "41980945",
"PMCID": "PMC13250166",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
14
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1371/journal.pgen.1012081
- ADNP regulates chromatin architecture and lineage fidelity during neural differentiation.Journal: PLoS geneticsIn common: mouse, 9 references
- [2] doi:10.21203/rs.3.rs-9927928/v1 [code]
- Genome-wide and allele-resolved maps of the radial architecture of the mouse genomeJournal: Research Square (preprint)In common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 4 references
- [3] doi:10.1101/gr.280394.124 [code]
- De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.Journal: Genome researchIn common: pysam, BEDTools, scikit-image, 3 other tools, 4 references
- [4] doi:10.1038/s42003-026-10462-y [code]
- SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.Journal: Communications biologyIn common: pysam, BEDTools, PyTorch, 3 other tools, genetics / omics, mouse, 3 references
- [5] doi:10.1093/bioinformatics/btag652 [code]
- mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.Journal: Bioinformatics (Oxford, England)In common: pysam, PyTorch Lightning, BEDTools, 4 other tools, genetics / omics, mouse, 1 reference
- [6] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: pysam, BEDTools, PyTorch, 3 other tools, genetics / omics, mouse, 2 references
- [7] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: pysam, BEDTools, SAMtools, 4 other tools, genetics / omics, 1 reference
- [8] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: pysam, BEDTools, SAMtools, 3 other tools, mouse, 2 references
- [9] doi:10.1038/s41592-026-03211-w [code]
- Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.Journal: Nature methodsIn common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 1 reference
- [10] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: pysam, BEDTools, SAMtools, 3 other tools, genetics / omics, mouse, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 54 scripts, and 15 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:ec2ef33e18a379d7…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
