High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND.
The 9 matches
- [1] § Methods › Basic architecture of scBOND › The encoders in scBOND ↔ sccross/models/layers.py, lines 24–142 · score 0.77 · neural network, negative slope, LeakyReLU, dropout rate, layers, module
- [2] § Methods › Basic architecture of scBOND › The translator in scBOND ↔ scBond/model_component.py, lines 525–641 · score 0.75 · ReLU, gating network, expert networks, Softmax, probability, activation
- [3] § Results › scBOND preserves cellular heterogeneity and improves the discrimination of similar cell types in early embryonic development ↔ scBond/calculate_cluster.py, lines 6–42 · score 0.73 · adjusted mutual information, normalized mutual information, adjusted Rand, homogeneity, AMI, NMI
- [4] § Results › scBOND preserves cellular heterogeneity and improves the discrimination of similar cell types in early embryonic development ↔ scBond/calculate_cluster.py, lines 6–42 · score 0.73 · adjusted mutual information, normalized mutual information, adjusted Rand, homogeneity, AMI, NMI
- [5] § Methods › The training procedure of scBOND ↔ scBond/bond.py, lines 428–577 · score 0.63 · KL divergence, loss function, lr, patience, discriminators, reconstruction
- [6] § Methods › Data collection and preprocessing ↔ scBond/bond.py, lines 131–214 · score 0.58 · highly variable genes, HVGs, imputation, median, min, preprocessing
- [7] § Methods › Basic architecture of scBOND › The translator in scBOND ↔ scBond/model_component.py, lines 525–641 · score 0.55 · gating network, expert networks, space, block, weights, latent
- [8] § Methods › Data collection and preprocessing ↔ scBond/data_processing.py, lines 205–297 · score 0.53 · min max, imputation, median, preprocessing, methylation, cell
- [9] § Methods › The training procedure of scBOND ↔ scBond/train_model.py, lines 607–666 · score 0.51 · reconstruction loss, Adam, KL, patience, optimize, discriminators
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 830 lines · 31 KB · MIT · 2 matches
- import torch
- import torch.nn as nn
- import torch.nn.functional as F
- class NetBlock(nn.Module):
- def __init__(
- self,
- nlayer: int,
- dim_list: list,
- act_list: list,
- dropout_rate: float,
- noise_rate: float
- ):
- """
- multiple layers netblock with specific layer counts, dimension, activations and dropout.
- Parameters
- ----------
- nlayer
- layer counts.
- dim_list
- dimension list, length equal to nlayer + 1.
- act_list
- activation list, length equal to nlayer + 1.
- dropout_rate
- rate of dropout.
- noise_rate
- rate of set part of input data to 0.
- """
- super(NetBlock, self).__init__()
- self.nlayer = nlayer
- self.noise_dropout = nn.Dropout(noise_rate)
- self.linear_list = nn.ModuleList()
- self.bn_list = nn.ModuleList()
- self.activation_list = nn.ModuleList()
- self.dropout_list = nn.ModuleList()
- for i in range(nlayer):
- self.linear_list.append(nn.Linear(dim_list[i], dim_list[i + 1]))
- nn.init.xavier_uniform_(self.linear_list[i].weight)
- self.bn_list.append(nn.BatchNorm1d(dim_list[i + 1]))
- self.activation_list.append(act_list[i])
- if not i == nlayer -1:
- self.dropout_list.append(nn.Dropout(dropout_rate))
- def forward(self, x):
- x = self.noise_dropout(x)
- for i in range(self.nlayer):
- x = self.linear_list[i](x)
- x = self.bn_list[i](x)
- x = self.activation_list[i](x)
- if not i == self.nlayer -1:
- """ don't use dropout for output to avoid loss calculate break down """
- x = self.dropout_list[i](x)
- return x
- class Split_Chrom_Encoder_block(nn.Module):
- def __init__(
- self,
- nlayer: int,
- dim_list: list,
- act_list: list,
- chrom_list: list,
- dropout_rate: float,
- noise_rate: float
- ):
- """
- MET encoder netblock with specific layer counts, dimension, activations and dropout.
- Parameters
- ----------
- nlayer
- layer counts.
- dim_list
- dimension list, length equal to nlayer + 1.
- act_list
- activation list, length equal to nlayer + 1.
- chrom_list
- list record the peaks count for each chrom, assert that sum of chrom list equal to dim_list[0].
- dropout_rate
- rate of dropout.
- noise_rate
- rate of set part of input data to 0.
- """
- super(Split_Chrom_Encoder_block, self).__init__()
- self.nlayer = nlayer
- self.chrom_list = chrom_list
- self.noise_dropout = nn.Dropout(noise_rate)
- self.linear_list = nn.ModuleList()
- self.bn_list = nn.ModuleList()
- self.activation_list = nn.ModuleList()
- self.dropout_list = nn.ModuleList()
- for i in range(nlayer):
- if i == 0:
- """first layer seperately forward for each chrom"""
- self.linear_list.append(nn.ModuleList())
- self.bn_list.append(nn.ModuleList())
- self.activation_list.append(nn.ModuleList())
- self.dropout_list.append(nn.ModuleList())
- for j in range(len(chrom_list)):
- self.linear_list[i].append(nn.Linear(chrom_list[j], dim_list[i + 1] // len(chrom_list)))
- nn.init.xavier_uniform_(self.linear_list[i][j].weight)
- self.bn_list[i].append(nn.BatchNorm1d(dim_list[i + 1] // len(chrom_list)))
- self.activation_list[i].append(act_list[i])
- self.dropout_list[i].append(nn.Dropout(dropout_rate))
- else:
- self.linear_list.append(nn.Linear(dim_list[i], dim_list[i + 1]))
- nn.init.xavier_uniform_(self.linear_list[i].weight)
- self.bn_list.append(nn.BatchNorm1d(dim_list[i + 1]))
- self.activation_list.append(act_list[i])
- if not i == nlayer -1:
- self.dropout_list.append(nn.Dropout(dropout_rate))
- def forward(self, x):
- x = self.noise_dropout(x)
- for i in range(self.nlayer):
- if i == 0:
- x = torch.split(x, self.chrom_list, dim = 1)
- temp = []
- for j in range(len(self.chrom_list)):
- temp.append(self.dropout_list[0][j](self.activation_list[0][j](self.bn_list[0][j](self.linear_list[0][j](x[j])))))
- x = torch.concat(temp, dim = 1)
- else:
- x = self.linear_list[i](x)
- x = self.bn_list[i](x)
- x = self.activation_list[i](x)
- if not i == self.nlayer -1:
- """ don't use dropout for output to avoid loss calculate break down """
- x = self.dropout_list[i](x)
- return x
- class Split_Chrom_Decoder_block(nn.Module):
- def __init__(
- self,
- nlayer: int,
- dim_list: list,
- act_list: list,
- chrom_list: list,
- dropout_rate: float,
- noise_rate: float
- ):
- """
- MET decoder netblock with specific layer counts, dimension, activations and dropout.
- Parameters
- ----------
- nlayer
- layer counts.
- dim_list
- dimension list, length equal to nlayer + 1.
- act_list
- activation list, length equal to nlayer + 1.
- chrom_list
- list record the peaks count for each chrom, assert that sum of chrom list equal to dim_list[end].
- dropout_rate
- rate of dropout.
- noise_rate
- rate of set part of input data to 0.
- """
- super(Split_Chrom_Decoder_block, self).__init__()
- self.nlayer = nlayer
- self.noise_dropout = nn.Dropout(noise_rate)
- self.chrom_list = chrom_list
- self.linear_list = nn.ModuleList()
- self.bn_list = nn.ModuleList()
- self.activation_list = nn.ModuleList()
- self.dropout_list = nn.ModuleList()
- for i in range(nlayer):
- if not i == nlayer -1:
- self.linear_list.append(nn.Linear(dim_list[i], dim_list[i + 1]))
- nn.init.xavier_uniform_(self.linear_list[i].weight)
- self.bn_list.append(nn.BatchNorm1d(dim_list[i + 1]))
- self.activation_list.append(act_list[i])
- self.dropout_list.append(nn.Dropout(dropout_rate))
- else:
- """last layer seperately forward for each chrom"""
- self.linear_list.append(nn.ModuleList())
- self.bn_list.append(nn.ModuleList())
- self.activation_list.append(nn.ModuleList())
- self.dropout_list.append(nn.ModuleList())
- for j in range(len(chrom_list)):
- self.linear_list[i].append(nn.Linear(dim_list[i] // len(chrom_list), chrom_list[j]))
- nn.init.xavier_uniform_(self.linear_list[i][j].weight)
- self.bn_list[i].append(nn.BatchNorm1d(chrom_list[j]))
- self.activation_list[i].append(act_list[i])
- def forward(self, x):
- x = self.noise_dropout(x)
- for i in range(self.nlayer):
- if not i == self.nlayer -1:
- x = self.linear_list[i](x)
- x = self.bn_list[i](x)
- x = self.activation_list[i](x)
- x = self.dropout_list[i](x)
- else:
- x = torch.chunk(x, len(self.chrom_list), dim = 1)
- temp = []
- for j in range(len(self.chrom_list)):
- temp.append(self.activation_list[i][j](self.bn_list[i][j](self.linear_list[i][j](x[j]))))
- x = torch.concat(temp, dim = 1)
- return x
- class SELayer(nn.Module):
- def __init__(self, channel, reduction=16):
- super(SELayer, self).__init__()
- self.fc = nn.Sequential(
- nn.Linear(channel, channel // reduction, bias=False),
- nn.ReLU(inplace=True),
- nn.Linear(channel // reduction, channel, bias=False),
- nn.Sigmoid()
- )
- def forward(self, x):
- b, c = x.size()
- y = self.fc(x)
- return x * y.view(b, c)
- class SelfAttention(nn.Module):
- def __init__(
- self,
- dim,
- num_heads=8,
- qkv_bias=False,
- attn_drop=0.,
- proj_drop=0.
- ):
- """
- Multi-head self-attention module that computes attention between input and context tensors.
- Parameters
- ----------
- dim: int
- Dimension of the input features and embedding.
- num_heads: int
- Number of parallel attention heads, default 8.
- qkv_bias: bool
- Whether to include bias in the query, key, and value linear projections, default False.
- attn_drop: float
- Dropout probability applied to attention weights, default 0.0.
- proj_drop: float
- Dropout probability applied to the output projection, default 0.0.
- """
- super().__init__()
- self.num_heads = num_heads
- self.scale = (dim // num_heads) ** -0.5
- self.q = nn.Linear(dim, dim, bias=qkv_bias)
- self.k = nn.Linear(dim, dim, bias=qkv_bias)
- self.v = nn.Linear(dim, dim, bias=qkv_bias)
- self.attn_drop = nn.Dropout(attn_drop)
- self.proj = nn.Linear(dim, dim)
- self.proj_drop = nn.Dropout(proj_drop)
- def forward(self, x, context):
- B, N, C = x.shape
- q = self.q(x).reshape(B, N, self.num_heads, C // self.num_heads).permute(0, 2, 1, 3)
- k = self.k(context).reshape(B, N, self.num_heads, C // self.num_heads).permute(0, 2, 1, 3)
- v = self.v(context).reshape(B, N, self.num_heads, C // self.num_heads).permute(0, 2, 1, 3)
- attn = (q @ k.transpose(-2, -1)) * self.scale
- attn = attn.softmax(dim=-1)
- attn = self.attn_drop(attn)
- x = (attn @ v).transpose(1, 2).reshape(B, N, C)
- x = self.proj(x)
- x = self.proj_drop(x)
- return x
- class Translator(nn.Module):
- def __init__(
- self,
- translator_input_dim_r: int,
- translator_input_dim_m: int,
- translator_embed_dim: int,
- translator_embed_act_list: list,
- num_experts: int = 4,
- reduction_ratio: int = 64,
- num_heads=8,
- attn_drop=0.1,
- proj_drop=0.1
- ):
- """
- Translator block with MoE mechanism used for translation between different omics.
- Parameters
- ----------
- translator_input_dim_r
- dimension of input from RNA encoder for translator.
- translator_input_dim_a
- dimension of input from MET encoder for translator.
- translator_embed_dim
- dimension of embedding space for translator.
- translator_embed_act_list
- activation list for translator, involving [mean_activation, log_var_activation, decoder_activation].
- num_experts
- number of expert networks.
- reduction_ratio
- The reduction ratio for the SE layer.
- num_heads: int
- Number of parallel attention heads, default 8.
- attn_drop: float
- Dropout probability applied to attention weights, default 0.0.
- proj_drop: float
- Dropout probability applied to the output projection, default 0.0.
- """
- super(Translator, self).__init__()
- mean_activation, log_var_activation, decoder_activation = translator_embed_act_list
- self.expert_networks_r_mu = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim_r, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- mean_activation
- ) for _ in range(num_experts)
- ])
- self.expert_networks_m_mu = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim_m, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- mean_activation
- ) for _ in range(num_experts)
- ])
- self.expert_networks_r_d = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim_r, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- log_var_activation
- ) for _ in range(num_experts)
- ])
- self.expert_networks_m_d = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim_m, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- log_var_activation
- ) for _ in range(num_experts)
- ])
- self.gating_network_r_mu = nn.Sequential(
- nn.Linear(translator_input_dim_r, 64),
- nn.ReLU(),
- nn.Linear(64, num_experts),
- nn.Softmax(dim=1)
- )
- self.gating_network_m_mu = nn.Sequential(
- nn.Linear(translator_input_dim_m, 64),
- nn.ReLU(),
- nn.Linear(64, num_experts),
- nn.Softmax(dim=1)
- )
- self.gating_network_r_d = nn.Sequential(
- nn.Linear(translator_input_dim_r, 64),
- nn.ReLU(),
- nn.Linear(64, num_experts),
- nn.Softmax(dim=1)
- )
- self.gating_network_m_d = nn.Sequential(
- nn.Linear(translator_input_dim_m, 64),
- nn.ReLU(),
- nn.Linear(64, num_experts),
- nn.Softmax(dim=1)
- )
- self.RNA_decoder_l = nn.Linear(translator_embed_dim, translator_input_dim_r)
- nn.init.xavier_uniform_(self.RNA_decoder_l.weight)
- self.RNA_decoder_bn = nn.BatchNorm1d(translator_input_dim_r)
- self.RNA_decoder_act = decoder_activation
- self.MET_decoder_l = nn.Linear(translator_embed_dim, translator_input_dim_m)
- nn.init.xavier_uniform_(self.MET_decoder_l.weight)
- self.MET_decoder_bn = nn.BatchNorm1d(translator_input_dim_m)
- self.MET_decoder_act = decoder_activation
- self.se_rna = SELayer(translator_embed_dim, reduction_ratio)
- self.se_met = SELayer(translator_embed_dim, reduction_ratio)
- self.self_attn_r2m = SelfAttention(
- dim=translator_embed_dim,
- num_heads=num_heads,
- attn_drop=attn_drop,
- proj_drop=proj_drop
- )
- self.self_attn_m2r = SelfAttention(
- dim=translator_embed_dim,
- num_heads=num_heads,
- attn_drop=attn_drop,
- proj_drop=proj_drop
- )
- def reparameterize(self, mu, sigma):
- sigma = torch.exp(sigma / 2)
- eps = torch.randn_like(sigma)
- return mu + eps * sigma
- def forward_with_RNA(self, x, forward_type):
- expert_outputs_r_mu = [expert_network_mu(x) for expert_network_mu in self.expert_networks_r_mu]
- gating_weights_r_mu = self.gating_network_r_mu(x)
- gating_weights_r_mu = gating_weights_r_mu.unsqueeze(2)
- expert_outputs_r_mu = torch.stack(expert_outputs_r_mu, dim=1)
- weighted_output_r_mu = (expert_outputs_r_mu * gating_weights_r_mu).sum(dim=1)
- expert_outputs_r_d = [expert_network_d(x) for expert_network_d in self.expert_networks_r_d]
- gating_weights_r_d = self.gating_network_r_d(x)
- gating_weights_r_d = gating_weights_r_d.unsqueeze(2)
- expert_outputs_r_d = torch.stack(expert_outputs_r_d, dim=1)
- weighted_output_r_d = (expert_outputs_r_d * gating_weights_r_d).sum(dim=1)
- if forward_type == 'test':
- latent_layer = weighted_output_r_mu
- elif forward_type == 'train':
- latent_layer = self.reparameterize(weighted_output_r_mu, weighted_output_r_d)
- latent_layer = self.se_rna(latent_layer)
- latent_layer_R_out = self.RNA_decoder_act(self.RNA_decoder_bn(self.RNA_decoder_l(latent_layer)))
- latent_layer_M_out = self.MET_decoder_act(self.MET_decoder_bn(self.MET_decoder_l(latent_layer)))
- return latent_layer_R_out, latent_layer_M_out, weighted_output_r_mu, weighted_output_r_d
- def forward_with_MET(self, x, forward_type):
- expert_outputs_m_mu = [expert_network_mu(x) for expert_network_mu in self.expert_networks_m_mu]
- gating_weights_m_mu = self.gating_network_m_mu(x)
- gating_weights_m_mu = gating_weights_m_mu.unsqueeze(2)
- expert_outputs_m_mu = torch.stack(expert_outputs_m_mu, dim=1)
- weighted_output_m_mu = (expert_outputs_m_mu * gating_weights_m_mu).sum(dim=1)
- expert_outputs_m_d = [expert_network_d(x) for expert_network_d in self.expert_networks_m_d]
- gating_weights_m_d = self.gating_network_m_d(x)
- gating_weights_m_d = gating_weights_m_d.unsqueeze(2)
- expert_outputs_m_d = torch.stack(expert_outputs_m_d, dim=1)
- weighted_output_m_d = (expert_outputs_m_d * gating_weights_m_d).sum(dim=1)
- x = x.unsqueeze(1)
- weighted_output_m_mu = weighted_output_m_mu.unsqueeze(1)
- self_attn_output = self.self_attn_m2r(x, x)
- weighted_output_m_mu = weighted_output_m_mu + self_attn_output # 残差连接
- weighted_output_m_mu = weighted_output_m_mu.squeeze(1)
- weighted_output_m_d = weighted_output_m_d.squeeze(1)
- if forward_type == 'test':
- latent_layer = weighted_output_m_mu
- elif forward_type == 'train':
- latent_layer = self.reparameterize(weighted_output_m_mu, weighted_output_m_d)
- latent_layer = self.se_met(latent_layer)
- latent_layer_R_out = self.RNA_decoder_act(self.RNA_decoder_bn(self.RNA_decoder_l(latent_layer)))
- latent_layer_M_out = self.MET_decoder_act(self.MET_decoder_bn(self.MET_decoder_l(latent_layer)))
- return latent_layer_R_out, latent_layer_M_out, weighted_output_m_mu, weighted_output_m_d
- def train_model(self, x, input_type):
- if input_type == 'RNA':
- return self.forward_with_RNA(x, 'train')
- elif input_type == 'MET':
- return self.forward_with_MET(x, 'train')
- def test_model(self, x, input_type):
- if input_type == 'RNA':
- return self.forward_with_RNA(x, 'test')
- elif input_type == 'MET':
- return self.forward_with_MET(x, 'test')
- class Single_Translator(nn.Module):
- def __init__(
- self,
- translator_input_dim: int,
- translator_embed_dim: int,
- translator_embed_act_list: list,
- num_experts_single: int = 6,
- reduction_ratio: int = 64,
- num_heads: int = 8,
- attn_drop: float = 0.1,
- proj_drop: float = 0.1
- ):
- """
- Single translator block with MoE mechanism used only for pretraining.
- Parameters
- ----------
- translator_input_dim
- dimension of input from encoder for translator.
- translator_embed_dim
- dimension of embedding space for translator.
- translator_embed_act_list
- activation list for translator, involving [mean_activation, log_var_activation, decoder_activation].
- num_experts_single
- number of expert networks, default 6.
- reduction_ratio
- The reduction ratio for the SE layer, default 64.
- num_heads: int
- Number of parallel attention heads, default 8.
- attn_drop: float
- Dropout probability applied to attention weights, default 0.0.
- proj_drop: float
- Dropout probability applied to the output projection, default 0.0.
- """
- super(Single_Translator, self).__init__()
- mean_activation, log_var_activation, decoder_activation = translator_embed_act_list
- self.expert_networks_mu = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- mean_activation
- ) for _ in range(num_experts_single)
- ])
- self.expert_networks_d = nn.ModuleList([
- nn.Sequential(
- nn.Linear(translator_input_dim, translator_embed_dim),
- nn.BatchNorm1d(translator_embed_dim),
- log_var_activation
- ) for _ in range(num_experts_single)
- ])
- self.gating_network = nn.Sequential(
- nn.Linear(translator_input_dim, 64),
- nn.ReLU(),
- nn.Linear(64, num_experts_single),
- nn.Softmax(dim=1)
- )
- self.decoder_l = nn.Linear(translator_embed_dim, translator_input_dim)
- nn.init.xavier_uniform_(self.decoder_l.weight)
- self.decoder_bn = nn.BatchNorm1d(translator_input_dim)
- self.decoder_act = decoder_activation
- self.se = SELayer(translator_embed_dim, reduction_ratio)
- self.self_attn = SelfAttention(
- dim=translator_embed_dim,
- num_heads=8,
- attn_drop=0.1,
- proj_drop=0.1
- )
- def reparameterize(self, mu, sigma):
- sigma = torch.exp(sigma / 2)
- eps = torch.randn_like(sigma)
- return mu + eps * sigma
- def forward(self, x, forward_type):
- expert_outputs_mu = [expert_network(x) for expert_network in self.expert_networks_mu]
- expert_outputs_d = [expert_network(x) for expert_network in self.expert_networks_d]
- gating_weights = self.gating_network(x)
- gating_weights = gating_weights.unsqueeze(2)
- expert_outputs_mu = torch.stack(expert_outputs_mu, dim=1)
- expert_outputs_d = torch.stack(expert_outputs_d, dim=1)
- weighted_output_mu = (expert_outputs_mu * gating_weights).sum(dim=1)
- weighted_output_d = (expert_outputs_d * gating_weights).sum(dim=1)
- weighted_output_mu = weighted_output_mu.unsqueeze(1)
- attn_output = self.self_attn(weighted_output_mu, weighted_output_mu)
- weighted_output_mu = weighted_output_mu + attn_output
- weighted_output_mu = weighted_output_mu.squeeze(1)
- if forward_type == 'test':
- latent_layer = weighted_output_mu
- elif forward_type == 'train':
- latent_layer = self.reparameterize(weighted_output_mu, weighted_output_d)
- latent_layer = self.se(latent_layer)
- latent_layer_out = self.decoder_act(self.decoder_bn(self.decoder_l(latent_layer)))
- return latent_layer_out, weighted_output_mu, weighted_output_d
- class Background_Translator(nn.Module):
- def __init__(
- self,
- translator_input_dim_r: int,
- translator_input_dim_m: int,
- translator_embed_dim: int,
- translator_embed_act_list: list,
- reduction_ratio: int = 64
- ):
- """
- Background Translator block with SE layer.
- Parameters
- ----------
- translator_input_dim_r
- dimension of input from RNA encoder for translator.
- translator_input_dim_m
- dimension of input from MET encoder for translator.
- translator_embed_dim
- dimension of embedding space for translator.
- translator_embed_act_list
- activation list for translator, involving [mean_activation, log_var_activation, decoder_activation].
- reduction_ratio
- The reduction ratio for the SE layer, default 64.
- """
- super(Background_Translator, self).__init__()
- mean_activation, log_var_activation, decoder_activation = translator_embed_act_list
- self.RNA_encoder_l_mu = nn.Linear(translator_input_dim_r, translator_embed_dim)
- nn.init.xavier_uniform_(self.RNA_encoder_l_mu.weight)
- self.RNA_encoder_bn_mu = nn.BatchNorm1d(translator_embed_dim)
- self.RNA_encoder_act_mu = mean_activation
- self.MET_encoder_l_mu = nn.Linear(translator_input_dim_m, translator_embed_dim)
- nn.init.xavier_uniform_(self.MET_encoder_l_mu.weight)
- self.MET_encoder_bn_mu = nn.BatchNorm1d(translator_embed_dim)
- self.MET_encoder_act_mu = mean_activation
- self.RNA_encoder_l_d = nn.Linear(translator_input_dim_r, translator_embed_dim)
- nn.init.xavier_uniform_(self.RNA_encoder_l_d.weight)
- self.RNA_encoder_bn_d = nn.BatchNorm1d(translator_embed_dim)
- self.RNA_encoder_act_d = log_var_activation
- self.MET_encoder_l_d = nn.Linear(translator_input_dim_m, translator_embed_dim)
- nn.init.xavier_uniform_(self.MET_encoder_l_d.weight)
- self.MET_encoder_bn_d = nn.BatchNorm1d(translator_embed_dim)
- self.MET_encoder_act_d = log_var_activation
- self.RNA_decoder_l = nn.Linear(translator_embed_dim, translator_input_dim_r)
- nn.init.xavier_uniform_(self.RNA_decoder_l.weight)
- self.RNA_decoder_bn = nn.BatchNorm1d(translator_input_dim_r)
- self.RNA_decoder_act = decoder_activation
- self.MET_decoder_l = nn.Linear(translator_embed_dim, translator_input_dim_m)
- nn.init.xavier_uniform_(self.MET_decoder_l.weight)
- self.MET_decoder_bn = nn.BatchNorm1d(translator_input_dim_m)
- self.MET_decoder_act = decoder_activation
- self.background_pro_alpha = nn.Parameter(torch.randn(1, translator_input_dim_m))
- self.background_pro_log_beta = nn.Parameter(torch.clamp(torch.randn(1, translator_input_dim_m), -10, 1))
- self.scale_parameters_l = nn.Linear(translator_embed_dim, translator_embed_dim)
- self.scale_parameters_bn = nn.BatchNorm1d(translator_embed_dim)
- self.scale_parameters_act = mean_activation
- self.pi_l = nn.Linear(translator_embed_dim, 1)
- self.pi_bn = nn.BatchNorm1d(1)
- self.pi_act = nn.Sigmoid()
- self.se_rna = SELayer(translator_embed_dim, reduction_ratio)
- self.se_met = SELayer(translator_embed_dim, reduction_ratio)
- def reparameterize(self, mu, sigma):
- sigma = torch.exp(sigma / 2)
- eps = torch.randn_like(sigma)
- return mu + eps * sigma
- def forward_with_RNA(self, x, forward_type):
- latent_layer_mu = self.RNA_encoder_act_mu(self.RNA_encoder_bn_mu(self.RNA_encoder_l_mu(x)))
- latent_layer_d = self.RNA_encoder_act_d(self.RNA_encoder_bn_d(self.RNA_encoder_l_d(x)))
- if forward_type == 'test':
- latent_layer = latent_layer_mu
- elif forward_type == 'train':
- latent_layer = self.reparameterize(latent_layer_mu, latent_layer_d)
- latent_layer = self.se_rna(latent_layer)
- latent_layer_R_out = self.RNA_decoder_act(self.RNA_decoder_bn(self.RNA_decoder_l(latent_layer)))
- latent_layer_M_out = self.MET_decoder_act(self.MET_decoder_bn(self.MET_decoder_l(latent_layer)))
- return latent_layer_R_out, latent_layer_M_out, latent_layer_mu, latent_layer_d
- def forward_with_MET(self, x, forward_type):
- latent_layer = self.forward_adt(x, forward_type)
- latent_layer = self.se_met(latent_layer)
- latent_layer_R_out = self.RNA_decoder_act(self.RNA_decoder_bn(self.RNA_decoder_l(latent_layer)))
- latent_layer_M_out = self.MET_decoder_act(self.MET_decoder_bn(self.MET_decoder_l(latent_layer)))
- return latent_layer_R_out, latent_layer_M_out, latent_layer, latent_layer
- def train_model(self, x, input_type):
- if input_type == 'RNA':
- return self.forward_with_RNA(x, 'train')
- elif input_type == 'MET':
- return self.forward_with_MET(x, 'train')
- def test_model(self, x, input_type):
- if input_type == 'RNA':
- return self.forward_with_RNA(x, 'test')
- elif input_type == 'MET':
- return self.forward_with_MET(x, 'test')
- class Background_Single_Translator(nn.Module):
- def __init__(
- self,
- translator_input_dim: int,
- translator_embed_dim: int,
- translator_embed_act_list: list,
- reduction_ratio: int = 64
- ):
- """
- Background Single Translator block with SE layer.
- Parameters
- ----------
- translator_input_dim
- dimension of input from encoder for translator.
- translator_embed_dim
- dimension of embedding space for translator.
- translator_embed_act_list
- activation list for translator, involving [mean_activation, log_var_activation, decoder_activation].
- reduction_ratio
- The reduction ratio for the SE layer, default 64.
- """
- super(Background_Single_Translator, self).__init__()
- mean_activation, log_var_activation, decoder_activation = translator_embed_act_list
- self.encoder_l_mu = nn.Linear(translator_input_dim, translator_embed_dim)
- nn.init.xavier_uniform_(self.encoder_l_mu.weight)
- self.encoder_bn_mu = nn.BatchNorm1d(translator_embed_dim)
- self.encoder_act_mu = mean_activation
- self.encoder_l_d = nn.Linear(translator_input_dim, translator_embed_dim)
- nn.init.xavier_uniform_(self.encoder_l_d.weight)
- self.encoder_bn_d = nn.BatchNorm1d(translator_embed_dim)
- self.encoder_act_d = log_var_activation
- self.decoder_l = nn.Linear(translator_embed_dim, translator_input_dim)
- nn.init.xavier_uniform_(self.decoder_l.weight)
- self.decoder_bn = nn.BatchNorm1d(translator_input_dim)
- self.decoder_act = decoder_activation
- self.background_pro_alpha = nn.Parameter(torch.randn(1, translator_input_dim))
- self.background_pro_log_beta = nn.Parameter(torch.clamp(torch.randn(1, translator_input_dim), -10, 1))
- self.scale_parameters_l = nn.Linear(translator_embed_dim, translator_embed_dim)
- self.scale_parameters_bn = nn.BatchNorm1d(translator_embed_dim)
- self.scale_parameters_act = mean_activation
- self.pi_l = nn.Linear(translator_embed_dim, 1)
- self.pi_bn = nn.BatchNorm1d(1)
- self.pi_act = nn.Sigmoid()
- self.se = SELayer(translator_embed_dim, reduction_ratio)
- def reparameterize(self, mu, sigma):
- sigma = torch.exp(sigma / 2)
- eps = torch.randn_like(sigma)
- return mu + eps * sigma
- def forward(self, x, forward_type):
- latent_layer = self.forward_adt(x, forward_type)
- latent_layer = self.se(latent_layer)
- latent_layer_out = self.decoder_act(self.decoder_bn(self.decoder_l(latent_layer)))
- return latent_layer_out, latent_layer, latent_layer
model_component.py at commit 22a2192, under MIT · at the source
Overview
- School of Mathematical Sciences and LPMC, Nankai University, Tianjin 300071, China
- College of Electronic Information and Optical Engineering, Nankai University, Tianjin 300350, China
- Academy for Advanced Interdisciplinary Studies, Nankai University, Tianjin 300071, China
Abstract
Single-cell multiomic sequencing technologies have offered unprecedented insights into cellular heterogeneity by jointly profiling gene expression and epigenetic landscapes at single-cell resolution. However, the application of these technologies remains limited owing to technical challenges and high costs. Computational approaches for cross-modality translation provide a promising solution to these limitations by enabling the inference of one modality from another. However, existing methods for cross-modality translation between single-cell RNA sequencing (scRNA-seq) and single-cell DNA methylation (scDNAm) data face limitations, including unidirectionality, inadequate modeling of context-specific DNA methylation–expression associations, neglect of biological relevance in evaluation, and poor performance in limited paired training data. To fill these gaps, we introduce scBOND, a bidirectional cross-modal translation framework tailored for scRNA-seq and scDNAm profiles. scBOND leverages a mixture-of-experts block to capture context-dependent regulatory patterns, while implementing self-attention mechanism and a feature recalibration module to enhance biological signal fidelity. Extensive experiments demonstrate scBOND consistently outperforms baseline methods in both translation directions, yielding high-accuracy translation while preserving cellular structure. In mouse embryonic data, scBOND preserves subtle, functionally significant differences between closely related cell types, which are undetected in the original data. Downstream analyses confirm that scBOND effectively recovers tissue-specific signals in human brain neurons. Moreover, using RNA-only data, we reconstruct scDNAm profiles and identify cell-type- and stage-specific regulatory mechanisms in oligodendrocyte lineage. To further improve model generalization in paired data-scarce scenarios, we propose scBOND-Aug, a variant of scBOND equipped with a biologically informed data augmentation strategy, which demonstrates superior results with limited paired data.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.
mcgilldinglab/scCross
586d415ae67c89225d12312c84588d02434caa0c, 9 August 2025Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
31 files
- data/
COVID-19/ , Python, 31 linespreprocess.py - data/
human_cell_atlas/ , Python, 151 linespreprocess.py - data/
matched_mouse_atheroscle , Python, 31 linesrotic_plaque_immune_cell s/ preprocess.py - data/
matched_mouse_cortex/ , Python, 26 linespreprocess.py - data/
matched_mouse_lymphonodu , Python, 28 liness/ preprocess.py - data/
unmatched_mouse_cortex/ , Python, 90 linespreprocess.py - docs/
COVID-19.ipynb , Jupyter, 165 lines - docs/
conf.py , Python, 89 lines - docs/
human_cell_atlas.ipynb , Jupyter, 116 lines - docs/
matched_mouse_atheroscle , Jupyter, 175 linesrotic_plaque_immune_cell s.ipynb - docs/
matched_mouse_cortex.ipy , Jupyter, 343 linesnb - docs/
matched_mouse_lymphonodu , Jupyter, 172 liness.ipynb - docs/
unmatched_mouse_cortex.i , Jupyter, 188 linespynb - examples/
COVID-19.ipynb , Jupyter, 165 lines - examples/
human_cell_atlas.ipynb , Jupyter, 116 lines - examples/
matched_mouse_atheroscle , Jupyter, 175 linesrotic_plaque_immune_cell s.ipynb - examples/
matched_mouse_cortex.ipy , Jupyter, 172 linesnb - examples/
matched_mouse_lymphonodu , Jupyter, 172 liness.ipynb - examples/
unmatched_mouse_cortex.i , Jupyter, 188 linespynb - sccross/
__init__.py , Python, 18 lines - sccross/
data.py , Python, 1,054 lines - sccross/
metrics.py , Python, 320 lines - sccross/
models/ , Python, 87 lines__init__.py - sccross/
models/ , Python, 356 linesdata.py - sccross/
models/ , Python, 640 lines, 1 matchlayers.py - sccross/
models/ , Python, 2,299 linessccross.py - sccross/
models/ , Python, 861 linesutils.py - sccross/
utils.py , Python, 1,030 lines - setup.py, Python, 11 lines
- LICENSE, License, 21 lines
- README.md, Text, 51 lines
tanlabcode/MAPLE.1.0
f707eaf880b5c373b81b559c4494a5502f4b7005, 20 September 2021Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
7 files
- R/
classes.R , R, 31 lines - R/
feature_and_response_pro , R, 381 linescessing.R - R/
hello.R , R, 18 lines - R/
met_matrix_processing.R , R, 692 lines - R/
prediction.R , R, 143 lines - R/
set_ops.R , R, 34 lines - README.md, Text, 88 lines
BioX-NKU/scBOND
22a21923897186ce3ca55e893a9693ae05583d1d, 22 April 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
24 files
- docs/
source/ , Jupyter, 137 linesTutorial/ RNA_DNAm_paired_basic/ .ipynb_checkpoints/ scBOND_usage-checkpoint. ipynb - docs/
source/ , Jupyter, 137 linesTutorial/ RNA_DNAm_paired_basic/ scBOND_usage.ipynb - docs/
source/ , Jupyter, 112 linesTutorial/ RNA_DNAm_variants/ .ipynb_checkpoints/ scBOND_aug_usage-checkpo int.ipynb - docs/
source/ , Jupyter, 123 linesTutorial/ RNA_DNAm_variants/ .ipynb_checkpoints/ scBOND_for_single_modali ty-checkpoint.ipynb - docs/
source/ , Jupyter, 112 linesTutorial/ RNA_DNAm_variants/ scBOND_aug_usage.ipynb - docs/
source/ , Jupyter, 123 linesTutorial/ RNA_DNAm_variants/ scBOND_for_single_modali ty.ipynb - docs/
source/ , Python, 95 linesconf.py - examples/
scBOND_aug_usage.ipynb , Jupyter, 112 lines - examples/
scBOND_for_single_modali , Jupyter, 123 linesty.ipynb - examples/
scBOND_usage.ipynb , Jupyter, 137 lines - scBond/
__init__.py , Python, 6 lines - scBond/
bond.py , Python, 786 lines, 2 matches - scBond/
calculate_cluster.py , Python, 42 lines, 1 match - scBond/
data_processing.py , Python, 297 lines, 1 match - scBond/
draw_cluster.py , Python, 132 lines - scBond/
logger.py , Python, 50 lines - scBond/
model_component.py , Python, 830 lines, 2 matches - scBond/
model_utlis.py , Python, 350 lines - scBond/
split_datasets.py , Python, 107 lines - scBond/
train_model.py , Python, 1,230 lines, 1 match - scBond/
version.py , Python, 1 line - setup.py, Python, 43 lines
- LICENSE, License, 21 lines
- README.md, Text, 165 lines
Zenodo 17699419
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
- 29 September 2026: the link answers (HTTP 200)
17 files
- example/
scBOND_aug_usage.ipynb , Jupyter, 112 lines - example/
scBOND_for_single_modal. , Jupyter, 123 linesipynb - example/
scBOND_usage.ipynb , Jupyter, 137 lines - scBond/
__init__.py , Python, 6 lines - scBond/
bond.py , Python, 786 lines - scBond/
calculate_cluster.py , Python, 42 lines, 1 match - scBond/
data_processing.py , Python, 297 lines - scBond/
draw_cluster.py , Python, 132 lines - scBond/
logger.py , Python, 50 lines - scBond/
model_component.py , Python, 830 lines - scBond/
model_utlis.py , Python, 350 lines - scBond/
split_datasets.py , Python, 107 lines - scBond/
train_model.py , Python, 1,230 lines - scBond/
version.py , Python, 1 line - setup.py, Python, 43 lines
- LICENSE.txt, License, 21 lines
- README.md, Text, 161 lines
Code availability
The MIT-licensed scBOND software, including detailed documents and tutorials, is freely available at GitHub (https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 72 scripts, each with its path and the digest of its content;
- 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Data links
- ncbi.nlm.nih.gov/
geo , NCBI; found in the text, “Data collection and preprocessing”
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 9 MeSH terms, 4 funders, 48 references.
Cite
This paper
Lang, K., Jia, C., Li, S., Guo, Y., Hu, D., & Chen, S. (2026). High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND. Genome research, 36(4), 769-784. https://
BibTeX
@article{lang2026high,
author = {Lang, Kehan and Jia, Chenyang and Li, Siyu and Guo, Yi and Hu, Dingjun and Chen, Shengquan},
title = {{High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND}},
journal = {Genome research},
year = {2026},
month = apr,
volume = {36},
number = {4},
pages = {769--784},
publisher = {Cold Spring Harbor Laboratory Press},
issn = {1088-9051},
doi = {10.1101/
url = {https://
pmid = {41887797},
pmcid = {PMC13138010}
}
RIS
TY - JOUR
AU - Lang, Kehan
AU - Jia, Chenyang
AU - Li, Siyu
AU - Guo, Yi
AU - Hu, Dingjun
AU - Chen, Shengquan
TI - High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND
T2 - Genome research
J2 - Genome Res
PY - 2026
DA - 2026/
VL - 36
IS - 4
SP - 769
EP - 784
SN - 1088-9051
PB - Cold Spring Harbor Laboratory Press
DO - 10.1101/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1101/
"type": "article-journal",
"title": "High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND",
"container-title": "Genome research",
"author": [
{
"family": "Lang",
"given": "Kehan"
},
{
"family": "Jia",
"given": "Chenyang"
},
{
"family": "Li",
"given": "Siyu"
},
{
"family": "Guo",
"given": "Yi"
},
{
"family": "Hu",
"given": "Dingjun"
},
{
"family": "Chen",
"given": "Shengquan"
}
],
"container-title-short":
"volume": "36",
"issue": "4",
"page": "769-784",
"DOI": "10.1101/
"PMID": "41887797",
"PMCID": "PMC13138010",
"ISSN": "1088-9051",
"publisher": "Cold Spring Harbor Laboratory Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
7
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-73171-4 [code]
- Dissecting epigenetic heterogeneity in single-cell DNA methylomes with a unified framework.Journal: Nature communicationsIn common: BEDTools, anndata, Scanpy, 8 other tools, genetics / omics, cellular / molecular, 7 references
- [2] doi:10.1038/s41467-026-71803-3 [code]
- Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.Journal: Nature communicationsIn common: BEDTools, anndata, Scanpy, 9 other tools, genetics / omics, mouse, cellular / molecular, 4 references
- [3] doi:10.1002/advs.77524 [code]
- MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data.Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 6 references
- [4] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: BEDTools, anndata, Scanpy, 8 other tools, genetics / omics, mouse, cellular / molecular, 2 references
- [5] doi:10.1093/bioinformatics/btag652 [code]
- mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.Journal: Bioinformatics (Oxford, England)In common: BEDTools, anndata, Scanpy, 7 other tools, genetics / omics, mouse, 4 references
- [6] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: anndata, Scanpy, NetworkX, 9 other tools, genetics / omics, cellular / molecular, 2 references
- [7] doi:10.1038/s41467-026-68596-w [code]
- Spatial cartography of human thymus enables the geopositioning of lineage transcription factors in rare mimetic thymic epithelial cells.Journal: Nature communicationsIn common: anndata, Scanpy, NetworkX, 10 other tools, genetics / omics, cellular / molecular, 1 reference
- [8] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: anndata, Scanpy, NetworkX, 9 other tools, genetics / omics, mouse, 2 references
- [9] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: BEDTools, anndata, Scanpy, 8 other tools, genetics / omics, mouse, cellular / molecular, 2 references
- [10] doi:10.1038/s42003-026-10462-y [code]
- SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.Journal: Communications biologyIn common: BEDTools, anndata, Scanpy, 7 other tools, genetics / omics, mouse, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 4 repositories of the authors' code, each at its verified commit and with its license, 72 scripts, and 9 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:2f20d4cf06b4730a…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
