MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data.
The 9 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Methods › Model Structure—Cross‐Attention Mechanism for Feature Fusion ↔ Script/model.py, lines 41–161 · score 0.71 · Layer normalization, cross attention, positional, query, linear, tensors
- [2] § Methods › The Overview of MethyAnno ↔ Script/model.py, lines 204–266 · score 0.54 · prototype subspace, cross attention, fuse, module, classification, model
- [3] § Methods › Composite Loss Function for Model Training ↔ Script/loss.py, lines 74–86 · score 0.53 · cross entropy, contrastive loss, vectors, subspace
- [4] § Methods › Model Structure—Cross‐Attention Mechanism for Feature Fusion ↔ Script/model.py, lines 41–161 · score 0.51 · FFN, residual, layer, query, linear, tensor
- [5] § Results › The Overview of MethyAnno ↔ Script/novel_discover.py, lines 116–215 · score 0.51 · Gaussian mixture model, GMM, component, density, threshold, discovery
- [6] § Results › The Overview of MethyAnno ↔ scMethyAnno/novel_discover.py, lines 117–216 · score 0.51 · Gaussian mixture model, GMM, component, density, threshold, discovery
- [7] § Methods › Model Structure—Identification of Novel Cell Types ↔ Script/novel_discover.py, lines 116–215 · score 0.50 · Gaussian mixture model, component, density, threshold, scores
- [8] § Methods › Model Structure—Identification of Novel Cell Types ↔ scMethyAnno/novel_discover.py, lines 117–216 · score 0.50 · Gaussian mixture model, component, density, threshold, scores
- [9] § Methods › Composite Loss Function for Model Training ↔ Script/config.py, the whole file · a weak match · score 0.50 · hyperparameter configurations, weights, prototypes, loss, subspace, training
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 266 lines · 10 KB · MIT · 3 matches
- # --- Python 标准库 ---
- import os
- import gc
- import random
- import warnings
- # --- 第三方核心科学计算库 ---
- import numpy as np
- import pandas as pd
- import scipy.stats # 只导入需要的子模块
- import torch
- import torch.nn as nn
- import torch.nn.functional as F
- # --- 脚本级别的设置 ---
- warnings.filterwarnings("ignore")
- gc.collect()
- def setup_seed(seed):
- """
- Set random seed.
- Parameters
- ----------
- seed
- Number to be set as random seed for reproducibility.
- """
- np.random.seed(seed)
- random.seed(seed)
- torch.manual_seed(seed)
- if torch.cuda.is_available():
- torch.cuda.manual_seed_all(seed)
- setup_seed(123)
- class CrossAttention(nn.Module):
- def __init__(self, query_dim, context_dim, dropout_rate, embedding_dim = 64, head_dim=64, num_heads=1):
- """
- 改进的交叉注意力模块,将特征维度视为序列元素。
- Args:
- query_feature_dim (int): 查询输入张量的原始特征维度 (例如,feat1的维度)
- context_feature_dim (int): 上下文输入张量的原始特征维度 (例如,feat2的维度)
- embedding_dim (int): 每个特征元素将被映射到的嵌入维度
- dropout_rate (float): Dropout比率
- head_dim (int): 每个注意力头的维度
- num_heads (int): 注意力头的数量
- """
- super().__init__()
- self.query_feature_dim = query_dim
- self.context_feature_dim = context_dim
- self.embedding_dim = embedding_dim
- self.num_heads = num_heads
- self.head_dim = head_dim
- self.scale = head_dim ** -0.5
- inner_dim = num_heads * head_dim # 每个头的维度 * 头数量
- # 线性层用于将 (batch_size, feature_dim) -> (batch_size, feature_dim, embedding_dim)
- self.query_to_embedding = nn.Linear(1, embedding_dim) # 每个独立的特征值映射到embedding
- self.context_to_embedding = nn.Linear(1, embedding_dim)
- # 用于生成Q, K, V的线性层
- self.to_q = nn.Linear(embedding_dim, inner_dim, bias=False)
- self.to_k = nn.Linear(embedding_dim, inner_dim, bias=False)
- self.to_v = nn.Linear(embedding_dim, inner_dim, bias=False)
- # 最终输出的线性层
- self.to_out = nn.Sequential(
- nn.Linear(inner_dim, embedding_dim), # 从合并的头维度映射回embedding_dim
- nn.Dropout(dropout_rate)
- )
- self.output_to_original_dim = nn.Linear(embedding_dim, 1)
- self.pos_enc_query = nn.Parameter(torch.randn(query_dim, embedding_dim))
- self.pos_enc_context = nn.Parameter(torch.randn(context_dim, embedding_dim))
- self.norm_query = nn.LayerNorm(embedding_dim)
- self.norm_context = nn.LayerNorm(embedding_dim)
- self.ffn = nn.Sequential(
- nn.LayerNorm(embedding_dim),
- nn.Linear(embedding_dim, embedding_dim * 4),
- nn.GELU(),
- nn.Dropout(dropout_rate),
- nn.Linear(embedding_dim * 4, embedding_dim)
- )
- def forward(self, query, context):
- """
- Args:
- query (torch.Tensor): 查询张量,形状为 [B, query_feature_dim]
- context (torch.Tensor): 上下文张量,形状为 [B, context_feature_dim]
- Returns:
- torch.Tensor: 更新后的查询张量,形状为 [B, query_feature_dim] (融合后,维度与原始query一致)
- """
- # --- 1. 将输入特征转换为序列形式并应用位置编码 ---
- # query: [B, query_feature_dim] -> [B, query_feature_dim, 1]
- query_expanded = query.unsqueeze(-1)
- # [B, query_feature_dim, 1] -> [B, query_feature_dim, embedding_dim]
- query_seq = self.query_to_embedding(query_expanded) + self.pos_enc_query
- # context: [B, context_feature_dim] -> [B, context_feature_dim, 1]
- context_expanded = context.unsqueeze(-1)
- # [B, context_feature_dim, 1] -> [B, context_feature_dim, embedding_dim]
- context_seq = self.context_to_embedding(context_expanded) + self.pos_enc_context
- residual_query_seq = query_seq
- # --- 2. 归一化输入序列 ---
- query_seq = self.norm_query(query_seq)
- context_seq = self.norm_context(context_seq)
- # --- 3. 计算 Q, K, V ---
- # Q: [B, query_feature_dim, embedding_dim] -> [B, query_feature_dim, inner_dim]
- # K, V: [B, context_feature_dim, embedding_dim] -> [B, context_feature_dim, inner_dim]
- q = self.to_q(query_seq)
- k = self.to_k(context_seq)
- v = self.to_v(context_seq)
- # [B, seq_len, inner_dim] -> [B, num_heads, seq_len, head_dim]
- q = q.view(q.shape[0], q.shape[1], self.num_heads, self.head_dim).transpose(1, 2)
- k = k.view(k.shape[0], k.shape[1], self.num_heads, self.head_dim).transpose(1, 2)
- v = v.view(v.shape[0], v.shape[1], self.num_heads, self.head_dim).transpose(1, 2)
- # --- 4. 计算注意力分数 ---
- # attention_scores: [B, num_heads, query_seq_len, context_seq_len]
- attention_scores = torch.matmul(q, k.transpose(-1, -2)) * self.scale
- attention_probs = attention_scores.softmax(dim=-1)
- # attention_output: [B, num_heads, query_seq_len, head_dim]
- attention_output = torch.matmul(attention_probs, v)
- # --- 5. 合并多头的结果 ---
- # [B, num_heads, query_seq_len, head_dim] -> [B, query_seq_len, num_heads, head_dim]
- attention_output = attention_output.transpose(1, 2).contiguous()
- # [B, query_seq_len, num_heads, head_dim] -> [B, query_seq_len, inner_dim]
- attention_output = attention_output.view(attention_output.shape[0], attention_output.shape[1], -1)
- # [B, query_seq_len, inner_dim] -> [B, query_seq_len, embedding_dim]
- attention_output = self.to_out(attention_output)
- # --- 6. 残差连接和FFN ---
- x = residual_query_seq + attention_output # [B, query_feature_dim, embedding_dim]
- x = x + self.ffn(x) # [B, query_feature_dim, embedding_dim]
- # --- 7. 将融合后的嵌入序列转换回原始特征维度 ---
- # [B, query_feature_dim, embedding_dim] -> [B, query_feature_dim, 1]
- final_output = self.output_to_original_dim(x)
- # [B, query_feature_dim, 1] -> [B, query_feature_dim]
- final_output = final_output.squeeze(-1)
- return final_output
- def projection_distance(x, subspace):
- # x: [B, H], subspace: [D, H]
- # 投影矩阵:P = U.T @ U
- # proj_x = x @ P.T
- P = subspace.T @ subspace # [H, H]
- x_proj = x @ P # [B, H]
- dist = F.mse_loss(x_proj, x, reduction='none').sum(dim=-1) # 投影误差
- return dist
- class PrototypeSubspace(nn.Module):
- def __init__(self, num_classes, subspace_dim, hidden_dim):
- super().__init__()
- self.num_classes = num_classes
- self.subspace_dim = subspace_dim
- self.hidden_dim = hidden_dim
- # 类别的子空间表示:[num_classes, subspace_dim, hidden_dim]
- self.subspaces = nn.Parameter(torch.randn(num_classes, subspace_dim, hidden_dim))
- nn.init.orthogonal_(self.subspaces.view(-1, hidden_dim)) # 初始化每个子空间为正交向量组
- def forward(self, x):
- """
- 输入:
- x: [B, hidden_dim] --- 样本嵌入表示
- 输出:
- logits: [B, num_classes] --- 类别相似性得分
- """
- B = x.size(0)
- subspaces = self.subspaces # [C, D, H]
- # 投影距离:计算 fused_feat 到每类原型子空间的投影误差
- logits = []
- for c in range(self.num_classes):
- dist = projection_distance(x, subspaces[c])
- logits.append(-dist)
- logits = torch.stack(logits, dim=1)
- return logits
- class MultiScalePrototypeModel(nn.Module):
- def __init__(self, input_dims, num_classes, hidden_dim, dropout_rate, subspace_dim):
- super().__init__()
- self.encoder1 = nn.Sequential(
- nn.Linear(input_dims[0], hidden_dim*4),
- nn.BatchNorm1d(hidden_dim*4),
- nn.GELU(),
- nn.Dropout(dropout_rate),
- nn.Linear(hidden_dim*4, hidden_dim*2),
- nn.BatchNorm1d(hidden_dim*2),
- nn.GELU(),
- nn.Dropout(dropout_rate),
- nn.Linear(hidden_dim*2, hidden_dim),
- nn.GELU()
- )
- self.encoder2 = nn.Sequential(
- nn.Linear(input_dims[1], hidden_dim*2),
- nn.BatchNorm1d(hidden_dim*2),
- nn.GELU(),
- nn.Dropout(dropout_rate),
- nn.Linear(hidden_dim*2, hidden_dim),
- nn.GELU()
- )
- self.projection1 = nn.Sequential(
- nn.Linear(hidden_dim, hidden_dim//2),
- nn.GELU(),
- nn.Linear(hidden_dim//2, hidden_dim//2)
- )
- self.projection2 = nn.Sequential(
- nn.Linear(hidden_dim, hidden_dim//2),
- nn.GELU(),
- nn.Linear(hidden_dim//2, hidden_dim//2)
- )
- # 1. 定义两个交叉注意力模块
- # 一个用于 feat1 查询 feat2
- self.cross_attn_1_to_2 = CrossAttention(query_dim=hidden_dim, context_dim=hidden_dim, dropout_rate=dropout_rate)
- # 另一个用于 feat2 查询 feat1
- self.cross_attn_2_to_1 = CrossAttention(query_dim=hidden_dim, context_dim=hidden_dim, dropout_rate=dropout_rate)
- self.classifier = PrototypeSubspace(num_classes=num_classes,
- subspace_dim=subspace_dim,
- hidden_dim=hidden_dim * 2)
- def forward(self, x1, x2):
- feat1 = self.encoder1(x1)
- feat2 = self.encoder2(x2)
- z1 = F.normalize(self.projection1(feat1), dim=-1)
- z2 = F.normalize(self.projection2(feat2), dim=-1)
- fused_feat1 = self.cross_attn_1_to_2(query=feat1, context=feat2)
- fused_feat2 = self.cross_attn_2_to_1(query=feat2, context=feat1)
- fused_intermediate = torch.cat([fused_feat1, fused_feat2], dim=-1)
- logits = self.classifier(fused_intermediate)
- return z1, z2, logits, fused_intermediate
model.py at commit b98fde4, under MIT · at the source
Overview
- School of Mathematical Sciences and LPMC, Nankai University, Tianjin, China
- College of Life Sciences, Nankai University, Tianjin, China
- Academy for Advanced Interdisciplinary Studies, Nankai University, Tianjin, China
Abstract
Single‐cell DNA methylation (scDNAm) sequencing provides unique insights into epigenetic heterogeneity and cell‐specific regulatory landscapes. However, accurate cell type annotation for scDNAm data remains challenging, as the distinct data distribution of scDNAm hinders the adaptation of annotation methods from other omics, and specialized annotation tools for scDNAm are currently lacking. Here, MethyAnno is proposed as an interpretable deep metric learning framework that leverages multi‐scale information for accurate cell type annotation of scDNAm data. Additionally, MethyAnno enables generalized category discovery in open‐set scenarios by utilizing density‐based clustering to automatically estimate the number of novel cell types, while simultaneously deciphering cell‐type‐specific epigenetic signatures for biological interpretability. Extensive experiments demonstrate that MethyAnno excels in cross‐dataset annotation and novel type discovery, showing exceptional robustness in few‐shot scenarios for rare cell types. Moreover, interpretability analysis in the human brain dataset correctly recovers the genetic link between Sst interneurons and epilepsy heritability, the association of OPC cells with Alzheimer's disease, as well as the regulatory role of Pvalb cells in synaptic plasticity. Taken together, these findings establish MethyAnno as a robust and biologically interpretable tool for accurate cell type annotation and downstream epigenetic analysis.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.
BioX-NKU/MethyAnno
b98fde428f459286e51b11a29fd8f0612d292d47, 16 August 2026Availability: 1 check, the latest on 26 September 2026: the link answers
- 26 September 2026: the link answers
22 files
- Script/
config.py — Python, 49 lines, 1 match - Script/
data_preprocess.py — Python, 249 lines - Script/
interpretation.py — Python, 185 lines - Script/
loss.py — Python, 121 lines, 1 match - Script/
main.py — Python, 264 lines - Script/
model.py — Python, 266 lines, 3 matches - Script/
novel_discover.py — Python, 254 lines, 2 matches - Script/
taining_and_testing_scri — Python, 34 linespt/ inter_data_prediction_ru n.py - Script/
taining_and_testing_scri — Python, 42 linespt/ novel_type_identificatio n_run.py - Script/
train_test.py — Python, 202 lines - scMethyAnno/
__init__.py — Python, 5 lines - scMethyAnno/
config.py — Python, 49 lines - scMethyAnno/
data_preprocess.py — Python, 250 lines - scMethyAnno/
interpretation.py — Python, 184 lines - scMethyAnno/
loss.py — Python, 121 lines - scMethyAnno/
main.py — Python, 266 lines - scMethyAnno/
model.py — Python, 266 lines - scMethyAnno/
novel_discover.py — Python, 255 lines, 2 matches - scMethyAnno/
train_test.py — Python, 202 lines - setup.py — Python, 3 lines
- LICENSE — License, 21 lines
- README.md — Text, 181 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 20 scripts, each with its path and the digest of its content;
- 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE215353 — at NCBI GEO; found in “Data Availability Statement”
Data Availability Statement
All datasets used in this study are publicly available from the National Center for Biotechnology Information (NCBI) Gene Expression Omnibus (GEO). For the human cerebral cortex datasets, the corresponding accession number is GSE215353 (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, pages, dates, 5 authors, 5 keywords, 3 funders, 48 references.
Cite
This paper
Jia, Y., Li, S., Tang, S., Gu, K., & Chen, S. (2026). MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data. Advanced science (Weinheim, Baden-Wurttemberg, Germany), e77524. https://
BibTeX
@article{jia2026methyann
author = {Jia, Yuhang and Li, Siyu and Tang, Songming and Gu, Keju and Chen, Shengquan},
title = {{MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data}},
journal = {Advanced science (Weinheim, Baden-Wurttemberg, Germany)},
year = {2026},
month = sep,
pages = {e77524},
publisher = {Wiley},
issn = {2198-3844},
doi = {10.1002/
url = {https://
pmid = {42681808},
pmcid = {PMC13534786}
}
RIS
TY - JOUR
AU - Jia, Yuhang
AU - Li, Siyu
AU - Tang, Songming
AU - Gu, Keju
AU - Chen, Shengquan
TI - MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data
T2 - Advanced science (Weinheim, Baden-Wurttemberg, Germany)
J2 - Adv Sci (Weinh)
PY - 2026
DA - 2026/
SP - e77524
SN - 2198-3844
PB - Wiley
DO - 10.1002/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1002/
"type": "article-journal",
"title": "MethyAnno: An Interpretable Automated Annotation Method Leveraging Multi-Scale Information and Metric Learning Framework for scDNAm Data",
"container-title": "Advanced science (Weinheim, Baden-Wurttemberg, Germany)",
"author": [
{
"family": "Jia",
"given": "Yuhang"
},
{
"family": "Li",
"given": "Siyu"
},
{
"family": "Tang",
"given": "Songming"
},
{
"family": "Gu",
"given": "Keju"
},
{
"family": "Chen",
"given": "Shengquan"
}
],
"container-title-short":
"page": "e77524",
"DOI": "10.1002/
"PMID": "42681808",
"PMCID": "PMC13534786",
"ISSN": "2198-3844",
"publisher": "Wiley",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-73171-4 [code]
- Dissecting epigenetic heterogeneity in single-cell DNA methylomes with a unified framework.Journal: Nature communicationsIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 10 references, author Shengquan Chen
- [2] doi:10.1101/gr.281350.125 [code]
- High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND.Journal: Genome researchIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 6 references
- [3] doi:10.1093/bioinformatics/btag652 [code]
- mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.Journal: Bioinformatics (Oxford, England)In common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 3 references
- [4] doi:10.1038/s41467-026-71331-0 [code]
- A multimodal approach for visualizing and identifying electrophysiological cell types in vivo.Journal: Nature communicationsIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 2 references
- [5] doi:10.1016/j.xgen.2026.101217 [code]
- ProtoCloud: A prototypical self-explaining model for single-cell analysis.Journal: Cell genomicsIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 1 reference
- [6] doi:10.1038/s42003-026-10462-y [code]
- SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.Journal: Communications biologyIn common: anndata, Scanpy, PyTorch, 5 other tools, genetics / omics, 2 references
- [7] doi:10.1016/j.isci.2026.117206 [code]
- ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.Journal: iScienceIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 1 reference
- [8] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 1 reference
- [9] doi:10.1093/bioinformatics/btag540 [code]
- Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.Journal: Bioinformatics (Oxford, England)In common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 1 reference
- [10] doi:10.21203/rs.3.rs-9676637/v1 [code]
- A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic TechnologiesJournal: Research Square (preprint)In common: anndata, Scanpy, PyTorch, 6 other tools, genetics / omics, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 20 scripts, and 9 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e87ea03bd170fa23…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
