OSCR

Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstruction.

Code ↔ Paper

4 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 4 matches
  1. [1] § Materials and methods › Feature fusion and attention mechanism ↔ model/AE-Trans.py, lines 189–321 · score 0.73 · positional encoding, ReLU, layer normalization, Transformer encoder, fused, linear
  2. [2] § Materials and methods › Training and testing ↔ model/AE-Trans.py, lines 581–620 · score 0.65 · learning rate scheduler, ReduceLROnPlateau, Adam, L2, optimizer, loss
  3. [3] § Materials and methods › Feature fusion and attention mechanism ↔ model/AE-Trans.py, lines 189–321 · score 0.60 · Transformer encoder layer, positional encoding, Layer Normalization, fusion, model
  4. [4] § Materials and methods › Training and testing ↔ model/AE-Trans.py, lines 496–535 · score 0.56 · reconstruction losses, classification loss, BCE, MSE, split, fold

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 842 lines · 35 KB · Apache-2.0 · 4 matches

  1. # -*- coding: utf-8 -*-
  2. """
  3. Created on Tue Oct 15 12:02:16 2024
  4. @author: Administrator
  5. """
  6. import pandas as pd
  7. import numpy as np
  8. import torch
  9. from sklearn.model_selection import train_test_split
  10. from sklearn.preprocessing import StandardScaler
  11. import torch.nn as nn
  12. import math
  13. import torch.optim as optim
  14. from sklearn.model_selection import KFold
  15. from sklearn.metrics import precision_score, recall_score, f1_score, roc_auc_score, accuracy_score
  16. import torch.optim.lr_scheduler as lr_scheduler
  17. from torch.autograd import Variable
  18. import argparse
  19. # 定义位置编码类
  20. class PositionalEncoding(nn.Module):
  21. def __init__(self, d_model, dropout=0.1, max_len=5000):
  22. super(PositionalEncoding, self).__init__()
  23. self.dropout = nn.Dropout(p=dropout)
  24. pe = torch.zeros(max_len, d_model)
  25. position = torch.arange(0, max_len, dtype=torch.float).unsqueeze(1)
  26. div_term = torch.exp(torch.arange(0, d_model, 2).float() * (-math.log(10000.0) / d_model))
  27. pe[:, 0::2] = torch.sin(position * div_term)
  28. pe[:, 1::2] = torch.cos(position * div_term)
  29. pe = pe.unsqueeze(0).transpose(0, 1)
  30. self.register_buffer('pe', pe)
  31. def forward(self, x):
  32. x = x + self.pe[:x.size(0), :]
  33. return self.dropout(x)
  34. # 定义多头注意力机制
  35. class MultiheadAttention(nn.Module):
  36. def __init__(self, embed_dim, num_heads, dropout=0.1):
  37. super(MultiheadAttention, self).__init__()
  38. self.embed_dim = embed_dim
  39. self.num_heads = num_heads
  40. self.head_dim = embed_dim // num_heads
  41. assert self.head_dim * num_heads == self.embed_dim, "embed_dim must be divisible by num_heads"
  42. self.q_proj = nn.Linear(embed_dim, embed_dim)
  43. self.k_proj = nn.Linear(embed_dim, embed_dim)
  44. self.v_proj = nn.Linear(embed_dim, embed_dim)
  45. self.out_proj = nn.Linear(embed_dim, embed_dim)
  46. self.dropout = nn.Dropout(dropout)
  47. def scaled_dot_product_attention(self, q, k, v, mask=None):
  48. d_k = q.size(-1)
  49. scores = torch.matmul(q, k.transpose(-2, -1)) / math.sqrt(d_k)
  50. if mask is not None:
  51. scores = scores.masked_fill(mask == 0, float('-inf'))
  52. attn = torch.softmax(scores, dim=-1)
  53. attn = self.dropout(attn)
  54. output = torch.matmul(attn, v)
  55. return output, attn
  56. def split_heads(self, x, batch_size):
  57. return x.view(batch_size, -1, self.num_heads, self.head_dim).transpose(1, 2)
  58. def forward(self, query, key, value, mask=None):
  59. batch_size = query.size(0)
  60. q = self.split_heads(self.q_proj(query), batch_size)
  61. k = self.split_heads(self.k_proj(key), batch_size)
  62. v = self.split_heads(self.v_proj(value), batch_size)
  63. attn_output, _ = self.scaled_dot_product_attention(q, k, v, mask)
  64. attn_output = attn_output.transpose(1, 2).contiguous().view(batch_size, -1, self.embed_dim)
  65. output = self.out_proj(attn_output)
  66. return output
  67. # 定义Transformer编码器层
  68. class MyTransformerEncoderLayer(nn.Module):
  69. def __init__(self, d_model, nhead, dim_feedforward=2048, dropout=0.1):
  70. super(MyTransformerEncoderLayer, self).__init__()
  71. self.self_attn = MultiheadAttention(d_model, nhead, dropout=dropout)
  72. self.linear1 = nn.Linear(d_model, dim_feedforward)
  73. self.dropout = nn.Dropout(dropout)
  74. self.linear2 = nn.Linear(dim_feedforward, d_model)
  75. self.norm1 = nn.LayerNorm(d_model)
  76. self.norm2 = nn.LayerNorm(d_model)
  77. self.dropout1 = nn.Dropout(dropout)
  78. self.dropout2 = nn.Dropout(dropout)
  79. self.activation = nn.ReLU()
  80. def forward(self, src, src_mask=None, src_key_padding_mask=None):
  81. src2 = self.self_attn(src, src, src, mask=src_mask)
  82. src = src + self.dropout1(src2)
  83. src = self.norm1(src)
  84. src2 = self.linear2(self.dropout(self.activation(self.linear1(src))))
  85. src = src + self.dropout2(src2)
  86. src = self.norm2(src)
  87. return src
  88. # 定义Transformer编码器
  89. class MyTransformerEncoder(nn.Module):
  90. def __init__(self, encoder_layer, num_layers, norm=None):
  91. super(MyTransformerEncoder, self).__init__()
  92. self.layers = nn.ModuleList([encoder_layer for _ in range(num_layers)])
  93. self.norm = norm
  94. def forward(self, src, mask=None, src_key_padding_mask=None):
  95. output = src
  96. for mod in self.layers:
  97. output = mod(output, src_mask=mask, src_key_padding_mask=src_key_padding_mask)
  98. if self.norm is not None:
  99. output = self.norm(output)
  100. return output
  101. # 定义Transformer解码器层
  102. class MyTransformerDecoderLayer(nn.Module):
  103. def __init__(self, d_model, nhead, dim_feedforward=256, dropout=0.1):
  104. super(MyTransformerDecoderLayer, self).__init__()
  105. self.self_attn = MultiheadAttention(d_model, nhead, dropout=dropout)
  106. self.multihead_attn = MultiheadAttention(d_model, nhead, dropout=dropout)
  107. self.linear1 = nn.Linear(d_model, dim_feedforward)
  108. self.dropout = nn.Dropout(dropout)
  109. self.linear2 = nn.Linear(dim_feedforward, d_model)
  110. self.norm1 = nn.LayerNorm(d_model)
  111. self.norm2 = nn.LayerNorm(d_model)
  112. self.norm3 = nn.LayerNorm(d_model)
  113. self.dropout1 = nn.Dropout(dropout)
  114. self.dropout2 = nn.Dropout(dropout)
  115. self.dropout3 = nn.Dropout(dropout)
  116. self.activation = nn.ReLU()
  117. def forward(self, tgt, memory, tgt_mask=None, memory_mask=None, tgt_key_padding_mask=None,
  118. memory_key_padding_mask=None):
  119. tgt2 = self.self_attn(tgt, tgt, tgt, mask=tgt_mask)
  120. tgt = tgt + self.dropout1(tgt2)
  121. tgt = self.norm1(tgt)
  122. tgt2 = self.multihead_attn(tgt, memory, memory, mask=memory_mask)
  123. tgt = tgt + self.dropout2(tgt2)
  124. tgt = self.norm2(tgt)
  125. tgt2 = self.linear2(self.dropout(self.activation(self.linear1(tgt))))
  126. tgt = tgt + self.dropout3(tgt2)
  127. tgt = self.norm3(tgt)
  128. return tgt
  129. # 定义Transformer解码器
  130. class MyTransformerDecoder(nn.Module):
  131. def __init__(self, decoder_layer, num_layers, norm=None):
  132. super(MyTransformerDecoder, self).__init__()
  133. self.layers = nn.ModuleList([decoder_layer for _ in range(num_layers)])
  134. self.norm = norm
  135. def forward(self, tgt, memory, tgt_mask=None, memory_mask=None, tgt_key_padding_mask=None,
  136. memory_key_padding_mask=None):
  137. output = tgt
  138. for mod in self.layers:
  139. output = mod(output, memory, tgt_mask=tgt_mask, memory_mask=memory_mask,
  140. tgt_key_padding_mask=tgt_key_padding_mask,
  141. memory_key_padding_mask=memory_key_padding_mask)
  142. if self.norm is not None:
  143. output = self.norm(output)
  144. return output
  145. # 定义随机掩码模块
  146. class FeatureMasking(nn.Module):
  147. def __init__(self, mask_num):
  148. super(FeatureMasking, self).__init__()
  149. self.mask_num = mask_num
  150. def forward(self, x):
  151. if self.mask_num > 0:
  152. x_clone = x.clone() # 克隆张量,避免原地操作
  153. batch_size, feature_dim = x_clone.size() # 处理二维张量
  154. mask_indices = torch.randperm(feature_dim)[:self.mask_num]
  155. x_clone[:, mask_indices] = 0
  156. return x_clone
  157. return x
  158. # 定义自动编码器模型,将掩码操作移到AE编码器之前,并增加两个并行的AE编码器和解码器
  159. class AutoencoderWithTransformer(nn.Module):
  160. def __init__(self, input_dim, encoding_dim, d_model=128, nhead=8, num_encoder_layers=3,
  161. num_decoder_layers=3, dim_feedforward=256, dropout=0.1, mask_num=0):
  162. super(AutoencoderWithTransformer, self).__init__()
  163. # 定义随机掩码模块
  164. self.masking = FeatureMasking(mask_num) # 掩码操作移到这里
  165. # 根据 input_dim 计算RNA和甲基化数据的拆分点
  166. self.rna_split_point = input_dim // 4 # RNA的前一半
  167. self.methylation_split_point = input_dim // 4 # 甲基化的前一半
  168. self.first_part_dim = self.rna_split_point + self.methylation_split_point
  169. self.second_part_dim = input_dim - self.first_part_dim # 第二部分为剩余的维度
  170. # 自动编码器的编码器部分(第一部分:RNA前一部分 + 甲基化前一部分)
  171. self.encoder1 = nn.Sequential(
  172. nn.Linear(self.first_part_dim, 4096),
  173. nn.ReLU(),
  174. nn.Linear(4096, 512),
  175. nn.ReLU(),
  176. nn.Linear(512, encoding_dim),
  177. nn.ReLU()
  178. )
  179. # 自动编码器的编码器部分(第二部分:RNA剩余部分 + 甲基化剩余部分)
  180. self.encoder2 = nn.Sequential(
  181. nn.Linear(self.second_part_dim, 4096),
  182. nn.ReLU(),
  183. nn.Linear(4096, 512),
  184. nn.ReLU(),
  185. nn.Linear(512, encoding_dim),
  186. nn.ReLU()
  187. )
  188. # 位置编码
  189. self.pos_embedding = PositionalEncoding(d_model=d_model, dropout=dropout)
  190. # Transformer部分(第一部分)
  191. transformer_encoder_layer1 = MyTransformerEncoderLayer(d_model, nhead, dim_feedforward, dropout)
  192. transformer_encoder_norm1 = nn.LayerNorm(d_model)
  193. self.transformer_encoder1 = MyTransformerEncoder(transformer_encoder_layer1, num_encoder_layers, transformer_encoder_norm1)
  194. # Transformer部分(第二部分)
  195. transformer_encoder_layer2 = MyTransformerEncoderLayer(d_model, nhead, dim_feedforward, dropout)
  196. transformer_encoder_norm2 = nn.LayerNorm(d_model)
  197. self.transformer_encoder2 = MyTransformerEncoder(transformer_encoder_layer2, num_encoder_layers, transformer_encoder_norm2)
  198. # 线性层用于降维到 d_model
  199. self.fusion_to_d_model = nn.Linear(encoding_dim * 2, d_model)
  200. # Transformer解码器部分(第一部分)
  201. transformer_decoder_layer1 = MyTransformerDecoderLayer(d_model, nhead, dim_feedforward, dropout)
  202. transformer_decoder_norm1 = nn.LayerNorm(d_model)
  203. self.transformer_decoder1 = MyTransformerDecoder(transformer_decoder_layer1, num_decoder_layers, transformer_decoder_norm1)
  204. # Transformer解码器部分(第二部分)
  205. transformer_decoder_layer2 = MyTransformerDecoderLayer(d_model, nhead, dim_feedforward, dropout)
  206. transformer_decoder_norm2 = nn.LayerNorm(d_model)
  207. self.transformer_decoder2 = MyTransformerDecoder(transformer_decoder_layer2, num_decoder_layers, transformer_decoder_norm2)
  208. # 自动编码器的解码器部分(第一部分)
  209. self.decoder1 = nn.Sequential(
  210. nn.Linear(encoding_dim, 512),
  211. nn.ReLU(),
  212. nn.Linear(512, 4096),
  213. nn.ReLU(),
  214. nn.Linear(4096, self.first_part_dim),
  215. nn.Sigmoid()
  216. )
  217. # 自动编码器的解码器部分(第二部分)
  218. self.decoder2 = nn.Sequential(
  219. nn.Linear(encoding_dim, 512),
  220. nn.ReLU(),
  221. nn.Linear(512, 4096),
  222. nn.ReLU(),
  223. nn.Linear(4096, self.second_part_dim),
  224. nn.Sigmoid()
  225. )
  226. # 分类器部分
  227. self.classifier = nn.Sequential(
  228. nn.Linear(encoding_dim, 32),
  229. nn.ReLU(),
  230. nn.Dropout(dropout),
  231. nn.Linear(32, 8),
  232. nn.ReLU(),
  233. nn.Dropout(dropout),
  234. nn.Linear(8, 1),
  235. nn.Sigmoid()
  236. )
  237. def forward(self, x):
  238. # 掩码操作在AE编码器之前
  239. input_dim = x.shape[1]
  240. x = self.masking(x)
  241. # RNA的前一半和甲基化的前一半
  242. x1 = torch.cat((x[:, :self.rna_split_point], x[:, input_dim // 2:input_dim // 2 + self.methylation_split_point]), dim=1)
  243. # RNA的剩余部分和甲基化的剩余部分
  244. x2 = torch.cat((x[:, self.rna_split_point:input_dim // 2], x[:, input_dim // 2 + self.methylation_split_point:]), dim=1)
  245. # AE编码器部分
  246. encoded1 = self.encoder1(x1)
  247. encoded2 = self.encoder2(x2)
  248. # Transformer部分
  249. encoded1 = self.pos_embedding(encoded1.unsqueeze(0)) # 添加伪序列维度
  250. encoded2 = self.pos_embedding(encoded2.unsqueeze(0)) # 添加伪序列维度
  251. transformer_encoder1 = self.transformer_encoder1(encoded1).squeeze(0) # 移除伪序列维度
  252. transformer_encoder2 = self.transformer_encoder2(encoded2).squeeze(0) # 移除伪序列维度
  253. # 融合编码结果
  254. fused_encoding = torch.cat((transformer_encoder1, transformer_encoder2), dim=1)
  255. # 通过线性层降维到 d_model
  256. fused_encoding_reduced = self.fusion_to_d_model(fused_encoding)
  257. # Transformer解码部分,使用降维后的向量作为输入
  258. transformer_decoder1 = self.transformer_decoder1(fused_encoding_reduced.unsqueeze(0), encoded1).squeeze(0)
  259. transformer_decoder2 = self.transformer_decoder2(fused_encoding_reduced.unsqueeze(0), encoded2).squeeze(0)
  260. # AE解码器部分
  261. decoded1 = self.decoder1(transformer_decoder1)
  262. decoded2 = self.decoder2(transformer_decoder2)
  263. # 分类任务
  264. classification_output = self.classifier(fused_encoding_reduced)
  265. return classification_output, decoded1, decoded2
  266. # 积分梯度函数
  267. def integrated_gradients(model, x, baseline=None, steps=50):
  268. if baseline is None:
  269. baseline = torch.zeros_like(x) # 使用全0向量作为基线
  270. # 梯度累加
  271. scaled_inputs = [baseline + (float(i) / steps) * (x - baseline) for i in range(steps + 1)]
  272. classifier_grads = []
  273. for scaled_input in scaled_inputs:
  274. scaled_input = Variable(scaled_input, requires_grad=True)
  275. # 前向传播计算输出
  276. logits, _, _ = model(scaled_input)
  277. model.zero_grad()
  278. # 如果是标量输出(如二分类任务)
  279. logit_target = logits[0] # 只取第一个 logit 值
  280. logit_target.backward(retain_graph=True)
  281. classifier_grads.append(scaled_input.grad.cpu().detach().numpy()) # 确保分离梯度
  282. # 计算平均梯度
  283. avg_classifier_grads = np.average(classifier_grads[:-1], axis=0)
  284. integrated_classifier_grad = (x.cpu().detach().numpy() - baseline.cpu().detach().numpy()) * avg_classifier_grads
  285. return integrated_classifier_grad
  286. # 计算总的平均积分梯度
  287. def get_average_integrated_gradients(model, X_test, device):
  288. total_grads = None
  289. num_samples = len(X_test)
  290. for i in range(num_samples):
  291. x = torch.tensor(X_test[i:i+1].clone().detach(), requires_grad=True, dtype=torch.float32).to(device)
  292. # 计算样本的积分梯度
  293. classifier_grad = integrated_gradients(model, x)
  294. # 累加梯度
  295. if total_grads is None:
  296. total_grads = classifier_grad
  297. else:
  298. total_grads += classifier_grad
  299. # 计算平均梯度
  300. avg_grads = total_grads / num_samples
  301. return avg_grads
  302. # 对积分梯度排序,返回三个排序结果
  303. def get_sorted_features(avg_grads, selected_features, num):
  304. # 构建DataFrame,用于排序
  305. attributions = pd.DataFrame({
  306. 'feature': selected_features,
  307. 'classifier_attribution': np.abs(np.mean(avg_grads, axis=0)) # 计算每个特征的平均积分梯度贡献
  308. })
  309. # 对特征按贡献值进行降序排序
  310. sorted_attributions = attributions.sort_values(by='classifier_attribution', ascending=False)
  311. # 获取总的前200个特征
  312. top_all_features = sorted_attributions.head(num)
  313. # 从前200个特征中筛选出RNA特征(没有Metht_前缀的)
  314. top_rna_features = top_all_features[top_all_features['feature'].apply(lambda x: not x.startswith('Methy_'))]
  315. # 从前200个特征中筛选出甲基化特征(有Metht_前缀的)
  316. top_methylation_features = top_all_features[top_all_features['feature'].apply(lambda x: x.startswith('Methy_'))]
  317. return top_all_features, top_rna_features, top_methylation_features
  318. # 主函数
  319. def main(epochs, file_path_gene, file_path_methy):
  320. # 读取RNA和甲基化数据
  321. rna_file_path = file_path_gene
  322. methylation_file_path = file_path_methy
  323. rna_data = pd.read_csv(rna_file_path)
  324. methylation_data = pd.read_csv(methylation_file_path)
  325. # 提取特征和标签
  326. X_rna = rna_data.iloc[:, 2:].values # RNA数据的特征
  327. y_rna = rna_data.iloc[:, 1].values # RNA数据的标签
  328. X_methylation = methylation_data.iloc[:, 2:].values # 甲基化数据的特征
  329. y_methylation = methylation_data.iloc[:, 1].values # 甲基化数据的标签
  330. # 1. 划分RNA数据为正负样本
  331. X_rna_pos = X_rna[y_rna == 1]
  332. X_rna_neg = X_rna[y_rna == 0]
  333. # 2. 划分甲基化数据为正负样本
  334. X_methylation_pos = X_methylation[y_methylation == 1]
  335. X_methylation_neg = X_methylation[y_methylation == 0]
  336. # 3. 分别将RNA正样本、RNA负样本、甲基化正样本和甲基化负样本划分为训练集和测试集
  337. X_rna_pos_train, X_rna_pos_test = train_test_split(X_rna_pos, test_size=0.2, random_state=42)
  338. X_rna_neg_train, X_rna_neg_test = train_test_split(X_rna_neg, test_size=0.2, random_state=42)
  339. X_methylation_pos_train, X_methylation_pos_test = train_test_split(X_methylation_pos, test_size=0.2, random_state=42)
  340. X_methylation_neg_train, X_methylation_neg_test = train_test_split(X_methylation_neg, test_size=0.2, random_state=42)
  341. # 4. 正样本训练集融合:每个RNA正样本训练集拼接每个甲基化正样本训练集
  342. positive_train_samples = np.array([np.hstack((rna_sample, meth_sample))
  343. for rna_sample in X_rna_pos_train
  344. for meth_sample in X_methylation_pos_train])
  345. # 5. 负样本训练集融合:每个RNA负样本训练集拼接每个甲基化负样本训练集
  346. negative_train_samples = np.array([np.hstack((rna_sample, meth_sample))
  347. for rna_sample in X_rna_neg_train
  348. for meth_sample in X_methylation_neg_train])
  349. # 6. 正样本测试集融合:每个RNA正样本测试集拼接每个甲基化正样本测试集
  350. positive_test_samples = np.array([np.hstack((rna_sample, meth_sample))
  351. for rna_sample in X_rna_pos_test
  352. for meth_sample in X_methylation_pos_test])
  353. # 7. 负样本测试集融合:每个RNA负样本测试集拼接每个甲基化负样本测试集
  354. negative_test_samples = np.array([np.hstack((rna_sample, meth_sample))
  355. for rna_sample in X_rna_neg_test
  356. for meth_sample in X_methylation_neg_test])
  357. # 8. 标签:正样本集为1,负样本集为0
  358. y_positive_train = np.ones(len(positive_train_samples))
  359. y_negative_train = np.zeros(len(negative_train_samples))
  360. y_positive_test = np.ones(len(positive_test_samples))
  361. y_negative_test = np.zeros(len(negative_test_samples))
  362. # 9. 合并训练集正负样本
  363. X_train_combined = np.vstack((positive_train_samples, negative_train_samples))
  364. y_train_combined = np.concatenate((y_positive_train, y_negative_train))
  365. # 10. 合并测试集正负样本
  366. X_test_combined = np.vstack((positive_test_samples, negative_test_samples))
  367. y_test_combined = np.concatenate((y_positive_test, y_negative_test))
  368. # 11. 打乱训练集
  369. train_indices = np.random.permutation(len(X_train_combined))
  370. X_train_combined = X_train_combined[train_indices]
  371. y_train_combined = y_train_combined[train_indices]
  372. # 对特征进行标准化处理
  373. scaler = StandardScaler()
  374. X_train_combined = scaler.fit_transform(X_train_combined)
  375. # 12. 打乱测试集
  376. test_indices = np.random.permutation(len(X_test_combined))
  377. X_test_combined = X_test_combined[test_indices]
  378. y_test_combined = y_test_combined[test_indices]
  379. X_test_combined = scaler.fit_transform(X_test_combined)
  380. # 13. 将数据转换为PyTorch的张量格式
  381. X_train = torch.tensor(X_train_combined, dtype=torch.float32)
  382. X_test = torch.tensor(X_test_combined, dtype=torch.float32)
  383. y_test = torch.tensor(y_test_combined, dtype=torch.float32)
  384. # 设置设备
  385. device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
  386. # 定义模型参数
  387. input_dim = X_train.shape[1]
  388. encoding_dim = 128
  389. d_model = 128
  390. nhead = 8
  391. num_encoder_layers = 3
  392. num_decoder_layers = 3
  393. dim_feedforward = 256
  394. dropout = 0.1
  395. mask_num = 15000 # 设定掩码特征的数量
  396. batch_size = 64
  397. X_train_combined_all = X_train_combined
  398. y_train_combined_all = y_train_combined
  399. # 定义损失函数
  400. classification_loss_fn = nn.BCELoss()
  401. reconstruction_loss_fn = nn.MSELoss()
  402. # 存储每一折的评估指标和AUC曲线数据
  403. fold_accuracies = []
  404. fold_precisions = []
  405. fold_recalls = []
  406. fold_f1s = []
  407. fold_aucs = []
  408. kf = KFold(n_splits=5, shuffle=True, random_state=42)
  409. # 五折交叉验证训练
  410. for fold, (train_idx, val_idx) in enumerate(kf.split(X_rna_pos_train)):
  411. print(f"Starting Fold {fold + 1}")
  412. # RNA正样本训练集划分为训练和验证集
  413. X_rna_pos_train_fold, X_rna_pos_val_fold = X_rna_pos_train[train_idx], X_rna_pos_train[val_idx]
  414. # 计算验证集比例
  415. val_ratio = len(val_idx) / len(X_rna_pos_train)
  416. # RNA负样本训练集按比例划分训练和验证集
  417. X_rna_neg_train_fold, X_rna_neg_val_fold = train_test_split(X_rna_neg_train, test_size=val_ratio, random_state=fold)
  418. # 甲基化正样本训练集按比例划分训练和验证集
  419. X_methylation_pos_train_fold, X_methylation_pos_val_fold = train_test_split(X_methylation_pos_train, test_size=val_ratio, random_state=fold)
  420. # 甲基化负样本训练集按比例划分训练和验证集
  421. X_methylation_neg_train_fold, X_methylation_neg_val_fold = train_test_split(X_methylation_neg_train, test_size=val_ratio, random_state=fold)
  422. # 融合训练集和验证集
  423. positive_train_samples = np.array([np.hstack((rna_sample, meth_sample))
  424. for rna_sample in X_rna_pos_train_fold
  425. for meth_sample in X_methylation_pos_train_fold])
  426. negative_train_samples = np.array([np.hstack((rna_sample, meth_sample))
  427. for rna_sample in X_rna_neg_train_fold
  428. for meth_sample in X_methylation_neg_train_fold])
  429. positive_val_samples = np.array([np.hstack((rna_sample, meth_sample))
  430. for rna_sample in X_rna_pos_val_fold
  431. for meth_sample in X_methylation_pos_val_fold])
  432. negative_val_samples = np.array([np.hstack((rna_sample, meth_sample))
  433. for rna_sample in X_rna_neg_val_fold
  434. for meth_sample in X_methylation_neg_val_fold])
  435. # 合并正负样本
  436. X_train_combined = np.vstack((positive_train_samples, negative_train_samples))
  437. y_train_combined = np.concatenate((np.ones(len(positive_train_samples)), np.zeros(len(negative_train_samples))))
  438. X_val_combined = np.vstack((positive_val_samples, negative_val_samples))
  439. y_val_combined = np.concatenate((np.ones(len(positive_val_samples)), np.zeros(len(negative_val_samples))))
  440. # 打乱训练集和验证集
  441. train_indices = np.random.permutation(len(X_train_combined))
  442. X_train_combined = X_train_combined[train_indices]
  443. y_train_combined = y_train_combined[train_indices]
  444. val_indices = np.random.permutation(len(X_val_combined))
  445. X_val_combined = X_val_combined[val_indices]
  446. y_val_combined = y_val_combined[val_indices]
  447. # 将数据转换为 DataLoader 格式
  448. train_dataset = torch.utils.data.TensorDataset(torch.tensor(X_train_combined, dtype=torch.float32), torch.tensor(y_train_combined, dtype=torch.float32))
  449. train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=True)
  450. val_dataset = torch.utils.data.TensorDataset(torch.tensor(X_val_combined, dtype=torch.float32), torch.tensor(y_val_combined, dtype=torch.float32))
  451. val_loader = torch.utils.data.DataLoader(val_dataset, batch_size=batch_size, shuffle=False)
  452. # 实例化模型
  453. model = AutoencoderWithTransformer(input_dim=input_dim, encoding_dim=encoding_dim, d_model=d_model, nhead=nhead,
  454. num_encoder_layers=num_encoder_layers, num_decoder_layers=num_decoder_layers,
  455. dim_feedforward=dim_feedforward, dropout=dropout, mask_num=mask_num)
  456. model.to(device)
  457. # 定义优化器并添加L2正则化
  458. optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)
  459. # 引入学习率调度器
  460. scheduler = lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.1, patience=5, verbose=True)
  461. # 训练每一折
  462. for epoch in range(epochs):
  463. model.train()
  464. total_loss = 0.0
  465. total_classification_loss = 0.0
  466. total_reconstruction_loss1 = 0.0
  467. total_reconstruction_loss2 = 0.0
  468. for batch_X, batch_y in train_loader:
  469. batch_X, batch_y = batch_X.to(device), batch_y.to(device)
  470. optimizer.zero_grad()
  471. # 前向传播
  472. logits, reconstructed1, reconstructed2 = model(batch_X)
  473. # 计算损失
  474. classification_loss = classification_loss_fn(logits, batch_y.unsqueeze(1))
  475. reconstruction_loss1 = reconstruction_loss_fn(reconstructed1, batch_X[:, :model.first_part_dim])
  476. reconstruction_loss2 = reconstruction_loss_fn(reconstructed2, batch_X[:, model.first_part_dim:])
  477. # 总损失 = 分类损失 + 两个重构损失
  478. loss = classification_loss + reconstruction_loss1 + reconstruction_loss2
  479. # 反向传播和优化
  480. loss.backward()
  481. optimizer.step()
  482. total_loss += loss.item()
  483. total_classification_loss += classification_loss.item()
  484. total_reconstruction_loss1 += reconstruction_loss1.item()
  485. total_reconstruction_loss2 += reconstruction_loss2.item()
  486. avg_loss = total_loss / len(train_loader)
  487. avg_classification_loss = total_classification_loss / len(train_loader)
  488. avg_reconstruction_loss1 = total_reconstruction_loss1 / len(train_loader)
  489. avg_reconstruction_loss2 = total_reconstruction_loss2 / len(train_loader)
  490. print(f'Epoch [{epoch+1}/{epochs}], Fold [{fold+1}], Classification Loss: {avg_classification_loss:.4f}, '
  491. f'Reconstruction Loss 1: {avg_reconstruction_loss1:.4f}, Reconstruction Loss 2: {avg_reconstruction_loss2:.4f}, Loss: {avg_loss:.4f}')
  492. # 更新学习率调度器
  493. scheduler.step(avg_loss)
  494. # 验证阶段
  495. model.eval()
  496. y_true_val = []
  497. y_pred_val = []
  498. y_scores_val = []
  499. with torch.no_grad():
  500. for batch_X, batch_y in val_loader:
  501. batch_X, batch_y = batch_X.to(device), batch_y.to(device)
  502. logits, _, _ = model(batch_X)
  503. predicted = (logits > 0.5).float()
  504. y_true_val.extend(batch_y.cpu().numpy())
  505. y_pred_val.extend(predicted.cpu().numpy())
  506. y_scores_val.extend(logits.cpu().numpy())
  507. # 计算评估指标
  508. accuracy = accuracy_score(y_true_val, y_pred_val)
  509. precision = precision_score(y_true_val, y_pred_val)
  510. recall = recall_score(y_true_val, y_pred_val)
  511. f1 = f1_score(y_true_val, y_pred_val)
  512. auc_value = roc_auc_score(y_true_val, y_scores_val)
  513. fold_accuracies.append(accuracy)
  514. fold_precisions.append(precision)
  515. fold_recalls.append(recall)
  516. fold_f1s.append(f1)
  517. fold_aucs.append(auc_value)
  518. print(f'Fold {fold+1} - Accuracy: {accuracy:.4f}, Precision: {precision:.4f}, Recall: {recall:.4f}, '
  519. f'F1-score: {f1:.4f}, AUC: {auc_value:.4f}')
  520. # 计算五折交叉验证的平均评估指标
  521. avg_accuracy = np.mean(fold_accuracies)
  522. avg_precision = np.mean(fold_precisions)
  523. avg_recall = np.mean(fold_recalls)
  524. avg_f1 = np.mean(fold_f1s)
  525. avg_auc = np.mean(fold_aucs)
  526. cv_resultdata = {
  527. "Evaluation indicators": ["Accuracy", "Precision", "Recall", "F1-Score", "AUC"],
  528. "Value": [round(avg_accuracy, 4), round(avg_precision, 4), round(avg_recall, 4), round(avg_f1, 4), round(avg_auc, 4)]
  529. }
  530. cv_df = pd.DataFrame(cv_resultdata)
  531. # cv_df.to_csv("../result/cv_resultdata.csv", index=False)
  532. # 将数据转换为 DataLoader 格式
  533. train_dataset = torch.utils.data.TensorDataset(torch.tensor(X_train_combined_all, dtype=torch.float32), torch.tensor(y_train_combined_all, dtype=torch.float32))
  534. train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=True)
  535. # 实例化模型
  536. model = AutoencoderWithTransformer(input_dim=input_dim, encoding_dim=encoding_dim, d_model=d_model, nhead=nhead,
  537. num_encoder_layers=num_encoder_layers, num_decoder_layers=num_decoder_layers,
  538. dim_feedforward=dim_feedforward, dropout=dropout, mask_num=mask_num)
  539. model.to(device)
  540. # 定义优化器并添加L2正则化
  541. optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)
  542. # 引入学习率调度器
  543. scheduler = lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.1, patience=5, verbose=True)
  544. # 训练每一折
  545. for epoch in range(epochs):
  546. model.train()
  547. total_loss = 0.0
  548. total_classification_loss = 0.0
  549. total_reconstruction_loss1 = 0.0
  550. total_reconstruction_loss2 = 0.0
  551. for batch_X, batch_y in train_loader:
  552. batch_X, batch_y = batch_X.to(device), batch_y.to(device)
  553. optimizer.zero_grad()
  554. # 前向传播
  555. logits, reconstructed1, reconstructed2 = model(batch_X)
  556. # RNA的前一半和甲基化的前一半
  557. original_part1 = torch.cat((batch_X[:, :model.rna_split_point],
  558. batch_X[:, input_dim // 2:input_dim // 2 + model.methylation_split_point]), dim=1)
  559. # RNA的剩余部分和甲基化的剩余部分
  560. original_part2 = torch.cat((batch_X[:, model.rna_split_point:input_dim // 2],
  561. batch_X[:, input_dim // 2 + model.methylation_split_point:]), dim=1)
  562. # 计算损失
  563. classification_loss = classification_loss_fn(logits, batch_y.unsqueeze(1))
  564. reconstruction_loss1 = reconstruction_loss_fn(reconstructed1, original_part1)
  565. reconstruction_loss2 = reconstruction_loss_fn(reconstructed2, original_part2)
  566. # 总损失 = 分类损失 + 两个重构损失
  567. loss = classification_loss + reconstruction_loss1 + reconstruction_loss2
  568. # 反向传播和优化
  569. loss.backward()
  570. optimizer.step()
  571. total_loss += loss.item()
  572. total_classification_loss += classification_loss.item()
  573. total_reconstruction_loss1 += reconstruction_loss1.item()
  574. total_reconstruction_loss2 += reconstruction_loss2.item()
  575. avg_loss = total_loss / len(train_loader)
  576. avg_classification_loss = total_classification_loss / len(train_loader)
  577. avg_reconstruction_loss1 = total_reconstruction_loss1 / len(train_loader)
  578. avg_reconstruction_loss2 = total_reconstruction_loss2 / len(train_loader)
  579. print(f'Epoch [{epoch+1}/{epochs}], Classification Loss: {avg_classification_loss:.4f}, '
  580. f'Reconstruction Loss 1: {avg_reconstruction_loss1:.4f}, Reconstruction Loss 2: {avg_reconstruction_loss2:.4f}, Loss: {avg_loss:.4f}')
  581. # 测试集评估并计算AUC
  582. print("\nTesting on total test set:")
  583. model.eval()
  584. y_true = []
  585. y_pred = []
  586. y_scores = []
  587. with torch.no_grad():
  588. for i in range(len(X_test)):
  589. batch_X = X_test[i].unsqueeze(0).to(device) # 添加维度以匹配批次大小
  590. batch_y = y_test[i].unsqueeze(0).to(device) # 添加维度以匹配批次大小
  591. logits, _, _ = model(batch_X)
  592. predicted = (logits > 0.5).float()
  593. y_true.append(batch_y.item())
  594. y_pred.append(predicted.item())
  595. y_scores.append(logits.item())
  596. # 计算测试集评估指标
  597. accuracy = accuracy_score(y_true, y_pred)
  598. precision = precision_score(y_true, y_pred)
  599. recall = recall_score(y_true, y_pred)
  600. f1 = f1_score(y_true, y_pred)
  601. auc_value = roc_auc_score(y_true, y_scores)
  602. test_resultdata = {
  603. "Evaluation indicators": ["Accuracy", "Precision", "Recall", "F1-Score", "AUC"],
  604. "Value": [round(accuracy, 4), round(precision, 4), round(recall, 4), round(f1, 4), round(auc_value, 4)]
  605. }
  606. test_df = pd.DataFrame(test_resultdata)
  607. # test_df.to_csv("../result/test_resultdata.csv", index=False)
  608. print(f'Test Accuracy: {accuracy:.4f}, Precision: {precision:.4f}, Recall: {recall:.4f}, F1-score: {f1:.4f}, AUC: {auc_value:.4f}')
  609. # # 提取特征名称(从第三列开始)
  610. # rna_features = rna_data.columns[2:] # RNA特征名称从第三列开始
  611. # methylation_features = methylation_data.columns[2:] # 甲基化特征名称从第三列开始
  612. # # 合并 RNA 和 甲基化数据的特征名称
  613. # selected_features = list(rna_features) + list(methylation_features)
  614. # # 计算测试集的平均积分梯度
  615. # avg_grads = get_average_integrated_gradients(model, X_test, device)
  616. # top_all_features, top_rna_features, top_methylation_features = get_sorted_features(avg_grads, selected_features, 1600)
  617. # top_all_features.to_csv("../result/top_features.csv", index=False)
  618. # top_rna_features.to_csv("../result/rna_top_features.csv", index=False)
  619. # top_methylation_features.to_csv("../result/methylation_top_features.csv", index=False)
  620. return test_df
  621. if __name__ == "__main__":
  622. parser = argparse.ArgumentParser(description="Autoencoder with Transformer")
  623. parser.add_argument('--epochs', type=int, default=20, help='Number of epochs for training')
  624. parser.add_argument('--file_path_gene', type=str, required=True, help='Path to the gene data file')
  625. parser.add_argument('--file_path_methy', type=str, required=True, help='Path to the methylation data file')
  626. args = parser.parse_args()
  627. result = main(args.epochs, args.file_path_gene, args.file_path_methy)
  628. print(result)

AE-Trans.py at commit 67ee18e, under Apache-2.0 · at the source

Overview

Authors: Kai Liao1,2,3,4, Danfeng Du5, Jiawei Li6, Jian Huang7, Xiaodan Fan1, Changshui Chen2, Shanshan Wu3, Bowei Yan8, Haibo Li1,2,3
ORCID iDs: Haibo Li
  1. The Central Laboratory of Birth Defects Prevention and Control, The Affiliated Women and Children’s Hospital of Ningbo University, Ningbo, China
  2. Ningbo Key Laboratory for the Prevention and Treatment of Embryogenic Diseases, The Affiliated Women and Children’s Hospital of Ningbo University, Ningbo, China
  3. Ningbo Key Laboratory of Genomic Medicine and Birth Defects Prevention, The Affiliated Women and Children’s Hospital of Ningbo University, Ningbo, China
  4. MOE Engineering Research Center of Gene Technology, School of Life Sciences, Fudan University, Shanghai, China
  5. Obstetrics and Gynecology Hospital of Fudan University, Shanghai, China
  6. School of Computer Science, Northwestern Polytechnical University, Shaanxi, China
  7. College of Computer Science, Chongqing University, Chongqing, China
  8. Institutes of Biomedical Sciences, Fudan University, Shanghai, China
Journal: PLoS computational biology, volume 22, issue 3, article e1014074
Dates: received 10 August 2025; accepted 26 February 2026; published online 12 March 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1371/journal.pcbi.1014074 · PMID 41818290 · PMCID PMC12994821 · OpenAlex W7135059389
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Alzheimer's / dementia (population)
Methods: Smoothing, state filtering, decompositions, Machine learning, Statistics
MeSH: Alzheimer Disease*, Algorithms, Computational Biology, DNA Methylation, Humans, Multiomics, Prefrontal Cortex, Transcriptome (* major topic)
Topic: Epigenetics and DNA Methylation (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Key Technology Breakthrough Program of Ningbo Sci-Tech Innovation YONGJIANG 2035 (2024Z221, 2024Z222, 2025Z160); Clinical Innovation Team Talent Project (CXTD202502005); Social Development Public Welfare Foundation of Ningbo (2022S035); Innovation Project of Distinguished Medical Team in Ningbo (2022020405); Ningbo Science and Technology project (2023Z178); Ningbo Medical and Health Brand Discipline (PPXK2024-06); General Project of Ningbo Public Welfare Research Program (2023S043); NINGBO Leading Medical &Health Discipline (2026-A34); Major Research Project of Ningbo Clinical Medical Research Center (2024L002)
Citations: not cited yet (Europe PMC); 40 references in the paper

Abstract

Alzheimer’s disease (AD) is a progressive neurodegenerative disorder with limited diagnostic tools and poorly understood molecular underpinnings. Although multi-omics technologies hold promise for early detection, integrating unpaired transcriptomic and epigenetic data remains a major challenge due to modality heterogeneity and small sample sizes. We present AE-Trans, an interpretable dual-channel Transformer framework that aligns RNA and DNA methylation data through cross-modal reconstruction and multi-head attention. AE-Trans achieves superior performance on prefrontal cortex datasets (accuracy = 0.9736, AUC = 0.9910) and demonstrates strong generalizability to external regions temporal cortex cohorts across brain regions (accuracy = 0.7389, AUC = 0.8432). To validate the performance within the same brain region, we tested AE-Trans on an external unpaired multi-omics dataset from the prefrontal cortex. Additionally, we validated the model on a paired multi-omics dataset to assess whether it could achieve good results in real-world scenarios. In the unpaired dataset from the external same brain region, AE-Trans achieved an accuracy of (accuracy = 0.87) and AUC of (AUC = 0.94), while in the real-world paired multi-omics dataset, the accuracy was (accuracy = 0.88) and AUC was (AUC = 0.93). These results demonstrate that AE-Trans not only validates well on external unpaired datasets, but also generalizes effectively to real-world multi-omics paired datasets, highlighting its robustness in practical applications. Through counterfactual integrated gradients, we identified key features associated with immune regulation, hormonal signaling, and neuronal metabolism. These were validated via pathway enrichment and logistic regression (AUC = 0.9749), confirming the biological relevance of model-derived markers. Furthermore, AE-Trans generalized well to two independent RNA datasets, where latent representations not only improved classification (AUCs = 0.92 and 0.89) but also stratified patients into subgroups with significantly different prognoses. These results highlight AE-Trans as a robust and explainable tool for multi-omics integration, supporting early diagnosis, biomarker discovery, and individualized risk prediction in Alzheimer’s disease.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.

bowei-color/AE-Trans

License: Apache-2.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 67ee18ed1c231053a5350891473c421e8bb04283, 22 February 2025
Languages: Python (1)
Size: 20 files, 1 script
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (1 file), pandas (1 file), PyTorch (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
3 files

Code availability

All data and code are publicly available at https://github.com/bowei-color/AE-Trans.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability

Source data are provided with this paper. Source data is publicly available on Gene Expression Omnibus by accession numbers GSE33000, GSE44770 and GSE80970. Our postprocessed form of the these publicly available data is available at https://doi.org/10.5281/zenodo.13933763.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 8 MeSH terms, 10 funders, 39 references.

Cite

This paper

Liao, K., Du, D., Li, J., Huang, J., Fan, X., Chen, C., Wu, S., Yan, B., & Li, H. (2026). Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstruction. PLoS computational biology, 22(3), e1014074. https://doi.org/10.1371/journal.pcbi.1014074

BibTeX

@article{liao2026interpretable,
author = {Liao, Kai and Du, Danfeng and Li, Jiawei and Huang, Jian and Fan, Xiaodan and Chen, Changshui and Wu, Shanshan and Yan, Bowei and Li, Haibo},
title = {{Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstruction}},
journal = {PLoS computational biology},
year = {2026},
month = mar,
volume = {22},
number = {3},
pages = {e1014074},
publisher = {PLOS},
issn = {1553-734X},
doi = {10.1371/journal.pcbi.1014074},
url = {https://doi.org/10.1371/journal.pcbi.1014074},
pmid = {41818290},
pmcid = {PMC12994821}
}

RIS

TY - JOUR
AU - Liao, Kai
AU - Du, Danfeng
AU - Li, Jiawei
AU - Huang, Jian
AU - Fan, Xiaodan
AU - Chen, Changshui
AU - Wu, Shanshan
AU - Yan, Bowei
AU - Li, Haibo
TI - Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstruction
T2 - PLoS computational biology
J2 - PLoS Comput Biol
PY - 2026
DA - 2026/03/12
VL - 22
IS - 3
SP - e1014074
SN - 1553-734X
PB - PLOS
DO - 10.1371/journal.pcbi.1014074
UR - https://doi.org/10.1371/journal.pcbi.1014074
LA - en
ER -

CSL-JSON

{
"id": "10.1371/journal.pcbi.1014074",
"type": "article-journal",
"title": "Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstruction",
"container-title": "PLoS computational biology",
"author": [
{
"family": "Liao",
"given": "Kai"
},
{
"family": "Du",
"given": "Danfeng"
},
{
"family": "Li",
"given": "Jiawei"
},
{
"family": "Huang",
"given": "Jian"
},
{
"family": "Fan",
"given": "Xiaodan"
},
{
"family": "Chen",
"given": "Changshui"
},
{
"family": "Wu",
"given": "Shanshan"
},
{
"family": "Yan",
"given": "Bowei"
},
{
"family": "Li",
"given": "Haibo"
}
],
"container-title-short": "PLoS Comput Biol",
"volume": "22",
"issue": "3",
"page": "e1014074",
"DOI": "10.1371/journal.pcbi.1014074",
"PMID": "41818290",
"PMCID": "PMC12994821",
"ISSN": "1553-734X",
"publisher": "PLOS",
"URL": "https://doi.org/10.1371/journal.pcbi.1014074",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3390/biomedicines14050998 [code]
Integrative Multi-Omics and Machine Learning Analysis Identifies Therapeutic Targets and Drug Repurposing Candidates for Alzheimer's Disease.
Journal: Biomedicines
In common: NCBI GEO GSE33000, Alzheimer's / dementia, genetics / omics, 2 references
[2] doi:10.3390/ijms27146196
Investigation of the Potential Neuroprotective Mechanisms of <i>Acalypha indica</i> Against Alzheimer's Disease by Integrated Bioinformatics Analysis.
Journal: International journal of molecular sciences
In common: NCBI GEO GSE33000, Alzheimer's / dementia, genetics / omics, 1 reference
[3] doi:10.1002/cns.71021
Multilayer Proteome and Metabolome-Based Validation Uncovers Combined Regulatory Roles and Predictive Values of 6 RNA Modifications and Cellular Senescence in Alzheimer's Disease.
Journal: CNS neuroscience & therapeutics
In common: Alzheimer's / dementia, genetics / omics, 3 references
[4] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: scikit-learn, pandas, NumPy, Alzheimer's / dementia, 2 references
[5] doi:10.1016/j.isci.2026.116336 [code]
Transcriptomic signatures of synaptic loss in Alzheimer's disease.
Journal: iScience
In common: scikit-learn, pandas, NumPy, Alzheimer's / dementia, genetics / omics, 2 references
[6] doi:10.1038/s41419-026-08791-1
The PM20D1-OLE pathway induces microglia rewiring to ameliorate Alzheimer disease.
Journal: Cell death & disease
In common: NCBI GEO GSE33000, Alzheimer's / dementia, genetics / omics
[7] doi:10.3389/fimmu.2026.1937738
Multi-regional transcriptomic analysis reveals early nociceptive and neuroinflammatory alterations in APP/PS1 mice.
Journal: Frontiers in immunology
In common: NCBI GEO GSE33000, Alzheimer's / dementia, genetics / omics
[8] doi:10.1186/s12859-026-06490-4 [code]
Tissueformer: extending single-cell foundation models to predict population-level phenotypes.
Journal: BMC bioinformatics
In common: PyTorch, scikit-learn, pandas, 1 other tool, genetics / omics, 1 reference
[9] doi:10.1038/s41467-026-74694-6 [code]
Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 1 other tool, genetics / omics, 1 reference
[10] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: PyTorch, scikit-learn, pandas, 1 other tool, genetics / omics, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.