OSCR

A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction.

Code ↔ Paper

17 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 17 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Materials and methods › Model construction ↔ operation/loss_func.py, lines 31–106 · score 0.87 · Negative Log Likelihood, deep survival prediction, Loss function, discrete survival, batch, network
  2. [2] § Materials and methods › Evaluation metrics ↔ biological_plausibility/scripts/01_download_databases.py, lines 560–640 · score 0.73 · OpenTargets, CancerMine, DGIdb, OncoKB, databases, Biological
  3. [3] § Materials and methods › Model construction ↔ operation/core_utils.py, lines 105–254 · score 0.71 · Adam, L1, Warmup, patience, epoch, optimize
  4. [4] § Results › DeepSHAP outperforms other XAI methods in discovery of prognostic factors from pan-cancer transcriptomic data ↔ biological_plausibility/scripts/01_download_databases.py, lines 281–324 · score 0.71 · Cell Carcinoma, Grade Glioma, Glioblastoma, Papillary, Renal, GBMLGG
  5. [5] § Results › LRP achieves the highest biological consistency with established cancer gene databases ↔ biological_plausibility/scripts/01_download_databases.py, lines 560–640 · score 0.70 · OpenTargets, CancerMine, DGIdb, OncoKB, database, BRCA
  6. [6] § Materials and methods › Dataset description ↔ datasets_csv/preprocessing_cancer_single.py, lines 58–166 · score 0.70 · Molecular Signatures Database, MSigDB, preprocessed, TCGA, censorship, genes
  7. [7] § Materials and methods › Software and environment ↔ biological_plausibility/scripts/01_download_databases.py, lines 450–557 · score 0.66 · Literature mined cancer, CancerMine, gene associations, Zenodo
  8. [8] § Materials and methods › Dataset description ↔ datasets_csv/preprocessing_no_normalization.py, lines 103–142 · score 0.65 · Molecular Signatures Database, MSigDB, preprocessed, cancer
  9. [9] § Results › DeepSHAP delivers the optimal comprehensive performance across prognostic, biological and stability metrics ↔ biological_plausibility/scripts/04_visualize_2.py, lines 1697–1771 · score 0.64 · DB supported Hits, Min Max normalization, XAI categories, radar, biological, Prognostic Factor
  10. [10] § Materials and methods › Software and environment ↔ biological_plausibility/scripts/01_download_databases.py, lines 376–447 · score 0.64 · Gene disease associations, Open Targets Platform, GraphQL, API
  11. [11] § Results › LRP achieves the highest biological consistency with established cancer gene databases ↔ biological_plausibility/scripts/02_calculate_gene_scores.py, lines 275–341 · score 0.63 · DGIdb, OpenTargets, CancerMine, OncoKB, BRCA, database
  12. [12] § Materials and methods › Interpretability framework ↔ operation/lrp_bootstrap_analysis.py, lines 18–69 · score 0.60 · AlphaDropout, bias, weighted, SELU, activations, layers
  13. [13] § Materials and methods › Interpretability framework ↔ operation/lrp_individual_analysis.py, lines 34–85 · score 0.60 · AlphaDropout, bias, weighted, SELU, activations, layers
  14. [14] § Materials and methods › Evaluation metrics ↔ biological_plausibility/scripts/02_calculate_gene_scores.py, lines 1–39 · score 0.59 · CancerMine, DGIdb, OncoKB, databases, Biological
  15. [15] § Results › DeepSHAP outperforms other XAI methods in discovery of prognostic factors from pan-cancer transcriptomic data ↔ biological_plausibility/scripts/database_loader.py, lines 19–37 · score 0.56 · Brain, Cell, Glioblastoma, Kidney, Papillary, Renal
  16. [16] § Materials and methods › Software and environment ↔ operation/command.sh, the whole file · a weak match · score 0.54 · DeepLIFT, DeepSHAP, Python, SNN, IG
  17. [17] § Materials and methods › Software and environment ↔ operation/evaluate_faithfulness.py, lines 335–383 · score 0.53 · GradientSHAP, DeepLIFT, DeepSHAP, IG

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 644 lines · 30 KB · MIT · 5 matches

  1. #!/usr/bin/env python3
  2. """
  3. Step 1: 下载并准备各数据库
  4. - OncoKB: 癌症基因临床分级 (需API Token)
  5. - CGC: Cancer Gene Census (需COSMIC账号)
  6. - DGIdb: 药物-基因互作 (需手动下载)
  7. - CIViC: 临床变异解读 (可自动下载)
  8. 运行方式:
  9. python scripts/01_download_databases.py
  10. python scripts/01_download_databases.py --skip_existing # 跳过已存在的文件
  11. """
  12. import os
  13. import sys
  14. import time
  15. import argparse
  16. import requests
  17. import pandas as pd
  18. from pathlib import Path
  19. from io import StringIO
  20. # 添加父目录到路径
  21. sys.path.insert(0, str(Path(__file__).parent.parent))
  22. from config import DATABASE_DIR, ONCOKB_API_TOKEN
  23. def check_file_valid(filepath, min_rows=5):
  24. """检查文件是否存在且有效(非HTML,有足够行数)"""
  25. if not filepath.exists():
  26. return False
  27. try:
  28. # 检查是否为HTML(下载失败的情况)
  29. with open(filepath, 'r') as f:
  30. first_line = f.readline()
  31. if first_line.strip().startswith('<!') or first_line.strip().startswith('<html'):
  32. return False
  33. # 检查行数
  34. df = pd.read_csv(filepath, sep='\t' if filepath.suffix == '.tsv' else ',', nrows=min_rows)
  35. return len(df) >= min_rows
  36. except:
  37. return False
  38. def write_version_file(data_file: Path, lines):
  39. """为数据文件生成/更新一个 *_VERSION.txt 说明文件
  40. 版本文件命名规则: basename + "_VERSION.txt",例如:
  41. - oncokb_genes.tsv -> oncokb_genes_VERSION.txt
  42. - opentargets_associations.tsv -> opentargets_associations_VERSION.txt
  43. """
  44. version_path = data_file.with_name(f"{data_file.stem}_VERSION.txt")
  45. timestamp = time.strftime("%Y-%m-%d %H:%M:%S")
  46. header = [f"Generated: {timestamp}"]
  47. content = header + list(lines)
  48. try:
  49. with open(version_path, "w", encoding="utf-8") as f:
  50. for line in content:
  51. f.write(str(line).rstrip("\n") + "\n")
  52. except Exception as e:
  53. print(f" ⚠ 写入版本说明文件失败: {version_path} ({e})")
  54. def download_oncokb():
  55. """
  56. 下载 OncoKB 癌症基因列表
  57. 注意: OncoKB 需要 API token,请先申请:
  58. https://www.oncokb.org/account/register
  59. """
  60. print("\n" + "="*60)
  61. print("下载 OncoKB 数据")
  62. print("="*60)
  63. output_file = DATABASE_DIR / "oncokb_genes.tsv"
  64. if ONCOKB_API_TOKEN:
  65. # 使用 API 获取完整数据
  66. headers = {"Authorization": f"Bearer {ONCOKB_API_TOKEN}"}
  67. url = "https://www.oncokb.org/api/v1/utils/allCuratedGenes"
  68. try:
  69. response = requests.get(url, headers=headers)
  70. response.raise_for_status()
  71. genes = response.json()
  72. df = pd.DataFrame(genes)
  73. df.to_csv(output_file, sep="\t", index=False)
  74. print(f"✓ 已保存 {len(df)} 个 OncoKB 基因到 {output_file}")
  75. write_version_file(output_file, [
  76. f"Source: {url}",
  77. "Description: OncoKB curated genes downloaded via official API",
  78. ])
  79. return True
  80. except Exception as e:
  81. print(f"✗ OncoKB API 请求失败: {e}")
  82. # 如果没有 API token,创建模板文件
  83. print("⚠ 未设置 ONCOKB_API_TOKEN 环境变量")
  84. print(" 请手动下载数据或设置环境变量后重试")
  85. print(" 申请地址: https://www.oncokb.org/account/register")
  86. # 创建模板文件,包含常见癌症基因
  87. template_genes = [
  88. {"hugoSymbol": "TP53", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
  89. {"hugoSymbol": "EGFR", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  90. {"hugoSymbol": "BRAF", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  91. {"hugoSymbol": "KRAS", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  92. {"hugoSymbol": "PIK3CA", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  93. {"hugoSymbol": "ERBB2", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  94. {"hugoSymbol": "ALK", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  95. {"hugoSymbol": "ROS1", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  96. {"hugoSymbol": "MET", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  97. {"hugoSymbol": "RET", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
  98. {"hugoSymbol": "BRCA1", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
  99. {"hugoSymbol": "BRCA2", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
  100. {"hugoSymbol": "ATM", "highestSensitiveLevel": "2", "oncogene": False, "tsg": True},
  101. {"hugoSymbol": "PTEN", "highestSensitiveLevel": "2", "oncogene": False, "tsg": True},
  102. {"hugoSymbol": "AKT1", "highestSensitiveLevel": "2", "oncogene": True, "tsg": False},
  103. ]
  104. print("\n ⚠ 创建模板文件(仅包含示例基因,请替换为完整数据)")
  105. df = pd.DataFrame(template_genes)
  106. df.to_csv(output_file, sep="\t", index=False)
  107. print(f" 模板已保存到 {output_file}")
  108. write_version_file(output_file, [
  109. "Source: local template (no ONCOKB_API_TOKEN)",
  110. "Description: example curated genes only, please replace with full OncoKB data",
  111. "Manual action: download full data from https://www.oncokb.org/ and overwrite this file",
  112. ])
  113. return False
  114. def download_dgidb():
  115. """
  116. 下载 DGIdb 药物-基因互作数据
  117. 注意: DGIdb 网站已改版 (2024),旧API失效,需手动下载
  118. 下载地址: https://dgidb.org/downloads
  119. """
  120. print("\n" + "="*60)
  121. print("下载 DGIdb 数据")
  122. print("="*60)
  123. output_file = DATABASE_DIR / "dgidb_interactions.tsv"
  124. # 检查是否已有有效文件
  125. if check_file_valid(output_file, min_rows=100):
  126. print(f" ✓ 已存在有效数据文件: {output_file}")
  127. df = pd.read_csv(output_file, sep="\t")
  128. print(f" 包含 {len(df)} 条记录")
  129. # 保留用户手动下载版本,不覆盖版本说明,只在缺失时创建一个简短说明
  130. if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
  131. write_version_file(output_file, [
  132. "Source: user-provided DGIdb interactions.tsv",
  133. "Description: manually downloaded from https://dgidb.org/downloads",
  134. ])
  135. return True
  136. # DGIdb 网站已改版,API不可用,需要手动下载
  137. print(" ⚠ DGIdb 网站已改版,需要手动下载数据")
  138. print("")
  139. print(" 【手动下载步骤】:")
  140. print(" 1. 浏览器访问: https://dgidb.org/downloads")
  141. print(" 2. 找到并下载: interactions.tsv")
  142. print(" 3. 将文件保存到:")
  143. print(f" {output_file}")
  144. print("")
  145. # 创建扩展模板(包含常见 FDA 批准靶点药物)
  146. print(" 创建临时模板文件(包含常见靶点药物)...")
  147. template = [
  148. # EGFR 抑制剂
  149. {"gene_name": "EGFR", "drug_name": "Erlotinib", "interaction_types": "inhibitor", "approved": True},
  150. {"gene_name": "EGFR", "drug_name": "Gefitinib", "interaction_types": "inhibitor", "approved": True},
  151. {"gene_name": "EGFR", "drug_name": "Osimertinib", "interaction_types": "inhibitor", "approved": True},
  152. {"gene_name": "EGFR", "drug_name": "Afatinib", "interaction_types": "inhibitor", "approved": True},
  153. # BRAF 抑制剂
  154. {"gene_name": "BRAF", "drug_name": "Vemurafenib", "interaction_types": "inhibitor", "approved": True},
  155. {"gene_name": "BRAF", "drug_name": "Dabrafenib", "interaction_types": "inhibitor", "approved": True},
  156. {"gene_name": "BRAF", "drug_name": "Encorafenib", "interaction_types": "inhibitor", "approved": True},
  157. # HER2/ERBB2 抑制剂
  158. {"gene_name": "ERBB2", "drug_name": "Trastuzumab", "interaction_types": "antibody", "approved": True},
  159. {"gene_name": "ERBB2", "drug_name": "Pertuzumab", "interaction_types": "antibody", "approved": True},
  160. {"gene_name": "ERBB2", "drug_name": "Lapatinib", "interaction_types": "inhibitor", "approved": True},
  161. {"gene_name": "ERBB2", "drug_name": "Neratinib", "interaction_types": "inhibitor", "approved": True},
  162. # ALK 抑制剂
  163. {"gene_name": "ALK", "drug_name": "Crizotinib", "interaction_types": "inhibitor", "approved": True},
  164. {"gene_name": "ALK", "drug_name": "Alectinib", "interaction_types": "inhibitor", "approved": True},
  165. {"gene_name": "ALK", "drug_name": "Brigatinib", "interaction_types": "inhibitor", "approved": True},
  166. {"gene_name": "ALK", "drug_name": "Lorlatinib", "interaction_types": "inhibitor", "approved": True},
  167. # PARP 抑制剂 (BRCA)
  168. {"gene_name": "BRCA1", "drug_name": "Olaparib", "interaction_types": "inhibitor", "approved": True},
  169. {"gene_name": "BRCA2", "drug_name": "Olaparib", "interaction_types": "inhibitor", "approved": True},
  170. {"gene_name": "BRCA1", "drug_name": "Rucaparib", "interaction_types": "inhibitor", "approved": True},
  171. {"gene_name": "BRCA2", "drug_name": "Rucaparib", "interaction_types": "inhibitor", "approved": True},
  172. {"gene_name": "BRCA1", "drug_name": "Niraparib", "interaction_types": "inhibitor", "approved": True},
  173. {"gene_name": "BRCA2", "drug_name": "Niraparib", "interaction_types": "inhibitor", "approved": True},
  174. # PIK3CA
  175. {"gene_name": "PIK3CA", "drug_name": "Alpelisib", "interaction_types": "inhibitor", "approved": True},
  176. # KRAS
  177. {"gene_name": "KRAS", "drug_name": "Sotorasib", "interaction_types": "inhibitor", "approved": True},
  178. {"gene_name": "KRAS", "drug_name": "Adagrasib", "interaction_types": "inhibitor", "approved": True},
  179. # RET
  180. {"gene_name": "RET", "drug_name": "Selpercatinib", "interaction_types": "inhibitor", "approved": True},
  181. {"gene_name": "RET", "drug_name": "Pralsetinib", "interaction_types": "inhibitor", "approved": True},
  182. # MET
  183. {"gene_name": "MET", "drug_name": "Capmatinib", "interaction_types": "inhibitor", "approved": True},
  184. {"gene_name": "MET", "drug_name": "Tepotinib", "interaction_types": "inhibitor", "approved": True},
  185. # ROS1
  186. {"gene_name": "ROS1", "drug_name": "Crizotinib", "interaction_types": "inhibitor", "approved": True},
  187. {"gene_name": "ROS1", "drug_name": "Entrectinib", "interaction_types": "inhibitor", "approved": True},
  188. # NTRK
  189. {"gene_name": "NTRK1", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
  190. {"gene_name": "NTRK2", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
  191. {"gene_name": "NTRK3", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
  192. {"gene_name": "NTRK1", "drug_name": "Entrectinib", "interaction_types": "inhibitor", "approved": True},
  193. # FGFR
  194. {"gene_name": "FGFR2", "drug_name": "Pemigatinib", "interaction_types": "inhibitor", "approved": True},
  195. {"gene_name": "FGFR2", "drug_name": "Erdafitinib", "interaction_types": "inhibitor", "approved": True},
  196. {"gene_name": "FGFR3", "drug_name": "Erdafitinib", "interaction_types": "inhibitor", "approved": True},
  197. # IDH
  198. {"gene_name": "IDH1", "drug_name": "Ivosidenib", "interaction_types": "inhibitor", "approved": True},
  199. {"gene_name": "IDH2", "drug_name": "Enasidenib", "interaction_types": "inhibitor", "approved": True},
  200. # BCR-ABL
  201. {"gene_name": "BCR", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
  202. {"gene_name": "ABL1", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
  203. {"gene_name": "ABL1", "drug_name": "Dasatinib", "interaction_types": "inhibitor", "approved": True},
  204. {"gene_name": "ABL1", "drug_name": "Nilotinib", "interaction_types": "inhibitor", "approved": True},
  205. # KIT
  206. {"gene_name": "KIT", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
  207. {"gene_name": "KIT", "drug_name": "Sunitinib", "interaction_types": "inhibitor", "approved": True},
  208. # PDGFRA
  209. {"gene_name": "PDGFRA", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
  210. {"gene_name": "PDGFRA", "drug_name": "Avapritinib", "interaction_types": "inhibitor", "approved": True},
  211. # FLT3
  212. {"gene_name": "FLT3", "drug_name": "Midostaurin", "interaction_types": "inhibitor", "approved": True},
  213. {"gene_name": "FLT3", "drug_name": "Gilteritinib", "interaction_types": "inhibitor", "approved": True},
  214. # BTK
  215. {"gene_name": "BTK", "drug_name": "Ibrutinib", "interaction_types": "inhibitor", "approved": True},
  216. {"gene_name": "BTK", "drug_name": "Acalabrutinib", "interaction_types": "inhibitor", "approved": True},
  217. # BCL2
  218. {"gene_name": "BCL2", "drug_name": "Venetoclax", "interaction_types": "inhibitor", "approved": True},
  219. # VEGFR
  220. {"gene_name": "KDR", "drug_name": "Bevacizumab", "interaction_types": "antibody", "approved": True},
  221. {"gene_name": "KDR", "drug_name": "Sorafenib", "interaction_types": "inhibitor", "approved": True},
  222. {"gene_name": "KDR", "drug_name": "Sunitinib", "interaction_types": "inhibitor", "approved": True},
  223. # ESR1 (雌激素受体)
  224. {"gene_name": "ESR1", "drug_name": "Tamoxifen", "interaction_types": "antagonist", "approved": True},
  225. {"gene_name": "ESR1", "drug_name": "Fulvestrant", "interaction_types": "antagonist", "approved": True},
  226. # AR (雄激素受体)
  227. {"gene_name": "AR", "drug_name": "Enzalutamide", "interaction_types": "antagonist", "approved": True},
  228. {"gene_name": "AR", "drug_name": "Abiraterone", "interaction_types": "inhibitor", "approved": True},
  229. # CDK4/6
  230. {"gene_name": "CDK4", "drug_name": "Palbociclib", "interaction_types": "inhibitor", "approved": True},
  231. {"gene_name": "CDK6", "drug_name": "Palbociclib", "interaction_types": "inhibitor", "approved": True},
  232. {"gene_name": "CDK4", "drug_name": "Ribociclib", "interaction_types": "inhibitor", "approved": True},
  233. {"gene_name": "CDK6", "drug_name": "Ribociclib", "interaction_types": "inhibitor", "approved": True},
  234. # MTOR
  235. {"gene_name": "MTOR", "drug_name": "Everolimus", "interaction_types": "inhibitor", "approved": True},
  236. {"gene_name": "MTOR", "drug_name": "Temsirolimus", "interaction_types": "inhibitor", "approved": True},
  237. # SMO (Hedgehog)
  238. {"gene_name": "SMO", "drug_name": "Vismodegib", "interaction_types": "inhibitor", "approved": True},
  239. {"gene_name": "SMO", "drug_name": "Sonidegib", "interaction_types": "inhibitor", "approved": True},
  240. ]
  241. df = pd.DataFrame(template)
  242. df.to_csv(output_file, sep="\t", index=False)
  243. print(f" ✓ 模板已保存: {output_file}")
  244. write_version_file(output_file, [
  245. "Source: local template (DGIdb manual download required)",
  246. "Description: contains common FDA-approved targeted agents only, please replace with full DGIdb interactions.tsv",
  247. "Manual action: download interactions.tsv from https://dgidb.org/downloads and overwrite this file",
  248. ])
  249. print(f" 包含 {len(template)} 条 FDA 批准靶向药物记录")
  250. print(" ⚠ 请尽快下载完整数据替换此模板")
  251. return False
  252. def download_opentargets():
  253. """
  254. 下载 Open Targets Platform 基因-疾病关联数据
  255. Open Targets API 免费无需注册
  256. 为每个 TCGA 癌症类型下载相关基因列表
  257. """
  258. print("\n" + "="*60)
  259. print("下载 Open Targets 数据")
  260. print("="*60)
  261. output_file = DATABASE_DIR / "opentargets_associations.tsv"
  262. # 检查是否已有有效文件
  263. if check_file_valid(output_file, min_rows=100):
  264. print(f" ✓ 已存在有效数据文件: {output_file}")
  265. df = pd.read_csv(output_file, sep="\t")
  266. print(f" 包含 {len(df)} 条记录")
  267. if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
  268. write_version_file(output_file, [
  269. "Source: user-provided or previously downloaded Open Targets associations",
  270. "Description: gene-disease associations for TCGA cancer types via EFO/MONDO IDs",
  271. "URL: https://api.platform.opentargets.org/api/v4/graphql",
  272. ])
  273. return True
  274. # TCGA 癌种到 Open Targets EFO/MONDO ID 的映射 (已通过 API 验证)
  275. TCGA_TO_EFO = {
  276. "BLCA": ("MONDO_0001187", "bladder carcinoma"),
  277. "BRCA": ("EFO_0000305", "breast carcinoma"),
  278. "COADREAD": ("MONDO_0005575", "colorectal cancer"),
  279. "GBMLGG": ("EFO_0000519", "glioblastoma"),
  280. "HNSC": ("EFO_0000181", "head and neck squamous cell carcinoma"),
  281. "KIRC": ("EFO_0000349", "renal cell carcinoma"),
  282. "KIRP": ("EFO_0000640", "papillary renal cell carcinoma"),
  283. "LGG": ("EFO_0005543", "low grade glioma"),
  284. "LIHC": ("EFO_0000182", "hepatocellular carcinoma"),
  285. "LUAD": ("EFO_0000571", "lung adenocarcinoma"),
  286. "LUSC": ("EFO_0000708", "lung squamous cell carcinoma"),
  287. "PAAD": ("EFO_1000044", "pancreatic adenocarcinoma"),
  288. "SKCM": ("EFO_0000389", "melanoma"),
  289. "STAD": ("EFO_0000503", "stomach carcinoma"),
  290. "UCEC": ("EFO_1001512", "endometrial carcinoma"),
  291. }
  292. # API 分页限制
  293. MAX_PAGE_SIZE = 2500 # API 最大允许 3000,保守设置 2500
  294. GRAPHQL_URL = "https://api.platform.opentargets.org/api/v4/graphql"
  295. # GraphQL 查询 - 获取疾病关联的所有靶点
  296. query = """
  297. query DiseaseAssociations($efoId: String!, $size: Int!) {
  298. disease(efoId: $efoId) {
  299. id
  300. name
  301. associatedTargets(page: { index: 0, size: $size }) {
  302. count
  303. rows {
  304. target {
  305. id
  306. approvedSymbol
  307. }
  308. score
  309. }
  310. }
  311. }
  312. }
  313. """
  314. all_associations = []
  315. print(f" 将下载 {len(TCGA_TO_EFO)} 种癌症类型的关联数据...")
  316. for tcga_code, (efo_id, disease_name) in TCGA_TO_EFO.items():
  317. print(f" {tcga_code} ({disease_name})...", end=" ", flush=True)
  318. try:
  319. # 分页下载关联数据
  320. page_index = 0
  321. total_rows = 0
  322. while True:
  323. response = requests.post(
  324. GRAPHQL_URL,
  325. json={
  326. "query": query,
  327. "variables": {"efoId": efo_id, "size": MAX_PAGE_SIZE}
  328. },
  329. timeout=120,
  330. headers={"Content-Type": "application/json"}
  331. )
  332. response.raise_for_status()
  333. data = response.json()
  334. # 检查是否有错误
  335. if "errors" in data:
  336. error_msg = data['errors'][0].get('message', '')[:100]
  337. print(f"API错误: {error_msg}")
  338. break
  339. # 安全解析
  340. disease_data = data.get("data", {})
  341. if disease_data is None:
  342. print("data 为空")
  343. break
  344. disease_info = disease_data.get("disease")
  345. if disease_info is None:
  346. print(f"未找到疾病 {efo_id}")
  347. break
  348. associated = disease_info.get("associatedTargets")
  349. if associated is None:
  350. print("无关联靶点")
  351. break
  352. rows = associated.get("rows", [])
  353. if not rows:
  354. break
  355. for row in rows:
  356. target = row.get("target", {}) or {}
  357. gene_symbol = target.get("approvedSymbol", "")
  358. if gene_symbol:
  359. all_associations.append({
  360. "tcga_code": tcga_code,
  361. "efo_id": efo_id,
  362. "disease_name": disease_name,
  363. "gene_symbol": gene_symbol,
  364. "ensembl_id": target.get("id", ""),
  365. "association_score": row.get("score", 0),
  366. })
  367. total_rows += len(rows)
  368. # 只取第一页 (2500 个基因已经足够覆盖 Top 100)
  369. break
  370. if total_rows > 0:
  371. print(f"{total_rows} 个基因")
  372. # 避免 API 限流
  373. time.sleep(0.3)
  374. except requests.exceptions.RequestException as e:
  375. print(f"网络错误: {e}")
  376. except Exception as e:
  377. print(f"失败: {type(e).__name__}: {e}")
  378. if all_associations:
  379. df = pd.DataFrame(all_associations)
  380. df.to_csv(output_file, sep="\t", index=False)
  381. print(f"\n ✓ 已保存 {len(df)} 条关联记录到 {output_file}")
  382. # 统计
  383. gene_count = df['gene_symbol'].nunique()
  384. print(f" 涵盖 {gene_count} 个独立基因")
  385. write_version_file(output_file, [
  386. "Source: Open Targets Platform GraphQL API",
  387. "Description: gene-disease associations for TCGA cancer types (first page up to 2500 targets per disease)",
  388. "URL: https://api.platform.opentargets.org/api/v4/graphql",
  389. ])
  390. return True
  391. else:
  392. print(" ✗ 未能下载任何数据")
  393. return False
  394. def download_cancermine():
  395. """
  396. 下载 CancerMine 文献挖掘数据
  397. CancerMine 是基于文献挖掘的癌症基因数据库,包含三种角色:
  398. - Driver: 驱动基因
  399. - Oncogene: 癌基因
  400. - Tumor_Suppressor: 抑癌基因
  401. 数据下载地址: https://zenodo.org/records/7689627
  402. (原网站 http://bionlp.bcgsc.ca/cancermine/ 已下线)
  403. """
  404. print("\n" + "="*60)
  405. print("下载 CancerMine 数据")
  406. print("="*60)
  407. output_file = DATABASE_DIR / "cancermine.tsv"
  408. # 检查是否已有有效文件
  409. if check_file_valid(output_file, min_rows=100):
  410. print(f" ✓ 已存在有效数据文件: {output_file}")
  411. df = pd.read_csv(output_file, sep="\t")
  412. print(f" 包含 {len(df)} 条记录")
  413. if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
  414. write_version_file(output_file, [
  415. "Source: user-provided or previously downloaded CancerMine data",
  416. "Description: literature-mined cancer gene associations (Driver/Oncogene/Tumor_Suppressor)",
  417. "URL: https://zenodo.org/records/7689627",
  418. ])
  419. return True
  420. # CancerMine 公开下载链接 (Zenodo - 原网站已下线)
  421. url = "https://zenodo.org/records/7689627/files/cancermine_collated.tsv?download=1"
  422. try:
  423. print(f" 下载: {url}")
  424. response = requests.get(url, timeout=300)
  425. response.raise_for_status()
  426. # 检查是否为有效TSV
  427. if response.text.strip().startswith('<!') or '<html' in response.text[:100].lower():
  428. print(" ✗ 返回HTML而非数据")
  429. raise Exception("Invalid response")
  430. df = pd.read_csv(StringIO(response.text), sep="\t")
  431. if len(df) > 100:
  432. df.to_csv(output_file, sep="\t", index=False)
  433. print(f" ✓ 已保存 {len(df)} 条 CancerMine 记录到 {output_file}")
  434. # 统计角色分布
  435. if 'role' in df.columns:
  436. role_counts = df['role'].value_counts()
  437. print(f" 角色分布:")
  438. for role, count in role_counts.items():
  439. print(f" - {role}: {count}")
  440. # 统计基因和癌症数量
  441. gene_col = 'gene_normalized' if 'gene_normalized' in df.columns else 'gene'
  442. cancer_col = 'cancer_normalized' if 'cancer_normalized' in df.columns else 'cancer'
  443. if gene_col in df.columns:
  444. print(f" 独立基因数: {df[gene_col].nunique()}")
  445. if cancer_col in df.columns:
  446. print(f" 独立癌症类型数: {df[cancer_col].nunique()}")
  447. write_version_file(output_file, [
  448. f"Source: {url}",
  449. "Description: CancerMine collated data - literature-mined cancer gene associations",
  450. "Roles: Driver, Oncogene, Tumor_Suppressor",
  451. f"Total records: {len(df)}",
  452. ])
  453. return True
  454. else:
  455. print(" ✗ 数据行数不足")
  456. except Exception as e:
  457. print(f" ✗ 下载失败: {e}")
  458. # 下载失败,创建模板
  459. print(" ⚠ 自动下载失败,创建模板文件...")
  460. print(" 请手动下载: https://zenodo.org/records/7689627")
  461. print(" 选择 cancermine_collated.tsv 并保存为 databases/cancermine.tsv")
  462. template = [
  463. {"gene_normalized": "TP53", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 100},
  464. {"gene_normalized": "BRCA1", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 80},
  465. {"gene_normalized": "BRCA2", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 70},
  466. {"gene_normalized": "EGFR", "cancer_normalized": "lung cancer", "role": "Oncogene", "citation_count": 90},
  467. {"gene_normalized": "KRAS", "cancer_normalized": "colorectal cancer", "role": "Driver", "citation_count": 85},
  468. {"gene_normalized": "BRAF", "cancer_normalized": "melanoma", "role": "Oncogene", "citation_count": 75},
  469. {"gene_normalized": "PIK3CA", "cancer_normalized": "breast cancer", "role": "Oncogene", "citation_count": 60},
  470. {"gene_normalized": "PTEN", "cancer_normalized": "prostate cancer", "role": "Tumor_Suppressor", "citation_count": 55},
  471. {"gene_normalized": "APC", "cancer_normalized": "colorectal cancer", "role": "Tumor_Suppressor", "citation_count": 50},
  472. {"gene_normalized": "IDH1", "cancer_normalized": "glioma", "role": "Driver", "citation_count": 45},
  473. {"gene_normalized": "ERBB2", "cancer_normalized": "breast cancer", "role": "Oncogene", "citation_count": 65},
  474. {"gene_normalized": "MYC", "cancer_normalized": "lymphoma", "role": "Oncogene", "citation_count": 70},
  475. {"gene_normalized": "RB1", "cancer_normalized": "retinoblastoma", "role": "Tumor_Suppressor", "citation_count": 40},
  476. {"gene_normalized": "VHL", "cancer_normalized": "kidney cancer", "role": "Tumor_Suppressor", "citation_count": 35},
  477. {"gene_normalized": "ALK", "cancer_normalized": "lung cancer", "role": "Oncogene", "citation_count": 55},
  478. ]
  479. df = pd.DataFrame(template)
  480. df.to_csv(output_file, sep="\t", index=False)
  481. print(f" 模板已保存到 {output_file}")
  482. write_version_file(output_file, [
  483. "Source: local template (CancerMine manual download required)",
  484. "Description: example cancer gene associations only, please replace with full CancerMine data",
  485. "Manual action: download cancermine_collated.tsv from http://bionlp.bcgsc.ca/cancermine/ and overwrite this file",
  486. ])
  487. return False
  488. def main():
  489. parser = argparse.ArgumentParser(description="下载生物学数据库")
  490. parser.add_argument('--skip_existing', action='store_true',
  491. help='跳过已存在的有效文件')
  492. args = parser.parse_args()
  493. print("="*60)
  494. print("XAI 生物学合理性评估 - 数据库下载")
  495. print("="*60)
  496. # 确保目录存在
  497. DATABASE_DIR.mkdir(parents=True, exist_ok=True)
  498. print(f"数据库目录: {DATABASE_DIR}")
  499. # 下载各数据库
  500. results = {
  501. "OncoKB": download_oncokb(),
  502. "DGIdb": download_dgidb(),
  503. "OpenTargets": download_opentargets(),
  504. "CancerMine": download_cancermine()
  505. }
  506. # 汇总
  507. print("\n" + "="*60)
  508. print("下载汇总")
  509. print("="*60)
  510. success_count = 0
  511. for db, success in results.items():
  512. if success:
  513. status = "✓ 完整数据"
  514. success_count += 1
  515. else:
  516. status = "⚠ 模板数据 (需替换)"
  517. print(f" {db}: {status}")
  518. print(f"\n 完整数据: {success_count}/6")
  519. # 检查文件大小
  520. print("\n" + "="*60)
  521. print("文件详情")
  522. print("="*60)
  523. files_info = [
  524. ("oncokb_genes.tsv", "OncoKB"),
  525. ("dgidb_interactions.tsv", "DGIdb"),
  526. ("opentargets_associations.tsv", "OpenTargets"),
  527. ("cancermine.tsv", "CancerMine"),
  528. ]
  529. for filename, db_name in files_info:
  530. filepath = DATABASE_DIR / filename
  531. if filepath.exists():
  532. try:
  533. sep = '\t' if filename.endswith('.tsv') else ','
  534. df = pd.read_csv(filepath, sep=sep)
  535. print(f" {db_name}: {len(df)} 条记录")
  536. except:
  537. print(f" {db_name}: 文件存在但无法读取")
  538. else:
  539. print(f" {db_name}: 文件不存在")
  540. # 下一步提示
  541. print("\n" + "="*60)
  542. print("下一步")
  543. print("="*60)
  544. if success_count < 6:
  545. print(" 需要补充的数据:")
  546. if not results["OncoKB"]:
  547. print(" - OncoKB: 申请 API Token → https://www.oncokb.org/account/register")
  548. if not results["DGIdb"]:
  549. print(" - DGIdb: 下载数据 → https://dgidb.org/downloads")
  550. if not results["OpenTargets"]:
  551. print(" - OpenTargets: 重新运行脚本 (API 免费)")
  552. if not results["CancerMine"]:
  553. print(" - CancerMine: 下载数据 → https://zenodo.org/records/7689627")
  554. print("")
  555. print(" 继续下一步 (可使用模板数据测试):")
  556. print(" python scripts/02_calculate_gene_scores.py --cancer BRCA --xai DeepLIFT")
  557. if __name__ == "__main__":
  558. main()

01_download_databases.py at commit 6d50a98, under MIT · at the source

Overview

Authors: Yiyi Zuo1, Shuting Yang2, Wenxue Zhao1
  1. Shenzhen Campus of Sun Yat-sen University, Molecular Cancer Research Center, School of Medicine, Shenzhen, China
  2. Sun Yat-sen University Sixth Affiliated Hospital, Department of Neurosurgery, Graceland Medical Center, Guangzhou, China
Journal: Frontiers in physiology, volume 17, article 1830956
Dates: received 15 March 2026; accepted 26 March 2026; published online 22 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fphys.2026.1830956 · PMID 42099920 · PMCID PMC13143651 · OpenAlex W7155191053
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), other condition (population), clinical / translational (subfield)
Methods: Statistics, Machine learning, Connectivity
Keywords: cancer, deep learning, explainable AI (XAI), survival prediction, transcriptomics
Journal subjects: Technology and Code
Topic: Explainable Artificial Intelligence (XAI) (Artificial Intelligence, Computer Science), according to OpenAlex
Funding: Basic and Applied Basic Research Foundation of Guangdong Province
Citations: cited by 2 papers (Europe PMC); 34 references in the paper

Abstract

Explainable Artificial Intelligence (XAI) holds the promise to compensate for the “black-box” nature of deep learning which impedes transcriptome-based cancer survival prediction. However, there is a lack of systematic benchmarking XAI frameworks tailored for high-dimensional survival data. To bridge this gap, we systematically evaluated six representative XAI methods in three main categories: gradient-based, propagation-based, and perturbation-based approaches by using a Self-Normalizing Neural Network (SNN) as the baseline survival model. 6,248 samples across 15 cancer types from The Cancer Genome Atlas (TCGA) was analysed in this evaluation with a unified framework we developed. The evaluation metrics encompassed three key dimensions: prognostic factor enrichment (univariate Cox regression significance), biological consistency (supported by four authoritative databases, including OpenTargets), and explanation stability (Kuncheva Index). Among the six XAI methods, we find that DeepSHAP achieved the best overall performance, identifying the highest number of statistically significant prognostic factors while maintaining superior explanation stability; LRP (Layer-wise Relevance Propagation) showed slightly lower prognostic specificity but the highest consensus with biological databases in capturing general cancer genes, making it suitable for validating biological plausibility. In contrast, the perturbation-based method, PFI (Permutation Feature Importance) exhibited systematic failure and extremely low stability due to its inability to handle feature collinearity in high-dimensional transcriptomic data. Furthermore, we identified explanation stability as a robust proxy for the biological validity of the XAI. Collectively, this study establishes an empirical framework for selecting trustworthy AI explanation tools for precision medicine.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.

ZYyli/xai-cancer-survival

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 6d50a980ab6a93475ff441b725970c9a554ef180, 1 April 2026
Languages: Python (44), Shell (1)
Size: 50 files, 45 scripts
Software Heritage: not archived
Found in: the text, “Software and environment”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: pandas (35 files), NumPy (34 files), Matplotlib (27 files), SciPy (26 files), PyTorch (23 files), statsmodels (19 files), seaborn (14 files), scikit-learn (7 files), SHAP (4 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
47 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 45 scripts, each with its path and the digest of its content;
  • 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The data presented in the study are deposited in the UCSC Xena repository. The RNA-seq datasets for the TCGA-COAD and TCGA-READ cohorts were obtained from the GDC Hub (https://xenabrowser.net/datapages/?dataset=TCGA-COAD.star_fpkm.tsv&host, https://xenabrowser.net/datapages/?dataset=TCGA-READ.star_fpkm.tsv&host). General access to other cancer datasets within the GDC hub is available at the UCSC Xena Browser (https://xenabrowser.net/datapages/). The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding authors.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 3 authors, 5 keywords, 1 funder, 30 references.

Cite

This paper

Zuo, Y., Yang, S., & Zhao, W. (2026). A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction. Frontiers in physiology, 17, 1830956. https://doi.org/10.3389/fphys.2026.1830956

BibTeX

@article{zuo2026systematic,
author = {Zuo, Yiyi and Yang, Shuting and Zhao, Wenxue},
title = {{A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction}},
journal = {Frontiers in physiology},
year = {2026},
month = apr,
volume = {17},
pages = {1830956},
publisher = {Frontiers Media SA},
issn = {1664-042X},
doi = {10.3389/fphys.2026.1830956},
url = {https://doi.org/10.3389/fphys.2026.1830956},
pmid = {42099920},
pmcid = {PMC13143651}
}

RIS

TY - JOUR
AU - Zuo, Yiyi
AU - Yang, Shuting
AU - Zhao, Wenxue
TI - A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction
T2 - Frontiers in physiology
J2 - Front Physiol
PY - 2026
DA - 2026/04/22
VL - 17
SP - 1830956
SN - 1664-042X
PB - Frontiers Media SA
DO - 10.3389/fphys.2026.1830956
UR - https://doi.org/10.3389/fphys.2026.1830956
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fphys.2026.1830956",
"type": "article-journal",
"title": "A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction",
"container-title": "Frontiers in physiology",
"author": [
{
"family": "Zuo",
"given": "Yiyi"
},
{
"family": "Yang",
"given": "Shuting"
},
{
"family": "Zhao",
"given": "Wenxue"
}
],
"container-title-short": "Front Physiol",
"volume": "17",
"page": "1830956",
"DOI": "10.3389/fphys.2026.1830956",
"PMID": "42099920",
"PMCID": "PMC13143651",
"ISSN": "1664-042X",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fphys.2026.1830956",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
22
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3389/fgene.2026.1799530 [code]
Sex-dependent prediction of autism.
Journal: Frontiers in genetics
In common: SHAP, statsmodels, seaborn, 5 other tools, genetics / omics, 1 reference
[2] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: SHAP, PyTorch, seaborn, 5 other tools, clinical / translational, 1 reference
[3] doi:10.1038/s43856-026-01606-6 [code]
Validation of remote multimodal AI screening for Parkinson disease across diverse settings.
Journal: Communications medicine
In common: SHAP, statsmodels, PyTorch, 6 other tools, clinical / translational
[4] doi:10.1016/j.isci.2026.116825 [code]
Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.
Journal: iScience
In common: SHAP, statsmodels, PyTorch, 6 other tools, genetics / omics
[5] doi:10.1186/s13059-026-04125-8 [code]
MLMarker: a machine learning framework for tissue inference and biomarker discovery.
Journal: Genome biology
In common: SHAP, statsmodels, PyTorch, 6 other tools, genetics / omics
[6] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: SHAP, statsmodels, PyTorch, 6 other tools
[7] doi:10.1523/eneuro.0362-25.2026 [code]
Similarities between &lt;i&gt;Ciona&lt;/i&gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
Journal: eNeuro
In common: SHAP, statsmodels, PyTorch, 6 other tools
[8] doi:10.1371/journal.pcbi.1014615 [code]
Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.
Journal: PLoS computational biology
In common: SHAP, statsmodels, PyTorch, 6 other tools
[9] doi:10.1093/nargab/lqag050 [code]
TSProm: deep learning framework to predict tissue-specific regulatory logic.
Journal: NAR genomics and bioinformatics
In common: SHAP, statsmodels, PyTorch, 6 other tools
[10] doi:10.1093/bioinformatics/btag213 [code]
Interpretable deep survival analysis of Alzheimer's disease via metabolic genetic variants.
Journal: Bioinformatics (Oxford, England)
In common: SHAP, PyTorch, seaborn, 5 other tools, clinical / translational, genetics / omics

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.