A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction.
The 17 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Materials and methods › Model construction ↔ operation/loss_func.py, lines 31–106 · score 0.87 · Negative Log Likelihood, deep survival prediction, Loss function, discrete survival, batch, network
- [2] § Materials and methods › Evaluation metrics ↔ biological_plausibility/scripts/01_download_databases.py, lines 560–640 · score 0.73 · OpenTargets, CancerMine, DGIdb, OncoKB, databases, Biological
- [3] § Materials and methods › Model construction ↔ operation/core_utils.py, lines 105–254 · score 0.71 · Adam, L1, Warmup, patience, epoch, optimize
- [4] § Results › DeepSHAP outperforms other XAI methods in discovery of prognostic factors from pan-cancer transcriptomic data ↔ biological_plausibility/scripts/01_download_databases.py, lines 281–324 · score 0.71 · Cell Carcinoma, Grade Glioma, Glioblastoma, Papillary, Renal, GBMLGG
- [5] § Results › LRP achieves the highest biological consistency with established cancer gene databases ↔ biological_plausibility/scripts/01_download_databases.py, lines 560–640 · score 0.70 · OpenTargets, CancerMine, DGIdb, OncoKB, database, BRCA
- [6] § Materials and methods › Dataset description ↔ datasets_csv/preprocessing_cancer_single.py, lines 58–166 · score 0.70 · Molecular Signatures Database, MSigDB, preprocessed, TCGA, censorship, genes
- [7] § Materials and methods › Software and environment ↔ biological_plausibility/scripts/01_download_databases.py, lines 450–557 · score 0.66 · Literature mined cancer, CancerMine, gene associations, Zenodo
- [8] § Materials and methods › Dataset description ↔ datasets_csv/preprocessing_no_normalization.py, lines 103–142 · score 0.65 · Molecular Signatures Database, MSigDB, preprocessed, cancer
- [9] § Results › DeepSHAP delivers the optimal comprehensive performance across prognostic, biological and stability metrics ↔ biological_plausibility/scripts/04_visualize_2.py, lines 1697–1771 · score 0.64 · DB supported Hits, Min Max normalization, XAI categories, radar, biological, Prognostic Factor
- [10] § Materials and methods › Software and environment ↔ biological_plausibility/scripts/01_download_databases.py, lines 376–447 · score 0.64 · Gene disease associations, Open Targets Platform, GraphQL, API
- [11] § Results › LRP achieves the highest biological consistency with established cancer gene databases ↔ biological_plausibility/scripts/02_calculate_gene_scores.py, lines 275–341 · score 0.63 · DGIdb, OpenTargets, CancerMine, OncoKB, BRCA, database
- [12] § Materials and methods › Interpretability framework ↔ operation/lrp_bootstrap_analysis.py, lines 18–69 · score 0.60 · AlphaDropout, bias, weighted, SELU, activations, layers
- [13] § Materials and methods › Interpretability framework ↔ operation/lrp_individual_analysis.py, lines 34–85 · score 0.60 · AlphaDropout, bias, weighted, SELU, activations, layers
- [14] § Materials and methods › Evaluation metrics ↔ biological_plausibility/scripts/02_calculate_gene_scores.py, lines 1–39 · score 0.59 · CancerMine, DGIdb, OncoKB, databases, Biological
- [15] § Results › DeepSHAP outperforms other XAI methods in discovery of prognostic factors from pan-cancer transcriptomic data ↔ biological_plausibility/scripts/database_loader.py, lines 19–37 · score 0.56 · Brain, Cell, Glioblastoma, Kidney, Papillary, Renal
- [16] § Materials and methods › Software and environment ↔ operation/command.sh, the whole file · a weak match · score 0.54 · DeepLIFT, DeepSHAP, Python, SNN, IG
- [17] § Materials and methods › Software and environment ↔ operation/evaluate_faithfulness.py, lines 335–383 · score 0.53 · GradientSHAP, DeepLIFT, DeepSHAP, IG
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 644 lines · 30 KB · MIT · 5 matches
- #!/usr/bin/env python3
- """
- Step 1: 下载并准备各数据库
- - OncoKB: 癌症基因临床分级 (需API Token)
- - CGC: Cancer Gene Census (需COSMIC账号)
- - DGIdb: 药物-基因互作 (需手动下载)
- - CIViC: 临床变异解读 (可自动下载)
- 运行方式:
- python scripts/01_download_databases.py
- python scripts/01_download_databases.py --skip_existing # 跳过已存在的文件
- """
- import os
- import sys
- import time
- import argparse
- import requests
- import pandas as pd
- from pathlib import Path
- from io import StringIO
- # 添加父目录到路径
- sys.path.insert(0, str(Path(__file__).parent.parent))
- from config import DATABASE_DIR, ONCOKB_API_TOKEN
- def check_file_valid(filepath, min_rows=5):
- """检查文件是否存在且有效(非HTML,有足够行数)"""
- if not filepath.exists():
- return False
- try:
- # 检查是否为HTML(下载失败的情况)
- with open(filepath, 'r') as f:
- first_line = f.readline()
- if first_line.strip().startswith('<!') or first_line.strip().startswith('<html'):
- return False
- # 检查行数
- df = pd.read_csv(filepath, sep='\t' if filepath.suffix == '.tsv' else ',', nrows=min_rows)
- return len(df) >= min_rows
- except:
- return False
- def write_version_file(data_file: Path, lines):
- """为数据文件生成/更新一个 *_VERSION.txt 说明文件
- 版本文件命名规则: basename + "_VERSION.txt",例如:
- - oncokb_genes.tsv -> oncokb_genes_VERSION.txt
- - opentargets_associations.tsv -> opentargets_associations_VERSION.txt
- """
- version_path = data_file.with_name(f"{data_file.stem}_VERSION.txt")
- timestamp = time.strftime("%Y-%m-%d %H:%M:%S")
- header = [f"Generated: {timestamp}"]
- content = header + list(lines)
- try:
- with open(version_path, "w", encoding="utf-8") as f:
- for line in content:
- f.write(str(line).rstrip("\n") + "\n")
- except Exception as e:
- print(f" ⚠ 写入版本说明文件失败: {version_path} ({e})")
- def download_oncokb():
- """
- 下载 OncoKB 癌症基因列表
- 注意: OncoKB 需要 API token,请先申请:
- https://www.oncokb.org/account/register
- """
- print("\n" + "="*60)
- print("下载 OncoKB 数据")
- print("="*60)
- output_file = DATABASE_DIR / "oncokb_genes.tsv"
- if ONCOKB_API_TOKEN:
- # 使用 API 获取完整数据
- headers = {"Authorization": f"Bearer {ONCOKB_API_TOKEN}"}
- url = "https://www.oncokb.org/api/v1/utils/allCuratedGenes"
- try:
- response = requests.get(url, headers=headers)
- response.raise_for_status()
- genes = response.json()
- df = pd.DataFrame(genes)
- df.to_csv(output_file, sep="\t", index=False)
- print(f"✓ 已保存 {len(df)} 个 OncoKB 基因到 {output_file}")
- write_version_file(output_file, [
- f"Source: {url}",
- "Description: OncoKB curated genes downloaded via official API",
- ])
- return True
- except Exception as e:
- print(f"✗ OncoKB API 请求失败: {e}")
- # 如果没有 API token,创建模板文件
- print("⚠ 未设置 ONCOKB_API_TOKEN 环境变量")
- print(" 请手动下载数据或设置环境变量后重试")
- print(" 申请地址: https://www.oncokb.org/account/register")
- # 创建模板文件,包含常见癌症基因
- template_genes = [
- {"hugoSymbol": "TP53", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
- {"hugoSymbol": "EGFR", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "BRAF", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "KRAS", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "PIK3CA", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "ERBB2", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "ALK", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "ROS1", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "MET", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "RET", "highestSensitiveLevel": "1", "oncogene": True, "tsg": False},
- {"hugoSymbol": "BRCA1", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
- {"hugoSymbol": "BRCA2", "highestSensitiveLevel": "1", "oncogene": False, "tsg": True},
- {"hugoSymbol": "ATM", "highestSensitiveLevel": "2", "oncogene": False, "tsg": True},
- {"hugoSymbol": "PTEN", "highestSensitiveLevel": "2", "oncogene": False, "tsg": True},
- {"hugoSymbol": "AKT1", "highestSensitiveLevel": "2", "oncogene": True, "tsg": False},
- ]
- print("\n ⚠ 创建模板文件(仅包含示例基因,请替换为完整数据)")
- df = pd.DataFrame(template_genes)
- df.to_csv(output_file, sep="\t", index=False)
- print(f" 模板已保存到 {output_file}")
- write_version_file(output_file, [
- "Source: local template (no ONCOKB_API_TOKEN)",
- "Description: example curated genes only, please replace with full OncoKB data",
- "Manual action: download full data from https://www.oncokb.org/ and overwrite this file",
- ])
- return False
- def download_dgidb():
- """
- 下载 DGIdb 药物-基因互作数据
- 注意: DGIdb 网站已改版 (2024),旧API失效,需手动下载
- 下载地址: https://dgidb.org/downloads
- """
- print("\n" + "="*60)
- print("下载 DGIdb 数据")
- print("="*60)
- output_file = DATABASE_DIR / "dgidb_interactions.tsv"
- # 检查是否已有有效文件
- if check_file_valid(output_file, min_rows=100):
- print(f" ✓ 已存在有效数据文件: {output_file}")
- df = pd.read_csv(output_file, sep="\t")
- print(f" 包含 {len(df)} 条记录")
- # 保留用户手动下载版本,不覆盖版本说明,只在缺失时创建一个简短说明
- if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
- write_version_file(output_file, [
- "Source: user-provided DGIdb interactions.tsv",
- "Description: manually downloaded from https://dgidb.org/downloads",
- ])
- return True
- # DGIdb 网站已改版,API不可用,需要手动下载
- print(" ⚠ DGIdb 网站已改版,需要手动下载数据")
- print("")
- print(" 【手动下载步骤】:")
- print(" 1. 浏览器访问: https://dgidb.org/downloads")
- print(" 2. 找到并下载: interactions.tsv")
- print(" 3. 将文件保存到:")
- print(f" {output_file}")
- print("")
- # 创建扩展模板(包含常见 FDA 批准靶点药物)
- print(" 创建临时模板文件(包含常见靶点药物)...")
- template = [
- # EGFR 抑制剂
- {"gene_name": "EGFR", "drug_name": "Erlotinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "EGFR", "drug_name": "Gefitinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "EGFR", "drug_name": "Osimertinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "EGFR", "drug_name": "Afatinib", "interaction_types": "inhibitor", "approved": True},
- # BRAF 抑制剂
- {"gene_name": "BRAF", "drug_name": "Vemurafenib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRAF", "drug_name": "Dabrafenib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRAF", "drug_name": "Encorafenib", "interaction_types": "inhibitor", "approved": True},
- # HER2/ERBB2 抑制剂
- {"gene_name": "ERBB2", "drug_name": "Trastuzumab", "interaction_types": "antibody", "approved": True},
- {"gene_name": "ERBB2", "drug_name": "Pertuzumab", "interaction_types": "antibody", "approved": True},
- {"gene_name": "ERBB2", "drug_name": "Lapatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ERBB2", "drug_name": "Neratinib", "interaction_types": "inhibitor", "approved": True},
- # ALK 抑制剂
- {"gene_name": "ALK", "drug_name": "Crizotinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ALK", "drug_name": "Alectinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ALK", "drug_name": "Brigatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ALK", "drug_name": "Lorlatinib", "interaction_types": "inhibitor", "approved": True},
- # PARP 抑制剂 (BRCA)
- {"gene_name": "BRCA1", "drug_name": "Olaparib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRCA2", "drug_name": "Olaparib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRCA1", "drug_name": "Rucaparib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRCA2", "drug_name": "Rucaparib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRCA1", "drug_name": "Niraparib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BRCA2", "drug_name": "Niraparib", "interaction_types": "inhibitor", "approved": True},
- # PIK3CA
- {"gene_name": "PIK3CA", "drug_name": "Alpelisib", "interaction_types": "inhibitor", "approved": True},
- # KRAS
- {"gene_name": "KRAS", "drug_name": "Sotorasib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "KRAS", "drug_name": "Adagrasib", "interaction_types": "inhibitor", "approved": True},
- # RET
- {"gene_name": "RET", "drug_name": "Selpercatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "RET", "drug_name": "Pralsetinib", "interaction_types": "inhibitor", "approved": True},
- # MET
- {"gene_name": "MET", "drug_name": "Capmatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "MET", "drug_name": "Tepotinib", "interaction_types": "inhibitor", "approved": True},
- # ROS1
- {"gene_name": "ROS1", "drug_name": "Crizotinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ROS1", "drug_name": "Entrectinib", "interaction_types": "inhibitor", "approved": True},
- # NTRK
- {"gene_name": "NTRK1", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "NTRK2", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "NTRK3", "drug_name": "Larotrectinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "NTRK1", "drug_name": "Entrectinib", "interaction_types": "inhibitor", "approved": True},
- # FGFR
- {"gene_name": "FGFR2", "drug_name": "Pemigatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "FGFR2", "drug_name": "Erdafitinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "FGFR3", "drug_name": "Erdafitinib", "interaction_types": "inhibitor", "approved": True},
- # IDH
- {"gene_name": "IDH1", "drug_name": "Ivosidenib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "IDH2", "drug_name": "Enasidenib", "interaction_types": "inhibitor", "approved": True},
- # BCR-ABL
- {"gene_name": "BCR", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ABL1", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ABL1", "drug_name": "Dasatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "ABL1", "drug_name": "Nilotinib", "interaction_types": "inhibitor", "approved": True},
- # KIT
- {"gene_name": "KIT", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "KIT", "drug_name": "Sunitinib", "interaction_types": "inhibitor", "approved": True},
- # PDGFRA
- {"gene_name": "PDGFRA", "drug_name": "Imatinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "PDGFRA", "drug_name": "Avapritinib", "interaction_types": "inhibitor", "approved": True},
- # FLT3
- {"gene_name": "FLT3", "drug_name": "Midostaurin", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "FLT3", "drug_name": "Gilteritinib", "interaction_types": "inhibitor", "approved": True},
- # BTK
- {"gene_name": "BTK", "drug_name": "Ibrutinib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "BTK", "drug_name": "Acalabrutinib", "interaction_types": "inhibitor", "approved": True},
- # BCL2
- {"gene_name": "BCL2", "drug_name": "Venetoclax", "interaction_types": "inhibitor", "approved": True},
- # VEGFR
- {"gene_name": "KDR", "drug_name": "Bevacizumab", "interaction_types": "antibody", "approved": True},
- {"gene_name": "KDR", "drug_name": "Sorafenib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "KDR", "drug_name": "Sunitinib", "interaction_types": "inhibitor", "approved": True},
- # ESR1 (雌激素受体)
- {"gene_name": "ESR1", "drug_name": "Tamoxifen", "interaction_types": "antagonist", "approved": True},
- {"gene_name": "ESR1", "drug_name": "Fulvestrant", "interaction_types": "antagonist", "approved": True},
- # AR (雄激素受体)
- {"gene_name": "AR", "drug_name": "Enzalutamide", "interaction_types": "antagonist", "approved": True},
- {"gene_name": "AR", "drug_name": "Abiraterone", "interaction_types": "inhibitor", "approved": True},
- # CDK4/6
- {"gene_name": "CDK4", "drug_name": "Palbociclib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "CDK6", "drug_name": "Palbociclib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "CDK4", "drug_name": "Ribociclib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "CDK6", "drug_name": "Ribociclib", "interaction_types": "inhibitor", "approved": True},
- # MTOR
- {"gene_name": "MTOR", "drug_name": "Everolimus", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "MTOR", "drug_name": "Temsirolimus", "interaction_types": "inhibitor", "approved": True},
- # SMO (Hedgehog)
- {"gene_name": "SMO", "drug_name": "Vismodegib", "interaction_types": "inhibitor", "approved": True},
- {"gene_name": "SMO", "drug_name": "Sonidegib", "interaction_types": "inhibitor", "approved": True},
- ]
- df = pd.DataFrame(template)
- df.to_csv(output_file, sep="\t", index=False)
- print(f" ✓ 模板已保存: {output_file}")
- write_version_file(output_file, [
- "Source: local template (DGIdb manual download required)",
- "Description: contains common FDA-approved targeted agents only, please replace with full DGIdb interactions.tsv",
- "Manual action: download interactions.tsv from https://dgidb.org/downloads and overwrite this file",
- ])
- print(f" 包含 {len(template)} 条 FDA 批准靶向药物记录")
- print(" ⚠ 请尽快下载完整数据替换此模板")
- return False
- def download_opentargets():
- """
- 下载 Open Targets Platform 基因-疾病关联数据
- Open Targets API 免费无需注册
- 为每个 TCGA 癌症类型下载相关基因列表
- """
- print("\n" + "="*60)
- print("下载 Open Targets 数据")
- print("="*60)
- output_file = DATABASE_DIR / "opentargets_associations.tsv"
- # 检查是否已有有效文件
- if check_file_valid(output_file, min_rows=100):
- print(f" ✓ 已存在有效数据文件: {output_file}")
- df = pd.read_csv(output_file, sep="\t")
- print(f" 包含 {len(df)} 条记录")
- if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
- write_version_file(output_file, [
- "Source: user-provided or previously downloaded Open Targets associations",
- "Description: gene-disease associations for TCGA cancer types via EFO/MONDO IDs",
- "URL: https://api.platform.opentargets.org/api/v4/graphql",
- ])
- return True
- # TCGA 癌种到 Open Targets EFO/MONDO ID 的映射 (已通过 API 验证)
- TCGA_TO_EFO = {
- "BLCA": ("MONDO_0001187", "bladder carcinoma"),
- "BRCA": ("EFO_0000305", "breast carcinoma"),
- "COADREAD": ("MONDO_0005575", "colorectal cancer"),
- "GBMLGG": ("EFO_0000519", "glioblastoma"),
- "HNSC": ("EFO_0000181", "head and neck squamous cell carcinoma"),
- "KIRC": ("EFO_0000349", "renal cell carcinoma"),
- "KIRP": ("EFO_0000640", "papillary renal cell carcinoma"),
- "LGG": ("EFO_0005543", "low grade glioma"),
- "LIHC": ("EFO_0000182", "hepatocellular carcinoma"),
- "LUAD": ("EFO_0000571", "lung adenocarcinoma"),
- "LUSC": ("EFO_0000708", "lung squamous cell carcinoma"),
- "PAAD": ("EFO_1000044", "pancreatic adenocarcinoma"),
- "SKCM": ("EFO_0000389", "melanoma"),
- "STAD": ("EFO_0000503", "stomach carcinoma"),
- "UCEC": ("EFO_1001512", "endometrial carcinoma"),
- }
- # API 分页限制
- MAX_PAGE_SIZE = 2500 # API 最大允许 3000,保守设置 2500
- GRAPHQL_URL = "https://api.platform.opentargets.org/api/v4/graphql"
- # GraphQL 查询 - 获取疾病关联的所有靶点
- query = """
- query DiseaseAssociations($efoId: String!, $size: Int!) {
- disease(efoId: $efoId) {
- id
- name
- associatedTargets(page: { index: 0, size: $size }) {
- count
- rows {
- target {
- id
- approvedSymbol
- }
- score
- }
- }
- }
- }
- """
- all_associations = []
- print(f" 将下载 {len(TCGA_TO_EFO)} 种癌症类型的关联数据...")
- for tcga_code, (efo_id, disease_name) in TCGA_TO_EFO.items():
- print(f" {tcga_code} ({disease_name})...", end=" ", flush=True)
- try:
- # 分页下载关联数据
- page_index = 0
- total_rows = 0
- while True:
- response = requests.post(
- GRAPHQL_URL,
- json={
- "query": query,
- "variables": {"efoId": efo_id, "size": MAX_PAGE_SIZE}
- },
- timeout=120,
- headers={"Content-Type": "application/json"}
- )
- response.raise_for_status()
- data = response.json()
- # 检查是否有错误
- if "errors" in data:
- error_msg = data['errors'][0].get('message', '')[:100]
- print(f"API错误: {error_msg}")
- break
- # 安全解析
- disease_data = data.get("data", {})
- if disease_data is None:
- print("data 为空")
- break
- disease_info = disease_data.get("disease")
- if disease_info is None:
- print(f"未找到疾病 {efo_id}")
- break
- associated = disease_info.get("associatedTargets")
- if associated is None:
- print("无关联靶点")
- break
- rows = associated.get("rows", [])
- if not rows:
- break
- for row in rows:
- target = row.get("target", {}) or {}
- gene_symbol = target.get("approvedSymbol", "")
- if gene_symbol:
- all_associations.append({
- "tcga_code": tcga_code,
- "efo_id": efo_id,
- "disease_name": disease_name,
- "gene_symbol": gene_symbol,
- "ensembl_id": target.get("id", ""),
- "association_score": row.get("score", 0),
- })
- total_rows += len(rows)
- # 只取第一页 (2500 个基因已经足够覆盖 Top 100)
- break
- if total_rows > 0:
- print(f"{total_rows} 个基因")
- # 避免 API 限流
- time.sleep(0.3)
- except requests.exceptions.RequestException as e:
- print(f"网络错误: {e}")
- except Exception as e:
- print(f"失败: {type(e).__name__}: {e}")
- if all_associations:
- df = pd.DataFrame(all_associations)
- df.to_csv(output_file, sep="\t", index=False)
- print(f"\n ✓ 已保存 {len(df)} 条关联记录到 {output_file}")
- # 统计
- gene_count = df['gene_symbol'].nunique()
- print(f" 涵盖 {gene_count} 个独立基因")
- write_version_file(output_file, [
- "Source: Open Targets Platform GraphQL API",
- "Description: gene-disease associations for TCGA cancer types (first page up to 2500 targets per disease)",
- "URL: https://api.platform.opentargets.org/api/v4/graphql",
- ])
- return True
- else:
- print(" ✗ 未能下载任何数据")
- return False
- def download_cancermine():
- """
- 下载 CancerMine 文献挖掘数据
- CancerMine 是基于文献挖掘的癌症基因数据库,包含三种角色:
- - Driver: 驱动基因
- - Oncogene: 癌基因
- - Tumor_Suppressor: 抑癌基因
- 数据下载地址: https://zenodo.org/records/7689627
- (原网站 http://bionlp.bcgsc.ca/cancermine/ 已下线)
- """
- print("\n" + "="*60)
- print("下载 CancerMine 数据")
- print("="*60)
- output_file = DATABASE_DIR / "cancermine.tsv"
- # 检查是否已有有效文件
- if check_file_valid(output_file, min_rows=100):
- print(f" ✓ 已存在有效数据文件: {output_file}")
- df = pd.read_csv(output_file, sep="\t")
- print(f" 包含 {len(df)} 条记录")
- if not (output_file.parent / f"{output_file.stem}_VERSION.txt").exists():
- write_version_file(output_file, [
- "Source: user-provided or previously downloaded CancerMine data",
- "Description: literature-mined cancer gene associations (Driver/Oncogene/Tumor_Suppressor)",
- "URL: https://zenodo.org/records/7689627",
- ])
- return True
- # CancerMine 公开下载链接 (Zenodo - 原网站已下线)
- url = "https://zenodo.org/records/7689627/files/cancermine_collated.tsv?download=1"
- try:
- print(f" 下载: {url}")
- response = requests.get(url, timeout=300)
- response.raise_for_status()
- # 检查是否为有效TSV
- if response.text.strip().startswith('<!') or '<html' in response.text[:100].lower():
- print(" ✗ 返回HTML而非数据")
- raise Exception("Invalid response")
- df = pd.read_csv(StringIO(response.text), sep="\t")
- if len(df) > 100:
- df.to_csv(output_file, sep="\t", index=False)
- print(f" ✓ 已保存 {len(df)} 条 CancerMine 记录到 {output_file}")
- # 统计角色分布
- if 'role' in df.columns:
- role_counts = df['role'].value_counts()
- print(f" 角色分布:")
- for role, count in role_counts.items():
- print(f" - {role}: {count}")
- # 统计基因和癌症数量
- gene_col = 'gene_normalized' if 'gene_normalized' in df.columns else 'gene'
- cancer_col = 'cancer_normalized' if 'cancer_normalized' in df.columns else 'cancer'
- if gene_col in df.columns:
- print(f" 独立基因数: {df[gene_col].nunique()}")
- if cancer_col in df.columns:
- print(f" 独立癌症类型数: {df[cancer_col].nunique()}")
- write_version_file(output_file, [
- f"Source: {url}",
- "Description: CancerMine collated data - literature-mined cancer gene associations",
- "Roles: Driver, Oncogene, Tumor_Suppressor",
- f"Total records: {len(df)}",
- ])
- return True
- else:
- print(" ✗ 数据行数不足")
- except Exception as e:
- print(f" ✗ 下载失败: {e}")
- # 下载失败,创建模板
- print(" ⚠ 自动下载失败,创建模板文件...")
- print(" 请手动下载: https://zenodo.org/records/7689627")
- print(" 选择 cancermine_collated.tsv 并保存为 databases/cancermine.tsv")
- template = [
- {"gene_normalized": "TP53", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 100},
- {"gene_normalized": "BRCA1", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 80},
- {"gene_normalized": "BRCA2", "cancer_normalized": "breast cancer", "role": "Tumor_Suppressor", "citation_count": 70},
- {"gene_normalized": "EGFR", "cancer_normalized": "lung cancer", "role": "Oncogene", "citation_count": 90},
- {"gene_normalized": "KRAS", "cancer_normalized": "colorectal cancer", "role": "Driver", "citation_count": 85},
- {"gene_normalized": "BRAF", "cancer_normalized": "melanoma", "role": "Oncogene", "citation_count": 75},
- {"gene_normalized": "PIK3CA", "cancer_normalized": "breast cancer", "role": "Oncogene", "citation_count": 60},
- {"gene_normalized": "PTEN", "cancer_normalized": "prostate cancer", "role": "Tumor_Suppressor", "citation_count": 55},
- {"gene_normalized": "APC", "cancer_normalized": "colorectal cancer", "role": "Tumor_Suppressor", "citation_count": 50},
- {"gene_normalized": "IDH1", "cancer_normalized": "glioma", "role": "Driver", "citation_count": 45},
- {"gene_normalized": "ERBB2", "cancer_normalized": "breast cancer", "role": "Oncogene", "citation_count": 65},
- {"gene_normalized": "MYC", "cancer_normalized": "lymphoma", "role": "Oncogene", "citation_count": 70},
- {"gene_normalized": "RB1", "cancer_normalized": "retinoblastoma", "role": "Tumor_Suppressor", "citation_count": 40},
- {"gene_normalized": "VHL", "cancer_normalized": "kidney cancer", "role": "Tumor_Suppressor", "citation_count": 35},
- {"gene_normalized": "ALK", "cancer_normalized": "lung cancer", "role": "Oncogene", "citation_count": 55},
- ]
- df = pd.DataFrame(template)
- df.to_csv(output_file, sep="\t", index=False)
- print(f" 模板已保存到 {output_file}")
- write_version_file(output_file, [
- "Source: local template (CancerMine manual download required)",
- "Description: example cancer gene associations only, please replace with full CancerMine data",
- "Manual action: download cancermine_collated.tsv from http://bionlp.bcgsc.ca/cancermine/ and overwrite this file",
- ])
- return False
- def main():
- parser = argparse.ArgumentParser(description="下载生物学数据库")
- parser.add_argument('--skip_existing', action='store_true',
- help='跳过已存在的有效文件')
- args = parser.parse_args()
- print("="*60)
- print("XAI 生物学合理性评估 - 数据库下载")
- print("="*60)
- # 确保目录存在
- DATABASE_DIR.mkdir(parents=True, exist_ok=True)
- print(f"数据库目录: {DATABASE_DIR}")
- # 下载各数据库
- results = {
- "OncoKB": download_oncokb(),
- "DGIdb": download_dgidb(),
- "OpenTargets": download_opentargets(),
- "CancerMine": download_cancermine()
- }
- # 汇总
- print("\n" + "="*60)
- print("下载汇总")
- print("="*60)
- success_count = 0
- for db, success in results.items():
- if success:
- status = "✓ 完整数据"
- success_count += 1
- else:
- status = "⚠ 模板数据 (需替换)"
- print(f" {db}: {status}")
- print(f"\n 完整数据: {success_count}/6")
- # 检查文件大小
- print("\n" + "="*60)
- print("文件详情")
- print("="*60)
- files_info = [
- ("oncokb_genes.tsv", "OncoKB"),
- ("dgidb_interactions.tsv", "DGIdb"),
- ("opentargets_associations.tsv", "OpenTargets"),
- ("cancermine.tsv", "CancerMine"),
- ]
- for filename, db_name in files_info:
- filepath = DATABASE_DIR / filename
- if filepath.exists():
- try:
- sep = '\t' if filename.endswith('.tsv') else ','
- df = pd.read_csv(filepath, sep=sep)
- print(f" {db_name}: {len(df)} 条记录")
- except:
- print(f" {db_name}: 文件存在但无法读取")
- else:
- print(f" {db_name}: 文件不存在")
- # 下一步提示
- print("\n" + "="*60)
- print("下一步")
- print("="*60)
- if success_count < 6:
- print(" 需要补充的数据:")
- if not results["OncoKB"]:
- print(" - OncoKB: 申请 API Token → https://www.oncokb.org/account/register")
- if not results["DGIdb"]:
- print(" - DGIdb: 下载数据 → https://dgidb.org/downloads")
- if not results["OpenTargets"]:
- print(" - OpenTargets: 重新运行脚本 (API 免费)")
- if not results["CancerMine"]:
- print(" - CancerMine: 下载数据 → https://zenodo.org/records/7689627")
- print("")
- print(" 继续下一步 (可使用模板数据测试):")
- print(" python scripts/02_calculate_gene_scores.py --cancer BRCA --xai DeepLIFT")
- if __name__ == "__main__":
- main()
01_download_databases.py at commit 6d50a98, under MIT · at the source
Overview
- Shenzhen Campus of Sun Yat-sen University, Molecular Cancer Research Center, School of Medicine, Shenzhen, China
- Sun Yat-sen University Sixth Affiliated Hospital, Department of Neurosurgery, Graceland Medical Center, Guangzhou, China
Abstract
Explainable Artificial Intelligence (XAI) holds the promise to compensate for the “black-box” nature of deep learning which impedes transcriptome-based cancer survival prediction. However, there is a lack of systematic benchmarking XAI frameworks tailored for high-dimensional survival data. To bridge this gap, we systematically evaluated six representative XAI methods in three main categories: gradient-based, propagation-based, and perturbation-based approaches by using a Self-Normalizing Neural Network (SNN) as the baseline survival model. 6,248 samples across 15 cancer types from The Cancer Genome Atlas (TCGA) was analysed in this evaluation with a unified framework we developed. The evaluation metrics encompassed three key dimensions: prognostic factor enrichment (univariate Cox regression significance), biological consistency (supported by four authoritative databases, including OpenTargets), and explanation stability (Kuncheva Index). Among the six XAI methods, we find that DeepSHAP achieved the best overall performance, identifying the highest number of statistically significant prognostic factors while maintaining superior explanation stability; LRP (Layer-wise Relevance Propagation) showed slightly lower prognostic specificity but the highest consensus with biological databases in capturing general cancer genes, making it suitable for validating biological plausibility. In contrast, the perturbation-based method, PFI (Permutation Feature Importance) exhibited systematic failure and extremely low stability due to its inability to handle feature collinearity in high-dimensional transcriptomic data. Furthermore, we identified explanation stability as a robust proxy for the biological validity of the XAI. Collectively, this study establishes an empirical framework for selecting trustworthy AI explanation tools for precision medicine.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.
ZYyli/xai-cancer-survival
6d50a980ab6a93475ff441b725970c9a554ef180, 1 April 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
47 files
- biological_plausibility/
config.py , Python, 81 lines - biological_plausibility/
scripts/ , Python, 644 lines, 5 matches01_download_databases.py - biological_plausibility/
scripts/ , Python, 344 lines, 2 matches02_calculate_gene_scores .py - biological_plausibility/
scripts/ , Python, 1,160 lines03_visualize_and_test.py - biological_plausibility/
scripts/ , Python, 1,774 lines, 1 match04_visualize_2.py - biological_plausibility/
scripts/ , Python, 1 line__init__.py - biological_plausibility/
scripts/ , Python, 369 lines, 1 matchdatabase_loader.py - datasets_csv/
preprocessing_cancer_sin , Python, 185 lines, 1 matchgle.py - datasets_csv/
preprocessing_no_normali , Python, 231 lines, 1 matchzation.py - datasets_csv/
split.py , Python, 37 lines - operation/
bootstrap_boxplot_analys , Python, 346 linesis.py - operation/
bootstrap_snn_evaluation , Python, 587 lines.py - operation/
boxplot_prognotic.py , Python, 865 lines - operation/
command.sh , Shell, 66 lines, 1 match - operation/
core_utils.py , Python, 475 lines, 1 match - operation/
corr_stability_cindex_he , Python, 690 linesatmap.py - operation/
corr_xai_cindex_heatmap. , Python, 251 linespy - operation/
dataset_survival.py , Python, 43 lines - operation/
deepLIFT_bootstrap_analy , Python, 876 linessis.py - operation/
deepLIFT_individual_anal , Python, 921 linesysis.py - operation/
deepshap_bootstrap_analy , Python, 894 linessis.py - operation/
deepshap_individual_anal , Python, 932 linesysis.py - operation/
evaluate_faithfulness.py , Python, 636 lines, 1 match - operation/
evaluate_plots_nested_cv , Python, 1,161 lines.py - operation/
feature_stability_analys , Python, 1,288 linesis.py - operation/
feature_stability_analys , Python, 1,364 linesis_bootstrap.py - operation/
file_utils.py , Python, 12 lines - operation/
ig_bootstrap_analysis.py , Python, 945 lines - operation/
ig_individual_analysis.p , Python, 922 linesy - operation/
knn_cpi_bootstrap_analys , Python, 626 linesis.py - operation/
knn_cpi_individual_analy , Python, 588 linessis.py - operation/
loss_func.py , Python, 106 lines, 1 match - operation/
lrp_bootstrap_analysis.p , Python, 912 lines, 1 matchy - operation/
lrp_individual_analysis. , Python, 966 lines, 1 matchpy - operation/
main.py , Python, 413 lines - operation/
model_genomic.py , Python, 54 lines - operation/
model_risk.py , Python, 78 lines - operation/
pfi_bootstrap_analysis.p , Python, 866 linesy - operation/
pfi_individual_analysis. , Python, 978 linespy - operation/
plot_cindex_correlation. , Python, 65 linespy - operation/
run_SNN.py , Python, 49 lines - operation/
shap_bootstrap_analysis. , Python, 871 linespy - operation/
shap_individual_analysis , Python, 925 lines.py - operation/
stability_visualization. , Python, 1,287 linespy - operation/
utils.py , Python, 58 lines - LICENSE, License, 21 lines
- README.md, Text, 35 lines
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 45 scripts, each with its path and the digest of its content;
- 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability statement
The data presented in the study are deposited in the UCSC Xena repository. The RNA-seq datasets for the TCGA-COAD and TCGA-READ cohorts were obtained from the GDC Hub (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 3 authors, 5 keywords, 1 funder, 30 references.
Cite
This paper
Zuo, Y., Yang, S., & Zhao, W. (2026). A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction. Frontiers in physiology, 17, 1830956. https://
BibTeX
@article{zuo2026systemat
author = {Zuo, Yiyi and Yang, Shuting and Zhao, Wenxue},
title = {{A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction}},
journal = {Frontiers in physiology},
year = {2026},
month = apr,
volume = {17},
pages = {1830956},
publisher = {Frontiers Media SA},
issn = {1664-042X},
doi = {10.3389/
url = {https://
pmid = {42099920},
pmcid = {PMC13143651}
}
RIS
TY - JOUR
AU - Zuo, Yiyi
AU - Yang, Shuting
AU - Zhao, Wenxue
TI - A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction
T2 - Frontiers in physiology
J2 - Front Physiol
PY - 2026
DA - 2026/
VL - 17
SP - 1830956
SN - 1664-042X
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "A systematic evaluation of explainable AI methods for high-dimensional transcriptome-based cancer survival prediction",
"container-title": "Frontiers in physiology",
"author": [
{
"family": "Zuo",
"given": "Yiyi"
},
{
"family": "Yang",
"given": "Shuting"
},
{
"family": "Zhao",
"given": "Wenxue"
}
],
"container-title-short":
"volume": "17",
"page": "1830956",
"DOI": "10.3389/
"PMID": "42099920",
"PMCID": "PMC13143651",
"ISSN": "1664-042X",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
22
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.3389/fgene.2026.1799530 [code]
- Sex-dependent prediction of autism.Journal: Frontiers in geneticsIn common: SHAP, statsmodels, seaborn, 5 other tools, genetics / omics, 1 reference
- [2] doi:10.1038/s41467-026-71555-0 [code]
- A deep representation learning model to predict response to vagus nerve stimulation.Journal: Nature communicationsIn common: SHAP, PyTorch, seaborn, 5 other tools, clinical / translational, 1 reference
- [3] doi:10.1038/s43856-026-01606-6 [code]
- Validation of remote multimodal AI screening for Parkinson disease across diverse settings.Journal: Communications medicineIn common: SHAP, statsmodels, PyTorch, 6 other tools, clinical / translational
- [4] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: SHAP, statsmodels, PyTorch, 6 other tools, genetics / omics
- [5] doi:10.1186/s13059-026-04125-8 [code]
- MLMarker: a machine learning framework for tissue inference and biomarker discovery.Journal: Genome biologyIn common: SHAP, statsmodels, PyTorch, 6 other tools, genetics / omics
- [6] doi:10.1016/j.isci.2026.115329 [code]
- Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.Journal: iScienceIn common: SHAP, statsmodels, seaborn, 5 other tools, clinical / translational, genetics / omics, other condition
- [7] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: SHAP, statsmodels, PyTorch, 6 other tools
- [8] doi:10.1523/eneuro.0362-25.2026 [code]
- Similarities between &
lt;i& gt;Ciona& lt;/ i& gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/ V Subdivision within a Hindbrain Precursor? Journal: eNeuroIn common: SHAP, statsmodels, PyTorch, 6 other tools - [9] doi:10.1371/journal.pcbi.1014615 [code]
- Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.Journal: PLoS computational biologyIn common: SHAP, statsmodels, PyTorch, 6 other tools
- [10] doi:10.1093/nargab/lqag050 [code]
- TSProm: deep learning framework to predict tissue-specific regulatory logic.Journal: NAR genomics and bioinformaticsIn common: SHAP, statsmodels, PyTorch, 6 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 45 scripts, and 17 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:5c9af88a3054f4ac…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
