OSCR

SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § 2 Materials and methods › 2.1 The SIMLINK method › 2.1.2 Model architecture ↔ layers.py, lines 235–268 · score 0.89 · TransH, QuatE, RotatE, TransE, DistMult, TransD
  2. [2] § 3 Results and discussion › 3.1 Comparison with state-of-the-art pathogenicity prediction methods ↔ utils.py, lines 241–289 · score 0.84 · PrDSM, SilVA, TraP, fathmm MKL, usDSM, DANN
  3. [3] § 3 Results and discussion › 3.1 Comparison with state-of-the-art pathogenicity prediction methods ↔ utils.py, lines 241–289 · score 0.51 · fathmm MKL, Eigen, DANN, MVP, CADD, error

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 347 lines · 13 KB · no license · 2 matches

  1. import numpy as np
  2. import pickle as pkl
  3. import scipy.sparse as sp
  4. import sys
  5. import tensorflow as tf
  6. import math
  7. import os
  8. import random
  9. from collections import Counter
  10. import logging
  11. import pandas as pd
  12. import shutil
  13. from sklearn.impute import SimpleImputer
  14. from sklearn.linear_model import LinearRegression
  15. from sklearn.metrics import mean_squared_error, r2_score, accuracy_score, roc_auc_score
  16. import joblib
  17. import argparse
  18. flags = tf.app.flags
  19. FLAGS = flags.FLAGS
  20. def create_exp_dir(path, scripts_to_save=None):
  21. path_split = path.split("/")
  22. path_i = "."
  23. for one_path in path_split:
  24. path_i += "/" + one_path
  25. if not os.path.exists(path_i):
  26. os.mkdir(path_i)
  27. print('Experiment dir : {}'.format(path_i))
  28. if scripts_to_save is not None:
  29. os.mkdir(os.path.join(path, 'scripts'))
  30. for script in scripts_to_save:
  31. dst_file = os.path.join(path, 'scripts', os.path.basename(script))
  32. shutil.copyfile(script, dst_file)
  33. def inverse_sum(adj):
  34. adj = sp.coo_matrix(adj)
  35. rowsum = np.array(adj.sum(1))
  36. d_inv_sqrt = np.power(rowsum, -1).flatten()
  37. d_inv_sqrt[np.isinf(d_inv_sqrt)] = 0.
  38. return d_inv_sqrt.reshape((-1, 1))
  39. def preprocess_adj(adj):
  40. ent_adj_invsum = inverse_sum(adj[0])
  41. rel_adj_invsum = inverse_sum(adj[1])
  42. return [ent_adj_invsum, rel_adj_invsum, adj[2]]
  43. def construct_feed_dict(features, support, placeholders):
  44. feed_dict = dict()
  45. feed_dict.update({placeholders['features']: features})
  46. if isinstance(support[0], list):
  47. for i in range(len(support)):
  48. feed_dict.update({placeholders['support'][i][j]: support[i][j] \
  49. for j in range(len(support[i]))})
  50. else:
  51. feed_dict.update({placeholders['support'][i]: support[i] \
  52. for i in range(len(support))})
  53. return feed_dict
  54. def loadfile(file, num=1):
  55. '''
  56. num: number of elements per row
  57. '''
  58. print('loading file ' + file)
  59. ret = []
  60. with open(file, "r", encoding='utf-8') as rf:
  61. for line in rf:
  62. th = line[:-1].split('\t')
  63. x = []
  64. for i in range(num):
  65. x.append(int(th[i]))
  66. ret.append(tuple(x))
  67. return ret
  68. def get_ent2id(files):
  69. ent2id = {}
  70. for file in files:
  71. with open(file, 'r', encoding='utf-8') as rf:
  72. for line in rf:
  73. th = line[:-1].split('\t')
  74. ent2id[th[1]] = int(th[0])
  75. return ent2id
  76. def get_extended_adj_auto(e, KG):
  77. nei_list = []
  78. ent_row, rel_row = [], []
  79. ent_col, rel_col = [], []
  80. ent_data, rel_data = [], []
  81. count = 0
  82. for tri in KG:
  83. nei_list.append([tri[0], tri[1], tri[2]])
  84. ent_row.append(tri[0])
  85. ent_col.append(count)
  86. ent_data.append(1.)
  87. ent_row.append(tri[2])
  88. ent_col.append(count)
  89. ent_data.append(1.)
  90. rel_row.append(tri[1])
  91. rel_col.append(count)
  92. rel_data.append(1.)
  93. count += 1
  94. ent_adj_ind = sp.coo_matrix((ent_data, (ent_row, ent_col)), shape=(e, count))
  95. rel_adj_ind = sp.coo_matrix((rel_data, (rel_row, rel_col)), shape=(max(rel_row)+1, count))
  96. return [ent_adj_ind, rel_adj_ind, np.array(nei_list)]
  97. def load_data_class(FLAGS):
  98. def analysis(A, y, train, test):
  99. for A_i in A:
  100. print(A_i.nonzero())
  101. exit()
  102. def to_KG(A):
  103. KG = []
  104. count = 0
  105. for A_i in A:
  106. idx = A_i.nonzero()
  107. for head, tail in zip(idx[0], idx[1]):
  108. KG.append([head, count, tail])
  109. if len(idx[0]) > 0:
  110. count += 1
  111. # print(KG[:100])
  112. return KG
  113. dirname = os.path.dirname(os.path.realpath(sys.argv[0]))
  114. raw_file = dirname + '/kgdata/class/' + FLAGS.dataset + '.pickle'
  115. pro_file = dirname + '/kgdata/class/' + FLAGS.dataset + 'pro.pickle'
  116. if not os.path.exists(pro_file):
  117. with open(dirname + '/kgdata/class/' + FLAGS.dataset + '.pickle', 'rb') as f:
  118. data = pkl.load(f)
  119. A = data['A']
  120. KG = to_KG(A)
  121. num_ent = A[0].shape[0]
  122. data["A"] = KG
  123. data["e"] = num_ent
  124. # analysis(A, y, train, test)
  125. with open(dirname + '/kgdata/class/' + FLAGS.dataset + 'pro.pickle', 'wb') as handle:
  126. pkl.dump(data, handle, protocol=pkl.HIGHEST_PROTOCOL)
  127. with open(dirname + '/kgdata/class/' + FLAGS.dataset + 'pro.pickle', 'rb') as f:
  128. data = pkl.load(f)
  129. KG = data["A"]
  130. # y: csr_sparse_matrix
  131. y = sp.csr_matrix(data['y']).astype(np.float32)
  132. train = data['train_idx']
  133. test = data['test_idx']
  134. num_ent = data["e"]
  135. if FLAGS.dataset in ["train_clinvar_2022_all_test_clinvar_20230326", "train_clinvar_2022_all_test_usDSM"]:
  136. random.shuffle(train)
  137. temp_train = train[:int(0.9*len(train))]
  138. valid = train[int(0.9*len(train)):]
  139. train = temp_train
  140. test = test
  141. logging.info("train {}, valid {}, test {}".format(len(train), len(valid), len(test)))
  142. else:
  143. valid = None
  144. adj = get_extended_adj_auto(num_ent, KG)
  145. return adj, num_ent, train, test, valid, y
  146. def load_data_align(FLAGS):
  147. names = [['ent_ids_1', 'ent_ids_2'], ['triples_1', 'triples_2'], ['ref_ent_ids']]
  148. if FLAGS.rel_align:
  149. names[1][1] = "triples_2_relaligned"
  150. for fns in names:
  151. for i in range(len(fns)):
  152. fns[i] = 'data/'+FLAGS.dataset+'/'+fns[i]
  153. Ent_files, Tri_files, align_file = names
  154. num_ent = len(set(loadfile(Ent_files[0], 1)) | set(loadfile(Ent_files[1], 1)))
  155. align_labels = loadfile(align_file[0], 2)
  156. num_align_labels = len(align_labels)
  157. np.random.shuffle(align_labels)
  158. if not FLAGS.valid:
  159. train = np.array(align_labels[:num_align_labels // 10 * FLAGS.seed])
  160. valid = None
  161. else:
  162. train = np.array(align_labels[:int(num_align_labels // 10 * (FLAGS.seed-1))])
  163. valid = align_labels[int(num_align_labels // 10 * (FLAGS.seed-1)): num_align_labels // 10 * FLAGS.seed]
  164. test = align_labels[num_align_labels // 10 * FLAGS.seed:]
  165. KG = loadfile(Tri_files[0], 3) + loadfile(Tri_files[1], 3)
  166. ent2id = get_ent2id([Ent_files[0], Ent_files[1]])
  167. adj = get_extended_adj_auto(num_ent, KG)
  168. return adj, num_ent, train, test, valid
  169. def load_data_rel_align(FLAGS):
  170. names = [['ent_ids_1', 'ent_ids_2'], ['triples_1', 'triples_2'], ['ref_ent_ids']]
  171. for fns in names:
  172. for i in range(len(fns)):
  173. fns[i] = 'data/'+FLAGS.dataset+'/'+fns[i]
  174. Ent_files, Tri_files, align_file = names
  175. num_ent = len(set(loadfile(Ent_files[0], 1)) | set(loadfile(Ent_files[1], 1)))
  176. align_labels = loadfile(align_file[0], 2)
  177. num_align_labels = len(align_labels)
  178. np.random.shuffle(align_labels)
  179. if not FLAGS.valid:
  180. train = np.array(align_labels[:num_align_labels // 10 * FLAGS.seed])
  181. valid = None
  182. else:
  183. train = np.array(align_labels[:int(num_align_labels // 10 * (FLAGS.seed-1))])
  184. valid = align_labels[int(num_align_labels // 10 * (FLAGS.seed-1)): num_align_labels // 10 * FLAGS.seed]
  185. test = align_labels[num_align_labels // 10 * FLAGS.seed:]
  186. KG = loadfile(Tri_files[0], 3) + loadfile(Tri_files[1], 3)
  187. ent2id = get_ent2id([Ent_files[0], Ent_files[1]])
  188. adj = get_extended_adj_auto(num_ent, KG)
  189. rel_align_labels = loadfile('data/'+FLAGS.dataset+"/ref_rel_ids", 2)
  190. num_rel_align_labels = len(rel_align_labels)
  191. np.random.shuffle(rel_align_labels)
  192. if not FLAGS.valid:
  193. train_rel = np.array(rel_align_labels[:num_rel_align_labels // 10 * FLAGS.rel_seed])
  194. valid_rel = None
  195. else:
  196. train_rel = np.array(rel_align_labels[:int(num_rel_align_labels // 10 * (FLAGS.rel_seed-1))])
  197. valid_rel = rel_align_labels[int(num_rel_align_labels // 10 * (FLAGS.rel_seed-1)): num_rel_align_labels // 10 * FLAGS.rel_seed]
  198. test_rel = rel_align_labels[num_rel_align_labels // 10 * FLAGS.rel_seed:]
  199. return adj, num_ent, train, test, valid, train_rel, test_rel, valid_rel
  200. def get_batch(datax, datay, batch_size):
  201. input_queue = tf.train.slice_input_producer([datax, datay], num_epochs=None, shuffle=False, capacity=32 )
  202. x_batch, y_batch = tf.train.batch(input_queue, batch_size=batch_size, num_threads=1, capacity=32, allow_smaller_final_batch=False)
  203. return x_batch, y_batch
  204. def load_and_preprocess_data(train_file_path, test_file_path, mode):
  205. """
  206. 加载和预处理训练集和测试集数据。
  207. 【修改】: 此函数现在返回在训练集上训练好的imputer对象,
  208. 以便在后续的预测中保持数据处理的一致性。
  209. """
  210. # 读取训练集和测试集CSV文件
  211. train_df = pd.read_csv(train_file_path)
  212. test_df = pd.read_csv(test_file_path)
  213. if mode == "missense":
  214. feature_columns = [
  215. 'BayesDel_addAF_rankscore', 'BayesDel_noAF_rankscore', 'CADD_raw_rankscore',
  216. 'CADD_raw_rankscore_hg19', 'ClinPred_rankscore', 'DANN_rankscore',
  217. 'DEOGEN2_rankscore', 'Eigen_PC_raw_coding_rankscore', 'Eigen_raw_coding_rankscore',
  218. 'FATHMM_converted_rankscore', 'LIST_S2_rankscore', 'M_CAP_rankscore',
  219. 'MPC_rankscore', 'MVP_rankscore', 'MetaLR_rankscore', 'MetaRNN_rankscore',
  220. 'MetaSVM_rankscore', 'MutPred_rankscore', 'MutationAssessor_rankscore',
  221. 'MutationTaster_converted_rankscore', 'PROVEAN_converted_rankscore',
  222. 'Polyphen2_HDIV_rankscore', 'Polyphen2_HVAR_rankscore', 'PrimateAI_rankscore',
  223. 'REVEL_rankscore', 'SIFT4G_converted_rankscore', 'SIFT_converted_rankscore',
  224. 'VEST4_rankscore', 'fathmm_MKL_coding_rankscore', 'fathmm_XF_coding_rankscore',
  225. 'phastCons100way_vertebrate_rankscore', 'phyloP100way_vertebrate_rankscore']
  226. else:
  227. feature_columns = ['usDSM', 'CADD_raw_rankscore', 'TraP', 'SilVA', 'fathmm_MKL_coding_rankscore', 'PrDSM', 'DANN_rankscore']
  228. # 使用两个数据集中都存在的列
  229. common_columns = [col for col in feature_columns if col in train_df.columns and col in test_df.columns]
  230. print(f"使用的共同特征数量: {len(common_columns)}")
  231. # 提取特征和标签
  232. X_train_raw = train_df[common_columns].apply(pd.to_numeric, errors='coerce')
  233. X_test_raw = test_df[common_columns].apply(pd.to_numeric, errors='coerce')
  234. # 检查目标变量列
  235. if 'True Label' not in train_df.columns or 'True Label' not in test_df.columns:
  236. raise ValueError("训练集和测试集中都必须包含 'True Label' 列")
  237. # 【优化】使用numpy.where进行向量化操作,比循环更高效
  238. y_train = np.where(train_df['True Label'] < 0.5, -1, 1)
  239. y_test = np.where(test_df['True Label'] < 0.5, -1, 1)
  240. # 【核心修改】创建imputer,在训练集上fit,然后分别转换训练集和测试集
  241. imputer = SimpleImputer(strategy='mean')
  242. X_train = imputer.fit_transform(X_train_raw)
  243. X_test = imputer.transform(X_test_raw) # 注意:这里只用transform
  244. # 【核心修改】返回训练好的imputer
  245. return X_train, X_test, y_train, y_test, common_columns, imputer
  246. def train_linear_model(X_train, X_test, y_train, y_test):
  247. """
  248. 训练线性回归模型并在测试集上评估。
  249. 【无修改】此函数逻辑正确。
  250. """
  251. # 创建并训练模型
  252. model = LinearRegression()
  253. model.fit(X_train, y_train)
  254. # 在训练集和测试集上预测
  255. y_train_pred = model.predict(X_train)
  256. y_test_pred = model.predict(X_test)
  257. # 评估模型
  258. train_mse = mean_squared_error(y_train, y_train_pred)
  259. test_mse = mean_squared_error(y_test, y_test_pred)
  260. train_r2 = r2_score(y_train, y_train_pred)
  261. test_r2 = r2_score(y_test, y_test_pred)
  262. print(f"模型评估结果:")
  263. print(f"训练集均方误差 (MSE): {train_mse:.4f}")
  264. print(f"测试集均方误差 (MSE): {test_mse:.4f}")
  265. print(f"训练集 R2 分数: {train_r2:.4f}")
  266. print(f"测试集 R2 分数: {test_r2:.4f}")
  267. # 返回模型本身,真实的测试标签,和对测试集的预测
  268. return model, y_test, y_test_pred
  269. def predict_new_data(model, imputer, new_data_file, feature_columns):
  270. """
  271. 【修改版】
  272. 对无标签的新数据进行预测,并输出0到1之间的概率分数。
  273. """
  274. # 加载新数据
  275. new_df = pd.read_csv(new_data_file)
  276. # 检查新数据是否包含所有必需的特征列
  277. if not all(col in new_df.columns for col in feature_columns):
  278. missing = [col for col in feature_columns if col not in new_df.columns]
  279. raise ValueError(f"新数据文件中缺少以下必需的特征列: {missing}")
  280. # 提取特征数据并处理
  281. X_new_raw = new_df[feature_columns].apply(pd.to_numeric, errors='coerce')
  282. X_new = imputer.transform(X_new_raw)
  283. # 步骤1: 从线性模型获取原始预测分数(和之前一样)
  284. raw_scores = model.predict(X_new)
  285. # 步骤2: 【核心修改】应用Sigmoid函数将原始分数转换为0-1之间的概率
  286. # Sigmoid(x) = 1 / (1 + exp(-x))
  287. predicted_probabilities = 1 / (1 + np.exp(-raw_scores))
  288. # 步骤3: 【核心修改】只返回概率分数
  289. return predicted_probabilities

utils.py at commit 3e26596, no license · at the source

Overview

Authors: Hong-Dong Li1,2, Chenlu Wang1, Dongfang Yan1, Wenkui Huang1, Zongxuan Li1, Shaokai Wang1,3
ORCID iDs: Hong-Dong Li
  1. School of Computer Science and Engineering, Central South University, Changsha, Hunan 410083, P.R. China
  2. Hunan Provincial Key Lab on Bioinformatics, Changsha, Hunan 410083, P.R. China
  3. Department of Mathematics, Hong Kong University of Science and Technology, Hong Kong, P.R. China
Journal: Bioinformatics (Oxford, England), volume 42, issue 8, article btag601
Dates: received 17 April 2026; accepted 4 August 2026; published online 10 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1093/bioinformatics/btag601 · PMID 42574509 · PMCID PMC13505626 · OpenAlex W7202139167
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), autism (population), methods / tools (subfield)
Methods: Machine learning
MeSH: Computational Biology*, Genetic Variation*, Models, Genetic*, Software*, Algorithms, Autism Spectrum Disorder, Humans (* major topic)
Topic: Genomics and Rare Diseases (Genetics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Natural Science Foundation of China (62302426, 62473382); Brain Science and Brain-like Intelligence Technology-National Science and Technology Major (2022ZD0213700); Hunan Provincial Natural Science Foundation of China (2025JJ20068)
Citations: not cited yet (Europe PMC); 42 references in the paper

Abstract

Motivation: Predicting variant pathogenicity is crucial for clinical genetics. Existing approaches face two primary limitations. First, biologically, data for pathogenicity prediction often lacks explicit modeling of the gene-variant-feature association structure. A single gene can harbor multiple variants, and each variant can be characterized by multiple features intrinsically associated with its parent gene. Current methods fail to explicitly model the gene-variant-feature association, thus limiting their performance. Second, methodologically, the variant-pathogenicity association is often assumed to comprise a linear component alongside a nonlinear one. However, current methods typically do not explicitly model the linear component, often failing to disentangle the linear component that might be better addressed with a linear approach.

Results: To overcome these limitations, we introduce simultaneous modeling of linear and nonlinear components of knowledge graph (SIMLINK). This novel approach leverages a knowledge graph to model gene-variant-feature associations and a linear model to isolate the linear component. We begin by constructing a variant-centered knowledge graph, comprising over 8 million triplets, which explicitly models the associations between genes, variants, and features. Subsequently, the linear and nonlinear components are learned using a combination of linear and graph neural networks. We train SIMLINK on ClinVar variants. Benchmarking experiments on independent test sets demonstrate its superior prediction on both missense and synonymous variants compared to state-of-the-art methods, including CADD and AlphaMissense. We evaluate the impact of allele frequencies on prediction performance. Applied to variants implicated in Autism Spectrum Disorder, SIMLINK effectively distinguished between high- and low-confidence variants, and critically, the genes harboring top-ranked variants are highly pathogenic.

Availability and implementation: The source code is freely available at https://github.com/Chen-LuWang/SIMLINK.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

Chen-LuWang/SIMLINK

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 3e2659637d8950362b344b7a0db2fe5e0ca74a4d, 5 August 2026
Languages: Python (6), Shell (1)
Size: 24 files, 7 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: TensorFlow (5 files), NumPy (3 files), scikit-learn (2 files), SciPy (2 files), pandas (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
8 files

Zenodo 21802211

License: CC-BY-4.0
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: TensorFlow (5 files), NumPy (3 files), scikit-learn (2 files), SciPy (2 files), pandas (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers (HTTP 200)
  • 26 September 2026: the link answers (HTTP 200)
8 files

Availability and implementation

The source code is freely available at https://github.com/Chen-LuWang/SIMLINK.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 14 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The source code of SIMLINK is freely available at GitHub (https://github.com/Chen-LuWang/SIMLINK). The software version (v1.0.0), benchmark resources, and related data used in this study have been archived in Zenodo (https://doi.org/10.5281/zenodo.21802211).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 7 MeSH terms, 3 funders, 40 references.

Cite

This paper

Li, H.-D., Wang, C., Yan, D., Huang, W., Li, Z., & Wang, S. (2026). SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure. Bioinformatics (Oxford, England), 42(8), btag601. https://doi.org/10.1093/bioinformatics/btag601

BibTeX

@article{li2026simlink,
author = {Li, Hong-Dong and Wang, Chenlu and Yan, Dongfang and Huang, Wenkui and Li, Zongxuan and Wang, Shaokai},
title = {{SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure}},
journal = {Bioinformatics (Oxford, England)},
year = {2026},
month = aug,
volume = {42},
number = {8},
pages = {btag601},
publisher = {Oxford University Press},
issn = {1367-4803},
doi = {10.1093/bioinformatics/btag601},
url = {https://doi.org/10.1093/bioinformatics/btag601},
pmid = {42574509},
pmcid = {PMC13505626}
}

RIS

TY - JOUR
AU - Li, Hong-Dong
AU - Wang, Chenlu
AU - Yan, Dongfang
AU - Huang, Wenkui
AU - Li, Zongxuan
AU - Wang, Shaokai
TI - SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure
T2 - Bioinformatics (Oxford, England)
J2 - Bioinformatics
PY - 2026
DA - 2026/08/01
VL - 42
IS - 8
SP - btag601
SN - 1367-4803
PB - Oxford University Press
DO - 10.1093/bioinformatics/btag601
UR - https://doi.org/10.1093/bioinformatics/btag601
LA - en
ER -

CSL-JSON

{
"id": "10.1093/bioinformatics/btag601",
"type": "article-journal",
"title": "SIMLINK enables accurate variant pathogenicity prediction through modeling the gene-variant-feature association structure",
"container-title": "Bioinformatics (Oxford, England)",
"author": [
{
"family": "Li",
"given": "Hong-Dong"
},
{
"family": "Wang",
"given": "Chenlu"
},
{
"family": "Yan",
"given": "Dongfang"
},
{
"family": "Huang",
"given": "Wenkui"
},
{
"family": "Li",
"given": "Zongxuan"
},
{
"family": "Wang",
"given": "Shaokai"
}
],
"container-title-short": "Bioinformatics",
"volume": "42",
"issue": "8",
"page": "btag601",
"DOI": "10.1093/bioinformatics/btag601",
"PMID": "42574509",
"PMCID": "PMC13505626",
"ISSN": "1367-4803",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/bioinformatics/btag601",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1126/sciadv.adq6577 [code]
Autism-like phenotypes and increased NMDAR2D expression in mice with KDM5B histone lysine demethylase deficiency.
Journal: Science advances
In common: pandas, SciPy, NumPy, autism, 4 references
[2] doi:10.1038/s41467-026-72598-z [code]
Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.
Journal: Nature communications
In common: pandas, SciPy, NumPy, 4 references
[3] doi:10.1101/gr.280394.124 [code]
De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.
Journal: Genome research
In common: TensorFlow, scikit-learn, pandas, 2 other tools, autism, 1 reference
[4] doi:10.1038/s42003-026-10380-z [code]
Identification of moderate effect size genes in autism spectrum disorder through a novel gene pairing approach.
Journal: Communications biology
In common: autism, 4 references
[5] doi:10.1038/s41586-026-10679-1 [code]
Cortical development dynamics across autism spectrum disorder mouse models.
Journal: Nature
In common: scikit-learn, pandas, SciPy, 1 other tool, autism, 2 references
[6] doi:10.1038/s41586-026-10515-6 [code]
An X-linked long non-coding RNA, PTCHD1-AS, and the core features of autism.
Journal: Nature
In common: pandas, SciPy, NumPy, autism, 2 references
[7] doi:10.1016/j.xhgg.2026.100652 [code]
CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling.
Journal: HGG advances
In common: pandas, SciPy, NumPy, autism, 2 references
[8] doi:10.1038/s41598-026-55163-y [code]
Autism spectrum disorder identification using machine learning models on MRI data.
Journal: Scientific reports
In common: TensorFlow, scikit-learn, pandas, 2 other tools, autism
[9] doi:10.1016/j.xgen.2026.101284 [code]
NERINE reveals rare variant associations in gene networks across phenotypes and implicates an SNCA-PRL-LRRK2 subnetwork in Parkinson's disease.
Journal: Cell genomics
In common: pandas, SciPy, NumPy, 2 references
[10] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: TensorFlow, scikit-learn, pandas, 2 other tools, methods / tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.