OSCR

Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Results › Feature importance analysis › Calculate feature importance using DeepSHAP ↔ Notebooks/DeepShap.ipynb, lines 281–343 · score 0.87 · gen.f1, gen.f3, gen.f4, gen.f2, img.f4, img.f2
  2. [2] § Results › Feature importance analysis › Visualize voxel-level importance on image using Grad-CAM ↔ Notebooks/DeepRisk_CNN_evaluation_SNP.ipynb, lines 2032–2050 · score 0.79 · Frontal lobe, Parietal lobe, Temporal lobe, cingulate gyri, brain regions, Insula
  3. [3] § Results › Feature importance analysis › Visualize voxel-level importance on image using Grad-CAM ↔ Notebooks/DeepRisk_CNN_evaluation_image_features.ipynb, lines 1793–1811 · score 0.79 · Frontal lobe, Parietal lobe, Temporal lobe, cingulate gyri, brain regions, Insula
  4. [4] § Results › Feature interaction analysis › Feature interaction plot ↔ Notebooks/DeepShap.ipynb, lines 281–343 · score 0.62 · gen.f4, gen.f2, img.f2, APOE, age
  5. [5] § Results › Feature importance analysis › Visualize voxel-level importance on image using Grad-CAM ↔ Notebooks/DeepRisk_CNN_evaluation_image_features.ipynb, lines 687–759 · score 0.55 · Grad CAM, attention map, heatmaps, Gradient, CNN, GM
  6. [6] § Results › Feature importance analysis › Visualize voxel-level importance on image using Grad-CAM ↔ Notebooks/DeepRisk_CNN_evaluation_SNP.ipynb, lines 688–760 · score 0.55 · Grad CAM, attention map, heatmaps, Gradient, CNN, GM

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 576 lines · 18 KB · no license · 2 matches

  1. # %%
  2. # !pip install shap
  3. # !pip install xgboost
  4. # %%
  5. from __future__ import print_function
  6. import keras
  7. import tensorflow as tf
  8. from tensorflow.keras.models import Sequential, Model
  9. from tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping
  10. from tensorflow.keras.layers import Dense, Activation, Flatten, Conv3D, MaxPooling3D, BatchNormalization, Dropout, GlobalAveragePooling3D
  11. from tensorflow.keras.layers import Input, concatenate, multiply, add, Reshape, Lambda
  12. from keras.datasets import mnist
  13. from keras.models import Sequential
  14. from keras.layers import Dense, Dropout, Flatten
  15. from keras.layers import Conv2D, MaxPooling2D
  16. from keras import backend as K
  17. import shap
  18. import pandas as pd
  19. import numpy as np
  20. import h5py
  21. import matplotlib
  22. matplotlib.use('Agg')
  23. import matplotlib.pyplot as plt
  24. import matplotlib.gridspec as gridspec
  25. import seaborn as sns
  26. %matplotlib inline
  27. # %%
  28. class LoadData:
  29. """
  30. Loading preprocessed data from .h5 file.
  31. Has to be similar to saving data function in data processing notebook.
  32. (Same names for datasets etc.)
  33. """
  34. def __init__(self, name):
  35. dataset_file = name+'.h5'
  36. f = h5py.File(DATASET_DIR+dataset_file, 'r')
  37. # f = h5py.File('/home/gennadyr/IPython/tests_john/version_age_3/models/'+dataset_file, 'r')
  38. self.fraction_train = f['fraction_train'][:]
  39. self.fraction_validation = f['fraction_validation'][:]
  40. self.fraction_test = f['fraction_test'][:]
  41. self.train_MRI_data = f['train_MRI_data'][:]
  42. self.validation_MRI_data = f['validation_MRI_data'][:]
  43. self.test_MRI_data = f['test_MRI_data'][:]
  44. gene_columns = f['gene_column_names'][:]
  45. train_gene_data = f['train_gene_data'][:]
  46. validation_gene_data = f['validation_gene_data'][:]
  47. test_gene_data = f['test_gene_data'][:]
  48. self.train_gene_data = pd.DataFrame(train_gene_data, columns=gene_columns)
  49. self.validation_gene_data = pd.DataFrame(validation_gene_data, columns=gene_columns)
  50. self.test_gene_data = pd.DataFrame(test_gene_data, columns=gene_columns)
  51. train_label_data1 = f['train_label_data1'][:]
  52. train_label_data2 = f['train_label_data2'][:]
  53. validation_label_data1 = f['validation_label_data1'][:]
  54. validation_label_data2 = f['validation_label_data2'][:]
  55. test_label_data1 = f['test_label_data1'][:]
  56. test_label_data2 = f['test_label_data2'][:]
  57. columns = f['label_column_names'][:]
  58. train_label_data2 = pd.DataFrame(train_label_data2)
  59. train_label_data2.columns = columns
  60. validation_label_data2 = pd.DataFrame(validation_label_data2)
  61. validation_label_data2.columns = columns
  62. test_label_data2 = pd.DataFrame(test_label_data2)
  63. test_label_data2.columns = columns
  64. train_label_data2['bigrfullname'] = train_label_data1
  65. validation_label_data2['bigrfullname'] = validation_label_data1
  66. test_label_data2['bigrfullname'] = test_label_data1
  67. self.train_label_data = train_label_data2
  68. self.validation_label_data = validation_label_data2
  69. self.test_label_data = test_label_data2
  70. f.close()
  71. print('Loaded datasets from '+DATASET_DIR+dataset_file)
  72. # %%
  73. def generate_riskset(event_times):
  74. """
  75. Generates the riskset for every individual. Riskset is the set of individuals that have a
  76. longer event time and are thus at risk of experiencing the event : Tj>=Ti
  77. Input:
  78. - label_data = dataframe with file name, event times and other labels that do not get used
  79. Output:
  80. - riskset = square matrix in which row i is the riskset of individual i compared to all
  81. individuals j. Entry is true if Tj>=Ti, so individual j is 'at risk'.
  82. """
  83. o = np.argsort(-event_times, kind="mergesort")
  84. n_samples = len(event_times)
  85. risk_set = np.zeros((n_samples, n_samples), dtype=np.bool_)
  86. for i_org, i_sort in enumerate(o):
  87. ti = event_times[i_sort]
  88. k = i_org
  89. while k < n_samples and ti == event_times[o[k]]:
  90. k += 1
  91. risk_set[i_sort, o[:k]] = True
  92. return risk_set
  93. # **Function to normalize risk scores**
  94. # In[36]:
  95. def safe_normalize(x):
  96. """Normalize risk scores to avoid exp underflowing.
  97. Note that only risk scores relative to each other matter.
  98. If minimum risk score is negative, we shift scores so minimum
  99. is at zero.
  100. """
  101. x_min = tf.reduce_min(x, axis=0)
  102. c = tf.zeros_like(x_min)
  103. norm = tf.where(x_min < 0, -x_min, c)
  104. return x + norm
  105. # **Function to calculate log of sum of exponent of predictions** (right hand side of equation 1.1)
  106. # In[37]:
  107. def logsumexp_masked(risk_scores, mask, axis = 0, keepdims= None):
  108. """
  109. Computes the log of the sum of the exponent of the predictions across `axis`
  110. for all entries where `mask` (riskset) is true:
  111. log(sum(e^h_j))
  112. where h_j are the predictions of patients at risk of developing dementia (T_j>=T_i)
  113. Inputs:
  114. - risk_scores = the predictions from the network of patients h_j
  115. - mask = a mask to select which patients are at risk
  116. Output:
  117. - output = right hand part of the NPLL (Negative Partial Log Likelihood)
  118. """
  119. risk_scores.shape.assert_same_rank(mask.shape)
  120. with tf.name_scope("logsumexp_masked"):
  121. risk_scores = tf.cast(risk_scores,tf.float32)
  122. mask_f = tf.cast(mask, risk_scores.dtype)
  123. risk_scores_masked = tf.math.multiply(risk_scores, mask_f)
  124. #for numerical stability, substract the maximum value
  125. #before taking the exponential
  126. amax = tf.reduce_max(risk_scores_masked, axis=axis, keepdims=True)
  127. risk_scores_shift = risk_scores_masked - amax
  128. exp_masked = tf.math.multiply(tf.math.exp(risk_scores_shift), mask_f)
  129. exp_sum = tf.reduce_sum(exp_masked, axis=axis, keepdims=True)
  130. #turn 0's to 1's to get rid of inf loss (log(0) = inf)
  131. condition = tf.not_equal(exp_sum, 0)
  132. exp_sum_clean = tf.where(condition, exp_sum, tf.ones_like(exp_sum))
  133. output = amax + tf.math.log(exp_sum_clean)
  134. if not keepdims:
  135. output = tf.squeeze(output, axis=axis)
  136. return output
  137. # **Custom loss function** (Negative Partial Log Likelihood)
  138. # In[38]:
  139. def CoxPH_loss(y_true, y_pred):
  140. """
  141. Calculates the Negative Partial Log Likelihood:
  142. L = sum(h_i - log(sum(e^h_j)))
  143. where;
  144. h_i = risk prediction of patient i
  145. h_j is risk prediction of patients j at risk of developing dementia (T_j>=T_i)
  146. Inputs:
  147. - y_true = label data composed of
  148. y_event: A 1 or 0 indicating if the patient developed dementia or not, and
  149. y_riskset:(set of patients j which are at risk dependent on patient i (Tj>=Ti)).
  150. - y_pred = the risk prediction of the network
  151. Output:
  152. - loss = the loss used to optimize the network
  153. """
  154. event = y_true[:,0]
  155. event = tf.reshape(event,(-1,1))
  156. riskset_loss = y_true[:,1:]
  157. predictions = y_pred
  158. predictions = tf.cast(predictions,tf.float32)
  159. riskset_loss = tf.cast(riskset_loss,tf.bool)
  160. event = tf.cast(event, predictions.dtype)
  161. predictions = safe_normalize(predictions)
  162. # with tf.name_scope("assertions"):
  163. # assertions = (
  164. # tf.debugging.assert_less_equal(event, 1.),
  165. # tf.debugging.assert_greater_equal(event, 0.),
  166. # tf.debugging.assert_type(riskset_loss, tf.bool)
  167. # )
  168. # move batch dimension to the end so predictions get broadcast
  169. # row-wise when multiplying by riskset
  170. pred_t = tf.transpose(predictions)
  171. # compute log of sum over risk set for each row
  172. rr = logsumexp_masked(pred_t, riskset_loss, axis=1, keepdims=True)
  173. # print(predictions.shape.as_list())
  174. # assert rr.shape.as_list() == predictions.shape.as_list()
  175. loss = tf.math.multiply(event, rr - predictions)
  176. return loss
  177. # %%
  178. ### 11 features
  179. def submodel():
  180. model = tf.keras.models.load_model(MODEL_DIR+'model_'+version+'.h5', custom_objects={'CoxPH_loss': CoxPH_loss})
  181. input1 = Input((11,))# MRI input
  182. x = input1
  183. for layer in model.layers[-4:-1]: # loop over convolutional layers from [1] and add to model
  184. # connect the layers
  185. x = layer(x)
  186. final= model.layers[-1](x)
  187. model = Model(inputs=[input1], outputs=final)
  188. adam_opt = keras.optimizers.Adam(lr=0.001, beta_1=0.9, beta_2=0.999, epsilon=1e-08, decay=0.01)
  189. model.compile(loss=CoxPH_loss, optimizer=adam_opt)
  190. return model
  191. # %%
  192. ### 7 features and SNPs
  193. def submodel():
  194. model = tf.keras.models.load_model(MODEL_DIR+'model_'+version+'.h5', custom_objects={'CoxPH_loss': CoxPH_loss})
  195. input1 = Input((7,))
  196. input2 = Input((76,))# MRI input
  197. x1= model.layers[-6](input2)
  198. x = concatenate([input1, x1])
  199. for layer in model.layers[-4:-1]: # loop over convolutional layers from [1] and add to model
  200. # connect the layersxx
  201. x = layer(x)
  202. final= model.layers[-1](x)
  203. model = Model(inputs=[input1,input2], outputs=final)
  204. adam_opt = keras.optimizers.Adam(lr=0.001, beta_1=0.9, beta_2=0.999, epsilon=1e-08, decay=0.01)
  205. model.compile(loss=CoxPH_loss, optimizer=adam_opt)
  206. return model
  207. # %%
  208. version = 'MRI_SNP_TFS_FS_l_1'
  209. MODEL_DIR = '/trinity/home/jyu/DeepSurvival/models/RS+MCICASES_FS_59_5cv/MRI/'
  210. DATASET_DIR = '/data/scratch/jyu/DeepSurvival/data/'
  211. prepdata_name = 'Total_fs_RS_cv_split_1'
  212. # %%
  213. model1=submodel()
  214. # %%
  215. data = LoadData(prepdata_name)
  216. #prep full RS datasets
  217. # Train
  218. train_MRI_set = data.train_MRI_data
  219. train_MRI_set=np.char.decode(train_MRI_set)
  220. train_label_set = data.train_label_data
  221. b=train_label_set['bigrfullname'].str.decode("utf-8")
  222. train_label_set['bigrfullname']=1
  223. train_label_set['bigrfullname']=b
  224. train_label_set = train_label_set.set_index('bigrfullname')
  225. train_label_set.columns=train_label_set.columns.str.decode("utf-8")
  226. train_label_set['ergoid']=train_label_set['ergoid'].astype('int')
  227. columnsname=['age']
  228. mean=train_label_set[columnsname].mean()
  229. std=train_label_set[columnsname].std()
  230. train_label_set[columnsname]=(train_label_set[columnsname]-mean)/std
  231. train_gene_data = data.train_gene_data
  232. train_gene_data.columns=train_gene_data.columns.str.decode("utf-8")
  233. train_gene_data['ergoid']=train_gene_data['ergoid'].astype('int')
  234. train_gene_data = train_gene_data.set_index('ergoid')
  235. test_MRI_set = data.test_MRI_data
  236. test_MRI_set=np.char.decode(test_MRI_set)
  237. test_label_set = data.test_label_data
  238. b=test_label_set['bigrfullname'].str.decode("utf-8")
  239. test_label_set['bigrfullname']=1
  240. test_label_set['bigrfullname']=b
  241. test_label_set = test_label_set.set_index('bigrfullname')
  242. test_label_set.columns=test_label_set.columns.str.decode("utf-8")
  243. test_label_set['ergoid']=test_label_set['ergoid'].astype('int')
  244. test_label_set[columnsname]=(test_label_set[columnsname]-mean)/std
  245. test_gene_data = data.test_gene_data
  246. test_gene_data.columns=test_gene_data.columns.str.decode("utf-8")
  247. test_gene_data['ergoid']=test_gene_data['ergoid'].astype('int')
  248. test_gene_data = test_gene_data.set_index('ergoid')
  249. train_label_set=train_label_set.merge(train_gene_data,on='ergoid')
  250. test_label_set=test_label_set.merge(test_gene_data,on='ergoid')
  251. train_label_set=train_label_set.sort_values(['age','ergoid']).drop_duplicates('ergoid')
  252. train=pd.read_csv('../Plot/output/features_5fold/train_features_TFS_1.csv')
  253. train=train.loc[train_label_set.index]
  254. test_label_set=test_label_set.sort_values(['ergoid','age']).drop_duplicates('ergoid')
  255. features=pd.read_csv('../Plot/output/features_5fold/test_features_TFS_1.csv')
  256. X=features[['img.f1','img.f2','img.f3','img.f4','sex','age','apoe','gen.f1','gen.f2','gen.f3','gen.f4']]
  257. X=X.loc[test_label_set.index]
  258. X[columnsname]=(X[columnsname]-mean)/std
  259. # %%
  260. # explain the model's predictions using SHAP
  261. # (same syntax works for LightGBM, CatBoost, scikit-learn, transformers, Spark, etc.)
  262. explainer = shap.DeepExplainer(model1,np.array(train))
  263. shap_values = explainer.shap_values(np.array(X))
  264. # %%
  265. # explain the model's predictions using SHAP
  266. # (same syntax works for LightGBM, CatBoost, scikit-learn, transformers, Spark, etc.)
  267. train=pd.concat([train[['img.f1','img.f2','img.f3','img.f4','sex','age','apoe']],train_label_set.iloc[:,-76:]],axis=1)
  268. X=pd.concat([X[['img.f1','img.f2','img.f3','img.f4','sex','age','apoe']],test_label_set.iloc[:,-76:]],axis=1)
  269. X1=X.iloc[:,-76:]
  270. X2=X.iloc[:,:7]
  271. explainer = shap.DeepExplainer(model1,[np.array(train.iloc[:,:7]),np.array(train.iloc[:,-76:])])
  272. shap_values = explainer.shap_values([np.array(X2),np.array(X1)])
  273. # %%
  274. shap.summary_plot(shap_values[0][1], X1, plot_type='dot',max_display=76, plot_size=(8,20),sort=True,color_bar_label='Feature value')
  275. # %%
  276. genetic=pd.read_csv('genetic_shap.csv')
  277. beta=pd.read_csv('/data/scratch/jyu/DeepSurvival/dataset/snp/NewGWAS-83snps.csv')[['SNP','BETA','P value']]
  278. beta=beta.set_index('SNP')
  279. # #beta['BETA'] = np.log(beta['BETA'])
  280. # beta=beta[~beta.index.isin(MAF_5[1])]
  281. # #beta['OR'] = np.exp(beta['BETA'])
  282. beta=beta.loc[genetic.Input]
  283. beta['SHAP_1']=genetic.shap1.tolist()
  284. beta['SHAP_all']=genetic.shapall.tolist()
  285. beta['SHAP_all1']=genetic.shapall1.tolist()
  286. #beta['MAF']=fsnp.tolist()
  287. beta['std']=std
  288. beta = beta.fillna(0)
  289. # %%
  290. fig, ax = plt.subplots(figsize=(6,6))
  291. p1=sns.scatterplot(x='std', y='SHAP_all',data=beta)
  292. # for line in range(0,beta.shape[0]):
  293. # p1.text(beta['std'][line]+0.001, beta.SHAP_all1[line]-0.001,
  294. # beta.index[line], horizontalalignment='left',
  295. # size=8, color='black', weight='semibold')
  296. plt.xlabel('Standard deviation of SNP allele')
  297. plt.ylabel('mean |SHAP value|')
  298. sns.despine()
  299. # %%
  300. fig, ax = plt.subplots(figsize=(6,6))
  301. p1=sns.scatterplot(x='SHAP_all', y='SHAP_1',data=beta)
  302. # for line in range(0,beta.shape[0]):
  303. # p1.text(beta.SHAP_all[line]+0.0001, beta.SHAP_1[line],
  304. # beta.index[line], horizontalalignment='left',
  305. # size=8, color='black', weight='semibold')
  306. plt.xlabel('mean |SHAP value| (all)')
  307. plt.ylabel('mean |SHAP value| (1)')
  308. sns.despine()
  309. # %%
  310. beta1=beta[beta['P value']>0.01]
  311. fig, ax = plt.subplots(figsize=(6,6))
  312. p1=sns.scatterplot(x='P value', y='SHAP_all',data=beta)
  313. for line in range(0,beta1.shape[0]):
  314. p1.text(beta1['P value'][line]+0.001, beta1.SHAP_all[line],
  315. beta1.index[line], horizontalalignment='left',
  316. size=8, color='black', weight='semibold')
  317. plt.xlabel('P value')
  318. plt.ylabel('mean |SHAP value|')
  319. sns.despine()
  320. # %%
  321. shap.summary_plot(shap_values[0], X, plot_type='dot', color_bar_label='Interaction strength',color='blue')
  322. # %%
  323. shap.summary_plot(shap_values[0], X, plot_type='bar')
  324. # %%
  325. shap.plots._waterfall.waterfall_legacy(explainer.expected_value[0].numpy(), shap_values[0][1], feature_names = X.columns)
  326. # %%
  327. shap.plots._waterfall.waterfall_legacy(1/(1+np.exp(-explainer.expected_value[0].numpy())), shap_values1[0][429], feature_names = X.columns)
  328. # %%
  329. shap.initjs()
  330. shap.force_plot(explainer.expected_value[0].numpy(), shap_values[0][1,:], X.iloc[1,:], link="logit")
  331. # %%
  332. shap.initjs()
  333. shap.force_plot(explainer.expected_value[0].numpy(), shap_values[0][429,:], X.iloc[429,:], link="logit")
  334. # %%
  335. X['Age']=std[0]*X['Age']+mean[0]
  336. # %%
  337. plt.figure(figsize=(6,5))
  338. sns.scatterplot(x=X['Age'].tolist(),y=shap_values[0][:,5])
  339. plt.xlabel('Age',size=14)
  340. # Set y-axis label
  341. plt.ylabel('SHAP value',size=14)
  342. # %%
  343. plt.figure(figsize=(6,5))
  344. sns.scatterplot(x=X['Img.F2'].tolist(),y=shap_values[0][:,1])
  345. plt.xlabel('Img.F2',size=14)
  346. # Set y-axis label
  347. plt.ylabel('SHAP value',size=14)
  348. # %%
  349. plt.figure(figsize=(6,5))
  350. sns.scatterplot(x=X['Img.F4'].tolist(),y=shap_values[0][:,3])
  351. plt.xlabel('Img.F4',size=14)
  352. # Set y-axis label
  353. plt.ylabel('SHAP value',size=14)
  354. # %%
  355. plt.figure(figsize=(6,5))
  356. sns.scatterplot(x=X['ApoE-ε4'].tolist(),y=shap_values[0][:,6])
  357. plt.xlabel('ApoE-ε4',size=14)
  358. # Set y-axis label
  359. plt.ylabel('SHAP value',size=14)
  360. # %%
  361. plt.figure(figsize=(6,5))
  362. sns.scatterplot(x=X['ApoE-ε4'].tolist(),y=shap_values[0][:,6])
  363. plt.xlabel('ApoE-ε4',size=14)
  364. # Set y-axis label
  365. plt.ylabel('SHAP value',size=14)
  366. # %%
  367. plt.figure(figsize=(6,5))
  368. sns.scatterplot(x=X['Gen.F1'].tolist(),y=shap_values[0][:,-4])
  369. plt.xlabel('Gen.F1',size=14)
  370. # Set y-axis label
  371. plt.ylabel('SHAP value',size=14)
  372. # %%
  373. plt.figure(figsize=(6,5))
  374. sns.scatterplot(x=X['Gen.F2'].tolist(),y=shap_values[0][:,-3])
  375. plt.xlabel('Gen.F2',size=14)
  376. # Set y-axis label
  377. plt.ylabel('SHAP value',size=14)
  378. # %%
  379. shap.dependence_plot("Age",shap_values[0], X)
  380. # %%
  381. shap.dependence_plot("ApoE-ε4",shap_values[0], X)
  382. # %%
  383. shap.dependence_plot("ApoE-ε4",shap_values[0][:,[1,6]], X.iloc[:,[1,6]])
  384. # %%
  385. shap.dependence_plot("Age",shap_values[0][:,[5,6]], X.iloc[:,[5,6]])
  386. # %%
  387. shap.dependence_plot("Age",shap_values[0][:,[1,5]], X.iloc[:,[1,5]])
  388. # %%
  389. shap.dependence_plot("Img.F2",shap_values[0][:,[1,5]], X.iloc[:,[1,5]])
  390. # %%
  391. shap.dependence_plot("Age",shap_values[0][:,[3,5]], X.iloc[:,[3,5]])
  392. # %%
  393. shap.dependence_plot("Img.F4",shap_values[0][:,[3,5]], X.iloc[:,[3,5]])
  394. # %%
  395. shap.dependence_plot("Img.F2",shap_values[0][:,[1,3]], X.iloc[:,[1,3]])
  396. # %%
  397. shap.dependence_plot("Img.F4",shap_values[0][:,[1,3]], X.iloc[:,[1,3]])
  398. # %%
  399. shap.dependence_plot("Img.F4",shap_values[0][:,[3,5]], X.iloc[:,[3,5]])
  400. # %%
  401. shap.dependence_plot("Gen.F2",shap_values[0][:,[8,5]], X.iloc[:,[8,5]])
  402. # %%
  403. shap.dependence_plot("Gen.F2",shap_values[0][:,[8,1]], X.iloc[:,[8,1]])
  404. # %%
  405. shap.dependence_plot("Gen.F4",shap_values[0][:,[-1,1]], X.iloc[:,[-1,1]])
  406. # %%
  407. shap.dependence_plot("Gen.F4",shap_values[0][:,[-1,3]], X.iloc[:,[-1,3]])
  408. # %%
  409. shap.dependence_plot("Gen.F4",shap_values[0][:,[-1,5]], X.iloc[:,[-1,5]])
  410. # %%
  411. shap.dependence_plot("ApoE-ε4",shap_values1[0][:,[5,6]], X.iloc[:,[5,6]])
  412. # %%
  413. shap.dependence_plot("ApoE-ε4",shap_values[0][:,[1,6]], X.iloc[:,[1,6]])
  414. # %%
  415. shap.dependence_plot("Gen.F2",shap_values[0][:,[1,8]], X.iloc[:,[1,8]])
  416. # %%
  417. shap.dependence_plot("Img.F2",shap_values[0][:,[1,8]], X.iloc[:,[1,8]])
  418. # %%
  419. shap.dependence_plot("Gen.F2",shap_values[0][:,[3,8]], X.iloc[:,[3,8]])
  420. # %%
  421. shap.dependence_plot("Img.F4",shap_values[0][:,[3,8]], X.iloc[:,[3,8]])
  422. # %%

DeepShap.ipynb at commit f50a5c1, no license · at the source

Overview

Authors: Jing Yu1,2,3, Mathijs T. Rosbergen1,2, Frank J. Wolters1,2, Esther E. Bron2, Meike W. Vernooij1,2, M. Arfan Ikram1,2, Gennady V. Roshchupkin1,2
  1. Department of Epidemiology, Erasmus MC University Medical Center,Dr. Molewaterplein 40, Rotterdam, 3015 GD Netherlands
  2. Department of Radiology & Nuclear Medicine, Erasmus MC University Medical Center,Dr. Molewaterplein 40, Rotterdam, 3015 GD Netherlands
  3. Department of Epidemiology, Department of Radiology & Nuclear Medicine, Erasmus MC University Medical Center,Dr. Molewaterplein 40, Rotterdam, 3015 GD Netherlands
Institutions: Erasmus MC (Netherlands)
Journal: Scientific reports, volume 16, issue 1, article 16787
Dates: received 15 September 2025; accepted 29 March 2026; published online 9 April 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41598-026-47047-y · PMID 41957153 · PMCID PMC13223248 · OpenAlex W4414420356
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), Alzheimer's / dementia (population), clinical / translational (subfield)
Methods: Statistics, Machine learning, Preprocessing
Keywords: Dementia risk prediction, Brain MRI, Genetics, Deep learning, Survival analysis, Cohort study, Computational biology and bioinformatics, Neurology, Neuroscience
MeSH: Dementia*, Aged, Aged, 80 and over, Alzheimer Disease, Convolutional Neural Networks, Deep Learning, Female, Humans, Magnetic Resonance Imaging, Male, Netherlands, Neural Networks, Computer, Neuroimaging, Polymorphism, Single Nucleotide, Proportional Hazards Models (* major topic)
Topic: Machine Learning in Healthcare (Artificial Intelligence, Computer Science), according to OpenAlex
Funding: China Scholarship Council (202006640010); BIRD-NL dementia prevention initiative grant (10510032120005); ZonMw Veni grant (1936320)
Citations: not cited yet (Europe PMC); 57 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

yyyuj/Deep_Survival

License: none: the authors keep all their rights
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: f50a5c1203b2625744793a992cf51fef2bfa9eb2, 6 September 2025
Languages: Shell (12), Python (12), Jupyter (4)
Size: 30 files, 28 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (requirements.txt), 4 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: TensorFlow (27 files), h5py (16 files), Matplotlib (16 files), NumPy (16 files), pandas (16 files), Keras (15 files), SciPy (15 files), NiBabel (14 files), seaborn (4 files), scikit-learn (1 file), SHAP (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
29 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41598-026-47047-y.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 28 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s41598-026-47047-y.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 9 keywords, 15 MeSH terms, 3 funders, 53 references.

Cite

This paper

Yu, J., Rosbergen, M. T., Wolters, F. J., Bron, E. E., Vernooij, M. W., Ikram, M. A., & Roshchupkin, G. V. (2026). Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study. Scientific reports, 16(1), 16787. https://doi.org/10.1038/s41598-026-47047-y

BibTeX

@article{yu2026imaging,
author = {Yu, Jing and Rosbergen, Mathijs T. and Wolters, Frank J. and Bron, Esther E. and Vernooij, Meike W. and Ikram, M. Arfan and Roshchupkin, Gennady V.},
title = {{Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study}},
journal = {Scientific reports},
year = {2026},
month = apr,
volume = {16},
number = {1},
pages = {16787},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-47047-y},
url = {https://doi.org/10.1038/s41598-026-47047-y},
pmid = {41957153},
pmcid = {PMC13223248}
}

RIS

TY - JOUR
AU - Yu, Jing
AU - Rosbergen, Mathijs T.
AU - Wolters, Frank J.
AU - Bron, Esther E.
AU - Vernooij, Meike W.
AU - Ikram, M. Arfan
AU - Roshchupkin, Gennady V.
TI - Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/04/09
VL - 16
IS - 1
SP - 16787
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-47047-y
UR - https://doi.org/10.1038/s41598-026-47047-y
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-47047-y",
"type": "article-journal",
"title": "Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study",
"container-title": "Scientific reports",
"author": [
{
"family": "Yu",
"given": "Jing"
},
{
"family": "Rosbergen",
"given": "Mathijs T."
},
{
"family": "Wolters",
"given": "Frank J."
},
{
"family": "Bron",
"given": "Esther E."
},
{
"family": "Vernooij",
"given": "Meike W."
},
{
"family": "Ikram",
"given": "M. Arfan"
},
{
"family": "Roshchupkin",
"given": "Gennady V."
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "16787",
"DOI": "10.1038/s41598-026-47047-y",
"PMID": "41957153",
"PMCID": "PMC13223248",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-47047-y",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
9
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1162/imag.a.1164 [code]
Bias and generalizability of brain age prediction models: A multi-cohort evaluation with anatomical and interpretability insights.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Keras, TensorFlow, NiBabel, 6 other tools, Alzheimer's / dementia, structural MRI / diffusion, 1 reference
[2] doi:10.1038/s41586-026-10631-3 [code]
A prognostic human brain network for diffuse midline glioma.
Journal: Nature
In common: Keras, TensorFlow, h5py, 7 other tools, clinical / translational
[3] doi:10.1038/s41398-026-04081-8 [code]
Functional system-specific brain aging across the Alzheimer's disease continuum.
Journal: Translational psychiatry
In common: Keras, TensorFlow, NiBabel, 6 other tools, Alzheimer's / dementia, structural MRI / diffusion, clinical / translational
[4] doi:10.64898/2026.05.06.26352540 [code]
Generating synthetic tau-PET scans in Alzheimer’s disease from MRI, blood biomarkers and demographics with deep learning
Journal: medRxiv (preprint)
In common: Keras, TensorFlow, NiBabel, 6 other tools, Alzheimer's / dementia, structural MRI / diffusion, clinical / translational
[5] doi:10.1080/07853890.2026.2685416 [code]
Pulmonary and cerebral damage in COVID-19 survivors: is there any association?
Journal: Annals of medicine
In common: Keras, TensorFlow, h5py, 6 other tools, structural MRI / diffusion, clinical / translational
[6] doi:10.1038/s41597-025-05174-7 [code]
A large-scale MEG and EEG dataset for object recognition in naturalistic scenes
Journal: n/a
In common: Keras, TensorFlow, h5py, 7 other tools
[7] doi:10.3389/fnsys.2026.1822122 [code]
Convergence-divergence circuits for multimodal integration of innate and learned opponent valences.
Journal: Frontiers in systems neuroscience
In common: Keras, TensorFlow, h5py, 7 other tools
[8] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: SHAP, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, clinical / translational, 1 reference
[9] doi:10.1186/s12880-026-02481-2 [code]
Deep learning-based neuroanatomical profiling reveals population-specific brain changes in multiple sclerosis: a large-scale Middle Eastern study.
Journal: BMC medical imaging
In common: Keras, TensorFlow, NiBabel, 6 other tools, structural MRI / diffusion, clinical / translational
[10] doi:10.1002/hbm.70602 [code]
Neuroimaging Correlates of Post-Stroke Pain After Ischemic Stroke: Secondary Analysis of the INSPiRE-TMS Trial.
Journal: Human brain mapping
In common: Keras, TensorFlow, h5py, 6 other tools, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.