OSCR

Local and global patterns support medical imaging as a biomarker of ageing.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Results › Lifestyle and Environment ↔ source/analysis/perform_phewas.py, lines 154–213 · score 0.96 · family history, sun exposure, mental health, sexual factors, social support, general health
  2. [2] § Results › Chronic Diseases ↔ source/analysis/disease_analysis.py, lines 51–78 · score 0.92 · ischemic heart disease, liver disease subjects, COPD subjects, MS subjects, scoliosis subjects, hypertension subjects
  3. [3] § Methods › Lifestyle and Environment ↔ source/analysis/perform_phewas.py, lines 320–399 · score 0.63 · rank normalisation, imaging phenotypes, variables, confounding, PheWAS, correlate
  4. [4] § Methods › Deep Learning Models ↔ source/cnn/model/resnet.py, lines 185–193 · score 0.61 · resnet18, ImageNet, pre trained, models
  5. [5] § Methods › Survival Analysis ↔ source/analysis/survival.py, lines 1–44 · score 0.60 · Cox proportional hazards, survival, death, female, brain, body
  6. [6] § Methods › Dataset › UK Biobank ↔ source/cnn/train_utils.py, lines 272–295 · score 0.57 · largest connected component, voxels, mask, training

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 612 lines · 24 KB · no license · 2 matches

  1. # %%
  2. import numpy as np
  3. import pandas as pd
  4. import ast
  5. import os
  6. from tqdm import tqdm
  7. import seaborn as sns
  8. import scipy.stats
  9. import math
  10. import csv
  11. import matplotlib.pyplot as plt
  12. import utils
  13. import paths
  14. #%%
  15. ############################################################################################################################################################################
  16. #### This script is used to generate the Manhattan plot and the correlation plot between non-imaging phenotypes and imaging phenotypes ###
  17. #### The script is adapted from Bai et al. 2020, "A population-based phenome-wide association study of cardiac and aortic structure and function", Nature medicine. ###
  18. ############################################################################################################################################################################
  19. #%%
  20. def normalise(x):
  21. return (x - np.mean(x)) / np.std(x)
  22. def rank_normalise(x):
  23. # Rank-based inverse normal transform
  24. # Please refer to the function inormal() in the FSLNets package
  25. # Get the rank of the values in x
  26. ri = np.argsort(np.argsort(x))
  27. # Correct for the ranks of repeated values
  28. # argsort assign different ranks for these values
  29. # We fill them with the same value
  30. u, inv_idx = np.unique(x, return_inverse=True)
  31. sii = np.sort(inv_idx)
  32. repeated_idx = np.unique(sii[np.diff(np.append(sii, 1)) == 0])
  33. for i in repeated_idx:
  34. ri[inv_idx == i] = np.mean(ri[inv_idx == i])
  35. # Perform inverse normal transform
  36. # ri + 1 so that the rank starts from 1, to be consistent with Karla's Matlab code
  37. # p squashes the rank into the range of [0, 1]
  38. # erfinv can generate a distribution from 2 * p - 1 with 0 mean and 1 standard deviation
  39. N = len(x)
  40. ri = ri + 1
  41. c = 3.0 / 8
  42. p = (ri - c) / (N - 2 * c + 1)
  43. y = math.sqrt(2) * scipy.special.erfinv(2 * p - 1)
  44. return y
  45. def p_adjust_fdr(p):
  46. # FDR correction for multiple testing
  47. # The implementation provides consistent results as R p.adjust function.
  48. # Use a new np array, otherwise p will be modified.
  49. p2 = np.zeros(p.shape, dtype=np.float32)
  50. idx = np.argsort(p)
  51. n = len(p)
  52. p2[idx] = (p[idx] * n) / np.arange(1, n + 1)
  53. p2[p2 > 1] = 1
  54. return p2
  55. def fdr_threshold(p, q):
  56. # p: vector of p-values
  57. # q: false discovery rate level
  58. #
  59. # return values
  60. # pID: p-value threshold based on independence or positive dependence
  61. # pN: nonparametric p-value threshold
  62. #
  63. # This function takes a vector of p-values and a FDR rate.
  64. # It returns two p-value thresholds, one based on an assumption of
  65. # independence or positive dependence, and one that makes no assumptions
  66. # about how the tests are correlated. For imaging data, an assumption of
  67. # positive dependence is reasonable, so it should be OK to use the first
  68. # (more sensitive) threshold.
  69. #
  70. # This implementation provides consistent results as the FDR function at
  71. # https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/nichols/software/fdr
  72. # https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/nichols/software/fdr/FDR.m
  73. p2 = p[~np.isnan(p)]
  74. p2 = np.sort(p2)
  75. n = len(p2)
  76. I = np.arange(1, n + 1)
  77. cVID = 1
  78. cVN = np.sum(1 / I)
  79. idx = np.nonzero(p2 <= ((I * q) / (n * cVID)))[0]
  80. pID = p2[np.max(idx)] if len(idx) >= 1 else 0
  81. idx = np.nonzero(p2 <= ((I * q) / (n * cVN)))[0]
  82. pN = p2[np.max(idx)] if len(idx) >= 1 else 0
  83. return pID, pN
  84. #%%
  85. ####################
  86. ### Loading Data ###
  87. ####################
  88. predictions_whole_body = pd.read_csv(paths.path_to_predictions_whole_body)
  89. predictions_brain = pd.read_csv(paths.path_to_predictions_brain)
  90. predictions_lungs = pd.read_csv(paths.organ_paths["lungs"])
  91. predictions_heart = pd.read_csv(paths.organ_paths["heart"])
  92. predictions_spine = pd.read_csv(paths.organ_paths["spine"])
  93. predictions_liver = pd.read_csv(paths.organ_paths["liver"])
  94. predictions_muscle = pd.read_csv(paths.organ_paths["muscle"])
  95. predictions_intestine = pd.read_csv(paths.organ_paths["intestine"])
  96. predictions_whole_body.rename(columns={"age_gap":"age_gap_whole_body"}, inplace=True)
  97. predictions_whole_body.rename(columns={"mean_prediction": "whole_body_pred"}, inplace=True)
  98. predictions_brain.rename(columns={"age_gap":"age_gap_brain"}, inplace=True)
  99. predictions_brain.rename(columns={"mean_prediction": "brain_pred"}, inplace=True)
  100. predictions_brain = predictions_brain[["eid", "age_gap_brain", "brain_pred"]]
  101. predictions_lungs.rename(columns={"age_gap":"age_gap_lungs"}, inplace=True)
  102. predictions_lungs.rename(columns={"mean_prediction": "lungs_pred"}, inplace=True)
  103. predictions_lungs = predictions_lungs[["eid", "age_gap_lungs", "lungs_pred"]]
  104. predictions_heart.rename(columns={"age_gap":"age_gap_heart"}, inplace=True)
  105. predictions_heart.rename(columns={"mean_prediction": "heart_pred"}, inplace=True)
  106. predictions_heart = predictions_heart[["eid", "age_gap_heart", "heart_pred"]]
  107. predictions_spine.rename(columns={"age_gap":"age_gap_spine"}, inplace=True)
  108. predictions_spine.rename(columns={"mean_prediction": "spine_pred"}, inplace=True)
  109. predictions_spine = predictions_spine[["eid", "age_gap_spine", "spine_pred"]]
  110. predictions_liver.rename(columns={"age_gap":"age_gap_liver"}, inplace=True)
  111. predictions_liver.rename(columns={"mean_prediction": "liver_pred"}, inplace=True)
  112. predictions_liver = predictions_liver[["eid", "age_gap_liver", "liver_pred"]]
  113. predictions_muscle.rename(columns={"age_gap":"age_gap_muscle"}, inplace=True)
  114. predictions_muscle.rename(columns={"mean_prediction": "muscle_pred"}, inplace=True)
  115. predictions_muscle = predictions_muscle[["eid", "age_gap_muscle", "muscle_pred"]]
  116. predictions_intestine.rename(columns={"age_gap":"age_gap_intestine"}, inplace=True)
  117. predictions_intestine.rename(columns={"mean_prediction": "intestine_pred"}, inplace=True)
  118. predictions_intestine = predictions_intestine[["eid", "age_gap_intestine", "intestine_pred"]]
  119. predictions = pd.merge(predictions_whole_body, predictions_brain, on="eid", how="left")
  120. predictions = pd.merge(predictions, predictions_lungs, on="eid", how="left")
  121. predictions = pd.merge(predictions, predictions_heart, on="eid", how="left")
  122. predictions = pd.merge(predictions, predictions_spine, on="eid", how="left")
  123. predictions = pd.merge(predictions, predictions_liver, on="eid", how="left")
  124. predictions = pd.merge(predictions, predictions_muscle, on="eid", how="left")
  125. predictions = pd.merge(predictions, predictions_intestine, on="eid", how="left")
  126. predictions = predictions[["eid", "age_gap_whole_body",
  127. "age_gap_brain",
  128. "whole_body_pred",
  129. "brain_pred",
  130. "age_gap_lungs", "lungs_pred",
  131. "age_gap_heart", "heart_pred",
  132. "age_gap_spine", "spine_pred",
  133. "age_gap_liver", "liver_pred",
  134. "age_gap_muscle", "muscle_pred",
  135. "age_gap_intestine", "intestine_pred"
  136. ]]
  137. #%%
  138. ################################################
  139. ### Merging data with non-imaging phenoytpes ###
  140. ################################################
  141. merged_data = pd.merge(predictions, pd.read_csv(paths.pain_path), on="eid", how="left")
  142. merged_data = pd.merge(merged_data, pd.read_csv(paths.sleep_path), on="eid", how="left")
  143. merged_data = pd.merge(merged_data, pd.read_csv(paths.sexual_factors_path), on="eid", how="left")
  144. merged_data = pd.merge(merged_data, pd.read_csv(paths.alcohol_path), on="eid", how="left")
  145. merged_data = pd.merge(merged_data, pd.read_csv(paths.diet_path), on="eid", how="left")
  146. merged_data = pd.merge(merged_data, pd.read_csv(paths.employment_path), on="eid", how="left")
  147. merged_data = pd.merge(merged_data, pd.read_csv(paths.family_history_path), on="eid", how="left")
  148. merged_data = pd.merge(merged_data, pd.read_csv(paths.female_sex_specific_path), on="eid", how="left")
  149. merged_data = pd.merge(merged_data, pd.read_csv(paths.general_health_path), on="eid", how="left")
  150. merged_data = pd.merge(merged_data, pd.read_csv(paths.male_sex_specific_path), on="eid", how="left")
  151. merged_data = pd.merge(merged_data, pd.read_csv(paths.medical_conditions_path), on="eid", how="left")
  152. merged_data = pd.merge(merged_data, pd.read_csv(paths.mental_health_path), on="eid", how="left")
  153. merged_data = pd.merge(merged_data, pd.read_csv(paths.smoking_path), on="eid", how="left")
  154. merged_data = pd.merge(merged_data, pd.read_csv(paths.social_support_path), on="eid", how="left")
  155. merged_data = pd.merge(merged_data, pd.read_csv(paths.sun_exposure_path), on="eid", how="left")
  156. merged_data = pd.merge(merged_data, pd.read_csv(paths.telomeres_path), on="eid", how="left")
  157. merged_data = pd.merge(merged_data, pd.read_csv(paths.month_of_birth_path), on="eid", how="left")
  158. merged_data = pd.merge(merged_data, pd.read_csv(paths.deaths_path, sep="\t"), on="eid", how="left")
  159. merged_data = pd.merge(merged_data, pd.read_csv(paths.body_composition_path), on="eid", how="left")
  160. met = pd.read_csv(paths.path_to_MET_features)
  161. merged_data = pd.merge(merged_data, met, on="eid", how="left")
  162. fiftythree = utils.get_feature("53-2.0")
  163. merged_data = pd.merge(merged_data, fiftythree, on="eid", how="left")
  164. loneliness = utils.get_feature("2020-2.0")
  165. merged_data = pd.merge(merged_data, loneliness, on="eid", how="left")
  166. year_of_birth = utils.get_feature("34-0.0")
  167. merged_data = pd.merge(merged_data, year_of_birth, on="eid", how="left")
  168. merged_data = merged_data.dropna(subset=["age_gap_lungs", "lungs_pred", "age_gap_heart", "heart_pred",
  169. "age_gap_spine", "spine_pred","age_gap_liver", "liver_pred",
  170. "age_gap_muscle", "muscle_pred",
  171. "age_gap_intestine", "intestine_pred",
  172. ])
  173. df = merged_data
  174. df.dropna(subset=["34-0.0", "52-0.0", "21002-2.0"], inplace=True)
  175. predictions = predictions[predictions["eid"].isin(df["eid"])]
  176. # %%
  177. # # # # # # # # # # # # # # # # # # # #
  178. # Step 3: confounding factors
  179. # # # # # # # # # # # # # # # # # # # #
  180. sex = df['31-0.0'].values
  181. # Age provided by UK Biobank (21003-2.0) seems to be floored, i.e. with half a year error.
  182. # To get more accurate age values, we calculate age by date.
  183. age = np.zeros(len(df))
  184. for i in range(len(df)):
  185. # Calculate age
  186. d1 = datetime.date(int(df.iloc[i]['34-0.0']), int(df.iloc[i]['52-0.0']), 15)
  187. s = df.iloc[i]['53-2.0']
  188. d2 = datetime.date(int(s[:4]), int(s[5:7]), int(s[8:]))
  189. age[i] = np.round((d2 - d1).days / 365.25, 1)
  190. weight = df['21002-2.0'].values
  191. bmi = df['21001-2.0'].values
  192. height = np.round(np.sqrt(weight / bmi) * 100)
  193. # %%
  194. valid_idx = ~np.isnan(age) & ~np.isnan(sex) & ~np.isnan(weight) & ~np.isnan(height)
  195. df = df[valid_idx]
  196. df_idp = predictions[valid_idx]
  197. sex = sex[valid_idx]
  198. age = age[valid_idx]
  199. sex_age = sex * age # Sex and age interaction
  200. weight = weight[valid_idx]
  201. height = height[valid_idx]
  202. conf = np.stack((sex, age, sex_age, weight, height), axis=1)
  203. df_conf = pd.DataFrame(conf, index=df.index, columns=['Sex', 'Age', 'Sex * Age', 'Weight', 'Height'])
  204. # %%
  205. # Remove confounding factors from df
  206. df = df.drop(columns=[
  207. '34-0.0',
  208. '52-0.0',
  209. '21002-2.0',
  210. 'whole_body_pred',
  211. 'brain_pred',
  212. 'lungs_pred',
  213. 'heart_pred',
  214. 'spine_pred',
  215. 'liver_pred',
  216. 'muscle_pred',
  217. 'intestine_pred',
  218. 'age_gap_brain',
  219. 'age_gap_whole_body',
  220. 'age_gap_lungs',
  221. 'age_gap_heart',
  222. 'age_gap_spine',
  223. 'age_gap_liver',
  224. "age_gap_muscle",
  225. "age_gap_intestine",
  226. 'date_of_death',
  227. 'ins_index',
  228. 'dsource',
  229. 'source',
  230. "eid",
  231. ])
  232. # %%
  233. # # # # # # # # # # # # # # # # # # # #
  234. # Step 4: clean, de-confound and normalise data
  235. # # # # # # # # # # # # # # # # # # # #
  236. # This part of the code was adapted from Karla Miller's Matlab code at
  237. # https://www.fmrib.ox.ac.uk/ukbiobank/gwaspaper/index.html
  238. # Step 4.1: cleaning
  239. n_subj, n_col = df.shape
  240. bad_vars = []
  241. for i in range(n_col):
  242. val = df.iloc[:, i]
  243. # Discard columns which not numbers
  244. if not np.issubdtype(df.dtypes[i], np.number):
  245. bad_vars += [i]
  246. continue
  247. # Assume negative values are invalid, set them to NaN
  248. # There are also a lot of empty values, which are already NaN
  249. val[val < 0] = np.nan
  250. df.iloc[:, i] = val
  251. # Valid indices
  252. valid_idx = ~np.isnan(val)
  253. # Discard columns with more than 90% missing data
  254. if np.sum(valid_idx) < (0.1 * n_subj):
  255. bad_vars += [i]
  256. continue
  257. # Discard columns with over 95% elements with the exactly same value
  258. val_unique, counts = np.unique(val[valid_idx], return_counts=True)
  259. if np.max(counts) >= (0.95 * np.sum(valid_idx)):
  260. bad_vars += [i]
  261. continue
  262. #%%
  263. merged_data["2644-2.0"]
  264. # merged_data.drop(columns=[["2644-2.0","20161-2.0"]], inplace=True)
  265. merged_data.drop(columns=["20161-2.0"], inplace=True)
  266. # %%
  267. for i in range(n_col):
  268. for j in range(i + 1, n_col):
  269. if i in bad_vars or j in bad_vars:
  270. continue
  271. # Discard columns with very high correlation
  272. val_i = df.iloc[:, i]
  273. val_j = df.iloc[:, j]
  274. valid_idx = ~np.isnan(val_i) & ~np.isnan(val_j)
  275. if np.sum(valid_idx) == 0:
  276. continue
  277. # print(val_i[valid_idx], val_j[valid_idx])
  278. cc, _ = scipy.stats.pearsonr(val_i[valid_idx], val_j[valid_idx])
  279. if cc > 0.9999:
  280. # Keep the column with more valid elements
  281. if np.sum(~np.isnan(val_i)) > np.sum(~np.isnan(val_j)):
  282. bad_vars += [j]
  283. else:
  284. bad_vars += [i]
  285. # %%
  286. # The cleaned data
  287. bad_vars = np.unique(bad_vars)
  288. keep_vars = sorted(list(set(np.arange(n_col)) - set(bad_vars)))
  289. df = df.iloc[:, keep_vars]
  290. print('{0} columns kept after data cleaning.'.format(df.shape[1]))
  291. # Step 4.2: normalise confounding factors
  292. conf = (conf - np.mean(conf, axis=0)) / np.std(conf, axis=0)
  293. #%%
  294. # Step 4.3: normalise non imaging phenotypes
  295. df_cont = pd.read_csv(paths.path_to_continuous_features, index_col=0)
  296. #%%
  297. n_col = df.shape[1]
  298. for i in range(n_col):
  299. val = df.iloc[:, i]
  300. valid_idx = ~np.isnan(val)
  301. x = val[valid_idx]
  302. # Field ID
  303. field_id = int(df.columns[i][1].split('-')[0])
  304. is_continuous = df_cont.loc[field_id]['continuous']
  305. if is_continuous:
  306. # If it is a continuous variable, perform standard normalisation.
  307. x = normalise(x)
  308. else:
  309. # If we are not sure whether it is a continuous or categorical variable, convert it
  310. # into a continuous variable using rank-based inverse normal transform.``
  311. x = rank_normalise(x)
  312. df.iloc[:, i][valid_idx] = x# The cleaned data
  313. #%%
  314. # Step 4.4: de-confound and normalise IDPs
  315. try:
  316. df_idp = df_idp.drop(columns=["eid"])
  317. df_idp = df_idp.drop(columns=["split"])
  318. except:
  319. print("Columns already dropped")
  320. n_row = conf.shape[1]
  321. n_col = df_idp.shape[1]
  322. beta = np.zeros((n_row, n_col))
  323. for i in range(n_col):
  324. val = df_idp.iloc[:, i]
  325. print(val)
  326. valid_idx = ~np.isnan(val)
  327. x = val[valid_idx]
  328. beta[:, i] = np.dot(np.linalg.pinv(conf[valid_idx]), x)
  329. x = x - np.dot(conf[valid_idx], beta[:, i])
  330. x = normalise(x)
  331. df_idp.iloc[:, i][valid_idx] = x
  332. df_idp = df_idp[[
  333. "age_gap_brain",
  334. "age_gap_whole_body",
  335. "age_gap_lungs",
  336. "age_gap_heart",
  337. "age_gap_spine",
  338. "age_gap_liver",
  339. "age_gap_muscle",
  340. "age_gap_intestine"
  341. ]]
  342. # %%
  343. # # # # # # # # # # # # # # # # # # # # #
  344. # # Step 5: uni-variate correlation studies
  345. # # # # # # # # # # # # # # # # # # # # #
  346. M = df_idp.shape[1]
  347. N = df.shape[1]
  348. corr = np.zeros((M, N))
  349. corr_p = np.zeros((M, N))
  350. # features = []
  351. for i in range(M):
  352. features = []
  353. for j in range(N):
  354. # Remove NaNs
  355. x = df_idp.iloc[:, i]
  356. y = df.iloc[:, j]
  357. valid_idx = ~np.isnan(x) & ~np.isnan(y)
  358. x = x[valid_idx]
  359. y = y[valid_idx]
  360. # Pearson correlation
  361. cc, p_val = scipy.stats.pearsonr(x, y)
  362. corr[i, j] = cc
  363. corr_p[i, j] = p_val
  364. features.append(df.columns[j])
  365. # For p-value of 0, assign it with the mininal positive floating value
  366. # so that we can calculate the logarithm for the Manhattan plot
  367. corr_p[corr_p == 0] = np.finfo(np.float64).tiny
  368. # Logarithm
  369. log_corr_p = - np.log10(corr_p)
  370. # Save the tables
  371. df_corr = pd.DataFrame(corr, index=df_idp.columns, columns=df.columns)
  372. df_p = pd.DataFrame(corr_p, index=df_idp.columns, columns=df.columns)
  373. df_log_p = pd.DataFrame(log_corr_p, index=df_idp.columns, columns=df.columns)
  374. corr = df_corr.values
  375. corr_p = df_p.values
  376. log_corr_p = df_log_p.values
  377. # Bonferroni correction
  378. M, N = corr.shape
  379. p_bonf = 0.05 / (M * N)
  380. # FDR correction
  381. p_fdr, _ = fdr_threshold(corr_p.flatten(), 0.05)
  382. # Number of phenotypes that is significantly associated with at least one of the IDPs
  383. print('p_bonf = {0}'.format(p_bonf))
  384. print('p_fdr = {0}'.format(p_fdr))
  385. print('Number of correlations reaching Bonferroni threshold = {0}'.format(np.sum(corr_p < p_bonf)))
  386. print('Number of correlations reaching FDR threshold = {0}'.format(np.sum(corr_p < p_fdr)))
  387. print('Number of phenotypes reaching Bonferroni threshold = {0}'.format(np.sum(np.sum(corr_p < p_bonf, axis=0) > 0)))
  388. print('Number of phenotypes reaching FDR threshold = {0}'.format(np.sum(np.sum(corr_p < p_fdr, axis=0) > 0)))
  389. # %%
  390. # # # # # # # # # # # # # # # # # # # #
  391. # Step 6: Manhattan plot
  392. # # # # # # # # # # # # # # # # # # # #
  393. # The category for each column
  394. # They should be in ascending order. Otherwise, we need to sort the category IDs.
  395. pain_columns = pd.read_csv(paths.pain_path).columns
  396. sleep_columns = pd.read_csv(paths.sleep_path).columns
  397. sexual_columns = pd.read_csv(paths.sexual_factors_path).columns
  398. alcohol_columns = pd.read_csv(paths.alcohol_path).columns
  399. diet_columns = pd.read_csv(paths.diet_path).columns
  400. employment_columns = pd.read_csv(paths.employment_path).columns
  401. family_history_columns = pd.read_csv(paths.family_history_path).columns
  402. smoking_columns = pd.read_csv(paths.smoking_path).columns
  403. female_specific_columns = pd.read_csv(paths.female_sex_specific_path).columns
  404. general_health_columns = pd.read_csv(paths.general_health_path).columns
  405. male_specific_columns = pd.read_csv(paths.male_sex_specific_path).columns
  406. medical_conditions_columns = pd.read_csv(paths.medical_conditions_path).columns
  407. mental_health_columns = pd.read_csv(paths.mental_health_path).columns
  408. social_support_columns = pd.read_csv(paths.social_support_path).columns
  409. sun_exposure_columns = pd.read_csv(paths.sun_exposure_path).columns
  410. telomeres_columns = pd.read_csv(paths.telomeres_path).columns
  411. body_composition_columns = pd.read_csv(paths.body_composition_path).columns
  412. met = pd.read_csv(paths.path_to_MET_features).columns
  413. def get_category(field_id):
  414. if field_id in pain_columns:
  415. return "pain"
  416. elif field_id in sleep_columns:
  417. return "sleep"
  418. elif field_id in sexual_columns:
  419. return "sexual"
  420. elif field_id in alcohol_columns:
  421. return "alcohol"
  422. elif field_id in diet_columns:
  423. return "diet"
  424. elif field_id in employment_columns:
  425. return "employment"
  426. elif field_id in family_history_columns:
  427. return "family_history"
  428. elif field_id in smoking_columns:
  429. return "smoking"
  430. elif field_id in female_specific_columns:
  431. return "female_specific"
  432. elif field_id in general_health_columns:
  433. return "general_health"
  434. elif field_id in male_specific_columns:
  435. return "male_specific"
  436. elif field_id in medical_conditions_columns:
  437. return "medical_conditions"
  438. elif field_id in mental_health_columns:
  439. return "mental_health"
  440. elif field_id in social_support_columns or field_id=="2020-2.0" or field_id=="2020-2.0" or field_id=="6160-2.0_y" or field_id=="6160-0.0":
  441. return "social_support"
  442. elif field_id in sun_exposure_columns:
  443. return "sun_exposure"
  444. elif field_id in telomeres_columns:
  445. return "telomeres"
  446. elif field_id in body_composition_columns:
  447. return "body_composition"
  448. elif field_id in met:
  449. return "physical activity"
  450. else:
  451. return "general"
  452. # The Manhattan plot
  453. table = []
  454. for i in range(M):
  455. if df_idp.columns[i] == 'age_gap_whole_body':
  456. s = 'whole_body'
  457. elif df_idp.columns[i] == 'age_gap_brain':
  458. s = 'brain'
  459. elif df_idp.columns[i] == 'age_gap_lungs':
  460. s = 'lungs'
  461. elif df_idp.columns[i] == 'age_gap_heart':
  462. s = 'heart'
  463. elif df_idp.columns[i] == 'age_gap_spine':
  464. s = 'spine'
  465. elif df_idp.columns[i] == 'age_gap_liver':
  466. s = 'liver'
  467. elif df_idp.columns[i] == 'age_gap_muscle':
  468. s = 'muscle'
  469. elif df_idp.columns[i] == 'age_gap_intestine':
  470. s = 'intestine'
  471. for j in range(N):
  472. line = [j, log_corr_p[i, j], corr[i, j], abs(corr[i, j]), s, features[j]]
  473. table += [line]
  474. df_table = pd.DataFrame(table, columns=['x', 'log_p', 'corr', 'Correlation', 'Modality', "Feature"])
  475. #%%
  476. category = np.array([get_category(field_id) for field_id in tqdm(df_table["Feature"])])
  477. df_table["Category"] = category
  478. df_table = df_table.sort_values(by=["Category","Feature"])
  479. #%%
  480. df_table = df_table.sort_values(by=["Category","Feature", "Modality"])
  481. #%%
  482. df_table["x2"] = np.repeat(np.arange(1, N+1), M)
  483. #%%
  484. df_table = df_table.sample(frac=1)
  485. #%%
  486. df_table.reset_index(drop=True, inplace=True)
  487. #%%
  488. pallette = {
  489. 'heart': '#D81B60',
  490. 'muscle': "#1E88E5",
  491. 'whole_body': "#FFBF00",
  492. 'liver': '#004D40',
  493. 'brain': '#979E41',
  494. 'spine': '#51D9B8',
  495. 'lungs': '#AA58B3',
  496. 'intestine': '#C56800',
  497. }
  498. #%%
  499. plt.figure(figsize=(16, 8))
  500. ax = sns.scatterplot(x='x2', y='log_p', hue='Modality', size='Correlation',
  501. sizes=(30, 200), size_norm=mpl.colors.Normalize(vmin=0, vmax=0.3),
  502. data=df_table, alpha=0.8, palette=pallette)
  503. plt.plot([0, N], [-math.log10(p_bonf), -math.log10(p_bonf)], 'k--', linewidth=1, alpha=0.9)
  504. plt.text(-1, -math.log10(p_bonf), 'Bonf', horizontalalignment='right', fontsize=15)
  505. handles, labels = ax.get_legend_handles_labels()
  506. # labels[-1] = '0.3'
  507. ax.legend(handles, labels)
  508. xticks = []
  509. xticklabels = []
  510. category_list = df_table["Category"].values
  511. unique_category = np.unique(category_list)
  512. print(unique_category)
  513. for c in unique_category:
  514. x = np.max(df_table[df_table["Category"]==c]["x2"]) + 0.5
  515. plt.plot([x, x], [0, 310], 'k-.', linewidth=0.5, alpha=0.2)
  516. xticks += [np.mean(df_table[df_table["Category"]==c]["x2"])]
  517. xticklabels += [c]
  518. plt.xlim(0, N)
  519. plt.ylim(0, 110)
  520. plt.xticks(xticks, xticklabels, fontsize=15, rotation=90)
  521. plt.yticks(fontsize=15)
  522. plt.xlabel('Non-imaging phenotypes', fontsize=15, fontweight='bold')
  523. plt.ylabel(r'$-\log_{10}(p)$', fontsize=15, fontweight='bold')
  524. fig = plt.gcf()
  525. fig.set_size_inches(16, 8)
  526. plt.tight_layout()
  527. plt.savefig('../../outputs/figures/manhattan_phewas.pdf', bbox_inches='tight')
  528. plt.show()
  529. # %%
  530. plt.figure(figsize=(16, 8))
  531. ax = sns.scatterplot(x='x2', y='corr', hue='Modality', size='Correlation',
  532. sizes=(30, 100), size_norm=mpl.colors.Normalize(vmin=0, vmax=0.2),
  533. data=df_table, alpha=0.8, palette=palette)
  534. plt.hlines(0, xmin=0, xmax=N*2, linewidth=1, alpha=0.8, color="gray", linestyle='--')
  535. handles, labels = ax.get_legend_handles_labels()
  536. # labels[-1] = '0.25'
  537. ax.legend(handles, labels)
  538. xticks = []
  539. xticklabels = []
  540. category_list = df_table["Category"].values
  541. unique_category = np.unique(category_list)
  542. print(unique_category)
  543. for c in unique_category:
  544. x = np.max(df_table[df_table["Category"]==c]["x2"]) + 0.5
  545. plt.plot([x, x], [-0.15, 0.21], 'k-.', linewidth=0.5)
  546. xticks += [np.mean(df_table[df_table["Category"]==c]["x2"])]
  547. xticklabels += [c]
  548. plt.xlim(0, N)
  549. plt.xticks(xticks, xticklabels, fontsize=15, rotation=90)
  550. plt.yticks(fontsize=15)
  551. plt.xlabel('Non-imaging phenotypes', fontsize=15, fontweight='bold')
  552. plt.ylabel('Correlation', fontsize=15, fontweight='bold')
  553. fig = plt.gcf()
  554. fig.set_size_inches(16, 8)
  555. plt.tight_layout()
  556. plt.savefig('../../outputs/figures/correlations_phewas.pdf', bbox_inches='tight')
  557. plt.show()

perform_phewas.py at commit a772251, no license · at the source

Overview

Authors: Tamara T. Mueller1, Sophie Starck1, Rozafë Llalloshi1, Georgios Kaissis1, Alexander Ziller1, Robert Graf1,2, Christopher Schlett3, Steffen Ringhof3, Fabian Bamberg3, Mark Wielpütz4, Henry Völzke4, Michael Leitzmann5, Thoralf Niendorf6, Thomas Keil7,8,9, Lilian Krist7, Tobias Pischon6, André Karch10, Klaus Berger10, Jan Kirschke2, Daniel Rueckert1,11, Rickmer Braren12,13,14
14 affiliations
  1. Lab for AI in Medicine and Healthcare, Technical University of Munich,Munich, Germany
  2. Department of Diagnostic and Interventional Neuroradiology, School of Medicine, TUM University Hospital,Munich, Germany
  3. Department of Diagnostic and Interventional Radiology, Medical Center, University of Freiburg, Faculty of Medicine,Freiburg, Germany
  4. Institut für Community Medicine, University Medicine Greifswald,Greifswald, Germany
  5. Institute for Epidemiology and Preventive Medicine, University of Regensburg,Regensburg, Germany
  6. Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine, Helmholtz Association,Berlin, Germany
  7. Institute of Social Medicine, Epidemiology and Health Economics, Charité - Universitätsmedizin Berlin,Berlin, Germany
  8. Institute of Clinical Epidemiology and Biometry, University of Würzburg,Würzburg, Germany
  9. State Institute of Health I, Bavarian Health and Food Safety Authority,Erlangen, Germany
  10. Institute of Epidemiology and Social Medicine, University of Münster,Münster, Germany
  11. Department of Computing, Imperial College London,London, United Kingdom
  12. Institute for Diagnostic and Interventional Radiology, Klinikum rechts der Isar, Technical University of Munich,Munich, Germany
  13. German Cancer Consortium (DKTK), Munich partner site,Heidelberg, Germany
  14. Department of Diagnostic and Interventional Radiology and Nuclear Medicine, University Medical Center Hamburg-Eppendorf,Hamburg, Germany
Journal: Communications medicine, volume 6, issue 1, article 335
Dates: received 2 June 2025; accepted 3 June 2026; published online 13 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s43856-026-01722-3 · PMID 42288701 · PMCID PMC13264634 · OpenAlex W4410519636
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), clinical / translational (subfield)
Methods: Connectivity, Statistics, Machine learning, fMRI & imaging
Keywords: Biomarkers
Topic: Health, Environment, Cognitive Aging (Health, Toxicology and Mutagenesis, Environmental Science), according to OpenAlex
Funding: European Research Council (884622)
Citations: not cited yet (Europe PMC); 54 references in the paper

Abstract

Background: Understanding human ageing across multiple organs is essential for characterising individual health trajectories and identifying abnormal ageing processes. Multi-organ imaging provides an opportunity to quantify biological ageing beyond chronological age. The aim of this study is to assess organ-specific and whole-body ageing patterns and their associations with disease and lifestyle factors.

Methods: In this large-scale study, we evaluate biological ageing patterns using 70,000 MRI scans from the UK Biobank and the German National Cohort. We employ 3D ResNet-18 models to predict chronological age from various body regions (brain, heart, liver, spine, lungs, muscle, and intestine) and the whole body. From these predictions, we derive “age gaps” relative to a strictly healthy reference cohort, which enables the identification of accelerated ageing patterns. We then evaluate associations with chronic diseases and lifestyle factors, and a virtual ageing framework was developed to explore counterfactual scenarios by substituting anatomical regions across subjects, quantifying local impacts on global biological age.

Results: Here we show significant associations between detected accelerated ageing and specific chronic diseases, including multiple sclerosis and chronic obstructive pulmonary disease, as well as lifestyle factors such as smoking and physical activity. Virtual substitution of anatomical regions demonstrates that local substitutions can influence global ageing patterns.

Conclusions: This study demonstrates that multi-organ imaging enables the detection of abnormal ageing patterns at both local and global levels. The presented framework provides a foundation for improved risk stratification and supports the development of personalised approaches to health assessment and disease prevention.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

tamaramueller/age_biomarkers

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: a772251fc863ac522f143f918919b7659ec45cb3, 21 April 2026
Languages: Python (35)
Size: 47 files, 35 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (source/environment.yaml)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: PyTorch (18 files), NumPy (16 files), pandas (11 files), NiBabel (7 files), SciPy (6 files), PyTorch Lightning (5 files), Matplotlib (5 files), seaborn (4 files), scikit-image (2 files), SimpleITK (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
36 files

Code availability

All files can be found in the corresponding source code repository at: https://github.com/tamaramueller/age_biomarkers/54.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 35 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

A were created with Biorender, detailed acknowledgement will be provided upon publication. Eligible researchers may access UK Biobank data on https://www.ukbiobank.ac.uk, upon registration. For this study, permission to access and analyse the UK Biobank data was approved under the application 87802. This project was conducted with data (Application No. NAKO-732) from the German National Cohort (NAKO) (https://www.nako.de). The NAKO is funded by the Federal Ministry of Education and Research (BMBF) [project funding reference numbers: 01ER1301A/B/C, 01ER1511D, 01ER1801A/B/C/D and 01ER2301A/B/C], federal states of Germany and the Helmholtz Association, the participating universities and the institutes of the Leibniz Association. We thank all participants who took part in the NAKO study and the staff of this research initiative. Figures 1 and 8A were created with Biorender, detailed acknowledgement will be provided upon publication. Figure 1 was created in BioRender. Mueller, T. (2025) https://BioRender.com/8ryioo7. Figure 8A was created in BioRender. Mueller, T. (2025) https://BioRender.com/8rm2rn2.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 21 authors, 1 keyword, 1 funder, 43 references.

Cite

This paper

Mueller, T. T., Starck, S., Llalloshi, R., Kaissis, G., Ziller, A., Graf, R., Schlett, C., Ringhof, S., Bamberg, F., Wielpütz, M., Völzke, H., Leitzmann, M., Niendorf, T., Keil, T., Krist, L., Pischon, T., Karch, A., Berger, K., Kirschke, J., . . . Braren, R. (2026). Local and global patterns support medical imaging as a biomarker of ageing. Communications medicine, 6(1), 335. https://doi.org/10.1038/s43856-026-01722-3

BibTeX

@article{mueller2026local,
author = {Mueller, Tamara T. and Starck, Sophie and Llalloshi, Rozafë and Kaissis, Georgios and Ziller, Alexander and Graf, Robert and Schlett, Christopher and Ringhof, Steffen and Bamberg, Fabian and Wielpütz, Mark and Völzke, Henry and Leitzmann, Michael and Niendorf, Thoralf and Keil, Thomas and Krist, Lilian and Pischon, Tobias and Karch, André and Berger, Klaus and Kirschke, Jan and Rueckert, Daniel and Braren, Rickmer},
title = {{Local and global patterns support medical imaging as a biomarker of ageing}},
journal = {Communications medicine},
year = {2026},
month = jun,
volume = {6},
number = {1},
pages = {335},
publisher = {Nature Publishing Group},
issn = {2730-664X},
doi = {10.1038/s43856-026-01722-3},
url = {https://doi.org/10.1038/s43856-026-01722-3},
pmid = {42288701},
pmcid = {PMC13264634}
}

RIS

TY - JOUR
AU - Mueller, Tamara T.
AU - Starck, Sophie
AU - Llalloshi, Rozafë
AU - Kaissis, Georgios
AU - Ziller, Alexander
AU - Graf, Robert
AU - Schlett, Christopher
AU - Ringhof, Steffen
AU - Bamberg, Fabian
AU - Wielpütz, Mark
AU - Völzke, Henry
AU - Leitzmann, Michael
AU - Niendorf, Thoralf
AU - Keil, Thomas
AU - Krist, Lilian
AU - Pischon, Tobias
AU - Karch, André
AU - Berger, Klaus
AU - Kirschke, Jan
AU - Rueckert, Daniel
AU - Braren, Rickmer
TI - Local and global patterns support medical imaging as a biomarker of ageing
T2 - Communications medicine
J2 - Commun Med (Lond)
PY - 2026
DA - 2026/06/13
VL - 6
IS - 1
SP - 335
SN - 2730-664X
PB - Nature Publishing Group
DO - 10.1038/s43856-026-01722-3
UR - https://doi.org/10.1038/s43856-026-01722-3
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s43856-026-01722-3",
"type": "article-journal",
"title": "Local and global patterns support medical imaging as a biomarker of ageing",
"container-title": "Communications medicine",
"author": [
{
"family": "Mueller",
"given": "Tamara T."
},
{
"family": "Starck",
"given": "Sophie"
},
{
"family": "Llalloshi",
"given": "Rozafë"
},
{
"family": "Kaissis",
"given": "Georgios"
},
{
"family": "Ziller",
"given": "Alexander"
},
{
"family": "Graf",
"given": "Robert"
},
{
"family": "Schlett",
"given": "Christopher"
},
{
"family": "Ringhof",
"given": "Steffen"
},
{
"family": "Bamberg",
"given": "Fabian"
},
{
"family": "Wielpütz",
"given": "Mark"
},
{
"family": "Völzke",
"given": "Henry"
},
{
"family": "Leitzmann",
"given": "Michael"
},
{
"family": "Niendorf",
"given": "Thoralf"
},
{
"family": "Keil",
"given": "Thomas"
},
{
"family": "Krist",
"given": "Lilian"
},
{
"family": "Pischon",
"given": "Tobias"
},
{
"family": "Karch",
"given": "André"
},
{
"family": "Berger",
"given": "Klaus"
},
{
"family": "Kirschke",
"given": "Jan"
},
{
"family": "Rueckert",
"given": "Daniel"
},
{
"family": "Braren",
"given": "Rickmer"
}
],
"container-title-short": "Commun Med (Lond)",
"volume": "6",
"issue": "1",
"page": "335",
"DOI": "10.1038/s43856-026-01722-3",
"PMID": "42288701",
"PMCID": "PMC13264634",
"ISSN": "2730-664X",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s43856-026-01722-3",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
13
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1002/hbm.70425 [code]
Brain Age Estimation on T2-FLAIR Scans for Application to Multiple Sclerosis.
Journal: Human brain mapping
In common: PyTorch, pandas, SciPy, 1 other tool, clinical / translational, 8 references
[2] doi:10.1186/s40708-026-00316-y [code]
Generalizable and explainable deep learning for brain MRI: a multi-cohort evaluation of 3D architectures for age and sex prediction.
Journal: Brain informatics
In common: PyTorch, seaborn, pandas, 3 other tools, 6 references
[3] doi:10.1162/imag.a.1164 [code]
Bias and generalizability of brain age prediction models: A multi-cohort evaluation with anatomical and interpretability insights.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: NiBabel, statsmodels, PyTorch, 5 other tools, 4 references
[4] doi:10.1038/s41380-026-03691-4 [code]
Breaking the norm: population-scale deviations of brain structure in depression and anxiety.
Journal: Molecular psychiatry
In common: NiBabel, statsmodels, PyTorch, 5 other tools, clinical / translational, author Tobias Pischon
[5] doi:10.1038/s41398-026-04131-1 [code]
Multimodal phenotypic classification of generalized anxiety and panic using structural MRI data and psychosocial factors: machine learning results from the German National Cohort (NAKO) study.
Journal: Translational psychiatry
In common: clinical / translational, 1 reference, 2 authors
[6] doi:10.1007/s12021-026-09817-x [code]
Circle of Willis-Guided Localization for Simultaneous Detection and Classification of Large Vessel Occlusions in Brain CTA.
Journal: Neuroinformatics
In common: PyTorch Lightning, SimpleITK, scikit-image, 7 other tools
[7] doi:10.1093/radadv/umag025 [code]
OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI.
Journal: Radiology advances
In common: SimpleITK, NiBabel, PyTorch, 4 other tools, 3 references
[8] doi:10.3389/frai.2026.1771088 [code]
Few-shot deployment of pretrained MRI transformers in brain imaging tasks.
Journal: Frontiers in artificial intelligence
In common: SimpleITK, scikit-image, NiBabel, 6 other tools, 1 reference
[9] doi:10.1038/s41398-026-04081-8 [code]
Functional system-specific brain aging across the Alzheimer's disease continuum.
Journal: Translational psychiatry
In common: NiBabel, statsmodels, seaborn, 4 other tools, clinical / translational, 3 references
[10] doi:10.1038/s42003-026-10205-z [code]
Source-space EEG alpha activity reveals brain age gaps due to neurodegeneration and disparity.
Journal: Communications biology
In common: statsmodels, seaborn, pandas, 3 other tools, 4 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.