Local and global patterns support medical imaging as a biomarker of ageing.
The 6 matches
- [1] § Results › Lifestyle and Environment ↔ source/analysis/perform_phewas.py, lines 154–213 · score 0.96 · family history, sun exposure, mental health, sexual factors, social support, general health
- [2] § Results › Chronic Diseases ↔ source/analysis/disease_analysis.py, lines 51–78 · score 0.92 · ischemic heart disease, liver disease subjects, COPD subjects, MS subjects, scoliosis subjects, hypertension subjects
- [3] § Methods › Lifestyle and Environment ↔ source/analysis/perform_phewas.py, lines 320–399 · score 0.63 · rank normalisation, imaging phenotypes, variables, confounding, PheWAS, correlate
- [4] § Methods › Deep Learning Models ↔ source/cnn/model/resnet.py, lines 185–193 · score 0.61 · resnet18, ImageNet, pre trained, models
- [5] § Methods › Survival Analysis ↔ source/analysis/survival.py, lines 1–44 · score 0.60 · Cox proportional hazards, survival, death, female, brain, body
- [6] § Methods › Dataset › UK Biobank ↔ source/cnn/train_utils.py, lines 272–295 · score 0.57 · largest connected component, voxels, mask, training
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 612 lines · 24 KB · no license · 2 matches
- # %%
- import numpy as np
- import pandas as pd
- import ast
- import os
- from tqdm import tqdm
- import seaborn as sns
- import scipy.stats
- import math
- import csv
- import matplotlib.pyplot as plt
- import utils
- import paths
- #%%
- ############################################################################################################################################################################
- #### This script is used to generate the Manhattan plot and the correlation plot between non-imaging phenotypes and imaging phenotypes ###
- #### The script is adapted from Bai et al. 2020, "A population-based phenome-wide association study of cardiac and aortic structure and function", Nature medicine. ###
- ############################################################################################################################################################################
- #%%
- def normalise(x):
- return (x - np.mean(x)) / np.std(x)
- def rank_normalise(x):
- # Rank-based inverse normal transform
- # Please refer to the function inormal() in the FSLNets package
- # Get the rank of the values in x
- ri = np.argsort(np.argsort(x))
- # Correct for the ranks of repeated values
- # argsort assign different ranks for these values
- # We fill them with the same value
- u, inv_idx = np.unique(x, return_inverse=True)
- sii = np.sort(inv_idx)
- repeated_idx = np.unique(sii[np.diff(np.append(sii, 1)) == 0])
- for i in repeated_idx:
- ri[inv_idx == i] = np.mean(ri[inv_idx == i])
- # Perform inverse normal transform
- # ri + 1 so that the rank starts from 1, to be consistent with Karla's Matlab code
- # p squashes the rank into the range of [0, 1]
- # erfinv can generate a distribution from 2 * p - 1 with 0 mean and 1 standard deviation
- N = len(x)
- ri = ri + 1
- c = 3.0 / 8
- p = (ri - c) / (N - 2 * c + 1)
- y = math.sqrt(2) * scipy.special.erfinv(2 * p - 1)
- return y
- def p_adjust_fdr(p):
- # FDR correction for multiple testing
- # The implementation provides consistent results as R p.adjust function.
- # Use a new np array, otherwise p will be modified.
- p2 = np.zeros(p.shape, dtype=np.float32)
- idx = np.argsort(p)
- n = len(p)
- p2[idx] = (p[idx] * n) / np.arange(1, n + 1)
- p2[p2 > 1] = 1
- return p2
- def fdr_threshold(p, q):
- # p: vector of p-values
- # q: false discovery rate level
- #
- # return values
- # pID: p-value threshold based on independence or positive dependence
- # pN: nonparametric p-value threshold
- #
- # This function takes a vector of p-values and a FDR rate.
- # It returns two p-value thresholds, one based on an assumption of
- # independence or positive dependence, and one that makes no assumptions
- # about how the tests are correlated. For imaging data, an assumption of
- # positive dependence is reasonable, so it should be OK to use the first
- # (more sensitive) threshold.
- #
- # This implementation provides consistent results as the FDR function at
- # https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/nichols/software/fdr
- # https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/nichols/software/fdr/FDR.m
- p2 = p[~np.isnan(p)]
- p2 = np.sort(p2)
- n = len(p2)
- I = np.arange(1, n + 1)
- cVID = 1
- cVN = np.sum(1 / I)
- idx = np.nonzero(p2 <= ((I * q) / (n * cVID)))[0]
- pID = p2[np.max(idx)] if len(idx) >= 1 else 0
- idx = np.nonzero(p2 <= ((I * q) / (n * cVN)))[0]
- pN = p2[np.max(idx)] if len(idx) >= 1 else 0
- return pID, pN
- #%%
- ####################
- ### Loading Data ###
- ####################
- predictions_whole_body = pd.read_csv(paths.path_to_predictions_whole_body)
- predictions_brain = pd.read_csv(paths.path_to_predictions_brain)
- predictions_lungs = pd.read_csv(paths.organ_paths["lungs"])
- predictions_heart = pd.read_csv(paths.organ_paths["heart"])
- predictions_spine = pd.read_csv(paths.organ_paths["spine"])
- predictions_liver = pd.read_csv(paths.organ_paths["liver"])
- predictions_muscle = pd.read_csv(paths.organ_paths["muscle"])
- predictions_intestine = pd.read_csv(paths.organ_paths["intestine"])
- predictions_whole_body.rename(columns={"age_gap":"age_gap_whole_body"}, inplace=True)
- predictions_whole_body.rename(columns={"mean_prediction": "whole_body_pred"}, inplace=True)
- predictions_brain.rename(columns={"age_gap":"age_gap_brain"}, inplace=True)
- predictions_brain.rename(columns={"mean_prediction": "brain_pred"}, inplace=True)
- predictions_brain = predictions_brain[["eid", "age_gap_brain", "brain_pred"]]
- predictions_lungs.rename(columns={"age_gap":"age_gap_lungs"}, inplace=True)
- predictions_lungs.rename(columns={"mean_prediction": "lungs_pred"}, inplace=True)
- predictions_lungs = predictions_lungs[["eid", "age_gap_lungs", "lungs_pred"]]
- predictions_heart.rename(columns={"age_gap":"age_gap_heart"}, inplace=True)
- predictions_heart.rename(columns={"mean_prediction": "heart_pred"}, inplace=True)
- predictions_heart = predictions_heart[["eid", "age_gap_heart", "heart_pred"]]
- predictions_spine.rename(columns={"age_gap":"age_gap_spine"}, inplace=True)
- predictions_spine.rename(columns={"mean_prediction": "spine_pred"}, inplace=True)
- predictions_spine = predictions_spine[["eid", "age_gap_spine", "spine_pred"]]
- predictions_liver.rename(columns={"age_gap":"age_gap_liver"}, inplace=True)
- predictions_liver.rename(columns={"mean_prediction": "liver_pred"}, inplace=True)
- predictions_liver = predictions_liver[["eid", "age_gap_liver", "liver_pred"]]
- predictions_muscle.rename(columns={"age_gap":"age_gap_muscle"}, inplace=True)
- predictions_muscle.rename(columns={"mean_prediction": "muscle_pred"}, inplace=True)
- predictions_muscle = predictions_muscle[["eid", "age_gap_muscle", "muscle_pred"]]
- predictions_intestine.rename(columns={"age_gap":"age_gap_intestine"}, inplace=True)
- predictions_intestine.rename(columns={"mean_prediction": "intestine_pred"}, inplace=True)
- predictions_intestine = predictions_intestine[["eid", "age_gap_intestine", "intestine_pred"]]
- predictions = pd.merge(predictions_whole_body, predictions_brain, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_lungs, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_heart, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_spine, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_liver, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_muscle, on="eid", how="left")
- predictions = pd.merge(predictions, predictions_intestine, on="eid", how="left")
- predictions = predictions[["eid", "age_gap_whole_body",
- "age_gap_brain",
- "whole_body_pred",
- "brain_pred",
- "age_gap_lungs", "lungs_pred",
- "age_gap_heart", "heart_pred",
- "age_gap_spine", "spine_pred",
- "age_gap_liver", "liver_pred",
- "age_gap_muscle", "muscle_pred",
- "age_gap_intestine", "intestine_pred"
- ]]
- #%%
- ################################################
- ### Merging data with non-imaging phenoytpes ###
- ################################################
- merged_data = pd.merge(predictions, pd.read_csv(paths.pain_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.sleep_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.sexual_factors_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.alcohol_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.diet_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.employment_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.family_history_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.female_sex_specific_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.general_health_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.male_sex_specific_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.medical_conditions_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.mental_health_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.smoking_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.social_support_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.sun_exposure_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.telomeres_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.month_of_birth_path), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.deaths_path, sep="\t"), on="eid", how="left")
- merged_data = pd.merge(merged_data, pd.read_csv(paths.body_composition_path), on="eid", how="left")
- met = pd.read_csv(paths.path_to_MET_features)
- merged_data = pd.merge(merged_data, met, on="eid", how="left")
- fiftythree = utils.get_feature("53-2.0")
- merged_data = pd.merge(merged_data, fiftythree, on="eid", how="left")
- loneliness = utils.get_feature("2020-2.0")
- merged_data = pd.merge(merged_data, loneliness, on="eid", how="left")
- year_of_birth = utils.get_feature("34-0.0")
- merged_data = pd.merge(merged_data, year_of_birth, on="eid", how="left")
- merged_data = merged_data.dropna(subset=["age_gap_lungs", "lungs_pred", "age_gap_heart", "heart_pred",
- "age_gap_spine", "spine_pred","age_gap_liver", "liver_pred",
- "age_gap_muscle", "muscle_pred",
- "age_gap_intestine", "intestine_pred",
- ])
- df = merged_data
- df.dropna(subset=["34-0.0", "52-0.0", "21002-2.0"], inplace=True)
- predictions = predictions[predictions["eid"].isin(df["eid"])]
- # %%
- # # # # # # # # # # # # # # # # # # # #
- # Step 3: confounding factors
- # # # # # # # # # # # # # # # # # # # #
- sex = df['31-0.0'].values
- # Age provided by UK Biobank (21003-2.0) seems to be floored, i.e. with half a year error.
- # To get more accurate age values, we calculate age by date.
- age = np.zeros(len(df))
- for i in range(len(df)):
- # Calculate age
- d1 = datetime.date(int(df.iloc[i]['34-0.0']), int(df.iloc[i]['52-0.0']), 15)
- s = df.iloc[i]['53-2.0']
- d2 = datetime.date(int(s[:4]), int(s[5:7]), int(s[8:]))
- age[i] = np.round((d2 - d1).days / 365.25, 1)
- weight = df['21002-2.0'].values
- bmi = df['21001-2.0'].values
- height = np.round(np.sqrt(weight / bmi) * 100)
- # %%
- valid_idx = ~np.isnan(age) & ~np.isnan(sex) & ~np.isnan(weight) & ~np.isnan(height)
- df = df[valid_idx]
- df_idp = predictions[valid_idx]
- sex = sex[valid_idx]
- age = age[valid_idx]
- sex_age = sex * age # Sex and age interaction
- weight = weight[valid_idx]
- height = height[valid_idx]
- conf = np.stack((sex, age, sex_age, weight, height), axis=1)
- df_conf = pd.DataFrame(conf, index=df.index, columns=['Sex', 'Age', 'Sex * Age', 'Weight', 'Height'])
- # %%
- # Remove confounding factors from df
- df = df.drop(columns=[
- '34-0.0',
- '52-0.0',
- '21002-2.0',
- 'whole_body_pred',
- 'brain_pred',
- 'lungs_pred',
- 'heart_pred',
- 'spine_pred',
- 'liver_pred',
- 'muscle_pred',
- 'intestine_pred',
- 'age_gap_brain',
- 'age_gap_whole_body',
- 'age_gap_lungs',
- 'age_gap_heart',
- 'age_gap_spine',
- 'age_gap_liver',
- "age_gap_muscle",
- "age_gap_intestine",
- 'date_of_death',
- 'ins_index',
- 'dsource',
- 'source',
- "eid",
- ])
- # %%
- # # # # # # # # # # # # # # # # # # # #
- # Step 4: clean, de-confound and normalise data
- # # # # # # # # # # # # # # # # # # # #
- # This part of the code was adapted from Karla Miller's Matlab code at
- # https://www.fmrib.ox.ac.uk/ukbiobank/gwaspaper/index.html
- # Step 4.1: cleaning
- n_subj, n_col = df.shape
- bad_vars = []
- for i in range(n_col):
- val = df.iloc[:, i]
- # Discard columns which not numbers
- if not np.issubdtype(df.dtypes[i], np.number):
- bad_vars += [i]
- continue
- # Assume negative values are invalid, set them to NaN
- # There are also a lot of empty values, which are already NaN
- val[val < 0] = np.nan
- df.iloc[:, i] = val
- # Valid indices
- valid_idx = ~np.isnan(val)
- # Discard columns with more than 90% missing data
- if np.sum(valid_idx) < (0.1 * n_subj):
- bad_vars += [i]
- continue
- # Discard columns with over 95% elements with the exactly same value
- val_unique, counts = np.unique(val[valid_idx], return_counts=True)
- if np.max(counts) >= (0.95 * np.sum(valid_idx)):
- bad_vars += [i]
- continue
- #%%
- merged_data["2644-2.0"]
- # merged_data.drop(columns=[["2644-2.0","20161-2.0"]], inplace=True)
- merged_data.drop(columns=["20161-2.0"], inplace=True)
- # %%
- for i in range(n_col):
- for j in range(i + 1, n_col):
- if i in bad_vars or j in bad_vars:
- continue
- # Discard columns with very high correlation
- val_i = df.iloc[:, i]
- val_j = df.iloc[:, j]
- valid_idx = ~np.isnan(val_i) & ~np.isnan(val_j)
- if np.sum(valid_idx) == 0:
- continue
- # print(val_i[valid_idx], val_j[valid_idx])
- cc, _ = scipy.stats.pearsonr(val_i[valid_idx], val_j[valid_idx])
- if cc > 0.9999:
- # Keep the column with more valid elements
- if np.sum(~np.isnan(val_i)) > np.sum(~np.isnan(val_j)):
- bad_vars += [j]
- else:
- bad_vars += [i]
- # %%
- # The cleaned data
- bad_vars = np.unique(bad_vars)
- keep_vars = sorted(list(set(np.arange(n_col)) - set(bad_vars)))
- df = df.iloc[:, keep_vars]
- print('{0} columns kept after data cleaning.'.format(df.shape[1]))
- # Step 4.2: normalise confounding factors
- conf = (conf - np.mean(conf, axis=0)) / np.std(conf, axis=0)
- #%%
- # Step 4.3: normalise non imaging phenotypes
- df_cont = pd.read_csv(paths.path_to_continuous_features, index_col=0)
- #%%
- n_col = df.shape[1]
- for i in range(n_col):
- val = df.iloc[:, i]
- valid_idx = ~np.isnan(val)
- x = val[valid_idx]
- # Field ID
- field_id = int(df.columns[i][1].split('-')[0])
- is_continuous = df_cont.loc[field_id]['continuous']
- if is_continuous:
- # If it is a continuous variable, perform standard normalisation.
- x = normalise(x)
- else:
- # If we are not sure whether it is a continuous or categorical variable, convert it
- # into a continuous variable using rank-based inverse normal transform.``
- x = rank_normalise(x)
- df.iloc[:, i][valid_idx] = x# The cleaned data
- #%%
- # Step 4.4: de-confound and normalise IDPs
- try:
- df_idp = df_idp.drop(columns=["eid"])
- df_idp = df_idp.drop(columns=["split"])
- except:
- print("Columns already dropped")
- n_row = conf.shape[1]
- n_col = df_idp.shape[1]
- beta = np.zeros((n_row, n_col))
- for i in range(n_col):
- val = df_idp.iloc[:, i]
- print(val)
- valid_idx = ~np.isnan(val)
- x = val[valid_idx]
- beta[:, i] = np.dot(np.linalg.pinv(conf[valid_idx]), x)
- x = x - np.dot(conf[valid_idx], beta[:, i])
- x = normalise(x)
- df_idp.iloc[:, i][valid_idx] = x
- df_idp = df_idp[[
- "age_gap_brain",
- "age_gap_whole_body",
- "age_gap_lungs",
- "age_gap_heart",
- "age_gap_spine",
- "age_gap_liver",
- "age_gap_muscle",
- "age_gap_intestine"
- ]]
- # %%
- # # # # # # # # # # # # # # # # # # # # #
- # # Step 5: uni-variate correlation studies
- # # # # # # # # # # # # # # # # # # # # #
- M = df_idp.shape[1]
- N = df.shape[1]
- corr = np.zeros((M, N))
- corr_p = np.zeros((M, N))
- # features = []
- for i in range(M):
- features = []
- for j in range(N):
- # Remove NaNs
- x = df_idp.iloc[:, i]
- y = df.iloc[:, j]
- valid_idx = ~np.isnan(x) & ~np.isnan(y)
- x = x[valid_idx]
- y = y[valid_idx]
- # Pearson correlation
- cc, p_val = scipy.stats.pearsonr(x, y)
- corr[i, j] = cc
- corr_p[i, j] = p_val
- features.append(df.columns[j])
- # For p-value of 0, assign it with the mininal positive floating value
- # so that we can calculate the logarithm for the Manhattan plot
- corr_p[corr_p == 0] = np.finfo(np.float64).tiny
- # Logarithm
- log_corr_p = - np.log10(corr_p)
- # Save the tables
- df_corr = pd.DataFrame(corr, index=df_idp.columns, columns=df.columns)
- df_p = pd.DataFrame(corr_p, index=df_idp.columns, columns=df.columns)
- df_log_p = pd.DataFrame(log_corr_p, index=df_idp.columns, columns=df.columns)
- corr = df_corr.values
- corr_p = df_p.values
- log_corr_p = df_log_p.values
- # Bonferroni correction
- M, N = corr.shape
- p_bonf = 0.05 / (M * N)
- # FDR correction
- p_fdr, _ = fdr_threshold(corr_p.flatten(), 0.05)
- # Number of phenotypes that is significantly associated with at least one of the IDPs
- print('p_bonf = {0}'.format(p_bonf))
- print('p_fdr = {0}'.format(p_fdr))
- print('Number of correlations reaching Bonferroni threshold = {0}'.format(np.sum(corr_p < p_bonf)))
- print('Number of correlations reaching FDR threshold = {0}'.format(np.sum(corr_p < p_fdr)))
- print('Number of phenotypes reaching Bonferroni threshold = {0}'.format(np.sum(np.sum(corr_p < p_bonf, axis=0) > 0)))
- print('Number of phenotypes reaching FDR threshold = {0}'.format(np.sum(np.sum(corr_p < p_fdr, axis=0) > 0)))
- # %%
- # # # # # # # # # # # # # # # # # # # #
- # Step 6: Manhattan plot
- # # # # # # # # # # # # # # # # # # # #
- # The category for each column
- # They should be in ascending order. Otherwise, we need to sort the category IDs.
- pain_columns = pd.read_csv(paths.pain_path).columns
- sleep_columns = pd.read_csv(paths.sleep_path).columns
- sexual_columns = pd.read_csv(paths.sexual_factors_path).columns
- alcohol_columns = pd.read_csv(paths.alcohol_path).columns
- diet_columns = pd.read_csv(paths.diet_path).columns
- employment_columns = pd.read_csv(paths.employment_path).columns
- family_history_columns = pd.read_csv(paths.family_history_path).columns
- smoking_columns = pd.read_csv(paths.smoking_path).columns
- female_specific_columns = pd.read_csv(paths.female_sex_specific_path).columns
- general_health_columns = pd.read_csv(paths.general_health_path).columns
- male_specific_columns = pd.read_csv(paths.male_sex_specific_path).columns
- medical_conditions_columns = pd.read_csv(paths.medical_conditions_path).columns
- mental_health_columns = pd.read_csv(paths.mental_health_path).columns
- social_support_columns = pd.read_csv(paths.social_support_path).columns
- sun_exposure_columns = pd.read_csv(paths.sun_exposure_path).columns
- telomeres_columns = pd.read_csv(paths.telomeres_path).columns
- body_composition_columns = pd.read_csv(paths.body_composition_path).columns
- met = pd.read_csv(paths.path_to_MET_features).columns
- def get_category(field_id):
- if field_id in pain_columns:
- return "pain"
- elif field_id in sleep_columns:
- return "sleep"
- elif field_id in sexual_columns:
- return "sexual"
- elif field_id in alcohol_columns:
- return "alcohol"
- elif field_id in diet_columns:
- return "diet"
- elif field_id in employment_columns:
- return "employment"
- elif field_id in family_history_columns:
- return "family_history"
- elif field_id in smoking_columns:
- return "smoking"
- elif field_id in female_specific_columns:
- return "female_specific"
- elif field_id in general_health_columns:
- return "general_health"
- elif field_id in male_specific_columns:
- return "male_specific"
- elif field_id in medical_conditions_columns:
- return "medical_conditions"
- elif field_id in mental_health_columns:
- return "mental_health"
- elif field_id in social_support_columns or field_id=="2020-2.0" or field_id=="2020-2.0" or field_id=="6160-2.0_y" or field_id=="6160-0.0":
- return "social_support"
- elif field_id in sun_exposure_columns:
- return "sun_exposure"
- elif field_id in telomeres_columns:
- return "telomeres"
- elif field_id in body_composition_columns:
- return "body_composition"
- elif field_id in met:
- return "physical activity"
- else:
- return "general"
- # The Manhattan plot
- table = []
- for i in range(M):
- if df_idp.columns[i] == 'age_gap_whole_body':
- s = 'whole_body'
- elif df_idp.columns[i] == 'age_gap_brain':
- s = 'brain'
- elif df_idp.columns[i] == 'age_gap_lungs':
- s = 'lungs'
- elif df_idp.columns[i] == 'age_gap_heart':
- s = 'heart'
- elif df_idp.columns[i] == 'age_gap_spine':
- s = 'spine'
- elif df_idp.columns[i] == 'age_gap_liver':
- s = 'liver'
- elif df_idp.columns[i] == 'age_gap_muscle':
- s = 'muscle'
- elif df_idp.columns[i] == 'age_gap_intestine':
- s = 'intestine'
- for j in range(N):
- line = [j, log_corr_p[i, j], corr[i, j], abs(corr[i, j]), s, features[j]]
- table += [line]
- df_table = pd.DataFrame(table, columns=['x', 'log_p', 'corr', 'Correlation', 'Modality', "Feature"])
- #%%
- category = np.array([get_category(field_id) for field_id in tqdm(df_table["Feature"])])
- df_table["Category"] = category
- df_table = df_table.sort_values(by=["Category","Feature"])
- #%%
- df_table = df_table.sort_values(by=["Category","Feature", "Modality"])
- #%%
- df_table["x2"] = np.repeat(np.arange(1, N+1), M)
- #%%
- df_table = df_table.sample(frac=1)
- #%%
- df_table.reset_index(drop=True, inplace=True)
- #%%
- pallette = {
- 'heart': '#D81B60',
- 'muscle': "#1E88E5",
- 'whole_body': "#FFBF00",
- 'liver': '#004D40',
- 'brain': '#979E41',
- 'spine': '#51D9B8',
- 'lungs': '#AA58B3',
- 'intestine': '#C56800',
- }
- #%%
- plt.figure(figsize=(16, 8))
- ax = sns.scatterplot(x='x2', y='log_p', hue='Modality', size='Correlation',
- sizes=(30, 200), size_norm=mpl.colors.Normalize(vmin=0, vmax=0.3),
- data=df_table, alpha=0.8, palette=pallette)
- plt.plot([0, N], [-math.log10(p_bonf), -math.log10(p_bonf)], 'k--', linewidth=1, alpha=0.9)
- plt.text(-1, -math.log10(p_bonf), 'Bonf', horizontalalignment='right', fontsize=15)
- handles, labels = ax.get_legend_handles_labels()
- # labels[-1] = '0.3'
- ax.legend(handles, labels)
- xticks = []
- xticklabels = []
- category_list = df_table["Category"].values
- unique_category = np.unique(category_list)
- print(unique_category)
- for c in unique_category:
- x = np.max(df_table[df_table["Category"]==c]["x2"]) + 0.5
- plt.plot([x, x], [0, 310], 'k-.', linewidth=0.5, alpha=0.2)
- xticks += [np.mean(df_table[df_table["Category"]==c]["x2"])]
- xticklabels += [c]
- plt.xlim(0, N)
- plt.ylim(0, 110)
- plt.xticks(xticks, xticklabels, fontsize=15, rotation=90)
- plt.yticks(fontsize=15)
- plt.xlabel('Non-imaging phenotypes', fontsize=15, fontweight='bold')
- plt.ylabel(r'$-\log_{10}(p)$', fontsize=15, fontweight='bold')
- fig = plt.gcf()
- fig.set_size_inches(16, 8)
- plt.tight_layout()
- plt.savefig('../../outputs/figures/manhattan_phewas.pdf', bbox_inches='tight')
- plt.show()
- # %%
- plt.figure(figsize=(16, 8))
- ax = sns.scatterplot(x='x2', y='corr', hue='Modality', size='Correlation',
- sizes=(30, 100), size_norm=mpl.colors.Normalize(vmin=0, vmax=0.2),
- data=df_table, alpha=0.8, palette=palette)
- plt.hlines(0, xmin=0, xmax=N*2, linewidth=1, alpha=0.8, color="gray", linestyle='--')
- handles, labels = ax.get_legend_handles_labels()
- # labels[-1] = '0.25'
- ax.legend(handles, labels)
- xticks = []
- xticklabels = []
- category_list = df_table["Category"].values
- unique_category = np.unique(category_list)
- print(unique_category)
- for c in unique_category:
- x = np.max(df_table[df_table["Category"]==c]["x2"]) + 0.5
- plt.plot([x, x], [-0.15, 0.21], 'k-.', linewidth=0.5)
- xticks += [np.mean(df_table[df_table["Category"]==c]["x2"])]
- xticklabels += [c]
- plt.xlim(0, N)
- plt.xticks(xticks, xticklabels, fontsize=15, rotation=90)
- plt.yticks(fontsize=15)
- plt.xlabel('Non-imaging phenotypes', fontsize=15, fontweight='bold')
- plt.ylabel('Correlation', fontsize=15, fontweight='bold')
- fig = plt.gcf()
- fig.set_size_inches(16, 8)
- plt.tight_layout()
- plt.savefig('../../outputs/figures/correlations_phewas.pdf', bbox_inches='tight')
- plt.show()
perform_phewas.py at commit a772251, no license · at the source
Overview
14 affiliations
- Lab for AI in Medicine and Healthcare, Technical University of Munich,Munich, Germany
- Department of Diagnostic and Interventional Neuroradiology, School of Medicine, TUM University Hospital,Munich, Germany
- Department of Diagnostic and Interventional Radiology, Medical Center, University of Freiburg, Faculty of Medicine,Freiburg, Germany
- Institut für Community Medicine, University Medicine Greifswald,Greifswald, Germany
- Institute for Epidemiology and Preventive Medicine, University of Regensburg,Regensburg, Germany
- Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine, Helmholtz Association,Berlin, Germany
- Institute of Social Medicine, Epidemiology and Health Economics, Charité - Universitätsmedizin Berlin,Berlin, Germany
- Institute of Clinical Epidemiology and Biometry, University of Würzburg,Würzburg, Germany
- State Institute of Health I, Bavarian Health and Food Safety Authority,Erlangen, Germany
- Institute of Epidemiology and Social Medicine, University of Münster,Münster, Germany
- Department of Computing, Imperial College London,London, United Kingdom
- Institute for Diagnostic and Interventional Radiology, Klinikum rechts der Isar, Technical University of Munich,Munich, Germany
- German Cancer Consortium (DKTK), Munich partner site,Heidelberg, Germany
- Department of Diagnostic and Interventional Radiology and Nuclear Medicine, University Medical Center Hamburg-Eppendorf,Hamburg, Germany
Abstract
Background: Understanding human ageing across multiple organs is essential for characterising individual health trajectories and identifying abnormal ageing processes. Multi-organ imaging provides an opportunity to quantify biological ageing beyond chronological age. The aim of this study is to assess organ-specific and whole-body ageing patterns and their associations with disease and lifestyle factors.
Methods: In this large-scale study, we evaluate biological ageing patterns using 70,000 MRI scans from the UK Biobank and the German National Cohort. We employ 3D ResNet-18 models to predict chronological age from various body regions (brain, heart, liver, spine, lungs, muscle, and intestine) and the whole body. From these predictions, we derive “age gaps” relative to a strictly healthy reference cohort, which enables the identification of accelerated ageing patterns. We then evaluate associations with chronic diseases and lifestyle factors, and a virtual ageing framework was developed to explore counterfactual scenarios by substituting anatomical regions across subjects, quantifying local impacts on global biological age.
Results: Here we show significant associations between detected accelerated ageing and specific chronic diseases, including multiple sclerosis and chronic obstructive pulmonary disease, as well as lifestyle factors such as smoking and physical activity. Virtual substitution of anatomical regions demonstrates that local substitutions can influence global ageing patterns.
Conclusions: This study demonstrates that multi-organ imaging enables the detection of abnormal ageing patterns at both local and global levels. The presented framework provides a foundation for improved risk stratification and supports the development of personalised approaches to health assessment and disease prevention.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.
tamaramueller/age_biomarkers
a772251fc863ac522f143f918919b7659ec45cb3, 21 April 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
36 files
- source/
analysis/ , Python, 1 line__init__.py - source/
analysis/ , Python, 28 linesconfig_diseases.py - source/
analysis/ , Python, 81 lines, 1 matchdisease_analysis.py - source/
analysis/ , Python, 12 lineslinear_correction.py - source/
analysis/ , Python, 17 linespaths.py - source/
analysis/ , Python, 612 lines, 2 matchesperform_phewas.py - source/
analysis/ , Python, 48 linesphewas_interactive_plot. py - source/
analysis/ , Python, 108 linespredict_frankensteins.py - source/
analysis/ , Python, 191 linespropensity_weighted_scor e.py - source/
analysis/ , Python, 1 lineregistration.py - source/
analysis/ , Python, 147 lines, 1 matchsurvival.py - source/
analysis/ , Python, 142 linesutils.py - source/
cnn/ , Python, 1 line__init__.py - source/
cnn/ , Python, 1 linedataset/ __init__.py - source/
cnn/ , Python, 87 linesdataset/ base_dataset.py - source/
cnn/ , Python, 74 linesdataset/ brain_3d_dataset.py - source/
cnn/ , Python, 66 linesdataset/ dataset_utils.py - source/
cnn/ , Python, 38 linesdataset/ helpers.py - source/
cnn/ , Python, 88 linesdataset/ whole_body_3d_dataset.py - source/
cnn/ , Python, 1 linemodel/ __init__.py - source/
cnn/ , Python, 93 linesmodel/ base_model.py - source/
cnn/ , Python, 434 linesmodel/ efficientnet.py - source/
cnn/ , Python, 403 linesmodel/ ensemble.py - source/
cnn/ , Python, 4 linesmodel/ helpers.py - source/
cnn/ , Python, 73 linesmodel/ model_with_features.py - source/
cnn/ , Python, 236 lines, 1 matchmodel/ resnet.py - source/
cnn/ , Python, 48 linesmodel/ resnet3D.py - source/
cnn/ , Python, 616 linesmodel/ utils.py - source/
cnn/ , Python, 205 linesmodel/ vgg16.py - source/
cnn/ , Python, 140 linespredict.py - source/
cnn/ , Python, 50 linestest.py - source/
cnn/ , Python, 89 linestrain.py - source/
cnn/ , Python, 402 lines, 1 matchtrain_utils.py - source/
cnn/ , Python, 57 linestrainer.py - source/
registration/ , Python, 302 linesregistration.py - README.md, Text, 59 lines
Code availability
All files can be found in the corresponding source code repository at: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 35 scripts, each with its path and the digest of its content;
- 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
A were created with Biorender, detailed acknowledgement will be provided upon publication. Eligible researchers may access UK Biobank data on https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 21 authors, 1 keyword, 1 funder, 43 references.
Cite
This paper
Mueller, T. T., Starck, S., Llalloshi, R., Kaissis, G., Ziller, A., Graf, R., Schlett, C., Ringhof, S., Bamberg, F., Wielpütz, M., Völzke, H., Leitzmann, M., Niendorf, T., Keil, T., Krist, L., Pischon, T., Karch, A., Berger, K., Kirschke, J., . . . Braren, R. (2026). Local and global patterns support medical imaging as a biomarker of ageing. Communications medicine, 6(1), 335. https://
BibTeX
@article{mueller2026loca
author = {Mueller, Tamara T. and Starck, Sophie and Llalloshi, Rozafë and Kaissis, Georgios and Ziller, Alexander and Graf, Robert and Schlett, Christopher and Ringhof, Steffen and Bamberg, Fabian and Wielpütz, Mark and Völzke, Henry and Leitzmann, Michael and Niendorf, Thoralf and Keil, Thomas and Krist, Lilian and Pischon, Tobias and Karch, André and Berger, Klaus and Kirschke, Jan and Rueckert, Daniel and Braren, Rickmer},
title = {{Local and global patterns support medical imaging as a biomarker of ageing}},
journal = {Communications medicine},
year = {2026},
month = jun,
volume = {6},
number = {1},
pages = {335},
publisher = {Nature Publishing Group},
issn = {2730-664X},
doi = {10.1038/
url = {https://
pmid = {42288701},
pmcid = {PMC13264634}
}
RIS
TY - JOUR
AU - Mueller, Tamara T.
AU - Starck, Sophie
AU - Llalloshi, Rozafë
AU - Kaissis, Georgios
AU - Ziller, Alexander
AU - Graf, Robert
AU - Schlett, Christopher
AU - Ringhof, Steffen
AU - Bamberg, Fabian
AU - Wielpütz, Mark
AU - Völzke, Henry
AU - Leitzmann, Michael
AU - Niendorf, Thoralf
AU - Keil, Thomas
AU - Krist, Lilian
AU - Pischon, Tobias
AU - Karch, André
AU - Berger, Klaus
AU - Kirschke, Jan
AU - Rueckert, Daniel
AU - Braren, Rickmer
TI - Local and global patterns support medical imaging as a biomarker of ageing
T2 - Communications medicine
J2 - Commun Med (Lond)
PY - 2026
DA - 2026/
VL - 6
IS - 1
SP - 335
SN - 2730-664X
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Local and global patterns support medical imaging as a biomarker of ageing",
"container-title": "Communications medicine",
"author": [
{
"family": "Mueller",
"given": "Tamara T."
},
{
"family": "Starck",
"given": "Sophie"
},
{
"family": "Llalloshi",
"given": "Rozafë"
},
{
"family": "Kaissis",
"given": "Georgios"
},
{
"family": "Ziller",
"given": "Alexander"
},
{
"family": "Graf",
"given": "Robert"
},
{
"family": "Schlett",
"given": "Christopher"
},
{
"family": "Ringhof",
"given": "Steffen"
},
{
"family": "Bamberg",
"given": "Fabian"
},
{
"family": "Wielpütz",
"given": "Mark"
},
{
"family": "Völzke",
"given": "Henry"
},
{
"family": "Leitzmann",
"given": "Michael"
},
{
"family": "Niendorf",
"given": "Thoralf"
},
{
"family": "Keil",
"given": "Thomas"
},
{
"family": "Krist",
"given": "Lilian"
},
{
"family": "Pischon",
"given": "Tobias"
},
{
"family": "Karch",
"given": "André"
},
{
"family": "Berger",
"given": "Klaus"
},
{
"family": "Kirschke",
"given": "Jan"
},
{
"family": "Rueckert",
"given": "Daniel"
},
{
"family": "Braren",
"given": "Rickmer"
}
],
"container-title-short":
"volume": "6",
"issue": "1",
"page": "335",
"DOI": "10.1038/
"PMID": "42288701",
"PMCID": "PMC13264634",
"ISSN": "2730-664X",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
13
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1002/hbm.70425 [code]
- Brain Age Estimation on T2-FLAIR Scans for Application to Multiple Sclerosis.Journal: Human brain mappingIn common: PyTorch, pandas, SciPy, 1 other tool, clinical / translational, 8 references
- [2] doi:10.1186/s40708-026-00316-y [code]
- Generalizable and explainable deep learning for brain MRI: a multi-cohort evaluation of 3D architectures for age and sex prediction.Journal: Brain informaticsIn common: PyTorch, seaborn, pandas, 3 other tools, 6 references
- [3] doi:10.1162/imag.a.1164 [code]
- Bias and generalizability of brain age prediction models: A multi-cohort evaluation with anatomical and interpretability insights.Journal: Imaging neuroscience (Cambridge, Mass.)In common: NiBabel, statsmodels, PyTorch, 5 other tools, 4 references
- [4] doi:10.1038/s41380-026-03691-4 [code]
- Breaking the norm: population-scale deviations of brain structure in depression and anxiety.Journal: Molecular psychiatryIn common: NiBabel, statsmodels, PyTorch, 5 other tools, clinical / translational, author Tobias Pischon
- [5] doi:10.1038/s41398-026-04131-1 [code]
- Multimodal phenotypic classification of generalized anxiety and panic using structural MRI data and psychosocial factors: machine learning results from the German National Cohort (NAKO) study.Journal: Translational psychiatryIn common: clinical / translational, 1 reference, 2 authors
- [6] doi:10.1007/s12021-026-09817-x [code]
- Circle of Willis-Guided Localization for Simultaneous Detection and Classification of Large Vessel Occlusions in Brain CTA.Journal: NeuroinformaticsIn common: PyTorch Lightning, SimpleITK, scikit-image, 7 other tools
- [7] doi:10.1093/radadv/umag025 [code]
- OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI.Journal: Radiology advancesIn common: SimpleITK, NiBabel, PyTorch, 4 other tools, 3 references
- [8] doi:10.3389/frai.2026.1771088 [code]
- Few-shot deployment of pretrained MRI transformers in brain imaging tasks.Journal: Frontiers in artificial intelligenceIn common: SimpleITK, scikit-image, NiBabel, 6 other tools, 1 reference
- [9] doi:10.1038/s41398-026-04081-8 [code]
- Functional system-specific brain aging across the Alzheimer's disease continuum.Journal: Translational psychiatryIn common: NiBabel, statsmodels, seaborn, 4 other tools, clinical / translational, 3 references
- [10] doi:10.1038/s42003-026-10205-z [code]
- Source-space EEG alpha activity reveals brain age gaps due to neurodegeneration and disparity.Journal: Communications biologyIn common: statsmodels, seaborn, pandas, 3 other tools, 4 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 35 scripts, and 6 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:494686a857704492…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
