OSCR

Breaking the norm: population-scale deviations of brain structure in depression and anxiety.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Materials and methods › Group definitions ↔ ukb_validation/directional_analysis.py, lines 342–419 · score 0.83 · 15–19, 15–21, 20–27, 10–14, 0–4, 5–9
  2. [2] § Results › Shift analysis ↔ normative_modeling/2_1_significance_testing.py, lines 1–37 · score 0.74 · Benjamini Hochberg correction, age squared, HC reference, healthy control, Mahalanobis distances, variable
  3. [3] § Materials and methods › Group definitions ↔ ukb_validation/genetic_and_shift.py, lines 423–483 · score 0.71 · 15–19, 20–27, 10–14, 5–9, scores, PHQ
  4. [4] § Materials and methods › Significance testing of mahalanobis distances ↔ normative_modeling/2_1_significance_testing.py, lines 1–37 · score 0.71 · Benjamini Hochberg, discovery rate, ANOVA, HCs, covariate, mahalanobis
  5. [5] § Materials and methods › Classification analysis ↔ ukb_validation/classification.py, lines 50–115 · score 0.62 · logistic regression, classifications, fold, ratio, elastic, stratified
  6. [6] § Materials and methods › Classification analysis ↔ ukb_validation/classification.py, lines 50–115 · score 0.61 · balanced accuracy, ROC, fold, AUC, BACC, stratified
  7. [7] § Materials and methods › Normative modeling using autoencoder ↔ normative_modeling/1_train_autoencoder.py, lines 234–316 · score 0.60 · fully connected autoencoder, Architectural, reconstruction, space, normative, trained
  8. [8] § Materials and methods › Normative modeling using autoencoder ↔ ukb_validation/directional_analysis.py, lines 66–148 · score 0.59 · fully connected autoencoder, Architectural, reconstruction, validation, space, normative
  9. [9] § Materials and methods › Shift analysis of mahalanobis distances ↔ normative_modeling/2_3_mahalanobis_shift_analysis.py, lines 21–105 · score 0.57 · covariance matrix, Mahalanobis distance, subset, HCs, Shift
  10. [10] § Materials and methods › Normative modeling using autoencoder ↔ normative_modeling/1_train_autoencoder.py, lines 351–384 · score 0.56 · trained autoencoder, age squared, linear, sex, errors, modeling
  11. [11] § Materials and methods › Shift analysis of mahalanobis distances ↔ normative_modeling/2_4_shift_analysis.py, lines 50–115 · score 0.56 · bootstrap resampling, confidence intervals, shift
  12. [12] § Results › Directional analysis ↔ ukb_validation/directional_analysis.py, lines 704–756 · score 0.53 · moderately severe, deviation vectors, mild, centroid, ellipses, magnitudes
  13. [13] § Materials and methods › Shift analysis of mahalanobis distances ↔ normative_modeling/2_4_shift_analysis.py, lines 140–192 · score 0.53 · weighted area, emphasizes, metric, quantified, shift, distance
  14. [14] § Materials and methods › Cohorts - UK biobank (UKB) ↔ ukb_validation/genetic_and_shift.py, lines 299–356 · score 0.53 · neurodegenerative diseases, volume, UKB, sex, biobank, age
  15. [15] § Materials and methods › Normative modeling using autoencoder ↔ normative_modeling/1_train_autoencoder.py, lines 443–501 · score 0.52 · min max, TIV, autoencoder, linearly, age, sex

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 850 lines · 30 KB · GPL-3.0 · 3 matches

  1. import torch.nn as nn
  2. import torch
  3. import joblib
  4. import matplotlib.pyplot as plt
  5. import numpy as np
  6. import pandas as pd
  7. import pickle
  8. from sklearn.preprocessing import StandardScaler
  9. from sklearn.preprocessing import OneHotEncoder, MinMaxScaler
  10. from sklearn.linear_model import LinearRegression
  11. from matplotlib.patches import Ellipse, Patch
  12. from matplotlib.lines import Line2D
  13. from shapely.geometry import MultiPoint
  14. class Deconfounder:
  15. def __init__(self):
  16. self.models = []
  17. def train(self, embeddings, confounders):
  18. """
  19. Train the deconfounding models on the provided embeddings and confounders.
  20. Args:
  21. embeddings (numpy.ndarray): The embeddings to be deconfounded (N x D).
  22. confounders (numpy.ndarray): The confounders (N x C).
  23. Returns:
  24. numpy.ndarray: Deconfounded embeddings for the training data.
  25. """
  26. self.models = []
  27. residuals = np.zeros_like(embeddings)
  28. for i in range(embeddings.shape[1]):
  29. model = LinearRegression()
  30. model.fit(confounders, embeddings[:, i])
  31. self.models.append(model)
  32. predicted = model.predict(confounders)
  33. residuals[:, i] = embeddings[:, i] - predicted
  34. return residuals
  35. def apply(self, embeddings, confounders):
  36. """
  37. Apply the trained deconfounding models on new embeddings and confounders.
  38. Args:
  39. embeddings (numpy.ndarray): The new embeddings to be deconfounded (N x D).
  40. confounders (numpy.ndarray): The confounders for the new embeddings (N x C).
  41. Returns:
  42. numpy.ndarray: Deconfounded embeddings.
  43. """
  44. if not self.models:
  45. raise ValueError("Deconfounding models have not been trained. Please train them first.")
  46. residuals = np.zeros_like(embeddings)
  47. for i, model in enumerate(self.models):
  48. predicted = model.predict(confounders)
  49. residuals[:, i] = embeddings[:, i] - predicted
  50. return residuals
  51. class Autoencoder(nn.Module):
  52. """
  53. Normative autoencoder model for structural MRI embeddings.
  54. This fully connected autoencoder compresses input features into a low-dimensional latent space
  55. and reconstructs them back, aiming to model normative (healthy) data distributions.
  56. Architecture:
  57. - Encoder: 4 linear layers with ReLU activations reducing to a low-dimensional latent space.
  58. - Decoder: 4 linear layers with ReLU activations reconstructing the input, ending with a Sigmoid activation.
  59. Methods
  60. -------
  61. encode(x):
  62. Encodes input into a lower-dimensional latent representation.
  63. decode(x):
  64. Decodes latent representations back to input space.
  65. forward(x, train=True):
  66. Full autoencoding pass; if train=False, returns the reconstruction only.
  67. """
  68. def __init__(self, input_dim=512, hidden_dim1=150, hidden_dim2=100, hidden_dim3=75, hidden_dim4=50):
  69. super(Autoencoder, self).__init__()
  70. # Encoder
  71. self.enc_linear1 = nn.Linear(input_dim, hidden_dim1)
  72. self.enc_relu1 = nn.ReLU()
  73. self.enc_linear2 = nn.Linear(hidden_dim1, hidden_dim2)
  74. self.enc_relu2 = nn.ReLU()
  75. self.enc_linear3 = nn.Linear(hidden_dim2, hidden_dim3)
  76. self.enc_relu3 = nn.ReLU()
  77. self.enc_linear4 = nn.Linear(hidden_dim3, hidden_dim4)
  78. self.enc_relu4 = nn.ReLU()
  79. # Decoder
  80. self.dec_linear1 = nn.Linear(hidden_dim4, hidden_dim3)
  81. self.dec_relu1 = nn.ReLU()
  82. self.dec_linear2 = nn.Linear(hidden_dim3, hidden_dim2)
  83. self.dec_relu2 = nn.ReLU()
  84. self.dec_linear3 = nn.Linear(hidden_dim2, hidden_dim1)
  85. self.dec_relu3 = nn.ReLU()
  86. self.dec_linear4 = nn.Linear(hidden_dim1, input_dim)
  87. self.dec_relu4 = nn.ReLU()
  88. self.sig = nn.Sigmoid()
  89. def encode(self, x):
  90. x = self.enc_linear1(x)
  91. x = self.enc_relu1(x)
  92. x = self.enc_linear2(x)
  93. x = self.enc_relu2(x)
  94. x = self.enc_linear3(x)
  95. x = self.enc_relu3(x)
  96. x = self.enc_linear4(x)
  97. x = self.enc_relu4(x)
  98. return x
  99. def decode(self, x):
  100. x = self.dec_linear1(x)
  101. x = self.dec_relu1(x)
  102. x = self.dec_linear2(x)
  103. x = self.dec_relu2(x)
  104. x = self.dec_linear3(x)
  105. x = self.dec_relu3(x)
  106. x = self.dec_linear4(x)
  107. x = self.sig(x)
  108. return x
  109. def forward(self, x, train=True):
  110. z = self.encode(x)
  111. decoded = self.decode(z)
  112. if not train:
  113. return decoded
  114. return decoded
  115. def assign_color_from_label(label):
  116. label_lower = label.lower()
  117. if "and_depression" in label_lower:
  118. return group_colors_map["GAD+PHQ-9"]
  119. elif "depression_ad" in label_lower:
  120. return group_colors_map["GAD/PHQ-9+AUDIT-C"]
  121. elif "audit" in label_lower or "aud_" in label_lower:
  122. return group_colors_map["AUDIT-C"]
  123. elif "ad_" in label_lower or "gad_" in label_lower or "ad_ic" in label_lower or "self_rep_anx" in label_lower:
  124. return group_colors_map["GAD"]
  125. elif "depression" in label_lower or "dep_ic" in label_lower or "self_rep_mdd" in label_lower:
  126. return group_colors_map["PHQ-9"]
  127. return "gray"
  128. def sample_ellipse_points(center, width, height, angle, n_points=200):
  129. t = np.linspace(0, 2 * np.pi, n_points)
  130. ellipse = np.column_stack([np.cos(t) * width / 2, np.sin(t) * height / 2])
  131. theta = np.radians(angle)
  132. R = np.array([[np.cos(theta), -np.sin(theta)], [np.sin(theta), np.cos(theta)]])
  133. ellipse_rot = ellipse @ R.T + center
  134. return ellipse_rot
  135. def encapsulate_meta_group(meta_color_list):
  136. all_points = []
  137. for group in strict_diag_cols:
  138. mask = df[group] == 1
  139. if mask.sum() == 0: continue
  140. group_color = assign_color_from_label(group)
  141. if group_color in meta_color_list:
  142. center, width, height, angle = bootstrap_ci_ellipse(residuals_2d[mask])
  143. points = sample_ellipse_points(center, width, height, angle)
  144. all_points.append(points)
  145. if all_points:
  146. all_points = np.vstack(all_points)
  147. hull = MultiPoint(all_points).convex_hull
  148. return hull
  149. return None
  150. def geom_median(X, eps=1e-6, max_iter=1000):
  151. """Weiszfeld algorithm for the spatial (geometric) median in 2D."""
  152. X = np.asarray(X, dtype=float)
  153. m = X.mean(axis=0)
  154. for _ in range(max_iter):
  155. d = np.linalg.norm(X - m, axis=1)
  156. if np.any(d < eps):
  157. return X[d.argmin()]
  158. w = 1.0 / d
  159. m_new = (X * w[:, None]).sum(axis=0) / w.sum()
  160. if np.linalg.norm(m_new - m) < eps:
  161. return m_new
  162. m = m_new
  163. return m
  164. def bootstrap_ci_ellipse(points_2d, n_bootstrap=1000):
  165. """
  166. Returns a 1-SD ellipse of the bootstrap-median distribution (NOT a CI).
  167. Interface unchanged: (center, width, height, angle).
  168. """
  169. X = np.asarray(points_2d, dtype=float)
  170. n = X.shape[0]
  171. meds = np.zeros((n_bootstrap, 2), dtype=float)
  172. for i in range(n_bootstrap):
  173. sample = X[np.random.randint(0, n, n)]
  174. meds[i] = geom_median(sample)
  175. center = geom_median(meds)
  176. cov = np.cov(meds, rowvar=False) + 1e-9 * np.eye(2)
  177. vals, vecs = np.linalg.eigh(cov)
  178. order = vals.argsort()[::-1]
  179. vals, vecs = vals[order], vecs[:, order]
  180. angle = np.degrees(np.arctan2(vecs[1, 0], vecs[0, 0]))
  181. # 1-SD ellipse: diameters = 2 * sqrt(eigenvalues)
  182. width, height = 2.0 * np.sqrt(vals[0]), 2.0 * np.sqrt(vals[1])
  183. return center, width, height, angle
  184. def bootstrap_ci(data, n_bootstrap=1000, ci=95):
  185. """
  186. Bootstraps the confidence interval for the median vector.
  187. """
  188. medians = []
  189. n_samples = data.shape[0]
  190. for _ in range(n_bootstrap):
  191. sample = data[np.random.choice(n_samples, n_samples, replace=True)]
  192. medians.append(np.median(sample, axis=0))
  193. medians = np.array(medians)
  194. lower = np.percentile(medians, (100 - ci) / 2, axis=0)
  195. upper = np.percentile(medians, 100 - (100 - ci) / 2, axis=0)
  196. return lower, upper
  197. def replace_terms(text):
  198. replacements = {
  199. "DEP_ICD10": "ICD-10",
  200. "ALZ_ICD10": "ND ICD-10",
  201. "depression_1_no_comor": "(1)",
  202. "depression_2_no_comor": "(2)",
  203. "depression_3_no_comor": "(3)",
  204. "depression_4_no_comor": "(4)",
  205. "AUD_ICD10": "AUD ICD-10",
  206. "diagnosis_audit_8_no_comor": "(8)",
  207. "diagnosis_audit_9_no_comor": "(9)",
  208. "diagnosis_audit_10_no_comor": "(>= 10)",
  209. "GAD_ICD10": "GAD ICD-10",
  210. "AD_ICD10": "ANX ICD-10",
  211. "ad_2_no_comor": "(2)",
  212. "ad_3_no_comor": "(3)",
  213. "ad_4_no_comor": "(4)",
  214. "ad_2_and_depression_2_comor": "(2)",
  215. "ad_3_and_depression_3_comor": "(3)",
  216. "ad_4_and_depression_4_comor": "(4)",
  217. "depression_ad_aud_2": "(2)",
  218. "depression_ad_aud_3": "(3)", "depression_ad_aud_4": "(4)",
  219. "depression_ad_audit_2": "MDD/GAD (2) and AUDIT-C >= 8", "depression_ad_audit_3": "and AUDIT-C >= 9",
  220. "depression_ad_audit_4": "and AUDIT-C >= 10"
  221. }
  222. for key, value in replacements.items():
  223. text = text.replace(key, value)
  224. return text
  225. df_prs_meta = pd.read_csv("/opt/notebooks/studies_info.tsv", sep="\t")
  226. hc_embeddings_anxiety_path = "/opt/notebooks/all_embdes_raw_40k.csv"
  227. df = pd.read_csv(hc_embeddings_anxiety_path)
  228. df["id_"] = df["id_"].astype(str)
  229. df['embedding'] = df['embedding'].apply(
  230. lambda x: [np.fromstring(x.strip('[]'), sep=' ')])
  231. freesurfer_path = "/opt/notebooks/aseg.volume.tsv"
  232. free = pd.read_csv(freesurfer_path, sep="\t")
  233. free.rename(columns={"Measure:volume": "Participant ID"}, inplace=True)
  234. free["Participant ID"] = free["Participant ID"].str.replace("sub-", "").astype(str)
  235. df["Participant ID"] = df["Participant ID"].astype(str)
  236. merged_df = df.merge(free, on=["Participant ID"], how="inner")
  237. columns_to_check = [
  238. "Sex",
  239. "Age when attended assessment centre | Instance 2",
  240. "EstimatedTotalIntraCranialVol",
  241. "UK Biobank assessment centre | Instance 2",
  242. "embedding"
  243. ]
  244. merged_df = merged_df.dropna(subset=columns_to_check)
  245. encoder = OneHotEncoder(drop='first', sparse_output=False)
  246. sex_labels = merged_df['Sex'].map({1: 0, 2: 1}).values.reshape(-1, 1)
  247. mean_age = pd.to_numeric(merged_df["Age when attended assessment centre | Instance 2"], errors="coerce").mean()
  248. ages = merged_df["Age when attended assessment centre | Instance 2"].values.reshape(-1, 1).astype(float)
  249. tiv = merged_df["EstimatedTotalIntraCranialVol"].values.reshape(-1, 1).astype(float)
  250. basis_uort = merged_df["UK Biobank assessment centre | Instance 2"].values.reshape(-1, 1)
  251. basis_uort_onehot = encoder.fit_transform(basis_uort.reshape(-1, 1))
  252. confounders = np.concatenate([sex_labels, ages, (ages - mean_age) ** 2, tiv], axis=1)
  253. all_embeddings = np.vstack(merged_df['embedding'].values)
  254. deconfounder = joblib.load("/opt/notebooks/first_deconfounder.joblib")
  255. embeddings = deconfounder.apply(all_embeddings, confounders)
  256. deconfounder = Deconfounder()
  257. deconfounder.train(embeddings, basis_uort_onehot)
  258. embeddings = deconfounder.apply(embeddings, basis_uort_onehot)
  259. merged_df["embeddings_deconfounded"] = list(embeddings)
  260. gad7 = pd.read_csv("/opt/notebooks/gad7_filtered.csv")
  261. only_one_missing = gad7.isnull().sum(axis=1) == 1
  262. more_than_one_missing = gad7.isnull().sum(axis=1) > 1
  263. gad7 = gad7[~more_than_one_missing].copy()
  264. gad7.loc[only_one_missing, :] = gad7.loc[only_one_missing, :].fillna(gad7.mode().iloc[0])
  265. gad7.rename(columns={'eid': 'Participant ID'}, inplace=True)
  266. gad7_mapping = {
  267. "Not at all": 0,
  268. "Several days": 1,
  269. "More than half the days": 2,
  270. "Nearly every day": 3,
  271. "Prefer not to answer": -3
  272. }
  273. for col in gad7.columns:
  274. if col != "Participant ID":
  275. gad7[col] = gad7[col].map(gad7_mapping)
  276. gad7 = gad7[~gad7.isin([-3]).any(axis=1)].dropna()
  277. gad7["gad7_sum"] = gad7.drop(columns=["Participant ID"]).sum(axis=1)
  278. gad7["ad_1"] = (gad7["gad7_sum"].between(0, 4)).astype(int)
  279. gad7["ad_2"] = (gad7["gad7_sum"].between(5, 9)).astype(int)
  280. gad7["ad_3"] = (gad7["gad7_sum"].between(10, 14)).astype(int)
  281. gad7["ad_4"] = (gad7["gad7_sum"].between(15, 21)).astype(int)
  282. merged_df["Participant ID"] = merged_df["Participant ID"].astype(str)
  283. gad7["Participant ID"] = gad7["Participant ID"].astype(str)
  284. merged_df = merged_df.merge(gad7, on='Participant ID', how="inner")
  285. phq = pd.read_csv("/opt/notebooks/PHQ9_filtered.csv")
  286. only_one_missing = phq.isnull().sum(axis=1) == 1
  287. more_than_one_missing = phq.isnull().sum(axis=1) > 1
  288. phq = phq[~more_than_one_missing].copy()
  289. phq.loc[only_one_missing, :] = phq.loc[only_one_missing, :].fillna(phq.mode().iloc[0])
  290. phq.rename(columns={'eid': 'Participant ID'}, inplace=True)
  291. phq_mapping = {
  292. "Not at all": 0,
  293. "Several days": 1,
  294. "More than half the days": 2,
  295. "Nearly every day": 3,
  296. "Prefer not to answer": -1
  297. }
  298. for col in phq.columns:
  299. if col != "Participant ID":
  300. phq[col] = phq[col].map(phq_mapping)
  301. phq = phq[~phq.isin([-1]).any(axis=1)].dropna()
  302. phq["phq9_sum"] = phq.drop(columns=["Participant ID"]).sum(axis=1)
  303. phq["depression_1"] = (phq["phq9_sum"].between(5, 9)).astype(int)
  304. phq["depression_2"] = (phq["phq9_sum"].between(10, 14)).astype(int)
  305. phq["depression_3"] = (phq["phq9_sum"].between(15, 19)).astype(int)
  306. phq["depression_4"] = (phq["phq9_sum"].between(20, 27)).astype(int)
  307. merged_df["Participant ID"] = merged_df["Participant ID"].astype(str)
  308. phq["Participant ID"] = phq["Participant ID"].astype(str)
  309. merged_df = merged_df.merge(phq, on='Participant ID', how="inner")
  310. alc = pd.read_csv("/opt/notebooks/alcohol_amount.csv")
  311. six = pd.read_csv("/opt/notebooks/alc_six_units.csv")
  312. six.rename(columns={'p29093': 'six_alc_frq'}, inplace=True)
  313. alc = pd.merge(six, alc, on="eid", how="inner")
  314. alc = alc.fillna(alc.mode().iloc[0])
  315. alcohol_per_day = {
  316. "1 or 2": 1.5,
  317. "3 or 4": 3.5,
  318. "5 or 6": 5.5,
  319. "7, 8 or 9": 8,
  320. "10 or more": 11
  321. }
  322. audit_freq_score = {
  323. "Never": 0,
  324. "Monthly or less": 1,
  325. "2 to 4 times a month": 2,
  326. "2 to 3 times a week": 3,
  327. "4 or more times a week": 4
  328. }
  329. audit_amount_score = {
  330. "1 or 2": 0,
  331. "3 or 4": 1,
  332. "5 or 6": 2,
  333. "7, 8 or 9": 3,
  334. "10 or more": 4
  335. }
  336. audit_six_score = {
  337. "Never": 0,
  338. "Less than monthly": 1,
  339. "Monthly": 2,
  340. "Weekly": 3,
  341. "Daily or almost daily": 4
  342. }
  343. # Compute AUDIT-C score
  344. alc["audit_freq"] = alc["drinking_frequency"].map(audit_freq_score).fillna(0)
  345. alc["audit_amount"] = alc["amount_per_day"].map(audit_amount_score).fillna(0)
  346. alc["audit_six"] = alc["six_alc_frq"].map(audit_six_score).fillna(0)
  347. alc["audit_c_score"] = alc["audit_freq"] + alc["audit_amount"] + alc["audit_six"]
  348. alc["diagnosis_audit_8"] = (alc["audit_c_score"] == 8).astype(int)
  349. alc["diagnosis_audit_9"] = (alc["audit_c_score"] == 9).astype(int)
  350. alc["diagnosis_audit_10"] = (alc["audit_c_score"] >= 10).astype(int)
  351. merged_df["Participant ID"] = merged_df["Participant ID"].astype(str)
  352. alc["Participant ID"] = alc["Participant ID"].astype(str)
  353. merged_df = merged_df.merge(alc, on='Participant ID', how="inner")
  354. icd = pd.read_csv("/opt/notebooks/ICD_10_Diagnosis.csv")
  355. icd.rename(columns={'eid': 'Participant ID'}, inplace=True)
  356. icd["GAD_ICD10"] = 0
  357. icd["AUD_ICD10"] = 0
  358. icd["ALZ_ICD10"] = 0
  359. icd["DEP_ICD10"] = 0
  360. icd["AD_ICD10"] = 0
  361. icd["has_anything"] = 0
  362. mask = icd.drop(columns=["Participant ID"]).astype(str).applymap(
  363. lambda x:
  364. ("anxiety" in x.lower())
  365. )
  366. icd.loc[mask.any(axis=1), "AD_ICD10"] = 1
  367. mask = icd.drop(columns=["Participant ID"]).astype(str).applymap(
  368. lambda x: ("generalized anxiety" in x.lower()) or
  369. ("generalised anxiety" in x.lower())
  370. )
  371. icd.loc[mask.any(axis=1), "GAD_ICD10"] = 1
  372. mask = icd.drop(columns=["Participant ID"]).astype(str).applymap(
  373. lambda x: ("alcohol use disorder" in x.lower()) or
  374. ("f10" in x.lower())
  375. )
  376. icd.loc[mask.any(axis=1), "AUD_ICD10"] = 1
  377. dep_mask = icd.drop(columns=["Participant ID"]).astype(str).applymap(
  378. lambda x: ("mdd" in x.lower()) or
  379. ("depression" in x.lower()) or
  380. ("f32" in x.lower()) or
  381. ("f33" in x.lower())
  382. )
  383. icd.loc[dep_mask.any(axis=1), "DEP_ICD10"] = 1
  384. merged_df["Participant ID"] = merged_df["Participant ID"].astype(str)
  385. icd["Participant ID"] = icd["Participant ID"].astype(str)
  386. merged_df = merged_df.merge(icd, on='Participant ID', how="inner")
  387. diag = pd.read_csv("/opt/notebooks/merged_multitarget_df_with_demographics.csv")
  388. filtered_columns = [col for col in diag.columns]
  389. diag.rename(columns={'p29058': 'frequency_a_attack'}, inplace=True)
  390. diag["Participant ID"] = diag["ID"]
  391. neurodegenerative_diseases = pd.read_csv("/opt/notebooks/neurodegenerative_disease_all.csv")
  392. diag = diag.merge(neurodegenerative_diseases, on="Participant ID", how="inner")
  393. diag = diag[diag["neurodegenerative_disease"] != 1].reset_index(drop=True)
  394. merged_df["Participant ID"] = merged_df["Participant ID"].astype(str)
  395. diag["Participant ID"] = diag["Participant ID"].astype(str)
  396. df = merged_df.merge(diag, on="Participant ID", how="inner")
  397. columns_to_check = [
  398. 'GAD_ICD10', 'DEP_ICD10', "ALZ_ICD10", 'depression_1', "ad_2", "ad_3", "ad_4", "diagnosis_audit_8",
  399. "diagnosis_audit_9", "diagnosis_audit_10",
  400. 'depression_2', 'depression_3', 'depression_4', 'AUD_ICD10',
  401. "AD_ICD10", "AUD_ICD10",
  402. ]
  403. columns_to_check_comors = ['depression_1', "ad_2", "ad_3", "ad_4", "diagnosis_audit_8", "diagnosis_audit_9",
  404. "diagnosis_audit_10",
  405. 'depression_2', 'depression_3', 'depression_4', "ALZ_ICD10"]
  406. coloumns_ro_check_comors_depr = ["ad_2", "ad_3", "ad_4", 'GAD_ICD10']
  407. df["is_healthy"] = (
  408. (df[columns_to_check] != 1).all(axis=1)
  409. )
  410. for col in columns_to_check_comors:
  411. df[f"{col}_no_comor"] = df[col].where(
  412. (df[col] == 1) & (df[columns_to_check_comors].sum(axis=1) == 1), 0
  413. )
  414. for threshold in [2, 3, 4]:
  415. ad_col = f"ad_{threshold}"
  416. dep_col = f"depression_{threshold}"
  417. comor_col = f"ad_{threshold}_and_depression_{threshold}_comor"
  418. if ad_col in df.columns and dep_col in df.columns:
  419. df[comor_col] = ((df[ad_col] == 1) & (df[dep_col] == 1)).astype(int)
  420. else:
  421. print(f"Missing column(s): {ad_col} or {dep_col}")
  422. df.to_csv("Diagnosis_tmp.csv")
  423. healthy_embeddings = np.stack(df.loc[df["is_healthy"], 'embeddings_deconfounded'].values)
  424. scaler = joblib.load("/opt/notebooks/scaler.pkl")
  425. df["embeddings_deconfounded"] = list(scaler.transform(np.stack(df['embeddings_deconfounded'].values)))
  426. deconfounded_embeddings = np.stack(df['embeddings_deconfounded'].values)
  427. mdd_cols = ['depression_1', 'depression_2', 'depression_3', 'depression_4', 'DEP_ICD10']
  428. anx_cols = ['ad_2', 'ad_3', 'ad_4', 'GAD_ICD10', 'AD_ICD10']
  429. aud_cols = ['diagnosis_audit_8', 'diagnosis_audit_9', 'diagnosis_audit_10', 'AUD_ICD10']
  430. mdd_df = pd.read_csv('bd_mdd_self_reported')
  431. df["Participant ID"] = df["Participant ID"].astype(str)
  432. mdd_df["eid"] = mdd_df["eid"].astype(str)
  433. merged_df = df.merge(mdd_df, left_on='Participant ID', right_on='eid', how='left')
  434. mdd_df = pd.read_csv("self_reported_mdd_anx.csv")
  435. mdd_df["eid"] = mdd_df["eid"].astype(str)
  436. merged_df = df.merge(mdd_df, left_on="Participant ID", right_on="eid", how="left")
  437. mask_mdd = merged_df["p20544"].str.contains("Depression", case=False, na=False)
  438. mask_anx = merged_df["p20544"].str.contains("Anxiety", case=False, na=False)
  439. if "is_healthy" not in df.columns:
  440. raise ValueError("Expected 'is_healthy' in smri_prs.csv")
  441. df.loc[mask_mdd | mask_anx, "is_healthy"] = 0
  442. df["self_rep_mdd"] = mask_mdd.astype(int)
  443. df["self_rep_anx"] = mask_anx.astype(int)
  444. input_dim = deconfounded_embeddings.shape[1]
  445. model = Autoencoder(input_dim)
  446. checkpoint = torch.load("/opt/notebooks/model_ae.pth", map_location=torch.device('cpu'))
  447. model.load_state_dict(checkpoint)
  448. model.eval()
  449. deconfounded_embeddings = torch.tensor(deconfounded_embeddings, dtype=torch.float32)
  450. with torch.no_grad():
  451. deconfounded_embeddings_reconstructions = model(deconfounded_embeddings)
  452. l1_loss = nn.MSELoss(reduction='none')
  453. all_l1s_raw = l1_loss(deconfounded_embeddings_reconstructions, deconfounded_embeddings)
  454. dev_raw = deconfounded_embeddings_reconstructions - deconfounded_embeddings
  455. all_l1s_raw = all_l1s_raw.detach().cpu().numpy()
  456. all_l1s = np.mean(all_l1s_raw, axis=1)
  457. df['l1_loss'] = all_l1s
  458. df['l1_loss_raw'] = list(all_l1s_raw)
  459. df['dev_loss_raw'] = list(dev_raw)
  460. # Build the confounder matrix (shape: [n_samples, 2])
  461. confounders_df = df[['Sex', 'Age when attended assessment centre | Instance 2']].copy()
  462. confounders_df['Sex'] = confounders_df['Sex'].map({1: 0, 2: 1}).astype(float)
  463. confounders_df = confounders_df.astype(float)
  464. confounder_values = confounders_df.values # shape: (n_samples, 2)
  465. # Load regression params
  466. with open("/opt/notebooks/linear_regression_params.pkl", "rb") as f:
  467. regression_params = pickle.load(f)
  468. coefficients = regression_params["coefficients"]
  469. intercepts = regression_params["intercepts"]
  470. dev_raw_np = dev_raw.detach().cpu().numpy() if isinstance(dev_raw, torch.Tensor) else dev_raw.copy()
  471. dev_raw_np = StandardScaler().fit_transform(dev_raw_np)
  472. predicted = confounder_values @ coefficients.T + intercepts
  473. residuals = dev_raw_np - predicted
  474. df["dev_loss_raw_resid"] = list(residuals)
  475. print(f"Final Data Count: {df.shape}")
  476. healthies = df[df['is_healthy'] == 1]
  477. healthies["label"] = 0
  478. print(f"Number of Healthies: {healthies.shape}")
  479. final_results = []
  480. diagnosis_to_test = [
  481. 'DEP_ICD10', 'depression_1_no_comor',
  482. 'depression_2_no_comor', 'depression_3_no_comor', 'depression_4_no_comor', 'AUD_ICD10',
  483. "diagnosis_audit_8_no_comor", "diagnosis_audit_9_no_comor", "diagnosis_audit_10_no_comor", "AD_ICD10", 'GAD_ICD10',
  484. "ad_2_no_comor", "ad_3_no_comor", "ad_4_no_comor",
  485. "ad_2_and_depression_2_comor", "ad_3_and_depression_3_comor", "ad_4_and_depression_4_comor", "depression_ad_audit_2",
  486. "depression_ad_audit_3", "depression_ad_audit_4", "depression_ad_aud_3",
  487. "depression_ad_aud_4", "depression_ad_aud_2", "depression_1", "depression_2", "depression_3", "depression_4",
  488. "ad_2", "ad_3", "ad_4"
  489. ]
  490. ad_base = [
  491. "ad_2_no_comor",
  492. "ad_3_no_comor",
  493. "ad_4_no_comor"
  494. ]
  495. comor_base = [
  496. "ad_2_and_depression_2_comor",
  497. "ad_3_and_depression_3_comor",
  498. "ad_4_and_depression_4_comor"
  499. ]
  500. depr_base = [
  501. 'depression_1_no_comor',
  502. 'depression_2_no_comor',
  503. 'depression_3_no_comor',
  504. 'depression_4_no_comor', 'self_rep_mdd'
  505. ]
  506. df["depression_ad_audit_2"] = (
  507. (((df["depression_2"] == 1) | (df["depression_3"] == 1) | (df["depression_4"] == 1)) |
  508. ((df["ad_2"] == 1) | (df["ad_3"] == 1) | (df["ad_4"] == 1))) &
  509. ((df["diagnosis_audit_8"] == 1) | (df["diagnosis_audit_9"] == 1) | (df["diagnosis_audit_10"] == 1))
  510. ).astype(int)
  511. df["depression_ad_audit_3"] = (
  512. (((df["depression_3"] == 1) | (df["depression_4"] == 1)) |
  513. ((df["ad_3"] == 1) | (df["ad_4"] == 1))) &
  514. ((df["diagnosis_audit_8"] == 1) | (df["diagnosis_audit_9"] == 1) | (df["diagnosis_audit_10"] == 1))
  515. ).astype(int)
  516. df["depression_ad_audit_4"] = (
  517. ((df["depression_4"] == 1) |
  518. (df["ad_4"] == 1)) &
  519. ((df["diagnosis_audit_8"] == 1) | (df["diagnosis_audit_9"] == 1) | (df["diagnosis_audit_10"] == 1))
  520. ).astype(int)
  521. df["depression_ad_aud_2"] = ((
  522. ((df["depression_2"] == 1) | (df["depression_3"] == 1) | (
  523. df["depression_4"] == 1)) |
  524. ((df["ad_2"] == 1) | (df["ad_3"] == 1) | (df["ad_4"] == 1))) &
  525. (df["AUD_ICD10"] == 1)
  526. ).astype(int)
  527. df["depression_ad_aud_3"] = (
  528. (((df["depression_3"] == 1) | (df["depression_4"] == 1)) |
  529. ((df["ad_3"] == 1) | (df["ad_4"] == 1))) &
  530. (df["AUD_ICD10"] == 1)
  531. ).astype(int)
  532. df["depression_ad_aud_4"] = (
  533. (df["depression_4"] == 1) |
  534. (df["ad_4"] == 1) &
  535. (df["AUD_ICD10"] == 1)
  536. ).astype(int)
  537. diagnosis_to_test.append("is_healthy")
  538. for col in diagnosis_to_test:
  539. subset = df[df[col] == 1]
  540. fraction_females = (subset['Sex'] == 2).mean()
  541. age_col = "Age when attended assessment centre | Instance 2"
  542. mean_age = subset[age_col].mean()
  543. std_age = subset[age_col].std()
  544. print(f"Diagnosis: {col}")
  545. print(f"Fraction of females: {fraction_females:.2%}")
  546. print(f"Mean age: {mean_age:.2f}, Std age: {std_age:.2f}\n")
  547. dev_dict = {}
  548. dev_dict_resid = {}
  549. dev_healthies = np.stack(healthies["dev_loss_raw"].to_list())
  550. mse_healthies = np.stack(healthies["l1_loss_raw"].to_list())
  551. for diagnosis in diagnosis_to_test:
  552. patients = df[
  553. (df[diagnosis] == 1)
  554. ]
  555. patients["label"] = 1
  556. dev_dict[diagnosis] = np.stack(patients["dev_loss_raw"].to_list())
  557. dev_dict_resid[diagnosis] = np.stack(patients["dev_loss_raw_resid"].to_list())
  558. dev_healthies = np.stack(healthies["dev_loss_raw_resid"].to_list())
  559. hc_embeds_combined = dev_healthies
  560. normative_centroid = np.mean(hc_embeds_combined, axis=0)
  561. hc_dev = hc_embeds_combined - normative_centroid
  562. color_list = plt.cm.tab10(np.linspace(0, 1, len(dev_dict)))
  563. deviation_vectors = {}
  564. scaling_factor = 5
  565. angles = []
  566. magnitudes = []
  567. deviation_vectors_256d = []
  568. deviation_vectors_256d_perc_25 = []
  569. deviation_vectors_256d_perc_75 = []
  570. std_angles = []
  571. std_magnitudes = []
  572. group_colors = []
  573. used_cols = []
  574. group_labels = {
  575. 'alkohol': ["diagnosis_audit_8_no_comor", "diagnosis_audit_9_no_comor", "diagnosis_audit_10_no_comor"],
  576. 'alkohol_custom': ["AUD_ICD10"],
  577. 'depression': ['depression_1_no_comor', 'depression_2_no_comor', 'depression_3_no_comor',
  578. 'depression_4_no_comor'],
  579. 'depression_custom': ["DEP_ICD10", "self_rep_mdd"],
  580. 'angst': ["ad_2_no_comor", "ad_3_no_comor", "ad_4_no_comor"],
  581. 'angst_icd10': ["AD_ICD10", "GAD_ICD10"],
  582. 'depression_and_angst': ["ad_2_and_depression_2_comor", "ad_3_and_depression_3_comor",
  583. "ad_4_and_depression_4_comor"],
  584. 'mdd_aud_gad': ["depression_ad_aud_2", "depression_ad_aud_3", "depression_ad_aud_4"]
  585. }
  586. general_mapping = {
  587. '1': 'mild',
  588. '2': 'moderate',
  589. '3': 'moderate severe',
  590. '4': 'severe'
  591. }
  592. with open("/opt/notebooks/pca_2d_projection_allpoints.pkl", "rb") as f:
  593. pca = pickle.load(f)
  594. residuals_2d = pca.transform(residuals)
  595. hc_mask = df["is_healthy"] == 1
  596. hc_embeds_2d = residuals_2d[hc_mask]
  597. patient_embeds_2d = residuals_2d[~hc_mask]
  598. hc_center, hc_width, hc_height, hc_angle = bootstrap_ci_ellipse(hc_embeds_2d)
  599. shift_vector = hc_center
  600. hc_embeds_2d -= shift_vector
  601. patient_embeds_2d -= shift_vector
  602. plot_groups = {
  603. 'alkohol': [
  604. 'diagnosis_audit_10_no_comor', 'diagnosis_audit_8_no_comor',
  605. 'diagnosis_audit_9_no_comor', 'AUD_ICD10'
  606. ],
  607. 'depression': [
  608. 'depression_1_no_comor', 'depression_2_no_comor',
  609. 'depression_3_no_comor', 'depression_4_no_comor', 'DEP_ICD10', 'self_rep_mdd'
  610. ],
  611. 'angst': [
  612. 'ad_4_no_comor', 'ad_3_no_comor', 'ad_2_no_comor', "AD_ICD10", "GAD_ICD10", "self_rep_anx"
  613. ],
  614. 'comor': [
  615. 'ad_2_and_depression_2_comor', 'ad_3_and_depression_3_comor',
  616. 'ad_4_and_depression_4_comor'
  617. ],
  618. 'aud_gad_mdd': [
  619. 'depression_ad_audit_2', 'depression_ad_audit_3',
  620. 'depression_ad_audit_4'
  621. ]
  622. }
  623. strict_diag_cols = [c for group in plot_groups.values() for c in group]
  624. group_colors_map = {
  625. "AUDIT-C": "blue",
  626. "GAD": "green",
  627. "PHQ-9": "red",
  628. "GAD+PHQ-9": "black",
  629. "GAD/PHQ-9+AUDIT-C": "purple"
  630. }
  631. fig, ax = plt.subplots(figsize=(12, 12), dpi=300)
  632. hc_center, hc_width, hc_height, hc_angle = bootstrap_ci_ellipse(hc_embeds_2d)
  633. hc_ellipse = Ellipse(xy=hc_center, width=hc_width, height=hc_height, angle=hc_angle,
  634. edgecolor='black', facecolor='black', linestyle='--', linewidth=1.5, alpha=0.05)
  635. ax.add_patch(hc_ellipse)
  636. line_styles = ['solid', 'dashed', 'dotted', 'dashdot']
  637. style_tracker = {c: 0 for c in list(group_colors_map.values()) + ["gray"]}
  638. for group in strict_diag_cols:
  639. mask = df[group] == 1
  640. if mask.sum() == 0: continue
  641. center, width, height, angle = bootstrap_ci_ellipse(residuals_2d[mask])
  642. color = assign_color_from_label(group)
  643. linestyle = line_styles[style_tracker[color] % len(line_styles)]
  644. style_tracker[color] += 1
  645. ellipse = Ellipse(xy=center, width=width, height=height, angle=angle,
  646. edgecolor=color, facecolor=color, alpha=0.05)
  647. ax.add_patch(ellipse)
  648. ax.arrow(
  649. 0, 0, center[0], center[1],
  650. head_width=0.05, head_length=0.05,
  651. fc=color, ec=color,
  652. linewidth=1.5
  653. )
  654. ax.add_patch(ellipse)
  655. ax.text(center[0], center[1], replace_terms(group), fontsize=20, color=color)
  656. aud_hull = encapsulate_meta_group(["blue"])
  657. internalizing_hull = encapsulate_meta_group(["red", "black"])
  658. gad_hull = encapsulate_meta_group(["green", "black"])
  659. mix_hull = encapsulate_meta_group(["purple"])
  660. for hull, color, label in [(mix_hull, "purple", ""), (aud_hull, "blue", "AUD meta"),
  661. (internalizing_hull, "red", "Internalizing meta"), (gad_hull, "green", "GAD")]:
  662. if hull:
  663. x, y = hull.exterior.xy
  664. ax.plot(x, y, linestyle=(0, (5, 5)), color=color, linewidth=2)
  665. legend_elements = [
  666. Patch(facecolor='blue', edgecolor='blue', label='AUDIT-C-only'),
  667. Patch(facecolor='green', edgecolor='green', label='GAD-7-only'),
  668. Patch(facecolor='red', edgecolor='red', label='PHQ-9-only'),
  669. Patch(facecolor='black', edgecolor='black', label='GAD-7 + PHQ-9'),
  670. Patch(facecolor='purple', edgecolor='purple', label='GAD-7/PHQ-9 + AUDIT-C'),
  671. Line2D([0], [0], linestyle='--', color='gray', label='HC median ellipse', linewidth=1.5)
  672. ]
  673. ax.legend(handles=legend_elements, loc='best', fontsize=18)
  674. ax.tick_params(axis='both', which='major', labelsize=18)
  675. ax.tick_params(axis='both', which='minor', labelsize=18)
  676. ax.axhline(0, color='black', linewidth=1)
  677. ax.axvline(0, color='black', linewidth=1)
  678. ax.set_xlabel("PC1", fontsize=18)
  679. ax.set_ylabel("PC2", fontsize=18)
  680. ax.set_xlim(-2, 2)
  681. ax.set_ylim(-3, 1.5)
  682. plt.tight_layout()
  683. plt.savefig("/opt/notebooks/arrow_plot_STRICT_2Dcentroids_validation.pdf", dpi=300)
  684. plt.savefig("/opt/notebooks/arrow_plot_STRICT_2Dcentroids_validation.png", dpi=300)
  685. plt.show()

directional_analysis.py at commit 11cd20b, under GPL-3.0 · at the source

Overview

Authors: Julius Wiegert1,2, Sebastián Marty-Lombardi1,2, Jailan Oweda1,2, Esra Lenz1,2, Peter Ahnert3, Klaus Berger4, Hermann Brenner5,6, Josef Frank7, Hans J. Grabe8, Karin Halina Greiser9, Johanna Klinger-König8, André Karch4, Michael Leitzmann10, Claudia Meinke-Franze11, Rafael Mikolajczyk12,13, Frauke Nees14,15,16, Thoralf Niendorf17, Oliver Sander18, Carsten Oliver Schmidt11, Steffi G. Riedel-Heller19
and 15 other authorsKerstin Ritter20,21,22, Annette Peters23,24,25, Tobias Pischon26,27,28, Stephanie Witt7,29, Johannes Nitsche1,2, Joonas Naamanka1,2,30, Sebastian Volkmer1,2, Antonia Mai1,2, Amrou Abas1,2, Xiuzhi Li1,2, Andreas Meyer-Lindenberg2, Tobias Gradinger1,2, Fabian Streit1,2,29, Urs Braun1,2, Emanuel Schwarz1,2,29
30 affiliations
  1. Hector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University,Mannheim, Germany
  2. Department of Psychiatry and Psychotherapy, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University,Mannheim, Germany
  3. Universität Leipzig, Institute for Medical Informatics, Statistics and Epidemiology,Leipzig, Germany
  4. Institute of Epidemiology and Social Medicine, University of Münster,Münster, Germany
  5. German Cancer Research Center (DKFZ),Heidelberg, Germany
  6. Network Aging Research, Heidelberg University,Heidelberg, Germany
  7. Department of Genetic Epidemiology in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University,Mannheim, Germany
  8. Department of Psychiatry and Psychotherapy, University Medicine Greifswald,Greifswald, Germany
  9. Division of Cancer Epidemiology, German Cancer Research Center (DKFZ),Heidelberg, Germany
  10. Department of Epidemiology and Preventive Medicine, University of Regensburg,Regensburg, Germany
  11. Institute for Community Medicine, University Medicine Greifswald,Greifswald, Germany
  12. Institute for Medical Epidemiology, Biometrics, and Informatics, Medical Faculty of the Martin Luther University Halle-Wittenberg,Halle (Saale), Germany
  13. German Center for Mental Health (DZPG), partner site Halle-Jena-Magdeburg,Halle (Saale), Germany
  14. Institute of Medical Psychology and Medical Sociology, University Medical Center Schleswig-Holstein, Kiel University,Kiel, Germany
  15. Department of Child and Adolescent Psychiatry and Psychotherapy, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University,Mannheim, Germany
  16. Institute of Medical Psychology, Faculty of Medicine, Ludwig-Maximilians-Universität München,Munich, Germany
  17. Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrueck Center for Molecular Medicine in the Helmholtz Association (MDC),Berlin, Germany
  18. Clinic for Rheumatology and Hiller Research Center, University Hospital, Heinrich-Heine-University Düsseldorf,Düsseldorf, Germany
  19. Institute of Social Medicine, Occupational Health and Public Health (ISAP), Leipzig University,Leipzig, Germany
  20. Department of Machine Learning, Hertie Institute for AI in Brain Health, University of Tübingen,Tübingen, Germany
  21. Department of Psychiatry and Neurosciences, Charité - Universitätsmedizin Berlin (corporate member of Freie Universität Berlin, Humboldt-Universität zu Berlin, and Berlin Institute of Health),Berlin, Germany
  22. Tübingen AI Center,Tübingen, Germany
  23. Institute of Epidemiology, Helmholtz Zentrum München - German Research Center for Environmental Health (GmbH),Neuherberg, Germany
  24. Chair of Epidemiology, Institute for Medical Information Processing, Biometry and Epidemiology, Medical Faculty, Ludwig-Maximilians-Universität München,Munich, Germany
  25. German Center for Mental Health (DZPG), partner site Munich,Munich, Germany
  26. Max Delbrück Center for Molecular Medicine in the Helmholtz Association (MDC), Molecular Epidemiology Research Group,Berlin, Germany
  27. Charité - Universitätsmedizin Berlin, corporate member of Freie Universität Berlin and Humboldt-Universität zu Berlin,Berlin, Germany
  28. Max Delbrück Center for Molecular Medicine in the Helmholtz Association (MDC), Biobank Technology Platform,Berlin, Germany
  29. German Center for Mental Health (DZPG), Partner Site Mannheim - Heidelberg - Ulm,Heidelberg, Germany
  30. SleepWell Research Program, Faculty of Medicine, University of Helsinki,Helsinki, Finland
Journal: Molecular psychiatry, volume 31, issue 10, pages 6034-6045
Dates: received 11 February 2026; accepted 3 June 2026; published online 18 July 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41380-026-03691-4 · PMID 42471446 · PMCID PMC13569451 · OpenAlex W7169665349
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), other condition (population), depression (population), clinical / translational (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning
Keywords: Diagnostic markers, Depression, Neuroscience
MeSH: Anxiety*, Brain*, Depression*, Adult, Aged, Anxiety Disorders, Brain Mapping, Cohort Studies, Female, Germany, Humans, Magnetic Resonance Imaging, Male, Middle Aged (* major topic)
Topic: Functional Brain Connectivity Studies (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: project was supported by the DZPG (German Centre for Mental Health Research) and by the BMBF (German Ministry of Education and Research) grant 01EE2303E. This work was supported by the Hector foundation II
Citations: not cited yet (Europe PMC); 55 references in the paper

Abstract

Structural brain alterations associated with depression and anxiety are subtle, heterogeneous, and difficult to characterize. We applied autoencoder-based normative modeling to contrastively learned structural MRI representations from two large population-based cohorts (German National Cohort, N ≈ 29,000; UK Biobank, N ≈ 25,000) to quantify individual deviations from normative brain structure across symptom dimensions of depression, anxiety, and, for contextualization, alcohol use.

Deviation magnitude increased with symptom severity for depressive and anxiety symptoms and was most pronounced in individuals with high alcohol use. Directional analyses revealed shared deviation patterns for depression and anxiety that were largely distinct from alcohol-related deviations, and these patterns generalized across cohorts. These affective-symptom-related patterns implicated distributed regional brain-structural variation. Individual deviation profiles improved classification of symptomatic status beyond demographic covariates, with gains concentrated at higher symptom severity.

Together, these findings indicate that affective symptoms are associated with reproducible, dimensional patterns of regional brain-structural deviation that extend beyond normative population variability, supporting transdiagnostic models of internalizing psychopathology.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

wiegertj/DeepNormativeModeling

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 11cd20b2a6f5a7d74601c9b677b2243e78b764f2, 9 October 2025
Languages: Python (20), Shell (1)
Size: 25 files, 21 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (16 files), pandas (16 files), scikit-learn (10 files), PyTorch (9 files), SciPy (9 files), Matplotlib (8 files), statsmodels (4 files), NiBabel (3 files), seaborn (2 files), FSL (1 file), Nilearn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
23 files

Code availability

All analysis were conducted in Python (version 3.12). All analysis scripts and the utilized packages are available via GitHub at https://github.com/wiegertj/DeepNormativeModeling.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 21 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

Access to and use of NAKO data and biosamples can be obtained via the electronic application portal (https://transfer.nako.de). Access to UK Biobank data requires application through the registration and application portal (http://ukbiobank.ac.uk/register-apply).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Springer Nature

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 35 authors, 3 keywords, 14 MeSH terms, 1 funder, 52 references.

Cite

This paper

Wiegert, J., Marty-Lombardi, S., Oweda, J., Lenz, E., Ahnert, P., Berger, K., Brenner, H., Frank, J., Grabe, H. J., Greiser, K. H., Klinger-König, J., Karch, A., Leitzmann, M., Meinke-Franze, C., Mikolajczyk, R., Nees, F., Niendorf, T., Sander, O., Schmidt, C. O., . . . Schwarz, E. (2026). Breaking the norm: population-scale deviations of brain structure in depression and anxiety. Molecular psychiatry, 31(10), 6034-6045. https://doi.org/10.1038/s41380-026-03691-4

BibTeX

@article{wiegert2026breaking,
author = {Wiegert, Julius and Marty-Lombardi, Sebastián and Oweda, Jailan and Lenz, Esra and Ahnert, Peter and Berger, Klaus and Brenner, Hermann and Frank, Josef and Grabe, Hans J. and Greiser, Karin Halina and Klinger-König, Johanna and Karch, André and Leitzmann, Michael and Meinke-Franze, Claudia and Mikolajczyk, Rafael and Nees, Frauke and Niendorf, Thoralf and Sander, Oliver and Schmidt, Carsten Oliver and Riedel-Heller, Steffi G. and Ritter, Kerstin and Peters, Annette and Pischon, Tobias and Witt, Stephanie and Nitsche, Johannes and Naamanka, Joonas and Volkmer, Sebastian and Mai, Antonia and Abas, Amrou and Li, Xiuzhi and Meyer-Lindenberg, Andreas and Gradinger, Tobias and Streit, Fabian and Braun, Urs and Schwarz, Emanuel},
title = {{Breaking the norm: population-scale deviations of brain structure in depression and anxiety}},
journal = {Molecular psychiatry},
year = {2026},
month = jul,
volume = {31},
number = {10},
pages = {6034--6045},
publisher = {Springer Nature},
issn = {1359-4184},
doi = {10.1038/s41380-026-03691-4},
url = {https://doi.org/10.1038/s41380-026-03691-4},
pmid = {42471446},
pmcid = {PMC13569451}
}

RIS

TY - JOUR
AU - Wiegert, Julius
AU - Marty-Lombardi, Sebastián
AU - Oweda, Jailan
AU - Lenz, Esra
AU - Ahnert, Peter
AU - Berger, Klaus
AU - Brenner, Hermann
AU - Frank, Josef
AU - Grabe, Hans J.
AU - Greiser, Karin Halina
AU - Klinger-König, Johanna
AU - Karch, André
AU - Leitzmann, Michael
AU - Meinke-Franze, Claudia
AU - Mikolajczyk, Rafael
AU - Nees, Frauke
AU - Niendorf, Thoralf
AU - Sander, Oliver
AU - Schmidt, Carsten Oliver
AU - Riedel-Heller, Steffi G.
AU - Ritter, Kerstin
AU - Peters, Annette
AU - Pischon, Tobias
AU - Witt, Stephanie
AU - Nitsche, Johannes
AU - Naamanka, Joonas
AU - Volkmer, Sebastian
AU - Mai, Antonia
AU - Abas, Amrou
AU - Li, Xiuzhi
AU - Meyer-Lindenberg, Andreas
AU - Gradinger, Tobias
AU - Streit, Fabian
AU - Braun, Urs
AU - Schwarz, Emanuel
TI - Breaking the norm: population-scale deviations of brain structure in depression and anxiety
T2 - Molecular psychiatry
J2 - Mol Psychiatry
PY - 2026
DA - 2026/07/18
VL - 31
IS - 10
SP - 6034
EP - 6045
SN - 1359-4184
PB - Springer Nature
DO - 10.1038/s41380-026-03691-4
UR - https://doi.org/10.1038/s41380-026-03691-4
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41380-026-03691-4",
"type": "article-journal",
"title": "Breaking the norm: population-scale deviations of brain structure in depression and anxiety",
"container-title": "Molecular psychiatry",
"author": [
{
"family": "Wiegert",
"given": "Julius"
},
{
"family": "Marty-Lombardi",
"given": "Sebastián"
},
{
"family": "Oweda",
"given": "Jailan"
},
{
"family": "Lenz",
"given": "Esra"
},
{
"family": "Ahnert",
"given": "Peter"
},
{
"family": "Berger",
"given": "Klaus"
},
{
"family": "Brenner",
"given": "Hermann"
},
{
"family": "Frank",
"given": "Josef"
},
{
"family": "Grabe",
"given": "Hans J."
},
{
"family": "Greiser",
"given": "Karin Halina"
},
{
"family": "Klinger-König",
"given": "Johanna"
},
{
"family": "Karch",
"given": "André"
},
{
"family": "Leitzmann",
"given": "Michael"
},
{
"family": "Meinke-Franze",
"given": "Claudia"
},
{
"family": "Mikolajczyk",
"given": "Rafael"
},
{
"family": "Nees",
"given": "Frauke"
},
{
"family": "Niendorf",
"given": "Thoralf"
},
{
"family": "Sander",
"given": "Oliver"
},
{
"family": "Schmidt",
"given": "Carsten Oliver"
},
{
"family": "Riedel-Heller",
"given": "Steffi G."
},
{
"family": "Ritter",
"given": "Kerstin"
},
{
"family": "Peters",
"given": "Annette"
},
{
"family": "Pischon",
"given": "Tobias"
},
{
"family": "Witt",
"given": "Stephanie"
},
{
"family": "Nitsche",
"given": "Johannes"
},
{
"family": "Naamanka",
"given": "Joonas"
},
{
"family": "Volkmer",
"given": "Sebastian"
},
{
"family": "Mai",
"given": "Antonia"
},
{
"family": "Abas",
"given": "Amrou"
},
{
"family": "Li",
"given": "Xiuzhi"
},
{
"family": "Meyer-Lindenberg",
"given": "Andreas"
},
{
"family": "Gradinger",
"given": "Tobias"
},
{
"family": "Streit",
"given": "Fabian"
},
{
"family": "Braun",
"given": "Urs"
},
{
"family": "Schwarz",
"given": "Emanuel"
}
],
"container-title-short": "Mol Psychiatry",
"volume": "31",
"issue": "10",
"page": "6034-6045",
"DOI": "10.1038/s41380-026-03691-4",
"PMID": "42471446",
"PMCID": "PMC13569451",
"ISSN": "1359-4184",
"publisher": "Springer Nature",
"URL": "https://doi.org/10.1038/s41380-026-03691-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
18
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41398-026-04131-1 [code]
Multimodal phenotypic classification of generalized anxiety and panic using structural MRI data and psychosocial factors: machine learning results from the German National Cohort (NAKO) study.
Journal: Translational psychiatry
In common: structural MRI / diffusion, clinical / translational, other condition, 2 references, 4 authors
[2] doi:10.1162/imag.a.1269 [code]
From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Nilearn, NiBabel, statsmodels, 6 other tools, 9 references
[3] doi:10.1038/s41398-026-04426-3 [code]
Resting-state brain age is associated with cognitive and sensorimotor abnormalities in schizophrenia spectrum disorders: a validation &amp; longitudinal study.
Journal: Translational psychiatry
In common: NiBabel, statsmodels, pandas, 1 other tool, structural MRI / diffusion, 1 reference, 3 authors
[4] doi:10.1038/s41398-026-04025-2 [code]
Brain energetic landscapes shape state dysregulation in major depressive disorder: a morphological network controllability perspective.
Journal: Translational psychiatry
In common: FSL, Nilearn, NiBabel, 7 other tools, depression, 2 references
[5] doi:10.1371/journal.pmed.1004809 [code]
Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data.
Journal: PLoS medicine
In common: NiBabel, statsmodels, seaborn, 5 other tools, structural MRI / diffusion, other condition, 4 references
[6] doi:10.64898/2026.08.13.26360304 [code]
Lifespan brain structural variation reveals shared organization across mental health conditions
Journal: medRxiv (preprint)
In common: Nilearn, NiBabel, statsmodels, 6 other tools, clinical / translational, 3 references
[7] doi:10.1038/s41593-026-02359-0 [code]
The cross-site reproducibility of MRI morphometric phenotypes in psychiatric disorders.
Journal: Nature neuroscience
In common: FSL, NiBabel, scikit-learn, 4 other tools, structural MRI / diffusion, 4 references
[8] doi:10.1038/s43856-026-01722-3 [code]
Local and global patterns support medical imaging as a biomarker of ageing.
Journal: Communications medicine
In common: NiBabel, statsmodels, PyTorch, 5 other tools, clinical / translational, author Tobias Pischon
[9] doi:10.1038/s44220-026-00680-y [code]
The neuroimaging correlates of depression established across six large-scale population datasets.
Journal: Nature. Mental health
In common: statsmodels, seaborn, scikit-learn, 4 other tools, depression, 4 references
[10] doi:10.1371/journal.pbio.3003856 [code]
Aging and metabolism contribute separately to brain-body health.
Journal: PLoS biology
In common: FSL, Nilearn, NiBabel, 8 other tools, structural MRI / diffusion, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.