OSCR

Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches
  1. [1] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/01 - model0_RS/model0.py, lines 1–59 · score 0.84 · feature vectors, retained variance, cross validation, imputed, median, bounded
  2. [2] § Materials and methods › Feature set construction ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 273–293 · score 0.83 · Stop Signal, sMRI, working memory, fMRI, Cued, RSS
  3. [3] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 1–58 · score 0.82 · retained variance, cross validation, rsfMRI, chosen, imputed, median
  4. [4] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/D_confusion_matrix_grid.py, lines 1–35 · score 0.76 · Row normalized confusion, best multimodal model, bACC, single modality models, best single modality, matrices
  5. [5] § Materials and methods › ML-framework ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 1–58 · score 0.73 · machine learning, structural MRI, rsfMRI, single modalities, SVMs, dummy
  6. [6] § Materials and methods › ML-framework ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 1–67 · score 0.71 · machine learning, structural MRI, rsfMRI, single modalities, SVMs, dummy
  7. [7] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 295–319 · score 0.67 · general eating behaviors, Eating unspecific, equal weights, DEBQ, FCQ, subscales
  8. [8] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 295–319 · score 0.66 · general depressive symptoms, Disease unspecific, weight fluctuations, BDI, Composite, domains
  9. [9] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/A_plot_heatmap_target_modality.py, lines 1–27 · score 0.65 · cross validated balanced, rsfMRI, balanced accuracy, Heatmap, intrinsic, bh
  10. [10] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/D_confusion_matrix_grid.py, lines 1–35 · score 0.64 · confusion matrices, bACC, best single modality, classes, classification, Figure 2
  11. [11] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/06 - plotting/FIG 04 - RegrResults/scripts/C_regr_univariate_results_targetXgroup.py, lines 55–74 · score 0.64 · Disease unspecific, owHC, nwHC, BDI, severity, scores
  12. [12] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 1–67 · score 0.63 · dummy baseline, fitting, fusion, concatenated, single modality, component
  13. [13] § Materials and methods › Participants ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 273–293 · score 0.60 · working memory, fMRI, age, sex, overweight, BMI
  14. [14] § Results › Unimodal regression ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 109–134 · score 0.53 · Eating unspecific, regression targets, signal, R2, variance, severity
  15. [15] § Results › Multimodal regression ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 125–144 · score 0.52 · Eating unspecific severity, disease unspecific, Single modality, R2, accuracy, weight
  16. [16] § Results › Unimodal regression ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 125–144 · score 0.52 · Eating unspecific, regression targets, signal, R2, variance, severity

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 776 lines · 41 KB · MIT · 4 matches

  1. # %% [markdown]
  2. # <h1 style="color: #1e88e5; font-weight: bold; margin-bottom: 5px;">NeuroBED_ML: DATA EXPLORATION NOTEBOOK</h1>
  3. # <hr style="border: 2px solid #cfd8dc; margin-top: 0; margin-bottom: 20px;">
  4. # %% [markdown]
  5. # <div style="color: #37474f; font-size: 16px;">
  6. # <p>
  7. # This notebook performs data exploration and additional preprocessing of the NeuroBED dataset in preparation for machine learning using the
  8. # <code style="font-size: 16px;">julearn</code> framework. At this point, the variables used for analyses have already been extracted from the raw data files and the data underwent minimal preprocessing (renaming variables, uniform naming convention for missing data etc.). Steps in this notebook include loading and cleaning data, handling missing values, removing outliers, and exporting structured outputs for modeling. The data will be written to file after each preprocessing step and stored here:
  9. # <code style="font-size: 16px;">/02 - data_exploration/output/</code>.
  10. # </p>
  11. #
  12. # <p>
  13. # The files used for analysis are additionally saved here:
  14. # <code style="font-size: 16px;">/03 - data4ML</code>.
  15. # </p>
  16. #
  17. # <p>
  18. # <code style="font-size: 16px;">allData_precleaned_step2.2c_Final.pkl</code>: All data<br>
  19. # <code style="font-size: 16px;">data_cats.pkl</code>: Info about category-variable mapping.
  20. # </p>
  21. #
  22. # <p style="font-size: 16px; margin-top: 1em;">
  23. # <strong>Please note that some variables have been renamed in the paper for clarity:</strong><br><br>
  24. # <span style="font-size: 16px;">fMRI&nbsp;&nbsp;&nbsp;= rsfMRI</span><br>
  25. # <span style="font-size: 16px;">MRI&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;= GMV</span><br>
  26. # <span style="font-size: 16px;">CON_BED = owHC</span><br>
  27. # <span style="font-size: 16px;">CON_BN&nbsp;&nbsp;&nbsp;= nwHC</span><br>
  28. # <span style="font-size: 16px;">Weight history = weight fluctuations and monitoring</span>
  29. # </p>
  30. # </div>
  31. # %% [markdown]
  32. # <div style="color: #37474f; border-left: 5px solid #607d8b; background-color: #f5f5f5; padding: 15px; margin-bottom: 20px; border-radius: 5px;">
  33. # <h3><strong style="color: #1e88e5;">Table of Contents</strong></h3>
  34. # <ul>
  35. # <li><a href="#section-0" style="color: #37474f;"><strong>0 Notebook Styling Params</strong></a></li>
  36. # <li><a href="#section-1" style="color: #37474f;"><strong>1 Imports</strong></a>
  37. # <ul>
  38. # <li><a href="#section-1-1" style="color: #37474f;">1.1 Variables</a></li>
  39. # <li><a href="#section-1-2" style="color: #37474f;">1.2 Load Data</a></li>
  40. # </ul>
  41. # </li>
  42. # <li><a href="#section-2" style="color: #37474f;"><strong>2 Preprocessing</strong></a>
  43. # <ul>
  44. # <li><a href="#section-2-1" style="color: #37474f;">2.1 Data Categorization</a></li>
  45. # <ul>
  46. # <li><a href="#section-2-1-1" style="color: #37474f;">2.1.1 Full Categorization</a></li>
  47. # <li><a href="#section-2-1-2" style="color: #37474f;">2.1.2 Model Categorization</a></li>
  48. # <ul>
  49. # <li><a href="#section-2-1-2.1" style="color: #37474f;">2.1.2.1 Features</a></li>
  50. # <li><a href="#section-2-1-2.2" style="color: #37474f;">2.1.2.2 Targets</a></li>
  51. # <li><a href="#section-2-1-2.3" style="color: #37474f;">2.1.2.2 Confounders & Others</a></li>
  52. # </ul>
  53. # </ul>
  54. # <li><a href="#section-2-2" style="color: #37474f;">2.2 Outliers and Missing Data</a></li>
  55. # <li><a href="#section-2-3" style="color: #37474f;">2.3 Save Preprocessed Data</a></li>
  56. # </ul>
  57. # </li>
  58. # <li><a href="#section-3" style="color: #37474f;"><strong>4 Data Exploration</strong></a>
  59. # <ul>
  60. # <li><a href="#section-3-1" style="color: #37474f;">4.1 Raw Data Inspection</a></li>
  61. # <li><a href="#section-3-2" style="color: #37474f;">4.2 Final Sample Inspection</a></li>
  62. # <ul>
  63. # <li><a href="#section-3-2" style="color: #37474f;">4.2.1 Features</a></li>
  64. # <li><a href="#section-3-2" style="color: #37474f;">4.2.2 Targets</a></li>
  65. # </ul>
  66. # </ul>
  67. # </li>
  68. # </ul>
  69. # </div>
  70. # %% [markdown]
  71. # | Step | Description | Object(s) | Input File(s) | Output File(s) |
  72. # |------|-------------|-----------|----------------|-----------------|
  73. # | 1 | Load input data (precleaned, selected variables) | `df_all`, `vars_all`, `rs_matrix` | `allData_precleaned.pkl`, `allVars_precleaned.pkl`, `NeuroBED_ML_AAL3Matrix_precleaned.pkl` | – |
  74. # | 2 | Convert all data columns to float, drop empty ones | `df_all`, `dropped_cols`, `colsb4` | – | `allData_precleaned_step1.pkl`,` allVars_precleaned_step1.pkl` |
  75. # | 2.1a | Generate nested modality-feature mapping | `data_cats_full` | `NeuroBED_ML_Vars.xlsx` | – |
  76. # | 2.1b | Generate final analysis categorization | `data_cats` | `data_cats_full` | - |
  77. # | 2.1c | Add severity indices | `df_all` | – | `allData_precleaned_step2.1a.pkl`, `data_cats.pkl`,`data_cats_full.pkl` |
  78. # | 2.2a | Calculate outliers/remove preselected subs | `zscore`, `zscore2`, `df_all2`,`allMissingData` | `df_all` | - |
  79. # | 2.2b | Replace outlier data with NaN| `df_all3`,`rs_matrix`| `df_all2`,`allMissingData` | - |
  80. # | 2.2c | Export Data and Cleaning summaries | `summary_df`,`allMissingData`,`df_all3`,`zscore`, `zscore2`,`rs_matrix` | – | `allData_precleaned_step2.2a_Z.pkl`, `allData_precleaned_step2.2b_Z.pkl`, `allData_precleaned_step2.2b.pkl`, `allData_precleaned_step2.2c_Final.pkl`, `subject_exclusion_summary.xlsx`, `data_cleaning_summary.csv`, `data_cleaning_summary.xlsx`, `AAL3Matrix_precleaned_step2.2c_Final.pkl`|
  81. # %% [markdown]
  82. # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">0. NOTEBOOK STYLING PARAMS</h3>
  83. # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
  84. # %%
  85. # Printing style
  86. from IPython.display import Markdown, display, Image,HTML
  87. from pprint import pprint
  88. def printmd(string):
  89. display(Markdown(string))
  90. # %% [markdown]
  91. # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">1. IMPORTS</h3>
  92. # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
  93. # %%
  94. import os, sys, random
  95. from pathlib import Path
  96. basepath = f"{os.getcwd()[:2]}/NeuroBed_ML/analysis/"
  97. if not basepath in sys.path:
  98. sys.path.insert(0,basepath) # to get code from other files #"E:/NeuroBed_ML/analysis/"
  99. sys.path.insert(1, f"{basepath}05 - utils/") # to get code from other files
  100. from NeuroBED_ML_helpers import *
  101. from LR_Cmaps import *
  102. import pandas as pd
  103. import numpy as np
  104. from scipy import stats
  105. from scipy.stats import pearsonr
  106. import matplotlib.pyplot as plt
  107. from itertools import combinations
  108. import matplotlib_inline
  109. import seaborn as sns
  110. import pickle,time
  111. from datetime import datetime
  112. # plotting parameters
  113. try:
  114. plt.style.use(f'{basepath}05 - utils/LR_ST2.mplstyle')
  115. matplotlib_inline.backend_inline.set_matplotlib_formats('retina')
  116. fonts = ['Poppins','Montserrat','Manrope','Figtree','Helvetica','Roboto','Quicksand','Arial','AvenirLTPro-Book']
  117. plt.rcParams['font.family'] = fonts[random.randint(0,len(fonts)-1)]
  118. except:
  119. plt.rcParams['font.family'] = 'Arial'
  120. print ('current font: ', plt.rcParams['font.family'][0])
  121. # %% [markdown]
  122. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">1.1 Variables</h4>
  123. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  124. # %% [markdown]
  125. # <p style="color: #37474f; font-size: 16px;">
  126. # This section sets up the configuration, including naming conventions and I/O structure, and then loads the data:
  127. # </p>
  128. #
  129. # <ul style="color: #37474f; font-size: 15px;">
  130. # <li><strong>Input Directory:</strong> Points to extracted and minimally cleaned data under <code>/01 - data/</code>.</li>
  131. # <li><strong>Output Directory:</strong> Created if it doesn't exist. Located under <code>/02 - data_exploration/output/</code>.</li>
  132. # <li><strong>Group Mapping:</strong> Defines numerical IDs for each group:
  133. # <ul>
  134. # <li><code>BED</code> → 1</li>
  135. # <li><code>BN</code> → 2</li>
  136. # <li><code>CON_BED (owHC)</code> → 3</li>
  137. # <li><code>CON_BN (nwHC)</code> → 4</li>
  138. # </ul>
  139. # </li>
  140. # </ul>
  141. #
  142. # <p style="color: #37474f; font-size: 15px;">
  143. # These group IDs are used for subject exclusion, stratified cross validation and plotting.
  144. # </p>
  145. # %%
  146. parent_dir = basepath
  147. in_dir = resolve_path('data')
  148. out_dir_ML = resolve_path('data4ML') # for output files that are used for ML models. creates dir if it does not exist
  149. out_dir_saf = resolve_path('data_exploration/output/additional_data_safeties') # safety outputs. creates dir if it does not exist
  150. out_dir_expl = resolve_path('data_exploration/output/data_cleaning') # exploration outputs. creates dir if it does not exist
  151. groupNames = {'BED':1,'CON_BED':3,'BN':2,'CON_BN':4}
  152. groupNames_rev = { v: k for k, v in groupNames.items()}
  153. saveFiles = False # save all output files from this notebook. Can be manually adjusted for each section
  154. # %% [markdown]
  155. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">1.2 Load Data</h4>
  156. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  157. # %% [markdown]
  158. # <p style="color: #37474f; font-size: 16px;">
  159. # The loaded dataset is already complete with all data types and pre-cleaned.
  160. # </p>
  161. # %%
  162. df_all = pd.read_pickle(f"{in_dir}/NeuroBED_ML_allData_precleaned.pkl") # import all data (except rsfMRI)
  163. rs_matrix = pd.read_pickle(f"{in_dir}/NeuroBED_ML_AAL3Matrix_precleaned.pkl") # import resting-state matrix (112x166x166)
  164. vars_all = pd.read_pickle(f"{in_dir}/NeuroBED_ML_allVars_precleaned.pkl") # variable info
  165. # %% [markdown]
  166. # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">2. PREPROCESSING</h3>
  167. # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
  168. # %% [markdown]
  169. # <p style="color: #37474f; font-size: 15px;">
  170. # Before the features are categorized, empty features are removed from the dataset. The remaining feature values are converted to float for compatibility with the ML pipeline. The result is saved here:
  171. # <code>allVars_precleaned_remEmptyCols.pkl/xlsx</code>
  172. # </p>
  173. # %%
  174. saveFile = saveFiles # use global parameter
  175. # %%
  176. # convert all data to float
  177. df_all = df_all.apply(pd.to_numeric, errors='coerce')
  178. # remove empty features. There are some MRI features where this happend
  179. colsb4 = list(df_all.columns)
  180. df_all.dropna(axis=1, how='all',inplace=True)
  181. dropped_cols = list(set(colsb4) - set(list(df_all.columns)))
  182. print ('dropped ',str(len(dropped_cols)), ' feature columns from the dataframe')
  183. print (dropped_cols)
  184. # remove these columns from the variable list
  185. vars_all = vars_all[~vars_all['feature'].isin(dropped_cols)]
  186. # %% [markdown]
  187. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.1 Data Categorization</h4>
  188. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  189. # %% [markdown]
  190. # <p style="color: #37474f; font-size: 15px;">
  191. # This section lists all variables (features), organized by modality and stimulus type, as defined in
  192. # <code>data/variable_info/NeuroBED_ML_Vars.xlsx</code> (e.g., <code>category1 → category2 → category3</code>).
  193. # <br><br>
  194. # Because there are many categories with different numbers of variables, we also combine related items into
  195. # clinically meaningful groups. These groups make the results easier to interpret and keep the analysis concise.
  196. # We save these groupings so they can be reused when training machine-learning models with different sets of variables.
  197. # </p>
  198. # %%
  199. saveFile = saveFiles # use global parameter
  200. # %% [markdown]
  201. # <h5 style="color: #37474f; font-weight: normal; font-size: 18px; margin-top: 30px;">2.1.1 Full Categorization</h4>
  202. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  203. # %%
  204. # Data categories and combinations
  205. data_cats_full = {}
  206. data_cats_full['cat1'] = vars_all.groupby(['category1'])['feature'].apply(list).to_dict()
  207. data_cats_full['cat2'] = vars_all.groupby(['category2'])['feature'].apply(list).to_dict()
  208. data_cats_full['cat3'] = vars_all.groupby(['category3'])['feature'].apply(list).to_dict()
  209. data_cats_full['cat1cat2'] = vars_all.groupby(['category1','category2'])['feature'].apply(list).to_dict()
  210. data_cats_full['cat1cat3'] = vars_all.groupby(['category1','category3'])['feature'].apply(list).to_dict()
  211. data_cats_full['cat2cat3'] = vars_all.groupby(['category2','category3'])['feature'].apply(list).to_dict()
  212. data_cats_full['cat1cat2cat3'] = vars_all.groupby(['category1','category2','category3'])['feature'].apply(list).to_dict()
  213. if saveFile: # save this too for later
  214. with open(str(in_dir)+'data_cats_full.pkl', 'wb') as f:
  215. pickle.dump(data_cats_full, f)
  216. vars_all.groupby(['category1','category2','category3']).describe()
  217. # task fMRI: 136+2=138 (AAL3 minus thalamus + AAL2 thalamus)
  218. # VBM: 166 (AAL3 only)
  219. # rs: 165x166/2 = 13695 (AAL3 only minus diagonal)
  220. # %% [markdown]
  221. # <h5 style="color: #37474f; font-weight: normal; font-size: 18px; margin-top: 30px;">2.1.1 Model Categorization</h4>
  222. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  223. # %% [markdown]
  224. # <p style="color: #37474f; font-size: 15px;">
  225. # This is the final grouping for the analyses.
  226. # </p>
  227. # %%
  228. data_cats = {'features': {},'targets': {},'confounders': {},'others': {}}
  229. # %% [markdown]
  230. # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.1 FEATURES</h4>
  231. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  232. # %%
  233. Image(filename='./plots/FeaturesOnly.png',width=300, height=150)
  234. # %%
  235. features = {'bh_neu': [('behavioral', 'CuedTask', 'all'),('behavioral', 'GNG', 'neu'),('behavioral', 'MID', 'neu'),('behavioral', 'RSS', 'all'),
  236. ('behavioral', 'StopSignal', 'all'),('behavioral', 'intelligence', 'all'),('behavioral', 'working memory', 'all')],
  237. 'bh_spec': [('behavioral', 'GNG', 'spec'),('behavioral', 'MID', 'spec')],
  238. 'blood': [('blood', 'all', 'all')],
  239. 'task_fMRI_spec':[('fMRI', 'GNG', 'spec'),('fMRI', 'MID', 'spec')],
  240. 'task_fMRI_neu': [('fMRI', 'GNG', 'neu'),('fMRI', 'MID', 'neu')],
  241. 'fMRI': [('fMRI', 'RS', 'all')],
  242. 'MRI' : [('sMRI', 'struct', 'all')]}
  243. data_cats['confounders'] = ['sex','age','BMI']
  244. data_cats['others'] = ['VAS_hungr_before','VAS_mood_before','VAS_hungr_after','VAS_mood after','hightest_weight',
  245. 'age_highest_weight','lowest_weight','age_lowest_weight','look','bodyshape_father','bodyshape_mother',
  246. 'first_time_overweight']
  247. for mod,subkeys in features.items():
  248. data_cats['features'][mod] = [feat for tup in features[mod] for feat in data_cats_full['cat1cat2cat3'][tup]]
  249. # Display the new category - feature mapping. Limit the display to x features
  250. displayDictJupyter(data_cats['features'],val_lim=30)
  251. # %% [markdown]
  252. # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.2 TARGETS</h4>
  253. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  254. # %% [markdown]
  255. # <p style="color: #37474f; font-size: 15px;"> To measure how severe the symptoms were across different areas, we created summary scores (composite indices) from selected questionnaire features. These scores are based on values that were first standardized to a common scale (ranging from 0 to 1). In some cases, features were given more or less weight to reflect their clinical importance within a domain. </p> <p style="color: #37474f; font-size: 15px;"> We evaluated five symptom areas: </p> <ol style="color: #37474f; font-size: 15px;"> <li><code><strong>Disease unspecific</strong></code> – general depressive symptoms (e.g., BDI)</li> <li><code><strong>Weight history</strong></code> – history of weight changes and frequency of self-weighing</li> <li><code><strong>Eating unspecific</strong></code> – general eating behaviors (DEBQ and FCQ subscales)</li> <li><code><strong>Eating specific</strong></code> – eating disorder–related symptoms (EDEQ total score)</li> <li><code><strong>Binge specific</strong></code> – frequency of binge eating episodes</li> </ol> <p style="color: #37474f; font-size: 15px;"> The severity scores were calculated in two steps: </p> <ul style="color: #37474f; font-size: 15px;"> <li>Each feature was rescaled to a 0–1 range so that all measures were comparable.</li> <li>The features were then averaged, either equally or with weights applied, to produce one score per symptom area.</li> </ul> <p style="color: #37474f; font-size: 15px;"> For the <strong>Eating unspecific</strong> area, we weighted the features so that the DEBQ (3 items) and the FCQ (2 items) each contributed equally (50% each). A similar approach was used for the <strong>Weight history</strong> area, with equal weight (50% each) given to weight fluctuations and number of weigh-ins. </p> <p style="color: #37474f; font-size: 15px;"> The final scores were added to the dataset with the suffix <code>_severity</code> (e.g., <code>Disease_unspecific_severity</code>). </p>
  256. # %%
  257. severity_idx = {'Disease_unspecific': ['BDI'],
  258. 'Weight_history': ['nr_gains_plus5','nr_gains_plus6to10','nr_gains_plus11to20','nr_gains_plus21to30','nr_gains_plus31to40',
  259. 'nr_gains_plus40','n_weighings'],
  260. 'Eating_unspecific': ['DEBQ_restrained_eating', 'DEBQ_emotional_eating', 'DEBQ_external_eating','G_FCQ_state', 'G_FCQ_trait'],
  261. 'Eating_specific': ['EDEQ_total_score'],
  262. 'Binge_specific': ['n_binges'],
  263. }
  264. weights = {'Eating_unspecific': 3*[0.5/3] + 2*[0.25], # weights of the individual factors to the total score (e.g. 50% DEBQ + 50% FCQ)
  265. 'Weight_history': 6*[0.5/6] + [0.55]}
  266. # calculate the severity indices
  267. suffix = '_severity'
  268. df_all = compute_weighted_severity_indices(df_all, severity_idx, weights,suffix=suffix)
  269. # Assign the new severity indices names to the data_cats dict
  270. data_cats['targets'] = {f"{key}{suffix}": value for key, value in severity_idx.items()}
  271. # %% [markdown]
  272. # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.3 CONFOUNDERS & OTHERS</h4>
  273. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  274. # %%
  275. data_cats['confounders'] = ['sex','age','BMI']
  276. data_cats['others'] = ['VAS_hungr_before','VAS_mood_before','VAS_hungr_after','VAS_mood after','hightest_weight',
  277. 'age_highest_weight','lowest_weight','age_lowest_weight','look','bodyshape_father','bodyshape_mother',
  278. 'first_time_overweight']
  279. # %% [markdown]
  280. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.2 Outliers and Missing Data</h4>
  281. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  282. # %% [markdown]
  283. #
  284. # %% [markdown]
  285. # <p style="color: #37474f; font-size: 15px;"> To maintain data quality across different types of measures, we flagged subjects with too much missing or unusable information. This included both missing values and extreme outliers (based on z-scores) across categories such as behavioral, blood, clinical, demographics, fMRI, questionnaires, and sMRI. </p> <p style="color: #37474f; font-size: 15px;"> Subjects were flagged for exclusion if they met either of these thresholds: </p> <ul style="color: #37474f; font-size: 15px;"> <li><strong>Missing data cutoff</strong>: more than 50% missing in a category</li> <li><strong>Z-score threshold</strong>: values above 10 (extreme outliers)</li> </ul> <p style="color: #37474f; font-size: 15px;"> In addition, 9 subjects were removed based on prior knowledge or manual review: <code>K020SJ</code>, <code>K030CN</code>, <code>K031KB</code>, <code>K089KP</code>, <code>K091AH</code>, <code>P010SV</code>, <code>P011GR</code>, <code>P022KB</code>, <code>P025JB</code>. </p> <p style="color: #37474f; font-size: 15px;"> A summary matrix (<code>data_cleaning_summary</code>) was created to track missing and flagged values for each subject and category. Exclusion was decided separately for each category and flagged if the proportion of missing data or outliers exceeded the set thresholds. </p> <p style="color: #37474f; font-size: 15px;"> Instead of removing entire subjects, only the flagged categories were set to <code>NaN</code>. This allowed subjects with partial data to still contribute to analyses in other categories. </p> <p style="color: #37474f; font-size: 15px;"> The following output files were saved: </p> <ul style="color: #37474f; font-size: 15px;"> <li><code>data_cleaning_summary.csv</code> / <code>.xlsx</code> – summary of missing values and outliers by subject and category</li> <li><code>subject_exclusion_summary.xlsx</code> – simplified version of the cleaning summary showing excluded subjects by category</li> <li><code>allData_precleaned_step2.2a_Z.pkl</code> – z-scored version of the main dataframe</li> <li><code>allData_precleaned_step2.2b_Z.pkl</code> – z-scored dataframe with preselected subjects removed</li> <li><code>allData_precleaned_step2.2b.pkl</code> – main dataframe with preselected subjects removed</li> <li><code>allData_precleaned_step2.2c_Final.pkl</code> – final dataframe with outlier data removed</li> </ul>
  286. # %% [markdown]
  287. # <p style="color: #37474f; font-size: 15px;">
  288. # (a) calculate outliers + remove preselected subjects
  289. # </p>
  290. # %%
  291. # features: multiple columns for each modality. Can take some missing data
  292. missing_cutoff = {'bh_neu': 0.5,'bh_spec': 0.5, 'blood': 0.5,'task_fMRI_spec': 0.5,'task_fMRI_neu': 0.5,'fMRI': 0.5, 'MRI': 0.5,
  293. # targets: only one column. so if the value is missing, the subject will be excluded for that regression analysis.
  294. 'Disease_unspecific_severity': 0.5, 'Weight_history_severity': 0.5,'Eating_unspecific_severity': 0.5,
  295. 'Eating_specific_severity': 0.5,'Binge_specific_severity':0.5}
  296. # features
  297. zthresholds = {'bh_neu': 10,'bh_spec': 10, 'blood': 10,'task_fMRI_spec': 10,'task_fMRI_neu': 10,'fMRI': 10, 'MRI': 10,
  298. # targets: Dont exclude anybody based on z scores here
  299. 'Disease_unspecific_severity': None, 'Weight_history_severity': None,'Eating_unspecific_severity': None,
  300. 'Eating_specific_severity': None,'Binge_specific_severity':None}
  301. subs2rem = ['K020SJ','K030CN','K031KB','K089KP','K091AH','P010SV','P011GR','P022KB','P025JB']
  302. # create a z scored version of that dataframe. Use only float columns
  303. zscore = pd.DataFrame(np.abs(stats.zscore(df_all.select_dtypes(include=["float"]),nan_policy='omit')),columns=df_all.columns,index=df_all.index)
  304. # Remove pre-selected subjects from both zscore and df_all
  305. zscore2 = zscore.drop(index=subs2rem, errors='ignore')
  306. df_all2 = df_all.drop(index=subs2rem, errors='ignore')
  307. allMissingData = pd.DataFrame(index = df_all2.index) # collect all information here
  308. missing_data_list = [] # Prepare a list to collect the missing data for each category combination
  309. # We check missing and outlier data only for features and targets
  310. for modality, features in getCatDataMapping(my_dict=data_cats).items():
  311. # Pre-compute thresholds to avoid repetitive dictionary lookups
  312. z_thresh = zthresholds[modality]
  313. missing_thresh = missing_cutoff[modality]
  314. # Vectorized calculations for the missing data and exclusion flags
  315. if z_thresh is not None:
  316. zscore_above_thresh = zscore2[features].gt(z_thresh).sum(axis=1)
  317. else: # Don't apply any z-threshold filtering
  318. zscore_above_thresh = pd.Series(0, index=zscore2.index)
  319. missing_features = zscore2[features].isnull().sum(axis=1)
  320. total_features = len(features)
  321. missing_total = (zscore_above_thresh + missing_features) / total_features
  322. exclusion_flag = (missing_total > missing_thresh)
  323. # Create a DataFrame for the current category and append to the list
  324. missing_data = pd.DataFrame({
  325. f'zthresh_{z_thresh}_{modality}': zscore_above_thresh,
  326. f'missing_{modality}': missing_features,
  327. f'ntotal_{modality}': total_features,
  328. f'missingtotal_{modality}': missing_total,
  329. f'exclusion_{modality}': exclusion_flag
  330. })
  331. missing_data_list.append(missing_data)
  332. # Concatenate all the missing data into a single DataFrame
  333. allMissingData = pd.concat(missing_data_list, axis=1)
  334. # %% [markdown]
  335. # <p style="color: #37474f; font-size: 15px;">
  336. # (b) replace outlier data with NaNs
  337. # </p>
  338. # %%
  339. # remove subject data that was flagged for exclusion (we do this individually for each category so subjects may remain in the dataframe)
  340. df_all3 = df_all2.copy() # removed subject data will be replaced with NaN and then excluded when we split the frame
  341. for modality, features in getCatDataMapping(my_dict=data_cats).items():
  342. missing_thresh = missing_cutoff[modality]
  343. curr_cols = allMissingData.filter(like = 'exclusion_'+modality) # list of flagged subjects
  344. if len(curr_cols) > 1: # if we have more than one exclusion flag per category, take the mean
  345. curr_flag = curr_cols.mean(axis=1) > missing_thresh
  346. else:
  347. curr_flag = curr_cols
  348. printmd('**' +modality+':** ' + str(len(curr_flag.loc[curr_flag==True])) +' subjects were excluded') # number of subjects to be exclusion
  349. printmd(str(curr_flag.loc[curr_flag==True])) # number of subjects to be exclusion
  350. printmd(str(len(curr_flag.loc[curr_flag==False])) +' subjects were used') # number of subjects to be exclusion
  351. # mark all flagged subject data as nan. This is done per datatype because subjects may have enough good data
  352. # of some datatypes and not enough of the others. So only the data of bad datatypes are removed but the subject stays in the analysis for now
  353. for val in features:
  354. df_all3[val]=np.where(curr_flag==True, np.nan, df_all3[val])
  355. # For resting-state data, do the same with the full matrix used for analyses
  356. if modality == 'fMRI':
  357. flagged_subs = curr_flag.loc[curr_flag==True].index.tolist() # if there is any data for these subjects, set it to NaN
  358. for key in flagged_subs:
  359. if key in rs_matrix:
  360. rs_matrix[key][:] = np.nan # Sets the existing array's values to NaN
  361. for group_name, sub_df in df_all3.groupby('group'):
  362. print(f"\nGroup: {groupNames_rev[group_name]}")
  363. print(f"Count: {len(sub_df)}")
  364. print("Subjects (index):")
  365. print(sub_df.index.tolist())
  366. print ('-----')
  367. # %% [markdown]
  368. # <p style="color: #37474f; font-size: 15px;">
  369. # c) manual data exclusion
  370. # </p>
  371. # %%
  372. df_all3.loc[df_all3['glucose'] < 60, 'glucose'] = np.nan # exclude glucose values below 60 (4 subjects)
  373. df_all3.loc[df_all3['triglyceride'] < 10, ['cholesterin','triglyceride']] = np.nan # exclude trygliceride values below 10 (3 subjects). Remove the same subjects for extremely low cholesterine values
  374. df_all3.loc['K098EW', data_cats['features']['blood']] = np.nan # this subjects has more than half of the blood values missing now, so we remove his blood data completely
  375. # check individual subjects
  376. #df_all3.loc['P104PN', data_cats['features']['blood']]
  377. #df_all3.groupby(['group','sex'])['sex'].describe()
  378. # %% [markdown]
  379. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.3 Save Preprocessed Data</h4>
  380. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  381. # %%
  382. saveFile = False
  383. # %%
  384. if saveFile:
  385. print ('saving preprocessed files...')
  386. # create an overview over the exclusions
  387. # Start with subject index and group
  388. summary_df = df_all[['group']].copy()
  389. summary_df = summary_df.loc[summary_df.index.union(subs2rem)] # Ensure pre-excluded are included
  390. # Add group name
  391. summary_df['group_name'] = summary_df['group'].map({v: k for k, v in groupNames.items()})
  392. # Add pre-exclusion flag
  393. summary_df['pre_excluded'] = summary_df.index.isin(subs2rem)
  394. # Add all exclusion flags from the main missing data check
  395. exclusion_cols = [col for col in allMissingData.columns if col.startswith('exclusion_')]
  396. summary_df = summary_df.join(allMissingData[exclusion_cols], how='left')
  397. # Fill NA exclusion flags with False for subjects not evaluated (e.g., pre-excluded)
  398. pd.set_option('future.no_silent_downcasting', True)
  399. summary_df[exclusion_cols] = summary_df[exclusion_cols].apply(lambda col: col.fillna(False).astype(bool))
  400. # (1) data for ML analyses
  401. df_all3.drop(columns=df_all3.filter(like='_RS_').columns).to_pickle(f'{out_dir_ML}/allData_precleaned_step2.2c_Final.pkl') # with outliers replaced by nans + rs columns dropped
  402. with open(f'{out_dir_ML}/data_cats.pkl', 'wb') as f:
  403. pickle.dump(data_cats, f)
  404. with open(f'{out_dir_ML}/AAL3Matrix_precleaned_step2.2c_Final.pkl', 'wb') as f:
  405. pickle.dump(rs_matrix, f)
  406. # (2) additional data safeties
  407. allMissingData.to_csv(f'{out_dir_expl}/data_cleaning_summary.csv', index=True)
  408. allMissingData.to_excel(f'{out_dir_expl}/data_cleaning_summary.xlsx', index=True)
  409. summary_df.to_excel(f'{out_dir_expl}/subject_exclusion_summary.xlsx', index=True)
  410. vars_all.to_pickle(f'{out_dir_saf}/allVars_precleaned_step1.pkl') # variable list without empty columns
  411. zscore.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2a_Z.pkl') # zscored version of df_all
  412. zscore2.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2b_Z.pkl') # with pre-selected subs removed
  413. df_all.to_pickle(f'{out_dir_saf}/allData_precleaned_step1.pkl') # data without empty columns
  414. df_all.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.1a.pkl') # added severity indices
  415. df_all2.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2b.pkl') # with pre-selected subs removed
  416. df_all3.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2c_Final.pkl') # with outliers replaced by nans
  417. # save both data categorizations
  418. with open(f'{out_dir_saf}/data_cats.pkl', 'wb') as f:
  419. pickle.dump(data_cats, f)
  420. with open(f'{out_dir_saf}/data_cats_full.pkl', 'wb') as f:
  421. pickle.dump(data_cats_full, f)
  422. print ('Done!')
  423. else:
  424. print ('Not saving files due to saveFile=False')
  425. # %% [markdown]
  426. # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">3 DATA EXPLORATION</h3>
  427. # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
  428. # %% [markdown]
  429. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.1 Raw Data Inspection</h4>
  430. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  431. # %% [markdown]
  432. # <p style="color: #37474f; font-size: 15px;"> For each modality, outliers were examined independently, regardless of how the data were grouped for later analysis. Outliers were also checked visually. In the behavioral data, outliers in error scores appeared as sharp peaks, usually when subjects made no errors. </p>
  433. # %%
  434. saveFile = True#saveFiles # use global parameter
  435. z_score_out = Path(out_dir_expl).parent.joinpath('z_thresholds')
  436. z_score_out.mkdir(exist_ok=True)
  437. # %%
  438. NonMRICats = [i for i in data_cats_full['cat1'] if 'MRI' not in i]
  439. for cat in NonMRICats[:]:
  440. curr_cols = vars_all.loc[vars_all['category1']==cat]['feature'].tolist()
  441. curr_data = df_all2[curr_cols]
  442. #print (curr_data.describe())
  443. printmd("**------------------------------ "+cat+" ------------------------------**")
  444. if saveFile:
  445. zscored = convert_display_zScore(curr_data,zthresh=3,saveStats=str(z_score_out)+'/BH_threshs'+cat+'.csv')
  446. else:
  447. zscored = convert_display_zScore(curr_data,zthresh=3,saveStats=None)
  448. # %%
  449. curr_cols = vars_all.loc[vars_all['category1']=='sMRI']['feature'].tolist()
  450. curr_data = df_all[curr_cols]
  451. printmd("**------------------------------ VBM ------------------------------**")
  452. #print (curr_data.describe()[:5])
  453. if saveFile:
  454. zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=str(z_score_out)+'/VBM_threshs.csv')
  455. else:
  456. zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=None)
  457. # %%
  458. TaskfMRICats = vars_all.loc[(vars_all['category1']=='fMRI')&(vars_all['category2'].isin(['MID', 'GNG']))]['category2'].unique()
  459. for cat in TaskfMRICats:
  460. curr_cols = vars_all.loc[(vars_all['category1']=='fMRI')&(vars_all['category2'].isin([cat]))]['feature'].tolist()
  461. curr_data = df_all[curr_cols]
  462. printmd("**------------------------------ "+cat+" ------------------------------**")
  463. #print (curr_data.describe())
  464. if saveFile:
  465. zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=str(z_score_out)+'/TaskfMRI_threshs'+cat+'.csv')
  466. else:
  467. zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=None)
  468. # %%
  469. curr_cols = vars_all.loc[vars_all['category2']=='RS']['feature'].tolist()
  470. curr_data = df_all[curr_cols]
  471. printmd("**------------------------------ Resting-State functional connectivity ------------------------------**")
  472. #print (curr_data.describe()[:5])
  473. if saveFile:
  474. zscored = convert_display_zScore(curr_data.iloc[:, :10],zthresh=3,saveStats=str(z_score_out)+'/RS_threshs.csv')
  475. else:
  476. zscored = convert_display_zScore(curr_data.iloc[:, :10],zthresh=3,saveStats=None)
  477. # %% [markdown]
  478. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2 Final Sample Inspection</h4>
  479. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  480. # %% [markdown]
  481. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2.1 Features</h4>
  482. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  483. # %% [markdown]
  484. # <p style="color: #37474f; font-size: 15px;"> We check the intercorrelation between the features of each modality for the entire sample. The code can be easily adjusted to check just for specific groups or across modalities </p>
  485. # %%
  486. saveFile = False#saveFiles # use global parameter
  487. corr_out = Path(out_dir_expl).parent.joinpath('correlations')
  488. corr_out.mkdir(exist_ok=True)
  489. # %% [markdown]
  490. # (a) Correlations
  491. # %%
  492. group_combs = {'BED':[1.0],'CON_BED':[3.0],'BN':[2.0],'CON_BN':[4.0],'PAT':[1.0,2.0],
  493. 'HC':[3.0,4.0],'O_weight':[1.0,3.0],'N_weight':[2.0,4.0],'All':[1.0,2.0,3.0,4.0]}
  494. group_combs = {'All':[1.0,2.0,3.0,4.0]} # do all just for now
  495. for curr_feat_name,curr_cols in data_cats['features'].items():
  496. if curr_feat_name == 'fMRI': # no need to run this for RS data because this is already a correlation matrix
  497. continue
  498. print (curr_feat_name)
  499. for group_name,group_idx in group_combs.items():
  500. print (group_name)
  501. curr_data = df_all3.loc[df_all3['group'].isin(group_idx)][curr_cols].dropna(axis = 0, how = 'all')
  502. # We save these as a file, so we dont have to run it again. We can plot it with the fromFile key
  503. # dont show table if the category has many features
  504. #printmd("**------------------------------" + curr_feat_name + " features Correlation Matrix ------------------------------**")
  505. showtable = False
  506. rho,pval,ax = corrMatrix(curr_data,table=showtable,plot=False,cmap="RdBu_r",
  507. saveFile=f'{str(corr_out)}/{curr_feat_name}_{group_name}')
  508. # %% [markdown]
  509. # <p style="color: #37474f; font-size: 15px;"> Here we check the level of significant corrations within feature modalities. </p>
  510. # %%
  511. for curr_mod_name,curr_cols in data_cats['features'].items():
  512. curr_data = df_all3[curr_cols].dropna(axis = 0, how = 'all')
  513. printmd(f"**------------------------------ {curr_mod_name} features Intercorrelations ------------------------------**")
  514. # Skip fMRI category for now
  515. if 'fMRI' not in curr_mod_name:
  516. # Get correlation matrix and p-values
  517. rho, pval, ax = corrMatrix(curr_data, table=False, plot=False, cmap="RdBu_r",
  518. saveFile = f'{str(corr_out)}/{curr_mod_name}_{group_name}',
  519. fromFile=True)
  520. # Extract significant correlations
  521. sign_pairs = [
  522. (col, i) for col in pval.columns for i in pval[pval[col] < 0.05].index if i != col
  523. ]
  524. sign_pairs = list(set(tuple(sorted(pair)) for pair in sign_pairs)) # Remove duplicates
  525. # Get all combinations
  526. combos = list(combinations(pval.columns, 2))
  527. print(f"{len(sign_pairs)} significant correlations out of {len(combos)} feature combinations")
  528. # Plot some example pairs
  529. col_ranges = list(range(0, len(sign_pairs), 6))[:2]
  530. for i in range(len(col_ranges) - 1):
  531. curr_pairs = sign_pairs[col_ranges[i]:col_ranges[i + 1]]
  532. # Create subplots with a 1x6 grid
  533. fig, ax = plt.subplots(1, 6, figsize=(11, 2))
  534. for pair_n, (x_col, y_col) in enumerate(curr_pairs):
  535. sns.regplot(
  536. x=x_col, y=y_col, ax=ax[pair_n], data=curr_data,
  537. scatter_kws={"color": "darkred", "alpha": 0.3, "s": 20},
  538. line_kws={"color": "k", "alpha": 0.8, "lw": 2}
  539. )
  540. plt.tight_layout()
  541. plt.show()
  542. plt.clf()
  543. plt.close()
  544. break # remove to plot everything
  545. # %% [markdown]
  546. # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2.1 Targets</h4>
  547. # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
  548. # %% [markdown]
  549. # <p style="color: #37474f; font-size: 15px;"> Here we check the level of significant corrations within feature modalities. </p>
  550. # %%
  551. #for group in range(1,5):
  552. curr_data = df_all3[data_cats['targets'].keys()]#.loc[df_reduced['group']==group]
  553. rho,pval,ax = corrMatrix(curr_data,table=True,plot=True)
  554. # %%
  555. from scipy.stats import f_oneway
  556. from statsmodels.stats.multicomp import pairwise_tukeyhsd
  557. from statannotations.Annotator import Annotator
  558. from statsmodels.formula.api import ols
  559. import statsmodels.api as sm
  560. plotting_df = df_all3.copy()
  561. plotting_df['named_group'] = plotting_df['group'].map(groupNames_rev)
  562. # Prepare long-form dataframe for two-way ANOVA
  563. severity_features = data_cats['targets'].keys()
  564. df_long = plotting_df.melt(id_vars=['named_group'], value_vars=severity_features,
  565. var_name='severity_index', value_name='value').dropna()
  566. # Run two-way ANOVA with interaction
  567. model = ols('value ~ C(named_group) * C(severity_index)', data=df_long).fit()
  568. anova_table = sm.stats.anova_lm(model, typ=2)
  569. print("\n Two-way ANOVA Results (Group x Severity Index):\n")
  570. print(anova_table)
  571. annotate = True # plot test significance
  572. for idx, cols in severity_idx.items():
  573. if annotate:
  574. plt.figure(figsize=(4, 6))
  575. else:
  576. plt.figure(figsize=(4, 4))
  577. cols = default_cols(n=2, random=True, n_set=None)
  578. print(cols)
  579. feature = idx + '_severity'
  580. df_plot = plotting_df.dropna(subset=[feature])
  581. # Violin plot
  582. sns.violinplot(
  583. data=df_plot,
  584. x='named_group',
  585. y=feature,
  586. inner='box',
  587. color='white',
  588. edgecolor='black'
  589. )
  590. # Swarm plot
  591. ax = sns.swarmplot(
  592. data=df_plot,
  593. x='named_group',
  594. y=feature,
  595. hue='named_group',
  596. alpha=0.9,
  597. palette=2 * cols,
  598. dodge=False,
  599. s=4.5,
  600. )
  601. # ANOVA
  602. groups = [group_df[feature].values for name, group_df in df_plot.groupby('named_group')]
  603. stat, p_anova = f_oneway(*groups)
  604. plt.title(f"{idx.replace('_', ' ').title()} Severity by Group\nANOVA p={p_anova:.3e}")
  605. # Post-hoc Tukey HSD if ANOVA is significant
  606. if p_anova < 0.05:
  607. tukey = pairwise_tukeyhsd(endog=df_plot[feature], groups=df_plot['named_group'], alpha=0.05)
  608. results_df = pd.DataFrame(data=tukey._results_table.data[1:], columns=tukey._results_table.data[0])
  609. # Filter for significant comparisons (p <= 0.05)
  610. sig_results = results_df[results_df['p-adj'] <= 0.05]
  611. # If any significant pairs exist, annotate them
  612. if not sig_results.empty:
  613. pairs = [(row['group1'], row['group2']) for _, row in sig_results.iterrows()]
  614. pvals = [row['p-adj'] for _, row in sig_results.iterrows()]
  615. if annotate:
  616. annotator = Annotator(ax, pairs, data=df_plot, x='named_group', y=feature)
  617. annotator.set_pvalues_and_annotate(pvalues=pvals)
  618. plt.xlabel("Group")
  619. plt.ylabel(idx)
  620. plt.legend([], [], frameon=False) # Remove redundant legend
  621. plt.tight_layout()
  622. plt.show()
  623. # %%

NeuroBED_ML_DataExploration.ipynb at commit c5e2120, under MIT · at the source

Overview

Authors: Lena Rommerskirchen1, Mandy Skunde1,2, Martin Bendszus3, Wolfgang Herzog1, Hans-Christoph Friederich1,4, Joe J Simon1,4
  1. Department of General Internal Medicine, Psychosomatics and Psychotherapy, Centre for Psychosocial Medicine, University Hospital Heidelberg, Heidelberg, Germany
  2. Institute of Pathology, University Hospital Heidelberg, Heidelberg, Germany
  3. Department of Neuroradiology, University Hospital Heidelberg, Heidelberg, Germany
  4. DZPG (German Centre for Mental Health – Partner Site Heidelberg/Mannheim/Ulm), Germany
Institutions: Heidelberg University (Germany); University Hospital Heidelberg (Germany)
Journal: Frontiers in neuroscience, volume 20, article 1803154
Dates: received 3 February 2026; accepted 20 March 2026; published online 13 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fnins.2026.1803154 · PMID 42051561 · PMCID PMC13111282 · OpenAlex W7154017790
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: fMRI (modality), other condition (population)
Methods: Spectral & time-frequency, Statistics, Smoothing, state filtering, decompositions, Machine learning, fMRI & imaging
Keywords: binge eating disorder (BED), bulimia nervosa (BN), fMRI, functional connectivity, machine learning (ML)
Topic: Eating Disorders and Behaviors (Clinical Psychology, Psychology), according to OpenAlex
Citations: not cited yet (Europe PMC); 96 references in the paper

Abstract

Introduction: Binge-type eating disorders, including bulimia nervosa (BN) and binge eating disorder (BED), are associated with both shared and disorder-specific neurobiological mechanisms across brain, behavior, and physiology. A clearer distinction between shared mechanisms and disorder-specific alterations may advance our understanding of binge-type eating pathology.

Methods: We applied a comprehensive multimodal machine learning framework to 110 participants (BN, BED, and age & weight matched controls), integrating task-based fMRI, intrinsic connectivity, voxel-based morphometry, neuropsychological assessments, and peripheral blood biomarkers. Both unimodal and multimodal machine learning models were trained to classify groups and to predict individual variation in symptom expression.

Results: Functional brain connectivity achieved the highest accuracy for diagnostic classification and symptom prediction (with a mean balanced classification accuracy (bACC) of 68.7%), whereas task-based fMRI with disorder-specific food stimuli and peripheral blood biomarkers best distinguished BN from BED (mean bACC of 87%). Multimodal models did not generally outperform the best unimodal approaches, except from modest gains in a limited set of regression targets.

Conclusions: These findings suggest that functional brain connectivity carries robust predictive information for transdiagnostic classification, whereas task-evoked activation patterns and peripheral biomarkers show stronger predictive utility for distinguishing BN from BED. Whether these modality-specific patterns reflect underlying neurobiological mechanisms remains to be established in future hypothesis-driven work. Identifying which modalities best represent shared vulnerability vs. symptom-type-dependent variation may help to provide a foundation for a more mechanistic understanding of these disorders.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

LenaRo09/NeuroBED_ML

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: c5e21208047398a4e06f72827bd96e102792448f, 26 March 2026
Languages: Python (33), Shell (3), Jupyter (1)
Size: 282 files, 37 scripts
Software Heritage: not archived
Found in: “Data availability statement”
Holds: README, license file, 1 notebook
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (31 files), pandas (30 files), Matplotlib (20 files), scikit-learn (8 files), SciPy (6 files), seaborn (6 files), statannotations (2 files), statsmodels (2 files), NiBabel (1 file), SHAP (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
39 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 37 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The datasets presented in this study can be found invonline repositories. The names of the repository/repositoriesvand accession number(s) can be found at: https://github.com/LenaRo09/NeuroBED_ML.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 5 keywords, 1 funder, 91 references.

Cite

This paper

Rommerskirchen, L., Skunde, M., Bendszus, M., Herzog, W., Friederich, H.-C., & Simon, J. J. (2026). Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders. Frontiers in neuroscience, 20, 1803154. https://doi.org/10.3389/fnins.2026.1803154

BibTeX

@article{rommerskirchen2026multimodal,
author = {Rommerskirchen, Lena and Skunde, Mandy and Bendszus, Martin and Herzog, Wolfgang and Friederich, Hans-Christoph and Simon, Joe J},
title = {{Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders}},
journal = {Frontiers in neuroscience},
year = {2026},
month = apr,
volume = {20},
pages = {1803154},
publisher = {Frontiers Media SA},
issn = {1662-4548},
doi = {10.3389/fnins.2026.1803154},
url = {https://doi.org/10.3389/fnins.2026.1803154},
pmid = {42051561},
pmcid = {PMC13111282}
}

RIS

TY - JOUR
AU - Rommerskirchen, Lena
AU - Skunde, Mandy
AU - Bendszus, Martin
AU - Herzog, Wolfgang
AU - Friederich, Hans-Christoph
AU - Simon, Joe J
TI - Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders
T2 - Frontiers in neuroscience
J2 - Front Neurosci
PY - 2026
DA - 2026/04/13
VL - 20
SP - 1803154
SN - 1662-4548
PB - Frontiers Media SA
DO - 10.3389/fnins.2026.1803154
UR - https://doi.org/10.3389/fnins.2026.1803154
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fnins.2026.1803154",
"type": "article-journal",
"title": "Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders",
"container-title": "Frontiers in neuroscience",
"author": [
{
"family": "Rommerskirchen",
"given": "Lena"
},
{
"family": "Skunde",
"given": "Mandy"
},
{
"family": "Bendszus",
"given": "Martin"
},
{
"family": "Herzog",
"given": "Wolfgang"
},
{
"family": "Friederich",
"given": "Hans-Christoph"
},
{
"family": "Simon",
"given": "Joe J"
}
],
"container-title-short": "Front Neurosci",
"volume": "20",
"page": "1803154",
"DOI": "10.3389/fnins.2026.1803154",
"PMID": "42051561",
"PMCID": "PMC13111282",
"ISSN": "1662-4548",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fnins.2026.1803154",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
13
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pmed.1004809 [code]
Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data.
Journal: PLoS medicine
In common: statannotations, NiBabel, statsmodels, 6 other tools, other condition, 3 references
[2] doi:10.1038/s43856-026-01518-5 [code]
Machine learning-based identification of abnormal functional connectivity in obesity across different metabolic states.
Journal: Communications medicine
In common: statannotations, XGBoost, statsmodels, 6 other tools, other condition, 2 references
[3] doi:10.1016/j.isci.2026.115329 [code]
Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.
Journal: iScience
In common: statannotations, SHAP, XGBoost, 8 other tools, other condition
[4] doi:10.21203/rs.3.rs-9914920/v1 [code]
Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group
Journal: Research Square (preprint)
In common: statannotations, SHAP, XGBoost, 8 other tools
[5] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: SHAP, XGBoost, NiBabel, 7 other tools, fMRI, other condition
[6] doi:10.1038/s41598-026-56688-y [code]
On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.
Journal: Scientific reports
In common: SHAP, XGBoost, NiBabel, 7 other tools, fMRI
[7] doi:10.1002/hbm.70469 [code]
VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.
Journal: Human brain mapping
In common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
[8] doi:10.21203/rs.3.rs-9326213/v1 [code]
Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brain
Journal: Research Square (preprint)
In common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
[9] doi:10.64898/2026.03.09.710558 [code]
Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brain
Journal: bioRxiv (preprint)
In common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
[10] doi:10.1038/s42003-026-10276-y [code]
The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.
Journal: Communications biology
In common: statannotations, NiBabel, statsmodels, 6 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.