Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders.
The 16 matches
- [1] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/01 - model0_RS/model0.py, lines 1–59 · score 0.84 · feature vectors, retained variance, cross validation, imputed, median, bounded
- [2] § Materials and methods › Feature set construction ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 273–293 · score 0.83 · Stop Signal, sMRI, working memory, fMRI, Cued, RSS
- [3] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 1–58 · score 0.82 · retained variance, cross validation, rsfMRI, chosen, imputed, median
- [4] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/D_confusion_matrix_grid.py, lines 1–35 · score 0.76 · Row normalized confusion, best multimodal model, bACC, single modality models, best single modality, matrices
- [5] § Materials and methods › ML-framework ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 1–58 · score 0.73 · machine learning, structural MRI, rsfMRI, single modalities, SVMs, dummy
- [6] § Materials and methods › ML-framework ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 1–67 · score 0.71 · machine learning, structural MRI, rsfMRI, single modalities, SVMs, dummy
- [7] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 295–319 · score 0.67 · general eating behaviors, Eating unspecific, equal weights, DEBQ, FCQ, subscales
- [8] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 295–319 · score 0.66 · general depressive symptoms, Disease unspecific, weight fluctuations, BDI, Composite, domains
- [9] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/A_plot_heatmap_target_modality.py, lines 1–27 · score 0.65 · cross validated balanced, rsfMRI, balanced accuracy, Heatmap, intrinsic, bh
- [10] § Results › Unimodal classification ↔ NeuroBED_ML/analysis/06 - plotting/FIG 02 - ClassResults/scripts/D_confusion_matrix_grid.py, lines 1–35 · score 0.64 · confusion matrices, bACC, best single modality, classes, classification, Figure 2
- [11] § Materials and methods › Targets ↔ NeuroBED_ML/analysis/06 - plotting/FIG 04 - RegrResults/scripts/C_regr_univariate_results_targetXgroup.py, lines 55–74 · score 0.64 · Disease unspecific, owHC, nwHC, BDI, severity, scores
- [12] § Materials and methods › Model evaluation ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 1–67 · score 0.63 · dummy baseline, fitting, fusion, concatenated, single modality, component
- [13] § Materials and methods › Participants ↔ NeuroBED_ML/analysis/02 - data_exploration/NeuroBED_ML_DataExploration.ipynb, lines 273–293 · score 0.60 · working memory, fMRI, age, sex, overweight, BMI
- [14] § Results › Unimodal regression ↔ NeuroBED_ML/analysis/04 - models/03 - model2_combinedModalities/model2.py, lines 109–134 · score 0.53 · Eating unspecific, regression targets, signal, R2, variance, severity
- [15] § Results › Multimodal regression ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 125–144 · score 0.52 · Eating unspecific severity, disease unspecific, Single modality, R2, accuracy, weight
- [16] § Results › Unimodal regression ↔ NeuroBED_ML/analysis/04 - models/02 - model1_singleModality/model1.py, lines 125–144 · score 0.52 · Eating unspecific, regression targets, signal, R2, variance, severity
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 776 lines · 41 KB · MIT · 4 matches
- # %% [markdown]
- # <h1 style="color: #1e88e5; font-weight: bold; margin-bottom: 5px;">NeuroBED_ML: DATA EXPLORATION NOTEBOOK</h1>
- # <hr style="border: 2px solid #cfd8dc; margin-top: 0; margin-bottom: 20px;">
- # %% [markdown]
- # <div style="color: #37474f; font-size: 16px;">
- # <p>
- # This notebook performs data exploration and additional preprocessing of the NeuroBED dataset in preparation for machine learning using the
- # <code style="font-size: 16px;">julearn</code> framework. At this point, the variables used for analyses have already been extracted from the raw data files and the data underwent minimal preprocessing (renaming variables, uniform naming convention for missing data etc.). Steps in this notebook include loading and cleaning data, handling missing values, removing outliers, and exporting structured outputs for modeling. The data will be written to file after each preprocessing step and stored here:
- # <code style="font-size: 16px;">/02 - data_exploration/output/</code>.
- # </p>
- #
- # <p>
- # The files used for analysis are additionally saved here:
- # <code style="font-size: 16px;">/03 - data4ML</code>.
- # </p>
- #
- # <p>
- # <code style="font-size: 16px;">allData_precleaned_step2.2c_Final.pkl</code>: All data<br>
- # <code style="font-size: 16px;">data_cats.pkl</code>: Info about category-variable mapping.
- # </p>
- #
- # <p style="font-size: 16px; margin-top: 1em;">
- # <strong>Please note that some variables have been renamed in the paper for clarity:</strong><br><br>
- # <span style="font-size: 16px;">fMRI = rsfMRI</span><br>
- # <span style="font-size: 16px;">MRI = GMV</span><br>
- # <span style="font-size: 16px;">CON_BED = owHC</span><br>
- # <span style="font-size: 16px;">CON_BN = nwHC</span><br>
- # <span style="font-size: 16px;">Weight history = weight fluctuations and monitoring</span>
- # </p>
- # </div>
- # %% [markdown]
- # <div style="color: #37474f; border-left: 5px solid #607d8b; background-color: #f5f5f5; padding: 15px; margin-bottom: 20px; border-radius: 5px;">
- # <h3><strong style="color: #1e88e5;">Table of Contents</strong></h3>
- # <ul>
- # <li><a href="#section-0" style="color: #37474f;"><strong>0 Notebook Styling Params</strong></a></li>
- # <li><a href="#section-1" style="color: #37474f;"><strong>1 Imports</strong></a>
- # <ul>
- # <li><a href="#section-1-1" style="color: #37474f;">1.1 Variables</a></li>
- # <li><a href="#section-1-2" style="color: #37474f;">1.2 Load Data</a></li>
- # </ul>
- # </li>
- # <li><a href="#section-2" style="color: #37474f;"><strong>2 Preprocessing</strong></a>
- # <ul>
- # <li><a href="#section-2-1" style="color: #37474f;">2.1 Data Categorization</a></li>
- # <ul>
- # <li><a href="#section-2-1-1" style="color: #37474f;">2.1.1 Full Categorization</a></li>
- # <li><a href="#section-2-1-2" style="color: #37474f;">2.1.2 Model Categorization</a></li>
- # <ul>
- # <li><a href="#section-2-1-2.1" style="color: #37474f;">2.1.2.1 Features</a></li>
- # <li><a href="#section-2-1-2.2" style="color: #37474f;">2.1.2.2 Targets</a></li>
- # <li><a href="#section-2-1-2.3" style="color: #37474f;">2.1.2.2 Confounders & Others</a></li>
- # </ul>
- # </ul>
- # <li><a href="#section-2-2" style="color: #37474f;">2.2 Outliers and Missing Data</a></li>
- # <li><a href="#section-2-3" style="color: #37474f;">2.3 Save Preprocessed Data</a></li>
- # </ul>
- # </li>
- # <li><a href="#section-3" style="color: #37474f;"><strong>4 Data Exploration</strong></a>
- # <ul>
- # <li><a href="#section-3-1" style="color: #37474f;">4.1 Raw Data Inspection</a></li>
- # <li><a href="#section-3-2" style="color: #37474f;">4.2 Final Sample Inspection</a></li>
- # <ul>
- # <li><a href="#section-3-2" style="color: #37474f;">4.2.1 Features</a></li>
- # <li><a href="#section-3-2" style="color: #37474f;">4.2.2 Targets</a></li>
- # </ul>
- # </ul>
- # </li>
- # </ul>
- # </div>
- # %% [markdown]
- # | Step | Description | Object(s) | Input File(s) | Output File(s) |
- # |------|-------------|-----------|----------------|-----------------|
- # | 1 | Load input data (precleaned, selected variables) | `df_all`, `vars_all`, `rs_matrix` | `allData_precleaned.pkl`, `allVars_precleaned.pkl`, `NeuroBED_ML_AAL3Matrix_precleaned.pkl` | – |
- # | 2 | Convert all data columns to float, drop empty ones | `df_all`, `dropped_cols`, `colsb4` | – | `allData_precleaned_step1.pkl`,` allVars_precleaned_step1.pkl` |
- # | 2.1a | Generate nested modality-feature mapping | `data_cats_full` | `NeuroBED_ML_Vars.xlsx` | – |
- # | 2.1b | Generate final analysis categorization | `data_cats` | `data_cats_full` | - |
- # | 2.1c | Add severity indices | `df_all` | – | `allData_precleaned_step2.1a.pkl`, `data_cats.pkl`,`data_cats_full.pkl` |
- # | 2.2a | Calculate outliers/remove preselected subs | `zscore`, `zscore2`, `df_all2`,`allMissingData` | `df_all` | - |
- # | 2.2b | Replace outlier data with NaN| `df_all3`,`rs_matrix`| `df_all2`,`allMissingData` | - |
- # | 2.2c | Export Data and Cleaning summaries | `summary_df`,`allMissingData`,`df_all3`,`zscore`, `zscore2`,`rs_matrix` | – | `allData_precleaned_step2.2a_Z.pkl`, `allData_precleaned_step2.2b_Z.pkl`, `allData_precleaned_step2.2b.pkl`, `allData_precleaned_step2.2c_Final.pkl`, `subject_exclusion_summary.xlsx`, `data_cleaning_summary.csv`, `data_cleaning_summary.xlsx`, `AAL3Matrix_precleaned_step2.2c_Final.pkl`|
- # %% [markdown]
- # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">0. NOTEBOOK STYLING PARAMS</h3>
- # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
- # %%
- # Printing style
- from IPython.display import Markdown, display, Image,HTML
- from pprint import pprint
- def printmd(string):
- display(Markdown(string))
- # %% [markdown]
- # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">1. IMPORTS</h3>
- # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
- # %%
- import os, sys, random
- from pathlib import Path
- basepath = f"{os.getcwd()[:2]}/NeuroBed_ML/analysis/"
- if not basepath in sys.path:
- sys.path.insert(0,basepath) # to get code from other files #"E:/NeuroBed_ML/analysis/"
- sys.path.insert(1, f"{basepath}05 - utils/") # to get code from other files
- from NeuroBED_ML_helpers import *
- from LR_Cmaps import *
- import pandas as pd
- import numpy as np
- from scipy import stats
- from scipy.stats import pearsonr
- import matplotlib.pyplot as plt
- from itertools import combinations
- import matplotlib_inline
- import seaborn as sns
- import pickle,time
- from datetime import datetime
- # plotting parameters
- try:
- plt.style.use(f'{basepath}05 - utils/LR_ST2.mplstyle')
- matplotlib_inline.backend_inline.set_matplotlib_formats('retina')
- fonts = ['Poppins','Montserrat','Manrope','Figtree','Helvetica','Roboto','Quicksand','Arial','AvenirLTPro-Book']
- plt.rcParams['font.family'] = fonts[random.randint(0,len(fonts)-1)]
- except:
- plt.rcParams['font.family'] = 'Arial'
- print ('current font: ', plt.rcParams['font.family'][0])
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">1.1 Variables</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 16px;">
- # This section sets up the configuration, including naming conventions and I/O structure, and then loads the data:
- # </p>
- #
- # <ul style="color: #37474f; font-size: 15px;">
- # <li><strong>Input Directory:</strong> Points to extracted and minimally cleaned data under <code>/01 - data/</code>.</li>
- # <li><strong>Output Directory:</strong> Created if it doesn't exist. Located under <code>/02 - data_exploration/output/</code>.</li>
- # <li><strong>Group Mapping:</strong> Defines numerical IDs for each group:
- # <ul>
- # <li><code>BED</code> → 1</li>
- # <li><code>BN</code> → 2</li>
- # <li><code>CON_BED (owHC)</code> → 3</li>
- # <li><code>CON_BN (nwHC)</code> → 4</li>
- # </ul>
- # </li>
- # </ul>
- #
- # <p style="color: #37474f; font-size: 15px;">
- # These group IDs are used for subject exclusion, stratified cross validation and plotting.
- # </p>
- # %%
- parent_dir = basepath
- in_dir = resolve_path('data')
- out_dir_ML = resolve_path('data4ML') # for output files that are used for ML models. creates dir if it does not exist
- out_dir_saf = resolve_path('data_exploration/output/additional_data_safeties') # safety outputs. creates dir if it does not exist
- out_dir_expl = resolve_path('data_exploration/output/data_cleaning') # exploration outputs. creates dir if it does not exist
- groupNames = {'BED':1,'CON_BED':3,'BN':2,'CON_BN':4}
- groupNames_rev = { v: k for k, v in groupNames.items()}
- saveFiles = False # save all output files from this notebook. Can be manually adjusted for each section
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">1.2 Load Data</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 16px;">
- # The loaded dataset is already complete with all data types and pre-cleaned.
- # </p>
- # %%
- df_all = pd.read_pickle(f"{in_dir}/NeuroBED_ML_allData_precleaned.pkl") # import all data (except rsfMRI)
- rs_matrix = pd.read_pickle(f"{in_dir}/NeuroBED_ML_AAL3Matrix_precleaned.pkl") # import resting-state matrix (112x166x166)
- vars_all = pd.read_pickle(f"{in_dir}/NeuroBED_ML_allVars_precleaned.pkl") # variable info
- # %% [markdown]
- # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">2. PREPROCESSING</h3>
- # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # Before the features are categorized, empty features are removed from the dataset. The remaining feature values are converted to float for compatibility with the ML pipeline. The result is saved here:
- # <code>allVars_precleaned_remEmptyCols.pkl/xlsx</code>
- # </p>
- # %%
- saveFile = saveFiles # use global parameter
- # %%
- # convert all data to float
- df_all = df_all.apply(pd.to_numeric, errors='coerce')
- # remove empty features. There are some MRI features where this happend
- colsb4 = list(df_all.columns)
- df_all.dropna(axis=1, how='all',inplace=True)
- dropped_cols = list(set(colsb4) - set(list(df_all.columns)))
- print ('dropped ',str(len(dropped_cols)), ' feature columns from the dataframe')
- print (dropped_cols)
- # remove these columns from the variable list
- vars_all = vars_all[~vars_all['feature'].isin(dropped_cols)]
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.1 Data Categorization</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # This section lists all variables (features), organized by modality and stimulus type, as defined in
- # <code>data/variable_info/NeuroBED_ML_Vars.xlsx</code> (e.g., <code>category1 → category2 → category3</code>).
- # <br><br>
- # Because there are many categories with different numbers of variables, we also combine related items into
- # clinically meaningful groups. These groups make the results easier to interpret and keep the analysis concise.
- # We save these groupings so they can be reused when training machine-learning models with different sets of variables.
- # </p>
- # %%
- saveFile = saveFiles # use global parameter
- # %% [markdown]
- # <h5 style="color: #37474f; font-weight: normal; font-size: 18px; margin-top: 30px;">2.1.1 Full Categorization</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %%
- # Data categories and combinations
- data_cats_full = {}
- data_cats_full['cat1'] = vars_all.groupby(['category1'])['feature'].apply(list).to_dict()
- data_cats_full['cat2'] = vars_all.groupby(['category2'])['feature'].apply(list).to_dict()
- data_cats_full['cat3'] = vars_all.groupby(['category3'])['feature'].apply(list).to_dict()
- data_cats_full['cat1cat2'] = vars_all.groupby(['category1','category2'])['feature'].apply(list).to_dict()
- data_cats_full['cat1cat3'] = vars_all.groupby(['category1','category3'])['feature'].apply(list).to_dict()
- data_cats_full['cat2cat3'] = vars_all.groupby(['category2','category3'])['feature'].apply(list).to_dict()
- data_cats_full['cat1cat2cat3'] = vars_all.groupby(['category1','category2','category3'])['feature'].apply(list).to_dict()
- if saveFile: # save this too for later
- with open(str(in_dir)+'data_cats_full.pkl', 'wb') as f:
- pickle.dump(data_cats_full, f)
- vars_all.groupby(['category1','category2','category3']).describe()
- # task fMRI: 136+2=138 (AAL3 minus thalamus + AAL2 thalamus)
- # VBM: 166 (AAL3 only)
- # rs: 165x166/2 = 13695 (AAL3 only minus diagonal)
- # %% [markdown]
- # <h5 style="color: #37474f; font-weight: normal; font-size: 18px; margin-top: 30px;">2.1.1 Model Categorization</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # This is the final grouping for the analyses.
- # </p>
- # %%
- data_cats = {'features': {},'targets': {},'confounders': {},'others': {}}
- # %% [markdown]
- # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.1 FEATURES</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %%
- Image(filename='./plots/FeaturesOnly.png',width=300, height=150)
- # %%
- features = {'bh_neu': [('behavioral', 'CuedTask', 'all'),('behavioral', 'GNG', 'neu'),('behavioral', 'MID', 'neu'),('behavioral', 'RSS', 'all'),
- ('behavioral', 'StopSignal', 'all'),('behavioral', 'intelligence', 'all'),('behavioral', 'working memory', 'all')],
- 'bh_spec': [('behavioral', 'GNG', 'spec'),('behavioral', 'MID', 'spec')],
- 'blood': [('blood', 'all', 'all')],
- 'task_fMRI_spec':[('fMRI', 'GNG', 'spec'),('fMRI', 'MID', 'spec')],
- 'task_fMRI_neu': [('fMRI', 'GNG', 'neu'),('fMRI', 'MID', 'neu')],
- 'fMRI': [('fMRI', 'RS', 'all')],
- 'MRI' : [('sMRI', 'struct', 'all')]}
- data_cats['confounders'] = ['sex','age','BMI']
- data_cats['others'] = ['VAS_hungr_before','VAS_mood_before','VAS_hungr_after','VAS_mood after','hightest_weight',
- 'age_highest_weight','lowest_weight','age_lowest_weight','look','bodyshape_father','bodyshape_mother',
- 'first_time_overweight']
- for mod,subkeys in features.items():
- data_cats['features'][mod] = [feat for tup in features[mod] for feat in data_cats_full['cat1cat2cat3'][tup]]
- # Display the new category - feature mapping. Limit the display to x features
- displayDictJupyter(data_cats['features'],val_lim=30)
- # %% [markdown]
- # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.2 TARGETS</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> To measure how severe the symptoms were across different areas, we created summary scores (composite indices) from selected questionnaire features. These scores are based on values that were first standardized to a common scale (ranging from 0 to 1). In some cases, features were given more or less weight to reflect their clinical importance within a domain. </p> <p style="color: #37474f; font-size: 15px;"> We evaluated five symptom areas: </p> <ol style="color: #37474f; font-size: 15px;"> <li><code><strong>Disease unspecific</strong></code> – general depressive symptoms (e.g., BDI)</li> <li><code><strong>Weight history</strong></code> – history of weight changes and frequency of self-weighing</li> <li><code><strong>Eating unspecific</strong></code> – general eating behaviors (DEBQ and FCQ subscales)</li> <li><code><strong>Eating specific</strong></code> – eating disorder–related symptoms (EDEQ total score)</li> <li><code><strong>Binge specific</strong></code> – frequency of binge eating episodes</li> </ol> <p style="color: #37474f; font-size: 15px;"> The severity scores were calculated in two steps: </p> <ul style="color: #37474f; font-size: 15px;"> <li>Each feature was rescaled to a 0–1 range so that all measures were comparable.</li> <li>The features were then averaged, either equally or with weights applied, to produce one score per symptom area.</li> </ul> <p style="color: #37474f; font-size: 15px;"> For the <strong>Eating unspecific</strong> area, we weighted the features so that the DEBQ (3 items) and the FCQ (2 items) each contributed equally (50% each). A similar approach was used for the <strong>Weight history</strong> area, with equal weight (50% each) given to weight fluctuations and number of weigh-ins. </p> <p style="color: #37474f; font-size: 15px;"> The final scores were added to the dataset with the suffix <code>_severity</code> (e.g., <code>Disease_unspecific_severity</code>). </p>
- # %%
- severity_idx = {'Disease_unspecific': ['BDI'],
- 'Weight_history': ['nr_gains_plus5','nr_gains_plus6to10','nr_gains_plus11to20','nr_gains_plus21to30','nr_gains_plus31to40',
- 'nr_gains_plus40','n_weighings'],
- 'Eating_unspecific': ['DEBQ_restrained_eating', 'DEBQ_emotional_eating', 'DEBQ_external_eating','G_FCQ_state', 'G_FCQ_trait'],
- 'Eating_specific': ['EDEQ_total_score'],
- 'Binge_specific': ['n_binges'],
- }
- weights = {'Eating_unspecific': 3*[0.5/3] + 2*[0.25], # weights of the individual factors to the total score (e.g. 50% DEBQ + 50% FCQ)
- 'Weight_history': 6*[0.5/6] + [0.55]}
- # calculate the severity indices
- suffix = '_severity'
- df_all = compute_weighted_severity_indices(df_all, severity_idx, weights,suffix=suffix)
- # Assign the new severity indices names to the data_cats dict
- data_cats['targets'] = {f"{key}{suffix}": value for key, value in severity_idx.items()}
- # %% [markdown]
- # <h6 style="color: #37474f; font-weight: normal; font-size: 15px; margin-top: 30px;">2.1.1.3 CONFOUNDERS & OTHERS</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %%
- data_cats['confounders'] = ['sex','age','BMI']
- data_cats['others'] = ['VAS_hungr_before','VAS_mood_before','VAS_hungr_after','VAS_mood after','hightest_weight',
- 'age_highest_weight','lowest_weight','age_lowest_weight','look','bodyshape_father','bodyshape_mother',
- 'first_time_overweight']
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.2 Outliers and Missing Data</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- #
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> To maintain data quality across different types of measures, we flagged subjects with too much missing or unusable information. This included both missing values and extreme outliers (based on z-scores) across categories such as behavioral, blood, clinical, demographics, fMRI, questionnaires, and sMRI. </p> <p style="color: #37474f; font-size: 15px;"> Subjects were flagged for exclusion if they met either of these thresholds: </p> <ul style="color: #37474f; font-size: 15px;"> <li><strong>Missing data cutoff</strong>: more than 50% missing in a category</li> <li><strong>Z-score threshold</strong>: values above 10 (extreme outliers)</li> </ul> <p style="color: #37474f; font-size: 15px;"> In addition, 9 subjects were removed based on prior knowledge or manual review: <code>K020SJ</code>, <code>K030CN</code>, <code>K031KB</code>, <code>K089KP</code>, <code>K091AH</code>, <code>P010SV</code>, <code>P011GR</code>, <code>P022KB</code>, <code>P025JB</code>. </p> <p style="color: #37474f; font-size: 15px;"> A summary matrix (<code>data_cleaning_summary</code>) was created to track missing and flagged values for each subject and category. Exclusion was decided separately for each category and flagged if the proportion of missing data or outliers exceeded the set thresholds. </p> <p style="color: #37474f; font-size: 15px;"> Instead of removing entire subjects, only the flagged categories were set to <code>NaN</code>. This allowed subjects with partial data to still contribute to analyses in other categories. </p> <p style="color: #37474f; font-size: 15px;"> The following output files were saved: </p> <ul style="color: #37474f; font-size: 15px;"> <li><code>data_cleaning_summary.csv</code> / <code>.xlsx</code> – summary of missing values and outliers by subject and category</li> <li><code>subject_exclusion_summary.xlsx</code> – simplified version of the cleaning summary showing excluded subjects by category</li> <li><code>allData_precleaned_step2.2a_Z.pkl</code> – z-scored version of the main dataframe</li> <li><code>allData_precleaned_step2.2b_Z.pkl</code> – z-scored dataframe with preselected subjects removed</li> <li><code>allData_precleaned_step2.2b.pkl</code> – main dataframe with preselected subjects removed</li> <li><code>allData_precleaned_step2.2c_Final.pkl</code> – final dataframe with outlier data removed</li> </ul>
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # (a) calculate outliers + remove preselected subjects
- # </p>
- # %%
- # features: multiple columns for each modality. Can take some missing data
- missing_cutoff = {'bh_neu': 0.5,'bh_spec': 0.5, 'blood': 0.5,'task_fMRI_spec': 0.5,'task_fMRI_neu': 0.5,'fMRI': 0.5, 'MRI': 0.5,
- # targets: only one column. so if the value is missing, the subject will be excluded for that regression analysis.
- 'Disease_unspecific_severity': 0.5, 'Weight_history_severity': 0.5,'Eating_unspecific_severity': 0.5,
- 'Eating_specific_severity': 0.5,'Binge_specific_severity':0.5}
- # features
- zthresholds = {'bh_neu': 10,'bh_spec': 10, 'blood': 10,'task_fMRI_spec': 10,'task_fMRI_neu': 10,'fMRI': 10, 'MRI': 10,
- # targets: Dont exclude anybody based on z scores here
- 'Disease_unspecific_severity': None, 'Weight_history_severity': None,'Eating_unspecific_severity': None,
- 'Eating_specific_severity': None,'Binge_specific_severity':None}
- subs2rem = ['K020SJ','K030CN','K031KB','K089KP','K091AH','P010SV','P011GR','P022KB','P025JB']
- # create a z scored version of that dataframe. Use only float columns
- zscore = pd.DataFrame(np.abs(stats.zscore(df_all.select_dtypes(include=["float"]),nan_policy='omit')),columns=df_all.columns,index=df_all.index)
- # Remove pre-selected subjects from both zscore and df_all
- zscore2 = zscore.drop(index=subs2rem, errors='ignore')
- df_all2 = df_all.drop(index=subs2rem, errors='ignore')
- allMissingData = pd.DataFrame(index = df_all2.index) # collect all information here
- missing_data_list = [] # Prepare a list to collect the missing data for each category combination
- # We check missing and outlier data only for features and targets
- for modality, features in getCatDataMapping(my_dict=data_cats).items():
- # Pre-compute thresholds to avoid repetitive dictionary lookups
- z_thresh = zthresholds[modality]
- missing_thresh = missing_cutoff[modality]
- # Vectorized calculations for the missing data and exclusion flags
- if z_thresh is not None:
- zscore_above_thresh = zscore2[features].gt(z_thresh).sum(axis=1)
- else: # Don't apply any z-threshold filtering
- zscore_above_thresh = pd.Series(0, index=zscore2.index)
- missing_features = zscore2[features].isnull().sum(axis=1)
- total_features = len(features)
- missing_total = (zscore_above_thresh + missing_features) / total_features
- exclusion_flag = (missing_total > missing_thresh)
- # Create a DataFrame for the current category and append to the list
- missing_data = pd.DataFrame({
- f'zthresh_{z_thresh}_{modality}': zscore_above_thresh,
- f'missing_{modality}': missing_features,
- f'ntotal_{modality}': total_features,
- f'missingtotal_{modality}': missing_total,
- f'exclusion_{modality}': exclusion_flag
- })
- missing_data_list.append(missing_data)
- # Concatenate all the missing data into a single DataFrame
- allMissingData = pd.concat(missing_data_list, axis=1)
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # (b) replace outlier data with NaNs
- # </p>
- # %%
- # remove subject data that was flagged for exclusion (we do this individually for each category so subjects may remain in the dataframe)
- df_all3 = df_all2.copy() # removed subject data will be replaced with NaN and then excluded when we split the frame
- for modality, features in getCatDataMapping(my_dict=data_cats).items():
- missing_thresh = missing_cutoff[modality]
- curr_cols = allMissingData.filter(like = 'exclusion_'+modality) # list of flagged subjects
- if len(curr_cols) > 1: # if we have more than one exclusion flag per category, take the mean
- curr_flag = curr_cols.mean(axis=1) > missing_thresh
- else:
- curr_flag = curr_cols
- printmd('**' +modality+':** ' + str(len(curr_flag.loc[curr_flag==True])) +' subjects were excluded') # number of subjects to be exclusion
- printmd(str(curr_flag.loc[curr_flag==True])) # number of subjects to be exclusion
- printmd(str(len(curr_flag.loc[curr_flag==False])) +' subjects were used') # number of subjects to be exclusion
- # mark all flagged subject data as nan. This is done per datatype because subjects may have enough good data
- # of some datatypes and not enough of the others. So only the data of bad datatypes are removed but the subject stays in the analysis for now
- for val in features:
- df_all3[val]=np.where(curr_flag==True, np.nan, df_all3[val])
- # For resting-state data, do the same with the full matrix used for analyses
- if modality == 'fMRI':
- flagged_subs = curr_flag.loc[curr_flag==True].index.tolist() # if there is any data for these subjects, set it to NaN
- for key in flagged_subs:
- if key in rs_matrix:
- rs_matrix[key][:] = np.nan # Sets the existing array's values to NaN
- for group_name, sub_df in df_all3.groupby('group'):
- print(f"\nGroup: {groupNames_rev[group_name]}")
- print(f"Count: {len(sub_df)}")
- print("Subjects (index):")
- print(sub_df.index.tolist())
- print ('-----')
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;">
- # c) manual data exclusion
- # </p>
- # %%
- df_all3.loc[df_all3['glucose'] < 60, 'glucose'] = np.nan # exclude glucose values below 60 (4 subjects)
- df_all3.loc[df_all3['triglyceride'] < 10, ['cholesterin','triglyceride']] = np.nan # exclude trygliceride values below 10 (3 subjects). Remove the same subjects for extremely low cholesterine values
- df_all3.loc['K098EW', data_cats['features']['blood']] = np.nan # this subjects has more than half of the blood values missing now, so we remove his blood data completely
- # check individual subjects
- #df_all3.loc['P104PN', data_cats['features']['blood']]
- #df_all3.groupby(['group','sex'])['sex'].describe()
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">2.3 Save Preprocessed Data</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %%
- saveFile = False
- # %%
- if saveFile:
- print ('saving preprocessed files...')
- # create an overview over the exclusions
- # Start with subject index and group
- summary_df = df_all[['group']].copy()
- summary_df = summary_df.loc[summary_df.index.union(subs2rem)] # Ensure pre-excluded are included
- # Add group name
- summary_df['group_name'] = summary_df['group'].map({v: k for k, v in groupNames.items()})
- # Add pre-exclusion flag
- summary_df['pre_excluded'] = summary_df.index.isin(subs2rem)
- # Add all exclusion flags from the main missing data check
- exclusion_cols = [col for col in allMissingData.columns if col.startswith('exclusion_')]
- summary_df = summary_df.join(allMissingData[exclusion_cols], how='left')
- # Fill NA exclusion flags with False for subjects not evaluated (e.g., pre-excluded)
- pd.set_option('future.no_silent_downcasting', True)
- summary_df[exclusion_cols] = summary_df[exclusion_cols].apply(lambda col: col.fillna(False).astype(bool))
- # (1) data for ML analyses
- df_all3.drop(columns=df_all3.filter(like='_RS_').columns).to_pickle(f'{out_dir_ML}/allData_precleaned_step2.2c_Final.pkl') # with outliers replaced by nans + rs columns dropped
- with open(f'{out_dir_ML}/data_cats.pkl', 'wb') as f:
- pickle.dump(data_cats, f)
- with open(f'{out_dir_ML}/AAL3Matrix_precleaned_step2.2c_Final.pkl', 'wb') as f:
- pickle.dump(rs_matrix, f)
- # (2) additional data safeties
- allMissingData.to_csv(f'{out_dir_expl}/data_cleaning_summary.csv', index=True)
- allMissingData.to_excel(f'{out_dir_expl}/data_cleaning_summary.xlsx', index=True)
- summary_df.to_excel(f'{out_dir_expl}/subject_exclusion_summary.xlsx', index=True)
- vars_all.to_pickle(f'{out_dir_saf}/allVars_precleaned_step1.pkl') # variable list without empty columns
- zscore.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2a_Z.pkl') # zscored version of df_all
- zscore2.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2b_Z.pkl') # with pre-selected subs removed
- df_all.to_pickle(f'{out_dir_saf}/allData_precleaned_step1.pkl') # data without empty columns
- df_all.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.1a.pkl') # added severity indices
- df_all2.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2b.pkl') # with pre-selected subs removed
- df_all3.to_pickle(f'{out_dir_saf}/allData_precleaned_step2.2c_Final.pkl') # with outliers replaced by nans
- # save both data categorizations
- with open(f'{out_dir_saf}/data_cats.pkl', 'wb') as f:
- pickle.dump(data_cats, f)
- with open(f'{out_dir_saf}/data_cats_full.pkl', 'wb') as f:
- pickle.dump(data_cats_full, f)
- print ('Done!')
- else:
- print ('Not saving files due to saveFile=False')
- # %% [markdown]
- # <h3 style="color: #1e88e5; font-weight: bold; font-size: 24px; margin-top: 50px;">3 DATA EXPLORATION</h3>
- # <hr style="border: 1px solid #cfd8dc; margin-top: 0; margin-bottom: 15px;">
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.1 Raw Data Inspection</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> For each modality, outliers were examined independently, regardless of how the data were grouped for later analysis. Outliers were also checked visually. In the behavioral data, outliers in error scores appeared as sharp peaks, usually when subjects made no errors. </p>
- # %%
- saveFile = True#saveFiles # use global parameter
- z_score_out = Path(out_dir_expl).parent.joinpath('z_thresholds')
- z_score_out.mkdir(exist_ok=True)
- # %%
- NonMRICats = [i for i in data_cats_full['cat1'] if 'MRI' not in i]
- for cat in NonMRICats[:]:
- curr_cols = vars_all.loc[vars_all['category1']==cat]['feature'].tolist()
- curr_data = df_all2[curr_cols]
- #print (curr_data.describe())
- printmd("**------------------------------ "+cat+" ------------------------------**")
- if saveFile:
- zscored = convert_display_zScore(curr_data,zthresh=3,saveStats=str(z_score_out)+'/BH_threshs'+cat+'.csv')
- else:
- zscored = convert_display_zScore(curr_data,zthresh=3,saveStats=None)
- # %%
- curr_cols = vars_all.loc[vars_all['category1']=='sMRI']['feature'].tolist()
- curr_data = df_all[curr_cols]
- printmd("**------------------------------ VBM ------------------------------**")
- #print (curr_data.describe()[:5])
- if saveFile:
- zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=str(z_score_out)+'/VBM_threshs.csv')
- else:
- zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=None)
- # %%
- TaskfMRICats = vars_all.loc[(vars_all['category1']=='fMRI')&(vars_all['category2'].isin(['MID', 'GNG']))]['category2'].unique()
- for cat in TaskfMRICats:
- curr_cols = vars_all.loc[(vars_all['category1']=='fMRI')&(vars_all['category2'].isin([cat]))]['feature'].tolist()
- curr_data = df_all[curr_cols]
- printmd("**------------------------------ "+cat+" ------------------------------**")
- #print (curr_data.describe())
- if saveFile:
- zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=str(z_score_out)+'/TaskfMRI_threshs'+cat+'.csv')
- else:
- zscored = convert_display_zScore(curr_data.iloc[:, :20],zthresh=3,saveStats=None)
- # %%
- curr_cols = vars_all.loc[vars_all['category2']=='RS']['feature'].tolist()
- curr_data = df_all[curr_cols]
- printmd("**------------------------------ Resting-State functional connectivity ------------------------------**")
- #print (curr_data.describe()[:5])
- if saveFile:
- zscored = convert_display_zScore(curr_data.iloc[:, :10],zthresh=3,saveStats=str(z_score_out)+'/RS_threshs.csv')
- else:
- zscored = convert_display_zScore(curr_data.iloc[:, :10],zthresh=3,saveStats=None)
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2 Final Sample Inspection</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2.1 Features</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> We check the intercorrelation between the features of each modality for the entire sample. The code can be easily adjusted to check just for specific groups or across modalities </p>
- # %%
- saveFile = False#saveFiles # use global parameter
- corr_out = Path(out_dir_expl).parent.joinpath('correlations')
- corr_out.mkdir(exist_ok=True)
- # %% [markdown]
- # (a) Correlations
- # %%
- group_combs = {'BED':[1.0],'CON_BED':[3.0],'BN':[2.0],'CON_BN':[4.0],'PAT':[1.0,2.0],
- 'HC':[3.0,4.0],'O_weight':[1.0,3.0],'N_weight':[2.0,4.0],'All':[1.0,2.0,3.0,4.0]}
- group_combs = {'All':[1.0,2.0,3.0,4.0]} # do all just for now
- for curr_feat_name,curr_cols in data_cats['features'].items():
- if curr_feat_name == 'fMRI': # no need to run this for RS data because this is already a correlation matrix
- continue
- print (curr_feat_name)
- for group_name,group_idx in group_combs.items():
- print (group_name)
- curr_data = df_all3.loc[df_all3['group'].isin(group_idx)][curr_cols].dropna(axis = 0, how = 'all')
- # We save these as a file, so we dont have to run it again. We can plot it with the fromFile key
- # dont show table if the category has many features
- #printmd("**------------------------------" + curr_feat_name + " features Correlation Matrix ------------------------------**")
- showtable = False
- rho,pval,ax = corrMatrix(curr_data,table=showtable,plot=False,cmap="RdBu_r",
- saveFile=f'{str(corr_out)}/{curr_feat_name}_{group_name}')
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> Here we check the level of significant corrations within feature modalities. </p>
- # %%
- for curr_mod_name,curr_cols in data_cats['features'].items():
- curr_data = df_all3[curr_cols].dropna(axis = 0, how = 'all')
- printmd(f"**------------------------------ {curr_mod_name} features Intercorrelations ------------------------------**")
- # Skip fMRI category for now
- if 'fMRI' not in curr_mod_name:
- # Get correlation matrix and p-values
- rho, pval, ax = corrMatrix(curr_data, table=False, plot=False, cmap="RdBu_r",
- saveFile = f'{str(corr_out)}/{curr_mod_name}_{group_name}',
- fromFile=True)
- # Extract significant correlations
- sign_pairs = [
- (col, i) for col in pval.columns for i in pval[pval[col] < 0.05].index if i != col
- ]
- sign_pairs = list(set(tuple(sorted(pair)) for pair in sign_pairs)) # Remove duplicates
- # Get all combinations
- combos = list(combinations(pval.columns, 2))
- print(f"{len(sign_pairs)} significant correlations out of {len(combos)} feature combinations")
- # Plot some example pairs
- col_ranges = list(range(0, len(sign_pairs), 6))[:2]
- for i in range(len(col_ranges) - 1):
- curr_pairs = sign_pairs[col_ranges[i]:col_ranges[i + 1]]
- # Create subplots with a 1x6 grid
- fig, ax = plt.subplots(1, 6, figsize=(11, 2))
- for pair_n, (x_col, y_col) in enumerate(curr_pairs):
- sns.regplot(
- x=x_col, y=y_col, ax=ax[pair_n], data=curr_data,
- scatter_kws={"color": "darkred", "alpha": 0.3, "s": 20},
- line_kws={"color": "k", "alpha": 0.8, "lw": 2}
- )
- plt.tight_layout()
- plt.show()
- plt.clf()
- plt.close()
- break # remove to plot everything
- # %% [markdown]
- # <h4 style="color: #37474f; font-weight: bold; font-size: 18px; margin-top: 30px;">3.2.1 Targets</h4>
- # <hr style="border: 0; height: 1px; background-color: #e0e0e0; margin-top: 4px; margin-bottom: 15px;">
- # %% [markdown]
- # <p style="color: #37474f; font-size: 15px;"> Here we check the level of significant corrations within feature modalities. </p>
- # %%
- #for group in range(1,5):
- curr_data = df_all3[data_cats['targets'].keys()]#.loc[df_reduced['group']==group]
- rho,pval,ax = corrMatrix(curr_data,table=True,plot=True)
- # %%
- from scipy.stats import f_oneway
- from statsmodels.stats.multicomp import pairwise_tukeyhsd
- from statannotations.Annotator import Annotator
- from statsmodels.formula.api import ols
- import statsmodels.api as sm
- plotting_df = df_all3.copy()
- plotting_df['named_group'] = plotting_df['group'].map(groupNames_rev)
- # Prepare long-form dataframe for two-way ANOVA
- severity_features = data_cats['targets'].keys()
- df_long = plotting_df.melt(id_vars=['named_group'], value_vars=severity_features,
- var_name='severity_index', value_name='value').dropna()
- # Run two-way ANOVA with interaction
- model = ols('value ~ C(named_group) * C(severity_index)', data=df_long).fit()
- anova_table = sm.stats.anova_lm(model, typ=2)
- print("\n Two-way ANOVA Results (Group x Severity Index):\n")
- print(anova_table)
- annotate = True # plot test significance
- for idx, cols in severity_idx.items():
- if annotate:
- plt.figure(figsize=(4, 6))
- else:
- plt.figure(figsize=(4, 4))
- cols = default_cols(n=2, random=True, n_set=None)
- print(cols)
- feature = idx + '_severity'
- df_plot = plotting_df.dropna(subset=[feature])
- # Violin plot
- sns.violinplot(
- data=df_plot,
- x='named_group',
- y=feature,
- inner='box',
- color='white',
- edgecolor='black'
- )
- # Swarm plot
- ax = sns.swarmplot(
- data=df_plot,
- x='named_group',
- y=feature,
- hue='named_group',
- alpha=0.9,
- palette=2 * cols,
- dodge=False,
- s=4.5,
- )
- # ANOVA
- groups = [group_df[feature].values for name, group_df in df_plot.groupby('named_group')]
- stat, p_anova = f_oneway(*groups)
- plt.title(f"{idx.replace('_', ' ').title()} Severity by Group\nANOVA p={p_anova:.3e}")
- # Post-hoc Tukey HSD if ANOVA is significant
- if p_anova < 0.05:
- tukey = pairwise_tukeyhsd(endog=df_plot[feature], groups=df_plot['named_group'], alpha=0.05)
- results_df = pd.DataFrame(data=tukey._results_table.data[1:], columns=tukey._results_table.data[0])
- # Filter for significant comparisons (p <= 0.05)
- sig_results = results_df[results_df['p-adj'] <= 0.05]
- # If any significant pairs exist, annotate them
- if not sig_results.empty:
- pairs = [(row['group1'], row['group2']) for _, row in sig_results.iterrows()]
- pvals = [row['p-adj'] for _, row in sig_results.iterrows()]
- if annotate:
- annotator = Annotator(ax, pairs, data=df_plot, x='named_group', y=feature)
- annotator.set_pvalues_and_annotate(pvalues=pvals)
- plt.xlabel("Group")
- plt.ylabel(idx)
- plt.legend([], [], frameon=False) # Remove redundant legend
- plt.tight_layout()
- plt.show()
- # %%
NeuroBED_ML_DataExploration.ipynb at commit c5e2120, under MIT · at the source
Overview
- Department of General Internal Medicine, Psychosomatics and Psychotherapy, Centre for Psychosocial Medicine, University Hospital Heidelberg, Heidelberg, Germany
- Institute of Pathology, University Hospital Heidelberg, Heidelberg, Germany
- Department of Neuroradiology, University Hospital Heidelberg, Heidelberg, Germany
- DZPG (German Centre for Mental Health – Partner Site Heidelberg/Mannheim/Ulm), Germany
Abstract
Introduction: Binge-type eating disorders, including bulimia nervosa (BN) and binge eating disorder (BED), are associated with both shared and disorder-specific neurobiological mechanisms across brain, behavior, and physiology. A clearer distinction between shared mechanisms and disorder-specific alterations may advance our understanding of binge-type eating pathology.
Methods: We applied a comprehensive multimodal machine learning framework to 110 participants (BN, BED, and age & weight matched controls), integrating task-based fMRI, intrinsic connectivity, voxel-based morphometry, neuropsychological assessments, and peripheral blood biomarkers. Both unimodal and multimodal machine learning models were trained to classify groups and to predict individual variation in symptom expression.
Results: Functional brain connectivity achieved the highest accuracy for diagnostic classification and symptom prediction (with a mean balanced classification accuracy (bACC) of 68.7%), whereas task-based fMRI with disorder-specific food stimuli and peripheral blood biomarkers best distinguished BN from BED (mean bACC of 87%). Multimodal models did not generally outperform the best unimodal approaches, except from modest gains in a limited set of regression targets.
Conclusions: These findings suggest that functional brain connectivity carries robust predictive information for transdiagnostic classification, whereas task-evoked activation patterns and peripheral biomarkers show stronger predictive utility for distinguishing BN from BED. Whether these modality-specific patterns reflect underlying neurobiological mechanisms remains to be established in future hypothesis-driven work. Identifying which modalities best represent shared vulnerability vs. symptom-type-dependent variation may help to provide a foundation for a more mechanistic understanding of these disorders.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.
LenaRo09/NeuroBED_ML
c5e21208047398a4e06f72827bd96e102792448f, 26 March 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
39 files
- NeuroBED_ML/
analysis/ , Jupyter, 776 lines, 4 matches02 - data_exploration/ NeuroBED_ML_DataExplorat ion.ipynb - NeuroBED_ML/
analysis/ , Python, 141 lines04 - models/ 01 - model0_RS/ 04 - scripts/ 01 - NeuroBED_ML_model0_aggRe sults.py - NeuroBED_ML/
analysis/ , Python, 120 lines04 - models/ 01 - model0_RS/ 04 - scripts/ 02 - AAL3_extractROIs2nifti.p y - NeuroBED_ML/
analysis/ , Python, 164 lines04 - models/ 01 - model0_RS/ 04 - scripts/ 03 - Enrichment+Jaccard.py - NeuroBED_ML/
analysis/ , Python, 358 lines, 1 match04 - models/ 01 - model0_RS/ model0.py - NeuroBED_ML/
analysis/ , Shell, 24 lines04 - models/ 01 - model0_RS/ model0.sh - NeuroBED_ML/
analysis/ , Python, 165 lines04 - models/ 02 - model1_singleModality/ 04 - scripts/ 01 - NeuroBED_ML_aggResults.p y - NeuroBED_ML/
analysis/ , Python, 115 lines04 - models/ 02 - model1_singleModality/ 04 - scripts/ 02 - NeuroBED_ML_cleanResults .py - NeuroBED_ML/
analysis/ , Python, 381 lines, 4 matches04 - models/ 02 - model1_singleModality/ model1.py - NeuroBED_ML/
analysis/ , Shell, 29 lines04 - models/ 02 - model1_singleModality/ model1.sh - NeuroBED_ML/
analysis/ , Python, 166 lines04 - models/ 03 - model2_combinedModalitie s/ 04 - scripts/ 01 - NeuroBED_ML_aggResults.p y - NeuroBED_ML/
analysis/ , Python, 116 lines04 - models/ 03 - model2_combinedModalitie s/ 04 - scripts/ 02 - NeuroBED_ML_cleanResults .py - NeuroBED_ML/
analysis/ , Python, 472 lines, 3 matches04 - models/ 03 - model2_combinedModalitie s/ model2.py - NeuroBED_ML/
analysis/ , Shell, 44 lines04 - models/ 03 - model2_combinedModalitie s/ model2.sh - NeuroBED_ML/
analysis/ , Python, 88 lines04 - models/ 04 - crossModel_results/ 01 - scripts/ 01 - NeuroBED_ML_crossModelag g.py - NeuroBED_ML/
analysis/ , Python, 240 lines04 - models/ 04 - crossModel_results/ 01 - scripts/ 02 - NeuroBED_ML_crossStat.py - NeuroBED_ML/
analysis/ , Python, 23 lines04 - models/ 04 - crossModel_results/ 01 - scripts/ Add1 - NeuroBED_ML_loadResults. py - NeuroBED_ML/
analysis/ , Python, 210 lines04 - models/ 05 - miscellaneous/ 01 - confounderBaseline/ NeuroBED_ML_confounderBa seline.py - NeuroBED_ML/
analysis/ , Python, 63 lines04 - models/ 05 - miscellaneous/ load_single_model.py - NeuroBED_ML/
analysis/ , Python, 532 lines05 - utils/ LR_Cmaps.py - NeuroBED_ML/
analysis/ , Python, 146 lines05 - utils/ NeuroBED_ML_classifiers. py - NeuroBED_ML/
analysis/ , Python, 797 lines05 - utils/ NeuroBED_ML_helpers.py - NeuroBED_ML/
analysis/ , Python, 1 line05 - utils/ __init__.py - NeuroBED_ML/
analysis/ , Python, 191 lines, 1 match06 - plotting/ FIG 02 - ClassResults/ scripts/ A_plot_heatmap_target_mo dality.py - NeuroBED_ML/
analysis/ , Python, 228 lines06 - plotting/ FIG 02 - ClassResults/ scripts/ B_scoreX#modality.py - NeuroBED_ML/
analysis/ , Python, 172 lines06 - plotting/ FIG 02 - ClassResults/ scripts/ C1_class_splitviolin_bes tsinglemulti.py - NeuroBED_ML/
analysis/ , Python, 173 lines06 - plotting/ FIG 02 - ClassResults/ scripts/ C2_delta_bACC.py.py - NeuroBED_ML/
analysis/ , Python, 286 lines, 2 matches06 - plotting/ FIG 02 - ClassResults/ scripts/ D_confusion_matrix_grid. py - NeuroBED_ML/
analysis/ , Python, 155 lines06 - plotting/ FIG 04 - RegrResults/ scripts/ A_regr_heatmap_targetXmo d.py - NeuroBED_ML/
analysis/ , Python, 193 lines06 - plotting/ FIG 04 - RegrResults/ scripts/ B_regr_split_violin_best SinglevsMulti.py - NeuroBED_ML/
analysis/ , Python, 170 lines, 1 match06 - plotting/ FIG 04 - RegrResults/ scripts/ C_regr_univariate_result s_targetXgroup.py - NeuroBED_ML/
analysis/ , Python, 153 lines06 - plotting/ FIG 04 - RegrResults/ scripts/ SM_FIG2_regr_heatmap_tar getXgroupXmod.py - NeuroBED_ML/
analysis/ , Python, 480 lines06 - plotting/ FIG SM01 - model0_SankeyJaccardViol in/ A_NeuroBED_ML_SankeyPlot .py - NeuroBED_ML/
analysis/ , Python, 635 lines06 - plotting/ FIG SM01 - model0_SankeyJaccardViol in/ B_Model0_CorrelationMatr ices.py - NeuroBED_ML/
analysis/ , Python, 217 lines06 - plotting/ FIG SM01 - model0_SankeyJaccardViol in/ C_Jackard_plotting.py - NeuroBED_ML/
analysis/ , Python, 180 lines06 - plotting/ FIG SM01 - model0_SankeyJaccardViol in/ D_model0_RS_violinTop1.p y - NeuroBED_ML/
analysis/ , Python, 185 lines06 - plotting/ FIG SM01 - model0_SankeyJaccardViol in/ split_heatmap.py - LICENSE, License, 21 lines
- README.md, Text, 69 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 37 scripts, each with its path and the digest of its content;
- 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability statement
The datasets presented in this study can be found invonline repositories. The names of the repository/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 6 authors, 5 keywords, 1 funder, 91 references.
Cite
This paper
Rommerskirchen, L., Skunde, M., Bendszus, M., Herzog, W., Friederich, H.-C., & Simon, J. J. (2026). Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders. Frontiers in neuroscience, 20, 1803154. https://
BibTeX
@article{rommerskirchen2
author = {Rommerskirchen, Lena and Skunde, Mandy and Bendszus, Martin and Herzog, Wolfgang and Friederich, Hans-Christoph and Simon, Joe J},
title = {{Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders}},
journal = {Frontiers in neuroscience},
year = {2026},
month = apr,
volume = {20},
pages = {1803154},
publisher = {Frontiers Media SA},
issn = {1662-4548},
doi = {10.3389/
url = {https://
pmid = {42051561},
pmcid = {PMC13111282}
}
RIS
TY - JOUR
AU - Rommerskirchen, Lena
AU - Skunde, Mandy
AU - Bendszus, Martin
AU - Herzog, Wolfgang
AU - Friederich, Hans-Christoph
AU - Simon, Joe J
TI - Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders
T2 - Frontiers in neuroscience
J2 - Front Neurosci
PY - 2026
DA - 2026/
VL - 20
SP - 1803154
SN - 1662-4548
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders",
"container-title": "Frontiers in neuroscience",
"author": [
{
"family": "Rommerskirchen",
"given": "Lena"
},
{
"family": "Skunde",
"given": "Mandy"
},
{
"family": "Bendszus",
"given": "Martin"
},
{
"family": "Herzog",
"given": "Wolfgang"
},
{
"family": "Friederich",
"given": "Hans-Christoph"
},
{
"family": "Simon",
"given": "Joe J"
}
],
"container-title-short":
"volume": "20",
"page": "1803154",
"DOI": "10.3389/
"PMID": "42051561",
"PMCID": "PMC13111282",
"ISSN": "1662-4548",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
13
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1371/journal.pmed.1004809 [code]
- Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data.Journal: PLoS medicineIn common: statannotations, NiBabel, statsmodels, 6 other tools, other condition, 3 references
- [2] doi:10.1038/s43856-026-01518-5 [code]
- Machine learning-based identification of abnormal functional connectivity in obesity across different metabolic states.Journal: Communications medicineIn common: statannotations, XGBoost, statsmodels, 6 other tools, other condition, 2 references
- [3] doi:10.1016/j.isci.2026.115329 [code]
- Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.Journal: iScienceIn common: statannotations, SHAP, XGBoost, 8 other tools, other condition
- [4] doi:10.21203/rs.3.rs-9914920/v1 [code]
- Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working GroupJournal: Research Square (preprint)In common: statannotations, SHAP, XGBoost, 8 other tools
- [5] doi:10.1016/j.nicl.2026.104012 [code]
- Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.Journal: NeuroImage. ClinicalIn common: SHAP, XGBoost, NiBabel, 7 other tools, fMRI, other condition
- [6] doi:10.1038/s41598-026-56688-y [code]
- On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.Journal: Scientific reportsIn common: SHAP, XGBoost, NiBabel, 7 other tools, fMRI
- [7] doi:10.1002/hbm.70469 [code]
- VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.Journal: Human brain mappingIn common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
- [8] doi:10.21203/rs.3.rs-9326213/v1 [code]
- Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brainJournal: Research Square (preprint)In common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
- [9] doi:10.64898/2026.03.09.710558 [code]
- Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brainJournal: bioRxiv (preprint)In common: statannotations, NiBabel, seaborn, 5 other tools, fMRI, 2 references
- [10] doi:10.1038/s42003-026-10276-y [code]
- The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.Journal: Communications biologyIn common: statannotations, NiBabel, statsmodels, 6 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 37 scripts, and 16 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:d68e73f93d2453fa…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
