Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder.
The 10 matches
- [1] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ notebooks/1_input_prep/3_EPOC_fMRI_graph_features/weighted_graph_functions.py, lines 306–323 · score 0.88 · global efficiency, clustering coefficient, correlation matrices, graph features, density, sigma
- [2] § Methods › Analysis ↔ src/MLpipeline/runMLpipelines.py, lines 50–195 · score 0.83 · random seeds, inner CV, outer CV, ML models, machine, stratified
- [3] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ notebooks/1_input_prep/3_EPOC_fMRI_graph_features/graph_features.py, lines 12–37 · score 0.82 · global efficiency, clustering coefficient, graph features, density, sigma, transitivity
- [4] § Methods › Analysis ↔ src/MLpipeline/MLpipeline.py, lines 236–342 · score 0.74 · inner CV folds, cross validation, stratified, splits, tuned, binary
- [5] § Methods › Feature importance ↔ data/results/classification_MRI_PCA_Remission/fMRI_OUTRemission_CONFSSex/20240515-1524/classification_Marija_MRI_PCA_Remission_20240515_1524.py, lines 13–83 · score 0.66 · linear SVM, gradient boosting, logistic regression, fitting, kernel, RBF
- [6] § Methods › Feature importance ↔ data/results/classification_MRI_PCA_Response/fMRI_OUTResponse/20240516-1018/classification_Marija_MRI_PCA_20240516_1018.py, lines 13–83 · score 0.66 · linear SVM, gradient boosting, logistic regression, fitting, kernel, RBF
- [7] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ src/MLpipeline/MLpipeline.py, lines 236–342 · score 0.63 · cross validated, scikit-learn, correlation, preprocessing, hyperparameter, inner
- [8] § Results › ML predictions ↔ notebooks/3_results/p_values_plotting/results_figures_EPOC.ipynb, lines 427–440 · score 0.56 · sMRI, Tabular OCD, fMRI, multimodal, PCA, graph
- [9] § Methods › Analysis ↔ data/results/classification_MRI_PCA_Remission/fMRI_OUTRemission_CONFSSex/20240515-1524/classification_Marija_MRI_PCA_Remission_20240515_1524.py, lines 13–83 · score 0.52 · gradient boosting, logistic regression, GB, LR, kernel, linear
- [10] § Methods › Analysis ↔ data/results/classification_Remission/MRI_OUTRemission_CONFSSex/20240202-1246/classification_Marija_MRI_Remission_20240202_1246.py, lines 13–83 · score 0.52 · gradient boosting, logistic regression, GB, LR, kernel, linear
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 424 lines · 19 KB · no license · 2 matches
- from os.path import join
- import h5py
- import numpy as np
- import pandas as pd
- import statistics as st
- from datetime import datetime
- from joblib import Parallel, delayed
- from sklearn.model_selection import GridSearchCV
- from sklearn.metrics import get_scorer, confusion_matrix
- from sklearn.model_selection import train_test_split, StratifiedKFold
- from sklearn.preprocessing import OneHotEncoder, StandardScaler
- from sklearn.utils._testing import ignore_warnings
- from sklearn.exceptions import ConvergenceWarning
- from sklearn.decomposition import PCA
- class MLpipeline:
- def __init__(self, parallelize=True, random_state=None):
- """ Define parameters for the MLLearn object.
- Args:
- "h5file" (str): Full path to the HDF5 file.
- "conf_list" (list of str): the list of confound names to load from the hdf5.
- random_state (int): A random state to get reproducible results
- """
- self.random_state = random_state
- self.parallelize = parallelize
- self.confs = {}
- self.X_test = None
- self.y_test = None
- self.confs_test = {}
- self.sub_ids = None
- self.sub_ids_test = None
- # These are defined to calculate later.
- self.func_auc = get_scorer("roc_auc")
- self.func_r2 = get_scorer('r2')
- def load_data(self, f, X="X", y="y", confs=[], group_confs=False):
- """ Load the data from the HDF5 file,
- The neuroimaging data/ input data is saved under "X", the label under "y" and
- the confounds from the dict self.confs
- If group_confs==True, groups all confounds into one numpy array and saves in
- self.confs["group"]
- """
- if isinstance(f, dict):
- df = f
- else:
- df = h5py.File(f, "r")
- self.X = np.array(df[X])
- self.y = np.array(df[y])
- self.sub_ids = np.array(df["i"])
- if df.get('scale_info'): #edit Marija
- self.scale_info = list(df["scale_info"]) #edit Marija
- if df.get('confs_scale_info'): #edit Marija
- self.confs_scale_info = list(df["confs_scale_info"]) #edit Marija
- # self.X needs to be flattened as sklearn expects 2D input.
- if self.X.ndim > 2:
- self.X = self.X.reshape([self.X.shape[0], np.prod(self.X[0].shape)])
- elif self.X.ndim == 1:
- self.X = self.X.reshape(-1,1) # sklearn expects 2D input
- # Load the confounding variables into a dictionary, if there are any
- for c in confs:
- if c in df.keys():
- v = np.array(df[c])
- self.confs[c] = v
- if group_confs:
- if "group" not in self.confs:
- self.confs["group"] = v
- else:
- self.confs["group"] = v + 100*self.confs["group"]
- else:
- print(f"[WARN]: confound {c} not found in hdf5 file.")
- def train_test_split(self, test_idx=[], test_size=0.25, stratify_by_conf=None):
- """ Split the data into a test set (_test) and a train+val set.
- Args:
- stratify_group (str, optional): A string that indicates a confound name.
- Get confounds list available from the dict self.confs.
- If a confound is selected, subjects will be stratified according
- to confound and outcome label. Defaults to None, in which case
- subjects are only stratified according to the outcome label.
- """
- # if test_idx are already provided then dont generate the test_idx yourself
- if not len(test_idx):
- # by default, always stratify by label
- stratify = self.y
- if stratify_by_conf is not None:
- stratify = self.y + 100*self.confs[stratify_by_conf]
- _, test_idx = train_test_split(range(len(self.X)),
- test_size=test_size,
- stratify=stratify,
- shuffle=True,
- random_state=self.random_state)
- test_mask = np.zeros(len(self.X), dtype=bool)
- test_mask[test_idx] = True
- self.X_test = self.X[test_mask]
- self.y_test = self.y[test_mask]
- self.sub_ids_test = self.sub_ids[test_mask]
- self.X = self.X[~test_mask]
- self.y = self.y[~test_mask]
- self.sub_ids = self.sub_ids[~test_mask]
- for c in self.confs:
- self.confs_test[c] = self.confs[c][test_mask]
- for c in self.confs:
- self.confs[c] = self.confs[c][~test_mask]
- self.n_samples_tv = len(self.y)
- self.n_samples_test = len(self.y_test)
- # fill missing values edit Marija
- def fill_missing_vals(self):
- """ Fill in missing values in outer CV loop."""
- for X_col in range(self.X.shape[1]):
- if np.isnan(self.X[:,X_col]).any(axis=0):
- if self.scale_info[X_col] == b'categorical':
- fill_val = st.mode(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
- self.X[:,X_col][np.isnan(self.X[:,X_col])] = fill_val
- elif self.scale_info[X_col] == b'numerical':
- fill_val = st.median(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
- #fill_val = np.nanmean(self.X[:,X_col])
- self.X[:,X_col][np.isnan(self.X[:,X_col])] = fill_val
- if np.isnan(self.X_test[:,X_col]).any(axis=0):
- if self.scale_info[X_col] == b'categorical':
- fill_val = st.mode(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
- self.X_test[:,X_col][np.isnan(self.X_test[:,X_col])] = fill_val
- elif self.scale_info[X_col] == b'numerical':
- fill_val = st.median(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
- #fill_val = np.nanmean(self.X[:,X_col])
- self.X_test[:,X_col][np.isnan(self.X_test[:,X_col])] = fill_val
- if self.confs:
- for X_col, c in enumerate(self.confs):
- if np.isnan(self.confs[c]).any(axis=0):
- if self.confs_scale_info[X_col] == b'categorical':
- fill_val = st.mode(self.confs[c][~np.isnan(self.confs[c])])
- self.confs[c][np.isnan(self.confs[c])] = fill_val
- elif self.confs_scale_info[X_col] == b'numerical':
- fill_val = st.median(self.confs[c][~np.isnan(self.confs[c])])
- self.confs[c][np.isnan(self.confs[c])] = fill_val
- if np.isnan(self.confs_test[c]).any(axis=0):
- if self.confs_scale_info[X_col] == b'categorical':
- fill_val = st.mode(self.confs[c][~np.isnan(self.confs[c])])
- self.confs_test[c][np.isnan(self.confs_test[c])] = fill_val
- elif self.confs_scale_info[X_col] == b'numerical':
- fill_val = st.median(self.confs[c][~np.isnan(self.confs[c])])
- self.confs_test[c][np.isnan(self.confs_test[c])] = fill_val
- def PCA_data(self, n_components):
- #X_train_scaled = (self.X - self.X.mean(axis=0)) / self.X.std(axis=0)
- #X_test_scaled = (self.X_test - self.X_test.mean(axis=0)) / self.X_test.std(axis=0)
- X_train_scaled = StandardScaler().fit_transform(self.X)
- X_test_scaled = StandardScaler().fit_transform(self.X_test)
- pca = PCA(n_components=n_components)
- #pca_test = PCA(n_components=n_components)
- X_train_pca = pca.fit_transform(X_train_scaled)
- X_test_pca = pca.transform(X_test_scaled)
- self.X = X_train_pca
- self.X_test = X_test_pca
- def transform_data(self, func):
- '''performs func.transform() (see sklearn structures)
- on either (a) all the data loaded, if performed before train_test_split()
- or on (b) the train+val subset, if performed after train_test_split()'''
- # first check if train_test_split has already been performed
- if self.X is not None:
- self.X = func.transform(self.X)
- self.y = func.transform(self.y)
- for c in self.confs:
- self.confs[c] = func.transform(self.confs[c])
- def print_data_size(self):
- if self.X is None:
- print("data is not loaded yet.. run MLpipeline.load_data()")
- else:
- print(f"Trainval data info:\t X ={self.X.shape} \t y ={self.y.shape} \t confs={self.confs.keys()}")
- if self.X_test is None:
- print("test data is not loaded yet.. run MLpipeline.load_data() and then MLpipeline.train_test_split()")
- else:
- print(f" Test data info:\t X ={self.X_test.shape} \t y ={self.y_test.shape} \t n(trainval)={self.n_samples_tv}, n(test)={self.n_samples_test}")
- def change_input_to_conf(self, c, onehot=True):
- """ Change the inputs of the model to a vector containing a confound.
- Args:
- c (str): The name of the confound from the list loaded in conf_dict.
- Get confounds list available from the dict self.confs.
- onehot (bool, optional): Whether to one-hot encode the new input. This should be
- enabled for linear/logistic regression models. Defaults to True.
- """
- assert self.y_test is not None, "self.train_test_split() should be run before calling self.change_input_to()"
- self.X = self.confs[c].reshape(-1, 1)
- self.X_test = self.confs_test[c].reshape(-1, 1)
- if onehot:
- self.X = OneHotEncoder(sparse=False).fit_transform(self.X)
- self.X_test = OneHotEncoder(sparse=False).fit_transform(self.X_test)
- def change_output_to_conf(self, c):
- """ Change the targets (outputs) of the classifier to a confound vector.
- Args:
- c (str): The name of the confound which is loaded in self.confs,
- Get confounds list using self.confs().
- """
- assert self.y_test is not None, "self.train_test_split() should be run before calling self.change_output_to()"
- self.y = self.confs[c]
- self.y_test = self.confs_test[c]
- @ignore_warnings(category=ConvergenceWarning)
- def run(self, pipe, grid, task_type='classification', scoring_list=None,
- n_splits=5, stratify_by_conf=None, conf_corr_params={},
- permute=0, verbose=2):
- """ The main function to run the classification.
- Args:
- pipe (sklearn Pipeline): a pipeline containing preprocessing steps
- and the classification model.
- grid (dict): A dict of hyperparameter lists that will be passed to
- sklearn's GridSearchCV() function as 'param_grid'
- permute (int, optional): Number of permutations to perform. Defaults to 0 i.e.
- no permutations are performed.
- Returns:
- results (dict): A dictionary containing classification metrics for the
- best parameters.
- best_params (dict): The best parameters found by grid search.
- """
- assert self.X_test is not None, "self.train_test_split() has not been run yet. \
- First split the data into train and test data"
- # create the inner cv folds on the trainval split
- cv = StratifiedKFold(n_splits, shuffle=True, random_state=self.random_state)
- # stratify the inner cv folds by the label and confound groups
- stratify = self.y
- if stratify_by_conf is not None:
- stratify = self.y + 100*self.confs[stratify_by_conf]
- cv_splits = cv.split(self.X, stratify)
- n_jobs = n_splits if (self.parallelize) else None
- # grid search for hyperparameter tuning with the inner cv
- gs = GridSearchCV(estimator=pipe, param_grid=grid, n_jobs=n_jobs,
- cv=cv_splits, scoring=scoring_list,
- return_train_score=True, refit=True, verbose=verbose)
- # fit the estimator on data
- gs = gs.fit(self.X, self.y, **conf_corr_params)
- self.estimator = gs.best_estimator_
- # store scores
- train_score = np.mean(gs.cv_results_["mean_train_score"])
- inner_cv_score = gs.best_score_ # mean cross-validated score of the best_estimator
- val_score = gs.score(self.X_test, self.y_test)
- val_lbls = self.y_test
- val_ids = self.sub_ids_test
- train_ids = self.sub_ids #edit Marija
- if "classification" in task_type:
- # save the predicted probability scores on the test subjects (for trying other metrics later)
- val_probs = np.around(gs.predict_proba(self.X_test), decimals=4)
- # Calculate AUC if label is binary
- roc_auc = np.nan
- if len(np.unique(self.y_test))==2:
- roc_auc = self.func_auc(self.estimator, self.X_test, self.y_test)
- elif "regression" in task_type:
- # save the predicted probability scores on the test subjects (for trying other metrics later)
- val_probs = np.around(gs.predict(self.X_test), decimals=4)
- # Calculate R^2 score if label is continuous
- r2_score = self.func_r2(self.estimator, self.X_test, self.y_test)
- # if permutation is requested, then calculate the test statistic after permuting y
- # TODO: permutation on regression
- results_pt = {}
- if permute:
- with Parallel(n_jobs=n_jobs) as parallel:
- # run parallel jobs on all cores at once
- pt_scores = parallel(delayed(
- MLpipeline._one_permutation)(
- self.X, self.y, self.X_test, self.y_test,
- pipe, grid, fit_params=conf_corr_params, confs_test=self.confs_test,
- n_splits=n_splits, score_func=scoring_list, score_func_auc=self.func_auc,
- # permute_y=False, permute_x=True, permute_y_test=False, permute_x_test=False
- ) for _ in range(permute))
- pt_scores = np.array(pt_scores)
- results_pt = {"permuted_test_score" : pt_scores[:,0].tolist()}
- if not np.isnan(pt_scores[:,1]).all():
- results_pt.update({"permuted_roc_auc": pt_scores[:,1].tolist()})
- results = {
- "model_metric" : scoring_list,
- "innerval_metric" : inner_cv_score,
- "m__params" : gs.best_params_,
- "train_metric" : train_score,
- "val_metric" : val_score,
- "val_ids" : val_ids.tolist(),
- "train_ids": train_ids.tolist(),
- "val_lbls" : val_lbls.tolist(),
- "val_preds" : val_probs.tolist(),
- **results_pt,
- }
- if "classification" in task_type:
- classification = {
- "other_metrics" : dict(roc_auc = roc_auc)
- }
- results.update(classification)
- elif "regression" in task_type:
- regression = {
- "other_metrics" : dict(r2_score = r2_score)
- }
- results.update(regression)
- return results
- @staticmethod
- def sensitivity_score(y_true, y_pred):
- tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
- return tp/(tp+fn)
- @staticmethod
- def specificity_score(y_true, y_pred):
- tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
- return tn/(tn+fp)
- @staticmethod
- def _shuffle(arr, groups=None):
- """
- Permute a vector or array y along axis 0.
- Args:
- arr (numpy array): The to-be-permuted array. The array will be permuted
- along axis 0.
- groups (str): If set, the permutations of y will only happen within members
- of the same group. Get confounds list available from the dict self.confs
- Returns:
- y_permuted: The permuted array y_permuted. """
- if groups is None:
- indices = np.random.permutation(len(arr))
- else:
- indices = np.arange(len(groups))
- for group in np.unique(groups):
- this_mask = (groups == group)
- indices[this_mask] = np.random.permutation(indices[this_mask])
- return arr[indices]
- @staticmethod
- def _one_permutation(X, y, X_test, y_test,
- estimator, grid, fit_params={}, confs_test=None,
- n_splits=5, score_func=None, score_func_auc=get_scorer("roc_auc"),
- permute_test=False):
- """ Run the standard gridsearch+cv pipeline once after permuting X or y
- and evaluate on the test set X_test, y_test
- Args:
- Returns:
- pt_score (float): Balanced accuracy score from permuted samples.
- pt_score_auc (float): ROC AUC score from permuted samples.
- d2_* (float): D2 scores for fitting the PBCC models with different
- independent variables *. See function pbcc().
- TODO: permutation on regression
- """
- # By default X is shuffled rather than the labels y as the relationship
- # between confound and label should be maintained for PBCC. refer Dinga et al., 2020.
- X = MLpipeline._shuffle(X)
- if permute_test:
- X_test = MLpipeline._shuffle(X_test)
- # create the inner crossvalidation folds on the trainval split
- # random state is not fixed because each permutation must be completely random
- cv = StratifiedKFold(n_splits, shuffle=True, random_state=None)
- # grid search for hyperparameter tuning with the inner cv
- gs = GridSearchCV(estimator, param_grid=grid, n_jobs=None,
- cv=cv, scoring=score_func,
- return_train_score=False, refit=True, verbose=0)
- # disable the random_state in the confound correction to get varying scores
- if "conf_corr_cb__cb_by" in fit_params:
- estimator["conf_corr_cb"].random_state = None
- # fit the estimator on data
- gs = gs.fit(X, y, **fit_params)
- pt_score = gs.score(X_test, y_test)
- # calc auc score
- pt_score_auc = np.nan
- if len(np.unique(y)) == 2:
- pt_score_auc = score_func_auc(gs.best_estimator_, X_test, y_test)
- return [pt_score, pt_score_auc]
MLpipeline.py at commit e5546dd, no license · at the source
Overview
- Department of Psychology, Humboldt-Universität zu Berlin,Berlin, Germany
- Hertie Institute for AI in Brain Health, Eberhard Karls Universität Tübingen,Tübingen, Germany
- Department of Psychiatry and Psychotherapy, Charité – Universitätsmedizin Berlin (corporate member of Freie Universität at Berlin, Humboldt-Universität zu Berlin, Berlin Institute of Health),Berlin, Germany
- Department of Medicine, MSB Medical School Berlin,Berlin, Germany
- Department of Psychology, MSB Medical School Berlin,Berlin, Germany
- Department of Psychology, Universität Hamburg,Hamburg, Germany
- Department of Educational Sciences and Psychology, TU Dortmund University,Dortmund, Germany
- Bernstein Center for Computational Neuroscience,Berlin, Germany
- Modelling of Cognitive Processes, Technical University of Berlin,Berlin, Germany
Abstract
Cognitive-behavioral therapy (CBT) is the first-line treatment for obsessive-compulsive disorder (OCD), yet a significant number of patients do not achieve remission or substantial symptom relief. This study aims to enhance the prediction of CBT outcomes in OCD by integrating demographic, clinical, and neuroimaging data using machine learning (ML) models. We conduct a comprehensive analysis on a well-characterized clinical sample, employing a rigorous validation scheme to avoid data leakage, and comparing multiple ML algorithms to minimize bias. Out of four different ML models trained on demographic and clinical data, structural MRI, and resting-state MRI functional connectivity data, no model was able to predict CBT success significantly above chance level in the present sample. Although clinical and demographic data enabled 64%-66% accuracy for predicting remission, this did not reach statistical significance after permutation testing. Pre-treatment symptom severity emerged numerically as the most promising predictor of remission, aligning with previous studies, but did not pass the significance threshold in the present study. Despite efforts to identify neuroimaging predictors, neither functional nor structural MRI features significantly contributed to the prediction models. These findings suggest that robust, individualized brain-based predictions for mental health outcomes remain challenging with the available data and sample size.
Supplementary Information: The online version contains supplementary material available at 10.1038/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 10 matches between paragraphs and lines of code.
ritterlab/EPOC_outcome_prediction
e5546dd2472dedcabd439711beaa84af48403b24, 3 January 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
47 files
- data/
results/ , Python, 219 lines, 2 matchesclassification_MRI_PCA_R emission/ fMRI_OUTRemission_CONFSS ex/ 20240515-1524/ classification_Marija_MR I_PCA_Remission_20240515 _1524.py - data/
results/ , Python, 219 lines, 1 matchclassification_MRI_PCA_R esponse/ fMRI_OUTResponse/ 20240516-1018/ classification_Marija_MR I_PCA_20240516_1018.py - data/
results/ , Python, 222 lines, 1 matchclassification_Remission / MRI_OUTRemission_CONFSSe x/ 20240202-1246/ classification_Marija_MR I_Remission_20240202_124 6.py - data/
results/ , Python, 227 linesclassification_Remission / TAB_OUTRemission_CONFSSe x/ 20240130-1045/ classification_Marija_CO NFSSex_20240130_1045.py - data/
results/ , Python, 228 linesclassification_Remission / TAB_noclin_OUTRemission_ CONFSSex/ 20240409-1437/ classification_Marija_CO NFSSex_20240409_1437.py - data/
results/ , Python, 223 linesclassification_Remission / TABnoclin_sMRI_OUTRemiss ion_CONFSSex/ 20240410-1124/ classification_Marija_MR I_Remission_20240410_112 4.py - data/
results/ , Python, 222 linesclassification_Remission / fMRI_OUTRemission_CONFSS ex/ 20240202-1251/ classification_Marija_MR I_Remission_20240202_125 1.py - data/
results/ , Python, 229 linesclassification_Remission / graph_fMRI_OUTRemission_ CONFSSex/ 20240411-1656/ classification_Marija_CO NFSSex_20240411_1656.py - data/
results/ , Python, 221 linesclassification_Response/ MRI_OUTResponse/ 20240106-1240/ classification_Marija_MR I_20240106_1240.py - data/
results/ , Python, 244 linesclassification_Response/ TAB_OUTResponse/ 20230906-1612/ classification_Marija_20 230906_1612.py - data/
results/ , Python, 245 linesclassification_Response/ TAB_noclin_OUTResponse/ 20240408-1422/ classification_Marija_20 240408_1422.py - data/
results/ , Python, 225 linesclassification_Response/ TABnoclin_sMRI_OUTRespon se/ 20240408-1513/ classification_Marija_MR I_20240408_1513.py - data/
results/ , Python, 221 linesclassification_Response/ fMRI_OUTResponse/ 20240115-1434/ classification_Marija_MR I_20240115_1434.py - data/
results/ , Python, 246 linesclassification_Response/ graph_fMRI_OUTResponse/ 20240411-1553/ classification_Marija_20 240411_1553.py - data/
results/ , Python, 245 linesclassification_Response/ sex_OUTResponse/ 20231117-1530/ classification_Marija_20 231117_1530.py - notebooks/
0_explorations/ , Jupyter, 1,266 linesEPOC_overview.ipynb - notebooks/
1_input_prep/ , Shell, 33 lines1_freesurfer_scripts/ 1_subj_level_reconall_EP OC.sh - notebooks/
1_input_prep/ , Shell, 20 lines1_freesurfer_scripts/ 2_loop_subjs_reconall_EP OC.sh - notebooks/
1_input_prep/ , Shell, 8 lines1_freesurfer_scripts/ 3_parallelize_reconall_E POC.sh - notebooks/
1_input_prep/ , Shell, 47 lines1_freesurfer_scripts/ 4_subj_level_map_BNA_EPO C.sh - notebooks/
1_input_prep/ , Shell, 15 lines1_freesurfer_scripts/ 5_loop_subjs_map_BNA_EPO C.sh - notebooks/
1_input_prep/ , Shell, 35 lines1_freesurfer_scripts/ 6_BNA_to_table_EPOC.sh - notebooks/
1_input_prep/ , Shell, 65 lines2_halfpipe/ submit.slurm.sh - notebooks/
1_input_prep/ , Python, 64 lines, 1 match3_EPOC_fMRI_graph_featur es/ graph_features.py - notebooks/
1_input_prep/ , Jupyter, 36 lines3_EPOC_fMRI_graph_featur es/ run_graph_features.ipynb - notebooks/
1_input_prep/ , Python, 324 lines, 1 match3_EPOC_fMRI_graph_featur es/ weighted_graph_functions .py - notebooks/
1_input_prep/ , Jupyter, 152 lines4_EPOC_create_input_file s/ 1_make_demographic_csv_f iles.ipynb - notebooks/
1_input_prep/ , Jupyter, 82 lines4_EPOC_create_input_file s/ 2_make_sMRI_csv_file.ipy nb - notebooks/
1_input_prep/ , Jupyter, 142 lines4_EPOC_create_input_file s/ 3_make_fMRI_csv_file.ipy nb - notebooks/
1_input_prep/ , Jupyter, 52 lines4_EPOC_create_input_file s/ 4_make_graph_fMRI_csv_fi le.ipynb - notebooks/
1_input_prep/ , Jupyter, 102 lines4_EPOC_create_input_file s/ 5_make_demografic_sMRI_c sv_file.ipynb - notebooks/
1_input_prep/ , Jupyter, 47 lines4_EPOC_create_input_file s/ 6_make_demografic_fMRI_c sv_file.ipynb - notebooks/
1_input_prep/ , Jupyter, 49 lines4_EPOC_create_input_file s/ 7_make_demografic_graph_ fMRI_csv_file.ipynb - notebooks/
1_input_prep/ , Jupyter, 246 lines4_EPOC_create_input_file s/ run_makeMLpipelineFiles. ipynb - notebooks/
2_run_model/ , Jupyter, 25 linesrun_epoc_ML.ipynb - notebooks/
3_results/ , Jupyter, 186 linesp_values_plotting/ calculate_p_values.ipynb - notebooks/
3_results/ , Jupyter, 532 lines, 1 matchp_values_plotting/ results_figures_EPOC.ipy nb - src/
MLpipeline/ , Python, 424 lines, 2 matchesMLpipeline.py - src/
MLpipeline/ , Python, 289 linesconfounds.py - src/
MLpipeline/ , Python, 312 lines, 1 matchrunMLpipelines.py - src/
configs.py , Python, 13 lines - src/
pipeline_configs/ , Python, 222 linesclassification_MRI_PCA_R emission.py - src/
pipeline_configs/ , Python, 222 linesclassification_MRI_PCA_R esponse.py - src/
pipeline_configs/ , Python, 227 linesclassification_Remission .py - src/
pipeline_configs/ , Python, 228 linesclassification_Response. py - src/
prep_h5/ , Python, 134 linesmakeMLpipelineFiles.py - README.md, Text, 35 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 46 scripts, each with its path and the digest of its content;
- 10 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
Data is not available to be shared due to data protection agreements. All code used for data analysis is provided under https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 12 authors, 10 keywords, 13 MeSH terms, 1 funder, 50 references.
Cite
This paper
Tochadse, M., Klawohn, J., Kaufmann, C., Grützmann, R., Riesel, A., Heinzel, S., Rane, R. P., da Silva, S. P., Wellan, S., Gijsen, S., Ritter, K., & Kathmann, N. (2026). Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder. Scientific reports, 16(1), 24542. https://
BibTeX
@article{tochadse2026eva
author = {Tochadse, Marija and Klawohn, Julia and Kaufmann, Christian and Grützmann, Rosa and Riesel, Anja and Heinzel, Stephan and Rane, Roshan Prakash and da Silva, Sofia Pereira and Wellan, Sarah and Gijsen, Sam and Ritter, Kerstin and Kathmann, Norbert},
title = {{Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder}},
journal = {Scientific reports},
year = {2026},
month = aug,
volume = {16},
number = {1},
pages = {24542},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/
url = {https://
pmid = {42570963},
pmcid = {PMC13452815}
}
RIS
TY - JOUR
AU - Tochadse, Marija
AU - Klawohn, Julia
AU - Kaufmann, Christian
AU - Grützmann, Rosa
AU - Riesel, Anja
AU - Heinzel, Stephan
AU - Rane, Roshan Prakash
AU - da Silva, Sofia Pereira
AU - Wellan, Sarah
AU - Gijsen, Sam
AU - Ritter, Kerstin
AU - Kathmann, Norbert
TI - Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/
VL - 16
IS - 1
SP - 24542
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder",
"container-title": "Scientific reports",
"author": [
{
"family": "Tochadse",
"given": "Marija"
},
{
"family": "Klawohn",
"given": "Julia"
},
{
"family": "Kaufmann",
"given": "Christian"
},
{
"family": "Grützmann",
"given": "Rosa"
},
{
"family": "Riesel",
"given": "Anja"
},
{
"family": "Heinzel",
"given": "Stephan"
},
{
"family": "Rane",
"given": "Roshan Prakash"
},
{
"family": "da Silva",
"given": "Sofia Pereira"
},
{
"family": "Wellan",
"given": "Sarah"
},
{
"family": "Gijsen",
"given": "Sam"
},
{
"family": "Ritter",
"given": "Kerstin"
},
{
"family": "Kathmann",
"given": "Norbert"
}
],
"container-title-short":
"volume": "16",
"issue": "1",
"page": "24542",
"DOI": "10.1038/
"PMID": "42570963",
"PMCID": "PMC13452815",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
8
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.nicl.2026.104012 [code]
- Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.Journal: NeuroImage. ClinicalIn common: imbalanced-learn, XGBoost, FreeSurfer, 8 other tools, structural MRI / diffusion, other condition, 1 reference
- [2] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: imbalanced-learn, XGBoost, NetworkX, 8 other tools
- [3] doi:10.7554/elife.110588 [code]
- Opening the black box toward a modular approach to spike sorting.Journal: eLifeIn common: imbalanced-learn, XGBoost, NetworkX, 7 other tools
- [4] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: imbalanced-learn, XGBoost, NetworkX, 7 other tools
- [5] doi:10.1038/s43856-026-01518-5 [code]
- Machine learning-based identification of abnormal functional connectivity in obesity across different metabolic states.Journal: Communications medicineIn common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, clinical / translational, other condition
- [6] doi:10.1038/s41598-026-56688-y [code]
- On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.Journal: Scientific reportsIn common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, 1 reference
- [7] doi:10.1016/j.isci.2026.115329 [code]
- Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.Journal: iScienceIn common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, clinical / translational, other condition
- [8] doi:10.1038/s41398-026-04081-8 [code]
- Functional system-specific brain aging across the Alzheimer's disease continuum.Journal: Translational psychiatryIn common: imbalanced-learn, FreeSurfer, statsmodels, 6 other tools, structural MRI / diffusion, clinical / translational
- [9] doi:10.1038/s42003-026-09794-6 [code]
- The role of dorsal anterior cingulate cortex in dynamic attitude changes in naturalistic settings.Journal: Communications biologyIn common: imbalanced-learn, NetworkX, statsmodels, 6 other tools, 1 reference
- [10] doi:10.64898/2026.03.04.709586 [code]
- Mapping Higher-Order Topology in OCD Brain Networks with Hodge LaplacianJournal: bioRxiv (preprint)In common: NetworkX, h5py, statsmodels, 6 other tools, other condition, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 46 scripts, and 10 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e6f036c41b3a21fb…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
