OSCR

Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder.

Code ↔ Paper

10 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 10 matches
  1. [1] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ notebooks/1_input_prep/3_EPOC_fMRI_graph_features/weighted_graph_functions.py, lines 306–323 · score 0.88 · global efficiency, clustering coefficient, correlation matrices, graph features, density, sigma
  2. [2] § Methods › Analysis ↔ src/MLpipeline/runMLpipelines.py, lines 50–195 · score 0.83 · random seeds, inner CV, outer CV, ML models, machine, stratified
  3. [3] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ notebooks/1_input_prep/3_EPOC_fMRI_graph_features/graph_features.py, lines 12–37 · score 0.82 · global efficiency, clustering coefficient, graph features, density, sigma, transitivity
  4. [4] § Methods › Analysis ↔ src/MLpipeline/MLpipeline.py, lines 236–342 · score 0.74 · inner CV folds, cross validation, stratified, splits, tuned, binary
  5. [5] § Methods › Feature importance ↔ data/results/classification_MRI_PCA_Remission/fMRI_OUTRemission_CONFSSex/20240515-1524/classification_Marija_MRI_PCA_Remission_20240515_1524.py, lines 13–83 · score 0.66 · linear SVM, gradient boosting, logistic regression, fitting, kernel, RBF
  6. [6] § Methods › Feature importance ↔ data/results/classification_MRI_PCA_Response/fMRI_OUTResponse/20240516-1018/classification_Marija_MRI_PCA_20240516_1018.py, lines 13–83 · score 0.66 · linear SVM, gradient boosting, logistic regression, fitting, kernel, RBF
  7. [7] § Methods › Neuroimaging data (acquisition & preprocessing) ↔ src/MLpipeline/MLpipeline.py, lines 236–342 · score 0.63 · cross validated, scikit-learn, correlation, preprocessing, hyperparameter, inner
  8. [8] § Results › ML predictions ↔ notebooks/3_results/p_values_plotting/results_figures_EPOC.ipynb, lines 427–440 · score 0.56 · sMRI, Tabular OCD, fMRI, multimodal, PCA, graph
  9. [9] § Methods › Analysis ↔ data/results/classification_MRI_PCA_Remission/fMRI_OUTRemission_CONFSSex/20240515-1524/classification_Marija_MRI_PCA_Remission_20240515_1524.py, lines 13–83 · score 0.52 · gradient boosting, logistic regression, GB, LR, kernel, linear
  10. [10] § Methods › Analysis ↔ data/results/classification_Remission/MRI_OUTRemission_CONFSSex/20240202-1246/classification_Marija_MRI_Remission_20240202_1246.py, lines 13–83 · score 0.52 · gradient boosting, logistic regression, GB, LR, kernel, linear

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 424 lines · 19 KB · no license · 2 matches

  1. from os.path import join
  2. import h5py
  3. import numpy as np
  4. import pandas as pd
  5. import statistics as st
  6. from datetime import datetime
  7. from joblib import Parallel, delayed
  8. from sklearn.model_selection import GridSearchCV
  9. from sklearn.metrics import get_scorer, confusion_matrix
  10. from sklearn.model_selection import train_test_split, StratifiedKFold
  11. from sklearn.preprocessing import OneHotEncoder, StandardScaler
  12. from sklearn.utils._testing import ignore_warnings
  13. from sklearn.exceptions import ConvergenceWarning
  14. from sklearn.decomposition import PCA
  15. class MLpipeline:
  16. def __init__(self, parallelize=True, random_state=None):
  17. """ Define parameters for the MLLearn object.
  18. Args:
  19. "h5file" (str): Full path to the HDF5 file.
  20. "conf_list" (list of str): the list of confound names to load from the hdf5.
  21. random_state (int): A random state to get reproducible results
  22. """
  23. self.random_state = random_state
  24. self.parallelize = parallelize
  25. self.confs = {}
  26. self.X_test = None
  27. self.y_test = None
  28. self.confs_test = {}
  29. self.sub_ids = None
  30. self.sub_ids_test = None
  31. # These are defined to calculate later.
  32. self.func_auc = get_scorer("roc_auc")
  33. self.func_r2 = get_scorer('r2')
  34. def load_data(self, f, X="X", y="y", confs=[], group_confs=False):
  35. """ Load the data from the HDF5 file,
  36. The neuroimaging data/ input data is saved under "X", the label under "y" and
  37. the confounds from the dict self.confs
  38. If group_confs==True, groups all confounds into one numpy array and saves in
  39. self.confs["group"]
  40. """
  41. if isinstance(f, dict):
  42. df = f
  43. else:
  44. df = h5py.File(f, "r")
  45. self.X = np.array(df[X])
  46. self.y = np.array(df[y])
  47. self.sub_ids = np.array(df["i"])
  48. if df.get('scale_info'): #edit Marija
  49. self.scale_info = list(df["scale_info"]) #edit Marija
  50. if df.get('confs_scale_info'): #edit Marija
  51. self.confs_scale_info = list(df["confs_scale_info"]) #edit Marija
  52. # self.X needs to be flattened as sklearn expects 2D input.
  53. if self.X.ndim > 2:
  54. self.X = self.X.reshape([self.X.shape[0], np.prod(self.X[0].shape)])
  55. elif self.X.ndim == 1:
  56. self.X = self.X.reshape(-1,1) # sklearn expects 2D input
  57. # Load the confounding variables into a dictionary, if there are any
  58. for c in confs:
  59. if c in df.keys():
  60. v = np.array(df[c])
  61. self.confs[c] = v
  62. if group_confs:
  63. if "group" not in self.confs:
  64. self.confs["group"] = v
  65. else:
  66. self.confs["group"] = v + 100*self.confs["group"]
  67. else:
  68. print(f"[WARN]: confound {c} not found in hdf5 file.")
  69. def train_test_split(self, test_idx=[], test_size=0.25, stratify_by_conf=None):
  70. """ Split the data into a test set (_test) and a train+val set.
  71. Args:
  72. stratify_group (str, optional): A string that indicates a confound name.
  73. Get confounds list available from the dict self.confs.
  74. If a confound is selected, subjects will be stratified according
  75. to confound and outcome label. Defaults to None, in which case
  76. subjects are only stratified according to the outcome label.
  77. """
  78. # if test_idx are already provided then dont generate the test_idx yourself
  79. if not len(test_idx):
  80. # by default, always stratify by label
  81. stratify = self.y
  82. if stratify_by_conf is not None:
  83. stratify = self.y + 100*self.confs[stratify_by_conf]
  84. _, test_idx = train_test_split(range(len(self.X)),
  85. test_size=test_size,
  86. stratify=stratify,
  87. shuffle=True,
  88. random_state=self.random_state)
  89. test_mask = np.zeros(len(self.X), dtype=bool)
  90. test_mask[test_idx] = True
  91. self.X_test = self.X[test_mask]
  92. self.y_test = self.y[test_mask]
  93. self.sub_ids_test = self.sub_ids[test_mask]
  94. self.X = self.X[~test_mask]
  95. self.y = self.y[~test_mask]
  96. self.sub_ids = self.sub_ids[~test_mask]
  97. for c in self.confs:
  98. self.confs_test[c] = self.confs[c][test_mask]
  99. for c in self.confs:
  100. self.confs[c] = self.confs[c][~test_mask]
  101. self.n_samples_tv = len(self.y)
  102. self.n_samples_test = len(self.y_test)
  103. # fill missing values edit Marija
  104. def fill_missing_vals(self):
  105. """ Fill in missing values in outer CV loop."""
  106. for X_col in range(self.X.shape[1]):
  107. if np.isnan(self.X[:,X_col]).any(axis=0):
  108. if self.scale_info[X_col] == b'categorical':
  109. fill_val = st.mode(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
  110. self.X[:,X_col][np.isnan(self.X[:,X_col])] = fill_val
  111. elif self.scale_info[X_col] == b'numerical':
  112. fill_val = st.median(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
  113. #fill_val = np.nanmean(self.X[:,X_col])
  114. self.X[:,X_col][np.isnan(self.X[:,X_col])] = fill_val
  115. if np.isnan(self.X_test[:,X_col]).any(axis=0):
  116. if self.scale_info[X_col] == b'categorical':
  117. fill_val = st.mode(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
  118. self.X_test[:,X_col][np.isnan(self.X_test[:,X_col])] = fill_val
  119. elif self.scale_info[X_col] == b'numerical':
  120. fill_val = st.median(self.X[:,X_col][~np.isnan(self.X[:,X_col])])
  121. #fill_val = np.nanmean(self.X[:,X_col])
  122. self.X_test[:,X_col][np.isnan(self.X_test[:,X_col])] = fill_val
  123. if self.confs:
  124. for X_col, c in enumerate(self.confs):
  125. if np.isnan(self.confs[c]).any(axis=0):
  126. if self.confs_scale_info[X_col] == b'categorical':
  127. fill_val = st.mode(self.confs[c][~np.isnan(self.confs[c])])
  128. self.confs[c][np.isnan(self.confs[c])] = fill_val
  129. elif self.confs_scale_info[X_col] == b'numerical':
  130. fill_val = st.median(self.confs[c][~np.isnan(self.confs[c])])
  131. self.confs[c][np.isnan(self.confs[c])] = fill_val
  132. if np.isnan(self.confs_test[c]).any(axis=0):
  133. if self.confs_scale_info[X_col] == b'categorical':
  134. fill_val = st.mode(self.confs[c][~np.isnan(self.confs[c])])
  135. self.confs_test[c][np.isnan(self.confs_test[c])] = fill_val
  136. elif self.confs_scale_info[X_col] == b'numerical':
  137. fill_val = st.median(self.confs[c][~np.isnan(self.confs[c])])
  138. self.confs_test[c][np.isnan(self.confs_test[c])] = fill_val
  139. def PCA_data(self, n_components):
  140. #X_train_scaled = (self.X - self.X.mean(axis=0)) / self.X.std(axis=0)
  141. #X_test_scaled = (self.X_test - self.X_test.mean(axis=0)) / self.X_test.std(axis=0)
  142. X_train_scaled = StandardScaler().fit_transform(self.X)
  143. X_test_scaled = StandardScaler().fit_transform(self.X_test)
  144. pca = PCA(n_components=n_components)
  145. #pca_test = PCA(n_components=n_components)
  146. X_train_pca = pca.fit_transform(X_train_scaled)
  147. X_test_pca = pca.transform(X_test_scaled)
  148. self.X = X_train_pca
  149. self.X_test = X_test_pca
  150. def transform_data(self, func):
  151. '''performs func.transform() (see sklearn structures)
  152. on either (a) all the data loaded, if performed before train_test_split()
  153. or on (b) the train+val subset, if performed after train_test_split()'''
  154. # first check if train_test_split has already been performed
  155. if self.X is not None:
  156. self.X = func.transform(self.X)
  157. self.y = func.transform(self.y)
  158. for c in self.confs:
  159. self.confs[c] = func.transform(self.confs[c])
  160. def print_data_size(self):
  161. if self.X is None:
  162. print("data is not loaded yet.. run MLpipeline.load_data()")
  163. else:
  164. print(f"Trainval data info:\t X ={self.X.shape} \t y ={self.y.shape} \t confs={self.confs.keys()}")
  165. if self.X_test is None:
  166. print("test data is not loaded yet.. run MLpipeline.load_data() and then MLpipeline.train_test_split()")
  167. else:
  168. print(f" Test data info:\t X ={self.X_test.shape} \t y ={self.y_test.shape} \t n(trainval)={self.n_samples_tv}, n(test)={self.n_samples_test}")
  169. def change_input_to_conf(self, c, onehot=True):
  170. """ Change the inputs of the model to a vector containing a confound.
  171. Args:
  172. c (str): The name of the confound from the list loaded in conf_dict.
  173. Get confounds list available from the dict self.confs.
  174. onehot (bool, optional): Whether to one-hot encode the new input. This should be
  175. enabled for linear/logistic regression models. Defaults to True.
  176. """
  177. assert self.y_test is not None, "self.train_test_split() should be run before calling self.change_input_to()"
  178. self.X = self.confs[c].reshape(-1, 1)
  179. self.X_test = self.confs_test[c].reshape(-1, 1)
  180. if onehot:
  181. self.X = OneHotEncoder(sparse=False).fit_transform(self.X)
  182. self.X_test = OneHotEncoder(sparse=False).fit_transform(self.X_test)
  183. def change_output_to_conf(self, c):
  184. """ Change the targets (outputs) of the classifier to a confound vector.
  185. Args:
  186. c (str): The name of the confound which is loaded in self.confs,
  187. Get confounds list using self.confs().
  188. """
  189. assert self.y_test is not None, "self.train_test_split() should be run before calling self.change_output_to()"
  190. self.y = self.confs[c]
  191. self.y_test = self.confs_test[c]
  192. @ignore_warnings(category=ConvergenceWarning)
  193. def run(self, pipe, grid, task_type='classification', scoring_list=None,
  194. n_splits=5, stratify_by_conf=None, conf_corr_params={},
  195. permute=0, verbose=2):
  196. """ The main function to run the classification.
  197. Args:
  198. pipe (sklearn Pipeline): a pipeline containing preprocessing steps
  199. and the classification model.
  200. grid (dict): A dict of hyperparameter lists that will be passed to
  201. sklearn's GridSearchCV() function as 'param_grid'
  202. permute (int, optional): Number of permutations to perform. Defaults to 0 i.e.
  203. no permutations are performed.
  204. Returns:
  205. results (dict): A dictionary containing classification metrics for the
  206. best parameters.
  207. best_params (dict): The best parameters found by grid search.
  208. """
  209. assert self.X_test is not None, "self.train_test_split() has not been run yet. \
  210. First split the data into train and test data"
  211. # create the inner cv folds on the trainval split
  212. cv = StratifiedKFold(n_splits, shuffle=True, random_state=self.random_state)
  213. # stratify the inner cv folds by the label and confound groups
  214. stratify = self.y
  215. if stratify_by_conf is not None:
  216. stratify = self.y + 100*self.confs[stratify_by_conf]
  217. cv_splits = cv.split(self.X, stratify)
  218. n_jobs = n_splits if (self.parallelize) else None
  219. # grid search for hyperparameter tuning with the inner cv
  220. gs = GridSearchCV(estimator=pipe, param_grid=grid, n_jobs=n_jobs,
  221. cv=cv_splits, scoring=scoring_list,
  222. return_train_score=True, refit=True, verbose=verbose)
  223. # fit the estimator on data
  224. gs = gs.fit(self.X, self.y, **conf_corr_params)
  225. self.estimator = gs.best_estimator_
  226. # store scores
  227. train_score = np.mean(gs.cv_results_["mean_train_score"])
  228. inner_cv_score = gs.best_score_ # mean cross-validated score of the best_estimator
  229. val_score = gs.score(self.X_test, self.y_test)
  230. val_lbls = self.y_test
  231. val_ids = self.sub_ids_test
  232. train_ids = self.sub_ids #edit Marija
  233. if "classification" in task_type:
  234. # save the predicted probability scores on the test subjects (for trying other metrics later)
  235. val_probs = np.around(gs.predict_proba(self.X_test), decimals=4)
  236. # Calculate AUC if label is binary
  237. roc_auc = np.nan
  238. if len(np.unique(self.y_test))==2:
  239. roc_auc = self.func_auc(self.estimator, self.X_test, self.y_test)
  240. elif "regression" in task_type:
  241. # save the predicted probability scores on the test subjects (for trying other metrics later)
  242. val_probs = np.around(gs.predict(self.X_test), decimals=4)
  243. # Calculate R^2 score if label is continuous
  244. r2_score = self.func_r2(self.estimator, self.X_test, self.y_test)
  245. # if permutation is requested, then calculate the test statistic after permuting y
  246. # TODO: permutation on regression
  247. results_pt = {}
  248. if permute:
  249. with Parallel(n_jobs=n_jobs) as parallel:
  250. # run parallel jobs on all cores at once
  251. pt_scores = parallel(delayed(
  252. MLpipeline._one_permutation)(
  253. self.X, self.y, self.X_test, self.y_test,
  254. pipe, grid, fit_params=conf_corr_params, confs_test=self.confs_test,
  255. n_splits=n_splits, score_func=scoring_list, score_func_auc=self.func_auc,
  256. # permute_y=False, permute_x=True, permute_y_test=False, permute_x_test=False
  257. ) for _ in range(permute))
  258. pt_scores = np.array(pt_scores)
  259. results_pt = {"permuted_test_score" : pt_scores[:,0].tolist()}
  260. if not np.isnan(pt_scores[:,1]).all():
  261. results_pt.update({"permuted_roc_auc": pt_scores[:,1].tolist()})
  262. results = {
  263. "model_metric" : scoring_list,
  264. "innerval_metric" : inner_cv_score,
  265. "m__params" : gs.best_params_,
  266. "train_metric" : train_score,
  267. "val_metric" : val_score,
  268. "val_ids" : val_ids.tolist(),
  269. "train_ids": train_ids.tolist(),
  270. "val_lbls" : val_lbls.tolist(),
  271. "val_preds" : val_probs.tolist(),
  272. **results_pt,
  273. }
  274. if "classification" in task_type:
  275. classification = {
  276. "other_metrics" : dict(roc_auc = roc_auc)
  277. }
  278. results.update(classification)
  279. elif "regression" in task_type:
  280. regression = {
  281. "other_metrics" : dict(r2_score = r2_score)
  282. }
  283. results.update(regression)
  284. return results
  285. @staticmethod
  286. def sensitivity_score(y_true, y_pred):
  287. tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
  288. return tp/(tp+fn)
  289. @staticmethod
  290. def specificity_score(y_true, y_pred):
  291. tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel()
  292. return tn/(tn+fp)
  293. @staticmethod
  294. def _shuffle(arr, groups=None):
  295. """
  296. Permute a vector or array y along axis 0.
  297. Args:
  298. arr (numpy array): The to-be-permuted array. The array will be permuted
  299. along axis 0.
  300. groups (str): If set, the permutations of y will only happen within members
  301. of the same group. Get confounds list available from the dict self.confs
  302. Returns:
  303. y_permuted: The permuted array y_permuted. """
  304. if groups is None:
  305. indices = np.random.permutation(len(arr))
  306. else:
  307. indices = np.arange(len(groups))
  308. for group in np.unique(groups):
  309. this_mask = (groups == group)
  310. indices[this_mask] = np.random.permutation(indices[this_mask])
  311. return arr[indices]
  312. @staticmethod
  313. def _one_permutation(X, y, X_test, y_test,
  314. estimator, grid, fit_params={}, confs_test=None,
  315. n_splits=5, score_func=None, score_func_auc=get_scorer("roc_auc"),
  316. permute_test=False):
  317. """ Run the standard gridsearch+cv pipeline once after permuting X or y
  318. and evaluate on the test set X_test, y_test
  319. Args:
  320. Returns:
  321. pt_score (float): Balanced accuracy score from permuted samples.
  322. pt_score_auc (float): ROC AUC score from permuted samples.
  323. d2_* (float): D2 scores for fitting the PBCC models with different
  324. independent variables *. See function pbcc().
  325. TODO: permutation on regression
  326. """
  327. # By default X is shuffled rather than the labels y as the relationship
  328. # between confound and label should be maintained for PBCC. refer Dinga et al., 2020.
  329. X = MLpipeline._shuffle(X)
  330. if permute_test:
  331. X_test = MLpipeline._shuffle(X_test)
  332. # create the inner crossvalidation folds on the trainval split
  333. # random state is not fixed because each permutation must be completely random
  334. cv = StratifiedKFold(n_splits, shuffle=True, random_state=None)
  335. # grid search for hyperparameter tuning with the inner cv
  336. gs = GridSearchCV(estimator, param_grid=grid, n_jobs=None,
  337. cv=cv, scoring=score_func,
  338. return_train_score=False, refit=True, verbose=0)
  339. # disable the random_state in the confound correction to get varying scores
  340. if "conf_corr_cb__cb_by" in fit_params:
  341. estimator["conf_corr_cb"].random_state = None
  342. # fit the estimator on data
  343. gs = gs.fit(X, y, **fit_params)
  344. pt_score = gs.score(X_test, y_test)
  345. # calc auc score
  346. pt_score_auc = np.nan
  347. if len(np.unique(y)) == 2:
  348. pt_score_auc = score_func_auc(gs.best_estimator_, X_test, y_test)
  349. return [pt_score, pt_score_auc]

MLpipeline.py at commit e5546dd, no license · at the source

Overview

Authors: Marija Tochadse1,2,3, Julia Klawohn1,4, Christian Kaufmann1, Rosa Grützmann1,5, Anja Riesel1,6, Stephan Heinzel1,7, Roshan Prakash Rane1,2,3, Sofia Pereira da Silva8,9, Sarah Wellan3, Sam Gijsen2,3, Kerstin Ritter2,3,8, Norbert Kathmann1
ORCID iDs: Marija Tochadse
  1. Department of Psychology, Humboldt-Universität zu Berlin,Berlin, Germany
  2. Hertie Institute for AI in Brain Health, Eberhard Karls Universität Tübingen,Tübingen, Germany
  3. Department of Psychiatry and Psychotherapy, Charité – Universitätsmedizin Berlin (corporate member of Freie Universität at Berlin, Humboldt-Universität zu Berlin, Berlin Institute of Health),Berlin, Germany
  4. Department of Medicine, MSB Medical School Berlin,Berlin, Germany
  5. Department of Psychology, MSB Medical School Berlin,Berlin, Germany
  6. Department of Psychology, Universität Hamburg,Hamburg, Germany
  7. Department of Educational Sciences and Psychology, TU Dortmund University,Dortmund, Germany
  8. Bernstein Center for Computational Neuroscience,Berlin, Germany
  9. Modelling of Cognitive Processes, Technical University of Berlin,Berlin, Germany
Journal: Scientific reports, volume 16, issue 1, article 24542
Dates: received 24 June 2025; accepted 24 July 2026; published online 8 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41598-026-64405-y · PMID 42570963 · PMCID PMC13452815 · OpenAlex W7201971269
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), other condition (population), clinical / translational (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Graphs, Preprocessing
Keywords: Cognitive-behavioral therapy (CBT), Obsessive–compulsive disorder (OCD), Machine Learning (ML), Neuroimaging predictors, Outcome prediction, Functional connectivity, Neuroscience, Psychology, Human behaviour, Predictive markers
MeSH: Cognitive Behavioral Therapy*, Neuroimaging*, Obsessive-Compulsive Disorder*, Adult, Brain, Female, Humans, Machine Learning, Magnetic Resonance Imaging, Male, Predictive Learning Models, Treatment Outcome, Young Adult (* major topic)
Topic: Obsessive-Compulsive Spectrum Disorders (Clinical Psychology, Psychology), according to OpenAlex
Funding: Charité - Universitätsmedizin Berlin (3093)
Citations: not cited yet (Europe PMC); 57 references in the paper

Abstract

Cognitive-behavioral therapy (CBT) is the first-line treatment for obsessive-compulsive disorder (OCD), yet a significant number of patients do not achieve remission or substantial symptom relief. This study aims to enhance the prediction of CBT outcomes in OCD by integrating demographic, clinical, and neuroimaging data using machine learning (ML) models. We conduct a comprehensive analysis on a well-characterized clinical sample, employing a rigorous validation scheme to avoid data leakage, and comparing multiple ML algorithms to minimize bias. Out of four different ML models trained on demographic and clinical data, structural MRI, and resting-state MRI functional connectivity data, no model was able to predict CBT success significantly above chance level in the present sample. Although clinical and demographic data enabled 64%-66% accuracy for predicting remission, this did not reach statistical significance after permutation testing. Pre-treatment symptom severity emerged numerically as the most promising predictor of remission, aligning with previous studies, but did not pass the significance threshold in the present study. Despite efforts to identify neuroimaging predictors, neither functional nor structural MRI features significantly contributed to the prediction models. These findings suggest that robust, individualized brain-based predictions for mental health outcomes remain challenging with the available data and sample size.

Supplementary Information: The online version contains supplementary material available at 10.1038/s41598-026-64405-y.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 10 matches between paragraphs and lines of code.

ritterlab/EPOC_outcome_prediction

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: e5546dd2472dedcabd439711beaa84af48403b24, 3 January 2025
Languages: Python (26), Jupyter (13), Shell (7)
Size: 66 files, 46 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, 13 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: scikit-learn (22 files), imbalanced-learn (19 files), XGBoost (19 files), NumPy (17 files), pandas (17 files), Matplotlib (12 files), SciPy (12 files), seaborn (11 files), h5py (5 files), FreeSurfer (3 files), NetworkX (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
47 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 46 scripts, each with its path and the digest of its content;
  • 10 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

Data is not available to be shared due to data protection agreements. All code used for data analysis is provided under https://github.com/ritterlab/EPOC_outcome_prediction.git.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 12 authors, 10 keywords, 13 MeSH terms, 1 funder, 50 references.

Cite

This paper

Tochadse, M., Klawohn, J., Kaufmann, C., Grützmann, R., Riesel, A., Heinzel, S., Rane, R. P., da Silva, S. P., Wellan, S., Gijsen, S., Ritter, K., & Kathmann, N. (2026). Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder. Scientific reports, 16(1), 24542. https://doi.org/10.1038/s41598-026-64405-y

BibTeX

@article{tochadse2026evaluating,
author = {Tochadse, Marija and Klawohn, Julia and Kaufmann, Christian and Grützmann, Rosa and Riesel, Anja and Heinzel, Stephan and Rane, Roshan Prakash and da Silva, Sofia Pereira and Wellan, Sarah and Gijsen, Sam and Ritter, Kerstin and Kathmann, Norbert},
title = {{Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder}},
journal = {Scientific reports},
year = {2026},
month = aug,
volume = {16},
number = {1},
pages = {24542},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-64405-y},
url = {https://doi.org/10.1038/s41598-026-64405-y},
pmid = {42570963},
pmcid = {PMC13452815}
}

RIS

TY - JOUR
AU - Tochadse, Marija
AU - Klawohn, Julia
AU - Kaufmann, Christian
AU - Grützmann, Rosa
AU - Riesel, Anja
AU - Heinzel, Stephan
AU - Rane, Roshan Prakash
AU - da Silva, Sofia Pereira
AU - Wellan, Sarah
AU - Gijsen, Sam
AU - Ritter, Kerstin
AU - Kathmann, Norbert
TI - Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/08/08
VL - 16
IS - 1
SP - 24542
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-64405-y
UR - https://doi.org/10.1038/s41598-026-64405-y
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-64405-y",
"type": "article-journal",
"title": "Evaluating clinical and neuroimaging predictors for cognitive-behavioral therapy outcome in obsessive-compulsive disorder",
"container-title": "Scientific reports",
"author": [
{
"family": "Tochadse",
"given": "Marija"
},
{
"family": "Klawohn",
"given": "Julia"
},
{
"family": "Kaufmann",
"given": "Christian"
},
{
"family": "Grützmann",
"given": "Rosa"
},
{
"family": "Riesel",
"given": "Anja"
},
{
"family": "Heinzel",
"given": "Stephan"
},
{
"family": "Rane",
"given": "Roshan Prakash"
},
{
"family": "da Silva",
"given": "Sofia Pereira"
},
{
"family": "Wellan",
"given": "Sarah"
},
{
"family": "Gijsen",
"given": "Sam"
},
{
"family": "Ritter",
"given": "Kerstin"
},
{
"family": "Kathmann",
"given": "Norbert"
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "24542",
"DOI": "10.1038/s41598-026-64405-y",
"PMID": "42570963",
"PMCID": "PMC13452815",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-64405-y",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
8
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: imbalanced-learn, XGBoost, FreeSurfer, 8 other tools, structural MRI / diffusion, other condition, 1 reference
[2] doi:10.1016/j.isci.2026.116825 [code]
Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.
Journal: iScience
In common: imbalanced-learn, XGBoost, NetworkX, 8 other tools
[3] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: imbalanced-learn, XGBoost, NetworkX, 7 other tools
[4] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: imbalanced-learn, XGBoost, NetworkX, 7 other tools
[5] doi:10.1038/s43856-026-01518-5 [code]
Machine learning-based identification of abnormal functional connectivity in obesity across different metabolic states.
Journal: Communications medicine
In common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, clinical / translational, other condition
[6] doi:10.1038/s41598-026-56688-y [code]
On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.
Journal: Scientific reports
In common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, 1 reference
[7] doi:10.1016/j.isci.2026.115329 [code]
Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.
Journal: iScience
In common: imbalanced-learn, XGBoost, statsmodels, 6 other tools, clinical / translational, other condition
[8] doi:10.1038/s41398-026-04081-8 [code]
Functional system-specific brain aging across the Alzheimer's disease continuum.
Journal: Translational psychiatry
In common: imbalanced-learn, FreeSurfer, statsmodels, 6 other tools, structural MRI / diffusion, clinical / translational
[9] doi:10.1038/s42003-026-09794-6 [code]
The role of dorsal anterior cingulate cortex in dynamic attitude changes in naturalistic settings.
Journal: Communications biology
In common: imbalanced-learn, NetworkX, statsmodels, 6 other tools, 1 reference
[10] doi:10.64898/2026.03.04.709586 [code]
Mapping Higher-Order Topology in OCD Brain Networks with Hodge Laplacian
Journal: bioRxiv (preprint)
In common: NetworkX, h5py, statsmodels, 6 other tools, other condition, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.