OSCR

Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group

Code ↔ Paper

17 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 17 matches
  1. [1] § Methods › Image acquisition and processing ↔ scripts/01_data_preparation/make_VETSA_dataset.py, lines 123–136 · score 0.84 · cerebellar cortex, cerebellar white matter, lateral ventricles, accumbens, amygdala, caudate
  2. [2] § Methods › Image acquisition and processing ↔ scripts/01_data_preparation/make_EMC_dataset.py, lines 22–33 · score 0.84 · cerebellar cortex, cerebellar white matter, lateral ventricles, accumbens, amygdala, caudate
  3. [3] § Methods › Model training and internal evaluation ↔ scripts/05_figures_tables/new_plot_cross_model_results.py, lines 210–251 · score 0.83 · ridge regression, Pearson correlation, random forest, Spearman correlation, Linear regression, XGBoost
  4. [4] § Methods › SHAP-based subgroup analysis ↔ scripts/02_model_training/autogluon/plot_SHAP_cluster_SHAP-IQ_AutoGluon_SHIP.py, lines 631–677 · score 0.74 · Gaussian mixture, spectral clustering, silhouette score, elbow, gap, embed
  5. [5] § Methods › Model training and internal evaluation ↔ scripts/02_model_training/xgboost/code_references.py, lines 995–1074 · score 0.74 · Optuna optimized, fitted transformers, linear regression, inner, XGBoost, outer
  6. [6] § Methods › SHAP-based subgroup analysis ↔ scripts/02_model_training/autogluon/plot_test_SHAP_cluster_SHAP-IQ_AutoGluon_SHIP.py, lines 949–1035 · score 0.70 · Gaussian mixture models, silhouette score, elbow, gap, spectral, embed
  7. [7] § Methods › Machine learning models ↔ scripts/05_figures_tables/new_plot_cross_model_results.py, lines 210–251 · score 0.69 · ridge regression, random forest, linear regression, XGBoost, rbf, SVM
  8. [8] § Methods › Model training and internal evaluation ↔ scripts/03_out_of_sample_validation/lib/AutoGluon_pipeline.py, lines 242–347 · score 0.69 · TabularPredictor, best quality, properties, bagged, preset, stack
  9. [9] § Methods › Model training and internal evaluation ↔ src/enigma_sleep_cognition/AutoGluon_pipeline.py, lines 242–349 · score 0.69 · TabularPredictor, best quality, properties, bagged, preset, stack
  10. [10] § Methods › Study participants and measurements › Memory test score ↔ scripts/01_data_preparation/make_VETSA_dataset.py, lines 39–70 · score 0.68 · Digit Span Backward, Digit Span Forward, backward raw, Sequencing, Letter, scores
  11. [11] § Methods › Out-of-cohort validation in independent cohorts ↔ scripts/05_figures_tables/new_plot_SHAP_brain.py, lines 59–115 · score 0.63 · Li ge, San Diego, XGBoost, lich, Pittsburgh, AutoGluon
  12. [12] § Methods › Out-of-cohort validation in independent cohorts ↔ scripts/05_figures_tables/new_plot_SHAP_brain_no_DK.py, lines 59–115 · score 0.63 · Li ge, San Diego, XGBoost, lich, Pittsburgh, AutoGluon
  13. [13] § Results › Overview of the framework and study cohort ↔ scripts/05_figures_tables/new_plot_SHAP_brain.py, lines 59–115 · score 0.60 · Li ge, San Diego, sleep duration, depressive scores, VETSA, lich
  14. [14] § Results › Overview of the framework and study cohort ↔ scripts/05_figures_tables/new_plot_SHAP_brain_no_DK.py, lines 59–115 · score 0.60 · Li ge, San Diego, sleep duration, depressive scores, VETSA, lich
  15. [15] § Methods › Study participants and measurements › Stroop test score ↔ scripts/01_data_preparation/make_VETSA_dataset.py, lines 39–70 · score 0.55 · Stroop Interference Norm, word, score
  16. [16] § Methods › Model training and internal evaluation ↔ src/enigma_sleep_cognition/AutoGluon_pipeline.py, lines 73–179 · score 0.53 · squared error, Linear regression, absolute error, refitted, AutoGluon, rooted
  17. [17] § Methods › Machine learning models ↔ src/enigma_sleep_cognition/AutoGluon_pipeline_log.py, lines 246–358 · score 0.53 · linear regression, XGBoost, stacking, tree, tabular, preprocessing

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 338 lines · 15 KB · MIT · 3 matches

  1. import os
  2. import sys
  3. sys.path.append('/data/project/sleep_ENIGMA_Cognition/Codes/ENIGMA_Sleep_Cognitive/Code/lib/')
  4. import utils
  5. import pandas as pd
  6. import matplotlib.pyplot as plt
  7. import seaborn as sns
  8. # %%
  9. raw_data_save_path = '/data/project/sleep_ENIGMA_Cognition/Codes/ENIGMA_Sleep_Cognitive/Data/raw_datasets/San Diago/ENIGMA_Sleep_Masoud_2024_10_14/files/'
  10. data_save_path = '/data/project/sleep_ENIGMA_Cognition/Codes/ENIGMA_Sleep_Cognitive/Data/'
  11. # %%
  12. df_demo_VETSA = pd.read_csv(raw_data_save_path + 'ESR_01_ENIGMA_Sleep_nonMRIdata_revised.csv')
  13. df_area_DK_VETSA = pd.read_csv(raw_data_save_path + 'ESR_02_UCSD_area_DK.csv')
  14. df_area_Schaefer_VETSA = pd.read_csv(raw_data_save_path + 'ESR_03_UCSD_area_Schaefer.csv')
  15. df_subcor_VETSA = pd.read_csv(raw_data_save_path + 'ESR_04_UCSD_subcortical_volume.csv')
  16. df_thickness_DK_VETSA = pd.read_csv(raw_data_save_path + 'ESR_05_UCSD_thickness_DK.csv')
  17. df_thickness_Schaefer_VETSA = pd.read_csv(raw_data_save_path + 'ESR_06_UCSD_thickness_Schaefer.csv')
  18. # %%
  19. # for all dataframes, rename 'CID' to 'Sub_ID', then sort by 'Sub_ID'
  20. df_demo_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  21. df_demo_VETSA.sort_values(by='Sub_ID', inplace=True)
  22. df_area_DK_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  23. df_area_DK_VETSA.sort_values(by='Sub_ID', inplace=True)
  24. df_area_Schaefer_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  25. df_area_Schaefer_VETSA.sort_values(by='Sub_ID', inplace=True)
  26. df_subcor_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  27. df_subcor_VETSA.sort_values(by='Sub_ID', inplace=True)
  28. df_thickness_DK_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  29. df_thickness_DK_VETSA.sort_values(by='Sub_ID', inplace=True)
  30. df_thickness_Schaefer_VETSA.rename(columns={'CID': 'Sub_ID'}, inplace=True)
  31. df_thickness_Schaefer_VETSA.sort_values(by='Sub_ID', inplace=True)
  32. # %%
  33. """
  34. Change column names for df_demo_VETSA:
  35. 'cesdtot_V3' to 'Depression_score'
  36. 'BMI_V3' to 'BMI'
  37. 'Age_V3' to 'Age_at_Scan'
  38. 'apoe2024' to 'APOE4'
  39. 'HRSSLEEP_V3' to 'Self_Sleep_Dur'
  40. 'sleepeff_V3' to 'Self_Sleep_Eff'
  41. 'DSFRAW_V3' to 'Digit Span Forward Raw'
  42. 'DSBRAW_V3' to 'Digit Span Backward Raw'
  43. 'STRWRAW_V3' to 'Stroop Raw Word Score'
  44. 'STRCRAW_V3' to 'Stroop Raw Color Score'
  45. 'STRCWRAW_V3' to 'Stroop Raw Color-Word Score'
  46. 'DSTOT_V3p' to 'Digit Span Total Trials Passed'
  47. 'LNTOT_V3p' to 'Letter-Number Sequencing Total Score'
  48. 'STRIT_V3p' to 'Stroop Interference Norm-Based T-Score'
  49. """
  50. df_demo_VETSA.rename(columns={'cesdtot_V3': 'Depression_score'}, inplace=True)
  51. df_demo_VETSA.rename(columns={'BMI_V3': 'BMI'}, inplace=True)
  52. df_demo_VETSA.rename(columns={'AGE_V3': 'Age_at_Scan'}, inplace=True)
  53. df_demo_VETSA.rename(columns={'apoe2024': 'APOE4'}, inplace=True)
  54. df_demo_VETSA.rename(columns={'HRSSLEEP_V3': 'Self_Sleep_Dur'}, inplace=True)
  55. df_demo_VETSA.rename(columns={'sleepeff_V3': 'Self_Sleep_Eff'}, inplace=True)
  56. df_demo_VETSA.rename(columns={'DSFRAW_V3': 'Digit Span Forward Raw'}, inplace=True)
  57. df_demo_VETSA.rename(columns={'DSBRAW_V3': 'Digit Span Backward Raw'}, inplace=True)
  58. df_demo_VETSA.rename(columns={'STRWRAW_V3': 'Stroop Raw Word Score'}, inplace=True)
  59. df_demo_VETSA.rename(columns={'STRCRAW_V3': 'Stroop Raw Color Score'}, inplace=True)
  60. df_demo_VETSA.rename(columns={'STRCWRAW_V3': 'Stroop Raw Color-Word Score'}, inplace=True)
  61. df_demo_VETSA.rename(columns={'DSTOT_V3p': 'Digit Span Total Trials Passed'}, inplace=True)
  62. df_demo_VETSA.rename(columns={'LNTOT_V3p': 'Letter-Number Sequencing Total Score'}, inplace=True)
  63. df_demo_VETSA.rename(columns={'STRIT_V3p': 'Stroop Interference Norm-Based T-Score'}, inplace=True)
  64. # %%
  65. # for df_demo_VETSA, df_thickness_DK_VETSA, df_thickness_Schaefer_VETSA, df_area_DK_VETSA, df_area_Schaefer_VETSA, df_subcor_VETSA, check whether 'Sub_ID' column has unique values
  66. # Create a set of Sub_IDs for each DataFrame
  67. sub_ids = {
  68. "df_demo_VETSA": set(df_demo_VETSA['Sub_ID']),
  69. "df_thickness_DK_VETSA": set(df_thickness_DK_VETSA['Sub_ID']),
  70. "df_thickness_Schaefer_VETSA": set(df_thickness_Schaefer_VETSA['Sub_ID']),
  71. "df_area_DK_VETSA": set(df_area_DK_VETSA['Sub_ID']),
  72. "df_area_Schaefer_VETSA": set(df_area_Schaefer_VETSA['Sub_ID']),
  73. "df_subcor_VETSA": set(df_subcor_VETSA['Sub_ID']),
  74. }
  75. # Find common Sub_IDs across all DataFrames
  76. common_sub_ids = set.intersection(*sub_ids.values())
  77. print(f"Common Sub_IDs across all DataFrames: {len(common_sub_ids)}")
  78. # Find unique Sub_IDs for each DataFrame
  79. unique_sub_ids = {key: value - common_sub_ids for key, value in sub_ids.items()}
  80. for df_name, unique_ids in unique_sub_ids.items():
  81. print(f"Unique Sub_IDs in {df_name}: {len(unique_ids)}")
  82. # Check which DataFrames have more or less subjects
  83. sub_id_counts = {df_name: len(ids) for df_name, ids in sub_ids.items()}
  84. print("Number of Sub_IDs in each DataFrame:")
  85. for df_name, count in sub_id_counts.items():
  86. print(f"{df_name}: {count}")
  87. # Identify differences between DataFrames
  88. for df1, ids1 in sub_ids.items():
  89. for df2, ids2 in sub_ids.items():
  90. if df1 != df2:
  91. extra_in_df1 = ids1 - ids2
  92. extra_in_df2 = ids2 - ids1
  93. print(f"Subjects in {df1} but not in {df2}: {len(extra_in_df1)}")
  94. print(f"Subjects in {df2} but not in {df1}: {len(extra_in_df2)}")
  95. # Identify and display the specific Sub_ID values
  96. for df1, ids1 in sub_ids.items():
  97. for df2, ids2 in sub_ids.items():
  98. if df1 != df2:
  99. extra_in_df1 = ids1 - ids2 # Sub_IDs in df1 but not in df2
  100. extra_in_df2 = ids2 - ids1 # Sub_IDs in df2 but not in df1
  101. if extra_in_df1:
  102. print(f"Sub_IDs in {df1} but not in {df2} ({len(extra_in_df1)}): {sorted(extra_in_df1)}")
  103. if extra_in_df2:
  104. print(f"Sub_IDs in {df2} but not in {df1} ({len(extra_in_df2)}): {sorted(extra_in_df2)}")
  105. # %%
  106. # remove subject 51011 from df_demo_VETSA
  107. df_demo_VETSA = df_demo_VETSA[df_demo_VETSA['Sub_ID'] != 51011]
  108. # %%
  109. """
  110. Rename of imaging tables
  111. """
  112. label_subcortical = ['Sub_ID', 'Left-Lateral-Ventricle', 'Left-Inf-Lat-Vent', 'Left-Cerebellum-White-Matter',
  113. 'Left-Cerebellum-Cortex', 'Left-Thalamus-Proper', 'Left-Caudate', 'Left-Putamen', 'Left-Pallidum',
  114. '3rd-Ventricle', '4th-Ventricle', 'Brain-Stem', 'Left-Hippocampus', 'Left-Amygdala', 'CSF',
  115. 'Left-Accumbens-area', 'Left-VentralDC', 'Left-vessel', 'Right-Lateral-Ventricle',
  116. 'Right-Inf-Lat-Vent', 'Right-Cerebellum-White-Matter', 'Right-Cerebellum-Cortex',
  117. 'Right-Thalamus-Proper', 'Right-Caudate', 'Right-Putamen', 'Right-Pallidum', 'Right-Hippocampus',
  118. 'Right-Amygdala', 'Right-Accumbens-area', 'Right-VentralDC', 'Right-vessel', '5th-Ventricle',
  119. 'CC_Posterior', 'CC_Mid_Posterior', 'CC_Central', 'CC_Mid_Anterior', 'CC_Anterior',
  120. 'EstimatedTotalIntraCranialVol']
  121. df_subcor_VETSA = df_subcor_VETSA[label_subcortical]
  122. # %%
  123. df_subcor_VETSA = df_subcor_VETSA.rename(columns={'Left-Thalamus-Proper': 'Left-Thalamus',
  124. 'Right-Thalamus-Proper': 'Right-Thalamus',
  125. 'subject_ID': 'Sub_ID'})
  126. # %%
  127. # rename cortical thickness
  128. # Schaefer
  129. filtered_columns = [col for col in df_thickness_Schaefer_VETSA.columns if
  130. col not in ['Sub_ID', 'BrainSegVolNotVent', 'eTIV']]
  131. renamed_schaefer_ct_df_columns = [col.replace('lh_7Networks_', '').replace('rh_7Networks_', '') for col in
  132. filtered_columns]
  133. # DK
  134. renamed_dk_ct_df_columns = [col for col in df_thickness_DK_VETSA.columns if
  135. col not in ['subject_ID', 'BrainSegVolNotVent', 'eTIV']]
  136. # rename surface area
  137. # Schaefer
  138. filtered_columns = [col for col in df_area_Schaefer_VETSA.columns if
  139. col not in ['subject_ID', 'BrainSegVolNotVent', 'eTIV', 'lh_WhiteSurfArea_area',
  140. 'rh_WhiteSurfArea_area']]
  141. renamed_schaefer_sa_df_columns = [col.replace('lh_7Networks_', '').replace('rh_7Networks_', '') for col in
  142. filtered_columns]
  143. # DK
  144. renamed_dk_sa_df_columns = [col for col in df_area_DK_VETSA.columns if
  145. col not in ['subject_ID', 'BrainSegVolNotVent', 'eTIV']]
  146. # %%
  147. # apply the renaming and filtering
  148. # Rename cortical thickness columns for Schaefer
  149. filtered_ct_columns_schaefer = [col for col in df_thickness_Schaefer_VETSA.columns if
  150. col not in ['Sub_ID', 'BrainSegVolNotVent', 'eTIV']]
  151. renamed_ct_columns_schaefer = [col.replace('lh_7Networks_', '').replace('rh_7Networks_', '') for col in
  152. filtered_ct_columns_schaefer]
  153. df_thickness_Schaefer_VETSA.rename(columns=dict(zip(filtered_ct_columns_schaefer, renamed_ct_columns_schaefer)), inplace=True)
  154. df_thickness_Schaefer_VETSA = df_thickness_Schaefer_VETSA[['Sub_ID'] + renamed_ct_columns_schaefer]
  155. # Rename cortical thickness columns for DK
  156. filtered_ct_columns_dk = [col for col in df_thickness_DK_VETSA.columns if
  157. col not in ['Sub_ID', 'BrainSegVolNotVent', 'eTIV']]
  158. # No specific replacement given for DK; keeping original names
  159. df_thickness_DK_VETSA = df_thickness_DK_VETSA[['Sub_ID'] + filtered_ct_columns_dk]
  160. # Rename surface area columns for Schaefer
  161. filtered_sa_columns_schaefer = [col for col in df_area_Schaefer_VETSA.columns if
  162. col not in ['Sub_ID', 'BrainSegVolNotVent', 'eTIV', 'lh_WhiteSurfArea_area', 'rh_WhiteSurfArea_area']]
  163. renamed_sa_columns_schaefer = [col.replace('lh_7Networks_', '').replace('rh_7Networks_', '') for col in
  164. filtered_sa_columns_schaefer]
  165. df_area_Schaefer_VETSA.rename(columns=dict(zip(filtered_sa_columns_schaefer, renamed_sa_columns_schaefer)), inplace=True)
  166. df_area_Schaefer_VETSA = df_area_Schaefer_VETSA[['Sub_ID'] + renamed_sa_columns_schaefer]
  167. # Rename surface area columns for DK
  168. filtered_sa_columns_dk = [col for col in df_area_DK_VETSA.columns if
  169. col not in ['Sub_ID', 'BrainSegVolNotVent', 'eTIV']]
  170. # No specific replacement given for DK; keeping original names
  171. df_area_DK_VETSA = df_area_DK_VETSA[['Sub_ID'] + filtered_sa_columns_dk]
  172. # %%
  173. # concatenate all dataframes by 'Sub_ID' column. Order of dataframes is df_demo_VETSA, df_thickness_DK_VETSA, df_thickness_Schaefer_VETSA, df_area_DK_VETSA, df_area_Schaefer_VETSA, df_subcor_VETSA
  174. # Sequentially merge all DataFrames by 'Sub_ID'
  175. df_combined = df_demo_VETSA.copy() # Start with df_demo_VETSA as the base
  176. # Merge with each subsequent DataFrame
  177. df_combined = df_combined.merge(df_thickness_DK_VETSA, on='Sub_ID', how='outer')
  178. df_combined = df_combined.merge(df_thickness_Schaefer_VETSA, on='Sub_ID', how='outer')
  179. df_combined = df_combined.merge(df_area_DK_VETSA, on='Sub_ID', how='outer')
  180. df_combined = df_combined.merge(df_area_Schaefer_VETSA, on='Sub_ID', how='outer')
  181. df_combined = df_combined.merge(df_subcor_VETSA, on='Sub_ID', how='outer')
  182. # %%
  183. # Transform the APOE4 column: 1 for carrier (if '4' is present), 2 for non-carrier
  184. # Safely transform the APOE4 column, handling NaN values
  185. def transform_apoe4(value):
  186. try:
  187. # Check if value is NaN
  188. if pd.isna(value):
  189. return None # Or any other representation for missing values, e.g., 'Unknown'
  190. # Check if '4' is present in the string
  191. return '1' if '4' in str(value) else '2'
  192. except Exception as e:
  193. print(f"Error processing value: {value}, Error: {e}")
  194. return None # Handle unexpected cases gracefully
  195. # Apply the function
  196. df_combined['APOE4'] = df_combined['APOE4'].apply(transform_apoe4)
  197. # %%
  198. # for df_combined, make a new column 'Stroop_Test', which is same as 'Stroop Interference Norm-Based T-Score'
  199. df_combined['Stroop_Test'] = df_combined['Stroop Interference Norm-Based T-Score']
  200. # %%
  201. # Get the list of columns
  202. columns = df_combined.columns.tolist()
  203. # Remove 'Stroop_Test' from the columns list
  204. columns.remove('Stroop_Test')
  205. # Find the index of 'Stroop Interference Norm-Based T-Score'
  206. index = columns.index('Stroop Interference Norm-Based T-Score')
  207. # Insert 'Stroop_Test' immediately after 'Stroop Interference Norm-Based T-Score'
  208. columns.insert(index + 1, 'Stroop_Test')
  209. # Reorder the DataFrame
  210. df_combined = df_combined[columns]
  211. # %%
  212. # description of 'Digit Span Forward Raw', 'Digit Span Backward Raw'
  213. df_combined['Digit Span Forward Raw'] = df_combined['Digit Span Forward Raw'].astype(float)
  214. df_combined['Digit Span Backward Raw'] = df_combined['Digit Span Backward Raw'].astype(float)
  215. print(df_combined['Digit Span Forward Raw'].describe())
  216. print(df_combined['Digit Span Backward Raw'].describe())
  217. # %%
  218. # add a column 'Memory_Test_Digit'. The value of this column is the sum of 'Digit Span Forward Raw' and 'Digit Span Backward Raw' then divided by 30
  219. df_combined['Memory_Test_Digit'] = (df_combined['Digit Span Forward Raw'] + df_combined['Digit Span Backward Raw']) / 30
  220. print(df_combined['Memory_Test_Digit'].describe())
  221. # %%
  222. # Get the list of columns
  223. columns = df_combined.columns.tolist()
  224. # Remove 'Stroop_Test' from the columns list
  225. columns.remove('Memory_Test_Digit')
  226. # Find the index of 'Stroop Interference Norm-Based T-Score'
  227. index = columns.index('Stroop_Test')
  228. # Insert 'Stroop_Test' immediately after 'Stroop Interference Norm-Based T-Score'
  229. columns.insert(index + 1, 'Memory_Test_Digit')
  230. # Reorder the DataFrame
  231. df_combined = df_combined[columns]
  232. # %%
  233. print(df_combined['Letter-Number Sequencing Total Score'].describe())
  234. # %%
  235. df_combined['Memory_Test_Letter'] = df_combined['Letter-Number Sequencing Total Score'] / 21
  236. # %%
  237. # Get the list of columns
  238. columns = df_combined.columns.tolist()
  239. # Remove 'Stroop_Test' from the columns list
  240. columns.remove('Memory_Test_Letter')
  241. # Find the index of 'Stroop Interference Norm-Based T-Score'
  242. index = columns.index('Memory_Test_Digit')
  243. # Insert 'Stroop_Test' immediately after 'Stroop Interference Norm-Based T-Score'
  244. columns.insert(index + 1, 'Memory_Test_Letter')
  245. # Reorder the DataFrame
  246. df_combined = df_combined[columns]
  247. # %%
  248. # add 'SEX' column to the dataframe, all values are 1
  249. df_combined['SEX'] = '1'
  250. # %%
  251. # save the df_combined to a csv file
  252. # df_combined.to_csv(data_save_path + 'VETSA_dataset_renamed.csv')
  253. # %%
  254. # load the saved csv file
  255. df_combined = pd.read_csv(data_save_path + 'VETSA_dataset_renamed.csv')
  256. # %%
  257. # print range of Stroop
  258. print(f"Range of Stroop_Test: {df_combined['Stroop_Test'].min()} - {df_combined['Stroop_Test'].max()}")
  259. # plot histogram of Stroop_Test
  260. plt.figure(figsize=(8, 6))
  261. sns.histplot(df_combined['Stroop_Test'].dropna(), bins=30, kde=True)
  262. plt.title('Histogram of Stroop_Test')
  263. plt.xlabel('Stroop_Test Score')
  264. plt.ylabel('Frequency')
  265. plt.show()
  266. min_val = df_combined['Stroop_Test'].min() # 20.0
  267. max_val = df_combined['Stroop_Test'].max() # 65.250882948
  268. # Transform the Stroop_Test so that smaller values correspond to better performance
  269. df_combined['Stroop_Test'] = (max_val + min_val) - df_combined['Stroop_Test']
  270. # Now the range will be inverted.
  271. print(f"New range of Stroop_Test: {df_combined['Stroop_Test'].min()} - {df_combined['Stroop_Test'].max()}")
  272. # plot histogram of transformed Stroop_Test
  273. plt.figure(figsize=(8, 6))
  274. sns.histplot(df_combined['Stroop_Test'].dropna(), bins=30, kde=True)
  275. plt.title('Histogram of Transformed Stroop_Test')
  276. plt.xlabel('Transformed Stroop_Test Score')
  277. plt.ylabel('Frequency')
  278. plt.show()
  279. # %%
  280. # save back
  281. # df_combined.to_csv(data_save_path + 'VETSA_dataset_renamed.csv')

make_VETSA_dataset.py at commit dbfb1ca, under MIT · at the source

Overview

Authors: Hanwen Bi1,2, Torbjörn Åkerstedt3,4, Robin Bülow5, Michele Deantoni6, Alexander Drzezga7,8,9, Jeremy A. Elman10,11, David Elmenhorst9,7, Eva-Maria Elmenhorst12,13, Ralf Ewert14, Christine Fennema-Notestine10,11, Fabio Ferrarelli15, Stefan Frenzel16, Charlotte von Gall17, Sarah Genon1,2, Hans J. Grabe16,18, Phuong Thuy Nguyen Ho19, Sanne J.W. Hoepel20, Felix Hoffstaedter1,2, Agustin Ibanez21,22,23,24,25, Neda Jahanshad26
and 29 other authorsAhmadreza Keihani15, Vincent Küppers7,1, Wen Liu1,2, Ahmad Mayeli15, Nasrin Mortazavi6, Julia Neitzel19,20,27, Gustav Nilsonne3,28, Matthew S. Panizzon10,11, Julia S. Rupp15, Amin Saberi1,29,2, Christina Schmidt6, Kai Spiegelhalder30, Beate Stubbe14, Sandra Tamm3, Sophia I. Thomopoulos26, Paul M. Thompson26, Sofie L. Valk1,29,2, Gilles Vandewalle6, Tina Thi Vo-Eckerle10,11, Henry Völzke31,32, Laura K. Waite1, Joseph Wexler33,3, Katharina Wittfeld16, Kaustubh R. Patil1,2, Antoine Weihs16,18, Simon B. Eickhoff1,2, Federico Raimondo1,2, Masoud Tahmasian1,2,7, ENIGMA-Sleep Working Group
33 affiliations
  1. Institute of Neuroscience and Medicine, Brain and Behaviour (INM-7), Research Center Jülich, Jülich, Germany
  2. Institute of Systems Neuroscience, Medical Faculty and University Hospital Düsseldorf, Heinrich Heine University, Düsseldorf, Germany
  3. Department of Clinical Neuroscience, Karolinska Institutet, Stockholm, Sweden
  4. Department of Psychology, Stockholm University, Stockholm, Sweden
  5. Institute of Diagnostic Radiology and Neuroradiology, University Medicine Greifswald, Greifswald, Germany
  6. GIGA-CRC-Human Imaging, University of Liège, Liège, Belgium
  7. Department of Nuclear Medicine, Faculty of Medicine and University Hospital Cologne, University of Cologne, Cologne, Germany
  8. German Center for Neurodegenerative Diseases (DZNE), Bonn-Cologne, Germany
  9. Institute for Neuroscience and Medicine, Molecular Organization of the Brain (INM-2), Research Center Jülich, Jülich, Germany
  10. Center for Behavior Genetics of Aging, Department of Psychiatry, University of California San Diego
  11. Department of Psychiatry, University of California San Diego
  12. Department of Sleep and Human Factors Research, German Aerospace Center, Cologne, Germany
  13. Institute for Occupational, Social and Environmental Medicine, Medical Faculty, RWTH Aachen University, Aachen, Germany
  14. Department of Internal Medicine B-Cardiology, Pneumology, Infectious Diseases, Intensive Care Medicine, University Medicine Greifswald
  15. Department of Psychiatry, University of Pittsburgh, Pittsburgh, PA, USA
  16. Department of Psychiatry and Psychotherapy, University Medicine Greifswald, Greifswald, Germany
  17. Institute of Anatomy II, Medical Faculty, Heinrich Heine University, Düsseldorf, Germany
  18. German Center for Neurodegenerative Diseases (DZNE), Site Rostock/Greifswald, Greifswald, Germany
  19. Department of Radiology and Nuclear Medicine, Erasmus University Medical Centre, Rotterdam, the Netherlands
  20. Department of Epidemiology, Erasmus MC University Medical Center Rotterdam, Rotterdam, Netherlands
  21. Latin American Brain Health Institute (BrainLat), Universidad Adolfo Ibáñez, Santiago, Chile
  22. Global Brain Health Institute, Trinity College Dublin, Dublin, Ireland
  23. Department of Biophysics, School of Medicine, Istanbul Medipol University, 34815, Istanbul, Türkiye
  24. Barcelonaβeta Brain Research Center (BBRC), Pasqual Maragall Foundation, 08005, Barcelona, Spain
  25. Cognitive Neuroscience Center (CNC), Universidad de San Andrés, Buenos Aires, Argentina
  26. Imaging Genetics Center, Mark and Mary Stevens Neuroimaging and Informatics Institute, Keck School of Medicine of the University of Southern California, Los Angeles, CA, USA
  27. Department of Epidemiology, Harvard T. H. Chan School of Public Health, Boston, Massachusetts
  28. International Globally Distributed Organization for Research and Education (IGDORE), Stockholm, Sweden
  29. Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany
  30. Department of Psychiatry and Psychotherapy, Medical Center – University of Freiburg, Faculty of Medicine, University of Freiburg, Germany
  31. Institute for Community Medicine, SHIP/Clinical-Epidemiological Research, University Medicine Greifswald, Greifswald, Germany
  32. German Centre for Cardiovascular Research (DZHK), Partner Site Greifswald, Greifswald, Germany
  33. Department of Psychology, Stanford University, Stanford, CA, USA
Dates: published online 9 June 2026
Type: Preprint · Language: English
License: CC BY
Identifiers: DOI 10.21203/rs.3.rs-9914920/v1 · OpenAlex W7164000357
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing
Topic: Sleep and related disorders (Experimental and Cognitive Psychology, Psychology), according to OpenAlex
Citations: not cited yet (Europe PMC); 81 references in the paper

Abstract

Group-level studies have highlighted the roles of aging, poor sleep, and brain atrophy in cognitive performance (CP) but have overlooked inter-individual variability. We predict CP from feature sets (demographic, subjective/objective sleep parameters, and regional brain morphometry) using multisite ENIGMA-Sleep data (n = 2,372). Linear and non-linear machine learning models were trained on the largest cohort (n = 845), and the best-performing models were validated on independent cohorts. Subsequently, based on the best-performing model on the largest cohort, we characterized feature importance and interactions across all cohorts. We observed that a combination of demographic, sleep, and brain parameters moderately predicted CP, with age emerging as the key predictor. Model explanations further suggested that age was the primary driver of prediction models, while sleep played a smaller role that varied across subgroups. These findings endorsed inter-individual variability and complex interaction between aging, sleep, brain, and CP.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.

OSF nhbkq

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 0 files, 0 scripts
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source: osf.io/nhbkq/

harveybi/ENIGMA-Sleep-Cognitive-Performance

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: dbfb1cae6e8de025b7fcdb82581695a57961f47b, 11 June 2026
Languages: Python (228), Shell (87)
Size: 427 files, 315 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (environment-analysis.yml, environment-brainviz.yml, pyproject.toml, requirements.txt, scripts/03_out_of_sample_validation/EMC/requirements_AutoGluon.txt, scripts/03_out_of_sample_validation/EMC/requirements_XGBoost.txt), documentation
Not found: CITATION.cff, tests, continuous integration
Tools: pandas (188 files), SHAP (107 files), Matplotlib (77 files), NumPy (75 files), seaborn (64 files), scikit-learn (56 files), SciPy (40 files), XGBoost (25 files), UMAP (14 files), statsmodels (10 files), CAT12 (8 files), FreeSurfer (6 files), Pingouin (4 files), neuromaps (3 files), NiBabel (3 files), Nilearn (3 files), statannotations (3 files), BrainSpace (2 files), netneurotools (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
317 files

Code availability

This study was preregistered on the Open Science Framework (OSF, https://osf.io/nhbkq/). All analysis tools are open-source in this study based on Python 3.10.18 using NumPy 2.1.3, pandas 2.3.3, SciPy 1.15.3, scikit-learn 1.7.2, XGBoost 3.0.5, AutoGluon 1.2.0, SHAP 0.49.1, SHAP-IQ 1.3.1, Skope-Rules 1.0.1, and UMAP-learn 0.5.7. Brain visualizations were generated based on Python 3.11.14 with NumPy 1.26.4, nilearn 0.12.1, nibabel 5.3.2, neuromaps 0.0.5, brainspace 0.1.20, and surfplot 0.2.0. All scripts for analyses in this study and Python environment specifications are provided in the GitHub repository (https://github.com/harveybi/ENIGMA-Sleep-Cognitive-Performance).

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 315 scripts, each with its path and the digest of its content;
  • 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The data used in this study were collected independently at participating sites and analyzed following the standardized ENIGMA protocols. Due to participant privacy considerations, ethical restrictions, and site-level data-sharing agreements, individual-level data cannot be made publicly available. Data requests will be reviewed by the respective site and must comply with local ethics approvals and data-use agreements. For the Greifswald SHIP-Trend cohort, data access applications should be submitted through the SHIP transfer portal at https://transfer.ship-med.uni-greifswald.de/FAIRequest/?lang=en and will be reviewed by the SHIP committee. For the San Diego cohort, instructions for data access requests are available on the Vietnam Era Twin Study of Aging (VETSA) website at https://psychiatry.ucsd.edu/research/programs-centers/vetsa/researchers.html. De-identified VETSA Wave 1, 2, and 3 data are publicly available through the National Archive of Computerized Data on Aging at https://www.icpsr.umich.edu/web/ICPSR/studies/38836. For the Stockholm cohort, anonymized data are available at https://openneuro.org/datasets/ds000201.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Language: n/a → en
  • Funding: added Wellcome Trust

Version 1, 27 September 2026: the first record

Recorded: type, journal, dates, 49 authors, 76 references.

Cite

This paper

Bi, H., Åkerstedt, T., Bülow, R., Deantoni, M., Drzezga, A., Elman, J. A., Elmenhorst, D., Elmenhorst, E.-M., Ewert, R., Fennema-Notestine, C., Ferrarelli, F., Frenzel, S., von Gall, C., Genon, S., Grabe, H. J., Nguyen Ho, P. T., Hoepel, S. J., Hoffstaedter, F., Ibanez, A., . . . ENIGMA-Sleep Working Group. (2026). Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group. Research Square (preprint). https://doi.org/10.21203/rs.3.rs-9914920/v1

BibTeX

@article{bi2026prediction,
author = {Bi, Hanwen and Åkerstedt, Torbjörn and Bülow, Robin and Deantoni, Michele and Drzezga, Alexander and Elman, Jeremy A. and Elmenhorst, David and Elmenhorst, Eva-Maria and Ewert, Ralf and Fennema-Notestine, Christine and Ferrarelli, Fabio and Frenzel, Stefan and von Gall, Charlotte and Genon, Sarah and Grabe, Hans J. and Nguyen Ho, Phuong Thuy and Hoepel, Sanne J.W. and Hoffstaedter, Felix and Ibanez, Agustin and Jahanshad, Neda and Keihani, Ahmadreza and Küppers, Vincent and Liu, Wen and Mayeli, Ahmad and Mortazavi, Nasrin and Neitzel, Julia and Nilsonne, Gustav and Panizzon, Matthew S. and Rupp, Julia S. and Saberi, Amin and Schmidt, Christina and Spiegelhalder, Kai and Stubbe, Beate and Tamm, Sandra and Thomopoulos, Sophia I. and Thompson, Paul M. and Valk, Sofie L. and Vandewalle, Gilles and Vo-Eckerle, Tina Thi and Völzke, Henry and Waite, Laura K. and Wexler, Joseph and Wittfeld, Katharina and Patil, Kaustubh R. and Weihs, Antoine and Eickhoff, Simon B. and Raimondo, Federico and Tahmasian, Masoud and {ENIGMA-Sleep Working Group}},
title = {{Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group}},
journal = {Research Square (preprint)},
year = {2026},
month = jun,
publisher = {Research Square},
issn = {2693-5015},
doi = {10.21203/rs.3.rs-9914920/v1},
url = {https://doi.org/10.21203/rs.3.rs-9914920/v1}
}

RIS

TY - JOUR
AU - Bi, Hanwen
AU - Åkerstedt, Torbjörn
AU - Bülow, Robin
AU - Deantoni, Michele
AU - Drzezga, Alexander
AU - Elman, Jeremy A.
AU - Elmenhorst, David
AU - Elmenhorst, Eva-Maria
AU - Ewert, Ralf
AU - Fennema-Notestine, Christine
AU - Ferrarelli, Fabio
AU - Frenzel, Stefan
AU - von Gall, Charlotte
AU - Genon, Sarah
AU - Grabe, Hans J.
AU - Nguyen Ho, Phuong Thuy
AU - Hoepel, Sanne J.W.
AU - Hoffstaedter, Felix
AU - Ibanez, Agustin
AU - Jahanshad, Neda
AU - Keihani, Ahmadreza
AU - Küppers, Vincent
AU - Liu, Wen
AU - Mayeli, Ahmad
AU - Mortazavi, Nasrin
AU - Neitzel, Julia
AU - Nilsonne, Gustav
AU - Panizzon, Matthew S.
AU - Rupp, Julia S.
AU - Saberi, Amin
AU - Schmidt, Christina
AU - Spiegelhalder, Kai
AU - Stubbe, Beate
AU - Tamm, Sandra
AU - Thomopoulos, Sophia I.
AU - Thompson, Paul M.
AU - Valk, Sofie L.
AU - Vandewalle, Gilles
AU - Vo-Eckerle, Tina Thi
AU - Völzke, Henry
AU - Waite, Laura K.
AU - Wexler, Joseph
AU - Wittfeld, Katharina
AU - Patil, Kaustubh R.
AU - Weihs, Antoine
AU - Eickhoff, Simon B.
AU - Raimondo, Federico
AU - Tahmasian, Masoud
AU - ENIGMA-Sleep Working Group
TI - Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group
T2 - Research Square (preprint)
J2 - Res Sq
PY - 2026
DA - 2026/06/09
SN - 2693-5015
PB - Research Square
DO - 10.21203/rs.3.rs-9914920/v1
UR - https://doi.org/10.21203/rs.3.rs-9914920/v1
LA - en
ER -

CSL-JSON

{
"id": "10.21203/rs.3.rs-9914920/v1",
"type": "article",
"title": "Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working Group",
"container-title": "Research Square (preprint)",
"author": [
{
"family": "Bi",
"given": "Hanwen"
},
{
"family": "Åkerstedt",
"given": "Torbjörn"
},
{
"family": "Bülow",
"given": "Robin"
},
{
"family": "Deantoni",
"given": "Michele"
},
{
"family": "Drzezga",
"given": "Alexander"
},
{
"family": "Elman",
"given": "Jeremy A."
},
{
"family": "Elmenhorst",
"given": "David"
},
{
"family": "Elmenhorst",
"given": "Eva-Maria"
},
{
"family": "Ewert",
"given": "Ralf"
},
{
"family": "Fennema-Notestine",
"given": "Christine"
},
{
"family": "Ferrarelli",
"given": "Fabio"
},
{
"family": "Frenzel",
"given": "Stefan"
},
{
"family": "von Gall",
"given": "Charlotte"
},
{
"family": "Genon",
"given": "Sarah"
},
{
"family": "Grabe",
"given": "Hans J."
},
{
"family": "Nguyen Ho",
"given": "Phuong Thuy"
},
{
"family": "Hoepel",
"given": "Sanne J.W."
},
{
"family": "Hoffstaedter",
"given": "Felix"
},
{
"family": "Ibanez",
"given": "Agustin"
},
{
"family": "Jahanshad",
"given": "Neda"
},
{
"family": "Keihani",
"given": "Ahmadreza"
},
{
"family": "Küppers",
"given": "Vincent"
},
{
"family": "Liu",
"given": "Wen"
},
{
"family": "Mayeli",
"given": "Ahmad"
},
{
"family": "Mortazavi",
"given": "Nasrin"
},
{
"family": "Neitzel",
"given": "Julia"
},
{
"family": "Nilsonne",
"given": "Gustav"
},
{
"family": "Panizzon",
"given": "Matthew S."
},
{
"family": "Rupp",
"given": "Julia S."
},
{
"family": "Saberi",
"given": "Amin"
},
{
"family": "Schmidt",
"given": "Christina"
},
{
"family": "Spiegelhalder",
"given": "Kai"
},
{
"family": "Stubbe",
"given": "Beate"
},
{
"family": "Tamm",
"given": "Sandra"
},
{
"family": "Thomopoulos",
"given": "Sophia I."
},
{
"family": "Thompson",
"given": "Paul M."
},
{
"family": "Valk",
"given": "Sofie L."
},
{
"family": "Vandewalle",
"given": "Gilles"
},
{
"family": "Vo-Eckerle",
"given": "Tina Thi"
},
{
"family": "Völzke",
"given": "Henry"
},
{
"family": "Waite",
"given": "Laura K."
},
{
"family": "Wexler",
"given": "Joseph"
},
{
"family": "Wittfeld",
"given": "Katharina"
},
{
"family": "Patil",
"given": "Kaustubh R."
},
{
"family": "Weihs",
"given": "Antoine"
},
{
"family": "Eickhoff",
"given": "Simon B."
},
{
"family": "Raimondo",
"given": "Federico"
},
{
"family": "Tahmasian",
"given": "Masoud"
},
{
"literal": "ENIGMA-Sleep Working Group"
}
],
"container-title-short": "Res Sq",
"DOI": "10.21203/rs.3.rs-9914920/v1",
"ISSN": "2693-5015",
"publisher": "Research Square",
"URL": "https://doi.org/10.21203/rs.3.rs-9914920/v1",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
9
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-71271-9 [code]
Exposome-wide patterns predict brain health in aging.
Journal: Nature communications
In common: SHAP, seaborn, scikit-learn, 4 other tools, 2 references, 4 authors
[2] doi:10.1038/s41467-026-74153-2 [code]
Regional, functional and transcriptomic decoding of multidimensional brain structure alterations in obsessive-compulsive disorder.
Journal: Nature communications
In common: neuromaps, BrainSpace, FreeSurfer, 7 other tools, 2 authors
[3] doi:10.64898/2026.08.13.26360304 [code]
Lifespan brain structural variation reveals shared organization across mental health conditions
Journal: medRxiv (preprint)
In common: BrainSpace, Nilearn, NiBabel, 7 other tools, 2 authors
[4] doi:10.1038/s41398-026-04189-x [code]
Decomposing neuroanatomical heterogeneity in depression: insights from an ENIGMA major depressive disorder working group study in 5146 individuals.
Journal: Translational psychiatry
In common: structural MRI / diffusion, 4 authors
[5] doi:10.1038/s41398-026-04078-3 [code]
Brain age prediction in generalized anxiety disorder using a convolutional neural network.
Journal: Translational psychiatry
In common: structural MRI / diffusion, 4 authors
[6] doi:10.1002/agm2.70073 [code]
Foreign Language Learning in Older Adults Modifies Resting-State Functional Connectivity Between the Subcortical Structures and the Cortex.
Journal: Aging medicine (Milton (N.S.W))
In common: netneurotools, neuromaps, BrainSpace, 10 other tools, 1 reference
[7] doi:10.1038/s42003-026-10276-y [code]
The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.
Journal: Communications biology
In common: CAT12, BrainSpace, statannotations, 10 other tools, structural MRI / diffusion
[8] doi:10.1371/journal.pbio.3003684 [code]
The retrieval of previously learned motor memories is facilitated by the reinstatement of default mode network manifold structures.
Journal: PLoS biology
In common: neuromaps, BrainSpace, Pingouin, 9 other tools, 2 references
[9] doi:10.7554/elife.103097 [code]
Canonical neurodevelopmental trajectories of structural and functional manifolds.
Journal: eLife
In common: netneurotools, neuromaps, BrainSpace, 8 other tools, structural MRI / diffusion, 1 reference
[10] doi:10.1371/journal.pbio.3003856 [code]
Aging and metabolism contribute separately to brain-body health.
Journal: PLoS biology
In common: netneurotools, neuromaps, FreeSurfer, 9 other tools, structural MRI / diffusion, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.