OSCR

A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.

Code ↔ Paper

17 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 17 matches
  1. [1] § Materials and methods › Neuroimaging preprocessing › Anatomical preprocessing ↔ MEG_preprocessing/MEG_preprocessing_driver_script.sh, lines 47–108 · score 0.90 · single shell boundary, Scalp surfaces, FreeSurfer, elements model, forward model, recon
  2. [2] § Materials and methods › Fitting classifiers ↔ barycenter_robustness/barycenter_robustness_checks.ipynb, lines 81–130 · score 0.70 · StratifiedGroupKFold, cross task classification, logistic regression, v1, Irrelevant, predict
  3. [3] § Materials and methods › Neuroimaging preprocessing › MEG preprocessing ↔ MEG_preprocessing/cogitate-msp1/coglib/meeg/qc/QC_processing_eeg.py, lines 179–221 · score 0.69 · muscle artifacts, Maxwell filter, gradiometer, magnetometer, threshold, preprocessed
  4. [4] § Materials and methods › Fitting classifiers ↔ classification/fit_pyspi_classifiers.py, lines 346–422 · score 0.67 · StratifiedGroupKFold, cross task classification, cross validated, fit, SD, train
  5. [5] § Materials and methods › Neuroimaging preprocessing › MEG preprocessing ↔ MEG_preprocessing/MEG_preprocessing_driver_script.sh, lines 47–108 · score 0.67 · event related field, brain region, pipeline, preprocessed, fit, MEG
  6. [6] § Materials and methods › Neuroimaging preprocessing › Anatomical preprocessing ↔ data_visualization/methods.ipynb, lines 26–76 · score 0.65 · intraparietal sulcus, prefrontal cortex, Category selective, inflated, fsaverage, surface
  7. [7] § Materials and methods › Neuroimaging preprocessing › Anatomical preprocessing ↔ data_visualization/methods.ipynb, lines 26–76 · score 0.62 · prefrontal cortex, category selective, parcel, Network, atlas, FreeSurfer
  8. [8] § Materials and methods › Fitting classifiers ↔ MEG_preprocessing/cogitate-msp1/coglib/ieeg/decoding/calibration.py, lines 18–89 · score 0.62 · support vector machine, cross validated, binary, regression, probability, fit
  9. [9] § Materials and methods › Fitting classifiers ↔ barycenter_robustness/barycenter_robustness_checks.ipynb, lines 153–244 · score 0.59 · robustness check, logistic regression, cross validated, fit, classifier, accuracy
  10. [10] § Materials and methods › Neuroimaging data acquisition and task paradigm ↔ MEG_preprocessing/cogitate-msp1/coglib/beh_et/behavior/quality_checks.py, lines 833–903 · score 0.59 · alarm rates, hit rates, behavioral, durations, stimuli
  11. [11] § Materials and methods › Theory-driven neural modeling › Model development and implementation ↔ modeling/CogitateModels.ipynb, lines 66–122 · score 0.56 · strong adaptation, stimulus offset, stimulus onset, simulated, CS, PFC
  12. [12] § Materials and methods › Fitting classifiers ↔ classification/fit_pyspi_classifiers.py, lines 346–422 · score 0.56 · cross task classification, cross validating, pyspi, fit, stratified, folds
  13. [13] § Results ↔ classification/fit_pyspi_classifiers.py, lines 550–631 · score 0.53 · standard deviation, Fitting classifiers, cross validation, pyspi, SD, Irrelevant
  14. [14] § Materials and methods › Neuroimaging data acquisition and task paradigm ↔ MEG_preprocessing/cogitate-msp1/coglib/fmri/logfiles_and_checks/01_exp1_create_events_tsv_file.py, lines 89–223 · score 0.53 · stimulus events, alarm, hit, letters, BIDS, Irrelevant
  15. [15] § Materials and methods › Neuroimaging preprocessing › Anatomical preprocessing ↔ data_visualization/classification_analysis_visualization.ipynb, lines 65–155 · score 0.53 · prefrontal cortex, Category selective, IPS, SPIs, V2, V1
  16. [16] § Materials and methods › Neuroimaging preprocessing › Anatomical preprocessing ↔ data_visualization/classification_analysis_visualization.ipynb, lines 65–155 · score 0.52 · prefrontal cortex, category selective, V2, V1, mapping, CS
  17. [17] § Materials and methods › Theory-driven neural modeling › Quantitative model evaluation ↔ modeling/CogitateModels.ipynb, lines 66–122 · score 0.50 · GNWT models, sweep, smaller, stimulus onset, noise, simulated

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 894 lines · 52 KB · no license · 3 matches

  1. from copy import deepcopy
  2. from glob import glob
  3. import os
  4. from os import path as op
  5. import numpy as np
  6. import pandas as pd
  7. import sys
  8. from sklearn import svm
  9. from sklearn.linear_model import LogisticRegression
  10. from sklearn.metrics import make_scorer, roc_auc_score, accuracy_score, balanced_accuracy_score
  11. from sklearn.model_selection import StratifiedGroupKFold, cross_validate, StratifiedKFold, LeaveOneOut, cross_val_predict
  12. from sklearn.pipeline import Pipeline
  13. from sklearn.utils import resample
  14. import itertools
  15. import argparse
  16. from joblib import Parallel, delayed
  17. # add path to classification analysis functions
  18. from mixed_sigmoid_normalisation import MixedSigmoidScaler
  19. parser=argparse.ArgumentParser()
  20. parser.add_argument('--bids_root',
  21. type=str,
  22. default='/project/hctsa/annie/data/Cogitate_Batch1/MEG_Data/',
  23. help='Path to the BIDS root directory')
  24. parser.add_argument('--n_jobs',
  25. type=int,
  26. default=1,
  27. help='Number of concurrent processing jobs')
  28. parser.add_argument('--SPI_directionality_file',
  29. type=str,
  30. default='/headnode1/abry4213/github/Cogitate_Connectivity_2024/feature_extraction/pyspi_SPI_info.csv',
  31. help='CSV file with SPI directionality info')
  32. parser.add_argument('--subject_ID',
  33. type=str,
  34. default=None,
  35. help='Subject for intra-subject classification [optional]')
  36. parser.add_argument('--classification_type',
  37. type=str,
  38. default='all',
  39. help='Whether to perform average and/or individual classification; default is all')
  40. parser.add_argument('--classifier',
  41. type=str,
  42. default='Logistic_Regression',
  43. help='Which type of classifier to use')
  44. opt=parser.parse_args()
  45. bids_root = opt.bids_root
  46. n_jobs = opt.n_jobs
  47. subject_ID = opt.subject_ID
  48. SPI_directionality_file = opt.SPI_directionality_file
  49. classification_type = opt.classification_type
  50. classifier = opt.classifier
  51. # Read in SPI directionality info
  52. SPI_directionality_info = pd.read_csv(SPI_directionality_file)
  53. # Load data paths
  54. pyspi_res_path = f"{bids_root}/derivatives/time_series_features"
  55. pyspi_res_path_averaged = f"{pyspi_res_path}/averaged_epochs"
  56. pyspi_res_path_individual = f"{pyspi_res_path}/individual_epochs"
  57. classification_res_path = f"{bids_root}/derivatives/classification_results"
  58. classification_res_path_averaged = f"{classification_res_path}/across_participants"
  59. classification_res_path_individual = f"{classification_res_path}/within_participants"
  60. # Make classification result directories
  61. os.makedirs(classification_res_path_averaged, exist_ok=True)
  62. os.makedirs(classification_res_path_individual, exist_ok=True)
  63. # Define classifier
  64. if classifier == "Linear_SVM":
  65. model = svm.SVC(C=1, class_weight='balanced', kernel='linear', random_state=127, probability=True)
  66. elif classifier == "Logistic_Regression":
  67. model = LogisticRegression(penalty='l1', C=1, solver='liblinear', class_weight='balanced', random_state=127)
  68. else:
  69. model = svm.SVC(C=1, class_weight='balanced', kernel='rbf', random_state=127, probability=True)
  70. pipe = Pipeline([('scaler', MixedSigmoidScaler(unit_variance=True)),
  71. ('model', model)])
  72. # Define scoring type
  73. scoring = {'accuracy': 'accuracy',
  74. 'balanced_accuracy': 'balanced_accuracy',
  75. 'AUC': make_scorer(roc_auc_score, response_method='predict_proba')}
  76. # meta-ROI comparisons
  77. meta_ROIs = ["Category_Selective", "IPS", "Prefrontal_Cortex", "V1_V2"]
  78. # Manually define combinations
  79. meta_roi_comparisons = [("Category_Selective", "IPS"),
  80. ("Category_Selective", "Prefrontal_Cortex"),
  81. ("Category_Selective", "V1_V2"),
  82. ("IPS", "Category_Selective"),
  83. ("Prefrontal_Cortex", "Category_Selective"),
  84. ("V1_V2", "Category_Selective")]
  85. # meta_roi_comparisons = list(itertools.permutations(meta_ROIs, 2))
  86. # Relevance type comparisons
  87. relevance_type_comparisons = ["Relevant-non-target", "Irrelevant"]
  88. # Stimulus presentation comparisons
  89. stimulus_presentation_comparisons = ["on", "off"]
  90. # Define all combinations for cross-task classification
  91. all_combos_for_cross_task = list(itertools.product(["relevant_to_irrelevant", "irrelevant_to_relevant"],
  92. ["on", "off"],
  93. meta_roi_comparisons))
  94. # Define cross-validators
  95. group_stratified_CV = StratifiedGroupKFold(n_splits = 10, shuffle = True, random_state=127)
  96. LOOCV = LeaveOneOut()
  97. SKF = StratifiedKFold(n_splits=5, shuffle=True, random_state=127)
  98. # Helper function for cross-task analysis
  99. def cross_task_classifier(direction, meta_roi_comparison, stimulus_presentation, pyspi_data):
  100. ROI_from, ROI_to = meta_roi_comparison
  101. # Filter pyspi data
  102. pyspi_data = (pyspi_data.query("meta_ROI_from == @ROI_from & meta_ROI_to == @ROI_to & stimulus_presentation == @stimulus_presentation")
  103. .reset_index(drop=True)
  104. .drop(columns=['index']))
  105. # All comparisons list
  106. cross_task_classification_results_list = []
  107. for SPI in pyspi_data.SPI.unique():
  108. # Extract this SPI
  109. this_SPI_data = pyspi_data.query(f"SPI == '{SPI}'")
  110. # Find overall number of rows
  111. num_rows = this_SPI_data.shape[0]
  112. # Extract SPI values
  113. this_column_data = this_SPI_data["value"]
  114. # Find number of NaN in this column
  115. num_NaN = this_column_data.isna().sum()
  116. prop_NaN = num_NaN / num_rows
  117. # Find mode and SD
  118. column_mode_max = this_column_data.value_counts().max()
  119. column_SD = this_column_data.std()
  120. # If 0% < num_NaN < 10%, impute by the mean of each component
  121. if 0 < prop_NaN < 0.1:
  122. values_imputed = (this_column_data
  123. .transform(lambda x: x.fillna(x.mean())))
  124. this_column_data = values_imputed
  125. print(f"Imputing column values for {SPI}")
  126. this_SPI_data["value"] = this_column_data
  127. # If there are:
  128. # - more than 10% NaN values;
  129. # - more than 90% of the values are the same; OR
  130. # - the standard deviation is less than 1*10**(-10)
  131. # then remove the column
  132. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  133. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  134. continue
  135. # Iterate over stimulus combos
  136. for this_combo in stimulus_type_comparisons:
  137. # Subset data to the corresponding stimulus pairs
  138. final_dataset_for_classification_this_combo = this_SPI_data.query(f"stimulus_type in {this_combo}")
  139. if direction == "relevant_to_irrelevant":
  140. train_df = final_dataset_for_classification_this_combo.query("relevance_type == 'Relevant-non-target'")
  141. test_df = final_dataset_for_classification_this_combo.query("relevance_type == 'Irrelevant'")
  142. else:
  143. train_df = final_dataset_for_classification_this_combo.query("relevance_type == 'Irrelevant'")
  144. test_df = final_dataset_for_classification_this_combo.query("relevance_type == 'Relevant-non-target'")
  145. # Make a deepcopy of the pipeline
  146. this_iter_pipe = deepcopy(pipe)
  147. # Fit classifier
  148. X_train = train_df.value.to_numpy().reshape(-1, 1)
  149. y_train = train_df.stimulus_type.to_numpy().reshape(-1, 1)
  150. X_test = test_df.value.to_numpy().reshape(-1, 1)
  151. y_test = test_df.stimulus_type.to_numpy().reshape(-1, 1)
  152. this_iter_pipe.fit(X_train, y_train)
  153. y_pred = this_iter_pipe.predict(X_test)
  154. # Compute accuracy, balanced accuracy, and AUC
  155. accuracy = accuracy_score(y_test, y_pred)
  156. balanced_accuracy = balanced_accuracy_score(y_test, y_pred)
  157. this_SPI_combo_df = pd.DataFrame({"SPI": [SPI],
  158. "classifier": [classifier],
  159. "meta_ROI_from": [ROI_from],
  160. "meta_ROI_to": [ROI_to],
  161. "cross_task_direction": [direction],
  162. "stimulus_presentation": [stimulus_presentation],
  163. "stimulus_combo": [this_combo],
  164. "accuracy": [accuracy],
  165. "balanced_accuracy": [balanced_accuracy]})
  166. # Append to growing results list
  167. cross_task_classification_results_list.append(this_SPI_combo_df)
  168. # Concatenate all results
  169. cross_task_classification_results_df = pd.concat(cross_task_classification_results_list)
  170. # Return results
  171. return cross_task_classification_results_df
  172. #################################################################################################
  173. # Classification across participants with averaged epochs
  174. #################################################################################################
  175. if classification_type == "averaged":
  176. # Load in pyspi results
  177. all_pyspi_res_list = []
  178. # for pyspi_res_file in os.listdir(pyspi_res_path_averaged):
  179. for pyspi_res_file in glob(f"{pyspi_res_path_averaged}/*all_pyspi_results_1000ms.csv"):
  180. pyspi_res = pd.read_csv(pyspi_res_file)
  181. # Reset index
  182. pyspi_res.reset_index(inplace=True, drop=True)
  183. pyspi_res['stimulus_type'] = pyspi_res['stimulus_type'].replace(False, 'false').replace('False', 'false')
  184. pyspi_res['relevance_type'] = pyspi_res['relevance_type'].replace("Relevant non-target", "Relevant-non-target")
  185. # Rename stimulus to stimulus_presentation if it is present
  186. if 'stimulus' in pyspi_res.columns:
  187. if 'stimulus_presentation' in pyspi_res.columns:
  188. pyspi_res.drop(columns=['stimulus'], inplace=True)
  189. else:
  190. pyspi_res = pyspi_res.rename(columns={'stimulus': 'stimulus_presentation'})
  191. all_pyspi_res_list.append(pyspi_res)
  192. all_pyspi_res = pd.concat(all_pyspi_res_list)
  193. # Stimulus type comparisons
  194. stimulus_types = all_pyspi_res.stimulus_type.unique().tolist()
  195. stimulus_type_comparisons = list(itertools.combinations(stimulus_types, 2))
  196. # Comparing between stimulus types
  197. if not os.path.isfile(f"{classification_res_path_averaged}/comparing_between_stimulus_types_{classifier}_classification_results_AUC.csv"):
  198. # All comparisons list
  199. comparing_between_stimulus_types_classification_results_list = []
  200. for relevance_type in relevance_type_comparisons:
  201. print("Relevance type:" + str(relevance_type))
  202. for stimulus_presentation in stimulus_presentation_comparisons:
  203. print("Stimulus presentation:" + str(stimulus_presentation))
  204. for SPI in all_pyspi_res.SPI.unique():
  205. # First, look at each meta-ROI pair separately
  206. for meta_roi_comparison in meta_roi_comparisons:
  207. print("ROI Comparison:" + str(meta_roi_comparison))
  208. ROI_from, ROI_to = meta_roi_comparison
  209. # Finally, we get to the final dataset
  210. roi_pair_wise_dataset_for_classification = (all_pyspi_res.query("meta_ROI_from == @ROI_from & meta_ROI_to == @ROI_to & relevance_type == @relevance_type & stimulus_presentation == @stimulus_presentation")
  211. .reset_index(drop=True)
  212. .drop(columns=['index']))
  213. # Extract this SPI
  214. this_SPI_data = roi_pair_wise_dataset_for_classification.query(f"SPI == '{SPI}'")
  215. # Find overall number of rows
  216. num_rows = this_SPI_data.shape[0]
  217. # Extract SPI values
  218. this_column_data = this_SPI_data["value"]
  219. # Find number of NaN in this column
  220. num_NaN = this_column_data.isna().sum()
  221. prop_NaN = num_NaN / num_rows
  222. # Find mode and SD
  223. column_mode_max = this_column_data.value_counts().max()
  224. column_SD = this_column_data.std()
  225. # If 0% < num_NaN < 10%, impute by the mean of each component
  226. if 0 < prop_NaN < 0.1:
  227. values_imputed = (this_column_data
  228. .transform(lambda x: x.fillna(x.mean())))
  229. this_column_data = values_imputed
  230. print(f"Imputing column values for {SPI}")
  231. this_SPI_data["value"] = this_column_data
  232. # If there are:
  233. # - more than 10% NaN values;
  234. # - more than 90% of the values are the same; OR
  235. # - the standard deviation is less than 1*10**(-10)
  236. # then remove the column
  237. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  238. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  239. continue
  240. # Start an empty list for the classification results
  241. SPI_combo_res_list = []
  242. # Iterate over stimulus combos
  243. for this_combo in stimulus_type_comparisons:
  244. # Subset data to the corresponding stimulus pairs
  245. final_dataset_for_classification_this_combo = this_SPI_data.query(f"stimulus_type in {this_combo}")
  246. # Fit classifier
  247. X = final_dataset_for_classification_this_combo.value.to_numpy().reshape(-1, 1)
  248. y = final_dataset_for_classification_this_combo.stimulus_type.to_numpy().reshape(-1, 1)
  249. groups = final_dataset_for_classification_this_combo.subject_ID.to_numpy().reshape(-1, 1)
  250. groups_flat = np.array([str(item[0]) for item in groups])
  251. # Make a deepcopy of the pipeline
  252. this_iter_pipe = deepcopy(pipe)
  253. this_classifier_res = cross_validate(this_iter_pipe, X, y, groups=groups_flat, cv=group_stratified_CV, scoring=scoring, n_jobs=n_jobs,
  254. return_estimator=False, return_train_score=False)
  255. this_SPI_combo_df = pd.DataFrame({"SPI": [SPI],
  256. "classifier": [classifier],
  257. "meta_ROI_from": [ROI_from],
  258. "meta_ROI_to": [ROI_to],
  259. "relevance_type": [relevance_type],
  260. "stimulus_presentation": [stimulus_presentation],
  261. "stimulus_combo": [this_combo],
  262. "accuracy": [this_classifier_res['test_accuracy'].mean()],
  263. "accuracy_SD": [this_classifier_res['test_accuracy'].std()],
  264. "AUC": [this_classifier_res['test_AUC'].mean()],
  265. "AUC_SD": [this_classifier_res['test_AUC'].std()]})
  266. # Append to growing results list
  267. comparing_between_stimulus_types_classification_results_list.append(this_SPI_combo_df)
  268. comparing_between_stimulus_types_classification_results = pd.concat(comparing_between_stimulus_types_classification_results_list).reset_index(drop=True)
  269. comparing_between_stimulus_types_classification_results.to_csv(f"{classification_res_path_averaged}/comparing_between_stimulus_types_{classifier}_classification_results_AUC.csv", index=False)
  270. # Comparing between relevance types
  271. if not os.path.isfile(f"{classification_res_path_averaged}/comparing_between_relevance_types_{classifier}_classification_results.csv"):
  272. # All comparisons list
  273. comparing_between_relevance_types_classification_results_list = []
  274. for meta_roi_comparison in meta_roi_comparisons:
  275. print("ROI Comparison:" + str(meta_roi_comparison))
  276. ROI_from, ROI_to = meta_roi_comparison
  277. for stimulus_presentation in stimulus_presentation_comparisons:
  278. print("Stimulus presentation:" + str(stimulus_presentation))
  279. # Finally, we get to the final dataset
  280. final_dataset_for_classification = all_pyspi_res.query("meta_ROI_from == @ROI_from & relevance_type in @relevance_type_comparisons and meta_ROI_to == @ROI_to & stimulus_presentation == @stimulus_presentation").reset_index(drop=True).drop(columns=['index'])
  281. for SPI in final_dataset_for_classification.SPI.unique():
  282. # Extract this SPI
  283. this_SPI_data = final_dataset_for_classification.query(f"SPI == '{SPI}'")
  284. # Find overall number of rows
  285. num_rows = this_SPI_data.shape[0]
  286. # Extract SPI values
  287. this_column_data = this_SPI_data["value"]
  288. # Find number of NaN in this column
  289. num_NaN = this_column_data.isna().sum()
  290. prop_NaN = num_NaN / num_rows
  291. # Find mode and SD
  292. column_mode_max = this_column_data.value_counts().max()
  293. column_SD = this_column_data.std()
  294. # If 0% < num_NaN < 10%, impute by the mean of each component
  295. if 0 < prop_NaN < 0.1:
  296. values_imputed = (this_column_data
  297. .transform(lambda x: x.fillna(x.mean())))
  298. this_column_data = values_imputed
  299. print(f"Imputing column values for {SPI}")
  300. this_SPI_data["value"] = this_column_data
  301. # If there are:
  302. # - more than 10% NaN values;
  303. # - more than 90% of the values are the same; OR
  304. # - the standard deviation is less than 1*10**(-10)
  305. # then remove the column
  306. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  307. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  308. continue
  309. # Start an empty list for the classification results
  310. SPI_combo_res_list = []
  311. # Fit classifier
  312. X = this_SPI_data.value.to_numpy().reshape(-1, 1)
  313. y = this_SPI_data.relevance_type.to_numpy().reshape(-1, 1)
  314. groups = this_SPI_data.subject_ID.to_numpy().reshape(-1, 1)
  315. groups_flat = np.array([str(item[0]) for item in groups])
  316. group_stratified_CV = StratifiedGroupKFold(n_splits = 10, shuffle = True, random_state=127)
  317. # Make a deepcopy of the pipeline
  318. this_iter_pipe = deepcopy(pipe)
  319. this_classifier_res = cross_validate(this_iter_pipe, X, y, groups=groups_flat, cv=group_stratified_CV, scoring=scoring, n_jobs=n_jobs,
  320. return_estimator=False, return_train_score=False)
  321. this_SPI_relevance_results_df = pd.DataFrame({"SPI": [SPI],
  322. "meta_ROI_from": [ROI_from],
  323. "meta_ROI_to": [ROI_to],
  324. "stimulus_presentation": [stimulus_presentation],
  325. "comparison": ["Relevant non-target vs. Irrelevant"],
  326. "accuracy": [this_classifier_res['test_accuracy'].mean()],
  327. "accuracy_SD": [this_classifier_res['test_accuracy'].std()]})
  328. # Append to growing results list
  329. comparing_between_relevance_types_classification_results_list.append(this_SPI_relevance_results_df)
  330. comparing_between_relevance_types_classification_results = pd.concat(comparing_between_relevance_types_classification_results_list).reset_index(drop=True)
  331. comparing_between_relevance_types_classification_results.to_csv(f"{classification_res_path_averaged }/comparing_between_relevance_types_{classifier}_classification_results.csv", index=False)
  332. # Cross-task learning
  333. if not os.path.isfile(f"{classification_res_path_averaged}/cross_task_{classifier}_classification_results.csv"):
  334. print("Starting cross-task classification")
  335. cross_task_classification_results_list = Parallel(n_jobs=int(n_jobs))(delayed(cross_task_classifier)(direction=direction,
  336. meta_roi_comparison=meta_roi_comparison,
  337. stimulus_presentation=stimulus_presentation,
  338. pyspi_data=all_pyspi_res)
  339. for direction, stimulus_presentation, meta_roi_comparison in all_combos_for_cross_task)
  340. cross_task_classification_results = pd.concat(cross_task_classification_results_list).reset_index(drop=True)
  341. cross_task_classification_results.to_csv(f"{classification_res_path_averaged}/cross_task_{classifier}_classification_results.csv", index=False)
  342. #################################################################################################
  343. # Classification across participants with individual epochs
  344. #################################################################################################
  345. if classification_type == "individual":
  346. # meta-ROI comparisons
  347. meta_ROIs = ["Category_Selective", "IPS", "Prefrontal_Cortex", "V1_V2"]
  348. meta_roi_comparisons = list(itertools.permutations(meta_ROIs, 2))
  349. # BY STIMULUS TYPE
  350. # Load in this subject's pyspi results
  351. if not op.isfile(f"{classification_res_path_individual}/sub-{subject_ID}_comparing_between_stimulus_types_{classifier}_classification_results.csv"):
  352. # Load in results
  353. individual_subject_pyspi_res = pd.read_csv(f"{pyspi_res_path_individual}/sub-{subject_ID}_ses-1_all_pyspi_results_individual_epochs_1000ms.csv")
  354. # Fix stimulus_type where False to 'false'
  355. individual_subject_pyspi_res['stimulus_type'] = individual_subject_pyspi_res['stimulus_type'].replace(False, 'false')
  356. individual_subject_pyspi_res['relevance_type'] = individual_subject_pyspi_res['relevance_type'].replace("Relevant non-target", "Relevant-non-target")
  357. # Relevance type comparisons
  358. relevance_type_comparisons = ["Relevant-non-target", "Irrelevant"]
  359. # Stimulus presentation comparisons
  360. stimulus_presentation_comparisons = individual_subject_pyspi_res.stimulus_presentation.unique().tolist()
  361. # Stimulus type comparisons
  362. stimulus_types = individual_subject_pyspi_res.stimulus_type.unique().tolist()
  363. stimulus_type_comparisons = list(itertools.combinations(stimulus_types, 2))
  364. # All comparisons list
  365. comparing_between_stimulus_types_classification_results_list = []
  366. for meta_roi_comparison in meta_roi_comparisons:
  367. print("ROI Comparison:" + str(meta_roi_comparison))
  368. ROI_from, ROI_to = meta_roi_comparison
  369. for relevance_type in relevance_type_comparisons:
  370. print("Relevance type:" + str(relevance_type))
  371. for stimulus_presentation in stimulus_presentation_comparisons:
  372. print("Stimulus presentation:" + str(stimulus_presentation))
  373. # Finally, we get to the final dataset
  374. final_dataset_for_classification = individual_subject_pyspi_res.query("meta_ROI_from == @ROI_from & meta_ROI_to == @ROI_to & relevance_type == @relevance_type & stimulus_presentation == @stimulus_presentation").reset_index(drop=True).drop(columns=['index'])
  375. for SPI in final_dataset_for_classification.SPI.unique():
  376. # Extract this SPI
  377. this_SPI_data = final_dataset_for_classification.query(f"SPI == '{SPI}'")
  378. # Find overall number of rows
  379. num_rows = this_SPI_data.shape[0]
  380. # Extract SPI values
  381. this_column_data = this_SPI_data["value"]
  382. # Find number of NaN in this column
  383. num_NaN = this_column_data.isna().sum()
  384. prop_NaN = num_NaN / num_rows
  385. # Find mode and SD
  386. column_mode_max = this_column_data.value_counts().max()
  387. column_SD = this_column_data.std()
  388. # If 0% < num_NaN < 10%, impute by the mean of each component
  389. if 0 < prop_NaN < 0.1:
  390. values_imputed = (this_column_data
  391. .transform(lambda x: x.fillna(x.mean())))
  392. this_column_data = values_imputed
  393. print(f"Imputing column values for {SPI}")
  394. this_SPI_data["value"] = this_column_data
  395. # If there are:
  396. # - more than 10% NaN values;
  397. # - more than 90% of the values are the same; OR
  398. # - the standard deviation is less than 1*10**(-10)
  399. # then remove the column
  400. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  401. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  402. continue
  403. # Start an empty list for the classification results
  404. SPI_combo_res_list = []
  405. # Iterate over stimulus combos
  406. for this_combo in stimulus_type_comparisons:
  407. # Subset data to the corresponding stimulus pairs
  408. final_dataset_for_classification_this_combo = this_SPI_data.query(f"stimulus_type in {this_combo}")
  409. # Fit classifier
  410. X = final_dataset_for_classification_this_combo.value.to_numpy().reshape(-1, 1)
  411. y = final_dataset_for_classification_this_combo.stimulus_type.to_numpy().reshape(-1, 1)
  412. stimulus_stratified_CV = StratifiedKFold(n_splits = 10, shuffle = True, random_state=127)
  413. # Make a deepcopy of the pipeline
  414. this_iter_pipe = deepcopy(pipe)
  415. this_classifier_res = cross_validate(this_iter_pipe, X, y, cv=stimulus_stratified_CV, scoring=scoring, n_jobs=n_jobs,
  416. return_estimator=False, return_train_score=False)
  417. this_SPI_combo_df = pd.DataFrame({subject_ID: ["sub-" + subject_ID],
  418. "SPI": [SPI],
  419. "meta_ROI_from": [ROI_from],
  420. "meta_ROI_to": [ROI_to],
  421. "relevance_type": [relevance_type],
  422. "stimulus_presentation": [stimulus_presentation],
  423. "stimulus_combo": [this_combo],
  424. "accuracy": [this_classifier_res['test_accuracy'].mean()],
  425. "accuracy_SD": [this_classifier_res['test_accuracy'].std()]})
  426. # Append to growing results list
  427. comparing_between_stimulus_types_classification_results_list.append(this_SPI_combo_df)
  428. comparing_between_stimulus_types_classification_results = pd.concat(comparing_between_stimulus_types_classification_results_list).reset_index(drop=True)
  429. comparing_between_stimulus_types_classification_results.to_csv(f"{classification_res_path_individual}/sub-{subject_ID}_comparing_between_stimulus_types_{classifier}_classification_results.csv", index=False)
  430. # Comparing between relevance types
  431. # BY RELEVANCE TYPE
  432. if not os.path.isfile(f"{classification_res_path_individual}/sub-{subject_ID}_comparing_between_relevance_types_{classifier}_classification_results.csv"):
  433. # Load in results
  434. individual_subject_pyspi_res = pd.read_csv(f"{pyspi_res_path_individual}/sub-{subject_ID}_ses-1_all_pyspi_results_individual_epochs_1000ms.csv")
  435. # Fix stimulus_type where False to 'false'
  436. individual_subject_pyspi_res['stimulus_type'] = individual_subject_pyspi_res['stimulus_type'].replace(False, 'false')
  437. # Relevance type comparisons
  438. relevance_type_comparisons = ["Relevant-non-target", "Irrelevant"]
  439. # Stimulus presentation comparisons
  440. stimulus_presentation_comparisons = individual_subject_pyspi_res.stimulus_presentation.unique().tolist()
  441. # All comparisons list
  442. comparing_between_relevance_types_classification_results_list = []
  443. for meta_roi_comparison in meta_roi_comparisons:
  444. print("ROI Comparison:" + str(meta_roi_comparison))
  445. ROI_from, ROI_to = meta_roi_comparison
  446. for stimulus_presentation in stimulus_presentation_comparisons:
  447. print("Stimulus presentation:" + str(stimulus_presentation))
  448. # Finally, we get to the final dataset
  449. final_dataset_for_classification = individual_subject_pyspi_res.query("meta_ROI_from == @ROI_from & relevance_type in @relevance_type_comparisons and meta_ROI_to == @ROI_to & stimulus_presentation == @stimulus_presentation").reset_index(drop=True).drop(columns=['index'])
  450. for SPI in final_dataset_for_classification.SPI.unique():
  451. # Extract this SPI
  452. this_SPI_data = final_dataset_for_classification.query(f"SPI == '{SPI}'")
  453. # Find overall number of rows
  454. num_rows = this_SPI_data.shape[0]
  455. # Extract SPI values
  456. this_column_data = this_SPI_data["value"]
  457. # Find number of NaN in this column
  458. num_NaN = this_column_data.isna().sum()
  459. prop_NaN = num_NaN / num_rows
  460. # Find mode and SD
  461. column_mode_max = this_column_data.value_counts().max()
  462. column_SD = this_column_data.std()
  463. # If 0% < num_NaN < 10%, impute by the mean of each component
  464. if 0 < prop_NaN < 0.1:
  465. values_imputed = (this_column_data
  466. .transform(lambda x: x.fillna(x.mean())))
  467. this_column_data = values_imputed
  468. print(f"Imputing column values for {SPI}")
  469. this_SPI_data["value"] = this_column_data
  470. # If there are:
  471. # - more than 10% NaN values;
  472. # - more than 90% of the values are the same; OR
  473. # - the standard deviation is less than 1*10**(-10)
  474. # then remove the column
  475. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  476. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  477. continue
  478. # Start an empty list for the classification results
  479. SPI_combo_res_list = []
  480. # Fit classifier
  481. X = this_SPI_data.value.to_numpy().reshape(-1, 1)
  482. y = this_SPI_data.relevance_type.to_numpy().reshape(-1, 1)
  483. stimulus_stratified_CV = StratifiedKFold(n_splits = 10, shuffle = True, random_state=127)
  484. # Make a deepcopy of the pipeline
  485. this_iter_pipe = deepcopy(pipe)
  486. this_classifier_res = cross_validate(this_iter_pipe, X, y, cv=stimulus_stratified_CV, scoring=scoring, n_jobs=n_jobs,
  487. return_estimator=False, return_train_score=False)
  488. this_SPI_relevance_results_df = pd.DataFrame({subject_ID: ["sub-" + subject_ID],
  489. "SPI": [SPI],
  490. "meta_ROI_from": [ROI_from],
  491. "meta_ROI_to": [ROI_to],
  492. "stimulus_presentation": [stimulus_presentation],
  493. "comparison": ["Relevant non-target vs. Irrelevant"],
  494. "accuracy": [this_classifier_res['test_accuracy'].mean()],
  495. "accuracy_SD": [this_classifier_res['test_accuracy'].std()]})
  496. # Append to growing results list
  497. comparing_between_relevance_types_classification_results_list.append(this_SPI_relevance_results_df)
  498. comparing_between_relevance_types_classification_results = pd.concat(comparing_between_relevance_types_classification_results_list).reset_index(drop=True)
  499. comparing_between_relevance_types_classification_results.to_csv(f"{classification_res_path_individual}/sub-{subject_ID}_comparing_between_relevance_types_{classifier}_classification_results.csv", index=False)
  500. #################################################################################################
  501. # Classification across participants with individual epochs
  502. #################################################################################################
  503. if classification_type == "individual_subsampled":
  504. # meta-ROI comparisons
  505. meta_ROIs = ["Category_Selective", "IPS", "Prefrontal_Cortex", "V1_V2"]
  506. meta_roi_comparisons = list(itertools.permutations(meta_ROIs, 2))
  507. # BY STIMULUS TYPE
  508. # Load in this subject's pyspi results
  509. if not op.isfile(f"{classification_res_path_individual}/sub-{subject_ID}_subsampled_comparing_between_stimulus_types_{classifier}_classification_results.csv"):
  510. # Load in results
  511. individual_subject_pyspi_res = pd.read_csv(f"{pyspi_res_path_individual}/sub-{subject_ID}_ses-1_all_pyspi_results_individual_epochs_1000ms.csv")
  512. # Fix stimulus_type where False to 'false'
  513. individual_subject_pyspi_res['stimulus_type'] = individual_subject_pyspi_res['stimulus_type'].replace(False, 'false')
  514. individual_subject_pyspi_res['relevance_type'] = individual_subject_pyspi_res['relevance_type'].replace("Relevant non-target", "Relevant-non-target")
  515. # Relevance type comparisons
  516. relevance_type_comparisons = ["Relevant-non-target", "Irrelevant"]
  517. # Stimulus presentation comparisons
  518. stimulus_presentation_comparisons = individual_subject_pyspi_res.stimulus_presentation.unique().tolist()
  519. # Stimulus type comparisons
  520. stimulus_types = individual_subject_pyspi_res.stimulus_type.unique().tolist()
  521. stimulus_type_comparisons = list(itertools.combinations(stimulus_types, 2))
  522. # All comparisons list
  523. comparing_between_stimulus_types_classification_results_list = []
  524. for meta_roi_comparison in meta_roi_comparisons:
  525. print("ROI Comparison:" + str(meta_roi_comparison))
  526. ROI_from, ROI_to = meta_roi_comparison
  527. for relevance_type in relevance_type_comparisons:
  528. print("Relevance type:" + str(relevance_type))
  529. for stimulus_presentation in stimulus_presentation_comparisons:
  530. print("Stimulus presentation:" + str(stimulus_presentation))
  531. # Finally, we get to the final dataset
  532. final_dataset_for_classification = individual_subject_pyspi_res.query("meta_ROI_from == @ROI_from & meta_ROI_to == @ROI_to & relevance_type == @relevance_type & stimulus_presentation == @stimulus_presentation").reset_index(drop=True).drop(columns=['index'])
  533. for SPI in final_dataset_for_classification.SPI.unique():
  534. # Extract this SPI
  535. this_SPI_data = final_dataset_for_classification.query(f"SPI == '{SPI}'")
  536. # Find overall number of rows
  537. num_rows = this_SPI_data.shape[0]
  538. # Extract SPI values
  539. this_column_data = this_SPI_data["value"]
  540. # Find number of NaN in this column
  541. num_NaN = this_column_data.isna().sum()
  542. prop_NaN = num_NaN / num_rows
  543. # Find mode and SD
  544. column_mode_max = this_column_data.value_counts().max()
  545. column_SD = this_column_data.std()
  546. # If 0% < num_NaN < 10%, impute by the mean of each component
  547. if 0 < prop_NaN < 0.1:
  548. values_imputed = (this_column_data
  549. .transform(lambda x: x.fillna(x.mean())))
  550. this_column_data = values_imputed
  551. print(f"Imputing column values for {SPI}")
  552. this_SPI_data["value"] = this_column_data
  553. # If there are:
  554. # - more than 10% NaN values;
  555. # - more than 90% of the values are the same; OR
  556. # - the standard deviation is less than 1*10**(-10)
  557. # then remove the column
  558. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  559. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  560. continue
  561. # Start an empty list for the classification results
  562. SPI_combo_res_list = []
  563. # Iterate over stimulus combos
  564. for this_combo in stimulus_type_comparisons:
  565. # Subset data to the corresponding stimulus pairs
  566. final_dataset_for_classification_this_combo = this_SPI_data.query(f"stimulus_type in {this_combo}")
  567. # Fit classifier
  568. X = final_dataset_for_classification_this_combo.value.to_numpy().reshape(-1, 1)
  569. y = final_dataset_for_classification_this_combo.stimulus_type.to_numpy().reshape(-1, 1)
  570. # Check if there are >20 samples in each class of y, without hard-coding the values of y
  571. if np.unique(y, return_counts=True)[1].min() > 20:
  572. print(f"Running subsampled classification for {this_combo}")
  573. classification_across_iters_list = []
  574. # For 50 iterations, randomly sample 20 samples from each class for leave-one-out classification
  575. for iter_num in range(100):
  576. # Randomly sample 20 samples from each class for leave-one-out classification
  577. X_resampled = []
  578. y_resampled = []
  579. for this_class in np.unique(y):
  580. X_resampled_class, y_resampled_class = resample(X[y == this_class], y[y == this_class], n_samples=20, replace=False, random_state=iter_num)
  581. X_resampled.append(X_resampled_class)
  582. y_resampled.append(y_resampled_class)
  583. X_resampled = np.concatenate(X_resampled).reshape(-1, 1)
  584. y_resampled = np.concatenate(y_resampled)
  585. y_pred = cross_val_predict(deepcopy(pipe), X_resampled, y_resampled, cv=LOOCV, n_jobs=n_jobs)
  586. y_pred_proba = cross_val_predict(deepcopy(pipe), X_resampled, y_resampled, cv=LOOCV, n_jobs=n_jobs, method='predict_proba')
  587. # Calculate classification results
  588. accuracy = accuracy_score(y_resampled, y_pred)
  589. balanced_accuracy = balanced_accuracy_score(y_resampled, y_pred)
  590. AUC = roc_auc_score(y_resampled, y_pred_proba[:, 1])
  591. # Combine into df
  592. iter_df = pd.DataFrame({"iter_num": iter_num + 1, "accuracy": accuracy, "balanced_accuracy": balanced_accuracy, "AUC": AUC}, index=[0])
  593. classification_across_iters_list.append(iter_df)
  594. # Combine all iterations
  595. classification_across_iters_df = pd.concat(classification_across_iters_list, ignore_index=True)
  596. this_SPI_combo_df = pd.DataFrame({subject_ID: ["sub-" + subject_ID],
  597. "SPI": [SPI],
  598. "meta_ROI_from": [ROI_from],
  599. "meta_ROI_to": [ROI_to],
  600. "relevance_type": [relevance_type],
  601. "stimulus_presentation": [stimulus_presentation],
  602. "stimulus_combo": [this_combo],
  603. "accuracy": [classification_across_iters_df['accuracy'].mean()],
  604. "balanced_accuracy": [classification_across_iters_df['balanced_accuracy'].mean()],
  605. "AUC": [classification_across_iters_df['AUC'].mean()]})
  606. # Append to growing results list
  607. comparing_between_stimulus_types_classification_results_list.append(this_SPI_combo_df)
  608. if len(comparing_between_stimulus_types_classification_results_list) > 0:
  609. comparing_between_stimulus_types_classification_results = pd.concat(comparing_between_stimulus_types_classification_results_list).reset_index(drop=True)
  610. comparing_between_stimulus_types_classification_results.to_csv(f"{classification_res_path_individual}/sub-{subject_ID}_subsampled_comparing_between_stimulus_types_{classifier}_classification_results.csv", index=False)
  611. # Comparing between relevance types
  612. # BY RELEVANCE TYPE
  613. if not os.path.isfile(f"{classification_res_path_individual}/sub-{subject_ID}_subsampled_comparing_between_relevance_types_{classifier}_classification_results.csv"):
  614. # Load in results
  615. individual_subject_pyspi_res = pd.read_csv(f"{pyspi_res_path_individual}/sub-{subject_ID}_ses-1_all_pyspi_results_individual_epochs_1000ms.csv")
  616. # Fix stimulus_type where False to 'false'
  617. individual_subject_pyspi_res['stimulus_type'] = individual_subject_pyspi_res['stimulus_type'].replace(False, 'false')
  618. # Relevance type comparisons
  619. relevance_type_comparisons = ["Relevant-non-target", "Irrelevant"]
  620. # Stimulus presentation comparisons
  621. stimulus_presentation_comparisons = individual_subject_pyspi_res.stimulus_presentation.unique().tolist()
  622. # All comparisons list
  623. comparing_between_relevance_types_classification_results_list = []
  624. for meta_roi_comparison in meta_roi_comparisons:
  625. print("ROI Comparison:" + str(meta_roi_comparison))
  626. ROI_from, ROI_to = meta_roi_comparison
  627. for stimulus_presentation in stimulus_presentation_comparisons:
  628. print("Stimulus presentation:" + str(stimulus_presentation))
  629. # Finally, we get to the final dataset
  630. final_dataset_for_classification = individual_subject_pyspi_res.query("meta_ROI_from == @ROI_from & relevance_type in @relevance_type_comparisons and meta_ROI_to == @ROI_to & stimulus_presentation == @stimulus_presentation").reset_index(drop=True).drop(columns=['index'])
  631. for SPI in final_dataset_for_classification.SPI.unique():
  632. # Extract this SPI
  633. this_SPI_data = final_dataset_for_classification.query(f"SPI == '{SPI}'")
  634. # Find overall number of rows
  635. num_rows = this_SPI_data.shape[0]
  636. # Extract SPI values
  637. this_column_data = this_SPI_data["value"]
  638. # Find number of NaN in this column
  639. num_NaN = this_column_data.isna().sum()
  640. prop_NaN = num_NaN / num_rows
  641. # Find mode and SD
  642. column_mode_max = this_column_data.value_counts().max()
  643. column_SD = this_column_data.std()
  644. # If 0% < num_NaN < 10%, impute by the mean of each component
  645. if 0 < prop_NaN < 0.1:
  646. values_imputed = (this_column_data
  647. .transform(lambda x: x.fillna(x.mean())))
  648. this_column_data = values_imputed
  649. print(f"Imputing column values for {SPI}")
  650. this_SPI_data["value"] = this_column_data
  651. # If there are:
  652. # - more than 10% NaN values;
  653. # - more than 90% of the values are the same; OR
  654. # - the standard deviation is less than 1*10**(-10)
  655. # then remove the column
  656. if prop_NaN > 0.1 or column_mode_max / num_rows > 0.9 or column_SD < 1*10**(-10):
  657. print(f"{SPI} has low SD: {column_SD}, and/or too many mode occurences: {column_mode_max} out of {num_rows}, and/or {100*prop_NaN}% NaN")
  658. continue
  659. # Start an empty list for the classification results
  660. SPI_combo_res_list = []
  661. # Fit classifier
  662. X = this_SPI_data.value.to_numpy().reshape(-1, 1)
  663. y = this_SPI_data.relevance_type.to_numpy().reshape(-1, 1)
  664. # Check if there are >20 samples in each class of y, without hard-coding the values of y
  665. if np.unique(y, return_counts=True)[1].min() > 20:
  666. print(f"Running subsampled classification for {this_combo}")
  667. classification_across_iters_list = []
  668. # For 100 iterations, randomly sample 20 samples from each class for leave-one-out classification
  669. for iter_num in range(100):
  670. # Randomly sample 20 samples from each class for leave-one-out classification
  671. X_resampled = []
  672. y_resampled = []
  673. for this_class in np.unique(y):
  674. X_resampled_class, y_resampled_class = resample(X[y == this_class], y[y == this_class], n_samples=20, replace=False, random_state=iter_num)
  675. X_resampled.append(X_resampled_class)
  676. y_resampled.append(y_resampled_class)
  677. X_resampled = np.concatenate(X_resampled).reshape(-1, 1)
  678. y_resampled = np.concatenate(y_resampled)
  679. y_pred = cross_val_predict(deepcopy(pipe), X_resampled, y_resampled, cv=LOOCV, n_jobs=n_jobs)
  680. y_pred_proba = cross_val_predict(deepcopy(pipe), X_resampled, y_resampled, cv=LOOCV, n_jobs=n_jobs, method='predict_proba')
  681. # Calculate classification results
  682. accuracy = accuracy_score(y_resampled, y_pred)
  683. balanced_accuracy = balanced_accuracy_score(y_resampled, y_pred)
  684. AUC = roc_auc_score(y_resampled, y_pred_proba[:, 1])
  685. # Combine into df
  686. iter_df = pd.DataFrame({"iter_num": iter_num + 1, "accuracy": accuracy, "balanced_accuracy": balanced_accuracy, "AUC": AUC}, index=[0])
  687. classification_across_iters_list.append(iter_df)
  688. # Combine all iterations
  689. classification_across_iters_df = pd.concat(classification_across_iters_list, ignore_index=True)
  690. this_SPI_relevance_results_df = pd.DataFrame({subject_ID: ["sub-" + subject_ID],
  691. "SPI": [SPI],
  692. "meta_ROI_from": [ROI_from],
  693. "meta_ROI_to": [ROI_to],
  694. "stimulus_presentation": [stimulus_presentation],
  695. "comparison": ["Relevant non-target vs. Irrelevant"],
  696. "accuracy": [classification_across_iters_df['accuracy'].mean()],
  697. "balanced_accuracy": [classification_across_iters_df['balanced_accuracy'].mean()],
  698. "AUC": [classification_across_iters_df['AUC'].mean()]})
  699. # Append to growing results list
  700. comparing_between_relevance_types_classification_results_list.append(this_SPI_relevance_results_df)
  701. if len(comparing_between_relevance_types_classification_results_list) > 0:
  702. comparing_between_relevance_types_classification_results = pd.concat(comparing_between_relevance_types_classification_results_list).reset_index(drop=True)
  703. comparing_between_relevance_types_classification_results.to_csv(f"{classification_res_path_individual}/sub-{subject_ID}_subsampled_comparing_between_relevance_types_{classifier}_classification_results.csv", index=False)

fit_pyspi_classifiers.py at commit cab648b, no license · at the source

Overview

  1. School of Physics, The University of Sydney, Camperdown, NSW, Australia
  2. Centre for Complex Systems, The University of Sydney, Camperdown, NSW, Australia
  3. Neuroscience Research Theme, School of Medical Sciences, Faculty of Medicine and Health, The University of Sydney, Sydney, NSW, Australia
Institutions: The University of Sydney (Australia)
Journal: Neuroscience of consciousness, volume 2026, issue 1, article niag029
Dates: received 7 April 2025; accepted 17 May 2026; published online 7 July 2026
Type: Research article · Language: English
License: CC BY-NC
Identifiers: DOI 10.1093/nc/niag029 · PMID 42415873 · PMCID PMC13338905 · OpenAlex W7167585600
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: MEG (modality), none (in silico) (organism), cognitive (subfield)
Methods: Connectivity, Smoothing, state filtering, decompositions, Machine learning, Statistics, Preprocessing
Keywords: complex systems, consciousness, visual perception, functional connectivity
Topic: Face Recognition and Perception (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 80 references in the paper

Abstract

Identifying the neural correlates of conscious visual perception remains a major challenge in neuroscience, requiring theories that bridge between subjective experience and measurable neural correlates. However, theoretical interpretation of empirical evidence is often post hoc and susceptible to confirmation bias. Building upon the adversarial collaboration mediated by the COGITATE Consortium, we present a generalizable approach for the data-driven identification, evaluation, and theoretical modeling of connectivity-based neural correlates of conscious visual perception. Using the same magnetoencephalography (MEG) dataset and accompanying pre-registered hypotheses from the COGITATE Consortium, we systematically compared 246 functional connectivity (FC) measures between regions predicted to underlie conscious vision by Integrated Information Theory (IIT) and/or Global Neuronal Workspace Theory (GNWT). We identified a family of FC measures based on the barycenter—tracking the ‘center of mass’ between two signals—as the top-performing stimulus decoding measures that generalize across regions central to predictions of both IIT and GNWT. To interpret these findings within a theoretical framework, we developed neural mass models that recapitulate the neural dynamics hypothesized to underlie conscious perception by each theory. Comparing simulated barycenter values from these models against empirically measured MEG data revealed that both the GNWT-based model, featuring delayed ignition dynamics, and the IIT-based model, which relied on synchronous sensory dynamics, captured the observed connectivity patterns. These results lend tentative support to GNWT, as the presence of ignition dynamics independent of task-demand conditions contradicts the predictions of IIT. Beyond dataset-specific conclusions and limitations, we introduce a framework for systematically identifying and testing candidate neural correlates of conscious visual perception in an unbiased and interpretable manner.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.

anniegbryant/MEG_functional_connectivity

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: cab648bbe2eab7093f9ab7b9a5fcddabf28e60a0, 28 January 2026
Languages: Python (223), MATLAB (31), Shell (19), Jupyter (12), R (3)
Size: 945 files, 288 scripts
Software Heritage: not archived
Found in: “Code and data availability”
Holds: README, environment (requirements.txt, MEG_preprocessing/cogitate-msp1/coglib/bayesFactor/requirements.txt, MEG_preprocessing/cogitate-msp1/coglib/xnat/requirements_xnat.txt, MEG_preprocessing/cogitate-msp1/coglib/ieeg/plotting_uniformization/requirements.txt), tests, 12 notebooks
Not found: license file, CITATION.cff, continuous integration, documentation
Tools: NumPy (149 files), pandas (101 files), Matplotlib (89 files), MNE-Python (81 files), SciPy (66 files), MNE-BIDS (39 files), scikit-learn (29 files), seaborn (28 files), NiBabel (15 files), Nilearn (12 files), statsmodels (12 files), tidyverse (11 files), cowplot (9 files), easystats (9 files), broom (8 files), patchwork (8 files), scikit-image (8 files), Pingouin (6 files), FreeSurfer (5 files), MNE-Connectivity (5 files), autoreject (4 files), ggpubr (4 files), PyPREP (4 files), SPM (4 files), afex (2 files), BayesFactor (2 files), emmeans (2 files), FSL (2 files), ggplot2 (2 files), lmerTest (2 files), xarray (2 files), AFNI (1 file), car (1 file), ggseg (1 file), Statistics and Machine Learning Toolbox (1 file), NetworkX (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
289 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 288 scripts, each with its path and the digest of its content;
  • 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Code and data availability

All MEG data analyzed in this study are openly available upon registration at https://doi.org/10.17617/1.wqa3-wk71 (http://dx.doi.org/10.17617/1.wqa3-wk71) (Liu et al. 2024). All code needed to reproduce our analyses and visuals are freely available in our GitHub repository at https://github.com/anniegbryant/MEG_functional_connectivity. Preprocessed empirical data are also shared as a Zenodo repository at https://doi.org/10.5281/zenodo.18294156 (http://dx.doi.org/10.5281/zenodo.18294156).

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 4 keywords, 68 references.

Cite

This paper

Bryant, A. G., & Whyte, C. J. (2026). A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception. Neuroscience of consciousness, 2026(1), niag029. https://doi.org/10.1093/nc/niag029

BibTeX

@article{bryant2026data,
author = {Bryant, Annie G and Whyte, Christopher J},
title = {{A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception}},
journal = {Neuroscience of consciousness},
year = {2026},
month = jul,
volume = {2026},
number = {1},
pages = {niag029},
publisher = {Oxford University Press},
issn = {2057-2107},
doi = {10.1093/nc/niag029},
url = {https://doi.org/10.1093/nc/niag029},
pmid = {42415873},
pmcid = {PMC13338905}
}

RIS

TY - JOUR
AU - Bryant, Annie G
AU - Whyte, Christopher J
TI - A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception
T2 - Neuroscience of consciousness
J2 - Neurosci Conscious
PY - 2026
DA - 2026/07/07
VL - 2026
IS - 1
SP - niag029
SN - 2057-2107
PB - Oxford University Press
DO - 10.1093/nc/niag029
UR - https://doi.org/10.1093/nc/niag029
LA - en
ER -

CSL-JSON

{
"id": "10.1093/nc/niag029",
"type": "article-journal",
"title": "A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception",
"container-title": "Neuroscience of consciousness",
"author": [
{
"family": "Bryant",
"given": "Annie G"
},
{
"family": "Whyte",
"given": "Christopher J"
}
],
"container-title-short": "Neurosci Conscious",
"volume": "2026",
"issue": "1",
"page": "niag029",
"DOI": "10.1093/nc/niag029",
"PMID": "42415873",
"PMCID": "PMC13338905",
"ISSN": "2057-2107",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/nc/niag029",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
7
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41597-026-07350-9 [code]
An open multi-center MEG-EEG dataset for studying conscious visual perception.
Journal: Scientific data
In common: PyPREP, MNE-Connectivity, MNE-BIDS, 27 other tools, MEG, 7 references
[2] doi:10.1038/s41597-026-07377-y [code]
An open-access multi-site fMRI dataset for investigating conscious visual perception.
Journal: Scientific data
In common: PyPREP, MNE-Connectivity, MNE-BIDS, 27 other tools, 5 references
[3] doi:10.1162/imag.a.1245 [code]
Towards precision EEG connectomics: Evaluating the benefits of dense sampling.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: MNE-Connectivity, autoreject, AFNI, 19 other tools, 2 references
[4] doi:10.1162/imag.a.1321 [code]
Phase similarity between similar objects indicates representational merging across retrieval training but not sleep.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: autoreject, Pingouin, easystats, 15 other tools, cognitive, 2 references
[5] doi:10.1002/hbm.70605 [code]
BrainEnrich: Revealing Biological Insights for Imaging-Derived Phenotypes Through Transcriptomic Enrichment.
Journal: Human brain mapping
In common: ggseg, easystats, broom, 17 other tools, 1 reference
[6] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: afex, Pingouin, car, 16 other tools, 2 references
[7] doi:10.1038/s41467-026-74824-0 [code]
Learning regularities in noise engages both neural predictive activity and representational changes.
Journal: Nature communications
In common: MNE-BIDS, autoreject, Pingouin, 12 other tools, MEG, cognitive, 3 references
[8] doi:10.1038/s41467-026-71830-0 [code]
Predicting individual differences of fear and cognitive learning and extinction.
Journal: Nature communications
In common: AFNI, car, broom, 14 other tools, cognitive, 1 reference
[9] doi:10.1093/cercor/bhag113 [code]
Long-term reliability and stability of parameterized resting state EEG: evidence from a five-year follow-up.
Journal: Cerebral cortex (New York, N.Y. : 1991)
In common: PyPREP, MNE-BIDS, easystats, 14 other tools, 1 reference
[10] doi:10.1038/s41467-026-73865-9 [code]
Histamine shapes the neurocomputational dynamics of human learning.
Journal: Nature communications
In common: BayesFactor, easystats, car, 13 other tools, cognitive

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.