OSCR

From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.

Code ↔ Paper

8 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 8 matches
  1. [1] § Early Studies of Normative Modeling ↔ meganorm/utils/IO.py, lines 34–496 · score 0.69 · EEG source, exponent, broadband, peak, aperiodic, alpha
  2. [2] § Contemporary Studies of Normative Modeling ↔ meganorm/src/normative_modeling.py, lines 498–603 · score 0.68 · hierarchical Bayesian regression, fitted normative, predict, batch, mapping, age
  3. [3] § Early Studies of Normative Modeling ↔ meganorm/utils/IO.py, lines 34–496 · score 0.66 · band power, spectral features, cross, head, norms, channels
  4. [4] § Early Studies of Normative Modeling ↔ notebooks/feature_extraction_tutorial.ipynb, lines 416–482 · score 0.56 · peak frequency, oscillatory, exponent, broadband, aperiodic, alpha
  5. [5] § What is Normative Modeling? ↔ meganorm/utils/nm.py, lines 206–343 · score 0.55 · evaluating model calibration, Ideally, predictive, error, covariates, scores
  6. [6] § What is Normative Modeling? ↔ meganorm/utils/nm.py, lines 204–341 · score 0.55 · evaluating model calibration, Ideally, predictive, error, covariates, scores
  7. [7] § Limitations of Contemporary Approaches ↔ meganorm/src/normative_modeling.py, lines 498–603 · score 0.51 · hierarchical Bayesian regression, neuroimaging, variables, covariates, brain
  8. [8] § Early Studies of Normative Modeling ↔ meganorm/src/mainParallel.py, lines 241–300 · score 0.51 · band power, spectral features, head, computational, demographic, channels

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,470 lines · 55 KB · GPL-3.0 · 2 matches

  1. import pickle
  2. import pandas as pd
  3. import json
  4. import os
  5. import re
  6. import glob
  7. from pathlib import Path
  8. import mne
  9. import numpy as np
  10. from typing import Literal, Dict, Tuple, List, Optional
  11. from typing import Optional
  12. import warnings
  13. from typing import Union
  14. from typing import ClassVar
  15. from pydantic import (
  16. BaseModel,
  17. Field,
  18. PositiveInt,
  19. confloat,
  20. conint,
  21. conlist,
  22. field_validator,
  23. NegativeInt,
  24. model_validator,
  25. )
  26. from mne.io.constants import FIFF
  27. class BandRatio(BaseModel):
  28. numerator: Literal["Delta", "Theta", "Alpha", "Beta", "Gamma"]
  29. denominator: Literal["Delta", "Theta", "Alpha", "Beta", "Gamma"]
  30. class Config(BaseModel):
  31. """
  32. Configuration for preprocessing, artifact correction, source localization,
  33. and spectral feature extraction of neurophysiological signals (MEG/EEG/OPM).
  34. Parameters
  35. ----------
  36. which_meg_session : int, default=0
  37. Index of the MEG session to process.
  38. which_layout : {"all", "lobe", None}, default="all"
  39. Sensor layout grouping used for reporting/analysis.
  40. which_sensor : {"mag", "grad", "meg", "eeg", "opm"}, default="meg"
  41. Sensor type to process.
  42. drop_noisy_flat_channel : bool, default=True
  43. Drop channels flagged as noisy or flat before further processing.
  44. Filtering & Resampling
  45. -----------------------
  46. cutoffFreqLow, cutoffFreqHigh : float, PositiveInt, default=1.0, 80
  47. Bandpass filter cutoff frequencies (Hz).
  48. resampling_rate : PositiveInt, default=1000
  49. Target sampling rate (Hz).
  50. digital_filter, notch_filter : bool, default=True
  51. Apply bandpass / line-noise notch filtering.
  52. apply_oversampled_temporal_projection : bool, default=True
  53. Apply oversampled temporal projection (OTP) denoising.
  54. apply_Head_movement_correction : bool, default=True
  55. Correct for head movement during recording.
  56. Head_movement_limit_from_mean : float, default=0.0015
  57. Maximum allowed deviation from mean head position (m).
  58. apply_chpi_filter : bool, default=False
  59. Filter out cHPI coil signals.
  60. ICA
  61. ---
  62. apply_ica : bool, default=True
  63. Apply ICA-based artifact correction.
  64. apply_ica_elbow_detection : bool, default=False
  65. Automatically select the number of ICA components via elbow detection.
  66. ica_n_component : PositiveInt or None, default=None
  67. Number of ICA components to compute.
  68. ica_max_iter : PositiveInt, default=800
  69. Maximum ICA iterations.
  70. ica_method : {"fastica", "infomax", "picard"}, default="fastica"
  71. ICA algorithm.
  72. ica_if_reject_by_annotation : bool, default=True
  73. Exclude annotated bad segments when fitting ICA.
  74. auto_ica_corr_thr : float, default=0.5
  75. Correlation threshold (0-1) for automatic ICA component rejection.
  76. GEDAI Artifact Removal
  77. -----------------------
  78. apply_gedai : bool, default=True
  79. Apply GEDAI-based artifact removal.
  80. gedai_method : {"both", "spectral", "broadband"}, default="both"
  81. GEDAI denoising strategy.
  82. sensai_method : {"optimize", "gridsearch"}, default="optimize"
  83. Parameter search strategy for SensAI.
  84. gedai_duration, gedai_overlap : float or int, default=12, 0.5
  85. Window duration (s) and overlap fraction for GEDAI.
  86. gedai_preliminary_broadband_noise_multiplier : float, default=6.0
  87. Noise multiplier for preliminary broadband detection.
  88. gedai_noise_multiplier : float, default=3.0
  89. Noise multiplier used in GEDAI thresholding.
  90. gedai_wavelet_type : str, default="haar"
  91. Wavelet family used for spectral GEDAI.
  92. gedai_wavelet_level : "auto", PositiveInt, or 0, default="auto"
  93. Wavelet decomposition level.
  94. gedai_wavelet_low_cutoff : float or None, default=None
  95. Low-frequency cutoff for wavelet-based denoising.
  96. gedai_epoch_size_in_cycles : PositiveInt, default=12
  97. Epoch size expressed in number of cycles.
  98. gedai_highpass_cutoff : float, default=0.1
  99. High-pass cutoff applied before GEDAI (Hz).
  100. Muscle Artifact Detection
  101. ---------------------------
  102. muscle_activity_thr : int, default=4
  103. Detection threshold for muscle artifacts.
  104. muscle_activity_min_length_good : float, default=0.1
  105. Minimum length of a clean segment retained after removal (s).
  106. muscle_activity_filter_freq : tuple[int, int], default=(110, 140)
  107. Frequency band used for muscle artifact detection (Hz).
  108. Environmental Noise Correction
  109. --------------------------------
  110. apply_environmental_noise_correction : bool, default=True
  111. Apply environmental/reference-based noise correction.
  112. ctf_gradient_comp_level : PositiveInt, default=3
  113. CTF gradient compensation level.
  114. apply_environmental_noise_ssp_with_eroom : bool, default=False
  115. Use empty-room SSP projectors for noise correction.
  116. apply_environmental_noise_ica_with_ref_meg : bool, default=True
  117. Use reference-MEG-guided ICA for environmental noise removal.
  118. environmental_noise_ica_with_ref_meg_thr : float, default=2.5
  119. Threshold for ref-MEG-guided ICA component rejection.
  120. environmental_noise_ica_with_ref_meg_method : {"together", "separate"}, default="separate"
  121. Whether to process reference channels jointly or separately.
  122. environmental_noise_ica_with_ref_meg_measure : {"zscore", "correlation"}, default="zscore"
  123. Metric used to score ICA components against reference channels.
  124. EEG Reference & Bad-Segment Rejection
  125. ----------------------------------------
  126. rereference_method : {"average", "REST", None}, default="average"
  127. EEG re-referencing scheme.
  128. bad_segment_removal_method : {"autoreject", "fixed_thr", None}, default="autoreject"
  129. Method for rejecting bad data segments.
  130. mag_var_threshold, grad_var_threshold, eeg_var_threshold : float
  131. Variance-based rejection thresholds per channel type.
  132. mag_flat_threshold, grad_flat_threshold, eeg_flat_threshold : float
  133. Flatline-detection thresholds per channel type.
  134. zscore_std_thresh : PositiveInt, default=15
  135. Z-score threshold for outlier rejection.
  136. autoreject_n_interpolates : list[int], default=[1, 4, 8, 16, 32]
  137. Candidate interpolation counts for Autoreject.
  138. autoreject_consensus_percs : list[float]
  139. Candidate consensus percentages for Autoreject (11 values, 0-1).
  140. autoreject_cv : int or "auto", default="auto"
  141. Cross-validation folds for Autoreject.
  142. autoreject_thresh_method : {"bayesian_optimization", "random_search"}, default="bayesian_optimization"
  143. Threshold search strategy for Autoreject.
  144. Segmentation
  145. ------------
  146. segments_tmin, segments_tmax : PositiveInt, NegativeInt, default=20, -20
  147. Segment start/end times relative to event (s).
  148. segments_length, segments_overlap : int, default=10, 2
  149. Segment length and overlap (s).
  150. Source Localization
  151. --------------------
  152. apply_source_localization : bool, default=False
  153. Whether to perform source localization.
  154. apply_empty_room_recording : bool, default=True
  155. Use empty-room recordings for noise covariance estimation.
  156. apply_mri_QC : bool, default=False
  157. Run quality control on MRI/FreeSurfer output.
  158. apply_mri_template : bool, default=False
  159. Use a template MRI instead of subject-specific anatomy.
  160. freesurfer_template_path, freesurfer_home, freesurfer_license : str or None
  161. Paths to FreeSurfer template derivatives, installation, and license.
  162. force_new_watershed_bem : bool, default=False
  163. Recompute the watershed BEM surfaces.
  164. gcaatlas : bool, default=True
  165. Use the GCA atlas for subcortical segmentation.
  166. SL_source_space : {"surface", "volumetric"}, default="volumetric"
  167. Type of source space.
  168. SL_conductivity : tuple[float, ...], default=(0.3,)
  169. Head-model layer conductivities (three values required for EEG).
  170. SL_inverse_operator : {"lcmv"}, default="lcmv"
  171. Inverse operator method.
  172. source_space_spacing : {"ico3"..."ico6", "oct5", "oct6"}, default="ico4"
  173. Source space resolution.
  174. source_space_spacing_number : {3, 4, 5, 6}, default=4
  175. Numeric resolution; must match `source_space_spacing`.
  176. coregisteration_final_n_iterations : int, default=20
  177. Iterations for the final coregistration refinement.
  178. coregisteration_final_nasion_weight : float, default=10.0
  179. Weight applied to the nasion fiducial during coregistration.
  180. covariance_method : str, default="empirical"
  181. Method for noise covariance estimation.
  182. Beamformer & Parcellation
  183. ----------------------------
  184. beamformer_pick_ori : {None, "normal", "max-power", "vector"}, default="max-power"
  185. Source orientation constraint.
  186. beamformer_weight_norm : {None, "unit-noise-gain", "nai", "unit-noise-gain-invariant"}, default="unit-noise-gain"
  187. Beamformer weight normalization.
  188. beamforme_depth : float, default=0.08
  189. Depth-weighting factor correcting for center-of-head bias.
  190. inverse_regularization_value : float, default=0.05
  191. Regularization applied to the data covariance matrix.
  192. apply_morphing : bool, default=False
  193. Morph source estimates to a common template brain.
  194. parcellation_parc : {None, "aparc.a2009s", "parac"}, default="aparc.a2009s"
  195. Predefined cortical parcellation.
  196. parcellation_annot_fname : Path or None
  197. Custom parcellation (.annot) file, used if `parcellation_parc` is None.
  198. PSD & Spectral Parametrization
  199. ---------------------------------
  200. psd_method : {"multitaper", "welch"}, default="welch"
  201. PSD estimation method.
  202. psd_n_overlap, psd_n_fft, psd_n_per_seg : PositiveInt, default=1, 2, 2
  203. Welch/multitaper PSD parameters.
  204. parametrization_method : {"fooof", "irasa"}, default="irasa"
  205. Method for separating aperiodic and periodic spectral components.
  206. irasa_hset : tuple[float, float, float], default=(1.05, 2.0, 0.05)
  207. Resampling factor range/step for IRASA.
  208. fooof_freq_range_low, fooof_freq_range_high : PositiveInt, default=3, 40
  209. Frequency range for FOOOF fitting (Hz).
  210. aperiodic_mode : {"knee", "fixed"}, default="knee"
  211. Aperiodic component model.
  212. fooof_peak_width_limits : list[float], default=[1.0, 12.0]
  213. Allowed peak width range (Hz).
  214. fooof_min_peak_height : int, default=0
  215. Minimum peak height for detection.
  216. fooof_peak_threshold : PositiveInt, default=2
  217. Peak detection threshold (in SD of the flattened spectrum).
  218. fooof_res_save_path : str or None
  219. Path to save FOOOF results.
  220. save_source_localized_epochs, save_psds : bool, default=False
  221. Persist intermediate source-localized epochs / PSDs to disk.
  222. Feature Extraction
  223. -------------------
  224. freq_bands : dict[str, tuple[int, int]]
  225. Canonical frequency band definitions (Theta, Alpha, Beta, Gamma).
  226. individualized_band_ranges : dict[str, tuple[int, int]]
  227. Per-band offsets (Hz) used to individualize canonical bands.
  228. power_band_ratios_list : list[BandRatio]
  229. Band-power ratios to compute (e.g. Theta/Beta).
  230. min_r_squared : float, default=0.9
  231. Minimum R² required to accept a spectral model fit.
  232. feature_categories : dict[str, bool]
  233. Flags selecting which feature families to extract (offset, exponent,
  234. peak parameters, canonical/individualized band power, band ratios,
  235. hemispheric asymmetry).
  236. Miscellaneous
  237. -------------
  238. random_state : int, default=42
  239. Random seed for reproducibility.
  240. Methods
  241. -------
  242. save(save_path, overwrite=False)
  243. Serialize the configuration to a JSON file.
  244. load(path)
  245. Load a configuration from a JSON file.
  246. Notes
  247. -----
  248. Model validators enforce cross-field consistency, e.g.: a three-layer
  249. `SL_conductivity` is required for EEG source localization; `beamformer_pick_ori
  250. == "vector"` requires `beamformer_weight_norm == "unit-noise-gain-invariant"`;
  251. `source_space_spacing` must match `source_space_spacing_number`; GEDAI
  252. parameters must be consistent with the chosen `gedai_method`; and MRI
  253. template use is mutually exclusive with MRI QC.
  254. """
  255. model_config = {"extra": "forbid"}
  256. which_meg_session: int = 0 # the first session
  257. which_layout: Literal["all", "lobe", None] = "all"
  258. which_sensor: Literal["mag", "grad", "meg", "eeg", "opm"] = "meg"
  259. drop_noisy_flat_channel: bool = True
  260. remove_nonfinite_segment_threshold: PositiveInt = 5
  261. # ICA
  262. apply_ica_elbow_detection: bool = False
  263. ica_n_component: Optional[PositiveInt] = None
  264. ica_max_iter: PositiveInt = 800
  265. ica_method: Literal["fastica", "infomax", "picard"] = "fastica"
  266. cutoffFreqLow: float = 1.0
  267. cutoffFreqHigh: PositiveInt = 80
  268. resampling_rate: PositiveInt = 1000
  269. digital_filter: bool = True
  270. notch_filter: bool = True
  271. apply_oversampled_temporal_projection: bool = True
  272. apply_Head_movement_correction: bool = True
  273. Head_movement_limit_from_mean: float = 0.0015
  274. apply_chpi_filter: bool = False
  275. # gedai settings
  276. apply_gedai: bool = True
  277. gedai_method: Literal["both", "spectral", "broadband"] = "both"
  278. sensai_method: Literal["optimize", "gridsearch"] = "optimize"
  279. gedai_duration: Union[float, int] = 12
  280. gedai_overlap: Union[float, int] = 0.5
  281. gedai_preliminary_broadband_noise_multiplier: float = 6.0
  282. gedai_noise_multiplier: float = 3.0
  283. gedai_wavelet_type: str = "haar"
  284. gedai_wavelet_level: Union[Literal["auto"], PositiveInt, Literal[0]] = "auto"
  285. gedai_wavelet_low_cutoff: Union[None, float] = None
  286. gedai_epoch_size_in_cycles: PositiveInt = 12
  287. gedai_highpass_cutoff: float = 0.1
  288. muscle_activity_thr: int = 4
  289. muscle_activity_min_length_good: float = 0.1
  290. muscle_activity_filter_freq: Tuple[int, int] = (110, 140)
  291. apply_environmental_noise_correction: bool = True
  292. same_environmental_noise_removal: bool = False
  293. ctf_gradient_comp_level: PositiveInt = 3
  294. apply_environmental_noise_ssp_with_eroom: bool = False
  295. apply_environmental_noise_ica_with_ref_meg: bool = True
  296. environmental_noise_ica_with_ref_meg_thr: float = 2.0
  297. ica_if_reject_by_annotation: bool = True
  298. environmental_noise_ica_with_ref_meg_method: Literal["together", "separate"] = (
  299. "separate"
  300. )
  301. environmental_noise_ica_with_ref_meg_measure: Literal["zscore", "correlation"] = (
  302. "zscore"
  303. )
  304. apply_ica: bool = True
  305. auto_ica_corr_thr: confloat(ge=0, le=1) = 0.5
  306. save_segmented_data: bool = False
  307. rereference_method: Literal["average", "REST", None] = "average"
  308. bad_segment_removal_method: Literal["autoreject", "fixed_thr", None] = "autoreject"
  309. mag_var_threshold: float = 5000e-15
  310. grad_var_threshold: float = 5000e-13
  311. eeg_var_threshold: float = 40e-6
  312. mag_flat_threshold: float = 10e-15
  313. grad_flat_threshold: float = 10e-13
  314. eeg_flat_threshold: float = 40e-6
  315. zscore_std_thresh: PositiveInt = 15
  316. segments_tmin: PositiveInt = 20
  317. segments_tmax: NegativeInt = -20
  318. segments_length: PositiveInt = 10
  319. segments_overlap: int = 2
  320. save_preprocessed_data: bool = True
  321. # autoreject
  322. autoreject_n_interpolates: List[int] = [1, 4, 8, 16, 32]
  323. autoreject_consensus_percs: List[float] = list(np.linspace(0, 1.0, 11))
  324. autoreject_cv: Union[int, Literal["auto"]] = "auto"
  325. autoreject_thresh_method: Literal["bayesian_optimization", "random_search"] = (
  326. "bayesian_optimization"
  327. )
  328. # Source localization
  329. apply_source_localization: bool = False
  330. apply_empty_room_recording: bool = True
  331. apply_mri_QC: bool = False
  332. take_screenshot_of_coregisteration: bool = True
  333. apply_mri_template: bool = False
  334. freesurfer_template_path: Optional[str] = None
  335. freesurfer_home: Optional[str] = None
  336. freesurfer_license: Optional[str] = None
  337. coregisteration_scale_mode: Literal["uniform", "3-axis", None] = None
  338. force_new_watershed_bem: bool = False
  339. gcaatlas: bool = True
  340. SL_source_space: Literal["surface", "volumetric"] = "volumetric"
  341. SL_conductivity: Tuple[float, ...] = (0.3,)
  342. SL_inverse_operator: Literal["lcmv"] = "lcmv"
  343. bem_preflood_parameter_space: List[int] = (10, 15, 20, 30, 35)
  344. bem_preflood: int = 25
  345. bem_gcaatlas: bool = True
  346. bem_max_erosion_pct: float = 15.0
  347. bem_suspect_erosion_pct: float = 0.2
  348. bem_max_fine_segmentation_iteration: int = 100
  349. bem_plot_orientations: Literal["coronal", "axial", "sagittal", None] = "coronal"
  350. # the spacing to use for source space specificatin
  351. source_space_spacing: Literal[
  352. "ico3", "ico4", "ico5", "ico6", "oct5", "oct6", "all"
  353. ] = "ico4"
  354. source_space_spacing_number: Literal[3, 4, 5, 6, None] = 4
  355. source_space_add_dist: Literal["patch", True, False] = "patch"
  356. save_transformation_FIF_file: bool = False
  357. coregisteration_final_n_iterations: int = 20
  358. coregisteration_final_nasion_weight: float = 10.0
  359. covariance_method: str = "empirical"
  360. # Determines whether to keep vectors for all source orientations
  361. # or to select a single fixed orientation, depending on the chosen algorithm.
  362. beamformer_pick_ori: Literal[None, "normal", "max-power", "vector"] = "max-power"
  363. beamformer_weight_norm: Literal[
  364. None, "unit-noise-gain", "nai", "unit-noise-gain-invariant"
  365. ] = "unit-noise-gain-invariant"
  366. # This parameter scales the activation to correct for head-center bias.
  367. beamforme_depth: confloat(ge=0, le=1) = 0.8
  368. # this is used for regularaizing the data covariance (shifting the matrix)
  369. inverse_regularization_value: confloat(ge=0, le=1) = 0.05
  370. apply_morphing: bool = False
  371. # the pacellation to use
  372. parcellation_parc: Literal[None, "aparc.a2009s", "parac"] = "aparc.a2009s"
  373. parcellation_mode: Literal["mean_flip", "mean", "auto", "pca_flip", "max"] = "auto"
  374. # A custom parcellation file
  375. parcellation_annot_fname: Optional[Path] = None
  376. # PSD
  377. psd_method: Literal["multitaper", "welch"] = "welch"
  378. psd_n_overlap: PositiveInt = 1
  379. psd_n_fft: PositiveInt = 2
  380. psd_n_per_seg: PositiveInt = 2
  381. parametrization_method: Literal["fooof", "irasa"] = "irasa"
  382. # PYRASA
  383. irasa_hset: Tuple[float, float, float] = (1.05, 2.0, 0.05)
  384. # FOOOF analysis
  385. fooof_freq_range_low: PositiveInt = 3
  386. fooof_freq_range_high: PositiveInt = 40
  387. aperiodic_mode: Literal["knee", "fixed"] = "knee"
  388. fooof_peak_width_limits: List[float] = [1.0, 12.0]
  389. fooof_min_peak_height: int = 0
  390. fooof_peak_threshold: PositiveInt = 2
  391. save_source_localized_epochs: bool = False
  392. save_psds: bool = False
  393. lowest_num_of_epochs: Optional[int] = None
  394. # Feature extraction
  395. freq_bands: Dict[str, Tuple[int, int]] = {
  396. "Theta": (3, 8),
  397. "Alpha": (8, 13),
  398. "Beta": (13, 30),
  399. "Gamma": (30, 40),
  400. }
  401. individualized_band_ranges: Dict[str, Tuple[int, int]] = {
  402. "Theta": (-2, 3),
  403. "Alpha": (-2, 3),
  404. "Beta": (-8, 9),
  405. "Gamma": (-5, 5),
  406. }
  407. power_band_ratios_list: List[BandRatio] = [
  408. BandRatio(numerator="Theta", denominator="Beta"),
  409. BandRatio(numerator="Theta", denominator="Alpha"),
  410. BandRatio(numerator="Alpha", denominator="Beta"),
  411. BandRatio(numerator="Delta", denominator="Beta"),
  412. BandRatio(numerator="Delta", denominator="Alpha"),
  413. BandRatio(numerator="Delta", denominator="Theta"),
  414. ]
  415. min_r_squared: confloat(ge=0, le=1) = 0.9
  416. feature_categories: Dict[str, bool] = {
  417. "Offset": True,
  418. "Exponent": True,
  419. "Peak_Center": False,
  420. "Peak_Power": False,
  421. "Peak_Width": False,
  422. "Adjusted_Canonical_Relative_Power": True,
  423. "Adjusted_Canonical_Absolute_Power": True,
  424. "Adjusted_Individualized_Relative_Power": False,
  425. "Adjusted_Individualized_Absolute_Power": False,
  426. "OriginalPSD_Canonical_Relative_Power": True,
  427. "OriginalPSD_Canonical_Absolute_Power": True,
  428. "OriginalPSD_Individualized_Relative_Power": False,
  429. "OriginalPSD_Individualized_Absolute_Power": False,
  430. "Adjusted_Band_Ratio": True,
  431. "OriginalPSD_Band_Ratio": True,
  432. "Hemispheric_Asymmetry_index": True,
  433. "Knee_Frequency": False,
  434. }
  435. fooof_res_save_path: Optional[str] = None
  436. random_state: int = 42
  437. @field_validator("muscle_activity_thr")
  438. def muscle_activity_thr_fv(cls, v):
  439. if v < 3:
  440. warnings.warn(
  441. "Select a higher threshold for muscle activity artifacts. Low values "
  442. "remove clean data."
  443. )
  444. return v
  445. @field_validator("muscle_activity_filter_freq")
  446. def muscle_activity_filter_freq_fv(cls, v):
  447. if v[0] < 100:
  448. warnings.warn(
  449. f"Muscle activity artifact affects higher frequencies than {v[0]} Hz."
  450. )
  451. return v
  452. @model_validator(mode="after")
  453. def SL_conductivity_mv(self):
  454. if (
  455. len(self.SL_conductivity) == 1
  456. and self.which_sensor == "eeg"
  457. and self.apply_source_localization
  458. ):
  459. raise ValueError(
  460. "In the case of EEG, you must have a three layers conductivity model due to volume conduction."
  461. )
  462. return self
  463. @model_validator(mode="after")
  464. def source_space_res(self):
  465. if self.source_space_spacing == "all":
  466. if self.source_space_spacing_number is not None:
  467. raise ValueError(
  468. "source_space_spacing_number must be None when source_space_spacing='all'"
  469. )
  470. return self
  471. if int(self.source_space_spacing[-1]) != self.source_space_spacing_number:
  472. raise ValueError(
  473. "The source_space_spacing and source_space_spacing_number should match"
  474. )
  475. return self
  476. @model_validator(mode="after")
  477. def beamformer_arg_check(self):
  478. if (
  479. self.beamformer_pick_ori == "vector"
  480. and self.beamformer_weight_norm != "unit-noise-gain-invariant"
  481. ):
  482. error_msg = (
  483. "If you wish to compute a vector beamformer, it is necessary to use"
  484. " unit-noise-gain-invariant for weight_norm argument. This is for addressing"
  485. " the center of head bias where deeper sources can have larger scale than"
  486. " superfacial sources."
  487. )
  488. raise ValueError(error_msg)
  489. return self
  490. @model_validator(mode="after")
  491. def center_head_bias_scale_check(self):
  492. if not self.beamforme_depth and self.SL_source_space == "volumetric":
  493. error_msg = (
  494. "If you want to use volumetric source space (interested in deeper sources),"
  495. " please define beamforme_depth as positive float number, i.e., 0.8. This is used to address"
  496. " the center of head bias."
  497. )
  498. raise ValueError(error_msg)
  499. return self
  500. @model_validator(mode="after")
  501. def ica_e_noise_removal(self):
  502. if (
  503. self.environmental_noise_ica_with_ref_meg_measure == "correlation"
  504. and not 0 < self.environmental_noise_ica_with_ref_meg_thr < 1
  505. ):
  506. error_mg = (
  507. "If the threshold method for removing environmental noise using "
  508. "ICA and ref-MEG is correlation, the corresponding measure must be between 0 and 1."
  509. )
  510. raise ValueError(error_mg)
  511. return self
  512. @model_validator(mode="after")
  513. def pacellation_checker(self):
  514. if not self.parcellation_parc and not self.parcellation_annot_fname:
  515. raise ValueError(
  516. "Parcellation should be passed. Otherwise pass a custom parcellation file (.annot)"
  517. )
  518. return self
  519. @model_validator(mode="after")
  520. def mri_template_check(self):
  521. if self.apply_mri_template and not self.freesurfer_template_path:
  522. err_msg = (
  523. "In order to apply template MRI for source localization, a path to already freesurfer-preprocessed"
  524. "derivatives is necessary"
  525. )
  526. raise ValueError(err_msg)
  527. return self
  528. @model_validator(mode="after")
  529. def either_template_mri_or_mri_qc(self):
  530. if self.apply_mri_QC and self.apply_mri_template:
  531. err_msg = (
  532. "You can not apply MRI QC on already preprocessed freesurfer template"
  533. )
  534. raise ValueError(err_msg)
  535. return self
  536. @model_validator(mode="after")
  537. def gedai_params_check(self):
  538. method = self.gedai_method
  539. wavelet_level = self.gedai_wavelet_level
  540. duration = self.gedai_duration
  541. broadband_multiplier = self.gedai_preliminary_broadband_noise_multiplier
  542. if method == "broadband" and wavelet_level != 0:
  543. raise ValueError("broadband method requires wavelet_level=0")
  544. if method == "broadband" and not duration:
  545. raise ValueError("broadband method requires gedai_duration")
  546. if method == "spectral" and wavelet_level == 0:
  547. raise ValueError("spectral method requires wavelet_level > 0")
  548. if method == "both" and not broadband_multiplier:
  549. raise ValueError(
  550. "both method requires gedai_preliminary_broadband_noise_multiplier"
  551. )
  552. return self
  553. def save(self, save_path: str, overwrite=False):
  554. "save the configurations to a JSON file"
  555. if os.path.exists(save_path) and overwrite == False:
  556. err_msg = f"A configuration file already exists in this directory: {save_path}. Set the overwrite to True."
  557. raise FileExistsError(err_msg)
  558. with open(save_path, "w") as file:
  559. file.write(self.model_dump_json(indent=4))
  560. @classmethod
  561. def load(cls, path: str):
  562. # Load configuration from a JSON file
  563. with open(path, "r") as file:
  564. cfg = json.load(file)
  565. return cls(**cfg)
  566. def _resolve_artemis_pos_file(path, pos_file, logger):
  567. """Return an explicit pos/dig file, or find one near the recording."""
  568. if pos_file:
  569. return pos_file
  570. # .pos usually lives in a sibling Headshape/ folder, not next to the run,
  571. # so widen the search one directory up before giving up.
  572. search_roots = [Path(path).parent, Path(path).parents[1]]
  573. for root in search_roots:
  574. matches = sorted(glob.glob(f"{root}/**/*.pos", recursive=True))
  575. if matches:
  576. if len(matches) > 1:
  577. logger.warning(
  578. f"{len(matches)} .pos files found under {root}; using {matches[0]}."
  579. )
  580. return matches[0]
  581. err_msg = (
  582. f"No .pos file found near the Artemis recording {path} "
  583. f"(searched {', '.join(str(r) for r in search_roots)})."
  584. )
  585. logger.error(err_msg)
  586. raise FileNotFoundError(err_msg)
  587. def _add_artemis_headshape(data, path, pos_file, logger):
  588. """Attach head-shape points to an already-converted Artemis FIF."""
  589. existing = data.info["dig"] or []
  590. n_extra = sum(1 for d in existing if d["kind"] == FIFF.FIFFV_POINT_EXTRA)
  591. if n_extra:
  592. logger.info(
  593. f"Artemis FIF already carries {n_extra} head-shape points; "
  594. "not adding any from a .pos/.fif file."
  595. )
  596. return data
  597. resolved = _resolve_artemis_pos_file(path, pos_file, logger)
  598. ext = str(resolved).split(".")[-1].lower()
  599. if ext == "pos":
  600. digs = mne.io.artemis123.utils._read_pos(fname=resolved)
  601. with data.info._unlock():
  602. data.info["dig"] = existing + digs
  603. data.info["dig"] = mne._fiff._digitization._format_dig_points(
  604. data.info["dig"]
  605. )
  606. logger.info(f"Head shape .pos file {resolved} was added to the Artemis FIF.")
  607. elif ext == "fif":
  608. dig = mne.channels.read_dig_fif(resolved)
  609. data.set_montage(dig, on_missing="warn")
  610. logger.info(f"Head shape .fif file {resolved} was added to the Artemis FIF.")
  611. else:
  612. logger.warning(
  613. f"pos_file '{resolved}' has unrecognized extension '.{ext}'; "
  614. "no head-shape/dig info was added to the recording."
  615. )
  616. return data
  617. def select_session_path(pattern, index=0):
  618. """
  619. Resolve a '*'-separated glob pattern to a single path.
  620. Returns None when `pattern` is falsy, so callers can pass optional
  621. CLI arguments straight through.
  622. """
  623. if not pattern:
  624. return None
  625. candidates = [p for p in pattern.split("*") if p]
  626. return candidates[index]
  627. def infer_device(path, device_type, which_sensor, logger):
  628. """Determine the acquisition device from an explicit override or the path."""
  629. MEG_DEVICE_BY_EXTENSION = {"ds": "CTF", "fif": "MEGIN", "bin": "ARTEMIS123"}
  630. if which_sensor == "eeg":
  631. # TODO: was originally path[0]. Check if this correction is correct.
  632. return path.split(".")[-1].upper()
  633. if device_type:
  634. return device_type.upper()
  635. if "4D" in path:
  636. return "BTI"
  637. extension = path.split(".")[-1]
  638. if extension in MEG_DEVICE_BY_EXTENSION:
  639. return MEG_DEVICE_BY_EXTENSION[extension]
  640. err_msg = "The provided MEG recording is not supported yet."
  641. logger.error(err_msg)
  642. raise ValueError(err_msg)
  643. def load_recording(
  644. device, path, empty_room_recording_path, configs, logger, pos_file=None
  645. ):
  646. """Load data"""
  647. if device == "CTF":
  648. data = mne.io.read_raw_ctf(path, preload=True)
  649. if pos_file:
  650. ext = pos_file.split(".")[-1]
  651. if ext == "pos":
  652. digs = mne.io.artemis123.utils._read_pos(fname=pos_file)
  653. with data.info._unlock():
  654. data.info["dig"] += digs
  655. data.info["dig"] = mne._fiff._digitization._format_dig_points(
  656. data.info["dig"]
  657. )
  658. logger.info(
  659. "Head shape .pos file was found and added to the CTF recording"
  660. )
  661. elif ext == "fif":
  662. dig = mne.channels.read_dig_fif(pos_file)
  663. data.set_montage(dig, on_missing="warn")
  664. logger.info(
  665. "Head shape .fif file was found and added to the CTF recording"
  666. )
  667. else:
  668. logger.warning(
  669. f"pos_file '{pos_file}' has unrecognized extension '.{ext}'; "
  670. "no head-shape/dig info was added to the recording."
  671. )
  672. if empty_room_recording_path and (
  673. configs.apply_source_localization
  674. or configs.apply_environmental_noise_ssp_with_eroom
  675. ):
  676. empty_room_recording = mne.io.read_raw_ctf(
  677. empty_room_recording_path, preload=True
  678. )
  679. logger.info("Empty room recording was found")
  680. elif not empty_room_recording_path and (
  681. configs.apply_source_localization
  682. or configs.apply_environmental_noise_ssp_with_eroom
  683. ):
  684. empty_room_recording = None
  685. logger.warning("No empty room recording was found")
  686. else:
  687. empty_room_recording = None
  688. elif device == "BTI":
  689. temp_hs_file_path = os.path.join(path, "hs_file")
  690. hs_file = temp_hs_file_path if os.path.exists(temp_hs_file_path) else None
  691. if not hs_file:
  692. convert = False
  693. else:
  694. convert = True
  695. data = mne.io.read_raw_bti(
  696. pdf_fname=os.path.join(path, "c,rfDC"),
  697. config_fname=os.path.join(path, "config"),
  698. head_shape_fname=hs_file,
  699. preload=True,
  700. convert=convert,
  701. )
  702. if empty_room_recording_path and configs.apply_source_localization:
  703. empty_room_recording = mne.io.read_raw_bti(
  704. pdf_fname=os.path.join(empty_room_recording_path, "c,rfDC"),
  705. config_fname=os.path.join(empty_room_recording_path, "config"),
  706. head_shape_fname=None,
  707. convert=convert,
  708. preload=True,
  709. )
  710. logger.info("Empty room recording was found")
  711. elif not empty_room_recording_path and configs.apply_source_localization:
  712. empty_room_recording = None
  713. logger.info("No empty room recording was found")
  714. else:
  715. empty_room_recording = None
  716. elif device == "ARTEMIS123":
  717. # Artemis recordings may arrive as the native .bin or as a FIF that was
  718. # converted beforehand. read_raw_artemis123 only accepts the .bin, so
  719. # anything already in FIF goes through the standard FIF reader.
  720. if str(path).lower().endswith(".fif"):
  721. data = mne.io.read_raw_fif(path, preload=True)
  722. if configs.apply_source_localization:
  723. data = _add_artemis_headshape(data, path, pos_file, logger)
  724. else:
  725. if configs.apply_source_localization:
  726. data = mne.io.read_raw_artemis123(
  727. path,
  728. preload=True,
  729. pos_fname=_resolve_artemis_pos_file(path, pos_file, logger),
  730. add_head_trans=True,
  731. )
  732. else:
  733. data = mne.io.read_raw_artemis123(
  734. path,
  735. preload=True,
  736. )
  737. if empty_room_recording_path and configs.apply_source_localization:
  738. if str(empty_room_recording_path).lower().endswith(".fif"):
  739. empty_room_recording = mne.io.read_raw_fif(
  740. empty_room_recording_path, preload=True
  741. )
  742. else:
  743. empty_room_recording = mne.io.read_raw_artemis123(
  744. empty_room_recording_path,
  745. preload=True,
  746. )
  747. logger.info("Empty room recording was found")
  748. elif not empty_room_recording_path and configs.apply_source_localization:
  749. empty_room_recording = None
  750. logger.info("No empty room recording was found")
  751. else:
  752. empty_room_recording = None
  753. else:
  754. data = mne.io.read_raw(path, preload=True)
  755. if empty_room_recording_path and configs.apply_source_localization:
  756. empty_room_recording = mne.io.read_raw(
  757. empty_room_recording_path, preload=True
  758. )
  759. logger.info("Empty room recording was found")
  760. elif not empty_room_recording_path and configs.apply_source_localization:
  761. empty_room_recording = None
  762. logger.info("No empty room recording was found")
  763. else:
  764. empty_room_recording = None
  765. return data, empty_room_recording
  766. def merge_fidp_demo(
  767. demographic_paths: list,
  768. features_dir: str,
  769. dataset_names: list,
  770. drop_columns: list = ["eyes"],
  771. ):
  772. """
  773. Merge demographic metadata and extracted features into a single DataFrame.
  774. This function loads demographic data and feature data,
  775. assigns a site label to each participant if missing, removes unnecessary columns,
  776. and merges demographic information with corresponding extracted features.
  777. Parameters
  778. ----------
  779. datasets_paths : list
  780. List of paths to the dataset directories containing demographic files
  781. ('participants_bids.tsv').
  782. features_dir : str
  783. Path to the directory containing the extracted features ('all_features.csv').
  784. dataset_names : list of str
  785. List of dataset names corresponding to each dataset path. Used to populate
  786. missing 'site' information if necessary.
  787. drop_columns : list of str, optional
  788. Columns to drop from the demographic data before merging. Default is ["eyes"].
  789. Returns
  790. -------
  791. data : pandas.DataFrame
  792. Merged DataFrame containing both demographic information and feature data,
  793. with participants indexed as strings.
  794. Raises
  795. ------
  796. FileNotFoundError
  797. If the 'participants_bids.tsv' file is missing in any of the dataset paths or
  798. the 'all_features.csv' file is missing in the provided features directory.
  799. """
  800. demographic_df = pd.DataFrame()
  801. for counter, demo_path in enumerate(demographic_paths):
  802. if not demo_path or not os.path.exists(demo_path):
  803. raise FileNotFoundError(
  804. f"The demographic file for dataset '{dataset_names[counter]}' was not "
  805. f"found at: {demo_path}. Set 'demographic_path' for this dataset, or "
  806. "create the file with 'make_demo_file_bids'."
  807. )
  808. demo = load_demographic_file(demo_path)
  809. if "site" not in demo.columns:
  810. demo["site"] = dataset_names[counter]
  811. demographic_df = pd.concat([demographic_df, demo], axis=0)
  812. # Drop unnecessary columns
  813. if drop_columns:
  814. demographic_df.drop(columns=drop_columns, errors="ignore", inplace=True)
  815. # Load features
  816. feature_path = os.path.join(features_dir, "all_features.csv")
  817. if not os.path.exists(feature_path):
  818. raise FileNotFoundError(
  819. f"The file 'all_features.csv' is missing in the directory: {features_dir}."
  820. )
  821. features_df = pd.read_csv(feature_path, index_col=0)
  822. features_df.index = features_df.index.astype(str)
  823. # Merge demographic and features
  824. data = demographic_df.join(features_df, how="inner")
  825. data.index.name = None
  826. return data
  827. def merge_datasets_with_glob(datasets):
  828. """
  829. Merges file paths across multiple datasets using glob pattern matching.
  830. This function walks through the provided datasets' base directories to find
  831. subject folders and file paths matching a specified task and file ending. It
  832. creates a dictionary mapping each subject to a glob pattern that can be used
  833. to aggregate files across multiple runs or sessions.
  834. Parameters
  835. ----------
  836. datasets : dict
  837. Dictionary where each key is a dataset name, and each value is a dictionary
  838. with the following keys:
  839. - "base_dir" (str): Base directory containing subject subdirectories.
  840. - "task" (str): Task keyword to search for in filenames.
  841. - "ending" (str): File ending (e.g., '.nii.gz') to filter relevant files.
  842. Returns
  843. -------
  844. dict
  845. A dictionary mapping subject IDs to a glob-style path string that aggregates
  846. all matching files for that subject. Only subjects with at least one matched
  847. file are included.
  848. Notes
  849. -----
  850. This function is designed to assist in scenarios where each subject may have
  851. multiple files (e.g., different runs or sessions), and the goal is to create
  852. a single pattern that can be used to load all related files for a subject.
  853. """
  854. def join_with_star(lst):
  855. if not lst:
  856. return None
  857. if len(lst) == 1:
  858. return lst[0] + "*"
  859. return "*".join(lst)
  860. subjects = {}
  861. for dataset_name, dataset_info in datasets.items():
  862. base_dir = dataset_info["base_dir"]
  863. task = dataset_info["task"]
  864. ending = dataset_info["ending"]
  865. device = dataset_info.get("device_type", None)
  866. line_freq = dataset_info.get("line_freq", 50)
  867. empty_room_task = dataset_info.get("empty_room_task", None)
  868. empty_room_path = dataset_info.get("empty_room_path", None)
  869. empty_room_ending = dataset_info.get("empty_room_ending", None)
  870. surfaces = dataset_info.get("surfaces_dir", None)
  871. event_file_path = dataset_info.get("event_file_path", None)
  872. event_file_task = dataset_info.get("event_file_task", None)
  873. event_file_ending = dataset_info.get("event_file_ending", None)
  874. event_of_interest = dataset_info.get("event_of_interest", None)
  875. trans_file_p = dataset_info.get("trans_path", None)
  876. pos_file_p = dataset_info.get("pos_path", None)
  877. pos_file_ending = dataset_info.get("pos_file_ending", None)
  878. annotation_p = dataset_info.get("annotation_path", None)
  879. annotaion_task_name = dataset_info.get("annotaion_task_name", None)
  880. annotation_ending = dataset_info.get("annotation_ending", None)
  881. layout_path = dataset_info.get("layout_path", None)
  882. demographic_path = dataset_info.get(
  883. "demographic_path", os.path.join(base_dir, "participants_bids.tsv")
  884. )
  885. dirs = [
  886. d for d in os.listdir(base_dir) if os.path.isdir(os.path.join(base_dir, d))
  887. ]
  888. for subj in dirs:
  889. # resting state data
  890. rs_record_paths = glob.glob(
  891. f"{base_dir}/{subj}/**/*{task}*{ending}", recursive=True
  892. )
  893. # empty room record
  894. if empty_room_task:
  895. er_record_paths = glob.glob(
  896. f"{empty_room_path}/{subj}/**/*{empty_room_task}*{empty_room_ending}",
  897. recursive=True,
  898. )
  899. else:
  900. er_record_paths = None
  901. # freesurfer surfaces
  902. if surfaces:
  903. if os.path.isdir(os.path.join(surfaces, subj)):
  904. surface = surfaces
  905. else:
  906. surface = None
  907. else:
  908. surface = None
  909. # event file
  910. if event_file_task and event_file_path:
  911. event_record_paths = glob.glob(
  912. f"{event_file_path}/{subj}/**/*{event_file_task}*{event_file_ending}",
  913. recursive=True,
  914. )
  915. else:
  916. event_record_paths = None
  917. # trans file
  918. if trans_file_p:
  919. trans_path = glob.glob(
  920. f"{trans_file_p}/{subj}/**/*-trans.fif", recursive=True
  921. )
  922. else:
  923. trans_path = None
  924. # pos file
  925. if pos_file_p and pos_file_ending:
  926. pos_path = glob.glob(
  927. f"{pos_file_p}/{subj}/**/*{pos_file_ending}", recursive=True
  928. )
  929. else:
  930. pos_path = None
  931. if annotation_p:
  932. annotation_path = glob.glob(
  933. f"{annotation_p}/*{subj}*/**/*{annotaion_task_name}*{annotation_ending}",
  934. recursive=True,
  935. )
  936. else:
  937. annotation_path = None
  938. subjects.update(
  939. {
  940. subj: {
  941. "rest_record": join_with_star(rs_record_paths),
  942. "device": device,
  943. "line_freq": str(line_freq),
  944. "empty_room_record": join_with_star(er_record_paths),
  945. "mri_surface": surface,
  946. "dataset_name": dataset_name,
  947. "event_record": join_with_star(event_record_paths),
  948. "event_of_interest": str(event_of_interest),
  949. "trans_path": join_with_star(trans_path),
  950. "pos_path": join_with_star(pos_path),
  951. "annotation_path": join_with_star(annotation_path),
  952. "layout_path": layout_path,
  953. "demographic_path": demographic_path,
  954. }
  955. }
  956. )
  957. return subjects
  958. def load_demographic_file(path, index_col=0):
  959. """Read a participants/demographic table (.tsv, .txt, .csv, .xlsx)."""
  960. ext = os.path.splitext(path)[1].lower()
  961. if ext in (".tsv", ".txt"):
  962. df = pd.read_csv(path, sep="\t", index_col=index_col)
  963. elif ext == ".csv":
  964. df = pd.read_csv(path, index_col=index_col)
  965. elif ext in (".xlsx", ".xls"):
  966. df = pd.read_excel(path, index_col=index_col)
  967. else:
  968. raise ValueError(f"Unsupported demographic file type: {path}")
  969. df.index = df.index.astype(str)
  970. return df
  971. def make_demo_file_bids(
  972. file_dir: str, save_dir: str, id_col: int, age_col: int, *columns
  973. ) -> None:
  974. """
  975. Convert formats of demographic data into a single format so it can be used
  976. in later stages.
  977. Parameters
  978. ----------
  979. file_dir : str
  980. Path to the input demographic file (supports CSV, TSV, or XLSX).
  981. save_dir : str
  982. Path where the BIDS-formatted demographic file will be saved (as TSV).
  983. id_col : int
  984. Column index containing the participant ID.
  985. age_col : int
  986. Column index containing participant age.
  987. *extra_columns : dict
  988. Additional column definitions. While age and participants id were defined
  989. using positional arguments, extra coulmn modification (e.g., sex and eyes
  990. condition) can be revised and converted to a single format across dataset
  991. using this function. Each dict can contain:
  992. - 'col_name': str, required name for the output column. This does not
  993. necessarly match the column name before being passed to this function.
  994. - 'col_id': int, index of the column that the revision should be applied to.
  995. - 'single_value': value to assign to all rows if no col_id and mapping are given.
  996. This can be helpful when all subjects in a dataset have the same properties
  997. e.g., eyes open condition.
  998. - 'mapping': dict, if single value is not defined, value mapping can be passed
  999. to map the initial values to the target values.
  1000. Returns
  1001. -------
  1002. None
  1003. """
  1004. for col in columns:
  1005. if col.get("single_value") and col.get("mapping"):
  1006. raise ValueError(
  1007. "'single_value' and 'mapping' can not be both defined. One of them must be None; see the documentation!"
  1008. )
  1009. # Load input file based on extension
  1010. if file_dir.endswith(".xlsx"):
  1011. df = pd.read_excel(file_dir)
  1012. elif file_dir.endswith(".csv"):
  1013. df = pd.read_csv(file_dir)
  1014. elif file_dir.endswith(".tsv"):
  1015. df = pd.read_csv(file_dir, sep="\t")
  1016. else:
  1017. raise ValueError(f"Unsupported file type for: {file_dir}")
  1018. # Initialize new dataframe with required fields
  1019. new_df = pd.DataFrame(
  1020. {"participant_id": df.iloc[:, id_col], "age": df.iloc[:, age_col]}
  1021. )
  1022. for col in columns:
  1023. col_name = col.get("col_name")
  1024. col_id = col.get("col_id")
  1025. mapping = col.get("mapping")
  1026. single_value = col.get("single_value")
  1027. if col_name is None:
  1028. raise ValueError("Each column dictionary must contain a 'col_name'.")
  1029. if col_id is not None:
  1030. new_df[col_name] = df.iloc[:, col_id]
  1031. if mapping:
  1032. new_df[col_name] = new_df[col_name].map(mapping)
  1033. elif single_value is not None:
  1034. new_df[col_name] = single_value
  1035. else:
  1036. raise ValueError(
  1037. f"Column '{col_name}' must have either 'col_id' or 'single_value'."
  1038. )
  1039. # Special case handling
  1040. if col_name == "diagnosis":
  1041. new_df[col_name] = new_df[col_name].fillna("nan")
  1042. # Remove duplicate participants
  1043. new_df = new_df.drop_duplicates(subset="participant_id", keep="first")
  1044. save_dir = os.path.join(save_dir, "participants_bids.tsv")
  1045. # Save as BIDS-compatible TSV
  1046. new_df.to_csv(save_dir, sep="\t", index=False)
  1047. def set_path(project_dir):
  1048. """
  1049. Create and initialize directory structure for a given project.
  1050. This function generates a set of predefined directories for
  1051. feature extraction and normative modeling workflows within the
  1052. specified project directory. If any of these directories do not
  1053. exist, they will be created. The function returns the path to the
  1054. features log directory.
  1055. Parameters
  1056. ----------
  1057. project_dir : str
  1058. Path to the root project directory where the folder structure
  1059. will be created.
  1060. Returns
  1061. -------
  1062. str
  1063. Absolute path to the 'log' directory inside the 'Features'
  1064. folder.
  1065. Notes
  1066. -----
  1067. The function creates the following directory structure:
  1068. - ``Features/``
  1069. - ``log/`` (for saving logs of feature extraction)
  1070. - ``temp/`` (for temporarily storing extracted features)
  1071. - ``figures/`` (for saving generated figures)
  1072. - ``Normative modeling/``
  1073. - ``Runs/`` (for saving model run outputs)
  1074. - ``Figures/`` (for visual outputs related to modeling)
  1075. - ``Models summary/`` (for summaries of model results)
  1076. """
  1077. def make_folder(path):
  1078. if not os.path.isdir(path):
  1079. os.makedirs(path)
  1080. # Feature extraction
  1081. features_dir = os.path.join(project_dir, "Features")
  1082. features_log_path = os.path.join(features_dir, "log_slurm_jobs")
  1083. features_temp_path = os.path.join(features_dir, "temp")
  1084. exluded_participants_path = os.path.join(features_dir, "excluded_participants")
  1085. saved_outputs_path = os.path.join(features_dir, "Saved_outputs")
  1086. save_epochs_path = os.path.join(saved_outputs_path, "Epochs")
  1087. save_psds_path = os.path.join(saved_outputs_path, "PSDs")
  1088. save_preprocessed_data = os.path.join(saved_outputs_path, "Preprocessed_data")
  1089. save_coregistration_QC_path = os.path.join(saved_outputs_path, "coregistration_QC")
  1090. save_auto_reject_plot = os.path.join(saved_outputs_path, "auto_reject_plot")
  1091. save_covariance_figures_path = os.path.join(
  1092. saved_outputs_path, "Covariance_figures"
  1093. )
  1094. save_BEM_figures_path = os.path.join(saved_outputs_path, "BEM_figures")
  1095. save_transformation_path = os.path.join(
  1096. saved_outputs_path, "transformation_FIF_file"
  1097. )
  1098. configurations = os.path.join(features_dir, "Configurations")
  1099. mri_templates = os.path.join(features_dir, "MRI_templates")
  1100. save_grouping_effect = os.path.join(saved_outputs_path, "Grouping_effects")
  1101. make_folder(features_dir)
  1102. make_folder(features_log_path)
  1103. make_folder(features_temp_path)
  1104. make_folder(saved_outputs_path)
  1105. make_folder(save_epochs_path)
  1106. make_folder(save_psds_path)
  1107. make_folder(save_coregistration_QC_path)
  1108. make_folder(save_covariance_figures_path)
  1109. make_folder(save_BEM_figures_path)
  1110. make_folder(save_transformation_path)
  1111. make_folder(configurations)
  1112. make_folder(exluded_participants_path)
  1113. make_folder(mri_templates)
  1114. make_folder(save_grouping_effect)
  1115. make_folder(save_preprocessed_data)
  1116. make_folder(save_auto_reject_plot)
  1117. # Normative models
  1118. nm_dir = os.path.join(project_dir, "Normative_models")
  1119. make_folder(nm_dir)
  1120. return features_dir, features_log_path
  1121. def clean_nan_columns(df, nan_threshold):
  1122. """
  1123. Remove columns with more NaNs than `nan_threshold`,
  1124. otherwise impute NaNs with the column median.
  1125. Parameters:
  1126. df (pd.DataFrame): Input dataframe
  1127. nan_threshold (int): Max allowed NaNs per column
  1128. Returns:
  1129. pd.DataFrame: Cleaned dataframe
  1130. """
  1131. df = df.copy()
  1132. for col in df.columns:
  1133. nan_count = df[col].isna().sum()
  1134. if nan_count > nan_threshold:
  1135. # Drop column
  1136. df.drop(columns=col, inplace=True)
  1137. else:
  1138. # Impute with median (only for numeric columns)
  1139. if pd.api.types.is_numeric_dtype(df[col]):
  1140. median_value = df[col].median()
  1141. df[col] = df[col].fillna(median_value)
  1142. return df
  1143. def find_other_mri_session(
  1144. base_mri_path, missing_mri_subjects, str_mri_ending, which_session
  1145. ):
  1146. """
  1147. Find an alternative MRI file for each subject at a given session index.
  1148. Searches recursively under each subject's directory for files matching
  1149. the given filename ending, then selects the file at the requested
  1150. session index. Intended for locating a fallback MRI session (e.g. a
  1151. second scan) for subjects whose primary MRI is missing or unusable.
  1152. Parameters
  1153. ----------
  1154. base_mri_path : str or Path
  1155. Root directory containing per-subject MRI folders.
  1156. missing_mri_subjects : iterable of str
  1157. Subject IDs to search for.
  1158. str_mri_ending : str
  1159. Filename suffix/pattern used to match MRI files (e.g. "T1w.nii.gz").
  1160. which_session : int
  1161. 1-based index of the session to select from each subject's matched
  1162. file list.
  1163. Returns
  1164. -------
  1165. dict[str, str]
  1166. Mapping of subject ID to the matched MRI file path. Subjects with
  1167. fewer than `which_session` matching files are omitted.
  1168. """
  1169. new_paths = {}
  1170. for subject in missing_mri_subjects:
  1171. mri_paths = glob.glob(
  1172. f"{base_mri_path}/{subject}/**/*{str_mri_ending}", recursive=True
  1173. )
  1174. if len(mri_paths) > which_session - 1:
  1175. new_paths.update({subject: mri_paths[which_session - 1]})
  1176. return new_paths
  1177. def find_failed_meg_subjects(log_path):
  1178. """
  1179. Identify subjects whose MEG processing failed, based on log files.
  1180. Scans the given directory for error log files (files with "err" in the
  1181. name) and flags any subject whose log contains the string "error".
  1182. Parameters
  1183. ----------
  1184. log_path : str or Path
  1185. Directory containing per-subject log files.
  1186. Returns
  1187. -------
  1188. set of str
  1189. Unique subject IDs (parsed from the log filename, before the first
  1190. underscore) whose logs indicate a processing error.
  1191. """
  1192. missing_meg_subjects = []
  1193. paths = os.scandir(log_path)
  1194. paths = list(filter(lambda x: "err" in x.name, paths))
  1195. for path in paths:
  1196. with open(path, "r") as f:
  1197. content = f.read()
  1198. if "error" in content:
  1199. subject = os.path.basename(path).split(".")[0].split("_")[0]
  1200. missing_meg_subjects.append(subject)
  1201. return set(missing_meg_subjects)
  1202. def find_other_meg_session(
  1203. base_meg_path, missing_meg_subjects, str_meg_ending, task_name, which_session
  1204. ):
  1205. """
  1206. Find an alternative MEG recording for each subject at a given session index.
  1207. Searches recursively under each subject's directory for files matching
  1208. the given task name and filename ending, then selects the file at the
  1209. requested session index. Intended for locating a fallback MEG session
  1210. for subjects whose primary recording is missing or failed processing.
  1211. Parameters
  1212. ----------
  1213. base_meg_path : str or Path
  1214. Root directory containing per-subject MEG folders.
  1215. missing_meg_subjects : iterable of str
  1216. Subject IDs to search for.
  1217. str_meg_ending : str
  1218. Filename suffix/pattern used to match MEG files (e.g. "raw.fif").
  1219. task_name : str
  1220. Task identifier expected to appear in the filename (e.g. "rest").
  1221. which_session : int
  1222. 1-based index of the session to select from each subject's matched
  1223. file list.
  1224. Returns
  1225. -------
  1226. dict[str, str]
  1227. Mapping of subject ID to the matched MEG file path. Subjects with
  1228. fewer than `which_session` matching files are omitted.
  1229. """
  1230. new_paths = {}
  1231. for subject in missing_meg_subjects:
  1232. rs_record_paths = glob.glob(
  1233. f"{base_meg_path}/{subject}/**/*{task_name}*{str_meg_ending}",
  1234. recursive=True,
  1235. )
  1236. if len(rs_record_paths) > which_session - 1:
  1237. new_paths.update({subject: rs_record_paths[which_session - 1]})
  1238. return new_paths
  1239. def check_demographic_format(df):
  1240. """Raise ValueError if a demographic table is not in the standard format."""
  1241. DEMOGRAPHIC_REQUIRED_COLUMNS = ("participant_id", "age", "sex", "eyes")
  1242. DEMOGRAPHIC_ALLOWED_SEX = ("Male", "Female")
  1243. if "participant_id" in df.columns:
  1244. ids = df["participant_id"]
  1245. else:
  1246. ids = df.index.to_series()
  1247. for col in DEMOGRAPHIC_REQUIRED_COLUMNS:
  1248. if col != "participant_id" and col not in df.columns:
  1249. raise ValueError(f"The demographic file is missing the '{col}' column.")
  1250. if not all(isinstance(v, str) for v in ids.dropna()):
  1251. raise ValueError("All participant IDs must be strings.")
  1252. if not pd.api.types.is_numeric_dtype(df["age"]):
  1253. raise ValueError("The 'age' column must hold integers or floats.")
  1254. unexpected = set(df["sex"].dropna().unique()) - set(DEMOGRAPHIC_ALLOWED_SEX)
  1255. if unexpected:
  1256. raise ValueError(
  1257. f"The 'sex' column holds {sorted(unexpected)}; "
  1258. f"only {list(DEMOGRAPHIC_ALLOWED_SEX)} are allowed."
  1259. )

IO.py at commit a1fa698, under GPL-3.0 · at the source

Overview

Authors: Francesco Antonio Mallus1,2, Yanwu Yang1,3,4, Richard Dinga5, Mostafa Seyed Kia6,7,8, Tijl Grootswagers9, Abele Michela10, Tomas Ros10,11, Thomas Wolfers1,12,13
13 affiliations
  1. Department of Psychiatry and Psychotherapy, University Hospital Tübingen, Tübingen, Germany
  2. Graduate Training Center of Neuroscience, International Max-Planck Research School, Tübingen, Germany
  3. Max-Planck Institute for Intelligent Systems, Tübingen, Germany
  4. German Center for Mental Health (DZPG), partner site Tübingen, Germany
  5. Department of Neurology, Jena University Hospital, Jena, Germany
  6. Center of Cognitive Science and Artificial Intelligence, Tilburg University, Tilburg, The Netherlands
  7. Donders Institute for Cognition, Brain and Behavior, Radboud University, Nijmegen, The Netherlands
  8. Department of Psychiatry, UMC Utrecht Brain Center, University Medical Center, Utrecht, The Netherlands
  9. The MARCS Institute for Brain, Behaviour and Development, Western Sydney University, Sydney, Australia
  10. CIBM Center for Biomedical Imaging, University of Geneva, Geneva, Switzerland
  11. Department of Clinical Neuroscience, University of Geneva, Geneva, Switzerland
  12. German Center for Mental Health (DZPG), Jena-Magdeburg-Halle, Jena, Germany
  13. Institute of Psychology, Friedrich Schiller University Jena, Jena, Germany
Journal: Imaging neuroscience (Cambridge, Mass.), volume 4, article IMAG.a.1269
Dates: received 20 October 2025; accepted 8 May 2026; published online 15 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1162/imag.a.1269 · PMID 42312085 · PMCID PMC13271153 · OpenAlex W7161968072
Open access: diamond, a free copy (OpenAlex)
Status: code verified
Categories: EEG (modality), MEG (modality), computational (subfield)
Keywords: normative modeling, electroencephalography (EEG), magnetoencephalography (MEG), psychiatric diagnostics, psychiatric disorders, mental disorders, neurological diseases
Topic: Functional Brain Connectivity Studies (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: Deutsche Forschungsgemeinschaft (13851350); Bundesministerium für Forschung und Technologie (01EQ2403G); Carl-Zeiss-Stiftung (P2022-00-087)
Citations: not cited yet (Europe PMC); 308 references in the paper

Abstract

Normative modeling has become a cornerstone of computational neuroscience, offering a powerful framework for detecting individual deviations from typical brain function. This review traces its trajectory in electrophysiology of the brain, from early studies in the 1970s, through a period of relative neglect, to its recent revival driven by machine learning advances and the availability of large-scale datasets. We provide a structured overview of this evolution, showing the shift from small, site-specific age-based models to increasingly harmonized approaches that integrate diverse biological and methodological innovations. Key studies are compared with respect to cohort composition, modeling strategies, and validation procedures, situating each within the broader arc of methodological progress. Despite this momentum, significant challenges remain, such as a lack of standardized practices, limited comparability across studies, and the need for richer integration of complex neurophysiological signals. Looking ahead, we argue that the future of electrophysiological normative modeling lies in scaling and unifying efforts, through international collaborations, standardized pipelines, and the incorporation of novel features. By coupling machine learning with both cross-sectional and longitudinal designs, the field can progress from proof-of-concept demonstrations to precise, individualized brain mapping. Ultimately, such advances will provide the foundation for clinical applications, enabling cost-effective personalized treatment monitoring and a more refined understanding of individual differences in complex brain disorders.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.

Zenodo 15441320

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (7 files), pandas (7 files), MNE-Python (5 files), SciPy (3 files), Matplotlib (2 files), seaborn (2 files), specparam (formerly FOOOF) (2 files), ICLabel (1 file), Plotly (1 file), scikit-learn (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
21 files
At the source:

ml4pnp/meganorm

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: a1fa6984c442e310098f3b3b0a73c1d880bb149a, 16 September 2026
Languages: Python (21), Jupyter (1)
Size: 60 files, 22 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, license file, environment (Dockerfile, pyproject.toml, requirements.txt, setup.py, docs/requirements.txt), continuous integration, documentation, 1 notebook
Not found: CITATION.cff, tests
Tools: NumPy (12 files), pandas (10 files), Matplotlib (7 files), MNE-Python (6 files), SciPy (5 files), NiBabel (3 files), scikit-learn (3 files), specparam (formerly FOOOF) (3 files), FreeSurfer (2 files), seaborn (2 files), ArviZ (1 file), autoreject (1 file), ICLabel (1 file), Nilearn (1 file), Pingouin (1 file), Plotly (1 file), PyMC (1 file), statsmodels (1 file), xarray (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
24 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 41 scripts, each with its path and the digest of its content;
  • 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data and Code Availability

This review used no new or pre-existing data beyond the cited manuscripts and their associated dataset information, all referenced accordingly. No code was developed for or is fundamental to the reproducibility of this study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 8 authors, 7 keywords, 3 funders, 293 references.

Cite

This paper

Mallus, F. A., Yang, Y., Dinga, R., Kia, M. S., Grootswagers, T., Michela, A., Ros, T., & Wolfers, T. (2026). From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals. Imaging neuroscience (Cambridge, Mass.), 4, IMAG.a.1269. https://doi.org/10.1162/imag.a.1269

BibTeX

@article{mallus2026early,
author = {Mallus, Francesco Antonio and Yang, Yanwu and Dinga, Richard and Kia, Mostafa Seyed and Grootswagers, Tijl and Michela, Abele and Ros, Tomas and Wolfers, Thomas},
title = {{From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals}},
journal = {Imaging neuroscience (Cambridge, Mass.)},
year = {2026},
month = jun,
volume = {4},
pages = {IMAG.a.1269},
publisher = {MIT Press},
issn = {2837-6056},
doi = {10.1162/imag.a.1269},
url = {https://doi.org/10.1162/imag.a.1269},
pmid = {42312085},
pmcid = {PMC13271153}
}

RIS

TY - JOUR
AU - Mallus, Francesco Antonio
AU - Yang, Yanwu
AU - Dinga, Richard
AU - Kia, Mostafa Seyed
AU - Grootswagers, Tijl
AU - Michela, Abele
AU - Ros, Tomas
AU - Wolfers, Thomas
TI - From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals
T2 - Imaging neuroscience (Cambridge, Mass.)
J2 - Imaging Neurosci (Camb)
PY - 2026
DA - 2026/06/15
VL - 4
SP - IMAG.a.1269
SN - 2837-6056
PB - MIT Press
DO - 10.1162/imag.a.1269
UR - https://doi.org/10.1162/imag.a.1269
LA - en
ER -

CSL-JSON

{
"id": "10.1162/imag.a.1269",
"type": "article-journal",
"title": "From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals",
"container-title": "Imaging neuroscience (Cambridge, Mass.)",
"author": [
{
"family": "Mallus",
"given": "Francesco Antonio"
},
{
"family": "Yang",
"given": "Yanwu"
},
{
"family": "Dinga",
"given": "Richard"
},
{
"family": "Kia",
"given": "Mostafa Seyed"
},
{
"family": "Grootswagers",
"given": "Tijl"
},
{
"family": "Michela",
"given": "Abele"
},
{
"family": "Ros",
"given": "Tomas"
},
{
"family": "Wolfers",
"given": "Thomas"
}
],
"container-title-short": "Imaging Neurosci (Camb)",
"volume": "4",
"page": "IMAG.a.1269",
"DOI": "10.1162/imag.a.1269",
"PMID": "42312085",
"PMCID": "PMC13271153",
"ISSN": "2837-6056",
"publisher": "MIT Press",
"URL": "https://doi.org/10.1162/imag.a.1269",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
15
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.64898/2026.03.31.26349848 [code]
Mapping Individual Neuroanatomical Alterations to Schizophrenia Psychopathology with Normative Modeling
Journal: medRxiv (preprint)
In common: FreeSurfer, seaborn, pandas, 2 other tools, computational, 15 references
[2] doi:10.1038/s41597-026-07350-9 [code]
An open multi-center MEG-EEG dataset for studying conscious visual perception.
Journal: Scientific data
In common: autoreject, xarray, Pingouin, 11 other tools, MEG, EEG, 5 references
[3] doi:10.7554/elife.100605 [code]
Age-related changes in ‘cortical’ 1/f dynamics are linked to cardiac activity
Journal: n/a
In common: ArviZ, PyMC, specparam (formerly FOOOF), 10 other tools, MEG, 3 references
[4] doi:10.64898/2026.08.13.26360304 [code]
Lifespan brain structural variation reveals shared organization across mental health conditions
Journal: medRxiv (preprint)
In common: Nilearn, statsmodels, NiBabel, 6 other tools, 11 references
[5] doi:10.1038/s41380-026-03691-4 [code]
Breaking the norm: population-scale deviations of brain structure in depression and anxiety.
Journal: Molecular psychiatry
In common: Nilearn, statsmodels, NiBabel, 6 other tools, 9 references
[6] doi:10.1038/s41467-026-72940-5 [code]
Cerebellar growth is associated with domain-specific cerebral maturation and socio-linguistic behavior.
Journal: Nature communications
In common: ArviZ, PyMC, xarray, 8 other tools, 5 references
[7] doi:10.1038/s41467-026-72875-x [code]
Lifespan normative modeling of brain microstructure.
Journal: Nature communications
In common: seaborn, scikit-learn, pandas, 3 other tools, 12 references
[8] doi:10.1038/s41597-026-07377-y [code]
An open-access multi-site fMRI dataset for investigating conscious visual perception.
Journal: Scientific data
In common: autoreject, xarray, Pingouin, 11 other tools, 1 reference
[9] doi:10.1093/nc/niag029 [code]
A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.
Journal: Neuroscience of consciousness
In common: autoreject, xarray, Pingouin, 11 other tools, MEG
[10] doi:10.1162/imag.a.1245 [code]
Towards precision EEG connectomics: Evaluating the benefits of dense sampling.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: autoreject, ICLabel, Pingouin, 10 other tools, EEG, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.