From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.
The 8 matches
- [1] § Early Studies of Normative Modeling ↔ meganorm/utils/IO.py, lines 34–496 · score 0.69 · EEG source, exponent, broadband, peak, aperiodic, alpha
- [2] § Contemporary Studies of Normative Modeling ↔ meganorm/src/normative_modeling.py, lines 498–603 · score 0.68 · hierarchical Bayesian regression, fitted normative, predict, batch, mapping, age
- [3] § Early Studies of Normative Modeling ↔ meganorm/utils/IO.py, lines 34–496 · score 0.66 · band power, spectral features, cross, head, norms, channels
- [4] § Early Studies of Normative Modeling ↔ notebooks/feature_extraction_tutorial.ipynb, lines 416–482 · score 0.56 · peak frequency, oscillatory, exponent, broadband, aperiodic, alpha
- [5] § What is Normative Modeling? ↔ meganorm/utils/nm.py, lines 206–343 · score 0.55 · evaluating model calibration, Ideally, predictive, error, covariates, scores
- [6] § What is Normative Modeling? ↔ meganorm/utils/nm.py, lines 204–341 · score 0.55 · evaluating model calibration, Ideally, predictive, error, covariates, scores
- [7] § Limitations of Contemporary Approaches ↔ meganorm/src/normative_modeling.py, lines 498–603 · score 0.51 · hierarchical Bayesian regression, neuroimaging, variables, covariates, brain
- [8] § Early Studies of Normative Modeling ↔ meganorm/src/mainParallel.py, lines 241–300 · score 0.51 · band power, spectral features, head, computational, demographic, channels
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 1,470 lines · 55 KB · GPL-3.0 · 2 matches
- import pickle
- import pandas as pd
- import json
- import os
- import re
- import glob
- from pathlib import Path
- import mne
- import numpy as np
- from typing import Literal, Dict, Tuple, List, Optional
- from typing import Optional
- import warnings
- from typing import Union
- from typing import ClassVar
- from pydantic import (
- BaseModel,
- Field,
- PositiveInt,
- confloat,
- conint,
- conlist,
- field_validator,
- NegativeInt,
- model_validator,
- )
- from mne.io.constants import FIFF
- class BandRatio(BaseModel):
- numerator: Literal["Delta", "Theta", "Alpha", "Beta", "Gamma"]
- denominator: Literal["Delta", "Theta", "Alpha", "Beta", "Gamma"]
- class Config(BaseModel):
- """
- Configuration for preprocessing, artifact correction, source localization,
- and spectral feature extraction of neurophysiological signals (MEG/EEG/OPM).
- Parameters
- ----------
- which_meg_session : int, default=0
- Index of the MEG session to process.
- which_layout : {"all", "lobe", None}, default="all"
- Sensor layout grouping used for reporting/analysis.
- which_sensor : {"mag", "grad", "meg", "eeg", "opm"}, default="meg"
- Sensor type to process.
- drop_noisy_flat_channel : bool, default=True
- Drop channels flagged as noisy or flat before further processing.
- Filtering & Resampling
- -----------------------
- cutoffFreqLow, cutoffFreqHigh : float, PositiveInt, default=1.0, 80
- Bandpass filter cutoff frequencies (Hz).
- resampling_rate : PositiveInt, default=1000
- Target sampling rate (Hz).
- digital_filter, notch_filter : bool, default=True
- Apply bandpass / line-noise notch filtering.
- apply_oversampled_temporal_projection : bool, default=True
- Apply oversampled temporal projection (OTP) denoising.
- apply_Head_movement_correction : bool, default=True
- Correct for head movement during recording.
- Head_movement_limit_from_mean : float, default=0.0015
- Maximum allowed deviation from mean head position (m).
- apply_chpi_filter : bool, default=False
- Filter out cHPI coil signals.
- ICA
- ---
- apply_ica : bool, default=True
- Apply ICA-based artifact correction.
- apply_ica_elbow_detection : bool, default=False
- Automatically select the number of ICA components via elbow detection.
- ica_n_component : PositiveInt or None, default=None
- Number of ICA components to compute.
- ica_max_iter : PositiveInt, default=800
- Maximum ICA iterations.
- ica_method : {"fastica", "infomax", "picard"}, default="fastica"
- ICA algorithm.
- ica_if_reject_by_annotation : bool, default=True
- Exclude annotated bad segments when fitting ICA.
- auto_ica_corr_thr : float, default=0.5
- Correlation threshold (0-1) for automatic ICA component rejection.
- GEDAI Artifact Removal
- -----------------------
- apply_gedai : bool, default=True
- Apply GEDAI-based artifact removal.
- gedai_method : {"both", "spectral", "broadband"}, default="both"
- GEDAI denoising strategy.
- sensai_method : {"optimize", "gridsearch"}, default="optimize"
- Parameter search strategy for SensAI.
- gedai_duration, gedai_overlap : float or int, default=12, 0.5
- Window duration (s) and overlap fraction for GEDAI.
- gedai_preliminary_broadband_noise_multiplier : float, default=6.0
- Noise multiplier for preliminary broadband detection.
- gedai_noise_multiplier : float, default=3.0
- Noise multiplier used in GEDAI thresholding.
- gedai_wavelet_type : str, default="haar"
- Wavelet family used for spectral GEDAI.
- gedai_wavelet_level : "auto", PositiveInt, or 0, default="auto"
- Wavelet decomposition level.
- gedai_wavelet_low_cutoff : float or None, default=None
- Low-frequency cutoff for wavelet-based denoising.
- gedai_epoch_size_in_cycles : PositiveInt, default=12
- Epoch size expressed in number of cycles.
- gedai_highpass_cutoff : float, default=0.1
- High-pass cutoff applied before GEDAI (Hz).
- Muscle Artifact Detection
- ---------------------------
- muscle_activity_thr : int, default=4
- Detection threshold for muscle artifacts.
- muscle_activity_min_length_good : float, default=0.1
- Minimum length of a clean segment retained after removal (s).
- muscle_activity_filter_freq : tuple[int, int], default=(110, 140)
- Frequency band used for muscle artifact detection (Hz).
- Environmental Noise Correction
- --------------------------------
- apply_environmental_noise_correction : bool, default=True
- Apply environmental/reference-based noise correction.
- ctf_gradient_comp_level : PositiveInt, default=3
- CTF gradient compensation level.
- apply_environmental_noise_ssp_with_eroom : bool, default=False
- Use empty-room SSP projectors for noise correction.
- apply_environmental_noise_ica_with_ref_meg : bool, default=True
- Use reference-MEG-guided ICA for environmental noise removal.
- environmental_noise_ica_with_ref_meg_thr : float, default=2.5
- Threshold for ref-MEG-guided ICA component rejection.
- environmental_noise_ica_with_ref_meg_method : {"together", "separate"}, default="separate"
- Whether to process reference channels jointly or separately.
- environmental_noise_ica_with_ref_meg_measure : {"zscore", "correlation"}, default="zscore"
- Metric used to score ICA components against reference channels.
- EEG Reference & Bad-Segment Rejection
- ----------------------------------------
- rereference_method : {"average", "REST", None}, default="average"
- EEG re-referencing scheme.
- bad_segment_removal_method : {"autoreject", "fixed_thr", None}, default="autoreject"
- Method for rejecting bad data segments.
- mag_var_threshold, grad_var_threshold, eeg_var_threshold : float
- Variance-based rejection thresholds per channel type.
- mag_flat_threshold, grad_flat_threshold, eeg_flat_threshold : float
- Flatline-detection thresholds per channel type.
- zscore_std_thresh : PositiveInt, default=15
- Z-score threshold for outlier rejection.
- autoreject_n_interpolates : list[int], default=[1, 4, 8, 16, 32]
- Candidate interpolation counts for Autoreject.
- autoreject_consensus_percs : list[float]
- Candidate consensus percentages for Autoreject (11 values, 0-1).
- autoreject_cv : int or "auto", default="auto"
- Cross-validation folds for Autoreject.
- autoreject_thresh_method : {"bayesian_optimization", "random_search"}, default="bayesian_optimization"
- Threshold search strategy for Autoreject.
- Segmentation
- ------------
- segments_tmin, segments_tmax : PositiveInt, NegativeInt, default=20, -20
- Segment start/end times relative to event (s).
- segments_length, segments_overlap : int, default=10, 2
- Segment length and overlap (s).
- Source Localization
- --------------------
- apply_source_localization : bool, default=False
- Whether to perform source localization.
- apply_empty_room_recording : bool, default=True
- Use empty-room recordings for noise covariance estimation.
- apply_mri_QC : bool, default=False
- Run quality control on MRI/FreeSurfer output.
- apply_mri_template : bool, default=False
- Use a template MRI instead of subject-specific anatomy.
- freesurfer_template_path, freesurfer_home, freesurfer_license : str or None
- Paths to FreeSurfer template derivatives, installation, and license.
- force_new_watershed_bem : bool, default=False
- Recompute the watershed BEM surfaces.
- gcaatlas : bool, default=True
- Use the GCA atlas for subcortical segmentation.
- SL_source_space : {"surface", "volumetric"}, default="volumetric"
- Type of source space.
- SL_conductivity : tuple[float, ...], default=(0.3,)
- Head-model layer conductivities (three values required for EEG).
- SL_inverse_operator : {"lcmv"}, default="lcmv"
- Inverse operator method.
- source_space_spacing : {"ico3"..."ico6", "oct5", "oct6"}, default="ico4"
- Source space resolution.
- source_space_spacing_number : {3, 4, 5, 6}, default=4
- Numeric resolution; must match `source_space_spacing`.
- coregisteration_final_n_iterations : int, default=20
- Iterations for the final coregistration refinement.
- coregisteration_final_nasion_weight : float, default=10.0
- Weight applied to the nasion fiducial during coregistration.
- covariance_method : str, default="empirical"
- Method for noise covariance estimation.
- Beamformer & Parcellation
- ----------------------------
- beamformer_pick_ori : {None, "normal", "max-power", "vector"}, default="max-power"
- Source orientation constraint.
- beamformer_weight_norm : {None, "unit-noise-gain", "nai", "unit-noise-gain-invariant"}, default="unit-noise-gain"
- Beamformer weight normalization.
- beamforme_depth : float, default=0.08
- Depth-weighting factor correcting for center-of-head bias.
- inverse_regularization_value : float, default=0.05
- Regularization applied to the data covariance matrix.
- apply_morphing : bool, default=False
- Morph source estimates to a common template brain.
- parcellation_parc : {None, "aparc.a2009s", "parac"}, default="aparc.a2009s"
- Predefined cortical parcellation.
- parcellation_annot_fname : Path or None
- Custom parcellation (.annot) file, used if `parcellation_parc` is None.
- PSD & Spectral Parametrization
- ---------------------------------
- psd_method : {"multitaper", "welch"}, default="welch"
- PSD estimation method.
- psd_n_overlap, psd_n_fft, psd_n_per_seg : PositiveInt, default=1, 2, 2
- Welch/multitaper PSD parameters.
- parametrization_method : {"fooof", "irasa"}, default="irasa"
- Method for separating aperiodic and periodic spectral components.
- irasa_hset : tuple[float, float, float], default=(1.05, 2.0, 0.05)
- Resampling factor range/step for IRASA.
- fooof_freq_range_low, fooof_freq_range_high : PositiveInt, default=3, 40
- Frequency range for FOOOF fitting (Hz).
- aperiodic_mode : {"knee", "fixed"}, default="knee"
- Aperiodic component model.
- fooof_peak_width_limits : list[float], default=[1.0, 12.0]
- Allowed peak width range (Hz).
- fooof_min_peak_height : int, default=0
- Minimum peak height for detection.
- fooof_peak_threshold : PositiveInt, default=2
- Peak detection threshold (in SD of the flattened spectrum).
- fooof_res_save_path : str or None
- Path to save FOOOF results.
- save_source_localized_epochs, save_psds : bool, default=False
- Persist intermediate source-localized epochs / PSDs to disk.
- Feature Extraction
- -------------------
- freq_bands : dict[str, tuple[int, int]]
- Canonical frequency band definitions (Theta, Alpha, Beta, Gamma).
- individualized_band_ranges : dict[str, tuple[int, int]]
- Per-band offsets (Hz) used to individualize canonical bands.
- power_band_ratios_list : list[BandRatio]
- Band-power ratios to compute (e.g. Theta/Beta).
- min_r_squared : float, default=0.9
- Minimum R² required to accept a spectral model fit.
- feature_categories : dict[str, bool]
- Flags selecting which feature families to extract (offset, exponent,
- peak parameters, canonical/individualized band power, band ratios,
- hemispheric asymmetry).
- Miscellaneous
- -------------
- random_state : int, default=42
- Random seed for reproducibility.
- Methods
- -------
- save(save_path, overwrite=False)
- Serialize the configuration to a JSON file.
- load(path)
- Load a configuration from a JSON file.
- Notes
- -----
- Model validators enforce cross-field consistency, e.g.: a three-layer
- `SL_conductivity` is required for EEG source localization; `beamformer_pick_ori
- == "vector"` requires `beamformer_weight_norm == "unit-noise-gain-invariant"`;
- `source_space_spacing` must match `source_space_spacing_number`; GEDAI
- parameters must be consistent with the chosen `gedai_method`; and MRI
- template use is mutually exclusive with MRI QC.
- """
- model_config = {"extra": "forbid"}
- which_meg_session: int = 0 # the first session
- which_layout: Literal["all", "lobe", None] = "all"
- which_sensor: Literal["mag", "grad", "meg", "eeg", "opm"] = "meg"
- drop_noisy_flat_channel: bool = True
- remove_nonfinite_segment_threshold: PositiveInt = 5
- # ICA
- apply_ica_elbow_detection: bool = False
- ica_n_component: Optional[PositiveInt] = None
- ica_max_iter: PositiveInt = 800
- ica_method: Literal["fastica", "infomax", "picard"] = "fastica"
- cutoffFreqLow: float = 1.0
- cutoffFreqHigh: PositiveInt = 80
- resampling_rate: PositiveInt = 1000
- digital_filter: bool = True
- notch_filter: bool = True
- apply_oversampled_temporal_projection: bool = True
- apply_Head_movement_correction: bool = True
- Head_movement_limit_from_mean: float = 0.0015
- apply_chpi_filter: bool = False
- # gedai settings
- apply_gedai: bool = True
- gedai_method: Literal["both", "spectral", "broadband"] = "both"
- sensai_method: Literal["optimize", "gridsearch"] = "optimize"
- gedai_duration: Union[float, int] = 12
- gedai_overlap: Union[float, int] = 0.5
- gedai_preliminary_broadband_noise_multiplier: float = 6.0
- gedai_noise_multiplier: float = 3.0
- gedai_wavelet_type: str = "haar"
- gedai_wavelet_level: Union[Literal["auto"], PositiveInt, Literal[0]] = "auto"
- gedai_wavelet_low_cutoff: Union[None, float] = None
- gedai_epoch_size_in_cycles: PositiveInt = 12
- gedai_highpass_cutoff: float = 0.1
- muscle_activity_thr: int = 4
- muscle_activity_min_length_good: float = 0.1
- muscle_activity_filter_freq: Tuple[int, int] = (110, 140)
- apply_environmental_noise_correction: bool = True
- same_environmental_noise_removal: bool = False
- ctf_gradient_comp_level: PositiveInt = 3
- apply_environmental_noise_ssp_with_eroom: bool = False
- apply_environmental_noise_ica_with_ref_meg: bool = True
- environmental_noise_ica_with_ref_meg_thr: float = 2.0
- ica_if_reject_by_annotation: bool = True
- environmental_noise_ica_with_ref_meg_method: Literal["together", "separate"] = (
- "separate"
- )
- environmental_noise_ica_with_ref_meg_measure: Literal["zscore", "correlation"] = (
- "zscore"
- )
- apply_ica: bool = True
- auto_ica_corr_thr: confloat(ge=0, le=1) = 0.5
- save_segmented_data: bool = False
- rereference_method: Literal["average", "REST", None] = "average"
- bad_segment_removal_method: Literal["autoreject", "fixed_thr", None] = "autoreject"
- mag_var_threshold: float = 5000e-15
- grad_var_threshold: float = 5000e-13
- eeg_var_threshold: float = 40e-6
- mag_flat_threshold: float = 10e-15
- grad_flat_threshold: float = 10e-13
- eeg_flat_threshold: float = 40e-6
- zscore_std_thresh: PositiveInt = 15
- segments_tmin: PositiveInt = 20
- segments_tmax: NegativeInt = -20
- segments_length: PositiveInt = 10
- segments_overlap: int = 2
- save_preprocessed_data: bool = True
- # autoreject
- autoreject_n_interpolates: List[int] = [1, 4, 8, 16, 32]
- autoreject_consensus_percs: List[float] = list(np.linspace(0, 1.0, 11))
- autoreject_cv: Union[int, Literal["auto"]] = "auto"
- autoreject_thresh_method: Literal["bayesian_optimization", "random_search"] = (
- "bayesian_optimization"
- )
- # Source localization
- apply_source_localization: bool = False
- apply_empty_room_recording: bool = True
- apply_mri_QC: bool = False
- take_screenshot_of_coregisteration: bool = True
- apply_mri_template: bool = False
- freesurfer_template_path: Optional[str] = None
- freesurfer_home: Optional[str] = None
- freesurfer_license: Optional[str] = None
- coregisteration_scale_mode: Literal["uniform", "3-axis", None] = None
- force_new_watershed_bem: bool = False
- gcaatlas: bool = True
- SL_source_space: Literal["surface", "volumetric"] = "volumetric"
- SL_conductivity: Tuple[float, ...] = (0.3,)
- SL_inverse_operator: Literal["lcmv"] = "lcmv"
- bem_preflood_parameter_space: List[int] = (10, 15, 20, 30, 35)
- bem_preflood: int = 25
- bem_gcaatlas: bool = True
- bem_max_erosion_pct: float = 15.0
- bem_suspect_erosion_pct: float = 0.2
- bem_max_fine_segmentation_iteration: int = 100
- bem_plot_orientations: Literal["coronal", "axial", "sagittal", None] = "coronal"
- # the spacing to use for source space specificatin
- source_space_spacing: Literal[
- "ico3", "ico4", "ico5", "ico6", "oct5", "oct6", "all"
- ] = "ico4"
- source_space_spacing_number: Literal[3, 4, 5, 6, None] = 4
- source_space_add_dist: Literal["patch", True, False] = "patch"
- save_transformation_FIF_file: bool = False
- coregisteration_final_n_iterations: int = 20
- coregisteration_final_nasion_weight: float = 10.0
- covariance_method: str = "empirical"
- # Determines whether to keep vectors for all source orientations
- # or to select a single fixed orientation, depending on the chosen algorithm.
- beamformer_pick_ori: Literal[None, "normal", "max-power", "vector"] = "max-power"
- beamformer_weight_norm: Literal[
- None, "unit-noise-gain", "nai", "unit-noise-gain-invariant"
- ] = "unit-noise-gain-invariant"
- # This parameter scales the activation to correct for head-center bias.
- beamforme_depth: confloat(ge=0, le=1) = 0.8
- # this is used for regularaizing the data covariance (shifting the matrix)
- inverse_regularization_value: confloat(ge=0, le=1) = 0.05
- apply_morphing: bool = False
- # the pacellation to use
- parcellation_parc: Literal[None, "aparc.a2009s", "parac"] = "aparc.a2009s"
- parcellation_mode: Literal["mean_flip", "mean", "auto", "pca_flip", "max"] = "auto"
- # A custom parcellation file
- parcellation_annot_fname: Optional[Path] = None
- # PSD
- psd_method: Literal["multitaper", "welch"] = "welch"
- psd_n_overlap: PositiveInt = 1
- psd_n_fft: PositiveInt = 2
- psd_n_per_seg: PositiveInt = 2
- parametrization_method: Literal["fooof", "irasa"] = "irasa"
- # PYRASA
- irasa_hset: Tuple[float, float, float] = (1.05, 2.0, 0.05)
- # FOOOF analysis
- fooof_freq_range_low: PositiveInt = 3
- fooof_freq_range_high: PositiveInt = 40
- aperiodic_mode: Literal["knee", "fixed"] = "knee"
- fooof_peak_width_limits: List[float] = [1.0, 12.0]
- fooof_min_peak_height: int = 0
- fooof_peak_threshold: PositiveInt = 2
- save_source_localized_epochs: bool = False
- save_psds: bool = False
- lowest_num_of_epochs: Optional[int] = None
- # Feature extraction
- freq_bands: Dict[str, Tuple[int, int]] = {
- "Theta": (3, 8),
- "Alpha": (8, 13),
- "Beta": (13, 30),
- "Gamma": (30, 40),
- }
- individualized_band_ranges: Dict[str, Tuple[int, int]] = {
- "Theta": (-2, 3),
- "Alpha": (-2, 3),
- "Beta": (-8, 9),
- "Gamma": (-5, 5),
- }
- power_band_ratios_list: List[BandRatio] = [
- BandRatio(numerator="Theta", denominator="Beta"),
- BandRatio(numerator="Theta", denominator="Alpha"),
- BandRatio(numerator="Alpha", denominator="Beta"),
- BandRatio(numerator="Delta", denominator="Beta"),
- BandRatio(numerator="Delta", denominator="Alpha"),
- BandRatio(numerator="Delta", denominator="Theta"),
- ]
- min_r_squared: confloat(ge=0, le=1) = 0.9
- feature_categories: Dict[str, bool] = {
- "Offset": True,
- "Exponent": True,
- "Peak_Center": False,
- "Peak_Power": False,
- "Peak_Width": False,
- "Adjusted_Canonical_Relative_Power": True,
- "Adjusted_Canonical_Absolute_Power": True,
- "Adjusted_Individualized_Relative_Power": False,
- "Adjusted_Individualized_Absolute_Power": False,
- "OriginalPSD_Canonical_Relative_Power": True,
- "OriginalPSD_Canonical_Absolute_Power": True,
- "OriginalPSD_Individualized_Relative_Power": False,
- "OriginalPSD_Individualized_Absolute_Power": False,
- "Adjusted_Band_Ratio": True,
- "OriginalPSD_Band_Ratio": True,
- "Hemispheric_Asymmetry_index": True,
- "Knee_Frequency": False,
- }
- fooof_res_save_path: Optional[str] = None
- random_state: int = 42
- @field_validator("muscle_activity_thr")
- def muscle_activity_thr_fv(cls, v):
- if v < 3:
- warnings.warn(
- "Select a higher threshold for muscle activity artifacts. Low values "
- "remove clean data."
- )
- return v
- @field_validator("muscle_activity_filter_freq")
- def muscle_activity_filter_freq_fv(cls, v):
- if v[0] < 100:
- warnings.warn(
- f"Muscle activity artifact affects higher frequencies than {v[0]} Hz."
- )
- return v
- @model_validator(mode="after")
- def SL_conductivity_mv(self):
- if (
- len(self.SL_conductivity) == 1
- and self.which_sensor == "eeg"
- and self.apply_source_localization
- ):
- raise ValueError(
- "In the case of EEG, you must have a three layers conductivity model due to volume conduction."
- )
- return self
- @model_validator(mode="after")
- def source_space_res(self):
- if self.source_space_spacing == "all":
- if self.source_space_spacing_number is not None:
- raise ValueError(
- "source_space_spacing_number must be None when source_space_spacing='all'"
- )
- return self
- if int(self.source_space_spacing[-1]) != self.source_space_spacing_number:
- raise ValueError(
- "The source_space_spacing and source_space_spacing_number should match"
- )
- return self
- @model_validator(mode="after")
- def beamformer_arg_check(self):
- if (
- self.beamformer_pick_ori == "vector"
- and self.beamformer_weight_norm != "unit-noise-gain-invariant"
- ):
- error_msg = (
- "If you wish to compute a vector beamformer, it is necessary to use"
- " unit-noise-gain-invariant for weight_norm argument. This is for addressing"
- " the center of head bias where deeper sources can have larger scale than"
- " superfacial sources."
- )
- raise ValueError(error_msg)
- return self
- @model_validator(mode="after")
- def center_head_bias_scale_check(self):
- if not self.beamforme_depth and self.SL_source_space == "volumetric":
- error_msg = (
- "If you want to use volumetric source space (interested in deeper sources),"
- " please define beamforme_depth as positive float number, i.e., 0.8. This is used to address"
- " the center of head bias."
- )
- raise ValueError(error_msg)
- return self
- @model_validator(mode="after")
- def ica_e_noise_removal(self):
- if (
- self.environmental_noise_ica_with_ref_meg_measure == "correlation"
- and not 0 < self.environmental_noise_ica_with_ref_meg_thr < 1
- ):
- error_mg = (
- "If the threshold method for removing environmental noise using "
- "ICA and ref-MEG is correlation, the corresponding measure must be between 0 and 1."
- )
- raise ValueError(error_mg)
- return self
- @model_validator(mode="after")
- def pacellation_checker(self):
- if not self.parcellation_parc and not self.parcellation_annot_fname:
- raise ValueError(
- "Parcellation should be passed. Otherwise pass a custom parcellation file (.annot)"
- )
- return self
- @model_validator(mode="after")
- def mri_template_check(self):
- if self.apply_mri_template and not self.freesurfer_template_path:
- err_msg = (
- "In order to apply template MRI for source localization, a path to already freesurfer-preprocessed"
- "derivatives is necessary"
- )
- raise ValueError(err_msg)
- return self
- @model_validator(mode="after")
- def either_template_mri_or_mri_qc(self):
- if self.apply_mri_QC and self.apply_mri_template:
- err_msg = (
- "You can not apply MRI QC on already preprocessed freesurfer template"
- )
- raise ValueError(err_msg)
- return self
- @model_validator(mode="after")
- def gedai_params_check(self):
- method = self.gedai_method
- wavelet_level = self.gedai_wavelet_level
- duration = self.gedai_duration
- broadband_multiplier = self.gedai_preliminary_broadband_noise_multiplier
- if method == "broadband" and wavelet_level != 0:
- raise ValueError("broadband method requires wavelet_level=0")
- if method == "broadband" and not duration:
- raise ValueError("broadband method requires gedai_duration")
- if method == "spectral" and wavelet_level == 0:
- raise ValueError("spectral method requires wavelet_level > 0")
- if method == "both" and not broadband_multiplier:
- raise ValueError(
- "both method requires gedai_preliminary_broadband_noise_multiplier"
- )
- return self
- def save(self, save_path: str, overwrite=False):
- "save the configurations to a JSON file"
- if os.path.exists(save_path) and overwrite == False:
- err_msg = f"A configuration file already exists in this directory: {save_path}. Set the overwrite to True."
- raise FileExistsError(err_msg)
- with open(save_path, "w") as file:
- file.write(self.model_dump_json(indent=4))
- @classmethod
- def load(cls, path: str):
- # Load configuration from a JSON file
- with open(path, "r") as file:
- cfg = json.load(file)
- return cls(**cfg)
- def _resolve_artemis_pos_file(path, pos_file, logger):
- """Return an explicit pos/dig file, or find one near the recording."""
- if pos_file:
- return pos_file
- # .pos usually lives in a sibling Headshape/ folder, not next to the run,
- # so widen the search one directory up before giving up.
- search_roots = [Path(path).parent, Path(path).parents[1]]
- for root in search_roots:
- matches = sorted(glob.glob(f"{root}/**/*.pos", recursive=True))
- if matches:
- if len(matches) > 1:
- logger.warning(
- f"{len(matches)} .pos files found under {root}; using {matches[0]}."
- )
- return matches[0]
- err_msg = (
- f"No .pos file found near the Artemis recording {path} "
- f"(searched {', '.join(str(r) for r in search_roots)})."
- )
- logger.error(err_msg)
- raise FileNotFoundError(err_msg)
- def _add_artemis_headshape(data, path, pos_file, logger):
- """Attach head-shape points to an already-converted Artemis FIF."""
- existing = data.info["dig"] or []
- n_extra = sum(1 for d in existing if d["kind"] == FIFF.FIFFV_POINT_EXTRA)
- if n_extra:
- logger.info(
- f"Artemis FIF already carries {n_extra} head-shape points; "
- "not adding any from a .pos/.fif file."
- )
- return data
- resolved = _resolve_artemis_pos_file(path, pos_file, logger)
- ext = str(resolved).split(".")[-1].lower()
- if ext == "pos":
- digs = mne.io.artemis123.utils._read_pos(fname=resolved)
- with data.info._unlock():
- data.info["dig"] = existing + digs
- data.info["dig"] = mne._fiff._digitization._format_dig_points(
- data.info["dig"]
- )
- logger.info(f"Head shape .pos file {resolved} was added to the Artemis FIF.")
- elif ext == "fif":
- dig = mne.channels.read_dig_fif(resolved)
- data.set_montage(dig, on_missing="warn")
- logger.info(f"Head shape .fif file {resolved} was added to the Artemis FIF.")
- else:
- logger.warning(
- f"pos_file '{resolved}' has unrecognized extension '.{ext}'; "
- "no head-shape/dig info was added to the recording."
- )
- return data
- def select_session_path(pattern, index=0):
- """
- Resolve a '*'-separated glob pattern to a single path.
- Returns None when `pattern` is falsy, so callers can pass optional
- CLI arguments straight through.
- """
- if not pattern:
- return None
- candidates = [p for p in pattern.split("*") if p]
- return candidates[index]
- def infer_device(path, device_type, which_sensor, logger):
- """Determine the acquisition device from an explicit override or the path."""
- MEG_DEVICE_BY_EXTENSION = {"ds": "CTF", "fif": "MEGIN", "bin": "ARTEMIS123"}
- if which_sensor == "eeg":
- # TODO: was originally path[0]. Check if this correction is correct.
- return path.split(".")[-1].upper()
- if device_type:
- return device_type.upper()
- if "4D" in path:
- return "BTI"
- extension = path.split(".")[-1]
- if extension in MEG_DEVICE_BY_EXTENSION:
- return MEG_DEVICE_BY_EXTENSION[extension]
- err_msg = "The provided MEG recording is not supported yet."
- logger.error(err_msg)
- raise ValueError(err_msg)
- def load_recording(
- device, path, empty_room_recording_path, configs, logger, pos_file=None
- ):
- """Load data"""
- if device == "CTF":
- data = mne.io.read_raw_ctf(path, preload=True)
- if pos_file:
- ext = pos_file.split(".")[-1]
- if ext == "pos":
- digs = mne.io.artemis123.utils._read_pos(fname=pos_file)
- with data.info._unlock():
- data.info["dig"] += digs
- data.info["dig"] = mne._fiff._digitization._format_dig_points(
- data.info["dig"]
- )
- logger.info(
- "Head shape .pos file was found and added to the CTF recording"
- )
- elif ext == "fif":
- dig = mne.channels.read_dig_fif(pos_file)
- data.set_montage(dig, on_missing="warn")
- logger.info(
- "Head shape .fif file was found and added to the CTF recording"
- )
- else:
- logger.warning(
- f"pos_file '{pos_file}' has unrecognized extension '.{ext}'; "
- "no head-shape/dig info was added to the recording."
- )
- if empty_room_recording_path and (
- configs.apply_source_localization
- or configs.apply_environmental_noise_ssp_with_eroom
- ):
- empty_room_recording = mne.io.read_raw_ctf(
- empty_room_recording_path, preload=True
- )
- logger.info("Empty room recording was found")
- elif not empty_room_recording_path and (
- configs.apply_source_localization
- or configs.apply_environmental_noise_ssp_with_eroom
- ):
- empty_room_recording = None
- logger.warning("No empty room recording was found")
- else:
- empty_room_recording = None
- elif device == "BTI":
- temp_hs_file_path = os.path.join(path, "hs_file")
- hs_file = temp_hs_file_path if os.path.exists(temp_hs_file_path) else None
- if not hs_file:
- convert = False
- else:
- convert = True
- data = mne.io.read_raw_bti(
- pdf_fname=os.path.join(path, "c,rfDC"),
- config_fname=os.path.join(path, "config"),
- head_shape_fname=hs_file,
- preload=True,
- convert=convert,
- )
- if empty_room_recording_path and configs.apply_source_localization:
- empty_room_recording = mne.io.read_raw_bti(
- pdf_fname=os.path.join(empty_room_recording_path, "c,rfDC"),
- config_fname=os.path.join(empty_room_recording_path, "config"),
- head_shape_fname=None,
- convert=convert,
- preload=True,
- )
- logger.info("Empty room recording was found")
- elif not empty_room_recording_path and configs.apply_source_localization:
- empty_room_recording = None
- logger.info("No empty room recording was found")
- else:
- empty_room_recording = None
- elif device == "ARTEMIS123":
- # Artemis recordings may arrive as the native .bin or as a FIF that was
- # converted beforehand. read_raw_artemis123 only accepts the .bin, so
- # anything already in FIF goes through the standard FIF reader.
- if str(path).lower().endswith(".fif"):
- data = mne.io.read_raw_fif(path, preload=True)
- if configs.apply_source_localization:
- data = _add_artemis_headshape(data, path, pos_file, logger)
- else:
- if configs.apply_source_localization:
- data = mne.io.read_raw_artemis123(
- path,
- preload=True,
- pos_fname=_resolve_artemis_pos_file(path, pos_file, logger),
- add_head_trans=True,
- )
- else:
- data = mne.io.read_raw_artemis123(
- path,
- preload=True,
- )
- if empty_room_recording_path and configs.apply_source_localization:
- if str(empty_room_recording_path).lower().endswith(".fif"):
- empty_room_recording = mne.io.read_raw_fif(
- empty_room_recording_path, preload=True
- )
- else:
- empty_room_recording = mne.io.read_raw_artemis123(
- empty_room_recording_path,
- preload=True,
- )
- logger.info("Empty room recording was found")
- elif not empty_room_recording_path and configs.apply_source_localization:
- empty_room_recording = None
- logger.info("No empty room recording was found")
- else:
- empty_room_recording = None
- else:
- data = mne.io.read_raw(path, preload=True)
- if empty_room_recording_path and configs.apply_source_localization:
- empty_room_recording = mne.io.read_raw(
- empty_room_recording_path, preload=True
- )
- logger.info("Empty room recording was found")
- elif not empty_room_recording_path and configs.apply_source_localization:
- empty_room_recording = None
- logger.info("No empty room recording was found")
- else:
- empty_room_recording = None
- return data, empty_room_recording
- def merge_fidp_demo(
- demographic_paths: list,
- features_dir: str,
- dataset_names: list,
- drop_columns: list = ["eyes"],
- ):
- """
- Merge demographic metadata and extracted features into a single DataFrame.
- This function loads demographic data and feature data,
- assigns a site label to each participant if missing, removes unnecessary columns,
- and merges demographic information with corresponding extracted features.
- Parameters
- ----------
- datasets_paths : list
- List of paths to the dataset directories containing demographic files
- ('participants_bids.tsv').
- features_dir : str
- Path to the directory containing the extracted features ('all_features.csv').
- dataset_names : list of str
- List of dataset names corresponding to each dataset path. Used to populate
- missing 'site' information if necessary.
- drop_columns : list of str, optional
- Columns to drop from the demographic data before merging. Default is ["eyes"].
- Returns
- -------
- data : pandas.DataFrame
- Merged DataFrame containing both demographic information and feature data,
- with participants indexed as strings.
- Raises
- ------
- FileNotFoundError
- If the 'participants_bids.tsv' file is missing in any of the dataset paths or
- the 'all_features.csv' file is missing in the provided features directory.
- """
- demographic_df = pd.DataFrame()
- for counter, demo_path in enumerate(demographic_paths):
- if not demo_path or not os.path.exists(demo_path):
- raise FileNotFoundError(
- f"The demographic file for dataset '{dataset_names[counter]}' was not "
- f"found at: {demo_path}. Set 'demographic_path' for this dataset, or "
- "create the file with 'make_demo_file_bids'."
- )
- demo = load_demographic_file(demo_path)
- if "site" not in demo.columns:
- demo["site"] = dataset_names[counter]
- demographic_df = pd.concat([demographic_df, demo], axis=0)
- # Drop unnecessary columns
- if drop_columns:
- demographic_df.drop(columns=drop_columns, errors="ignore", inplace=True)
- # Load features
- feature_path = os.path.join(features_dir, "all_features.csv")
- if not os.path.exists(feature_path):
- raise FileNotFoundError(
- f"The file 'all_features.csv' is missing in the directory: {features_dir}."
- )
- features_df = pd.read_csv(feature_path, index_col=0)
- features_df.index = features_df.index.astype(str)
- # Merge demographic and features
- data = demographic_df.join(features_df, how="inner")
- data.index.name = None
- return data
- def merge_datasets_with_glob(datasets):
- """
- Merges file paths across multiple datasets using glob pattern matching.
- This function walks through the provided datasets' base directories to find
- subject folders and file paths matching a specified task and file ending. It
- creates a dictionary mapping each subject to a glob pattern that can be used
- to aggregate files across multiple runs or sessions.
- Parameters
- ----------
- datasets : dict
- Dictionary where each key is a dataset name, and each value is a dictionary
- with the following keys:
- - "base_dir" (str): Base directory containing subject subdirectories.
- - "task" (str): Task keyword to search for in filenames.
- - "ending" (str): File ending (e.g., '.nii.gz') to filter relevant files.
- Returns
- -------
- dict
- A dictionary mapping subject IDs to a glob-style path string that aggregates
- all matching files for that subject. Only subjects with at least one matched
- file are included.
- Notes
- -----
- This function is designed to assist in scenarios where each subject may have
- multiple files (e.g., different runs or sessions), and the goal is to create
- a single pattern that can be used to load all related files for a subject.
- """
- def join_with_star(lst):
- if not lst:
- return None
- if len(lst) == 1:
- return lst[0] + "*"
- return "*".join(lst)
- subjects = {}
- for dataset_name, dataset_info in datasets.items():
- base_dir = dataset_info["base_dir"]
- task = dataset_info["task"]
- ending = dataset_info["ending"]
- device = dataset_info.get("device_type", None)
- line_freq = dataset_info.get("line_freq", 50)
- empty_room_task = dataset_info.get("empty_room_task", None)
- empty_room_path = dataset_info.get("empty_room_path", None)
- empty_room_ending = dataset_info.get("empty_room_ending", None)
- surfaces = dataset_info.get("surfaces_dir", None)
- event_file_path = dataset_info.get("event_file_path", None)
- event_file_task = dataset_info.get("event_file_task", None)
- event_file_ending = dataset_info.get("event_file_ending", None)
- event_of_interest = dataset_info.get("event_of_interest", None)
- trans_file_p = dataset_info.get("trans_path", None)
- pos_file_p = dataset_info.get("pos_path", None)
- pos_file_ending = dataset_info.get("pos_file_ending", None)
- annotation_p = dataset_info.get("annotation_path", None)
- annotaion_task_name = dataset_info.get("annotaion_task_name", None)
- annotation_ending = dataset_info.get("annotation_ending", None)
- layout_path = dataset_info.get("layout_path", None)
- demographic_path = dataset_info.get(
- "demographic_path", os.path.join(base_dir, "participants_bids.tsv")
- )
- dirs = [
- d for d in os.listdir(base_dir) if os.path.isdir(os.path.join(base_dir, d))
- ]
- for subj in dirs:
- # resting state data
- rs_record_paths = glob.glob(
- f"{base_dir}/{subj}/**/*{task}*{ending}", recursive=True
- )
- # empty room record
- if empty_room_task:
- er_record_paths = glob.glob(
- f"{empty_room_path}/{subj}/**/*{empty_room_task}*{empty_room_ending}",
- recursive=True,
- )
- else:
- er_record_paths = None
- # freesurfer surfaces
- if surfaces:
- if os.path.isdir(os.path.join(surfaces, subj)):
- surface = surfaces
- else:
- surface = None
- else:
- surface = None
- # event file
- if event_file_task and event_file_path:
- event_record_paths = glob.glob(
- f"{event_file_path}/{subj}/**/*{event_file_task}*{event_file_ending}",
- recursive=True,
- )
- else:
- event_record_paths = None
- # trans file
- if trans_file_p:
- trans_path = glob.glob(
- f"{trans_file_p}/{subj}/**/*-trans.fif", recursive=True
- )
- else:
- trans_path = None
- # pos file
- if pos_file_p and pos_file_ending:
- pos_path = glob.glob(
- f"{pos_file_p}/{subj}/**/*{pos_file_ending}", recursive=True
- )
- else:
- pos_path = None
- if annotation_p:
- annotation_path = glob.glob(
- f"{annotation_p}/*{subj}*/**/*{annotaion_task_name}*{annotation_ending}",
- recursive=True,
- )
- else:
- annotation_path = None
- subjects.update(
- {
- subj: {
- "rest_record": join_with_star(rs_record_paths),
- "device": device,
- "line_freq": str(line_freq),
- "empty_room_record": join_with_star(er_record_paths),
- "mri_surface": surface,
- "dataset_name": dataset_name,
- "event_record": join_with_star(event_record_paths),
- "event_of_interest": str(event_of_interest),
- "trans_path": join_with_star(trans_path),
- "pos_path": join_with_star(pos_path),
- "annotation_path": join_with_star(annotation_path),
- "layout_path": layout_path,
- "demographic_path": demographic_path,
- }
- }
- )
- return subjects
- def load_demographic_file(path, index_col=0):
- """Read a participants/demographic table (.tsv, .txt, .csv, .xlsx)."""
- ext = os.path.splitext(path)[1].lower()
- if ext in (".tsv", ".txt"):
- df = pd.read_csv(path, sep="\t", index_col=index_col)
- elif ext == ".csv":
- df = pd.read_csv(path, index_col=index_col)
- elif ext in (".xlsx", ".xls"):
- df = pd.read_excel(path, index_col=index_col)
- else:
- raise ValueError(f"Unsupported demographic file type: {path}")
- df.index = df.index.astype(str)
- return df
- def make_demo_file_bids(
- file_dir: str, save_dir: str, id_col: int, age_col: int, *columns
- ) -> None:
- """
- Convert formats of demographic data into a single format so it can be used
- in later stages.
- Parameters
- ----------
- file_dir : str
- Path to the input demographic file (supports CSV, TSV, or XLSX).
- save_dir : str
- Path where the BIDS-formatted demographic file will be saved (as TSV).
- id_col : int
- Column index containing the participant ID.
- age_col : int
- Column index containing participant age.
- *extra_columns : dict
- Additional column definitions. While age and participants id were defined
- using positional arguments, extra coulmn modification (e.g., sex and eyes
- condition) can be revised and converted to a single format across dataset
- using this function. Each dict can contain:
- - 'col_name': str, required name for the output column. This does not
- necessarly match the column name before being passed to this function.
- - 'col_id': int, index of the column that the revision should be applied to.
- - 'single_value': value to assign to all rows if no col_id and mapping are given.
- This can be helpful when all subjects in a dataset have the same properties
- e.g., eyes open condition.
- - 'mapping': dict, if single value is not defined, value mapping can be passed
- to map the initial values to the target values.
- Returns
- -------
- None
- """
- for col in columns:
- if col.get("single_value") and col.get("mapping"):
- raise ValueError(
- "'single_value' and 'mapping' can not be both defined. One of them must be None; see the documentation!"
- )
- # Load input file based on extension
- if file_dir.endswith(".xlsx"):
- df = pd.read_excel(file_dir)
- elif file_dir.endswith(".csv"):
- df = pd.read_csv(file_dir)
- elif file_dir.endswith(".tsv"):
- df = pd.read_csv(file_dir, sep="\t")
- else:
- raise ValueError(f"Unsupported file type for: {file_dir}")
- # Initialize new dataframe with required fields
- new_df = pd.DataFrame(
- {"participant_id": df.iloc[:, id_col], "age": df.iloc[:, age_col]}
- )
- for col in columns:
- col_name = col.get("col_name")
- col_id = col.get("col_id")
- mapping = col.get("mapping")
- single_value = col.get("single_value")
- if col_name is None:
- raise ValueError("Each column dictionary must contain a 'col_name'.")
- if col_id is not None:
- new_df[col_name] = df.iloc[:, col_id]
- if mapping:
- new_df[col_name] = new_df[col_name].map(mapping)
- elif single_value is not None:
- new_df[col_name] = single_value
- else:
- raise ValueError(
- f"Column '{col_name}' must have either 'col_id' or 'single_value'."
- )
- # Special case handling
- if col_name == "diagnosis":
- new_df[col_name] = new_df[col_name].fillna("nan")
- # Remove duplicate participants
- new_df = new_df.drop_duplicates(subset="participant_id", keep="first")
- save_dir = os.path.join(save_dir, "participants_bids.tsv")
- # Save as BIDS-compatible TSV
- new_df.to_csv(save_dir, sep="\t", index=False)
- def set_path(project_dir):
- """
- Create and initialize directory structure for a given project.
- This function generates a set of predefined directories for
- feature extraction and normative modeling workflows within the
- specified project directory. If any of these directories do not
- exist, they will be created. The function returns the path to the
- features log directory.
- Parameters
- ----------
- project_dir : str
- Path to the root project directory where the folder structure
- will be created.
- Returns
- -------
- str
- Absolute path to the 'log' directory inside the 'Features'
- folder.
- Notes
- -----
- The function creates the following directory structure:
- - ``Features/``
- - ``log/`` (for saving logs of feature extraction)
- - ``temp/`` (for temporarily storing extracted features)
- - ``figures/`` (for saving generated figures)
- - ``Normative modeling/``
- - ``Runs/`` (for saving model run outputs)
- - ``Figures/`` (for visual outputs related to modeling)
- - ``Models summary/`` (for summaries of model results)
- """
- def make_folder(path):
- if not os.path.isdir(path):
- os.makedirs(path)
- # Feature extraction
- features_dir = os.path.join(project_dir, "Features")
- features_log_path = os.path.join(features_dir, "log_slurm_jobs")
- features_temp_path = os.path.join(features_dir, "temp")
- exluded_participants_path = os.path.join(features_dir, "excluded_participants")
- saved_outputs_path = os.path.join(features_dir, "Saved_outputs")
- save_epochs_path = os.path.join(saved_outputs_path, "Epochs")
- save_psds_path = os.path.join(saved_outputs_path, "PSDs")
- save_preprocessed_data = os.path.join(saved_outputs_path, "Preprocessed_data")
- save_coregistration_QC_path = os.path.join(saved_outputs_path, "coregistration_QC")
- save_auto_reject_plot = os.path.join(saved_outputs_path, "auto_reject_plot")
- save_covariance_figures_path = os.path.join(
- saved_outputs_path, "Covariance_figures"
- )
- save_BEM_figures_path = os.path.join(saved_outputs_path, "BEM_figures")
- save_transformation_path = os.path.join(
- saved_outputs_path, "transformation_FIF_file"
- )
- configurations = os.path.join(features_dir, "Configurations")
- mri_templates = os.path.join(features_dir, "MRI_templates")
- save_grouping_effect = os.path.join(saved_outputs_path, "Grouping_effects")
- make_folder(features_dir)
- make_folder(features_log_path)
- make_folder(features_temp_path)
- make_folder(saved_outputs_path)
- make_folder(save_epochs_path)
- make_folder(save_psds_path)
- make_folder(save_coregistration_QC_path)
- make_folder(save_covariance_figures_path)
- make_folder(save_BEM_figures_path)
- make_folder(save_transformation_path)
- make_folder(configurations)
- make_folder(exluded_participants_path)
- make_folder(mri_templates)
- make_folder(save_grouping_effect)
- make_folder(save_preprocessed_data)
- make_folder(save_auto_reject_plot)
- # Normative models
- nm_dir = os.path.join(project_dir, "Normative_models")
- make_folder(nm_dir)
- return features_dir, features_log_path
- def clean_nan_columns(df, nan_threshold):
- """
- Remove columns with more NaNs than `nan_threshold`,
- otherwise impute NaNs with the column median.
- Parameters:
- df (pd.DataFrame): Input dataframe
- nan_threshold (int): Max allowed NaNs per column
- Returns:
- pd.DataFrame: Cleaned dataframe
- """
- df = df.copy()
- for col in df.columns:
- nan_count = df[col].isna().sum()
- if nan_count > nan_threshold:
- # Drop column
- df.drop(columns=col, inplace=True)
- else:
- # Impute with median (only for numeric columns)
- if pd.api.types.is_numeric_dtype(df[col]):
- median_value = df[col].median()
- df[col] = df[col].fillna(median_value)
- return df
- def find_other_mri_session(
- base_mri_path, missing_mri_subjects, str_mri_ending, which_session
- ):
- """
- Find an alternative MRI file for each subject at a given session index.
- Searches recursively under each subject's directory for files matching
- the given filename ending, then selects the file at the requested
- session index. Intended for locating a fallback MRI session (e.g. a
- second scan) for subjects whose primary MRI is missing or unusable.
- Parameters
- ----------
- base_mri_path : str or Path
- Root directory containing per-subject MRI folders.
- missing_mri_subjects : iterable of str
- Subject IDs to search for.
- str_mri_ending : str
- Filename suffix/pattern used to match MRI files (e.g. "T1w.nii.gz").
- which_session : int
- 1-based index of the session to select from each subject's matched
- file list.
- Returns
- -------
- dict[str, str]
- Mapping of subject ID to the matched MRI file path. Subjects with
- fewer than `which_session` matching files are omitted.
- """
- new_paths = {}
- for subject in missing_mri_subjects:
- mri_paths = glob.glob(
- f"{base_mri_path}/{subject}/**/*{str_mri_ending}", recursive=True
- )
- if len(mri_paths) > which_session - 1:
- new_paths.update({subject: mri_paths[which_session - 1]})
- return new_paths
- def find_failed_meg_subjects(log_path):
- """
- Identify subjects whose MEG processing failed, based on log files.
- Scans the given directory for error log files (files with "err" in the
- name) and flags any subject whose log contains the string "error".
- Parameters
- ----------
- log_path : str or Path
- Directory containing per-subject log files.
- Returns
- -------
- set of str
- Unique subject IDs (parsed from the log filename, before the first
- underscore) whose logs indicate a processing error.
- """
- missing_meg_subjects = []
- paths = os.scandir(log_path)
- paths = list(filter(lambda x: "err" in x.name, paths))
- for path in paths:
- with open(path, "r") as f:
- content = f.read()
- if "error" in content:
- subject = os.path.basename(path).split(".")[0].split("_")[0]
- missing_meg_subjects.append(subject)
- return set(missing_meg_subjects)
- def find_other_meg_session(
- base_meg_path, missing_meg_subjects, str_meg_ending, task_name, which_session
- ):
- """
- Find an alternative MEG recording for each subject at a given session index.
- Searches recursively under each subject's directory for files matching
- the given task name and filename ending, then selects the file at the
- requested session index. Intended for locating a fallback MEG session
- for subjects whose primary recording is missing or failed processing.
- Parameters
- ----------
- base_meg_path : str or Path
- Root directory containing per-subject MEG folders.
- missing_meg_subjects : iterable of str
- Subject IDs to search for.
- str_meg_ending : str
- Filename suffix/pattern used to match MEG files (e.g. "raw.fif").
- task_name : str
- Task identifier expected to appear in the filename (e.g. "rest").
- which_session : int
- 1-based index of the session to select from each subject's matched
- file list.
- Returns
- -------
- dict[str, str]
- Mapping of subject ID to the matched MEG file path. Subjects with
- fewer than `which_session` matching files are omitted.
- """
- new_paths = {}
- for subject in missing_meg_subjects:
- rs_record_paths = glob.glob(
- f"{base_meg_path}/{subject}/**/*{task_name}*{str_meg_ending}",
- recursive=True,
- )
- if len(rs_record_paths) > which_session - 1:
- new_paths.update({subject: rs_record_paths[which_session - 1]})
- return new_paths
- def check_demographic_format(df):
- """Raise ValueError if a demographic table is not in the standard format."""
- DEMOGRAPHIC_REQUIRED_COLUMNS = ("participant_id", "age", "sex", "eyes")
- DEMOGRAPHIC_ALLOWED_SEX = ("Male", "Female")
- if "participant_id" in df.columns:
- ids = df["participant_id"]
- else:
- ids = df.index.to_series()
- for col in DEMOGRAPHIC_REQUIRED_COLUMNS:
- if col != "participant_id" and col not in df.columns:
- raise ValueError(f"The demographic file is missing the '{col}' column.")
- if not all(isinstance(v, str) for v in ids.dropna()):
- raise ValueError("All participant IDs must be strings.")
- if not pd.api.types.is_numeric_dtype(df["age"]):
- raise ValueError("The 'age' column must hold integers or floats.")
- unexpected = set(df["sex"].dropna().unique()) - set(DEMOGRAPHIC_ALLOWED_SEX)
- if unexpected:
- raise ValueError(
- f"The 'sex' column holds {sorted(unexpected)}; "
- f"only {list(DEMOGRAPHIC_ALLOWED_SEX)} are allowed."
- )
IO.py at commit a1fa698, under GPL-3.0 · at the source
Overview
13 affiliations
- Department of Psychiatry and Psychotherapy, University Hospital Tübingen, Tübingen, Germany
- Graduate Training Center of Neuroscience, International Max-Planck Research School, Tübingen, Germany
- Max-Planck Institute for Intelligent Systems, Tübingen, Germany
- German Center for Mental Health (DZPG), partner site Tübingen, Germany
- Department of Neurology, Jena University Hospital, Jena, Germany
- Center of Cognitive Science and Artificial Intelligence, Tilburg University, Tilburg, The Netherlands
- Donders Institute for Cognition, Brain and Behavior, Radboud University, Nijmegen, The Netherlands
- Department of Psychiatry, UMC Utrecht Brain Center, University Medical Center, Utrecht, The Netherlands
- The MARCS Institute for Brain, Behaviour and Development, Western Sydney University, Sydney, Australia
- CIBM Center for Biomedical Imaging, University of Geneva, Geneva, Switzerland
- Department of Clinical Neuroscience, University of Geneva, Geneva, Switzerland
- German Center for Mental Health (DZPG), Jena-Magdeburg-Halle, Jena, Germany
- Institute of Psychology, Friedrich Schiller University Jena, Jena, Germany
Abstract
Normative modeling has become a cornerstone of computational neuroscience, offering a powerful framework for detecting individual deviations from typical brain function. This review traces its trajectory in electrophysiology of the brain, from early studies in the 1970s, through a period of relative neglect, to its recent revival driven by machine learning advances and the availability of large-scale datasets. We provide a structured overview of this evolution, showing the shift from small, site-specific age-based models to increasingly harmonized approaches that integrate diverse biological and methodological innovations. Key studies are compared with respect to cohort composition, modeling strategies, and validation procedures, situating each within the broader arc of methodological progress. Despite this momentum, significant challenges remain, such as a lack of standardized practices, limited comparability across studies, and the need for richer integration of complex neurophysiological signals. Looking ahead, we argue that the future of electrophysiological normative modeling lies in scaling and unifying efforts, through international collaborations, standardized pipelines, and the incorporation of novel features. By coupling machine learning with both cross-sectional and longitudinal designs, the field can progress from proof-of-concept demonstrations to precise, individualized brain mapping. Ultimately, such advances will provide the foundation for clinical applications, enabling cost-effective personalized treatment monitoring and a more refined understanding of individual differences in complex brain disorders.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.
Zenodo 15441320
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
21 files
- docs/
source/ , Python, 44 linesconf.py - meganorm/
__init__.py , Python, 6 lines - meganorm/
_version.py , Python, 1 line - meganorm/
layouts/ , Python, 1 line__init__.py - meganorm/
layouts/ , Python, 1,427 lineslayouts.py - meganorm/
plots/ , Python, 1 line__init__.py - meganorm/
plots/ , Python, 1,611 linesplots.py - meganorm/
src/ , Python, 4 lines__init__.py - meganorm/
src/ , Python, 856 linesfeatureExtraction.py - meganorm/
src/ , Python, 197 linesmainParallel.py - meganorm/
src/ , Python, 594 linespreprocess.py - meganorm/
src/ , Python, 222 linespsdParameterize.py - meganorm/
utils/ , Python, 550 linesEEGlab.py - meganorm/
utils/ , Python, 652 linesIO.py - meganorm/
utils/ , Python, 5 lines__init__.py - meganorm/
utils/ , Python, 281 linesfreesurfer.py - meganorm/
utils/ , Python, 1,067 lines, 1 matchnm.py - meganorm/
utils/ , Python, 471 linesparallel.py - setup.py, Python, 37 lines
- LICENSE, License, 674 lines
- README.md, Text, 141 lines
ml4pnp/meganorm
a1fa6984c442e310098f3b3b0a73c1d880bb149a, 16 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
24 files
- docs/
source/ , Python, 58 linesconf.py - meganorm/
__init__.py , Python, 6 lines - meganorm/
_version.py , Python, 1 line - meganorm/
layouts/ , Python, 1 line__init__.py - meganorm/
layouts/ , Python, 1,427 lineslayouts.py - meganorm/
plots/ , Python, 1 line__init__.py - meganorm/
plots/ , Python, 3,143 linesplots.py - meganorm/
src/ , Python, 4 lines__init__.py - meganorm/
src/ , Python, 1,196 linesfeatureExtraction.py - meganorm/
src/ , Python, 635 lines, 1 matchmainParallel.py - meganorm/
src/ , Python, 774 lines, 2 matchesnormative_modeling.py - meganorm/
src/ , Python, 3,053 linespreprocess.py - meganorm/
src/ , Python, 412 linespsdParameterize.py - meganorm/
src/ , Python, 1,738 linessource_localization.py - meganorm/
utils/ , Python, 1,470 lines, 2 matchesIO.py - meganorm/
utils/ , Python, 4 lines__init__.py - meganorm/
utils/ , Python, 313 linesdata_specific_utils.py - meganorm/
utils/ , Python, 965 linesfreesurfer.py - meganorm/
utils/ , Python, 927 lines, 1 matchnm.py - meganorm/
utils/ , Python, 911 linesparallel.py - notebooks/
feature_extraction_tutor , Jupyter, 533 lines, 1 matchial.ipynb - setup.py, Python, 40 lines
- LICENSE, License, 674 lines
- README.md, Text, 161 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 41 scripts, each with its path and the digest of its content;
- 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data and Code Availability
This review used no new or pre-existing data beyond the cited manuscripts and their associated dataset information, all referenced accordingly. No code was developed for or is fundamental to the reproducibility of this study.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 8 authors, 7 keywords, 3 funders, 293 references.
Cite
This paper
Mallus, F. A., Yang, Y., Dinga, R., Kia, M. S., Grootswagers, T., Michela, A., Ros, T., & Wolfers, T. (2026). From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals. Imaging neuroscience (Cambridge, Mass.), 4, IMAG.a.1269. https://
BibTeX
@article{mallus2026early
author = {Mallus, Francesco Antonio and Yang, Yanwu and Dinga, Richard and Kia, Mostafa Seyed and Grootswagers, Tijl and Michela, Abele and Ros, Tomas and Wolfers, Thomas},
title = {{From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals}},
journal = {Imaging neuroscience (Cambridge, Mass.)},
year = {2026},
month = jun,
volume = {4},
pages = {IMAG.a.1269},
publisher = {MIT Press},
issn = {2837-6056},
doi = {10.1162/
url = {https://
pmid = {42312085},
pmcid = {PMC13271153}
}
RIS
TY - JOUR
AU - Mallus, Francesco Antonio
AU - Yang, Yanwu
AU - Dinga, Richard
AU - Kia, Mostafa Seyed
AU - Grootswagers, Tijl
AU - Michela, Abele
AU - Ros, Tomas
AU - Wolfers, Thomas
TI - From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals
T2 - Imaging neuroscience (Cambridge, Mass.)
J2 - Imaging Neurosci (Camb)
PY - 2026
DA - 2026/
VL - 4
SP - IMAG.a.1269
SN - 2837-6056
PB - MIT Press
DO - 10.1162/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1162/
"type": "article-journal",
"title": "From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals",
"container-title": "Imaging neuroscience (Cambridge, Mass.)",
"author": [
{
"family": "Mallus",
"given": "Francesco Antonio"
},
{
"family": "Yang",
"given": "Yanwu"
},
{
"family": "Dinga",
"given": "Richard"
},
{
"family": "Kia",
"given": "Mostafa Seyed"
},
{
"family": "Grootswagers",
"given": "Tijl"
},
{
"family": "Michela",
"given": "Abele"
},
{
"family": "Ros",
"given": "Tomas"
},
{
"family": "Wolfers",
"given": "Thomas"
}
],
"container-title-short":
"volume": "4",
"page": "IMAG.a.1269",
"DOI": "10.1162/
"PMID": "42312085",
"PMCID": "PMC13271153",
"ISSN": "2837-6056",
"publisher": "MIT Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
15
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.64898/2026.03.31.26349848 [code]
- Mapping Individual Neuroanatomical Alterations to Schizophrenia Psychopathology with Normative ModelingJournal: medRxiv (preprint)In common: FreeSurfer, seaborn, pandas, 2 other tools, computational, 15 references
- [2] doi:10.1038/s41597-026-07350-9 [code]
- An open multi-center MEG-EEG dataset for studying conscious visual perception.Journal: Scientific dataIn common: autoreject, xarray, Pingouin, 11 other tools, MEG, EEG, 5 references
- [3] doi:10.7554/elife.100605 [code]
- Age-related changes in ‘cortical’ 1/
f dynamics are linked to cardiac activity Journal: n/aIn common: ArviZ, PyMC, specparam (formerly FOOOF), 10 other tools, MEG, 3 references - [4] doi:10.64898/2026.08.13.26360304 [code]
- Lifespan brain structural variation reveals shared organization across mental health conditionsJournal: medRxiv (preprint)In common: Nilearn, statsmodels, NiBabel, 6 other tools, 11 references
- [5] doi:10.1038/s41380-026-03691-4 [code]
- Breaking the norm: population-scale deviations of brain structure in depression and anxiety.Journal: Molecular psychiatryIn common: Nilearn, statsmodels, NiBabel, 6 other tools, 9 references
- [6] doi:10.1038/s41467-026-72940-5 [code]
- Cerebellar growth is associated with domain-specific cerebral maturation and socio-linguistic behavior.Journal: Nature communicationsIn common: ArviZ, PyMC, xarray, 8 other tools, 5 references
- [7] doi:10.1038/s41467-026-72875-x [code]
- Lifespan normative modeling of brain microstructure.Journal: Nature communicationsIn common: seaborn, scikit-learn, pandas, 3 other tools, 12 references
- [8] doi:10.1038/s41597-026-07377-y [code]
- An open-access multi-site fMRI dataset for investigating conscious visual perception.Journal: Scientific dataIn common: autoreject, xarray, Pingouin, 11 other tools, 1 reference
- [9] doi:10.1093/nc/niag029 [code]
- A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.Journal: Neuroscience of consciousnessIn common: autoreject, xarray, Pingouin, 11 other tools, MEG
- [10] doi:10.1162/imag.a.1245 [code]
- Towards precision EEG connectomics: Evaluating the benefits of dense sampling.Journal: Imaging neuroscience (Cambridge, Mass.)In common: autoreject, ICLabel, Pingouin, 10 other tools, EEG, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 41 scripts, and 8 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:9f7e4b1f16ff34d0…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
