Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.
The 25 matches
- [1] § 3. Materials and Methods › 3.2. Feature Extraction ↔ experiments/gen_latex_tables.py, lines 286–309 · score 0.99 · Lempel Ziv complexity, Higuchi fractal dimension, median threshold binarization, Hilbert phase, Hann window, Hjorth parameters
- [2] § 3. Materials and Methods › 3.2. Feature Extraction ↔ experiments/featurelib.py, lines 1–25 · score 0.98 · Lempel Ziv complexity, Higuchi fractal dimension, median threshold binarization, Hann window, log asymmetries, permutation entropy
- [3] § 3. Materials and Methods › 3.3. Label Binarization and Evaluation Protocols ↔ experiments/run_riemann.py, lines 1–34 · score 0.93 · Riemannian tangent space, trial band limited, domain adaptation, Euclidean Alignment, GroupKFold, covariance
- [4] § 4. Results › 4.2. Participant-Independent Performance ↔ experiments/gen_latex_tables.py, lines 463–492 · score 0.88 · soft voting ensembles, Riemannian tangent space, Euclidean Alignment, feature selection, ceiling, nested
- [5] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/gen_latex_tables.py, lines 85–105 · score 0.84 · class recall, validation split, Decision thresholds, PR AUC, confidence intervals, ROC AUC
- [6] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/gen_latex_tables.py, lines 223–239 · score 0.79 · pairwise Jaccard overlap, Feature selection stability, independent folds, XGBoost, Kuncheva
- [7] § 4. Results › 4.5. Personalized Models and Multimodal Fusion ↔ experiments/gen_latex_tables.py, lines 494–520 · score 0.79 · peripheral autonomic, interpretable peripheral, EEG yields, fused, weak, modality
- [8] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/run_stability_importance.py, lines 1–32 · score 0.78 · pairwise Jaccard, Feature selection stability, permutation importance, XGBoost, Kuncheva, held
- [9] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/run_final_opt.py, lines 1–33 · score 0.76 · soft voting ensemble, mutual information, feature selection, inner, bootstrapping, AUC
- [10] § 4. Results › 4.6. Effect Sizes, Ablations, and Feature Importance ↔ experiments/gen_latex_tables.py, lines 245–259 · score 0.75 · unreliable indicator, predictor relevance, permutation importance, XGBoost, Spearman, correlation
- [11] § 4. Results › 4.3. A Participant-Specific Signature in the Feature Space ↔ experiments/gen_latex_tables.py, lines 311–356 · score 0.75 · emotion decoding, variance decomposition, emotional state, random forest, qEEG, ratio
- [12] § 4. Results › 4.6. Effect Sizes, Ablations, and Feature Importance ↔ experiments/gen_latex_tables.py, lines 286–309 · score 0.74 · Lempel Ziv complexity, Frontal alpha asymmetry, permutation entropy, gamma, beta, theta
- [13] § 4. Results › 4.6. Effect Sizes, Ablations, and Feature Importance ↔ experiments/run_stability_importance.py, lines 93–140 · score 0.69 · permutation importance, model agnostic, XGBoost, tree, Spearman, held
- [14] § 4. Results › 4.6. Effect Sizes, Ablations, and Feature Importance ↔ experiments/featurelib.py, lines 1–25 · score 0.67 · Lempel Ziv complexity, permutation entropy, gamma, beta, theta, asymmetry
- [15] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/run_stats.py, lines 1–30 · score 0.67 · Benjamini Hochberg, Kruskal Wallis, Bonferroni, Wilcoxon, bootstrap, Model
- [16] § 4. Results › 4.5. Personalized Models and Multimodal Fusion ↔ experiments/run_multimodal.py, lines 1–22 · score 0.66 · peripheral autonomic, Adding interpretable peripheral, fused, cortical, EMG, respiration
- [17] § 4. Results › 4.2. Participant-Independent Performance ↔ experiments/run_riemann.py, lines 1–34 · score 0.64 · Riemannian tangent space, Euclidean Alignment, compact, vector, channel, protocol
- [18] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/run_stats.py, lines 1–30 · score 0.64 · Wilcoxon signed rank, Kruskal Wallis, ANOVA, bootstrap, class, DEAP
- [19] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/gen_latex_tables.py, lines 178–215 · score 0.63 · Wilcoxon signed rank, Kruskal Wallis, inflation, bootstrap
- [20] § 3. Materials and Methods › 3.3. Label Binarization and Evaluation Protocols ↔ experiments/gen_latex_tables.py, lines 463–492 · score 0.63 · Riemannian tangent space, Euclidean Alignment, classifier, model
- [21] § 4. Results › 4.1. Effect of the Evaluation Protocol ↔ experiments/run_revision_stats.py, lines 1–35 · score 0.63 · near duplicate, participant overlap, training fold, leakage, split, channels
- [22] § 4. Results › 4.6. Effect Sizes, Ablations, and Feature Importance ↔ experiments/gen_latex_tables.py, lines 124–140 · score 0.61 · feature family dominates, Functional connectivity, measurable, Ablations, channel, valence
- [23] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/gen_latex_tables.py, lines 85–105 · score 0.56 · bootstrap confidence interval, threshold tuning, ROC AUC, Classifiers, modeling
- [24] § 4. Results › 4.1. Effect of the Evaluation Protocol ↔ experiments/gen_latex_tables.py, lines 42–70 · score 0.55 · class imbalance, GroupKFold, XGBoost, binarization, median, AUCs
- [25] § 3. Materials and Methods › 3.4. Classifiers and Statistical Analysis ↔ experiments/run_personalization.py, lines 131–166 · score 0.51 · logistic regression, class weights, ROC AUC, Wilcoxon, Classifiers
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 520 lines · 23 KB · no license · 13 matches
- """Emit English LaTeX result tables (booktabs) from computed JSONs/CSVs into
- overelaf-paper/revised/tables.tex. Every number is pulled from results/ so the
- manuscript cannot drift from the experiments."""
- from __future__ import annotations
- import json
- from pathlib import Path
- import numpy as np
- import pandas as pd
- HERE = Path(__file__).resolve().parent
- R = HERE / "results"
- OUT = HERE.parent / "overelaf-paper/revised"
- OUT.mkdir(parents=True, exist_ok=True)
- def load(n): return json.load(open(R / n))
- def ci(d): return f"{d['value']:.3f} [{d['ci'][0]:.3f}, {d['ci'][1]:.3f}]"
- blocks = []
- # ---- Table: protocol / leakage ----
- ce = load("core_eval.json")["leakage_demo"]
- t = r"""\begin{table}[H]
- \caption{Effect of the evaluation protocol on apparent performance. The same
- XGBoost model and 6-channel feature set are evaluated three ways; only the data
- partition changes. Window-level pooling allows overlapping windows from the same
- subject into both train and test, inflating AUC by roughly 0.20.\label{tab:protocol}}
- \centering
- \begin{tabular}{lcc}
- \toprule
- \textbf{Evaluation protocol} & \textbf{Valence AUC} & \textbf{Arousal AUC} \\
- \midrule
- Window-level, pooled (random $k$-fold) & %.3f & %.3f \\
- Window-level, subject-grouped & %.3f & %.3f \\
- Trial-level, subject-grouped + subject-norm & %.3f & %.3f \\
- \bottomrule
- \end{tabular}
- \end{table}""" % (
- ce["valence"]["window_pooled_leaky_auc"][0], ce["arousal"]["window_pooled_leaky_auc"][0],
- ce["valence"]["window_grouped_auc"][0], ce["arousal"]["window_grouped_auc"][0],
- ce["valence"]["trial_grouped_subjnorm_auc"], ce["arousal"]["trial_grouped_subjnorm_auc"])
- blocks.append(t)
- # ---- Table: labeling sweep ----
- sw = load("core_eval.json")["labeling_sweep"]
- names = {"fixed5": "Fixed threshold ($>5$)", "global_median": "Global median",
- "subject_median": "Per-subject median", "margin": "Margin ($|r-5|\\ge 2$)"}
- rows = []
- for s in ["fixed5", "global_median", "subject_median", "margin"]:
- v = sw["valence"][s]; a = sw["arousal"][s]
- rows.append(f"{names[s]} & {v['class_balance']:.2f} & {v['n_trials']} & "
- f"{v['models']['xgb']['auc']['value']:.3f} & "
- f"{a['class_balance']:.2f} & {a['n_trials']} & "
- f"{a['models']['xgb']['auc']['value']:.3f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Sensitivity to the binarization scheme (XGBoost, subject-independent
- GroupKFold, subject-normalized). The fixed threshold of 5 produces severe class
- imbalance and near-chance AUC; per-subject median balances the classes and gives
- the best stable valence performance. Margin labels reduce $N$ and inflate
- variance.\label{tab:labeling}}
- \centering
- \begin{tabular}{lcccccc}
- \toprule
- & \multicolumn{3}{c}{\textbf{Valence}} & \multicolumn{3}{c}{\textbf{Arousal}} \\
- \cmidrule(lr){2-4}\cmidrule(lr){5-7}
- \textbf{Labeling scheme} & Balance & $N$ & AUC & Balance & $N$ & AUC \\
- \midrule
- """ + "\n".join(rows) + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: subject-independent classification (headline) ----
- hl = load("headline.json")
- mlabel = {"logreg": "Logistic Regression", "linsvm": "Linear SVM", "rbfsvm": "RBF SVM",
- "rf": "Random Forest", "xgb": "XGBoost", "majority": "Majority baseline"}
- def hl_rows(key):
- d = hl[key]["models"]; out = []
- for m in ["majority", "logreg", "linsvm", "rbfsvm", "rf", "xgb"]:
- prot = "loso" if m != "majority" else None
- mk = f"{m}_loso" if prot else "majority"
- v = d[mk]
- out.append(f"{mlabel[m]} & {ci(v['auc'])} & {v['pr_auc']['value']:.3f} & "
- f"{v['bacc']['value']:.3f} & {v['recall_pos']['value']:.2f} & {v['recall_neg']['value']:.2f} \\\\")
- return "\n".join(out)
- t = r"""\begin{table}[H]
- \caption{Subject-independent classification (leave-one-subject-out, per-subject
- median labels, per-subject normalization, decision threshold tuned by Youden's
- $J$ on an inner validation split). Brackets give 95\% bootstrap confidence
- intervals obtained by resampling subjects. Recall$_+$/Recall$_-$ are per-class
- recall.\label{tab:classification}}
- \centering
- \footnotesize
- \begin{tabular}{lccccc}
- \toprule
- \textbf{Model} & \textbf{ROC-AUC [95\% CI]} & \textbf{PR-AUC} & \textbf{Bal.\ Acc.} & \textbf{Recall$_+$} & \textbf{Recall$_-$} \\
- \midrule
- \multicolumn{6}{c}{\textit{Valence}} \\
- """ + hl_rows("valence_subject_median") + r"""
- \midrule
- \multicolumn{6}{c}{\textit{Arousal}} \\
- """ + hl_rows("arousal_subject_median") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: ablations ----
- ab = load("ablations.json")
- def fam_rows(tgt):
- out = []
- fam6 = ab["family_6ch"][tgt]; fam32 = ab["family_32ch"][tgt]
- famnames = {"spectral": "Spectral only", "nonlinear": "Nonlinear only",
- "connectivity": "Connectivity only", "no_connectivity": "All $-$ connectivity",
- "all": "All features"}
- for f in ["spectral", "nonlinear", "all"]:
- d6 = fam6[f]
- out.append(f"{famnames[f]} & {d6['n_features']} & {d6['results']['xgb']['auc']['value']:.3f} & "
- + (f"{fam32[f]['n_features']} & {fam32[f]['results']['xgb']['auc']['value']:.3f}" if f in fam32 else "-- & --") + r" \\")
- for f in ["connectivity", "no_connectivity"]:
- if f in fam32:
- d = fam32[f]
- out.append(f"{famnames[f]} & -- & -- & {d['n_features']} & {d['results']['xgb']['auc']['value']:.3f} \\\\")
- return "\n".join(out)
- t = r"""\begin{table}[H]
- \caption{Feature-family ablation (XGBoost, subject-independent GroupKFold,
- per-subject median labels). No family dominates; functional connectivity (PLV and
- coherence) adds no measurable value over local features, and connectivity-only is
- at chance.\label{tab:ablation_family}}
- \centering
- \begin{tabular}{lcccc}
- \toprule
- & \multicolumn{2}{c}{\textbf{6-channel set}} & \multicolumn{2}{c}{\textbf{32-channel set}} \\
- \cmidrule(lr){2-3}\cmidrule(lr){4-5}
- \textbf{Feature family} & $d$ & Valence AUC & $d$ & Valence AUC \\
- \midrule
- """ + fam_rows("valence") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- def chan_rows(tgt):
- cn = {"frontal": "Frontal (F3,F4,AF3,AF4)", "parietal": "Parietal (P3,P4)",
- "all6": "All 6 channels", "all32": "All 32 channels"}
- ch = ab["channel"][tgt]
- return "\n".join(f"{cn[c]} & {ch[c]['n_features']} & {ch[c]['results']['xgb']['auc']['value']:.3f} \\\\"
- for c in ["frontal", "parietal", "all6", "all32"])
- t = r"""\begin{table}[H]
- \caption{Channel ablation (XGBoost, subject-independent GroupKFold, per-subject
- median labels). For valence the parietal pair alone matches the full montage and
- exceeds the frontal channels; the full 32-channel set offers little
- gain.\label{tab:ablation_channel}}
- \centering
- \begin{tabular}{lcc|cc}
- \toprule
- & \multicolumn{2}{c}{\textbf{Valence}} & \multicolumn{2}{c}{\textbf{Arousal}} \\
- \textbf{Channel subset} & $d$ & AUC & $d$ & AUC \\
- \midrule
- """ + "\n".join(
- f"{ {'frontal':'Frontal (F3,F4,AF3,AF4)','parietal':'Parietal (P3,P4)','all6':'All 6 channels','all32':'All 32 channels'}[c] } & "
- f"{ab['channel']['valence'][c]['n_features']} & {ab['channel']['valence'][c]['results']['xgb']['auc']['value']:.3f} & "
- f"{ab['channel']['arousal'][c]['n_features']} & {ab['channel']['arousal'][c]['results']['xgb']['auc']['value']:.3f} \\\\"
- for c in ["frontal", "parietal", "all6", "all32"]) + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: effect sizes / within-subject (replaces old Tables 3 & 4) ----
- def eff_rows(tgt):
- df = pd.read_csv(R / f"stats_{tgt}_subject_median.csv").sort_values("p_subj_wilcoxon").head(10)
- out = []
- for _, r0 in df.iterrows():
- feat = r0["feature"].replace("_", r"\_")
- out.append(f"{feat} & {r0['eps2']:.4f} [{r0['eps2_lo']:.4f}, {r0['eps2_hi']:.4f}] & "
- f"{r0['cohens_dz_subj']:+.2f} & {r0['p_subj_wilcoxon']:.1e} & {r0['subj_fdr']:.3f} & {r0['p_kw_window']:.1e} \\\\")
- return "\n".join(out)
- t = r"""\begin{table}[H]
- \caption{Top qEEG features for \textbf{valence} ranked by the within-subject paired
- test (per-subject median split, $n=32$ subjects). $\epsilon^2$ is the trial-level
- Kruskal--Wallis effect size with 95\% bootstrap CI (resampling subjects); $d_z$ is
- the within-subject paired effect size; $p_{\text{subj}}$ is the Wilcoxon
- signed-rank $p$-value across subjects (FDR-corrected). The last column shows the
- window-level Kruskal--Wallis $p$, illustrating the inflation from pseudo-replication
- (small $\epsilon^2$ yet $p<10^{-18}$).\label{tab:effect_valence}}
- \centering
- \footnotesize
- \begin{tabular}{lccccc}
- \toprule
- \textbf{Feature} & $\epsilon^2$ [95\% CI] & $d_z$ & $p_{\text{subj}}$ & FDR & $p_{\text{KW,window}}$ \\
- \midrule
- """ + eff_rows("valence") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- t = t.replace("valence", "arousal").replace("\\textbf{arousal}", "\\textbf{arousal}").replace("tab:effect_arousal", "tab:effect_arousal")
- # rebuild arousal table cleanly
- t = r"""\begin{table}[H]
- \caption{Top qEEG features for \textbf{arousal} ranked by the within-subject paired
- test (per-subject median split, $n=32$). Columns as in
- Table~\ref{tab:effect_valence}. After FDR correction no arousal feature reaches
- significance, although several (theta/alpha ratio, permutation entropy, Hjorth)
- show the expected direction.\label{tab:effect_arousal}}
- \centering
- \footnotesize
- \begin{tabular}{lccccc}
- \toprule
- \textbf{Feature} & $\epsilon^2$ [95\% CI] & $d_z$ & $p_{\text{subj}}$ & FDR & $p_{\text{KW,window}}$ \\
- \midrule
- """ + eff_rows("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: stability + importance ----
- si = load("stability_importance.json")
- def stab_block(tgt):
- s10 = si[tgt]["stability"]["K10"]; s20 = si[tgt]["stability"]["K20"]
- return (f"{tgt.capitalize()} & {s10['kuncheva']:.3f} [{s10['kuncheva_ci'][0]:.2f}, {s10['kuncheva_ci'][1]:.2f}] & "
- f"{s10['mean_jaccard']:.3f} & {s20['kuncheva']:.3f} & {s20['mean_jaccard']:.3f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Feature-selection stability across the 10 subject-independent folds,
- measured by the Kuncheva consistency index (chance-corrected) and mean pairwise
- Jaccard overlap of the top-$K$ XGBoost-gain features. $K$ was fixed a priori.
- Stability is modest, and lower for arousal.\label{tab:stability}}
- \centering
- \begin{tabular}{lcccc}
- \toprule
- & \multicolumn{2}{c}{\textbf{Top-10}} & \multicolumn{2}{c}{\textbf{Top-20}} \\
- \cmidrule(lr){2-3}\cmidrule(lr){4-5}
- \textbf{Target} & Kuncheva [95\% CI] & Jaccard & Kuncheva & Jaccard \\
- \midrule
- """ + stab_block("valence") + "\n" + stab_block("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- def imp_block(tgt):
- a = si[tgt]["importance"]["agreement"]
- return (f"{tgt.capitalize()} & {a['spearman_gain_shap']:.2f} & {a['spearman_gain_perm']:.2f} & "
- f"{a['spearman_shap_perm']:.2f} & {a['jaccard_top10_gain_shap']:.2f} & {a['jaccard_top10_gain_perm']:.2f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Agreement between feature-importance methods (Spearman rank correlation
- and top-10 Jaccard overlap). XGBoost gain correlates only moderately with SHAP and
- weakly with permutation importance, confirming that gain alone is an unreliable
- indicator of predictor relevance.\label{tab:importance}}
- \centering
- \begin{tabular}{lccccc}
- \toprule
- \textbf{Target} & $\rho_{\text{gain,SHAP}}$ & $\rho_{\text{gain,perm}}$ & $\rho_{\text{SHAP,perm}}$ & $J_{10}^{\text{gain,SHAP}}$ & $J_{10}^{\text{gain,perm}}$ \\
- \midrule
- """ + imp_block("valence") + "\n" + imp_block("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: cross-dataset ----
- cd = load("crossdataset.json")
- def cdrow(tgt, model):
- r = cd[tgt]["subject_median"]
- return (f"{tgt.capitalize()} ({model}) & {r['within_deap_'+model]['auc']['value']:.3f} & "
- f"{r['within_dreamer_'+model]['auc']['value']:.3f} & "
- f"{r['deap2dreamer_'+model]['auc']['value']:.3f} & {r['dreamer2deap_'+model]['auc']['value']:.3f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Cross-dataset robustness on the four EEG channels common to DEAP and
- DREAMER (AF3, AF4, F3, F4) with an identical feature definition and per-subject
- median labels. Within-dataset values are subject-independent (GroupKFold for DEAP,
- LOSO for DREAMER); transfer columns train on one dataset and test on the other.
- Cross-dataset transfer is close to chance.\label{tab:crossdataset}}
- \centering
- \begin{tabular}{lcccc}
- \toprule
- \textbf{Target (model)} & Within DEAP & Within DREAMER & DEAP$\to$DREAMER & DREAMER$\to$DEAP \\
- \midrule
- """ + cdrow("valence", "logreg") + "\n" + cdrow("valence", "xgb") + "\n" + \
- cdrow("arousal", "logreg") + "\n" + cdrow("arousal", "xgb") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: feature-extraction hyperparameters ----
- t = r"""\begin{table}[H]
- \caption{Feature-extraction hyperparameters (identical for DEAP and DREAMER).\label{tab:hyperparams}}
- \centering
- \begin{tabular}{ll}
- \toprule
- \textbf{Parameter} & \textbf{Value} \\
- \midrule
- Bandpass filter & 0.5--40 Hz, 4th-order Butterworth (zero-phase) \\
- Epoch length / overlap & 4 s / 2 s (50\%) \\
- Welch PSD & Hann window, nperseg $=$ epoch, 50\% segment overlap \\
- Frequency bands & $\delta$ 1--4, $\theta$ 4--8, $\alpha$ 8--13, $\beta$ 13--30, $\gamma$ 30--45 Hz \\
- Relative power denominator & total power 1--45 Hz \\
- Permutation entropy & orders $m\in\{3,5,7\}$, delay $\tau=1$, normalized \\
- Sample entropy & $m=2$, tolerance $r=0.2\cdot\mathrm{SD}$ \\
- Lempel--Ziv complexity & median-threshold binarization, normalized \\
- Higuchi fractal dimension & $k_{\max}=10$ \\
- Hjorth parameters & activity, mobility, complexity \\
- Frontal alpha asymmetry & $\ln P_\alpha^{\text{right}}-\ln P_\alpha^{\text{left}}$, pairs (F4,F3), (AF4,AF3) \\
- PLV / coherence (32-ch only) & per band, Hilbert phase, selected pairs \\
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # ---- Table: subject fingerprinting + variance decomposition ----
- try:
- P = load("personalization.json")
- def fv_row(d):
- f = P[d]["valence"]["fingerprint"]; v = P[d]["valence"]["variance"]
- return (f"{d} & {f['subject_id_acc']:.3f} & {f['subject_id_chance']:.3f} & "
- f"{f['emotion_loso_auc']:.3f} & {v['eta2_subject_median']:.3f} & "
- f"{v['eta2_emotion_median']:.4f} & {v['ratio_subject_over_emotion']:.0f}$\\times$ \\\\")
- t = r"""\begin{table}[H]
- \caption{The qEEG features encode subject identity far more than emotional state.
- Subject-identity decoding (random forest, 5-fold) versus cross-subject emotion
- decoding (LOSO), and the median per-feature variance explained by subject vs.\ by
- emotion class.\label{tab:fingerprint}}
- \centering
- \begin{tabular}{lcccccc}
- \toprule
- \textbf{Dataset} & \makecell{Subject-ID\\acc.} & \makecell{ID\\chance} & \makecell{Emotion\\AUC (LOSO)} & $\eta^2_{\text{subj}}$ & $\eta^2_{\text{emo}}$ & ratio \\
- \midrule
- """ + fv_row("DEAP") + "\n" + fv_row("DREAMER") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- # reliability vs idiosyncrasy
- def rel_row(d, t_):
- r = P[d][t_]["reliability"]
- return (f"{d} / {t_} & {r['within_subject_reliability']:.3f} & "
- f"{r['between_subject_similarity']:.3f} & {r['mean_signflip_rate']:.2f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Emotion effects are more consistent within a person than across people.
- Within-subject split-half reliability and between-subject similarity of the
- per-subject emotion-effect vector (cosine), and the mean per-feature sign-flip rate
- across subjects (0 = all agree, 0.5 = random).\label{tab:reliability}}
- \centering
- \begin{tabular}{lccc}
- \toprule
- \textbf{Dataset / target} & Within-subj.\ reliability & Between-subj.\ similarity & Sign-flip rate \\
- \midrule
- """ + "\n".join(rel_row(d, t_) for d in ["DEAP", "DREAMER"] for t_ in ["valence", "arousal"]) + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("personalization tables skipped:", e)
- # ---- Table: cross-dataset biomarker replication ----
- try:
- br = load("biomarker_replication.json")
- def br_row(t_):
- x = br["targets"][t_]
- return (f"{t_.capitalize()} & {x['deap_sig_count']}/{br['n_common_features']} & "
- f"{x['dreamer_sig_count']}/{br['n_common_features']} & {x['n_replicated']} & "
- f"{x['overall_sign_agreement']:.2f} \\\\")
- t = r"""\begin{table}[H]
- \caption{Cross-dataset within-subject biomarker replication on the channels common
- to DEAP and DREAMER (AF3, AF4, F3, F4). Features FDR-significant within-subject in
- each dataset, the number replicating in both with matching sign, and the overall
- effect-direction agreement between datasets (0.5 = chance).\label{tab:replication}}
- \centering
- \begin{tabular}{lcccc}
- \toprule
- \textbf{Target} & DEAP FDR-sig & DREAMER FDR-sig & Replicated (both) & Sign agreement \\
- \midrule
- """ + br_row("valence") + "\n" + br_row("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("replication table skipped:", e)
- # ---- Table: personalized vs universal vs subject-dependent ----
- try:
- hl = load("headline.json"); sd = load("subject_dependent.json")
- def pu_row(t_):
- uni = max(hl[f"{t_}_subject_median"]["models"][f"{m}_loso"]["auc"]["value"]
- for m in ["rf", "rbfsvm", "xgb", "logreg"])
- sdv = max(sd[t_]["subject_median"][m]["auc"] for m in ["logreg", "rf", "xgb"])
- return f"{t_.capitalize()} & {uni:.3f} & {sdv:.3f} & {sdv-uni:+.3f} \\\\"
- t = r"""\begin{table}[H]
- \caption{Personalized vs.\ universal modelling (per-subject median labels, best model
- per cell). The universal model is trained on all other subjects (LOSO); the
- personalized model is trained on the subject's own data (leak-free within-subject
- CV). A model built from a subject's own data exceeds one built from 31 other
- subjects.\label{tab:personalized}}
- \centering
- \begin{tabular}{lccc}
- \toprule
- \textbf{Target} & Universal (LOSO) AUC & Personalized (within-subj.) AUC & $\Delta$ \\
- \midrule
- """ + pu_row("valence") + "\n" + pu_row("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("personalized table skipped:", e)
- # ---- Table: calibration learning curve (optional) ----
- try:
- cal = load("calibration.json")
- ks = sorted(int(k) for k in cal["valence"]["rf"])
- head = " & ".join(f"$k$={k}" for k in ks)
- def cal_row(t_, m):
- return f"{t_.capitalize()} ({m}) & " + " & ".join(f"{cal[t_][m][str(k)]['auc']:.3f}" for k in ks) + r" \\"
- t = r"""\begin{table}[H]
- \caption{Subject-adaptive calibration: LOSO ROC-AUC as $k$ labelled trials from the
- target subject are added to the training set. A few target-subject trials yield only
- marginal gains, indicating that meaningful personalization requires substantial
- per-subject data rather than light calibration.\label{tab:calibration}}
- \centering
- \begin{tabular}{l""" + "c" * len(ks) + r"""}
- \toprule
- \textbf{Target (model)} & """ + head + r""" \\
- \midrule
- """ + "\n".join(cal_row(t_, m) for t_ in ["valence", "arousal"] for m in ["logreg", "rf"]) + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("calibration table skipped (job may still be running):", e)
- # ---- Table: identity-confound (centerpiece) ----
- try:
- cf = load("confound_full.json")
- def cf_row(ds, t_):
- c = cf[ds]["fixed"][t_]
- return (f"{ds} & {t_.capitalize()} & {c['A_pooled_eeg']:.3f} & {c['B_identity_only']:.3f} & "
- f"{c['B_recovers_pct']:.0f}\\% & {c['D_subject_independent']:.3f} \\\\")
- t = r"""\begin{table}[H]
- \caption{The reported performance is largely subject identity, not emotion (fixed
- threshold, window-pooled, XGBoost). \textbf{Pooled EEG} is the commonly used leaky
- protocol; \textbf{Identity-only} uses no EEG at all---it predicts each test window
- with its subject's training-set positive rate; \textbf{Subject-independent} respects
- subject boundaries. An identity-only predictor recovers most of the pooled AUC,
- which collapses to chance once subjects are separated.\label{tab:confound}}
- \centering
- \begin{tabular}{llcccc}
- \toprule
- \textbf{Dataset} & \textbf{Target} & Pooled EEG & Identity-only (no EEG) & \% recovered & Subject-indep.\ \\
- \midrule
- """ + "\n".join(cf_row(ds, t_) for ds in ["DEAP", "DREAMER"] for t_ in ["valence", "arousal"]) + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.insert(2, t) # place right after protocol + labeling tables
- except Exception as e:
- print("confound table skipped:", e)
- # ---- Table: best valid universal model (final optimization) ----
- try:
- fo = load("final_opt.json")
- def best_uni(t_):
- best = ("", -1)
- for ch in fo:
- for meth, r in fo[ch][t_].items():
- if r["auc_subjmean"] > best[1]:
- best = (f"{ch}/{meth}", r["auc_subjmean"])
- return best
- bv, bvv = best_uni("valence"); ba, bav = best_uni("arousal")
- t = r"""\begin{table}[H]
- \caption{Best subject-independent (universal) performance after exhaustive
- optimization---nested-tuned XGBoost/RF, top-$k$ feature selection, soft-voting
- ensembles, and Riemannian tangent-space classification with Euclidean alignment.
- No configuration exceeds a modest ceiling, indicating the limit is the
- subject-specificity of the signal rather than the model.\label{tab:bestuniversal}}
- \centering
- \begin{tabular}{lcc}
- \toprule
- \textbf{Target} & Best configuration & Best universal AUC \\
- \midrule
- Valence & """ + bv.replace("_", r"\_") + f" & {bvv:.3f}" + r""" \\
- Arousal & """ + ba.replace("_", r"\_") + f" & {bav:.3f}" + r""" \\
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("best-universal table skipped:", e)
- # ---- Table: multimodal fusion ----
- try:
- mm = load("multimodal.json")
- def mm_row(t_):
- e = mm[t_]["eeg"]["rf_loso"]["auc"]; p = mm[t_]["peripheral"]["rf_loso"]["auc"]; f = mm[t_]["fused"]["rf_loso"]["auc"]
- return f"{t_.capitalize()} & {e['value']:.3f} & {p['value']:.3f} & {f['value']:.3f} \\\\"
- t = r"""\begin{table}[H]
- \caption{Multimodal fusion does not rescue subject-independent performance (LOSO,
- random forest, per-subject median labels). Interpretable peripheral autonomic
- features (HRV, EDA/GSR, respiration, EMG, temperature) are themselves weak and
- subject-specific; fusing them with EEG yields no reliable gain, extending the
- ``personalized, not universal'' conclusion across modalities.\label{tab:multimodal}}
- \centering
- \begin{tabular}{lccc}
- \toprule
- \textbf{Target} & EEG only & Peripheral only & Fused \\
- \midrule
- """ + mm_row("valence") + "\n" + mm_row("arousal") + r"""
- \bottomrule
- \end{tabular}
- \end{table}"""
- blocks.append(t)
- except Exception as e:
- print("multimodal table skipped:", e)
- (OUT / "tables.tex").write_text("\n\n".join(blocks) + "\n")
- print(f"Wrote {len(blocks)} tables -> {OUT/'tables.tex'}")
gen_latex_tables.py at commit e603254, no license · at the source
Overview
Abstract
Quantitative EEG features such as frontal alpha asymmetry, spectral ratios and signal-complexity measures are often presented as interpretable biomarkers of emotion. Such claims require the markers to generalize across individuals, yet common evaluation protocols allow overlapping epochs and recordings from the same participants to appear in both training and test sets. We re-evaluated qEEG-based valence and arousal recognition on DEAP and DREAMER under trial-grouped, participant-independent,
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 25 matches between paragraphs and lines of code.
ema-pandilova/qeeg-emotion-pipeline
e60325483181fce04ada3aa5ace4fdc27ce99267, 14 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
81 files
- experiments/
correct_all_label_tables , Python, 47 lines.py - experiments/
diff_results.py , Python, 47 lines - experiments/
evallib.py , Python, 221 lines - experiments/
extract_deap.py , Python, 68 lines - experiments/
extract_dreamer.py , Python, 63 lines - experiments/
extract_peripheral_deap. , Python, 96 linespy - experiments/
featurelib.py , Python, 155 lines, 2 matches - experiments/
fix_labels_and_check.py , Python, 51 lines - experiments/
gen_confound_fig.py , Python, 33 lines - experiments/
gen_figures.py , Python, 197 lines - experiments/
gen_latex_tables.py , Python, 520 lines, 13 matches - experiments/
gen_personalization_figu , Python, 124 linesres.py - experiments/
gen_tables.py , Python, 87 lines - experiments/
make_leakage_figure.py , Python, 129 lines - experiments/
run_ablations.py , Python, 72 lines - experiments/
run_all_corrected.sh , Shell, 35 lines - experiments/
run_biomarker_replicatio , Python, 115 linesn.py - experiments/
run_calibration.py , Python, 84 lines - experiments/
run_confound.py , Python, 106 lines - experiments/
run_confound_ext.py , Python, 93 lines - experiments/
run_core_eval.py , Python, 127 lines - experiments/
run_crossdataset.py , Python, 120 lines - experiments/
run_final_opt.py , Python, 109 lines, 1 match - experiments/
run_headline.py , Python, 58 lines - experiments/
run_multimodal.py , Python, 57 lines, 1 match - experiments/
run_parallel_corrected.s , Shell, 28 linesh - experiments/
run_personalization.py , Python, 196 lines, 1 match - experiments/
run_revision_stats.py , Python, 169 lines, 1 match - experiments/
run_riemann.py , Python, 152 lines, 2 matches - experiments/
run_stability_importance , Python, 148 lines, 2 matches.py - experiments/
run_stats.py , Python, 174 lines, 2 matches - experiments/
run_subject_dependent.py , Python, 86 lines - experiments/
validate_tex.py , Python, 22 lines - scripts/
compute_feature_stabilit , Python, 310 linesy.py - scripts/
create_feature_table.py , Python, 175 lines - scripts/
create_feature_table_sim , Python, 79 linesple.py - scripts/
create_final_feature_tab , Python, 39 linesle.py - scripts/
eval_baselines.py , Python, 448 lines - scripts/
eval_cnn_topomap.py , Python, 404 lines - scripts/
eval_features_stats.py , Python, 114 lines - scripts/
eval_improved_pipeline.p , Python, 363 linesy - scripts/
eval_subject_dependent.p , Python, 158 linesy - scripts/
eval_xgb_enhanced.py , Python, 536 lines - scripts/
eval_xgb_groupcv.py , Python, 135 lines - scripts/
eval_xgb_improved.py , Python, 450 lines - scripts/
eval_xgb_plots.py , Python, 144 lines - scripts/
explore_data.py , Python, 140 lines - scripts/
extract_features.py , Python, 89 lines - scripts/
extract_features_enhance , Python, 137 linesd.py - scripts/
generate_comparison_figu , Python, 149 linesre.py - scripts/
generate_latex_tables.py , Python, 281 lines - scripts/
generate_thesis_figures. , Python, 482 linespy - scripts/
make_results_plots.py , Python, 281 lines - scripts/
plot_difference_topomaps , Python, 267 lines.py - scripts/
plot_xgb_importance.py , Python, 91 lines - scripts/
run_enhanced_pipeline.py , Python, 144 lines - scripts/
train_lstm_fuzzy.py , Python, 580 lines - scripts/
train_xgb.py , Python, 93 lines - scripts/
train_xgb_enhanced.py , Python, 315 lines - scripts/
train_xgb_loso.py , Python, 108 lines - scripts/
train_xgb_loso_enhanced. , Python, 381 linespy - scripts/
update_lstm_fuzzy_mae.py , Python, 113 lines - src/
__init__.py , Python, 22 lines - src/
config.py , Python, 62 lines - src/
datasets/ , Python, 3 lines__init__.py - src/
datasets/ , Python, 134 linesdeap.py - src/
features/ , Python, 3 lines__init__.py - src/
features/ , Python, 180 linesconnectivity.py - src/
features/ , Python, 125 lineslinear.py - src/
features/ , Python, 95 linesnonlinear.py - src/
features/ , Python, 321 linestemporal.py - src/
labels.py , Python, 80 lines - src/
models/ , Python, 16 lines__init__.py - src/
models/ , Python, 131 lineslstm_fuzzy_regressor.py - src/
models/ , Python, 38 linesxgb_classifier.py - src/
models/ , Python, 284 linesxgb_classifier_enhanced. py - src/
paths.py , Python, 15 lines - src/
stats.py , Python, 84 lines - src/
utils.py , Python, 16 lines - src/
viz.py , Python, 146 lines - README.md, Text, 123 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 80 scripts, each with its path and the digest of its content;
- 25 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:546113, at Zenodo; found in “Data Availability Statement”
Data Availability Statement
DEAP is available from https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 9 keywords, 8 MeSH terms, 4 funders, 42 references.
Cite
This paper
Pandilova, E., Stojmenski, A., Chorbev, I., Petrov, M., Kitanovski, I., & Trajanov, D. (2026). Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER. Sensors (Basel, Switzerland), 26(17), 5327. https://
BibTeX
@article{pandilova2026su
author = {Pandilova, Ema and Stojmenski, Aleksandar and Chorbev, Ivan and Petrov, Marko and Kitanovski, Ivan and Trajanov, Dimitar},
title = {{Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER}},
journal = {Sensors (Basel, Switzerland)},
year = {2026},
month = aug,
volume = {26},
number = {17},
pages = {5327},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1424-8220},
doi = {10.3390/
url = {https://
pmid = {42739948},
pmcid = {PMC13567904}
}
RIS
TY - JOUR
AU - Pandilova, Ema
AU - Stojmenski, Aleksandar
AU - Chorbev, Ivan
AU - Petrov, Marko
AU - Kitanovski, Ivan
AU - Trajanov, Dimitar
TI - Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER
T2 - Sensors (Basel, Switzerland)
J2 - Sensors (Basel)
PY - 2026
DA - 2026/
VL - 26
IS - 17
SP - 5327
SN - 1424-8220
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER",
"container-title": "Sensors (Basel, Switzerland)",
"author": [
{
"family": "Pandilova",
"given": "Ema"
},
{
"family": "Stojmenski",
"given": "Aleksandar"
},
{
"family": "Chorbev",
"given": "Ivan"
},
{
"family": "Petrov",
"given": "Marko"
},
{
"family": "Kitanovski",
"given": "Ivan"
},
{
"family": "Trajanov",
"given": "Dimitar"
}
],
"container-title-short":
"volume": "26",
"issue": "17",
"page": "5327",
"DOI": "10.3390/
"PMID": "42739948",
"PMCID": "PMC13567904",
"ISSN": "1424-8220",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
22
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: imbalanced-learn, SHAP, XGBoost, 9 other tools
- [2] doi:10.3389/fpsyg.2026.1774068 [code]
- Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.Journal: Frontiers in psychologyIn common: pyRiemann, XGBoost, Keras, 8 other tools, EEG, cognitive
- [3] doi:10.3390/brainsci16070716
- RGB-Style Input Representations for EEG: Evaluating Spatial Concatenation Versus Band-Wise Stacking in Deep Emotion Recognition.Journal: Brain sciencesIn common: Zenodo 546113, EEG, cognitive, 6 references
- [4] doi:10.1038/s41598-026-52330-z [code]
- SHAP analysis of an improved EEG-based mental workload classification framework: utilizing data augmentation and explainable AI.Journal: Scientific reportsIn common: imbalanced-learn, SHAP, Keras, 7 other tools, EEG
- [5] doi:10.1371/journal.pone.0347671 [code]
- RMETNet: A cross-subject motor imagery EEG signal classification model based on TSLANet and riemannian geometry features.Journal: PloS oneIn common: pyRiemann, imbalanced-learn, TensorFlow, 7 other tools, EEG
- [6] doi:10.1371/journal.pcbi.1014615 [code]
- Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.Journal: PLoS computational biologyIn common: SHAP, XGBoost, Keras, 8 other tools
- [7] doi:10.3389/fnins.2026.1810609
- A dual-branch network with brain region-constrained attention for EEG emotion recognition.Journal: Frontiers in neuroscienceIn common: Zenodo 546113, EEG, cognitive, 5 references
- [8] doi:10.1038/s41598-026-55163-y [code]
- Autism spectrum disorder identification using machine learning models on MRI data.Journal: Scientific reportsIn common: imbalanced-learn, XGBoost, Keras, 7 other tools
- [9] doi:10.1186/s13059-026-04125-8 [code]
- MLMarker: a machine learning framework for tissue inference and biomarker discovery.Journal: Genome biologyIn common: imbalanced-learn, SHAP, XGBoost, 7 other tools
- [10] doi:10.1038/s41598-026-56688-y [code]
- On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.Journal: Scientific reportsIn common: imbalanced-learn, SHAP, XGBoost, 7 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 80 scripts, and 25 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e8e7b54354a4f2b2…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
