OSCR

Structured multi-domain EEG descriptors with phase-based connectivity for lie and truth detection.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Method details › Feature extraction framework: structured multi-domain eeg descriptors with phase-based connectivity ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 126–161 · score 0.85 · power spectral density, 12–30 Hz, 8–12 Hz, 1–4 Hz, 4–8 Hz, Welch
  2. [2] § Method details ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 1–46 · score 0.75 · overlapping sliding window, multi domain descriptor, meta correlation, pipeline, segmentation, fold
  3. [3] § Method details › Algorithmic workflow of structured multi-domain eeg descriptors with phase-based connectivity ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 1–46 · score 0.65 · overlapping sliding window, Structured Multi Domain, EEG Descriptor, segmentation, preprocessing
  4. [4] § Method details › Algorithmic workflow of structured multi-domain eeg descriptors with phase-based connectivity ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 126–161 · score 0.64 · power spectral density, kurtosis, bands, Welch, Hjorth, Fractal
  5. [5] § Method details › Evaluation and testing scenario ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 462–546 · score 0.62 · F1 score, ROC, ACC, AUC, recall, predictions
  6. [6] § Method validation › Classification performance ↔ lie-truth-eeg-descriptor/run_liewaves_descriptor_pipeline.py, lines 387–421 · score 0.50 · Random Forest, RBF, kernel, configuration, XGB, RF

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 808 lines · 30 KB · no license · 6 matches

  1. """
  2. Structured Multi-Domain EEG Descriptor Pipeline for Lie and Truth Classification
  3. This script implements the preprocessing, window segmentation, feature extraction,
  4. classification, and evaluation pipeline used in the MethodsX manuscript.
  5. The implementation includes:
  6. 1. deterministic EEG preprocessing,
  7. 2. overlapping sliding-window segmentation,
  8. 3. structured multi-domain descriptor extraction,
  9. 4. fold-wise normalization,
  10. 5. window-level and grouped validation,
  11. 6. meta-correlation ablation analysis.
  12. The LieWaves dataset is not redistributed in this repository.
  13. Users should download the dataset from the official Mendeley Data repository.
  14. """
  15. import os, re, time, json
  16. import numpy as np
  17. import pandas as pd
  18. from itertools import combinations
  19. from collections import defaultdict
  20. from scipy.signal import welch, butter, filtfilt, iirnotch, hilbert
  21. from scipy.stats import skew, kurtosis, ttest_rel
  22. from sklearn.model_selection import StratifiedKFold, GroupKFold, LeaveOneGroupOut
  23. from sklearn.preprocessing import StandardScaler
  24. from sklearn.metrics import (
  25. accuracy_score, recall_score, f1_score, roc_auc_score,
  26. confusion_matrix, ConfusionMatrixDisplay
  27. )
  28. from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
  29. from sklearn.ensemble import RandomForestClassifier
  30. from sklearn.svm import SVC
  31. try:
  32. from xgboost import XGBClassifier
  33. HAS_XGB = True
  34. except Exception:
  35. HAS_XGB = False
  36. import matplotlib.pyplot as plt
  37. import warnings
  38. warnings.filterwarnings("ignore")
  39. # ========
  40. # CONFIG
  41. # ========
  42. DATA_DIR = "./data/LieWaves"
  43. RAW_DIR = os.path.join(DATA_DIR, "Raw")
  44. SUBJECT_STIMULI_XLSX = os.path.join(DATA_DIR, "Subject_Stimuli.xlsx")
  45. CHAN_ORDER = ["EEG.AF3", "EEG.T7", "EEG.Pz", "EEG.T8", "EEG.AF4"]
  46. SF = 128
  47. WINDOW_SIZE = 384
  48. STEP_SIZE = 16
  49. N_SPLITS = 5
  50. RANDOM_STATE = 42
  51. # Validation settings
  52. RUN_WINDOW_LEVEL_SKF = True
  53. RUN_SUBJECT_GROUPK = False
  54. RUN_SESSION_GROUPK = False
  55. RUN_LOSO = False
  56. RUN_NON_OVERLAP = False
  57. RUN_META_ABLATION = True
  58. # SVM is included in the main window-level evaluation.
  59. # It is disabled by default for optional grouped validation to reduce runtime.
  60. INCLUDE_SVM_IN_STRICT_VALIDATION = False
  61. INCLUDE_SVM_IN_WINDOW_LEVEL = True
  62. # Output
  63. OUT_DIR = "./results"
  64. FIG_DIR = os.path.join(OUT_DIR, "figures")
  65. TAB_DIR = os.path.join(OUT_DIR, "tables")
  66. os.makedirs(FIG_DIR, exist_ok=True)
  67. os.makedirs(TAB_DIR, exist_ok=True)
  68. # LDA plot options
  69. LDA_JITTER = 0.02
  70. LDA_POINT_SIZE = 10
  71. LDA_ALPHA = 0.75
  72. LDA_MAX_POINTS = 20000
  73. # ============================================================
  74. # Filtering functions
  75. # ============================================================
  76. def bandpass(x):
  77. b, a = butter(4, [0.5, 45], btype='bandpass', fs=SF)
  78. return filtfilt(b, a, x)
  79. def notch(x):
  80. b, a = iirnotch(50, 30, fs=SF)
  81. return filtfilt(b, a, x)
  82. def preprocess(arr):
  83. out = np.zeros_like(arr, dtype=np.float32)
  84. for ch in range(arr.shape[1]):
  85. out[:, ch] = bandpass(notch(arr[:, ch]))
  86. return out
  87. # ============================================================
  88. # Feature extraction functions
  89. # ============================================================
  90. def hjorth(x):
  91. dx = np.diff(x); ddx = np.diff(dx)
  92. v0, v1, v2 = np.var(x), np.var(dx), np.var(ddx)
  93. act = v0
  94. mob = np.sqrt(v1/v0) if v0 > 0 else 0
  95. comp = np.sqrt(v2/v1)/mob if (v1 > 0 and mob > 0) else 0
  96. return act, mob, comp
  97. def wpli(x, y):
  98. ax, ay = hilbert(x), hilbert(y)
  99. im = np.imag(ax * np.conj(ay))
  100. return float(np.abs(np.mean(im)) / (np.mean(np.abs(im)) + 1e-6))
  101. def fast_corr(a, b):
  102. am = a - a.mean()
  103. bm = b - b.mean()
  104. return float(np.dot(am, bm) / (np.std(a) * np.std(b) * len(a) + 1e-6))
  105. def extract_descriptor(window):
  106. C = window.shape[0]
  107. temporal, fractal, spectral = [], [], []
  108. for ch in range(C):
  109. x = (window[ch] - np.mean(window[ch])) / (np.std(window[ch]) + 1e-6)
  110. feats = [np.mean(x), np.std(x), skew(x), kurtosis(x)]
  111. feats += list(hjorth(x))
  112. temporal += feats
  113. N = len(x); L = np.sum(np.abs(np.diff(x)))
  114. fractal.append(float(np.log(N) / (np.log(N) + np.log(L/N) + 1e-6)))
  115. f, psd = welch(x, SF, nperseg=256)
  116. bands = [(1,4),(4,8),(8,12),(12,30)]
  117. abs_bp = np.array([np.sum(psd[(f>=b0) & (f<=b1)]) for b0,b1 in bands], dtype=float)
  118. rel_bp = abs_bp / (np.sum(abs_bp) + 1e-6)
  119. sent = -np.sum((psd/np.sum(psd)) * np.log(psd/np.sum(psd) + 1e-12))
  120. spectral += abs_bp.tolist() + rel_bp.tolist() + [float(sent)]
  121. spatial = []
  122. for i, j in combinations(range(C), 2):
  123. spatial.append(float(np.mean(window[i]) - np.mean(window[j])))
  124. spatial.append(wpli(window[i], window[j]))
  125. fTS = np.array(temporal + spatial, dtype=np.float32)
  126. fCorr = np.array([fast_corr(fTS, np.log1p(np.abs(fTS)))], dtype=np.float32)
  127. fF = np.array(fractal, dtype=np.float32)
  128. fSP = np.array(spectral, dtype=np.float32)
  129. # Feature order:
  130. # temporal + spatial + meta-correlation + fractal + spectral
  131. # 35 + 20 + 1 + 5 + 45 = 106 features.
  132. return np.concatenate([fTS, fCorr, fF, fSP]).astype(np.float32)
  133. # ============================================================
  134. # Feature group definitions
  135. # ============================================================
  136. def group_slices(n_channels=5):
  137. n_temporal = n_channels * 7
  138. n_pairs = n_channels * (n_channels - 1) // 2
  139. n_spatial = n_pairs * 2
  140. n_fTS = n_temporal + n_spatial
  141. n_fCorr = 1
  142. n_fF = n_channels
  143. n_fSP = n_channels * 9
  144. s_fTS = slice(0, n_fTS)
  145. s_fCorr = slice(s_fTS.stop, s_fTS.stop + n_fCorr)
  146. s_fF = slice(s_fCorr.stop, s_fCorr.stop + n_fF)
  147. s_fSP = slice(s_fF.stop, s_fF.stop + n_fSP)
  148. s_full = slice(0, s_fSP.stop)
  149. return {"fTS": s_fTS, "fCorr": s_fCorr, "fF": s_fF, "fSP": s_fSP, "full": s_full}
  150. SLICES = group_slices(n_channels=len(CHAN_ORDER))
  151. # ============================================================
  152. # FEATURE NAMES — ORIGINAL LOGIC KEPT
  153. # ============================================================
  154. def build_feature_names():
  155. names = []
  156. tnames = ["mean","std","skew","kurt","hjorth_act","hjorth_mob","hjorth_comp"]
  157. for ch in CHAN_ORDER:
  158. for tn in tnames:
  159. names.append(f"temporal_{ch}_{tn}")
  160. pairs = list(combinations(range(len(CHAN_ORDER)), 2))
  161. for (i,j) in pairs:
  162. names.append(f"spatial_asym_{CHAN_ORDER[i]}-{CHAN_ORDER[j]}")
  163. names.append(f"spatial_wpli_{CHAN_ORDER[i]}-{CHAN_ORDER[j]}")
  164. names.append("fCorr_corr(fTS,log1p(|fTS|))")
  165. for ch in CHAN_ORDER:
  166. names.append(f"fractal_{ch}")
  167. bnames = ["delta","theta","alpha","beta"]
  168. for ch in CHAN_ORDER:
  169. for bn in bnames:
  170. names.append(f"spectral_absbp_{ch}_{bn}")
  171. for bn in bnames:
  172. names.append(f"spectral_relbp_{ch}_{bn}")
  173. names.append(f"spectral_entropy_{ch}")
  174. return names
  175. FEAT_NAMES = build_feature_names()
  176. # ============================================================
  177. # Data loading and metadata construction
  178. # ============================================================
  179. def load_label_map():
  180. df = pd.read_excel(SUBJECT_STIMULI_XLSX)
  181. df.columns = [str(c).strip().upper() for c in df.columns]
  182. df = df[['SUBJECT','SESSION','LIE/TRUTH']]
  183. df['SUBJECT'] = df['SUBJECT'].astype(str).str.replace('S','', regex=False).astype(int)
  184. df['SESSION'] = df['SESSION'].astype(str).str.replace('S','', regex=False).astype(int)
  185. df['KEY'] = list(zip(df['SUBJECT'], df['SESSION']))
  186. label_map = dict(zip(df['KEY'], df['LIE/TRUTH']))
  187. print("Total labels loaded:", len(label_map))
  188. return label_map
  189. def parse_filename(f):
  190. m = re.match(r"S(\d+)S(\d+)\.csv", f)
  191. return (int(m.group(1)), int(m.group(2))) if m else None
  192. def sliding_window(data, step_size=STEP_SIZE):
  193. return [data[:, i:i+WINDOW_SIZE] for i in range(0, data.shape[1]-WINDOW_SIZE+1, step_size)]
  194. def load_data(step_size=STEP_SIZE):
  195. label_map = load_label_map()
  196. X, y = [], []
  197. subjects, sessions, file_groups = [], [], []
  198. filenames, window_indices = [], []
  199. for f in sorted(os.listdir(RAW_DIR)):
  200. key = parse_filename(f)
  201. if key is None or key not in label_map:
  202. continue
  203. subject_id, session_id = key
  204. df = pd.read_csv(os.path.join(RAW_DIR, f))[CHAN_ORDER]
  205. arr = preprocess(df.values).T # (C,T)
  206. windows = sliding_window(arr, step_size=step_size)
  207. for wi, w in enumerate(windows):
  208. X.append(extract_descriptor(w))
  209. y.append(int(label_map[key]))
  210. subjects.append(subject_id)
  211. sessions.append(session_id)
  212. file_groups.append(f"S{subject_id:02d}_Sess{session_id:02d}")
  213. filenames.append(f)
  214. window_indices.append(wi)
  215. X = np.array(X, dtype=np.float32)
  216. y = np.array(y, dtype=int)
  217. meta = pd.DataFrame({
  218. "subject": subjects,
  219. "session": sessions,
  220. "file_group": file_groups,
  221. "filename": filenames,
  222. "window_index": window_indices,
  223. "label": y
  224. })
  225. return X, y, meta
  226. # ============================================================
  227. # Plotting and result saving utilities
  228. # ============================================================
  229. def save_confusion_matrix(y_true, y_pred, title, outpath, labels=("Lie (0)", "Truth (1)")):
  230. cm = confusion_matrix(y_true, y_pred, labels=[0,1])
  231. cm_norm = cm.astype(float) / np.clip(cm.sum(axis=1, keepdims=True), 1, None)
  232. disp = ConfusionMatrixDisplay(confusion_matrix=cm_norm, display_labels=list(labels))
  233. disp.plot(cmap="Blues", colorbar=True)
  234. plt.title(title)
  235. plt.tight_layout()
  236. plt.savefig(outpath, dpi=300)
  237. plt.close()
  238. def paired_ttest_matrix(fold_metrics, metric_key, out_csv, out_png, title):
  239. model_names = list(fold_metrics.keys())
  240. n = len(model_names)
  241. pmat = np.ones((n, n), dtype=float)
  242. for i in range(n):
  243. for j in range(n):
  244. if i == j:
  245. pmat[i, j] = np.nan
  246. elif i < j:
  247. a = np.array(fold_metrics[model_names[i]][metric_key], dtype=float)
  248. b = np.array(fold_metrics[model_names[j]][metric_key], dtype=float)
  249. if len(a) == len(b) and len(a) > 1:
  250. p = ttest_rel(a, b).pvalue
  251. else:
  252. p = np.nan
  253. pmat[i, j] = p
  254. pmat[j, i] = p
  255. dfp = pd.DataFrame(pmat, index=model_names, columns=model_names)
  256. dfp.to_csv(out_csv, index=True)
  257. plt.figure()
  258. im = plt.imshow(pmat, aspect="auto")
  259. plt.colorbar(im)
  260. plt.xticks(range(n), model_names)
  261. plt.yticks(range(n), model_names)
  262. plt.title(title)
  263. for i in range(n):
  264. for j in range(n):
  265. txt = "-" if np.isnan(pmat[i, j]) else f"{pmat[i, j]:.3g}"
  266. plt.text(j, i, txt, ha="center", va="center")
  267. plt.tight_layout()
  268. plt.savefig(out_png, dpi=300)
  269. plt.close()
  270. def lda_plot_like_example(X, y, out_png,
  271. max_points=LDA_MAX_POINTS,
  272. jitter=LDA_JITTER,
  273. seed=RANDOM_STATE):
  274. rng = np.random.default_rng(seed)
  275. if (max_points is not None) and (X.shape[0] > max_points):
  276. idx = rng.choice(X.shape[0], size=max_points, replace=False)
  277. Xp = X[idx]
  278. yp = y[idx]
  279. else:
  280. Xp = X
  281. yp = y
  282. Xs = StandardScaler().fit_transform(Xp)
  283. lda = LinearDiscriminantAnalysis(n_components=1)
  284. z = lda.fit_transform(Xs, yp).ravel()
  285. y0 = (rng.random(np.sum(yp == 0)) - 0.5) * jitter
  286. y1 = (rng.random(np.sum(yp == 1)) - 0.5) * jitter
  287. plt.figure(figsize=(8, 4))
  288. plt.scatter(z[yp == 0], y0, s=LDA_POINT_SIZE, alpha=LDA_ALPHA, label="Lie (0)")
  289. plt.scatter(z[yp == 1], y1, s=LDA_POINT_SIZE, alpha=LDA_ALPHA, label="Truth (1)")
  290. plt.yticks([])
  291. plt.xlabel("LDA projection (LD1)")
  292. plt.title("LDA Separability View — Lie vs Truth")
  293. plt.legend()
  294. plt.tight_layout()
  295. plt.savefig(out_png, dpi=300)
  296. plt.close()
  297. def save_barh(values, names, title, out_png, topk=10):
  298. idx = np.argsort(values)[::-1][:topk]
  299. vals = values[idx][::-1]
  300. nms = [names[i] for i in idx][::-1]
  301. plt.figure(figsize=(8, 5))
  302. plt.barh(range(len(vals)), vals)
  303. plt.yticks(range(len(vals)), nms)
  304. plt.title(title)
  305. plt.tight_layout()
  306. plt.savefig(out_png, dpi=300)
  307. plt.close()
  308. def rf_importance_by_group(mean_imp, slices_dict, out_csv, out_png):
  309. rows = []
  310. for g in ["fTS", "fCorr", "fF", "fSP", "full"]:
  311. sl = slices_dict[g]
  312. rows.append((g, float(np.sum(mean_imp[sl]))))
  313. df = pd.DataFrame(rows, columns=["Group", "Total_importance"])
  314. df.to_csv(out_csv, index=False)
  315. plt.figure()
  316. plt.bar(df["Group"], df["Total_importance"])
  317. plt.title("RF Importance by Feature Group (Sum of Importances)")
  318. plt.tight_layout()
  319. plt.savefig(out_png, dpi=300)
  320. plt.close()
  321. # ============================================================
  322. # Classifier configuration
  323. # ============================================================
  324. def make_models(include_svm=True):
  325. models = {
  326. "RF": RandomForestClassifier(
  327. n_estimators=900,
  328. max_depth=None,
  329. min_samples_leaf=1,
  330. n_jobs=-1,
  331. random_state=RANDOM_STATE
  332. )
  333. }
  334. if HAS_XGB:
  335. models["XGB"] = XGBClassifier(
  336. n_estimators=1200,
  337. max_depth=8,
  338. learning_rate=0.015,
  339. subsample=0.85,
  340. colsample_bytree=0.85,
  341. reg_lambda=1,
  342. eval_metric="logloss",
  343. random_state=RANDOM_STATE
  344. )
  345. if include_svm:
  346. models["SVM"] = SVC(
  347. C=50,
  348. gamma="auto",
  349. kernel="rbf",
  350. probability=True,
  351. random_state=RANDOM_STATE
  352. )
  353. return models
  354. # ============================================================
  355. # FEATURE SETS FOR ABLATION
  356. # ============================================================
  357. def get_feature_sets(X):
  358. idx_full = np.arange(X.shape[1])
  359. idx_no_meta = np.setdiff1d(idx_full, np.arange(SLICES["fCorr"].start, SLICES["fCorr"].stop))
  360. feature_sets = {
  361. "full_106": idx_full,
  362. "without_meta_corr": idx_no_meta,
  363. }
  364. return feature_sets
  365. # ============================================================
  366. # OVERLAP / LEAKAGE DIAGNOSTICS
  367. # ============================================================
  368. def save_overlap_report(step_size, meta, out_json, out_csv):
  369. overlap_samples = WINDOW_SIZE - step_size
  370. overlap_percent = overlap_samples / WINDOW_SIZE * 100 if step_size < WINDOW_SIZE else 0.0
  371. report = {
  372. "window_size_samples": WINDOW_SIZE,
  373. "step_size_samples": step_size,
  374. "sampling_rate_hz": SF,
  375. "window_duration_seconds": WINDOW_SIZE / SF,
  376. "step_duration_seconds": step_size / SF,
  377. "overlap_samples": max(overlap_samples, 0),
  378. "overlap_percent_adjacent_windows": overlap_percent,
  379. "n_windows": int(len(meta)),
  380. "n_subjects": int(meta["subject"].nunique()),
  381. "n_file_groups": int(meta["file_group"].nunique()),
  382. "class_counts": {str(k): int(v) for k, v in meta["label"].value_counts().sort_index().items()}
  383. }
  384. with open(out_json, "w") as f:
  385. json.dump(report, f, indent=2)
  386. meta.groupby(["subject", "label"]).size().reset_index(name="n_windows").to_csv(out_csv, index=False)
  387. print("\n================ Overlap / Leakage Diagnostic ================")
  388. print(json.dumps(report, indent=2))
  389. print("==============================================================\n")
  390. return report
  391. # ============================================================
  392. # CV EVALUATOR — ADDED FOR REUSABILITY
  393. # ============================================================
  394. def evaluate_cv(X, y, cv, cv_name, out_prefix, groups=None, include_svm=True, feature_indices=None):
  395. if feature_indices is None:
  396. feature_indices = np.arange(X.shape[1])
  397. X_eval = X[:, feature_indices]
  398. models = make_models(include_svm=include_svm)
  399. fold_metrics = {m: defaultdict(list) for m in models}
  400. all_true = {m: [] for m in models}
  401. all_pred = {m: [] for m in models}
  402. rf_importances = []
  403. fold_rows = []
  404. if groups is None:
  405. split_iter = cv.split(X_eval, y)
  406. else:
  407. split_iter = cv.split(X_eval, y, groups=groups)
  408. for fold, (tr, te) in enumerate(split_iter, start=1):
  409. scaler = StandardScaler()
  410. Xtr = scaler.fit_transform(X_eval[tr])
  411. Xte = scaler.transform(X_eval[te])
  412. ytr, yte = y[tr], y[te]
  413. fold_group_info = ""
  414. if groups is not None:
  415. test_groups = sorted(set(np.array(groups)[te].tolist()))
  416. fold_group_info = ",".join(map(str, test_groups[:10]))
  417. if len(test_groups) > 10:
  418. fold_group_info += f",...(+{len(test_groups)-10})"
  419. for name, clf in models.items():
  420. t0 = time.perf_counter()
  421. clf.fit(Xtr, ytr)
  422. t1 = time.perf_counter()
  423. pred = clf.predict(Xte)
  424. if hasattr(clf, "predict_proba"):
  425. proba = clf.predict_proba(Xte)[:, 1]
  426. else:
  427. proba = pred.astype(float)
  428. t2 = time.perf_counter()
  429. acc = accuracy_score(yte, pred)
  430. rec = recall_score(yte, pred, average="binary", zero_division=0)
  431. f1 = f1_score(yte, pred, zero_division=0)
  432. try:
  433. auc = roc_auc_score(yte, proba)
  434. except Exception:
  435. auc = np.nan
  436. fold_metrics[name]["acc"].append(acc)
  437. fold_metrics[name]["recall"].append(rec)
  438. fold_metrics[name]["f1"].append(f1)
  439. fold_metrics[name]["auc"].append(auc)
  440. fold_metrics[name]["train_time"].append(t1 - t0)
  441. fold_metrics[name]["infer_time"].append(t2 - t1)
  442. all_true[name].extend(yte.tolist())
  443. all_pred[name].extend(pred.tolist())
  444. fold_rows.append({
  445. "cv_name": cv_name,
  446. "fold": fold,
  447. "model": name,
  448. "n_train": int(len(tr)),
  449. "n_test": int(len(te)),
  450. "test_groups": fold_group_info,
  451. "acc": acc,
  452. "recall": rec,
  453. "f1": f1,
  454. "auc": auc,
  455. "train_time": t1 - t0,
  456. "infer_time": t2 - t1
  457. })
  458. if name == "RF" and hasattr(clf, "feature_importances_"):
  459. # Importance corresponds to selected feature indices.
  460. imp_full = np.zeros(X.shape[1], dtype=float)
  461. imp_full[feature_indices] = clf.feature_importances_
  462. rf_importances.append(imp_full.copy())
  463. print(f"[{cv_name}] Fold {fold} done. Test size={len(te)}")
  464. # Summary
  465. summary_rows = []
  466. for name in models:
  467. acc = np.array(fold_metrics[name]["acc"], dtype=float)
  468. rec = np.array(fold_metrics[name]["recall"], dtype=float)
  469. f1 = np.array(fold_metrics[name]["f1"], dtype=float)
  470. auc = np.array(fold_metrics[name]["auc"], dtype=float)
  471. trt = np.array(fold_metrics[name]["train_time"], dtype=float)
  472. inf = np.array(fold_metrics[name]["infer_time"], dtype=float)
  473. summary_rows.append({
  474. "CV": cv_name,
  475. "Feature_Set": out_prefix,
  476. "Model": name,
  477. "ACC_mean(%)": float(np.nanmean(acc)*100),
  478. "ACC_std(%)": float(np.nanstd(acc)*100),
  479. "Recall_mean(%)": float(np.nanmean(rec)*100),
  480. "Recall_std(%)": float(np.nanstd(rec)*100),
  481. "F1_mean(%)": float(np.nanmean(f1)*100),
  482. "F1_std(%)": float(np.nanstd(f1)*100),
  483. "AUC_mean(%)": float(np.nanmean(auc)*100),
  484. "AUC_std(%)": float(np.nanstd(auc)*100),
  485. "Train_mean(s)": float(np.nanmean(trt)),
  486. "Train_std(s)": float(np.nanstd(trt)),
  487. "Infer_mean(s)": float(np.nanmean(inf)),
  488. "Infer_std(s)": float(np.nanstd(inf)),
  489. })
  490. df_summary = pd.DataFrame(summary_rows)
  491. df_folds = pd.DataFrame(fold_rows)
  492. df_summary.to_csv(os.path.join(TAB_DIR, f"summary_{out_prefix}_{cv_name}.csv"), index=False)
  493. df_folds.to_csv(os.path.join(TAB_DIR, f"fold_metrics_{out_prefix}_{cv_name}.csv"), index=False)
  494. print(f"\n================ Summary: {cv_name} | {out_prefix} ================")
  495. for r in summary_rows:
  496. print(f"\nModel: {r['Model']}")
  497. print(f"ACC : {r['ACC_mean(%)']:.2f}% ± {r['ACC_std(%)']:.2f}%")
  498. print(f"Recall: {r['Recall_mean(%)']:.2f}% ± {r['Recall_std(%)']:.2f}%")
  499. print(f"F1 : {r['F1_mean(%)']:.2f}% ± {r['F1_std(%)']:.2f}%")
  500. print(f"AUC : {r['AUC_mean(%)']:.2f}% ± {r['AUC_std(%)']:.2f}%")
  501. print(f"Train : {r['Train_mean(s)']:.4f}s ± {r['Train_std(s)']:.4f}s")
  502. print(f"Infer : {r['Infer_mean(s)']:.4f}s ± {r['Infer_std(s)']:.4f}s")
  503. print("==============================================================\n")
  504. # Confusion matrices
  505. for name in models:
  506. save_confusion_matrix(
  507. np.array(all_true[name]), np.array(all_pred[name]),
  508. title=f"LieWaves — {cv_name} Confusion ({name})",
  509. outpath=os.path.join(FIG_DIR, f"confusion_{out_prefix}_{cv_name}_{name}.png"),
  510. labels=("Lie (0)", "Truth (1)")
  511. )
  512. # Paired t-test only when at least 2 folds and comparable models
  513. try:
  514. paired_ttest_matrix(
  515. fold_metrics,
  516. metric_key="acc",
  517. out_csv=os.path.join(TAB_DIR, f"pvalues_acc_{out_prefix}_{cv_name}.csv"),
  518. out_png=os.path.join(FIG_DIR, f"pvalues_acc_{out_prefix}_{cv_name}.png"),
  519. title=f"Paired t-test p-values ({cv_name} ACC)"
  520. )
  521. except Exception as e:
  522. print(f"[WARN] Paired t-test skipped for {cv_name}: {e}")
  523. # RF feature importance
  524. if len(rf_importances) > 0:
  525. mean_imp = np.mean(np.vstack(rf_importances), axis=0)
  526. feat_names = FEAT_NAMES if len(FEAT_NAMES) == X.shape[1] else [f"f{i}" for i in range(X.shape[1])]
  527. df_fi = pd.DataFrame({"feature": feat_names, "importance": mean_imp})
  528. df_fi.sort_values("importance", ascending=False).to_csv(
  529. os.path.join(TAB_DIR, f"rf_feature_importance_all_{out_prefix}_{cv_name}.csv"),
  530. index=False
  531. )
  532. save_barh(
  533. mean_imp, feat_names,
  534. title=f"RF Feature Importance Top 10 — {cv_name}",
  535. out_png=os.path.join(FIG_DIR, f"rf_feature_importance_top10_{out_prefix}_{cv_name}.png"),
  536. topk=10
  537. )
  538. rf_importance_by_group(
  539. mean_imp, SLICES,
  540. out_csv=os.path.join(TAB_DIR, f"rf_importance_by_group_{out_prefix}_{cv_name}.csv"),
  541. out_png=os.path.join(FIG_DIR, f"rf_importance_by_group_{out_prefix}_{cv_name}.png")
  542. )
  543. return df_summary, df_folds
  544. # ============================================================
  545. # MAIN
  546. # ============================================================
  547. if __name__ == "__main__":
  548. # Save config for reproducibility / MethodsX code availability
  549. config = {
  550. "DATA_DIR": DATA_DIR,
  551. "RAW_DIR": RAW_DIR,
  552. "SUBJECT_STIMULI_XLSX": SUBJECT_STIMULI_XLSX,
  553. "CHAN_ORDER": CHAN_ORDER,
  554. "SF": SF,
  555. "WINDOW_SIZE": WINDOW_SIZE,
  556. "STEP_SIZE": STEP_SIZE,
  557. "N_SPLITS": N_SPLITS,
  558. "RANDOM_STATE": RANDOM_STATE,
  559. "RUN_WINDOW_LEVEL_SKF": RUN_WINDOW_LEVEL_SKF,
  560. "RUN_SUBJECT_GROUPK": RUN_SUBJECT_GROUPK,
  561. "RUN_SESSION_GROUPK": RUN_SESSION_GROUPK,
  562. "RUN_LOSO": RUN_LOSO,
  563. "RUN_NON_OVERLAP": RUN_NON_OVERLAP,
  564. "RUN_META_ABLATION": RUN_META_ABLATION,
  565. "INCLUDE_SVM_IN_STRICT_VALIDATION": INCLUDE_SVM_IN_STRICT_VALIDATION,
  566. "INCLUDE_SVM_IN_WINDOW_LEVEL": INCLUDE_SVM_IN_WINDOW_LEVEL,
  567. "HAS_XGB": HAS_XGB
  568. }
  569. with open(os.path.join(TAB_DIR, "run_config.json"), "w") as f:
  570. json.dump(config, f, indent=2)
  571. # 1) Load original overlapping-window dataset
  572. X, y, meta = load_data(step_size=STEP_SIZE)
  573. print("Total samples:", len(X))
  574. print("Dataset shape:", X.shape)
  575. print("Metadata shape:", meta.shape)
  576. meta.to_csv(os.path.join(TAB_DIR, "window_metadata_original_step16.csv"), index=False)
  577. if len(FEAT_NAMES) != X.shape[1]:
  578. print(f"[WARN] Feature name count {len(FEAT_NAMES)} != X features {X.shape[1]}. Using generic names.")
  579. FEAT_NAMES = [f"f{i}" for i in range(X.shape[1])]
  580. save_overlap_report(
  581. step_size=STEP_SIZE,
  582. meta=meta,
  583. out_json=os.path.join(TAB_DIR, "overlap_leakage_report_step16.json"),
  584. out_csv=os.path.join(TAB_DIR, "windows_per_subject_label_step16.csv")
  585. )
  586. # LDA plot for original full descriptor
  587. lda_plot_like_example(X, y, os.path.join(FIG_DIR, "lda_separability_original_step16.png"))
  588. # Save group definitions
  589. group_def = pd.DataFrame([
  590. {"Group": "fTS", "start": SLICES["fTS"].start, "stop": SLICES["fTS"].stop, "dim": SLICES["fTS"].stop - SLICES["fTS"].start},
  591. {"Group": "fCorr", "start": SLICES["fCorr"].start, "stop": SLICES["fCorr"].stop, "dim": SLICES["fCorr"].stop - SLICES["fCorr"].start},
  592. {"Group": "fF", "start": SLICES["fF"].start, "stop": SLICES["fF"].stop, "dim": SLICES["fF"].stop - SLICES["fF"].start},
  593. {"Group": "fSP", "start": SLICES["fSP"].start, "stop": SLICES["fSP"].stop, "dim": SLICES["fSP"].stop - SLICES["fSP"].start},
  594. {"Group": "full", "start": SLICES["full"].start, "stop": SLICES["full"].stop, "dim": SLICES["full"].stop - SLICES["full"].start},
  595. ])
  596. group_def.to_csv(os.path.join(TAB_DIR, "feature_group_slices.csv"), index=False)
  597. # 2) Define feature sets for full and meta-correlation ablation
  598. feature_sets = get_feature_sets(X) if RUN_META_ABLATION else {"full_106": np.arange(X.shape[1])}
  599. all_summaries = []
  600. # 3) Window-level StratifiedKFold evaluation
  601. if RUN_WINDOW_LEVEL_SKF:
  602. for fs_name, fs_idx in feature_sets.items():
  603. # Evaluate the full descriptor and the descriptor without meta-correlation.
  604. cv = StratifiedKFold(n_splits=N_SPLITS, shuffle=True, random_state=RANDOM_STATE)
  605. summary, _ = evaluate_cv(
  606. X, y,
  607. cv=cv,
  608. cv_name="window_stratified5fold",
  609. out_prefix=fs_name,
  610. groups=None,
  611. include_svm=INCLUDE_SVM_IN_WINDOW_LEVEL,
  612. feature_indices=fs_idx
  613. )
  614. all_summaries.append(summary)
  615. # 4) Optional subject-wise GroupKFold validation
  616. if RUN_SUBJECT_GROUPK:
  617. for fs_name, fs_idx in feature_sets.items():
  618. cv = GroupKFold(n_splits=N_SPLITS)
  619. summary, _ = evaluate_cv(
  620. X, y,
  621. cv=cv,
  622. cv_name="subject_group5fold",
  623. out_prefix=fs_name,
  624. groups=meta["subject"].values,
  625. include_svm=INCLUDE_SVM_IN_STRICT_VALIDATION,
  626. feature_indices=fs_idx
  627. )
  628. all_summaries.append(summary)
  629. # 5) Optional session/file-wise GroupKFold
  630. if RUN_SESSION_GROUPK:
  631. for fs_name, fs_idx in feature_sets.items():
  632. cv = GroupKFold(n_splits=N_SPLITS)
  633. summary, _ = evaluate_cv(
  634. X, y,
  635. cv=cv,
  636. cv_name="session_group5fold",
  637. out_prefix=fs_name,
  638. groups=meta["file_group"].values,
  639. include_svm=INCLUDE_SVM_IN_STRICT_VALIDATION,
  640. feature_indices=fs_idx
  641. )
  642. all_summaries.append(summary)
  643. # 6) Optional Leave-One-Subject-Out. This can be slow.
  644. if RUN_LOSO:
  645. for fs_name, fs_idx in feature_sets.items():
  646. cv = LeaveOneGroupOut()
  647. summary, _ = evaluate_cv(
  648. X, y,
  649. cv=cv,
  650. cv_name="LOSO_subject",
  651. out_prefix=fs_name,
  652. groups=meta["subject"].values,
  653. include_svm=INCLUDE_SVM_IN_STRICT_VALIDATION,
  654. feature_indices=fs_idx
  655. )
  656. all_summaries.append(summary)
  657. # 7) Optional non-overlapping windows sensitivity analysis
  658. if RUN_NON_OVERLAP:
  659. X_no, y_no, meta_no = load_data(step_size=WINDOW_SIZE)
  660. meta_no.to_csv(os.path.join(TAB_DIR, "window_metadata_nonoverlap_step384.csv"), index=False)
  661. save_overlap_report(
  662. step_size=WINDOW_SIZE,
  663. meta=meta_no,
  664. out_json=os.path.join(TAB_DIR, "overlap_leakage_report_step384.json"),
  665. out_csv=os.path.join(TAB_DIR, "windows_per_subject_label_step384.csv")
  666. )
  667. lda_plot_like_example(X_no, y_no, os.path.join(FIG_DIR, "lda_separability_nonoverlap_step384.png"))
  668. feature_sets_no = get_feature_sets(X_no) if RUN_META_ABLATION else {"full_106": np.arange(X_no.shape[1])}
  669. for fs_name, fs_idx in feature_sets_no.items():
  670. cv = StratifiedKFold(n_splits=N_SPLITS, shuffle=True, random_state=RANDOM_STATE)
  671. summary, _ = evaluate_cv(
  672. X_no, y_no,
  673. cv=cv,
  674. cv_name="nonoverlap_window_stratified5fold",
  675. out_prefix=fs_name,
  676. groups=None,
  677. include_svm=INCLUDE_SVM_IN_WINDOW_LEVEL,
  678. feature_indices=fs_idx
  679. )
  680. all_summaries.append(summary)
  681. cvg = GroupKFold(n_splits=N_SPLITS)
  682. summary, _ = evaluate_cv(
  683. X_no, y_no,
  684. cv=cvg,
  685. cv_name="nonoverlap_subject_group5fold",
  686. out_prefix=fs_name,
  687. groups=meta_no["subject"].values,
  688. include_svm=INCLUDE_SVM_IN_STRICT_VALIDATION,
  689. feature_indices=fs_idx
  690. )
  691. all_summaries.append(summary)
  692. # 8) Save combined summary
  693. if len(all_summaries) > 0:
  694. df_all = pd.concat(all_summaries, ignore_index=True)
  695. df_all.to_csv(os.path.join(TAB_DIR, "ALL_SUMMARY_COMBINED.csv"), index=False)
  696. print("\n================ ALL SUMMARY COMBINED ================")
  697. print(df_all[["CV", "Feature_Set", "Model", "ACC_mean(%)", "ACC_std(%)", "Recall_mean(%)", "F1_mean(%)", "AUC_mean(%)"]])
  698. print("======================================================\n")
  699. print(f"✅ All figures saved to: {FIG_DIR}")
  700. print(f"✅ All tables saved to: {TAB_DIR}")

run_liewaves_descriptor_pipeline.py at commit d4dec23, no license · at the source

Overview

Authors: Dwi Utari Surya1, Sholeh Hadi Pramono2, Panca Mudjirahardjo2, Muhammad Aziz Muslim2, Cries Avian2, Mahdin Rohmatillah2
ORCID iDs: Dwi Utari Surya
  1. Department of Creative and Digital Industry, Faculty of Vocational Studies, Universitas Brawijaya, 65145, Indonesia
  2. Department of Electrical Engineering, Faculty of Engineering, Universitas Brawijaya, Malang, 65145, Indonesia
Institutions: University of Brawijaya (Indonesia)
Journal: MethodsX, volume 17, article 104077
Dates: received 6 March 2026; accepted 27 July 2026; published online 28 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.mex.2026.104077 · PMID 42571525 · PMCID PMC13452325 · OpenAlex W7171544820
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: EEG (modality), methods / tools (subfield)
Methods: Spectral & time-frequency, Preprocessing, Statistics, Smoothing, state filtering, decompositions, Machine learning, Evoked potentials, Connectivity, Complexity, Physiology & signal measures
Keywords: EEG-based Lie and Truth detection, Multi-domain descriptors, Phase connectivity, Handcrafted EEG features
Journal subjects: Neuroscience
Topic: EEG and Brain-Computer Interfaces (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 35 references in the paper

Abstract

Wearable EEG-based Lie and Truth detection requires structured descriptor engineering to achieve reproducible results and ensure method transparency. The existing approaches use loosely integrated features alongside hidden modeling methods, which create challenges for understanding their results and verifying research progress across different studies. The existing research gap requires the development of a deterministic descriptor framework to document wearable EEG systems.

The research presents the Structured Multi-Domain EEG Descriptor with Phase-Based Connectivity as a reproducible framework that uses five-channel wearable EEG signals to classify binary Lie-Truth. The method emphasizes coordinated descriptor organization and controlled evaluation, summarized as follows: • The structured integration of temporal statistics, fractal complexity, spectral power distributions, and phase-based connectivity within a unified descriptor representation augmented by meta-correlation modelling. • The processing framework provides detailed specifications for preprocessing, window-level segmentation, descriptor extraction, fold-wise normalization, classifier evaluation, and computational profiling. • A controlled evaluation strategy for examining statistical consistency, descriptor separability, domain-level feature contribution, and workstation-based computational feasibility. Validation on the publicly available LieWaves dataset demonstrates stable fold-wise behaviour and consistent classifier responses under controlled conditions.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

dutarisurya/lie-truth-eeg-descriptor

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: d4dec2337ffb134a0783ce81b172b039ea48b890, 29 April 2026
Languages: Python (1)
Size: 4 files, 1 script
Software Heritage: not archived
Found in: the text
Holds: README, environment (lie-truth-eeg-descriptor/requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
2 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

LieWaves data are available at Mendeley Data (https://doi.org/10.17632/5gzxb2bzs2.2). Code: https://github.com/dutarisurya/lie-truth-eeg-descriptor.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 4 keywords, 20 references.

Cite

This paper

Surya, D. U., Pramono, S. H., Mudjirahardjo, P., Muslim, M. A., Avian, C., & Rohmatillah, M. (2026). Structured multi-domain EEG descriptors with phase-based connectivity for lie and truth detection. MethodsX, 17, 104077. https://doi.org/10.1016/j.mex.2026.104077

BibTeX

@article{surya2026structured,
author = {Surya, Dwi Utari and Pramono, Sholeh Hadi and Mudjirahardjo, Panca and Muslim, Muhammad Aziz and Avian, Cries and Rohmatillah, Mahdin},
title = {{Structured multi-domain EEG descriptors with phase-based connectivity for lie and truth detection}},
journal = {MethodsX},
year = {2026},
month = jul,
volume = {17},
pages = {104077},
publisher = {Elsevier},
issn = {2215-0161},
doi = {10.1016/j.mex.2026.104077},
url = {https://doi.org/10.1016/j.mex.2026.104077},
pmid = {42571525},
pmcid = {PMC13452325}
}

RIS

TY - JOUR
AU - Surya, Dwi Utari
AU - Pramono, Sholeh Hadi
AU - Mudjirahardjo, Panca
AU - Muslim, Muhammad Aziz
AU - Avian, Cries
AU - Rohmatillah, Mahdin
TI - Structured multi-domain EEG descriptors with phase-based connectivity for lie and truth detection
T2 - MethodsX
J2 - MethodsX
PY - 2026
DA - 2026/07/28
VL - 17
SP - 104077
SN - 2215-0161
PB - Elsevier
DO - 10.1016/j.mex.2026.104077
UR - https://doi.org/10.1016/j.mex.2026.104077
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.mex.2026.104077",
"type": "article-journal",
"title": "Structured multi-domain EEG descriptors with phase-based connectivity for lie and truth detection",
"container-title": "MethodsX",
"author": [
{
"family": "Surya",
"given": "Dwi Utari"
},
{
"family": "Pramono",
"given": "Sholeh Hadi"
},
{
"family": "Mudjirahardjo",
"given": "Panca"
},
{
"family": "Muslim",
"given": "Muhammad Aziz"
},
{
"family": "Avian",
"given": "Cries"
},
{
"family": "Rohmatillah",
"given": "Mahdin"
}
],
"container-title-short": "MethodsX",
"volume": "17",
"page": "104077",
"DOI": "10.1016/j.mex.2026.104077",
"PMID": "42571525",
"PMCID": "PMC13452325",
"ISSN": "2215-0161",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.mex.2026.104077",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
28
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3389/fpsyg.2026.1774068 [code]
Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.
Journal: Frontiers in psychology
In common: XGBoost, scikit-learn, pandas, 3 other tools, EEG, 1 reference
[2] doi:10.3390/s26134045
Brain Signal for Secure EEG Biometric Authentication: A Comprehensive Survey.
Journal: Sensors (Basel, Switzerland)
In common: methods / tools, EEG, 4 references
[3] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: XGBoost, scikit-learn, pandas, 3 other tools, methods / tools
[4] doi:10.1126/sciadv.aed3650 [code]
Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &lt;i&gt;m/z&lt;/i&gt; mapping and exploration.
Journal: Science advances
In common: XGBoost, scikit-learn, pandas, 3 other tools, methods / tools
[5] doi:10.1186/s13059-026-04125-8 [code]
MLMarker: a machine learning framework for tissue inference and biomarker discovery.
Journal: Genome biology
In common: XGBoost, scikit-learn, pandas, 3 other tools, methods / tools
[6] doi:10.1038/s41586-026-10658-6 [code]
An AI system to help scientists write expert-level empirical software.
Journal: Nature
In common: XGBoost, scikit-learn, pandas, 3 other tools, methods / tools
[7] doi:10.3390/s26175327 [code]
Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.
Journal: Sensors (Basel, Switzerland)
In common: XGBoost, scikit-learn, pandas, 3 other tools, EEG
[8] doi:10.1007/s10548-026-01238-y [code]
Topographic Reorganization of EEG Complexity During Visual Mental Imagery: Insights from Lempel-Ziv Complexity in High-Density EEG.
Journal: Brain topography
In common: XGBoost, scikit-learn, pandas, 3 other tools, EEG
[9] doi:10.3389/fnhum.2026.1869918 [code]
Single-subject auditory ERP-BCI performance enhancement in ALS via an AI coding assistant prompt.
Journal: Frontiers in human neuroscience
In common: XGBoost, scikit-learn, pandas, 3 other tools, EEG
[10] doi:10.1093/brain/awaf412 [code]
Multimodal multicentre investigation of diagnostic and prognostic markers in disorders of consciousness.
Journal: Brain : a journal of neurology
In common: XGBoost, scikit-learn, pandas, 3 other tools, EEG

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.