OSCR

PhysioMotion Artifact: A task-driven EEG dataset with point-wise motion artifact annotations.

Code ↔ Paper

13 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 13 matches
  1. [1] § Methods › Preprocessing and Annotations ↔ preprocess.py, lines 20–63 · score 0.87 · bandpass filter, notch filter, 0.5–150 Hz, Bipolar referencing, bipolar montages, preprocessed
  2. [2] § Usage Notes ↔ classification/model_M.py, lines 4–63 · score 0.81 · NumPy, PyTorch, EEG artifact classification, GPUtil, Python, Hyperopt
  3. [3] § Technical Validation › Model Input ↔ classification/model_M.py, lines 4–63 · score 0.78 · GroupKFold, inner folds, outer folds, nested, tuned, hyperparameters
  4. [4] § Technical Validation › Results ↔ classification/model_B.py, lines 468–490 · score 0.76 · standard classification metrics, F1 score, confusion matrices, Precision, Recall, binary
  5. [5] § Technical Validation › Results ↔ classification/model_M.py, lines 70–123 · score 0.75 · swallow_eyebrow, tongue_eyebrow, class weighted, motion, chew, training
  6. [6] § Technical Validation › Model Input ↔ classification/model_B.py, lines 4–58 · score 0.70 · GroupKFold, inner folds, nested, tuned, hyperparameters, outer
  7. [7] § Technical Validation › Results ↔ classification/model_M.py, lines 457–466 · score 0.64 · F1 score, confusion matrices, multi class, metrics, classification, model
  8. [8] § Technical Validation › Model Architectures ↔ classification/model_B.py, lines 637–711 · score 0.62 · cross entropy loss, class weighted, optimized, training, model
  9. [9] § Technical Validation › Model Architectures ↔ classification/model_B.py, lines 637–711 · score 0.59 · cross entropy loss, class weighted, optimized, Model
  10. [10] § Technical Validation › Model Architectures ↔ classification/model_B.py, lines 299–358 · score 0.55 · Transformer encoders, CNN Transformer, layers, model, classifiers
  11. [11] § Technical Validation › Model Architectures ↔ classification/model_M.py, lines 283–335 · score 0.55 · Transformer encoders, CNN Transformer, layers, model, classifiers
  12. [12] § Technical Validation › Hyperparameter Optimization ↔ classification/model_B.py, lines 754–834 · score 0.54 · weight decay, Transformer layers, batch, CNN, Optimization
  13. [13] § Data Records › Preprocessed Data ↔ preprocess.py, lines 20–63 · score 0.51 · BIDS format, MNE, preprocessing, raw, EEG

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 873 lines · 31 KB · MIT · 6 matches

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. """
  4. Nested subject-independent GroupKFold training with Hyperopt for EEG artifact binary classification.
  5. Key features
  6. - Subject-level nested CV (outer = evaluation, inner = hyperparameter tuning)
  7. - Hyperopt (TPE) tuning on inner-fold mean validation loss
  8. - Robust CUDA OOM handling (a trial that OOMs is assigned a very bad loss and skipped)
  9. - Detailed logging to both console and a log file
  10. Notes
  11. - This script assumes each .pkl file contains a list of dicts, each dict includes:
  12. - "eeg_data": array-like shaped (C, T) (e.g., 34 x 375)
  13. - "artifact_types": list[str], e.g. ["close_base"] or ["blink", ...]
  14. - Labels are mapped to a binary task:
  15. - ["close_base"] -> "close_base"
  16. - otherwise -> "artifact"
  17. """
  18. import os
  19. import re
  20. import sys
  21. import math
  22. import time
  23. import pickle
  24. import random
  25. import logging
  26. import warnings
  27. from dataclasses import dataclass
  28. from collections import Counter, defaultdict
  29. from typing import Dict, List, Tuple
  30. import numpy as np
  31. import torch
  32. import torch.nn as nn
  33. import torch.optim as optim
  34. from torch.utils.data import Dataset, DataLoader, Subset
  35. from sklearn.model_selection import GroupKFold
  36. from sklearn.metrics import (
  37. classification_report,
  38. confusion_matrix,
  39. accuracy_score,
  40. balanced_accuracy_score,
  41. f1_score,
  42. precision_score,
  43. recall_score,
  44. )
  45. from hyperopt import fmin, tpe, hp, Trials, STATUS_OK
  46. try:
  47. import GPUtil
  48. except Exception:
  49. GPUtil = None
  50. # -------------------------
  51. # Configuration
  52. # -------------------------
  53. @dataclass
  54. class Config:
  55. # Data
  56. data_dir: str = "/home/mnt_disk1/Motion_Artifact_han1/Slidingwindow"
  57. top2_labels: Tuple[str, str] = ("close_base", "artifact")
  58. # Nested CV
  59. outer_k: int = 5
  60. inner_k: int = 3
  61. max_evals: int = 30
  62. seed: int = 42
  63. # Training
  64. final_num_epochs: int = 15
  65. final_patience: int = 3
  66. inner_num_epochs: int = 10
  67. inner_patience: int = 2
  68. # DataLoader
  69. num_workers: int = 2
  70. pin_memory: bool = True
  71. # Logging
  72. log_filename: str = "nested_groupkfold_verbose_log.txt"
  73. log_every_epoch: bool = True
  74. log_inner_fold_detail: bool = True
  75. log_trial_params: bool = True
  76. show_tqdm: bool = False # keep False for cleaner logs
  77. # GPU selection / OOM safety
  78. min_free_mb: int = 6000
  79. oom_bad_loss: float = 1e9
  80. # Hyperopt search space (OOM-safer default)
  81. # (You can widen it, but OOM risk increases.)
  82. batch_size_choices: Tuple[int, ...] = (64, 128)
  83. transformer_layers_choices: Tuple[int, ...] = (1, 2)
  84. transformer_nhead_choices: Tuple[int, ...] = (2, 4, 8)
  85. num_conv_blocks_choices: Tuple[int, ...] = (3, 4, 5)
  86. # Class-weight scaling search ranges
  87. cw0_low: float = 0.5
  88. cw0_high: float = 0.7
  89. cw1_low: float = 1.1
  90. cw1_high: float = 1.4
  91. CFG = Config()
  92. warnings.filterwarnings("ignore")
  93. # -------------------------
  94. # Logging utilities
  95. # -------------------------
  96. def setup_logger(log_filename: str) -> logging.Logger:
  97. logger = logging.getLogger("nested_cv")
  98. logger.setLevel(logging.INFO)
  99. logger.handlers.clear()
  100. fmt = logging.Formatter("%(asctime)s [%(levelname)s] %(message)s")
  101. fh = logging.FileHandler(log_filename, mode="w")
  102. fh.setFormatter(fmt)
  103. logger.addHandler(fh)
  104. sh = logging.StreamHandler(sys.stdout)
  105. sh.setFormatter(fmt)
  106. logger.addHandler(sh)
  107. return logger
  108. logger = setup_logger(CFG.log_filename)
  109. # -------------------------
  110. # Reproducibility
  111. # -------------------------
  112. def set_seed(seed: int) -> None:
  113. random.seed(seed)
  114. np.random.seed(seed)
  115. torch.manual_seed(seed)
  116. torch.cuda.manual_seed_all(seed)
  117. # Determinism (may reduce throughput)
  118. torch.backends.cudnn.deterministic = True
  119. torch.backends.cudnn.benchmark = False
  120. # -------------------------
  121. # GPU selection & cleanup
  122. # -------------------------
  123. def select_free_gpu(min_free_mb: int = 6000) -> torch.device:
  124. """
  125. Select the GPU with the largest free memory.
  126. If free memory is below threshold, fall back to CPU (safer for long tuning runs).
  127. """
  128. if (GPUtil is None) or (not torch.cuda.is_available()):
  129. logger.info("GPUtil not available or CUDA not available; using CPU.")
  130. return torch.device("cpu")
  131. gpus = GPUtil.getGPUs()
  132. if len(gpus) == 0:
  133. logger.info("No GPUs detected by GPUtil; using CPU.")
  134. return torch.device("cpu")
  135. best = max(gpus, key=lambda g: g.memoryFree)
  136. logger.info("GPU candidates: " + " | ".join([f"id={g.id}, free={g.memoryFree:.0f}MB" for g in gpus]))
  137. logger.info(f"Selected GPU: id={best.id} name={best.name}, free={best.memoryFree:.0f}MB")
  138. if best.memoryFree < min_free_mb:
  139. logger.warning(
  140. f"Best GPU free memory {best.memoryFree:.0f}MB < {min_free_mb}MB. "
  141. f"Falling back to CPU to reduce OOM risk. (You may lower min_free_mb to force GPU.)"
  142. )
  143. return torch.device("cpu")
  144. return torch.device(f"cuda:{best.id}")
  145. def cleanup_cuda() -> None:
  146. """Best-effort CUDA cache cleanup."""
  147. if torch.cuda.is_available():
  148. torch.cuda.empty_cache()
  149. try:
  150. torch.cuda.ipc_collect()
  151. except Exception:
  152. pass
  153. # -------------------------
  154. # Data loading
  155. # -------------------------
  156. SUBJECT_RE = re.compile(r"(sub\d+)_run\d+\.pkl$", re.IGNORECASE)
  157. def parse_subject_id(filename: str) -> str:
  158. """
  159. Parse subject id from filename like:
  160. sub1_run01.pkl -> "sub1"
  161. """
  162. m = SUBJECT_RE.search(filename)
  163. if not m:
  164. raise ValueError(
  165. f"Cannot parse subject_id from filename: {filename}. Expected pattern like sub1_run01.pkl"
  166. )
  167. return m.group(1).lower()
  168. def load_all_samples_with_subject(data_dir: str) -> List[dict]:
  169. """
  170. Load all samples from all .pkl files under data_dir.
  171. Adds 'subject_id' into each sample dict.
  172. """
  173. if not os.path.isdir(data_dir):
  174. raise FileNotFoundError(f"DATA_DIR not found: {data_dir}")
  175. pkl_files = sorted([f for f in os.listdir(data_dir) if f.endswith(".pkl")])
  176. if len(pkl_files) == 0:
  177. raise FileNotFoundError(f"No .pkl files found in {data_dir}")
  178. samples: List[dict] = []
  179. for fname in pkl_files:
  180. subject_id = parse_subject_id(fname)
  181. path = os.path.join(data_dir, fname)
  182. with open(path, "rb") as f:
  183. arr = pickle.load(f)
  184. # arr is expected to be list[dict]
  185. for s in arr:
  186. if isinstance(s, dict) and ("eeg_data" in s) and ("artifact_types" in s):
  187. d = dict(s)
  188. d["subject_id"] = subject_id
  189. samples.append(d)
  190. return samples
  191. # -------------------------
  192. # Dataset
  193. # -------------------------
  194. class EEGArtifactDataset(Dataset):
  195. """
  196. Binary dataset:
  197. label = "close_base" if artifact_types == ["close_base"] else "artifact"
  198. """
  199. def __init__(self, samples: List[dict], label_to_idx: Dict[str, int], scale: float = 1e6):
  200. super().__init__()
  201. self.label_to_idx = label_to_idx
  202. self.scale = float(scale)
  203. filtered = []
  204. for s in samples:
  205. if ("eeg_data" not in s) or ("artifact_types" not in s) or ("subject_id" not in s):
  206. continue
  207. label = "close_base" if s["artifact_types"] == ["close_base"] else "artifact"
  208. if label not in label_to_idx:
  209. continue
  210. d = dict(s)
  211. d["label"] = label
  212. filtered.append(d)
  213. self.samples = filtered
  214. logger.info(f"Filtered dataset size: {len(self.samples)}")
  215. label_counter = Counter([s["label"] for s in self.samples])
  216. subj_counter = Counter([s["subject_id"] for s in self.samples])
  217. logger.info(f"Label distribution: {label_counter}")
  218. logger.info(
  219. f"Num subjects: {len(subj_counter)} | "
  220. f"subject sample range: min={min(subj_counter.values())}, max={max(subj_counter.values())}"
  221. )
  222. def __len__(self) -> int:
  223. return len(self.samples)
  224. def __getitem__(self, idx: int):
  225. s = self.samples[idx]
  226. x = np.asarray(s["eeg_data"], dtype=np.float32) * self.scale # scale to microvolts
  227. y = self.label_to_idx[s["label"]]
  228. return torch.tensor(x, dtype=torch.float32), torch.tensor(y, dtype=torch.long)
  229. # -------------------------
  230. # Model
  231. # -------------------------
  232. class CNNTransformerModel(nn.Module):
  233. """
  234. CNN + Transformer encoder baseline.
  235. Input: (B, C, T) -> treated as (B, 1, C, T) to use 2D conv.
  236. """
  237. def __init__(
  238. self,
  239. num_classes: int = 2,
  240. cnn_filters: List[int] = None,
  241. transformer_nhead: int = 8,
  242. transformer_ff: int = 1024,
  243. transformer_layers: int = 1,
  244. dropout_p: float = 0.5,
  245. ):
  246. super().__init__()
  247. if cnn_filters is None:
  248. cnn_filters = [8, 16, 32, 64, 128]
  249. convs = []
  250. in_ch = 1
  251. for out_ch in cnn_filters:
  252. convs.append(
  253. nn.Sequential(
  254. nn.Conv2d(in_ch, out_ch, kernel_size=3, padding=1),
  255. nn.BatchNorm2d(out_ch),
  256. nn.ReLU(),
  257. nn.MaxPool2d(kernel_size=2, stride=2),
  258. )
  259. )
  260. in_ch = out_ch
  261. self.cnn = nn.Sequential(*convs)
  262. enc_layer = nn.TransformerEncoderLayer(
  263. d_model=cnn_filters[-1],
  264. nhead=transformer_nhead,
  265. dim_feedforward=transformer_ff,
  266. dropout=0.1,
  267. activation="relu",
  268. batch_first=True,
  269. )
  270. self.transformer = nn.TransformerEncoder(enc_layer, num_layers=transformer_layers)
  271. self.global_pool = nn.AdaptiveAvgPool1d(1)
  272. self.fc1 = nn.Linear(cnn_filters[-1], 100)
  273. self.fc2 = nn.Linear(100, 100)
  274. self.dropout = nn.Dropout(p=dropout_p)
  275. self.out = nn.Linear(100, num_classes)
  276. def forward(self, x: torch.Tensor) -> torch.Tensor:
  277. x = x.unsqueeze(1) # (B,1,C,T)
  278. x = self.cnn(x) # (B,Ch,H,W)
  279. x = x.flatten(2).transpose(1, 2) # (B,seq,Ch)
  280. x = self.transformer(x) # (B,seq,Ch)
  281. x = x.transpose(1, 2) # (B,Ch,seq)
  282. x = self.global_pool(x).squeeze(-1) # (B,Ch)
  283. x = torch.relu(self.fc1(x))
  284. x = torch.relu(self.fc2(x))
  285. x = self.dropout(x)
  286. return self.out(x)
  287. # -------------------------
  288. # Training / evaluation
  289. # -------------------------
  290. def compute_class_weights_from_indices(
  291. dataset: EEGArtifactDataset,
  292. indices: np.ndarray,
  293. label_to_idx: Dict[str, int]
  294. ) -> np.ndarray:
  295. """Compute inverse-frequency weights from a subset of indices."""
  296. counts = np.zeros(len(label_to_idx), dtype=np.float64)
  297. for i in indices:
  298. lbl = dataset.samples[int(i)]["label"]
  299. counts[label_to_idx[lbl]] += 1.0
  300. counts = np.maximum(counts, 1.0)
  301. total = float(counts.sum())
  302. weights = total / counts
  303. return weights.astype(np.float32)
  304. def train_one_run(
  305. model: nn.Module,
  306. criterion: nn.Module,
  307. optimizer: optim.Optimizer,
  308. train_loader: DataLoader,
  309. val_loader: DataLoader,
  310. device: torch.device,
  311. num_epochs: int,
  312. patience: int,
  313. run_name: str,
  314. log_every_epoch: bool,
  315. ) -> Tuple[float, float]:
  316. """
  317. Train with early stopping on validation loss.
  318. Returns: (best_val_loss, best_val_acc_at_end_epoch)
  319. """
  320. best_val_loss = float("inf")
  321. epochs_wo_improve = 0
  322. best_state = None
  323. best_val_acc = 0.0
  324. for epoch in range(1, num_epochs + 1):
  325. model.train()
  326. train_loss_sum = 0.0
  327. train_correct = 0
  328. train_n = 0
  329. for x, y in train_loader:
  330. x, y = x.to(device), y.to(device)
  331. optimizer.zero_grad(set_to_none=True)
  332. logits = model(x)
  333. loss = criterion(logits, y)
  334. loss.backward()
  335. optimizer.step()
  336. bs = x.size(0)
  337. train_loss_sum += float(loss.item()) * bs
  338. train_correct += int((logits.argmax(1) == y).sum().item())
  339. train_n += bs
  340. avg_train_loss = train_loss_sum / max(1, train_n)
  341. train_acc = train_correct / max(1, train_n)
  342. model.eval()
  343. val_loss_sum = 0.0
  344. val_correct = 0
  345. val_n = 0
  346. with torch.no_grad():
  347. for x, y in val_loader:
  348. x, y = x.to(device), y.to(device)
  349. logits = model(x)
  350. loss = criterion(logits, y)
  351. bs = x.size(0)
  352. val_loss_sum += float(loss.item()) * bs
  353. val_correct += int((logits.argmax(1) == y).sum().item())
  354. val_n += bs
  355. avg_val_loss = val_loss_sum / max(1, val_n)
  356. val_acc = val_correct / max(1, val_n)
  357. if log_every_epoch:
  358. logger.info(
  359. f"{run_name} | Epoch {epoch}/{num_epochs} "
  360. f"| train_loss={avg_train_loss:.4f}, train_acc={train_acc:.4f} "
  361. f"| val_loss={avg_val_loss:.4f}, val_acc={val_acc:.4f}"
  362. )
  363. if avg_val_loss < best_val_loss:
  364. best_val_loss = avg_val_loss
  365. best_val_acc = val_acc
  366. epochs_wo_improve = 0
  367. # Save CPU copy for safety
  368. best_state = {k: v.detach().cpu().clone() for k, v in model.state_dict().items()}
  369. else:
  370. epochs_wo_improve += 1
  371. if epochs_wo_improve >= patience:
  372. if log_every_epoch:
  373. logger.info(f"{run_name} | Early stopping at epoch {epoch}.")
  374. break
  375. if best_state is not None:
  376. model.load_state_dict(best_state)
  377. return float(best_val_loss), float(best_val_acc)
  378. def evaluate_model(model: nn.Module, loader: DataLoader, device: torch.device, labels: List[str]) -> Dict:
  379. """Evaluate on loader and return standard classification metrics + confusion matrix."""
  380. model.eval()
  381. preds, trues = [], []
  382. with torch.no_grad():
  383. for x, y in loader:
  384. x = x.to(device)
  385. logits = model(x)
  386. preds.extend(logits.argmax(1).cpu().numpy().tolist())
  387. trues.extend(y.numpy().tolist())
  388. y_true = np.array(trues, dtype=np.int64)
  389. y_pred = np.array(preds, dtype=np.int64)
  390. return {
  391. "acc": accuracy_score(y_true, y_pred),
  392. "bacc": balanced_accuracy_score(y_true, y_pred),
  393. "f1": f1_score(y_true, y_pred, average="binary"),
  394. "precision": precision_score(y_true, y_pred, average="binary", zero_division=0),
  395. "recall": recall_score(y_true, y_pred, average="binary", zero_division=0),
  396. "cm": confusion_matrix(y_true, y_pred),
  397. "report": classification_report(y_true, y_pred, target_names=labels, zero_division=0),
  398. }
  399. # -------------------------
  400. # Summary stats (mean/std/95% CI)
  401. # -------------------------
  402. def t_critical_975(df: int) -> float:
  403. """
  404. Two-sided 95% CI critical t value (0.975 quantile).
  405. For df > 30, use normal approximation 1.96.
  406. """
  407. table = {
  408. 1: 12.706, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571,
  409. 6: 2.447, 7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228,
  410. 11: 2.201, 12: 2.179, 13: 2.160, 14: 2.145, 15: 2.131,
  411. 16: 2.120, 17: 2.110, 18: 2.101, 19: 2.093, 20: 2.086,
  412. 21: 2.080, 22: 2.074, 23: 2.069, 24: 2.064, 25: 2.060,
  413. 26: 2.056, 27: 2.052, 28: 2.048, 29: 2.045, 30: 2.042
  414. }
  415. if df <= 0:
  416. return float("nan")
  417. return float(table.get(df, 1.96))
  418. def mean_std_ci(values: List[float]) -> Tuple[float, float, Tuple[float, float]]:
  419. v = np.asarray(values, dtype=np.float64)
  420. mean = float(v.mean())
  421. std = float(v.std(ddof=1)) if len(v) > 1 else 0.0
  422. n = len(v)
  423. if n > 1:
  424. tcrit = t_critical_975(n - 1)
  425. half = tcrit * std / math.sqrt(n)
  426. else:
  427. half = 0.0
  428. return mean, std, (mean - half, mean + half)
  429. # -------------------------
  430. # Hyperopt search space
  431. # -------------------------
  432. def build_search_space(cfg: Config) -> Dict:
  433. """
  434. Hyperopt search space. Keep it relatively safe to avoid OOM.
  435. """
  436. return {
  437. "lr": hp.loguniform("lr", np.log(1e-5), np.log(1e-3)),
  438. "batch_size": hp.choice("batch_size", list(cfg.batch_size_choices)),
  439. "transformer_layers": hp.choice("transformer_layers", list(cfg.transformer_layers_choices)),
  440. "transformer_nhead": hp.choice("transformer_nhead", list(cfg.transformer_nhead_choices)),
  441. "num_conv_blocks": hp.choice("num_conv_blocks", list(cfg.num_conv_blocks_choices)),
  442. "weight_decay": hp.loguniform("weight_decay", np.log(1e-6), np.log(1e-3)),
  443. "class_weight_scaling_0": hp.uniform("class_weight_scaling_0", cfg.cw0_low, cfg.cw0_high),
  444. "class_weight_scaling_1": hp.uniform("class_weight_scaling_1", cfg.cw1_low, cfg.cw1_high),
  445. }
  446. def decode_best_params(best: Dict, cfg: Config) -> Dict:
  447. """Decode hp.choice indices into actual values."""
  448. bs_choices = list(cfg.batch_size_choices)
  449. lyr_choices = list(cfg.transformer_layers_choices)
  450. head_choices = list(cfg.transformer_nhead_choices)
  451. conv_choices = list(cfg.num_conv_blocks_choices)
  452. return {
  453. "lr": float(best["lr"]),
  454. "batch_size": int(bs_choices[int(best["batch_size"])]),
  455. "transformer_layers": int(lyr_choices[int(best["transformer_layers"])]),
  456. "transformer_nhead": int(head_choices[int(best["transformer_nhead"])]),
  457. "num_conv_blocks": int(conv_choices[int(best["num_conv_blocks"])]),
  458. "weight_decay": float(best["weight_decay"]),
  459. "class_weight_scaling": [float(best["class_weight_scaling_0"]), float(best["class_weight_scaling_1"])],
  460. }
  461. # -------------------------
  462. # Main: nested CV
  463. # -------------------------
  464. def main(cfg: Config) -> None:
  465. set_seed(cfg.seed)
  466. device = select_free_gpu(min_free_mb=cfg.min_free_mb)
  467. logger.info("########################")
  468. logger.info("Nested GroupKFold subject-level CV (OOM-safe)")
  469. logger.info(f"DATA_DIR={cfg.data_dir}")
  470. logger.info(f"OUTER_K={cfg.outer_k}, INNER_K={cfg.inner_k}, MAX_EVALS={cfg.max_evals}")
  471. logger.info("Tip: to reduce CUDA fragmentation, you may set:")
  472. logger.info(" export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True")
  473. logger.info("########################")
  474. raw_samples = load_all_samples_with_subject(cfg.data_dir)
  475. logger.info(f"Total samples loaded (raw): {len(raw_samples)}")
  476. label_to_idx = {lbl: i for i, lbl in enumerate(cfg.top2_labels)}
  477. dataset = EEGArtifactDataset(raw_samples, label_to_idx)
  478. groups = np.array([s["subject_id"] for s in dataset.samples])
  479. unique_subjects = sorted(set(groups.tolist()))
  480. logger.info(f"Unique subjects ({len(unique_subjects)}): {unique_subjects}")
  481. space = build_search_space(cfg)
  482. outer_cv = GroupKFold(n_splits=min(cfg.outer_k, len(unique_subjects)))
  483. outer_metrics = defaultdict(list)
  484. for outer_fold, (trainval_idx, test_idx) in enumerate(
  485. outer_cv.split(np.zeros(len(dataset)), np.zeros(len(dataset)), groups=groups),
  486. start=1
  487. ):
  488. trainval_idx = np.array(trainval_idx, dtype=np.int64)
  489. test_idx = np.array(test_idx, dtype=np.int64)
  490. trainval_groups = groups[trainval_idx]
  491. test_groups = groups[test_idx]
  492. logger.info("=" * 90)
  493. logger.info(f"[Outer Fold {outer_fold}/{outer_cv.n_splits}]")
  494. logger.info(f"TrainVal subjects ({len(set(trainval_groups))}): {sorted(set(trainval_groups.tolist()))}")
  495. logger.info(f"Test subjects ({len(set(test_groups))}): {sorted(set(test_groups.tolist()))}")
  496. logger.info(f"TrainVal samples={len(trainval_idx)} | Test samples={len(test_idx)}")
  497. inner_n_splits = min(cfg.inner_k, len(set(trainval_groups.tolist())))
  498. if inner_n_splits < 2:
  499. raise ValueError("Not enough subjects in TrainVal split for inner GroupKFold. Reduce OUTER_K/INNER_K.")
  500. inner_cv = GroupKFold(n_splits=inner_n_splits)
  501. trial_counter = {"i": 0}
  502. best_seen = {"loss": float("inf")}
  503. def objective(params: Dict) -> Dict:
  504. trial_counter["i"] += 1
  505. t0 = time.time()
  506. lr = float(params["lr"])
  507. batch_size = int(params["batch_size"])
  508. transformer_layers = int(params["transformer_layers"])
  509. transformer_nhead = int(params["transformer_nhead"])
  510. num_conv_blocks = int(params["num_conv_blocks"])
  511. weight_decay = float(params["weight_decay"])
  512. class_weight_scaling = [float(params["class_weight_scaling_0"]), float(params["class_weight_scaling_1"])]
  513. base_filters = [8, 16, 32, 64, 128]
  514. cnn_filters = base_filters[:num_conv_blocks]
  515. # Constraint: d_model must be divisible by nhead
  516. if cnn_filters[-1] % transformer_nhead != 0:
  517. logger.info(
  518. f"[Outer {outer_fold}] Trial {trial_counter['i']} INVALID (d_model%head!=0) => loss={cfg.oom_bad_loss}"
  519. )
  520. return {"loss": cfg.oom_bad_loss, "status": STATUS_OK}
  521. if cfg.log_trial_params:
  522. logger.info(f"[Outer {outer_fold}] Trial {trial_counter['i']}/{cfg.max_evals} params={params}")
  523. fold_losses = []
  524. fold_accs = []
  525. try:
  526. for inner_fold, (inner_tr_rel, inner_va_rel) in enumerate(
  527. inner_cv.split(np.zeros(len(trainval_idx)), np.zeros(len(trainval_idx)), groups=trainval_groups),
  528. start=1
  529. ):
  530. inner_tr_idx = trainval_idx[inner_tr_rel]
  531. inner_va_idx = trainval_idx[inner_va_rel]
  532. # class weights computed from inner-train only
  533. base_w = compute_class_weights_from_indices(dataset, inner_tr_idx, label_to_idx)
  534. adjusted = base_w * np.array(class_weight_scaling, dtype=np.float32)
  535. class_w_tensor = torch.tensor(adjusted, dtype=torch.float32, device=device)
  536. tr_loader = DataLoader(
  537. Subset(dataset, inner_tr_idx.tolist()),
  538. batch_size=batch_size,
  539. shuffle=True,
  540. num_workers=cfg.num_workers,
  541. pin_memory=cfg.pin_memory,
  542. )
  543. va_loader = DataLoader(
  544. Subset(dataset, inner_va_idx.tolist()),
  545. batch_size=batch_size,
  546. shuffle=False,
  547. num_workers=cfg.num_workers,
  548. pin_memory=cfg.pin_memory,
  549. )
  550. model = CNNTransformerModel(
  551. num_classes=len(label_to_idx),
  552. cnn_filters=cnn_filters,
  553. transformer_layers=transformer_layers,
  554. transformer_nhead=transformer_nhead,
  555. ).to(device)
  556. criterion = nn.CrossEntropyLoss(weight=class_w_tensor)
  557. optimizer = optim.Adam(model.parameters(), lr=lr, weight_decay=weight_decay)
  558. run_name = f"[Outer {outer_fold}][Trial {trial_counter['i']}][Inner {inner_fold}]"
  559. val_loss, val_acc = train_one_run(
  560. model=model,
  561. criterion=criterion,
  562. optimizer=optimizer,
  563. train_loader=tr_loader,
  564. val_loader=va_loader,
  565. device=device,
  566. num_epochs=cfg.inner_num_epochs,
  567. patience=cfg.inner_patience,
  568. run_name=run_name,
  569. log_every_epoch=cfg.log_every_epoch,
  570. )
  571. fold_losses.append(val_loss)
  572. fold_accs.append(val_acc)
  573. if cfg.log_inner_fold_detail:
  574. subs_tr = sorted(set(trainval_groups[inner_tr_rel].tolist()))
  575. subs_va = sorted(set(trainval_groups[inner_va_rel].tolist()))
  576. logger.info(
  577. f"{run_name} | done | val_loss={val_loss:.4f}, val_acc={val_acc:.4f} "
  578. f"| tr_subs={subs_tr} | va_subs={subs_va}"
  579. )
  580. # release memory
  581. del model, optimizer, criterion, tr_loader, va_loader
  582. cleanup_cuda()
  583. except torch.cuda.OutOfMemoryError as e:
  584. logger.warning(
  585. f"[Outer {outer_fold}] Trial {trial_counter['i']} OOM => mark as bad loss. "
  586. f"err={str(e)[:160]}"
  587. )
  588. cleanup_cuda()
  589. dt = time.time() - t0
  590. logger.info(f"[Outer {outer_fold}] Trial {trial_counter['i']} finished (OOM) | time={dt:.1f}s")
  591. return {"loss": cfg.oom_bad_loss, "status": STATUS_OK}
  592. mean_loss = float(np.mean(fold_losses)) if len(fold_losses) > 0 else cfg.oom_bad_loss
  593. mean_acc = float(np.mean(fold_accs)) if len(fold_accs) > 0 else 0.0
  594. dt = time.time() - t0
  595. best_seen["loss"] = min(best_seen["loss"], mean_loss)
  596. logger.info(
  597. f"[Outer {outer_fold}] Trial {trial_counter['i']} finished "
  598. f"| inner_mean_loss={mean_loss:.6f}, inner_mean_acc={mean_acc:.4f} "
  599. f"| best_loss_so_far={best_seen['loss']:.6f} | time={dt:.1f}s"
  600. )
  601. return {"loss": mean_loss, "status": STATUS_OK}
  602. logger.info(f"[Outer Fold {outer_fold}] Hyperopt tuning (inner CV mean val_loss) ...")
  603. trials = Trials()
  604. best = fmin(
  605. fn=objective,
  606. space=space,
  607. algo=tpe.suggest,
  608. max_evals=cfg.max_evals,
  609. trials=trials,
  610. rstate=np.random.default_rng(cfg.seed + outer_fold),
  611. )
  612. best_params = decode_best_params(best, cfg)
  613. logger.info(f"[Outer Fold {outer_fold}] Best hyperparams: {best_params}")
  614. # ----- Final training: trainval -> test evaluation -----
  615. base_filters = [8, 16, 32, 64, 128]
  616. cnn_filters = base_filters[:best_params["num_conv_blocks"]]
  617. base_w = compute_class_weights_from_indices(dataset, trainval_idx, label_to_idx)
  618. adjusted = base_w * np.array(best_params["class_weight_scaling"], dtype=np.float32)
  619. class_w_tensor = torch.tensor(adjusted, dtype=torch.float32, device=device)
  620. train_loader = DataLoader(
  621. Subset(dataset, trainval_idx.tolist()),
  622. batch_size=best_params["batch_size"],
  623. shuffle=True,
  624. num_workers=cfg.num_workers,
  625. pin_memory=cfg.pin_memory,
  626. )
  627. test_loader = DataLoader(
  628. Subset(dataset, test_idx.tolist()),
  629. batch_size=best_params["batch_size"],
  630. shuffle=False,
  631. num_workers=cfg.num_workers,
  632. pin_memory=cfg.pin_memory,
  633. )
  634. model = CNNTransformerModel(
  635. num_classes=len(label_to_idx),
  636. cnn_filters=cnn_filters,
  637. transformer_layers=best_params["transformer_layers"],
  638. transformer_nhead=best_params["transformer_nhead"],
  639. ).to(device)
  640. criterion = nn.CrossEntropyLoss(weight=class_w_tensor)
  641. optimizer = optim.Adam(model.parameters(), lr=best_params["lr"], weight_decay=best_params["weight_decay"])
  642. logger.info(f"[Outer Fold {outer_fold}] Final training on TrainVal, then evaluate on Test ...")
  643. try:
  644. _ = train_one_run(
  645. model=model,
  646. criterion=criterion,
  647. optimizer=optimizer,
  648. train_loader=train_loader,
  649. val_loader=test_loader, # keep your original behavior
  650. device=device,
  651. num_epochs=cfg.final_num_epochs,
  652. patience=cfg.final_patience,
  653. run_name=f"[Outer {outer_fold}][FINAL]",
  654. log_every_epoch=True,
  655. )
  656. except torch.cuda.OutOfMemoryError as e:
  657. logger.error(f"[Outer Fold {outer_fold}] FINAL training OOM. err={str(e)[:200]}")
  658. cleanup_cuda()
  659. # minimal retry strategy: smaller batch size
  660. retry_bs = 32
  661. logger.warning(f"[Outer Fold {outer_fold}] Retry FINAL with batch_size={retry_bs} ...")
  662. train_loader = DataLoader(
  663. Subset(dataset, trainval_idx.tolist()),
  664. batch_size=retry_bs,
  665. shuffle=True,
  666. num_workers=cfg.num_workers,
  667. pin_memory=cfg.pin_memory,
  668. )
  669. test_loader = DataLoader(
  670. Subset(dataset, test_idx.tolist()),
  671. batch_size=retry_bs,
  672. shuffle=False,
  673. num_workers=cfg.num_workers,
  674. pin_memory=cfg.pin_memory,
  675. )
  676. optimizer = optim.Adam(model.parameters(), lr=best_params["lr"], weight_decay=best_params["weight_decay"])
  677. _ = train_one_run(
  678. model=model,
  679. criterion=criterion,
  680. optimizer=optimizer,
  681. train_loader=train_loader,
  682. val_loader=test_loader,
  683. device=device,
  684. num_epochs=cfg.final_num_epochs,
  685. patience=cfg.final_patience,
  686. run_name=f"[Outer {outer_fold}][FINAL][RETRY_BS32]",
  687. log_every_epoch=True,
  688. )
  689. metrics = evaluate_model(model, test_loader, device, labels=list(cfg.top2_labels))
  690. logger.info(f"[Outer Fold {outer_fold}] TEST metrics:")
  691. logger.info(f" ACC = {metrics['acc']:.4f}")
  692. logger.info(f" BACC = {metrics['bacc']:.4f}")
  693. logger.info(f" F1 = {metrics['f1']:.4f}")
  694. logger.info(f" Prec = {metrics['precision']:.4f}")
  695. logger.info(f" Rec = {metrics['recall']:.4f}")
  696. logger.info(" Classification Report:\n" + metrics["report"])
  697. logger.info(" Confusion Matrix:\n" + str(metrics["cm"]))
  698. ckpt_path = f"nested_groupkfold_verbose_fold{outer_fold}_model.pth"
  699. torch.save(model.state_dict(), ckpt_path)
  700. logger.info(f"[Outer Fold {outer_fold}] Saved model checkpoint: {ckpt_path}")
  701. outer_metrics["acc"].append(float(metrics["acc"]))
  702. outer_metrics["bacc"].append(float(metrics["bacc"]))
  703. outer_metrics["f1"].append(float(metrics["f1"]))
  704. outer_metrics["precision"].append(float(metrics["precision"]))
  705. outer_metrics["recall"].append(float(metrics["recall"]))
  706. del model, optimizer, criterion, train_loader, test_loader
  707. cleanup_cuda()
  708. # ----- Summary -----
  709. logger.info("=" * 90)
  710. logger.info("Nested GroupKFold Summary (Outer-fold variability)")
  711. for k in ["acc", "bacc", "f1", "precision", "recall"]:
  712. mean, std, ci = mean_std_ci(outer_metrics[k])
  713. logger.info(f"{k.upper():>9s}: mean={mean:.4f}, std={std:.4f}, 95%CI=({ci[0]:.4f}, {ci[1]:.4f})")
  714. logger.info("-" * 90)
  715. for k in ["acc", "bacc", "f1", "precision", "recall"]:
  716. logger.info(f"{k.upper():>9s} per-fold: {['{:.4f}'.format(x) for x in outer_metrics[k]]}")
  717. logger.info("Done.")
  718. if __name__ == "__main__":
  719. main(CFG)

model_B.py at commit 93a8363, under MIT · at the source

Overview

Authors: Chunfeng Yang1, Jiangwei Yu2, Aonan He1, Wentao Xiang3, Xi Wang1, Guangquan Zhou2, Yudong Zhang1, Miao Cao4, Yang Chen1, Juan M Gorriz5
  1. Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Jiangsu Provincial Joint International Research Laboratory of Medical Information Processing, School of Computer Science and Engineering, Southeast University and Centre de Recherche en Information Biomédicale Sino-français (CRIBs), 2 Sipailou, Nanjing, 210096 Jiangsu China
  2. School of Biological Science and Medical Engineering, Southeast University, Nanjing, Jiangsu 210096 China
  3. Jiangsu Province Engineering Research Center for Smart Wearable and Rehabilitation Devices, School of Biomedical Engineering and Informatics, Nanjing Medical University, Nanjing, 211166 Jiangsu China
  4. School of Health Sciences, Swinburne University of Technology, John Street, Hawthorn, Victoria 3122 Australia
  5. Data Science and Computational Intelligence Institute, University of Granada, Granada, 52005 Spain
Journal: Scientific data, volume 13, issue 1, article 807
Dates: received 19 September 2025; accepted 19 March 2026; published online 9 April 2026
Type: Data paper · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41597-026-07120-7 · PMID 41957382 · PMCID PMC13223301 · OpenAlex W7152408264
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: EEG (modality), human (organism), methods / tools (subfield)
Methods: Spectral & time-frequency, Preprocessing, Physiology & signal measures
Keywords: Data acquisition, Data publication and archiving, Machine learning, Neural decoding
MeSH: Artifacts*, Electroencephalography*, Convolutional Neural Networks, Humans, Movement (* major topic)
Topic: EEG and Brain-Computer Interfaces (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Citations: not cited yet (Europe PMC); 34 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.

JiangweiYu221/PhysioMotion_Artifact

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 93a8363d2ec62604f120ffde14c85b6b3112134d, 8 March 2026
Languages: Python (6)
Size: 9 files, 6 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (5 files), MNE-BIDS (3 files), pandas (3 files), Matplotlib (2 files), MNE-Python (2 files), PyTorch (2 files), scikit-learn (2 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
8 files

neuracle/neuracle-api

License: LGPL-2.1
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 4477c996c50d5e1c28cde702c2be52e3f6c80002, 8 December 2022
Languages: MATLAB (7), Python (7)
Size: 16 files, 14 scripts
Software Heritage: archived
Found in: the text, “Usage Notes”
Holds: license file
Not found: README, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (4 files), Psychtoolbox (2 files), Matplotlib (1 file), MNE-Python (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
15 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41597-026-07120-7.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 20 scripts, each with its path and the digest of its content;
  • 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41597-026-07120-7.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 4 keywords, 5 MeSH terms, 2 funders, 24 references.

Cite

This paper

Yang, C., Yu, J., He, A., Xiang, W., Wang, X., Zhou, G., Zhang, Y., Cao, M., Chen, Y., & Gorriz, J. M. (2026). PhysioMotion Artifact: A task-driven EEG dataset with point-wise motion artifact annotations. Scientific data, 13(1), 807. https://doi.org/10.1038/s41597-026-07120-7

BibTeX

@article{yang2026physiomotion,
author = {Yang, Chunfeng and Yu, Jiangwei and He, Aonan and Xiang, Wentao and Wang, Xi and Zhou, Guangquan and Zhang, Yudong and Cao, Miao and Chen, Yang and Gorriz, Juan M},
title = {{PhysioMotion Artifact: A task-driven EEG dataset with point-wise motion artifact annotations}},
journal = {Scientific data},
year = {2026},
month = apr,
volume = {13},
number = {1},
pages = {807},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/s41597-026-07120-7},
url = {https://doi.org/10.1038/s41597-026-07120-7},
pmid = {41957382},
pmcid = {PMC13223301}
}

RIS

TY - JOUR
AU - Yang, Chunfeng
AU - Yu, Jiangwei
AU - He, Aonan
AU - Xiang, Wentao
AU - Wang, Xi
AU - Zhou, Guangquan
AU - Zhang, Yudong
AU - Cao, Miao
AU - Chen, Yang
AU - Gorriz, Juan M
TI - PhysioMotion Artifact: A task-driven EEG dataset with point-wise motion artifact annotations
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/04/09
VL - 13
IS - 1
SP - 807
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/s41597-026-07120-7
UR - https://doi.org/10.1038/s41597-026-07120-7
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41597-026-07120-7",
"type": "article-journal",
"title": "PhysioMotion Artifact: A task-driven EEG dataset with point-wise motion artifact annotations",
"container-title": "Scientific data",
"author": [
{
"family": "Yang",
"given": "Chunfeng"
},
{
"family": "Yu",
"given": "Jiangwei"
},
{
"family": "He",
"given": "Aonan"
},
{
"family": "Xiang",
"given": "Wentao"
},
{
"family": "Wang",
"given": "Xi"
},
{
"family": "Zhou",
"given": "Guangquan"
},
{
"family": "Zhang",
"given": "Yudong"
},
{
"family": "Cao",
"given": "Miao"
},
{
"family": "Chen",
"given": "Yang"
},
{
"family": "Gorriz",
"given": "Juan M"
}
],
"container-title-short": "Sci Data",
"volume": "13",
"issue": "1",
"page": "807",
"DOI": "10.1038/s41597-026-07120-7",
"PMID": "41957382",
"PMCID": "PMC13223301",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41597-026-07120-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
9
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41597-025-05174-7 [code]
A large-scale MEG and EEG dataset for object recognition in naturalistic scenes
Journal: n/a
In common: MNE-BIDS, Psychtoolbox, MNE-Python, 5 other tools, methods / tools, EEG, 1 reference
[2] doi:10.1093/cercor/bhag113 [code]
Long-term reliability and stability of parameterized resting state EEG: evidence from a five-year follow-up.
Journal: Cerebral cortex (New York, N.Y. : 1991)
In common: MNE-BIDS, MNE-Python, scikit-learn, 3 other tools, methods / tools, EEG, 2 references
[3] doi:10.1038/s41598-026-56070-y [code]
SSDLabeler: realistic semi-synthetic data generation for multi-label artifact classification in EEG.
Journal: Scientific reports
In common: PyTorch, scikit-learn, pandas, 2 other tools, EEG, 3 references
[4] doi:10.1007/s12021-026-09803-3 [code]
NeuroFusion: A Unified Framework for Generalized Visual Stimulus Decoding from fMRI Across Datasets and Subjects.
Journal: Neuroinformatics
In common: MNE-BIDS, MNE-Python, PyTorch, 4 other tools, methods / tools
[5] doi:10.3390/bioengineering13070820 [code]
Benchmarking Multimodal Workload Classification: Effects of Modality, Validation Protocol, and Segmentation Contrast on an Open Graded-Arithmetic Dataset.
Journal: Bioengineering (Basel, Switzerland)
In common: MNE-Python, PyTorch, scikit-learn, 3 other tools, methods / tools, EEG, 2 references
[6] doi:10.1371/journal.pone.0343722 [code]
Comprehensive methodology for sample enrichment in EEG biomarker studies for Alzheimer's risk classification.
Journal: PloS one
In common: MNE-BIDS, MNE-Python, scikit-learn, 3 other tools, EEG, 1 reference
[7] doi:10.1097/j.pain.0000000000004044 [code]
No effect of rhythmic visual stimulation on experimental pain perception.
Journal: Pain
In common: MNE-BIDS, MNE-Python, pandas, 2 other tools, EEG, 2 references
[8] doi:10.3758/s13428-026-02997-z [code]
PyLossless: A non-destructive EEG processing pipeline.
Journal: Behavior research methods
In common: MNE-BIDS, MNE-Python, pandas, 1 other tool, methods / tools, EEG, 2 references
[9] doi:10.1038/s41598-026-47627-y [code]
QuantumNeuroXAI: a quantum-inspired deep learning framework with explainability for brain signal analysis and neurological disorder detection.
Journal: Scientific reports
In common: MNE-Python, PyTorch, scikit-learn, 3 other tools, methods / tools, EEG, 1 reference
[10] doi:10.1038/s41597-026-07377-y [code]
An open-access multi-site fMRI dataset for investigating conscious visual perception.
Journal: Scientific data
In common: MNE-BIDS, MNE-Python, scikit-learn, 3 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.