OSCR

Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Methods › Classifiers ↔ ibc_benchmark/models.py, lines 22–75 · score 0.91 · convolutional blocks, pooling layer, max pooling, batch normalisation, ReLU, SpectralCNN
  2. [2] § Results › Additional learned-model analysis ↔ experiments/mlp_raw_vs_combined_clean.py, lines 1–28 · score 0.74 · feature fusion, Raw MLP, model capacity, Combined MLP, raw spectra, seed
  3. [3] § Methods › Performance metrics ↔ notebooks/06_synthetic_roc_from_experiments.ipynb, the whole file · a weak match · score 0.71 · synthetic scores, ROC curves, calibrated, KNN, metrics, accuracy
  4. [4] § Methods › Classifiers ↔ notebooks/05_embedded_profiling_and_reports.ipynb, lines 341–419 · score 0.69 · AdamW, weight decay, ReLU, lightweight, dropout, batch
  5. [5] § Methods › Feature extraction ↔ ibc_benchmark/features.py, lines 35–80 · score 0.69 · discrete wavelet transform, standard deviation, coefficient, entropy, periodisation, energy
  6. [6] § Methods › Classifiers ↔ notebooks/05_embedded_profiling_and_reports.ipynb, lines 341–419 · score 0.67 · AdamW, weight decay, ReLU, hidden, Dropout, MLP
  7. [7] § Results › Additional learned-model analysis ↔ experiments/mlp_raw_vs_combined.py, lines 1–29 · score 0.65 · feature fusion, model capacity, raw spectra, matching, MLP, seed
  8. [8] § Methods › Classifiers ↔ ibc_benchmark/models.py, lines 78–101 · score 0.63 · hidden layer, ReLU, activations, network, Dropout, MLP
  9. [9] § Methods › Performance metrics ↔ models/train_evaluate.py, lines 50–82 · score 0.63 · ROC curves, FN, FP, TN, TP, EER
  10. [10] § Methods › Pre-processing ↔ ibc_benchmark/data.py, lines 79–144 · score 0.63 · experiment_id, remove rows, outliers, freq, filtered
  11. [11] § Methods › Classifiers ↔ models.py, lines 1–31 · score 0.63 · LightGBM, SpectralCNN, Random Forest, KNN, embedded, MLP
  12. [12] § Methods › MCU latency and energy profiling ↔ notebooks/05_embedded_profiling_and_reports.ipynb, lines 1–79 · score 0.62 · Cortex M4, board, voltage, profile, MCU, device
  13. [13] § Methods › Leakage-free evaluation protocol ↔ experiments/mlp_raw_vs_combined.py, lines 1–29 · score 0.57 · subject_id, enrollment, protocol, query, seeds, training
  14. [14] § Methods › Leakage-free evaluation protocol ↔ ibc_benchmark/data.py, lines 147–173 · score 0.55 · GroupShuffleSplit, subject_id, seeds, Leakage, training

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 444 lines · 17 KB · MIT · 3 matches

  1. # %%
  2. # Notebook: Frequency Analysis & On-Device Simulation (Cortex-M4) — Improved
  3. # Python >= 3.9
  4. # Requires: numpy, pandas, pywt, matplotlib, seaborn, scikit-learn, scipy, (optional) torch
  5. # =========================
  6. # Cell 0: Environment & Config
  7. # =========================
  8. import os
  9. from pathlib import Path
  10. from typing import Tuple, Dict, List, Optional
  11. import numpy as np
  12. import pandas as pd
  13. import matplotlib.pyplot as plt
  14. import seaborn as sns
  15. import pywt
  16. from sklearn.ensemble import RandomForestClassifier
  17. from sklearn.model_selection import StratifiedShuffleSplit
  18. from sklearn.preprocessing import StandardScaler
  19. from sklearn.decomposition import PCA
  20. from sklearn.metrics import accuracy_score
  21. from scipy import signal
  22. sns.set_context("talk")
  23. sns.set_style("whitegrid")
  24. SEED = 42
  25. rng = np.random.default_rng(SEED)
  26. CONFIG = {
  27. # Data & processing
  28. "data_dir": "data",
  29. "proc_dir": "data/processed",
  30. "results_dir": "results",
  31. "fig_dir": "results/figures",
  32. "out_len": 256, # resampled spectrum length
  33. "downsample_to": None, # e.g., 128 or None
  34. "standardize": True, # z-score with train stats
  35. "pca_components": None, # e.g., 128 or None (applied to raw only)
  36. "val_ratio": 0.2,
  37. # DWT
  38. "dwt_wavelet": "db4",
  39. "dwt_level": 2,
  40. "cache_dwt": True,
  41. # Visualization
  42. "scalogram_num": 6, # how many validation samples to visualize
  43. "scalogram_scales": 64,
  44. "scalogram_wavelet": "morl",
  45. "topk_importance": 20,
  46. # Baseline models
  47. "run_rf": True,
  48. "rf_n_estimators": 500,
  49. "rf_n_jobs": -1,
  50. "run_mlp": False, # optional quick MLP on combined features
  51. "mlp_hidden": 256,
  52. "mlp_epochs": 300,
  53. "mlp_lr": 5e-4,
  54. "mlp_wd": 1e-3,
  55. # MCU simulation (Cortex-M4 defaults; tune per board)
  56. "mcu_clock_hz": 96e6,
  57. "mcu_voltage_v": 3.3,
  58. "mcu_current_a": 0.035,
  59. # Reproducibility
  60. "seed": SEED,
  61. }
  62. DATA_DIR = Path(CONFIG["data_dir"])
  63. PROC_DIR = Path(CONFIG["proc_dir"])
  64. RES_DIR = Path(CONFIG["results_dir"])
  65. FIG_DIR = Path(CONFIG["fig_dir"])
  66. for p in [PROC_DIR, RES_DIR, FIG_DIR]:
  67. p.mkdir(parents=True, exist_ok=True)
  68. # =========================
  69. # Cell 1: Data Loading with Fallback
  70. # =========================
  71. def parse_gain_columns(df: pd.DataFrame, family_prefix: str) -> Tuple[np.ndarray, List[str]]:
  72. """Parse rx_gain columns and return (freqs, column_names) sorted by frequency."""
  73. cols, freqs = [], []
  74. for c in df.columns:
  75. if c.startswith(family_prefix + "_f_"):
  76. f = c.split("_f_")[-1]
  77. if f.isdigit():
  78. freqs.append(float(f)); cols.append(c)
  79. order = np.argsort(freqs)
  80. return np.asarray(freqs, dtype=float)[order], [cols[i] for i in order]
  81. def select_gain_family(df: pd.DataFrame) -> Tuple[str, np.ndarray, List[str]]:
  82. """Choose gain family with lower zero/NaN ratio across rows."""
  83. f1, c1 = parse_gain_columns(df, "rx_gain_50")
  84. f2, c2 = parse_gain_columns(df, "rx_gain_1M")
  85. def score(cols):
  86. if len(cols) == 0: return np.inf
  87. vals = df[cols].replace([np.inf, -np.inf], np.nan)
  88. return float(((vals == 0) | vals.isna()).sum(axis=1).mean())
  89. s1, s2 = score(c1), score(c2)
  90. if s1 < s2 and len(c1) > 0:
  91. return "rx_gain_50", f1, c1
  92. return "rx_gain_1M", f2, c2
  93. def resample_spectrum(row_vals: np.ndarray, in_freqs: np.ndarray, out_len: int = 256) -> np.ndarray:
  94. """Linear interpolation to fixed length; expects dB-domain magnitudes."""
  95. fmin, fmax = float(in_freqs.min()), float(in_freqs.max())
  96. f_out = np.linspace(fmin, fmax, out_len, dtype=float)
  97. return np.interp(f_out, in_freqs, row_vals.astype(float))
  98. def load_or_build_processed(out_len: int = 256) -> Tuple[np.ndarray, np.ndarray]:
  99. """Load processed X(out_len) and y, or build from all_measurements.csv."""
  100. x_csv = PROC_DIR / "ibc_processed.csv"
  101. y_csv = DATA_DIR / "labels_filtered.csv"
  102. if x_csv.exists() and y_csv.exists():
  103. X = pd.read_csv(x_csv).values.astype(np.float32)
  104. y = pd.read_csv(y_csv)["subject_id"].astype(int).values
  105. return X, y
  106. raw_csvs = list(DATA_DIR.glob("all_measurements.csv"))
  107. if not raw_csvs:
  108. raise FileNotFoundError("Provide data/processed/ibc_processed.csv or data/all_measurements.csv")
  109. df = pd.read_csv(raw_csvs[0])
  110. assert "subject_id" in df.columns, "subject_id missing in all_measurements.csv"
  111. family, freqs, cols = select_gain_family(df)
  112. vals = df[cols].replace([np.inf, -np.inf], np.nan)
  113. mask = ~vals.isna().any(axis=1)
  114. df = df.loc[mask].reset_index(drop=True)
  115. vals = vals.loc[mask].reset_index(drop=True)
  116. y = df["subject_id"].astype(int).values
  117. X = np.vstack([resample_spectrum(vals.iloc[i].values, freqs, out_len=out_len)
  118. for i in range(len(vals))]).astype(np.float32)
  119. pd.DataFrame(X, columns=[f"f{i}" for i in range(X.shape[1])]).to_csv(x_csv, index=False)
  120. pd.DataFrame({"subject_id": y}).to_csv(y_csv, index=False)
  121. return X, y
  122. X_raw, y = load_or_build_processed(out_len=CONFIG["out_len"])
  123. print("Processed:", X_raw.shape, y.shape)
  124. # =========================
  125. # Cell 2: Split & Preprocess (leakage-free)
  126. # =========================
  127. sss = StratifiedShuffleSplit(n_splits=1, test_size=CONFIG["val_ratio"], random_state=CONFIG["seed"])
  128. tr_idx, va_idx = next(sss.split(X_raw, y))
  129. X_tr_raw, X_va_raw = X_raw[tr_idx], X_raw[va_idx]
  130. y_tr, y_va = y[tr_idx], y[va_idx]
  131. print("Split:", X_tr_raw.shape, X_va_raw.shape, len(np.unique(y_tr)), len(np.unique(y_va)))
  132. # Optional downsample (e.g., 256 -> 128) BEFORE standardization
  133. def maybe_downsample(X: np.ndarray, to_len: Optional[int]) -> np.ndarray:
  134. if to_len is None or to_len == X.shape[1]:
  135. return X
  136. # Use Fourier resampling for smooth decimation
  137. return signal.resample(X, num=to_len, axis=1)
  138. X_tr_ds = maybe_downsample(X_tr_raw, CONFIG["downsample_to"])
  139. X_va_ds = maybe_downsample(X_va_raw, CONFIG["downsample_to"])
  140. print("After downsample:", X_tr_ds.shape, X_va_ds.shape)
  141. # Standardization with train stats
  142. if CONFIG["standardize"]:
  143. scaler = StandardScaler(with_mean=True, with_std=True)
  144. X_tr = scaler.fit_transform(X_tr_ds)
  145. X_va = scaler.transform(X_va_ds)
  146. else:
  147. X_tr, X_va = X_tr_ds, X_va_ds
  148. # Optional PCA on RAW (retain DWT over standardized raw)
  149. def maybe_pca_fit_transform(Xtr: np.ndarray, Xva: np.ndarray, n_comp: Optional[int]):
  150. if n_comp is None:
  151. return Xtr, Xva, None
  152. pca = PCA(n_components=n_comp, random_state=CONFIG["seed"])
  153. return pca.fit_transform(Xtr), pca.transform(Xva), pca
  154. X_tr_raw_pca, X_va_raw_pca, pca_model = maybe_pca_fit_transform(X_tr, X_va, CONFIG["pca_components"])
  155. print("Raw/PCA shapes:", X_tr.shape, X_va.shape, X_tr_raw_pca.shape, X_va_raw_pca.shape)
  156. # =========================
  157. # Cell 3: DWT(db4, L2) feature extraction (with caching)
  158. # =========================
  159. def dwt_stats(X: np.ndarray, wavelet: str = "db4", level: int = 2) -> np.ndarray:
  160. """Compute DWT stats per sample: energy, entropy, mean, std for each coeff array."""
  161. feats = []
  162. for row in X:
  163. coeffs = pywt.wavedec(row, wavelet, level=level, mode="periodization")
  164. fvec = []
  165. for c in coeffs:
  166. e = float(np.sum(c**2))
  167. p = (c**2) / (e + 1e-12)
  168. ent = float(-np.sum(p * np.log2(np.clip(p, 1e-12, 1.0))))
  169. fvec.extend([e, ent, float(np.mean(c)), float(np.std(c))])
  170. feats.append(fvec)
  171. return np.asarray(feats, dtype=np.float32)
  172. def dwt_cached(Xtr: np.ndarray, Xva: np.ndarray, wavelet: str, level: int) -> Tuple[np.ndarray, np.ndarray]:
  173. key = f"dwt_{wavelet}_L{level}_std{int(CONFIG['standardize'])}_len{Xtr.shape[1]}"
  174. ftr = PROC_DIR / f"{key}_train.npy"
  175. fva = PROC_DIR / f"{key}_val.npy"
  176. if CONFIG["cache_dwt"] and ftr.exists() and fva.exists():
  177. return np.load(ftr), np.load(fva)
  178. X_tr_dwt = dwt_stats(Xtr, wavelet=wavelet, level=level)
  179. X_va_dwt = dwt_stats(Xva, wavelet=wavelet, level=level)
  180. if CONFIG["cache_dwt"]:
  181. np.save(ftr, X_tr_dwt); np.save(fva, X_va_dwt)
  182. return X_tr_dwt, X_va_dwt
  183. X_tr_dwt, X_va_dwt = dwt_cached(X_tr, X_va, CONFIG["dwt_wavelet"], CONFIG["dwt_level"])
  184. print("DWT shapes:", X_tr_dwt.shape, X_va_dwt.shape)
  185. # Combined features (raw/PCA + DWT)
  186. X_tr_comb = np.hstack([X_tr_raw_pca, X_tr_dwt])
  187. X_va_comb = np.hstack([X_va_raw_pca, X_va_dwt])
  188. print("Combined shapes:", X_tr_comb.shape, X_va_comb.shape)
  189. # =========================
  190. # Cell 4: CWT Scalogram Grid
  191. # =========================
  192. def cwt_scalogram(x: np.ndarray, wavelet: str = "morl", num_scales: int = 64) -> np.ndarray:
  193. scales = np.linspace(1, num_scales, num_scales)
  194. coef, freqs = pywt.cwt(x, scales, wavelet)
  195. power = np.abs(coef)**2
  196. return power.astype(np.float32)
  197. def save_scalogram_grid(Xv: np.ndarray, k: int, wavelet: str, num_scales: int, out_path: Path):
  198. idx = rng.choice(np.arange(Xv.shape[0]), size=min(k, Xv.shape[0]), replace=False)
  199. n = len(idx); cols = min(3, n); rows = int(np.ceil(n / cols))
  200. fig, axes = plt.subplots(rows, cols, figsize=(4*cols, 3.2*rows), squeeze=False)
  201. for ax, i in zip(axes.flat, idx):
  202. S = cwt_scalogram(Xv[i], wavelet=wavelet, num_scales=num_scales)
  203. sns.heatmap(S, cmap="magma", cbar=False, ax=ax)
  204. ax.set_title(f"val idx {i}")
  205. ax.set_xlabel("Freq idx"); ax.set_ylabel("Scales")
  206. for ax in axes.flat[n:]:
  207. ax.axis("off")
  208. fig.suptitle("CWT Scalogram Grid")
  209. fig.tight_layout()
  210. fig.savefig(out_path, dpi=200)
  211. plt.close(fig)
  212. save_scalogram_grid(X_va, CONFIG["scalogram_num"], CONFIG["scalogram_wavelet"],
  213. CONFIG["scalogram_scales"], FIG_DIR / "scalogram_grid.png")
  214. # =========================
  215. # Cell 5: RF Feature Importance on DWT stats
  216. # =========================
  217. if CONFIG["run_rf"]:
  218. rf = RandomForestClassifier(
  219. n_estimators=CONFIG["rf_n_estimators"], max_depth=None,
  220. random_state=SEED, n_jobs=CONFIG["rf_n_jobs"]
  221. )
  222. rf.fit(X_tr_dwt, y_tr)
  223. imp = rf.feature_importances_
  224. order = np.argsort(imp)[::-1]
  225. topk = CONFIG["topk_importance"]
  226. # Bar plot
  227. plt.figure(figsize=(10, 4))
  228. plt.bar(np.arange(min(topk, len(imp))), imp[order][:topk])
  229. plt.title(f"Top-{topk} DWT Feature Importances (RF Gini)")
  230. plt.xlabel("DWT feature index (energy/entropy/mean/std per coeff)")
  231. plt.ylabel("Gini importance")
  232. plt.tight_layout()
  233. plt.savefig(FIG_DIR / "dwt_importances_topk.png", dpi=200)
  234. plt.close()
  235. # CSV of importances
  236. pd.DataFrame({
  237. "feat_idx": order,
  238. "importance": imp[order]
  239. }).to_csv(RES_DIR / "dwt_feature_importances.csv", index=False)
  240. # =========================
  241. # Cell 6: MCU On-Device Simulation (Cortex-M4) — Stage-wise
  242. # =========================
  243. def cycles_zscore(n: int, d: int) -> int:
  244. # simple model: subtract mean, divide std, load/store per element
  245. C_ADD, C_DIV, C_MEM = 1, 3, 2
  246. return int(n * d * (C_ADD + C_DIV + C_MEM))
  247. def cycles_dwt_db4_l2(n: int, d: int) -> int:
  248. # first-order estimator: lifting-like passes across ~2d with aggregated constants
  249. C_MUL, C_ADD, C_MEM = 1, 1, 2
  250. a = 20 # aggregated factor (tunable)
  251. dwt = int(a * 2 * d * (C_MUL + C_ADD + C_MEM))
  252. stats = int(10 * d * (C_ADD + C_MUL + C_MEM)) # energy/entropy/mean/std
  253. return (dwt + stats) * n
  254. def time_energy(cycles: int, f_hz: float, v: float, i: float) -> Tuple[float, float]:
  255. t = cycles / f_hz
  256. e = v * i * t
  257. return t, e
  258. def simulate_pipeline(n_samples: int, d_spec: int, clk: float, v: float, i: float) -> Dict[str, float]:
  259. c_z = cycles_zscore(n_samples, d_spec) if CONFIG["standardize"] else 0
  260. c_dwt = cycles_dwt_db4_l2(n_samples, d_spec)
  261. c_tot = c_z + c_dwt
  262. t, e = time_energy(c_tot, clk, v, i)
  263. return {
  264. "cycles_total": float(c_tot),
  265. "time_s": float(t),
  266. "energy_J": float(e),
  267. "cycles_zscore": float(c_z),
  268. "cycles_dwt": float(c_dwt),
  269. "clock_hz": float(clk),
  270. "voltage_v": float(v),
  271. "current_a": float(i),
  272. "n_samples": int(n_samples),
  273. "d_spec": int(d_spec),
  274. }
  275. mcu = simulate_pipeline(
  276. n_samples=1,
  277. d_spec=X_tr.shape[1],
  278. clk=CONFIG["mcu_clock_hz"],
  279. v=CONFIG["mcu_voltage_v"],
  280. i=CONFIG["mcu_current_a"],
  281. )
  282. print("MCU Simulation (1 sample):", mcu)
  283. # Also simulate a small batch (e.g., 50 samples)
  284. mcu_batch = simulate_pipeline(
  285. n_samples=50,
  286. d_spec=X_tr.shape[1],
  287. clk=CONFIG["mcu_clock_hz"],
  288. v=CONFIG["mcu_voltage_v"],
  289. i=CONFIG["mcu_current_a"],
  290. )
  291. # =========================
  292. # Cell 7: Lightweight Baselines (RF combined, optional MLP)
  293. # =========================
  294. results = []
  295. if CONFIG["run_rf"]:
  296. clf = RandomForestClassifier(n_estimators=CONFIG["rf_n_estimators"], random_state=SEED, n_jobs=CONFIG["rf_n_jobs"])
  297. clf.fit(X_tr_comb, y_tr)
  298. y_pred = clf.predict(X_va_comb)
  299. acc = accuracy_score(y_va, y_pred)
  300. print("RF(combined) val accuracy:", acc)
  301. results.append({"model": "rf_combined", "val_acc": float(acc)})
  302. if CONFIG["run_mlp"]:
  303. try:
  304. import torch, torch.nn as nn, torch.nn.functional as F
  305. from torch.utils.data import TensorDataset, DataLoader
  306. DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
  307. class MLP(nn.Module):
  308. def __init__(self, in_dim, n_classes, hidden=256, p=0.3):
  309. super().__init__()
  310. self.net = nn.Sequential(
  311. nn.Linear(in_dim, hidden),
  312. nn.ReLU(inplace=True),
  313. nn.Dropout(p),
  314. nn.Linear(hidden, hidden),
  315. nn.ReLU(inplace=True),
  316. nn.Dropout(p),
  317. nn.Linear(hidden, n_classes),
  318. )
  319. def forward(self, x): return self.net(x)
  320. C = len(np.unique(y_tr))
  321. D = X_tr_comb.shape[1]
  322. model = MLP(D, C, hidden=CONFIG["mlp_hidden"], p=0.3).to(DEVICE)
  323. opt = torch.optim.AdamW(model.parameters(), lr=CONFIG["mlp_lr"], weight_decay=CONFIG["mlp_wd"])
  324. crit = nn.CrossEntropyLoss()
  325. Xt = torch.from_numpy(X_tr_comb).float().to(DEVICE)
  326. yt = torch.from_numpy(pd.factorize(y_tr)[0]).long().to(DEVICE)
  327. Xv = torch.from_numpy(X_va_comb).float().to(DEVICE)
  328. yv = torch.from_numpy(pd.factorize(y_va, sort=True)[0]).long().to(DEVICE)
  329. # Align val encoding to train classes
  330. # Build mapping from original labels to 0..C-1 using train labels
  331. classes = np.unique(y_tr)
  332. to_index = {int(c): i for i, c in enumerate(classes)}
  333. y_tr_enc = np.vectorize(lambda t: to_index[int(t)])(y_tr)
  334. y_va_enc = np.vectorize(lambda t: to_index[int(t)])(y_va)
  335. yt = torch.from_numpy(y_tr_enc).long().to(DEVICE)
  336. yv = torch.from_numpy(y_va_enc).long().to(DEVICE)
  337. dl = DataLoader(TensorDataset(Xt, yt), batch_size=64, shuffle=True, drop_last=False)
  338. best_val = 1e9; best_state=None
  339. for epoch in range(1, CONFIG["mlp_epochs"]+1):
  340. model.train(); run=0.0
  341. for xb, yb in dl:
  342. opt.zero_grad(); logits = model(xb); loss = crit(logits, yb)
  343. loss.backward(); opt.step(); run += loss.item() * xb.size(0)
  344. tr_loss = run / len(dl.dataset)
  345. model.eval()
  346. with torch.no_grad():
  347. va_loss = crit(model(Xv), yv).item()
  348. preds = model(Xv).argmax(dim=1).cpu().numpy()
  349. va_acc = (preds == y_va_enc).mean()
  350. if va_loss < best_val:
  351. best_val = va_loss
  352. best_state = {k: v.detach().cpu().clone() for k, v in model.state_dict().items()}
  353. if epoch % 50 == 0 or epoch == 1:
  354. print(f"[MLP {epoch:03d}] tr={tr_loss:.4f} val={va_loss:.4f} acc={va_acc:.4f}")
  355. if best_state is not None:
  356. model.load_state_dict(best_state)
  357. with torch.no_grad():
  358. preds = model(Xv).argmax(dim=1).cpu().numpy()
  359. acc = (preds == y_va_enc).mean()
  360. print("MLP(combined) val accuracy:", acc)
  361. results.append({"model": "mlp_combined", "val_acc": float(acc)})
  362. except Exception as e:
  363. print("MLP failed:", e)
  364. # =========================
  365. # Cell 8: Save Reports
  366. # =========================
  367. report = {
  368. "n_train": int(X_tr.shape[0]), "n_val": int(X_va.shape[0]),
  369. "d_raw": int(X_tr.shape[1]), "d_dwt": int(X_tr_dwt.shape[1]),
  370. "d_combined": int(X_tr_comb.shape[1]),
  371. "standardize": bool(CONFIG["standardize"]),
  372. "downsample_to": int(CONFIG["downsample_to"]) if CONFIG["downsample_to"] is not None else None,
  373. "pca_components": int(CONFIG["pca_components"]) if CONFIG["pca_components"] is not None else None,
  374. "dwt_wavelet": CONFIG["dwt_wavelet"], "dwt_level": int(CONFIG["dwt_level"]),
  375. "rf_ran": bool(CONFIG["run_rf"]), "mlp_ran": bool(CONFIG["run_mlp"]),
  376. "mcu_sim_1": mcu, "mcu_sim_50": mcu_batch,
  377. }
  378. pd.DataFrame(results).to_csv(RES_DIR / "baseline_results.csv", index=False)
  379. pd.Series(report, dtype=object).to_json(RES_DIR / "ondevice_simulation_report.json", indent=2)
  380. print("Saved:")
  381. print(" - Figures: scalogram_grid.png, dwt_importances_topk.png")
  382. print(" - Reports: dwt_feature_importances.csv, baseline_results.csv, ondevice_simulation_report.json")
  383. # %%

05_embedded_profiling_and_reports.ipynb at commit 18235d5, under MIT · at the source

Overview

Authors: Seungmin Jin1, Mikhail M. Komarov1
  1. HSE University, Graduate School of Business,Moscow, Russia 101000
Journal: Scientific reports, volume 16, issue 1, article 24421
Dates: received 9 October 2025; accepted 25 May 2026; published online 28 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41598-026-55584-9 · PMID 42209669 · PMCID PMC13448553 · OpenAlex W7162668449
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: other (modality), human (organism), methods / tools (subfield)
Methods: Connectivity, Statistics, Machine learning, Preprocessing, fMRI & imaging
Keywords: Computational biology and bioinformatics, Engineering, Mathematics and computing
MeSH: Biometric Identification*, Wavelet Analysis*, Wearable Electronic Devices*, Algorithms, Benchmarking, Humans, Random Forest, Reproducibility of Results (* major topic)
Topic: Wireless Body Area Networks (Biomedical Engineering, Engineering), according to OpenAlex
Funding: Russian Science Foundation (RSF) (24-19-00299)
Citations: not cited yet (Europe PMC); 31 references in the paper

Abstract

Intrabody communication (IBC) channels offer physiological diversity that may support future wearable biometric identification. Recent reports of over 99 per cent identification accuracy have frequently resulted from data leakage, where samples from the same subject are seen in both training and evaluation, yielding inflated and unreliable metrics. In this work, we establish a public, leakage-free benchmark for IBC biometrics built on a 30-subject open dataset, using strict subject-wise 80/20 splits repeated five times to ensure reproducibility. We systematically compare frequency-domain and time-frequency representations, including resampled spectra, discrete wavelet transform (DWT) statistics, and their fusion. Under the subject-wise embedded-friendly benchmark, the strongest classical configuration, Scattering + LightGBM, reaches 54.0 per cent accuracy, while db4-DWT and lifting-based wavelet statistics with Random Forest improve over the Simple-3 baseline (49.3 and 51.6 per cent versus 39.0 per cent). Separately, closed-set neural analyses provide exploratory upper bounds rather than leakage-free subject-wise results: a Raw MLP reaches 83.7 per cent accuracy, whereas adding DWT statistics does not improve this result (81.2 per cent for Combined MLP), and SpectralCNN reaches 74 per cent. Confusion matrix analysis reveals that residual errors are concentrated among subject pairs with statistically overlapping signatures, suggesting the presence of intrinsically hard users and a potential biometric ceiling for this modality. Embedded profiling on an STM32F446RE Cortex-M4 microcontroller indicates that lifting-based wavelet features enable low-latency, low-energy scoring, requiring approximately 0.55 ms and 18 micro-J per 256-point spectrum for Lift-bior feature extraction plus Random Forest inference (versus approx. 33 micro-J for the equivalent db4-DWT pipeline). All code, data split scripts, and Jupyter notebooks are released open source to facilitate reproducibility and enable rigorous future comparisons.

Supplementary Information: The online version contains supplementary material available at https://doi.org/10.1038/s41598-026-55584-9.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

dryjins/ibc-wavelet-benchmark

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 18235d5e195935711400e9c02754e5241f83efcd, 14 May 2026
Languages: Python (20), Jupyter (8)
Size: 72 files, 28 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (requirements.txt), 8 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (23 files), pandas (20 files), scikit-learn (17 files), PyTorch (10 files), SciPy (7 files), Matplotlib (6 files), PyWavelets (6 files), seaborn (3 files), LightGBM (2 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
30 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 28 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The dataset analysed during this study is publicly available in the Zenodo repository at https://zenodo.org/records/8214497. All source code, unit tests, and notebooks required to reproduce the experiments and figures are openly available in a public GitHub repository at https://github.com/dryjins/ibc-wavelet-benchmark. This research was supported in part through computational resources of HPC facilities at HSE University.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 3 keywords, 8 MeSH terms, 1 funder, 27 references.

Cite

This paper

Jin, S., & Komarov, M. M. (2026). Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification. Scientific reports, 16(1), 24421. https://doi.org/10.1038/s41598-026-55584-9

BibTeX

@article{jin2026reproducible,
author = {Jin, Seungmin and Komarov, Mikhail M.},
title = {{Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification}},
journal = {Scientific reports},
year = {2026},
month = may,
volume = {16},
number = {1},
pages = {24421},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-55584-9},
url = {https://doi.org/10.1038/s41598-026-55584-9},
pmid = {42209669},
pmcid = {PMC13448553}
}

RIS

TY - JOUR
AU - Jin, Seungmin
AU - Komarov, Mikhail M.
TI - Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/05/28
VL - 16
IS - 1
SP - 24421
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-55584-9
UR - https://doi.org/10.1038/s41598-026-55584-9
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-55584-9",
"type": "article-journal",
"title": "Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification",
"container-title": "Scientific reports",
"author": [
{
"family": "Jin",
"given": "Seungmin"
},
{
"family": "Komarov",
"given": "Mikhail M."
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "24421",
"DOI": "10.1038/s41598-026-55584-9",
"PMID": "42209669",
"PMCID": "PMC13448553",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-55584-9",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
28
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pgen.1012242 [code]
Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.
Journal: PLoS genetics
In common: LightGBM, PyTorch, seaborn, 5 other tools, 1 reference
[2] doi:10.1371/journal.pone.0354510 [code]
Personalized adaptive virtual reality experience driven by electroencephalography-based pain recognition.
Journal: PloS one
In common: LightGBM, PyWavelets, seaborn, 4 other tools
[3] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: LightGBM, PyTorch, seaborn, 5 other tools, methods / tools
[4] doi:10.1002/dad2.70430 [code]
Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research.
Journal: Alzheimer's & dementia (Amsterdam, Netherlands)
In common: LightGBM, PyTorch, seaborn, 5 other tools, methods / tools
[5] doi:10.1038/s41598-026-68186-2 [code]
NeuroStream: spectral-spatio-temporal deep learning for visual stimulus classification from EEG.
Journal: Scientific reports
In common: PyWavelets, PyTorch, seaborn, 5 other tools, methods / tools
[6] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: PyWavelets, PyTorch, seaborn, 5 other tools, methods / tools
[7] doi:10.1186/s13244-026-02365-7 [code]
Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion.
Journal: Insights into imaging
In common: LightGBM, PyTorch, seaborn, 5 other tools
[8] doi:10.1038/s41467-026-73996-z [code]
Genetic architecture of white matter microstructure captured by unsupervised deep representation learning of fractional anisotropy maps.
Journal: Nature communications
In common: LightGBM, PyTorch, seaborn, 5 other tools
[9] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: LightGBM, PyTorch, seaborn, 5 other tools
[10] doi:10.1093/bib/bbag118 [code]
Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.
Journal: Briefings in bioinformatics
In common: LightGBM, PyTorch, seaborn, 5 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.