OSCR

Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches
  1. [1] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/modules/models.py, lines 19–35 · score 0.82 · sensor channel separability, ReLU, Cross sensor, axis, mixing, BN
  2. [2] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/modules/training.py, lines 21–63 · score 0.77 · cosine annealing, weight decay, AdamW, cross entropy, min, loss
  3. [3] § 2. Materials and Methods › 2.7. Domain–Adaptation Task Definitions ↔ code/eval_drift_taskB.py, lines 1–36 · score 0.74 · cumulative training scenario, DRCA fair, fine tuning, KD DM, target batch, held
  4. [4] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/analysis_main.py, lines 167–290 · score 0.74 · KD DM unsup, DRCA fair, fine tuning, unsupervised, pseudo, temperature
  5. [5] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/eval_compression.py, lines 31–129 · score 0.72 · ReLU, pointwise, FC, DSConv, activations, BN
  6. [6] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/run_litesensor_t1.py, lines 64–105 · score 0.72 · weight decay, AdamW, cross entropy, smoothing, cosine, validation
  7. [7] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_combo2.py, lines 259–338 · score 0.72 · stratified random, cross entropy, fine tuning, KD DM, target batch, KL
  8. [8] § 2. Materials and Methods › 2.5. Multi-Stage Compression › 2.5.1. INT8 Post-Training Quantization ↔ code/run_t3_qat.py, lines 1–31 · score 0.63 · Quantization aware training, QAT, fine tuning, PTQ, Post, weights
  9. [9] § 3. Results › 3.4. Sensor Drift Compensation › 3.4.1. Task A: Laboratory Re-Calibration Scenario ↔ code/run_kdm_v4.py, lines 115–239 · score 0.61 · KD DM unsup, DRCA fair, Random Forest, target batch, NC, chronological
  10. [10] § 3. Results › 3.4. Sensor Drift Compensation › 3.4.1. Task A: Laboratory Re-Calibration Scenario ↔ code/analysis_main.py, lines 167–290 · score 0.60 · KD DM unsup, DRCA fair, Random Forest, NC, chronological, fraction
  11. [11] § 3. Results › 3.3. CBRN Simulant Classification Performance ↔ code/eval_classification.py, lines 61–146 · score 0.59 · confusion matrix, macro F1, recall, CWA, precision, scores
  12. [12] § 2. Materials and Methods › 2.5. Multi-Stage Compression › 2.5.1. INT8 Post-Training Quantization ↔ code/eval_compression.py, lines 1–21 · score 0.57 · TensorFlow, INT8 quantized, activations, weights, Lite
  13. [13] § 2. Materials and Methods › 2.3. Preprocessing Pipeline ↔ code/modules/data_loader.py, lines 128–140 · score 0.54 · Feature selection, Random Forest, Gini, Batch
  14. [14] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_combo2.py, lines 259–338 · score 0.53 · soft target, fine tuning, KD DM, ensemble, logits, student
  15. [15] § 2. Materials and Methods › 2.7. Domain–Adaptation Task Definitions ↔ code/run_kdm_v4.py, lines 115–239 · score 0.53 · KD DM unsup, DRCA fair, target batch, NC, fraction
  16. [16] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_s1_30splits.py, lines 18–118 · score 0.53 · member ensemble, fine tuning, KD DM, supervision, student, teacher

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 431 lines · 18 KB · Apache-2.0 · 2 matches

  1. """Re-experiment orchestrator (CHUNK MODE) — Molecules MS-4274054 R1 revision.
  2. Sandbox bash call has a 45 s limit, so each invocation runs **ONE** chunk and
  3. exits, persisting state to disk. Subsequent calls resume.
  4. Chunk identifiers
  5. p2_s{seed_idx} : Phase 2 — all 4 archs × 1 seed × Batch 1
  6. p3_b{batch_id}_s{seed_idx} : Phase 3 — 1 target batch × 1 seed × all 5 conditions
  7. p4 : Phase 4 — compression breakdown + sanity + summary
  8. State files
  9. results/progress.json : list of completed chunks
  10. results/raw_results_benchmark.csv : appended after each p2 chunk
  11. results/raw_results_drift.csv : appended after each p3 chunk
  12. results/experiment_summary.json : written by p4
  13. Usage
  14. python analysis_main.py --resume # auto-pick next chunk
  15. python analysis_main.py --chunk p2_s0 --device cpu # explicit chunk
  16. python analysis_main.py --chunk p3_b2_s0
  17. python analysis_main.py --chunk p4
  18. """
  19. from __future__ import annotations
  20. import argparse
  21. import csv
  22. import json
  23. import os
  24. import random
  25. import sys
  26. import time
  27. from pathlib import Path
  28. from typing import List
  29. import numpy as np
  30. import pandas as pd
  31. import torch
  32. import yaml
  33. ROOT = Path(__file__).resolve().parent
  34. sys.path.insert(0, str(ROOT))
  35. from modules import data_loader as DL
  36. from modules import models as M
  37. from modules import compression as C
  38. from modules import kd_dm as KD
  39. from modules import evaluate as EV
  40. from modules import training as TR
  41. def set_seed(seed: int):
  42. random.seed(seed); np.random.seed(seed); torch.manual_seed(seed)
  43. torch.backends.cudnn.deterministic = True
  44. def reshape_for_sensors(X: np.ndarray) -> np.ndarray:
  45. return DL.reshape_for_sensors(X, n_sensors=16, n_feat_per_sensor=8)
  46. # =================================================================
  47. # Chunk planning
  48. # =================================================================
  49. def all_chunks(cfg: dict) -> List[str]:
  50. seeds = cfg["statistics"]["seeds"]
  51. chunks = [f"p2_s{i}" for i in range(len(seeds))]
  52. for bid in range(2, cfg["dataset"]["n_batches"] + 1):
  53. for i in range(len(seeds)):
  54. chunks.append(f"p3_b{bid}_s{i}")
  55. chunks.append("p4")
  56. return chunks
  57. def load_progress(path: Path) -> set:
  58. if path.exists():
  59. return set(json.loads(path.read_text()).get("done", []))
  60. return set()
  61. def save_progress(path: Path, done: set):
  62. path.write_text(json.dumps({"done": sorted(done),
  63. "ts": time.strftime("%Y-%m-%dT%H:%M:%S")},
  64. indent=2))
  65. # =================================================================
  66. # CSV append helpers
  67. # =================================================================
  68. BENCH_HEADERS = ["chunk_id", "seed", "model", "value", "macro_f1", "params",
  69. "size_fp32_kB", "size_int8_kB", "infer_ms", "train_s", "best_val_acc"]
  70. DRIFT_HEADERS = ["chunk_id", "seed", "target_batch", "method", "variant",
  71. "acc", "macro_f1", "n_test"]
  72. def append_csv(path: Path, headers: list, rows: list[dict]):
  73. new = not path.exists()
  74. with path.open("a", newline="") as f:
  75. w = csv.DictWriter(f, fieldnames=headers)
  76. if new: w.writeheader()
  77. for r in rows:
  78. w.writerow({k: r.get(k, "") for k in headers})
  79. # =================================================================
  80. # Phase 2 chunk
  81. # =================================================================
  82. def run_p2_chunk(seed_idx: int, cfg: dict, batches, device: str) -> list[dict]:
  83. seed = cfg["statistics"]["seeds"][seed_idx]
  84. set_seed(seed)
  85. b1 = batches[0]
  86. n = len(b1.y)
  87. n_test = max(2, int(0.30 * n))
  88. Xtr_full = b1.X[:n - n_test]; ytr_full = b1.y[:n - n_test]
  89. Xte = b1.X[n - n_test:]; yte = b1.y[n - n_test:]
  90. mu, sigma = DL.fit_standardiser(Xtr_full)
  91. Xtr_full = DL.standardise(Xtr_full, mu, sigma)
  92. Xte = DL.standardise(Xte, mu, sigma)
  93. idx = np.arange(len(ytr_full)); np.random.shuffle(idx); v = int(0.20 * len(idx))
  94. val_i, tr_i = idx[:v], idx[v:]
  95. Xtr = reshape_for_sensors(Xtr_full[tr_i]); ytr = ytr_full[tr_i] - 1
  96. Xv = reshape_for_sensors(Xtr_full[val_i]); yv = ytr_full[val_i] - 1
  97. Xtest = reshape_for_sensors(Xte); ytest = yte - 1
  98. model_specs = {
  99. "litesensor": cfg["models"]["liteSensor"],
  100. "mobilenet1d": cfg["models"]["mobileNet1D"],
  101. "inceptiontime1d": cfg["models"]["inceptionTime1D"],
  102. "resnet1d": cfg["models"]["resNet1D"],
  103. }
  104. n_classes = cfg["dataset"]["n_classes"]
  105. rows = []
  106. for name, mc in model_specs.items():
  107. model = M.build(name, n_classes=n_classes, in_ch=16, **mc)
  108. t0 = time.time()
  109. res = TR.train_supervised(
  110. model, Xtr, ytr, Xv, yv,
  111. epochs=cfg["training"]["epochs_source"],
  112. lr=cfg["training"]["lr"],
  113. batch_size=cfg["training"]["batch_size_source"],
  114. weight_decay=cfg["training"]["weight_decay"],
  115. device=device,
  116. )
  117. train_s = time.time() - t0
  118. model.load_state_dict(res.best_state)
  119. yhat = TR.predict(model, Xtest, device=device)
  120. t1 = time.time()
  121. for _ in range(3):
  122. _ = TR.predict(model, Xtest[:1], device=device)
  123. inf_ms = (time.time() - t1) / 3 * 1000
  124. n_params = M.count_parameters(model)
  125. rows.append({
  126. "chunk_id": f"p2_s{seed_idx}",
  127. "seed": seed, "model": name,
  128. "value": EV.accuracy(ytest, yhat),
  129. "macro_f1": EV.macro_f1(ytest, yhat),
  130. "params": n_params,
  131. "size_fp32_kB": round(n_params * 4 / 1024, 2),
  132. "size_int8_kB": round(n_params / 1024, 2),
  133. "infer_ms": round(inf_ms, 2),
  134. "train_s": round(train_s, 1),
  135. "best_val_acc": res.best_val_acc,
  136. })
  137. print(f" {name:18s} acc={rows[-1]['value']:.4f} f1={rows[-1]['macro_f1']:.4f} "
  138. f"params={n_params:,} inf={inf_ms:.2f}ms train={train_s:.1f}s")
  139. return rows
  140. # =================================================================
  141. # Phase 3 chunk
  142. # =================================================================
  143. def run_p3_chunk(batch_id: int, seed_idx: int, cfg: dict, batches, device: str) -> list[dict]:
  144. seed = cfg["statistics"]["seeds"][seed_idx]
  145. set_seed(seed)
  146. b1 = batches[0]
  147. n = len(b1.y); n_test = max(2, int(0.30 * n))
  148. Xtr_b1, ytr_b1 = b1.X[:n - n_test], b1.y[:n - n_test]
  149. mu, sigma = DL.fit_standardiser(Xtr_b1)
  150. Xtr_b1n = DL.standardise(Xtr_b1, mu, sigma)
  151. cfg_lite = cfg["models"]["liteSensor"]
  152. n_classes = cfg["dataset"]["n_classes"]
  153. epochs_src = cfg["training"]["epochs_source"]
  154. epochs_kd = cfg["training"]["epochs_kd_finetune"]
  155. Xtr = reshape_for_sensors(Xtr_b1n); ytr = ytr_b1 - 1
  156. idx = np.arange(len(ytr)); np.random.shuffle(idx); v = int(0.20 * len(idx))
  157. teacher = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
  158. res = TR.train_supervised(
  159. teacher, Xtr[idx[v:]], ytr[idx[v:]], Xtr[idx[:v]], ytr[idx[:v]],
  160. epochs=epochs_src, lr=cfg["training"]["lr"],
  161. batch_size=cfg["training"]["batch_size_source"], device=device,
  162. )
  163. teacher.load_state_dict(res.best_state)
  164. batch = batches[batch_id - 1] # batch_id 2..10 → index 1..9
  165. X_b = DL.standardise(batch.X, mu, sigma)
  166. rows = []
  167. # NC
  168. il, iu, ite = DL.stratified_chronological_split(
  169. DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
  170. target_labeled_fraction=0.20,
  171. test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
  172. rng=np.random.default_rng(seed),
  173. )
  174. Xte_t = reshape_for_sensors(X_b[ite]); yte_t = batch.y[ite] - 1
  175. yhat_nc = TR.predict(teacher, Xte_t, device=device)
  176. rows.append(_drow(seed_idx, seed, batch.batch_id, "NC", "supervised_00",
  177. EV.accuracy(yte_t, yhat_nc), EV.macro_f1(yte_t, yhat_nc), len(yte_t)))
  178. # KD-DM ablation
  179. for frac in cfg["splits"]["target_labeled_fraction"]:
  180. il_f, iu_f, ite_f = DL.stratified_chronological_split(
  181. DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
  182. target_labeled_fraction=frac,
  183. test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
  184. rng=np.random.default_rng(seed + int(frac * 100)),
  185. )
  186. if len(il_f) == 0:
  187. continue # no labeled samples for this batch
  188. Xl = reshape_for_sensors(X_b[il_f])
  189. yl = torch.as_tensor(batch.y[il_f] - 1, dtype=torch.long)
  190. Xu = reshape_for_sensors(X_b[iu_f]) if len(iu_f) > 0 else np.zeros((0, 16, 8), dtype=np.float32)
  191. student = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
  192. student.load_state_dict(teacher.state_dict())
  193. student, _ = KD.fine_tune_kd_dm(
  194. teacher, student,
  195. torch.as_tensor(Xl, dtype=torch.float32),
  196. yl,
  197. torch.as_tensor(Xu, dtype=torch.float32),
  198. epochs=epochs_kd, alpha=cfg["kd_dm"]["alpha"], T=cfg["kd_dm"]["temperature"],
  199. lr=cfg["training"]["lr"], device=device,
  200. batch_size=cfg["training"]["batch_size_target"],
  201. )
  202. Xte_test = reshape_for_sensors(X_b[ite_f]); yte_test = batch.y[ite_f] - 1
  203. yhat = TR.predict(student, Xte_test, device=device)
  204. rows.append(_drow(seed_idx, seed, batch.batch_id,
  205. f"KD-DM-{int(frac*100):02d}", f"supervised_{int(frac*100):02d}",
  206. EV.accuracy(yte_test, yhat), EV.macro_f1(yte_test, yhat), len(yte_test)))
  207. # KD-DM-unsup
  208. il0, iu0, ite0 = DL.stratified_chronological_split(
  209. DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
  210. target_labeled_fraction=0.0,
  211. test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
  212. rng=np.random.default_rng(seed + 999),
  213. )
  214. Xall_unsup = reshape_for_sensors(np.concatenate([X_b[il0], X_b[iu0]], axis=0))
  215. student = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
  216. student.load_state_dict(teacher.state_dict())
  217. student, _ = KD.fine_tune_kd_dm(
  218. teacher, student,
  219. torch.zeros(0, 16, 8), torch.zeros(0, dtype=torch.long),
  220. torch.as_tensor(Xall_unsup, dtype=torch.float32),
  221. epochs=epochs_kd, alpha=0.0, T=cfg["kd_dm"]["temperature"],
  222. lr=cfg["training"]["lr"], device=device,
  223. pseudo_label_threshold=0.7,
  224. batch_size=cfg["training"]["batch_size_target"],
  225. )
  226. Xte_test = reshape_for_sensors(X_b[ite0]); yte_test = batch.y[ite0] - 1
  227. yhat = TR.predict(student, Xte_test, device=device)
  228. rows.append(_drow(seed_idx, seed, batch.batch_id, "KD-DM-unsup", "unsupervised",
  229. EV.accuracy(yte_test, yhat), EV.macro_f1(yte_test, yhat), len(yte_test)))
  230. # DRCA-fair
  231. from sklearn.ensemble import RandomForestClassifier
  232. il20, iu20, ite20 = DL.stratified_chronological_split(
  233. DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
  234. target_labeled_fraction=0.20,
  235. test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
  236. rng=np.random.default_rng(seed + 555),
  237. )
  238. if len(il20) > 0:
  239. clf, _ = KD.drca_fair_adapt(
  240. lambda: RandomForestClassifier(n_estimators=80, random_state=seed),
  241. teacher,
  242. reshape_for_sensors(Xtr_b1n), ytr_b1 - 1,
  243. reshape_for_sensors(X_b[il20]), batch.y[il20] - 1,
  244. reshape_for_sensors(X_b[iu20]) if len(iu20) > 0 else np.zeros((0,16,8),dtype=np.float32),
  245. device=device,
  246. )
  247. with torch.no_grad():
  248. Ft_test = teacher(torch.as_tensor(reshape_for_sensors(X_b[ite20]),
  249. dtype=torch.float32, device=device))
  250. Ft_test = (Ft_test[1] if isinstance(Ft_test, tuple) else Ft_test).cpu().numpy()
  251. yhat_drca = clf.predict(Ft_test)
  252. yte_drca = batch.y[ite20] - 1
  253. rows.append(_drow(seed_idx, seed, batch.batch_id, "DRCA-fair-20", "supervised_20",
  254. EV.accuracy(yte_drca, yhat_drca),
  255. EV.macro_f1(yte_drca, yhat_drca), len(yte_drca)))
  256. return rows
  257. def _drow(seed_idx, seed, target_batch, method, variant, acc, f1, n):
  258. return {"chunk_id": f"p3_b{target_batch}_s{seed_idx}",
  259. "seed": seed, "target_batch": target_batch, "method": method,
  260. "variant": variant, "acc": float(acc), "macro_f1": float(f1), "n_test": int(n)}
  261. # =================================================================
  262. # Phase 4 — final summary
  263. # =================================================================
  264. def run_p4(cfg: dict, results_dir: Path) -> dict:
  265. n_classes = cfg["dataset"]["n_classes"]
  266. fp32 = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg["models"]["liteSensor"])
  267. pruned = C.structured_l1_prune(fp32, sparsity=cfg["compression"]["prune_target_sparsity"])
  268. breakdown = C.precision_vs_sparsity_breakdown(fp32, pruned)
  269. df = pd.read_csv(results_dir / "raw_results_drift.csv")
  270. nc_per = df[df.method == "NC"].groupby("target_batch").acc.mean().sort_index().values.tolist()
  271. mono = EV.sanity_monotonicity(nc_per, "down")
  272. base = {"expected": 1.0 / n_classes,
  273. "observed": float(df[df.method == "NC"].acc.mean()),
  274. "passed": bool(df[df.method == "NC"].acc.mean() >= 1.0 / n_classes - 0.05)}
  275. cross = EV.sanity_cross_condition(
  276. df[df.method.str.startswith("KD-DM-20")].acc.mean(),
  277. df[df.method == "DRCA-fair-20"].acc.mean(),
  278. )
  279. summary = {
  280. "version": "v1_reduced_2026-05-10",
  281. "n_seeds": len(cfg["statistics"]["seeds"]),
  282. "n_target_batches": int(df["target_batch"].nunique()),
  283. "compression_breakdown_R3_14": breakdown,
  284. "sanity_checks_GATE4": {
  285. "monotonicity_NC": mono,
  286. "baseline_plausibility": base,
  287. "cross_condition": cross,
  288. },
  289. "drift_method_summary": (
  290. df.groupby("method").agg(
  291. mean_acc=("acc", "mean"), std_acc=("acc", "std"),
  292. mean_f1=("macro_f1", "mean"), n=("acc", "size"),
  293. ).round(4).reset_index().to_dict(orient="records")
  294. ),
  295. "submitted_claims_to_revise": {
  296. "78_percent_reduction": (
  297. f"submitted=78%; revised: precision={breakdown['precision_reduction_pct']}% "
  298. f"+ structural_sparsity={breakdown['sparsity_reduction_pct']}%"
  299. ),
  300. "Task_A_100_percent": "submitted=~100% (random splits); see raw_results_drift.csv",
  301. "MCU_scale": "edge-board-scale (Cortex-A class)",
  302. "LiteSensor_params": (
  303. f"submitted=47,392 params/41 kB; "
  304. f"measured={breakdown['size_fp32_kB']} kB FP32, "
  305. f"{breakdown['size_int8_pruned_kB']} kB INT8 pruned"
  306. ),
  307. },
  308. }
  309. return summary
  310. # =================================================================
  311. # Main
  312. # =================================================================
  313. def main():
  314. ap = argparse.ArgumentParser()
  315. ap.add_argument("--config", default="config.yaml")
  316. ap.add_argument("--device", default="cpu")
  317. ap.add_argument("--chunk", default=None,
  318. help="explicit chunk id (e.g., p2_s0, p3_b2_s0, p4)")
  319. ap.add_argument("--resume", action="store_true",
  320. help="auto-pick next pending chunk")
  321. args = ap.parse_args()
  322. with open(args.config) as f:
  323. cfg = yaml.safe_load(f)
  324. results_dir = (ROOT / cfg["paths"]["results_dir"]).resolve()
  325. results_dir.mkdir(parents=True, exist_ok=True)
  326. progress_file = results_dir / "progress.json"
  327. bench_csv = results_dir / "raw_results_benchmark.csv"
  328. drift_csv = results_dir / "raw_results_drift.csv"
  329. done = load_progress(progress_file)
  330. plan = all_chunks(cfg)
  331. if args.resume and not args.chunk:
  332. pending = [c for c in plan if c not in done]
  333. if not pending:
  334. print("ALL CHUNKS DONE."); return
  335. args.chunk = pending[0]
  336. print(f"[RESUME] next pending chunk: {args.chunk} ({len(done)}/{len(plan)} done)")
  337. if not args.chunk:
  338. print("No chunk specified. Use --chunk or --resume."); return
  339. if args.chunk in done:
  340. print(f"chunk {args.chunk} already done — skipping"); return
  341. raw_dir = (ROOT / cfg["paths"]["raw_data_dir"]).resolve()
  342. print(f"[chunk={args.chunk}] loading batches ...")
  343. batches = DL.load_all_batches(str(raw_dir), n_batches=cfg["dataset"]["n_batches"])
  344. t0 = time.time()
  345. if args.chunk.startswith("p2_s"):
  346. seed_idx = int(args.chunk.split("_s")[1])
  347. rows = run_p2_chunk(seed_idx, cfg, batches, args.device)
  348. append_csv(bench_csv, BENCH_HEADERS, rows)
  349. print(f"[chunk={args.chunk}] benchmark rows appended: {len(rows)}")
  350. elif args.chunk.startswith("p3_"):
  351. # p3_b{batch_id}_s{seed_idx}
  352. parts = args.chunk.split("_")
  353. batch_id = int(parts[1][1:]); seed_idx = int(parts[2][1:])
  354. rows = run_p3_chunk(batch_id, seed_idx, cfg, batches, args.device)
  355. append_csv(drift_csv, DRIFT_HEADERS, rows)
  356. print(f"[chunk={args.chunk}] drift rows appended: {len(rows)}")
  357. elif args.chunk == "p4":
  358. # require all p2 + p3 chunks done
  359. prereq = [c for c in plan if c != "p4" and c not in done]
  360. if prereq:
  361. print(f"p4 blocked — pending: {prereq[:5]}..."); return
  362. summary = run_p4(cfg, results_dir)
  363. out = results_dir / "experiment_summary.json"
  364. out.write_text(json.dumps(summary, indent=2, ensure_ascii=False,
  365. default=lambda o: bool(o) if hasattr(o, "__bool__") else str(o)))
  366. print(f"summary written: {out}")
  367. print(json.dumps(summary, indent=2,
  368. default=lambda o: bool(o) if hasattr(o, "__bool__") else str(o))[:1500])
  369. else:
  370. print(f"unknown chunk: {args.chunk}"); return
  371. done.add(args.chunk)
  372. save_progress(progress_file, done)
  373. elapsed = time.time() - t0
  374. print(f"[chunk={args.chunk}] done in {elapsed:.1f}s ({len(done)}/{len(plan)})")
  375. if __name__ == "__main__":
  376. main()

analysis_main.py at commit c102c14, under Apache-2.0 · at the source

Overview

Authors: Soohwan Kim1, Myeongsik Shin1, Ku Kang1, Doo-Hee Lee1, David G Churchill2,3, Yoon Jeong Jang1
  1. CBRN Defense Research Institute, Seoul 06796, Republic of Korea; (S.K.); (M.S.); (K.K.); (D.-H.L.)
  2. Department of Chemistry, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
  3. Therapeutic Bioengineering Section, KAIST Institute for Health Science and Technology (KIHST), Daejeon 34141, Republic of Korea
Journal: Molecules (Basel, Switzerland), volume 31, issue 11, article 1884
Dates: received 7 April 2026; accepted 25 May 2026; published online 1 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/molecules31111884 · PMID 42280188 · PMCID PMC13257611 · OpenAlex W7163049842
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: methods / tools (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing
Keywords: CBRN detection, gas sensor array, edge AI, TinyML, knowledge distillation, sensor drift compensation
Topic: Advanced Chemical Sensor Technologies (Biomedical Engineering, Engineering), according to OpenAlex
Citations: not cited yet (Europe PMC); 33 references in the paper

Abstract

Resource-constrained edge processors deployed on unmanned aerial vehicles and wearable platforms require compact, drift-robust gas classification models for a range of environmental and security monitoring applications, including CBRN-motivated scenarios. Existing approaches rely on server-grade architectures incompatible with edge-board-scale deployment, or on classifiers that chemically degrade severely under long-term sensor drift. Each UCI gas class was mapped to a CBRN behavioral category based on physicochemical analogy (molecular functional group, vapor pressure, and metal-oxide semiconductor (MOS) cross-sensitivity pattern), following established precedent. Analyzed were Ammonia (NH3), Acetaldehyde (CH3CHO), Acetone ((CH3)2CO), Ethylene (C2H4), Ethanol (C2H5OH), Toluene (C6H5CH3). We propose herein an end-to-end pipeline integrating a novel 1-D convolutional neural network with depth-wise separable convolutions (LiteSensor-Net), INT8 post-training quantization, structured magnitude pruning, and a knowledge-distillation domain-adaptation module (KD–DM) for sensor drift compensation. Using the UCI Gas Sensor Array Drift Dataset (13,910 measurements; 16 metal-oxide sensors; six analyte gases; a 36-month work span). LiteSensor-Net achieved accuracy = 92.63 ± 2.02%, macro-F1 = 0.898, model size = 5.99 kB INT8 pruned, inference latency = 6.3 ms, RAM footprint = 31.7 kB, and energy per inference = 0.04 mJ (all metrics on Raspberry Pi 4B, ARM Cortex-A72). Under chronological forward-chaining evaluation, KD–DM–20 achieved 47.91 ± 18.79% mean accuracy over Batches 2–10, representing a +9.25 pp improvement over uncompensated NC (38.66%). A six-metric benchmark framework—accuracy, macro-F1, model size, inference latency, RAM footprint, and energy per inference—is introduced to standardize edge-AI gas classifier evaluation. The proposed pipeline provides an open-source, deployable foundation for edge-class gas classification systems, with CBRN detection as a motivating application. Full operational validation on certified chemical simulants remains as future work.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

bisu9082/LiteSensor-Net

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c102c1466471fcf980787b6a57958f438d693dd0, 16 May 2026
Languages: Python (31), Shell (3)
Size: 100 files, 34 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, license file, CITATION.cff, environment (code/environment.yml, code/pyproject.toml, code/requirements.txt), tests, documentation
Not found: continuous integration
Tools: NumPy (26 files), PyTorch (26 files), scikit-learn (10 files), pandas (4 files), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
36 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 34 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The UCI Gas Sensor Array Drift Dataset is publicly available at https://doi.org/10.24432/C5RP6W. The Scientific Data long-term drift dataset is at https://doi.org/10.1038/s41597-025-05993-8. The public GitHub repository (https://github.com/bisu9082/LiteSensor-Net, accessed on 24 May 2026), implemented in Python 3.10 with PyTorch 2.4.0 and TensorFlow Lite runtime ≥ 2.14, contains the training/evaluation code, raw UCI batch files, raw result CSV/JSON outputs, configuration files, and reproducibility documentation required to reproduce the revised benchmark, compression ablation, drift-compensation experiments, and calibration analyses.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 6 keywords, 30 references.

Cite

This paper

Kim, S., Shin, M., Kang, K., Lee, D.-H., Churchill, D. G., & Jang, Y. J. (2026). Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping. Molecules (Basel, Switzerland), 31(11), 1884. https://doi.org/10.3390/molecules31111884

BibTeX

@article{kim2026drift,
author = {Kim, Soohwan and Shin, Myeongsik and Kang, Ku and Lee, Doo-Hee and Churchill, David G and Jang, Yoon Jeong},
title = {{Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping}},
journal = {Molecules (Basel, Switzerland)},
year = {2026},
month = jun,
volume = {31},
number = {11},
pages = {1884},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1420-3049},
doi = {10.3390/molecules31111884},
url = {https://doi.org/10.3390/molecules31111884},
pmid = {42280188},
pmcid = {PMC13257611}
}

RIS

TY - JOUR
AU - Kim, Soohwan
AU - Shin, Myeongsik
AU - Kang, Ku
AU - Lee, Doo-Hee
AU - Churchill, David G
AU - Jang, Yoon Jeong
TI - Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping
T2 - Molecules (Basel, Switzerland)
J2 - Molecules
PY - 2026
DA - 2026/06/01
VL - 31
IS - 11
SP - 1884
SN - 1420-3049
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/molecules31111884
UR - https://doi.org/10.3390/molecules31111884
LA - en
ER -

CSL-JSON

{
"id": "10.3390/molecules31111884",
"type": "article-journal",
"title": "Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping",
"container-title": "Molecules (Basel, Switzerland)",
"author": [
{
"family": "Kim",
"given": "Soohwan"
},
{
"family": "Shin",
"given": "Myeongsik"
},
{
"family": "Kang",
"given": "Ku"
},
{
"family": "Lee",
"given": "Doo-Hee"
},
{
"family": "Churchill",
"given": "David G"
},
{
"family": "Jang",
"given": "Yoon Jeong"
}
],
"container-title-short": "Molecules",
"volume": "31",
"issue": "11",
"page": "1884",
"DOI": "10.3390/molecules31111884",
"PMID": "42280188",
"PMCID": "PMC13257611",
"ISSN": "1420-3049",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/molecules31111884",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.isci.2026.117206 [code]
ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.
Journal: iScience
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[2] doi:10.1038/s41746-026-02778-0 [code]
Trust-gated synthetic EEG augmentation reduces performance drops when generalizing to new patients.
Journal: NPJ digital medicine
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[3] doi:10.3390/diagnostics16111588 [code]
Bridging Annotation Gaps: Hierarchical Self-Support Learning for Brain Tumor Segmentation.
Journal: Diagnostics (Basel, Switzerland)
In common: PyTorch, SciPy, NumPy, methods / tools, 1 reference
[4] doi:10.1038/s41598-026-54785-6 [code]
Lightweight deep learning model for nonconvulsive status epilepticus diagnosis using EEG time-frequency analysis.
Journal: Scientific reports
In common: PyTorch, SciPy, NumPy, 1 reference
[5] doi:10.1038/s41598-026-54840-2 [code]
Predictive metacognition: a neuro-computational framework for self-monitoring in large language models.
Journal: Scientific reports
In common: PyTorch, pandas, NumPy, 1 reference
[6] doi:10.3390/jimaging12060233 [code]
Brain Tumor Classification in MRI Images Using Combined Transfer Learning and Convolutional Neural Networks.
Journal: Journal of imaging
In common: scikit-learn, pandas, NumPy, 1 reference
[7] doi:10.1016/j.patter.2026.101564 [code]
Spacing effect improves generalization in biological and artificial systems.
Journal: Patterns (New York, N.Y.)
In common: PyTorch, NumPy, 1 reference
[8] doi:10.1038/s41598-026-50603-1
A multi-class framework for face mask compliance detection using lightweight deep learning models.
Journal: Scientific reports
In common: methods / tools, 1 reference
[9] doi:
NeuroTrustNet: a cost-effective multimodal ensemble framework for brain tumor classification under cross-dataset variability
Journal: Frontiers in artificial intelligence
In common: methods / tools, 1 reference
[10] doi:10.3389/frai.2026.1849571
NeuroPlast: a learnable activation function evaluated under knowledge distillation for medical image classification.
Journal: Frontiers in artificial intelligence
In common: methods / tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.