Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping.
The 16 matches
- [1] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/modules/models.py, lines 19–35 · score 0.82 · sensor channel separability, ReLU, Cross sensor, axis, mixing, BN
- [2] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/modules/training.py, lines 21–63 · score 0.77 · cosine annealing, weight decay, AdamW, cross entropy, min, loss
- [3] § 2. Materials and Methods › 2.7. Domain–Adaptation Task Definitions ↔ code/eval_drift_taskB.py, lines 1–36 · score 0.74 · cumulative training scenario, DRCA fair, fine tuning, KD DM, target batch, held
- [4] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/analysis_main.py, lines 167–290 · score 0.74 · KD DM unsup, DRCA fair, fine tuning, unsupervised, pseudo, temperature
- [5] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/eval_compression.py, lines 31–129 · score 0.72 · ReLU, pointwise, FC, DSConv, activations, BN
- [6] § 2. Materials and Methods › 2.4. LiteSensor-Net Architecture ↔ code/run_litesensor_t1.py, lines 64–105 · score 0.72 · weight decay, AdamW, cross entropy, smoothing, cosine, validation
- [7] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_combo2.py, lines 259–338 · score 0.72 · stratified random, cross entropy, fine tuning, KD DM, target batch, KL
- [8] § 2. Materials and Methods › 2.5. Multi-Stage Compression › 2.5.1. INT8 Post-Training Quantization ↔ code/run_t3_qat.py, lines 1–31 · score 0.63 · Quantization aware training, QAT, fine tuning, PTQ, Post, weights
- [9] § 3. Results › 3.4. Sensor Drift Compensation › 3.4.1. Task A: Laboratory Re-Calibration Scenario ↔ code/run_kdm_v4.py, lines 115–239 · score 0.61 · KD DM unsup, DRCA fair, Random Forest, target batch, NC, chronological
- [10] § 3. Results › 3.4. Sensor Drift Compensation › 3.4.1. Task A: Laboratory Re-Calibration Scenario ↔ code/analysis_main.py, lines 167–290 · score 0.60 · KD DM unsup, DRCA fair, Random Forest, NC, chronological, fraction
- [11] § 3. Results › 3.3. CBRN Simulant Classification Performance ↔ code/eval_classification.py, lines 61–146 · score 0.59 · confusion matrix, macro F1, recall, CWA, precision, scores
- [12] § 2. Materials and Methods › 2.5. Multi-Stage Compression › 2.5.1. INT8 Post-Training Quantization ↔ code/eval_compression.py, lines 1–21 · score 0.57 · TensorFlow, INT8 quantized, activations, weights, Lite
- [13] § 2. Materials and Methods › 2.3. Preprocessing Pipeline ↔ code/modules/data_loader.py, lines 128–140 · score 0.54 · Feature selection, Random Forest, Gini, Batch
- [14] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_combo2.py, lines 259–338 · score 0.53 · soft target, fine tuning, KD DM, ensemble, logits, student
- [15] § 2. Materials and Methods › 2.7. Domain–Adaptation Task Definitions ↔ code/run_kdm_v4.py, lines 115–239 · score 0.53 · KD DM unsup, DRCA fair, target batch, NC, fraction
- [16] § 2. Materials and Methods › 2.6. Knowledge-Distillation Drift-Compensation Module ↔ code/run_s1_30splits.py, lines 18–118 · score 0.53 · member ensemble, fine tuning, KD DM, supervision, student, teacher
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 431 lines · 18 KB · Apache-2.0 · 2 matches
- """Re-experiment orchestrator (CHUNK MODE) — Molecules MS-4274054 R1 revision.
- Sandbox bash call has a 45 s limit, so each invocation runs **ONE** chunk and
- exits, persisting state to disk. Subsequent calls resume.
- Chunk identifiers
- p2_s{seed_idx} : Phase 2 — all 4 archs × 1 seed × Batch 1
- p3_b{batch_id}_s{seed_idx} : Phase 3 — 1 target batch × 1 seed × all 5 conditions
- p4 : Phase 4 — compression breakdown + sanity + summary
- State files
- results/progress.json : list of completed chunks
- results/raw_results_benchmark.csv : appended after each p2 chunk
- results/raw_results_drift.csv : appended after each p3 chunk
- results/experiment_summary.json : written by p4
- Usage
- python analysis_main.py --resume # auto-pick next chunk
- python analysis_main.py --chunk p2_s0 --device cpu # explicit chunk
- python analysis_main.py --chunk p3_b2_s0
- python analysis_main.py --chunk p4
- """
- from __future__ import annotations
- import argparse
- import csv
- import json
- import os
- import random
- import sys
- import time
- from pathlib import Path
- from typing import List
- import numpy as np
- import pandas as pd
- import torch
- import yaml
- ROOT = Path(__file__).resolve().parent
- sys.path.insert(0, str(ROOT))
- from modules import data_loader as DL
- from modules import models as M
- from modules import compression as C
- from modules import kd_dm as KD
- from modules import evaluate as EV
- from modules import training as TR
- def set_seed(seed: int):
- random.seed(seed); np.random.seed(seed); torch.manual_seed(seed)
- torch.backends.cudnn.deterministic = True
- def reshape_for_sensors(X: np.ndarray) -> np.ndarray:
- return DL.reshape_for_sensors(X, n_sensors=16, n_feat_per_sensor=8)
- # =================================================================
- # Chunk planning
- # =================================================================
- def all_chunks(cfg: dict) -> List[str]:
- seeds = cfg["statistics"]["seeds"]
- chunks = [f"p2_s{i}" for i in range(len(seeds))]
- for bid in range(2, cfg["dataset"]["n_batches"] + 1):
- for i in range(len(seeds)):
- chunks.append(f"p3_b{bid}_s{i}")
- chunks.append("p4")
- return chunks
- def load_progress(path: Path) -> set:
- if path.exists():
- return set(json.loads(path.read_text()).get("done", []))
- return set()
- def save_progress(path: Path, done: set):
- path.write_text(json.dumps({"done": sorted(done),
- "ts": time.strftime("%Y-%m-%dT%H:%M:%S")},
- indent=2))
- # =================================================================
- # CSV append helpers
- # =================================================================
- BENCH_HEADERS = ["chunk_id", "seed", "model", "value", "macro_f1", "params",
- "size_fp32_kB", "size_int8_kB", "infer_ms", "train_s", "best_val_acc"]
- DRIFT_HEADERS = ["chunk_id", "seed", "target_batch", "method", "variant",
- "acc", "macro_f1", "n_test"]
- def append_csv(path: Path, headers: list, rows: list[dict]):
- new = not path.exists()
- with path.open("a", newline="") as f:
- w = csv.DictWriter(f, fieldnames=headers)
- if new: w.writeheader()
- for r in rows:
- w.writerow({k: r.get(k, "") for k in headers})
- # =================================================================
- # Phase 2 chunk
- # =================================================================
- def run_p2_chunk(seed_idx: int, cfg: dict, batches, device: str) -> list[dict]:
- seed = cfg["statistics"]["seeds"][seed_idx]
- set_seed(seed)
- b1 = batches[0]
- n = len(b1.y)
- n_test = max(2, int(0.30 * n))
- Xtr_full = b1.X[:n - n_test]; ytr_full = b1.y[:n - n_test]
- Xte = b1.X[n - n_test:]; yte = b1.y[n - n_test:]
- mu, sigma = DL.fit_standardiser(Xtr_full)
- Xtr_full = DL.standardise(Xtr_full, mu, sigma)
- Xte = DL.standardise(Xte, mu, sigma)
- idx = np.arange(len(ytr_full)); np.random.shuffle(idx); v = int(0.20 * len(idx))
- val_i, tr_i = idx[:v], idx[v:]
- Xtr = reshape_for_sensors(Xtr_full[tr_i]); ytr = ytr_full[tr_i] - 1
- Xv = reshape_for_sensors(Xtr_full[val_i]); yv = ytr_full[val_i] - 1
- Xtest = reshape_for_sensors(Xte); ytest = yte - 1
- model_specs = {
- "litesensor": cfg["models"]["liteSensor"],
- "mobilenet1d": cfg["models"]["mobileNet1D"],
- "inceptiontime1d": cfg["models"]["inceptionTime1D"],
- "resnet1d": cfg["models"]["resNet1D"],
- }
- n_classes = cfg["dataset"]["n_classes"]
- rows = []
- for name, mc in model_specs.items():
- model = M.build(name, n_classes=n_classes, in_ch=16, **mc)
- t0 = time.time()
- res = TR.train_supervised(
- model, Xtr, ytr, Xv, yv,
- epochs=cfg["training"]["epochs_source"],
- lr=cfg["training"]["lr"],
- batch_size=cfg["training"]["batch_size_source"],
- weight_decay=cfg["training"]["weight_decay"],
- device=device,
- )
- train_s = time.time() - t0
- model.load_state_dict(res.best_state)
- yhat = TR.predict(model, Xtest, device=device)
- t1 = time.time()
- for _ in range(3):
- _ = TR.predict(model, Xtest[:1], device=device)
- inf_ms = (time.time() - t1) / 3 * 1000
- n_params = M.count_parameters(model)
- rows.append({
- "chunk_id": f"p2_s{seed_idx}",
- "seed": seed, "model": name,
- "value": EV.accuracy(ytest, yhat),
- "macro_f1": EV.macro_f1(ytest, yhat),
- "params": n_params,
- "size_fp32_kB": round(n_params * 4 / 1024, 2),
- "size_int8_kB": round(n_params / 1024, 2),
- "infer_ms": round(inf_ms, 2),
- "train_s": round(train_s, 1),
- "best_val_acc": res.best_val_acc,
- })
- print(f" {name:18s} acc={rows[-1]['value']:.4f} f1={rows[-1]['macro_f1']:.4f} "
- f"params={n_params:,} inf={inf_ms:.2f}ms train={train_s:.1f}s")
- return rows
- # =================================================================
- # Phase 3 chunk
- # =================================================================
- def run_p3_chunk(batch_id: int, seed_idx: int, cfg: dict, batches, device: str) -> list[dict]:
- seed = cfg["statistics"]["seeds"][seed_idx]
- set_seed(seed)
- b1 = batches[0]
- n = len(b1.y); n_test = max(2, int(0.30 * n))
- Xtr_b1, ytr_b1 = b1.X[:n - n_test], b1.y[:n - n_test]
- mu, sigma = DL.fit_standardiser(Xtr_b1)
- Xtr_b1n = DL.standardise(Xtr_b1, mu, sigma)
- cfg_lite = cfg["models"]["liteSensor"]
- n_classes = cfg["dataset"]["n_classes"]
- epochs_src = cfg["training"]["epochs_source"]
- epochs_kd = cfg["training"]["epochs_kd_finetune"]
- Xtr = reshape_for_sensors(Xtr_b1n); ytr = ytr_b1 - 1
- idx = np.arange(len(ytr)); np.random.shuffle(idx); v = int(0.20 * len(idx))
- teacher = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
- res = TR.train_supervised(
- teacher, Xtr[idx[v:]], ytr[idx[v:]], Xtr[idx[:v]], ytr[idx[:v]],
- epochs=epochs_src, lr=cfg["training"]["lr"],
- batch_size=cfg["training"]["batch_size_source"], device=device,
- )
- teacher.load_state_dict(res.best_state)
- batch = batches[batch_id - 1] # batch_id 2..10 → index 1..9
- X_b = DL.standardise(batch.X, mu, sigma)
- rows = []
- # NC
- il, iu, ite = DL.stratified_chronological_split(
- DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
- target_labeled_fraction=0.20,
- test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
- rng=np.random.default_rng(seed),
- )
- Xte_t = reshape_for_sensors(X_b[ite]); yte_t = batch.y[ite] - 1
- yhat_nc = TR.predict(teacher, Xte_t, device=device)
- rows.append(_drow(seed_idx, seed, batch.batch_id, "NC", "supervised_00",
- EV.accuracy(yte_t, yhat_nc), EV.macro_f1(yte_t, yhat_nc), len(yte_t)))
- # KD-DM ablation
- for frac in cfg["splits"]["target_labeled_fraction"]:
- il_f, iu_f, ite_f = DL.stratified_chronological_split(
- DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
- target_labeled_fraction=frac,
- test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
- rng=np.random.default_rng(seed + int(frac * 100)),
- )
- if len(il_f) == 0:
- continue # no labeled samples for this batch
- Xl = reshape_for_sensors(X_b[il_f])
- yl = torch.as_tensor(batch.y[il_f] - 1, dtype=torch.long)
- Xu = reshape_for_sensors(X_b[iu_f]) if len(iu_f) > 0 else np.zeros((0, 16, 8), dtype=np.float32)
- student = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
- student.load_state_dict(teacher.state_dict())
- student, _ = KD.fine_tune_kd_dm(
- teacher, student,
- torch.as_tensor(Xl, dtype=torch.float32),
- yl,
- torch.as_tensor(Xu, dtype=torch.float32),
- epochs=epochs_kd, alpha=cfg["kd_dm"]["alpha"], T=cfg["kd_dm"]["temperature"],
- lr=cfg["training"]["lr"], device=device,
- batch_size=cfg["training"]["batch_size_target"],
- )
- Xte_test = reshape_for_sensors(X_b[ite_f]); yte_test = batch.y[ite_f] - 1
- yhat = TR.predict(student, Xte_test, device=device)
- rows.append(_drow(seed_idx, seed, batch.batch_id,
- f"KD-DM-{int(frac*100):02d}", f"supervised_{int(frac*100):02d}",
- EV.accuracy(yte_test, yhat), EV.macro_f1(yte_test, yhat), len(yte_test)))
- # KD-DM-unsup
- il0, iu0, ite0 = DL.stratified_chronological_split(
- DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
- target_labeled_fraction=0.0,
- test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
- rng=np.random.default_rng(seed + 999),
- )
- Xall_unsup = reshape_for_sensors(np.concatenate([X_b[il0], X_b[iu0]], axis=0))
- student = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg_lite)
- student.load_state_dict(teacher.state_dict())
- student, _ = KD.fine_tune_kd_dm(
- teacher, student,
- torch.zeros(0, 16, 8), torch.zeros(0, dtype=torch.long),
- torch.as_tensor(Xall_unsup, dtype=torch.float32),
- epochs=epochs_kd, alpha=0.0, T=cfg["kd_dm"]["temperature"],
- lr=cfg["training"]["lr"], device=device,
- pseudo_label_threshold=0.7,
- batch_size=cfg["training"]["batch_size_target"],
- )
- Xte_test = reshape_for_sensors(X_b[ite0]); yte_test = batch.y[ite0] - 1
- yhat = TR.predict(student, Xte_test, device=device)
- rows.append(_drow(seed_idx, seed, batch.batch_id, "KD-DM-unsup", "unsupervised",
- EV.accuracy(yte_test, yhat), EV.macro_f1(yte_test, yhat), len(yte_test)))
- # DRCA-fair
- from sklearn.ensemble import RandomForestClassifier
- il20, iu20, ite20 = DL.stratified_chronological_split(
- DL.BatchData(batch.batch_id, X_b, batch.y, batch.order),
- target_labeled_fraction=0.20,
- test_holdout_fraction=cfg["splits"]["test_holdout_fraction"],
- rng=np.random.default_rng(seed + 555),
- )
- if len(il20) > 0:
- clf, _ = KD.drca_fair_adapt(
- lambda: RandomForestClassifier(n_estimators=80, random_state=seed),
- teacher,
- reshape_for_sensors(Xtr_b1n), ytr_b1 - 1,
- reshape_for_sensors(X_b[il20]), batch.y[il20] - 1,
- reshape_for_sensors(X_b[iu20]) if len(iu20) > 0 else np.zeros((0,16,8),dtype=np.float32),
- device=device,
- )
- with torch.no_grad():
- Ft_test = teacher(torch.as_tensor(reshape_for_sensors(X_b[ite20]),
- dtype=torch.float32, device=device))
- Ft_test = (Ft_test[1] if isinstance(Ft_test, tuple) else Ft_test).cpu().numpy()
- yhat_drca = clf.predict(Ft_test)
- yte_drca = batch.y[ite20] - 1
- rows.append(_drow(seed_idx, seed, batch.batch_id, "DRCA-fair-20", "supervised_20",
- EV.accuracy(yte_drca, yhat_drca),
- EV.macro_f1(yte_drca, yhat_drca), len(yte_drca)))
- return rows
- def _drow(seed_idx, seed, target_batch, method, variant, acc, f1, n):
- return {"chunk_id": f"p3_b{target_batch}_s{seed_idx}",
- "seed": seed, "target_batch": target_batch, "method": method,
- "variant": variant, "acc": float(acc), "macro_f1": float(f1), "n_test": int(n)}
- # =================================================================
- # Phase 4 — final summary
- # =================================================================
- def run_p4(cfg: dict, results_dir: Path) -> dict:
- n_classes = cfg["dataset"]["n_classes"]
- fp32 = M.build("litesensor", n_classes=n_classes, in_ch=16, **cfg["models"]["liteSensor"])
- pruned = C.structured_l1_prune(fp32, sparsity=cfg["compression"]["prune_target_sparsity"])
- breakdown = C.precision_vs_sparsity_breakdown(fp32, pruned)
- df = pd.read_csv(results_dir / "raw_results_drift.csv")
- nc_per = df[df.method == "NC"].groupby("target_batch").acc.mean().sort_index().values.tolist()
- mono = EV.sanity_monotonicity(nc_per, "down")
- base = {"expected": 1.0 / n_classes,
- "observed": float(df[df.method == "NC"].acc.mean()),
- "passed": bool(df[df.method == "NC"].acc.mean() >= 1.0 / n_classes - 0.05)}
- cross = EV.sanity_cross_condition(
- df[df.method.str.startswith("KD-DM-20")].acc.mean(),
- df[df.method == "DRCA-fair-20"].acc.mean(),
- )
- summary = {
- "version": "v1_reduced_2026-05-10",
- "n_seeds": len(cfg["statistics"]["seeds"]),
- "n_target_batches": int(df["target_batch"].nunique()),
- "compression_breakdown_R3_14": breakdown,
- "sanity_checks_GATE4": {
- "monotonicity_NC": mono,
- "baseline_plausibility": base,
- "cross_condition": cross,
- },
- "drift_method_summary": (
- df.groupby("method").agg(
- mean_acc=("acc", "mean"), std_acc=("acc", "std"),
- mean_f1=("macro_f1", "mean"), n=("acc", "size"),
- ).round(4).reset_index().to_dict(orient="records")
- ),
- "submitted_claims_to_revise": {
- "78_percent_reduction": (
- f"submitted=78%; revised: precision={breakdown['precision_reduction_pct']}% "
- f"+ structural_sparsity={breakdown['sparsity_reduction_pct']}%"
- ),
- "Task_A_100_percent": "submitted=~100% (random splits); see raw_results_drift.csv",
- "MCU_scale": "edge-board-scale (Cortex-A class)",
- "LiteSensor_params": (
- f"submitted=47,392 params/41 kB; "
- f"measured={breakdown['size_fp32_kB']} kB FP32, "
- f"{breakdown['size_int8_pruned_kB']} kB INT8 pruned"
- ),
- },
- }
- return summary
- # =================================================================
- # Main
- # =================================================================
- def main():
- ap = argparse.ArgumentParser()
- ap.add_argument("--config", default="config.yaml")
- ap.add_argument("--device", default="cpu")
- ap.add_argument("--chunk", default=None,
- help="explicit chunk id (e.g., p2_s0, p3_b2_s0, p4)")
- ap.add_argument("--resume", action="store_true",
- help="auto-pick next pending chunk")
- args = ap.parse_args()
- with open(args.config) as f:
- cfg = yaml.safe_load(f)
- results_dir = (ROOT / cfg["paths"]["results_dir"]).resolve()
- results_dir.mkdir(parents=True, exist_ok=True)
- progress_file = results_dir / "progress.json"
- bench_csv = results_dir / "raw_results_benchmark.csv"
- drift_csv = results_dir / "raw_results_drift.csv"
- done = load_progress(progress_file)
- plan = all_chunks(cfg)
- if args.resume and not args.chunk:
- pending = [c for c in plan if c not in done]
- if not pending:
- print("ALL CHUNKS DONE."); return
- args.chunk = pending[0]
- print(f"[RESUME] next pending chunk: {args.chunk} ({len(done)}/{len(plan)} done)")
- if not args.chunk:
- print("No chunk specified. Use --chunk or --resume."); return
- if args.chunk in done:
- print(f"chunk {args.chunk} already done — skipping"); return
- raw_dir = (ROOT / cfg["paths"]["raw_data_dir"]).resolve()
- print(f"[chunk={args.chunk}] loading batches ...")
- batches = DL.load_all_batches(str(raw_dir), n_batches=cfg["dataset"]["n_batches"])
- t0 = time.time()
- if args.chunk.startswith("p2_s"):
- seed_idx = int(args.chunk.split("_s")[1])
- rows = run_p2_chunk(seed_idx, cfg, batches, args.device)
- append_csv(bench_csv, BENCH_HEADERS, rows)
- print(f"[chunk={args.chunk}] benchmark rows appended: {len(rows)}")
- elif args.chunk.startswith("p3_"):
- # p3_b{batch_id}_s{seed_idx}
- parts = args.chunk.split("_")
- batch_id = int(parts[1][1:]); seed_idx = int(parts[2][1:])
- rows = run_p3_chunk(batch_id, seed_idx, cfg, batches, args.device)
- append_csv(drift_csv, DRIFT_HEADERS, rows)
- print(f"[chunk={args.chunk}] drift rows appended: {len(rows)}")
- elif args.chunk == "p4":
- # require all p2 + p3 chunks done
- prereq = [c for c in plan if c != "p4" and c not in done]
- if prereq:
- print(f"p4 blocked — pending: {prereq[:5]}..."); return
- summary = run_p4(cfg, results_dir)
- out = results_dir / "experiment_summary.json"
- out.write_text(json.dumps(summary, indent=2, ensure_ascii=False,
- default=lambda o: bool(o) if hasattr(o, "__bool__") else str(o)))
- print(f"summary written: {out}")
- print(json.dumps(summary, indent=2,
- default=lambda o: bool(o) if hasattr(o, "__bool__") else str(o))[:1500])
- else:
- print(f"unknown chunk: {args.chunk}"); return
- done.add(args.chunk)
- save_progress(progress_file, done)
- elapsed = time.time() - t0
- print(f"[chunk={args.chunk}] done in {elapsed:.1f}s ({len(done)}/{len(plan)})")
- if __name__ == "__main__":
- main()
analysis_main.py at commit c102c14, under Apache-2.0 · at the source
Overview
- CBRN Defense Research Institute, Seoul 06796, Republic of Korea; (S.K.); (M.S.); (K.K.); (D.-H.L.)
- Department of Chemistry, Korea Advanced Institute of Science and Technology (KAIST), Daejeon 34141, Republic of Korea
- Therapeutic Bioengineering Section, KAIST Institute for Health Science and Technology (KIHST), Daejeon 34141, Republic of Korea
Abstract
Resource-constrained edge processors deployed on unmanned aerial vehicles and wearable platforms require compact, drift-robust gas classification models for a range of environmental and security monitoring applications, including CBRN-motivated scenarios. Existing approaches rely on server-grade architectures incompatible with edge-board-scale deployment, or on classifiers that chemically degrade severely under long-term sensor drift. Each UCI gas class was mapped to a CBRN behavioral category based on physicochemical analogy (molecular functional group, vapor pressure, and metal-oxide semiconductor (MOS) cross-sensitivity pattern), following established precedent. Analyzed were Ammonia (NH3), Acetaldehyde (CH3CHO), Acetone ((CH3)2CO), Ethylene (C2H4), Ethanol (C2H5OH), Toluene (C6H5CH3). We propose herein an end-to-end pipeline integrating a novel 1-D convolutional neural network with depth-wise separable convolutions (LiteSensor-Net), INT8 post-training quantization, structured magnitude pruning, and a knowledge-distillation domain-adaptation module (KD–DM) for sensor drift compensation. Using the UCI Gas Sensor Array Drift Dataset (13,910 measurements; 16 metal-oxide sensors; six analyte gases; a 36-month work span). LiteSensor-Net achieved accuracy = 92.63 ± 2.02%, macro-F1 = 0.898, model size = 5.99 kB INT8 pruned, inference latency = 6.3 ms, RAM footprint = 31.7 kB, and energy per inference = 0.04 mJ (all metrics on Raspberry Pi 4B, ARM Cortex-A72). Under chronological forward-chaining evaluation, KD–DM–20 achieved 47.91 ± 18.79% mean accuracy over Batches 2–10, representing a +9.25 pp improvement over uncompensated NC (38.66%). A six-metric benchmark framework—accuracy, macro-F1, model size, inference latency, RAM footprint, and energy per inference—is introduced to standardize edge-AI gas classifier evaluation. The proposed pipeline provides an open-source, deployable foundation for edge-class gas classification systems, with CBRN detection as a motivating application. Full operational validation on certified chemical simulants remains as future work.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.
bisu9082/LiteSensor-Net
c102c1466471fcf980787b6a57958f438d693dd0, 16 May 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
36 files
- code/
analysis_main.py , Python, 431 lines, 2 matches - code/
eval_benchmark.py , Python, 162 lines - code/
eval_calibration.py , Python, 120 lines - code/
eval_classification.py , Python, 150 lines, 1 match - code/
eval_compression.py , Python, 133 lines, 2 matches - code/
eval_drift_taskA.py , Python, 431 lines - code/
eval_drift_taskB.py , Python, 244 lines, 1 match - code/
modules/ , Python, 1 line__init__.py - code/
modules/ , Python, 95 linescompression.py - code/
modules/ , Python, 182 lines, 1 matchdata_loader.py - code/
modules/ , Python, 97 linesevaluate.py - code/
modules/ , Python, 133 lineskd_dm.py - code/
modules/ , Python, 109 lineskd_dm_v4.py - code/
modules/ , Python, 212 lines, 1 matchmodels.py - code/
modules/ , Python, 89 lines, 1 matchtraining.py - code/
run_combo2.py , Python, 342 lines, 2 matches - code/
run_e_ece_tempscale.py , Python, 120 lines - code/
run_kdm_v4.py , Python, 243 lines, 2 matches - code/
run_litesensor_t1.py , Python, 200 lines, 1 match - code/
run_m1_flops_cycles.py , Python, 98 lines - code/
run_m2_deployment_artifa , Python, 133 linesct.py - code/
run_random_split_kdm.py , Python, 100 lines - code/
run_s1_30splits.py , Python, 122 lines, 1 match - code/
run_t3_qat.py , Python, 202 lines, 1 match - code/
run_t4_loo.py , Python, 150 lines - code/
run_t67.py , Python, 150 lines - code/
run_t8_worner.py , Python, 252 lines - code/
run_table3_correction.py , Python, 162 lines - code/
scripts/ , Shell, 38 linesdownload_data.sh - code/
scripts/ , Python, 36 linesmcu_deployment/ npz_to_c_array.py - code/
scripts/ , Shell, 64 linesreproduce_all_paper.sh - code/
scripts/ , Shell, 58 linesreproduce_table3.sh - code/
tests/ , Python, 1 line__init__.py - code/
tests/ , Python, 128 linestest_smoke.py - LICENSE, License, 19 lines
- README.md, Text, 98 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 34 scripts, each with its path and the digest of its content;
- 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
The UCI Gas Sensor Array Drift Dataset is publicly available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 6 keywords, 30 references.
Cite
This paper
Kim, S., Shin, M., Kang, K., Lee, D.-H., Churchill, D. G., & Jang, Y. J. (2026). Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping. Molecules (Basel, Switzerland), 31(11), 1884. https://
BibTeX
@article{kim2026drift,
author = {Kim, Soohwan and Shin, Myeongsik and Kang, Ku and Lee, Doo-Hee and Churchill, David G and Jang, Yoon Jeong},
title = {{Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping}},
journal = {Molecules (Basel, Switzerland)},
year = {2026},
month = jun,
volume = {31},
number = {11},
pages = {1884},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1420-3049},
doi = {10.3390/
url = {https://
pmid = {42280188},
pmcid = {PMC13257611}
}
RIS
TY - JOUR
AU - Kim, Soohwan
AU - Shin, Myeongsik
AU - Kang, Ku
AU - Lee, Doo-Hee
AU - Churchill, David G
AU - Jang, Yoon Jeong
TI - Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping
T2 - Molecules (Basel, Switzerland)
J2 - Molecules
PY - 2026
DA - 2026/
VL - 31
IS - 11
SP - 1884
SN - 1420-3049
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "Drift-Robust Lightweight Deep Learning on Open Gas Sensor Benchmarks: A Reproducible Architecture Study with CBRN Applicability Mapping",
"container-title": "Molecules (Basel, Switzerland)",
"author": [
{
"family": "Kim",
"given": "Soohwan"
},
{
"family": "Shin",
"given": "Myeongsik"
},
{
"family": "Kang",
"given": "Ku"
},
{
"family": "Lee",
"given": "Doo-Hee"
},
{
"family": "Churchill",
"given": "David G"
},
{
"family": "Jang",
"given": "Yoon Jeong"
}
],
"container-title-short":
"volume": "31",
"issue": "11",
"page": "1884",
"DOI": "10.3390/
"PMID": "42280188",
"PMCID": "PMC13257611",
"ISSN": "1420-3049",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.isci.2026.117206 [code]
- ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.Journal: iScienceIn common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
- [2] doi:10.1038/s41746-026-02778-0 [code]
- Trust-gated synthetic EEG augmentation reduces performance drops when generalizing to new patients.Journal: NPJ digital medicineIn common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
- [3] doi:10.3390/diagnostics16111588 [code]
- Bridging Annotation Gaps: Hierarchical Self-Support Learning for Brain Tumor Segmentation.Journal: Diagnostics (Basel, Switzerland)In common: PyTorch, SciPy, NumPy, methods / tools, 1 reference
- [4] doi:10.1038/s41598-026-54785-6 [code]
- Lightweight deep learning model for nonconvulsive status epilepticus diagnosis using EEG time-frequency analysis.Journal: Scientific reportsIn common: PyTorch, SciPy, NumPy, 1 reference
- [5] doi:10.1038/s41598-026-54840-2 [code]
- Predictive metacognition: a neuro-computational framework for self-monitoring in large language models.Journal: Scientific reportsIn common: PyTorch, pandas, NumPy, 1 reference
- [6] doi:10.3390/jimaging12060233 [code]
- Brain Tumor Classification in MRI Images Using Combined Transfer Learning and Convolutional Neural Networks.Journal: Journal of imagingIn common: scikit-learn, pandas, NumPy, 1 reference
- [7] doi:10.1016/j.patter.2026.101564 [code]
- Spacing effect improves generalization in biological and artificial systems.Journal: Patterns (New York, N.Y.)In common: PyTorch, NumPy, 1 reference
- [8] doi:10.1038/s41598-026-50603-1
- A multi-class framework for face mask compliance detection using lightweight deep learning models.Journal: Scientific reportsIn common: methods / tools, 1 reference
- [9] doi:
- NeuroTrustNet: a cost-effective multimodal ensemble framework for brain tumor classification under cross-dataset variabilityJournal: Frontiers in artificial intelligenceIn common: methods / tools, 1 reference
- [10] doi:10.3389/frai.2026.1849571
- NeuroPlast: a learnable activation function evaluated under knowledge distillation for medical image classification.Journal: Frontiers in artificial intelligenceIn common: methods / tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 34 scripts, and 16 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:878874e9bc91790e…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
