OSCR

Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled Radio-Frequency Measurement-Replay Testbed.

Code ↔ Paper

40 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 40 matches
  1. [1] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ paper_aligned_results_and_audit/analysis/scripts/plot_revision_results.py, lines 890–939 · score 0.97 · frozen scan stratified, seed confidence intervals, occur twice, error bars, transfer scans, Pilot scans
  2. [2] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ paper_aligned_results_and_audit/analysis/scripts/plot_revision_results.py, lines 890–939 · score 0.96 · cross configuration evaluations, horizontal bars, Predeclared primary paired, comparison primary family, HNP DQN minus, paired training seeds
  3. [3] § 4. Experiments and Results Analysis › 4.6. Post Hoc Analysis of Expanded-Feature Weights ↔ paper_aligned_results_and_audit/analysis/scripts/audit_polynomial_branch_weights.py, lines 1–47 · score 0.93 · branch weights, descriptive raw weight, upstream activation scales, polynomial branch, descriptive audit, ignores activation frequency
  4. [4] § 3. HNP-DQN Architecture › 3.2. Capacity-Matched Conventional Control ↔ experiment_v2/src/models.py, lines 120–181 · score 0.85 · BatchNorm, hidden widths, ReLU, matched MLP DDQN, dueling head, HNP DQN
  5. [5] § 3. HNP-DQN Architecture › 3.2. Capacity-Matched Conventional Control ↔ experiment_v2/src/models.py, lines 120–181 · score 0.84 · post expansion, hidden widths, dueling head, LayerNorm, capacity matched, causal
  6. [6] § 4. Experiments and Results Analysis › 4.6. Post Hoc Analysis of Expanded-Feature Weights ↔ paper_aligned_results_and_audit/analysis/scripts/audit_polynomial_branch_weights.py, lines 326–376 · score 0.82 · absolute outgoing weight, squared minus linear, Frobenius norms, branch, jammer mode, audited
  7. [7] § 4. Experiments and Results Analysis › 4.4. Model Capacity and Computational Cost ↔ experiment_v2/src/runner.py, lines 394–411 · score 0.82 · temporary workspaces, excludes activations, Persistent tensor, registered buffers, optimizer state, allocator
  8. [8] § 4. Experiments and Results Analysis › 4.2. Endpoints and Statistical Analysis ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 1305–1359 · score 0.81 · predeclared primary family, exploratory Holm family, flip enumeration, fixed evaluation trajectories, paired training seeds, Holm correction
  9. [9] § 2. Dataset and Environment Modeling › 2.4. Agent Observation Representation ↔ experiment_v2/src/baselines.py, lines 1–24 · score 0.81 · jammer aware greedy, schedule aware sweep, ordinary observation, phase, baselines, oracle
  10. [10] § 2. Dataset and Environment Modeling › 2.1. RF Jamming Dataset ↔ experiment_v2/audit_data.py, lines 75–152 · score 0.79 · distance shift, power shift, Scan IDs, audits, dBm, split
  11. [11] § 4. Experiments and Results Analysis › 4.2. Endpoints and Statistical Analysis ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 35–110 · score 0.79 · jammer aware greedy, clairvoyant oracle, capacity matched MLP, schedule aware, HNP DQN, hysteresis
  12. [12] § 4. Experiments and Results Analysis › 4.4. Model Capacity and Computational Cost ↔ experiment_v2/profile_model_costs.py, lines 10–68 · score 0.78 · allocator overhead, temporary workspaces, excludes activations, registered buffers, pacing, thread
  13. [13] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ experiment_v2/src/baselines.py, lines 255–297 · score 0.77 · preceding step, Schedule aware sweeping, safe channel, RF observation, swept, phase
  14. [14] § 3. HNP-DQN Architecture › 3.1. Proposed Architecture ↔ paper_aligned_results_and_audit/analysis/scripts/plot_hnp_architecture.py, lines 123–261 · score 0.77 · TD objective, ReLU, online network, target network, linear, expansion
  15. [15] § 4. Experiments and Results Analysis › 4.4. Model Capacity and Computational Cost ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 707–775 · score 0.74 · GPU inference, persistent tensor, serialized state, trainable parameters, repetition, memory
  16. [16] § 4. Experiments and Results Analysis › 4.2. Endpoints and Statistical Analysis ↔ paper_aligned_results_and_audit/analysis/scripts/build_revision_artifacts.py, lines 836–955 · score 0.73 · flip enumeration, Holm family, paired seed, Cohen, trained seed, jammer mode
  17. [17] § 3. HNP-DQN Architecture › 3.1. Proposed Architecture ↔ paper_aligned_results_and_audit/analysis/scripts/plot_hnp_architecture.py, lines 123–261 · score 0.73 · online network, Polyak update, target network, RF entries, MSE, concatenates
  18. [18] § 2. Dataset and Environment Modeling › 2.6. Reward Design and Performance Objectives ↔ experiment_v2/src/env.py, lines 281–331 · score 0.72 · successful retention, successful switch, switching cost, dominant, event, collision
  19. [19] § 4. Experiments and Results Analysis › 4.4. Model Capacity and Computational Cost ↔ experiment_v2/src/runner.py, lines 436–491 · score 0.72 · reports trainable parameters, persistent tensor, serialized state, GPU, latency, CPU
  20. [20] § 2. Dataset and Environment Modeling › 2.4. Agent Observation Representation ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 35–110 · score 0.72 · jammer aware greedy, schedule aware sweep, oracle, privileged, ordinary, clairvoyant
  21. [21] § 3. HNP-DQN Architecture › 3.1. Proposed Architecture ↔ paper_aligned_results_and_audit/analysis/scripts/map_manuscript_locations.py, lines 47–106 · score 0.71 · neural feature extractor, HNP DQN combines, channel encoding, RF summaries, LayerNorm, expansion
  22. [22] § 3. HNP-DQN Architecture › 3.4. Agent Training Procedure ↔ experiment_v2/src/runner.py, lines 161–226 · score 0.70 · terminal state, stored transition, absorbing, mask, truncation, bootstrap
  23. [23] § 2. Dataset and Environment Modeling › 2.6. Reward Design and Performance Objectives ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 1490–1549 · score 0.69 · Supporting endpoints, inferential unit, fixed trajectories, independently trained seed, family, Holm
  24. [24] § 2. Dataset and Environment Modeling › 2.1. RF Jamming Dataset ↔ paper_aligned_results_and_audit/analysis/scripts/plot_training_snr_direction.py, lines 1–25 · score 0.69 · maximum quality, snr field, minimum interference, raw, scan, training
  25. [25] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ paper_aligned_results_and_audit/analysis/scripts/map_manuscript_locations.py, lines 107–166 · score 0.67 · Cross model TD, loss magnitude, TD loss, bootstrap targets, polynomial mapping, LayerNorm
  26. [26] § 3. HNP-DQN Architecture › 3.4. Agent Training Procedure ↔ paper_aligned_results_and_audit/analysis/scripts/map_manuscript_locations.py, lines 167–201 · score 0.67 · absorbing environmental terminal, limit truncation, step boundary, terminal state, Agent, Training
  27. [27] § 4. Experiments and Results Analysis › 4.5. Ablation Study ↔ paper_aligned_results_and_audit/analysis/scripts/build_revision_artifacts.py, lines 1–44 · score 0.66 · exploratory Holm family, primary comparison family, full HNP, capacity matched, training seed, variant
  28. [28] § 2. Dataset and Environment Modeling › 2.1. RF Jamming Dataset ↔ paper_aligned_results_and_audit/analysis/scripts/plot_training_snr_direction.py, lines 1–25 · score 0.66 · fixed jammer source, snr field, Development training, audit, quality, pilot
  29. [29] § 3. HNP-DQN Architecture › 3.4. Agent Training Procedure ↔ agents/DQN_agents/HNP_DQN.py, lines 65–100 · score 0.66 · soft update, replay buffer, target network, greedy, optimize, batch
  30. [30] § 3. HNP-DQN Architecture › 3.4. Agent Training Procedure ↔ experiment_v2/src/agent.py, lines 372–409 · score 0.65 · soft update, online network, target network, optimize, batch, replay
  31. [31] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ experiment_v2/src/baselines.py, lines 1–24 · score 0.65 · Jammer aware greedy, upper bound, jammed channel, privileged, quality
  32. [32] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ experiment_v2/src/baselines.py, lines 356–467 · score 0.64 · dynamic programming, upper bound, switching cost, DP, Clairvoyant, horizon
  33. [33] § 3. HNP-DQN Architecture › 3.3. Training and Target-Update Procedure ↔ experiment_v2/src/agent.py, lines 372–409 · score 0.64 · online network selects, target network evaluates, prediction, gradient, Training
  34. [34] § 3. HNP-DQN Architecture › 3.1. Proposed Architecture ↔ experiment_v2/src/models.py, lines 58–118 · score 0.64 · polynomial expansion, LayerNorm, concatenates, hidden, module, heads
  35. [35] § 4. Experiments and Results Analysis › 4.4. Model Capacity and Computational Cost ↔ paper_aligned_results_and_audit/analysis/scripts/summarize_revision_results.py, lines 1081–1152 · score 0.64 · serialized state, persistent parameter, KiB, median, engineering, HNP DQN
  36. [36] § 3. HNP-DQN Architecture › 3.3. Training and Target-Update Procedure ↔ agents/DQN_agents/DQN_With_Fixed_Q_Targets.py, lines 6–33 · score 0.60 · learning iteration, soft update, target network, Training
  37. [37] § 4. Experiments and Results Analysis › 4.1. Experimental Setup and Parameter Configuration ↔ paper_aligned_results_and_audit/analysis/scripts/map_manuscript_locations.py, lines 47–106 · score 0.60 · formal endpoint, scan stratified, learned policy, fixed trajectories, jammer mode, auditing
  38. [38] § 3. HNP-DQN Architecture › 3.1. Proposed Architecture › 3.1.2. Polynomial Mapping Layer ↔ experiment_v2/src/models.py, lines 58–118 · score 0.58 · polynomial basis, cross feature, expansion, squares, layer, dimensional
  39. [39] § 2. Dataset and Environment Modeling › 2.6. Reward Design and Performance Objectives ↔ paper_aligned_results_and_audit/analysis/scripts/map_manuscript_locations.py, lines 167–201 · score 0.57 · reward prioritizes interference, switching indicator, avoidance, Objectives
  40. [40] § 4. Experiments and Results Analysis › 4.3. Comparative Experiments ↔ experiment_v2/src/baselines.py, lines 150–249 · score 0.56 · post hoc, quality scores, hysteresis, reset, threshold, policy

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,629 lines · 74 KB · no license · 6 matches

  1. #!/usr/bin/env python3
  2. """Validate frozen revision artifacts and prepare auditable result text.
  3. This is a post-processing gate, not an experiment runner. It accepts one
  4. explicit canonical primary-run directory, one explicit directory produced by
  5. ``build_revision_artifacts.py``, and one explicit isolated cost-profile
  6. directory. It does not discover runs. It validates every input before it
  7. creates either output file.
  8. The generated prose is deliberately conservative. Mechanical labels report
  9. only whether the paired confidence interval and Holm-adjusted test jointly
  10. support an advantage, support a disadvantage, or are inconclusive. Practical
  11. importance, narrative emphasis, and the final language/caption audit remain
  12. human decisions.
  13. """
  14. from __future__ import annotations
  15. import argparse
  16. from datetime import datetime, timezone
  17. import hashlib
  18. import json
  19. import math
  20. from pathlib import Path
  21. import re
  22. import sys
  23. from typing import Any, Iterable, Mapping, Sequence
  24. import numpy as np
  25. import pandas as pd
  26. import build_revision_artifacts as gate
  27. SCHEMA_VERSION = 1
  28. EXPECTED_MODELS = ("hnp", "matched")
  29. EXPECTED_CONDITIONS = gate.EXPECTED_CONDITIONS
  30. EXPECTED_JAMMERS = gate.EXPECTED_JAMMERS
  31. EXPECTED_TRAIN_SEEDS = gate.EXPECTED_TRAIN_SEEDS
  32. EXPECTED_EVAL_SEEDS = gate.EXPECTED_EVAL_SEEDS
  33. EXPECTED_PRIMARY_COLUMNS = (
  34. "condition",
  35. "jammer_mode",
  36. "method_a",
  37. "method_b",
  38. "metric",
  39. "n_pairs",
  40. "mean_a",
  41. "sd_a",
  42. "mean_b",
  43. "sd_b",
  44. "mean_difference",
  45. "ci95_low",
  46. "ci95_high",
  47. "effect_dz",
  48. "p_exact",
  49. "p_holm",
  50. )
  51. EXPECTED_VARIANTS = (
  52. "no_polynomial",
  53. "no_layernorm",
  54. "no_dueling",
  55. "hnp_gamma0",
  56. )
  57. EXPECTED_BASELINES = {
  58. "sweeping": (
  59. "threshold",
  60. "max_quality",
  61. "stay",
  62. "random",
  63. "schedule_sweep",
  64. "jammer_greedy",
  65. "clairvoyant_oracle",
  66. ),
  67. "random": (
  68. "threshold",
  69. "max_quality",
  70. "stay",
  71. "random",
  72. "jammer_greedy",
  73. "clairvoyant_oracle",
  74. ),
  75. }
  76. ORDINARY_BASELINES = {
  77. "sweeping": ("threshold", "max_quality", "stay", "random", "schedule_sweep"),
  78. "random": ("threshold", "max_quality", "stay", "random"),
  79. }
  80. METHOD_LABELS = {
  81. "hnp": "HNP-DQN",
  82. "matched": "capacity-matched MLP-DDQN",
  83. "threshold": "threshold/hysteresis",
  84. "max_quality": "maximum measured quality",
  85. "stay": "stay",
  86. "random": "random selection",
  87. "schedule_sweep": "schedule-aware sweep rule",
  88. "jammer_greedy": "jammer-aware greedy reference (privileged)",
  89. "clairvoyant_oracle": "clairvoyant DP upper bound (privileged)",
  90. "no_polynomial": "No-Polynomial",
  91. "no_layernorm": "No-LayerNorm",
  92. "no_dueling": "No-Dueling",
  93. "hnp_gamma0": r"HNP with $\gamma=0$",
  94. }
  95. CONDITION_LABELS = {
  96. "within_condition_pilot": "20 cm/10 dBm transparent pilot",
  97. "distance_shift_40cm_10dBm": "40 cm/10 dBm distance transfer",
  98. "power_shift_20cm_5dBm": "20 cm/5 dBm power transfer",
  99. }
  100. MANUSCRIPT_PATTERN = re.compile(r"\[VERIFIED RESULT:\s*(.*?)\]", re.DOTALL)
  101. RESPONSE_PATTERN = re.compile(r"\[\[VERIFIED_RESULT:(.*?)\]\]", re.DOTALL)
  102. HEX64 = re.compile(r"^[0-9a-f]{64}$")
  103. class ValidationError(gate.ValidationError):
  104. """Raised before any output is written when an artifact is not canonical."""
  105. def _sha256(path: Path) -> str:
  106. digest = hashlib.sha256()
  107. try:
  108. with path.open("rb") as handle:
  109. for chunk in iter(lambda: handle.read(1024 * 1024), b""):
  110. digest.update(chunk)
  111. except OSError as exc:
  112. raise ValidationError(f"cannot hash {path}: {exc}") from exc
  113. return digest.hexdigest()
  114. def _load_csv(path: Path, label: str) -> pd.DataFrame:
  115. if not path.is_file():
  116. raise ValidationError(f"required {label} is missing: {path}")
  117. try:
  118. return pd.read_csv(path)
  119. except Exception as exc:
  120. raise ValidationError(f"cannot parse {label} {path}: {exc}") from exc
  121. def _load_json(path: Path, label: str) -> Any:
  122. if not path.is_file():
  123. raise ValidationError(f"required {label} is missing: {path}")
  124. try:
  125. return json.loads(path.read_text(encoding="utf-8"))
  126. except (OSError, json.JSONDecodeError) as exc:
  127. raise ValidationError(f"cannot parse {label} {path}: {exc}") from exc
  128. def _require_columns(frame: pd.DataFrame, columns: Iterable[str], label: str) -> None:
  129. missing = sorted(set(columns) - set(frame.columns))
  130. if missing:
  131. raise ValidationError(f"{label} is missing columns: {missing}")
  132. def _json_safe(value: Any) -> Any:
  133. if isinstance(value, Path):
  134. return str(value)
  135. if isinstance(value, np.generic):
  136. value = value.item()
  137. if isinstance(value, float) and not math.isfinite(value):
  138. return None
  139. if isinstance(value, Mapping):
  140. return {str(key): _json_safe(item) for key, item in value.items()}
  141. if isinstance(value, (list, tuple)):
  142. return [_json_safe(item) for item in value]
  143. return value
  144. def _normalised_scalar(value: Any) -> Any:
  145. if isinstance(value, np.generic):
  146. return value.item()
  147. return value
  148. def _equal_scalar(first: Any, second: Any) -> bool:
  149. first = _normalised_scalar(first)
  150. second = _normalised_scalar(second)
  151. if pd.isna(first) and pd.isna(second):
  152. return True
  153. if isinstance(first, (bool, np.bool_)) or isinstance(second, (bool, np.bool_)):
  154. return str(first).lower() == str(second).lower()
  155. if isinstance(first, (int, float, np.number)) and isinstance(
  156. second, (int, float, np.number)
  157. ):
  158. return bool(
  159. np.isclose(float(first), float(second), rtol=1e-10, atol=1e-10, equal_nan=True)
  160. )
  161. return str(first) == str(second)
  162. def _assert_frame_matches(
  163. actual: pd.DataFrame,
  164. expected: pd.DataFrame,
  165. *,
  166. keys: Sequence[str],
  167. label: str,
  168. exact_columns: bool = True,
  169. ) -> None:
  170. """Compare two keyed tables, including every reported statistic."""
  171. if exact_columns and set(actual.columns) != set(expected.columns):
  172. raise ValidationError(
  173. f"{label} column set differs; missing={sorted(set(expected.columns)-set(actual.columns))}, "
  174. f"extra={sorted(set(actual.columns)-set(expected.columns))}"
  175. )
  176. _require_columns(actual, keys, label)
  177. _require_columns(expected, keys, f"recomputed {label}")
  178. if actual.duplicated(list(keys)).any() or expected.duplicated(list(keys)).any():
  179. raise ValidationError(f"{label} has duplicate comparison keys")
  180. if len(actual) != len(expected):
  181. raise ValidationError(
  182. f"{label} has {len(actual)} rows but recomputation has {len(expected)}"
  183. )
  184. left = actual.sort_values(list(keys)).reset_index(drop=True)
  185. right = expected.sort_values(list(keys)).reset_index(drop=True)
  186. columns = list(expected.columns)
  187. for row_index in range(len(right)):
  188. for column in columns:
  189. if not _equal_scalar(left.iloc[row_index][column], right.iloc[row_index][column]):
  190. key = tuple(right.iloc[row_index][item] for item in keys)
  191. raise ValidationError(
  192. f"{label} mismatch for {key}, column {column}: "
  193. f"reported={left.iloc[row_index][column]!r}, "
  194. f"recomputed={right.iloc[row_index][column]!r}"
  195. )
  196. def _learned_rows(summary: pd.DataFrame, method: str, condition: str, jammer: str) -> pd.DataFrame:
  197. rows = summary[
  198. (summary["condition"] == condition)
  199. & (summary["jammer_mode"] == jammer)
  200. & (summary["method"] == method)
  201. & summary["train_seed"].notna()
  202. ].copy()
  203. rows["train_seed"] = rows["train_seed"].astype(int)
  204. rows = rows.sort_values("train_seed")
  205. seeds = tuple(int(value) for value in rows["train_seed"])
  206. if seeds != EXPECTED_TRAIN_SEEDS or rows.duplicated("train_seed").any():
  207. raise ValidationError(
  208. f"{condition}/{jammer}/{method} requires exactly seeds {EXPECTED_TRAIN_SEEDS}; "
  209. f"got {seeds}"
  210. )
  211. return rows
  212. def _primary_paired_row(
  213. first: np.ndarray,
  214. second: np.ndarray,
  215. *,
  216. condition: str,
  217. jammer: str,
  218. comparator: str,
  219. ) -> dict[str, Any]:
  220. """Mirror ``experiment_v2/src/statistics.py`` for an independent check.
  221. The primary runner uses NumPy's default ``isclose`` tolerance for the
  222. zero-variance Cohen-dz edge case, whereas the exploratory artifact builder
  223. intentionally has its own tighter helper. Keeping that distinction here
  224. prevents a legitimate primary CSV from being rejected only in a nearly
  225. constant-difference edge case.
  226. """
  227. first = np.asarray(first, dtype=float).reshape(-1)
  228. second = np.asarray(second, dtype=float).reshape(-1)
  229. if first.shape != second.shape or first.size != 10:
  230. raise ValidationError(f"primary comparison {comparator} needs 10 paired seeds")
  231. if not np.isfinite(first).all() or not np.isfinite(second).all():
  232. raise ValidationError(f"primary comparison {comparator} has non-finite values")
  233. differences = first - second
  234. mean_difference = float(differences.mean())
  235. sd_difference = float(differences.std(ddof=1))
  236. margin = (
  237. float(gate.student_t.ppf(0.975, differences.size - 1))
  238. * sd_difference
  239. / math.sqrt(differences.size)
  240. )
  241. if np.isclose(sd_difference, 0.0):
  242. effect = 0.0 if np.isclose(mean_difference, 0.0) else float(
  243. np.copysign(np.inf, mean_difference)
  244. )
  245. else:
  246. effect = mean_difference / sd_difference
  247. return {
  248. "condition": condition,
  249. "jammer_mode": jammer,
  250. "method_a": "hnp",
  251. "method_b": comparator,
  252. "metric": "return",
  253. "n_pairs": 10,
  254. "mean_a": float(first.mean()),
  255. "sd_a": float(first.std(ddof=1)),
  256. "mean_b": float(second.mean()),
  257. "sd_b": float(second.std(ddof=1)),
  258. "mean_difference": mean_difference,
  259. "ci95_low": mean_difference - margin,
  260. "ci95_high": mean_difference + margin,
  261. "effect_dz": effect,
  262. "p_exact": gate._exact_sign_flip_pvalue(differences),
  263. "p_holm": float("nan"),
  264. }
  265. def _recompute_primary_comparisons(seed_summary: pd.DataFrame) -> pd.DataFrame:
  266. """Recompute the 12 declared rows from seed-level endpoints."""
  267. _require_columns(
  268. seed_summary,
  269. {
  270. "condition",
  271. "jammer_mode",
  272. "method",
  273. "train_seed",
  274. "return",
  275. },
  276. "primary seed_summary",
  277. )
  278. rows: list[dict[str, Any]] = []
  279. for jammer in EXPECTED_JAMMERS:
  280. heuristic = "schedule_sweep" if jammer == "sweeping" else "threshold"
  281. for condition in EXPECTED_CONDITIONS:
  282. hnp = _learned_rows(seed_summary, "hnp", condition, jammer)
  283. for comparator in (heuristic, "matched"):
  284. if comparator == "matched":
  285. other = _learned_rows(seed_summary, comparator, condition, jammer)
  286. joined = hnp[["train_seed", "return"]].merge(
  287. other[["train_seed", "return"]],
  288. on="train_seed",
  289. how="outer",
  290. validate="one_to_one",
  291. suffixes=("_hnp", "_other"),
  292. indicator=True,
  293. )
  294. if set(joined["_merge"]) != {"both"}:
  295. raise ValidationError(
  296. f"unpaired HNP/matched seeds for {condition}/{jammer}"
  297. )
  298. joined = joined.sort_values("train_seed")
  299. first = joined["return_hnp"].to_numpy(dtype=float)
  300. second = joined["return_other"].to_numpy(dtype=float)
  301. else:
  302. baseline = seed_summary[
  303. (seed_summary["condition"] == condition)
  304. & (seed_summary["jammer_mode"] == jammer)
  305. & (seed_summary["method"] == comparator)
  306. & seed_summary["train_seed"].isna()
  307. ]
  308. if len(baseline) != 1:
  309. raise ValidationError(
  310. f"{condition}/{jammer}/{comparator} requires one fixed-trajectory mean; "
  311. f"got {len(baseline)}"
  312. )
  313. first = hnp["return"].to_numpy(dtype=float)
  314. second = np.full(first.shape, float(baseline.iloc[0]["return"]))
  315. row = _primary_paired_row(
  316. first,
  317. second,
  318. condition=condition,
  319. jammer=jammer,
  320. comparator=comparator,
  321. )
  322. rows.append(row)
  323. frame = pd.DataFrame(rows)
  324. if len(frame) != 12:
  325. raise ValidationError(f"primary comparison recomputation produced {len(frame)} rows")
  326. frame["p_holm"] = gate._holm_adjust(frame["p_exact"].to_numpy(dtype=float))
  327. return frame[list(EXPECTED_PRIMARY_COLUMNS)]
  328. def _validate_primary_comparisons(
  329. reported: pd.DataFrame, seed_summary: pd.DataFrame
  330. ) -> pd.DataFrame:
  331. if set(reported.columns) != set(EXPECTED_PRIMARY_COLUMNS):
  332. raise ValidationError(
  333. "primary_comparisons.csv does not have the exact declared schema; "
  334. f"expected={list(EXPECTED_PRIMARY_COLUMNS)}, got={list(reported.columns)}"
  335. )
  336. expected = _recompute_primary_comparisons(seed_summary)
  337. _assert_frame_matches(
  338. reported,
  339. expected,
  340. keys=("condition", "jammer_mode", "method_b"),
  341. label="primary_comparisons.csv",
  342. )
  343. return _canonical_primary_order(reported)
  344. def _canonical_primary_order(frame: pd.DataFrame) -> pd.DataFrame:
  345. result = frame.copy()
  346. result["_jammer"] = result["jammer_mode"].map(
  347. {name: index for index, name in enumerate(EXPECTED_JAMMERS)}
  348. )
  349. result["_condition"] = result["condition"].map(
  350. {name: index for index, name in enumerate(EXPECTED_CONDITIONS)}
  351. )
  352. result["_comparator"] = [
  353. 1 if method == "matched" else 0 for method in result["method_b"]
  354. ]
  355. return (
  356. result.sort_values(["_jammer", "_condition", "_comparator"])
  357. .drop(columns=["_jammer", "_condition", "_comparator"])
  358. .reset_index(drop=True)
  359. )
  360. def _validate_primary_run(primary_dir: Path) -> tuple[gate.ValidatedRun, pd.DataFrame]:
  361. run = gate._validate_one_run(
  362. gate.RunSpec("primary", Path(primary_dir), EXPECTED_MODELS, 0.95)
  363. )
  364. comparisons = _load_csv(run.spec.path / "primary_comparisons.csv", "primary comparisons")
  365. return run, _validate_primary_comparisons(comparisons, run.seed_summary)
  366. def _manifest_run_matches(record: Mapping[str, Any], run: gate.ValidatedRun) -> None:
  367. expected_scalars = {
  368. "path": str(run.spec.path),
  369. "expected_models": list(run.spec.expected_models),
  370. "gamma": float(run.config["gamma"]),
  371. "config_sha256": run.hashes["config_sha256"],
  372. "core_code_bundle_sha256": run.hashes["core_code_bundle_sha256"],
  373. "predeclared_schedule_sha256": run.hashes["predeclared_schedule_sha256"],
  374. "normalizer_sha256": run.hashes["normalizer_sha256"],
  375. "input_artifact_sha256": run.hashes["input_artifact_sha256"],
  376. "checkpoint_sha256": run.hashes["checkpoint_sha256"],
  377. "validation_counts": run.validation_counts,
  378. }
  379. for key, expected in expected_scalars.items():
  380. if _json_safe(record.get(key)) != _json_safe(expected):
  381. raise ValidationError(
  382. f"result_manifest run {run.spec.label} has stale or mismatched {key}"
  383. )
  384. def _validate_ablation_directory(
  385. ablation_dir: Path, primary: gate.ValidatedRun
  386. ) -> tuple[pd.DataFrame, dict[str, gate.ValidatedRun], dict[str, Any]]:
  387. root = Path(ablation_dir).expanduser().resolve()
  388. if not root.is_dir():
  389. raise ValidationError(f"ablation artifact directory does not exist: {root}")
  390. manifest_path = root / "result_manifest.json"
  391. manifest = _load_json(manifest_path, "ablation result manifest")
  392. if not isinstance(manifest, Mapping):
  393. raise ValidationError("ablation result_manifest.json must contain an object")
  394. if manifest.get("schema_version") != gate.SCHEMA_VERSION or manifest.get("status") != "validated":
  395. raise ValidationError("ablation manifest is not a validated schema-v1 artifact")
  396. run_records = manifest.get("runs")
  397. if not isinstance(run_records, Mapping) or set(run_records) != {
  398. "primary",
  399. "gamma0",
  400. "no_polynomial",
  401. "no_layernorm",
  402. "no_dueling",
  403. }:
  404. raise ValidationError("ablation manifest must contain exactly the five declared run roles")
  405. specs = {
  406. "primary": gate.RunSpec("primary", Path(run_records["primary"]["path"]), ("hnp", "matched"), 0.95),
  407. "gamma0": gate.RunSpec(
  408. "gamma0", Path(run_records["gamma0"]["path"]), ("hnp",), 0.0,
  409. variant_method="hnp", variant_label="hnp_gamma0"
  410. ),
  411. "no_polynomial": gate.RunSpec(
  412. "no_polynomial", Path(run_records["no_polynomial"]["path"]), ("no_polynomial",), 0.95,
  413. variant_method="no_polynomial", variant_label="no_polynomial"
  414. ),
  415. "no_layernorm": gate.RunSpec(
  416. "no_layernorm", Path(run_records["no_layernorm"]["path"]), ("no_layernorm",), 0.95,
  417. variant_method="no_layernorm", variant_label="no_layernorm"
  418. ),
  419. "no_dueling": gate.RunSpec(
  420. "no_dueling", Path(run_records["no_dueling"]["path"]), ("no_dueling",), 0.95,
  421. variant_method="no_dueling", variant_label="no_dueling"
  422. ),
  423. }
  424. if specs["primary"].path.expanduser().resolve() != primary.spec.path:
  425. raise ValidationError(
  426. "ablation manifest primary path is not the explicitly supplied canonical primary run"
  427. )
  428. runs: dict[str, gate.ValidatedRun] = {"primary": primary}
  429. for label in ("gamma0", "no_polynomial", "no_layernorm", "no_dueling"):
  430. runs[label] = gate._validate_one_run(specs[label])
  431. cross = gate._validate_cross_run_provenance(list(runs.values()))
  432. if _json_safe(manifest.get("cross_run_provenance")) != _json_safe(cross):
  433. raise ValidationError("ablation manifest cross-run provenance is stale")
  434. for label, run in runs.items():
  435. record = run_records.get(label)
  436. if not isinstance(record, Mapping):
  437. raise ValidationError(f"ablation manifest run record {label} is malformed")
  438. _manifest_run_matches(record, run)
  439. output_record = manifest.get("outputs", {}).get("ablation_comparisons")
  440. if not isinstance(output_record, Mapping):
  441. raise ValidationError("ablation manifest lacks the comparison output record")
  442. if output_record.get("path") != "ablation_comparisons.csv":
  443. raise ValidationError("ablation comparison path must be ablation_comparisons.csv")
  444. comparison_path = root / "ablation_comparisons.csv"
  445. if str(output_record.get("sha256", "")).lower() != _sha256(comparison_path):
  446. raise ValidationError("ablation comparison hash does not match the manifest")
  447. reported = _load_csv(comparison_path, "ablation comparisons")
  448. if int(output_record.get("rows", -1)) != len(reported) or list(
  449. output_record.get("columns", [])
  450. ) != list(reported.columns):
  451. raise ValidationError("ablation comparison row/column manifest is stale")
  452. expected = gate._build_ablation_comparisons(
  453. runs["primary"],
  454. [runs["gamma0"], runs["no_polynomial"], runs["no_layernorm"], runs["no_dueling"]],
  455. )
  456. _assert_frame_matches(
  457. reported,
  458. expected,
  459. keys=("condition", "jammer_mode", "method_b"),
  460. label="ablation_comparisons.csv",
  461. )
  462. protocol = manifest.get("protocol", {})
  463. expected_protocol = {
  464. "train_seeds": list(EXPECTED_TRAIN_SEEDS),
  465. "eval_seeds": list(EXPECTED_EVAL_SEEDS),
  466. "trajectories_per_method_training_seed_condition_jammer": 20,
  467. "conditions": list(EXPECTED_CONDITIONS),
  468. "jammer_modes": list(EXPECTED_JAMMERS),
  469. "metric": "return",
  470. "holm_family": gate.EXPLORATORY_FAMILY,
  471. "holm_family_size": 24,
  472. }
  473. for key, value in expected_protocol.items():
  474. if protocol.get(key) != value:
  475. raise ValidationError(f"ablation manifest protocol field {key} is invalid")
  476. return expected.reset_index(drop=True), runs, dict(manifest)
  477. def _records_match_csv(frame: pd.DataFrame, records: Sequence[Mapping[str, Any]]) -> None:
  478. union = set().union(*(record.keys() for record in records)) if records else set()
  479. if set(frame.columns) != union:
  480. raise ValidationError(
  481. f"cost CSV/JSON column union differs; csv-only={sorted(set(frame.columns)-union)}, "
  482. f"json-only={sorted(union-set(frame.columns))}"
  483. )
  484. if len(frame) != len(records):
  485. raise ValidationError("cost CSV and JSON have different row counts")
  486. for index, record in enumerate(records):
  487. for column in frame.columns:
  488. expected = record.get(column)
  489. actual = frame.iloc[index][column]
  490. if not _equal_scalar(actual, expected):
  491. raise ValidationError(
  492. f"cost CSV/JSON mismatch at row {index}, column {column}: "
  493. f"csv={actual!r}, json={expected!r}"
  494. )
  495. def _finite_nonnegative(frame: pd.DataFrame, columns: Sequence[str], label: str) -> None:
  496. _require_columns(frame, columns, label)
  497. values = frame[list(columns)].to_numpy(dtype=float)
  498. if not np.isfinite(values).all() or (values < 0).any():
  499. raise ValidationError(f"{label} contains non-finite or negative values in {columns}")
  500. def _int_set(values: Iterable[Any], label: str) -> set[int]:
  501. try:
  502. converted = {int(value) for value in values}
  503. except (TypeError, ValueError) as exc:
  504. raise ValidationError(f"{label} must contain integers") from exc
  505. return converted
  506. def _bool_value(value: Any) -> bool:
  507. if isinstance(value, (bool, np.bool_)):
  508. return bool(value)
  509. if isinstance(value, str) and value.lower() in {"true", "false"}:
  510. return value.lower() == "true"
  511. raise ValidationError(f"expected Boolean value, got {value!r}")
  512. def _primary_checkpoint_records(primary: gate.ValidatedRun) -> dict[tuple[str, str, int], dict[str, Any]]:
  513. records: dict[tuple[str, str, int], dict[str, Any]] = {}
  514. for record in primary.freeze.get("checkpoints", []):
  515. key = (str(record["model"]), str(record["jammer_mode"]), int(record["train_seed"]))
  516. path = primary.spec.path / str(record["path"])
  517. records[key] = {
  518. "sha256": str(record["sha256"]).lower(),
  519. "bytes": path.stat().st_size,
  520. }
  521. return records
  522. def _validate_cost_profile(
  523. cost_dir: Path, primary: gate.ValidatedRun
  524. ) -> tuple[pd.DataFrame, pd.DataFrame, dict[str, Any], dict[str, str]]:
  525. root = Path(cost_dir).expanduser().resolve()
  526. if not root.is_dir():
  527. raise ValidationError(f"cost profile directory does not exist: {root}")
  528. csv_path = root / "isolated_model_costs.csv"
  529. json_path = root / "isolated_model_costs.json"
  530. frame = _load_csv(csv_path, "isolated model cost CSV")
  531. payload = _load_json(json_path, "isolated model cost JSON")
  532. if not isinstance(payload, Mapping) or not isinstance(payload.get("metadata"), Mapping):
  533. raise ValidationError("isolated cost JSON must contain metadata")
  534. training_records = payload.get("isolated_training")
  535. inference_records = payload.get("canonical_checkpoint_inference")
  536. if not isinstance(training_records, list) or not isinstance(inference_records, list):
  537. raise ValidationError("isolated cost JSON record arrays are missing")
  538. records = [*training_records, *inference_records]
  539. if not all(isinstance(record, Mapping) for record in records):
  540. raise ValidationError("isolated cost JSON contains a malformed record")
  541. _records_match_csv(frame, records)
  542. training = frame[frame["record_type"] == "isolated_training_wall_time"].copy()
  543. inference = frame[frame["record_type"] == "canonical_checkpoint_inference"].copy()
  544. if len(training) != 6 or len(inference) != 40 or len(frame) != 46:
  545. raise ValidationError(
  546. f"isolated cost profile requires 6 training and 40 inference rows; "
  547. f"got {len(training)} and {len(inference)}"
  548. )
  549. if set(frame["record_type"]) != {
  550. "isolated_training_wall_time",
  551. "canonical_checkpoint_inference",
  552. }:
  553. raise ValidationError("isolated cost profile contains an unknown record type")
  554. metadata = dict(payload["metadata"])
  555. if metadata.get("purpose") != "isolated engineering cost comparison; no policy-performance inference":
  556. raise ValidationError("cost metadata purpose does not preserve non-inferential scope")
  557. if _int_set(metadata.get("profiling_seeds", ()), "profiling_seeds") != {91001, 91002, 91003}:
  558. raise ValidationError("cost metadata must contain profiling seeds 91001--91003")
  559. expected_training_order = [
  560. [91001, "hnp"],
  561. [91001, "matched"],
  562. [91002, "matched"],
  563. [91002, "hnp"],
  564. [91003, "hnp"],
  565. [91003, "matched"],
  566. ]
  567. if metadata.get("training_order") != expected_training_order:
  568. raise ValidationError("cost metadata does not contain the predeclared counterbalanced order")
  569. profiling_seed_role = str(metadata.get("profiling_seed_role", ""))
  570. if "not units for policy-performance inference" not in profiling_seed_role:
  571. raise ValidationError("cost metadata does not label profiling seeds as non-inferential")
  572. if str(metadata.get("data_scope", "")) != (
  573. "development 20cm/10dBm train scans 0-6; validation scan 7 loaded only for provenance"
  574. ):
  575. raise ValidationError("cost metadata development-only data scope is invalid")
  576. if "pilot scans 8-9" not in str(metadata.get("excluded_data", "")) or "cross-configuration" not in str(
  577. metadata.get("excluded_data", "")
  578. ):
  579. raise ValidationError("cost metadata does not explicitly exclude pilot and transfer data")
  580. if int(metadata.get("latency_warmup", -1)) != 200 or int(
  581. metadata.get("latency_repetitions", -1)
  582. ) != 1000:
  583. raise ValidationError("cost latency protocol must use 200 warm-ups and 1000 repetitions")
  584. if int(metadata.get("cpu_intraop_threads", -1)) != 1 or int(
  585. metadata.get("cpu_interop_threads", -1)
  586. ) != 1:
  587. raise ValidationError("cost profile must use one intra-op and one inter-op CPU thread")
  588. memory_scope = str(metadata.get("memory_scope", ""))
  589. if "excludes activations" not in memory_scope or "allocator overhead" not in memory_scope:
  590. raise ValidationError("cost metadata memory scope is incomplete")
  591. training_keys = set(
  592. (str(row.model), int(row.profiling_seed))
  593. for row in training[["model", "profiling_seed"]].itertuples(index=False)
  594. )
  595. expected_training = {(model, seed) for model in EXPECTED_MODELS for seed in (91001, 91002, 91003)}
  596. if training_keys != expected_training or training.duplicated(["model", "profiling_seed"]).any():
  597. raise ValidationError("isolated training profile is incomplete or duplicated")
  598. if _int_set(training["profile_order"], "training profile_order") != set(range(6)):
  599. raise ValidationError("isolated training profile order must be exactly 0--5")
  600. ordered_training = training.sort_values("profile_order")
  601. observed_training_order = [
  602. [int(row.profiling_seed), str(row.model)]
  603. for row in ordered_training[["profiling_seed", "model"]].itertuples(index=False)
  604. ]
  605. if observed_training_order != expected_training_order:
  606. raise ValidationError("training rows do not follow the predeclared counterbalanced order")
  607. for row in training.itertuples(index=False):
  608. if str(row.jammer_mode) != "sweeping":
  609. raise ValidationError("isolated training timing must use only sweeping mode")
  610. if int(row.train_episodes) != 200 or int(row.episode_length) != 100 or int(row.environment_steps) != 20000:
  611. raise ValidationError("isolated training timing does not use the exact 200x100 formal budget")
  612. if int(row.optimizer_updates) <= 0:
  613. raise ValidationError("isolated training timing has no optimizer updates")
  614. if int(row.cpu_intraop_threads) != 1 or int(row.cpu_interop_threads) != 1:
  615. raise ValidationError("isolated training row is not single-threaded")
  616. if str(row.training_data_scope) != "20cm/10dBm train scans 0-6 only" or not _bool_value(
  617. row.validation_loaded_not_used_for_training
  618. ):
  619. raise ValidationError("isolated training row has an invalid data scope")
  620. if str(row.profiling_seed_role) != profiling_seed_role:
  621. raise ValidationError("profiling seed was not labelled non-inferential")
  622. if str(row.inference_persistent_tensor_bytes_scope) != memory_scope:
  623. raise ValidationError("training-row persistent-tensor scope differs from metadata")
  624. _finite_nonnegative(
  625. training,
  626. (
  627. "agent_initialization_wall_seconds",
  628. "training_loop_wall_seconds",
  629. "agent_init_plus_training_wall_seconds",
  630. "trainable_parameters",
  631. "parameter_bytes",
  632. "registered_buffer_bytes",
  633. "inference_persistent_tensor_bytes",
  634. ),
  635. "isolated training rows",
  636. )
  637. if not np.allclose(
  638. training["agent_initialization_wall_seconds"].to_numpy(float)
  639. + training["training_loop_wall_seconds"].to_numpy(float),
  640. training["agent_init_plus_training_wall_seconds"].to_numpy(float),
  641. rtol=1e-8,
  642. atol=1e-8,
  643. ):
  644. raise ValidationError("initialization + training wall time does not add up")
  645. checkpoint_records = _primary_checkpoint_records(primary)
  646. inference_keys = set(
  647. (str(row.model), str(row.jammer_mode), int(row.canonical_train_seed))
  648. for row in inference[["model", "jammer_mode", "canonical_train_seed"]].itertuples(index=False)
  649. )
  650. if inference_keys != set(checkpoint_records) or inference.duplicated(
  651. ["model", "jammer_mode", "canonical_train_seed"]
  652. ).any():
  653. raise ValidationError("cost inference rows do not cover all 40 canonical primary checkpoints")
  654. if _int_set(inference["profile_order"], "inference profile_order") != set(range(40)):
  655. raise ValidationError("inference profile order must be exactly 0--39")
  656. cpu_columns = (
  657. "batch1_cpu_latency_ms_median",
  658. "batch1_cpu_latency_ms_mean",
  659. "batch1_cpu_latency_ms_p05",
  660. "batch1_cpu_latency_ms_p95",
  661. )
  662. _finite_nonnegative(
  663. inference,
  664. (*cpu_columns, "trainable_parameters", "parameter_bytes", "registered_buffer_bytes", "inference_persistent_tensor_bytes", "serialized_state_bytes", "checkpoint_bytes"),
  665. "canonical inference rows",
  666. )
  667. cuda_available = _bool_value(metadata.get("cuda_available"))
  668. gpu_columns = (
  669. "batch1_gpu_latency_ms_median",
  670. "batch1_gpu_latency_ms_mean",
  671. "batch1_gpu_latency_ms_p05",
  672. "batch1_gpu_latency_ms_p95",
  673. )
  674. if cuda_available:
  675. _finite_nonnegative(inference, gpu_columns, "canonical GPU inference rows")
  676. elif not inference[list(gpu_columns)].isna().all().all():
  677. raise ValidationError("GPU latency is present although metadata says CUDA was unavailable")
  678. if (inference["batch1_cpu_latency_ms_p05"] > inference["batch1_cpu_latency_ms_median"]).any() or (
  679. inference["batch1_cpu_latency_ms_median"] > inference["batch1_cpu_latency_ms_p95"]
  680. ).any():
  681. raise ValidationError("CPU latency quantiles are reversed")
  682. if cuda_available and (
  683. (inference["batch1_gpu_latency_ms_p05"] > inference["batch1_gpu_latency_ms_median"]).any()
  684. or (inference["batch1_gpu_latency_ms_median"] > inference["batch1_gpu_latency_ms_p95"]).any()
  685. ):
  686. raise ValidationError("GPU latency quantiles are reversed")
  687. for row in inference.itertuples(index=False):
  688. key = (str(row.model), str(row.jammer_mode), int(row.canonical_train_seed))
  689. expected = checkpoint_records[key]
  690. if str(row.checkpoint_sha256).lower() != expected["sha256"] or int(row.checkpoint_bytes) != expected["bytes"]:
  691. raise ValidationError(f"cost checkpoint provenance mismatch for {key}")
  692. if int(row.latency_warmup) != 200 or int(row.latency_repetitions) != 1000:
  693. raise ValidationError(f"cost latency schedule mismatch for {key}")
  694. if str(row.profiling_seed_role) != profiling_seed_role:
  695. raise ValidationError(f"cost inference row has an invalid profiling-seed role for {key}")
  696. if str(row.latency_input) != "batch-1 all-zero normalized 48D observation":
  697. raise ValidationError(f"cost latency input mismatch for {key}")
  698. expected_synchronization = (
  699. "per-forward torch.cuda.synchronize" if cuda_available else "CPU call return"
  700. )
  701. if str(row.latency_synchronization) != expected_synchronization:
  702. raise ValidationError(f"cost latency synchronization mismatch for {key}")
  703. if str(row.inference_persistent_tensor_bytes_scope) != memory_scope:
  704. raise ValidationError(f"cost memory scope mismatch for {key}")
  705. if str(row.serialized_state_bytes_role) != "storage size, not memory" or str(
  706. row.checkpoint_bytes_role
  707. ) != "storage size, not memory":
  708. raise ValidationError(f"storage bytes are mislabelled for {key}")
  709. if int(row.parameter_bytes) + int(row.registered_buffer_bytes) != int(
  710. row.inference_persistent_tensor_bytes
  711. ):
  712. raise ValidationError(f"persistent tensor byte accounting mismatch for {key}")
  713. for model in EXPECTED_MODELS:
  714. model_rows = inference[inference["model"] == model]
  715. for column in (
  716. "trainable_parameters",
  717. "parameter_bytes",
  718. "registered_buffer_bytes",
  719. "inference_persistent_tensor_bytes",
  720. "serialized_state_bytes",
  721. ):
  722. if model_rows[column].nunique(dropna=False) != 1:
  723. raise ValidationError(f"{model} {column} varies across canonical checkpoints")
  724. training_model = training[training["model"] == model]
  725. if set(training_model["trainable_parameters"].astype(int)) != set(
  726. model_rows["trainable_parameters"].astype(int)
  727. ):
  728. raise ValidationError(f"{model} parameter count differs between training and inference profiles")
  729. hnp_parameters = int(inference[inference["model"] == "hnp"].iloc[0]["trainable_parameters"])
  730. matched_parameters = int(inference[inference["model"] == "matched"].iloc[0]["trainable_parameters"])
  731. if abs(hnp_parameters - matched_parameters) / hnp_parameters > 0.01:
  732. raise ValidationError("the named capacity-matched MLP differs from HNP by more than 1% in parameters")
  733. return training, inference, metadata, {
  734. "isolated_model_costs.csv": _sha256(csv_path),
  735. "isolated_model_costs.json": _sha256(json_path),
  736. }
  737. def _classification(row: Mapping[str, Any]) -> str:
  738. low = float(row["ci95_low"])
  739. high = float(row["ci95_high"])
  740. p_holm = float(row["p_holm"])
  741. if low > 0.0 and p_holm < 0.05:
  742. return "supported_hnp_advantage"
  743. if high < 0.0 and p_holm < 0.05:
  744. return "supported_hnp_disadvantage"
  745. return "inconclusive"
  746. def _number(value: float, digits: int = 2) -> str:
  747. value = float(value)
  748. if math.isinf(value):
  749. return r"+\infty" if value > 0 else r"-\infty"
  750. if math.isnan(value):
  751. return "NA"
  752. rounded = 0.0 if abs(value) < 0.5 * (10 ** -digits) else value
  753. return f"{rounded:.{digits}f}"
  754. def _p_tex(value: float) -> str:
  755. value = float(value)
  756. if value < 0.001:
  757. return r"$p_{\mathrm{Holm}}<0.001$"
  758. return rf"$p_{{\mathrm{{Holm}}}}={value:.3f}$"
  759. def _comparison_record(row: Mapping[str, Any], family: str) -> dict[str, Any]:
  760. classification = _classification(row)
  761. condition = str(row["condition"])
  762. jammer = str(row["jammer_mode"])
  763. comparator = str(row["method_b"])
  764. difference_tex = (
  765. rf"${_number(row['mean_difference'])}$ "
  766. rf"[95\% CI ${_number(row['ci95_low'])}$, ${_number(row['ci95_high'])}$]"
  767. )
  768. effect_tex = rf"Cohen's $d_z={_number(row['effect_dz'])}$"
  769. inference_tex = f"{_p_tex(row['p_holm'])}; {classification.replace('_', ' ')}"
  770. direction = {
  771. "supported_hnp_advantage": "supported an HNP-DQN advantage",
  772. "supported_hnp_disadvantage": f"supported an advantage for {METHOD_LABELS[comparator]}",
  773. "inconclusive": "was inconclusive (not evidence of equivalence)",
  774. }[classification]
  775. sentence = (
  776. f"For the {CONDITION_LABELS[condition]} under {jammer} jamming, the paired "
  777. f"HNP-DQN minus {METHOD_LABELS[comparator]} return difference was "
  778. f"{_number(row['mean_difference'])} (95% CI {_number(row['ci95_low'])} to "
  779. f"{_number(row['ci95_high'])}; Cohen's dz={_number(row['effect_dz'])}; "
  780. f"Holm-adjusted p={float(row['p_holm']):.3f}), which {direction}."
  781. )
  782. # Keep non-finite effect sizes available to the renderer as +/- infinity;
  783. # the final JSON serializer converts them to null while the adjacent
  784. # ``effect_tex`` retains the explicit human-readable value.
  785. raw = dict(row)
  786. return {
  787. "family": family,
  788. "condition": condition,
  789. "condition_label": CONDITION_LABELS[condition],
  790. "jammer_mode": jammer,
  791. "method_b": comparator,
  792. "method_b_label": METHOD_LABELS[comparator],
  793. "classification": classification,
  794. "raw": raw,
  795. "difference_ci_tex": difference_tex,
  796. "effect_tex": effect_tex,
  797. "inference_tex": inference_tex,
  798. "english_sentence": sentence,
  799. }
  800. def _summarise_classifications(records: Sequence[Mapping[str, Any]]) -> dict[str, int]:
  801. return {
  802. label: sum(record["classification"] == label for record in records)
  803. for label in (
  804. "supported_hnp_advantage",
  805. "supported_hnp_disadvantage",
  806. "inconclusive",
  807. )
  808. }
  809. def _descriptive_performance(run: gate.ValidatedRun) -> list[dict[str, Any]]:
  810. summary = run.seed_summary
  811. evaluation = run.evaluation
  812. output: list[dict[str, Any]] = []
  813. for condition in EXPECTED_CONDITIONS:
  814. for jammer in EXPECTED_JAMMERS:
  815. setting = summary[
  816. (summary["condition"] == condition) & (summary["jammer_mode"] == jammer)
  817. ]
  818. expected_methods = set(EXPECTED_MODELS) | set(EXPECTED_BASELINES[jammer])
  819. if set(setting["method"]) != expected_methods:
  820. raise ValidationError(
  821. f"{condition}/{jammer} method set differs; "
  822. f"missing={sorted(expected_methods-set(setting['method']))}, "
  823. f"extra={sorted(set(setting['method'])-expected_methods)}"
  824. )
  825. methods: list[dict[str, Any]] = []
  826. for method in (*EXPECTED_MODELS, *EXPECTED_BASELINES[jammer]):
  827. rows = setting[setting["method"] == method]
  828. if method in EXPECTED_MODELS:
  829. learned = _learned_rows(summary, method, condition, jammer)
  830. record = {
  831. "method": method,
  832. "method_label": METHOD_LABELS[method],
  833. "role": "learned",
  834. "n_training_seeds": 10,
  835. }
  836. for metric in ("return", "collision_rate", "switch_rate"):
  837. values = learned[metric].to_numpy(dtype=float)
  838. record[metric] = {
  839. "mean": float(values.mean()),
  840. "sd_across_training_seeds": float(values.std(ddof=1)),
  841. }
  842. else:
  843. baseline = rows[rows["train_seed"].isna()]
  844. if len(baseline) != 1:
  845. raise ValidationError(
  846. f"{condition}/{jammer}/{method} needs one descriptive baseline row"
  847. )
  848. episode_rows = evaluation[
  849. (evaluation["condition"] == condition)
  850. & (evaluation["jammer_mode"] == jammer)
  851. & (evaluation["method"] == method)
  852. & evaluation["train_seed"].isna()
  853. ]
  854. if len(episode_rows) != 20:
  855. raise ValidationError(
  856. f"{condition}/{jammer}/{method} needs 20 fixed trajectories"
  857. )
  858. record = {
  859. "method": method,
  860. "method_label": METHOD_LABELS[method],
  861. "role": (
  862. "privileged"
  863. if method in {"jammer_greedy", "clairvoyant_oracle"}
  864. else "ordinary_or_schedule_aware"
  865. ),
  866. "n_fixed_trajectories": 20,
  867. }
  868. for metric in ("return", "collision_rate", "switch_rate"):
  869. values = episode_rows[metric].to_numpy(dtype=float)
  870. record[metric] = {
  871. "mean": float(values.mean()),
  872. "sd_across_fixed_trajectories_descriptive_only": float(values.std(ddof=1)),
  873. }
  874. methods.append(record)
  875. return_tex = "; ".join(
  876. f"{item['method_label']}: ${_number(item['return']['mean'])}"
  877. + (
  878. rf" \pm {_number(item['return']['sd_across_training_seeds'])}$"
  879. if item["role"] == "learned"
  880. else "$ (fixed-trajectory mean)"
  881. )
  882. for item in methods
  883. )
  884. rates_tex = "; ".join(
  885. f"{item['method_label']}: collision ${_number(100*item['collision_rate']['mean'])}\\%$, "
  886. f"switch ${_number(100*item['switch_rate']['mean'])}\\%$"
  887. for item in methods
  888. )
  889. output.append(
  890. {
  891. "condition": condition,
  892. "condition_label": CONDITION_LABELS[condition],
  893. "jammer_mode": jammer,
  894. "methods": methods,
  895. "returns_tex": return_tex,
  896. "collision_switch_tex": rates_tex,
  897. "uncertainty_note": (
  898. "Learned entries are mean +/- SD across 10 training seeds; "
  899. "fixed references are 20-trajectory point summaries and are not inferential replicates."
  900. ),
  901. }
  902. )
  903. return output
  904. def _strongest_ordinary(performance: Sequence[Mapping[str, Any]]) -> list[dict[str, Any]]:
  905. output: list[dict[str, Any]] = []
  906. for setting in performance:
  907. allowed = set(ORDINARY_BASELINES[setting["jammer_mode"]])
  908. candidates = [item for item in setting["methods"] if item["method"] in allowed]
  909. best_value = max(float(item["return"]["mean"]) for item in candidates)
  910. winners = [item for item in candidates if math.isclose(float(item["return"]["mean"]), best_value)]
  911. hnp = next(item for item in setting["methods"] if item["method"] == "hnp")
  912. predeclared = "schedule_sweep" if setting["jammer_mode"] == "sweeping" else "threshold"
  913. winner = winners[0]
  914. output.append(
  915. {
  916. "condition": setting["condition"],
  917. "jammer_mode": setting["jammer_mode"],
  918. "posthoc_best_observed_methods": [item["method"] for item in winners],
  919. "predeclared_simple_comparator": predeclared,
  920. "posthoc_winner_matches_predeclared": predeclared in {item["method"] for item in winners},
  921. "hnp_minus_best_observed_return": float(hnp["return"]["mean"] - best_value),
  922. "hnp_minus_best_observed_collision_percentage_points": float(
  923. 100 * (hnp["collision_rate"]["mean"] - winner["collision_rate"]["mean"])
  924. ),
  925. "hnp_minus_best_observed_switch_percentage_points": float(
  926. 100 * (hnp["switch_rate"]["mean"] - winner["switch_rate"]["mean"])
  927. ),
  928. "warning": (
  929. "This is a descriptive post-hoc ranking. It must not inherit the adjusted p-value "
  930. "of the predeclared schedule-aware/threshold comparison."
  931. ),
  932. }
  933. )
  934. return output
  935. def _ablation_supporting_endpoints(
  936. runs: Mapping[str, gate.ValidatedRun]
  937. ) -> list[dict[str, Any]]:
  938. primary = gate._learned_seed_rows(runs["primary"], "hnp")
  939. variant_specs = (
  940. ("gamma0", "hnp", "hnp_gamma0"),
  941. ("no_polynomial", "no_polynomial", "no_polynomial"),
  942. ("no_layernorm", "no_layernorm", "no_layernorm"),
  943. ("no_dueling", "no_dueling", "no_dueling"),
  944. )
  945. rows: list[dict[str, Any]] = []
  946. keys = ["condition", "jammer_mode", "train_seed"]
  947. for run_label, method, variant_label in variant_specs:
  948. variant = gate._learned_seed_rows(runs[run_label], method)
  949. joined = primary[keys + ["collision_rate", "switch_rate"]].merge(
  950. variant[keys + ["collision_rate", "switch_rate"]],
  951. on=keys,
  952. how="outer",
  953. validate="one_to_one",
  954. indicator=True,
  955. suffixes=("_hnp", "_variant"),
  956. )
  957. if set(joined["_merge"]) != {"both"}:
  958. raise ValidationError(f"supporting endpoint seed mismatch for {variant_label}")
  959. for (condition, jammer), group in joined.groupby(["condition", "jammer_mode"], sort=False):
  960. if tuple(sorted(group["train_seed"].astype(int))) != EXPECTED_TRAIN_SEEDS:
  961. raise ValidationError(f"supporting endpoint schedule incomplete for {variant_label}")
  962. rows.append(
  963. {
  964. "condition": str(condition),
  965. "jammer_mode": str(jammer),
  966. "method_b": variant_label,
  967. "mean_hnp_minus_variant_collision_percentage_points": float(
  968. 100 * (group["collision_rate_hnp"] - group["collision_rate_variant"]).mean()
  969. ),
  970. "mean_hnp_minus_variant_switch_percentage_points": float(
  971. 100 * (group["switch_rate_hnp"] - group["switch_rate_variant"]).mean()
  972. ),
  973. "scope": "descriptive paired-seed supporting endpoints; not added to the return Holm family",
  974. }
  975. )
  976. return rows
  977. def _distribution(values: Sequence[float]) -> dict[str, float | int]:
  978. data = np.asarray(values, dtype=float)
  979. if data.size == 0 or not np.isfinite(data).all():
  980. raise ValidationError("cost summary requires finite non-empty values")
  981. return {
  982. "n": int(data.size),
  983. "median": float(np.median(data)),
  984. "q1": float(np.quantile(data, 0.25)),
  985. "q3": float(np.quantile(data, 0.75)),
  986. "min": float(data.min()),
  987. "max": float(data.max()),
  988. }
  989. def _median_iqr_tex(summary: Mapping[str, Any], unit: str) -> str:
  990. return (
  991. rf"${_number(summary['median'], 3)}$ [{_number(summary['q1'], 3)}, "
  992. rf"{_number(summary['q3'], 3)}] {unit}"
  993. )
  994. def _cost_summary(
  995. training: pd.DataFrame, inference: pd.DataFrame, metadata: Mapping[str, Any]
  996. ) -> dict[str, Any]:
  997. models: dict[str, Any] = {}
  998. for model in EXPECTED_MODELS:
  999. train = training[training["model"] == model]
  1000. infer = inference[inference["model"] == model]
  1001. first = infer.iloc[0]
  1002. gpu_values = infer["batch1_gpu_latency_ms_median"].dropna().to_numpy(float)
  1003. model_record = {
  1004. "trainable_parameters": int(first["trainable_parameters"]),
  1005. "parameter_bytes": int(first["parameter_bytes"]),
  1006. "registered_buffer_bytes": int(first["registered_buffer_bytes"]),
  1007. "inference_persistent_tensor_bytes": int(first["inference_persistent_tensor_bytes"]),
  1008. "inference_persistent_tensor_bytes_scope": str(first["inference_persistent_tensor_bytes_scope"]),
  1009. "serialized_state_bytes": _distribution(infer["serialized_state_bytes"]),
  1010. "checkpoint_bytes": _distribution(infer["checkpoint_bytes"]),
  1011. "batch1_cpu_latency_ms_median_across_checkpoints": _distribution(
  1012. infer["batch1_cpu_latency_ms_median"]
  1013. ),
  1014. "batch1_gpu_latency_ms_median_across_checkpoints": (
  1015. _distribution(gpu_values) if gpu_values.size else None
  1016. ),
  1017. "training_loop_wall_seconds_across_engineering_seeds": _distribution(
  1018. train["training_loop_wall_seconds"]
  1019. ),
  1020. }
  1021. persistent_kib = model_record["inference_persistent_tensor_bytes"] / 1024
  1022. state = model_record["serialized_state_bytes"]
  1023. checkpoint = model_record["checkpoint_bytes"]
  1024. cpu = model_record["batch1_cpu_latency_ms_median_across_checkpoints"]
  1025. gpu = model_record["batch1_gpu_latency_ms_median_across_checkpoints"]
  1026. wall = model_record["training_loop_wall_seconds_across_engineering_seeds"]
  1027. model_record["capacity_tex"] = (
  1028. f"{model_record['trainable_parameters']:,} trainable parameters; "
  1029. rf"${_number(persistent_kib)}$ KiB persistent parameter/buffer tensors"
  1030. )
  1031. model_record["storage_tex"] = (
  1032. rf"state ${_number(state['median']/1024)}$ [{_number(state['q1']/1024)}, "
  1033. rf"{_number(state['q3']/1024)}] KiB; checkpoint "
  1034. rf"${_number(checkpoint['median']/1024)}$ [{_number(checkpoint['q1']/1024)}, "
  1035. rf"{_number(checkpoint['q3']/1024)}] KiB"
  1036. )
  1037. latency_parts = [f"CPU {_median_iqr_tex(cpu, 'ms')}"]
  1038. latency_parts.append(
  1039. f"GPU {_median_iqr_tex(gpu, 'ms')}" if gpu else "GPU not available on the profiling system"
  1040. )
  1041. latency_parts.append(f"training {_median_iqr_tex(wall, 's')}")
  1042. model_record["latency_training_tex"] = "; ".join(latency_parts)
  1043. models[model] = model_record
  1044. hnp_n = models["hnp"]["trainable_parameters"]
  1045. matched_n = models["matched"]["trainable_parameters"]
  1046. gap = matched_n - hnp_n
  1047. relative = 100 * gap / hnp_n
  1048. models["matched"]["capacity_gap_from_hnp"] = {
  1049. "absolute_parameters": gap,
  1050. "percent_of_hnp": relative,
  1051. "tex": f"{gap:+,} parameters ({relative:+.2f}\\% relative to HNP-DQN)",
  1052. }
  1053. return {
  1054. "models": models,
  1055. "protocol": {
  1056. "training_seed_role": metadata["profiling_seed_role"],
  1057. "training_timing_n_per_model": 3,
  1058. "inference_checkpoint_n_per_model": 20,
  1059. "latency_warmup": 200,
  1060. "latency_repetitions": 1000,
  1061. "cpu_threads": 1,
  1062. "cuda_available": _bool_value(metadata["cuda_available"]),
  1063. "memory_scope": metadata["memory_scope"],
  1064. },
  1065. }
  1066. def _stability_summary(run: gate.ValidatedRun) -> dict[str, Any]:
  1067. records: list[dict[str, Any]] = []
  1068. for method in EXPECTED_MODELS:
  1069. for condition in EXPECTED_CONDITIONS:
  1070. for jammer in EXPECTED_JAMMERS:
  1071. rows = _learned_rows(run.seed_summary, method, condition, jammer)
  1072. records.append(
  1073. {
  1074. "method": method,
  1075. "condition": condition,
  1076. "jammer_mode": jammer,
  1077. "return_sd_across_training_seeds": float(rows["return"].std(ddof=1)),
  1078. }
  1079. )
  1080. ranges = {}
  1081. for method in EXPECTED_MODELS:
  1082. values = [item["return_sd_across_training_seeds"] for item in records if item["method"] == method]
  1083. ranges[method] = {"min": min(values), "max": max(values)}
  1084. sentence = (
  1085. f"All 10 planned canonical training seeds were present for every learned-model and jammer-mode "
  1086. f"combination. Across the six held-out settings, seed-level return SD ranged from "
  1087. f"{_number(ranges['hnp']['min'])} to {_number(ranges['hnp']['max'])} for HNP-DQN and from "
  1088. f"{_number(ranges['matched']['min'])} to {_number(ranges['matched']['max'])} for the matched MLP. "
  1089. "The frozen artifacts establish completeness of the canonical runs but do not enumerate failed "
  1090. "exploratory or pre-canonical attempts; no convergence/AUC endpoint or cross-model TD-loss "
  1091. "stability claim is made."
  1092. )
  1093. return {"setting_records": records, "ranges": ranges, "english_sentence": sentence}
  1094. def _placeholder_inventory(path: Path, pattern: re.Pattern[str], label: str) -> list[dict[str, Any]]:
  1095. if not path.is_file():
  1096. raise ValidationError(f"{label} file is missing: {path}")
  1097. text = path.read_text(encoding="utf-8")
  1098. matches = list(pattern.finditer(text))
  1099. return [
  1100. {
  1101. "index": index,
  1102. "line": text.count("\n", 0, match.start()) + 1,
  1103. "raw": match.group(0),
  1104. "label": " ".join(match.group(1).split()),
  1105. }
  1106. for index, match in enumerate(matches, start=1)
  1107. ]
  1108. def _placeholder_map(
  1109. inventory: Sequence[Mapping[str, Any]],
  1110. performance: Sequence[Mapping[str, Any]],
  1111. primary_records: Sequence[Mapping[str, Any]],
  1112. ablation_records: Sequence[Mapping[str, Any]],
  1113. costs: Mapping[str, Any],
  1114. stability: Mapping[str, Any],
  1115. primary_summary_sentence: str,
  1116. ablation_summary_sentence: str,
  1117. ) -> list[dict[str, Any]]:
  1118. if len(inventory) != 99:
  1119. raise ValidationError(
  1120. f"current manuscript mapping requires exactly 99 VERIFIED RESULT markers; got {len(inventory)}"
  1121. )
  1122. mapping: list[dict[str, Any]] = []
  1123. auto: dict[int, tuple[str, str | None]] = {
  1124. 1: ("primary/composite", None),
  1125. 2: ("primary matched rows + cost", None),
  1126. 3: ("all validated outcomes", None),
  1127. 4: ("validated protocol", "10 independently trained seeds and 20 fixed evaluation trajectories per condition and jammer mode"),
  1128. 5: ("validated protocol", "10"),
  1129. 6: (
  1130. "validated protocol",
  1131. "10 training seeds; 20 fixed trajectories per setting; two-sided 95\\% Student-$t$ interval for paired seed differences; Cohen's $d_z$; two-sided exact sign-flip enumeration; one primary Holm family over 12 rows",
  1132. ),
  1133. 7: ("manual figure insertion", None),
  1134. 20: ("predeclared and descriptive comparator audit", None),
  1135. 57: ("all 12 primary rows", primary_summary_sentence),
  1136. 58: ("validated seed dispersion", stability["english_sentence"]),
  1137. 65: ("matched performance + cost", None),
  1138. 96: ("24 exploratory rows + supporting endpoints", ablation_summary_sentence),
  1139. 97: ("separate checkpoint weight-audit artifact", None),
  1140. 98: ("all primary rows", primary_summary_sentence),
  1141. 99: ("all ablation rows", ablation_summary_sentence),
  1142. }
  1143. # Table 4: two placeholders for each setting in manuscript order.
  1144. for offset, setting in enumerate(performance):
  1145. auto[8 + 2 * offset] = ("seed_summary/evaluation_episodes", setting["returns_tex"])
  1146. auto[9 + 2 * offset] = ("seed_summary/evaluation_episodes", setting["collision_switch_tex"])
  1147. # Table 5: difference, effect, inference for 12 rows.
  1148. for row_index, record in enumerate(primary_records):
  1149. start = 21 + 3 * row_index
  1150. auto[start] = ("primary_comparisons", record["difference_ci_tex"])
  1151. auto[start + 1] = ("primary_comparisons", record["effect_tex"])
  1152. auto[start + 2] = ("primary_comparisons", record["inference_tex"])
  1153. hnp = costs["models"]["hnp"]
  1154. matched = costs["models"]["matched"]
  1155. auto[59] = ("isolated cost profile", hnp["capacity_tex"])
  1156. auto[60] = ("isolated cost profile", hnp["storage_tex"])
  1157. auto[61] = ("isolated cost profile", hnp["latency_training_tex"])
  1158. auto[62] = (
  1159. "isolated cost profile",
  1160. matched["capacity_tex"] + "; " + matched["capacity_gap_from_hnp"]["tex"],
  1161. )
  1162. auto[63] = ("isolated cost profile", matched["storage_tex"])
  1163. auto[64] = ("isolated cost profile", matched["latency_training_tex"])
  1164. # Table 6 order: variant, jammer, then the three conditions. Matched rows
  1165. # come from the primary family; all other variants use the exploratory family.
  1166. primary_index = {
  1167. (item["method_b"], item["jammer_mode"], item["condition"]): item
  1168. for item in primary_records
  1169. }
  1170. ablation_index = {
  1171. (item["method_b"], item["jammer_mode"], item["condition"]): item
  1172. for item in ablation_records
  1173. }
  1174. table6_variants = ("no_polynomial", "no_layernorm", "no_dueling", "matched", "hnp_gamma0")
  1175. position = 66
  1176. for variant in table6_variants:
  1177. for jammer in EXPECTED_JAMMERS:
  1178. for condition in EXPECTED_CONDITIONS:
  1179. index = primary_index if variant == "matched" else ablation_index
  1180. record = index[(variant, jammer, condition)]
  1181. family = "primary" if variant == "matched" else "exploratory ablation"
  1182. auto[position] = (family, record["difference_ci_tex"])
  1183. position += 1
  1184. for item in inventory:
  1185. source, suggestion = auto[item["index"]]
  1186. mapping.append(
  1187. {
  1188. **dict(item),
  1189. "source": source,
  1190. "status": "auto_formatted" if suggestion is not None else "human_review_required",
  1191. "suggested_tex": suggestion,
  1192. }
  1193. )
  1194. return mapping
  1195. def _records_sentence(records: Sequence[Mapping[str, Any]]) -> str:
  1196. return " ".join(str(record["english_sentence"]) for record in records)
  1197. def _summary_sentence(records: Sequence[Mapping[str, Any]], label: str) -> str:
  1198. counts = _summarise_classifications(records)
  1199. return (
  1200. f"Across the {len(records)} {label}, {counts['supported_hnp_advantage']} supported an "
  1201. f"HNP-DQN advantage, {counts['supported_hnp_disadvantage']} supported the comparator/variant, "
  1202. f"and {counts['inconclusive']} were inconclusive under the joint CI-plus-Holm rule. "
  1203. "This count is a mechanical audit, not a practical-significance judgment; every setting-specific "
  1204. "estimate and null or adverse result must remain visible."
  1205. )
  1206. def _response_clauses(
  1207. primary_records: Sequence[Mapping[str, Any]],
  1208. ablation_records: Sequence[Mapping[str, Any]],
  1209. costs: Mapping[str, Any],
  1210. stability: Mapping[str, Any],
  1211. ) -> dict[str, str]:
  1212. simple = [item for item in primary_records if item["method_b"] != "matched"]
  1213. matched = [item for item in primary_records if item["method_b"] == "matched"]
  1214. gamma0 = [item for item in ablation_records if item["method_b"] == "hnp_gamma0"]
  1215. primary_summary = _summary_sentence(primary_records, "predeclared primary comparisons")
  1216. ablation_summary = _summary_sentence(ablation_records, "exploratory ablation comparisons")
  1217. simple_summary = _summary_sentence(simple, "predeclared simple-comparator comparisons")
  1218. matched_summary = _summary_sentence(matched, "capacity-matched comparisons")
  1219. gamma_summary = _summary_sentence(gamma0, "HNP-minus-gamma-zero comparisons")
  1220. hnp = costs["models"]["hnp"]
  1221. mlp = costs["models"]["matched"]
  1222. cost_sentence = (
  1223. f"HNP-DQN used {hnp['trainable_parameters']:,} trainable parameters and the matched MLP used "
  1224. f"{mlp['trainable_parameters']:,} ({mlp['capacity_gap_from_hnp']['percent_of_hnp']:+.2f}% "
  1225. f"relative to HNP). HNP cost: {hnp['capacity_tex']}; {hnp['storage_tex']}; "
  1226. f"{hnp['latency_training_tex']}. Matched-MLP cost: {mlp['capacity_tex']}; "
  1227. f"{mlp['storage_tex']}; {mlp['latency_training_tex']}. Timing summaries are engineering "
  1228. "medians [IQR], not policy-performance inference, and persistent tensors are not peak memory."
  1229. )
  1230. return {
  1231. "one_paragraph_overall_outcome_with_uncertainty": (
  1232. primary_summary + " " + cost_sentence
  1233. ),
  1234. "myopic_vs_hnp_result_with_ci": _records_sentence(gamma0),
  1235. "sequentiality_claim_consistent_with_outcome": gamma_summary,
  1236. "hnp_vs_strongest_heuristic_effect_ci_adjusted_p": _records_sentence(simple),
  1237. "heuristic_comparison_claim_consistent_with_outcome": simple_summary,
  1238. "number_of_training_seeds": "10 independently trained seeds per learned method and jammer mode",
  1239. "number_of_fixed_evaluation_trajectories": (
  1240. "20 fixed trajectories per physical condition and jammer mode (10 replicates per pilot scan; "
  1241. "two replicates per transfer scan)"
  1242. ),
  1243. "exact_test_details": (
  1244. "paired training-seed differences, two-sided 95% Student-t confidence intervals, Cohen's dz, "
  1245. "two-sided exact sign-flip enumeration, and Holm correction over the 12 predeclared primary rows"
  1246. ),
  1247. "primary_paired_statistics_holm": _records_sentence(primary_records),
  1248. "ablation_paired_statistics_holm": (
  1249. _records_sentence(ablation_records)
  1250. + " These 24 return comparisons form one separate exploratory Holm family; the matched MLP remains in the primary family."
  1251. ),
  1252. "model_cost_summary": cost_sentence,
  1253. "hnp_vs_matched_mlp_effect_ci_adjusted_p": _records_sentence(matched),
  1254. "capacity_control_claim_consistent_with_outcome": matched_summary,
  1255. "directly_comparable_stability_metrics_and_results": stability["english_sentence"],
  1256. "language_and_caption_audit": (
  1257. "MANUAL: complete only after inserting all values and inspecting every claim, table, caption, and compiled page."
  1258. ),
  1259. "r3_statistics_ablation_cost_summary": primary_summary + " " + ablation_summary + " " + cost_sentence,
  1260. }
  1261. def _render_response_markdown(
  1262. payload: Mapping[str, Any], clauses: Mapping[str, str]
  1263. ) -> str:
  1264. lines = [
  1265. "# Validated revision-result fill sheet",
  1266. "",
  1267. "> Generated only after strict run, schedule, checkpoint, seed, comparison, manifest, and cost-profile validation.",
  1268. "> Mechanical support labels require both a directionally excluding 95% CI and Holm-adjusted p<0.05.",
  1269. "> Inconclusive does not mean equivalent. Practical significance and final narrative emphasis require human review.",
  1270. "",
  1271. "## Validation receipt",
  1272. "",
  1273. f"- Generated UTC: `{payload['generated_utc']}`",
  1274. f"- Primary run: `{payload['provenance']['primary_dir']}`",
  1275. f"- Ablation artifacts: `{payload['provenance']['ablation_dir']}`",
  1276. f"- Cost profile: `{payload['provenance']['cost_profile_dir']}`",
  1277. "- Inferential unit: one independently trained seed; 10 paired seeds per comparison.",
  1278. "- Fixed evaluation cases: 20 scan-stratified trajectories per condition and jammer mode.",
  1279. "- Primary multiplicity family: 12 predeclared return comparisons.",
  1280. "- Exploratory multiplicity family: 24 HNP-minus-variant return comparisons, kept separate.",
  1281. "",
  1282. "## Response-placeholder clauses",
  1283. "",
  1284. ]
  1285. for name, text in clauses.items():
  1286. lines.extend([f"### `{name}`", "", text, ""])
  1287. lines.extend(
  1288. [
  1289. "## Primary comparison audit (all 12 rows)",
  1290. "",
  1291. "| Condition | Jammer | Comparator | HNP-minus-comparator difference (95% CI) | dz | Holm p | Mechanical label |",
  1292. "|---|---|---|---:|---:|---:|---|",
  1293. ]
  1294. )
  1295. for item in payload["primary_comparisons"]:
  1296. raw = item["raw"]
  1297. lines.append(
  1298. f"| {item['condition_label']} | {item['jammer_mode']} | {item['method_b_label']} | "
  1299. f"{_number(raw['mean_difference'])} [{_number(raw['ci95_low'])}, {_number(raw['ci95_high'])}] | "
  1300. f"{_number(raw['effect_dz'])} | {float(raw['p_holm']):.3f} | {item['classification']} |"
  1301. )
  1302. lines.extend(
  1303. [
  1304. "",
  1305. "## Exploratory ablation audit (all 24 rows)",
  1306. "",
  1307. "| Condition | Jammer | Variant | HNP-minus-variant difference (95% CI) | dz | Holm p | Mechanical label |",
  1308. "|---|---|---|---:|---:|---:|---|",
  1309. ]
  1310. )
  1311. for item in payload["ablation_comparisons"]:
  1312. raw = item["raw"]
  1313. lines.append(
  1314. f"| {item['condition_label']} | {item['jammer_mode']} | {item['method_b_label']} | "
  1315. f"{_number(raw['mean_difference'])} [{_number(raw['ci95_low'])}, {_number(raw['ci95_high'])}] | "
  1316. f"{_number(raw['effect_dz'])} | {float(raw['p_holm']):.3f} | {item['classification']} |"
  1317. )
  1318. lines.extend(["", "## Required human decisions", ""])
  1319. for item in payload["human_review_required"]:
  1320. lines.append(f"- {item}")
  1321. lines.append("")
  1322. return "\n".join(lines)
  1323. def summarize_revision_results(
  1324. *,
  1325. primary_dir: Path,
  1326. ablation_dir: Path,
  1327. cost_profile_dir: Path,
  1328. output_dir: Path,
  1329. manuscript_template: Path,
  1330. response_draft: Path,
  1331. overwrite: bool = False,
  1332. ) -> dict[str, Any]:
  1333. """Validate all inputs, then atomically prepare the two fill-sheet files."""
  1334. target = Path(output_dir).expanduser().resolve()
  1335. manuscript_output = target / "results_for_manuscript.json"
  1336. response_output = target / "results_for_response.md"
  1337. if not overwrite and (manuscript_output.exists() or response_output.exists()):
  1338. raise ValidationError(
  1339. f"refusing to overwrite an existing fill sheet in {target}; use --overwrite"
  1340. )
  1341. # All reads, recomputations, and prose construction happen before mkdir/write.
  1342. primary, primary_frame = _validate_primary_run(Path(primary_dir))
  1343. ablation_frame, runs, ablation_manifest = _validate_ablation_directory(
  1344. Path(ablation_dir), primary
  1345. )
  1346. training_costs, inference_costs, cost_metadata, cost_hashes = _validate_cost_profile(
  1347. Path(cost_profile_dir), primary
  1348. )
  1349. manuscript_inventory = _placeholder_inventory(
  1350. Path(manuscript_template).expanduser().resolve(), MANUSCRIPT_PATTERN, "manuscript template"
  1351. )
  1352. response_inventory = _placeholder_inventory(
  1353. Path(response_draft).expanduser().resolve(), RESPONSE_PATTERN, "response draft"
  1354. )
  1355. if len(response_inventory) != 17:
  1356. raise ValidationError(
  1357. f"current response mapping requires exactly 17 VERIFIED_RESULT markers; got {len(response_inventory)}"
  1358. )
  1359. response_names = {item["label"] for item in response_inventory}
  1360. expected_response_names = {
  1361. "...",
  1362. "one_paragraph_overall_outcome_with_uncertainty",
  1363. "myopic_vs_hnp_result_with_ci",
  1364. "sequentiality_claim_consistent_with_outcome",
  1365. "hnp_vs_strongest_heuristic_effect_ci_adjusted_p",
  1366. "heuristic_comparison_claim_consistent_with_outcome",
  1367. "number_of_training_seeds",
  1368. "number_of_fixed_evaluation_trajectories",
  1369. "exact_test_details",
  1370. "primary_paired_statistics_holm",
  1371. "ablation_paired_statistics_holm",
  1372. "model_cost_summary",
  1373. "hnp_vs_matched_mlp_effect_ci_adjusted_p",
  1374. "capacity_control_claim_consistent_with_outcome",
  1375. "directly_comparable_stability_metrics_and_results",
  1376. "language_and_caption_audit",
  1377. "r3_statistics_ablation_cost_summary",
  1378. }
  1379. if response_names != expected_response_names:
  1380. raise ValidationError(
  1381. f"response VERIFIED_RESULT names differ; missing={sorted(expected_response_names-response_names)}, "
  1382. f"extra={sorted(response_names-expected_response_names)}"
  1383. )
  1384. primary_records = [
  1385. _comparison_record(row, "primary_holm_12")
  1386. for row in primary_frame.to_dict(orient="records")
  1387. ]
  1388. ablation_records = [
  1389. _comparison_record(row, gate.EXPLORATORY_FAMILY)
  1390. for row in ablation_frame.to_dict(orient="records")
  1391. ]
  1392. performance = _descriptive_performance(primary)
  1393. strongest = _strongest_ordinary(performance)
  1394. supporting = _ablation_supporting_endpoints(runs)
  1395. costs = _cost_summary(training_costs, inference_costs, cost_metadata)
  1396. stability = _stability_summary(primary)
  1397. primary_summary_sentence = _summary_sentence(
  1398. primary_records, "predeclared primary comparisons"
  1399. )
  1400. ablation_summary_sentence = _summary_sentence(
  1401. ablation_records, "exploratory ablation comparisons"
  1402. )
  1403. manuscript_map = _placeholder_map(
  1404. manuscript_inventory,
  1405. performance,
  1406. primary_records,
  1407. ablation_records,
  1408. costs,
  1409. stability,
  1410. primary_summary_sentence,
  1411. ablation_summary_sentence,
  1412. )
  1413. clauses = _response_clauses(primary_records, ablation_records, costs, stability)
  1414. payload: dict[str, Any] = {
  1415. "schema_version": SCHEMA_VERSION,
  1416. "generated_utc": datetime.now(timezone.utc).isoformat(),
  1417. "status": "validated_fill_sheet_not_final_author_interpretation",
  1418. "provenance": {
  1419. "primary_dir": str(primary.spec.path),
  1420. "ablation_dir": str(Path(ablation_dir).expanduser().resolve()),
  1421. "cost_profile_dir": str(Path(cost_profile_dir).expanduser().resolve()),
  1422. "primary_input_artifact_sha256": primary.hashes["input_artifact_sha256"],
  1423. "primary_comparisons_sha256": _sha256(primary.spec.path / "primary_comparisons.csv"),
  1424. "ablation_manifest_sha256": _sha256(Path(ablation_dir).expanduser().resolve() / "result_manifest.json"),
  1425. "ablation_comparisons_sha256": ablation_manifest["outputs"]["ablation_comparisons"]["sha256"],
  1426. "cost_profile_sha256": cost_hashes,
  1427. },
  1428. "protocol": {
  1429. "train_seeds": list(EXPECTED_TRAIN_SEEDS),
  1430. "eval_seeds": list(EXPECTED_EVAL_SEEDS),
  1431. "fixed_trajectories_per_setting": 20,
  1432. "pilot_scan_allocation": "2 scans x 10 replicates",
  1433. "transfer_scan_allocation": "10 scans x 2 replicates",
  1434. "inferential_unit": "independently trained seed after averaging its 20 fixed trajectories",
  1435. "primary_holm_family_size": 12,
  1436. "exploratory_ablation_holm_family_size": 24,
  1437. },
  1438. "primary_comparisons": primary_records,
  1439. "primary_classification_counts": _summarise_classifications(primary_records),
  1440. "primary_summary_english": primary_summary_sentence,
  1441. "descriptive_performance": performance,
  1442. "posthoc_ordinary_comparator_audit": strongest,
  1443. "ablation_comparisons": ablation_records,
  1444. "ablation_classification_counts": _summarise_classifications(ablation_records),
  1445. "ablation_summary_english": ablation_summary_sentence,
  1446. "ablation_supporting_endpoints": supporting,
  1447. "costs": costs,
  1448. "stability": stability,
  1449. "response_placeholder_clauses": clauses,
  1450. "manuscript_placeholder_map": manuscript_map,
  1451. "placeholder_inventory": {
  1452. "manuscript": manuscript_inventory,
  1453. "response": response_inventory,
  1454. },
  1455. "human_review_required": [
  1456. "Choose the abstract's limited set of principal numerical results without significance cherry-picking.",
  1457. "Judge practical importance; the script only classifies CI/Holm direction and never defines a smallest important effect.",
  1458. "Decide the sequentiality claim from all six gamma-zero rows, retaining null or adverse settings.",
  1459. "Decide component-necessity wording from all 24 exploratory rows; do not turn inconclusive results into proof of necessity or equivalence.",
  1460. "If the descriptively best ordinary rule differs from the predeclared comparator, report that ranking without borrowing the predeclared adjusted p-value.",
  1461. "Complete the manuscript-wide language and caption audit after all values and figures are inserted and the document is compiled.",
  1462. "Insert the main figure and its final caption manually.",
  1463. "The polynomial-branch weight subsection requires its separate frozen-checkpoint audit artifact; these three input directories do not supply it.",
  1464. "No switching-cost sensitivity, MAC/FLOP, peak-memory, acquisition-session, new-site, or hardware-benefit claim can be generated from these artifacts.",
  1465. "Cost timing is an engineering profile on the recorded hardware and is not a policy-performance inferential endpoint.",
  1466. ],
  1467. }
  1468. markdown = _render_response_markdown(payload, clauses)
  1469. target.mkdir(parents=True, exist_ok=True)
  1470. manuscript_output.write_text(
  1471. json.dumps(_json_safe(payload), indent=2, ensure_ascii=False, sort_keys=True) + "\n",
  1472. encoding="utf-8",
  1473. )
  1474. response_output.write_text(markdown, encoding="utf-8")
  1475. return payload
  1476. def _parser() -> argparse.ArgumentParser:
  1477. project = Path(__file__).resolve().parents[1]
  1478. parser = argparse.ArgumentParser(
  1479. description=(
  1480. "Strictly validate canonical primary, ablation, and isolated cost artifacts, "
  1481. "then build auditable manuscript/response fill sheets."
  1482. )
  1483. )
  1484. parser.add_argument("--primary-dir", type=Path, required=True)
  1485. parser.add_argument("--ablation-dir", type=Path, required=True)
  1486. parser.add_argument("--cost-profile-dir", type=Path, required=True)
  1487. parser.add_argument("--output-dir", type=Path, required=True)
  1488. parser.add_argument(
  1489. "--manuscript-template", type=Path, default=project / "manuscript" / "template.tex"
  1490. )
  1491. parser.add_argument(
  1492. "--response-draft", type=Path, default=project / "response" / "response_draft.md"
  1493. )
  1494. parser.add_argument("--overwrite", action="store_true")
  1495. return parser
  1496. def main(argv: Sequence[str] | None = None) -> int:
  1497. args = _parser().parse_args(argv)
  1498. try:
  1499. payload = summarize_revision_results(
  1500. primary_dir=args.primary_dir,
  1501. ablation_dir=args.ablation_dir,
  1502. cost_profile_dir=args.cost_profile_dir,
  1503. output_dir=args.output_dir,
  1504. manuscript_template=args.manuscript_template,
  1505. response_draft=args.response_draft,
  1506. overwrite=args.overwrite,
  1507. )
  1508. except (ValidationError, gate.ValidationError) as exc:
  1509. print(f"validation failed: {exc}", file=sys.stderr)
  1510. return 2
  1511. print(
  1512. f"validated {len(payload['primary_comparisons'])} primary and "
  1513. f"{len(payload['ablation_comparisons'])} exploratory comparisons; wrote fill sheets"
  1514. )
  1515. return 0
  1516. if __name__ == "__main__":
  1517. raise SystemExit(main())

summarize_revision_results.py at commit 6b62ded, no license · at the source

Overview

Authors: Yuxuan Pan1, Ying Yan1,2, Dingyi Sun3, Zhenyu Li4, Zhixuan Zhang5, Jun Cai2, Dapeng Chen2, Qi Wu6, Zongyuan Shen7
  1. Reading Academy, Nanjing University of Information Science and Technology, Nanjing 210044, China
  2. School of Automation, Nanjing University of Information Science and Technology, Nanjing 210044, China; (J.C.); (D.C.)
  3. School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing 210044, China
  4. Nanjing Research Institute of Electronics Technology, Nanjing 210039, China
  5. School of Electronics & Information Engineering, Nanjing University of Information Science and Technology, Nanjing 210044, China
  6. Department of Automation, Shanghai Jiao Tong University, Shanghai 200240, China
  7. College of Information Science and Technology, Jinan University, Guangzhou 510632, China
Journal: Sensors (Basel, Switzerland), volume 26, issue 17, article 5501
Dates: received 8 July 2026; accepted 28 August 2026; published online 30 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/s26175501 · PMID 42740121 · PMCID PMC13568378 · OpenAlex W7204802221
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: cellular / molecular (subfield)
Methods: Statistics, Machine learning
Keywords: HNP-DQN, cognitive radar, adaptive spectrum selection, deep reinforcement learning, polynomial feature expansion, channel switching cost
Topic: Wireless Signal Modulation Classification (Artificial Intelligence, Computer Science), according to OpenAlex
Citations: not cited yet (Europe PMC); 39 references in the paper

Abstract

Switching-aware spectrum selection requires balancing interference avoidance against retuning costs. This paper evaluates a hybrid neural–polynomial deep Q-network (HNP-DQN) in an eight-channel, measurement-driven 5 GHz testbed with exogenous sweeping and random interference. The architecture combines learned latent features with an element-wise second-order expansion, layer normalization (LayerNorm), and a dueling Double Deep Q-Network (DDQN) backbone. Files are separated before window construction, and the evaluation includes a pre-inspected pilot and two outcome-uninspected distance and power transfers. The comparison includes strong deterministic rules and a capacity-matched multilayer perceptron (MLP) trained with the same DDQN procedure. All methods are evaluated on common trajectories with paired inference across 10 independently trained seeds and Holm correction. Under deterministic sweeping, the schedule-aware rule was given the declared initial phase and one-channel-per-step direction and used an internal step counter to track the deterministic progression. These schedule variables and the corresponding step index are absent from HNP-DQN’s 48-dimensional observation. The rule matched the trajectory-wise dynamic-programming upper bound and outperformed HNP-DQN in all three settings (differences calculated as HNP-DQN minus the rule: −1.91, −3.51, and −1.97; all adjusted p=0.023). This information-asymmetric operational comparison shows the advantage attainable when the declared sweep specification and its progression are directly exploited; however, it does not provide a matched-information architectural ranking. By contrast, under random jamming, all comparisons between HNP-DQN and the threshold rule were inconclusive after correction, as were all six capacity-matched comparisons between HNP-DQN and the MLP. Of the 24 exploratory ablation comparisons, one favored the γ=0 variant in the random pilot, whereas the other 23 were inconclusive. Accordingly, this paper provides an information-aware, approximately parameter-matched reference for evaluating when explicit knowledge, observation-based control, or additional learning complexity is justified within the declared measurement-replay scope.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.

TheRyans520/HNPDQN

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 6b62ded98eda36646f4174bc2cb3bc671e1bf75d, 5 September 2026
Languages: Python (65), Shell (1)
Size: 4,913 files, 66 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, environment (requirements.txt, experiment_v2/requirements.txt), tests, continuous integration
Not found: license file, CITATION.cff, documentation
Tools: NumPy (41 files), PyTorch (19 files), pandas (18 files), Matplotlib (10 files), SciPy (4 files), seaborn (2 files), Pillow (1 file), TensorFlow (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
67 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 66 scripts, each with its path and the digest of its content;
  • 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The source code developed for this paper is publicly available at https://github.com/TheRyans520/HNPDQN (accessed on 27 August 2026). The RF Jamming dataset used in this paper is publicly available and documented by Ali et al. in “RF Jamming Dataset: A Wireless Spectral Scan Approach for Malicious Interference Detection”.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 6 keywords, 31 references.

Cite

This paper

Pan, Y., Yan, Y., Sun, D., Li, Z., Zhang, Z., Cai, J., Chen, D., Wu, Q., & Shen, Z. (2026). Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled Radio-Frequency Measurement-Replay Testbed. Sensors (Basel, Switzerland), 26(17), 5501. https://doi.org/10.3390/s26175501

BibTeX

@article{pan2026evaluation,
author = {Pan, Yuxuan and Yan, Ying and Sun, Dingyi and Li, Zhenyu and Zhang, Zhixuan and Cai, Jun and Chen, Dapeng and Wu, Qi and Shen, Zongyuan},
title = {{Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled Radio-Frequency Measurement-Replay Testbed}},
journal = {Sensors (Basel, Switzerland)},
year = {2026},
month = aug,
volume = {26},
number = {17},
pages = {5501},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1424-8220},
doi = {10.3390/s26175501},
url = {https://doi.org/10.3390/s26175501},
pmid = {42740121},
pmcid = {PMC13568378}
}

RIS

TY - JOUR
AU - Pan, Yuxuan
AU - Yan, Ying
AU - Sun, Dingyi
AU - Li, Zhenyu
AU - Zhang, Zhixuan
AU - Cai, Jun
AU - Chen, Dapeng
AU - Wu, Qi
AU - Shen, Zongyuan
TI - Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled Radio-Frequency Measurement-Replay Testbed
T2 - Sensors (Basel, Switzerland)
J2 - Sensors (Basel)
PY - 2026
DA - 2026/08/30
VL - 26
IS - 17
SP - 5501
SN - 1424-8220
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/s26175501
UR - https://doi.org/10.3390/s26175501
LA - en
ER -

CSL-JSON

{
"id": "10.3390/s26175501",
"type": "article-journal",
"title": "Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled Radio-Frequency Measurement-Replay Testbed",
"container-title": "Sensors (Basel, Switzerland)",
"author": [
{
"family": "Pan",
"given": "Yuxuan"
},
{
"family": "Yan",
"given": "Ying"
},
{
"family": "Sun",
"given": "Dingyi"
},
{
"family": "Li",
"given": "Zhenyu"
},
{
"family": "Zhang",
"given": "Zhixuan"
},
{
"family": "Cai",
"given": "Jun"
},
{
"family": "Chen",
"given": "Dapeng"
},
{
"family": "Wu",
"given": "Qi"
},
{
"family": "Shen",
"given": "Zongyuan"
}
],
"container-title-short": "Sensors (Basel)",
"volume": "26",
"issue": "17",
"page": "5501",
"DOI": "10.3390/s26175501",
"PMID": "42740121",
"PMCID": "PMC13568378",
"ISSN": "1424-8220",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/s26175501",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
30
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3390/s26113488
Towards Interpretable Seizure Detection: An Excitation/Inhibition Dynamic Polynomial Network Framework for Electroencephalography.
Journal: Sensors (Basel, Switzerland)
In common: 2 authors
[2] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: TensorFlow, Pillow, PyTorch, 5 other tools, cellular / molecular
[3] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[4] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[5] doi:10.1371/journal.pcbi.1014571 [code]
SynAPSeg: A novel dataset and image analysis framework for deep learning-based synapse detection and quantification.
Journal: PLoS computational biology
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[6] doi:10.1364/boe.605322 [code]
Generalized plaque digitization framework for multi-dimensional mesoscopic images.
Journal: Biomedical optics express
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[7] doi:10.1523/eneuro.0023-26.2026 [code]
Real-Time Segmentation and Classification of Birdsong Syllables for Learning Experiments.
Journal: eNeuro
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[8] doi:10.1038/s41467-026-74358-5 [code]
Brain-inspired spatial intelligence for embodied agents.
Journal: Nature communications
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[9] doi:10.1364/boe.600665 [code]
NeuroSeg-MF: robust neuron segmentation in two-photon Ca&lt;sup&gt;2+&lt;/sup&gt; imaging using multi-feature fusion and detection-guided SAM.
Journal: Biomedical optics express
In common: TensorFlow, Pillow, PyTorch, 5 other tools
[10] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: TensorFlow, Pillow, PyTorch, 5 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.