OSCR

Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes

Code ↔ Paper

1 match between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 1 match
  1. [1] § Material and Methods › BN training › Variable discretization and algorithm ↔ layerbn/bn_utils.py, lines 161–224 · score 0.64 · DiscreteTypeProcessor, pyAgrum, enforce, K2, Bayesian, bins

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 648 lines · 26 KB · MIT · 1 match

  1. """Bayesian network structure learning, bootstrapping, and scenarios.
  2. Everything cohort-specific is a parameter: which layers to exclude, which
  3. layer names mark the outcome and dropout roles, the score, the seed. No
  4. cohort's conventions are assumed.
  5. These functions take the layer map, the discretisation and the layer role
  6. patterns as separate arguments, which means a caller has to keep them
  7. consistent with the spec at every call site. `layerbn.analysis.Analysis`
  8. does that for you and is the recommended entry point; use these directly
  9. when you need control it does not expose.
  10. """
  11. from __future__ import annotations
  12. import random
  13. from collections import defaultdict
  14. from collections.abc import Callable, Iterable, Mapping, Sequence
  15. from itertools import product
  16. from typing import Any
  17. import numpy as np
  18. import pandas as pd
  19. from layerbn.discretisation import state_for_value
  20. # Layer-name substrings that identify semantic roles. Override when your
  21. # cohort names its layers differently.
  22. DEFAULT_OUTCOME_LAYER_PATTERNS: tuple[str, ...] = ("Outcomes",)
  23. DEFAULT_DROPOUT_LAYER_PATTERNS: tuple[str, ...] = ("Dropout",)
  24. def _match_any(layer_name: str, patterns: Sequence[str]) -> bool:
  25. return any(p in layer_name for p in patterns)
  26. def collect_outcome_vars(
  27. layer_map: Mapping[str, Sequence[str]],
  28. *,
  29. outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
  30. dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
  31. ) -> tuple[list[str], list[str]]:
  32. """Return `(outcome_vars, dropout_vars)` inferred from layer names."""
  33. outcome_vars: list[str] = []
  34. dropout_vars: list[str] = []
  35. for layer, variables in layer_map.items():
  36. if _match_any(layer, dropout_patterns):
  37. dropout_vars.extend(variables)
  38. elif _match_any(layer, outcome_patterns):
  39. outcome_vars.extend(variables)
  40. return outcome_vars, dropout_vars
  41. def configure_guided_structure(
  42. learner: Any,
  43. layer_map_filtered: Mapping[str, Sequence[str]],
  44. variables_to_keep: Sequence[str],
  45. outcomes: Sequence[str],
  46. *,
  47. dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
  48. within_layers: bool = True,
  49. arcs_between_outcomes: bool = False,
  50. selection_parents: str = "outcomes",
  51. forbidden_pairs: Iterable[tuple[str, str]] = (),
  52. mandatory_pairs: Iterable[tuple[str, str]] = (),
  53. no_parents: Iterable[str] = (),
  54. no_children: Iterable[str] = (),
  55. ) -> tuple[list[str], list[str]]:
  56. """Restrict the learner to arcs that go downward through the layers.
  57. Rules always encoded:
  58. * Arcs may only go from an earlier layer to the same or a later layer.
  59. Rules that the keyword arguments control, whose defaults reproduce the
  60. behaviour this function had before they existed:
  61. * `within_layers` — whether two variables in the same layer may be
  62. connected. False makes each layer internally independent.
  63. * `arcs_between_outcomes` — whether one outcome may point at another.
  64. * `selection_parents` — `"outcomes"` allows arcs into a dropout-layer
  65. variable only from an outcome, on the grounds that dropout follows
  66. the outcomes rather than the covariates. `"any"` drops that
  67. restriction and lets the layer order alone decide.
  68. Additional structural knowledge:
  69. * `forbidden_pairs` — `(parent, child)` arcs to rule out even though
  70. the layer order permits them.
  71. * `mandatory_pairs` — arcs the learner must include. These must be
  72. consistent with everything above; a pair that is both mandatory and
  73. forbidden raises `ValueError` rather than reaching pyAgrum, whose own
  74. message does not say which pair was at fault.
  75. * `no_parents` / `no_children` — variables pinned as roots or sinks.
  76. Returns `(dropout_vars, true_outcomes)` for the caller to reuse when
  77. adding hard outcome→dropout arcs after learning.
  78. """
  79. if selection_parents not in ("outcomes", "any"):
  80. raise ValueError(
  81. f"selection_parents must be 'outcomes' or 'any', got {selection_parents!r}")
  82. layer_keys = list(layer_map_filtered.keys())
  83. dropout_vars = [v for k in layer_keys if _match_any(k, dropout_patterns)
  84. for v in layer_map_filtered[k]]
  85. true_outcomes = [v for v in outcomes if v not in dropout_vars]
  86. keep = set(variables_to_keep)
  87. forbidden_pairs = set(forbidden_pairs)
  88. mandatory_pairs = set(mandatory_pairs)
  89. no_parents = set(no_parents)
  90. no_children = set(no_children)
  91. conflict = forbidden_pairs & mandatory_pairs
  92. if conflict:
  93. raise ValueError(
  94. f"these arcs are both mandatory and forbidden: {sorted(conflict)}")
  95. allowed_arcs: list[tuple[str, str]] = []
  96. for i, from_key in enumerate(layer_keys):
  97. for j, to_key in enumerate(layer_keys[i:], start=i):
  98. if i == j and not within_layers:
  99. continue
  100. for parent, child in product(layer_map_filtered[from_key],
  101. layer_map_filtered[to_key]):
  102. if child in dropout_vars and selection_parents == "outcomes":
  103. if parent in outcomes:
  104. allowed_arcs.append((parent, child))
  105. continue
  106. if parent in keep and child in keep:
  107. allowed_arcs.append((parent, child))
  108. allowed_set = set(allowed_arcs) - forbidden_pairs
  109. for parent in variables_to_keep:
  110. for child in variables_to_keep:
  111. if parent != child and (parent, child) not in allowed_set:
  112. learner.addForbiddenArc(parent, child)
  113. if not arcs_between_outcomes:
  114. for parent in true_outcomes:
  115. for child in true_outcomes:
  116. if parent != child:
  117. learner.addForbiddenArc(parent, child)
  118. # Node-level constraints are applied after the pairwise ones so that a
  119. # variable pinned as a root or a sink stays that way regardless of what
  120. # the layer order allowed.
  121. for variable in sorted(no_parents & keep):
  122. learner.addNoParentNode(variable)
  123. for variable in sorted(no_children & keep):
  124. learner.addNoChildrenNode(variable)
  125. # Mandatory arcs come last: pyAgrum resolves them against the
  126. # constraints already registered.
  127. for parent, child in sorted(mandatory_pairs):
  128. if parent in keep and child in keep:
  129. learner.addMandatoryArc(parent, child)
  130. return dropout_vars, true_outcomes
  131. def build_bn(
  132. df: pd.DataFrame,
  133. outcomes: Sequence[str],
  134. layer_map: Mapping[str, Sequence[str]],
  135. type_processor: Any,
  136. *,
  137. score: str = "K2",
  138. use_tabu: bool = True,
  139. random_seed: int = 42,
  140. enforce_structure: bool = True,
  141. fixed_template: Any = None,
  142. use_smoothing: bool = False,
  143. exclude_layers: Iterable[str] = (),
  144. max_indegree: int = 5,
  145. outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
  146. dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
  147. within_layers: bool = True,
  148. arcs_between_outcomes: bool = False,
  149. selection_parents: str = "outcomes",
  150. forbidden_pairs: Iterable[tuple[str, str]] = (),
  151. mandatory_pairs: Iterable[tuple[str, str]] = (),
  152. no_parents: Iterable[str] = (),
  153. no_children: Iterable[str] = (),
  154. ) -> Any:
  155. """Learn a layered Bayesian network with pyAgrum.
  156. Parameters
  157. ----------
  158. df : DataFrame
  159. Fully-imputed input data. Columns not needed as outcomes and not
  160. listed in `exclude_layers` become learnable nodes.
  161. outcomes : sequence of str
  162. Outcome variable names to keep in the network. Any *other* variable
  163. that belongs to an outcome/dropout layer is dropped before learning.
  164. layer_map : mapping
  165. `{layer_name: [variable, ...]}` defining the expert layering.
  166. type_processor : DiscreteTypeProcessor
  167. Constructed via `layerbn.discretisation.make_type_processor`.
  168. score : {"K2", "BIC", "BDeu"}
  169. Structure-learning score.
  170. fixed_template : optional
  171. Use a pre-computed discretisation template so bootstrap resamples
  172. share bin edges. Build once with `type_processor.discretizedTemplate(df)`.
  173. exclude_layers : iterable of str
  174. Layer names whose variables should be dropped before learning, for
  175. instance to leave out a layer measured on only part of the cohort.
  176. use_smoothing : bool
  177. Add a Laplace prior, so that a parent configuration nobody in the
  178. sample happens to occupy gets a small probability rather than zero.
  179. Always on for BIC, which carries no prior of its own and whose
  180. zero cells otherwise break inference on bootstrap resamples.
  181. Ignored for K2 and BDeu: both already contain an implicit prior,
  182. and adding a second one double-counts it. aGrUM reports that as
  183. "the K2 score already contains a different 'implicit' prior ...
  184. the learning will probably be biased", which is exactly what would
  185. happen, so the request is dropped rather than honoured.
  186. within_layers, arcs_between_outcomes, selection_parents
  187. The three structural conventions. See `configure_guided_structure`;
  188. the defaults reproduce the behaviour that used to be fixed in code.
  189. forbidden_pairs, mandatory_pairs, no_parents, no_children
  190. Extra structural knowledge, as `(parent, child)` pairs and variable
  191. names. `layerbn.spec` builds these from a spec's `constraints`
  192. block, and `layerbn.analysis.Analysis` passes them in for you.
  193. Notes
  194. -----
  195. * When `enforce_structure=True`, `configure_guided_structure` restricts
  196. arcs to obey the layer order and any constraints given.
  197. * When `enforce_structure=False` every constraint is skipped, including
  198. the layer order. That is a genuinely unconstrained search, useful as a
  199. comparison but not as the analysis.
  200. * After learning, hard outcome→dropout arcs are added if not already
  201. present, so scenario inference always sees the dropout mechanism.
  202. """
  203. import pyagrum as gum
  204. random.seed(random_seed)
  205. np.random.seed(random_seed)
  206. exclude_layers = set(exclude_layers)
  207. layer_map_filtered = {k: v for k, v in layer_map.items() if k not in exclude_layers}
  208. excluded_vars = [v for k in exclude_layers for v in layer_map.get(k, [])]
  209. outcome_vars_all, dropout_vars_all = collect_outcome_vars(
  210. layer_map, outcome_patterns=outcome_patterns, dropout_patterns=dropout_patterns,
  211. )
  212. outcome_layer_vars = outcome_vars_all + dropout_vars_all
  213. drop_outcomes = [v for v in outcome_layer_vars if v not in outcomes]
  214. df_reduced = df.drop(columns=drop_outcomes + excluded_vars, errors="ignore")
  215. variables_to_keep = df_reduced.columns.tolist()
  216. template = fixed_template if fixed_template is not None \
  217. else type_processor.discretizedTemplate(df_reduced)
  218. learner = gum.BNLearner(df_reduced, template)
  219. score = score.upper()
  220. if score == "K2":
  221. learner.useScoreK2()
  222. elif score == "BIC":
  223. learner.useScoreBIC()
  224. elif score == "BDEU":
  225. learner.useScoreBDeu(ess=1)
  226. else:
  227. raise ValueError(f"Unsupported score {score!r}: choose K2, BIC, or BDeu")
  228. # Only BIC has no prior of its own. K2 and BDeu each carry an implicit
  229. # one, and stacking a smoothing prior on top of it biases the score --
  230. # aGrUM emits a notification saying so for every fit, which under a
  231. # bootstrap means hundreds of them. See `use_smoothing` above.
  232. if score == "BIC":
  233. learner.useSmoothingPrior()
  234. if use_tabu:
  235. learner.useLocalSearchWithTabuList()
  236. else:
  237. learner.useGreedyHillClimbing()
  238. if enforce_structure:
  239. dropout_vars, true_outcomes = configure_guided_structure(
  240. learner, layer_map_filtered, variables_to_keep, outcomes,
  241. dropout_patterns=dropout_patterns,
  242. within_layers=within_layers,
  243. arcs_between_outcomes=arcs_between_outcomes,
  244. selection_parents=selection_parents,
  245. forbidden_pairs=forbidden_pairs,
  246. mandatory_pairs=mandatory_pairs,
  247. no_parents=no_parents,
  248. no_children=no_children,
  249. )
  250. else:
  251. # Same pair `configure_guided_structure` would have returned, without
  252. # touching the learner. NB: `collect_outcome_vars` returns
  253. # (outcome_vars, dropout_vars) — the opposite order — so it must not be
  254. # unpacked into (dropout_vars, true_outcomes) directly.
  255. dropout_vars = [v for k in layer_map_filtered
  256. if _match_any(k, dropout_patterns)
  257. for v in layer_map_filtered[k]]
  258. true_outcomes = [v for v in outcomes if v not in dropout_vars]
  259. learner.setMaxIndegree(max_indegree)
  260. bn = learner.learnBN()
  261. if enforce_structure:
  262. # Connect every outcome to the selection layer, so scenario inference
  263. # always sees the dropout mechanism even when the score did not pick
  264. # the arc up. An arc the caller forbade explicitly is left out: the
  265. # spec's promise is that constraints narrow what the layer order
  266. # allows, and silently reinstating one here would break it.
  267. forbidden = set(forbidden_pairs)
  268. try:
  269. names = set(bn.names())
  270. for dropout in (n for n in dropout_vars if n in names):
  271. for outcome in (n for n in true_outcomes if n in names):
  272. if (outcome, dropout) in forbidden:
  273. continue
  274. out_id, drop_id = bn.idFromName(outcome), bn.idFromName(dropout)
  275. if not bn.existsArc(out_id, drop_id):
  276. bn.addArc(out_id, drop_id)
  277. except Exception:
  278. pass # gum.InvalidArgument etc. — arc already exists / var missing
  279. return bn
  280. def bootstrap_edge_frequencies(
  281. df: pd.DataFrame,
  282. *,
  283. outcomes: Sequence[str],
  284. layer_map: Mapping[str, Sequence[str]],
  285. type_processor: Any,
  286. build_bn_func: Callable[..., Any] = build_bn,
  287. n_bootstraps: int = 100,
  288. random_seed: int = 42,
  289. **build_kwargs: Any,
  290. ) -> dict[tuple[str, str], float]:
  291. """Fit `build_bn_func` on `n_bootstraps` resamples; return per-arc frequencies.
  292. Extra keyword arguments are forwarded to `build_bn_func` unchanged so
  293. callers can pin `score`, `use_smoothing`, `fixed_template`, etc.
  294. Notes
  295. -----
  296. The denominator is `n_bootstraps`, **not** the number of resamples that
  297. fitted successfully. A failed fit therefore lowers every frequency rather
  298. than being excluded from the calculation. These frequencies are reported
  299. as the `f=NN%` arc labels in published figures, so changing the
  300. denominator would silently restate them.
  301. """
  302. counts: dict[tuple[str, str], int] = defaultdict(int)
  303. successes = 0
  304. for i in range(n_bootstraps):
  305. seed = random_seed + i
  306. boot = df.sample(frac=1.0, replace=True, random_state=seed)
  307. try:
  308. bn = build_bn_func(
  309. boot, outcomes=outcomes, layer_map=layer_map,
  310. type_processor=type_processor, random_seed=seed, **build_kwargs,
  311. )
  312. except Exception as exc:
  313. print(f"Bootstrap {i} failed: {exc}")
  314. continue
  315. successes += 1
  316. for parent_id, child_id in bn.arcs():
  317. counts[(bn.variable(parent_id).name(), bn.variable(child_id).name())] += 1
  318. if successes < n_bootstraps:
  319. print(f"Note: {n_bootstraps - successes}/{n_bootstraps} bootstrap fits "
  320. "failed; frequencies are still divided by n_bootstraps.")
  321. return {edge: count / n_bootstraps for edge, count in counts.items()}
  322. def bootstrap_scenario_risks(
  323. df: pd.DataFrame,
  324. *,
  325. scenario_profiles: Sequence[tuple[str, Mapping[str, Any]]],
  326. target_outcomes: Sequence[tuple[str, str]],
  327. build_bn_func: Callable[..., Any] = build_bn,
  328. layer_map: Mapping[str, Sequence[str]],
  329. type_processor: Any,
  330. outcomes_for_learning: Sequence[str],
  331. n_bootstraps: int = 200,
  332. random_seed: int = 42,
  333. **build_kwargs: Any,
  334. ) -> pd.DataFrame:
  335. """Compute posterior probabilities under a set of patient profiles.
  336. Parameters
  337. ----------
  338. scenario_profiles : sequence of (label, evidence_dict)
  339. Each dict maps a variable name to the evidence for it, passed to
  340. `LazyPropagation.setEvidence` unchanged. Prefer **state indices**:
  341. pyAgrum reads a bare integer as a state index, a numeric *string*
  342. as a value to place in a bin, and rejects an interval label such as
  343. `'(45;64.6['` outright. `layerbn.analysis.Analysis.resolve_profile`
  344. converts any of these to indices and reports what it could not
  345. match; using it avoids the failure mode below.
  346. target_outcomes : sequence of (var_name, display_label)
  347. outcomes_for_learning : sequence of str
  348. Passed to `build_bn_func` as the fixed set of outcomes so bootstrap
  349. networks are structurally comparable across resamples.
  350. Notes
  351. -----
  352. Evidence that pyAgrum rejects raises `gum.InvalidArgument`, which this
  353. function catches and skips. The affected scenario then produces no rows
  354. at all rather than an error, so check that every profile you passed
  355. appears in the result.
  356. Pass `fixed_template` (via `**build_kwargs`) when the profiles refer to
  357. discretised variables. Without it every resample recomputes its own bin
  358. edges, and a fixed profile silently refers to a different group of
  359. people in each resample.
  360. """
  361. import pyagrum as gum
  362. samples: dict[tuple[str, str, str], list[float]] = defaultdict(list)
  363. label_cache: dict[str, list[str]] = {}
  364. failed = 0
  365. for i in range(n_bootstraps):
  366. seed = random_seed + i
  367. boot = df.sample(frac=1.0, replace=True, random_state=seed)
  368. try:
  369. bn = build_bn_func(
  370. boot, outcomes=list(outcomes_for_learning), layer_map=layer_map,
  371. type_processor=type_processor, random_seed=seed, **build_kwargs,
  372. )
  373. except Exception:
  374. failed += 1
  375. continue
  376. for scenario_name, evidence in scenario_profiles:
  377. ie = gum.LazyPropagation(bn)
  378. try:
  379. ie.setEvidence(dict(evidence))
  380. except gum.InvalidArgument:
  381. continue
  382. for outcome_name, _ in target_outcomes:
  383. if outcome_name not in label_cache:
  384. label_cache[outcome_name] = list(
  385. bn.variable(bn.idFromName(outcome_name)).labels()
  386. )
  387. posterior = ie.posterior(outcome_name)
  388. for idx, state in enumerate(label_cache[outcome_name]):
  389. samples[(scenario_name, outcome_name, state)].append(float(posterior[idx]))
  390. outcome_labels = dict(target_outcomes)
  391. rows = []
  392. for (scenario_name, outcome_name, state), values in samples.items():
  393. arr = np.asarray(values, float)
  394. arr = arr[np.isfinite(arr)]
  395. if arr.size == 0:
  396. continue
  397. rows.append({
  398. "Scenario": scenario_name,
  399. "Outcome": outcome_labels.get(outcome_name, outcome_name),
  400. "Category": state,
  401. "Mean probability": arr.mean(),
  402. "Std": arr.std(ddof=1) if arr.size > 1 else 0.0,
  403. "CI 2.5%": np.quantile(arr, 0.025),
  404. "CI 97.5%": np.quantile(arr, 0.975),
  405. "Bootstraps": int(arr.size),
  406. })
  407. if failed:
  408. print(f"Warning: {failed}/{n_bootstraps} bootstrap fits failed and were skipped.")
  409. return (pd.DataFrame(rows)
  410. .sort_values(["Scenario", "Outcome", "Category"])
  411. .reset_index(drop=True))
  412. def descendants(bn: Any, start_id: int) -> set[int]:
  413. """All nodes reachable from `start_id` by following directed arcs."""
  414. seen: set[int] = set()
  415. stack = [start_id]
  416. while stack:
  417. for child in bn.children(stack.pop()):
  418. if child not in seen:
  419. seen.add(child)
  420. stack.append(child)
  421. return seen
  422. def bootstrap_knob_sweep(
  423. df: pd.DataFrame,
  424. *,
  425. base_profile: Mapping[str, Any],
  426. knob: str,
  427. outcomes: Sequence[str],
  428. outcomes_for_learning: Sequence[str],
  429. layer_map: Mapping[str, Sequence[str]],
  430. type_processor: Any,
  431. build_bn_func: Callable[..., Any] = build_bn,
  432. n_bootstraps: int = 200,
  433. random_seed: int = 42,
  434. verbose: bool = True,
  435. **build_kwargs: Any,
  436. ) -> tuple[pd.DataFrame, dict[str, Any]]:
  437. """Sensitivity sweep: fix a patient, vary one variable across its states.
  438. Returns `(sweep_df, meta)` where `sweep_df` has one row per
  439. (outcome, knob state, outcome state) with bootstrap mean and 95% CI.
  440. `meta` carries the resolved knob states and outcome states.
  441. Notes
  442. -----
  443. Warns if any variable in `base_profile` lies on a directed path from
  444. `knob` to an outcome (overadjustment / blocked mediator). Uses a fixed
  445. discretisation template so knob states mean the same thing across
  446. resamples.
  447. """
  448. import pyagrum as gum
  449. fixed_template = type_processor.discretizedTemplate(
  450. _reduced_for_template(
  451. df, layer_map, outcomes_for_learning,
  452. build_kwargs.get("exclude_layers", ()),
  453. outcome_patterns=build_kwargs.get(
  454. "outcome_patterns", DEFAULT_OUTCOME_LAYER_PATTERNS),
  455. dropout_patterns=build_kwargs.get(
  456. "dropout_patterns", DEFAULT_DROPOUT_LAYER_PATTERNS),
  457. )
  458. )
  459. # `use_smoothing` is deliberately NOT forced on here. It only takes
  460. # effect under BIC (see `build_bn`), and forcing it meant every sweep
  461. # resample was learned under a different prior from the network being
  462. # reported -- so the interval described a procedure nobody ran.
  463. build_kwargs = {**build_kwargs, "fixed_template": fixed_template}
  464. ref_bn = build_bn_func(
  465. df, outcomes=list(outcomes_for_learning), layer_map=layer_map,
  466. type_processor=type_processor, random_seed=random_seed, **build_kwargs,
  467. )
  468. def labels_of(name: str) -> list[str]:
  469. return list(ref_bn.variable(ref_bn.idFromName(name)).labels())
  470. if knob in base_profile:
  471. raise ValueError("base_profile must not contain the knob variable.")
  472. applied, rejected, snapped = {}, {}, {}
  473. for var, val in base_profile.items():
  474. if var not in ref_bn.names():
  475. rejected[var] = val
  476. continue
  477. state = state_for_value(labels_of(var), val)
  478. if state is None:
  479. rejected[var] = val
  480. else:
  481. applied[var] = state
  482. if str(val) != state:
  483. snapped[var] = f"{val} -> {state!r}"
  484. if verbose:
  485. print("Fixed patient — evidence applied:", applied or "(none)")
  486. if snapped:
  487. print("Fixed patient — snapped to bin:", snapped)
  488. if rejected:
  489. print("Fixed patient — REJECTED:", rejected,
  490. "(unknown variable or value not matchable — check the discretised template).")
  491. knob_descendants = {
  492. ref_bn.variable(i).name()
  493. for i in descendants(ref_bn, ref_bn.idFromName(knob))
  494. }
  495. blocked = [v for v in applied if v in knob_descendants]
  496. if blocked and verbose:
  497. print(f"WARNING: base_profile fixes {blocked}, which lie downstream of "
  498. f"{knob!r}. Holding a mediator constant blocks part of the knob's "
  499. "effect on the outcomes.")
  500. applied_idx = {v: labels_of(v).index(s) for v, s in applied.items()}
  501. knob_states = labels_of(knob)
  502. outcome_states = {o: labels_of(o) for o in outcomes}
  503. samples: dict[tuple[str, str, str], list[float]] = defaultdict(list)
  504. failed = 0
  505. for i in range(n_bootstraps):
  506. seed = random_seed + i
  507. boot = df.sample(frac=1.0, replace=True, random_state=seed)
  508. try:
  509. bn = build_bn_func(
  510. boot, outcomes=list(outcomes_for_learning), layer_map=layer_map,
  511. type_processor=type_processor, random_seed=seed, **build_kwargs,
  512. )
  513. except Exception:
  514. failed += 1
  515. continue
  516. for k_idx, k_state in enumerate(knob_states):
  517. try:
  518. ie = gum.LazyPropagation(bn)
  519. ie.setEvidence({**applied_idx, knob: k_idx})
  520. post = {o: ie.posterior(o) for o in outcomes}
  521. except Exception:
  522. continue
  523. for outcome in outcomes:
  524. for s_idx, s_label in enumerate(outcome_states[outcome]):
  525. samples[(outcome, k_state, s_label)].append(float(post[outcome][s_idx]))
  526. rows = []
  527. for (outcome, k_state, s_label), values in samples.items():
  528. arr = np.asarray(values, float)
  529. arr = arr[np.isfinite(arr)]
  530. if arr.size == 0:
  531. continue
  532. rows.append({
  533. "Outcome": outcome,
  534. knob: k_state,
  535. "Outcome state": s_label,
  536. "P": arr.mean(),
  537. "CI_low": np.quantile(arr, 0.025),
  538. "CI_high": np.quantile(arr, 0.975),
  539. "Bootstraps": int(arr.size),
  540. })
  541. if verbose and failed:
  542. print(f"Note: {failed}/{n_bootstraps} bootstrap fits failed and were skipped.")
  543. if not rows:
  544. raise RuntimeError("No valid scenarios were produced — check base_profile.")
  545. return pd.DataFrame(rows), {
  546. "knob": knob, "knob_states": knob_states,
  547. "outcomes": list(outcomes), "outcome_states": outcome_states,
  548. }
  549. def _reduced_for_template(
  550. df: pd.DataFrame,
  551. layer_map: Mapping[str, Sequence[str]],
  552. outcomes_for_learning: Sequence[str],
  553. exclude_layers: Iterable[str],
  554. *,
  555. outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
  556. dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
  557. ) -> pd.DataFrame:
  558. """Reproduce `build_bn`'s column reduction so one template covers all bootstraps.
  559. The patterns must be the same ones `build_bn` will be called with. If they
  560. are not, this function and `build_bn` disagree about which columns are
  561. outcomes, the template ends up describing a different set of variables
  562. than the learner is given, and `BNLearner` raises.
  563. """
  564. excluded = [v for k in exclude_layers for v in layer_map.get(k, [])]
  565. outcome_all, dropout_all = collect_outcome_vars(
  566. layer_map, outcome_patterns=outcome_patterns, dropout_patterns=dropout_patterns,
  567. )
  568. drop_outcomes = [v for v in outcome_all + dropout_all if v not in outcomes_for_learning]
  569. return df.drop(columns=drop_outcomes + excluded, errors="ignore")

bn_utils.py at commit fd5d083, under MIT · at the source

Overview

Authors: L. Malin Overmars1, Cor Allaart2,3, Esther E. Bron4, Hans-Peter Brunner La Rocca5, Jeroen de Bresser6, Majon Muller7, Matthias J.P. van Osch6, Charlotte Teunissen8, Betty M Tijms9, Frank J. Wolters4,10, Geert Jan Biessels1
  1. Department of Neurology and Neurosurgery, UMC Utrecht Brain Center, University Medical Center Utrecht, Utrecht, the Netherlands
  2. Department of Cardiology, Institute for Cardiovascular Research, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
  3. Department of Cardiology, Radboud University Medical Center, Radboud University, Nijmegen, The Netherlands
  4. Department of Radiology and Nuclear Medicine, Erasmus Medical Center, Rotterdam, the Netherlands
  5. Department of Cardiology, Cardiovascular Research Institute Maastricht, Maastricht University Medical Center, Maastricht, the Netherlands
  6. Department of Radiology, Leiden University Medical Center, Leiden, the Netherlands
  7. Department of Internal Medicine, Section Geriatrics, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
  8. Department of Clinical Chemistry, Neurochemistry Laboratory, Amsterdam Neuroscience, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
  9. Alzheimer Center Amsterdam, Department of Neurology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
  10. Department of Epidemiology, Erasmus MC University Medical Centre, Rotterdam, the Netherlands
Dates: published online 4 June 2026
Type: Preprint
License: CC BY-NC-ND
Identifiers: DOI 10.64898/2026.06.03.26354793 · OpenAlex W7163507635
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), Alzheimer's / dementia (population), stroke (population), clinical / translational (subfield)
Methods: Connectivity, Machine learning, Statistics
Keywords: Vascular cognitive impairment, Small vessel disease, BN, Prognosis, Explainability
Topic: Dementia and Cognitive Impairment Research (Psychiatry and Mental health, Medicine), according to OpenAlex
Funding: ZonMw (10510032120003)
Citations: not cited yet (Europe PMC); 40 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.

umcu/vci-bayes

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: fd5d083d2d20afde79d57d6c88e5c4e1017eece4, 24 August 2026
Languages: Python (19), Jupyter (1)
Size: 32 files, 20 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, license file, CITATION.cff, environment (pyproject.toml), tests, continuous integration, documentation, 1 notebook
Tools: pandas (8 files), NumPy (4 files), Matplotlib (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
22 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 20 scripts, each with its path and the digest of its content;
  • 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it points to the authors' code: umcu/vci-bayes
  • it says that the data are available on request

Read it in the paper: doi.org/10.64898/2026.06.03.26354793.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Authors: added L. Malin Overmars (0000-0001-7086-0864); Geert Jan Biessels (0000-0001-6862-2496); removed L. Malin Overmars; Geert Jan Biessels

Version 1, 27 September 2026: the first record

Recorded: type, journal, dates, 11 authors, 5 keywords, 1 funder, 36 references.

Cite

This paper

Overmars, L. M., Allaart, C., Bron, E. E., Brunner La Rocca, H.-P., de Bresser, J., Muller, M., van Osch, M. J., Teunissen, C., Tijms, B. M., Wolters, F. J., & Biessels, G. J. (2026). Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes. medRxiv (preprint). https://doi.org/10.64898/2026.06.03.26354793

BibTeX

@article{overmars2026bayesian,
author = {Overmars, L. Malin and Allaart, Cor and Bron, Esther E. and Brunner La Rocca, Hans-Peter and de Bresser, Jeroen and Muller, Majon and van Osch, Matthias J.P. and Teunissen, Charlotte and Tijms, Betty M and Wolters, Frank J. and Biessels, Geert Jan},
title = {{Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes}},
journal = {medRxiv (preprint)},
year = {2026},
month = jun,
publisher = {medRxiv},
doi = {10.64898/2026.06.03.26354793},
url = {https://doi.org/10.64898/2026.06.03.26354793}
}

RIS

TY - JOUR
AU - Overmars, L. Malin
AU - Allaart, Cor
AU - Bron, Esther E.
AU - Brunner La Rocca, Hans-Peter
AU - de Bresser, Jeroen
AU - Muller, Majon
AU - van Osch, Matthias J.P.
AU - Teunissen, Charlotte
AU - Tijms, Betty M
AU - Wolters, Frank J.
AU - Biessels, Geert Jan
TI - Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes
T2 - medRxiv (preprint)
J2 - medRxiv
PY - 2026
DA - 2026/06/04
PB - medRxiv
DO - 10.64898/2026.06.03.26354793
UR - https://doi.org/10.64898/2026.06.03.26354793
ER -

CSL-JSON

{
"id": "10.64898/2026.06.03.26354793",
"type": "article",
"title": "Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes",
"container-title": "medRxiv (preprint)",
"author": [
{
"family": "Overmars",
"given": "L. Malin"
},
{
"family": "Allaart",
"given": "Cor"
},
{
"family": "Bron",
"given": "Esther E."
},
{
"family": "Brunner La Rocca",
"given": "Hans-Peter"
},
{
"family": "de Bresser",
"given": "Jeroen"
},
{
"family": "Muller",
"given": "Majon"
},
{
"family": "van Osch",
"given": "Matthias J.P."
},
{
"family": "Teunissen",
"given": "Charlotte"
},
{
"family": "Tijms",
"given": "Betty M"
},
{
"family": "Wolters",
"given": "Frank J."
},
{
"family": "Biessels",
"given": "Geert Jan"
}
],
"container-title-short": "medRxiv",
"DOI": "10.64898/2026.06.03.26354793",
"publisher": "medRxiv",
"URL": "https://doi.org/10.64898/2026.06.03.26354793",
"issued": {
"date-parts": [
[
2026,
6,
4
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/braincomms/fcag257 [code]
Neuroplasticity and immune system are related to altered grey matter networks: a cohort study in sporadic Alzheimer's disease.
Journal: Brain communications
In common: Alzheimer's / dementia, 2 authors
[2] doi:10.1038/s41598-026-47047-y [code]
Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study.
Journal: Scientific reports
In common: scikit-learn, pandas, Matplotlib, 1 other tool, Alzheimer's / dementia, clinical / translational, author Frank J. Wolters
[3] doi:10.1038/s41467-026-71682-8 [code]
GWAS meta-analysis of cerebrospinal fluid Alzheimer's biomarkers reveals loci regulating lipids, brain volume and autophagy.
Journal: Nature communications
In common: scikit-learn, pandas, NumPy, Alzheimer's / dementia, author Betty M Tijms
[4] doi:10.1093/braincomms/fcag196 [code]
Music processing in behavioural variant frontotemporal dementia and Alzheimer's disease: a functional MRI study.
Journal: Brain communications
In common: Alzheimer's / dementia, author Betty M Tijms
[5] doi:10.1212/wnl.0000000000218472 [code]
Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.
Journal: Neurology
In common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, Alzheimer's / dementia, clinical / translational
[6] doi:10.1038/s41746-026-02803-2 [code]
A clinical neuroimaging platform for rapid, automated lesion detection and personalized post-stroke outcome prediction.
Journal: NPJ digital medicine
In common: pandas, NumPy, stroke, clinical / translational, 1 reference
[7] doi:10.2463/mrms.mp.2024-0149 [code]
Image Distortion Correction for Diffusion MR Imaging Using a Transformer-based U-Net.
Journal: Magnetic resonance in medical sciences : MRMS : an official journal of Japan Society of Magnetic Resonance in Medicine
In common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, clinical / translational
[8] doi:10.1038/s41598-026-47814-x [code]
Predicting post-stroke functional outcome using explainable machine learning and integrated data.
Journal: Scientific reports
In common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, clinical / translational
[9] doi:10.1002/alz.71530 [code]
Differential associations of plasma biomarkers with Alzheimer's disease and small vessel disease: A multimodal imaging study.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: pandas, Matplotlib, NumPy, stroke, Alzheimer's / dementia, clinical / translational
[10] doi:10.1038/s41467-026-76812-w [code]
Assessing molecular, cellular and transcriptomic bases of laminar perfusion and cytoarchitecture coupling in the human cortex.
Journal: Nature communications
In common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.