Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes
The 1 match
- [1] § Material and Methods › BN training › Variable discretization and algorithm ↔ layerbn/bn_utils.py, lines 161–224 · score 0.64 · DiscreteTypeProcessor, pyAgrum, enforce, K2, Bayesian, bins
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 648 lines · 26 KB · MIT · 1 match
- """Bayesian network structure learning, bootstrapping, and scenarios.
- Everything cohort-specific is a parameter: which layers to exclude, which
- layer names mark the outcome and dropout roles, the score, the seed. No
- cohort's conventions are assumed.
- These functions take the layer map, the discretisation and the layer role
- patterns as separate arguments, which means a caller has to keep them
- consistent with the spec at every call site. `layerbn.analysis.Analysis`
- does that for you and is the recommended entry point; use these directly
- when you need control it does not expose.
- """
- from __future__ import annotations
- import random
- from collections import defaultdict
- from collections.abc import Callable, Iterable, Mapping, Sequence
- from itertools import product
- from typing import Any
- import numpy as np
- import pandas as pd
- from layerbn.discretisation import state_for_value
- # Layer-name substrings that identify semantic roles. Override when your
- # cohort names its layers differently.
- DEFAULT_OUTCOME_LAYER_PATTERNS: tuple[str, ...] = ("Outcomes",)
- DEFAULT_DROPOUT_LAYER_PATTERNS: tuple[str, ...] = ("Dropout",)
- def _match_any(layer_name: str, patterns: Sequence[str]) -> bool:
- return any(p in layer_name for p in patterns)
- def collect_outcome_vars(
- layer_map: Mapping[str, Sequence[str]],
- *,
- outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
- dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
- ) -> tuple[list[str], list[str]]:
- """Return `(outcome_vars, dropout_vars)` inferred from layer names."""
- outcome_vars: list[str] = []
- dropout_vars: list[str] = []
- for layer, variables in layer_map.items():
- if _match_any(layer, dropout_patterns):
- dropout_vars.extend(variables)
- elif _match_any(layer, outcome_patterns):
- outcome_vars.extend(variables)
- return outcome_vars, dropout_vars
- def configure_guided_structure(
- learner: Any,
- layer_map_filtered: Mapping[str, Sequence[str]],
- variables_to_keep: Sequence[str],
- outcomes: Sequence[str],
- *,
- dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
- within_layers: bool = True,
- arcs_between_outcomes: bool = False,
- selection_parents: str = "outcomes",
- forbidden_pairs: Iterable[tuple[str, str]] = (),
- mandatory_pairs: Iterable[tuple[str, str]] = (),
- no_parents: Iterable[str] = (),
- no_children: Iterable[str] = (),
- ) -> tuple[list[str], list[str]]:
- """Restrict the learner to arcs that go downward through the layers.
- Rules always encoded:
- * Arcs may only go from an earlier layer to the same or a later layer.
- Rules that the keyword arguments control, whose defaults reproduce the
- behaviour this function had before they existed:
- * `within_layers` — whether two variables in the same layer may be
- connected. False makes each layer internally independent.
- * `arcs_between_outcomes` — whether one outcome may point at another.
- * `selection_parents` — `"outcomes"` allows arcs into a dropout-layer
- variable only from an outcome, on the grounds that dropout follows
- the outcomes rather than the covariates. `"any"` drops that
- restriction and lets the layer order alone decide.
- Additional structural knowledge:
- * `forbidden_pairs` — `(parent, child)` arcs to rule out even though
- the layer order permits them.
- * `mandatory_pairs` — arcs the learner must include. These must be
- consistent with everything above; a pair that is both mandatory and
- forbidden raises `ValueError` rather than reaching pyAgrum, whose own
- message does not say which pair was at fault.
- * `no_parents` / `no_children` — variables pinned as roots or sinks.
- Returns `(dropout_vars, true_outcomes)` for the caller to reuse when
- adding hard outcome→dropout arcs after learning.
- """
- if selection_parents not in ("outcomes", "any"):
- raise ValueError(
- f"selection_parents must be 'outcomes' or 'any', got {selection_parents!r}")
- layer_keys = list(layer_map_filtered.keys())
- dropout_vars = [v for k in layer_keys if _match_any(k, dropout_patterns)
- for v in layer_map_filtered[k]]
- true_outcomes = [v for v in outcomes if v not in dropout_vars]
- keep = set(variables_to_keep)
- forbidden_pairs = set(forbidden_pairs)
- mandatory_pairs = set(mandatory_pairs)
- no_parents = set(no_parents)
- no_children = set(no_children)
- conflict = forbidden_pairs & mandatory_pairs
- if conflict:
- raise ValueError(
- f"these arcs are both mandatory and forbidden: {sorted(conflict)}")
- allowed_arcs: list[tuple[str, str]] = []
- for i, from_key in enumerate(layer_keys):
- for j, to_key in enumerate(layer_keys[i:], start=i):
- if i == j and not within_layers:
- continue
- for parent, child in product(layer_map_filtered[from_key],
- layer_map_filtered[to_key]):
- if child in dropout_vars and selection_parents == "outcomes":
- if parent in outcomes:
- allowed_arcs.append((parent, child))
- continue
- if parent in keep and child in keep:
- allowed_arcs.append((parent, child))
- allowed_set = set(allowed_arcs) - forbidden_pairs
- for parent in variables_to_keep:
- for child in variables_to_keep:
- if parent != child and (parent, child) not in allowed_set:
- learner.addForbiddenArc(parent, child)
- if not arcs_between_outcomes:
- for parent in true_outcomes:
- for child in true_outcomes:
- if parent != child:
- learner.addForbiddenArc(parent, child)
- # Node-level constraints are applied after the pairwise ones so that a
- # variable pinned as a root or a sink stays that way regardless of what
- # the layer order allowed.
- for variable in sorted(no_parents & keep):
- learner.addNoParentNode(variable)
- for variable in sorted(no_children & keep):
- learner.addNoChildrenNode(variable)
- # Mandatory arcs come last: pyAgrum resolves them against the
- # constraints already registered.
- for parent, child in sorted(mandatory_pairs):
- if parent in keep and child in keep:
- learner.addMandatoryArc(parent, child)
- return dropout_vars, true_outcomes
- def build_bn(
- df: pd.DataFrame,
- outcomes: Sequence[str],
- layer_map: Mapping[str, Sequence[str]],
- type_processor: Any,
- *,
- score: str = "K2",
- use_tabu: bool = True,
- random_seed: int = 42,
- enforce_structure: bool = True,
- fixed_template: Any = None,
- use_smoothing: bool = False,
- exclude_layers: Iterable[str] = (),
- max_indegree: int = 5,
- outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
- dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
- within_layers: bool = True,
- arcs_between_outcomes: bool = False,
- selection_parents: str = "outcomes",
- forbidden_pairs: Iterable[tuple[str, str]] = (),
- mandatory_pairs: Iterable[tuple[str, str]] = (),
- no_parents: Iterable[str] = (),
- no_children: Iterable[str] = (),
- ) -> Any:
- """Learn a layered Bayesian network with pyAgrum.
- Parameters
- ----------
- df : DataFrame
- Fully-imputed input data. Columns not needed as outcomes and not
- listed in `exclude_layers` become learnable nodes.
- outcomes : sequence of str
- Outcome variable names to keep in the network. Any *other* variable
- that belongs to an outcome/dropout layer is dropped before learning.
- layer_map : mapping
- `{layer_name: [variable, ...]}` defining the expert layering.
- type_processor : DiscreteTypeProcessor
- Constructed via `layerbn.discretisation.make_type_processor`.
- score : {"K2", "BIC", "BDeu"}
- Structure-learning score.
- fixed_template : optional
- Use a pre-computed discretisation template so bootstrap resamples
- share bin edges. Build once with `type_processor.discretizedTemplate(df)`.
- exclude_layers : iterable of str
- Layer names whose variables should be dropped before learning, for
- instance to leave out a layer measured on only part of the cohort.
- use_smoothing : bool
- Add a Laplace prior, so that a parent configuration nobody in the
- sample happens to occupy gets a small probability rather than zero.
- Always on for BIC, which carries no prior of its own and whose
- zero cells otherwise break inference on bootstrap resamples.
- Ignored for K2 and BDeu: both already contain an implicit prior,
- and adding a second one double-counts it. aGrUM reports that as
- "the K2 score already contains a different 'implicit' prior ...
- the learning will probably be biased", which is exactly what would
- happen, so the request is dropped rather than honoured.
- within_layers, arcs_between_outcomes, selection_parents
- The three structural conventions. See `configure_guided_structure`;
- the defaults reproduce the behaviour that used to be fixed in code.
- forbidden_pairs, mandatory_pairs, no_parents, no_children
- Extra structural knowledge, as `(parent, child)` pairs and variable
- names. `layerbn.spec` builds these from a spec's `constraints`
- block, and `layerbn.analysis.Analysis` passes them in for you.
- Notes
- -----
- * When `enforce_structure=True`, `configure_guided_structure` restricts
- arcs to obey the layer order and any constraints given.
- * When `enforce_structure=False` every constraint is skipped, including
- the layer order. That is a genuinely unconstrained search, useful as a
- comparison but not as the analysis.
- * After learning, hard outcome→dropout arcs are added if not already
- present, so scenario inference always sees the dropout mechanism.
- """
- import pyagrum as gum
- random.seed(random_seed)
- np.random.seed(random_seed)
- exclude_layers = set(exclude_layers)
- layer_map_filtered = {k: v for k, v in layer_map.items() if k not in exclude_layers}
- excluded_vars = [v for k in exclude_layers for v in layer_map.get(k, [])]
- outcome_vars_all, dropout_vars_all = collect_outcome_vars(
- layer_map, outcome_patterns=outcome_patterns, dropout_patterns=dropout_patterns,
- )
- outcome_layer_vars = outcome_vars_all + dropout_vars_all
- drop_outcomes = [v for v in outcome_layer_vars if v not in outcomes]
- df_reduced = df.drop(columns=drop_outcomes + excluded_vars, errors="ignore")
- variables_to_keep = df_reduced.columns.tolist()
- template = fixed_template if fixed_template is not None \
- else type_processor.discretizedTemplate(df_reduced)
- learner = gum.BNLearner(df_reduced, template)
- score = score.upper()
- if score == "K2":
- learner.useScoreK2()
- elif score == "BIC":
- learner.useScoreBIC()
- elif score == "BDEU":
- learner.useScoreBDeu(ess=1)
- else:
- raise ValueError(f"Unsupported score {score!r}: choose K2, BIC, or BDeu")
- # Only BIC has no prior of its own. K2 and BDeu each carry an implicit
- # one, and stacking a smoothing prior on top of it biases the score --
- # aGrUM emits a notification saying so for every fit, which under a
- # bootstrap means hundreds of them. See `use_smoothing` above.
- if score == "BIC":
- learner.useSmoothingPrior()
- if use_tabu:
- learner.useLocalSearchWithTabuList()
- else:
- learner.useGreedyHillClimbing()
- if enforce_structure:
- dropout_vars, true_outcomes = configure_guided_structure(
- learner, layer_map_filtered, variables_to_keep, outcomes,
- dropout_patterns=dropout_patterns,
- within_layers=within_layers,
- arcs_between_outcomes=arcs_between_outcomes,
- selection_parents=selection_parents,
- forbidden_pairs=forbidden_pairs,
- mandatory_pairs=mandatory_pairs,
- no_parents=no_parents,
- no_children=no_children,
- )
- else:
- # Same pair `configure_guided_structure` would have returned, without
- # touching the learner. NB: `collect_outcome_vars` returns
- # (outcome_vars, dropout_vars) — the opposite order — so it must not be
- # unpacked into (dropout_vars, true_outcomes) directly.
- dropout_vars = [v for k in layer_map_filtered
- if _match_any(k, dropout_patterns)
- for v in layer_map_filtered[k]]
- true_outcomes = [v for v in outcomes if v not in dropout_vars]
- learner.setMaxIndegree(max_indegree)
- bn = learner.learnBN()
- if enforce_structure:
- # Connect every outcome to the selection layer, so scenario inference
- # always sees the dropout mechanism even when the score did not pick
- # the arc up. An arc the caller forbade explicitly is left out: the
- # spec's promise is that constraints narrow what the layer order
- # allows, and silently reinstating one here would break it.
- forbidden = set(forbidden_pairs)
- try:
- names = set(bn.names())
- for dropout in (n for n in dropout_vars if n in names):
- for outcome in (n for n in true_outcomes if n in names):
- if (outcome, dropout) in forbidden:
- continue
- out_id, drop_id = bn.idFromName(outcome), bn.idFromName(dropout)
- if not bn.existsArc(out_id, drop_id):
- bn.addArc(out_id, drop_id)
- except Exception:
- pass # gum.InvalidArgument etc. — arc already exists / var missing
- return bn
- def bootstrap_edge_frequencies(
- df: pd.DataFrame,
- *,
- outcomes: Sequence[str],
- layer_map: Mapping[str, Sequence[str]],
- type_processor: Any,
- build_bn_func: Callable[..., Any] = build_bn,
- n_bootstraps: int = 100,
- random_seed: int = 42,
- **build_kwargs: Any,
- ) -> dict[tuple[str, str], float]:
- """Fit `build_bn_func` on `n_bootstraps` resamples; return per-arc frequencies.
- Extra keyword arguments are forwarded to `build_bn_func` unchanged so
- callers can pin `score`, `use_smoothing`, `fixed_template`, etc.
- Notes
- -----
- The denominator is `n_bootstraps`, **not** the number of resamples that
- fitted successfully. A failed fit therefore lowers every frequency rather
- than being excluded from the calculation. These frequencies are reported
- as the `f=NN%` arc labels in published figures, so changing the
- denominator would silently restate them.
- """
- counts: dict[tuple[str, str], int] = defaultdict(int)
- successes = 0
- for i in range(n_bootstraps):
- seed = random_seed + i
- boot = df.sample(frac=1.0, replace=True, random_state=seed)
- try:
- bn = build_bn_func(
- boot, outcomes=outcomes, layer_map=layer_map,
- type_processor=type_processor, random_seed=seed, **build_kwargs,
- )
- except Exception as exc:
- print(f"Bootstrap {i} failed: {exc}")
- continue
- successes += 1
- for parent_id, child_id in bn.arcs():
- counts[(bn.variable(parent_id).name(), bn.variable(child_id).name())] += 1
- if successes < n_bootstraps:
- print(f"Note: {n_bootstraps - successes}/{n_bootstraps} bootstrap fits "
- "failed; frequencies are still divided by n_bootstraps.")
- return {edge: count / n_bootstraps for edge, count in counts.items()}
- def bootstrap_scenario_risks(
- df: pd.DataFrame,
- *,
- scenario_profiles: Sequence[tuple[str, Mapping[str, Any]]],
- target_outcomes: Sequence[tuple[str, str]],
- build_bn_func: Callable[..., Any] = build_bn,
- layer_map: Mapping[str, Sequence[str]],
- type_processor: Any,
- outcomes_for_learning: Sequence[str],
- n_bootstraps: int = 200,
- random_seed: int = 42,
- **build_kwargs: Any,
- ) -> pd.DataFrame:
- """Compute posterior probabilities under a set of patient profiles.
- Parameters
- ----------
- scenario_profiles : sequence of (label, evidence_dict)
- Each dict maps a variable name to the evidence for it, passed to
- `LazyPropagation.setEvidence` unchanged. Prefer **state indices**:
- pyAgrum reads a bare integer as a state index, a numeric *string*
- as a value to place in a bin, and rejects an interval label such as
- `'(45;64.6['` outright. `layerbn.analysis.Analysis.resolve_profile`
- converts any of these to indices and reports what it could not
- match; using it avoids the failure mode below.
- target_outcomes : sequence of (var_name, display_label)
- outcomes_for_learning : sequence of str
- Passed to `build_bn_func` as the fixed set of outcomes so bootstrap
- networks are structurally comparable across resamples.
- Notes
- -----
- Evidence that pyAgrum rejects raises `gum.InvalidArgument`, which this
- function catches and skips. The affected scenario then produces no rows
- at all rather than an error, so check that every profile you passed
- appears in the result.
- Pass `fixed_template` (via `**build_kwargs`) when the profiles refer to
- discretised variables. Without it every resample recomputes its own bin
- edges, and a fixed profile silently refers to a different group of
- people in each resample.
- """
- import pyagrum as gum
- samples: dict[tuple[str, str, str], list[float]] = defaultdict(list)
- label_cache: dict[str, list[str]] = {}
- failed = 0
- for i in range(n_bootstraps):
- seed = random_seed + i
- boot = df.sample(frac=1.0, replace=True, random_state=seed)
- try:
- bn = build_bn_func(
- boot, outcomes=list(outcomes_for_learning), layer_map=layer_map,
- type_processor=type_processor, random_seed=seed, **build_kwargs,
- )
- except Exception:
- failed += 1
- continue
- for scenario_name, evidence in scenario_profiles:
- ie = gum.LazyPropagation(bn)
- try:
- ie.setEvidence(dict(evidence))
- except gum.InvalidArgument:
- continue
- for outcome_name, _ in target_outcomes:
- if outcome_name not in label_cache:
- label_cache[outcome_name] = list(
- bn.variable(bn.idFromName(outcome_name)).labels()
- )
- posterior = ie.posterior(outcome_name)
- for idx, state in enumerate(label_cache[outcome_name]):
- samples[(scenario_name, outcome_name, state)].append(float(posterior[idx]))
- outcome_labels = dict(target_outcomes)
- rows = []
- for (scenario_name, outcome_name, state), values in samples.items():
- arr = np.asarray(values, float)
- arr = arr[np.isfinite(arr)]
- if arr.size == 0:
- continue
- rows.append({
- "Scenario": scenario_name,
- "Outcome": outcome_labels.get(outcome_name, outcome_name),
- "Category": state,
- "Mean probability": arr.mean(),
- "Std": arr.std(ddof=1) if arr.size > 1 else 0.0,
- "CI 2.5%": np.quantile(arr, 0.025),
- "CI 97.5%": np.quantile(arr, 0.975),
- "Bootstraps": int(arr.size),
- })
- if failed:
- print(f"Warning: {failed}/{n_bootstraps} bootstrap fits failed and were skipped.")
- return (pd.DataFrame(rows)
- .sort_values(["Scenario", "Outcome", "Category"])
- .reset_index(drop=True))
- def descendants(bn: Any, start_id: int) -> set[int]:
- """All nodes reachable from `start_id` by following directed arcs."""
- seen: set[int] = set()
- stack = [start_id]
- while stack:
- for child in bn.children(stack.pop()):
- if child not in seen:
- seen.add(child)
- stack.append(child)
- return seen
- def bootstrap_knob_sweep(
- df: pd.DataFrame,
- *,
- base_profile: Mapping[str, Any],
- knob: str,
- outcomes: Sequence[str],
- outcomes_for_learning: Sequence[str],
- layer_map: Mapping[str, Sequence[str]],
- type_processor: Any,
- build_bn_func: Callable[..., Any] = build_bn,
- n_bootstraps: int = 200,
- random_seed: int = 42,
- verbose: bool = True,
- **build_kwargs: Any,
- ) -> tuple[pd.DataFrame, dict[str, Any]]:
- """Sensitivity sweep: fix a patient, vary one variable across its states.
- Returns `(sweep_df, meta)` where `sweep_df` has one row per
- (outcome, knob state, outcome state) with bootstrap mean and 95% CI.
- `meta` carries the resolved knob states and outcome states.
- Notes
- -----
- Warns if any variable in `base_profile` lies on a directed path from
- `knob` to an outcome (overadjustment / blocked mediator). Uses a fixed
- discretisation template so knob states mean the same thing across
- resamples.
- """
- import pyagrum as gum
- fixed_template = type_processor.discretizedTemplate(
- _reduced_for_template(
- df, layer_map, outcomes_for_learning,
- build_kwargs.get("exclude_layers", ()),
- outcome_patterns=build_kwargs.get(
- "outcome_patterns", DEFAULT_OUTCOME_LAYER_PATTERNS),
- dropout_patterns=build_kwargs.get(
- "dropout_patterns", DEFAULT_DROPOUT_LAYER_PATTERNS),
- )
- )
- # `use_smoothing` is deliberately NOT forced on here. It only takes
- # effect under BIC (see `build_bn`), and forcing it meant every sweep
- # resample was learned under a different prior from the network being
- # reported -- so the interval described a procedure nobody ran.
- build_kwargs = {**build_kwargs, "fixed_template": fixed_template}
- ref_bn = build_bn_func(
- df, outcomes=list(outcomes_for_learning), layer_map=layer_map,
- type_processor=type_processor, random_seed=random_seed, **build_kwargs,
- )
- def labels_of(name: str) -> list[str]:
- return list(ref_bn.variable(ref_bn.idFromName(name)).labels())
- if knob in base_profile:
- raise ValueError("base_profile must not contain the knob variable.")
- applied, rejected, snapped = {}, {}, {}
- for var, val in base_profile.items():
- if var not in ref_bn.names():
- rejected[var] = val
- continue
- state = state_for_value(labels_of(var), val)
- if state is None:
- rejected[var] = val
- else:
- applied[var] = state
- if str(val) != state:
- snapped[var] = f"{val} -> {state!r}"
- if verbose:
- print("Fixed patient — evidence applied:", applied or "(none)")
- if snapped:
- print("Fixed patient — snapped to bin:", snapped)
- if rejected:
- print("Fixed patient — REJECTED:", rejected,
- "(unknown variable or value not matchable — check the discretised template).")
- knob_descendants = {
- ref_bn.variable(i).name()
- for i in descendants(ref_bn, ref_bn.idFromName(knob))
- }
- blocked = [v for v in applied if v in knob_descendants]
- if blocked and verbose:
- print(f"WARNING: base_profile fixes {blocked}, which lie downstream of "
- f"{knob!r}. Holding a mediator constant blocks part of the knob's "
- "effect on the outcomes.")
- applied_idx = {v: labels_of(v).index(s) for v, s in applied.items()}
- knob_states = labels_of(knob)
- outcome_states = {o: labels_of(o) for o in outcomes}
- samples: dict[tuple[str, str, str], list[float]] = defaultdict(list)
- failed = 0
- for i in range(n_bootstraps):
- seed = random_seed + i
- boot = df.sample(frac=1.0, replace=True, random_state=seed)
- try:
- bn = build_bn_func(
- boot, outcomes=list(outcomes_for_learning), layer_map=layer_map,
- type_processor=type_processor, random_seed=seed, **build_kwargs,
- )
- except Exception:
- failed += 1
- continue
- for k_idx, k_state in enumerate(knob_states):
- try:
- ie = gum.LazyPropagation(bn)
- ie.setEvidence({**applied_idx, knob: k_idx})
- post = {o: ie.posterior(o) for o in outcomes}
- except Exception:
- continue
- for outcome in outcomes:
- for s_idx, s_label in enumerate(outcome_states[outcome]):
- samples[(outcome, k_state, s_label)].append(float(post[outcome][s_idx]))
- rows = []
- for (outcome, k_state, s_label), values in samples.items():
- arr = np.asarray(values, float)
- arr = arr[np.isfinite(arr)]
- if arr.size == 0:
- continue
- rows.append({
- "Outcome": outcome,
- knob: k_state,
- "Outcome state": s_label,
- "P": arr.mean(),
- "CI_low": np.quantile(arr, 0.025),
- "CI_high": np.quantile(arr, 0.975),
- "Bootstraps": int(arr.size),
- })
- if verbose and failed:
- print(f"Note: {failed}/{n_bootstraps} bootstrap fits failed and were skipped.")
- if not rows:
- raise RuntimeError("No valid scenarios were produced — check base_profile.")
- return pd.DataFrame(rows), {
- "knob": knob, "knob_states": knob_states,
- "outcomes": list(outcomes), "outcome_states": outcome_states,
- }
- def _reduced_for_template(
- df: pd.DataFrame,
- layer_map: Mapping[str, Sequence[str]],
- outcomes_for_learning: Sequence[str],
- exclude_layers: Iterable[str],
- *,
- outcome_patterns: Sequence[str] = DEFAULT_OUTCOME_LAYER_PATTERNS,
- dropout_patterns: Sequence[str] = DEFAULT_DROPOUT_LAYER_PATTERNS,
- ) -> pd.DataFrame:
- """Reproduce `build_bn`'s column reduction so one template covers all bootstraps.
- The patterns must be the same ones `build_bn` will be called with. If they
- are not, this function and `build_bn` disagree about which columns are
- outcomes, the template ends up describing a different set of variables
- than the learner is given, and `BNLearner` raises.
- """
- excluded = [v for k in exclude_layers for v in layer_map.get(k, [])]
- outcome_all, dropout_all = collect_outcome_vars(
- layer_map, outcome_patterns=outcome_patterns, dropout_patterns=dropout_patterns,
- )
- drop_outcomes = [v for v in outcome_all + dropout_all if v not in outcomes_for_learning]
- return df.drop(columns=drop_outcomes + excluded, errors="ignore")
bn_utils.py at commit fd5d083, under MIT · at the source
Overview
- Department of Neurology and Neurosurgery, UMC Utrecht Brain Center, University Medical Center Utrecht, Utrecht, the Netherlands
- Department of Cardiology, Institute for Cardiovascular Research, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
- Department of Cardiology, Radboud University Medical Center, Radboud University, Nijmegen, The Netherlands
- Department of Radiology and Nuclear Medicine, Erasmus Medical Center, Rotterdam, the Netherlands
- Department of Cardiology, Cardiovascular Research Institute Maastricht, Maastricht University Medical Center, Maastricht, the Netherlands
- Department of Radiology, Leiden University Medical Center, Leiden, the Netherlands
- Department of Internal Medicine, Section Geriatrics, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
- Department of Clinical Chemistry, Neurochemistry Laboratory, Amsterdam Neuroscience, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
- Alzheimer Center Amsterdam, Department of Neurology, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, the Netherlands
- Department of Epidemiology, Erasmus MC University Medical Centre, Rotterdam, the Netherlands
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.
umcu/vci-bayes
fd5d083d2d20afde79d57d6c88e5c4e1017eece4, 24 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
22 files
- layerbn/
__init__.py — Python, 38 lines - layerbn/
__main__.py — Python, 139 lines - layerbn/
analysis.py — Python, 767 lines - layerbn/
bn_utils.py — Python, 648 lines, 1 match - layerbn/
config.py — Python, 169 lines - layerbn/
demo.py — Python, 186 lines - layerbn/
discretisation.py — Python, 98 lines - layerbn/
inference.py — Python, 84 lines - layerbn/
plotting.py — Python, 157 lines - layerbn/
preprocess.py — Python, 112 lines - layerbn/
spec.py — Python, 805 lines - layerbn/
templates/ — Jupyter, 341 linesanalysis.ipynb - tests/
__init__.py — Python, 1 line - tests/
conftest.py — Python, 73 lines - tests/
test_cli.py — Python, 93 lines - tests/
test_config.py — Python, 69 lines - tests/
test_learning.py — Python, 201 lines - tests/
test_notebook.py — Python, 94 lines - tests/
test_reproducibility.py — Python, 92 lines - tests/
test_spec.py — Python, 160 lines - LICENSE — License, 21 lines
- README.md — Text, 272 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 20 scripts, each with its path and the digest of its content;
- 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Code and data availability statement
The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: umcu/
vci-bayes - it says that the data are available on request
Read it in the paper: doi.org/10.64898/2026.06.03.26354793.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added L. Malin Overmars (0000-0001-7086-0864); Geert Jan Biessels (0000-0001-6862-2496); removed L. Malin Overmars; Geert Jan Biessels
Version 1, 27 September 2026: the first record
Recorded: type, journal, dates, 11 authors, 5 keywords, 1 funder, 36 references.
Cite
This paper
Overmars, L. M., Allaart, C., Bron, E. E., Brunner La Rocca, H.-P., de Bresser, J., Muller, M., van Osch, M. J., Teunissen, C., Tijms, B. M., Wolters, F. J., & Biessels, G. J. (2026). Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes. medRxiv (preprint). https://
BibTeX
@article{overmars2026bay
author = {Overmars, L. Malin and Allaart, Cor and Bron, Esther E. and Brunner La Rocca, Hans-Peter and de Bresser, Jeroen and Muller, Majon and van Osch, Matthias J.P. and Teunissen, Charlotte and Tijms, Betty M and Wolters, Frank J. and Biessels, Geert Jan},
title = {{Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes}},
journal = {medRxiv (preprint)},
year = {2026},
month = jun,
publisher = {medRxiv},
doi = {10.64898/
url = {https://
}
RIS
TY - JOUR
AU - Overmars, L. Malin
AU - Allaart, Cor
AU - Bron, Esther E.
AU - Brunner La Rocca, Hans-Peter
AU - de Bresser, Jeroen
AU - Muller, Majon
AU - van Osch, Matthias J.P.
AU - Teunissen, Charlotte
AU - Tijms, Betty M
AU - Wolters, Frank J.
AU - Biessels, Geert Jan
TI - Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes
T2 - medRxiv (preprint)
J2 - medRxiv
PY - 2026
DA - 2026/
PB - medRxiv
DO - 10.64898/
UR - https://
ER -
CSL-JSON
{
"id": "10.64898/
"type": "article",
"title": "Bayesian networks to estimate prognosis in vascular cognitive impairment and small vessel disease: integrated analyses of interdependent contributors to multiple outcomes",
"container-title": "medRxiv (preprint)",
"author": [
{
"family": "Overmars",
"given": "L. Malin"
},
{
"family": "Allaart",
"given": "Cor"
},
{
"family": "Bron",
"given": "Esther E."
},
{
"family": "Brunner La Rocca",
"given": "Hans-Peter"
},
{
"family": "de Bresser",
"given": "Jeroen"
},
{
"family": "Muller",
"given": "Majon"
},
{
"family": "van Osch",
"given": "Matthias J.P."
},
{
"family": "Teunissen",
"given": "Charlotte"
},
{
"family": "Tijms",
"given": "Betty M"
},
{
"family": "Wolters",
"given": "Frank J."
},
{
"family": "Biessels",
"given": "Geert Jan"
}
],
"container-title-short":
"DOI": "10.64898/
"publisher": "medRxiv",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
6,
4
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1093/braincomms/fcag257 [code]
- Neuroplasticity and immune system are related to altered grey matter networks: a cohort study in sporadic Alzheimer's disease.Journal: Brain communicationsIn common: Alzheimer's / dementia, 2 authors
- [2] doi:10.1038/s41598-026-47047-y [code]
- Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study.Journal: Scientific reportsIn common: scikit-learn, pandas, Matplotlib, 1 other tool, Alzheimer's / dementia, clinical / translational, author Frank J. Wolters
- [3] doi:10.1093/braincomms/fcag196 [code]
- Music processing in behavioural variant frontotemporal dementia and Alzheimer's disease: a functional MRI study.Journal: Brain communicationsIn common: Alzheimer's / dementia, author Betty M Tijms
- [4] doi:10.1212/wnl.0000000000218472 [code]
- Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.Journal: NeurologyIn common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, Alzheimer's / dementia, clinical / translational
- [5] doi:10.1038/s41746-026-02803-2 [code]
- A clinical neuroimaging platform for rapid, automated lesion detection and personalized post-stroke outcome prediction.Journal: NPJ digital medicineIn common: pandas, NumPy, stroke, clinical / translational, 1 reference
- [6] doi:10.2463/mrms.mp.2024-0149 [code]
- Image Distortion Correction for Diffusion MR Imaging Using a Transformer-based U-Net.Journal: Magnetic resonance in medical sciences : MRMS : an official journal of Japan Society of Magnetic Resonance in MedicineIn common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, clinical / translational
- [7] doi:10.1038/s41598-026-47814-x [code]
- Predicting post-stroke functional outcome using explainable machine learning and integrated data.Journal: Scientific reportsIn common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke, clinical / translational
- [8] doi:10.1002/alz.71530 [code]
- Differential associations of plasma biomarkers with Alzheimer's disease and small vessel disease: A multimodal imaging study.Journal: Alzheimer's & dementia : the journal of the Alzheimer's AssociationIn common: pandas, Matplotlib, NumPy, stroke, Alzheimer's / dementia, clinical / translational
- [9] doi:10.1038/s41467-026-76812-w [code]
- Assessing molecular, cellular and transcriptomic bases of laminar perfusion and cytoarchitecture coupling in the human cortex.Journal: Nature communicationsIn common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke
- [10] doi:10.1016/j.nicl.2026.104001 [code]
- Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation.Journal: NeuroImage. ClinicalIn common: scikit-learn, pandas, Matplotlib, 1 other tool, stroke
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 20 scripts, and 1 match between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:80a72bacb831381a…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
