Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation.
The 28 matches
- [1] § Methods › Statistical analysis ↔ Phase_2_a/1_2a_bland_altman_test_temp.py, lines 919–984 · score 0.94 · 0.75–0.9, Intraclass Correlation Coefficient, Bland Altman, 0.5–0.75, equivalence margin, way mixed
- [2] § Methods › LOCATE adaptive thresholding ↔ 1_preprare_locate.py, lines 4–85 · score 0.91 · unbiased lesion probability, LOCally Adaptive Threshold, fold cross validation, LOCATE training, training subjects, leakage
- [3] § Methods › Lesion delineation, removal, and inpainting ↔ 6_train_bianca.py, lines 1–59 · score 0.85 · NAWM intensities, lesion_filling, FLAIR images, lesion voxel, intensity distribution, FSL
- [4] § Results › Phase II-B: Volume-based agreement (n = 211) ↔ Phase_2_b/2_2b_volume_difference_lesion_bonf.py, lines 1–61 · score 0.81 · 0.14 mL, 0.07 mL, 0.29 mL, 0.06 mL, marginal, 0.04 mL
- [5] § Methods › Statistical analysis ↔ Phase_2_a/5_2a_shapley_values.py, lines 1–59 · score 0.79 · TreeExplainer, max depth, Random Forest, bag, SHAP, trees
- [6] § Methods › Phase I: Parameter optimization › Cluster-based analysis ↔ 4_Analyze_cluster_metrics.py, lines 43–116 · score 0.77 · ground truth masks, Minimum cluster, Dice coefficient, complementing, recall, detection
- [7] § Methods › Preprocessing pipeline ↔ 1_preprare_locate.py, lines 4–85 · score 0.75 · ventricle distance map, white matter mask, adaptive thresholding, pipeline, LOCATE, segmentation
- [8] § Methods › Statistical analysis ↔ Phase_2_b/5_2b_shapley_values.py, lines 1–63 · score 0.75 · TreeExplainer, max depth, Random Forest, SHAP, trees, MAE
- [9] § Results › Population characteristics ↔ Phase_2_b/0_2b_population.py, lines 1–79 · score 0.74 · Prisma fit, 3–10, Scanner distribution, BeLOVE, lesion volume, ARWMC
- [10] § Methods › Datasets and scanner distribution ↔ Phase_2_a/0_2a_population.py, lines 1–85 · score 0.72 · full cohort, ground truth WMH, Prisma fit, WMH masks, BeLOVE, Trio
- [11] § Methods › Lesion delineation, removal, and inpainting ↔ 3_threshold_analysis_parallel_temp.py, lines 1–78 · score 0.70 · lesion_filling, lesion voxel, zero filled, FSL, physiologically, NAWM
- [12] § Methods › Phase I: Parameter optimization › Cluster-based analysis ↔ 4_plot_cluster_grid_search.py, lines 1–48 · score 0.69 · Optimal minimum cluster, ground truth masks, recall, filtering, predicted, voxel
- [13] § Methods › Study design and datasets ↔ 3_threshold_analysis_parallel_temp.py, lines 1–78 · score 0.69 · lesion_filling, lesion voxels, zero filled, cross validation, FSL, NAWM
- [14] § Results › Population characteristics ↔ Phase_2_a/0_2a_population.py, lines 1–85 · score 0.69 · LOCATE training pool, ground truth WMH, Prisma fit, scanner balanced, BeLOVE, Population
- [15] § Results › Phase II-B: Volume-based agreement (n = 211) ↔ Phase_2_b/3_2b_volume_difference_scanner_bonf.py, lines 1–59 · score 0.69 · Prisma fit, 14.67 mL, 15.58 mL, largest, partly, baseline
- [16] § Methods › Study design and datasets ↔ 5_create_5_fold_cross_validation_sets.py, lines 1–63 · score 0.67 · lesion_filling, cross validation, zero filled, FSL, NAWM, CV
- [17] § Methods › Preprocessing pipeline › GE scanner exclusion ↔ Phase_1/GE_COMPARE/2_GE_analysis.py, lines 1–86 · score 0.64 · GE inclusion, GE subjects, medium, Philips, stratified, negligible
- [18] § Results › Phase II-A: preprocessing effect assessment (n = 89) ↔ Phase_2_a/5_2a_shapley_values.py, lines 61–120 · score 0.64 · profiles, SHAP, explanatory, Forest, lesion volume, MAE
- [19] § Methods › Phase I: Parameter optimization › Training condition comparison ↔ 4_Analyze_cluster_metrics.py, lines 407–555 · score 0.64 · preprocessing affects, Wilcoxon signed rank, training preprocessing, bootstrapped, Bonferroni, Cliff
- [20] § Methods › Cross-validation strategy ↔ 4_cluster_size_grid_search.py, lines 1–84 · score 0.63 · LOO CV, cross validation, leakage, mixed, seeds, stratified
- [21] § Results › Phase I: parameter optimization (n = 89) › Threshold analysis ↔ Phase_1/5FCV_SET/1_voxel_treshhold_analysis.py, lines 1–61 · score 0.63 · disadvantage, b0, medium, adaptive, Pairwise, Wilcoxon
- [22] § Methods › Preprocessing pipeline ↔ Phase_2_b/0_2b_population.py, lines 1–79 · score 0.63 · HD BET, fsl_anat, Brain, mask, segmentation, Preprocessing
- [23] § Methods › Cross-validation strategy ↔ 5_create_5_fold_cross_validation_sets.py, lines 1–63 · score 0.62 · fold cross validation, LOO CV, seeds, stratified, severity, scanner
- [24] § Results › Phase II-A: preprocessing effect assessment (n = 89) ↔ Phase_2_a/1_2a_bland_altman_test_temp.py, lines 1–64 · score 0.60 · variance, attributable, TOST, opposing, marginally, shifts
- [25] § Results › Population characteristics ↔ Phase_1/1_phase_1_belove_k_bianca_population.py, lines 62–125 · score 0.60 · Prisma fit, ground truth WMH, 6.96 mL, BeLOVE, cutoffs, IQR
- [26] § Results › Phase I: parameter optimization (n = 89) › Cluster-based analysis ↔ 4_Analyze_cluster_metrics.py, lines 43–116 · score 0.58 · complemented voxel, minimum cluster, detection, precision, Cliff, metrics
- [27] § Methods › Statistical analysis ↔ 4_Analyze_cluster_metrics.py, lines 670–767 · score 0.54 · Mann Whitney, Wilcoxon signed rank, BeLOVE, Cliff, Delta
- [28] § Methods › Lesion delineation, removal, and inpainting ↔ Phase_2_a/8_2a_correlation_dice_2.py, lines 203–232 · score 0.52 · intracranial hemorrhage, ischemic infarcts, lacunes, inpainting, lesions
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 771 lines · 33 KB · no license · 4 matches
- #!/usr/bin/env python3
- # -*- coding: utf-8 -*-
- """
- Cluster-Level Analysis of BIANCA WMH Segmentation
- ===================================================
- @author: temuuleu
- # =============================================================
- # REVIEWER RESPONSE (R1, Comment 2; R5, #17)
- # =============================================================
- #
- # R1 raised the concern that global Dice alone cannot distinguish
- # genuine sensitivity gains from spatially distributed false positives.
- # R5 (#17) noted that precision values were not contextualized.
- #
- # This script produces all cluster-level analyses for the manuscript:
- #
- # MAIN TEXT:
- # 1. Voxel vs Cluster dissociation (Figure: Dice vs Cluster F1 scatter)
- # 2. Cluster Precision vs Recall (Figure: bar chart)
- # 3. R vs NR condition comparison (Figure: boxplot + Kruskal-Wallis)
- # 4. Cliff's Delta for R vs NR on cluster-level
- #
- # SUPPLEMENTAL:
- # 5. Post-hoc Wilcoxon with Bonferroni correction
- # 6. Per-dataset breakdown (BeLOVE vs Challenge, Mann-Whitney U)
- # 7. GT cluster count distribution
- # 8. Summary tables (mean, SD, median per condition x threshold)
- #
- # NOT INCLUDED (not requested by reviewers):
- # - Per-scanner stratification (already covered in voxel-level analysis)
- # - Per-severity stratification (already covered in voxel-level analysis)
- # - Full 9x3 condition matrix (identical values, no added information)
- #
- # Methodology:
- # - Connected-component labeling: 26-connectivity
- # - Minimum cluster size: threshold- and dataset-specific (grid search)
- # - Overlap criterion: any-overlap (>=1 voxel), MICCAI 2017
- # - Filtering: prediction only; GT retained without filtering
- # - Evaluated: 10 seeds x 5-fold stratified CV x 89 subjects
- # =============================================================
- # =============================================================
- # PAPER TEXT: METHODS
- # =============================================================
- #
- # 2.X Lesion-level evaluation
- #
- # To complement voxel-level Dice coefficients, we computed lesion-level
- # precision, recall, and F1 score using connected-component analysis.
- # Binary segmentation masks were decomposed into spatially contiguous
- # clusters using 26-connectivity (scipy.ndimage.label). Predicted
- # clusters smaller than a threshold-specific minimum cluster size
- # (determined via grid search; Supplemental Table X) were removed to
- # exclude spurious components; ground truth masks were not filtered.
- # Consistent with the MICCAI 2017 WMH Segmentation Challenge (Kuijf
- # et al., 2019), a predicted cluster was classified as a true positive
- # if it overlapped with any ground truth cluster by at least one voxel
- # (any-overlap criterion). Lesion-level precision was defined as the
- # proportion of predicted clusters classified as true positives, and
- # lesion-level recall as the proportion of ground truth clusters
- # detected. Lesion-level F1 was computed as the harmonic mean of
- # precision and recall.
- #
- # To assess the effect of training preprocessing on cluster-level
- # detection, Kruskal-Wallis H-tests were performed for each threshold
- # and metric. Where the omnibus test was significant (alpha = 0.05),
- # post-hoc pairwise Wilcoxon signed-rank tests were conducted with
- # Bonferroni correction (corrected alpha = 0.05/3 = 0.0167). Effect
- # sizes were quantified using Cliff's Delta, with |delta| >= 0.28 as
- # the threshold for a meaningful difference.
- # =============================================================
- # =============================================================
- # PAPER TEXT: RESULTS
- # =============================================================
- #
- # 3.X Lesion-level evaluation
- #
- # Cluster-level F1 scores (0.855-0.877) substantially exceeded
- # voxel-level Dice coefficients (0.540-0.567), indicating that BIANCA
- # detected the majority of WMH clusters but delineated their boundaries
- # imprecisely (Supplemental Figure X). Lesion-level precision was
- # consistently high across all thresholding approaches (0.881-0.928),
- # confirming that the majority of predicted clusters corresponded to
- # genuine WMH regions. Lesion-level recall ranged from 0.849 to 0.905,
- # with B+L achieving the highest detection rate.
- #
- # Kruskal-Wallis tests revealed no significant differences between
- # training preprocessing conditions on any cluster-level metric
- # (all p > 0.92; Supplemental Table X), confirming that lesion removal
- # does not affect lesion detection performance. This finding, combined
- # with the high and stable cluster-level precision (0.881-0.928),
- # addresses the concern that higher WMH volumes after lesion removal
- # might reflect spatially distributed false positives rather than
- # genuine sensitivity gains.
- # =============================================================
- """
- import matplotlib
- matplotlib.use('Agg')
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import seaborn as sns
- import os
- import glob
- from scipy.stats import wilcoxon, mannwhitneyu, kruskal
- try:
- from cliffs_delta import cliffs_delta
- HAS_CLIFFS = True
- except ImportError:
- HAS_CLIFFS = False
- print("WARNING: cliffs_delta not installed. Install with: pip install cliffs_delta")
- # =============================================================
- # CONFIG
- # =============================================================
- SEED_DIR = "Phase_1/5FCV_SET/bianca_result_dir"
- OUTPUT_DIR = "Phase_1/Cluster_analysis_results"
- PLOT_DIR = os.path.join(OUTPUT_DIR, "plots")
- # Labels for plots (LaTeX subscript) vs Excel (plain text)
- TH_PLOT = {'85': r'B$_{0.85}$', '90': r'B$_{0.90}$', 'locate': 'B+L'}
- TH_PLAIN = {'85': 'B_0.85', '90': 'B_0.90', 'locate': 'B+L'}
- COND_LABELS = {
- 'non_removed': 'Train Non Removed',
- 'removed': 'Train Removed',
- 'filled': 'Train Inpainted',
- }
- # All metrics: voxel-level + cluster-level
- VOXEL_METRICS = ['dice_score', 'sensitivity', 'precision']
- CLUSTER_METRICS = ['lesion_f1', 'lesion_precision', 'lesion_recall']
- ALL_METRICS = VOXEL_METRICS + CLUSTER_METRICS
- # Bootstrap CI parameters (consistent with manuscript: Efron & Tibshirani, 1993)
- N_BOOTSTRAP = 1000
- CI_ALPHA = 0.05 # 95% CI
- RANDOM_SEED = 42
- os.makedirs(OUTPUT_DIR, exist_ok=True)
- os.makedirs(PLOT_DIR, exist_ok=True)
- # =============================================================
- # HELPER: Excel export with metadata
- # =============================================================
- def save_excel(filepath, dataframe, title, description):
- """Save DataFrame to Excel with descriptive metadata header."""
- from openpyxl import Workbook
- from openpyxl.styles import Font, PatternFill, Alignment, Border, Side
- from openpyxl.utils import get_column_letter
- wb = Workbook()
- ws = wb.active
- n_cols = len(dataframe.columns)
- merge_width = max(n_cols, 6)
- thin = Border(*(Side(style='thin', color='B0B0B0'),) * 4)
- # Title
- row = 1
- ws.merge_cells(start_row=row, start_column=1, end_row=row, end_column=merge_width)
- c = ws.cell(row=row, column=1, value=title)
- c.font = Font(name='Arial', bold=True, size=11, color='FFFFFF')
- c.fill = PatternFill('solid', fgColor='2F5496')
- c.alignment = Alignment(horizontal='left', vertical='center', wrap_text=True)
- # Description
- for line in description:
- row += 1
- ws.merge_cells(start_row=row, start_column=1, end_row=row, end_column=merge_width)
- c = ws.cell(row=row, column=1, value=line)
- c.font = Font(name='Arial', size=9, italic=True, color='333333')
- c.fill = PatternFill('solid', fgColor='F2F2F2')
- c.alignment = Alignment(horizontal='left', vertical='center', wrap_text=True)
- row += 2 # empty separator
- # Headers
- for ci, col in enumerate(dataframe.columns, 1):
- c = ws.cell(row=row, column=ci, value=col)
- c.font = Font(name='Arial', bold=True, size=10)
- c.fill = PatternFill('solid', fgColor='D6E4F0')
- c.alignment = Alignment(horizontal='center', vertical='center')
- c.border = thin
- # Data
- for _, data_row in dataframe.iterrows():
- row += 1
- for ci, val in enumerate(data_row, 1):
- c = ws.cell(row=row, column=ci)
- if isinstance(val, (np.bool_, bool)):
- c.value = str(val)
- elif isinstance(val, float) and not np.isnan(val):
- c.value = val
- c.number_format = '0.0000'
- else:
- c.value = val
- c.font = Font(name='Arial', size=10)
- c.alignment = Alignment(horizontal='center')
- c.border = thin
- # Auto-width
- for ci in range(1, n_cols + 1):
- max_len = max(len(str(dataframe.columns[ci-1])),
- dataframe.iloc[:, ci-1].astype(str).str.len().max() if len(dataframe) > 0 else 0)
- ws.column_dimensions[get_column_letter(ci)].width = min(max_len + 3, 35)
- wb.save(filepath)
- # =============================================================
- # HELPER: Bootstrap 95% confidence intervals
- # =============================================================
- def bootstrap_ci(values, n_boot=N_BOOTSTRAP, alpha=CI_ALPHA, seed=RANDOM_SEED):
- """
- Compute bootstrap 95% CI for the mean.
- Same-size resampling with replacement, 1000 iterations,
- consistent with manuscript methodology (Efron & Tibshirani, 1993).
- Returns (mean, ci_lower, ci_upper)
- """
- rng = np.random.default_rng(seed)
- values = np.asarray(values)
- n = len(values)
- boot_means = np.array([
- rng.choice(values, size=n, replace=True).mean()
- for _ in range(n_boot)
- ])
- ci_lo = np.percentile(boot_means, 100 * alpha / 2)
- ci_hi = np.percentile(boot_means, 100 * (1 - alpha / 2))
- return round(values.mean(), 4), round(ci_lo, 4), round(ci_hi, 4)
- def fmt_ci(mean, lo, hi):
- """Format as 'mean [lo, hi]' for tables."""
- return f"{mean:.3f} [{lo:.3f}, {hi:.3f}]"
- # =============================================================
- # 1. LOAD DATA
- # =============================================================
- def load_data(seed_dir=SEED_DIR):
- files = sorted(glob.glob(os.path.join(seed_dir, "bianca_metrics_seed_*.xlsx")))
- if not files:
- files = sorted(glob.glob("bianca_metrics_seed_*.xlsx"))
- if not files:
- raise FileNotFoundError("No bianca_metrics_seed_*.xlsx files found.")
- df = pd.concat([pd.read_excel(f) for f in files], ignore_index=True)
- df['threshold'] = df['threshold'].astype(str)
- df['dataset'] = df['subject'].apply(
- lambda x: 'BeLOVE' if x.startswith('belove') else 'Challenge')
- print(f"Loaded {len(files)} seeds: {len(df)} rows, "
- f"{df['subject'].nunique()} subjects")
- return df
- # =============================================================
- # 2. VOXEL vs CLUSTER DISSOCIATION (Main Text Figure + Table)
- # =============================================================
- def voxel_vs_cluster(df):
- """
- Core analysis addressing R1 Comment 2:
- Cluster F1 >> Dice demonstrates that BIANCA detects most lesions
- but delineates boundaries imprecisely. High cluster precision
- confirms that detected regions are genuine WMH, not distributed FP.
- Reports both voxel-level (Dice, sensitivity, precision) and
- cluster-level (F1, precision, recall) with bootstrap 95% CIs.
- """
- filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
- per_sub = filled.groupby(['subject', 'threshold', 'dataset']).agg(
- {m: 'mean' for m in ALL_METRICS}).reset_index()
- # Summary table with bootstrap CIs
- rows = []
- print("\n" + "=" * 70)
- print("VOXEL vs CLUSTER PERFORMANCE (train=filled, test=filled)")
- print("Bootstrap 95% CI (1000 iterations)")
- print("=" * 70)
- for th in ['85', '90', 'locate']:
- sub = per_sub[per_sub['threshold'] == th]
- row = {'threshold': TH_PLAIN[th]}
- # Voxel-level metrics with CIs
- for m, label in [('dice_score', 'Dice'),
- ('sensitivity', 'Vox Sens'),
- ('precision', 'Vox Prec')]:
- mean, lo, hi = bootstrap_ci(sub[m].values)
- row[f'{m}_mean'] = mean
- row[f'{m}_sd'] = round(sub[m].std(), 4)
- row[f'{m}_ci_lower'] = lo
- row[f'{m}_ci_upper'] = hi
- # Cluster-level metrics with CIs
- for m, label in [('lesion_f1', 'Clust F1'),
- ('lesion_precision', 'Clust P'),
- ('lesion_recall', 'Clust R')]:
- mean, lo, hi = bootstrap_ci(sub[m].values)
- row[f'{m}_mean'] = mean
- row[f'{m}_sd'] = round(sub[m].std(), 4)
- row[f'{m}_ci_lower'] = lo
- row[f'{m}_ci_upper'] = hi
- row['f1_minus_dice'] = round(
- row['lesion_f1_mean'] - row['dice_score_mean'], 4)
- rows.append(row)
- print(f"\n {TH_PLOT[th]}:")
- print(f" Voxel: Dice={fmt_ci(row['dice_score_mean'], row['dice_score_ci_lower'], row['dice_score_ci_upper'])} "
- f"Sens={fmt_ci(row['sensitivity_mean'], row['sensitivity_ci_lower'], row['sensitivity_ci_upper'])} "
- f"Prec={fmt_ci(row['precision_mean'], row['precision_ci_lower'], row['precision_ci_upper'])}")
- print(f" Cluster: F1={fmt_ci(row['lesion_f1_mean'], row['lesion_f1_ci_lower'], row['lesion_f1_ci_upper'])} "
- f"P={fmt_ci(row['lesion_precision_mean'], row['lesion_precision_ci_lower'], row['lesion_precision_ci_upper'])} "
- f"R={fmt_ci(row['lesion_recall_mean'], row['lesion_recall_ci_lower'], row['lesion_recall_ci_upper'])}")
- print(f" Gap (cF1 - Dice) = {row['f1_minus_dice']:+.3f}")
- # Scatter plot: Dice vs Cluster F1
- colors = {'85': '#1f77b4', '90': '#ff7f0e', 'locate': '#2ca02c'}
- fig, ax = plt.subplots(figsize=(8, 7))
- for th in ['85', '90', 'locate']:
- sub = per_sub[per_sub['threshold'] == th]
- ax.scatter(sub['dice_score'], sub['lesion_f1'], alpha=0.5, s=30,
- color=colors[th], label=TH_PLOT[th])
- ax.plot([0, 1], [0, 1], 'k--', alpha=0.3, lw=1)
- ax.set_xlabel('Voxel-level Dice', fontsize=12)
- ax.set_ylabel('Cluster-level F1', fontsize=12)
- ax.legend(fontsize=10)
- ax.grid(True, alpha=0.2)
- ax.set_xlim(0, 1); ax.set_ylim(0, 1)
- fig.tight_layout()
- fig.savefig(os.path.join(PLOT_DIR, "dice_vs_cluster_f1_scatter.png"), dpi=300)
- plt.close()
- print(" Saved: dice_vs_cluster_f1_scatter.png")
- return pd.DataFrame(rows)
- # =============================================================
- # 3. PRECISION vs RECALL: Voxel + Cluster (Main Text Figure)
- # =============================================================
- def precision_recall_bars(df):
- """
- Addresses R5 #17: Precision values contextualized.
- Shows voxel-level (Dice, Sens, Prec) vs cluster-level (F1, P, R)
- side by side, demonstrating that moderate voxel precision coexists
- with high cluster precision.
- """
- filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
- thresholds = ['85', '90', 'locate']
- fig, axes = plt.subplots(1, 2, figsize=(14, 5.5))
- # Panel A: Voxel-level
- vox_metrics = [('dice_score', 'Dice'), ('sensitivity', 'Sensitivity'), ('precision', 'Precision')]
- x = np.arange(len(thresholds))
- w = 0.25
- vox_colors = ['#5B9BD5', '#70AD47', '#ED7D31']
- for i, (m, label) in enumerate(vox_metrics):
- means = [filled[filled['threshold'] == t][m].mean() for t in thresholds]
- sds = [filled[filled['threshold'] == t][m].std() for t in thresholds]
- bars = axes[0].bar(x + i * w - w, means, w, yerr=sds, label=label,
- color=vox_colors[i], capsize=3, alpha=0.85)
- for bar in bars:
- axes[0].text(bar.get_x() + bar.get_width()/2., bar.get_height() + 0.02,
- f'{bar.get_height():.2f}', ha='center', va='bottom', fontsize=8)
- axes[0].set_xticks(x)
- axes[0].set_xticklabels([TH_PLOT[t] for t in thresholds], fontsize=11)
- axes[0].set_ylabel('Score', fontsize=11)
- axes[0].set_title('Voxel-level', fontsize=12, fontweight='bold')
- axes[0].legend(fontsize=9)
- axes[0].set_ylim(0, 1.15)
- axes[0].grid(True, alpha=0.2, axis='y')
- # Panel B: Cluster-level
- clu_metrics = [('lesion_f1', 'F1'), ('lesion_precision', 'Precision'), ('lesion_recall', 'Recall')]
- clu_colors = ['#5B9BD5', '#70AD47', '#ED7D31']
- for i, (m, label) in enumerate(clu_metrics):
- means = [filled[filled['threshold'] == t][m].mean() for t in thresholds]
- sds = [filled[filled['threshold'] == t][m].std() for t in thresholds]
- bars = axes[1].bar(x + i * w - w, means, w, yerr=sds, label=label,
- color=clu_colors[i], capsize=3, alpha=0.85)
- for bar in bars:
- axes[1].text(bar.get_x() + bar.get_width()/2., bar.get_height() + 0.02,
- f'{bar.get_height():.2f}', ha='center', va='bottom', fontsize=8)
- axes[1].set_xticks(x)
- axes[1].set_xticklabels([TH_PLOT[t] for t in thresholds], fontsize=11)
- axes[1].set_title('Cluster-level', fontsize=12, fontweight='bold')
- axes[1].legend(fontsize=9)
- axes[1].set_ylim(0, 1.15)
- axes[1].grid(True, alpha=0.2, axis='y')
- fig.tight_layout()
- fig.savefig(os.path.join(PLOT_DIR, "voxel_vs_cluster_bars.png"), dpi=300)
- plt.close()
- print(" Saved: voxel_vs_cluster_bars.png")
- # =============================================================
- # 4. CONDITION COMPARISON: Kruskal-Wallis + Wilcoxon + Cliff's Delta
- # (Main Text Results + Supplemental Tables)
- # =============================================================
- def condition_comparison(df):
- """
- Core analysis: Does training preprocessing affect cluster-level detection?
- 1. Kruskal-Wallis H-test (omnibus, 3 training conditions)
- 2. Post-hoc Wilcoxon signed-rank (Bonferroni alpha = 0.05/3)
- 3. Cliff's Delta (|delta| >= 0.28 for meaningful difference)
- Test condition fixed to 'filled' (inpainted).
- Per-subject means averaged across seeds/folds for independence.
- """
- COMPARISONS = [
- ('non_removed', 'removed', 'Train Non Removed vs Train Removed'),
- ('non_removed', 'filled', 'Train Non Removed vs Train Inpainted'),
- ('removed', 'filled', 'Train Removed vs Train Inpainted'),
- ]
- BONF_ALPHA = 0.05 / len(COMPARISONS)
- train_conds = ['non_removed', 'removed', 'filled']
- omnibus_rows = []
- posthoc_rows = []
- print("\n" + "=" * 70)
- print("TRAIN CONDITION COMPARISON (cluster-level, test=filled)")
- print(f"Omnibus: Kruskal-Wallis | Post-hoc: Wilcoxon (Bonferroni alpha={BONF_ALPHA:.4f})")
- print("=" * 70)
- for th in ['85', '90', 'locate']:
- # Per-subject averages per condition
- cond_data = {}
- for tc in train_conds:
- cond_data[tc] = df[
- (df['train_condition'] == tc) &
- (df['test_condition'] == 'filled') &
- (df['threshold'] == th)
- ].groupby("subject")[ALL_METRICS].mean()
- print(f"\n --- {TH_PLOT[th]} ---")
- for metric in ALL_METRICS:
- groups = [cond_data[tc][metric].values for tc in train_conds]
- try:
- h_stat, kw_p = kruskal(*groups)
- except ValueError:
- h_stat, kw_p = 0.0, 1.0
- omnibus_rows.append({
- 'threshold': TH_PLAIN[th],
- 'metric': metric,
- 'groups': 'Non Removed, Removed, Inpainted',
- 'H_statistic': round(h_stat, 4),
- 'p_kruskal_wallis': round(kw_p, 4),
- 'significant_alpha_0.05': kw_p < 0.05,
- 'n_per_group': len(groups[0]),
- })
- if metric == 'lesion_f1':
- sig = "SIGNIFICANT" if kw_p < 0.05 else "n.s."
- print(f" Kruskal-Wallis (F1): H={h_stat:.4f}, p={kw_p:.4f} ({sig})")
- # Post-hoc pairwise
- for cond_a, cond_b, label in COMPARISONS:
- merged = cond_data[cond_a].merge(
- cond_data[cond_b], on='subject', suffixes=('_a', '_b'))
- vals_a = merged[f'{metric}_a'].values
- vals_b = merged[f'{metric}_b'].values
- try:
- _, p_w = wilcoxon(vals_a, vals_b)
- except ValueError:
- p_w = 1.0
- d_val, d_size = (np.nan, 'N/A')
- if HAS_CLIFFS:
- d_val, d_size = cliffs_delta(vals_b, vals_a)
- # Bootstrap CIs for each condition mean
- m_a, lo_a, hi_a = bootstrap_ci(vals_a)
- m_b, lo_b, hi_b = bootstrap_ci(vals_b)
- diff_vals = vals_b - vals_a
- m_d, lo_d, hi_d = bootstrap_ci(diff_vals)
- posthoc_rows.append({
- 'threshold': TH_PLAIN[th],
- 'comparison': label,
- 'metric': metric,
- 'mean_a': m_a,
- 'ci_lower_a': lo_a,
- 'ci_upper_a': hi_a,
- 'mean_b': m_b,
- 'ci_lower_b': lo_b,
- 'ci_upper_b': hi_b,
- 'mean_diff': m_d,
- 'diff_ci_lower': lo_d,
- 'diff_ci_upper': hi_d,
- 'kruskal_wallis_p': round(kw_p, 4),
- 'omnibus_significant': kw_p < 0.05,
- 'p_wilcoxon': round(p_w, 4),
- 'bonferroni_alpha': round(BONF_ALPHA, 4),
- 'significant_bonferroni': (kw_p < 0.05) and (p_w < BONF_ALPHA),
- 'cliffs_delta': round(d_val, 4) if not np.isnan(d_val) else np.nan,
- 'effect_size_category': d_size,
- 'n_subjects': len(merged),
- })
- # Print post-hoc F1 summary
- for cond_a, cond_b, label in COMPARISONS:
- merged = cond_data[cond_a].merge(
- cond_data[cond_b], on='subject', suffixes=('_a', '_b'))
- f1_diff = merged['lesion_f1_b'].mean() - merged['lesion_f1_a'].mean()
- try:
- _, pw = wilcoxon(merged['lesion_f1_a'].values, merged['lesion_f1_b'].values)
- except ValueError:
- pw = 1.0
- d_str = ""
- if HAS_CLIFFS:
- d, s = cliffs_delta(merged['lesion_f1_b'].values, merged['lesion_f1_a'].values)
- d_str = f" delta={d:.4f} ({s})"
- print(f" {label}: dF1={f1_diff:+.4f} p={pw:.4f}{d_str}")
- # Boxplot (Main Text Figure)
- filled_test = df[df['test_condition'] == 'filled']
- per_sub = filled_test.groupby(
- ['subject', 'train_condition', 'threshold']).agg({'lesion_f1': 'mean'}).reset_index()
- per_sub['train_label'] = per_sub['train_condition'].map(COND_LABELS)
- fig, axes = plt.subplots(1, 3, figsize=(15, 5), sharey=True)
- order = list(COND_LABELS.values())
- for i, th in enumerate(['85', '90', 'locate']):
- sub = per_sub[per_sub['threshold'] == th]
- sns.boxplot(data=sub, x='train_label', y='lesion_f1', order=order,
- hue='train_label', hue_order=order, legend=False,
- ax=axes[i], palette='Set2', width=0.6)
- axes[i].set_title(TH_PLOT[th], fontsize=12, fontweight='bold')
- axes[i].set_xlabel('')
- axes[i].set_ylim(0, 1.05)
- axes[i].grid(True, alpha=0.2, axis='y')
- axes[i].tick_params(axis='x', rotation=15)
- axes[0].set_ylabel('Cluster-level F1', fontsize=11)
- fig.tight_layout()
- fig.savefig(os.path.join(PLOT_DIR, "cluster_f1_condition_boxplot.png"), dpi=300)
- plt.close()
- print(" Saved: cluster_f1_condition_boxplot.png")
- return pd.DataFrame(omnibus_rows), pd.DataFrame(posthoc_rows)
- # =============================================================
- # 5. PER-DATASET: BeLOVE vs Challenge (Supplemental)
- # =============================================================
- def per_dataset(df):
- """
- Supplemental analysis: in-distribution (BeLOVE) vs
- out-of-distribution (Challenge) performance.
- Mann-Whitney U for unpaired comparison.
- """
- filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
- # Descriptive
- desc_rows = []
- for ds in ['BeLOVE', 'Challenge']:
- dsdf = filled[filled['dataset'] == ds]
- n = dsdf['subject'].nunique()
- for th in ['85', '90', 'locate']:
- sub = dsdf[dsdf['threshold'] == th]
- per_sub = sub.groupby('subject')[ALL_METRICS].mean()
- row = {'dataset': ds, 'n': n, 'threshold': TH_PLAIN[th]}
- for m in ALL_METRICS:
- mean, lo, hi = bootstrap_ci(per_sub[m].values)
- row[f'{m}_mean'] = mean
- row[f'{m}_sd'] = round(per_sub[m].std(), 4)
- row[f'{m}_ci_lower'] = lo
- row[f'{m}_ci_upper'] = hi
- row['n_pred_mean'] = round(sub['n_pred_clusters'].mean(), 1)
- row['n_gt_mean'] = round(sub['n_gt_clusters'].mean(), 1)
- desc_rows.append(row)
- # Statistical comparison
- stat_rows = []
- print("\n" + "=" * 70)
- print("PER-DATASET (Mann-Whitney U: BeLOVE vs Challenge)")
- print("=" * 70)
- for th in ['85', '90', 'locate']:
- bel = filled[(filled['dataset'] == 'BeLOVE') & (filled['threshold'] == th)].groupby("subject")[ALL_METRICS].mean()
- cha = filled[(filled['dataset'] == 'Challenge') & (filled['threshold'] == th)].groupby("subject")[ALL_METRICS].mean()
- for m in ALL_METRICS:
- _, p = mannwhitneyu(bel[m].values, cha[m].values, alternative='two-sided')
- d_val, d_size = (np.nan, 'N/A')
- if HAS_CLIFFS:
- d_val, d_size = cliffs_delta(bel[m].values, cha[m].values)
- m_bel, lo_bel, hi_bel = bootstrap_ci(bel[m].values)
- m_cha, lo_cha, hi_cha = bootstrap_ci(cha[m].values)
- stat_rows.append({
- 'threshold': TH_PLAIN[th], 'metric': m,
- 'mean_belove': m_bel,
- 'ci_belove': f'[{lo_bel:.4f}, {hi_bel:.4f}]',
- 'mean_challenge': m_cha,
- 'ci_challenge': f'[{lo_cha:.4f}, {hi_cha:.4f}]',
- 'diff': round(m_bel - m_cha, 4),
- 'p_mann_whitney': round(p, 4),
- 'cliffs_delta': round(d_val, 4) if not np.isnan(d_val) else np.nan,
- 'effect_size': d_size,
- })
- if m == 'lesion_f1':
- print(f" {TH_PLAIN[th]} F1: BeLOVE={bel[m].mean():.4f} vs "
- f"Challenge={cha[m].mean():.4f} p={p:.4f}")
- return pd.DataFrame(desc_rows), pd.DataFrame(stat_rows)
- # =============================================================
- # 6. GT CLUSTER DISTRIBUTION (Supplemental)
- # =============================================================
- def gt_distribution(df):
- """Descriptive statistics of GT cluster counts (26-connectivity)."""
- filled = df[(df['train_condition'] == 'filled') &
- (df['test_condition'] == 'filled') &
- (df['threshold'] == '85')]
- gt = filled.groupby('subject')['n_gt_clusters'].mean()
- rows = []
- for label, data in [('All', gt),
- ('BeLOVE', filled[filled['dataset'] == 'BeLOVE'].groupby('subject')['n_gt_clusters'].mean()),
- ('Challenge', filled[filled['dataset'] == 'Challenge'].groupby('subject')['n_gt_clusters'].mean())]:
- rows.append({
- 'group': label, 'n': len(data),
- 'mean': round(data.mean(), 1), 'sd': round(data.std(), 1),
- 'median': round(data.median(), 1),
- 'min': int(data.min()), 'max': int(data.max()),
- 'q25': int(data.quantile(0.25)), 'q75': int(data.quantile(0.75)),
- })
- print("\n" + "=" * 70)
- print("GT CLUSTER DISTRIBUTION (26-connectivity)")
- print("=" * 70)
- for r in rows:
- print(f" {r['group']:>10s}: mean={r['mean']}, median={r['median']}, "
- f"range={r['min']}-{r['max']}")
- return pd.DataFrame(rows)
- # =============================================================
- # 7. SUMMARY TABLE (Supplemental)
- # =============================================================
- def summary_table(df):
- """Mean, SD, median per train x test x threshold (all 27 combos)."""
- metric_cols = ['dice_score', 'sensitivity', 'precision',
- 'lesion_f1', 'lesion_precision', 'lesion_recall',
- 'n_pred_clusters', 'n_gt_clusters']
- group_cols = ['train_condition', 'test_condition', 'threshold']
- summary = df.groupby(group_cols)[metric_cols].agg(['mean', 'std', 'median']).round(4)
- # Flatten MultiIndex columns
- flat = summary.reset_index()
- flat.columns = ['_'.join(str(c) for c in col).rstrip('_') for col in flat.columns]
- return flat
- # =============================================================
- # MAIN
- # =============================================================
- def main():
- print("=" * 70)
- print("CLUSTER-LEVEL ANALYSIS OF BIANCA WMH SEGMENTATION")
- print("=" * 70)
- df = load_data()
- # --- MAIN TEXT ---
- dissoc_df = voxel_vs_cluster(df)
- precision_recall_bars(df)
- omnibus_df, posthoc_df = condition_comparison(df)
- # --- SUPPLEMENTAL ---
- dataset_desc_df, dataset_stat_df = per_dataset(df)
- gt_df = gt_distribution(df)
- summary_df = summary_table(df)
- # =============================================================
- # SAVE
- # =============================================================
- print(f"\nSaving to {OUTPUT_DIR}/...")
- # Main Text tables
- save_excel(
- os.path.join(OUTPUT_DIR, "voxel_vs_cluster_dissociation.xlsx"),
- dissoc_df,
- "Voxel-Level vs Cluster-Level Performance Dissociation",
- ["Addresses R1 Comment 2: Cluster F1 substantially exceeds Dice, indicating that BIANCA detects most",
- "WMH clusters but delineates boundaries imprecisely. High cluster precision confirms detected regions",
- "are genuine WMH, not distributed false positives.",
- "Condition: train=filled, test=filled. Per-subject means across 10 seeds x 5-fold CV (n=89).",
- "Cluster parameters: 26-connectivity, any-overlap (MICCAI 2017), prediction-only MCS filtering."])
- save_excel(
- os.path.join(OUTPUT_DIR, "kruskal_wallis_train_condition.xlsx"),
- omnibus_df,
- "Kruskal-Wallis H-Test: Effect of Training Preprocessing on Cluster-Level Metrics",
- ["Question: Does training with Non Removed vs Removed vs Inpainted images affect lesion detection?",
- "Design: 3 groups (train conditions), test condition fixed to 'filled' (inpainted).",
- "Test: Kruskal-Wallis H-test (non-parametric omnibus, k=3 independent groups).",
- "Null hypothesis: All three training conditions produce equal cluster-level metrics.",
- "Significance: alpha = 0.05.",
- "Data: Per-subject means averaged across 10 seeds x 5-fold stratified CV (n=89 per group).",
- "If significant: see posthoc_wilcoxon_bonferroni.xlsx for pairwise comparisons."])
- save_excel(
- os.path.join(OUTPUT_DIR, "posthoc_wilcoxon_bonferroni.xlsx"),
- posthoc_df,
- "Post-hoc Wilcoxon Signed-Rank Tests with Bonferroni Correction",
- ["Pairwise comparisons of training conditions (paired by subject).",
- "Test: Wilcoxon signed-rank (non-parametric, paired).",
- "Correction: Bonferroni (3 comparisons, corrected alpha = 0.05/3 = 0.0167).",
- "Effect size: Cliff's Delta (|d| >= 0.28 = meaningful, per manuscript conventions).",
- "Column 'omnibus_significant': Kruskal-Wallis was significant for this metric.",
- "Column 'significant_bonferroni': True only if BOTH omnibus AND post-hoc significant.",
- "Note: Post-hoc tests interpretable only when omnibus is significant."])
- # Supplemental tables
- save_excel(
- os.path.join(OUTPUT_DIR, "per_dataset_descriptive.xlsx"),
- dataset_desc_df,
- "Cluster-Level Metrics by Dataset: BeLOVE vs WMH Segmentation Challenge",
- ["BeLOVE (n=58): in-distribution (included in cross-validated training).",
- "Challenge (n=31): out-of-distribution (external validation dataset).",
- "Condition: train=filled, test=filled."])
- save_excel(
- os.path.join(OUTPUT_DIR, "per_dataset_mann_whitney.xlsx"),
- dataset_stat_df,
- "Mann-Whitney U Test: BeLOVE vs Challenge Comparison",
- ["Test: Mann-Whitney U (non-parametric, unpaired, 2 groups).",
- "Significance: alpha = 0.05.",
- "Effect size: Cliff's Delta.",
- "Per-subject means (BeLOVE n=58, Challenge n=31)."])
- save_excel(
- os.path.join(OUTPUT_DIR, "gt_cluster_distribution.xlsx"),
- gt_df,
- "Ground Truth Cluster Count Distribution (26-Connectivity)",
- ["Descriptive statistics of the number of GT clusters per subject.",
- "26-connectivity used (consistent with all cluster-level analyses).",
- "No size filtering applied to ground truth (expert-delineated lesions)."])
- save_excel(
- os.path.join(OUTPUT_DIR, "summary_all_conditions.xlsx"),
- summary_df,
- "Full Summary: Mean, SD, Median per Train x Test x Threshold",
- ["All 3 train x 3 test x 3 threshold = 27 combinations.",
- "Statistics: mean, standard deviation, median for each metric.",
- "Averaged across 10 seeds x 5-fold stratified CV x 89 subjects."])
- # Raw data
- df.to_excel(os.path.join(OUTPUT_DIR, "all_seeds_raw.xlsx"), index=False)
- print(f"\nDone. Results: {OUTPUT_DIR}/ Plots: {PLOT_DIR}/")
- if __name__ == "__main__":
- main()
4_Analyze_cluster_metrics.py at commit 1f27300, no license · at the source
Overview
- Center for Stroke Research Berlin, Charité-Universitätsmedizin Berlin, Corporate Member of Freie Universität Berlin and Humboldt-Universität zu Berlin, Berlin, Germany
- Berlin Institute of Health (BIH) at Charité-Universitätsmedizin Berlin, Berlin, Germany
- German Centre for Cardiovascular Research (DZHK), Partner site Berlin, Berlin, Germany
- Department of Neurology, Charité-Universitätsmedizin Berlin, Berlin, Germany
- Department of Cardiology, Charité-Universitätsmedizin Berlin, Berlin, Germany
- German Center for Neurodegenerative Diseases (DZNE), Partner site Berlin, Berlin, Germany
- German Center for Mental Health (DZPG), Partner site Berlin, Berlin, Germany
- Philips Clinical Science, Hamburg, Germany
Abstract
Background: White matter hyperintensity (WMH) segmentation using BIANCA (Brain Intensity AbNormality Classification Algorithm) in stroke populations is complicated by vascular lesions that share T2-hyperintense signal characteristics with WMH. Whether preprocessing decisions in the treatment of lesions affect segmentation accuracy has not yet been systematically evaluated in cerebrovascular cohorts involving multiple scanners.
Methods: We compared fixed probability thresholds and locally adaptive thresholding with LOCATE (LOCally Adaptive Threshold Estimation), and three lesion-handling approaches: lesions present (non removed), replaced with zero intensities (removed), and replaced with normal-appearing white matter intensities (NAWM; inpainted), using the BeLOVE cohort (Berlin Longterm Observation of Vascular Events) and the WMH Segmentation Challenge dataset. Phase I (n = 89) optimized thresholding via stratified 5-fold cross-validation. Phase II assessed preprocessing effects on segmentation accuracy (Phase II-A, n = 89) and volume agreement (Phase II-B, n = 211).
Results: BIANCA with LOCATE adaptive thresholding achieved moderate segmentation overlap (mean Dice 0.567), with lesion-level detection exceeding an F1 of 0.85. Preprocessing effects were statistically detectable but negligible in magnitude, with near-perfect agreement between all conditions. Stroke lesion volume was the highest-ranked predictor of volume differences between conditions; these scaled with lesion size but remained negligible in magnitude across all subgroups.
Conclusions: BIANCA with LOCATE achieved moderate WMH segmentation performance with the best sensitivity-precision trade-off in this multi-scanner cerebrovascular cohort. Preprocessing effects were negligible at the group level. However, large lesions distort the FLAIR intensity distribution on which BIANCA relies for classification, which justifies lesion removal as a recommended preprocessing step.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 28 matches between paragraphs and lines of code.
mendeltem/Bianca_Evaluation
1f27300aa772a82509608f194a7558f4d79b4b70, 20 March 2026Availability: 1 check, the latest on 28 September 2026: the link answers
- 28 September 2026: the link answers
35 files
- 0_Severity_definitions.p
y , Python, 376 lines - 0_severity_population.py
, Python, 258 lines - 1_preprare_locate.py, Python, 724 lines, 2 matches
- 2_compare_GE_analyse.py, Python, 1,299 lines
- 3_threshold_analysis_all
_seeds.py , Python, 707 lines - 3_threshold_analysis_par
allel_temp.py , Python, 523 lines, 2 matches - 4_Analyze_cluster_metric
s.py , Python, 771 lines, 4 matches - 4_cluster_size_grid_sear
ch.py , Python, 880 lines, 1 match - 4_plot_cluster_grid_sear
ch.py , Python, 372 lines, 1 match - 5_create_5_fold_cross_va
lidation_sets.py , Python, 1,152 lines, 2 matches - 5_population_5fcv.py, Python, 92 lines
- 6_train_bianca.py, Python, 201 lines, 1 match
- 7_run_removal.py, Python, 681 lines
- 7_run_removal.sh, Shell, 30 lines
- 8_histogramm_removal.py, Python, 526 lines
- Phase_1/
1_phase_1_belove_k_bianc , Python, 792 lines, 1 matcha_population.py - Phase_1/
2_severity.py , Python, 228 lines - Phase_1/
5FCV_SET/ , Python, 1,003 lines, 1 match1_voxel_treshhold_analys is.py - Phase_1/
5FCV_SET/ , Python, 731 lines2_cluster_dice_score_ana lysis.py - Phase_1/
5FCV_SET/ , Python, 1,077 lines3_compare_train_conditio ns.py - Phase_1/
5FCV_SET/ , Python, 1,080 lines4_compare_test_condition s.py - Phase_1/
GE_COMPARE/ , Python, 489 lines, 1 match2_GE_analysis.py - Phase_2_a/
0_2a_population.py , Python, 226 lines, 2 matches - Phase_2_a/
1_2a_bland_altman_test.p , Python, 967 linesy - Phase_2_a/
1_2a_bland_altman_test_t , Python, 984 lines, 2 matchesemp.py - Phase_2_a/
4_2a_volumes.py , Python, 478 lines - Phase_2_a/
5_2a_shapley_values.py , Python, 619 lines, 2 matches - Phase_2_a/
8_2a_correlation_dice_2. , Python, 818 lines, 1 matchpy - Phase_2_b/
0_2b_population.py , Python, 344 lines, 2 matches - Phase_2_b/
2_2b_volume_difference_l , Python, 588 lines, 1 matchesion_bonf.py - Phase_2_b/
3_2b_volume_difference_s , Python, 472 lines, 1 matchcanner_bonf.py - Phase_2_b/
5_2b_shapley_values.py , Python, 359 lines, 1 match - Phase_2_b/
7_2b_lesion_vol_dic.py , Python, 549 lines - Phase_2_b/
8_correlation_analysis.p , Python, 657 linesy - README.md, Text, 81 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 34 scripts, each with its path and the digest of its content;
- 28 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
Analysis scripts for WMH segmentation, preprocessing, statistical testing and figure generation are available at https://
Source imaging data cannot be shared publicly due to participant privacy restrictions under the European General Data Protection Regulation (GDPR). Processed, de-identified derivative data can be obtained from the corresponding author upon reasonable request and completion of appropriate data sharing agreements with the Center for Stroke Research Berlin.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 28 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 9 authors, 7 keywords, 10 MeSH terms, 3 funders, 47 references.
Cite
This paper
Temuulen, U., Mekle, R., Galinovic, I., Weber, J. E., Landmesser, U., Kelle, S., Endres, M., Stehning, C., & Villringer, K. (2026). Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation. NeuroImage. Clinical, 50, 104001. https://
BibTeX
@article{temuulen2026eff
author = {Temuulen, Uchralt and Mekle, Ralf and Galinovic, Ivana and Weber, Joachim E and Landmesser, Ulf and Kelle, Sebastian and Endres, Matthias and Stehning, Christian and Villringer, Kersten},
title = {{Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation}},
journal = {NeuroImage. Clinical},
year = {2026},
month = may,
volume = {50},
pages = {104001},
publisher = {Elsevier},
issn = {2213-1582},
doi = {10.1016/
url = {https://
pmid = {42119523},
pmcid = {PMC13195355}
}
RIS
TY - JOUR
AU - Temuulen, Uchralt
AU - Mekle, Ralf
AU - Galinovic, Ivana
AU - Weber, Joachim E
AU - Landmesser, Ulf
AU - Kelle, Sebastian
AU - Endres, Matthias
AU - Stehning, Christian
AU - Villringer, Kersten
TI - Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation
T2 - NeuroImage. Clinical
J2 - Neuroimage Clin
PY - 2026
DA - 2026/
VL - 50
SP - 104001
SN - 2213-1582
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation",
"container-title": "NeuroImage. Clinical",
"author": [
{
"family": "Temuulen",
"given": "Uchralt"
},
{
"family": "Mekle",
"given": "Ralf"
},
{
"family": "Galinovic",
"given": "Ivana"
},
{
"family": "Weber",
"given": "Joachim E"
},
{
"family": "Landmesser",
"given": "Ulf"
},
{
"family": "Kelle",
"given": "Sebastian"
},
{
"family": "Endres",
"given": "Matthias"
},
{
"family": "Stehning",
"given": "Christian"
},
{
"family": "Villringer",
"given": "Kersten"
}
],
"container-title-short":
"volume": "50",
"page": "104001",
"DOI": "10.1016/
"PMID": "42119523",
"PMCID": "PMC13195355",
"ISSN": "2213-1582",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
11
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s12938-026-01555-0 [code]
- Incorporating normal periventricular changes for enhanced pathological white matter hyperintensity segmentation: on multiclass deep learning approaches.Journal: Biomedical engineering onlineIn common: seaborn, scikit-learn, pandas, 3 other tools, stroke, methods / tools, structural MRI / diffusion, 5 references
- [2] doi:10.1212/wnl.0000000000218472 [code]
- Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.Journal: NeurologyIn common: SHAP, NiBabel, scikit-learn, 4 other tools, stroke, structural MRI / diffusion, 3 references
- [3] doi:10.1111/ene.70678 [code]
- Who Falls After a Stroke? Evidence From a Prospective Stroke Cohort.Journal: European journal of neurologyIn common: NiBabel, seaborn, pandas, 3 other tools, stroke, 1 reference, author Matthias Endres
- [4] doi:10.1371/journal.pbio.3003856 [code]
- Aging and metabolism contribute separately to brain-body health.Journal: PLoS biologyIn common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 3 references
- [5] doi:10.1038/s41467-026-71555-0 [code]
- A deep representation learning model to predict response to vagus nerve stimulation.Journal: Nature communicationsIn common: SHAP, FSL, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
- [6] doi:10.1162/imag.a.1252 [code]
- Does the brain's E:I balance really shape long-range temporal correlations? Lessons learned from 3T MRI.Journal: Imaging neuroscience (Cambridge, Mass.)In common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 3 references
- [7] doi:10.1002/hbm.70602 [code]
- Neuroimaging Correlates of Post-Stroke Pain After Ischemic Stroke: Secondary Analysis of the INSPiRE-TMS Trial.Journal: Human brain mappingIn common: NiBabel, seaborn, pandas, 3 other tools, stroke, author Matthias Endres
- [8] doi:10.1186/s12880-026-02481-2 [code]
- Deep learning-based neuroanatomical profiling reveals population-specific brain changes in multiple sclerosis: a large-scale Middle Eastern study.Journal: BMC medical imagingIn common: FSL, Pillow, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
- [9] doi:10.1002/ana.78203 [code]
- AI-Driven Mapping of Seizure Spread Patterns.Journal: Annals of neurologyIn common: FSL, Pillow, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
- [10] doi:10.3389/fnins.2026.1870124 [code]
- An end-to-end pipeline for automated fetal brain segmentation and biometry from 3D SSFP MRI.Journal: Frontiers in neuroscienceIn common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 34 scripts, and 28 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:dc59e2752e2f7912…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
