OSCR

Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation.

Code ↔ Paper

28 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 28 matches
  1. [1] § Methods › Statistical analysis ↔ Phase_2_a/1_2a_bland_altman_test_temp.py, lines 919–984 · score 0.94 · 0.75–0.9, Intraclass Correlation Coefficient, Bland Altman, 0.5–0.75, equivalence margin, way mixed
  2. [2] § Methods › LOCATE adaptive thresholding ↔ 1_preprare_locate.py, lines 4–85 · score 0.91 · unbiased lesion probability, LOCally Adaptive Threshold, fold cross validation, LOCATE training, training subjects, leakage
  3. [3] § Methods › Lesion delineation, removal, and inpainting ↔ 6_train_bianca.py, lines 1–59 · score 0.85 · NAWM intensities, lesion_filling, FLAIR images, lesion voxel, intensity distribution, FSL
  4. [4] § Results › Phase II-B: Volume-based agreement (n = 211) ↔ Phase_2_b/2_2b_volume_difference_lesion_bonf.py, lines 1–61 · score 0.81 · 0.14 mL, 0.07 mL, 0.29 mL, 0.06 mL, marginal, 0.04 mL
  5. [5] § Methods › Statistical analysis ↔ Phase_2_a/5_2a_shapley_values.py, lines 1–59 · score 0.79 · TreeExplainer, max depth, Random Forest, bag, SHAP, trees
  6. [6] § Methods › Phase I: Parameter optimization › Cluster-based analysis ↔ 4_Analyze_cluster_metrics.py, lines 43–116 · score 0.77 · ground truth masks, Minimum cluster, Dice coefficient, complementing, recall, detection
  7. [7] § Methods › Preprocessing pipeline ↔ 1_preprare_locate.py, lines 4–85 · score 0.75 · ventricle distance map, white matter mask, adaptive thresholding, pipeline, LOCATE, segmentation
  8. [8] § Methods › Statistical analysis ↔ Phase_2_b/5_2b_shapley_values.py, lines 1–63 · score 0.75 · TreeExplainer, max depth, Random Forest, SHAP, trees, MAE
  9. [9] § Results › Population characteristics ↔ Phase_2_b/0_2b_population.py, lines 1–79 · score 0.74 · Prisma fit, 3–10, Scanner distribution, BeLOVE, lesion volume, ARWMC
  10. [10] § Methods › Datasets and scanner distribution ↔ Phase_2_a/0_2a_population.py, lines 1–85 · score 0.72 · full cohort, ground truth WMH, Prisma fit, WMH masks, BeLOVE, Trio
  11. [11] § Methods › Lesion delineation, removal, and inpainting ↔ 3_threshold_analysis_parallel_temp.py, lines 1–78 · score 0.70 · lesion_filling, lesion voxel, zero filled, FSL, physiologically, NAWM
  12. [12] § Methods › Phase I: Parameter optimization › Cluster-based analysis ↔ 4_plot_cluster_grid_search.py, lines 1–48 · score 0.69 · Optimal minimum cluster, ground truth masks, recall, filtering, predicted, voxel
  13. [13] § Methods › Study design and datasets ↔ 3_threshold_analysis_parallel_temp.py, lines 1–78 · score 0.69 · lesion_filling, lesion voxels, zero filled, cross validation, FSL, NAWM
  14. [14] § Results › Population characteristics ↔ Phase_2_a/0_2a_population.py, lines 1–85 · score 0.69 · LOCATE training pool, ground truth WMH, Prisma fit, scanner balanced, BeLOVE, Population
  15. [15] § Results › Phase II-B: Volume-based agreement (n = 211) ↔ Phase_2_b/3_2b_volume_difference_scanner_bonf.py, lines 1–59 · score 0.69 · Prisma fit, 14.67 mL, 15.58 mL, largest, partly, baseline
  16. [16] § Methods › Study design and datasets ↔ 5_create_5_fold_cross_validation_sets.py, lines 1–63 · score 0.67 · lesion_filling, cross validation, zero filled, FSL, NAWM, CV
  17. [17] § Methods › Preprocessing pipeline › GE scanner exclusion ↔ Phase_1/GE_COMPARE/2_GE_analysis.py, lines 1–86 · score 0.64 · GE inclusion, GE subjects, medium, Philips, stratified, negligible
  18. [18] § Results › Phase II-A: preprocessing effect assessment (n = 89) ↔ Phase_2_a/5_2a_shapley_values.py, lines 61–120 · score 0.64 · profiles, SHAP, explanatory, Forest, lesion volume, MAE
  19. [19] § Methods › Phase I: Parameter optimization › Training condition comparison ↔ 4_Analyze_cluster_metrics.py, lines 407–555 · score 0.64 · preprocessing affects, Wilcoxon signed rank, training preprocessing, bootstrapped, Bonferroni, Cliff
  20. [20] § Methods › Cross-validation strategy ↔ 4_cluster_size_grid_search.py, lines 1–84 · score 0.63 · LOO CV, cross validation, leakage, mixed, seeds, stratified
  21. [21] § Results › Phase I: parameter optimization (n = 89) › Threshold analysis ↔ Phase_1/5FCV_SET/1_voxel_treshhold_analysis.py, lines 1–61 · score 0.63 · disadvantage, b0, medium, adaptive, Pairwise, Wilcoxon
  22. [22] § Methods › Preprocessing pipeline ↔ Phase_2_b/0_2b_population.py, lines 1–79 · score 0.63 · HD BET, fsl_anat, Brain, mask, segmentation, Preprocessing
  23. [23] § Methods › Cross-validation strategy ↔ 5_create_5_fold_cross_validation_sets.py, lines 1–63 · score 0.62 · fold cross validation, LOO CV, seeds, stratified, severity, scanner
  24. [24] § Results › Phase II-A: preprocessing effect assessment (n = 89) ↔ Phase_2_a/1_2a_bland_altman_test_temp.py, lines 1–64 · score 0.60 · variance, attributable, TOST, opposing, marginally, shifts
  25. [25] § Results › Population characteristics ↔ Phase_1/1_phase_1_belove_k_bianca_population.py, lines 62–125 · score 0.60 · Prisma fit, ground truth WMH, 6.96 mL, BeLOVE, cutoffs, IQR
  26. [26] § Results › Phase I: parameter optimization (n = 89) › Cluster-based analysis ↔ 4_Analyze_cluster_metrics.py, lines 43–116 · score 0.58 · complemented voxel, minimum cluster, detection, precision, Cliff, metrics
  27. [27] § Methods › Statistical analysis ↔ 4_Analyze_cluster_metrics.py, lines 670–767 · score 0.54 · Mann Whitney, Wilcoxon signed rank, BeLOVE, Cliff, Delta
  28. [28] § Methods › Lesion delineation, removal, and inpainting ↔ Phase_2_a/8_2a_correlation_dice_2.py, lines 203–232 · score 0.52 · intracranial hemorrhage, ischemic infarcts, lacunes, inpainting, lesions

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 771 lines · 33 KB · no license · 4 matches

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. """
  4. Cluster-Level Analysis of BIANCA WMH Segmentation
  5. ===================================================
  6. @author: temuuleu
  7. # =============================================================
  8. # REVIEWER RESPONSE (R1, Comment 2; R5, #17)
  9. # =============================================================
  10. #
  11. # R1 raised the concern that global Dice alone cannot distinguish
  12. # genuine sensitivity gains from spatially distributed false positives.
  13. # R5 (#17) noted that precision values were not contextualized.
  14. #
  15. # This script produces all cluster-level analyses for the manuscript:
  16. #
  17. # MAIN TEXT:
  18. # 1. Voxel vs Cluster dissociation (Figure: Dice vs Cluster F1 scatter)
  19. # 2. Cluster Precision vs Recall (Figure: bar chart)
  20. # 3. R vs NR condition comparison (Figure: boxplot + Kruskal-Wallis)
  21. # 4. Cliff's Delta for R vs NR on cluster-level
  22. #
  23. # SUPPLEMENTAL:
  24. # 5. Post-hoc Wilcoxon with Bonferroni correction
  25. # 6. Per-dataset breakdown (BeLOVE vs Challenge, Mann-Whitney U)
  26. # 7. GT cluster count distribution
  27. # 8. Summary tables (mean, SD, median per condition x threshold)
  28. #
  29. # NOT INCLUDED (not requested by reviewers):
  30. # - Per-scanner stratification (already covered in voxel-level analysis)
  31. # - Per-severity stratification (already covered in voxel-level analysis)
  32. # - Full 9x3 condition matrix (identical values, no added information)
  33. #
  34. # Methodology:
  35. # - Connected-component labeling: 26-connectivity
  36. # - Minimum cluster size: threshold- and dataset-specific (grid search)
  37. # - Overlap criterion: any-overlap (>=1 voxel), MICCAI 2017
  38. # - Filtering: prediction only; GT retained without filtering
  39. # - Evaluated: 10 seeds x 5-fold stratified CV x 89 subjects
  40. # =============================================================
  41. # =============================================================
  42. # PAPER TEXT: METHODS
  43. # =============================================================
  44. #
  45. # 2.X Lesion-level evaluation
  46. #
  47. # To complement voxel-level Dice coefficients, we computed lesion-level
  48. # precision, recall, and F1 score using connected-component analysis.
  49. # Binary segmentation masks were decomposed into spatially contiguous
  50. # clusters using 26-connectivity (scipy.ndimage.label). Predicted
  51. # clusters smaller than a threshold-specific minimum cluster size
  52. # (determined via grid search; Supplemental Table X) were removed to
  53. # exclude spurious components; ground truth masks were not filtered.
  54. # Consistent with the MICCAI 2017 WMH Segmentation Challenge (Kuijf
  55. # et al., 2019), a predicted cluster was classified as a true positive
  56. # if it overlapped with any ground truth cluster by at least one voxel
  57. # (any-overlap criterion). Lesion-level precision was defined as the
  58. # proportion of predicted clusters classified as true positives, and
  59. # lesion-level recall as the proportion of ground truth clusters
  60. # detected. Lesion-level F1 was computed as the harmonic mean of
  61. # precision and recall.
  62. #
  63. # To assess the effect of training preprocessing on cluster-level
  64. # detection, Kruskal-Wallis H-tests were performed for each threshold
  65. # and metric. Where the omnibus test was significant (alpha = 0.05),
  66. # post-hoc pairwise Wilcoxon signed-rank tests were conducted with
  67. # Bonferroni correction (corrected alpha = 0.05/3 = 0.0167). Effect
  68. # sizes were quantified using Cliff's Delta, with |delta| >= 0.28 as
  69. # the threshold for a meaningful difference.
  70. # =============================================================
  71. # =============================================================
  72. # PAPER TEXT: RESULTS
  73. # =============================================================
  74. #
  75. # 3.X Lesion-level evaluation
  76. #
  77. # Cluster-level F1 scores (0.855-0.877) substantially exceeded
  78. # voxel-level Dice coefficients (0.540-0.567), indicating that BIANCA
  79. # detected the majority of WMH clusters but delineated their boundaries
  80. # imprecisely (Supplemental Figure X). Lesion-level precision was
  81. # consistently high across all thresholding approaches (0.881-0.928),
  82. # confirming that the majority of predicted clusters corresponded to
  83. # genuine WMH regions. Lesion-level recall ranged from 0.849 to 0.905,
  84. # with B+L achieving the highest detection rate.
  85. #
  86. # Kruskal-Wallis tests revealed no significant differences between
  87. # training preprocessing conditions on any cluster-level metric
  88. # (all p > 0.92; Supplemental Table X), confirming that lesion removal
  89. # does not affect lesion detection performance. This finding, combined
  90. # with the high and stable cluster-level precision (0.881-0.928),
  91. # addresses the concern that higher WMH volumes after lesion removal
  92. # might reflect spatially distributed false positives rather than
  93. # genuine sensitivity gains.
  94. # =============================================================
  95. """
  96. import matplotlib
  97. matplotlib.use('Agg')
  98. import pandas as pd
  99. import numpy as np
  100. import matplotlib.pyplot as plt
  101. import seaborn as sns
  102. import os
  103. import glob
  104. from scipy.stats import wilcoxon, mannwhitneyu, kruskal
  105. try:
  106. from cliffs_delta import cliffs_delta
  107. HAS_CLIFFS = True
  108. except ImportError:
  109. HAS_CLIFFS = False
  110. print("WARNING: cliffs_delta not installed. Install with: pip install cliffs_delta")
  111. # =============================================================
  112. # CONFIG
  113. # =============================================================
  114. SEED_DIR = "Phase_1/5FCV_SET/bianca_result_dir"
  115. OUTPUT_DIR = "Phase_1/Cluster_analysis_results"
  116. PLOT_DIR = os.path.join(OUTPUT_DIR, "plots")
  117. # Labels for plots (LaTeX subscript) vs Excel (plain text)
  118. TH_PLOT = {'85': r'B$_{0.85}$', '90': r'B$_{0.90}$', 'locate': 'B+L'}
  119. TH_PLAIN = {'85': 'B_0.85', '90': 'B_0.90', 'locate': 'B+L'}
  120. COND_LABELS = {
  121. 'non_removed': 'Train Non Removed',
  122. 'removed': 'Train Removed',
  123. 'filled': 'Train Inpainted',
  124. }
  125. # All metrics: voxel-level + cluster-level
  126. VOXEL_METRICS = ['dice_score', 'sensitivity', 'precision']
  127. CLUSTER_METRICS = ['lesion_f1', 'lesion_precision', 'lesion_recall']
  128. ALL_METRICS = VOXEL_METRICS + CLUSTER_METRICS
  129. # Bootstrap CI parameters (consistent with manuscript: Efron & Tibshirani, 1993)
  130. N_BOOTSTRAP = 1000
  131. CI_ALPHA = 0.05 # 95% CI
  132. RANDOM_SEED = 42
  133. os.makedirs(OUTPUT_DIR, exist_ok=True)
  134. os.makedirs(PLOT_DIR, exist_ok=True)
  135. # =============================================================
  136. # HELPER: Excel export with metadata
  137. # =============================================================
  138. def save_excel(filepath, dataframe, title, description):
  139. """Save DataFrame to Excel with descriptive metadata header."""
  140. from openpyxl import Workbook
  141. from openpyxl.styles import Font, PatternFill, Alignment, Border, Side
  142. from openpyxl.utils import get_column_letter
  143. wb = Workbook()
  144. ws = wb.active
  145. n_cols = len(dataframe.columns)
  146. merge_width = max(n_cols, 6)
  147. thin = Border(*(Side(style='thin', color='B0B0B0'),) * 4)
  148. # Title
  149. row = 1
  150. ws.merge_cells(start_row=row, start_column=1, end_row=row, end_column=merge_width)
  151. c = ws.cell(row=row, column=1, value=title)
  152. c.font = Font(name='Arial', bold=True, size=11, color='FFFFFF')
  153. c.fill = PatternFill('solid', fgColor='2F5496')
  154. c.alignment = Alignment(horizontal='left', vertical='center', wrap_text=True)
  155. # Description
  156. for line in description:
  157. row += 1
  158. ws.merge_cells(start_row=row, start_column=1, end_row=row, end_column=merge_width)
  159. c = ws.cell(row=row, column=1, value=line)
  160. c.font = Font(name='Arial', size=9, italic=True, color='333333')
  161. c.fill = PatternFill('solid', fgColor='F2F2F2')
  162. c.alignment = Alignment(horizontal='left', vertical='center', wrap_text=True)
  163. row += 2 # empty separator
  164. # Headers
  165. for ci, col in enumerate(dataframe.columns, 1):
  166. c = ws.cell(row=row, column=ci, value=col)
  167. c.font = Font(name='Arial', bold=True, size=10)
  168. c.fill = PatternFill('solid', fgColor='D6E4F0')
  169. c.alignment = Alignment(horizontal='center', vertical='center')
  170. c.border = thin
  171. # Data
  172. for _, data_row in dataframe.iterrows():
  173. row += 1
  174. for ci, val in enumerate(data_row, 1):
  175. c = ws.cell(row=row, column=ci)
  176. if isinstance(val, (np.bool_, bool)):
  177. c.value = str(val)
  178. elif isinstance(val, float) and not np.isnan(val):
  179. c.value = val
  180. c.number_format = '0.0000'
  181. else:
  182. c.value = val
  183. c.font = Font(name='Arial', size=10)
  184. c.alignment = Alignment(horizontal='center')
  185. c.border = thin
  186. # Auto-width
  187. for ci in range(1, n_cols + 1):
  188. max_len = max(len(str(dataframe.columns[ci-1])),
  189. dataframe.iloc[:, ci-1].astype(str).str.len().max() if len(dataframe) > 0 else 0)
  190. ws.column_dimensions[get_column_letter(ci)].width = min(max_len + 3, 35)
  191. wb.save(filepath)
  192. # =============================================================
  193. # HELPER: Bootstrap 95% confidence intervals
  194. # =============================================================
  195. def bootstrap_ci(values, n_boot=N_BOOTSTRAP, alpha=CI_ALPHA, seed=RANDOM_SEED):
  196. """
  197. Compute bootstrap 95% CI for the mean.
  198. Same-size resampling with replacement, 1000 iterations,
  199. consistent with manuscript methodology (Efron & Tibshirani, 1993).
  200. Returns (mean, ci_lower, ci_upper)
  201. """
  202. rng = np.random.default_rng(seed)
  203. values = np.asarray(values)
  204. n = len(values)
  205. boot_means = np.array([
  206. rng.choice(values, size=n, replace=True).mean()
  207. for _ in range(n_boot)
  208. ])
  209. ci_lo = np.percentile(boot_means, 100 * alpha / 2)
  210. ci_hi = np.percentile(boot_means, 100 * (1 - alpha / 2))
  211. return round(values.mean(), 4), round(ci_lo, 4), round(ci_hi, 4)
  212. def fmt_ci(mean, lo, hi):
  213. """Format as 'mean [lo, hi]' for tables."""
  214. return f"{mean:.3f} [{lo:.3f}, {hi:.3f}]"
  215. # =============================================================
  216. # 1. LOAD DATA
  217. # =============================================================
  218. def load_data(seed_dir=SEED_DIR):
  219. files = sorted(glob.glob(os.path.join(seed_dir, "bianca_metrics_seed_*.xlsx")))
  220. if not files:
  221. files = sorted(glob.glob("bianca_metrics_seed_*.xlsx"))
  222. if not files:
  223. raise FileNotFoundError("No bianca_metrics_seed_*.xlsx files found.")
  224. df = pd.concat([pd.read_excel(f) for f in files], ignore_index=True)
  225. df['threshold'] = df['threshold'].astype(str)
  226. df['dataset'] = df['subject'].apply(
  227. lambda x: 'BeLOVE' if x.startswith('belove') else 'Challenge')
  228. print(f"Loaded {len(files)} seeds: {len(df)} rows, "
  229. f"{df['subject'].nunique()} subjects")
  230. return df
  231. # =============================================================
  232. # 2. VOXEL vs CLUSTER DISSOCIATION (Main Text Figure + Table)
  233. # =============================================================
  234. def voxel_vs_cluster(df):
  235. """
  236. Core analysis addressing R1 Comment 2:
  237. Cluster F1 >> Dice demonstrates that BIANCA detects most lesions
  238. but delineates boundaries imprecisely. High cluster precision
  239. confirms that detected regions are genuine WMH, not distributed FP.
  240. Reports both voxel-level (Dice, sensitivity, precision) and
  241. cluster-level (F1, precision, recall) with bootstrap 95% CIs.
  242. """
  243. filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
  244. per_sub = filled.groupby(['subject', 'threshold', 'dataset']).agg(
  245. {m: 'mean' for m in ALL_METRICS}).reset_index()
  246. # Summary table with bootstrap CIs
  247. rows = []
  248. print("\n" + "=" * 70)
  249. print("VOXEL vs CLUSTER PERFORMANCE (train=filled, test=filled)")
  250. print("Bootstrap 95% CI (1000 iterations)")
  251. print("=" * 70)
  252. for th in ['85', '90', 'locate']:
  253. sub = per_sub[per_sub['threshold'] == th]
  254. row = {'threshold': TH_PLAIN[th]}
  255. # Voxel-level metrics with CIs
  256. for m, label in [('dice_score', 'Dice'),
  257. ('sensitivity', 'Vox Sens'),
  258. ('precision', 'Vox Prec')]:
  259. mean, lo, hi = bootstrap_ci(sub[m].values)
  260. row[f'{m}_mean'] = mean
  261. row[f'{m}_sd'] = round(sub[m].std(), 4)
  262. row[f'{m}_ci_lower'] = lo
  263. row[f'{m}_ci_upper'] = hi
  264. # Cluster-level metrics with CIs
  265. for m, label in [('lesion_f1', 'Clust F1'),
  266. ('lesion_precision', 'Clust P'),
  267. ('lesion_recall', 'Clust R')]:
  268. mean, lo, hi = bootstrap_ci(sub[m].values)
  269. row[f'{m}_mean'] = mean
  270. row[f'{m}_sd'] = round(sub[m].std(), 4)
  271. row[f'{m}_ci_lower'] = lo
  272. row[f'{m}_ci_upper'] = hi
  273. row['f1_minus_dice'] = round(
  274. row['lesion_f1_mean'] - row['dice_score_mean'], 4)
  275. rows.append(row)
  276. print(f"\n {TH_PLOT[th]}:")
  277. print(f" Voxel: Dice={fmt_ci(row['dice_score_mean'], row['dice_score_ci_lower'], row['dice_score_ci_upper'])} "
  278. f"Sens={fmt_ci(row['sensitivity_mean'], row['sensitivity_ci_lower'], row['sensitivity_ci_upper'])} "
  279. f"Prec={fmt_ci(row['precision_mean'], row['precision_ci_lower'], row['precision_ci_upper'])}")
  280. print(f" Cluster: F1={fmt_ci(row['lesion_f1_mean'], row['lesion_f1_ci_lower'], row['lesion_f1_ci_upper'])} "
  281. f"P={fmt_ci(row['lesion_precision_mean'], row['lesion_precision_ci_lower'], row['lesion_precision_ci_upper'])} "
  282. f"R={fmt_ci(row['lesion_recall_mean'], row['lesion_recall_ci_lower'], row['lesion_recall_ci_upper'])}")
  283. print(f" Gap (cF1 - Dice) = {row['f1_minus_dice']:+.3f}")
  284. # Scatter plot: Dice vs Cluster F1
  285. colors = {'85': '#1f77b4', '90': '#ff7f0e', 'locate': '#2ca02c'}
  286. fig, ax = plt.subplots(figsize=(8, 7))
  287. for th in ['85', '90', 'locate']:
  288. sub = per_sub[per_sub['threshold'] == th]
  289. ax.scatter(sub['dice_score'], sub['lesion_f1'], alpha=0.5, s=30,
  290. color=colors[th], label=TH_PLOT[th])
  291. ax.plot([0, 1], [0, 1], 'k--', alpha=0.3, lw=1)
  292. ax.set_xlabel('Voxel-level Dice', fontsize=12)
  293. ax.set_ylabel('Cluster-level F1', fontsize=12)
  294. ax.legend(fontsize=10)
  295. ax.grid(True, alpha=0.2)
  296. ax.set_xlim(0, 1); ax.set_ylim(0, 1)
  297. fig.tight_layout()
  298. fig.savefig(os.path.join(PLOT_DIR, "dice_vs_cluster_f1_scatter.png"), dpi=300)
  299. plt.close()
  300. print(" Saved: dice_vs_cluster_f1_scatter.png")
  301. return pd.DataFrame(rows)
  302. # =============================================================
  303. # 3. PRECISION vs RECALL: Voxel + Cluster (Main Text Figure)
  304. # =============================================================
  305. def precision_recall_bars(df):
  306. """
  307. Addresses R5 #17: Precision values contextualized.
  308. Shows voxel-level (Dice, Sens, Prec) vs cluster-level (F1, P, R)
  309. side by side, demonstrating that moderate voxel precision coexists
  310. with high cluster precision.
  311. """
  312. filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
  313. thresholds = ['85', '90', 'locate']
  314. fig, axes = plt.subplots(1, 2, figsize=(14, 5.5))
  315. # Panel A: Voxel-level
  316. vox_metrics = [('dice_score', 'Dice'), ('sensitivity', 'Sensitivity'), ('precision', 'Precision')]
  317. x = np.arange(len(thresholds))
  318. w = 0.25
  319. vox_colors = ['#5B9BD5', '#70AD47', '#ED7D31']
  320. for i, (m, label) in enumerate(vox_metrics):
  321. means = [filled[filled['threshold'] == t][m].mean() for t in thresholds]
  322. sds = [filled[filled['threshold'] == t][m].std() for t in thresholds]
  323. bars = axes[0].bar(x + i * w - w, means, w, yerr=sds, label=label,
  324. color=vox_colors[i], capsize=3, alpha=0.85)
  325. for bar in bars:
  326. axes[0].text(bar.get_x() + bar.get_width()/2., bar.get_height() + 0.02,
  327. f'{bar.get_height():.2f}', ha='center', va='bottom', fontsize=8)
  328. axes[0].set_xticks(x)
  329. axes[0].set_xticklabels([TH_PLOT[t] for t in thresholds], fontsize=11)
  330. axes[0].set_ylabel('Score', fontsize=11)
  331. axes[0].set_title('Voxel-level', fontsize=12, fontweight='bold')
  332. axes[0].legend(fontsize=9)
  333. axes[0].set_ylim(0, 1.15)
  334. axes[0].grid(True, alpha=0.2, axis='y')
  335. # Panel B: Cluster-level
  336. clu_metrics = [('lesion_f1', 'F1'), ('lesion_precision', 'Precision'), ('lesion_recall', 'Recall')]
  337. clu_colors = ['#5B9BD5', '#70AD47', '#ED7D31']
  338. for i, (m, label) in enumerate(clu_metrics):
  339. means = [filled[filled['threshold'] == t][m].mean() for t in thresholds]
  340. sds = [filled[filled['threshold'] == t][m].std() for t in thresholds]
  341. bars = axes[1].bar(x + i * w - w, means, w, yerr=sds, label=label,
  342. color=clu_colors[i], capsize=3, alpha=0.85)
  343. for bar in bars:
  344. axes[1].text(bar.get_x() + bar.get_width()/2., bar.get_height() + 0.02,
  345. f'{bar.get_height():.2f}', ha='center', va='bottom', fontsize=8)
  346. axes[1].set_xticks(x)
  347. axes[1].set_xticklabels([TH_PLOT[t] for t in thresholds], fontsize=11)
  348. axes[1].set_title('Cluster-level', fontsize=12, fontweight='bold')
  349. axes[1].legend(fontsize=9)
  350. axes[1].set_ylim(0, 1.15)
  351. axes[1].grid(True, alpha=0.2, axis='y')
  352. fig.tight_layout()
  353. fig.savefig(os.path.join(PLOT_DIR, "voxel_vs_cluster_bars.png"), dpi=300)
  354. plt.close()
  355. print(" Saved: voxel_vs_cluster_bars.png")
  356. # =============================================================
  357. # 4. CONDITION COMPARISON: Kruskal-Wallis + Wilcoxon + Cliff's Delta
  358. # (Main Text Results + Supplemental Tables)
  359. # =============================================================
  360. def condition_comparison(df):
  361. """
  362. Core analysis: Does training preprocessing affect cluster-level detection?
  363. 1. Kruskal-Wallis H-test (omnibus, 3 training conditions)
  364. 2. Post-hoc Wilcoxon signed-rank (Bonferroni alpha = 0.05/3)
  365. 3. Cliff's Delta (|delta| >= 0.28 for meaningful difference)
  366. Test condition fixed to 'filled' (inpainted).
  367. Per-subject means averaged across seeds/folds for independence.
  368. """
  369. COMPARISONS = [
  370. ('non_removed', 'removed', 'Train Non Removed vs Train Removed'),
  371. ('non_removed', 'filled', 'Train Non Removed vs Train Inpainted'),
  372. ('removed', 'filled', 'Train Removed vs Train Inpainted'),
  373. ]
  374. BONF_ALPHA = 0.05 / len(COMPARISONS)
  375. train_conds = ['non_removed', 'removed', 'filled']
  376. omnibus_rows = []
  377. posthoc_rows = []
  378. print("\n" + "=" * 70)
  379. print("TRAIN CONDITION COMPARISON (cluster-level, test=filled)")
  380. print(f"Omnibus: Kruskal-Wallis | Post-hoc: Wilcoxon (Bonferroni alpha={BONF_ALPHA:.4f})")
  381. print("=" * 70)
  382. for th in ['85', '90', 'locate']:
  383. # Per-subject averages per condition
  384. cond_data = {}
  385. for tc in train_conds:
  386. cond_data[tc] = df[
  387. (df['train_condition'] == tc) &
  388. (df['test_condition'] == 'filled') &
  389. (df['threshold'] == th)
  390. ].groupby("subject")[ALL_METRICS].mean()
  391. print(f"\n --- {TH_PLOT[th]} ---")
  392. for metric in ALL_METRICS:
  393. groups = [cond_data[tc][metric].values for tc in train_conds]
  394. try:
  395. h_stat, kw_p = kruskal(*groups)
  396. except ValueError:
  397. h_stat, kw_p = 0.0, 1.0
  398. omnibus_rows.append({
  399. 'threshold': TH_PLAIN[th],
  400. 'metric': metric,
  401. 'groups': 'Non Removed, Removed, Inpainted',
  402. 'H_statistic': round(h_stat, 4),
  403. 'p_kruskal_wallis': round(kw_p, 4),
  404. 'significant_alpha_0.05': kw_p < 0.05,
  405. 'n_per_group': len(groups[0]),
  406. })
  407. if metric == 'lesion_f1':
  408. sig = "SIGNIFICANT" if kw_p < 0.05 else "n.s."
  409. print(f" Kruskal-Wallis (F1): H={h_stat:.4f}, p={kw_p:.4f} ({sig})")
  410. # Post-hoc pairwise
  411. for cond_a, cond_b, label in COMPARISONS:
  412. merged = cond_data[cond_a].merge(
  413. cond_data[cond_b], on='subject', suffixes=('_a', '_b'))
  414. vals_a = merged[f'{metric}_a'].values
  415. vals_b = merged[f'{metric}_b'].values
  416. try:
  417. _, p_w = wilcoxon(vals_a, vals_b)
  418. except ValueError:
  419. p_w = 1.0
  420. d_val, d_size = (np.nan, 'N/A')
  421. if HAS_CLIFFS:
  422. d_val, d_size = cliffs_delta(vals_b, vals_a)
  423. # Bootstrap CIs for each condition mean
  424. m_a, lo_a, hi_a = bootstrap_ci(vals_a)
  425. m_b, lo_b, hi_b = bootstrap_ci(vals_b)
  426. diff_vals = vals_b - vals_a
  427. m_d, lo_d, hi_d = bootstrap_ci(diff_vals)
  428. posthoc_rows.append({
  429. 'threshold': TH_PLAIN[th],
  430. 'comparison': label,
  431. 'metric': metric,
  432. 'mean_a': m_a,
  433. 'ci_lower_a': lo_a,
  434. 'ci_upper_a': hi_a,
  435. 'mean_b': m_b,
  436. 'ci_lower_b': lo_b,
  437. 'ci_upper_b': hi_b,
  438. 'mean_diff': m_d,
  439. 'diff_ci_lower': lo_d,
  440. 'diff_ci_upper': hi_d,
  441. 'kruskal_wallis_p': round(kw_p, 4),
  442. 'omnibus_significant': kw_p < 0.05,
  443. 'p_wilcoxon': round(p_w, 4),
  444. 'bonferroni_alpha': round(BONF_ALPHA, 4),
  445. 'significant_bonferroni': (kw_p < 0.05) and (p_w < BONF_ALPHA),
  446. 'cliffs_delta': round(d_val, 4) if not np.isnan(d_val) else np.nan,
  447. 'effect_size_category': d_size,
  448. 'n_subjects': len(merged),
  449. })
  450. # Print post-hoc F1 summary
  451. for cond_a, cond_b, label in COMPARISONS:
  452. merged = cond_data[cond_a].merge(
  453. cond_data[cond_b], on='subject', suffixes=('_a', '_b'))
  454. f1_diff = merged['lesion_f1_b'].mean() - merged['lesion_f1_a'].mean()
  455. try:
  456. _, pw = wilcoxon(merged['lesion_f1_a'].values, merged['lesion_f1_b'].values)
  457. except ValueError:
  458. pw = 1.0
  459. d_str = ""
  460. if HAS_CLIFFS:
  461. d, s = cliffs_delta(merged['lesion_f1_b'].values, merged['lesion_f1_a'].values)
  462. d_str = f" delta={d:.4f} ({s})"
  463. print(f" {label}: dF1={f1_diff:+.4f} p={pw:.4f}{d_str}")
  464. # Boxplot (Main Text Figure)
  465. filled_test = df[df['test_condition'] == 'filled']
  466. per_sub = filled_test.groupby(
  467. ['subject', 'train_condition', 'threshold']).agg({'lesion_f1': 'mean'}).reset_index()
  468. per_sub['train_label'] = per_sub['train_condition'].map(COND_LABELS)
  469. fig, axes = plt.subplots(1, 3, figsize=(15, 5), sharey=True)
  470. order = list(COND_LABELS.values())
  471. for i, th in enumerate(['85', '90', 'locate']):
  472. sub = per_sub[per_sub['threshold'] == th]
  473. sns.boxplot(data=sub, x='train_label', y='lesion_f1', order=order,
  474. hue='train_label', hue_order=order, legend=False,
  475. ax=axes[i], palette='Set2', width=0.6)
  476. axes[i].set_title(TH_PLOT[th], fontsize=12, fontweight='bold')
  477. axes[i].set_xlabel('')
  478. axes[i].set_ylim(0, 1.05)
  479. axes[i].grid(True, alpha=0.2, axis='y')
  480. axes[i].tick_params(axis='x', rotation=15)
  481. axes[0].set_ylabel('Cluster-level F1', fontsize=11)
  482. fig.tight_layout()
  483. fig.savefig(os.path.join(PLOT_DIR, "cluster_f1_condition_boxplot.png"), dpi=300)
  484. plt.close()
  485. print(" Saved: cluster_f1_condition_boxplot.png")
  486. return pd.DataFrame(omnibus_rows), pd.DataFrame(posthoc_rows)
  487. # =============================================================
  488. # 5. PER-DATASET: BeLOVE vs Challenge (Supplemental)
  489. # =============================================================
  490. def per_dataset(df):
  491. """
  492. Supplemental analysis: in-distribution (BeLOVE) vs
  493. out-of-distribution (Challenge) performance.
  494. Mann-Whitney U for unpaired comparison.
  495. """
  496. filled = df[(df['train_condition'] == 'filled') & (df['test_condition'] == 'filled')]
  497. # Descriptive
  498. desc_rows = []
  499. for ds in ['BeLOVE', 'Challenge']:
  500. dsdf = filled[filled['dataset'] == ds]
  501. n = dsdf['subject'].nunique()
  502. for th in ['85', '90', 'locate']:
  503. sub = dsdf[dsdf['threshold'] == th]
  504. per_sub = sub.groupby('subject')[ALL_METRICS].mean()
  505. row = {'dataset': ds, 'n': n, 'threshold': TH_PLAIN[th]}
  506. for m in ALL_METRICS:
  507. mean, lo, hi = bootstrap_ci(per_sub[m].values)
  508. row[f'{m}_mean'] = mean
  509. row[f'{m}_sd'] = round(per_sub[m].std(), 4)
  510. row[f'{m}_ci_lower'] = lo
  511. row[f'{m}_ci_upper'] = hi
  512. row['n_pred_mean'] = round(sub['n_pred_clusters'].mean(), 1)
  513. row['n_gt_mean'] = round(sub['n_gt_clusters'].mean(), 1)
  514. desc_rows.append(row)
  515. # Statistical comparison
  516. stat_rows = []
  517. print("\n" + "=" * 70)
  518. print("PER-DATASET (Mann-Whitney U: BeLOVE vs Challenge)")
  519. print("=" * 70)
  520. for th in ['85', '90', 'locate']:
  521. bel = filled[(filled['dataset'] == 'BeLOVE') & (filled['threshold'] == th)].groupby("subject")[ALL_METRICS].mean()
  522. cha = filled[(filled['dataset'] == 'Challenge') & (filled['threshold'] == th)].groupby("subject")[ALL_METRICS].mean()
  523. for m in ALL_METRICS:
  524. _, p = mannwhitneyu(bel[m].values, cha[m].values, alternative='two-sided')
  525. d_val, d_size = (np.nan, 'N/A')
  526. if HAS_CLIFFS:
  527. d_val, d_size = cliffs_delta(bel[m].values, cha[m].values)
  528. m_bel, lo_bel, hi_bel = bootstrap_ci(bel[m].values)
  529. m_cha, lo_cha, hi_cha = bootstrap_ci(cha[m].values)
  530. stat_rows.append({
  531. 'threshold': TH_PLAIN[th], 'metric': m,
  532. 'mean_belove': m_bel,
  533. 'ci_belove': f'[{lo_bel:.4f}, {hi_bel:.4f}]',
  534. 'mean_challenge': m_cha,
  535. 'ci_challenge': f'[{lo_cha:.4f}, {hi_cha:.4f}]',
  536. 'diff': round(m_bel - m_cha, 4),
  537. 'p_mann_whitney': round(p, 4),
  538. 'cliffs_delta': round(d_val, 4) if not np.isnan(d_val) else np.nan,
  539. 'effect_size': d_size,
  540. })
  541. if m == 'lesion_f1':
  542. print(f" {TH_PLAIN[th]} F1: BeLOVE={bel[m].mean():.4f} vs "
  543. f"Challenge={cha[m].mean():.4f} p={p:.4f}")
  544. return pd.DataFrame(desc_rows), pd.DataFrame(stat_rows)
  545. # =============================================================
  546. # 6. GT CLUSTER DISTRIBUTION (Supplemental)
  547. # =============================================================
  548. def gt_distribution(df):
  549. """Descriptive statistics of GT cluster counts (26-connectivity)."""
  550. filled = df[(df['train_condition'] == 'filled') &
  551. (df['test_condition'] == 'filled') &
  552. (df['threshold'] == '85')]
  553. gt = filled.groupby('subject')['n_gt_clusters'].mean()
  554. rows = []
  555. for label, data in [('All', gt),
  556. ('BeLOVE', filled[filled['dataset'] == 'BeLOVE'].groupby('subject')['n_gt_clusters'].mean()),
  557. ('Challenge', filled[filled['dataset'] == 'Challenge'].groupby('subject')['n_gt_clusters'].mean())]:
  558. rows.append({
  559. 'group': label, 'n': len(data),
  560. 'mean': round(data.mean(), 1), 'sd': round(data.std(), 1),
  561. 'median': round(data.median(), 1),
  562. 'min': int(data.min()), 'max': int(data.max()),
  563. 'q25': int(data.quantile(0.25)), 'q75': int(data.quantile(0.75)),
  564. })
  565. print("\n" + "=" * 70)
  566. print("GT CLUSTER DISTRIBUTION (26-connectivity)")
  567. print("=" * 70)
  568. for r in rows:
  569. print(f" {r['group']:>10s}: mean={r['mean']}, median={r['median']}, "
  570. f"range={r['min']}-{r['max']}")
  571. return pd.DataFrame(rows)
  572. # =============================================================
  573. # 7. SUMMARY TABLE (Supplemental)
  574. # =============================================================
  575. def summary_table(df):
  576. """Mean, SD, median per train x test x threshold (all 27 combos)."""
  577. metric_cols = ['dice_score', 'sensitivity', 'precision',
  578. 'lesion_f1', 'lesion_precision', 'lesion_recall',
  579. 'n_pred_clusters', 'n_gt_clusters']
  580. group_cols = ['train_condition', 'test_condition', 'threshold']
  581. summary = df.groupby(group_cols)[metric_cols].agg(['mean', 'std', 'median']).round(4)
  582. # Flatten MultiIndex columns
  583. flat = summary.reset_index()
  584. flat.columns = ['_'.join(str(c) for c in col).rstrip('_') for col in flat.columns]
  585. return flat
  586. # =============================================================
  587. # MAIN
  588. # =============================================================
  589. def main():
  590. print("=" * 70)
  591. print("CLUSTER-LEVEL ANALYSIS OF BIANCA WMH SEGMENTATION")
  592. print("=" * 70)
  593. df = load_data()
  594. # --- MAIN TEXT ---
  595. dissoc_df = voxel_vs_cluster(df)
  596. precision_recall_bars(df)
  597. omnibus_df, posthoc_df = condition_comparison(df)
  598. # --- SUPPLEMENTAL ---
  599. dataset_desc_df, dataset_stat_df = per_dataset(df)
  600. gt_df = gt_distribution(df)
  601. summary_df = summary_table(df)
  602. # =============================================================
  603. # SAVE
  604. # =============================================================
  605. print(f"\nSaving to {OUTPUT_DIR}/...")
  606. # Main Text tables
  607. save_excel(
  608. os.path.join(OUTPUT_DIR, "voxel_vs_cluster_dissociation.xlsx"),
  609. dissoc_df,
  610. "Voxel-Level vs Cluster-Level Performance Dissociation",
  611. ["Addresses R1 Comment 2: Cluster F1 substantially exceeds Dice, indicating that BIANCA detects most",
  612. "WMH clusters but delineates boundaries imprecisely. High cluster precision confirms detected regions",
  613. "are genuine WMH, not distributed false positives.",
  614. "Condition: train=filled, test=filled. Per-subject means across 10 seeds x 5-fold CV (n=89).",
  615. "Cluster parameters: 26-connectivity, any-overlap (MICCAI 2017), prediction-only MCS filtering."])
  616. save_excel(
  617. os.path.join(OUTPUT_DIR, "kruskal_wallis_train_condition.xlsx"),
  618. omnibus_df,
  619. "Kruskal-Wallis H-Test: Effect of Training Preprocessing on Cluster-Level Metrics",
  620. ["Question: Does training with Non Removed vs Removed vs Inpainted images affect lesion detection?",
  621. "Design: 3 groups (train conditions), test condition fixed to 'filled' (inpainted).",
  622. "Test: Kruskal-Wallis H-test (non-parametric omnibus, k=3 independent groups).",
  623. "Null hypothesis: All three training conditions produce equal cluster-level metrics.",
  624. "Significance: alpha = 0.05.",
  625. "Data: Per-subject means averaged across 10 seeds x 5-fold stratified CV (n=89 per group).",
  626. "If significant: see posthoc_wilcoxon_bonferroni.xlsx for pairwise comparisons."])
  627. save_excel(
  628. os.path.join(OUTPUT_DIR, "posthoc_wilcoxon_bonferroni.xlsx"),
  629. posthoc_df,
  630. "Post-hoc Wilcoxon Signed-Rank Tests with Bonferroni Correction",
  631. ["Pairwise comparisons of training conditions (paired by subject).",
  632. "Test: Wilcoxon signed-rank (non-parametric, paired).",
  633. "Correction: Bonferroni (3 comparisons, corrected alpha = 0.05/3 = 0.0167).",
  634. "Effect size: Cliff's Delta (|d| >= 0.28 = meaningful, per manuscript conventions).",
  635. "Column 'omnibus_significant': Kruskal-Wallis was significant for this metric.",
  636. "Column 'significant_bonferroni': True only if BOTH omnibus AND post-hoc significant.",
  637. "Note: Post-hoc tests interpretable only when omnibus is significant."])
  638. # Supplemental tables
  639. save_excel(
  640. os.path.join(OUTPUT_DIR, "per_dataset_descriptive.xlsx"),
  641. dataset_desc_df,
  642. "Cluster-Level Metrics by Dataset: BeLOVE vs WMH Segmentation Challenge",
  643. ["BeLOVE (n=58): in-distribution (included in cross-validated training).",
  644. "Challenge (n=31): out-of-distribution (external validation dataset).",
  645. "Condition: train=filled, test=filled."])
  646. save_excel(
  647. os.path.join(OUTPUT_DIR, "per_dataset_mann_whitney.xlsx"),
  648. dataset_stat_df,
  649. "Mann-Whitney U Test: BeLOVE vs Challenge Comparison",
  650. ["Test: Mann-Whitney U (non-parametric, unpaired, 2 groups).",
  651. "Significance: alpha = 0.05.",
  652. "Effect size: Cliff's Delta.",
  653. "Per-subject means (BeLOVE n=58, Challenge n=31)."])
  654. save_excel(
  655. os.path.join(OUTPUT_DIR, "gt_cluster_distribution.xlsx"),
  656. gt_df,
  657. "Ground Truth Cluster Count Distribution (26-Connectivity)",
  658. ["Descriptive statistics of the number of GT clusters per subject.",
  659. "26-connectivity used (consistent with all cluster-level analyses).",
  660. "No size filtering applied to ground truth (expert-delineated lesions)."])
  661. save_excel(
  662. os.path.join(OUTPUT_DIR, "summary_all_conditions.xlsx"),
  663. summary_df,
  664. "Full Summary: Mean, SD, Median per Train x Test x Threshold",
  665. ["All 3 train x 3 test x 3 threshold = 27 combinations.",
  666. "Statistics: mean, standard deviation, median for each metric.",
  667. "Averaged across 10 seeds x 5-fold stratified CV x 89 subjects."])
  668. # Raw data
  669. df.to_excel(os.path.join(OUTPUT_DIR, "all_seeds_raw.xlsx"), index=False)
  670. print(f"\nDone. Results: {OUTPUT_DIR}/ Plots: {PLOT_DIR}/")
  671. if __name__ == "__main__":
  672. main()

4_Analyze_cluster_metrics.py at commit 1f27300, no license · at the source

Overview

Authors: Uchralt Temuulen1, Ralf Mekle1, Ivana Galinovic1, Joachim E Weber1,2,3,4, Ulf Landmesser2,3,5, Sebastian Kelle3, Matthias Endres1,3,4,6,7, Christian Stehning8, Kersten Villringer1
  1. Center for Stroke Research Berlin, Charité-Universitätsmedizin Berlin, Corporate Member of Freie Universität Berlin and Humboldt-Universität zu Berlin, Berlin, Germany
  2. Berlin Institute of Health (BIH) at Charité-Universitätsmedizin Berlin, Berlin, Germany
  3. German Centre for Cardiovascular Research (DZHK), Partner site Berlin, Berlin, Germany
  4. Department of Neurology, Charité-Universitätsmedizin Berlin, Berlin, Germany
  5. Department of Cardiology, Charité-Universitätsmedizin Berlin, Berlin, Germany
  6. German Center for Neurodegenerative Diseases (DZNE), Partner site Berlin, Berlin, Germany
  7. German Center for Mental Health (DZPG), Partner site Berlin, Berlin, Germany
  8. Philips Clinical Science, Hamburg, Germany
Journal: NeuroImage. Clinical, volume 50, article 104001
Dates: received 15 January 2026; accepted 1 May 2026; published online 11 May 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.nicl.2026.104001 · PMID 42119523 · PMCID PMC13195355 · OpenAlex W7161065280
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), stroke (population), methods / tools (subfield)
Methods: Statistics, Machine learning, Connectivity, Single-unit activity, calcium imaging
Keywords: White matterhyperintensities, BIANCA, LOCATE, Automated segmentation, Stroke, Cerebral small vessel disease, FLAIR
MeSH: Algorithms*, Brain*, Image Interpretation, Computer-Assisted*, Image Processing, Computer-Assisted*, Magnetic Resonance Imaging*, Stroke*, White Matter*, Female, Humans, Male (* major topic)
Topic: Acute Ischemic Stroke Management (Epidemiology, Medicine), according to OpenAlex
Funding: German Research Foundation; BMBF Bonn; Deutsches Zentrum für Herz-Kreislauf-Forschung eV
Citations: not cited yet (Europe PMC); 60 references in the paper

Abstract

Background: White matter hyperintensity (WMH) segmentation using BIANCA (Brain Intensity AbNormality Classification Algorithm) in stroke populations is complicated by vascular lesions that share T2-hyperintense signal characteristics with WMH. Whether preprocessing decisions in the treatment of lesions affect segmentation accuracy has not yet been systematically evaluated in cerebrovascular cohorts involving multiple scanners.

Methods: We compared fixed probability thresholds and locally adaptive thresholding with LOCATE (LOCally Adaptive Threshold Estimation), and three lesion-handling approaches: lesions present (non removed), replaced with zero intensities (removed), and replaced with normal-appearing white matter intensities (NAWM; inpainted), using the BeLOVE cohort (Berlin Longterm Observation of Vascular Events) and the WMH Segmentation Challenge dataset. Phase I (n = 89) optimized thresholding via stratified 5-fold cross-validation. Phase II assessed preprocessing effects on segmentation accuracy (Phase II-A, n = 89) and volume agreement (Phase II-B, n = 211).

Results: BIANCA with LOCATE adaptive thresholding achieved moderate segmentation overlap (mean Dice 0.567), with lesion-level detection exceeding an F1 of 0.85. Preprocessing effects were statistically detectable but negligible in magnitude, with near-perfect agreement between all conditions. Stroke lesion volume was the highest-ranked predictor of volume differences between conditions; these scaled with lesion size but remained negligible in magnitude across all subgroups.

Conclusions: BIANCA with LOCATE achieved moderate WMH segmentation performance with the best sensitivity-precision trade-off in this multi-scanner cerebrovascular cohort. Preprocessing effects were negligible at the group level. However, large lesions distort the FLAIR intensity distribution on which BIANCA relies for classification, which justifies lesion removal as a recommended preprocessing step.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 28 matches between paragraphs and lines of code.

mendeltem/Bianca_Evaluation

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 1f27300aa772a82509608f194a7558f4d79b4b70, 20 March 2026
Languages: Python (33), Shell (1)
Size: 40 files, 34 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: pandas (32 files), NumPy (31 files), SciPy (21 files), Matplotlib (20 files), FSL (7 files), Pillow (6 files), scikit-learn (6 files), seaborn (4 files), NiBabel (3 files), SHAP (2 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
35 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 34 scripts, each with its path and the digest of its content;
  • 28 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

Analysis scripts for WMH segmentation, preprocessing, statistical testing and figure generation are available at https://github.com/mendeltem/Bianca_Evaluation under the MIT License.

Source imaging data cannot be shared publicly due to participant privacy restrictions under the European General Data Protection Regulation (GDPR). Processed, de-identified derivative data can be obtained from the corresponding author upon reasonable request and completion of appropriate data sharing agreements with the Center for Stroke Research Berlin.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 9 authors, 7 keywords, 10 MeSH terms, 3 funders, 47 references.

Cite

This paper

Temuulen, U., Mekle, R., Galinovic, I., Weber, J. E., Landmesser, U., Kelle, S., Endres, M., Stehning, C., & Villringer, K. (2026). Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation. NeuroImage. Clinical, 50, 104001. https://doi.org/10.1016/j.nicl.2026.104001

BibTeX

@article{temuulen2026effect,
author = {Temuulen, Uchralt and Mekle, Ralf and Galinovic, Ivana and Weber, Joachim E and Landmesser, Ulf and Kelle, Sebastian and Endres, Matthias and Stehning, Christian and Villringer, Kersten},
title = {{Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation}},
journal = {NeuroImage. Clinical},
year = {2026},
month = may,
volume = {50},
pages = {104001},
publisher = {Elsevier},
issn = {2213-1582},
doi = {10.1016/j.nicl.2026.104001},
url = {https://doi.org/10.1016/j.nicl.2026.104001},
pmid = {42119523},
pmcid = {PMC13195355}
}

RIS

TY - JOUR
AU - Temuulen, Uchralt
AU - Mekle, Ralf
AU - Galinovic, Ivana
AU - Weber, Joachim E
AU - Landmesser, Ulf
AU - Kelle, Sebastian
AU - Endres, Matthias
AU - Stehning, Christian
AU - Villringer, Kersten
TI - Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation
T2 - NeuroImage. Clinical
J2 - Neuroimage Clin
PY - 2026
DA - 2026/05/11
VL - 50
SP - 104001
SN - 2213-1582
PB - Elsevier
DO - 10.1016/j.nicl.2026.104001
UR - https://doi.org/10.1016/j.nicl.2026.104001
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.nicl.2026.104001",
"type": "article-journal",
"title": "Effect of vascular lesion preprocessing on Brain Intensity AbNormality Classification Algorithm (BIANCA) white matter hyperintensity segmentation",
"container-title": "NeuroImage. Clinical",
"author": [
{
"family": "Temuulen",
"given": "Uchralt"
},
{
"family": "Mekle",
"given": "Ralf"
},
{
"family": "Galinovic",
"given": "Ivana"
},
{
"family": "Weber",
"given": "Joachim E"
},
{
"family": "Landmesser",
"given": "Ulf"
},
{
"family": "Kelle",
"given": "Sebastian"
},
{
"family": "Endres",
"given": "Matthias"
},
{
"family": "Stehning",
"given": "Christian"
},
{
"family": "Villringer",
"given": "Kersten"
}
],
"container-title-short": "Neuroimage Clin",
"volume": "50",
"page": "104001",
"DOI": "10.1016/j.nicl.2026.104001",
"PMID": "42119523",
"PMCID": "PMC13195355",
"ISSN": "2213-1582",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.nicl.2026.104001",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
11
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s12938-026-01555-0 [code]
Incorporating normal periventricular changes for enhanced pathological white matter hyperintensity segmentation: on multiclass deep learning approaches.
Journal: Biomedical engineering online
In common: seaborn, scikit-learn, pandas, 3 other tools, stroke, methods / tools, structural MRI / diffusion, 5 references
[2] doi:10.1212/wnl.0000000000218472 [code]
Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.
Journal: Neurology
In common: SHAP, NiBabel, scikit-learn, 4 other tools, stroke, structural MRI / diffusion, 3 references
[3] doi:10.1111/ene.70678 [code]
Who Falls After a Stroke? Evidence From a Prospective Stroke Cohort.
Journal: European journal of neurology
In common: NiBabel, seaborn, pandas, 3 other tools, stroke, 1 reference, author Matthias Endres
[4] doi:10.1371/journal.pbio.3003856 [code]
Aging and metabolism contribute separately to brain-body health.
Journal: PLoS biology
In common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 3 references
[5] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: SHAP, FSL, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
[6] doi:10.1162/imag.a.1252 [code]
Does the brain's E:I balance really shape long-range temporal correlations? Lessons learned from 3T MRI.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 3 references
[7] doi:10.1002/hbm.70602 [code]
Neuroimaging Correlates of Post-Stroke Pain After Ischemic Stroke: Secondary Analysis of the INSPiRE-TMS Trial.
Journal: Human brain mapping
In common: NiBabel, seaborn, pandas, 3 other tools, stroke, author Matthias Endres
[8] doi:10.1186/s12880-026-02481-2 [code]
Deep learning-based neuroanatomical profiling reveals population-specific brain changes in multiple sclerosis: a large-scale Middle Eastern study.
Journal: BMC medical imaging
In common: FSL, Pillow, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
[9] doi:10.1002/ana.78203 [code]
AI-Driven Mapping of Seizure Spread Patterns.
Journal: Annals of neurology
In common: FSL, Pillow, NiBabel, 6 other tools, structural MRI / diffusion, 1 reference
[10] doi:10.3389/fnins.2026.1870124 [code]
An end-to-end pipeline for automated fetal brain segmentation and biometry from 3D SSFP MRI.
Journal: Frontiers in neuroscience
In common: FSL, NiBabel, seaborn, 5 other tools, structural MRI / diffusion, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.