OSCR

Data quality biases normative models derived from fetal brain MRI.

Code ↔ Paper

8 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 8 matches
  1. [1] § Methods › Data processing › Segmentation ↔ scripts/plots_and_analyses/supp_plot_iterative_removal.r, lines 1–80 · score 0.76 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, brainstem, thalamus
  2. [2] § Methods › Data processing › Segmentation ↔ scripts/plots_and_analyses/1_plot_main.r, lines 1–39 · score 0.75 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, brainstem, thalamus
  3. [3] § Results › Poor-quality images can bias normative trajectories ↔ notebooks/1_harmonize_data.ipynb, lines 539–558 · score 0.74 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, cerebellar volume, brainstem
  4. [4] § Results › Poor-quality images can bias normative trajectories ↔ scripts/plots_and_analyses/1_plot_main.r, lines 1–39 · score 0.74 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, cerebellar volume, brainstem
  5. [5] § Methods › Data processing › Super-resolution reconstruction ↔ fetpype/pipelines/full_pipeline.py, lines 26–104 · score 0.65 · bias field correction, brain extraction, denoising, resolution, preprocessing, stacks
  6. [6] § Results › Poor-quality images can bias normative trajectories ↔ scripts/plots_and_analyses/supp_plot_iterative_removal.r, lines 1–80 · score 0.60 · lateral ventricles volume, white matter volume, cerebellar volume, confidence, bootstrap, centile
  7. [7] § Methods › Analysis of the impact of image quality on normative models › Harmonization on volumetric measurements ↔ notebooks/1_harmonize_data.ipynb, lines 197–283 · score 0.58 · ComBat, GroundTruth, GAM, smoothing, spline, harmonized
  8. [8] § Methods › Analysis of the impact of image quality on normative models › QC subgroups definition ↔ notebooks/1_harmonize_data.ipynb, lines 36–86 · score 0.50 · global quality score, excellent quality, cohorts, scans

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 814 lines · 28 KB · no license · 3 matches

  1. # %% [markdown]
  2. # # Harmonization and data description
  3. # This python notebook loads the raw data, plots some demographics, runs the harmonization and saves the harmonized CSV files to be used in the rest of the analyses.
  4. #
  5. # Here is how the notebook is organized.
  6. #
  7. # 1. **Demographics plotting:** Producing basic descriptive figures of the data (Figure 3)
  8. # 2. **Harmonization:** Run and save the harmonization
  9. # 3. **Detrending analysis:** Plot and analyze the influence of data quality on the raw trajectories (Figure 4)
  10. # %%
  11. import pandas as pd
  12. import numpy as np
  13. import warnings
  14. import matplotlib.pyplot as plt
  15. import seaborn as sns
  16. from sklearn.preprocessing import StandardScaler
  17. import os
  18. import scipy
  19. from neuroHarmonize import harmonizationLearn, harmonizationApply
  20. from sklearn.metrics import r2_score, mean_squared_error
  21. warnings.filterwarnings("ignore")
  22. %load_ext autoreload
  23. %autoreload 2
  24. # %%
  25. input_file = "../data/raw_data/finalized_df_qcglobal_and_z.csv"
  26. output_dir = "../data/harmonization"
  27. plot_dir = output_dir
  28. os.makedirs(output_dir, exist_ok=True)
  29. data_df = pd.read_csv(input_file)
  30. # %% [markdown]
  31. # ## 1. Demographics plotting
  32. # %%
  33. f, ax = plt.subplots(2, figsize=(8, 5), gridspec_kw={"height_ratios": [1, 3]}, sharex=True)
  34. plt.subplots_adjust(hspace=0.01)
  35. n_excl = len(data_df[data_df["qcglobal"] < 0.95])
  36. n_poor = len(data_df[(data_df["qcglobal"] >= 0.95) & (data_df["qcglobal"] < 1.95)])
  37. n_acc = len(data_df[(data_df["qcglobal"] >= 1.95) & (data_df["qcglobal"] < 2.7)])
  38. n_excellent = len(data_df[data_df["qcglobal"] >= 2.7])
  39. print(f"Number of excluded scans (qcglobal < 0.95): {n_excl}")
  40. print(f"Number of poor quality scans (0.95 <= qcglobal < 1.95): {n_poor}")
  41. print(f"Number of acceptable quality scans (1.95 <= qcglobal < 2.7): {n_acc}")
  42. print(f"Number of excellent quality scans (qcglobal >= 2.7): {n_excellent}")
  43. for i, a in enumerate(ax):
  44. if i==0:
  45. max_ = len(data_df["cohort"].unique())
  46. else:
  47. max_ = 65
  48. a.axvline(x=0.95, color="black", linestyle="--")
  49. a.axvline(x=1.95, color="black", linestyle="--")
  50. a.axvline(x=2.7, color="black", linestyle="--")
  51. a.fill_betweenx(y=[-1, max_], x1=0, x2=0.95, color="lightgray", alpha=0.5)
  52. a.fill_betweenx(y=[-1, max_], x1=0.95, x2=1.95, color="gray", alpha=0.5)
  53. a.fill_betweenx(y=[-1, max_], x1=1.95, x2=2.7, color="lightgray", alpha=0.5)
  54. a.fill_betweenx(y=[-1, max_], x1=2.7, x2=4, color="gray", alpha=0.5)
  55. # Select bins to finish at 1, 2, 2.7
  56. bins = np.arange(0, data_df["qcglobal"].max()+0.1, 0.1)
  57. sns.histplot(data=data_df, ax=ax[1], x="qcglobal", hue="cohort", multiple="stack", bins=bins)
  58. # Add ontop a kind of boxplot to show quartiles for each cohort, it's a second plot with same x axis
  59. sns.boxplot(data=data_df, ax=ax[0], x="qcglobal", y="cohort", hue="cohort")
  60. # For ax[0], remove xlabel and ylabel and xticks
  61. ax[0].set_xlabel("")
  62. ax[0].set_ylabel("")
  63. ax[1].set_xlabel("Global quality score")
  64. ax[0].set_xticks([])
  65. ax[0].set_yticks([])
  66. ax[1].set_ylim(0, 65)
  67. ax[1].set_xlim(0,4)
  68. ax[0].set_xlim(0,4)
  69. ax[1].annotate("Exclude Quality", xy=(0.475, max_), xytext=(0.475, max_-5), fontsize=12, color="black", ha="center")
  70. ax[1].annotate(f"(n={n_excl})", xy=(0.475, max_-10), xytext=(0.475, max_-8), fontsize=10, color="black", ha="center")
  71. ax[1].annotate("Poor Quality", xy=(1.45, max_), xytext=(1.45, max_-5), fontsize=12, color="black", ha="center")
  72. ax[1].annotate(f"(n={n_poor})", xy=(1.45, max_-10), xytext=(1.45, max_-9), fontsize=10, color="black", ha="center")
  73. ax[1].annotate("Acceptable\nQuality", xy=(2.325, max_), xytext=(2.325, max_-9), fontsize=12, color="black", ha="center")
  74. ax[1].annotate(f"(n={n_acc})", xy=(2.325, max_-10), xytext=(2.325, max_-12.5), fontsize=10, color="black", ha="center")
  75. ax[1].annotate("Excellent Quality", xy=(4, max_), xytext=(3.3, max_-5), fontsize=12, color="black", ha="center")
  76. ax[1].annotate(f"(n={n_excellent})", xy=(4, max_-10), xytext=(3.3, max_-8), fontsize=10, color="black", ha="center")
  77. # Across the two plots, add a vertical line at qcglobal = 0.95, one at 1.95 and one at 2.7 and fill the background between 0 and 0.95 with lightgray, between 0.95 and 1.95 with gray, between 1.95 and 2.7 with darkgray, and above 2.7 with black
  78. # %%
  79. f, ax = plt.subplots(2, figsize=(5, 6), gridspec_kw={"height_ratios": [1, 3]}, sharex=True)
  80. plt.subplots_adjust(hspace=0.01)
  81. n_excl = len(data_df[data_df["qcglobal"] < 0.95])
  82. n_poor = len(data_df[(data_df["qcglobal"] >= 0.95) & (data_df["qcglobal"] < 1.95)])
  83. n_acc = len(data_df[(data_df["qcglobal"] >= 1.95) & (data_df["qcglobal"] < 2.7)])
  84. n_excellent = len(data_df[data_df["qcglobal"] >= 2.7])
  85. print(f"Number of excluded scans (qcglobal < 0.95): {n_excl}")
  86. print(f"Number of poor quality scans (0.95 <= qcglobal < 1.95): {n_poor}")
  87. print(f"Number of acceptable quality scans (1.95 <= qcglobal < 2.7): {n_acc}")
  88. print(f"Number of excellent quality scans (qcglobal >= 2.7): {n_excellent}")
  89. # Select bins to finish at 1, 2, 2.7
  90. bins = np.arange(20,39, 1)
  91. sns.histplot(data=data_df, ax=ax[1], x="age_at_scan", hue="cohort", multiple="stack", bins=bins)
  92. # Add ontop a kind of boxplot to show quartiles for each cohort, it's a second plot with same x axis
  93. sns.boxplot(data=data_df, ax=ax[0], x="age_at_scan", y="cohort", hue="cohort")
  94. # For ax[0], remove xlabel and ylabel and xticks
  95. ax[0].set_xlabel("")
  96. ax[0].set_ylabel("")
  97. ax[1].set_xlabel("Gestational age (weeks)")
  98. # Remove the lengends from ax[0]
  99. ax[1].get_legend().remove()
  100. # %%
  101. data_df['age_bin'] = pd.cut(data_df['age_at_scan'], bins=np.arange(20,40,2))
  102. plt.figure(figsize=(6,6))
  103. # Global boxplot (light gray behind)
  104. sns.boxplot(
  105. data=data_df,
  106. x='age_bin',
  107. y='qcglobal',
  108. color='white',
  109. showfliers=False,
  110. boxprops=dict(facecolor='lightgray', edgecolor='black', linewidth=1.5,alpha=0.7),
  111. whiskerprops=dict(color='black', linewidth=1.5),
  112. capprops=dict(color='black', linewidth=0.),
  113. medianprops=dict(color='black', linewidth=1.5)
  114. )
  115. # Overlay cohort-specific boxplots with transparency
  116. sns.boxplot(
  117. data=data_df,
  118. x='age_bin',
  119. y='qcglobal',
  120. hue='cohort',
  121. showfliers=False,
  122. dodge=True,
  123. palette='tab10',
  124. linewidth=0.5,
  125. boxprops=dict(alpha=1.)
  126. )
  127. plt.xticks(rotation=45)
  128. plt.xlabel("Age at scan (binned)")
  129. plt.ylabel("Global quality score")
  130. plt.xticks(fontsize=14)
  131. plt.yticks(fontsize=14)
  132. plt.ylabel("Global quality score", fontsize=14)
  133. plt.xlabel("Gestational age (weeks)", fontsize=14)
  134. sns.set_style("ticks") # adds small ticks
  135. sns.set_context("talk", font_scale=1.1)
  136. plt.legend(title="Cohort", bbox_to_anchor=(1.05, 1), loc="upper left", fontsize=10)
  137. plt.tight_layout()
  138. plt.show()
  139. # %% [markdown]
  140. # ## 2. Harmonization
  141. # %%
  142. ## Set up thresholds for different quality groups
  143. data_df["all"] = data_df["qcglobal"] >= 0
  144. data_df["poor_plus"] = data_df["qcglobal"] >= 0.95
  145. data_df["accept_plus"] = data_df["qcglobal"] >= 1.95
  146. data_df["ground_truth"] = data_df["qcglobal"] >= 2.7
  147. data_df["age_rounded"] = round(data_df["age_at_scan"], 2)
  148. data_df["qc_group"] = pd.Categorical(
  149. np.where(data_df["qcglobal"] < 0.95, "excluded",
  150. np.where(data_df["qcglobal"] < 1.95, "poor",
  151. np.where(data_df["qcglobal"] < 2.7, "acceptable", "ground truth")))
  152. )
  153. data_df['batch'] = pd.Categorical(data_df["cohort"]).codes.astype(str) #convert cohort variable to BATCH variable for analysis. Can change to "scanner"
  154. cohort_match = dict(data_df[["batch", "cohort"]].drop_duplicates().values)
  155. # %%
  156. ## Set up the renaming of the different variables.
  157. brain_cols = [
  158. 'CSF.Volume', 'cGrey.Matter.Volume', 'Lateral.Ventricles.Volume', "Brainstem.Volume", 'Cerebellum.Volume', 'Basal.Ganglia.Volume', 'Thalamus.Volume', 'White.Matter.Volume', 'CSF_total', 'eCSF_Ventricles_total'
  159. ]
  160. cols_name = [
  161. "eCSF (no ventricles) Volume",
  162. "Cortical Gray Matter Volume",
  163. "Lateral Ventricles Volume",
  164. "Brainstem Volume",
  165. "Cerebellum Volume",
  166. "Basal Ganglia Volume",
  167. "Thalamus Volume",
  168. "White Matter Volume",
  169. "CSF total Volume",
  170. "eCSF Volume"
  171. ]
  172. # %%
  173. dfs = []
  174. # QC category order + display labels (match R: qc_order <- c("GroundTruth", "AcceptPlus", "PoorPlus", "All"))
  175. qc_order = ["GroundTruth", "AcceptPlus", "PoorPlus", "All"]
  176. categories = ["ground_truth", "accept_plus", "poor_plus", "all"]
  177. category_display = dict(zip(categories, qc_order))
  178. trajs = []
  179. ga_by_category = {}
  180. def get_covars(df):
  181. covars = pd.DataFrame({
  182. #"ICV": df["ICV"],
  183. "age": df["age_at_scan"],
  184. "SITE":df["batch"]
  185. })
  186. return covars
  187. def predict_gam_on_grid(model, ga_grid):
  188. smooth_model = model['smooth_model']
  189. bs = smooth_model['bsplines_constructor']
  190. B_hat = model['B_hat']
  191. df_gam_train = smooth_model['df_gam']
  192. n_grid = len(ga_grid)
  193. site_cols = [c for c in df_gam_train.columns if c.startswith('x')]
  194. cov_cols = [c for c in df_gam_train.columns if c.startswith('c')]
  195. param_cols = site_cols + cov_cols
  196. X_param = np.zeros((n_grid, len(param_cols)))
  197. for j, col in enumerate(param_cols):
  198. if col in cov_cols:
  199. X_param[:, j] = df_gam_train[col].mean()
  200. else:
  201. X_param[:, j] = 0
  202. X_spline = bs.transform(ga_grid.reshape(-1, 1))
  203. X_new = np.concatenate([X_param, X_spline], axis=1)
  204. return X_new @ B_hat + model["stand_mean"][:,0] # (n_grid, n_features)
  205. y_preds = []
  206. models = []
  207. site_trajectories = []
  208. ga_grid = np.linspace(22, 38, 100) # adjust to your GA range
  209. for d in categories:
  210. print(f"Harmonizing for {d} subgroup")
  211. df2 = data_df.copy()
  212. harmonized_df = df2.copy()
  213. data_df2 = df2.copy()
  214. harmo_train = data_df2[data_df2[d] == True] #change here if want to do by scanner
  215. data_train = data_df2[data_df2[d] == True]
  216. data_test = data_df2[data_df2[d] == False]
  217. print(len(data_train), len(data_test))
  218. covars_train = get_covars(data_train)
  219. covars_test = get_covars(data_test)
  220. model, data_train_harmo = harmonizationLearn(np.array(harmo_train[brain_cols]), covars_train, smooth_terms=["age"])
  221. harmonized_df.loc[harmonized_df[d] == True, [c+"_harm" for c in brain_cols]] = data_train_harmo
  222. if len(data_test) > 0:
  223. data_test_harmo = harmonizationApply(np.array(data_test[brain_cols]), covars_test,model)
  224. harmonized_df.loc[harmonized_df[d] == False, [c+"_harm" for c in brain_cols]] = data_test_harmo
  225. y_pred = predict_gam_on_grid(model, ga_grid)
  226. grand_mean = model['stand_mean'][:, 0] # (n_features,)
  227. # The biological trajectory (same for all sites)
  228. trajectory = y_pred # (n_grid, n_features)
  229. # Site-specific shift in natural units
  230. gamma_natural = model['gamma_star'] * np.sqrt(model['var_pooled']).T # (n_sites, n_features)
  231. # Full signal per site across GA grid
  232. site_dict = {}
  233. for i, site in enumerate(model['SITE_labels']):
  234. site_trajectory = trajectory - gamma_natural[i, :] # (n_grid, n_features)
  235. site = cohort_match[str(site)]
  236. site_dict[site] = site_trajectory
  237. site_trajectories.append(site_dict)
  238. y_preds.append(y_pred)
  239. models.append(model)
  240. dfs.append(harmonized_df)
  241. harmonized_df.to_csv(os.path.join(output_dir, f"harmonized_data_{d}.csv"), index=False)
  242. # %% [markdown]
  243. # ## 3. Detrending analysis
  244. # %%
  245. sns.set_theme(style="whitegrid")
  246. import matplotlib.colors as mcolors
  247. import matplotlib as mpl
  248. import matplotlib.cm as mcm
  249. from matplotlib.lines import Line2D
  250. # ---- Options ----
  251. show_fit = True # set to False to hide the spline fit line
  252. #brain_cols = ['eCSF_Ventricles_total', 'cGrey.Matter.Volume', 'Basal.Ganglia.Volume', "Brainstem.Volume", 'Thalamus.Volume' ]
  253. #cols_name = ["eCSF Volume", "Cortical Gray Matter Volume", "Basal Ganglia Volume", "Brainstem Volume", "Thalamus Volume"]
  254. n_features = len(brain_cols)
  255. fig, axes = plt.subplots(
  256. n_features, 3,
  257. figsize=(13, n_features * 2.8),
  258. gridspec_kw={"width_ratios": [3, 3, 1]},
  259. )
  260. # ---------- Manual axis sharing ----------
  261. # All GA panels (cols 0 & 1) share the same x-axis
  262. for i in range(1, n_features):
  263. axes[i, 0].sharex(axes[0, 0])
  264. axes[i, 1].sharex(axes[0, 0])
  265. axes[0, 1].sharex(axes[0, 0])
  266. # QC marginal panels (col 2) share x among themselves
  267. for i in range(1, n_features):
  268. axes[i, 2].sharex(axes[0, 2])
  269. # col 1 (detrended) and col 2 (QC marginal) are NOT y-linked:
  270. # each detrended panel keeps its own independent y-axis.
  271. # Hide x tick labels on all but the bottom row
  272. for i in range(n_features - 1):
  273. axes[i, 0].tick_params(labelbottom=False)
  274. axes[i, 1].tick_params(labelbottom=False)
  275. axes[i, 2].tick_params(labelbottom=False)
  276. # ---------- Column titles ----------
  277. axes[0, 0].set_title("Volume [mm3]", fontsize=11, fontweight="bold")
  278. axes[0, 1].set_title("Detrended Volume [mm3]", fontsize=11, fontweight="bold")
  279. # ---------- Color / marker setup ----------
  280. qc_min = harmonized_df["qcglobal"].min()
  281. qc_max = harmonized_df["qcglobal"].max()
  282. norm = mcolors.Normalize(vmin=qc_min, vmax=qc_max)
  283. cmap = mpl.colormaps["viridis_r"]
  284. cohorts = sorted(harmonized_df["cohort"].unique())
  285. _marker_cycle = ["o", "s", "^", "D", "v", "P", "X", "h", "*", "+"]
  286. cohort_marker = {c: _marker_cycle[i % len(_marker_cycle)] for i, c in enumerate(cohorts)}
  287. # ---------- Per-feature rows ----------
  288. hdr = f"{'Feature':<40} {'beta_lin':>10} {'R2_lin':>8} {'p_lin':>12} {'a (quad)':>10} {'b (quad)':>10} {'R2_quad':>8}"
  289. print(hdr)
  290. print("-" * len(hdr))
  291. for idx in range(n_features):
  292. df2 = harmonized_df.copy()
  293. col = brain_cols[idx]
  294. name = cols_name[idx]
  295. ax_vol = axes[idx, 0]
  296. ax_det = axes[idx, 1]
  297. ax_qc = axes[idx, 2]
  298. # ---- Spline fit on high-QC subset ----
  299. df_sel = df2[df2["qcglobal"] >= 2.75] #2.75
  300. x_hq = df_sel["age_at_scan"].values
  301. y_hq = df_sel[col].values
  302. order = np.argsort(x_hq)
  303. x_sorted = x_hq[order]; y_sorted = y_hq[order]
  304. xl, yl = [], []
  305. for i in range(len(x_sorted)):
  306. if i == 0 or xl[-1] != x_sorted[i]:
  307. xl.append(x_sorted[i]); yl.append([y_sorted[i]])
  308. else:
  309. yl[-1].append(y_sorted[i])
  310. yl = [np.mean(v) for v in yl]
  311. spline = scipy.interpolate.make_smoothing_spline(xl, yl)
  312. x_fit_s = np.linspace(np.min(xl), np.max(xl), 400)
  313. y_fit_s = spline(x_fit_s)
  314. df2["spl_trend"] = spline(df2["age_at_scan"])
  315. df2["residual"] = df2[col] - df2["spl_trend"]
  316. # ---- Scatter: one call per cohort so markers differ ----
  317. for cohort in cohorts:
  318. mask = df2["cohort"] == cohort
  319. subset = df2[mask]
  320. colors = cmap(norm(subset["qcglobal"].values))
  321. kw = dict(c=colors, marker=cohort_marker[cohort],
  322. alpha=0.7, s=20, linewidths=0)
  323. ax_vol.scatter(subset["age_at_scan"], subset[col], **kw)
  324. ax_det.scatter(subset["age_at_scan"], subset["residual"], **kw)
  325. if show_fit:
  326. ax_vol.plot(x_fit_s, y_fit_s, color="black", lw=1.5)
  327. ax_vol.set_ylabel(name, fontsize=11)
  328. ax_det.set_ylabel("")
  329. vol_lo = df2[col].quantile(0.01)
  330. vol_hi = df2[col].quantile(0.99)
  331. ymin = max(0, vol_lo*0.8)
  332. ax_vol.set_ylim(ymin, vol_hi)
  333. ax_det.axhline(0, color="black", lw=1, ls="--")
  334. resid_lo = df2["residual"].quantile(0.01)
  335. resid_hi = df2["residual"].quantile(0.99)
  336. ax_det.set_ylim(resid_lo, resid_hi)
  337. # ---- Marginal QC panel ----
  338. n_bins = 20
  339. bins = np.linspace(resid_lo, resid_hi, n_bins + 1)
  340. centers = 0.5 * (bins[:-1] + bins[1:])
  341. df2["resid_bin"] = pd.cut(df2["residual"], bins)
  342. qc_mean = df2.groupby("resid_bin", observed=False)["qcglobal"].mean()
  343. qc_std = df2.groupby("resid_bin", observed=False)["qcglobal"].std()
  344. # Linear + quadratic regression (printed only, not drawn)
  345. reg_df = df2[["residual", "qcglobal"]].dropna()
  346. reg_mask = np.isfinite(reg_df["residual"]) & np.isfinite(reg_df["qcglobal"])
  347. reg_df = reg_df.loc[reg_mask]
  348. x_obs = reg_df["residual"].to_numpy()
  349. y_obs = reg_df["qcglobal"].to_numpy()
  350. lr = scipy.stats.linregress(x_obs, y_obs)
  351. pred = lr.slope * x_obs + lr.intercept
  352. a2, b1, c0 = np.polyfit(x_obs, y_obs, deg=2)
  353. y_hat_quad = a2 * x_obs**2 + b1 * x_obs + c0
  354. r2_quad = r2_score(y_obs, y_hat_quad)
  355. print(f"{name:<40} {lr.slope:>10.4f} {lr.rvalue**2:>8.4f} {lr.pvalue:>12.4e} "
  356. f"{a2:>10.4f} {b1:>10.4f} {r2_quad:>8.4f}")
  357. # Mean \u00b1 SD only (no regression line overlay)
  358. ax_qc.plot(qc_mean, centers, color="black", lw=2)
  359. ax_qc.fill_betweenx(centers, qc_mean - qc_std, qc_mean + qc_std,
  360. color="gray", alpha=0.3)
  361. ax_qc.axhline(0, color="black", lw=2, ls="--")
  362. ax_qc.set_xlim(qc_min, qc_max)
  363. ax_qc.set_ylim(resid_lo, resid_hi)
  364. ax_qc.set_yticks([])
  365. ax_qc.tick_params(labelleft=False)
  366. ax_qc.grid(True, axis="x", alpha=0.2)
  367. # ---------- x-labels on bottom row ----------
  368. axes[-1, 0].set_xlabel("GA (weeks)", fontsize=10)
  369. axes[-1, 1].set_xlabel("GA (weeks)", fontsize=10)
  370. axes[-1, 2].set_xlabel("QC", fontsize=10)
  371. # ---------- Reserve space: right margin for colorbar, bottom for legend ----------
  372. plt.tight_layout(rect=[0, 0.03, 0.8, 1])
  373. # ---------- Colorbar ----------
  374. cbar_ax = fig.add_axes([0.87, 0.38, 0.018, 0.25])
  375. sm = mcm.ScalarMappable(cmap=cmap, norm=norm)
  376. sm.set_array([])
  377. cbar = fig.colorbar(sm, cax=cbar_ax)
  378. cbar.set_label("QC Global", rotation=270, labelpad=14, fontsize=10)
  379. # ---------- Site / marker legend \u2014 horizontal, below the plot ----------
  380. legend_handles = [
  381. Line2D(
  382. [0], [0],
  383. marker=cohort_marker[c], color="w",
  384. markerfacecolor="gray", markeredgecolor="gray",
  385. markersize=7, linestyle="None",
  386. label=c,
  387. )
  388. for c in cohorts
  389. ]
  390. fig.legend(
  391. handles=legend_handles,
  392. title="Site",
  393. loc="lower center",
  394. bbox_to_anchor=(0.42, 0.0),
  395. ncol=len(cohorts),
  396. fontsize=8,
  397. title_fontsize=9,
  398. framealpha=0.9,
  399. )
  400. # ---------- Save high-resolution ----------
  401. save_path = os.path.join(plot_dir, "qc_residual_scatter_linear.svg")
  402. fig.savefig(save_path, dpi=300, bbox_inches="tight")
  403. print(f"\nFigure saved to: {save_path}")
  404. plt.show()
  405. # %% [markdown]
  406. # ## 4. Harmonization and QC relationship (Supplementary Section S3)
  407. # %%
  408. # --- Multi-panel visualization for cgrey matter + site effects ---
  409. # Produces TWO figure blocks:
  410. # A1) QC rows × site columns for cgrey curves (site-adjusted), baseline dashed.
  411. # A2) 2 panels: shift=(g* s_k / d*) and scale=(1/d*) across QC levels (one line per site).
  412. # B1) 4 site panels overlay QC curves.
  413. # B2) 2 panels: shift/scale across sites (one line per QC level).
  414. from matplotlib.ticker import AutoMinorLocator
  415. fontsize = 12
  416. # ---------- config ----------
  417. feature_idx = 1
  418. ga_grid = np.linspace(22, 38, 100)
  419. n_qc = len(models)
  420. qc_labels = categories if "categories" in globals() else [f"QC_{i}" for i in range(n_qc)]
  421. n_sites = int(models[0]["gamma_star"].shape[0])
  422. site_indices = list(range(n_sites))
  423. site_labels = (
  424. [cohort_match.get(str(i), str(i)) for i in site_indices]
  425. if "cohort_match" in globals()
  426. else [str(i) for i in site_indices]
  427. )
  428. # consistent, "R-y" palettes
  429. qc_colors = mpl.colormaps["gray"](np.linspace(0.0, 0.6, n_qc))
  430. site_colors = mpl.colormaps["tab10"](np.arange(n_sites) % 10)
  431. def _calc_site_terms(model, site_idx, feature_idx, ga_grid):
  432. m_at_ga = predict_gam_on_grid(model, ga_grid)
  433. f_k = m_at_ga[:, feature_idx]
  434. s_k = float(np.sqrt(model["var_pooled"][feature_idx, 0]))
  435. g_star = float(model["gamma_star"][site_idx, feature_idx])
  436. d_star = float(np.sqrt(model["delta_star"][site_idx, feature_idx]))
  437. shift = g_star * s_k / d_star
  438. scale = 1.0 / d_star
  439. return f_k, shift, scale
  440. def _site_adjusted_curve(model, site_idx, feature_idx, ga_grid):
  441. f_k, shift, scale = _calc_site_terms(model, site_idx, feature_idx, ga_grid)
  442. return f_k - shift, f_k, shift, scale
  443. # %%
  444. plot_dict = {
  445. "eCSF_Ventricles_total" : "eCSF Volume",
  446. "cGrey.Matter.Volume" : "Cortical Gray Matter Volume",
  447. "White.Matter.Volume" : "White Matter Volume",
  448. "Lateral.Ventricles.Volume" : "Lateral Ventricles Volume",
  449. "Cerebellum.Volume" : "Cerebellum Volume",
  450. "Basal.Ganglia.Volume" : "Basal Ganglia Volume",
  451. "Brainstem.Volume" : "Brainstem Volume",
  452. "Thalamus.Volume" : "Thalamus Volume"
  453. }
  454. plot_values = list(plot_dict.values())
  455. plot_keys = list(plot_dict.keys())
  456. fig, ax = plt.subplots(len(plot_values), 3, figsize=(12, 20), sharex=False)
  457. qc_x = np.arange(n_qc)
  458. qc_names = ["GroundTruth", "AcceptPlus", "PoorPlus", "All"]
  459. qc_linestyle = ["-", "--", "-.", ":"]
  460. # ---- Column 0: average trend ----
  461. for feature_idx in range(len(plot_dict)):
  462. brain_cols_idx = brain_cols.index(plot_keys[feature_idx])
  463. for qi, model in enumerate(models):
  464. y_adj0, y_base0, *_ = _site_adjusted_curve(model, site_indices[0], brain_cols_idx, ga_grid)
  465. ax[feature_idx, 0].plot(
  466. ga_grid,
  467. y_base0,
  468. linestyle=qc_linestyle[qi],
  469. lw=1.4,
  470. color=qc_colors[qi],
  471. label=qc_names[qi],
  472. alpha=0.9,
  473. )
  474. ax[feature_idx, 0].set_ylabel(plot_values[feature_idx])
  475. ax[feature_idx, 0].grid(True)
  476. ax[feature_idx, 0].set_xticks([25, 30, 35])
  477. if feature_idx == len(plot_values) - 1:
  478. ax[feature_idx, 0].set_xlabel("GA")
  479. ax[feature_idx, 0].set_xticklabels([25, 30, 35], rotation=30, ha="right")
  480. else:
  481. ax[feature_idx, 0].set_xticklabels([])
  482. ax[0, 0].set_title("Average trend")
  483. # ---- Columns 1-2: shift / scale terms ----
  484. for feature_idx in range(len(plot_dict)):
  485. shift_mat = np.zeros((n_qc, n_sites), dtype=float)
  486. scale_mat = np.zeros((n_qc, n_sites), dtype=float)
  487. brain_cols_idx = brain_cols.index(plot_keys[feature_idx])
  488. for qi, model in enumerate(models):
  489. for sj, site_idx in enumerate(site_indices):
  490. _f, sh, sc = _calc_site_terms(model, site_idx, brain_cols_idx, ga_grid)
  491. shift_mat[qi, sj] = sh
  492. scale_mat[qi, sj] = sc
  493. for sj, site_lab in enumerate(site_labels):
  494. ax[feature_idx, 1].plot(
  495. qc_x[::-1],
  496. shift_mat[:, sj],
  497. marker="o",
  498. lw=2,
  499. color=site_colors[sj],
  500. label=site_lab,
  501. )
  502. ax[feature_idx, 1].axhline(0, color="black", lw=0.5, linestyle="--")
  503. ax[feature_idx, 2].plot(
  504. qc_x[::-1],
  505. scale_mat[:, sj],
  506. marker="o",
  507. lw=2,
  508. color=site_colors[sj],
  509. label=site_lab,
  510. )
  511. ax[feature_idx, 2].axhline(1, color="black", lw=0.5, linestyle="--")
  512. if feature_idx == 0:
  513. ax[feature_idx, 1].set_title("Shift")
  514. ax[feature_idx, 2].set_title("Scale")
  515. ax[feature_idx, 1].set_xticks(qc_x[::-1])
  516. ax[feature_idx, 2].set_xticks(qc_x[::-1])
  517. if feature_idx == len(plot_values) - 1:
  518. ax[feature_idx, 1].set_xticklabels([str(x) for x in qc_names], rotation=30, ha="right")
  519. ax[feature_idx, 2].set_xticklabels([str(x) for x in qc_names], rotation=30, ha="right")
  520. else:
  521. ax[feature_idx, 1].set_xticklabels([])
  522. ax[feature_idx, 2].set_xticklabels([])
  523. # ---- Formatting: consistent ylabel x-position ----
  524. # Lock all left-column ylabels at the same x coordinate so tick label width changes don't shift them.
  525. for r in range(len(plot_values)):
  526. ax[r, 0].yaxis.set_label_coords(-0.22, 0.5)
  527. # ---- Remove per-axis legends and add one global legend below ----
  528. for r in range(len(plot_values)):
  529. for c in range(3):
  530. leg = ax[r, c].get_legend()
  531. if leg is not None:
  532. leg.remove()
  533. # Collect handles/labels (QC lines from [0,0], site lines from [0,1])
  534. qc_handles, qc_labels = ax[0, 0].get_legend_handles_labels()
  535. site_handles, site_labels2 = ax[0, 1].get_legend_handles_labels()
  536. # Deduplicate (separately) while preserving order
  537. seen_qc = set()
  538. qc_h = []
  539. qc_l = []
  540. for h, lab in zip(qc_handles, qc_labels):
  541. if lab not in seen_qc:
  542. seen_qc.add(lab)
  543. qc_h.append(h)
  544. qc_l.append(lab)
  545. seen_site = set()
  546. site_h = []
  547. site_l = []
  548. for h, lab in zip(site_handles, site_labels2):
  549. if lab not in seen_site:
  550. seen_site.add(lab)
  551. site_h.append(h)
  552. site_l.append(lab)
  553. # Make space for legends at bottom
  554. fig.tight_layout(rect=[0, 0.10, 1, 1])
  555. # Two grouped legends (QC level on left, Site on right)
  556. leg_qc = fig.legend(
  557. qc_h,
  558. qc_l,
  559. title="QC level",
  560. loc="lower left",
  561. bbox_to_anchor=(0.02, 0.08),
  562. ncol=min(len(qc_l), 4),
  563. frameon=False,
  564. handlelength=3,
  565. columnspacing=1.2,
  566. )
  567. leg_site = fig.legend(
  568. site_h,
  569. site_l,
  570. title="Site",
  571. loc="lower right",
  572. bbox_to_anchor=(0.98, 0.08),
  573. ncol=min(len(site_l), 4),
  574. frameon=False,
  575. handlelength=2.5,
  576. columnspacing=1.2,
  577. )
  578. # Ensure both legends are kept
  579. fig.add_artist(leg_qc)
  580. fig.add_artist(leg_site)
  581. fig.show()
  582. # %%
  583. idx = 5
  584. struc = brain_cols[idx]
  585. harm_col = f"{struc}_harm"
  586. # ----- Plot 1: For a given site, plot category-wise site_trajectory -----
  587. site_names = list(site_trajectories[0].keys()) # extracted from first entry
  588. selected_site = site_names[1] # CHANGE to your site of interest
  589. ga_grid = np.linspace(22, 38, 100)
  590. model = models[-1]
  591. model2 = models[0]
  592. m_at_ga = predict_gam_on_grid(model, ga_grid) # (n_grid, n_features)
  593. m_at_ga2 = predict_gam_on_grid(model2, ga_grid) # (n_grid, n_features)
  594. plt.figure(figsize=(8, 5))
  595. for y_pred, site_dict, d in zip(y_preds, site_trajectories, categories):
  596. if selected_site in site_dict:
  597. traj = site_dict[selected_site][:, idx]
  598. plt.plot(ga_grid, traj, label=f'{d}')
  599. # Add observed data for the given site
  600. plt.scatter(
  601. data_df2["age_at_scan"],
  602. data_df2[struc],
  603. c=data_df2["qcglobal"],
  604. #marker=data_df2["batch"].astype(int),
  605. s=8, alpha=0.5, label="Observed"
  606. )
  607. plt.plot(ga_grid, m_at_ga[:, idx], "k--")
  608. plt.plot(ga_grid, m_at_ga2[:, idx], "k:")
  609. plt.title(f"Site Trajectory for {struc} at site {selected_site} (category-wise)")
  610. plt.xlabel("age_at_scan")
  611. plt.ylabel(f"{struc} (harmonized units)")
  612. plt.legend(title="Quality Category")
  613. plt.grid(True)
  614. plt.show()
  615. # %%
  616. def compute_E1_errors(y_preds):
  617. E1_errors = []
  618. E1_error_norm = []
  619. for y_pred, d in zip(y_preds[1:], categories[1:]):
  620. e1 = np.mean(y_pred - y_preds[0],axis=0)
  621. E1_errors.append(e1)
  622. norm = np.mean(y_preds[0],axis=0)
  623. E1_error_norm.append(e1/norm)
  624. return E1_errors, E1_error_norm
  625. errors, errors_norm = compute_E1_errors(y_preds)
  626. category_labels = categories[1:]
  627. n_categories = len(category_labels)
  628. n_structures = len(plot_dict)
  629. error_array = np.vstack(errors_norm)
  630. error_array = np.vstack([error_array[:, brain_cols.index(key)] for key in plot_dict.keys()]).T
  631. x = np.arange(n_structures)
  632. width = 0.25
  633. color_map = {
  634. "accept_plus": "#FF7256", # coral1
  635. "poor_plus": "#458B00", # chartreuse4
  636. "all": "#6495ED", # cornflowerblue
  637. }
  638. labels_str = {
  639. "accept_plus": "AcceptPlus",
  640. "poor_plus": "PoorPlus",
  641. "all": "All",
  642. }
  643. fig, ax = plt.subplots(figsize=(7, 5))
  644. for i, cat in enumerate(category_labels):
  645. ax.bar(x + (i - (n_categories-1)/2)*width, error_array[i], width,
  646. label=labels_str[cat], color=color_map[cat], edgecolor='none')
  647. ax.axhline(0, color='black', lw=1.2, ls='--')
  648. ax.set_xticks(x)
  649. ax.set_xticklabels(plot_dict.values(), rotation=60, ha="right")
  650. ax.set_ylabel('Mean E1 Error')
  651. ax.set_title('Mean E1 Error over average trend\n across harmonization categories')
  652. ax.legend(title='Quality Group', bbox_to_anchor=(1.00, 1.04), loc='upper left')
  653. plt.tight_layout()
  654. plt.show()
  655. # %% [markdown]
  656. # ## Sort harmonize
  657. # For the experiment with sorting and removing iteratively data.
  658. # %%
  659. # Sort data_df by qcglobal ascending
  660. data_df = data_df.sort_values(by="qcglobal", ascending=True).reset_index(drop=True)
  661. # %%
  662. dfs = []
  663. os.makedirs(os.path.join(output_dir,"iterative"), exist_ok=True)
  664. for idx in range(0, 605, 5):
  665. print(f"Harmonizing for indices {idx} onwards")
  666. df2 = data_df[idx:].copy()
  667. harmonized_df = df2.copy()
  668. data_df2 = df2.copy()
  669. covars = get_covars(data_df2)
  670. scaler = StandardScaler()
  671. data_train_scaled = scaler.fit_transform(data_df2[brain_cols])
  672. model, data_train_harmo = harmonizationLearn(data_train_scaled, covars, smooth_terms=["age"])
  673. data_train_harmo = scaler.inverse_transform(data_train_harmo)
  674. harmonized_df[[c+"_harm" for c in brain_cols]] = data_train_harmo
  675. harmonized_df.to_csv(os.path.join(output_dir,"iterative", f"harmonized_data_{idx}.csv"), index=False)
  676. dfs.append(harmonized_df)
  677. # %%

1_harmonize_data.ipynb at commit 24d93cb, no license · at the source

Overview

13 affiliations
  1. CIBM Center for Biomedical Imaging, Lausanne, Switzerland
  2. Department of Medical Radiology, Lausanne University Hospital and University of Lausanne, Lausanne, Switzerland
  3. Institut de Neurosciences de la Timone, CNRS, Aix-Marseille Université, Marseille, France
  4. BCN MedTech, Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain
  5. APHM, Service de Neuroradiologie Diagnostique et Interventionnelle, Hôpital de la Timone, Aix-Marseille University, Marseille, France
  6. APHM, Service de Neurologie Pédiatrique, Hôpital de la Timone, Aix-Marseille University, Marseille, France
  7. Aix Marseille University, Inserm, Marseille, France
  8. BCNatal | Fetal Medicine Research Center (Hospital Clínic and Hospital Sant Joan de Déu, Universitat de Barcelona), Barcelona, Spain
  9. Institut d’Investigacions Biomèdiques August Pi i Sunyer (IDIBAPS), Barcelona, Spain
  10. Centre for Biomedical Research on Rare Diseases (CIBERER), Barcelona, Spain
  11. Department Woman-Mother-Child, Lausanne University Hospital and University of Lausanne, Lausanne, Switzerland
  12. School of Health Sciences (HESAV), HES-SO, University of Applied Sciences and Arts Western Switzerland, Lausanne, Switzerland
  13. ICREA, Barcelona, Spain
Journal: Imaging neuroscience (Cambridge, Mass.), volume 4, article IMAG.a.1337
Dates: received 22 January 2026; accepted 17 July 2026; published online 14 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1162/imag.a.1337 · PMID 42609916 · PMCID PMC13479347 · OpenAlex W7171370507
Open access: diamond, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality)
Methods: Connectivity, Statistics, Physiology & signal measures
Keywords: fetal, brain MRI, normative modeling, image artifacts, quality control
Topic: Fetal and Pediatric Neurological Disorders (Pediatrics, Perinatology and Child Health, Medicine), according to OpenAlex
Citations: not cited yet (Europe PMC); 42 references in the paper

Abstract

Normative modeling is increasingly used to characterize typical growth trajectories and identify atypical neurodevelopment, including early brain development using magnetic resonance imaging (MRI) acquired before birth. Recent work has emphasized the importance of large sample sizes for accurate and robust centile estimation. In this study, we investigate how image quality influences fetal brain normative models, a critical factor in this context where MRI is acquired on a moving fetus in utero. Using a multi-centric cohort of 635 fetal MRI scans, we applied a standardized visual quality control (QC) protocol with continuous quality ratings. We fit normative models for multiple brain structures under progressively relaxed QC stringency, and quantified the deviations in centile estimates relative to a high-quality reference subgroup. Our results showed that including lower-quality data systematically biased normative centiles, with the strongest effects observed in the outer centiles, particularly the lower tail (1st–10th). Bias increased progressively as QC stringency was relaxed and could not be attributed solely to the number of scans used to fit the models. Quality-induced bias was structure dependent, and often not visually apparent at the segmentation level. These findings highlight that image quality is an important source of bias in normative fetal brain modeling, and that increasing sample size at the expense of quality may systematically affect centile estimates, potentially jeopardizing the utility of the model.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.

fetpype/fetpype

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 5c99de9a3b35734b2fa71b5fcde2920cb9b1d93b, 3 September 2026
Languages: Python (27)
Size: 99 files, 27 scripts
Software Heritage: archived
Found in: the end of the paper
Holds: README, license file, CITATION.cff, environment (pyproject.toml), tests, continuous integration, documentation
Tools: Nipype (10 files), NiBabel (2 files), NumPy (2 files), FSL (1 file), pandas (1 file), PyBIDS (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
29 files

fetpype/normative_modeling_data_quality

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 24d93cbe628568ac5ff3df680bd6046fdd8e393e, 22 July 2026
Languages: R (10), Jupyter (1), Python (1)
Size: 13 files, 12 scripts
Software Heritage: not archived
Found in: “Data and Code Availability”
Holds: README, 1 notebook
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (10 files), ggplot2 (9 files), patchwork (6 files), cowplot (4 files), mgcv (4 files), neuroHarmonize (2 files), NumPy (2 files), pandas (2 files), scikit-learn (2 files), Matplotlib (1 file), nlme (1 file), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
13 files

Zenodo 15696638

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file, 0 scripts
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 39 scripts, each with its path and the digest of its content;
  • 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data and Code Availability

Tabular data (estimated volumes data and corresponding data quality for each scan) are available on a Zenodo repository at https://doi.org/10.5281/zenodo.21485414. The code is publicly available on GitHub at https://github.com/fetpype/normative_modeling_data_quality. Sharing of the reconstructed images and segmentations requires data transfer agreements with each of the centers.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 16 authors, 5 keywords, 4 funders, 42 references.

Cite

This paper

Sanchez, T., Mihailov, A., Martí-Juan, G., Girard, N., Manchon, A., Milh, M., Eixarch, E., Dunet, V., Koob, M., Pomar, L., Sichitiu, J., González Ballester, M. A., Camara, O., Piella, G., Bach Cuadra, M., & Auzias, G. (2026). Data quality biases normative models derived from fetal brain MRI. Imaging neuroscience (Cambridge, Mass.), 4, IMAG.a.1337. https://doi.org/10.1162/imag.a.1337

BibTeX

@article{sanchez2026data,
author = {Sanchez, Thomas and Mihailov, Angeline and Martí-Juan, Gerard and Girard, Nadine and Manchon, Aurélie and Milh, Mathieu and Eixarch, Elisenda and Dunet, Vincent and Koob, Mériam and Pomar, Léo and Sichitiu, Joanna and González Ballester, Miguel A. and Camara, Oscar and Piella, Gemma and Bach Cuadra, Meritxell and Auzias, Guillaume},
title = {{Data quality biases normative models derived from fetal brain MRI}},
journal = {Imaging neuroscience (Cambridge, Mass.)},
year = {2026},
month = aug,
volume = {4},
pages = {IMAG.a.1337},
publisher = {MIT Press},
issn = {2837-6056},
doi = {10.1162/imag.a.1337},
url = {https://doi.org/10.1162/imag.a.1337},
pmid = {42609916},
pmcid = {PMC13479347}
}

RIS

TY - JOUR
AU - Sanchez, Thomas
AU - Mihailov, Angeline
AU - Martí-Juan, Gerard
AU - Girard, Nadine
AU - Manchon, Aurélie
AU - Milh, Mathieu
AU - Eixarch, Elisenda
AU - Dunet, Vincent
AU - Koob, Mériam
AU - Pomar, Léo
AU - Sichitiu, Joanna
AU - González Ballester, Miguel A.
AU - Camara, Oscar
AU - Piella, Gemma
AU - Bach Cuadra, Meritxell
AU - Auzias, Guillaume
TI - Data quality biases normative models derived from fetal brain MRI
T2 - Imaging neuroscience (Cambridge, Mass.)
J2 - Imaging Neurosci (Camb)
PY - 2026
DA - 2026/08/14
VL - 4
SP - IMAG.a.1337
SN - 2837-6056
PB - MIT Press
DO - 10.1162/imag.a.1337
UR - https://doi.org/10.1162/imag.a.1337
LA - en
ER -

CSL-JSON

{
"id": "10.1162/imag.a.1337",
"type": "article-journal",
"title": "Data quality biases normative models derived from fetal brain MRI",
"container-title": "Imaging neuroscience (Cambridge, Mass.)",
"author": [
{
"family": "Sanchez",
"given": "Thomas"
},
{
"family": "Mihailov",
"given": "Angeline"
},
{
"family": "Martí-Juan",
"given": "Gerard"
},
{
"family": "Girard",
"given": "Nadine"
},
{
"family": "Manchon",
"given": "Aurélie"
},
{
"family": "Milh",
"given": "Mathieu"
},
{
"family": "Eixarch",
"given": "Elisenda"
},
{
"family": "Dunet",
"given": "Vincent"
},
{
"family": "Koob",
"given": "Mériam"
},
{
"family": "Pomar",
"given": "Léo"
},
{
"family": "Sichitiu",
"given": "Joanna"
},
{
"family": "González Ballester",
"given": "Miguel A."
},
{
"family": "Camara",
"given": "Oscar"
},
{
"family": "Piella",
"given": "Gemma"
},
{
"family": "Bach Cuadra",
"given": "Meritxell"
},
{
"family": "Auzias",
"given": "Guillaume"
}
],
"container-title-short": "Imaging Neurosci (Camb)",
"volume": "4",
"page": "IMAG.a.1337",
"DOI": "10.1162/imag.a.1337",
"PMID": "42609916",
"PMCID": "PMC13479347",
"ISSN": "2837-6056",
"publisher": "MIT Press",
"URL": "https://doi.org/10.1162/imag.a.1337",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-73072-6 [code]
Mapping the spatiotemporal continuum of structural connectivity development across the human connectome in youth.
Journal: Nature communications
In common: neuroHarmonize, mgcv, patchwork, 5 other tools, structural MRI / diffusion, 4 references
[2] doi:10.1162/imag.a.1269 [code]
From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: NiBabel, seaborn, scikit-learn, 4 other tools, 7 references
[3] doi:10.1162/imag.a.1347 [code]
Neural and behavioural correlates of theory of mind reasoning in five-year-old children born preterm.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: PyBIDS, Nipype, FSL, 9 other tools
[4] doi:10.1162/imag.a.1245 [code]
Towards precision EEG connectomics: Evaluating the benefits of dense sampling.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Nipype, FSL, cowplot, 9 other tools
[5] doi:10.1002/hbm.70605 [code]
BrainEnrich: Revealing Biological Insights for Imaging-Derived Phenotypes Through Transcriptomic Enrichment.
Journal: Human brain mapping
In common: nlme, cowplot, patchwork, 9 other tools, 1 reference
[6] doi:10.1016/j.dcn.2026.101765 [code]
Fusiform face area development correlates with development in higher-order social brain regions.
Journal: Developmental cognitive neuroscience
In common: PyBIDS, Nipype, FSL, 8 other tools
[7] doi:10.1038/s42003-026-10276-y [code]
The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.
Journal: Communications biology
In common: Nipype, mgcv, patchwork, 8 other tools, structural MRI / diffusion
[8] doi:10.1162/imag.a.1276 [code]
High-resolution whole-brain magnetic resonance spectroscopic imaging in youth at risk for psychosis.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: FSL, NiBabel, seaborn, 5 other tools, author Meritxell Bach Cuadra
[9] doi:10.1371/journal.pone.0343722 [code]
Comprehensive methodology for sample enrichment in EEG biomarker studies for Alzheimer's risk classification.
Journal: PloS one
In common: neuroHarmonize, PyBIDS, ggplot2, 6 other tools, 1 reference
[10] doi:10.3389/fnins.2026.1873417 [code]
Functional brain network alterations in a rat model of dental malocclusion.
Journal: Frontiers in neuroscience
In common: PyBIDS, Nipype, FSL, 7 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.