Data quality biases normative models derived from fetal brain MRI.
The 8 matches
- [1] § Methods › Data processing › Segmentation ↔ scripts/plots_and_analyses/supp_plot_iterative_removal.r, lines 1–80 · score 0.76 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, brainstem, thalamus
- [2] § Methods › Data processing › Segmentation ↔ scripts/plots_and_analyses/1_plot_main.r, lines 1–39 · score 0.75 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, brainstem, thalamus
- [3] § Results › Poor-quality images can bias normative trajectories ↔ notebooks/1_harmonize_data.ipynb, lines 539–558 · score 0.74 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, cerebellar volume, brainstem
- [4] § Results › Poor-quality images can bias normative trajectories ↔ scripts/plots_and_analyses/1_plot_main.r, lines 1–39 · score 0.74 · cortical gray matter, eCSF, basal ganglia, lateral ventricles, cerebellar volume, brainstem
- [5] § Methods › Data processing › Super-resolution reconstruction ↔ fetpype/pipelines/full_pipeline.py, lines 26–104 · score 0.65 · bias field correction, brain extraction, denoising, resolution, preprocessing, stacks
- [6] § Results › Poor-quality images can bias normative trajectories ↔ scripts/plots_and_analyses/supp_plot_iterative_removal.r, lines 1–80 · score 0.60 · lateral ventricles volume, white matter volume, cerebellar volume, confidence, bootstrap, centile
- [7] § Methods › Analysis of the impact of image quality on normative models › Harmonization on volumetric measurements ↔ notebooks/1_harmonize_data.ipynb, lines 197–283 · score 0.58 · ComBat, GroundTruth, GAM, smoothing, spline, harmonized
- [8] § Methods › Analysis of the impact of image quality on normative models › QC subgroups definition ↔ notebooks/1_harmonize_data.ipynb, lines 36–86 · score 0.50 · global quality score, excellent quality, cohorts, scans
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 814 lines · 28 KB · no license · 3 matches
- # %% [markdown]
- # # Harmonization and data description
- # This python notebook loads the raw data, plots some demographics, runs the harmonization and saves the harmonized CSV files to be used in the rest of the analyses.
- #
- # Here is how the notebook is organized.
- #
- # 1. **Demographics plotting:** Producing basic descriptive figures of the data (Figure 3)
- # 2. **Harmonization:** Run and save the harmonization
- # 3. **Detrending analysis:** Plot and analyze the influence of data quality on the raw trajectories (Figure 4)
- # %%
- import pandas as pd
- import numpy as np
- import warnings
- import matplotlib.pyplot as plt
- import seaborn as sns
- from sklearn.preprocessing import StandardScaler
- import os
- import scipy
- from neuroHarmonize import harmonizationLearn, harmonizationApply
- from sklearn.metrics import r2_score, mean_squared_error
- warnings.filterwarnings("ignore")
- %load_ext autoreload
- %autoreload 2
- # %%
- input_file = "../data/raw_data/finalized_df_qcglobal_and_z.csv"
- output_dir = "../data/harmonization"
- plot_dir = output_dir
- os.makedirs(output_dir, exist_ok=True)
- data_df = pd.read_csv(input_file)
- # %% [markdown]
- # ## 1. Demographics plotting
- # %%
- f, ax = plt.subplots(2, figsize=(8, 5), gridspec_kw={"height_ratios": [1, 3]}, sharex=True)
- plt.subplots_adjust(hspace=0.01)
- n_excl = len(data_df[data_df["qcglobal"] < 0.95])
- n_poor = len(data_df[(data_df["qcglobal"] >= 0.95) & (data_df["qcglobal"] < 1.95)])
- n_acc = len(data_df[(data_df["qcglobal"] >= 1.95) & (data_df["qcglobal"] < 2.7)])
- n_excellent = len(data_df[data_df["qcglobal"] >= 2.7])
- print(f"Number of excluded scans (qcglobal < 0.95): {n_excl}")
- print(f"Number of poor quality scans (0.95 <= qcglobal < 1.95): {n_poor}")
- print(f"Number of acceptable quality scans (1.95 <= qcglobal < 2.7): {n_acc}")
- print(f"Number of excellent quality scans (qcglobal >= 2.7): {n_excellent}")
- for i, a in enumerate(ax):
- if i==0:
- max_ = len(data_df["cohort"].unique())
- else:
- max_ = 65
- a.axvline(x=0.95, color="black", linestyle="--")
- a.axvline(x=1.95, color="black", linestyle="--")
- a.axvline(x=2.7, color="black", linestyle="--")
- a.fill_betweenx(y=[-1, max_], x1=0, x2=0.95, color="lightgray", alpha=0.5)
- a.fill_betweenx(y=[-1, max_], x1=0.95, x2=1.95, color="gray", alpha=0.5)
- a.fill_betweenx(y=[-1, max_], x1=1.95, x2=2.7, color="lightgray", alpha=0.5)
- a.fill_betweenx(y=[-1, max_], x1=2.7, x2=4, color="gray", alpha=0.5)
- # Select bins to finish at 1, 2, 2.7
- bins = np.arange(0, data_df["qcglobal"].max()+0.1, 0.1)
- sns.histplot(data=data_df, ax=ax[1], x="qcglobal", hue="cohort", multiple="stack", bins=bins)
- # Add ontop a kind of boxplot to show quartiles for each cohort, it's a second plot with same x axis
- sns.boxplot(data=data_df, ax=ax[0], x="qcglobal", y="cohort", hue="cohort")
- # For ax[0], remove xlabel and ylabel and xticks
- ax[0].set_xlabel("")
- ax[0].set_ylabel("")
- ax[1].set_xlabel("Global quality score")
- ax[0].set_xticks([])
- ax[0].set_yticks([])
- ax[1].set_ylim(0, 65)
- ax[1].set_xlim(0,4)
- ax[0].set_xlim(0,4)
- ax[1].annotate("Exclude Quality", xy=(0.475, max_), xytext=(0.475, max_-5), fontsize=12, color="black", ha="center")
- ax[1].annotate(f"(n={n_excl})", xy=(0.475, max_-10), xytext=(0.475, max_-8), fontsize=10, color="black", ha="center")
- ax[1].annotate("Poor Quality", xy=(1.45, max_), xytext=(1.45, max_-5), fontsize=12, color="black", ha="center")
- ax[1].annotate(f"(n={n_poor})", xy=(1.45, max_-10), xytext=(1.45, max_-9), fontsize=10, color="black", ha="center")
- ax[1].annotate("Acceptable\nQuality", xy=(2.325, max_), xytext=(2.325, max_-9), fontsize=12, color="black", ha="center")
- ax[1].annotate(f"(n={n_acc})", xy=(2.325, max_-10), xytext=(2.325, max_-12.5), fontsize=10, color="black", ha="center")
- ax[1].annotate("Excellent Quality", xy=(4, max_), xytext=(3.3, max_-5), fontsize=12, color="black", ha="center")
- ax[1].annotate(f"(n={n_excellent})", xy=(4, max_-10), xytext=(3.3, max_-8), fontsize=10, color="black", ha="center")
- # Across the two plots, add a vertical line at qcglobal = 0.95, one at 1.95 and one at 2.7 and fill the background between 0 and 0.95 with lightgray, between 0.95 and 1.95 with gray, between 1.95 and 2.7 with darkgray, and above 2.7 with black
- # %%
- f, ax = plt.subplots(2, figsize=(5, 6), gridspec_kw={"height_ratios": [1, 3]}, sharex=True)
- plt.subplots_adjust(hspace=0.01)
- n_excl = len(data_df[data_df["qcglobal"] < 0.95])
- n_poor = len(data_df[(data_df["qcglobal"] >= 0.95) & (data_df["qcglobal"] < 1.95)])
- n_acc = len(data_df[(data_df["qcglobal"] >= 1.95) & (data_df["qcglobal"] < 2.7)])
- n_excellent = len(data_df[data_df["qcglobal"] >= 2.7])
- print(f"Number of excluded scans (qcglobal < 0.95): {n_excl}")
- print(f"Number of poor quality scans (0.95 <= qcglobal < 1.95): {n_poor}")
- print(f"Number of acceptable quality scans (1.95 <= qcglobal < 2.7): {n_acc}")
- print(f"Number of excellent quality scans (qcglobal >= 2.7): {n_excellent}")
- # Select bins to finish at 1, 2, 2.7
- bins = np.arange(20,39, 1)
- sns.histplot(data=data_df, ax=ax[1], x="age_at_scan", hue="cohort", multiple="stack", bins=bins)
- # Add ontop a kind of boxplot to show quartiles for each cohort, it's a second plot with same x axis
- sns.boxplot(data=data_df, ax=ax[0], x="age_at_scan", y="cohort", hue="cohort")
- # For ax[0], remove xlabel and ylabel and xticks
- ax[0].set_xlabel("")
- ax[0].set_ylabel("")
- ax[1].set_xlabel("Gestational age (weeks)")
- # Remove the lengends from ax[0]
- ax[1].get_legend().remove()
- # %%
- data_df['age_bin'] = pd.cut(data_df['age_at_scan'], bins=np.arange(20,40,2))
- plt.figure(figsize=(6,6))
- # Global boxplot (light gray behind)
- sns.boxplot(
- data=data_df,
- x='age_bin',
- y='qcglobal',
- color='white',
- showfliers=False,
- boxprops=dict(facecolor='lightgray', edgecolor='black', linewidth=1.5,alpha=0.7),
- whiskerprops=dict(color='black', linewidth=1.5),
- capprops=dict(color='black', linewidth=0.),
- medianprops=dict(color='black', linewidth=1.5)
- )
- # Overlay cohort-specific boxplots with transparency
- sns.boxplot(
- data=data_df,
- x='age_bin',
- y='qcglobal',
- hue='cohort',
- showfliers=False,
- dodge=True,
- palette='tab10',
- linewidth=0.5,
- boxprops=dict(alpha=1.)
- )
- plt.xticks(rotation=45)
- plt.xlabel("Age at scan (binned)")
- plt.ylabel("Global quality score")
- plt.xticks(fontsize=14)
- plt.yticks(fontsize=14)
- plt.ylabel("Global quality score", fontsize=14)
- plt.xlabel("Gestational age (weeks)", fontsize=14)
- sns.set_style("ticks") # adds small ticks
- sns.set_context("talk", font_scale=1.1)
- plt.legend(title="Cohort", bbox_to_anchor=(1.05, 1), loc="upper left", fontsize=10)
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # ## 2. Harmonization
- # %%
- ## Set up thresholds for different quality groups
- data_df["all"] = data_df["qcglobal"] >= 0
- data_df["poor_plus"] = data_df["qcglobal"] >= 0.95
- data_df["accept_plus"] = data_df["qcglobal"] >= 1.95
- data_df["ground_truth"] = data_df["qcglobal"] >= 2.7
- data_df["age_rounded"] = round(data_df["age_at_scan"], 2)
- data_df["qc_group"] = pd.Categorical(
- np.where(data_df["qcglobal"] < 0.95, "excluded",
- np.where(data_df["qcglobal"] < 1.95, "poor",
- np.where(data_df["qcglobal"] < 2.7, "acceptable", "ground truth")))
- )
- data_df['batch'] = pd.Categorical(data_df["cohort"]).codes.astype(str) #convert cohort variable to BATCH variable for analysis. Can change to "scanner"
- cohort_match = dict(data_df[["batch", "cohort"]].drop_duplicates().values)
- # %%
- ## Set up the renaming of the different variables.
- brain_cols = [
- 'CSF.Volume', 'cGrey.Matter.Volume', 'Lateral.Ventricles.Volume', "Brainstem.Volume", 'Cerebellum.Volume', 'Basal.Ganglia.Volume', 'Thalamus.Volume', 'White.Matter.Volume', 'CSF_total', 'eCSF_Ventricles_total'
- ]
- cols_name = [
- "eCSF (no ventricles) Volume",
- "Cortical Gray Matter Volume",
- "Lateral Ventricles Volume",
- "Brainstem Volume",
- "Cerebellum Volume",
- "Basal Ganglia Volume",
- "Thalamus Volume",
- "White Matter Volume",
- "CSF total Volume",
- "eCSF Volume"
- ]
- # %%
- dfs = []
- # QC category order + display labels (match R: qc_order <- c("GroundTruth", "AcceptPlus", "PoorPlus", "All"))
- qc_order = ["GroundTruth", "AcceptPlus", "PoorPlus", "All"]
- categories = ["ground_truth", "accept_plus", "poor_plus", "all"]
- category_display = dict(zip(categories, qc_order))
- trajs = []
- ga_by_category = {}
- def get_covars(df):
- covars = pd.DataFrame({
- #"ICV": df["ICV"],
- "age": df["age_at_scan"],
- "SITE":df["batch"]
- })
- return covars
- def predict_gam_on_grid(model, ga_grid):
- smooth_model = model['smooth_model']
- bs = smooth_model['bsplines_constructor']
- B_hat = model['B_hat']
- df_gam_train = smooth_model['df_gam']
- n_grid = len(ga_grid)
- site_cols = [c for c in df_gam_train.columns if c.startswith('x')]
- cov_cols = [c for c in df_gam_train.columns if c.startswith('c')]
- param_cols = site_cols + cov_cols
- X_param = np.zeros((n_grid, len(param_cols)))
- for j, col in enumerate(param_cols):
- if col in cov_cols:
- X_param[:, j] = df_gam_train[col].mean()
- else:
- X_param[:, j] = 0
- X_spline = bs.transform(ga_grid.reshape(-1, 1))
- X_new = np.concatenate([X_param, X_spline], axis=1)
- return X_new @ B_hat + model["stand_mean"][:,0] # (n_grid, n_features)
- y_preds = []
- models = []
- site_trajectories = []
- ga_grid = np.linspace(22, 38, 100) # adjust to your GA range
- for d in categories:
- print(f"Harmonizing for {d} subgroup")
- df2 = data_df.copy()
- harmonized_df = df2.copy()
- data_df2 = df2.copy()
- harmo_train = data_df2[data_df2[d] == True] #change here if want to do by scanner
- data_train = data_df2[data_df2[d] == True]
- data_test = data_df2[data_df2[d] == False]
- print(len(data_train), len(data_test))
- covars_train = get_covars(data_train)
- covars_test = get_covars(data_test)
- model, data_train_harmo = harmonizationLearn(np.array(harmo_train[brain_cols]), covars_train, smooth_terms=["age"])
- harmonized_df.loc[harmonized_df[d] == True, [c+"_harm" for c in brain_cols]] = data_train_harmo
- if len(data_test) > 0:
- data_test_harmo = harmonizationApply(np.array(data_test[brain_cols]), covars_test,model)
- harmonized_df.loc[harmonized_df[d] == False, [c+"_harm" for c in brain_cols]] = data_test_harmo
- y_pred = predict_gam_on_grid(model, ga_grid)
- grand_mean = model['stand_mean'][:, 0] # (n_features,)
- # The biological trajectory (same for all sites)
- trajectory = y_pred # (n_grid, n_features)
- # Site-specific shift in natural units
- gamma_natural = model['gamma_star'] * np.sqrt(model['var_pooled']).T # (n_sites, n_features)
- # Full signal per site across GA grid
- site_dict = {}
- for i, site in enumerate(model['SITE_labels']):
- site_trajectory = trajectory - gamma_natural[i, :] # (n_grid, n_features)
- site = cohort_match[str(site)]
- site_dict[site] = site_trajectory
- site_trajectories.append(site_dict)
- y_preds.append(y_pred)
- models.append(model)
- dfs.append(harmonized_df)
- harmonized_df.to_csv(os.path.join(output_dir, f"harmonized_data_{d}.csv"), index=False)
- # %% [markdown]
- # ## 3. Detrending analysis
- # %%
- sns.set_theme(style="whitegrid")
- import matplotlib.colors as mcolors
- import matplotlib as mpl
- import matplotlib.cm as mcm
- from matplotlib.lines import Line2D
- # ---- Options ----
- show_fit = True # set to False to hide the spline fit line
- #brain_cols = ['eCSF_Ventricles_total', 'cGrey.Matter.Volume', 'Basal.Ganglia.Volume', "Brainstem.Volume", 'Thalamus.Volume' ]
- #cols_name = ["eCSF Volume", "Cortical Gray Matter Volume", "Basal Ganglia Volume", "Brainstem Volume", "Thalamus Volume"]
- n_features = len(brain_cols)
- fig, axes = plt.subplots(
- n_features, 3,
- figsize=(13, n_features * 2.8),
- gridspec_kw={"width_ratios": [3, 3, 1]},
- )
- # ---------- Manual axis sharing ----------
- # All GA panels (cols 0 & 1) share the same x-axis
- for i in range(1, n_features):
- axes[i, 0].sharex(axes[0, 0])
- axes[i, 1].sharex(axes[0, 0])
- axes[0, 1].sharex(axes[0, 0])
- # QC marginal panels (col 2) share x among themselves
- for i in range(1, n_features):
- axes[i, 2].sharex(axes[0, 2])
- # col 1 (detrended) and col 2 (QC marginal) are NOT y-linked:
- # each detrended panel keeps its own independent y-axis.
- # Hide x tick labels on all but the bottom row
- for i in range(n_features - 1):
- axes[i, 0].tick_params(labelbottom=False)
- axes[i, 1].tick_params(labelbottom=False)
- axes[i, 2].tick_params(labelbottom=False)
- # ---------- Column titles ----------
- axes[0, 0].set_title("Volume [mm3]", fontsize=11, fontweight="bold")
- axes[0, 1].set_title("Detrended Volume [mm3]", fontsize=11, fontweight="bold")
- # ---------- Color / marker setup ----------
- qc_min = harmonized_df["qcglobal"].min()
- qc_max = harmonized_df["qcglobal"].max()
- norm = mcolors.Normalize(vmin=qc_min, vmax=qc_max)
- cmap = mpl.colormaps["viridis_r"]
- cohorts = sorted(harmonized_df["cohort"].unique())
- _marker_cycle = ["o", "s", "^", "D", "v", "P", "X", "h", "*", "+"]
- cohort_marker = {c: _marker_cycle[i % len(_marker_cycle)] for i, c in enumerate(cohorts)}
- # ---------- Per-feature rows ----------
- hdr = f"{'Feature':<40} {'beta_lin':>10} {'R2_lin':>8} {'p_lin':>12} {'a (quad)':>10} {'b (quad)':>10} {'R2_quad':>8}"
- print(hdr)
- print("-" * len(hdr))
- for idx in range(n_features):
- df2 = harmonized_df.copy()
- col = brain_cols[idx]
- name = cols_name[idx]
- ax_vol = axes[idx, 0]
- ax_det = axes[idx, 1]
- ax_qc = axes[idx, 2]
- # ---- Spline fit on high-QC subset ----
- df_sel = df2[df2["qcglobal"] >= 2.75] #2.75
- x_hq = df_sel["age_at_scan"].values
- y_hq = df_sel[col].values
- order = np.argsort(x_hq)
- x_sorted = x_hq[order]; y_sorted = y_hq[order]
- xl, yl = [], []
- for i in range(len(x_sorted)):
- if i == 0 or xl[-1] != x_sorted[i]:
- xl.append(x_sorted[i]); yl.append([y_sorted[i]])
- else:
- yl[-1].append(y_sorted[i])
- yl = [np.mean(v) for v in yl]
- spline = scipy.interpolate.make_smoothing_spline(xl, yl)
- x_fit_s = np.linspace(np.min(xl), np.max(xl), 400)
- y_fit_s = spline(x_fit_s)
- df2["spl_trend"] = spline(df2["age_at_scan"])
- df2["residual"] = df2[col] - df2["spl_trend"]
- # ---- Scatter: one call per cohort so markers differ ----
- for cohort in cohorts:
- mask = df2["cohort"] == cohort
- subset = df2[mask]
- colors = cmap(norm(subset["qcglobal"].values))
- kw = dict(c=colors, marker=cohort_marker[cohort],
- alpha=0.7, s=20, linewidths=0)
- ax_vol.scatter(subset["age_at_scan"], subset[col], **kw)
- ax_det.scatter(subset["age_at_scan"], subset["residual"], **kw)
- if show_fit:
- ax_vol.plot(x_fit_s, y_fit_s, color="black", lw=1.5)
- ax_vol.set_ylabel(name, fontsize=11)
- ax_det.set_ylabel("")
- vol_lo = df2[col].quantile(0.01)
- vol_hi = df2[col].quantile(0.99)
- ymin = max(0, vol_lo*0.8)
- ax_vol.set_ylim(ymin, vol_hi)
- ax_det.axhline(0, color="black", lw=1, ls="--")
- resid_lo = df2["residual"].quantile(0.01)
- resid_hi = df2["residual"].quantile(0.99)
- ax_det.set_ylim(resid_lo, resid_hi)
- # ---- Marginal QC panel ----
- n_bins = 20
- bins = np.linspace(resid_lo, resid_hi, n_bins + 1)
- centers = 0.5 * (bins[:-1] + bins[1:])
- df2["resid_bin"] = pd.cut(df2["residual"], bins)
- qc_mean = df2.groupby("resid_bin", observed=False)["qcglobal"].mean()
- qc_std = df2.groupby("resid_bin", observed=False)["qcglobal"].std()
- # Linear + quadratic regression (printed only, not drawn)
- reg_df = df2[["residual", "qcglobal"]].dropna()
- reg_mask = np.isfinite(reg_df["residual"]) & np.isfinite(reg_df["qcglobal"])
- reg_df = reg_df.loc[reg_mask]
- x_obs = reg_df["residual"].to_numpy()
- y_obs = reg_df["qcglobal"].to_numpy()
- lr = scipy.stats.linregress(x_obs, y_obs)
- pred = lr.slope * x_obs + lr.intercept
- a2, b1, c0 = np.polyfit(x_obs, y_obs, deg=2)
- y_hat_quad = a2 * x_obs**2 + b1 * x_obs + c0
- r2_quad = r2_score(y_obs, y_hat_quad)
- print(f"{name:<40} {lr.slope:>10.4f} {lr.rvalue**2:>8.4f} {lr.pvalue:>12.4e} "
- f"{a2:>10.4f} {b1:>10.4f} {r2_quad:>8.4f}")
- # Mean \u00b1 SD only (no regression line overlay)
- ax_qc.plot(qc_mean, centers, color="black", lw=2)
- ax_qc.fill_betweenx(centers, qc_mean - qc_std, qc_mean + qc_std,
- color="gray", alpha=0.3)
- ax_qc.axhline(0, color="black", lw=2, ls="--")
- ax_qc.set_xlim(qc_min, qc_max)
- ax_qc.set_ylim(resid_lo, resid_hi)
- ax_qc.set_yticks([])
- ax_qc.tick_params(labelleft=False)
- ax_qc.grid(True, axis="x", alpha=0.2)
- # ---------- x-labels on bottom row ----------
- axes[-1, 0].set_xlabel("GA (weeks)", fontsize=10)
- axes[-1, 1].set_xlabel("GA (weeks)", fontsize=10)
- axes[-1, 2].set_xlabel("QC", fontsize=10)
- # ---------- Reserve space: right margin for colorbar, bottom for legend ----------
- plt.tight_layout(rect=[0, 0.03, 0.8, 1])
- # ---------- Colorbar ----------
- cbar_ax = fig.add_axes([0.87, 0.38, 0.018, 0.25])
- sm = mcm.ScalarMappable(cmap=cmap, norm=norm)
- sm.set_array([])
- cbar = fig.colorbar(sm, cax=cbar_ax)
- cbar.set_label("QC Global", rotation=270, labelpad=14, fontsize=10)
- # ---------- Site / marker legend \u2014 horizontal, below the plot ----------
- legend_handles = [
- Line2D(
- [0], [0],
- marker=cohort_marker[c], color="w",
- markerfacecolor="gray", markeredgecolor="gray",
- markersize=7, linestyle="None",
- label=c,
- )
- for c in cohorts
- ]
- fig.legend(
- handles=legend_handles,
- title="Site",
- loc="lower center",
- bbox_to_anchor=(0.42, 0.0),
- ncol=len(cohorts),
- fontsize=8,
- title_fontsize=9,
- framealpha=0.9,
- )
- # ---------- Save high-resolution ----------
- save_path = os.path.join(plot_dir, "qc_residual_scatter_linear.svg")
- fig.savefig(save_path, dpi=300, bbox_inches="tight")
- print(f"\nFigure saved to: {save_path}")
- plt.show()
- # %% [markdown]
- # ## 4. Harmonization and QC relationship (Supplementary Section S3)
- # %%
- # --- Multi-panel visualization for cgrey matter + site effects ---
- # Produces TWO figure blocks:
- # A1) QC rows × site columns for cgrey curves (site-adjusted), baseline dashed.
- # A2) 2 panels: shift=(g* s_k / d*) and scale=(1/d*) across QC levels (one line per site).
- # B1) 4 site panels overlay QC curves.
- # B2) 2 panels: shift/scale across sites (one line per QC level).
- from matplotlib.ticker import AutoMinorLocator
- fontsize = 12
- # ---------- config ----------
- feature_idx = 1
- ga_grid = np.linspace(22, 38, 100)
- n_qc = len(models)
- qc_labels = categories if "categories" in globals() else [f"QC_{i}" for i in range(n_qc)]
- n_sites = int(models[0]["gamma_star"].shape[0])
- site_indices = list(range(n_sites))
- site_labels = (
- [cohort_match.get(str(i), str(i)) for i in site_indices]
- if "cohort_match" in globals()
- else [str(i) for i in site_indices]
- )
- # consistent, "R-y" palettes
- qc_colors = mpl.colormaps["gray"](np.linspace(0.0, 0.6, n_qc))
- site_colors = mpl.colormaps["tab10"](np.arange(n_sites) % 10)
- def _calc_site_terms(model, site_idx, feature_idx, ga_grid):
- m_at_ga = predict_gam_on_grid(model, ga_grid)
- f_k = m_at_ga[:, feature_idx]
- s_k = float(np.sqrt(model["var_pooled"][feature_idx, 0]))
- g_star = float(model["gamma_star"][site_idx, feature_idx])
- d_star = float(np.sqrt(model["delta_star"][site_idx, feature_idx]))
- shift = g_star * s_k / d_star
- scale = 1.0 / d_star
- return f_k, shift, scale
- def _site_adjusted_curve(model, site_idx, feature_idx, ga_grid):
- f_k, shift, scale = _calc_site_terms(model, site_idx, feature_idx, ga_grid)
- return f_k - shift, f_k, shift, scale
- # %%
- plot_dict = {
- "eCSF_Ventricles_total" : "eCSF Volume",
- "cGrey.Matter.Volume" : "Cortical Gray Matter Volume",
- "White.Matter.Volume" : "White Matter Volume",
- "Lateral.Ventricles.Volume" : "Lateral Ventricles Volume",
- "Cerebellum.Volume" : "Cerebellum Volume",
- "Basal.Ganglia.Volume" : "Basal Ganglia Volume",
- "Brainstem.Volume" : "Brainstem Volume",
- "Thalamus.Volume" : "Thalamus Volume"
- }
- plot_values = list(plot_dict.values())
- plot_keys = list(plot_dict.keys())
- fig, ax = plt.subplots(len(plot_values), 3, figsize=(12, 20), sharex=False)
- qc_x = np.arange(n_qc)
- qc_names = ["GroundTruth", "AcceptPlus", "PoorPlus", "All"]
- qc_linestyle = ["-", "--", "-.", ":"]
- # ---- Column 0: average trend ----
- for feature_idx in range(len(plot_dict)):
- brain_cols_idx = brain_cols.index(plot_keys[feature_idx])
- for qi, model in enumerate(models):
- y_adj0, y_base0, *_ = _site_adjusted_curve(model, site_indices[0], brain_cols_idx, ga_grid)
- ax[feature_idx, 0].plot(
- ga_grid,
- y_base0,
- linestyle=qc_linestyle[qi],
- lw=1.4,
- color=qc_colors[qi],
- label=qc_names[qi],
- alpha=0.9,
- )
- ax[feature_idx, 0].set_ylabel(plot_values[feature_idx])
- ax[feature_idx, 0].grid(True)
- ax[feature_idx, 0].set_xticks([25, 30, 35])
- if feature_idx == len(plot_values) - 1:
- ax[feature_idx, 0].set_xlabel("GA")
- ax[feature_idx, 0].set_xticklabels([25, 30, 35], rotation=30, ha="right")
- else:
- ax[feature_idx, 0].set_xticklabels([])
- ax[0, 0].set_title("Average trend")
- # ---- Columns 1-2: shift / scale terms ----
- for feature_idx in range(len(plot_dict)):
- shift_mat = np.zeros((n_qc, n_sites), dtype=float)
- scale_mat = np.zeros((n_qc, n_sites), dtype=float)
- brain_cols_idx = brain_cols.index(plot_keys[feature_idx])
- for qi, model in enumerate(models):
- for sj, site_idx in enumerate(site_indices):
- _f, sh, sc = _calc_site_terms(model, site_idx, brain_cols_idx, ga_grid)
- shift_mat[qi, sj] = sh
- scale_mat[qi, sj] = sc
- for sj, site_lab in enumerate(site_labels):
- ax[feature_idx, 1].plot(
- qc_x[::-1],
- shift_mat[:, sj],
- marker="o",
- lw=2,
- color=site_colors[sj],
- label=site_lab,
- )
- ax[feature_idx, 1].axhline(0, color="black", lw=0.5, linestyle="--")
- ax[feature_idx, 2].plot(
- qc_x[::-1],
- scale_mat[:, sj],
- marker="o",
- lw=2,
- color=site_colors[sj],
- label=site_lab,
- )
- ax[feature_idx, 2].axhline(1, color="black", lw=0.5, linestyle="--")
- if feature_idx == 0:
- ax[feature_idx, 1].set_title("Shift")
- ax[feature_idx, 2].set_title("Scale")
- ax[feature_idx, 1].set_xticks(qc_x[::-1])
- ax[feature_idx, 2].set_xticks(qc_x[::-1])
- if feature_idx == len(plot_values) - 1:
- ax[feature_idx, 1].set_xticklabels([str(x) for x in qc_names], rotation=30, ha="right")
- ax[feature_idx, 2].set_xticklabels([str(x) for x in qc_names], rotation=30, ha="right")
- else:
- ax[feature_idx, 1].set_xticklabels([])
- ax[feature_idx, 2].set_xticklabels([])
- # ---- Formatting: consistent ylabel x-position ----
- # Lock all left-column ylabels at the same x coordinate so tick label width changes don't shift them.
- for r in range(len(plot_values)):
- ax[r, 0].yaxis.set_label_coords(-0.22, 0.5)
- # ---- Remove per-axis legends and add one global legend below ----
- for r in range(len(plot_values)):
- for c in range(3):
- leg = ax[r, c].get_legend()
- if leg is not None:
- leg.remove()
- # Collect handles/labels (QC lines from [0,0], site lines from [0,1])
- qc_handles, qc_labels = ax[0, 0].get_legend_handles_labels()
- site_handles, site_labels2 = ax[0, 1].get_legend_handles_labels()
- # Deduplicate (separately) while preserving order
- seen_qc = set()
- qc_h = []
- qc_l = []
- for h, lab in zip(qc_handles, qc_labels):
- if lab not in seen_qc:
- seen_qc.add(lab)
- qc_h.append(h)
- qc_l.append(lab)
- seen_site = set()
- site_h = []
- site_l = []
- for h, lab in zip(site_handles, site_labels2):
- if lab not in seen_site:
- seen_site.add(lab)
- site_h.append(h)
- site_l.append(lab)
- # Make space for legends at bottom
- fig.tight_layout(rect=[0, 0.10, 1, 1])
- # Two grouped legends (QC level on left, Site on right)
- leg_qc = fig.legend(
- qc_h,
- qc_l,
- title="QC level",
- loc="lower left",
- bbox_to_anchor=(0.02, 0.08),
- ncol=min(len(qc_l), 4),
- frameon=False,
- handlelength=3,
- columnspacing=1.2,
- )
- leg_site = fig.legend(
- site_h,
- site_l,
- title="Site",
- loc="lower right",
- bbox_to_anchor=(0.98, 0.08),
- ncol=min(len(site_l), 4),
- frameon=False,
- handlelength=2.5,
- columnspacing=1.2,
- )
- # Ensure both legends are kept
- fig.add_artist(leg_qc)
- fig.add_artist(leg_site)
- fig.show()
- # %%
- idx = 5
- struc = brain_cols[idx]
- harm_col = f"{struc}_harm"
- # ----- Plot 1: For a given site, plot category-wise site_trajectory -----
- site_names = list(site_trajectories[0].keys()) # extracted from first entry
- selected_site = site_names[1] # CHANGE to your site of interest
- ga_grid = np.linspace(22, 38, 100)
- model = models[-1]
- model2 = models[0]
- m_at_ga = predict_gam_on_grid(model, ga_grid) # (n_grid, n_features)
- m_at_ga2 = predict_gam_on_grid(model2, ga_grid) # (n_grid, n_features)
- plt.figure(figsize=(8, 5))
- for y_pred, site_dict, d in zip(y_preds, site_trajectories, categories):
- if selected_site in site_dict:
- traj = site_dict[selected_site][:, idx]
- plt.plot(ga_grid, traj, label=f'{d}')
- # Add observed data for the given site
- plt.scatter(
- data_df2["age_at_scan"],
- data_df2[struc],
- c=data_df2["qcglobal"],
- #marker=data_df2["batch"].astype(int),
- s=8, alpha=0.5, label="Observed"
- )
- plt.plot(ga_grid, m_at_ga[:, idx], "k--")
- plt.plot(ga_grid, m_at_ga2[:, idx], "k:")
- plt.title(f"Site Trajectory for {struc} at site {selected_site} (category-wise)")
- plt.xlabel("age_at_scan")
- plt.ylabel(f"{struc} (harmonized units)")
- plt.legend(title="Quality Category")
- plt.grid(True)
- plt.show()
- # %%
- def compute_E1_errors(y_preds):
- E1_errors = []
- E1_error_norm = []
- for y_pred, d in zip(y_preds[1:], categories[1:]):
- e1 = np.mean(y_pred - y_preds[0],axis=0)
- E1_errors.append(e1)
- norm = np.mean(y_preds[0],axis=0)
- E1_error_norm.append(e1/norm)
- return E1_errors, E1_error_norm
- errors, errors_norm = compute_E1_errors(y_preds)
- category_labels = categories[1:]
- n_categories = len(category_labels)
- n_structures = len(plot_dict)
- error_array = np.vstack(errors_norm)
- error_array = np.vstack([error_array[:, brain_cols.index(key)] for key in plot_dict.keys()]).T
- x = np.arange(n_structures)
- width = 0.25
- color_map = {
- "accept_plus": "#FF7256", # coral1
- "poor_plus": "#458B00", # chartreuse4
- "all": "#6495ED", # cornflowerblue
- }
- labels_str = {
- "accept_plus": "AcceptPlus",
- "poor_plus": "PoorPlus",
- "all": "All",
- }
- fig, ax = plt.subplots(figsize=(7, 5))
- for i, cat in enumerate(category_labels):
- ax.bar(x + (i - (n_categories-1)/2)*width, error_array[i], width,
- label=labels_str[cat], color=color_map[cat], edgecolor='none')
- ax.axhline(0, color='black', lw=1.2, ls='--')
- ax.set_xticks(x)
- ax.set_xticklabels(plot_dict.values(), rotation=60, ha="right")
- ax.set_ylabel('Mean E1 Error')
- ax.set_title('Mean E1 Error over average trend\n across harmonization categories')
- ax.legend(title='Quality Group', bbox_to_anchor=(1.00, 1.04), loc='upper left')
- plt.tight_layout()
- plt.show()
- # %% [markdown]
- # ## Sort harmonize
- # For the experiment with sorting and removing iteratively data.
- # %%
- # Sort data_df by qcglobal ascending
- data_df = data_df.sort_values(by="qcglobal", ascending=True).reset_index(drop=True)
- # %%
- dfs = []
- os.makedirs(os.path.join(output_dir,"iterative"), exist_ok=True)
- for idx in range(0, 605, 5):
- print(f"Harmonizing for indices {idx} onwards")
- df2 = data_df[idx:].copy()
- harmonized_df = df2.copy()
- data_df2 = df2.copy()
- covars = get_covars(data_df2)
- scaler = StandardScaler()
- data_train_scaled = scaler.fit_transform(data_df2[brain_cols])
- model, data_train_harmo = harmonizationLearn(data_train_scaled, covars, smooth_terms=["age"])
- data_train_harmo = scaler.inverse_transform(data_train_harmo)
- harmonized_df[[c+"_harm" for c in brain_cols]] = data_train_harmo
- harmonized_df.to_csv(os.path.join(output_dir,"iterative", f"harmonized_data_{idx}.csv"), index=False)
- dfs.append(harmonized_df)
- # %%
1_harmonize_data.ipynb at commit 24d93cb, no license · at the source
Overview
13 affiliations
- CIBM Center for Biomedical Imaging, Lausanne, Switzerland
- Department of Medical Radiology, Lausanne University Hospital and University of Lausanne, Lausanne, Switzerland
- Institut de Neurosciences de la Timone, CNRS, Aix-Marseille Université, Marseille, France
- BCN MedTech, Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain
- APHM, Service de Neuroradiologie Diagnostique et Interventionnelle, Hôpital de la Timone, Aix-Marseille University, Marseille, France
- APHM, Service de Neurologie Pédiatrique, Hôpital de la Timone, Aix-Marseille University, Marseille, France
- Aix Marseille University, Inserm, Marseille, France
- BCNatal | Fetal Medicine Research Center (Hospital Clínic and Hospital Sant Joan de Déu, Universitat de Barcelona), Barcelona, Spain
- Institut d’Investigacions Biomèdiques August Pi i Sunyer (IDIBAPS), Barcelona, Spain
- Centre for Biomedical Research on Rare Diseases (CIBERER), Barcelona, Spain
- Department Woman-Mother-Child, Lausanne University Hospital and University of Lausanne, Lausanne, Switzerland
- School of Health Sciences (HESAV), HES-SO, University of Applied Sciences and Arts Western Switzerland, Lausanne, Switzerland
- ICREA, Barcelona, Spain
Abstract
Normative modeling is increasingly used to characterize typical growth trajectories and identify atypical neurodevelopment, including early brain development using magnetic resonance imaging (MRI) acquired before birth. Recent work has emphasized the importance of large sample sizes for accurate and robust centile estimation. In this study, we investigate how image quality influences fetal brain normative models, a critical factor in this context where MRI is acquired on a moving fetus in utero. Using a multi-centric cohort of 635 fetal MRI scans, we applied a standardized visual quality control (QC) protocol with continuous quality ratings. We fit normative models for multiple brain structures under progressively relaxed QC stringency, and quantified the deviations in centile estimates relative to a high-quality reference subgroup. Our results showed that including lower-quality data systematically biased normative centiles, with the strongest effects observed in the outer centiles, particularly the lower tail (1st–10th). Bias increased progressively as QC stringency was relaxed and could not be attributed solely to the number of scans used to fit the models. Quality-induced bias was structure dependent, and often not visually apparent at the segmentation level. These findings highlight that image quality is an important source of bias in normative fetal brain modeling, and that increasing sample size at the expense of quality may systematically affect centile estimates, potentially jeopardizing the utility of the model.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.
fetpype/fetpype
5c99de9a3b35734b2fa71b5fcde2920cb9b1d93b, 3 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
29 files
- fetpype/
__init__.py , Python, 19 lines - fetpype/
_version.py , Python, 1 line - fetpype/
definitions.py , Python, 60 lines - fetpype/
nodes/ , Python, 1 line__init__.py - fetpype/
nodes/ , Python, 102 linesdhcp.py - fetpype/
nodes/ , Python, 675 linespreprocessing.py - fetpype/
nodes/ , Python, 291 linesreconstruction.py - fetpype/
nodes/ , Python, 96 linessegmentation.py - fetpype/
nodes/ , Python, 95 linessurface_extraction.py - fetpype/
nodes/ , Python, 61 linesutils.py - fetpype/
pipelines/ , Python, 1 line__init__.py - fetpype/
pipelines/ , Python, 785 lines, 1 matchfull_pipeline.py - fetpype/
utils/ , Python, 1 line__init__.py - fetpype/
utils/ , Python, 289 lineslogging.py - fetpype/
utils/ , Python, 765 linesutils_bids.py - fetpype/
utils/ , Python, 196 linesutils_docker.py - fetpype/
workflows/ , Python, 1 line__init__.py - fetpype/
workflows/ , Python, 293 linespipeline_fet.py - fetpype/
workflows/ , Python, 320 linespipeline_fet_dhcp.py - fetpype/
workflows/ , Python, 196 linespipeline_rec.py - fetpype/
workflows/ , Python, 212 linespipeline_seg.py - fetpype/
workflows/ , Python, 218 linespipeline_surf.py - fetpype/
workflows/ , Python, 241 linesutils.py - tests/
__init__.py , Python, 1 line - tests/
conftest.py , Python, 129 lines - tests/
test_pipelines.py , Python, 71 lines - tests/
test_utils_bids.py , Python, 366 lines - LICENSE, License, 21 lines
- README.md, Text, 34 lines
fetpype/normative_modeling_data_quality
24d93cbe628568ac5ff3df680bd6046fdd8e393e, 22 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
13 files
- notebooks/
1_harmonize_data.ipynb , Jupyter, 814 lines, 3 matches - scripts/
1_run_main_analysis.r , R, 183 lines - scripts/
2_run_fixed_harmonizatio , R, 180 linesn_ablation.r - scripts/
3_run_bootstrapping_expe , R, 172 linesriment.r - scripts/
plots_and_analyses/ , R, 254 lines0_plot_harmonization_eff ect.r - scripts/
plots_and_analyses/ , R, 559 lines, 2 matches1_plot_main.r - scripts/
plots_and_analyses/ , R, 309 lines2_plot_fixed_harmonizati on.r - scripts/
plots_and_analyses/ , R, 195 lines3_plot_bootstrapping.r - scripts/
plots_and_analyses/ , R, 71 lines3_statistical_analysis_b ootstrapping.r - scripts/
plots_and_analyses/ , R, 186 lines, 2 matchessupp_plot_iterative_remo val.r - scripts/
supp_harmonize_remove_co , Python, 260 lineshort.py - scripts/
supp_run_iterative_remov , R, 171 linesal.r - README.md, Text, 42 lines
Zenodo 15696638
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 39 scripts, each with its path and the digest of its content;
- 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:21485413, at Zenodo; found in DataCite
- zenodo:21485414, at Zenodo; found in “Data and Code Availability”
Data and Code Availability
Tabular data (estimated volumes data and corresponding data quality for each scan) are available on a Zenodo repository at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 16 authors, 5 keywords, 4 funders, 42 references.
Cite
This paper
Sanchez, T., Mihailov, A., Martí-Juan, G., Girard, N., Manchon, A., Milh, M., Eixarch, E., Dunet, V., Koob, M., Pomar, L., Sichitiu, J., González Ballester, M. A., Camara, O., Piella, G., Bach Cuadra, M., & Auzias, G. (2026). Data quality biases normative models derived from fetal brain MRI. Imaging neuroscience (Cambridge, Mass.), 4, IMAG.a.1337. https://
BibTeX
@article{sanchez2026data
author = {Sanchez, Thomas and Mihailov, Angeline and Martí-Juan, Gerard and Girard, Nadine and Manchon, Aurélie and Milh, Mathieu and Eixarch, Elisenda and Dunet, Vincent and Koob, Mériam and Pomar, Léo and Sichitiu, Joanna and González Ballester, Miguel A. and Camara, Oscar and Piella, Gemma and Bach Cuadra, Meritxell and Auzias, Guillaume},
title = {{Data quality biases normative models derived from fetal brain MRI}},
journal = {Imaging neuroscience (Cambridge, Mass.)},
year = {2026},
month = aug,
volume = {4},
pages = {IMAG.a.1337},
publisher = {MIT Press},
issn = {2837-6056},
doi = {10.1162/
url = {https://
pmid = {42609916},
pmcid = {PMC13479347}
}
RIS
TY - JOUR
AU - Sanchez, Thomas
AU - Mihailov, Angeline
AU - Martí-Juan, Gerard
AU - Girard, Nadine
AU - Manchon, Aurélie
AU - Milh, Mathieu
AU - Eixarch, Elisenda
AU - Dunet, Vincent
AU - Koob, Mériam
AU - Pomar, Léo
AU - Sichitiu, Joanna
AU - González Ballester, Miguel A.
AU - Camara, Oscar
AU - Piella, Gemma
AU - Bach Cuadra, Meritxell
AU - Auzias, Guillaume
TI - Data quality biases normative models derived from fetal brain MRI
T2 - Imaging neuroscience (Cambridge, Mass.)
J2 - Imaging Neurosci (Camb)
PY - 2026
DA - 2026/
VL - 4
SP - IMAG.a.1337
SN - 2837-6056
PB - MIT Press
DO - 10.1162/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1162/
"type": "article-journal",
"title": "Data quality biases normative models derived from fetal brain MRI",
"container-title": "Imaging neuroscience (Cambridge, Mass.)",
"author": [
{
"family": "Sanchez",
"given": "Thomas"
},
{
"family": "Mihailov",
"given": "Angeline"
},
{
"family": "Martí-Juan",
"given": "Gerard"
},
{
"family": "Girard",
"given": "Nadine"
},
{
"family": "Manchon",
"given": "Aurélie"
},
{
"family": "Milh",
"given": "Mathieu"
},
{
"family": "Eixarch",
"given": "Elisenda"
},
{
"family": "Dunet",
"given": "Vincent"
},
{
"family": "Koob",
"given": "Mériam"
},
{
"family": "Pomar",
"given": "Léo"
},
{
"family": "Sichitiu",
"given": "Joanna"
},
{
"family": "González Ballester",
"given": "Miguel A."
},
{
"family": "Camara",
"given": "Oscar"
},
{
"family": "Piella",
"given": "Gemma"
},
{
"family": "Bach Cuadra",
"given": "Meritxell"
},
{
"family": "Auzias",
"given": "Guillaume"
}
],
"container-title-short":
"volume": "4",
"page": "IMAG.a.1337",
"DOI": "10.1162/
"PMID": "42609916",
"PMCID": "PMC13479347",
"ISSN": "2837-6056",
"publisher": "MIT Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-73072-6 [code]
- Mapping the spatiotemporal continuum of structural connectivity development across the human connectome in youth.Journal: Nature communicationsIn common: neuroHarmonize, mgcv, patchwork, 5 other tools, structural MRI / diffusion, 4 references
- [2] doi:10.1162/imag.a.1269 [code]
- From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.Journal: Imaging neuroscience (Cambridge, Mass.)In common: NiBabel, seaborn, scikit-learn, 4 other tools, 7 references
- [3] doi:10.1162/imag.a.1347 [code]
- Neural and behavioural correlates of theory of mind reasoning in five-year-old children born preterm.Journal: Imaging neuroscience (Cambridge, Mass.)In common: PyBIDS, Nipype, FSL, 9 other tools
- [4] doi:10.1162/imag.a.1245 [code]
- Towards precision EEG connectomics: Evaluating the benefits of dense sampling.Journal: Imaging neuroscience (Cambridge, Mass.)In common: Nipype, FSL, cowplot, 9 other tools
- [5] doi:10.1002/hbm.70605 [code]
- BrainEnrich: Revealing Biological Insights for Imaging-Derived Phenotypes Through Transcriptomic Enrichment.Journal: Human brain mappingIn common: nlme, cowplot, patchwork, 9 other tools, 1 reference
- [6] doi:10.1016/j.dcn.2026.101765 [code]
- Fusiform face area development correlates with development in higher-order social brain regions.Journal: Developmental cognitive neuroscienceIn common: PyBIDS, Nipype, FSL, 8 other tools
- [7] doi:10.1038/s42003-026-10276-y [code]
- The cellular correlates and adolescent reorganisation of cortical myelination networks in the common marmoset.Journal: Communications biologyIn common: Nipype, mgcv, patchwork, 8 other tools, structural MRI / diffusion
- [8] doi:10.1162/imag.a.1276 [code]
- High-resolution whole-brain magnetic resonance spectroscopic imaging in youth at risk for psychosis.Journal: Imaging neuroscience (Cambridge, Mass.)In common: FSL, NiBabel, seaborn, 5 other tools, author Meritxell Bach Cuadra
- [9] doi:10.1371/journal.pone.0343722 [code]
- Comprehensive methodology for sample enrichment in EEG biomarker studies for Alzheimer's risk classification.Journal: PloS oneIn common: neuroHarmonize, PyBIDS, ggplot2, 6 other tools, 1 reference
- [10] doi:10.3389/fnins.2026.1873417 [code]
- Functional brain network alterations in a rat model of dental malocclusion.Journal: Frontiers in neuroscienceIn common: PyBIDS, Nipype, FSL, 7 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 39 scripts, and 8 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:d56ee56d527aacba…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
