Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data.
The 8 matches
- [1] § Methods › Machine learning classification › Classification pipelines. ↔ src/an_heterogeneity/tools/model_training.py, lines 107–212 · score 0.75 · model training, Hyperparameter optimization, cross validation, scikit-learn, nested, fold
- [2] § Methods › Machine learning classification › Hyperparameters optimization and performance estimation. ↔ src/an_heterogeneity/tools/model_pipelines.py, lines 36–71 · score 0.68 · precision recall, model pipeline, PR AUC, optimization, curve, classes
- [3] § Methods › Machine learning classification › Hyperparameters optimization and performance estimation. ↔ src/an_heterogeneity/model_evaluation/cross_validation.py, lines 43–158 · score 0.58 · class ratios, cross validation, trained, curve, PR, ROC
- [4] § Methods › Image acquisition and processing ↔ src/an_heterogeneity/tools/load_parse_neuromaps.py, lines 33–51 · score 0.58 · Desikan Killiany, FreeSurfer, atlas
- [5] § Methods › Normative modeling ↔ src/an_heterogeneity/tools/normative_tools.py, lines 873–919 · score 0.57 · infra normal, regional CT, supranormal, deviation, normative, SA
- [6] § Results › z-scores from CentileBrain normative model › AN versus HC. ↔ src/an_heterogeneity/tools/normative_tools.py, lines 873–919 · score 0.55 · subcortical volume, cortical thickness, surface area, threshold, supranormal, infranormal
- [7] § Results › Univariate comparisons ↔ src/an_heterogeneity/ANHC_univariate.ipynb, lines 461–472 · score 0.51 · FDR correction, CI, SD, ventricles, Univariate
- [8] § Results › Machine learning classification ↔ src/an_heterogeneity/model_evaluation/performance_metric_visualization.py, lines 89–144 · score 0.50 · precision recall, PR AUC, baseline, metrics
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 1,056 lines · 44 KB · no license · 2 matches
- import os
- import pandas as pd
- import numpy as np
- import seaborn as sns
- from docx import Document
- from docx.shared import Pt
- from docx.enum.text import WD_PARAGRAPH_ALIGNMENT
- from docx.oxml import OxmlElement
- from docx.oxml.ns import qn
- import scipy.stats as stats
- from statsmodels.stats.proportion import proportions_ztest
- from statsmodels.stats.multitest import multipletests
- import matplotlib.pyplot as plt
- import matplotlib.colorbar as colorbar
- import matplotlib.colors as colors
- from PIL import Image, ImageDraw, ImageFont
- # MOST IMPORTANT FUNCTION:
- def compute_deviation_percentages(df, thr=1.96):
- """
- Compute the percentage of supra- and infranormal deviations (z > 1.96 and z < -1.96, respectively)
- for each column in a dataframe, along with p-values, FDR-corrected p-values, and FDR thresholds.
- Args:
- df (pd.DataFrame): A dataframe containing z-scores for participants (rows) and measures (columns).
- Returns:
- pd.DataFrame: A dataframe with columns:
- ['Measure', 'Supra (%)', 'Infra (%)', 'Supra p-value', 'Infra p-value',
- 'Supra FDR p-value', 'Infra FDR p-value', 'FDR Threshold Supra (%)',
- 'FDR Threshold Infra (%)']
- """
- # Expected proportion of supra/infranormal values under the null hypothesis
- pthr = {1.96:0.025, 1.28:0.05, 0.84:0.1}
- expected_proportion = pthr[thr]
- # Total number of participants (rows)
- n = len(df)
- # Calculate the percentage threshold for uncorrected significance
- uncorrected_threshold = stats.binom.ppf(0.95, n, expected_proportion) / n * 100
- # Initialize results
- results = []
- supra_p_values = []
- infra_p_values = []
- # Iterate through each column
- for col in df.columns:
- # Calculate supra- and infranormal percentages
- supra_percentage = (df[col] > thr).mean() * 100
- infra_percentage = (df[col] < -thr).mean() * 100
- # Calculate p-values for supra- and infranormal percentages
- supra_count = (df[col] > thr).sum()
- infra_count = (df[col] < -thr).sum()
- supra_p_value = stats.binomtest(supra_count, n, expected_proportion, alternative='greater').pvalue
- infra_p_value = stats.binomtest(infra_count, n, expected_proportion, alternative='greater').pvalue
- supra_p_values.append(supra_p_value)
- infra_p_values.append(infra_p_value)
- # Append results for the column
- results.append({
- 'Measure': col,
- 'Supra (%)': supra_percentage,
- 'Infra (%)': infra_percentage,
- 'Supra p-value': supra_p_value,
- 'Infra p-value': infra_p_value
- })
- # Correct p-values using FDR
- supra_fdr_corrected = multipletests(supra_p_values, alpha=0.05, method='fdr_bh')[1]
- infra_fdr_corrected = multipletests(infra_p_values, alpha=0.05, method='fdr_bh')[1]
- # Compute FDR thresholds dynamically
- fdr_threshold_supra = min(
- [results[i]['Supra (%)'] for i in range(len(results)) if supra_fdr_corrected[i] < 0.05],
- default=100
- )
- fdr_threshold_infra = min(
- [results[i]['Infra (%)'] for i in range(len(results)) if infra_fdr_corrected[i] < 0.05],
- default=100
- )
- # Add FDR-corrected p-values and thresholds to results
- for i, row in enumerate(results):
- row['Supra FDR p-value'] = supra_fdr_corrected[i]
- row['Infra FDR p-value'] = infra_fdr_corrected[i]
- row['FDR Threshold Supra (%)'] = fdr_threshold_supra
- row['FDR Threshold Infra (%)'] = fdr_threshold_infra
- # Convert results to a dataframe
- results_df = pd.DataFrame(results)
- return results_df
- # Example Usage:
- # results_ANbp = compute_deviation_percentages(df_T1.loc[(df_T1.dx==1) & (df_T1.subtype==1),thicknesses], thr=1.96)
- def test_global_heterogeneity_mean_sd(z_matrix, n_bootstrap=20000, random_state=42, verbose=True):
- """
- Assess whether the mean SD across ROIs (columns) is higher than expected under null using bootstrap.
- Parameters:
- - z_matrix: pd.DataFrame, z-score matrix (subjects x ROIs)
- - n_bootstrap: int, number of bootstrap samples
- - random_state: int, random seed for reproducibility
- - verbose: bool, if True prints result
- Returns:
- - observed_mean_sd: float, mean SD across ROIs in actual data
- - p_value: float, one-sided p-value
- - null_distribution: np.array, mean SDs from null model
- """
- np.random.seed(random_state)
- n_subjects, n_rois = z_matrix.shape
- # 1. Compute observed mean SD across ROIs
- observed_sds = z_matrix.std(axis=0, ddof=1)
- observed_mean_sd = observed_sds.mean()
- # 2. Bootstrap null distribution under standard normal
- null_distribution = []
- for _ in range(n_bootstrap):
- synthetic_data = np.random.normal(loc=0, scale=1, size=(n_subjects, n_rois))
- synthetic_sds = synthetic_data.std(axis=0, ddof=1)
- null_distribution.append(np.mean(synthetic_sds))
- null_distribution = np.array(null_distribution)
- # 3. Compute one-sided p-value (observed > null)
- p_value = np.mean(null_distribution >= observed_mean_sd)
- if verbose:
- print(f"Observed mean SD across ROIs: {observed_mean_sd:.4f}")
- print(f"Mean of null distribution: {null_distribution.mean():.4f}")
- print(f"One-sided p-value: {p_value:.4f}")
- return observed_mean_sd, p_value, null_distribution
- # OUTDATED FUNCTION COMPARING THE FREQUENCY OF EXTREME Z-SCORES BETWEEN GROUPS, BUT THIS DOES NOT MAKE MUCH SENSE SINCE IT
- # IS BETTER TO COMPARE THE FREQUENCY OF EXTREME Z-SCORES BETWEEN AN AND THE NORMATIVE REFERENCE
- def group_compare_extreme_zs(df, feats, group='dx', labels_dict={'HC': 0, 'AN': 1}, thr=1.96, mult_comp=None):
- # this function assesses whether there are significant differences for each measure in feats in the
- # proportion of supra/infra-normal deviations (it plots a table with p-values computed according to Fisher, Fisher Chi2, and Z tests)
- results = []
- p_values_pos_fisher_chi2 = []
- p_values_neg_fisher_chi2 = []
- p_values_pos_ztest = []
- p_values_neg_ztest = []
- stats_pos_ztest = []
- stats_neg_ztest = []
- fisher_used_pos = []
- fisher_used_neg = []
- for feature in feats:
- row = {}
- for label, group_value in labels_dict.items():
- row[f'{label}+'] = ((df[df[group] == group_value][feature] > thr).mean() * 100)
- row[f'{label}-'] = ((df[df[group] == group_value][feature] < -thr).mean() * 100)
- # Count of outliers and non-outliers for Fisher/Chi-Square Test
- pos_outliers_group = (df[df[group] == 1][feature] > thr).sum()
- pos_total_group = (df[group] == 1).sum()
- pos_outliers_control = (df[df[group] == 0][feature] > thr).sum()
- pos_total_control = (df[group] == 0).sum()
- neg_outliers_group = (df[df[group] == 1][feature] < -thr).sum()
- neg_total_group = (df[group] == 1).sum()
- neg_outliers_control = (df[df[group] == 0][feature] < -thr).sum()
- neg_total_control = (df[group] == 0).sum()
- contingency_table_pos = [[pos_outliers_group, pos_outliers_control],
- [pos_total_group - pos_outliers_group, pos_total_control - pos_outliers_control]]
- contingency_table_neg = [[neg_outliers_group, neg_outliers_control],
- [neg_total_group - neg_outliers_group, neg_total_control - neg_outliers_control]]
- # Check and perform Fisher's Exact Test or Chi-Square Test
- fisher_flag_pos = min(pos_outliers_group, pos_outliers_control, pos_total_group - pos_outliers_group,
- pos_total_control - pos_outliers_control) < 10
- fisher_flag_neg = min(neg_outliers_group, neg_outliers_control, neg_total_group - neg_outliers_group,
- neg_total_control - neg_outliers_control) < 10
- fisher_used_pos.append(int(fisher_flag_pos))
- fisher_used_neg.append(int(fisher_flag_neg))
- if fisher_flag_pos:
- _, p_pos_fisher_chi2 = stats.fisher_exact(contingency_table_pos)
- else:
- _, p_pos_fisher_chi2, _, _ = stats.chi2_contingency(contingency_table_pos)
- if fisher_flag_neg:
- _, p_neg_fisher_chi2 = stats.fisher_exact(contingency_table_neg)
- else:
- _, p_neg_fisher_chi2, _, _ = stats.chi2_contingency(contingency_table_neg)
- # Two Proportions Z-Test (with check for valid inputs)
- if any(n > 0 for n in [pos_outliers_group, pos_outliers_control]):
- stat_pos_ztest, p_pos_ztest = proportions_ztest([pos_outliers_group, pos_outliers_control],
- [pos_total_group, pos_total_control])
- else:
- stat_pos_ztest = 0
- p_pos_ztest = 1
- if any(n > 0 for n in [neg_outliers_group, neg_outliers_control]):
- stat_neg_ztest, p_neg_ztest = proportions_ztest([neg_outliers_group, neg_outliers_control],
- [neg_total_group, neg_total_control])
- else:
- stat_neg_ztest = 0
- p_neg_ztest = 1
- p_values_pos_fisher_chi2.append(p_pos_fisher_chi2)
- p_values_neg_fisher_chi2.append(p_neg_fisher_chi2)
- p_values_pos_ztest.append(p_pos_ztest)
- p_values_neg_ztest.append(p_neg_ztest)
- stats_pos_ztest.append(stat_pos_ztest)
- stats_neg_ztest.append(stat_neg_ztest)
- results.append(row)
- # Multiple comparisons correction
- if mult_comp == 'fdr':
- corrected_p_values_pos_fisher_chi2 = multipletests(p_values_pos_fisher_chi2, method='fdr_bh')[1]
- corrected_p_values_neg_fisher_chi2 = multipletests(p_values_neg_fisher_chi2, method='fdr_bh')[1]
- corrected_p_values_pos_ztest = multipletests(p_values_pos_ztest, method='fdr_bh')[1]
- corrected_p_values_neg_ztest = multipletests(p_values_neg_ztest, method='fdr_bh')[1]
- else:
- corrected_p_values_pos_fisher_chi2 = p_values_pos_fisher_chi2
- corrected_p_values_neg_fisher_chi2 = p_values_neg_fisher_chi2
- corrected_p_values_pos_ztest = p_values_pos_ztest
- corrected_p_values_neg_ztest = p_values_neg_ztest
- for i, row in enumerate(results):
- row['Fisher_Chi2_Pos_Test'] = corrected_p_values_pos_fisher_chi2[i]
- row['Fisher_Chi2_Neg_Test'] = corrected_p_values_neg_fisher_chi2[i]
- row['ZTest_Pos_Test'] = corrected_p_values_pos_ztest[i]
- row['ZTest_Neg_Test'] = corrected_p_values_neg_ztest[i]
- row['ZTest_Pos_Stat'] = stats_pos_ztest[i]
- row['ZTest_Neg_Stat'] = stats_neg_ztest[i]
- row['Fisher_Pos'] = fisher_used_pos[i]
- row['Fisher_Neg'] = fisher_used_neg[i]
- return pd.DataFrame(results, index=feats)
- # Example usage:
- # result_df = group_compare_extreme_zs(df, feats)
- # print(result_df)
- def set_cell_border(cell, **kwargs):
- """
- Set cell border to the specified styles.
- Usage: set_cell_border(cell, top="single", bottom="single", start="single", end="single")
- """
- tc = cell._element
- tcPr = tc.get_or_add_tcPr()
- for edge in ('top', 'start', 'bottom', 'end'):
- if edge in kwargs:
- tag = f'w:{edge}'
- element = OxmlElement(tag)
- element.set(qn('w:val'), kwargs[edge])
- element.set(qn('w:sz'), '4') # Size 4 for thin border
- element.set(qn('w:space'), '0')
- element.set(qn('w:color'), 'auto')
- tcPr.append(element)
- def format_number(value):
- """
- Format the number to have at most 4 digits after the comma.
- """
- if isinstance(value, (int, float)):
- return f"{value:.2f}"
- return str(value)
- def dataframe_to_word(df, filename):
- # Create a new Document
- doc = Document()
- # Add a table with an extra column for the index
- table = doc.add_table(rows=df.shape[0] + 1, cols=df.shape[1] + 1)
- # Add the header row
- table.cell(0, 0).text = "Index"
- set_cell_border(table.cell(0, 0), bottom="single")
- for j, column in enumerate(df.columns):
- table.cell(0, j + 1).text = column
- set_cell_border(table.cell(0, j + 1), bottom="single")
- # Add the rest of the DataFrame rows including the index
- for i in range(df.shape[0]):
- table.cell(i + 1, 0).text = str(df.index[i])
- set_cell_border(table.cell(i + 1, 0), right="single")
- for j in range(df.shape[1]):
- table.cell(i + 1, j + 1).text = format_number(df.iat[i, j])
- # Save the document
- doc.save(filename)
- print(f"Table saved to {filename}")
- def analyze_deviations(dataframe, columns, group_column, min_infrasupra=1, thr=1.96):
- """
- Analyze the percentage of participants with supranormal (z > 1.96) or infranormal (z < -1.96) deviations
- across specified columns, grouped by a specific group column. Includes chi-squared tests for group differences.
- Args:
- dataframe (pd.DataFrame): The input dataframe containing participant data.
- columns (list): List of column names to analyze for deviations.
- group_column (str): Name of the column indicating group membership.
- min_infrasupra (int or float, optional): Minimum number of deviations required.
- If between 0 and 0.1, interprets it as a p-value and computes the threshold
- using the binomial distribution.
- Returns:
- str: A textual summary of the results.
- """
- # Define thresholds for deviations
- nc = len(columns) # Number of columns (regions of interest)
- pthr = {1.96:0.025, 1.28:0.05, 0.84:0.1}
- if 0 < min_infrasupra < 0.1:
- p = pthr[thr] # Probability of |z| > 1.96
- mu = nc * p
- sigma = np.sqrt(nc * p * (1 - p))
- threshold = int(binom.ppf(1 - min_infrasupra / 2, n=nc, p=p))
- threshold_text = f"at least {threshold}"
- else:
- threshold = min_infrasupra
- threshold_text = f"at least {threshold}"
- # Create masks for deviations
- supranormal_mask = (dataframe[columns] > thr).sum(axis=1) >= threshold
- infranormal_mask = (dataframe[columns] < -thr).sum(axis=1) >= threshold
- any_deviation_mask = supranormal_mask | infranormal_mask
- # Group data by the specified group column
- grouped = dataframe.groupby(group_column)
- # Calculate percentages for each group
- results = {}
- for group, group_data in grouped:
- total = len(group_data)
- if total == 0:
- results[group] = {
- "any_percentage": 0, "any_count": 0,
- "supra_percentage": 0, "supra_count": 0,
- "infra_percentage": 0, "infra_count": 0,
- "total": 0
- }
- continue
- count_with_any_deviation = any_deviation_mask.loc[group_data.index].sum()
- count_with_supra_deviation = supranormal_mask.loc[group_data.index].sum()
- count_with_infra_deviation = infranormal_mask.loc[group_data.index].sum()
- results[group] = {
- "any_percentage": (count_with_any_deviation / total) * 100,
- "any_count": count_with_any_deviation,
- "supra_percentage": (count_with_supra_deviation / total) * 100,
- "supra_count": count_with_supra_deviation,
- "infra_percentage": (count_with_infra_deviation / total) * 100,
- "infra_count": count_with_infra_deviation,
- "total": total
- }
- # Create contingency tables for chi-squared tests
- contingency_table_any = []
- contingency_table_supra = []
- contingency_table_infra = []
- for group in results:
- # Any deviation
- count_with_any_deviation = results[group]["any_count"]
- count_without_any_deviation = results[group]["total"] - count_with_any_deviation
- contingency_table_any.append([count_with_any_deviation, count_without_any_deviation])
- # Supranormal deviation
- count_with_supra_deviation = results[group]["supra_count"]
- count_without_supra_deviation = results[group]["total"] - count_with_supra_deviation
- contingency_table_supra.append([count_with_supra_deviation, count_without_supra_deviation])
- # Infranormal deviation
- count_with_infra_deviation = results[group]["infra_count"]
- count_without_infra_deviation = results[group]["total"] - count_with_infra_deviation
- contingency_table_infra.append([count_with_infra_deviation, count_without_infra_deviation])
- # Perform chi-squared tests
- chi2_any, p_value_any, _, _ = stats.chi2_contingency(contingency_table_any)
- chi2_supra, p_value_supra, _, _ = stats.chi2_contingency(contingency_table_supra)
- chi2_infra, p_value_infra, _, _ = stats.chi2_contingency(contingency_table_infra)
- # Build the textual output
- output = []
- output.append("Analysis of deviations:")
- for group, sts in results.items():
- output.append(
- f"- In group '{group}':\n"
- f" * {sts['any_percentage']:.2f}% ({sts['any_count']} out of {sts['total']}) participants had {threshold_text} supranormal or infranormal deviation.\n"
- f" * {sts['supra_percentage']:.2f}% ({sts['supra_count']} out of {sts['total']}) participants had {threshold_text} supranormal deviation.\n"
- f" * {sts['infra_percentage']:.2f}% ({sts['infra_count']} out of {sts['total']}) participants had {threshold_text} infranormal deviation."
- )
- output.append(f"\nChi-squared test results:")
- output.append(f"- Any deviation:\n * Chi-squared statistic: {chi2_any:.2f}\n * p-value: {p_value_any:.4f}")
- if p_value_any < 0.05:
- output.append(" * The difference in percentages between groups for any deviation is statistically significant.")
- else:
- output.append(" * The difference in percentages between groups for any deviation is not statistically significant.")
- output.append(f"- Supranormal deviation:\n * Chi-squared statistic: {chi2_supra:.2f}\n * p-value: {p_value_supra:.4f}")
- if p_value_supra < 0.05:
- output.append(" * The difference in percentages between groups for supranormal deviations is statistically significant.")
- else:
- output.append(" * The difference in percentages between groups for supranormal deviations is not statistically significant.")
- output.append(f"- Infranormal deviation:\n * Chi-squared statistic: {chi2_infra:.2f}\n * p-value: {p_value_infra:.4f}")
- if p_value_infra < 0.05:
- output.append(" * The difference in percentages between groups for infranormal deviations is statistically significant.")
- else:
- output.append(" * The difference in percentages between groups for infranormal deviations is not statistically significant.")
- return "\n".join(output)
- #############################
- # Unuseful stuff?
- import pandas as pd
- import statsmodels.formula.api as smf
- from statsmodels.stats.multitest import multipletests
- def fit_models_and_extract_values(df, measures, model_terms, multcomp=None, filter_significant=0,
- vars_of_interest=None):
- """
- Fits a linear model for each measure and extracts p-values and t-values for specified factors, with optional filtering.
- Args:
- df (DataFrame): The DataFrame containing the data.
- measures (list): List of measures to fit models for.
- model_terms (list): List of model terms to include in the formula.
- multcomp (str, optional): Multiple comparison correction method. Defaults to None.
- filter_significant (int, optional): If 1, filters results based on significance of vars_of_interest.
- vars_of_interest (list, optional): Variables to check for significance if filtering.
- Returns:
- DataFrame: A DataFrame with each measure and corresponding p-values and t-values for factors.
- """
- results = []
- for measure in measures:
- # Construct the formula
- formula_terms = ' + '.join(model_terms)
- formula = f'{measure} ~ {formula_terms}'
- model = smf.ols(formula, data=df).fit()
- # Extract p-values and t-values
- relevant_terms = [term for term in model_terms if term in model.params.index]
- pvals = model.pvalues[relevant_terms].to_dict()
- tvals = model.tvalues[relevant_terms].to_dict()
- # Combine and add measure
- combined_dict = {**{f'{term}_pval': pvals[term] for term in relevant_terms},
- **{f'{term}_tval': tvals[term] for term in relevant_terms},
- 'measure': measure}
- results.append(combined_dict)
- results_df = pd.DataFrame(results)
- # Multiple comparisons correction
- if multcomp:
- pval_cols = [f'{term}_pval' for term in relevant_terms]
- corrected_pvals = multipletests(results_df[pval_cols].values.flatten(),
- method=multcomp)[1]
- results_df[pval_cols] = corrected_pvals.reshape(results_df[pval_cols].shape)
- # Filter for significant effects of interest
- if filter_significant and vars_of_interest:
- filter_conditions = [(results_df[f'{var}_pval'] < 0.05) for var in vars_of_interest]
- combined_condition = filter_conditions[0]
- for condition in filter_conditions[1:]:
- combined_condition |= condition
- results_df = results_df[combined_condition]
- return results_df
- def set_column_width(column, width_cm):
- # Adjust column width in cm
- for cell in column.cells:
- tc = cell._element
- tcPr = tc.get_or_add_tcPr()
- tcW = OxmlElement('w:tcW')
- tcW.set(qn('w:w'), str(int(width_cm * 567))) # 1cm ≈ 567 units in docx format
- tcW.set(qn('w:type'), 'dxa')
- tcPr.append(tcW)
- def create_word_table_from_dataframe(df, dest_dir, filename):
- # Initialize the document
- doc = Document()
- # Get the covariates by splitting column names (assuming X_pval and X_tval structure)
- covariates = sorted(set(col.split('_')[0] for col in df.columns))
- # Create the table: rows = number of rows in df + 2 (for headers), columns = 1 (feature names) + 2 * number of covariates
- table = doc.add_table(rows=len(df) + 2, cols=1 + 2 * len(covariates))
- # First row: feature names and covariate names
- table.cell(0, 0).text = 'Feature'
- for i, covariate in enumerate(covariates):
- cell = table.cell(0, 1 + 2 * i)
- cell.text = covariate
- cell.merge(table.cell(0, 1 + 2 * i + 1)) # Merge the two columns for covariate name
- # Make covariate name bold and center it
- for paragraph in cell.paragraphs:
- paragraph.alignment = WD_PARAGRAPH_ALIGNMENT.CENTER
- for run in paragraph.runs:
- run.bold = True
- # Second row: 'p' and 't' labels for each covariate
- table.cell(1, 0).text = '' # Blank cell for the feature names column header
- for i, covariate in enumerate(covariates):
- table.cell(1, 1 + 2 * i).text = 'p'
- table.cell(1, 1 + 2 * i + 1).text = 't'
- # Remaining rows: fill in the feature names, p-values, and t-values from the DataFrame
- for row_idx, feature_name in enumerate(df.index):
- # Insert feature name in the first column
- table.cell(row_idx + 2, 0).text = str(feature_name)
- # Insert p-values and t-values in the remaining columns
- for i, covariate in enumerate(covariates):
- pval_col = f'{covariate}_pval'
- tval_col = f'{covariate}_tval'
- # Format p-values to 4 decimal places and t-values to 2 decimal places
- pval = f'{df.iloc[row_idx][pval_col]:.4f}'
- tval = f'{df.iloc[row_idx][tval_col]:.2f}'
- # Insert formatted values into the table
- table.cell(row_idx + 2, 1 + 2 * i).text = pval
- table.cell(row_idx + 2, 1 + 2 * i + 1).text = tval
- # Adjust font size for the entire table (optional, for better readability)
- for row in table.rows:
- for cell in row.cells:
- for paragraph in cell.paragraphs:
- for run in paragraph.runs:
- run.font.size = Pt(10)
- # Set the column width for a more compact layout (width in cm)
- set_column_width(table.columns[0], 2.5) # Feature names column width
- for i in range(1, len(covariates) * 2 + 1):
- set_column_width(table.columns[i], 1.5) # Covariate columns width (for both p and t)
- # Save the document
- output_path = os.path.join(dest_dir, f'{filename}.docx')
- doc.save(output_path)
- return output_path
- import pandas as pd
- import statsmodels.formula.api as smf
- from statsmodels.stats.multitest import multipletests
- def fit_models_and_extract_values(df, measures, model_terms, multcomp=None, filter_significant=0,
- vars_of_interest=None):
- """
- Fits a linear model for each measure and extracts p-values and t-values for specified factors, with optional filtering.
- Args:
- df (DataFrame): The DataFrame containing the data.
- measures (list): List of measures to fit models for.
- model_terms (list): List of model terms to include in the formula.
- multcomp (str, optional): Multiple comparison correction method. Defaults to None.
- filter_significant (int, optional): If 1, filters results based on significance of vars_of_interest.
- vars_of_interest (list, optional): Variables to check for significance if filtering.
- Returns:
- DataFrame: A DataFrame with each measure and corresponding p-values and t-values for factors.
- """
- results = []
- for measure in measures:
- # Construct the formula
- formula_terms = ' + '.join(model_terms)
- formula = f'{measure} ~ {formula_terms}'
- model = smf.ols(formula, data=df).fit()
- # Extract p-values and t-values
- relevant_terms = [term for term in model_terms if term in model.params.index]
- pvals = model.pvalues[relevant_terms].to_dict()
- tvals = model.tvalues[relevant_terms].to_dict()
- # Combine and add measure
- combined_dict = {**{f'{term}_pval': pvals[term] for term in relevant_terms},
- **{f'{term}_tval': tvals[term] for term in relevant_terms},
- 'measure': measure}
- results.append(combined_dict)
- results_df = pd.DataFrame(results)
- # Multiple comparisons correction
- if multcomp:
- pval_cols = [f'{term}_pval' for term in relevant_terms]
- corrected_pvals = multipletests(results_df[pval_cols].values.flatten(),
- method=multcomp)[1]
- results_df[pval_cols] = corrected_pvals.reshape(results_df[pval_cols].shape)
- # Filter for significant effects of interest
- if filter_significant and vars_of_interest:
- filter_conditions = [(results_df[f'{var}_pval'] < 0.05) for var in vars_of_interest]
- combined_condition = filter_conditions[0]
- for condition in filter_conditions[1:]:
- combined_condition |= condition
- results_df = results_df[combined_condition]
- return results_df
- def set_column_width(column, width_cm):
- # Adjust column width in cm
- for cell in column.cells:
- tc = cell._element
- tcPr = tc.get_or_add_tcPr()
- tcW = OxmlElement('w:tcW')
- tcW.set(qn('w:w'), str(int(width_cm * 567))) # 1cm ≈ 567 units in docx format
- tcW.set(qn('w:type'), 'dxa')
- tcPr.append(tcW)
- def create_word_table_from_dataframe(df, dest_dir, filename):
- # Initialize the document
- doc = Document()
- # Get the covariates by splitting column names (assuming X_pval and X_tval structure)
- covariates = sorted(set(col.split('_')[0] for col in df.columns))
- # Create the table: rows = number of rows in df + 2 (for headers), columns = 1 (feature names) + 2 * number of covariates
- table = doc.add_table(rows=len(df) + 2, cols=1 + 2 * len(covariates))
- # First row: feature names and covariate names
- table.cell(0, 0).text = 'Feature'
- for i, covariate in enumerate(covariates):
- cell = table.cell(0, 1 + 2 * i)
- cell.text = covariate
- cell.merge(table.cell(0, 1 + 2 * i + 1)) # Merge the two columns for covariate name
- # Make covariate name bold and center it
- for paragraph in cell.paragraphs:
- paragraph.alignment = WD_PARAGRAPH_ALIGNMENT.CENTER
- for run in paragraph.runs:
- run.bold = True
- # Second row: 'p' and 't' labels for each covariate
- table.cell(1, 0).text = '' # Blank cell for the feature names column header
- for i, covariate in enumerate(covariates):
- table.cell(1, 1 + 2 * i).text = 'p'
- table.cell(1, 1 + 2 * i + 1).text = 't'
- # Remaining rows: fill in the feature names, p-values, and t-values from the DataFrame
- for row_idx, feature_name in enumerate(df.index):
- # Insert feature name in the first column
- table.cell(row_idx + 2, 0).text = str(feature_name)
- # Insert p-values and t-values in the remaining columns
- for i, covariate in enumerate(covariates):
- pval_col = f'{covariate}_pval'
- tval_col = f'{covariate}_tval'
- # Format p-values to 4 decimal places and t-values to 2 decimal places
- pval = f'{df.iloc[row_idx][pval_col]:.4f}'
- tval = f'{df.iloc[row_idx][tval_col]:.2f}'
- # Insert formatted values into the table
- table.cell(row_idx + 2, 1 + 2 * i).text = pval
- table.cell(row_idx + 2, 1 + 2 * i + 1).text = tval
- # Adjust font size for the entire table (optional, for better readability)
- for row in table.rows:
- for cell in row.cells:
- for paragraph in cell.paragraphs:
- for run in paragraph.runs:
- run.font.size = Pt(10)
- # Set the column width for a more compact layout (width in cm)
- set_column_width(table.columns[0], 2.5) # Feature names column width
- for i in range(1, len(covariates) * 2 + 1):
- set_column_width(table.columns[i], 1.5) # Covariate columns width (for both p and t)
- # Save the document
- output_path = os.path.join(dest_dir, f'{filename}.docx')
- doc.save(output_path)
- return output_path
- def create_combined_violin_plots(df, group_col, feature_cols, plotdir):
- sns.set(style="whitegrid", font_scale=1) # Set style and adjust font size
- num_features = len(feature_cols)
- # Determine the number of rows needed for the subplot grid
- num_rows = (num_features + 1) // 2
- plt.figure(figsize=(13, 4 * num_rows)) # Adjust the figure size as needed
- # Melt the DataFrame to have feature names as a categorical variable
- df_melted = df.melt(id_vars=group_col, value_vars=feature_cols, var_name='Feature', value_name='z-score')
- sns.violinplot(x='z-score', y='Feature', hue=group_col, data=df_melted, split=True, inner='quartile', orient="h")
- plt.title('z-scores distributions')
- plt.xlabel('z-score')
- # Add total counts for each site
- pvals=[]
- for s in df_melted['Feature'].unique():
- group1 = df_melted[(df_melted['Feature'] == s) & (df_melted[group_col] == 0)]['z-score']
- group2 = df_melted[(df_melted['Feature'] == s) & (df_melted[group_col] == 1)]['z-score']
- _, p_value = stats.ttest_ind(group1, group2, alternative='two-sided')
- #_, p_value = stats.mannwhitneyu(group1, group2, alternative='two-sided')
- pvals.append(p_value)
- corrected_pvals = multipletests(pvals, method='fdr_bh')[1]
- i=0
- for s in df_melted['Feature'].unique():
- plt.text(df_melted['z-score'].max() + 5, i, f'{s}: p = {corrected_pvals[i]:.4f}', va='center')
- i=i+1
- plt.tight_layout()
- plt.savefig(f'{plotdir}/{group_col}_combined_violin_plots.tif', format='tif', dpi=300)
- plt.close()
- # Example usage
- # df = pd.read_csv('your_data.csv') # Load your dataframe
- # group_col = 'dx' # Column name for the group (dx in your case)
- # feature_cols = df.columns.drop(group_col) # List of feature columns
- # plotdir = 'path_to_save_plots' # Directory where you want to save the plots
- # create_combined_violin_plots(df, group_col, feature_cols, plotdir)
- def calculate_group_stats(df, group_var, measure_dict, labels_dict, thr = 1.96):
- # Create a copy of the DataFrame to avoid SettingWithCopyWarning
- df = df.copy()
- # Calculate Average Deviation Scores for each measure group
- for group_name, measures in measure_dict.items():
- df[f'avg_dev_{group_name}'] = df[measures].mean(axis=1)
- # Calculate Global Average Deviation Score
- group_avg_columns = [f'avg_dev_{group_name}' for group_name in measure_dict.keys()]
- df['avg_dev_global'] = df[group_avg_columns].mean(axis=1)
- # Initialize the result DataFrame
- result_df = pd.DataFrame()
- # Iterate over measure groups and compute statistics
- for group_name in list(measure_dict.keys()) + ['global']:
- # Count occurrences for each group
- counts_positive = df.groupby(group_var)[f'avg_dev_{group_name}'].apply(lambda x: (x > thr).sum())
- counts_negative = df.groupby(group_var)[f'avg_dev_{group_name}'].apply(lambda x: (x < -thr).sum())
- # Perform Proportions Tests
- groups = df[group_var].unique()
- if len(groups) != 2:
- raise ValueError("The function currently supports exactly two groups for comparison.")
- count_pos = [counts_positive[group] for group in groups]
- count_neg = [counts_negative[group] for group in groups]
- nobs = [df[df[group_var] == group].shape[0] for group in groups]
- stat_pos, pval_pos = proportions_ztest(count_pos, nobs)
- stat_neg, pval_neg = proportions_ztest(count_neg, nobs)
- # Append the statistics to the result DataFrame
- for g, gl in labels_dict.items():
- result_df.loc[group_name, f'{g}_count_>{thr}'] = counts_positive[gl]
- result_df.loc[group_name, f'{g}_count_<-{thr}'] = counts_negative[gl]
- result_df.loc[group_name, f'p_value_>{thr}'] = pval_pos
- result_df.loc[group_name, f'p_value_<-{thr}'] = pval_neg
- return result_df
- # Example usage
- # df = [your_dataframe_here]
- # group_var = 'subtype'
- # measure_dict = {
- # 'Thicknesses': ['thickness_measure1', 'thickness_measure2'],
- # 'Volumes': ['volume_measure1', 'volume_measure2'],
- # 'Surfaces': ['surface_measure1', 'surface_measure2']
- # }
- # labels_dict = {'HC': 'Healthy Control', 'AN': 'Affected Group'}
- # result_df = calculate_group_stats(df, group_var, measure_dict, labels_dict)
- import numpy as np
- from scipy.stats import binom
- def print_group_percentages(dataframe, columns, group_column, min_infrasupra=1, thr = 1.96):
- """
- Prints the percentage of participants in each group with at least a specified number
- of supra or infranormal deviations.
- Parameters:
- dataframe (pd.DataFrame): DataFrame containing the data.
- columns (list): List of column names for the z-scores of ROIs.
- group_column (str): Column name for group labels.
- min_infrasupra (int or float, optional): Minimum number of deviations required.
- If between 0 and 0.1, interprets it as a p-value and computes the threshold
- using the binomial distribution.
- Returns:
- None
- """
- nc = len(columns) # Number of columns (regions of interest)
- pthr = {1.96:0.05, 1.28:0.1, 0.84:0.2}
- # Handle case where min_infrasupra is a p-value
- if 0 < min_infrasupra < 0.1:
- p = pthr[thr] # Probability of |z| > 1.96
- mu = nc * p
- sigma = np.sqrt(nc * p * (1 - p))
- # Compute the threshold corresponding to the (min_infrasupra / 2)% tail
- k = binom.ppf(1 - min_infrasupra / 2, n=nc, p=p)
- threshold = int(k)
- threshold_text = f"at least {threshold}"
- else:
- threshold = min_infrasupra
- threshold_text = f"at least {threshold}"
- for group in dataframe[group_column].unique():
- group_data = dataframe[dataframe[group_column] == group][columns]
- # Count supra and infra deviations for each participant
- supra_deviations = (group_data > thr).sum(axis=1)
- infra_deviations = (group_data < -thr).sum(axis=1)
- total_deviations = supra_deviations + infra_deviations
- total_percentage = (total_deviations >= threshold).mean() * 100
- supra_percentage = (supra_deviations >= threshold).mean() * 100
- infra_percentage = (infra_deviations >= threshold).mean() * 100
- total_count = (total_deviations >= threshold).sum()
- supra_count = (supra_deviations >= threshold).sum()
- infra_count = (infra_deviations >= threshold).sum()
- n_participants = len(group_data)
- print(f"- In group '{group}':")
- print(f" * {total_percentage:.2f}% ({total_count} out of {n_participants}) participants had {threshold_text} supranormal or infranormal deviation.")
- print(f" * {supra_percentage:.2f}% ({supra_count} out of {n_participants}) participants had {threshold_text} supranormal deviation.")
- print(f" * {infra_percentage:.2f}% ({infra_count} out of {n_participants}) participants had {threshold_text} infranormal deviation.")
- def summarize_deviations(df, thicknesses, surfaces, volumes, group_column='dx', group_labels={0:'HC', 1:'AN'}, thr = 1.96 ):
- """
- Summarizes the percentage of participants with supra-/infra-normal deviations in CT, SA, and SV measures for AN and HC groups.
- Args:
- df (pd.DataFrame): The dataframe containing participant data.
- thicknesses (list): List of column names for cortical thickness measures.
- surfaces (list): List of column names for surface area measures.
- volumes (list): List of column names for subcortical volume measures.
- Returns:
- str: A single summary sentence.
- """
- # Define thresholds
- supranormal_threshold = thr
- infranormal_threshold = -thr
- # Define groups
- an_group = df[df[group_column] == 1]
- hc_group = df[df[group_column] == 0]
- # Helper function to calculate percentages
- def calculate_percentages(group, measures):
- supra_mask = (group[measures] > supranormal_threshold).any(axis=1)
- infra_mask = (group[measures] < infranormal_threshold).any(axis=1)
- supra_percentage = supra_mask.mean() * 100
- infra_percentage = infra_mask.mean() * 100
- return supra_percentage, infra_percentage
- # Calculate for each measure type and group
- an_ct_supra, an_ct_infra = calculate_percentages(an_group, thicknesses)
- an_sa_supra, an_sa_infra = calculate_percentages(an_group, surfaces)
- an_sv_supra, an_sv_infra = calculate_percentages(an_group, volumes)
- hc_ct_supra, hc_ct_infra = calculate_percentages(hc_group, thicknesses)
- hc_sa_supra, hc_sa_infra = calculate_percentages(hc_group, surfaces)
- hc_sv_supra, hc_sv_infra = calculate_percentages(hc_group, volumes)
- # Create the summary sentence
- summary = (
- f"In the whole {group_labels[1]} group, supra-/infra-normal z-scores were observed in {an_ct_supra:.2f}%/{an_ct_infra:.2f}% of regional CT measures, "
- f"{an_sa_supra:.2f}%/{an_sa_infra:.2f}% of regional SA measures, and {an_sv_supra:.2f}%/{an_sv_infra:.2f}% of regional SV measures. "
- f"The percentages in the {group_labels[0]} group were {hc_ct_supra:.2f}%/{hc_ct_infra:.2f}%, {hc_sa_supra:.2f}%/{hc_sa_infra:.2f}%, "
- f"and {hc_sv_supra:.2f}%/{hc_sv_infra:.2f}%, respectively."
- )
- return summary
- def summarize_deviation_percentages(df, thicknesses, surfaces, volumes, group_column='dx', group_labels={0: 'HC', 1: 'AN'}, thr=1.96):
- """
- Summarizes the percentage of supra-normal and infra-normal values across all measures (thickness, surface, volume) for specified groups.
- Args:
- df (pd.DataFrame): The dataframe containing participant data.
- thicknesses (list): List of column names for cortical thickness measures.
- surfaces (list): List of column names for surface area measures.
- volumes (list): List of column names for subcortical volume measures.
- group_column (str): Name of the column indicating group membership.
- group_labels (dict): Dictionary mapping group values to labels.
- Returns:
- str: A summary of the percentages of supra-normal and infra-normal values for each group.
- """
- # Define thresholds
- supranormal_threshold = thr
- infranormal_threshold = -thr
- # Initialize results dictionary
- results = {}
- for group_value, group_label in group_labels.items():
- # Filter data for the current group
- group_data = df[df[group_column] == group_value]
- # Combine all measures
- measures = group_data[thicknesses + surfaces + volumes].values.flatten()
- measures = measures[~np.isnan(measures)] # Remove NaN values
- # Calculate percentages
- supra_percentage = np.mean(measures > supranormal_threshold) * 100
- infra_percentage = np.mean(measures < infranormal_threshold) * 100
- # Store results
- results[group_label] = {
- 'supra': supra_percentage,
- 'infra': infra_percentage
- }
- # Build summary text
- summary_lines = []
- for group_label, percentages in results.items():
- summary_lines.append(
- f"In the {group_label} group, {percentages['supra']:.2f}% of values are supra-normal (>{thr}) "
- f"and {percentages['infra']:.2f}% of values are infra-normal (<-{thr})."
- )
- return " \n".join(summary_lines)
- # Example usage:
- # summary = summarize_deviation_percentages(df, thicknesses, surfaces, volumes)
- # print(summary)
- def identify_columns_with_extreme_values(dataframe, threshold=1.96):
- """
- Identify columns with at least one participant with z > threshold or z < -threshold.
- Parameters:
- dataframe (pd.DataFrame): DataFrame containing the z-scores.
- threshold (float): Threshold for identifying extreme values (default is 1.96).
- Returns:
- dict: Dictionary with lists of columns for z > threshold and z < -threshold.
- """
- columns_z_gt_threshold = dataframe.columns[(dataframe > threshold).any()]
- columns_z_lt_threshold = dataframe.columns[(dataframe < -threshold).any()]
- return {
- "z_greater_than_threshold": list(columns_z_gt_threshold),
- "z_less_than_threshold": list(columns_z_lt_threshold)
- }
- #
- # Example usage
- #colswextremes = identify_columns_with_extreme_values(df_T1.loc[df_T1.dx==0.0,thicknesses], threshold=1.96)
- def save_custom_colorbar(cmap_name,
- vmin,
- vmax,
- ticks,
- out_path,
- orientation='vertical',
- label=None,
- dpi=300,
- background='black',
- tick_color='white',
- font_size=10,
- bar_width=0.3,
- bar_height=6):
- """
- Saves a styled standalone colorbar with control over size, orientation, and appearance.
- Parameters:
- cmap_name (str): Colormap name (registered in matplotlib).
- vmin, vmax (float): Value range for colorbar.
- ticks (list): List of tick values.
- out_path (str): Output file path (.tif, .png, etc.).
- orientation (str): 'vertical' or 'horizontal'.
- label (str): Optional axis label.
- dpi (int): Output resolution.
- background (str): Background color (e.g. 'black').
- tick_color (str): Tick label color.
- font_size (int): Font size of ticks.
- bar_width (float): Width in inches (thickness of the bar).
- bar_height (float): Height in inches (for vertical) or width (for horizontal).
- """
- if orientation == 'vertical':
- fig, ax = plt.subplots(figsize=(bar_width, bar_height))
- fig.subplots_adjust(left=0.3, right=0.7)
- else:
- fig, ax = plt.subplots(figsize=(bar_height, bar_width)) # horizontal: width=bar_height
- fig.subplots_adjust(bottom=0.3, top=0.7)
- # Background color
- fig.patch.set_facecolor(background)
- ax.set_facecolor(background)
- # Create colorbar
- norm = colors.Normalize(vmin=vmin, vmax=vmax)
- cb = colorbar.ColorbarBase(ax,
- cmap=plt.get_cmap(cmap_name),
- norm=norm,
- orientation=orientation,
- ticks=ticks)
- cb.ax.tick_params(labelsize=font_size, colors=tick_color)
- if label:
- cb.set_label(label, color=tick_color)
- # Save figure
- plt.savefig(out_path, bbox_inches='tight', dpi=dpi, facecolor=background)
- plt.close()
normative_tools.py, no license · at the source
Overview
and 16 other authors
Luca Lavagnino21, Christina E Wierenga12,13, Amanda Bischoff-Grethe12,13, Amy E Miles22, Allan Kaplan22, Aristotle Voineskos22, Paul A M Smeets23,24, Annemarie A van Elburg25,26, Unna Danner25,26, Sophia I Thomopoulos27, Laura Berner28, Neda Jahanshad27, Sophia Frangou3,28, Joseph A King1, Paul Thompson27, Stefan Ehrlich1,2929 affiliations
- Translational Developmental Neuroscience Section, Division of Psychological and Social Medicine and Developmental Neurosciences, Faculty of Medicine, Technische Universität Dresden, Dresden, Germany
- Maurice Wohl Clinical Neuroscience Institute, Department of Psychological Medicine, Institute of Psychiatry, Psychology and Neuroscience, King’s College London, London, United Kingdom
- Djavad Mowafaghian Centre for Brain Health, University of British Columbia, Vancouver, British Columbia, Canada
- Centre de recherche CHU Sainte Justine, Department of Psychiatry and Addictology, University of Montreal, Montreal, Québec, Canada
- Department of Child Health and Development, Norwegian Institute of Public Health, Oslo, Norway
- Department of Neurosciences ‘Rita Levi Montalcini’, University of Turin, Turin, Italy
- Eating Disorders Center for Treatment and Research, University of Turin, Turin, Italy
- PROMENTA Research Center, Department of Psychology, University of Oslo, Oslo, Norway
- Division of Mental Health and Substance Abuse, Diakonhjemmet Hospital, Oslo, Norway
- Centre for Research in Eating and Weight Disorders, Institute of Psychitry, Psychology and Neuroscience, King’s College London, London, United Kingdom
- Department of Neuroimaging, Institute of Psychiatry, Psychology and Neuroscience, King’s College London, London, United Kingdom
- Department of Psychiatry, University of California San Diego, La Jolla, California, United States of America
- Eating Disorders Center for Treatment and Research, University of California San Diego, La Jolla, California, United States of America
- Department of Child and Adolescent Psychiatry, Bielefeld University, Medical School and University Medical Center OWL, Protestant Hospital of the Bethel Foundation, Bielefeld, Germany
- Department of Child and Adolescent Psychiatry, University Clinic Erlangen, Erlangen, Germany
- Institute of Experimental and Clinical Pharmacology and Toxicology, Emil Fischer Center, University of Erlangen-Nuremberg, Erlangen, Germany
- Department of Neuroradiology, University of Erlangen-Nuremberg, Erlangen, Germany
- FAU NeW - Research Center for New Bioactive Compounds, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany
- Centre for Psychosocial Medicine, Department of General Internal Medicine and Psychosomatics, University Hospital Heidelberg, Heidelberg, Germany
- Padova Neuroscience Center, Department of Neurosciences, University of Padova, Padova, Italy
- Department of Psychiatry and Behavioral Sciences, University of Texas Health Science Center, Houston, Texas, United States of America
- Campbell Family Mental Health Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada
- UMC Utrecht Brain Center, Utrecht University, Utrecht, the Netherlands
- Division of Human Nutrition and Health, Wageningen University, Wageningen, the Netherlands
- Altrecht Eating Disorders Rintveld, Altrecht Mental Health Institute, Zeist, the Netherlands
- Faculty of Social Sciences, Utrecht University, Utrecht, the Netherlands
- Imaging Genetics Center, Stevens Institute for Neuroimaging and Informatics, Keck USC School of Medicine, Marina del Rey, California, United States of America
- Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, New York, United States of America
- Eating Disorders Research and Treatment Center, Department of Child and Adolescent Psychiatry, Faculty of Medicine, Dresden University of Technology, Dresden, Germany
Abstract
Background: In a recent coordinated meta-analysis of neuroimaging data, we reported gray matter (GM) alterations in acutely underweight patients with anorexia nervosa (AN). Here, we extend these findings by examining individual variation in brain structure within AN, individual-level differentiation between AN and healthy controls (HC), and differences between AN subtypes, with potential relevance for understanding clinical heterogeneity.
Methods and findings: We analyzed individual-level data from 11 international sites in the ENIGMA Eating Disorders Working Group, including 570 female participants with AN and 739 HC. We examined cortical thickness, cortical surface area and subcortical volumes in AN versus HC using three complementary approaches: (i) group-level differences in a mega-analysis correcting for age effects, (ii) frequencies of extreme deviations (infra-/
Mega-analyses reinforced previous meta-analytic findings of pronounced and widespread GM deficits in AN compared to HC. Normative modelling revealed that the frequency of infranormal z-scores (23/
Conclusion: Using a mega-analytic approach, we confirm widespread GM deficits in AN, show that these alterations are (in some patients) extreme, and demonstrate that they enable robust classification with superior performance compared to most MRI-based psychiatric classification studies. The absence of differences between AN subtypes may reflect shared neurobiology, though other imaging modalities may reveal distinctions beyond brain structure.
Reproduced under the paper's license (CC0), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.
OSF xrjkf
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
- 28 September 2026: the link answers (HTTP 200)
45 files
- src/
an_heterogeneity/ , Jupyter, 174 linesANHC_VisualizeNormativeM odeling.ipynb - src/
an_heterogeneity/ , Jupyter, 558 linesANHC_classification_Comb atGAM.ipynb - src/
an_heterogeneity/ , Jupyter, 566 linesANHC_classification_LSsO .ipynb - src/
an_heterogeneity/ , Jupyter, 120 linesANHC_demographics.ipynb - src/
an_heterogeneity/ , Jupyter, 594 lines, 1 matchANHC_univariate.ipynb - src/
an_heterogeneity/ , Jupyter, 371 linesVisualizeNormativeModeli ng.ipynb - src/
an_heterogeneity/ , Python, 20 linesanalysis_utilities/ customClassifiers.py - src/
an_heterogeneity/ , Python, 66 linesanalysis_utilities/ preparation.py - src/
an_heterogeneity/ , Python, 22 linesconfig/ global_variables.py - src/
an_heterogeneity/ , Jupyter, 759 linesdata_join.ipynb - src/
an_heterogeneity/ , Python, 45 linesdata_utilities/ data_IO.py - src/
an_heterogeneity/ , Python, 136 linesdata_utilities/ data_formatting.py - src/
an_heterogeneity/ , Python, 130 linesdata_utilities/ data_join_utils.py - src/
an_heterogeneity/ , Python, 3 linesmodel_evaluation/ __init__.py - src/
an_heterogeneity/ , Python, 172 lines, 1 matchmodel_evaluation/ cross_validation.py - src/
an_heterogeneity/ , Python, 294 linesmodel_evaluation/ modeling_functions.py - src/
an_heterogeneity/ , Python, 475 lines, 1 matchmodel_evaluation/ performance_metric_visua lization.py - src/
an_heterogeneity/ , Python, 201 linesmodel_evaluation/ permutation_test.py - src/
an_heterogeneity/ , Python, 88 linesmodel_evaluation/ unused_functions.py - src/
an_heterogeneity/ , Jupyter, 174 linessubtype_VisualizeNormati veModeling.ipynb - src/
an_heterogeneity/ , Jupyter, 668 linessubtype_classification_C ombatGAM.ipynb - src/
an_heterogeneity/ , Jupyter, 671 linessubtype_classification_L SsO.ipynb - src/
an_heterogeneity/ , Jupyter, 94 linessubtype_demographics.ipy nb - src/
an_heterogeneity/ , Jupyter, 252 linessubtype_univariate.ipynb - src/
an_heterogeneity/ , Python, 111 linestaurus_scripts/ hyperparam_tuning.py - src/
an_heterogeneity/ , Python, 593 linestaurus_scripts/ ncv_permutation_test_bar nard.py - src/
an_heterogeneity/ , Python, 149 linestaurus_scripts/ run_ncv_taurus_LSsO.py - src/
an_heterogeneity/ , Python, 178 linestaurus_scripts/ run_ncv_taurus_centBrain .py - src/
an_heterogeneity/ , Python, 146 linestaurus_scripts/ run_ncv_taurus_combat.py - src/
an_heterogeneity/ , Python, 221 linestaurus_scripts/ run_perms_taurus_LSsO.py - src/
an_heterogeneity/ , Python, 231 linestaurus_scripts/ run_perms_taurus_centBra in.py - src/
an_heterogeneity/ , Python, 203 linestaurus_scripts/ run_perms_taurus_combat. py - src/
an_heterogeneity/ , Python, 1 linetools/ __init__.py - src/
an_heterogeneity/ , Python, 315 linestools/ data_handling.py - src/
an_heterogeneity/ , Python, 141 linestools/ dinga.py - src/
an_heterogeneity/ , Python, 171 linestools/ evaluate_ncv_results.py - src/
an_heterogeneity/ , Python, 277 lines, 1 matchtools/ load_parse_neuromaps.py - src/
an_heterogeneity/ , Python, 328 lines, 1 matchtools/ model_pipelines.py - src/
an_heterogeneity/ , Python, 49 linestools/ model_pipelines_306090.p y - src/
an_heterogeneity/ , Python, 505 lines, 1 matchtools/ model_training.py - src/
an_heterogeneity/ , Python, 1,056 lines, 2 matchestools/ normative_tools.py - src/
an_heterogeneity/ , Python, 683 linestools/ normative_tools_paper.py - src/
an_heterogeneity/ , Python, 202 linestools/ permutation_test.py - src/
an_heterogeneity/ , Python, 576 linestools/ visualization.py - src/
an_heterogeneity/ , Jupyter, 814 linesz_scores_analysis.ipynb
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 45 scripts, each with its path and the digest of its content;
- 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability
Individual-level data underlying the findings of this study cannot be shared publicly because they are governed by site-specific ethical approvals and national and institutional data protection regulations at the 11 contributing ENIGMA Eating Disorders Working Group sites. The study authors are not the legal custodians of these data. Data access inquiries must be directed to the relevant institutional data custodian or ethics/
Reproduced under the paper's license (CC0), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 28 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 36 authors, 11 MeSH terms, 13 funders, 88 references.
Cite
This paper
Bernardoni, F., Arold, D., Schoppik, L., Bahnsen, K., Ge, R., Moreau, C., Bang, L., D’Agata, F., Abbate-Daga, G., Tamnes, C. K., Campbell, I., O’Daly, O., Schmidt, U., Frank, G., Horndasch, S., Hess, A., Dörfler, A., Friederich, H.-C., Simon, J., . . . Ehrlich, S. (2026). Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data. PLoS medicine, 23(5), e1004809. https://
BibTeX
@article{bernardoni2026b
author = {Bernardoni, Fabio and Arold, Dominic and Schoppik, Luis and Bahnsen, Klaas and Ge, Ruiyang and Moreau, Clara and Bang, Lasse and D’Agata, Federico and Abbate-Daga, Giovanni and Tamnes, Christian K and Campbell, Iain and O’Daly, Owen and Schmidt, Ulrike and Frank, Guido and Horndasch, Stefanie and Hess, Andreas and Dörfler, Arnd and Friederich, Hans-Christoph and Simon, Joe and Favaro, Angela and Lavagnino, Luca and Wierenga, Christina E and Bischoff-Grethe, Amanda and Miles, Amy E and Kaplan, Allan and Voineskos, Aristotle and Smeets, Paul A M and van Elburg, Annemarie A and Danner, Unna and Thomopoulos, Sophia I and Berner, Laura and Jahanshad, Neda and Frangou, Sophia and King, Joseph A and Thompson, Paul and Ehrlich, Stefan},
title = {{Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data}},
journal = {PLoS medicine},
year = {2026},
month = may,
volume = {23},
number = {5},
pages = {e1004809},
publisher = {PLOS},
issn = {1549-1277},
doi = {10.1371/
url = {https://
pmid = {42160333},
pmcid = {PMC13215615}
}
RIS
TY - JOUR
AU - Bernardoni, Fabio
AU - Arold, Dominic
AU - Schoppik, Luis
AU - Bahnsen, Klaas
AU - Ge, Ruiyang
AU - Moreau, Clara
AU - Bang, Lasse
AU - D’Agata, Federico
AU - Abbate-Daga, Giovanni
AU - Tamnes, Christian K
AU - Campbell, Iain
AU - O’Daly, Owen
AU - Schmidt, Ulrike
AU - Frank, Guido
AU - Horndasch, Stefanie
AU - Hess, Andreas
AU - Dörfler, Arnd
AU - Friederich, Hans-Christoph
AU - Simon, Joe
AU - Favaro, Angela
AU - Lavagnino, Luca
AU - Wierenga, Christina E
AU - Bischoff-Grethe, Amanda
AU - Miles, Amy E
AU - Kaplan, Allan
AU - Voineskos, Aristotle
AU - Smeets, Paul A M
AU - van Elburg, Annemarie A
AU - Danner, Unna
AU - Thomopoulos, Sophia I
AU - Berner, Laura
AU - Jahanshad, Neda
AU - Frangou, Sophia
AU - King, Joseph A
AU - Thompson, Paul
AU - Ehrlich, Stefan
TI - Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data
T2 - PLoS medicine
J2 - PLoS Med
PY - 2026
DA - 2026/
VL - 23
IS - 5
SP - e1004809
SN - 1549-1277
PB - PLOS
DO - 10.1371/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1371/
"type": "article-journal",
"title": "Brain morphology in Anorexia Nervosa and its subtypes: A multi-cohort study of individual participant data",
"container-title": "PLoS medicine",
"author": [
{
"family": "Bernardoni",
"given": "Fabio"
},
{
"family": "Arold",
"given": "Dominic"
},
{
"family": "Schoppik",
"given": "Luis"
},
{
"family": "Bahnsen",
"given": "Klaas"
},
{
"family": "Ge",
"given": "Ruiyang"
},
{
"family": "Moreau",
"given": "Clara"
},
{
"family": "Bang",
"given": "Lasse"
},
{
"family": "D’Agata",
"given": "Federico"
},
{
"family": "Abbate-Daga",
"given": "Giovanni"
},
{
"family": "Tamnes",
"given": "Christian K"
},
{
"family": "Campbell",
"given": "Iain"
},
{
"family": "O’Daly",
"given": "Owen"
},
{
"family": "Schmidt",
"given": "Ulrike"
},
{
"family": "Frank",
"given": "Guido"
},
{
"family": "Horndasch",
"given": "Stefanie"
},
{
"family": "Hess",
"given": "Andreas"
},
{
"family": "Dörfler",
"given": "Arnd"
},
{
"family": "Friederich",
"given": "Hans-Christoph"
},
{
"family": "Simon",
"given": "Joe"
},
{
"family": "Favaro",
"given": "Angela"
},
{
"family": "Lavagnino",
"given": "Luca"
},
{
"family": "Wierenga",
"given": "Christina E"
},
{
"family": "Bischoff-Grethe",
"given": "Amanda"
},
{
"family": "Miles",
"given": "Amy E"
},
{
"family": "Kaplan",
"given": "Allan"
},
{
"family": "Voineskos",
"given": "Aristotle"
},
{
"family": "Smeets",
"given": "Paul A M"
},
{
"family": "van Elburg",
"given": "Annemarie A"
},
{
"family": "Danner",
"given": "Unna"
},
{
"family": "Thomopoulos",
"given": "Sophia I"
},
{
"family": "Berner",
"given": "Laura"
},
{
"family": "Jahanshad",
"given": "Neda"
},
{
"family": "Frangou",
"given": "Sophia"
},
{
"family": "King",
"given": "Joseph A"
},
{
"family": "Thompson",
"given": "Paul"
},
{
"family": "Ehrlich",
"given": "Stefan"
}
],
"container-title-short":
"volume": "23",
"issue": "5",
"page": "e1004809",
"DOI": "10.1371/
"PMID": "42160333",
"PMCID": "PMC13215615",
"ISSN": "1549-1277",
"publisher": "PLOS",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
20
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1073/pnas.2521055123 [code]
- Empirical validation of race-neutral normative brain morphometry models across ethnoracially diverse populations.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: neuroHarmonize, pandas, NumPy, structural MRI / diffusion, 7 references, author Ruiyang Ge
- [2] doi:10.1162/imag.a.1269 [code]
- From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.Journal: Imaging neuroscience (Cambridge, Mass.)In common: Plotly, NiBabel, statsmodels, 6 other tools, 6 references
- [3] doi:10.1002/aur.70243
- Autism and Cortical Thickness Deviation From Neurotypical Controls: Evidence for a Spatial Association With Serotonin Receptors.Journal: Autism research : official journal of the International Society for Autism ResearchIn common: structural MRI / diffusion, 3 authors
- [4] doi:10.3389/fnins.2026.1803154 [code]
- Multimodal machine learning reveals neurobiological signatures of binge-type eating disorders.Journal: Frontiers in neuroscienceIn common: statannotations, NiBabel, statsmodels, 6 other tools, other condition, 3 references
- [5] doi:10.1002/alz.71649 [code]
- Postmortem brain MRI reveals differential associations of subcortical and limbic volumes with cortical thinning and neurodegenerative pathologies.Journal: Alzheimer's & dementia : the journal of the Alzheimer's AssociationIn common: statannotations, Plotly, Pillow, 8 other tools, structural MRI / diffusion, other condition, 1 reference
- [6] doi:10.1038/s41380-026-03691-4 [code]
- Breaking the norm: population-scale deviations of brain structure in depression and anxiety.Journal: Molecular psychiatryIn common: NiBabel, statsmodels, seaborn, 5 other tools, structural MRI / diffusion, other condition, 4 references
- [7] doi:10.21203/rs.3.rs-9326213/v1 [code]
- Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brainJournal: Research Square (preprint)In common: neuromaps, statannotations, Plotly, 8 other tools
- [8] doi:10.64898/2026.03.09.710558 [code]
- Multi-task fMRI outperforms resting-state fMRI for revealing task-invariant organization of the human brainJournal: bioRxiv (preprint)In common: neuromaps, statannotations, Plotly, 8 other tools
- [9] doi:10.21203/rs.3.rs-9914920/v1 [code]
- Prediction of cognitive performance by demographics, sleep, and brain morphometry: machine learning findings from ENIGMA-Sleep Working GroupJournal: Research Square (preprint)In common: neuromaps, statannotations, NiBabel, 7 other tools, structural MRI / diffusion, 1 reference
- [10] doi:10.1038/s41380-026-03497-4 [code]
- Transcriptome-informed brain cartography of polygenic risk and association with brain structure in major psychiatric disorders.Journal: Molecular psychiatryIn common: neuromaps, NiBabel, statsmodels, 4 other tools, other condition, 4 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 45 scripts, and 8 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:a44976508b90d5c8…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
