APOE ε4 as a predictor of cognitive decline and its interaction with hippocampal volume in Alzheimer's disease.
The 20 matches
- [1] § Materials and methods › Variables ↔ notebooks2/01_data_pipeline.ipynb, lines 288–360 · score 0.87 · ST120SV, ST29SV, ST30SV, FreeSurfer, ICV adjusted, HIPPO_ICV_ADJ
- [2] § Materials and methods › Statistical analysis › Linear mixed-effects models ↔ notebooks2/03_lme_analysis.ipynb, lines 279–343 · score 0.78 · likelihood ratio, APOE4 interaction terms, way interaction term, reduced model, nested, LRT
- [3] § Materials and methods › Variables ↔ notebooks2/02_descriptive_stats.ipynb, lines 198–248 · score 0.73 · 0–18, 0–85, 0–30, ADAS Cog13, CDR SB, worse
- [4] § Materials and methods › Statistical analysis › Survival analysis ↔ notebooks2/04_survival_sensitivity.ipynb, lines 119–162 · score 0.72 · Kaplan Meier curves, log rank, APOE4_DOSE, lifelines, fitted, Survival
- [5] § Materials and methods › Statistical analysis › Sensitivity analyses ↔ notebooks2/04_survival_sensitivity.ipynb, lines 381–421 · score 0.68 · selection bias, multi visit, single visit, LME, sensitivity, MMSE
- [6] § Materials and methods › Statistical analysis › Model assumptions ↔ notebooks2/03_lme_analysis.ipynb, lines 199–268 · score 0.66 · Shapiro Wilk, LME models, Residuals, SD, MMSE
- [7] § Materials and methods › Statistical analysis › Linear mixed-effects models ↔ notebooks2/04_survival_sensitivity.ipynb, lines 212–297 · score 0.65 · random intercept, random slope, HIPPO_ICV_ADJ, formulated, baseline diagnosis, DX
- [8] § Materials and methods › Statistical analysis › Descriptive statistics ↔ notebooks2/02_descriptive_stats.ipynb, lines 45–153 · score 0.65 · chi squared, Kruskal Wallis, Baseline characteristics, SD, variables, dose
- [9] § Materials and methods › Statistical analysis › Linear mixed-effects models ↔ notebooks2/03_lme_analysis.ipynb, lines 77–152 · score 0.62 · random intercepts, random slopes, statsmodels, LME, fitted
- [10] § Materials and methods › Statistical analysis › Linear mixed-effects models ↔ notebooks2/04_survival_sensitivity.ipynb, lines 212–297 · score 0.62 · random intercepts, random slopes, statsmodels, fitted, LME, visit
- [11] § Results › APOE ε4 × hippocampal volume × time interaction ↔ notebooks2/03_lme_analysis.ipynb, lines 279–343 · score 0.60 · LME model, way interaction term, APOE4_DOSE, HIPPO_ICV_ADJ, MMSE
- [12] § Results › Survival analysis: conversion risk ↔ notebooks2/04_survival_sensitivity.ipynb, lines 119–162 · score 0.60 · Kaplan Meier curves, log rank, CN MCI, Survival, event, conversion
- [13] § Materials and methods › Workflow overview ↔ notebooks2/04_survival_sensitivity.ipynb, lines 1–59 · score 0.60 · Kaplan Meier curves, robust standard errors, MMSE ceiling, survival, Cox, outliers
- [14] § Materials and methods › Preprocessing and quality control ↔ notebooks2/01_data_pipeline.ipynb, lines 288–360 · score 0.58 · FreeSurfer, ICV adjusted, scans, pipeline, Raw, VISCODE2
- [15] § Results › Baseline characteristics ↔ notebooks2/02_descriptive_stats.ipynb, lines 45–153 · score 0.57 · Kruskal Wallis, ADAS Cog13, CDR SB, baseline characteristics, NPI, SD
- [16] § Materials and methods › Statistical analysis › Linear mixed-effects models ↔ notebooks2/03_lme_analysis.ipynb, lines 77–152 · score 0.55 · random intercept, random slope, formulated, HIPPO_ICV_ADJ, DX, model
- [17] § Results › Survival analysis: conversion risk ↔ notebooks2/04_survival_sensitivity.ipynb, lines 164–210 · score 0.53 · Cox proportional hazards, AD conversion, survival, fitted, predictor, events
- [18] § Results › Survival analysis: conversion risk ↔ notebooks2/04_survival_sensitivity.ipynb, lines 164–210 · score 0.52 · hazard ratios, Cox model, Survival, conversion, interaction, dose
- [19] § Results › Sensitivity analyses ↔ notebooks2/04_survival_sensitivity.ipynb, lines 483–514 · score 0.51 · Cluster robust, standard errors, OLS, fitted, LME, Sensitivity
- [20] § Results › APOE ε4 × hippocampal volume × time interaction ↔ notebooks2/02_descriptive_stats.ipynb, lines 385–416 · score 0.51 · MMSE score, ICV adjusted hippocampal, Regression, hippocampal volume, correlation, baseline
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 570 lines · 23 KB · MIT · 9 matches
- # %% [markdown]
- # # Survival Analysis & Sensitivity Analyses
- # **Study:** APOE ε4 × Hippocampal Volume × Cognitive Decline
- #
- # This notebook covers:
- # 1. **Survival/Conversion Analysis** — Cox PH model for CN → MCI/AD conversion
- # 2. **Kaplan-Meier Curves** by APOE dose
- # 3. **Sensitivity Analysis 1** — Stratified by baseline diagnosis (CN, MCI, AD)
- # 4. **Sensitivity Analysis 2** — Excluding outliers (±3 SD hippocampal volume)
- # 5. **Sensitivity Analysis 3** — Subjects with ≥2 visits only
- # 6. **Sensitivity Analysis 4** — Robust standard errors (MMSE ceiling effect)
- #
- # Requires: `reports/ADNI_Complete_Cases.csv` and `reports/ADNI_Baseline_Analysis.csv` from `01_data_pipeline.ipynb`
- # %%
- import pandas as pd
- import numpy as np
- import matplotlib.pyplot as plt
- import matplotlib.gridspec as gridspec
- import seaborn as sns
- import statsmodels.formula.api as smf
- import statsmodels.api as sm
- from scipy import stats
- import os
- import warnings
- warnings.filterwarnings('ignore')
- # Try to import lifelines for survival analysis
- try:
- from lifelines import CoxPHFitter, KaplanMeierFitter
- from lifelines.statistics import logrank_test, multivariate_logrank_test
- LIFELINES = True
- print('✓ lifelines available')
- except ImportError:
- LIFELINES = False
- print('⚠ lifelines not installed — install with: pip install lifelines')
- print(' Survival analysis will be skipped')
- BASE = '/media/faizaan/4TB/1_DATA_PROJECTS/Projects/Multimodel_study'
- REPORTS = os.path.join(BASE, 'reports')
- plt.rcParams.update({'figure.dpi': 120, 'font.size': 11,
- 'axes.spines.top': False, 'axes.spines.right': False})
- APOE_COLORS = {0: '#2196F3', 1: '#FF9800', 2: '#F44336'}
- # Load data
- complete = pd.read_csv(os.path.join(REPORTS, 'ADNI_Complete_Cases.csv'))
- baseline = pd.read_csv(os.path.join(REPORTS, 'ADNI_Baseline_Analysis.csv'))
- for df in [complete, baseline]:
- df['APOE4_DOSE'] = pd.to_numeric(df['APOE4_DOSE'], errors='coerce')
- df['HIPPO_ICV_ADJ'] = pd.to_numeric(df['HIPPO_ICV_ADJ'], errors='coerce')
- df['YEARS_FROM_BL'] = pd.to_numeric(df['YEARS_FROM_BL'], errors='coerce')
- df['MMSCORE'] = pd.to_numeric(df['MMSCORE'], errors='coerce')
- if 'SEX' in df.columns:
- df['SEX_MALE'] = (df['SEX'] == 'Male').astype(float)
- print(f'Complete: {complete["RID"].nunique()} subjects, {len(complete):,} obs')
- print(f'Baseline: {len(baseline)} subjects')
- # %% [markdown]
- # ## 1. Survival Analysis — CN → MCI/AD Conversion
- # Using pre-computed survival variables (`Event`, `Time`) from `baseline_subj.csv`.
- # %%
- # Prepare survival dataset
- # Keep CN subjects at baseline (those at risk of converting)
- surv_cols_needed = ['RID', 'APOE4_DOSE', 'HIPPO_ICV_ADJ', 'HIPPO_ICV_Z',
- 'AGE', 'SEX_MALE', 'PTEDUCAT', 'BL_DX_LABEL']
- event_cols = ['Event', 'Time', 'TimeMonths', 'TIME_YEARS']
- available_surv = [c for c in surv_cols_needed + event_cols if c in baseline.columns]
- surv_df = baseline[available_surv].copy()
- # Check survival columns
- has_event = 'Event' in surv_df.columns
- has_time = 'TIME_YEARS' in surv_df.columns
- print(f'Survival columns available: Event={has_event}, Time={has_time}')
- if has_event and has_time:
- # TIME_YEARS already converted in NB01 (was in months, now years)
- surv_df = surv_df.dropna(subset=['Event', 'TIME_YEARS', 'APOE4_DOSE', 'HIPPO_ICV_ADJ'])
- surv_df['TIME_YEARS'] = surv_df['TIME_YEARS'].clip(lower=0.01) # avoid zero-time
- print(f'\nSurvival dataset: {len(surv_df)} subjects')
- print(f'Events (conversions): {surv_df["Event"].sum():.0f} ({surv_df["Event"].mean()*100:.1f}%)')
- print(f'Follow-up (years): median={surv_df["TIME_YEARS"].median():.1f}, max={surv_df["TIME_YEARS"].max():.1f}')
- print('\nEvents by APOE dose:')
- print(surv_df.groupby('APOE4_DOSE')[['Event']].agg(['sum', 'count']))
- else:
- print('\n⚠ Survival variables (Event, Time) not found in baseline data.')
- print(' Will derive conversion status from longitudinal diagnosis data.')
- # Derive from longitudinal data
- if 'BL_DX_LABEL' in complete.columns and 'DX_LABEL' in complete.columns:
- cn_subjects = complete[complete['BL_DX_LABEL'] == 'CN']['RID'].unique()
- conversions = []
- for rid in cn_subjects:
- sub = complete[complete['RID'] == rid].sort_values('YEARS_FROM_BL')
- converted = (sub['DX_LABEL'].isin(['MCI', 'AD'])).any()
- if converted:
- conv_time = sub[sub['DX_LABEL'].isin(['MCI', 'AD'])]['YEARS_FROM_BL'].min()
- else:
- conv_time = sub['YEARS_FROM_BL'].max()
- conversions.append({'RID': rid, 'Event': int(converted), 'Time': conv_time})
- conv_df = pd.DataFrame(conversions)
- surv_df = baseline[baseline['RID'].isin(cn_subjects)].merge(conv_df, on='RID', how='inner')
- surv_df = surv_df.dropna(subset=['Event', 'Time', 'APOE4_DOSE', 'HIPPO_ICV_ADJ'])
- has_event = True
- has_time = True
- print(f'Derived conversion: {len(surv_df)} CN subjects, {surv_df["Event"].sum():.0f} conversions')
- else:
- print('Cannot derive conversion status — no DX_LABEL in longitudinal data')
- has_event = False
- # %% [markdown]
- # ### 1.1 Kaplan-Meier Curves by APOE Dose
- # %%
- if LIFELINES and has_event and 'surv_df' in locals() and len(surv_df) > 10:
- fig, ax = plt.subplots(figsize=(10, 6))
- kmfs = {}
- for dose in [0, 1, 2]:
- sub = surv_df[surv_df['APOE4_DOSE'] == dose]
- if len(sub) < 5: continue
- kmf = KaplanMeierFitter()
- kmf.fit(sub['TIME_YEARS'], sub['Event'], label=f'{dose} ε4 alleles (n={len(sub)})')
- kmf.plot_survival_function(ax=ax, color=APOE_COLORS[dose], ci_show=True, ci_alpha=0.15)
- kmfs[dose] = kmf
- ax.set_xlabel('Time (years)', fontsize=12)
- ax.set_ylabel('Conversion-free survival probability', fontsize=12)
- ax.set_title('Kaplan-Meier Survival Curves\nCN → MCI/AD Conversion by APOE ε4 Dose', fontweight='bold')
- ax.set_ylim(0, 1.05)
- ax.legend(title='APOE ε4 dose')
- # Log-rank test
- if len(kmfs) >= 2:
- try:
- from lifelines.statistics import multivariate_logrank_test
- results_lr = multivariate_logrank_test(surv_df['TIME_YEARS'], surv_df['APOE4_DOSE'], surv_df['Event'])
- p_lr = results_lr.p_value
- ax.text(0.98, 0.97, f'Log-rank p={"<0.001" if p_lr<0.001 else f"{p_lr:.3f}"}',
- transform=ax.transAxes, ha='right', va='top', fontsize=10,
- bbox=dict(boxstyle='round', facecolor='white', alpha=0.8))
- print(f'Log-rank test: p={p_lr:.4f}')
- except Exception as e:
- print(f'Log-rank test failed: {e}')
- plt.tight_layout()
- plt.savefig(os.path.join(REPORTS, 'Fig7_KaplanMeier.png'), dpi=150, bbox_inches='tight')
- plt.show()
- print('✓ Saved: Fig7_KaplanMeier.png')
- elif not LIFELINES:
- print('lifelines not available — install with: pip install lifelines')
- else:
- print('Insufficient survival data for KM curves')
- # %% [markdown]
- # ### 1.2 Cox Proportional Hazards Model
- # %%
- if LIFELINES and has_event and 'surv_df' in locals() and len(surv_df) > 10:
- # Build Cox model predictors
- cox_vars = ['APOE4_DOSE', 'HIPPO_ICV_ADJ']
- for v in ['AGE', 'SEX_MALE', 'PTEDUCAT', 'GDTOTAL']:
- if v in surv_df.columns:
- cox_vars.append(v)
- # Add APOE×Hippo interaction
- surv_df['APOE4_x_HIPPO'] = surv_df['APOE4_DOSE'] * surv_df['HIPPO_ICV_ADJ']
- cox_vars.append('APOE4_x_HIPPO')
- cox_cols = ['TIME_YEARS', 'Event'] + cox_vars
- cox_data = surv_df[[c for c in cox_cols if c in surv_df.columns]].dropna().copy()
- print(f'Cox model data: {len(cox_data)} subjects, {cox_data["Event"].sum():.0f} events')
- print(f'Predictors: {[c for c in cox_vars if c in cox_data.columns]}')
- cph = CoxPHFitter(penalizer=0.1) # small regularization for stability
- try:
- cph.fit(cox_data, duration_col='TIME_YEARS', event_col='Event')
- cph.print_summary()
- # Save Cox results
- cox_summary = cph.summary
- cox_summary.to_csv(os.path.join(REPORTS, 'Cox_PH_Results.csv'))
- print('\n✓ Saved: Cox_PH_Results.csv')
- # Forest plot of hazard ratios
- fig, ax = plt.subplots(figsize=(9, 5))
- cph.plot(ax=ax)
- ax.set_title('Cox PH Model — Hazard Ratios (95% CI)\nOutcome: CN → MCI/AD Conversion', fontweight='bold')
- plt.tight_layout()
- plt.savefig(os.path.join(REPORTS, 'Fig8_Cox_HazardRatios.png'), dpi=150, bbox_inches='tight')
- plt.show()
- print('✓ Saved: Fig8_Cox_HazardRatios.png')
- except Exception as e:
- print(f'Cox model failed: {e}')
- elif not LIFELINES:
- print('lifelines not available — skipping Cox model')
- else:
- print('Insufficient data for Cox model')
- # %% [markdown]
- # ## 2. Sensitivity Analysis 1 — Stratified by Baseline Diagnosis
- # Run LME models separately for CN, MCI, and AD subgroups.
- # %%
- def build_covariates(df):
- cov_parts = []
- if 'AGE' in df.columns: cov_parts.append('AGE')
- if 'SEX_MALE' in df.columns: cov_parts.append('SEX_MALE')
- if 'PTEDUCAT' in df.columns: cov_parts.append('PTEDUCAT')
- if 'GDTOTAL' in df.columns: cov_parts.append('GDTOTAL')
- return ' + '.join(cov_parts)
- def run_lme_simple(data, outcome, cov_str, label):
- """Fit LME with random intercept only (robust to small samples)"""
- if outcome not in data.columns:
- return None
- data_clean = data.dropna(subset=[outcome, 'YEARS_FROM_BL', 'APOE4_DOSE', 'HIPPO_ICV_ADJ']).copy()
- if 'GDTOTAL' in data_clean.columns:
- data_clean['GDTOTAL'] = data_clean['GDTOTAL'].fillna(0)
- # Need ≥2 visits per subject
- vc = data_clean.groupby('RID').size()
- data_clean = data_clean[data_clean['RID'].isin(vc[vc>=2].index)]
- n_sub = data_clean['RID'].nunique()
- if n_sub < 10:
- print(f' {label}: only {n_sub} subjects — too few for LME, skipping')
- return None
- formula = f"{outcome} ~ YEARS_FROM_BL * APOE4_DOSE * HIPPO_ICV_ADJ" + \
- (f" + {cov_str}" if cov_str else "")
- # Align data with Patsy's NaN-dropping to prevent statsmodels index mismatch
- import patsy as _patsy
- try:
- _, _X = _patsy.dmatrices(formula, data=data_clean, return_type='dataframe')
- data_clean = data_clean.loc[_X.index].reset_index(drop=True)
- n_sub = data_clean['RID'].nunique()
- except Exception:
- data_clean = data_clean.reset_index(drop=True)
- try:
- # Try random slope first
- m = smf.mixedlm(formula, data=data_clean, groups=data_clean['RID'],
- re_formula='~YEARS_FROM_BL')
- r = m.fit(method='lbfgs', maxiter=300, disp=False)
- if not np.isnan(r.llf):
- print(f' {label}: n={n_sub}, random slope converged (AIC={r.aic:.1f})')
- return r
- except:
- pass
- try:
- m = smf.mixedlm(formula, data=data_clean, groups=data_clean['RID'])
- r = m.fit(method='lbfgs', maxiter=300, disp=False)
- print(f' {label}: n={n_sub}, random intercept only (AIC={r.aic:.1f})')
- return r
- except Exception as e:
- print(f' {label}: FAILED — {str(e)[:60]}')
- return None
- print('=== SENSITIVITY 1: Stratified by Baseline Diagnosis ===')
- print('MMSE models by BL diagnosis:\n')
- strat_results = {}
- for dx in ['CN', 'MCI', 'AD']:
- if 'BL_DX_LABEL' not in complete.columns:
- print(f' BL_DX_LABEL not available — cannot stratify')
- break
- sub = complete[complete['BL_DX_LABEL'] == dx].copy()
- cov = build_covariates(sub)
- r = run_lme_simple(sub, 'MMSCORE', cov, f'MMSE-{dx}')
- strat_results[dx] = r
- print('\nResults summary:')
- for dx, r in strat_results.items():
- if r is not None:
- # Focus on key interaction terms
- key_terms = ['YEARS_FROM_BL', 'APOE4_DOSE', 'HIPPO_ICV_ADJ',
- 'YEARS_FROM_BL:APOE4_DOSE', 'YEARS_FROM_BL:HIPPO_ICV_ADJ']
- print(f'\n {dx} stratum:')
- for t in key_terms:
- if t in r.fe_params:
- p = r.pvalues[t]
- print(f' {t:<35} β={r.fe_params[t]:>7.3f}, p={"<0.001" if p<0.001 else f"{p:.3f}"}')
- # %%
- # Compare interaction coefficients across diagnostic groups (forest-style comparison)
- if any(v is not None for v in strat_results.values()):
- target_term = 'YEARS_FROM_BL:APOE4_DOSE'
- alt_terms = ['YEARS_FROM_BL:APOE4_DOSE', 'YEARS_FROM_BL:HIPPO_ICV_ADJ', 'APOE4_DOSE']
- comparison_rows = []
- for dx, r in strat_results.items():
- if r is None: continue
- for term in alt_terms:
- if term in r.fe_params:
- comparison_rows.append({
- 'Diagnosis': dx, 'Term': term,
- 'Coef': r.fe_params[term],
- 'CI_lo': r.conf_int().loc[term, 0],
- 'CI_hi': r.conf_int().loc[term, 1],
- 'P_value': r.pvalues[term]
- })
- if comparison_rows:
- comp_df = pd.DataFrame(comparison_rows)
- comp_df.to_csv(os.path.join(REPORTS, 'Sensitivity1_Stratified_Dx.csv'), index=False)
- print('\n✓ Saved: Sensitivity1_Stratified_Dx.csv')
- # Plot comparison
- for term in alt_terms:
- sub = comp_df[comp_df['Term'] == term]
- if len(sub) < 2: continue
- fig, ax = plt.subplots(figsize=(7, 4))
- colors_dx = {'CN': '#4CAF50', 'MCI': '#FF9800', 'AD': '#F44336'}
- for _, row in sub.iterrows():
- dx = row['Diagnosis']
- ax.errorbar(row['Coef'], dx,
- xerr=[[row['Coef']-row['CI_lo']], [row['CI_hi']-row['Coef']]],
- fmt='o', color=colors_dx.get(dx, 'grey'), capsize=5, markersize=8)
- ax.axvline(0, color='black', linewidth=1, linestyle='--', alpha=0.7)
- ax.set_xlabel('Coefficient (95% CI)')
- ax.set_title(f'Stratified Analysis: {term}\nMMSE Outcome', fontweight='bold')
- plt.tight_layout()
- safe_term = term.replace(':', '_').replace('/', '_')
- plt.savefig(os.path.join(REPORTS, f'FigStrat_{safe_term}.png'), dpi=150, bbox_inches='tight')
- plt.show()
- break # just first term
- # %% [markdown]
- # ## 3. Sensitivity Analysis 2 — Exclude Hippocampal Outliers (±3 SD)
- # %%
- print('=== SENSITIVITY 2: Exclude Hippocampal Outliers (±3 SD) ===')
- hippo_mean = complete['HIPPO_ICV_ADJ'].mean()
- hippo_sd = complete['HIPPO_ICV_ADJ'].std()
- low_cut = hippo_mean - 3*hippo_sd
- high_cut = hippo_mean + 3*hippo_sd
- print(f'Hippocampal mean: {hippo_mean:.4f}, SD: {hippo_sd:.4f}')
- print(f'Exclusion range: < {low_cut:.4f} or > {high_cut:.4f}')
- outlier_mask = (complete['HIPPO_ICV_ADJ'] < low_cut) | (complete['HIPPO_ICV_ADJ'] > high_cut)
- n_outliers = outlier_mask.sum()
- print(f'Outliers: {n_outliers} observations ({n_outliers/len(complete)*100:.1f}%)')
- no_outlier_df = complete[~outlier_mask].copy()
- print(f'\nPost-exclusion: {no_outlier_df["RID"].nunique()} subjects, {len(no_outlier_df):,} obs')
- cov_no = build_covariates(no_outlier_df)
- if 'GDTOTAL' in no_outlier_df.columns:
- no_outlier_df['GDTOTAL'] = no_outlier_df['GDTOTAL'].fillna(0)
- r_no_outlier = run_lme_simple(no_outlier_df, 'MMSCORE', cov_no, 'No-outlier MMSE')
- if r_no_outlier is not None:
- print('\nKey terms (no-outlier model):')
- key_terms = ['YEARS_FROM_BL', 'APOE4_DOSE', 'HIPPO_ICV_ADJ',
- 'YEARS_FROM_BL:APOE4_DOSE', 'YEARS_FROM_BL:HIPPO_ICV_ADJ',
- 'YEARS_FROM_BL:APOE4_DOSE:HIPPO_ICV_ADJ']
- for t in key_terms:
- if t in r_no_outlier.fe_params:
- p = r_no_outlier.pvalues[t]
- print(f' {t:<40} β={r_no_outlier.fe_params[t]:>8.4f}, p={"<0.001" if p<0.001 else f"{p:.3f}"}')
- # %% [markdown]
- # ## 4. Sensitivity Analysis 3 — Subjects with ≥2 Visits Only
- # %%
- print('=== SENSITIVITY 3: ≥2 Visits Only ===')
- visit_counts = complete.groupby('RID').size()
- multi_rids = visit_counts[visit_counts >= 2].index
- multi_df = complete[complete['RID'].isin(multi_rids)].copy()
- print(f'Original: {complete["RID"].nunique()} subjects')
- print(f'≥2 visits: {multi_df["RID"].nunique()} subjects ({multi_df["RID"].nunique()/complete["RID"].nunique()*100:.1f}%)')
- print(f'Mean visits: {visit_counts[multi_rids].mean():.1f}')
- # Compare demographics between single vs multi-visit (check selection bias)
- single_rids = visit_counts[visit_counts == 1].index
- if len(single_rids) > 0:
- bl_multi = baseline[baseline['RID'].isin(multi_rids)]
- bl_single = baseline[baseline['RID'].isin(single_rids)]
- print('\nSelection check — Multi-visit vs Single-visit subjects:')
- for var in ['AGE', 'MMSCORE', 'APOE4_DOSE', 'HIPPO_ICV_ADJ']:
- if var in bl_multi.columns:
- m_m = bl_multi[var].mean()
- m_s = bl_single[var].mean()
- _, p = stats.ttest_ind(bl_multi[var].dropna(), bl_single[var].dropna())
- print(f' {var:<18} Multi: {m_m:.2f} Single: {m_s:.2f} p={"<0.001" if p<0.001 else f"{p:.3f}"}')
- cov_multi = build_covariates(multi_df)
- if 'GDTOTAL' in multi_df.columns:
- multi_df['GDTOTAL'] = multi_df['GDTOTAL'].fillna(0)
- r_multi = run_lme_simple(multi_df, 'MMSCORE', cov_multi, '≥2-visit MMSE')
- if r_multi is not None:
- print('\nKey terms (≥2 visits model):')
- for t in ['YEARS_FROM_BL:APOE4_DOSE', 'YEARS_FROM_BL:HIPPO_ICV_ADJ',
- 'YEARS_FROM_BL:APOE4_DOSE:HIPPO_ICV_ADJ']:
- if t in r_multi.fe_params:
- p = r_multi.pvalues[t]
- print(f' {t:<45} β={r_multi.fe_params[t]:>8.4f}, p={"<0.001" if p<0.001 else f"{p:.3f}"}')
- # %% [markdown]
- # ## 5. Sensitivity Analysis 4 — Robust Standard Errors
- # MMSE has ceiling effects in CN subjects. Use robust SE to account for non-normality (Reviewer 2.3).
- # %%
- print('=== SENSITIVITY 4: MMSE Ceiling Effect & Robust Errors ===')
- # Show MMSE distribution
- fig, axes = plt.subplots(1, 3, figsize=(15, 4))
- # Overall MMSE distribution at baseline
- ax = axes[0]
- bl_mmse = baseline['MMSCORE'].dropna()
- ax.hist(bl_mmse, bins=20, color='steelblue', edgecolor='white', density=True)
- ax.axvline(28, color='red', linestyle='--', alpha=0.7, label='Ceiling zone (≥28)')
- ax.set_xlabel('MMSE score')
- ax.set_ylabel('Density')
- ax.set_title('MMSE Distribution (Baseline)\nNote ceiling effect near 30', fontweight='bold')
- ax.legend()
- ceil_pct = (bl_mmse >= 28).mean() * 100
- ax.text(0.05, 0.95, f'{ceil_pct:.1f}% ≥ 28', transform=ax.transAxes, va='top')
- # MMSE by APOE dose at baseline
- ax = axes[1]
- for dose in [0, 1, 2]:
- sub = baseline[baseline['APOE4_DOSE'] == dose]['MMSCORE'].dropna()
- ax.hist(sub, bins=15, alpha=0.5, color=APOE_COLORS[dose],
- label=f'APOE{dose} (n={len(sub)})', edgecolor='none', density=True)
- ax.set_xlabel('MMSE score')
- ax.set_ylabel('Density')
- ax.set_title('MMSE by APOE Dose (Baseline)', fontweight='bold')
- ax.legend(fontsize=9)
- # MMSE × Hippocampus scatter (all visits)
- ax = axes[2]
- if 'HIPPO_ICV_ADJ' in complete.columns:
- sample = complete.dropna(subset=['MMSCORE', 'HIPPO_ICV_ADJ']).sample(min(1000, len(complete)))
- ax.scatter(sample['HIPPO_ICV_ADJ'], sample['MMSCORE'],
- alpha=0.2, s=8, color='steelblue')
- slope, intercept, r, p, _ = stats.linregress(sample['HIPPO_ICV_ADJ'], sample['MMSCORE'])
- x = np.linspace(sample['HIPPO_ICV_ADJ'].min(), sample['HIPPO_ICV_ADJ'].max(), 100)
- ax.plot(x, slope*x + intercept, color='red', linewidth=2)
- ax.set_xlabel('ICV-adjusted hippocampal volume')
- ax.set_ylabel('MMSE score')
- ax.set_title(f'MMSE vs Hippocampus\nr={r:.2f}, p={"<0.001" if p<0.001 else f"{p:.3f}"}', fontweight='bold')
- plt.tight_layout()
- plt.savefig(os.path.join(REPORTS, 'FigSensitivity_MMSE_Distribution.png'), dpi=150, bbox_inches='tight')
- plt.show()
- print('✓ Saved: FigSensitivity_MMSE_Distribution.png')
- # Apply floor/ceiling sensitivity analysis
- print(f'\nMMSE ceiling analysis:')
- print(f' Overall: {(bl_mmse >= 28).mean()*100:.1f}% at ceiling (≥28)')
- if 'BL_DX_LABEL' in baseline.columns:
- for dx in ['CN', 'MCI', 'AD']:
- sub = baseline[baseline['BL_DX_LABEL'] == dx]['MMSCORE'].dropna()
- if len(sub) > 0:
- print(f' {dx}: {(sub >= 28).mean()*100:.1f}% at ceiling')
- # %%
- # Robust standard errors via statsmodels OLS with cluster-robust SE
- # (as sensitivity to LME)
- print('\nRobust SE approach — pooled OLS with cluster-robust SE (sensitivity check):')
- robust_data = complete.dropna(subset=['MMSCORE', 'YEARS_FROM_BL', 'APOE4_DOSE', 'HIPPO_ICV_ADJ']).copy()
- if 'GDTOTAL' in robust_data.columns:
- robust_data['GDTOTAL'] = robust_data['GDTOTAL'].fillna(0)
- cov_parts = [v for v in ['AGE', 'SEX_MALE', 'PTEDUCAT', 'GDTOTAL'] if v in robust_data.columns]
- if 'BL_DX_LABEL' in robust_data.columns:
- cov_parts.append('C(BL_DX_LABEL)')
- formula_ols = ("MMSCORE ~ YEARS_FROM_BL * APOE4_DOSE * HIPPO_ICV_ADJ" +
- (f" + {' + '.join(cov_parts)}" if cov_parts else ""))
- try:
- import patsy
- y, X = patsy.dmatrices(formula_ols, data=robust_data, return_type='dataframe')
- # Align groups to patsy's post-NaN-drop index
- robust_aligned = robust_data.loc[y.index].reset_index(drop=True)
- ols_model = sm.OLS(y, X)
- # Cluster-robust SE (clusters = subjects)
- ols_result_robust = ols_model.fit(cov_type='cluster', cov_kwds={'groups': robust_aligned['RID'].values})
- print(ols_result_robust.summary().tables[1])
- # Save
- with open(os.path.join(REPORTS, 'Sensitivity4_Robust_SE_MMSE.txt'), 'w') as f:
- f.write(ols_result_robust.summary().as_text())
- print('\n✓ Saved: Sensitivity4_Robust_SE_MMSE.txt')
- except Exception as e:
- print(f'Robust SE failed: {e}')
- # %% [markdown]
- # ## 6. Sensitivity Results Summary Table
- # %%
- # Compile key interaction terms across sensitivity analyses
- from importlib import import_module
- print('=== SENSITIVITY ANALYSIS SUMMARY ===')
- print('Key interaction term: Time × APOE4_DOSE × HIPPO_ICV_ADJ (3-way)\n')
- def extract_term(result, term):
- if result is None or term not in result.fe_params: return None, None, None
- coef = result.fe_params[term]
- ci_lo = result.conf_int().loc[term, 0]
- ci_hi = result.conf_int().loc[term, 1]
- p = result.pvalues[term]
- return coef, (ci_lo, ci_hi), p
- summary_rows = []
- for label, res in [
- ('No outliers (±3 SD)', r_no_outlier),
- ('≥2 visits only', r_multi),
- ('CN subgroup', strat_results.get('CN')),
- ('MCI subgroup', strat_results.get('MCI')),
- ('AD subgroup', strat_results.get('AD')),
- ]:
- for term in ['YEARS_FROM_BL:APOE4_DOSE', 'YEARS_FROM_BL:HIPPO_ICV_ADJ']:
- coef, ci, p = extract_term(res, term)
- if coef is not None:
- summary_rows.append({
- 'Analysis': label, 'Term': term,
- 'Coef': round(coef, 4),
- 'CI_95': f'[{ci[0]:.3f}, {ci[1]:.3f}]',
- 'P_value': '<0.001' if p < 0.001 else f'{p:.3f}',
- 'Sig': '***' if p<0.001 else '**' if p<0.01 else '*' if p<0.05 else ''
- })
- if summary_rows:
- sens_summary = pd.DataFrame(summary_rows)
- print(sens_summary.to_string(index=False))
- sens_summary.to_csv(os.path.join(REPORTS, 'Sensitivity_Summary_All.csv'), index=False)
- print('\n✓ Saved: Sensitivity_Summary_All.csv')
- else:
- print('No sensitivity results available to compile')
- # %% [markdown]
- # ## 7. Output Summary
- # %%
- print('SURVIVAL & SENSITIVITY ANALYSIS COMPLETE')
- print('='*60)
- print('\nFiles in reports/:')
- for f in sorted(os.listdir(REPORTS)):
- size = os.path.getsize(os.path.join(REPORTS, f)) / 1024
- print(f' {f:<50} {size:>8.1f} KB')
04_survival_sensitivity.ipynb at commit dc0e2c2, under MIT · at the source
Overview
Abstract
Background/
Methods: This study analyzed data from 2,417 Alzheimer’s Disease Neuroimaging Initiative participants with complete APOE genotypes, intracranial volume-adjusted hippocampal volumes, and longitudinal cognitive assessments spanning a mean follow-up of 4.2 years and up to 19.3 years with an average of 4.9 visits per participant. Linear mixed-effects models with random intercepts and slopes per subject estimated cognitive trajectories across the Mini-Mental State Examination (MMSE), Clinical Dementia Rating Sum of Boxes (CDR-SB), and Alzheimer’s Disease Assessment Scale Cognitive Subscale 13 (ADAS-Cog13) as a function of time, APOE ε4 dose, and ICV-adjusted hippocampal volume, including their three-way interaction and adjusting for age, sex, education, baseline diagnosis, and depression. Cox proportional hazards models were used to assess conversion risk.
Results: A clear APOE ε4 dose–response gradient was observed at baseline across all cognitive and hippocampal measures (all p < 0.001). Linear mixed-effects models revealed a significant three-way interaction of time × APOE ε4 dose × hippocampal volume on MMSE (β = −0.79, 95% CI [−1.51, −0.08], p = 0.030) and CDR-SB (β = +0.47, 95% CI [+0.03, +0.91], p = 0.037) trajectories, both significant under Bonferroni correction (α = 0.017), indicating that APOE ε4 amplifies the association between smaller hippocampal volume and faster cognitive deterioration over time. The time × hippocampal volume interaction was confirmed as highly significant by likelihood ratio test (LR = 3712.99, p < 0.001). Cox proportional hazards analyses of 845 conversion events showed that each additional ε4 allele conferred a 48% increase in conversion risk (HR = 1.48, 95% CI [1.29, 1.71], p < 0.001). Sensitivity analyses across diagnostic strata, after outlier exclusion, and in multi-visit subsamples confirmed the robustness of hippocampal volume effects.
Discussion/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 20 matches between paragraphs and lines of code.
FaizaanFazal/adni-apoe4-hippocampus-cognitive-decline
dc0e2c2273170867e2b47f128e6bf30cc42e9b0c, 22 April 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
8 files
- notebooks2/
01_data_pipeline.ipynb , Jupyter, 630 lines, 2 matches - notebooks2/
02_descriptive_stats.ipy , Jupyter, 430 lines, 4 matchesnb - notebooks2/
03_lme_analysis.ipynb , Jupyter, 572 lines, 5 matches - notebooks2/
04_survival_sensitivity. , Jupyter, 570 lines, 9 matchesipynb - notebooks2/
05_datailed_counts.ipynb , Jupyter, 314 lines - notebooks2/
05_detailed_counts.py , Python, 312 lines - LICENSE, License, 21 lines
- Readme.md, Text, 208 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 6 scripts, each with its path and the digest of its content;
- 20 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data and code availability
The data used in this study were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu) and are available to qualified researchers upon request. The analysis code supporting the findings of this study will be made available by the authors upon acceptance of the manuscript.
Reproduced under the paper's license (CC BY), from the paper cited above.
Data availability statement
The data used in this study were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database () and are available to qualified researchers upon application to ADNI. The analysis code supporting the findings of this study is publicly available at: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 4 authors, 8 keywords, 4 funders, 15 references.
Cite
This paper
Khan, F. F., Kim, J.-H., Kim, J.-I., & Kwon, G.-R. (2026). APOE ε4 as a predictor of cognitive decline and its interaction with hippocampal volume in Alzheimer's disease. Frontiers in aging neuroscience, 18, 1730265. https://
BibTeX
@article{khan2026apoe,
author = {Khan, Faizaan Fazal and Kim, Jun-Hyung and Kim, Ji-In and Kwon, Goo-Rak},
title = {{APOE ε4 as a predictor of cognitive decline and its interaction with hippocampal volume in Alzheimer's disease}},
journal = {Frontiers in aging neuroscience},
year = {2026},
month = apr,
volume = {18},
pages = {1730265},
publisher = {Frontiers Media SA},
issn = {1663-4365},
doi = {10.3389/
url = {https://
pmid = {42100481},
pmcid = {PMC13144130}
}
RIS
TY - JOUR
AU - Khan, Faizaan Fazal
AU - Kim, Jun-Hyung
AU - Kim, Ji-In
AU - Kwon, Goo-Rak
TI - APOE ε4 as a predictor of cognitive decline and its interaction with hippocampal volume in Alzheimer's disease
T2 - Frontiers in aging neuroscience
J2 - Front Aging Neurosci
PY - 2026
DA - 2026/
VL - 18
SP - 1730265
SN - 1663-4365
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "APOE ε4 as a predictor of cognitive decline and its interaction with hippocampal volume in Alzheimer's disease",
"container-title": "Frontiers in aging neuroscience",
"author": [
{
"family": "Khan",
"given": "Faizaan Fazal"
},
{
"family": "Kim",
"given": "Jun-Hyung"
},
{
"family": "Kim",
"given": "Ji-In"
},
{
"family": "Kwon",
"given": "Goo-Rak"
}
],
"container-title-short":
"volume": "18",
"page": "1730265",
"DOI": "10.3389/
"PMID": "42100481",
"PMCID": "PMC13144130",
"ISSN": "1663-4365",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
22
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41398-026-04143-x [code]
- Neighborhood opportunity, white matter, and cognition in a large pediatric neuroimaging study.Journal: Translational psychiatryIn common: statsmodels, seaborn, pandas, 3 other tools, structural MRI / diffusion, 1 reference
- [2] doi:10.1016/j.cell.2026.05.047 [code]
- An emergent disease-associated motor neuron state precedes cell death in ALS.Journal: CellIn common: statsmodels, seaborn, pandas, 3 other tools, genetics / omics, 1 reference
- [3] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: statsmodels, seaborn, pandas, 3 other tools, genetics / omics, 1 reference
- [4] doi:10.1016/j.isci.2026.116825 [code]
- Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.Journal: iScienceIn common: statsmodels, seaborn, pandas, 3 other tools, genetics / omics, 1 reference
- [5] doi:10.1038/s41514-026-00390-w [code]
- Mild cognitive impairment cases affect the predictive power of Alzheimer's disease diagnostic models using routine clinical variables.Journal: npj agingIn common: seaborn, pandas, SciPy, 2 other tools, Alzheimer's / dementia, clinical / translational, 1 reference
- [6] doi:10.1073/pnas.2516601123 [code]
- Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: seaborn, pandas, SciPy, 2 other tools, structural MRI / diffusion, clinical / translational, 1 reference
- [7] doi:10.1186/s40708-026-00316-y [code]
- Generalizable and explainable deep learning for brain MRI: a multi-cohort evaluation of 3D architectures for age and sex prediction.Journal: Brain informaticsIn common: seaborn, pandas, SciPy, 2 other tools, structural MRI / diffusion, 1 reference
- [8] doi:10.1162/imag.a.1278 [code]
- Young and old adult brains experience opposite effects of acute sleep restriction on the functional connectivity network.Journal: Imaging neuroscience (Cambridge, Mass.)In common: statsmodels, seaborn, pandas, 3 other tools, 1 reference
- [9] doi:10.1093/cercor/bhag077 [code]
- The longitudinal development of intrinsic timescales in infancy and their relation to alpha brain rhythm.Journal: Cerebral cortex (New York, N.Y. : 1991)In common: statsmodels, seaborn, pandas, 3 other tools, 1 reference
- [10] doi:10.1162/imag.a.1246 [code]
- A 3.5-minute-long reading-based fMRI localizer for the language network.Journal: Imaging neuroscience (Cambridge, Mass.)In common: statsmodels, seaborn, pandas, 3 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 6 scripts, and 20 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e45c00f604b6a0aa…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
