Genetic drivers of progression in Alzheimer's disease are distinct from disease risk.
The 3 matches
- [1] § Methods › Genomic quality control and imputation ↔ Progression_Genetic_Modelling/Genetic QC/genetic_qc_consolidated.R, lines 1–59 · score 0.85 · Hardy Weinberg Equilibrium, minor allele frequency, missing genotyping, deviations, QC, quality
- [2] § Methods › Patient filtering ↔ Progression_Genetic_Modelling/Clinical Data Modelling and Filtering/comprehensive_analysis.qmd, lines 84–105 · score 0.70 · Pittsburgh compound, positive patients, CSF, AV45, PiB, filtered
- [3] § Methods › Mixed-effects modelling ↔ Progression_Genetic_Modelling/Clinical Data Modelling and Filtering/comprehensive_analysis.qmd, lines 150–186 · score 0.56 · random intercept, baseline MMSE, PC1, PC2, education, age
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Quarto · 1,071 lines · 36 KB · no license · 2 matches
- ---
- title: "Exploratory Data Analysis, Model Optimisation and Clinical Filtering"
- author:
- - "CCOHEN"
- - "MSHOAI"
- date: today
- format:
- html:
- theme: cosmo
- toc: true
- toc-depth: 3
- toc-location: left
- number-sections: true
- code-fold: true
- code-tools: true
- df-print: kable
- fig-width: 10
- fig-height: 6
- embed-resources: true
- execute:
- warning: false
- message: false
- cache: false
- ---
- # Introduction
- This document presents a comprehensive analysis of longitudinal cognitive decline in Alzheimer's Disease (AD) patients from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The analysis includes:
- - Quality control and patient filtering based on amyloid-beta biomarkers
- - Exploratory data analysis and descriptive statistics
- - Mixed-effects model optimization and comparison
- - Integration with polygenic risk scores
- # Data Loading and Preparation
- ```{r setup, warning=FALSE, message=FALSE}
- # Load required libraries
- library(tidyverse)
- library(tseries)
- library(visdat)
- library(ggplot2)
- library(dplyr)
- library(gridExtra)
- library(Amelia)
- library(corrplot)
- library(knitr)
- library(lme4)
- library(lmerTest)
- library(performance)
- library(fitdistrplus)
- library(glmmTMB)
- # Set options for output formatting
- options(knitr.kable.NA = '', digits = 2)
- ```
- ```{r load_data}
- # Read ADNIMERGE clinical data
- adnimerge <- readRDS("./QCjune/clinicaladni/adnimerge.rds")
- # Select relevant columns
- dataset <- adnimerge[, c(4, 9, 10, 11, 114, 27, 69, 61, 8, 18, 110, 17, 109, 20, 105, 14, 15, 1)]
- colnames(dataset) <- c("Patient_ID", "Age", "Gender", "Education", "Months_since_bl",
- "MMSE", "MMSE.bl", "DX", "DX.bl", "AV45", "AV45.bl", "PIB", "PIB.bl",
- "CSF_AB", "CSF_AB.bl", "Marital", "APOE4", "RID")
- # Clean CSF Amyloid-beta values
- dataset$CSF_AB.bl <- gsub(">", "", dataset$CSF_AB.bl)
- dataset$CSF_AB.bl <- gsub("<", "", dataset$CSF_AB.bl)
- dataset$CSF_AB.bl <- as.numeric(dataset$CSF_AB.bl)
- dataset$CSF_AB <- gsub(">", "", dataset$CSF_AB)
- dataset$CSF_AB <- gsub("<", "", dataset$CSF_AB)
- dataset$CSF_AB <- as.numeric(dataset$CSF_AB)
- ```
- ## Initial Dataset Summary
- ```{r initial_summary}
- cat("Total unique patients:", length(unique(dataset$Patient_ID)), "\n")
- cat("Total datapoints:", nrow(dataset), "\n")
- ```
- # Patient Filtering Pipeline
- ## Step 1: Amyloid-Beta Positive Selection {#step1}
- Patients are classified as amyloid-beta positive if they meet at least one of the following criteria at baseline:
- - **AV45 (Florbetapir PET)**: ≥ 1.1
- - **PIB (Pittsburgh Compound B PET)**: ≥ 1.5
- - **CSF Amyloid-beta**: ≤ 550 pg/mL
- ```{r filter_amyloid}
- # Define thresholds
- th_amyl_AV45 <- 1.1
- th_amyl_PIB <- 1.5
- th_amyl_CSF <- 550
- # Filter for amyloid-beta positive patients
- df_amyloid <- subset(dataset, AV45.bl >= th_amyl_AV45 | PIB.bl >= th_amyl_PIB | CSF_AB.bl <= th_amyl_CSF)
- cat("Amyloid-beta positive patients:", length(unique(df_amyloid$Patient_ID)), "\n")
- cat("Number of datapoints:", nrow(df_amyloid), "\n")
- ```
- ### Amyloid-Beta Distribution
- ```{r amyloid_dist, fig.height=4}
- # Create binary amyloid status variable
- dataset$Amyloid_Status <- ifelse(
- dataset$AV45.bl >= th_amyl_AV45 | dataset$PIB.bl >= th_amyl_PIB | dataset$CSF_AB.bl <= th_amyl_CSF,
- "Positive", "Negative"
- )
- # Plot distribution
- ggplot(dataset %>% group_by(Patient_ID) %>% slice(1), aes(x = Amyloid_Status, fill = Amyloid_Status)) +
- geom_bar() +
- geom_text(stat='count', aes(label=..count..), vjust=-0.5) +
- labs(title = "Distribution of Amyloid-Beta Status",
- x = "Amyloid-Beta Status", y = "Number of Patients") +
- scale_fill_manual(values = c("Positive" = "#E74C3C", "Negative" = "#3498DB")) +
- theme_minimal() +
- theme(legend.position = "none", text = element_text(size = 12))
- ```
- # Model Optimization on Amyloid-Positive Cohort
- Before applying additional filtering, we test multiple mixed-effects models on the amyloid-positive cohort to identify the optimal model structure.
- ```{r prepare_model_data}
- # Prepare data for modeling (amyloid+ only, at this stage)
- df_model1 <- df_amyloid
- df_model1$Gender <- as.factor(df_model1$Gender)
- df_model1$Months_since_bl <- as.numeric(as.character(df_model1$Months_since_bl))
- # For this initial modeling, we need PC1 and PC2
- # Read PCA data
- pcs <- read.table("./QCjune/adnimerge_genomes/FINALADNI_nohapmap_pca.eigenvec")[, 2:4]
- colnames(pcs) <- c("PTID", "PC1", "PC2")
- # Merge with clinical data
- df_model1$PTID <- as.character(df_model1$Patient_ID)
- df_model1_pcs <- inner_join(df_model1, pcs, by = "PTID", multiple = "all")
- cat("Amyloid+ patients with genetic data for modeling:", length(unique(df_model1_pcs$Patient_ID)), "\n")
- cat("Datapoints for modeling:", nrow(df_model1_pcs), "\n")
- ```
- ## Model 1-9 Comparison
- We test nine hierarchical models of increasing complexity:
- ```{r model_comparison, results='hide'}
- # Model 1: Time only
- model1 <- lm(MMSE ~ Months_since_bl, data = df_model1_pcs)
- # Model 2: Time + Baseline age
- model2 <- lm(MMSE ~ Months_since_bl + Age, data = df_model1_pcs)
- # Model 3: Time + Current age
- model3 <- lm(MMSE ~ Months_since_bl + Age, data = df_model1_pcs)
- # Model 4: Time + Baseline MMSE + Baseline age
- model4 <- lm(MMSE ~ MMSE.bl + Months_since_bl + Age, data = df_model1_pcs)
- # Model 5: Random intercept
- model5 <- lmer(MMSE ~ MMSE.bl + Age + Months_since_bl + (1|Patient_ID),
- data = df_model1_pcs, REML = TRUE)
- # Model 6: Random intercept + Education
- model6 <- lmer(MMSE ~ MMSE.bl + Age + Months_since_bl + (1|Patient_ID) + Education,
- data = df_model1_pcs, REML = TRUE)
- # Model 7: Random intercept + Education + Gender
- model7 <- lmer(MMSE ~ MMSE.bl + Age + Months_since_bl + (1|Patient_ID) + Education + Gender,
- data = df_model1_pcs, REML = TRUE)
- # Model 8: Random intercept + Education + Gender + PCs
- model8 <- lmer(MMSE ~ MMSE.bl + Age + Months_since_bl + (1|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_model1_pcs, REML = TRUE)
- # Model 9: Random slope and intercept + Education + Gender + PCs
- model9 <- lmer(MMSE ~ MMSE.bl + Age + Months_since_bl + (Months_since_bl|Patient_ID) +
- Education + Gender + PC1 + PC2, data = df_model1_pcs, REML = TRUE)
- ```
- ```{r model_performance}
- # Compile model performance metrics
- models <- list(model1, model2, model3, model4, model5, model6, model7, model8, model9)
- model_names <- paste0("Model ", 1:9)
- # Get performance metrics manually
- aics <- sapply(models, AIC)
- bics <- sapply(models, BIC)
- r2_vals <- sapply(models, function(m) {
- r2_result <- tryCatch(r2(m), error = function(e) list(R2 = NA))
- if ("R2" %in% names(r2_result)) return(r2_result$R2)
- if ("R2_conditional" %in% names(r2_result)) return(r2_result$R2_conditional)
- return(NA)
- })
- performance_table <- data.frame(
- Model = model_names,
- AIC = aics,
- BIC = bics,
- R2 = r2_vals
- )
- kable(performance_table,
- caption = "Performance Comparison of Models 1-9 on Amyloid-Positive Cohort",
- digits = 2)
- ```
- ::: {.callout-note}
- ## Best Model Selection
- Based on AIC, BIC, and R² metrics, **Model 9** (random slopes and intercepts with full covariates) provides the best fit. This model will be used for subsequent analyses after full filtering.
- :::
- # Continued Patient Filtering
- ## Step 2: Remove Missing MMSE and Require ≥3 Datapoints
- ```{r filter_mmse}
- df <- df_amyloid
- cat("Patients before filtering:", length(unique(df$Patient_ID)), "\n")
- cat("Datapoints before filtering:", nrow(df), "\n\n")
- # Remove rows with missing MMSE
- df <- df[-which(is.na(df$MMSE)), ]
- # Require at least 3 datapoints per patient
- Keep <- df %>%
- group_by(Patient_ID) %>%
- summarize(datapoints = n()) %>%
- filter(datapoints >= 3)
- df <- df[df$Patient_ID %in% Keep$Patient_ID, ]
- cat("Patients after filtering:", length(unique(df$Patient_ID)), "\n")
- cat("Datapoints after filtering:", nrow(df), "\n")
- ```
- ## Step 3: Remove "Always Cognitively Normal" Patients
- ```{r filter_always_cn}
- # Identify patients who are CN at baseline
- PTID_CNbl <- unique(df[df$Months_since_bl == 0 & df$DX == "CN", ]$Patient_ID)
- PTID_CNbl_df <- df[df$Patient_ID %in% PTID_CNbl, ]
- # Find those who progress to MCI or Dementia
- PTID_CNthenMCI <- unique(PTID_CNbl_df[which(PTID_CNbl_df$DX == "MCI" | PTID_CNbl_df$DX == "Dementia"), ]$Patient_ID)
- # Remove always-CN patients
- PTID_alwaysCN <- PTID_CNbl[-which(PTID_CNbl %in% PTID_CNthenMCI)]
- df_clean <- df[-which(df$Patient_ID %in% PTID_alwaysCN), ]
- cat("Patients not always cognitively normal:", length(unique(df_clean$Patient_ID)), "\n")
- cat("Datapoints:", nrow(df_clean), "\n")
- ```
- ## Step 4: Remove Patients with Minimum MMSE ≥ 28
- ```{r filter_min_mmse}
- # Calculate minimum MMSE per patient
- minMMSEperpatient <- df_clean %>%
- group_by(Patient_ID) %>%
- summarise(minMMSE = min(MMSE))
- min28 <- minMMSEperpatient$Patient_ID[which(minMMSEperpatient$minMMSE >= 28)]
- df_clean <- df_clean[!df_clean$Patient_ID %in% min28, ]
- cat("Patients with minimum MMSE < 28:", length(unique(df_clean$Patient_ID)), "\n")
- cat("Datapoints:", nrow(df_clean), "\n")
- ```
- ## Step 5: Remove Patients with Last MMSE 28-30 OR Last DX = CN
- ```{r filter_last_status}
- # Calculate last MMSE and DX per patient
- lastMMSE <- c()
- lastDX <- c()
- for (i in unique(df_clean$Patient_ID)) {
- d <- subset(df_clean, Patient_ID == i)
- d_sort <- d[order(d$Months_since_bl, decreasing = FALSE), ]
- lastMMSE <- c(lastMMSE, d_sort$MMSE[length(d_sort$MMSE)])
- lastDX <- c(lastDX, d_sort$DX[length(d_sort$DX)])
- }
- # Remove patients with last MMSE 28-30
- df_lastMMSE30 <- subset(df_clean, Patient_ID %in% unique(df_clean$Patient_ID)[which(lastMMSE == 30 | lastMMSE == 29 | lastMMSE == 28)])
- df_clean <- df_clean[!df_clean$Patient_ID %in% df_lastMMSE30$Patient_ID, ]
- # Remove patients with last DX = CN
- df_lastDXCN <- subset(df_clean, Patient_ID %in% unique(df_clean$Patient_ID)[which(lastDX == "CN")])
- df_clean <- df_clean[!df_clean$Patient_ID %in% df_lastDXCN$Patient_ID, ]
- cat("Patients without ceiling effects:", length(unique(df_clean$Patient_ID)), "\n")
- cat("Datapoints:", nrow(df_clean), "\n")
- ```
- ## Step 6: Remove Datapoints Before MMSE = 30
- For patients who achieved MMSE of 30 at any point, we remove all observations before the last occurrence of MMSE = 30. This focuses the analysis on the decline phase.
- ```{r filter_before_30}
- before_count <- nrow(df_clean)
- df_clean$dpid <- seq(1, nrow(df_clean))
- rm_rows <- c()
- for (i in unique(df_clean$Patient_ID)) {
- d <- subset(df_clean, Patient_ID == i)
- d_sort <- d[order(d$Months_since_bl, decreasing = FALSE), ]
- instances <- c()
- for (n in 1:(length(d_sort$MMSE) - 1)) {
- if (d_sort$MMSE[n] == 30) {
- instances <- c(instances, n)
- }
- }
- if (length(instances) > 0) {
- rm <- d_sort$dpid[1:(max(instances) - 1)]
- rm_rows <- c(rm_rows, rm)
- }
- }
- df_clean <- df_clean[!df_clean$dpid %in% rm_rows, ]
- df_clean$dpid <- NULL
- cat("Datapoints removed:", before_count - nrow(df_clean), "\n")
- cat("Patients remaining:", length(unique(df_clean$Patient_ID)), "\n")
- cat("Datapoints remaining:", nrow(df_clean), "\n")
- ```
- ## Step 7: Keep Only Last 5 Observations Per Patient
- This critical step focuses the analysis on recent disease trajectory.
- ```{r filter_last_5}
- before_count <- nrow(df_clean)
- df_clean$dpid <- seq(1, nrow(df_clean))
- rm_rows <- c()
- for (i in unique(df_clean$Patient_ID)) {
- d <- subset(df_clean, Patient_ID == i)
- d_sort <- d[order(d$Months_since_bl, decreasing = FALSE), ]
- if (length(d$dpid) > 5) {
- rm <- d_sort$dpid[1:(length(d_sort$dpid) - 5)]
- rm_rows <- c(rm_rows, rm)
- }
- }
- df_clean <- df_clean[!df_clean$dpid %in% rm_rows, ]
- df_clean$dpid <- NULL
- cat("Datapoints removed:", before_count - nrow(df_clean), "\n")
- cat("Datapoints remaining:", nrow(df_clean), "\n")
- ```
- ## Step 8: Require ≥3 Datapoints (After Pruning)
- ```{r filter_3pts_final}
- Keep <- df_clean %>%
- group_by(Patient_ID) %>%
- summarize(datapoints = n()) %>%
- filter(datapoints >= 3)
- df_clean <- df_clean[df_clean$Patient_ID %in% Keep$Patient_ID, ]
- cat("Final clinically filtered patients:", length(unique(df_clean$Patient_ID)), "\n")
- cat("Final datapoints:", nrow(df_clean), "\n")
- ```
- ::: {.callout-important}
- ## Filtering Summary
- Starting with **`r length(unique(dataset$Patient_ID))`** patients, the filtering pipeline identified **`r length(unique(df_clean$Patient_ID))`** patients with:
- - Confirmed amyloid-beta pathology
- - Sufficient longitudinal data (3-5 recent observations)
- - Evidence of cognitive decline
- :::
- # Descriptive Statistics
- ## Missing Data Visualization
- ```{r missing_data, fig.height=8}
- # Visualize missing data patterns
- missmap(df_clean[, c("MMSE", "Age", "Education", "Gender", "DX", "AV45", "PIB", "CSF_AB", "APOE4")],
- main = "Missing Data Map - Filtered Cohort",
- col = c("lightcoral", "lightblue"),
- legend = TRUE,
- x.cex = 0.8,
- y.cex = 0.6)
- ```
- ## Baseline Characteristics
- ```{r baseline_char}
- # Get baseline characteristics (one row per patient)
- df_baseline <- df_clean %>%
- group_by(Patient_ID) %>%
- arrange(Months_since_bl) %>%
- slice(1) %>%
- ungroup()
- # Summary statistics
- baseline_stats <- data.frame(
- Variable = c("Age", "Education", "MMSE"),
- Mean = c(mean(df_baseline$Age, na.rm = TRUE),
- mean(df_baseline$Education, na.rm = TRUE),
- mean(df_baseline$MMSE, na.rm = TRUE)),
- SD = c(sd(df_baseline$Age, na.rm = TRUE),
- sd(df_baseline$Education, na.rm = TRUE),
- sd(df_baseline$MMSE, na.rm = TRUE)),
- Median = c(median(df_baseline$Age, na.rm = TRUE),
- median(df_baseline$Education, na.rm = TRUE),
- median(df_baseline$MMSE, na.rm = TRUE)),
- Min = c(min(df_baseline$Age, na.rm = TRUE),
- min(df_baseline$Education, na.rm = TRUE),
- min(df_baseline$MMSE, na.rm = TRUE)),
- Max = c(max(df_baseline$Age, na.rm = TRUE),
- max(df_baseline$Education, na.rm = TRUE),
- max(df_baseline$MMSE, na.rm = TRUE))
- )
- kable(baseline_stats, caption = "Baseline Characteristics (Continuous Variables)", digits = 2)
- # Categorical variables
- cat_summary <- data.frame(
- Variable = c("Gender (Male)", "Gender (Female)",
- "DX.bl (AD)", "DX.bl (MCI)", "DX.bl (CN)", "DX.bl (Dementia)"),
- N = c(sum(df_baseline$Gender == "Male", na.rm = TRUE),
- sum(df_baseline$Gender == "Female", na.rm = TRUE),
- sum(df_baseline$DX.bl == "AD", na.rm = TRUE),
- sum(df_baseline$DX.bl == "MCI" | df_baseline$DX.bl == "EMCI" | df_baseline$DX.bl == "LMCI", na.rm = TRUE),
- sum(df_baseline$DX.bl == "CN", na.rm = TRUE),
- sum(df_baseline$DX.bl == "Dementia", na.rm = TRUE)),
- Percentage = c(100 * sum(df_baseline$Gender == "Male", na.rm = TRUE) / nrow(df_baseline),
- 100 * sum(df_baseline$Gender == "Female", na.rm = TRUE) / nrow(df_baseline),
- 100 * sum(df_baseline$DX.bl == "AD", na.rm = TRUE) / nrow(df_baseline),
- 100 * sum(df_baseline$DX.bl == "MCI" | df_baseline$DX.bl == "EMCI" | df_baseline$DX.bl == "LMCI", na.rm = TRUE) / nrow(df_baseline),
- 100 * sum(df_baseline$DX.bl == "CN", na.rm = TRUE) / nrow(df_baseline),
- 100 * sum(df_baseline$DX.bl == "Dementia", na.rm = TRUE) / nrow(df_baseline))
- )
- kable(cat_summary, caption = "Baseline Characteristics (Categorical Variables)", digits = 2)
- ```
- ## Age Distribution by Diagnosis and Gender
- ```{r age_distributions, fig.height=8}
- # Plot age distribution by baseline diagnosis
- p1 <- ggplot(df_baseline, aes(x = Age, fill = DX.bl)) +
- geom_histogram(binwidth = 5, alpha = 0.7) +
- labs(title = "Age Distribution by Baseline Diagnosis",
- x = "Age (years)", y = "Count", fill = "Baseline DX") +
- theme_minimal() +
- theme(text = element_text(size = 11))
- # Plot age distribution by gender
- p2 <- ggplot(df_baseline, aes(x = Age, fill = Gender)) +
- geom_histogram(binwidth = 5, alpha = 0.7, position = "identity") +
- labs(title = "Age Distribution by Gender",
- x = "Age (years)", y = "Count", fill = "Gender") +
- scale_fill_manual(values = c("Male" = "#3498DB", "Female" = "#E74C3C")) +
- theme_minimal() +
- theme(text = element_text(size = 11))
- # Boxplot of age by gender with t-test
- t_test_result <- t.test(Age ~ Gender, data = df_baseline)
- p3 <- ggplot(df_baseline, aes(x = Gender, y = Age, fill = Gender)) +
- geom_boxplot(alpha = 0.7) +
- labs(title = paste0("Age by Gender (p = ", format(t_test_result$p.value, digits = 3, scientific = TRUE), ")"),
- x = "Gender", y = "Age (years)") +
- scale_fill_manual(values = c("Male" = "#3498DB", "Female" = "#E74C3C")) +
- theme_minimal() +
- theme(legend.position = "none", text = element_text(size = 11))
- # MMSE trajectory over time by gender
- p4 <- ggplot(df_clean, aes(x = Months_since_bl, y = MMSE, color = Gender)) +
- geom_point(alpha = 0.3, size = 0.8) +
- geom_smooth(method = "lm", se = TRUE, size = 1.2) +
- labs(title = "MMSE Trajectory Over Time by Gender",
- x = "Months Since Baseline", y = "MMSE Score", color = "Gender") +
- scale_color_manual(values = c("Male" = "#3498DB", "Female" = "#E74C3C")) +
- theme_minimal() +
- theme(text = element_text(size = 11))
- # Arrange plots
- grid.arrange(p1, p2, p3, p4, ncol = 2)
- ```
- ## Age and Progression Analysis
- ### Progression Rate by Age Group
- ```{r age_progression, fig.height=6}
- # Calculate progression slopes by age groups
- agerange <- c(50, 60, 65, 70, 75, 80, 85, 95)
- slopes <- c()
- age_group <- c()
- for (i in 1:(length(agerange) - 1)) {
- low <- agerange[i]
- high <- agerange[i + 1]
- dfa <- df_clean[which(df_clean$Age >= low & df_clean$Age < high), ]
- for (id in unique(dfa$Patient_ID)) {
- d <- dfa[which(dfa$Patient_ID == id), ]
- if (nrow(d) >= 2) {
- slopes <- c(slopes, coef(lm(MMSE ~ Months_since_bl, data = d))[2])
- age_group <- c(age_group, paste(low, "to", high))
- }
- }
- }
- age_prog_df <- data.frame(Age_Group = age_group, Progression_Rate = slopes)
- # Box plot
- ggplot(age_prog_df, aes(x = Age_Group, y = Progression_Rate)) +
- geom_boxplot(fill = "#3498DB", alpha = 0.7) +
- geom_hline(yintercept = 0, linetype = "dashed", color = "red") +
- labs(title = "Cognitive Decline Rate by Age Group",
- x = "Age Group (years)", y = "MMSE Change per Month",
- caption = paste("ANOVA p-value:", format(anova(lm(Progression_Rate ~ Age_Group, data = age_prog_df))$Pr[1], digits = 3))) +
- theme_minimal() +
- theme(axis.text.x = element_text(angle = 45, hjust = 1), text = element_text(size = 11))
- ```
- ### Progression by Age of Onset
- ```{r age_onset, fig.height=4}
- # Analyze progression by age of onset (for dementia patients)
- df_dementia <- df_clean[which(df_clean$DX == "Dementia"), ]
- slopes_dem <- c()
- onset_age <- c()
- for (i in unique(df_dementia$Patient_ID)) {
- d <- subset(df_dementia, Patient_ID == i)
- if (nrow(d) >= 2) {
- slopes_dem <- c(slopes_dem, coef(lm(MMSE ~ Months_since_bl, data = d))[2])
- onset_age <- c(onset_age, min(d$Age))
- }
- }
- onset_df <- data.frame(Onset_Age = onset_age, Progression_Rate = slopes_dem)
- onset_df$Onset_Category <- cut(onset_df$Onset_Age,
- breaks = c(-Inf, 65, 75, Inf),
- labels = c("EOAD (<65)", "MOAD (65-75)", "LOAD (>75)"))
- # Plot
- p1 <- ggplot(onset_df, aes(x = Onset_Age, y = Progression_Rate)) +
- geom_point(alpha = 0.6, color = "#E74C3C") +
- geom_smooth(method = "lm", se = TRUE, color = "#3498DB") +
- labs(title = "Progression Rate vs Age of Onset",
- x = "Age of Onset (years)", y = "MMSE Change per Month") +
- theme_minimal()
- p2 <- ggplot(onset_df, aes(x = Onset_Category, y = Progression_Rate, fill = Onset_Category)) +
- geom_boxplot(alpha = 0.7) +
- labs(title = "Progression by Onset Category",
- x = "Age of Onset Category", y = "MMSE Change per Month") +
- scale_fill_brewer(palette = "Set2") +
- theme_minimal() +
- theme(legend.position = "none")
- grid.arrange(p1, p2, ncol = 2)
- # Statistical test
- lm_onset <- lm(Progression_Rate ~ Onset_Age, data = onset_df)
- cat("\nLinear regression p-value:", format(summary(lm_onset)$coefficients[2, 4], digits = 3), "\n")
- ```
- # Final Model on Filtered Cohort
- ## Data Preparation
- ```{r prepare_final_model}
- # Calculate baseline values
- df_clean <- df_clean[order(df_clean$Patient_ID, df_clean$Months_since_bl), ]
- Age_bl <- c()
- MMSE_bl <- c()
- DX_bl <- c()
- AV45_bl <- c()
- PIB_bl <- c()
- CSF_AB_bl <- c()
- prograte <- c()
- visits <- c()
- for (i in unique(df_clean$Patient_ID)) {
- d <- subset(df_clean, Patient_ID == i)
- datapoints <- nrow(d)
- visits <- c(visits, rep(datapoints, datapoints))
- # Calculate progression rate
- if (datapoints >= 2) {
- prograte <- c(prograte, rep(coef(lm(MMSE ~ Months_since_bl, data = d))[2], datapoints))
- } else {
- prograte <- c(prograte, rep(NA, datapoints))
- }
- Age_bl <- c(Age_bl, rep(min(d$Age), datapoints))
- MMSE_bl <- c(MMSE_bl, rep(d$MMSE[1], datapoints))
- DX_bl <- c(DX_bl, rep(d$DX[1], datapoints))
- AV45_bl <- c(AV45_bl, rep(d$AV45[1], datapoints))
- PIB_bl <- c(PIB_bl, rep(d$PIB[1], datapoints))
- CSF_AB_bl <- c(CSF_AB_bl, rep(d$CSF_AB[1], datapoints))
- }
- df_clean$Age_bl <- Age_bl
- df_clean$MMSE_bl <- MMSE_bl
- df_clean$DX_bl <- DX_bl
- df_clean$AV45_bl_calc <- AV45_bl
- df_clean$PIB_bl_calc <- PIB_bl
- df_clean$CSF_AB_bl_calc <- CSF_AB_bl
- df_clean$prograte <- prograte
- df_clean$visits <- visits
- # Remove old baseline columns
- df_clean$MMSE.bl <- NULL
- df_clean$CSF_AB.bl <- NULL
- df_clean$PIB.bl <- NULL
- df_clean$AV45.bl <- NULL
- df_clean$DX.bl <- NULL
- # Convert variables
- df_clean$Months_since_bl <- as.numeric(as.character(df_clean$Months_since_bl))
- df_clean$Gender <- as.factor(df_clean$Gender)
- df_clean$PTID <- as.character(df_clean$Patient_ID)
- # Merge with PCA data
- pcs <- read.table("./QCjune/adnimerge_genomes/FINALADNI_nohapmap_pca.eigenvec")[, 2:4]
- colnames(pcs) <- c("PTID", "PC1", "PC2")
- df_final <- inner_join(df_clean, pcs, by = "PTID", multiple = "all")
- cat("Patients with clinical and genetic data:", length(unique(df_final$Patient_ID)), "\n")
- cat("Final datapoints for modeling:", nrow(df_final), "\n")
- ```
- ## Optimal Model: Random Slopes and Intercepts
- Based on the model comparison in @step1, we fit Model 9 to the filtered cohort:
- ```{r final_model}
- final_model <- lmer(MMSE ~ MMSE_bl + Age_bl + Months_since_bl + (Months_since_bl|Patient_ID) +
- Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE)
- summary(final_model)
- ```
- ## Model Diagnostics
- ```{r model_diagnostics, fig.height=8}
- # Shapiro-Wilk test for normality of residuals
- shapiro_result <- shapiro.test(resid(final_model))
- cat("Shapiro-Wilk test for normality:\n")
- cat(" W =", format(shapiro_result$statistic, digits = 4), "\n")
- cat(" p-value =", format(shapiro_result$p.value, digits = 4), "\n\n")
- # Homogeneity of variances test
- hom_var_test <- function(m) {
- p <- coef(summary(lm(abs(resid(m)) ~ fitted(m))))[2, 4]
- return(p)
- }
- hom_p <- hom_var_test(final_model)
- cat("Homogeneity of variance test:\n")
- cat(" p-value =", format(hom_p, digits = 4), "\n")
- cat(" ", ifelse(hom_p < 0.05, "❌ Assumption violated", "✓ Assumption met"), "\n\n")
- # Model performance
- model_perf <- performance::r2(final_model)
- cat("Model Performance:\n")
- cat(" R² Conditional:", format(model_perf$R2_conditional, digits = 3), "\n")
- cat(" R² Marginal:", format(model_perf$R2_marginal, digits = 3), "\n")
- cat(" AIC:", format(AIC(final_model), digits = 1), "\n")
- cat(" BIC:", format(BIC(final_model), digits = 1), "\n")
- ```
- ### Random Effects Visualization
- ```{r random_effects, fig.height=6}
- # Extract random effects
- ranef_data <- as.data.frame(ranef(final_model)$Patient_ID)
- ranef_data$Patient_ID <- rownames(ranef_data)
- colnames(ranef_data)[1:2] <- c("Intercept", "Slope")
- # Merge with baseline diagnosis
- group_info <- df_final %>%
- dplyr::select(Patient_ID, DX_bl) %>%
- dplyr::distinct()
- ranef_with_group <- merge(ranef_data, group_info, by = "Patient_ID")
- # Plot random slopes by baseline diagnosis
- ggplot(ranef_with_group, aes(x = Slope, fill = DX_bl)) +
- geom_density(alpha = 0.6) +
- geom_vline(xintercept = 0, linetype = "dashed", color = "black", size = 1) +
- labs(title = "Distribution of Individual Progression Rates by Baseline Diagnosis",
- x = "Random Slope (MMSE change per month)",
- y = "Density",
- fill = "Baseline DX") +
- scale_fill_brewer(palette = "Set2") +
- theme_minimal() +
- theme(text = element_text(size = 12))
- ```
- # Alternative Model Specifications
- ## Testing MMSE_bl vs DX_bl as Baseline Predictor
- ```{r baseline_predictor_comparison, results='hide'}
- # Model with MMSE_bl (may have convergence issues)
- model_mmse <- glmmTMB(MMSE ~ MMSE_bl + Age_bl + Months_since_bl * APOE4 +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE)
- # Model with MMSE_bl and BFGS optimizer
- model_mmse_bfgs <- glmmTMB(MMSE ~ MMSE_bl + Age_bl + Months_since_bl * APOE4 +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE,
- control = glmmTMBControl(optimizer = optim, optArgs = list(method = "BFGS")))
- # Model with DX_bl
- model_dx <- glmmTMB(MMSE ~ DX_bl + Age_bl + Months_since_bl * APOE4 +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE)
- # Model with DX_bl and BFGS optimizer
- model_dx_bfgs <- glmmTMB(MMSE ~ DX_bl + Age_bl + Months_since_bl * APOE4 +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE,
- control = glmmTMBControl(optimizer = optim, optArgs = list(method = "BFGS")))
- ```
- ```{r baseline_comparison}
- models_baseline <- list(model_mmse, model_mmse_bfgs, model_dx, model_dx_bfgs)
- model_names_baseline <- c("MMSE_bl (default)", "MMSE_bl (BFGS)", "DX_bl (default)", "DX_bl (BFGS)")
- # Performance metrics
- aics <- sapply(models_baseline, AIC)
- r2cond <- sapply(models_baseline, function(m) r2(m)$R2_conditional)
- r2marg <- sapply(models_baseline, function(m) r2(m)$R2_marginal)
- shap_w <- sapply(models_baseline, function(m) shapiro.test(resid(m))$statistic)
- shap_p <- sapply(models_baseline, function(m) shapiro.test(resid(m))$p.value)
- homosced <- sapply(models_baseline, function(m) hom_var_test(m))
- perform_baseline <- data.frame(
- Model = model_names_baseline,
- AIC = aics,
- R2_Conditional = r2cond,
- R2_Marginal = r2marg,
- Shapiro_W = shap_w,
- Shapiro_p = shap_p,
- Homoscedasticity_p = homosced
- )
- kable(perform_baseline, caption = "Model Comparison: MMSE_bl vs DX_bl as Baseline Predictor", digits = 2)
- ```
- # Polygenic Score Analysis
- ## Polygenic Hazard Score (PHS)
- ```{r phs_analysis}
- # Read PHS data
- PHS <- read.csv("./QCjune/clinicaladni/DESIKANLAB_18Jun2024.csv")
- df_final$RID <- as.integer(df_final$RID)
- df_PHS <- inner_join(df_final, PHS, by = "RID", multiple = "all")
- cat("Patients with PHS data:", length(unique(df_PHS$Patient_ID)), "\n")
- cat("Datapoints:", nrow(df_PHS), "\n")
- # Model with PHS
- model_PHS <- glmmTMB(MMSE ~ DX_bl + Age_bl + PHS * Months_since_bl +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_PHS, REML = TRUE)
- # Extract coefficients
- phs_coef <- summary(model_PHS)$coef$cond
- kable(phs_coef, caption = "PHS Model Coefficients", digits = 4)
- ```
- ## Polygenic Risk Scores (PRS)
- ```{r prs_analysis}
- # Read PRS data
- PRS <- read.csv("./QCjune/clinicaladni/PRS_andrealtmann.csv")[, 1:5]
- df_PRS <- inner_join(df_final, PRS, by = "RID", multiple = "all")
- cat("Patients with PRS data:", length(unique(df_PRS$Patient_ID)), "\n")
- cat("Datapoints:", nrow(df_PRS), "\n")
- # Models with different PRS
- model_PRS1 <- glmmTMB(MMSE ~ DX_bl + Age_bl + PRS1 * Months_since_bl +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_PRS, REML = TRUE)
- model_PRS2 <- glmmTMB(MMSE ~ DX_bl + Age_bl + PRS2 * Months_since_bl +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_PRS, REML = TRUE)
- model_PRScs <- glmmTMB(MMSE ~ DX_bl + Age_bl + PRScs_auto * Months_since_bl +
- (1 + Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_PRS, REML = TRUE)
- # Summary table
- prs_results <- data.frame(
- Model = c("PRS1", "PRS2", "PRScs_auto"),
- AIC = c(AIC(model_PRS1), AIC(model_PRS2), AIC(model_PRScs)),
- PRS_Interaction_Beta = c(
- summary(model_PRS1)$coef$cond["PRS1:Months_since_bl", "Estimate"],
- summary(model_PRS2)$coef$cond["PRS2:Months_since_bl", "Estimate"],
- summary(model_PRScs)$coef$cond["PRScs_auto:Months_since_bl", "Estimate"]
- ),
- PRS_Interaction_p = c(
- summary(model_PRS1)$coef$cond["PRS1:Months_since_bl", "Pr(>|z|)"],
- summary(model_PRS2)$coef$cond["PRS2:Months_since_bl", "Pr(>|z|)"],
- summary(model_PRScs)$coef$cond["PRScs_auto:Months_since_bl", "Pr(>|z|)"]
- )
- )
- kable(prs_results, caption = "Polygenic Risk Score Models - Interaction Effects", digits = 4)
- ```
- # Advanced Model Testing
- ## Data Normalization
- Testing whether data transformation improves model fit:
- ```{r normalization}
- # Reflect and transform MMSE_bl (beta-distributed)
- MMSE_bl_refl <- max(df_final$MMSE_bl + 1) - df_final$MMSE_bl
- MMSE_bl_norm <- sqrt(MMSE_bl_refl)
- df_final$MMSE_bl_norm <- MMSE_bl_norm
- # Reflect and transform MMSE
- MMSE_refl <- max(df_final$MMSE + 1) - df_final$MMSE
- MMSE_norm <- sqrt(MMSE_refl)
- df_final$MMSE_norm <- MMSE_norm
- # Normalized model
- model_norm <- lmer(MMSE_norm ~ MMSE_bl_norm + Age_bl + Months_since_bl +
- (Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_final, REML = TRUE)
- ```
- ## Data Scaling
- ```{r scaling}
- # Scale numeric variables
- df_scaled <- df_final
- numeric_cols <- sapply(df_scaled, is.numeric)
- for (col in names(df_scaled)[numeric_cols]) {
- if (col != "Patient_ID" && col != "RID") {
- df_scaled[[col]] <- round((df_scaled[[col]] - mean(df_scaled[[col]], na.rm = TRUE)) /
- sd(df_scaled[[col]], na.rm = TRUE), 3)
- }
- }
- # Scaled model
- model_scaled <- lmer(MMSE ~ MMSE_bl + Age_bl + Months_since_bl +
- (Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_scaled, REML = TRUE)
- # Normalized and scaled model
- model_norm_scaled <- lmer(MMSE_norm ~ MMSE_bl_norm + Age_bl + Months_since_bl +
- (Months_since_bl|Patient_ID) + Education + Gender + PC1 + PC2,
- data = df_scaled, REML = TRUE)
- ```
- ## Comprehensive Model Comparison
- ```{r model_comparison_all}
- models_all <- list(final_model, model_scaled, model_norm, model_norm_scaled)
- model_names_all <- c("Original", "Scaled", "Normalized", "Normalized + Scaled")
- aics <- sapply(models_all, AIC)
- bics <- sapply(models_all, BIC)
- r2cond <- sapply(models_all, function(m) r2(m)$R2_conditional)
- r2marg <- sapply(models_all, function(m) r2(m)$R2_marginal)
- shap_w <- sapply(models_all, function(m) shapiro.test(resid(m))$statistic)
- shap_p <- sapply(models_all, function(m) shapiro.test(resid(m))$p.value)
- homosced <- sapply(models_all, function(m) hom_var_test(m))
- comparison_table <- data.frame(
- Model = model_names_all,
- AIC = aics,
- BIC = bics,
- R2_Conditional = r2cond,
- R2_Marginal = r2marg,
- Shapiro_W = shap_w,
- Shapiro_p = shap_p,
- Homoscedasticity_p = homosced
- )
- kable(comparison_table,
- caption = "Comprehensive Model Comparison: Transformation and Scaling Effects",
- digits = 4)
- ```
- ::: {.callout-tip}
- ## Model Selection Summary
- Lower AIC/BIC values indicate better model fit. Higher R² values indicate more variance explained. The Shapiro-Wilk test assesses normality of residuals (p > 0.05 desired). Homoscedasticity test assesses equal variance (p > 0.05 desired).
- :::
- # Patient Summary Table
- ## Generate Comprehensive Summary
- ```{r patient_summary_table}
- # Read FAM file for genotype status
- fam_data <- read.table("./QCjune/adnimerge_genomes/adnimerge.fam", header = FALSE)
- fam_ptids <- as.character(fam_data[, 2])
- # Get list of filtered patients
- filtered_patients <- unique(df_clean$Patient_ID)
- # Create summary for ALL patients in original dataset
- unique_patients <- unique(dataset$Patient_ID)
- summary_list <- list()
- for (patient in unique_patients) {
- patient_data <- dataset[dataset$Patient_ID == patient, ]
- patient_data <- patient_data[order(patient_data$Months_since_bl), ]
- # Basic info
- ptid <- as.character(patient)
- n_obs <- nrow(patient_data)
- n_mmse_obs <- sum(!is.na(patient_data$MMSE))
- # Baseline diagnosis
- dx_baseline <- as.character(patient_data$DX.bl[1])
- if (is.na(dx_baseline) || dx_baseline == "") dx_baseline <- "NA"
- # Last diagnosis
- max_idx <- which.max(patient_data$Months_since_bl)
- dx_last <- as.character(patient_data$DX[max_idx])
- if (is.na(dx_last) || dx_last == "") dx_last <- "NA"
- # Last MMSE
- last_mmse <- patient_data$MMSE[max_idx]
- if (is.na(last_mmse)) last_mmse <- "NA" else last_mmse <- as.character(last_mmse)
- # In FAM file
- in_fam <- ifelse(any(grepl(ptid, fam_ptids, fixed = TRUE)), "Yes", "No")
- # Amyloid-beta status
- av45_bl <- patient_data$AV45.bl[1]
- pib_bl <- patient_data$PIB.bl[1]
- csf_ab_bl <- patient_data$CSF_AB.bl[1]
- abeta_pos <- FALSE
- if (!is.na(av45_bl) && av45_bl >= th_amyl_AV45) abeta_pos <- TRUE
- if (!is.na(pib_bl) && pib_bl >= th_amyl_PIB) abeta_pos <- TRUE
- if (!is.na(csf_ab_bl) && csf_ab_bl <= th_amyl_CSF) abeta_pos <- TRUE
- abeta_status <- ifelse(abeta_pos, "AbetaPos", "AbetaNeg")
- # DX trajectory
- dx_trajectory <- sapply(patient_data$DX, function(x) {
- if (is.na(x) || x == "") return("NA")
- return(as.character(x))
- })
- dx_trajectory_str <- paste(dx_trajectory, collapse = "_")
- # Always CN
- dx_values_no_na <- patient_data$DX[!is.na(patient_data$DX) & patient_data$DX != ""]
- always_cn <- ""
- if (length(dx_values_no_na) > 0 && all(dx_values_no_na == "CN")) {
- always_cn <- "alwaysCN"
- }
- # Survives filtering
- survives <- ifelse(ptid %in% filtered_patients, "Yes", "No")
- summary_list[[patient]] <- data.frame(
- PTID = ptid,
- N_Observations = n_obs,
- N_MMSE_Observations = n_mmse_obs,
- DX_Baseline = dx_baseline,
- DX_Last = dx_last,
- Last_MMSE = last_mmse,
- In_FAM_File = in_fam,
- AbetaPos = abeta_status,
- DX_Trajectory = dx_trajectory_str,
- AlwaysCN = always_cn,
- Survives_Filtering = survives,
- stringsAsFactors = FALSE
- )
- }
- # Combine into data frame
- summary_table <- bind_rows(summary_list)
- summary_table <- summary_table[order(summary_table$PTID), ]
- # Save table
- write.csv(summary_table, "./patient_summary_table_complete.csv", row.names = FALSE, quote = TRUE)
- cat("Total patients in summary table:", nrow(summary_table), "\n")
- cat("Patients surviving filtering:", sum(summary_table$Survives_Filtering == "Yes"), "\n")
- cat("Patients not surviving filtering:", sum(summary_table$Survives_Filtering == "No"), "\n")
- ```
- ## Summary Table Preview
- ```{r summary_preview}
- kable(head(summary_table, 20),
- caption = "Patient Summary Table (First 20 Patients)",
- digits = 2)
- ```
- ## Filtering Statistics
- ```{r filtering_stats}
- filter_stats <- data.frame(
- Category = c("Total Patients", "Amyloid-Beta Positive", "After MMSE Filtering",
- "After Clinical Filtering", "With Genotype Data", "Final Cohort"),
- Count = c(
- length(unique(dataset$Patient_ID)),
- length(unique(df_amyloid$Patient_ID)),
- length(unique(df$Patient_ID)),
- length(unique(df_clean$Patient_ID)),
- length(unique(df_final$Patient_ID)),
- length(unique(df_final$Patient_ID))
- ),
- Percentage = c(
- 100,
- 100 * length(unique(df_amyloid$Patient_ID)) / length(unique(dataset$Patient_ID)),
- 100 * length(unique(df$Patient_ID)) / length(unique(dataset$Patient_ID)),
- 100 * length(unique(df_clean$Patient_ID)) / length(unique(dataset$Patient_ID)),
- 100 * length(unique(df_final$Patient_ID)) / length(unique(dataset$Patient_ID)),
- 100 * length(unique(df_final$Patient_ID)) / length(unique(dataset$Patient_ID))
- )
- )
- kable(filter_stats, caption = "Filtering Pipeline Statistics", digits = 2)
- ```
- # Conclusions
- This comprehensive analysis identified **`r length(unique(df_final$Patient_ID))` patients** with:
- - Confirmed amyloid-beta pathology
- - Sufficient longitudinal cognitive data (3-5 observations)
- - Evidence of cognitive decline
- - High-quality genotype data
- The optimal model (random slopes and intercepts with full covariates) achieved:
- - **R² Conditional**: `r format(r2(final_model)$R2_conditional, digits = 3)`
- - **R² Marginal**: `r format(r2(final_model)$R2_marginal, digits = 3)`
- - **AIC**: `r format(AIC(final_model), digits = 1)`
- Polygenic scores (PHS and PRS) were successfully integrated and show no effect on progression.
- # Session Information
- ```{r session_info}
- sessionInfo()
- ```
comprehensive_analysis.qmd at commit 0bd2f20, no license · at the source
Overview
- Department of Neurodegenerative Disease, UCL Queen Square Institute of Neurology,London, UK
- Wellcome Sanger Institute,Cambridge, UK
- Centre for Precision Health, Edith Cowan University,Joondalup, WA Australia
- Collaborative Genomics and Translation Group, School of Medical and Health Sciences, Edith Cowan University,Joondalup, WA Australia
- Dementia Research Institute, University College London,London, UK
- Max-Planck-Institute for Intelligent Systems,Tubingen, Germany
- UK Dementia Research Institute, Imperial College London,London, UK
- Department of Brain Sciences, Imperial College London,London, UK
- Florey Department of Neuroscience and Mental Health, University of Melbourne,Melbourne, Australia
- CogState Ltd,Melbourne, Australia
- Reta Lila Weston Institute of Neurological Studies,London, UK
Abstract
Background: Recent trials in Alzheimer’s disease (AD) demonstrate encouraging outcomes. These trials target risk mechanisms identified through genetic analysis whilst directly aiming to reduce progression rates. Evidence from other neurodegenerative diseases suggests the genetics of progression is distinct from risk of disease. To expand these initial successes and improve clinical outcomes further we need to understand genetics of progression of disease. These can be deduced through rigorous analysis of meticulously phenotyped longitudinal cohorts. In this study we first looked at known genetic drivers of risk, namely polygenic risk scores for AD and APOE‑ε4, to assess their role in progression. This was then extended to a genome wide association analysis to identify the role of other genetic variants in progression of AD.
Methods: A total of 387 individuals with genetic data, amyloid positivity, and in active decline (ADNI (n = 222) and AIBL(n = 165)) were used to perform generalised mixed effects linear model genome wide association studies of longitudinal cognitive decline as measured by mini mental state examination (MMSE). The resulting summary statistics were subjected to functional annotation, and colocalisation analyses.
Results: Established AD risk factors, including APOE‑ε4 dosage and polygenic risk scores, were not associated with disease progression in amyloid positive individuals who are actively declining. A mixed effects GWAS meta-analysis revealed one genome-wide significant locus on chromosome 22 (rs78369883) and several nominally significant loci linked with AD progression. Functional annotation, finemapping, and colocalisation analyses implicated genes primarily involved in immune response, neurodegeneration (including tau pathology), brain resilience, and neurogenesis. These progression-related genes were significantly enriched in neuronal-interferon-micr
Conclusion: These findings enhance our understanding of the biological underpinnings of AD progression, opening new avenues for therapeutic intervention.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.
MaryamShoai/Codes
0bd2f20200b3fcd0d8403aada2da72be3393bb58, 16 October 2025Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
7 files
- Progression_Genetic_Mode
lling/ , Quarto, 1,071 lines, 2 matchesClinical Data Modelling and Filtering/ comprehensive_analysis.q md - Progression_Genetic_Mode
lling/ , R, 656 lines, 1 matchGenetic QC/ genetic_qc_consolidated. R - Progression_Genetic_Mode
lling/ , R, 126 linesModelling in parallel/ 1_data_partition.R - Progression_Genetic_Mode
lling/ , R, 90 linesModelling in parallel/ 2_data_modelling.R - Progression_Genetic_Mode
lling/ , R, 112 linesModelling in parallel/ 3_parallel_modelling.R - Progression_Measures/
Forestplots_Alz_dem_prog , R, 417 linesression_measures.R - README.md, Text, 21 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 6 scripts, each with its path and the digest of its content;
- 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
The codes used for running the mixed effect models and the full resulting summary statistics are available from: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 16 authors, 5 keywords, 14 MeSH terms, 8 funders, 103 references.
Cite
This paper
Cohen, C. E., Fernandez, S., Yaman, U., Ehyaei, A. R., Kodosaki, E., Askarova, A., Porter, T., O’Brien, E., Australian Imaging Biomarkers and Lifestyle Study, Alzheimer’s Disease Neuroimaging Initiative, Maruff, P., Nott, A., Hardy, J. A., Laws, S. M., A. Salih, D., & Shoai, M. (2026). Genetic drivers of progression in Alzheimer's disease are distinct from disease risk. Alzheimer's research & therapy, 18(1), 142. https://
BibTeX
@article{cohen2026geneti
author = {Cohen, Celeste E. and Fernandez, Shane and Yaman, Umran and Ehyaei, Ahmad R. and Kodosaki, Eleftheria and Askarova, Aydan and Porter, Tenielle and O’Brien, Eleanor and {Australian Imaging Biomarkers and Lifestyle Study} and {Alzheimer’s Disease Neuroimaging Initiative} and Maruff, Paul and Nott, Alexi and Hardy, John A. and Laws, Simon M. and A. Salih, Dervis and Shoai, Maryam},
title = {{Genetic drivers of progression in Alzheimer's disease are distinct from disease risk}},
journal = {Alzheimer's research \& therapy},
year = {2026},
month = apr,
volume = {18},
number = {1},
pages = {142},
publisher = {BMC},
issn = {1758-9193},
doi = {10.1186/
url = {https://
pmid = {42035106},
pmcid = {PMC13248444}
}
RIS
TY - JOUR
AU - Cohen, Celeste E.
AU - Fernandez, Shane
AU - Yaman, Umran
AU - Ehyaei, Ahmad R.
AU - Kodosaki, Eleftheria
AU - Askarova, Aydan
AU - Porter, Tenielle
AU - O’Brien, Eleanor
AU - Australian Imaging Biomarkers and Lifestyle Study
AU - Alzheimer’s Disease Neuroimaging Initiative
AU - Maruff, Paul
AU - Nott, Alexi
AU - Hardy, John A.
AU - Laws, Simon M.
AU - A. Salih, Dervis
AU - Shoai, Maryam
TI - Genetic drivers of progression in Alzheimer's disease are distinct from disease risk
T2 - Alzheimer's research & therapy
J2 - Alzheimers Res Ther
PY - 2026
DA - 2026/
VL - 18
IS - 1
SP - 142
SN - 1758-9193
PB - BMC
DO - 10.1186/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1186/
"type": "article-journal",
"title": "Genetic drivers of progression in Alzheimer's disease are distinct from disease risk",
"container-title": "Alzheimer's research & therapy",
"author": [
{
"family": "Cohen",
"given": "Celeste E."
},
{
"family": "Fernandez",
"given": "Shane"
},
{
"family": "Yaman",
"given": "Umran"
},
{
"family": "Ehyaei",
"given": "Ahmad R."
},
{
"family": "Kodosaki",
"given": "Eleftheria"
},
{
"family": "Askarova",
"given": "Aydan"
},
{
"family": "Porter",
"given": "Tenielle"
},
{
"family": "O’Brien",
"given": "Eleanor"
},
{
"literal": "Australian Imaging Biomarkers and Lifestyle Study"
},
{
"literal": "Alzheimer’s Disease Neuroimaging Initiative"
},
{
"family": "Maruff",
"given": "Paul"
},
{
"family": "Nott",
"given": "Alexi"
},
{
"family": "Hardy",
"given": "John A."
},
{
"family": "Laws",
"given": "Simon M."
},
{
"family": "A. Salih",
"given": "Dervis"
},
{
"family": "Shoai",
"given": "Maryam"
}
],
"container-title-short":
"volume": "18",
"issue": "1",
"page": "142",
"DOI": "10.1186/
"PMID": "42035106",
"PMCID": "PMC13248444",
"ISSN": "1758-9193",
"publisher": "BMC",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
25
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41380-026-03686-1 [code]
- Early oligodendrocyte dysfunction signature in Alzheimer's disease: Insights from DNA methylomics and transcriptomics.Journal: Molecular psychiatryIn common: data.table, ggplot2, tidyverse, Alzheimer's / dementia, genetics / omics, 5 references, 2 authors
- [2] doi:10.1038/s41562-026-02486-5 [code]
- Genome-wide association studies of infant and toddler temperament in European and multi-ancestry populations.Journal: Nature human behaviourIn common: metafor, nlme, data.table, 2 other tools, genetics / omics, 7 references
- [3] doi:10.1016/j.celrep.2026.117505 [code]
- Impaired spatial coding and neuronal hyperactivity in the medial entorhinal cortex of aged APP knock-in mice.Journal: Cell reportsIn common: glmmTMB, metafor, nlme, 6 other tools, Alzheimer's / dementia
- [4] doi:10.1371/journal.pgen.1012170 [code]
- Genome wide association study meta-analysis of neuropathologic lesions of Alzheimer's disease and related dementias in a multi-site autopsy cohort.Journal: PLoS geneticsIn common: data.table, ggplot2, tidyverse, Alzheimer's / dementia, clinical / translational, genetics / omics, 8 references
- [5] doi:10.1002/hbm.70605 [code]
- BrainEnrich: Revealing Biological Insights for Imaging-Derived Phenotypes Through Transcriptomic Enrichment.Journal: Human brain mappingIn common: nlme, reticulate, easystats, 5 other tools, genetics / omics, 1 reference
- [6] doi:10.1038/s41467-026-71682-8 [code]
- GWAS meta-analysis of cerebrospinal fluid Alzheimer's biomarkers reveals loci regulating lipids, brain volume and autophagy.Journal: Nature communicationsIn common: data.table, ggplot2, tidyverse, Alzheimer's / dementia, genetics / omics, 7 references
- [7] doi:10.1038/s41467-026-73865-9 [code]
- Histamine shapes the neurocomputational dynamics of human learning.Journal: Nature communicationsIn common: metafor, easystats, emmeans, 6 other tools
- [8] doi:10.1073/pnas.2606871123 [code]
- Oxytocin modulates the neurocomputational mechanisms engaged in learning rank relationships in social networks.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: nlme, easystats, emmeans, 6 other tools
- [9] doi:10.1186/s40168-026-02342-8 [code]
- Impacts of host genetics on gut microbiome composition in Alzheimer's disease.Journal: MicrobiomeIn common: metafor, cowplot, data.table, 2 other tools, Alzheimer's / dementia, genetics / omics, 4 references
- [10] doi:10.1038/s41467-026-74753-y [code]
- A human-specific microRNA controls the timing of excitatory synaptogenesis.Journal: Nature communicationsIn common: reticulate, easystats, emmeans, 6 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 6 scripts, and 3 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:fa7e3c934db6a700…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
