OSCR

Divergent patterns of cognitive decline in preclinical Alzheimer's disease: Implications for secondary prevention trials.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches
  1. [1] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 994–1088 · score 1.00 · comparatively younger amyloid, survivor bias, faster declining classes, smaller available sample, latent class severity, numerically younger
  2. [2] § METHODS › Statistical analysis › Latent class mixed‐effects models ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 432–530 · score 1.00 · stimulus version administered, generic assay unit, florbetapir cortical standardized, Absolute calibration, Bayesian Information Criterion, mass concentration
  3. [3] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 911–943 · score 0.99 · amyloid positron emission, selected baseline variables, represent quartiles, plasma phosphorylated tau, assay signal, volume normalized
  4. [4] § METHODS › Conduct of A4 and LEARN studies ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 214–265 · score 0.99 · fluid biomarker collection, elevated brain amyloid, safety monitoring, double blind, baseline clinical assessment, regular cognitive evaluations
  5. [5] § RESULTS › Longitudinal biomarker progression across latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 1090–1133 · score 0.99 · increasing amyloid burden, rapid biomarker progression, PET tau signals, compared longitudinal changes, temporal dissociation, steeper cognitive decline
  6. [6] § RESULTS › Longitudinal biomarker progression across latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 1135–1195 · score 0.99 · medial temporal lobe, clear biomarker progression, Longitudinal biomarker trajectories, linear mixed, natural cubic spline, random slopes
  7. [7] § RESULTS › Post hoc regression tree ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 1090–1133 · score 0.99 · aggregate multiple trees, tree algorithm selects, optimally tuned classification, hierarchical decision process, random forests, correlated variables
  8. [8] § METHODS › Statistical analysis › Latent class mixed‐effects models ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 432–530 · score 0.99 · caudal middle frontal, frontal pole, superior parietal, inferior parietal, tau pathology improved, entorhinal cortex
  9. [9] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 994–1088 · score 0.98 · positive coefficients indicate, negligible remaining, Higher PACC scores, 0.095–0.15, 0.24–0.35, 1.3–2.8
  10. [10] § METHODS › Conduct of A4 and LEARN studies ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 214–265 · score 0.93 · observational cohort enrolling, visit scheduling, quality control, harmonized procedures, identical clinical, amyloid PET positivity
  11. [11] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 817–903 · score 0.87 · confidence intervals, individual participant trajectories, latent class derived, Higher PACC scores, Preclinical Alzheimer Cognitive, indicate better cognitive
  12. [12] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 1383–1474 · score 0.77 · class membership submodels, odds ratios, tau PET model, standardized coefficients, Continuous predictors, baseline predictor
  13. [13] § RESULTS › Characterization and prediction of latent classes ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 775–797 · score 0.73 · CDR Global scores, Clinical Dementia Rating, CDR progression, slow decliners, fast decliners, zero
  14. [14] § METHODS › Statistical analysis › Ten‐fold cross‐validation of latent class predictions ↔ vignettes/articles/A4LEARN-Latent-Classes.qmd, lines 1878–1932 · score 0.57 · precision recall curve, fold cross validation, ROC, AUPRC, latent class, predict

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Quarto · 2,852 lines · 136 KB · no license · 14 matches

  1. ---
  2. title: "Divergent patterns of cognitive decline in preclinical Alzheimer's disease: implications for secondary prevention trials"
  3. author:
  4. - name: Runpeng Li
  5. email: [email hidden]
  6. affiliations: ATRI
  7. - name: Oliver Langford
  8. affiliations:
  9. - ref: ATRI
  10. - name: Philip S. Insel
  11. affiliations:
  12. - ref: UCSF
  13. - name: Reisa A. Sperling
  14. affiliations:
  15. - ref: MGH
  16. - name: Rema Raman
  17. affiliations:
  18. - ref: ATRI
  19. - name: Paul S. Aisen
  20. affiliations:
  21. - ref: ATRI
  22. - name: Michael C. Donohue
  23. email: [email hidden]
  24. correspondingauthor: true
  25. affiliations:
  26. - ref: ATRI
  27. affiliations:
  28. - id: ATRI
  29. name: USC Epstein Family Alzheimer's Therapeutic Research Institute, University of Southern California
  30. city: San Diego
  31. state: CA
  32. - id: UCSF
  33. name: Department of Psychiatry, University of California, San Francisco
  34. city: San Francisco
  35. state: CA
  36. - id: MGH
  37. name: Center for Alzheimer Research and Treatment, Brigham and Women's Hospital, Massachusetts General Hospital, Harvard Medical School
  38. city: Boston
  39. state: MA
  40. abstract: |
  41. **Introduction:** Biomarkers identify Alzheimer disease pathology in cognitively unimpaired adults, but the timing and rate of cognitive decline vary widely. This study aimed to identify subgroups of cognitive decline and baseline predictors of heterogeneity in preclinical progression.
  42. **Methods:** Data were drawn from the A4 Study, which enrolled amyloid positive participants, and the LEARN Study, which enrolled amyloid negative individuals. Latent class mixed effects models identified cognitive trajectory classes. Associations between class membership and demographic, clinical, and biomarker variables were evaluated. The primary outcome was change in the Preclinical Alzheimer Cognitive Composite.
  43. **Results:** Three trajectory classes were identified: stable, slow decliners, and fast decliners. Higher P-tau217, smaller hippocampal volume, and elevated tau PET were associated with declining classes. About 70 percent of amyloid positive individuals were stable.
  44. **Discussion:** Latent class modeling reveals substantial heterogeneity in preclinical trajectories with important implications for prevention trial design.
  45. keywords:
  46. - Latent Class Mixed Model
  47. - Preclinical Alzheimer's Disease
  48. - Cognitive decline prediction
  49. format:
  50. html:
  51. code-fold: true
  52. toc: true
  53. pdf:
  54. documentclass: article
  55. geometry:
  56. - margin=1in
  57. docx: default
  58. editor: source
  59. prefer-html: true
  60. always_allow_html: true
  61. date: "`r Sys.Date()`"
  62. bibliography: A4LEARN-Latent-Classes.bib
  63. csl: american-medical-association.csl
  64. nocite: |
  65. @petersen2016association, @Sperling2023, @donohue2017association, @parent2023longitudinal, @Sperling2024
  66. crossref:
  67. custom:
  68. - kind: float
  69. key: suppfig
  70. latex-env: suppfig
  71. reference-prefix: Supplementary Figure S
  72. space-before-numbering: false
  73. latex-list-of-description: Supplementary Figure
  74. - kind: float
  75. key: supptbl
  76. latex-env: supptbl
  77. reference-prefix: Supplementary Table S
  78. space-before-numbering: false
  79. latex-list-of-description: Supplementary Table
  80. ---
  81. # Introduction
  82. ```{r setup, include=FALSE}
  83. # rmarkdown::render('A4LEARN-Latent-Classes.qmd')
  84. UPDATELCMM <- FALSE
  85. UPDATELCMMCV <- FALSE
  86. UPDATETREECV <- FALSE
  87. library(A4LEARN)
  88. library(lcmm)
  89. library(emmeans)
  90. library(tidyverse)
  91. library(assertr)
  92. library(knitr)
  93. library(ggforce)
  94. library(patchwork)
  95. library(pROC)
  96. library(PRROC)
  97. library(ggsci)
  98. library(caret)
  99. library(arsenal)
  100. library(kableExtra)
  101. library(gridExtra)
  102. library(rpart)
  103. library(rpart.plot)
  104. library(tidymodels)
  105. library(vip)
  106. library(nlme)
  107. options(htmltools.dir.version = FALSE, knitr.kable.NA = '')
  108. opts_chunk$set(
  109. echo = FALSE,
  110. tidy=FALSE,
  111. collapse = TRUE,
  112. fig.path = '',
  113. fig.width = 6,
  114. fig.height = 4,
  115. fig.align = "center",
  116. message = FALSE,
  117. warning = FALSE,
  118. out.extra = '',
  119. out.width='\\linewidth',
  120. fig.align = 'center',
  121. crop = TRUE,
  122. fig.pos = '!h',
  123. cache=TRUE,
  124. comment = '',
  125. fig.dpi = 300)
  126. TX_level_sola <- A4LEARN::SUBJINFO %>%
  127. filter(stringr::str_detect(TX, "Sola")) %>% .$TX %>% unique()
  128. TX_level_placebo <- A4LEARN::SUBJINFO %>%
  129. filter(!stringr::str_detect(TX, "Sola")) %>% .$TX %>% unique()
  130. TX_levels <- c(TX_level_placebo, TX_level_sola)
  131. scale_vec <- function(x, ...){
  132. y <- scale(x, ...)
  133. sc <- attr(y, "scaled:center")
  134. ss <- attr(y, "scaled:scale")
  135. y <- y[,1]
  136. attr(y, "scaled:center") <- sc
  137. attr(y, "scaled:scale") <- ss
  138. y
  139. }
  140. ```
  141. ```{r ggplot-setup}
  142. theme_jama <- function(base_size = 10) {
  143. theme_bw(base_size = base_size) %+replace%
  144. theme(
  145. # Overall plot settings
  146. plot.title = element_text(size = base_size + 1, face = "bold",
  147. hjust = 0, margin = margin(b = 10)),
  148. plot.subtitle = element_text(size = base_size, hjust = 0,
  149. margin = margin(b = 10)),
  150. plot.caption = element_text(size = base_size - 1, hjust = 1,
  151. margin = margin(t = 10)),
  152. # Axis settings
  153. axis.title = element_text(size = base_size, face = "bold"),
  154. axis.text = element_text(size = base_size - 1, color = "black"),
  155. axis.line = element_line(color = "black", linewidth = 0.5),
  156. axis.ticks = element_line(color = "black", linewidth = 0.5),
  157. # Panel settings
  158. panel.background = element_rect(fill = "white", color = NA),
  159. panel.border = element_blank(),
  160. panel.grid.major = element_blank(),
  161. panel.grid.minor = element_blank(),
  162. # Legend settings
  163. legend.title = element_text(size = base_size, face = "bold"),
  164. legend.text = element_text(size = base_size - 1),
  165. legend.key = element_rect(fill = "white", color = NA),
  166. legend.background = element_rect(fill = "transparent", color = NA),
  167. legend.position = "bottom",
  168. # Strip settings (for facets)
  169. strip.background = element_blank(),
  170. strip.text = element_text(size = base_size - 2, face = "bold",
  171. vjust = 1, color = "black"),
  172. # Remove plot background
  173. plot.background = element_rect(fill = "white", color = NA)
  174. )
  175. }
  176. # JAMA-appropriate color palette (colorblind-friendly)
  177. jama_colors <- c("#374E55", "#DF8F44", "#00A1D5", "#B24745",
  178. "#79AF97", "#6A6599", "#80796B")
  179. # Scale functions for easy use
  180. scale_color_jama <- function(...) {
  181. scale_color_manual(values = jama_colors, ...)
  182. }
  183. scale_fill_jama <- function(...) {
  184. scale_fill_manual(values = jama_colors, ...)
  185. }
  186. # Example usage:
  187. # ggplot(data, aes(x = x, y = y, color = group)) +
  188. # geom_point() +
  189. # theme_jama() +
  190. # scale_color_jama()
  191. ```
  192. Elevated levels of brain amyloid in cognitively unimpaired older adults are associated with subsequent cognitive decline and increased risk of clinical progression.[@petersen2016association; @Sperling2023; @donohue2017association; @parent2023longitudinal; @Sperling2024] The Anti-Amyloid Treatment in Asymptomatic Alzheimer's Disease (A4) Study, a randomized trial of solanezumab in cognitively unimpaired individuals with elevated amyloid PET, demonstrated group-level decline on a cognitive composite but no treatment benefit.[@Sperling2023] Follow-up analyses from A4 and its companion observational study of amyloid-negative individuals, the Longitudinal Evaluation of Amyloid Risk and Neurodegeneration (LEARN), further confirmed that higher baseline amyloid PET and plasma phosphorylated tau (P‐tau217) levels are associated with faster cognitive decline and increased functional progression.[@Sperling2024]
  193. Despite these consistent group-level associations, cognitive trajectories among amyloid-positive cognitively unimpaired individuals remain highly variable. Conventional longitudinal models assume that individual cognitive trajectories are randomly scattered about a single mean trend for given covariate values, potentially obscuring meaningful heterogeneity. To better characterize this heterogeneity, we applied a Latent Class Mixed-Effects Model (LCMM) [@proust2017estimation] to A4 and LEARN data. This approach identifies unobserved subgroups of participants with distinct longitudinal patterns of change and allows for evaluation of baseline biomarkers and demographic factors as predictors of class membership.
  194. Latent class growth model approaches have been applied several times in the literature to identify distinct classes of late-life cognitive trajectories. A systematic review by Wu et al.[@wu2020distinct] identified 37 investigations and these consistently identified three classes: stable, slow decline, and fast decline. However, many of these papers lacked more recently developed biomarkers of Alzheimer's pathology. One exception is the study by Teipel et al.[@teipel2018effect] included amyloid PET as a predictor in a cohort of $N=265$ individuals followed for two years. Villeneuve et al.[@VILLENEUVE2019553] identified latent classes of Amsterdam Instrumental-Activities-of-Daily-Living questionnaire using amyloid PET and FDG PET in N=289 participants over three years. Here, we consider latent classes of cognition predicted by APOE genotype, amyloid PET, P‐tau217, hippocampal atrophy, and tau PET in a cohort of $N=1,629$ individuals followed for up to seven years in A4 and LEARN. Furthermore, we explore practical implications for classification algorithms and clinical trials.
  195. # Methods
  196. ## Conduct of the A4 and LEARN Studies
  197. The A4 and LEARN studies have been previously described[@Sperling2023; @Sperling2024], but we briefly summarize key elements. The A4 Study was a multicenter, randomized, double-blind, placebo-controlled secondary prevention trial enrolling cognitively unimpaired older adults aged 65–85 years with elevated brain amyloid on florbetapir PET. Participants were recruited from 67 sites in the United States, Canada, Australia, and Japan and underwent standardized screening, baseline clinical assessment, biomarker acquisition, and neuropsychological testing. After eligibility confirmation, participants were randomized to receive solanezumab or placebo and were assessed longitudinally with regular cognitive evaluations, safety monitoring, and follow-up imaging and fluid biomarker collection per protocol.
  198. The LEARN study was conducted in parallel as an observational cohort enrolling individuals who met all A4 screening criteria except for amyloid PET positivity. LEARN participants completed identical clinical and cognitive assessments, allowing direct comparison of trajectories between amyloid-positive and amyloid-negative individuals. Both studies followed harmonized procedures for data collection, visit scheduling, and quality control to ensure comparability of longitudinal outcomes.
  199. All participants of both studies provided written informed consent, and study protocols were approved by the institutional review boards or ethics committees at each participating site.
  200. ```{r tau-mri-data-prep}
  201. tau_pet_data <- A4LEARN::imaging_Tau_PET_Stanford %>%
  202. select(BID, bi_entorhinal, bi_inferiortemporal, bi_inferiorparietal,
  203. bi_posteriorcingulate, bi_caudalmiddlefrontal,
  204. bi_middletemporal, bi_superiorparietal, bi_frontalpole) %>%
  205. assertr::verify(!duplicated(BID)) %>%
  206. rowwise() %>%
  207. mutate(Tau_PET = mean(c(bi_entorhinal, bi_inferiortemporal,
  208. bi_inferiorparietal, bi_posteriorcingulate, bi_caudalmiddlefrontal,
  209. bi_middletemporal, bi_superiorparietal, bi_frontalpole), na.rm=TRUE)) %>%
  210. ungroup()
  211. mri_data <- A4LEARN::imaging_volumetric_mri %>%
  212. filter((VISCODE == 4 | is.na(VISCODE))) %>%
  213. filter(!is.na(RightHippocampus)) %>%
  214. filter(!duplicated(BID, fromLast = TRUE)) %>%
  215. select(BID, VISCODE, SUBSTUDY,
  216. LeftEntorhinal, RightEntorhinal,
  217. LeftHippocampus, RightHippocampus, IntraCranialVolume) %>%
  218. filter(LeftEntorhinal != -4, RightEntorhinal != -4,
  219. LeftHippocampus != -4, RightHippocampus != -4) %>%
  220. mutate(
  221. Entorhinal = LeftEntorhinal + RightEntorhinal,
  222. Hippocampus = LeftHippocampus + RightHippocampus)
  223. # "Residualize" hippocampus relative to ICV
  224. # Y ~ int + scale(ICV) * beta + e
  225. # Y - scale(ICV)*beta
  226. # int + e
  227. mri_icv_mod <- lm(Hippocampus ~ scale(IntraCranialVolume), data = mri_data)
  228. mri_data$Hipp_res <- mri_icv_mod$coefficients[["(Intercept)"]] +
  229. mri_icv_mod$residuals
  230. mri_data$Hipp_atrophy <- scale(-mri_data$Hipp_res)[,1]
  231. # with(mri_data, plot(Hippocampus, Hipp_res))
  232. # with(mri_data, plot(Hippocampus, Hipp_atrophy))
  233. ```
  234. ```{r pacc-prep-data, include = FALSE}
  235. ## A4 data prep ----
  236. qs_pacc <- A4LEARN::ADQS %>% filter(QSTESTCD %in% "PACC") %>%
  237. left_join(A4LEARN::ptdemog %>%
  238. select(BID, Sex = PTGENDER, PTRACE, PTETHNIC), by='BID') %>%
  239. left_join(A4LEARN::cdr %>%
  240. select(BID, MEMORY, CDOLEEVENT) %>%
  241. filter(!is.na(CDOLEEVENT)), by = 'BID') %>%
  242. mutate(Y = QSSTRESN, Y_ch = QSCHANGE, Y_bl = QSBLRES,
  243. QSVERSION = as.factor(QSVERSION),
  244. Group = case_when(
  245. SUBSTUDY == 'LEARN' ~ 'LEARN',
  246. TRUE ~ as.character(TX)) %>%
  247. factor(levels = c('LEARN', as.character(TX_levels))),
  248. APOEe4 = factor(AAPOEGNPRSNFLG, levels = 0:1,
  249. labels = c('non-carrier', 'carrier')),
  250. Female = case_when(
  251. Sex == 'Female' ~ 1,
  252. TRUE ~ 0),
  253. Race = factor(PTRACE,
  254. levels = 1:6,
  255. labels = c(
  256. "Am. Indian or Alaska Native",
  257. "Asian",
  258. "Native Hawaiian or Other PI",
  259. "Black or African Am.",
  260. "White",
  261. "Unknown or Not Reported")),
  262. Ethnicity = PTETHNIC,
  263. Sola = case_when(
  264. TX == "Solanezumab" ~ 1,
  265. TRUE ~ 0),
  266. `CDR Progressor` = case_when(
  267. CDOLEEVENT == 1 ~ 'CDR Progressor',
  268. TRUE ~ 'CDR Non-progressor') %>%
  269. factor(levels = c('CDR Non-progressor', 'CDR Progressor')),
  270. CDMEM = cut(MEMORY, breaks = c(-1, 0, 2), labels = c("0", ">0"))) %>%
  271. left_join(A4LEARN::biomarker_pTau217 %>%
  272. filter(TESTCD == 'PTAU217' & !is.na(ORRESRAW)) %>%
  273. arrange(BID, VISCODE) %>%
  274. filter(!duplicated(BID)) %>%
  275. select(BID, Ptau217 = ORRESRAW), by = 'BID') %>%
  276. left_join(tau_pet_data, by = 'BID') %>%
  277. left_join(mri_data %>%
  278. select(BID, IntraCranialVolume, Entorhinal, Hippocampus, Hipp_atrophy),
  279. by = 'BID') %>%
  280. arrange(BID, ADURW) %>%
  281. select(BID, SUBSTUDY, EPOCH, VISITCD, ASEQNCS, ASEQMMRM, AVISIT, ADURW, AMYLCENT,
  282. Y, Y_ch, Y_bl, TX, Group, AGEYR, Sex, Female, Race, Ethnicity, EDCCNTU,
  283. QSVERSION, AAPOEGNPRSNFLG, APOEe4, SUVRCER, Ptau217, Sola, Tau_PET,
  284. IntraCranialVolume, Entorhinal, Hippocampus, Hipp_atrophy,
  285. MEMORY, CDMEM, `CDR Progressor`) %>%
  286. ungroup() %>%
  287. mutate(id = as.numeric(as.factor(BID))) %>%
  288. filter(!is.na(Y), !is.na(ADURW), !is.na(Ptau217),
  289. !is.na(APOEe4), !is.na(Hipp_atrophy))
  290. # PACC scores can be negative. Do a shift before boxcox
  291. min_Y <- min(qs_pacc$Y, na.rm = TRUE)
  292. shift_constant <- abs(min_Y) + 1
  293. qs_pacc <- qs_pacc %>%
  294. mutate(Y_shifted = Y + shift_constant)
  295. boxcox_result <- MASS::boxcox(Y_shifted ~ 1, data = qs_pacc, plotit = TRUE)
  296. lambda_optimal <- boxcox_result$x[which.max(boxcox_result$y)]
  297. boxcox_f <- function(y) ((y+shift_constant)^lambda_optimal - 1) / lambda_optimal
  298. qs_pacc <- qs_pacc %>%
  299. mutate(Y_boxcox = (Y_shifted^lambda_optimal - 1) / lambda_optimal)
  300. reverse_boxcox <- function(pred_boxcox, lambda=lambda_optimal,
  301. shift = shift_constant, center = 0, scale = 1) {
  302. pred_boxcox <- (pred_boxcox * scale) + center
  303. if (lambda == 0) {
  304. return(exp(pred_boxcox) - shift)
  305. } else {
  306. return((pred_boxcox * lambda + 1)^(1 / lambda) - shift)
  307. }
  308. }
  309. if(min(abs(qs_pacc$Y - reverse_boxcox(qs_pacc$Y_boxcox)))>0){
  310. stop("reverse_boxcox not working properly")
  311. }
  312. qs_pacc <- qs_pacc %>%
  313. mutate(
  314. Z = scale_vec(Y_boxcox),
  315. AGEYR_z = scale_vec(AGEYR),
  316. EDCCNTU_z = scale_vec(EDCCNTU),
  317. Hipp_atrophy_z = scale_vec(Hipp_atrophy),
  318. Ptau217_z = scale_vec(Ptau217),
  319. SUVRCER_z = scale_vec(SUVRCER),
  320. Tau_PET_z = scale_vec(Tau_PET))
  321. if(min(abs(qs_pacc$Y - reverse_boxcox(qs_pacc$Z,
  322. center = attr(qs_pacc$Z, "scaled:center"),
  323. scale = attr(qs_pacc$Z, "scaled:scale"))))>0)
  324. {
  325. stop("reverse_boxcox not working properly")
  326. }
  327. base_model_data <- qs_pacc %>%
  328. select(id, Z, ADURW, Female, AAPOEGNPRSNFLG, QSVERSION, Sola,
  329. AGEYR_z, EDCCNTU_z, SUVRCER_z, Ptau217_z, Hipp_atrophy_z)
  330. if(nrow(base_model_data) != nrow(base_model_data %>% na.omit()))
  331. stop("Unexpected missing data")
  332. tau_model_data <- qs_pacc %>%
  333. select(id, Y_boxcox, ADURW, Female, AAPOEGNPRSNFLG, QSVERSION, Sola,
  334. AGEYR, EDCCNTU, SUVRCER, Ptau217, Hipp_atrophy, Tau_PET) %>%
  335. na.omit() %>%
  336. mutate(
  337. Z = scale(Y_boxcox),
  338. AGEYR_z = scale(AGEYR),
  339. EDCCNTU_z = scale(EDCCNTU),
  340. Hipp_atrophy_z = scale(Hipp_atrophy),
  341. Ptau217_z = scale(Ptau217),
  342. SUVRCER_z = scale(SUVRCER),
  343. Tau_PET_z = scale(Tau_PET)) %>%
  344. select(id, Z, ADURW, Female, AAPOEGNPRSNFLG, QSVERSION, Sola,
  345. AGEYR_z, EDCCNTU_z, SUVRCER_z, Ptau217_z, Hipp_atrophy_z, Tau_PET_z)
  346. timeby_cont <- seq(0, max(qs_pacc$ADURW), by=24)
  347. timeby_cont_plot <- seq(0, max(qs_pacc$ADURW), by=8)
  348. xlimit_cont <- c(min(timeby_cont) - 5, max(timeby_cont) + 5)
  349. ```
  350. ```{r spline_functions, warning=FALSE}
  351. ## spline functions 2df ----
  352. ns21 <- function(t){
  353. as.numeric(predict(splines::ns(qs_pacc$ADURW, df=2,
  354. Boundary.knots = c(0, max(qs_pacc$ADURW))), t)[,1])
  355. }
  356. ns22 <- function(t){
  357. as.numeric(predict(splines::ns(qs_pacc$ADURW, df=2,
  358. Boundary.knots = c(0, max(qs_pacc$ADURW))), t)[,2])
  359. }
  360. # required for mclapply
  361. assign("ns21", ns21, envir = .GlobalEnv)
  362. assign("ns22", ns22, envir = .GlobalEnv)
  363. ## spline functions 3df ----
  364. ns31 <- function(t){
  365. as.numeric(predict(splines::ns(qs_pacc$ADURW, df=3,
  366. Boundary.knots = c(0, max(qs_pacc$ADURW))), t)[,1])
  367. }
  368. ns32 <- function(t){
  369. as.numeric(predict(splines::ns(qs_pacc$ADURW, df=3,
  370. Boundary.knots = c(0, max(qs_pacc$ADURW))), t)[,2])
  371. }
  372. ns33 <- function(t){
  373. as.numeric(predict(splines::ns(qs_pacc$ADURW, df=3,
  374. Boundary.knots = c(0, max(qs_pacc$ADURW))), t)[,3])
  375. }
  376. # required for mclapply
  377. assign("ns31", ns21, envir = .GlobalEnv)
  378. assign("ns32", ns22, envir = .GlobalEnv)
  379. assign("ns33", ns22, envir = .GlobalEnv)
  380. ```
  381. ## Statistical Analysis
  382. ### Latent Class Mixed-Effects Models
  383. Latent class mixed-effects models (LCMMs) extend traditional mixed-effects models by identifying unobserved subgroups, or latent classes, of individuals who follow distinct longitudinal trajectories. In the context of cognitive decline in Alzheimer's disease, LCMMs allow for the estimation of population-level trends while capturing individual variability and uncovering hidden subpopulations that may progress at different rates.
  384. Analyses were conducted using harmonized data from the A4 and LEARN studies, accessed through the `A4LEARN` R data package (version 1.1.20250808 [@a4studydata; @donohue2025alzheimer]). This package provides curated datasets and metadata derived from the A4 and LEARN clinical studies for reproducible statistical analysis. Code to reproduce the results from this paper are available from the `A4LEARN` GitHub repository [@a4learn_vignettes].
  385. The dependent variable was the Box-Cox-transformed Preclinical Alzheimer Cognitive Composite (PACC)[@Donohue2014] score at each visit. The Box-Cox transformation parameter ($\lambda$) was selected to maximize the likelihood that the transformed data approximated normality. Time was modeled as a continuous variable representing years since baseline.
  386. Natural cubic spline basis functions with one, two or three degrees of freedom were evaluated to model nonlinear trajectories, with boundary knots at baseline and the maximum follow-up time and an interior knots at the median or tertiles of observation time. The choice of degrees of freedom and the number of classes (also one, two, or three) was determined by model fit criteria (Bayesian Information Criterion \[BIC\] and Integrated Complete Likelihood \[ICL\]).[@biernacki2002assessing] Each model included class-specific spline-based time effects, as well as subject-specific random intercepts. The effect of baseline covariates were shared across latent classes and included: randomized to solanezumab (1 for the active group, 0 for placebo or LEARN), plasma P-tau217 concentration, florbetapir cortical standardized uptake value ratio (SUVr), APOE $\epsilon 4$ carrier status, sex, age, education, hippocampal atrophy, and PACC test stimulus version administered (which alternated per study protocol). Plasma P-tau217 concentrations were measured using an immunoassay developed by Eli Lilly and reported in arbitrary units per milliliter (U/mL), where "U" denotes a generic assay unit proportional to signal intensity. Absolute calibration against mass concentration (e.g., pg/mL) was not available. Hippocampal atrophy measures were derived by first residualizing with respect to intracranial volume, and then standardizing to mean zero and variance one (i.e. z-scoring).
  387. The LCMM includes two linked submodels. The longitudinal submodel captures within-class trajectories of cognitive change over time. Each class had its own set of spline coefficients, allowing class-specific shapes of decline. The class membership submodel defines the probability of belonging to each latent class as a multinomial logistic function of baseline covariates (which includes all covariates listed above except spline terms and PACC test version).
  388. In a secondary analysis, we fit the model to the subset of participants with baseline flortaucipir (tau) PET imaging. Tau PET was summarized as the mean SUVr across eight cortical regions: entorhinal cortex, inferior temporal, inferior parietal, posterior cingulate, caudal middle frontal, middle temporal, superior parietal, and frontal pole. This composite measure was included as an additional covariate in both the longitudinal and class membership submodels to assess whether tau pathology improved prediction of cognitive decline patterns or latent class assignment.
  389. To assess the reliability and potential predictive utility of the latent class membership model, we examined the posterior probabilities of class assignment for each participant. High posterior probabilities indicate greater certainty in latent class classification. We summarized these distributions across classes and computed the mean posterior probability within each group as an index of classification confidence.
  390. ### Power Analysis Stratified by Latent Class
  391. The design of preclinical Alzheimer’s clinical trials have typically assumed a homogeneous population. To explore the impact of heterogeneity on clinical trials in preclinical Alzheimer’s disease, we conducted power calculations to estimate power for trials used pilot parameters from different latent classes of cognitive decline among amyloid-positive participants. We emphasize this analysis is not intended to support class‑specific or stratified trials. Instead, the intent is to clarify how latent trajectory classes, if present in a trial population, may impact statistical power, sample‑size requirements, and interpretation of treatment effects.
  392. Effect size estimates were derived from longitudinal models using natural cubic splines applied to Preclinical Alzheimer Cognitive Composite (PACC) scores.[@donohue2023natural] Separate models were fit for (a) the stable class, (b) the decliner classes, and (c) a reference group of amyloid-negative stable individuals. The models included fixed effects for time (two degrees of freedom), APOE $\epsilon 4$ carrier status, sex, age, education, plasma P-tau217, and florbetapir PET. Residuals were assumed to be normally distributed with an unstructured covariance matrix.
  393. For each group, we extracted mean PACC and residual variance from the model at two time points: 2 years and 4 years, representing potential trial durations. These estimates were used to compute the power to detect treatment effects. Power was approximated using two-sample t-test calculations, assuming 500 participants per arm, attrition rates of 10% at 2 years and 20% at 4 years, and a two-sided alpha of 0.05.
  394. ### Ten-fold cross-validation of latent class predictions
  395. To further evaluate prospective discrimination, we performed ten-fold cross-validation stratified by P-tau217 and latent class. For each of the ten folds, 90% of the data were used to re-train the LCMM, and the re-trained model was used to predict latent classes for the 10% held-out test set using only baseline data. Model performance was summarized using the area under the precision-recall curve (AUPRC) for discrimination of each class versus all others. This approach can be more informative than receiver operating characteristic (ROC) curves when group sizes are imbalance.[@saito2015precision]
  396. ### Regression Tree Analysis for Latent Class Characterization
  397. To evaluate whether latent classes could be distinguished by baseline characteristics through a combination of binary decision rules, we conducted a post-hoc classification tree analysis [@rpart; @breiman2017classification]. Class assignments from the optimal latent class mixed model were used as the outcome variable, with baseline demographic and clinical characteristics as predictors (randomized to solanezumab, plasma P-tau217 concentration, florbetapir PET, APOE $\epsilon 4$ carrier status, sex, age, education, hippocampal atrophy, and PACC). We separately considered a regression tree with tau PET as an additional predictor.
  398. Hyperparameter tuning was performed using 10-fold cross-validation. The optimal hyperparameters were selected based on maximum cross-validated accuracy. The final tree structure was visualized to illustrate the hierarchy of decision rules for class assignment. Variable importance was calculated based on the total reduction in node impurity attributed to splits on each predictor. This approach provides an interpretable framework for understanding the multivariate profile of each latent class and assessing whether classes can be reliably separated using clinically available baseline information.
  399. All analyses were conducted using R (version 4.5.2; R Foundation for Statistical Computing) with the `lcmm` package (version 2.2.1).
  400. # Results
  401. ## Study Participants
  402. A total of $N=1,629$ participants were included in the analytic sample for the base model (without tau PET), comprising $N=1,110$ from the A4 Study and $N=519$ from the LEARN study. The mean (standard deviation \[SD\]) age at baseline was 71.46 (4.69) years, and 60.2% were female. The median follow-up time was 6.0 years (interquartile range 3.9 to 7.0). Baseline characteristics stratified by latent class membership are presented in @tbl-baseline-base-characteristics. A total of $N=427$ individuals had tau PET data and were submitted to the tau PET model. The tau PET subset included $N=372$ from the A4 Study and $N=55$ from the LEARN study. Baseline characteristics for the tau PET subset are presented in @supptbl-baseline-tau-characteristics.
  403. ```{r lcmm-mri, eval = UPDATELCMM}
  404. ## 2 df spline ----
  405. m1 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  406. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  407. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  408. random = ~ 1, subject = 'id', ng = 1,
  409. data = base_model_data)
  410. m2 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  411. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  412. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  413. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  414. random = ~ 1, subject = 'id', ng = 2, B = m1,
  415. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  416. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  417. data = base_model_data)
  418. m3 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  419. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  420. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  421. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  422. random = ~ 1, subject = 'id', ng = 3, B = m1,
  423. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  424. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  425. data = base_model_data)
  426. ## 3 df spline ----
  427. m31 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  428. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  429. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  430. random = ~ 1, subject = 'id', ng = 1,
  431. data = base_model_data)
  432. m32 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  433. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  434. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  435. mixture = ~ (I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW))),
  436. random = ~ 1, subject = 'id', ng = 2, B = m31,
  437. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  438. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  439. data = base_model_data)
  440. m33 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  441. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  442. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  443. mixture = ~ (I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW))),
  444. random = ~ 1, subject = 'id', ng = 3, B = m31,
  445. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  446. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  447. data = base_model_data)
  448. ## save ----
  449. save(m1, m2, m3, m31, m32, m33,
  450. file='a4learn-pacc-lcmm-base.rdata')
  451. ```
  452. ```{r lcmm-tau, eval = UPDATELCMM}
  453. ## 2 df spline ----
  454. mt1 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  455. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  456. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  457. random = ~ 1, subject = 'id', ng = 1,
  458. data = tau_model_data)
  459. mt2 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  460. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  461. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  462. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  463. random = ~ 1, subject = 'id', ng = 2, B = mt1,
  464. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  465. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  466. data = tau_model_data)
  467. mt3 <- hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  468. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  469. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  470. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  471. random = ~ 1, subject = 'id', ng = 3, B = mt1,
  472. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  473. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  474. data = tau_model_data)
  475. ## 3 df spline ----
  476. mt31 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  477. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  478. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  479. random = ~ 1, subject = 'id', ng = 1,
  480. data = tau_model_data)
  481. mt32 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  482. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  483. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  484. mixture = ~ (I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW))),
  485. random = ~ 1, subject = 'id', ng = 2, B = mt31,
  486. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  487. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  488. data = tau_model_data)
  489. mt33 <- hlme(Z ~ I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW)) +
  490. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  491. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  492. mixture = ~ (I(ns31(ADURW)) + I(ns32(ADURW)) + I(ns33(ADURW))),
  493. random = ~ 1, subject = 'id', ng = 3, B = mt31,
  494. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  495. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  496. data = tau_model_data)
  497. ## save ----
  498. save(mt1, mt2, mt3, mt31, mt32, mt33,
  499. file='a4learn-pacc-lcmm-tau.rdata')
  500. ```
  501. ```{r load-lcmm-results, include = FALSE, eval = UPDATELCMM}
  502. load('a4learn-pacc-lcmm-base.rdata')
  503. summarytable(m1, m2, m3, m31, m32, m33,
  504. which = c("G", "loglik", "conv", "AIC", "BIC", "entropy", "ICL", "%class"))
  505. summaryplot(m1, m2, m3, which = "BIC")
  506. summaryplot(m31, m32, m33, which = "BIC")
  507. # make stable last so that it is the reference group
  508. plot(predictY(m3, newdata = base_model_data,
  509. var.time="ADURW"),legend.loc="right",bty="l")
  510. m.best.mri <- permut(m3, order=c(2,1,3))
  511. plot(predictY(m.best.mri, newdata = base_model_data,
  512. var.time="ADURW"),legend.loc="right",bty="l")
  513. load('a4learn-pacc-lcmm-tau.rdata')
  514. summarytable(mt1, mt2, mt3, mt31, mt32, mt33,
  515. which = c("G", "loglik", "conv", "AIC", "BIC", "entropy", "ICL", "%class"))
  516. summaryplot(mt1, mt2, mt3, which = "BIC")
  517. summaryplot(mt31, mt32, mt33, which = "BIC")
  518. # make stable last so that it is the reference group
  519. plot(predictY(mt3, newdata = tau_model_data,
  520. var.time="ADURW"),legend.loc="right",bty="l")
  521. m.best.tau <- permut(mt3, order=c(3,1,2))
  522. plot(predictY(m.best.tau, newdata = tau_model_data,
  523. var.time="ADURW"),legend.loc="right",bty="l")
  524. save(m.best.mri, m.best.tau, file = "a4learn-pacc-lcmm-best-models.rdata")
  525. ```
  526. ```{r attach-class-labels}
  527. load('a4learn-pacc-lcmm-base.rdata')
  528. load('a4learn-pacc-lcmm-tau.rdata')
  529. load("a4learn-pacc-lcmm-best-models.rdata")
  530. qs_pacc <- qs_pacc %>%
  531. left_join(m.best.mri$pprob %>% as_tibble() %>% select(id, class),
  532. by = 'id') %>%
  533. mutate(`Class (Base)` =
  534. case_when(
  535. class == 1 ~ 'Slow decliner',
  536. class == 2 ~ 'Fast decliner',
  537. class == 3 ~ 'Stable',) %>%
  538. factor(levels = c('Stable', 'Slow decliner', 'Fast decliner'))) %>%
  539. select(-class) %>%
  540. left_join(m.best.tau$pprob %>% as_tibble() %>% select(id, class),
  541. by = 'id') %>%
  542. mutate(`Class (Tau PET)` =
  543. case_when(
  544. class == 1 ~ 'Slow decliner',
  545. class == 2 ~ 'Fast decliner',
  546. class == 3 ~ 'Stable',) %>%
  547. factor(levels = c('Stable', 'Slow decliner', 'Fast decliner'))) %>%
  548. select(-class)
  549. ```
  550. ```{r long-biomarkers-prep}
  551. amypet <- qs_pacc %>%
  552. arrange(BID, ADURW) %>%
  553. filter(!duplicated(BID)) %>%
  554. right_join(A4LEARN::imaging_SUVR_amyloid %>%
  555. filter(brain_region == 'Composite_Summary') %>%
  556. mutate(Centiloid = 183.07 * suvr_cer - 177.26) %>%
  557. dplyr::select(-SUBSTUDY), by='BID') %>%
  558. left_join(A4LEARN::ADQS %>%
  559. dplyr::select(BID, VISCODE=VISITCD, QSDTC_DAYS_T0) %>%
  560. mutate(VISCODE = as.numeric(VISCODE)) %>%
  561. dplyr::filter(!duplicated(paste(BID, VISCODE))),
  562. by = c('BID', 'VISCODE')) %>%
  563. mutate(scan_date_DAYS_T0 = case_when(
  564. is.na(scan_date_DAYS_T0) ~ QSDTC_DAYS_T0,
  565. TRUE ~ scan_date_DAYS_T0
  566. )) %>% dplyr::select(-QSDTC_DAYS_T0) %>%
  567. mutate(Weeks = scan_date_DAYS_T0/7)
  568. ptau217 <- qs_pacc %>%
  569. arrange(BID, ADURW) %>%
  570. filter(!duplicated(BID)) %>%
  571. right_join(A4LEARN::biomarker_pTau217 %>%
  572. dplyr::select(-SUBSTUDY), by='BID') %>%
  573. left_join(A4LEARN::ADQS %>%
  574. dplyr::select(BID, VISCODE=VISITCD, QSDTC_DAYS_T0) %>%
  575. mutate(VISCODE = as.numeric(VISCODE)) %>%
  576. dplyr::filter(!duplicated(paste(BID, VISCODE))),
  577. by = c('BID', 'VISCODE')) %>%
  578. mutate(COLLECTION_DATE_DAYS_T0 = case_when(
  579. is.na(COLLECTION_DATE_DAYS_T0) ~ QSDTC_DAYS_T0,
  580. TRUE ~ COLLECTION_DATE_DAYS_T0
  581. )) %>% dplyr::select(-QSDTC_DAYS_T0) %>%
  582. mutate(Weeks = COLLECTION_DATE_DAYS_T0/7)
  583. dd_flortaucipir <- A4LEARN::imaging_SUVR_tau %>%
  584. filter(str_detect(brain_region, "VOI")) %>%
  585. dplyr::select(BID, scan_date_DAYS_T0, VISCODE, brain_region, suvr_crus) %>%
  586. left_join(A4LEARN::SV %>%
  587. dplyr::select(BID, VISCODE=VISITCD, SVSTDTC_DAYS_T0) %>%
  588. mutate(VISCODE = as.numeric(VISCODE)) %>%
  589. dplyr::filter(!duplicated(paste(BID, VISCODE))),
  590. by = c('BID', 'VISCODE')) %>%
  591. mutate(scan_date_DAYS_T0 = case_when(
  592. is.na(scan_date_DAYS_T0) ~ SVSTDTC_DAYS_T0,
  593. TRUE ~ scan_date_DAYS_T0
  594. )) %>% dplyr::select(-SVSTDTC_DAYS_T0) %>%
  595. pivot_wider(names_from = brain_region, values_from = suvr_crus) %>%
  596. mutate(
  597. # Early (MTL)
  598. MTL = (Amygdala_lh_VOI * 258 + Amygdala_rh_VOI * 272 +
  599. entorhinal_lh_VOI * 256 + entorhinal_rh_VOI * 241 +
  600. parahippocampal_lh_VOI * 328 + parahippocampal_rh_VOI * 316)
  601. / 1671,
  602. # Middle (Neocortical)
  603. Neocortical = (fusiform_lh_VOI * 1495 + fusiform_rh_VOI * 1487 +
  604. inferiortemporal_lh_VOI * 1771 + inferiortemporal_rh_VOI * 1715 +
  605. middletemporal_lh_VOI * 1737 + middletemporal_rh_VOI * 1816 +
  606. inferiorparietal_lh_VOI * 1927 + inferiorparietal_rh_VOI * 2214)
  607. / 14162) %>%
  608. pivot_longer(MTL:Neocortical, values_to = "VALUE", names_to = "PARAMETER") %>%
  609. mutate(VALUE = round(VALUE, digits = 2))
  610. taupet <- qs_pacc %>%
  611. arrange(BID, ADURW) %>%
  612. filter(!duplicated(BID)) %>%
  613. right_join(dd_flortaucipir, by='BID') %>%
  614. mutate(Weeks = scan_date_DAYS_T0/7)
  615. mri_data_l <- A4LEARN::imaging_volumetric_mri %>%
  616. filter(!is.na(RightHippocampus)) %>%
  617. dplyr::select(BID, Date_DAYS_T0, VISCODE, SUBSTUDY,
  618. LeftEntorhinal, RightEntorhinal,
  619. LeftHippocampus, RightHippocampus, IntraCranialVolume) %>%
  620. filter(LeftEntorhinal != -4, RightEntorhinal != -4,
  621. LeftHippocampus != -4, RightHippocampus != -4) %>%
  622. mutate(
  623. Entorhinal = LeftEntorhinal + RightEntorhinal,
  624. Hippocampus = LeftHippocampus + RightHippocampus) %>%
  625. left_join(qs_pacc %>%
  626. dplyr::select(BID, Group, `Class (Base)`, `Class (Tau PET)`) %>%
  627. distinct(.keep_all = TRUE), by = 'BID') %>%
  628. filter(!is.na(Group)) %>%
  629. mutate(Weeks = Date_DAYS_T0/7)
  630. mri_icv_mod_l <- lme(Hippocampus ~ scale(IntraCranialVolume) + Weeks,
  631. random = ~1 | BID, data = mri_data_l)
  632. mri_data_l$Y <- mri_data_l$Hippocampus -
  633. mri_icv_mod$coefficients[['scale(IntraCranialVolume)']] *
  634. scale(mri_data_l$IntraCranialVolume)
  635. ```
  636. ```{r long-biomarkers}
  637. long_biomarker <- amypet %>%
  638. filter(!is.na(Group), !is.na(Weeks), !is.na(Centiloid)) %>%
  639. select(BID, Group, `Class (Base)`, Weeks, Y = Centiloid) %>%
  640. mutate(Outcome = 'Amyloid PET (CL)') %>%
  641. bind_rows(ptau217 %>%
  642. filter(!is.na(Group), !is.na(ORRESRAW)) %>%
  643. select(BID, Group, `Class (Base)`, Weeks, Y = ORRESRAW) %>%
  644. mutate(Outcome = 'P-tau217 (U/ml)')) %>%
  645. bind_rows(taupet %>%
  646. filter(!is.na(Group), !is.na(VALUE), PARAMETER == 'MTL') %>%
  647. select(BID, Group, `Class (Base)`, Weeks, Y = VALUE) %>%
  648. mutate(Outcome = 'MTL tau (SUVr)')) %>%
  649. bind_rows(taupet %>%
  650. filter(!is.na(Group), !is.na(VALUE), PARAMETER == 'Neocortical') %>%
  651. select(BID, Group, `Class (Base)`, Weeks, Y = VALUE) %>%
  652. mutate(Outcome = 'Neocortical tau (SUVr)')) %>%
  653. bind_rows(mri_data_l %>%
  654. as_tibble() %>%
  655. filter(!is.na(Group), !is.na(Weeks), !is.na(Y)) %>%
  656. select(BID, Group, `Class (Base)`, Weeks, Y) %>%
  657. mutate(Outcome = 'Hippocampus (cc)')) %>%
  658. mutate(
  659. Group = case_when(
  660. Group == 'LEARN' & `Class (Base)` == "Stable" ~
  661. 'A\u03B2- stable',
  662. `Class (Base)` == "Stable" ~ 'A\u03B2+ stable',
  663. TRUE ~ `Class (Base)`) %>%
  664. factor(levels = c('Aβ- stable', 'Aβ+ stable',
  665. 'Slow decliner', 'Fast decliner')),
  666. Outcome = factor(Outcome, levels = c("Amyloid PET (CL)",
  667. "P-tau217 (U/ml)", "MTL tau (SUVr)", "Neocortical tau (SUVr)",
  668. "Hippocampus (cc)")))
  669. ```
  670. ```{r tbl-baseline-base-characteristics}
  671. #| results: asis
  672. #| tbl-colwidths: [30,15,15,15,15,10]
  673. #| tbl-cap: "**Baseline Demographic, Clinical, and Biomarker Characteristics by Latent Class of Cognitive Decline.** Baseline characteristics of participants classified by the latent class mixed model (LCMM) of Preclinical Alzheimer Cognitive Composite (PACC) trajectories. Classes represent distinct longitudinal cognitive patterns: stable, slow decliner, and fast decliner. Continuous variables are presented as mean (SD); categorical variables as No. (%). **Abbreviations:** APOE = apolipoprotein E; PET = positron emission tomography; P-tau217 = plasma phosphorylated tau 217; SUVr = standardized uptake value ratio; U/mL = arbitrary units proportional to assay signal; Am. = American; PI = Pacific Islander; CDR = Clinical Dementia Rating; Prog. = Progressor. **Footnotes:** Amyloid PET and tau PET SUVr values represent mean cortical uptake relative to cerebellar reference region. Hippocampal atrophy are residualized for intracranial volume and z-scored. Education reported in years of formal schooling. CDR Progressors were observed to have a CDR Global score greater than zero at two consecutive visits, or their last visit."
  674. mylabels <- list(Y = "PACC, mean (SD)", AMYLCENT = "Amyloid PET, mean (SD), CL",
  675. EDCCNTU = 'Education, mean (SD), y', TX = 'Group, No. (%)',
  676. AGEYR = 'Age, mean (SD), y', Ptau217 = "P-tau217, mean (SD), U/ml",
  677. Tau_PET = "Tau PET, mean (SD), SUVr",
  678. Hipp_atrophy_z = "Hipp. atrophy, mean (SD), z-score", Sex = "Sex, No. (%)",
  679. Race = "Race, No. (%)", Ethnicity = "Ethnicity, No. (%)",
  680. APOEe4 = "APOEε4, No. (%)", 'CDR Progressor' = "CDR Prog., No. (%)")
  681. tableby(`Class (Base)` ~ Group + AGEYR + Sex + Race + Ethnicity + EDCCNTU +
  682. APOEe4 + Y + AMYLCENT + Ptau217 + Hipp_atrophy_z + Tau_PET +
  683. `CDR Progressor`,
  684. data = qs_pacc %>% filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)),
  685. digits = 2, test.always = TRUE, numeric.stats = c("Nmiss", "meansd"),
  686. numeric.simplify = TRUE, cat.simplify = FALSE) %>%
  687. summary(labelTranslations = mylabels, stats.labels = list(Nmiss = 'N missing'), text = TRUE) %>%
  688. as.data.frame() %>%
  689. kbl("pipe", booktabs = T)
  690. ```
  691. ```{r followup-summary, include = FALSE}
  692. base_model_data %>%
  693. arrange(id, desc(ADURW)) %>%
  694. filter(!duplicated(id)) %>% mutate(Years = ADURW * 7 / 365.25) %>%
  695. pull(Years) %>%
  696. summary()
  697. tableby(`Class (Base)` ~ CDMEM,
  698. data = qs_pacc %>%
  699. filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)) %>%
  700. filter(`CDR Progressor` == 'CDR Progressor'),
  701. digits = 2, test.always = TRUE, numeric.stats = c("Nmiss", "meansd"),
  702. numeric.simplify = TRUE, cat.simplify = FALSE) %>%
  703. summary(labelTranslations = mylabels, stats.labels = list(Nmiss = 'N missing'), text = TRUE) %>%
  704. as.data.frame() %>%
  705. kbl("pipe", booktabs = T)
  706. ```
  707. ## Characterization and Prediction of Latent Classes
  708. ```{r fig-long-spaghetti-pacc-mri}
  709. #| fig-cap: "**Individual and mean PACC trajectories by latent class.** Left panel (A): Spaghetti plot of individual participant trajectories on the Preclinical Alzheimer Cognitive Composite (PACC), colored by latent class derived from the latent class mixed model (LCMM). Each line represents one participant's observed scores over time. Right panel (B): Estimated mean PACC trajectories for each latent class, with shaded regions indicating 95% confidence intervals. Higher PACC scores indicate better cognitive performance."
  710. data_pred_mri <- qs_pacc %>%
  711. select(`Class (Base)`, id, Female, Sola, AAPOEGNPRSNFLG, SUVRCER_z, Ptau217_z,
  712. AGEYR_z, EDCCNTU_z, Hipp_atrophy_z) %>%
  713. distinct(.keep_all = TRUE) %>%
  714. pivot_longer(cols = Female:Hipp_atrophy_z, names_to = 'variable', values_to = 'value') %>%
  715. group_by(variable, `Class (Base)`) %>%
  716. summarise(mean = mean(value)) %>%
  717. pivot_wider(names_from = variable, values_from = mean) %>%
  718. mutate(QSVERSION = sort(unique(qs_pacc$QSVERSION))[1]) %>%
  719. cross_join(tibble(ADURW = seq(0, 425, by = 25)))
  720. pred_stable <-
  721. predictY(m.best.mri, data_pred_mri %>% filter(`Class (Base)` == 'Stable'),
  722. var.time = "ADURW", draws = TRUE)$pred %>%
  723. as_tibble() %>%
  724. bind_cols(data_pred_mri %>% filter(`Class (Base)` == 'Stable') %>%
  725. select(`Class (Base)`, Week = ADURW)) %>%
  726. select(`Class (Base)`, Week, prediction = Ypred_class3,
  727. lower = lower.Ypred_class3, upper = upper.Ypred_class3)
  728. pred_slow <-
  729. predictY(m.best.mri, data_pred_mri %>% filter(`Class (Base)` == 'Slow decliner'),
  730. var.time = "ADURW", draws = TRUE)$pred %>%
  731. as_tibble() %>%
  732. bind_cols(data_pred_mri %>% filter(`Class (Base)` == 'Slow decliner') %>%
  733. select(`Class (Base)`, Week = ADURW)) %>%
  734. select(`Class (Base)`, Week, prediction = Ypred_class1,
  735. lower = lower.Ypred_class1, upper = upper.Ypred_class1)
  736. pred_fast <-
  737. predictY(m.best.mri, data_pred_mri %>% filter(`Class (Base)` == 'Fast decliner'),
  738. var.time = "ADURW", draws = TRUE)$pred %>%
  739. as_tibble() %>%
  740. bind_cols(data_pred_mri %>% filter(`Class (Base)` == 'Fast decliner') %>%
  741. select(`Class (Base)`, Week = ADURW)) %>%
  742. select(`Class (Base)`, Week, prediction = Ypred_class2,
  743. lower = lower.Ypred_class2, upper = upper.Ypred_class2)
  744. pred <- bind_rows(pred_stable, pred_slow, pred_fast) %>%
  745. mutate(
  746. Years = Week * 7 / 365.25,
  747. prediction_raw = reverse_boxcox(prediction,
  748. center = attr(base_model_data$Z, "scaled:center"),
  749. scale = attr(base_model_data$Z, "scaled:scale")),
  750. lower_raw = reverse_boxcox(lower,
  751. center = attr(base_model_data$Z, "scaled:center"),
  752. scale = attr(base_model_data$Z, "scaled:scale")),
  753. upper_raw = reverse_boxcox(upper,
  754. center = attr(base_model_data$Z, "scaled:center"),
  755. scale = attr(base_model_data$Z, "scaled:scale")))
  756. yrange <- qs_pacc %>% filter(!is.na(`Class (Base)`)) %>%
  757. pull(Y) %>% range()
  758. p1 <- qs_pacc %>% filter(!is.na(`Class (Base)`)) %>%
  759. mutate(Years = ADURW * 7 / 365.25) %>%
  760. ggplot(aes(x = Years, y=Y, group = id)) +
  761. geom_line(aes(color = `Class (Base)`), alpha = 0.2) +
  762. scale_x_continuous(breaks = 0:8) +
  763. ylab('PACC') +
  764. ylim(yrange) +
  765. guides(colour = guide_legend(override.aes = list(alpha = 1)))+
  766. guides(color = "none") +
  767. theme_jama() +
  768. scale_color_jama() +
  769. scale_fill_jama()
  770. p2 <- ggplot(pred, aes(x=Years, y=prediction_raw, group = `Class (Base)`)) +
  771. geom_line(aes(color = `Class (Base)`)) +
  772. geom_ribbon(aes(ymin = lower_raw, ymax = upper_raw, fill = `Class (Base)`), alpha = 0.2) +
  773. ylab('Predicted PACC (95% CI)') +
  774. scale_x_continuous(breaks = 0:8) +
  775. ylim(yrange) +
  776. guides(color = guide_legend(title = "PACC class",
  777. override.aes = list(size = 5)), fill = "none") +
  778. theme_jama() +
  779. theme(legend.position = 'inside', legend.position.inside = c(0.3, 0.2)) +
  780. scale_color_jama() +
  781. scale_fill_jama()
  782. p1 + p2 + plot_annotation(tag_levels = "A")
  783. ```
  784. ```{r prediction-results-for-text, include = FALSE}
  785. pred %>%
  786. mutate(Year = Week * 7 / 365.25) %>%
  787. filter(Week %in% c(0, 312))
  788. ```
  789. ```{r fig-sina-baseline, results='asis', fig.width = 6, fig.height = 6.5}
  790. #| fig-cap: "**Distribution of baseline demographic and biomarker variables by latent class of cognitive decline.** Sina plots show the distribution of selected baseline variables among latent classes identified by the latent class mixed model (LCMM) of Preclinical Alzheimer Cognitive Composite (PACC) trajectories: stable, slow decliner, and fast decliner. Variables include age, baseline PACC score, plasma phosphorylated tau 217 (P-tau217), amyloid positron emission tomography (PET) standardized uptake value ratio (SUVr), tau PET SUVr (subset with tau PET available), and hippocampal atrophy (volume normalized to intracranial volume). Each dot represents an individual participant; box plots represent quartiles. Higher PACC scores indicate better cognitive performance. **Abbreviations:** PET = positron emission tomography; P-tau217 = plasma phosphorylated tau 217; SUVr = standardized uptake value ratio; U/mL = arbitrary units proportional to assay signal."
  791. pd <- qs_pacc %>%
  792. filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)) %>%
  793. select(`PACC Class` = `Class (Base)`, Y, Ptau217, AMYLCENT, AGEYR, Tau_PET,
  794. Hipp_atrophy_z) %>%
  795. filter(!is.na(`PACC Class`)) %>%
  796. pivot_longer(Y:Hipp_atrophy_z) %>%
  797. mutate(
  798. name = case_when(
  799. name == 'Y' ~ "PACC",
  800. name == 'AMYLCENT' ~ "Amyloid PET (CL)",
  801. name == 'EDCCNTU' ~ 'Education (years)',
  802. name == 'AGEYR' ~ 'Age (years)',
  803. name == 'Ptau217' ~ "P-tau217 (U/ml)",
  804. name == 'Tau_PET' ~ "Tau PET (SUVr)",
  805. name == 'Hipp_atrophy_z' ~ "Hipp. atrophy (z-score)") %>%
  806. factor(levels = c('Age (years)', 'PACC', 'P-tau217 (U/ml)',
  807. 'Amyloid PET (CL)', 'Tau PET (SUVr)', 'Hipp. atrophy (z-score)')))
  808. ggplot(pd, aes(y=value, x=`PACC Class`, color=`PACC Class`)) +
  809. geom_sina() +
  810. geom_boxplot(width = 0.3, alpha = 0.3, outlier.shape = NA,
  811. color = "black", linewidth = 0.6) +
  812. facet_wrap(vars(name), scales = 'free_y') +
  813. theme_jama() +
  814. theme(legend.position = 'none',
  815. axis.text.x = element_text(angle = 45, hjust = 1, vjust = 1),
  816. strip.text = element_text(size = 7, vjust = 1)) +
  817. scale_color_jama() +
  818. xlab('') + ylab('')
  819. ```
  820. ```{r lcmm-summary-function}
  821. summarize_lcmm <- function(x) {
  822. # Get indices
  823. NPROB <- x$N[1]
  824. NEF <- x$N[2]
  825. NVC <- x$N[3]
  826. NW <- x$N[4]
  827. ncor <- x$N[5]
  828. NPM <- length(x$best)
  829. # Get parameter estimates
  830. id <- 1:NPM
  831. indice <- rep(id*(id+1)/2)
  832. se <- sqrt(x$V[indice])
  833. coef <- x$best
  834. # Calculate statistics
  835. z_stat <- coef / se
  836. p_value <- 2 * (1 - pnorm(abs(z_stat)))
  837. ci_lower <- coef - 1.96 * se
  838. ci_upper <- coef + 1.96 * se
  839. # Parameter names
  840. param_names <- names(coef)
  841. tmp <- tibble(
  842. Parameter = param_names,
  843. Estimate = coef,
  844. SE = se,
  845. lwr = ci_lower,
  846. upr = ci_upper,
  847. OR = exp(coef),
  848. OR.lwr = exp(ci_lower),
  849. OR.upr = exp(ci_upper),
  850. `P Value` = p_value)
  851. # Identify longitudinal parameters (adjust based on your model)
  852. memb_indices <- 1:NPROB
  853. long_indices <- (NPROB+1):(NPROB+NEF)
  854. list(
  855. memb_sum = tmp[memb_indices, ] %>%
  856. select(Parameter, OR, OR.lwr, OR.upr, `P Value`),
  857. long_sum = tmp[long_indices, ] %>%
  858. select(Parameter, Estimate, SE, lwr, upr, `P Value`))
  859. }
  860. ```
  861. Individual and base LCMM (excluding Tau PET) mean trajectories are shown in @fig-long-spaghetti-pacc-mri. Model selection criteria identified the model with two degrees of freedom and three latent classes, which we labeled as stable (77%), slow decliner (16%), and fast decliner (7%). The next best fitting model had three classes and three degrees of freedom ($\Delta$ BIC = 22.2, $\Delta$ ICL = 22.2). The best fitting Tau PET LCMM had two degrees of freedom and three latent classes (@suppfig-long-spaghetti-pacc-tau-pet) and the next best fitting model had two classes and two degrees of freedom ($\Delta$ BIC = 376.4, $\Delta$ ICL = 336.9). The two models largely agreed on individual classifications. Only 7/427 (1.6%) individuals were reclassified from one model to the next, and none of those were fast decliners. Posterior class probabilities (@suppfig-postprobs) indicated good separation (mean posterior probability, 0.94), indicating model confidence in class assignment.
  862. The three classes are significantly separated starting immediately at baseline. From the base model, participants classified as fast decliners started with a mean -0.98 (95% confidence interval \[CI\] -1.44 to -0.52) PACC points and declined to a mean of -15.8 (95% CI -16.7 to -15.0) at six years. In contrast stable individuals started with a mean 0.52 (95% CI 0.37 to 0.67) PACC points and improved to a mean of 1.16 (95% CI 1.00 to 1.31) at six years. Slow decliners were intermediate between the other two classes starting with a mean -0.13 (95% CI -0.44 to 0.17) PACC points and declining to a mean -4.74 (95% CI -5.22 to -4.26).
  863. Clinical relevance of the PACC latent classes was supported by differences in functional outcomes. The proportion of participants who were observed to progress on the Clinical Dementia Rating (CDR) Global score increased monotonically across latent classes, with the lowest rate among stable individuals (22.8%) and progressively higher rates among slow- (71.9%) and fast-decliners (88.2%) (@tbl-baseline-base-characteristics). However, discordance between PACC latent class labels and observed CDR progression is common. While CDR Global scores greater than zero are possible without an indication of memory issues, this explains a very little of the discordance. Among CDR progressors, the proportion of individuals with a CDR Memory Box score of zero was 8.4% of PACC stable individuals, 5.5% of slow decliners, and 4.8% of fast decliners.
  864. The distribution of continuous baseline variables by latent class is shown in @fig-sina-baseline. @suppfig-lcmm-coefs summarizes standardized coefficients and 95% CIs from both components of the LCMM, and @supptbl-lcmm-raw-coefs shows these coefficients on the original scale of the covariates. @suppfig-lcmm-coefs Panels A and C depicts odds ratios (ORs) for the class membership submodels from the base model and tau PET model. Continuous predictors are standardized so that ORs reflect the change in odds of belonging to the fast- or slow declining class compared with the stable class per one SD increase in each continuous baseline predictor.
  865. From the base model (Panel A), higher plasma P-tau217 levels (OR 3.2, 95% CI 2.4 to 4.1), elevated amyloid PET centiloid (OR 1.4, 95% CI 1.0 to 1.9), and greater hippocampal atrophy scores (OR 3.9, 95% CI 2.7 to 5.6) were each associated with increased odds of belonging to faster-declining classes. Interestingly, P-tau217 and hippocampal atrophy had numerically stronger effects than amyloid PET. Also of interest, fast decliners were numerically younger (mean 73.1 years, SD 4.7), on average, than slow decliners (mean 74.1 years, SD 5.2) and increased age was associated with membership in the slow declining group (OR 1.4, 95% CI 1.1 to 1.7) but not the fast declining group (OR 0.93, 95% CI 0.69 to 1.2). A possible explanation for the non‑monotonic relationship between age and latent class severity (slow decliners are oldest followed by fast decliners, then stable) is that age exerts two opposing influences in this cohort. On one hand, older age is associated with a greater likelihood of belonging to the slow‑declining class, consistent with the well‑established role of age as a risk factor for cognitive decline. On the other hand, some comparatively younger amyloid‑positive individuals may have more aggressive or rapidly progressive disease biology, which attenuates the age association with membership in the fast‑declining class. In addition, survivor bias may contribute to the slow‑declining class having the highest mean age, as individuals with more aggressive forms of preclinical AD are less likely to be represented at older ages. Tau PET model results are largely similar, but confidence intervals are wider, likely due to smaller available sample ($N=427$ vs $N=1,629$). Increased tau PET was associated with increased odds of slow- (OR = 1.8, 95% CI 1.2 to 2.8) and fast-declining groups (OR = 3.6, 95% CI 2.1 to 6.1).
  866. @suppfig-lcmm-coefs Panels B and D presents standardized coefficients from the longitudinal submodels, representing associations between each covariate and the PACC outcome. Higher PACC scores correspond to better cognitive performance and continuous variables are standardized and oriented so that positive coefficients indicate better performance. We see that females had increased risk of being fast decliners (OR 2.5, 95% CI 1.5 to 4.4), but had better PACC performance for a given class membership (0.30 PACC points, 95% CI 0.24 to 0.35). APOE $\epsilon 4$ carriage was associated with increased risk of both slow (OR 1.9, 95% CI 1.3 to 2.8) and fast (OR 2.9, 95% CI 1.6 to 5.3) declining class versus stable class, but negligible remaining effect on PACC. Education had no effect on class membership, but was associated with better PACC performance (0.12 PACC points per SD of education, 95% CI 0.095 to 0.15).
  867. ## Power Analysis Stratified by Latent Class
  868. ```{r fig-clda-power, fig.width = 7, fig.height = 7/1.6}
  869. #| fig-cap: "**Trends for power analysis stratified by latent class.** Mean trajectories were estimated from longitudinal models using natural cubic splines (two degrees of freedom) applied to Preclinical Alzheimer Cognitive Composite (PACC) scores for (a) two decliner classes, (b) amyloid-positive stable participants, and (c) a reference group of amyloid-negative stable individuals representing the maximum possible treatment benefit. For decliners, power calculations assume a treatment effect equal to 20% of the maximum possible benefit at 2 and 4 years. For amyloid-positive stable participants, power calculations assume a treatment effect equal to 100% of the maximum possible benefit at 2 and 4 years. Dashed lines represent projected mean PACC scores under treatment; crosses indicate time points used for power calculations."
  870. clda_data <- qs_pacc %>%
  871. filter(EPOCH %in% c("BLINDED TREATMENT", "Analysis visit <= 66") |
  872. AVISIT == "006") %>%
  873. mutate(ASEQNCS = as.factor(ASEQNCS), id = as.factor(id))
  874. clda_decline <- mmrm::mmrm(
  875. Y ~ I(ns21(ADURW)) + I(ns22(ADURW)) + QSVERSION + Female + Sola +
  876. AAPOEGNPRSNFLG + Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z +
  877. us(ASEQNCS | id),
  878. data = clda_data %>% filter(`Class (Base)` != 'Stable'))
  879. clda_stable_a4 <- mmrm::mmrm(
  880. Y ~ I(ns21(ADURW)) + I(ns22(ADURW)) + QSVERSION + Female + Sola +
  881. AAPOEGNPRSNFLG + Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z +
  882. us(ASEQNCS | id),
  883. data = clda_data %>% filter(`Class (Base)` == 'Stable', SUBSTUDY == 'A4'))
  884. clda_stable_learn <- mmrm::mmrm(
  885. Y ~ I(ns21(ADURW)) + I(ns22(ADURW)) + QSVERSION + Female + Sola +
  886. AAPOEGNPRSNFLG + Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z +
  887. us(ASEQNCS | id),
  888. data = clda_data %>% filter(`Class (Base)` == 'Stable', SUBSTUDY == 'LEARN'))
  889. #vis 36 w120 (2.3 years)
  890. #vis 66 w240 (4.6 years)
  891. em_stable_learn <- ref_grid(clda_stable_learn,
  892. at = list(ADURW = seq(0, 5, by=0.5)*(365.25/7))) %>%
  893. emmeans(specs = "ADURW") %>%
  894. as_tibble() %>%
  895. mutate(Group = 'Stable Amyloid-') %>%
  896. bind_cols(tibble(Variance = diag(clda_stable_learn$cov)))
  897. em_stable_a4 <- ref_grid(clda_stable_a4,
  898. at = list(ADURW = seq(0, 5, by=0.5)*(365.25/7))) %>%
  899. emmeans(specs = "ADURW") %>%
  900. as_tibble() %>%
  901. mutate(Group = 'Stable Amyloid+') %>%
  902. bind_cols(tibble(Variance = diag(clda_stable_a4$cov))) %>%
  903. left_join(em_stable_learn %>% select(ADURW, upper_limit = emmean), by = c("ADURW"))
  904. em_decline <- ref_grid(clda_decline,
  905. at = list(ADURW = seq(0, 5, by=0.5)*(365.25/7))) %>%
  906. emmeans(specs = "ADURW") %>%
  907. as_tibble() %>%
  908. mutate(Group = 'Decliners') %>%
  909. bind_cols(tibble(Variance = diag(clda_decline$cov))) %>%
  910. left_join(em_stable_learn %>% select(ADURW, upper_limit = emmean), by = c("ADURW"))
  911. em_all <- bind_rows(em_decline, em_stable_a4, em_stable_learn) %>%
  912. mutate(Years = ADURW / (365.25/7))
  913. em_effect <- em_all %>%
  914. filter(Years == 0 | Years >= 2) %>%
  915. mutate(
  916. `Delta (%)` = case_when(
  917. Group == 'Decliners' ~ 20,
  918. Group == 'Stable Amyloid+' ~ 100),
  919. Delta = `Delta (%)`/100 * (upper_limit - emmean),
  920. Treated = case_when(
  921. Years == 0 ~ emmean,
  922. TRUE ~ emmean + Delta))
  923. ggplot(em_all, aes(x=Years, y=emmean)) +
  924. geom_line(aes(color=Group), size = 1.5) +
  925. geom_line(data=em_effect, aes(x=Years, y=Treated, color=Group),
  926. linetype='dashed', size = 1.5) +
  927. geom_point(data=em_effect %>% filter(Years %in% c(2,4)),
  928. aes(x=Years, y=Treated, color=Group),
  929. shape=3, stroke = 2) +
  930. ylab("Mean PACC Score") +
  931. theme_jama() +
  932. scale_color_manual(values = c("#B24745", "#79AF97", "#6A6599"),
  933. guide = guide_legend(override.aes = list(shape = NA))) +
  934. theme(
  935. legend.position = 'inside',
  936. legend.position.inside = c(0.2, 0.2))
  937. ```
  938. Power analyses stratified by latent class revealed marked differences in cognitive trajectories and variability across classes (see @supptbl-clda-power). The amyloid-negative stable group had the highest mean (SD) PACC scores over time: +1.14 (2.11) at 2 years and +1.20 (2.25) at 4 years. We considered this the maximum possible treatment benefit (see @fig-clda-power). The stable amyloid-positive group had means of +0.88 (2.19) at 2 years and +0.96 (2.34) at 4 years. With 500 participants per arm, a trial in amyloid-positive stable individuals would have only 44% power at 2 years and 30% at 4 years to detect even 100% of the maximum benefit (0.26 PACC points at 2 years and 0.24 at 4 years). In contrast, a trial in amyloid-positive declining individuals would have 93% power at 2 years and 98% at 4 years to detect just 20% of the maximum benefit (0.79 PACC points at 2 years and 1.40 at 4 years).
  939. These findings suggest that stable amyloid-positive individuals, although they may have the greatest potential long-term benefit from a disease-modifying drug, contribute little to the statistical power of a preclinical Alzheimer's disease trial to detect an effect on the PACC. Instead, power is primarily driven by declining individuals, who represent a minority of the amyloid-positive population. Enrolling large numbers of stable individuals may dilute the overall treatment effect and reduce the trial's ability to detect meaningful cognitive benefits. These results underscore the need to account for latent classes of decline when designing preclinical Alzheimer's disease trials. Again, the goal of this analysis was not to suggest conducting trials in homogeneous populations. Rather, we aimed to highlight the complexity and potential inefficiency introduced when multiple latent classes with distinct trajectories are present within a trial cohort.
  940. ## Predictive Performance and Cross-Validation
  941. The overall cross-validated accuracy was 0.80 (95% CI 0.79 to 0.82) for the base model and 0.756 (95% CI 0.73 to 0.78) for the tau PET model. However, this accuracy rate is inflated by the large number of stable individuals (1,257/1,629 = 77.2%), which are relatively easy for the model to identify. AUPRCs from the base model were 0.95, 0.41, and 0.49 for non-, slow-, and fast decliners versus the rest; and 0.70 for either slow- or fast decliner vs stable (@suppfig-cross-val-base-pr). AUPRCs from the tau PET model were 0.92, 0.31, and 0.56 for non-, slow-, and fast decliners versus the rest; and 0.71 for either slow- or fast decliner vs stable (@suppfig-cross-val-tau-pet-pr).
  942. ## Post-hoc Regression Tree
  943. The optimally-tuned classification tree without tau PET achieved a mean balanced accuracy of 0.63 (standard error 0.007) in 10-fold cross-validation (@suppfig-tree-base). Variable importance analysis identified baseline P-tau217 (56.2%), baseline PACC (15.5%), amyloid PET (13.1%), and hippocampal atrophy (10.0%) as most important; though amyloid PET was not included in the tree, likely due to its correlation with P-tau217. The tree structure revealed a hierarchical decision process with participants with baseline P-tau217 \< 0.29 were classified as stable individuals, though this miss-classifies $N=155$ decliners. Those with P-tau217 $\geq$ 0.29 are further parsed by hippocampal atrophy, PACC, and P-tau217. Only one leaf contains a majority of fast decliners, and it contains only 1% of the initial sample. Note that although P‑tau217 was selected instead of amyloid PET in the regression tree, amyloid PET is still recognized as an important predictor. This occurs because the tree algorithm selects the single best classifier among correlated variables (such as P‑tau217 and amyloid PET). Methods that aggregate multiple trees (e.g., random forests) could retain both biomarkers and leverage any independent predictive value each contributes.
  944. While the tree demonstrates the relative importance and utility of the predictors, it also demonstrates a high degree of miss-classification (31.4%). Cross-validated balanced accuracy is improved by the addition of tau PET (mean 0.69, standard error 0.02; @suppfig-tree-tau-pet). Tau PET has the second largest variable importance (22.5%), between P-tau217 (29.8%) and hippocampal atrophy (17.4%). The tree with tau PET results in a larger proportion of leaves which are majority fast decliners (8.5% versus 1% without tau PET).
  945. ## Longitudinal Biomarker Progression Across Latent Classes
  946. To further characterize biological differences among the latent trajectory groups, we compared longitudinal changes in amyloid PET, plasma P-tau217, tau PET (medial temporal and neocortical), and hippocampal volume across the four classes (@fig-long-biomarkers). As expected, slow and fast decliners showed the most rapid biomarker progression across modalities, consistent with their steeper cognitive decline. In contrast, Aβ– stable individuals demonstrated minimal change over time in all biomarkers, supporting their classification as biologically and clinically stable. Notably, Aβ+ stable individuals, despite showing PACC trajectories that closely resembled Aβ– stable individuals, exhibited clear evidence of biomarker progression, including increasing amyloid burden, rising plasma and PET tau signals, and accelerating hippocampal atrophy. These findings suggest that many Aβ+ stable individuals may represent an earlier stage of preclinical disease and could transition into declining classes with extended follow-up, highlighting the temporal dissociation between biomarker progression and short-term cognitive change.
  947. ```{r long-biomarkers-lme-fit}
  948. Means <- lapply(unique(long_biomarker$Outcome), function(cc){
  949. tmp <- filter(long_biomarker, Outcome == cc)
  950. if(cc == "P-tau217 (U/ml)"){
  951. tmp$Y.log <- log10(tmp$Y)
  952. fit <- lme(Y.log ~ (ns21(Weeks) + ns22(Weeks))*Group,
  953. random = ~ Weeks | BID, data = tmp)
  954. return(emmeans(fit, ~ Weeks | Group,
  955. at = list(Weeks = (0:5)*52.17857,
  956. Group = unique(tmp$Group))) %>%
  957. as_tibble() %>%
  958. mutate(Outcome = cc,
  959. emmean = 10^emmean,
  960. lower = 10^lower.CL,
  961. upper = 10^upper.CL))
  962. }else{
  963. fit <- lme(Y ~ (ns21(Weeks) + ns22(Weeks))*Group,
  964. random = ~ Weeks | BID, data = tmp)
  965. return(emmeans(fit, ~ Weeks | Group,
  966. at = list(Weeks = (0:5)*52.17857,
  967. Group = unique(tmp$Group))) %>%
  968. as_tibble() %>%
  969. mutate(Outcome = cc))
  970. }
  971. }) %>% bind_rows()
  972. ```
  973. ```{r fig-long-biomarkers, fig.width = 8, fig.height = 8}
  974. #| fig-cap: "**Longitudinal biomarker trajectories by PACC latent class.** Each panel displays individual biomarker trajectories (light lines) and the estimated mean trajectory (bold black line) from a linear mixed-effects model with random slopes and natural cubic spline fixed effects (2 degrees of freedom). Columns correspond to Aβ– stable, Aβ+ stable, slow decliner, and fast decliner. Rows represent five longitudinal biomarkers: amyloid PET (Centiloids), plasma P-tau217 (U/mL), medial temporal lobe (MTL) tau PET SUVR, neocortical tau PET SUVR, and hippocampal volume. While Aβ– and Aβ+ stable individuals demonstrate similarly stable cognitive trajectories, the Aβ+ stable individuals exhibit clear biomarker progression across amyloid, tau, and neurodegeneration measures, suggesting that many may be in an earlier stage of preclinical disease and could transition into declining classes with longer follow-up."
  975. OC <- levels(long_biomarker$Outcome)
  976. names(OC) <- OC
  977. pp <- lapply(OC, function(cc){
  978. tmp <- filter(long_biomarker, Outcome == cc)
  979. mean.tmp <- filter(Means, Outcome == cc)
  980. p <- ggplot(tmp, aes(x = Weeks/52.17857, y = Y)) +
  981. facet_grid(. ~ Group, scales = 'free_y') +
  982. geom_line(aes(group=BID, color=`Class (Base)`), alpha = 0.05) +
  983. geom_point(data = tmp %>%
  984. filter(Outcome == 'Hippocampus (cc)' & Group == 'Aβ- stable'),
  985. aes(color=`Class (Base)`), alpha = 0.1) +
  986. geom_line(data = mean.tmp %>% filter(!(Outcome == 'Hippocampus (cc)' &
  987. Group == 'Aβ- stable')),
  988. aes(y = emmean, x = Weeks/52.17857), color="white",
  989. linewidth = 1.5) +
  990. geom_line(data = mean.tmp %>% filter(!(Outcome == 'Hippocampus (cc)' &
  991. Group == 'Aβ- stable')),
  992. aes(y = emmean, x = Weeks/52.17857), color="black",
  993. linewidth = 1) +
  994. ylab(cc) +
  995. xlab('Years') +
  996. coord_cartesian(xlim = c(0,8)) +
  997. theme_jama() +
  998. theme(legend.position = 'none') +
  999. scale_color_jama() +
  1000. scale_fill_jama()
  1001. if(cc %in% c("Neocortical tau (SUVr)", "MTL tau (SUVr)"))
  1002. p <- p + coord_cartesian(xlim = c(0,8), ylim = c(1, 1.75))
  1003. if(cc == "Amyloid PET (CL)")
  1004. p <- p + coord_cartesian(xlim = c(0,8), ylim = c(-50, 150))
  1005. if(cc == "P-tau217 (U/ml)")
  1006. p <- p + coord_cartesian(xlim = c(0,8), ylim = c(0, 0.75))
  1007. if(cc == "Hippocampus (cc)")
  1008. p <- p + coord_cartesian(xlim = c(0,8), ylim = c(4, 8)) +
  1009. # t, r, b, l
  1010. theme(plot.margin=unit(c(0,0,0,0.9),"lines"))
  1011. if(cc == "Hippocampus (cc)") p <- p + xlab('Years') else p <- p + xlab('')
  1012. if(cc != "Amyloid PET (CL)") p <- p + theme(
  1013. strip.background = element_blank(),
  1014. strip.text.x = element_blank())
  1015. if(cc != "Hippocampus (cc)") p <- p + theme(
  1016. axis.line.x = element_blank(),
  1017. axis.title.x = element_blank(),
  1018. axis.text.x = element_blank(),
  1019. axis.ticks.x = element_blank())
  1020. p
  1021. })
  1022. grid.arrange(grobs=pp, ncol = 1)
  1023. ```
  1024. # Discussion
  1025. In this latent class mixed-model analysis of cognitively unimpaired older adults, we identified distinct trajectories of PACC performance and evaluated the extent to which plasma and imaging biomarkers prospectively distinguished individuals who remained cognitively stable from those who showed slow or rapid decline. Several findings highlight challenges for prognostic modeling in the preclinical stage of Alzheimer disease and have direct implications for secondary prevention trial design.
  1026. ## Available Predictors Only Partially Discriminate Future Decline
  1027. Although plasma P-tau217 and amyloid PET were each associated with latent class membership, neither biomarker alone, or in combination, achieved high accuracy in predicting of cognitive decline. A key observation was the high frequency of cognitive stability among amyloid-positive participants. As shown in @tbl-baseline-base-characteristics, 776 of 1,110 amyloid-positive A4 participants (69.9%) were classified as stable individuals, indicating that substantial amyloid burden does not necessarily translate into measurable cognitive change over the timescales typical of prevention trials. However, this is consistent with a pre-symptomatic phase of disease believed to last as long as 15 years.
  1028. Compared to the 30.1% of amyloid-positive individuals classified as decliners, a small minority of amyloid-negative LEARN participants were classified as decliners. Thirty-eight of 519 (7.3%) amyloid-negative individuals were classified as slow or fast decliners, suggesting that some causes of decline may reflect non–AD processes, early AD pathology below PET detection thresholds at baseline that accumulated over the 5 year period, or cognitive measurement variability. These findings reinforce that amyloid positivity, while necessary for defining the AD pathophysiologic continuum, is an insufficient standalone predictor of short- to intermediate-term cognitive decline.
  1029. ## Relative Contributions of Plasma P-tau217 and Tau PET
  1030. Plasma P-tau217 demonstrated stronger prognostic associations than amyloid PET, consistent with biomarker models positioning tau abnormalities closer to symptom onset. This observation aligns with prior reports, showing improved prognostic accuracy with plasma tau markers.[@Sperling2024] However, classification tree analyses indicated that no single P-tau217 threshold meaningfully enriched for decliners without simultaneously excluding many true decliners, underscoring the limitations of threshold-based enrichment strategies.
  1031. Tau PET provided incremental predictive value beyond plasma P-tau217 and amyloid PET, although gains were modest. Because tau PET captures later-stage pathologic changes, these results are biologically plausible. Nevertheless, the logistical and financial limitations of tau PET constrain its feasibility as a routine screening tool in large-scale prevention trials.
  1032. ## Implications for Secondary Prevention Trial Design
  1033. **Participant Selection.** Because most amyloid-positive individuals remain cognitively stable, eligibility based solely on amyloid positivity will enroll many stable participants, reducing power for cognitive endpoints. Plasma P-tau217 may improve enrichment but cannot reliably isolate decliners. More complex multimodal strategies or repeated biomarker assessments may be needed for adequate prognostic discrimination.
  1034. **Designing Trials for Heterogeneous Cohorts.** Traditional power calculations assume a homogeneous population and uniform treatment effect size, overlooking substantial heterogeneity in cognitive trajectories. Fast decliners may show larger measurable effects but often represent individuals further along the disease continuum and exhibit greater variability, while stable individuals pose a different challenge. Their stability limits detectable cognitive benefit even if biological effects occur, requiring treatments to produce gains beyond natural improvement. Our analysis shows that variability across classes directly influences minimum detectable effect sizes, meaning that single-effect-size assumptions risk underpowered studies or inefficient resource allocation. More realistic simulation-based approaches that incorporate class-specific trajectories and differential treatment effects can improve trial design by informing enrichment strategies, optimizing sample sizes, and aligning outcomes with subgroup trajectories.
  1035. **Added Value of Tau PET.** Tau PET provided additional predictive information but its incremental value must be weighed against feasibility. Tau-PET-based staging may help identify individuals closer to clinical transition, though widespread implementation in prevention trials remains challenging.
  1036. **Outcome Selection.** Limited cognitive change in stable individuals raises questions about the sensitivity of traditional cognitive outcomes at this early stage. Digital measures of learning curves and biological endpoints such as plasma or tau PET may offer more responsive markers of treatment effect and serve as secondary or exploratory outcomes. Given evidence that tau PET correlates more strongly with cognitive decline than amyloid biomarkers, longitudinal tau PET or plasma P-tau217 may be useful for monitoring treatment response or detecting early trajectory divergence. Future work applying latent class or trajectory models to biomarker data may clarify whether biomarker-defined progressors can be identified earlier or more reliably than cognitive progressors.
  1037. **Early Intervention.** Anti-amyloid therapies show greater benefit in symptomatic individuals with less tau pathology, supporting the theory that earlier intervention yields greater benefit. However, if presymptomatic decline unfolds gradually over a decade or more, trials starting early may struggle to demonstrate cognitive benefit within feasible windows. Refining and validating sensitive digital and biomarker-based endpoints will be essential for advancing preclinical therapy development. Reliable measures of target engagement and disease progression could improve enrichment, increase power, and enable detection of treatment effects even among cognitively stable participants.
  1038. ## Overall Interpretation
  1039. In summary, data-driven cognitive trajectories among initially unimpaired older adults exhibited substantial heterogeneity, and widely used biomarkers, including plasma P-tau217 and amyloid PET, were only partially effective in predicting who would decline over the time frame of a prevention trial. These limitations highlight the need for improved prognostic tools, potentially incorporating longitudinal biomarker dynamics, multimodal risk models, or digital assessments. Until then, secondary prevention trials must account for substantial heterogeneity in cognitive trajectories when planning enrollment, power calculations, trial simulations, and analytic strategies.
  1040. ## Strengths and Limitations
  1041. This study has several strengths. We leveraged a large, deeply phenotyped cohort of cognitively unimpaired older adults with longitudinal cognitive assessments and harmonized biomarker data, using the A4LEARN R package to standardize and reproducibly analyze data across the A4 and LEARN cohorts. The latent class mixed-model approach allowed us to identify data-driven cognitive trajectories rather than relying on predefined thresholds or single-outcome definitions of decline. The inclusion of plasma P-tau217, amyloid PET, and, within a subset, tau PET enabled evaluation of multiple biomarkers spanning distinct phases of Alzheimer disease pathophysiology.
  1042. This study also has limitations. First, although latent class methods improve characterization of heterogeneity, class assignments are probabilistic and sensitive to model specification. Second, follow-up duration, while substantial, may still be insufficient to capture long-term cognitive change among stable individuals. Third, tau PET was available only in a subset, limiting statistical power for comparisons involving this biomarker and potentially reducing generalizability. Fourth, although we examined a broad panel of covariates, unmeasured factors, including comorbidities, lifestyle variables, and scanner-related differences, may influence both biomarker levels and cognitive outcomes. Finally, this is highly selected sample of individuals who were either eligible for the A4 study, or ineligible owing only to low amyloid PET signal, and may not fully represent community-dwelling older adults.
  1043. # Conclusions
  1044. Among cognitively unimpaired older adults, cognitive trajectories showed marked heterogeneity, and widely used biomarkers, including plasma P-tau217 and amyloid PET, were only partially effective in distinguishing who would decline. Tau PET provided modest additional predictive value but remains impractical for widespread enrichment. These findings highlight the challenge of prospectively identifying decliners in preclinical Alzheimer's disease and underscore the need for improved multimodal prognostic tools. Secondary prevention trials should account for substantial heterogeneity in cognitive trajectories when developing enrollment strategies, powering cognitive endpoints, and interpreting treatment effects.
  1045. \clearpage
  1046. # Acknowledgements
  1047. The authors would like to thank the A4 and LEARN Study Teams and site principal investigators and staff. Special gratitude to the A4 and LEARN participants and their study partners, without whom these studies would not be possible.
  1048. **Funding:** The A4 and LEARN Studies were supported by a public-private-philanthropic partnership which included funding from the National Institute of Aging of the National Institutes of Health (R01 AG063689, U19AG010483 and U24AG057437), Eli Lilly (also the supplier of active medication and placebo), the Alzheimer's Association, the Accelerating Medicines Partnership through the Foundation for the National Institutes of Health, the GHR Foundation, the Davis Alzheimer Prevention Program, the Yugilbar Foundation, an anonymous foundation, and additional private donors to Brigham and Women's Hospital, with in-kind support from Avid Radiopharmaceuticals, Cogstate, Albert Einstein College of Medicine and the Foundation for Neurologic Diseases. Other support was provided by grants from the Epstein Family Foundation.
  1049. **Conflict of Interest Disclosures:** Dr Aisen reported personal fees from Merck, Biogen, Roche, AbbVie, ImmunoBrain Checkpoint, Bristol Myers Squibb, and Neurimmune and grants from Eisai outside the submitted work. Dr Sperling reported consulting fees from AbbVie, AC Immune, Acumen, Alector, Apellis, Biohaven, Bristol Myers Squibb, Genentech, Janssen, Nervgen, Oligomerixg, Prothena, Roche, Vigil Neuroscience, Ionis, and Vaxxinity outside the submitted work. Dr Donohue reported personal fees from Roche (consultant) and Johnson & Johnson (spouse is full-time employee) outside the submitted work. Dr Raman reported grants from American Heart Association, Gates Ventures, Eisai, Alzheimer's Association. No other disclosures were reported.
  1050. # Supplementary Material
  1051. ::: {#suppfig-lcmm-coefs}
  1052. ```{r suppfig-lcmm-coefs, fig.width = 7, fig.height = 6}
  1053. ## mri model ----
  1054. raw_names <- c('Hipp_atrophy', 'Ptau217', 'SUVRCER', 'AGEYR',
  1055. 'EDCCNTU', 'AAPOEGNPRSNFLG', 'Female', 'Sola')
  1056. out_names <- c("Hipp. atrophy", "P-tau217", "Amyloid PET", 'Age',
  1057. 'Education', "APOE e4", 'Female', 'Solanezumab')
  1058. tmp.or <- summarize_lcmm(m.best.mri)$memb_sum %>%
  1059. mutate(
  1060. Parameter = gsub(' class1', ' (slow)', Parameter),
  1061. Parameter = gsub(' class2', ' (fast)', Parameter),
  1062. Parameter = gsub('_z', '', Parameter)) %>%
  1063. filter(!grepl("intercept", Parameter)) %>%
  1064. mutate(
  1065. Parameter = factor(Parameter,
  1066. levels = paste(rep(raw_names, each=2), c('(fast)', '(slow)')),
  1067. labels = paste(rep(out_names, each=2), c('(fast)', '(slow)'))))
  1068. tbl.or.mri.raw <- tmp.or
  1069. tmp.long <- summarize_lcmm(m.best.mri)$long_sum %>%
  1070. mutate(
  1071. Parameter = gsub(' class1', ' (slow)', Parameter),
  1072. Parameter = gsub(' class2', ' (fast)', Parameter),
  1073. Parameter = gsub(' class3', ' (non)', Parameter),
  1074. Parameter = gsub('_z', '', Parameter)) %>%
  1075. filter(!grepl("intercept", Parameter)) %>%
  1076. filter(!grepl("I(ns", Parameter, fixed = TRUE)) %>%
  1077. filter(!grepl("QSVERSION", Parameter, fixed = TRUE)) %>%
  1078. mutate(
  1079. Parameter = factor(Parameter,
  1080. levels = raw_names,
  1081. labels = out_names))
  1082. tbl.long.mri.raw <- tmp.long
  1083. pA <- ggplot(tmp.or, aes(x = OR, y = Parameter)) +
  1084. geom_vline(xintercept = 1, linetype = "solid",
  1085. color = "black", linewidth = 0.5) +
  1086. geom_segment(aes(x = OR.lwr, xend = OR.upr,
  1087. y = Parameter, yend = Parameter),
  1088. linewidth = 0.8, color = "black") +
  1089. geom_point(size = 3, shape = 15, color = "black") +
  1090. labs(x = "OR (95% CI)",
  1091. y = NULL,
  1092. title = NULL) +
  1093. theme_jama() +
  1094. xlim(0,7)
  1095. pB <- ggplot(tmp.long, aes(x = Estimate, y = Parameter)) +
  1096. geom_vline(xintercept = 0, linetype = "solid",
  1097. color = "black", linewidth = 0.5) +
  1098. geom_segment(aes(x = lwr, xend = upr,
  1099. y = Parameter, yend = Parameter),
  1100. linewidth = 0.8, color = "black") +
  1101. geom_point(size = 3, shape = 15, color = "black") +
  1102. labs(x = "Standardized Coefficient (95% CI)",
  1103. y = NULL,
  1104. title = NULL) +
  1105. theme_jama() +
  1106. xlim(-0.3, 0.5)
  1107. ## tau pet model ----
  1108. raw_names <- c('Tau_PET', 'Hipp_atrophy', 'Ptau217', 'SUVRCER', 'AGEYR',
  1109. 'EDCCNTU', 'AAPOEGNPRSNFLG', 'Female', 'Sola')
  1110. out_names <- c('Tau PET', "Hipp. atrophy", "P-tau217", "Amyloid PET", 'Age',
  1111. 'Education', "APOE e4", 'Female', 'Solanezumab')
  1112. tmp.or <- summarize_lcmm(m.best.tau)$memb_sum %>%
  1113. mutate(
  1114. Parameter = gsub(' class1', ' (slow)', Parameter),
  1115. Parameter = gsub(' class2', ' (fast)', Parameter),
  1116. Parameter = gsub('_z', '', Parameter)) %>%
  1117. filter(!grepl("intercept", Parameter)) %>%
  1118. mutate(
  1119. Parameter = factor(Parameter,
  1120. levels = paste(rep(raw_names, each=2), c('(fast)', '(slow)')),
  1121. labels = paste(rep(out_names, each=2), c('(fast)', '(slow)'))))
  1122. tmp.long <- summarize_lcmm(m.best.tau)$long_sum %>%
  1123. mutate(
  1124. Parameter = gsub(' class1', ' (slow)', Parameter),
  1125. Parameter = gsub(' class2', ' (fast)', Parameter),
  1126. Parameter = gsub(' class3', ' (non)', Parameter),
  1127. Parameter = gsub('_z', '', Parameter)) %>%
  1128. filter(!grepl("intercept", Parameter)) %>%
  1129. filter(!grepl("I(ns", Parameter, fixed = TRUE)) %>%
  1130. filter(!grepl("QSVERSION", Parameter, fixed = TRUE)) %>%
  1131. mutate(
  1132. Parameter = factor(Parameter,
  1133. levels = raw_names,
  1134. labels = out_names))
  1135. tbl.or.tau.raw <- tmp.or
  1136. tbl.long.tau.raw <- tmp.long
  1137. pC <- ggplot(tmp.or, aes(x = OR, y = Parameter)) +
  1138. geom_vline(xintercept = 1, linetype = "solid",
  1139. color = "black", linewidth = 0.5) +
  1140. geom_segment(aes(x = OR.lwr, xend = OR.upr,
  1141. y = Parameter, yend = Parameter),
  1142. linewidth = 0.8, color = "black") +
  1143. geom_point(size = 3, shape = 15, color = "black") +
  1144. labs(x = "OR (95% CI)",
  1145. y = NULL,
  1146. title = NULL) +
  1147. theme_jama() +
  1148. xlim(0,7)
  1149. pD <- ggplot(tmp.long, aes(x = Estimate, y = Parameter)) +
  1150. geom_vline(xintercept = 0, linetype = "solid",
  1151. color = "black", linewidth = 0.5) +
  1152. geom_segment(aes(x = lwr, xend = upr,
  1153. y = Parameter, yend = Parameter),
  1154. linewidth = 0.8, color = "black") +
  1155. geom_point(size = 3, shape = 15, color = "black") +
  1156. labs(x = "Standardized Coefficient (95% CI)",
  1157. y = NULL,
  1158. title = NULL) +
  1159. theme_jama() +
  1160. xlim(-0.3, 0.5)
  1161. ## arrange plot ----
  1162. pA + pB + pC + pD +
  1163. plot_annotation(
  1164. tag_levels = "A",
  1165. tag_prefix = "",
  1166. tag_suffix = ""
  1167. )
  1168. ```
  1169. **Association of baseline predictors with latent class membership and longitudinal cognitive decline in models with and without tau PET.** Panels A and B display results from the latent class mixed model (LCMM) excluding tau PET; panels C and D show corresponding results from the model including tau PET as an additional predictor. (A, C) Odds ratios (95% CIs) from the LCMM class-membership submodel represent the relative likelihood of belonging to the slow- or fast-declining classes relative to stable class. (B, D) Standardized coefficients (95% CIs) from the LCMM longitudinal submodel show associations between baseline predictors and rate of change in the Preclinical Alzheimer Cognitive Composite (PACC). All continuous predictors and the PACC outcome were standardized (z-scored) before analysis to enable comparison of effect magnitudes. Higher PACC scores indicate better cognitive performance; thus, larger positive coefficients correspond to slower decline or better preservation of cognition over time.
  1170. :::
  1171. ::: {#supptbl-lcmm-raw-coefs}
  1172. ```{r supptbl-lcmm-raw-coefs}
  1173. #| results: asis
  1174. #| tbl-colwidths: [2,10,2,10,5,10,5]
  1175. format.ci <- function(x, digits = 2, ...){
  1176. fx <- format(x, digits = digits, ...)
  1177. paste0(fx[1], " (", fx[2], ", ", fx[3], ")")
  1178. }
  1179. tbl.comb.long <- bind_rows(
  1180. tbl.long.mri.raw %>% mutate(Model = 'Base model', Part = "Longitudidnal"),
  1181. tbl.long.tau.raw %>% mutate(Model = 'Tau PET model', Part = "Longitudidnal")) %>%
  1182. select(Model, Part, Parameter, Estimate, lwr, upr, `P Value`) %>%
  1183. bind_rows(
  1184. bind_rows(
  1185. tbl.or.mri.raw %>% mutate(Model = 'Base model', Part = "Class membership"),
  1186. tbl.or.tau.raw %>% mutate(Model = 'Tau PET model', Part = "Class membership")) %>%
  1187. select(Model, Part, Parameter, Estimate = OR, lwr = OR.lwr, upr = OR.upr, `P Value`)) %>%
  1188. mutate(
  1189. Estimate = case_when(
  1190. grepl("P-tau217", Parameter) ~ Estimate * attr(qs_pacc$Ptau217_z, "scaled:scale"),
  1191. grepl("Amyloid PET", Parameter) ~ Estimate * attr(qs_pacc$SUVRCER_z, "scaled:scale"),
  1192. grepl("Age", Parameter) ~ Estimate * attr(qs_pacc$AGEYR_z, "scaled:scale"),
  1193. grepl("Education", Parameter) ~ Estimate * attr(qs_pacc$EDCCNTU_z, "scaled:scale"),
  1194. grepl("Hipp. atrophy", Parameter) ~ Estimate * attr(qs_pacc$Hipp_atrophy_z, "scaled:scale"),
  1195. grepl("Tau PET", Parameter) ~ Estimate * attr(qs_pacc$Tau_PET_z, "scaled:scale"),
  1196. TRUE ~ Estimate),
  1197. lwr = case_when(
  1198. grepl("P-tau217", Parameter) ~ lwr * attr(qs_pacc$Ptau217_z, "scaled:scale"),
  1199. grepl("Amyloid PET", Parameter) ~ lwr * attr(qs_pacc$SUVRCER_z, "scaled:scale"),
  1200. grepl("Age", Parameter) ~ lwr * attr(qs_pacc$AGEYR_z, "scaled:scale"),
  1201. grepl("Education", Parameter) ~ lwr * attr(qs_pacc$EDCCNTU_z, "scaled:scale"),
  1202. grepl("Hipp. atrophy", Parameter) ~ lwr * attr(qs_pacc$Hipp_atrophy_z, "scaled:scale"),
  1203. grepl("Tau PET", Parameter) ~ lwr * attr(qs_pacc$Tau_PET_z, "scaled:scale"),
  1204. TRUE ~ lwr),
  1205. upr = case_when(
  1206. grepl("P-tau217", Parameter) ~ upr * attr(qs_pacc$Ptau217_z, "scaled:scale"),
  1207. grepl("Amyloid PET", Parameter) ~ upr * attr(qs_pacc$SUVRCER_z, "scaled:scale"),
  1208. grepl("Age", Parameter) ~ upr * attr(qs_pacc$AGEYR_z, "scaled:scale"),
  1209. grepl("Education", Parameter) ~ upr * attr(qs_pacc$EDCCNTU_z, "scaled:scale"),
  1210. grepl("Hipp. atrophy", Parameter) ~ upr * attr(qs_pacc$Hipp_atrophy_z, "scaled:scale"),
  1211. grepl("Tau PET", Parameter) ~ upr * attr(qs_pacc$Tau_PET_z, "scaled:scale"),
  1212. TRUE ~ upr),
  1213. Parameter = as.character(Parameter),
  1214. Parameter = gsub("P-tau217", "P-tau217 (U/ml)", Parameter),
  1215. Parameter = gsub("Amyloid PET", "Amyloid PET (SUVr)", Parameter),
  1216. Parameter = gsub("Age", "Age (yrs)", Parameter),
  1217. Parameter = gsub("Education", "Education (yrs)", Parameter),
  1218. Parameter = gsub("Hipp. atrophy", "Hipp. atrophy (cc$^3$)", Parameter),
  1219. Parameter = gsub("Tau PET", "Tau PET (SUVr)", Parameter),
  1220. `P Value` = Hmisc::format.pval(`P Value`, digits = 3, eps = 0.001, nsmall=3)) %>%
  1221. rowwise() %>%
  1222. mutate(`Estimate (95% CI)` = format.ci(c(Estimate, lwr, upr))) %>%
  1223. select(Model, Part, Parameter, `Estimate (95% CI)`, `P Value`)
  1224. tbl.comb <- left_join(
  1225. tbl.comb.long %>% filter(Model == 'Base model', Part == "Class membership") %>%
  1226. mutate(Parameter0 = gsub(' (slow)', '', Parameter, fixed = TRUE)),
  1227. tbl.comb.long %>% filter(Model == 'Base model', Part == "Longitudidnal") %>%
  1228. rename(Parameter0 = Parameter),
  1229. by = "Parameter0") %>%
  1230. bind_rows(
  1231. left_join(
  1232. tbl.comb.long %>% filter(Model == 'Tau PET model', Part == "Class membership") %>%
  1233. mutate(Parameter0 = gsub(' (slow)', '', Parameter, fixed = TRUE)),
  1234. tbl.comb.long %>% filter(Model == 'Tau PET model', Part == "Longitudidnal") %>%
  1235. rename(Parameter0 = Parameter),
  1236. by = "Parameter0")) %>%
  1237. mutate(Class = case_when(
  1238. grepl("(slow)", Parameter) ~ "slow",
  1239. grepl("(fast)", Parameter) ~ "fast"),
  1240. Parameter = gsub(" (slow)", '', Parameter, fixed = TRUE),
  1241. Parameter = gsub(" (fast)", '', Parameter, fixed = TRUE)) %>%
  1242. select(Model = Model.x, Parameter, Class, `Estimate (95% CI).x`, `P Value.x`,
  1243. `Estimate (95% CI).y`, `P Value.y`)
  1244. tbl.comb %>%
  1245. kable(col.names = c('Model', 'Parameter', 'Class',
  1246. 'Estimate (95% CI)', 'P Value', 'Estimate (95% CI)', 'P Value'),
  1247. escape = FALSE, font_size = 5) %>%
  1248. collapse_rows(c(1,2)) %>%
  1249. add_header_above(
  1250. c(" " = 2, "Class membership model" = 3,
  1251. "Longitudinal model" = 2)) %>%
  1252. kable_styling()
  1253. ```
  1254. **Association of baseline predictors (on the _raw_ scale) with latent class membership and longitudinal cognitive decline in models with and without tau PET.** The top rows display results from the latent class mixed model (LCMM) excluding tau PET, while the bottom rows show corresponding results from the model including tau PET as an additional predictor. The left side shows odds ratios (95% CIs) from the LCMM class-membership submodel represent the relative likelihood of belonging to the slow- or fast-declining classes relative to stable class. The right side shows coefficients (95% CIs) from the LCMM longitudinal submodel show associations between baseline predictors and the Preclinical Alzheimer Cognitive Composite (PACC). All continuous predictors are on the raw scale. Higher PACC scores indicate better cognitive performance; thus, larger positive coefficients correspond to slower decline or better preservation of cognition over time.
  1255. :::
  1256. ::: {#suppfig-tree-base}
  1257. ```{r suppfig-tree-base}
  1258. knitr::include_graphics("suppfig-tree-base.png")
  1259. ```
  1260. {{< include _suppfig-tree-base-cap.txt >}}
  1261. :::
  1262. ::: {#supptbl-baseline-tau-characteristics}
  1263. ```{r supptbl-baseline-tau-characteristics}
  1264. #| results: asis
  1265. #| tbl-colwidths: [30,15,15,15,15,10]
  1266. tableby(`Class (Tau PET)` ~ Group + AGEYR + Sex + Race + Ethnicity + EDCCNTU +
  1267. APOEe4 + Y + AMYLCENT + Ptau217 + Hipp_atrophy_z + Tau_PET +
  1268. `CDR Progressor`,
  1269. data = qs_pacc %>% filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)),
  1270. digits = 2, test.always = TRUE, numeric.stats = c("Nmiss", "meansd"),
  1271. numeric.simplify = TRUE, cat.simplify = FALSE) %>%
  1272. summary(labelTranslations = mylabels, stats.labels = list(Nmiss = 'N missing'), text = TRUE) %>%
  1273. as.data.frame() %>%
  1274. kbl("pipe", booktabs = T)
  1275. ```
  1276. **Baseline demographic, clinical, and biomarker characteristics by latent class of cognitive decline with tau PET.** Baseline characteristics of participants classified by the latent class mixed model (LCMM) of Preclinical Alzheimer Cognitive Composite (PACC) trajectories. Classes represent distinct longitudinal cognitive patterns: stable, slow decliner, and fast decliner. Continuous variables are presented as mean (SD); categorical variables as No. (%). **Abbreviations:** APOE = apolipoprotein E; PET = positron emission tomography; P-tau217 = plasma phosphorylated tau 217; SUVr = standardized uptake value ratio; U/mL = arbitrary units proportional to assay signal; Am. = American; PI = Pacific Islander; CDR = Clinical Dementia Rating; Prog. = Progressor. **Footnotes:** Amyloid PET and tau PET SUVr values represent mean cortical uptake relative to cerebellar reference region. Hippocampal atrophy are residualized for intracranial volume and z-scored. Education reported in years of formal schooling. CDR Progressors were observed to have a CDR Global score greater than zero at two consecutive visits, or their last visit.
  1277. :::
  1278. ::: {#suppfig-long-spaghetti-pacc-tau-pet}
  1279. ```{r suppfig-long-spaghetti-pacc-tau-pet}
  1280. data_pred_tau <- qs_pacc %>%
  1281. select(`Class (Tau PET)`, id, Female, Sola, AAPOEGNPRSNFLG, SUVRCER_z, Ptau217_z,
  1282. AGEYR_z, EDCCNTU_z, Tau_PET_z, Hipp_atrophy_z) %>%
  1283. distinct(.keep_all = TRUE) %>%
  1284. pivot_longer(cols = Female:Hipp_atrophy_z, names_to = 'variable', values_to = 'value') %>%
  1285. group_by(variable, `Class (Tau PET)`) %>%
  1286. summarise(mean = mean(value)) %>%
  1287. pivot_wider(names_from = variable, values_from = mean) %>%
  1288. mutate(QSVERSION = sort(unique(qs_pacc$QSVERSION))[1]) %>%
  1289. cross_join(tibble(ADURW = seq(0, 425, by = 25)))
  1290. pred_stable <-
  1291. predictY(m.best.tau, data_pred_tau %>% filter(`Class (Tau PET)` == 'Stable'),
  1292. var.time = "ADURW", draws = TRUE)$pred %>%
  1293. as_tibble() %>%
  1294. bind_cols(data_pred_tau %>% filter(`Class (Tau PET)` == 'Stable') %>%
  1295. select(`Class (Tau PET)`, Week = ADURW)) %>%
  1296. select(`Class (Tau PET)`, Week, prediction = Ypred_class3,
  1297. lower = lower.Ypred_class3, upper = upper.Ypred_class3)
  1298. pred_slow <-
  1299. predictY(m.best.tau, data_pred_tau %>% filter(`Class (Tau PET)` == 'Slow decliner'),
  1300. var.time = "ADURW", draws = TRUE)$pred %>%
  1301. as_tibble() %>%
  1302. bind_cols(data_pred_tau %>% filter(`Class (Tau PET)` == 'Slow decliner') %>%
  1303. select(`Class (Tau PET)`, Week = ADURW)) %>%
  1304. select(`Class (Tau PET)`, Week, prediction = Ypred_class1,
  1305. lower = lower.Ypred_class1, upper = upper.Ypred_class1)
  1306. pred_fast <-
  1307. predictY(m.best.tau, data_pred_tau %>% filter(`Class (Tau PET)` == 'Fast decliner'),
  1308. var.time = "ADURW", draws = TRUE)$pred %>%
  1309. as_tibble() %>%
  1310. bind_cols(data_pred_tau %>% filter(`Class (Tau PET)` == 'Fast decliner') %>%
  1311. select(`Class (Tau PET)`, Week = ADURW)) %>%
  1312. select(`Class (Tau PET)`, Week, prediction = Ypred_class2,
  1313. lower = lower.Ypred_class2, upper = upper.Ypred_class2)
  1314. pred <- bind_rows(pred_stable, pred_slow, pred_fast) %>%
  1315. mutate(
  1316. Years = Week * 7 / 365.25,
  1317. prediction_raw = reverse_boxcox(prediction,
  1318. center = attr(base_model_data$Z, "scaled:center"),
  1319. scale = attr(base_model_data$Z, "scaled:scale")),
  1320. lower_raw = reverse_boxcox(lower,
  1321. center = attr(base_model_data$Z, "scaled:center"),
  1322. scale = attr(base_model_data$Z, "scaled:scale")),
  1323. upper_raw = reverse_boxcox(upper,
  1324. center = attr(base_model_data$Z, "scaled:center"),
  1325. scale = attr(base_model_data$Z, "scaled:scale")))
  1326. yrange <- qs_pacc %>% filter(!is.na(`Class (Tau PET)`)) %>%
  1327. pull(Y) %>% range()
  1328. p1 <- qs_pacc %>% filter(!is.na(`Class (Tau PET)`)) %>%
  1329. mutate(Years = ADURW * 7 / 365.25) %>%
  1330. ggplot(aes(x = Years, y=Y, group = id)) +
  1331. geom_line(aes(color = `Class (Tau PET)`), alpha = 0.2) +
  1332. scale_x_continuous(breaks = 0:8) +
  1333. ylab('PACC') +
  1334. ylim(yrange) +
  1335. guides(colour = guide_legend(override.aes = list(alpha = 1)))+
  1336. guides(color = "none") +
  1337. theme_jama() +
  1338. scale_color_jama() +
  1339. scale_fill_jama()
  1340. p2 <- ggplot(pred, aes(x=Years, y=prediction_raw, group = `Class (Tau PET)`)) +
  1341. geom_line(aes(color = `Class (Tau PET)`)) +
  1342. geom_ribbon(aes(ymin = lower_raw, ymax = upper_raw, fill = `Class (Tau PET)`), alpha = 0.2) +
  1343. ylab('Predicted PACC (95% CI)') +
  1344. scale_x_continuous(breaks = 0:8) +
  1345. ylim(yrange) +
  1346. guides(color = guide_legend(title = "PACC class",
  1347. override.aes = list(size = 5)), fill = "none") +
  1348. theme_jama() +
  1349. theme(legend.position = 'inside', legend.position.inside = c(0.3, 0.2)) +
  1350. scale_color_jama() +
  1351. scale_fill_jama()
  1352. p1 + p2 + plot_annotation(tag_levels = "A")
  1353. ```
  1354. **Individual and mean PACC trajectories by latent class model with tau PET.** Left panel (A): Spaghetti plot of individual participant trajectories on the Preclinical Alzheimer Cognitive Composite (PACC), colored by latent class derived from the latent class mixed model (LCMM) which included tau PET as a predictor. Each line represents one participant's observed scores over time. Right panel (B): Estimated mean PACC trajectories for each latent class, with shaded regions indicating 95% confidence intervals. Higher PACC scores indicate better cognitive performance. **Abbreviations:** PET = positron emission tomography.
  1355. :::
  1356. ::: {#suppfig-postprobs}
  1357. ```{r postprobs-cap}
  1358. posterior_probs <- m.best.mri$pprob %>%
  1359. pivot_longer(prob1:prob3) %>%
  1360. filter(name == paste0("prob", class)) %>%
  1361. mutate(
  1362. `PACC class` = case_when(
  1363. name == 'prob1' ~ 'Slow decliner',
  1364. name == 'prob2' ~ 'Fast decliner',
  1365. name == 'prob3' ~ 'Stable') %>%
  1366. factor(levels = c('Stable', 'Slow decliner', 'Fast decliner')))
  1367. pp_sum <- posterior_probs %>%
  1368. group_by(`PACC class`) %>%
  1369. summarise(mean = mean(value)) %>%
  1370. mutate(mean = format(round(mean, digits = 2)))
  1371. pp_sum_overall <- posterior_probs %>%
  1372. summarise(mean = mean(value)) %>%
  1373. mutate(mean = format(round(mean, digits = 2)))
  1374. ```
  1375. ```{r suppfig-postprobs, results='asis'}
  1376. ggplot(posterior_probs, aes(x = `PACC class`, y = value, color = `PACC class`)) +
  1377. geom_sina() +
  1378. geom_boxplot(width = 0.3, alpha = 0.3, outlier.shape = NA,
  1379. color = "black", linewidth = 0.6) +
  1380. theme_jama() +
  1381. theme(legend.position = 'none',
  1382. axis.text.x = element_text(angle = 45, hjust = 1, vjust = 1)) +
  1383. scale_color_jama() +
  1384. xlab('') +
  1385. ylab('Posterior class membership probability') +
  1386. ylim(0,1)
  1387. ```
  1388. **Distribution of posterior class membership probabilities by assigned latent class, indicating model confidence in class assignment.** This plot shows the degree of certainty in class assignments, a good indicator of model reliability. Mean posterior probabilities were `r pp_sum[[1,2]]`, `r pp_sum[[2,2]]`, and `r pp_sum[[3,2]]` for non-, slow-, and fast decliners. With most participants having high posterior probabilities ($>0.8$) for their assigned class, class separation is strong and predictive discrimination is good.
  1389. :::
  1390. \clearpage
  1391. ```{r cross-val-base-setup}
  1392. # Cross-Validated Classification Performance for LCMM
  1393. # Create strata ----
  1394. subject_ids <- qs_pacc %>%
  1395. select(id, Ptau217, `Class (Base)`) %>%
  1396. distinct() %>%
  1397. mutate(
  1398. Ptau217_f = case_when(
  1399. Ptau217 < 0.192 ~ 'low ptau',
  1400. TRUE ~ 'high ptau'),
  1401. strata = paste(Ptau217_f, `Class (Base)`))
  1402. # 10-FOLD CROSS-VALIDATION SETUP ----
  1403. n_folds <- 10
  1404. n_classes <- 3
  1405. # Create folds ----
  1406. set.seed(20251111)
  1407. folds <- createFolds(subject_ids$strata, k = n_folds, list = TRUE)
  1408. ```
  1409. ```{r cross-val-base-model-fits, eval = UPDATELCMMCV}
  1410. model_fits <- parallel::mclapply(1:n_folds, function(fold){
  1411. train_subjects <- subject_ids$id[-folds[[fold]]]
  1412. train_data <- base_model_data %>% filter(id %in% train_subjects)
  1413. hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  1414. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  1415. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  1416. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  1417. random = ~ 1, subject = 'id', ng = 3, B = m1,
  1418. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  1419. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z,
  1420. data = train_data)},
  1421. mc.cores = 5)
  1422. save(model_fits, file = 'cross-val-base-model-fits.rdata')
  1423. ```
  1424. ```{r cross-val-base-aggregate, include = FALSE}
  1425. load('cross-val-base-model-fits.rdata')
  1426. # Initialize storage for results ----
  1427. cv_results <- list()
  1428. all_predictions <- data.frame()
  1429. roc_data <- list()
  1430. for(fold in 1:n_folds) {
  1431. # Split data by subjects (not observations)
  1432. test_subjects <- subject_ids$id[folds[[fold]]]
  1433. train_subjects <- subject_ids$id[-folds[[fold]]]
  1434. train_data <- base_model_data %>% filter(id %in% train_subjects)
  1435. # Use only baseline for test data ----
  1436. test_data <- base_model_data %>% filter(id %in% test_subjects) %>%
  1437. arrange(id, ADURW) %>%
  1438. filter(!duplicated(id))
  1439. lcmm_train <- model_fits[[fold]]
  1440. # Predict on test set
  1441. # Get predictions for test subjects ----
  1442. test_post <- predictClass(lcmm_train, newdata = test_data)
  1443. # Label classes non, slow, fast decliners ----
  1444. labs_data <- predictY(lcmm_train,
  1445. newdata = base_model_data[1, ] %>% mutate(ADURW = 240),
  1446. var.time="ADURW")$pred
  1447. fast <- which.min(labs_data[1,])
  1448. non <- which.max(labs_data[1,])
  1449. slow <- setdiff(1:3, c(fast, non))
  1450. labs <- tibble(
  1451. class_raw = colnames(labs_data)[c(non, slow, fast)],
  1452. class_number = gsub('Ypred_class', '', class_raw) %>% as.numeric(),
  1453. predicted_class = factor(levels(qs_pacc$`Class (Base)`),
  1454. levels = levels(qs_pacc$`Class (Base)`)))
  1455. # Extract predictions
  1456. fold_predictions <- test_post %>%
  1457. rename(class_number = class) %>%
  1458. left_join(labs, by = 'class_number') %>%
  1459. left_join(qs_pacc %>%
  1460. select(id, `Class (Base)`) %>%
  1461. distinct(), by = 'id') %>%
  1462. mutate(fold = fold)
  1463. # with(fold_predictions, table(`Class (Base)`, predicted_class))
  1464. # Store results
  1465. all_predictions <- bind_rows(all_predictions, fold_predictions)
  1466. # Calculate fold-specific metrics
  1467. prob_cols <- grep("^prob", names(fold_predictions), value = TRUE)
  1468. # Store for ROC data (one-vs-rest approach)
  1469. for(k in 1:n_classes) {
  1470. roc_data[[length(roc_data) + 1]] <- data.frame(
  1471. fold = fold,
  1472. class_number = k,
  1473. class_label = labs %>% filter(class_number == k) %>%
  1474. pull(predicted_class),
  1475. predicted_class = fold_predictions$predicted_class,
  1476. true_class = fold_predictions$`Class (Base)`,
  1477. predicted_prob = fold_predictions[[paste0("prob", k)]]
  1478. )
  1479. }
  1480. # Calculate accuracy for this fold
  1481. accuracy <- mean(fold_predictions$predicted_class == fold_predictions$`Class (Base)`)
  1482. cv_results[[fold]] <- list(
  1483. fold = fold,
  1484. accuracy = accuracy,
  1485. n_test = nrow(test_post)
  1486. )
  1487. }
  1488. # Convert results to data frame ----
  1489. cv_summary <- do.call(rbind, lapply(cv_results, as.data.frame))
  1490. cat("=== Overall Performance Metrics ===\n")
  1491. cat(sprintf("Mean Accuracy: %.3f (SD = %.3f)\n",
  1492. mean(cv_summary$accuracy), sd(cv_summary$accuracy)))
  1493. cat(sprintf("95%% CI: [%.3f, %.3f]\n",
  1494. mean(cv_summary$accuracy) - 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds),
  1495. mean(cv_summary$accuracy) + 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds)))
  1496. # Confusion matrix
  1497. confusion_matrix <- table(
  1498. Predicted = all_predictions$predicted_class,
  1499. Actual = all_predictions$`Class (Base)`
  1500. )
  1501. cat("\n=== Confusion Matrix ===\n")
  1502. print(confusion_matrix)
  1503. # Calculate per-class metrics
  1504. cat("\n=== Per-Class Metrics ===\n")
  1505. for(k in 1:n_classes) {
  1506. tp <- confusion_matrix[k, k]
  1507. fp <- sum(confusion_matrix[k, ]) - tp
  1508. fn <- sum(confusion_matrix[, k]) - tp
  1509. tn <- sum(confusion_matrix) - tp - fp - fn
  1510. sensitivity <- tp / (tp + fn)
  1511. specificity <- tn / (tn + fp)
  1512. ppv <- tp / (tp + fp)
  1513. npv <- tn / (tn + fn)
  1514. cat(sprintf("\nClass %d:\n", k))
  1515. cat(sprintf(" Sensitivity: %.3f\n", sensitivity))
  1516. cat(sprintf(" Specificity: %.3f\n", specificity))
  1517. cat(sprintf(" PPV: %.3f\n", ppv))
  1518. cat(sprintf(" NPV: %.3f\n", npv))
  1519. }
  1520. ```
  1521. ::: {#suppfig-cross-val-base-pr}
  1522. ```{r suppfig-cross-val-base-pr}
  1523. pr_df <- do.call(rbind, roc_data)
  1524. # Calculate PR curves and AUCs for each class
  1525. pr_list <- list()
  1526. pr_auc_values <- numeric(n_classes)
  1527. for(k in 1:n_classes) {
  1528. lab <- levels(pr_df$class_label)[k]
  1529. bin_class_data <- pr_df %>%
  1530. filter(class_label == lab) %>%
  1531. mutate(
  1532. pred_bin_class = case_when(
  1533. predicted_class == lab ~ 1,
  1534. TRUE ~ 0),
  1535. true_bin_class = case_when(
  1536. true_class == lab ~ 1,
  1537. TRUE ~ 0))
  1538. pr_obj <- pr.curve(
  1539. scores.class0 = bin_class_data %>% pull(predicted_prob),
  1540. weights.class0 = bin_class_data %>% pull(true_bin_class),
  1541. curve = TRUE)
  1542. pr_list[[k]] <- data.frame(
  1543. recall = pr_obj$curve[,1],
  1544. precision = pr_obj$curve[,2],
  1545. threshold = pr_obj$curve[,3],
  1546. class = lab,
  1547. auc = pr_obj$auc.integral)
  1548. pr_auc_values[k] <- pr_obj$auc.integral
  1549. }
  1550. # Either fast- or slow ----
  1551. bin_class_data <- pr_df %>%
  1552. filter(class_label == "Stable") %>%
  1553. mutate(
  1554. pred_bin_class = case_when(
  1555. predicted_class == "Stable" ~ 0,
  1556. TRUE ~ 1),
  1557. true_bin_class = case_when(
  1558. true_class == "Stable" ~ 0,
  1559. TRUE ~ 1))
  1560. pr_obj <- pr.curve(
  1561. scores.class0 = 1-bin_class_data %>% pull(predicted_prob),
  1562. weights.class0 = bin_class_data %>% pull(true_bin_class),
  1563. curve = TRUE)
  1564. pr_either <- data.frame(
  1565. recall = pr_obj$curve[,1],
  1566. precision = pr_obj$curve[,2],
  1567. threshold = pr_obj$curve[,3],
  1568. class = "Either decliner",
  1569. auc = pr_obj$auc.integral)
  1570. pr_plot_data <- do.call(rbind, pr_list) %>%
  1571. bind_rows(pr_either) %>%
  1572. mutate(Class = paste0(class, " (AUPRC = ",
  1573. format(round(auc, 3), nsmall = 3), ")"))
  1574. pr_plot_data$Class <- factor(pr_plot_data$Class,
  1575. levels = unique(pr_plot_data$Class))
  1576. ggplot(pr_plot_data,
  1577. aes(x = recall, y = precision, color = Class)) +
  1578. geom_line(linewidth = 1.2) +
  1579. geom_abline(intercept = 1, slope = -1, linetype = "dashed",
  1580. color = "gray50", linewidth = 0.5) +
  1581. theme_jama() +
  1582. theme(legend.position = 'right') +
  1583. scale_color_jama() +
  1584. labs(
  1585. x = "Recall (TP/(TP+FN))",
  1586. y = "Precision (TP/(TP+FP))",
  1587. title = "Cross-Validated PR Curves for LCMM Class Prediction",
  1588. subtitle = sprintf("10-fold cross-validation (N = %d subjects)",
  1589. length(unique(base_model_data$id)))) +
  1590. coord_equal()
  1591. ```
  1592. **Cross-validated precision-recall curves for latent class membership model without tau PET.** Precision-recall curves summarizing 10-fold cross-validated classification performance of the latent class membership submodel based on baseline demographics and biomarkers, excluding tau PET. Curves depict the ability to distinguish the given class from the other two. The area under the precision-recall curve (AUPRC) quantifies overall predictive performance, emphasizing recall (or sensitivity, TP/(TP+FN)) and precision (TP/(TP+FP)) for less prevalent classes. **Abbreviations:** AUPRC = area under the precision-recall curve; TP = true positive rate; FN = false negative rate; FP = false positive rate.
  1593. :::
  1594. ```{r cross-val-base-roc, eval = FALSE}
  1595. roc_df <- do.call(rbind, roc_data)
  1596. # Calculate ROC curves and AUCs for each class
  1597. roc_list <- list()
  1598. auc_values <- numeric(n_classes)
  1599. for(k in 1:n_classes) {
  1600. lab <- levels(roc_df$class_label)[k]
  1601. bin_class_data <- roc_df %>%
  1602. filter(class_label == lab) %>%
  1603. mutate(
  1604. pred_bin_class = case_when(
  1605. predicted_class == lab ~ 1,
  1606. TRUE ~ 0),
  1607. true_bin_class = case_when(
  1608. true_class == lab ~ 1,
  1609. TRUE ~ 0))
  1610. roc_obj <- roc(bin_class_data$true_bin_class, bin_class_data$predicted_prob,
  1611. quiet = TRUE)
  1612. roc_list[[k]] <- data.frame(
  1613. sensitivity = roc_obj$sensitivities,
  1614. specificity = roc_obj$specificities,
  1615. class = lab,
  1616. auc = as.numeric(auc(roc_obj))
  1617. )
  1618. auc_values[k] <- as.numeric(auc(roc_obj))
  1619. }
  1620. roc_plot_data <- do.call(rbind, roc_list) %>%
  1621. mutate(Class = paste0(class, " (AUC = ",
  1622. format(round(auc, 3), nsmall = 3), ")"))
  1623. roc_plot_data$Class <- factor(roc_plot_data$Class,
  1624. levels = unique(roc_plot_data$Class))
  1625. ggplot(roc_plot_data,
  1626. aes(x = 1 - specificity, y = sensitivity, color = Class)) +
  1627. geom_line(linewidth = 1.2) +
  1628. geom_abline(intercept = 0, slope = 1, linetype = "dashed",
  1629. color = "gray50", linewidth = 0.5) +
  1630. theme_jama() +
  1631. theme(legend.position = 'inside', legend.position.inside = c(0.8, 0.2)) +
  1632. scale_color_jama() +
  1633. labs(
  1634. x = "1 - Specificity (False Positive Rate)",
  1635. y = "Sensitivity (True Positive Rate)",
  1636. title = "Cross-Validated ROC Curves for LCMM Class Prediction",
  1637. subtitle = sprintf("10-fold cross-validation (N = %d subjects)",
  1638. length(unique(base_model_data$id)))) +
  1639. coord_equal()
  1640. ```
  1641. ```{r cross-val-base-accuracy, eval = FALSE}
  1642. ggplot(cv_summary, aes(x = factor(fold), y = accuracy)) +
  1643. geom_point(size = 3, shape = 15, color = "black") +
  1644. geom_hline(yintercept = mean(cv_summary$accuracy),
  1645. linetype = "dashed", color = "darkred", linewidth = 1) +
  1646. # Add text labels
  1647. geom_text(aes(label = sprintf("%.2f", accuracy)),
  1648. vjust = -0.8, size = 3) +
  1649. # Add mean accuracy annotation
  1650. annotate("text", x = n_folds - 1, y = mean(cv_summary$accuracy) - 0.1,
  1651. label = sprintf("Mean = %.3f", mean(cv_summary$accuracy)),
  1652. color = "darkred", fontface = "bold") +
  1653. labs(
  1654. x = "Cross-Validation Fold",
  1655. y = "Classification Accuracy",
  1656. title = "Classification Accuracy Across 10 Cross-Validation Folds",
  1657. subtitle = sprintf("Overall accuracy: %.3f (95%% CI: %.3f-%.3f)",
  1658. mean(cv_summary$accuracy),
  1659. mean(cv_summary$accuracy) - 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds),
  1660. mean(cv_summary$accuracy) + 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds))
  1661. ) +
  1662. scale_y_continuous(limits = c(0, 1), breaks = seq(0, 1, 0.2)) +
  1663. theme_jama()
  1664. ```
  1665. ```{r cross-val-base-confusion-heatmap, eval = FALSE}
  1666. confusion_df <- as.data.frame(confusion_matrix) %>%
  1667. mutate(
  1668. Proportion = Freq / sum(Freq),
  1669. RowProp = ave(Freq, Predicted, FUN = function(x) x / sum(x))
  1670. )
  1671. ggplot(confusion_df,
  1672. aes(x = Actual, y = Predicted, fill = RowProp)) +
  1673. geom_tile(color = "white", linewidth = 1) +
  1674. geom_text(aes(label = Freq), size = 5, fontface = "bold") +
  1675. scale_fill_gradient(low = "white", high = "#56B4E9",
  1676. name = "Proportion\n(by row)",
  1677. labels = scales::percent) +
  1678. labs(
  1679. x = "Actual Class",
  1680. y = "Predicted Class",
  1681. title = "Confusion Matrix: Cross-Validated LCMM Class Predictions",
  1682. subtitle = sprintf("Overall accuracy: %.1f%%",
  1683. 100 * mean(cv_summary$accuracy))
  1684. ) +
  1685. coord_equal() +
  1686. theme_minimal(base_size = 11) +
  1687. theme(
  1688. plot.title = element_text(size = 12, face = "bold"),
  1689. plot.subtitle = element_text(size = 10),
  1690. axis.text = element_text(size = 10, face = "bold"),
  1691. legend.position = "right",
  1692. panel.grid = element_blank()
  1693. )
  1694. ```
  1695. ```{r cross-val-base-pre-class-perf, eval = FALSE}
  1696. # Calculate per-class metrics for plotting
  1697. class_metrics <- data.frame()
  1698. for(k in 1:n_classes) {
  1699. tp <- confusion_matrix[k, k]
  1700. fp <- sum(confusion_matrix[k, ]) - tp
  1701. fn <- sum(confusion_matrix[, k]) - tp
  1702. tn <- sum(confusion_matrix) - tp - fp - fn
  1703. class_metrics <- rbind(class_metrics, data.frame(
  1704. Class = colnames(confusion_matrix)[k],
  1705. Metric = c("Sensitivity", "Specificity", "PPV", "NPV"),
  1706. Value = c(
  1707. tp / (tp + fn),
  1708. tn / (tn + fp),
  1709. tp / (tp + fp),
  1710. tn / (tn + fn)
  1711. )
  1712. ))
  1713. }
  1714. class_metrics$Class <- factor(class_metrics$Class,
  1715. levels = colnames(confusion_matrix))
  1716. ggplot(class_metrics,
  1717. aes(x = Class, y = Value, fill = Metric)) +
  1718. geom_col(position = "dodge", color = "black", width = 0.7) +
  1719. geom_hline(yintercept = 0.8, linetype = "dashed",
  1720. color = "darkred", linewidth = 0.5) +
  1721. scale_fill_manual(values = c("Sensitivity" = "#E69F00",
  1722. "Specificity" = "#56B4E9",
  1723. "PPV" = "#009E73",
  1724. "NPV" = "#F0E442")) +
  1725. labs(
  1726. x = "Latent Class",
  1727. y = "Metric Value",
  1728. title = "Per-Class Performance Metrics from Cross-Validation",
  1729. fill = "Metric"
  1730. ) +
  1731. scale_y_continuous(limits = c(0, 1), breaks = seq(0, 1, 0.2)) +
  1732. theme_jama() +
  1733. theme(legend.position = "top")
  1734. ```
  1735. \clearpage
  1736. ```{r cross-val-tau-pet-setup}
  1737. # Cross-Validated Classification Performance for LCMM
  1738. # Create strata ----
  1739. subject_ids <- qs_pacc %>%
  1740. filter(!is.na(Tau_PET)) %>%
  1741. select(id, Tau_PET, `Class (Base)`) %>%
  1742. distinct() %>%
  1743. mutate(
  1744. Tau_PET_f = case_when(
  1745. Tau_PET < 1.091732 ~ 'low ptau',
  1746. TRUE ~ 'high ptau'),
  1747. strata = paste(Tau_PET_f, `Class (Base)`))
  1748. # 10-FOLD CROSS-VALIDATION SETUP ----
  1749. n_folds <- 10
  1750. n_classes <- 3
  1751. # Create folds ----
  1752. set.seed(20251111)
  1753. folds <- createFolds(subject_ids$strata, k = n_folds, list = TRUE)
  1754. ```
  1755. ```{r cross-val-tau-model-fits, eval = UPDATELCMMCV}
  1756. model_fits_tau <- parallel::mclapply(1:n_folds, function(fold){
  1757. train_subjects <- subject_ids$id[-folds[[fold]]]
  1758. train_data <- tau_model_data %>% filter(id %in% train_subjects)
  1759. hlme(Z ~ I(ns21(ADURW)) + I(ns22(ADURW)) +
  1760. Female + Sola + AAPOEGNPRSNFLG + QSVERSION +
  1761. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  1762. mixture = ~ (I(ns21(ADURW)) + I(ns22(ADURW))),
  1763. random = ~ 1, subject = 'id', ng = 3, B = mt1,
  1764. classmb = ~ Female + Sola + AAPOEGNPRSNFLG +
  1765. Ptau217_z + SUVRCER_z + AGEYR_z + EDCCNTU_z + Hipp_atrophy_z + Tau_PET_z,
  1766. data = train_data)},
  1767. mc.cores = 5)
  1768. save(model_fits_tau, file = 'cross-val-tau-pet-model-fits.rdata')
  1769. ```
  1770. ```{r cross-val-tau-aggregate, include = FALSE}
  1771. load('cross-val-tau-pet-model-fits.rdata')
  1772. # Initialize storage for results ----
  1773. cv_results <- list()
  1774. all_predictions <- data.frame()
  1775. roc_data <- list()
  1776. for(fold in 1:n_folds) {
  1777. # Split data by subjects (not observations)
  1778. test_subjects <- subject_ids$id[folds[[fold]]]
  1779. train_subjects <- subject_ids$id[-folds[[fold]]]
  1780. train_data <- tau_model_data %>% filter(id %in% train_subjects)
  1781. # Use only baseline for test data ----
  1782. test_data <- tau_model_data %>% filter(id %in% test_subjects) %>%
  1783. arrange(id, ADURW) %>%
  1784. filter(!duplicated(id))
  1785. lcmm_train <- model_fits_tau[[fold]]
  1786. # Predict on test set
  1787. # Get predictions for test subjects ----
  1788. test_post <- predictClass(lcmm_train, newdata = test_data)
  1789. # Label classes non, slow, fast decliners ----
  1790. labs_data <- predictY(lcmm_train,
  1791. newdata = tau_model_data[1, ] %>% mutate(ADURW = 240),
  1792. var.time="ADURW")$pred
  1793. fast <- which.min(labs_data[1,])
  1794. non <- which.max(labs_data[1,])
  1795. slow <- setdiff(1:3, c(fast, non))
  1796. labs <- tibble(
  1797. class_raw = colnames(labs_data)[c(non, slow, fast)],
  1798. class_number = gsub('Ypred_class', '', class_raw) %>% as.numeric(),
  1799. predicted_class = factor(c('Stable', 'Slow decliner', 'Fast decliner'),
  1800. levels = levels(qs_pacc$`Class (Base)`)))
  1801. # Extract predictions
  1802. fold_predictions <- test_post %>%
  1803. rename(class_number = class) %>%
  1804. left_join(labs, by = 'class_number') %>%
  1805. left_join(qs_pacc %>%
  1806. select(id, `Class (Base)`) %>%
  1807. distinct(), by = 'id') %>%
  1808. mutate(fold = fold)
  1809. # with(fold_predictions, table(`Class (Base)`, predicted_class))
  1810. # Store results
  1811. all_predictions <- bind_rows(all_predictions, fold_predictions)
  1812. # Calculate fold-specific metrics
  1813. prob_cols <- grep("^prob", names(fold_predictions), value = TRUE)
  1814. # Store for ROC data (one-vs-rest approach)
  1815. for(k in 1:n_classes) {
  1816. roc_data[[length(roc_data) + 1]] <- data.frame(
  1817. fold = fold,
  1818. class_number = k,
  1819. class_label = labs %>% filter(class_number == k) %>%
  1820. pull(predicted_class),
  1821. predicted_class = fold_predictions$predicted_class,
  1822. true_class = fold_predictions$`Class (Base)`,
  1823. predicted_prob = fold_predictions[[paste0("prob", k)]]
  1824. )
  1825. }
  1826. # Calculate accuracy for this fold
  1827. accuracy <- mean(fold_predictions$predicted_class == fold_predictions$`Class (Base)`)
  1828. cv_results[[fold]] <- list(
  1829. fold = fold,
  1830. accuracy = accuracy,
  1831. n_test = nrow(test_post)
  1832. )
  1833. }
  1834. # Convert results to data frame ----
  1835. cv_summary <- do.call(rbind, lapply(cv_results, as.data.frame))
  1836. cat("=== Overall Performance Metrics ===\n")
  1837. cat(sprintf("Mean Accuracy: %.3f (SD = %.3f)\n",
  1838. mean(cv_summary$accuracy), sd(cv_summary$accuracy)))
  1839. cat(sprintf("95%% CI: [%.3f, %.3f]\n",
  1840. mean(cv_summary$accuracy) - 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds),
  1841. mean(cv_summary$accuracy) + 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds)))
  1842. # Confusion matrix
  1843. confusion_matrix <- table(
  1844. Predicted = all_predictions$predicted_class,
  1845. Actual = all_predictions$`Class (Base)`
  1846. )
  1847. cat("\n=== Confusion Matrix ===\n")
  1848. print(confusion_matrix)
  1849. # Calculate per-class metrics
  1850. cat("\n=== Per-Class Metrics ===\n")
  1851. for(k in 1:n_classes) {
  1852. tp <- confusion_matrix[k, k]
  1853. fp <- sum(confusion_matrix[k, ]) - tp
  1854. fn <- sum(confusion_matrix[, k]) - tp
  1855. tn <- sum(confusion_matrix) - tp - fp - fn
  1856. sensitivity <- tp / (tp + fn)
  1857. specificity <- tn / (tn + fp)
  1858. ppv <- tp / (tp + fp)
  1859. npv <- tn / (tn + fn)
  1860. cat(sprintf("\nClass %d:\n", k))
  1861. cat(sprintf(" Sensitivity: %.3f\n", sensitivity))
  1862. cat(sprintf(" Specificity: %.3f\n", specificity))
  1863. cat(sprintf(" PPV: %.3f\n", ppv))
  1864. cat(sprintf(" NPV: %.3f\n", npv))
  1865. }
  1866. ```
  1867. ::: {#suppfig-cross-val-tau-pet-pr}
  1868. ```{r suppfig-cross-val-tau-pet-pr}
  1869. pr_df <- do.call(rbind, roc_data)
  1870. # Calculate PR curves and AUCs for each class
  1871. pr_list <- list()
  1872. pr_auc_values <- numeric(n_classes)
  1873. for(k in 1:n_classes) {
  1874. lab <- levels(pr_df$class_label)[k]
  1875. bin_class_data <- pr_df %>%
  1876. filter(class_label == lab) %>%
  1877. mutate(
  1878. pred_bin_class = case_when(
  1879. predicted_class == lab ~ 1,
  1880. TRUE ~ 0),
  1881. true_bin_class = case_when(
  1882. true_class == lab ~ 1,
  1883. TRUE ~ 0))
  1884. pr_obj <- pr.curve(
  1885. scores.class0 = bin_class_data %>% pull(predicted_prob),
  1886. weights.class0 = bin_class_data %>% pull(true_bin_class),
  1887. curve = TRUE)
  1888. pr_list[[k]] <- data.frame(
  1889. recall = pr_obj$curve[,1],
  1890. precision = pr_obj$curve[,2],
  1891. threshold = pr_obj$curve[,3],
  1892. class = lab,
  1893. auc = pr_obj$auc.integral)
  1894. pr_auc_values[k] <- pr_obj$auc.integral
  1895. }
  1896. # Either fast- or slow ----
  1897. bin_class_data <- pr_df %>%
  1898. filter(class_label == "Stable") %>%
  1899. mutate(
  1900. pred_bin_class = case_when(
  1901. predicted_class == "Stable" ~ 0,
  1902. TRUE ~ 1),
  1903. true_bin_class = case_when(
  1904. true_class == "Stable" ~ 0,
  1905. TRUE ~ 1))
  1906. pr_obj <- pr.curve(
  1907. scores.class0 = 1-bin_class_data %>% pull(predicted_prob),
  1908. weights.class0 = bin_class_data %>% pull(true_bin_class),
  1909. curve = TRUE)
  1910. pr_either <- data.frame(
  1911. recall = pr_obj$curve[,1],
  1912. precision = pr_obj$curve[,2],
  1913. threshold = pr_obj$curve[,3],
  1914. class = "Either decliner",
  1915. auc = pr_obj$auc.integral)
  1916. pr_plot_data <- do.call(rbind, pr_list) %>%
  1917. bind_rows(pr_either) %>%
  1918. mutate(Class = paste0(class, " (AUPRC = ",
  1919. format(round(auc, 3), nsmall = 3), ")"))
  1920. pr_plot_data$Class <- factor(pr_plot_data$Class,
  1921. levels = unique(pr_plot_data$Class))
  1922. ggplot(pr_plot_data,
  1923. aes(x = recall, y = precision, color = Class)) +
  1924. geom_line(linewidth = 1.2) +
  1925. geom_abline(intercept = 1, slope = -1, linetype = "dashed",
  1926. color = "gray50", linewidth = 0.5) +
  1927. theme_jama() +
  1928. theme(legend.position = 'right') +
  1929. scale_color_jama() +
  1930. labs(
  1931. x = "Recall (TP/(TP+FN))",
  1932. y = "Precision (TP/(TP+FP))",
  1933. title = "Cross-Validated PR Curves for LCMM Class Prediction",
  1934. subtitle = sprintf("10-fold cross-validation (N = %d subjects)",
  1935. length(unique(tau_model_data$id)))) +
  1936. coord_equal()
  1937. ```
  1938. **Cross-validated precision-recall curves for latent class membership model with tau PET.** Precision-recall curves summarizing 10-fold cross-validated classification performance of the latent class membership submodel based on baseline demographics and biomarkers, excluding tau PET. Curves depict the ability to distinguish the given class from the other two. The area under the precision-recall curve (AUPRC) quantifies overall predictive performance, emphasizing recall (or sensitivity, TP/(TP+FN)) and precision (TP/(TP+FP)) for less prevalent classes. **Abbreviations:** AUPRC = area under the precision-recall curve; TP = true positive rate; FN = false negative rate; FP = false positive rate.
  1939. :::
  1940. \clearpage
  1941. ```{r cross-val-tau-roc, eval = FALSE}
  1942. roc_df <- do.call(rbind, roc_data)
  1943. # Calculate ROC curves and AUCs for each class
  1944. roc_list <- list()
  1945. auc_values <- numeric(n_classes)
  1946. for(k in 1:n_classes) {
  1947. lab <- unique(roc_df$class_label)[k]
  1948. bin_class_data <- roc_df %>%
  1949. filter(class_label == lab) %>%
  1950. mutate(
  1951. pred_bin_class = case_when(
  1952. predicted_class == lab ~ 1,
  1953. TRUE ~ 0),
  1954. true_bin_class = case_when(
  1955. true_class == lab ~ 1,
  1956. TRUE ~ 0))
  1957. roc_obj <- roc(bin_class_data$true_bin_class, bin_class_data$predicted_prob,
  1958. quiet = TRUE)
  1959. roc_list[[k]] <- data.frame(
  1960. sensitivity = roc_obj$sensitivities,
  1961. specificity = roc_obj$specificities,
  1962. class = lab,
  1963. auc = as.numeric(auc(roc_obj))
  1964. )
  1965. auc_values[k] <- as.numeric(auc(roc_obj))
  1966. }
  1967. roc_plot_data <- do.call(rbind, roc_list) %>%
  1968. # mutate(Class = paste0(class, " (AUC = ",
  1969. # format(round(auc, 3), nsmall = 3), ")"))
  1970. # table(roc_plot_data$Class)
  1971. mutate(Class = factor(paste0(class, " (AUC = ",
  1972. format(round(auc, 3), nsmall = 3), ")"),
  1973. levels = c("Stable (AUC = 0.841)",
  1974. "Slow decliner (AUC = 0.730)",
  1975. "Fast decliner (AUC = 0.872)")))
  1976. ggplot(roc_plot_data,
  1977. aes(x = 1 - specificity, y = sensitivity, color = Class)) +
  1978. geom_line(linewidth = 1.2) +
  1979. geom_abline(intercept = 0, slope = 1, linetype = "dashed",
  1980. color = "gray50", linewidth = 0.5) +
  1981. theme_jama() +
  1982. theme(legend.position = 'inside', legend.position.inside = c(0.8, 0.2)) +
  1983. scale_color_jama() +
  1984. labs(
  1985. x = "1 - Specificity (False Positive Rate)",
  1986. y = "Sensitivity (True Positive Rate)",
  1987. title = "Cross-Validated ROC Curves for LCMM Class Prediction",
  1988. subtitle = sprintf("10-fold cross-validation (N = %d subjects)",
  1989. length(unique(tau_model_data$id)))) +
  1990. coord_equal()
  1991. ```
  1992. ```{r cross-val-tau-accuracy, eval = FALSE}
  1993. ggplot(cv_summary, aes(x = factor(fold), y = accuracy)) +
  1994. geom_point(size = 3, shape = 15, color = "black") +
  1995. geom_hline(yintercept = mean(cv_summary$accuracy),
  1996. linetype = "dashed", color = "darkred", linewidth = 1) +
  1997. # Add text labels
  1998. geom_text(aes(label = sprintf("%.2f", accuracy)),
  1999. vjust = -0.8, size = 3) +
  2000. # Add mean accuracy annotation
  2001. annotate("text", x = n_folds - 1, y = mean(cv_summary$accuracy) - 0.1,
  2002. label = sprintf("Mean = %.3f", mean(cv_summary$accuracy)),
  2003. color = "darkred", fontface = "bold") +
  2004. labs(
  2005. x = "Cross-Validation Fold",
  2006. y = "Classification Accuracy",
  2007. title = "Classification Accuracy Across 10 Cross-Validation Folds",
  2008. subtitle = sprintf("Overall accuracy: %.3f (95%% CI: %.3f-%.3f)",
  2009. mean(cv_summary$accuracy),
  2010. mean(cv_summary$accuracy) - 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds),
  2011. mean(cv_summary$accuracy) + 1.96 * sd(cv_summary$accuracy) / sqrt(n_folds))
  2012. ) +
  2013. scale_y_continuous(limits = c(0, 1), breaks = seq(0, 1, 0.2)) +
  2014. theme_jama()
  2015. ```
  2016. ```{r cross-val-tau-confusion-heatmap, eval = FALSE}
  2017. confusion_df <- as.data.frame(confusion_matrix) %>%
  2018. mutate(
  2019. Proportion = Freq / sum(Freq),
  2020. RowProp = ave(Freq, Predicted, FUN = function(x) x / sum(x))
  2021. )
  2022. ggplot(confusion_df,
  2023. aes(x = Actual, y = Predicted, fill = RowProp)) +
  2024. geom_tile(color = "white", linewidth = 1) +
  2025. geom_text(aes(label = Freq), size = 5, fontface = "bold") +
  2026. scale_fill_gradient(low = "white", high = "#56B4E9",
  2027. name = "Proportion\n(by row)",
  2028. labels = scales::percent) +
  2029. labs(
  2030. x = "Actual Class",
  2031. y = "Predicted Class",
  2032. title = "Confusion Matrix: Cross-Validated LCMM Class Predictions",
  2033. subtitle = sprintf("Overall accuracy: %.1f%%",
  2034. 100 * mean(cv_summary$accuracy))
  2035. ) +
  2036. coord_equal() +
  2037. theme_minimal(base_size = 11) +
  2038. theme(
  2039. plot.title = element_text(size = 12, face = "bold"),
  2040. plot.subtitle = element_text(size = 10),
  2041. axis.text = element_text(size = 10, face = "bold"),
  2042. legend.position = "right",
  2043. panel.grid = element_blank()
  2044. )
  2045. ```
  2046. ```{r cross-val-tau-pre-class-perf, eval = FALSE}
  2047. # Calculate per-class metrics for plotting
  2048. class_metrics <- data.frame()
  2049. for(k in 1:n_classes) {
  2050. tp <- confusion_matrix[k, k]
  2051. fp <- sum(confusion_matrix[k, ]) - tp
  2052. fn <- sum(confusion_matrix[, k]) - tp
  2053. tn <- sum(confusion_matrix) - tp - fp - fn
  2054. class_metrics <- rbind(class_metrics, data.frame(
  2055. Class = colnames(confusion_matrix)[k],
  2056. Metric = c("Sensitivity", "Specificity", "PPV", "NPV"),
  2057. Value = c(
  2058. tp / (tp + fn),
  2059. tn / (tn + fp),
  2060. tp / (tp + fp),
  2061. tn / (tn + fn)
  2062. )
  2063. ))
  2064. }
  2065. class_metrics$Class <- factor(class_metrics$Class,
  2066. levels = colnames(confusion_matrix))
  2067. ggplot(class_metrics,
  2068. aes(x = Class, y = Value, fill = Metric)) +
  2069. geom_col(position = "dodge", color = "black", width = 0.7) +
  2070. geom_hline(yintercept = 0.8, linetype = "dashed",
  2071. color = "darkred", linewidth = 0.5) +
  2072. scale_fill_manual(values = c("Sensitivity" = "#E69F00",
  2073. "Specificity" = "#56B4E9",
  2074. "PPV" = "#009E73",
  2075. "NPV" = "#F0E442")) +
  2076. labs(
  2077. x = "Latent Class",
  2078. y = "Metric Value",
  2079. title = "Per-Class Performance Metrics from Cross-Validation",
  2080. fill = "Metric"
  2081. ) +
  2082. scale_y_continuous(limits = c(0, 1), breaks = seq(0, 1, 0.2)) +
  2083. theme_jama() +
  2084. theme(legend.position = "top")
  2085. ```
  2086. ```{r rpart_tune_setup, include=FALSE}
  2087. # Tuning Tree Hyperparameters with tidymodels ----
  2088. tree_data <- qs_pacc %>%
  2089. filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)) %>%
  2090. select(Class = `Class (Base)`, Solanezumab = Sola, Age = AGEYR, Sex,
  2091. Education = EDCCNTU, APOEe4, PACC = Y, `Amyloid PET` = AMYLCENT, Ptau217,
  2092. `Hipp atrophy` = Hipp_atrophy_z)
  2093. # Set seed for reproducibility
  2094. set.seed(20251113)
  2095. # Create cross-validation folds (stratified by class)
  2096. cv_folds <- vfold_cv(tree_data, v = 10, strata = Class)
  2097. # Define the model specification with tuning parameters
  2098. tree_spec <- decision_tree(
  2099. cost_complexity = tune(), # cp parameter
  2100. tree_depth = tune(), # maxdepth parameter
  2101. min_n = tune() # minsplit parameter
  2102. ) %>%
  2103. set_engine("rpart") %>%
  2104. set_mode("classification")
  2105. # Create recipe (preprocessing)
  2106. tree_recipe <- recipe(
  2107. Class ~ Solanezumab + Age + Sex + Education + APOEe4 +
  2108. PACC + `Amyloid PET` + Ptau217 + `Hipp atrophy`,
  2109. data = tree_data
  2110. ) %>%
  2111. step_zv(all_predictors()) # Remove zero-variance predictors if any
  2112. # Create workflow
  2113. tree_workflow <- workflow() %>%
  2114. add_model(tree_spec) %>%
  2115. add_recipe(tree_recipe)
  2116. # Define tuning grid
  2117. # Option 1: Regular grid
  2118. tree_grid <- grid_regular(
  2119. cost_complexity(range = c(-5, -1)), # log10 scale: 0.00001 to 0.1
  2120. tree_depth(range = c(1, 4)),
  2121. min_n(range = c(5, 40)),
  2122. levels = 5 # 5 values per parameter = 125 combinations
  2123. )
  2124. ```
  2125. ```{r rpart_tune, eval = UPDATETREECV}
  2126. # Perform tuning ----
  2127. tune_results <- tree_workflow %>%
  2128. tune_grid(
  2129. resamples = cv_folds,
  2130. grid = tree_grid,
  2131. metrics = metric_set(yardstick::accuracy, roc_auc, f_meas, bal_accuracy),
  2132. control = control_grid(save_pred = TRUE, verbose = FALSE)
  2133. )
  2134. # Best models ----
  2135. top_models <- show_best(tune_results, metric = "bal_accuracy", n = 5)
  2136. # Select best model (by accuracy)
  2137. best_params <- select_best(tune_results, metric = "bal_accuracy")
  2138. # Finalize workflow with best parameters
  2139. final_workflow <- tree_workflow %>%
  2140. finalize_workflow(best_params)
  2141. # Fit final model on full dataset
  2142. # final_fit <- final_workflow %>%
  2143. # fit(data = tree_data)
  2144. # Extract the fitted model for plotting
  2145. tree_model <- rpart(Class ~ Solanezumab + Age + Sex + Education + APOEe4 +
  2146. PACC + `Amyloid PET` + Ptau217 + `Hipp atrophy`,
  2147. data = tree_data,
  2148. method = "class",
  2149. control = rpart.control(
  2150. cp = best_params$cost_complexity,
  2151. minsplit = best_params$min_n,
  2152. minbucket = 10, # Minimum observations in terminal node
  2153. maxdepth = best_params$tree_depth)
  2154. )
  2155. save(tune_results, top_models, best_params, final_workflow, tree_model,
  2156. file = 'tree_cv_base_results.rdata')
  2157. ```
  2158. ```{r suppfig-tree-base-cap}
  2159. load('tree_cv_base_results.rdata')
  2160. cat("**Classification tree analysis for latent class assignment from Base Model.** Panel (A) shows relative importance of baseline predictors for class separation. Panel (B) shows classification tree showing binary decision rules for assigning subjects to latent classes based on predictors. Terminal nodes show predicted class, number of subjects, and percentage of total sample. The optimal tree complexity was determined using 10-fold cross-validation with the 1-standard error rule, resulting in a tree with ", sum(tree_model$frame$var != '<leaf>'), " splits (mean balanced accuracy ", format(round(top_models[[1, 'mean']], 2), nsmall = 2), ", standard error ", round(top_models[[1, 'std_err']], 3), ").", sep = '',
  2161. file = "_suppfig-tree-base-cap.txt")
  2162. ```
  2163. ```{r rpart_acc}
  2164. tree_predictions <- predict(tree_model, type = "class")
  2165. tree_probs <- predict(tree_model, type = "prob")
  2166. # Calculate accuracy
  2167. tree_accuracy <- mean(tree_predictions == tree_data$Class)
  2168. ```
  2169. ```{r Variable_importance}
  2170. # Variable importance ----
  2171. var_imp <- tree_model$variable.importance
  2172. if(length(var_imp) > 0) {
  2173. var_imp_df <- data.frame(
  2174. Variable = names(var_imp),
  2175. Importance = var_imp,
  2176. Relative_Importance = 100 * var_imp / sum(var_imp)
  2177. ) %>%
  2178. arrange(desc(Importance))
  2179. # print(var_imp_df, row.names = FALSE)
  2180. } else {
  2181. cat("No splits in tree - all subjects assigned to single class\n")
  2182. }
  2183. ```
  2184. ```{r Confusion_matrix, include=FALSE}
  2185. # Confusion matrix
  2186. conf_matrix <- confusionMatrix(tree_predictions, tree_data$Class)
  2187. print(conf_matrix$table)
  2188. print(conf_matrix$byClass[, c("Sensitivity", "Specificity", "Pos Pred Value")])
  2189. ```
  2190. ```{r make-suppfig-tree-base, include=FALSE}
  2191. png("suppfig-tree-base.png", width = 4.5*1.25, height = 5.75*1.2, units = "in", res = 300)
  2192. # Set up 2-panel layout
  2193. layout(matrix(c(1, 2.5, 2.5), nrow = 3, ncol = 1))
  2194. # Panel A: Variable Importance (left panel - narrower)
  2195. par(mar = c(5, 7, 4, 1))
  2196. # Sort by importance
  2197. var_order <- order(var_imp_df$Importance, decreasing = FALSE)
  2198. var_names <- var_imp_df$Variable[var_order]
  2199. var_values <- var_imp_df$Relative_Importance[var_order]
  2200. # Create horizontal barplot
  2201. bp <- barplot(var_values,
  2202. horiz = TRUE,
  2203. names.arg = var_names,
  2204. las = 1,
  2205. col = "#56B4E9",
  2206. border = "black",
  2207. xlim = c(0, max(var_values) * 1.1),
  2208. xlab = "Relative Importance (%)",
  2209. main = "A. Variable Importance",
  2210. cex.main = 1,
  2211. cex.lab = 1.1,
  2212. cex.axis = 1,
  2213. cex.names = 1)
  2214. # Add value labels
  2215. text(var_values + max(var_values) * 0.02, bp,
  2216. labels = sprintf("%.1f%%", var_values),
  2217. pos = 4, cex = 0.9)
  2218. # Add gridlines
  2219. abline(v = seq(0, max(var_values), by = 10), col = "gray90", lty = 1)
  2220. # Panel B: Classification Tree (right panel)
  2221. par(mar = c(2, 1, 4, 2))
  2222. rpart.plot(
  2223. tree_model,
  2224. type = 4, # All labels, beneath nodes
  2225. extra = 101, # Show class and probability
  2226. under = TRUE,
  2227. fallen.leaves = TRUE,
  2228. branch = 0.5,
  2229. box.palette = as.list(c("lightgrey", jama_colors[2:3])), # Color by class
  2230. cex = 0.75,
  2231. family = "sans",
  2232. faclen = 0,
  2233. split.cex = 0.8,
  2234. split.box.col = "white",
  2235. split.border.col = "black",
  2236. main = "B. Classification Tree",
  2237. cex.main = 1
  2238. )
  2239. dev.off()
  2240. ```
  2241. ```{r rpart_tau_pet_tune_setup, include=FALSE}
  2242. # Tuning Tree Hyperparameters with tidymodels ----
  2243. tree_data <- qs_pacc %>%
  2244. filter(VISITCD == '006' & !is.na(Ptau217) & !is.na(APOEe4)) %>%
  2245. select(Class = `Class (Tau PET)`, Solanezumab = Sola, Age = AGEYR, Sex,
  2246. Education = EDCCNTU, APOEe4, PACC = Y, `Amyloid PET` = AMYLCENT,
  2247. Ptau217, `Hipp atrophy` = Hipp_atrophy_z, `Tau PET` = Tau_PET) %>%
  2248. na.omit()
  2249. # Set seed for reproducibility
  2250. set.seed(20251113)
  2251. # Create cross-validation folds (stratified by class)
  2252. cv_folds <- vfold_cv(tree_data, v = 10, strata = Class)
  2253. # Define the model specification with tuning parameters
  2254. tree_spec <- decision_tree(
  2255. cost_complexity = tune(), # cp parameter
  2256. tree_depth = tune(), # maxdepth parameter
  2257. min_n = tune() # minsplit parameter
  2258. ) %>%
  2259. set_engine("rpart") %>%
  2260. set_mode("classification")
  2261. # Create recipe (preprocessing)
  2262. tree_recipe <- recipe(
  2263. Class ~ Solanezumab + Age + Sex + Education + APOEe4 +
  2264. PACC + `Amyloid PET` + Ptau217 + `Hipp atrophy` + `Tau PET`,
  2265. data = tree_data
  2266. ) %>%
  2267. step_zv(all_predictors()) # Remove zero-variance predictors if any
  2268. # Create workflow
  2269. tree_workflow <- workflow() %>%
  2270. add_model(tree_spec) %>%
  2271. add_recipe(tree_recipe)
  2272. # Define tuning grid
  2273. # Option 1: Regular grid
  2274. tree_grid <- grid_regular(
  2275. cost_complexity(range = c(-5, -1)), # log10 scale: 0.00001 to 0.1
  2276. tree_depth(range = c(1, 4)),
  2277. min_n(range = c(5, 40)),
  2278. levels = 5 # 5 values per parameter = 125 combinations
  2279. )
  2280. ```
  2281. ```{r rpart_tau_pet_tune, eval = UPDATETREECV}
  2282. # Perform tuning ----
  2283. tune_results <- tree_workflow %>%
  2284. tune_grid(
  2285. resamples = cv_folds,
  2286. grid = tree_grid,
  2287. metrics = metric_set(yardstick::accuracy, roc_auc, f_meas, bal_accuracy),
  2288. control = control_grid(save_pred = TRUE, verbose = FALSE)
  2289. )
  2290. # Best models ----
  2291. top_models <- show_best(tune_results, metric = "bal_accuracy", n = 5)
  2292. # Select best model (by accuracy)
  2293. best_params <- select_best(tune_results, metric = "bal_accuracy")
  2294. # Finalize workflow with best parameters
  2295. final_workflow <- tree_workflow %>%
  2296. finalize_workflow(best_params)
  2297. # Fit final model on full dataset
  2298. # final_fit <- final_workflow %>%
  2299. # fit(data = tree_data)
  2300. # Extract the fitted model for plotting
  2301. tree_model <- rpart(Class ~ Solanezumab + Age + Sex + Education + APOEe4 +
  2302. PACC + `Amyloid PET` + Ptau217 + `Hipp atrophy` + `Tau PET`,
  2303. data = tree_data,
  2304. method = "class",
  2305. control = rpart.control(
  2306. cp = best_params$cost_complexity,
  2307. minsplit = best_params$min_n,
  2308. minbucket = 10, # Minimum observations in terminal node
  2309. maxdepth = best_params$tree_depth)
  2310. )
  2311. save(tune_results, top_models, best_params, final_workflow, tree_model,
  2312. file = 'tree_cv_tau_pet_results.rdata')
  2313. ```
  2314. ```{r rpart_tau_pet_acc}
  2315. load('tree_cv_tau_pet_results.rdata')
  2316. tree_predictions <- predict(tree_model, type = "class")
  2317. tree_probs <- predict(tree_model, type = "prob")
  2318. # Calculate accuracy
  2319. tree_accuracy <- mean(tree_predictions == tree_data$Class)
  2320. ```
  2321. ```{r rpart_tau_pet_Variable_importance}
  2322. # Variable importance ----
  2323. var_imp <- tree_model$variable.importance
  2324. if(length(var_imp) > 0) {
  2325. var_imp_df <- data.frame(
  2326. Variable = names(var_imp),
  2327. Importance = var_imp,
  2328. Relative_Importance = 100 * var_imp / sum(var_imp)
  2329. ) %>%
  2330. arrange(desc(Importance))
  2331. # print(var_imp_df, row.names = FALSE)
  2332. } else {
  2333. cat("No splits in tree - all subjects assigned to single class\n")
  2334. }
  2335. ```
  2336. ```{r rpart_tau_pet_Confusion_matrix, include=FALSE}
  2337. # Confusion matrix
  2338. conf_matrix <- confusionMatrix(tree_predictions, tree_data$Class)
  2339. print(conf_matrix$table)
  2340. print(conf_matrix$byClass[, c("Sensitivity", "Specificity", "Pos Pred Value")])
  2341. cat("**Classification tree analysis for latent class assignment from Tau PET Model.** (A) Relative importance of baseline characteristics for class separation. (B) Classification tree showing binary decision rules for assigning subjects to latent classes based on baseline characteristics. Terminal nodes show predicted class, number of subjects, and percentage of total sample. The optimal tree complexity was determined using 10-fold cross-validation with the 1-standard error rule, resulting in a tree with ",
  2342. sum(tree_model$frame$var != '<leaf>'), " splits (mean balanced accuracy ",
  2343. format(round(top_models[[1, 'mean']], 2), nsmall = 2),
  2344. ", standard error ",
  2345. round(top_models[[1, 'std_err']], 3), ").",
  2346. file = "_suppfig-tree-tau-pet-cap.txt")
  2347. ```
  2348. ::: {#suppfig-tree-tau-pet}
  2349. ```{r suppfig-tree-tau-pet, fig.width = 4.5*1.1, fig.height = 5.75*1.1}
  2350. # Set up 2-panel layout
  2351. layout(matrix(c(1, 2, 2), nrow = 3, ncol = 1))
  2352. # Panel A: Variable Importance (left panel - narrower)
  2353. par(mar = c(5, 7, 4, 1))
  2354. # Sort by importance
  2355. var_order <- order(var_imp_df$Importance, decreasing = FALSE)
  2356. var_names <- var_imp_df$Variable[var_order]
  2357. var_values <- var_imp_df$Relative_Importance[var_order]
  2358. # Create horizontal barplot
  2359. bp <- barplot(var_values,
  2360. horiz = TRUE,
  2361. names.arg = var_names,
  2362. las = 1,
  2363. col = "#56B4E9",
  2364. border = "black",
  2365. xlim = c(0, max(var_values) * 1.1),
  2366. xlab = "Relative Importance (%)",
  2367. main = "A. Variable Importance",
  2368. cex.main = 1,
  2369. cex.lab = 1.1,
  2370. cex.axis = 1,
  2371. cex.names = 1)
  2372. # Add value labels
  2373. text(var_values + max(var_values) * 0.02, bp,
  2374. labels = sprintf("%.1f%%", var_values),
  2375. pos = 4, cex = 0.9)
  2376. # Add gridlines
  2377. abline(v = seq(0, max(var_values), by = 10), col = "gray90", lty = 1)
  2378. # Panel B: Classification Tree (right panel)
  2379. par(mar = c(2, 1, 4, 2))
  2380. rpart.plot(
  2381. tree_model,
  2382. type = 4, # All labels, beneath nodes
  2383. extra = 101, # Show class and probability
  2384. under = TRUE,
  2385. fallen.leaves = TRUE,
  2386. digits = 3,
  2387. branch = 0.5,
  2388. box.palette = as.list(c("lightgrey", jama_colors[2:3])), # Color by class
  2389. cex = 0.75,
  2390. family = "sans",
  2391. faclen = 0,
  2392. split.cex = 0.8,
  2393. split.box.col = "white",
  2394. split.border.col = "black",
  2395. main = "B. Classification Tree",
  2396. cex.main = 1
  2397. )
  2398. ```
  2399. {{< include _suppfig-tree-tau-pet-cap.txt >}}
  2400. :::
  2401. \clearpage
  2402. ::: {#supptbl-clda-power}
  2403. ```{r supptbl-clda-power}
  2404. #| results: asis
  2405. #| tbl-colwidths: [10,10,10,10,10,10,10]
  2406. em_effect %>%
  2407. filter(Years %in% c(2, 4), Group != 'Stable Amyloid-') %>%
  2408. mutate(
  2409. N = case_when(
  2410. Years == 2 ~ 500*(1-0.1),
  2411. Years == 4 ~ 500*(1-0.2)),
  2412. `Standard deviation` = sqrt(Variance)) %>%
  2413. rowwise() %>%
  2414. mutate(
  2415. `Power (%)` =
  2416. power.t.test(delta = Delta, n=N, sd=`Standard deviation`)$power*100) %>%
  2417. select(Group, `Follow-up (yrs)` = Years, `Mean PACC untreated` = emmean,
  2418. `Standard deviation`, `Delta (%)`, `Delta (PACC points)` = Delta, `Power (%)`) %>%
  2419. kbl("pipe", booktabs = T, digit = 2)
  2420. ```
  2421. **Power for latent class-specific clinical trials.** Effect size estimates were derived from longitudinal models using natural cubic splines (two degrees of freedom) applied to Preclinical Alzheimer Cognitive Composite (PACC) scores among amyloid-positive A4 participants in (a) the stable class and (b) the two decliner classes. The table reports mean PACC and residual variance at 2 and 4 years, representing expected control group trajectories for two potential trial durations. Power was approximated using two-sample t-test calculations assuming 500 participants per group, with 10% attrition at 2 years and 20% at 4 years. Effect size is expressed both as absolute PACC change and as a percentage of the maximum possible improvement, defined as the mean PACC among stable amyloid-negative LEARN participants at the corresponding time point.
  2422. :::
  2423. \clearpage
  2424. # References {.unnumbered}
  2425. ::: {#refs}
  2426. :::

A4LEARN-Latent-Classes.qmd at commit 4b113dc, no license · at the source

Overview

Authors: Runpeng Li1, Oliver Langford1, Philip S Insel2, Reisa A Sperling3, Rema Raman1, Paul S Aisen1, Michael C Donohue1
  1. USC Epstein Family Alzheimer's Therapeutic Research Institute, University of Southern California, San Diego, California, USA
  2. Department of Psychiatry, University of California, San Francisco, California, USA
  3. Center for Alzheimer Research and Treatment, Brigham and Women's Hospital, Massachusetts General Hospital, Harvard Medical School, Boston, Massachusetts, USA
Institutions: University of Southern California (United States); University of California, San Francisco (United States); Brigham and Women's Hospital (United States); Harvard University (United States); Massachusetts General Hospital (United States)
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association, volume 22, issue 4, article e71366
Dates: received 26 January 2026; accepted 11 March 2026; published online 21 April 2026; in print April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/alz.71366 · PMID 42015324 · PMCID PMC13099594 · OpenAlex W7155186346
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: PET / SPECT (modality), human (organism), Alzheimer's / dementia (population), clinical / translational (subfield)
Methods: Machine learning, Statistics
Keywords: clinical trial design, cognitive decline, latent class mixed‐effects model, preclinical Alzheimer's disease
MeSH: Alzheimer Disease*, Cognitive Dysfunction*, Secondary Prevention*, Aged, Amyloid beta-Peptides, Biomarkers, Disease Progression, Female, Hippocampus, Humans, Longitudinal Studies, Male, Neuropsychological Tests, Positron-Emission Tomography, Prodromal Symptoms, tau Proteins (* major topic)
Topic: Dementia and Cognitive Impairment Research (Psychiatry and Mental health, Medicine), according to OpenAlex
Funding: Alzheimer’s Association; Foundation for the National Institutes of Health; GHR Foundation; NIA NIH HHS (U24AG057437, R01 AG063689, R01 AG079142, U19 AG010483, U24 AG057437); Alzheimer&apos;s Association; Avid Radiopharmaceuticals; Eli Lilly and Company; National Institute on Aging (U24AG057437); Epstein Family Foundation
Citations: cited by 1 paper (Europe PMC); 18 references in the paper

Abstract

INTRODUCTION: Biomarkers identify Alzheimer's disease pathology in cognitively unimpaired adults, but the timing and rate of cognitive decline vary widely. This study aimed to identify subgroups of cognitive decline and baseline predictors of heterogeneity in preclinical progression.

METHODS: Data were drawn from the Anti‐Amyloid Treatment in Asymptomatic Alzheimer's Disease Study, which enrolled amyloid beta‐positive (Aβ+) participants, and the Longitudinal Evaluation of Amyloid Risk and Neurodegeneration (LEARN) Study, which enrolled amyloid beta‐negative (Aβ−) individuals. Latent class mixed‐effects models identified cognitive trajectory classes. Associations between class membership and demographic, clinical, and biomarker variables were evaluated. The primary outcome was change in the Preclinical Alzheimer Cognitive Composite.

RESULTS: Three trajectory classes were identified: stable, slow decliners, and fast decliners. Higher phosphorylated tau at 217 (p‐tau217), smaller hippocampal volume, and elevated tau positron emission tomography were associated with declining classes. About 70% of Aβ+ individuals were stable.

DISCUSSION: Latent class modeling reveals substantial heterogeneity in preclinical trajectories with important implications for prevention trial design.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

atri-biostats/a4learn

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 4b113dc0ce7d5fb66d967f8168e5ea9c9d2cbbbc, 6 March 2026
Languages: R (7), Quarto (1)
Size: 23 files, 8 scripts
Software Heritage: not archived
Found in: the references
Holds: README, environment (DESCRIPTION), documentation, 5 notebooks
Not found: license file, CITATION.cff, tests, continuous integration
Tools: tidyverse (6 files), emmeans (2 files), nlme (2 files), caret (1 file), patchwork (1 file), pROC (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
9 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 8 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 4 keywords, 16 MeSH terms, 9 funders, 13 references.

Cite

This paper

Li, R., Langford, O., Insel, P. S., Sperling, R. A., Raman, R., Aisen, P. S., & Donohue, M. C. (2026). Divergent patterns of cognitive decline in preclinical Alzheimer's disease: Implications for secondary prevention trials. Alzheimer's & dementia : the journal of the Alzheimer's Association, 22(4), e71366. https://doi.org/10.1002/alz.71366

BibTeX

@article{li2026divergent,
author = {Li, Runpeng and Langford, Oliver and Insel, Philip S and Sperling, Reisa A and Raman, Rema and Aisen, Paul S and Donohue, Michael C},
title = {{Divergent patterns of cognitive decline in preclinical Alzheimer's disease: Implications for secondary prevention trials}},
journal = {Alzheimer's \& dementia : the journal of the Alzheimer's Association},
year = {2026},
month = apr,
volume = {22},
number = {4},
pages = {e71366},
publisher = {Wiley},
issn = {1552-5260},
doi = {10.1002/alz.71366},
url = {https://doi.org/10.1002/alz.71366},
pmid = {42015324},
pmcid = {PMC13099594}
}

RIS

TY - JOUR
AU - Li, Runpeng
AU - Langford, Oliver
AU - Insel, Philip S
AU - Sperling, Reisa A
AU - Raman, Rema
AU - Aisen, Paul S
AU - Donohue, Michael C
TI - Divergent patterns of cognitive decline in preclinical Alzheimer's disease: Implications for secondary prevention trials
T2 - Alzheimer's & dementia : the journal of the Alzheimer's Association
J2 - Alzheimers Dement
PY - 2026
DA - 2026/04/01
VL - 22
IS - 4
SP - e71366
SN - 1552-5260
PB - Wiley
DO - 10.1002/alz.71366
UR - https://doi.org/10.1002/alz.71366
LA - en
ER -

CSL-JSON

{
"id": "10.1002/alz.71366",
"type": "article-journal",
"title": "Divergent patterns of cognitive decline in preclinical Alzheimer's disease: Implications for secondary prevention trials",
"container-title": "Alzheimer's & dementia : the journal of the Alzheimer's Association",
"author": [
{
"family": "Li",
"given": "Runpeng"
},
{
"family": "Langford",
"given": "Oliver"
},
{
"family": "Insel",
"given": "Philip S"
},
{
"family": "Sperling",
"given": "Reisa A"
},
{
"family": "Raman",
"given": "Rema"
},
{
"family": "Aisen",
"given": "Paul S"
},
{
"family": "Donohue",
"given": "Michael C"
}
],
"container-title-short": "Alzheimers Dement",
"volume": "22",
"issue": "4",
"page": "e71366",
"DOI": "10.1002/alz.71366",
"PMID": "42015324",
"PMCID": "PMC13099594",
"ISSN": "1552-5260",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/alz.71366",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1126/sciadv.aec9291 [code]
Computational mechanisms of perception in autism revealed using games inspired by rodent operant tasks.
Journal: Science advances
In common: pROC, nlme, caret, 3 other tools
[2] doi:10.1016/j.celrep.2026.117298 [code]
Midbrain endocannabinoids actuate dopamine-based action selection.
Journal: Cell reports
In common: nlme, caret, emmeans, 2 other tools
[3] doi:10.1093/braincomms/fcag236 [code]
Dynamic, state-dependent characteristics of cognitive fluctuations in Lewy body dementia: a magnetoencephalography study.
Journal: Brain communications
In common: pROC, caret, patchwork, 1 other tool, Alzheimer's / dementia
[4] doi:10.1093/neuonc/noag128 [code]
Spatially-resolved single-cell imaging of melanoma brain metastases identifies localized immune patterns predictive of immune checkpoint blockade response.
Journal: Neuro-oncology
In common: pROC, caret, patchwork, 1 other tool, clinical / translational
[5] doi:10.1038/s41467-026-77170-3 [code]
DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.
Journal: Nature communications
In common: pROC, caret, emmeans, 1 other tool
[6] doi:10.3390/ijms27125533 [code]
Discovery-Driven Plasma Proteomics Identifies a Multi-Protein Signature for Amyloid PET Positivity: A Machine Learning Analysis of the Bio-Hermes Cohort.
Journal: International journal of molecular sciences
In common: pROC, caret, tidyverse, PET / SPECT, Alzheimer's / dementia
[7] doi:10.3390/ijms27156925 [code]
XGBoost-SHAP Interpretable Modeling Identifies and Validates an Eight-Gene Biomarker for Hepatic Encephalopathy Risk Prediction in Cirrhosis.
Journal: International journal of molecular sciences
In common: pROC, caret, patchwork, 1 other tool
[8] doi:10.7717/peerj.21426 [code]
Integrated transcriptomic identification and validation reveal key autophagy-associated biomarkers in sleep deprivation.
Journal: PeerJ
In common: pROC, caret, patchwork, 1 other tool
[9] doi:10.3390/ijms27104466 [code]
Uncovering the Key Circuit FOSL2/FOS/EGR3/EGR1, Contributing to the Hyperexcitability of Excitatory Neurons in the Epileptic Temporal Cortex and Hippocampus.
Journal: International journal of molecular sciences
In common: pROC, caret, patchwork, 1 other tool
[10] doi:10.1002/alz.71567 [code]
Associations of dementia polyexposure scores to Alzheimer's disease endophenotypes in a diverse population.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: pROC, patchwork, tidyverse, PET / SPECT, Alzheimer's / dementia, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.