Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease.
The 16 matches
- [1] § Methods › DNA methylation profiling ↔ R/DNAm-preprocessing.R, lines 107–150 · score 0.85 · bisulfite conversion rates, rmSNPandCH, cross hybridizing, sex chromosomes, genomic, technical
- [2] § Methods › Functional enrichment and visualization ↔ R/differential_analyses.R, lines 723–852 · score 0.81 · normalized enrichment score, MSigDB, fold change, NES, BH FDR, fgsea
- [3] § Methods › Functional enrichment and visualization ↔ R/MOFA.R, lines 361–502 · score 0.81 · normalized enrichment score, MSigDB, NES, ranked, fgsea, ORA
- [4] § Methods › Amyloid PET and MRI-based atrophy measures ↔ R/phenotype.R, lines 538–597 · score 0.74 · cubic age, meta ROI cortical, cortical thickness, Hippocampal volumes, ICV, ADNI
- [5] § Methods › Gene expression profiling ↔ R/phenotype.R, lines 599–637 · score 0.69 · Affymetrix Human Genome, U219, chip, Gene expression, ADNI, probes
- [6] § Methods › Network and multi-omics analyses ↔ R/WGCNA.R, lines 845–940 · score 0.69 · principal component, module eigengene, limma, linear, fit, matrices
- [7] § Methods › Survival analysis for MCI to AD progression ↔ R/WGCNA.R, lines 1418–1563 · score 0.60 · module assignments, original modules, MEs, AD progression, predictors, covariates
- [8] § Methods › DNA methylation profiling ↔ R/phenotype.R, lines 1131–1173 · score 0.58 · ADNI Genetics Core, DNA Methylation, QC
- [9] § Results › Co-methylation network analysis reveals immune-related dysregulation in AD ↔ R/wgcna2cytoscape.R, lines 151–274 · score 0.56 · co methylation network, CpGs, Hub gene, WGCNA, modules
- [10] § Results › Gene co-expression network analysis (WGCNA) identifies robust immune networks ↔ R/WGCNA.R, lines 608–633 · score 0.55 · S13a, S13c, WGCNA, women, gene
- [11] § Methods › Network and multi-omics analyses ↔ R/MOFA.R, lines 35–89 · score 0.55 · MOFA model, limma, trained, variance, linear, matched
- [12] § Methods › DNA methylation profiling ↔ R/phenotype.R, lines 1393–1471 · score 0.55 · immune cell, DNA methylation, EpiDISH, blood, profiling
- [13] § Results › Participant characteristics ↔ R/tables.R, lines 18–85 · score 0.55 · amyloid PET positive, hippocampal volume, S1a, carriers, variables, plasma
- [14] § Results › Gene co-expression network analysis (WGCNA) identifies robust immune networks ↔ R/WGCNA.R, lines 446–485 · score 0.54 · module membership, gene co expression, gene modules, MM, WGCNA, network
- [15] § Methods › Sensitivity analyses ↔ R/tables.R, lines 18–85 · score 0.54 · amyloid positive, carrier status, amyloid negative, amyloid PET, MCI, CU
- [16] § Results › AD-related gene modules are linked to AD endophenotypes ↔ R/phenotype.R, lines 236–281 · score 0.53 · mini mental state, examination, cognitive, MMSE
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 1,471 lines · 65 KB · no license · 5 matches
- library(SummarizedExperiment)
- library(tidyverse)
- library(lme4)
- library(EpiDISH)
- # Define Paths & Functions ####
- base_dir <- ".../data"
- proc_dir <- ".../processed"
- idat_dir <- ".../IDAT"
- addZeros <- function(x){
- sprintf("%04d", as.integer(x))
- }
- # Define function to match tables at same time or within 1-year of gene expression profiling
- gene_within1yr <- function(df1, df2_Edate, df2_VISCODE = FALSE) {
- df1 <- df1
- # If table contains both VISCODE and EXAMDATE, use both to determine closest match
- if (df2_VISCODE){
- df1 <- df1 %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(GeneX_Edate, {{ df2_Edate }}), "years"))) %>%
- mutate(Edate_match = ifelse(date_diff <= 1, 1, 0),
- VISCODE_match = ifelse(VISCODE == VISCODE2, 1, 0)) %>%
- group_by(RID) %>%
- mutate(min_match_date = min(date_diff)) %>%
- # Same time = VISCODE_match; within 1-year = Edate_match
- mutate(match_priority = case_when(
- VISCODE_match == 1 ~ 1,
- Edate_match == 1 ~ 2,
- TRUE ~ 3
- )) %>%
- # Only keep matches
- filter(match_priority <= 2) %>%
- # If multiple matches, keep min date diff
- filter(min_match_date == date_diff) %>%
- ungroup() %>%
- distinct(RID, VISCODE, .keep_all = T) %>%
- dplyr::select(-c(Edate_match, VISCODE_match, match_priority, date_diff,
- min_match_date, VISCODE2))
- } else{
- # If table is missing VISCODE, use EXAMDATE only
- df1 <- df1 %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(GeneX_Edate, {{ df2_Edate }}), "years"))) %>%
- mutate(Edate_match = ifelse(date_diff <= 1, 1, 0)) %>%
- group_by(RID) %>%
- mutate(min_match_date = min(date_diff)) %>%
- filter(Edate_match == 1) %>%
- filter(min_match_date == date_diff) %>%
- ungroup() %>%
- distinct(RID, VISCODE, .keep_all = T) %>%
- dplyr::select(-c(Edate_match, date_diff, min_match_date))
- }
- }
- # Define function to match tables at same time or within 1-year of DNA methylation profiling
- dnam_within1yr <- function(df1, df2_Edate, df2_VISCODE = FALSE) {
- df1 <- df1
- # If table contains both VISCODE and EXAMDATE, use both to determine closest match
- if (df2_VISCODE){
- df1 <- df1 %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(DNAm_Edate, {{ df2_Edate }}), "years"))) %>%
- mutate(Edate_match = ifelse(date_diff <= 1, 1, 0),
- VISCODE_match = ifelse(VISCODE == VISCODE2, 1, 0)) %>%
- group_by(RID, DNAm_Edate) %>%
- mutate(min_match_date = min(date_diff)) %>%
- # Same time = VISCODE_match; within 1-year = Edate_match
- mutate(match_priority = case_when(
- VISCODE_match == 1 ~ 1,
- Edate_match == 1 ~ 2,
- TRUE ~ 3
- )) %>%
- # Only keep matches
- filter(match_priority <= 2) %>%
- # If multiple matches, keep min date diff
- filter(min_match_date == date_diff) %>%
- ungroup() %>%
- distinct(RID, barcodes, .keep_all = T) %>%
- dplyr::select(-c(Edate_match, VISCODE_match, match_priority, date_diff,
- min_match_date, VISCODE2))
- } else{
- # If table is missing VISCODE, use EXAMDATE only
- df1 <- df1 %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(DNAm_Edate, {{ df2_Edate }}), "years"))) %>%
- mutate(Edate_match = ifelse(date_diff <= 1, 1, 0)) %>%
- group_by(RID, DNAm_Edate) %>%
- mutate(min_match_date = min(date_diff)) %>%
- filter(Edate_match == 1) %>%
- filter(min_match_date == date_diff) %>%
- ungroup() %>%
- distinct(RID, barcodes, .keep_all = T) %>%
- dplyr::select(-c(Edate_match, date_diff, min_match_date))
- }
- }
- #______________________________________Prepare ADNI Tables Prior to Merging With -Omics Tables______________________________________#####
- message("Now reading in and preparing ADNI tables prior to merging with -omics metadata...")
- ## Demographics Table ####
- # Read in and prepare demographics table
- demog <- read.csv(file.path(base_dir, "All_Subjects_PTDEMOG_28Oct2024.csv"), stringsAsFactors = FALSE)
- adni_merge <- read.csv(file.path(base_dir, "ADNIMERGE.csv"), stringsAsFactors = FALSE)
- adni_merge_unique <- adni_merge %>% distinct_at(vars(RID), .keep_all = TRUE)
- demog_unique <- demog %>% distinct_at(vars(RID), .keep_all = TRUE)
- demog_merged <- left_join(adni_merge_unique, demog_unique %>%
- dplyr::select(RID, PTDOB, PTDOBYY), by = "RID")
- demog_merged <- demog_merged %>%
- dplyr::select(RID, ORIGPROT, AGE, PTDOB, PTGENDER,
- PTEDUCAT, PTRACCAT, PTETHCAT) %>%
- dplyr::rename(AGE_bl = AGE,
- PTDOBMMYY = PTDOB)
- demog_merged <- demog_merged %>%
- mutate(day = 1,
- PTDOB = as.character(paste(day, PTDOBMMYY, sep = "/")))
- demog_merged$PTDOB <- as.Date(demog_merged$PTDOB, format = "%d/%m/%Y")
- demog_merged <- demog_merged %>%
- dplyr::select(-c(PTDOBMMYY, day))
- demog_merged <- relocate(demog_merged, "PTDOB", .before = "AGE_bl")
- demog_merged$RID <- addZeros(demog_merged$RID)
- ## Diagnosis Table ####
- # Read in and prepare diagnosis table
- dx <- read.csv(file.path(base_dir, "All_Subjects_DXSUM_22Nov2024.csv"), stringsAsFactors = FALSE)
- dx <- dx %>%
- dplyr::select(RID, DIAGNOSIS, EXAMDATE, VISCODE2) %>%
- dplyr::rename(DX_Edate = EXAMDATE,
- VISCODE = VISCODE2) %>%
- mutate(DXSUM = case_when(
- (DIAGNOSIS == 1) ~ "CU",
- (DIAGNOSIS == 2) ~ "MCI",
- (DIAGNOSIS == 3) ~ "Dementia"),
- VISCODE = case_when(
- VISCODE == "sc" ~ "bl",
- TRUE ~ VISCODE)
- ) %>% filter(!is.na(DXSUM))
- dx$RID <- addZeros(dx$RID)
- dx <- dx %>%
- arrange(RID, DX_Edate) %>%
- group_by(RID) %>%
- mutate(DX_last = DXSUM[[length(DXSUM)]]) %>%
- ungroup() %>%
- mutate(DX_Edate = as.Date(DX_Edate, "%Y-%m-%d")) %>%
- dplyr::select(RID, VISCODE, DX_Edate, DXSUM, DX_last)
- ## Registry Table ####
- # Read in and prepare registry table
- registry <- read.csv(file.path(base_dir, "REGISTRY_28Jul2023.csv"), stringsAsFactor = FALSE)
- registry$RID <- addZeros(registry$RID)
- registry <- registry %>% dplyr::select(RID, VISCODE, VISCODE2)
- registry$VISCODE <- ifelse(registry$VISCODE == "v03", "bl", registry$VISCODE)
- registry$VISCODE <- ifelse(registry$VISCODE == "v04", "m03", registry$VISCODE)
- registry$VISCODE <- ifelse(registry$VISCODE == "v05", "m06", registry$VISCODE)
- registry <- registry %>% distinct_at(vars(RID, VISCODE), .keep_all = TRUE)
- ## Neurofilament Light Chain (NfL) Table ####
- # Read in and prepare NfL table
- nfl_data <- read.csv(file.path(base_dir, "ADNI_BLENNOWPLASMANFLLONG_10_03_18_22Apr2025.csv"), stringsAsFactor = FALSE, check.names = FALSE)
- # Recode with correct VISCODE2
- nfl_data$VISCODE2[which(nfl_data$RID=="1097"&nfl_data$EXAMDATE=="2013-01-31")]<-"m72"
- nfl_data <- nfl_data %>%
- group_by(RID,EXAMDATE) %>%
- mutate(PLASMA_NFL = mean(PLASMA_NFL)) %>%
- dplyr::distinct_at(vars(RID,EXAMDATE), .keep_all = TRUE) %>%
- ungroup()
- nfl_data <- nfl_data %>%
- mutate(VISCODE2 = case_when(
- VISCODE2 == "sc" ~ "bl",
- TRUE ~ VISCODE2
- )) %>%
- dplyr::select(RID, VISCODE2, EXAMDATE, PLASMA_NFL) %>%
- dplyr::rename(NfL_Edate = EXAMDATE, NfL_B = PLASMA_NFL)
- nfl_data$NfL_Edate <- as.Date(nfl_data$NfL_Edate, "%Y-%m-%d")
- nfl_data$RID <- addZeros(nfl_data$RID)
- ## Estimated CSF AB/TAU-Positivity Onset Table ####
- # Read in and prepare estimated CSF AB/TAU positivity onset table
- csf_age <- read.csv(file.path(base_dir, "estimated_csf_ages.csv"), stringsAsFactor = FALSE, check.names = FALSE)
- colnames(csf_age)[2] <- "AB_AGE"
- colnames(csf_age)[3] <- "TAU_AGE"
- csf_age$AB_AGE <- ifelse(csf_age$AB_AGE == "NaN", NA, csf_age$AB_AGE)
- csf_age$TAU_AGE <- ifelse(csf_age$TAU_AGE == "NaN", NA, csf_age$TAU_AGE)
- csf_age$RID <- addZeros(csf_age$RID)
- csf_age <- csf_age[!duplicated(csf_age$RID),]
- ## APOE Genotypes Table ####
- # Read in and prepare APOE genotypes table
- apoeres <- read.csv(file.path(base_dir, "APOERES_24Aug2023.csv"), stringsAsFactor = FALSE, check.names = FALSE)
- apoeres$APOE4 <- as.numeric(apply(apoeres[, c('APGEN1', 'APGEN2')]==4, 1, sum)>=1)
- apoeres$APOE2 <- as.numeric(apply(apoeres[, c('APGEN1', 'APGEN2')]==2, 1, sum)>=1)
- apoeres$APOE3 <- as.numeric(apply(apoeres[, c('APGEN1', 'APGEN2')]==3, 1, sum)>=1)
- apoeres$APOE4_n <- apply(apoeres[, c('APGEN1', 'APGEN2')], 1, function(row) {
- if (any(row == 4)) {
- if (all(row == 4)) {
- return(2)
- } else {
- return(1)
- }
- } else {
- return(0)
- }
- })
- apoeres$APOE2_n <- apply(apoeres[, c('APGEN1', 'APGEN2')], 1, function(row) {
- if (any(row == 2)) {
- if (all(row == 2)) {
- return(2)
- } else {
- return(1)
- }
- } else {
- return(0)
- }
- })
- apoeres$APOE3_n <- apply(apoeres[, c('APGEN1', 'APGEN2')], 1, function(row) {
- if (any(row == 3)) {
- if (all(row == 3)) {
- return(2)
- } else {
- return(1)
- }
- } else {
- return(0)
- }
- })
- apoeres_df <- apoeres[c("RID", "APOE4", "APOE2", "APOE3", "APOE4_n", "APOE2_n", "APOE3_n")]
- apoeres_df$RID <- addZeros(apoeres_df$RID)
- ## Cognitive Performance Tables ####
- # Read in and prepare cognitive performance tables
- # Clinical Dementia Rating (CDR)
- cdr <- read.csv(file.path(base_dir, "All_Subjects_CDR_22Nov2024.csv"), stringsAsFactors = FALSE)
- cdr <- cdr[cdr$CDGLOBAL>=0,]
- cdr$CDRSB <- rowSums(cdr[,c('CDMEMORY', 'CDORIENT', 'CDJUDGE', 'CDCOMMUN', 'CDHOME', 'CDCARE')])
- cdr <- cdr %>%
- mutate(VISCODE2 = case_when(
- VISCODE2 == "sc" ~ "bl",
- TRUE ~ VISCODE2
- )) %>%
- dplyr::select(RID, VISCODE2, VISDATE, CDRSB) %>%
- dplyr::rename(VISCODE = VISCODE2)
- cdr$RID <- addZeros(cdr$RID)
- cdr$CDR_Edate <- as.Date(cdr$VISDATE, "%Y-%m-%d")
- # 13-Item Alzheimer's Disease Assessment Scale–Cognitive Subscale (ADAS13)
- adas <- read.csv(file.path(base_dir, "All_Subjects_ADAS_ADNIGO23_22Nov2024.csv"), stringsAsFactors = FALSE)
- adas <- adas %>%
- dplyr::select(RID, VISCODE2, TOTAL13) %>%
- dplyr::rename(VISCODE = VISCODE2) %>%
- dplyr::rename(ADAS13 = TOTAL13) %>%
- group_by(RID, VISCODE) %>%
- mutate(across(.cols = dplyr::where(is.numeric), .fns=~mean(.,na.rm = TRUE))) %>%
- distinct_at(vars(RID), .keep_all = TRUE)
- adas$RID <- addZeros(adas$RID)
- adas$ADAS13 <- ifelse(adas$ADAS13 == "NaN", NA, adas$ADAS13)
- # Mini-Mental State Examination (MMSE)
- mmse <- read.csv(file.path(base_dir, "All_Subjects_MMSE_22Nov2024.csv"), stringsAsFactors = FALSE)
- mmse <- mmse %>%
- dplyr::select(RID, VISCODE2, VISDATE, MMSCORE) %>%
- dplyr::rename(MMSE = MMSCORE) %>%
- mutate(VISCODE2 = case_when(
- VISCODE2 == "sc" ~ "bl",
- TRUE ~ VISCODE2
- )) %>%
- group_by(RID, VISCODE2) %>%
- mutate(across(.cols = dplyr::where(is.numeric), .fns=~mean(., na.rm = TRUE))) %>%
- distinct_at(vars(RID), .keep_all = TRUE) %>%
- ungroup()
- mmse$MMSE_Edate <- as.Date(mmse$VISDATE, "%Y-%m-%d")
- mmse <- mmse %>% dplyr::select(-VISDATE)
- mmse$MMSE <- ifelse(mmse$MMSE == "NaN", NA, mmse$MMSE)
- mmse$RID <- addZeros(mmse$RID)
- ## CSF Biomarkers Table ####
- # Read in and prepare CSF biomarkers table
- upenn_csf <- read.csv(file.path(base_dir, "All_Subjects_UPENNBIOMK_ROCHE_ELECSYS_23Nov2024.csv"), stringsAsFactors = FALSE)
- upenn_csf <- upenn_csf %>%
- arrange(desc(RUNDATE)) %>%
- distinct_at(vars(RID, VISCODE2), .keep_all = TRUE) #keep most recent replicate
- upenn_csf <- upenn_csf %>%
- dplyr::rename(ABETA40_csf = ABETA40,
- ABETA42_csf = ABETA42,
- tTAU_csf = TAU,
- pTAU181_csf = PTAU)
- upenn_csf <- upenn_csf %>%
- mutate(ABETA42_csf = round(ABETA42_csf),
- ptau_pos_csf = pTAU181_csf > 24,
- amyloid_pos_csf = ABETA42_csf < 980,
- ptau_ab_ratio_csf = (as.numeric(pTAU181_csf) / as.numeric(ABETA42_csf)),
- ad_pathology_pos_csf = ptau_ab_ratio_csf > 0.025)
- upenn_csf <- upenn_csf %>%
- mutate(across(.cols=c(ABETA42_csf, ABETA40_csf, tTAU_csf,pTAU181_csf, ptau_ab_ratio_csf), .fns=as.numeric))
- upenn_csf <- upenn_csf %>%
- mutate(across(.cols=c(ptau_pos_csf, amyloid_pos_csf, ad_pathology_pos_csf), .fns=as.numeric))
- # reassign records with VISCODE2 = "UNK" to correct VISCODE2
- upenn_csf$VISCODE2[upenn_csf$RID==89 & upenn_csf$VISCODE2=="UNK"]<-"m162"
- upenn_csf$VISCODE2[upenn_csf$RID==232 & upenn_csf$VISCODE2=="UNK"]<-"m36"
- upenn_csf$VISCODE2[upenn_csf$RID==726 & upenn_csf$VISCODE2=="UNK"]<-"m24"
- upenn_csf$VISCODE2[upenn_csf$RID==790 & upenn_csf$VISCODE2=="UNK"]<-"m18"
- upenn_csf$VISCODE2[upenn_csf$RID==1200 & upenn_csf$VISCODE2=="UNK"]<-"m36"
- upenn_csf$VISCODE2[upenn_csf$RID==2373 & upenn_csf$VISCODE2=="UNK"]<-"m120"
- upenn_csf$VISCODE2[upenn_csf$RID==4216 & upenn_csf$VISCODE2=="UNK"]<-"m102"
- upenn_csf$VISCODE2[upenn_csf$RID==4393 & upenn_csf$VISCODE2=="UNK"]<-"m96"
- upenn_csf$VISCODE2[upenn_csf$RID==4643 & upenn_csf$VISCODE2=="UNK"]<-"m72"
- upenn_csf$VISCODE2[upenn_csf$RID==5140 & upenn_csf$VISCODE2=="UNK"]<-"m84"
- upenn_csf$VISCODE2[upenn_csf$RID==6161 & upenn_csf$VISCODE2=="UNK"]<-"bl"
- upenn_csf$VISCODE2[upenn_csf$RID==6880 & upenn_csf$VISCODE2=="UNK"]<-"bl"
- upenn_csf$VISCODE2[upenn_csf$RID==6906 & upenn_csf$VISCODE2=="UNK"]<-"bl"
- upenn_csf <- upenn_csf %>%
- dplyr::select(RID, VISCODE2, EXAMDATE, ABETA40_csf, ABETA42_csf, tTAU_csf, pTAU181_csf,
- ptau_pos_csf, amyloid_pos_csf, ptau_ab_ratio_csf, ad_pathology_pos_csf) %>%
- dplyr::rename(VISCODE = VISCODE2,
- CSF_Edate = EXAMDATE)
- upenn_csf$CSF_Edate <- as.Date(upenn_csf$CSF_Edate, "%Y-%m-%d")
- upenn_csf$RID <- addZeros(upenn_csf$RID)
- # Identify the first/last negative/positive dates of CSF amyloid
- last_negative_date_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(amyloid_pos_csf == 0) %>%
- mutate(last_AB_neg_date = max(CSF_Edate)) %>%
- dplyr::select(RID, last_AB_neg_date) %>%
- ungroup() %>%
- distinct()
- first_positive_date_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(amyloid_pos_csf == 1) %>%
- mutate(first_AB_pos_date = min(CSF_Edate)) %>%
- dplyr::select(RID, first_AB_pos_date) %>%
- ungroup() %>%
- distinct()
- last_negative_date2_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(ptau_pos_csf == 0) %>%
- mutate(last_TAU_neg_date2 = max(CSF_Edate)) %>%
- dplyr::select(RID, last_TAU_neg_date2) %>%
- ungroup() %>%
- distinct()
- first_positive_date2_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(ptau_pos_csf == 1) %>%
- mutate(first_TAU_pos_date2 = min(CSF_Edate)) %>%
- dplyr::select(RID, first_TAU_pos_date2) %>%
- ungroup() %>%
- distinct()
- last_negative_date3_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(ad_pathology_pos_csf == 0) %>%
- mutate(last_AD_neg_date3 = max(CSF_Edate)) %>%
- dplyr::select(RID, last_AD_neg_date3) %>%
- ungroup() %>%
- distinct()
- first_positive_date3_csf <- upenn_csf %>%
- group_by(RID) %>%
- filter(ad_pathology_pos_csf == 1) %>%
- mutate(first_AD_pos_date3 = min(CSF_Edate)) %>%
- dplyr::select(RID, first_AD_pos_date3) %>%
- ungroup() %>%
- distinct()
- ## Amyloid PET Table ####
- # Read in and prepare amyloid PET table
- ab_pet <- read.csv(file.path(base_dir, "All_Subjects_UCBERKELEY_AMY_6MM_23Nov2024.csv"), stringsAsFactors = FALSE)
- ab_pet$RID <- addZeros(ab_pet$RID)
- ab_pet <- ab_pet %>%
- dplyr::rename(PET_Edate = SCANDATE,
- amyloid_pos_pet_cross = AMYLOID_STATUS,
- amyloid_pos_pet_long = AMYLOID_STATUS_COMPOSITE_REF) %>%
- mutate(amyloid_pos_pet_centiloid = case_when(CENTILOIDS > 22 ~ 1,
- TRUE ~ 0)) %>%
- dplyr::select(RID, PET_Edate, IMAGE_RESOLUTION, TRACER, SUMMARY_SUVR, CENTILOIDS, amyloid_pos_pet_cross,
- amyloid_pos_pet_long, amyloid_pos_pet_centiloid)
- ab_pet$PET_Edate <- as.Date(ab_pet$PET_Edate, "%Y-%m-%d")
- # Identify the first/last negative/positive dates of amyloid PET
- last_negative_date_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_cross == 0) %>%
- mutate(last_AB_neg_date = max(PET_Edate)) %>%
- dplyr::select(RID, last_AB_neg_date) %>%
- ungroup() %>%
- distinct()
- first_positive_date_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_cross == 1) %>%
- mutate(first_AB_pos_date = min(PET_Edate)) %>%
- dplyr::select(RID, first_AB_pos_date) %>%
- ungroup() %>%
- distinct()
- last_negative_date2_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_long == 0) %>%
- mutate(last_AB_neg_date2 = max(PET_Edate)) %>%
- dplyr::select(RID, last_AB_neg_date2) %>%
- ungroup() %>%
- distinct()
- first_positive_date2_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_long == 1) %>%
- mutate(first_AB_pos_date2 = min(PET_Edate)) %>%
- dplyr::select(RID, first_AB_pos_date2) %>%
- ungroup() %>%
- distinct()
- last_negative_date3_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_centiloid == 0) %>%
- mutate(last_AB_neg_date3 = max(PET_Edate)) %>%
- dplyr::select(RID, last_AB_neg_date3) %>%
- ungroup() %>%
- distinct()
- first_positive_date3_pet <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_centiloid == 1) %>%
- mutate(first_AB_pos_date3 = min(PET_Edate)) %>%
- dplyr::select(RID, first_AB_pos_date3) %>%
- ungroup() %>%
- distinct()
- ## FreeSurfer MRI Table ####
- # Get amyloid negative RIDs that were consistently negative for every scan they had
- ab_negs <- ab_pet %>%
- group_by(RID) %>%
- filter(amyloid_pos_pet_cross == 0) %>%
- mutate(first_a_neg_date_pet = min(PET_Edate)) %>%
- dplyr::select(RID, first_a_neg_date_pet) %>%
- ungroup() %>%
- distinct()
- # # Read in and prepare cross-sectional FreeSurfer MRI table
- fs_cs <- read.csv(file.path(base_dir, "UCSF_FS_CS_5.1_2019.csv"), stringsAsFactors = FALSE)
- fs_cs <- fs_cs %>%
- filter(OVERALLQC == "Pass") %>%
- dplyr::select(-VISCODE) %>%
- dplyr::rename(ICV = ST10CV,
- VISCODE = VISCODE2) %>%
- group_by(RID, VISCODE) %>%
- mutate(n = n()) %>%
- filter(
- # Keep all single-scan visits
- n == 1 |
- # For multiple scans: keep Non-Accelerated if available
- (n > 1 & IMAGETYPE == "Non-Accelerated T1") |
- # If no Non-Accelerated, then keep Accelerated
- (n > 1 & !any(IMAGETYPE == "Non-Accelerated T1") & IMAGETYPE == "Accelerated T1")
- ) %>%
- ungroup() %>%
- dplyr::select(-n)
- fs_cs$MRIcs_Edate <- as.Date(fs_cs$EXAMDATE, "%Y-%m-%d")
- fs_cs$RID <- addZeros(fs_cs$RID)
- fs_cs <- fs_cs %>% dplyr::select(-EXAMDATE)
- fs_cs <- relocate(fs_cs, "MRIcs_Edate", .after = "VISCODE")
- # Get demographics and closest diagnosis information
- fs_cs <- merge(fs_cs, demog_merged, by = "RID", all.x = TRUE) %>%
- mutate(age = round(lubridate::time_length(difftime(MRIcs_Edate, PTDOB), "years"), digits = 1))
- fs_cs <- merge(fs_cs, dx %>%
- dplyr::select(-VISCODE) %>%
- filter(!is.na(DX_Edate)), by = "RID") %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(MRIcs_Edate, DX_Edate), "years"))) %>%
- group_by(RID, VISCODE) %>%
- mutate(min_match_date = min(date_diff)) %>%
- ungroup() %>%
- filter(min_match_date == date_diff) %>%
- distinct()
- # For CU only, get earliest amyloid negative scan per RID
- fs_cs_temp <- fs_cs %>%
- filter(DXSUM == "CU")
- fs_cs_temp <- fs_cs_temp %>%
- left_join(ab_negs, by = "RID") %>%
- mutate(diff_time = abs(lubridate::time_length(difftime(MRIcs_Edate, first_a_neg_date_pet), "years"))) %>%
- group_by(RID) %>%
- filter(diff_time == min(diff_time)) %>%
- ungroup()
- fs_cs_temp <- fs_cs_temp %>%
- group_by(RID) %>%
- filter(row_number() == 1) %>%
- ungroup()
- # ICV adjustment for ROI volumes
- # Create list of volumes that need to be ICV-adjusted
- volumes_cs <- fs_cs_temp %>%
- dplyr::select(ICV | contains("CV") | contains("SV"))
- volumes_cs <- volumes_cs %>%
- dplyr::select(-ICV)
- # Perform ICV adjustment across all ROI volumes
- results_cs <- c()
- for (roi_num in 1:length(names(volumes_cs))){
- current_roi <- names(volumes_cs)[roi_num]
- results_cs <- append(results_cs, current_roi)
- }
- for (i in results_cs){
- fit <- lm(fs_cs_temp[[i]] ~ ICV, data = fs_cs_temp, na.action = na.exclude)
- fs_cs[paste0(i,'_ICV')] <- fs_cs[paste0(i)] - (fs_cs$ICV - mean(fs_cs_temp$ICV, na.rm = T)) * fit$coefficients[2]
- }
- fs_cs <- fs_cs %>%
- dplyr::select(-names(volumes_cs))
- # Cubic age adjustment for ROI volumes (captures normative aging)
- # Obtain list of variables that need to be age-adjusted
- # Add "_ICV" suffix to CV and SV in temp to match
- colnames(fs_cs_temp) <- ifelse(grepl("CV|SV", colnames(fs_cs_temp)),
- paste0(colnames(fs_cs_temp), "_ICV"),
- colnames(fs_cs_temp))
- fs_cs_temp <- fs_cs_temp %>% dplyr::rename(ICV = ICV_ICV)
- features_cs <- fs_cs_temp %>%
- dplyr::select(ICV | contains("CV") | contains("SV"))
- features_cs <- features_cs %>%
- dplyr::select(-ICV)
- # Perform age adjustment across all ROI volumes
- results_cs <- c()
- for (roi_num in 1:length(names(features_cs))){
- current_roi <- names(features_cs)[roi_num]
- results_cs <- append(results_cs, current_roi)
- }
- # Adjust for cubic age
- for (i in results_cs){
- fit <- lm(fs_cs_temp[[i]] ~ age + poly(age, 3, raw = TRUE)[,"3"], data = fs_cs_temp, na.action = na.exclude)
- fs_cs[paste0(i,'_age')] <- fs_cs[paste0(i)] - (fs_cs$age - mean(fs_cs_temp$age, na.rm = T)) * fit$coefficients[2]
- }
- # Compute global measure of hippocampal volume after ICV and age adjustments
- fs_cs <- fs_cs %>%
- dplyr::select(-names(features_cs)) %>%
- dplyr::rename(HVa_L = ST29SV_ICV_age,
- HVa_R = ST88SV_ICV_age) %>%
- mutate(HVa = HVa_L + HVa_R)
- # Compute meta-ROI cortical thickness
- fs_cs <- fs_cs %>%
- mutate(SA_total_L = ST24SA + ST32SA + ST40SA + ST26SA,
- SA_total_R = ST83SA + ST91SA + ST99SA + ST85SA,
- meta_ROI_L = (ST24TA*ST24SA + ST32TA*ST32SA + ST40TA*ST40SA + ST26TA*ST26SA) / SA_total_L,
- meta_ROI_R = (ST83TA*ST83SA + ST91TA*ST91SA + ST99TA*ST99SA + ST85TA*ST85SA) / SA_total_R,
- meta_ROI = (meta_ROI_L * SA_total_L + meta_ROI_R * SA_total_R) / (SA_total_L + SA_total_R))
- fs_cs <- fs_cs %>%
- dplyr::select(RID, VISCODE, STATUS, IMAGETYPE, MRIcs_Edate, HVa_L, HVa_R, HVa, meta_ROI_L, meta_ROI_R, meta_ROI)
- #______________________________________Merge Prepared ADNI Tables With Gene Expression Metadata______________________________________#####
- ## Read in and Preprocess Raw Gene Expression Data ####
- message("Now reading in and preparing raw gene expression data...")
- raw_exp <- read.csv(file.path(base_dir, "ADNI_Gene_Expression_Profile.csv"), stringsAsFactors = FALSE)
- # Isolate and minimally format gene expression-level data
- # Extract RIDs from PTID
- id <- as.data.frame(t(raw_exp[raw_exp$Phase == "SubjectID", ]))
- id$RID <- gsub("^.*_(\\d{4})$", "\\1", id[,1])
- # Extract expression-level data
- exp <- raw_exp
- colnames(exp) <- exp[8,]
- exp <- exp[9:nrow(exp),]
- # Update column names to RIDs
- colnames(exp)[4:ncol(exp)] <- id$RID[4:nrow(id)]
- # Combine gene symbol and probe set ID
- exp$GeneSet <- paste(exp$Symbol, exp$ProbeSet, sep = " / ")
- # Restrict to the expression data
- row.names(exp) <- NULL
- exp <- exp[, names(exp) != ""]
- exp <- exp %>%
- tibble::column_to_rownames("GeneSet") %>%
- dplyr::select(-c(ProbeSet, LocusLink, Symbol, ))
- # Ensure expression data is numeric
- exp <- exp %>%
- mutate(dplyr::across(everything(), ~ as.numeric(as.character(.x))))
- # Sort by RID
- exp <- exp[, order(names(exp))]
- ### Update Gene Symbols Using Affymetrix & MSigDB Annotation Files ####
- # Ensures that we maximize overlap with MSigDB's gene sets for functional enrichment analyses
- # There are gene symbols that were unintentionally converted to dates (e.g., 1-Sep = SEPTIN 1)
- message("Now updating the gene symbols in the gene expression matrix using Affymetrix and MSigDB annotation files...")
- # Read in and prepare Affymetrix's annotation file
- anno <- data.table::fread(file.path(base_dir, "AFFY-U219_GPL13667-15572.txt"), header = T, sep = "\t")
- anno <- anno %>%
- mutate(`Gene Symbol` = ifelse(`Gene Symbol` == "---" | `Gene Symbol` == "", NA, `Gene Symbol`),
- `Gene Symbol` = ifelse(`Gene Symbol` == "Schr2q13", "SEPTIN10", `Gene Symbol`),
- `Gene Symbol` = gsub("///", "||", `Gene Symbol`))
- # Read in and prepare a modified version of Affymetrix's annotation file tailored to MSigDB
- chip <- read.table(file.path(base_dir, "Affymetrix_Human_Genome_U219_MSigDB.v.7.4_custom.chip"), sep = "\t",
- comment.char = "", quote = "", stringsAsFactors = FALSE,
- fill = TRUE, header = T)
- chip_collapsed <- chip %>%
- group_by(Probe.Set.ID) %>%
- summarise(Gene.Symbol.New = paste(unique(Gene.Symbol), collapse = " || "), .groups = "drop")
- # Further refine the reformatted gene expression matrix
- exp$GeneSet <- row.names(exp)
- exp <- exp %>%
- mutate(Symbol = sub(" /.*", "", GeneSet),
- ProbeSet = sub(".*/ ", "", GeneSet)) %>%
- relocate(Symbol, .after = GeneSet) %>%
- relocate(ProbeSet, .after = Symbol)
- # Update gene symbols with the following priority as needed:
- # (1): MSigDB version, (2) Affymetrix version, and (3) ADNI version
- exp_final <- left_join(exp, anno %>%
- dplyr::select(c(ID, `Gene Symbol`)) %>%
- rename(ProbeSet = ID,
- Affy_Symbol = `Gene Symbol`), by = "ProbeSet")
- exp_final <- left_join(exp_final, chip_collapsed %>%
- dplyr::select(c(Probe.Set.ID, Gene.Symbol.New)) %>%
- rename(ProbeSet = Probe.Set.ID,
- MSigDB_Symbol = Gene.Symbol.New), by = "ProbeSet") %>%
- relocate(ProbeSet, .before = Symbol)
- # Apply the naming priority when updating the gene symbols
- exp_final <- exp_final %>%
- mutate(
- MSigDB_Symbol = case_when(
- is.na(MSigDB_Symbol) & !is.na(Affy_Symbol) ~ Affy_Symbol,
- is.na(MSigDB_Symbol) & is.na(Affy_Symbol) ~ Symbol,
- TRUE ~ MSigDB_Symbol
- ),
- # Remove all asterisks
- MSigDB_Symbol = gsub("\\*", "", MSigDB_Symbol),
- # Manually correct symbol for SEPTIN which was still in date format (e.g., 1-SEP)
- # Symbols for other months (Dec & Mar) were fixed using annotation files
- MSigDB_Symbol = ifelse(grepl("^[0-9]+-Sep$", MSigDB_Symbol, ignore.case=T),
- paste0("SEPTIN", sub("-Sep$", "", MSigDB_Symbol, ignore.case=T)),
- MSigDB_Symbol),
- Combined = paste(MSigDB_Symbol, ProbeSet, sep = " / ")
- ) %>%
- column_to_rownames('Combined') %>%
- dplyr::select(-c(GeneSet, ProbeSet, Symbol, Affy_Symbol, MSigDB_Symbol))
- message("Now saving the final gene expression matrix for analysis...")
- # Save the updated gene expression matrix for analysis
- write.csv(exp_final, file.path(base_dir, "gene_expression_final.csv"), row.names = T)
- # Clear up workspace
- rm(exp, exp_final, id, chip, chip_collapsed, anno)
- ## Create and Prepare Gene Expression Metadata Table ####
- # Extract gene expression metadata
- exp_pheno <- raw_exp[1:7,]
- exp_pheno <- exp_pheno %>%
- dplyr::select(where(~ !all(. == "")))
- exp_pheno <- rbind(colnames(exp_pheno), exp_pheno)
- exp_pheno <- as.data.frame(t(exp_pheno))
- colnames(exp_pheno) <- exp_pheno[1,]
- exp_pheno <- exp_pheno[-1, ]
- exp_pheno <- exp_pheno %>%
- dplyr::rename(VISCODE = Visit,
- PTID = SubjectID,
- GeneX_YearDrawn = YearofCollection,
- AffyPlate = `Affy Plate`)
- exp_pheno$RID <- gsub("^.*_(\\d{4})$", "\\1", exp_pheno$PTID)
- exp_pheno <- relocate(exp_pheno, "RID", .before = "Phase")
- # Fix one case of RIN written as "[5.1]" and ensure RIN is numeric
- exp_pheno$RIN[exp_pheno$RID == "0307"] <- "5.1"
- exp_pheno$RIN <- as.numeric(exp_pheno$RIN)
- # Sort by RID
- exp_pheno <- exp_pheno[order(exp_pheno$RID), ]
- rownames(exp_pheno) <- NULL
- exp_pheno$index <- seq.int(nrow(exp_pheno))
- ### Combine ADNI Tables with Gene Expression Metadata ####
- # Main demographics
- exp_pheno_merged <- left_join(demog_merged, exp_pheno, by = "RID")
- exp_pheno_merged <- exp_pheno_merged[!is.na(exp_pheno_merged$GeneX_YearDrawn),]
- # Update VISCODES using registry
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v03", "bl", exp_pheno_merged$VISCODE)
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v04", "m03", exp_pheno_merged$VISCODE)
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v05", "m06", exp_pheno_merged$VISCODE)
- exp_pheno_merged <- left_join(exp_pheno_merged, registry, by = c("RID","VISCODE"))
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v06", exp_pheno_merged$VISCODE2, exp_pheno_merged$VISCODE)
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v11", exp_pheno_merged$VISCODE2, exp_pheno_merged$VISCODE)
- # One case of v02, so use bl instead since it is only 11 days from v02
- exp_pheno_merged$VISCODE <- ifelse(exp_pheno_merged$VISCODE == "v02", "bl", exp_pheno_merged$VISCODE)
- exp_pheno_merged <- exp_pheno_merged %>%
- dplyr::select(-c(VISCODE2))
- # Approximate EXAMDATE for gene expression profiling from ADNIMERGE to calculate age
- # This is necessary because the gene expression metadata only came with the year of collection
- adni_merge2 <- adni_merge %>% dplyr::select(RID, VISCODE, EXAMDATE)
- adni_merge2$RID <- as.character(adni_merge2$RID)
- adni_merge2$RID <- addZeros(adni_merge2$RID)
- exp_pheno_merged <- left_join(exp_pheno_merged, adni_merge2, by = c("RID","VISCODE"))
- exp_pheno_merged <- exp_pheno_merged %>%
- dplyr::rename(GeneX_Edate = EXAMDATE)
- exp_pheno_merged$GeneX_Edate <- as.Date(exp_pheno_merged$GeneX_Edate,"%Y-%m-%d")
- exp_pheno_merged$AGE <- round(time_length(difftime(exp_pheno_merged$GeneX_Edate,
- exp_pheno_merged$PTDOB),
- "years"), digits = 1)
- exp_pheno_merged <- relocate(exp_pheno_merged, "AGE", .before = "AGE_bl")
- exp_pheno_merged <- relocate(exp_pheno_merged, "index", .before = "RID")
- # APOE genotypes
- exp_pheno_merged <- left_join(exp_pheno_merged, apoeres_df, by = "RID")
- # CSF AB/TAU estimated age of positivity onset
- exp_pheno_merged <- left_join(exp_pheno_merged, csf_age, by = "RID")
- exp_pheno_merged <- exp_pheno_merged %>%
- mutate(AB_time = AGE - AB_AGE,
- TAU_time = AGE - TAU_AGE)
- # Diagnosis
- exp_pheno_merged1 <- merge(exp_pheno_merged, dx %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- exp_pheno_merged2 <- gene_within1yr(exp_pheno_merged1, DX_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged, exp_pheno_merged2, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged3 <- bind_rows(exp_pheno_merged2, unmatched)
- exp_pheno_merged3 <- exp_pheno_merged3 %>%
- distinct(index, .keep_all = TRUE)
- # Compute time2event and MCI progressor status
- mci <- left_join(dx, exp_pheno_merged3 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- mci <- mci %>%
- group_by(RID) %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(GeneX_Edate, DX_Edate), "years"))) %>%
- group_by(RID, GeneX_Edate) %>%
- mutate(min_match_date = min(date_diff)) %>%
- ungroup()
- first <- mci %>%
- group_by(RID) %>%
- mutate(DX_bl = ifelse(min_match_date == date_diff, DXSUM, NA),
- bl_DX_date = min(DX_Edate[min_match_date == date_diff], na.rm = T),
- bl_DX_date = as.Date(bl_DX_date, "%Y-%m-%d")) %>%
- filter(!is.na(DX_bl)) %>%
- ungroup() %>%
- dplyr::select(RID, DX_bl, bl_DX_date)
- mci <- left_join(mci, first, by = "RID")
- mci <- mci %>%
- arrange(RID, DX_Edate) %>%
- group_by(RID) %>%
- mutate(last_DX_date = max(DX_Edate, na.rm = T),
- last_DX_date = as.Date(last_DX_date, "%Y-%m-%d"),
- window = DX_Edate >= GeneX_Edate & DX_Edate <= last_DX_date,
- window = ifelse(window == FALSE & bl_DX_date <= GeneX_Edate & min_match_date == date_diff, TRUE, window)) %>%
- filter(window == TRUE) %>%
- mutate(first_AD_date = min(DX_Edate[DXSUM == "Dementia" & window], na.rm = T),
- first_AD_date = as.Date(first_AD_date, "%Y-%m-%d"),
- last_MCI_date = max(DX_Edate[DXSUM == "MCI" & window], na.rm = T),
- last_MCI_date = as.Date(last_MCI_date, "%Y-%m-%d"),
- swap = DX_bl == "MCI" & window & first_AD_date <= last_MCI_date,
- AD_flag = any(DXSUM == "Dementia" & window),
- Progressor = case_when(
- DX_bl == "MCI" & AD_flag & swap ~ NA,
- DX_bl == "MCI" & AD_flag ~ 1 ,
- DX_bl == "MCI" & !AD_flag ~ 0,
- TRUE ~ NA
- )) %>%
- ungroup()
- mci[sapply(mci, is.infinite)] <- NA
- mci <- mci %>%
- group_by(RID) %>%
- mutate(time2event = ifelse(Progressor == 1,
- lubridate::time_length(difftime(first_AD_date, GeneX_Edate), "years"),
- lubridate::time_length(difftime(last_DX_date, GeneX_Edate), "years"))) %>%
- filter(time2event >= 0)
- exp_pheno_merged3 <- left_join(exp_pheno_merged3, mci %>%
- dplyr::select(RID, bl_DX_date, last_DX_date, first_AD_date, last_MCI_date, Progressor, time2event), by = "RID")
- exp_pheno_merged3 <- exp_pheno_merged3 %>%
- distinct(index, .keep_all = TRUE)
- exp_pheno_merged3 <- exp_pheno_merged3 %>%
- mutate(DXSUM2 = case_when(
- (DXSUM == "CU") ~ "CU",
- (DXSUM == "Dementia") ~ "Dementia",
- (Progressor == 0) ~ "Non-Progressor",
- (Progressor == 1) ~ "Progressor MCI",
- TRUE ~ DXSUM
- ))
- # Cognitive performance
- cdr <- left_join(cdr, exp_pheno_merged3 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- cdr <-cdr %>%
- arrange(RID, CDR_Edate) %>%
- group_by(RID) %>%
- mutate(CDR_time=lubridate::time_length(difftime(CDR_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(CDR_Edate))) %>%
- ungroup()
- # Compute individualized slopes of CDRSB
- model <- lmer(CDRSB ~ CDR_time + (CDR_time | RID), subset = n > 1, data = cdr, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "CDRSB_slope"
- slopes$RID <- row.names(slopes)
- cdr <- left_join(cdr, slopes, by = "RID")
- cdr <- cdr %>%
- dplyr::select(-c(n, GeneX_Edate, CDR_time, VISDATE))
- cdr_slopes <- cdr %>%
- dplyr::select(RID, CDRSB_slope) %>%
- distinct()
- exp_pheno_merged4 <- merge(exp_pheno_merged3, cdr %>%
- dplyr::select(-CDRSB_slope) %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- exp_pheno_merged5 <- gene_within1yr(exp_pheno_merged4, CDR_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged3, exp_pheno_merged5, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged6 <- bind_rows(exp_pheno_merged5, unmatched)
- exp_pheno_merged6 <- exp_pheno_merged6[order(exp_pheno_merged6$RID), ]
- rownames(exp_pheno_merged6) <- NULL
- exp_pheno_merged6 <- left_join(exp_pheno_merged6, cdr_slopes, by = "RID")
- exp_pheno_merged6 <- left_join(exp_pheno_merged6, adas, by = c("RID", "VISCODE"))
- mmse <- left_join(mmse, exp_pheno_merged6 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- mmse <- mmse %>%
- arrange(RID, MMSE_Edate) %>%
- group_by(RID) %>%
- mutate(MMSE_time=lubridate::time_length(difftime(MMSE_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(MMSE_Edate))) %>%
- ungroup()
- # Compute individualized slopes of MMSE
- model <- lmer(MMSE ~ MMSE_time + (MMSE_time | RID), subset = n > 1, data = mmse, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "MMSE_slope"
- slopes$RID <- row.names(slopes)
- mmse <- left_join(mmse, slopes, by = "RID")
- mmse <- mmse %>%
- dplyr::select(-c(n, GeneX_Edate, MMSE_time))
- mmse_slopes <- mmse %>%
- dplyr::select(RID, MMSE_slope) %>%
- distinct()
- exp_pheno_merged7 <- merge(exp_pheno_merged6, mmse %>%
- dplyr::select(-MMSE_slope), by = "RID", all.x = TRUE)
- exp_pheno_merged8 <- gene_within1yr(exp_pheno_merged7, MMSE_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged6, exp_pheno_merged8, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged9 <- bind_rows(exp_pheno_merged8, unmatched)
- exp_pheno_merged9 <- exp_pheno_merged9[order(exp_pheno_merged9$RID), ]
- rownames(exp_pheno_merged9) <- NULL
- exp_pheno_merged9 <- left_join(exp_pheno_merged9, mmse_slopes, by = "RID")
- # Plasma NfL
- nfl_data <- left_join(nfl_data, exp_pheno_merged9 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- nfl_data <- nfl_data %>%
- arrange(RID, NfL_Edate) %>%
- group_by(RID) %>%
- mutate(NfL_time=lubridate::time_length(difftime(NfL_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(NfL_Edate))) %>%
- ungroup()
- # Compute individualized slopes of NfL
- model <- lmer(NfL_B ~ NfL_time + (NfL_time | RID), subset = n > 1, data = nfl_data, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "NfL_B_slope"
- slopes$RID <- row.names(slopes)
- nfl_data <- left_join(nfl_data, slopes, by = "RID")
- nfl_data <- nfl_data %>%
- dplyr::select(-c(n, GeneX_Edate, NfL_time))
- nfl_slopes <- nfl_data %>%
- dplyr::select(RID, NfL_B_slope) %>%
- distinct()
- exp_pheno_merged10 <- merge(exp_pheno_merged9, nfl_data %>%
- dplyr::select(-NfL_B_slope), by = "RID", all.x = TRUE)
- exp_pheno_merged11 <- gene_within1yr(exp_pheno_merged10, NfL_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged9, exp_pheno_merged11, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged12 <- bind_rows(exp_pheno_merged11, unmatched)
- exp_pheno_merged12 <- exp_pheno_merged12[order(exp_pheno_merged12$RID), ]
- rownames(exp_pheno_merged12) <- NULL
- exp_pheno_merged12 <- left_join(exp_pheno_merged12, nfl_slopes, by = "RID")
- # Amyloid PET
- ab_pet <- left_join(ab_pet, exp_pheno_merged12 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- ab_pet <- ab_pet %>%
- arrange(RID, PET_Edate) %>%
- group_by(RID) %>%
- mutate(PET_time=lubridate::time_length(difftime(PET_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(PET_Edate))) %>%
- ungroup()
- # Compute individualized slopes of amyloid centiloids
- model <- lmer(CENTILOIDS ~ PET_time + (PET_time | RID), subset = n > 1, data = ab_pet, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "CENTILOIDS_slope"
- slopes$RID <- row.names(slopes)
- ab_pet <- left_join(ab_pet, slopes, by = "RID")
- ab_pet <- ab_pet %>%
- dplyr::select(-c(n, GeneX_Edate, PET_time))
- ab_slopes <- ab_pet %>%
- dplyr::select(RID, CENTILOIDS_slope) %>%
- distinct()
- exp_pheno_merged13 <- merge(exp_pheno_merged12, ab_pet %>%
- dplyr::select(-CENTILOIDS_slope), by = "RID", all.x = TRUE)
- exp_pheno_merged14 <- gene_within1yr(exp_pheno_merged13, PET_Edate, df2_VISCODE = FALSE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged12, exp_pheno_merged14, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged15 <- bind_rows(exp_pheno_merged14, unmatched)
- exp_pheno_merged15 <- exp_pheno_merged15[order(exp_pheno_merged15$RID), ]
- rownames(exp_pheno_merged15) <- NULL
- exp_pheno_merged15 <- left_join(exp_pheno_merged15, ab_slopes, by = "RID")
- exp_pheno_merged15 <- merge(exp_pheno_merged15, last_negative_date_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- merge(exp_pheno_merged15, first_positive_date_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- merge(exp_pheno_merged15, last_negative_date2_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- merge(exp_pheno_merged15, first_positive_date2_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- merge(exp_pheno_merged15, last_negative_date3_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- merge(exp_pheno_merged15, first_positive_date3_pet, by = "RID", all.x = TRUE)
- exp_pheno_merged15 <- exp_pheno_merged15[order(exp_pheno_merged15$RID), ]
- rownames(exp_pheno_merged15) <- NULL
- # Ensure amyloid PET positivity follows expected logic
- # Prioritize using available data, then apply expected logic where appropriate
- # (e.g., can progress from negative to positive; once positive stays positive)
- exp_pheno_merged16 <- exp_pheno_merged15 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_cross = case_when(
- !is.na(last_AB_neg_date) & !is.na(first_AB_pos_date) & is.na(amyloid_pos_pet_cross) & first_AB_pos_date < last_AB_neg_date ~ amyloid_pos_pet_cross,
- !is.na(last_AB_neg_date) & is.na(amyloid_pos_pet_cross) & GeneX_Edate <= last_AB_neg_date ~ 0,
- !is.na(first_AB_pos_date) & is.na(amyloid_pos_pet_cross) & GeneX_Edate >= first_AB_pos_date ~ 1,
- TRUE ~ amyloid_pos_pet_cross)) %>%
- ungroup()
- exp_pheno_merged16 <- exp_pheno_merged16 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_long = case_when(
- !is.na(last_AB_neg_date2) & !is.na(first_AB_pos_date2) & is.na(amyloid_pos_pet_long) & first_AB_pos_date2 < last_AB_neg_date2 ~ amyloid_pos_pet_long,
- !is.na(last_AB_neg_date2) & is.na(amyloid_pos_pet_long) & GeneX_Edate <= last_AB_neg_date2 ~ 0,
- !is.na(first_AB_pos_date2) & is.na(amyloid_pos_pet_long) & GeneX_Edate >= first_AB_pos_date2 ~ 1,
- TRUE ~ amyloid_pos_pet_long)) %>%
- ungroup()
- exp_pheno_merged16 <- exp_pheno_merged16 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_centiloid = case_when(
- !is.na(last_AB_neg_date3) & !is.na(first_AB_pos_date3) & is.na(amyloid_pos_pet_centiloid) & first_AB_pos_date3 < last_AB_neg_date3 ~ amyloid_pos_pet_centiloid,
- !is.na(last_AB_neg_date3) & is.na(amyloid_pos_pet_centiloid) & GeneX_Edate <= last_AB_neg_date3 ~ 0,
- !is.na(first_AB_pos_date3) & is.na(amyloid_pos_pet_centiloid) & GeneX_Edate >= first_AB_pos_date3 ~ 1,
- TRUE ~ amyloid_pos_pet_centiloid)) %>%
- ungroup() %>%
- dplyr::select(-c("last_AB_neg_date", "first_AB_pos_date", "last_AB_neg_date2", "first_AB_pos_date2", "last_AB_neg_date3", "first_AB_pos_date3"))
- # CSF biomarkers
- upenn_csf <- left_join(upenn_csf, exp_pheno_merged16 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- upenn_csf <- upenn_csf %>%
- arrange(RID, CSF_Edate) %>%
- group_by(RID) %>%
- mutate(CSF_time=lubridate::time_length(difftime(CSF_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(CSF_Edate))) %>%
- ungroup()
- # Compute individualized slopes of CSF p-tau181
- model <- lmer(pTAU181_csf ~ CSF_time + (CSF_time | RID), subset = n > 1, data = upenn_csf, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "pTAU181_csf_slope"
- slopes$RID <- row.names(slopes)
- upenn_csf <- left_join(upenn_csf, slopes, by = "RID")
- upenn_csf <- upenn_csf %>%
- dplyr::select(-c(n, GeneX_Edate, CSF_time))
- csf_slopes <- upenn_csf %>%
- dplyr::select(RID, pTAU181_csf_slope) %>%
- distinct()
- exp_pheno_merged17 <- merge(exp_pheno_merged16, upenn_csf %>%
- dplyr::select(-pTAU181_csf_slope) %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- exp_pheno_merged18 <- gene_within1yr(exp_pheno_merged17, CSF_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged16, exp_pheno_merged18, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged19 <- bind_rows(exp_pheno_merged18, unmatched)
- exp_pheno_merged19 <- exp_pheno_merged19[order(exp_pheno_merged19$RID), ]
- rownames(exp_pheno_merged19) <- NULL
- exp_pheno_merged19 <- left_join(exp_pheno_merged19, csf_slopes, by = "RID")
- exp_pheno_merged19 <- merge(exp_pheno_merged19, last_negative_date_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- merge(exp_pheno_merged19, first_positive_date_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- merge(exp_pheno_merged19, last_negative_date2_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- merge(exp_pheno_merged19, first_positive_date2_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- merge(exp_pheno_merged19, last_negative_date3_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- merge(exp_pheno_merged19, first_positive_date3_csf, by = "RID", all.x = TRUE)
- exp_pheno_merged19 <- exp_pheno_merged19[order(exp_pheno_merged19$RID), ]
- rownames(exp_pheno_merged19) <- NULL
- # Ensure CSF biomarker positivity follows expected logic
- exp_pheno_merged20 <- exp_pheno_merged19 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_csf = case_when(
- !is.na(last_AB_neg_date) & !is.na(first_AB_pos_date) & is.na(amyloid_pos_csf) & first_AB_pos_date < last_AB_neg_date ~ amyloid_pos_csf,
- !is.na(last_AB_neg_date) & is.na(amyloid_pos_csf) & GeneX_Edate <= last_AB_neg_date ~ 0,
- !is.na(first_AB_pos_date) & is.na(amyloid_pos_csf) & GeneX_Edate >= first_AB_pos_date ~ 1,
- TRUE ~ amyloid_pos_csf)) %>%
- ungroup()
- exp_pheno_merged20 <- exp_pheno_merged20 %>%
- group_by(RID) %>%
- mutate(ptau_pos_csf = case_when(
- !is.na(last_TAU_neg_date2) & !is.na(first_TAU_pos_date2) & is.na(ptau_pos_csf) & first_TAU_pos_date2 < last_TAU_neg_date2 ~ ptau_pos_csf,
- !is.na(last_TAU_neg_date2) & is.na(ptau_pos_csf) & GeneX_Edate <= last_TAU_neg_date2 ~ 0,
- !is.na(first_TAU_pos_date2) & is.na(ptau_pos_csf) & GeneX_Edate >= first_TAU_pos_date2 ~ 1,
- TRUE ~ ptau_pos_csf)) %>%
- ungroup()
- exp_pheno_merged20 <- exp_pheno_merged20 %>%
- group_by(RID) %>%
- mutate(ad_pathology_pos_csf = case_when(
- !is.na(last_AD_neg_date3) & !is.na(first_AD_pos_date3) & is.na(ad_pathology_pos_csf) & first_AD_pos_date3 < last_AD_neg_date3 ~ ad_pathology_pos_csf,
- !is.na(last_AD_neg_date3) & is.na(ad_pathology_pos_csf) & GeneX_Edate <= last_AD_neg_date3 ~ 0,
- !is.na(first_AD_pos_date3) & is.na(ad_pathology_pos_csf) & GeneX_Edate >= first_AD_pos_date3 ~ 1,
- TRUE ~ ad_pathology_pos_csf)) %>%
- ungroup() %>%
- dplyr::select(-c("last_AB_neg_date", "first_AB_pos_date", "last_TAU_neg_date2", "first_TAU_pos_date2", "last_AD_neg_date3", "first_AD_pos_date3", "index"))
- # Hippocampal volume and meta-ROI cortical thickness
- fs_cs <- left_join(fs_cs, exp_pheno_merged20 %>%
- dplyr::select(RID, GeneX_Edate), by = "RID")
- fs_cs <- fs_cs %>%
- arrange(RID, MRIcs_Edate) %>%
- group_by(RID) %>%
- mutate(MRI_time=lubridate::time_length(difftime(MRIcs_Edate, GeneX_Edate), "years")) %>%
- mutate(n = length(unique(MRIcs_Edate))) %>%
- ungroup()
- # Compute individualized slopes of hippocampal volume
- model <- lmer(HVa ~ MRI_time + (MRI_time | RID), subset = n > 1, data = fs_cs, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "HVa_slope"
- slopes$RID <- row.names(slopes)
- fs_cs <- left_join(fs_cs, slopes, by = "RID")
- # Compute individualized slopes of meta-ROI cortical thickness
- model <- lmer(meta_ROI ~ MRI_time + (MRI_time | RID), subset = n > 1, data = fs_cs, na.action = na.omit)
- slopes <- ranef(model)$RID[2]
- colnames(slopes)[1] <- "meta_ROI_slope"
- slopes$RID <- row.names(slopes)
- fs_cs <- left_join(fs_cs, slopes, by = "RID")
- fs_slopes <- fs_cs %>%
- dplyr::select(RID, HVa_slope, meta_ROI_slope) %>%
- distinct()
- fs_cs <- fs_cs %>%
- dplyr::select(-c(n, GeneX_Edate, MRI_time, HVa_slope, meta_ROI_slope))
- exp_pheno_merged21 <- merge(exp_pheno_merged20, fs_cs %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- exp_pheno_merged22 <- gene_within1yr(exp_pheno_merged21, MRIcs_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(exp_pheno_merged20, exp_pheno_merged22, by = c("RID", "GeneX_Edate"))
- exp_pheno_merged23 <- bind_rows(exp_pheno_merged22, unmatched)
- exp_pheno_merged23 <- exp_pheno_merged23[order(exp_pheno_merged23$RID), ]
- rownames(exp_pheno_merged23) <- NULL
- exp_pheno_merged23 <- left_join(exp_pheno_merged23, fs_slopes, by = "RID")
- message("Now saving final gene expression phenotype data for analysis...")
- # Save final gene expression phenotype data for analysis
- write.csv(exp_pheno_merged23, file.path(base_dir, "GeneX_pheno_final.csv"), row.names = F)
- #______________________________________Merge Prepared ADNI Tables With DNA Methylation Metadata______________________________________#####
- ## Create and Prepare DNA Methylation Metadata Table ####
- message("Now reading in and preparing DNA methylation metadata table...")
- dnam_pheno <- read.csv(file.path(base_dir, "Sample_Sheet_rv1.csv"), stringsAsFactor = FALSE)
- # Exclude the 14 samples not in metadata (failed QC by ADNI Genetics Core)
- dnam_pheno <- dnam_pheno[!is.na(dnam_pheno$RID), ]
- dnam_pheno <- dnam_pheno %>%
- dplyr::rename(DNAm_Edate = Edate,
- DNAm_DateDrawn = DateDrawn) %>%
- dplyr::select(-Basename)
- dnam_pheno$RID <- addZeros(dnam_pheno$RID)
- dnam_pheno$DNAm_Edate <- as.Date(dnam_pheno$DNAm_Edate, "%m/%d/%Y")
- dnam_pheno$DNAm_DateDrawn <- as.Date(dnam_pheno$DNAm_DateDrawn, "%m/%d/%Y")
- ### Combine ADNI Tables with DNA Methylation Metadata ####
- # Main demographics
- dnam_pheno_merged <- left_join(demog_merged, dnam_pheno, by = "RID")
- dnam_pheno_merged <- dnam_pheno_merged[!is.na(dnam_pheno_merged$barcodes),]
- dnam_pheno_merged$AGE <- round(time_length(difftime(dnam_pheno_merged$DNAm_DateDrawn,
- dnam_pheno_merged$PTDOB),
- "years"), digits = 1)
- dnam_pheno_merged <- relocate(dnam_pheno_merged, "AGE", .before = "AGE_bl")
- dnam_pheno_merged <- relocate(dnam_pheno_merged, "barcodes", .after = "RID")
- # APOE genotypes
- dnam_pheno_merged <- left_join(dnam_pheno_merged, apoeres_df, by = "RID")
- # CSF AB/TAU estimated age of positivity onset
- dnam_pheno_merged <- left_join(dnam_pheno_merged, csf_age, by = "RID")
- dnam_pheno_merged <- dnam_pheno_merged %>%
- mutate(AB_time = AGE - AB_AGE,
- TAU_time = AGE - TAU_AGE)
- # Diagnosis
- dnam_pheno_merged1 <- merge(dnam_pheno_merged, dx, by = "RID", all.x = TRUE)
- dnam_pheno_merged2 <- dnam_within1yr(dnam_pheno_merged1, DX_Edate, df2_VISCODE = FALSE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged, dnam_pheno_merged2, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged3 <- bind_rows(dnam_pheno_merged2, unmatched)
- dnam_pheno_merged3 <- dnam_pheno_merged3 %>%
- distinct(barcodes, .keep_all = TRUE)
- # Compute time2event and MCI progressor status
- mci <- left_join(dx, dnam_pheno_merged3 %>%
- dplyr::select(RID, DNAm_Edate), by = "RID")
- mci <- mci %>%
- group_by(RID, DNAm_Edate) %>%
- mutate(date_diff = abs(lubridate::time_length(difftime(DNAm_Edate, DX_Edate), "years")),
- min_match_date = min(date_diff)) %>%
- ungroup()
- first <- mci %>%
- arrange(RID, DNAm_Edate) %>%
- group_by(RID) %>%
- mutate(DX_bl = DXSUM[min_match_date == date_diff][1],
- bl_DX_date = DX_Edate[min_match_date == date_diff][1],
- bl_DX_date = as.Date(bl_DX_date, "%Y-%m-%d")) %>%
- filter(!is.na(DX_bl)) %>%
- ungroup() %>%
- distinct(RID, .keep_all = T) %>%
- dplyr::select(RID, DX_bl, bl_DX_date)
- mci <- left_join(mci, first, by = "RID")
- mci <- mci %>%
- arrange(RID, DNAm_Edate, DX_Edate) %>%
- group_by(RID) %>%
- mutate(last_DX_date = max(DX_Edate, na.rm = T),
- last_DX_date = as.Date(last_DX_date, "%Y-%m-%d"),
- window = DX_Edate >= DNAm_Edate[[1]] & DX_Edate <= last_DX_date) %>%
- filter(window == TRUE) %>%
- mutate(first_AD_date = min(DX_Edate[DXSUM == "Dementia" & window], na.rm = T),
- first_AD_date = as.Date(first_AD_date, "%Y-%m-%d"),
- last_MCI_date = max(DX_Edate[DXSUM == "MCI" & window], na.rm = T),
- last_MCI_date = as.Date(last_MCI_date, "%Y-%m-%d"),
- swap = DX_bl == "MCI" & window & first_AD_date <= last_MCI_date,
- AD_flag = any(DXSUM == "Dementia" & window),
- Progressor = case_when(
- DX_bl == "MCI" & AD_flag & swap ~ NA,
- DX_bl == "MCI" & AD_flag ~ 1 ,
- DX_bl == "MCI" & !AD_flag ~ 0,
- TRUE ~ NA
- )) %>%
- ungroup()
- mci[sapply(mci, is.infinite)] <- NA
- mci <- mci %>%
- group_by(RID) %>%
- mutate(time2event = ifelse(Progressor == 1,
- lubridate::time_length(difftime(first_AD_date, DNAm_Edate), "years"),
- lubridate::time_length(difftime(last_DX_date, DNAm_Edate), "years"))) %>%
- filter(time2event >= 0)
- dnam_pheno_merged3 <- left_join(dnam_pheno_merged3, mci %>%
- dplyr::select(RID, DX_bl, bl_DX_date, last_DX_date, first_AD_date, last_MCI_date, Progressor, time2event), by = "RID")
- dnam_pheno_merged3 <- dnam_pheno_merged3 %>%
- distinct(barcodes, .keep_all = TRUE)
- dnam_pheno_merged3 <- dnam_pheno_merged3 %>%
- mutate(DXSUM2 = case_when(
- (Progressor == 0) ~ "Non-Progressor",
- (Progressor == 1) ~ "Progressor MCI",
- (DXSUM == "CU") ~ "CU",
- (DXSUM == "Dementia") ~ "Dementia",
- TRUE ~ DXSUM
- ))
- # Cognitive performance & individualized slopes of CDRSB and MMSE
- dnam_pheno_merged4 <- merge(dnam_pheno_merged3, cdr %>%
- dplyr::select(-CDRSB_slope) %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- dnam_pheno_merged5 <- dnam_within1yr(dnam_pheno_merged4, CDR_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged3, dnam_pheno_merged5, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged6 <- bind_rows(dnam_pheno_merged5, unmatched)
- dnam_pheno_merged6 <- dnam_pheno_merged6[order(dnam_pheno_merged6$RID, dnam_pheno_merged6$DNAm_Edate), ]
- rownames(dnam_pheno_merged6) <- NULL
- dnam_pheno_merged6 <- left_join(dnam_pheno_merged6, cdr_slopes, by = "RID")
- dnam_pheno_merged6 <- left_join(dnam_pheno_merged6, adas, by = c("RID", "VISCODE"))
- dnam_pheno_merged7 <- merge(dnam_pheno_merged6, mmse %>%
- dplyr::select(-MMSE_slope), by = "RID", all.x = TRUE)
- dnam_pheno_merged8 <- dnam_within1yr(dnam_pheno_merged7, MMSE_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged6, dnam_pheno_merged8, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged9 <- bind_rows(dnam_pheno_merged8, unmatched)
- dnam_pheno_merged9 <- dnam_pheno_merged9[order(dnam_pheno_merged9$RID, dnam_pheno_merged9$DNAm_Edate), ]
- rownames(dnam_pheno_merged9) <- NULL
- dnam_pheno_merged9 <- left_join(dnam_pheno_merged9, mmse_slopes, by = "RID")
- # Plasma NfL & individualized slopes
- dnam_pheno_merged10 <- merge(dnam_pheno_merged9, nfl_data %>%
- dplyr::select(-NfL_B_slope), by = "RID", all.x = TRUE)
- dnam_pheno_merged11 <- dnam_within1yr(dnam_pheno_merged10, NfL_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged9, dnam_pheno_merged11, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged12 <- bind_rows(dnam_pheno_merged11, unmatched)
- dnam_pheno_merged12 <- dnam_pheno_merged12[order(dnam_pheno_merged12$RID, dnam_pheno_merged12$DNAm_Edate), ]
- rownames(dnam_pheno_merged12) <- NULL
- dnam_pheno_merged12 <- left_join(dnam_pheno_merged12, nfl_slopes, by = "RID")
- # Amyloid PET & individualized slopes of centiloids
- dnam_pheno_merged13 <- merge(dnam_pheno_merged12, ab_pet %>%
- dplyr::select(-CENTILOIDS_slope), by = "RID", all.x = TRUE)
- dnam_pheno_merged14 <- dnam_within1yr(dnam_pheno_merged13, PET_Edate, df2_VISCODE = FALSE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged12, dnam_pheno_merged14, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged15 <- bind_rows(dnam_pheno_merged14, unmatched)
- dnam_pheno_merged15 <- dnam_pheno_merged15[order(dnam_pheno_merged15$RID, dnam_pheno_merged15$DNAm_Edate), ]
- rownames(dnam_pheno_merged15) <- NULL
- dnam_pheno_merged15 <- left_join(dnam_pheno_merged15, ab_slopes, by = "RID")
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, last_negative_date_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, first_positive_date_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, last_negative_date2_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, first_positive_date2_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, last_negative_date3_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- merge(dnam_pheno_merged15, first_positive_date3_pet, by = "RID", all.x = TRUE)
- dnam_pheno_merged15 <- dnam_pheno_merged15[order(dnam_pheno_merged15$RID, dnam_pheno_merged15$DNAm_Edate), ]
- rownames(dnam_pheno_merged15) <- NULL
- # Ensure amyloid PET positivity follows expected logic
- dnam_pheno_merged16 <- dnam_pheno_merged15 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_cross = case_when(
- !is.na(last_AB_neg_date) & !is.na(first_AB_pos_date) & is.na(amyloid_pos_pet_cross) & first_AB_pos_date < last_AB_neg_date ~ amyloid_pos_pet_cross,
- !is.na(last_AB_neg_date) & is.na(amyloid_pos_pet_cross) & DNAm_Edate <= last_AB_neg_date ~ 0,
- !is.na(first_AB_pos_date) & is.na(amyloid_pos_pet_cross) & DNAm_Edate >= first_AB_pos_date ~ 1,
- TRUE ~ amyloid_pos_pet_cross)) %>%
- ungroup()
- dnam_pheno_merged16 <- dnam_pheno_merged16 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_long = case_when(
- !is.na(last_AB_neg_date2) & !is.na(first_AB_pos_date2) & is.na(amyloid_pos_pet_long) & first_AB_pos_date2 < last_AB_neg_date2 ~ amyloid_pos_pet_long,
- !is.na(last_AB_neg_date2) & is.na(amyloid_pos_pet_long) & DNAm_Edate <= last_AB_neg_date2 ~ 0,
- !is.na(first_AB_pos_date2) & is.na(amyloid_pos_pet_long) & DNAm_Edate >= first_AB_pos_date2 ~ 1,
- TRUE ~ amyloid_pos_pet_long)) %>%
- ungroup()
- dnam_pheno_merged16 <- dnam_pheno_merged16 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_pet_centiloid = case_when(
- !is.na(last_AB_neg_date3) & !is.na(first_AB_pos_date3) & is.na(amyloid_pos_pet_centiloid) & first_AB_pos_date3 < last_AB_neg_date3 ~ amyloid_pos_pet_centiloid,
- !is.na(last_AB_neg_date3) & is.na(amyloid_pos_pet_centiloid) & DNAm_Edate <= last_AB_neg_date3 ~ 0,
- !is.na(first_AB_pos_date3) & is.na(amyloid_pos_pet_centiloid) & DNAm_Edate >= first_AB_pos_date3 ~ 1,
- TRUE ~ amyloid_pos_pet_centiloid)) %>%
- ungroup() %>%
- dplyr::select(-c("last_AB_neg_date", "first_AB_pos_date", "last_AB_neg_date2", "first_AB_pos_date2", "last_AB_neg_date3", "first_AB_pos_date3"))
- # CSF biomarkers & individualized slopes of p-tau181
- dnam_pheno_merged17 <- merge(dnam_pheno_merged16, upenn_csf %>%
- dplyr::select(-pTAU181_csf_slope) %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- dnam_pheno_merged18 <- dnam_within1yr(dnam_pheno_merged17, CSF_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged16, dnam_pheno_merged18, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged19 <- bind_rows(dnam_pheno_merged18, unmatched)
- dnam_pheno_merged19 <- dnam_pheno_merged19[order(dnam_pheno_merged19$RID, dnam_pheno_merged19$DNAm_Edate), ]
- rownames(dnam_pheno_merged19) <- NULL
- dnam_pheno_merged19 <- left_join(dnam_pheno_merged19, csf_slopes, by = "RID")
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, last_negative_date_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, first_positive_date_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, last_negative_date2_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, first_positive_date2_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, last_negative_date3_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- merge(dnam_pheno_merged19, first_positive_date3_csf, by = "RID", all.x = TRUE)
- dnam_pheno_merged19 <- dnam_pheno_merged19[order(dnam_pheno_merged19$RID, dnam_pheno_merged19$DNAm_Edate), ]
- rownames(dnam_pheno_merged19) <- NULL
- # Ensure CSF biomarker positivity follows expected logic
- dnam_pheno_merged20 <- dnam_pheno_merged19 %>%
- group_by(RID) %>%
- mutate(amyloid_pos_csf = case_when(
- !is.na(last_AB_neg_date) & !is.na(first_AB_pos_date) & is.na(amyloid_pos_csf) & first_AB_pos_date < last_AB_neg_date ~ amyloid_pos_csf,
- !is.na(last_AB_neg_date) & is.na(amyloid_pos_csf) & DNAm_Edate <= last_AB_neg_date ~ 0,
- !is.na(first_AB_pos_date) & is.na(amyloid_pos_csf) & DNAm_Edate >= first_AB_pos_date ~ 1,
- TRUE ~ amyloid_pos_csf)) %>%
- ungroup()
- dnam_pheno_merged20 <- dnam_pheno_merged20 %>%
- group_by(RID) %>%
- mutate(ptau_pos_csf = case_when(
- !is.na(last_TAU_neg_date2) & !is.na(first_TAU_pos_date2) & is.na(ptau_pos_csf) & first_TAU_pos_date2 < last_TAU_neg_date2 ~ ptau_pos_csf,
- !is.na(last_TAU_neg_date2) & is.na(ptau_pos_csf) & DNAm_Edate <= last_TAU_neg_date2 ~ 0,
- !is.na(first_TAU_pos_date2) & is.na(ptau_pos_csf) & DNAm_Edate >= first_TAU_pos_date2 ~ 1,
- TRUE ~ ptau_pos_csf)) %>%
- ungroup()
- dnam_pheno_merged20 <- dnam_pheno_merged20 %>%
- group_by(RID) %>%
- mutate(ad_pathology_pos_csf = case_when(
- !is.na(last_AD_neg_date3) & !is.na(first_AD_pos_date3) & is.na(ad_pathology_pos_csf) & first_AD_pos_date3 < last_AD_neg_date3 ~ ad_pathology_pos_csf,
- !is.na(last_AD_neg_date3) & is.na(ad_pathology_pos_csf) & DNAm_Edate <= last_AD_neg_date3 ~ 0,
- !is.na(first_AD_pos_date3) & is.na(ad_pathology_pos_csf) & DNAm_Edate >= first_AD_pos_date3 ~ 1,
- TRUE ~ ad_pathology_pos_csf)) %>%
- ungroup() %>%
- dplyr::select(-c("last_AB_neg_date", "first_AB_pos_date", "last_TAU_neg_date2", "first_TAU_pos_date2", "last_AD_neg_date3", "first_AD_pos_date3"))
- # Hippocampal volume and meta-ROI cortical thickness & individualized slopes
- dnam_pheno_merged21 <- merge(dnam_pheno_merged20, fs_cs %>%
- dplyr::rename(VISCODE2 = VISCODE), by = "RID", all.x = TRUE)
- dnam_pheno_merged22 <- dnam_within1yr(dnam_pheno_merged21, MRIcs_Edate, df2_VISCODE = TRUE)
- # Get back rest of the data
- unmatched <- anti_join(dnam_pheno_merged20, dnam_pheno_merged22, by = c("RID", "DNAm_Edate"))
- dnam_pheno_merged23 <- bind_rows(dnam_pheno_merged22, unmatched)
- dnam_pheno_merged23 <- dnam_pheno_merged23[order(dnam_pheno_merged23$RID, dnam_pheno_merged23$DNAm_Edate), ]
- rownames(dnam_pheno_merged23) <- NULL
- dnam_pheno_merged23 <- left_join(dnam_pheno_merged23, fs_slopes, by = "RID")
- # Estimate immune cell-type proportions
- adni_se <- readRDS(file.path(proc_dir, "dasen_se_all_rv1.RDS"))
- beta <- as.matrix(assays(adni_se)$DNAm)
- rm(adni_se)
- data(centDHSbloodDMC.m)
- BloodFrac.m <- epidish(beta.m = beta, ref.m = centDHSbloodDMC.m, method = "RPC")$estF
- cell_types <- as.data.frame(BloodFrac.m)
- cell_types <- tibble::rownames_to_column(cell_types, "barcodes")
- cell_types$Gran <- cell_types$Neutro + cell_types$Eosino
- dnam_pheno_merged24 <- left_join(dnam_pheno_merged23, cell_types, by = "barcodes")
- message("Now saving final (longitudinal) DNA methylation phenotype data...")
- # Save final DNA methylation phenotype data for analysis
- write.csv(dnam_pheno_merged24, file.path(base_dir, "DNAm_pheno_final.csv"), row.names = F)
- ## Create Subset of DNA Methylation Phenotype Data Aligned to Time of Gene Expression Profiling ####
- # Matched at same time-point or within 1-year of gene expression profiling
- message("Now creating a cross-sectional subset of DNA methylation phenotype data based on time of gene expression profiling...")
- # Read in phenotype data from both -omics that we just created
- exp_pheno <- read.csv(file.path(base_dir, "GeneX_pheno_final.csv"), stringsAsFactor = FALSE, check.names = FALSE)
- exp_pheno$RID <- addZeros(exp_pheno$RID)
- dnam_pheno <- read.csv(file.path(base_dir, "DNAm_pheno_final.csv"), stringsAsFactor = FALSE)
- dnam_pheno$RID <- addZeros(dnam_pheno$RID)
- # Identify the cross-sectional matched time-point for DNA methylation based on time of gene expression profiling
- merged <- merge(dnam_pheno %>%
- dplyr::select(-c(DXSUM2, Progressor, time2event)) %>%
- dplyr::rename(DNAm_DX_bl = DX_bl),
- exp_pheno %>%
- dplyr::select(RID, VISCODE, GeneX_Edate, DX_Edate, DXSUM, DXSUM2, Progressor, time2event) %>%
- dplyr::rename(GeneX_VISCODE = VISCODE,
- GeneX_DXSUM = DXSUM,
- GeneX_DX_Edate = DX_Edate), by = "RID", all.x = TRUE) %>%
- # Use time of gene expression profiling as the anchor
- mutate(date_diff = abs(lubridate::time_length(difftime(DNAm_Edate, GeneX_Edate), "years"))) %>%
- mutate(Edate_match = ifelse(date_diff <= 1, 1, 0),
- VISCODE_match = ifelse(VISCODE == GeneX_VISCODE, 1, 0)) %>%
- group_by(RID) %>%
- mutate(min_match_date = min(date_diff)) %>%
- # Same time = VISCODE_match; within 1-year = Edate_match
- mutate(match_priority = case_when(
- VISCODE_match == 1 ~ 1,
- Edate_match == 1 ~ 2,
- TRUE ~ 3
- )) %>%
- # Only keep matches
- filter(match_priority <= 2) %>%
- # If multiple matches, keep min date diff
- filter(min_match_date == date_diff) %>%
- # Replace any missing diagnosis, EXAMDATE, and VISCODE using the ones from gene expression pheno data
- mutate(DXSUM = ifelse(is.na(DXSUM) & date_diff == 0 & !is.na(GeneX_DXSUM), GeneX_DXSUM, DXSUM),
- VISCODE = ifelse(is.na(VISCODE) & date_diff == 0 & !is.na(GeneX_VISCODE), GeneX_VISCODE, VISCODE),
- DX_Edate = ifelse(is.na(DX_Edate) & date_diff == 0 & !is.na(GeneX_DX_Edate), GeneX_DX_Edate, DX_Edate)
- ) %>%
- # Filter missing B, a convenient proxy for the replicates that were excluded during original DNA methylation preprocessing
- filter(!is.na(B)) %>%
- ungroup() %>%
- dplyr::select(-c(Edate_match, VISCODE_match, match_priority, min_match_date, date_diff,
- GeneX_DXSUM, GeneX_DX_Edate))
- message("Now saving the final DNA methylation phenotype data for analysis...")
- # Save the final DNA methylation phenotype data (same time or within 1-year of gene expression) for cross-sectional analysis
- write.csv(merged, file.path(base_dir, "DNAm_pheno_final_shared.csv"), row.names = F)
phenotype.R at commit 5874e91, no license · at the source
Overview
- Department of Radiology and Biomedical Imaging, University of California,San Francisco, CA USA
- UC Berkeley - UCSF Graduate Program in Bioengineering,Berkeley, CA USA
- Department of Medicine, Division of Rheumatology, University of California,San Francisco, CA USA
- CoLabs, University of California,San Francisco, CA USA
- Northern California Institute for Research and Education (NCIRE),San Francisco, CA USA
- Bakar Computational Health Sciences Institute, University of California,San Francisco, CA USA
- Department of Pediatrics, University of California,San Francisco, CA USA
Abstract
Leveraging multi-omics to better understand the molecular signatures and pathways underlying Alzheimer’s disease (AD) pathogenesis is critical for early diagnosis and disease modifying interventions. We performed peripheral blood transcriptome (N = 669) and epigenome microarray analyses (N = 553) on non-Hispanic white participants from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) to identify molecular signatures of AD. We identified specific transcripts (e.g., MAPK14, GM2A, CD177) and co-expression networks that were dysregulated in AD, marked by a strong influence of APOE ε4 genotype, and characterized by a consistent pattern of immune activation, inflammation, and metabolic suppression. Further, these peripheral signatures were linked to central AD pathology (amyloid PET, CSF p-tau181) and neurodegeneration (plasma NfL, regional atrophy), with two genes, MXD3 and NR4A1, identified as protective against progression from MCI to AD. Our work emphasizes the importance of APOE genotypes in AD pathophysiology and highlights potential targets for biomarker discovery and personalized therapeutic strategies.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.
cind/adni-mrna-dnam-analysis
5874e916559bb984d4614f4299af1611056a7fa8, 19 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
17 files
- R/
DNAm-preprocessing.R , R, 197 lines, 1 match - R/
DNAm-runWGCNA-All.R , R, 80 lines - R/
DNAm-runWGCNA-Men.R , R, 80 lines - R/
DNAm-runWGCNA-Women.R , R, 80 lines - R/
GeneX-runWGCNA-All.R , R, 80 lines - R/
GeneX-runWGCNA-Men.R , R, 80 lines - R/
GeneX-runWGCNA-Women.R , R, 80 lines - R/
MOFA.R , R, 891 lines, 2 matches - R/
WGCNA.R , R, 1,590 lines, 4 matches - R/
covarRegress.R , R, 662 lines - R/
differential_analyses.R , R, 1,262 lines, 1 match - R/
install_packages.R , R, 117 lines - R/
phenotype.R , R, 1,471 lines, 5 matches - R/
runMOFA.R , R, 96 lines - R/
tables.R , R, 234 lines, 2 matches - R/
wgcna2cytoscape.R , R, 296 lines, 1 match - README.md, Text, 28 lines
Code availability
The code used in this manuscript can be found at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 16 scripts, each with its path and the digest of its content;
- 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
Data used in the preparation of this article were obtained from the ADNI database (adni.loni.usc.edu). As such, the investigators within the ADNI contributed to the design and implementation of ADNI and/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Springer Science+Business Media
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 5 keywords, 2 funders, 79 references.
Cite
This paper
Mitchell, B. A., Hausle, I., Smith, S., Thropp, P., Sirota, M., & Tosun, D. (2026). Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease. NPJ dementia, 2(1), 74. https://
BibTeX
@article{mitchell2026per
author = {Mitchell, Brendan A. and Hausle, Isabella and Smith, Sara and Thropp, Pamela and Sirota, Marina and Tosun, Duygu},
title = {{Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease}},
journal = {NPJ dementia},
year = {2026},
month = aug,
volume = {2},
number = {1},
pages = {74},
publisher = {Springer Science+Business Media},
issn = {3005-1940},
doi = {10.1038/
url = {https://
pmid = {42687894},
pmcid = {PMC13529590}
}
RIS
TY - JOUR
AU - Mitchell, Brendan A.
AU - Hausle, Isabella
AU - Smith, Sara
AU - Thropp, Pamela
AU - Sirota, Marina
AU - Tosun, Duygu
TI - Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease
T2 - NPJ dementia
J2 - NPJ Dement
PY - 2026
DA - 2026/
VL - 2
IS - 1
SP - 74
SN - 3005-1940
PB - Springer Science+Business Media
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease",
"container-title": "NPJ dementia",
"author": [
{
"family": "Mitchell",
"given": "Brendan A."
},
{
"family": "Hausle",
"given": "Isabella"
},
{
"family": "Smith",
"given": "Sara"
},
{
"family": "Thropp",
"given": "Pamela"
},
{
"family": "Sirota",
"given": "Marina"
},
{
"family": "Tosun",
"given": "Duygu"
}
],
"container-title-short":
"volume": "2",
"issue": "1",
"page": "74",
"DOI": "10.1038/
"PMID": "42687894",
"PMCID": "PMC13529590",
"ISSN": "3005-1940",
"publisher": "Springer Science+Business Media",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
31
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102682 [code]
- TET CpG sequence-context-specifi
c DNA demethylation shapes progression of IDH-mutant gliomas. Journal: Cell reports. MedicineIn common: survival, limma, UMAP, 9 other tools, genetics / omics, 2 references - [2] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: survival, WGCNA, limma, 10 other tools, genetics / omics
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: survival, WGCNA, limma, 10 other tools
- [4] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: WGCNA, limma, UMAP, 8 other tools, genetics / omics, 2 references
- [5] doi:10.1016/j.isci.2026.115657 [code]
- Integration of machine learning to develop a disulfidptosis model for predicting glioma prognosis, immunotherapy response, and drug.Journal: iScienceIn common: survival, WGCNA, limma, 8 other tools
- [6] doi:10.1038/s41593-026-02367-0 [code]
- A reproducible three-dimensional model of human brain tissue to investigate physiological and disease-associated microglia phenotypes.Journal: Nature neuroscienceIn common: WGCNA, limma, clusterProfiler, 6 other tools, Alzheimer's / dementia, 2 references
- [7] doi:10.3390/ijms27104466 [code]
- Uncovering the Key Circuit FOSL2/
FOS/ EGR3/ EGR1, Contributing to the Hyperexcitability of Excitatory Neurons in the Epileptic Temporal Cortex and Hippocampus. Journal: International journal of molecular sciencesIn common: limma, UMAP, clusterProfiler, 7 other tools, genetics / omics, 1 reference - [8] doi:10.1038/s41380-026-03629-w [code]
- Maternal fasting during early gestation induces epigenetic alterations and schizophrenia-related phenotypes.Journal: Molecular psychiatryIn common: WGCNA, limma, clusterProfiler, 4 other tools, genetics / omics, 3 references
- [9] doi:10.1038/s41467-026-77170-3 [code]
- DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.Journal: Nature communicationsIn common: survival, limma, broom, 6 other tools, genetics / omics
- [10] doi:10.1038/s41467-026-73305-8 [code]
- Comparative analysis of the cellular landscape in mammalian striatum.Journal: Nature communicationsIn common: WGCNA, clusterProfiler, ComplexHeatmap, 5 other tools, genetics / omics, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 16 scripts, and 16 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:6213dcbde0383835…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
