The aging epigenome: integrative analyses reveal intersection with Alzheimer's disease.
The 6 matches
- [1] § Methods › Differentially methylated regions analysis ↔ code/utility/coMethDMR_aux.R, lines 151–199 · score 0.63 · linear regression, co methylated, medians, errors, models, CpGs
- [2] § Methods › Pathway analysis ↔ code/utility/pathway.R, lines 69–184 · score 0.61 · methylRRA, methylGSA, pathways, UCSC, enriched, enrichment
- [3] § Methods › Inflation assessment and correction ↔ code/utility/annotation_and_bacon.R, lines 483–626 · score 0.57 · inflation factors, bacon corrected, bias, genomic
- [4] § Methods › Inflation assessment and correction ↔ code/utility/annotation_and_bacon.R, lines 483–626 · score 0.57 · Genomic inflation factors, lambda, bias, bacon
- [5] § Methods › Association of DNA methylation at individual CpGs with chronological age ↔ code/dmr/coMethDMR.Rmd, lines 212–240 · score 0.54 · CD4T, Gran, Mono, NK, sex, beta
- [6] § Methods › Pre-processing of DNA methylation data ↔ code/dmr/coMethDMR.Rmd, lines 187–210 · score 0.50 · bisulfite conversion, status, EPIC, FHS, DNAm
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 629 lines · 24 KB · MIT · 2 matches
- #######################################################################################################
- # =================================================================================================== #
- # Function for annotation and inflation adjusted
- # =================================================================================================== #
- #######################################################################################################
- library(SummarizedExperiment)
- library(tidyverse)
- library(bacon)
- library(GWASTools)
- library(minfi)
- library(rGREAT)
- # ===================================================================================================
- # Annotation Single CpGs
- # ===================================================================================================
- # Function: get_anno_gr
- # ------------------------------------------------------------------------------
- # Description:
- # Retrieves annotation information as a GRanges object for the specified methylation array and genome.
- #
- # Parameters:
- # array : Character. An input array (e.g., methylation array) for which genomic annotations
- # are required. Supported options include "HM450", "EPICv1", "EPIC", and "EPICv2".
- # genome : Character. A string specifying the genome build used for annotation, such as "hg19"
- # or "hg38".
- # dir.data.aux : (Optional) Character. A directory path containing auxiliary data used by the
- # annotation function, particularly for reading external manifest files when using
- # genome "hg38".
- #
- # Returns:
- # A GRanges object containing the genomic annotation data.
- get_anno_gr <- function(array = "HM450",
- genome = "hg19",
- dir.data.aux = NULL) {
- if(array == "HM450"){
- library(IlluminaHumanMethylation450kanno.ilmn12.hg19)
- anno <- minfi::getAnnotation(IlluminaHumanMethylation450kanno.ilmn12.hg19)
- }
- if(array %in% c("EPICv1", "EPIC")){
- if(genome == "hg19") {
- library(IlluminaHumanMethylationEPICanno.ilm10b4.hg19)
- anno <- minfi::getAnnotation(IlluminaHumanMethylationEPICanno.ilm10b4.hg19)
- }
- if(genome == "hg38") {
- anno <- read_csv(
- file.path(dir.data.aux, "infinium-methylationepic-v-1-0-b5-manifest-file.csv"),
- show_col_types = F,
- skip = 7
- )
- }
- }
- if(array == "EPICv2") {
- library(IlluminaHumanMethylationEPICv2anno.20a1.hg38)
- anno <- minfi::getAnnotation(IlluminaHumanMethylationEPICv2anno.20a1.hg38)
- }
- if(array %in% c("EPICv1", "EPIC") & genome == "hg38") {
- anno.gr <- anno %>% makeGRangesFromDataFrame(
- seqnames.field = "CHR_hg38",
- start.field = "Start_hg38", end.field = "End_hg38",
- strand.field = "Strand_hg38",
- keep.extra.columns = T,
- na.rm = T
- )
- } else {
- anno.gr <- anno %>% makeGRangesFromDataFrame(
- start.field = "pos", end.field = "pos", keep.extra.columns = T
- )
- }
- anno.gr
- }
- # ------------------------------------------------------------------------------
- # Function: annotate_results
- # ------------------------------------------------------------------------------
- # Description:
- # Annotates a CpG results data frame by merging genomic annotation (from get_anno_gr) with
- # additional GREAT analysis results.
- #
- # Parameters:
- # result : Data frame. A table containing CpG IDs and associated statistics to annotate.
- # array : Character. The methylation array type (e.g., "HM450", "EPICv1", "EPIC").
- # genome : Character. The genome version ("hg19" or "hg38").
- # dir.data.aux : Character. Directory path for auxiliary data files containing precomputed annotations.
- # save : Logical. Whether to save the annotated result as a CSV file.
- # dir.save : Character. Directory path where the annotated result file will be saved.
- # prefix : Character. A prefix used for naming the output file.
- #
- # Returns:
- # A data frame with additional columns for genomic annotation and GREAT analysis.
- annotate_results <- function(result,
- array = "HM450",
- genome = "hg19",
- dir.data.aux = NULL,
- save = T,
- dir.save = NULL,
- prefix = "Framingham"){
- # load(file.path(dir.data.aux,"E073_15_coreMarks_segments.rda"))
- # load(file.path(dir.data.aux,"meta_analysis_cpgs.rda"))
- anno.gr <- get_anno_gr(array = array, genome = genome, dir.data.aux = dir.data.aux)
- anno_df <- as.data.frame(anno.gr)
- if(genome == "hg19") {
- if(array == "HM450"){
- load(file.path(dir.data.aux,"great_HM450_array_annotation.rda"))
- }
- if(array %in% c("EPICv1", "EPIC")){
- load(file.path(dir.data.aux,"great_EPIC_array_annotation.rda"))
- }
- result <- cbind(
- result,
- anno_df[result$cpg,c("seqnames", "start", "end", "width", "Relation_to_Island", "UCSC_RefGene_Name", "UCSC_RefGene_Group")]
- )
- result <- dplyr::left_join(result, great, by = c("seqnames","start","end","cpg"))
- }
- # Add annotation
- if(genome == "hg38"){
- if(array %in% c("EPICv1", "EPIC")) {
- load(file.path(dir.data.aux,"great_EPIC_array_annotation.rda"))
- anno_df <- anno_df %>%
- mutate(cpg = Name) %>%
- dplyr::select(cpg, seqnames, start, end, width,
- UCSC_RefGene_Group, UCSC_RefGene_Name, Relation_to_UCSC_CpG_Island) %>%
- unique()
- } else {
- load(file.path(dir.data.aux,"great_EPICv2_array_annotation.rda"))
- anno_df <- anno_df %>%
- mutate(cpg = gsub("_.*", "",anno_df$Name)) %>%
- dplyr::select(cpg, seqnames, start, end, width,
- UCSC_RefGene_Group, UCSC_RefGene_Name, Relation_to_Island) %>%
- unique()
- }
- result <- left_join(
- result,
- anno_df
- )
- result <- dplyr::left_join(result, great %>% ungroup %>% dplyr::select(cpg, GREAT_annotation))
- }
- if(save){
- write_csv(
- result,
- file.path(dir.save, paste0(prefix, "_annotated_results.csv"))
- )
- }
- return(result)
- }
- # ===================================================================================================
- # Annotation DMR
- # ===================================================================================================
- # Annotate the regions with E073 15-core marks segmentation states
- # ------------------------------------------------------------------------------
- # Description:
- # Annotates regions with ChromHMM segmentation states based on E073 15-core marks segmentation data.
- #
- # Parameters:
- # result : Data frame. Regions to be annotated.
- # ChmmModels.gr: GRanges object. ChromHMM segmentation data containing state information.
- # region_var : Character. The column name in 'result' representing region identifiers.
- #
- # Returns:
- # The original data frame with an added column "E073_15_coreMarks_segments_state" for segmentation states.
- annotate_coreMarks_segments <- function(result, ChmmModels.gr, region_var) {
- # Create a GRanges object from the result data frame
- result.gr <- result %>%
- makeGRangesFromDataFrame(start.field = "start",
- end.field = "end",
- seqnames.field = "seqnames")
- # Find overlaps between the result regions and the ChromHMM models
- hits <- findOverlaps(result.gr, ChmmModels.gr) %>% as.data.frame()
- hits$state <- ChmmModels.gr$state[hits$subjectHits]
- hits$region <- result[[region_var]][hits$queryHits]
- # Match the state information to each region based on the 'region' string
- result$E073_15_coreMarks_segments_state <- hits$state[match(result[[region_var]], hits$region)]
- return(result)
- }
- # ------------------------------------------------------------------------------
- # Function: annotate_great
- # ------------------------------------------------------------------------------
- # Description:
- # Annotates regions with gene associations using GREAT analysis.
- #
- # Parameters:
- # result : Data frame. Regions to be annotated.
- # genome : Character. The genome version ("hg19" or "hg38") used in GREAT analysis.
- #
- # Returns:
- # A data frame with an additional column "GREAT_annotation" containing GREAT-derived gene information.
- # ------------------------------------------------------------------------------
- annotate_great <- function(result, genome) {
- # Create a GRanges object from the result data frame
- result.gr <- result %>%
- makeGRangesFromDataFrame(start.field = "start",
- end.field = "end",
- seqnames.field = "seqnames")
- # Submit the GREAT job and retrieve gene associations
- job <- submitGreatJob(result.gr, species = genome)
- regionsToGenes_gr <- rGREAT::getRegionGeneAssociations(job)
- regionsToGenes <- as.data.frame(regionsToGenes_gr)
- # Create annotation strings for each region
- GREAT_annotation <- lapply(seq_len(length(regionsToGenes$annotated_genes)), function(i) {
- g <- ifelse(regionsToGenes$dist_to_TSS[[i]] > 0,
- paste0(regionsToGenes$annotated_genes[[i]], " (+", regionsToGenes$dist_to_TSS[[i]], ")"),
- paste0(regionsToGenes$annotated_genes[[i]], " (", regionsToGenes$dist_to_TSS[[i]], ")"))
- paste0(g, collapse = ";")
- })
- # Select key columns from GREAT output and combine with the annotation strings
- great <- dplyr::select(regionsToGenes, seqnames, start, end, width)
- great <- data.frame(great, GREAT_annotation = unlist(GREAT_annotation))
- # Merge the GREAT annotation with the original result data frame
- result <- dplyr::left_join(result, great, by = c("seqnames", "start", "end"))
- return(result)
- }
- # ------------------------------------------------------------------------------
- # Function: annotate_region
- # ------------------------------------------------------------------------------
- # Description:
- # Annotates regions with CpG UCSC information by summarizing overlapping CpG details.
- #
- # Parameters:
- # result : Data frame. Regions to be annotated.
- # array : Character. The methylation array type (e.g., "HM450", "EPIC").
- # genome : Character. The genome version ("hg19" or "hg38").
- # dir.data.aux: Character. Directory path for auxiliary annotation data.
- # cpg : Vector. List of CpG identifiers to be considered.
- # region_var : Character. (Optional) Column name in 'result' that contains region identifiers. Default is "region".
- # cores : Numeric. Number of CPU cores to use for parallel processing. Default is 30.
- #
- # Returns:
- # A data frame with additional columns summarizing UCSC CpG annotation information.
- annotate_region <- function(result, array, genome, dir.data.aux, cpg, region_var = "region", cores = 30) {
- anno.gr <- get_anno_gr(array = array, genome = genome, dir.data.aux = dir.data.aux)
- anno_df <- as.data.frame(anno.gr)
- if(genome == "hg38" & array %in% c("EPIC","EPICv1")) island_var <- "Relation_to_UCSC_CpG_Island"
- else island_var <- "Relation_to_Island"
- # Create a GRanges object from the result data frame
- result.gr <- result %>%
- makeGRangesFromDataFrame(start.field = "start",
- end.field = "end",
- seqnames.field = "seqnames")
- doParallel::registerDoParallel(cores)
- anno <- plyr::ldply(
- 1:length(result.gr),
- .fun = function (i) {
- rg <- result.gr[i]
- overlap <- findOverlaps(rg, anno.gr)
- hit <- subjectHits(overlap)
- anno_df_sub <- anno_df[hit,]
- cpgs <- intersect(anno_df_sub$Name, cpg)
- anno_df_sub <- anno_df_sub[match(cpgs, anno_df_sub$Name),]
- # cpg in region
- cpgs_in_region <- paste(cpgs, collapse = ",")
- # UCSC_RefGene_Name
- UCSC_RefGene_Name <- paste0(unique(anno_df_sub[,"UCSC_RefGene_Name"]), collapse = ";")
- # UCSC_RefGene_Accession
- UCSC_RefGene_Accession <- paste0(unique(anno_df_sub[,"UCSC_RefGene_Accession"]), collapse = ";")
- # UCSC_RefGene_Group
- UCSC_RefGene_Group <- paste0(unique(anno_df_sub[,"UCSC_RefGene_Group"]), collapse = ";")
- # Relation_to_Island
- Relation_to_Island <- paste0(unique(anno_df_sub[,island_var]), collapse = ";")
- df <- data.frame(
- num_probes = length(cpgs),
- UCSC_RefGene_Name = UCSC_RefGene_Name,
- UCSC_RefGene_Accession = UCSC_RefGene_Accession,
- UCSC_RefGene_Group = UCSC_RefGene_Group,
- Relation_to_Island = Relation_to_Island,
- cpgs_in_region = cpgs_in_region
- )
- df[df == "NA"] <- NA
- df
- }, .parallel = T
- )
- cbind(result, anno)
- }
- # ------------------------------------------------------------------------------
- # Function: annotate_enhancer
- # ------------------------------------------------------------------------------
- # Description:
- # Annotates regions with enhancer overlap information using an external enhancer dataset.
- #
- # Parameters:
- # result : Data frame. Regions to be annotated.
- # nasser.enhancer.gr: GRanges object. Enhancer regions with associated cell type information.
- # cpg : Vector. List of CpG identifiers used in the analysis.
- # genome : Character. The genome version ("hg19" or "hg38").
- # array : Character. The methylation array type.
- #
- # Returns:
- # A data frame with two additional columns:
- # - nasser_is_enhancer: Logical indicating enhancer overlap.
- # - nasser_is_enhancer_cell_types: Character string of associated cell types.
- annotate_enhancer <- function(result, nasser.enhancer.gr, cpg, genome, array) {
- # Create a GRanges object from the result data frame
- result.gr <- result %>%
- makeGRangesFromDataFrame(start.field = "start",
- end.field = "end",
- seqnames.field = "seqnames")
- # Find overlaps between the result regions and the enhancer regions
- hits <- findOverlaps(result.gr, nasser.enhancer.gr) %>% as.data.frame()
- # Initialize the enhancer annotation columns
- result$nasser_is_enhancer <- FALSE
- result$nasser_is_enhancer[unique(hits$queryHits)] <- TRUE
- result$nasser_is_enhancer_cell_types <- NA
- result$nasser_is_enhancer_cell_types[unique(hits$queryHits)] <- sapply(unique(hits$queryHits), function(x) {
- paste(unique(nasser.enhancer.gr$CellType[hits$subjectHits[hits$queryHits %in% x]]), collapse = ",")
- })
- return(result)
- }
- annotate_chmm <- function(result, dir.data.aux = dir.data.aux) {
- load(file.path(dir.data.aux,"E073_15_coreMarks_segments.rda"))
- message("Annotating E073_15_coreMarks_segments")
- result$region <- paste0(result$seqnames,":",result$start,"-", result$end)
- result$start <- as.numeric(result$start)
- result$end <- as.numeric(result$end)
- result.gr <- result %>% makeGRangesFromDataFrame(
- start.field = "start",
- end.field = "end",
- seqnames.field = "seqnames"
- )
- hits <- findOverlaps(result.gr, ChmmModels.gr) %>% as.data.frame()
- hits$state <- ChmmModels.gr$state[hits$subjectHits]
- hits$region <- result$region[hits$queryHits]
- result$E073_15_coreMarks_segments_state <- hits$state[match(result$region,hits$region)]
- result$region <- NULL
- result
- }
- # ------------------------------------------------------------------------------
- # Function: add_dmr_annotation
- # ------------------------------------------------------------------------------
- # Description:
- # Integrates multiple annotation methods (region, enhancer, coreMarks segmentation, GREAT) for DMRs.
- #
- # Parameters:
- # result : Data frame. DMRs (differentially methylated regions) to be annotated.
- # cpg : Vector. List of CpG identifiers used for regional annotation.
- # dir.data.aux : Character. Directory path for auxiliary data files (e.g., segmentation, enhancer data).
- # region_var : (Optional) Character. Column name in 'result' containing region identifiers.
- # If NULL, a region identifier is created using chromosome, start, and end.
- # array : Character. The methylation array type (default "EPIC").
- # genome : Character. Genome version ("hg19" or "hg38").
- # annotate : Vector of characters. Specifies which annotation types to apply. Options include:
- # "region", "enhancer", "E073_15_coreMarks_segments", "GREAT".
- #
- # Returns:
- # A data frame with multiple annotation columns added.
- add_dmr_annotation <- function(result,
- cpg,
- dir.data.aux,
- region_var = NULL,
- array = "EPIC",
- genome = "hg19",
- annotate = c("region", "enhancer", "E073_15_coreMarks_segments", "GREAT")) {
- # Prepare the result data frame for genomic range creation
- if(is.null(region_var)) {
- result$seqnames <- paste0("chr", result$`#chrom`)
- result$region <- paste0(result$seqnames, ":", result$start, "-", result$end)
- region_var <- "region"
- } else {
- region <- str_split(result[[region_var]], ":|-", simplify = T)
- result$seqnames <- region[,1]
- result$start <- region[,2]
- result$end <- region[,3]
- }
- result$start <- as.numeric(result$start)
- result$end <- as.numeric(result$end)
- # Annotate with E073_15_coreMarks_segments if selected.
- if ("E073_15_coreMarks_segments" %in% annotate) {
- load(file.path(dir.data.aux, "E073_15_coreMarks_segments.rda"))
- message("Annotating E073_15_coreMarks_segments")
- result <- annotate_coreMarks_segments(result, ChmmModels.gr, region_var)
- }
- # Annotate with GREAT if selected.
- if ("GREAT" %in% annotate) {
- message("Annotating GREAT")
- result <- annotate_great(result, genome)
- }
- # Annotate with enhancer if selected.
- if ("enhancer" %in% annotate) {
- # Load enhancer data and filter
- data <- readr::read_tsv(file.path(dir.data.aux, "AllPredictions.AvgHiC.ABC0.015.minus150.ForABCPaperV3.txt.gz"),
- show_col_types = F)
- CellType.selected <- readxl::read_xlsx(file.path(dir.data.aux, "Nassser study selected biosamples.xlsx"),
- col_names = FALSE) %>% dplyr::pull(1)
- data.filtered <- data %>%
- dplyr::filter(CellType %in% CellType.selected) %>%
- dplyr::filter(!isSelfPromoter) %>%
- dplyr::filter(class != "promoter")
- nasser.enhancer.gr <- data.filtered %>%
- makeGRangesFromDataFrame(start.field = "start",
- end.field = "end",
- seqnames.field = "chr",
- keep.extra.columns = TRUE)
- message("Annotating enhancer")
- result <- annotate_enhancer(result, nasser.enhancer.gr, cpg, genome, array)
- }
- # Annotate with UCSC (island annotation) if selected.
- if ("region" %in% annotate) {
- message("Annotating region")
- result <- annotate_region(result, array, genome, dir.data.aux, cpg, region_var = region_var)
- }
- return(result)
- }
- # ===================================================================================================
- # Bacon correction
- # ===================================================================================================
- # Function: bacon_adj
- # ------------------------------------------------------------------------------
- # Description:
- # Adjusts for bias and genomic inflation in EWAS data using the bacon method.
- # Supports correction using either z-scores or effect sizes with standard errors.
- #
- # Parameters:
- # data : Data frame. The EWAS results containing statistics to be adjusted.
- # est_var : Character. Column name for effect size estimates.
- # z_var : Character. Column name for z-scores.
- # std_var : Character. Column name for standard errors.
- # use_z : Logical. If TRUE, bacon correction is performed using z-scores only; otherwise, effect sizes and SEs are used.
- # save : Logical. Whether to save the bacon-corrected data and inflation statistics to files.
- # dir.save: Character. Directory path where the output files will be saved.
- # prefix : Character. A prefix used for naming the output files.
- #
- # Returns:
- # A list containing:
- # - data.with.inflation: The corrected EWAS data frame.
- # - bacon.obj : The bacon object from the initial analysis.
- # - inflation.stat : A data frame summarizing inflation and bias statistics.
- bacon_adj <- function(data, est_var, z_var, std_var,
- use_z = F,
- save = F,
- dir.save = NULL,
- prefix = "Framingham"){
- ### 1. Compute genomic inflation factor before bacon adjustment
- data <- data %>% mutate(
- chisq = get(z_var)^2
- )
- # inflation factor - last term is median from chisq distrn with 1 df
- inflationFactor <- median(data$chisq,na.rm = TRUE) / qchisq(0.5, 1)
- print("lambda")
- print(inflationFactor)
- ### 2. bacon analysis
- if(use_z){
- ### 2. bacon analysis
- z_scores <- data[[z_var]]
- bc <- bacon(
- teststatistics = z_scores,
- na.exclude = TRUE,
- verbose = F
- )
- # inflation factor
- print("lambda.bacon")
- print(inflation(bc))
- # bias
- print("estimate bias")
- print(bias(bc))
- print("estimates")
- print(bacon::estimates(bc))
- ### 3. Create final dataset
- data.with.inflation <- data %>% mutate(
- zScore.bacon = tstat(bc)[,1],
- pValue.bacon.z = pval(bc)[,1],
- fdr.bacon.z = p.adjust(pval(bc), method = "fdr"),
- ) %>% mutate(z.value = z_scores)
- print("o After bacon correction")
- print("Conventional lambda")
- lambda.con <- median((data.with.inflation$zScore.bacon) ^ 2,na.rm = TRUE)/qchisq(0.5, 1)
- print(lambda.con)
- # percent_null <- trunc ( bacon::estimates(bc)[1]*100, digits = 0)
- # percent_1 <- trunc ( bacon::estimates(bc)[2]*100, digits = 0 )
- # percent_2 <- 100 - percent_null - percent_1
- bc2 <- bacon(
- teststatistics = data.with.inflation$zScore.bacon,
- na.exclude = TRUE,
- priors = list(
- sigma = list(alpha = 1.28, beta = 0.36),
- mu = list(lambda = c(0, 3, -3), tau = c(1000, 100, 100)),
- epsilon = list(gamma = c(90, 5, 5)))
- )
- } else {
- est <- data[[est_var]]
- se <- data[[std_var]]
- bc <- bacon(
- teststatistics = NULL,
- effectsizes = est,
- standarderrors = se,
- na.exclude = TRUE,
- verbose = F
- )
- # posteriors(bc)
- # inflation factor
- print("lambda.bacon")
- print(inflation(bc))
- # bias
- print("estimate bias")
- print(bias(bc))
- print("estimates")
- print(bacon::estimates(bc))
- ### 3. Create final dataset
- data.with.inflation <- data.frame(
- data,
- Estimate.bacon = bacon::es(bc),
- StdErr.bacon = bacon::se(bc),
- pValue.bacon = pval(bc),
- fdr.bacon = p.adjust(pval(bc), method = "fdr"),
- stringsAsFactors = FALSE
- )
- print("o After bacon correction")
- print("Conventional lambda")
- lambda.con <- median((data.with.inflation$Estimate.bacon/data.with.inflation$StdErr.bacon) ^ 2,na.rm = TRUE)/qchisq(0.5, 1)
- print(lambda.con)
- # percent_null <- trunc ( estimates(bc)[1]*100, digits = 0)
- # percent_1 <- trunc ( estimates(bc)[2]*100, digits = 0 )
- # percent_2 <- 100 - percent_null - percent_1
- bc2 <- bacon(
- teststatistics = NULL,
- effectsizes = data.with.inflation$Estimate.bacon,
- standarderrors = data.with.inflation$StdErr.bacon,
- na.exclude = TRUE,
- priors = list(
- sigma = list(alpha = 1.28, beta = 0.36),
- mu = list(lambda = c(0, 3, -3), tau = c(1000, 100, 100)),
- epsilon = list(gamma = c(99, .5, .5)))
- )
- }
- print("inflation")
- print(inflation(bc2))
- print("estimates")
- print(bacon::estimates(bc2))
- data.with.inflation$chisq <- NULL
- inflation.stat <- data.frame(
- "Inflation.org" = inflationFactor,
- "Inflation.bacon" = inflation(bc),
- "Bias.bacon" = bias(bc),
- "Inflation.after.correction" = lambda.con,
- "Inflation.bacon.after.correction" = inflation(bc2),
- "Bias.bacon.after.correction" = bias(bc2)
- )
- if(save){
- readr::write_csv(
- data.with.inflation,
- file.path(dir.save, paste0(prefix, "_bacon_correction.csv"))
- )
- writexl::write_xlsx(
- inflation.stat,
- file.path(dir.save, paste0(prefix, "_inflation_stats.xlsx"))
- )
- }
- return(
- list(
- "data.with.inflation" = data.with.inflation,
- "bacon.obj" = bc,
- "inflation.stat" = inflation.stat
- )
- )
- }
annotation_and_bacon.R at commit d2cc5c5, under MIT · at the source
Overview
- Division of Biostatistics, Department of Public Health Sciences, University of Miami, Miller School of Medicine, Miami, FL 33136 USA
- Dr. John T Macdonald Foundation Department of Human Genetics, University of Miami, Miller School of Medicine, Miami, FL 33136 USA
- John P. Hussman Institute for Human Genomics, University of Miami Miller School of Medicine, Miami, FL 33136 USA
- Sylvester Comprehensive Cancer Center, University of Miami, Miller School of Medicine, Miami, FL 33136 USA
- Soffer Clinical Research Ctr, University of Miami Miller School of Medicine, 1120 NW 14Th St, Miami, FL 33136 USA
Abstract
Aging is the strongest risk factor for Alzheimer’s disease (AD), yet the role of age-associated DNA methylation (DNAm) changes in blood and their relevance to AD remains poorly understood. We performed a meta-analysis of blood DNAm samples from 475 dementia-free subjects aged over 65 years across two independent cohorts, the Framingham Heart Study (FHS) at Exam 9 and the Alzheimer’s Disease Neuroimaging Initiative (ADNI). We adjusted for sex and immune cell-type proportions and corrected batch effects and genomic inflation. Integrative analyses included pathway enrichment, mQTL analysis, colocalization with Alzheimer’s disease and related dementia (ADRD) GWAS summary statistics, brain-blood DNAm correlations, and comparison to independent AD methylation studies. We identified 3758 CpGs and 556 differentially methylated regions (DMRs) consistently associated with chronological age in both cohorts at a 5% false discovery rate. Our pathway enrichment analyses highlighted metabolic regulation and synaptic signaling, processes previously implicated in Alzheimer’s disease. Colocalization with ADRD GWAS summary statistics identified 32 genomic regions consistent with shared genetic signals for DNAm and ADRD risk. Roughly one-third of aging-associated CpGs overlapped CpGs associated with AD or AD neuropathology in external studies. Finally, we prioritized nine promoter CpGs (including those located in PDE1B, ELOVL2, and PODXL2) showing strong positive blood-to-brain methylation concordance and external AD associations, nominating them as candidate blood-based biomarkers. Our study demonstrated that late-life aging signatures in blood DNAm converge on processes implicated in AD and intersect with dementia genetics. A small set of CpGs with blood-brain concordance and external AD support offers promising candidate blood-based biomarkers for future validation.
Supplementary Information: The online version contains supplementary material available at 10.1007/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.
TransBioInfoLab/AD-Aging-blood-sample-analysis
d2cc5c5d37b2512bdcdc11a402a38a9bd6c7fcdc, 5 June 2025Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
16 files
- code/
analysis/ , R, 461 linesassociation_test.Rmd - code/
analysis/ , R, 59 linesmanhattan_plot.R - code/
brain_blood_corr/ , R, 318 linesbrain_blood_corr.Rmd - code/
compare_study/ , R, 304 linescompare_study_miami_ad.R md - code/
dmr/ , R, 552 lines, 2 matchescoMethDMR.Rmd - code/
dmr/ , R, 67 linescombp_annotation.R - code/
dnam_vs_rna/ , R, 244 linesdnam_vs_rna.Rmd - code/
session_info.R , R, 74 lines - code/
utility/ , R, 629 lines, 2 matchesannotation_and_bacon.R - code/
utility/ , R, 199 lines, 1 matchcoMethDMR_aux.R - code/
utility/ , R, 692 linescpg_test.R - code/
utility/ , R, 137 linesmeta.R - code/
utility/ , R, 234 lines, 1 matchpathway.R - code/
utility/ , R, 301 linesplot.R - LICENSE, License, 21 lines
- README.md, Text, 77 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 14 scripts, each with its path and the digest of its content;
- 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
The ADNI and Framingham Heart Study datasets can be accessed from http://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 4 keywords, 11 MeSH terms, 2 funders, 75 references.
Cite
This paper
Zhang, W., Lukacsovich, D., Young, J. I., Gomez, L., Schmidt, M. A., Kunkle, B. W., Chen, X., Martin, E. R., & Wang, L. (2026). The aging epigenome: integrative analyses reveal intersection with Alzheimer's disease. GeroScience, 48(3), 3185-3203. https://
BibTeX
@article{zhang2026aging,
author = {Zhang, Wei and Lukacsovich, David and Young, Juan I and Gomez, Lissette and Schmidt, Michael A and Kunkle, Brian W and Chen, Xi and Martin, Eden R and Wang, Lily},
title = {{The aging epigenome: integrative analyses reveal intersection with Alzheimer's disease}},
journal = {GeroScience},
year = {2026},
month = apr,
volume = {48},
number = {3},
pages = {3185--3203},
publisher = {Springer},
issn = {2509-2715},
doi = {10.1007/
url = {https://
pmid = {41975027},
pmcid = {PMC13356189}
}
RIS
TY - JOUR
AU - Zhang, Wei
AU - Lukacsovich, David
AU - Young, Juan I
AU - Gomez, Lissette
AU - Schmidt, Michael A
AU - Kunkle, Brian W
AU - Chen, Xi
AU - Martin, Eden R
AU - Wang, Lily
TI - The aging epigenome: integrative analyses reveal intersection with Alzheimer's disease
T2 - GeroScience
J2 - Geroscience
PY - 2026
DA - 2026/
VL - 48
IS - 3
SP - 3185
EP - 3203
SN - 2509-2715
PB - Springer
DO - 10.1007/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1007/
"type": "article-journal",
"title": "The aging epigenome: integrative analyses reveal intersection with Alzheimer's disease",
"container-title": "GeroScience",
"author": [
{
"family": "Zhang",
"given": "Wei"
},
{
"family": "Lukacsovich",
"given": "David"
},
{
"family": "Young",
"given": "Juan I"
},
{
"family": "Gomez",
"given": "Lissette"
},
{
"family": "Schmidt",
"given": "Michael A"
},
{
"family": "Kunkle",
"given": "Brian W"
},
{
"family": "Chen",
"given": "Xi"
},
{
"family": "Martin",
"given": "Eden R"
},
{
"family": "Wang",
"given": "Lily"
}
],
"container-title-short":
"volume": "48",
"issue": "3",
"page": "3185-3203",
"DOI": "10.1007/
"PMID": "41975027",
"PMCID": "PMC13356189",
"ISSN": "2509-2715",
"publisher": "Springer",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
14
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s13073-026-01698-8 [code]
- From aging to Alzheimer's disease: concordant brain DNA methylation changes in late life.Journal: Genome medicineIn common: survival, ggpubr, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics, 25 references, author Lily Wang
- [2] doi:10.1002/trc2.70257 [code]
- Blood DNA methylation signature of cognitive reserve moderates the association between CSF tau pathology and memory in prodromal Alzheimer's disease.Journal: Alzheimer's & dementia (New York, N. Y.)In common: lmerTest, lme4, ggpubr, 2 other tools, Alzheimer's / dementia, genetics / omics, 2 references, author Lily Wang
- [3] doi:10.1038/s41586-026-10877-x [code]
- Human brain organoids record the passage of time over multiple years.Journal: NatureIn common: lmerTest, lme4, cowplot, 4 other tools, genetics / omics, 2 references
- [4] doi:10.1038/s41467-026-73865-9 [code]
- Histamine shapes the neurocomputational dynamics of human learning.Journal: Nature communicationsIn common: metafor, lmerTest, lme4, 5 other tools
- [5] doi:10.1038/s44400-026-00100-z [code]
- Peripheral blood microarray-based transcriptomic and epigenetic analyses identify immune, inflammation, and metabolic dysregulation in Alzheimer's disease.Journal: NPJ dementiaIn common: survival, lme4, ggpubr, 2 other tools, Alzheimer's / dementia, genetics / omics, 2 references
- [6] doi:10.1186/s13195-026-02036-1 [code]
- Genetic drivers of progression in Alzheimer's disease are distinct from disease risk.Journal: Alzheimer's research & therapyIn common: metafor, lmerTest, lme4, 4 other tools, Alzheimer's / dementia, genetics / omics
- [7] doi:10.1002/ece3.73881 [code]
- Epigenetic Aging in Brain Tissue of the Self-Fertilizing Vertebrate, &
lt;i& gt;Kryptolebias marmoratus& lt;/ i& gt;. Journal: Ecology and evolutionIn common: ggpubr, ggplot2, tidyverse, genetics / omics, 5 references - [8] doi:10.1038/s41588-026-02722-8 [code]
- A multiancestry polygenic risk score for Alzheimer's disease is associated with cognitive decline and neuropathological hallmarks in diverse populations.Journal: Nature geneticsIn common: survival, lmerTest, lme4, 3 other tools, Alzheimer's / dementia, genetics / omics, 1 reference
- [9] doi:10.1038/s41467-026-77170-3 [code]
- DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.Journal: Nature communicationsIn common: survival, lmerTest, lme4, 4 other tools, genetics / omics
- [10] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: survival, lme4, cowplot, 4 other tools, genetics / omics
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 14 scripts, and 6 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:ddf8159a2b4f6442…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
