Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program.
The 8 matches
- [1] § Methods › Phenotype Ascertainment ↔ Code_Files/AoU_or_calculation.Rmd, lines 288–357 · score 0.97 · acute myocardial infarction, chronic ischemic heart, chronic rheumatic heart, heart failure, pulmonary embolism, angina pectoris
- [2] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation.Rmd, lines 18–84 · score 0.93 · annual income, alcohol consumption, physical activity, smoking status, diabetes status, vigorous
- [3] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 16–94 · score 0.93 · annual income, alcohol consumption, physical activity, smoking status, diabetes status, vigorous
- [4] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation.Rmd, lines 18–84 · score 0.78 · physical activity, ethnic background, alcohol, Depression, income, Age
- [5] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 16–94 · score 0.78 · physical activity, ethnic background, alcohol, Depression, income, Age
- [6] § Results › Cardiovascular Disease Subtypes and AD Prevalence ↔ Code_Files/AoU_or_calculation.Rmd, lines 288–357 · score 0.75 · chronic ischemic heart, heart failure, pulmonary embolism, cerebral infarction, hypertension, hypotension
- [7] § Methods › Genetic Analysis ↔ Code_Files/genetic_analysis.Rmd, lines 158–186 · score 0.72 · heart rate, heart function, AD GWAS, amyloid, tau, dementia
- [8] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 607–657 · score 0.51 · logistic regression, Chinese, Unknown, Asian, strata, ethnic
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 761 lines · 22 KB · no license · 3 matches
- ---
- title: "OR Calculations"
- author: "Aili Toyli"
- date: "`r Sys.Date()`"
- output: html_document
- ---
- ```{r setup, include=FALSE}
- knitr::opts_chunk$set(echo = TRUE)
- library(tidyverse)
- #set working path
- PATH <- "~/"
- ```
- ##Read in OR datafile and prepare variables for analysis
- ```{r}
- #read in file
- cvd <- read.csv(PATH + "or_clean.csv")
- #clean dataframe
- cvd <- cvd %>%
- #remove unneeded columns
- select(-c(eid, p41270)) %>%
- #label covariates for better readability
- mutate(Age = p21003_i0,
- Sex = factor(p31, labels = c('Female', 'Male')),
- `Age completed full time education` = p845_i0,
- BMI = p21001_i0) %>%
- #Create readable smoking status levels
- mutate(`Smoking Status` = factor(p20116_i0, labels = c('Prefer not to answer', 'Never', 'Previous', 'Current'))) %>%
- #Label all types of depression/bipolar as yes, no if otherwise
- mutate(`Depression/Bipolar Disorder` = factor(p20126_i0, labels = c('No', rep('Yes', 5)))) %>%
- #Replace non response encodings from physical activity data with NA
- mutate(`Days of moderate physical activity each week` = replace(p884_i0, p884_i0 %in% c(-1, -3), NA)) %>%
- mutate(`Days of vigorous physical activity each week` = replace(p904_i0, p904_i0 %in% c(-1, -3), NA)) %>%
- #Label alcohol consumption levels
- mutate(`Alcohol Consumption` = factor(p1558_i0, labels = c('Daily or almost daily', '3-4 times a week', '1-2 times a week', '1-3 times a month', 'Special occasions only', 'Never', 'Prefer not to answer'))) %>%
- #Label annual income levels
- mutate(`Annual Household Income` = factor(p738_i0, labels = c('<£18,000', '£18,000-£30,999', '£31,000-£51,999', '£52,000-£100,000', '>£100,000', 'Do not know', 'Prefer not to answer'))) %>%
- #Save diabetes status as factor
- mutate(Diabetes = factor(Diabetes, labels = c('No', 'Yes'))) %>%
- #Replace spelling error on afib
- rename(Atrial.Fibrillation = Atrial.Fibriliation) %>%
- #replace blockage and afib with new arrhythmia category
- mutate(Cardiac.Arrhythmia = ifelse(Atrial.Fibrillation | Blockage, TRUE, FALSE)) %>%
- select(-c(Atrial.Fibrillation, Blockage)) %>%
- #Reorder columns to have AD followed by CVD, then covariates
- select(Alzheimers.Disease, Hypertension, Hypotension, Angina.Pectoris,
- Acute.Myocardial.Infarction, Pulmonary.Embolism, Cardiac.Arrhythmia,
- Heart.Failure, Chronic.Rheumatic.Heart.Disease,
- Chronic.Ischemic.Heart.Disease, Cerebral.Infarction, everything())
- #Adjust ethnicity encodings to be more generalized
- cvd$`Ethnic background` <- as.character(cvd$p21000_i0)
- cvd$`Ethnic background` <- ifelse(startsWith(cvd$`Ethnic background`, '1'), 'White',
- ifelse(startsWith(cvd$`Ethnic background`, '2'), 'Mixed',
- ifelse(startsWith(cvd$`Ethnic background`, '3'), 'Asian',
- ifelse(startsWith(cvd$`Ethnic background`, '4'), 'Black',
- ifelse(cvd$`Ethnic background`=='5', 'Chinese',
- ifelse(cvd$`Ethnic background`=='6', 'Other',
- ifelse(cvd$`Ethnic background`=='-3', 'Prefer not to answer',
- 'Unknown')))))))
- #drop original columns containing covariate data
- cvd<- cvd %>% select(-matches("^p[0-9]+"))
- #categorize ethnic background as a factor variable
- cvd$`Ethnic background` <- factor(cvd$`Ethnic background`)
- library(dplyr)
- library(caret)
- # 1. Impute missing categorical variables as "Prefer not to answer"
- cvd_clean <- cvd %>%
- mutate(across(where(is.factor), ~ {
- fct <- .
- # Add the level if not present
- if (!"Prefer not to answer" %in% levels(fct)) {
- levels(fct) <- c(levels(fct), "Prefer not to answer")
- }
- # Replace NA
- fct[is.na(fct)] <- "Prefer not to answer"
- fct
- }))
- #create baseline dataframe with mean imputation
- cvd_clean_mean <- cvd_clean %>%
- mutate(across(where(is.numeric), ~ {
- ifelse(is.na(.), mean(., na.rm = TRUE), .)
- }))
- ```
- ## Code for baseline OR calculation
- ```{r}
- #cvd data frame should have alzheimers in first column, then cvd subtypes
- #make sure eid is removed
- #some of the index numbers would need to be adjusted if looking at a different number of subtypes
- #create empty dataframe to store results
- output <- data.frame(cvd = character(10), est = numeric(10),
- lower.lim = numeric(10), upper.lim = numeric(10))
- #run regression for each subtype
- for(i in 1:10){
- #put name of disease in output dataframe
- output$cvd[i] <- names(cvd_clean_mean)[i+1]
- #save formula
- formula <- as.formula(paste0('`', names(cvd_clean_mean)[1], '` ~ `', names(cvd_clean_mean)[i+1], '` + ',
- paste(paste0('`', names(cvd_clean_mean)[12:ncol(cvd_clean_mean)], '`'), collapse = ' + '))
- )
- #run regression
- mylogit <- glm(formula, family = 'binomial', data = cvd_clean_mean)
- #save logistic regression summary
- sum_mylogit <- summary(mylogit)
- #save odds ratio and confidence interval
- or_confint <- exp(cbind(OR = sum_mylogit$coefficients[2, 1],
- LL = sum_mylogit$coefficients[2, 1]
- - 1.96 * sum_mylogit$coefficients[2, 2],
- UL = sum_mylogit$coefficients[2, 1]
- + 1.96 * sum_mylogit$coefficients[2, 2]))
- #save in output table
- output[i, 2] <- or_confint[1, 1]
- output[i, 3] <- or_confint[1, 2]
- output[i, 4] <- or_confint[1, 3]
- }
- #visualize with a forest plot
- plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
- geom_point(shape = 21, fill = "blue", size = 5) +
- geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
- geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
- annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
- vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
- coord_flip() +
- labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
- title = "Forest Plot of Odds Ratios\nfor AD with CVD Subtypes",
- subtitle = 'UK Biobank Cohort\nMean Imputation')+
- theme(axis.text.x = element_text(size = 12),
- axis.text.y = element_text(size = 16),
- plot.title = element_text(hjust = 0.5, size = 30),
- plot.subtitle = element_text(hjust = 0.5),
- axis.title = element_text(size = 20),
- legend.text= element_text(size = 12),
- legend.title = element_text(size = 20),
- panel.background = element_rect(fill = "white"))
- plot
- #save plot
- ggsave(PATH + 'or_plot_mean_UKB.png', plot = plot,
- width = 9, height = 9)
- #save datafile
- write.csv(output, PATH + 'or_results_mean.csv')
- ```
- ##Impute missing numeric data with mice, run analysis on multiple datasets
- ```{r}
- #load library
- library(mice)
- #inspect data
- md.pattern(cvd_clean)
- #perform imputation
- set.seed(123)
- library(parallel)
- cl <- makeCluster(4)
- imp <- mice(cvd_clean, m = 5, maxit = 5, cluster = cl)
- stopCluster(cl)
- ##get logistic regression odds ratio estimates
- library(mice)
- library(dplyr)
- ##--------------------------------------------------
- ## 1. Identify variables (by name, not index)
- ##--------------------------------------------------
- all_vars <- names(imp$data)
- outcome <- all_vars[1]
- exposures <- all_vars[2:11]
- covars <- all_vars[12:length(all_vars)]
- ##--------------------------------------------------
- ## 2. Storage for pooled results
- ##--------------------------------------------------
- output <- data.frame(
- cvd = exposures,
- OR = NA_real_,
- lower.lim = NA_real_,
- upper.lim = NA_real_
- )
- ##--------------------------------------------------
- ## 3. Parallelize odds ratio calculation
- ##--------------------------------------------------
- library(parallel)
- cl <- makeCluster(detectCores() - 1)
- clusterExport(
- cl,
- varlist = c("imp", "exposures", "outcome", "covars"),
- envir = environment()
- )
- clusterEvalQ(cl, {
- library(mice)
- })
- results <- parLapply(
- cl,
- seq_along(exposures),
- function(i) {
- cvd_var <- exposures[i]
- fit_imp <- with(
- imp,
- glm(
- as.formula(
- paste0(
- "`", outcome, "` ~ `", cvd_var, "` + ",
- paste0("`", covars, "`", collapse = " + ")
- )
- ),
- family = binomial
- )
- )
- pooled_sum <- summary(pool(fit_imp))
- coef_row <- which(pooled_sum$term == paste0(cvd_var, "TRUE"))
- if (length(coef_row) == 0) {
- return(c(NA, NA, NA))
- }
- beta <- pooled_sum$estimate[coef_row]
- se <- pooled_sum$std.error[coef_row]
- c(
- exp(beta),
- exp(beta - 1.96 * se),
- exp(beta + 1.96 * se)
- )
- }
- )
- stopCluster(cl)
- output <- data.frame(
- exposure = exposures,
- OR = sapply(results, `[`, 1),
- lower.lim = sapply(results, `[`, 2),
- upper.lim = sapply(results, `[`, 3)
- )
- output
- ## Visualize results
- #visualize with a forest plot
- plot <- ggplot(output, aes(x = reorder(exposure, OR), y = OR)) +
- geom_point(shape = 21, fill = "blue", size = 5) +
- geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
- geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
- annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
- vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
- coord_flip() +
- labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
- title = "Forest Plot of Odds Ratios\nfor AD with CVD Subtypes",
- subtitle = 'UK Biobank Cohort\nMICE Imputation')+
- theme(axis.text.x = element_text(size = 12),
- axis.text.y = element_text(size = 16),
- plot.title = element_text(hjust = 0.5, size = 30),
- plot.subtitle = element_text(hjust = 0.5),
- axis.title = element_text(size = 20),
- legend.text= element_text(size = 12),
- legend.title = element_text(size = 20),
- panel.background = element_rect(fill = "white"))
- plot
- #save plot
- ggsave(PATH + 'or_plot_mice_UKB.png', plot = plot,
- width = 9, height = 9)
- #save datafile
- write.csv(output, PATH + 'or_results_mice.csv')
- ```
- ## Calculating crosstabs
- ```{r}
- library(dplyr)
- library(purrr)
- # Variables of interest
- vars <- names(cvd)[2:11] # your CVD indicators (0/1 or FALSE/TRUE)
- # Function to compute one row per CVD subtype
- build_row <- function(var) {
- # Define vectors
- cvd_var <- cvd[[var]]
- ad_var <- cvd$Alzheimers.Disease
- # Prevalence of CVD subtype
- n_yes <- sum(cvd_var == TRUE, na.rm = TRUE)
- n_no <- sum(cvd_var == FALSE, na.rm = TRUE)
- n_total <- n_yes + n_no
- prev_str <- paste0(
- n_yes, " (", round(n_yes / n_total * 100, 1), "%)"
- )
- # AD among CVD=yes
- ad_yes_cvd_yes <- sum(ad_var == TRUE & cvd_var == TRUE, na.rm = TRUE)
- ad_rate_yes <- paste0(
- ad_yes_cvd_yes,
- " (", round(ad_yes_cvd_yes / n_yes * 100, 1), "%)"
- )
- # AD among CVD=no
- ad_yes_cvd_no <- sum(ad_var == TRUE & cvd_var == FALSE, na.rm = TRUE)
- ad_rate_no <- paste0(
- ad_yes_cvd_no,
- " (", round(ad_yes_cvd_no / n_no * 100, 1), "%)"
- )
- # Contingency table + p-value
- tbl <- table(cvd_var, ad_var)
- # Use Fisher's exact when needed
- pval <- if (any(tbl < 5)) {
- fisher.test(tbl)$p.value
- } else {
- chisq.test(tbl)$p.value
- }
- tibble(
- CVD_subtype = var,
- Prevalence = prev_str,
- AD_among_CVD_yes = ad_rate_yes,
- AD_among_CVD_no = ad_rate_no,
- p_value = signif(pval, 3)
- )
- }
- # Build the final table
- final_table <- map_dfr(vars, build_row)
- final_table
- # Write output
- write.csv(final_table,
- PATH + "ad_cvd_crosstabs.csv",
- row.names = FALSE)
- ```
- ## Firth regression to account for case-control imbalances
- ```{r}
- ## ============================================================
- ## Parallelized Firth Logistic Regression Across CVD Subtypes
- ## ============================================================
- ## -----------------------------
- ## 0. Libraries
- ## -----------------------------
- library(logistf)
- library(parallel)
- library(ggplot2)
- ## -----------------------------
- ## 1. Setup
- ## -----------------------------
- # outcome in first column, CVD subtypes next
- outcome_var <- names(cvd_clean_mean)[1]
- exposures <- names(cvd_clean_mean)[2:11] # 10 CVD subtypes
- covars <- names(cvd_clean_mean)[12:ncol(cvd_clean_mean)]
- # initialize output
- output <- data.frame(
- cvd = exposures,
- est = NA_real_,
- lower.lim = NA_real_,
- upper.lim = NA_real_,
- stringsAsFactors = FALSE
- )
- ## -----------------------------
- ## 2. Parallel backend
- ## -----------------------------
- ncores <- max(1, detectCores() - 2)
- cl <- makeCluster(ncores)
- clusterExport(
- cl,
- varlist = c("cvd_clean_mean", "outcome_var", "exposures", "covars"),
- envir = environment()
- )
- clusterEvalQ(cl, {
- library(logistf)
- })
- ## Non parallel firth regression
- results <- lapply(
- seq_along(exposures),
- function(i) {
- exposure <- exposures[i]
- message("Starting Firth regression for: ", exposure)
- # build formula
- form <- as.formula(
- paste0(
- "`", outcome_var, "` ~ `", exposure, "` + ",
- paste(paste0("`", covars, "`"), collapse = " + ")
- )
- )
- # fit Firth model
- fit <- try(
- logistf(
- formula = form,
- data = cvd_clean_mean,
- plconf = 2
- ),
- silent = TRUE
- )
- if (inherits(fit, "try-error")) {
- message("FAILED: ", exposure)
- return(c(NA, NA, NA))
- }
- coef_idx <- which(names(fit$coefficients) == paste0(exposure, "TRUE"))
- if (length(coef_idx) == 0) {
- message("Coefficient not found: ", exposure)
- return(c(NA, NA, NA))
- }
- beta <- fit$coefficients[coef_idx]
- ci_low <- fit$ci.lower[coef_idx]
- ci_up <- fit$ci.upper[coef_idx]
- message("Finished Firth regression for: ", exposure)
- c(
- exp(beta),
- exp(ci_low),
- exp(ci_up)
- )
- }
- )
- ## -----------------------------
- ## 4. Collect results
- ## -----------------------------
- output$est <- sapply(results, `[`, 1)
- output$lower.lim <- sapply(results, `[`, 2)
- output$upper.lim <- sapply(results, `[`, 3)
- ## -----------------------------
- ## 5. Forest plot
- ## -----------------------------
- plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
- geom_point(shape = 21, fill = "blue", size = 5) +
- geom_errorbar(
- aes(ymin = lower.lim, ymax = upper.lim),
- width = 0.15,
- linewidth = 0.75
- ) +
- geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
- coord_flip() +
- labs(
- x = "CVD Subtype",
- y = "Odds Ratio (95% CI)",
- title = "Firth Odds Ratios\nfor AD with CVD Subtypes",
- subtitle = "UK Biobank Cohort\nMean Imputation"
- ) +
- theme(
- axis.text.x = element_text(size = 12),
- axis.text.y = element_text(size = 16),
- plot.title = element_text(hjust = 0.5, size = 30),
- plot.subtitle = element_text(hjust = 0.5),
- axis.title = element_text(size = 20),
- panel.background = element_rect(fill = "white")
- )
- print(plot)
- ## -----------------------------
- ## 6. Save outputs
- ## -----------------------------
- ggsave(
- PATH + "or_plot_firth_UKB.png",
- plot = plot,
- width = 9,
- height = 9
- )
- write.csv(
- output,
- PATH + "or_results_firth_UKB.csv",
- row.names = FALSE
- )
- ```
- ## OR calculation with comorbidity adjustment
- ```{r}
- #cvd data frame should have alzheimers in first column, then cvd subtypes
- #make sure eid is removed
- #some of the index numbers would need to be adjusted if looking at a different number of subtypes
- #create empty dataframe to store results
- output <- data.frame(cvd = character(10), est = numeric(10),
- lower.lim = numeric(10), upper.lim = numeric(10))
- #run regression for each subtype
- for(i in 1:10){
- #put name of disease in output dataframe
- output$cvd[i] <- names(cvd_clean_mean)[i+1]
- #save formula
- formula <- as.formula(paste0('`', names(cvd_clean_mean)[1], '` ~ `', names(cvd_clean_mean)[i+1], '` + .')
- )
- #run regression
- mylogit <- glm(formula, family = 'binomial', data = cvd_clean_mean)
- #save logistic regression summary
- sum_mylogit <- summary(mylogit)
- #save odds ratio and confidence interval
- or_confint <- exp(cbind(OR = sum_mylogit$coefficients[2, 1],
- LL = sum_mylogit$coefficients[2, 1]
- - 1.96 * sum_mylogit$coefficients[2, 2],
- UL = sum_mylogit$coefficients[2, 1]
- + 1.96 * sum_mylogit$coefficients[2, 2]))
- #save in output table
- output[i, 2] <- or_confint[1, 1]
- output[i, 3] <- or_confint[1, 2]
- output[i, 4] <- or_confint[1, 3]
- }
- #visualize with a forest plot
- plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
- geom_point(shape = 21, fill = "blue", size = 5) +
- geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
- geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
- annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
- vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
- coord_flip() +
- labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
- title = "Comorbidity-Adjusted\nOdds Ratios",
- subtitle = 'UK Biobank Cohort')+
- theme(axis.text.x = element_text(size = 12),
- axis.text.y = element_text(size = 16),
- plot.title = element_text(hjust = 0.5, size = 30),
- plot.subtitle = element_text(hjust = 0.5),
- axis.title = element_text(size = 20),
- legend.text= element_text(size = 12),
- legend.title = element_text(size = 20),
- panel.background = element_rect(fill = "white"))
- plot
- #save plot
- ggsave(PATH + 'or_plot_comorbid_UKB.png', plot = plot,
- width = 9, height = 9)
- #save datafile
- write.csv(output, PATH + 'or_results_comorbid.csv')
- ```
- ## OR code with ethnically stratified sampling
- ```{r}
- ############################################################
- ## Stratified Logistic Regression for AD ~ CVD Subtypes
- ## Stratified by Ethnicity4 (UK Biobank)
- ############################################################
- library(dplyr)
- library(forcats)
- library(ggplot2)
- set.seed(123)
- #--------------------------------------------------
- # 1. CREATE ETHNICITY STRATA
- #--------------------------------------------------
- cvd_clean <- cvd_clean %>%
- mutate(
- Ethnicity4 = fct_collapse(
- `Ethnic background`,
- White = "White",
- Black = "Black",
- Asian = c("Chinese", "Asian"),
- Other.Unknown = c("Mixed", "Other", "Unknown", "Prefer not to answer")
- )
- ) %>%
- mutate(Ethnicity4 = droplevels(as.factor(Ethnicity4))) %>%
- select(-`Ethnic background`)
- #--------------------------------------------------
- # 2. VARIABLE SETUP
- #--------------------------------------------------
- outcome_var <- names(cvd_clean)[1] # AD outcome
- exposure_vars <- names(cvd_clean)[2:11] # CVD subtypes
- covariates <- names(cvd_clean)[12:ncol(cvd_clean)] # comorbidities
- #--------------------------------------------------
- # 3. OUTPUT DATAFRAME (LONG FORMAT)
- #--------------------------------------------------
- output <- expand.grid(
- cvd = exposure_vars,
- Ethnicity4 = levels(cvd_clean$Ethnicity4),
- stringsAsFactors = FALSE
- )
- output$est <- NA_real_
- output$lower.lim <- NA_real_
- output$upper.lim <- NA_real_
- #--------------------------------------------------
- # 4. STRATIFIED LOGISTIC REGRESSION
- #--------------------------------------------------
- row_id <- 1
- for (eth in levels(cvd_clean$Ethnicity4)) {
- data_eth <- cvd_clean %>%
- filter(Ethnicity4 == eth)
- # Skip strata with very small sample size
- if (nrow(data_eth) < 50) next
- for (exposure in exposure_vars) {
- # Drop covariates with no variation in this stratum
- cov_eth <- covariates[
- sapply(data_eth[covariates], function(x) length(unique(x)) > 1)
- ]
- # Build formula with backticks
- formula <- as.formula(
- paste0(
- "`", outcome_var, "` ~ `", exposure, "` + ",
- paste(paste0("`", cov_eth, "`"), collapse = " + ")
- )
- )
- # Fit logistic regression
- fit <- glm(
- formula,
- family = binomial,
- data = data_eth
- )
- coef_summary <- summary(fit)$coefficients
- # Exposure coefficient (always row 2)
- beta <- coef_summary[2, 1]
- se <- coef_summary[2, 2]
- # OR and 95% CI
- or_ci <- exp(c(
- beta,
- beta - 1.96 * se,
- beta + 1.96 * se
- ))
- output[row_id, c("est", "lower.lim", "upper.lim")] <- or_ci
- row_id <- row_id + 1
- }
- }
- # Remove rows that were skipped
- output <- output %>% filter(!is.na(est))
- # Remove rows where the upper 95% CI exceeds 10
- output_filtered <- output %>%
- filter(upper.lim <= 10)
- #--------------------------------------------------
- # 5. FOREST PLOT (FACETED BY ETHNICITY)
- #--------------------------------------------------
- plot <- ggplot(output_filtered, aes(x = reorder(cvd, est), y = est)) +
- geom_point(size = 3, color = "blue") +
- geom_errorbar(
- aes(ymin = lower.lim, ymax = upper.lim),
- width = 0.15
- ) +
- geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
- coord_flip() +
- facet_wrap(~ Ethnicity4) +
- labs(
- x = "CVD Subtype",
- y = "Odds Ratio (95% CI)",
- title = "Stratified Odds Ratios for AD by Ethnicity",
- subtitle = "Adjusted Logistic Regression (UK Biobank)"
- ) +
- theme_bw(base_size = 14)
- print(plot)
- #--------------------------------------------------
- # 6. SAVE OUTPUTS
- #--------------------------------------------------
- ggsave(
- PATH + "or_plot_stratified_ethnicity_UKB.png",
- plot = plot,
- width = 12,
- height = 8
- )
- write.csv(
- output,
- PATH + "or_results_stratified_ethnicity_UKB.csv",
- row.names = FALSE
- )
- ```
UKB_or_calculation_revisions.Rmd at commit 6ac93bd, no license · at the source
Overview
- Department of Mathematical Sciences, Michigan Technological University, Houghton, MI, USA
- Department of Computer Science, Kennesaw State University, Marietta, GA, USA
- Division of Biomedical Informatics and Genomics, Tulane Center of Biomedical Informatics and Genomics, Deming Department of Medicine, Tulane University, New Orleans, LA, USA
- Department of Kinesiology and Integrative Physiology, Michigan Technological University, Houghton, MI, USA
- Department of Applied Computing, Michigan Technological University, Houghton, MI, USA
- Center for Biocomputing and Digital Health, Institute of Computing and Cybersystems, and Health Research Institute, Michigan Technological University, Houghton, MI, USA
Abstract
Background: Cardiovascular disease (CVD) and Alzheimer's disease (AD) are major public health concerns that share overlapping risk factors and potential mechanistic pathways. Although vascular contributions to cognitive decline are well documented, the specific relationships between AD and different CVD subtypes remain poorly understood.
Methods: In this cross‐sectional study, we examined associations between AD and 11 CVD subtypes using logistic regression models in 2 large biobanks: the UK Biobank (n=
Results: Most CVD subtypes were significantly associated with AD in both cohorts. Hypotension had the strongest and most consistent association, although it has been comparatively understudied in AD research. Strong associations were also consistently observed between AD and hypertension and cerebral infarction. Notably, acute myocardial infarction was not significantly linked to AD. Genetic analyses revealed shared loci between AD‐ and CVD‐related traits, particularly in regions near APOE, MAPT, and genes influencing myocardial structure and vascular function.
Conclusions: This study identifies subtype‐specific CVD associations with AD across 2 diverse cohorts and highlights shared genetic architecture underlying heart–brain interactions. These findings underscore the importance of vascular health in AD risk and suggest that certain CVD subtypes, especially hypotension, may play underrecognized roles in cognitive decline.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.
MIILab-MTU/CVDSubtypesADCorrelation
6ac93bda9b81f5902390838815361ea52885ad6f, 12 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
10 files
- Code_Files/
AoU Analysis/ , Jupyter, 124 linesAoU_comorbid.ipynb - Code_Files/
AoU Analysis/ , Jupyter, 213 linesAoU_firth.ipynb - Code_Files/
AoU Analysis/ , Jupyter, 188 linesAoU_mice.ipynb - Code_Files/
AoU Analysis/ , R, 177 linesaou_stratified.R - Code_Files/
AoU Analysis/ , R, 70 linescrosstabs.R - Code_Files/
AoU_or_calculation.Rmd , R, 1,092 lines, 2 matches - Code_Files/
UKB_or_calculation.Rmd , R, 292 lines, 2 matches - Code_Files/
UKB_or_calculation_revis , R, 761 lines, 3 matchesions.Rmd - Code_Files/
genetic_analysis.Rmd , R, 259 lines, 1 match - README.md, Text, 3 lines
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 9 scripts, each with its path and the digest of its content;
- 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 8 authors, 7 keywords, 17 MeSH terms, 2 funders, 50 references.
Cite
This paper
Toyli, A., Zhao, C., Su, K., Shen, H., Deng, H., Chen, Q., Sha, Q., & Zhou, W. (2026). Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program. Journal of the American Heart Association, 15(12), e046172. https://
BibTeX
@article{toyli2026cardio
author = {Toyli, Aili and Zhao, Chen and Su, Kuan‐Jui and Shen, Hui and Deng, Hong‐Wen and Chen, Qing‐Hui and Sha, Qiuying and Zhou, Weihua},
title = {{Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program}},
journal = {Journal of the American Heart Association},
year = {2026},
month = jun,
volume = {15},
number = {12},
pages = {e046172},
publisher = {Wiley},
issn = {2047-9980},
doi = {10.1161/
url = {https://
pmid = {42267709},
pmcid = {PMC13323167}
}
RIS
TY - JOUR
AU - Toyli, Aili
AU - Zhao, Chen
AU - Su, Kuan‐Jui
AU - Shen, Hui
AU - Deng, Hong‐Wen
AU - Chen, Qing‐Hui
AU - Sha, Qiuying
AU - Zhou, Weihua
TI - Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program
T2 - Journal of the American Heart Association
J2 - J Am Heart Assoc
PY - 2026
DA - 2026/
VL - 15
IS - 12
SP - e046172
SN - 2047-9980
PB - Wiley
DO - 10.1161/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1161/
"type": "article-journal",
"title": "Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program",
"container-title": "Journal of the American Heart Association",
"author": [
{
"family": "Toyli",
"given": "Aili"
},
{
"family": "Zhao",
"given": "Chen"
},
{
"family": "Su",
"given": "Kuan‐Jui"
},
{
"family": "Shen",
"given": "Hui"
},
{
"family": "Deng",
"given": "Hong‐Wen"
},
{
"family": "Chen",
"given": "Qing‐Hui"
},
{
"family": "Sha",
"given": "Qiuying"
},
{
"family": "Zhou",
"given": "Weihua"
}
],
"container-title-short":
"volume": "15",
"issue": "12",
"page": "e046172",
"DOI": "10.1161/
"PMID": "42267709",
"PMCID": "PMC13323167",
"ISSN": "2047-9980",
"publisher": "Wiley",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s40168-026-02342-8 [code]
- Impacts of host genetics on gut microbiome composition in Alzheimer's disease.Journal: MicrobiomeIn common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics, 1 reference
- [2] doi:10.1038/s41467-026-70374-7 [code]
- scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies.Journal: Nature communicationsIn common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics
- [3] doi:10.1038/s41597-026-06971-4 [code]
- Human neuronal differentiation under Aβ exposure: a single-cell transcriptomic and epigenomic dataset.Journal: Scientific dataIn common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics
- [4] doi:10.1093/neuonc/noag128 [code]
- Spatially-resolved single-cell imaging of melanoma brain metastases identifies localized immune patterns predictive of immune checkpoint blockade response.Journal: Neuro-oncologyIn common: caret, data.table, ggplot2, 1 other tool, clinical / translational, genetics / omics
- [5] doi:10.1111/cts.70577 [code]
- Machine Learning-Based Prediction of Drug-Induced QTc Changes in a Large Finnish Biobank Cohort.Journal: Clinical and translational scienceIn common: caret, data.table, ggplot2, 1 other tool, clinical / translational, genetics / omics
- [6] doi:10.1111/acel.70499 [code]
- Brain Aging Mediating Heart Imaging-Derived Phenotypes and Mental and Nervous System Disorders.Journal: Aging cellIn common: caret, data.table, tidyverse, 1 reference
- [7] doi:10.1093/bib/bbag175 [code]
- Uncovering causal relationships in single-cell omic studies with causarray.Journal: Briefings in bioinformaticsIn common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia
- [8] doi:10.1038/s41398-026-04131-1 [code]
- Multimodal phenotypic classification of generalized anxiety and panic using structural MRI data and psychosocial factors: machine learning results from the German National Cohort (NAKO) study.Journal: Translational psychiatryIn common: caret, data.table, ggplot2, 1 other tool, clinical / translational
- [9] doi:10.1038/s41588-026-02722-8 [code]
- A multiancestry polygenic risk score for Alzheimer's disease is associated with cognitive decline and neuropathological hallmarks in diverse populations.Journal: Nature geneticsIn common: data.table, ggplot2, tidyverse, Alzheimer's / dementia, genetics / omics, 1 reference
- [10] doi:10.1038/s41467-026-77170-3 [code]
- DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.Journal: Nature communicationsIn common: caret, data.table, ggplot2, 1 other tool, genetics / omics
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 9 scripts, and 8 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:6744821f81fc64d1…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
