OSCR

Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program.

Code ↔ Paper

8 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 8 matches
  1. [1] § Methods › Phenotype Ascertainment ↔ Code_Files/AoU_or_calculation.Rmd, lines 288–357 · score 0.97 · acute myocardial infarction, chronic ischemic heart, chronic rheumatic heart, heart failure, pulmonary embolism, angina pectoris
  2. [2] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation.Rmd, lines 18–84 · score 0.93 · annual income, alcohol consumption, physical activity, smoking status, diabetes status, vigorous
  3. [3] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 16–94 · score 0.93 · annual income, alcohol consumption, physical activity, smoking status, diabetes status, vigorous
  4. [4] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation.Rmd, lines 18–84 · score 0.78 · physical activity, ethnic background, alcohol, Depression, income, Age
  5. [5] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 16–94 · score 0.78 · physical activity, ethnic background, alcohol, Depression, income, Age
  6. [6] § Results › Cardiovascular Disease Subtypes and AD Prevalence ↔ Code_Files/AoU_or_calculation.Rmd, lines 288–357 · score 0.75 · chronic ischemic heart, heart failure, pulmonary embolism, cerebral infarction, hypertension, hypotension
  7. [7] § Methods › Genetic Analysis ↔ Code_Files/genetic_analysis.Rmd, lines 158–186 · score 0.72 · heart rate, heart function, AD GWAS, amyloid, tau, dementia
  8. [8] § Methods › Statistical Analysis ↔ Code_Files/UKB_or_calculation_revisions.Rmd, lines 607–657 · score 0.51 · logistic regression, Chinese, Unknown, Asian, strata, ethnic

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R Markdown · 761 lines · 22 KB · no license · 3 matches

  1. ---
  2. title: "OR Calculations"
  3. author: "Aili Toyli"
  4. date: "`r Sys.Date()`"
  5. output: html_document
  6. ---
  7. ```{r setup, include=FALSE}
  8. knitr::opts_chunk$set(echo = TRUE)
  9. library(tidyverse)
  10. #set working path
  11. PATH <- "~/"
  12. ```
  13. ##Read in OR datafile and prepare variables for analysis
  14. ```{r}
  15. #read in file
  16. cvd <- read.csv(PATH + "or_clean.csv")
  17. #clean dataframe
  18. cvd <- cvd %>%
  19. #remove unneeded columns
  20. select(-c(eid, p41270)) %>%
  21. #label covariates for better readability
  22. mutate(Age = p21003_i0,
  23. Sex = factor(p31, labels = c('Female', 'Male')),
  24. `Age completed full time education` = p845_i0,
  25. BMI = p21001_i0) %>%
  26. #Create readable smoking status levels
  27. mutate(`Smoking Status` = factor(p20116_i0, labels = c('Prefer not to answer', 'Never', 'Previous', 'Current'))) %>%
  28. #Label all types of depression/bipolar as yes, no if otherwise
  29. mutate(`Depression/Bipolar Disorder` = factor(p20126_i0, labels = c('No', rep('Yes', 5)))) %>%
  30. #Replace non response encodings from physical activity data with NA
  31. mutate(`Days of moderate physical activity each week` = replace(p884_i0, p884_i0 %in% c(-1, -3), NA)) %>%
  32. mutate(`Days of vigorous physical activity each week` = replace(p904_i0, p904_i0 %in% c(-1, -3), NA)) %>%
  33. #Label alcohol consumption levels
  34. mutate(`Alcohol Consumption` = factor(p1558_i0, labels = c('Daily or almost daily', '3-4 times a week', '1-2 times a week', '1-3 times a month', 'Special occasions only', 'Never', 'Prefer not to answer'))) %>%
  35. #Label annual income levels
  36. mutate(`Annual Household Income` = factor(p738_i0, labels = c('<£18,000', '£18,000-£30,999', '£31,000-£51,999', '£52,000-£100,000', '>£100,000', 'Do not know', 'Prefer not to answer'))) %>%
  37. #Save diabetes status as factor
  38. mutate(Diabetes = factor(Diabetes, labels = c('No', 'Yes'))) %>%
  39. #Replace spelling error on afib
  40. rename(Atrial.Fibrillation = Atrial.Fibriliation) %>%
  41. #replace blockage and afib with new arrhythmia category
  42. mutate(Cardiac.Arrhythmia = ifelse(Atrial.Fibrillation | Blockage, TRUE, FALSE)) %>%
  43. select(-c(Atrial.Fibrillation, Blockage)) %>%
  44. #Reorder columns to have AD followed by CVD, then covariates
  45. select(Alzheimers.Disease, Hypertension, Hypotension, Angina.Pectoris,
  46. Acute.Myocardial.Infarction, Pulmonary.Embolism, Cardiac.Arrhythmia,
  47. Heart.Failure, Chronic.Rheumatic.Heart.Disease,
  48. Chronic.Ischemic.Heart.Disease, Cerebral.Infarction, everything())
  49. #Adjust ethnicity encodings to be more generalized
  50. cvd$`Ethnic background` <- as.character(cvd$p21000_i0)
  51. cvd$`Ethnic background` <- ifelse(startsWith(cvd$`Ethnic background`, '1'), 'White',
  52. ifelse(startsWith(cvd$`Ethnic background`, '2'), 'Mixed',
  53. ifelse(startsWith(cvd$`Ethnic background`, '3'), 'Asian',
  54. ifelse(startsWith(cvd$`Ethnic background`, '4'), 'Black',
  55. ifelse(cvd$`Ethnic background`=='5', 'Chinese',
  56. ifelse(cvd$`Ethnic background`=='6', 'Other',
  57. ifelse(cvd$`Ethnic background`=='-3', 'Prefer not to answer',
  58. 'Unknown')))))))
  59. #drop original columns containing covariate data
  60. cvd<- cvd %>% select(-matches("^p[0-9]+"))
  61. #categorize ethnic background as a factor variable
  62. cvd$`Ethnic background` <- factor(cvd$`Ethnic background`)
  63. library(dplyr)
  64. library(caret)
  65. # 1. Impute missing categorical variables as "Prefer not to answer"
  66. cvd_clean <- cvd %>%
  67. mutate(across(where(is.factor), ~ {
  68. fct <- .
  69. # Add the level if not present
  70. if (!"Prefer not to answer" %in% levels(fct)) {
  71. levels(fct) <- c(levels(fct), "Prefer not to answer")
  72. }
  73. # Replace NA
  74. fct[is.na(fct)] <- "Prefer not to answer"
  75. fct
  76. }))
  77. #create baseline dataframe with mean imputation
  78. cvd_clean_mean <- cvd_clean %>%
  79. mutate(across(where(is.numeric), ~ {
  80. ifelse(is.na(.), mean(., na.rm = TRUE), .)
  81. }))
  82. ```
  83. ## Code for baseline OR calculation
  84. ```{r}
  85. #cvd data frame should have alzheimers in first column, then cvd subtypes
  86. #make sure eid is removed
  87. #some of the index numbers would need to be adjusted if looking at a different number of subtypes
  88. #create empty dataframe to store results
  89. output <- data.frame(cvd = character(10), est = numeric(10),
  90. lower.lim = numeric(10), upper.lim = numeric(10))
  91. #run regression for each subtype
  92. for(i in 1:10){
  93. #put name of disease in output dataframe
  94. output$cvd[i] <- names(cvd_clean_mean)[i+1]
  95. #save formula
  96. formula <- as.formula(paste0('`', names(cvd_clean_mean)[1], '` ~ `', names(cvd_clean_mean)[i+1], '` + ',
  97. paste(paste0('`', names(cvd_clean_mean)[12:ncol(cvd_clean_mean)], '`'), collapse = ' + '))
  98. )
  99. #run regression
  100. mylogit <- glm(formula, family = 'binomial', data = cvd_clean_mean)
  101. #save logistic regression summary
  102. sum_mylogit <- summary(mylogit)
  103. #save odds ratio and confidence interval
  104. or_confint <- exp(cbind(OR = sum_mylogit$coefficients[2, 1],
  105. LL = sum_mylogit$coefficients[2, 1]
  106. - 1.96 * sum_mylogit$coefficients[2, 2],
  107. UL = sum_mylogit$coefficients[2, 1]
  108. + 1.96 * sum_mylogit$coefficients[2, 2]))
  109. #save in output table
  110. output[i, 2] <- or_confint[1, 1]
  111. output[i, 3] <- or_confint[1, 2]
  112. output[i, 4] <- or_confint[1, 3]
  113. }
  114. #visualize with a forest plot
  115. plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
  116. geom_point(shape = 21, fill = "blue", size = 5) +
  117. geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
  118. geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
  119. annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
  120. vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
  121. coord_flip() +
  122. labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
  123. title = "Forest Plot of Odds Ratios\nfor AD with CVD Subtypes",
  124. subtitle = 'UK Biobank Cohort\nMean Imputation')+
  125. theme(axis.text.x = element_text(size = 12),
  126. axis.text.y = element_text(size = 16),
  127. plot.title = element_text(hjust = 0.5, size = 30),
  128. plot.subtitle = element_text(hjust = 0.5),
  129. axis.title = element_text(size = 20),
  130. legend.text= element_text(size = 12),
  131. legend.title = element_text(size = 20),
  132. panel.background = element_rect(fill = "white"))
  133. plot
  134. #save plot
  135. ggsave(PATH + 'or_plot_mean_UKB.png', plot = plot,
  136. width = 9, height = 9)
  137. #save datafile
  138. write.csv(output, PATH + 'or_results_mean.csv')
  139. ```
  140. ##Impute missing numeric data with mice, run analysis on multiple datasets
  141. ```{r}
  142. #load library
  143. library(mice)
  144. #inspect data
  145. md.pattern(cvd_clean)
  146. #perform imputation
  147. set.seed(123)
  148. library(parallel)
  149. cl <- makeCluster(4)
  150. imp <- mice(cvd_clean, m = 5, maxit = 5, cluster = cl)
  151. stopCluster(cl)
  152. ##get logistic regression odds ratio estimates
  153. library(mice)
  154. library(dplyr)
  155. ##--------------------------------------------------
  156. ## 1. Identify variables (by name, not index)
  157. ##--------------------------------------------------
  158. all_vars <- names(imp$data)
  159. outcome <- all_vars[1]
  160. exposures <- all_vars[2:11]
  161. covars <- all_vars[12:length(all_vars)]
  162. ##--------------------------------------------------
  163. ## 2. Storage for pooled results
  164. ##--------------------------------------------------
  165. output <- data.frame(
  166. cvd = exposures,
  167. OR = NA_real_,
  168. lower.lim = NA_real_,
  169. upper.lim = NA_real_
  170. )
  171. ##--------------------------------------------------
  172. ## 3. Parallelize odds ratio calculation
  173. ##--------------------------------------------------
  174. library(parallel)
  175. cl <- makeCluster(detectCores() - 1)
  176. clusterExport(
  177. cl,
  178. varlist = c("imp", "exposures", "outcome", "covars"),
  179. envir = environment()
  180. )
  181. clusterEvalQ(cl, {
  182. library(mice)
  183. })
  184. results <- parLapply(
  185. cl,
  186. seq_along(exposures),
  187. function(i) {
  188. cvd_var <- exposures[i]
  189. fit_imp <- with(
  190. imp,
  191. glm(
  192. as.formula(
  193. paste0(
  194. "`", outcome, "` ~ `", cvd_var, "` + ",
  195. paste0("`", covars, "`", collapse = " + ")
  196. )
  197. ),
  198. family = binomial
  199. )
  200. )
  201. pooled_sum <- summary(pool(fit_imp))
  202. coef_row <- which(pooled_sum$term == paste0(cvd_var, "TRUE"))
  203. if (length(coef_row) == 0) {
  204. return(c(NA, NA, NA))
  205. }
  206. beta <- pooled_sum$estimate[coef_row]
  207. se <- pooled_sum$std.error[coef_row]
  208. c(
  209. exp(beta),
  210. exp(beta - 1.96 * se),
  211. exp(beta + 1.96 * se)
  212. )
  213. }
  214. )
  215. stopCluster(cl)
  216. output <- data.frame(
  217. exposure = exposures,
  218. OR = sapply(results, `[`, 1),
  219. lower.lim = sapply(results, `[`, 2),
  220. upper.lim = sapply(results, `[`, 3)
  221. )
  222. output
  223. ## Visualize results
  224. #visualize with a forest plot
  225. plot <- ggplot(output, aes(x = reorder(exposure, OR), y = OR)) +
  226. geom_point(shape = 21, fill = "blue", size = 5) +
  227. geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
  228. geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
  229. annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
  230. vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
  231. coord_flip() +
  232. labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
  233. title = "Forest Plot of Odds Ratios\nfor AD with CVD Subtypes",
  234. subtitle = 'UK Biobank Cohort\nMICE Imputation')+
  235. theme(axis.text.x = element_text(size = 12),
  236. axis.text.y = element_text(size = 16),
  237. plot.title = element_text(hjust = 0.5, size = 30),
  238. plot.subtitle = element_text(hjust = 0.5),
  239. axis.title = element_text(size = 20),
  240. legend.text= element_text(size = 12),
  241. legend.title = element_text(size = 20),
  242. panel.background = element_rect(fill = "white"))
  243. plot
  244. #save plot
  245. ggsave(PATH + 'or_plot_mice_UKB.png', plot = plot,
  246. width = 9, height = 9)
  247. #save datafile
  248. write.csv(output, PATH + 'or_results_mice.csv')
  249. ```
  250. ## Calculating crosstabs
  251. ```{r}
  252. library(dplyr)
  253. library(purrr)
  254. # Variables of interest
  255. vars <- names(cvd)[2:11] # your CVD indicators (0/1 or FALSE/TRUE)
  256. # Function to compute one row per CVD subtype
  257. build_row <- function(var) {
  258. # Define vectors
  259. cvd_var <- cvd[[var]]
  260. ad_var <- cvd$Alzheimers.Disease
  261. # Prevalence of CVD subtype
  262. n_yes <- sum(cvd_var == TRUE, na.rm = TRUE)
  263. n_no <- sum(cvd_var == FALSE, na.rm = TRUE)
  264. n_total <- n_yes + n_no
  265. prev_str <- paste0(
  266. n_yes, " (", round(n_yes / n_total * 100, 1), "%)"
  267. )
  268. # AD among CVD=yes
  269. ad_yes_cvd_yes <- sum(ad_var == TRUE & cvd_var == TRUE, na.rm = TRUE)
  270. ad_rate_yes <- paste0(
  271. ad_yes_cvd_yes,
  272. " (", round(ad_yes_cvd_yes / n_yes * 100, 1), "%)"
  273. )
  274. # AD among CVD=no
  275. ad_yes_cvd_no <- sum(ad_var == TRUE & cvd_var == FALSE, na.rm = TRUE)
  276. ad_rate_no <- paste0(
  277. ad_yes_cvd_no,
  278. " (", round(ad_yes_cvd_no / n_no * 100, 1), "%)"
  279. )
  280. # Contingency table + p-value
  281. tbl <- table(cvd_var, ad_var)
  282. # Use Fisher's exact when needed
  283. pval <- if (any(tbl < 5)) {
  284. fisher.test(tbl)$p.value
  285. } else {
  286. chisq.test(tbl)$p.value
  287. }
  288. tibble(
  289. CVD_subtype = var,
  290. Prevalence = prev_str,
  291. AD_among_CVD_yes = ad_rate_yes,
  292. AD_among_CVD_no = ad_rate_no,
  293. p_value = signif(pval, 3)
  294. )
  295. }
  296. # Build the final table
  297. final_table <- map_dfr(vars, build_row)
  298. final_table
  299. # Write output
  300. write.csv(final_table,
  301. PATH + "ad_cvd_crosstabs.csv",
  302. row.names = FALSE)
  303. ```
  304. ## Firth regression to account for case-control imbalances
  305. ```{r}
  306. ## ============================================================
  307. ## Parallelized Firth Logistic Regression Across CVD Subtypes
  308. ## ============================================================
  309. ## -----------------------------
  310. ## 0. Libraries
  311. ## -----------------------------
  312. library(logistf)
  313. library(parallel)
  314. library(ggplot2)
  315. ## -----------------------------
  316. ## 1. Setup
  317. ## -----------------------------
  318. # outcome in first column, CVD subtypes next
  319. outcome_var <- names(cvd_clean_mean)[1]
  320. exposures <- names(cvd_clean_mean)[2:11] # 10 CVD subtypes
  321. covars <- names(cvd_clean_mean)[12:ncol(cvd_clean_mean)]
  322. # initialize output
  323. output <- data.frame(
  324. cvd = exposures,
  325. est = NA_real_,
  326. lower.lim = NA_real_,
  327. upper.lim = NA_real_,
  328. stringsAsFactors = FALSE
  329. )
  330. ## -----------------------------
  331. ## 2. Parallel backend
  332. ## -----------------------------
  333. ncores <- max(1, detectCores() - 2)
  334. cl <- makeCluster(ncores)
  335. clusterExport(
  336. cl,
  337. varlist = c("cvd_clean_mean", "outcome_var", "exposures", "covars"),
  338. envir = environment()
  339. )
  340. clusterEvalQ(cl, {
  341. library(logistf)
  342. })
  343. ## Non parallel firth regression
  344. results <- lapply(
  345. seq_along(exposures),
  346. function(i) {
  347. exposure <- exposures[i]
  348. message("Starting Firth regression for: ", exposure)
  349. # build formula
  350. form <- as.formula(
  351. paste0(
  352. "`", outcome_var, "` ~ `", exposure, "` + ",
  353. paste(paste0("`", covars, "`"), collapse = " + ")
  354. )
  355. )
  356. # fit Firth model
  357. fit <- try(
  358. logistf(
  359. formula = form,
  360. data = cvd_clean_mean,
  361. plconf = 2
  362. ),
  363. silent = TRUE
  364. )
  365. if (inherits(fit, "try-error")) {
  366. message("FAILED: ", exposure)
  367. return(c(NA, NA, NA))
  368. }
  369. coef_idx <- which(names(fit$coefficients) == paste0(exposure, "TRUE"))
  370. if (length(coef_idx) == 0) {
  371. message("Coefficient not found: ", exposure)
  372. return(c(NA, NA, NA))
  373. }
  374. beta <- fit$coefficients[coef_idx]
  375. ci_low <- fit$ci.lower[coef_idx]
  376. ci_up <- fit$ci.upper[coef_idx]
  377. message("Finished Firth regression for: ", exposure)
  378. c(
  379. exp(beta),
  380. exp(ci_low),
  381. exp(ci_up)
  382. )
  383. }
  384. )
  385. ## -----------------------------
  386. ## 4. Collect results
  387. ## -----------------------------
  388. output$est <- sapply(results, `[`, 1)
  389. output$lower.lim <- sapply(results, `[`, 2)
  390. output$upper.lim <- sapply(results, `[`, 3)
  391. ## -----------------------------
  392. ## 5. Forest plot
  393. ## -----------------------------
  394. plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
  395. geom_point(shape = 21, fill = "blue", size = 5) +
  396. geom_errorbar(
  397. aes(ymin = lower.lim, ymax = upper.lim),
  398. width = 0.15,
  399. linewidth = 0.75
  400. ) +
  401. geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
  402. coord_flip() +
  403. labs(
  404. x = "CVD Subtype",
  405. y = "Odds Ratio (95% CI)",
  406. title = "Firth Odds Ratios\nfor AD with CVD Subtypes",
  407. subtitle = "UK Biobank Cohort\nMean Imputation"
  408. ) +
  409. theme(
  410. axis.text.x = element_text(size = 12),
  411. axis.text.y = element_text(size = 16),
  412. plot.title = element_text(hjust = 0.5, size = 30),
  413. plot.subtitle = element_text(hjust = 0.5),
  414. axis.title = element_text(size = 20),
  415. panel.background = element_rect(fill = "white")
  416. )
  417. print(plot)
  418. ## -----------------------------
  419. ## 6. Save outputs
  420. ## -----------------------------
  421. ggsave(
  422. PATH + "or_plot_firth_UKB.png",
  423. plot = plot,
  424. width = 9,
  425. height = 9
  426. )
  427. write.csv(
  428. output,
  429. PATH + "or_results_firth_UKB.csv",
  430. row.names = FALSE
  431. )
  432. ```
  433. ## OR calculation with comorbidity adjustment
  434. ```{r}
  435. #cvd data frame should have alzheimers in first column, then cvd subtypes
  436. #make sure eid is removed
  437. #some of the index numbers would need to be adjusted if looking at a different number of subtypes
  438. #create empty dataframe to store results
  439. output <- data.frame(cvd = character(10), est = numeric(10),
  440. lower.lim = numeric(10), upper.lim = numeric(10))
  441. #run regression for each subtype
  442. for(i in 1:10){
  443. #put name of disease in output dataframe
  444. output$cvd[i] <- names(cvd_clean_mean)[i+1]
  445. #save formula
  446. formula <- as.formula(paste0('`', names(cvd_clean_mean)[1], '` ~ `', names(cvd_clean_mean)[i+1], '` + .')
  447. )
  448. #run regression
  449. mylogit <- glm(formula, family = 'binomial', data = cvd_clean_mean)
  450. #save logistic regression summary
  451. sum_mylogit <- summary(mylogit)
  452. #save odds ratio and confidence interval
  453. or_confint <- exp(cbind(OR = sum_mylogit$coefficients[2, 1],
  454. LL = sum_mylogit$coefficients[2, 1]
  455. - 1.96 * sum_mylogit$coefficients[2, 2],
  456. UL = sum_mylogit$coefficients[2, 1]
  457. + 1.96 * sum_mylogit$coefficients[2, 2]))
  458. #save in output table
  459. output[i, 2] <- or_confint[1, 1]
  460. output[i, 3] <- or_confint[1, 2]
  461. output[i, 4] <- or_confint[1, 3]
  462. }
  463. #visualize with a forest plot
  464. plot <- ggplot(output, aes(x = reorder(cvd, est), y = est)) +
  465. geom_point(shape = 21, fill = "blue", size = 5) +
  466. geom_errorbar(aes(ymin = lower.lim, ymax = upper.lim), width = 0.15, linewidth = 0.75, color = 'black') +
  467. geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
  468. annotate("text", x = 10.35, y = 1.52, label = "Significance\nThreshold ",
  469. vjust = -0.5, hjust = 0.95, color = "red", size = 5) +
  470. coord_flip() +
  471. labs(x = "CVD Subtype", y = "Odds Ratio (95% CI)",
  472. title = "Comorbidity-Adjusted\nOdds Ratios",
  473. subtitle = 'UK Biobank Cohort')+
  474. theme(axis.text.x = element_text(size = 12),
  475. axis.text.y = element_text(size = 16),
  476. plot.title = element_text(hjust = 0.5, size = 30),
  477. plot.subtitle = element_text(hjust = 0.5),
  478. axis.title = element_text(size = 20),
  479. legend.text= element_text(size = 12),
  480. legend.title = element_text(size = 20),
  481. panel.background = element_rect(fill = "white"))
  482. plot
  483. #save plot
  484. ggsave(PATH + 'or_plot_comorbid_UKB.png', plot = plot,
  485. width = 9, height = 9)
  486. #save datafile
  487. write.csv(output, PATH + 'or_results_comorbid.csv')
  488. ```
  489. ## OR code with ethnically stratified sampling
  490. ```{r}
  491. ############################################################
  492. ## Stratified Logistic Regression for AD ~ CVD Subtypes
  493. ## Stratified by Ethnicity4 (UK Biobank)
  494. ############################################################
  495. library(dplyr)
  496. library(forcats)
  497. library(ggplot2)
  498. set.seed(123)
  499. #--------------------------------------------------
  500. # 1. CREATE ETHNICITY STRATA
  501. #--------------------------------------------------
  502. cvd_clean <- cvd_clean %>%
  503. mutate(
  504. Ethnicity4 = fct_collapse(
  505. `Ethnic background`,
  506. White = "White",
  507. Black = "Black",
  508. Asian = c("Chinese", "Asian"),
  509. Other.Unknown = c("Mixed", "Other", "Unknown", "Prefer not to answer")
  510. )
  511. ) %>%
  512. mutate(Ethnicity4 = droplevels(as.factor(Ethnicity4))) %>%
  513. select(-`Ethnic background`)
  514. #--------------------------------------------------
  515. # 2. VARIABLE SETUP
  516. #--------------------------------------------------
  517. outcome_var <- names(cvd_clean)[1] # AD outcome
  518. exposure_vars <- names(cvd_clean)[2:11] # CVD subtypes
  519. covariates <- names(cvd_clean)[12:ncol(cvd_clean)] # comorbidities
  520. #--------------------------------------------------
  521. # 3. OUTPUT DATAFRAME (LONG FORMAT)
  522. #--------------------------------------------------
  523. output <- expand.grid(
  524. cvd = exposure_vars,
  525. Ethnicity4 = levels(cvd_clean$Ethnicity4),
  526. stringsAsFactors = FALSE
  527. )
  528. output$est <- NA_real_
  529. output$lower.lim <- NA_real_
  530. output$upper.lim <- NA_real_
  531. #--------------------------------------------------
  532. # 4. STRATIFIED LOGISTIC REGRESSION
  533. #--------------------------------------------------
  534. row_id <- 1
  535. for (eth in levels(cvd_clean$Ethnicity4)) {
  536. data_eth <- cvd_clean %>%
  537. filter(Ethnicity4 == eth)
  538. # Skip strata with very small sample size
  539. if (nrow(data_eth) < 50) next
  540. for (exposure in exposure_vars) {
  541. # Drop covariates with no variation in this stratum
  542. cov_eth <- covariates[
  543. sapply(data_eth[covariates], function(x) length(unique(x)) > 1)
  544. ]
  545. # Build formula with backticks
  546. formula <- as.formula(
  547. paste0(
  548. "`", outcome_var, "` ~ `", exposure, "` + ",
  549. paste(paste0("`", cov_eth, "`"), collapse = " + ")
  550. )
  551. )
  552. # Fit logistic regression
  553. fit <- glm(
  554. formula,
  555. family = binomial,
  556. data = data_eth
  557. )
  558. coef_summary <- summary(fit)$coefficients
  559. # Exposure coefficient (always row 2)
  560. beta <- coef_summary[2, 1]
  561. se <- coef_summary[2, 2]
  562. # OR and 95% CI
  563. or_ci <- exp(c(
  564. beta,
  565. beta - 1.96 * se,
  566. beta + 1.96 * se
  567. ))
  568. output[row_id, c("est", "lower.lim", "upper.lim")] <- or_ci
  569. row_id <- row_id + 1
  570. }
  571. }
  572. # Remove rows that were skipped
  573. output <- output %>% filter(!is.na(est))
  574. # Remove rows where the upper 95% CI exceeds 10
  575. output_filtered <- output %>%
  576. filter(upper.lim <= 10)
  577. #--------------------------------------------------
  578. # 5. FOREST PLOT (FACETED BY ETHNICITY)
  579. #--------------------------------------------------
  580. plot <- ggplot(output_filtered, aes(x = reorder(cvd, est), y = est)) +
  581. geom_point(size = 3, color = "blue") +
  582. geom_errorbar(
  583. aes(ymin = lower.lim, ymax = upper.lim),
  584. width = 0.15
  585. ) +
  586. geom_hline(yintercept = 1, linetype = "dashed", color = "red") +
  587. coord_flip() +
  588. facet_wrap(~ Ethnicity4) +
  589. labs(
  590. x = "CVD Subtype",
  591. y = "Odds Ratio (95% CI)",
  592. title = "Stratified Odds Ratios for AD by Ethnicity",
  593. subtitle = "Adjusted Logistic Regression (UK Biobank)"
  594. ) +
  595. theme_bw(base_size = 14)
  596. print(plot)
  597. #--------------------------------------------------
  598. # 6. SAVE OUTPUTS
  599. #--------------------------------------------------
  600. ggsave(
  601. PATH + "or_plot_stratified_ethnicity_UKB.png",
  602. plot = plot,
  603. width = 12,
  604. height = 8
  605. )
  606. write.csv(
  607. output,
  608. PATH + "or_results_stratified_ethnicity_UKB.csv",
  609. row.names = FALSE
  610. )
  611. ```

UKB_or_calculation_revisions.Rmd at commit 6ac93bd, no license · at the source

Overview

Authors: Aili Toyli1, Chen Zhao2, Kuan‐Jui Su3, Hui Shen3, Hong‐Wen Deng3, Qing‐Hui Chen4, Qiuying Sha1, Weihua Zhou5,6
  1. Department of Mathematical Sciences, Michigan Technological University, Houghton, MI, USA
  2. Department of Computer Science, Kennesaw State University, Marietta, GA, USA
  3. Division of Biomedical Informatics and Genomics, Tulane Center of Biomedical Informatics and Genomics, Deming Department of Medicine, Tulane University, New Orleans, LA, USA
  4. Department of Kinesiology and Integrative Physiology, Michigan Technological University, Houghton, MI, USA
  5. Department of Applied Computing, Michigan Technological University, Houghton, MI, USA
  6. Center for Biocomputing and Digital Health, Institute of Computing and Cybersystems, and Health Research Institute, Michigan Technological University, Houghton, MI, USA
Institutions: Michigan Technological University (United States); Kennesaw State University (United States); Tulane University (United States)
Journal: Journal of the American Heart Association, volume 15, issue 12, article e046172
Dates: received 28 August 2025; accepted 11 March 2026; published online 10 June 2026; in print June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1161/jaha.125.046172 · PMID 42267709 · PMCID PMC13323167 · OpenAlex W7164209024
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Alzheimer's / dementia (population), stroke (population), clinical / translational (subfield)
Methods: Machine learning
Keywords: Alzheimer's disease, cardiovascular disease, cerebral infarction, heart–brain axis, hypotension, Risk Factors, Aging
MeSH: Alzheimer Disease*, Cardiovascular Diseases*, Aged, Biological Specimen Banks, Cross-Sectional Studies, Female, Genetic Predisposition to Disease, Genome-Wide Association Study, Humans, Male, Middle Aged, Phenotype, Polymorphism, Single Nucleotide, Risk Factors, UK Biobank, United Kingdom, United States (* major topic)
Topic: Genetic Associations and Epidemiology (Genetics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NHLBI NIH HHS (R15 HL172198, R15 HL173852); NIA NIH HHS (U19 AG055373)
Citations: not cited yet (Europe PMC); 56 references in the paper

Abstract

Background: Cardiovascular disease (CVD) and Alzheimer's disease (AD) are major public health concerns that share overlapping risk factors and potential mechanistic pathways. Although vascular contributions to cognitive decline are well documented, the specific relationships between AD and different CVD subtypes remain poorly understood.

Methods: In this cross‐sectional study, we examined associations between AD and 11 CVD subtypes using logistic regression models in 2 large biobanks: the UK Biobank (n=502 133) and the All of Us Research Program (n=287 011). Models were adjusted for demographic, lifestyle, and clinical covariates. We also explored genetic overlap between AD and CVD traits through proximity‐based analysis of significant single nucleotide variants (P<5 × 10−8) using genome‐wide association study data.

Results: Most CVD subtypes were significantly associated with AD in both cohorts. Hypotension had the strongest and most consistent association, although it has been comparatively understudied in AD research. Strong associations were also consistently observed between AD and hypertension and cerebral infarction. Notably, acute myocardial infarction was not significantly linked to AD. Genetic analyses revealed shared loci between AD‐ and CVD‐related traits, particularly in regions near APOE, MAPT, and genes influencing myocardial structure and vascular function.

Conclusions: This study identifies subtype‐specific CVD associations with AD across 2 diverse cohorts and highlights shared genetic architecture underlying heart–brain interactions. These findings underscore the importance of vascular health in AD risk and suggest that certain CVD subtypes, especially hypotension, may play underrecognized roles in cognitive decline.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 8 matches between paragraphs and lines of code.

MIILab-MTU/CVDSubtypesADCorrelation

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 6ac93bda9b81f5902390838815361ea52885ad6f, 12 June 2026
Languages: R (6), Jupyter (3)
Size: 35 files, 9 scripts
Software Heritage: not archived
Found in: the text, “Study Population”
Holds: README, 7 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (7 files), ggplot2 (3 files), caret (1 file), data.table (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
10 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 9 scripts, each with its path and the digest of its content;
  • 8 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 8 authors, 7 keywords, 17 MeSH terms, 2 funders, 50 references.

Cite

This paper

Toyli, A., Zhao, C., Su, K., Shen, H., Deng, H., Chen, Q., Sha, Q., & Zhou, W. (2026). Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program. Journal of the American Heart Association, 15(12), e046172. https://doi.org/10.1161/jaha.125.046172

BibTeX

@article{toyli2026cardiovascular,
author = {Toyli, Aili and Zhao, Chen and Su, Kuan‐Jui and Shen, Hui and Deng, Hong‐Wen and Chen, Qing‐Hui and Sha, Qiuying and Zhou, Weihua},
title = {{Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program}},
journal = {Journal of the American Heart Association},
year = {2026},
month = jun,
volume = {15},
number = {12},
pages = {e046172},
publisher = {Wiley},
issn = {2047-9980},
doi = {10.1161/jaha.125.046172},
url = {https://doi.org/10.1161/jaha.125.046172},
pmid = {42267709},
pmcid = {PMC13323167}
}

RIS

TY - JOUR
AU - Toyli, Aili
AU - Zhao, Chen
AU - Su, Kuan‐Jui
AU - Shen, Hui
AU - Deng, Hong‐Wen
AU - Chen, Qing‐Hui
AU - Sha, Qiuying
AU - Zhou, Weihua
TI - Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program
T2 - Journal of the American Heart Association
J2 - J Am Heart Assoc
PY - 2026
DA - 2026/06/10
VL - 15
IS - 12
SP - e046172
SN - 2047-9980
PB - Wiley
DO - 10.1161/jaha.125.046172
UR - https://doi.org/10.1161/jaha.125.046172
LA - en
ER -

CSL-JSON

{
"id": "10.1161/jaha.125.046172",
"type": "article-journal",
"title": "Cardiovascular Disease Subtypes and Alzheimer's Disease: Phenotypic and Genetic Associations in the UK Biobank and All of Us Research Program",
"container-title": "Journal of the American Heart Association",
"author": [
{
"family": "Toyli",
"given": "Aili"
},
{
"family": "Zhao",
"given": "Chen"
},
{
"family": "Su",
"given": "Kuan‐Jui"
},
{
"family": "Shen",
"given": "Hui"
},
{
"family": "Deng",
"given": "Hong‐Wen"
},
{
"family": "Chen",
"given": "Qing‐Hui"
},
{
"family": "Sha",
"given": "Qiuying"
},
{
"family": "Zhou",
"given": "Weihua"
}
],
"container-title-short": "J Am Heart Assoc",
"volume": "15",
"issue": "12",
"page": "e046172",
"DOI": "10.1161/jaha.125.046172",
"PMID": "42267709",
"PMCID": "PMC13323167",
"ISSN": "2047-9980",
"publisher": "Wiley",
"URL": "https://doi.org/10.1161/jaha.125.046172",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s40168-026-02342-8 [code]
Impacts of host genetics on gut microbiome composition in Alzheimer's disease.
Journal: Microbiome
In common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics, 1 reference
[2] doi:10.1038/s41467-026-70374-7 [code]
scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies.
Journal: Nature communications
In common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics
[3] doi:10.1038/s41597-026-06971-4 [code]
Human neuronal differentiation under Aβ exposure: a single-cell transcriptomic and epigenomic dataset.
Journal: Scientific data
In common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics
[4] doi:10.1093/neuonc/noag128 [code]
Spatially-resolved single-cell imaging of melanoma brain metastases identifies localized immune patterns predictive of immune checkpoint blockade response.
Journal: Neuro-oncology
In common: caret, data.table, ggplot2, 1 other tool, clinical / translational, genetics / omics
[5] doi:10.1111/cts.70577 [code]
Machine Learning-Based Prediction of Drug-Induced QTc Changes in a Large Finnish Biobank Cohort.
Journal: Clinical and translational science
In common: caret, data.table, ggplot2, 1 other tool, clinical / translational, genetics / omics
[6] doi:10.1111/acel.70499 [code]
Brain Aging Mediating Heart Imaging-Derived Phenotypes and Mental and Nervous System Disorders.
Journal: Aging cell
In common: caret, data.table, tidyverse, 1 reference
[7] doi:10.1093/bib/bbag175 [code]
Uncovering causal relationships in single-cell omic studies with causarray.
Journal: Briefings in bioinformatics
In common: caret, data.table, ggplot2, 1 other tool, Alzheimer's / dementia
[8] doi:10.1038/s41398-026-04131-1 [code]
Multimodal phenotypic classification of generalized anxiety and panic using structural MRI data and psychosocial factors: machine learning results from the German National Cohort (NAKO) study.
Journal: Translational psychiatry
In common: caret, data.table, ggplot2, 1 other tool, clinical / translational
[9] doi:10.1038/s41588-026-02722-8 [code]
A multiancestry polygenic risk score for Alzheimer's disease is associated with cognitive decline and neuropathological hallmarks in diverse populations.
Journal: Nature genetics
In common: data.table, ggplot2, tidyverse, Alzheimer's / dementia, genetics / omics, 1 reference
[10] doi:10.1038/s41467-026-77170-3 [code]
DNA methylation profiling identifies long-range epigenetic silencing of clustered protocadherins as a key determinant of meningioma progression.
Journal: Nature communications
In common: caret, data.table, ggplot2, 1 other tool, genetics / omics

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.