OSCR

Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § METHODS › Privacy evaluation ↔ R/benchmark.plots.R, lines 273–317 · score 0.67 · privacy loss, hitting rate, identifiability risk, ADR, NNDR, NNAA
  2. [2] § RESULTS › Privacy analysis ↔ R/bmk_ctgan.R, the whole file · a weak match · score 0.60 · confidence interval, hitting rate, identifiability risk, recall, MIA, CTGAN
  3. [3] § RESULTS › Privacy analysis ↔ R/bmk_degrees.R, lines 1–72 · score 0.59 · confidence interval, hitting rate, identifiability risk, recall, MIA, score
  4. [4] § METHODS › Generating synthetic tabular data ↔ python/bayesian_synthesizing.py, lines 17–59 · score 0.56 · correlated attribute mode, Bayesian networks, root, privacy
  5. [5] § METHODS › Utility evaluation › Missing values ↔ python/train_svm.py, lines 108–116 · score 0.56 · IterativeImputer, SimpleImputer, imputed, training
  6. [6] § METHODS › Utility evaluation › Missing values ↔ python/main_ctgan.py, lines 31–43 · score 0.56 · IterativeImputer, SimpleImputer, imputed, CTGAN
  7. [7] § METHODS › Privacy evaluation ↔ python/eval_benchmark.py, lines 111–140 · score 0.53 · SynthEval, hitting rate, NNDR, benchmarking, NNAA, MIA

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 717 lines · 28 KB · MIT · 1 match

  1. library(dplyr)
  2. library(Rmisc)
  3. library(ggplot2)
  4. library(ggpubr)
  5. library(stringr)
  6. library(ggforce)
  7. library(paletteer)
  8. library(ggsci)
  9. ##### ADNI #####
  10. epsilons <- c(5, 10, 50, 100, 200, NA)
  11. samples <- c(rep(100, length(epsilons)-length(which(is.na(epsilons)))), 18)
  12. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/adni/bmk_new_deg2_eps", ifelse(is.na(epsilons), "zero", epsilons), ".csv")
  13. # Read CSVs in a loop
  14. res_list <- list()
  15. for (i in seq_along(epsilons)) {
  16. df <- read.csv(file_paths[i])
  17. df$Epsilon <- epsilons[i]
  18. df$samples <- samples[i]
  19. res_list[[i]] <- df
  20. }
  21. # Combine all data frames
  22. bmk_ds <- bind_rows(res_list)
  23. util_cols <- c("mutual_inf_diff_value", "ks_tvd_stat_value",
  24. "frac_ks_sigs_value","avg_F1_diff_value", "avg_F1_diff_hout_value",
  25. "nnaa_value")
  26. priv_cols <- c("priv_loss_nndr_value", "priv_loss_nnaa_value","hit_rate_value",
  27. "avg_nndr_value", "eps_identif_risk_value", "priv_loss_eps_value",
  28. "mia_recall_value", "att_discl_risk_value")
  29. cols <- c(util_cols, priv_cols)
  30. bmk_ds$Epsilon <- bmk_ds %>%
  31. select(Epsilon) %>%
  32. mutate_all(~replace(., is.na(.), 0)) %>%
  33. mutate(Epsilon = factor(Epsilon, levels = as.character(unique(Epsilon)))) %>%
  34. pull(Epsilon)
  35. # CTGAN
  36. epochs <- c(750)
  37. setting <- c("default", "optim")
  38. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/adni/bmk_new_ctgan_", setting, "_epochs_", epochs, ".csv")
  39. #file_paths <- c(file_paths, "~/Python/WASP-DDLS/SE-benchmark/bmk_ctgan_epochs_100.csv")
  40. # Read CSVs in a loop
  41. res_list <- list()
  42. for (i in seq_along(file_paths)) {
  43. df <- read.csv(file_paths[i])
  44. df$Epochs <- epochs[i]
  45. df$samples <- 100
  46. res_list[[i]] <- df
  47. }
  48. # Combine all data frames
  49. bmk_ctgan <- bind_rows(res_list)
  50. bmk_ctgan$Epochs <- as.factor(bmk_ctgan$Epochs)
  51. # Synthpop
  52. bmk_synthpop <- read.csv("~/Python/WASP-DDLS/SE-benchmark/adni/bmk_synthpop_2.csv") %>% mutate(
  53. Epochs = 0,
  54. samples = 100
  55. )
  56. # TabPFN
  57. temps <- c('1.0', '0.75', '0.5', '0.25')
  58. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/adni/bmk_new_tabpfn_t_", temps, ".csv")
  59. # Read CSVs in a loop
  60. res_list <- list()
  61. for (i in seq_along(file_paths)) {
  62. df <- read.csv(file_paths[i])
  63. df$temp <- temps[i]
  64. df$samples <- 100
  65. res_list[[i]] <- df
  66. }
  67. # Combine all data frames
  68. bmk_tabpfn <- bind_rows(res_list)
  69. bmk_tabpfn$temp <- as.factor(bmk_tabpfn$temp)
  70. label_list = c("DS (5)", "DS (10)","DS (50)","DS (100)",
  71. "DS (200)", "DS w/o DP", "CTGAN (default)", "CTGAN (optimal)",
  72. "Synthpop",
  73. "TabPFN (t=1.0)", "TabPFN (t=0.75)", "TabPFN (t=0.5)", "TabPFN (t=0.25)")
  74. util_df_adni <- rbind(select(bmk_ds, dataset, model, all_of(util_cols)),
  75. select(bmk_ctgan, dataset, model, all_of(util_cols)),
  76. select(bmk_synthpop, dataset, model, all_of(util_cols)),
  77. select(bmk_tabpfn, dataset, model, all_of(util_cols))) |>
  78. dplyr::mutate(avg_F1_diff_value = abs(avg_F1_diff_value),
  79. avg_F1_diff_hout_value = abs(avg_F1_diff_hout_value)) |>
  80. dplyr::mutate(mutual_inf_diff = 1 - tanh(mutual_inf_diff_value),
  81. ks_tvd_stat = 1 - ks_tvd_stat_value,
  82. frac_ks_sigs = 1 - frac_ks_sigs_value,
  83. avg_F1_diff = 1 - abs(avg_F1_diff_value),
  84. avg_F1_diff_hout = 1 - abs(avg_F1_diff_hout_value),
  85. nnaa = 1 - nnaa_value)
  86. util_df_adni$util_score <- util_df_adni |> dplyr::select(c(mutual_inf_diff, ks_tvd_stat, frac_ks_sigs, avg_F1_diff, avg_F1_diff_hout, nnaa)) |> rowMeans()
  87. util_df_adni$Method <- factor(util_df_adni$dataset, levels = unique(util_df_adni$dataset),
  88. labels=label_list)
  89. priv_df_adni <- rbind(select(bmk_ds, dataset, model, all_of(priv_cols)),
  90. select(bmk_ctgan, dataset, model, all_of(priv_cols)),
  91. select(bmk_synthpop, dataset, model, all_of(priv_cols)),
  92. select(bmk_tabpfn, dataset, model, all_of(priv_cols))) |>
  93. dplyr::mutate(priv_loss_nndr = 1 - abs(priv_loss_nndr_value),
  94. priv_loss_nnaa = 1 - abs(priv_loss_nnaa_value),
  95. priv_loss_eps = 1 - abs(priv_loss_eps_value),
  96. hit_rate = 1 - hit_rate_value,
  97. eps_identif_risk = 1 - eps_identif_risk_value,
  98. mia_recall = 1 - mia_recall_value,
  99. att_discl_risk = 1 - att_discl_risk_value)
  100. priv_df_adni$priv_score <- priv_df_adni |> dplyr::select(c(priv_loss_nndr, priv_loss_nnaa, priv_loss_eps, hit_rate, eps_identif_risk, att_discl_risk)) |> rowMeans()
  101. priv_df_adni$Method <- factor(priv_df_adni$dataset, levels = unique(priv_df_adni$dataset),
  102. labels=label_list)
  103. up_df_adni <- util_df_adni |> dplyr::select(c(Method, model, util_score)) |>
  104. dplyr::inner_join(priv_df_adni |> dplyr::select(c(Method, model, priv_score)), by = c("Method", "model"))
  105. colormap <- c(
  106. paletteer_d("rcartocolor::BluYl")[1:6],
  107. paletteer_d("beyonce::X58")[3:4],
  108. "#663399",
  109. paletteer_d("ggsci::cyan_material")[c(1,3,5,7)]
  110. )
  111. up_plot_adni <- ggplot(data = up_df_adni, aes(x=util_score, y=priv_score, fill=Method)) +
  112. geom_point(pch=21, color = "black", size=3) +
  113. labs(title = "ADNI", x = "Utility", y = "Privacy") +
  114. #scale_fill_paletteer_d("rcartocolor::BluYl") +
  115. scale_fill_manual(values = colormap) +
  116. theme_minimal() +
  117. ylim(0.7,1.0) + xlim(0.3,0.85)
  118. plot(up_plot_adni)
  119. #ggsave("~/R/DDLS-plots/up_plot.png", up_plot, width = 7, height = 5, units = "in", dpi = 500)
  120. ###### A4 #########
  121. epsilons <- c(5, 10, 50, 100, 200, NA)
  122. samples <- c(rep(100, length(epsilons)-length(which(is.na(epsilons)))), 23)
  123. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/a4/bmk_new_deg2_eps", ifelse(is.na(epsilons), "zero", epsilons), ".csv")
  124. # Read CSVs in a loop
  125. res_list <- list()
  126. for (i in seq_along(epsilons)) {
  127. df <- read.csv(file_paths[i])
  128. df$Epsilon <- epsilons[i]
  129. df$samples <- samples[i]
  130. res_list[[i]] <- df
  131. }
  132. # Combine all data frames
  133. bmk_ds <- bind_rows(res_list)
  134. util_cols <- c("mutual_inf_diff_value", "ks_tvd_stat_value",
  135. "frac_ks_sigs_value","avg_F1_diff_value", "avg_F1_diff_hout_value",
  136. "nnaa_value")
  137. priv_cols <- c("priv_loss_nndr_value", "priv_loss_nnaa_value","hit_rate_value",
  138. "avg_nndr_value", "eps_identif_risk_value", "priv_loss_eps_value",
  139. "mia_recall_value", "att_discl_risk_value")
  140. cols <- c(util_cols, priv_cols)
  141. bmk_ds$Epsilon <- bmk_ds %>%
  142. select(Epsilon) %>%
  143. mutate_all(~replace(., is.na(.), 0)) %>%
  144. mutate(Epsilon = factor(Epsilon, levels = as.character(unique(Epsilon)))) %>%
  145. pull(Epsilon)
  146. # CTGAN
  147. epochs <- c(750)
  148. setting <- c("default", "optim")
  149. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/a4/bmk_new_ctgan_", setting, "_epochs_", epochs, ".csv")
  150. #file_paths <- c(file_paths, "~/Python/WASP-DDLS/SE-benchmark/bmk_ctgan_epochs_100.csv")
  151. # Read CSVs in a loop
  152. res_list <- list()
  153. for (i in seq_along(file_paths)) {
  154. df <- read.csv(file_paths[i])
  155. df$Epochs <- epochs[i]
  156. df$samples <- 100
  157. res_list[[i]] <- df
  158. }
  159. # Combine all data frames
  160. bmk_ctgan <- bind_rows(res_list)
  161. bmk_ctgan$Epochs <- as.factor(bmk_ctgan$Epochs)
  162. # Synthpop
  163. bmk_synthpop <- read.csv("~/Python/WASP-DDLS/SE-benchmark/a4/bmk_new_synthpop.csv") %>% mutate(
  164. Epochs = 0,
  165. samples = 100
  166. )
  167. # TabPFN
  168. temps <- c('1.0')
  169. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/a4/bmk_new_tabpfn_t_", temps, ".csv")
  170. # Read CSVs in a loop
  171. res_list <- list()
  172. for (i in seq_along(file_paths)) {
  173. df <- read.csv(file_paths[i])
  174. df$temp <- temps[i]
  175. df$samples <- 100
  176. res_list[[i]] <- df
  177. }
  178. # Combine all data frames
  179. bmk_tabpfn <- bind_rows(res_list)
  180. bmk_tabpfn$temp <- as.factor(bmk_tabpfn$temp)
  181. label_list = c("DS (5)", "DS (10)", "DS (50)", "DS (100)", "DS (200)", "DS w/o DP",
  182. "CTGAN (default)", "CTGAN (optimal)",
  183. "Synthpop", "TabPFN (t=1.0")
  184. util_df_a4 <- rbind(select(bmk_ds, dataset, model, all_of(util_cols)),
  185. select(bmk_ctgan, dataset, model, all_of(util_cols)),
  186. select(bmk_synthpop, dataset, model, all_of(util_cols)),
  187. select(bmk_tabpfn, dataset, model, all_of(util_cols))
  188. ) |>
  189. dplyr::mutate(avg_F1_diff_value = abs(avg_F1_diff_value),
  190. avg_F1_diff_hout_value = abs(avg_F1_diff_hout_value)) |>
  191. dplyr::mutate(mutual_inf_diff = 1 - tanh(mutual_inf_diff_value),
  192. ks_tvd_stat = 1 - ks_tvd_stat_value,
  193. frac_ks_sigs = 1 - frac_ks_sigs_value,
  194. avg_F1_diff = 1 - abs(avg_F1_diff_value),
  195. avg_F1_diff_hout = 1 - abs(avg_F1_diff_hout_value),
  196. nnaa = 1 - nnaa_value)
  197. util_df_a4$util_score <- util_df_a4 |> dplyr::select(c(mutual_inf_diff, ks_tvd_stat, frac_ks_sigs, avg_F1_diff, avg_F1_diff_hout, nnaa)) |> rowMeans()
  198. util_df_a4$Method <- factor(util_df_a4$dataset, levels = unique(util_df_a4$dataset),
  199. labels=label_list)
  200. priv_df_a4 <- rbind(select(bmk_ds, dataset, model, all_of(priv_cols)),
  201. select(bmk_ctgan, dataset, model, all_of(priv_cols)),
  202. select(bmk_synthpop, dataset, model, all_of(priv_cols)),
  203. select(bmk_tabpfn, dataset, model, all_of(priv_cols))
  204. ) |>
  205. dplyr::mutate(priv_loss_nndr = 1 - abs(priv_loss_nndr_value),
  206. priv_loss_nnaa = 1 - abs(priv_loss_nnaa_value),
  207. priv_loss_eps = 1 - abs(priv_loss_eps_value),
  208. hit_rate = 1 - hit_rate_value,
  209. eps_identif_risk = 1 - eps_identif_risk_value,
  210. mia_recall = 1 - mia_recall_value,
  211. att_discl_risk = 1 - att_discl_risk_value)
  212. priv_df_a4$priv_score <- priv_df_a4 |> dplyr::select(c(priv_loss_nndr, priv_loss_nnaa, priv_loss_eps, hit_rate, eps_identif_risk, att_discl_risk)) |> rowMeans()
  213. priv_df_a4$Method <- factor(priv_df_a4$dataset, levels = unique(priv_df_a4$dataset),
  214. labels=label_list)
  215. up_df_a4 <- util_df_a4 |> dplyr::select(c(Method, model, util_score)) |>
  216. dplyr::inner_join(priv_df_a4 |> dplyr::select(c(Method, model, priv_score)), by = c("Method", "model"))
  217. up_df_a4 <- up_df_a4 %>% mutate(Framework = as.factor(gsub(" .*", "", Method)))
  218. #up_df$Method <- factor(up_df$dataset, levels = unique(up_df$dataset),
  219. # labels=c("DS (eps=5)", "DS (eps=10)", "DS (eps=25)","DS (eps=50)","DS (eps=100)",
  220. # "DS (eps=200)", "DS w/o DP", "CTGAN (10 epochs)",
  221. # "CTGAN (50 epochs)", "CTGAN (100 epochs)", "Synthpop"))
  222. colormap <- c(
  223. paletteer_d("rcartocolor::BluYl")[1:6],
  224. paletteer_d("beyonce::X58")[3:4],
  225. "#663399",
  226. paletteer_d("ggsci::cyan_material")[c(1,3,5,7)]
  227. )
  228. up_plot_a4 <- ggplot(data = up_df_a4, aes(x=util_score, y=priv_score, fill=Method)) +
  229. geom_point(pch=21, color="black", size=3) +
  230. labs(title = "A4", x = "Utility", y = "Privacy") +
  231. #scale_fill_paletteer_d("rcartocolor::BluYl") +
  232. scale_fill_manual(values = colormap) +
  233. theme_minimal() +
  234. ylim(0.7,1.0) + xlim(0.3,0.85)
  235. #plot(up_plot_a4)
  236. up_plot <- ggarrange(up_plot_adni, up_plot_a4, nrow=2, common.legend = TRUE, legend = "right")
  237. plot(up_plot)
  238. #ggsave("~/R/DDLS-plots/up_plot.jpg", up_plot, width = 7, height = 10, units = "in", dpi = 1000)
  239. ###### GRIDS OF METRICS ######
  240. library(lme4)
  241. # Privacy grid
  242. privs <- c("priv_loss_nndr_value", "priv_loss_nnaa_value", "eps_identif_risk_value", "hit_rate_value", "mia_recall_value", "att_discl_risk_value")
  243. priv_titels <- c(
  244. priv_loss_nndr_value = "NNDR privacy loss",
  245. priv_loss_nnaa_value = "NNAA privacy loss",
  246. eps_identif_risk_value = "Epsilon identifiability",
  247. hit_rate_value = "Hitting rate",
  248. mia_recall_value = "MIA",
  249. att_discl_risk_value = "ADR"
  250. )
  251. priv_df_adni <- priv_df_adni %>% mutate(eps_identif_risk_value = eps_identif_risk_value*100,
  252. hit_rate_value = hit_rate_value*100)
  253. priv_df_a4 <- priv_df_a4 %>% mutate(eps_identif_risk_value = eps_identif_risk_value*100,
  254. hit_rate_value = hit_rate_value*100)
  255. get_privacy_plot <- function(priv_df, p) {
  256. pp <- ggplot(data = priv_df, aes(x = Method, y = !!sym(p), fill = Method)) +
  257. labs(title = priv_titels[p],
  258. x = "",
  259. y = "") +
  260. scale_fill_manual(values = colormap) +
  261. theme_classic() +
  262. theme(axis.text.x = element_text(angle = 30, vjust = 1, hjust=1, size = 8))
  263. if(p %in% c("eps_identif_risk_value", "hit_rate_value")) {
  264. yval = 9.0
  265. pp <- pp + geom_hline(yintercept = yval, linetype = "dashed", color = "darkred") +
  266. geom_rect(
  267. xmin = -Inf, xmax = Inf,
  268. ymin = yval, ymax = Inf,
  269. fill = "grey80",
  270. alpha = 0.5
  271. )
  272. } else if(p %in% c("priv_loss_nndr_value", "priv_loss_nnaa_value")) {
  273. mu = mean(priv_df[[p]], na.rm=T)
  274. #m = length(unique(priv_df$Method))
  275. #n = nrow(priv_df)/m
  276. #se = sd(priv_df[[p]], na.rm=T)/sqrt(n)
  277. #t_crit <- qt(1 - 0.01 / (2 * m), df = round(n)-1)
  278. form <- as.formula(paste(p, "~ 1 + (1 | Method)"))
  279. fit <- lmer(form, data = priv_df)
  280. cis <- confint(fit, level = 0.99, parm = "(Intercept)")
  281. upper = cis[2]
  282. lower = cis[1]
  283. pp <- pp +
  284. geom_hline(yintercept = mu, linetype = "solid", color = "lightgray") +
  285. geom_hline(yintercept = upper, linetype = "dashed", color = "darkred", alpha=0.5) +
  286. geom_hline(yintercept = lower, linetype = "dashed", color = "darkred", alpha=0.5) +
  287. geom_rect(xmin = -Inf, xmax = Inf, ymin = upper, ymax = Inf, fill = "grey80", alpha = 0.2) +
  288. geom_rect(xmin = -Inf, xmax = Inf, ymin = -Inf, ymax = lower, fill = "grey80", alpha = 0.2)
  289. } else if(p %in% c("mia_recall_value", "att_discl_risk_value")) {
  290. yval = 0.5
  291. pp <- pp + geom_hline(yintercept = yval, linetype = "dashed", color = "darkred") +
  292. geom_rect(
  293. xmin = -Inf, xmax = Inf,
  294. ymin = yval, ymax = Inf,
  295. fill = "grey80",
  296. alpha = 0.5
  297. )
  298. }
  299. pp <- pp + geom_hline(yintercept = 0.0, linetype = "dotted", color = "black") +
  300. geom_boxplot()
  301. return(pp)
  302. }
  303. pPlotlist_adni <- list()
  304. pPlotlist_a4 <- list()
  305. for (p in privs) {
  306. pp_adni <- get_privacy_plot(priv_df_adni, p)
  307. pp_a4 <- get_privacy_plot(priv_df_a4, p)
  308. pPlotlist_adni[[paste0(p, "_adni")]] <- pp_adni
  309. pPlotlist_a4[[paste0(p, "_a4")]] <- pp_a4
  310. }
  311. priv_grid_adni <- ggarrange(plotlist = pPlotlist_adni, ncol = 2, nrow = 3, common.legend = TRUE, legend = "right")
  312. plot(priv_grid_adni)
  313. priv_grid_a4 <- ggarrange(plotlist = pPlotlist_a4, ncol = 2, nrow = 3, common.legend = TRUE, legend = "right")
  314. plot(priv_grid_a4)
  315. ggsave("~/R/DDLS-plots/priv_grid_adni.tiff", priv_grid_adni, width = 10, height = 10, units = "in", dpi = 500)
  316. ggsave("~/R/DDLS-plots/priv_grid_a4.tiff", priv_grid_a4, width = 10, height = 10, units = "in", dpi = 500)
  317. #ggsave("~/R/DDLS-plots/util_grid.png", util_grid, width = 10, height = 10, units = "in", dpi = 500)
  318. #ggsave("~/R/DDLS-plots/priv_grid.png", priv_grid, width = 10, height = 10, units = "in", dpi = 500)
  319. utils <- c("mutual_inf_diff_value", "ks_tvd_stat_value", "frac_ks_sigs_value", "avg_F1_diff_value", "avg_F1_diff_hout_value", "nnaa_value")
  320. # Mapping column names
  321. util_titles <- c(
  322. mutual_inf_diff_value = "MI diff",
  323. ks_tvd_stat_value = "KS/TVD stat",
  324. frac_ks_sigs_value = "Frac. KS sigs",
  325. avg_F1_diff_value = "Avg. Diff. F1",
  326. avg_F1_diff_hout_value = "Avg. Diff. F1 (test)",
  327. nnaa_value = "NNAA value"
  328. )
  329. get_util_plot <- function(util_df, u) {
  330. up <- ggplot(data = util_df, aes(x = Method, y = !!sym(u), fill = Method)) +
  331. labs(title = util_titles[u],
  332. x = "",
  333. y = "") +
  334. scale_fill_manual(values = colormap) +
  335. theme_classic() +
  336. theme(axis.text.x = element_text(angle = 30, vjust = 1, hjust=1, size = 8))
  337. if (u == 'nnaa_value') {
  338. mu = mean(util_df[[u]], na.rm=T)
  339. #m = length(unique(priv_df$Method))
  340. #n = nrow(priv_df)/m
  341. #se = sd(priv_df[[p]], na.rm=T)/sqrt(n)
  342. #t_crit <- qt(1 - 0.01 / (2 * m), df = round(n)-1)
  343. form <- as.formula(paste(u, "~ 1 + (1 | Method)"))
  344. fit <- lmer(form, data = util_df)
  345. cis <- confint(fit, level = 0.99, parm = "(Intercept)")
  346. upper = cis[2]
  347. lower = cis[1]
  348. up <- up +
  349. geom_hline(yintercept = mu, linetype = "solid", color = "lightgray") +
  350. geom_hline(yintercept = upper, linetype = "dashed", color = "darkred", alpha=0.5) +
  351. geom_hline(yintercept = lower, linetype = "dashed", color = "darkred", alpha=0.5) +
  352. geom_rect(xmin = -Inf, xmax = Inf, ymin = upper, ymax = Inf, fill = "grey80", alpha = 0.2) +
  353. geom_rect(xmin = -Inf, xmax = Inf, ymin = -Inf, ymax = lower, fill = "grey80", alpha = 0.2)
  354. } else if (u == "frac_ks_sigs_value"){
  355. up <- up + geom_hline(yintercept = 1.0, linetype = "solid", color = "gray")
  356. }
  357. up <- up + geom_boxplot()
  358. return(up)
  359. }
  360. uPlotlist_adni <- list()
  361. uPlotlist_a4 <- list()
  362. for (u in utils) {
  363. uPlotlist_adni[[u]] <- get_util_plot(util_df_adni, u)
  364. uPlotlist_a4[[u]] <- get_util_plot(util_df_a4, u)
  365. }
  366. util_grid_adni <- ggarrange(plotlist = uPlotlist_adni, ncol = 2, nrow = 3, common.legend = TRUE, legend = "right")
  367. plot(util_grid_adni)
  368. util_grid_a4 <- ggarrange(plotlist = uPlotlist_a4, ncol = 2, nrow = 3, common.legend = TRUE, legend = "right")
  369. plot(util_grid_a4)
  370. ggsave("~/R/DDLS-plots/util_grid_adni.tiff", util_grid_adni, width = 10, height = 10, units = "in", dpi = 500)
  371. ggsave("~/R/DDLS-plots/util_grid_a4.tiff", util_grid_a4, width = 10, height = 10, units = "in", dpi = 500)
  372. ##### Get specific values #####
  373. summary_stats_adni <- up_df_adni %>%
  374. group_by(Method) %>%
  375. dplyr::summarize(
  376. mean_u = mean(util_score),
  377. mean_p = mean(priv_score),
  378. sd_u = sd(util_score),
  379. sd_p = sd(priv_score),
  380. num = n(),
  381. ci_margin_u = qt(0.975, df = n() - 1) * sd(util_score) / sqrt(n()),
  382. ci_margin_p = qt(0.975, df = n() - 1) * sd(priv_score) / sqrt(n()),
  383. ci_lwr_u = mean_u - ci_margin_u,
  384. ci_upr_u = mean_u + ci_margin_u,
  385. ci_lwr_p = mean_p - ci_margin_p,
  386. ci_upr_p = mean_p + ci_margin_p
  387. )
  388. summary_stats_a4 <- up_df_a4 %>%
  389. group_by(Method) %>%
  390. dplyr::summarize(
  391. mean_u = mean(util_score),
  392. mean_p = mean(priv_score),
  393. ci_margin_u = qt(0.975, df = n() - 1) * sd(util_score) / sqrt(n()),
  394. ci_margin_p = qt(0.975, df = n() - 1) * sd(priv_score) / sqrt(n()),
  395. ci_lwr_u = mean_u - ci_margin_u,
  396. ci_upr_u = mean_u + ci_margin_u,
  397. ci_lwr_p = mean_p - ci_margin_p,
  398. ci_upr_p = mean_p + ci_margin_p
  399. )
  400. priv_df_adni %>% group_by(dataset) %>%
  401. dplyr::summarise(Mean = mean(eps_identif_risk_value),
  402. SD = sd(eps_identif_risk_value),
  403. N = n(),
  404. margin = qt(0.975, df = n() - 1) * sd(eps_identif_risk_value) / sqrt(n()),
  405. lower = Mean - margin,
  406. upper = Mean + margin) %>% filter(upper > 0.09)
  407. priv_df_a4 %>% group_by(dataset) %>%
  408. dplyr::summarise(Mean = mean(eps_identif_risk_value),
  409. SD = sd(eps_identif_risk_value),
  410. N = n(),
  411. margin = qt(0.975, df = n() - 1) * sd(eps_identif_risk_value) / sqrt(n()),
  412. lower = Mean - margin,
  413. upper = Mean + margin) %>% filter(upper > 0.09)
  414. priv_df_adni %>% group_by(dataset) %>%
  415. dplyr::summarise(Mean = mean(att_discl_risk_value),
  416. SD = sd(att_discl_risk_value),
  417. N = n(),
  418. margin = qt(0.975, df = n() - 1) * sd(att_discl_risk_value) / sqrt(n()),
  419. lower = Mean - margin,
  420. upper = Mean + margin) %>% filter(upper > 0.1)
  421. priv_df_a4 %>% group_by(dataset) %>%
  422. dplyr::summarise(Mean = mean(att_discl_risk_value),
  423. SD = sd(att_discl_risk_value),
  424. N = n(),
  425. margin = qt(0.975, df = n() - 1) * sd(att_discl_risk_value) / sqrt(n()),
  426. lower = Mean - margin,
  427. upper = Mean + margin) %>% filter(upper > 0.1)
  428. priv_df_adni %>% group_by(dataset) %>%
  429. dplyr::summarise(Mean = mean(mia_recall_value),
  430. SD = sd(mia_recall_value),
  431. N = n(),
  432. margin = qt(0.975, df = n() - 1) * sd(mia_recall_value) / sqrt(n()),
  433. lower = Mean - margin,
  434. upper = Mean + margin) %>% filter(upper > 0.1)
  435. priv_df_a4 %>% group_by(dataset) %>%
  436. dplyr::summarise(Mean = mean(mia_recall_value),
  437. SD = sd(mia_recall_value),
  438. N = n(),
  439. margin = qt(0.975, df = n() - 1) * sd(mia_recall_value) / sqrt(n()),
  440. lower = Mean - margin,
  441. upper = Mean + margin) %>% filter(upper > 0.1)
  442. # Bonus experiment
  443. #### ADNI+ ####
  444. epsilons <- c(100, NA)
  445. samples <- c(100, 18)
  446. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/adni_plus/bmk_new_deg2_eps", ifelse(is.na(epsilons), "zero", epsilons), ".csv")
  447. # Read CSVs in a loop
  448. res_list <- list()
  449. for (i in seq_along(epsilons)) {
  450. df <- read.csv(file_paths[i])
  451. df$Epsilon <- epsilons[i]
  452. df$samples <- samples[i]
  453. res_list[[i]] <- df
  454. }
  455. # Combine all data frames
  456. bmk_ds <- bind_rows(res_list)
  457. util_cols <- c("mutual_inf_diff_value", "ks_tvd_stat_value",
  458. "frac_ks_sigs_value","avg_F1_diff_value", "avg_F1_diff_hout_value",
  459. "nnaa_value")
  460. priv_cols <- c("priv_loss_nndr_value", "priv_loss_nnaa_value","hit_rate_value",
  461. "avg_nndr_value", "eps_identif_risk_value", "priv_loss_eps_value",
  462. "mia_recall_value", "att_discl_risk_value")
  463. cols <- c(util_cols, priv_cols)
  464. bmk_ds$Epsilon <- bmk_ds %>%
  465. select(Epsilon) %>%
  466. mutate_all(~replace(., is.na(.), 0)) %>%
  467. mutate(Epsilon = factor(Epsilon, levels = as.character(unique(Epsilon)))) %>%
  468. pull(Epsilon)
  469. # CTGAN
  470. epochs <- c(750)
  471. setting <- c("default")
  472. file_paths <- paste0("~/Python/WASP-DDLS/SE-benchmark/adni_plus/bmk_new_ctgan_", setting, "_epochs_", epochs, ".csv")
  473. #file_paths <- c(file_paths, "~/Python/WASP-DDLS/SE-benchmark/bmk_ctgan_epochs_100.csv")
  474. # Read CSVs in a loop
  475. res_list <- list()
  476. for (i in seq_along(file_paths)) {
  477. df <- read.csv(file_paths[i])
  478. df$Epochs <- epochs[i]
  479. df$samples <- 100
  480. res_list[[i]] <- df
  481. }
  482. # Combine all data frames
  483. bmk_ctgan <- bind_rows(res_list)
  484. bmk_ctgan$Epochs <- as.factor(bmk_ctgan$Epochs)
  485. # Synthpop
  486. bmk_synthpop <- read.csv("~/Python/WASP-DDLS/SE-benchmark/adni_plus/bmk_new_synthpop.csv") %>% mutate(
  487. Epochs = 0,
  488. samples = 100
  489. )
  490. # TabPFN
  491. bmk_tabpfn <- read.csv("~/Python/WASP-DDLS/SE-benchmark/adni_plus/bmk_new_tabpfn_t_1.0.csv")
  492. label_list = c("DS (100)", "DS w/o DP", "CTGAN (default)", "Synthpop", "TabPFN (t=1.0)")
  493. util_df_adni_plus <- rbind(select(bmk_ds, dataset, model, all_of(util_cols)),
  494. select(bmk_ctgan, dataset, model, all_of(util_cols)),
  495. select(bmk_synthpop, dataset, model, all_of(util_cols)),
  496. select(bmk_tabpfn, dataset, model, all_of(util_cols))
  497. ) |>
  498. dplyr::mutate(avg_F1_diff_value = abs(avg_F1_diff_value),
  499. avg_F1_diff_hout_value = abs(avg_F1_diff_hout_value)) |>
  500. dplyr::mutate(mutual_inf_diff = 1 - tanh(mutual_inf_diff_value),
  501. ks_tvd_stat = 1 - ks_tvd_stat_value,
  502. frac_ks_sigs = 1 - frac_ks_sigs_value,
  503. avg_F1_diff = 1 - abs(avg_F1_diff_value),
  504. avg_F1_diff_hout = 1 - abs(avg_F1_diff_hout_value),
  505. nnaa = 1 - nnaa_value)
  506. util_df_adni_plus$util_score <- util_df_adni_plus |> dplyr::select(c(mutual_inf_diff, ks_tvd_stat, frac_ks_sigs, avg_F1_diff, avg_F1_diff_hout, nnaa)) |> rowMeans()
  507. util_df_adni_plus$Method <- factor(util_df_adni_plus$dataset, levels = unique(util_df_adni_plus$dataset),
  508. labels=label_list)
  509. priv_df_adni_plus <- rbind(select(bmk_ds, dataset, model, all_of(priv_cols)),
  510. select(bmk_ctgan, dataset, model, all_of(priv_cols)),
  511. select(bmk_synthpop, dataset, model, all_of(priv_cols)),
  512. select(bmk_tabpfn, dataset, model, all_of(priv_cols))
  513. ) |>
  514. dplyr::mutate(priv_loss_nndr = 1 - abs(priv_loss_nndr_value),
  515. priv_loss_nnaa = 1 - abs(priv_loss_nnaa_value),
  516. priv_loss_eps = 1 - abs(priv_loss_eps_value),
  517. hit_rate = 1 - hit_rate_value,
  518. eps_identif_risk = 1 - eps_identif_risk_value,
  519. mia_recall = 1 - mia_recall_value,
  520. att_discl_risk = 1 - att_discl_risk_value)
  521. priv_df_adni_plus$priv_score <- priv_df_adni_plus |> dplyr::select(c(priv_loss_nndr, priv_loss_nnaa, priv_loss_eps, hit_rate, eps_identif_risk, att_discl_risk)) |> rowMeans()
  522. priv_df_adni_plus$Method <- factor(priv_df_adni_plus$dataset, levels = unique(priv_df_adni_plus$dataset),
  523. labels=label_list)
  524. up_df_adni_plus <- util_df_adni_plus |> dplyr::select(c(Method, model, util_score)) |>
  525. dplyr::inner_join(priv_df_adni_plus |> dplyr::select(c(Method, model, priv_score)), by = c("Method", "model"))
  526. colormap <- c(
  527. paletteer_d("rcartocolor::BluYl")[1:2],
  528. paletteer_d("beyonce::X58")[3],
  529. "#663399",
  530. paletteer_d("ggsci::cyan_material")[c(1,3,5,7)]
  531. )
  532. up_plot_adni_plus <- ggplot(data = up_df_adni_plus, aes(x=util_score, y=priv_score, fill=Method)) +
  533. labs(title = "ADNI+", x = "Utility", y = "Privacy") +
  534. #scale_fill_paletteer_d("rcartocolor::BluYl") +
  535. scale_fill_manual(values = colormap) +
  536. theme_minimal()
  537. #ylim(0.7,1.0) + xlim(0.3,0.85)
  538. up_df_adni_ <- up_df_adni %>% filter(Method %in% up_df_adni_plus$Method)
  539. up_plot_adni_plus <- up_plot_adni_plus +
  540. geom_point(pch=21, color="gray", size = 3, data=up_df_adni_, alpha=0.75) +
  541. geom_point(pch=21, color = "black", size=3)
  542. plot(up_plot_adni_plus)
  543. up_plot_adni_plus <- up_plot_adni_plus + theme(plot.margin = margin(5.5, 125, 5.5, 125))
  544. up_plot_all <- ggarrange(up_plot_adni, up_plot_a4, ncol=2, common.legend = TRUE, legend = "right")
  545. up_plot_all <- ggarrange(up_plot_all, up_plot_adni_plus, nrow=2, common.legend = FALSE)
  546. plot(up_plot_all)
  547. ##### Saving and stuff ####
  548. ggsave("~/R/DDLS-plots/up_plot_all.tiff", up_plot_all, width = 10, height = 9, units = "in", dpi = 500)
  549. summary_stats_adni_plus <- up_df_adni_plus %>%
  550. group_by(Method) %>%
  551. dplyr::summarize(
  552. mean_u = mean(util_score),
  553. mean_p = mean(priv_score),
  554. sd_u = sd(util_score),
  555. sd_p = sd(priv_score),
  556. num = n(),
  557. ci_margin_u = qt(0.975, df = n() - 1) * sd(util_score) / sqrt(n()),
  558. ci_margin_p = qt(0.975, df = n() - 1) * sd(priv_score) / sqrt(n()),
  559. ci_lwr_u = mean_u - ci_margin_u,
  560. ci_upr_u = mean_u + ci_margin_u,
  561. ci_lwr_p = mean_p - ci_margin_p,
  562. ci_upr_p = mean_p + ci_margin_p
  563. )
  564. res <- summary_stats_adni %>% inner_join(summary_stats_adni_plus, by='Method', suffix = c('.adni', '.adni_plus')) %>%
  565. mutate(ci.bound_u = (mean_u.adni - mean_u.adni_plus) - 1.96 * sqrt((sd_u.adni^2/num.adni) + (sd_u.adni^2/num.adni)),
  566. ci.bound_p = (mean_p.adni - mean_p.adni_plus) - 1.96 * sqrt((sd_p.adni^2/num.adni) + (sd_p.adni^2/num.adni)))
  567. res <- data.frame()
  568. for(method in unique(up_df_adni_plus$Method)) {
  569. a <- up_df_adni_plus %>% filter(Method == method)
  570. b <- up_df_adni %>% filter(Method == method)
  571. t_u <- t.test(a$util_score, b$util_score, alternative = "two.sided", var.equal=F)
  572. t_p <- t.test(a$priv_score, b$priv_score, alternative = "two.sided", var.equal=F)
  573. res <- rbind(res, data.frame(Method = method,
  574. #t_stat_u = t_u$statistic,
  575. p_value_u = t_u$p.value * nrow(a),
  576. mean_u_x = t_u$estimate[1],
  577. mean_u_y = t_u$estimate[2],
  578. #t_stat_p = t_p$statistic,
  579. p_value_p = t_p$p.value * nrow(a),
  580. mean_p_x = t_p$estimate[1],
  581. mean_p_y = t_p$estimate[2]
  582. ))
  583. }

benchmark.plots.R at commit 299d5c6, under MIT · at the source

Overview

Authors: Filip Winzell1, Ida Arvidsson1, Niels Christian Overgaard1, Anders Heyden1, Kalle Åström1, Linda Karlsson2, Jacob W Vogel3, Oskar Hansson4, Niklas Mattsson‐Carlgren3,4,5, for the Alzheimer's Disease Neuroimaging Initiative
  1. Centre for Mathematical Sciences, Lund University, Lund, Sweden
  2. Clinical Memory Research Unit, Department of Clinical Sciences in Malmö, Lund University, Lund, Sweden
  3. Department of Clinical Sciences Malmö, SciLifeLab, Lund University, Lund, Sweden
  4. Wallenberg Center for Molecular Medicine, Lund University, Lund, Sweden
  5. Memory Clinic, Skåne University Hospital, Malmö, Sweden
Journal: Alzheimer's & dementia (Amsterdam, Netherlands), volume 18, issue 3, article e70430
Dates: received 20 October 2025; accepted 18 June 2026; published online 10 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/dad2.70430 · PMID 42582267 · PMCID PMC13457355 · OpenAlex W7202110940
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), Alzheimer's / dementia (population), methods / tools (subfield)
Methods: Machine learning, Statistics, Connectivity
Keywords: Alzheimer's disease, machine learning, synthetic data, tabular data
Topic: Privacy-Preserving Technologies in Data (Artificial Intelligence, Computer Science), according to OpenAlex
Funding: European Research Council (101108819); NIA NIH HHS (R01 AG083740, U01 AG024904)
Citations: not cited yet (Europe PMC); 28 references in the paper

Abstract

INTRODUCTION: The scarcity of large, clinically relevant cohorts is becoming a bottleneck in Alzheimer's disease (AD) research, as their sensitive nature makes open data sharing difficult. Privacy‐preserving synthetic datasets generated with machine learning may help address this challenge.

METHODS: We compared five frameworks for generating synthetic tabular data from the Alzheimer's Disease Neuroimaging Initiative and Anti‐Amyloid Treatment in Asymptomatic Alzheimer's Disease cohorts, with a set of empirical privacy and utility metrics. Two of the methods, DataSynthesizer and TableDiffusion, provide ε‐differential privacy guarantees.

RESULTS: Methods with differential privacy achieved high privacy ratings but low levels of utility. Deep learning methods like Tabular Prior‐data Fitted Network (TabPFN) and Conditional Generative Adversarial Network (CTGAN) also showed high privacy with limited utility. In contrast, non‐private DataSynthesizer and Synthpop offered higher utility at a cost of lower privacy.

DISCUSSION: The evaluated methods demonstrated a clear trade‐off between privacy and utility. High privacy was generally associated with insufficient utility, highlighting the need for further research into synthetic data generation for AD.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

fwinzell/synthetic_ad

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 299d5c63736fcf624ab4aa5a67f43723c116be7d, 12 February 2026
Languages: R (15), Python (15)
Size: 48 files, 30 scripts
Software Heritage: not archived
Found in: the text, “INTRODUCTION”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: pandas (14 files), ggplot2 (13 files), NumPy (13 files), tidyverse (13 files), ggpubr (12 files), Matplotlib (8 files), scikit-learn (5 files), NetworkX (2 files), seaborn (2 files), LightGBM (1 file), lme4 (1 file), Plotly (1 file), PyTorch (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
32 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 30 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 4 keywords, 2 funders, 26 references.

Cite

This paper

Winzell, F., Arvidsson, I., Overgaard, N. C., Heyden, A., Åström, K., Karlsson, L., Vogel, J. W., Hansson, O., Mattsson‐Carlgren, N., & for the Alzheimer's Disease Neuroimaging Initiative. (2026). Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research. Alzheimer's & dementia (Amsterdam, Netherlands), 18(3), e70430. https://doi.org/10.1002/dad2.70430

BibTeX

@article{winzell2026benchmarking,
author = {Winzell, Filip and Arvidsson, Ida and Overgaard, Niels Christian and Heyden, Anders and Åström, Kalle and Karlsson, Linda and Vogel, Jacob W and Hansson, Oskar and Mattsson‐Carlgren, Niklas and {for the Alzheimer's Disease Neuroimaging Initiative}},
title = {{Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research}},
journal = {Alzheimer's \& dementia (Amsterdam, Netherlands)},
year = {2026},
month = jul,
volume = {18},
number = {3},
pages = {e70430},
publisher = {Wiley},
issn = {2352-8729},
doi = {10.1002/dad2.70430},
url = {https://doi.org/10.1002/dad2.70430},
pmid = {42582267},
pmcid = {PMC13457355}
}

RIS

TY - JOUR
AU - Winzell, Filip
AU - Arvidsson, Ida
AU - Overgaard, Niels Christian
AU - Heyden, Anders
AU - Åström, Kalle
AU - Karlsson, Linda
AU - Vogel, Jacob W
AU - Hansson, Oskar
AU - Mattsson‐Carlgren, Niklas
AU - for the Alzheimer's Disease Neuroimaging Initiative
TI - Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research
T2 - Alzheimer's & dementia (Amsterdam, Netherlands)
J2 - Alzheimers Dement (Amst)
PY - 2026
DA - 2026/07/01
VL - 18
IS - 3
SP - e70430
SN - 2352-8729
PB - Wiley
DO - 10.1002/dad2.70430
UR - https://doi.org/10.1002/dad2.70430
LA - en
ER -

CSL-JSON

{
"id": "10.1002/dad2.70430",
"type": "article-journal",
"title": "Benchmarking privacy and utility in synthetic tabular cohorts for Alzheimer's disease research",
"container-title": "Alzheimer's & dementia (Amsterdam, Netherlands)",
"author": [
{
"family": "Winzell",
"given": "Filip"
},
{
"family": "Arvidsson",
"given": "Ida"
},
{
"family": "Overgaard",
"given": "Niels Christian"
},
{
"family": "Heyden",
"given": "Anders"
},
{
"family": "Åström",
"given": "Kalle"
},
{
"family": "Karlsson",
"given": "Linda"
},
{
"family": "Vogel",
"given": "Jacob W"
},
{
"family": "Hansson",
"given": "Oskar"
},
{
"family": "Mattsson‐Carlgren",
"given": "Niklas"
},
{
"literal": "for the Alzheimer's Disease Neuroimaging Initiative"
}
],
"container-title-short": "Alzheimers Dement (Amst)",
"volume": "18",
"issue": "3",
"page": "e70430",
"DOI": "10.1002/dad2.70430",
"PMID": "42582267",
"PMCID": "PMC13457355",
"ISSN": "2352-8729",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/dad2.70430",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: NetworkX, Plotly, lme4, 10 other tools
[2] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: LightGBM, NetworkX, Plotly, 9 other tools
[3] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: NetworkX, Plotly, lme4, 9 other tools
[4] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: NetworkX, Plotly, lme4, 9 other tools
[5] doi:10.1038/s41586-026-10735-w [code]
Distributed control circuits across a brain-and-cord connectome.
Journal: Nature
In common: NetworkX, Plotly, ggpubr, 9 other tools
[6] doi:10.1038/s41467-026-73996-z [code]
Genetic architecture of white matter microstructure captured by unsupervised deep representation learning of fractional anisotropy maps.
Journal: Nature communications
In common: LightGBM, Plotly, PyTorch, 8 other tools
[7] doi:10.1093/braincomms/fcag176 [code]
Tau topography subtypes account for clinical heterogeneity and longitudinal trajectories in early-onset Alzheimer's disease.
Journal: Brain communications
In common: Plotly, lme4, ggpubr, 8 other tools, Alzheimer's / dementia
[8] doi:10.1038/s41514-026-00443-0 [code]
Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease.
Journal: npj aging
In common: NetworkX, Plotly, ggpubr, 8 other tools, Alzheimer's / dementia
[9] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: NetworkX, Plotly, ggpubr, 8 other tools
[10] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: NetworkX, lme4, ggpubr, 8 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.