OSCR

Personalized machine learning-guided radiation dose escalation in newly diagnosed glioblastoma: prospective pilot study.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches
  1. [1] § Methods › Statistical analysis ↔ statistical_analysis/qc_and_os_analysis.qmd, lines 427–500 · score 0.87 · Cox proportional hazards, propensity score, PPRT intervention, MatchIt, nearest, wildtype
  2. [2] § Results › Overall survival outcomes ↔ statistical_analysis/qc_and_os_analysis.qmd, lines 427–500 · score 0.87 · proportional hazards assumption, Cox proportional hazards, Schoenfeld residuals, underwent PPRT intervention, 0.17–0.69, PH
  3. [3] § Results › Progression-free survival ↔ statistical_analysis/pfs_analysis.qmd, lines 202–269 · score 0.87 · Kaplan Meier, Cox proportional hazards, slower progression, log rank, robust standard errors, 0.13–0.61
  4. [4] § Methods › Propensity matched historical control group ↔ statistical_analysis/qc_and_os_analysis.qmd, lines 238–295 · score 0.70 · greedy nearest neighbor, MGMTp methylation status, propensity score, age, sex, variables
  5. [5] § Methods › Study design ↔ statistical_analysis/pfs_analysis.qmd, lines 53–118 · score 0.54 · progression free survival, disease progression, PFS, surgery, events, Status

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Quarto · 502 lines · 21 KB · no license · 3 matches

  1. ---
  2. title: "PPRT Survival Analysis"
  3. description: "To evaluate if there is a significant difference in **overall** survival between the PPRT intervention and control group"
  4. date: last-modified
  5. date-format: "[Last Updated:] MMM DD, YYYY"
  6. format:
  7. html:
  8. self-contained: true
  9. title-block-banner: true
  10. title-block-banner-color: "#FFFFFF"
  11. toc: true
  12. toc-location: left
  13. number-sections: true
  14. code-fold: true
  15. code-tools:
  16. source: true #hide the source option
  17. toggle: true
  18. code-copy: true #option to copy code
  19. code-overflow: scroll
  20. df-print: kable
  21. code-block-bg: true
  22. code-block-border-left: "#31BAE9"
  23. editor: source #open in source editor by default
  24. theme: flatly
  25. echo: false
  26. execute:
  27. warning: false
  28. message: false
  29. ---
  30. # ------------------------------------------------------------------
  31. # NOTE: Patient identifiers have been removed for privacy.
  32. # Any identifiers appearing in the original analysis have been
  33. # replaced with "XXXX" in this shared version of the script.
  34. # This does not affect analytic results.
  35. # ------------------------------------------------------------------
  36. ```{r load_packages}
  37. #| label: load-packages
  38. #| include: false
  39. # Remove environment to avoid potential errors or conflicts
  40. rm(list = ls())
  41. library(tidyverse) #data wrangling, viz
  42. library(readxl) #reading excel sheets
  43. library(gtsummary) #for creating table 1
  44. library(DT) # for datatables
  45. library(survival) #for survival analysis
  46. library(survminer) #for visualizing kaplan-meier plots
  47. library(tab) #for generating tables for statistical reports
  48. library(lubridate) #date manipulation
  49. library(MatchIt) #for PSM matching
  50. library(lmtest)
  51. library(sandwich)
  52. library(writexl)
  53. library(geepack) #for GEE
  54. ```
  55. ```{r intervention}
  56. #updated survival data as of 03/20/25
  57. intervention <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/ACC_Trial_11_20_2024_To_Amy.xlsx", sheet = 2) %>%
  58. janitor::clean_names() %>%
  59. mutate(survival_all = if_else(alive == "n", survival, survival_till_3_20_2025)) %>%
  60. mutate(survival_all = survival) %>%
  61. mutate(year = factor(year(dos))) %>%
  62. filter(is.na(remove) | remove != "Optune device") %>% #remove optune devices (tumor treating fields)
  63. select(cbicaid_preop, age, sex, alive, survival_all, mgmt, year) %>%
  64. rename(survival = survival_all) %>%
  65. mutate(group = "intervention") %>%
  66. #mutate(baseline_id = "000")
  67. mutate(baseline_id = cbicaid_preop) %>%
  68. select(age:baseline_id)
  69. ```
  70. ```{r selected_controls}
  71. ## control subjects that were QCed previously who meet the criterion
  72. sel_controls <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/20240812_ACC_clinical.xlsx", sheet = 2) %>%
  73. janitor::clean_names() %>%
  74. filter(is.na(remove)) %>% #Remove the patients that did not meet the inclusion criteria
  75. pull(baseline_id) %>%
  76. unique()
  77. sept_alive <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/potential_controls_9_9_2024_To Amy.xlsx", sheet = 2) %>%
  78. janitor::clean_names() %>%
  79. filter(is.na(remove)) %>% #N=23 patients did not meet the inclusion criteria
  80. pull(baseline_id) %>%
  81. unique()
  82. selected_controls <- c(sel_controls, sept_alive)
  83. #remove 8 patients from the list of selected controls that does not meet inclusion criteria based on quality cont
  84. remove_from_ctrl <- c("XXXX", "XXXX",
  85. "XXXX", "XXXX",
  86. "XXXX", "XXXX",
  87. "XXXX", "XXXX")
  88. selected_controls <- selected_controls[!selected_controls %in% remove_from_ctrl]
  89. #remove interim objects to avoid workplace confusion
  90. rm(sel_controls, sept_alive)
  91. ```
  92. ## Identify all the control subjects that should be removed (i.e., does not meet criterion)
  93. ```{r remove}
  94. ## Iteration 1; the n=5 patients that we needed to remove
  95. ctr_rm <- c("XXXX",
  96. "XXXX",
  97. "XXXX",
  98. "XXXX",
  99. "XXXX")
  100. ## Iteration 2; n=15 subjects that did not meet the inclusion criteria
  101. need_to_remove <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/Control_Group_17_June2024_REMOVE.xlsx", sheet = "sheet2") %>%
  102. janitor::clean_names() %>%
  103. filter(remove == "Yes") %>% #N=15 patients did not meet the inclusion criteria
  104. pull(baseline_id)
  105. ## Iteration 3; additional n=8 subjects that we need to remove
  106. need_to_remove_july <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/07022024_ACC_clinical_REMOVE.xlsx", sheet = "sheet2") %>%
  107. janitor::clean_names() %>%
  108. filter(remove == "Yes") %>%
  109. pull(baseline_id)
  110. ## Iteration 4;
  111. need_to_remove_august <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/20240812_ACC_clinical.xlsx", sheet = 2) %>%
  112. janitor::clean_names() %>%
  113. filter(remove == "yes") %>%
  114. pull(baseline_id)
  115. ## Iteration 5
  116. need_to_remove_september <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/potential_controls_9_9_2024_To Amy.xlsx", sheet = 2) %>%
  117. janitor::clean_names() %>%
  118. filter(remove == "yes") %>% #N=23 patients did not meet the inclusion criteria
  119. pull(baseline_id)
  120. # `remove` contains a list of ids that should not be used as potential controls since criteria not meet
  121. remove <- c(ctr_rm, need_to_remove, need_to_remove_july,
  122. need_to_remove_august, need_to_remove_september,
  123. remove_from_ctrl) %>% unique()
  124. rm(ctr_rm, need_to_remove, need_to_remove_july, need_to_remove_august, need_to_remove_september)
  125. ```
  126. ## Qualified controls after QC
  127. ```{r cntrls_combined}
  128. ## This is the old list of "qualified controls" that meet the inclusion criteria
  129. # idh is wildtype; eor is GTR, survival > 90 days, methylated or unmethylated group
  130. qualified_cntr_old <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collaborator_projects/Hamed_ACC/UPenn.xlsx") %>%
  131. janitor::clean_names() %>%
  132. mutate(across(baseline_id:alive, ~str_sub(., 2, -2))) %>%
  133. mutate(across(rec:kps, ~str_sub(., 2, -2))) %>%
  134. filter(idh == "Wildtype" & eor == "GTR") %>%
  135. filter(survival >= 90) %>% #exclude patients who passed away before completing radiation therapy
  136. mutate(group = "control") %>% #create a control variable
  137. mutate(year = factor(str_sub(baseline_id, 6,9))) %>%
  138. filter(mgmt %in% c("Methylated", "Unmethylated")) %>%
  139. mutate(age = as.numeric(age)) %>%
  140. select(age, sex, alive, survival, mgmt, year, group, baseline_id)
  141. #### CURRENT CONTROL###
  142. # note: "Positive" ="Methylated"; "Not Detected" = "Unmethylated"
  143. # Control group
  144. # exclude patients who passed away before completing radiation therapy (approximately 90 days)
  145. # list of qualified control patients that are DECEASED
  146. control_new_deceased <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/potential_control_patients.xlsx", sheet = 1) %>%
  147. janitor::clean_names() %>%
  148. filter(deceased == "yes") %>% #only look at deceased patients for now
  149. filter(idh == "Wildtype" & eor == 3) %>% #EOR of 3 is GTR
  150. mutate(mgmt = if_else(mgmt %in% c("positive", "Positive"),
  151. "Methylated", "Unmethylated")) %>%
  152. filter(survival >= 90) %>%
  153. mutate(group = "control") %>% #create indicator variable
  154. mutate(year = factor(str_sub(baseline_id, 6,9))) %>%
  155. mutate(alive = if_else(deceased == "yes", "n", "y")) %>%
  156. select(age, sex, alive, survival, mgmt, year, group, baseline_id)
  157. # list of control patients that were still ALIVE
  158. control_new_alive <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/potential_control_patients.xlsx",
  159. sheet = 1) %>%
  160. janitor::clean_names() %>%
  161. filter(is.na(deceased)) %>% #look at alive patients
  162. filter(idh == "Wildtype" & eor == 3) %>% #EOR of 3 is GTR
  163. mutate(mgmt = if_else(mgmt %in% c("positive", "Positive"),
  164. "Methylated", "Unmethylated")) %>%
  165. #filter(mgmt == "Methylated") %>% ## ONLY look at the methylated patients
  166. mutate(group = "control") %>% #create indicator variable
  167. mutate(year = factor(str_sub(baseline_id, 6,9))) %>%
  168. # mutate(alive = if_else(deceased == "yes", "n", "y")) %>%
  169. mutate(alive = "y") %>%
  170. select(age, sex, alive, survival, mgmt, year, group, baseline_id)
  171. ####### combine all the controls from above and remove the unqualified lesions #########
  172. cntrls_combined <- rbind(qualified_cntr_old, control_new_deceased, control_new_alive) %>%
  173. filter(!baseline_id %in% remove)
  174. rm(qualified_cntr_old, control_new_deceased, control_new_alive)
  175. ```
  176. ## Combine intervention and control data (comp_data)
  177. ```{r comp_data, eval=T}
  178. set.seed(123)
  179. # NOTE: qualifed controls contains remaining controls that meet the diagnostic criteria, after subtracting the list of previously QC'ed controls
  180. cntrls_qualified <- cntrls_combined %>%
  181. filter(!baseline_id %in% selected_controls) %>%
  182. slice_sample(prop = 0)
  183. # The final control subjects list inlcude qualified controls from previous iterations and new patients
  184. control <- rbind(cntrls_combined %>% filter(baseline_id %in% selected_controls),
  185. cntrls_qualified) %>%
  186. mutate(survival = case_when(
  187. is.na(survival) & baseline_id == "XXXX" ~ 1070,
  188. is.na(survival) & baseline_id == "XXXX" ~ 650,
  189. is.na(survival) & baseline_id == "XXXX" ~ 530,
  190. is.na(survival) & baseline_id == "XXXX" ~ 647,
  191. TRUE ~ survival
  192. ))
  193. #NOTE: look at `potential_controls_9_9_2024_To Amy.xlsx` sheet3 to see how to obtain the survival #'s above. Briefly, we calculated the # of days between baseline date and the date of last follow-up (the survival status is unknown for these patients).
  194. #remove from workspace to avoid confusion
  195. rm(cntrls_combined, cntrls_qualified)
  196. ```
  197. ```{r}
  198. ######### COMBINE INTERVENTION AND CONTROL DATA to create final data frame #######
  199. comp_data <- rbind(intervention, control) %>%
  200. mutate(status = if_else(alive == "y", 0, 1)) #0=alive; 1=dead
  201. ```
  202. ::: callout-tip
  203. ## Note
  204. Currently we have n=17 patients in the intervention group and we want to match each subject to 4 controls; thus, we should have 17x4=68 control subjects.
  205. :::
  206. ---
  207. ## Check if data is balanced after matching (Table 1)
  208. The propensity score were estimated using logistic regression based on age, sex, and MGMTp methylation status. All subjects in the intervention group were matched to 4 control subjects who did not have PPRT using one-to-four nearest neighbour matching.
  209. In matchit(), setting method = "nearest" performs greedy nearest neighbor matching. A distance is computed between each treated unit and each control unit, and, one by one, each treated unit is assigned a control unit as a match. The matching is "greedy" in the sense that there is no action taken to optimize an overall criterion; each match is selected without considering the other matches that may occur subsequently.
  210. ```{r PSM}
  211. # matchit create treatment and control groups balanced on included covariates
  212. match_obj <- matchit(factor(group) ~ age + sex + mgmt,
  213. data = comp_data,
  214. method = "nearest",
  215. distance ="glm",
  216. ratio = 4, #how many control units should be matched to each treated unit in k:1 matching
  217. replace = F) #control units can only be matched to ONE treated unit each
  218. #check if data is balanced between control and intervention group
  219. check_balance <- summary(match_obj)
  220. check_balance[["sum.matched"]] %>% data.frame() %>% knitr::kable()
  221. matched_data <- get_matches(match_obj) %>%
  222. arrange(subclass) %>%
  223. mutate(sex = ifelse(sex == "M", 1, 0),
  224. mgmt = ifelse(mgmt == "Methylated", 1, 0))
  225. #save the matched data
  226. #mat_data <- match.data(match_obj)
  227. #saveRDS(mat_data, "./matched_data_w_id.rds")
  228. # create table 1
  229. get_matches(match_obj) %>%
  230. select(age, sex, mgmt, group) %>% #select the variables of interest
  231. tbl_summary(by = group,
  232. statistic = list(all_continuous() ~ "{mean}({sd})",
  233. all_categorical() ~ "{n}({p}%)")) %>%
  234. modify_header(label ~ "**Variable**") %>%
  235. modify_spanning_header(c("stat_1", "stat_2") ~ "**Intervention**") %>%
  236. bold_labels() #%>%
  237. #add_p(pvalue_fun = function(x) style_pvalue(x, digits = 2))
  238. ######## EXPORT SOURCE DATA FOR TABLE 1 ########
  239. # tab1_data <- get_matches(match_obj) %>%
  240. # select(subclass, group, age, sex, mgmt)
  241. #
  242. # write_xlsx(tab1_data, "~/Downloads/Tab1_data.xlsx")
  243. ################################################
  244. #check distribution of age
  245. ks_result <- ks.test(matched_data$age[matched_data$group == "control"],
  246. matched_data$age[matched_data$group == "intervention"])
  247. ks_result
  248. ```
  249. ### Assess the matching by fitting GEE models
  250. Use GEE models to account for potential clustering within matched set (assume a exchangeable within-set correlation structure). Report parameter significance using Wald test.
  251. ```{r}
  252. # 1) Is there a difference in age between control and treatment group?
  253. #age is a normally-distributed response
  254. mod <- geeglm(age ~ group,
  255. family = gaussian, #normally distributed response
  256. data = matched_data,
  257. id = subclass,
  258. corstr = "exch") #within-subject observations are equally correlated
  259. summary(mod)
  260. #difference in mean age between intervention and control groups --> large p-value -> there is NO significant difference in age between the groups
  261. # 2) Is there a significant difference in the proportion of males and females between the control and treatment groups?
  262. mod_sex <- geeglm(sex ~ group,
  263. family = binomial, #because the response is binary
  264. data = matched_data,
  265. id = subclass,
  266. corstr = "exch") #within-subject observations are equally correlated
  267. summary(mod_sex)
  268. #3) whether the probability of mgmt = 1 differs between the control and treatment groups, accounting for within-subclass correlations?
  269. mod_mgmt <- geeglm(mgmt ~ group,
  270. family = binomial, #because the response is binary
  271. data = matched_data,
  272. id = subclass,
  273. corstr = "exch") #within-subject observations are equally correlated
  274. # no strong evidence that mgmt differs between groups
  275. summary(mod_mgmt)
  276. ```
  277. ::: callout-important
  278. In total, we have `r nrow(intervention)` patients in the intervention group and `r nrow(intervention)*4` patients in the control group after PSM, for a total of `r nrow(intervention)*5` patients. From the results shown above, we can see that there is good balance between the controls and intervention group in terms of age, sex, and MGMTp methylation status (p>.05). All patients had gross total resection (GTR) and IDH-widtype glioblastoma.
  279. :::
  280. ## Log-rank tests
  281. ```{r, eval=T}
  282. fit <- survfit(Surv(survival, status) ~ group, #0=alive, 1=dead
  283. data = matched_data)
  284. # ordinary log-rank tests
  285. survdiff(Surv(survival, status) ~ group, #0=alive, 1=dead
  286. data = matched_data)
  287. #med survival
  288. median_surv <- data.frame(surv_median(fit)) %>%
  289. select(strata, median) %>%
  290. rename(median_day = median) %>%
  291. mutate(median_mon = round(median_day/30,2))
  292. median_surv
  293. # Show the median survival time
  294. survfit(Surv(survival, status) ~ group, data = matched_data) %>%
  295. tbl_survfit(
  296. probs = 0.5, #survival quantile to return
  297. label_header = "**Median survival (95% CI)**",
  298. )
  299. # kaplan-meier plot
  300. ggsurvplot(fit,
  301. data = matched_data,
  302. xscale = "d_m", #convert days to months
  303. break.time.by=365.25*0.75, #break the x-scale
  304. xlim = c(0, max(matched_data$survival) + 5), #limit the x-axis
  305. surv.median.line = "hv",
  306. #censor.shape = 124,
  307. cumcens= F,
  308. cumevents = F,
  309. palette = c("#f16c23", "#2b6a99"),
  310. xlab = "Time (months)",
  311. ylab = "Survival probability",
  312. legend.title = "Group",
  313. legend.labs = c("Control", "PPRT"),
  314. #pval = T, #p-value
  315. #pval.method = T,
  316. conf.int = T,
  317. risk.table = T,
  318. tables.height = 0.2,
  319. tables.theme = theme_cleantable(),
  320. ggtheme = theme_classic())
  321. ######## EXPORT DATA USED FOR K-M PLOT #############
  322. # saved_data <- matched_data %>%
  323. # select(baseline_id, group, survival, status, subclass) %>%
  324. # arrange(baseline_id)
  325. # write_xlsx(saved_data, "~/Downloads/Figure2_SourceData_OS.xlsx")
  326. ####################################################
  327. ```
  328. ::: callout-important
  329. The median overall survival was `r median_surv$median_mon[2]` months for patients in the intervention PPRT group and `r median_surv$median_mon[1]` months for patients in the control standard of care group.
  330. The log-rank test found a significant difference in survival between patients in the control and treatment group, indicating distinct survival experiences between patients with and without PPRT (p<0.05).
  331. :::
  332. ## Cox proportional hazards model with robust variance estimator
  333. ```{r, eval=T}
  334. #change reference group
  335. matched_data2 <- get_matches(match_obj) %>%
  336. mutate(group = factor(group, levels = c("control", "intervention")),
  337. mgmt = factor(mgmt, levels = c("Unmethylated", "Methylated")))
  338. # A) standard cox regression
  339. # cox_model_std <- coxph(Surv(survival, status) ~ group,
  340. # data = matched_data2)
  341. #
  342. # tabcoxph(cox_model_std, formatp.list = list(decimals = 4, lowerbound = 0.0001, leading0 = F))
  343. # B) cox model with robust variance estimator
  344. cox_model_robust <- coxph(Surv(survival, status) ~ group,
  345. cluster = subclass, #account for clustering within matched sets
  346. data = matched_data2)
  347. # option 2 output
  348. tabcoxph(cox_model_robust, formatp.list = list(decimals = 4, lowerbound = 0.0001, leading0 = F))
  349. # Test the Proportional Hazards Assumption of a Cox Regression
  350. cox.zph(cox_model_robust)
  351. ```
  352. ::: callout-important
  353. Based on the Cox proportional hazards regression (with robust standard errors) output above, death rate from glioblastoma for patients in the PPRT intervention group is **0.34** times the risk of patients in the standard control group. In other words, the analysis showed a significant improvement in survival times among patients who underwent PPRT intervention compared to those that did not (HR=0.34; 95% CI: 0.17-0.69; p=.003).
  354. To test the proportional hazard (PH) assumption, we can use a global Schoenfeld test, a form of goodness-of-fit test. If the PH assumption is true, then the Schoenfeld residuals should be independent of time. As shown above, the test is NOT statistically significant (GLOBAL p>.05). Therefore, the proportional hazard assumption holds.
  355. :::
  356. ## Intent-to-treat (ITT) analysis
  357. Let's use data from the 20 intervention subjects who were enrolled in PPRT, including the 3 subjects treated with tumor treating fields (TTFields) therapy.
  358. ```{r}
  359. intervention <- read_excel("/Volumes/TOSHIBA/work_mac/Documents/collab_yr2/Hamed_SNO_0524/july24/ACC_Trial_11_20_2024_To_Amy.xlsx",
  360. sheet = 2) %>%
  361. janitor::clean_names() %>%
  362. mutate(survival_all = if_else(alive == "n", survival, survival_till_3_20_2025)) %>%
  363. mutate(survival_all = survival) %>%
  364. mutate(year = factor(year(dos))) %>%
  365. #filter(is.na(remove) | remove != "Optune device") %>% #remove optune devices (tumor treating fields)
  366. select(age, sex, alive, survival_all, mgmt, year) %>%
  367. rename(survival = survival_all) %>%
  368. mutate(group = "intervention") %>%
  369. mutate(baseline_id = "000")
  370. comp_data <- rbind(intervention, control) %>%
  371. mutate(status = if_else(alive == "y", 0, 1))
  372. # perform propensity score matching (PSM) with the original control patients
  373. match_obj <- matchit(factor(group) ~ age + sex + mgmt, #eor = GTR and IDH = wildtype
  374. data = comp_data,
  375. method = "nearest",
  376. distance ="glm",
  377. ratio = 4, #how many control units should be matched to each treated unit in k:1 matching
  378. replace = F) #control units can only be matched to ONE treated unit each
  379. #Table 1 for the new matches
  380. # get_matches(match_obj) %>%
  381. # select(age, sex, mgmt, group) %>% #select the variables of interest
  382. # tbl_summary(by = group,
  383. # statistic = list(all_continuous() ~ "{mean}({sd})",
  384. # all_categorical() ~ "{n}({p}%)")) %>%
  385. # modify_header(label ~ "**Variable**") %>%
  386. # #modify_spanning_header(c("stat_1", "stat_2") ~ "**Intervention**") %>%
  387. # bold_labels() %>%
  388. # add_p(pvalue_fun = function(x) style_pvalue(x, digits = 2))
  389. matched_data <- get_matches(match_obj) %>%
  390. arrange(subclass) %>%
  391. mutate(sex = ifelse(sex == "M", 1, 0),
  392. mgmt = ifelse(mgmt == "Methylated", 1, 0))
  393. #median survival
  394. survfit(Surv(survival, status) ~ group, data = matched_data,
  395. #conf.type = "log" #indicate the confidence interval type
  396. ) %>%
  397. tbl_survfit(
  398. probs = 0.5, #survival quantile to return
  399. label_header = "**Median survival (95% CI)**",
  400. )
  401. # Cox PH hazard model w/robust estimators
  402. matched_data2 <- get_matches(match_obj) %>%
  403. mutate(group = factor(group, levels = c("control", "intervention")),
  404. mgmt = factor(mgmt, levels = c("Unmethylated", "Methylated")))
  405. cox_model_robust <- coxph(Surv(survival, status) ~ group,
  406. cluster = subclass, #account for clustering within matched sets
  407. data = matched_data2)
  408. # option 2 output
  409. tabcoxph(cox_model_robust, formatp.list = list(decimals = 4, lowerbound = 0.0001, leading0 = F))
  410. # Test the Proportional Hazards Assumption of a Cox Regression
  411. cox.zph(cox_model_robust)
  412. ```
  413. When we include all of the intervention subjects, the effect of the PPRT intervention is still significant (HR=0.4; 95% CI: 0.21-0.78; p=.007).

qc_and_os_analysis.qmd at commit 20b68d5, no license · at the source

Overview

Authors: Hamed Akbari1, Suyash Mohan2,3,4, Fang Liu5,6, Cecilia Jiang7, Anahita Fathi Kazerooni8,9, Melanie Berger7, Saima Rathore10,11,12, Michel Bilello3,4, Chiharu Sako2,3,4, Jose Garcia2,3,4, Steven Brem8,13, Spyridon Bakas14,15, Elizabeth Mamourian3,4, Abigail Pepin7, Paul James7, Jay Dorsey7, Ashish Singh3,4, Drew Parker2,3,4,16, Ragini Verma2,3,4,16, Quy Cao5,6
and 11 other authorsStephen J Bagley13,17, Arati Desai13,17, Donald M O’Rourke8,13, Russell T Shinohara2,3,5,6, MacLean P Nasrallah2,3,13,18, Boon-Keng Kevin Teo7, Goldie Kurtz7, Gaurav Shukla3,19, Michelle Alonso-Basanta7, Robert A Lustig7, Christos Davatzikos2,3,4
19 affiliations
  1. Department of Bioengineering, School of Engineering, Santa Clara University, Santa Clara, CA USA
  2. Center for Data Science and AI for Integrated Diagnostics (AI2D), University of Pennsylvania, Philadelphia, PA USA
  3. Center for Biomedical Image Computing and Analytics (CBICA), University of Pennsylvania, Philadelphia, PA USA
  4. Department of Radiology, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  5. Penn Statistics in Imaging and Visualization Center, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  6. Center for Clinical Epidemiology and Biostatistics, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  7. Department of Radiation Oncology, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  8. Department of Neurosurgery, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  9. Center for Data-Driven Discovery in Biomedicine (D3b), Children’s Hospital of Philadelphia, Philadelphia, PA USA
  10. Department of Neurology, Emory University School of Medicine, Atlanta, GA USA
  11. Goizueta Alzheimer’s Disease Research Center, Emory University School of Medicine, Atlanta, GA USA
  12. Department of Biomedical Informatics, Emory University School of Medicine, Atlanta, GA USA
  13. GBM Translational Center of Excellence, Abramson Cancer Center, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  14. Department of Pathology and Laboratory Medicine, Indiana University School of Medicine, Indianapolis, IN USA
  15. Department of Radiology and Imaging Sciences, Indiana University School of Medicine, Indianapolis, IN USA
  16. Diffusion and Connectomics in Precision Healthcare Research (DiCIPHR), Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  17. Department of Medicine, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  18. Department of Pathology & Laboratory Medicine, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA USA
  19. Department of Radiation Oncology, Christiana Care Health System, Newark, DE USA
Institutions: Santa Clara University (United States); University of Pennsylvania (United States); Children's Hospital of Philadelphia (United States); Emory University (United States); Abramson Cancer Center (United States); Indiana University School of Medicine; Christiana Care Health System (United States); Christiana Hospital (United States)
Journal: Nature communications, volume 17, issue 1, article 6413
Dates: received 16 May 2025; accepted 10 April 2026; published online 13 May 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41467-026-72545-y · PMID 42129213 · PMCID PMC13377092 · OpenAlex W7161026812
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), other condition (population), clinical / translational (subfield)
Methods: Machine learning, Statistics
Keywords: Translational research, Phase I trials, Radiotherapy, CNS cancer
MeSH: Brain Neoplasms*, Glioblastoma*, Machine Learning*, Precision Medicine*, Adult, Aged, Antineoplastic Agents, Alkylating, Chemoradiotherapy, Female, Humans, Male, Middle Aged, Neoplasm Recurrence, Local, Pilot Projects, Progression-Free Survival, Prospective Studies, Radiotherapy Dosage, Temozolomide (* major topic)
Topic: Glioma Diagnosis and Treatment (Genetics, Medicine), according to OpenAlex
Funding: Funding: This clinical trial was supported by the Penn Abramson Cancer Center.
Citations: not cited yet (Europe PMC); 46 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

CBICA/AI-guided-radiation-dose-escalation

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 20b68d56e0cb20e29cbdb3ca191b8418b9ce052e, 5 March 2026
Languages: Shell (6), Quarto (2), Python (1), R (1)
Size: 17 files, 10 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, 2 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: survival (3 files), tidyverse (3 files), NiBabel (1 file), NumPy (1 file), SciPy (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
11 files

Zenodo 18920301

License: CC-BY-4.0
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: survival (3 files), tidyverse (3 files), NiBabel (1 file), NumPy (1 file), SciPy (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
  • 28 September 2026: the link answers (HTTP 200)
11 files
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41467-026-72545-y.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 20 scripts, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s41467-026-72545-y.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 31 authors, 4 keywords, 18 MeSH terms, 1 funder, 38 references.

Cite

This paper

Akbari, H., Mohan, S., Liu, F., Jiang, C., Fathi Kazerooni, A., Berger, M., Rathore, S., Bilello, M., Sako, C., Garcia, J., Brem, S., Bakas, S., Mamourian, E., Pepin, A., James, P., Dorsey, J., Singh, A., Parker, D., Verma, R., . . . Davatzikos, C. (2026). Personalized machine learning-guided radiation dose escalation in newly diagnosed glioblastoma: prospective pilot study. Nature communications, 17(1), 6413. https://doi.org/10.1038/s41467-026-72545-y

BibTeX

@article{akbari2026personalized,
author = {Akbari, Hamed and Mohan, Suyash and Liu, Fang and Jiang, Cecilia and Fathi Kazerooni, Anahita and Berger, Melanie and Rathore, Saima and Bilello, Michel and Sako, Chiharu and Garcia, Jose and Brem, Steven and Bakas, Spyridon and Mamourian, Elizabeth and Pepin, Abigail and James, Paul and Dorsey, Jay and Singh, Ashish and Parker, Drew and Verma, Ragini and Cao, Quy and Bagley, Stephen J and Desai, Arati and O’Rourke, Donald M and Shinohara, Russell T and Nasrallah, MacLean P and Teo, Boon-Keng Kevin and Kurtz, Goldie and Shukla, Gaurav and Alonso-Basanta, Michelle and Lustig, Robert A and Davatzikos, Christos},
title = {{Personalized machine learning-guided radiation dose escalation in newly diagnosed glioblastoma: prospective pilot study}},
journal = {Nature communications},
year = {2026},
month = may,
volume = {17},
number = {1},
pages = {6413},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-72545-y},
url = {https://doi.org/10.1038/s41467-026-72545-y},
pmid = {42129213},
pmcid = {PMC13377092}
}

RIS

TY - JOUR
AU - Akbari, Hamed
AU - Mohan, Suyash
AU - Liu, Fang
AU - Jiang, Cecilia
AU - Fathi Kazerooni, Anahita
AU - Berger, Melanie
AU - Rathore, Saima
AU - Bilello, Michel
AU - Sako, Chiharu
AU - Garcia, Jose
AU - Brem, Steven
AU - Bakas, Spyridon
AU - Mamourian, Elizabeth
AU - Pepin, Abigail
AU - James, Paul
AU - Dorsey, Jay
AU - Singh, Ashish
AU - Parker, Drew
AU - Verma, Ragini
AU - Cao, Quy
AU - Bagley, Stephen J
AU - Desai, Arati
AU - O’Rourke, Donald M
AU - Shinohara, Russell T
AU - Nasrallah, MacLean P
AU - Teo, Boon-Keng Kevin
AU - Kurtz, Goldie
AU - Shukla, Gaurav
AU - Alonso-Basanta, Michelle
AU - Lustig, Robert A
AU - Davatzikos, Christos
TI - Personalized machine learning-guided radiation dose escalation in newly diagnosed glioblastoma: prospective pilot study
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/05/13
VL - 17
IS - 1
SP - 6413
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-72545-y
UR - https://doi.org/10.1038/s41467-026-72545-y
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-72545-y",
"type": "article-journal",
"title": "Personalized machine learning-guided radiation dose escalation in newly diagnosed glioblastoma: prospective pilot study",
"container-title": "Nature communications",
"author": [
{
"family": "Akbari",
"given": "Hamed"
},
{
"family": "Mohan",
"given": "Suyash"
},
{
"family": "Liu",
"given": "Fang"
},
{
"family": "Jiang",
"given": "Cecilia"
},
{
"family": "Fathi Kazerooni",
"given": "Anahita"
},
{
"family": "Berger",
"given": "Melanie"
},
{
"family": "Rathore",
"given": "Saima"
},
{
"family": "Bilello",
"given": "Michel"
},
{
"family": "Sako",
"given": "Chiharu"
},
{
"family": "Garcia",
"given": "Jose"
},
{
"family": "Brem",
"given": "Steven"
},
{
"family": "Bakas",
"given": "Spyridon"
},
{
"family": "Mamourian",
"given": "Elizabeth"
},
{
"family": "Pepin",
"given": "Abigail"
},
{
"family": "James",
"given": "Paul"
},
{
"family": "Dorsey",
"given": "Jay"
},
{
"family": "Singh",
"given": "Ashish"
},
{
"family": "Parker",
"given": "Drew"
},
{
"family": "Verma",
"given": "Ragini"
},
{
"family": "Cao",
"given": "Quy"
},
{
"family": "Bagley",
"given": "Stephen J"
},
{
"family": "Desai",
"given": "Arati"
},
{
"family": "O’Rourke",
"given": "Donald M"
},
{
"family": "Shinohara",
"given": "Russell T"
},
{
"family": "Nasrallah",
"given": "MacLean P"
},
{
"family": "Teo",
"given": "Boon-Keng Kevin"
},
{
"family": "Kurtz",
"given": "Goldie"
},
{
"family": "Shukla",
"given": "Gaurav"
},
{
"family": "Alonso-Basanta",
"given": "Michelle"
},
{
"family": "Lustig",
"given": "Robert A"
},
{
"family": "Davatzikos",
"given": "Christos"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "6413",
"DOI": "10.1038/s41467-026-72545-y",
"PMID": "42129213",
"PMCID": "PMC13377092",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-72545-y",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
13
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.isci.2026.115657 [code]
Integration of machine learning to develop a disulfidptosis model for predicting glioma prognosis, immunotherapy response, and drug.
Journal: iScience
In common: survival, tidyverse, SciPy, 1 other tool, clinical / translational, other condition, 1 reference
[2] doi:10.1016/j.isci.2026.115329 [code]
Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.
Journal: iScience
In common: survival, NiBabel, tidyverse, 2 other tools, clinical / translational, other condition
[3] doi:10.1158/1078-0432.ccr-25-4419 [code]
Framework for Statistical Parametric Mapping of the Interactions between Glioblastoma Location, Treatment, Prognostic Variables, and Survival Using a Phase III Trial.
Journal: Clinical cancer research : an official journal of the American Association for Cancer Research
In common: NiBabel, SciPy, NumPy, clinical / translational, other condition, 2 references
[4] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: survival, NiBabel, tidyverse, 2 other tools, other condition
[5] doi:10.1016/j.cell.2026.05.026 [code]
The critical role of the endogenous immune compartment after CAR T cell therapy in recurrent GBM.
Journal: Cell
In common: survival, tidyverse, SciPy, 1 other tool, other condition, 1 reference
[6] doi:10.1016/j.xcrm.2026.102682 [code]
TET CpG sequence-context-specific DNA demethylation shapes progression of IDH-mutant gliomas.
Journal: Cell reports. Medicine
In common: survival, tidyverse, NumPy, other condition, 1 reference
[7] doi:10.1038/s41467-026-74515-w [code]
Advancing fair and explainable machine learning for neuroimaging dementia pattern classification in multi-racial and multi-ethnic populations.
Journal: Nature communications
In common: NumPy, author Christos Davatzikos
[8] doi:10.1038/s41467-026-72091-7 [code]
Coupled cross-sectional and longitudinal non-negative matrix factorization reveals dominant brain aging trajectories in 48,949 individuals.
Journal: Nature communications
In common: NumPy, author Christos Davatzikos
[9] doi:10.1038/s41467-026-70287-5 [code]
Cortical-limbic circuit dynamics of approach-avoidance conflict in humans.
Journal: Nature communications
In common: survival, NiBabel, tidyverse, 2 other tools
[10] doi:10.1126/sciadv.aee2305 [code]
Prediction of mild cognitive impairment progression using time-sensitive multimodal biomarkers.
Journal: Science advances
In common: survival, NiBabel, tidyverse, 1 other tool, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.