OSCR

JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews

Code ↔ Paper

4 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 4 matches
  1. [1] § Methods › Text data pre-processing › Step 4: Modelling ↔ scripts/Hypertension h2o.R, lines 452–540 · score 0.63 · h2o.deeplearning, cross validate, training samples, folds, classified, models
  2. [2] § Methods › Text data pre-processing › Step 4: Modelling ↔ scripts/Anticoag atrial fib h2o.R, lines 446–529 · score 0.63 · h2o.deeplearning, cross validate, training samples, folds, classified, models
  3. [3] § Methods › Text data pre-processing ↔ scripts/Anticoag 2 h2o.R, lines 239–321 · score 0.59 · TF IDF, PCA, SMART, ngrams, tokenised, words
  4. [4] § Methods › Text data pre-processing ↔ scripts/Sedat behaviour children h2o.R, lines 238–309 · score 0.59 · TF IDF, PCA, SMART, ngrams, tokenised, words

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 764 lines · 32 KB · no license · 1 match

  1. ## libraries ####
  2. library(tidyverse)
  3. library(recipes)
  4. library(textrecipes)
  5. library(h2o)
  6. library(future.apply)
  7. library(tibble)
  8. library(dplyr)
  9. library(ggpubr)
  10. library(PRROC)
  11. # The two functions below allowed me to do hyperparameter tuning.
  12. # The first one, ending in 'parallel', allows me to open several parallel sessions
  13. # and test each hyperparameter combination in parallel. Requires a TON of RAM, so
  14. # I use it carefully.
  15. # The second one, ending in repeats, will test each combination in sequence. Takes
  16. # longer, but it works.
  17. run_each5_with_repeats_parallel <- function(df, n, epochs, hiddenunis, activ, stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate, repeats, n_workers) {
  18. # remove contents of the h2o server if one is open already
  19. if (tryCatch({ h2o.clusterStatus(); TRUE }, error = function(e) FALSE)) {
  20. h2o.removeAll(timeout_secs = 0)
  21. }
  22. # 1. Build all param combinations × repeats
  23. param_combinations <- expand.grid(
  24. epochs = epochs,
  25. hidden_units = hiddenunis,
  26. activ = activ,
  27. stop_rounds = stop_rounds,
  28. stop_tol = stop_tol,
  29. rates_anneal = rates_anneal,
  30. min_batch = min_batch,
  31. l2 = l2,
  32. rate = rate,
  33. stringsAsFactors = FALSE
  34. )
  35. jobs <- do.call(rbind, lapply(seq_len(nrow(param_combinations)), function(i) {
  36. cbind(param_combinations[i, , drop=FALSE], repeat_run = seq_len(repeats), ID = i)
  37. }))
  38. jobs <- as.data.frame(jobs)
  39. # 2. Function to run a single job
  40. run_single <- function(job_row) {
  41. epochs <- as.numeric(job_row["epochs"])
  42. stop_rounds <- as.numeric(job_row["stop_rounds"])
  43. stop_tol <- as.numeric(job_row["stop_tol"])
  44. rates_anneal <- as.numeric(job_row["rates_anneal"])
  45. min_batch <- as.numeric(job_row["min_batch"])
  46. hidden_units <- as.character(job_row["hidden_units"])
  47. activ <- as.character(job_row["activ"])
  48. repeat_run <- as.numeric(job_row["repeat_run"])
  49. l2 <- as.numeric(job_row["l2"])
  50. rate <- as.numeric(job_row["rate"])
  51. ID <- as.numeric(job_row["ID"])
  52. set.seed(123 + ID + repeat_run) # More variability per run
  53. start_time <- Sys.time()
  54. res <- each5.2(df, n, epochs, hidden_units, activ, stop_rounds, stop_tol,rates_anneal,min_batch, l2, rate)
  55. end_time <- Sys.time()
  56. res %>%
  57. mutate(
  58. ID = ID,
  59. repeat_run = repeat_run,
  60. configs = paste(epochs, hidden_units, activ, stop_rounds, stop_tol ,rates_anneal,min_batch,l2,rate, sep = ", "),
  61. elapsed_time = as.numeric(difftime(end_time, start_time, units = "secs"))
  62. )
  63. }
  64. # 3. Set up parallel plan
  65. future::plan(multisession, workers = 12)
  66. plan()
  67. # 4. Run in parallel!
  68. results_list <- future.apply::future_lapply(
  69. split(jobs, seq_len(nrow(jobs))),
  70. function(job) run_single(job[1, ])
  71. )
  72. # 5. Combine all results into one tibble
  73. results <- dplyr::bind_rows(results_list)
  74. return(results)
  75. }
  76. run_each5_with_repeats <- function(df, n, epochs, hiddenunis, activ, stop_rounds, stop_tol,rates_anneal,min_batch, l2, rate, repeats) {
  77. # remove contents of the h2o server if one is open already
  78. if (tryCatch({ h2o.clusterStatus(); TRUE }, error = function(e) FALSE)) {
  79. h2o.removeAll(timeout_secs = 0)
  80. }
  81. elapsed_time <- numeric() # Initialize as numeric vector
  82. param_combinations <- expand.grid(
  83. epochs = epochs,
  84. hidden_units = hiddenunis,
  85. activ = activ,
  86. stop_rounds = stop_rounds,
  87. stop_tol = stop_tol,
  88. rates_anneal = rates_anneal,
  89. min_batch = min_batch,
  90. l2 = l2,
  91. rate = rate
  92. )
  93. # Initialize results as an empty tibble
  94. results <- tibble()
  95. total_runs <- nrow(param_combinations) * repeats # Total number of iterations
  96. current_run <- 0 # Initialize counter for completed runs
  97. # Iterate over each combination of parameters
  98. for (i in seq_len(nrow(param_combinations))) {
  99. print(paste(i, '/', nrow(param_combinations)))
  100. epochs <- param_combinations$epochs[i]
  101. hidden_units <- as.character(param_combinations$hidden_units[i])
  102. activ <- param_combinations$activ[i]
  103. stop_rounds <- param_combinations$stop_rounds[i]
  104. stop_tol <- param_combinations$stop_tol[i]
  105. rates_anneal <- param_combinations$rates_anneal[i]
  106. min_batch <- param_combinations$min_batch[i]
  107. l2 <- param_combinations$l2[i]
  108. rate <- param_combinations$rate[i]
  109. set.seed(123)
  110. for (repeat_idx in seq_len(repeats)) {
  111. current_run <- current_run + 1 # Increment the completed runs counter
  112. start_time <- Sys.time() # Record start time
  113. current_result <- each5.2(df, n, epochs, hidden_units, activ, stop_rounds, stop_tol, rates_anneal, min_batch, l2, rate) %>%
  114. mutate(
  115. ID = i,
  116. repeat_run = repeat_idx,
  117. configs = paste(epochs, hidden_units, activ, stop_rounds, stop_tol, rates_anneal, min_batch, l2, rate, sep = ", ")
  118. )
  119. # Append the current result to the results tibble
  120. results <- bind_rows(results, current_result)
  121. end_time <- Sys.time() # Record end time
  122. # Calculate elapsed time for this run
  123. elapsed_time <- c(elapsed_time, as.numeric(difftime(end_time, start_time, units = "secs")))
  124. mean_time <- mean(elapsed_time, na.rm = TRUE)
  125. remaining_runs <- total_runs - current_run
  126. estimated_remaining_time <- mean_time * remaining_runs
  127. # Convert estimated time to minutes and seconds
  128. estimated_minutes <- floor(estimated_remaining_time / 60)
  129. estimated_seconds <- round(estimated_remaining_time %% 60)
  130. # Print estimated time remaining
  131. print(paste(
  132. "Estimated remaining time:",
  133. estimated_minutes, "min", estimated_seconds, "sec"
  134. ))
  135. }
  136. }
  137. return(results)
  138. }
  139. ## ROC-PR FUNCTION####
  140. # calc_aucpr is a function designed to calculate aucpr for our analyses. It'll be used in the 'summaries' section.
  141. calc_aucpr <- function(df_iter) {
  142. # df_iter = subset of results for a single iteration
  143. # True labels as 1/0
  144. truth <- ifelse(df_iter$finaldecision == "Include", 1, 0)
  145. # Predicted scores (probabilities)
  146. scores <- df_iter$newpred
  147. # PRROC needs the scores of the positive class separately:
  148. pr <- pr.curve(
  149. scores.class0 = scores[truth == 0],
  150. scores.class1 = scores[truth == 1],
  151. curve = FALSE
  152. )
  153. return(pr$auc.integral)
  154. }
  155. ##hypertension data clean #####
  156. # In this section we read the file in 'data' containing our abstracts, decisions and the LLM responses.
  157. # We break the responses into columns, clean them, create the numberical variables (PICOS scores).
  158. dfwithPICOS <- read_rds('dfgpt.stress.complete.rds')
  159. dfPICOSfinal <- dfwithPICOS %>%
  160. separate_wider_delim(cols = GPT_Response, delim = ',', names = c('review', 'P', 'I', 'C', 'O', 'S'), cols_remove = F) %>%
  161. mutate(review = chartr("[],012345 '","...........", review),
  162. review = chartr("- ","..", review),
  163. review = gsub('[.]','',review),
  164. P = chartr("[],012345 '","...........", P),
  165. P = chartr("- ","..", P),
  166. P = gsub('[.]','',P),
  167. I = chartr("[],012345 '","...........", I),
  168. I = chartr("- ","..", I),
  169. I = gsub('[.]','',I),
  170. C = chartr("[],012345 '","...........", C),
  171. C = chartr("- ","..", C),
  172. C = gsub('[.]','',C),
  173. O = chartr("[],012345 '","...........", O),
  174. O = chartr("- ","..", O),
  175. O = gsub('[.]','',O),
  176. S = chartr("[],012345 '","...........", S),
  177. S = chartr("- ","..", S),
  178. S = gsub('[.]','',S)) %>%
  179. mutate(Pn = case_when(P =='Yes'~ 1, ### create numerical variables with PICOS score
  180. P == 'No' ~ 0,
  181. P == 'Uncertain' ~ 0.5),
  182. In = case_when(I =='Yes'~ 1,
  183. I == 'No' ~ 0,
  184. I == 'Uncertain' ~ 0.5),
  185. Cn = case_when(C =='Yes'~ 1,
  186. C == 'No' ~ 0,
  187. C == 'Uncertain' ~ 0.5),
  188. On = case_when(O =='Yes'~ 1,
  189. O == 'No' ~ 0,
  190. O == 'Uncertain' ~ 0.5),
  191. Sn = case_when(S =='Yes'~ 1,
  192. S == 'No' ~ 0,
  193. S == 'Uncertain' ~ 0.5)) %>%
  194. dplyr::select(key, keywords, title, abstract, authors, year, review, P, I, C, O, S, Pn, In,Cn, On, Sn, decision, finaldecision) %>%
  195. mutate(totalscore = rowSums(.[c('Pn','In','Cn','On','Sn')])) %>%
  196. filter(key != 'ardiologie et Pneumologie de Quebec' & key != 'iomedical Science') %>%
  197. distinct(key, .keep_all = T)
  198. ## PREPARE DATA #####
  199. # In this section we clean the text variables and prepare them for tokenisation and TF-IDF generation.
  200. dftoken <- dfPICOSfinal %>%
  201. select(key, title, authors, abstract, totalscore, review, P, I, C, O, S,Pn, In,Cn, On, Sn, decision, finaldecision) %>%
  202. mutate(abstractsub = gsub('-', ' ', abstract)) %>% # Remove hyphens
  203. mutate(abstractsub = iconv(abstractsub, from = "", to = "ASCII//TRANSLIT")) %>% # Remove/convert weird chars
  204. mutate(abstractsub = gsub("–|—|â€|’|“|‘|•|−", " ", abstractsub)) %>% # Remove common mojibake
  205. mutate(abstractsub = gsub('[[:punct:]]', '', abstractsub)) %>% # Remove punctuation
  206. mutate(abstractsub = tolower(abstractsub)) %>%
  207. mutate(abstractsub = gsub('[0-9]+', ' ', abstractsub)) %>%
  208. mutate(abstractsub = gsub("\\s+", " ", abstractsub)) %>% # Normalize whitespace
  209. mutate(titlesub = gsub('-', ' ', title)) %>% # Remove hyphens
  210. mutate(titlesub = iconv(titlesub, from = "", to = "ASCII//TRANSLIT")) %>% # Remove/convert weird chars
  211. mutate(titlesub = gsub("–|—|â€|’|“|‘|•|−", " ", titlesub)) %>% # Remove common mojibake
  212. mutate(titlesub = gsub('[[:punct:]]', '', titlesub)) %>% # Remove punctuation
  213. mutate(titlesub = tolower(titlesub)) %>%
  214. mutate(titlesub = gsub('[0-9]+', ' ', titlesub)) %>%
  215. mutate(titlesub = gsub("\\s+", " ", titlesub)) %>% # Normalize whitespace
  216. mutate_at(vars(P, I, C, O, S), funs(factor(., levels = c('Uncertain', 'Yes', 'No') ))) %>%
  217. mutate(reviewn = case_when(review == 'No' ~ 1,
  218. review == 'Yes' ~ 0 ))
  219. rm(dfPICOSfinal)
  220. rm(dfwithPICOS)
  221. tokenization_recipe <- recipe(decision ~ key + abstract + authors + abstractsub + P + I + C + O + S + totalscore + review + finaldecision, data = dftoken) %>%
  222. update_role(key, new_role = "id") %>% #Mark 'key' as an ID
  223. update_role(abstract, new_role = "id") %>%
  224. update_role(authors, new_role = "id") %>%
  225. update_role(finaldecision, new_role = "id") %>%
  226. step_tokenize(abstractsub, token = "words") %>%
  227. step_stopwords(abstractsub, stopword_source = "smart") %>%
  228. #step_stem(abstractsub) %>%
  229. step_ngram(abstractsub, num_tokens = 3, min_num_tokens = 1) %>%
  230. step_tokenfilter(abstractsub, max_tokens = 7000) %>%
  231. step_tfidf(abstractsub) %>%
  232. #step_pca(starts_with('tfidf'),threshold = .80, prefix = 'KPCa') %>%
  233. #step_tokenize(titlesub, token = "words") %>%
  234. #step_stopwords(titlesub, stopword_source = "smart") %>%
  235. #step_stem(titlesub) %>%
  236. #step_ngram(titlesub, num_tokens = 3, min_num_tokens = 1) %>%
  237. #step_tokenfilter(titlesub, max_tokens = 3000 ) %>%
  238. #step_tfidf(titlesub) %>%
  239. #step_pca(c(reviewn, totalscore,keys1, score1), num_comp = 1, keep_original_cols = T) %>%
  240. step_pca(starts_with('tfidf'),threshold = .90, prefix = 'KPCa') %>%
  241. step_center(starts_with('KPC')) %>%
  242. step_range(totalscore, min = 0, max = 1) %>%
  243. step_dummy( all_of(c('P', 'I', 'C','O','S', 'review')), one_hot = T)
  244. prepped_recipe <- prep(tokenization_recipe, training = dftoken, retain = TRUE) # Preprocess and retain
  245. baked_df <- bake(prepped_recipe, new_data = NULL)
  246. dir.create('baked', showWarnings = F, recursive = T)
  247. saveRDS(baked_df, 'baked/hypertension final.rds')
  248. ##START FROM HERE WITH THE BAKED READY####
  249. # Sometimes, 'baking' the dataframe takes some time depending on your hardware. I preferred to save
  250. # the dataframe and reload it here below. Also this is where I include the weights.
  251. baked_df <- read_rds('baked/hypertension final.rds') %>%
  252. mutate(authors = ifelse(is.na(authors),'no author',authors )) %>%
  253. mutate(finaldecision = as.factor(finaldecision)) %>%
  254. mutate(weightsc = ifelse(decision == "Include", 40,1))
  255. ggplot(baked_df, aes( x = totalscore * 5)) +
  256. geom_histogram( aes(fill = finaldecision, y = ..density..)) +
  257. theme_classic() +
  258. theme(legend.position = "top") +
  259. labs(x = 'PICOS score', fill = "FT decision")
  260. ## new function 5 each time #####
  261. # each5.2 performs iterative active-screening with an H2O deep-learning classifier.
  262. # Workflow per round:
  263. # 1) Select new records to review (top/bottom scoring records plus a small buffer),
  264. # then add them to the cumulative labeled sample (`samplefull`).
  265. # 2) Build training folds by repeating Include cases across folds and sampling Exclude
  266. # cases per fold to keep class balance for cross-validation.
  267. # 3) Train a Bernoulli H2O deep-learning model with the provided hyperparameters.
  268. # 4) Score unscreened records with each CV model, average Include probabilities,
  269. # normalize to `newpred`, and assign Include/Exclude labels using `threshold`.
  270. # 5) Save per-round predictions and error-tracking fields, and mark low-probability
  271. # records as candidates to exclude in later sampling.
  272. # Early stopping: after the minimum rounds, stop when too few records remain above
  273. # threshold for consecutive rounds. Returns one data frame with all rounds combined.
  274. each5.2 <- function(df1, max, epoc, hidu, activ, stop_rounds, stop_tol, rates_anneal,min_batch, l2, rate) {
  275. # Initialize variables for Stage 2
  276. results <- vector("list", max)
  277. exclude_keys <- data.frame(NA)
  278. samplefull <- data.frame()
  279. preds <- df1
  280. min_rounds <- 71
  281. prefix <- paste0("worker", Sys.getpid(), "_", as.integer(Sys.time()), "_")
  282. threshold <- 0.4
  283. min_hold <- 1
  284. since_change <- min_hold
  285. for (i in seq_len(max)) {
  286. time1 <- Sys.time()
  287. if(i > 1) {
  288. # inc_rate <- (include_count + exclude_count) / nrow(df1)
  289. # print(inc_rate)
  290. #
  291. # cap <- dplyr::case_when(
  292. # #inc_rate > 0.30 ~ 0.75,
  293. # #inc_rate > 0.25 ~ 0.70,
  294. # #inc_rate > 0.20 ~ 0.60,
  295. # #inc_rate > 0.15 ~ 0.55,
  296. # inc_rate > 0.10 ~ 0.5,
  297. # inc_rate > 0.05 ~ 0.45,
  298. # TRUE ~ threshold
  299. # )
  300. #
  301. #
  302. # old <- threshold
  303. #
  304. # if (since_change >= min_hold && cap > threshold) {
  305. # threshold <- cap # jump directly to the mapped cap
  306. # since_change <- 1 # or set to 1 if you prefer counting this iter as "held"
  307. # } else {
  308. # since_change <- since_change + 1
  309. # }
  310. # print(paste('threshold =', threshold))
  311. keys_to_exclude <- exclude_keys$key
  312. keys_in_sample <- samplefull$key
  313. preds <- preds %>%
  314. suppressMessages(left_join(.,df1))
  315. sampled_df <- preds %>%
  316. ungroup() %>%
  317. filter(!key %in% keys_in_sample) %>%
  318. filter(!key %in% keys_to_exclude) %>%
  319. arrange(desc(Include)) %>%
  320. slice_head(n=15) %>%
  321. bind_rows(preds %>%
  322. ungroup() %>%
  323. filter(!key %in% keys_in_sample) %>%
  324. filter(!key %in% keys_to_exclude) %>%
  325. arrange(desc(Include)) %>%
  326. slice_tail(n=8) )%>%
  327. bind_rows(preds %>%
  328. ungroup() %>%
  329. filter(!key %in% keys_in_sample) %>%
  330. filter(key %in% keys_to_exclude) %>%
  331. arrange(desc(Include)) %>%
  332. slice_head(n=7)) %>%
  333. select(-predict, -Include, -Exclude, -new, -incpred, -excpred, -newpred, - thresh)
  334. }else{
  335. sampled_df <- preds %>%
  336. arrange(desc(totalscore)) %>%
  337. slice(c((nrow(.)-4):nrow(.), 1:20)) %>%
  338. bind_rows(preds %>%
  339. arrange(desc(totalscore)) %>%
  340. filter(totalscore*5 <= 2.5) %>% slice(1:5))
  341. localH2O = h2o.init(ip="localhost", port = 54321,
  342. startH2O = TRUE, nthreads=1, max_mem_size = '8G')
  343. prefix <- paste0("worker", Sys.getpid(), "_", as.integer(Sys.time()), "_")
  344. }
  345. samplefull <- samplefull %>%
  346. bind_rows(sampled_df %>% mutate(iter = i)) %>%
  347. distinct()
  348. includes <- samplefull %>% filter(decision == "Include")
  349. excludes <- samplefull %>% filter(decision == "Exclude")
  350. n_excl_per_fold <- nrow(includes) # or set any number you prefer per fold
  351. # 1) 3 repeats of each Include, labeled folds 1, 2, 3
  352. includes3 <- includes %>%
  353. mutate(.k = 1L) %>%
  354. left_join(data.frame(folds = 1:3, .k = 1L), by = ".k",relationship = "many-to-many") %>%
  355. select(-.k)
  356. # 2) Randomize all Excludes and partition into 3 (no overlap / no replacement)
  357. excl_folds <- lapply(1:3, function(f) {
  358. excludes %>%
  359. slice_sample(n = min(n_excl_per_fold, nrow(.)), replace = FALSE) %>%
  360. mutate(folds = f)
  361. }) %>% bind_rows()
  362. # 3) Combine
  363. samplefull2 <- bind_rows(includes3, excl_folds)
  364. h2otrain <- h2o.na_omit(as.h2o(samplefull2, destination_frame = paste0(prefix, 'train1',i)))
  365. test <- df1 %>%
  366. filter(!(key %in% samplefull$key))
  367. h2otest <- h2o.na_omit(as.h2o(test, destination_frame = paste0(prefix, 'preddata',i)))
  368. h2otest <- h2o.na_omit(h2otest)
  369. include_count <- samplefull %>%
  370. filter(decision == "Include") %>%
  371. nrow()
  372. exclude_count <- samplefull %>%
  373. filter(decision == "Exclude") %>%
  374. nrow()
  375. # Update the recipe to use RELEVANCE as the outcome
  376. y <- "decision"
  377. x <- names(df1)[c( 7:(ncol(df1)-1))]
  378. model = h2o.deeplearning(x=x,
  379. y=y,
  380. training_frame=h2otrain,
  381. distribution = "bernoulli",
  382. activation = as.character(activ),
  383. hidden = eval(parse(text = hidu)),
  384. epochs = epoc, l1 = 0, l2 = l2,rate = rate,
  385. loss = "CrossEntropy",
  386. initial_weight_distribution = c("UniformAdaptive"),
  387. ignore_const_cols = F, seed = 123,
  388. adaptive_rate = F, standardize = F,
  389. nesterov_accelerated_gradient = T,
  390. fast_mode = T,
  391. reproducible = T,
  392. overwrite_with_best_model = T,
  393. score_training_samples = 500, score_validation_samples = 500,
  394. stopping_metric = "mean_per_class_error",classification_stop = -1,
  395. stopping_rounds = stop_rounds, stopping_tolerance = stop_tol,
  396. model_id = paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate, sep = "_" ),
  397. #nfolds = 3, fold_assignment = 'Stratified',
  398. fold_column = 'folds',
  399. rate_annealing = rates_anneal,# rate_decay = 0.8,
  400. keep_cross_validation_models = T,
  401. keep_cross_validation_predictions = T, weights_column = 'weightsc',
  402. mini_batch_size = min_batch#, balance_classes = T,
  403. #max_after_balance_size = max(.1,include_count/exclude_count),
  404. #class_sampling_factors = c(1, exclude_count / include_count)
  405. )
  406. modelcv1 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_1',sep = '_'))
  407. predweird1 <- predict(modelcv1, newdata = h2otest )
  408. predweird1.2 <- bind_cols(test, as.data.frame(predweird1)) %>%
  409. mutate(pp = 1) %>%
  410. select(-starts_with('KPC'), -contains('_'))
  411. modelcv2 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_2',sep = '_'))
  412. predweird2 <- predict(modelcv2, newdata = h2otest )
  413. predweird2.2 <- bind_cols(test, as.data.frame(predweird2)) %>%
  414. mutate(pp = 2) %>%
  415. select(-starts_with('KPC'), -contains('_'))
  416. modelcv3 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_3',sep = '_'))
  417. predweird3 <- predict(modelcv3, newdata = h2otest )
  418. predweird3.2 <- bind_cols(test, as.data.frame(predweird3)) %>%
  419. mutate(pp = 3) %>%
  420. select(-starts_with('KPC'), -contains('_'))
  421. all <- bind_rows(predweird1.2,
  422. predweird2.2,
  423. predweird3.2) %>%
  424. group_by(key, decision, finaldecision) %>%
  425. summarise_at(vars(Include), funs(mean(., na.rm =T))) %>%
  426. ungroup() %>%
  427. mutate(Exclude = 1-Include,
  428. newpred = (Include - min(Include)) / (max(Include) - min(Include) ),
  429. predict = ifelse(newpred >= 0.5, "Include", "Exclude"))
  430. preds <- test %>%
  431. merge(.,all, .by = key) %>%
  432. mutate(new = i) %>%
  433. mutate(incpred = include_count, # Add the "Include" count
  434. excpred = exclude_count) %>%
  435. mutate(thresh = threshold)
  436. my_keys <- h2o.ls()[,1]
  437. h2o.rm(my_keys[grepl(paste0("^", prefix), my_keys)], cascade = TRUE)
  438. gc()
  439. time2 <- Sys.time()
  440. exclude_keys <- preds %>%
  441. ungroup() %>% filter(Include < threshold) %>%
  442. select(key, new)
  443. # Add the "Exclude" count
  444. results[[i]] <- preds %>%
  445. select(-starts_with('KPC'), -starts_with('tfidf'), -starts_with(c('P_','I_', 'C_', 'O_', 'S_')))
  446. sumtextPICOS <- (bind_rows(results)) %>%
  447. mutate(newclass = ifelse(Include < thresh,"Exclude", "Include")) %>%
  448. mutate(inc.correct = ifelse(finaldecision == "Include" & newclass == "Include",1,0) ) %>%
  449. mutate(exc.correct = ifelse(finaldecision == "Exclude" & newclass == "Exclude",1,0) ) %>%
  450. mutate(exc.incorrect = ifelse(finaldecision == "Include" & newclass == "Exclude",1,0) ) %>%
  451. mutate(inc.incorrect = ifelse(finaldecision == "Exclude" & newclass == "Include",1,0) ) %>%
  452. group_by(., new) %>% #, thresh) %>%
  453. summarise_at(vars(inc.correct,exc.correct,inc.incorrect, exc.incorrect), funs (sum(.,na.rm= T))) %>%
  454. merge(.,bind_rows(results) %>% group_by(., new) %>%
  455. summarise_at(vars(incpred, excpred), funs(max(.))), .by = c('configs', 'new')) %>%
  456. arrange(-desc(new))
  457. print(ggarrange(nrow = 2,ncol = 1, ggplot(data = sumtextPICOS, aes(x = new, y = inc.incorrect)) +
  458. geom_line(colour = 'blue') +
  459. geom_line(colour = 'red', aes(y = exc.incorrect)) +
  460. geom_point(pch = 21, colour = 'blue') +
  461. geom_point(pch = 21, colour = 'red', aes(y = exc.incorrect)) +
  462. geom_text(vjust = -0.5, size = 3 ,aes(y = exc.incorrect, label = exc.incorrect))+
  463. geom_text(vjust = -0.5, size = 3 ,aes(y = inc.incorrect, label = inc.incorrect))+
  464. #facet_wrap(~configs) +
  465. #scale_x_continuous(breaks = c(0,2,4,6,8,10)) +
  466. scale_y_continuous(limits = c(0,max(sumtextPICOS$inc.incorrect)),name = 'Blue = Included Incorrectly', sec.axis = sec_axis(~. , name = "Red = Excluded Incorrectly")) +
  467. labs(x = 'Rounds') +
  468. theme_classic(),
  469. ggplot(data = subset(preds), aes(x = (Include))) +
  470. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  471. scale_fill_manual(values = c('black', 'red')) +
  472. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  473. #facet_wrap(~new* configs, ncol = 2, scale = 'free')+
  474. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  475. theme_classic() +
  476. theme(legend.position = "top")))
  477. if (min_rounds <= i && sum(preds$Include >= threshold, na.rm = T) <= max(30, nrow(df1) / 100)) {
  478. return(bind_rows(results))
  479. break
  480. }
  481. gc()
  482. print((time2 - time1))
  483. print(i)
  484. }
  485. return(bind_rows(results))
  486. }
  487. ##run it here#####
  488. # This block sets one hyperparameter configuration and runs the full
  489. # iterative simulation pipeline.
  490. # - The vectors below are the model settings explored by
  491. # `run_each5_with_repeats` (single values here = one configuration).
  492. # - `TIMES` controls how many full repeats are executed.
  493. # - The returned object `results` contains per-round predictions/metrics.
  494. # - Results are persisted to `results/resultsanticoag2.rds` so the
  495. # summarisation section can be rerun without retraining.
  496. epochs <- c(100)
  497. hiddenunis <- c('c(50,25,10,5)')
  498. activ <- c('Tanh')
  499. stop_rounds <- c(3)
  500. stop_tol <- c(1e-5)
  501. rates_anneal <- c(0.001)
  502. min_batch <- c(1)
  503. l2 <- c(0.02575)
  504. rate <- c(0.001)
  505. TIMES <- 1
  506. results <- run_each5_with_repeats(baked_df, 71, epochs,hiddenunis,activ, stop_rounds, stop_tol,rates_anneal, min_batch,l2,rate, TIMES)
  507. dir.create('results', showWarnings = F, recursive = T)
  508. saveRDS(results, 'results/resultshypertension.rds')
  509. ##summarise#####
  510. results <- read_rds('results/resultshypertension.rds')
  511. EXCINC <- results %>%
  512. group_by(configs) %>%
  513. filter(finaldecision == "Include" & newpred < thresh) %>%
  514. select(-starts_with(c('P_','I_', 'C_', 'O_', 'S_', 'KPC', 'review_')))
  515. sumtextPICOS <- results %>%
  516. mutate(newclass = ifelse(Include < thresh,"Exclude", "Include")) %>%
  517. mutate(inc.correct = ifelse(finaldecision == "Include" & newclass == "Include",1,0) ) %>%
  518. mutate(exc.correct = ifelse(finaldecision == "Exclude" & newclass == "Exclude",1,0) ) %>%
  519. mutate(exc.incorrect = ifelse(finaldecision == "Include" & newclass == "Exclude",1,0) ) %>%
  520. mutate(inc.incorrect = ifelse(finaldecision == "Exclude" & newclass == "Include",1,0) ) %>%
  521. group_by(., configs, new, ID, thresh) %>% #, thresh) %>%
  522. summarise_at(vars(inc.correct,exc.correct,inc.incorrect, exc.incorrect), funs (sum(.,na.rm= T))) %>%
  523. merge(.,results %>% group_by(., configs, new, ID,thresh) %>%
  524. summarise_at(vars(incpred, excpred), funs(max(.))), .by = c('configs', 'new')) %>%
  525. mutate(reads = incpred+ excpred) %>%
  526. mutate(percread = (incpred + excpred )/ nrow(baked_df) * 100 ,
  527. percsave = exc.correct/nrow(baked_df) * 100,
  528. ratio = incpred/excpred * 100) %>%
  529. merge(.,results %>%
  530. group_by(configs, new, ID,thresh) %>%
  531. summarise(
  532. aucpr = calc_aucpr(cur_data())
  533. ),.by = new) %>%
  534. merge(.,results %>%
  535. group_by(new, finaldecision, configs) %>%
  536. summarise(foundft = (193 - n() ) / 193 * 100) %>%
  537. filter(finaldecision == "Include") %>%
  538. select(-finaldecision))%>%
  539. mutate(recall = inc.correct / (inc.correct + exc.incorrect) * 100,
  540. recall_cumm = (sum(baked_df$finaldecision == "Include") - exc.incorrect ) / ((sum(baked_df$finaldecision == "Include") - exc.incorrect ) + exc.incorrect) * 100,
  541. specificity = exc.correct / (exc.correct + inc.incorrect) * 100) %>%
  542. arrange(-desc(new)) ;View(sumtextPICOS)
  543. calclong <- sumtextPICOS %>%
  544. mutate(aucpr = aucpr*100) %>%
  545. bind_rows(data.frame(new = 0, percread = 0, specificity = 0, foundft =0, recall=0, recall_cumm = 0)) %>%
  546. pivot_longer(cols =c( specificity, foundft, recall, recall_cumm)) %>%
  547. mutate(name = factor(name))
  548. ggplot(data = sumtextPICOS, aes(x = (new*30 / nrow(baked_df)* 100), y = (inc.incorrect + inc.correct ))) +
  549. geom_line(colour = 'blue') +
  550. geom_line(colour = 'red', aes(y = exc.incorrect)) +
  551. geom_point(pch = 21, colour = 'blue') +
  552. geom_point(pch = 21, colour = 'red', aes(y = exc.incorrect)) +
  553. geom_text(vjust = -0.5, size = 3 ,aes(y = exc.incorrect, label = exc.incorrect))+
  554. geom_text(vjust = -0.5, size = 3 ,aes(y = inc.incorrect + inc.correct , label = inc.incorrect + inc.correct ))+
  555. scale_y_continuous(name = 'Blue = Suggested Includes', sec.axis = sec_axis(~. , name = "Red = Excluded Incorrectly")) +
  556. labs(x = '% of studies read') +
  557. theme_classic() +
  558. facet_wrap(~configs)
  559. ggplot(data = calclong, aes(x = percread, y = value)) +
  560. geom_line(aes(colour = name), size = 1, alpha= 0.5) +
  561. geom_abline(slope = 0, intercept =100, linetype=3) +
  562. scale_y_continuous(limits = c(0,100),name = 'Performance value') +
  563. labs(x = '% of studies read', colour = "Metric") +
  564. scale_color_manual(labels = c("% of includes identified", "Iteration Recall" ,"Joint recall", "Specificity"), values= c('red', 'blue', 'green', 'purple'))+
  565. scale_x_continuous(expand=c(0, 0)) +
  566. geom_vline(xintercept = 28.2030620, linetype = 2, colour = 'orange')+
  567. geom_vline(xintercept = 22.9653505, linetype = 3, colour = 'gray20')+
  568. theme_classic()
  569. #ggsave('Hypertension.png', width = 6, height = 3, dpi = 300)
  570. summary(sumtextPICOS$aucpr)
  571. ## Histograms ####
  572. ggplot(data = subset(results, new== 1| new == 35 | new == 70), aes(x = (Include))) +
  573. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  574. scale_fill_manual(values = c('black', 'red')) +
  575. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  576. facet_wrap(~new ,ncol = 3, scale = 'free')+
  577. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  578. theme_classic() +
  579. theme(legend.position = "top") + scale_x_continuous(limits = c(0,1))
  580. #ggsave('hypertension distrib.png', width = 10, height = 3.3, dpi = 300)
  581. probs1 <- results%>%
  582. arrange((newpred)) %>%
  583. mutate(p = pmin(pmax(newpred, 1e-6), 1 - 1e-6)) %>%
  584. mutate(x = qlogis(p))
  585. ggplot(subset(probs1, new == 1), aes(x, p)) +
  586. stat_function(fun = plogis, xlim = range(probs1$x), size = 0.5, color = "gray", linetype = 2) +
  587. geom_point(pch = 21, alpha = 1, size = 1.5, aes(fill = finaldecision)) +
  588. geom_point(data = subset(probs1, finaldecision == "Include" & new == 1),pch = 21, alpha = 1, size = 1.5, fill = 'red' ) +
  589. scale_fill_manual(values = c('black', 'red')) +
  590. geom_vline(xintercept = qlogis(0.4), linetype = 2) +
  591. coord_cartesian(ylim = c(0, 1)) +
  592. theme_minimal(base_size = 14) +
  593. labs(x = "logit(probability)", y = "Probability", fill = 'Known decisions') +
  594. theme(legend.position = "top") +
  595. facet_wrap(~new)
  596. ggplot(data = subset(results, new > 10 & new < 21), aes(x = (newpred))) +
  597. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  598. scale_fill_manual(values = c('black', 'red')) +
  599. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  600. facet_wrap(~new* configs, ncol = 2, scale = 'free')+
  601. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  602. theme_classic() +
  603. theme(legend.position = "top")
  604. ggplot(data = subset(results, new > 20 & new < 31), aes(x = (newpred))) +
  605. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  606. scale_fill_manual(values = c('black', 'red')) +
  607. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  608. facet_wrap(~new* configs, ncol = 2, scale = 'free')+
  609. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  610. theme_classic() +
  611. theme(legend.position = "top")
  612. ggplot(data = subset(results, new > 30 & new < 41), aes(x = (newpred))) +
  613. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  614. scale_fill_manual(values = c('black', 'red')) +
  615. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  616. facet_wrap(~new* configs, ncol = 2, scale = 'free')+
  617. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  618. theme_classic() +
  619. theme(legend.position = "top")
  620. ggplot(data = subset(results, new > 40 & new < 51), aes(x = (Include))) +
  621. geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
  622. scale_fill_manual(values = c('black', 'red')) +
  623. geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
  624. facet_wrap(~new* configs, ncol = 2, scale = 'free')+
  625. labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
  626. theme_classic() +
  627. theme(legend.position = "top")
  628. baked_df %>%
  629. mutate(scorecat = ifelse(totalscore*5 <3.5, 'low', 'high'))%>%
  630. group_by( finaldecision) %>%
  631. #filter(totalscore *5 >= 3.5) %>%
  632. summarise(min(totalscore * 5), max(totalscore*5))

Hypertension h2o.R at commit d69656c, no license · at the source

Overview

  1. Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, UK
  2. NIHR Bristol Evidence Synthesis Group, University of Bristol, Bristol, UK
  3. Department of Psychology, University of Bath, Bath, UK
  4. School of Health, Robert Gordon University, Aberdeen, UK
  5. Applied Physiology and Nutrition Research Group – School of Physical Education and Sport and Faculdade de Medicina FMUSP, Universidade de São Paulo, São Paulo, Brazil
  6. NIHR Applied Research Collaboration West (ARC West) at University Hospitals Bristol and Weston NHS Foundation Trust, Bristol, UK
Dates: published online 9 April 2026
Type: Preprint
License: CC BY-NC-ND
Identifiers: DOI 10.64898/2026.04.08.26350384 · OpenAlex W7152543262
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: clinical / translational (subfield)
Methods: Connectivity, Smoothing, state filtering, decompositions, Machine learning, Statistics
Topic: Artificial Intelligence in Healthcare and Education (Health Informatics, Medicine), according to OpenAlex
Funding: National Institute for Health Research (NIHR) (NIHR153861, NIHR203807, NIHR168894)
Citations: not cited yet (Europe PMC); 34 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.

gabsbarreto/JARVIS-R

License: none: the authors keep all their rights
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: d69656c9d5f763d24d4400d1be41eb74bf73a1e8, 28 July 2026
Languages: R (8)
Size: 18 files, 8 scripts
Software Heritage: not archived
Found in: the text, “Evaluation using archived systematic review data”
Holds: README, CITATION.cff
Not found: license file, environment file, tests, continuous integration, documentation
Tools: ggpubr (8 files), tidyverse (8 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
9 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 8 scripts, each with its path and the digest of its content;
  • 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, journal, dates, 8 authors, 1 funder, 25 references.

Cite

This paper

Barreto, G. H. C., Burke, C., Davies, P., Halicka, M., Paterson, C., Swinton, P., Saunders, B., & Higgins, J. (2026). JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews. medRxiv (preprint). https://doi.org/10.64898/2026.04.08.26350384

BibTeX

@article{barreto2026jarvis,
author = {Barreto, G. H. C. and Burke, C. and Davies, P and Halicka, M. and Paterson, C. and Swinton, P. and Saunders, B. and Higgins, J.P.T.},
title = {{JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews}},
journal = {medRxiv (preprint)},
year = {2026},
month = apr,
publisher = {medRxiv},
doi = {10.64898/2026.04.08.26350384},
url = {https://doi.org/10.64898/2026.04.08.26350384}
}

RIS

TY - JOUR
AU - Barreto, G. H. C.
AU - Burke, C.
AU - Davies, P
AU - Halicka, M.
AU - Paterson, C.
AU - Swinton, P.
AU - Saunders, B.
AU - Higgins, J.P.T.
TI - JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews
T2 - medRxiv (preprint)
J2 - medRxiv
PY - 2026
DA - 2026/04/09
PB - medRxiv
DO - 10.64898/2026.04.08.26350384
UR - https://doi.org/10.64898/2026.04.08.26350384
ER -

CSL-JSON

{
"id": "10.64898/2026.04.08.26350384",
"type": "article",
"title": "JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews",
"container-title": "medRxiv (preprint)",
"author": [
{
"family": "Barreto",
"given": "G. H. C."
},
{
"family": "Burke",
"given": "C."
},
{
"family": "Davies",
"given": "P"
},
{
"family": "Halicka",
"given": "M."
},
{
"family": "Paterson",
"given": "C."
},
{
"family": "Swinton",
"given": "P."
},
{
"family": "Saunders",
"given": "B."
},
{
"family": "Higgins",
"given": "J.P.T."
}
],
"container-title-short": "medRxiv",
"DOI": "10.64898/2026.04.08.26350384",
"publisher": "medRxiv",
"URL": "https://doi.org/10.64898/2026.04.08.26350384",
"issued": {
"date-parts": [
[
2026,
4,
9
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1128/msystems.00416-26 [code]
Integrative multicohort analysis reveals consistent sex differences in gut microbiota of multiple sclerosis patients.
Journal: mSystems
In common: ggpubr, tidyverse, clinical / translational, 1 reference
[2] doi:10.1093/geront/gnaf277 [code]
What characterizes the exceptional cognition of superagers? A systematic review of multidomain biomarkers of successful cognitive aging.
Journal: The Gerontologist
In common: ggpubr, tidyverse, clinical / translational, 1 reference
[3] doi:10.1007/s00415-026-14028-0 [code]
Personality change after traumatic brain injury: a systematic review and meta-analysis.
Journal: Journal of neurology
In common: tidyverse, clinical / translational, 1 reference
[4] doi:10.1093/aje/kwag135 [code]
Depressive symptoms and neuroimaging markers of brain aging in an ethno-racially diverse sample: a Bayesian analysis.
Journal: American journal of epidemiology
In common: ggpubr, tidyverse, clinical / translational
[5] doi:10.1093/brain/awag039 [code]
Mapping the causal chain from genetic risk variants to lipid dysmetabolism in Parkinson's disease.
Journal: Brain : a journal of neurology
In common: ggpubr, tidyverse, clinical / translational
[6] doi:10.1016/j.isci.2026.116517 [code]
Spatially resolved transcriptomics in human brain metastases identifies macrophage-tumor interactions associated with survival.
Journal: iScience
In common: ggpubr, tidyverse, clinical / translational
[7] doi:10.1126/sciadv.aee2305 [code]
Prediction of mild cognitive impairment progression using time-sensitive multimodal biomarkers.
Journal: Science advances
In common: ggpubr, tidyverse, clinical / translational
[8] doi:10.3389/fnins.2026.1858005 [code]
Targeted stool metabolomics suggests exploratory catecholamine- and tryptophan-linked metabolic features in autism spectrum disorder.
Journal: Frontiers in neuroscience
In common: ggpubr, tidyverse, clinical / translational
[9] doi:10.1038/s41380-026-03694-1 [code]
Targeting cortico-striatal-amygdalar networks via theta-band frontoparietal synchronization in opioid use disorder: a randomized tACS-fMRI Trial.
Journal: Molecular psychiatry
In common: ggpubr, tidyverse, clinical / translational
[10] doi:10.1186/s13073-026-01698-8 [code]
From aging to Alzheimer's disease: concordant brain DNA methylation changes in late life.
Journal: Genome medicine
In common: ggpubr, tidyverse, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.