JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews
The 4 matches
- [1] § Methods › Text data pre-processing › Step 4: Modelling ↔ scripts/Hypertension h2o.R, lines 452–540 · score 0.63 · h2o.deeplearning, cross validate, training samples, folds, classified, models
- [2] § Methods › Text data pre-processing › Step 4: Modelling ↔ scripts/Anticoag atrial fib h2o.R, lines 446–529 · score 0.63 · h2o.deeplearning, cross validate, training samples, folds, classified, models
- [3] § Methods › Text data pre-processing ↔ scripts/Anticoag 2 h2o.R, lines 239–321 · score 0.59 · TF IDF, PCA, SMART, ngrams, tokenised, words
- [4] § Methods › Text data pre-processing ↔ scripts/Sedat behaviour children h2o.R, lines 238–309 · score 0.59 · TF IDF, PCA, SMART, ngrams, tokenised, words
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 764 lines · 32 KB · no license · 1 match
- ## libraries ####
- library(tidyverse)
- library(recipes)
- library(textrecipes)
- library(h2o)
- library(future.apply)
- library(tibble)
- library(dplyr)
- library(ggpubr)
- library(PRROC)
- # The two functions below allowed me to do hyperparameter tuning.
- # The first one, ending in 'parallel', allows me to open several parallel sessions
- # and test each hyperparameter combination in parallel. Requires a TON of RAM, so
- # I use it carefully.
- # The second one, ending in repeats, will test each combination in sequence. Takes
- # longer, but it works.
- run_each5_with_repeats_parallel <- function(df, n, epochs, hiddenunis, activ, stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate, repeats, n_workers) {
- # remove contents of the h2o server if one is open already
- if (tryCatch({ h2o.clusterStatus(); TRUE }, error = function(e) FALSE)) {
- h2o.removeAll(timeout_secs = 0)
- }
- # 1. Build all param combinations × repeats
- param_combinations <- expand.grid(
- epochs = epochs,
- hidden_units = hiddenunis,
- activ = activ,
- stop_rounds = stop_rounds,
- stop_tol = stop_tol,
- rates_anneal = rates_anneal,
- min_batch = min_batch,
- l2 = l2,
- rate = rate,
- stringsAsFactors = FALSE
- )
- jobs <- do.call(rbind, lapply(seq_len(nrow(param_combinations)), function(i) {
- cbind(param_combinations[i, , drop=FALSE], repeat_run = seq_len(repeats), ID = i)
- }))
- jobs <- as.data.frame(jobs)
- # 2. Function to run a single job
- run_single <- function(job_row) {
- epochs <- as.numeric(job_row["epochs"])
- stop_rounds <- as.numeric(job_row["stop_rounds"])
- stop_tol <- as.numeric(job_row["stop_tol"])
- rates_anneal <- as.numeric(job_row["rates_anneal"])
- min_batch <- as.numeric(job_row["min_batch"])
- hidden_units <- as.character(job_row["hidden_units"])
- activ <- as.character(job_row["activ"])
- repeat_run <- as.numeric(job_row["repeat_run"])
- l2 <- as.numeric(job_row["l2"])
- rate <- as.numeric(job_row["rate"])
- ID <- as.numeric(job_row["ID"])
- set.seed(123 + ID + repeat_run) # More variability per run
- start_time <- Sys.time()
- res <- each5.2(df, n, epochs, hidden_units, activ, stop_rounds, stop_tol,rates_anneal,min_batch, l2, rate)
- end_time <- Sys.time()
- res %>%
- mutate(
- ID = ID,
- repeat_run = repeat_run,
- configs = paste(epochs, hidden_units, activ, stop_rounds, stop_tol ,rates_anneal,min_batch,l2,rate, sep = ", "),
- elapsed_time = as.numeric(difftime(end_time, start_time, units = "secs"))
- )
- }
- # 3. Set up parallel plan
- future::plan(multisession, workers = 12)
- plan()
- # 4. Run in parallel!
- results_list <- future.apply::future_lapply(
- split(jobs, seq_len(nrow(jobs))),
- function(job) run_single(job[1, ])
- )
- # 5. Combine all results into one tibble
- results <- dplyr::bind_rows(results_list)
- return(results)
- }
- run_each5_with_repeats <- function(df, n, epochs, hiddenunis, activ, stop_rounds, stop_tol,rates_anneal,min_batch, l2, rate, repeats) {
- # remove contents of the h2o server if one is open already
- if (tryCatch({ h2o.clusterStatus(); TRUE }, error = function(e) FALSE)) {
- h2o.removeAll(timeout_secs = 0)
- }
- elapsed_time <- numeric() # Initialize as numeric vector
- param_combinations <- expand.grid(
- epochs = epochs,
- hidden_units = hiddenunis,
- activ = activ,
- stop_rounds = stop_rounds,
- stop_tol = stop_tol,
- rates_anneal = rates_anneal,
- min_batch = min_batch,
- l2 = l2,
- rate = rate
- )
- # Initialize results as an empty tibble
- results <- tibble()
- total_runs <- nrow(param_combinations) * repeats # Total number of iterations
- current_run <- 0 # Initialize counter for completed runs
- # Iterate over each combination of parameters
- for (i in seq_len(nrow(param_combinations))) {
- print(paste(i, '/', nrow(param_combinations)))
- epochs <- param_combinations$epochs[i]
- hidden_units <- as.character(param_combinations$hidden_units[i])
- activ <- param_combinations$activ[i]
- stop_rounds <- param_combinations$stop_rounds[i]
- stop_tol <- param_combinations$stop_tol[i]
- rates_anneal <- param_combinations$rates_anneal[i]
- min_batch <- param_combinations$min_batch[i]
- l2 <- param_combinations$l2[i]
- rate <- param_combinations$rate[i]
- set.seed(123)
- for (repeat_idx in seq_len(repeats)) {
- current_run <- current_run + 1 # Increment the completed runs counter
- start_time <- Sys.time() # Record start time
- current_result <- each5.2(df, n, epochs, hidden_units, activ, stop_rounds, stop_tol, rates_anneal, min_batch, l2, rate) %>%
- mutate(
- ID = i,
- repeat_run = repeat_idx,
- configs = paste(epochs, hidden_units, activ, stop_rounds, stop_tol, rates_anneal, min_batch, l2, rate, sep = ", ")
- )
- # Append the current result to the results tibble
- results <- bind_rows(results, current_result)
- end_time <- Sys.time() # Record end time
- # Calculate elapsed time for this run
- elapsed_time <- c(elapsed_time, as.numeric(difftime(end_time, start_time, units = "secs")))
- mean_time <- mean(elapsed_time, na.rm = TRUE)
- remaining_runs <- total_runs - current_run
- estimated_remaining_time <- mean_time * remaining_runs
- # Convert estimated time to minutes and seconds
- estimated_minutes <- floor(estimated_remaining_time / 60)
- estimated_seconds <- round(estimated_remaining_time %% 60)
- # Print estimated time remaining
- print(paste(
- "Estimated remaining time:",
- estimated_minutes, "min", estimated_seconds, "sec"
- ))
- }
- }
- return(results)
- }
- ## ROC-PR FUNCTION####
- # calc_aucpr is a function designed to calculate aucpr for our analyses. It'll be used in the 'summaries' section.
- calc_aucpr <- function(df_iter) {
- # df_iter = subset of results for a single iteration
- # True labels as 1/0
- truth <- ifelse(df_iter$finaldecision == "Include", 1, 0)
- # Predicted scores (probabilities)
- scores <- df_iter$newpred
- # PRROC needs the scores of the positive class separately:
- pr <- pr.curve(
- scores.class0 = scores[truth == 0],
- scores.class1 = scores[truth == 1],
- curve = FALSE
- )
- return(pr$auc.integral)
- }
- ##hypertension data clean #####
- # In this section we read the file in 'data' containing our abstracts, decisions and the LLM responses.
- # We break the responses into columns, clean them, create the numberical variables (PICOS scores).
- dfwithPICOS <- read_rds('dfgpt.stress.complete.rds')
- dfPICOSfinal <- dfwithPICOS %>%
- separate_wider_delim(cols = GPT_Response, delim = ',', names = c('review', 'P', 'I', 'C', 'O', 'S'), cols_remove = F) %>%
- mutate(review = chartr("[],012345 '","...........", review),
- review = chartr("- ","..", review),
- review = gsub('[.]','',review),
- P = chartr("[],012345 '","...........", P),
- P = chartr("- ","..", P),
- P = gsub('[.]','',P),
- I = chartr("[],012345 '","...........", I),
- I = chartr("- ","..", I),
- I = gsub('[.]','',I),
- C = chartr("[],012345 '","...........", C),
- C = chartr("- ","..", C),
- C = gsub('[.]','',C),
- O = chartr("[],012345 '","...........", O),
- O = chartr("- ","..", O),
- O = gsub('[.]','',O),
- S = chartr("[],012345 '","...........", S),
- S = chartr("- ","..", S),
- S = gsub('[.]','',S)) %>%
- mutate(Pn = case_when(P =='Yes'~ 1, ### create numerical variables with PICOS score
- P == 'No' ~ 0,
- P == 'Uncertain' ~ 0.5),
- In = case_when(I =='Yes'~ 1,
- I == 'No' ~ 0,
- I == 'Uncertain' ~ 0.5),
- Cn = case_when(C =='Yes'~ 1,
- C == 'No' ~ 0,
- C == 'Uncertain' ~ 0.5),
- On = case_when(O =='Yes'~ 1,
- O == 'No' ~ 0,
- O == 'Uncertain' ~ 0.5),
- Sn = case_when(S =='Yes'~ 1,
- S == 'No' ~ 0,
- S == 'Uncertain' ~ 0.5)) %>%
- dplyr::select(key, keywords, title, abstract, authors, year, review, P, I, C, O, S, Pn, In,Cn, On, Sn, decision, finaldecision) %>%
- mutate(totalscore = rowSums(.[c('Pn','In','Cn','On','Sn')])) %>%
- filter(key != 'ardiologie et Pneumologie de Quebec' & key != 'iomedical Science') %>%
- distinct(key, .keep_all = T)
- ## PREPARE DATA #####
- # In this section we clean the text variables and prepare them for tokenisation and TF-IDF generation.
- dftoken <- dfPICOSfinal %>%
- select(key, title, authors, abstract, totalscore, review, P, I, C, O, S,Pn, In,Cn, On, Sn, decision, finaldecision) %>%
- mutate(abstractsub = gsub('-', ' ', abstract)) %>% # Remove hyphens
- mutate(abstractsub = iconv(abstractsub, from = "", to = "ASCII//TRANSLIT")) %>% # Remove/convert weird chars
- mutate(abstractsub = gsub("–|—|â€|’|“|‘|•|−", " ", abstractsub)) %>% # Remove common mojibake
- mutate(abstractsub = gsub('[[:punct:]]', '', abstractsub)) %>% # Remove punctuation
- mutate(abstractsub = tolower(abstractsub)) %>%
- mutate(abstractsub = gsub('[0-9]+', ' ', abstractsub)) %>%
- mutate(abstractsub = gsub("\\s+", " ", abstractsub)) %>% # Normalize whitespace
- mutate(titlesub = gsub('-', ' ', title)) %>% # Remove hyphens
- mutate(titlesub = iconv(titlesub, from = "", to = "ASCII//TRANSLIT")) %>% # Remove/convert weird chars
- mutate(titlesub = gsub("–|—|â€|’|“|‘|•|−", " ", titlesub)) %>% # Remove common mojibake
- mutate(titlesub = gsub('[[:punct:]]', '', titlesub)) %>% # Remove punctuation
- mutate(titlesub = tolower(titlesub)) %>%
- mutate(titlesub = gsub('[0-9]+', ' ', titlesub)) %>%
- mutate(titlesub = gsub("\\s+", " ", titlesub)) %>% # Normalize whitespace
- mutate_at(vars(P, I, C, O, S), funs(factor(., levels = c('Uncertain', 'Yes', 'No') ))) %>%
- mutate(reviewn = case_when(review == 'No' ~ 1,
- review == 'Yes' ~ 0 ))
- rm(dfPICOSfinal)
- rm(dfwithPICOS)
- tokenization_recipe <- recipe(decision ~ key + abstract + authors + abstractsub + P + I + C + O + S + totalscore + review + finaldecision, data = dftoken) %>%
- update_role(key, new_role = "id") %>% #Mark 'key' as an ID
- update_role(abstract, new_role = "id") %>%
- update_role(authors, new_role = "id") %>%
- update_role(finaldecision, new_role = "id") %>%
- step_tokenize(abstractsub, token = "words") %>%
- step_stopwords(abstractsub, stopword_source = "smart") %>%
- #step_stem(abstractsub) %>%
- step_ngram(abstractsub, num_tokens = 3, min_num_tokens = 1) %>%
- step_tokenfilter(abstractsub, max_tokens = 7000) %>%
- step_tfidf(abstractsub) %>%
- #step_pca(starts_with('tfidf'),threshold = .80, prefix = 'KPCa') %>%
- #step_tokenize(titlesub, token = "words") %>%
- #step_stopwords(titlesub, stopword_source = "smart") %>%
- #step_stem(titlesub) %>%
- #step_ngram(titlesub, num_tokens = 3, min_num_tokens = 1) %>%
- #step_tokenfilter(titlesub, max_tokens = 3000 ) %>%
- #step_tfidf(titlesub) %>%
- #step_pca(c(reviewn, totalscore,keys1, score1), num_comp = 1, keep_original_cols = T) %>%
- step_pca(starts_with('tfidf'),threshold = .90, prefix = 'KPCa') %>%
- step_center(starts_with('KPC')) %>%
- step_range(totalscore, min = 0, max = 1) %>%
- step_dummy( all_of(c('P', 'I', 'C','O','S', 'review')), one_hot = T)
- prepped_recipe <- prep(tokenization_recipe, training = dftoken, retain = TRUE) # Preprocess and retain
- baked_df <- bake(prepped_recipe, new_data = NULL)
- dir.create('baked', showWarnings = F, recursive = T)
- saveRDS(baked_df, 'baked/hypertension final.rds')
- ##START FROM HERE WITH THE BAKED READY####
- # Sometimes, 'baking' the dataframe takes some time depending on your hardware. I preferred to save
- # the dataframe and reload it here below. Also this is where I include the weights.
- baked_df <- read_rds('baked/hypertension final.rds') %>%
- mutate(authors = ifelse(is.na(authors),'no author',authors )) %>%
- mutate(finaldecision = as.factor(finaldecision)) %>%
- mutate(weightsc = ifelse(decision == "Include", 40,1))
- ggplot(baked_df, aes( x = totalscore * 5)) +
- geom_histogram( aes(fill = finaldecision, y = ..density..)) +
- theme_classic() +
- theme(legend.position = "top") +
- labs(x = 'PICOS score', fill = "FT decision")
- ## new function 5 each time #####
- # each5.2 performs iterative active-screening with an H2O deep-learning classifier.
- # Workflow per round:
- # 1) Select new records to review (top/bottom scoring records plus a small buffer),
- # then add them to the cumulative labeled sample (`samplefull`).
- # 2) Build training folds by repeating Include cases across folds and sampling Exclude
- # cases per fold to keep class balance for cross-validation.
- # 3) Train a Bernoulli H2O deep-learning model with the provided hyperparameters.
- # 4) Score unscreened records with each CV model, average Include probabilities,
- # normalize to `newpred`, and assign Include/Exclude labels using `threshold`.
- # 5) Save per-round predictions and error-tracking fields, and mark low-probability
- # records as candidates to exclude in later sampling.
- # Early stopping: after the minimum rounds, stop when too few records remain above
- # threshold for consecutive rounds. Returns one data frame with all rounds combined.
- each5.2 <- function(df1, max, epoc, hidu, activ, stop_rounds, stop_tol, rates_anneal,min_batch, l2, rate) {
- # Initialize variables for Stage 2
- results <- vector("list", max)
- exclude_keys <- data.frame(NA)
- samplefull <- data.frame()
- preds <- df1
- min_rounds <- 71
- prefix <- paste0("worker", Sys.getpid(), "_", as.integer(Sys.time()), "_")
- threshold <- 0.4
- min_hold <- 1
- since_change <- min_hold
- for (i in seq_len(max)) {
- time1 <- Sys.time()
- if(i > 1) {
- # inc_rate <- (include_count + exclude_count) / nrow(df1)
- # print(inc_rate)
- #
- # cap <- dplyr::case_when(
- # #inc_rate > 0.30 ~ 0.75,
- # #inc_rate > 0.25 ~ 0.70,
- # #inc_rate > 0.20 ~ 0.60,
- # #inc_rate > 0.15 ~ 0.55,
- # inc_rate > 0.10 ~ 0.5,
- # inc_rate > 0.05 ~ 0.45,
- # TRUE ~ threshold
- # )
- #
- #
- # old <- threshold
- #
- # if (since_change >= min_hold && cap > threshold) {
- # threshold <- cap # jump directly to the mapped cap
- # since_change <- 1 # or set to 1 if you prefer counting this iter as "held"
- # } else {
- # since_change <- since_change + 1
- # }
- # print(paste('threshold =', threshold))
- keys_to_exclude <- exclude_keys$key
- keys_in_sample <- samplefull$key
- preds <- preds %>%
- suppressMessages(left_join(.,df1))
- sampled_df <- preds %>%
- ungroup() %>%
- filter(!key %in% keys_in_sample) %>%
- filter(!key %in% keys_to_exclude) %>%
- arrange(desc(Include)) %>%
- slice_head(n=15) %>%
- bind_rows(preds %>%
- ungroup() %>%
- filter(!key %in% keys_in_sample) %>%
- filter(!key %in% keys_to_exclude) %>%
- arrange(desc(Include)) %>%
- slice_tail(n=8) )%>%
- bind_rows(preds %>%
- ungroup() %>%
- filter(!key %in% keys_in_sample) %>%
- filter(key %in% keys_to_exclude) %>%
- arrange(desc(Include)) %>%
- slice_head(n=7)) %>%
- select(-predict, -Include, -Exclude, -new, -incpred, -excpred, -newpred, - thresh)
- }else{
- sampled_df <- preds %>%
- arrange(desc(totalscore)) %>%
- slice(c((nrow(.)-4):nrow(.), 1:20)) %>%
- bind_rows(preds %>%
- arrange(desc(totalscore)) %>%
- filter(totalscore*5 <= 2.5) %>% slice(1:5))
- localH2O = h2o.init(ip="localhost", port = 54321,
- startH2O = TRUE, nthreads=1, max_mem_size = '8G')
- prefix <- paste0("worker", Sys.getpid(), "_", as.integer(Sys.time()), "_")
- }
- samplefull <- samplefull %>%
- bind_rows(sampled_df %>% mutate(iter = i)) %>%
- distinct()
- includes <- samplefull %>% filter(decision == "Include")
- excludes <- samplefull %>% filter(decision == "Exclude")
- n_excl_per_fold <- nrow(includes) # or set any number you prefer per fold
- # 1) 3 repeats of each Include, labeled folds 1, 2, 3
- includes3 <- includes %>%
- mutate(.k = 1L) %>%
- left_join(data.frame(folds = 1:3, .k = 1L), by = ".k",relationship = "many-to-many") %>%
- select(-.k)
- # 2) Randomize all Excludes and partition into 3 (no overlap / no replacement)
- excl_folds <- lapply(1:3, function(f) {
- excludes %>%
- slice_sample(n = min(n_excl_per_fold, nrow(.)), replace = FALSE) %>%
- mutate(folds = f)
- }) %>% bind_rows()
- # 3) Combine
- samplefull2 <- bind_rows(includes3, excl_folds)
- h2otrain <- h2o.na_omit(as.h2o(samplefull2, destination_frame = paste0(prefix, 'train1',i)))
- test <- df1 %>%
- filter(!(key %in% samplefull$key))
- h2otest <- h2o.na_omit(as.h2o(test, destination_frame = paste0(prefix, 'preddata',i)))
- h2otest <- h2o.na_omit(h2otest)
- include_count <- samplefull %>%
- filter(decision == "Include") %>%
- nrow()
- exclude_count <- samplefull %>%
- filter(decision == "Exclude") %>%
- nrow()
- # Update the recipe to use RELEVANCE as the outcome
- y <- "decision"
- x <- names(df1)[c( 7:(ncol(df1)-1))]
- model = h2o.deeplearning(x=x,
- y=y,
- training_frame=h2otrain,
- distribution = "bernoulli",
- activation = as.character(activ),
- hidden = eval(parse(text = hidu)),
- epochs = epoc, l1 = 0, l2 = l2,rate = rate,
- loss = "CrossEntropy",
- initial_weight_distribution = c("UniformAdaptive"),
- ignore_const_cols = F, seed = 123,
- adaptive_rate = F, standardize = F,
- nesterov_accelerated_gradient = T,
- fast_mode = T,
- reproducible = T,
- overwrite_with_best_model = T,
- score_training_samples = 500, score_validation_samples = 500,
- stopping_metric = "mean_per_class_error",classification_stop = -1,
- stopping_rounds = stop_rounds, stopping_tolerance = stop_tol,
- model_id = paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate, sep = "_" ),
- #nfolds = 3, fold_assignment = 'Stratified',
- fold_column = 'folds',
- rate_annealing = rates_anneal,# rate_decay = 0.8,
- keep_cross_validation_models = T,
- keep_cross_validation_predictions = T, weights_column = 'weightsc',
- mini_batch_size = min_batch#, balance_classes = T,
- #max_after_balance_size = max(.1,include_count/exclude_count),
- #class_sampling_factors = c(1, exclude_count / include_count)
- )
- modelcv1 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_1',sep = '_'))
- predweird1 <- predict(modelcv1, newdata = h2otest )
- predweird1.2 <- bind_cols(test, as.data.frame(predweird1)) %>%
- mutate(pp = 1) %>%
- select(-starts_with('KPC'), -contains('_'))
- modelcv2 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_2',sep = '_'))
- predweird2 <- predict(modelcv2, newdata = h2otest )
- predweird2.2 <- bind_cols(test, as.data.frame(predweird2)) %>%
- mutate(pp = 2) %>%
- select(-starts_with('KPC'), -contains('_'))
- modelcv3 <- h2o.getModel(paste(i, epoc, as.character(activ), stop_rounds, stop_tol, rates_anneal,min_batch,l2,rate,'cv_3',sep = '_'))
- predweird3 <- predict(modelcv3, newdata = h2otest )
- predweird3.2 <- bind_cols(test, as.data.frame(predweird3)) %>%
- mutate(pp = 3) %>%
- select(-starts_with('KPC'), -contains('_'))
- all <- bind_rows(predweird1.2,
- predweird2.2,
- predweird3.2) %>%
- group_by(key, decision, finaldecision) %>%
- summarise_at(vars(Include), funs(mean(., na.rm =T))) %>%
- ungroup() %>%
- mutate(Exclude = 1-Include,
- newpred = (Include - min(Include)) / (max(Include) - min(Include) ),
- predict = ifelse(newpred >= 0.5, "Include", "Exclude"))
- preds <- test %>%
- merge(.,all, .by = key) %>%
- mutate(new = i) %>%
- mutate(incpred = include_count, # Add the "Include" count
- excpred = exclude_count) %>%
- mutate(thresh = threshold)
- my_keys <- h2o.ls()[,1]
- h2o.rm(my_keys[grepl(paste0("^", prefix), my_keys)], cascade = TRUE)
- gc()
- time2 <- Sys.time()
- exclude_keys <- preds %>%
- ungroup() %>% filter(Include < threshold) %>%
- select(key, new)
- # Add the "Exclude" count
- results[[i]] <- preds %>%
- select(-starts_with('KPC'), -starts_with('tfidf'), -starts_with(c('P_','I_', 'C_', 'O_', 'S_')))
- sumtextPICOS <- (bind_rows(results)) %>%
- mutate(newclass = ifelse(Include < thresh,"Exclude", "Include")) %>%
- mutate(inc.correct = ifelse(finaldecision == "Include" & newclass == "Include",1,0) ) %>%
- mutate(exc.correct = ifelse(finaldecision == "Exclude" & newclass == "Exclude",1,0) ) %>%
- mutate(exc.incorrect = ifelse(finaldecision == "Include" & newclass == "Exclude",1,0) ) %>%
- mutate(inc.incorrect = ifelse(finaldecision == "Exclude" & newclass == "Include",1,0) ) %>%
- group_by(., new) %>% #, thresh) %>%
- summarise_at(vars(inc.correct,exc.correct,inc.incorrect, exc.incorrect), funs (sum(.,na.rm= T))) %>%
- merge(.,bind_rows(results) %>% group_by(., new) %>%
- summarise_at(vars(incpred, excpred), funs(max(.))), .by = c('configs', 'new')) %>%
- arrange(-desc(new))
- print(ggarrange(nrow = 2,ncol = 1, ggplot(data = sumtextPICOS, aes(x = new, y = inc.incorrect)) +
- geom_line(colour = 'blue') +
- geom_line(colour = 'red', aes(y = exc.incorrect)) +
- geom_point(pch = 21, colour = 'blue') +
- geom_point(pch = 21, colour = 'red', aes(y = exc.incorrect)) +
- geom_text(vjust = -0.5, size = 3 ,aes(y = exc.incorrect, label = exc.incorrect))+
- geom_text(vjust = -0.5, size = 3 ,aes(y = inc.incorrect, label = inc.incorrect))+
- #facet_wrap(~configs) +
- #scale_x_continuous(breaks = c(0,2,4,6,8,10)) +
- scale_y_continuous(limits = c(0,max(sumtextPICOS$inc.incorrect)),name = 'Blue = Included Incorrectly', sec.axis = sec_axis(~. , name = "Red = Excluded Incorrectly")) +
- labs(x = 'Rounds') +
- theme_classic(),
- ggplot(data = subset(preds), aes(x = (Include))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- #facet_wrap(~new* configs, ncol = 2, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top")))
- if (min_rounds <= i && sum(preds$Include >= threshold, na.rm = T) <= max(30, nrow(df1) / 100)) {
- return(bind_rows(results))
- break
- }
- gc()
- print((time2 - time1))
- print(i)
- }
- return(bind_rows(results))
- }
- ##run it here#####
- # This block sets one hyperparameter configuration and runs the full
- # iterative simulation pipeline.
- # - The vectors below are the model settings explored by
- # `run_each5_with_repeats` (single values here = one configuration).
- # - `TIMES` controls how many full repeats are executed.
- # - The returned object `results` contains per-round predictions/metrics.
- # - Results are persisted to `results/resultsanticoag2.rds` so the
- # summarisation section can be rerun without retraining.
- epochs <- c(100)
- hiddenunis <- c('c(50,25,10,5)')
- activ <- c('Tanh')
- stop_rounds <- c(3)
- stop_tol <- c(1e-5)
- rates_anneal <- c(0.001)
- min_batch <- c(1)
- l2 <- c(0.02575)
- rate <- c(0.001)
- TIMES <- 1
- results <- run_each5_with_repeats(baked_df, 71, epochs,hiddenunis,activ, stop_rounds, stop_tol,rates_anneal, min_batch,l2,rate, TIMES)
- dir.create('results', showWarnings = F, recursive = T)
- saveRDS(results, 'results/resultshypertension.rds')
- ##summarise#####
- results <- read_rds('results/resultshypertension.rds')
- EXCINC <- results %>%
- group_by(configs) %>%
- filter(finaldecision == "Include" & newpred < thresh) %>%
- select(-starts_with(c('P_','I_', 'C_', 'O_', 'S_', 'KPC', 'review_')))
- sumtextPICOS <- results %>%
- mutate(newclass = ifelse(Include < thresh,"Exclude", "Include")) %>%
- mutate(inc.correct = ifelse(finaldecision == "Include" & newclass == "Include",1,0) ) %>%
- mutate(exc.correct = ifelse(finaldecision == "Exclude" & newclass == "Exclude",1,0) ) %>%
- mutate(exc.incorrect = ifelse(finaldecision == "Include" & newclass == "Exclude",1,0) ) %>%
- mutate(inc.incorrect = ifelse(finaldecision == "Exclude" & newclass == "Include",1,0) ) %>%
- group_by(., configs, new, ID, thresh) %>% #, thresh) %>%
- summarise_at(vars(inc.correct,exc.correct,inc.incorrect, exc.incorrect), funs (sum(.,na.rm= T))) %>%
- merge(.,results %>% group_by(., configs, new, ID,thresh) %>%
- summarise_at(vars(incpred, excpred), funs(max(.))), .by = c('configs', 'new')) %>%
- mutate(reads = incpred+ excpred) %>%
- mutate(percread = (incpred + excpred )/ nrow(baked_df) * 100 ,
- percsave = exc.correct/nrow(baked_df) * 100,
- ratio = incpred/excpred * 100) %>%
- merge(.,results %>%
- group_by(configs, new, ID,thresh) %>%
- summarise(
- aucpr = calc_aucpr(cur_data())
- ),.by = new) %>%
- merge(.,results %>%
- group_by(new, finaldecision, configs) %>%
- summarise(foundft = (193 - n() ) / 193 * 100) %>%
- filter(finaldecision == "Include") %>%
- select(-finaldecision))%>%
- mutate(recall = inc.correct / (inc.correct + exc.incorrect) * 100,
- recall_cumm = (sum(baked_df$finaldecision == "Include") - exc.incorrect ) / ((sum(baked_df$finaldecision == "Include") - exc.incorrect ) + exc.incorrect) * 100,
- specificity = exc.correct / (exc.correct + inc.incorrect) * 100) %>%
- arrange(-desc(new)) ;View(sumtextPICOS)
- calclong <- sumtextPICOS %>%
- mutate(aucpr = aucpr*100) %>%
- bind_rows(data.frame(new = 0, percread = 0, specificity = 0, foundft =0, recall=0, recall_cumm = 0)) %>%
- pivot_longer(cols =c( specificity, foundft, recall, recall_cumm)) %>%
- mutate(name = factor(name))
- ggplot(data = sumtextPICOS, aes(x = (new*30 / nrow(baked_df)* 100), y = (inc.incorrect + inc.correct ))) +
- geom_line(colour = 'blue') +
- geom_line(colour = 'red', aes(y = exc.incorrect)) +
- geom_point(pch = 21, colour = 'blue') +
- geom_point(pch = 21, colour = 'red', aes(y = exc.incorrect)) +
- geom_text(vjust = -0.5, size = 3 ,aes(y = exc.incorrect, label = exc.incorrect))+
- geom_text(vjust = -0.5, size = 3 ,aes(y = inc.incorrect + inc.correct , label = inc.incorrect + inc.correct ))+
- scale_y_continuous(name = 'Blue = Suggested Includes', sec.axis = sec_axis(~. , name = "Red = Excluded Incorrectly")) +
- labs(x = '% of studies read') +
- theme_classic() +
- facet_wrap(~configs)
- ggplot(data = calclong, aes(x = percread, y = value)) +
- geom_line(aes(colour = name), size = 1, alpha= 0.5) +
- geom_abline(slope = 0, intercept =100, linetype=3) +
- scale_y_continuous(limits = c(0,100),name = 'Performance value') +
- labs(x = '% of studies read', colour = "Metric") +
- scale_color_manual(labels = c("% of includes identified", "Iteration Recall" ,"Joint recall", "Specificity"), values= c('red', 'blue', 'green', 'purple'))+
- scale_x_continuous(expand=c(0, 0)) +
- geom_vline(xintercept = 28.2030620, linetype = 2, colour = 'orange')+
- geom_vline(xintercept = 22.9653505, linetype = 3, colour = 'gray20')+
- theme_classic()
- #ggsave('Hypertension.png', width = 6, height = 3, dpi = 300)
- summary(sumtextPICOS$aucpr)
- ## Histograms ####
- ggplot(data = subset(results, new== 1| new == 35 | new == 70), aes(x = (Include))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- facet_wrap(~new ,ncol = 3, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top") + scale_x_continuous(limits = c(0,1))
- #ggsave('hypertension distrib.png', width = 10, height = 3.3, dpi = 300)
- probs1 <- results%>%
- arrange((newpred)) %>%
- mutate(p = pmin(pmax(newpred, 1e-6), 1 - 1e-6)) %>%
- mutate(x = qlogis(p))
- ggplot(subset(probs1, new == 1), aes(x, p)) +
- stat_function(fun = plogis, xlim = range(probs1$x), size = 0.5, color = "gray", linetype = 2) +
- geom_point(pch = 21, alpha = 1, size = 1.5, aes(fill = finaldecision)) +
- geom_point(data = subset(probs1, finaldecision == "Include" & new == 1),pch = 21, alpha = 1, size = 1.5, fill = 'red' ) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(xintercept = qlogis(0.4), linetype = 2) +
- coord_cartesian(ylim = c(0, 1)) +
- theme_minimal(base_size = 14) +
- labs(x = "logit(probability)", y = "Probability", fill = 'Known decisions') +
- theme(legend.position = "top") +
- facet_wrap(~new)
- ggplot(data = subset(results, new > 10 & new < 21), aes(x = (newpred))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- facet_wrap(~new* configs, ncol = 2, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top")
- ggplot(data = subset(results, new > 20 & new < 31), aes(x = (newpred))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- facet_wrap(~new* configs, ncol = 2, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top")
- ggplot(data = subset(results, new > 30 & new < 41), aes(x = (newpred))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- facet_wrap(~new* configs, ncol = 2, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top")
- ggplot(data = subset(results, new > 40 & new < 51), aes(x = (Include))) +
- geom_histogram( aes(fill = finaldecision,y = after_stat(density))) +
- scale_fill_manual(values = c('black', 'red')) +
- geom_vline(aes(xintercept = thresh), colour = 'blue', linetype = 2) +
- facet_wrap(~new* configs, ncol = 2, scale = 'free')+
- labs(x = "Predicted Probability ", y = "Probability density",fill = 'Known decision') +
- theme_classic() +
- theme(legend.position = "top")
- baked_df %>%
- mutate(scorecat = ifelse(totalscore*5 <3.5, 'low', 'high'))%>%
- group_by( finaldecision) %>%
- #filter(totalscore *5 >= 3.5) %>%
- summarise(min(totalscore * 5), max(totalscore*5))
Hypertension h2o.R at commit d69656c, no license · at the source
Overview
- Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, UK
- NIHR Bristol Evidence Synthesis Group, University of Bristol, Bristol, UK
- Department of Psychology, University of Bath, Bath, UK
- School of Health, Robert Gordon University, Aberdeen, UK
- Applied Physiology and Nutrition Research Group – School of Physical Education and Sport and Faculdade de Medicina FMUSP, Universidade de São Paulo, São Paulo, Brazil
- NIHR Applied Research Collaboration West (ARC West) at University Hospitals Bristol and Weston NHS Foundation Trust, Bristol, UK
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.
gabsbarreto/JARVIS-R
d69656c9d5f763d24d4400d1be41eb74bf73a1e8, 28 July 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
9 files
- requirements.R, R, 22 lines
- scripts/
Anticoag 2 h2o.R , R, 703 lines, 1 match - scripts/
Anticoag atrial fib h2o.R , R, 731 lines, 1 match - scripts/
CYP1A2 h2o.R , R, 727 lines - scripts/
Hypertension h2o.R , R, 764 lines, 1 match - scripts/
PO-UR h2o.R , R, 706 lines - scripts/
Sedat behaviour children h2o.R , R, 707 lines, 1 match - scripts/
cocaine h2o.R , R, 709 lines - README.md, Text, 88 lines
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 8 scripts, each with its path and the digest of its content;
- 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, journal, dates, 8 authors, 1 funder, 25 references.
Cite
This paper
Barreto, G. H. C., Burke, C., Davies, P., Halicka, M., Paterson, C., Swinton, P., Saunders, B., & Higgins, J. (2026). JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews. medRxiv (preprint). https://
BibTeX
@article{barreto2026jarv
author = {Barreto, G. H. C. and Burke, C. and Davies, P and Halicka, M. and Paterson, C. and Swinton, P. and Saunders, B. and Higgins, J.P.T.},
title = {{JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews}},
journal = {medRxiv (preprint)},
year = {2026},
month = apr,
publisher = {medRxiv},
doi = {10.64898/
url = {https://
}
RIS
TY - JOUR
AU - Barreto, G. H. C.
AU - Burke, C.
AU - Davies, P
AU - Halicka, M.
AU - Paterson, C.
AU - Swinton, P.
AU - Saunders, B.
AU - Higgins, J.P.T.
TI - JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews
T2 - medRxiv (preprint)
J2 - medRxiv
PY - 2026
DA - 2026/
PB - medRxiv
DO - 10.64898/
UR - https://
ER -
CSL-JSON
{
"id": "10.64898/
"type": "article",
"title": "JARVIS, should this study be selected for full-text screening? Performance of a Joint AI-ReViewer Interactive Screening tool for systematic reviews",
"container-title": "medRxiv (preprint)",
"author": [
{
"family": "Barreto",
"given": "G. H. C."
},
{
"family": "Burke",
"given": "C."
},
{
"family": "Davies",
"given": "P"
},
{
"family": "Halicka",
"given": "M."
},
{
"family": "Paterson",
"given": "C."
},
{
"family": "Swinton",
"given": "P."
},
{
"family": "Saunders",
"given": "B."
},
{
"family": "Higgins",
"given": "J.P.T."
}
],
"container-title-short":
"DOI": "10.64898/
"publisher": "medRxiv",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
4,
9
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1128/msystems.00416-26 [code]
- Integrative multicohort analysis reveals consistent sex differences in gut microbiota of multiple sclerosis patients.Journal: mSystemsIn common: ggpubr, tidyverse, clinical / translational, 1 reference
- [2] doi:10.1093/geront/gnaf277 [code]
- What characterizes the exceptional cognition of superagers? A systematic review of multidomain biomarkers of successful cognitive aging.Journal: The GerontologistIn common: ggpubr, tidyverse, clinical / translational, 1 reference
- [3] doi:10.1007/s00415-026-14028-0 [code]
- Personality change after traumatic brain injury: a systematic review and meta-analysis.Journal: Journal of neurologyIn common: tidyverse, clinical / translational, 1 reference
- [4] doi:10.1093/aje/kwag135 [code]
- Depressive symptoms and neuroimaging markers of brain aging in an ethno-racially diverse sample: a Bayesian analysis.Journal: American journal of epidemiologyIn common: ggpubr, tidyverse, clinical / translational
- [5] doi:10.1093/brain/awag039 [code]
- Mapping the causal chain from genetic risk variants to lipid dysmetabolism in Parkinson's disease.Journal: Brain : a journal of neurologyIn common: ggpubr, tidyverse, clinical / translational
- [6] doi:10.1016/j.isci.2026.116517 [code]
- Spatially resolved transcriptomics in human brain metastases identifies macrophage-tumor interactions associated with survival.Journal: iScienceIn common: ggpubr, tidyverse, clinical / translational
- [7] doi:10.1126/sciadv.aee2305 [code]
- Prediction of mild cognitive impairment progression using time-sensitive multimodal biomarkers.Journal: Science advancesIn common: ggpubr, tidyverse, clinical / translational
- [8] doi:10.3389/fnins.2026.1858005 [code]
- Targeted stool metabolomics suggests exploratory catecholamine- and tryptophan-linked metabolic features in autism spectrum disorder.Journal: Frontiers in neuroscienceIn common: ggpubr, tidyverse, clinical / translational
- [9] doi:10.1038/s41380-026-03694-1 [code]
- Targeting cortico-striatal-amygdal
ar networks via theta-band frontoparietal synchronization in opioid use disorder: a randomized tACS-fMRI Trial. Journal: Molecular psychiatryIn common: ggpubr, tidyverse, clinical / translational - [10] doi:10.1186/s13073-026-01698-8 [code]
- From aging to Alzheimer's disease: concordant brain DNA methylation changes in late life.Journal: Genome medicineIn common: ggpubr, tidyverse, clinical / translational
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 8 scripts, and 4 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:334ea0942a9baac2…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
