CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling.
The 17 matches · 3 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
- [1] § Material and methods › RNA-seq data analysis ↔ RNA-seq/runSTAR.sh, the whole file · a weak match · score 0.98 · GeneCounts, alignEndsType, alignIntronMax, outFilterMismatchNoverLmax, outSAMunmapped, quantMode
- [2] § Material and methods › Enrichment analysis ↔ RNA-seq/RNAseq.analysis/R/read_pathway_db.R, the whole file · a weak match · score 0.85 · Human Phenotype Ontology, canonical pathways, transcription factor, mSigDB, SynGO, hs
- [3] § Material and methods › ATAC-seq analysis › TF footprint ↔ ATAC-seq/Transcription_factor_footprint/5. call_TF_bindings.sh, the whole file · a weak match · score 0.80 · TF footprint, TF binding, merged peak, ATAC seq, transcription factor, motifs
- [4] § Material and methods › Co-expression analysis ↔ RNA-seq/co-expression_analysis_allsamples.R, lines 53–95 · score 0.75 · Soft power, Module membership, Co expression, eigengene, outlier, network
- [5] § Material and methods › ATAC-seq analysis › Data processing ↔ atac-seq-pipeline-1.5.4.zip/src/encode_task_qc_report.py, lines 277–421 · score 0.75 · alignment quality, GC bias, fragment length, PCR, duplicates, mitochondria
- [6] § Material and methods › ATAC-seq analysis › Differential accessibility ↔ ATAC-seq/Differential_accessibility/diff_peaks_v2.R, lines 91–153 · score 0.74 · dba.contrast, model design, Treatment, peaks, DARs, accessibility
- [7] § Material and methods › ATAC-seq analysis › Direct and indirect POGZ regulation ↔ ATAC-seq/Transcription_factor_footprint/tf_binding.Rmd, lines 181–248 · score 0.74 · predicted bound, log2 fold change, binding site, DTF, DEL, promoter
- [8] § Material and methods › Co-expression analysis ↔ RNA-seq/RNAseq.analysis/R/co_expression_functions.R, lines 125–159 · score 0.74 · Soft power, scale free topology, Co expression, fit, outlier, network
- [9] § Material and methods › ATAC-seq analysis › Data processing ↔ atac-seq-pipeline-1.1.7.zip/src/run_ataqc.py, lines 1470–1518 · score 0.74 · GC bias, naive overlap, overlapped peaks, duplicates, quality, fraction
- [10] § Material and methods › ATAC-seq analysis › Differential accessibility ↔ ATAC-seq/Differential_accessibility/differential_accessibility.Rmd, lines 72–121 · score 0.68 · dba.contrast, DiffBind, Treatment, DARs, accessibility, DEL
- [11] § Material and methods › ATAC-seq analysis › TF footprint ↔ ATAC-seq/Transcription_factor_footprint/tf_binding.Rmd, lines 181–248 · score 0.67 · TF binding, transcription factor, bound, footprint, predict, motifs
- [12] § Material and methods › Enrichment analysis ↔ RNA-seq/RNAseq.analysis/R/enrichment.R, lines 53–123 · score 0.60 · co expression, KEGG, REACTOME, SynGO, Ontology, enrichments
- [13] § Material and methods › RNA-seq data analysis ↔ RNA-seq/rnaseq_analysis_iN_final.Rmd, lines 32–56 · score 0.58 · wt c6, gene annotations, RNA seq, STAR, Ensembl, alignments
- [14] § Material and methods › ATAC-seq analysis › Direct and indirect POGZ regulation ↔ ATAC-seq/Differential_accessibility/differential_accessibility.Rmd, lines 393–464 · score 0.57 · log2 fold change, POGZ target, cortex, DEL, binding, regulated
- [15] § Results › POGZ direct regulatory targets modulated in edited NSCs are associated with chromatin regulation, cell cycle, and RNA processing ↔ ATAC-seq/Differential_accessibility/differential_accessibility.Rmd, lines 393–464 · score 0.57 · NSC DEGs, overlapping genes, POGZ target genes, binding, regulation, enrichment
- [16] § Material and methods › Differential expression analysis ↔ RNA-seq/RNAseq.analysis/R/rnaseq_analysis_functions.R, lines 623–684 · score 0.56 · DESeq2, RNA seq, weighted, SVA, meta, scores
- [17] § Results › Indirect effects are mediated by other TFs ↔ ATAC-seq/Transcription_factor_footprint/tf_binding.Rmd, lines 147–179 · score 0.53 · TF footprint, TF binding, probability, motif, scored, seq
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 896 lines · 64 KB · no license · 3 matches
- ---
- title: "POGZ TF binding"
- author: "Yating Liu"
- date: "3/2/2023"
- output: html_document
- ---
- ```{r setup, include=FALSE}
- knitr::opts_knit$set(root.dir = "~/Documents/Talkowski/POGZ_project/ATAC_seq")
- knitr::opts_chunk$set(echo = TRUE)
- ```
- ```{r message=FALSE}
- library(ggplot2)
- library(dplyr)
- library(stringr)
- library(biomaRt)
- library(GenomicRanges)
- options(stringsAsFactors = F)
- ```
- ## Get differential TF binding bw DEL vs WT from TF footprinting
- ### Read iN TF footprinting results into a list
- ```{r}
- if (!file.exists("TF_footprint/iN_bindetect_results.rds")) {
- iN_bindetect_results <- list()
- iN_bindetect_results[["iN_GM_DEL2_vs_WT2"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM8330_Feb2021.DEL2_vs_WT2.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_GM_DEL1_vs_WT1"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM8330_Mar2021.DEL1_vs_WT1.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_MGH_DEL2_vs_WT2"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_MGH_Feb2021.DEL2_vs_WT2.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_MGH_DEL1_vs_WT1"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_MGH_Mar2021.DEL1_vs_WT1.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_GM_DEL_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM.DEL_vs_WT.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_MGH_DEL_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_MGH.DEL_vs_WT.conservative/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_GM_del1del2_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM8330_Sep2020.DEL1DEL2_vs_WT.conservative_with_pogz_motifs/bindetect_results.txt", header = T)
- iN_bindetect_results[["iN_GM_del1_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM8330_Sep2020.DEL1_vs_WT.conservative_with_pogz_motifs/bindetect_results.txt", header = T)
- saveRDS(iN_bindetect_results, "TF_footprint/iN_bindetect_results.rds")
- }
- iN_bindetect_results <- readRDS("TF_footprint/iN_bindetect_results.rds")
- ```
- ### Read NSC TF footprinting results into a list
- ```{r}
- if (!file.exists("TF_footprint/NSC_bindetect_results.rds")) {
- NSC_bindetect_results <- list()
- NSC_bindetect_results[["NSC_GM_DEL1_vs_WT1"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM8330_Nov2020.DEL1_vs_WT1.conservative/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_GM_DEL2_vs_WT2"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM8330_Nov2020.DEL2_vs_WT2.conservative/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_MGH_DEL1_vs_WT1"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_MGH_Nov2020.DEL1_vs_WT1.conservative/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_MGH_DEL2_vs_WT2"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_MGH_Nov2020.DEL2_vs_WT2.conservative/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_GM_del1del2_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM8330_Aug2020.DEL1DEL2_vs_WT.conservative_with_pogz_motifs/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_GM_del1_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM8330_Aug2020.DEL1_vs_WT.conservative_with_pogz_motifs/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_GM_DEL_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM.DEL_vs_WT.conservative/bindetect_results.txt", header = T)
- NSC_bindetect_results[["NSC_MGH_DEL_vs_WT"]] <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_MGH.DEL_vs_WT.conservative/bindetect_results.txt", header = T)
- saveRDS(NSC_bindetect_results, "TF_footprint/NSC_bindetect_results.rds")
- }
- NSC_bindetect_results <- readRDS("TF_footprint/NSC_bindetect_results.rds")
- ```
- ### Function to get differential TF bindings
- threshold is set according to the TOBIAS paper: https://www.nature.com/articles/s41467-020-18035-1
- All TFs with -log10(p-value) above the 95% quantile or differential binding scores smaller/larger than the 5%
- and 95% quantiles (top 5% in each direction) are colored and shown with labels.
- ```{r}
- identify_differential_TF <- function(bindetect_results, by = 0.05) {
- pvalue_col <- colnames(bindetect_results)[str_detect(colnames(bindetect_results), "_pvalue")]
- score_change_col <- colnames(bindetect_results)[str_detect(colnames(bindetect_results), "_change")]
- sig_pvalue <- quantile(-log10(bindetect_results[[pvalue_col]]), probs = seq(0,1,by))
- sig_pvalue <- unname(sig_pvalue[length(sig_pvalue)-1])
- score_change <- quantile(bindetect_results[[score_change_col]], probs = seq(0,1,by))
- score_change_small <- unname(score_change[2])
- score_change_large <- unname(score_change[length(score_change)-1])
- bindetect_results[["sig_change_up"]] <- bindetect_results[[score_change_col]] > 0 & (bindetect_results[[score_change_col]] > score_change_large | -log10(bindetect_results[[pvalue_col]]) > sig_pvalue)
- bindetect_results[["sig_change_down"]] <- bindetect_results[[score_change_col]] < 0 & (bindetect_results[[score_change_col]] < score_change_small | -log10(bindetect_results[[pvalue_col]]) > sig_pvalue)
- bindetect_results <- bindetect_results %>% mutate(sig_change = case_when(sig_change_up == TRUE ~ "up",
- sig_change_down == TRUE ~"down",
- TRUE ~"NA"))
- #print(str(bindetect_results))
- p <- ggplot(bindetect_results, aes(x = get(score_change_col), y = -log10(get(pvalue_col)), color = sig_change)) +
- geom_point() +
- scale_color_manual(breaks = c("up", "down", "NA"), values = c("red", "blue", "grey")) +
- labs(x = "score_change", y = "-log10(pvalue)") +
- geom_hline(yintercept = sig_pvalue, color = "purple", linetype = "dotted") +
- geom_vline(xintercept = score_change_small, color = "blue", linetype = "dotted") +
- geom_vline(xintercept = score_change_large, color = "red", linetype = "dotted")
- print(p)
- return (bindetect_results)
- }
- ```
- ## Overlap function
- ```{r}
- overlap_tf <- function(tf1, tf2, tf1_name, tf2_name) {
- up <- intersect(tf1[tf1$sig_change == "up",]$motif_id, tf2[tf2$sig_change == "up",]$motif_id)
- down <- intersect(tf1[tf1$sig_change == "down",]$motif_id, tf2[tf2$sig_change == "down",]$motif_id)
- up_down <- intersect(tf1[tf1$sig_change == "up",]$motif_id, tf2[tf2$sig_change == "down",]$motif_id)
- down_up <- intersect(tf1[tf1$sig_change == "down",]$motif_id, tf2[tf2$sig_change == "up",]$motif_id)
- summary_stat <- matrix(0, nrow = 2, ncol = 2)
- rownames(summary_stat) <- c(paste0(tf1_name, "_up"), paste0(tf1_name, "_down"))
- colnames(summary_stat) <- c(paste0(tf2_name, "_up"), paste0(tf2_name, "_down"))
- summary_stat[paste0(tf1_name, "_up"),paste0(tf2_name, "_up")] <- length(up)
- summary_stat[paste0(tf1_name, "_down"),paste0(tf2_name, "_down")] <- length(down)
- summary_stat[paste0(tf1_name, "_up"),paste0(tf2_name, "_down")] <- length(up_down)
- summary_stat[paste0(tf1_name, "_down"),paste0(tf2_name, "_up")] <- length(down_up)
- print(summary_stat)
- return (list(up = up, down = down, up_down = up_down, down_up = down_up))
- }
- ```
- ### iN DEL vs WT
- ```{r}
- iN_GM_DEL_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_GM_DEL_vs_WT)
- iN_MGH_DEL_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_MGH_DEL_vs_WT)
- # iN GM het vs iN MGH het
- iN_overlapping_DEL_vs_WT_GM_MGH <- overlap_tf(iN_GM_DEL_vs_WT, iN_MGH_DEL_vs_WT, "GM", "MGH")
- saveRDS(iN_overlapping_DEL_vs_WT_GM_MGH, "TF_footprint/iN_GM_MGH_overlapped_DTF.rds")
- iN_GM_del1del2_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_GM_del1del2_vs_WT)
- # iN GM het vs comp het
- overlapping_DEL1DEL2_vs_WT_DEL_vs_WT <- overlap_tf(iN_GM_del1del2_vs_WT, iN_GM_DEL_vs_WT, "iN_GM_del1del2_vs_WT", "iN_GM_DEL_vs_WT")
- ```
- ### NSC DEL vs WT
- ```{r}
- NSC_GM_DEL_vs_WT <- identify_differential_TF(NSC_bindetect_results$NSC_GM_DEL_vs_WT)
- NSC_MGH_DEL_vs_WT <- identify_differential_TF(NSC_bindetect_results$NSC_MGH_DEL_vs_WT)
- NSC_overlapping_DEL_vs_WT_GM_MGH <- overlap_tf(NSC_GM_DEL_vs_WT, NSC_MGH_DEL_vs_WT, "GM", "MGH")
- saveRDS(NSC_overlapping_DEL_vs_WT_GM_MGH, "TF_footprint/NSC_GM_MGH_overlapped_DTF.rds")
- NSC_GM_del1del2_vs_WT <- identify_differential_TF(NSC_bindetect_results$NSC_GM_del1del2_vs_WT)
- # NSC GM het vs comp het
- overlapping_DEL1_vs_WT1_DEL2_vs_WT2 <- overlap_tf(NSC_GM_del1del2_vs_WT, NSC_GM_DEL_vs_WT, "NSC_GM_del1del2_vs_WT", "NSC_GM_DEL_vs_WT")
- ```
- #### POGZ motif change and pvalue
- ```{r}
- pogz_motif_stat <- data.frame()
- bindetect_lst <- list(iN_GM_DEL_vs_WT = iN_bindetect_results$iN_GM_DEL_vs_WT,
- iN_MGH_DEL_vs_WT = iN_bindetect_results$iN_MGH_DEL_vs_WT,
- NSC_GM_DEL_vs_WT = NSC_bindetect_results$NSC_GM_DEL_vs_WT,
- NSC_MGH_DEL_vs_WT = NSC_bindetect_results$NSC_MGH_DEL_vs_WT)
- for (bindetect_name in names(bindetect_lst)) {
- message(bindetect_name)
- bindetect_results <- bindetect_lst[[bindetect_name]]
- by = 0.05
- pvalue_col <- colnames(bindetect_results)[str_detect(colnames(bindetect_results), "_pvalue")]
- score_change_col <- colnames(bindetect_results)[str_detect(colnames(bindetect_results), "_change")]
- bindetect_results[["log10_pvalue"]] <- -log10(bindetect_results[[pvalue_col]])
- sig_pvalue <- quantile(-log10(bindetect_results[[pvalue_col]]), probs = seq(0,1,by))
- sig_pvalue_95 <- unname(sig_pvalue[length(sig_pvalue)-1])
- sig_pvalue_90 <- unname(sig_pvalue[length(sig_pvalue)-2])
- score_change <- quantile(bindetect_results[[score_change_col]], probs = seq(0,1,by))
- score_change_small <- unname(score_change[2])
- score_change_large <- unname(score_change[length(score_change)-1])
- colnames(bindetect_results) <- str_remove_all(colnames(bindetect_results), "_GM|_MGH")
- score_change_col <- colnames(bindetect_results)[str_detect(colnames(bindetect_results), "_change")]
- pogz_motif_stat <- rbind(pogz_motif_stat, data.frame(bindetect_results[str_detect(bindetect_results$name, "pogz"),],
- cell_type = bindetect_name,
- log10_pvalue_percentile = sapply(bindetect_results[str_detect(bindetect_results$name, "pogz"),c("log10_pvalue")], function(x) {return(names(tail(sig_pvalue[x > sig_pvalue],n=1)))}),
- score_change_percentile = sapply(bindetect_results[str_detect(bindetect_results$name, "pogz"),c(score_change_col)], function(x) {return(names(tail(score_change[x > score_change],n=1)))}),
- log10_pvalue_95_percentile = sig_pvalue_95,
- score_change_5_percentile = score_change_small,
- score_change_95_percentile = score_change_large))
- }
- write_xlsx(pogz_motif_stat, "TF_footprint/pogz_motif_stat.xlsx")
- ```
- ## Get differential binding TFs target genes
- ### Function for get target genes for a table of motif ids and name, in the end overlap these target genes with DEG results
- ```{r}
- ## Get promoters of all genes
- require(ensembldb)
- edb <- EnsDb("data/Homo_sapiens.GRCh38.92.sqlite")
- allpromoters <- promoters(edb, upstream = 1500, downstream = 500)
- allpromoters <- allpromoters[seqnames(allpromoters) %in% c(1:22, "X", "Y")]
- allpromoters <- keepStandardChromosomes(allpromoters)
- gene_annotation <- readRDS("../iN_05_2021/results/gene_annotation.rds")
- get_TF_target_genes <- function(motif_info, express_colname, allpromoters, gene_annotation, DTFB_path, del_vs_wt_final_degs, prefix) {
- # get motif directory name
- overlapped_TF_folder <- paste(motif_info$name, motif_info$motif_id, sep = "_")
- overlapped_TF_folder <- str_remove(overlapped_TF_folder, "::")
- overlapped_TF_folder <- str_remove_all(overlapped_TF_folder, "\\(|\\)")
- genes_vs_TF_overlapping <- data.frame()
- for (i in 1:length(overlapped_TF_folder)) {
- bind_sites <- read.table(file.path(DTFB_path, overlapped_TF_folder[i], paste0(overlapped_TF_folder[i],"_overview.txt")), header = T)
- bind_sites <- bind_sites[bind_sites$TFBS_chr %in% c(1:22, "X", "Y"),]
- # if the TF is up-regulated, we use the TF bind regions that specific to DEL, otherwise use TF bind region specific to WT
- del_bind_col <- which(str_detect(colnames(bind_sites), "DEL.*_bound"))
- wt_bind_col <- which(str_detect(colnames(bind_sites), "WT.*_bound"))
- # Filter bind sites by keeping the sites that are predicted bound within at least one condition (WT or DEL)
- bind_sites <- bind_sites[bind_sites[[del_bind_col]] + bind_sites[[wt_bind_col]] > 0,]
- # Filter annotated bind sites that have the log2 fold change between the footprint scores of the two conditions > 90% quantile or < 10% quantile
- log2fc_col <- which(str_detect(colnames(bind_sites), "_log2fc"))
- log2fc_quantile <- quantile(bind_sites[[log2fc_col]], prob = seq(0,1, by = 0.05))
- # print(log2fc_quantile)
- bind_sites <- bind_sites[bind_sites[[log2fc_col]] < log2fc_quantile[["10%"]] | bind_sites[[log2fc_col]] > log2fc_quantile[["90%"]],]
- message(motif_info[i,]$name, ": number of bind sites in top quantitle: ", nrow(bind_sites))
- # if (motif_info[i,]$sig_change == "up") {
- # bind_sites <- bind_sites[bind_sites[[del_bind_col]] == 1 & bind_sites[[wt_bind_col]] == 0,]
- # } else {
- # bind_sites <- bind_sites[bind_sites[[del_bind_col]] == 0 & bind_sites[[wt_bind_col]] == 1,]
- # }
- bind_sites_gr <- GRanges(seqnames = bind_sites$TFBS_chr, ranges = IRanges(start = bind_sites$TFBS_start, end = bind_sites$TFBS_end), strand = bind_sites$TFBS_strand)
- overlapped_genes <- unique(allpromoters[queryHits(findOverlaps(allpromoters, bind_sites_gr)),]$gene_id)
- if (length(overlapped_genes) > 0) {
- overlapped_genes <- gene_annotation[overlapped_genes,]
- overlapped_genes <- cbind(overlapped_genes, motif_info[i,])
- # overlapped_genes$motif_id <- motif_info[i,]$motif_id
- # overlapped_genes$name <- motif_info[i,]$name
- # overlapped_genes$sig_change <- motif_info[i,]$sig_change
- genes_vs_TF_overlapping <- rbind(genes_vs_TF_overlapping, overlapped_genes)
- }
- }
- write.csv(genes_vs_TF_overlapping, paste0(prefix, "TF_target_genes.csv"), row.names = F, quote = F)
- # only pick the DTF that are expressed
- message("Total num DTF: ", length(unique(genes_vs_TF_overlapping$motif_id)))
- message("Expressed DTF: ", length(unique(genes_vs_TF_overlapping[genes_vs_TF_overlapping[[express_colname]] == "Yes",]$motif_id)))
- genes_vs_TF_overlapping <- genes_vs_TF_overlapping[genes_vs_TF_overlapping[[express_colname]] == "Yes",]
- # Overlap TF targets with DEGs
- TF_genes <- genes_vs_TF_overlapping %>%
- dplyr::select(ensembl_id, symbol, gene_type, motif_id, name, sig_change) %>%
- group_by(ensembl_id, symbol, gene_type) %>%
- summarise(motif_id = paste(motif_id, collapse = ";"), name = paste(name, collapse = ";"),sig_change = paste(sig_change, collapse = ";"))
- TF_degs <- left_join(TF_genes[TF_genes$ensembl_id %in% intersect(TF_genes$ensembl_id, del_vs_wt_final_degs$ensembl_id),], del_vs_wt_final_degs[, which(str_detect(colnames(del_vs_wt_final_degs), "ensembl_id|padj|log2FC"))], by="ensembl_id")
- write.csv(TF_genes, paste0(prefix, "TF_genes.csv"), row.names = F, quote = F)
- write.csv(TF_degs, paste0(prefix, "TF_degs.csv"), row.names = F, quote = F)
- return (TF_degs)
- }
- ```
- ## all motif expression information (whether a motif is expressed in iN or NSC)
- ```{r}
- all_motifs_expression_info <- readxl::read_excel("../POGZ_final/TF_footprinting/all_motifs_whether_expressed.xlsx")
- ```
- ## iN DTF target genes
- ```{r}
- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/del_vs_wt_overlapped_GM_MGH.csv")
- ```
- ### Get differential binding GM TFs target genes
- ```{r}
- DTFB_GM_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM.DEL_vs_WT.conservative"
- DTFB_GM <- iN_GM_DEL_vs_WT[iN_GM_DEL_vs_WT$sig_change != "NA",]
- DTFB_GM <- left_join(DTFB_GM, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- writexl::write_xlsx(DTFB_GM, "TF_footprint/iN_GM_DTF.xlsx")
- DTFB_GM_DEGs <- get_TF_target_genes(DTFB_GM, express_colname = "iN_expression", allpromoters, gene_annotation, DTFB_GM_path, del_vs_wt_final_degs, "TF_footprint/iN_GM_")
- ```
- ### Get differential binding MGH TFs target genes
- ```{r}
- DTFB_MGH_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_MGH.DEL_vs_WT.conservative"
- DTFB_MGH <- iN_MGH_DEL_vs_WT[iN_MGH_DEL_vs_WT$sig_change != "NA",]
- DTFB_MGH <- left_join(DTFB_MGH, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- writexl::write_xlsx(DTFB_MGH, "TF_footprint/iN_MGH_DTF.xlsx")
- DTFB_MGH_DEGs <- get_TF_target_genes(DTFB_MGH, express_colname = "iN_expression", allpromoters, gene_annotation, DTFB_MGH_path, del_vs_wt_final_degs, "TF_footprint/iN_MGH_")
- ```
- ### Get overlapped TF targets DEGs across background
- ```{r}
- del_vs_wt_GM <- read.csv("../POGZ_paper/results/DEL_vs_WT_GM__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- del_vs_wt_MGH <- read.csv("../POGZ_paper/results/DEL_vs_WT_MGH__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- analyzed_genes <- intersect(del_vs_wt_GM$ensembl_id, del_vs_wt_MGH$ensembl_id)
- DTFB_GM <- readxl::read_excel("TF_footprint/iN_GM_DTF.xlsx")
- DTFB_MGH <- readxl::read_excel("TF_footprint/iN_MGH_DTF.xlsx")
- DTFB_GM_genes <- read.csv("TF_footprint/iN_GM_TF_genes.csv")
- DTFB_MGH_genes <- read.csv("TF_footprint/iN_MGH_TF_genes.csv")
- DTFB_GM_DEGs <- read.csv("TF_footprint/iN_GM_TF_degs.csv")
- DTFB_GM_DEGs$deg_regulation <- ifelse(DTFB_GM_DEGs$log2FC_GM>0, "up", "down")
- DTFB_MGH_DEGs <- read.csv("TF_footprint/iN_MGH_TF_degs.csv")
- overlapped_tf_degs <- intersect(DTFB_MGH_DEGs$ensembl_id, DTFB_GM_DEGs$ensembl_id)
- overlapped_tf_degs_info <- left_join(DTFB_GM_DEGs[, c("ensembl_id","symbol","gene_type","motif_id","name","sig_change", "deg_regulation")], DTFB_MGH_DEGs[, c("ensembl_id","symbol","gene_type","motif_id","name","sig_change")], by = c("ensembl_id"="ensembl_id", "symbol"="symbol", "gene_type"="gene_type"), suffix = c(".GM", ".MGH"))
- overlapped_tf_degs_info <- overlapped_tf_degs_info[complete.cases(overlapped_tf_degs_info),]
- write.csv(overlapped_tf_degs_info, "TF_footprint/iN_GM_MGH_overlapped_tf_degs_info.csv", row.names = F)
- dtf_target_genes_gm <- length(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes))
- dtf_target_genes_mgh <- length(intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes))
- overlap_dtf_targets <- length(intersect(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes), intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes)))
- fisher_test <- fisher.test(matrix(c(overlap_dtf_targets, dtf_target_genes_gm-overlap_dtf_targets, dtf_target_genes_mgh-overlap_dtf_targets, length(analyzed_genes)-dtf_target_genes_gm-dtf_target_genes_mgh+overlap_dtf_targets), nrow = 2, ncol = 2), alternative = "greater")
- message("Number of GM DTF: ", nrow(DTFB_GM),
- "\nNumber of GM DTF that are expressed: ", sum(DTFB_GM$iN_expression == "Yes"),
- "\nNumber of MGH DTF: ", nrow(DTFB_MGH),
- "\nNumber of MGH DTF that are expressed: ", sum(DTFB_MGH$iN_expression == "Yes"))
- message("Number of DTF target genes in GM: ", length(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes)),
- "\nNumber of DTF target genes in MGH: ", length(intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes)),
- "\nNumber of DTF target genes in both GM and MGH: ", length(intersect(DTFB_GM_genes$ensembl_id,intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes))),
- "\nNum overlap DTF target genes in GM vs MGH ", overlap_dtf_targets,
- " (fisher test pvalue = ", fisher_test$p.value, ")",
- "\nNumber of DTF target DEGs in GM: ", nrow(DTFB_GM_DEGs), " up = ", sum(DTFB_GM_DEGs$log2FC_GM>0), " down = ", sum(DTFB_GM_DEGs$log2FC_GM<0),
- "\nNumber of DTF target DEGs in MGH: ", nrow(DTFB_MGH_DEGs)," up = ", sum(DTFB_MGH_DEGs$log2FC_GM>0), " down = ", sum(DTFB_MGH_DEGs$log2FC_GM<0),
- "\nNumber of overlapped DTF target DEGs: ", nrow(overlapped_tf_degs_info), " up = ", sum(overlapped_tf_degs_info$deg_regulation=="up"), " down = ", sum(overlapped_tf_degs_info$deg_regulation=="down"))
- ```
- ## iN comp DTF target genes
- ```{r}
- del_vs_wt_final_degs <- read.csv("../POGZ_iN/results/8330_del1_del2_vs_8330_WT_threshold_0.5.csv")
- del_vs_wt_final_degs$log2FC <- del_vs_wt_final_degs$log2FoldChange
- del_vs_wt_final_degs <- del_vs_wt_final_degs[del_vs_wt_final_degs$padj<0.1,]
- ```
- ### Get differential binding GM TFs target genes
- ```{r}
- iN_GM_del1del2_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_GM_del1del2_vs_WT)
- DTFB_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM8330_Sep2020.DEL1DEL2_vs_WT.conservative_with_pogz_motifs"
- DTFB_comp <- iN_GM_del1del2_vs_WT[iN_GM_del1del2_vs_WT$sig_change != "NA",]
- DTFB_comp <- left_join(DTFB_comp, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- DTFB_DEGs <- get_TF_target_genes(DTFB_comp, express_colname = "iN_comp_het_expression", allpromoters, gene_annotation, DTFB_path, del_vs_wt_final_degs, "TF_footprint/iN_comp_")
- ```
- ## NSC DTF target genes
- ```{r}
- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/NSC_del_vs_wt_overlapped_GM_MGH.csv")
- ```
- ### Get differential binding GM TFs target genes
- ```{r}
- DTFB_GM_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM.DEL_vs_WT.conservative"
- DTFB_GM <- NSC_GM_DEL_vs_WT[NSC_GM_DEL_vs_WT$sig_change != "NA",]
- DTFB_GM <- left_join(DTFB_GM, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- writexl::write_xlsx(DTFB_GM, "TF_footprint/NSC_GM_DTF.xlsx")
- DTFB_GM_DEGs <- get_TF_target_genes(DTFB_GM, express_colname = "NSC_expression", allpromoters, gene_annotation, DTFB_GM_path, del_vs_wt_final_degs, "TF_footprint/NSC_GM_")
- ```
- ### Get differential binding MGH TFs target genes
- ```{r}
- DTFB_MGH_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_MGH.DEL_vs_WT.conservative"
- DTFB_MGH <- NSC_MGH_DEL_vs_WT[NSC_MGH_DEL_vs_WT$sig_change != "NA",]
- DTFB_MGH <- left_join(DTFB_MGH, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- writexl::write_xlsx(DTFB_MGH, "TF_footprint/NSC_MGH_DTF.xlsx")
- DTFB_MGH_DEGs <- get_TF_target_genes(DTFB_MGH, express_colname = "NSC_expression", allpromoters, gene_annotation, DTFB_MGH_path, del_vs_wt_final_degs, "TF_footprint/NSC_MGH_")
- ```
- ### Get overlapped TF targets DEGs across background
- ```{r}
- del_vs_wt_GM <- read.table("../NSC_03_2020/meta_analysis/POGZ_8330del1_2.metaAnalysis.txt", header = T)
- del_vs_wt_MGH <- read.table("../NSC_03_2020/meta_analysis/POGZ_MGHdel1_2.metaAnalysis.txt", header = T)
- analyzed_genes <- intersect(del_vs_wt_GM$ensemblid, del_vs_wt_MGH$ensemblid)
- DTFB_GM <- readxl::read_excel("TF_footprint/NSC_GM_DTF.xlsx")
- DTFB_MGH <- readxl::read_excel("TF_footprint/NSC_MGH_DTF.xlsx")
- DTFB_GM_genes <- read.csv("TF_footprint/NSC_GM_TF_genes.csv")
- DTFB_MGH_genes <- read.csv("TF_footprint/NSC_MGH_TF_genes.csv")
- DTFB_GM_DEGs <- read.csv("TF_footprint/NSC_GM_TF_degs.csv")
- DTFB_GM_DEGs$deg_regulation <- ifelse(DTFB_GM_DEGs$log2FC_GM>0, "up", "down")
- DTFB_MGH_DEGs <- read.csv("TF_footprint/NSC_MGH_TF_degs.csv")
- DTFB_MGH_DEGs$deg_regulation <- ifelse(DTFB_MGH_DEGs$log2FC_GM>0, "up", "down")
- overlapped_tf_degs <- intersect(DTFB_MGH_DEGs$ensembl_id, DTFB_GM_DEGs$ensembl_id)
- overlapped_tf_degs_info <- left_join(DTFB_GM_DEGs[, c("ensembl_id","symbol","gene_type","motif_id","name","sig_change", "deg_regulation")], DTFB_MGH_DEGs[, c("ensembl_id","symbol","gene_type","motif_id","name","sig_change")], by = c("ensembl_id"="ensembl_id", "symbol"="symbol", "gene_type"="gene_type"), suffix = c(".GM", ".MGH"))
- overlapped_tf_degs_info <- overlapped_tf_degs_info[complete.cases(overlapped_tf_degs_info),]
- write.csv(overlapped_tf_degs_info, "TF_footprint/NSC_GM_MGH_overlapped_tf_degs_info.csv", row.names = F)
- dtf_target_genes_gm <- length(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes))
- dtf_target_genes_mgh <- length(intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes))
- overlap_dtf_targets <- length(intersect(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes), intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes)))
- fisher_test <- fisher.test(matrix(c(overlap_dtf_targets, dtf_target_genes_gm-overlap_dtf_targets, dtf_target_genes_mgh-overlap_dtf_targets, length(analyzed_genes)-dtf_target_genes_gm-dtf_target_genes_mgh+overlap_dtf_targets), nrow = 2, ncol = 2), alternative = "greater")
- message("Number of GM DTF: ", nrow(DTFB_GM),
- "\nNumber of GM DTF that are expressed: ", sum(DTFB_GM$NSC_expression == "Yes"),
- "\nNumber of MGH DTF: ", nrow(DTFB_MGH),
- "\nNumber of MGH DTF that are expressed: ", sum(DTFB_MGH$NSC_expression == "Yes"))
- message("Number of DTF target genes in GM: ", length(intersect(DTFB_GM_genes$ensembl_id, analyzed_genes)),
- "\nNumber of DTF target genes in MGH: ", length(intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes)),
- "\nNumber of DTF target genes in both GM and MGH: ", length(intersect(DTFB_GM_genes$ensembl_id,intersect(DTFB_MGH_genes$ensembl_id, analyzed_genes))),
- "\n Num overlap DTF target genes in GM vs MGH ", overlap_dtf_targets,
- " (fisher test pvalue = ", fisher_test$p.value, ")",
- "\nNumber of DTF target DEGs in GM: ", nrow(DTFB_GM_DEGs), " up = ", sum(DTFB_GM_DEGs$log2FC_GM>0), " down = ", sum(DTFB_GM_DEGs$log2FC_GM<0),
- "\nNumber of DTF target DEGs in MGH: ", nrow(DTFB_MGH_DEGs)," up = ", sum(DTFB_MGH_DEGs$log2FC_GM>0), " down = ", sum(DTFB_MGH_DEGs$log2FC_GM<0),
- "\nNumber of overlapped DTF target DEGs: ", nrow(overlapped_tf_degs_info), " up = ", sum(overlapped_tf_degs_info$deg_regulation=="up"), " down = ", sum(overlapped_tf_degs_info$deg_regulation=="down"))
- ```
- ## NSC comp DTF target genes
- ```{r}
- del_vs_wt_final_degs <- read.csv("../POGZ_NSC/results/8330_del1_del2_vs_8330_WT_threshold_0.5.csv")
- del_vs_wt_final_degs$log2FC <- del_vs_wt_final_degs$log2FoldChange
- del_vs_wt_final_degs <- del_vs_wt_final_degs[del_vs_wt_final_degs$padj<0.1,]
- ```
- ### Get differential binding GM TFs target genes
- ```{r}
- NSC_GM_del1del2_vs_WT <- identify_differential_TF(NSC_bindetect_results$NSC_GM_del1del2_vs_WT)
- DTFB_path <- "/Volumes/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.NSC_GM8330_Aug2020.DEL1DEL2_vs_WT.conservative_with_pogz_motifs"
- DTFB_comp <- NSC_GM_del1del2_vs_WT[NSC_GM_del1del2_vs_WT$sig_change != "NA",]
- DTFB_comp <- left_join(DTFB_comp, all_motifs_expression_info[, -c(which(colnames(all_motifs_expression_info) == "name"))], by = "motif_id")
- DTFB_DEGs <- get_TF_target_genes(DTFB_comp, express_colname = "NSC_comp_het_expression", allpromoters, gene_annotation, DTFB_path, del_vs_wt_final_degs, "TF_footprint/NSC_comp_")
- ```
- ## Check whether promoter targets and DEGs are significant overlapped
- ### iN intersect with symbol
- ```{r}
- del_vs_wt_GM <- read.csv("../POGZ_paper/results/DEL_vs_WT_GM__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- del_vs_wt_MGH <- read.csv("../POGZ_paper/results/DEL_vs_WT_MGH__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- analyzed_genes <- intersect(toupper(del_vs_wt_GM$symbol), toupper(del_vs_wt_MGH$symbol))
- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/del_vs_wt_overlapped_GM_MGH.csv")
- final_degs <- toupper(del_vs_wt_final_degs$symbol)
- DTFB_GM_genes <- read.csv("TF_footprint/iN_GM_TF_genes.csv")
- DTFB_MGH_genes <- read.csv("TF_footprint/iN_MGH_TF_genes.csv")
- DTFB_GM_DEGs <- read.csv("TF_footprint/iN_GM_TF_degs.csv")
- DTFB_MGH_DEGs <- read.csv("TF_footprint/iN_MGH_TF_degs.csv")
- # iN_GM
- target_genes <- intersect(toupper(DTFB_GM_genes$symbol), analyzed_genes)
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- require(VennDiagram)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "iN_GM_promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/iN_GM_promoter_targes_vs_degs.png")
- # iN_MGH
- target_genes <- intersect(toupper(DTFB_MGH_genes$symbol), analyzed_genes)
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "iN_MGH_promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/iN_MGH_promoter_targes_vs_degs.png")
- # iN GM and MGH overlapped DTFP
- target_genes <- intersect(toupper(DTFB_GM_genes$symbol), intersect(toupper(DTFB_MGH_genes$symbol), analyzed_genes))
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "iN_GM_and_MGH_overlapped promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/iN_GM_MGH_overlapped_promoter_targes_vs_degs.png")
- ```
- ### NSC intersect with symbol
- ```{r}
- del_vs_wt_GM <- read.table("../NSC_03_2020/meta_analysis/POGZ_8330del1_2.metaAnalysis.txt", header = T)
- del_vs_wt_MGH <- read.table("../NSC_03_2020/meta_analysis/POGZ_MGHdel1_2.metaAnalysis.txt", header = T)
- analyzed_genes <- intersect(toupper(del_vs_wt_GM$symbol), toupper(del_vs_wt_MGH$symbol))
- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/NSC_del_vs_wt_overlapped_GM_MGH.csv")
- final_degs <- toupper(del_vs_wt_final_degs$symbol)
- DTFB_GM_genes <- read.csv("TF_footprint/NSC_GM_TF_genes.csv")
- DTFB_MGH_genes <- read.csv("TF_footprint/NSC_MGH_TF_genes.csv")
- DTFB_GM_DEGs <- read.csv("TF_footprint/NSC_GM_TF_degs.csv")
- DTFB_MGH_DEGs <- read.csv("TF_footprint/NSC_MGH_TF_degs.csv")
- # NSC_GM
- target_genes <- intersect(toupper(DTFB_GM_genes$symbol), analyzed_genes)
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- require(VennDiagram)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "NSC_GM_promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/NSC_GM_promoter_targes_vs_degs.png")
- # NSC_MGH
- target_genes <- intersect(toupper(DTFB_MGH_genes$symbol), analyzed_genes)
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "NSC_MGH_promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/NSC_MGH_promoter_targes_vs_degs.png")
- # NSC GM and MGH overlapped DTFP
- target_genes <- intersect(toupper(DTFB_GM_genes$symbol), intersect(toupper(DTFB_MGH_genes$symbol), analyzed_genes))
- target_degs <- intersect(target_genes, final_degs)
- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater")
- message("target genes = ", length(target_genes), " target degs = ", length(target_degs))
- print(targets_vs_degs)
- venn.diagram(list("promoter_targets" = target_genes, "DEGs" = final_degs), main = "NSC_GM_and_MGH_overlapped promoter_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/NSC_GM_MGH_overlapped_promoter_targes_vs_degs.png")
- ```
- ## Check if POGZ a DTF
- ```{r}
- DTF_lst <- list(iN_GM_DTF = readxl::read_excel("TF_footprint/iN_GM_DTF.xlsx"),
- iN_MGH_DTF = readxl::read_excel("TF_footprint/iN_MGH_DTF.xlsx"),
- NSC_GM_DTF = readxl::read_excel("TF_footprint/NSC_GM_DTF.xlsx"),
- NSC_MGH_DTF = readxl::read_excel("TF_footprint/NSC_MGH_DTF.xlsx"),
- iN_comp_het_DTF <- readxl::read_excel("TF_footprint/"))
- lapply(1:length(DTF_lst), function(i) {message(names(DTF_lst)[i], ": ", DTF_lst[[i]][str_detect(DTF_lst[[i]]$name, "POGZ|pogz"),]$name)})
- ```
- ## Run enrichment analysis on DTFP from NSC and iN
- ```{r}
- source("../../MyTools/RNAseq.analysis/R/enrichment.R")
- load("../../MyTools/enrichment_analysis/data/pathway_db/pathway_db.Rdata")
- iN_del_vs_wt_GM <- read.csv("../POGZ_paper/results/DEL_vs_WT_GM__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- rownames(iN_del_vs_wt_GM) <- iN_del_vs_wt_GM$ensembl_id
- iN_del_vs_wt_MGH <- read.csv("../POGZ_paper/results/DEL_vs_WT_MGH__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1)
- NSC_del_vs_wt_GM <- read.table("../NSC_03_2020/meta_analysis/POGZ_8330del1_2.metaAnalysis.txt", sep = "\t", header = T)
- NSC_del_vs_wt_MGH <- read.table("../NSC_03_2020/meta_analysis/POGZ_MGHdel1_2.metaAnalysis.txt", sep = "\t", header = T)
- iN_DTFB_GM_DEGs <- read.csv("TF_footprint/iN_GM_TF_degs.csv")
- iN_DTFB_MGH_DEGs <- read.csv("TF_footprint/iN_MGH_TF_degs.csv")
- NSC_DTFB_GM_DEGs <- read.csv("TF_footprint/NSC_GM_TF_degs.csv")
- NSC_DTFB_MGH_DEGs <- read.csv("TF_footprint/NSC_MGH_TF_degs.csv")
- iN_DTFB_GM_DEGs_sig_terms <- run_enrichment_on_genelist(list(iN_DTFB_GM_DEGs=iN_DTFB_GM_DEGs$symbol), iN_del_vs_wt_GM$symbol, outfolder = "TF_footprint/enrichment_analysis/iN_DTFB_GM_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/iN_DTFB_GM_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- iN_DTFB_MGH_DEGs_sig_terms <- run_enrichment_on_genelist(list(iN_DTFB_MGH_DEGs=iN_DTFB_MGH_DEGs$symbol), iN_del_vs_wt_MGH$symbol, outfolder = "TF_footprint/enrichment_analysis/iN_DTFB_MGH_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/iN_DTFB_MGH_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- iN_GM_MGH_overlapped_DTFB_DEGs_sig_terms <- run_enrichment_on_genelist(list(iN_GM_MGH_overlapped_DTFB_DEGs=intersect(iN_DTFB_GM_DEGs$symbol, iN_DTFB_MGH_DEGs$symbol)), intersect(iN_del_vs_wt_GM$symbol,iN_del_vs_wt_MGH$symbol), outfolder = "TF_footprint/enrichment_analysis/iN_GM_MGH_overlapped_DTFB_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/iN_GM_MGH_overlapped_DTFB_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- NSC_DTFB_GM_DEGs_sig_terms <- run_enrichment_on_genelist(list(NSC_DTFB_GM_DEGs=NSC_DTFB_GM_DEGs$symbol), NSC_del_vs_wt_MGH$symbol, gene_to_remove = "POGZ", outfolder = "TF_footprint/enrichment_analysis/NSC_DTFB_GM_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/NSC_DTFB_GM_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- NSC_DTFB_MGH_DEGs_sig_terms <- run_enrichment_on_genelist(list(NSC_DTFB_MGH_DEGs=NSC_DTFB_MGH_DEGs$symbol), NSC_del_vs_wt_MGH$symbol, gene_to_remove = "POGZ", outfolder = "TF_footprint/enrichment_analysis/NSC_DTFB_MGH_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/NSC_DTFB_MGH_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- NSC_GM_MGH_overlapped_DTFB_DEGs_sig_terms <- run_enrichment_on_genelist(list(NSC_GM_MGH_overlapped_DTFB_DEGs=intersect(NSC_DTFB_GM_DEGs$symbol,NSC_DTFB_MGH_DEGs$symbol)), intersect(NSC_del_vs_wt_GM$symbol,NSC_del_vs_wt_MGH$symbol), gene_to_remove = "POGZ", outfolder = "TF_footprint/enrichment_analysis/NSC_GM_MGH_overlapped_DTFB_DEGs", outplotfolder = "TF_footprint/enrichment_analysis/NSC_GM_MGH_overlapped_DTFB_DEGs/plots", skip_enrichment_analysis=F, fdr_threshold=0.1)
- sig_terms_lst <- list(iN_DTFB_GM_DEGs_sig_terms = iN_DTFB_GM_DEGs_sig_terms,
- iN_DTFB_MGH_DEGs_sig_terms =iN_DTFB_MGH_DEGs_sig_terms,
- iN_GM_MGH_overlapped_DTFB_DEGs_sig_terms = iN_GM_MGH_overlapped_DTFB_DEGs_sig_terms,
- NSC_DTFB_GM_DEGs_sig_terms = NSC_DTFB_GM_DEGs_sig_terms,
- NSC_DTFB_MGH_DEGs_sig_terms = NSC_DTFB_MGH_DEGs_sig_terms,
- NSC_GM_MGH_overlapped_DTFB_DEGs_sig_terms = NSC_GM_MGH_overlapped_DTFB_DEGs_sig_terms)
- for (datset in names(sig_terms_lst)) {
- sig_terms <- sig_terms_lst[[datset]]
- if (nrow(sig_terms) == 0) {
- next
- }
- sig_terms_top10 <- slice_max(sig_terms, order_by = logBH, n = 10)
- sig_terms_top10$name <- str_wrap(sig_terms_top10$name, width = 50)
- sig_terms_top10$name <- factor(sig_terms_top10$name, levels = sig_terms_top10[order(sig_terms_top10$logBH),]$name)
- ggplot(sig_terms_top10, aes(x = name, y = logBH)) +
- geom_col(fill = "grey60") +
- coord_flip() +
- geom_text(aes(label = significant_genelist), hjust = 1) +
- labs(x = "", title = paste0("Top 10 sig terms for ", str_remove(datset, "_sig_terms"))) +
- theme(axis.text.y = element_text(size = 12), plot.title.position = "plot")
- ggsave(file.path("TF_footprint/enrichment_analysis", paste0("top10_sig_terms_", str_remove(datset, "_sig_terms"), ".pdf")), height = 7, width = 6)
- }
- ```
- <!-- ## Get differential binding TFs target genes version2 -->
- <!-- ### Get DTF for GM and MGH, find the DTF results folder path and write them to a text file -->
- <!-- ```{r} -->
- <!-- # Get DTF -->
- <!-- iN_bindetect_results <- readRDS("TF_footprint/iN_bindetect_results.rds") -->
- <!-- iN_GM_DEL_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_GM_DEL_vs_WT) -->
- <!-- iN_GM_DTF <- iN_GM_DEL_vs_WT[iN_GM_DEL_vs_WT$sig_change != "NA",] -->
- <!-- iN_MGH_DEL_vs_WT <- identify_differential_TF(iN_bindetect_results$iN_MGH_DEL_vs_WT) -->
- <!-- iN_MGH_DTF <- iN_MGH_DEL_vs_WT[iN_MGH_DEL_vs_WT$sig_change != "NA",] -->
- <!-- DTFB_GM_path <- "/data/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_GM.DEL_vs_WT.conservative" -->
- <!-- DTFB_MGH_path <- "/data/talkowski/Samples/POGZ/ATAC/TF_footprint/DTFB.iN_MGH.DEL_vs_WT.conservative" -->
- <!-- get_TF_folders <- function(motif_info, DTFB_path, output) { -->
- <!-- # get motif directory name -->
- <!-- TF_folder <- paste(motif_info$name, motif_info$motif_id, sep = "_") -->
- <!-- TF_folder <- str_remove(TF_folder, "::") -->
- <!-- TF_folder <- str_remove_all(TF_folder, "\\(|\\)") -->
- <!-- TF_folder <- file.path(DTFB_path, TF_folder) -->
- <!-- write.table(TF_folder, output, quote = F, row.names = F, col.names = F, sep = "\t") -->
- <!-- } -->
- <!-- get_TF_folders(iN_GM_DTF, DTFB_GM_path, output = "/Volumes/talkowski/Samples/POGZ/ATAC/region_annotation/DTF_annotation/data/iN_GM_DTF.txt") -->
- <!-- get_TF_folders(iN_MGH_DTF, DTFB_MGH_path, output = "/Volumes/talkowski/Samples/POGZ/ATAC/region_annotation/DTF_annotation/data/iN_MGH_DTF.txt") -->
- <!-- # Annotate DTF bind sites with Ensembl regulatory features and promters-1500to500bp at /data/talkowski/Samples/POGZ/ATAC/region_annotation/DTF_annotation -->
- <!-- ``` -->
- <!-- ### Downstream analysis on the bind sites are predicted bound within at least one condition (take a bit long time) -->
- <!-- ```{r eval=FALSE} -->
- <!-- stat_summary <- data.frame() -->
- <!-- promoter_overlapped_bind_sites_all_DTFs <- data.frame() -->
- <!-- enhancer_overlapped_bind_sites_all_DTFs <- data.frame() -->
- <!-- DTFs <- read.table("/Volumes/talkowski/Samples/POGZ/ATAC/region_annotation/DTF_annotation/data/iN_DTF.txt", col.names = "folder") -->
- <!-- for (DTF in DTFs$folder) { -->
- <!-- print(DTF) -->
- <!-- bind_site_file <- str_replace(file.path(DTF, paste0(basename(DTF), "_overview.txt")), "/data", "/Volumes") -->
- <!-- original_bind_sites <- read.table(bind_site_file, sep = "\t", header = T) -->
- <!-- DTF_stat <- data.frame(motif = basename(DTF), total_bind_sites = nrow(original_bind_sites)) -->
- <!-- prefix <- str_remove(basename(dirname(DTF)),"DEL_vs_WT.conservative") -->
- <!-- annotated_file <- list.files("/Volumes/talkowski/Samples/POGZ/ATAC/region_annotation/DTF_annotation/results_iN", paste0(prefix, basename(DTF), ".sorted.bed"), full.names = T) -->
- <!-- bind_sites <- read.table(annotated_file, sep = "\t") -->
- <!-- DTF_stat$num_annotated_sites <- nrow(unique(bind_sites[, c(1:3,6)])) -->
- <!-- # the bind sites are predicted bound within at least one condition -->
- <!-- bind_sites <- bind_sites[(bind_sites$V19 + bind_sites$V20) > 0,] -->
- <!-- DTF_stat$num_bound_sites <- nrow(unique(bind_sites[, c(1:3,6)])) -->
- <!-- bind_sites$activity <- "" -->
- <!-- bind_sites[bind_sites$V30 != ".",]$activity <- str_remove(sapply(str_split(bind_sites[bind_sites$V30 != ".",]$V30,";"),`[[`, 1),"activity=") -->
- <!-- bind_sites$regulatory_feature_id <- "" -->
- <!-- bind_sites[bind_sites$V30 != ".",]$regulatory_feature_id <- str_remove(sapply(str_split(bind_sites[bind_sites$V30 != ".",]$V30,";"),`[[`, 7),"regulatory_feature_stable_id=") -->
- <!-- # the bind sites that are annotated as promoters -->
- <!-- # remove the false positive which the cell specify annotation activity=NA -->
- <!-- promoter_overlapped_bind_sites <- unique(bind_sites[str_detect(bind_sites$V24,"promoter$|promoter2000bp") & bind_sites$activity != "NA",c(1:6, 17:21,24,which(colnames(bind_sites) %in% c("activity","regulatory_feature_id")))]) -->
- <!-- DTF_stat$num_overlapped_promoters <- nrow(unique(promoter_overlapped_bind_sites[,c(1:11)])) -->
- <!-- # the bind sites that are annotated as enhancers -->
- <!-- enhancer_overlapped_bind_sites <- unique(bind_sites[bind_sites$V24 == "enhancer" & bind_sites$activity != "NA",c(1:6, 17:21,24,which(colnames(bind_sites) %in% c("activity","regulatory_feature_id")))]) -->
- <!-- DTF_stat$num_overlapped_enhancers <- nrow(unique(enhancer_overlapped_bind_sites[,c(1:11)])) -->
- <!-- colnames(promoter_overlapped_bind_sites) <- c("chr", "start", "end", "motif", "motif_score", "strand", "DEL_score", "WT_score", "DEL_bound", "WT_bound", "DEL_vs_WT_log2fc", "annotation", "activity", "regulatory_feature_id") -->
- <!-- colnames(enhancer_overlapped_bind_sites) <- c("chr", "start", "end", "motif", "motif_score", "strand", "DEL_score", "WT_score", "DEL_bound", "WT_bound", "DEL_vs_WT_log2fc", "annotation", "activity", "regulatory_feature_id") -->
- <!-- # add background info -->
- <!-- background <- str_remove(str_remove(basename(dirname(DTF)), ".DEL_vs_WT.conservative"), "DTFB.") -->
- <!-- promoter_overlapped_bind_sites$background <- background -->
- <!-- enhancer_overlapped_bind_sites$background <- background -->
- <!-- DTF_stat$background <- background -->
- <!-- promoter_overlapped_bind_sites_all_DTFs <- rbind(promoter_overlapped_bind_sites_all_DTFs, promoter_overlapped_bind_sites) -->
- <!-- enhancer_overlapped_bind_sites_all_DTFs <- rbind(enhancer_overlapped_bind_sites_all_DTFs, enhancer_overlapped_bind_sites) -->
- <!-- stat_summary <- rbind(stat_summary, DTF_stat) -->
- <!-- } -->
- <!-- writexl::write_xlsx(stat_summary, "TF_footprint/annotation_analysis/DTF_annotation_stat_summary.xlsx") -->
- <!-- writexl::write_xlsx(promoter_overlapped_bind_sites_all_DTFs, "TF_footprint/annotation_analysis/promoter_overlapped_bind_sites_all_DTFs_with_annotation.xlsx") -->
- <!-- writexl::write_xlsx(enhancer_overlapped_bind_sites_all_DTFs, "TF_footprint/annotation_analysis/enhancer_overlapped_bind_sites_all_DTFs_with_annotation.xlsx") -->
- <!-- writexl::write_xlsx(unique(promoter_overlapped_bind_sites_all_DTFs[,c(1:11,15)]), "TF_footprint/annotation_analysis/promoter_overlapped_bind_sites_all_DTFs.xlsx") -->
- <!-- writexl::write_xlsx(unique(enhancer_overlapped_bind_sites_all_DTFs[,c(1:11,15)]), "TF_footprint/annotation_analysis/enhancer_overlapped_bind_sites_all_DTFs.xlsx") -->
- <!-- ``` -->
- <!-- ### Plot overall stat -->
- <!-- ```{r} -->
- <!-- stat_summary <- readxl::read_excel("TF_footprint/annotation_analysis/DTF_annotation_stat_summary.xlsx") -->
- <!-- stat_summary_long <- tidyr::pivot_longer(stat_summary, names_to = "category", values_to = "Number", cols = total_bind_sites:num_overlapped_enhancers) -->
- <!-- ggplot(stat_summary_long, aes(x = category, y = Number)) + -->
- <!-- geom_boxplot() + -->
- <!-- labs(y = "Total number of bind sites", x = "") + -->
- <!-- facet_wrap(~background) + -->
- <!-- theme(axis.text.x = element_text(angle = 45, hjust=1)) -->
- <!-- stat_summary_long_tbl <- stat_summary_long %>% group_by(background, category) %>% summarise(num_motifs = n(), mean_num_bind_sites=round(mean(Number)), total_num_bind_sites=sum(Number)) -->
- <!-- stat_summary_tbl <- tidyr::pivot_wider(stat_summary_long_tbl, names_from = category, values_from = c("mean_num_bind_sites", "total_num_bind_sites")) -->
- <!-- writexl::write_xlsx(stat_summary_tbl, "TF_footprint/annotation_analysis/stat_summary_tbl.xlsx") -->
- <!-- ``` -->
- <!-- ### find promoter target genes -->
- <!-- ```{r} -->
- <!-- require(ensembldb) -->
- <!-- edb <- EnsDb("data/Homo_sapiens.GRCh38.92.sqlite") -->
- <!-- allpromoters <- promoters(edb, upstream = 1500, downstream = 500) -->
- <!-- gene_annotation <- readRDS("../iN_05_2021/results/gene_annotation.rds") -->
- <!-- gene_annotation$symbol <- toupper(gene_annotation$symbol) -->
- <!-- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/del_vs_wt_overlapped_GM_MGH.csv") -->
- <!-- del_vs_wt_final_degs$symbol <- toupper(del_vs_wt_final_degs$symbol) -->
- <!-- overlapped_bind_sites <- readxl::read_excel("TF_footprint/annotation_analysis/promoter_overlapped_bind_sites_all_DTFs.xlsx") -->
- <!-- identify_promoter_target_genes <- function(overlapped_bind_sites, background, promoter_regions, gene_annotation, final_degs) { -->
- <!-- bind_sites <- overlapped_bind_sites[overlapped_bind_sites$background == background,] -->
- <!-- message("number of bind sites: ", nrow(bind_sites)) -->
- <!-- log2fc_quantile <- quantile(bind_sites$DEL_vs_WT_log2fc, prob = seq(0,1, by = 0.05)) -->
- <!-- print(log2fc_quantile) -->
- <!-- bind_sites <- bind_sites[bind_sites$DEL_vs_WT_log2fc < log2fc_quantile[["10%"]] | bind_sites$DEL_vs_WT_log2fc > log2fc_quantile[["90%"]],] -->
- <!-- message("number of bind sites in top quantitle: ", nrow(bind_sites)) -->
- <!-- bind_sites_gr <- GRanges(seqnames = bind_sites$chr, ranges = IRanges(start = bind_sites$start, end = bind_sites$end), strand = bind_sites$strand) -->
- <!-- bind_sites_vs_promoters <- bind_sites[queryHits(findOverlaps(bind_sites_gr, allpromoters)),] -->
- <!-- bind_sites_vs_promoters$target_gene_id <- allpromoters[subjectHits(findOverlaps(bind_sites_gr, allpromoters)),]$gene_id -->
- <!-- bind_sites_vs_promoters <- left_join(bind_sites_vs_promoters, gene_annotation[, c("ensembl_id", "symbol", "gene_type")], by = c("target_gene_id" = "ensembl_id")) -->
- <!-- bind_sites_vs_promoters <- left_join(bind_sites_vs_promoters, final_degs[, c("ensembl_id", "padj", "log2FoldChange")], by = c("target_gene_id" = "ensembl_id")) -->
- <!-- writexl::write_xlsx(bind_sites_vs_promoters, file.path("TF_footprint/annotation_analysis", paste0(background, "_bind_sites_vs_promoters.xlsx"))) -->
- <!-- return (bind_sites_vs_promoters) -->
- <!-- } -->
- <!-- iN_GM_promoter_targets <- identify_promoter_target_genes(overlapped_bind_sites, "iN_GM", allpromoters, gene_annotation, del_vs_wt_final_degs) -->
- <!-- iN_MGH_promoter_targets <- identify_promoter_target_genes(overlapped_bind_sites, "iN_MGH", allpromoters, gene_annotation, del_vs_wt_final_degs) -->
- <!-- ``` -->
- <!-- ### Run enrichment analysis for promoter target DEGs -->
- <!-- ```{r} -->
- <!-- source("../../MyTools/RNAseq.analysis/R/enrichment.R") -->
- <!-- outfolder <- "TF_footprint/annotation_analysis/promoter_targets_enrichment_analysis" -->
- <!-- if (!dir.exists(outfolder)) { -->
- <!-- dir.create(outfolder, recursive = T) -->
- <!-- } -->
- <!-- iN_gene_table <- data.frame(ensembl_id = intersect(del_vs_wt_GM$ensembl_id, del_vs_wt_MGH$ensembl_id)) -->
- <!-- iN_gene_table$symbol <- gene_annotation[iN_gene_table$ensembl_id,]$symbol -->
- <!-- iN_gene_table$gene_type <- gene_annotation[iN_gene_table$ensembl_id,]$gene_type -->
- <!-- # set padj and log2FoldChange for promoter target degs -->
- <!-- iN_GM_promoter_targets_DEGs <- iN_gene_table -->
- <!-- iN_GM_promoter_targets_DEGs <- left_join(iN_GM_promoter_targets_DEGs, del_vs_wt_final_degs[del_vs_wt_final_degs$symbol %in% iN_GM_promoter_targets$symbol,c("ensembl_id", "padj", "log2FoldChange")], by = "ensembl_id") -->
- <!-- iN_GM_promoter_targets_DEGs$padj[is.na(iN_GM_promoter_targets_DEGs$padj)] <- 1 -->
- <!-- iN_GM_promoter_targets_DEGs$log2FoldChange[is.na(iN_GM_promoter_targets_DEGs$log2FoldChange)] <- 0 -->
- <!-- # set padj and log2FoldChange for promoter target degs -->
- <!-- iN_MGH_promoter_targets_DEGs <- iN_gene_table -->
- <!-- iN_MGH_promoter_targets_DEGs <- left_join(iN_MGH_promoter_targets_DEGs, del_vs_wt_final_degs[del_vs_wt_final_degs$symbol %in% iN_MGH_promoter_targets$symbol,c("ensembl_id", "padj", "log2FoldChange")], by = "ensembl_id") -->
- <!-- iN_MGH_promoter_targets_DEGs$padj[is.na(iN_MGH_promoter_targets_DEGs$padj)] <- 1 -->
- <!-- iN_MGH_promoter_targets_DEGs$log2FoldChange[is.na(iN_MGH_promoter_targets_DEGs$log2FoldChange)] <- 0 -->
- <!-- promoter_targets <- list(iN_GM_promoter_targets_DEGs = iN_GM_promoter_targets_DEGs, -->
- <!-- iN_MGH_promoter_targets_DEGs = iN_MGH_promoter_targets_DEGs) -->
- <!-- thresholds <- as.list(rep(0.1, length(promoter_targets))) -->
- <!-- names(thresholds) <- names(promoter_targets) -->
- <!-- protein_coding_option <- "yes" -->
- <!-- option <- "all_fdr" -->
- <!-- gene_lst_to_remove <- c(all_genes = "XXX") -->
- <!-- load("../../MyTools/enrichment_analysis/data/pathway_db/pathway_db.Rdata") -->
- <!-- top_enrichment_terms <- run_enrichment(promoter_targets, thresholds, protein_coding_option, option, gene_lst_to_remove, no = 20, outfolder = outfolder, outplotfolder = file.path(outfolder, "plots")) -->
- <!-- saveRDS(top_enrichment_terms, file.path(outfolder, "top_enrichment_terms.rds")) -->
- <!-- enrichment_summary(top_enrichment_terms,input_names = names(promoter_targets), fdr = 0.1, output = file.path(outfolder, "sig_enrichment_terms.xlsx")) -->
- <!-- ``` -->
- <!-- ### plot significant terms for promoter target DEGs -->
- <!-- ```{r} -->
- <!-- sig_enrichment_terms <- readxl::read_excel("TF_footprint/annotation_analysis/promoter_targets_enrichment_analysis/sig_enrichment_terms.xlsx") -->
- <!-- iN_GM_sig <- sig_enrichment_terms[sig_enrichment_terms$module == "iN_GM_promoter_targets_DEGs",] -->
- <!-- iN_GM_sig$name <- make.unique(iN_GM_sig$name) -->
- <!-- iN_GM_sig <- iN_GM_sig[order(iN_GM_sig$DB,iN_GM_sig$logBH),] -->
- <!-- iN_GM_sig$name <- factor(iN_GM_sig$name, levels = iN_GM_sig$name) -->
- <!-- ggplot(iN_GM_sig, aes(x = name, y = logBH, fill = DB)) + -->
- <!-- geom_col(position = "dodge") + -->
- <!-- geom_text(aes(label=paste0(up_sig,"/",down_sig)), position = position_stack(vjust = .5)) + -->
- <!-- coord_flip() + -->
- <!-- theme(axis.title.x = element_blank(), axis.text.x = element_text(size = 6)) -->
- <!-- ggsave("TF_footprint/annotation_analysis/promoter_targets_enrichment_analysis/iN_GM_sig_terms.pdf",height = 10, width = 10) -->
- <!-- iN_MGH_sig <- sig_enrichment_terms[sig_enrichment_terms$module == "iN_MGH_promoter_targets_DEGs",] -->
- <!-- iN_MGH_sig$name <- make.unique(iN_MGH_sig$name) -->
- <!-- iN_MGH_sig <- iN_MGH_sig[order(iN_MGH_sig$DB,iN_MGH_sig$logBH),] -->
- <!-- iN_MGH_sig$name <- factor(iN_MGH_sig$name, levels = iN_MGH_sig$name) -->
- <!-- ggplot(iN_MGH_sig, aes(x = name, y = logBH, fill = DB)) + -->
- <!-- geom_col(position = "dodge") + -->
- <!-- geom_text(aes(label=paste0(up_sig,"/",down_sig)), position = position_stack(vjust = .5), size = 3) + -->
- <!-- coord_flip() + -->
- <!-- theme(axis.title.x = element_blank(), axis.text.x = element_text(size = 6)) -->
- <!-- ggsave("TF_footprint/annotation_analysis/promoter_targets_enrichment_analysis/iN_MGH_sig_terms.pdf", height = 10, width = 10) -->
- <!-- ``` -->
- <!-- ### Make a table for DTF promoter target degs and their associate DTF -->
- <!-- ```{r} -->
- <!-- iN_GM_promoter_targets <- readxl::read_excel("TF_footprint/annotation_analysis/iN_GM_bind_sites_vs_promoters.xlsx") -->
- <!-- iN_GM_promoter_targets_DEGs <- unique(iN_GM_promoter_targets[!is.na(iN_GM_promoter_targets$padj),]) -->
- <!-- iN_GM_promoter_targets_DEGs <- iN_GM_promoter_targets_DEGs[order(iN_GM_promoter_targets_DEGs$symbol),] -->
- <!-- writexl::write_xlsx(iN_GM_promoter_targets_DEGs, "TF_footprint/annotation_analysis/iN_GM_promoter_targes_degs_with_DTF.xlsx") -->
- <!-- iN_MGH_promoter_targets <- readxl::read_excel("TF_footprint/annotation_analysis/iN_MGH_bind_sites_vs_promoters.xlsx") -->
- <!-- iN_MGH_promoter_targets_DEGs <- unique(iN_MGH_promoter_targets[!is.na(iN_MGH_promoter_targets$padj),]) -->
- <!-- iN_MGH_promoter_targets_DEGs <- iN_MGH_promoter_targets_DEGs[order(iN_MGH_promoter_targets_DEGs$symbol),] -->
- <!-- writexl::write_xlsx(iN_MGH_promoter_targets_DEGs, "TF_footprint/annotation_analysis/iN_MGH_promoter_targets_degs_with_DTF.xlsx") -->
- <!-- ``` -->
- <!-- ### Find promoter target DEGs enrichment terms that are overlapped in GM vs MGH -->
- <!-- ```{r} -->
- <!-- sig_enrichment_terms <- readxl::read_excel("TF_footprint/annotation_analysis_iN/promoter_targets_enrichment_analysis/sig_enrichment_terms.xlsx") -->
- <!-- db_names <- unique(sig_enrichment_terms$DB) -->
- <!-- overlapped_terms_tbl <- data.frame() -->
- <!-- for (db_name in db_names) { -->
- <!-- GM_sig_terms <- sig_enrichment_terms[sig_enrichment_terms$module == "iN_GM_promoter_targets_DEGs" & sig_enrichment_terms$DB == db_name,] -->
- <!-- MGH_sig_terms <- sig_enrichment_terms[sig_enrichment_terms$module == "iN_MGH_promoter_targets_DEGs" & sig_enrichment_terms$DB == db_name,] -->
- <!-- overlapped_terms <- intersect(GM_sig_terms$name, MGH_sig_terms$name) -->
- <!-- if (length(overlapped_terms) > 0) { -->
- <!-- overlapped_terms_tbl <- rbind(overlapped_terms_tbl, data.frame(DB = db_name, name = overlapped_terms, -->
- <!-- up_sig_GM = GM_sig_terms[GM_sig_terms$name %in% overlapped_terms,]$up_sig, -->
- <!-- down_sig_GM = GM_sig_terms[GM_sig_terms$name %in% overlapped_terms,]$down_sig, -->
- <!-- up_sig_MGH = MGH_sig_terms[MGH_sig_terms$name %in% overlapped_terms,]$up_sig, -->
- <!-- down_sig_MGH = MGH_sig_terms[MGH_sig_terms$name %in% overlapped_terms,]$down_sig)) -->
- <!-- } -->
- <!-- } -->
- <!-- writexl::write_xlsx(overlapped_terms_tbl, "TF_footprint/annotation_analysis_iN/promoter_targets_enrichment_analysis/sig_enrichment_terms_shared_in_GM_MGH.xlsx") -->
- <!-- ``` -->
- <!-- ### find enhancer target genes/ontology on GREAT -->
- <!-- ```{r} -->
- <!-- identify_enhancer_target_regions <- function(overlapped_bind_sites, background) { -->
- <!-- bind_sites <- overlapped_bind_sites[overlapped_bind_sites$background == background,] -->
- <!-- message("number of bind sites: ", nrow(bind_sites)) -->
- <!-- log2fc_quantile <- quantile(bind_sites$DEL_vs_WT_log2fc, prob = seq(0,1, by = 0.05)) -->
- <!-- print(log2fc_quantile) -->
- <!-- bind_sites <- bind_sites[bind_sites$DEL_vs_WT_log2fc < log2fc_quantile[["10%"]] | bind_sites$DEL_vs_WT_log2fc > log2fc_quantile[["90%"]],] -->
- <!-- message("number of bind sites in top quantitle: ", nrow(bind_sites)) -->
- <!-- bind_sites$chr <- str_c("chr", bind_sites$chr) -->
- <!-- bind_sites$start <- bind_sites$start-1 -->
- <!-- bind_sites$score <- round(bind_sites$motif_score) -->
- <!-- bind_sites$name <- paste(bind_sites$chr, bind_sites$start, bind_sites$end, bind_sites$motif, sep = "_") -->
- <!-- write.table(unique(bind_sites[, c("chr", "start", "end", "name", "score", "strand")]), file.path("TF_footprint/annotation_analysis/enhancer_analysis", paste0(background, "_enhancer_overlapped_bind_sites.bed")), sep = "\t", quote = F, col.names = F, row.names = F) -->
- <!-- return(bind_sites) -->
- <!-- } -->
- <!-- overlapped_bind_sites <- readxl::read_excel("TF_footprint/annotation_analysis/enhancer_overlapped_bind_sites_all_DTFs.xlsx") -->
- <!-- iN_GM_enhancer_targets <- identify_enhancer_target_regions(overlapped_bind_sites, "iN_GM") -->
- <!-- iN_MGH_enhancer_targets <- identify_enhancer_target_regions(overlapped_bind_sites, "iN_MGH") -->
- <!-- # upload bed files to GREAT and export region-gene association table -->
- <!-- ``` -->
- <!-- ### Read region-gene association table, get enhancer target genes -->
- <!-- ```{r} -->
- <!-- iN_GM_enhancer_target_genes <- read.table("TF_footprint/annotation_analysis/enhancer_analysis/iN_GM_enhancer_targets.txt", sep = "\t", col.names = c("region", "gene")) -->
- <!-- iN_GM_enhancer_target_gene_symbols <- unique(unlist(lapply(str_split(iN_GM_enhancer_target_genes[iN_GM_enhancer_target_genes$gene != "NONE", ]$gene,","), str_remove_all, "\\s+|\\(.*\\)"))) -->
- <!-- iN_MGH_enhancer_target_genes <- read.table("TF_footprint/annotation_analysis/enhancer_analysis/iN_MGH_enhancer_targets.txt", sep = "\t", col.names = c("region", "gene")) -->
- <!-- iN_MGH_enhancer_target_gene_symbols <- unique(unlist(lapply(str_split(iN_MGH_enhancer_target_genes[iN_MGH_enhancer_target_genes$gene != "NONE", ]$gene,","), str_remove_all, "\\s+|\\(.*\\)"))) -->
- <!-- ``` -->
- <!-- ### Check whether enhancer targets and DEGs are significant overlapped -->
- <!-- ```{r} -->
- <!-- del_vs_wt_GM <- read.csv("../POGZ_paper/results/DEL_vs_WT_GM__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1) -->
- <!-- del_vs_wt_MGH <- read.csv("../POGZ_paper/results/DEL_vs_WT_MGH__simple_model_sva/DEL_vs_WT_threshold_0.5_sv.csv", row.names = 1) -->
- <!-- analyzed_genes <- intersect(toupper(del_vs_wt_GM$symbol), toupper(del_vs_wt_MGH$symbol)) -->
- <!-- del_vs_wt_final_degs <- read.csv("../POGZ_paper/results/del_vs_wt_overlapped_GM_MGH.csv") -->
- <!-- del_vs_wt_final_degs$symbol <- toupper(del_vs_wt_final_degs$symbol) -->
- <!-- final_degs <- del_vs_wt_final_degs$symbol -->
- <!-- # iN_GM -->
- <!-- target_genes <- intersect(iN_GM_enhancer_target_gene_symbols, analyzed_genes) -->
- <!-- target_degs <- intersect(target_genes, final_degs) -->
- <!-- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater") -->
- <!-- require(VennDiagram) -->
- <!-- venn.diagram(list("enhancer_targets" = target_genes, "DEGs" = final_degs), main = "iN_GM_enhancer_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/iN_GM_enhancer_targes_vs_degs.png") -->
- <!-- # iN_MGH -->
- <!-- target_genes <- intersect(iN_MGH_enhancer_target_gene_symbols, analyzed_genes) -->
- <!-- target_degs <- intersect(target_genes, final_degs) -->
- <!-- targets_vs_degs <- fisher.test(matrix(c(length(target_degs), length(target_genes)-length(target_degs), length(final_degs)-length(target_degs), length(analyzed_genes)-length(final_degs)-length(target_genes)+length(target_degs)), nrow=2),alternative = "greater") -->
- <!-- venn.diagram(list("enhancer_targets" = target_genes, "DEGs" = final_degs), main = "iN_MGH_enhancer_targets vs DEGs", sub = paste0("p.value = ", format(targets_vs_degs$p.value, digits=3)), filename = "TF_footprint/annotation_analysis/iN_MGH_enhancer_targes_vs_degs.png") -->
- <!-- ``` -->
- <!-- ### Run enrichment analysis for enhancer target DEGs -->
- <!-- ```{r} -->
- <!-- source("../../MyTools/RNAseq.analysis/R/enrichment.R") -->
- <!-- outfolder <- "TF_footprint/annotation_analysis/enhancer_targets_enrichment_analysis" -->
- <!-- if (!dir.exists(outfolder)) { -->
- <!-- dir.create(outfolder, recursive = T) -->
- <!-- } -->
- <!-- gene_annotation <- readRDS("../iN_05_2021/results/gene_annotation.rds") -->
- <!-- gene_annotation$symbol <- toupper(gene_annotation$symbol) -->
- <!-- iN_gene_table <- data.frame(ensembl_id = intersect(del_vs_wt_GM$ensembl_id, del_vs_wt_MGH$ensembl_id)) -->
- <!-- iN_gene_table$symbol <- gene_annotation[iN_gene_table$ensembl_id,]$symbol -->
- <!-- iN_gene_table$gene_type <- gene_annotation[iN_gene_table$ensembl_id,]$gene_type -->
- <!-- # set padj and log2FoldChange for enhancer target genes MGH -->
- <!-- iN_GM_enhancer_targets <- iN_gene_table -->
- <!-- iN_GM_enhancer_targets$padj <- 1 -->
- <!-- iN_GM_enhancer_targets[iN_GM_enhancer_targets$symbol %in% iN_GM_enhancer_target_gene_symbols,]$padj <- 0 -->
- <!-- iN_GM_enhancer_targets$log2FoldChange <- 1 -->
- <!-- # set padj and log2FoldChange for enhancer target genes MGH -->
- <!-- iN_MGH_enhancer_targets <- iN_gene_table -->
- <!-- iN_MGH_enhancer_targets$padj <- 1 -->
- <!-- iN_MGH_enhancer_targets[iN_MGH_enhancer_targets$symbol %in% iN_MGH_enhancer_target_gene_symbols,]$padj <- 0 -->
- <!-- iN_MGH_enhancer_targets$log2FoldChange <- 1 -->
- <!-- # set padj and log2FoldChange for enhancer target degs -->
- <!-- iN_GM_enhancer_targets_DEGs <- iN_gene_table -->
- <!-- iN_GM_enhancer_targets_DEGs <- left_join(iN_GM_enhancer_targets_DEGs, del_vs_wt_final_degs[del_vs_wt_final_degs$symbol %in% iN_GM_enhancer_target_gene_symbols,c("ensembl_id", "padj", "log2FoldChange")], by = "ensembl_id") -->
- <!-- iN_GM_enhancer_targets_DEGs$padj[is.na(iN_GM_enhancer_targets_DEGs$padj)] <- 1 -->
- <!-- iN_GM_enhancer_targets_DEGs$log2FoldChange[is.na(iN_GM_enhancer_targets_DEGs$log2FoldChange)] <- 0 -->
- <!-- # set padj and log2FoldChange for enhancer target degs -->
- <!-- iN_MGH_enhancer_targets_DEGs <- iN_gene_table -->
- <!-- iN_MGH_enhancer_targets_DEGs <- left_join(iN_MGH_enhancer_targets_DEGs, del_vs_wt_final_degs[del_vs_wt_final_degs$symbol %in% iN_MGH_enhancer_target_gene_symbols,c("ensembl_id", "padj", "log2FoldChange")], by = "ensembl_id") -->
- <!-- iN_MGH_enhancer_targets_DEGs$padj[is.na(iN_MGH_enhancer_targets_DEGs$padj)] <- 1 -->
- <!-- iN_MGH_enhancer_targets_DEGs$log2FoldChange[is.na(iN_MGH_enhancer_targets_DEGs$log2FoldChange)] <- 0 -->
- <!-- enhancer_targets <- list(iN_GM_enhancer_targets = iN_GM_enhancer_targets, -->
- <!-- iN_MGH_enhancer_targets = iN_MGH_enhancer_targets, -->
- <!-- iN_GM_enhancer_targets_DEGs = iN_GM_enhancer_targets_DEGs, -->
- <!-- iN_MGH_enhancer_targets_DEGs = iN_MGH_enhancer_targets_DEGs) -->
- <!-- thresholds <- as.list(rep(0.1, length(enhancer_targets))) -->
- <!-- names(thresholds) <- names(enhancer_targets) -->
- <!-- protein_coding_option <- "yes" -->
- <!-- option <- "all_fdr" -->
- <!-- gene_lst_to_remove <- c(all_genes = "XXX") -->
- <!-- load("../../MyTools/enrichment_analysis/data/pathway_db/pathway_db.Rdata") -->
- <!-- top_enrichment_terms <- run_enrichment(enhancer_targets, thresholds, protein_coding_option, option, gene_lst_to_remove, no = 20, outfolder = outfolder, outplotfolder = file.path(outfolder, "plots")) -->
- <!-- saveRDS(top_enrichment_terms, file.path(outfolder, "top_enrichment_terms.rds")) -->
- <!-- enrichment_summary(top_enrichment_terms[str_detect(names(top_enrichment_terms), "_DEGs")],input_names = c("iN_GM_enhancer_targets_DEGs","iN_MGH_enhancer_targets_DEGs"), fdr = 0.1, output = file.path(outfolder, "enhancer_DEGs_sig_enrichment_terms.xlsx")) -->
- <!-- enrichment_summary(top_enrichment_terms[!str_detect(names(top_enrichment_terms), "_DEGs")],input_names = c("iN_GM_enhancer_targets","iN_MGH_enhancer_targets"), fdr = 0.1, output = file.path(outfolder, "enhaner_genes_sig_enrichment_terms.xlsx")) -->
- <!-- ``` -->
tf_binding.Rmd at commit 29381f6, no license · at the source
Overview
17 affiliations
- Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA
- Department of Neurology, Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA
- Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, MA, USA
- Program in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA
- Program in Biological and Biomedical Sciences, Harvard Medical School, Boston, MA, USA
- Division of Genetics and Genomics, Boston Children’s Hospital, Boston, MA, USA
- Analytic and Translational Genetics Unit, Department of Medicine, Massachusetts General Hospital, Boston, MA, USA
- Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY, USA
- Nash Family Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA
- Mount Sinai Center for Transformative Disease Modeling, Icahn School of Medicine at Mount Sinai, New York, NY, USA
- Friedman Brain Institute, Icahn School of Medicine at Mount Sinai, New York, NY, USA
- Icahn Institute for Data Science and Genomic Technology, Icahn School of Medicine at Mount Sinai, New York, NY, USA
- Division of Genetic Medicine, Department of Medicine, Vanderbilt Genetics Institute, Vanderbilt University Medical Center, Medical Center Dr., Nashville, TN 1211, USA
- Department of Biomedical Informatics and Department of Psychiatry and Behavioral Sciences, Vanderbilt University Medical Center, Medical Center Dr., Nashville, TN 1211, USA
- Department of Psychiatry, Yale University New Haven, New Haven, CT, USA
- Department of Genetics, Blavatnik Institute, Harvard Medical School, Boston, MA, USA
- Harvard Stem Cell Institute, Harvard University, Cambridge, MA, USA
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repositories
Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.
Zenodo 3564813
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
147 files
- atac-seq-pipeline-1.1.7.
zip/ , Shell, 255 linesconda/ build_genome_data.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 98 linesconda/ install_dependencies.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 28 linesconda/ uninstall_dependencies.s h - atac-seq-pipeline-1.1.7.
zip/ , Shell, 35 linesconda/ update_conda_env.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 69 linesexamples/ local/ ENCSR356KRQ_subsampled_s ge_conda.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 66 linesexamples/ local/ ENCSR356KRQ_subsampled_s ge_singularity.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 65 linesexamples/ local/ ENCSR356KRQ_subsampled_s lurm_conda.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 62 linesexamples/ local/ ENCSR356KRQ_subsampled_s lurm_singularity.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 65 linesexamples/ scg/ ENCSR356KRQ_subsampled_s cg_conda.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 62 linesexamples/ scg/ ENCSR356KRQ_subsampled_s cg_singularity.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 65 linesexamples/ sherlock/ ENCSR356KRQ_subsampled_s herlock_conda.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 62 linesexamples/ sherlock/ ENCSR356KRQ_subsampled_s herlock_singularity.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 201 linesgenome/ download_genome_data.sh - atac-seq-pipeline-1.1.7.
zip/ , Python, 73 linessrc/ assign_multimappers.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 76 linessrc/ detect_adapter.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 465 linessrc/ encode_ataqc.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 156 linessrc/ encode_bam2ta.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 80 linessrc/ encode_blacklist_filter. py - atac-seq-pipeline-1.1.7.
zip/ , Python, 209 linessrc/ encode_bowtie2.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 285 linessrc/ encode_common.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 403 linessrc/ encode_common_genomic.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 211 linessrc/ encode_common_html.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 467 linessrc/ encode_common_log_parser .py - atac-seq-pipeline-1.1.7.
zip/ , Python, 81 linessrc/ encode_count_signal_trac k.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 391 linessrc/ encode_filter.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 107 linessrc/ encode_frip.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 169 linessrc/ encode_idr.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 219 linessrc/ encode_macs2_atac.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 161 linessrc/ encode_macs2_signal_trac k_atac.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 143 linessrc/ encode_naive_overlap.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 59 linessrc/ encode_pool_ta.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 617 linessrc/ encode_qc_report.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 160 linessrc/ encode_reproducibility_q c.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 132 linessrc/ encode_spr.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 265 linessrc/ encode_trim_adapter.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 102 linessrc/ encode_xcor.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 1,704 lines, 1 matchsrc/ run_ataqc.py - atac-seq-pipeline-1.1.7.
zip/ , Shell, 15 linestest/ run_cromwell_server_on_g c.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 4 linestest/ test_task/ download_hg38_fasta_for_ test_ataqc.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 55 linestest/ test_task/ test.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 11 linestest/ test_task/ test_all.sh - atac-seq-pipeline-1.1.7.
zip/ , Shell, 49 linestest/ test_workflow/ test_atac.sh - atac-seq-pipeline-1.1.7.
zip/ , Python, 307 linesutils/ qc_jsons_to_tsv/ qc_jsons_to_tsv.py - atac-seq-pipeline-1.1.7.
zip/ , Python, 135 linesutils/ resumer/ resumer.py - atac-seq-pipeline-1.4.2.
zip/ , Shell, 255 linesconda/ build_genome_data.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 74 linesconda/ config_conda_env.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 21 linesconda/ config_conda_env_py3.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 39 linesconda/ install_dependencies.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 28 linesconda/ uninstall_dependencies.s h - atac-seq-pipeline-1.4.2.
zip/ , Shell, 34 linesconda/ update_conda_env.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 69 linesexamples/ local/ ENCSR356KRQ_subsampled_s ge_conda.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 66 linesexamples/ local/ ENCSR356KRQ_subsampled_s ge_singularity.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 65 linesexamples/ local/ ENCSR356KRQ_subsampled_s lurm_conda.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 62 linesexamples/ local/ ENCSR356KRQ_subsampled_s lurm_singularity.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 66 linesexamples/ scg/ ENCSR356KRQ_subsampled_s cg_conda.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 62 linesexamples/ scg/ ENCSR356KRQ_subsampled_s cg_singularity.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 65 linesexamples/ sherlock/ ENCSR356KRQ_subsampled_s herlock_conda.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 62 linesexamples/ sherlock/ ENCSR356KRQ_subsampled_s herlock_singularity.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 201 linesgenome/ download_genome_data.sh - atac-seq-pipeline-1.4.2.
zip/ , Python, 73 linessrc/ assign_multimappers.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 76 linessrc/ detect_adapter.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 469 linessrc/ encode_ataqc.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 156 linessrc/ encode_bam2ta.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 80 linessrc/ encode_blacklist_filter. py - atac-seq-pipeline-1.4.2.
zip/ , Python, 214 linessrc/ encode_bowtie2.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 285 linessrc/ encode_common.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 406 linessrc/ encode_common_genomic.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 211 linessrc/ encode_common_html.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 467 linessrc/ encode_common_log_parser .py - atac-seq-pipeline-1.4.2.
zip/ , Python, 81 linessrc/ encode_count_signal_trac k.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 391 linessrc/ encode_filter.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 107 linessrc/ encode_frip.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 169 linessrc/ encode_idr.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 134 linessrc/ encode_macs2_atac.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 163 linessrc/ encode_macs2_signal_trac k_atac.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 143 linessrc/ encode_naive_overlap.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 59 linessrc/ encode_pool_ta.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 633 linessrc/ encode_qc_report.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 160 linessrc/ encode_reproducibility_q c.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 132 linessrc/ encode_spr.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 265 linessrc/ encode_trim_adapter.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 144 linessrc/ encode_xcor.py - atac-seq-pipeline-1.4.2.
zip/ , Python, 1,735 linessrc/ run_ataqc.py - atac-seq-pipeline-1.4.2.
zip/ , Shell, 15 linestest/ run_cromwell_server_on_g c.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 4 linestest/ test_task/ download_hg38_fasta_for_ test_ataqc.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 55 linestest/ test_task/ test.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 11 linestest/ test_task/ test_all.sh - atac-seq-pipeline-1.4.2.
zip/ , Shell, 50 linestest/ test_workflow/ test_atac.sh - atac-seq-pipeline-1.4.2.
zip/ , Python, 326 linesutils/ qc_jsons_to_tsv/ qc_jsons_to_tsv.py - atac-seq-pipeline-1.5.4.
zip/ , Shell, 48 linesdev/ build_on_dx.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 15 linesdev/ test/ run_cromwell_server_on_g c.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 4 linesdev/ test/ test_task/ download_hg38_fasta_for_ test_ataqc.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 55 linesdev/ test/ test_task/ test.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 12 linesdev/ test/ test_task/ test_all.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 55 linesdev/ test/ test_workflow/ test_atac.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 320 linesscripts/ build_genome_data.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 265 linesscripts/ download_genome_data.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 82 linesscripts/ install_conda_env.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 10 linesscripts/ uninstall_conda_env.sh - atac-seq-pipeline-1.5.4.
zip/ , Shell, 32 linesscripts/ update_conda_env.sh - atac-seq-pipeline-1.5.4.
zip/ , Python, 73 linessrc/ assign_multimappers.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 84 linessrc/ detect_adapter.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 115 linessrc/ encode_lib_blacklist_fil ter.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 359 linessrc/ encode_lib_common.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 114 linessrc/ encode_lib_frip.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 568 linessrc/ encode_lib_genomic.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 660 linessrc/ encode_lib_log_parser.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 261 linessrc/ encode_lib_qc_category.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 109 linessrc/ encode_task_annot_enrich .py - atac-seq-pipeline-1.5.4.
zip/ , Python, 159 linessrc/ encode_task_bam2ta.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 192 linessrc/ encode_task_bowtie2.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 226 linessrc/ encode_task_bwa.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 127 linessrc/ encode_task_choose_ctl.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 137 linessrc/ encode_task_compare_sign al_to_roadmap.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 86 linessrc/ encode_task_count_signal _track.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 364 linessrc/ encode_task_filter.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 86 linessrc/ encode_task_frac_mito.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 244 linessrc/ encode_task_fraglen_stat _pe.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 143 linessrc/ encode_task_gc_bias.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 180 linessrc/ encode_task_idr.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 133 linessrc/ encode_task_jsd.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 118 linessrc/ encode_task_macs2_atac.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 125 linessrc/ encode_task_macs2_chip.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 171 linessrc/ encode_task_macs2_signal _track_atac.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 181 linessrc/ encode_task_macs2_signal _track_chip.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 104 linessrc/ encode_task_merge_fastq. py - atac-seq-pipeline-1.5.4.
zip/ , Python, 152 linessrc/ encode_task_overlap.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 71 linessrc/ encode_task_pool_ta.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 88 linessrc/ encode_task_post_align.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 88 linessrc/ encode_task_post_call_pe ak_atac.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 91 linessrc/ encode_task_post_call_pe ak_chip.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 165 linessrc/ encode_task_preseq.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 927 lines, 1 matchsrc/ encode_task_qc_report.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 185 linessrc/ encode_task_reproducibil ity.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 103 linessrc/ encode_task_spp.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 143 linessrc/ encode_task_spr.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 232 linessrc/ encode_task_trim_adapter .py - atac-seq-pipeline-1.5.4.
zip/ , Python, 73 linessrc/ encode_task_trim_fastq.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 178 linessrc/ encode_task_tss_enrich.p y - atac-seq-pipeline-1.5.4.
zip/ , Python, 156 linessrc/ encode_task_xcor.py - atac-seq-pipeline-1.5.4.
zip/ , Python, 255 linessrc/ trimfastq.py - atac-seq-pipeline-1.1.7.
zip/ , License, 21 linesLICENSE - atac-seq-pipeline-1.1.7.
zip/ , Text, 67 linesREADME.md - atac-seq-pipeline-1.4.2.
zip/ , License, 21 linesLICENSE - atac-seq-pipeline-1.4.2.
zip/ , Text, 99 linesREADME.md - atac-seq-pipeline-1.5.4.
zip/ , License, 21 linesLICENSE - atac-seq-pipeline-1.5.4.
zip/ , Text, 81 linesREADME.md
talkowski-lab/POGZ
29381f6e1324f375ed9a8a75f09dee9231ab0d85, 22 May 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
24 files
- ATAC-seq/
Differential_accessibili , R, 153 lines, 1 matchty/ diff_peaks_v2.R - ATAC-seq/
Differential_accessibili , R, 763 lines, 3 matchesty/ differential_accessibili ty.Rmd - ATAC-seq/
Transcription_factor_foo , Shell, 15 linestprint/ 1. merge_ATACpeaks.sh - ATAC-seq/
Transcription_factor_foo , Shell, 13 linestprint/ 2. get_bamfile_list.sh - ATAC-seq/
Transcription_factor_foo , Shell, 16 linestprint/ 3. call_TF_footprints.sh - ATAC-seq/
Transcription_factor_foo , Shell, 34 linestprint/ 4. iN_merge_bigwigs.sh - ATAC-seq/
Transcription_factor_foo , Shell, 15 lines, 1 matchtprint/ 5. call_TF_bindings.sh - ATAC-seq/
Transcription_factor_foo , R, 896 lines, 3 matchestprint/ tf_binding.Rmd - RNA-seq/
RNAseq.analysis/ , R, 452 linesR/ DESeq2_POGZ_NSC.R - RNA-seq/
RNAseq.analysis/ , R, 134 linesR/ adjust_ModuleEigenegene_ pvalues.R - RNA-seq/
RNAseq.analysis/ , R, 868 lines, 1 matchR/ co_expression_functions. R - RNA-seq/
RNAseq.analysis/ , R, 505 lines, 1 matchR/ enrichment.R - RNA-seq/
RNAseq.analysis/ , R, 416 linesR/ functions_POGZ.R - RNA-seq/
RNAseq.analysis/ , R, 65 linesR/ metaAnalysis_NSC.R - RNA-seq/
RNAseq.analysis/ , R, 70 lines, 1 matchR/ read_pathway_db.R - RNA-seq/
RNAseq.analysis/ , R, 685 lines, 1 matchR/ rnaseq_analysis_function s.R - RNA-seq/
RNAseq.analysis/ , R, 212 linesR/ rnaseq_qc_functions.R - RNA-seq/
RNAseq.analysis/ , R, 529 linesR/ runWGCNA.R - RNA-seq/
co-expression_analysis_a , R, 95 lines, 1 matchllsamples.R - RNA-seq/
co-expression_data.R , R, 102 lines - RNA-seq/
rnaseq_analysis_iN_final , R, 216 lines, 1 match.Rmd - RNA-seq/
runSTAR.sh , Shell, 44 lines, 1 match - RNA-seq/
runTrimmomatic.sh , Shell, 7 lines - README.md, Text, 158 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 164 scripts, each with its path and the digest of its content;
- 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE204778, at NCBI GEO; found in “Data and code availability”
Code and data availability statement
The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to a dataset: NCBI GEO GSE204778
- it points to the authors' code: talkowski-lab/
POGZ
Read it in the paper: doi.org/10.1016/j.xhgg.2026.100652.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 21 authors, 5 keywords, 8 funders, 72 references.
Cite
This paper
Moyses-Oliveira, M., Liu, Y., Erdin, S., Gao, D., Bhavsar, R., Mohajeri, K., O’Keefe, K., Boone, P. M., Xavier, G., Liao, C., Li, A., Yadav, R., Salani, M., Lucente, D., Currall, B., de Esch, C. E., Tai, D. J., Ruderfer, D., Brennand, K. J., . . . Talkowski, M. E. (2026). CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling. HGG advances, 7(4), 100652. https://
BibTeX
@article{moysesoliveira2
author = {Moyses-Oliveira, Mariana and Liu, Yating and Erdin, Serkan and Gao, Dadi and Bhavsar, Riya and Mohajeri, Kiana and O’Keefe, Kathryn and Boone, Philip M and Xavier, Gabriela and Liao, Calwing and Li, Aiqun and Yadav, Rachita and Salani, Monica and Lucente, Diane and Currall, Benjamin and de Esch, Celine EF and Tai, Derek JC and Ruderfer, Douglas and Brennand, Kristen J and Gusella, James F and Talkowski, Michael E},
title = {{CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling}},
journal = {HGG advances},
year = {2026},
month = jul,
volume = {7},
number = {4},
pages = {100652},
publisher = {Elsevier},
issn = {2666-2477},
doi = {10.1016/
url = {https://
pmid = {42433021},
pmcid = {PMC13429924}
}
RIS
TY - JOUR
AU - Moyses-Oliveira, Mariana
AU - Liu, Yating
AU - Erdin, Serkan
AU - Gao, Dadi
AU - Bhavsar, Riya
AU - Mohajeri, Kiana
AU - O’Keefe, Kathryn
AU - Boone, Philip M
AU - Xavier, Gabriela
AU - Liao, Calwing
AU - Li, Aiqun
AU - Yadav, Rachita
AU - Salani, Monica
AU - Lucente, Diane
AU - Currall, Benjamin
AU - de Esch, Celine EF
AU - Tai, Derek JC
AU - Ruderfer, Douglas
AU - Brennand, Kristen J
AU - Gusella, James F
AU - Talkowski, Michael E
TI - CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling
T2 - HGG advances
J2 - HGG Adv
PY - 2026
DA - 2026/
VL - 7
IS - 4
SP - 100652
SN - 2666-2477
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling",
"container-title": "HGG advances",
"author": [
{
"family": "Moyses-Oliveira",
"given": "Mariana"
},
{
"family": "Liu",
"given": "Yating"
},
{
"family": "Erdin",
"given": "Serkan"
},
{
"family": "Gao",
"given": "Dadi"
},
{
"family": "Bhavsar",
"given": "Riya"
},
{
"family": "Mohajeri",
"given": "Kiana"
},
{
"family": "O’Keefe",
"given": "Kathryn"
},
{
"family": "Boone",
"given": "Philip M"
},
{
"family": "Xavier",
"given": "Gabriela"
},
{
"family": "Liao",
"given": "Calwing"
},
{
"family": "Li",
"given": "Aiqun"
},
{
"family": "Yadav",
"given": "Rachita"
},
{
"family": "Salani",
"given": "Monica"
},
{
"family": "Lucente",
"given": "Diane"
},
{
"family": "Currall",
"given": "Benjamin"
},
{
"family": "de Esch",
"given": "Celine EF"
},
{
"family": "Tai",
"given": "Derek JC"
},
{
"family": "Ruderfer",
"given": "Douglas"
},
{
"family": "Brennand",
"given": "Kristen J"
},
{
"family": "Gusella",
"given": "James F"
},
{
"family": "Talkowski",
"given": "Michael E"
}
],
"container-title-short":
"volume": "7",
"issue": "4",
"page": "100652",
"DOI": "10.1016/
"PMID": "42433021",
"PMCID": "PMC13429924",
"ISSN": "2666-2477",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: STAR, pysam, WGCNA, 14 other tools, autism, 4 references
- [2] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: STAR, pysam, BEDTools, 13 other tools, cellular / molecular, 3 references
- [3] doi:10.1126/sciadv.aed2952 [code]
- Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.Journal: Science advancesIn common: STAR, pysam, BEDTools, 12 other tools, cellular / molecular, 3 references
- [4] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: STAR, WGCNA, SAMtools, 12 other tools, 3 references
- [5] doi:10.1038/s41467-026-69944-6 [code]
- Multi-modal dissection of cell-type specific TDP-43 pathology in the motor cortex.Journal: Nature communicationsIn common: pysam, BEDTools, SAMtools, 11 other tools, 5 references
- [6] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: pysam, BEDTools, SAMtools, 12 other tools, cellular / molecular, 2 references
- [7] doi:10.21203/rs.3.rs-9927928/v1 [code]
- Genome-wide and allele-resolved maps of the radial architecture of the mouse genomeJournal: Research Square (preprint)In common: STAR, pysam, BEDTools, 12 other tools, 2 references
- [8] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: pysam, WGCNA, edgeR, 11 other tools, cellular / molecular, 3 references
- [9] doi:10.1038/s41467-026-72598-z [code]
- Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.Journal: Nature communicationsIn common: WGCNA, BEDTools, SAMtools, 10 other tools, 4 references
- [10] doi:10.1038/s41467-026-71790-5 [code]
- Recurrent DNA break clusters drive replication-stress-induc
ed copy number variants and genome diversification. Journal: Nature communicationsIn common: pysam, BEDTools, SAMtools, 12 other tools, cellular / molecular
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 164 scripts, and 17 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:76685983a0d356a9…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
