KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease.
The 27 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Results › Inference of cCRE-gene associations across microglial states ↔ Scripts/cCRE-TargetGene_analysis.R, lines 1098–1184 · score 0.83 · GTPase, neutrophil degranulation, neuronal system, RHO, phase, cycle
- [2] § Results › Cis-regulatory landscapes of iTF-microglia across differentiation and activation ↔ Scripts/cCRE_analysis.R, lines 227–315 · score 0.83 · CA CTCF, CA H3K4me3, CA TF, dELS, pELS, cCRE
- [3] § Results › Cis-regulatory landscapes of iTF-microglia across differentiation and activation ↔ Scripts/TF_monaLisa_analysis.R, lines 199–242 · score 0.82 · CA CTCF, CA H3K4me3, CA TF, dELS, pELS, cCRE
- [4] § Results › Cis-regulatory landscapes of iTF-microglia across differentiation and activation ↔ Scripts/cCRE_analysis.R, lines 419–462 · score 0.79 · cell cycle, dELS, pELS, cytokine signaling, chemokine, translation
- [5] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of candidate cis-regulatory elements (cCREs) ↔ Scripts/TF_monaLisa_analysis.R, lines 41–81 · score 0.78 · regionCounts, windowCounts, background bins, H3K27ac, TMM, windows
- [6] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of target genes ↔ Scripts/cCRE-TargetGene_analysis.R, lines 1098–1184 · score 0.77 · Reactome pathway, Brain microglia links, target genes, DegCre, Closest, iPSCs
- [7] § Results › KOLF2.1J iTF-microglia transcriptome resembles other in vitro microglia models and brain microglia ↔ Scripts/RNAseq_analysis.R, lines 640–679 · score 0.76 · iHPCs, iMGLs, fetal microglia, CD16, WTC11, CD14
- [8] § Results › Inflammatory activation of iTF-microglia by LPS + IFNγ ↔ Scripts/cCRE_analysis.R, lines 419–462 · score 0.66 · innate immune, neutrophil degranulation, interleukin, phagocytosis, iPSCs, Pathway
- [9] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of candidate cis-regulatory elements (cCREs) ↔ Scripts/LDSC_bed_files_20250429.R, lines 84–162 · score 0.64 · GenomicRanges, findOverlaps, H3K27ac, cCREs, consensus, filtered
- [10] § Results › Cis-regulatory landscapes of iTF-microglia across differentiation and activation ↔ Scripts/TF_monaLisa_analysis.R, lines 244–284 · score 0.64 · dELS, pELS, ChIP, cCREs, PLS, promoters
- [11] § Results › Inflammatory activation of iTF-microglia by LPS + IFNγ ↔ Scripts/RNAseq_analysis.R, lines 170–233 · score 0.63 · HLA DRB1, CD40, CXCL9, GBP2, IDO1, CXCL10
- [12] § Results › Inflammatory activation of iTF-microglia by LPS + IFNγ ↔ Scripts/TF_monaLisa_analysis.R, lines 1061–1087 · score 0.63 · HLA DRB1, CD40, CXCL9, GBP2, IDO1, CXCL10
- [13] § STAR★Methods › Quantification and statistical analysis › Association of TRNs with candidate disease-risk variants ↔ Scripts/monalisa_b.R, lines 40–82 · score 0.62 · AD GWAS, risk variants, LD, H3K27ac, monaLisa, ATAC
- [14] § Results › KOLF2.1J iTF-microglia transcriptome resembles other in vitro microglia models and brain microglia ↔ Scripts/RNAseq_analysis.R, lines 640–679 · score 0.61 · iHPCs, fetal microglia, DESeq2, adult, KOLF2.1J, Monocytes
- [15] § STAR★Methods › Quantification and statistical analysis › Association of TRNs with candidate disease-risk variants ↔ Scripts/motifbreakR_script.R, lines 41–112 · score 0.61 · motifbreakR, ic, JASPAR, PWMs, candidate, motifs
- [16] § Results › Inference of cCRE-gene associations across microglial states ↔ Scripts/cCRE-TargetGene_analysis.R, lines 1200–1263 · score 0.60 · vivo microglial, Brain microglia links, DegCre, closest, distance, enhancer
- [17] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of candidate cis-regulatory elements (cCREs) ↔ Scripts/cCRE_disease_variant_annotation.R, lines 581–625 · score 0.60 · GenomicRanges, findOverlaps, cCRE, filtered, overlapping, cell
- [18] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of TFs in TRNs ↔ Scripts/monalisa_b.R, lines 211–253 · score 0.59 · findMotifHits, scanning, PWMs, monaLisa, JASPAR2024, csaw
- [19] § Results › Inference of cCRE-gene associations across microglial states ↔ Scripts/cCRE-TargetGene_analysis.R, lines 1524–1594 · score 0.57 · brain microglia links, cCRE, DegCre, promoter, TSS, enhancers
- [20] § STAR★Methods › Quantification and statistical analysis › ATAC- and ChIP-seq processing and peak calling ↔ src/encode_task_overlap.py, lines 132–174 · score 0.56 · naive overlap, ATAC seq, shift, blacklist, ChIP, filtered
- [21] § STAR★Methods › Quantification and statistical analysis › Nomination of transcriptional regulatory networks (TRNs) › Identification of TFs in TRNs ↔ Scripts/TF_monaLisa_analysis.R, lines 453–495 · score 0.56 · findMotifHits, scanning, PWMs, JASPAR2024, csaw, matrices
- [22] § Results › Inference of cCRE-gene associations across microglial states ↔ Scripts/cCRE-TargetGene_analysis.R, lines 1524–1594 · score 0.56 · cCRE, target genes, DegCre, brain microglia, ratio, TSS
- [23] § STAR★Methods › Quantification and statistical analysis › RNA-seq meta-analysis ↔ Scripts/motifbreakR_script.R, lines 464–533 · score 0.55 · Variance stabilizing transformation, VST, iPSCs, clustering, Gene, microglia
- [24] § STAR★Methods › Quantification and statistical analysis › Association of TRNs with candidate disease-risk variants ↔ Scripts/MAGMA_expression_matrix_format.R, the whole file · a weak match · score 0.54 · GWAS summary, MAGMA, winsorized, log2, TPM, subsets
- [25] § Results › iTF-microglia TRNs are enriched for neurodegenerative and autoimmune risk variants ↔ Scripts/LDSC_heatmap.R, lines 136–195 · score 0.54 · REL, H3K27ac, IBD, SCZ, SLE, cCREs
- [26] § STAR★Methods › Quantification and statistical analysis › RNA-seq meta-analysis ↔ Scripts/RNAseq_analysis.R, lines 681–720 · score 0.52 · Variance stabilizing transformation, VST, clustering
- [27] § STAR★Methods › Quantification and statistical analysis › ATAC- and ChIP-seq processing and peak calling ↔ src/encode_task_qc_report.py, lines 458–526 · score 0.50 · seq pipeline, mitochondrial, Picard, quality, QC, ENCODE
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R · 1,594 lines · 62 KB · MIT · 5 matches
- # Script to compare target genes in iTF cells
- # open R
- srun -n 10 -p interactive --pty /bin/bash
- module load cluster/gsl/1.15
- module load cluster/htslib/1.9-devel
- module load cluster/build_tools/llvm/8.0
- module load cluster/util/gcc/8.2
- module load hdf5_18/1.8.21
- conda activate r_env
- #conda deactivate
- R
- library(GenomicRanges)
- library(AnnotationHub)
- library(data.table)
- library(dplyr)
- library(tidyr)
- library(rtracklayer)
- library(tibble)
- library(stringr)
- library(DESeq2)
- library(org.Hs.eg.db)
- library(ggplot2)
- # Set work directory
- setwd("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/")
- # Load data
- iTF_CREs_ENCODE <- readRDS("batch2/Tables/cCREs_overlap_iTF_ATAC_H3K27ac_overlap.rds")
- consensus_CREs <- readRDS("batch2/Tables/cCREs_overlap_consensus_peaks.rds")
- gse <- readRDS("batch2/Tables/gse.rds")
- load("/cluster/home/broberts/ref_genomes/STAR_Gencode_v46/gencodeV46TSSGR.rda")
- Gencode.V46 <- finalTSsGR
- # Load and filter gene expression data
- coldata <- read.csv("/cluster/home/irodriguez/Myers_Lab/20240416_Novogene_RNAseq_Library_iTF_iMGLs_Nick_paper/R_analysis/iTF-Microglia_BulkRNA_SampleSheet_phenotype.csv")
- tpm <- assay(gse, "abundance")
- coldata_subset <- coldata[c(13:15, 19:24),]
- tpm_subset <- tpm[, c(13:15, 19:24)]
- colnames(tpm_subset) <- coldata_subset$Sample
- rownames(tpm_subset) <- sapply(strsplit(rownames(tpm_subset), "[.]"), `[[`, 1)
- # Compute row-wise means
- expr_means <- data.frame(
- iPSC = rowMeans(tpm_subset[, 1:3]),
- Microglia = rowMeans(tpm_subset[, 4:6]),
- Microglia_LPS_IFNG = rowMeans(tpm_subset[, 7:9])
- )
- # Filter genes based on expression
- filtered_genes <- rownames(expr_means)[rowSums(expr_means > 1) > 0]
- # Filter and process GENCODE annotations
- gtf_df_filtered <- as.data.frame(Gencode.V46) %>%
- filter(GeneID %in% filtered_genes) %>%
- mutate(
- promoter_start = ifelse(strand == "+", pmax(start - 2000, 1), end - 200),
- promoter_end = ifelse(strand == "+", start + 200, pmax(end + 2000, 1))
- )
- Gencode.V46_tss <- GRanges(gtf_df_filtered)
- # Promoters
- promoter_df <- unique(gtf_df_filtered[,c(1,10:11,7)])
- Gencode.V46_promoter <- GRanges(promoter_df)
- # Find closest TSS and distances
- closest_tss_indices <- nearest(consensus_CREs, Gencode.V46_tss)
- dists <- distanceToNearest(consensus_CREs, Gencode.V46_tss)
- consensus_CREs$closest_tss <- mcols(Gencode.V46_tss[closest_tss_indices])$GeneSymb
- consensus_CREs$distance_tss <- mcols(dists)$distance
- # Overlap with cell-type specific peaks
- overlap_gr_list <- GRangesList(
- lapply(names(iTF_CREs_ENCODE), function(name) {
- gr_current <- iTF_CREs_ENCODE[[name]]
- overlaps <- findOverlaps(consensus_CREs, gr_current)
- overlapping_ranges <- consensus_CREs[queryHits(overlaps)]
- mcols(overlapping_ranges)$cell.type <- mcols(gr_current)$cell.type[subjectHits(overlaps)]
- unique(overlapping_ranges)
- })
- )
- # Convert results to data frames
- iTF_cCREs_df <- as.data.frame(consensus_CREs)
- overlap_dfs <- lapply(overlap_gr_list, as.data.frame)
- names(overlap_dfs) <- names(iTF_CREs_ENCODE)
- iTF_iPSCs_cCREs_df <- overlap_dfs$iTF_iPSCs
- iTF_Microglia_cCREs_df <- overlap_dfs$iTF_Microglia
- iTF_Microglia.LPS.IFNG_cCREs_df <- overlap_dfs$iTF_Microglia.LPS.IFNG
- # Merge keeping all rows in df1
- merged_iTF_cCREs_df <- merge(iTF_cCREs_df, iTF_iPSCs_cCREs_df, by = intersect(names(iTF_cCREs_df), names(iTF_iPSCs_cCREs_df)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia_cCREs_df, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia_cCREs_df)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia.LPS.IFNG_cCREs_df, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia.LPS.IFNG_cCREs_df)), all.x = TRUE)
- # Rename columns
- merged_iTF_cCREs_df <- merged_iTF_cCREs_df %>% rename(
- iPSC_cCRE = cell.type.x,
- Microglia_cCRE = cell.type.y,
- Microglia_LPS_IFNG_cCRE = cell.type
- ) %>% unique()
- ########################
- # HiC loops Processing
- ########################
- library(mariner)
- library(InteractionSet)
- # Define file path
- dir_loop_data <- "/cluster/projects/ADFTD/HiC/Ivan_10-21-24/encode_pipeline_outputs"
- loopFiles <- list.files(path = dir_loop_data, pattern = "loops_30.bedpe.gz", recursive = TRUE, full.names = TRUE)
- # Read and sort loops file
- loops_sorted <- fread(loopFiles[7], header = FALSE) %>%
- as.data.frame() %>%
- arrange(V1, V2) %>%
- mutate(name = paste0("loop", row_number()))
- # Convert to GInteractions and clear metadata
- loops_GI <- as_ginteractions(loops_sorted)
- mcols(loops_GI) <- NULL
- # Find overlaps
- overlap_anchor1 <- findOverlaps(anchors(loops_GI, type = "first"), Gencode.V46_promoter)
- overlap_anchor2 <- findOverlaps(anchors(loops_GI, type = "second"), Gencode.V46_promoter)
- # Create data.frames of overlaps
- df_olap1 <- data.frame(
- anchor_id = queryHits(overlap_anchor1),
- gene = Gencode.V46_promoter$GeneSymb[subjectHits(overlap_anchor1)]
- )
- df_olap2 <- data.frame(
- anchor_id = queryHits(overlap_anchor2),
- gene = Gencode.V46_promoter$GeneSymb[subjectHits(overlap_anchor2)]
- )
- # Collapse overlapping gene names
- collapsed_genes1 <- aggregate(gene ~ anchor_id, df_olap1, function(x) paste(unique(x), collapse = ","))
- collapsed_genes2 <- aggregate(gene ~ anchor_id, df_olap2, function(x) paste(unique(x), collapse = ","))
- # Initialize columns
- mcols(loops_GI)$hic_anchor1_gene <- NA_character_
- mcols(loops_GI)$hic_anchor2_gene <- NA_character_
- # Assign collapsed gene names
- mcols(loops_GI)$hic_anchor1_gene[collapsed_genes1$anchor_id] <-
- collapsed_genes1$gene[match(collapsed_genes1$anchor_id, collapsed_genes1$anchor_id)]
- mcols(loops_GI)$hic_anchor2_gene[collapsed_genes2$anchor_id] <-
- collapsed_genes2$gene[match(collapsed_genes2$anchor_id, collapsed_genes2$anchor_id)]
- # Overlap with cCREs
- consensus_CREs_ranges <- consensus_CREs
- mcols(consensus_CREs_ranges) <- NULL
- # Overlap cCREs with anchor1
- overlap_cCRE_anchor1 <- findOverlaps(consensus_CREs_ranges,anchors(loops_GI, type = "first"))
- # Overlap cCREs with anchor2
- overlap_cCRE_anchor2 <- findOverlaps(consensus_CREs_ranges,anchors(loops_GI, type = "second"))
- # Add gene_name to anchor1 overlaps
- consensus_CREs_anchor1 <- consensus_CREs_ranges
- mcols(consensus_CREs_anchor1)$hic_anchor_gene <- NA
- mcols(consensus_CREs_anchor1)$hic_anchor_gene[queryHits(overlap_cCRE_anchor1)] <- loops_GI$hic_anchor2_gene[subjectHits(overlap_cCRE_anchor1)]
- # Add gene_name to anchor2 overlaps
- consensus_CREs_anchor2 <- consensus_CREs_ranges
- mcols(consensus_CREs_anchor2)$hic_anchor_gene <- NA
- mcols(consensus_CREs_anchor2)$hic_anchor_gene[queryHits(overlap_cCRE_anchor2)] <- loops_GI$hic_anchor1_gene[subjectHits(overlap_cCRE_anchor2)]
- # Create datafr
- consensus_CREs_anchor1_df <- as.data.frame(consensus_CREs_anchor1)
- consensus_CREs_anchor2_df <- as.data.frame(consensus_CREs_anchor2)
- # Assuming df1 and df2 are the dataframes you want to rbind
- df_combined <- bind_rows(consensus_CREs_anchor1_df, consensus_CREs_anchor1_df)
- # Remove NA from 'hic_anchor_gene' and collapse by ',' (comma separator)
- cCRE_hic_genes <- df_combined %>%
- filter(!is.na(hic_anchor_gene)) %>%
- group_by(across(-hic_anchor_gene)) %>% # Group by all columns except 'hic_anchor_gene'
- summarise(hic_anchor_gene = paste(unique(na.omit(hic_anchor_gene)), collapse = ","), .groups = 'drop')
- # Result
- cCRE_hic_genes <- as.data.frame(cCRE_hic_genes)
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, cCRE_hic_genes, by = intersect(names(iTF_cCREs_df), names(cCRE_hic_genes)), all.x = TRUE)
- ########################
- #DegCre results
- ########################
- # Define the directory
- folder_path <- "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables"
- # List all files with the "_degCreGInter.rd" suffix
- #rd_files <- list.files(folder_path, pattern = "_H3K27ac_overlapping_peaks_degCreGInter.rd$", full.names = TRUE)
- #rd_files <- list.files(folder_path, pattern = "_ATAC_overlapping_peaks_degCreGInter.rd$", full.names = TRUE)
- rd_files <- list.files(folder_path, pattern = "_ATAC_overlapping_peaks_diff_degCreGInter.rd$", full.names = TRUE)
- rd_files <- list.files(folder_path, pattern = "_H3K27ac_overlapping_peaks_stim_degCreGInter.rd$", full.names = TRUE)
- # Loop through the list of files and load each one
- for (file in rd_files) {
- # Extract the file name without extension to use as a variable name
- file_name <- gsub("_H3K27ac_overlapping_peaks_stim_degCreGInter.rd$", "", basename(file))
- # Load the file
- temp_object_name <- load(file)
- # Retrieve the object using `get`
- assign(file_name, get(temp_object_name))
- # Optionally remove the temporary object
- rm(list = temp_object_name)
- }
- # After the loop, the variables corresponding to the files will be available in the environment
- # Extract the second anchor (the second column of the anchor matrix)
- iTF_Microglia_vs_iPSCs_CRE <- anchors(csaw_iTF_Microglia_vs_iTF_iPSCs, type = "second")
- iTF_Microglia_vs_iPSCs_CRE_df <- as.data.frame(iTF_Microglia_vs_iPSCs_CRE)
- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE <- anchors(csaw_iTF_Microglia.LPS.IFNG_vs_iTF_Microglia, type = "second")
- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE_df <- as.data.frame(iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE)
- # Extract the metadata columns: Deg_GeneSymb and Cre_direction
- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE_meta_data <- mcols(csaw_iTF_Microglia.LPS.IFNG_vs_iTF_Microglia)[, c("Deg_GeneSymb", "Cre_direction", "assocDist")]
- iTF_Microglia_vs_iPSCs_CRE_meta_data <- mcols(csaw_iTF_Microglia_vs_iTF_iPSCs)[, c("Deg_GeneSymb", "Cre_direction", "assocDist")]
- # Combine the second anchor and the selected metadata columns into a data frame
- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_DegCre <- unique(cbind(iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE_df, iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_CRE_meta_data))
- iTF_Microglia_vs_iPSCs_DegCre <- unique(cbind(iTF_Microglia_vs_iPSCs_CRE_df, iTF_Microglia_vs_iPSCs_CRE_meta_data))
- # Separate by cell type
- iTF_Microglia.LPS.IFNG_DegCre <- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_DegCre[iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_DegCre$Cre_direction == "up",]
- iTF_Microglia.untreated_DegCre <- iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_DegCre[iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_DegCre$Cre_direction == "down",]
- iTF_Microglia.diff_DegCre <- iTF_Microglia_vs_iPSCs_DegCre[iTF_Microglia_vs_iPSCs_DegCre$Cre_direction == "up",]
- iTF_iPSCs_DegCre <- iTF_Microglia_vs_iPSCs_DegCre[iTF_Microglia_vs_iPSCs_DegCre$Cre_direction == "down",]
- # Bind cell types
- iTF_Microglia.LPS.IFNG_DegCre <- unique(iTF_Microglia.LPS.IFNG_DegCre[,c(1:6,8)])
- iTF_Microglia.untreated_DegCre <- unique(iTF_Microglia.untreated_DegCre[,c(1:6,8)])
- iTF_Microglia.diff_DegCre <- unique(iTF_Microglia.diff_DegCre[,c(1:6,8)])
- iTF_iPSCs_DegCre <- unique(iTF_iPSCs_DegCre[,c(1:6,8)])
- # Assuming df is your dataframe
- iTF_Microglia.LPS.IFNG_DegCre <- iTF_Microglia.LPS.IFNG_DegCre %>%
- group_by(seqnames, start, end) %>% # Group by seqnames and start
- summarise(
- DegCre_iTF_Microglia.LPS.IFNG = paste(sort(unique(na.omit(Deg_GeneSymb))), collapse = ","), # Sort and collapse unique Deg_GeneSymb
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(seqnames, start, end) %>%
- as.data.frame()
- iTF_Microglia.untreated_DegCre <- iTF_Microglia.untreated_DegCre %>%
- group_by(seqnames, start, end) %>% # Group by seqnames and start
- summarise(
- DegCre_iTF_Microglia.untreated = paste(sort(unique(na.omit(Deg_GeneSymb))), collapse = ","), # Collapse unique Deg_GeneSymb
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(seqnames, start, end) %>% as.data.frame()
- iTF_Microglia.diff_DegCre <- iTF_Microglia.diff_DegCre %>%
- group_by(seqnames, start, end) %>% # Group by seqnames and start
- summarise(
- DegCre_iTF_Microglia.diff = paste(sort(unique(na.omit(Deg_GeneSymb))), collapse = ","), # Collapse unique Deg_GeneSymb
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(seqnames, start, end) %>% as.data.frame()
- iTF_iPSCs_DegCre <- iTF_iPSCs_DegCre %>%
- group_by(seqnames, start, end) %>% # Group by seqnames and start
- summarise(
- DegCre_iTF_iPSCs = paste(sort(unique(na.omit(Deg_GeneSymb))), collapse = ","), # Collapse unique Deg_GeneSymb
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(seqnames, start, end) %>% as.data.frame()
- # Make Granges
- iTF_iPSCs_DegCre_gr <- GRanges(iTF_iPSCs_DegCre)
- iTF_Microglia.diff_DegCre_gr <- GRanges(iTF_Microglia.diff_DegCre)
- iTF_Microglia.untreated_DegCre_gr <- GRanges(iTF_Microglia.untreated_DegCre)
- iTF_Microglia.LPS.IFNG_DegCre_gr <- GRanges(iTF_Microglia.LPS.IFNG_DegCre)
- # Make GrangesList
- DegCre_iTF_list <- list(iTF_iPSCs_DegCre_gr,iTF_Microglia.diff_DegCre_gr,
- iTF_Microglia.untreated_DegCre_gr,iTF_Microglia.LPS.IFNG_DegCre_gr)
- #DegCre_iTF_list_grl <- GRangesList(DegCre_iTF_list)
- # overlap
- overlap_with_metadata <- function(query, subject) {
- # Find overlaps between query and subject
- hits <- findOverlaps(query, subject)
- # Extract overlapping ranges
- query_overlaps <- query[queryHits(hits)]
- subject_overlaps <- subject[subjectHits(hits)]
- # Combine metadata columns
- combined_mcols <- cbind(mcols(query_overlaps), mcols(subject_overlaps))
- # Create a new GRanges object with combined metadata
- result <- query_overlaps
- mcols(result) <- combined_mcols
- return(result)
- }
- # Assuming ABC_Mac_list is a named list of GRanges objects
- overlapped_list <- lapply(DegCre_iTF_list, function(gr) overlap_with_metadata(consensus_CREs_ranges, gr))
- iTF_iPSCs_DegCre <- as.data.frame(overlapped_list[[1]])
- iTF_Microglia.diff_DegCre <- as.data.frame(overlapped_list[[2]])
- iTF_Microglia.untreated_DegCre <- as.data.frame(overlapped_list[[3]])
- iTF_Microglia.LPS.IFNG_DegCre <- as.data.frame(overlapped_list[[4]])
- iTF_iPSCs_DegCre <- unique(iTF_iPSCs_DegCre[,c(1:3,6)])
- iTF_Microglia.diff_DegCre <- unique(iTF_Microglia.diff_DegCre[,c(1:3,6)])
- iTF_Microglia.untreated_DegCre <- unique(iTF_Microglia.untreated_DegCre[,c(1:3,6)])
- iTF_Microglia.LPS.IFNG_DegCre <- unique(iTF_Microglia.LPS.IFNG_DegCre[,c(1:3,6)])
- # Collapse CREs
- # Group by seqnames, start, and end, then collapse
- iTF_iPSCs_DegCre_collapsed <- iTF_iPSCs_DegCre %>%
- group_by(seqnames, start, end) %>%
- summarise(DegCre_iTF_iPSCs = paste(
- unique(sort(unlist(strsplit(DegCre_iTF_iPSCs, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- iTF_Microglia_DegCre.diff_collapsed <- iTF_Microglia.diff_DegCre %>%
- group_by(seqnames, start, end) %>%
- summarise(DegCre_iTF_Microglia.diff = paste(
- unique(sort(unlist(strsplit(DegCre_iTF_Microglia.diff, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- iTF_Microglia_DegCre.untreated_collapsed <- iTF_Microglia.untreated_DegCre %>%
- group_by(seqnames, start, end) %>%
- summarise(DegCre_iTF_Microglia.untreated = paste(
- unique(sort(unlist(strsplit(DegCre_iTF_Microglia.untreated, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- iTF_Microglia_DegCre.LPS.IFNG_collapsed <- iTF_Microglia.LPS.IFNG_DegCre %>%
- group_by(seqnames, start, end) %>%
- summarise(DegCre_iTF_Microglia.LPS.IFNG = paste(
- unique(sort(unlist(strsplit(DegCre_iTF_Microglia.LPS.IFNG, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_iPSCs_DegCre_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_iPSCs_DegCre_collapsed)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia_DegCre.diff_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia_DegCre.diff_collapsed)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia_DegCre.untreated_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia_DegCre.untreated_collapsed)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia_DegCre.LPS.IFNG_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia_DegCre.LPS.IFNG_collapsed)), all.x = TRUE)
- ########################
- #ABC results
- ########################
- # Define the list of file paths
- file_paths <- c(
- "/cluster/home/irodriguez/software/ABC-Enhancer-Gene-Prediction/results_KOLF2.1J-iTF_HiC_20250423/iTF_iPSCs_hic_all/Predictions/EnhancerPredictionsAllPutative.tsv.gz",
- "/cluster/home/irodriguez/software/ABC-Enhancer-Gene-Prediction/results_KOLF2.1J-iTF_HiC_20250423/iTF_Microglia_hic_all/Predictions/EnhancerPredictionsAllPutative.tsv.gz",
- "/cluster/home/irodriguez/software/ABC-Enhancer-Gene-Prediction/results_KOLF2.1J-iTF_HiC_20250423/iTF_Microglia.LPS.IFNG_hic_all/Predictions/EnhancerPredictionsAllPutative.tsv.gz"
- )
- # Apply the function to each file path
- cell_types <- sapply(strsplit(unlist(file_paths),"[/]"),"[[",8)
- # Read all files into a list
- ABC_list <- lapply(file_paths, function(fp) fread(fp, sep = "\t", header = TRUE))
- names(ABC_list) <- cell_types
- # ABC score threshold based on doi:10.1038/s41586-021-03446-x
- ABC_filtered <- lapply(ABC_list, function(df) {
- df[df$class == "promoter" & df$ABC.Score >= 0.1 |
- df$class != "promoter" & df$ABC.Score >= 0.015, ]
- })
- ABC_filtered_subset <- lapply(ABC_filtered, function(x) {
- unique(x[,c(1:3,11,31)])
- })
- iTF_iPSCs_ABC <- as.data.frame(ABC_filtered_subset$iTF_iPSCs_hic_all)
- iTF_Microglia_ABC <- as.data.frame(ABC_filtered_subset$iTF_Microglia_hic_all)
- iTF_Microglia.LPS.IFNG_ABC <- as.data.frame(ABC_filtered_subset$iTF_Microglia.LPS.IFNG_hic_all)
- # Assuming df is your dataframe
- iTF_Microglia.LPS.IFNG_ABC <- iTF_Microglia.LPS.IFNG_ABC %>%
- group_by(chr, start, end) %>% # Group by chr and start
- summarise(
- ABC_iTF_Microglia.LPS.IFNG = paste(sort(unique(na.omit(TargetGene))), collapse = ","), # Sort and collapse unique TargetGene
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(chr, start, end) %>%
- as.data.frame()
- iTF_Microglia_ABC <- iTF_Microglia_ABC %>%
- group_by(chr, start, end) %>% # Group by chr and start
- summarise(
- ABC_iTF_Microglia = paste(sort(unique(na.omit(TargetGene))), collapse = ","), # Collapse unique TargetGene
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(chr, start, end) %>% as.data.frame()
- iTF_iPSCs_ABC <- iTF_iPSCs_ABC %>%
- group_by(chr, start, end) %>% # Group by chr and start
- summarise(
- ABC_iTF_iPSCs = paste(sort(unique(na.omit(TargetGene))), collapse = ","), # Collapse unique TargetGene
- .groups = 'drop' # Drop grouping structure in the result
- ) %>%
- arrange(chr, start, end) %>% as.data.frame()
- # Make Granges
- iTF_iPSCs_ABC_gr <- GRanges(iTF_iPSCs_ABC)
- iTF_Microglia_ABC_gr <- GRanges(iTF_Microglia_ABC)
- iTF_Microglia.LPS.IFNG_ABC_gr <- GRanges(iTF_Microglia.LPS.IFNG_ABC)
- # Make GrangesList
- ABC_iTF_list <- list(iTF_iPSCs_ABC_gr,iTF_Microglia_ABC_gr,iTF_Microglia.LPS.IFNG_ABC_gr)
- # overlap
- overlap_with_metadata <- function(query, subject) {
- # Find overlaps between query and subject
- hits <- findOverlaps(query, subject)
- # Extract overlapping ranges
- query_overlaps <- query[queryHits(hits)]
- subject_overlaps <- subject[subjectHits(hits)]
- # Combine metadata columns
- combined_mcols <- cbind(mcols(query_overlaps), mcols(subject_overlaps))
- # Create a new GRanges object with combined metadata
- result <- query_overlaps
- mcols(result) <- combined_mcols
- return(result)
- }
- # Assuming ABC_iTF_list is a named list of GRanges objects
- overlapped_list <- lapply(ABC_iTF_list, function(gr) overlap_with_metadata(consensus_CREs_ranges, gr))
- iTF_iPSCs_ABC <- as.data.frame(overlapped_list[[1]])
- iTF_Microglia_ABC <- as.data.frame(overlapped_list[[2]])
- iTF_Microglia.LPS.IFNG_ABC <- as.data.frame(overlapped_list[[3]])
- iTF_iPSCs_ABC <- unique(iTF_iPSCs_ABC[,c(1:3,6)])
- iTF_Microglia_ABC <- unique(iTF_Microglia_ABC[,c(1:3,6)])
- iTF_Microglia.LPS.IFNG_ABC <- unique(iTF_Microglia.LPS.IFNG_ABC[,c(1:3,6)])
- # Collapse CREs
- # Group by seqnames, start, and end, then collapse ABC_THP1_Mac_0h
- iTF_iPSCs_ABC_collapsed <- iTF_iPSCs_ABC %>%
- group_by(seqnames, start, end) %>%
- summarise(ABC_iTF_iPSCs = paste(
- unique(sort(unlist(strsplit(ABC_iTF_iPSCs, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- iTF_Microglia_ABC_collapsed <- iTF_Microglia_ABC %>%
- group_by(seqnames, start, end) %>%
- summarise(ABC_iTF_Microglia = paste(
- unique(sort(unlist(strsplit(ABC_iTF_Microglia, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- iTF_Microglia.LPS.IFNG_ABC_collapsed <- iTF_Microglia.LPS.IFNG_ABC %>%
- group_by(seqnames, start, end) %>%
- summarise(ABC_iTF_Microglia.LPS.IFNG = paste(
- unique(sort(unlist(strsplit(ABC_iTF_Microglia.LPS.IFNG, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_iPSCs_ABC_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_iPSCs_ABC_collapsed)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia_ABC_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia_ABC_collapsed)), all.x = TRUE)
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_Microglia.LPS.IFNG_ABC_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_Microglia.LPS.IFNG_ABC_collapsed)), all.x = TRUE)
- #Save files
- saveRDS(merged_iTF_cCREs_df,file="/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- #merged_iTF_cCREs_df <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/Tables/iTF_CREs_target_genes.rds")
- #consensus_CREs <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/candidate_CREs/iTF_CREs_all_ENCODE_ABC_peaks.rds")
- #iTF_cCREs_df <- as.data.frame(consensus_CREs)
- #######################
- #snMULTIomics AD Myers lab
- #######################
- # Load links
- links <- read.csv("/cluster/home/irodriguez/Myers_Lab/20230327_AD_variants_TF_trios_snMultiomics/mmc6_feature_links.csv")
- # Subset trios
- links_subset <- unique(links[,c(1:3,5,9,8)])
- links_subset_MG <- links_subset[grep("Microglia",links_subset$CT),]
- links_subset_MG_genes <- unique(links_subset_MG$gene)
- # Summarize by unique seqnames, start, and end
- links_subset_MG_summary <- links_subset_MG %>%
- group_by(seqnames, start, end) %>%
- summarise(
- gene = paste(gene, collapse = ","),
- group = paste(group, collapse = ","),
- .groups = "drop"
- )
- links_subset_MG_summary <- unique(links_subset_MG_summary)
- #colnames(links_subset_MG_summary) <- c("seqnames","start","end","Peak2gene.Anderson2023")
- colnames(links_subset_MG_summary) <- c("seqnames","start","end","Peak2gene.Anderson2023","Phenotype.Anderson2023")
- # Create GRanges objects
- links_GR <- GRanges(links_subset_MG_summary)
- #overlap with SNPs
- olap_links <- findOverlaps(consensus_CREs_ranges,links_GR)
- # add the metacolumns to the GR object
- consensus_CREs_links <- consensus_CREs_ranges[queryHits(olap_links)]
- mcols(consensus_CREs_links) <- cbind(mcols(consensus_CREs_links), mcols(links_GR[subjectHits(olap_links)]))
- # Create df from the GR overlap
- iTF_cCREs_links <- data.frame(consensus_CREs_links)
- iTF_cCREs_links <- unique(iTF_cCREs_links[,c(1:3,6:7)])
- # Collapse CREs
- # Group by seqnames, start, and end
- iTF_cCREs_links_collapsed <- iTF_cCREs_links %>%
- group_by(seqnames, start, end) %>%
- summarise(Peak2gene.Anderson2023 = paste(
- unlist(strsplit(Peak2gene.Anderson2023, ",")), collapse = ","),
- Phenotype.Anderson2023 = paste(
- unlist(strsplit(Phenotype.Anderson2023, ",")), collapse = ","),
- .groups = "drop") %>% as.data.frame()
- iTF_cCREs_links_collapsed <- unique(iTF_cCREs_links_collapsed)
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_cCREs_links_collapsed, by = intersect(names(iTF_cCREs_df), names(iTF_cCREs_links_collapsed)), all.x = TRUE)
- #Save files
- saveRDS(merged_iTF_cCREs_df,file="/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- #merged_iTF_cCREs_df <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- #########################################
- # Nott PLACseq brain
- #########################################
- library(readxl)
- read_excel_allsheets <- function(filename, tibble = FALSE) {
- # I prefer straight data.frames
- # but if you like tidyverse tibbles (the default with read_excel)
- # then just pass tibble = TRUE
- sheets <- readxl::excel_sheets(filename)
- x <- lapply(sheets, function(X) readxl::read_excel(filename, sheet = X))
- if(!tibble) x <- lapply(x, as.data.frame)
- names(x) <- sheets
- x
- }
- # Load PLACseq files
- Nott_PLAC <- read_excel_allsheets("/cluster/home/irodriguez/Myers_Lab/GWAS_AD_summary_statistics/variant_annotation/Nott_files/aay0793-nott-table-s5.xlsx")
- names(Nott_PLAC)
- PLAC_mic <- as.data.frame(Nott_PLAC[[2]])
- colnames(PLAC_mic) <- PLAC_mic[2,]
- PLAC_mic <- PLAC_mic[-c(1:2),1:6]
- PLACseq <- unique(PLAC_mic)
- PLACseq$start1 <- as.numeric(PLACseq$start1)
- PLACseq$start2 <- as.numeric(PLACseq$start2)
- # Convert to GInteractions and clear metadata
- loops_GI <- as_ginteractions(PLACseq)
- mcols(loops_GI) <- NULL
- # Liftover loops
- # transformation to hg38 coordinates
- ch = import.chain("/cluster/home/irodriguez/software/hg19ToHg38.over.chain")
- str(ch[[1]])
- # Perform liftOver without flattening
- lifted1_list <- liftOver(anchors(loops_GI, type = "first"), ch)
- lifted2_list <- liftOver(anchors(loops_GI, type = "second"), ch)
- # Get indices where exactly one mapping was returned
- valid_idx <- elementNROWS(lifted1_list) == 1 & elementNROWS(lifted2_list) == 1
- # Subset to valid interactions
- loops_valid <- loops_GI[valid_idx]
- # Extract the lifted anchors
- lifted1 <- unlist(lifted1_list[valid_idx])
- lifted2 <- unlist(lifted2_list[valid_idx])
- lifted_loops_GI <- GInteractions(anchor1 = lifted1,
- anchor2 = lifted2,
- metadata = mcols(loops_valid),
- mode = "strict") # optional
- # Find overlaps
- overlap_anchor1 <- findOverlaps(anchors(lifted_loops_GI, type = "first"), Gencode.V46_promoter)
- overlap_anchor2 <- findOverlaps(anchors(lifted_loops_GI, type = "second"), Gencode.V46_promoter)
- # Create data.frames of overlaps
- df_olap1 <- data.frame(
- anchor_id = queryHits(overlap_anchor1),
- gene = Gencode.V46_promoter$GeneSymb[subjectHits(overlap_anchor1)]
- )
- df_olap2 <- data.frame(
- anchor_id = queryHits(overlap_anchor2),
- gene = Gencode.V46_promoter$GeneSymb[subjectHits(overlap_anchor2)]
- )
- # Collapse overlapping gene names
- collapsed_genes1 <- aggregate(gene ~ anchor_id, df_olap1, function(x) paste(unique(x), collapse = ","))
- collapsed_genes2 <- aggregate(gene ~ anchor_id, df_olap2, function(x) paste(unique(x), collapse = ","))
- # Initialize columns
- mcols(lifted_loops_GI)$PLAC_anchor1_gene <- NA_character_
- mcols(lifted_loops_GI)$PLAC_anchor2_gene <- NA_character_
- # Assign collapsed gene names
- mcols(lifted_loops_GI)$PLAC_anchor1_gene[collapsed_genes1$anchor_id] <-
- collapsed_genes1$gene[match(collapsed_genes1$anchor_id, collapsed_genes1$anchor_id)]
- mcols(lifted_loops_GI)$PLAC_anchor2_gene[collapsed_genes2$anchor_id] <-
- collapsed_genes2$gene[match(collapsed_genes2$anchor_id, collapsed_genes2$anchor_id)]
- # Overlap cCREs with anchor1
- overlap_cCRE_anchor1 <- findOverlaps(consensus_CREs_ranges,anchors(lifted_loops_GI, type = "first"))
- # Overlap cCREs with anchor2
- overlap_cCRE_anchor2 <- findOverlaps(consensus_CREs_ranges,anchors(lifted_loops_GI, type = "second"))
- # Add gene_name to anchor1 overlaps
- consensus_CREs_anchor1 <- consensus_CREs_ranges
- mcols(consensus_CREs_anchor1)$Nott_PLAC_anchor_gene <- NA
- mcols(consensus_CREs_anchor1)$Nott_PLAC_anchor_gene[queryHits(overlap_cCRE_anchor1)] <- lifted_loops_GI$PLAC_anchor2_gene[subjectHits(overlap_cCRE_anchor1)]
- # Add gene_name to anchor2 overlaps
- consensus_CREs_anchor2 <- consensus_CREs_ranges
- mcols(consensus_CREs_anchor2)$Nott_PLAC_anchor_gene <- NA
- mcols(consensus_CREs_anchor2)$Nott_PLAC_anchor_gene[queryHits(overlap_cCRE_anchor2)] <- lifted_loops_GI$PLAC_anchor1_gene[subjectHits(overlap_cCRE_anchor2)]
- # Create datafr
- consensus_CREs_anchor1_df <- as.data.frame(consensus_CREs_anchor1)
- consensus_CREs_anchor2_df <- as.data.frame(consensus_CREs_anchor2)
- # Assuming df1 and df2 are the dataframes you want to rbind
- df_combined <- bind_rows(consensus_CREs_anchor1_df, consensus_CREs_anchor1_df)
- # Remove NA from 'hic_anchor_gene' and collapse by ',' (comma separator)
- cCRE_hic_genes <- df_combined %>%
- filter(!is.na(Nott_PLAC_anchor_gene)) %>%
- group_by(across(-Nott_PLAC_anchor_gene)) %>% # Group by all columns except 'hic_anchor_gene'
- summarise(Nott_PLAC_anchor_gene = paste(unique(na.omit(Nott_PLAC_anchor_gene)), collapse = ","), .groups = 'drop')
- # Result
- cCRE_hic_genes <- as.data.frame(cCRE_hic_genes)
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, cCRE_hic_genes, by = intersect(names(iTF_cCREs_df), names(cCRE_hic_genes)), all.x = TRUE)
- #Save files
- saveRDS(merged_iTF_cCREs_df,file="/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- #############
- # monaLisa TF
- #############
- # monaLisa
- monaLisa_iTF_Microglia.LPS.IFNG_vs_iTF_Microglia <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/cCRE_specific_monaLisa_9050_iTF_Microglia.LPS.IFNG_vs_iTF_Microglia_H3K27ac_overlapping_peaks.rds")
- monaLisa_iTF_Microglia_vs_iTF_iPSC <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/cCRE_specific_monaLisa_9054_iTF_Microglia_vs_iTF_iPSCs_ATAC_overlapping_peaks.rds")
- grl <- c(monaLisa_iTF_Microglia.LPS.IFNG_vs_iTF_Microglia,monaLisa_iTF_Microglia_vs_iTF_iPSC)
- grl <- GRangesList(lapply(names(grl), function(nm) {
- gr <- grl[[nm]]
- mcols(gr)$TF <- nm
- gr
- }))
- unlisted_gr <- unlist(grl)
- # Remove row names
- names(unlisted_gr) <- NULL
- # Convert to data.frame without row names
- df <- as.data.frame(unlisted_gr, row.names = NULL)
- df <- unique(df[,c(1:5,15)])
- # Arrange the data frame by seqnames, start, end, and TF
- df_sorted <- df %>%
- arrange(seqnames, start, end, TF)
- df_summarized <- df_sorted %>%
- group_by(seqnames, start, end, width, strand) %>%
- summarise(
- monaLisa.TF = paste(unique(TF), collapse = ","),
- .groups = 'drop'
- )
- monalisa_gr <- GRanges(df_summarized)
- mcols(monalisa_gr)$monaLisa.TF <- toupper(mcols(monalisa_gr)$monaLisa.TF)
- # Find overlaps
- overlaps <- findOverlaps(consensus_CREs_ranges, monalisa_gr)
- # Create a new column with overlapping names
- iTF_cCREs_TF <- consensus_CREs_ranges[queryHits(overlaps)]
- mcols(iTF_cCREs_TF) <- cbind(mcols(iTF_cCREs_TF), mcols(monalisa_gr[subjectHits(overlaps)]))
- iTF_cCREs_TF_df <- as.data.frame(iTF_cCREs_TF)
- iTF_cCREs_TF_df <- unique(iTF_cCREs_TF_df)
- iTF_cCREs_TF_df_collapsed <- iTF_cCREs_TF_df %>%
- group_by(seqnames, start, end) %>%
- summarise(monaLisa.TF = paste(
- unique(sort(unlist(strsplit(monaLisa.TF, ",")))),
- collapse = ","
- ), .groups = "drop") %>% as.data.frame()
- # Merge dataframe
- merged_iTF_cCREs_df <- merge(merged_iTF_cCREs_df, iTF_cCREs_TF_df_collapsed,
- by = intersect(names(iTF_cCREs_df), names(iTF_cCREs_TF_df_collapsed)),
- all.x = TRUE)
- ########################
- # Filter target genes based on their expression
- ########################
- # Load and filter gene expression data
- coldata <- read.csv("/cluster/home/irodriguez/Myers_Lab/20240416_Novogene_RNAseq_Library_iTF_iMGLs_Nick_paper/R_analysis/iTF-Microglia_BulkRNA_SampleSheet_phenotype.csv")
- gse <- readRDS("batch2/Tables/gse.rds")
- tpm <- assay(gse, "abundance")
- coldata_subset <- coldata[c(13:15, 19:24),]
- tpm_subset <- tpm[, c(13:15, 19:24)]
- colnames(tpm_subset) <- coldata_subset$Sample
- #rownames(tpm_subset) <- sapply(strsplit(rownames(tpm_subset), "[.]"), `[[`, 1)
- # Replace ensgid to gene symbol
- gene_map <- readRDS("Tables/Gencode_v46_gene_map.rds")
- # Create a lookup table
- gene_lookup <- setNames(gene_map$gene_name, gene_map$gene_id)
- # Replace ENSG IDs with gene symbols where possible
- new_rownames <- gene_lookup[rownames(tpm_subset)]
- # Keep ENSG IDs for missing symbols
- new_rownames[is.na(new_rownames)] <- rownames(tpm_subset)[is.na(new_rownames)]
- # Assign new row names
- rownames(tpm_subset) <- new_rownames
- # Split columns into groups
- iPSC_cols <- tpm_subset[,1:3]
- Microglia_cols <- tpm_subset[,4:6]
- Microglia_LPS_IFNG_cols <- tpm_subset[,7:9]
- # Compute row-wise means for each group
- iPSC_mean <- rowMeans(iPSC_cols)
- Microglia_mean <- rowMeans(Microglia_cols)
- Microglia_LPS_IFNG_mean <- rowMeans(Microglia_LPS_IFNG_cols)
- # Get rows where any group has an average expression > 1
- iPSC_filtered_rows <- (iPSC_mean > 1)
- Microglia_filtered_rows <- (Microglia_mean > 1)
- Microglia_LPS_IFNG_filtered_rows <- (Microglia_LPS_IFNG_mean > 1)
- # Extract gene names from row names
- iPSC_filtered_genes <- rownames(iPSC_cols)[iPSC_filtered_rows]
- Microglia_filtered_genes <- rownames(Microglia_cols)[Microglia_filtered_rows]
- Microglia_LPS_IFNG_filtered_genes <- rownames(Microglia_LPS_IFNG_cols)[Microglia_LPS_IFNG_filtered_rows]
- filtered_genes <- c(iPSC_filtered_genes,Microglia_filtered_genes,Microglia_LPS_IFNG_filtered_genes)
- filtered_genes <- unique(filtered_genes)
- # Define gene list
- gene_list <- iPSC_filtered_genes #iPSCs
- gene_list <- Microglia_filtered_genes #Microglia
- gene_list <- Microglia_LPS_IFNG_filtered_genes #Microglia_LPS_IFNG
- gene_list <- filtered_genes #all
- # Define columns that need filtering
- gene_cols <- c("DegCre_iTF_iPSCs", "ABC_iTF_iPSCs") # iPSCs
- gene_cols <- c("DegCre_iTF_Microglia.diff", "DegCre_iTF_Microglia.untreated", "ABC_iTF_Microglia") # Microglia
- gene_cols <- c("DegCre_iTF_Microglia.LPS.IFNG", "ABC_iTF_Microglia.LPS.IFNG") # MicrogliaLPS.IFNG
- gene_cols <- c("monaLisa.TF") # filtered_genes
- # Function to filter gene symbols, handling "::" cases
- filter_gene_symbols <- function(col) {
- ifelse(is.na(col), NA,
- sapply(strsplit(col, ","), function(genes) {
- # Process each gene entry
- filtered_genes <- sapply(genes, function(g) {
- # Split by "::" and check if any part is in gene_list
- parts <- unlist(strsplit(g, "::"))
- if (any(parts %in% gene_list)) g else NA
- })
- # Remove NA entries and collapse back
- paste(na.omit(filtered_genes), collapse = ",")
- })
- )
- }
- # Apply the filtering function to the specified columns
- #filtered_iTF_cCREs_df <- merged_iTF_cCREs_df
- filtered_iTF_cCREs_df <- filtered_iTF_cCREs_df %>%
- mutate(across(all_of(gene_cols), filter_gene_symbols))
- #Save files
- saveRDS(filtered_iTF_cCREs_df,file="/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- filtered_iTF_cCREs_df <- readRDS("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_CREs_target_genes.rds")
- filtered_iTF_cCREs_df %>%
- filter(if_any(everything(), ~ grepl("^GRN", .)))
- ########################
- # Count cCREs and target genes
- ########################
- # Format dataframe
- iTF_cCREs_split_df <- filtered_iTF_cCREs_df
- iTF_cCREs_split_df <- unique(iTF_cCREs_split_df[,c(1:3,6,12:19,21)])
- # Remove"Nott_PLAC_anchor_gene" column
- iTF_cCREs_split_df <- unique(iTF_cCREs_split_df[,-13])
- # List of columns to separate
- cols_to_separate <- c(
- "closest_tss",
- "DegCre_iTF_iPSCs",
- "DegCre_iTF_Microglia.diff",
- "DegCre_iTF_Microglia.untreated",
- "DegCre_iTF_Microglia.LPS.IFNG",
- "ABC_iTF_iPSCs",
- "ABC_iTF_Microglia",
- "ABC_iTF_Microglia.LPS.IFNG",
- "Peak2gene.Anderson2023")
- cols_to_separate <- c(
- "closest_tss",
- "DegCre_iTF_iPSCs",
- "DegCre_iTF_Microglia.diff",
- "DegCre_iTF_Microglia.untreated",
- "DegCre_iTF_Microglia.LPS.IFNG",
- "ABC_iTF_iPSCs",
- "ABC_iTF_Microglia",
- "ABC_iTF_Microglia.LPS.IFNG",
- "Peak2gene.Anderson2023",
- "Nott_PLAC_anchor_gene")
- # Apply separate_rows to each specified column
- iTF_cCREs_split_long <- iTF_cCREs_split_df
- # Convert to data.table
- iTF_cCREs_split_long_dt <- as.data.table(iTF_cCREs_split_long)
- # Function to split columns into long format
- split_columns_to_long <- function(dt, cols_to_separate) {
- # Initialize an empty list to store the long-format data
- long_dt_list <- list()
- # Iterate through each column you want to separate
- for (col in cols_to_separate) {
- # For each column, split the comma-separated values into separate rows
- temp_dt <- dt[!is.na(get(col)) & get(col) != "",
- .(seqnames, start, end,
- target_gene = unlist(tstrsplit(get(col), ",", fixed = TRUE)),
- Mapping = col)
- ]
- # Append the processed data.table chunk to the list
- long_dt_list[[col]] <- temp_dt
- }
- # Combine all the chunks into one final long-format data.table
- final_long_dt <- rbindlist(long_dt_list, fill = TRUE)
- # Order the final data.table by seqnames and start
- setorder(final_long_dt, seqnames, start)
- return(final_long_dt)
- }
- # Apply the function to your data.table
- iTF_cCREs_split_long_dt <- split_columns_to_long(iTF_cCREs_split_long_dt, cols_to_separate)
- iTF_cCREs_split_long_dt <-unique(iTF_cCREs_split_long_dt)
- # View the first few rows of the transformed data
- head(iTF_cCREs_split_long_dt)
- # Remove NA
- iTF_cCREs_clean <- iTF_cCREs_split_long_dt[complete.cases(iTF_cCREs_split_long_dt),]
- iTF_cCREs_clean <- iTF_cCREs_clean %>% distinct()
- iTF_cCREs_clean <- iTF_cCREs_clean %>%
- mutate(cell.type = case_when(
- grepl("iTF_iPSCs", Mapping) ~ "iTF-iPSCs",
- grepl("iTF_Microglia$", Mapping) ~ "iTF-Microglia",
- grepl("iTF_Microglia.diff", Mapping) ~ "iTF-Microglia",
- grepl("iTF_Microglia.untreated", Mapping) ~ "iTF-Microglia",
- grepl("LPS.IFNG", Mapping) ~ "iTF-Microglia LPS+IFNG",
- grepl("Anderson2023", Mapping) ~ "Brain microglia",
- TRUE ~ NA_character_ # For rows that don't match any condition
- ))
- iTF_cCREs_clean <- iTF_cCREs_clean %>%
- mutate(mapping = case_when(
- grepl("DegCre", Mapping) ~ "DegCre",
- grepl("ABC_iTF", Mapping) ~ "ABC",
- grepl("closest_tss", Mapping) ~ "closest tss",
- grepl("Peak2gene", Mapping) ~ "Brain microglia link",
- TRUE ~ NA_character_ # For rows that don't match any condition
- ))
- iTF_cCREs_clean <- iTF_cCREs_clean %>%
- mutate(mapping = case_when(
- grepl("DegCre", Mapping) ~ "DegCre",
- grepl("ABC_iTF", Mapping) ~ "ABC",
- grepl("closest_tss", Mapping) ~ "closest tss",
- grepl("Peak2gene", Mapping) ~ "Brain microglia link",
- grepl("Nott", Mapping) ~ "PLAC-seq (Nott)",
- TRUE ~ NA_character_ # For rows that don't match any condition
- ))
- iTF_cCREs_clean$cell.type <- as.factor(iTF_cCREs_clean$cell.type)
- iTF_cCREs_clean$mapping <- as.factor(iTF_cCREs_clean$mapping)
- #iTF_cCREs_clean <- unique(iTF_cCREs_clean[,-5])
- iTF_cCREs_class <- iTF_cCREs_clean
- # Split the data by cell.type
- cell_type_list <- split(iTF_cCREs_clean, iTF_cCREs_clean$cell.type)
- iTF_cCREs_class_gr <- GRanges(iTF_cCREs_class)
- # Overlap with specific cell peaks
- overlap_with_metadata <- function(query, subject) {
- # Find overlaps between query and subject
- hits <- findOverlaps(query, subject)
- # Extract overlapping ranges
- query_overlaps <- query[queryHits(hits)]
- subject_overlaps <- subject[subjectHits(hits)]
- # Combine metadata columns
- combined_mcols <- cbind(mcols(query_overlaps), mcols(subject_overlaps))
- # Create a new GRanges object with combined metadata
- result <- query_overlaps
- mcols(result) <- combined_mcols
- return(result)
- }
- # Assuming ABC_iTF_list is a named list of GRanges objects
- overlapped_list <- lapply(iTF_CREs_ENCODE, function(gr) overlap_with_metadata(iTF_cCREs_class_gr, gr))
- iTF_iPSCs_class <- as.data.frame(overlapped_list$iTF_iPSCs)
- iTF_Microglia_class <- as.data.frame(overlapped_list$iTF_Microglia)
- iTF_Microglia.LPS.IFNG_class <- as.data.frame(overlapped_list$iTF_Microglia.LPS.IFNG)
- iTF_iPSCs_closest.tss <- iTF_iPSCs_class[iTF_iPSCs_class$mapping == "closest tss",]
- iTF_iPSCs_DegCre <- iTF_iPSCs_class[iTF_iPSCs_class$cell.type == "iTF-iPSCs" & iTF_iPSCs_class$mapping == "DegCre",]
- iTF_iPSCs_ABC <- iTF_iPSCs_class[iTF_iPSCs_class$cell.type == "iTF-iPSCs" & iTF_iPSCs_class$mapping == "ABC",]
- iTF_iPSCs_Brain.MG.link <- iTF_iPSCs_class[iTF_iPSCs_class$mapping == "Brain microglia link",]
- #iTF_iPSCs_PLAC <- iTF_iPSCs_class[iTF_iPSCs_class$mapping == "PLAC-seq (Nott)",]
- iTF_Microglia_closest.tss <- iTF_Microglia_class[iTF_Microglia_class$mapping == "closest tss",]
- iTF_Microglia.diff_DegCre <- iTF_Microglia_class[iTF_Microglia_class$Mapping == "DegCre_iTF_Microglia.diff",]
- iTF_Microglia.unt_DegCre <- iTF_Microglia_class[iTF_Microglia_class$Mapping == "DegCre_iTF_Microglia.untreated",]
- iTF_Microglia_ABC <- iTF_Microglia_class[iTF_Microglia_class$Mapping == "ABC_iTF_Microglia",]
- iTF_Microglia_Brain.MG.link <- iTF_Microglia_class[iTF_Microglia_class$mapping == "Brain microglia link",]
- #iTF_Microglia_PLAC <- iTF_Microglia_class[iTF_Microglia_class$mapping == "PLAC-seq (Nott)",]
- iTF_Microglia.LPS.IFNG_closest.tss <- iTF_Microglia.LPS.IFNG_class[iTF_Microglia.LPS.IFNG_class$Mapping == "closest_tss",]
- iTF_Microglia.LPS.IFNG_DegCre <- iTF_Microglia.LPS.IFNG_class[iTF_Microglia.LPS.IFNG_class$Mapping == "DegCre_iTF_Microglia.LPS.IFNG",]
- iTF_Microglia.LPS.IFNG_ABC <- iTF_Microglia.LPS.IFNG_class[iTF_Microglia.LPS.IFNG_class$Mapping == "ABC_iTF_Microglia.LPS.IFNG",]
- iTF_Microglia.LPS.IFNG_Brain.MG.link <- iTF_Microglia.LPS.IFNG_class[iTF_Microglia.LPS.IFNG_class$mapping == "Brain microglia link",]
- #iTF_Microglia.LPS.IFNG_PLAC <- iTF_Microglia.LPS.IFNG_class[iTF_Microglia.LPS.IFNG_class$mapping == "PLAC-seq (Nott)",]
- dim(unique(iTF_iPSCs_closest.tss[,1:3]))
- length(unique(iTF_iPSCs_closest.tss[,6]))
- dim(unique(iTF_iPSCs_DegCre[,1:3])) #42753
- length(unique(iTF_iPSCs_DegCre[,6])) #6645
- dim(unique(iTF_iPSCs_ABC[,1:3])) #39486
- length(unique(iTF_iPSCs_ABC[,6])) #14453
- dim(unique(iTF_iPSCs_Brain.MG.link[,1:3])) #16175
- length(unique(iTF_iPSCs_Brain.MG.link[,6])) #11613
- dim(unique(iTF_iPSCs_PLAC[,1:3])) #16175
- length(unique(iTF_iPSCs_PLAC[,6])) #11613
- dim(unique(iTF_Microglia_closest.tss[,1:3])) #62011
- length(unique(iTF_Microglia_closest.tss[,6])) #26899
- dim(unique(iTF_Microglia.diff_DegCre[,1:3])) #22377
- length(unique(iTF_Microglia.diff_DegCre[,6])) #9885
- dim(unique(iTF_Microglia.unt_DegCre[,1:3])) #22377
- length(unique(iTF_Microglia.unt_DegCre[,6])) #9885
- dim(unique(iTF_Microglia_ABC[,1:3])) #43971
- length(unique(iTF_Microglia_ABC[,6])) #15724
- dim(unique(iTF_Microglia_Brain.MG.link[,1:3])) #23788
- length(unique(iTF_Microglia_Brain.MG.link[,6])) #12240
- dim(unique(iTF_Microglia_PLAC[,1:3])) #16175
- length(unique(iTF_Microglia_PLAC[,6])) #11613
- dim(unique(iTF_Microglia.LPS.IFNG_closest.tss[,1:3])) #62470
- length(unique(iTF_Microglia.LPS.IFNG_closest.tss[,6])) #26387
- dim(unique(iTF_Microglia.LPS.IFNG_DegCre[,1:3])) #26298
- length(unique(iTF_Microglia.LPS.IFNG_DegCre[,6])) #10020
- dim(unique(iTF_Microglia.LPS.IFNG_ABC[,1:3])) #39913
- length(unique(iTF_Microglia.LPS.IFNG_ABC[,6])) #15695
- dim(unique(iTF_Microglia.LPS.IFNG_Brain.MG.link[,1:3])) #23861
- length(unique(iTF_Microglia.LPS.IFNG_Brain.MG.link[,6])) #12191
- dim(unique(iTF_Microglia.LPS.IFNG_PLAC[,1:3])) #16175
- length(unique(iTF_Microglia.LPS.IFNG_PLAC[,6])) #11613
- ########################
- #Cluster profiler target genes
- ########################
- # Load required packages
- library(clusterProfiler)
- library(ReactomePA)
- library(ggplot2)
- library(dplyr)
- library(ComplexUpset)
- library(purrr)
- # Consolidate gene lists
- gene_lists <- list(
- iTF_iPSCs = list(
- `closest TSS` = unique(iTF_iPSCs_closest.tss[, 6]),
- DegCre = unique(iTF_iPSCs_DegCre[, 6]),
- ABC = unique(iTF_iPSCs_ABC[, 6]),
- `Brain microglia link` = unique(iTF_iPSCs_Brain.MG.link[, 6])#,
- #`PLAC-seq (Nott)` = unique(iTF_iPSCs_PLAC[, 6])
- ),
- iTF_Microglia = list(
- `closest TSS` = unique(iTF_Microglia_closest.tss[, 6]),
- DegCre.diff = unique(iTF_Microglia.diff_DegCre[, 6]),
- DegCre.unt = unique(iTF_Microglia.unt_DegCre[, 6]),
- ABC = unique(iTF_Microglia_ABC[, 6]),
- `Brain microglia link` = unique(iTF_Microglia_Brain.MG.link[, 6])#,
- #`PLAC-seq (Nott)` = unique(iTF_Microglia_PLAC[, 6])
- ),
- iTF_Microglia_LPS_IFNG = list(
- `closest TSS` = unique(iTF_Microglia.LPS.IFNG_closest.tss[, 6]),
- DegCre = unique(iTF_Microglia.LPS.IFNG_DegCre[, 6]),
- ABC = unique(iTF_Microglia.LPS.IFNG_ABC[, 6]),
- `Brain microglia link` = unique(iTF_Microglia.LPS.IFNG_Brain.MG.link[, 6])#,
- #`PLAC-seq (Nott)` = unique(iTF_Microglia.LPS.IFNG_PLAC[, 6])
- )
- )
- # Function to map gene names/IDs to Entrez IDs
- convert_to_entrez <- function(gene_vector, gene_map) {
- gene_df <- data.frame(gene_name = gene_vector, stringsAsFactors = FALSE)
- # Ensure gene_map has unique mappings
- gene_map_filtered <- gene_map %>%
- select(gene_name, entrez_id) %>%
- distinct() %>%
- filter(!is.na(entrez_id)) # Remove NA Entrez IDs
- # Perform the mapping
- mapped_genes <- left_join(gene_df, gene_map_filtered, by = "gene_name")
- # Return only Entrez IDs (removing NAs)
- return(na.omit(mapped_genes$entrez_id))
- }
- # Apply function to each list element
- gene_lists_entrez <- map_depth(gene_lists, 2, ~convert_to_entrez(.x, gene_map))
- # Check structure of converted lists
- str(gene_lists_entrez)
- # Perform Reactome enrichment analysis and save results separately
- csv_file_paths <- list()
- for (cell_type in names(gene_lists_entrez)) {
- reactome_results <- compareCluster(
- geneCluster = gene_lists_entrez[[cell_type]],
- fun = "enrichPathway",
- pvalueCutoff = 0.05,
- readable = TRUE
- )
- # Save results to CSV
- csv_file_path <- paste0("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/Reactome_Enrichment_target_gene_", cell_type, ".csv")
- write.csv(reactome_results@compareClusterResult, csv_file_path, row.names = FALSE)
- csv_file_paths[[cell_type]] <- csv_file_path
- }
- # Combine all results for dot plot
- all_reactome_results <- do.call(rbind, lapply(csv_file_paths, read.csv))
- all_reactome_results$mapped.cell <- ifelse(grepl("iTF_iPSCs", rownames(all_reactome_results)), "iTF-iPSCs",
- ifelse(grepl("iTF_Microglia_LPS_IFNG", rownames(all_reactome_results)), "iTF-Microglia\nLPS+IFNG",
- ifelse(grepl("iTF_Microglia.", rownames(all_reactome_results)), "iTF-Microglia", NA)))
- # Define pathways of interest
- pathways <- c(
- "RHO GTPase cycle","Translation","S Phase",
- "Transcriptional regulation of pluripotent stem cells","Neuronal System",
- "Neutrophil degranulation","Inflammasomes","Signaling by Interleukins",
- "Interferon gamma signaling")
- # Subset Reactome results
- reactome_results_subset <- all_reactome_results %>% filter(Description %in% pathways)
- # Ensure Cluster is a factor with the desired order
- reactome_results_subset$Cluster <- factor(reactome_results_subset$Cluster, levels = c("closest TSS", "DegCre",
- "DegCre.diff","DegCre.unt",
- "ABC", "Brain microglia link"#,
- #"PLAC-seq (Nott)"
- ))
- # Order pathways as a factor
- reactome_results_subset$Description <- factor(reactome_results_subset$Description, levels = rev(pathways))
- pdf("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Reactome_Pathway_target_genes_DotPlot.pdf", width = 11, height = 9)
- #pdf("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Reactome_Pathway_target_genes_DotPlot.pdf", width = 10, height = 12)
- #pdf("/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Reactome_Pathway_target_genes_DotPlot.pdf", width = 14, height = 12)
- ggplot(reactome_results_subset, aes(x = Cluster, y = Description, size = Count, fill = p.adjust)) +
- geom_point(shape = 21, color = "black", stroke = 0.5) + # Black outline for each dot
- scale_fill_gradient(low = "blue", high = "red") +
- facet_wrap(~ mapped.cell, scales = "free_x") +
- theme_minimal() +
- labs(title = NULL, x = NULL, y = NULL) +
- theme(
- text = element_text(size = 16),
- axis.text.x = element_text(size = 14, angle = 45, hjust = 1),
- axis.text.y = element_text(size = 16),
- plot.title = element_text(size = 18, face = "bold"),
- strip.background = element_rect(fill = "grey80", color = "black"), # Grey background with black border for facet title
- strip.text = element_text(size = 16, face = "bold"),
- panel.border = element_rect(color = "black", fill = NA, linewidth = 1) # Outline around each facet
- )
- dev.off()
- ########################
- #Overlap CRE-gene pair models
- ########################
- iTF_cCREs_split_df <- unique(filtered_iTF_cCREs_df[, c(1:3,6,12:19,21)])
- iTF_cCREs_split_df <- unique(filtered_iTF_cCREs_df[, c(1:3,6,12:19)])
- # List of columns to separate
- cols_to_separate <- c(
- "closest_tss", "DegCre_iTF_iPSCs", "DegCre_iTF_Microglia.diff",
- "DegCre_iTF_Microglia.untreated",
- "DegCre_iTF_Microglia.LPS.IFNG", "ABC_iTF_iPSCs", "ABC_iTF_Microglia",
- "ABC_iTF_Microglia.LPS.IFNG", "Peak2gene.Anderson2023"#, "Nott_PLAC_anchor_gene"
- )
- # Convert to data.table
- iTF_cCREs_split_long_dt <- as.data.table(iTF_cCREs_split_df)
- # Function to split columns into long format
- split_columns_to_long <- function(dt, cols_to_separate) {
- long_dt_list <- lapply(cols_to_separate, function(col) {
- dt[!is.na(get(col)) & get(col) != "", .(
- seqnames, start, end,
- target_gene = unlist(tstrsplit(get(col), ",", fixed = TRUE)),
- Mapping = col
- )]
- })
- final_long_dt <- rbindlist(long_dt_list, fill = TRUE)
- setorder(final_long_dt, seqnames, start)
- unique(final_long_dt)
- }
- # Apply function
- iTF_cCREs_clean <- split_columns_to_long(iTF_cCREs_split_long_dt, cols_to_separate)
- iTF_cCREs_clean <- iTF_cCREs_clean[complete.cases(iTF_cCREs_clean),] %>% distinct()
- # Process promoters
- gencode_promoters <- gtf_df_filtered %>%
- dplyr::select(seqnames, promoter_start, promoter_end, strand, GeneID, GeneSymb) %>%
- unique() %>%
- filter(seqnames != "chrM")
- promoters_V46 <- GRanges(gencode_promoters[, c("seqnames", "promoter_start", "promoter_end", "GeneSymb")])
- mcols(promoters_V46)$class <- "promoter"
- # Merge with promoters
- iTF_cCREs_clean_promoter <- merge(iTF_cCREs_clean, gencode_promoters,
- by.x = "target_gene", by.y = "GeneSymb",
- allow.cartesian = TRUE) %>%
- dplyr::select(seqnames = seqnames.x, start, end, target_gene, Mapping)
- # Convert to GRanges
- iTF_cCREs_clean_gr <- GRanges(iTF_cCREs_clean_promoter)
- # Find overlaps
- olap_2KB <- findOverlaps(iTF_cCREs_clean_gr, promoters_V46)
- cCREs_2KB <- iTF_cCREs_clean_gr[queryHits(olap_2KB)]
- mcols(cCREs_2KB) <- cbind(mcols(cCREs_2KB), mcols(promoters_V46[subjectHits(olap_2KB)]))
- # Convert to DataFrame
- cCREs_2KB_df <- as.data.frame(cCREs_2KB) %>%
- dplyr::select(seqnames, start, end, GeneSymb) %>%
- distinct() %>%
- mutate(class = "promoter") %>%
- arrange(seqnames, start)
- # Save results
- saveRDS(cCREs_2KB_df, "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Tables/iTF_cCREs_promoter_2KB.rds")
- # Merge and classify for distance
- iTF_cCREs_clean_prom <- merge(as.data.frame(iTF_cCREs_clean_gr), cCREs_2KB_df,
- by = c("seqnames", "start", "end"), all.x = TRUE) %>%
- mutate(
- cCreType = case_when(
- class == "promoter" & target_gene == GeneSymb ~ "SelfPromoter",
- class == "promoter" & target_gene != GeneSymb ~ "OtherPromoter",
- TRUE ~ "enhancer"
- ),
- cell.type = case_when(
- grepl("iTF_iPSCs", Mapping) ~ "iTF-iPSCs",
- grepl("iTF_Microglia$", Mapping) ~ "iTF-Microglia",
- grepl("LPS.IFNG", Mapping) ~ "iTF-Microglia LPS+IFNG",
- grepl("Anderson2023", Mapping) ~ "Brain microglia",
- #grepl("Nott_PLAC", Mapping) ~ "Ex vivo microglia",
- TRUE ~ NA_character_
- ),
- mapping = case_when(
- grepl("DegCre", Mapping) ~ "DegCre",
- grepl("ABC_iTF", Mapping) ~ "ABC",
- grepl("closest_tss", Mapping) ~ "closest TSS",
- grepl("Peak2gene", Mapping) ~ "Brain microglia link",
- #grepl("Nott", Mapping) ~ "PLAC-seq (Nott)",
- TRUE ~ NA_character_
- )
- ) %>%
- mutate(across(c(cCreType, cell.type, mapping), as.factor))
- # Merge and classify for upset plot
- iTF_cCREs_clean_prom <- merge(as.data.frame(iTF_cCREs_clean_gr), cCREs_2KB_df,
- by = c("seqnames", "start", "end"), all.x = TRUE) %>%
- mutate(
- cCreType = case_when(
- class == "promoter" & target_gene == GeneSymb ~ "SelfPromoter",
- class == "promoter" & target_gene != GeneSymb ~ "OtherPromoter",
- TRUE ~ "enhancer"
- )
- ) %>%
- mutate(
- cell.type = case_when(
- grepl("iTF_iPSCs", Mapping) ~ "iTF-iPSCs",
- grepl("iTF_Microglia$", Mapping) ~ "iTF-Microglia",
- grepl("iTF_Microglia.diff", Mapping) ~ "iTF-Microglia",
- grepl("iTF_Microglia.untreated", Mapping) ~ "iTF-Microglia",
- grepl("LPS.IFNG", Mapping) ~ "iTF-Microglia LPS+IFNG",
- grepl("Anderson2023", Mapping) ~ "Brain microglia",
- #grepl("Nott_PLAC", Mapping) ~ "Ex vivo microglia",
- TRUE ~ NA_character_
- ),
- mapping = case_when(
- grepl("DegCre", Mapping) ~ "DegCre",
- grepl("ABC_iTF", Mapping) ~ "ABC",
- grepl("closest_tss", Mapping) ~ "closest TSS",
- grepl("Peak2gene", Mapping) ~ "Brain microglia link",
- #grepl("Nott", Mapping) ~ "PLAC-seq (Nott)",
- TRUE ~ NA_character_
- )
- ) %>%
- mutate(across(c(cCreType, cell.type, mapping), as.factor))
- # Convert to GRanges
- iTF_cCREs_class_gr <- GRanges(iTF_cCREs_clean_prom[, c(1:7, 10, 12)])
- # Overlap function
- overlap_with_metadata <- function(query, subject) {
- hits <- findOverlaps(query, subject)
- query_overlaps <- query[queryHits(hits)]
- subject_overlaps <- subject[subjectHits(hits)]
- result <- query_overlaps
- mcols(result) <- cbind(mcols(query_overlaps), mcols(subject_overlaps))
- result
- }
- # Process overlaps
- overlapped_list <- lapply(iTF_CREs_ENCODE, function(gr) overlap_with_metadata(iTF_cCREs_class_gr, gr))
- # Extract specific cell type classes
- iTF_iPSCs_class <- as.data.frame(overlapped_list$iTF_iPSCs) %>%
- filter(Mapping %in% c(#"Nott_PLAC_anchor_gene",
- "DegCre_iTF_iPSCs", "ABC_iTF_iPSCs", "Peak2gene.Anderson2023", "closest_tss")) %>%
- distinct()
- iTF_Microglia_class <- as.data.frame(overlapped_list$iTF_Microglia) %>%
- filter(Mapping %in% c(#"Nott_PLAC_anchor_gene",
- "DegCre_iTF_Microglia.diff", "DegCre_iTF_Microglia.untreated", "ABC_iTF_Microglia", "Peak2gene.Anderson2023", "closest_tss")) %>%
- distinct()
- iTF_Microglia.LPS.IFNG_class <- as.data.frame(overlapped_list$iTF_Microglia.LPS.IFNG) %>%
- filter(Mapping %in% c(#"Nott_PLAC_anchor_gene",
- "DegCre_iTF_Microglia.LPS.IFNG", "ABC_iTF_Microglia.LPS.IFNG", "Peak2gene.Anderson2023", "closest_tss")) %>%
- distinct()
- # Function to process data
- process_data <- function(df, dataset_name) {
- df <- df %>%
- group_by(across(-c(cCreType))) %>%
- filter(!(cCreType == "OtherPromoter" & "SelfPromoter" %in% cCreType)) %>%
- ungroup() %>%
- mutate(CREtoGene = paste(seqnames, start, end, target_gene, sep = "_"),
- Dataset = dataset_name) %>%
- as.data.frame()
- return(df)
- }
- # Process all datasets
- iTF_cCRE_gene_pair_iPSCs <- process_data(iTF_iPSCs_class, "iTF-iPSCs")
- iTF_cCRE_gene_pair_Microglia <- process_data(iTF_Microglia_class, "iTF-Microglia")
- iTF_cCRE_gene_pair_Microglia_LPS_IFNG <- process_data(iTF_Microglia.LPS.IFNG_class, "iTF-Microglia LPS+IFNG")
- # Combine datasets
- combined_df <- bind_rows(iTF_cCRE_gene_pair_iPSCs,
- iTF_cCRE_gene_pair_Microglia,
- iTF_cCRE_gene_pair_Microglia_LPS_IFNG)
- # Compute distances with correct sign
- cCRE_gr <- GRanges(
- seqnames = combined_df$seqnames,
- ranges = IRanges(start = combined_df$start, end = combined_df$end),
- strand = "*",
- target_gene = combined_df$target_gene
- )
- # Match TSS positions based on the target gene
- matched_tss <- Gencode.V46[match(cCRE_gr$target_gene, Gencode.V46$GeneSymb)]
- # Compute raw distances
- distances <- distance(cCRE_gr, matched_tss)
- # Determine the direction based on gene strand
- strand_info <- as.character(strand(matched_tss))
- # Adjust distances: negative for upstream (on + strand), positive for downstream
- signed_distances <- ifelse(
- strand_info == "+",
- start(cCRE_gr) - start(matched_tss),
- start(matched_tss) - start(cCRE_gr)
- )
- # Assign distances back
- combined_df$distance_to_TSS <- distances
- combined_df$signed_distances <- signed_distances
- # Change sign in distances based on strand
- combined_df$distance_to_TSS <- ifelse(
- combined_df$distance_to_TSS == 0,
- 0,
- ifelse(
- combined_df$signed_distances < 0,
- abs(combined_df$distance_to_TSS) * -1,
- abs(combined_df$distance_to_TSS)
- )
- )
- # Mapping column transformation
- combined_df <- combined_df %>%
- mutate(Mapping = case_when(
- grepl("closest", Mapping) ~ "closest TSS",
- grepl("DegCre", Mapping) ~ "DegCre",
- grepl("ABC", Mapping) ~ "ABC",
- grepl("Anderson2023", Mapping) ~ "Brain microglia link",
- #grepl("Nott", Mapping) ~ "PLAC-seq (Nott)",
- TRUE ~ NA_character_
- ))
- combined_df <- unique(combined_df[,-7])
- # Order
- orderMapping <- c(#"PLAC-seq (Nott)",
- "DegCre","ABC","Brain microglia link", "closest TSS")
- combined_df$mapping <- factor(combined_df$mapping, levels = rev(orderMapping))
- # Define colors
- custom_colors <- c(
- #"PLAC-seq (Nott)" = "#000000",
- "Brain microglia link" = "#009E73",
- "DegCre" = "#F0E442",
- "ABC" = "#0072B2",
- "closest TSS" = "#D55E00"
- )
- # Define your custom linetypes per mapping
- custom_linetypes <- c(
- "DegCre" = "dotdash",
- "ABC" = "dashed",
- "closest TSS" = "solid",
- "Brain microglia link" = "dotted"#,
- #"PLAC-seq (Nott)" = "twodash"
- )
- # Create plot
- library(scales) # For custom transformation
- # Define output files
- Distance_plot_file_log <- "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Distance_CREs_target_genes_log.pdf"
- # Define a signed log transformation
- signed_log_trans <- trans_new(
- name = "signed_log10",
- transform = function(x) sign(x) * log10(abs(x) + 1),
- inverse = function(x) sign(x) * (10^abs(x) - 1)
- )
- # Manually define tick positions (log scale, symmetric around zero)
- log_ticks <- c(-1e5, -1e3, 0, 1e3, 1e5)
- # Create plot
- pdf(Distance_plot_file_log, width=12, height=5)
- # Loop over each unique dataset to generate and plot individually
- ggplot(combined_df, aes(x = distance_to_TSS, color = mapping, linetype = mapping)) +
- geom_density(linewidth = 1.2) +
- facet_wrap(~ Dataset) +
- scale_x_continuous(trans = signed_log_trans, breaks = log_ticks) +
- scale_color_manual(values = custom_colors) +
- scale_linetype_manual(values = custom_linetypes) +
- theme_minimal() +
- theme(
- text = element_text(size = 16),
- axis.text.x = element_text(angle = 45, hjust = 1, size = 14),
- axis.text.y = element_text(size = 14),
- strip.text = element_text(size = 18, face = "bold"),
- legend.text = element_text(size = 14)
- ) +
- labs(
- x = "cCRE distance to target TSS",
- y = "Density",
- color = "Mapping",
- linetype = "Mapping"
- )
- dev.off()
- #######################
- # Overlap target genes
- #######################
- library(ComplexUpset)
- library(ggplot2)
- # Convert to the appropriate format
- upset_data <- iTF_cCRE_gene_pair_iPSCs %>%
- mutate(CREtoGene = factor(CREtoGene))
- upset_data <- iTF_cCRE_gene_pair_Microglia %>%
- mutate(CREtoGene = factor(CREtoGene))
- upset_data <- iTF_cCRE_gene_pair_Microglia_LPS_IFNG %>%
- mutate(CREtoGene = factor(CREtoGene))
- # Upset plot
- upset_data_subset <- unique(upset_data[,c(1:5,8:11)])
- upset_data_collapsed <- upset_data_subset %>%
- group_by(seqnames, start, end, width, strand, cCreType, cell.type, CREtoGene) %>%
- summarise(mapping = paste(mapping, collapse = ","), .groups = "drop")
- upset_data_collapsed <- upset_data_collapsed %>%
- mutate(
- `closest TSS` = ifelse(grepl("closest", mapping), 1, 0),
- ABC = ifelse(grepl("ABC", mapping), 1, 0),
- DegCre = ifelse(grepl("DegCre", mapping), 1, 0),
- `Brain microglia link` = ifelse(grepl("Brain", mapping), 1, 0)#,
- #`PLAC-seq (Nott)` = ifelse(grepl("Nott", mapping), 1, 0),
- ) %>% as.data.frame()
- cats <- colnames(upset_data_collapsed)[10:13]
- orderBind <- c("SelfPromoter","OtherPromoter","enhancer")
- orderBind <- rev(orderBind)
- upset_data_collapsed$cCreType <- factor(upset_data_collapsed$cCreType,levels = orderBind)
- Upset_plot_file = "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Upset_CREs_target_genes_iTF_iPSCs.pdf"
- Upset_plot_file = "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Upset_CREs_target_genes_iTF_Microglia.pdf"
- Upset_plot_file = "/cluster/home/irodriguez/Myers_Lab/KOLF2.1J-iTF_iMGLs_Nick_paper/manuscript/batch2/Figures/Upset_CREs_target_genes_iTF_Microglia_LPS_IFNG.pdf"
- library(grid)
- pdf(Upset_plot_file, width = 10, height = 5)
- p <- upset(
- upset_data_collapsed,
- cats,
- n_intersections = 10,
- # Custom themes
- themes = upset_modify_themes(list(
- 'intersections_matrix' = theme(text = element_text(size = 16))
- )),
- # Customize intersection size bar
- base_annotations = list(
- 'cCREs-TargetGene' = intersection_size(
- counts = FALSE,
- mapping = aes(fill = cCreType)
- ) +
- scale_fill_manual(values = c(
- "SelfPromoter" = "red1",
- "OtherPromoter" = "purple",
- "enhancer" = "gold"
- )) +
- theme(
- text = element_text(size = 18),
- axis.text.y = element_text(size = 14),
- legend.position = "top",
- panel.grid = element_blank()
- )
- ),
- # Customize set size bar
- set_sizes = (
- upset_set_size() +
- theme(
- axis.text.x = element_text(angle = 90, hjust = 1, vjust = 0.5, size = 16),
- text = element_blank()
- )
- ),
- width_ratio = 0.2,
- # Highlight specific sets
- queries = list(
- #upset_query(set = 'PLAC-seq (Nott)', fill = '#000000'),
- upset_query(set = 'Brain microglia link', fill = '#009E73'),
- upset_query(set = 'ABC', fill = '#0072B2'),
- upset_query(set = 'DegCre', fill = '#F0E442'),
- upset_query(set = 'closest TSS', fill = '#D55E00')
- )
- )
- # Ensure the plot is printed to the device and only one page is used
- print(p)
- dev.off()
- # Calculate average number of genes per cCRE
- genes_per_cCRE <- upset_data %>%
- group_by(Mapping, seqnames, start, end) %>%
- summarise(num_genes = n_distinct(target_gene), .groups = "drop") %>%
- group_by(Mapping) %>%
- summarise(avg_genes_per_cCRE = mean(num_genes))
- # Calculate average number of cCREs per gene
- cCREs_per_gene <- upset_data %>%
- group_by(Mapping, target_gene) %>%
- summarise(num_cCREs = n_distinct(paste(seqnames, start, end, sep = "_")), .groups = "drop") %>%
- group_by(Mapping) %>%
- summarise(avg_cCREs_per_gene = mean(num_cCREs))
- # Combine both results
- result_iPSC <- full_join(genes_per_cCRE, cCREs_per_gene, by = "Mapping")
- result_Microglia <- full_join(genes_per_cCRE, cCREs_per_gene, by = "Mapping")
- result_LPS_IFNG <- full_join(genes_per_cCRE, cCREs_per_gene, by = "Mapping")
cCRE-TargetGene_analysis.R at commit 83d4caa, under MIT · at the source
Overview
- HudsonAlpha Institute for Biotechnology, Huntsville, AL 35806, USA
- University of Alabama at Birmingham, Birmingham, AL 35294, USA
Abstract
Understanding transcriptional regulatory networks (TRNs) in microglia is essential for elucidating mechanisms underlying central nervous system (CNS) disorders. Human induced pluripotent stem cell (iPSC)-derived models enable mechanistic studies of microglia but often suffer from variability across lines. Here, we use the standardized KOLF2.1J iPSC line, engineered to inducibly express six transcription factors that allow rapid generation of microglia-like cells (iTF-microglia). We profile TRNs under homeostatic and inflammatory conditions and show that iTF-microglia resemble primary brain microglia at transcriptomic and epigenomic levels. Integrative analyses identify microglia-enriched candidate cis-regulatory elements (cCREs) and reveal dynamic enhancer remodeling during differentiation and stimulation with lipopolysaccharide (LPS) or interferon-gamma (IFNγ), involving NF-κB, IRF, and STAT transcription factors. TRNs active in iTF-microglia are enriched for genetic variants linked to Alzheimer’s disease and related CNS disorders. These findings establish KOLF2.1J iTF-microglia as a reproducible, genetically tractable system for dissecting microglial gene regulation and TRN remodeling in disease.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 27 matches between paragraphs and lines of code.
HudsonAlpha/iTF-Microglia
83d4caaea2aab82e4da550cf5a5f377e5ed611d3, 11 September 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
16 files
- Scripts/
DESeq2_tss_annotation_fo , R, 80 linesr_DegCre.R - Scripts/
DegCre_analysis_20250427 , R, 169 lines.R - Scripts/
LDSC_bed_files_20250429. , R, 194 lines, 1 matchR - Scripts/
LDSC_heatmap.R , R, 333 lines, 1 match - Scripts/
MAGMA_expression_matrix_ , R, 84 lines, 1 matchformat.R - Scripts/
MAGMA_plot.R , R, 128 lines - Scripts/
Partitioned_heritability , R, 83 lines_files.R - Scripts/
RNAseq_analysis.R , R, 987 lines, 4 matches - Scripts/
TF_monaLisa_analysis.R , R, 1,238 lines, 5 matches - Scripts/
cCRE-TargetGene_analysis , R, 1,594 lines, 5 matches.R - Scripts/
cCRE_analysis.R , R, 825 lines, 3 matches - Scripts/
cCRE_disease_variant_ann , R, 1,866 lines, 1 matchotation.R - Scripts/
monalisa_b.R , R, 928 lines, 2 matches - Scripts/
motifbreakR_script.R , R, 535 lines, 2 matches - LICENSE, License, 21 lines
- README.md, Text, 32 lines
ENCODE-DCC/atac-seq-pipeline
47ba8dff9c332e24b48e767303e9fcac98589cf2, 15 February 2024Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
56 files
- dev/
build_on_dx_dockerhub.sh , Shell, 60 lines - dev/
test/ , Shell, 15 linesrun_cromwell_server_on_g c.sh - dev/
test/ , Python, 1 linetest_py/ __init__.py - dev/
test/ , Shell, 3 linestest_workflow/ ref_output/ sync.sh - scripts/
build_genome_data.sh , Shell, 351 lines - scripts/
download_genome_data.sh , Shell, 297 lines - scripts/
install_conda_env.sh , Shell, 88 lines - scripts/
uninstall_conda_env.sh , Shell, 12 lines - scripts/
update_conda_env.sh , Shell, 21 lines - src/
assign_multimappers.py , Python, 73 lines - src/
detect_adapter.py , Python, 84 lines - src/
dev_check_sync_atac.sh , Shell, 21 lines - src/
encode_lib_blacklist_fil , Python, 117 linester.py - src/
encode_lib_common.py , Python, 399 lines - src/
encode_lib_frip.py , Python, 114 lines - src/
encode_lib_genomic.py , Python, 767 lines - src/
encode_lib_log_parser.py , Python, 660 lines - src/
encode_lib_qc_category.p , Python, 259 linesy - src/
encode_task_annot_enrich , Python, 109 lines.py - src/
encode_task_bam2ta.py , Python, 163 lines - src/
encode_task_bam_to_pbam. , Python, 64 linespy - src/
encode_task_bowtie2.py , Python, 192 lines - src/
encode_task_bwa.py , Python, 280 lines - src/
encode_task_choose_ctl.p , Python, 181 linesy - src/
encode_task_compare_sign , Python, 137 linesal_to_roadmap.py - src/
encode_task_count_signal , Python, 115 lines_track.py - src/
encode_task_filter.py , Python, 438 lines - src/
encode_task_frac_mito.py , Python, 86 lines - src/
encode_task_fraglen_stat , Python, 244 lines_pe.py - src/
encode_task_gc_bias.py , Python, 143 lines - src/
encode_task_idr.py , Python, 213 lines - src/
encode_task_jsd.py , Python, 133 lines - src/
encode_task_macs2_atac.p , Python, 139 linesy - src/
encode_task_macs2_chip.p , Python, 172 linesy - src/
encode_task_macs2_signal , Python, 190 lines_track_atac.py - src/
encode_task_macs2_signal , Python, 215 lines_track_chip.py - src/
encode_task_merge_fastq. , Python, 105 linespy - src/
encode_task_overlap.py , Python, 178 lines, 1 match - src/
encode_task_pool_ta.py , Python, 76 lines - src/
encode_task_post_align.p , Python, 92 linesy - src/
encode_task_post_call_pe , Python, 104 linesak_atac.py - src/
encode_task_post_call_pe , Python, 107 linesak_chip.py - src/
encode_task_preseq.py , Python, 174 lines - src/
encode_task_qc_report.py , Python, 941 lines, 1 match - src/
encode_task_reproducibil , Python, 204 linesity.py - src/
encode_task_spp.py , Python, 133 lines - src/
encode_task_spr.py , Python, 176 lines - src/
encode_task_subsample_ct , Python, 66 linesl.py - src/
encode_task_trim_adapter , Python, 214 lines.py - src/
encode_task_trim_fastq.p , Python, 73 linesy - src/
encode_task_trimmomatic. , Python, 204 linespy - src/
encode_task_tss_enrich.p , Python, 178 linesy - src/
encode_task_xcor.py , Python, 156 lines - src/
trimfastq.py , Python, 255 lines - LICENSE, License, 21 lines
- README.md, Text, 184 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 68 scripts, each with its path and the digest of its content;
- 27 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE299850, at NCBI GEO; found in the text, “Key resources table”
Data and code availability
• RNA-sequencing, ATAC-sequencing, ChIP-sequencing and Hi-C data have been deposited at GEO database under GEO Superseries: GSE299850 and are publicly available. • The analysis scripts used to generate the images and statistical output in this study are available on GitHub (https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added J. Nicholas Cochran (0000-0002-9852-5504); removed J. Nicholas Cochran
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 3 keywords, 6 funders, 99 references, 4 RRIDs.
Cite
This paper
Rogers, B. B., Anderson, A. G., Rodriguez-Nunez, I., Bartley, S. C., Johnston, S. Q., Taylor, J. W., Meadows, S. K., Newberry, K. M., Myers, R. M., & Cochran, J. N. (2026). KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease. iScience, 29(7), 116412. https://
BibTeX
@article{rogers2026kolf2
author = {Rogers, Brianne B. and Anderson, Ashlyn G. and Rodriguez-Nunez, Ivan and Bartley, Samuel C. and Johnston, S. Quinn and Taylor, Jared W. and Meadows, Sarah K. and Newberry, Kimberly M. and Myers, Richard M. and Cochran, J. Nicholas},
title = {{KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease}},
journal = {iScience},
year = {2026},
month = jun,
volume = {29},
number = {7},
pages = {116412},
publisher = {Elsevier},
issn = {2589-0042},
doi = {10.1016/
url = {https://
pmid = {42338472},
pmcid = {PMC13285636}
}
RIS
TY - JOUR
AU - Rogers, Brianne B.
AU - Anderson, Ashlyn G.
AU - Rodriguez-Nunez, Ivan
AU - Bartley, Samuel C.
AU - Johnston, S. Quinn
AU - Taylor, Jared W.
AU - Meadows, Sarah K.
AU - Newberry, Kimberly M.
AU - Myers, Richard M.
AU - Cochran, J. Nicholas
TI - KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease
T2 - iScience
J2 - iScience
PY - 2026
DA - 2026/
VL - 29
IS - 7
SP - 116412
SN - 2589-0042
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease",
"container-title": "iScience",
"author": [
{
"family": "Rogers",
"given": "Brianne B."
},
{
"family": "Anderson",
"given": "Ashlyn G."
},
{
"family": "Rodriguez-Nunez",
"given": "Ivan"
},
{
"family": "Bartley",
"given": "Samuel C."
},
{
"family": "Johnston",
"given": "S. Quinn"
},
{
"family": "Taylor",
"given": "Jared W."
},
{
"family": "Meadows",
"given": "Sarah K."
},
{
"family": "Newberry",
"given": "Kimberly M."
},
{
"family": "Myers",
"given": "Richard M."
},
{
"family": "Cochran",
"given": "J. Nicholas"
}
],
"container-title-short":
"volume": "29",
"issue": "7",
"page": "116412",
"DOI": "10.1016/
"PMID": "42338472",
"PMCID": "PMC13285636",
"ISSN": "2589-0042",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
16
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-73007-1 [code]
- Single-nucleus epigenomic dysregulation unmasks genetic risk-associated neurodegenerative glia states.Journal: Nature communicationsIn common: SAMtools, circlize, ComplexHeatmap, 7 other tools, genetics / omics, 13 references
- [2] doi:10.1038/s41593-026-02367-0 [code]
- A reproducible three-dimensional model of human brain tissue to investigate physiological and disease-associated microglia phenotypes.Journal: Nature neuroscienceIn common: edgeR, circlize, DESeq2, 7 other tools, 7 references
- [3] doi:10.1038/s41386-026-02406-1 [code]
- Functional genomic profiling of schizophrenia-associated
genes reveals key microglial regulators. Journal: Neuropsychopharmacology : official publication of the American College of NeuropsychopharmacologyIn common: SAMtools, circlize, DESeq2, 7 other tools, genetics / omics, 6 references - [4] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: BEDTools, SAMtools, edgeR, 13 other tools, 1 reference
- [5] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: SAMtools, edgeR, circlize, 12 other tools, genetics / omics, 2 references
- [6] doi:10.1101/gr.281113.125 [code]
- Single-nucleus multiomic profiling of the aging mouse substantia nigra reveals conserved gene alterations linked to Parkinson's disease.Journal: Genome researchIn common: BEDTools, edgeR, circlize, 9 other tools, genetics / omics, 3 references
- [7] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: BEDTools, SAMtools, edgeR, 11 other tools, genetics / omics, 2 references
- [8] doi:10.1038/s41467-026-69944-6 [code]
- Multi-modal dissection of cell-type specific TDP-43 pathology in the motor cortex.Journal: Nature communicationsIn common: BEDTools, SAMtools, circlize, 12 other tools, genetics / omics, 1 reference
- [9] doi:10.1038/s41467-026-71790-5 [code]
- Recurrent DNA break clusters drive replication-stress-induc
ed copy number variants and genome diversification. Journal: Nature communicationsIn common: BEDTools, SAMtools, edgeR, 12 other tools, genetics / omics - [10] doi:10.1038/s41467-026-72598-z [code]
- Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.Journal: Nature communicationsIn common: BEDTools, SAMtools, DESeq2, 11 other tools, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 68 scripts, and 27 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e467ae8b62ec735f…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
