The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation.
The 5 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Results › Reduced association of the PAF1 complex with the elongating RNA Pol II in the absence of SETD2 ↔ chromID/chromID_analysis.Rmd, lines 342–363 · score 0.81 · nTurbo, log2 FC, LFQ intensity, log2 fold change, TurboID, enriched proteins
- [2] § Methods › PolyA RNA-sequencing and differential gene expression analysis ↔ rnaseq/analysis/RNA-seq-analysis-DGE-GSEA.Rmd, lines 201–246 · score 0.76 · filterByExpr, glmTreat, edgeR, GSEA, rnaseq, model
- [3] § Methods › ChromID and label-free MS data acquisition and analysis ↔ chromID/chromID_analysis.Rmd, lines 19–57 · score 0.73 · MaxQuant, biotin incubation, bait, MS, Proteus, log2
- [4] § Methods › PolyA RNA-sequencing and differential gene expression analysis ↔ rnaseq/mapping_rnaseq_nfcore/rnaseq_nextflow_submission.sh, the whole file · a weak match · score 0.71 · gencode.vM35.annotation.gtf, rnaseq, Salmon, STAR, nf, trimgalore
- [5] § Results › Reduced association of the PAF1 complex with the elongating RNA Pol II in the absence of SETD2 ↔ chromID/chromID_analysis.Rmd, lines 265–303 · score 0.53 · proteins enriched, TurboID, Setd2 KO, NLS, SRI, WT
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 371 lines · 19 KB · no license · 3 matches
- ---
- title: "Ambrosi_ChromID_Analysis"
- author: "Ramon Pfaendler"
- date: "`r format(Sys.time(), '%d %B, %Y')`"
- output:
- html_document:
- theme: spacelab
- highlight: tango
- df_print: paged
- toc: true
- toc_float: true
- code_folding: "hide"
- fig_crop: false
- code_download: true
- editor_options:
- chunk_output_type: console
- ---
- ## Setup {.tabset .tabset-fade}
- ```{r setup, include=TRUE, warning=F, message=F}
- library(tidyr)
- library(tidyverse)
- # the knitr package is required to set an a different working directory to e.g. a different root where data is stored (while keeping the Rmd file stored in a different place simultaneously!)
- library(knitr)
- library(ggplot2)
- library(tidyverse)
- library(plotly)
- library(ggrepel)
- library(gridExtra)
- library(circlize)
- library(ComplexHeatmap)
- library(pheatmap)
- library(proteus)
- library(limma)
- devtools::install_github("bartongroup/Proteus", build_opts= c("--no-resave-data", "--no-manual"), build_vignettes=TRUE)
- # new function that creates a volcano plot more easily than just writing it every time
- volcano_rp <- function(data,comparison, timepoint, cellline,expnumber ){
- ggplot() + geom_point(data, mapping = aes(text = names,x = logFC, y = -log10(adj.P.Val), colour = significant)) +
- xlim(c(-12,12)) +geom_text_repel(data %>% filter(significant == "TRUE"), mapping = aes(y = -log10(adj.P.Val), x = logFC, label = names), cex = 2.5, colour = "darkblue", max.overlaps = 50, segment.linetype = 'dotted', segment.alpha = 0.6) +
- scale_color_manual(values = c("palegreen3","darkblue")) +
- ggtitle(paste(comparison," \n(Biotin Incubation = ",timepoint,")\n(",cellline,";", expnumber,")")) +
- xlab("enrichment log2(bait/control)") +
- ylab("-log10 adj.p-value")+annotate("text", x= -8.75, y = 8, label = "adj.p-val = 0.05", cex = 4) + theme_bw(base_size = 15)
- }
- # load the evidence.txt file - this is the output from MaxQuant
- # Change the path to the correct location :-)
- evi_sri<- proteus::readEvidenceFile("/Users/ramonpfaendler/Documents/UZH/BaubecLab/MS/rpx77_3/rpx77_3_evidence.txt")
- ```
- ## SRI-ChromID Data {.tabset .tabset-fade}
- ### Loading the data, playing around with some EDA
- ```{r}
- meta_sri <- as.data.frame(unique(evi_sri$experiment))
- colnames(meta_sri) <- "experiment"
- #create measure column (just repeat "intensity" for each sample, as intensities of the peptides are available in the evidence.txt)
- meta_sri$measure <- rep("Intensity",length(meta_sri$experiment))
- # create sample column ("evade" unnecessary characters in the front, in this case, first the rpx.. is removed, followed by the removal of numbers (with \\d )and the letter S)
- gsub("\\d\\d_S\\d\\d\\d\\d\\d\\d_","",unique(meta_sri$experiment)) %>% gsub("(TurboID_\\d).*","\\1",.) %>% gsub ("\\d\\d\\d\\d\\d\\d_\\d","",.) %>% gsub("rpx77_3_","",.) -> meta_sri$sample
- # create condition column, to do this, arrange the dataframe first (arrange by sample, to have it properly ordered!), to make it easier later!
- meta_sri %>% arrange(., sample) -> meta_sri
- meta_sri$condition[1:16] <- c(rep("WT - NLS-TurboID",4), rep("Setd2-KO - NLS-TurboID",4), rep("Setd2-KO - SRI-TurboID",4), rep("WT - SRI-TurboID",4))
- # add replicate number and column
- meta_sri %>% group_by(condition) %>% mutate(replicate = row_number()) -> meta_sri
- # concatenate strings to generate simpler sample ID than the one coming from FGCZ
- meta_sri %>% unite(., sample_new,condition, replicate, sep = "_") %>% .$sample_new -> sample_new
- meta_sri$sample <- sample_new
- ```
- ## Creation of peptide and protein dataset using all 16 samples {.tabset .tabset-fade}
- ### Create a peptide and a protein datast by merging evidence and meta data
- This section is based on the proteus package. Check also the other tab that uses similar visualisations of the same data.
- ```{r}
- # merge the files
- pepdat_sri <- proteus::makePeptideTable(evi_sri, meta_sri)
- # plot peptide counts per condition (of non-zero peptide intensities per sample)
- proteus::plotCount(pepdat_sri)
- prodat_sri <- proteus::makeProteinTable(pepdat_sri)
- proteus::plotCount(prodat_sri)
- # on the protein level, the samples seem to be quite similar to each other
- # next one can normalize data: this way, the data is normalized the way that the median
- # intensity in all samples is the same
- prodat.med_sri <- proteus::normalizeData(prodat_sri)
- # plot data as violin plots prior normalisation
- proteus::plotSampleDistributions(prodat_sri, title="Not normalised", fill="condition", method="violin")
- proteus::plotSampleDistributions(prodat.med_sri, title="Median normalised", fill="condition", method="violin")
- #with median normalisation, all samples seem to be in the same range which, I think, makes it plausible to keep all of them.
- ```
- ### Plotting independently of Proteus Package
- In this section, I tried to visualise some of the qc plots independently of the proteus package. This gives more flexibility, for example to render visualiations to log2 space or other properties. (log2 space might be interesting later to check if proteins only present in one condition are found on the lower or higher level of the intensity distribution.)
- ```{r}
- # to create violin plots, one needs to reshape the data into long format
- prodat_sri$tab %>% reshape2::melt(., id.vars = NULL) -> prodat_reshaped
- # add condition column, which can later be used for visualisation purposes
- prodat_reshaped$condition <- prodat_reshaped$Var2 %>% as.vector(.) %>% gsub(".{2}$", "",.)
- prodat_reshaped$condition %>% unique(.) # sanity check
- ggplot(prodat_reshaped, aes(x = Var2, y = log2(value), fill = condition )) + geom_violin() +theme_light() + theme(axis.text.x = element_text(angle = 90, hjust = 1,
- vjust = 0.5, size =8)) + labs(fill = "Condition", x = "Sample", y = "log2(LFQ_intensity)")+ ggtitle("Non-normalised LFQ Intensitites") + geom_boxplot(width = 0.1, outlier.shape = NA, show.legend = F)+ scale_fill_brewer(palette = "Pastel2")
- # to create violin plots, one needs to reshape the data into long format
- prodat.med_sri$tab %>% reshape2::melt(., id.vars = NULL) -> prodat_med_reshaped
- # add condition column, which can later be used for visualisation purposes
- prodat_med_reshaped$condition <- prodat_reshaped$Var2 %>% as.vector(.) %>% gsub(".{2}$", "",.)
- prodat_med_reshaped$condition %>% unique(.) # sanity check
- ggplot(prodat_med_reshaped, aes(x = Var2, y = log2(value), fill = condition )) + geom_violin() +theme_light() + theme(axis.text.x = element_text(angle = 90, hjust = 1,
- vjust = 0.5, size =8)) + labs(fill = "Condition", x = "Sample", y = "log2(LFQ_intensity)") + ggtitle("Median-normalised LFQ Intensitites") + geom_boxplot(width = 0.1, outlier.shape = NA, show.legend = F) + scale_fill_brewer(palette = "Pastel2")
- ```
- ## Differential Enrichment Calculations {.tabset .tabset-fade}
- In this section, different conditions are compared to each other. The proteus package provides limma-based enrichment calculations. Here, it is important to note, that proteins that are only found in one group and are entirely missing in the control group (or vice versa) will not show a logFC or p-value, since the linear model cannot find a value to compare it to! To do this, imputation-based methods, such as Perseus, would be needed.
- For the individual contrasts, we've always used an FDR of 0.05 as significance-threshold. (The FDR is based on a Benjamini-Hochberg adjusted p-value.)
- ### nTurboID vs. SETD2KO nTurboID
- ```{r}
- meta_sri$condition %>% unique()
- ctrls_rpx77 <- proteus::limmaDE(prodat.med_sri, conditions=c( "WT - NLS-TurboID","Setd2-KO - NLS-TurboID" ), sig.level = 0.05)
- #20 NAs
- ctrls_rpx77$names <- gsub('sp.*\\|','',ctrls_rpx77$protein)
- ctrls_rpx77$names <- gsub('_MOUSE','',ctrls_rpx77$names)
- #assign uniprot id's before rendering the names of the majority protein ids
- names_up <- (gsub('sp.|*\\*','',ctrls_rpx77$protein))
- ctrls_rpx77$UniprotID <- str_extract(names_up,"\\w+|")
- #use new function of volcano plots specified earlier in the script
- volcano_rp(ctrls_rpx77, comparison = paste(unique(meta_sri$condition)[1], "vs.", unique(meta_sri$condition)[2]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_ctrls_rpx77
- p_ctrls_rpx77
- # grid.arrange(p_sri_rpx77, p_sri_setd2ko_rpx77)
- library(plotly)
- # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
- ggplotly(p_ctrls_rpx77, tooltip = c("text","x","y"))
- ```
- ## ChromID analysis based on merged controls {.tabset .tabset-fade}
- ### Updated meta-data to unify nTurbo controls in one condition
- Given that the the control conditions in both backgrounds did not display any differential enrichment, we subsequently merged the control conditions from the WT and SETD2 -/- genetic backgrounds and perform all the DE analysis compared to these merged control samples.
- ```{r}
- meta_sri2 <- meta_sri
- meta_sri2$condition[1:16] <- c(rep("NLS-TurboID",8), rep("Setd2-KO - SRI-TurboID",4), rep("WT - SRI-TurboID",4))
- #updated conditions column, sample column remains the same to see which one's which later on!
- # merge the files
- pepdat_sri2 <- proteus::makePeptideTable(evi_sri, meta_sri2)
- prodat_sri2 <- proteus::makeProteinTable(pepdat_sri2)
- prodat.med_sri2 <- proteus::normalizeData(prodat_sri2)
- ```
- ### DE - SRI WT vs. nTurboID (HA36CB1, RPX77)
- ```{r}
- meta_sri2$condition
- #make sure to use the normalised version here
- sri_rpx77_2 <- proteus::limmaDE(prodat.med_sri2, conditions=c( "WT - SRI-TurboID","NLS-TurboID" ), sig.level = 0.05)
- # 32 NAs; 7 less than in non-grouped comparison
- sri_rpx77_2$names <- gsub('sp.*\\|','',sri_rpx77_2$protein)
- sri_rpx77_2$names <- gsub('_MOUSE','',sri_rpx77_2$names)
- #assign uniprot id's before rendering the names of the majority protein ids
- # names_up <- (gsub('sp.|*\\*','',sri_rpx77_2$protein))
- sri_rpx77_2$UniprotID <- str_extract(names_up,"\\w+|")
- #use new function of volcano plots specified earlier in the script
- volcano_rp(sri_rpx77_2, comparison = paste(unique(meta_sri2$condition)[3], "vs.", unique(meta_sri$condition)[1]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_sri_rpx77_2
- p_sri_rpx77_2
- # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
- ggplotly(p_sri_rpx77_2, tooltip = c("text","x","y"))
- ```
- ### DE - SRI WT vs. nTurboID (HA36CB1 SETD2KO, RPX77)
- ```{r}
- sri_setd2ko_rpx77_2 <- proteus::limmaDE(prodat.med_sri2, conditions=c( "Setd2-KO - SRI-TurboID","NLS-TurboID" ), sig.level = 0.05)
- #22 NAs; 7 less than in non-grouped comparison
- sri_setd2ko_rpx77_2$names <- gsub('sp.*\\|','',sri_setd2ko_rpx77_2$protein)
- sri_setd2ko_rpx77_2$names <- gsub('_MOUSE','',sri_setd2ko_rpx77_2$names)
- #assign uniprot id's before rendering the names of the majority protein ids
- sri_setd2ko_rpx77_2$UniprotID <- str_extract(names_up,"\\w+|")
- #use new function of volcano plots specified earlier in the script
- volcano_rp(sri_setd2ko_rpx77_2, comparison = paste(unique(meta_sri2$condition)[2], "vs.", unique(meta_sri2$condition)[1]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_sri_setd2ko_rpx77_2
- p_sri_setd2ko_rpx77_2
- # grid.arrange(p_sri_rpx77, p_sri_setd2ko_rpx77)
- library(plotly)
- # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
- ggplotly(p_sri_setd2ko_rpx77_2, tooltip = c("text","x","y"))
- ```
- ### Proteins enriched compared to grouped control
- In this section, proteins that were significantly enriched (or depleted for SRI vs SRI comparision) are displayed using the median normalised LFQ values.
- ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
- c(sri_rpx77_2 %>% filter(.$significant ==T & .$logFC >0) %>% .$protein %>% as.character()) %>% unique(.) -> sig_HA36CB1_2
- c(sri_setd2ko_rpx77_2 %>% filter(.$significant ==T & .$logFC >0) %>% .$protein %>% as.character()) %>% unique(.) -> sig_SETD2KO_2
- top_anno <- HeatmapAnnotation(condition = colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("Setd2-KO - NLS-TurboID","Setd2-KO - SRI-TurboID","WT - SRI-TurboID")
- ) %>% lapply(., function(x) if(identical(x, character(0))) "WT - NLS-TurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) , col = list(
- condition = structure(names= colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("Setd2-KO - NLS-TurboID","Setd2-KO - SRI-TurboID","WT - SRI-TurboID")
- ) %>% lapply(., function(x) if(identical(x, character(0))) "WT - NLS-TurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) %>% unique(),c('#b3e2cd','#fdcdac','#e6f5c9','#f4cae4'))))
- #c(setdiff(sig_HA36CB1_2, sig_SETD2KO_2),setdiff(sig_SETD2KO_2, sig_HA36CB1_2)) %>% unique(.)
- as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_HA36CB1_2) %>% column_to_rownames('protein')) -> mat_ha36_2
- as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_SETD2KO_2) %>% column_to_rownames('protein')) -> mat_setd2ko_2
- #top_anno2 <- HeatmapAnnotation(condition = colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("SETD2KO_nTurboID","SETD2KO_SRI-WT_TurboID","SRI-WT_TurboID")) %>% lapply(., function(x) if(identical(x, character(0))) "nTurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) , col = list( condition = structure(names= colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("SETD2KO_nTurboID","SETD2KO_SRI-WT_TurboID","SRI-WT_TurboID")) %>% lapply(., function(x) if(identical(x, character(0))) "nTurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) %>% unique(),c("#B3E2CD", "#FDCDAC", "#CBD5E8", "#F4CAE4"))))
- rownames(mat_ha36_2) <- gsub('.*\\|','',rownames(mat_ha36_2)) %>% gsub('_MOUSE','',.)
- Heatmap(mat_ha36_2,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_HA36CB1_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 6),gap = unit(3, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = "log2(LFQ intensity)"), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(mat_ha36_2)*unit(5,"mm"), height = nrow(mat_ha36_2)*unit(2.5,"mm"), row_title = "Proteins enriched in HA36CB1", row_title_gp = gpar(fontsize = 10)) ->h_ha36_2
- rownames(mat_setd2ko_2) <- gsub('.*\\|','',rownames(mat_setd2ko_2)) %>% gsub('_MOUSE','',.)
- Heatmap(mat_setd2ko_2,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_SETD2KO_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 6),gap = unit(3, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = "log2(LFQ intensity)"), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(mat_setd2ko_2)*unit(5,"mm"), height = nrow(mat_setd2ko_2)*unit(2.5,"mm"), row_title = "Proteins enriched in SETD2KO", row_title_gp = gpar(fontsize = 10)) ->h_setd2ko_2
- h_list_2 = h_setd2ko_2 %v% h_ha36_2
- draw(h_list_2, merge_legend = T)
- ```
- ### Combined Heatmaps
- The same as above, just with an additional row annotation to plot the heatmap only once.
- ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
- sig_2 <- c(sig_HA36CB1_2,sig_SETD2KO_2) %>% unique(.)
- as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_2) %>% column_to_rownames('protein')) -> another_mat
- Heatmap(another_mat,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 3.5),gap = unit(2, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = paste("log2(LFQ intensity)\nn Proteins enriched in HA36CB1 =",length(sig_2))), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(another_mat)*unit(2.5,"mm"), height = nrow(another_mat)*unit(1,"mm"), row_title = "Proteins enriched in either foreground", row_title_gp = gpar(fontsize = 10)) ->h_another
- draw(h_another+ rowAnnotation(HA36CB1 = ifelse(sig_2 %in% sig_HA36CB1_2 & !(sig_2 %in% sig_SETD2KO_2), "yes", "no"), col = list(HA36CB1 = c("yes" = 'grey50', "no"= "white"))) +rowAnnotation(SETD2KO = ifelse(sig_2 %in% sig_SETD2KO_2 & !(sig_2 %in% sig_HA36CB1_2), "yes", "no"), col = list(SETD2KO = c("yes" = 'grey80', "no"= "white"))) + rowAnnotation(both = ifelse(sig_2 %in% sig_HA36CB1_2 & (sig_2 %in% sig_SETD2KO_2), "yes", "no"), col = list(both = c("yes" = 'darkgreen', "no"= "white"))), merge_legend = T)
- # try to get rowannotation with all info of where proteins enriched in one
- anno_df <- as.data.frame(ifelse(sig_2 %in% sig_HA36CB1_2 & !(sig_2 %in% sig_SETD2KO_2), "HA36CB1", NA))
- colnames(anno_df)[1] <- "ha36cb1"
- anno_df$setd2ko <- ifelse(sig_2 %in% sig_SETD2KO_2 & !(sig_2 %in% sig_HA36CB1_2), "SETD2KO", NA)
- anno_df$both <- ifelse(sig_2 %in% sig_HA36CB1_2 & (sig_2 %in% sig_SETD2KO_2), "both", NA)
- # weird workaround that seems to work later on
- within(anno_df, ha36cb1 <- ifelse(is.na(ha36cb1), setd2ko,ha36cb1)) ->anno_df2
- within(anno_df2, ha36cb1 <- ifelse(is.na(ha36cb1), both,ha36cb1)) ->anno_df3
- anno_df3$ha36cb1[anno_df3$ha36cb1 == 1] <- "HA36CB1"
- draw(h_another + rowAnnotation(Enrichment = (anno_df3$ha36cb1) ,col = list(Enrichment = c("HA36CB1" = "#7FC97F", "SETD2KO" = "#BEAED4", "both" = "#FDC086"))), merge_legend = T)
- ```
- ### Heatmaps of log2-FC
- Heatmap visualisation of all significanly enriched proteins in log2 fold-change world.
- ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
- sri_rpx77_2 %>% filter(., .$protein %in% sig_2) %>% column_to_rownames('names') ->df_1
- sri_setd2ko_rpx77_2 %>% filter(., .$protein %in% sig_2) %>% column_to_rownames('names') -> df_2
- cbind(df_1, df_2$logFC) -> comb_df
- colnames(comb_df)
- colnames(comb_df)[2] = "log2FC_HA36CB1_SRI-WT-TurboID"
- colnames(comb_df)[14] = "log2FC_SETD2KO_SRI-WT-TurboID"
- Heatmap(as.matrix(comb_df[,c(2,14)]), col = colorRamp2(c(-4, 0, 4), c("#2166ac","#f7f7f7","#b2182b")), show_column_dend = FALSE, name = paste("log2(fc)\nn Proteins =",length(comb_df$protein), ""),row_names_side = "left", row_dend_side = "right", row_dend_width = unit(4, "cm"), row_names_gp = gpar(fontsize = 4), clustering_method_rows = "ward.D", gap = unit(2, "mm"), column_names_gp = gpar(fontsize = 4), heatmap_legend_param = list(title = paste("log2(fc(LFQ intensity))\nn Proteins =",length(comb_df$protein)), title_gp = gpar(fontsize = 5), labels_gp = gpar(fontsize = 5)), width = ncol(comb_df)*unit(1,"mm"), height = nrow(comb_df)*unit(1.4,"mm"), row_title = "log2(FC) vs. nTurbo-grouped controls", row_title_gp = gpar(fontsize = 10))
- ```
- ## Session Info
- ```{r}
- sessionInfo()$R.version
- installed.packages()[names(sessionInfo()$otherPkgs), "Version"]
- ```
chromID_analysis.Rmd at commit 9d6f50c, no license · at the source
Overview
- Department of Molecular Mechanisms of Disease, University of Zurich, Zurich, Switzerland
- Life Science Zurich Graduate School, University of Zurich and ETH Zurich, Zurich, Switzerland
- Present Address: Institute of Molecular Systems Biology, ETH Zurich, Zurich, Switzerland
- Genome Biology and Epigenetics, Institute of Biodynamics and Biocomplexity, Department of Biology, Utrecht University, Utrecht, The Netherlands
- Present Address: Institute of Experimental Immunology, University of Zurich, Zurich, Switzerland
Abstract
Chromatin modifications are essential for mammalian development, and their aberrant deposition is associated with human disease. While the mechanisms that deposit and remove histone modifications have been largely elucidated, their roles in regulating gene activity during cellular differentiation are yet to be fully understood. Here, we performed a deletion screen to identify stage-specific requirements of chromatin regulators during neuronal differentiation of mouse embryonic stem cells. We show that the H3K36me3 methyltransferase SETD2 is required for the establishment of neuronal gene expression during late stages of differentiation but is dispensable in mature neurons. Notably, this function is independent of its histone methyltransferase activity. Instead, SETD2 promotes interactions between the PAF1 complex and elongating RNA Pol II, suggesting a role in supporting efficient transcription of neuronal genes.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.
FelixKrueger/TrimGalore
c6528e54512e0e388a392d36475291bcbf0eb0dd, 27 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
31 files
- build.rs, Rust, 76 lines
- docs/
scripts/ , Python, 344 linesgenerate-benchmark-chart s.py - docs/
src/ , TypeScript, 7 linescontent.config.ts - docs/
src/ , TypeScript, 167 linespages/ og/ [...route].ts - plans/
06252026_ubam-input-supp , Rust, 323 linesort/ spikes/ spike1-recordsource/ src/ main.rs - plans/
06252026_ubam-input-supp , Shell, 59 linesort/ spikes/ spike2-paired-ordering/ build_test_bams.sh - plans/
06252026_ubam-input-supp , Rust, 107 linesort/ spikes/ spike2-paired-ordering/ deinterleave_sketch.rs - scripts/
benchmark.sh , Shell, 178 lines - src/
adapter.rs , Rust, 911 lines - src/
alignment.rs , Rust, 704 lines - src/
bam.rs , Rust, 1,672 lines - src/
cli.rs , Rust, 1,667 lines - src/
clump.rs , Rust, 567 lines - src/
demux.rs , Rust, 404 lines - src/
fastq.rs , Rust, 895 lines - src/
fastqc.rs , Rust, 172 lines - src/
filters.rs , Rust, 253 lines - src/
format.rs , Rust, 206 lines - src/
io.rs , Rust, 549 lines - src/
lib.rs , Rust, 16 lines - src/
main.rs , Rust, 2,029 lines - src/
parallel.rs , Rust, 2,350 lines - src/
quality.rs , Rust, 507 lines - src/
report.rs , Rust, 1,764 lines - src/
specialty.rs , Rust, 787 lines - src/
trimmer.rs , Rust, 1,172 lines - tests/
integration_passthrough. , Rust, 205 linesrs - tests/
integration_ubam.rs , Rust, 258 lines - tests/
integration_ubam_out.rs , Rust, 536 lines - LICENSE, License, 674 lines
- README.md, Text, 191 lines
BaubecLab/Setd2
9d6f50ce61cf1d521f4f2889a56d8e3a2a03192f, 17 September 2025Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
6 files
- chip_seq/
mapping/ , Shell, 102 lineschip_seq_input.sh - chip_seq/
mapping/ , Shell, 34 lineschip_seq_slurm_submissio n.sh - chromID/
chromID_analysis.Rmd , R, 371 lines, 3 matches - rnaseq/
analysis/ , R, 397 lines, 1 matchRNA-seq-analysis-DGE-GSE A.Rmd - rnaseq/
mapping_rnaseq_nfcore/ , Shell, 38 lines, 1 matchrnaseq_nextflow_submissi on.sh - README.md, Text, 19 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 34 scripts, each with its path and the digest of its content;
- 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE278955, at NCBI GEO; found in “Data availability”
Other data links
- ebi.ac.uk/
biostudies/ , EMBL-EBI; found in the text, “Author contributions”sourcedata
Data availability
High-throughput sequencing data obtained from ChIP-seq, ATAC-seq, RNA-seq, and single-cell RNA-seq studies were deposited to NCBI Gene Expression Omnibus under the following accessions: GSE278955 (https://
The source data of this paper are collected in the following database record: biostudies:S-SCDT-10_103
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 13 authors, 5 keywords, 10 MeSH terms, 2 funders, 78 references.
Cite
This paper
Ambrosi, C., Pfaendler, R., Eleftheriou, K., Butz, S., Recchia, D., Bao, X., Cardoso da Silva, R., Kupfer, N., Lagerwaard, I. M., Vlaming, H., Schmolka, N., Bhardwaj, V., & Baubec, T. (2026). The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation. The EMBO journal, 45(10), 3430-3443. https://
BibTeX
@article{ambrosi2026h3k3
author = {Ambrosi, Christina and Pfaendler, Ramon and Eleftheriou, Kristeli and Butz, Stefan and Recchia, Davide and Bao, Xue and Cardoso da Silva, Richard and Kupfer, Niklas and Lagerwaard, Ilse M and Vlaming, Hanneke and Schmolka, Nina and Bhardwaj, Vivek and Baubec, Tuncay},
title = {{The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation}},
journal = {The EMBO journal},
year = {2026},
month = apr,
volume = {45},
number = {10},
pages = {3430--3443},
publisher = {Nature Publishing Group},
issn = {0261-4189},
doi = {10.1038/
url = {https://
pmid = {41963556},
pmcid = {PMC13187324}
}
RIS
TY - JOUR
AU - Ambrosi, Christina
AU - Pfaendler, Ramon
AU - Eleftheriou, Kristeli
AU - Butz, Stefan
AU - Recchia, Davide
AU - Bao, Xue
AU - Cardoso da Silva, Richard
AU - Kupfer, Niklas
AU - Lagerwaard, Ilse M
AU - Vlaming, Hanneke
AU - Schmolka, Nina
AU - Bhardwaj, Vivek
AU - Baubec, Tuncay
TI - The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation
T2 - The EMBO journal
J2 - EMBO J
PY - 2026
DA - 2026/
VL - 45
IS - 10
SP - 3430
EP - 3443
SN - 0261-4189
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation",
"container-title": "The EMBO journal",
"author": [
{
"family": "Ambrosi",
"given": "Christina"
},
{
"family": "Pfaendler",
"given": "Ramon"
},
{
"family": "Eleftheriou",
"given": "Kristeli"
},
{
"family": "Butz",
"given": "Stefan"
},
{
"family": "Recchia",
"given": "Davide"
},
{
"family": "Bao",
"given": "Xue"
},
{
"family": "Cardoso da Silva",
"given": "Richard"
},
{
"family": "Kupfer",
"given": "Niklas"
},
{
"family": "Lagerwaard",
"given": "Ilse M"
},
{
"family": "Vlaming",
"given": "Hanneke"
},
{
"family": "Schmolka",
"given": "Nina"
},
{
"family": "Bhardwaj",
"given": "Vivek"
},
{
"family": "Baubec",
"given": "Tuncay"
}
],
"container-title-short":
"volume": "45",
"issue": "10",
"page": "3430-3443",
"DOI": "10.1038/
"PMID": "41963556",
"PMCID": "PMC13187324",
"ISSN": "0261-4189",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s11689-026-09713-0 [code]
- DRP1 mutations associated with EMPF1 encephalopathy perturb the transcriptional profile and maturation of cortical neurons.Journal: Journal of neurodevelopmental disordersIn common: edgeR, limma, circlize, 8 other tools, 4 references
- [2] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: SAMtools, edgeR, limma, 9 other tools, 1 reference
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: SAMtools, edgeR, limma, 9 other tools, mouse
- [4] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: edgeR, limma, circlize, 8 other tools, 2 references
- [5] doi:10.1038/s41467-026-69944-6 [code]
- Multi-modal dissection of cell-type specific TDP-43 pathology in the motor cortex.Journal: Nature communicationsIn common: deepTools, SAMtools, circlize, 7 other tools, 2 references
- [6] doi:10.1038/s41467-026-74753-y [code]
- A human-specific microRNA controls the timing of excitatory synaptogenesis.Journal: Nature communicationsIn common: SAMtools, edgeR, limma, 7 other tools, 2 references
- [7] doi:10.1016/j.xcrm.2026.102682 [code]
- TET CpG sequence-context-specifi
c DNA demethylation shapes progression of IDH-mutant gliomas. Journal: Cell reports. MedicineIn common: Nextflow, edgeR, limma, 7 other tools, 1 reference - [8] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: SAMtools, edgeR, circlize, 7 other tools, mouse, 2 references
- [9] doi:10.1111/adb.70179 [code]
- Transcriptional Response to Chronic Long-Access Fentanyl Self-Administration in Rat Habenula and Amygdala.Journal: Addiction biologyIn common: Nextflow, edgeR, limma, 7 other tools
- [10] doi:10.1016/j.cpblue.2026.100007 [code]
- An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.Journal: Cell press blueIn common: SAMtools, edgeR, limma, 8 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 34 scripts, and 5 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:feae5a72109e7a94…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
