OSCR

The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Results › Reduced association of the PAF1 complex with the elongating RNA Pol II in the absence of SETD2 ↔ chromID/chromID_analysis.Rmd, lines 342–363 · score 0.81 · nTurbo, log2 FC, LFQ intensity, log2 fold change, TurboID, enriched proteins
  2. [2] § Methods › PolyA RNA-sequencing and differential gene expression analysis ↔ rnaseq/analysis/RNA-seq-analysis-DGE-GSEA.Rmd, lines 201–246 · score 0.76 · filterByExpr, glmTreat, edgeR, GSEA, rnaseq, model
  3. [3] § Methods › ChromID and label-free MS data acquisition and analysis ↔ chromID/chromID_analysis.Rmd, lines 19–57 · score 0.73 · MaxQuant, biotin incubation, bait, MS, Proteus, log2
  4. [4] § Methods › PolyA RNA-sequencing and differential gene expression analysis ↔ rnaseq/mapping_rnaseq_nfcore/rnaseq_nextflow_submission.sh, the whole file · a weak match · score 0.71 · gencode.vM35.annotation.gtf, rnaseq, Salmon, STAR, nf, trimgalore
  5. [5] § Results › Reduced association of the PAF1 complex with the elongating RNA Pol II in the absence of SETD2 ↔ chromID/chromID_analysis.Rmd, lines 265–303 · score 0.53 · proteins enriched, TurboID, Setd2 KO, NLS, SRI, WT

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R Markdown · 371 lines · 19 KB · no license · 3 matches

  1. ---
  2. title: "Ambrosi_ChromID_Analysis"
  3. author: "Ramon Pfaendler"
  4. date: "`r format(Sys.time(), '%d %B, %Y')`"
  5. output:
  6. html_document:
  7. theme: spacelab
  8. highlight: tango
  9. df_print: paged
  10. toc: true
  11. toc_float: true
  12. code_folding: "hide"
  13. fig_crop: false
  14. code_download: true
  15. editor_options:
  16. chunk_output_type: console
  17. ---
  18. ## Setup {.tabset .tabset-fade}
  19. ```{r setup, include=TRUE, warning=F, message=F}
  20. library(tidyr)
  21. library(tidyverse)
  22. # the knitr package is required to set an a different working directory to e.g. a different root where data is stored (while keeping the Rmd file stored in a different place simultaneously!)
  23. library(knitr)
  24. library(ggplot2)
  25. library(tidyverse)
  26. library(plotly)
  27. library(ggrepel)
  28. library(gridExtra)
  29. library(circlize)
  30. library(ComplexHeatmap)
  31. library(pheatmap)
  32. library(proteus)
  33. library(limma)
  34. devtools::install_github("bartongroup/Proteus", build_opts= c("--no-resave-data", "--no-manual"), build_vignettes=TRUE)
  35. # new function that creates a volcano plot more easily than just writing it every time
  36. volcano_rp <- function(data,comparison, timepoint, cellline,expnumber ){
  37. ggplot() + geom_point(data, mapping = aes(text = names,x = logFC, y = -log10(adj.P.Val), colour = significant)) +
  38. xlim(c(-12,12)) +geom_text_repel(data %>% filter(significant == "TRUE"), mapping = aes(y = -log10(adj.P.Val), x = logFC, label = names), cex = 2.5, colour = "darkblue", max.overlaps = 50, segment.linetype = 'dotted', segment.alpha = 0.6) +
  39. scale_color_manual(values = c("palegreen3","darkblue")) +
  40. ggtitle(paste(comparison," \n(Biotin Incubation = ",timepoint,")\n(",cellline,";", expnumber,")")) +
  41. xlab("enrichment log2(bait/control)") +
  42. ylab("-log10 adj.p-value")+annotate("text", x= -8.75, y = 8, label = "adj.p-val = 0.05", cex = 4) + theme_bw(base_size = 15)
  43. }
  44. # load the evidence.txt file - this is the output from MaxQuant
  45. # Change the path to the correct location :-)
  46. evi_sri<- proteus::readEvidenceFile("/Users/ramonpfaendler/Documents/UZH/BaubecLab/MS/rpx77_3/rpx77_3_evidence.txt")
  47. ```
  48. ## SRI-ChromID Data {.tabset .tabset-fade}
  49. ### Loading the data, playing around with some EDA
  50. ```{r}
  51. meta_sri <- as.data.frame(unique(evi_sri$experiment))
  52. colnames(meta_sri) <- "experiment"
  53. #create measure column (just repeat "intensity" for each sample, as intensities of the peptides are available in the evidence.txt)
  54. meta_sri$measure <- rep("Intensity",length(meta_sri$experiment))
  55. # create sample column ("evade" unnecessary characters in the front, in this case, first the rpx.. is removed, followed by the removal of numbers (with \\d )and the letter S)
  56. gsub("\\d\\d_S\\d\\d\\d\\d\\d\\d_","",unique(meta_sri$experiment)) %>% gsub("(TurboID_\\d).*","\\1",.) %>% gsub ("\\d\\d\\d\\d\\d\\d_\\d","",.) %>% gsub("rpx77_3_","",.) -> meta_sri$sample
  57. # create condition column, to do this, arrange the dataframe first (arrange by sample, to have it properly ordered!), to make it easier later!
  58. meta_sri %>% arrange(., sample) -> meta_sri
  59. meta_sri$condition[1:16] <- c(rep("WT - NLS-TurboID",4), rep("Setd2-KO - NLS-TurboID",4), rep("Setd2-KO - SRI-TurboID",4), rep("WT - SRI-TurboID",4))
  60. # add replicate number and column
  61. meta_sri %>% group_by(condition) %>% mutate(replicate = row_number()) -> meta_sri
  62. # concatenate strings to generate simpler sample ID than the one coming from FGCZ
  63. meta_sri %>% unite(., sample_new,condition, replicate, sep = "_") %>% .$sample_new -> sample_new
  64. meta_sri$sample <- sample_new
  65. ```
  66. ## Creation of peptide and protein dataset using all 16 samples {.tabset .tabset-fade}
  67. ### Create a peptide and a protein datast by merging evidence and meta data
  68. This section is based on the proteus package. Check also the other tab that uses similar visualisations of the same data.
  69. ```{r}
  70. # merge the files
  71. pepdat_sri <- proteus::makePeptideTable(evi_sri, meta_sri)
  72. # plot peptide counts per condition (of non-zero peptide intensities per sample)
  73. proteus::plotCount(pepdat_sri)
  74. prodat_sri <- proteus::makeProteinTable(pepdat_sri)
  75. proteus::plotCount(prodat_sri)
  76. # on the protein level, the samples seem to be quite similar to each other
  77. # next one can normalize data: this way, the data is normalized the way that the median
  78. # intensity in all samples is the same
  79. prodat.med_sri <- proteus::normalizeData(prodat_sri)
  80. # plot data as violin plots prior normalisation
  81. proteus::plotSampleDistributions(prodat_sri, title="Not normalised", fill="condition", method="violin")
  82. proteus::plotSampleDistributions(prodat.med_sri, title="Median normalised", fill="condition", method="violin")
  83. #with median normalisation, all samples seem to be in the same range which, I think, makes it plausible to keep all of them.
  84. ```
  85. ### Plotting independently of Proteus Package
  86. In this section, I tried to visualise some of the qc plots independently of the proteus package. This gives more flexibility, for example to render visualiations to log2 space or other properties. (log2 space might be interesting later to check if proteins only present in one condition are found on the lower or higher level of the intensity distribution.)
  87. ```{r}
  88. # to create violin plots, one needs to reshape the data into long format
  89. prodat_sri$tab %>% reshape2::melt(., id.vars = NULL) -> prodat_reshaped
  90. # add condition column, which can later be used for visualisation purposes
  91. prodat_reshaped$condition <- prodat_reshaped$Var2 %>% as.vector(.) %>% gsub(".{2}$", "",.)
  92. prodat_reshaped$condition %>% unique(.) # sanity check
  93. ggplot(prodat_reshaped, aes(x = Var2, y = log2(value), fill = condition )) + geom_violin() +theme_light() + theme(axis.text.x = element_text(angle = 90, hjust = 1,
  94. vjust = 0.5, size =8)) + labs(fill = "Condition", x = "Sample", y = "log2(LFQ_intensity)")+ ggtitle("Non-normalised LFQ Intensitites") + geom_boxplot(width = 0.1, outlier.shape = NA, show.legend = F)+ scale_fill_brewer(palette = "Pastel2")
  95. # to create violin plots, one needs to reshape the data into long format
  96. prodat.med_sri$tab %>% reshape2::melt(., id.vars = NULL) -> prodat_med_reshaped
  97. # add condition column, which can later be used for visualisation purposes
  98. prodat_med_reshaped$condition <- prodat_reshaped$Var2 %>% as.vector(.) %>% gsub(".{2}$", "",.)
  99. prodat_med_reshaped$condition %>% unique(.) # sanity check
  100. ggplot(prodat_med_reshaped, aes(x = Var2, y = log2(value), fill = condition )) + geom_violin() +theme_light() + theme(axis.text.x = element_text(angle = 90, hjust = 1,
  101. vjust = 0.5, size =8)) + labs(fill = "Condition", x = "Sample", y = "log2(LFQ_intensity)") + ggtitle("Median-normalised LFQ Intensitites") + geom_boxplot(width = 0.1, outlier.shape = NA, show.legend = F) + scale_fill_brewer(palette = "Pastel2")
  102. ```
  103. ## Differential Enrichment Calculations {.tabset .tabset-fade}
  104. In this section, different conditions are compared to each other. The proteus package provides limma-based enrichment calculations. Here, it is important to note, that proteins that are only found in one group and are entirely missing in the control group (or vice versa) will not show a logFC or p-value, since the linear model cannot find a value to compare it to! To do this, imputation-based methods, such as Perseus, would be needed.
  105. For the individual contrasts, we've always used an FDR of 0.05 as significance-threshold. (The FDR is based on a Benjamini-Hochberg adjusted p-value.)
  106. ### nTurboID vs. SETD2KO nTurboID
  107. ```{r}
  108. meta_sri$condition %>% unique()
  109. ctrls_rpx77 <- proteus::limmaDE(prodat.med_sri, conditions=c( "WT - NLS-TurboID","Setd2-KO - NLS-TurboID" ), sig.level = 0.05)
  110. #20 NAs
  111. ctrls_rpx77$names <- gsub('sp.*\\|','',ctrls_rpx77$protein)
  112. ctrls_rpx77$names <- gsub('_MOUSE','',ctrls_rpx77$names)
  113. #assign uniprot id's before rendering the names of the majority protein ids
  114. names_up <- (gsub('sp.|*\\*','',ctrls_rpx77$protein))
  115. ctrls_rpx77$UniprotID <- str_extract(names_up,"\\w+|")
  116. #use new function of volcano plots specified earlier in the script
  117. volcano_rp(ctrls_rpx77, comparison = paste(unique(meta_sri$condition)[1], "vs.", unique(meta_sri$condition)[2]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_ctrls_rpx77
  118. p_ctrls_rpx77
  119. # grid.arrange(p_sri_rpx77, p_sri_setd2ko_rpx77)
  120. library(plotly)
  121. # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
  122. ggplotly(p_ctrls_rpx77, tooltip = c("text","x","y"))
  123. ```
  124. ## ChromID analysis based on merged controls {.tabset .tabset-fade}
  125. ### Updated meta-data to unify nTurbo controls in one condition
  126. Given that the the control conditions in both backgrounds did not display any differential enrichment, we subsequently merged the control conditions from the WT and SETD2 -/- genetic backgrounds and perform all the DE analysis compared to these merged control samples.
  127. ```{r}
  128. meta_sri2 <- meta_sri
  129. meta_sri2$condition[1:16] <- c(rep("NLS-TurboID",8), rep("Setd2-KO - SRI-TurboID",4), rep("WT - SRI-TurboID",4))
  130. #updated conditions column, sample column remains the same to see which one's which later on!
  131. # merge the files
  132. pepdat_sri2 <- proteus::makePeptideTable(evi_sri, meta_sri2)
  133. prodat_sri2 <- proteus::makeProteinTable(pepdat_sri2)
  134. prodat.med_sri2 <- proteus::normalizeData(prodat_sri2)
  135. ```
  136. ### DE - SRI WT vs. nTurboID (HA36CB1, RPX77)
  137. ```{r}
  138. meta_sri2$condition
  139. #make sure to use the normalised version here
  140. sri_rpx77_2 <- proteus::limmaDE(prodat.med_sri2, conditions=c( "WT - SRI-TurboID","NLS-TurboID" ), sig.level = 0.05)
  141. # 32 NAs; 7 less than in non-grouped comparison
  142. sri_rpx77_2$names <- gsub('sp.*\\|','',sri_rpx77_2$protein)
  143. sri_rpx77_2$names <- gsub('_MOUSE','',sri_rpx77_2$names)
  144. #assign uniprot id's before rendering the names of the majority protein ids
  145. # names_up <- (gsub('sp.|*\\*','',sri_rpx77_2$protein))
  146. sri_rpx77_2$UniprotID <- str_extract(names_up,"\\w+|")
  147. #use new function of volcano plots specified earlier in the script
  148. volcano_rp(sri_rpx77_2, comparison = paste(unique(meta_sri2$condition)[3], "vs.", unique(meta_sri$condition)[1]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_sri_rpx77_2
  149. p_sri_rpx77_2
  150. # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
  151. ggplotly(p_sri_rpx77_2, tooltip = c("text","x","y"))
  152. ```
  153. ### DE - SRI WT vs. nTurboID (HA36CB1 SETD2KO, RPX77)
  154. ```{r}
  155. sri_setd2ko_rpx77_2 <- proteus::limmaDE(prodat.med_sri2, conditions=c( "Setd2-KO - SRI-TurboID","NLS-TurboID" ), sig.level = 0.05)
  156. #22 NAs; 7 less than in non-grouped comparison
  157. sri_setd2ko_rpx77_2$names <- gsub('sp.*\\|','',sri_setd2ko_rpx77_2$protein)
  158. sri_setd2ko_rpx77_2$names <- gsub('_MOUSE','',sri_setd2ko_rpx77_2$names)
  159. #assign uniprot id's before rendering the names of the majority protein ids
  160. sri_setd2ko_rpx77_2$UniprotID <- str_extract(names_up,"\\w+|")
  161. #use new function of volcano plots specified earlier in the script
  162. volcano_rp(sri_setd2ko_rpx77_2, comparison = paste(unique(meta_sri2$condition)[2], "vs.", unique(meta_sri2$condition)[1]), timepoint = "12 hours", cellline = "mNPCs", expnumber = "RPX77") ->p_sri_setd2ko_rpx77_2
  163. p_sri_setd2ko_rpx77_2
  164. # grid.arrange(p_sri_rpx77, p_sri_setd2ko_rpx77)
  165. library(plotly)
  166. # to plot the name of the protein later in the plotly plot, specify an additional aes variable (here: text = Majority.protein.IDs), this one can be used later in the ggplotly command by using tooltip!
  167. ggplotly(p_sri_setd2ko_rpx77_2, tooltip = c("text","x","y"))
  168. ```
  169. ### Proteins enriched compared to grouped control
  170. In this section, proteins that were significantly enriched (or depleted for SRI vs SRI comparision) are displayed using the median normalised LFQ values.
  171. ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
  172. c(sri_rpx77_2 %>% filter(.$significant ==T & .$logFC >0) %>% .$protein %>% as.character()) %>% unique(.) -> sig_HA36CB1_2
  173. c(sri_setd2ko_rpx77_2 %>% filter(.$significant ==T & .$logFC >0) %>% .$protein %>% as.character()) %>% unique(.) -> sig_SETD2KO_2
  174. top_anno <- HeatmapAnnotation(condition = colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("Setd2-KO - NLS-TurboID","Setd2-KO - SRI-TurboID","WT - SRI-TurboID")
  175. ) %>% lapply(., function(x) if(identical(x, character(0))) "WT - NLS-TurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) , col = list(
  176. condition = structure(names= colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("Setd2-KO - NLS-TurboID","Setd2-KO - SRI-TurboID","WT - SRI-TurboID")
  177. ) %>% lapply(., function(x) if(identical(x, character(0))) "WT - NLS-TurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) %>% unique(),c('#b3e2cd','#fdcdac','#e6f5c9','#f4cae4'))))
  178. #c(setdiff(sig_HA36CB1_2, sig_SETD2KO_2),setdiff(sig_SETD2KO_2, sig_HA36CB1_2)) %>% unique(.)
  179. as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_HA36CB1_2) %>% column_to_rownames('protein')) -> mat_ha36_2
  180. as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_SETD2KO_2) %>% column_to_rownames('protein')) -> mat_setd2ko_2
  181. #top_anno2 <- HeatmapAnnotation(condition = colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("SETD2KO_nTurboID","SETD2KO_SRI-WT_TurboID","SRI-WT_TurboID")) %>% lapply(., function(x) if(identical(x, character(0))) "nTurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) , col = list( condition = structure(names= colnames(prodat.med_sri$tab) %>% map(function(x, y) y[str_detect(x, y)], c("SETD2KO_nTurboID","SETD2KO_SRI-WT_TurboID","SRI-WT_TurboID")) %>% lapply(., function(x) if(identical(x, character(0))) "nTurboID" else x) %>% lapply(., function(x) x[1]) %>% unlist(.) %>% unique(),c("#B3E2CD", "#FDCDAC", "#CBD5E8", "#F4CAE4"))))
  182. rownames(mat_ha36_2) <- gsub('.*\\|','',rownames(mat_ha36_2)) %>% gsub('_MOUSE','',.)
  183. Heatmap(mat_ha36_2,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_HA36CB1_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 6),gap = unit(3, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = "log2(LFQ intensity)"), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(mat_ha36_2)*unit(5,"mm"), height = nrow(mat_ha36_2)*unit(2.5,"mm"), row_title = "Proteins enriched in HA36CB1", row_title_gp = gpar(fontsize = 10)) ->h_ha36_2
  184. rownames(mat_setd2ko_2) <- gsub('.*\\|','',rownames(mat_setd2ko_2)) %>% gsub('_MOUSE','',.)
  185. Heatmap(mat_setd2ko_2,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_SETD2KO_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 6),gap = unit(3, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = "log2(LFQ intensity)"), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(mat_setd2ko_2)*unit(5,"mm"), height = nrow(mat_setd2ko_2)*unit(2.5,"mm"), row_title = "Proteins enriched in SETD2KO", row_title_gp = gpar(fontsize = 10)) ->h_setd2ko_2
  186. h_list_2 = h_setd2ko_2 %v% h_ha36_2
  187. draw(h_list_2, merge_legend = T)
  188. ```
  189. ### Combined Heatmaps
  190. The same as above, just with an additional row annotation to plot the heatmap only once.
  191. ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
  192. sig_2 <- c(sig_HA36CB1_2,sig_SETD2KO_2) %>% unique(.)
  193. as.matrix(prodat.med_sri2$tab %>% log2(.) %>% as.data.frame(.)%>% rownames_to_column('protein') %>% filter(.$protein %in% sig_2) %>% column_to_rownames('protein')) -> another_mat
  194. Heatmap(another_mat,col = colorRamp2((c(12,22,32)), c("white", "lightskyblue1", "firebrick")), show_column_dend = T, name = paste("log2(fc)\nn Proteins =",length(sig_2), ""),row_names_side = "left", row_dend_side = "right", row_names_gp = gpar(fontsize = 3.5),gap = unit(2, "mm"), column_names_gp = gpar(fontsize = 6),heatmap_legend_param = list(title = paste("log2(LFQ intensity)\nn Proteins enriched in HA36CB1 =",length(sig_2))), cluster_columns = T, na_col = "white", top_annotation = top_anno, cluster_rows = T, width = ncol(another_mat)*unit(2.5,"mm"), height = nrow(another_mat)*unit(1,"mm"), row_title = "Proteins enriched in either foreground", row_title_gp = gpar(fontsize = 10)) ->h_another
  195. draw(h_another+ rowAnnotation(HA36CB1 = ifelse(sig_2 %in% sig_HA36CB1_2 & !(sig_2 %in% sig_SETD2KO_2), "yes", "no"), col = list(HA36CB1 = c("yes" = 'grey50', "no"= "white"))) +rowAnnotation(SETD2KO = ifelse(sig_2 %in% sig_SETD2KO_2 & !(sig_2 %in% sig_HA36CB1_2), "yes", "no"), col = list(SETD2KO = c("yes" = 'grey80', "no"= "white"))) + rowAnnotation(both = ifelse(sig_2 %in% sig_HA36CB1_2 & (sig_2 %in% sig_SETD2KO_2), "yes", "no"), col = list(both = c("yes" = 'darkgreen', "no"= "white"))), merge_legend = T)
  196. # try to get rowannotation with all info of where proteins enriched in one
  197. anno_df <- as.data.frame(ifelse(sig_2 %in% sig_HA36CB1_2 & !(sig_2 %in% sig_SETD2KO_2), "HA36CB1", NA))
  198. colnames(anno_df)[1] <- "ha36cb1"
  199. anno_df$setd2ko <- ifelse(sig_2 %in% sig_SETD2KO_2 & !(sig_2 %in% sig_HA36CB1_2), "SETD2KO", NA)
  200. anno_df$both <- ifelse(sig_2 %in% sig_HA36CB1_2 & (sig_2 %in% sig_SETD2KO_2), "both", NA)
  201. # weird workaround that seems to work later on
  202. within(anno_df, ha36cb1 <- ifelse(is.na(ha36cb1), setd2ko,ha36cb1)) ->anno_df2
  203. within(anno_df2, ha36cb1 <- ifelse(is.na(ha36cb1), both,ha36cb1)) ->anno_df3
  204. anno_df3$ha36cb1[anno_df3$ha36cb1 == 1] <- "HA36CB1"
  205. draw(h_another + rowAnnotation(Enrichment = (anno_df3$ha36cb1) ,col = list(Enrichment = c("HA36CB1" = "#7FC97F", "SETD2KO" = "#BEAED4", "both" = "#FDC086"))), merge_legend = T)
  206. ```
  207. ### Heatmaps of log2-FC
  208. Heatmap visualisation of all significanly enriched proteins in log2 fold-change world.
  209. ```{r,fig.height = 9, fig.width = 11, fig.align = "center"}
  210. sri_rpx77_2 %>% filter(., .$protein %in% sig_2) %>% column_to_rownames('names') ->df_1
  211. sri_setd2ko_rpx77_2 %>% filter(., .$protein %in% sig_2) %>% column_to_rownames('names') -> df_2
  212. cbind(df_1, df_2$logFC) -> comb_df
  213. colnames(comb_df)
  214. colnames(comb_df)[2] = "log2FC_HA36CB1_SRI-WT-TurboID"
  215. colnames(comb_df)[14] = "log2FC_SETD2KO_SRI-WT-TurboID"
  216. Heatmap(as.matrix(comb_df[,c(2,14)]), col = colorRamp2(c(-4, 0, 4), c("#2166ac","#f7f7f7","#b2182b")), show_column_dend = FALSE, name = paste("log2(fc)\nn Proteins =",length(comb_df$protein), ""),row_names_side = "left", row_dend_side = "right", row_dend_width = unit(4, "cm"), row_names_gp = gpar(fontsize = 4), clustering_method_rows = "ward.D", gap = unit(2, "mm"), column_names_gp = gpar(fontsize = 4), heatmap_legend_param = list(title = paste("log2(fc(LFQ intensity))\nn Proteins =",length(comb_df$protein)), title_gp = gpar(fontsize = 5), labels_gp = gpar(fontsize = 5)), width = ncol(comb_df)*unit(1,"mm"), height = nrow(comb_df)*unit(1.4,"mm"), row_title = "log2(FC) vs. nTurbo-grouped controls", row_title_gp = gpar(fontsize = 10))
  217. ```
  218. ## Session Info
  219. ```{r}
  220. sessionInfo()$R.version
  221. installed.packages()[names(sessionInfo()$otherPkgs), "Version"]
  222. ```

chromID_analysis.Rmd at commit 9d6f50c, no license · at the source

Overview

Authors: Christina Ambrosi1,2, Ramon Pfaendler1,2,3, Kristeli Eleftheriou4, Stefan Butz1,2, Davide Recchia1,2,4, Xue Bao4, Richard Cardoso da Silva4, Niklas Kupfer4, Ilse M Lagerwaard4, Hanneke Vlaming4, Nina Schmolka1,5, Vivek Bhardwaj4, Tuncay Baubec1,4
  1. Department of Molecular Mechanisms of Disease, University of Zurich, Zurich, Switzerland
  2. Life Science Zurich Graduate School, University of Zurich and ETH Zurich, Zurich, Switzerland
  3. Present Address: Institute of Molecular Systems Biology, ETH Zurich, Zurich, Switzerland
  4. Genome Biology and Epigenetics, Institute of Biodynamics and Biocomplexity, Department of Biology, Utrecht University, Utrecht, The Netherlands
  5. Present Address: Institute of Experimental Immunology, University of Zurich, Zurich, Switzerland
Institutions: University of Zurich (Switzerland); ETH Zurich (Switzerland); Life Science Zurich (Switzerland); Institute for Molecular Systems Biology (Switzerland); Utrecht University (Netherlands)
Journal: The EMBO journal, volume 45, issue 10, pages 3430-3443
Dates: received 27 February 2025; accepted 11 March 2026; published online 10 April 2026; in print May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s44318-026-00768-2 · PMID 41963556 · PMCID PMC13187324 · OpenAlex W7153067057
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Connectivity
Keywords: Chromatin, H3K36me3, SETD2, Transcription & Genomics, Development
MeSH: Cell Differentiation*, Histone-Lysine N-Methyltransferase*, Neurons*, Nuclear Proteins*, RNA Polymerase II*, Animals, Carrier Proteins, Histones, Mice, Mouse Embryonic Stem Cells (* major topic)
Topic: RNA modifications and cancer (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: EC | Horizon Europe | Excellent Science | HORIZON EUROPE European Research Council (ERC) (865094); Swiss National Science Foundation (183722, 180354, 186012)
Citations: not cited yet (Europe PMC); 79 references in the paper

Abstract

Chromatin modifications are essential for mammalian development, and their aberrant deposition is associated with human disease. While the mechanisms that deposit and remove histone modifications have been largely elucidated, their roles in regulating gene activity during cellular differentiation are yet to be fully understood. Here, we performed a deletion screen to identify stage-specific requirements of chromatin regulators during neuronal differentiation of mouse embryonic stem cells. We show that the H3K36me3 methyltransferase SETD2 is required for the establishment of neuronal gene expression during late stages of differentiation but is dispensable in mature neurons. Notably, this function is independent of its histone methyltransferase activity. Instead, SETD2 promotes interactions between the PAF1 complex and elongating RNA Pol II, suggesting a role in supporting efficient transcription of neuronal genes.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

FelixKrueger/TrimGalore

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c6528e54512e0e388a392d36475291bcbf0eb0dd, 27 June 2026
Languages: Rust (24), TypeScript (2), Shell (2), Python (1)
Size: 267 files, 29 scripts
Software Heritage: not archived
Found in: the resources table
Holds: README, license file, CITATION.cff, environment (Dockerfile), tests, continuous integration, documentation
Tools: Matplotlib (1 file), SAMtools (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
31 files

BaubecLab/Setd2

License: none: the authors keep all their rights
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 9d6f50ce61cf1d521f4f2889a56d8e3a2a03192f, 17 September 2025
Languages: Shell (3), R (2)
Size: 11 files, 5 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (chip_seq/mapping/conda_env_ngs.yaml, rnaseq/mapping_rnaseq_nfcore/conda_env_nextflow.yaml), 2 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: ggplot2 (2 files), limma (2 files), tidyverse (2 files), circlize (1 file), clusterProfiler (1 file), ComplexHeatmap (1 file), deepTools (1 file), edgeR (1 file), Nextflow (1 file), pheatmap (1 file), Plotly (1 file), reshape2 (1 file), SAMtools (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
6 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 34 scripts, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Other data links

Data availability

High-throughput sequencing data obtained from ChIP-seq, ATAC-seq, RNA-seq, and single-cell RNA-seq studies were deposited to NCBI Gene Expression Omnibus under the following accessions: GSE278955 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE278955) (RNA-seq), GSE278956 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE278956) (ChIP-seq), GSE278957 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE278957) (ATAC-seq), GSE321694 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE321694) (scRNA-seq). Scripts used to analyze genomics and proteomics data were deposited to https://github.com/BaubecLab/Setd2.

The source data of this paper are collected in the following database record: biostudies:S-SCDT-10_1038-S44318-026-00768-2 (https://www.ebi.ac.uk/biostudies/sourcedata/studies/S-SCDT-10_1038-S44318-026-00768-2).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 13 authors, 5 keywords, 10 MeSH terms, 2 funders, 78 references.

Cite

This paper

Ambrosi, C., Pfaendler, R., Eleftheriou, K., Butz, S., Recchia, D., Bao, X., Cardoso da Silva, R., Kupfer, N., Lagerwaard, I. M., Vlaming, H., Schmolka, N., Bhardwaj, V., & Baubec, T. (2026). The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation. The EMBO journal, 45(10), 3430-3443. https://doi.org/10.1038/s44318-026-00768-2

BibTeX

@article{ambrosi2026h3k36me3,
author = {Ambrosi, Christina and Pfaendler, Ramon and Eleftheriou, Kristeli and Butz, Stefan and Recchia, Davide and Bao, Xue and Cardoso da Silva, Richard and Kupfer, Niklas and Lagerwaard, Ilse M and Vlaming, Hanneke and Schmolka, Nina and Bhardwaj, Vivek and Baubec, Tuncay},
title = {{The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation}},
journal = {The EMBO journal},
year = {2026},
month = apr,
volume = {45},
number = {10},
pages = {3430--3443},
publisher = {Nature Publishing Group},
issn = {0261-4189},
doi = {10.1038/s44318-026-00768-2},
url = {https://doi.org/10.1038/s44318-026-00768-2},
pmid = {41963556},
pmcid = {PMC13187324}
}

RIS

TY - JOUR
AU - Ambrosi, Christina
AU - Pfaendler, Ramon
AU - Eleftheriou, Kristeli
AU - Butz, Stefan
AU - Recchia, Davide
AU - Bao, Xue
AU - Cardoso da Silva, Richard
AU - Kupfer, Niklas
AU - Lagerwaard, Ilse M
AU - Vlaming, Hanneke
AU - Schmolka, Nina
AU - Bhardwaj, Vivek
AU - Baubec, Tuncay
TI - The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation
T2 - The EMBO journal
J2 - EMBO J
PY - 2026
DA - 2026/04/10
VL - 45
IS - 10
SP - 3430
EP - 3443
SN - 0261-4189
PB - Nature Publishing Group
DO - 10.1038/s44318-026-00768-2
UR - https://doi.org/10.1038/s44318-026-00768-2
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s44318-026-00768-2",
"type": "article-journal",
"title": "The H3K36me3 methyltransferase SETD2 contributes to PAF1C interactions with RNA Pol II and is required for neuronal differentiation",
"container-title": "The EMBO journal",
"author": [
{
"family": "Ambrosi",
"given": "Christina"
},
{
"family": "Pfaendler",
"given": "Ramon"
},
{
"family": "Eleftheriou",
"given": "Kristeli"
},
{
"family": "Butz",
"given": "Stefan"
},
{
"family": "Recchia",
"given": "Davide"
},
{
"family": "Bao",
"given": "Xue"
},
{
"family": "Cardoso da Silva",
"given": "Richard"
},
{
"family": "Kupfer",
"given": "Niklas"
},
{
"family": "Lagerwaard",
"given": "Ilse M"
},
{
"family": "Vlaming",
"given": "Hanneke"
},
{
"family": "Schmolka",
"given": "Nina"
},
{
"family": "Bhardwaj",
"given": "Vivek"
},
{
"family": "Baubec",
"given": "Tuncay"
}
],
"container-title-short": "EMBO J",
"volume": "45",
"issue": "10",
"page": "3430-3443",
"DOI": "10.1038/s44318-026-00768-2",
"PMID": "41963556",
"PMCID": "PMC13187324",
"ISSN": "0261-4189",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s44318-026-00768-2",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s11689-026-09713-0 [code]
DRP1 mutations associated with EMPF1 encephalopathy perturb the transcriptional profile and maturation of cortical neurons.
Journal: Journal of neurodevelopmental disorders
In common: edgeR, limma, circlize, 8 other tools, 4 references
[2] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: SAMtools, edgeR, limma, 9 other tools, 1 reference
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: SAMtools, edgeR, limma, 9 other tools, mouse
[4] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: edgeR, limma, circlize, 8 other tools, 2 references
[5] doi:10.1038/s41467-026-69944-6 [code]
Multi-modal dissection of cell-type specific TDP-43 pathology in the motor cortex.
Journal: Nature communications
In common: deepTools, SAMtools, circlize, 7 other tools, 2 references
[6] doi:10.1038/s41467-026-74753-y [code]
A human-specific microRNA controls the timing of excitatory synaptogenesis.
Journal: Nature communications
In common: SAMtools, edgeR, limma, 7 other tools, 2 references
[7] doi:10.1016/j.xcrm.2026.102682 [code]
TET CpG sequence-context-specific DNA demethylation shapes progression of IDH-mutant gliomas.
Journal: Cell reports. Medicine
In common: Nextflow, edgeR, limma, 7 other tools, 1 reference
[8] doi:10.1038/s41586-026-10512-9 [code]
Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.
Journal: Nature
In common: SAMtools, edgeR, circlize, 7 other tools, mouse, 2 references
[9] doi:10.1111/adb.70179 [code]
Transcriptional Response to Chronic Long-Access Fentanyl Self-Administration in Rat Habenula and Amygdala.
Journal: Addiction biology
In common: Nextflow, edgeR, limma, 7 other tools
[10] doi:10.1016/j.cpblue.2026.100007 [code]
An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.
Journal: Cell press blue
In common: SAMtools, edgeR, limma, 8 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.