OSCR

Determinants of functional burden pleiotropy and gene dosage responses across human traits.

Code ↔ Paper

21 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 21 matches
  1. [1] § Results › Between-trait CNV burden correlations and comparison to other variant classes ↔ Figure_Generation/Figure5.Rmd, lines 7–89 · score 0.98 · forced vital capacity, visuospatial working memory, squares regression, single nucleotide polymorphism, edge thickness, function SNV burden
  2. [2] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure5.Rmd, lines 7–89 · score 0.89 · forced vital capacity, visuospatial working memory, single nucleotide polymorphism, Townsend deprivation, educational attainment, fluid intelligence
  3. [3] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Figure_Generation/Figure4.Rmd, lines 7–28 · score 0.84 · top decile LOEUF, Functional pleiotropy normalized, Constraint gene percentages, intolerant genes, body cell, pleiotropy compared
  4. [4] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 10–43 · score 0.84 · top decile LOEUF, Functional pleiotropy normalized, Constraint gene percentages, intolerant genes, pleiotropy compared, body cell
  5. [5] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Fig2.Rmd, lines 41–161 · score 0.76 · medulla oblongata, locus coeruleus, substantia nigra, activity, pons, cortex
  6. [6] § Methods › Genetic constraint metrics ↔ Figure_Generation/Figure4.Rmd, lines 168–211 · score 0.76 · pHaplo, pTriplo, gene length, CDS, het, missense
  7. [7] § Methods › Genetic constraint metrics ↔ Fig3.Rmd, lines 160–227 · score 0.75 · pHaplo, pTriplo, gene length, CDS, het, missense
  8. [8] § Results › Gene dosage effects across whole-body and whole-brain cell types ↔ Figure_Generation/Figure5.Rmd, lines 239–293 · score 0.72 · HbA1c, Standing Height, Heart Rate, trait categories, BMD, HDL
  9. [9] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Fig2.Rmd, lines 41–161 · score 0.69 · HbA1c, standing height, heart rate, BMD, HDL, FVC
  10. [10] § Methods › CNV- burden correlations ↔ Functional_Burden_Association_Analysis/FunBurd_Multitrait_EffectSizes_Computation_Script_On_ToyDatasets.ipynb, lines 1–75 · score 0.66 · logistic regression, binary trait, continuous trait, covariates, burden, CNV
  11. [11] § Results › Monotonic versus non-monotonic gene dosage responses across traits ↔ Figure_Generation/Figure6.Rmd, lines 231–248 · score 0.62 · Wilcoxon rank sum, gene dosage responses, brain traits, monotonic, Box, Figure 6
  12. [12] § Methods › Replication in the All of Us cohort ↔ Figure_Generation/Figure2.Rmd, lines 600–696 · score 0.62 · Standing Height, Heart Rate, AoU, cohorts, Platelet, UKBB
  13. [13] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure2.Rmd, lines 52–93 · score 0.59 · locus coeruleus, substantia nigra, pons, cortex, brain, tissue
  14. [14] § Methods › Replication in the All of Us cohort ↔ Figure_Generation/Figure3.Rmd, lines 514–572 · score 0.58 · Standing Height, Heart Rate, AoU, Platelet, UKBB, BMI
  15. [15] § Methods › Normative constraint modeling ↔ Figure_Generation/Figure4.Rmd, lines 231–251 · score 0.57 · normative model, constraint fraction, functional pleiotropy, genetic constraint
  16. [16] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure3.Rmd, lines 52–96 · score 0.57 · blood assays, mental health, physical, reproductive, cognitive, tissues
  17. [17] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 511–565 · score 0.53 · Wilcox Ranksum, brain gene, normative, modeling, pleiotropy, FDR
  18. [18] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 10–43 · score 0.52 · LOEUF top decile, constraint metric, genetic constraint, constrained genes, FDR corrected, brain gene
  19. [19] § Methods › CNV- burden correlations ↔ Functional_Burden_Association_Analysis/FunBurd_Multitrait_EffectSizes_Computation_Script_On_ToyDatasets.ipynb, lines 1–75 · score 0.52 · logistic regression, binary trait, burden, CNV
  20. [20] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Figure_Generation/Figure4.Rmd, lines 7–28 · score 0.52 · LOEUF top decile, constraint metric, genetic constraint, FDR corrected, brain gene, constrained genes
  21. [21] § Methods › Permutation preserving Jaccard distance (P-Jaccard) ↔ Permutation_PreservingJaccardDistance_P_Jaccard/PJaccard_Computation.ipynb, lines 104–144 · score 0.50 · Jaccard distance, permutations, matrix, correlations

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R Markdown · 377 lines · 18 KB · MIT · 4 matches

  1. ---
  2. title: "Fig4"
  3. author: "Kuldeep Kumar"
  4. output: github_document
  5. ---
  6. ## Fig. 4: Dissecting pleiotropy, gene function, and genetic constraint.
  7. #### -- Figure legend -- ####
  8. **Legend:** (A) Correlation between the fraction of constraint genes (LOEUF top-decile) within a gene set and the number of traits showing significant associations with that gene set for deletions and duplications. Each data point is a gene set. Y-axis: the percentage of significantly associated traits for a variant (functional pleiotropy); X-axis: the percentage of top decile LOEUF genes within the gene set. (B) Box plots show the distribution of constraint gene percentages for all Brain and Non-Brain gene sets (172). Each point represents a gene set, and asterisks (*) indicate statistically significant differences between groups. (C) Shows the same information on correlations in panel (A) for different constraint metrics. (D) Functional pleiotropy, normalized by the fraction of genetic constraint at the tissue level, is shown for 27 brain and 33 non-brain gene sets – deletions are presented on the left and duplications on the right. X-axis: represents the proportion of intolerant genes for the different gene sets. Y-axis: functional pleiotropy normalized by genetic constraint (centile). I.e., the 50th centile shows median functional pleiotropy computed across 100 randomly sampled gene sets. Circles and triangles represent brain and non-brain gene sets, respectively. Gray shaded ribbons indicating 25th–75th (ribbon 1), 10th–90th (ribbon 2), and 5th–95th (ribbon 3) centiles. This is followed by violin plots showing the distribution of normalized functional pleiotropy across brain and non-brain traits. Orange and green asterisks demonstrate significantly (FDR-corrected, q < 0.05) increased functional pleiotropy compared to what is expected for a gene set with a comparable fraction of genetic constraint. The grey star shows a significant difference between the brain and non-brain gene sets, normalized functional pleiotropy. (E) The same analyses at the whole-body cell type level, including 7 brain and 74 non-brain gene sets.
  9. #### Libraries ####
  10. ```{r setup, message=FALSE, warning=FALSE}
  11. library(ggplot2)
  12. library(ggprism)
  13. library(ggpubr)
  14. library(dplyr)
  15. library(tidyr)
  16. library(patchwork)
  17. library(knitr)
  18. library(openxlsx)
  19. library(svglite)
  20. # Define Core Colors and Sizes
  21. brain_colorcode <- "#d95f02"
  22. nonbrain_colorcode <- "#66a61e"
  23. in_base_size_ggprism <- 18
  24. ```
  25. #### Panel A: Correlation between constraint and Del/Dup/GWAS functional pleiotropy
  26. Description: Scatter plots evaluating the relationship between the fraction of constraint genes (LOEUF) and the percentage of significantly associated traits. ####
  27. ##### Panel A Data #####
  28. ```{r, fig.width=7, fig.height=3 ,dpi=100 }
  29. ##### Panel A Data #####
  30. # 1. Load the Pre-computed Correlation & P-Jaccard Statistics
  31. load("df_constraint_pleiotropy_corr.RData")
  32. stats_panel_a <- df_constraint_pleiotropy_corr %>%
  33. # Filter for the specific constraint measure used in Panel A
  34. filter(constraint == "measure_LEOUF_topdecile") %>%
  35. # Extract BOTH the correlation and the adjusted Jaccard p-value
  36. select(stat, cor, adj_pvalue_Jaccard)
  37. # 2. Load the Geneset-level plotting coordinates
  38. load("df_geneset_level_stats.RData")
  39. df_fp <- df_geneset_level_stats %>%
  40. mutate(
  41. # Convert constraint to percentage for the X-axis
  42. constraint_pct = constraint_measure_LEOUF_topdecile * 100
  43. ) %>%
  44. select(Geneset, Geneset_Type, Geneset_Cat, constraint_pct,
  45. functional_pleiotropy_deletion, functional_pleiotropy_duplication, functional_pleiotropy_GWASenrichment)
  46. kable(head(df_fp), format = "markdown")
  47. kable(stats_panel_a, align = "c", caption = "Exact Stats Extracted for Panel A")
  48. ```
  49. ##### Panel A Figure #####
  50. ```{r, fig.width=9, fig.height=3.5, dpi=300}
  51. # Function to plot scatter and map the exact 'cor' and 'p-jaccard' from the stats dataframe
  52. plot_pleio_scatter_jaccard <- function(data, y_var, color_hex, stat_label, stats_df) {
  53. # Map BOTH correlation and p-value from the stats dataframe
  54. row_match <- stats_df %>% filter(stat == stat_label)
  55. cor_val <- row_match$cor
  56. pval_adj <- row_match$adj_pvalue_Jaccard
  57. # Format the label (add asterisk if FDR < 0.05)
  58. cor_label <- paste0("r=", round(cor_val, 2), ifelse(pval_adj < 0.05, "*", ""))
  59. ggplot(data, aes(x = constraint_pct, y = .data[[y_var]])) +
  60. geom_point(color = color_hex, size = 2, alpha = 1) +
  61. # Add the extracted correlation label to the plot
  62. annotate("text", x = Inf, y = Inf, hjust = 1.1, vjust = 1.5, label = cor_label, size = 6, fontface="bold") +
  63. geom_smooth(method = "lm", formula=y ~ x, se = FALSE, linewidth = 1.5, fullrange = TRUE, color = "black") +
  64. scale_y_continuous(limits = c(0, 75)) +
  65. theme_prism(axis_text_angle = 90, base_size = 14) +
  66. labs(x = "% constraint genes", y = "% functional pleiotropy") +
  67. theme(plot.title = element_text(hjust = 0.5, face = "bold"))
  68. }
  69. # Generate the three plots using their exact stat labels
  70. p_scatter_del <- plot_pleio_scatter_jaccard(
  71. df_fp, "functional_pleiotropy_deletion", "red3",
  72. "functional_burden_pleiotropy_deletion", stats_panel_a
  73. )
  74. p_scatter_dup <- plot_pleio_scatter_jaccard(
  75. df_fp, "functional_pleiotropy_duplication", "royalblue",
  76. "functional_burden_pleiotropy_duplication", stats_panel_a
  77. ) + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
  78. p_scatter_gwas <- plot_pleio_scatter_jaccard(
  79. df_fp, "functional_pleiotropy_GWASenrichment", "darkgreen",
  80. "functional_pleiotropy_GWASenrichment", stats_panel_a
  81. ) + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
  82. # Layout the final figure
  83. p_scatter_Fig4A <- (p_scatter_del + p_scatter_dup + p_scatter_gwas) + plot_layout(widths = c(1, 1, 1))
  84. print(p_scatter_Fig4A)
  85. ```
  86. ##### Panel A Stats #####
  87. ```{r, fig.width=18, fig.height=7 ,dpi=800 }
  88. kable(head(stats_panel_a), format = "markdown")
  89. ```
  90. #### Panel B: Distribution of constraint across Brain and Non-Brain gene sets
  91. Description: Box plots comparing the percentage of constraint genes between functional categories.
  92. ##### Panel B Data #####
  93. ```{r, fig.width=18, fig.height=7 ,dpi=800 }
  94. df_box <- df_fp %>%
  95. select(Geneset, Geneset_Cat, constraint_pct) %>%
  96. mutate(Geneset_Cat = factor(Geneset_Cat, levels = c("Brain", "NonBrain")))
  97. kable(head(df_box), format = "markdown")
  98. ```
  99. ##### Panel B Figure #####
  100. ```{r, fig.width=7, fig.height=4, dpi=400}
  101. p_box_Fig4B <- ggplot(df_box, aes(x = Geneset_Cat, y = constraint_pct, fill = Geneset_Cat)) +
  102. # Boxplot with black outline
  103. geom_boxplot(outlier.shape = NA, width = 0.5, alpha = 0.5, linewidth = 0.8, color = "black") +
  104. # Jitter points: shape 21 enables 'fill' aesthetic
  105. geom_jitter(aes(color=Geneset_Cat,shape = Geneset_Cat), stroke = 0.4,
  106. width = 0.15, size = 2, alpha = 0.9) +
  107. scale_fill_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
  108. scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
  109. scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
  110. theme_prism(base_size = in_base_size_ggprism) +
  111. theme(
  112. legend.position = "none",
  113. axis.text.x = element_text(size = 14, face = "bold"),
  114. axis.text.y = element_text(size = 14, face = "bold"),
  115. axis.title.y = element_text(size = 16, face = "bold")
  116. ) +
  117. labs(x = NULL, y = "% constraint genes") +
  118. coord_flip()
  119. print(p_box_Fig4B)
  120. ```
  121. ##### Panel B Stats #####
  122. ```{r, fig.width=9, fig.height=5, dpi=400}
  123. res_wilcox_B <- wilcox.test(
  124. df_box %>% filter(Geneset_Cat == "Brain") %>% pull(constraint_pct),
  125. df_box %>% filter(Geneset_Cat == "NonBrain") %>% pull(constraint_pct)
  126. )
  127. stats_panel_b <- data.frame(
  128. Test = "Wilcoxon Rank Sum Test",
  129. Comparison = "Brain vs Non-Brain Constraint (%)",
  130. W_Statistic = res_wilcox_B$statistic,
  131. P_Value = format(res_wilcox_B$p.value, digits = 4)
  132. )
  133. kable(stats_panel_b, align = "c", caption = "Statistical Comparison of Constraint Percentages")
  134. ```
  135. #### Panel C: Constraint vs Functional pleiotropy across metrics
  136. Description: Evaluates functional pleiotropy correlations across seven different established metrics of genetic constraint.
  137. ##### Panel C Data & Stats #####
  138. ```{r, fig.width=5, fig.height=9, dpi=400}
  139. load("df_constraint_pleiotropy_corr.RData")
  140. # Map the pre-computed dataframe directly to plotting coordinates
  141. df_metrics_cor <- df_constraint_pleiotropy_corr %>%
  142. mutate(
  143. # Map the raw constraint strings to clean metric labels
  144. Metric = case_when(
  145. grepl("LEOUF", constraint) ~ "LOEUF",
  146. grepl("missense_Z", constraint) ~ "missense Z",
  147. grepl("s_het", constraint) ~ "s-het",
  148. grepl("CDS", constraint) ~ "constraint CDS",
  149. grepl("pHaplo", constraint) ~ "pHaplo",
  150. grepl("pTriplo", constraint) ~ "pTriplo",
  151. grepl("GeneLenBP", constraint) ~ "gene length",
  152. TRUE ~ constraint
  153. ),
  154. # Map the stat column to Del, Dup, GWAS
  155. Type = case_when(
  156. grepl("deletion", stat) ~ "Del",
  157. grepl("duplication", stat) ~ "Dup",
  158. grepl("GWASenrichment", stat) ~ "GWAS",
  159. TRUE ~ stat
  160. ),
  161. # Assign significance using the Jaccard adjusted p-value directly
  162. sig = ifelse(adj_pvalue_Jaccard < 0.05, "q<0.05", "n.s.")
  163. ) %>%
  164. # Set factors for proper plot ordering
  165. mutate(
  166. Type = factor(Type, levels = c("Del", "Dup", "GWAS")),
  167. sig = factor(sig, levels = c("q<0.05", "n.s."))
  168. ) %>%
  169. select(Metric, Type, cor, pvalue_Jaccard, adj_pvalue_Jaccard, sig)
  170. # Order Y-axis metrics based on Deletion correlation strength
  171. metric_order <- df_metrics_cor %>% filter(Type == "Del") %>% arrange(cor) %>% pull(Metric)
  172. df_metrics_cor$Metric <- factor(df_metrics_cor$Metric, levels = metric_order)
  173. kable(head(df_metrics_cor), format = "markdown")
  174. ```
  175. ##### Panel C Figure #####
  176. ```{r, fig.width=6, fig.height=7, dpi=90}
  177. p_Fig4C <- ggplot(df_metrics_cor, aes(y = Metric, x = cor, group = Type)) +
  178. geom_point(aes(color = Type, fill = Type, shape = sig), size = 5, position = position_dodge(width = 0.6)) +
  179. scale_fill_manual(values = c("Del" = "red3", "Dup" = "royalblue", "GWAS" = "darkgreen")) +
  180. scale_color_manual(values = c("Del" = "red3", "Dup" = "royalblue", "GWAS" = "darkgreen")) +
  181. scale_shape_manual(values = c("q<0.05" = 21, "n.s." = 4)) +
  182. geom_vline(xintercept = 0, linetype = "longdash", color = "black") +
  183. geom_vline(xintercept = 0.4, linetype = "dotted", color = "black") +
  184. theme_prism(base_size = in_base_size_ggprism) +
  185. theme(legend.position = c(0.85, 0.2)) +
  186. labs(x = "Correlation", y = NULL) +
  187. coord_cartesian(xlim = c(0, 0.7))
  188. print(p_Fig4C)
  189. ```
  190. #### Panel D & E: Normative Null Models (Tissue & Cell Type)
  191. Description: Functional pleiotropy normalized by the fraction of genetic constraint, plotted against null distributions for Tissue (Fantom60) and Cell Type (HPA81).
  192. ##### Panel D & E Data #####
  193. ```{r, fig.width=6, fig.height=7, dpi=90}
  194. # 1. Extract and format the exact plotting coordinates from geneset level stats
  195. # No external null distribution data is required.
  196. df_normative <- df_geneset_level_stats %>%
  197. mutate(
  198. # Convert constraint fraction to percentage for the X-axis
  199. constraint_pct = constraint_measure_LEOUF_topdecile * 100
  200. ) %>%
  201. select(Geneset, Geneset_Type, Geneset_Cat, constraint_pct,
  202. normative_model_centile_del_pleiotropy, normative_model_centile_dup_pleiotropy)
  203. # Ensure Geneset_Cat is a factor for consistent color mapping
  204. df_normative$Geneset_Cat <- factor(df_normative$Geneset_Cat, levels = c("Brain", "NonBrain"))
  205. kable(head(df_normative), format = "markdown")
  206. ```
  207. ##### Panel D & E Figure #####
  208. ```{r, fig.width=18, fig.height=9, dpi=90}
  209. # --- 1. Compute Stats First (Needed for Dynamic Annotations) ---
  210. compute_normative_stats <- function(data, gset_type, metric_col, variant_label) {
  211. sub_df <- data %>% filter(Geneset_Type == gset_type)
  212. brain_vals <- sub_df %>% filter(Geneset_Cat == "Brain") %>% pull(.data[[metric_col]])
  213. nbrain_vals <- sub_df %>% filter(Geneset_Cat == "NonBrain") %>% pull(.data[[metric_col]])
  214. # Unified Wilcoxon Rank-sum Tests
  215. p_brain_50 <- wilcox.test(brain_vals, mu = 50,exact = FALSE)$p.value
  216. p_nbrain_50 <- wilcox.test(nbrain_vals, mu = 50,exact = FALSE)$p.value
  217. p_b_vs_nb <- wilcox.test(brain_vals, nbrain_vals,exact = FALSE)$p.value
  218. return(data.frame(Geneset_Type = gset_type, Variant = variant_label,
  219. P_Brain_vs_50 = p_brain_50, P_NonBrain_vs_50 = p_nbrain_50, P_Brain_vs_NonBrain = p_b_vs_nb))
  220. }
  221. # Run across all categories to get raw p-values
  222. df_stats_raw <- bind_rows(
  223. compute_normative_stats(df_normative, "Tissue_Fantom60", "normative_model_centile_del_pleiotropy", "Del"),
  224. compute_normative_stats(df_normative, "Tissue_Fantom60", "normative_model_centile_dup_pleiotropy", "Dup"),
  225. compute_normative_stats(df_normative, "SC_HPA81", "normative_model_centile_del_pleiotropy", "Del"),
  226. compute_normative_stats(df_normative, "SC_HPA81", "normative_model_centile_dup_pleiotropy", "Dup")
  227. )
  228. # 1. Pool all 8 tests for 'shift from 50%' and adjust FDR
  229. pooled_vs_50_pvals <- c(df_stats_raw$P_Brain_vs_50, df_stats_raw$P_NonBrain_vs_50)
  230. pooled_vs_50_fdr <- p.adjust(pooled_vs_50_pvals, method = "fdr")
  231. # 2. Pool all 4 tests for 'brain vs non-brain' and adjust FDR
  232. pooled_b_vs_nb_fdr <- p.adjust(df_stats_raw$P_Brain_vs_NonBrain, method = "fdr")
  233. # Map the correctly pooled FDR values back
  234. df_stats_DE_final <- df_stats_raw %>%
  235. mutate(
  236. # Indices 1:4 belong to Brain, indices 5:8 belong to NonBrain
  237. FDR_Brain_vs_50 = pooled_vs_50_fdr[1:4],
  238. FDR_NonBrain_vs_50 = pooled_vs_50_fdr[5:8],
  239. FDR_Brain_vs_NonBrain = pooled_b_vs_nb_fdr
  240. )
  241. # --- 2. Helper Function to Build Individual Components ---
  242. build_normative_plots <- function(data_obs, gset_type, var_y, variant_type, stats_df, y_label = NULL, hide_y = FALSE) {
  243. df_sub <- data_obs %>% filter(Geneset_Type == gset_type)
  244. plot_stats <- stats_df %>% filter(Geneset_Type == gset_type, Variant == variant_type)
  245. # Scatter Plot (Using manual annotations for the Centile bands instead of a null dataset)
  246. p_scatter <- ggplot(df_sub, aes(x = constraint_pct, y = .data[[var_y]])) +
  247. # Background Centile Ribbons
  248. annotate("rect", ymin = 5, ymax = 95, xmin = -Inf, xmax = Inf, fill = "grey80", alpha = 0.4) +
  249. annotate("rect", ymin = 10, ymax = 90, xmin = -Inf, xmax = Inf, fill = "grey70", alpha = 0.4) +
  250. annotate("rect", ymin = 25, ymax = 75, xmin = -Inf, xmax = Inf, fill = "grey60", alpha = 0.4) +
  251. geom_hline(yintercept = 50, color = "black", linewidth = 1) +
  252. # Scatter Points
  253. geom_point(aes(color = Geneset_Cat, shape = Geneset_Cat), size = 3) +
  254. scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
  255. scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
  256. scale_x_continuous(expand = c(0, 0), limits = c(0, 22), breaks = seq(0, 22, by = 5)) +
  257. scale_y_continuous(limits = c(0, 125), breaks = c(10, 25, 50, 75, 90)) +
  258. theme_prism(base_size = in_base_size_ggprism, base_line_size = in_base_size_ggprism/24) +
  259. theme(legend.position = "none") +
  260. labs(x = "% constraint genes", y = y_label)
  261. if (hide_y) {
  262. p_scatter <- p_scatter + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
  263. }
  264. # Violin Plot
  265. p_violin <- ggplot(df_sub, aes(x = Geneset_Cat, y = .data[[var_y]], color = Geneset_Cat,shape=Geneset_Cat)) +
  266. geom_hline(yintercept = 50, linetype = "dotted", color = "black", linewidth = 2) +
  267. geom_violin(fill = "white", alpha = 1, trim = TRUE, linewidth = 0.8) +
  268. geom_point(position = position_jitter(seed = 1, width = 0.2), size = 1.5, alpha = 0.8) +
  269. stat_summary(color = "black", fun.data = "mean_cl_boot", geom = "pointrange", size = 0.7) +
  270. scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
  271. scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
  272. scale_x_discrete(labels = c("Brain", "NonBrain")) +
  273. scale_y_continuous(limits = c(0, 125), breaks = c(10, 25, 50, 75, 90)) +
  274. theme_prism(base_size = in_base_size_ggprism, base_line_size = in_base_size_ggprism/24) +
  275. theme(legend.position = "none", axis.title.x = element_blank(), axis.title.y = element_blank(),
  276. axis.line.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
  277. # Add Significance Annotations dynamically based on the stats
  278. sig_size <- 10
  279. if (plot_stats$FDR_Brain_vs_50 < 0.05) {
  280. p_violin <- p_violin + annotate("text", x = 1, y = 105, label = "*", size = sig_size, color = brain_colorcode, fontface = "bold")
  281. }
  282. if (plot_stats$FDR_NonBrain_vs_50 < 0.05) {
  283. p_violin <- p_violin + annotate("text", x = 2, y = 105, label = "*", size = sig_size, color = nonbrain_colorcode, fontface = "bold")
  284. }
  285. if (plot_stats$FDR_Brain_vs_NonBrain < 0.05) {
  286. p_violin <- p_violin +
  287. annotate("segment", x = 1, xend = 2, y = 115, yend = 115, color = "grey50", linewidth = 1) +
  288. annotate("segment", x = 1, xend = 1, y = 110, yend = 115, color = "grey50", linewidth = 1) +
  289. annotate("segment", x = 2, xend = 2, y = 110, yend = 115, color = "grey50", linewidth = 1) +
  290. annotate("text", x = 1.5, y = 118, label = "*", size = sig_size, color = "grey50", fontface = "bold")
  291. }
  292. return(p_scatter + p_violin + plot_layout(widths = c(0.72, 0.28)))
  293. }
  294. # --- 3. Assemble Final Composite Panels ---
  295. y_label_text <- "functional pleiotropy\nnormalized by constraint (centiles)"
  296. pD_Del <- build_normative_plots(df_normative, "Tissue_Fantom60", "normative_model_centile_del_pleiotropy", "Del", df_stats_DE_final, y_label = y_label_text)
  297. pD_Dup <- build_normative_plots(df_normative, "Tissue_Fantom60", "normative_model_centile_dup_pleiotropy", "Dup", df_stats_DE_final, hide_y = TRUE)
  298. row_D <- pD_Del | pD_Dup
  299. pE_Del <- build_normative_plots(df_normative, "SC_HPA81", "normative_model_centile_del_pleiotropy", "Del", df_stats_DE_final, y_label = y_label_text)
  300. pE_Dup <- build_normative_plots(df_normative, "SC_HPA81", "normative_model_centile_dup_pleiotropy", "Dup", df_stats_DE_final, hide_y = TRUE)
  301. row_E <- pE_Del | pE_Dup
  302. p_single_stacked <- row_D / row_E
  303. print(p_single_stacked)
  304. ```
  305. ##### Panel D & E Stats #####
  306. ```{r, fig.width=6, fig.height=7, dpi=90}
  307. # --- Simplified and Robust Stats Calculation ---
  308. # Display the stats dataframe that was computed and utilized in the Figure chunk
  309. kable(df_stats_DE_final %>% select(Geneset_Type, Variant, starts_with("FDR"), starts_with("P")), align = "c", caption = "FDR-Adjusted Wilcoxon Tests for Normative Centiles (Panels D & E)")
  310. ```

Figure4.Rmd at commit 4415dfe, under MIT · at the source

Overview

  1. Centre de recherche Azrieli, CHU Sainte-Justine and University of Montréal, Montréal, QC Canada
  2. Department of Psychiatr, Harvard Medical School, Boston, MA USA
  3. Tommy Fuss Center for Neuropsychiatric Disease Research, Boston Children’s Hospital, Boston, MA USA
  4. Department of Biomedical and Health Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA USA
  5. Lifespan Brain Institute, Children’s Hospital of Philadelphia, and Penn Medicine, Philadelphia, PA USA
  6. Department of Genetics, University of Pennsylvania, Philadelphia, PA USA
  7. The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, ON Canada
  8. Department of Psychiatry, University of California San Diego, La Jolla, CA USA
  9. Lady Davis Institute for Medical Research, Jewish General Hospital, Montreal, QC Canada
  10. Gerald Bronfman Department of Oncology, Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, QC Canada
  11. Department of Molecular Genetics, University of Toronto, Toronto, ON Canada
  12. Mila – Quebec AI Institute, University of Montréal, Montreal, QC Canada
Journal: Nature communications, volume 17, issue 1, article 9798
Dates: received 6 November 2025; accepted 31 July 2026; published online 14 August 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41467-026-76676-0 · PMID 42736290 · PMCID PMC13575219 · OpenAlex W7203458115
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), cellular / molecular (subfield)
Methods: Statistics, Machine learning, Preprocessing, Connectivity
Keywords: Rare variants, Functional genomics, Structural variation
MeSH: Gene Dosage*, Genetic Pleiotropy*, DNA Copy Number Variations, Genome-Wide Association Study, Humans, Phenotype, Polymorphism, Single Nucleotide (* major topic)
Topic: Growth Hormone and Insulin-like Growth Factors (Endocrinology, Diabetes and Metabolism, Medicine), according to OpenAlex
Citations: cited by 1 paper (Europe PMC); 91 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 21 matches between paragraphs and lines of code.

martineaujeanlouis.github.io/mind-genesparallelcnv

License: none: the authors keep all their rights
State: the link is dead, verified on 27 September 2026
Evidence: found in the paper
Software Heritage: not checked
Found in: the text, “Genotyping and CNV calling”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link is dead (HTTP 404)
  • 27 September 2026: the link is dead (HTTP 404)

linnarsson-lab/adult-human-brain

License: BSD-2-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 2b5aa12dffb8cc1bcb3979c409cb6ac907546b2e, 5 June 2026
Languages: Python (96), Jupyter (17)
Size: 130 files, 113 scripts
Software Heritage: archived
Found in: “Data availability”
Holds: README, license file, environment (setup.cfg, setup.py), 17 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (74 files), SciPy (35 files), Matplotlib (34 files), scikit-learn (26 files), pandas (12 files), NetworkX (8 files), seaborn (7 files), Numba (6 files), igraph (3 files), statsmodels (1 file), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
97 files

SayehKazem/FunBurd

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 4415dfed4e8b276b374f1384041f310332fd34b8, 19 August 2026
Languages: R (5), Jupyter (2), Python (1), Shell (1)
Size: 318 files, 9 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, CITATION.cff, 7 notebooks
Not found: environment file, tests, continuous integration, documentation
Tools: ggplot2 (5 files), ggpubr (5 files), patchwork (5 files), tidyverse (5 files), NumPy (3 files), pandas (3 files), statsmodels (3 files), circlize (2 files), ComplexHeatmap (2 files), BrainSMASH (1 file), igraph (1 file), Matplotlib (1 file), NiBabel (1 file), reshape2 (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
11 files

Zenodo 17038354

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: ggplot2 (4 files), ggpubr (4 files), tidyverse (4 files), reshape2 (2 files), circlize (1 file), ComplexHeatmap (1 file), data.table (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
6 files
At the source:

funburd.readthedocs.io

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: the link answers
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41467-026-76676-0.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 5 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 108 scripts, each with its path and the digest of its content;
  • 21 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41467-026-76676-0.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Funding: added Fondation Brain Canada; Compute Canada; Canada First Research Excellence Fund; Institut de Valorisation des Données; National Institutes of Health: 1u01mh119690-01; Canadian Institutes of Health Research: CIHR_400528

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 20 authors, 3 keywords, 7 MeSH terms, 86 references.

Cite

This paper

Kazem, S., Kumar, K., Yang, J., Benitiere, F., Huguet, G., Mollon, J., Renne, T., Schultz, L. M., Knowles, E. E. M., Engchuan, W., Shanta, O., Thiruvahindrapuram, B., MacDonald, J. R., Greenwood, C. M. T., Scherer, S. W., Almasy, L., Sebat, J., Glahn, D. C., Dumas, G., & Jacquemont, S. (2026). Determinants of functional burden pleiotropy and gene dosage responses across human traits. Nature communications, 17(1), 9798. https://doi.org/10.1038/s41467-026-76676-0

BibTeX

@article{kazem2026determinants,
author = {Kazem, Sayeh and Kumar, Kuldeep and Yang, Jane and Benitiere, Florian and Huguet, Guillaume and Mollon, Josephine and Renne, Thomas and Schultz, Laura M and Knowles, Emma E M and Engchuan, Worrawat and Shanta, Omar and Thiruvahindrapuram, Bhooma and MacDonald, Jeffrey R and Greenwood, Celia M T and Scherer, Stephen W and Almasy, Laura and Sebat, Jonathan and Glahn, David C and Dumas, Guillaume and Jacquemont, Sébastien},
title = {{Determinants of functional burden pleiotropy and gene dosage responses across human traits}},
journal = {Nature communications},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {9798},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-76676-0},
url = {https://doi.org/10.1038/s41467-026-76676-0},
pmid = {42736290},
pmcid = {PMC13575219}
}

RIS

TY - JOUR
AU - Kazem, Sayeh
AU - Kumar, Kuldeep
AU - Yang, Jane
AU - Benitiere, Florian
AU - Huguet, Guillaume
AU - Mollon, Josephine
AU - Renne, Thomas
AU - Schultz, Laura M
AU - Knowles, Emma E M
AU - Engchuan, Worrawat
AU - Shanta, Omar
AU - Thiruvahindrapuram, Bhooma
AU - MacDonald, Jeffrey R
AU - Greenwood, Celia M T
AU - Scherer, Stephen W
AU - Almasy, Laura
AU - Sebat, Jonathan
AU - Glahn, David C
AU - Dumas, Guillaume
AU - Jacquemont, Sébastien
TI - Determinants of functional burden pleiotropy and gene dosage responses across human traits
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/08/14
VL - 17
IS - 1
SP - 9798
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-76676-0
UR - https://doi.org/10.1038/s41467-026-76676-0
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-76676-0",
"type": "article-journal",
"title": "Determinants of functional burden pleiotropy and gene dosage responses across human traits",
"container-title": "Nature communications",
"author": [
{
"family": "Kazem",
"given": "Sayeh"
},
{
"family": "Kumar",
"given": "Kuldeep"
},
{
"family": "Yang",
"given": "Jane"
},
{
"family": "Benitiere",
"given": "Florian"
},
{
"family": "Huguet",
"given": "Guillaume"
},
{
"family": "Mollon",
"given": "Josephine"
},
{
"family": "Renne",
"given": "Thomas"
},
{
"family": "Schultz",
"given": "Laura M"
},
{
"family": "Knowles",
"given": "Emma E M"
},
{
"family": "Engchuan",
"given": "Worrawat"
},
{
"family": "Shanta",
"given": "Omar"
},
{
"family": "Thiruvahindrapuram",
"given": "Bhooma"
},
{
"family": "MacDonald",
"given": "Jeffrey R"
},
{
"family": "Greenwood",
"given": "Celia M T"
},
{
"family": "Scherer",
"given": "Stephen W"
},
{
"family": "Almasy",
"given": "Laura"
},
{
"family": "Sebat",
"given": "Jonathan"
},
{
"family": "Glahn",
"given": "David C"
},
{
"family": "Dumas",
"given": "Guillaume"
},
{
"family": "Jacquemont",
"given": "Sébastien"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "9798",
"DOI": "10.1038/s41467-026-76676-0",
"PMID": "42736290",
"PMCID": "PMC13575219",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-76676-0",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.21203/rs.3.rs-9246968/v1 [code]
Copy number variants reveal divergent genetic and diagnostic cortical signatures across psychiatric disorders
Journal: Research Square (preprint)
In common: circlize, ComplexHeatmap, ggpubr, 4 other tools, genetics / omics, 6 references, 6 authors
[2] doi:10.1038/s41586-026-10515-6 [code]
An X-linked long non-coding RNA, PTCHD1-AS, and the core features of autism.
Journal: Nature
In common: pandas, SciPy, Matplotlib, 1 other tool, genetics / omics, cellular / molecular, 4 authors
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: UMAP, igraph, circlize, 16 other tools, cellular / molecular
[4] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: UMAP, igraph, circlize, 15 other tools, genetics / omics, cellular / molecular
[5] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: UMAP, igraph, circlize, 14 other tools, genetics / omics, cellular / molecular, 1 reference
[6] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: UMAP, igraph, circlize, 14 other tools, genetics / omics, cellular / molecular
[7] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: UMAP, igraph, circlize, 14 other tools, genetics / omics
[8] doi:10.1038/s41380-026-03571-x [code]
Convergent coexpression reveals shared biological mechanisms underlying common and rare variant risk in six neuropsychiatric disorders.
Journal: Molecular psychiatry
In common: igraph, reshape2, data.table, 8 other tools, genetics / omics, cellular / molecular, 5 references
[9] doi:10.1038/s41586-026-10735-w [code]
Distributed control circuits across a brain-and-cord connectome.
Journal: Nature
In common: UMAP, igraph, circlize, 14 other tools
[10] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: circlize, Numba, ComplexHeatmap, 13 other tools, genetics / omics, cellular / molecular, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.