Determinants of functional burden pleiotropy and gene dosage responses across human traits.
The 21 matches
- [1] § Results › Between-trait CNV burden correlations and comparison to other variant classes ↔ Figure_Generation/Figure5.Rmd, lines 7–89 · score 0.98 · forced vital capacity, visuospatial working memory, squares regression, single nucleotide polymorphism, edge thickness, function SNV burden
- [2] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure5.Rmd, lines 7–89 · score 0.89 · forced vital capacity, visuospatial working memory, single nucleotide polymorphism, Townsend deprivation, educational attainment, fluid intelligence
- [3] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Figure_Generation/Figure4.Rmd, lines 7–28 · score 0.84 · top decile LOEUF, Functional pleiotropy normalized, Constraint gene percentages, intolerant genes, body cell, pleiotropy compared
- [4] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 10–43 · score 0.84 · top decile LOEUF, Functional pleiotropy normalized, Constraint gene percentages, intolerant genes, pleiotropy compared, body cell
- [5] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Fig2.Rmd, lines 41–161 · score 0.76 · medulla oblongata, locus coeruleus, substantia nigra, activity, pons, cortex
- [6] § Methods › Genetic constraint metrics ↔ Figure_Generation/Figure4.Rmd, lines 168–211 · score 0.76 · pHaplo, pTriplo, gene length, CDS, het, missense
- [7] § Methods › Genetic constraint metrics ↔ Fig3.Rmd, lines 160–227 · score 0.75 · pHaplo, pTriplo, gene length, CDS, het, missense
- [8] § Results › Gene dosage effects across whole-body and whole-brain cell types ↔ Figure_Generation/Figure5.Rmd, lines 239–293 · score 0.72 · HbA1c, Standing Height, Heart Rate, trait categories, BMD, HDL
- [9] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Fig2.Rmd, lines 41–161 · score 0.69 · HbA1c, standing height, heart rate, BMD, HDL, FVC
- [10] § Methods › CNV- burden correlations ↔ Functional_Burden_Association_Analysis/FunBurd_Multitrait_EffectSizes_Computation_Script_On_ToyDatasets.ipynb, lines 1–75 · score 0.66 · logistic regression, binary trait, continuous trait, covariates, burden, CNV
- [11] § Results › Monotonic versus non-monotonic gene dosage responses across traits ↔ Figure_Generation/Figure6.Rmd, lines 231–248 · score 0.62 · Wilcoxon rank sum, gene dosage responses, brain traits, monotonic, Box, Figure 6
- [12] § Methods › Replication in the All of Us cohort ↔ Figure_Generation/Figure2.Rmd, lines 600–696 · score 0.62 · Standing Height, Heart Rate, AoU, cohorts, Platelet, UKBB
- [13] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure2.Rmd, lines 52–93 · score 0.59 · locus coeruleus, substantia nigra, pons, cortex, brain, tissue
- [14] § Methods › Replication in the All of Us cohort ↔ Figure_Generation/Figure3.Rmd, lines 514–572 · score 0.58 · Standing Height, Heart Rate, AoU, Platelet, UKBB, BMI
- [15] § Methods › Normative constraint modeling ↔ Figure_Generation/Figure4.Rmd, lines 231–251 · score 0.57 · normative model, constraint fraction, functional pleiotropy, genetic constraint
- [16] § Results › Functional burden pleiotropy is higher for genes assigned to brain tissue ↔ Figure_Generation/Figure3.Rmd, lines 52–96 · score 0.57 · blood assays, mental health, physical, reproductive, cognitive, tissues
- [17] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 511–565 · score 0.53 · Wilcox Ranksum, brain gene, normative, modeling, pleiotropy, FDR
- [18] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Fig3.Rmd, lines 10–43 · score 0.52 · LOEUF top decile, constraint metric, genetic constraint, constrained genes, FDR corrected, brain gene
- [19] § Methods › CNV- burden correlations ↔ Functional_Burden_Association_Analysis/FunBurd_Multitrait_EffectSizes_Computation_Script_On_ToyDatasets.ipynb, lines 1–75 · score 0.52 · logistic regression, binary trait, burden, CNV
- [20] § Results › Dissecting functional burden pleiotropy, gene function, and genetic constraint ↔ Figure_Generation/Figure4.Rmd, lines 7–28 · score 0.52 · LOEUF top decile, constraint metric, genetic constraint, FDR corrected, brain gene, constrained genes
- [21] § Methods › Permutation preserving Jaccard distance (P-Jaccard) ↔ Permutation_PreservingJaccardDistance_P_Jaccard/PJaccard_Computation.ipynb, lines 104–144 · score 0.50 · Jaccard distance, permutations, matrix, correlations
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 377 lines · 18 KB · MIT · 4 matches
- ---
- title: "Fig4"
- author: "Kuldeep Kumar"
- output: github_document
- ---
- ## Fig. 4: Dissecting pleiotropy, gene function, and genetic constraint.
- #### -- Figure legend -- ####
- **Legend:** (A) Correlation between the fraction of constraint genes (LOEUF top-decile) within a gene set and the number of traits showing significant associations with that gene set for deletions and duplications. Each data point is a gene set. Y-axis: the percentage of significantly associated traits for a variant (functional pleiotropy); X-axis: the percentage of top decile LOEUF genes within the gene set. (B) Box plots show the distribution of constraint gene percentages for all Brain and Non-Brain gene sets (172). Each point represents a gene set, and asterisks (*) indicate statistically significant differences between groups. (C) Shows the same information on correlations in panel (A) for different constraint metrics. (D) Functional pleiotropy, normalized by the fraction of genetic constraint at the tissue level, is shown for 27 brain and 33 non-brain gene sets – deletions are presented on the left and duplications on the right. X-axis: represents the proportion of intolerant genes for the different gene sets. Y-axis: functional pleiotropy normalized by genetic constraint (centile). I.e., the 50th centile shows median functional pleiotropy computed across 100 randomly sampled gene sets. Circles and triangles represent brain and non-brain gene sets, respectively. Gray shaded ribbons indicating 25th–75th (ribbon 1), 10th–90th (ribbon 2), and 5th–95th (ribbon 3) centiles. This is followed by violin plots showing the distribution of normalized functional pleiotropy across brain and non-brain traits. Orange and green asterisks demonstrate significantly (FDR-corrected, q < 0.05) increased functional pleiotropy compared to what is expected for a gene set with a comparable fraction of genetic constraint. The grey star shows a significant difference between the brain and non-brain gene sets, normalized functional pleiotropy. (E) The same analyses at the whole-body cell type level, including 7 brain and 74 non-brain gene sets.
- #### Libraries ####
- ```{r setup, message=FALSE, warning=FALSE}
- library(ggplot2)
- library(ggprism)
- library(ggpubr)
- library(dplyr)
- library(tidyr)
- library(patchwork)
- library(knitr)
- library(openxlsx)
- library(svglite)
- # Define Core Colors and Sizes
- brain_colorcode <- "#d95f02"
- nonbrain_colorcode <- "#66a61e"
- in_base_size_ggprism <- 18
- ```
- #### Panel A: Correlation between constraint and Del/Dup/GWAS functional pleiotropy
- Description: Scatter plots evaluating the relationship between the fraction of constraint genes (LOEUF) and the percentage of significantly associated traits. ####
- ##### Panel A Data #####
- ```{r, fig.width=7, fig.height=3 ,dpi=100 }
- ##### Panel A Data #####
- # 1. Load the Pre-computed Correlation & P-Jaccard Statistics
- load("df_constraint_pleiotropy_corr.RData")
- stats_panel_a <- df_constraint_pleiotropy_corr %>%
- # Filter for the specific constraint measure used in Panel A
- filter(constraint == "measure_LEOUF_topdecile") %>%
- # Extract BOTH the correlation and the adjusted Jaccard p-value
- select(stat, cor, adj_pvalue_Jaccard)
- # 2. Load the Geneset-level plotting coordinates
- load("df_geneset_level_stats.RData")
- df_fp <- df_geneset_level_stats %>%
- mutate(
- # Convert constraint to percentage for the X-axis
- constraint_pct = constraint_measure_LEOUF_topdecile * 100
- ) %>%
- select(Geneset, Geneset_Type, Geneset_Cat, constraint_pct,
- functional_pleiotropy_deletion, functional_pleiotropy_duplication, functional_pleiotropy_GWASenrichment)
- kable(head(df_fp), format = "markdown")
- kable(stats_panel_a, align = "c", caption = "Exact Stats Extracted for Panel A")
- ```
- ##### Panel A Figure #####
- ```{r, fig.width=9, fig.height=3.5, dpi=300}
- # Function to plot scatter and map the exact 'cor' and 'p-jaccard' from the stats dataframe
- plot_pleio_scatter_jaccard <- function(data, y_var, color_hex, stat_label, stats_df) {
- # Map BOTH correlation and p-value from the stats dataframe
- row_match <- stats_df %>% filter(stat == stat_label)
- cor_val <- row_match$cor
- pval_adj <- row_match$adj_pvalue_Jaccard
- # Format the label (add asterisk if FDR < 0.05)
- cor_label <- paste0("r=", round(cor_val, 2), ifelse(pval_adj < 0.05, "*", ""))
- ggplot(data, aes(x = constraint_pct, y = .data[[y_var]])) +
- geom_point(color = color_hex, size = 2, alpha = 1) +
- # Add the extracted correlation label to the plot
- annotate("text", x = Inf, y = Inf, hjust = 1.1, vjust = 1.5, label = cor_label, size = 6, fontface="bold") +
- geom_smooth(method = "lm", formula=y ~ x, se = FALSE, linewidth = 1.5, fullrange = TRUE, color = "black") +
- scale_y_continuous(limits = c(0, 75)) +
- theme_prism(axis_text_angle = 90, base_size = 14) +
- labs(x = "% constraint genes", y = "% functional pleiotropy") +
- theme(plot.title = element_text(hjust = 0.5, face = "bold"))
- }
- # Generate the three plots using their exact stat labels
- p_scatter_del <- plot_pleio_scatter_jaccard(
- df_fp, "functional_pleiotropy_deletion", "red3",
- "functional_burden_pleiotropy_deletion", stats_panel_a
- )
- p_scatter_dup <- plot_pleio_scatter_jaccard(
- df_fp, "functional_pleiotropy_duplication", "royalblue",
- "functional_burden_pleiotropy_duplication", stats_panel_a
- ) + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
- p_scatter_gwas <- plot_pleio_scatter_jaccard(
- df_fp, "functional_pleiotropy_GWASenrichment", "darkgreen",
- "functional_pleiotropy_GWASenrichment", stats_panel_a
- ) + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
- # Layout the final figure
- p_scatter_Fig4A <- (p_scatter_del + p_scatter_dup + p_scatter_gwas) + plot_layout(widths = c(1, 1, 1))
- print(p_scatter_Fig4A)
- ```
- ##### Panel A Stats #####
- ```{r, fig.width=18, fig.height=7 ,dpi=800 }
- kable(head(stats_panel_a), format = "markdown")
- ```
- #### Panel B: Distribution of constraint across Brain and Non-Brain gene sets
- Description: Box plots comparing the percentage of constraint genes between functional categories.
- ##### Panel B Data #####
- ```{r, fig.width=18, fig.height=7 ,dpi=800 }
- df_box <- df_fp %>%
- select(Geneset, Geneset_Cat, constraint_pct) %>%
- mutate(Geneset_Cat = factor(Geneset_Cat, levels = c("Brain", "NonBrain")))
- kable(head(df_box), format = "markdown")
- ```
- ##### Panel B Figure #####
- ```{r, fig.width=7, fig.height=4, dpi=400}
- p_box_Fig4B <- ggplot(df_box, aes(x = Geneset_Cat, y = constraint_pct, fill = Geneset_Cat)) +
- # Boxplot with black outline
- geom_boxplot(outlier.shape = NA, width = 0.5, alpha = 0.5, linewidth = 0.8, color = "black") +
- # Jitter points: shape 21 enables 'fill' aesthetic
- geom_jitter(aes(color=Geneset_Cat,shape = Geneset_Cat), stroke = 0.4,
- width = 0.15, size = 2, alpha = 0.9) +
- scale_fill_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
- scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
- scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
- theme_prism(base_size = in_base_size_ggprism) +
- theme(
- legend.position = "none",
- axis.text.x = element_text(size = 14, face = "bold"),
- axis.text.y = element_text(size = 14, face = "bold"),
- axis.title.y = element_text(size = 16, face = "bold")
- ) +
- labs(x = NULL, y = "% constraint genes") +
- coord_flip()
- print(p_box_Fig4B)
- ```
- ##### Panel B Stats #####
- ```{r, fig.width=9, fig.height=5, dpi=400}
- res_wilcox_B <- wilcox.test(
- df_box %>% filter(Geneset_Cat == "Brain") %>% pull(constraint_pct),
- df_box %>% filter(Geneset_Cat == "NonBrain") %>% pull(constraint_pct)
- )
- stats_panel_b <- data.frame(
- Test = "Wilcoxon Rank Sum Test",
- Comparison = "Brain vs Non-Brain Constraint (%)",
- W_Statistic = res_wilcox_B$statistic,
- P_Value = format(res_wilcox_B$p.value, digits = 4)
- )
- kable(stats_panel_b, align = "c", caption = "Statistical Comparison of Constraint Percentages")
- ```
- #### Panel C: Constraint vs Functional pleiotropy across metrics
- Description: Evaluates functional pleiotropy correlations across seven different established metrics of genetic constraint.
- ##### Panel C Data & Stats #####
- ```{r, fig.width=5, fig.height=9, dpi=400}
- load("df_constraint_pleiotropy_corr.RData")
- # Map the pre-computed dataframe directly to plotting coordinates
- df_metrics_cor <- df_constraint_pleiotropy_corr %>%
- mutate(
- # Map the raw constraint strings to clean metric labels
- Metric = case_when(
- grepl("LEOUF", constraint) ~ "LOEUF",
- grepl("missense_Z", constraint) ~ "missense Z",
- grepl("s_het", constraint) ~ "s-het",
- grepl("CDS", constraint) ~ "constraint CDS",
- grepl("pHaplo", constraint) ~ "pHaplo",
- grepl("pTriplo", constraint) ~ "pTriplo",
- grepl("GeneLenBP", constraint) ~ "gene length",
- TRUE ~ constraint
- ),
- # Map the stat column to Del, Dup, GWAS
- Type = case_when(
- grepl("deletion", stat) ~ "Del",
- grepl("duplication", stat) ~ "Dup",
- grepl("GWASenrichment", stat) ~ "GWAS",
- TRUE ~ stat
- ),
- # Assign significance using the Jaccard adjusted p-value directly
- sig = ifelse(adj_pvalue_Jaccard < 0.05, "q<0.05", "n.s.")
- ) %>%
- # Set factors for proper plot ordering
- mutate(
- Type = factor(Type, levels = c("Del", "Dup", "GWAS")),
- sig = factor(sig, levels = c("q<0.05", "n.s."))
- ) %>%
- select(Metric, Type, cor, pvalue_Jaccard, adj_pvalue_Jaccard, sig)
- # Order Y-axis metrics based on Deletion correlation strength
- metric_order <- df_metrics_cor %>% filter(Type == "Del") %>% arrange(cor) %>% pull(Metric)
- df_metrics_cor$Metric <- factor(df_metrics_cor$Metric, levels = metric_order)
- kable(head(df_metrics_cor), format = "markdown")
- ```
- ##### Panel C Figure #####
- ```{r, fig.width=6, fig.height=7, dpi=90}
- p_Fig4C <- ggplot(df_metrics_cor, aes(y = Metric, x = cor, group = Type)) +
- geom_point(aes(color = Type, fill = Type, shape = sig), size = 5, position = position_dodge(width = 0.6)) +
- scale_fill_manual(values = c("Del" = "red3", "Dup" = "royalblue", "GWAS" = "darkgreen")) +
- scale_color_manual(values = c("Del" = "red3", "Dup" = "royalblue", "GWAS" = "darkgreen")) +
- scale_shape_manual(values = c("q<0.05" = 21, "n.s." = 4)) +
- geom_vline(xintercept = 0, linetype = "longdash", color = "black") +
- geom_vline(xintercept = 0.4, linetype = "dotted", color = "black") +
- theme_prism(base_size = in_base_size_ggprism) +
- theme(legend.position = c(0.85, 0.2)) +
- labs(x = "Correlation", y = NULL) +
- coord_cartesian(xlim = c(0, 0.7))
- print(p_Fig4C)
- ```
- #### Panel D & E: Normative Null Models (Tissue & Cell Type)
- Description: Functional pleiotropy normalized by the fraction of genetic constraint, plotted against null distributions for Tissue (Fantom60) and Cell Type (HPA81).
- ##### Panel D & E Data #####
- ```{r, fig.width=6, fig.height=7, dpi=90}
- # 1. Extract and format the exact plotting coordinates from geneset level stats
- # No external null distribution data is required.
- df_normative <- df_geneset_level_stats %>%
- mutate(
- # Convert constraint fraction to percentage for the X-axis
- constraint_pct = constraint_measure_LEOUF_topdecile * 100
- ) %>%
- select(Geneset, Geneset_Type, Geneset_Cat, constraint_pct,
- normative_model_centile_del_pleiotropy, normative_model_centile_dup_pleiotropy)
- # Ensure Geneset_Cat is a factor for consistent color mapping
- df_normative$Geneset_Cat <- factor(df_normative$Geneset_Cat, levels = c("Brain", "NonBrain"))
- kable(head(df_normative), format = "markdown")
- ```
- ##### Panel D & E Figure #####
- ```{r, fig.width=18, fig.height=9, dpi=90}
- # --- 1. Compute Stats First (Needed for Dynamic Annotations) ---
- compute_normative_stats <- function(data, gset_type, metric_col, variant_label) {
- sub_df <- data %>% filter(Geneset_Type == gset_type)
- brain_vals <- sub_df %>% filter(Geneset_Cat == "Brain") %>% pull(.data[[metric_col]])
- nbrain_vals <- sub_df %>% filter(Geneset_Cat == "NonBrain") %>% pull(.data[[metric_col]])
- # Unified Wilcoxon Rank-sum Tests
- p_brain_50 <- wilcox.test(brain_vals, mu = 50,exact = FALSE)$p.value
- p_nbrain_50 <- wilcox.test(nbrain_vals, mu = 50,exact = FALSE)$p.value
- p_b_vs_nb <- wilcox.test(brain_vals, nbrain_vals,exact = FALSE)$p.value
- return(data.frame(Geneset_Type = gset_type, Variant = variant_label,
- P_Brain_vs_50 = p_brain_50, P_NonBrain_vs_50 = p_nbrain_50, P_Brain_vs_NonBrain = p_b_vs_nb))
- }
- # Run across all categories to get raw p-values
- df_stats_raw <- bind_rows(
- compute_normative_stats(df_normative, "Tissue_Fantom60", "normative_model_centile_del_pleiotropy", "Del"),
- compute_normative_stats(df_normative, "Tissue_Fantom60", "normative_model_centile_dup_pleiotropy", "Dup"),
- compute_normative_stats(df_normative, "SC_HPA81", "normative_model_centile_del_pleiotropy", "Del"),
- compute_normative_stats(df_normative, "SC_HPA81", "normative_model_centile_dup_pleiotropy", "Dup")
- )
- # 1. Pool all 8 tests for 'shift from 50%' and adjust FDR
- pooled_vs_50_pvals <- c(df_stats_raw$P_Brain_vs_50, df_stats_raw$P_NonBrain_vs_50)
- pooled_vs_50_fdr <- p.adjust(pooled_vs_50_pvals, method = "fdr")
- # 2. Pool all 4 tests for 'brain vs non-brain' and adjust FDR
- pooled_b_vs_nb_fdr <- p.adjust(df_stats_raw$P_Brain_vs_NonBrain, method = "fdr")
- # Map the correctly pooled FDR values back
- df_stats_DE_final <- df_stats_raw %>%
- mutate(
- # Indices 1:4 belong to Brain, indices 5:8 belong to NonBrain
- FDR_Brain_vs_50 = pooled_vs_50_fdr[1:4],
- FDR_NonBrain_vs_50 = pooled_vs_50_fdr[5:8],
- FDR_Brain_vs_NonBrain = pooled_b_vs_nb_fdr
- )
- # --- 2. Helper Function to Build Individual Components ---
- build_normative_plots <- function(data_obs, gset_type, var_y, variant_type, stats_df, y_label = NULL, hide_y = FALSE) {
- df_sub <- data_obs %>% filter(Geneset_Type == gset_type)
- plot_stats <- stats_df %>% filter(Geneset_Type == gset_type, Variant == variant_type)
- # Scatter Plot (Using manual annotations for the Centile bands instead of a null dataset)
- p_scatter <- ggplot(df_sub, aes(x = constraint_pct, y = .data[[var_y]])) +
- # Background Centile Ribbons
- annotate("rect", ymin = 5, ymax = 95, xmin = -Inf, xmax = Inf, fill = "grey80", alpha = 0.4) +
- annotate("rect", ymin = 10, ymax = 90, xmin = -Inf, xmax = Inf, fill = "grey70", alpha = 0.4) +
- annotate("rect", ymin = 25, ymax = 75, xmin = -Inf, xmax = Inf, fill = "grey60", alpha = 0.4) +
- geom_hline(yintercept = 50, color = "black", linewidth = 1) +
- # Scatter Points
- geom_point(aes(color = Geneset_Cat, shape = Geneset_Cat), size = 3) +
- scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
- scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
- scale_x_continuous(expand = c(0, 0), limits = c(0, 22), breaks = seq(0, 22, by = 5)) +
- scale_y_continuous(limits = c(0, 125), breaks = c(10, 25, 50, 75, 90)) +
- theme_prism(base_size = in_base_size_ggprism, base_line_size = in_base_size_ggprism/24) +
- theme(legend.position = "none") +
- labs(x = "% constraint genes", y = y_label)
- if (hide_y) {
- p_scatter <- p_scatter + theme(axis.title.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
- }
- # Violin Plot
- p_violin <- ggplot(df_sub, aes(x = Geneset_Cat, y = .data[[var_y]], color = Geneset_Cat,shape=Geneset_Cat)) +
- geom_hline(yintercept = 50, linetype = "dotted", color = "black", linewidth = 2) +
- geom_violin(fill = "white", alpha = 1, trim = TRUE, linewidth = 0.8) +
- geom_point(position = position_jitter(seed = 1, width = 0.2), size = 1.5, alpha = 0.8) +
- stat_summary(color = "black", fun.data = "mean_cl_boot", geom = "pointrange", size = 0.7) +
- scale_color_manual(values = c("Brain" = brain_colorcode, "NonBrain" = nonbrain_colorcode)) +
- scale_shape_manual(values = c("Brain" = 16, "NonBrain" = 17)) +
- scale_x_discrete(labels = c("Brain", "NonBrain")) +
- scale_y_continuous(limits = c(0, 125), breaks = c(10, 25, 50, 75, 90)) +
- theme_prism(base_size = in_base_size_ggprism, base_line_size = in_base_size_ggprism/24) +
- theme(legend.position = "none", axis.title.x = element_blank(), axis.title.y = element_blank(),
- axis.line.y = element_blank(), axis.text.y = element_blank(), axis.ticks.y = element_blank())
- # Add Significance Annotations dynamically based on the stats
- sig_size <- 10
- if (plot_stats$FDR_Brain_vs_50 < 0.05) {
- p_violin <- p_violin + annotate("text", x = 1, y = 105, label = "*", size = sig_size, color = brain_colorcode, fontface = "bold")
- }
- if (plot_stats$FDR_NonBrain_vs_50 < 0.05) {
- p_violin <- p_violin + annotate("text", x = 2, y = 105, label = "*", size = sig_size, color = nonbrain_colorcode, fontface = "bold")
- }
- if (plot_stats$FDR_Brain_vs_NonBrain < 0.05) {
- p_violin <- p_violin +
- annotate("segment", x = 1, xend = 2, y = 115, yend = 115, color = "grey50", linewidth = 1) +
- annotate("segment", x = 1, xend = 1, y = 110, yend = 115, color = "grey50", linewidth = 1) +
- annotate("segment", x = 2, xend = 2, y = 110, yend = 115, color = "grey50", linewidth = 1) +
- annotate("text", x = 1.5, y = 118, label = "*", size = sig_size, color = "grey50", fontface = "bold")
- }
- return(p_scatter + p_violin + plot_layout(widths = c(0.72, 0.28)))
- }
- # --- 3. Assemble Final Composite Panels ---
- y_label_text <- "functional pleiotropy\nnormalized by constraint (centiles)"
- pD_Del <- build_normative_plots(df_normative, "Tissue_Fantom60", "normative_model_centile_del_pleiotropy", "Del", df_stats_DE_final, y_label = y_label_text)
- pD_Dup <- build_normative_plots(df_normative, "Tissue_Fantom60", "normative_model_centile_dup_pleiotropy", "Dup", df_stats_DE_final, hide_y = TRUE)
- row_D <- pD_Del | pD_Dup
- pE_Del <- build_normative_plots(df_normative, "SC_HPA81", "normative_model_centile_del_pleiotropy", "Del", df_stats_DE_final, y_label = y_label_text)
- pE_Dup <- build_normative_plots(df_normative, "SC_HPA81", "normative_model_centile_dup_pleiotropy", "Dup", df_stats_DE_final, hide_y = TRUE)
- row_E <- pE_Del | pE_Dup
- p_single_stacked <- row_D / row_E
- print(p_single_stacked)
- ```
- ##### Panel D & E Stats #####
- ```{r, fig.width=6, fig.height=7, dpi=90}
- # --- Simplified and Robust Stats Calculation ---
- # Display the stats dataframe that was computed and utilized in the Figure chunk
- kable(df_stats_DE_final %>% select(Geneset_Type, Variant, starts_with("FDR"), starts_with("P")), align = "c", caption = "FDR-Adjusted Wilcoxon Tests for Normative Centiles (Panels D & E)")
- ```
Figure4.Rmd at commit 4415dfe, under MIT · at the source
Overview
- Centre de recherche Azrieli, CHU Sainte-Justine and University of Montréal, Montréal, QC Canada
- Department of Psychiatr, Harvard Medical School, Boston, MA USA
- Tommy Fuss Center for Neuropsychiatric Disease Research, Boston Children’s Hospital, Boston, MA USA
- Department of Biomedical and Health Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA USA
- Lifespan Brain Institute, Children’s Hospital of Philadelphia, and Penn Medicine, Philadelphia, PA USA
- Department of Genetics, University of Pennsylvania, Philadelphia, PA USA
- The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, ON Canada
- Department of Psychiatry, University of California San Diego, La Jolla, CA USA
- Lady Davis Institute for Medical Research, Jewish General Hospital, Montreal, QC Canada
- Gerald Bronfman Department of Oncology, Department of Epidemiology, Biostatistics and Occupational Health, McGill University, Montreal, QC Canada
- Department of Molecular Genetics, University of Toronto, Toronto, ON Canada
- Mila – Quebec AI Institute, University of Montréal, Montreal, QC Canada
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repositories
Its files are read in the Code ↔ Paper reader above, with 21 matches between paragraphs and lines of code.
martineaujeanlouis.github.io/mind-genesparallelcnv
Availability: 1 check, the latest on 27 September 2026: the link is dead (HTTP 404)
- 27 September 2026: the link is dead (HTTP 404)
linnarsson-lab/adult-human-brain
2b5aa12dffb8cc1bcb3979c409cb6ac907546b2e, 5 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
97 files
- cytograph/
__init__.py , Python, 15 lines - cytograph/
_version.py , Python, 1 line - cytograph/
annotation/ , Python, 4 lines__init__.py - cytograph/
annotation/ , Python, 142 linesauto_annotator.py - cytograph/
annotation/ , Python, 104 linesauto_auto_annotator.py - cytograph/
annotation/ , Python, 35 linescell_cycle_annotator.py - cytograph/
clustering/ , Python, 4 lines__init__.py - cytograph/
clustering/ , Python, 64 linescluster_validator.py - cytograph/
clustering/ , Python, 241 linespolished_louvain.py - cytograph/
clustering/ , Python, 59 linespolished_surprise.py - cytograph/
clustering/ , Python, 89 linesunpolished_louvain.py - cytograph/
decomposition/ , Python, 252 linesHPF.py - cytograph/
decomposition/ , Python, 303 linesHPF_accel.py - cytograph/
decomposition/ , Python, 3 lines__init__.py - cytograph/
decomposition/ , Python, 59 linesincremental_residuals_pc a.py - cytograph/
decomposition/ , Python, 99 linespca.py - cytograph/
embedding/ , Python, 1 line__init__.py - cytograph/
embedding/ , Python, 115 linesart_of_tsne.py - cytograph/
embedding/ , Python, 34 linescustom_init.py - cytograph/
enrichment/ , Python, 8 lines__init__.py - cytograph/
enrichment/ , Python, 86 linesbinary_differential_expr ession.py - cytograph/
enrichment/ , Python, 65 linesenrichment.py - cytograph/
enrichment/ , Python, 69 linesfeature_selection_by_dev iance.py - cytograph/
enrichment/ , Python, 144 linesfeature_selection_by_enr ichment.py - cytograph/
enrichment/ , Python, 185 linesfeature_selection_by_mul tilevel_enrichment.py - cytograph/
enrichment/ , Python, 56 linesfeature_selection_by_var iance.py - cytograph/
enrichment/ , Python, 75 linesgsea.py - cytograph/
enrichment/ , Python, 49 linesneighborhood_enrichment. py - cytograph/
enrichment/ , Python, 58 linestrinarizer.py - cytograph/
manifold/ , Python, 3 lines__init__.py - cytograph/
manifold/ , Python, 366 linesbalanced_knn.py - cytograph/
manifold/ , Python, 57 linesgraph_skeletonizer.py - cytograph/
manifold/ , Python, 103 linespoisson_pooling.py - cytograph/
metrics/ , Python, 2 lines__init__.py - cytograph/
metrics/ , Python, 47 lineslaplacian_score.py - cytograph/
metrics/ , Python, 68 linespoisson_proximity.py - cytograph/
metrics/ , Python, 142 linesspecial_metrics.py - cytograph/
pipeline/ , Python, 6 lines__init__.py - cytograph/
pipeline/ , Python, 128 linesaggregator.py - cytograph/
pipeline/ , Python, 516 linescommands.py - cytograph/
pipeline/ , Python, 109 linesconfig.py - cytograph/
pipeline/ , Python, 211 linescytograph.py - cytograph/
pipeline/ , Python, 245 linesengine.py - cytograph/
pipeline/ , Python, 165 linespunchcards.py - cytograph/
pipeline/ , Python, 20 linesutils.py - cytograph/
pipeline/ , Python, 519 linesworkflow.py - cytograph/
plotting/ , Python, 29 linesTF_heatmap.py - cytograph/
plotting/ , Python, 20 lines__init__.py - cytograph/
plotting/ , Python, 88 linesbatch_covariates.py - cytograph/
plotting/ , Python, 47 linesbuckets.py - cytograph/
plotting/ , Python, 26 linescell_cycle.py - cytograph/
plotting/ , Python, 72 linescolors.py - cytograph/
plotting/ , Python, 31 linesdecision_boundary.py - cytograph/
plotting/ , Python, 51 linesdendrogram.py - cytograph/
plotting/ , Python, 135 linesdoublets_plots.py - cytograph/
plotting/ , Python, 46 linesembedded_velocity.py - cytograph/
plotting/ , Python, 30 linesfactors.py - cytograph/
plotting/ , Python, 42 linesgene_velocity.py - cytograph/
plotting/ , Python, 166 linesheatmap.py - cytograph/
plotting/ , Python, 95 linesmanifold.py - cytograph/
plotting/ , Python, 26 linesmarkerheatmap.py - cytograph/
plotting/ , Python, 47 linesmetromap.py - cytograph/
plotting/ , Python, 12 linesmidpoint_normalize.py - cytograph/
plotting/ , Python, 91 linespunchcard_selection.py - cytograph/
plotting/ , Python, 120 linesqc_plots.py - cytograph/
plotting/ , Python, 85 linesradius_characteristics.p y - cytograph/
plotting/ , Python, 56 linesscatter.py - cytograph/
plotting/ , Python, 35 linesumi_genes.py - cytograph/
postprocessing/ , Python, 2 lines__init__.py - cytograph/
postprocessing/ , Python, 77 linesmerge_subset.py - cytograph/
postprocessing/ , Python, 186 linessplit_subset.py - cytograph/
preprocessing/ , Python, 4 lines__init__.py - cytograph/
preprocessing/ , Python, 173 linesdoublet_finder.py - cytograph/
preprocessing/ , Python, 104 linesgenotyping.py - cytograph/
preprocessing/ , Python, 68 linesnormalizer.py - cytograph/
preprocessing/ , Python, 24 linesqc_functions.py - cytograph/
preprocessing/ , Python, 26 linesutils.py - cytograph/
species/ , Python, 1 line__init__.py - cytograph/
species/ , Python, 1,801 lineshuman.py - cytograph/
species/ , Python, 1,508 linesmouse.py - cytograph/
species/ , Python, 154 linesspecies.py - cytograph/
utils.py , Python, 104 lines - cytograph/
visualization/ , Python, 2 lines__init__.py - cytograph/
visualization/ , Python, 281 linescolors.py - cytograph/
visualization/ , Python, 214 linesplot_overview.py - cytograph/
visualization/ , Python, 134 linesscatter.py - notebooks/
Preprint/ , Jupyter, 1,638 linesFigure1.ipynb - notebooks/
Preprint/ , Jupyter, 504 linesFigure2.ipynb - notebooks/
Preprint/ , Jupyter, 1,726 linesFigure3.ipynb - notebooks/
Preprint/ , Jupyter, 1,116 linesFigure5.ipynb - notebooks/
Preprint/ , Jupyter, 288 linesFigure5_Diffxpy.ipynb - notebooks/
Preprint/ , Jupyter, 605 linesFigure6.ipynb - notebooks/
Preprint/ , Jupyter, 513 linesFigureS1.ipynb - notebooks/
Preprint/ , Jupyter, 456 linesFigureS10.ipynb - notebooks/
Preprint/ , Jupyter, 482 linesFigureS7.ipynb - repository limit reached (2,000 files or 30 MB): the rest is at the source (18 files)
- LICENSE, License, 25 lines
- README.md, Text, 75 lines
SayehKazem/FunBurd
4415dfed4e8b276b374f1384041f310332fd34b8, 19 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
11 files
- Figure_Generation/
Figure2.Rmd , R, 910 lines, 2 matches - Figure_Generation/
Figure3.Rmd , R, 696 lines, 2 matches - Figure_Generation/
Figure4.Rmd , R, 377 lines, 4 matches - Figure_Generation/
Figure5.Rmd , R, 541 lines, 3 matches - Figure_Generation/
Figure6.Rmd , R, 431 lines, 1 match - Functional_Burden_Associ
ation_Analysis/ , Jupyter, 398 lines, 2 matchesFunBurd_Multitrait_Effec tSizes_Computation_Scrip t_On_ToyDatasets.ipynb - Functional_Burden_Associ
ation_Analysis/ , Python, 325 linesFunBurd_Multitrait_Effec tSizes_Computation_Scrip t_On_ToyDatasets.py - Functional_Burden_Associ
ation_Analysis/ , Shell, 33 linesSlurm_FunBurd.sh - Permutation_PreservingJa
ccardDistance_P_Jaccard/ , Jupyter, 303 lines, 1 matchPJaccard_Computation.ipy nb - LICENSE, License, 21 lines
- README.md, Text, 135 lines
Zenodo 17038354
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
6 files
funburd.readthedocs.io
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
Code availability statement
The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: funburd.readthedocs.io, linnarsson-lab/
adult-human-brain , SayehKazem/FunBurd
Read it in the paper: doi.org/10.1038/s41467-026-76676-0.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 5 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 108 scripts, each with its path and the digest of its content;
- 21 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- ukbiobank.ac.uk/
enable-your-research/ , at UK Biobank; found in “Data availability”apply-for-access
Code and data availability statement
The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to a dataset: ukbiobank.ac.uk/
enable-your-research/ apply-for-access - it points to the authors' code: funburd.readthedocs.io, linnarsson-lab/
adult-human-brain , SayehKazem/FunBurd
Read it in the paper: doi.org/10.1038/s41467-026-76676-0.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Funding: added Fondation Brain Canada; Compute Canada; Canada First Research Excellence Fund; Institut de Valorisation des Données; National Institutes of Health: 1u01mh119690-01; Canadian Institutes of Health Research: CIHR_400528
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 20 authors, 3 keywords, 7 MeSH terms, 86 references.
Cite
This paper
Kazem, S., Kumar, K., Yang, J., Benitiere, F., Huguet, G., Mollon, J., Renne, T., Schultz, L. M., Knowles, E. E. M., Engchuan, W., Shanta, O., Thiruvahindrapuram, B., MacDonald, J. R., Greenwood, C. M. T., Scherer, S. W., Almasy, L., Sebat, J., Glahn, D. C., Dumas, G., & Jacquemont, S. (2026). Determinants of functional burden pleiotropy and gene dosage responses across human traits. Nature communications, 17(1), 9798. https://
BibTeX
@article{kazem2026determ
author = {Kazem, Sayeh and Kumar, Kuldeep and Yang, Jane and Benitiere, Florian and Huguet, Guillaume and Mollon, Josephine and Renne, Thomas and Schultz, Laura M and Knowles, Emma E M and Engchuan, Worrawat and Shanta, Omar and Thiruvahindrapuram, Bhooma and MacDonald, Jeffrey R and Greenwood, Celia M T and Scherer, Stephen W and Almasy, Laura and Sebat, Jonathan and Glahn, David C and Dumas, Guillaume and Jacquemont, Sébastien},
title = {{Determinants of functional burden pleiotropy and gene dosage responses across human traits}},
journal = {Nature communications},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {9798},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42736290},
pmcid = {PMC13575219}
}
RIS
TY - JOUR
AU - Kazem, Sayeh
AU - Kumar, Kuldeep
AU - Yang, Jane
AU - Benitiere, Florian
AU - Huguet, Guillaume
AU - Mollon, Josephine
AU - Renne, Thomas
AU - Schultz, Laura M
AU - Knowles, Emma E M
AU - Engchuan, Worrawat
AU - Shanta, Omar
AU - Thiruvahindrapuram, Bhooma
AU - MacDonald, Jeffrey R
AU - Greenwood, Celia M T
AU - Scherer, Stephen W
AU - Almasy, Laura
AU - Sebat, Jonathan
AU - Glahn, David C
AU - Dumas, Guillaume
AU - Jacquemont, Sébastien
TI - Determinants of functional burden pleiotropy and gene dosage responses across human traits
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 9798
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Determinants of functional burden pleiotropy and gene dosage responses across human traits",
"container-title": "Nature communications",
"author": [
{
"family": "Kazem",
"given": "Sayeh"
},
{
"family": "Kumar",
"given": "Kuldeep"
},
{
"family": "Yang",
"given": "Jane"
},
{
"family": "Benitiere",
"given": "Florian"
},
{
"family": "Huguet",
"given": "Guillaume"
},
{
"family": "Mollon",
"given": "Josephine"
},
{
"family": "Renne",
"given": "Thomas"
},
{
"family": "Schultz",
"given": "Laura M"
},
{
"family": "Knowles",
"given": "Emma E M"
},
{
"family": "Engchuan",
"given": "Worrawat"
},
{
"family": "Shanta",
"given": "Omar"
},
{
"family": "Thiruvahindrapuram",
"given": "Bhooma"
},
{
"family": "MacDonald",
"given": "Jeffrey R"
},
{
"family": "Greenwood",
"given": "Celia M T"
},
{
"family": "Scherer",
"given": "Stephen W"
},
{
"family": "Almasy",
"given": "Laura"
},
{
"family": "Sebat",
"given": "Jonathan"
},
{
"family": "Glahn",
"given": "David C"
},
{
"family": "Dumas",
"given": "Guillaume"
},
{
"family": "Jacquemont",
"given": "Sébastien"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "9798",
"DOI": "10.1038/
"PMID": "42736290",
"PMCID": "PMC13575219",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.21203/rs.3.rs-9246968/v1 [code]
- Copy number variants reveal divergent genetic and diagnostic cortical signatures across psychiatric disordersJournal: Research Square (preprint)In common: circlize, ComplexHeatmap, ggpubr, 4 other tools, genetics / omics, 6 references, 6 authors
- [2] doi:10.1038/s41586-026-10515-6 [code]
- An X-linked long non-coding RNA, PTCHD1-AS, and the core features of autism.Journal: NatureIn common: pandas, SciPy, Matplotlib, 1 other tool, genetics / omics, cellular / molecular, 4 authors
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: UMAP, igraph, circlize, 16 other tools, cellular / molecular
- [4] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: UMAP, igraph, circlize, 15 other tools, genetics / omics, cellular / molecular
- [5] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: UMAP, igraph, circlize, 14 other tools, genetics / omics, cellular / molecular, 1 reference
- [6] doi:10.1016/j.xcrm.2026.102651 [code]
- Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.Journal: Cell reports. MedicineIn common: UMAP, igraph, circlize, 14 other tools, genetics / omics, cellular / molecular
- [7] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: UMAP, igraph, circlize, 14 other tools, genetics / omics
- [8] doi:10.1038/s41380-026-03571-x [code]
- Convergent coexpression reveals shared biological mechanisms underlying common and rare variant risk in six neuropsychiatric disorders.Journal: Molecular psychiatryIn common: igraph, reshape2, data.table, 8 other tools, genetics / omics, cellular / molecular, 5 references
- [9] doi:10.1038/s41586-026-10735-w [code]
- Distributed control circuits across a brain-and-cord connectome.Journal: NatureIn common: UMAP, igraph, circlize, 14 other tools
- [10] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: circlize, Numba, ComplexHeatmap, 13 other tools, genetics / omics, cellular / molecular, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 5 repositories of the authors' code, each at its verified commit and with its license, 108 scripts, and 21 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:51ea8774736ae343…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
