OSCR

CroCoNet: a framework for the quantitative comparison of gene regulatory networks across species.

Code ↔ Paper

40 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 40 matches · 5 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Cell lines ↔ 4.paper_figures_and_tables/suppl.tables.R, lines 32–83 · score 0.98 · peripheral blood mononuclear, urine derived stem, dermal fibroblasts, integration free Sendai, OSKM vectors, RNA seq
  2. [2] § Methods › Quantification of cross-species divergence ↔ R/data.R, lines 482–563 · score 0.94 · weighted linear model, upper boundary, inversely proportional, lower boundary, prediction interval, identify modules
  3. [3] § Methods › Quantification of cross-species divergence ↔ R/findConservedDivergedModules.R, lines 1–72 · score 0.94 · upper boundary, inversely proportional, lower boundary, dependent variable, prediction interval, linear model
  4. [4] § Results › Consensus network and module assignment ↔ R/pruneModules.R, lines 162–294 · score 0.90 · cumulative sum curves, median module, knee point, regulator target adjacency, pruning step, intramodular connectivity
  5. [5] § Methods › Quantification of cross-species divergence ↔ R/plotConservedDivergedModules.R, lines 1–60 · score 0.87 · lower boundary, prediction interval, linear model, jackknife versions, subtree length, module tree
  6. [6] § Methods › Target gene contribution ↔ R/findConservedDivergedTargets.R, lines 1–61 · score 0.87 · jackknifed module version, diverged target genes, target gene contribution, linear model, subtree length, original module
  7. [7] § Methods › Quantification of cross-species divergence ↔ R/findConservedDivergedModules.R, lines 1–72 · score 0.87 · externally studentized residuals, Benjamini Hochberg, discovery rate, robust subset, diverged modules, FDR
  8. [8] § Methods › POU5F1 perturbations using CRISPRi ↔ 2.validations/2.8.POU5F1_CRISPRi/2.8.9.stemness_scores_and_knockdown_efficiencies.R, lines 1–40 · score 0.84 · knockdown efficiencies, gRNAs, POU5F1 perturbed, control cells, validation, species
  9. [9] § Results › Lineage-specific divergence ↔ vignettes/CroCoNet.Rmd, lines 1032–1052 · score 0.83 · recent common ancestor, human MRCA, minimal requirement, branch leading, assess module divergence, human replicates form
  10. [10] § Methods › Quantification of cross-species divergence ↔ R/data.R, lines 482–563 · score 0.83 · externally studentized residuals, Benjamini Hochberg, discovery rate, FDR, POU5F1, cutoff
  11. [11] § Methods › Module assignment ↔ R/pruneModules.R, lines 162–294 · score 0.83 · cumulative sum curve, knee point, Pruning steps, pruned module, initial module, iterative
  12. [12] § Results › Applying CroCoNet to primate scRNA-seq data ↔ R/pruneModules.R, lines 64–159 · score 0.82 · cumulative sum curve, knee point, regulator target adjacency, intramodular connectivity, edge weights, pruned modules
  13. [13] § Results › Quantifying module divergence ↔ 4.paper_figures_and_tables/figure4.R, lines 83–162 · score 0.81 · negatively correlated POU5F1, stemness score, positively correlated, POU5F1 module member, network genes, module genes
  14. [14] § Results › Quantifying module divergence ↔ R/plotTreeStats.R, lines 112–257 · score 0.81 · linear regression model, neighbor joining trees, branch lengths, trees represent, module topology, preservation scores
  15. [15] § Methods › Network inference ↔ vignettes/CroCoNet.Rmd, lines 113–145 · score 0.80 · network inference, GRNBoost2, inferred networks, network versions, importance scores, potential regulators
  16. [16] § Methods › Inferring binding potential and binding site divergence of the regulators ↔ 2.validations/2.3.binding_site_enrichment_and_divergence/2.3.4.ATAC_seq_mapping.sh, the whole file · a weak match · score 0.79 · NPC samples, gorGor6, human NPC, ATAC seq, BWA, MEM2
  17. [17] § Methods › Inferring binding potential and binding site divergence of the regulators ↔ 2.validations/2.3.binding_site_enrichment_and_divergence/2.3.18.associate_peaks_to_gene.R, lines 82–118 · score 0.79 · gorGor6, cynomolgus NPC, gorilla NPCs, human NPC, LiftOver, ATAC seq
  18. [18] § Results › Applying CroCoNet to primate scRNA-seq data ↔ R/data.R, lines 41–93 · score 0.76 · neural progenitor cells, GRNBoost2, scRNA, cynomolgus macaque, day, trajectory
  19. [19] § Results › Preservation statistics to compare module topologies ↔ R/comparePresStats.R, lines 1–39 · score 0.75 · fine grained topology, intramodular connectivities, module topologies, cor adj, CroCoNet, adjacencies
  20. [20] § Results › Preservation statistics to compare module topologies ↔ R/calculatePresStats.R, lines 1–44 · score 0.74 · joint module assignment, preserved module, intramodular connectivities, cor adj, consensus network, WGCNA
  21. [21] § Results › Rewiring of the POU5F1 module ↔ 2.validations/2.7.POU5F1_LTR7_enrichment/2.7.3.calculate_enrichment_near_POU5F1_module_members.R, lines 50–96 · score 0.72 · near POU5F1 module, POU5F1 module member, LTR7 elements, iPSCs, genes expressed, bound
  22. [22] § Methods › Module preservation, tree reconstruction and module filtering ↔ R/filterModuleTrees.R, the whole file · a weak match · score 0.71 · probability density functions, bivariate normal distributions, random modules, tree, filtering
  23. [23] § Methods › Experimental procedures and data processing ↔ R/data.R, lines 96–137 · score 0.71 · gorGor6, macFas6, cynomolgus macaque, transferred, Liftoff, GENCODE
  24. [24] § Methods › Module preservation, tree reconstruction and module filtering ↔ R/plotTreeStats.R, lines 112–257 · score 0.71 · neighbor joining tree, chosen statistic, subtree length, preservation scores, species diversity, reconstruction
  25. [25] § Results › Quantifying module divergence ↔ vignettes/CroCoNet.Rmd, lines 927–950 · score 0.69 · positive score indicates, target gene weakens, negative score, divergence score, Target gene contribution, conserved modules
  26. [26] § Methods › Inferring binding potential and binding site divergence of the regulators ↔ 2.validations/2.3.binding_site_enrichment_and_divergence/2.3.21.summarize_motif_scores_per_gene.Rmd, lines 1157–1196 · score 0.69 · peaks associated, motif scores, gene associations, ATAC seq, Cluster, random modules
  27. [27] § Methods › Inferring binding potential and binding site divergence of the regulators ↔ 4.paper_figures_and_tables/suppl.tables.R, lines 128–169 · score 0.68 · binding site divergence, human cynomolgus, human gorilla, species pair, median, conserved
  28. [28] § Methods › Network inference ↔ 1.neural_differentiation_dataset/1.2.network_inference/1.2.1.prepare_data.R, the whole file · a weak match · score 0.67 · randomized quantile residual, network inference, transformations, matrix, regulators, genes
  29. [29] § Methods › Module preservation, tree reconstruction and module filtering ↔ R/data.R, lines 344–387 · score 0.66 · neighbor joining tree, distance matrix, subtree length, preservation scores, species diversity, infer
  30. [30] § Methods › Inferring binding potential and binding site divergence of the regulators ↔ 2.validations/2.7.POU5F1_LTR7_enrichment/2.7.3.calculate_enrichment_near_POU5F1_module_members.R, lines 1–48 · score 0.66 · ATAC seq peaks, iPSCs, gene associations, kb, Active, validation
  31. [31] § Methods › Network inference ↔ R/addDirectionality.R, the whole file · a weak match · score 0.66 · correlation coefficients, expression matrices, gene pairs, Spearman, inference, cells
  32. [32] § Methods › Module preservation, tree reconstruction and module filtering ↔ R/data.R, lines 390–435 · score 0.66 · cor.adj, intramodular connectivities, cor.kIM, module preservation, correlation, quantified
  33. [33] § Methods › Experimental procedures and data processing ↔ 3.brain_dataset/3.1.data_preparation/3.1.2.filtering.R, lines 45–85 · score 0.66 · H18.30.001, atypical cell, donor, composition, metadata, brain
  34. [34] § Results › Rewiring of the POU5F1 module ↔ 4.paper_figures_and_tables/suppl.figureS12_S13.R, lines 161–229 · score 0.65 · POU5F1 binding, LTR7 elements, binding sites, S12, S13, HERVH
  35. [35] § Results › Preservation statistics to compare module topologies ↔ R/comparePresStats.R, lines 1–39 · score 0.65 · fine grained, phylogenetic information, cor adj, CroCoNet, signal, topologies
  36. [36] § Methods › Cell lines ↔ 1.neural_differentiation_dataset/1.3.CroCoNet_analysis/1.3.1.CroCoNet_analysis.R, lines 43–81 · score 0.62 · Homo sapiens, Macaca fascicularis, cynomolgus, gorilla, human
  37. [37] § Results › Preservation statistics to compare module topologies ↔ 4.paper_figures_and_tables/suppl.figureS6.R, the whole file · a weak match · score 0.60 · cor.adj, phylogenetic information, cor.kIM, S6, inference, neural
  38. [38] § Methods › POU5F1 perturbations using CRISPRi ↔ 2.validations/2.8.POU5F1_CRISPRi/2.8.8.QC_and_filtering.R, lines 258–334 · score 0.60 · dCas9, gRNAs, POU5F1, cynomolgus, cell, human
  39. [39] § Results › Overview of the CroCoNet pipeline ↔ R/reconstructTrees.R, lines 101–184 · score 0.59 · neighbor joining tree, branch lengths, distance matrix, preservation score, CroCoNet, pipeline
  40. [40] § Results › Consensus network and module assignment ↔ 4.paper_figures_and_tables/figure2.R, lines 322–365 · score 0.59 · PAX6 ChIP seq, ChIP seq enrichment, NANOG, fraction, motifs, peak

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 563 lines · 49 KB · GPL-3.0 · 6 matches

  1. #' Transcriptional regulators
  2. #'
  3. #' 7 important transcriptional regulators during early neural differentiation that are used as cores to assemble modules around them.
  4. #'
  5. #' @format A character vector of 7 elements.
  6. "regulators"
  7. #' All genes
  8. #'
  9. #' The 300 genes that are used for the network inference.
  10. #'
  11. #' @format A character vector of 300 elements.
  12. "genes"
  13. #' JASPAR 2024 vertebrate core transcriptional regulators
  14. #'
  15. #' Transcriptional regulators that have at least 1 annotated motif in the JASPAR 2024 vertebrate core collection.
  16. #'
  17. #' @format A character vector of 784 elements.
  18. "jaspar_core_TRs"
  19. #' JASPAR 2024 unvalidated transcriptional regulators
  20. #'
  21. #' Transcriptional regulators that have at least 1 annotated motif in the JASPAR 2024 unvalidated collection.
  22. #'
  23. #' @format A character vector of 590 elements.
  24. "jaspar_unvalidated_TRs"
  25. #' IMAGE transcriptional regulators
  26. #'
  27. #' Transcriptional regulators that have at least 1 annotated motif in the IMAGE database (Madsen et al. 2018).
  28. #'
  29. #' @format A character vector of 1351 elements.
  30. "image_TRs"
  31. #' Replicate-species conversion
  32. #'
  33. #' A data frame that specifies which replicate belongs to which species.
  34. #'
  35. #' @format A data frame with 9 rows and 2 columns:
  36. #' \describe{
  37. #' \item{replicate}{Name of the replicate/cell line.}
  38. #' \item{species}{Name of the species.}
  39. #' }
  40. "replicate2species"
  41. #' SCE object of the primate neural differentiation dataset
  42. #'
  43. #' A subset of the primate neural differentiation scRNA-seq dataset in an SCE format. The data was collected during the early neural differentiation of human, gorilla and cynomolgus macaque iPS cells with 3 human, 2 gorilla and 4 cynomolgus cell lines (replicates). Using a directed differentiation protocol, cells were differentiated into neural progenitor cells (NPCs) over the course of 9 days, and scRNA-seq data was obtained at six time points (days 0, 1, 3, 5, 7 and 9) during this process. The SCE object contains the raw and log-normalized counts as well as the metadata for 300 genes and 900 cells (100 cells per replicate).
  44. #'
  45. #' @format An SCE object with 300 rows, 900 columns, 9 metadata columns and 2 assays.
  46. #'
  47. #' Metadata columns:
  48. #' \describe{
  49. #' \item{species}{Name of the species.}
  50. #' \item{replicate}{Name of the replicate/cell line.}
  51. #' \item{day}{The day when the cell was collected.}
  52. #' \item{n_UMIs}{Number of UMIs detected.}
  53. #' \item{n_genes}{Number of genes detected.}
  54. #' \item{perc_mito}{Percent of mitochondrial reads.}
  55. #' \item{sizeFactor}{Size factor for scaling normalization, calculated first per replicate using [scran::computeSumFactors] and [scran::quickCluster], then adjusted by [batchelor::multiBatchNorm] to remove systematic differences in covergae across replicates.}
  56. #' \item{pseudotime}{Pseudotime inferred by [SCORPIUS::infer_trajectory].}
  57. #' \item{cell_type}{Cell type labels predicted by [SingleR::classifySingleR] using the embryoid body dataset from Rhodes et al. 2022 as reference.}
  58. #' }
  59. #' Assays:
  60. #' \describe{
  61. #' \item{counts}{Raw counts.}
  62. #' \item{logcounts}{Log-normalized counts created by first calculating size factors per replicate using [scran::computeSumFactors] and [scran::quickCluster], then adjusted them to remove systematic differences in covergae across replicates and log-normalizing by [batchelor::multiBatchNorm].}
  63. #' }
  64. "sce"
  65. #' List of raw networks
  66. #'
  67. #' List of networks per replicate with raw edge weights. The networks were inferred using GRNBoost2 based on a subset of the primate neural differentiation scRNA-seq dataset. All 300 genes in the subsetted data were used as potential regulators. To circumvent the stochastic nature of the algorithm, GRNBoost2 was run 10 times on the same count matrices, then the results were averaged across runs, and rarely occurring edges were removed altogether. In addition, edges inferred between the same gene pair but in opposite directions were also averaged.
  68. #'
  69. #' @format A named list of 9 [igraph] objects. Each network contains 300 nodes, and has 1 node and 2 edge attributes:
  70. #'
  71. #' Node attributes:
  72. #' \describe{
  73. #' \item{name}{Name of the node (gene).}
  74. #' }
  75. #' Edge attributes:
  76. #' \describe{
  77. #' \item{weight}{Edge weight, the importance score calculate by GRNBoost2.}
  78. #' }
  79. "network_list_raw"
  80. #' List of genomic annotations
  81. #'
  82. #' List of genomic annotations per species as GRanges objects. The human annotation is the GTF of Hg38 GENCODE release 32 primary assembly, while the gorilla and cynomolgus macaque annotations were created by transferring the human annotation onto the gorGor6 and macFas6 genomes via the tool Liftoff (https://github.com/agshumate/Liftoff). Each annotation was subsetted for the 300 genes that feature in this example dataset.
  83. #'
  84. #' @format A named list of 3 GRanges objects. Each object has 8 columns.
  85. #'
  86. #' Attributes:
  87. #' \describe{
  88. #' \item{seqnames}{Chromosome or contig name.}
  89. #' \item{start}{Genomic start location.}
  90. #' \item{end}{Genomic end location.}
  91. #' \item{width}{Width of the feature in base pairs.}
  92. #' \item{strand}{Genomic strand ("+" or "-").}
  93. #' \item{source}{The prediction program or public database where the annotations came from.}
  94. #' \item{type}{Feature type (gene, transcript, exon, CDS, UTR, start_codon, stop_codon or Selenocysteine).}
  95. #' \item{score}{The degree of confidence in the feature's existence and coordinates.}
  96. #' \item{phase}{One of '0', '1' or '2'. '0' means that the first base of the feature is the first base of a codon, '1' that the second base is the first base of a codon, and so on.}
  97. #' \item{gene_id}{Unique identifier of the gene.}
  98. #' \item{gene_name}{Name of the gene.}
  99. #' \item{transcript_id}{Unique identifier of the transcript.}
  100. #' \item{transcript_name}{Name of the transcript.}
  101. #' }
  102. #' @source <ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_32/gencode.v32.primary_assembly.annotation.gtf.gz>, <https://hgdownload.soe.ucsc.edu/goldenPath/gorGor6/bigZips/gorGor6.fa.gz>, <https://ftp.ensembl.org/pub/release-109/fasta/macaca_fascicularis/dna/Macaca_fascicularis.Macaca_fascicularis_6.0.dna_sm.toplevel.fa.gz>
  103. "gtf_list"
  104. #' List of networks
  105. #'
  106. #' List of networks per replicate with edge weights re-scaled between 0 and 1, and gene pairs (edges) with overlapping annotations removed. The networks were inferred using GRNBoost2 based on a subset of the primate neural differentiation scRNA-seq dataset. All 300 genes in the subsetted data were used as potential regulators. To circumvent the stochastic nature of the algorithm, GRNBoost2 was run 10 times on the same count matrices, then the results were averaged across runs, and rarely occurring edges were removed altogether. In addition, edges inferred between the same gene pair but in opposite directions were also averaged. Edge weights were scaled by the maximum edge weight across all replicates. Gene pairs that have overlapping annotations in any of the species' genomes were removed from all networks.
  107. #'
  108. #' @format A named list of 9 [igraph] objects. Each network contains 300 nodes, and has 1 node attribute and 3 edge attributes:
  109. #'
  110. #' Node attributes:
  111. #' \describe{
  112. #' \item{name}{Name of the node (gene).}
  113. #' }
  114. #' Edge attributes:
  115. #' \describe{
  116. #' \item{weight}{Edge weight, the importance score calculate by GRNBoost2 rescaled between 0 and 1.}
  117. #' \item{genomic_dist}{Numeric, the genomic distance of the 2 genes that form the edge (Inf if the 2 genes are annotated on different chromosomes/contigs).}
  118. #' }
  119. "network_list"
  120. #' Phylogenetic tree
  121. #'
  122. #' Rooted phylogenetic tree of 3 primate species: human (Homo sapiens), gorilla (Gorilla gorilla) and cynomolgus macaque (Macaca Fascicularis). The tree was created by subsetting the mammalian tree with the best estimates of branch lengths from Bininda-Edmons et al. 2007.
  123. #'
  124. #' @format A [phylo] object with 4 edges and 3 nodes.
  125. #' @source Olaf R. P. Bininda-Emonds, Marcel Cardillo, Kate E. Jones, Ross D. E. MacPhee, Robin M. D. Beck, Richard Grenyer, Samantha A. Price, Rutger A. Vos, John L. Gittleman & Andy Purvis. "The delayed rise of present-day mammals" Nature 446, 507-512(29 March 2007). doi:10.1038/nature05634
  126. "tree"
  127. #' Consensus network
  128. #'
  129. #' Consensus network of the 9 primate replicates in the example dataset. For each edge, the consensus adjacency was calculated as the weighted average of replicate-wise adjacencies using weights that correct for 1) the phylogenetic distances between species and 2) the different numbers of replicates per species. If an edge was not detected in certain replicate, the adjacency of that replicate was regarded as 0 for the calculation of the consensus. The directionality of each edge was determined based on a modified Spearman's correlation between the corresponding 2 genes' expression profiles (positive expression correlation - activating interaction, negative expression correlation - repressing interaction). The correlations were calculated per replicate, then the mean correlation was taken across all replicates.
  130. #'
  131. #' @format An [igraph] object with 300 nodes, and has 1 node attribute and 3 edge attributes:
  132. #'
  133. #' Node attributes:
  134. #' \describe{
  135. #' \item{name}{Name of the node (gene).}
  136. #' }
  137. #' Edge attributes:
  138. #' \describe{
  139. #' \item{weight}{Consensus edge weight/adjacency, the weighted average of replicate-wise adjacencies.}
  140. #' \item{rho}{Approximate Spearman's correlation coefficient of the 2 genes' expression profiles that form the edge.}
  141. #' \item{p.adj}{BH-corrected approximate p-value of rho.}
  142. #' \item{direction}{Direction of the interaction between the 2 genes that form the edge ("+" or "-").}
  143. #' }
  144. "consensus_network"
  145. #' Initial modules
  146. #'
  147. #' Initial modules created by assigning the top 250 targets to each of 7 transcriptional regulators involved in the early neuronal differentiation of primates (\code{regulators}). For each regulator, the top 250 targets were selected based edge weights in the consensus network: all targets of the regulator were ranked based on their edge weight to the regulator (regulator-taregt adjacency) and the 250 targets with the highest regulator-target adjacencies were kept.
  148. #'
  149. #' @format A data frame with 1750 rows and 8 columns:
  150. #' \describe{
  151. #' \item{regulator}{Character, transcriptional regulator.}
  152. #' \item{target}{Target gene of the transcriptional regulator (member of the regulator's initial module).}
  153. #' \item{weight}{Consensus edge weight/adjacency, the weighted average of replicate-wise adjacencies.}
  154. #' \item{rho}{Approximate Spearman's correlation coefficient of the 2 genes' expression profiles that form the edge.}
  155. #' \item{p.adj}{BH-corrected approximate p-value of rho.}
  156. #' \item{direction}{Direction of the interaction between the 2 genes that form the edge ("+" or "-").}
  157. #' }
  158. "initial_modules"
  159. #' Pruned modules
  160. #'
  161. #' Pruned modules created by the dynamic pruning of the initial modules via the "UIK_adj_kIM" method. The initial module members were filtered in successive steps based on their adjacency to the regulator and their intramodular connectivitiy alternately. In each step, the cumulative sum curve based on one of these two characteristics was calculated per module, then the targets below the knee point of the curve were kept. This process was continued until the module sizes became as small as possible without the median falling below a pre-defined minimum of 20 genes.
  162. #'
  163. #' @format A data frame with 225 rows and 9 columns:
  164. #' \describe{
  165. #' \item{regulator}{Character, transcriptional regulator.}
  166. #' \item{target}{Character, target gene of the transcriptional regulator (member of the regulator's pruned module).}
  167. #' \item{weight}{Numeric, consensus edge weight/adjacency, the weighted average of replicate-wise adjacencies.}
  168. #' \item{rho}{Approximate Spearman's correlation coefficient of the 2 genes' expression profiles that form the edge.}
  169. #' \item{p.adj}{BH-corrected approximate p-value of rho.}
  170. #' \item{direction}{Character specifying the direction of regulation between the regulator and the target, either "+" or "-".}
  171. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  172. #' }
  173. "pruned_modules"
  174. #' Random modules
  175. #'
  176. #' Random modules that match the actual (pruned) modules in size. These random modules have the same regulators and contain the same number of target genes as the actual modules, but the target genes were randomly drawn from all genes in the network.
  177. #'
  178. #' @format A data frame with 225 rows and 3 columns:
  179. #' \describe{
  180. #' \item{regulator}{Character, transcriptional regulator.}
  181. #' \item{target}{Member of the regulator's random module.}
  182. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  183. #' }
  184. "random_modules"
  185. #' Eigengenes
  186. #'
  187. #' Eigengenes of the pruned modules calculated using the activated targets in each module. An eigengene summarizes the expression profile of the module as a whole, mathematically it is the first principal component of the module expression data (i.e. the scaled and centered logcounts subsetted for the activated targets and the regulators of the given module). In this case, the principal component was taken across all cells irrespective of species.
  188. #'
  189. #' @format A data frame with 6300 rows and 8 columns:
  190. #'\describe{
  191. #' \item{cell}{Character, the cell barcode.}
  192. #' \item{species}{Character, the name of the species.}
  193. #' \item{pseudotime}{Numeric, inferred pseudotime.}
  194. #' \item{cell_type}{Character, cell type annotation.}
  195. #' \item{module}{Character, transcriptional regulator and direction of regulation (in this case always nameOfRegulator(+)).}
  196. #' \item{eigengene}{Numeric, the eigengene (i.e. the first principal component of the scaled and centered logcounts) of the module.}
  197. #' \item{mean_expr}{Numeric, the mean of the scaled and centered logcounts across all genes in the module.}
  198. #' \item{regulator_expr}{Numeric, the scaled and centered logcounts of the regulator.}
  199. #' }
  200. "eigengenes"
  201. #' Eigengenes per species
  202. #'
  203. #' Eigengenes of the pruned modules calculated per species using the activated targets in each module. An eigengene summarizes the expression profile of the module as a whole, mathematically it is the first principal component of the module expression data (i.e. the scaled and centered logcounts subsetted for the activated targets and the regulator of the given module). In this case, the principal component was calculated for each species separately.
  204. #'
  205. #' @format A data frame with 6300 rows and 8 columns:
  206. #'\describe{
  207. #' \item{cell}{Character, the cell barcode.}
  208. #' \item{species}{Character, the name of the species.}
  209. #' \item{pseudotime}{Numeric, inferred pseudotime.}
  210. #' \item{cell_type}{Character, cell type annotation.}
  211. #' \item{module}{Character, transcriptional regulator and direction of regulation (in this case always nameOfRegulator(+)).}
  212. #' \item{eigengene}{Numeric, the eigengene (i.e. the first principal component of the scaled and centered logcounts) of the module.}
  213. #' \item{mean_expr}{Numeric, the mean of the scaled and centered logcounts across all genes in the module.}
  214. #' \item{regulator_expr}{Numeric, the scaled and centered logcounts of the regulator.}
  215. #' }
  216. "eigengenes_per_species"
  217. #' Preservation statistics of the original and jackknifed pruned modules
  218. #'
  219. #' Correlation of intramodular connectivities (cor_kIM) per replicate pair for the original and all jackknifed versions of the pruned modules. The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). The preservation statistic cor.kIM was then calculated for the original module as well as each jackknife module version by comparing each replicate to all others, both within and across species. cor.kIM quantifies how well the connectivity patterns are preserved between the networks of two replicates, mathematically it is the correlation of the intramodular connectivities per module member gene in the network of the 1st replicate VS the intramodular connectivities per module member gene in the network of the 2nd replicate.
  220. #'
  221. #' @format A data frame with 8352 rows and 8 columns:
  222. #' \describe{
  223. #' \item{regulator}{Character, transcriptional regulator.}
  224. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  225. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  226. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  227. #' \item{gene_removed}{Character, the name of the gene removed by jackknifing (NA in case of module type 'orig').}
  228. #' \item{replicate1, replicate2}{Character, the names of the replicates compared.}
  229. #' \item{species1, species2}{Character, the names of the species \code{replicate1} and \code{replicate2} belong to, respectively.}
  230. #' \item{cor_adj}{Numeric, correlation of adjacencies.}
  231. #' \item{cor_kIM}{Numeric, correlation of intramodular connectivities.}
  232. #' }
  233. "pres_stats_jk"
  234. #' Preservation statistics of the the original and jackknifed random modules
  235. #'
  236. #' Correlation of intramodular connectivities (cor_kIM) per replicate pair for the original and all jackknifed versions of the random modules. The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). The preservation statistic cor.kIM was then calculated for the original module as well as each jackknife module version by comparing each replicate to all others, both within and across species. cor.kIM quantifies how well the connectivity patterns are preserved between the networks of two replicates, mathematically it is the correlation of the intramodular connectivities per module member gene in the network of the 1st replicate VS the intramodular connectivities per module member gene in the network of the 2nd replicate.
  237. #'
  238. #' @format A data frame with 8352 rows and 8 columns:
  239. #' \describe{
  240. #' \item{regulator}{Character, transcriptional regulator.}
  241. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  242. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  243. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  244. #' \item{gene_removed}{Character, the name of the gene removed by jackknifing (NA in case of module type 'orig').}
  245. #' \item{replicate1, replicate2}{Character the names of the replicates compared.}
  246. #' \item{species1, species2}{Character, the names of the species \code{replicate1} and \code{replicate2} belong to, respectively.}
  247. #' \item{cor_adj}{Numeric, correlation of adjacencies.}
  248. #' \item{cor_kIM}{Numeric, correlation of intramodular connectivities.}
  249. #' }
  250. "random_pres_stats_jk"
  251. #' Distance measures of the original and jackknifed pruned modules
  252. #'
  253. #' Distance measures per replicate pair for the original and all jackknifed versions of the pruned modules. The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). The distance measures were calculated based on the correlation of intramodular connectivities: \eqn{dist = \frac{1 - cor.kIM}{2}} (a correlation of 1 corresponds to a distance of 0, whereas a correlation of -1 corresponds to a distance of 1). Each element of the list corresponds a jackknifed/original module version and contains the distance measures between all possible pairs of replicates for this module version.
  254. #'
  255. #' @format A named list with 232 elements containing the distance measures per (original or jackknifed) module version. Each element is a data frame with 72 rows and 10 columns:
  256. #' \describe{
  257. #' \item{regulator}{Character, transcriptional regulator.}
  258. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  259. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  260. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  261. #' \item{gene_removed}{Character, the name of the gene removed by jackknifing (NA in case of module type 'orig').}
  262. #' \item{replicate1, replicate2}{Character the names of the replicates compared.}
  263. #' \item{species1, species2}{Character, the names of the species \code{replicate1} and \code{replicate2} belong to, respectively.}
  264. #' \item{dist}{Numeric, distance measure ranging from 0 to 1, calculated based on the correlation of intramodular connectivities.}
  265. #' }
  266. "dist_jk"
  267. #' Distance measures of the pruned modules
  268. #'
  269. #' Distance measures per replicate pair for the original (i.e. not jackknifed) pruned modules.
  270. #'
  271. #' The distance measures were calculated based on the correlation of intramodular connectivities: \deqn{dist = \frac{1 - cor.kIM}{2}} (a correlation of 1 corresponds to a distance of 0, whereas a correlation of -1 corresponds to a distance of 1). Each element of the list corresponds to a module and contains the distance measures between all possible pairs of replicates for this module.
  272. #'
  273. #' @format A named list with 12 elements containing the distance measures per module. Each element is a data frame with 72 rows and 7 columns:
  274. #' \describe{
  275. #' \item{regulator}{Character, transcriptional regulator.}
  276. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  277. #' \item{replicate1, replicate2}{Character the names of the replicates compared.}
  278. #' \item{species1, species2}{Character, the names of the species \code{replicate1} and \code{replicate2} belong to, respectively.}
  279. #' \item{dist}{Numeric, distance measure ranging from 0 to 1, calculated based on the correlation of intramodular connectivities.}
  280. #' }
  281. "dist"
  282. #' Distance measures of the the original and jackknifed random modules
  283. #'
  284. #' Distance measures per replicate pair for the original and all jackknifed versions of the random modules.
  285. #'
  286. #' The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). The distance measures were calculated based on the correlation of intramodular connectivities: \deqn{dist = \frac{1 - cor.kIM}{2}} (a correlation of 1 corresponds to a distance of 0, whereas a correlation of -1 corresponds to a distance of 1). Each element of the list corresponds a jackknife module version and contains the distance measures between all possible pairs of replicates for this module version.
  287. #'
  288. #' @format A named list with 232 elements containing the distance measures per (original or jackknifed) module version. Each element is a data frame with 72 rows and 10 columns:
  289. #' \describe{
  290. #' \item{regulator}{Character, transcriptional regulator.}
  291. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  292. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  293. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  294. #' \item{gene_removed}{Character, the name of the gene removed by jackknifing (NA in case of module type 'orig').}
  295. #' \item{replicate1, replicate2}{Character the names of the replicates compared.}
  296. #' \item{species1, species2}{Character, the names of the species \code{replicate1} and \code{replicate2} belong to, respectively.}
  297. #' \item{dist}{Numeric, distance measure ranging from 0 to 1, calculated based on the correlation of intramodular connectivities.}
  298. #' }
  299. "random_dist_jk"
  300. #' Trees of the original and jackknifed pruned modules
  301. #'
  302. #' Neighbor-joining trees representing the similarities of connectivity patterns across the 9 primate replicates for the original and all jackknifed versions of the pruned modules. The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). For each of these module versions, the trees were inferred based on the preservation statistic cor.kIM (correlation of intramodular connectivities): first, the preservation scores were calculated between all possible replicate pairs, then they were converted into a distance matrix of replicates, and finally trees were reconstructed based on this distance matrix using the neighbor-joining algorithm. The result is a single tree for each original or jackknife module where the tips represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between these replicates.
  303. #' @format A named list with 232 elements containing the neighbor-joining trees as [phylo] objects.
  304. "trees_jk"
  305. #' Trees of the pruned modules
  306. #'
  307. #' Neighbor-joining trees representing the similarities of connectivity patterns across the 9 primate replicates for the original (i.e. not jackknifed) pruned modules. For each of the modules, the trees were inferred based on the preservation statistic cor.kIM (correlation of intramodular connectivities): first, the preservation scores were calculated between all possible replicate pairs, then they were converted into a distance matrix of replicates, and finally trees were reconstructed based on this distance matrix using the neighbor-joining algorithm. The result is a single tree per module where the tips represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between these replicates.
  308. #' @format A named list with 12 elements containing the neighbor-joining trees as [phylo] objects.
  309. "trees"
  310. #' Trees of the original and jackknifed random modules
  311. #'
  312. #' Neighbor-joining trees representing the similarities of connectivity patterns across the 9 primate replicates for the original and all jackknifed versions of the random modules. The jackknifed versions of the modules were created by removing each target gene assigned to a module (the regulators were never excluded). For each of these module versions, the trees were inferred based on the preservation statistic cor.kIM (correlation of intramodular connectivities): first, the preservation scores were calculated between all possible replicate pairs, then they were converted into a distance matrix of replicates, and finally trees were reconstructed based on this distance matrix using the neighbor-joining algorithm. The result is a single tree for each original or jackknife module where the tips represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between these replicates.
  313. #' @format A named list with 232 elements containing the neighbor-joining trees as [phylo] objects.
  314. "random_trees_jk"
  315. #' Tree-based statistics of the original and jackknifed pruned modules
  316. #'
  317. #' Tree-based statistics calculated for the original and all jackknifed versions of the pruned modules. For each module version, a tree was reconstructed based on the module preservation scores within and across species. The tips of resulting tree represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between the replicates. Branches between replicates of different species carry information about cross-species differences, while the branches between replicates of the same species carry information about the within-species diversity.
  318. #' @format A data frame with 232 rows and 16 columns:
  319. #' \describe{
  320. #' \item{regulator}{Character, transcriptional regulator.}
  321. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  322. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  323. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  324. #' \item{gene_removed}{The name of the gene removed by jackknifing (NA in case of module type 'orig').}
  325. #' \item{within_species_diversity}{Numeric, the sum of the human, gorilla and cynomolgus diversity.}
  326. #' \item{total_tree_length}{Numeric, the sum of all branch lengths in the tree.}
  327. #' \item{human_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the human tips.}
  328. #' \item{gorilla_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the gorilla tips.}
  329. #' \item{cynomolgus_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the cynomolgus tips.}
  330. #' \item{human_monophyl}{Logical indicating whether the tree is monophyletic for the human replicates.}
  331. #' \item{gorilla_monophyl}{Logical indicating whether the tree is monophyletic for the gorilla replicates.}
  332. #' \item{cynomolgus_monophyl}{Logical indicating whether the tree is monophyletic for the cynomolgus replicates.}
  333. #' \item{human_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the human replicates and includes the internal branch connecting the human replicates to the rest of the tree. NA if the tree is not monophyletic for the human replicates.}
  334. #' \item{gorilla_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the gorilla replicates and includes the internal branch connecting the gorilla replicates to the rest of the tree. NA if the tree is not monophyletic for the gorilla replicates.}
  335. #' \item{cynomolgus_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the cynomolgus replicates and includes the internal branch connecting the cynomolgus replicates to the rest of the tree. NA if the tree is not monophyletic for the cynomolgus replicates.}
  336. #' }
  337. "tree_stats_jk"
  338. #' Tree-based statistics of the original and jackknifed random modules
  339. #'
  340. #' Tree-based statistics calculated for the original and all jackknifed versions of the random modules. For each module version, a tree was reconstructed based on the module preservation scores within and across species. The tips of resulting tree represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between the replicates. Branches between replicates of different species carry information about cross-species differences, while the branches between replicates of the same species carry information about the within-species diversity.
  341. #' @format A data frame with 232 rows and 16 columns:
  342. #' \describe{
  343. #' \item{regulator}{Character, transcriptional regulator.}
  344. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  345. #' \item{type}{Character, module type (orig = original or jk = jackknifed).}
  346. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk' and nameOfRegulator_orig in case of module type 'orig').}
  347. #' \item{gene_removed}{The name of the gene removed by jackknifing (NA in case of module type 'orig').}
  348. #' \item{within_species_diversity}{Numeric, the sum of the human, gorilla and cynomolgus diversity.}
  349. #' \item{total_tree_length}{Numeric, the sum of all branch lengths in the tree.}
  350. #' \item{human_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the human tips.}
  351. #' \item{gorilla_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the gorilla tips.}
  352. #' \item{cynomolgus_diversity}{Numeric, the sum of the branch lengths in the subtree that contains only the cynomolgus tips.}
  353. #' \item{human_monophyl}{Logical indicating whether the tree is monophyletic for the human replicates.}
  354. #' \item{gorilla_monophyl}{Logical indicating whether the tree is monophyletic for the gorilla replicates.}
  355. #' \item{cynomolgus_monophyl}{Logical indicating whether the tree is monophyletic for the cynomolgus replicates.}
  356. #' \item{human_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the human replicates and includes the internal branch connecting the human replicates to the rest of the tree. NA if the tree is not monophyletic for the human replicates.}
  357. #' \item{gorilla_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the gorilla replicates and includes the internal branch connecting the gorilla replicates to the rest of the tree. NA if the tree is not monophyletic for the gorilla replicates.}
  358. #' \item{cynomolgus_subtree_length}{Numeric, the sum of the branch lengths in the subtree that is defined by the cynomolgus replicates and includes the internal branch connecting the cynomolgus replicates to the rest of the tree. NA if the tree is not monophyletic for the cynomolgus replicates.}
  359. #' }
  360. "random_tree_stats_jk"
  361. #' Preservation statistics of the pruned modules
  362. #'
  363. #' Correlation of intramodular connectivities (cor_kIM) per replicate pair and module. The preservation statistic cor.kIM quantifies how well the connectivity patterns are preserved between the networks of two replicates, mathematically it is the correlation of the intramodular connectivities per module member gene in the network of the 1st replicate VS the intramodular connectivities per module member gene in the network of the 2nd replicate.
  364. #' This statistic was calculated for all possible jackknifed versions of the modules, each of which was created by removing a target gene assigned to the given module (the regulators were never excluded). Each jackknifed module was compared between all posible pairs of replicates, both within and across species, resulting in a cor.kIM value per jackknifed module version and replicate pair. Finally, the cor.kIM values were summarized per module and replicate pair by taking the median and its 95\% confidence interval across all jackknifed module versions.
  365. #'
  366. #' @format A data frame with 252 rows and 10 columns:
  367. #' \describe{
  368. #' \item{regulator}{Character, transcriptional regulator.}
  369. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  370. #' \item{replicate1, replicate2}{The names of the replicates compared.}
  371. #' \item{species1, species2}{The names of the species 'replicate1' and 'replicate2' belongs to, respectively.}
  372. #' \item{cor_kIM}{The median of cor.kIM across all jackknifed versions of the module.}
  373. #' \item{var_cor_kIM}{The variance of cor.kIM across all jackknifed versions of the module.}
  374. #' \item{lwr_cor_kIM}{The lower bound of the 95\% confidence interval of cor.kIM calculated by jackknifing.}
  375. #' \item{upr_cor_kIM}{The upper bound of the 95\% confidence interval of cor.kIM calculated by jackknifing.}
  376. #' \item{cor_adj}{The median of cor.adj across all jackknifed versions of the module.}
  377. #' \item{var_cor_adj}{The variance of cor.adj across all jackknifed versions of the module.}
  378. #' \item{lwr_cor_adj}{The lower bound of the 95\% confidence interval of cor.adj calculated by jackknifing.}
  379. #' \item{upr_cor_adj}{The upper bound of the 95\% confidence interval of cor.adj calculated by jackknifing.}
  380. #' }
  381. "pres_stats"
  382. #' Preservation statistics of the random modules
  383. #'
  384. #' Correlation of intramodular connectivities (cor_kIM) per replicate pair and random module. The preservation statistic cor.kIM quantifies how well the connectivity patterns are preserved between the networks of two replicates, mathematically it is the correlation of the intramodular connectivities per module member gene in the network of the 1st replicate VS the intramodular connectivities per module member gene in the network of the 2nd replicate.
  385. #' This statistic was calculated for all possible jackknifed versions of the modules, each of which was created by removing a target gene assigned to the given module (the regulators were never excluded). Each jackknifed module was compared between all posible pairs of replicates, both within and across species, resulting in a cor.kIM value per jackknifed module version and replicate pair. Finally, the cor.kIM values were summarized per module and replicate pair by taking the median and its 95\% confidence interval across all jackknifed module versions.
  386. #'
  387. #' @format A data frame with 252 rows and 10 columns:
  388. #' \describe{
  389. #' \item{regulator}{Character, transcriptional regulator.}
  390. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  391. #' \item{replicate1, replicate2}{The names of the replicates compared.}
  392. #' \item{species1, species2}{The names of the species 'replicate1' and 'replicate2' belongs to, respectively.}
  393. #' \item{cor_kIM}{The median of cor.kIM across all jackknifed versions of the module.}
  394. #' \item{var_cor_kIM}{The variance of cor.kIM across all jackknifed versions of the module.}
  395. #' \item{lwr_cor_kIM}{The lower bound of the 95\% confidence interval of cor.kIM calculated by jackknifing.}
  396. #' \item{upr_cor_kIM}{The upper bound of the 95\% confidence interval of cor.kIM calculated by jackknifing.}
  397. #' \item{cor_adj}{The median of cor.adj across all jackknifed versions of the module.}
  398. #' \item{var_cor_adj}{The variance of cor.adj across all jackknifed versions of the module.}
  399. #' \item{lwr_cor_adj}{The lower bound of the 95\% confidence interval of cor.adj calculated by jackknifing.}
  400. #' \item{upr_cor_adj}{The upper bound of the 95\% confidence interval of cor.adj calculated by jackknifing.}
  401. #' }
  402. "random_pres_stats"
  403. #' Tree-based statistics of the pruned modules
  404. #'
  405. #' Tree-based statistics per module. The trees were reconstructed based on module preservation scores (cor.kIM) within and across species. The tips of resulting tree represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between the replicates. The total tree length is the sum of all branch lengths in the tree and it measures module variability both within and across species, while the within-species_diversity is the sum of the human, gorilla and cynomolgus within-species branch lengths and it measures module variability within species only.
  406. #' These statistics were calculated for all possible jackknifed versions of the modules, each of which was created by removing a target gene assigned to the given module (the regulators were never excluded). Then the statistics were summarized per module by taking the median and its 95\% confidence interval across all jackknifed module versions.
  407. #'
  408. #' @format A data frame with 7 rows and 10 columns:
  409. #' \describe{
  410. #' \item{regulator}{Character, transcriptional regulator.}
  411. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  412. #' \item{total_tree_length}{The median of total tree lengths across all jackknifed versions of the module.}
  413. #' \item{var_total_tree_length}{The variance of total tree lengths across all jackknifed versions of the module.}
  414. #' \item{lwr_total_tree_length}{The lower bound of the 95\% confidence interval of the total tree length calculated by jackknifing.}
  415. #' \item{upr_total_tree_length}{The upper bound of the 95\% confidence interval of the total tree length calculated by jackknifing.}
  416. #' \item{within_species_diversity}{The median of within-species_diversity values across all jackknifed versions of the module.}
  417. #' \item{var_within_species_diversity}{The variance of within-species_diversity values across all jackknifed versions of the module.}
  418. #' \item{lwr_within_species_diversity}{The lower bound of the 95\% confidence interval of the within-species_diversity calculated by jackknifing.}
  419. #' \item{upr_within_species_diversity}{The upper bound of the 95\% confidence interval of the within-species_diversity calculated by jackknifing.}
  420. #' }
  421. "tree_stats"
  422. #' Tree-based statistics of the random modules
  423. #'
  424. #' Tree-based statistics per random module. The trees were reconstructed based on module preservation scores (cor.kIM) within and across species. The tips of resulting tree represent the replicates and the branch lengths represent the dissimilarity of connectivity patterns between the replicates. The total tree length is the sum of all branch lengths in the tree and it measures module variability both within and across species, while the within-species_diversity is the sum of the human, gorilla and cynomolgus within-species branch lengths and it measures module variability within species only.
  425. #' These statistics were calculated for all possible jackknifed versions of the modules, each of which was created by removing a target gene assigned to the given module (the regulators were never excluded). Then the statistics were summarized per module by taking the median and its 95\% confidence interval across all jackknifed module versions.
  426. #'
  427. #' @format A data frame with 7 rows and 10 columns:
  428. #' \describe{
  429. #' \item{regulator}{Character, transcriptional regulator.}
  430. #' \item{module_size}{Module size, the numer of target genes assigned to a regulator.}
  431. #' \item{total_tree_length}{The median of total tree lengths across all jackknifed versions of the module.}
  432. #' \item{var_total_tree_length}{The variance of total tree lengths across all jackknifed versions of the module.}
  433. #' \item{lwr_total_tree_length}{The lower bound of the 95\% confidence interval of the total tree length calculated by jackknifing.}
  434. #' \item{upr_total_tree_length}{The upper bound of the 95\% confidence interval of the total tree length calculated by jackknifing.}
  435. #' \item{within_species_diversity}{The median of within-species_diversity values across all jackknifed versions of the module.}
  436. #' \item{var_within_species_diversity}{The variance of within-species_diversity values across all jackknifed versions of the module.}
  437. #' \item{lwr_within_species_diversity}{The lower bound of the 95\% confidence interval of the within-species_diversity calculated by jackknifing.}
  438. #' \item{upr_within_species_diversity}{The upper bound of the 95\% confidence interval of the within-species_diversity calculated by jackknifing.}
  439. #' }
  440. "random_tree_stats"
  441. #' Linear model for the characterization of overall module conservation
  442. #'
  443. #' A weighted linear model between the total tree length and within-species diversity of the 12 modules in the subsetted early neuronal differentiation dataset. It captures the general relationship between the two tree-based statistics in this dataset and thus also describes the expected total tree length for any given within-species diversity. By identifying modules that differ the most from this expectation, i.e. identifying the outlier data points, it can be used to pinpoint conserved and overall diverged modules. In combination with jackknifing, it can also be used to identify target genes within these modules that contribute the most to the conservation/divergence. The linear model was fit by calling \code{\link{lm}} with weights inversely proportional to the variance of the total tree lengths.
  444. #'
  445. #' @format An object of class \code{\link{lm}} fitted on the object \code{tree_stats} with the formula "total_tree_length ~ within_species_diversity".
  446. "lm_overall"
  447. #' Cross-species conservation measures per module
  448. #'
  449. #' Cross-species conservation measures of the 12 modules in the subsetted early neuronal differentiation dataset. Modules that were found to be conserved or diverged overall (across all species) are labelled as "conserved" or "diverged" in the column \code{conservation}.
  450. #'
  451. #' To determine whether a module as whole is conserved or diverged overall, module trees were reconstructed from pairwise preservation scores between replicates and based on these trees 2 useful statistics were calculated for each module: the total tree length and the within-species diversity (for details please see \code{\link{calculatePresStats}}, \code{\link{reconstructTrees}}, \code{\link{calculateTreeStats}}). After fitting a weighted linear model between the total tree length and within-species diversity values of all modules, a module was considered diverged if it fell above the prediction interval of the regression line, while a module was considered conserved if it fell below the prediction interval of the regression line (for details please see \code{\link{fitTreeStatsLm}} and \code{\link{findConservedDivergedModules}}. The degree of conservation/divergence can be further compared between the modules categorized as conserved/diverged using 2 measures, the residual and the t-score.
  452. #'
  453. #' @format A data frame with 12 rows and 16 columns:
  454. #' \describe{
  455. #' \item{focus}{Character, the focus of interest in terms of cross-species conservation ("overall").}
  456. #' \item{regulator}{Character, transcriptional regulator.}
  457. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  458. #' \item{total_tree_length}{Numeric, total tree length per module (the median across all jackknife versions of the module).}
  459. #' \item{lwr_total_tree_length}{Numeric, the lower bound of the confidence interval of the total tree length calculated based on the jackknifed versions of the module.}
  460. #' \item{upr_total_tree_length}{Numeric, the upper bound of the confidence interval of the total tree length calculated based on the jackknifed versions of the module.}
  461. #' \item{within_species_diversity}{Numeric, within-species diveristy per module (the median across all jackknife versions of the module).}
  462. #' \item{lwr_within_species_diversity}{Numeric, the lower bound of the confidence interval of the within-species diversity calculated based on the jackknifed versions of the module.}
  463. #' \item{upr_within_species_diversity}{Numeric, the upper bound of the confidence interval of the within-species diversity calculated based on the jackknifed versions of the module.}
  464. #' \item{fit}{Numeric, the fitted total tree length at the within-species diversity value of the module.}
  465. #' \item{lwr_fit}{Numeric, the lower bound of the prediction interval of the fit.}
  466. #' \item{upr_fit}{Numeric, the upper bound of the prediction interval of the fit.}
  467. #' \item{residual}{Numeric, the residual of the module in the linear model. It is calculated as the difference between the observed and expected (fitted) total tree lengths.}
  468. #' \item{studentized_residual}{Numeric, the externally studentized residual of the module. This statistic measures how strongly a module deviates from the fitted regression line relative to the expected variability.}
  469. #' \item{p_value}{Numeric, two-sided p-value associated with the externally studentized residual.}
  470. #' \item{fdr}{Numeric, false discovery rate obtained by adjusting the p-values across all modules using the Benjamini–Hochberg method.}
  471. #' \item{category}{Character, classification of the module based on the prediction interval: "diverged" if a module has a higher total tree length than the upper boundary of the prediction interval, "conserved" if a module has a lower total tree length than the lower boundary of the prediction interval, and "within_expectation" otherwise.}
  472. #' \item{robust}{Logical, indicates whether the module passes an additional robustness filter based on the false discovery rate (FDR). Modules with an FDR below \code{fdr_cutoff} are marked as \code{TRUE}.}
  473. #' }
  474. "module_conservation_overall"
  475. #' Cross-species conservation measures of the target genes in the POU5F1 module
  476. #'
  477. #' Cross-species conservation measures of the 21 target genes in the POU5F1 module in the subsetted early neuronal differentiation dataset.
  478. #'
  479. #' To determine which target genes are the most responsible for the divergence of the POU5F1 module, the same statistics were be used as for the identification of conserved/diverged modules but in combination with jackknifing. This means that each target gene in the module was removed and all statistics, including the tree-based statistics that carry information about cross-species conservation, were re-calculated (please see \code{\link{calculatePresStats}}, \code{\link{reconstructTrees}} and \code{\link{calculateTreeStats}}). Our working hypothesis is that if removing a target gene from the diverged POU5F1 module makes the module more conserved, then this gene contributed to the divergence of the original module in the first place. This can be quantified using the residuals compared to the regression line between the total tree lengths and within-species diversities of the tree representations across all modules (please see \code{\link{fitTreeStatsLm}} and \code{\link{findConservedDivergedTargets}}. The jackknifed versions of the POU5F1 module that have the lowest residuals are the most conserved, therefore the corresponding target genes are considered the most diverged.
  480. #'
  481. #' @format A data frame with 23 rows and 12 columns:
  482. #' \describe{
  483. #' \item{focus}{Character, the focus of interest in terms of cross-species conservation ("overall").}
  484. #' \item{regulator}{Character, transcriptional regulator.}
  485. #' \item{module_size}{Integer, the numer of target genes assigned to a regulator.}
  486. #' \item{type}{Character, module type (orig = original, jk = jackknifed or summary = the summary across all jackknife module versions).}
  487. #' \item{id}{Character, the unique ID of the module version (format: nameOfRegulator_jk_nameOfGeneRemoved in case of module type 'jk', nameOfRegulator_orig in case of module type 'orig' and nameOfRegulator_summary in case of module type 'summary').}
  488. #' \item{gene_removed}{Character, the name of the gene removed by jackknifing (NA in case of module types 'orig' and 'summary').}
  489. #' \item{total_tree_length}{Numeric, total tree length per jackknife module version.}
  490. #' \item{within_species_diversity}{Numeric, within-species diversity per jackknife module version.}
  491. #' \item{fit}{Numeric, the fitted total tree length at the within-species diversity value of a jackknife module version.}
  492. #' \item{lwr_fit}{Numeric, the lower bound of the prediction interval of the fit.}
  493. #' \item{upr_fit}{Numeric, the upper bound of the prediction interval of the fit.}
  494. #' \item{residual}{Numeric, the residual of the jackknife module version in the linear model. It is calculated as the difference between the observed and expected (fitted) total tree lengths.}
  495. #' \item{contribution}{Numeric, the contribution score of the removed target gene.}
  496. #' }
  497. "POU5F1_target_conservation"

data.R at commit c0bbc8a, under GPL-3.0 · at the source

Overview

Authors: Anita Térmeg1, Vladyslav Storozhuk1, Zane Kliesmete1, Fiona C. Edenhofer1, Johanna Geuder1, Tamina Dietl2, Beate Vieth1, Philipp Janssen1, Daniel Richter1, Boyan Bonev2, Ines Hellmann1
  1. Anthropology and Human Genomics, Ludwig-Maximilians-Universität München,82152, Planegg, Germany
  2. Research Unit Brain Epigenomics, Helmholtz Center Munich,81377, Munich, Germany
Journal: Genome biology, volume 27, issue 1, article 228
Dates: received 17 November 2025; accepted 5 June 2026; published online 15 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s13059-026-04152-5 · PMID 42458527 · PMCID PMC13371563 · OpenAlex W7168237161
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: Connectivity, Statistics
MeSH: Gene Regulatory Networks*, Software*, Animals, Humans, Species Specificity (* major topic)
Journal subjects: Methodology
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: European Research Council Consolidator Grant (EpiCortex, 101044469); Deutsche Forschungsgemeinschaft (DFG) (458888224); BioHPC hosted at Leibniz Rechenzentrum Munich funded by the DFG (INST 86/2050-1 FUGG); Ludwig-Maximilians-Universität München (1024)
Citations: not cited yet (Europe PMC); 119 references in the paper

Abstract

To understand phenotypic evolution, it is essential to investigate the underlying gene regulatory networks (GRNs). However, most comparative GRN analyzes remain descriptive due to the low signal-to-noise ratio inherent in single-cell transcriptomics data. To address this, we introduce CroCoNet (Cross-species Comparison of Networks), an R-package for quantitative GRN comparison across species. CroCoNet builds comparable network modules centered on putative regulators and compares module topologies within and between species, distinguishing true evolutionary divergence from technical and biological confounders. We demonstrate its utility by comparing early neural differentiation across primates and validating results with a CRISPRi analysis of the diverged POU5F1 module.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.

anita-termeg/gRNA-design

License: none: the authors keep all their rights
State: the link is dead, verified on 27 September 2026
Evidence: found in the paper
Software Heritage: not archived
Found in: the text, “POU5F1 perturbations using CRISPRi”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link is dead
  • 27 September 2026: the link is dead

Hellmann-Lab/CroCoNet

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c0bbc8a98747f33d2dfdd4251129c277df2ea502, 17 March 2026
Languages: R (37), Python (1), Shell (1)
Size: 345 files, 39 scripts
Software Heritage: not archived
Found in: the references
Holds: README, license file, environment (DESCRIPTION), continuous integration, documentation, 2 notebooks
Not found: CITATION.cff, tests
Tools: tidyverse (27 files), ggplot2 (11 files), data.table (10 files), igraph (10 files), patchwork (4 files), SingleCellExperiment (3 files), clusterProfiler (1 file), ggpubr (1 file), pandas (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
41 files

Hellmann-Lab/CroCoNet-analyses

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 4fd700f49c35ad41c161b80e6f021682c1826507, 1 June 2026
Languages: R (75), Shell (37), Python (1)
Size: 120 files, 113 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, 5 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (66 files), patchwork (22 files), SingleCellExperiment (18 files), cowplot (17 files), ggplot2 (17 files), data.table (7 files), igraph (6 files), SAMtools (6 files), broom (5 files), ggpubr (5 files), BEDTools (3 files), limma (3 files), Seurat (3 files), Cell Ranger (2 files), pandas (2 files), DESeq2 (1 file), Harmony (1 file), Numba (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
115 files

Zenodo 20343560

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 5 files
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (66 files), patchwork (22 files), SingleCellExperiment (18 files), cowplot (17 files), ggplot2 (17 files), data.table (7 files), igraph (6 files), SAMtools (6 files), broom (5 files), ggpubr (5 files), BEDTools (3 files), limma (3 files), Seurat (3 files), Cell Ranger (2 files), pandas (2 files), DESeq2 (1 file), Harmony (1 file), Numba (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
115 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 265 scripts, each with its path and the digest of its content;
  • 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The CroCoNet R-package is available at https://github.com/Hellmann-Lab/CroCoNet [108] under the GPL-3.0 license, with detailed documentation and a step-by-step tutorial at https://hellmann-lab.github.io/CroCoNet/. The code for all analyses in this manuscript is available at https://github.com/Hellmann-Lab/CroCoNet-analyses [109] under the GPL-3.0 license. A stable version of the code and the related data files have been deposited to the linked Zenodo archive [110].

Datasets generated for this study have been deposited in public repositories. The neural differentiation scRNA-seq data have been deposited in ArrayExpress under the accession number E-MTAB-15695 [111]. The ATAC-seq data of the cell lines G1c2, G1c3 and C1c3 have been deposited in ArrayExpress under accession number E-MTAB-15654 [112]. The long-read RNA-seq data have been deposited in GEO under the accession number GSE326106 [113]. The single-cell CRISPRi data, profiling the effects of the POU5F1 perturbation, have been deposited in GEO under the accession number GSE325993 [114].

All other datasets were previously published and reused for this study. The raw data for the primate brain snRNA-seq dataset are available at the Neuroscience Multi-omic Data Archive (NeMO) under the ID nemo:dat-net1412 [115] and the corresponding processed data can be downloaded from the NeMO Archive repository ”Great Ape MTG Analysis”. The ATAC-seq data of the cell lines H1c2, H2c1, C1c1 and C2c1 are available at ArrayExpress under the accession number E-MTAB-13373 [116]. The POU5F1 and NANOG ChIP-seq data are available at GEO under the accession numbers GSE69646 [117] and GSE99627 [118], respectively. The PAX6 ChIP-seq data are available at the Genome Sequence Archive (GSA) of the China National Center for Bioinformation (CNCB-NGDC) under the accession number CRA002482 [119]. The Hi-C data were published under the DOI 10.1101/2025.03.11.642620 (https://doi.org/10.1101/2025.03.11.642620) [61].

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 11 authors, 5 MeSH terms, 4 funders, 107 references.

Cite

This paper

Térmeg, A., Storozhuk, V., Kliesmete, Z., Edenhofer, F. C., Geuder, J., Dietl, T., Vieth, B., Janssen, P., Richter, D., Bonev, B., & Hellmann, I. (2026). CroCoNet: a framework for the quantitative comparison of gene regulatory networks across species. Genome biology, 27(1), 228. https://doi.org/10.1186/s13059-026-04152-5

BibTeX

@article{termeg2026croconet,
author = {Térmeg, Anita and Storozhuk, Vladyslav and Kliesmete, Zane and Edenhofer, Fiona C. and Geuder, Johanna and Dietl, Tamina and Vieth, Beate and Janssen, Philipp and Richter, Daniel and Bonev, Boyan and Hellmann, Ines},
title = {{CroCoNet: a framework for the quantitative comparison of gene regulatory networks across species}},
journal = {Genome biology},
year = {2026},
month = jul,
volume = {27},
number = {1},
pages = {228},
publisher = {BMC},
issn = {1474-7596},
doi = {10.1186/s13059-026-04152-5},
url = {https://doi.org/10.1186/s13059-026-04152-5},
pmid = {42458527},
pmcid = {PMC13371563}
}

RIS

TY - JOUR
AU - Térmeg, Anita
AU - Storozhuk, Vladyslav
AU - Kliesmete, Zane
AU - Edenhofer, Fiona C.
AU - Geuder, Johanna
AU - Dietl, Tamina
AU - Vieth, Beate
AU - Janssen, Philipp
AU - Richter, Daniel
AU - Bonev, Boyan
AU - Hellmann, Ines
TI - CroCoNet: a framework for the quantitative comparison of gene regulatory networks across species
T2 - Genome biology
J2 - Genome Biol
PY - 2026
DA - 2026/07/15
VL - 27
IS - 1
SP - 228
SN - 1474-7596
PB - BMC
DO - 10.1186/s13059-026-04152-5
UR - https://doi.org/10.1186/s13059-026-04152-5
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s13059-026-04152-5",
"type": "article-journal",
"title": "CroCoNet: a framework for the quantitative comparison of gene regulatory networks across species",
"container-title": "Genome biology",
"author": [
{
"family": "Térmeg",
"given": "Anita"
},
{
"family": "Storozhuk",
"given": "Vladyslav"
},
{
"family": "Kliesmete",
"given": "Zane"
},
{
"family": "Edenhofer",
"given": "Fiona C."
},
{
"family": "Geuder",
"given": "Johanna"
},
{
"family": "Dietl",
"given": "Tamina"
},
{
"family": "Vieth",
"given": "Beate"
},
{
"family": "Janssen",
"given": "Philipp"
},
{
"family": "Richter",
"given": "Daniel"
},
{
"family": "Bonev",
"given": "Boyan"
},
{
"family": "Hellmann",
"given": "Ines"
}
],
"container-title-short": "Genome Biol",
"volume": "27",
"issue": "1",
"page": "228",
"DOI": "10.1186/s13059-026-04152-5",
"PMID": "42458527",
"PMCID": "PMC13371563",
"ISSN": "1474-7596",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s13059-026-04152-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
15
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: Harmony, SingleCellExperiment, limma, 13 other tools, 3 references
[2] doi:10.1038/s41467-026-73305-8 [code]
Comparative analysis of the cellular landscape in mammalian striatum.
Journal: Nature communications
In common: Cell Ranger, Harmony, SingleCellExperiment, 11 other tools, 2 references
[3] doi:10.1038/s41586-026-10214-2 [code]
Multidimensional profiling of heterogeneity in supratentorial ependymomas.
Journal: Nature
In common: Cell Ranger, Harmony, SingleCellExperiment, 12 other tools
[4] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: BEDTools, Harmony, igraph, 11 other tools, 2 references
[5] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: BEDTools, Harmony, SingleCellExperiment, 12 other tools
[6] doi:10.1038/s41398-026-04200-5 [code]
Postmortem brain single-nucleus and bulk gene expression analyses identify shared and distinct abnormalities in bipolar disorder and major depressive disorder.
Journal: Translational psychiatry
In common: Harmony, SingleCellExperiment, limma, 9 other tools, 3 references
[7] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: SAMtools, limma, igraph, 12 other tools
[8] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: Harmony, SingleCellExperiment, igraph, 12 other tools
[9] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: Harmony, SingleCellExperiment, limma, 12 other tools
[10] doi:10.1038/s41586-026-10612-6 [code]
Acquired genetic and cell-state changes in IDH-mutant glioma progression.
Journal: Nature
In common: BEDTools, Harmony, SingleCellExperiment, 11 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.