OSCR

DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Methods › Weighted gene correlation network analysis (WGCNA) ↔ analysis/WGCNA/WGCNA_analysis.R, lines 1–72 · score 0.95 · soft thresholding powers, scale free topology, blockwiseModules, deepSplit, variable probes, WGCNA
  2. [2] § Methods › Cell enrichment analysis ↔ analysis/EWCE.R, lines 1–39 · score 0.92 · Read10X, SummarizedExperiment, colData, cell annotation, Seurat, GSE152058
  3. [3] § Methods › Illumina EPIC array profiling and data quality control ↔ QC/EPIC_array_QC.R, lines 215–304 · score 0.90 · wateRmelon, quantile normalize, dasen function, SNP probes, maximum correlation, beadcount
  4. [4] § Methods › Weighted gene correlation network analysis (WGCNA) ↔ analysis/WGCNA/Preprocessing.R, lines 72–111 · score 0.75 · Euclidean distance, variable probes, variance, PC, lowest, Outlier
  5. [5] § Results › Genes annotated to HD-associated modules have significantly enriched expression in disease affected neuronal subtypes in the striatum ↔ analysis/EWCE.R, lines 41–126 · score 0.70 · Cil Ependymal, Sec Ependymal, standard deviations, PV, interneuron, bootstrapped
  6. [6] § Results › DNA methylation variation in HD may be enriched in the vicinity of the HTT gene ↔ analysis/BrownsMethod.R, lines 1–59 · score 0.68 · C3orf35, genomic regions, chr3, chr4, GRK4, HTT
  7. [7] § Methods › Gene ontological enrichment analysis ↔ analysis/WGCNA/OntologyPathway.R, lines 41–82 · score 0.57 · KEGG terms, GO terms, gometh, Ontology, pathway, CpG
  8. [8] § Results › Highly connected probes show strong association with HD status in the striatum modules ↔ analysis/WGCNA/WGCNA_analysis.R, lines 319–379 · score 0.56 · hub genes, module membership, hub probes, MM
  9. [9] § Results › HD co-methylated networks are mostly independent of HD associated genetic variation ↔ analysis/BrownsMethod.R, lines 1–59 · score 0.55 · ANKRD34B, FAM193A, LETM1
  10. [10] § Methods › Association of modules with traits ↔ analysis/WGCNA/WGCNA_analysis.R, lines 220–273 · score 0.55 · module eigengene, module membership, Spearman, MM, traits, methylation
  11. [11] § Results › HD co-methylated networks are mostly independent of HD associated genetic variation ↔ analysis/WGCNA/fishers_gene_enrichment.R, the whole file · a weak match · score 0.54 · odds ratio, Fisher, CpGs, enrichment, modules, probes
  12. [12] § Methods › Epigenome-wide association study ↔ analysis/WGCNA/Preprocessing.R, lines 72–111 · score 0.52 · Principal component, PCs, prcomp, threshold
  13. [13] § Methods › Epigenome-wide association study ↔ QC/EPIC_array_QC.R, lines 320–371 · score 0.51 · CETS package, glia, neuronal, sex, cell, models
  14. [14] § Methods › Subjects and samples ↔ analysis/WGCNA/WGCNA_analysis.R, lines 1–72 · score 0.51 · CBB, LNDBB, MBB, OBB, Cambridge, London

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 379 lines · 14 KB · no license · 4 matches

  1. library(WGCNA)
  2. library("tidyverse")
  3. library(plyr)
  4. library(dplyr)
  5. library(stringr)
  6. library(corrplot)
  7. library(cowplot)
  8. library(parallel)
  9. library(lm.beta)
  10. library(utils)
  11. library(tibble)
  12. library(reshape2)
  13. library(gridExtra)
  14. library("ggplotify")
  15. ### load dat file containing only the most variable probes and with outlying samples removed
  16. load(file="dat.RData")
  17. ### load pheno file for trait association
  18. pheno <-read.csv("pheno.csv", header = T, stringsAsFactors = F, row.names = 1)
  19. ### Identify the soft threshold power ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  20. enableWGCNAThreads(8)
  21. powers = 1:20
  22. ### Call the network topology analysis function
  23. sft = pickSoftThreshold(dat, powerVector = powers, verbose = 5, networkType = "unsigned")
  24. ### plot results
  25. ## Scale-free topology fit index as a function of the soft-thresholding power
  26. plot(sft$fitIndices[,1], -sign(sft$fitIndices[,3])*sft$fitIndices[,2], xlab = "Soft Threshold (power)", ylab = "Scale Free Topology Model Fit, Unsigned R^2", type = "n", main = paste("Scale independence"),ylim=c(0.2,1))
  27. text(sft$fitIndices[,1], -sign(sft$fitIndices[,3])*sft$fitIndices[,2], labels= powers, cex = 0.8, col = "red")
  28. ##this line corresponds to using an R^2 cut off of h
  29. abline(h = 0.9, col = "red")
  30. abline(h = 0.8, col = "orange")
  31. ##Mean connectivity as
  32. plot(sft$fitIndices[,1], sft$fitIndices[,5], xlab = "Soft Threshold (power)", ylab = "Mean Connectivity", type = "n", main = paste("Mean connectivity"))
  33. text(sft$fitIndices[,1], sft$fitIndices[,5], labels = powers, cex=0.8, col = "red")
  34. ### Generate the modules ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  35. softPower = value;
  36. enableWGCNAThreads(32)
  37. bwnet = blockwiseModules(dat, maxBlockSize = 10000, power = softPower, TOMType = "unsigned", deepSplit = 0,
  38. minModuleSize = 100, reassignThreshold = 0, mergeCutHeight = 0.25, numericLabels = TRUE, saveTOMs= FALSE, verbose = 3)
  39. save(bwnet, file = "blockwiseMods.Rdata")
  40. ### Trait association ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  41. ### match pheno to samples in dat (i.e. remove outliers from pheno) and reduce to traits to test
  42. pheno <- pheno[match(rownames(dat), pheno_Striatum$Basename),c("Sample_ID", "Basename", "Phenotype", "Sex", "Age","prop","Institute_London", "Institute_Manchester", "Institute_Cambridge", "Institute_Oxford")]
  43. pheno<- pheno%>%
  44. rename("NeuN" = "prop",
  45. "LNDBB" = "Institute_London",
  46. "MBB" = "Institute_Manchester",
  47. "CBB" = "Institute_Cambridge",
  48. "OBB" = "Institute_Oxford")
  49. ####### Function to perform module trait associations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  50. # tVec: Vector of labels for trait assocation, must be in form binary or quantitative
  51. # tFile: Data frame with tVec labels as colnames (pheno file)
  52. # mFile: module file containing module eigengenes
  53. # rDat: "Raw" base gene expression / methylation matrix used to construct module
  54. traitModCor <- function(tVec,tFile,mFile,rDat){
  55. message(paste("You've started the function, at least"))
  56. # Define your variables
  57. nGenes = ncol(rDat)
  58. nSamples= nrow(rDat)
  59. moduleColors = labels2colors(mFile$colors)
  60. MEs0= moduleEigengenes(rDat, moduleColors)$eigengenes
  61. MEs = orderMEs(MEs0)
  62. message(paste("You've gotten to the pre-check"))
  63. #
  64. if(!identical(rownames(rDat),rownames(MEs))){
  65. stop("Raw data matrix and module eigengenes are discordant")
  66. }
  67. if(length(unique(tVec %in% colnames(tFile))) > 1 | unique(tVec %in% colnames(tFile)) == FALSE){
  68. stop("Label vector and trait file are discordant")
  69. }
  70. message(paste("You've gotten to the results file creation"))
  71. moduleTraitCor <- matrix(data = NA, ncol = length(tVec), nrow = ncol(MEs))
  72. colnames(moduleTraitCor)<- tVec
  73. rownames(moduleTraitCor) <- colnames(MEs)
  74. moduleTraitPvalue <- matrix(data = NA, ncol = length(tVec), nrow = ncol(MEs))
  75. colnames(moduleTraitPvalue)<- tVec
  76. rownames(moduleTraitPvalue) <- colnames(MEs)
  77. message(paste("You've gotten to the correlation loop"))
  78. for(t in tVec){
  79. if(length(unique(tFile[,t])) > 2){ # if Binary run spearman cor
  80. for(m in colnames(MEs)){
  81. try(res <-cor.test(MEs[,m] , as.numeric(tFile[,t]),method = "pearson"), silent = TRUE)
  82. if(class(res) != "try-error")
  83. moduleTraitCor[m,t]<-as.numeric(res$estimate)
  84. moduleTraitPvalue[m,t]<-res$p.value
  85. }
  86. } else { # do the same as above but if quantitative then run different pearson cor
  87. for(m in colnames(MEs)){
  88. try(res <-cor.test(MEs[,m] , as.numeric(tFile[,t]),method = "spearman"), silent = TRUE)
  89. if(class(res) != "try-error")
  90. moduleTraitCor[m,t]<-as.numeric(res$estimate)
  91. moduleTraitPvalue[m,t]<-res$p.value
  92. }
  93. }
  94. message(paste("Testing",t))
  95. }
  96. assign("moduleTraitCor", moduleTraitCor, envir = parent.frame() )
  97. assign("moduleTraitPvalue", moduleTraitPvalue, envir = parent.frame() )
  98. }
  99. names(bwnet)
  100. moduleColors = labels2colors(bwnet$colors)
  101. labs <- colnames(pheno[c(3:10)])
  102. traitModCor(tVec = labs, tFile = pheno, mFile = bwnet, rDat = dat)
  103. ### remove confounded modules (those associated with confounding variables)
  104. moduleTraitPvalue <- as.data.frame(moduleTraitPvalue)
  105. rem <- which(moduleTraitPvalue$Sex < 0.05 | moduleTraitPvalue$Age < 0.05 | moduleTraitPvalue$NeuN < 0.05|
  106. moduleTraitPvalue$LNDBB < 0.05 | moduleTraitPvalue$MBB < 0.05 | moduleTraitPvalue$CBB < 0.05 | moduleTraitPvalue$OBB < 0.05 | rownames(moduleTraitPvalue) == "MEgrey")
  107. #Remove confounded mocules
  108. moduleTraitPvalue <- moduleTraitPvalue[-rem,]
  109. moduleTraitCor<- moduleTraitCor[-rem,]
  110. moduleTraitPvalue <- as.matrix(moduleTraitPvalue)
  111. ### create Heatmap to display correlations and their p-values ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  112. textMatrix = paste(signif(moduleTraitCor, 2), "\n(",
  113. signif(moduleTraitPvalue, 1), ")", sep = "");
  114. dim(textMatrix) = dim(moduleTraitCor)
  115. par(mar = c(3,6.5,3,1));
  116. # Display the correlation values within a heatmap plot
  117. labeledHeatmap(Matrix = moduleTraitCor,
  118. xLabels = labs,
  119. yLabels = rownames(moduleTraitCor),
  120. ySymbols = rownames(moduleTraitCor),
  121. colorLabels = FALSE,
  122. colors = blueWhiteRed(50),
  123. textMatrix = textMatrix,
  124. setStdMargins = FALSE,
  125. cex.text = 0.6,
  126. cex.lab = 0.6,
  127. zlim = c(-1,1),
  128. main = paste("Module-trait relationships"))
  129. ### t.test ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  130. ### get the modules that are not assocaited with coVars
  131. modKeep <- rownames(moduleTraitPvalue)
  132. MEs0= moduleEigengenes(dat, moduleColors)$eigengenes
  133. MEs = orderMEs(MEs0)
  134. dim(MEs)
  135. t(MEs)->MET
  136. moduleCounts <- as.data.frame(table(moduleColors))
  137. ### t test function
  138. ttest <- function( row, Diag){
  139. ttest.result <- t.test( row ~ factor(Diag) )
  140. return(ttest.result$p.value)
  141. }
  142. Diag <- factor(pheno$Phenotype)
  143. diagtest <- t(apply( MET, 1,ttest, Diag)) ### apply t-test function
  144. diagtest <- t(diagtest)
  145. ###filter out modules associated with coVars
  146. diagtest <- as.data.frame(diagtest[row.names(diagtest) %in% modKeep, ], row.names = modKeep,)
  147. diagtest <- as.matrix(diagtest)
  148. pOrder <- order(diagtest[,1])
  149. diagtest2 <- as.data.frame(diagtest[pOrder,])
  150. head(diagtest2)
  151. colnames(diagtest2)[colnames(diagtest2) == 'diagtest[pOrder, ]'] <- 'pValues'
  152. as.data.frame(MET)->MET
  153. as.matrix(MET['ME_selected_moule_colour',])-> ME_module
  154. t(ME_module)->ME_module
  155. ### box plot of module eigengenes
  156. ggplot(pheno, aes(x = as.factor(Phenotype), y = ME_module, colour = as.factor(Phenotype),fill = as.factor(Phenotype)))+
  157. geom_boxplot(color = "black",lwd = 0.8, outlier.shape = NA)+
  158. geom_jitter(width = 0.3, shape = 21, colour = "black", fill = "black")+
  159. theme_bw()+
  160. ylab("Module eigengene")+
  161. xlab(NULL)+
  162. ggtitle("red")+
  163. theme(legend.position = "none")+
  164. theme(plot.title = element_text(hjust = 0.5, size = 18),
  165. axis.title.y = element_text(size = 16),
  166. axis.text.x = element_text(size = 14),
  167. axis.text.y = element_text(size = 14))+
  168. scale_x_discrete(labels=c("1" = "Control",
  169. "2" = "HD"))
  170. ### module memebership analysis ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  171. ##### change bwnet.XXX as necessary
  172. ##### change 'dat' as necesssary
  173. ##### change pheno as necessary
  174. ##### change module of interest as necessary
  175. moduleColors = labels2colors(bwnet$colors)
  176. table(moduleColors)
  177. MEs0= moduleEigengenes(dat, moduleColors)$eigengenes
  178. MEs = orderMEs(MEs0)
  179. rownames(MEs) == rownames(dat)
  180. weight = as.data.frame(as.numeric(pheno$Phenotype));
  181. names(weight) = "Phenotype"
  182. # names (colors) of the modules
  183. modNames = substring(names(MEs), 3)
  184. geneModuleMembership = as.data.frame(cor(dat, MEs, use = "p"));
  185. MMPvalue = as.data.frame(corPvalueStudent(as.matrix(geneModuleMembership), nrow(pheno)));
  186. names(geneModuleMembership) = paste("MM", modNames, sep="");
  187. names(MMPvalue) = paste("p.MM", modNames, sep="");
  188. geneTraitSignificance = as.data.frame(cor(dat, weight, use = "p",method = "spearman"));
  189. GSPvalue = as.data.frame(corPvalueStudent(as.matrix(geneTraitSignificance), nrow(pheno)));
  190. names(geneTraitSignificance) = paste("GS.", names(weight), sep="");
  191. names(GSPvalue) = paste("p.GS.", names(weight), sep="");
  192. head(MMPvalue)
  193. dg <- which(moduleColors == "MEcol") #### take module of interest
  194. dg_names<- names(bwnet$colors[dg])
  195. all_names <- names(bwnet$colors)
  196. epicManifest<-read.csv("MethylationEPIC_v-1-0_B4.csv", skip = 7)
  197. rownames(epicManifest)<-epicManifest[,1]
  198. plotMM <- data.frame(row.names = rownames(GSPvalue),moduleMembership = geneModuleMembership$MMcol,MMPvalue = MMPvalue$p.MMcol,
  199. geneCor = geneTraitSignificance$GS.Phenotype,geneP = GSPvalue$p.GS.Phenotype,epicManifest[colnames(dat[,all_names]),])
  200. ### make summary with rank for indivdual modul then add with EPIC annotation
  201. plotMM <- plotMM[dg_names,] ### cpg names in the module dg names is first generated above for the trait associatons
  202. #### add in MMrank and GSrank before EPIC annotation
  203. plotMM <- plotMM %>%
  204. add_column(MMrank = rank(abs(plotMM$moduleMembership)),
  205. GSrank = rank(-log10(plotMM$geneP)),
  206. .before = "IlmnID")
  207. ### summative rank
  208. sumrank = plotMM$MMrank + plotMM$GSrank
  209. plotMM <- plotMM %>%
  210. add_column(sumrank = sumrank,
  211. .before = "IlmnID")
  212. dim(plotMM)
  213. ### run correlation test
  214. est <- cor.test(abs(plotMM$moduleMembership),-log10(plotMM$geneP))$estimate ### negative log ten
  215. pval <- cor.test(abs(plotMM$moduleMembership),-log10(plotMM$geneP))$p.value
  216. MMPScorrelation<-data.frame(est,pval)
  217. ### plot ###
  218. ggplot(plotMM, aes(x = abs(moduleMembership), y = -log10(geneP), col = sumrank))+
  219. geom_point()+
  220. geom_hline(yintercept = -log10(0.05))+
  221. geom_vline(xintercept = 0.8)+
  222. geom_smooth(method = "lm")+
  223. theme_bw()+
  224. ylab("Probe Significance for HD (-log10(p))")+
  225. xlab("Module Membership")+
  226. ggtitle("lavenderblush3", subtitle = "corr = 0.458, p =1.91e-12")+
  227. theme(plot.title = element_text( size = 22),
  228. plot.subtitle = element_text(size = 18),
  229. axis.title.y = element_text(size = 18),
  230. axis.title.x = element_text(size = 18),
  231. axis.text.x = element_text(size = 16),
  232. axis.text.y = element_text(size = 16))+
  233. theme(legend.position = c(0.1, 0.85),
  234. legend.text = element_text(size = 14),
  235. legend.title = element_text(size = 14))+
  236. scale_color_gradient(low="palevioletred2", high="palevioletred4")
  237. #### Hub genes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
  238. top_MM <- plotMM[order(plotMM[, 7],decreasing = TRUE), ] ##### orders membership summary file by MM
  239. ### select hub probes
  240. hub_probes <- subset(top_MM, abs(moduleMembership) > 0.8 & geneP < 0.05)
  241. testDF <- hub_probes
  242. testDF$UCSC_RefGene_Name <- as.character(testDF$UCSC_RefGene_Name)
  243. #### seperate out UCSC multiple gene names into seprate columns
  244. pb <- txtProgressBar(min = 0, # Minimum value of the progress bar
  245. max = length(rownames(testDF)), # Maximum value of the progress bar
  246. style = 3, # Progress bar style (also available style = 1 and style = 2)
  247. width = 50, # Progress bar width. Defaults to getOption("width")
  248. char = "=")
  249. for(i in 1:length(rownames(testDF))){
  250. cpgIndex <- rownames(testDF)[i]
  251. test1 <- as.character(testDF[cpgIndex,"UCSC_RefGene_Name"])
  252. if(str_detect(test1,";") == TRUE){
  253. strIndex <- as.data.frame(str_locate_all(test1,";"))$end
  254. length(strIndex)
  255. strIndex <- c(0,strIndex,str_length(test1))
  256. storage <- c()
  257. for(y in strIndex){
  258. if(which(strIndex == y) == 2){
  259. storage <- append(storage,str_sub(test1,
  260. start = strIndex[which(strIndex == y) -1],
  261. end = y - 1))
  262. }else if(y >0 & y != max(strIndex)){
  263. storage <-append(storage,str_sub(test1,
  264. start = strIndex[which(strIndex == y) -1]+ 1,
  265. end = y - 1))
  266. }else if(y == max(strIndex)){
  267. storage <-append(storage,str_sub(test1,
  268. start = strIndex[which(strIndex == y) -1]+ 1))
  269. }
  270. }
  271. sumGene <- as.character(unique(storage))
  272. if(length(sumGene) == 1){
  273. testDF[cpgIndex,"UCSC1"] <- sumGene
  274. }else if(length(sumGene) == 2){
  275. testDF[cpgIndex,c("UCSC1","UCSC2")] <- sumGene[1:2]
  276. }else if(length(sumGene) > 2){
  277. testDF[cpgIndex,c("UCSC1","UCSC2","UCSC3")] <- sumGene[1:3]
  278. }
  279. } else if(str_detect(test1,";") == FALSE){
  280. testDF[cpgIndex,"UCSC1"] <- as.character(testDF[cpgIndex,"UCSC_RefGene_Name"])
  281. }
  282. setTxtProgressBar(pb, i)
  283. }
  284. write.table(testDF, file = "hub_probes.txt")
  285. probes_per_gene <-table(testDF$UCSC1)
  286. probes_per_gene_sort1 <- probes_per_gene[order(probes_per_gene)]
  287. write.table(probes_per_gene_sort1, file = "hub_probes_per_gene.txt")

WGCNA_analysis.R at commit 02aab78, no license · at the source

Overview

Authors: Gregory Wheildon1, Adam R. Smith1, Luke Weymouth1, Joshua Harvey1, Morteza Kouhsar1, Lachlan F. MacBean1, Claire Troakes2, Ehsan Pishva1,3, Rebecca G. Smith1, Katie Lunnon1
ORCID iDs: Gregory Wheildon
  1. Department of Clinical and Biomedical Sciences, Faculty of Health and Life Sciences, University of Exeter,Exeter, UK
  2. Institute of Psychiatry, Psychology & Neuroscience (IoPPN), King’s College London,De Crespigny Park, London, UK
  3. Department of Psychiatry and Neuropsychology, Mental Health and Neuroscience Research Institute (MHeNS), Maastricht University,Maastricht, The Netherlands
Institutions: University of Exeter (United Kingdom); King's College London (United Kingdom); Maastricht University (Netherlands)
Journal: Clinical epigenetics, volume 18, issue 1, article 92
Dates: received 16 May 2025; accepted 3 February 2026; published online 26 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s13148-026-02082-4 · PMID 42185880 · PMCID PMC13202909 · OpenAlex W7162316320
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), other condition (population), systems (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Connectivity, Spectral & time-frequency
Keywords: Brain, Cerebellum, DNA methylation, Epigenetics, Epigenome-wide association study (EWAS), Entorhinal cortex, Huntington’s disease (HD), Illumina infinium methylation EPIC v1.0 array, Striatum, Weighted gene correlation network analysis (WGCNA)
MeSH: Corpus Striatum*, DNA Methylation*, Huntington Disease*, Adult, Cerebellum, CpG Islands, Entorhinal Cortex, Epigenesis, Genetic, Female, Humans, Male, Middle Aged, Trinucleotide Repeat Expansion (* major topic)
Topic: Genetic Neurodegenerative Diseases (Cellular and Molecular Neuroscience, Neuroscience), according to OpenAlex
Funding: BRACE; UK Medical Research Council (MR/Y014685/1)
Citations: not cited yet (Europe PMC); 91 references in the paper

Abstract

Background: Huntington’s disease is caused by a trinucleotide CAG repeat expansion in the HTT gene. Despite displaying autosomal dominance, phenotypic variation exists amongst mutation carriers, in particular relating to the age that symptoms first occur. This variation is primarily driven by an inverse relationship between CAG expansion size and age of symptom onset. However, the majority of variation in age of onset that is independent of CAG repeat length is thought to be driven by environmental influences. Since DNA methylation can be altered by environmental factors, and as methylomic variation is reported in other neurodegenerative diseases, it may offer a potential mechanism underlying disease manifestation.

Results: We utilized the Illumina EPIC v1 methylation array to profile DNA methylation in 120 samples, including three distinct brain regions (striatum, entorhinal cortex and cerebellum) in 20 Huntington’s disease and 22 control donors. We identified seven Bonferroni-significant differentially methylated CpGs within the striatum along with 27 differentially methylated regions, annotated to genes involved in physiological processes known to be disrupted in HD such as the urea cycle and metabolism. Weighted gene correlation network analysis identified modules of co-methylated CpGs that were associated with Huntington’s disease, with ontological analyses showing enrichment in disease relevant processes. Furthermore, integration of single-nuclei RNA sequencing data highlighted that genes annotated to these modules are enriched in striatal spiny projection neurons, the primary cell types affected in the disease.

Conclusions: Here, we present the first epigenome-wide association study of Huntington’s disease conducted in the striatum, the primary region of neuropathology, along with matched entorhinal cortex and cerebellum on the Illumina EPIC v1 array. Our results suggest that DNA methylation is altered at loci associated with Huntington’s disease in disease relevant regions and cell types and strengthens evidence for areas of potential therapeutic intervention.

Supplementary Information: The online version contains supplementary material available at 10.1186/s13148-026-02082-4.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

UoE-Dementia-Genomics/HD-DNAmeth

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 02aab7843554df7bcb8ea700c14fc87cccb42edf, 14 March 2025
Languages: R (9)
Size: 10 files, 9 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: tidyverse (9 files), cowplot (4 files), WGCNA (3 files), ggplot2 (2 files), clusterProfiler (1 file), reshape2 (1 file), Seurat (1 file), SingleCellExperiment (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
10 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 9 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The datasets generated and analyzed during the current study are available in the Gene Expression Omnibus (GEO) repository (GSE297210). Analytical scripts used in this manuscript are available at https://github.com/UoE-Dementia-Genomics/HD-DNAmeth.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 10 keywords, 13 MeSH terms, 2 funders, 87 references.

Cite

This paper

Wheildon, G., Smith, A. R., Weymouth, L., Harvey, J., Kouhsar, M., MacBean, L. F., Troakes, C., Pishva, E., Smith, R. G., & Lunnon, K. (2026). DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum. Clinical epigenetics, 18(1), 92. https://doi.org/10.1186/s13148-026-02082-4

BibTeX

@article{wheildon2026dna,
author = {Wheildon, Gregory and Smith, Adam R. and Weymouth, Luke and Harvey, Joshua and Kouhsar, Morteza and MacBean, Lachlan F. and Troakes, Claire and Pishva, Ehsan and Smith, Rebecca G. and Lunnon, Katie},
title = {{DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum}},
journal = {Clinical epigenetics},
year = {2026},
month = may,
volume = {18},
number = {1},
pages = {92},
publisher = {BMC},
issn = {1868-7075},
doi = {10.1186/s13148-026-02082-4},
url = {https://doi.org/10.1186/s13148-026-02082-4},
pmid = {42185880},
pmcid = {PMC13202909}
}

RIS

TY - JOUR
AU - Wheildon, Gregory
AU - Smith, Adam R.
AU - Weymouth, Luke
AU - Harvey, Joshua
AU - Kouhsar, Morteza
AU - MacBean, Lachlan F.
AU - Troakes, Claire
AU - Pishva, Ehsan
AU - Smith, Rebecca G.
AU - Lunnon, Katie
TI - DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum
T2 - Clinical epigenetics
J2 - Clin Epigenetics
PY - 2026
DA - 2026/05/26
VL - 18
IS - 1
SP - 92
SN - 1868-7075
PB - BMC
DO - 10.1186/s13148-026-02082-4
UR - https://doi.org/10.1186/s13148-026-02082-4
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s13148-026-02082-4",
"type": "article-journal",
"title": "DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum",
"container-title": "Clinical epigenetics",
"author": [
{
"family": "Wheildon",
"given": "Gregory"
},
{
"family": "Smith",
"given": "Adam R."
},
{
"family": "Weymouth",
"given": "Luke"
},
{
"family": "Harvey",
"given": "Joshua"
},
{
"family": "Kouhsar",
"given": "Morteza"
},
{
"family": "MacBean",
"given": "Lachlan F."
},
{
"family": "Troakes",
"given": "Claire"
},
{
"family": "Pishva",
"given": "Ehsan"
},
{
"family": "Smith",
"given": "Rebecca G."
},
{
"family": "Lunnon",
"given": "Katie"
}
],
"container-title-short": "Clin Epigenetics",
"volume": "18",
"issue": "1",
"page": "92",
"DOI": "10.1186/s13148-026-02082-4",
"PMID": "42185880",
"PMCID": "PMC13202909",
"ISSN": "1868-7075",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s13148-026-02082-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
26
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41380-026-03686-1 [code]
Early oligodendrocyte dysfunction signature in Alzheimer's disease: Insights from DNA methylomics and transcriptomics.
Journal: Molecular psychiatry
In common: WGCNA, SingleCellExperiment, Seurat, 2 other tools, genetics / omics, 8 references
[2] doi:10.1186/s13073-026-01698-8 [code]
From aging to Alzheimer's disease: concordant brain DNA methylation changes in late life.
Journal: Genome medicine
In common: ggplot2, tidyverse, genetics / omics, 11 references
[3] doi:10.1038/s41467-026-73305-8 [code]
Comparative analysis of the cellular landscape in mammalian striatum.
Journal: Nature communications
In common: WGCNA, SingleCellExperiment, clusterProfiler, 5 other tools, genetics / omics, 4 references
[4] doi:10.1038/s41467-026-68864-9 [code]
Integrative epigenomic landscape of Alzheimer's Disease brains reveals oligodendrocyte molecular perturbations associated with tau.
Journal: Nature communications
In common: genetics / omics, 10 references
[5] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: WGCNA, clusterProfiler, cowplot, 3 other tools, genetics / omics, other condition, 3 references
[6] doi:10.1186/s13024-026-00960-2 [code]
Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction.
Journal: Molecular neurodegeneration
In common: systems, genetics / omics, other condition, 8 references
[7] doi:10.1038/s41398-026-04200-5 [code]
Postmortem brain single-nucleus and bulk gene expression analyses identify shared and distinct abnormalities in bipolar disorder and major depressive disorder.
Journal: Translational psychiatry
In common: WGCNA, SingleCellExperiment, clusterProfiler, 5 other tools, genetics / omics, 1 reference
[8] doi:10.1038/s41593-026-02367-0 [code]
A reproducible three-dimensional model of human brain tissue to investigate physiological and disease-associated microglia phenotypes.
Journal: Nature neuroscience
In common: WGCNA, SingleCellExperiment, clusterProfiler, 5 other tools, 1 reference
[9] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: WGCNA, SingleCellExperiment, clusterProfiler, 5 other tools, genetics / omics, other condition
[10] doi:10.1038/s44318-026-00806-z [code]
Interspecific diversity in the neuronal composition of the mammalian cortex arises from heterochrony in neurogenesis.
Journal: The EMBO journal
In common: WGCNA, SingleCellExperiment, clusterProfiler, 5 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.