Characterising the motif composition and allele length distribution of <i>ZFHX3</i> GGC repeat expansions in amyotrophic lateral sclerosis
The 1 match
- [1] § Results › Variability of repeat motif compositions and configurations in ZFHX3 ↔ 01_fig3_motif_composition.Rmd, lines 132–199 · score 0.58 · pure glycine, GGC motifs, motif composition, AGT, GAC, GGT
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
R Markdown · 352 lines · 9.1 KB · no license · 1 match
- ---
- title: "01_fig3_motif_composition"
- author: "Zoe Zussa"
- date: "2026-02-09"
- output: html_document
- ---
- This script is to recreate figure 3 from the publication "Characterising the motif composition and allele length distribution of ZFHX3 GGC repeat expansions in amyotrophic lateral sclerosis" by Zussa et al., 2026.
- Data to replicate this analysis is titled
- ZFHX3_Genotypes_Motif_Comp_Australian_ALS_2026-03-09.csv
- ZFHX3_Genotypes_Motif_Comp_MinE_ALS_Controls_2026-04-08.csv
- Load packages
- ```{r}
- library(tidyr)
- library(tidyverse)
- library(dplyr)
- library(data.table)
- library(plotrix) # needed for piecharts
- ```
- Load in data files
- ```{r}
- # Australian ALS motif compositions
- mq_motif<- fread("/Users/MQ10007635/Desktop/ZFHX3_Genotypes_Motif_Comp_Australian_ALS_2026-03-09.csv")
- # project MinE control motif compositions
- mine_motif<- fread("/Users/MQ10007635/Desktop/ZFHX3_Genotypes_Motif_Comp_MinE_ALS_Controls_2026-04-08.csv")
- ```
- Preparing data for plotting
- ```{r}
- # subset project MinE for controls with motif composition
- mine_motif <- subset(mine_motif, !is.na(motif_comp_allele_1))
- # pivot dfs
- mq_motif <- mq_motif %>%
- pivot_longer(
- cols = c(rep_allele_1, rep_allele_2, motif_comp_allele_1, motif_comp_allele_2),
- names_to = c(".value", "allele"),
- names_pattern = "(rep|motif_comp)_allele_(.)"
- ) %>%
- rename(seq = motif_comp) %>%
- select(rep, seq)
- mine_motif <- mine_motif %>%
- pivot_longer(
- cols = c(rep_allele_1, rep_allele_2, motif_comp_allele_1, motif_comp_allele_2),
- names_to = c(".value", "allele"),
- names_pattern = "(rep|motif_comp)_allele_(.)"
- ) %>%
- rename(seq = motif_comp) %>%
- select(rep, seq)
- # count how many times the motif composition occurs
- # counting case sequences
- counts_cases <- mq_motif %>%
- count(seq) %>%
- rename(seq_comp = seq, cases = n)
- # counting control sequences
- counts_controls <- mine_motif %>%
- count(seq) %>%
- rename(seq_comp = seq, controls = n)
- # combining counts
- counts_seq_comp <- full_join(
- counts_cases,
- counts_controls,
- by = "seq_comp"
- ) %>%
- mutate(
- cases = ifelse(is.na(cases), 0, cases),
- controls = ifelse(is.na(controls), 0, controls)
- )
- # get proportions of each motif per motif comp for pie charts
- # cases
- motif_list_cases <- mq_motif %>%
- # extract motifs structured like GGC(4)
- mutate(parts = str_extract_all(seq, "[A-Z]+\\(\\d+\\)")) %>%
- unnest(parts) %>%
- # separate motif and count
- extract(
- parts,
- into = c("motif", "count"),
- regex = "([A-Z]+)\\((\\d+)\\)",
- convert = TRUE
- ) %>%
- group_by(seq, motif) %>%
- # getting repeat size from count
- summarise(count = sum(count), .groups = "drop") %>%
- group_by(seq) %>%
- mutate(prop = count / sum(count)) %>%
- summarise(motifs = list(setNames(prop, motif))) %>%
- deframe()
- # extract out compositions without GGC motif for second column of pie charts
- motif_list_cases_noGGC <- lapply(motif_list_cases, function(x) {
- x2 <- x[names(x) != "GGC"]
- if(length(x2) == 0) return(NULL)
- x2 / sum(x2)
- })
- # controls
- motif_list_controls <- mine_motif %>%
- # extract motifs structured like GGC(4)
- mutate(parts = str_extract_all(seq, "[A-Z]+\\(\\d+\\)")) %>%
- unnest(parts) %>%
- # separate motif and count
- extract(
- parts,
- into = c("motif", "count"),
- regex = "([A-Z]+)\\((\\d+)\\)",
- convert = TRUE
- ) %>%
- group_by(seq, motif) %>%
- # getting repeat size from count
- summarise(count = sum(count), .groups = "drop") %>%
- group_by(seq) %>%
- mutate(prop = count / sum(count)) %>%
- summarise(motifs = list(setNames(prop, motif))) %>%
- deframe()
- # extract out compositions without GGC motif for second column of pie charts
- motif_list_controls_noGGC <- lapply(motif_list_controls, function(x) {
- x2 <- x[names(x) != "GGC"]
- if(length(x2) == 0) return(NULL)
- x2 / sum(x2)
- })
- # merging case and control
- all_seq <- counts_seq_comp$seq_comp
- # all motifs (first pie chart column)
- motif_list_all <- lapply(all_seq, function(s) {
- if(!is.null(motif_list_cases[[s]])) {
- motif_list_cases[[s]]
- } else if(!is.null(motif_list_controls[[s]])) {
- motif_list_controls[[s]]
- } else {
- NULL
- }
- })
- names(motif_list_all) <- all_seq
- # all motifs excluding GGC (second pie chart column)
- motif_list_all_noGGC <- lapply(all_seq, function(s) {
- if(!is.null(motif_list_cases_noGGC[[s]])) {
- motif_list_cases_noGGC[[s]]
- } else if(!is.null(motif_list_controls_noGGC[[s]])) {
- motif_list_controls_noGGC[[s]]
- } else {
- NULL
- }
- })
- names(motif_list_all_noGGC) <- all_seq
- # extracting repeat sizes from cases and controls
- repeat_all <- bind_rows(
- mq_motif %>% distinct(seq, rep),
- mine_motif %>% distinct(seq, rep)
- ) %>% distinct(seq, rep)
- repeat_size <- repeat_all$rep[
- match(all_seq, repeat_all$seq)]
- stopifnot(!any(is.na(repeat_size)))
- # ordering motif compositions by repeat size and then total counts
- ord <- order(
- repeat_size,
- counts_seq_comp$cases + counts_seq_comp$controls)
- counts_seq_comp <- counts_seq_comp[ord, ]
- seq_comp_names <- counts_seq_comp$seq_comp
- repeat_size <- repeat_size[ord]
- motif_list_all <- motif_list_all[seq_comp_names]
- motif_list_all_noGGC <- motif_list_all_noGGC[seq_comp_names]
- # identifying pure glycine sequences
- seqs_only_GGC_GGT <- sapply(motif_list_all, function(x) {
- non_canonical <- x[names(x) %in% c("AGT","AGC","GAC","AAT")]
- all(non_canonical == 0 | is.na(non_canonical))})
- # adding in asterisks to sequences extracted above (assists when plotting)
- seq_comp_names_marked <- seq_comp_names
- seq_comp_names_marked[seqs_only_GGC_GGT] <- paste0("*", seq_comp_names_marked[seqs_only_GGC_GGT])
- names_to_plot <- seq_comp_names_marked
- names_to_plot[startsWith(names_to_plot, "*")] <- ""
- ```
- Plotting figure 3
- ```{r}
- # log scale for counts
- plot_matrix <- rbind(
- Cases = log10(counts_seq_comp$cases + 1),
- Controls = log10(counts_seq_comp$controls + 1))
- # setting pie chart colous
- motif_colors <- c(
- GGC = "#AAAEB0",
- GGT = "#4D6291",
- AGT = "#9A68A4",
- AGC = "#9D2D52",
- GAC = "#006E4A",
- AAT = "#4F1259"
- )
- # starting plot
- pdf("ZFHX3_motif_comp.pdf", width = 15, height = 18)
- par(mar = c(6,40,1,1))
- # replace 0 with NA so 0 values are not plotted
- # this gets rid of the weird line when values are equal to 0
- plot_matrix[plot_matrix == 0] <- NA
- # define consistent spacing
- gap_unit <- max(colSums(plot_matrix, na.rm = TRUE)) * 0.035
- # column positions
- pie_x1 <- -3 * gap_unit
- pie_x2 <- -2 * gap_unit
- repeat_x <- -1 * gap_unit
- # grouped bar chart with cases and controls
- bp <- barplot(
- plot_matrix,
- horiz = TRUE,
- beside = TRUE,
- names.arg = names_to_plot,
- las = 1,
- space = c(0,1),
- col = c("hotpink4","plum3"),
- cex.names = 1,
- xaxt = "n",
- xlim = c(
- pie_x1 - gap_unit *0.45,
- max(plot_matrix, na.rm = TRUE) * 1.1
- )
- )
- # adding x axis labels
- tick_counts <- c(0,1,2,5,10,20,50,100,500,1500) # manual scale
- axis(1, at = log10(tick_counts + 1), labels = tick_counts)
- # centering x axis label to bars not entire plot window
- usr <- par("usr") # returns plot boundaries
- x_center <- mean(c(0, usr[2]))# getting mean of max number
- mtext(
- "Counts of motif compositions (log)",
- side = 1,
- line = 3,
- at = x_center
- )
- # getting midpoint for each grouped bar
- bp_mid <- colMeans(bp)
- # bolding pure glycine compositions
- asterisk_idx <- which(startsWith(seq_comp_names_marked, "*"))
- normal_label_x <- par("usr")[1] -
- 0.01 * max(colSums(plot_matrix, na.rm = TRUE))
- text(
- x = normal_label_x,
- y = bp_mid[asterisk_idx],
- labels = seq_comp_names_marked[asterisk_idx],
- pos = 2,
- xpd = TRUE,
- font = 2,
- cex = 1,
- adj = 0
- )
- # plotting pie charts using midpoint for grouped bars
- for(i in seq_along(motif_list_all)){
- motifs_all <- motif_list_all[[i]]
- motifs_noGGC <- motif_list_all_noGGC[[i]]
- if(!is.null(motifs_all) && length(motifs_all) > 0){
- floating.pie(
- xpos = pie_x1,
- ypos = bp_mid[i],
- x = motifs_all,
- col = motif_colors[names(motifs_all)],
- radius = 0.090,
- startpos = 0
- )
- }
- if(!is.null(motifs_noGGC) && length(motifs_noGGC) > 0){
- floating.pie(
- xpos = pie_x2,
- ypos = bp_mid[i],
- x = motifs_noGGC,
- col = motif_colors[names(motifs_noGGC)],
- radius = 0.090,
- startpos = 0
- )
- }
- }
- # printing repeat sizes next to pie charts
- repeat_x <- mean(c(pie_x2, 0))
- text(x = repeat_x, y = bp_mid, labels = repeat_size, cex = 1.1, adj = 0.5)
- # setting motif legend
- motif_aas <- c(GGC="Gly", GGT="Gly", AGT="Ser", AGC="Ser", GAC="Asp", AAT="Asn")
- legend_labels <- paste0(names(motif_colors), " (", motif_aas[names(motif_colors)], ")")
- par(xpd = NA) # allow drawing in margins
- # plotting motif legend
- legend(
- x = par("usr")[1] - 0.6 * max(plot_matrix, na.rm = TRUE), # shift left
- y = min(bp) - 1.7, # shift down under y-axis labels
- legend = legend_labels,
- fill = motif_colors,
- title = "Motifs",
- bty = "o",
- cex = 1,
- ncol = 3,
- x.intersp = 0.2,
- y.intersp = 0.9
- )
- # plotting case and control figure legend
- legend(
- x = par("usr")[2] - 0.05 * diff(par("usr")[1:2]), # slightly left from right edge
- y = par("usr")[4] - 0.05 * diff(par("usr")[3:4]), # slightly down from top edge
- legend = c("Case", "Control"),
- fill = c("hotpink4", "plum3"),
- bty = "o",
- cex = 1,
- xjust = 1, # right-align the legend box at x
- yjust = 1 # top-align the legend box at y
- )
- dev.off()
- ```
01_fig3_motif_composition.Rmd at commit 494d51c, no license · at the source
Overview
- Macquarie University Motor Neuron Disease Research Centre, Faculty of Medicine, Health and Human Sciences, Macquarie University, Sydney, NSW, Australia
- Department of Neurology, UMC Utrecht Brain Center, Utrecht University, 3584 CX, Utrecht, The Netherlands
- Macquarie University Health Neurology, Macquarie University, Sydney, NSW, Australia
- Department of Neuropathology, The University of Sydney, Sydney, NSW, Australia
- Molecular Medicine Laboratory, Concord Repatriation General Hospital, Sydney, NSW 2139, Australia
- Northcott Neuroscience Laboratory, ANZAC Research Institute, Sydney, NSW 2139, Australia
- Neuroscience Research Australia, and the University of NSW, Sydney, NSW, Australia
Abstract
Background and objectives: A pathogenic GGC repeat expansion in the zinc finger homeobox 3 (ZFHX3) gene, encoding a pure polyglycine tract, is the cause of spinocerebellar ataxia type 4 (SCA4). Intermediate expansions of other SCA loci contribute to the risk of amyotrophic lateral sclerosis (ALS), a fatal neurodegenerative disease involving the progressive loss of motor neurons. There is increasing awareness of the role of short tandem repeat (STR) motif composition and configuration in disease pathogenicity. Given the genetic pleiotropy between ALS and SCA, this study aimed to evaluate whether ZFHX3 GGC expansions were associated with ALS and to characterise repeat motif composition.
Methods: ExpansionHunter v5 was used to genotype ZFHX3 GGC repeat sizes in short-read whole genome sequencing data from people with ALS and healthy controls of European ancestry. Repeat sizes were visually inspected using REViewer v2. Repeat motif configurations of Australian ALS cases and healthy controls were manually derived from REViewer images. Receiver operating characteristic (ROC) curve analysis and Youden’s J statistic were performed to find a candidate repeat size threshold for association testing. Fisher’s exact tests were performed to evaluate the associations of repeat size and motif composition with disease status.
Results: Analysis of 5,785 people with ALS and 7,982 healthy controls found no association between ZFHX3 GGC repeat expansions and disease risk. Fifty unique repeat motif compositions were identified across 802 people with ALS and 800 healthy controls. Of these, eleven distinct configurations coded a pure polyglycine tract which, when expanded, is canonical to SCA4, though no association with ALS was found.
Discussion: Although no association was observed between ZFHX3 GGC repeat expansions and ALS, this study established the dynamic nature of ZFHX3 repeat motif composition and configuration. Unique motif compositions were identified both within and between repeat sizes, including the presence of pure polyglycine repeats. Consideration of repeat motif composition and configuration, in addition to repeat allele length, may be important for assessing neurodegenerative disease risk.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.
mq-mnd/grp_williams/zfhx3_analysis_publication_2026
494d51cde152943bb1f3fb9b34cc82f79d9ec34f, 8 April 2026Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
3 files
- 01_fig3_motif_compositio
n.Rmd , R, 352 lines, 1 match - 02_efig1_roc.Rmd, R, 111 lines
- README.md, Text, 23 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 2 scripts, each with its path and the digest of its content;
- 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:18931240, at Zenodo; found in “Data availability”
Data availability
Individual-level ZFHX3 repeat allele size data for 5,785 people with ALS and 7,982 healthy controls plus individual-level motif composition data for the subset of 802 Australian cases with ALS and 800 healthy controls is available at Zenodo (DOI: 10.5281/
Code written in R is available in R Markdown workbooks in a GitLab repository: https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, journal, dates, 16 authors, 6 funders, 34 references.
Cite
This paper
Zussa, Z. N., Smith, A. N., van Vugt, J. J., O’Shaughnessy, D. S., Grima, N., Moi Fat, S. C., Blair, I. P., Rowe, D. B., Pamphlett, R., Nicholson, G. A., Kiernan, M. C., van Rheenen, W., Veldink, J., Project MinE ALS sequencing consortium, Williams, K. L., & Henden, L. (2026). Characterising the motif composition and allele length distribution of <i>ZFHX3</
BibTeX
@article{zussa2026charac
author = {Zussa, Zoe N. and Smith, Andrew N. and van Vugt, Joke J.F.A and O’Shaughnessy, Daniel S. and Grima, Natalie and Moi Fat, Sandrine Chan and Blair, Ian P. and Rowe, Dominic B and Pamphlett, Roger and Nicholson, Garth A. and Kiernan, Matthew C.K and van Rheenen, Wouter and Veldink, Jan and {Project MinE ALS sequencing consortium} and Williams, Kelly L. and Henden, Lyndal},
title = {{Characterising the motif composition and allele length distribution of <i>ZFHX3</
journal = {medRxiv (preprint)},
year = {2026},
month = mar,
publisher = {medRxiv},
doi = {10.64898/
url = {https://
}
RIS
TY - JOUR
AU - Zussa, Zoe N.
AU - Smith, Andrew N.
AU - van Vugt, Joke J.F.A
AU - O’Shaughnessy, Daniel S.
AU - Grima, Natalie
AU - Moi Fat, Sandrine Chan
AU - Blair, Ian P.
AU - Rowe, Dominic B
AU - Pamphlett, Roger
AU - Nicholson, Garth A.
AU - Kiernan, Matthew C.K
AU - van Rheenen, Wouter
AU - Veldink, Jan
AU - Project MinE ALS sequencing consortium
AU - Williams, Kelly L.
AU - Henden, Lyndal
TI - Characterising the motif composition and allele length distribution of <i>ZFHX3</
T2 - medRxiv (preprint)
J2 - medRxiv
PY - 2026
DA - 2026/
PB - medRxiv
DO - 10.64898/
UR - https://
ER -
CSL-JSON
{
"id": "10.64898/
"type": "article",
"title": "Characterising the motif composition and allele length distribution of <i>ZFHX3</
"container-title": "medRxiv (preprint)",
"author": [
{
"family": "Zussa",
"given": "Zoe N."
},
{
"family": "Smith",
"given": "Andrew N."
},
{
"family": "van Vugt",
"given": "Joke J.F.A"
},
{
"family": "O’Shaughnessy",
"given": "Daniel S."
},
{
"family": "Grima",
"given": "Natalie"
},
{
"family": "Moi Fat",
"given": "Sandrine Chan"
},
{
"family": "Blair",
"given": "Ian P."
},
{
"family": "Rowe",
"given": "Dominic B"
},
{
"family": "Pamphlett",
"given": "Roger"
},
{
"family": "Nicholson",
"given": "Garth A."
},
{
"family": "Kiernan",
"given": "Matthew C.K"
},
{
"family": "van Rheenen",
"given": "Wouter"
},
{
"family": "Veldink",
"given": "Jan"
},
{
"literal": "Project MinE ALS sequencing consortium"
},
{
"family": "Williams",
"given": "Kelly L."
},
{
"family": "Henden",
"given": "Lyndal"
}
],
"container-title-short":
"DOI": "10.64898/
"publisher": "medRxiv",
"URL": "https://
"issued": {
"date-parts": [
[
2026,
3,
10
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1136/bmjopen-2025-110906
- Strategic Amyotrophic Lateral Sclerosis Australia-Systems Genomics Consortium (SALSA-SGC): cohort profile.Journal: BMJ openIn common: other condition, 1 reference, 3 authors
- [2] doi:10.1038/s41586-026-10345-6 [code]
- Population-scale repeat expansions elucidate disease risk and brain atrophy.Journal: NatureIn common: data.table, tidyverse, genetics / omics, other condition, 3 references
- [3] doi:10.1155/humu/4225263
- The Charcot-Marie-Tooth Neuropathy (CMTX3) Complex Structural Variation Causes Differential SOX3 Spatiotemporal Expression.Journal: Human mutationIn common: genetics / omics, other condition, author Garth A. Nicholson
- [4] doi:10.1038/s41467-026-73902-7 [code]
- GWAS on short tandem repeats identifies genetic mechanisms in Alzheimer's disease.Journal: Nature communicationsIn common: data.table, tidyverse, genetics / omics, 2 references
- [5] doi:10.1016/j.mocell.2026.100363 [code]
- Tandem repeats in human brain evolution and disease susceptibility.Journal: Molecules and cellsIn common: 3 references
- [6] doi:10.1038/s41380-026-03574-8
- Genome-wide tandem repeat expansions modify schizophrenia risk in the presence of a 22q11.2 deletion.Journal: Molecular psychiatryIn common: genetics / omics, 3 references
- [7] doi:10.1038/s41467-026-72598-z [code]
- Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.Journal: Nature communicationsIn common: data.table, tidyverse, other condition, 1 reference
- [8] doi:10.1093/braincomms/fcag322 [code]
- Genomic insights into stroke recovery: cross-phenotype associations.Journal: Brain communicationsIn common: data.table, tidyverse, genetics / omics, 1 reference
- [9] doi:10.1371/journal.pcbi.1014422 [code]
- Deciphering cell type-specific causal genetic effects on brain imaging-derived phenotypes and disorders with single-cell Mendelian randomization.Journal: PLoS computational biologyIn common: data.table, tidyverse, genetics / omics, 1 reference
- [10] doi:10.1186/s13195-026-02036-1 [code]
- Genetic drivers of progression in Alzheimer's disease are distinct from disease risk.Journal: Alzheimer's research & therapyIn common: data.table, tidyverse, genetics / omics, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 2 scripts, and 1 match between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:f6a22154dcaf524d…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
