Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa.
The 9 matches
- [1] § Materials and methods › Measures › Neuroimaging outcomes › Neuroimaging processing ↔ app/main.sh, lines 1–40 · score 0.99 · callosal parcellations, ANTs Atropos, supratentorial tissue, Subcortical grey matter, template space, native space
- [2] § Materials and methods › Statistical analysis › Sample characteristics ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 5522–5562 · score 0.81 · duplicated demographic, chi squared, avoid skewing, clinical observations, multiple scans acquired, categorical variables
- [3] § Materials and methods › Statistical analysis › Neuroimaging regions of interest ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 2253–2296 · score 0.75 · mid anterior, mid posterior, basal ganglia, corpus callosum, brain volume, caudate nucleus
- [4] § Results › Primary analysis: antenatal maternal anaemia | HF and ULF MRI › Modelling: antenatal maternal anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 2253–2296 · score 0.72 · linear growth trajectories, logarithmic transformation, linear relationship, absolute age, corpus callosum, LME models
- [5] § Materials and methods › Statistical analysis › Secondary analysis: postnatal child anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 6863–6884 · score 0.69 · postnatal exposures, regional child brain, postnatal child anaemia, volumes remained, regional brain volumes, predictor
- [6] § Materials and methods › Measures › Anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 5382–5416 · score 0.68 · Minimum child haemoglobin, corresponding age, 3–24, multiple scans, diagnosis, child anaemia status
- [7] § Materials and methods › Statistical analysis › Primary analysis: antenatal maternal anaemia status › Statistical modelling ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 338–363 · score 0.66 · absolute regional brain, linear mixed, individual infant, volume trajectories, ULF subsamples, fitted
- [8] § Materials and methods › Measures › Contextual measures ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 76–171 · score 0.63 · maternal education, maternal employment, household income, maternal HIV, variables, sex
- [9] § Results › Primary analysis: antenatal maternal anaemia | HF and ULF MRI › Exploratory analyses: antenatal maternal anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 4870–4903 · score 0.61 · basal ganglia, rapid growth, growth trajectory, corpus callosum, LOESS curves, brain volumes
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Quarto · 6,955 lines · 196 KB · no license · 8 matches
- ---
- title: "Khula_AnaemiaImagingAnalysis"
- format: pdf
- editor: visual
- ---
- ------------------------------------------------------------------------
- ------------------------------------------------------------------------
- \*This is the final script used in the analyses for the following research: "**Antenatal Maternal Anaemia and Infant Brain Structure: 3 T and 64 mT MRI Findings from South Africa**"
- *Details of the datasets:*
- - Both high-field (HF; 3 T) and ultra-low-field (ULF; 64 mT) datasets are in long format with some infants having multiple scans across study timepoints (3mo, 6mo, 12mo, 18mo, 24mo).
- - These datasets represent the final samples after the exclusion of 1) scans that failed quality check (QC), 2) infants outside of the age range (\>25.5 months or \<2.5 months), and 3) statistical outliers.
- - HF: All acquired scans (*n* = 264) –\> passed QC (*n* = 247) –\> within the defined age range (*n* = 239) –\>removal of outliers (*n* = 232)
- - ULF: All acquired scans (*n* = 553) –\> passed QC (*n* = 456) —\> within the defined age range (*n* = 439) –\> removal out statistical outliers (*n* = 418)
- - The script for the HF analyses is presented first, followed by the script for the ULF analyses.
- #### --------------------------------------------------------------------------
- **Install Packages**
- ```{r}
- # install packages if needed
- # install.packages('pacman')
- library(pacman)
- p_load(readr, readxl, tidyverse, magrittr, writexl, stringr, dplyr)
- ```
- ------------------------------------------------------------------------
- [***High-Field (HF) Analyses***]{.underline}
- ------------------------------------------------------------------------
- **Import HF Dataset**
- ```{r}
- # import dataset
- HF_ida_final <- read_excel("HF_ida_final.xlsx")
- ```
- **Create New Variables for Total Brain Volumes Across Regions of Interest (ROIs)**
- ```{r}
- # compute total brain volumes for ROIs
- # basal ganglia
- HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
- HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
- # corpus callosum
- HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
- ```
- **Create New Age Variable (Days to Months):**
- ```{r}
- # convert child age at scan from days to months
- # average number of days in a month (accounting for leap years): 30.44
- HF_ida_final <- HF_ida_final %>%
- mutate(scan_age_months = hf_age / 30.44)
- ```
- **Creating Meaningful Categorical Variables**
- ```{r}
- # recode categorical variables from numeric to factor
- # maternal anaemia status
- HF_ida_final$maternal_preg_anaemia_status <- factor(
- HF_ida_final$maternal_preg_anaemia_status,
- levels = c(0, 1),
- labels = c("no_mat_anaemia", "mat_anaemia")
- )
- # maternal ID status
- HF_ida_final$maternal_preg_adj_brinda_id <- factor(
- HF_ida_final$maternal_preg_adj_brinda_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # maternal ID status
- HF_ida_final$maternal_preg_unadjusted_id <- factor(
- HF_ida_final$maternal_preg_unadjusted_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # maternal trimester Hb
- HF_ida_final$maternal_trimester_hb <- factor(
- HF_ida_final$maternal_trimester_hb,
- levels = c(1, 2, 3),
- labels = c("first", "second", "third")
- )
- # child sex
- HF_ida_final$child_sex <- factor(
- HF_ida_final$child_sex,
- levels = c(0, 1),
- labels = c("girl", "boy")
- )
- # maternal pae
- HF_ida_final$maternal_pae <- factor(
- HF_ida_final$maternal_pae,
- levels = c(0, 1),
- labels = c("no_pae", "pae")
- )
- #maternal pte
- HF_ida_final$maternal_pte <- factor(
- HF_ida_final$maternal_pte,
- levels = c(0, 1),
- labels = c("no_pte", "pte")
- )
- # maternal hiv
- HF_ida_final$maternal_hiv <- factor(
- HF_ida_final$maternal_hiv,
- levels = c(0, 1),
- labels = c("no_hiv", "hiv")
- )
- # maternal dep
- HF_ida_final$maternal_dep <- factor(
- HF_ida_final$maternal_dep,
- levels = c(0, 1),
- labels = c("no_dep", "dep")
- )
- # maternal employment
- HF_ida_final$mom_employment_en_dichotomous <- factor(
- HF_ida_final$mom_employment_en_dichotomous,
- levels = c(0, 1),
- labels = c("unemployed", "employed")
- )
- # household income
- HF_ida_final$income_household_en <- factor(
- HF_ida_final$income_household_en,
- levels = c(1, 2, 3, 4),
- labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
- )
- #maternal education
- HF_ida_final$mom_edu_en <- factor(
- HF_ida_final$mom_edu_en,
- levels = c(2, 3, 4, 5, 6),
- labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
- )
- # baby hiv
- HF_ida_final$baby_hiv <- factor(
- HF_ida_final$baby_hiv,
- levels = c(0, 1),
- labels = c("no_child_hiv", "child_hiv")
- )
- ```
- **Sample Distribution: Availability of Data Across Study Timepoints**
- ```{r}
- # sample over 12 months with maternal anaemia data
- sample_over_12M <- HF_ida_final %>%
- filter(scan_age_months > 12,
- !is.na(maternal_preg_anaemia_status))
- # sample under 12 months with maternal anaemia data
- sample_under_12M <- HF_ida_final %>%
- filter(scan_age_months < 12,
- !is.na(maternal_preg_anaemia_status))
- # sample over 18 months with maternal anaemia data
- sample_over_18 <- HF_ida_final %>%
- filter(scan_age_months > 18,
- !is.na(maternal_preg_anaemia_status))
- # sample over 24 months with maternal anaemia data
- sample_over_24 <- HF_ida_final %>%
- filter(scan_age_months > 24,
- !is.na(maternal_preg_anaemia_status))
- ```
- ```{r}
- # HF sample size
- # unique observations in HF dataset
- # excluding repeated measures
- library(dplyr)
- HF_ida_final %>%
- summarise(n_unique_subjects = n_distinct(subject_id))
- ```
- ------------------------------------------------------------------------
- **Primary Analysis: Antenatal Maternal Anaemia Status (HF Sample)**
- ------------------------------------------------------------------------
- **Sample Size: For infants with Antenatal Maternal Anemia Data**
- Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
- - The sample of unique participants was used to report on sample characteristics
- - The full sample with repeated measures was used for analyses including linear mixed effects (LME) models to include children with multiple scans from different study timepoints.
- ```{r}
- # total sample size
- # repeated measures included
- # sample with antenatal maternal anaemia data
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- summarise(n = n())
- ```
- ```{r}
- # unique sample size
- # repeated measures excluded
- # number of unique participants in sample
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include subjects with anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep unique subject IDs
- summarise(n_total = n()) %>% # count unique subjects
- print()
- ```
- ```{r}
- # prevalence of antenatal maternal anaemia
- # excluding repeated measures
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per mother
- group_by(maternal_preg_anaemia_status) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # antenatal maternal anaemia severity
- # unique subject IDs: Based on true prevalence
- # excluding repeated measures
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_severity)) %>% # exclude missing severity
- group_by(subject_id) %>% # group by mother
- summarise(
- maternal_preg_anaemia_severity = first(
- maternal_preg_anaemia_severity[maternal_preg_anaemia_severity != ""]
- )
- ) %>%
- group_by(maternal_preg_anaemia_severity) %>% # count unique mothers by severity
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- **Repeated Measures Data Availability**
- ```{r}
- # repeated measures count
- # unique subject IDs
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # keep only rows with anaemia status
- count(subject_id) %>% # count rows per subject
- count(n) %>% # count how many subjects had n rows
- rename(times_measured = n, number_of_subjects = nn) %>%
- mutate(percent = round(100 * number_of_subjects / sum(number_of_subjects), 1)) # add % column
- ```
- ```{r}
- # HF Figure 2
- # plot of data availability with repeated measures
- library(ggplot2)
- library(dplyr)
- Figure_2_HF<- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- mutate(
- timepoint = factor(timepoint, levels = c("3M", "6M", "12M", "18M", "24M")),
- subject_id = factor(subject_id),
- anaemia_group = as.factor(maternal_preg_anaemia_status)
- ) %>%
- ggplot(aes(x = timepoint, y = subject_id, group = subject_id, color = anaemia_group)) +
- geom_line(alpha = 0.5, linewidth = 0.5) + # updated for ggplot2 ≥3.4.0
- geom_point(size = 2) +
- scale_color_manual(
- values = c("red", "blue"),
- labels = c("Maternal Anaemia", "No Maternal Anaemia")
- ) +
- scale_x_discrete(expand = expansion(mult = 0.05)) +
- scale_y_discrete(expand = expansion(mult = 0.02)) +
- coord_cartesian(clip = "off") + # ensures no points are clipped at the edges
- labs(
- title = "Participant Data Availability by Timepoint",
- x = "Timepoint",
- y = "Participants",
- color = "Maternal Anaemia"
- ) +
- theme_minimal(base_size = 13) +
- theme(
- axis.text.y = element_blank(),
- axis.ticks.y = element_blank(),
- panel.grid.major.y = element_blank()
- )
- print(Figure_2_HF)
- ```
- **Sample Characteristics and Group Differences: Antenatal Maternal Anaemia Status**
- Sample characteristics were assessed based on the number of unique subjects in each subsample, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
- However, given the repeated measures study design with individual infant data across multiple study visits in both HF and ULF subsamples, linear mixed effects (LME) models were fitted to account for individual differences in assessing the association between antenatal maternal anaemia status and absolute regional brain volume trajectories
- - **Categorical Variables**
- *Trimester of Maternal Hb Measurement*
- ```{r}
- # trimester Hb was collected from mothers
- # excluding repeated measures
- # unique subject IDs: Based on true prevalence
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia
- group_by(subject_id) %>% # keep one row per mother
- summarise(maternal_trimester_hb = first(maternal_trimester_hb)) %>%
- group_by(maternal_trimester_hb) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in trimester of Hb collection
- # chi-squared: trimester Hb
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # Filter unique subjects with non-missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # Create contingency table: Trimester by Maternal Anaemia Status
- tbl <- table(unique_data$maternal_trimester_hb, unique_data$maternal_preg_anaemia_status)
- # Print contingency table
- cat("Contingency Table: Trimester by Maternal Anaemia Status\n")
- print(format(tbl, nsmall = 0))
- # Print column percentages
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # Run appropriate test based on expected cell counts
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Child Sex at Birth*
- ```{r}
- # child sex at birth (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per mother/child
- group_by(child_sex) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in child sex by antenatal maternal anaemia status
- # chi-squared: child sex
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # keep only unique subjects and non-missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$child_sex,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Child Sex by Maternal Anaemia Status\n")
- print(format(round(tbl, 2), nsmall = 2))
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Prenatal Alcohol Exposure (PAE)*
- ```{r}
- # pae prevalence (full sample)
- # unique subject IDs: based on true prevalence
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include subjects with anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep only unique subjects
- group_by(maternal_pae) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pae by antenatal maternal anaemia status
- # chi-squared: pae
- # unique subject IDs: based on true prevalence
- # filter unique subjects and exclude missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$maternal_pae,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal PAE by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Prenatal Tobacco Exposure (PTE)*
- ```{r}
- # pte prevalence (full sample)
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summarise maternal pte
- unique_data %>%
- group_by(maternal_pte) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pte by antenatal maternal anaemia status
- # chi-squared: pte
- # unique subject IDs: based on true prevalence
- # filter unique subjects and exclude missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$maternal_pte,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal PTE by Maternal Anaemia Status (Unique Subjects)\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test based on expected counts
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal HIV Infection*
- ```{r}
- # materal hiv prevalence (full sample)
- # unique subject IDs: based on true prevalence
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal hiv by antenatal maternal anaemia status
- # chi-squared: hiv
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$maternal_hiv,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal HIV by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Depression*
- ```{r}
- # maternal depression prevalence (full sample)
- # unique subject IDs: based on true prevalence
- HF_ida_final %>%
- filter(!is.na(maternal_dep)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_dep) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differenced by maternal depression status
- # chi-squared: maternal depression
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$maternal_dep,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal DEP by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Employment*
- ```{r}
- # maternal employment (full sample)
- # unique subject IDs: based on true prevalence
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep unique subjects
- group_by(mom_employment_en_dichotomous) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal employment by antenatal maternal anaemia status
- # chi-squared: maternal employment
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$mom_employment_en_dichotomous,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal Employment by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Education*
- ```{r}
- # maternal education (full sample)
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summarise maternal education
- unique_data %>%
- group_by(mom_edu_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal education by antenatal maternal anaemia status
- # chi-squared: maternal education
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$mom_edu_en,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Maternal Education by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Household Income*
- ```{r}
- # household income (full sample)
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summarise household income
- unique_data %>%
- group_by(income_household_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in household income by antenatal maternal anaemia status
- # chi-squared: household income
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$income_household_en,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Household Income by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Infant HIV Infection*
- ```{r}
- # infant hiv prevalence (full sample)
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summarise counts and percentages
- unique_data %>%
- group_by(baby_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in infant hiv by antenatal maternal anaemia status
- # chi-squared infant hiv
- # unique subject IDs: based on true prevalence
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # create contingency table
- tbl <- table(
- unique_data$baby_hiv,
- unique_data$maternal_preg_anaemia_status
- )
- # print contingency table
- cat("Contingency Table: Baby HIV by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- - **Continuous Variables**
- ```{r}
- #install packages if needed
- #install.packages('car')
- library("car")
- ```
- *Maternal Age at Enrolment*
- ```{r}
- # maternal age at enrolment (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(mom_age_en)) %>% # exclude missing anaemia or age
- group_by(subject_id) %>% # keep unique mothers only
- summarise(
- mom_age_en = first(mom_age_en) # assume age is constant per mother
- ) %>%
- summarise(
- mean_mom_age = mean(mom_age_en, na.rm = TRUE),
- sd_mom_age = sd(mom_age_en, na.rm = TRUE),
- min_mom_age = min(mom_age_en, na.rm = TRUE),
- max_mom_age = max(mom_age_en, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in maternal age at enrolment by antenatal maternal anaemia status
- # ttest: maternal age at enrolment
- # unique subject IDs: based on true prevalence
- library("dplyr")
- library("car")
- # keep only unique subjects with non-missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats with exactly 2 decimals
- summary_stats <- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats table
- print(summary_stats)
- # levene's Test for equality of variances
- levene_result <- leveneTest(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically select Student's or Welch's t-test
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # round t-test output (except p-value)
- t_result$estimate <- round(t_result$estimate, 2)
- t_result$statistic <- round(t_result$statistic, 2)
- t_result$conf.int <- round(t_result$conf.int, 2)
- # keep p-value as is, or round to 4 decimals for clean display
- t_result$p.value <- round(t_result$p.value, 4)
- # print t-test result
- print(t_result)
- ```
- *Child Age at Scan*
- Include repeated measures to account for different ages for scans acquired across study timepoints from the same infant
- ```{r}
- # child age at scan for full group
- # include repeated measures: different ages at every scan acquired (full sample)
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- summarise(
- mean_child_age = mean(scan_age_months, na.rm = TRUE),
- sd_child_age = sd(scan_age_months, na.rm = TRUE),
- min_child_age = min(scan_age_months, na.rm = TRUE),
- max_child_age = max(scan_age_months, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # child age at scan
- # full sample including repeated measures
- # t-test: child age at scan
- library(dplyr)
- library(car)
- # filter once and store
- analysis_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status))
- # summary statistics with exactly 2 decimals displayed
- summary_stats <- analysis_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's Test for equality of variances
- levene_result <- leveneTest(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = analysis_data
- )
- print(levene_result)
- # extract Levene p-value (rounded just for reporting)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test based on Levene
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = analysis_data,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = analysis_data,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Child Gestational Age (GA) at Birth*
- ```{r}
- # child ga at birth for full group
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics for the full sample
- unique_data %>%
- summarise(
- mean_child_ga = mean(ga_weeks, na.rm = TRUE),
- sd_child_ga = sd(ga_weeks, na.rm = TRUE),
- min_child_ga = min(ga_weeks, na.rm = TRUE),
- max_child_ga = max(ga_weeks, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child ga at birth
- # ttest: child ga
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics with exactly 2 decimals displayed
- summary_stats <- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test
- levene_result <- leveneTest(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- # round levene's p-value just for display
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Child Birth Weight (BW)*
- ```{r}
- # child bw (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics for the full sample
- unique_data %>%
- summarise(
- mean_bw = mean(bw, na.rm = TRUE),
- sd_bw = sd(bw, na.rm = TRUE),
- min_bw = min(bw, na.rm = TRUE),
- max_bw = max(bw, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child bw
- # ttest: child bw
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by maternal anaemia status with exactly 2 decimals displayed
- summary_stats <- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
- sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
- min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
- max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's Test for homogeneity of variance
- levene_result <- leveneTest(bw ~ maternal_preg_anaemia_status, data = unique_data)
- print(levene_result)
- # extract Levene p-value (rounded for reporting)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # t-test comparing birthweight by anaemia status
- # Automatically use equal variances if Levene p > 0.05
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- bw ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- bw ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # Print t-test result in full precision
- print(t_result)
- ```
- *Child Birth Length (BL)*
- ```{r}
- # child bl (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics for birth length
- unique_data %>%
- summarise(
- mean_birthlength = mean(birth_length, na.rm = TRUE),
- sd_birthlength = sd(birth_length, na.rm = TRUE),
- min_birthlength = min(birth_length, na.rm = TRUE),
- max_birthlength = max(birth_length, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child birthlength
- # ttest: child bl
- # exclude repeated measures: unique id
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics with exactly 2 decimals displayed
- summary_stats <- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
- sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
- min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
- max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's Test
- levene_result <- leveneTest(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Birth Head Circumference (HC)*
- ```{r}
- # child birth hc (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics for birth head circumference
- unique_data %>%
- summarise(
- mean_hc = mean(birth_hc, na.rm = TRUE),
- sd_hc = sd(birth_hc, na.rm = TRUE),
- min_hc = min(birth_hc, na.rm = TRUE),
- max_hc = max(birth_hc, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child birth hc
- # ttest: child hc
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for unique subjects with non-missing maternal anaemia status
- unique_data <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics with exactly 2 decimals displayed
- summary_stats <- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's Test
- levene_result <- leveneTest(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Gestational Age at Maternal Hb Measurement*
- ```{r}
- # ga at Hb measurement (full sample)
- # unique subject IDs: based on true prevalence
- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(maternal_min_hb_ga_roundown)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- n = n(),
- mean_ga_hb = round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- sd_ga_hb = round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- min_ga_hb = round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- q1_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.25, na.rm = TRUE), 2),
- median_ga_hb = round(median(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- q3_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.75, na.rm = TRUE), 2),
- max_ga_hb = round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2)
- )
- ```
- ```{r}
- # group differences in ga at Hb measurement by antenatal maternal anaemia status
- # ttest: ga at Hb measurement
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter non-missing and keep one row per subject
- HF_ida_final_unique <- HF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(maternal_min_hb_ga_roundown)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics with exactly 2 decimals displayed
- summary_stats <- HF_ida_final_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_ga = format(round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- sd_ga = format(round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- min_ga = format(round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- max_ga = format(round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test
- levene_result <- leveneTest(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = HF_ida_final_unique
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = HF_ida_final_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = HF_ida_final_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Minimum Maternal Hb measurement in Trimester 1*
- ```{r}
- # min Hb across groups in trimester 1
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 1, non-missing anaemia status and min Hb, keep unique subjects
- trim1_data_unique <- HF_ida_final %>%
- filter(
- maternal_trimester_hb == "first",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim1_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Minimum Maternal Hb in trimester 2*
- ```{r}
- # min Hb across groups in trimester 2
- # unique subject IDs: Based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 2, non-missing anaemia status and min Hb, keep unique subjects
- trim2_data_unique <- HF_ida_final %>%
- filter(
- maternal_trimester_hb == "second",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim2_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Minimum Maternal Hb in Trimester 3*
- ```{r}
- # min Hb across groups in trimester 3
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 3, non-missing anaemia status and min Hb, keep unique subjects
- trim3_data_unique <- HF_ida_final %>%
- filter(
- maternal_trimester_hb == "third",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim3_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique
- )
- print(levene_result)
- # extract levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- **Exploratory Analyses: Plotted Data using LOESS Curves**
- Absolute Volumes for ROIs plotted against child age at scan. Includes repeated measures to capture all scans collected across study timepoints.
- 1. Full sample
- 2. By group (Key: Antenatal Maternal Anaemia Status)
- *LOESS Curves: ICV*
- ```{r}
- # icv (full sample)
- # including repeated measures
- library(ggplot2)
- icv_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = hf_mm_icv)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "ICV vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total ICV"
- ) +
- theme_minimal()
- icv_hf_sample
- ```
- ```{r}
- # icv by group (antenatal maternal anaemia status)
- # including repeated measures
- library(ggplot2)
- library(dplyr)
- icv_hf_group <- ggplot(
- data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = hf_mm_icv,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # make ribbons slightly transparent
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4", # blue line
- "mat_anaemia" = "#d62728" # red line
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total ICV Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total ICV Volume (mm3)",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- icv_hf_group
- ```
- *LOESS Curves: Putamen*
- ```{r}
- # putamen (full sample)
- # including repeated measures
- library(ggplot2)
- putamen_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_putamen)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "Putamen vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Putamen Volume"
- ) +
- theme_minimal()
- putamen_hf_sample
- ```
- ```{r}
- # putamen by group (antenatal maternal anaemia status)
- # including repeated measures
- library(ggplot2)
- library(dplyr)
- putamen_hf_group <- ggplot(
- data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_putamen,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # slightly transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4", # blue line
- "mat_anaemia" = "#d62728" # red line
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total Putamen Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Putamen Volume",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- putamen_hf_group
- ```
- *LOESS Curves: Caudate Nucleus*
- ```{r}
- # caudate Nucleus (full sample)
- # including repeated measures
- library(ggplot2)
- caudate_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_caudate)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "Caudate vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Caudate Volume"
- ) +
- theme_minimal()
- caudate_hf_sample
- ```
- ```{r}
- # caudate nucleus by group (antenatal maternal anaemia status)
- # including repeated measures
- library(ggplot2)
- library(dplyr)
- caudate_hf_group <- ggplot(
- data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_caudate,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # light, semi-transparent ribbon
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4", # blue line
- "mat_anaemia" = "#d62728" # red line
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total Caudate Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Caudate Volume",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- caudate_hf_group
- ```
- *LOESS Curves: Corpus Callosum*
- ```{r}
- # corpus callosum (full sample)
- # including repeated measures
- library(ggplot2)
- cc_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_cc)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "CC vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Corpus Callosum Volume"
- ) +
- theme_minimal()
- cc_hf_sample
- ```
- ```{r}
- # corpus callosum by group (antenatal maternal anaemia status)
- # including repeated measures
- library(ggplot2)
- library(dplyr)
- cc_hf_group <- ggplot(
- data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_cc,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # make ribbons slightly transparent
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4", # original blue line
- "mat_anaemia" = "#d62728" # original red line
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue for ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red for ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- scale_y_continuous(
- n.breaks = 6
- ) +
- labs(
- title = "Total CC Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Corpus Callosum Volume (mm³)",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- cc_hf_group
- ```
- **Zero-Order Correlation Matrices: Antenatal Maternal Anaemia Status**
- ```{r}
- # install packages if needed
- # install and load Hmisc
- if (!require(Hmisc)) {
- install.packages("Hmisc")
- library(Hmisc)
- } else {
- library(Hmisc)}
- ```
- Convert variables to numeric for correlation matrices
- ```{r}
- # convert variables back from categorical to numeric
- HF_ida_final$income_household_en <- as.numeric(HF_ida_final$income_household_en)
- HF_ida_final$maternal_preg_anaemia_status_numeric <- as.numeric(as.character(
- factor(HF_ida_final$maternal_preg_anaemia_status,
- levels = c("no_mat_anaemia", "mat_anaemia"),
- labels = c(0, 1))
- ))
- HF_ida_final$maternal_hiv <- as.numeric(as.character(
- factor(HF_ida_final$maternal_hiv,
- levels = c("no_hiv", "mat_hiv"),
- labels = c(0, 1))
- ))
- HF_ida_final$child_sex <- as.numeric(as.character(
- factor(HF_ida_final$child_sex,
- levels = c("girl", "boy"),
- labels = c(0, 1))
- ))
- ```
- \*\* The variable for "antenatal maternal anaemia status" remained a categorical variable even after running the code above. Therefore, I recalled the dataset to ensure that this was numeric (as initially defined) for assessing correlations:
- ```{r}
- # recall HF dataset
- HF_ida_final <- read_excel("HF_ida_final.xlsx")
- # create new variables for total brain volumes across ROIs
- # basal ganglia
- HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
- HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
- # corpus callosum
- HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
- # create new age variable (days to months):
- # convert days to months
- # average Number of days in a month (accounting for leap years): 30.44
- HF_ida_final <- HF_ida_final %>%
- mutate(scan_age_months = hf_age / 30.44)
- ```
- **Correlation Matrices**
- ```{r}
- # correlation matrix for corpus callosum
- # filter dataset
- filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_cc")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for caudate nucleus
- # filter dataset
- filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_caudate")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for putamen
- # filter dataset
- filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_putamen")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for ICV
- # filter dataset
- filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "hf_mm_icv")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- **Multicollinearity**
- ```{r}
- # correlation between child age at scan and ICV
- library(dplyr)
- correlation_test <- cor.test(
- HF_ida_final$scan_age_months[!is.na(HF_ida_final$maternal_preg_anaemia_status)],
- HF_ida_final$hf_mm_icv[!is.na(HF_ida_final$maternal_preg_anaemia_status)])
- print(correlation_test)
- ```
- Decision to exclude ICV from the models for regional brain volumes due to high levels of multicollinearity (*r* \> 0.80) between child age at scan and ICV. Only child age at scan was included in the models below.
- **Linear Mixed Effects (LME) Models**
- ```{r}
- # install packages if needed
- # install.packages("lme4")
- library(lme4)
- # install.packages("lmerTest")
- library(lmerTest)
- # install.packages("parameters")
- library(parameters)
- # install.packages("emmeans")
- library(emmeans)
- ```
- Based on the exploratory analyses (plotted data; LOESS curves), absolute age was used for models with an observed linear relationship. Log (age) was used for models with an observed non-linear relationship.
- In the HF dataset:
- - Corpus callosum: Absolute age - Linear growth trajectory on plot
- - ICV, Putamen, Caudate Nucleus: Log (age) - Growth trajectory characterised by rapid growth and plateauing on plot
- For the HF sample, no interactions (child age at scan \* antenatal maternal anaemia status) were explored due to limited power. We kept the models simple and only assessed main effects.
- **LME Model variables**
- - Recall the dataset to reset variables
- - Create variabels for total ROI volumes and child age at scan in months
- - Ensure that all categorical variables (e.g. anaemia status, child sex, maternal HIV etc) are factors
- - Create variable for the logarithmic transformation of child age in months
- ```{r}
- # recall dataset
- HF_ida_final <- read_excel("HF_ida_final.xlsx")
- # create new variables for total brain volumes across ROIs
- # basal ganglia
- HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
- HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
- # corpus callosum
- HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
- # create new age variable (days to months):
- # convert days to months
- # average Number of days in a month (accounting for leap years): 30.44
- HF_ida_final <- HF_ida_final %>%
- mutate(scan_age_months = hf_age / 30.44)
- ```
- ```{r}
- # recode variables from numeric to factor
- # maternal anaemia status
- HF_ida_final$maternal_preg_anaemia_status <- factor(
- HF_ida_final$maternal_preg_anaemia_status,
- levels = c(0, 1),
- labels = c("no_mat_anaemia", "mat_anaemia")
- )
- # maternal ID status
- HF_ida_final$maternal_preg_adj_brinda_id <- factor(
- HF_ida_final$maternal_preg_adj_brinda_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # maternal ID status
- HF_ida_final$maternal_preg_unadjusted_id <- factor(
- HF_ida_final$maternal_preg_unadjusted_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # maternal trimester Hb
- HF_ida_final$maternal_trimester_hb <- factor(
- HF_ida_final$maternal_trimester_hb,
- levels = c(1, 2, 3),
- labels = c("first", "second", "third")
- )
- # child sex
- HF_ida_final$child_sex <- factor(
- HF_ida_final$child_sex,
- levels = c(0, 1),
- labels = c("girl", "boy")
- )
- # maternal pae
- HF_ida_final$maternal_pae <- factor(
- HF_ida_final$maternal_pae,
- levels = c(0, 1),
- labels = c("no_pae", "pae")
- )
- # maternal pte
- HF_ida_final$maternal_pte <- factor(
- HF_ida_final$maternal_pte,
- levels = c(0, 1),
- labels = c("no_pte", "pte")
- )
- # maternal hiv
- HF_ida_final$maternal_hiv <- factor(
- HF_ida_final$maternal_hiv,
- levels = c(0, 1),
- labels = c("no_hiv", "hiv")
- )
- # maternal dep
- HF_ida_final$maternal_dep <- factor(
- HF_ida_final$maternal_dep,
- levels = c(0, 1),
- labels = c("no_dep", "dep")
- )
- # maternal employment
- HF_ida_final$mom_employment_en_dichotomous <- factor(
- HF_ida_final$mom_employment_en_dichotomous,
- levels = c(0, 1),
- labels = c("unemployed", "employed")
- )
- # household income
- HF_ida_final$income_household_en <- factor(
- HF_ida_final$income_household_en,
- levels = c(1, 2, 3, 4),
- labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
- )
- # maternal education
- HF_ida_final$mom_edu_en <- factor(
- HF_ida_final$mom_edu_en,
- levels = c(2, 3, 4, 5, 6),
- labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
- )
- # baby hiv
- HF_ida_final$baby_hiv <- factor(
- HF_ida_final$baby_hiv,
- levels = c(0, 1),
- labels = c("no_child_hiv", "child_hiv")
- )
- ```
- ```{r}
- # create a log(age) variable for use in models of brain regions with non-linear growth
- HF_ida_final$log_age <- log(HF_ida_final$scan_age_months)
- ```
- *LME: Corpus Callosum*
- ```{r}
- # LME for corpus callosum
- # no interaction between maternal anaemia and age
- # absolute age for linear growth
- # model summary
- hf_model_cc <- lmer(total_cc ~ maternal_preg_anaemia_status + scan_age_months + child_sex + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_cc)
- # beta values
- standardize_parameters(hf_model_cc, method="refit")
- ```
- ```{r}
- # emms for corpus callosum volume by antenatal maternal anaemia status
- # based on final adjusted model
- library(emmeans)
- emm_cc <- emmeans(hf_model_cc, ~ maternal_preg_anaemia_status)
- # ----- EMM table with exactly 2 decimals -----
- emm_cc_df <- as.data.frame(emm_cc)
- emm_cc_df$emmean <- format(round(emm_cc_df$emmean, 2), nsmall = 2)
- emm_cc_df$SE <- format(round(emm_cc_df$SE, 2), nsmall = 2)
- emm_cc_df$lower.CL <- format(round(emm_cc_df$lower.CL, 2), nsmall = 2)
- emm_cc_df$upper.CL <- format(round(emm_cc_df$upper.CL, 2), nsmall = 2)
- print(emm_cc_df)
- # ----- contrast results -----
- contrast_results_cc <- contrast(emm_cc, method = "revpairwise", adjust = "holm")
- contrast_df_cc <- as.data.frame(summary(contrast_results_cc, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df_cc$estimate <- format(round(contrast_df_cc$estimate, 2), nsmall = 2)
- print(contrast_df_cc)
- ```
- *LME: ICV*
- ```{r}
- # LME for icv
- # no interaction between maternal anaemia and age
- # log (scan age) as growth non-linear
- # model summary
- hf_model_icv <- lmer(hf_mm_icv ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_icv)
- # beta values
- standardize_parameters(hf_model_icv)
- ```
- ```{r}
- # emms for icv by antenatal maternal anaemia status
- # based on final adjusted model
- library(emmeans)
- emm_icv <- emmeans(hf_model_icv, ~ maternal_preg_anaemia_status)
- # ----- EMM table with exactly 2 decimals -----
- emm_icv_df <- as.data.frame(emm_icv)
- emm_icv_df$emmean <- format(round(emm_icv_df$emmean, 2), nsmall = 2)
- emm_icv_df$SE <- format(round(emm_icv_df$SE, 2), nsmall = 2)
- emm_icv_df$lower.CL <- format(round(emm_icv_df$lower.CL, 2), nsmall = 2)
- emm_icv_df$upper.CL <- format(round(emm_icv_df$upper.CL, 2), nsmall = 2)
- print(emm_icv_df)
- # ----- contrast results -----
- contrast_results_icv <- contrast(emm_icv, method = "revpairwise", adjust = "holm")
- contrast_df_icv <- as.data.frame(summary(contrast_results_icv, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df_icv$estimate <- format(round(contrast_df_icv$estimate, 2), nsmall = 2)
- print(contrast_df_icv)
- ```
- *LME: Caudate Nucleus*
- ```{r}
- # lme for caudate nucleus
- # no interaction between maternal anaemia and age
- # log (scan age) as growth non-linear
- hf_model_caudate <- lmer(total_caudate ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_caudate)
- # beta values
- standardize_parameters(hf_model_caudate)
- ```
- ```{r}
- # emms for caudate nucleus by antenatal maternal anaemia status
- # based on final adjusted model
- library(emmeans)
- emm_caudate <- emmeans(hf_model_caudate, ~ maternal_preg_anaemia_status)
- # ----- emm table with exactly 2 decimals -----
- emm_caudate_df <- as.data.frame(emm_caudate)
- emm_caudate_df$emmean <- format(round(emm_caudate_df$emmean, 2), nsmall = 2)
- emm_caudate_df$SE <- format(round(emm_caudate_df$SE, 2), nsmall = 2)
- emm_caudate_df$lower.CL <- format(round(emm_caudate_df$lower.CL, 2), nsmall = 2)
- emm_caudate_df$upper.CL <- format(round(emm_caudate_df$upper.CL, 2), nsmall = 2)
- print(emm_caudate_df)
- # ----- contrast results -----
- contrast_results_caudate <- contrast(emm_caudate, method = "revpairwise", adjust = "holm")
- contrast_df_caudate <- as.data.frame(summary(contrast_results_caudate, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df_caudate$estimate <- format(round(contrast_df_caudate$estimate, 2), nsmall = 2)
- print(contrast_df_caudate)
- ```
- *LME: Putamen*
- ```{r}
- # lme for putamen
- # no interaction between maternal anaemia and age
- # log (scan age) as growth non-linear
- hf_model_putamen <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_putamen)
- # beta values
- standardize_parameters(hf_model_putamen)
- ```
- ```{r}
- # emms for caudate nucleus by antenatal maternal anaemia status
- # based on final adjusted model
- library(emmeans)
- emm_putamen <- emmeans(hf_model_putamen, ~ maternal_preg_anaemia_status)
- # ----- emm table with exactly 2 decimals -----
- emm_putamen_df <- as.data.frame(emm_putamen)
- emm_putamen_df$emmean <- format(round(emm_putamen_df$emmean, 2), nsmall = 2)
- emm_putamen_df$SE <- format(round(emm_putamen_df$SE, 2), nsmall = 2)
- emm_putamen_df$lower.CL <- format(round(emm_putamen_df$lower.CL, 2), nsmall = 2)
- emm_putamen_df$upper.CL <- format(round(emm_putamen_df$upper.CL, 2), nsmall = 2)
- print(emm_putamen_df)
- # ----- contrast results -----
- contrast_results_putamen <- contrast(emm_putamen, method = "revpairwise", adjust = "holm")
- contrast_df_putamen <- as.data.frame(summary(contrast_results_putamen, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df_putamen$estimate <- format(round(contrast_df_putamen$estimate, 2), nsmall = 2)
- print(contrast_df_putamen)
- ```
- **Sensitivity Analyses**
- Based on the final models for each ROI (established in primary analyses): With the inclusion of maternal HIV (potential confounder)
- ```{r}
- # sensitivity analysis for corpus callosum
- # including maternal hiv
- # excluding interaction between maternal anaemia and age
- hf_model_cc_sens <- lmer(total_cc ~ maternal_preg_anaemia_status + scan_age_months + child_sex + maternal_hiv + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_cc_sens)
- standardize_parameters(hf_model_cc_sens)
- ```
- ```{r}
- # sensitivity analysis for icv
- # including maternal hiv
- # excluding interaction between maternal anaemia and age
- hf_model_icv_sens <- lmer(hf_mm_icv ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_icv_sens)
- standardize_parameters(hf_model_icv_sens)
- ```
- ```{r}
- # sensitivity analysis for putamen
- # including maternal hiv
- # excluding interaction between maternal anaemia and age
- hf_model_putamen_sens <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_putamen_sens)
- standardize_parameters(hf_model_putamen_sens)
- ```
- ```{r}
- # sensitivity analysis for caudate nucleus
- # including maternal hiv
- # excluding interaction between maternal anaemia and age
- hf_model_caudate_sens <- lmer(total_caudate ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = HF_ida_final)
- summary(hf_model_caudate_sens)
- standardize_parameters(hf_model_caudate_sens)
- ```
- ------------------------------------------------------------------------
- [***Ultra-Low-Field (ULF) Analyses***]{.underline}
- ------------------------------------------------------------------------
- **Import ULF Dataset**
- ```{r}
- ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
- ```
- **Create New Variables for Total Brain Volumes across ROIs**
- ```{r}
- # basal ganglia
- ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
- ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
- # corpus callosum
- ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
- ```
- **Create New Age Variable (days to months)**
- ```{r}
- # convert days to months
- # average number of days in a month (accounting for leap years): 30.44
- ULF_ida_final <- ULF_ida_final %>%
- mutate(scan_age_months = hyp_age / 30.44)
- ```
- **Create Meaningful Categorical Variables**
- ```{r}
- # convert from numeric to factor
- # maternal anaemia status
- ULF_ida_final$maternal_preg_anaemia_status <- factor(
- ULF_ida_final$maternal_preg_anaemia_status,
- levels = c(0, 1),
- labels = c("no_mat_anaemia", "mat_anaemia")
- )
- # maternal anaemia severity
- ULF_ida_final$maternal_preg_anaemia_severity <- factor(
- ULF_ida_final$maternal_preg_anaemia_severity,
- levels = c(1, 2, 3),
- labels = c("mild", "moderate", "severe")
- )
- # child status
- ULF_ida_final$cor_child_anaemia <- factor(
- ULF_ida_final$cor_child_anaemia,
- levels = c(0, 1),
- labels = c("no_child_anaemia", "child_anaemia")
- )
- # maternal ID status
- ULF_ida_final$maternal_preg_adj_brinda_id <- factor(
- ULF_ida_final$maternal_preg_adj_brinda_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # child ID status
- ULF_ida_final$cor_child_id <- factor(
- ULF_ida_final$cor_child_id,
- levels = c(0, 1),
- labels = c("no_child_id", "child_id")
- )
- # maternal trimester hb
- ULF_ida_final$maternal_trimester_hb <- factor(
- ULF_ida_final$maternal_trimester_hb,
- levels = c(1, 2, 3),
- labels = c("first", "second", "third")
- )
- # child sex
- ULF_ida_final$child_sex <- factor(
- ULF_ida_final$child_sex,
- levels = c(0, 1),
- labels = c("girl", "boy")
- )
- # maternal pae
- ULF_ida_final$maternal_pae <- factor(
- ULF_ida_final$maternal_pae,
- levels = c(0, 1),
- labels = c("no_pae", "pae")
- )
- # maternal pte
- ULF_ida_final$maternal_pte <- factor(
- ULF_ida_final$maternal_pte,
- levels = c(0, 1),
- labels = c("no_pte", "pte")
- )
- # maternal hiv
- ULF_ida_final$maternal_hiv <- factor(
- ULF_ida_final$maternal_hiv,
- levels = c(0, 1),
- labels = c("no_hiv", "hiv")
- )
- # maternal dep
- ULF_ida_final$maternal_dep <- factor(
- ULF_ida_final$maternal_dep,
- levels = c(0, 1),
- labels = c("no_dep", "dep")
- )
- # maternal employment
- ULF_ida_final$mom_employment_en_dichotomous <- factor(
- ULF_ida_final$mom_employment_en_dichotomous,
- levels = c(0, 1),
- labels = c("unemployed", "employed")
- )
- # household income
- ULF_ida_final$income_household_en <- factor(
- ULF_ida_final$income_household_en,
- levels = c(1, 2, 3, 4),
- labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
- )
- # maternal education
- ULF_ida_final$mom_edu_en <- factor(
- ULF_ida_final$mom_edu_en,
- levels = c(2, 3, 4, 5, 6),
- labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
- )
- # baby hiv
- ULF_ida_final$baby_hiv <- factor(
- ULF_ida_final$baby_hiv,
- levels = c(0, 1),
- labels = c("no_child_hiv", "child_hiv")
- )
- ```
- **Data Distribution across Timepoint**
- ```{r}
- # sample over 12 months with maternal anaemia data
- sample_over_12M <- ULF_ida_final %>%
- filter(scan_age_months > 12,
- !is.na(maternal_preg_anaemia_status))
- # sample under 12 months with maternal anaemia data
- sample_under_12M <- ULF_ida_final %>%
- filter(scan_age_months < 12,
- !is.na(maternal_preg_anaemia_status))
- # sample over 18 months with maternal anaemia data
- sample_over_18 <- ULF_ida_final %>%
- filter(scan_age_months > 18,
- !is.na(maternal_preg_anaemia_status))
- # sample over 24 months with maternal anaemia data
- sample_over_24 <- ULF_ida_final %>%
- filter(scan_age_months > 24,
- !is.na(maternal_preg_anaemia_status))
- ```
- **Overlap Between HF and ULF Samples**
- No. of unique subjects for full ULF neuroimaging sample:
- ```{r}
- # unique observations in ULF dataset
- # excluding repeated measures
- library(dplyr)
- ULF_ida_final %>%
- summarise(n_unique_subjects = n_distinct(subject_id))
- ```
- Overlap in unique subjects between full HF and ULF samples:
- ```{r}
- # no. of children in common between HF and ULF neuroimaging datasets
- # find overlapping unique subject_ids
- overlap_subjects <- intersect(HF_ida_final$subject_id, ULF_ida_final$subject_id)
- # Count how many there are
- num_overlap <- length(overlap_subjects)
- # Print the number
- num_overlap
- ```
- For neuroimaging sample with maternal anaemia data available:
- ```{r}
- # no. of unique children in common between HF and ULF neuroimaging datasets with maternal anaemia data
- # filter dataframes to keep only rows with non-missing maternal_preg_anaemia_status
- HF_filtered <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
- ULF_filtered <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
- # find overlapping subject_ids among the filtered sets
- overlap_subjects <- intersect(HF_filtered$subject_id, ULF_filtered$subject_id)
- # count the number of overlapping subjects
- num_overlap <- length(overlap_subjects)
- # print the result
- num_overlap
- ```
- ------------------------------------------------------------------------
- **Primary Analysis: Antenatal Maternal Anaemia Status (ULF Sample)**
- ------------------------------------------------------------------------
- **Sample Size: For infants with Antenatal Maternal Anemia Data**
- Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
- ```{r}
- # total sample size
- # including repeated measures
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
- summarise(n = n())
- ```
- ```{r}
- # repeated measures excluded
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- summarise(unique_subjects = n_distinct(subject_id))
- ```
- ```{r}
- # prevalence of antenatal maternal anaemia
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # Exclude missing data
- group_by(subject_id) %>%
- summarise(
- maternal_preg_anaemia_status = if_else(
- any(maternal_preg_anaemia_status == "mat_anaemia", na.rm = TRUE),
- "mat_anaemia",
- "no_mat_anaemia"
- )
- ) %>%
- count(maternal_preg_anaemia_status) %>%
- mutate(
- prevalence_percent = n / sum(n) * 100
- )
- ```
- ```{r}
- # antenatal maternal anaemia severity
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_severity)) %>% # exclude missing severity
- group_by(subject_id) %>% # group by mother
- summarise(
- maternal_preg_anaemia_severity = first(
- maternal_preg_anaemia_severity[maternal_preg_anaemia_severity != ""]
- )
- ) %>%
- group_by(maternal_preg_anaemia_severity) %>% # count unique mothers by severity
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- **Data Availability with Repeated Measures**
- ```{r}
- # repeated measures count
- # unique ID
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # keep only rows with anaemia status
- count(subject_id) %>% # count rows per subject
- count(n) %>% # count how many subjects had n rows
- rename(times_measured = n, number_of_subjects = nn) %>%
- mutate(percent = round(100 * number_of_subjects / sum(number_of_subjects), 1)) # Add % column
- ```
- ```{r}
- # figure 2
- library(ggplot2)
- library(dplyr)
- Figure_2_ULF <-ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- mutate(
- timepoint = factor(timepoint, levels = c("3M", "6M", "12M", "18M", "24M")), # enforce order
- subject_id = factor(subject_id),
- anaemia_group = as.factor(maternal_preg_anaemia_status)
- ) %>%
- ggplot(aes(x = timepoint, y = subject_id, group = subject_id, color = anaemia_group)) +
- geom_line(alpha = 0.5, linewidth = 0.5) + # use linewidth instead of deprecated size
- geom_point(size = 2) +
- scale_color_manual(
- values = c("red", "blue"),
- labels = c("Maternal Anaemia", "No Maternal Anaemia")
- ) +
- labs(
- title = "Participant Data Availability by Timepoint",
- x = "Timepoint",
- y = "Participants",
- color = "Maternal Anaemia"
- ) +
- theme_minimal(base_size = 13) +
- theme(
- axis.text.y = element_blank(),
- axis.ticks.y = element_blank(),
- panel.grid.major.y = element_blank()
- ) +
- # <-- add padding so points are not cut off
- scale_y_discrete(expand = expansion(mult = c(0.02, 0.02))) +
- scale_x_discrete(expand = expansion(mult = c(0.02, 0.02)))
- print(Figure_2_ULF)
- ```
- **Sample Characteristics and Group Differences: Antenatal Maternal Anaemia Status**
- Sample characteristics were assessed based on the number of unique subjects in each subsample, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
- However, given the repeated measures study design with individual infant data across multiple study visits in both HF and ULF subsamples, linear mixed effects (LME) models were fitted to account for individual differences in assessing the association between antenatal maternal anaemia status and absolute regional brain volume trajectories
- - **Categorical Variables**
- *Trimester of Maternal Hb Measurement*
- ```{r}
- # trimester Hb was collected from mothers
- # excluding repeated measures
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia
- group_by(subject_id) %>% # keep unique mothers only
- summarise(
- maternal_trimester_hb = first(maternal_trimester_hb) # assume first record per mother
- ) %>%
- group_by(maternal_trimester_hb) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in trimester of Hb collection
- # Chi-squared: trimester Hb
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # run contingency table + test without creating a new dataset
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia status
- group_by(subject_id) %>% # ensure one row per mother
- summarise(
- maternal_trimester_hb = first(maternal_trimester_hb),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$maternal_trimester_hb, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Trimester by Maternal Anaemia Status\n")
- print(format(round(tbl, 2), nsmall = 2))
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Child Sex at Birth*
- ```{r}
- # child sex at birth (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per child
- summarise(
- child_sex = first(child_sex) # assume sex is consistent per child
- ) %>%
- group_by(child_sex) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in child sex by antenatal maternal anaemia status
- # chi-squared: child sex
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # run contingency table + test without creating a new dataset
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per child
- summarise(
- child_sex = first(child_sex),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$child_sex, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Child Sex by Maternal Anaemia Status\n")
- print(format(round(tbl, 2), nsmall = 2))
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Prenatal Alcohol Exposure (PAE)*
- ```{r}
- # pae prevalence (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_pae = first(maternal_pae) # assume pae status is consistent per mother
- ) %>%
- group_by(maternal_pae) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pae by antenatal maternal anaemia status
- # chi-squared: pae
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # build contingency table and run test inline (unique mothers only)
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_pae = first(maternal_pae),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$maternal_pae, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Maternal PAE by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Prenatal Tobacco Exposure (PTE)*
- ```{r}
- # pte prevalence (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_pte = first(maternal_pte) # assume pte status is consistent per mother
- ) %>%
- group_by(maternal_pte) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pte by antenatal maternal anaemia status
- # chi-squared: pte
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # build contingency table and run test inline (unique mothers only)
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_pte = first(maternal_pte),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$maternal_pte, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Maternal PTE by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Maternal HIV Infection*
- ```{r}
- # materal hiv prevalence (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_hiv = first(maternal_hiv) # assume HIV status is consistent per mother
- ) %>%
- group_by(maternal_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal hiv by antenatal maternal anaemia status
- # chi-squared: hiv
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # build contingency table and run test inline (unique mothers only)
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_hiv = first(maternal_hiv),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$maternal_hiv, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Maternal HIV by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Maternal Depression*
- ```{r}
- # maternal depression prevalence (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_dep) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differenced by maternal depression status
- # chi-squared: maternal depression
- # unique subject IDs: based on true prevalence
- library(dplyr)
- # build contingency table and run test inline (unique mothers only)
- tbl <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
- group_by(subject_id) %>% # ensure one record per mother
- summarise(
- maternal_dep = first(maternal_dep),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
- ) %>%
- {table(.$maternal_dep, .$maternal_preg_anaemia_status)} # build contingency table inline
- # print contingency table
- cat("Contingency Table: Maternal Depression by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- chisq_result <- suppressWarnings(chisq.test(tbl))
- expected <- chisq_result$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Maternal Employment*
- ```{r}
- # maternal employment (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia status
- group_by(subject_id) %>% # ensure one record per mother
- summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous)) %>%
- group_by(mom_employment_en_dichotomous) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal employment by antenatal maternal anaemia status
- # chi-squared: maternal employment
- # unique subject IDs: based on true prevalence
- # create contingency table inline using first record per mother
- tbl <- table(
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- group_by(subject_id) %>%
- summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
- pull(mom_employment_en_dichotomous),
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- group_by(subject_id) %>%
- summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous),
- maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
- pull(maternal_preg_anaemia_status)
- )
- # print contingency table
- cat("Contingency Table: Maternal Employment by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Education*
- ```{r}
- # maternal education (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # Include only mothers with anaemia data
- group_by(subject_id) %>% # Ensure one record per mother
- summarise(mom_edu_en = first(mom_edu_en)) %>% # Take first value per mother
- group_by(mom_edu_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal education by antenatal maternal anaemia status
- # chi-squared: maternal education
- # unique subject IDs: based on true prevalence
- # create contingency table inline, one record per mother
- tbl <- table(
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- group_by(subject_id) %>% # Ensure unique mother
- summarise(mom_edu_en = first(mom_edu_en)) %>%
- pull(mom_edu_en),
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- group_by(subject_id) %>%
- summarise(maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
- pull(maternal_preg_anaemia_status)
- )
- # print contingency table
- cat("Contingency Table: Maternal Education by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Household Income*
- ```{r}
- # household income (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include mothers with anaemia data
- group_by(subject_id) %>% # ensure each mother counted once
- summarise(income_household_en = first(income_household_en)) %>%
- group_by(income_household_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in household income by antenatal maternal anaemia status
- # chi-squared: household income
- # unique subject IDs: based on true prevalence
- # create contingency table counting each subject only once
- tbl <- table(
- ULF_ida_final$income_household_en[!duplicated(ULF_ida_final$subject_id) &
- !is.na(ULF_ida_final$maternal_preg_anaemia_status)],
- ULF_ida_final$maternal_preg_anaemia_status[!duplicated(ULF_ida_final$subject_id) &
- !is.na(ULF_ida_final$maternal_preg_anaemia_status)]
- )
- # print contingency table
- cat("Contingency Table: Household Income by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Infant HIV Infection*
- ```{r}
- # infant hiv prevalence (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- filter(!duplicated(subject_id)) %>% # count each subject once
- group_by(baby_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in infant hiv by antenatal maternal anaemia status
- # chi-squared: infant hiv
- # unique subject IDs: based on true prevalence
- # create contingency table counting each subject only once
- tbl <- table(
- ULF_ida_final$baby_hiv[!is.na(ULF_ida_final$maternal_preg_anaemia_status) & !duplicated(ULF_ida_final$subject_id)],
- ULF_ida_final$maternal_preg_anaemia_status[!is.na(ULF_ida_final$maternal_preg_anaemia_status) & !duplicated(ULF_ida_final$subject_id)]
- )
- # print contingency table
- cat("Contingency Table: Infant HIV by Maternal Anaemia Status\n")
- print(tbl)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- - **Continuous Variables**
- ```{r}
- #install packages if needed
- #install.packages('car')
- library(car)
- ```
- *Maternal Age at Enrolment*
- ```{r}
- # maternal age at enrolment (full sample)
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(mom_age_en)) %>% # exclude missing anaemia or age
- group_by(subject_id) %>% # keep unique mothers only
- summarise(
- mom_age_en = first(mom_age_en) # assume maternal age at enrolment is constant
- ) %>%
- summarise(
- mean_mom_age = mean(mom_age_en, na.rm = TRUE),
- sd_mom_age = sd(mom_age_en, na.rm = TRUE),
- min_mom_age = min(mom_age_en, na.rm = TRUE),
- max_mom_age = max(mom_age_en, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in maternal age at enrolment by antenatal maternal anaemia status
- # ttest: maternal age at enrolment
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing anaemia status
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(mom_age_en)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats (force exactly 2 decimals for display)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- mom_age_en ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Child Age at Scan:*
- Include repeated measures to account for different ages for scans acquired across study timepoints from the same infant
- ```{r}
- # child age at scan for full group
- # include repeated measures as child is a different age at every scan acquired (full sample)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% #only include PIDs with antenatal maternal anaemia data
- summarise(
- mean_child_age = mean(scan_age_months, na.rm = TRUE),
- sd_child_age = sd(scan_age_months, na.rm = TRUE),
- min_child_age = min(scan_age_months, na.rm = TRUE),
- max_child_age = max(scan_age_months, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # child age at scan
- # full sample including repeated measures
- # t-test: child age at scan
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing anaemia status and scan age
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(scan_age_months))
- # summary stats (force exactly 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- scan_age_months ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Child Gestational Age (GA) at Birth*
- ```{r}
- # child ga at birth for full group
- # unique subject IDs: based on true prevalence
- library(dplyr)
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>%
- distinct(subject_id, .keep_all = TRUE) %>% # keep only one row per child
- summarise(
- mean_child_ga = mean(ga_weeks, na.rm = TRUE),
- sd_child_ga = sd(ga_weeks, na.rm = TRUE),
- min_child_ga = min(ga_weeks, na.rm = TRUE),
- max_child_ga = max(ga_weeks, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child ga at birth
- # ttest: child ga at birth
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing maternal anaemia status and GA
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(ga_weeks)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats by maternal anaemia status (force 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # Decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- ga_weeks ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Child Birth Weight (BW)*
- ```{r}
- # child bw (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with maternal anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # ensure one record per child
- summarise(
- mean_bw = mean(bw, na.rm = TRUE),
- sd_bw = sd(bw, na.rm = TRUE),
- min_bw = min(bw, na.rm = TRUE),
- max_bw = max(bw, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child bw
- # ttest: child bw
- # exclude repeated measures:unique ID
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing maternal anaemia status and birthweight
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(bw)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats by maternal anaemia status (force 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
- sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
- min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
- max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- bw ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- bw ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- bw ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Child Birth Length (BL)*
- ```{r}
- # child bl (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # ensure only one record per child
- summarise(
- mean_birthlength = mean(birth_length, na.rm = TRUE),
- sd_birthlength = sd(birth_length, na.rm = TRUE),
- min_birthlength = min(birth_length, na.rm = TRUE),
- max_birthlength = max(birth_length, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # child bl
- # unique subject IDs: based on true prevalence
- # t-test: child bl
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing maternal anaemia status and birth length
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(birth_length)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats by maternal anaemia status (force 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
- sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
- min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
- max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- birth_length ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Child Birth Head Circumference (HC)*
- ```{r}
- # child birth hc (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with anaemia data
- distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per unique child
- summarise(
- mean_hc = mean(birth_hc, na.rm = TRUE),
- sd_hc = sd(birth_hc, na.rm = TRUE),
- min_hc = min(birth_hc, na.rm = TRUE),
- max_hc = max(birth_hc, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in birth hc by antenatal maternal anaemia status
- # t-test: child hc at birth
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing maternal anaemia status and birth HC
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(birth_hc)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats by maternal anaemia status (force 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # Decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- birth_hc ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Gestational Age (GA) at Maternal Hb Measurement*
- ```{r}
- # ga at min Hb measurement (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(maternal_min_hb_ga_roundown)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- n = n(),
- mean_ga_hb = round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- sd_ga_hb = round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- min_ga_hb = round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- q1_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.25, na.rm = TRUE), 2),
- median_ga_hb = round(median(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
- q3_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.75, na.rm = TRUE), 2),
- max_ga_hb = round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2)
- )
- ```
- ```{r}
- # group differences in ga at minimum maternal Hb by antenatal maternal anaemia status
- # t-test: ga at min Hb measurement
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # keep only unique subjects with non-missing maternal anaemia status and GA at min Hb
- unique_data <- ULF_ida_final %>%
- filter(!is.na(maternal_preg_anaemia_status),
- !is.na(maternal_min_hb_ga_roundown)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary stats by maternal anaemia status (force 2 decimals)
- unique_data %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_ga_hb = format(round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- sd_ga_hb = format(round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- min_ga_hb = format(round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- max_ga_hb = format(round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- ) %>%
- print()
- # Levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = unique_data
- )
- print(levene_result)
- # extract p-value from Levene’s Test
- levene_p <- levene_result$`Pr(>F)`[1]
- # decide automatically which t-test to use
- if (levene_p > 0.05) {
- cat("\nLevene’s test not significant (p =", round(levene_p, 4),
- ") → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = TRUE
- )
- } else {
- cat("\nLevene’s test significant (p =", round(levene_p, 4),
- ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
- data = unique_data,
- var.equal = FALSE
- )
- }
- # print t-test results (full precision)
- print(t_result)
- ```
- *Minimum Maternal Hb measurement in Trimester 1*
- ```{r}
- # maternal min Hb across groups in trimester 1: by antenatal maternal anaemia status
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 1, non-missing anaemia status and min Hb, keep unique subjects
- trim1_data_unique <- ULF_ida_final %>%
- filter(
- maternal_trimester_hb == "first",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim1_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(), # ← missing comma fixed here
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim1_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Minimum Maternal Hb measurement in Trimester 2*
- ```{r}
- # maternal minimum Hb in trimester 2 by antenatal anaemia status
- # t-test
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 2, non-missing anaemia status and min Hb, keep unique subjects
- trim2_data_unique <- ULF_ida_final %>%
- filter(
- maternal_trimester_hb == "second",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim2_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique
- )
- print(levene_result)
- # extract levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim2_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Minimum Maternal Hb measurement in Trimester 3*
- ```{r}
- # maternal minimum Hb in trimester 3 by antenatal anaemia status
- # t-test
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # filter for trimester 3, non-missing anaemia status and min Hb, keep unique subjects
- trim3_data_unique <- ULF_ida_final %>%
- filter(
- maternal_trimester_hb == "third",
- !is.na(maternal_preg_anaemia_status),
- !is.na(maternal_preg_min_hb)
- ) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics by anaemia status with exactly 2 decimals displayed
- summary_stats <- trim3_data_unique %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- # print summary stats
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique
- )
- print(levene_result)
- # extract Levene p-value (rounded for reporting only)
- levene_p <- round(levene_result$`Pr(>F)`[1], 4)
- cat("\nLevene’s test p-value =", levene_p, "\n")
- # automatically choose correct t-test
- if (levene_p > 0.05) {
- cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique,
- var.equal = TRUE
- )
- } else {
- cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
- t_result <- t.test(
- maternal_preg_min_hb ~ maternal_preg_anaemia_status,
- data = trim3_data_unique,
- var.equal = FALSE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- **Exploratory Analyses: Plotted Data using LOESS Curves**
- Absolute Volumes for ROIs plotted against child age at scan. Includes repeated measures to capture all scans collected across study timepoints.
- 1. Full sample
- 2. By group (Key: Antenatal Maternal Anaemia Status)
- *LOESS Curves: ICV*
- ```{r}
- # icv (full sample)
- # including repeated measures
- library(ggplot2)
- icv_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = mm_hyp_icv)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "ICV vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total ICV"
- ) +
- theme_minimal()
- icv_ulf_sample
- ```
- ```{r}
- # icv by group (antenatal maternal anaemia status)
- # including repeated measures
- icv_ulf_group <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = mm_hyp_icv,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # slightly transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4",
- "mat_anaemia" = "#d62728"
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
- ),
- guide = "none"
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- scale_y_continuous(
- limits = c(600000, NA),
- breaks = seq(600000, max(ULF_ida_final$mm_hyp_icv, na.rm = TRUE), by = 200000)
- ) +
- labs(
- title = "Total ICV Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total ICV Volume (mm³)",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- icv_ulf_group
- ```
- *LOESS Curve: Putamen*
- ```{r}
- # putamen (full sample)
- # including repeated measures
- library(ggplot2)
- putamen_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_putamen)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "Putamen vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Putamen Volume"
- ) +
- theme_minimal()
- putamen_ulf_sample
- ```
- ```{r}
- # putamen by group (antenatal maternal anaemia status)
- # including repeated measures
- putamen_ulf_group <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_putamen,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # slightly transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4",
- "mat_anaemia" = "#d62728"
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
- ),
- guide = "none"
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total Putamen Volume vs. Scan Age",
- x = "Scan Age (months)",
- y = "Total Putamen Volume",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- putamen_ulf_group
- ```
- *LOESS Curves: Caudate Nucleus*
- ```{r}
- # caudate nucleus (full sample)
- # including repeated measures
- library(ggplot2)
- caudate_hf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_caudate)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "Caudate vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Caudate Volume"
- ) +
- theme_minimal()
- caudate_hf_sample
- ```
- ```{r}
- # caudate nucleus by group (antenatal maternal anaemia status)
- # including repeated measures
- caudate_ulf_group <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_caudate,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # slightly transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4",
- "mat_anaemia" = "#d62728"
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
- ),
- guide = "none"
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total Caudate Volume vs. Scan Age",
- x = "Scan Age (months)",
- y = "Total Caudate Volume",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- caudate_ulf_group
- ```
- *LOESS Curves: Corpus Callosum*
- ```{r}
- # corpus callosum (full sample)
- # including repeated measures
- library(ggplot2)
- cc_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_cc)) +
- geom_point(alpha = 0.6) +
- geom_smooth(method = "loess", se = TRUE, color = "blue") +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
- ) +
- labs(
- title = "CC vs. Child Age",
- x = "Child Age at Scan (Months)",
- y = "Total Corpus Callosum Volume"
- ) +
- theme_minimal()
- cc_ulf_sample
- ```
- ```{r}
- # corpus callosum by group (antenatal maternal anaemia status)
- # including repeated measures
- cc_ulf_group <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
- aes(
- x = scan_age_months,
- y = total_cc,
- color = maternal_preg_anaemia_status,
- fill = maternal_preg_anaemia_status
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # slightly transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_mat_anaemia" = "#1f77b4",
- "mat_anaemia" = "#d62728"
- ),
- name = "Maternal Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
- "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
- ),
- guide = "none"
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- scale_y_continuous(
- n.breaks = 6
- ) +
- labs(
- title = "Total CC Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Corpus Callosum Volume (mm³)",
- color = "Maternal Anaemia Status"
- ) +
- theme_minimal()
- cc_ulf_group
- ```
- **Zero-Order Correlation Matrices: Antenatal Maternal Anaemia Status**
- ```{r}
- # install packages if needed
- # install and load Hmisc
- if (!require(Hmisc)) {
- install.packages("Hmisc")
- library(Hmisc)
- } else {
- library(Hmisc)}
- ```
- Convert variables to numeric for correlations
- ```{r}
- # convert variables back from categorical to numeric in ULF_ida_final
- # household income
- ULF_ida_final$income_household_en <- as.numeric(ULF_ida_final$income_household_en)
- # maternal anaemia status: 0 = no anaemia, 1 = anaemia
- ULF_ida_final$maternal_preg_anaemia_status_numeric <- as.numeric(as.character(
- factor(ULF_ida_final$maternal_preg_anaemia_status,
- levels = c("no_mat_anaemia", "mat_anaemia"),
- labels = c(0, 1))
- ))
- # maternal HIV status: 0 = no HIV, 1 = HIV
- ULF_ida_final$maternal_hiv <- as.numeric(as.character(
- factor(ULF_ida_final$maternal_hiv,
- levels = c("no_hiv", "mat_hiv"),
- labels = c(0, 1))
- ))
- # child sex: 0 = girl, 1 = boy
- ULF_ida_final$child_sex <- as.numeric(as.character(
- factor(ULF_ida_final$child_sex,
- levels = c("girl", "boy"),
- labels = c(0, 1))
- ))
- ```
- The variable for "antenatal maternal anaemia status" remained a categorical variable even after running the code above. Therefore, I recalled the dataset to ensure that this was numeric (as initially defined) for modelling:
- ```{r}
- # recall dataset
- # load dataset
- ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
- # create New Variables for Total Brain Volumes across ROIs
- # basal ganglia
- ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
- ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
- # corpus callosum
- ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
- # convert days to months
- # average Number of days in a month (accounting for leap years): 30.44
- ULF_ida_final <- ULF_ida_final %>%
- mutate(scan_age_months = hyp_age / 30.44)
- ```
- **Correlation Matrices**
- ```{r}
- # correlation matrix for corpus callosum
- # filter dataset
- filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_cc")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for caudate nucleus
- # filter dataset
- filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_caudate")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for putamen
- # filter dataset
- filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_putamen")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- ```{r}
- # correlation matrix for icv
- # filter dataset
- filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
- # select relevant columns
- vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "mm_hyp_icv")
- data_selected <- filtered_data[, vars]
- # convert categorical variables to numeric (for correlation)
- data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
- data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
- data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
- # compute correlation matrix with significance
- rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
- # rcorr_result$r contains correlation coefficients
- print("Correlation matrix:")
- print(rcorr_result$r)
- # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
- print("P-value matrix:")
- print(rcorr_result$P)
- ```
- **Multicollinearity**
- ```{r}
- library(dplyr)
- correlation_test <- cor.test(
- ULF_ida_final$scan_age_months[!is.na(ULF_ida_final$maternal_preg_anaemia_status)],
- ULF_ida_final$mm_hyp_icv[!is.na(ULF_ida_final$maternal_preg_anaemia_status)])
- print(correlation_test)
- ```
- Decision to exclude ICV from the models for regional brain volumes due to high levels of multicollinearity between child age at scan. Only child age at scan was included in the models below.
- **Linear Mixed Effects (LME) Models**
- ```{r}
- # install packages if needed
- #install.packages("lme4")
- library(lme4)
- #install.packages("lmerTest")
- library(lmerTest)
- #install.packages("parameters")
- library(parameters)
- #install.packages("emmeans")
- library(emmeans)
- ```
- Based on the exploratory analyses (LOESS curves; plotted data), absolute age was used for models with an observed linear relationship. Log (age) was used for models with an observed non-linear relationship.
- - ICV, Putamen, Caudate Nucleus, Corpus Callosm: Log (age) - Growth trajectory characterised by rapid growth and plateauing on plot
- For the ULF sample, interactions (child age at scan \* antenatal maternal anaemia status) were explored as there was sufficient power. If the interaction was not significant, the model was re-run without an interaction effect and only main effects were interpreted.
- **LME Model variables**
- - Ensure that all categorical variables (anaemia status, child sex, maternal HIV status etc) are factors
- - Create variable for the logarithmic transformation of child age in months
- ```{r}
- # recall dataset so that all input variables are categorical
- # load dataset
- ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
- # create New Variables for Total Brain Volumes across ROIs
- # basal ganglia
- ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
- ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
- # corpus callosum
- ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
- # create age variable
- # convert days to months
- # average Number of days in a month (accounting for leap years): 30.44
- ULF_ida_final <- ULF_ida_final %>%
- mutate(scan_age_months = hyp_age / 30.44)
- ```
- ```{r}
- # convert categorical variables form numeric to factor
- # maternal anaemia status
- ULF_ida_final$maternal_preg_anaemia_status <- factor(
- ULF_ida_final$maternal_preg_anaemia_status,
- levels = c(0, 1),
- labels = c("no_mat_anaemia", "mat_anaemia")
- )
- # maternal anaemia severity
- ULF_ida_final$maternal_preg_anaemia_severity <- factor(
- ULF_ida_final$maternal_preg_anaemia_severity,
- levels = c(1, 2, 3),
- labels = c("mild", "moderate", "severe")
- )
- # child status
- ULF_ida_final$cor_child_anaemia <- factor(
- ULF_ida_final$cor_child_anaemia,
- levels = c(0, 1),
- labels = c("no_child_anaemia", "child_anaemia")
- )
- # maternal ID status
- ULF_ida_final$maternal_preg_adj_brinda_id <- factor(
- ULF_ida_final$maternal_preg_adj_brinda_id,
- levels = c(0, 1),
- labels = c("no_mat_id", "mat_id")
- )
- # child ID status
- ULF_ida_final$cor_child_id <- factor(
- ULF_ida_final$cor_child_id,
- levels = c(0, 1),
- labels = c("no_child_id", "child_id")
- )
- # maternal trimester hb
- ULF_ida_final$maternal_trimester_hb <- factor(
- ULF_ida_final$maternal_trimester_hb,
- levels = c(1, 2, 3),
- labels = c("first", "second", "third")
- )
- # child sex
- ULF_ida_final$child_sex <- factor(
- ULF_ida_final$child_sex,
- levels = c(0, 1),
- labels = c("girl", "boy")
- )
- # maternal pae
- ULF_ida_final$maternal_pae <- factor(
- ULF_ida_final$maternal_pae,
- levels = c(0, 1),
- labels = c("no_pae", "pae")
- )
- # maternal pte
- ULF_ida_final$maternal_pte <- factor(
- ULF_ida_final$maternal_pte,
- levels = c(0, 1),
- labels = c("no_pte", "pte")
- )
- # maternal hiv
- ULF_ida_final$maternal_hiv <- factor(
- ULF_ida_final$maternal_hiv,
- levels = c(0, 1),
- labels = c("no_hiv", "hiv")
- )
- # maternal dep
- ULF_ida_final$maternal_dep <- factor(
- ULF_ida_final$maternal_dep,
- levels = c(0, 1),
- labels = c("no_dep", "dep")
- )
- # maternal employment
- ULF_ida_final$mom_employment_en_dichotomous <- factor(
- ULF_ida_final$mom_employment_en_dichotomous,
- levels = c(0, 1),
- labels = c("unemployed", "employed")
- )
- # household income
- ULF_ida_final$income_household_en <- factor(
- ULF_ida_final$income_household_en,
- levels = c(1, 2, 3, 4),
- labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
- )
- # maternal education
- ULF_ida_final$mom_edu_en <- factor(
- ULF_ida_final$mom_edu_en,
- levels = c(2, 3, 4, 5, 6),
- labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
- )
- #baby hiv
- ULF_ida_final$baby_hiv <- factor(
- ULF_ida_final$baby_hiv,
- levels = c(0, 1),
- labels = c("no_child_hiv", "child_hiv")
- )
- ```
- ```{r}
- # create a log(age) variable for use in models of brain regions with non-linear growth
- ULF_ida_final$log_age <- log(ULF_ida_final$scan_age_months)
- ```
- *LME: Corpus Callosum*
- ```{r}
- # lme for corpus callosum
- # including interaction between maternal anaemia and age
- # log (age): non-linear growth
- #model summary
- ulf_model_cc_interaction <- lmer(total_cc ~ maternal_preg_anaemia_status * log_age +
- child_sex +
- (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_cc_interaction)
- # beta values
- standardize_parameters(ulf_model_cc_interaction)
- ```
- | |
- |-----|
- | |
- ```{r}
- # emms for corpus callosum volume by antenatal maternal anaemia status
- # based on final adjusted model with interaction between maternal anaemia and age
- # by timepoint for significant interaction
- library(emmeans)
- library(dplyr)
- # step 1: Define raw ages and their logs for EMMs
- raw_ages <- c(3, 6, 12, 18, 24)
- log_ages <- log(raw_ages)
- # step 2: Get estimated marginal means at these log ages
- emm <- emmeans(
- ulf_model_cc_interaction,
- specs = ~ maternal_preg_anaemia_status | log_age,
- at = list(log_age = log_ages)
- )
- # step 3: Convert to data frame and add back-transformed age for interpretability
- emm_cc <- as.data.frame(emm)
- emm_cc$age_months <- round(exp(emm_cc$log_age), 0)
- print(emm_cc)
- # step 4: Run pairwise contrasts (differences between anaemia groups) at each age
- contrast_cc <- contrast(
- emm,
- method = "revpairwise",
- adjust = "holm",
- infer = TRUE
- ) %>%
- as.data.frame()
- # step 5: Add back-transformed age for contrasts
- contrast_cc$age_months <- round(exp(contrast_cc$log_age), 0)
- print(contrast_cc)
- ```
- *LME: ICV*
- ```{r}
- # lme for icv
- # including nteraction between maternal anaemia and age
- # log (age): non-linear growth
- # model summary
- ulf_model_icv_interaction <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_icv_interaction)
- # beta values
- standardize_parameters(ulf_model_icv_interaction)
- ```
- Interaction was not significant, so re-ran the mdoel without it and interpreted main effects:
- ```{r}
- # lme for icv
- # excluding interaction between maternal anaemia and age
- # log (age): non-linear growth
- # model summary
- ulf_model_icv_no_interaction <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_icv_no_interaction)
- # beta values
- standardize_parameters(ulf_model_icv_no_interaction)
- ```
- ```{r}
- # emms for icv volume by antenatal maternal anaemia status
- # based on final adjusted model with no interaction
- library(emmeans)
- emm_icv <- emmeans(ulf_model_icv_no_interaction, ~ maternal_preg_anaemia_status)
- # ----- EMM table with 2 decimals -----
- emm_icv_df <- as.data.frame(emm_icv)
- emm_icv_df$emmean <- format(round(emm_icv_df$emmean, 2), nsmall = 2)
- emm_icv_df$SE <- format(round(emm_icv_df$SE, 2), nsmall = 2)
- emm_icv_df$lower.CL <- format(round(emm_icv_df$lower.CL, 2), nsmall = 2)
- emm_icv_df$upper.CL <- format(round(emm_icv_df$upper.CL, 2), nsmall = 2)
- print(emm_icv_df)
- # ----- contrast results -----
- contrast_results <- contrast(emm_icv, method = "revpairwise", adjust = "none")
- contrast_df <- as.data.frame(summary(contrast_results, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df$estimate <- format(round(contrast_df$estimate, 2), nsmall = 2)
- print(contrast_df)
- ```
- *LME: Caudate Nucleus*
- ```{r}
- # lme for caudate nucleus
- # icluding interaction between maternal anaemia and age
- # log (age): non-linear growth
- # model summary
- ulf_model_caudate_interaction <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_caudate_interaction)
- # beta values
- standardize_parameters(ulf_model_caudate_interaction)
- ```
- ```{r}
- # emms for caudate volume by antenatal maternal anaemia status
- # based on final adjusted model with interaction
- # by timepoint for significant interaction
- library(emmeans)
- library(dplyr)
- # step 1: Define raw ages and their logs for EMMs
- raw_ages <- c(3, 6, 12, 18, 24)
- log_ages <- log(raw_ages)
- # step 2: Get estimated marginal means at these log ages
- emm <- emmeans(
- ulf_model_caudate_interaction,
- specs = ~ maternal_preg_anaemia_status | log_age,
- at = list(log_age = log_ages)
- )
- # step 3: Convert to data frame and add back-transformed age for interpretability
- emm_caudate <- as.data.frame(emm)
- emm_caudate$age_months <- round(exp(emm_caudate$log_age), 0)
- print(emm_caudate)
- # step 4: Run pairwise contrasts (differences between anaemia groups) at each age
- contrast_caudate <- contrast(emm, method = "revpairwise", adjust = "holm", infer = TRUE) %>% as.data.frame()
- # step 5: Add back-transformed age for contrasts (will be in 'log_age' column)
- contrast_caudate$age_months <- round(exp(contrast_caudate$log_age), 0)
- print(contrast_caudate)
- ```
- *LME: Putamen*
- ```{r}
- # lme for putamen
- # including interaction between maternal anaemia and age
- # log (age): non-linear growth
- # model summary
- ulf_model_putamen_interaction <- lmer(total_putamen ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_putamen_interaction)
- # beta values
- standardize_parameters(ulf_model_putamen_interaction)
- ```
- Interaction was not significant, so re-ran the model without it and interpreted main effects:
- ```{r}
- # lme for putamen
- # excluding interaction between maternal anaemia and age
- # model summary
- ulf_model_putamen_no_interaction <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_putamen_no_interaction)
- # beta values
- standardize_parameters(ulf_model_putamen_no_interaction)
- ```
- ```{r}
- # emms for putamen volume by antenatal maternal anaemia status
- # based on final adjusted model with no interaction
- library(emmeans)
- emm_putamen <- emmeans(ulf_model_putamen_no_interaction, ~ maternal_preg_anaemia_status)
- # ----- emm table with exactly 2 decimals -----
- emm_putamen_df <- as.data.frame(emm_putamen)
- emm_putamen_df$emmean <- format(round(emm_putamen_df$emmean, 2), nsmall = 2)
- emm_putamen_df$SE <- format(round(emm_putamen_df$SE, 2), nsmall = 2)
- emm_putamen_df$lower.CL <- format(round(emm_putamen_df$lower.CL, 2), nsmall = 2)
- emm_putamen_df$upper.CL <- format(round(emm_putamen_df$upper.CL, 2), nsmall = 2)
- print(emm_putamen_df)
- # ----- contrast results -----
- contrast_results_putamen <- contrast(emm_putamen, method = "revpairwise", adjust = "none")
- contrast_df_putamen <- as.data.frame(summary(contrast_results_putamen, infer = TRUE))
- # force 2 decimals for estimate
- contrast_df_putamen$estimate <- format(round(contrast_df_putamen$estimate, 2), nsmall = 2)
- print(contrast_df_putamen)
- ```
- **Sensitivity Analyses**
- Based on the final models for each ROI (established in primary analyses): With the inclusion of maternal HIV (potential confounder)
- - ICV: Final model was one without an interaction
- - Putamen: FInal model was without an interaction
- - Corpus Callosum: Final model was with an interaction
- - Caudate Nucleus: Final model was with an interaction
- ```{r}
- # sensitivity analysis for ICV
- # final model: excluding interaction between maternal anaemia and age
- # including maternal hiv
- # model summary
- ulf_model_icv_no_interaction_sens <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_icv_no_interaction_sens)
- # beta values
- standardize_parameters(ulf_model_icv_no_interaction_sens)
- ```
- ```{r}
- # sensitivity analysis for corpus callosum
- # final model: including interaction between maternal anaemia and age
- # including maternal hiv
- # model summary
- ulf_model_cc_interaction_sens <- lmer(total_cc ~ maternal_preg_anaemia_status * scan_age_months + child_sex + maternal_hiv + (1 | subject_id), data = ULF_ida_final)
- summary(ulf_model_cc_interaction_sens)
- # beta values
- standardize_parameters(ulf_model_cc_interaction_sens)
- ```
- ```{r}
- # sensitivity analysis for putamen
- # final model: excluding interaction between maternal anaemia and age
- # including maternal hiv
- # model summary
- ulf_model_putamen_no_interaction_sens <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_putamen_no_interaction_sens)
- # beta values
- standardize_parameters(ulf_model_putamen_no_interaction_sens)
- ```
- ```{r}
- # sensitivity analysis for caudate nucleus
- # final model: including interaction between maternal anaemia and age
- # including maternal hiv
- # model summary
- ulf_model_caudate_interaction_sens <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + maternal_hiv + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_caudate_interaction_sens)
- # beta values
- standardize_parameters(ulf_model_caudate_interaction_sens)
- ```
- ------------------------------------------------------------------------
- **Secondary Analysis: Child Anaemia (ULF Sample)**
- ------------------------------------------------------------------------
- **Sample Size: For infants with Child Anemia Data**
- Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
- ```{r}
- # total sample size
- # sample iwth child anaemia data
- # repeated measures included
- ULF_ida_final %>%
- filter(!is.na(cor_child_anaemia)) %>% # only include PIDs with child anaemia data
- summarise(n = n())
- ```
- ```{r}
- # sample with child anaemia data
- # repeated measures excluded
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(cor_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE) %>% #only include PIDs with child anaemia data
- summarise(n = n())
- length(unique(ULF_ida_final$subject_id)) # includes chidlren with no anaemia data
- ```
- **Child Anaemia Variables**
- An overall classification for child anaemia status was determined based on the diagnosis of anaemia at least once across study visits. This was used for the assessment of group differences in sample characteristics between infants who had ever been anaemic versus infants who had never been anaemic between 3 and 24 months of age. Minimum child haemoglobin and corresponding age across study visits were also determined for each infant and reported on across both groups. For modelling with the repeated measures subsample (multiple scans per child in some instances), child anaemia status was assessed based on the most severe presentation of anaemia as indicated by the minimum haemoglobin measure observed between birth and the time of scan.
- Each infant has multiple haemoglobin measurements from across study timepoints.
- - "Overall Anaemia" variable: For the purpose of assessing group differences in sample characteristics based on unqiue subject IDs (excluding repeated measures), we created a binary variable for child anaemia status based on being anaemic at least once between 3 and 24 months. i.e. "overall anaemia" = "anaemic at any timepoint"
- - Child anaemia status in repeated measures models (LMEs): Every scan has a corresponding child anaemia status variable (cor_child_anaemia) based on the minimum Hb measurement (cor_child_hb) between birth and the time of scan and the corresponding age (cor_child_hb_age) at the time of this measurement. i.e. In a repeated measures sample:
- - 3-month scan: anaemia status based on Hb at 3 months
- - 6-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3 and 6 months)
- - 12-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, and 12 months)
- - 18-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, 12, and 18 months)
- - 24-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, 12, 18, and 24 months)
- - Minimum haemoglobin measurement: This was computed to compare Hb across groups when assessing sample characteristics based on unique subjects (i.e. excluding repeated measures). We identified the minimum Hb measurement (and corresponding age) across study timepoints - i.e. across multiple cor_child_hb measures for children with multiple scans in a repeated measures sample.
- **Creating New Binary Variable for Overall Anaemia Status**
- ```{r}
- # overall child anaemia status variable
- # step 1: Create new dataframe with only child anaemia data. Select only anaemia variables interested in
- # step 2: Pivot to wide format so each child ID appears once only (see all timepoints for each child)
- dem_child <- ULF_ida_final %>% #create new dataframe
- filter(!is.na(cor_child_anaemia)) %>% #filter so excluding children with no child anaemia data
- select(subject_id, timepoint, cor_child_hb, cor_child_anaemia) %>% #select variables interested in
- pivot_wider(id_cols = subject_id, names_from = timepoint, values_from = c(cor_child_hb,cor_child_anaemia)) #convert to wide format so can see anaemia data fro all timepoints
- ```
- ```{r}
- # overall child anaemia status variable
- # step 3: create overall anaemia variable - i.e. "if anaemic at any timepoint"
- dem_child <- dem_child %>%
- mutate(overall_child_anaemia = if_else(if_any(c(cor_child_anaemia_3M, cor_child_anaemia_6M, cor_child_anaemia_12M, cor_child_anaemia_18M, cor_child_anaemia_24M), ~ !is.na(.x) & .x == "child_anaemia"), "overall_child_anaemia", "no_overall_child_anaemia")) %>%
- select(subject_id, overall_child_anaemia) #select and save the overall anaemia status
- # step 4: Merge the overall_child_anaemia variable/column back into long format of the original dataframe - i.e. "ULF_ida_final"
- ULF_ida_final <- left_join(ULF_ida_final, dem_child, by = 'subject_id')
- ```
- ```{r}
- # overall child anaemia status variable
- # sanity check that have the correct IDs
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE) %>% #only include PIDs with child anaemia data
- summarise(n = n())
- length(unique(ULF_ida_final$subject_id)) # inlcudes chidlren with no anaemia data
- ```
- **Creating Variable for Minimum Hb Across Study Timepoints**
- ```{r}
- # minimum child Hb variable
- # step 1: Create new dataframe
- # step 2: Pivot to wide format
- dem_child_1 <- ULF_ida_final %>% #create new dataframe
- filter(!is.na(cor_child_anaemia)) %>% #filter so excluding children with no child anaemia data
- select(subject_id, timepoint, cor_child_hb, cor_child_hb_age, cor_child_anaemia) %>% #select variables interested in
- pivot_wider(id_cols = subject_id, names_from = timepoint, values_from = c(cor_child_hb,cor_child_hb_age, cor_child_anaemia)) #convert to wide format so can see anaemia data for all timepoints
- ```
- ```{r}
- # step 3: Create min hb variable - min hb from across timepoints
- library(dplyr)
- dem_child_1 <- dem_child_1 %>%
- # create a new variable that represents the minimum child Hb across timepoints
- mutate(
- min_child_hb = pmin(cor_child_hb_3M, cor_child_hb_6M, cor_child_hb_12M, cor_child_hb_18M, cor_child_hb_24M, na.rm = TRUE),
- # create a new variable that represents the age at the time of the minimum Hb
- min_hb_age = case_when(
- min_child_hb == cor_child_hb_3M ~ cor_child_hb_age_3M,
- min_child_hb == cor_child_hb_6M ~ cor_child_hb_age_6M,
- min_child_hb == cor_child_hb_12M ~ cor_child_hb_age_12M,
- min_child_hb == cor_child_hb_18M ~ cor_child_hb_age_18M,
- min_child_hb == cor_child_hb_24M ~ cor_child_hb_age_24M
- )
- ) %>%
- # select subject_id, min_child_hb, and min_hb_age variables
- select(subject_id, min_child_hb, min_hb_age)
- # view the resulting dataframe
- head(dem_child_1)
- ```
- ```{r}
- # step 4: merge in the min Hb and cor age at min hb variable back into long format original dataframe "ULF_ida_final"
- ULF_ida_final <- left_join(ULF_ida_final, dem_child_1, by = 'subject_id')
- ```
- **Anaemia Prevalence**
- ```{r}
- # child anaemia prevalence
- # based on new variable for overall anaemia status
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(overall_child_anaemia) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # maternal anaemia prevalence
- # for infants with child anaemia data
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_preg_anaemia_status) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- **Sample Characteristics and Group Differences: Child Anaemia Status**
- Sample characteristics were assessed based on the number of unique subjects with child anaemia (haemoglobin) data, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
- An overall classification for child anaemia status was determined based on the diagnosis of anaemia at least once across study visits. This was used for the assessment of group differences in sample characteristics between infants who had ever been anaemic versus infants who had never been anaemic between 3 and 24 months of age. Minimum child haemoglobin and corresponding age across study visits were also determined for each infant and reported on across both groups. For modelling with the repeated measures subsample (multiple scans per child in some instances), child anaemia status was assessed based on the most severe presentation of anaemia as indicated by the minimum haemoglobin measure observed between birth and the time of scan.
- - **Categorical Variables**
- *Maternal Anaemia*
- ```{r}
- # group differences in maternal anaemia by child anaemia status (using overall child anaemia variable created)
- # unique subject IDs: based on true prevalence
- # filter only for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # create contingency table: maternal anaemia vs child anaemia
- tbl <- table(filtered$maternal_preg_anaemia_status, filtered$overall_child_anaemia)
- # print table
- cat("Contingency Table: Maternal anaemia status by Child Anaemia Status\n")
- print(tbl)
- # column percentages (percent within child anaemia status group)
- cat("\nColumn Percentages (% within child anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # choose appropriate test
- chisq_result <- chisq.test(tbl)
- if (any(chisq_result$expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq_result)
- }
- ```
- *Child Sex at Birth*
- ```{r}
- # child sex at birth (full sample)
- # unique subject IDs: Based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(child_sex) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in child sex by child anaemia status
- # chi-squared: child sex
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: child sex vs child anaemia
- tbl <- table(filtered$child_sex, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Child Sex by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each anaemia status group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Prenatal Alcohol Exposure (PAE)*
- ```{r}
- # pae (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_pae) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pae by child anaemia status
- # chi-squared: pae
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal pae vs child anaemia
- tbl <- table(filtered$maternal_pae, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Maternal PAE by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Prenatal Tobacco Exposure (PTE)*
- ```{r}
- # pte (full sample)
- # excluding repeated measures: unique id
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_pte) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in pte by child anaemia status
- # chi-squared: pte
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal PTE vs child anaemia
- tbl <- table(filtered$maternal_pte, filtered$overall_child_anaemia)
- # 3.print contingency table
- cat("Contingency Table: Maternal PTE by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal HIV*
- ```{r}
- # maternal hiv (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal hiv by child anaemia status
- # chi-squared: maternal hiv
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal HIV vs child anaemia
- tbl <- table(filtered$maternal_hiv, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Maternal HIV by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Depression*
- ```{r}
- # maternal depression (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(maternal_dep) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal depression by child anaemia status
- # chi-squared: maternal depression
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal DEP vs child anaemia
- tbl <- table(filtered$maternal_dep, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Maternal DEP by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Employment*
- ```{r}
- # maternal employment (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(mom_employment_en_dichotomous) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal employment by child anaemia status
- # chi-squared: maternal employment
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal employment vs child anaemia
- tbl <- table(filtered$mom_employment_en_dichotomous, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Maternal Employment by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Maternal Education*
- ```{r}
- # maternal education (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(mom_edu_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in maternal education by child anaemia status
- # chi-squared: maternal education
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: maternal education vs child anaemia
- tbl <- table(filtered$mom_edu_en, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Maternal Education by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Household Income*
- ```{r}
- # household income (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(income_household_en) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in household income by child anaemia status
- # chi-squared: household income
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: household income vs child anaemia
- tbl <- table(filtered$income_household_en, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Household Income by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- *Infant HIV*
- ```{r}
- # infant hiv (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE) %>%
- group_by(baby_hiv) %>%
- summarise(n = n()) %>%
- mutate(percent = round(100 * n / sum(n), 2))
- ```
- ```{r}
- # group differences in infant hiv by child anaemia status
- # chi-squared: infant hiv
- # unique subject IDs: based on true prevalence
- # 1. filter for non-missing child anaemia status and keep unique subjects
- filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
- filtered <- filtered[!duplicated(filtered$subject_id), ]
- # 2. create contingency table: infant HIV vs child anaemia
- tbl <- table(filtered$baby_hiv, filtered$overall_child_anaemia)
- # 3. print contingency table
- cat("Contingency Table: Infant HIV by Child Anaemia Status\n")
- print(tbl)
- # 4. column percentages (% within each child anaemia group)
- cat("\nColumn Percentages (% of each anaemia status group):\n")
- print(round(100 * prop.table(tbl, margin = 2), 2))
- # 5. run appropriate test
- expected <- chisq.test(tbl)$expected
- if (any(expected < 5)) {
- cat("\nFisher's Exact Test (due to small expected cell counts):\n")
- print(fisher.test(tbl))
- } else {
- cat("\nChi-squared Test:\n")
- print(chisq.test(tbl))
- }
- ```
- - **Continuous Variables**
- ```{r}
- #install packages if needed
- #install.packages('car')
- library (car)
- ```
- *Minimum Haemoglobin Measure*
- ```{r}
- # group differences in minimum Hb by child anaemia status
- # ttest: min Hb
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car) # for Levene's Test
- # create cleaned dataset once (avoids repeating code)
- child_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: min child haemoglobin by child anaemia status
- summary_stats <- child_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene’s test for homogeneity of variances
- levene_result <- leveneTest(min_child_hb ~ overall_child_anaemia, data = child_data)
- print(levene_result)
- # extract Levene p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # run appropriate t-test (full precision output)
- if (levene_p < 0.05) {
- cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
- t_result <- t.test(
- min_child_hb ~ overall_child_anaemia,
- data = child_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
- t_result <- t.test(
- min_child_hb ~ overall_child_anaemia,
- data = child_data,
- var.equal = TRUE
- )
- }
- # print t-test results with full precision
- print(t_result)
- ```
- *Child Age at Minimum Haemoglobin Measurement*
- ```{r}
- # group differences in age at Hb measurement by child anaemia status
- # ttest: age at Hb measurement
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car) # for Levene's Test
- # prepare cleaned dataset once
- child_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: child age at Hb measurement by anaemia status
- summary_stats <- child_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene’s test for homogeneity of variances
- levene_result <- leveneTest(min_hb_age ~ overall_child_anaemia, data = child_data)
- print(levene_result)
- # extract Levene p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # run appropriate t-test (full precision output)
- if (levene_p < 0.05) {
- cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
- t_result <- t.test(
- min_hb_age ~ overall_child_anaemia,
- data = child_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
- t_result <- t.test(
- min_hb_age ~ overall_child_anaemia,
- data = child_data,
- var.equal = TRUE
- )
- }
- # Print t-test results in full precision
- print(t_result)
- ```
- *Maternal Age at Enrolment*
- ```{r}
- # maternal age at enrolment for full group
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- mean_mom_age = mean(mom_age_en, na.rm = TRUE),
- sd_mom_age = sd(mom_age_en, na.rm = TRUE),
- min_mom_age = min(mom_age_en, na.rm = TRUE),
- max_mom_age = max(mom_age_en, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in maternal age at enrolment by child anaemia status
- # ttest: maternal age at enrolment
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car) # for Levene's Test
- # deduplicate and filter data once
- child_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: min child Hb by anaemia status (2 decimals)
- summary_stats <- child_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_hb = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- sd_hb = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- min_hb = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- max_hb = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene’s test for homogeneity of variances
- levene_result <- leveneTest(mom_age_en ~ overall_child_anaemia, data = child_data)
- print(levene_result)
- # extract Levene p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # run t-test (full precision)
- if (levene_p < 0.05) {
- cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
- t_result <- t.test(mom_age_en ~ overall_child_anaemia,
- data = child_data,
- var.equal = FALSE)
- } else {
- cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
- t_result <- t.test(mom_age_en ~ overall_child_anaemia,
- data = child_data,
- var.equal = TRUE)
- }
- # print t-test results
- print(t_result)
- ```
- *Child Age at Scan*
- ```{r}
- # child age at scan (full sample)
- # include repeated measures - for all scans across study timepoints
- ULF_ida_final %>%
- filter(!is.na(cor_child_anaemia)) %>% # only include PIDs with child anaemia data
- summarise(
- mean_child_age = mean(scan_age_months, na.rm = TRUE),
- sd_child_age = sd(scan_age_months, na.rm = TRUE),
- min_child_age = min(scan_age_months, na.rm = TRUE),
- max_child_age = max(scan_age_months, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in child age at scan by child anaemia status
- # ttest: child age at scan
- # all scans included
- library(dplyr)
- library(car)
- # deduplicate and filter data once
- child_data <- ULF_ida_final %>%
- filter(!is.na(cor_child_anaemia))
- # summary statistics: scan age by child anaemia status (2 decimals)
- summary_stats <- child_data %>%
- group_by(cor_child_anaemia) %>%
- summarise(
- n = n(),
- mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene's Test for homogeneity of variance
- levene_result <- leveneTest(scan_age_months ~ cor_child_anaemia, data = child_data)
- print(levene_result)
- # extract Levene p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically choose correct t-test
- if (levene_p < 0.05) {
- cat("\nLevene's test significant (p =", round(levene_p, 4), ") → Using Welch t-test (var.equal = FALSE)\n\n")
- t_result <- t.test(scan_age_months ~ cor_child_anaemia,
- data = child_data,
- var.equal = FALSE)
- } else {
- cat("\nLevene's test not significant (p =", round(levene_p, 4), ") → Using Student t-test (var.equal = TRUE)\n\n")
- t_result <- t.test(scan_age_months ~ cor_child_anaemia,
- data = child_data,
- var.equal = TRUE)
- }
- # print t-test results in full precision
- print(t_result)
- ```
- *Child Gestational Age (GA) at Birth*
- ```{r}
- # child ga at birth (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- mean_child_ga = mean(ga_weeks, na.rm = TRUE),
- sd_child_ga = sd(ga_weeks, na.rm = TRUE),
- min_child_ga = min(ga_weeks, na.rm = TRUE),
- max_child_ga = max(ga_weeks, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in ga at birth by child anaemia status
- # ttest: ga
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car) # for Levene's Test
- # deduplicate and filter once
- child_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: GA by child anaemia status (2 decimals)
- summary_stats <- child_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene's Test for homogeneity of variance
- levene_result <- leveneTest(ga_weeks ~ overall_child_anaemia, data = child_data)
- print(levene_result)
- # extract Levene p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically choose correct t-test
- if (levene_p < 0.05) {
- cat("\nLevene's test significant (p =", round(levene_p, 4),
- ") → Using Welch t-test (var.equal = FALSE)\n\n")
- t_result <- t.test(
- ga_weeks ~ overall_child_anaemia,
- data = child_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test not significant (p =", round(levene_p, 4),
- ") → Using Student t-test (var.equal = TRUE)\n\n")
- t_result <- t.test(
- ga_weeks ~ overall_child_anaemia,
- data = child_data,
- var.equal = TRUE
- )
- }
- # print t-test result in full precision
- print(t_result)
- ```
- *Infant Birth Weight (BW)*
- ```{r}
- # child bw (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- mean_bw = mean(bw, na.rm = TRUE),
- sd_bw = sd(bw, na.rm = TRUE),
- min_bw = min(bw, na.rm = TRUE),
- max_bw = max(bw, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in bw by child anaemia status
- # ttest: bw
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # deduplicate and filter once
- birth_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: birth weight by child anaemia status (2 decimals)
- summary_stats <- birth_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
- sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
- min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
- max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(bw ~ overall_child_anaemia, data = birth_data)
- print(levene_result)
- # extract p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically choose correct t-test
- if (levene_p < 0.05) {
- cat("\nLevene's test significant (p =", round(levene_p, 4),
- ") → Using Welch t-test (var.equal = FALSE)\n\n")
- t_test_result <- t.test(
- bw ~ overall_child_anaemia,
- data = birth_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test not significant (p =", round(levene_p, 4),
- ") → Using Student t-test (var.equal = TRUE)\n\n")
- t_test_result <- t.test(
- bw ~ overall_child_anaemia,
- data = birth_data,
- var.equal = TRUE
- )
- }
- # print t-test result in full precision
- print(t_test_result)
- ```
- *Infant Birth Length (BL)*
- ```{r}
- # child bl (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- mean_birthlength = mean(birth_length, na.rm = TRUE),
- sd_birthlength = sd(birth_length, na.rm = TRUE),
- min_birthlength = min(birth_length, na.rm = TRUE),
- max_birthlength = max(birth_length, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in bl by child anaemia status
- # ttest: bl
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # deduplicate and filter once
- birth_length_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: birth length by child anaemia status (2 decimals)
- summary_stats <- birth_length_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
- sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
- min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
- max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(birth_length ~ overall_child_anaemia, data = birth_length_data)
- print(levene_result)
- # extract p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically choose correct t-test
- if (levene_p < 0.05) {
- cat("\nLevene's test significant (p =", round(levene_p, 4),
- ") → Using Welch t-test (var.equal = FALSE)\n\n")
- t_test_result <- t.test(
- birth_length ~ overall_child_anaemia,
- data = birth_length_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test not significant (p =", round(levene_p, 4),
- ") → Using Student t-test (var.equal = TRUE)\n\n")
- t_test_result <- t.test(
- birth_length ~ overall_child_anaemia,
- data = birth_length_data,
- var.equal = TRUE
- )
- }
- # print t-test result in full precision
- print(t_test_result)
- ```
- *Infant Head Circumference (HC)*
- ```{r}
- # child birth hc (full sample)
- # unique subject IDs: based on true prevalence
- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
- distinct(subject_id, .keep_all = TRUE) %>%
- summarise(
- mean_hc = mean(birth_hc, na.rm = TRUE),
- sd_hc = sd(birth_hc, na.rm = TRUE),
- min_hc = min(birth_hc, na.rm = TRUE),
- max_hc = max(birth_hc, na.rm = TRUE),
- n = n()
- )
- ```
- ```{r}
- # group differences in hc at birth by child anaemia status
- # ttest: hc at birth
- # unique subject IDs: based on true prevalence
- library(dplyr)
- library(car)
- # deduplicate and filter once
- birth_hc_data <- ULF_ida_final %>%
- filter(!is.na(overall_child_anaemia)) %>%
- distinct(subject_id, .keep_all = TRUE)
- # summary statistics: birth head circumference by child anaemia status (2 decimals)
- summary_stats <- birth_hc_data %>%
- group_by(overall_child_anaemia) %>%
- summarise(
- n = n(),
- mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
- .groups = "drop"
- )
- print(summary_stats)
- # levene's test for homogeneity of variance
- levene_result <- leveneTest(birth_hc ~ overall_child_anaemia, data = birth_hc_data)
- print(levene_result)
- # extract p-value
- levene_p <- levene_result$`Pr(>F)`[1]
- # automatically choose correct t-test
- if (levene_p < 0.05) {
- cat("\nLevene's test significant (p =", round(levene_p, 4),
- ") → Using Welch t-test (var.equal = FALSE)\n\n")
- t_test_result <- t.test(
- birth_hc ~ overall_child_anaemia,
- data = birth_hc_data,
- var.equal = FALSE
- )
- } else {
- cat("\nLevene's test not significant (p =", round(levene_p, 4),
- ") → Using Student t-test (var.equal = TRUE)\n\n")
- t_test_result <- t.test(
- birth_hc ~ overall_child_anaemia,
- data = birth_hc_data,
- var.equal = TRUE
- )
- }
- # print t-test result in full precision
- print(t_test_result)
- ```
- **Exploratory Analyses: Plotted Data using LOESS Curves**
- *LOESS Curve: ICV*
- ```{r}
- # icv
- library(ggplot2)
- library(dplyr)
- icv_ulf_child <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
- aes(
- x = scan_age_months,
- y = mm_hyp_icv,
- color = cor_child_anaemia,
- fill = cor_child_anaemia
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE,
- level = 0.95,
- alpha = 0.3 # semi-transparent ribbon
- ) +
- scale_color_manual(
- values = c(
- "no_child_anaemia" = "#1f77b4", # dark blue line
- "child_anaemia" = "#d62728" # dark red line
- ),
- name = "Child Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
- "child_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 10)
- ) +
- labs(
- title = "Total ICV Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total ICV Volume (mm3)",
- color = "Child Anaemia Status"
- ) +
- theme_minimal()
- icv_ulf_child
- ```
- *LOESS Curve: Corpus Callosum*
- ```{r}
- # corpus callosum
- cc_ulf_child <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
- aes(
- x = scan_age_months,
- y = total_cc,
- color = cor_child_anaemia,
- fill = cor_child_anaemia
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE, # show 95% confidence ribbons
- level = 0.95,
- alpha = 0.3 # make ribbons semi-transparent
- ) +
- scale_color_manual(
- values = c(
- "no_child_anaemia" = "#1f77b4", # dark blue
- "child_anaemia" = "#d62728" # dark red
- ),
- name = "Child Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
- "child_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- scale_y_continuous(
- breaks = seq(
- floor(min(ULF_ida_final$total_cc, na.rm = TRUE) / 10) * 10,
- ceiling(max(ULF_ida_final$total_cc, na.rm = TRUE) / 10) * 10,
- by = 4000
- )
- ) +
- labs(
- title = "Total CC Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total CC Volume (mm3)",
- color = "Child Anaemia Status"
- ) +
- theme_minimal()
- cc_ulf_child
- ```
- *LOESS Curve: Caudate Nucleus*
- ```{r}
- # caudate nucleus
- caudate_ulf_child <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
- aes(
- x = scan_age_months,
- y = total_caudate,
- color = cor_child_anaemia,
- fill = cor_child_anaemia
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE, # show 95% confidence ribbons
- level = 0.95,
- alpha = 0.3 # semi-transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_child_anaemia" = "#1f77b4", # dark blue
- "child_anaemia" = "#d62728" # dark red
- ),
- name = "Child Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
- "child_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- scale_y_continuous(
- n.breaks = 5 # adjust y-axis ticks
- ) +
- labs(
- title = "Total Caudate Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Caudate Volume (mm³)",
- color = "Child Anaemia Status"
- ) +
- theme_minimal()
- caudate_ulf_child
- ```
- *LOESS Curve: Putamen*
- ```{r}
- # putamen
- putamen_ulf_child <- ggplot(
- data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
- aes(
- x = scan_age_months,
- y = total_putamen,
- color = cor_child_anaemia,
- fill = cor_child_anaemia
- )
- ) +
- geom_point() +
- geom_smooth(
- method = "loess",
- se = TRUE, # show 95% confidence ribbons
- level = 0.95,
- alpha = 0.3 # semi-transparent ribbons
- ) +
- scale_color_manual(
- values = c(
- "no_child_anaemia" = "#1f77b4", # dark blue
- "child_anaemia" = "#d62728" # dark red
- ),
- name = "Child Anaemia Status"
- ) +
- scale_fill_manual(
- values = c(
- "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
- "child_anaemia" = "#f7b0b0" # lighter red ribbon
- ),
- guide = "none" # hide fill legend
- ) +
- scale_x_continuous(
- breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
- ) +
- labs(
- title = "Total Putamen Volume vs. Scan Age",
- x = "Scan Age (Months)",
- y = "Total Putamen Volume (mm³)",
- color = "Child Anaemia Status"
- ) +
- theme_minimal()
- putamen_ulf_child
- ```
- **Linear Mixed Effects (LME) Models: Child Anaemia Status Only**
- First, the LMEs described above for maternal anaemia were replicated with postnatal child anaemia as the primary predictor of interest. This series of models were run to assess the independent associations between child anaemia and regional child brain volumes.
- Based on the LOESS curves, there were no observed age-related differences between groups based on child anaemia status. Therefore, interactions were excluded to simplify the models.
- --\> Main effects only (with child anaemia as the primary predictor)
- Log (age) was used for all models based on non-linear growth trajectories
- ```{r}
- # lme: icv
- # main effects: child anaemia
- # log (age)
- # model summary
- ulf_model_icv_child <- lmer(mm_hyp_icv ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_icv_child)
- # beta values
- standardize_parameters(ulf_model_icv_child)
- ```
- ```{r}
- # lme: cc
- # main effects: child anaemia
- # log (age)
- # model summary
- ulf_model_cc_child <- lmer(total_cc ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id), data = ULF_ida_final)
- summary(ulf_model_cc_child)
- # beta values
- standardize_parameters(ulf_model_cc_child)
- ```
- ```{r}
- # putamen
- # main effects: child anaemia
- # log (age)
- # model summary
- ulf_model_putamen_child <- lmer(total_putamen ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_putamen_child)
- # beta values
- standardize_parameters(ulf_model_putamen_child)
- ```
- ```{r}
- # lme: caudate nucleus
- # main effects: child anaemia
- # log (age)
- # model summary
- ulf_model_caudate_child <- lmer(total_caudate ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_caudate_child)
- #to get beta values
- standardize_parameters(ulf_model_caudate_child)
- ```
- **Linear Mixed Effects (LME) Models: Maternal Anaemia and Child Anaemia**
- Second, the final models from the primary analyses assessing the impact of antenatal maternal anaemia status on regional brain volumes were repeated with the addition of postnatal child anaemia status as a covariate. This assessed the relative contributions of antenatal versus postnatal exposures, and examined whether the observed associations between maternal anaemia and regional child brain volumes remained after controlling for child anaemia.
- Log (age) was used for all models based on non-linear growth trajectories
- ```{r}
- # lme: ICV
- # predictors: maternal anaemia and child anaemia
- # based on final maternal anaemia model (primary analysis) for icv: exclude interaction between maternal anaemia and age
- # model summary
- ulf_model_icv_mat_child <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + cor_child_anaemia + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_icv_mat_child)
- # beta values
- standardize_parameters(ulf_model_icv_mat_child)
- ```
- ```{r}
- # lme: corpus callosum
- # predictors: maternal anaemia and child anaemia
- # based on final maternal anaemia model (primary analysis) for corpus callosum : include interaction between maternal anaemia and age
- # model summary
- ulf_model_cc_mat_child <- lmer(total_cc ~ maternal_preg_anaemia_status * log_age + child_sex + cor_child_anaemia + (1 | subject_id), data = ULF_ida_final)
- summary(ulf_model_cc_mat_child)
- # beta values
- standardize_parameters(ulf_model_cc_mat_child)
- ```
- ```{r}
- # lme: putamen
- # predictors: maternal anaemia and child anaemia
- # based on final maternal anaemia model (primary analysis) for putamen: exclude interaction between maternal anaemia and age
- # model summary
- ulf_model_putamen_mat_child <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + cor_child_anaemia + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_putamen_mat_child)
- # beta values
- standardize_parameters(ulf_model_putamen_mat_child)
- ```
- ```{r}
- # lme: caudate nucleus
- # predictors: maternal anaemia and child anaemia
- # based on final maternal anaemia model (primary analysis) for caudate nucleus: include interaction between maternal anaemia and age
- # model summary
- ulf_model_caudate_mat_child <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + cor_child_anaemia + (1 | subject_id),
- data = ULF_ida_final)
- summary(ulf_model_caudate_mat_child)
- # beta values
- standardize_parameters(ulf_model_caudate_mat_child)
- ```
- ------------------------------------------------------------------------
- **Tertiary Analysis: Maternal Iron Deficiency (ULF sample)**
- ------------------------------------------------------------------------
- **Sample Size: For infants with Maternal ID Data**
- ```{r}
- # total sample size
- # including repeated measures
- ULF_ida_final %>%
- filter(!is.na(maternal_preg_adj_brinda_id)) %>% #only include PIDs with antenatal maternal ID data
- summarise(n = n())
- ```
- ```{r}
- # total sample size
- # Excluding repeated measures: unique id
- # unique
Khula_AnaemiaImagingAnalysis.qmd at commit b50e304, no license · at the source
Overview
- Department of Paediatrics and Child Health, Red Cross War Memorial Children’s Hospital, University of Cape Town, Cape Town, 7700, South Africa
- Neuroscience Institute, University of Cape Town, Cape Town, 7925, South Africa
- Centre for Neuroimaging Sciences, Department of Neuroimaging, King's College London, London, England, UK
- Research Department of Early Life Imaging, School of Biomedical Engineering and Imaging Sciences, King’s College London, London, England, UK
- Department for Forensic and Neurodevelopmental Sciences, Institute of Psychiatry, Psychology and Neuroscience, King’s College London, London, England, UK
- Maternal, Newborn, Child Nutrition and Health (MNCH) Discovery and Translational (D&T) Sciences Program, Gates Foundation, Seattle, USA
- Medical Research Council (MRC) Centre for Neurodevelopmental Disorders, King’s College London, London, England, UK
- South African Medical Research Council (SAMRC), Unit on Risk and Resilience in Mental Disorders, Department of Psychiatry, University of Cape Town, Cape Town, 7925, South Africa
- Hawkes Institute and Department of Computer Science, University College London, London, England, UK
- Cardiff University Brain Imaging Research Centre (CUBRIC), School of Psychology, Cardiff University, Cardiff, Wales, UK
Abstract
With the evolution of ultra-low-field MRI and the recognition of antenatal maternal anaemia as an important driver of altered neurodevelopment in toddlers and children, it is critical to determine whether these effects are detectable at ultra-low-field (64 mT) in infancy. The aim of this study was to assess the impact of antenatal maternal anaemia on infant brain structure across the first 2 years of life, using high-field (3 T) and ultra-low-field (64 mT) MRI. This neuroimaging substudy was embedded within Khula, an observational population-based birth cohort in South Africa. Pregnant women were enrolled antenatally and postnatally. Mother-child dyads (n = 394) were followed prospectively with a subsample attending neuroimaging at ∼3, 6, 12, 18 and 24 months of age. Anaemia was classified using World Health Organization thresholds, and neuroimaging data were processed using MiniMORPH. Linear mixed-effects models were used to investigate associations between antenatal maternal anaemia status and absolute regional infant brain volumes using high-field and ultra-low-field MRI. In repeated measures high-field (n = 195) and ultra-low-field (n = 341) infant neuroimaging subsamples, the prevalence of antenatal maternal anaemia was 28.24% (37/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.
UNITY-Physics
Availability: 1 check, the latest on 26 September 2026: the link answers (HTTP 200)
- 26 September 2026: the link answers (HTTP 200)
UNITY-Physics/fw-minimorph
08c7398e42c88031da40ae84941214aa6b4a0299, 7 April 2026Availability: 1 check, the latest on 26 September 2026: the link answers
- 26 September 2026: the link answers
13 files
- app/
main.sh , Shell, 335 lines, 1 match - interactive-run.sh, Shell, 23 lines
- run.py, Python, 54 lines
- utils/
Inspect_segmentations.py , Python, 158 lines - utils/
__init__.py , Python, 2 lines - utils/
command_line.py , Python, 193 lines - utils/
context.py , Python, 874 lines - utils/
generate_command.py , Python, 73 lines - utils/
join_data.py , Python, 37 lines - utils/
parser.py , Python, 69 lines - utils/
tests/ , Shell, 45 linesunit_test.sh - LICENSE, License, 21 lines
- README.md, Text, 173 lines
JERingshaw/Khula-Anaemia-Imaging-Analysis
b50e304453da1c77f7617057f487c5c3334a68b8, 24 March 2026Availability: 1 check, the latest on 26 September 2026: the link answers
- 26 September 2026: the link answers
2 files
- Khula_AnaemiaImagingAnal
ysis.qmd , Quarto, 6,955 lines, 8 matches - README.md, Text, 2 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 12 scripts, each with its path and the digest of its content;
- 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
The de-identified data that support the findings of this study are available from the authors upon reasonable request as per the Khula South Africa study guidelines. All software used in the development of Flywheel gears for the UNITY project has been shared as an open-source resource on GitHub (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 17 authors, 5 keywords, 14 funders, 48 references.
Cite
This paper
Ringshaw, J. E., Zieff, M. R., Bourke, N. J., Casella, C., Bradford, L. E., Williams, S. R., Herr, D., Miles, M., Bennallick, C., Khula South Africa Data Collection Team, Deoni, S., O’Muircheartaigh, J., Stein, D. J., Alexander, D. C., Jones, D. K., Williams, S. C. R., & Donald, K. A. (2026). Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa. Brain communications, 8(5), fcag318. https://
BibTeX
@article{ringshaw2026ant
author = {Ringshaw, Jessica E and Zieff, Michal R and Bourke, Niall J and Casella, Chiara and Bradford, Layla E and Williams, Simone R and Herr, Donna and Miles, Marlie and Bennallick, Carly and {Khula South Africa Data Collection Team} and Deoni, Sean and O’Muircheartaigh, Jonathan and Stein, Dan J and Alexander, Daniel C and Jones, Derek K and Williams, Steven C R and Donald, Kirsten A},
title = {{Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa}},
journal = {Brain communications},
year = {2026},
month = sep,
volume = {8},
number = {5},
pages = {fcag318},
publisher = {Oxford University Press},
issn = {2632-1297},
doi = {10.1093/
url = {https://
pmid = {42713352},
pmcid = {PMC13553114}
}
RIS
TY - JOUR
AU - Ringshaw, Jessica E
AU - Zieff, Michal R
AU - Bourke, Niall J
AU - Casella, Chiara
AU - Bradford, Layla E
AU - Williams, Simone R
AU - Herr, Donna
AU - Miles, Marlie
AU - Bennallick, Carly
AU - Khula South Africa Data Collection Team
AU - Deoni, Sean
AU - O’Muircheartaigh, Jonathan
AU - Stein, Dan J
AU - Alexander, Daniel C
AU - Jones, Derek K
AU - Williams, Steven C R
AU - Donald, Kirsten A
TI - Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa
T2 - Brain communications
J2 - Brain Commun
PY - 2026
DA - 2026/
VL - 8
IS - 5
SP - fcag318
SN - 2632-1297
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa",
"container-title": "Brain communications",
"author": [
{
"family": "Ringshaw",
"given": "Jessica E"
},
{
"family": "Zieff",
"given": "Michal R"
},
{
"family": "Bourke",
"given": "Niall J"
},
{
"family": "Casella",
"given": "Chiara"
},
{
"family": "Bradford",
"given": "Layla E"
},
{
"family": "Williams",
"given": "Simone R"
},
{
"family": "Herr",
"given": "Donna"
},
{
"family": "Miles",
"given": "Marlie"
},
{
"family": "Bennallick",
"given": "Carly"
},
{
"literal": "Khula South Africa Data Collection Team"
},
{
"family": "Deoni",
"given": "Sean"
},
{
"family": "O’Muircheartaigh",
"given": "Jonathan"
},
{
"family": "Stein",
"given": "Dan J"
},
{
"family": "Alexander",
"given": "Daniel C"
},
{
"family": "Jones",
"given": "Derek K"
},
{
"family": "Williams",
"given": "Steven C R"
},
{
"family": "Donald",
"given": "Kirsten A"
}
],
"container-title-short":
"volume": "8",
"issue": "5",
"page": "fcag318",
"DOI": "10.1093/
"PMID": "42713352",
"PMCID": "PMC13553114",
"ISSN": "2632-1297",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
8
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1002/hbm.70547 [code]
- miniMORPH: A Morphometry Pipeline for Low-Field MRI in Infants.Journal: Human brain mappingIn common: ANTs, FreeSurfer, FSL, 5 other tools, developmental, structural MRI / diffusion, 9 references, 3 authors
- [2] doi:10.1038/s41597-026-07350-9 [code]
- An open multi-center MEG-EEG dataset for studying conscious visual perception.Journal: Scientific dataIn common: easystats, car, emmeans, 9 other tools, structural MRI / diffusion
- [3] doi:10.1038/s41597-026-07377-y [code]
- An open-access multi-site fMRI dataset for investigating conscious visual perception.Journal: Scientific dataIn common: easystats, car, emmeans, 9 other tools
- [4] doi:10.1162/imag.a.1347 [code]
- Neural and behavioural correlates of theory of mind reasoning in five-year-old children born preterm.Journal: Imaging neuroscience (Cambridge, Mass.)In common: ANTs, easystats, FreeSurfer, 9 other tools
- [5] doi:10.1073/pnas.2603114123 [code]
- The human hippocampus can pattern separate memories by meaning.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: easystats, emmeans, lmerTest, 8 other tools, 2 references
- [6] doi:10.1093/nc/niag029 [code]
- A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.Journal: Neuroscience of consciousnessIn common: easystats, car, emmeans, 9 other tools
- [7] doi:10.1093/cercor/bhag132 [code]
- Spatiotemporal white-matter development across early childhood.Journal: Cerebral cortex (New York, N.Y. : 1991)In common: ANTs, lmerTest, FSL, 7 other tools, developmental, structural MRI / diffusion, 2 references
- [8] doi:10.1038/s41467-026-73865-9 [code]
- Histamine shapes the neurocomputational dynamics of human learning.Journal: Nature communicationsIn common: easystats, car, emmeans, 8 other tools, 1 reference
- [9] doi:10.1038/s41467-026-71830-0 [code]
- Predicting individual differences of fear and cognitive learning and extinction.Journal: Nature communicationsIn common: ANTs, car, emmeans, 8 other tools
- [10] doi:10.1002/hbm.70512 [code]
- Precision Imaging for Intraindividual Investigation of the Reward Response.Journal: Human brain mappingIn common: ANTs, easystats, lmerTest, 8 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 12 scripts, and 9 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:65679b20deb1491f…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
