OSCR

Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa.

Code ↔ Paper

9 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 9 matches
  1. [1] § Materials and methods › Measures › Neuroimaging outcomes › Neuroimaging processing ↔ app/main.sh, lines 1–40 · score 0.99 · callosal parcellations, ANTs Atropos, supratentorial tissue, Subcortical grey matter, template space, native space
  2. [2] § Materials and methods › Statistical analysis › Sample characteristics ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 5522–5562 · score 0.81 · duplicated demographic, chi squared, avoid skewing, clinical observations, multiple scans acquired, categorical variables
  3. [3] § Materials and methods › Statistical analysis › Neuroimaging regions of interest ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 2253–2296 · score 0.75 · mid anterior, mid posterior, basal ganglia, corpus callosum, brain volume, caudate nucleus
  4. [4] § Results › Primary analysis: antenatal maternal anaemia | HF and ULF MRI › Modelling: antenatal maternal anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 2253–2296 · score 0.72 · linear growth trajectories, logarithmic transformation, linear relationship, absolute age, corpus callosum, LME models
  5. [5] § Materials and methods › Statistical analysis › Secondary analysis: postnatal child anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 6863–6884 · score 0.69 · postnatal exposures, regional child brain, postnatal child anaemia, volumes remained, regional brain volumes, predictor
  6. [6] § Materials and methods › Measures › Anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 5382–5416 · score 0.68 · Minimum child haemoglobin, corresponding age, 3–24, multiple scans, diagnosis, child anaemia status
  7. [7] § Materials and methods › Statistical analysis › Primary analysis: antenatal maternal anaemia status › Statistical modelling ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 338–363 · score 0.66 · absolute regional brain, linear mixed, individual infant, volume trajectories, ULF subsamples, fitted
  8. [8] § Materials and methods › Measures › Contextual measures ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 76–171 · score 0.63 · maternal education, maternal employment, household income, maternal HIV, variables, sex
  9. [9] § Results › Primary analysis: antenatal maternal anaemia | HF and ULF MRI › Exploratory analyses: antenatal maternal anaemia status ↔ Khula_AnaemiaImagingAnalysis.qmd, lines 4870–4903 · score 0.61 · basal ganglia, rapid growth, growth trajectory, corpus callosum, LOESS curves, brain volumes

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Quarto · 6,955 lines · 196 KB · no license · 8 matches

  1. ---
  2. title: "Khula_AnaemiaImagingAnalysis"
  3. format: pdf
  4. editor: visual
  5. ---
  6. ------------------------------------------------------------------------
  7. ------------------------------------------------------------------------
  8. \*This is the final script used in the analyses for the following research: "**Antenatal Maternal Anaemia and Infant Brain Structure: 3 T and 64 mT MRI Findings from South Africa**"
  9. *Details of the datasets:*
  10. - Both high-field (HF; 3 T) and ultra-low-field (ULF; 64 mT) datasets are in long format with some infants having multiple scans across study timepoints (3mo, 6mo, 12mo, 18mo, 24mo).
  11. - These datasets represent the final samples after the exclusion of 1) scans that failed quality check (QC), 2) infants outside of the age range (\>25.5 months or \<2.5 months), and 3) statistical outliers.
  12. - HF: All acquired scans (*n* = 264) –\> passed QC (*n* = 247) –\> within the defined age range (*n* = 239) –\>removal of outliers (*n* = 232)
  13. - ULF: All acquired scans (*n* = 553) –\> passed QC (*n* = 456) —\> within the defined age range (*n* = 439) –\> removal out statistical outliers (*n* = 418)
  14. - The script for the HF analyses is presented first, followed by the script for the ULF analyses.
  15. #### --------------------------------------------------------------------------
  16. **Install Packages**
  17. ```{r}
  18. # install packages if needed
  19. # install.packages('pacman')
  20. library(pacman)
  21. p_load(readr, readxl, tidyverse, magrittr, writexl, stringr, dplyr)
  22. ```
  23. ------------------------------------------------------------------------
  24. [***High-Field (HF) Analyses***]{.underline}
  25. ------------------------------------------------------------------------
  26. **Import HF Dataset**
  27. ```{r}
  28. # import dataset
  29. HF_ida_final <- read_excel("HF_ida_final.xlsx")
  30. ```
  31. **Create New Variables for Total Brain Volumes Across Regions of Interest (ROIs)**
  32. ```{r}
  33. # compute total brain volumes for ROIs
  34. # basal ganglia
  35. HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
  36. HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
  37. # corpus callosum
  38. HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
  39. ```
  40. **Create New Age Variable (Days to Months):**
  41. ```{r}
  42. # convert child age at scan from days to months
  43. # average number of days in a month (accounting for leap years): 30.44
  44. HF_ida_final <- HF_ida_final %>%
  45. mutate(scan_age_months = hf_age / 30.44)
  46. ```
  47. **Creating Meaningful Categorical Variables**
  48. ```{r}
  49. # recode categorical variables from numeric to factor
  50. # maternal anaemia status
  51. HF_ida_final$maternal_preg_anaemia_status <- factor(
  52. HF_ida_final$maternal_preg_anaemia_status,
  53. levels = c(0, 1),
  54. labels = c("no_mat_anaemia", "mat_anaemia")
  55. )
  56. # maternal ID status
  57. HF_ida_final$maternal_preg_adj_brinda_id <- factor(
  58. HF_ida_final$maternal_preg_adj_brinda_id,
  59. levels = c(0, 1),
  60. labels = c("no_mat_id", "mat_id")
  61. )
  62. # maternal ID status
  63. HF_ida_final$maternal_preg_unadjusted_id <- factor(
  64. HF_ida_final$maternal_preg_unadjusted_id,
  65. levels = c(0, 1),
  66. labels = c("no_mat_id", "mat_id")
  67. )
  68. # maternal trimester Hb
  69. HF_ida_final$maternal_trimester_hb <- factor(
  70. HF_ida_final$maternal_trimester_hb,
  71. levels = c(1, 2, 3),
  72. labels = c("first", "second", "third")
  73. )
  74. # child sex
  75. HF_ida_final$child_sex <- factor(
  76. HF_ida_final$child_sex,
  77. levels = c(0, 1),
  78. labels = c("girl", "boy")
  79. )
  80. # maternal pae
  81. HF_ida_final$maternal_pae <- factor(
  82. HF_ida_final$maternal_pae,
  83. levels = c(0, 1),
  84. labels = c("no_pae", "pae")
  85. )
  86. #maternal pte
  87. HF_ida_final$maternal_pte <- factor(
  88. HF_ida_final$maternal_pte,
  89. levels = c(0, 1),
  90. labels = c("no_pte", "pte")
  91. )
  92. # maternal hiv
  93. HF_ida_final$maternal_hiv <- factor(
  94. HF_ida_final$maternal_hiv,
  95. levels = c(0, 1),
  96. labels = c("no_hiv", "hiv")
  97. )
  98. # maternal dep
  99. HF_ida_final$maternal_dep <- factor(
  100. HF_ida_final$maternal_dep,
  101. levels = c(0, 1),
  102. labels = c("no_dep", "dep")
  103. )
  104. # maternal employment
  105. HF_ida_final$mom_employment_en_dichotomous <- factor(
  106. HF_ida_final$mom_employment_en_dichotomous,
  107. levels = c(0, 1),
  108. labels = c("unemployed", "employed")
  109. )
  110. # household income
  111. HF_ida_final$income_household_en <- factor(
  112. HF_ida_final$income_household_en,
  113. levels = c(1, 2, 3, 4),
  114. labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
  115. )
  116. #maternal education
  117. HF_ida_final$mom_edu_en <- factor(
  118. HF_ida_final$mom_edu_en,
  119. levels = c(2, 3, 4, 5, 6),
  120. labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
  121. )
  122. # baby hiv
  123. HF_ida_final$baby_hiv <- factor(
  124. HF_ida_final$baby_hiv,
  125. levels = c(0, 1),
  126. labels = c("no_child_hiv", "child_hiv")
  127. )
  128. ```
  129. **Sample Distribution: Availability of Data Across Study Timepoints**
  130. ```{r}
  131. # sample over 12 months with maternal anaemia data
  132. sample_over_12M <- HF_ida_final %>%
  133. filter(scan_age_months > 12,
  134. !is.na(maternal_preg_anaemia_status))
  135. # sample under 12 months with maternal anaemia data
  136. sample_under_12M <- HF_ida_final %>%
  137. filter(scan_age_months < 12,
  138. !is.na(maternal_preg_anaemia_status))
  139. # sample over 18 months with maternal anaemia data
  140. sample_over_18 <- HF_ida_final %>%
  141. filter(scan_age_months > 18,
  142. !is.na(maternal_preg_anaemia_status))
  143. # sample over 24 months with maternal anaemia data
  144. sample_over_24 <- HF_ida_final %>%
  145. filter(scan_age_months > 24,
  146. !is.na(maternal_preg_anaemia_status))
  147. ```
  148. ```{r}
  149. # HF sample size
  150. # unique observations in HF dataset
  151. # excluding repeated measures
  152. library(dplyr)
  153. HF_ida_final %>%
  154. summarise(n_unique_subjects = n_distinct(subject_id))
  155. ```
  156. ------------------------------------------------------------------------
  157. **Primary Analysis: Antenatal Maternal Anaemia Status (HF Sample)**
  158. ------------------------------------------------------------------------
  159. **Sample Size: For infants with Antenatal Maternal Anemia Data**
  160. Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
  161. - The sample of unique participants was used to report on sample characteristics
  162. - The full sample with repeated measures was used for analyses including linear mixed effects (LME) models to include children with multiple scans from different study timepoints.
  163. ```{r}
  164. # total sample size
  165. # repeated measures included
  166. # sample with antenatal maternal anaemia data
  167. HF_ida_final %>%
  168. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  169. summarise(n = n())
  170. ```
  171. ```{r}
  172. # unique sample size
  173. # repeated measures excluded
  174. # number of unique participants in sample
  175. library(dplyr)
  176. HF_ida_final %>%
  177. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include subjects with anaemia data
  178. distinct(subject_id, .keep_all = TRUE) %>% # keep unique subject IDs
  179. summarise(n_total = n()) %>% # count unique subjects
  180. print()
  181. ```
  182. ```{r}
  183. # prevalence of antenatal maternal anaemia
  184. # excluding repeated measures
  185. library(dplyr)
  186. HF_ida_final %>%
  187. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  188. distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per mother
  189. group_by(maternal_preg_anaemia_status) %>%
  190. summarise(n = n()) %>%
  191. mutate(percent = round(100 * n / sum(n), 2))
  192. ```
  193. ```{r}
  194. # antenatal maternal anaemia severity
  195. # unique subject IDs: Based on true prevalence
  196. # excluding repeated measures
  197. library(dplyr)
  198. HF_ida_final %>%
  199. filter(!is.na(maternal_preg_anaemia_severity)) %>% # exclude missing severity
  200. group_by(subject_id) %>% # group by mother
  201. summarise(
  202. maternal_preg_anaemia_severity = first(
  203. maternal_preg_anaemia_severity[maternal_preg_anaemia_severity != ""]
  204. )
  205. ) %>%
  206. group_by(maternal_preg_anaemia_severity) %>% # count unique mothers by severity
  207. summarise(n = n()) %>%
  208. mutate(percent = round(100 * n / sum(n), 2))
  209. ```
  210. **Repeated Measures Data Availability**
  211. ```{r}
  212. # repeated measures count
  213. # unique subject IDs
  214. library(dplyr)
  215. HF_ida_final %>%
  216. filter(!is.na(maternal_preg_anaemia_status)) %>% # keep only rows with anaemia status
  217. count(subject_id) %>% # count rows per subject
  218. count(n) %>% # count how many subjects had n rows
  219. rename(times_measured = n, number_of_subjects = nn) %>%
  220. mutate(percent = round(100 * number_of_subjects / sum(number_of_subjects), 1)) # add % column
  221. ```
  222. ```{r}
  223. # HF Figure 2
  224. # plot of data availability with repeated measures
  225. library(ggplot2)
  226. library(dplyr)
  227. Figure_2_HF<- HF_ida_final %>%
  228. filter(!is.na(maternal_preg_anaemia_status)) %>%
  229. mutate(
  230. timepoint = factor(timepoint, levels = c("3M", "6M", "12M", "18M", "24M")),
  231. subject_id = factor(subject_id),
  232. anaemia_group = as.factor(maternal_preg_anaemia_status)
  233. ) %>%
  234. ggplot(aes(x = timepoint, y = subject_id, group = subject_id, color = anaemia_group)) +
  235. geom_line(alpha = 0.5, linewidth = 0.5) + # updated for ggplot2 ≥3.4.0
  236. geom_point(size = 2) +
  237. scale_color_manual(
  238. values = c("red", "blue"),
  239. labels = c("Maternal Anaemia", "No Maternal Anaemia")
  240. ) +
  241. scale_x_discrete(expand = expansion(mult = 0.05)) +
  242. scale_y_discrete(expand = expansion(mult = 0.02)) +
  243. coord_cartesian(clip = "off") + # ensures no points are clipped at the edges
  244. labs(
  245. title = "Participant Data Availability by Timepoint",
  246. x = "Timepoint",
  247. y = "Participants",
  248. color = "Maternal Anaemia"
  249. ) +
  250. theme_minimal(base_size = 13) +
  251. theme(
  252. axis.text.y = element_blank(),
  253. axis.ticks.y = element_blank(),
  254. panel.grid.major.y = element_blank()
  255. )
  256. print(Figure_2_HF)
  257. ```
  258. **Sample Characteristics and Group Differences: Antenatal Maternal Anaemia Status**
  259. Sample characteristics were assessed based on the number of unique subjects in each subsample, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
  260. However, given the repeated measures study design with individual infant data across multiple study visits in both HF and ULF subsamples, linear mixed effects (LME) models were fitted to account for individual differences in assessing the association between antenatal maternal anaemia status and absolute regional brain volume trajectories
  261. - **Categorical Variables**
  262. *Trimester of Maternal Hb Measurement*
  263. ```{r}
  264. # trimester Hb was collected from mothers
  265. # excluding repeated measures
  266. # unique subject IDs: Based on true prevalence
  267. library(dplyr)
  268. HF_ida_final %>%
  269. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia
  270. group_by(subject_id) %>% # keep one row per mother
  271. summarise(maternal_trimester_hb = first(maternal_trimester_hb)) %>%
  272. group_by(maternal_trimester_hb) %>%
  273. summarise(n = n()) %>%
  274. mutate(percent = round(100 * n / sum(n), 2))
  275. ```
  276. ```{r}
  277. # group differences in trimester of Hb collection
  278. # chi-squared: trimester Hb
  279. # unique subject IDs: based on true prevalence
  280. library(dplyr)
  281. # Filter unique subjects with non-missing anaemia status
  282. unique_data <- HF_ida_final %>%
  283. filter(!is.na(maternal_preg_anaemia_status)) %>%
  284. distinct(subject_id, .keep_all = TRUE)
  285. # Create contingency table: Trimester by Maternal Anaemia Status
  286. tbl <- table(unique_data$maternal_trimester_hb, unique_data$maternal_preg_anaemia_status)
  287. # Print contingency table
  288. cat("Contingency Table: Trimester by Maternal Anaemia Status\n")
  289. print(format(tbl, nsmall = 0))
  290. # Print column percentages
  291. cat("\nColumn Percentages (% of each anaemia status group):\n")
  292. print(round(100 * prop.table(tbl, margin = 2), 2))
  293. # Run appropriate test based on expected cell counts
  294. expected <- chisq.test(tbl)$expected
  295. if (any(expected < 5)) {
  296. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  297. print(fisher.test(tbl))
  298. } else {
  299. cat("\nChi-squared Test:\n")
  300. print(chisq.test(tbl))
  301. }
  302. ```
  303. *Child Sex at Birth*
  304. ```{r}
  305. # child sex at birth (full sample)
  306. # unique subject IDs: based on true prevalence
  307. library(dplyr)
  308. HF_ida_final %>%
  309. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  310. distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per mother/child
  311. group_by(child_sex) %>%
  312. summarise(n = n()) %>%
  313. mutate(percent = round(100 * n / sum(n), 2))
  314. ```
  315. ```{r}
  316. # group differences in child sex by antenatal maternal anaemia status
  317. # chi-squared: child sex
  318. # unique subject IDs: based on true prevalence
  319. library(dplyr)
  320. # keep only unique subjects and non-missing anaemia status
  321. unique_data <- HF_ida_final %>%
  322. filter(!is.na(maternal_preg_anaemia_status)) %>%
  323. distinct(subject_id, .keep_all = TRUE)
  324. # create contingency table
  325. tbl <- table(
  326. unique_data$child_sex,
  327. unique_data$maternal_preg_anaemia_status
  328. )
  329. # print contingency table
  330. cat("Contingency Table: Child Sex by Maternal Anaemia Status\n")
  331. print(format(round(tbl, 2), nsmall = 2))
  332. cat("\nColumn Percentages (% of each anaemia status group):\n")
  333. print(round(100 * prop.table(tbl, margin = 2), 2))
  334. # run appropriate test
  335. expected <- chisq.test(tbl)$expected
  336. if (any(expected < 5)) {
  337. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  338. print(fisher.test(tbl))
  339. } else {
  340. cat("\nChi-squared Test:\n")
  341. print(chisq.test(tbl))
  342. }
  343. ```
  344. *Prenatal Alcohol Exposure (PAE)*
  345. ```{r}
  346. # pae prevalence (full sample)
  347. # unique subject IDs: based on true prevalence
  348. HF_ida_final %>%
  349. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include subjects with anaemia data
  350. distinct(subject_id, .keep_all = TRUE) %>% # keep only unique subjects
  351. group_by(maternal_pae) %>%
  352. summarise(n = n()) %>%
  353. mutate(percent = round(100 * n / sum(n), 2))
  354. ```
  355. ```{r}
  356. # group differences in pae by antenatal maternal anaemia status
  357. # chi-squared: pae
  358. # unique subject IDs: based on true prevalence
  359. # filter unique subjects and exclude missing anaemia status
  360. unique_data <- HF_ida_final %>%
  361. filter(!is.na(maternal_preg_anaemia_status)) %>%
  362. distinct(subject_id, .keep_all = TRUE)
  363. # create contingency table
  364. tbl <- table(
  365. unique_data$maternal_pae,
  366. unique_data$maternal_preg_anaemia_status
  367. )
  368. # print contingency table
  369. cat("Contingency Table: Maternal PAE by Maternal Anaemia Status\n")
  370. print(tbl)
  371. cat("\nColumn Percentages (% of each anaemia status group):\n")
  372. print(round(100 * prop.table(tbl, margin = 2), 2))
  373. # run appropriate test
  374. expected <- chisq.test(tbl)$expected
  375. if (any(expected < 5)) {
  376. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  377. print(fisher.test(tbl))
  378. } else {
  379. cat("\nChi-squared Test:\n")
  380. print(chisq.test(tbl))
  381. }
  382. ```
  383. *Prenatal Tobacco Exposure (PTE)*
  384. ```{r}
  385. # pte prevalence (full sample)
  386. # unique subject IDs: based on true prevalence
  387. unique_data <- HF_ida_final %>%
  388. filter(!is.na(maternal_preg_anaemia_status)) %>%
  389. distinct(subject_id, .keep_all = TRUE)
  390. # summarise maternal pte
  391. unique_data %>%
  392. group_by(maternal_pte) %>%
  393. summarise(n = n()) %>%
  394. mutate(percent = round(100 * n / sum(n), 2))
  395. ```
  396. ```{r}
  397. # group differences in pte by antenatal maternal anaemia status
  398. # chi-squared: pte
  399. # unique subject IDs: based on true prevalence
  400. # filter unique subjects and exclude missing anaemia status
  401. unique_data <- HF_ida_final %>%
  402. filter(!is.na(maternal_preg_anaemia_status)) %>%
  403. distinct(subject_id, .keep_all = TRUE)
  404. # create contingency table
  405. tbl <- table(
  406. unique_data$maternal_pte,
  407. unique_data$maternal_preg_anaemia_status
  408. )
  409. # print contingency table
  410. cat("Contingency Table: Maternal PTE by Maternal Anaemia Status (Unique Subjects)\n")
  411. print(tbl)
  412. cat("\nColumn Percentages (% of each anaemia status group):\n")
  413. print(round(100 * prop.table(tbl, margin = 2), 2))
  414. # run appropriate test based on expected counts
  415. expected <- chisq.test(tbl)$expected
  416. if (any(expected < 5)) {
  417. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  418. print(fisher.test(tbl))
  419. } else {
  420. cat("\nChi-squared Test:\n")
  421. print(chisq.test(tbl))
  422. }
  423. ```
  424. *Maternal HIV Infection*
  425. ```{r}
  426. # materal hiv prevalence (full sample)
  427. # unique subject IDs: based on true prevalence
  428. HF_ida_final %>%
  429. filter(!is.na(maternal_preg_anaemia_status)) %>%
  430. distinct(subject_id, .keep_all = TRUE) %>%
  431. group_by(maternal_hiv) %>%
  432. summarise(n = n()) %>%
  433. mutate(percent = round(100 * n / sum(n), 2))
  434. ```
  435. ```{r}
  436. # group differences in maternal hiv by antenatal maternal anaemia status
  437. # chi-squared: hiv
  438. # unique subject IDs: based on true prevalence
  439. unique_data <- HF_ida_final %>%
  440. filter(!is.na(maternal_preg_anaemia_status)) %>%
  441. distinct(subject_id, .keep_all = TRUE)
  442. # create contingency table
  443. tbl <- table(
  444. unique_data$maternal_hiv,
  445. unique_data$maternal_preg_anaemia_status
  446. )
  447. # print contingency table
  448. cat("Contingency Table: Maternal HIV by Maternal Anaemia Status\n")
  449. print(tbl)
  450. cat("\nColumn Percentages (% of each anaemia status group):\n")
  451. print(round(100 * prop.table(tbl, margin = 2), 2))
  452. # run appropriate test
  453. expected <- chisq.test(tbl)$expected
  454. if (any(expected < 5)) {
  455. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  456. print(fisher.test(tbl))
  457. } else {
  458. cat("\nChi-squared Test:\n")
  459. print(chisq.test(tbl))
  460. }
  461. ```
  462. *Maternal Depression*
  463. ```{r}
  464. # maternal depression prevalence (full sample)
  465. # unique subject IDs: based on true prevalence
  466. HF_ida_final %>%
  467. filter(!is.na(maternal_dep)) %>%
  468. distinct(subject_id, .keep_all = TRUE) %>%
  469. group_by(maternal_dep) %>%
  470. summarise(n = n()) %>%
  471. mutate(percent = round(100 * n / sum(n), 2))
  472. ```
  473. ```{r}
  474. # group differenced by maternal depression status
  475. # chi-squared: maternal depression
  476. # unique subject IDs: based on true prevalence
  477. unique_data <- HF_ida_final %>%
  478. filter(!is.na(maternal_preg_anaemia_status)) %>%
  479. distinct(subject_id, .keep_all = TRUE)
  480. # create contingency table
  481. tbl <- table(
  482. unique_data$maternal_dep,
  483. unique_data$maternal_preg_anaemia_status
  484. )
  485. # print contingency table
  486. cat("Contingency Table: Maternal DEP by Maternal Anaemia Status\n")
  487. print(tbl)
  488. cat("\nColumn Percentages (% of each anaemia status group):\n")
  489. print(round(100 * prop.table(tbl, margin = 2), 2))
  490. # run appropriate test
  491. expected <- chisq.test(tbl)$expected
  492. if (any(expected < 5)) {
  493. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  494. print(fisher.test(tbl))
  495. } else {
  496. cat("\nChi-squared Test:\n")
  497. print(chisq.test(tbl))
  498. }
  499. ```
  500. *Maternal Employment*
  501. ```{r}
  502. # maternal employment (full sample)
  503. # unique subject IDs: based on true prevalence
  504. HF_ida_final %>%
  505. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  506. distinct(subject_id, .keep_all = TRUE) %>% # keep unique subjects
  507. group_by(mom_employment_en_dichotomous) %>%
  508. summarise(n = n()) %>%
  509. mutate(percent = round(100 * n / sum(n), 2))
  510. ```
  511. ```{r}
  512. # group differences in maternal employment by antenatal maternal anaemia status
  513. # chi-squared: maternal employment
  514. # unique subject IDs: based on true prevalence
  515. unique_data <- HF_ida_final %>%
  516. filter(!is.na(maternal_preg_anaemia_status)) %>%
  517. distinct(subject_id, .keep_all = TRUE)
  518. # create contingency table
  519. tbl <- table(
  520. unique_data$mom_employment_en_dichotomous,
  521. unique_data$maternal_preg_anaemia_status
  522. )
  523. # print contingency table
  524. cat("Contingency Table: Maternal Employment by Maternal Anaemia Status\n")
  525. print(tbl)
  526. cat("\nColumn Percentages (% of each anaemia status group):\n")
  527. print(round(100 * prop.table(tbl, margin = 2), 2))
  528. # run appropriate test
  529. expected <- chisq.test(tbl)$expected
  530. if (any(expected < 5)) {
  531. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  532. print(fisher.test(tbl))
  533. } else {
  534. cat("\nChi-squared Test:\n")
  535. print(chisq.test(tbl))
  536. }
  537. ```
  538. *Maternal Education*
  539. ```{r}
  540. # maternal education (full sample)
  541. # unique subject IDs: based on true prevalence
  542. unique_data <- HF_ida_final %>%
  543. filter(!is.na(maternal_preg_anaemia_status)) %>%
  544. distinct(subject_id, .keep_all = TRUE)
  545. # summarise maternal education
  546. unique_data %>%
  547. group_by(mom_edu_en) %>%
  548. summarise(n = n()) %>%
  549. mutate(percent = round(100 * n / sum(n), 2))
  550. ```
  551. ```{r}
  552. # group differences in maternal education by antenatal maternal anaemia status
  553. # chi-squared: maternal education
  554. # unique subject IDs: based on true prevalence
  555. unique_data <- HF_ida_final %>%
  556. filter(!is.na(maternal_preg_anaemia_status)) %>%
  557. distinct(subject_id, .keep_all = TRUE)
  558. # create contingency table
  559. tbl <- table(
  560. unique_data$mom_edu_en,
  561. unique_data$maternal_preg_anaemia_status
  562. )
  563. # print contingency table
  564. cat("Contingency Table: Maternal Education by Maternal Anaemia Status\n")
  565. print(tbl)
  566. cat("\nColumn Percentages (% of each anaemia status group):\n")
  567. print(round(100 * prop.table(tbl, margin = 2), 2))
  568. # run appropriate test
  569. expected <- chisq.test(tbl)$expected
  570. if (any(expected < 5)) {
  571. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  572. print(fisher.test(tbl))
  573. } else {
  574. cat("\nChi-squared Test:\n")
  575. print(chisq.test(tbl))
  576. }
  577. ```
  578. *Household Income*
  579. ```{r}
  580. # household income (full sample)
  581. # unique subject IDs: based on true prevalence
  582. unique_data <- HF_ida_final %>%
  583. filter(!is.na(maternal_preg_anaemia_status)) %>%
  584. distinct(subject_id, .keep_all = TRUE)
  585. # summarise household income
  586. unique_data %>%
  587. group_by(income_household_en) %>%
  588. summarise(n = n()) %>%
  589. mutate(percent = round(100 * n / sum(n), 2))
  590. ```
  591. ```{r}
  592. # group differences in household income by antenatal maternal anaemia status
  593. # chi-squared: household income
  594. # unique subject IDs: based on true prevalence
  595. unique_data <- HF_ida_final %>%
  596. filter(!is.na(maternal_preg_anaemia_status)) %>%
  597. distinct(subject_id, .keep_all = TRUE)
  598. # create contingency table
  599. tbl <- table(
  600. unique_data$income_household_en,
  601. unique_data$maternal_preg_anaemia_status
  602. )
  603. # print contingency table
  604. cat("Contingency Table: Household Income by Maternal Anaemia Status\n")
  605. print(tbl)
  606. cat("\nColumn Percentages (% of each anaemia status group):\n")
  607. print(round(100 * prop.table(tbl, margin = 2), 2))
  608. # run appropriate test
  609. expected <- chisq.test(tbl)$expected
  610. if (any(expected < 5)) {
  611. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  612. print(fisher.test(tbl))
  613. } else {
  614. cat("\nChi-squared Test:\n")
  615. print(chisq.test(tbl))
  616. }
  617. ```
  618. *Infant HIV Infection*
  619. ```{r}
  620. # infant hiv prevalence (full sample)
  621. # unique subject IDs: based on true prevalence
  622. unique_data <- HF_ida_final %>%
  623. filter(!is.na(maternal_preg_anaemia_status)) %>%
  624. distinct(subject_id, .keep_all = TRUE)
  625. # summarise counts and percentages
  626. unique_data %>%
  627. group_by(baby_hiv) %>%
  628. summarise(n = n()) %>%
  629. mutate(percent = round(100 * n / sum(n), 2))
  630. ```
  631. ```{r}
  632. # group differences in infant hiv by antenatal maternal anaemia status
  633. # chi-squared infant hiv
  634. # unique subject IDs: based on true prevalence
  635. unique_data <- HF_ida_final %>%
  636. filter(!is.na(maternal_preg_anaemia_status)) %>%
  637. distinct(subject_id, .keep_all = TRUE)
  638. # create contingency table
  639. tbl <- table(
  640. unique_data$baby_hiv,
  641. unique_data$maternal_preg_anaemia_status
  642. )
  643. # print contingency table
  644. cat("Contingency Table: Baby HIV by Maternal Anaemia Status\n")
  645. print(tbl)
  646. cat("\nColumn Percentages (% of each anaemia status group):\n")
  647. print(round(100 * prop.table(tbl, margin = 2), 2))
  648. # run appropriate test
  649. expected <- chisq.test(tbl)$expected
  650. if (any(expected < 5)) {
  651. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  652. print(fisher.test(tbl))
  653. } else {
  654. cat("\nChi-squared Test:\n")
  655. print(chisq.test(tbl))
  656. }
  657. ```
  658. - **Continuous Variables**
  659. ```{r}
  660. #install packages if needed
  661. #install.packages('car')
  662. library("car")
  663. ```
  664. *Maternal Age at Enrolment*
  665. ```{r}
  666. # maternal age at enrolment (full sample)
  667. # unique subject IDs: based on true prevalence
  668. library(dplyr)
  669. HF_ida_final %>%
  670. filter(!is.na(maternal_preg_anaemia_status),
  671. !is.na(mom_age_en)) %>% # exclude missing anaemia or age
  672. group_by(subject_id) %>% # keep unique mothers only
  673. summarise(
  674. mom_age_en = first(mom_age_en) # assume age is constant per mother
  675. ) %>%
  676. summarise(
  677. mean_mom_age = mean(mom_age_en, na.rm = TRUE),
  678. sd_mom_age = sd(mom_age_en, na.rm = TRUE),
  679. min_mom_age = min(mom_age_en, na.rm = TRUE),
  680. max_mom_age = max(mom_age_en, na.rm = TRUE),
  681. n = n()
  682. )
  683. ```
  684. ```{r}
  685. # group differences in maternal age at enrolment by antenatal maternal anaemia status
  686. # ttest: maternal age at enrolment
  687. # unique subject IDs: based on true prevalence
  688. library("dplyr")
  689. library("car")
  690. # keep only unique subjects with non-missing anaemia status
  691. unique_data <- HF_ida_final %>%
  692. filter(!is.na(maternal_preg_anaemia_status)) %>%
  693. distinct(subject_id, .keep_all = TRUE)
  694. # summary stats with exactly 2 decimals
  695. summary_stats <- unique_data %>%
  696. group_by(maternal_preg_anaemia_status) %>%
  697. summarise(
  698. n = n(),
  699. mean_age = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  700. sd_age = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  701. min_age = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  702. max_age = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  703. .groups = "drop"
  704. )
  705. # print summary stats table
  706. print(summary_stats)
  707. # levene's Test for equality of variances
  708. levene_result <- leveneTest(
  709. mom_age_en ~ maternal_preg_anaemia_status,
  710. data = unique_data
  711. )
  712. print(levene_result)
  713. # extract p-value from Levene’s Test
  714. levene_p <- levene_result$`Pr(>F)`[1]
  715. # automatically select Student's or Welch's t-test
  716. if (levene_p > 0.05) {
  717. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  718. ") → Using Student's t-test (equal variances assumed)\n\n")
  719. t_result <- t.test(
  720. mom_age_en ~ maternal_preg_anaemia_status,
  721. data = unique_data,
  722. var.equal = TRUE
  723. )
  724. } else {
  725. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  726. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  727. t_result <- t.test(
  728. mom_age_en ~ maternal_preg_anaemia_status,
  729. data = unique_data,
  730. var.equal = FALSE
  731. )
  732. }
  733. # round t-test output (except p-value)
  734. t_result$estimate <- round(t_result$estimate, 2)
  735. t_result$statistic <- round(t_result$statistic, 2)
  736. t_result$conf.int <- round(t_result$conf.int, 2)
  737. # keep p-value as is, or round to 4 decimals for clean display
  738. t_result$p.value <- round(t_result$p.value, 4)
  739. # print t-test result
  740. print(t_result)
  741. ```
  742. *Child Age at Scan*
  743. Include repeated measures to account for different ages for scans acquired across study timepoints from the same infant
  744. ```{r}
  745. # child age at scan for full group
  746. # include repeated measures: different ages at every scan acquired (full sample)
  747. HF_ida_final %>%
  748. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  749. summarise(
  750. mean_child_age = mean(scan_age_months, na.rm = TRUE),
  751. sd_child_age = sd(scan_age_months, na.rm = TRUE),
  752. min_child_age = min(scan_age_months, na.rm = TRUE),
  753. max_child_age = max(scan_age_months, na.rm = TRUE),
  754. n = n()
  755. )
  756. ```
  757. ```{r}
  758. # child age at scan
  759. # full sample including repeated measures
  760. # t-test: child age at scan
  761. library(dplyr)
  762. library(car)
  763. # filter once and store
  764. analysis_data <- HF_ida_final %>%
  765. filter(!is.na(maternal_preg_anaemia_status))
  766. # summary statistics with exactly 2 decimals displayed
  767. summary_stats <- analysis_data %>%
  768. group_by(maternal_preg_anaemia_status) %>%
  769. summarise(
  770. n = n(),
  771. mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  772. sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  773. min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  774. max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  775. .groups = "drop"
  776. )
  777. # print summary stats
  778. print(summary_stats)
  779. # levene's Test for equality of variances
  780. levene_result <- leveneTest(
  781. scan_age_months ~ maternal_preg_anaemia_status,
  782. data = analysis_data
  783. )
  784. print(levene_result)
  785. # extract Levene p-value (rounded just for reporting)
  786. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  787. cat("\nLevene’s test p-value =", levene_p, "\n")
  788. # automatically choose correct t-test based on Levene
  789. if (levene_p > 0.05) {
  790. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  791. t_result <- t.test(
  792. scan_age_months ~ maternal_preg_anaemia_status,
  793. data = analysis_data,
  794. var.equal = TRUE
  795. )
  796. } else {
  797. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  798. t_result <- t.test(
  799. scan_age_months ~ maternal_preg_anaemia_status,
  800. data = analysis_data,
  801. var.equal = FALSE
  802. )
  803. }
  804. # print t-test result in full precision
  805. print(t_result)
  806. ```
  807. *Child Gestational Age (GA) at Birth*
  808. ```{r}
  809. # child ga at birth for full group
  810. # unique subject IDs: based on true prevalence
  811. library(dplyr)
  812. # filter for unique subjects with non-missing maternal anaemia status
  813. unique_data <- HF_ida_final %>%
  814. filter(!is.na(maternal_preg_anaemia_status)) %>%
  815. distinct(subject_id, .keep_all = TRUE)
  816. # summary statistics for the full sample
  817. unique_data %>%
  818. summarise(
  819. mean_child_ga = mean(ga_weeks, na.rm = TRUE),
  820. sd_child_ga = sd(ga_weeks, na.rm = TRUE),
  821. min_child_ga = min(ga_weeks, na.rm = TRUE),
  822. max_child_ga = max(ga_weeks, na.rm = TRUE),
  823. n = n()
  824. )
  825. ```
  826. ```{r}
  827. # group differences in child ga at birth
  828. # ttest: child ga
  829. # unique subject IDs: based on true prevalence
  830. library(dplyr)
  831. library(car)
  832. # keep only unique subjects with non-missing anaemia status
  833. unique_data <- HF_ida_final %>%
  834. filter(!is.na(maternal_preg_anaemia_status)) %>%
  835. distinct(subject_id, .keep_all = TRUE)
  836. # summary statistics with exactly 2 decimals displayed
  837. summary_stats <- unique_data %>%
  838. group_by(maternal_preg_anaemia_status) %>%
  839. summarise(
  840. n = n(),
  841. mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  842. sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  843. min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  844. max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  845. .groups = "drop"
  846. )
  847. # print summary stats
  848. print(summary_stats)
  849. # levene's test
  850. levene_result <- leveneTest(
  851. ga_weeks ~ maternal_preg_anaemia_status,
  852. data = unique_data
  853. )
  854. # round levene's p-value just for display
  855. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  856. cat("\nLevene’s test p-value =", levene_p, "\n")
  857. # automatically choose correct t-test
  858. if (levene_p > 0.05) {
  859. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  860. t_result <- t.test(
  861. ga_weeks ~ maternal_preg_anaemia_status,
  862. data = unique_data,
  863. var.equal = TRUE
  864. )
  865. } else {
  866. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  867. t_result <- t.test(
  868. ga_weeks ~ maternal_preg_anaemia_status,
  869. data = unique_data,
  870. var.equal = FALSE
  871. )
  872. }
  873. # print t-test result in full precision
  874. print(t_result)
  875. ```
  876. *Child Birth Weight (BW)*
  877. ```{r}
  878. # child bw (full sample)
  879. # unique subject IDs: based on true prevalence
  880. library(dplyr)
  881. # filter for unique subjects with non-missing maternal anaemia status
  882. unique_data <- HF_ida_final %>%
  883. filter(!is.na(maternal_preg_anaemia_status)) %>%
  884. distinct(subject_id, .keep_all = TRUE)
  885. # summary statistics for the full sample
  886. unique_data %>%
  887. summarise(
  888. mean_bw = mean(bw, na.rm = TRUE),
  889. sd_bw = sd(bw, na.rm = TRUE),
  890. min_bw = min(bw, na.rm = TRUE),
  891. max_bw = max(bw, na.rm = TRUE),
  892. n = n()
  893. )
  894. ```
  895. ```{r}
  896. # group differences in child bw
  897. # ttest: child bw
  898. # unique subject IDs: based on true prevalence
  899. library(dplyr)
  900. library(car)
  901. # filter for unique subjects with non-missing maternal anaemia status
  902. unique_data <- HF_ida_final %>%
  903. filter(!is.na(maternal_preg_anaemia_status)) %>%
  904. distinct(subject_id, .keep_all = TRUE)
  905. # summary statistics by maternal anaemia status with exactly 2 decimals displayed
  906. summary_stats <- unique_data %>%
  907. group_by(maternal_preg_anaemia_status) %>%
  908. summarise(
  909. n = n(),
  910. mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
  911. sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
  912. min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
  913. max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
  914. .groups = "drop"
  915. )
  916. # print summary stats
  917. print(summary_stats)
  918. # levene's Test for homogeneity of variance
  919. levene_result <- leveneTest(bw ~ maternal_preg_anaemia_status, data = unique_data)
  920. print(levene_result)
  921. # extract Levene p-value (rounded for reporting)
  922. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  923. cat("\nLevene’s test p-value =", levene_p, "\n")
  924. # t-test comparing birthweight by anaemia status
  925. # Automatically use equal variances if Levene p > 0.05
  926. if (levene_p > 0.05) {
  927. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  928. t_result <- t.test(
  929. bw ~ maternal_preg_anaemia_status,
  930. data = unique_data,
  931. var.equal = TRUE
  932. )
  933. } else {
  934. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  935. t_result <- t.test(
  936. bw ~ maternal_preg_anaemia_status,
  937. data = unique_data,
  938. var.equal = FALSE
  939. )
  940. }
  941. # Print t-test result in full precision
  942. print(t_result)
  943. ```
  944. *Child Birth Length (BL)*
  945. ```{r}
  946. # child bl (full sample)
  947. # unique subject IDs: based on true prevalence
  948. library(dplyr)
  949. # filter for unique subjects with non-missing maternal anaemia status
  950. unique_data <- HF_ida_final %>%
  951. filter(!is.na(maternal_preg_anaemia_status)) %>%
  952. distinct(subject_id, .keep_all = TRUE)
  953. # summary statistics for birth length
  954. unique_data %>%
  955. summarise(
  956. mean_birthlength = mean(birth_length, na.rm = TRUE),
  957. sd_birthlength = sd(birth_length, na.rm = TRUE),
  958. min_birthlength = min(birth_length, na.rm = TRUE),
  959. max_birthlength = max(birth_length, na.rm = TRUE),
  960. n = n()
  961. )
  962. ```
  963. ```{r}
  964. # group differences in child birthlength
  965. # ttest: child bl
  966. # exclude repeated measures: unique id
  967. library(dplyr)
  968. library(car)
  969. # keep only unique subjects with non-missing anaemia status
  970. unique_data <- HF_ida_final %>%
  971. filter(!is.na(maternal_preg_anaemia_status)) %>%
  972. distinct(subject_id, .keep_all = TRUE)
  973. # summary statistics with exactly 2 decimals displayed
  974. summary_stats <- unique_data %>%
  975. group_by(maternal_preg_anaemia_status) %>%
  976. summarise(
  977. n = n(),
  978. mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
  979. sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
  980. min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
  981. max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
  982. .groups = "drop"
  983. )
  984. # print summary stats
  985. print(summary_stats)
  986. # levene's Test
  987. levene_result <- leveneTest(
  988. birth_length ~ maternal_preg_anaemia_status,
  989. data = unique_data
  990. )
  991. print(levene_result)
  992. # extract Levene p-value (rounded for reporting only)
  993. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  994. cat("\nLevene’s test p-value =", levene_p, "\n")
  995. # automatically choose correct t-test
  996. if (levene_p > 0.05) {
  997. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  998. t_result <- t.test(
  999. birth_length ~ maternal_preg_anaemia_status,
  1000. data = unique_data,
  1001. var.equal = TRUE
  1002. )
  1003. } else {
  1004. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1005. t_result <- t.test(
  1006. birth_length ~ maternal_preg_anaemia_status,
  1007. data = unique_data,
  1008. var.equal = FALSE
  1009. )
  1010. }
  1011. # print t-test result in full precision
  1012. print(t_result)
  1013. ```
  1014. *Birth Head Circumference (HC)*
  1015. ```{r}
  1016. # child birth hc (full sample)
  1017. # unique subject IDs: based on true prevalence
  1018. library(dplyr)
  1019. # filter for unique subjects with non-missing maternal anaemia status
  1020. unique_data <- HF_ida_final %>%
  1021. filter(!is.na(maternal_preg_anaemia_status)) %>%
  1022. distinct(subject_id, .keep_all = TRUE)
  1023. # summary statistics for birth head circumference
  1024. unique_data %>%
  1025. summarise(
  1026. mean_hc = mean(birth_hc, na.rm = TRUE),
  1027. sd_hc = sd(birth_hc, na.rm = TRUE),
  1028. min_hc = min(birth_hc, na.rm = TRUE),
  1029. max_hc = max(birth_hc, na.rm = TRUE),
  1030. n = n()
  1031. )
  1032. ```
  1033. ```{r}
  1034. # group differences in child birth hc
  1035. # ttest: child hc
  1036. # unique subject IDs: based on true prevalence
  1037. library(dplyr)
  1038. library(car)
  1039. # filter for unique subjects with non-missing maternal anaemia status
  1040. unique_data <- HF_ida_final %>%
  1041. filter(!is.na(maternal_preg_anaemia_status)) %>%
  1042. distinct(subject_id, .keep_all = TRUE)
  1043. # summary statistics with exactly 2 decimals displayed
  1044. summary_stats <- unique_data %>%
  1045. group_by(maternal_preg_anaemia_status) %>%
  1046. summarise(
  1047. n = n(),
  1048. mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  1049. sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  1050. min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  1051. max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  1052. .groups = "drop"
  1053. )
  1054. # print summary stats
  1055. print(summary_stats)
  1056. # levene's Test
  1057. levene_result <- leveneTest(
  1058. birth_hc ~ maternal_preg_anaemia_status,
  1059. data = unique_data
  1060. )
  1061. print(levene_result)
  1062. # extract Levene p-value (rounded for reporting only)
  1063. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  1064. cat("\nLevene’s test p-value =", levene_p, "\n")
  1065. # automatically choose correct t-test
  1066. if (levene_p > 0.05) {
  1067. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  1068. t_result <- t.test(
  1069. birth_hc ~ maternal_preg_anaemia_status,
  1070. data = unique_data,
  1071. var.equal = TRUE
  1072. )
  1073. } else {
  1074. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1075. t_result <- t.test(
  1076. birth_hc ~ maternal_preg_anaemia_status,
  1077. data = unique_data,
  1078. var.equal = FALSE
  1079. )
  1080. }
  1081. # print t-test result in full precision
  1082. print(t_result)
  1083. ```
  1084. *Gestational Age at Maternal Hb Measurement*
  1085. ```{r}
  1086. # ga at Hb measurement (full sample)
  1087. # unique subject IDs: based on true prevalence
  1088. HF_ida_final %>%
  1089. filter(!is.na(maternal_preg_anaemia_status),
  1090. !is.na(maternal_min_hb_ga_roundown)) %>%
  1091. distinct(subject_id, .keep_all = TRUE) %>%
  1092. summarise(
  1093. n = n(),
  1094. mean_ga_hb = round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  1095. sd_ga_hb = round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  1096. min_ga_hb = round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  1097. q1_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.25, na.rm = TRUE), 2),
  1098. median_ga_hb = round(median(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  1099. q3_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.75, na.rm = TRUE), 2),
  1100. max_ga_hb = round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2)
  1101. )
  1102. ```
  1103. ```{r}
  1104. # group differences in ga at Hb measurement by antenatal maternal anaemia status
  1105. # ttest: ga at Hb measurement
  1106. # unique subject IDs: based on true prevalence
  1107. library(dplyr)
  1108. library(car)
  1109. # filter non-missing and keep one row per subject
  1110. HF_ida_final_unique <- HF_ida_final %>%
  1111. filter(!is.na(maternal_preg_anaemia_status),
  1112. !is.na(maternal_min_hb_ga_roundown)) %>%
  1113. distinct(subject_id, .keep_all = TRUE)
  1114. # summary statistics with exactly 2 decimals displayed
  1115. summary_stats <- HF_ida_final_unique %>%
  1116. group_by(maternal_preg_anaemia_status) %>%
  1117. summarise(
  1118. n = n(),
  1119. mean_ga = format(round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  1120. sd_ga = format(round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  1121. min_ga = format(round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  1122. max_ga = format(round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  1123. .groups = "drop"
  1124. )
  1125. # print summary stats
  1126. print(summary_stats)
  1127. # levene's test
  1128. levene_result <- leveneTest(
  1129. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  1130. data = HF_ida_final_unique
  1131. )
  1132. print(levene_result)
  1133. # extract Levene p-value (rounded for reporting)
  1134. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  1135. cat("\nLevene’s test p-value =", levene_p, "\n")
  1136. # automatically choose correct t-test
  1137. if (levene_p > 0.05) {
  1138. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  1139. t_result <- t.test(
  1140. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  1141. data = HF_ida_final_unique,
  1142. var.equal = TRUE
  1143. )
  1144. } else {
  1145. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1146. t_result <- t.test(
  1147. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  1148. data = HF_ida_final_unique,
  1149. var.equal = FALSE
  1150. )
  1151. }
  1152. # print t-test result in full precision
  1153. print(t_result)
  1154. ```
  1155. *Minimum Maternal Hb measurement in Trimester 1*
  1156. ```{r}
  1157. # min Hb across groups in trimester 1
  1158. # unique subject IDs: based on true prevalence
  1159. library(dplyr)
  1160. library(car)
  1161. # filter for trimester 1, non-missing anaemia status and min Hb, keep unique subjects
  1162. trim1_data_unique <- HF_ida_final %>%
  1163. filter(
  1164. maternal_trimester_hb == "first",
  1165. !is.na(maternal_preg_anaemia_status),
  1166. !is.na(maternal_preg_min_hb)
  1167. ) %>%
  1168. distinct(subject_id, .keep_all = TRUE)
  1169. # summary statistics by anaemia status with exactly 2 decimals displayed
  1170. summary_stats <- trim1_data_unique %>%
  1171. group_by(maternal_preg_anaemia_status) %>%
  1172. summarise(
  1173. n = n(),
  1174. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1175. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1176. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1177. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1178. .groups = "drop"
  1179. )
  1180. # print summary stats
  1181. print(summary_stats)
  1182. # levene's test for homogeneity of variance
  1183. levene_result <- leveneTest(
  1184. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1185. data = trim1_data_unique
  1186. )
  1187. print(levene_result)
  1188. # extract Levene p-value (rounded for reporting only)
  1189. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  1190. cat("\nLevene’s test p-value =", levene_p, "\n")
  1191. # automatically choose correct t-test
  1192. if (levene_p > 0.05) {
  1193. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  1194. t_result <- t.test(
  1195. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1196. data = trim1_data_unique,
  1197. var.equal = TRUE
  1198. )
  1199. } else {
  1200. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1201. t_result <- t.test(
  1202. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1203. data = trim1_data_unique,
  1204. var.equal = FALSE
  1205. )
  1206. }
  1207. # print t-test result in full precision
  1208. print(t_result)
  1209. ```
  1210. *Minimum Maternal Hb in trimester 2*
  1211. ```{r}
  1212. # min Hb across groups in trimester 2
  1213. # unique subject IDs: Based on true prevalence
  1214. library(dplyr)
  1215. library(car)
  1216. # filter for trimester 2, non-missing anaemia status and min Hb, keep unique subjects
  1217. trim2_data_unique <- HF_ida_final %>%
  1218. filter(
  1219. maternal_trimester_hb == "second",
  1220. !is.na(maternal_preg_anaemia_status),
  1221. !is.na(maternal_preg_min_hb)
  1222. ) %>%
  1223. distinct(subject_id, .keep_all = TRUE)
  1224. # summary statistics by anaemia status with exactly 2 decimals displayed
  1225. summary_stats <- trim2_data_unique %>%
  1226. group_by(maternal_preg_anaemia_status) %>%
  1227. summarise(
  1228. n = n(),
  1229. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1230. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1231. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1232. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1233. .groups = "drop"
  1234. )
  1235. # print summary stats
  1236. print(summary_stats)
  1237. # levene's test for homogeneity of variance
  1238. levene_result <- leveneTest(
  1239. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1240. data = trim2_data_unique
  1241. )
  1242. print(levene_result)
  1243. # extract Levene p-value (rounded for reporting only)
  1244. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  1245. cat("\nLevene’s test p-value =", levene_p, "\n")
  1246. # automatically choose correct t-test
  1247. if (levene_p > 0.05) {
  1248. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  1249. t_result <- t.test(
  1250. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1251. data = trim2_data_unique,
  1252. var.equal = TRUE
  1253. )
  1254. } else {
  1255. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1256. t_result <- t.test(
  1257. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1258. data = trim2_data_unique,
  1259. var.equal = FALSE
  1260. )
  1261. }
  1262. # print t-test result in full precision
  1263. print(t_result)
  1264. ```
  1265. *Minimum Maternal Hb in Trimester 3*
  1266. ```{r}
  1267. # min Hb across groups in trimester 3
  1268. # unique subject IDs: based on true prevalence
  1269. library(dplyr)
  1270. library(car)
  1271. # filter for trimester 3, non-missing anaemia status and min Hb, keep unique subjects
  1272. trim3_data_unique <- HF_ida_final %>%
  1273. filter(
  1274. maternal_trimester_hb == "third",
  1275. !is.na(maternal_preg_anaemia_status),
  1276. !is.na(maternal_preg_min_hb)
  1277. ) %>%
  1278. distinct(subject_id, .keep_all = TRUE)
  1279. # summary statistics by anaemia status with exactly 2 decimals displayed
  1280. summary_stats <- trim3_data_unique %>%
  1281. group_by(maternal_preg_anaemia_status) %>%
  1282. summarise(
  1283. n = n(),
  1284. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1285. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1286. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1287. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  1288. .groups = "drop"
  1289. )
  1290. # print summary stats
  1291. print(summary_stats)
  1292. # levene's test for homogeneity of variance
  1293. levene_result <- leveneTest(
  1294. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1295. data = trim3_data_unique
  1296. )
  1297. print(levene_result)
  1298. # extract levene p-value (rounded for reporting only)
  1299. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  1300. cat("\nLevene’s test p-value =", levene_p, "\n")
  1301. # automatically choose correct t-test
  1302. if (levene_p > 0.05) {
  1303. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  1304. t_result <- t.test(
  1305. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1306. data = trim3_data_unique,
  1307. var.equal = TRUE
  1308. )
  1309. } else {
  1310. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  1311. t_result <- t.test(
  1312. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  1313. data = trim3_data_unique,
  1314. var.equal = FALSE
  1315. )
  1316. }
  1317. # print t-test result in full precision
  1318. print(t_result)
  1319. ```
  1320. **Exploratory Analyses: Plotted Data using LOESS Curves**
  1321. Absolute Volumes for ROIs plotted against child age at scan. Includes repeated measures to capture all scans collected across study timepoints.
  1322. 1. Full sample
  1323. 2. By group (Key: Antenatal Maternal Anaemia Status)
  1324. *LOESS Curves: ICV*
  1325. ```{r}
  1326. # icv (full sample)
  1327. # including repeated measures
  1328. library(ggplot2)
  1329. icv_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = hf_mm_icv)) +
  1330. geom_point(alpha = 0.6) +
  1331. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  1332. scale_x_continuous(
  1333. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  1334. ) +
  1335. labs(
  1336. title = "ICV vs. Child Age",
  1337. x = "Child Age at Scan (Months)",
  1338. y = "Total ICV"
  1339. ) +
  1340. theme_minimal()
  1341. icv_hf_sample
  1342. ```
  1343. ```{r}
  1344. # icv by group (antenatal maternal anaemia status)
  1345. # including repeated measures
  1346. library(ggplot2)
  1347. library(dplyr)
  1348. icv_hf_group <- ggplot(
  1349. data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  1350. aes(
  1351. x = scan_age_months,
  1352. y = hf_mm_icv,
  1353. color = maternal_preg_anaemia_status,
  1354. fill = maternal_preg_anaemia_status
  1355. )
  1356. ) +
  1357. geom_point() +
  1358. geom_smooth(
  1359. method = "loess",
  1360. se = TRUE,
  1361. level = 0.95,
  1362. alpha = 0.3 # make ribbons slightly transparent
  1363. ) +
  1364. scale_color_manual(
  1365. values = c(
  1366. "no_mat_anaemia" = "#1f77b4", # blue line
  1367. "mat_anaemia" = "#d62728" # red line
  1368. ),
  1369. name = "Maternal Anaemia Status"
  1370. ) +
  1371. scale_fill_manual(
  1372. values = c(
  1373. "no_mat_anaemia" = "#c6d9f1", # lighter blue ribbon
  1374. "mat_anaemia" = "#f7b0b0" # lighter red ribbon
  1375. ),
  1376. guide = "none" # hide fill legend
  1377. ) +
  1378. scale_x_continuous(
  1379. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  1380. ) +
  1381. labs(
  1382. title = "Total ICV Volume vs. Scan Age",
  1383. x = "Scan Age (Months)",
  1384. y = "Total ICV Volume (mm3)",
  1385. color = "Maternal Anaemia Status"
  1386. ) +
  1387. theme_minimal()
  1388. icv_hf_group
  1389. ```
  1390. *LOESS Curves: Putamen*
  1391. ```{r}
  1392. # putamen (full sample)
  1393. # including repeated measures
  1394. library(ggplot2)
  1395. putamen_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_putamen)) +
  1396. geom_point(alpha = 0.6) +
  1397. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  1398. scale_x_continuous(
  1399. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  1400. ) +
  1401. labs(
  1402. title = "Putamen vs. Child Age",
  1403. x = "Child Age at Scan (Months)",
  1404. y = "Total Putamen Volume"
  1405. ) +
  1406. theme_minimal()
  1407. putamen_hf_sample
  1408. ```
  1409. ```{r}
  1410. # putamen by group (antenatal maternal anaemia status)
  1411. # including repeated measures
  1412. library(ggplot2)
  1413. library(dplyr)
  1414. putamen_hf_group <- ggplot(
  1415. data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  1416. aes(
  1417. x = scan_age_months,
  1418. y = total_putamen,
  1419. color = maternal_preg_anaemia_status,
  1420. fill = maternal_preg_anaemia_status
  1421. )
  1422. ) +
  1423. geom_point() +
  1424. geom_smooth(
  1425. method = "loess",
  1426. se = TRUE,
  1427. level = 0.95,
  1428. alpha = 0.3 # slightly transparent ribbons
  1429. ) +
  1430. scale_color_manual(
  1431. values = c(
  1432. "no_mat_anaemia" = "#1f77b4", # blue line
  1433. "mat_anaemia" = "#d62728" # red line
  1434. ),
  1435. name = "Maternal Anaemia Status"
  1436. ) +
  1437. scale_fill_manual(
  1438. values = c(
  1439. "no_mat_anaemia" = "#c6d9f1", # lighter blue ribbon
  1440. "mat_anaemia" = "#f7b0b0" # lighter red ribbon
  1441. ),
  1442. guide = "none" # hide fill legend
  1443. ) +
  1444. scale_x_continuous(
  1445. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  1446. ) +
  1447. labs(
  1448. title = "Total Putamen Volume vs. Scan Age",
  1449. x = "Scan Age (Months)",
  1450. y = "Total Putamen Volume",
  1451. color = "Maternal Anaemia Status"
  1452. ) +
  1453. theme_minimal()
  1454. putamen_hf_group
  1455. ```
  1456. *LOESS Curves: Caudate Nucleus*
  1457. ```{r}
  1458. # caudate Nucleus (full sample)
  1459. # including repeated measures
  1460. library(ggplot2)
  1461. caudate_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_caudate)) +
  1462. geom_point(alpha = 0.6) +
  1463. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  1464. scale_x_continuous(
  1465. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  1466. ) +
  1467. labs(
  1468. title = "Caudate vs. Child Age",
  1469. x = "Child Age at Scan (Months)",
  1470. y = "Total Caudate Volume"
  1471. ) +
  1472. theme_minimal()
  1473. caudate_hf_sample
  1474. ```
  1475. ```{r}
  1476. # caudate nucleus by group (antenatal maternal anaemia status)
  1477. # including repeated measures
  1478. library(ggplot2)
  1479. library(dplyr)
  1480. caudate_hf_group <- ggplot(
  1481. data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  1482. aes(
  1483. x = scan_age_months,
  1484. y = total_caudate,
  1485. color = maternal_preg_anaemia_status,
  1486. fill = maternal_preg_anaemia_status
  1487. )
  1488. ) +
  1489. geom_point() +
  1490. geom_smooth(
  1491. method = "loess",
  1492. se = TRUE,
  1493. level = 0.95,
  1494. alpha = 0.3 # light, semi-transparent ribbon
  1495. ) +
  1496. scale_color_manual(
  1497. values = c(
  1498. "no_mat_anaemia" = "#1f77b4", # blue line
  1499. "mat_anaemia" = "#d62728" # red line
  1500. ),
  1501. name = "Maternal Anaemia Status"
  1502. ) +
  1503. scale_fill_manual(
  1504. values = c(
  1505. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
  1506. "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
  1507. ),
  1508. guide = "none" # hide fill legend
  1509. ) +
  1510. scale_x_continuous(
  1511. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  1512. ) +
  1513. labs(
  1514. title = "Total Caudate Volume vs. Scan Age",
  1515. x = "Scan Age (Months)",
  1516. y = "Total Caudate Volume",
  1517. color = "Maternal Anaemia Status"
  1518. ) +
  1519. theme_minimal()
  1520. caudate_hf_group
  1521. ```
  1522. *LOESS Curves: Corpus Callosum*
  1523. ```{r}
  1524. # corpus callosum (full sample)
  1525. # including repeated measures
  1526. library(ggplot2)
  1527. cc_hf_sample <- ggplot(HF_ida_final, aes(x = scan_age_months, y = total_cc)) +
  1528. geom_point(alpha = 0.6) +
  1529. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  1530. scale_x_continuous(
  1531. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  1532. ) +
  1533. labs(
  1534. title = "CC vs. Child Age",
  1535. x = "Child Age at Scan (Months)",
  1536. y = "Total Corpus Callosum Volume"
  1537. ) +
  1538. theme_minimal()
  1539. cc_hf_sample
  1540. ```
  1541. ```{r}
  1542. # corpus callosum by group (antenatal maternal anaemia status)
  1543. # including repeated measures
  1544. library(ggplot2)
  1545. library(dplyr)
  1546. cc_hf_group <- ggplot(
  1547. data = HF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  1548. aes(
  1549. x = scan_age_months,
  1550. y = total_cc,
  1551. color = maternal_preg_anaemia_status,
  1552. fill = maternal_preg_anaemia_status
  1553. )
  1554. ) +
  1555. geom_point() +
  1556. geom_smooth(
  1557. method = "loess",
  1558. se = TRUE,
  1559. level = 0.95,
  1560. alpha = 0.3 # make ribbons slightly transparent
  1561. ) +
  1562. scale_color_manual(
  1563. values = c(
  1564. "no_mat_anaemia" = "#1f77b4", # original blue line
  1565. "mat_anaemia" = "#d62728" # original red line
  1566. ),
  1567. name = "Maternal Anaemia Status"
  1568. ) +
  1569. scale_fill_manual(
  1570. values = c(
  1571. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue for ribbon
  1572. "mat_anaemia" = "#f7b0b0" # slightly lighter red for ribbon
  1573. ),
  1574. guide = "none" # hide fill legend
  1575. ) +
  1576. scale_x_continuous(
  1577. breaks = seq(0, max(HF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  1578. ) +
  1579. scale_y_continuous(
  1580. n.breaks = 6
  1581. ) +
  1582. labs(
  1583. title = "Total CC Volume vs. Scan Age",
  1584. x = "Scan Age (Months)",
  1585. y = "Total Corpus Callosum Volume (mm³)",
  1586. color = "Maternal Anaemia Status"
  1587. ) +
  1588. theme_minimal()
  1589. cc_hf_group
  1590. ```
  1591. **Zero-Order Correlation Matrices: Antenatal Maternal Anaemia Status**
  1592. ```{r}
  1593. # install packages if needed
  1594. # install and load Hmisc
  1595. if (!require(Hmisc)) {
  1596. install.packages("Hmisc")
  1597. library(Hmisc)
  1598. } else {
  1599. library(Hmisc)}
  1600. ```
  1601. Convert variables to numeric for correlation matrices
  1602. ```{r}
  1603. # convert variables back from categorical to numeric
  1604. HF_ida_final$income_household_en <- as.numeric(HF_ida_final$income_household_en)
  1605. HF_ida_final$maternal_preg_anaemia_status_numeric <- as.numeric(as.character(
  1606. factor(HF_ida_final$maternal_preg_anaemia_status,
  1607. levels = c("no_mat_anaemia", "mat_anaemia"),
  1608. labels = c(0, 1))
  1609. ))
  1610. HF_ida_final$maternal_hiv <- as.numeric(as.character(
  1611. factor(HF_ida_final$maternal_hiv,
  1612. levels = c("no_hiv", "mat_hiv"),
  1613. labels = c(0, 1))
  1614. ))
  1615. HF_ida_final$child_sex <- as.numeric(as.character(
  1616. factor(HF_ida_final$child_sex,
  1617. levels = c("girl", "boy"),
  1618. labels = c(0, 1))
  1619. ))
  1620. ```
  1621. \*\* The variable for "antenatal maternal anaemia status" remained a categorical variable even after running the code above. Therefore, I recalled the dataset to ensure that this was numeric (as initially defined) for assessing correlations:
  1622. ```{r}
  1623. # recall HF dataset
  1624. HF_ida_final <- read_excel("HF_ida_final.xlsx")
  1625. # create new variables for total brain volumes across ROIs
  1626. # basal ganglia
  1627. HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
  1628. HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
  1629. # corpus callosum
  1630. HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
  1631. # create new age variable (days to months):
  1632. # convert days to months
  1633. # average Number of days in a month (accounting for leap years): 30.44
  1634. HF_ida_final <- HF_ida_final %>%
  1635. mutate(scan_age_months = hf_age / 30.44)
  1636. ```
  1637. **Correlation Matrices**
  1638. ```{r}
  1639. # correlation matrix for corpus callosum
  1640. # filter dataset
  1641. filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
  1642. # select relevant columns
  1643. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_cc")
  1644. data_selected <- filtered_data[, vars]
  1645. # convert categorical variables to numeric (for correlation)
  1646. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  1647. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  1648. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  1649. # compute correlation matrix with significance
  1650. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  1651. # rcorr_result$r contains correlation coefficients
  1652. print("Correlation matrix:")
  1653. print(rcorr_result$r)
  1654. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  1655. print("P-value matrix:")
  1656. print(rcorr_result$P)
  1657. ```
  1658. ```{r}
  1659. # correlation matrix for caudate nucleus
  1660. # filter dataset
  1661. filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
  1662. # select relevant columns
  1663. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_caudate")
  1664. data_selected <- filtered_data[, vars]
  1665. # convert categorical variables to numeric (for correlation)
  1666. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  1667. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  1668. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  1669. # compute correlation matrix with significance
  1670. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  1671. # rcorr_result$r contains correlation coefficients
  1672. print("Correlation matrix:")
  1673. print(rcorr_result$r)
  1674. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  1675. print("P-value matrix:")
  1676. print(rcorr_result$P)
  1677. ```
  1678. ```{r}
  1679. # correlation matrix for putamen
  1680. # filter dataset
  1681. filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
  1682. # select relevant columns
  1683. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_putamen")
  1684. data_selected <- filtered_data[, vars]
  1685. # convert categorical variables to numeric (for correlation)
  1686. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  1687. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  1688. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  1689. # compute correlation matrix with significance
  1690. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  1691. # rcorr_result$r contains correlation coefficients
  1692. print("Correlation matrix:")
  1693. print(rcorr_result$r)
  1694. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  1695. print("P-value matrix:")
  1696. print(rcorr_result$P)
  1697. ```
  1698. ```{r}
  1699. # correlation matrix for ICV
  1700. # filter dataset
  1701. filtered_data <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
  1702. # select relevant columns
  1703. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "hf_mm_icv")
  1704. data_selected <- filtered_data[, vars]
  1705. # convert categorical variables to numeric (for correlation)
  1706. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  1707. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  1708. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  1709. # compute correlation matrix with significance
  1710. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  1711. # rcorr_result$r contains correlation coefficients
  1712. print("Correlation matrix:")
  1713. print(rcorr_result$r)
  1714. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  1715. print("P-value matrix:")
  1716. print(rcorr_result$P)
  1717. ```
  1718. **Multicollinearity**
  1719. ```{r}
  1720. # correlation between child age at scan and ICV
  1721. library(dplyr)
  1722. correlation_test <- cor.test(
  1723. HF_ida_final$scan_age_months[!is.na(HF_ida_final$maternal_preg_anaemia_status)],
  1724. HF_ida_final$hf_mm_icv[!is.na(HF_ida_final$maternal_preg_anaemia_status)])
  1725. print(correlation_test)
  1726. ```
  1727. Decision to exclude ICV from the models for regional brain volumes due to high levels of multicollinearity (*r* \> 0.80) between child age at scan and ICV. Only child age at scan was included in the models below.
  1728. **Linear Mixed Effects (LME) Models**
  1729. ```{r}
  1730. # install packages if needed
  1731. # install.packages("lme4")
  1732. library(lme4)
  1733. # install.packages("lmerTest")
  1734. library(lmerTest)
  1735. # install.packages("parameters")
  1736. library(parameters)
  1737. # install.packages("emmeans")
  1738. library(emmeans)
  1739. ```
  1740. Based on the exploratory analyses (plotted data; LOESS curves), absolute age was used for models with an observed linear relationship. Log (age) was used for models with an observed non-linear relationship.
  1741. In the HF dataset:
  1742. - Corpus callosum: Absolute age - Linear growth trajectory on plot
  1743. - ICV, Putamen, Caudate Nucleus: Log (age) - Growth trajectory characterised by rapid growth and plateauing on plot
  1744. For the HF sample, no interactions (child age at scan \* antenatal maternal anaemia status) were explored due to limited power. We kept the models simple and only assessed main effects.
  1745. **LME Model variables**
  1746. - Recall the dataset to reset variables
  1747. - Create variabels for total ROI volumes and child age at scan in months
  1748. - Ensure that all categorical variables (e.g. anaemia status, child sex, maternal HIV etc) are factors
  1749. - Create variable for the logarithmic transformation of child age in months
  1750. ```{r}
  1751. # recall dataset
  1752. HF_ida_final <- read_excel("HF_ida_final.xlsx")
  1753. # create new variables for total brain volumes across ROIs
  1754. # basal ganglia
  1755. HF_ida_final$total_putamen <- HF_ida_final$hf_mm_left_putamen + HF_ida_final$hf_mm_right_putamen
  1756. HF_ida_final$total_caudate <- HF_ida_final$hf_mm_left_caudate + HF_ida_final$hf_mm_right_caudate
  1757. # corpus callosum
  1758. HF_ida_final$total_cc <- HF_ida_final$hf_mm_posterior_callosum + HF_ida_final$hf_mm_mid_posterior_callosum + HF_ida_final$hf_mm_central_callosum + HF_ida_final$hf_mm_mid_anterior_callosum + HF_ida_final$hf_mm_anterior_callosum
  1759. # create new age variable (days to months):
  1760. # convert days to months
  1761. # average Number of days in a month (accounting for leap years): 30.44
  1762. HF_ida_final <- HF_ida_final %>%
  1763. mutate(scan_age_months = hf_age / 30.44)
  1764. ```
  1765. ```{r}
  1766. # recode variables from numeric to factor
  1767. # maternal anaemia status
  1768. HF_ida_final$maternal_preg_anaemia_status <- factor(
  1769. HF_ida_final$maternal_preg_anaemia_status,
  1770. levels = c(0, 1),
  1771. labels = c("no_mat_anaemia", "mat_anaemia")
  1772. )
  1773. # maternal ID status
  1774. HF_ida_final$maternal_preg_adj_brinda_id <- factor(
  1775. HF_ida_final$maternal_preg_adj_brinda_id,
  1776. levels = c(0, 1),
  1777. labels = c("no_mat_id", "mat_id")
  1778. )
  1779. # maternal ID status
  1780. HF_ida_final$maternal_preg_unadjusted_id <- factor(
  1781. HF_ida_final$maternal_preg_unadjusted_id,
  1782. levels = c(0, 1),
  1783. labels = c("no_mat_id", "mat_id")
  1784. )
  1785. # maternal trimester Hb
  1786. HF_ida_final$maternal_trimester_hb <- factor(
  1787. HF_ida_final$maternal_trimester_hb,
  1788. levels = c(1, 2, 3),
  1789. labels = c("first", "second", "third")
  1790. )
  1791. # child sex
  1792. HF_ida_final$child_sex <- factor(
  1793. HF_ida_final$child_sex,
  1794. levels = c(0, 1),
  1795. labels = c("girl", "boy")
  1796. )
  1797. # maternal pae
  1798. HF_ida_final$maternal_pae <- factor(
  1799. HF_ida_final$maternal_pae,
  1800. levels = c(0, 1),
  1801. labels = c("no_pae", "pae")
  1802. )
  1803. # maternal pte
  1804. HF_ida_final$maternal_pte <- factor(
  1805. HF_ida_final$maternal_pte,
  1806. levels = c(0, 1),
  1807. labels = c("no_pte", "pte")
  1808. )
  1809. # maternal hiv
  1810. HF_ida_final$maternal_hiv <- factor(
  1811. HF_ida_final$maternal_hiv,
  1812. levels = c(0, 1),
  1813. labels = c("no_hiv", "hiv")
  1814. )
  1815. # maternal dep
  1816. HF_ida_final$maternal_dep <- factor(
  1817. HF_ida_final$maternal_dep,
  1818. levels = c(0, 1),
  1819. labels = c("no_dep", "dep")
  1820. )
  1821. # maternal employment
  1822. HF_ida_final$mom_employment_en_dichotomous <- factor(
  1823. HF_ida_final$mom_employment_en_dichotomous,
  1824. levels = c(0, 1),
  1825. labels = c("unemployed", "employed")
  1826. )
  1827. # household income
  1828. HF_ida_final$income_household_en <- factor(
  1829. HF_ida_final$income_household_en,
  1830. levels = c(1, 2, 3, 4),
  1831. labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
  1832. )
  1833. # maternal education
  1834. HF_ida_final$mom_edu_en <- factor(
  1835. HF_ida_final$mom_edu_en,
  1836. levels = c(2, 3, 4, 5, 6),
  1837. labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
  1838. )
  1839. # baby hiv
  1840. HF_ida_final$baby_hiv <- factor(
  1841. HF_ida_final$baby_hiv,
  1842. levels = c(0, 1),
  1843. labels = c("no_child_hiv", "child_hiv")
  1844. )
  1845. ```
  1846. ```{r}
  1847. # create a log(age) variable for use in models of brain regions with non-linear growth
  1848. HF_ida_final$log_age <- log(HF_ida_final$scan_age_months)
  1849. ```
  1850. *LME: Corpus Callosum*
  1851. ```{r}
  1852. # LME for corpus callosum
  1853. # no interaction between maternal anaemia and age
  1854. # absolute age for linear growth
  1855. # model summary
  1856. hf_model_cc <- lmer(total_cc ~ maternal_preg_anaemia_status + scan_age_months + child_sex + (1 | subject_id),
  1857. data = HF_ida_final)
  1858. summary(hf_model_cc)
  1859. # beta values
  1860. standardize_parameters(hf_model_cc, method="refit")
  1861. ```
  1862. ```{r}
  1863. # emms for corpus callosum volume by antenatal maternal anaemia status
  1864. # based on final adjusted model
  1865. library(emmeans)
  1866. emm_cc <- emmeans(hf_model_cc, ~ maternal_preg_anaemia_status)
  1867. # ----- EMM table with exactly 2 decimals -----
  1868. emm_cc_df <- as.data.frame(emm_cc)
  1869. emm_cc_df$emmean <- format(round(emm_cc_df$emmean, 2), nsmall = 2)
  1870. emm_cc_df$SE <- format(round(emm_cc_df$SE, 2), nsmall = 2)
  1871. emm_cc_df$lower.CL <- format(round(emm_cc_df$lower.CL, 2), nsmall = 2)
  1872. emm_cc_df$upper.CL <- format(round(emm_cc_df$upper.CL, 2), nsmall = 2)
  1873. print(emm_cc_df)
  1874. # ----- contrast results -----
  1875. contrast_results_cc <- contrast(emm_cc, method = "revpairwise", adjust = "holm")
  1876. contrast_df_cc <- as.data.frame(summary(contrast_results_cc, infer = TRUE))
  1877. # force 2 decimals for estimate
  1878. contrast_df_cc$estimate <- format(round(contrast_df_cc$estimate, 2), nsmall = 2)
  1879. print(contrast_df_cc)
  1880. ```
  1881. *LME: ICV*
  1882. ```{r}
  1883. # LME for icv
  1884. # no interaction between maternal anaemia and age
  1885. # log (scan age) as growth non-linear
  1886. # model summary
  1887. hf_model_icv <- lmer(hf_mm_icv ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
  1888. data = HF_ida_final)
  1889. summary(hf_model_icv)
  1890. # beta values
  1891. standardize_parameters(hf_model_icv)
  1892. ```
  1893. ```{r}
  1894. # emms for icv by antenatal maternal anaemia status
  1895. # based on final adjusted model
  1896. library(emmeans)
  1897. emm_icv <- emmeans(hf_model_icv, ~ maternal_preg_anaemia_status)
  1898. # ----- EMM table with exactly 2 decimals -----
  1899. emm_icv_df <- as.data.frame(emm_icv)
  1900. emm_icv_df$emmean <- format(round(emm_icv_df$emmean, 2), nsmall = 2)
  1901. emm_icv_df$SE <- format(round(emm_icv_df$SE, 2), nsmall = 2)
  1902. emm_icv_df$lower.CL <- format(round(emm_icv_df$lower.CL, 2), nsmall = 2)
  1903. emm_icv_df$upper.CL <- format(round(emm_icv_df$upper.CL, 2), nsmall = 2)
  1904. print(emm_icv_df)
  1905. # ----- contrast results -----
  1906. contrast_results_icv <- contrast(emm_icv, method = "revpairwise", adjust = "holm")
  1907. contrast_df_icv <- as.data.frame(summary(contrast_results_icv, infer = TRUE))
  1908. # force 2 decimals for estimate
  1909. contrast_df_icv$estimate <- format(round(contrast_df_icv$estimate, 2), nsmall = 2)
  1910. print(contrast_df_icv)
  1911. ```
  1912. *LME: Caudate Nucleus*
  1913. ```{r}
  1914. # lme for caudate nucleus
  1915. # no interaction between maternal anaemia and age
  1916. # log (scan age) as growth non-linear
  1917. hf_model_caudate <- lmer(total_caudate ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
  1918. data = HF_ida_final)
  1919. summary(hf_model_caudate)
  1920. # beta values
  1921. standardize_parameters(hf_model_caudate)
  1922. ```
  1923. ```{r}
  1924. # emms for caudate nucleus by antenatal maternal anaemia status
  1925. # based on final adjusted model
  1926. library(emmeans)
  1927. emm_caudate <- emmeans(hf_model_caudate, ~ maternal_preg_anaemia_status)
  1928. # ----- emm table with exactly 2 decimals -----
  1929. emm_caudate_df <- as.data.frame(emm_caudate)
  1930. emm_caudate_df$emmean <- format(round(emm_caudate_df$emmean, 2), nsmall = 2)
  1931. emm_caudate_df$SE <- format(round(emm_caudate_df$SE, 2), nsmall = 2)
  1932. emm_caudate_df$lower.CL <- format(round(emm_caudate_df$lower.CL, 2), nsmall = 2)
  1933. emm_caudate_df$upper.CL <- format(round(emm_caudate_df$upper.CL, 2), nsmall = 2)
  1934. print(emm_caudate_df)
  1935. # ----- contrast results -----
  1936. contrast_results_caudate <- contrast(emm_caudate, method = "revpairwise", adjust = "holm")
  1937. contrast_df_caudate <- as.data.frame(summary(contrast_results_caudate, infer = TRUE))
  1938. # force 2 decimals for estimate
  1939. contrast_df_caudate$estimate <- format(round(contrast_df_caudate$estimate, 2), nsmall = 2)
  1940. print(contrast_df_caudate)
  1941. ```
  1942. *LME: Putamen*
  1943. ```{r}
  1944. # lme for putamen
  1945. # no interaction between maternal anaemia and age
  1946. # log (scan age) as growth non-linear
  1947. hf_model_putamen <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
  1948. data = HF_ida_final)
  1949. summary(hf_model_putamen)
  1950. # beta values
  1951. standardize_parameters(hf_model_putamen)
  1952. ```
  1953. ```{r}
  1954. # emms for caudate nucleus by antenatal maternal anaemia status
  1955. # based on final adjusted model
  1956. library(emmeans)
  1957. emm_putamen <- emmeans(hf_model_putamen, ~ maternal_preg_anaemia_status)
  1958. # ----- emm table with exactly 2 decimals -----
  1959. emm_putamen_df <- as.data.frame(emm_putamen)
  1960. emm_putamen_df$emmean <- format(round(emm_putamen_df$emmean, 2), nsmall = 2)
  1961. emm_putamen_df$SE <- format(round(emm_putamen_df$SE, 2), nsmall = 2)
  1962. emm_putamen_df$lower.CL <- format(round(emm_putamen_df$lower.CL, 2), nsmall = 2)
  1963. emm_putamen_df$upper.CL <- format(round(emm_putamen_df$upper.CL, 2), nsmall = 2)
  1964. print(emm_putamen_df)
  1965. # ----- contrast results -----
  1966. contrast_results_putamen <- contrast(emm_putamen, method = "revpairwise", adjust = "holm")
  1967. contrast_df_putamen <- as.data.frame(summary(contrast_results_putamen, infer = TRUE))
  1968. # force 2 decimals for estimate
  1969. contrast_df_putamen$estimate <- format(round(contrast_df_putamen$estimate, 2), nsmall = 2)
  1970. print(contrast_df_putamen)
  1971. ```
  1972. **Sensitivity Analyses**
  1973. Based on the final models for each ROI (established in primary analyses): With the inclusion of maternal HIV (potential confounder)
  1974. ```{r}
  1975. # sensitivity analysis for corpus callosum
  1976. # including maternal hiv
  1977. # excluding interaction between maternal anaemia and age
  1978. hf_model_cc_sens <- lmer(total_cc ~ maternal_preg_anaemia_status + scan_age_months + child_sex + maternal_hiv + (1 | subject_id),
  1979. data = HF_ida_final)
  1980. summary(hf_model_cc_sens)
  1981. standardize_parameters(hf_model_cc_sens)
  1982. ```
  1983. ```{r}
  1984. # sensitivity analysis for icv
  1985. # including maternal hiv
  1986. # excluding interaction between maternal anaemia and age
  1987. hf_model_icv_sens <- lmer(hf_mm_icv ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
  1988. data = HF_ida_final)
  1989. summary(hf_model_icv_sens)
  1990. standardize_parameters(hf_model_icv_sens)
  1991. ```
  1992. ```{r}
  1993. # sensitivity analysis for putamen
  1994. # including maternal hiv
  1995. # excluding interaction between maternal anaemia and age
  1996. hf_model_putamen_sens <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
  1997. data = HF_ida_final)
  1998. summary(hf_model_putamen_sens)
  1999. standardize_parameters(hf_model_putamen_sens)
  2000. ```
  2001. ```{r}
  2002. # sensitivity analysis for caudate nucleus
  2003. # including maternal hiv
  2004. # excluding interaction between maternal anaemia and age
  2005. hf_model_caudate_sens <- lmer(total_caudate ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
  2006. data = HF_ida_final)
  2007. summary(hf_model_caudate_sens)
  2008. standardize_parameters(hf_model_caudate_sens)
  2009. ```
  2010. ------------------------------------------------------------------------
  2011. [***Ultra-Low-Field (ULF) Analyses***]{.underline}
  2012. ------------------------------------------------------------------------
  2013. **Import ULF Dataset**
  2014. ```{r}
  2015. ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
  2016. ```
  2017. **Create New Variables for Total Brain Volumes across ROIs**
  2018. ```{r}
  2019. # basal ganglia
  2020. ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
  2021. ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
  2022. # corpus callosum
  2023. ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
  2024. ```
  2025. **Create New Age Variable (days to months)**
  2026. ```{r}
  2027. # convert days to months
  2028. # average number of days in a month (accounting for leap years): 30.44
  2029. ULF_ida_final <- ULF_ida_final %>%
  2030. mutate(scan_age_months = hyp_age / 30.44)
  2031. ```
  2032. **Create Meaningful Categorical Variables**
  2033. ```{r}
  2034. # convert from numeric to factor
  2035. # maternal anaemia status
  2036. ULF_ida_final$maternal_preg_anaemia_status <- factor(
  2037. ULF_ida_final$maternal_preg_anaemia_status,
  2038. levels = c(0, 1),
  2039. labels = c("no_mat_anaemia", "mat_anaemia")
  2040. )
  2041. # maternal anaemia severity
  2042. ULF_ida_final$maternal_preg_anaemia_severity <- factor(
  2043. ULF_ida_final$maternal_preg_anaemia_severity,
  2044. levels = c(1, 2, 3),
  2045. labels = c("mild", "moderate", "severe")
  2046. )
  2047. # child status
  2048. ULF_ida_final$cor_child_anaemia <- factor(
  2049. ULF_ida_final$cor_child_anaemia,
  2050. levels = c(0, 1),
  2051. labels = c("no_child_anaemia", "child_anaemia")
  2052. )
  2053. # maternal ID status
  2054. ULF_ida_final$maternal_preg_adj_brinda_id <- factor(
  2055. ULF_ida_final$maternal_preg_adj_brinda_id,
  2056. levels = c(0, 1),
  2057. labels = c("no_mat_id", "mat_id")
  2058. )
  2059. # child ID status
  2060. ULF_ida_final$cor_child_id <- factor(
  2061. ULF_ida_final$cor_child_id,
  2062. levels = c(0, 1),
  2063. labels = c("no_child_id", "child_id")
  2064. )
  2065. # maternal trimester hb
  2066. ULF_ida_final$maternal_trimester_hb <- factor(
  2067. ULF_ida_final$maternal_trimester_hb,
  2068. levels = c(1, 2, 3),
  2069. labels = c("first", "second", "third")
  2070. )
  2071. # child sex
  2072. ULF_ida_final$child_sex <- factor(
  2073. ULF_ida_final$child_sex,
  2074. levels = c(0, 1),
  2075. labels = c("girl", "boy")
  2076. )
  2077. # maternal pae
  2078. ULF_ida_final$maternal_pae <- factor(
  2079. ULF_ida_final$maternal_pae,
  2080. levels = c(0, 1),
  2081. labels = c("no_pae", "pae")
  2082. )
  2083. # maternal pte
  2084. ULF_ida_final$maternal_pte <- factor(
  2085. ULF_ida_final$maternal_pte,
  2086. levels = c(0, 1),
  2087. labels = c("no_pte", "pte")
  2088. )
  2089. # maternal hiv
  2090. ULF_ida_final$maternal_hiv <- factor(
  2091. ULF_ida_final$maternal_hiv,
  2092. levels = c(0, 1),
  2093. labels = c("no_hiv", "hiv")
  2094. )
  2095. # maternal dep
  2096. ULF_ida_final$maternal_dep <- factor(
  2097. ULF_ida_final$maternal_dep,
  2098. levels = c(0, 1),
  2099. labels = c("no_dep", "dep")
  2100. )
  2101. # maternal employment
  2102. ULF_ida_final$mom_employment_en_dichotomous <- factor(
  2103. ULF_ida_final$mom_employment_en_dichotomous,
  2104. levels = c(0, 1),
  2105. labels = c("unemployed", "employed")
  2106. )
  2107. # household income
  2108. ULF_ida_final$income_household_en <- factor(
  2109. ULF_ida_final$income_household_en,
  2110. levels = c(1, 2, 3, 4),
  2111. labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
  2112. )
  2113. # maternal education
  2114. ULF_ida_final$mom_edu_en <- factor(
  2115. ULF_ida_final$mom_edu_en,
  2116. levels = c(2, 3, 4, 5, 6),
  2117. labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
  2118. )
  2119. # baby hiv
  2120. ULF_ida_final$baby_hiv <- factor(
  2121. ULF_ida_final$baby_hiv,
  2122. levels = c(0, 1),
  2123. labels = c("no_child_hiv", "child_hiv")
  2124. )
  2125. ```
  2126. **Data Distribution across Timepoint**
  2127. ```{r}
  2128. # sample over 12 months with maternal anaemia data
  2129. sample_over_12M <- ULF_ida_final %>%
  2130. filter(scan_age_months > 12,
  2131. !is.na(maternal_preg_anaemia_status))
  2132. # sample under 12 months with maternal anaemia data
  2133. sample_under_12M <- ULF_ida_final %>%
  2134. filter(scan_age_months < 12,
  2135. !is.na(maternal_preg_anaemia_status))
  2136. # sample over 18 months with maternal anaemia data
  2137. sample_over_18 <- ULF_ida_final %>%
  2138. filter(scan_age_months > 18,
  2139. !is.na(maternal_preg_anaemia_status))
  2140. # sample over 24 months with maternal anaemia data
  2141. sample_over_24 <- ULF_ida_final %>%
  2142. filter(scan_age_months > 24,
  2143. !is.na(maternal_preg_anaemia_status))
  2144. ```
  2145. **Overlap Between HF and ULF Samples**
  2146. No. of unique subjects for full ULF neuroimaging sample:
  2147. ```{r}
  2148. # unique observations in ULF dataset
  2149. # excluding repeated measures
  2150. library(dplyr)
  2151. ULF_ida_final %>%
  2152. summarise(n_unique_subjects = n_distinct(subject_id))
  2153. ```
  2154. Overlap in unique subjects between full HF and ULF samples:
  2155. ```{r}
  2156. # no. of children in common between HF and ULF neuroimaging datasets
  2157. # find overlapping unique subject_ids
  2158. overlap_subjects <- intersect(HF_ida_final$subject_id, ULF_ida_final$subject_id)
  2159. # Count how many there are
  2160. num_overlap <- length(overlap_subjects)
  2161. # Print the number
  2162. num_overlap
  2163. ```
  2164. For neuroimaging sample with maternal anaemia data available:
  2165. ```{r}
  2166. # no. of unique children in common between HF and ULF neuroimaging datasets with maternal anaemia data
  2167. # filter dataframes to keep only rows with non-missing maternal_preg_anaemia_status
  2168. HF_filtered <- HF_ida_final[!is.na(HF_ida_final$maternal_preg_anaemia_status), ]
  2169. ULF_filtered <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
  2170. # find overlapping subject_ids among the filtered sets
  2171. overlap_subjects <- intersect(HF_filtered$subject_id, ULF_filtered$subject_id)
  2172. # count the number of overlapping subjects
  2173. num_overlap <- length(overlap_subjects)
  2174. # print the result
  2175. num_overlap
  2176. ```
  2177. ------------------------------------------------------------------------
  2178. **Primary Analysis: Antenatal Maternal Anaemia Status (ULF Sample)**
  2179. ------------------------------------------------------------------------
  2180. **Sample Size: For infants with Antenatal Maternal Anemia Data**
  2181. Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
  2182. ```{r}
  2183. # total sample size
  2184. # including repeated measures
  2185. ULF_ida_final %>%
  2186. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include PIDs with antenatal maternal anaemia data
  2187. summarise(n = n())
  2188. ```
  2189. ```{r}
  2190. # repeated measures excluded
  2191. # unique subject IDs: based on true prevalence
  2192. library(dplyr)
  2193. ULF_ida_final %>%
  2194. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2195. summarise(unique_subjects = n_distinct(subject_id))
  2196. ```
  2197. ```{r}
  2198. # prevalence of antenatal maternal anaemia
  2199. # unique subject IDs: based on true prevalence
  2200. ULF_ida_final %>%
  2201. filter(!is.na(maternal_preg_anaemia_status)) %>% # Exclude missing data
  2202. group_by(subject_id) %>%
  2203. summarise(
  2204. maternal_preg_anaemia_status = if_else(
  2205. any(maternal_preg_anaemia_status == "mat_anaemia", na.rm = TRUE),
  2206. "mat_anaemia",
  2207. "no_mat_anaemia"
  2208. )
  2209. ) %>%
  2210. count(maternal_preg_anaemia_status) %>%
  2211. mutate(
  2212. prevalence_percent = n / sum(n) * 100
  2213. )
  2214. ```
  2215. ```{r}
  2216. # antenatal maternal anaemia severity
  2217. # unique subject IDs: based on true prevalence
  2218. library(dplyr)
  2219. ULF_ida_final %>%
  2220. filter(!is.na(maternal_preg_anaemia_severity)) %>% # exclude missing severity
  2221. group_by(subject_id) %>% # group by mother
  2222. summarise(
  2223. maternal_preg_anaemia_severity = first(
  2224. maternal_preg_anaemia_severity[maternal_preg_anaemia_severity != ""]
  2225. )
  2226. ) %>%
  2227. group_by(maternal_preg_anaemia_severity) %>% # count unique mothers by severity
  2228. summarise(n = n()) %>%
  2229. mutate(percent = round(100 * n / sum(n), 2))
  2230. ```
  2231. **Data Availability with Repeated Measures**
  2232. ```{r}
  2233. # repeated measures count
  2234. # unique ID
  2235. library(dplyr)
  2236. ULF_ida_final %>%
  2237. filter(!is.na(maternal_preg_anaemia_status)) %>% # keep only rows with anaemia status
  2238. count(subject_id) %>% # count rows per subject
  2239. count(n) %>% # count how many subjects had n rows
  2240. rename(times_measured = n, number_of_subjects = nn) %>%
  2241. mutate(percent = round(100 * number_of_subjects / sum(number_of_subjects), 1)) # Add % column
  2242. ```
  2243. ```{r}
  2244. # figure 2
  2245. library(ggplot2)
  2246. library(dplyr)
  2247. Figure_2_ULF <-ULF_ida_final %>%
  2248. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2249. mutate(
  2250. timepoint = factor(timepoint, levels = c("3M", "6M", "12M", "18M", "24M")), # enforce order
  2251. subject_id = factor(subject_id),
  2252. anaemia_group = as.factor(maternal_preg_anaemia_status)
  2253. ) %>%
  2254. ggplot(aes(x = timepoint, y = subject_id, group = subject_id, color = anaemia_group)) +
  2255. geom_line(alpha = 0.5, linewidth = 0.5) + # use linewidth instead of deprecated size
  2256. geom_point(size = 2) +
  2257. scale_color_manual(
  2258. values = c("red", "blue"),
  2259. labels = c("Maternal Anaemia", "No Maternal Anaemia")
  2260. ) +
  2261. labs(
  2262. title = "Participant Data Availability by Timepoint",
  2263. x = "Timepoint",
  2264. y = "Participants",
  2265. color = "Maternal Anaemia"
  2266. ) +
  2267. theme_minimal(base_size = 13) +
  2268. theme(
  2269. axis.text.y = element_blank(),
  2270. axis.ticks.y = element_blank(),
  2271. panel.grid.major.y = element_blank()
  2272. ) +
  2273. # <-- add padding so points are not cut off
  2274. scale_y_discrete(expand = expansion(mult = c(0.02, 0.02))) +
  2275. scale_x_discrete(expand = expansion(mult = c(0.02, 0.02)))
  2276. print(Figure_2_ULF)
  2277. ```
  2278. **Sample Characteristics and Group Differences: Antenatal Maternal Anaemia Status**
  2279. Sample characteristics were assessed based on the number of unique subjects in each subsample, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
  2280. However, given the repeated measures study design with individual infant data across multiple study visits in both HF and ULF subsamples, linear mixed effects (LME) models were fitted to account for individual differences in assessing the association between antenatal maternal anaemia status and absolute regional brain volume trajectories
  2281. - **Categorical Variables**
  2282. *Trimester of Maternal Hb Measurement*
  2283. ```{r}
  2284. # trimester Hb was collected from mothers
  2285. # excluding repeated measures
  2286. # unique subject IDs: based on true prevalence
  2287. library(dplyr)
  2288. ULF_ida_final %>%
  2289. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia
  2290. group_by(subject_id) %>% # keep unique mothers only
  2291. summarise(
  2292. maternal_trimester_hb = first(maternal_trimester_hb) # assume first record per mother
  2293. ) %>%
  2294. group_by(maternal_trimester_hb) %>%
  2295. summarise(n = n()) %>%
  2296. mutate(percent = round(100 * n / sum(n), 2))
  2297. ```
  2298. ```{r}
  2299. # group differences in trimester of Hb collection
  2300. # Chi-squared: trimester Hb
  2301. # unique subject IDs: based on true prevalence
  2302. library(dplyr)
  2303. # run contingency table + test without creating a new dataset
  2304. tbl <- ULF_ida_final %>%
  2305. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia status
  2306. group_by(subject_id) %>% # ensure one row per mother
  2307. summarise(
  2308. maternal_trimester_hb = first(maternal_trimester_hb),
  2309. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2310. ) %>%
  2311. {table(.$maternal_trimester_hb, .$maternal_preg_anaemia_status)} # build contingency table inline
  2312. # print contingency table
  2313. cat("Contingency Table: Trimester by Maternal Anaemia Status\n")
  2314. print(format(round(tbl, 2), nsmall = 2))
  2315. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2316. print(round(100 * prop.table(tbl, margin = 2), 2))
  2317. # run appropriate test
  2318. chisq_result <- suppressWarnings(chisq.test(tbl))
  2319. expected <- chisq_result$expected
  2320. if (any(expected < 5)) {
  2321. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2322. print(fisher.test(tbl))
  2323. } else {
  2324. cat("\nChi-squared Test:\n")
  2325. print(chisq_result)
  2326. }
  2327. ```
  2328. *Child Sex at Birth*
  2329. ```{r}
  2330. # child sex at birth (full sample)
  2331. # unique subject IDs: based on true prevalence
  2332. library(dplyr)
  2333. ULF_ida_final %>%
  2334. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2335. group_by(subject_id) %>% # ensure one record per child
  2336. summarise(
  2337. child_sex = first(child_sex) # assume sex is consistent per child
  2338. ) %>%
  2339. group_by(child_sex) %>%
  2340. summarise(n = n()) %>%
  2341. mutate(percent = round(100 * n / sum(n), 2))
  2342. ```
  2343. ```{r}
  2344. # group differences in child sex by antenatal maternal anaemia status
  2345. # chi-squared: child sex
  2346. # unique subject IDs: based on true prevalence
  2347. library(dplyr)
  2348. # run contingency table + test without creating a new dataset
  2349. tbl <- ULF_ida_final %>%
  2350. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2351. group_by(subject_id) %>% # ensure one record per child
  2352. summarise(
  2353. child_sex = first(child_sex),
  2354. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2355. ) %>%
  2356. {table(.$child_sex, .$maternal_preg_anaemia_status)} # build contingency table inline
  2357. # print contingency table
  2358. cat("Contingency Table: Child Sex by Maternal Anaemia Status\n")
  2359. print(format(round(tbl, 2), nsmall = 2))
  2360. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2361. print(round(100 * prop.table(tbl, margin = 2), 2))
  2362. # run appropriate test
  2363. chisq_result <- suppressWarnings(chisq.test(tbl))
  2364. expected <- chisq_result$expected
  2365. if (any(expected < 5)) {
  2366. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2367. print(fisher.test(tbl))
  2368. } else {
  2369. cat("\nChi-squared Test:\n")
  2370. print(chisq_result)
  2371. }
  2372. ```
  2373. *Prenatal Alcohol Exposure (PAE)*
  2374. ```{r}
  2375. # pae prevalence (full sample)
  2376. # unique subject IDs: based on true prevalence
  2377. library(dplyr)
  2378. ULF_ida_final %>%
  2379. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2380. group_by(subject_id) %>% # ensure one record per mother
  2381. summarise(
  2382. maternal_pae = first(maternal_pae) # assume pae status is consistent per mother
  2383. ) %>%
  2384. group_by(maternal_pae) %>%
  2385. summarise(n = n()) %>%
  2386. mutate(percent = round(100 * n / sum(n), 2))
  2387. ```
  2388. ```{r}
  2389. # group differences in pae by antenatal maternal anaemia status
  2390. # chi-squared: pae
  2391. # unique subject IDs: based on true prevalence
  2392. library(dplyr)
  2393. # build contingency table and run test inline (unique mothers only)
  2394. tbl <- ULF_ida_final %>%
  2395. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2396. group_by(subject_id) %>% # ensure one record per mother
  2397. summarise(
  2398. maternal_pae = first(maternal_pae),
  2399. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2400. ) %>%
  2401. {table(.$maternal_pae, .$maternal_preg_anaemia_status)} # build contingency table inline
  2402. # print contingency table
  2403. cat("Contingency Table: Maternal PAE by Maternal Anaemia Status\n")
  2404. print(tbl)
  2405. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2406. print(round(100 * prop.table(tbl, margin = 2), 2))
  2407. # run appropriate test
  2408. chisq_result <- suppressWarnings(chisq.test(tbl))
  2409. expected <- chisq_result$expected
  2410. if (any(expected < 5)) {
  2411. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2412. print(fisher.test(tbl))
  2413. } else {
  2414. cat("\nChi-squared Test:\n")
  2415. print(chisq_result)
  2416. }
  2417. ```
  2418. *Prenatal Tobacco Exposure (PTE)*
  2419. ```{r}
  2420. # pte prevalence (full sample)
  2421. # unique subject IDs: based on true prevalence
  2422. library(dplyr)
  2423. ULF_ida_final %>%
  2424. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2425. group_by(subject_id) %>% # ensure one record per mother
  2426. summarise(
  2427. maternal_pte = first(maternal_pte) # assume pte status is consistent per mother
  2428. ) %>%
  2429. group_by(maternal_pte) %>%
  2430. summarise(n = n()) %>%
  2431. mutate(percent = round(100 * n / sum(n), 2))
  2432. ```
  2433. ```{r}
  2434. # group differences in pte by antenatal maternal anaemia status
  2435. # chi-squared: pte
  2436. # unique subject IDs: based on true prevalence
  2437. library(dplyr)
  2438. # build contingency table and run test inline (unique mothers only)
  2439. tbl <- ULF_ida_final %>%
  2440. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2441. group_by(subject_id) %>% # ensure one record per mother
  2442. summarise(
  2443. maternal_pte = first(maternal_pte),
  2444. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2445. ) %>%
  2446. {table(.$maternal_pte, .$maternal_preg_anaemia_status)} # build contingency table inline
  2447. # print contingency table
  2448. cat("Contingency Table: Maternal PTE by Maternal Anaemia Status\n")
  2449. print(tbl)
  2450. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2451. print(round(100 * prop.table(tbl, margin = 2), 2))
  2452. # run appropriate test
  2453. chisq_result <- suppressWarnings(chisq.test(tbl))
  2454. expected <- chisq_result$expected
  2455. if (any(expected < 5)) {
  2456. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2457. print(fisher.test(tbl))
  2458. } else {
  2459. cat("\nChi-squared Test:\n")
  2460. print(chisq_result)
  2461. }
  2462. ```
  2463. *Maternal HIV Infection*
  2464. ```{r}
  2465. # materal hiv prevalence (full sample)
  2466. # unique subject IDs: based on true prevalence
  2467. ULF_ida_final %>%
  2468. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2469. group_by(subject_id) %>% # ensure one record per mother
  2470. summarise(
  2471. maternal_hiv = first(maternal_hiv) # assume HIV status is consistent per mother
  2472. ) %>%
  2473. group_by(maternal_hiv) %>%
  2474. summarise(n = n()) %>%
  2475. mutate(percent = round(100 * n / sum(n), 2))
  2476. ```
  2477. ```{r}
  2478. # group differences in maternal hiv by antenatal maternal anaemia status
  2479. # chi-squared: hiv
  2480. # unique subject IDs: based on true prevalence
  2481. library(dplyr)
  2482. # build contingency table and run test inline (unique mothers only)
  2483. tbl <- ULF_ida_final %>%
  2484. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2485. group_by(subject_id) %>% # ensure one record per mother
  2486. summarise(
  2487. maternal_hiv = first(maternal_hiv),
  2488. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2489. ) %>%
  2490. {table(.$maternal_hiv, .$maternal_preg_anaemia_status)} # build contingency table inline
  2491. # print contingency table
  2492. cat("Contingency Table: Maternal HIV by Maternal Anaemia Status\n")
  2493. print(tbl)
  2494. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2495. print(round(100 * prop.table(tbl, margin = 2), 2))
  2496. # run appropriate test
  2497. chisq_result <- suppressWarnings(chisq.test(tbl))
  2498. expected <- chisq_result$expected
  2499. if (any(expected < 5)) {
  2500. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2501. print(fisher.test(tbl))
  2502. } else {
  2503. cat("\nChi-squared Test:\n")
  2504. print(chisq_result)
  2505. }
  2506. ```
  2507. *Maternal Depression*
  2508. ```{r}
  2509. # maternal depression prevalence (full sample)
  2510. # unique subject IDs: based on true prevalence
  2511. ULF_ida_final %>%
  2512. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2513. distinct(subject_id, .keep_all = TRUE) %>%
  2514. group_by(maternal_dep) %>%
  2515. summarise(n = n()) %>%
  2516. mutate(percent = round(100 * n / sum(n), 2))
  2517. ```
  2518. ```{r}
  2519. # group differenced by maternal depression status
  2520. # chi-squared: maternal depression
  2521. # unique subject IDs: based on true prevalence
  2522. library(dplyr)
  2523. # build contingency table and run test inline (unique mothers only)
  2524. tbl <- ULF_ida_final %>%
  2525. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia data
  2526. group_by(subject_id) %>% # ensure one record per mother
  2527. summarise(
  2528. maternal_dep = first(maternal_dep),
  2529. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)
  2530. ) %>%
  2531. {table(.$maternal_dep, .$maternal_preg_anaemia_status)} # build contingency table inline
  2532. # print contingency table
  2533. cat("Contingency Table: Maternal Depression by Maternal Anaemia Status\n")
  2534. print(tbl)
  2535. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2536. print(round(100 * prop.table(tbl, margin = 2), 2))
  2537. # run appropriate test
  2538. chisq_result <- suppressWarnings(chisq.test(tbl))
  2539. expected <- chisq_result$expected
  2540. if (any(expected < 5)) {
  2541. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2542. print(fisher.test(tbl))
  2543. } else {
  2544. cat("\nChi-squared Test:\n")
  2545. print(chisq_result)
  2546. }
  2547. ```
  2548. *Maternal Employment*
  2549. ```{r}
  2550. # maternal employment (full sample)
  2551. # unique subject IDs: based on true prevalence
  2552. library(dplyr)
  2553. ULF_ida_final %>%
  2554. filter(!is.na(maternal_preg_anaemia_status)) %>% # exclude missing anaemia status
  2555. group_by(subject_id) %>% # ensure one record per mother
  2556. summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous)) %>%
  2557. group_by(mom_employment_en_dichotomous) %>%
  2558. summarise(n = n()) %>%
  2559. mutate(percent = round(100 * n / sum(n), 2))
  2560. ```
  2561. ```{r}
  2562. # group differences in maternal employment by antenatal maternal anaemia status
  2563. # chi-squared: maternal employment
  2564. # unique subject IDs: based on true prevalence
  2565. # create contingency table inline using first record per mother
  2566. tbl <- table(
  2567. ULF_ida_final %>%
  2568. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2569. group_by(subject_id) %>%
  2570. summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous),
  2571. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
  2572. pull(mom_employment_en_dichotomous),
  2573. ULF_ida_final %>%
  2574. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2575. group_by(subject_id) %>%
  2576. summarise(mom_employment_en_dichotomous = first(mom_employment_en_dichotomous),
  2577. maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
  2578. pull(maternal_preg_anaemia_status)
  2579. )
  2580. # print contingency table
  2581. cat("Contingency Table: Maternal Employment by Maternal Anaemia Status\n")
  2582. print(tbl)
  2583. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2584. print(round(100 * prop.table(tbl, margin = 2), 2))
  2585. # run appropriate test
  2586. expected <- chisq.test(tbl)$expected
  2587. if (any(expected < 5)) {
  2588. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2589. print(fisher.test(tbl))
  2590. } else {
  2591. cat("\nChi-squared Test:\n")
  2592. print(chisq.test(tbl))
  2593. }
  2594. ```
  2595. *Maternal Education*
  2596. ```{r}
  2597. # maternal education (full sample)
  2598. # unique subject IDs: based on true prevalence
  2599. ULF_ida_final %>%
  2600. filter(!is.na(maternal_preg_anaemia_status)) %>% # Include only mothers with anaemia data
  2601. group_by(subject_id) %>% # Ensure one record per mother
  2602. summarise(mom_edu_en = first(mom_edu_en)) %>% # Take first value per mother
  2603. group_by(mom_edu_en) %>%
  2604. summarise(n = n()) %>%
  2605. mutate(percent = round(100 * n / sum(n), 2))
  2606. ```
  2607. ```{r}
  2608. # group differences in maternal education by antenatal maternal anaemia status
  2609. # chi-squared: maternal education
  2610. # unique subject IDs: based on true prevalence
  2611. # create contingency table inline, one record per mother
  2612. tbl <- table(
  2613. ULF_ida_final %>%
  2614. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2615. group_by(subject_id) %>% # Ensure unique mother
  2616. summarise(mom_edu_en = first(mom_edu_en)) %>%
  2617. pull(mom_edu_en),
  2618. ULF_ida_final %>%
  2619. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2620. group_by(subject_id) %>%
  2621. summarise(maternal_preg_anaemia_status = first(maternal_preg_anaemia_status)) %>%
  2622. pull(maternal_preg_anaemia_status)
  2623. )
  2624. # print contingency table
  2625. cat("Contingency Table: Maternal Education by Maternal Anaemia Status\n")
  2626. print(tbl)
  2627. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2628. print(round(100 * prop.table(tbl, margin = 2), 2))
  2629. # run appropriate test
  2630. expected <- chisq.test(tbl)$expected
  2631. if (any(expected < 5)) {
  2632. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2633. print(fisher.test(tbl))
  2634. } else {
  2635. cat("\nChi-squared Test:\n")
  2636. print(chisq.test(tbl))
  2637. }
  2638. ```
  2639. *Household Income*
  2640. ```{r}
  2641. # household income (full sample)
  2642. # unique subject IDs: based on true prevalence
  2643. ULF_ida_final %>%
  2644. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include mothers with anaemia data
  2645. group_by(subject_id) %>% # ensure each mother counted once
  2646. summarise(income_household_en = first(income_household_en)) %>%
  2647. group_by(income_household_en) %>%
  2648. summarise(n = n()) %>%
  2649. mutate(percent = round(100 * n / sum(n), 2))
  2650. ```
  2651. ```{r}
  2652. # group differences in household income by antenatal maternal anaemia status
  2653. # chi-squared: household income
  2654. # unique subject IDs: based on true prevalence
  2655. # create contingency table counting each subject only once
  2656. tbl <- table(
  2657. ULF_ida_final$income_household_en[!duplicated(ULF_ida_final$subject_id) &
  2658. !is.na(ULF_ida_final$maternal_preg_anaemia_status)],
  2659. ULF_ida_final$maternal_preg_anaemia_status[!duplicated(ULF_ida_final$subject_id) &
  2660. !is.na(ULF_ida_final$maternal_preg_anaemia_status)]
  2661. )
  2662. # print contingency table
  2663. cat("Contingency Table: Household Income by Maternal Anaemia Status\n")
  2664. print(tbl)
  2665. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2666. print(round(100 * prop.table(tbl, margin = 2), 2))
  2667. # run appropriate test
  2668. expected <- chisq.test(tbl)$expected
  2669. if (any(expected < 5)) {
  2670. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2671. print(fisher.test(tbl))
  2672. } else {
  2673. cat("\nChi-squared Test:\n")
  2674. print(chisq.test(tbl))
  2675. }
  2676. ```
  2677. *Infant HIV Infection*
  2678. ```{r}
  2679. # infant hiv prevalence (full sample)
  2680. # unique subject IDs: based on true prevalence
  2681. ULF_ida_final %>%
  2682. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2683. filter(!duplicated(subject_id)) %>% # count each subject once
  2684. group_by(baby_hiv) %>%
  2685. summarise(n = n()) %>%
  2686. mutate(percent = round(100 * n / sum(n), 2))
  2687. ```
  2688. ```{r}
  2689. # group differences in infant hiv by antenatal maternal anaemia status
  2690. # chi-squared: infant hiv
  2691. # unique subject IDs: based on true prevalence
  2692. # create contingency table counting each subject only once
  2693. tbl <- table(
  2694. ULF_ida_final$baby_hiv[!is.na(ULF_ida_final$maternal_preg_anaemia_status) & !duplicated(ULF_ida_final$subject_id)],
  2695. ULF_ida_final$maternal_preg_anaemia_status[!is.na(ULF_ida_final$maternal_preg_anaemia_status) & !duplicated(ULF_ida_final$subject_id)]
  2696. )
  2697. # print contingency table
  2698. cat("Contingency Table: Infant HIV by Maternal Anaemia Status\n")
  2699. print(tbl)
  2700. cat("\nColumn Percentages (% of each anaemia status group):\n")
  2701. print(round(100 * prop.table(tbl, margin = 2), 2))
  2702. # run appropriate test
  2703. expected <- chisq.test(tbl)$expected
  2704. if (any(expected < 5)) {
  2705. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  2706. print(fisher.test(tbl))
  2707. } else {
  2708. cat("\nChi-squared Test:\n")
  2709. print(chisq.test(tbl))
  2710. }
  2711. ```
  2712. - **Continuous Variables**
  2713. ```{r}
  2714. #install packages if needed
  2715. #install.packages('car')
  2716. library(car)
  2717. ```
  2718. *Maternal Age at Enrolment*
  2719. ```{r}
  2720. # maternal age at enrolment (full sample)
  2721. # unique subject IDs: based on true prevalence
  2722. library(dplyr)
  2723. ULF_ida_final %>%
  2724. filter(!is.na(maternal_preg_anaemia_status),
  2725. !is.na(mom_age_en)) %>% # exclude missing anaemia or age
  2726. group_by(subject_id) %>% # keep unique mothers only
  2727. summarise(
  2728. mom_age_en = first(mom_age_en) # assume maternal age at enrolment is constant
  2729. ) %>%
  2730. summarise(
  2731. mean_mom_age = mean(mom_age_en, na.rm = TRUE),
  2732. sd_mom_age = sd(mom_age_en, na.rm = TRUE),
  2733. min_mom_age = min(mom_age_en, na.rm = TRUE),
  2734. max_mom_age = max(mom_age_en, na.rm = TRUE),
  2735. n = n()
  2736. )
  2737. ```
  2738. ```{r}
  2739. # group differences in maternal age at enrolment by antenatal maternal anaemia status
  2740. # ttest: maternal age at enrolment
  2741. # unique subject IDs: based on true prevalence
  2742. library(dplyr)
  2743. library(car)
  2744. # keep only unique subjects with non-missing anaemia status
  2745. unique_data <- ULF_ida_final %>%
  2746. filter(!is.na(maternal_preg_anaemia_status),
  2747. !is.na(mom_age_en)) %>%
  2748. distinct(subject_id, .keep_all = TRUE)
  2749. # summary stats (force exactly 2 decimals for display)
  2750. unique_data %>%
  2751. group_by(maternal_preg_anaemia_status) %>%
  2752. summarise(
  2753. n = n(),
  2754. mean_age = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  2755. sd_age = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  2756. min_age = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  2757. max_age = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  2758. .groups = "drop"
  2759. ) %>%
  2760. print()
  2761. # levene's test for homogeneity of variance
  2762. levene_result <- leveneTest(
  2763. mom_age_en ~ maternal_preg_anaemia_status,
  2764. data = unique_data
  2765. )
  2766. print(levene_result)
  2767. # extract p-value from Levene’s Test
  2768. levene_p <- levene_result$`Pr(>F)`[1]
  2769. # decide automatically which t-test to use
  2770. if (levene_p > 0.05) {
  2771. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  2772. ") → Using Student's t-test (equal variances assumed)\n\n")
  2773. t_result <- t.test(
  2774. mom_age_en ~ maternal_preg_anaemia_status,
  2775. data = unique_data,
  2776. var.equal = TRUE
  2777. )
  2778. } else {
  2779. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  2780. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  2781. t_result <- t.test(
  2782. mom_age_en ~ maternal_preg_anaemia_status,
  2783. data = unique_data,
  2784. var.equal = FALSE
  2785. )
  2786. }
  2787. # print t-test results (full precision)
  2788. print(t_result)
  2789. ```
  2790. *Child Age at Scan:*
  2791. Include repeated measures to account for different ages for scans acquired across study timepoints from the same infant
  2792. ```{r}
  2793. # child age at scan for full group
  2794. # include repeated measures as child is a different age at every scan acquired (full sample)
  2795. ULF_ida_final %>%
  2796. filter(!is.na(maternal_preg_anaemia_status)) %>% #only include PIDs with antenatal maternal anaemia data
  2797. summarise(
  2798. mean_child_age = mean(scan_age_months, na.rm = TRUE),
  2799. sd_child_age = sd(scan_age_months, na.rm = TRUE),
  2800. min_child_age = min(scan_age_months, na.rm = TRUE),
  2801. max_child_age = max(scan_age_months, na.rm = TRUE),
  2802. n = n()
  2803. )
  2804. ```
  2805. ```{r}
  2806. # child age at scan
  2807. # full sample including repeated measures
  2808. # t-test: child age at scan
  2809. library(dplyr)
  2810. library(car)
  2811. # keep only unique subjects with non-missing anaemia status and scan age
  2812. unique_data <- ULF_ida_final %>%
  2813. filter(!is.na(maternal_preg_anaemia_status),
  2814. !is.na(scan_age_months))
  2815. # summary stats (force exactly 2 decimals)
  2816. unique_data %>%
  2817. group_by(maternal_preg_anaemia_status) %>%
  2818. summarise(
  2819. n = n(),
  2820. mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  2821. sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  2822. min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  2823. max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  2824. .groups = "drop"
  2825. ) %>%
  2826. print()
  2827. # levene's test for homogeneity of variance
  2828. levene_result <- leveneTest(
  2829. scan_age_months ~ maternal_preg_anaemia_status,
  2830. data = unique_data
  2831. )
  2832. print(levene_result)
  2833. # extract p-value from Levene’s Test
  2834. levene_p <- levene_result$`Pr(>F)`[1]
  2835. # decide automatically which t-test to use
  2836. if (levene_p > 0.05) {
  2837. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  2838. ") → Using Student's t-test (equal variances assumed)\n\n")
  2839. t_result <- t.test(
  2840. scan_age_months ~ maternal_preg_anaemia_status,
  2841. data = unique_data,
  2842. var.equal = TRUE
  2843. )
  2844. } else {
  2845. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  2846. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  2847. t_result <- t.test(
  2848. scan_age_months ~ maternal_preg_anaemia_status,
  2849. data = unique_data,
  2850. var.equal = FALSE
  2851. )
  2852. }
  2853. # print t-test results (full precision)
  2854. print(t_result)
  2855. ```
  2856. *Child Gestational Age (GA) at Birth*
  2857. ```{r}
  2858. # child ga at birth for full group
  2859. # unique subject IDs: based on true prevalence
  2860. library(dplyr)
  2861. ULF_ida_final %>%
  2862. filter(!is.na(maternal_preg_anaemia_status)) %>%
  2863. distinct(subject_id, .keep_all = TRUE) %>% # keep only one row per child
  2864. summarise(
  2865. mean_child_ga = mean(ga_weeks, na.rm = TRUE),
  2866. sd_child_ga = sd(ga_weeks, na.rm = TRUE),
  2867. min_child_ga = min(ga_weeks, na.rm = TRUE),
  2868. max_child_ga = max(ga_weeks, na.rm = TRUE),
  2869. n = n()
  2870. )
  2871. ```
  2872. ```{r}
  2873. # group differences in child ga at birth
  2874. # ttest: child ga at birth
  2875. # unique subject IDs: based on true prevalence
  2876. library(dplyr)
  2877. library(car)
  2878. # keep only unique subjects with non-missing maternal anaemia status and GA
  2879. unique_data <- ULF_ida_final %>%
  2880. filter(!is.na(maternal_preg_anaemia_status),
  2881. !is.na(ga_weeks)) %>%
  2882. distinct(subject_id, .keep_all = TRUE)
  2883. # summary stats by maternal anaemia status (force 2 decimals)
  2884. unique_data %>%
  2885. group_by(maternal_preg_anaemia_status) %>%
  2886. summarise(
  2887. n = n(),
  2888. mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  2889. sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  2890. min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  2891. max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  2892. .groups = "drop"
  2893. ) %>%
  2894. print()
  2895. # levene's test for homogeneity of variance
  2896. levene_result <- leveneTest(
  2897. ga_weeks ~ maternal_preg_anaemia_status,
  2898. data = unique_data
  2899. )
  2900. print(levene_result)
  2901. # extract p-value from Levene’s Test
  2902. levene_p <- levene_result$`Pr(>F)`[1]
  2903. # Decide automatically which t-test to use
  2904. if (levene_p > 0.05) {
  2905. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  2906. ") → Using Student's t-test (equal variances assumed)\n\n")
  2907. t_result <- t.test(
  2908. ga_weeks ~ maternal_preg_anaemia_status,
  2909. data = unique_data,
  2910. var.equal = TRUE
  2911. )
  2912. } else {
  2913. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  2914. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  2915. t_result <- t.test(
  2916. ga_weeks ~ maternal_preg_anaemia_status,
  2917. data = unique_data,
  2918. var.equal = FALSE
  2919. )
  2920. }
  2921. # print t-test results (full precision)
  2922. print(t_result)
  2923. ```
  2924. *Child Birth Weight (BW)*
  2925. ```{r}
  2926. # child bw (full sample)
  2927. # unique subject IDs: based on true prevalence
  2928. ULF_ida_final %>%
  2929. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with maternal anaemia data
  2930. distinct(subject_id, .keep_all = TRUE) %>% # ensure one record per child
  2931. summarise(
  2932. mean_bw = mean(bw, na.rm = TRUE),
  2933. sd_bw = sd(bw, na.rm = TRUE),
  2934. min_bw = min(bw, na.rm = TRUE),
  2935. max_bw = max(bw, na.rm = TRUE),
  2936. n = n()
  2937. )
  2938. ```
  2939. ```{r}
  2940. # group differences in child bw
  2941. # ttest: child bw
  2942. # exclude repeated measures:unique ID
  2943. library(dplyr)
  2944. library(car)
  2945. # keep only unique subjects with non-missing maternal anaemia status and birthweight
  2946. unique_data <- ULF_ida_final %>%
  2947. filter(!is.na(maternal_preg_anaemia_status),
  2948. !is.na(bw)) %>%
  2949. distinct(subject_id, .keep_all = TRUE)
  2950. # summary stats by maternal anaemia status (force 2 decimals)
  2951. unique_data %>%
  2952. group_by(maternal_preg_anaemia_status) %>%
  2953. summarise(
  2954. n = n(),
  2955. mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
  2956. sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
  2957. min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
  2958. max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
  2959. .groups = "drop"
  2960. ) %>%
  2961. print()
  2962. # levene's test for homogeneity of variance
  2963. levene_result <- leveneTest(
  2964. bw ~ maternal_preg_anaemia_status,
  2965. data = unique_data
  2966. )
  2967. print(levene_result)
  2968. # extract p-value from Levene’s Test
  2969. levene_p <- levene_result$`Pr(>F)`[1]
  2970. # decide automatically which t-test to use
  2971. if (levene_p > 0.05) {
  2972. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  2973. ") → Using Student's t-test (equal variances assumed)\n\n")
  2974. t_result <- t.test(
  2975. bw ~ maternal_preg_anaemia_status,
  2976. data = unique_data,
  2977. var.equal = TRUE
  2978. )
  2979. } else {
  2980. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  2981. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  2982. t_result <- t.test(
  2983. bw ~ maternal_preg_anaemia_status,
  2984. data = unique_data,
  2985. var.equal = FALSE
  2986. )
  2987. }
  2988. # print t-test results (full precision)
  2989. print(t_result)
  2990. ```
  2991. *Child Birth Length (BL)*
  2992. ```{r}
  2993. # child bl (full sample)
  2994. # unique subject IDs: based on true prevalence
  2995. ULF_ida_final %>%
  2996. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with anaemia data
  2997. distinct(subject_id, .keep_all = TRUE) %>% # ensure only one record per child
  2998. summarise(
  2999. mean_birthlength = mean(birth_length, na.rm = TRUE),
  3000. sd_birthlength = sd(birth_length, na.rm = TRUE),
  3001. min_birthlength = min(birth_length, na.rm = TRUE),
  3002. max_birthlength = max(birth_length, na.rm = TRUE),
  3003. n = n()
  3004. )
  3005. ```
  3006. ```{r}
  3007. # child bl
  3008. # unique subject IDs: based on true prevalence
  3009. # t-test: child bl
  3010. library(dplyr)
  3011. library(car)
  3012. # keep only unique subjects with non-missing maternal anaemia status and birth length
  3013. unique_data <- ULF_ida_final %>%
  3014. filter(!is.na(maternal_preg_anaemia_status),
  3015. !is.na(birth_length)) %>%
  3016. distinct(subject_id, .keep_all = TRUE)
  3017. # summary stats by maternal anaemia status (force 2 decimals)
  3018. unique_data %>%
  3019. group_by(maternal_preg_anaemia_status) %>%
  3020. summarise(
  3021. n = n(),
  3022. mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
  3023. sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
  3024. min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
  3025. max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
  3026. .groups = "drop"
  3027. ) %>%
  3028. print()
  3029. # levene's test for homogeneity of variance
  3030. levene_result <- leveneTest(
  3031. birth_length ~ maternal_preg_anaemia_status,
  3032. data = unique_data
  3033. )
  3034. print(levene_result)
  3035. # extract p-value from Levene’s Test
  3036. levene_p <- levene_result$`Pr(>F)`[1]
  3037. # decide automatically which t-test to use
  3038. if (levene_p > 0.05) {
  3039. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  3040. ") → Using Student's t-test (equal variances assumed)\n\n")
  3041. t_result <- t.test(
  3042. birth_length ~ maternal_preg_anaemia_status,
  3043. data = unique_data,
  3044. var.equal = TRUE
  3045. )
  3046. } else {
  3047. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  3048. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3049. t_result <- t.test(
  3050. birth_length ~ maternal_preg_anaemia_status,
  3051. data = unique_data,
  3052. var.equal = FALSE
  3053. )
  3054. }
  3055. # print t-test results (full precision)
  3056. print(t_result)
  3057. ```
  3058. *Child Birth Head Circumference (HC)*
  3059. ```{r}
  3060. # child birth hc (full sample)
  3061. # unique subject IDs: based on true prevalence
  3062. ULF_ida_final %>%
  3063. filter(!is.na(maternal_preg_anaemia_status)) %>% # only include records with anaemia data
  3064. distinct(subject_id, .keep_all = TRUE) %>% # keep only one record per unique child
  3065. summarise(
  3066. mean_hc = mean(birth_hc, na.rm = TRUE),
  3067. sd_hc = sd(birth_hc, na.rm = TRUE),
  3068. min_hc = min(birth_hc, na.rm = TRUE),
  3069. max_hc = max(birth_hc, na.rm = TRUE),
  3070. n = n()
  3071. )
  3072. ```
  3073. ```{r}
  3074. # group differences in birth hc by antenatal maternal anaemia status
  3075. # t-test: child hc at birth
  3076. # unique subject IDs: based on true prevalence
  3077. library(dplyr)
  3078. library(car)
  3079. # keep only unique subjects with non-missing maternal anaemia status and birth HC
  3080. unique_data <- ULF_ida_final %>%
  3081. filter(!is.na(maternal_preg_anaemia_status),
  3082. !is.na(birth_hc)) %>%
  3083. distinct(subject_id, .keep_all = TRUE)
  3084. # summary stats by maternal anaemia status (force 2 decimals)
  3085. unique_data %>%
  3086. group_by(maternal_preg_anaemia_status) %>%
  3087. summarise(
  3088. n = n(),
  3089. mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  3090. sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  3091. min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  3092. max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  3093. .groups = "drop"
  3094. ) %>%
  3095. print()
  3096. # levene's test for homogeneity of variance
  3097. levene_result <- leveneTest(
  3098. birth_hc ~ maternal_preg_anaemia_status,
  3099. data = unique_data
  3100. )
  3101. print(levene_result)
  3102. # extract p-value from Levene’s Test
  3103. levene_p <- levene_result$`Pr(>F)`[1]
  3104. # Decide automatically which t-test to use
  3105. if (levene_p > 0.05) {
  3106. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  3107. ") → Using Student's t-test (equal variances assumed)\n\n")
  3108. t_result <- t.test(
  3109. birth_hc ~ maternal_preg_anaemia_status,
  3110. data = unique_data,
  3111. var.equal = TRUE
  3112. )
  3113. } else {
  3114. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  3115. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3116. t_result <- t.test(
  3117. birth_hc ~ maternal_preg_anaemia_status,
  3118. data = unique_data,
  3119. var.equal = FALSE
  3120. )
  3121. }
  3122. # print t-test results (full precision)
  3123. print(t_result)
  3124. ```
  3125. *Gestational Age (GA) at Maternal Hb Measurement*
  3126. ```{r}
  3127. # ga at min Hb measurement (full sample)
  3128. # unique subject IDs: based on true prevalence
  3129. ULF_ida_final %>%
  3130. filter(!is.na(maternal_preg_anaemia_status),
  3131. !is.na(maternal_min_hb_ga_roundown)) %>%
  3132. distinct(subject_id, .keep_all = TRUE) %>%
  3133. summarise(
  3134. n = n(),
  3135. mean_ga_hb = round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  3136. sd_ga_hb = round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  3137. min_ga_hb = round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  3138. q1_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.25, na.rm = TRUE), 2),
  3139. median_ga_hb = round(median(maternal_min_hb_ga_roundown, na.rm = TRUE), 2),
  3140. q3_ga_hb = round(quantile(maternal_min_hb_ga_roundown, 0.75, na.rm = TRUE), 2),
  3141. max_ga_hb = round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2)
  3142. )
  3143. ```
  3144. ```{r}
  3145. # group differences in ga at minimum maternal Hb by antenatal maternal anaemia status
  3146. # t-test: ga at min Hb measurement
  3147. # unique subject IDs: based on true prevalence
  3148. library(dplyr)
  3149. library(car)
  3150. # keep only unique subjects with non-missing maternal anaemia status and GA at min Hb
  3151. unique_data <- ULF_ida_final %>%
  3152. filter(!is.na(maternal_preg_anaemia_status),
  3153. !is.na(maternal_min_hb_ga_roundown)) %>%
  3154. distinct(subject_id, .keep_all = TRUE)
  3155. # summary stats by maternal anaemia status (force 2 decimals)
  3156. unique_data %>%
  3157. group_by(maternal_preg_anaemia_status) %>%
  3158. summarise(
  3159. n = n(),
  3160. mean_ga_hb = format(round(mean(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  3161. sd_ga_hb = format(round(sd(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  3162. min_ga_hb = format(round(min(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  3163. max_ga_hb = format(round(max(maternal_min_hb_ga_roundown, na.rm = TRUE), 2), nsmall = 2),
  3164. .groups = "drop"
  3165. ) %>%
  3166. print()
  3167. # Levene's test for homogeneity of variance
  3168. levene_result <- leveneTest(
  3169. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  3170. data = unique_data
  3171. )
  3172. print(levene_result)
  3173. # extract p-value from Levene’s Test
  3174. levene_p <- levene_result$`Pr(>F)`[1]
  3175. # decide automatically which t-test to use
  3176. if (levene_p > 0.05) {
  3177. cat("\nLevene’s test not significant (p =", round(levene_p, 4),
  3178. ") → Using Student's t-test (equal variances assumed)\n\n")
  3179. t_result <- t.test(
  3180. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  3181. data = unique_data,
  3182. var.equal = TRUE
  3183. )
  3184. } else {
  3185. cat("\nLevene’s test significant (p =", round(levene_p, 4),
  3186. ") → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3187. t_result <- t.test(
  3188. maternal_min_hb_ga_roundown ~ maternal_preg_anaemia_status,
  3189. data = unique_data,
  3190. var.equal = FALSE
  3191. )
  3192. }
  3193. # print t-test results (full precision)
  3194. print(t_result)
  3195. ```
  3196. *Minimum Maternal Hb measurement in Trimester 1*
  3197. ```{r}
  3198. # maternal min Hb across groups in trimester 1: by antenatal maternal anaemia status
  3199. # unique subject IDs: based on true prevalence
  3200. library(dplyr)
  3201. library(car)
  3202. # filter for trimester 1, non-missing anaemia status and min Hb, keep unique subjects
  3203. trim1_data_unique <- ULF_ida_final %>%
  3204. filter(
  3205. maternal_trimester_hb == "first",
  3206. !is.na(maternal_preg_anaemia_status),
  3207. !is.na(maternal_preg_min_hb)
  3208. ) %>%
  3209. distinct(subject_id, .keep_all = TRUE)
  3210. # summary statistics by anaemia status with exactly 2 decimals displayed
  3211. summary_stats <- trim1_data_unique %>%
  3212. group_by(maternal_preg_anaemia_status) %>%
  3213. summarise(
  3214. n = n(), # ← missing comma fixed here
  3215. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3216. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3217. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3218. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3219. .groups = "drop"
  3220. )
  3221. # print summary stats
  3222. print(summary_stats)
  3223. # levene's test for homogeneity of variance
  3224. levene_result <- leveneTest(
  3225. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3226. data = trim1_data_unique
  3227. )
  3228. print(levene_result)
  3229. # extract Levene p-value (rounded for reporting only)
  3230. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  3231. cat("\nLevene’s test p-value =", levene_p, "\n")
  3232. # automatically choose correct t-test
  3233. if (levene_p > 0.05) {
  3234. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  3235. t_result <- t.test(
  3236. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3237. data = trim1_data_unique,
  3238. var.equal = TRUE
  3239. )
  3240. } else {
  3241. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3242. t_result <- t.test(
  3243. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3244. data = trim1_data_unique,
  3245. var.equal = FALSE
  3246. )
  3247. }
  3248. # print t-test result in full precision
  3249. print(t_result)
  3250. ```
  3251. *Minimum Maternal Hb measurement in Trimester 2*
  3252. ```{r}
  3253. # maternal minimum Hb in trimester 2 by antenatal anaemia status
  3254. # t-test
  3255. # unique subject IDs: based on true prevalence
  3256. library(dplyr)
  3257. library(car)
  3258. # filter for trimester 2, non-missing anaemia status and min Hb, keep unique subjects
  3259. trim2_data_unique <- ULF_ida_final %>%
  3260. filter(
  3261. maternal_trimester_hb == "second",
  3262. !is.na(maternal_preg_anaemia_status),
  3263. !is.na(maternal_preg_min_hb)
  3264. ) %>%
  3265. distinct(subject_id, .keep_all = TRUE)
  3266. # summary statistics by anaemia status with exactly 2 decimals displayed
  3267. summary_stats <- trim2_data_unique %>%
  3268. group_by(maternal_preg_anaemia_status) %>%
  3269. summarise(
  3270. n = n(),
  3271. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3272. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3273. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3274. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3275. .groups = "drop"
  3276. )
  3277. # print summary stats
  3278. print(summary_stats)
  3279. # levene's test for homogeneity of variance
  3280. levene_result <- leveneTest(
  3281. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3282. data = trim2_data_unique
  3283. )
  3284. print(levene_result)
  3285. # extract levene p-value (rounded for reporting only)
  3286. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  3287. cat("\nLevene’s test p-value =", levene_p, "\n")
  3288. # automatically choose correct t-test
  3289. if (levene_p > 0.05) {
  3290. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  3291. t_result <- t.test(
  3292. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3293. data = trim2_data_unique,
  3294. var.equal = TRUE
  3295. )
  3296. } else {
  3297. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3298. t_result <- t.test(
  3299. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3300. data = trim2_data_unique,
  3301. var.equal = FALSE
  3302. )
  3303. }
  3304. # print t-test result in full precision
  3305. print(t_result)
  3306. ```
  3307. *Minimum Maternal Hb measurement in Trimester 3*
  3308. ```{r}
  3309. # maternal minimum Hb in trimester 3 by antenatal anaemia status
  3310. # t-test
  3311. # unique subject IDs: based on true prevalence
  3312. library(dplyr)
  3313. library(car)
  3314. # filter for trimester 3, non-missing anaemia status and min Hb, keep unique subjects
  3315. trim3_data_unique <- ULF_ida_final %>%
  3316. filter(
  3317. maternal_trimester_hb == "third",
  3318. !is.na(maternal_preg_anaemia_status),
  3319. !is.na(maternal_preg_min_hb)
  3320. ) %>%
  3321. distinct(subject_id, .keep_all = TRUE)
  3322. # summary statistics by anaemia status with exactly 2 decimals displayed
  3323. summary_stats <- trim3_data_unique %>%
  3324. group_by(maternal_preg_anaemia_status) %>%
  3325. summarise(
  3326. n = n(),
  3327. mean_hb = format(round(mean(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3328. sd_hb = format(round(sd(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3329. min_hb = format(round(min(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3330. max_hb = format(round(max(maternal_preg_min_hb, na.rm = TRUE), 2), nsmall = 2),
  3331. .groups = "drop"
  3332. )
  3333. # print summary stats
  3334. print(summary_stats)
  3335. # levene's test for homogeneity of variance
  3336. levene_result <- leveneTest(
  3337. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3338. data = trim3_data_unique
  3339. )
  3340. print(levene_result)
  3341. # extract Levene p-value (rounded for reporting only)
  3342. levene_p <- round(levene_result$`Pr(>F)`[1], 4)
  3343. cat("\nLevene’s test p-value =", levene_p, "\n")
  3344. # automatically choose correct t-test
  3345. if (levene_p > 0.05) {
  3346. cat("Levene’s test not significant → Using Student's t-test (equal variances assumed)\n\n")
  3347. t_result <- t.test(
  3348. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3349. data = trim3_data_unique,
  3350. var.equal = TRUE
  3351. )
  3352. } else {
  3353. cat("Levene’s test significant → Using Welch's t-test (equal variances NOT assumed)\n\n")
  3354. t_result <- t.test(
  3355. maternal_preg_min_hb ~ maternal_preg_anaemia_status,
  3356. data = trim3_data_unique,
  3357. var.equal = FALSE
  3358. )
  3359. }
  3360. # print t-test result in full precision
  3361. print(t_result)
  3362. ```
  3363. **Exploratory Analyses: Plotted Data using LOESS Curves**
  3364. Absolute Volumes for ROIs plotted against child age at scan. Includes repeated measures to capture all scans collected across study timepoints.
  3365. 1. Full sample
  3366. 2. By group (Key: Antenatal Maternal Anaemia Status)
  3367. *LOESS Curves: ICV*
  3368. ```{r}
  3369. # icv (full sample)
  3370. # including repeated measures
  3371. library(ggplot2)
  3372. icv_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = mm_hyp_icv)) +
  3373. geom_point(alpha = 0.6) +
  3374. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  3375. scale_x_continuous(
  3376. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  3377. ) +
  3378. labs(
  3379. title = "ICV vs. Child Age",
  3380. x = "Child Age at Scan (Months)",
  3381. y = "Total ICV"
  3382. ) +
  3383. theme_minimal()
  3384. icv_ulf_sample
  3385. ```
  3386. ```{r}
  3387. # icv by group (antenatal maternal anaemia status)
  3388. # including repeated measures
  3389. icv_ulf_group <- ggplot(
  3390. data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  3391. aes(
  3392. x = scan_age_months,
  3393. y = mm_hyp_icv,
  3394. color = maternal_preg_anaemia_status,
  3395. fill = maternal_preg_anaemia_status
  3396. )
  3397. ) +
  3398. geom_point() +
  3399. geom_smooth(
  3400. method = "loess",
  3401. se = TRUE,
  3402. level = 0.95,
  3403. alpha = 0.3 # slightly transparent ribbons
  3404. ) +
  3405. scale_color_manual(
  3406. values = c(
  3407. "no_mat_anaemia" = "#1f77b4",
  3408. "mat_anaemia" = "#d62728"
  3409. ),
  3410. name = "Maternal Anaemia Status"
  3411. ) +
  3412. scale_fill_manual(
  3413. values = c(
  3414. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
  3415. "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
  3416. ),
  3417. guide = "none"
  3418. ) +
  3419. scale_x_continuous(
  3420. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  3421. ) +
  3422. scale_y_continuous(
  3423. limits = c(600000, NA),
  3424. breaks = seq(600000, max(ULF_ida_final$mm_hyp_icv, na.rm = TRUE), by = 200000)
  3425. ) +
  3426. labs(
  3427. title = "Total ICV Volume vs. Scan Age",
  3428. x = "Scan Age (Months)",
  3429. y = "Total ICV Volume (mm³)",
  3430. color = "Maternal Anaemia Status"
  3431. ) +
  3432. theme_minimal()
  3433. icv_ulf_group
  3434. ```
  3435. *LOESS Curve: Putamen*
  3436. ```{r}
  3437. # putamen (full sample)
  3438. # including repeated measures
  3439. library(ggplot2)
  3440. putamen_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_putamen)) +
  3441. geom_point(alpha = 0.6) +
  3442. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  3443. scale_x_continuous(
  3444. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  3445. ) +
  3446. labs(
  3447. title = "Putamen vs. Child Age",
  3448. x = "Child Age at Scan (Months)",
  3449. y = "Total Putamen Volume"
  3450. ) +
  3451. theme_minimal()
  3452. putamen_ulf_sample
  3453. ```
  3454. ```{r}
  3455. # putamen by group (antenatal maternal anaemia status)
  3456. # including repeated measures
  3457. putamen_ulf_group <- ggplot(
  3458. data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  3459. aes(
  3460. x = scan_age_months,
  3461. y = total_putamen,
  3462. color = maternal_preg_anaemia_status,
  3463. fill = maternal_preg_anaemia_status
  3464. )
  3465. ) +
  3466. geom_point() +
  3467. geom_smooth(
  3468. method = "loess",
  3469. se = TRUE,
  3470. level = 0.95,
  3471. alpha = 0.3 # slightly transparent ribbons
  3472. ) +
  3473. scale_color_manual(
  3474. values = c(
  3475. "no_mat_anaemia" = "#1f77b4",
  3476. "mat_anaemia" = "#d62728"
  3477. ),
  3478. name = "Maternal Anaemia Status"
  3479. ) +
  3480. scale_fill_manual(
  3481. values = c(
  3482. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
  3483. "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
  3484. ),
  3485. guide = "none"
  3486. ) +
  3487. scale_x_continuous(
  3488. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  3489. ) +
  3490. labs(
  3491. title = "Total Putamen Volume vs. Scan Age",
  3492. x = "Scan Age (months)",
  3493. y = "Total Putamen Volume",
  3494. color = "Maternal Anaemia Status"
  3495. ) +
  3496. theme_minimal()
  3497. putamen_ulf_group
  3498. ```
  3499. *LOESS Curves: Caudate Nucleus*
  3500. ```{r}
  3501. # caudate nucleus (full sample)
  3502. # including repeated measures
  3503. library(ggplot2)
  3504. caudate_hf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_caudate)) +
  3505. geom_point(alpha = 0.6) +
  3506. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  3507. scale_x_continuous(
  3508. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  3509. ) +
  3510. labs(
  3511. title = "Caudate vs. Child Age",
  3512. x = "Child Age at Scan (Months)",
  3513. y = "Total Caudate Volume"
  3514. ) +
  3515. theme_minimal()
  3516. caudate_hf_sample
  3517. ```
  3518. ```{r}
  3519. # caudate nucleus by group (antenatal maternal anaemia status)
  3520. # including repeated measures
  3521. caudate_ulf_group <- ggplot(
  3522. data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  3523. aes(
  3524. x = scan_age_months,
  3525. y = total_caudate,
  3526. color = maternal_preg_anaemia_status,
  3527. fill = maternal_preg_anaemia_status
  3528. )
  3529. ) +
  3530. geom_point() +
  3531. geom_smooth(
  3532. method = "loess",
  3533. se = TRUE,
  3534. level = 0.95,
  3535. alpha = 0.3 # slightly transparent ribbons
  3536. ) +
  3537. scale_color_manual(
  3538. values = c(
  3539. "no_mat_anaemia" = "#1f77b4",
  3540. "mat_anaemia" = "#d62728"
  3541. ),
  3542. name = "Maternal Anaemia Status"
  3543. ) +
  3544. scale_fill_manual(
  3545. values = c(
  3546. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
  3547. "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
  3548. ),
  3549. guide = "none"
  3550. ) +
  3551. scale_x_continuous(
  3552. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  3553. ) +
  3554. labs(
  3555. title = "Total Caudate Volume vs. Scan Age",
  3556. x = "Scan Age (months)",
  3557. y = "Total Caudate Volume",
  3558. color = "Maternal Anaemia Status"
  3559. ) +
  3560. theme_minimal()
  3561. caudate_ulf_group
  3562. ```
  3563. *LOESS Curves: Corpus Callosum*
  3564. ```{r}
  3565. # corpus callosum (full sample)
  3566. # including repeated measures
  3567. library(ggplot2)
  3568. cc_ulf_sample <- ggplot(ULF_ida_final, aes(x = scan_age_months, y = total_cc)) +
  3569. geom_point(alpha = 0.6) +
  3570. geom_smooth(method = "loess", se = TRUE, color = "blue") +
  3571. scale_x_continuous(
  3572. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3) # x-axis every 3 months
  3573. ) +
  3574. labs(
  3575. title = "CC vs. Child Age",
  3576. x = "Child Age at Scan (Months)",
  3577. y = "Total Corpus Callosum Volume"
  3578. ) +
  3579. theme_minimal()
  3580. cc_ulf_sample
  3581. ```
  3582. ```{r}
  3583. # corpus callosum by group (antenatal maternal anaemia status)
  3584. # including repeated measures
  3585. cc_ulf_group <- ggplot(
  3586. data = ULF_ida_final %>% filter(!is.na(maternal_preg_anaemia_status)),
  3587. aes(
  3588. x = scan_age_months,
  3589. y = total_cc,
  3590. color = maternal_preg_anaemia_status,
  3591. fill = maternal_preg_anaemia_status
  3592. )
  3593. ) +
  3594. geom_point() +
  3595. geom_smooth(
  3596. method = "loess",
  3597. se = TRUE,
  3598. level = 0.95,
  3599. alpha = 0.3 # slightly transparent ribbons
  3600. ) +
  3601. scale_color_manual(
  3602. values = c(
  3603. "no_mat_anaemia" = "#1f77b4",
  3604. "mat_anaemia" = "#d62728"
  3605. ),
  3606. name = "Maternal Anaemia Status"
  3607. ) +
  3608. scale_fill_manual(
  3609. values = c(
  3610. "no_mat_anaemia" = "#c6d9f1", # slightly lighter blue ribbon
  3611. "mat_anaemia" = "#f7b0b0" # slightly lighter red ribbon
  3612. ),
  3613. guide = "none"
  3614. ) +
  3615. scale_x_continuous(
  3616. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  3617. ) +
  3618. scale_y_continuous(
  3619. n.breaks = 6
  3620. ) +
  3621. labs(
  3622. title = "Total CC Volume vs. Scan Age",
  3623. x = "Scan Age (Months)",
  3624. y = "Total Corpus Callosum Volume (mm³)",
  3625. color = "Maternal Anaemia Status"
  3626. ) +
  3627. theme_minimal()
  3628. cc_ulf_group
  3629. ```
  3630. **Zero-Order Correlation Matrices: Antenatal Maternal Anaemia Status**
  3631. ```{r}
  3632. # install packages if needed
  3633. # install and load Hmisc
  3634. if (!require(Hmisc)) {
  3635. install.packages("Hmisc")
  3636. library(Hmisc)
  3637. } else {
  3638. library(Hmisc)}
  3639. ```
  3640. Convert variables to numeric for correlations
  3641. ```{r}
  3642. # convert variables back from categorical to numeric in ULF_ida_final
  3643. # household income
  3644. ULF_ida_final$income_household_en <- as.numeric(ULF_ida_final$income_household_en)
  3645. # maternal anaemia status: 0 = no anaemia, 1 = anaemia
  3646. ULF_ida_final$maternal_preg_anaemia_status_numeric <- as.numeric(as.character(
  3647. factor(ULF_ida_final$maternal_preg_anaemia_status,
  3648. levels = c("no_mat_anaemia", "mat_anaemia"),
  3649. labels = c(0, 1))
  3650. ))
  3651. # maternal HIV status: 0 = no HIV, 1 = HIV
  3652. ULF_ida_final$maternal_hiv <- as.numeric(as.character(
  3653. factor(ULF_ida_final$maternal_hiv,
  3654. levels = c("no_hiv", "mat_hiv"),
  3655. labels = c(0, 1))
  3656. ))
  3657. # child sex: 0 = girl, 1 = boy
  3658. ULF_ida_final$child_sex <- as.numeric(as.character(
  3659. factor(ULF_ida_final$child_sex,
  3660. levels = c("girl", "boy"),
  3661. labels = c(0, 1))
  3662. ))
  3663. ```
  3664. The variable for "antenatal maternal anaemia status" remained a categorical variable even after running the code above. Therefore, I recalled the dataset to ensure that this was numeric (as initially defined) for modelling:
  3665. ```{r}
  3666. # recall dataset
  3667. # load dataset
  3668. ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
  3669. # create New Variables for Total Brain Volumes across ROIs
  3670. # basal ganglia
  3671. ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
  3672. ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
  3673. # corpus callosum
  3674. ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
  3675. # convert days to months
  3676. # average Number of days in a month (accounting for leap years): 30.44
  3677. ULF_ida_final <- ULF_ida_final %>%
  3678. mutate(scan_age_months = hyp_age / 30.44)
  3679. ```
  3680. **Correlation Matrices**
  3681. ```{r}
  3682. # correlation matrix for corpus callosum
  3683. # filter dataset
  3684. filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
  3685. # select relevant columns
  3686. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_cc")
  3687. data_selected <- filtered_data[, vars]
  3688. # convert categorical variables to numeric (for correlation)
  3689. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  3690. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  3691. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  3692. # compute correlation matrix with significance
  3693. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  3694. # rcorr_result$r contains correlation coefficients
  3695. print("Correlation matrix:")
  3696. print(rcorr_result$r)
  3697. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  3698. print("P-value matrix:")
  3699. print(rcorr_result$P)
  3700. ```
  3701. ```{r}
  3702. # correlation matrix for caudate nucleus
  3703. # filter dataset
  3704. filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
  3705. # select relevant columns
  3706. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_caudate")
  3707. data_selected <- filtered_data[, vars]
  3708. # convert categorical variables to numeric (for correlation)
  3709. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  3710. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  3711. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  3712. # compute correlation matrix with significance
  3713. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  3714. # rcorr_result$r contains correlation coefficients
  3715. print("Correlation matrix:")
  3716. print(rcorr_result$r)
  3717. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  3718. print("P-value matrix:")
  3719. print(rcorr_result$P)
  3720. ```
  3721. ```{r}
  3722. # correlation matrix for putamen
  3723. # filter dataset
  3724. filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
  3725. # select relevant columns
  3726. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "total_putamen")
  3727. data_selected <- filtered_data[, vars]
  3728. # convert categorical variables to numeric (for correlation)
  3729. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  3730. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  3731. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  3732. # compute correlation matrix with significance
  3733. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  3734. # rcorr_result$r contains correlation coefficients
  3735. print("Correlation matrix:")
  3736. print(rcorr_result$r)
  3737. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  3738. print("P-value matrix:")
  3739. print(rcorr_result$P)
  3740. ```
  3741. ```{r}
  3742. # correlation matrix for icv
  3743. # filter dataset
  3744. filtered_data <- ULF_ida_final[!is.na(ULF_ida_final$maternal_preg_anaemia_status), ]
  3745. # select relevant columns
  3746. vars <- c("child_sex", "scan_age_months", "income_household_en", "maternal_hiv", "maternal_preg_anaemia_status", "mm_hyp_icv")
  3747. data_selected <- filtered_data[, vars]
  3748. # convert categorical variables to numeric (for correlation)
  3749. data_selected$child_sex <- as.numeric(as.factor(data_selected$child_sex))
  3750. data_selected$maternal_hiv <- as.numeric(as.factor(data_selected$maternal_hiv))
  3751. data_selected$maternal_preg_anaemia_status <- as.numeric(as.factor(data_selected$maternal_preg_anaemia_status))
  3752. # compute correlation matrix with significance
  3753. rcorr_result <- Hmisc::rcorr(as.matrix(data_selected))
  3754. # rcorr_result$r contains correlation coefficients
  3755. print("Correlation matrix:")
  3756. print(rcorr_result$r)
  3757. # rcorr_result$P contains p-values for testing the null hypothesis of no correlation
  3758. print("P-value matrix:")
  3759. print(rcorr_result$P)
  3760. ```
  3761. **Multicollinearity**
  3762. ```{r}
  3763. library(dplyr)
  3764. correlation_test <- cor.test(
  3765. ULF_ida_final$scan_age_months[!is.na(ULF_ida_final$maternal_preg_anaemia_status)],
  3766. ULF_ida_final$mm_hyp_icv[!is.na(ULF_ida_final$maternal_preg_anaemia_status)])
  3767. print(correlation_test)
  3768. ```
  3769. Decision to exclude ICV from the models for regional brain volumes due to high levels of multicollinearity between child age at scan. Only child age at scan was included in the models below.
  3770. **Linear Mixed Effects (LME) Models**
  3771. ```{r}
  3772. # install packages if needed
  3773. #install.packages("lme4")
  3774. library(lme4)
  3775. #install.packages("lmerTest")
  3776. library(lmerTest)
  3777. #install.packages("parameters")
  3778. library(parameters)
  3779. #install.packages("emmeans")
  3780. library(emmeans)
  3781. ```
  3782. Based on the exploratory analyses (LOESS curves; plotted data), absolute age was used for models with an observed linear relationship. Log (age) was used for models with an observed non-linear relationship.
  3783. - ICV, Putamen, Caudate Nucleus, Corpus Callosm: Log (age) - Growth trajectory characterised by rapid growth and plateauing on plot
  3784. For the ULF sample, interactions (child age at scan \* antenatal maternal anaemia status) were explored as there was sufficient power. If the interaction was not significant, the model was re-run without an interaction effect and only main effects were interpreted.
  3785. **LME Model variables**
  3786. - Ensure that all categorical variables (anaemia status, child sex, maternal HIV status etc) are factors
  3787. - Create variable for the logarithmic transformation of child age in months
  3788. ```{r}
  3789. # recall dataset so that all input variables are categorical
  3790. # load dataset
  3791. ULF_ida_final <- read_excel("ULF_ida_final.xlsx")
  3792. # create New Variables for Total Brain Volumes across ROIs
  3793. # basal ganglia
  3794. ULF_ida_final$total_putamen <- ULF_ida_final$mm_hyp_left_putamen + ULF_ida_final$mm_hyp_right_putamen
  3795. ULF_ida_final$total_caudate <- ULF_ida_final$mm_hyp_left_caudate + ULF_ida_final$mm_hyp_right_caudate
  3796. # corpus callosum
  3797. ULF_ida_final$total_cc <- ULF_ida_final$mm_hyp_posterior_callosum + ULF_ida_final$mm_hyp_mid_posterior_callosum + ULF_ida_final$mm_hyp_central_callosum + ULF_ida_final$mm_hyp_mid_anterior_callosum + ULF_ida_final$mm_hyp_anterior_callosum
  3798. # create age variable
  3799. # convert days to months
  3800. # average Number of days in a month (accounting for leap years): 30.44
  3801. ULF_ida_final <- ULF_ida_final %>%
  3802. mutate(scan_age_months = hyp_age / 30.44)
  3803. ```
  3804. ```{r}
  3805. # convert categorical variables form numeric to factor
  3806. # maternal anaemia status
  3807. ULF_ida_final$maternal_preg_anaemia_status <- factor(
  3808. ULF_ida_final$maternal_preg_anaemia_status,
  3809. levels = c(0, 1),
  3810. labels = c("no_mat_anaemia", "mat_anaemia")
  3811. )
  3812. # maternal anaemia severity
  3813. ULF_ida_final$maternal_preg_anaemia_severity <- factor(
  3814. ULF_ida_final$maternal_preg_anaemia_severity,
  3815. levels = c(1, 2, 3),
  3816. labels = c("mild", "moderate", "severe")
  3817. )
  3818. # child status
  3819. ULF_ida_final$cor_child_anaemia <- factor(
  3820. ULF_ida_final$cor_child_anaemia,
  3821. levels = c(0, 1),
  3822. labels = c("no_child_anaemia", "child_anaemia")
  3823. )
  3824. # maternal ID status
  3825. ULF_ida_final$maternal_preg_adj_brinda_id <- factor(
  3826. ULF_ida_final$maternal_preg_adj_brinda_id,
  3827. levels = c(0, 1),
  3828. labels = c("no_mat_id", "mat_id")
  3829. )
  3830. # child ID status
  3831. ULF_ida_final$cor_child_id <- factor(
  3832. ULF_ida_final$cor_child_id,
  3833. levels = c(0, 1),
  3834. labels = c("no_child_id", "child_id")
  3835. )
  3836. # maternal trimester hb
  3837. ULF_ida_final$maternal_trimester_hb <- factor(
  3838. ULF_ida_final$maternal_trimester_hb,
  3839. levels = c(1, 2, 3),
  3840. labels = c("first", "second", "third")
  3841. )
  3842. # child sex
  3843. ULF_ida_final$child_sex <- factor(
  3844. ULF_ida_final$child_sex,
  3845. levels = c(0, 1),
  3846. labels = c("girl", "boy")
  3847. )
  3848. # maternal pae
  3849. ULF_ida_final$maternal_pae <- factor(
  3850. ULF_ida_final$maternal_pae,
  3851. levels = c(0, 1),
  3852. labels = c("no_pae", "pae")
  3853. )
  3854. # maternal pte
  3855. ULF_ida_final$maternal_pte <- factor(
  3856. ULF_ida_final$maternal_pte,
  3857. levels = c(0, 1),
  3858. labels = c("no_pte", "pte")
  3859. )
  3860. # maternal hiv
  3861. ULF_ida_final$maternal_hiv <- factor(
  3862. ULF_ida_final$maternal_hiv,
  3863. levels = c(0, 1),
  3864. labels = c("no_hiv", "hiv")
  3865. )
  3866. # maternal dep
  3867. ULF_ida_final$maternal_dep <- factor(
  3868. ULF_ida_final$maternal_dep,
  3869. levels = c(0, 1),
  3870. labels = c("no_dep", "dep")
  3871. )
  3872. # maternal employment
  3873. ULF_ida_final$mom_employment_en_dichotomous <- factor(
  3874. ULF_ida_final$mom_employment_en_dichotomous,
  3875. levels = c(0, 1),
  3876. labels = c("unemployed", "employed")
  3877. )
  3878. # household income
  3879. ULF_ida_final$income_household_en <- factor(
  3880. ULF_ida_final$income_household_en,
  3881. levels = c(1, 2, 3, 4),
  3882. labels = c("Less than R1000 per month", "R1000-R5000 per monthy", " R5000-R10 000 per month", "More than R10 000 per month" )
  3883. )
  3884. # maternal education
  3885. ULF_ida_final$mom_edu_en <- factor(
  3886. ULF_ida_final$mom_edu_en,
  3887. levels = c(2, 3, 4, 5, 6),
  3888. labels = c("primary", "some_secondary", "complete_secondary", "some_tertiary", "completed_tertiary" )
  3889. )
  3890. #baby hiv
  3891. ULF_ida_final$baby_hiv <- factor(
  3892. ULF_ida_final$baby_hiv,
  3893. levels = c(0, 1),
  3894. labels = c("no_child_hiv", "child_hiv")
  3895. )
  3896. ```
  3897. ```{r}
  3898. # create a log(age) variable for use in models of brain regions with non-linear growth
  3899. ULF_ida_final$log_age <- log(ULF_ida_final$scan_age_months)
  3900. ```
  3901. *LME: Corpus Callosum*
  3902. ```{r}
  3903. # lme for corpus callosum
  3904. # including interaction between maternal anaemia and age
  3905. # log (age): non-linear growth
  3906. #model summary
  3907. ulf_model_cc_interaction <- lmer(total_cc ~ maternal_preg_anaemia_status * log_age +
  3908. child_sex +
  3909. (1 | subject_id),
  3910. data = ULF_ida_final)
  3911. summary(ulf_model_cc_interaction)
  3912. # beta values
  3913. standardize_parameters(ulf_model_cc_interaction)
  3914. ```
  3915. | |
  3916. |-----|
  3917. | |
  3918. ```{r}
  3919. # emms for corpus callosum volume by antenatal maternal anaemia status
  3920. # based on final adjusted model with interaction between maternal anaemia and age
  3921. # by timepoint for significant interaction
  3922. library(emmeans)
  3923. library(dplyr)
  3924. # step 1: Define raw ages and their logs for EMMs
  3925. raw_ages <- c(3, 6, 12, 18, 24)
  3926. log_ages <- log(raw_ages)
  3927. # step 2: Get estimated marginal means at these log ages
  3928. emm <- emmeans(
  3929. ulf_model_cc_interaction,
  3930. specs = ~ maternal_preg_anaemia_status | log_age,
  3931. at = list(log_age = log_ages)
  3932. )
  3933. # step 3: Convert to data frame and add back-transformed age for interpretability
  3934. emm_cc <- as.data.frame(emm)
  3935. emm_cc$age_months <- round(exp(emm_cc$log_age), 0)
  3936. print(emm_cc)
  3937. # step 4: Run pairwise contrasts (differences between anaemia groups) at each age
  3938. contrast_cc <- contrast(
  3939. emm,
  3940. method = "revpairwise",
  3941. adjust = "holm",
  3942. infer = TRUE
  3943. ) %>%
  3944. as.data.frame()
  3945. # step 5: Add back-transformed age for contrasts
  3946. contrast_cc$age_months <- round(exp(contrast_cc$log_age), 0)
  3947. print(contrast_cc)
  3948. ```
  3949. *LME: ICV*
  3950. ```{r}
  3951. # lme for icv
  3952. # including nteraction between maternal anaemia and age
  3953. # log (age): non-linear growth
  3954. # model summary
  3955. ulf_model_icv_interaction <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
  3956. data = ULF_ida_final)
  3957. summary(ulf_model_icv_interaction)
  3958. # beta values
  3959. standardize_parameters(ulf_model_icv_interaction)
  3960. ```
  3961. Interaction was not significant, so re-ran the mdoel without it and interpreted main effects:
  3962. ```{r}
  3963. # lme for icv
  3964. # excluding interaction between maternal anaemia and age
  3965. # log (age): non-linear growth
  3966. # model summary
  3967. ulf_model_icv_no_interaction <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
  3968. data = ULF_ida_final)
  3969. summary(ulf_model_icv_no_interaction)
  3970. # beta values
  3971. standardize_parameters(ulf_model_icv_no_interaction)
  3972. ```
  3973. ```{r}
  3974. # emms for icv volume by antenatal maternal anaemia status
  3975. # based on final adjusted model with no interaction
  3976. library(emmeans)
  3977. emm_icv <- emmeans(ulf_model_icv_no_interaction, ~ maternal_preg_anaemia_status)
  3978. # ----- EMM table with 2 decimals -----
  3979. emm_icv_df <- as.data.frame(emm_icv)
  3980. emm_icv_df$emmean <- format(round(emm_icv_df$emmean, 2), nsmall = 2)
  3981. emm_icv_df$SE <- format(round(emm_icv_df$SE, 2), nsmall = 2)
  3982. emm_icv_df$lower.CL <- format(round(emm_icv_df$lower.CL, 2), nsmall = 2)
  3983. emm_icv_df$upper.CL <- format(round(emm_icv_df$upper.CL, 2), nsmall = 2)
  3984. print(emm_icv_df)
  3985. # ----- contrast results -----
  3986. contrast_results <- contrast(emm_icv, method = "revpairwise", adjust = "none")
  3987. contrast_df <- as.data.frame(summary(contrast_results, infer = TRUE))
  3988. # force 2 decimals for estimate
  3989. contrast_df$estimate <- format(round(contrast_df$estimate, 2), nsmall = 2)
  3990. print(contrast_df)
  3991. ```
  3992. *LME: Caudate Nucleus*
  3993. ```{r}
  3994. # lme for caudate nucleus
  3995. # icluding interaction between maternal anaemia and age
  3996. # log (age): non-linear growth
  3997. # model summary
  3998. ulf_model_caudate_interaction <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
  3999. data = ULF_ida_final)
  4000. summary(ulf_model_caudate_interaction)
  4001. # beta values
  4002. standardize_parameters(ulf_model_caudate_interaction)
  4003. ```
  4004. ```{r}
  4005. # emms for caudate volume by antenatal maternal anaemia status
  4006. # based on final adjusted model with interaction
  4007. # by timepoint for significant interaction
  4008. library(emmeans)
  4009. library(dplyr)
  4010. # step 1: Define raw ages and their logs for EMMs
  4011. raw_ages <- c(3, 6, 12, 18, 24)
  4012. log_ages <- log(raw_ages)
  4013. # step 2: Get estimated marginal means at these log ages
  4014. emm <- emmeans(
  4015. ulf_model_caudate_interaction,
  4016. specs = ~ maternal_preg_anaemia_status | log_age,
  4017. at = list(log_age = log_ages)
  4018. )
  4019. # step 3: Convert to data frame and add back-transformed age for interpretability
  4020. emm_caudate <- as.data.frame(emm)
  4021. emm_caudate$age_months <- round(exp(emm_caudate$log_age), 0)
  4022. print(emm_caudate)
  4023. # step 4: Run pairwise contrasts (differences between anaemia groups) at each age
  4024. contrast_caudate <- contrast(emm, method = "revpairwise", adjust = "holm", infer = TRUE) %>% as.data.frame()
  4025. # step 5: Add back-transformed age for contrasts (will be in 'log_age' column)
  4026. contrast_caudate$age_months <- round(exp(contrast_caudate$log_age), 0)
  4027. print(contrast_caudate)
  4028. ```
  4029. *LME: Putamen*
  4030. ```{r}
  4031. # lme for putamen
  4032. # including interaction between maternal anaemia and age
  4033. # log (age): non-linear growth
  4034. # model summary
  4035. ulf_model_putamen_interaction <- lmer(total_putamen ~ maternal_preg_anaemia_status * log_age + child_sex + (1 | subject_id),
  4036. data = ULF_ida_final)
  4037. summary(ulf_model_putamen_interaction)
  4038. # beta values
  4039. standardize_parameters(ulf_model_putamen_interaction)
  4040. ```
  4041. Interaction was not significant, so re-ran the model without it and interpreted main effects:
  4042. ```{r}
  4043. # lme for putamen
  4044. # excluding interaction between maternal anaemia and age
  4045. # model summary
  4046. ulf_model_putamen_no_interaction <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + (1 | subject_id),
  4047. data = ULF_ida_final)
  4048. summary(ulf_model_putamen_no_interaction)
  4049. # beta values
  4050. standardize_parameters(ulf_model_putamen_no_interaction)
  4051. ```
  4052. ```{r}
  4053. # emms for putamen volume by antenatal maternal anaemia status
  4054. # based on final adjusted model with no interaction
  4055. library(emmeans)
  4056. emm_putamen <- emmeans(ulf_model_putamen_no_interaction, ~ maternal_preg_anaemia_status)
  4057. # ----- emm table with exactly 2 decimals -----
  4058. emm_putamen_df <- as.data.frame(emm_putamen)
  4059. emm_putamen_df$emmean <- format(round(emm_putamen_df$emmean, 2), nsmall = 2)
  4060. emm_putamen_df$SE <- format(round(emm_putamen_df$SE, 2), nsmall = 2)
  4061. emm_putamen_df$lower.CL <- format(round(emm_putamen_df$lower.CL, 2), nsmall = 2)
  4062. emm_putamen_df$upper.CL <- format(round(emm_putamen_df$upper.CL, 2), nsmall = 2)
  4063. print(emm_putamen_df)
  4064. # ----- contrast results -----
  4065. contrast_results_putamen <- contrast(emm_putamen, method = "revpairwise", adjust = "none")
  4066. contrast_df_putamen <- as.data.frame(summary(contrast_results_putamen, infer = TRUE))
  4067. # force 2 decimals for estimate
  4068. contrast_df_putamen$estimate <- format(round(contrast_df_putamen$estimate, 2), nsmall = 2)
  4069. print(contrast_df_putamen)
  4070. ```
  4071. **Sensitivity Analyses**
  4072. Based on the final models for each ROI (established in primary analyses): With the inclusion of maternal HIV (potential confounder)
  4073. - ICV: Final model was one without an interaction
  4074. - Putamen: FInal model was without an interaction
  4075. - Corpus Callosum: Final model was with an interaction
  4076. - Caudate Nucleus: Final model was with an interaction
  4077. ```{r}
  4078. # sensitivity analysis for ICV
  4079. # final model: excluding interaction between maternal anaemia and age
  4080. # including maternal hiv
  4081. # model summary
  4082. ulf_model_icv_no_interaction_sens <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
  4083. data = ULF_ida_final)
  4084. summary(ulf_model_icv_no_interaction_sens)
  4085. # beta values
  4086. standardize_parameters(ulf_model_icv_no_interaction_sens)
  4087. ```
  4088. ```{r}
  4089. # sensitivity analysis for corpus callosum
  4090. # final model: including interaction between maternal anaemia and age
  4091. # including maternal hiv
  4092. # model summary
  4093. ulf_model_cc_interaction_sens <- lmer(total_cc ~ maternal_preg_anaemia_status * scan_age_months + child_sex + maternal_hiv + (1 | subject_id), data = ULF_ida_final)
  4094. summary(ulf_model_cc_interaction_sens)
  4095. # beta values
  4096. standardize_parameters(ulf_model_cc_interaction_sens)
  4097. ```
  4098. ```{r}
  4099. # sensitivity analysis for putamen
  4100. # final model: excluding interaction between maternal anaemia and age
  4101. # including maternal hiv
  4102. # model summary
  4103. ulf_model_putamen_no_interaction_sens <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + maternal_hiv + (1 | subject_id),
  4104. data = ULF_ida_final)
  4105. summary(ulf_model_putamen_no_interaction_sens)
  4106. # beta values
  4107. standardize_parameters(ulf_model_putamen_no_interaction_sens)
  4108. ```
  4109. ```{r}
  4110. # sensitivity analysis for caudate nucleus
  4111. # final model: including interaction between maternal anaemia and age
  4112. # including maternal hiv
  4113. # model summary
  4114. ulf_model_caudate_interaction_sens <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + maternal_hiv + (1 | subject_id),
  4115. data = ULF_ida_final)
  4116. summary(ulf_model_caudate_interaction_sens)
  4117. # beta values
  4118. standardize_parameters(ulf_model_caudate_interaction_sens)
  4119. ```
  4120. ------------------------------------------------------------------------
  4121. **Secondary Analysis: Child Anaemia (ULF Sample)**
  4122. ------------------------------------------------------------------------
  4123. **Sample Size: For infants with Child Anemia Data**
  4124. Note: Sample sizes were calculated with 1) the inclusion of repeated measures (multiple scans across study timepoints), and for 2) unique subject IDs excluding repeated measures.
  4125. ```{r}
  4126. # total sample size
  4127. # sample iwth child anaemia data
  4128. # repeated measures included
  4129. ULF_ida_final %>%
  4130. filter(!is.na(cor_child_anaemia)) %>% # only include PIDs with child anaemia data
  4131. summarise(n = n())
  4132. ```
  4133. ```{r}
  4134. # sample with child anaemia data
  4135. # repeated measures excluded
  4136. # unique subject IDs: based on true prevalence
  4137. ULF_ida_final %>%
  4138. filter(!is.na(cor_child_anaemia)) %>%
  4139. distinct(subject_id, .keep_all = TRUE) %>% #only include PIDs with child anaemia data
  4140. summarise(n = n())
  4141. length(unique(ULF_ida_final$subject_id)) # includes chidlren with no anaemia data
  4142. ```
  4143. **Child Anaemia Variables**
  4144. An overall classification for child anaemia status was determined based on the diagnosis of anaemia at least once across study visits. This was used for the assessment of group differences in sample characteristics between infants who had ever been anaemic versus infants who had never been anaemic between 3 and 24 months of age. Minimum child haemoglobin and corresponding age across study visits were also determined for each infant and reported on across both groups. For modelling with the repeated measures subsample (multiple scans per child in some instances), child anaemia status was assessed based on the most severe presentation of anaemia as indicated by the minimum haemoglobin measure observed between birth and the time of scan.
  4145. Each infant has multiple haemoglobin measurements from across study timepoints.
  4146. - "Overall Anaemia" variable: For the purpose of assessing group differences in sample characteristics based on unqiue subject IDs (excluding repeated measures), we created a binary variable for child anaemia status based on being anaemic at least once between 3 and 24 months. i.e. "overall anaemia" = "anaemic at any timepoint"
  4147. - Child anaemia status in repeated measures models (LMEs): Every scan has a corresponding child anaemia status variable (cor_child_anaemia) based on the minimum Hb measurement (cor_child_hb) between birth and the time of scan and the corresponding age (cor_child_hb_age) at the time of this measurement. i.e. In a repeated measures sample:
  4148. - 3-month scan: anaemia status based on Hb at 3 months
  4149. - 6-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3 and 6 months)
  4150. - 12-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, and 12 months)
  4151. - 18-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, 12, and 18 months)
  4152. - 24-month scan: anaemia status based on min Hb between birth and time of scan (i.e. between measures at 3, 6, 12, 18, and 24 months)
  4153. - Minimum haemoglobin measurement: This was computed to compare Hb across groups when assessing sample characteristics based on unique subjects (i.e. excluding repeated measures). We identified the minimum Hb measurement (and corresponding age) across study timepoints - i.e. across multiple cor_child_hb measures for children with multiple scans in a repeated measures sample.
  4154. **Creating New Binary Variable for Overall Anaemia Status**
  4155. ```{r}
  4156. # overall child anaemia status variable
  4157. # step 1: Create new dataframe with only child anaemia data. Select only anaemia variables interested in
  4158. # step 2: Pivot to wide format so each child ID appears once only (see all timepoints for each child)
  4159. dem_child <- ULF_ida_final %>% #create new dataframe
  4160. filter(!is.na(cor_child_anaemia)) %>% #filter so excluding children with no child anaemia data
  4161. select(subject_id, timepoint, cor_child_hb, cor_child_anaemia) %>% #select variables interested in
  4162. pivot_wider(id_cols = subject_id, names_from = timepoint, values_from = c(cor_child_hb,cor_child_anaemia)) #convert to wide format so can see anaemia data fro all timepoints
  4163. ```
  4164. ```{r}
  4165. # overall child anaemia status variable
  4166. # step 3: create overall anaemia variable - i.e. "if anaemic at any timepoint"
  4167. dem_child <- dem_child %>%
  4168. mutate(overall_child_anaemia = if_else(if_any(c(cor_child_anaemia_3M, cor_child_anaemia_6M, cor_child_anaemia_12M, cor_child_anaemia_18M, cor_child_anaemia_24M), ~ !is.na(.x) & .x == "child_anaemia"), "overall_child_anaemia", "no_overall_child_anaemia")) %>%
  4169. select(subject_id, overall_child_anaemia) #select and save the overall anaemia status
  4170. # step 4: Merge the overall_child_anaemia variable/column back into long format of the original dataframe - i.e. "ULF_ida_final"
  4171. ULF_ida_final <- left_join(ULF_ida_final, dem_child, by = 'subject_id')
  4172. ```
  4173. ```{r}
  4174. # overall child anaemia status variable
  4175. # sanity check that have the correct IDs
  4176. ULF_ida_final %>%
  4177. filter(!is.na(overall_child_anaemia)) %>%
  4178. distinct(subject_id, .keep_all = TRUE) %>% #only include PIDs with child anaemia data
  4179. summarise(n = n())
  4180. length(unique(ULF_ida_final$subject_id)) # inlcudes chidlren with no anaemia data
  4181. ```
  4182. **Creating Variable for Minimum Hb Across Study Timepoints**
  4183. ```{r}
  4184. # minimum child Hb variable
  4185. # step 1: Create new dataframe
  4186. # step 2: Pivot to wide format
  4187. dem_child_1 <- ULF_ida_final %>% #create new dataframe
  4188. filter(!is.na(cor_child_anaemia)) %>% #filter so excluding children with no child anaemia data
  4189. select(subject_id, timepoint, cor_child_hb, cor_child_hb_age, cor_child_anaemia) %>% #select variables interested in
  4190. pivot_wider(id_cols = subject_id, names_from = timepoint, values_from = c(cor_child_hb,cor_child_hb_age, cor_child_anaemia)) #convert to wide format so can see anaemia data for all timepoints
  4191. ```
  4192. ```{r}
  4193. # step 3: Create min hb variable - min hb from across timepoints
  4194. library(dplyr)
  4195. dem_child_1 <- dem_child_1 %>%
  4196. # create a new variable that represents the minimum child Hb across timepoints
  4197. mutate(
  4198. min_child_hb = pmin(cor_child_hb_3M, cor_child_hb_6M, cor_child_hb_12M, cor_child_hb_18M, cor_child_hb_24M, na.rm = TRUE),
  4199. # create a new variable that represents the age at the time of the minimum Hb
  4200. min_hb_age = case_when(
  4201. min_child_hb == cor_child_hb_3M ~ cor_child_hb_age_3M,
  4202. min_child_hb == cor_child_hb_6M ~ cor_child_hb_age_6M,
  4203. min_child_hb == cor_child_hb_12M ~ cor_child_hb_age_12M,
  4204. min_child_hb == cor_child_hb_18M ~ cor_child_hb_age_18M,
  4205. min_child_hb == cor_child_hb_24M ~ cor_child_hb_age_24M
  4206. )
  4207. ) %>%
  4208. # select subject_id, min_child_hb, and min_hb_age variables
  4209. select(subject_id, min_child_hb, min_hb_age)
  4210. # view the resulting dataframe
  4211. head(dem_child_1)
  4212. ```
  4213. ```{r}
  4214. # step 4: merge in the min Hb and cor age at min hb variable back into long format original dataframe "ULF_ida_final"
  4215. ULF_ida_final <- left_join(ULF_ida_final, dem_child_1, by = 'subject_id')
  4216. ```
  4217. **Anaemia Prevalence**
  4218. ```{r}
  4219. # child anaemia prevalence
  4220. # based on new variable for overall anaemia status
  4221. # unique subject IDs: based on true prevalence
  4222. ULF_ida_final %>%
  4223. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4224. distinct(subject_id, .keep_all = TRUE) %>%
  4225. group_by(overall_child_anaemia) %>%
  4226. summarise(n = n()) %>%
  4227. mutate(percent = round(100 * n / sum(n), 2))
  4228. ```
  4229. ```{r}
  4230. # maternal anaemia prevalence
  4231. # for infants with child anaemia data
  4232. # unique subject IDs: based on true prevalence
  4233. ULF_ida_final %>%
  4234. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4235. distinct(subject_id, .keep_all = TRUE) %>%
  4236. group_by(maternal_preg_anaemia_status) %>%
  4237. summarise(n = n()) %>%
  4238. mutate(percent = round(100 * n / sum(n), 2))
  4239. ```
  4240. **Sample Characteristics and Group Differences: Child Anaemia Status**
  4241. Sample characteristics were assessed based on the number of unique subjects with child anaemia (haemoglobin) data, excluding repeated measures data for children with multiple scans. This approach was used to avoid skewing results with duplicated demographic and clinical observations for the same infants. The only exception to this approach was child age at scan which was considered for the full repeated measures subsample, with a proportion of children having multiple scans acquired at different ages.
  4242. An overall classification for child anaemia status was determined based on the diagnosis of anaemia at least once across study visits. This was used for the assessment of group differences in sample characteristics between infants who had ever been anaemic versus infants who had never been anaemic between 3 and 24 months of age. Minimum child haemoglobin and corresponding age across study visits were also determined for each infant and reported on across both groups. For modelling with the repeated measures subsample (multiple scans per child in some instances), child anaemia status was assessed based on the most severe presentation of anaemia as indicated by the minimum haemoglobin measure observed between birth and the time of scan.
  4243. - **Categorical Variables**
  4244. *Maternal Anaemia*
  4245. ```{r}
  4246. # group differences in maternal anaemia by child anaemia status (using overall child anaemia variable created)
  4247. # unique subject IDs: based on true prevalence
  4248. # filter only for non-missing child anaemia status and keep unique subjects
  4249. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4250. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4251. # create contingency table: maternal anaemia vs child anaemia
  4252. tbl <- table(filtered$maternal_preg_anaemia_status, filtered$overall_child_anaemia)
  4253. # print table
  4254. cat("Contingency Table: Maternal anaemia status by Child Anaemia Status\n")
  4255. print(tbl)
  4256. # column percentages (percent within child anaemia status group)
  4257. cat("\nColumn Percentages (% within child anaemia status group):\n")
  4258. print(round(100 * prop.table(tbl, margin = 2), 2))
  4259. # choose appropriate test
  4260. chisq_result <- chisq.test(tbl)
  4261. if (any(chisq_result$expected < 5)) {
  4262. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4263. print(fisher.test(tbl))
  4264. } else {
  4265. cat("\nChi-squared Test:\n")
  4266. print(chisq_result)
  4267. }
  4268. ```
  4269. *Child Sex at Birth*
  4270. ```{r}
  4271. # child sex at birth (full sample)
  4272. # unique subject IDs: Based on true prevalence
  4273. ULF_ida_final %>%
  4274. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4275. distinct(subject_id, .keep_all = TRUE) %>%
  4276. group_by(child_sex) %>%
  4277. summarise(n = n()) %>%
  4278. mutate(percent = round(100 * n / sum(n), 2))
  4279. ```
  4280. ```{r}
  4281. # group differences in child sex by child anaemia status
  4282. # chi-squared: child sex
  4283. # unique subject IDs: based on true prevalence
  4284. # 1. filter for non-missing child anaemia status and keep unique subjects
  4285. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4286. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4287. # 2. create contingency table: child sex vs child anaemia
  4288. tbl <- table(filtered$child_sex, filtered$overall_child_anaemia)
  4289. # 3. print contingency table
  4290. cat("Contingency Table: Child Sex by Child Anaemia Status\n")
  4291. print(tbl)
  4292. # 4. column percentages (% within each anaemia status group)
  4293. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4294. print(round(100 * prop.table(tbl, margin = 2), 2))
  4295. # 5. run appropriate test
  4296. expected <- chisq.test(tbl)$expected
  4297. if (any(expected < 5)) {
  4298. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4299. print(fisher.test(tbl))
  4300. } else {
  4301. cat("\nChi-squared Test:\n")
  4302. print(chisq.test(tbl))
  4303. }
  4304. ```
  4305. *Prenatal Alcohol Exposure (PAE)*
  4306. ```{r}
  4307. # pae (full sample)
  4308. # unique subject IDs: based on true prevalence
  4309. ULF_ida_final %>%
  4310. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4311. distinct(subject_id, .keep_all = TRUE) %>%
  4312. group_by(maternal_pae) %>%
  4313. summarise(n = n()) %>%
  4314. mutate(percent = round(100 * n / sum(n), 2))
  4315. ```
  4316. ```{r}
  4317. # group differences in pae by child anaemia status
  4318. # chi-squared: pae
  4319. # unique subject IDs: based on true prevalence
  4320. # 1. filter for non-missing child anaemia status and keep unique subjects
  4321. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4322. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4323. # 2. create contingency table: maternal pae vs child anaemia
  4324. tbl <- table(filtered$maternal_pae, filtered$overall_child_anaemia)
  4325. # 3. print contingency table
  4326. cat("Contingency Table: Maternal PAE by Child Anaemia Status\n")
  4327. print(tbl)
  4328. # 4. column percentages (% within each child anaemia group)
  4329. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4330. print(round(100 * prop.table(tbl, margin = 2), 2))
  4331. # 5. run appropriate test
  4332. expected <- chisq.test(tbl)$expected
  4333. if (any(expected < 5)) {
  4334. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4335. print(fisher.test(tbl))
  4336. } else {
  4337. cat("\nChi-squared Test:\n")
  4338. print(chisq.test(tbl))
  4339. }
  4340. ```
  4341. *Prenatal Tobacco Exposure (PTE)*
  4342. ```{r}
  4343. # pte (full sample)
  4344. # excluding repeated measures: unique id
  4345. ULF_ida_final %>%
  4346. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4347. distinct(subject_id, .keep_all = TRUE) %>%
  4348. group_by(maternal_pte) %>%
  4349. summarise(n = n()) %>%
  4350. mutate(percent = round(100 * n / sum(n), 2))
  4351. ```
  4352. ```{r}
  4353. # group differences in pte by child anaemia status
  4354. # chi-squared: pte
  4355. # unique subject IDs: based on true prevalence
  4356. # 1. filter for non-missing child anaemia status and keep unique subjects
  4357. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4358. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4359. # 2. create contingency table: maternal PTE vs child anaemia
  4360. tbl <- table(filtered$maternal_pte, filtered$overall_child_anaemia)
  4361. # 3.print contingency table
  4362. cat("Contingency Table: Maternal PTE by Child Anaemia Status\n")
  4363. print(tbl)
  4364. # 4. column percentages (% within each child anaemia group)
  4365. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4366. print(round(100 * prop.table(tbl, margin = 2), 2))
  4367. # 5. run appropriate test
  4368. expected <- chisq.test(tbl)$expected
  4369. if (any(expected < 5)) {
  4370. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4371. print(fisher.test(tbl))
  4372. } else {
  4373. cat("\nChi-squared Test:\n")
  4374. print(chisq.test(tbl))
  4375. }
  4376. ```
  4377. *Maternal HIV*
  4378. ```{r}
  4379. # maternal hiv (full sample)
  4380. # unique subject IDs: based on true prevalence
  4381. ULF_ida_final %>%
  4382. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4383. distinct(subject_id, .keep_all = TRUE) %>%
  4384. group_by(maternal_hiv) %>%
  4385. summarise(n = n()) %>%
  4386. mutate(percent = round(100 * n / sum(n), 2))
  4387. ```
  4388. ```{r}
  4389. # group differences in maternal hiv by child anaemia status
  4390. # chi-squared: maternal hiv
  4391. # unique subject IDs: based on true prevalence
  4392. # 1. filter for non-missing child anaemia status and keep unique subjects
  4393. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4394. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4395. # 2. create contingency table: maternal HIV vs child anaemia
  4396. tbl <- table(filtered$maternal_hiv, filtered$overall_child_anaemia)
  4397. # 3. print contingency table
  4398. cat("Contingency Table: Maternal HIV by Child Anaemia Status\n")
  4399. print(tbl)
  4400. # 4. column percentages (% within each child anaemia group)
  4401. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4402. print(round(100 * prop.table(tbl, margin = 2), 2))
  4403. # 5. run appropriate test
  4404. expected <- chisq.test(tbl)$expected
  4405. if (any(expected < 5)) {
  4406. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4407. print(fisher.test(tbl))
  4408. } else {
  4409. cat("\nChi-squared Test:\n")
  4410. print(chisq.test(tbl))
  4411. }
  4412. ```
  4413. *Maternal Depression*
  4414. ```{r}
  4415. # maternal depression (full sample)
  4416. # unique subject IDs: based on true prevalence
  4417. ULF_ida_final %>%
  4418. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4419. distinct(subject_id, .keep_all = TRUE) %>%
  4420. group_by(maternal_dep) %>%
  4421. summarise(n = n()) %>%
  4422. mutate(percent = round(100 * n / sum(n), 2))
  4423. ```
  4424. ```{r}
  4425. # group differences in maternal depression by child anaemia status
  4426. # chi-squared: maternal depression
  4427. # unique subject IDs: based on true prevalence
  4428. # 1. filter for non-missing child anaemia status and keep unique subjects
  4429. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4430. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4431. # 2. create contingency table: maternal DEP vs child anaemia
  4432. tbl <- table(filtered$maternal_dep, filtered$overall_child_anaemia)
  4433. # 3. print contingency table
  4434. cat("Contingency Table: Maternal DEP by Child Anaemia Status\n")
  4435. print(tbl)
  4436. # 4. column percentages (% within each child anaemia group)
  4437. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4438. print(round(100 * prop.table(tbl, margin = 2), 2))
  4439. # 5. run appropriate test
  4440. expected <- chisq.test(tbl)$expected
  4441. if (any(expected < 5)) {
  4442. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4443. print(fisher.test(tbl))
  4444. } else {
  4445. cat("\nChi-squared Test:\n")
  4446. print(chisq.test(tbl))
  4447. }
  4448. ```
  4449. *Maternal Employment*
  4450. ```{r}
  4451. # maternal employment (full sample)
  4452. # unique subject IDs: based on true prevalence
  4453. ULF_ida_final %>%
  4454. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4455. distinct(subject_id, .keep_all = TRUE) %>%
  4456. group_by(mom_employment_en_dichotomous) %>%
  4457. summarise(n = n()) %>%
  4458. mutate(percent = round(100 * n / sum(n), 2))
  4459. ```
  4460. ```{r}
  4461. # group differences in maternal employment by child anaemia status
  4462. # chi-squared: maternal employment
  4463. # unique subject IDs: based on true prevalence
  4464. # 1. filter for non-missing child anaemia status and keep unique subjects
  4465. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4466. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4467. # 2. create contingency table: maternal employment vs child anaemia
  4468. tbl <- table(filtered$mom_employment_en_dichotomous, filtered$overall_child_anaemia)
  4469. # 3. print contingency table
  4470. cat("Contingency Table: Maternal Employment by Child Anaemia Status\n")
  4471. print(tbl)
  4472. # 4. column percentages (% within each child anaemia group)
  4473. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4474. print(round(100 * prop.table(tbl, margin = 2), 2))
  4475. # 5. run appropriate test
  4476. expected <- chisq.test(tbl)$expected
  4477. if (any(expected < 5)) {
  4478. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4479. print(fisher.test(tbl))
  4480. } else {
  4481. cat("\nChi-squared Test:\n")
  4482. print(chisq.test(tbl))
  4483. }
  4484. ```
  4485. *Maternal Education*
  4486. ```{r}
  4487. # maternal education (full sample)
  4488. # unique subject IDs: based on true prevalence
  4489. ULF_ida_final %>%
  4490. filter(!is.na(overall_child_anaemia)) %>%
  4491. distinct(subject_id, .keep_all = TRUE) %>%
  4492. group_by(mom_edu_en) %>%
  4493. summarise(n = n()) %>%
  4494. mutate(percent = round(100 * n / sum(n), 2))
  4495. ```
  4496. ```{r}
  4497. # group differences in maternal education by child anaemia status
  4498. # chi-squared: maternal education
  4499. # unique subject IDs: based on true prevalence
  4500. # 1. filter for non-missing child anaemia status and keep unique subjects
  4501. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4502. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4503. # 2. create contingency table: maternal education vs child anaemia
  4504. tbl <- table(filtered$mom_edu_en, filtered$overall_child_anaemia)
  4505. # 3. print contingency table
  4506. cat("Contingency Table: Maternal Education by Child Anaemia Status\n")
  4507. print(tbl)
  4508. # 4. column percentages (% within each child anaemia group)
  4509. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4510. print(round(100 * prop.table(tbl, margin = 2), 2))
  4511. # 5. run appropriate test
  4512. expected <- chisq.test(tbl)$expected
  4513. if (any(expected < 5)) {
  4514. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4515. print(fisher.test(tbl))
  4516. } else {
  4517. cat("\nChi-squared Test:\n")
  4518. print(chisq.test(tbl))
  4519. }
  4520. ```
  4521. *Household Income*
  4522. ```{r}
  4523. # household income (full sample)
  4524. # unique subject IDs: based on true prevalence
  4525. ULF_ida_final %>%
  4526. filter(!is.na(overall_child_anaemia)) %>%
  4527. distinct(subject_id, .keep_all = TRUE) %>%
  4528. group_by(income_household_en) %>%
  4529. summarise(n = n()) %>%
  4530. mutate(percent = round(100 * n / sum(n), 2))
  4531. ```
  4532. ```{r}
  4533. # group differences in household income by child anaemia status
  4534. # chi-squared: household income
  4535. # unique subject IDs: based on true prevalence
  4536. # 1. filter for non-missing child anaemia status and keep unique subjects
  4537. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4538. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4539. # 2. create contingency table: household income vs child anaemia
  4540. tbl <- table(filtered$income_household_en, filtered$overall_child_anaemia)
  4541. # 3. print contingency table
  4542. cat("Contingency Table: Household Income by Child Anaemia Status\n")
  4543. print(tbl)
  4544. # 4. column percentages (% within each child anaemia group)
  4545. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4546. print(round(100 * prop.table(tbl, margin = 2), 2))
  4547. # 5. run appropriate test
  4548. expected <- chisq.test(tbl)$expected
  4549. if (any(expected < 5)) {
  4550. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4551. print(fisher.test(tbl))
  4552. } else {
  4553. cat("\nChi-squared Test:\n")
  4554. print(chisq.test(tbl))
  4555. }
  4556. ```
  4557. *Infant HIV*
  4558. ```{r}
  4559. # infant hiv (full sample)
  4560. # unique subject IDs: based on true prevalence
  4561. ULF_ida_final %>%
  4562. filter(!is.na(overall_child_anaemia)) %>%
  4563. distinct(subject_id, .keep_all = TRUE) %>%
  4564. group_by(baby_hiv) %>%
  4565. summarise(n = n()) %>%
  4566. mutate(percent = round(100 * n / sum(n), 2))
  4567. ```
  4568. ```{r}
  4569. # group differences in infant hiv by child anaemia status
  4570. # chi-squared: infant hiv
  4571. # unique subject IDs: based on true prevalence
  4572. # 1. filter for non-missing child anaemia status and keep unique subjects
  4573. filtered <- ULF_ida_final[!is.na(ULF_ida_final$overall_child_anaemia), ]
  4574. filtered <- filtered[!duplicated(filtered$subject_id), ]
  4575. # 2. create contingency table: infant HIV vs child anaemia
  4576. tbl <- table(filtered$baby_hiv, filtered$overall_child_anaemia)
  4577. # 3. print contingency table
  4578. cat("Contingency Table: Infant HIV by Child Anaemia Status\n")
  4579. print(tbl)
  4580. # 4. column percentages (% within each child anaemia group)
  4581. cat("\nColumn Percentages (% of each anaemia status group):\n")
  4582. print(round(100 * prop.table(tbl, margin = 2), 2))
  4583. # 5. run appropriate test
  4584. expected <- chisq.test(tbl)$expected
  4585. if (any(expected < 5)) {
  4586. cat("\nFisher's Exact Test (due to small expected cell counts):\n")
  4587. print(fisher.test(tbl))
  4588. } else {
  4589. cat("\nChi-squared Test:\n")
  4590. print(chisq.test(tbl))
  4591. }
  4592. ```
  4593. - **Continuous Variables**
  4594. ```{r}
  4595. #install packages if needed
  4596. #install.packages('car')
  4597. library (car)
  4598. ```
  4599. *Minimum Haemoglobin Measure*
  4600. ```{r}
  4601. # group differences in minimum Hb by child anaemia status
  4602. # ttest: min Hb
  4603. # unique subject IDs: based on true prevalence
  4604. library(dplyr)
  4605. library(car) # for Levene's Test
  4606. # create cleaned dataset once (avoids repeating code)
  4607. child_data <- ULF_ida_final %>%
  4608. filter(!is.na(overall_child_anaemia)) %>%
  4609. distinct(subject_id, .keep_all = TRUE)
  4610. # summary statistics: min child haemoglobin by child anaemia status
  4611. summary_stats <- child_data %>%
  4612. group_by(overall_child_anaemia) %>%
  4613. summarise(
  4614. n = n(),
  4615. mean_hb = format(round(mean(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
  4616. sd_hb = format(round(sd(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
  4617. min_hb = format(round(min(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
  4618. max_hb = format(round(max(min_child_hb, na.rm = TRUE), 2), nsmall = 2),
  4619. .groups = "drop"
  4620. )
  4621. print(summary_stats)
  4622. # levene’s test for homogeneity of variances
  4623. levene_result <- leveneTest(min_child_hb ~ overall_child_anaemia, data = child_data)
  4624. print(levene_result)
  4625. # extract Levene p-value
  4626. levene_p <- levene_result$`Pr(>F)`[1]
  4627. # run appropriate t-test (full precision output)
  4628. if (levene_p < 0.05) {
  4629. cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
  4630. t_result <- t.test(
  4631. min_child_hb ~ overall_child_anaemia,
  4632. data = child_data,
  4633. var.equal = FALSE
  4634. )
  4635. } else {
  4636. cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
  4637. t_result <- t.test(
  4638. min_child_hb ~ overall_child_anaemia,
  4639. data = child_data,
  4640. var.equal = TRUE
  4641. )
  4642. }
  4643. # print t-test results with full precision
  4644. print(t_result)
  4645. ```
  4646. *Child Age at Minimum Haemoglobin Measurement*
  4647. ```{r}
  4648. # group differences in age at Hb measurement by child anaemia status
  4649. # ttest: age at Hb measurement
  4650. # unique subject IDs: based on true prevalence
  4651. library(dplyr)
  4652. library(car) # for Levene's Test
  4653. # prepare cleaned dataset once
  4654. child_data <- ULF_ida_final %>%
  4655. filter(!is.na(overall_child_anaemia)) %>%
  4656. distinct(subject_id, .keep_all = TRUE)
  4657. # summary statistics: child age at Hb measurement by anaemia status
  4658. summary_stats <- child_data %>%
  4659. group_by(overall_child_anaemia) %>%
  4660. summarise(
  4661. n = n(),
  4662. mean_age = format(round(mean(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
  4663. sd_age = format(round(sd(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
  4664. min_age = format(round(min(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
  4665. max_age = format(round(max(min_hb_age, na.rm = TRUE), 2), nsmall = 2),
  4666. .groups = "drop"
  4667. )
  4668. print(summary_stats)
  4669. # levene’s test for homogeneity of variances
  4670. levene_result <- leveneTest(min_hb_age ~ overall_child_anaemia, data = child_data)
  4671. print(levene_result)
  4672. # extract Levene p-value
  4673. levene_p <- levene_result$`Pr(>F)`[1]
  4674. # run appropriate t-test (full precision output)
  4675. if (levene_p < 0.05) {
  4676. cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
  4677. t_result <- t.test(
  4678. min_hb_age ~ overall_child_anaemia,
  4679. data = child_data,
  4680. var.equal = FALSE
  4681. )
  4682. } else {
  4683. cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
  4684. t_result <- t.test(
  4685. min_hb_age ~ overall_child_anaemia,
  4686. data = child_data,
  4687. var.equal = TRUE
  4688. )
  4689. }
  4690. # Print t-test results in full precision
  4691. print(t_result)
  4692. ```
  4693. *Maternal Age at Enrolment*
  4694. ```{r}
  4695. # maternal age at enrolment for full group
  4696. # unique subject IDs: based on true prevalence
  4697. ULF_ida_final %>%
  4698. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4699. distinct(subject_id, .keep_all = TRUE) %>%
  4700. summarise(
  4701. mean_mom_age = mean(mom_age_en, na.rm = TRUE),
  4702. sd_mom_age = sd(mom_age_en, na.rm = TRUE),
  4703. min_mom_age = min(mom_age_en, na.rm = TRUE),
  4704. max_mom_age = max(mom_age_en, na.rm = TRUE),
  4705. n = n()
  4706. )
  4707. ```
  4708. ```{r}
  4709. # group differences in maternal age at enrolment by child anaemia status
  4710. # ttest: maternal age at enrolment
  4711. # unique subject IDs: based on true prevalence
  4712. library(dplyr)
  4713. library(car) # for Levene's Test
  4714. # deduplicate and filter data once
  4715. child_data <- ULF_ida_final %>%
  4716. filter(!is.na(overall_child_anaemia)) %>%
  4717. distinct(subject_id, .keep_all = TRUE)
  4718. # summary statistics: min child Hb by anaemia status (2 decimals)
  4719. summary_stats <- child_data %>%
  4720. group_by(overall_child_anaemia) %>%
  4721. summarise(
  4722. n = n(),
  4723. mean_hb = format(round(mean(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  4724. sd_hb = format(round(sd(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  4725. min_hb = format(round(min(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  4726. max_hb = format(round(max(mom_age_en, na.rm = TRUE), 2), nsmall = 2),
  4727. .groups = "drop"
  4728. )
  4729. print(summary_stats)
  4730. # levene’s test for homogeneity of variances
  4731. levene_result <- leveneTest(mom_age_en ~ overall_child_anaemia, data = child_data)
  4732. print(levene_result)
  4733. # extract Levene p-value
  4734. levene_p <- levene_result$`Pr(>F)`[1]
  4735. # run t-test (full precision)
  4736. if (levene_p < 0.05) {
  4737. cat("\nLevene's test is significant — using Welch t-test (unequal variances):\n")
  4738. t_result <- t.test(mom_age_en ~ overall_child_anaemia,
  4739. data = child_data,
  4740. var.equal = FALSE)
  4741. } else {
  4742. cat("\nLevene's test is not significant — using Student's t-test (equal variances):\n")
  4743. t_result <- t.test(mom_age_en ~ overall_child_anaemia,
  4744. data = child_data,
  4745. var.equal = TRUE)
  4746. }
  4747. # print t-test results
  4748. print(t_result)
  4749. ```
  4750. *Child Age at Scan*
  4751. ```{r}
  4752. # child age at scan (full sample)
  4753. # include repeated measures - for all scans across study timepoints
  4754. ULF_ida_final %>%
  4755. filter(!is.na(cor_child_anaemia)) %>% # only include PIDs with child anaemia data
  4756. summarise(
  4757. mean_child_age = mean(scan_age_months, na.rm = TRUE),
  4758. sd_child_age = sd(scan_age_months, na.rm = TRUE),
  4759. min_child_age = min(scan_age_months, na.rm = TRUE),
  4760. max_child_age = max(scan_age_months, na.rm = TRUE),
  4761. n = n()
  4762. )
  4763. ```
  4764. ```{r}
  4765. # group differences in child age at scan by child anaemia status
  4766. # ttest: child age at scan
  4767. # all scans included
  4768. library(dplyr)
  4769. library(car)
  4770. # deduplicate and filter data once
  4771. child_data <- ULF_ida_final %>%
  4772. filter(!is.na(cor_child_anaemia))
  4773. # summary statistics: scan age by child anaemia status (2 decimals)
  4774. summary_stats <- child_data %>%
  4775. group_by(cor_child_anaemia) %>%
  4776. summarise(
  4777. n = n(),
  4778. mean_age = format(round(mean(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  4779. sd_age = format(round(sd(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  4780. min_age = format(round(min(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  4781. max_age = format(round(max(scan_age_months, na.rm = TRUE), 2), nsmall = 2),
  4782. .groups = "drop"
  4783. )
  4784. print(summary_stats)
  4785. # levene's Test for homogeneity of variance
  4786. levene_result <- leveneTest(scan_age_months ~ cor_child_anaemia, data = child_data)
  4787. print(levene_result)
  4788. # extract Levene p-value
  4789. levene_p <- levene_result$`Pr(>F)`[1]
  4790. # automatically choose correct t-test
  4791. if (levene_p < 0.05) {
  4792. cat("\nLevene's test significant (p =", round(levene_p, 4), ") → Using Welch t-test (var.equal = FALSE)\n\n")
  4793. t_result <- t.test(scan_age_months ~ cor_child_anaemia,
  4794. data = child_data,
  4795. var.equal = FALSE)
  4796. } else {
  4797. cat("\nLevene's test not significant (p =", round(levene_p, 4), ") → Using Student t-test (var.equal = TRUE)\n\n")
  4798. t_result <- t.test(scan_age_months ~ cor_child_anaemia,
  4799. data = child_data,
  4800. var.equal = TRUE)
  4801. }
  4802. # print t-test results in full precision
  4803. print(t_result)
  4804. ```
  4805. *Child Gestational Age (GA) at Birth*
  4806. ```{r}
  4807. # child ga at birth (full sample)
  4808. # unique subject IDs: based on true prevalence
  4809. ULF_ida_final %>%
  4810. filter(!is.na(overall_child_anaemia)) %>% # only include PIDs with child anaemia data
  4811. distinct(subject_id, .keep_all = TRUE) %>%
  4812. summarise(
  4813. mean_child_ga = mean(ga_weeks, na.rm = TRUE),
  4814. sd_child_ga = sd(ga_weeks, na.rm = TRUE),
  4815. min_child_ga = min(ga_weeks, na.rm = TRUE),
  4816. max_child_ga = max(ga_weeks, na.rm = TRUE),
  4817. n = n()
  4818. )
  4819. ```
  4820. ```{r}
  4821. # group differences in ga at birth by child anaemia status
  4822. # ttest: ga
  4823. # unique subject IDs: based on true prevalence
  4824. library(dplyr)
  4825. library(car) # for Levene's Test
  4826. # deduplicate and filter once
  4827. child_data <- ULF_ida_final %>%
  4828. filter(!is.na(overall_child_anaemia)) %>%
  4829. distinct(subject_id, .keep_all = TRUE)
  4830. # summary statistics: GA by child anaemia status (2 decimals)
  4831. summary_stats <- child_data %>%
  4832. group_by(overall_child_anaemia) %>%
  4833. summarise(
  4834. n = n(),
  4835. mean_ga = format(round(mean(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  4836. sd_ga = format(round(sd(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  4837. min_ga = format(round(min(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  4838. max_ga = format(round(max(ga_weeks, na.rm = TRUE), 2), nsmall = 2),
  4839. .groups = "drop"
  4840. )
  4841. print(summary_stats)
  4842. # levene's Test for homogeneity of variance
  4843. levene_result <- leveneTest(ga_weeks ~ overall_child_anaemia, data = child_data)
  4844. print(levene_result)
  4845. # extract Levene p-value
  4846. levene_p <- levene_result$`Pr(>F)`[1]
  4847. # automatically choose correct t-test
  4848. if (levene_p < 0.05) {
  4849. cat("\nLevene's test significant (p =", round(levene_p, 4),
  4850. ") → Using Welch t-test (var.equal = FALSE)\n\n")
  4851. t_result <- t.test(
  4852. ga_weeks ~ overall_child_anaemia,
  4853. data = child_data,
  4854. var.equal = FALSE
  4855. )
  4856. } else {
  4857. cat("\nLevene's test not significant (p =", round(levene_p, 4),
  4858. ") → Using Student t-test (var.equal = TRUE)\n\n")
  4859. t_result <- t.test(
  4860. ga_weeks ~ overall_child_anaemia,
  4861. data = child_data,
  4862. var.equal = TRUE
  4863. )
  4864. }
  4865. # print t-test result in full precision
  4866. print(t_result)
  4867. ```
  4868. *Infant Birth Weight (BW)*
  4869. ```{r}
  4870. # child bw (full sample)
  4871. # unique subject IDs: based on true prevalence
  4872. ULF_ida_final %>%
  4873. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4874. distinct(subject_id, .keep_all = TRUE) %>%
  4875. summarise(
  4876. mean_bw = mean(bw, na.rm = TRUE),
  4877. sd_bw = sd(bw, na.rm = TRUE),
  4878. min_bw = min(bw, na.rm = TRUE),
  4879. max_bw = max(bw, na.rm = TRUE),
  4880. n = n()
  4881. )
  4882. ```
  4883. ```{r}
  4884. # group differences in bw by child anaemia status
  4885. # ttest: bw
  4886. # unique subject IDs: based on true prevalence
  4887. library(dplyr)
  4888. library(car)
  4889. # deduplicate and filter once
  4890. birth_data <- ULF_ida_final %>%
  4891. filter(!is.na(overall_child_anaemia)) %>%
  4892. distinct(subject_id, .keep_all = TRUE)
  4893. # summary statistics: birth weight by child anaemia status (2 decimals)
  4894. summary_stats <- birth_data %>%
  4895. group_by(overall_child_anaemia) %>%
  4896. summarise(
  4897. n = n(),
  4898. mean_bw = format(round(mean(bw, na.rm = TRUE), 2), nsmall = 2),
  4899. sd_bw = format(round(sd(bw, na.rm = TRUE), 2), nsmall = 2),
  4900. min_bw = format(round(min(bw, na.rm = TRUE), 2), nsmall = 2),
  4901. max_bw = format(round(max(bw, na.rm = TRUE), 2), nsmall = 2),
  4902. .groups = "drop"
  4903. )
  4904. print(summary_stats)
  4905. # levene's test for homogeneity of variance
  4906. levene_result <- leveneTest(bw ~ overall_child_anaemia, data = birth_data)
  4907. print(levene_result)
  4908. # extract p-value
  4909. levene_p <- levene_result$`Pr(>F)`[1]
  4910. # automatically choose correct t-test
  4911. if (levene_p < 0.05) {
  4912. cat("\nLevene's test significant (p =", round(levene_p, 4),
  4913. ") → Using Welch t-test (var.equal = FALSE)\n\n")
  4914. t_test_result <- t.test(
  4915. bw ~ overall_child_anaemia,
  4916. data = birth_data,
  4917. var.equal = FALSE
  4918. )
  4919. } else {
  4920. cat("\nLevene's test not significant (p =", round(levene_p, 4),
  4921. ") → Using Student t-test (var.equal = TRUE)\n\n")
  4922. t_test_result <- t.test(
  4923. bw ~ overall_child_anaemia,
  4924. data = birth_data,
  4925. var.equal = TRUE
  4926. )
  4927. }
  4928. # print t-test result in full precision
  4929. print(t_test_result)
  4930. ```
  4931. *Infant Birth Length (BL)*
  4932. ```{r}
  4933. # child bl (full sample)
  4934. # unique subject IDs: based on true prevalence
  4935. ULF_ida_final %>%
  4936. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  4937. distinct(subject_id, .keep_all = TRUE) %>%
  4938. summarise(
  4939. mean_birthlength = mean(birth_length, na.rm = TRUE),
  4940. sd_birthlength = sd(birth_length, na.rm = TRUE),
  4941. min_birthlength = min(birth_length, na.rm = TRUE),
  4942. max_birthlength = max(birth_length, na.rm = TRUE),
  4943. n = n()
  4944. )
  4945. ```
  4946. ```{r}
  4947. # group differences in bl by child anaemia status
  4948. # ttest: bl
  4949. # unique subject IDs: based on true prevalence
  4950. library(dplyr)
  4951. library(car)
  4952. # deduplicate and filter once
  4953. birth_length_data <- ULF_ida_final %>%
  4954. filter(!is.na(overall_child_anaemia)) %>%
  4955. distinct(subject_id, .keep_all = TRUE)
  4956. # summary statistics: birth length by child anaemia status (2 decimals)
  4957. summary_stats <- birth_length_data %>%
  4958. group_by(overall_child_anaemia) %>%
  4959. summarise(
  4960. n = n(),
  4961. mean_bl = format(round(mean(birth_length, na.rm = TRUE), 2), nsmall = 2),
  4962. sd_bl = format(round(sd(birth_length, na.rm = TRUE), 2), nsmall = 2),
  4963. min_bl = format(round(min(birth_length, na.rm = TRUE), 2), nsmall = 2),
  4964. max_bl = format(round(max(birth_length, na.rm = TRUE), 2), nsmall = 2),
  4965. .groups = "drop"
  4966. )
  4967. print(summary_stats)
  4968. # levene's test for homogeneity of variance
  4969. levene_result <- leveneTest(birth_length ~ overall_child_anaemia, data = birth_length_data)
  4970. print(levene_result)
  4971. # extract p-value
  4972. levene_p <- levene_result$`Pr(>F)`[1]
  4973. # automatically choose correct t-test
  4974. if (levene_p < 0.05) {
  4975. cat("\nLevene's test significant (p =", round(levene_p, 4),
  4976. ") → Using Welch t-test (var.equal = FALSE)\n\n")
  4977. t_test_result <- t.test(
  4978. birth_length ~ overall_child_anaemia,
  4979. data = birth_length_data,
  4980. var.equal = FALSE
  4981. )
  4982. } else {
  4983. cat("\nLevene's test not significant (p =", round(levene_p, 4),
  4984. ") → Using Student t-test (var.equal = TRUE)\n\n")
  4985. t_test_result <- t.test(
  4986. birth_length ~ overall_child_anaemia,
  4987. data = birth_length_data,
  4988. var.equal = TRUE
  4989. )
  4990. }
  4991. # print t-test result in full precision
  4992. print(t_test_result)
  4993. ```
  4994. *Infant Head Circumference (HC)*
  4995. ```{r}
  4996. # child birth hc (full sample)
  4997. # unique subject IDs: based on true prevalence
  4998. ULF_ida_final %>%
  4999. filter(!is.na(overall_child_anaemia)) %>% #only include PIDs with child anaemia data
  5000. distinct(subject_id, .keep_all = TRUE) %>%
  5001. summarise(
  5002. mean_hc = mean(birth_hc, na.rm = TRUE),
  5003. sd_hc = sd(birth_hc, na.rm = TRUE),
  5004. min_hc = min(birth_hc, na.rm = TRUE),
  5005. max_hc = max(birth_hc, na.rm = TRUE),
  5006. n = n()
  5007. )
  5008. ```
  5009. ```{r}
  5010. # group differences in hc at birth by child anaemia status
  5011. # ttest: hc at birth
  5012. # unique subject IDs: based on true prevalence
  5013. library(dplyr)
  5014. library(car)
  5015. # deduplicate and filter once
  5016. birth_hc_data <- ULF_ida_final %>%
  5017. filter(!is.na(overall_child_anaemia)) %>%
  5018. distinct(subject_id, .keep_all = TRUE)
  5019. # summary statistics: birth head circumference by child anaemia status (2 decimals)
  5020. summary_stats <- birth_hc_data %>%
  5021. group_by(overall_child_anaemia) %>%
  5022. summarise(
  5023. n = n(),
  5024. mean_hc = format(round(mean(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  5025. sd_hc = format(round(sd(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  5026. min_hc = format(round(min(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  5027. max_hc = format(round(max(birth_hc, na.rm = TRUE), 2), nsmall = 2),
  5028. .groups = "drop"
  5029. )
  5030. print(summary_stats)
  5031. # levene's test for homogeneity of variance
  5032. levene_result <- leveneTest(birth_hc ~ overall_child_anaemia, data = birth_hc_data)
  5033. print(levene_result)
  5034. # extract p-value
  5035. levene_p <- levene_result$`Pr(>F)`[1]
  5036. # automatically choose correct t-test
  5037. if (levene_p < 0.05) {
  5038. cat("\nLevene's test significant (p =", round(levene_p, 4),
  5039. ") → Using Welch t-test (var.equal = FALSE)\n\n")
  5040. t_test_result <- t.test(
  5041. birth_hc ~ overall_child_anaemia,
  5042. data = birth_hc_data,
  5043. var.equal = FALSE
  5044. )
  5045. } else {
  5046. cat("\nLevene's test not significant (p =", round(levene_p, 4),
  5047. ") → Using Student t-test (var.equal = TRUE)\n\n")
  5048. t_test_result <- t.test(
  5049. birth_hc ~ overall_child_anaemia,
  5050. data = birth_hc_data,
  5051. var.equal = TRUE
  5052. )
  5053. }
  5054. # print t-test result in full precision
  5055. print(t_test_result)
  5056. ```
  5057. **Exploratory Analyses: Plotted Data using LOESS Curves**
  5058. *LOESS Curve: ICV*
  5059. ```{r}
  5060. # icv
  5061. library(ggplot2)
  5062. library(dplyr)
  5063. icv_ulf_child <- ggplot(
  5064. data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
  5065. aes(
  5066. x = scan_age_months,
  5067. y = mm_hyp_icv,
  5068. color = cor_child_anaemia,
  5069. fill = cor_child_anaemia
  5070. )
  5071. ) +
  5072. geom_point() +
  5073. geom_smooth(
  5074. method = "loess",
  5075. se = TRUE,
  5076. level = 0.95,
  5077. alpha = 0.3 # semi-transparent ribbon
  5078. ) +
  5079. scale_color_manual(
  5080. values = c(
  5081. "no_child_anaemia" = "#1f77b4", # dark blue line
  5082. "child_anaemia" = "#d62728" # dark red line
  5083. ),
  5084. name = "Child Anaemia Status"
  5085. ) +
  5086. scale_fill_manual(
  5087. values = c(
  5088. "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
  5089. "child_anaemia" = "#f7b0b0" # lighter red ribbon
  5090. ),
  5091. guide = "none" # hide fill legend
  5092. ) +
  5093. scale_x_continuous(
  5094. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 10)
  5095. ) +
  5096. labs(
  5097. title = "Total ICV Volume vs. Scan Age",
  5098. x = "Scan Age (Months)",
  5099. y = "Total ICV Volume (mm3)",
  5100. color = "Child Anaemia Status"
  5101. ) +
  5102. theme_minimal()
  5103. icv_ulf_child
  5104. ```
  5105. *LOESS Curve: Corpus Callosum*
  5106. ```{r}
  5107. # corpus callosum
  5108. cc_ulf_child <- ggplot(
  5109. data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
  5110. aes(
  5111. x = scan_age_months,
  5112. y = total_cc,
  5113. color = cor_child_anaemia,
  5114. fill = cor_child_anaemia
  5115. )
  5116. ) +
  5117. geom_point() +
  5118. geom_smooth(
  5119. method = "loess",
  5120. se = TRUE, # show 95% confidence ribbons
  5121. level = 0.95,
  5122. alpha = 0.3 # make ribbons semi-transparent
  5123. ) +
  5124. scale_color_manual(
  5125. values = c(
  5126. "no_child_anaemia" = "#1f77b4", # dark blue
  5127. "child_anaemia" = "#d62728" # dark red
  5128. ),
  5129. name = "Child Anaemia Status"
  5130. ) +
  5131. scale_fill_manual(
  5132. values = c(
  5133. "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
  5134. "child_anaemia" = "#f7b0b0" # lighter red ribbon
  5135. ),
  5136. guide = "none" # hide fill legend
  5137. ) +
  5138. scale_x_continuous(
  5139. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  5140. ) +
  5141. scale_y_continuous(
  5142. breaks = seq(
  5143. floor(min(ULF_ida_final$total_cc, na.rm = TRUE) / 10) * 10,
  5144. ceiling(max(ULF_ida_final$total_cc, na.rm = TRUE) / 10) * 10,
  5145. by = 4000
  5146. )
  5147. ) +
  5148. labs(
  5149. title = "Total CC Volume vs. Scan Age",
  5150. x = "Scan Age (Months)",
  5151. y = "Total CC Volume (mm3)",
  5152. color = "Child Anaemia Status"
  5153. ) +
  5154. theme_minimal()
  5155. cc_ulf_child
  5156. ```
  5157. *LOESS Curve: Caudate Nucleus*
  5158. ```{r}
  5159. # caudate nucleus
  5160. caudate_ulf_child <- ggplot(
  5161. data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
  5162. aes(
  5163. x = scan_age_months,
  5164. y = total_caudate,
  5165. color = cor_child_anaemia,
  5166. fill = cor_child_anaemia
  5167. )
  5168. ) +
  5169. geom_point() +
  5170. geom_smooth(
  5171. method = "loess",
  5172. se = TRUE, # show 95% confidence ribbons
  5173. level = 0.95,
  5174. alpha = 0.3 # semi-transparent ribbons
  5175. ) +
  5176. scale_color_manual(
  5177. values = c(
  5178. "no_child_anaemia" = "#1f77b4", # dark blue
  5179. "child_anaemia" = "#d62728" # dark red
  5180. ),
  5181. name = "Child Anaemia Status"
  5182. ) +
  5183. scale_fill_manual(
  5184. values = c(
  5185. "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
  5186. "child_anaemia" = "#f7b0b0" # lighter red ribbon
  5187. ),
  5188. guide = "none" # hide fill legend
  5189. ) +
  5190. scale_x_continuous(
  5191. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  5192. ) +
  5193. scale_y_continuous(
  5194. n.breaks = 5 # adjust y-axis ticks
  5195. ) +
  5196. labs(
  5197. title = "Total Caudate Volume vs. Scan Age",
  5198. x = "Scan Age (Months)",
  5199. y = "Total Caudate Volume (mm³)",
  5200. color = "Child Anaemia Status"
  5201. ) +
  5202. theme_minimal()
  5203. caudate_ulf_child
  5204. ```
  5205. *LOESS Curve: Putamen*
  5206. ```{r}
  5207. # putamen
  5208. putamen_ulf_child <- ggplot(
  5209. data = ULF_ida_final %>% filter(!is.na(cor_child_anaemia)),
  5210. aes(
  5211. x = scan_age_months,
  5212. y = total_putamen,
  5213. color = cor_child_anaemia,
  5214. fill = cor_child_anaemia
  5215. )
  5216. ) +
  5217. geom_point() +
  5218. geom_smooth(
  5219. method = "loess",
  5220. se = TRUE, # show 95% confidence ribbons
  5221. level = 0.95,
  5222. alpha = 0.3 # semi-transparent ribbons
  5223. ) +
  5224. scale_color_manual(
  5225. values = c(
  5226. "no_child_anaemia" = "#1f77b4", # dark blue
  5227. "child_anaemia" = "#d62728" # dark red
  5228. ),
  5229. name = "Child Anaemia Status"
  5230. ) +
  5231. scale_fill_manual(
  5232. values = c(
  5233. "no_child_anaemia" = "#c6d9f1", # lighter blue ribbon
  5234. "child_anaemia" = "#f7b0b0" # lighter red ribbon
  5235. ),
  5236. guide = "none" # hide fill legend
  5237. ) +
  5238. scale_x_continuous(
  5239. breaks = seq(0, max(ULF_ida_final$scan_age_months, na.rm = TRUE), by = 3)
  5240. ) +
  5241. labs(
  5242. title = "Total Putamen Volume vs. Scan Age",
  5243. x = "Scan Age (Months)",
  5244. y = "Total Putamen Volume (mm³)",
  5245. color = "Child Anaemia Status"
  5246. ) +
  5247. theme_minimal()
  5248. putamen_ulf_child
  5249. ```
  5250. **Linear Mixed Effects (LME) Models: Child Anaemia Status Only**
  5251. First, the LMEs described above for maternal anaemia were replicated with postnatal child anaemia as the primary predictor of interest. This series of models were run to assess the independent associations between child anaemia and regional child brain volumes.
  5252. Based on the LOESS curves, there were no observed age-related differences between groups based on child anaemia status. Therefore, interactions were excluded to simplify the models.
  5253. --\> Main effects only (with child anaemia as the primary predictor)
  5254. Log (age) was used for all models based on non-linear growth trajectories
  5255. ```{r}
  5256. # lme: icv
  5257. # main effects: child anaemia
  5258. # log (age)
  5259. # model summary
  5260. ulf_model_icv_child <- lmer(mm_hyp_icv ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
  5261. data = ULF_ida_final)
  5262. summary(ulf_model_icv_child)
  5263. # beta values
  5264. standardize_parameters(ulf_model_icv_child)
  5265. ```
  5266. ```{r}
  5267. # lme: cc
  5268. # main effects: child anaemia
  5269. # log (age)
  5270. # model summary
  5271. ulf_model_cc_child <- lmer(total_cc ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id), data = ULF_ida_final)
  5272. summary(ulf_model_cc_child)
  5273. # beta values
  5274. standardize_parameters(ulf_model_cc_child)
  5275. ```
  5276. ```{r}
  5277. # putamen
  5278. # main effects: child anaemia
  5279. # log (age)
  5280. # model summary
  5281. ulf_model_putamen_child <- lmer(total_putamen ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
  5282. data = ULF_ida_final)
  5283. summary(ulf_model_putamen_child)
  5284. # beta values
  5285. standardize_parameters(ulf_model_putamen_child)
  5286. ```
  5287. ```{r}
  5288. # lme: caudate nucleus
  5289. # main effects: child anaemia
  5290. # log (age)
  5291. # model summary
  5292. ulf_model_caudate_child <- lmer(total_caudate ~ cor_child_anaemia + log_age + child_sex + (1 | subject_id),
  5293. data = ULF_ida_final)
  5294. summary(ulf_model_caudate_child)
  5295. #to get beta values
  5296. standardize_parameters(ulf_model_caudate_child)
  5297. ```
  5298. **Linear Mixed Effects (LME) Models: Maternal Anaemia and Child Anaemia**
  5299. Second, the final models from the primary analyses assessing the impact of antenatal maternal anaemia status on regional brain volumes were repeated with the addition of postnatal child anaemia status as a covariate. This assessed the relative contributions of antenatal versus postnatal exposures, and examined whether the observed associations between maternal anaemia and regional child brain volumes remained after controlling for child anaemia.
  5300. Log (age) was used for all models based on non-linear growth trajectories
  5301. ```{r}
  5302. # lme: ICV
  5303. # predictors: maternal anaemia and child anaemia
  5304. # based on final maternal anaemia model (primary analysis) for icv: exclude interaction between maternal anaemia and age
  5305. # model summary
  5306. ulf_model_icv_mat_child <- lmer(mm_hyp_icv ~ maternal_preg_anaemia_status + log_age + child_sex + cor_child_anaemia + (1 | subject_id),
  5307. data = ULF_ida_final)
  5308. summary(ulf_model_icv_mat_child)
  5309. # beta values
  5310. standardize_parameters(ulf_model_icv_mat_child)
  5311. ```
  5312. ```{r}
  5313. # lme: corpus callosum
  5314. # predictors: maternal anaemia and child anaemia
  5315. # based on final maternal anaemia model (primary analysis) for corpus callosum : include interaction between maternal anaemia and age
  5316. # model summary
  5317. ulf_model_cc_mat_child <- lmer(total_cc ~ maternal_preg_anaemia_status * log_age + child_sex + cor_child_anaemia + (1 | subject_id), data = ULF_ida_final)
  5318. summary(ulf_model_cc_mat_child)
  5319. # beta values
  5320. standardize_parameters(ulf_model_cc_mat_child)
  5321. ```
  5322. ```{r}
  5323. # lme: putamen
  5324. # predictors: maternal anaemia and child anaemia
  5325. # based on final maternal anaemia model (primary analysis) for putamen: exclude interaction between maternal anaemia and age
  5326. # model summary
  5327. ulf_model_putamen_mat_child <- lmer(total_putamen ~ maternal_preg_anaemia_status + log_age + child_sex + cor_child_anaemia + (1 | subject_id),
  5328. data = ULF_ida_final)
  5329. summary(ulf_model_putamen_mat_child)
  5330. # beta values
  5331. standardize_parameters(ulf_model_putamen_mat_child)
  5332. ```
  5333. ```{r}
  5334. # lme: caudate nucleus
  5335. # predictors: maternal anaemia and child anaemia
  5336. # based on final maternal anaemia model (primary analysis) for caudate nucleus: include interaction between maternal anaemia and age
  5337. # model summary
  5338. ulf_model_caudate_mat_child <- lmer(total_caudate ~ maternal_preg_anaemia_status * log_age + child_sex + cor_child_anaemia + (1 | subject_id),
  5339. data = ULF_ida_final)
  5340. summary(ulf_model_caudate_mat_child)
  5341. # beta values
  5342. standardize_parameters(ulf_model_caudate_mat_child)
  5343. ```
  5344. ------------------------------------------------------------------------
  5345. **Tertiary Analysis: Maternal Iron Deficiency (ULF sample)**
  5346. ------------------------------------------------------------------------
  5347. **Sample Size: For infants with Maternal ID Data**
  5348. ```{r}
  5349. # total sample size
  5350. # including repeated measures
  5351. ULF_ida_final %>%
  5352. filter(!is.na(maternal_preg_adj_brinda_id)) %>% #only include PIDs with antenatal maternal ID data
  5353. summarise(n = n())
  5354. ```
  5355. ```{r}
  5356. # total sample size
  5357. # Excluding repeated measures: unique id
  5358. # unique

Khula_AnaemiaImagingAnalysis.qmd at commit b50e304, no license · at the source

Overview

Authors: Jessica E Ringshaw1,2,3, Michal R Zieff1, Niall J Bourke3, Chiara Casella4,5, Layla E Bradford1,2, Simone R Williams1,2, Donna Herr1, Marlie Miles1,2, Carly Bennallick3, Khula South Africa Data Collection Team, Sean Deoni6, Jonathan O’Muircheartaigh4,5,7, Dan J Stein2,8, Daniel C Alexander9, Derek K Jones10, Steven C R Williams3, Kirsten A Donald1,2
  1. Department of Paediatrics and Child Health, Red Cross War Memorial Children’s Hospital, University of Cape Town, Cape Town, 7700, South Africa
  2. Neuroscience Institute, University of Cape Town, Cape Town, 7925, South Africa
  3. Centre for Neuroimaging Sciences, Department of Neuroimaging, King's College London, London, England, UK
  4. Research Department of Early Life Imaging, School of Biomedical Engineering and Imaging Sciences, King’s College London, London, England, UK
  5. Department for Forensic and Neurodevelopmental Sciences, Institute of Psychiatry, Psychology and Neuroscience, King’s College London, London, England, UK
  6. Maternal, Newborn, Child Nutrition and Health (MNCH) Discovery and Translational (D&T) Sciences Program, Gates Foundation, Seattle, USA
  7. Medical Research Council (MRC) Centre for Neurodevelopmental Disorders, King’s College London, London, England, UK
  8. South African Medical Research Council (SAMRC), Unit on Risk and Resilience in Mental Disorders, Department of Psychiatry, University of Cape Town, Cape Town, 7925, South Africa
  9. Hawkes Institute and Department of Computer Science, University College London, London, England, UK
  10. Cardiff University Brain Imaging Research Centre (CUBRIC), School of Psychology, Cardiff University, Cardiff, Wales, UK
Institutions: University of Cape Town (South Africa); King's College London (United Kingdom); Red Cross War Memorial Children's Hospital (South Africa); Gates Foundation (United States); MRC Centre for Neurodevelopmental Disorders (United Kingdom); Medical Research Council (United Kingdom); South African Medical Research Council (South Africa); University College London (United Kingdom); Cardiff University (United Kingdom)
Journal: Brain communications, volume 8, issue 5, article fcag318
Dates: received 8 December 2025; accepted 4 August 2026; published online 8 September 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1093/braincomms/fcag318 · PMID 42713352 · PMCID PMC13553114 · OpenAlex W7212063535
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), developmental (subfield)
Methods: Connectivity, Statistics, Preprocessing, fMRI & imaging
Keywords: neuroimaging, high-field MRI, ultra-low-field MRI, antenatal maternal anaemia, child brain structure
Topic: Advanced MRI Techniques and Applications (Radiology, Nuclear Medicine and Imaging, Medicine), according to OpenAlex
Funding: National Institute for Health and Care Research; Bill & Melinda Gates Foundation (INV-023509); Gates Foundation (INV-090982, INV-047888, INV-023509); UK Foreign, Commonwealth, and Development Office; Wellcome Trust Research Enrichment Public Engagement (224287/Z/21/A); Wellcome Trust International Training Fellowship (224287/Z/21/Z); Developing Excellence in Leadership, Training, and Science in Africa (Del-22-002); European and Developing Countries Clinical Trials Partnership 2; Wellcome Leap 1kD; Wellcome Trust (224287/Z/21/Z, 224287/Z/21/A); The First 1000 Days (222076/Z/20/Z); European Union; Maudsley Biomedical Research Centre; Science for Africa Foundation
Citations: not cited yet (Europe PMC); 57 references in the paper

Abstract

With the evolution of ultra-low-field MRI and the recognition of antenatal maternal anaemia as an important driver of altered neurodevelopment in toddlers and children, it is critical to determine whether these effects are detectable at ultra-low-field (64 mT) in infancy. The aim of this study was to assess the impact of antenatal maternal anaemia on infant brain structure across the first 2 years of life, using high-field (3 T) and ultra-low-field (64 mT) MRI. This neuroimaging substudy was embedded within Khula, an observational population-based birth cohort in South Africa. Pregnant women were enrolled antenatally and postnatally. Mother-child dyads (n = 394) were followed prospectively with a subsample attending neuroimaging at ∼3, 6, 12, 18 and 24 months of age. Anaemia was classified using World Health Organization thresholds, and neuroimaging data were processed using MiniMORPH. Linear mixed-effects models were used to investigate associations between antenatal maternal anaemia status and absolute regional infant brain volumes using high-field and ultra-low-field MRI. In repeated measures high-field (n = 195) and ultra-low-field (n = 341) infant neuroimaging subsamples, the prevalence of antenatal maternal anaemia was 28.24% (37/131) and 29.76% (61/205), respectively. Maternal anaemia in pregnancy was associated with altered child brain structure across both MRI systems, with group differences becoming detectable at ∼12 months. In the ultra-low-field subsample, infants born to anaemic mothers had 3.77% smaller intracranial volume (β = −0.24, P = 0.004) and 3.32% smaller putamen volumes (β = −0.18, P = 0.040) across the first 2 years of life. The interaction between antenatal maternal anaemia and age was significant for the caudate nucleus (β = −0.13, P = 0.038) and corpus callosum (β = −0.15, P = 0.007). Antenatal maternal anaemia was associated with 3.70% and 4.29% smaller caudate nucleus volumes at 18 and 24 months of age, respectively. Similarly, infants born to anaemic mothers had 4.39% smaller corpus callosum volumes by 12 months and 6.27% smaller corpus callosum volumes by 24 months. Postnatal child anaemia and antenatal maternal iron deficiency status were not associated with total or regional child brain volumes in the ultra-low-field subsample from this cohort. Maternal anaemia remained a robust predictor of volume differences in sensitivity analyses. This study is the first to demonstrate that the impact of maternal anaemia in pregnancy on child brain structure is detectable as early as infancy. The implications of this research are 2-fold: (i) informing the feasibility of ultra-low-field MRI in low- and middle-income countries and (ii) the timing and optimization of targeted recommendations for anaemia management in practice and policy.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 9 matches between paragraphs and lines of code.

UNITY-Physics

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: the link answers
Software Heritage: not checked
Found in: the text, “Neuroimaging processing”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 26 September 2026: the link answers (HTTP 200)
  • 26 September 2026: the link answers (HTTP 200)

UNITY-Physics/fw-minimorph

License: MIT
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 08c7398e42c88031da40ae84941214aa6b4a0299, 7 April 2026
Languages: Python (8), Shell (3)
Size: 35 files, 11 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (Dockerfile), tests, documentation
Not found: CITATION.cff, continuous integration
Tools: FSL (2 files), ANTs (1 file), FreeSurfer (1 file), Matplotlib (1 file), NiBabel (1 file), NumPy (1 file), pandas (1 file), Pillow (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
13 files

JERingshaw/Khula-Anaemia-Imaging-Analysis

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: b50e304453da1c77f7617057f487c5c3334a68b8, 24 March 2026
Languages: Quarto (1)
Size: 3 files, 1 script
Software Heritage: not archived
Found in: “Data availability”
Holds: README, 1 notebook
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: car (1 file), easystats (1 file), emmeans (1 file), ggplot2 (1 file), lme4 (1 file), lmerTest (1 file), tidyverse (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
2 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 12 scripts, each with its path and the digest of its content;
  • 9 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The de-identified data that support the findings of this study are available from the authors upon reasonable request as per the Khula South Africa study guidelines. All software used in the development of Flywheel gears for the UNITY project has been shared as an open-source resource on GitHub (https://github.com/UNITY-Physics). This includes the MiniMORPH processing pipeline (https://github.com/UNITY-Physics/fw-minimorph) and the R script used for statistical analyses (https://github.com/JERingshaw/Khula-Anaemia-Imaging-Analysis.git)

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 17 authors, 5 keywords, 14 funders, 48 references.

Cite

This paper

Ringshaw, J. E., Zieff, M. R., Bourke, N. J., Casella, C., Bradford, L. E., Williams, S. R., Herr, D., Miles, M., Bennallick, C., Khula South Africa Data Collection Team, Deoni, S., O’Muircheartaigh, J., Stein, D. J., Alexander, D. C., Jones, D. K., Williams, S. C. R., & Donald, K. A. (2026). Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa. Brain communications, 8(5), fcag318. https://doi.org/10.1093/braincomms/fcag318

BibTeX

@article{ringshaw2026antenatal,
author = {Ringshaw, Jessica E and Zieff, Michal R and Bourke, Niall J and Casella, Chiara and Bradford, Layla E and Williams, Simone R and Herr, Donna and Miles, Marlie and Bennallick, Carly and {Khula South Africa Data Collection Team} and Deoni, Sean and O’Muircheartaigh, Jonathan and Stein, Dan J and Alexander, Daniel C and Jones, Derek K and Williams, Steven C R and Donald, Kirsten A},
title = {{Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa}},
journal = {Brain communications},
year = {2026},
month = sep,
volume = {8},
number = {5},
pages = {fcag318},
publisher = {Oxford University Press},
issn = {2632-1297},
doi = {10.1093/braincomms/fcag318},
url = {https://doi.org/10.1093/braincomms/fcag318},
pmid = {42713352},
pmcid = {PMC13553114}
}

RIS

TY - JOUR
AU - Ringshaw, Jessica E
AU - Zieff, Michal R
AU - Bourke, Niall J
AU - Casella, Chiara
AU - Bradford, Layla E
AU - Williams, Simone R
AU - Herr, Donna
AU - Miles, Marlie
AU - Bennallick, Carly
AU - Khula South Africa Data Collection Team
AU - Deoni, Sean
AU - O’Muircheartaigh, Jonathan
AU - Stein, Dan J
AU - Alexander, Daniel C
AU - Jones, Derek K
AU - Williams, Steven C R
AU - Donald, Kirsten A
TI - Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa
T2 - Brain communications
J2 - Brain Commun
PY - 2026
DA - 2026/09/08
VL - 8
IS - 5
SP - fcag318
SN - 2632-1297
PB - Oxford University Press
DO - 10.1093/braincomms/fcag318
UR - https://doi.org/10.1093/braincomms/fcag318
LA - en
ER -

CSL-JSON

{
"id": "10.1093/braincomms/fcag318",
"type": "article-journal",
"title": "Antenatal maternal anaemia and infant brain structure: high-field (3 T) and ultra-low-field (64 mT) MRI findings from South Africa",
"container-title": "Brain communications",
"author": [
{
"family": "Ringshaw",
"given": "Jessica E"
},
{
"family": "Zieff",
"given": "Michal R"
},
{
"family": "Bourke",
"given": "Niall J"
},
{
"family": "Casella",
"given": "Chiara"
},
{
"family": "Bradford",
"given": "Layla E"
},
{
"family": "Williams",
"given": "Simone R"
},
{
"family": "Herr",
"given": "Donna"
},
{
"family": "Miles",
"given": "Marlie"
},
{
"family": "Bennallick",
"given": "Carly"
},
{
"literal": "Khula South Africa Data Collection Team"
},
{
"family": "Deoni",
"given": "Sean"
},
{
"family": "O’Muircheartaigh",
"given": "Jonathan"
},
{
"family": "Stein",
"given": "Dan J"
},
{
"family": "Alexander",
"given": "Daniel C"
},
{
"family": "Jones",
"given": "Derek K"
},
{
"family": "Williams",
"given": "Steven C R"
},
{
"family": "Donald",
"given": "Kirsten A"
}
],
"container-title-short": "Brain Commun",
"volume": "8",
"issue": "5",
"page": "fcag318",
"DOI": "10.1093/braincomms/fcag318",
"PMID": "42713352",
"PMCID": "PMC13553114",
"ISSN": "2632-1297",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/braincomms/fcag318",
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
8
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1002/hbm.70547 [code]
miniMORPH: A Morphometry Pipeline for Low-Field MRI in Infants.
Journal: Human brain mapping
In common: ANTs, FreeSurfer, FSL, 5 other tools, developmental, structural MRI / diffusion, 9 references, 3 authors
[2] doi:10.1038/s41597-026-07350-9 [code]
An open multi-center MEG-EEG dataset for studying conscious visual perception.
Journal: Scientific data
In common: easystats, car, emmeans, 9 other tools, structural MRI / diffusion
[3] doi:10.1038/s41597-026-07377-y [code]
An open-access multi-site fMRI dataset for investigating conscious visual perception.
Journal: Scientific data
In common: easystats, car, emmeans, 9 other tools
[4] doi:10.1162/imag.a.1347 [code]
Neural and behavioural correlates of theory of mind reasoning in five-year-old children born preterm.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: ANTs, easystats, FreeSurfer, 9 other tools
[5] doi:10.1073/pnas.2603114123 [code]
The human hippocampus can pattern separate memories by meaning.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: easystats, emmeans, lmerTest, 8 other tools, 2 references
[6] doi:10.1093/nc/niag029 [code]
A data-driven approach to identifying and evaluating connectivity-based neural correlates of conscious visual perception.
Journal: Neuroscience of consciousness
In common: easystats, car, emmeans, 9 other tools
[7] doi:10.1093/cercor/bhag132 [code]
Spatiotemporal white-matter development across early childhood.
Journal: Cerebral cortex (New York, N.Y. : 1991)
In common: ANTs, lmerTest, FSL, 7 other tools, developmental, structural MRI / diffusion, 2 references
[8] doi:10.1038/s41467-026-73865-9 [code]
Histamine shapes the neurocomputational dynamics of human learning.
Journal: Nature communications
In common: easystats, car, emmeans, 8 other tools, 1 reference
[9] doi:10.1038/s41467-026-71830-0 [code]
Predicting individual differences of fear and cognitive learning and extinction.
Journal: Nature communications
In common: ANTs, car, emmeans, 8 other tools
[10] doi:10.1002/hbm.70512 [code]
Precision Imaging for Intraindividual Investigation of the Reward Response.
Journal: Human brain mapping
In common: ANTs, easystats, lmerTest, 8 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.