OSCR

Haplotype-resolved DNA methylation at the <i>APOE</i> locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk.

Code ↔ Paper

4 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 4 matches
  1. [1] § Methods › OLS linear regression methylation analyses ↔ scripts/linear_regression_genes.py, lines 45–75 · score 0.82 · independent variable, linear regression, brain bank, squares, discovery, OLS
  2. [2] § Methods › Expression linear regression analysis ↔ scripts/linear_regression_genes.py, lines 45–75 · score 0.75 · independent variable, brain bank, regression, OLS, BH, PMI
  3. [3] § Methods › OLS linear regression methylation analyses ↔ scripts/linear_regression_peaks.py, lines 52–72 · score 0.71 · independent variable, linear regression, squares, OLS, Covariates, model
  4. [4] § Methods › Expression linear regression analysis ↔ make_pcs_stepwise.py, lines 57–108 · score 0.60 · Genetic PCs, Gene expression, transformed, PCA, PLINK

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 75 lines · 2.8 KB · other · 2 matches

  1. import scipy
  2. import numpy as np
  3. import pandas as pd
  4. import matplotlib.pyplot as plt
  5. import seaborn as sns
  6. import scanpy as sc
  7. import decoupler as dc
  8. import statsmodels.api as sm
  9. from statsmodels.stats.multitest import multipletests
  10. from patsy import dmatrices
  11. adata = sc.read_h5ad(snakemake.input.merged_rna_anndata)
  12. sample_key = snakemake.params.sample_key
  13. disease_param = snakemake.params.disease_param
  14. # RNA pseudobulk
  15. pdata = dc.get_pseudobulk(
  16. adata,
  17. sample_col=sample_key,
  18. groups_col='celltype',
  19. layer='counts',
  20. mode='sum',
  21. min_cells=1,
  22. min_counts=1
  23. )
  24. pdata.layers['pcounts'] = pdata.X
  25. sc.pp.normalize_total(pdata, target_sum=1e6, max_fraction = 0.001, key_added='cpm', layer='pcounts', copy=False)
  26. # CSV pseudobulk for QTLs
  27. adata_df = pd.DataFrame(pdata.X)
  28. sample_cell = pdata.obs[[sample_key, 'celltype']]
  29. adata_df.columns = pdata.var_names.to_list()
  30. adata_df.index = sample_cell.index
  31. adata_df = pd.merge(left=sample_cell, right=adata_df, left_index=True, right_index=True)
  32. # CSV pseudobulk for linear regression (includes covariates)
  33. adata_df = pd.DataFrame(pdata.X)
  34. sample_cell = pdata.obs[snakemake.params.design_factors]
  35. adata_df.columns = pdata.var_names.to_list()
  36. adata_df.index = sample_cell.index
  37. adata_df = pd.merge(left=sample_cell, right=adata_df, left_index=True, right_index=True)
  38. # Write out pseudobulk data
  39. adata_df.to_csv(snakemake.output.rna_pseudobulk)
  40. gene_data = []
  41. for cell_type in adata_df['celltype'].drop_duplicates():
  42. cohort_df = adata_df[(adata_df['celltype'] == cell_type)]
  43. for gene in pdata.var_names.to_list():
  44. # Create independent variable Series object, save as X
  45. _, X = dmatrices(f'{disease_param} ~ Sex_numeric + PMI + Brain_bank_numeric', data=cohort_df, return_type='dataframe')
  46. # Create dependent variable Series object, save as y
  47. y = cohort_df[gene].values
  48. # Fit the model
  49. model = sm.OLS(y, X).fit(disp=0)
  50. # Add values to the dataframe
  51. gene_results = {
  52. 'celltype': cell_type,
  53. 'gene': gene,
  54. 'slope': model.params[disease_param],
  55. 'p_value': model.pvalues[disease_param],
  56. 'standard_error': model.bse[disease_param],
  57. 'r_squared': model.rsquared,
  58. 'adjusted_r_squared': model.rsquared_adj
  59. }
  60. gene_data.append(gene_results)
  61. regression_df = pd.DataFrame(gene_data)
  62. # Adjust p-value
  63. for cell_type in adata_df['celltype'].drop_duplicates():
  64. regression_df.loc[regression_df['celltype'] == cell_type, 'p_value_bh'] = scipy.stats.false_discovery_control(regression_df.loc[regression_df['cell_type'] == cell_type, 'p_value'].fillna(1))
  65. regression_df['-log10(p-value_bh)'] = -np.log10(regression_df['p_value_bh'])
  66. regression_df.to_csv(snakemake.output.cell_gene_regression)

linear_regression_genes.py at commit dce3d79, under other · at the source

Overview

Authors: Rylee M Genner1,2, Melissa Meredith3,4, Kensuke Daida1,5, Abraham Moller1, Cory Weller1,4, Alexis Ayuketah1, Pilar Alvarez Jerez1,6,7, Stuart Akeson8, Laksh Malik1, Breeana Baker1, Cedric Kouam1, Kimberly Paquette1, Adam Catching4, Sarah Bromberek1, Fangle Hu1, Xylena Reed1, Stefano Marenco9, Pavan Auluck9, Ajeet Mandal9, Benedict Paten3
and 7 other authorsJ Raphael Gibbs5, Miten Jain1,8, Mark R Cookson1,5, Andrew B Singleton1,5,7, Mike Nalls1,4, Cornelis Blauwendraat1,5, Kimberley J Billingsley1
ORCID iDs: Adam Catching
  1. Center for Alzheimer’s and Related Dementias, National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD USA
  2. Department of Biology, Johns Hopkins University, Baltimore, MD USA
  3. UC Santa Cruz Genomics Institute, Santa Cruz, CA USA
  4. DataTecnica LLC, Washington, DC USA
  5. Laboratory of Neurogenetics, National Institute on Aging, National Institutes of Health, Bethesda, MD USA
  6. Department of Neurodegenerative Disease, UCL Queen Square Institute of Neurology, University College London, London, UK
  7. Global Parkinson’s Genetics Program (GP2), Chevy Chase, MD USA
  8. Department of Bioengineering, Northeastern University, Boston, MA USA
  9. Human Brain Collection Core, Division of Intramural Research, National Institute of Mental Health, NIH, Bethesda, MD USA
Journal: NPJ dementia, volume 2, issue 1, article 45
Dates: received 1 July 2025; accepted 22 April 2026; published online 19 June 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s44400-026-00094-8 · PMID 42327426 · PMCID PMC13282175 · OpenAlex W7165124428
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), Alzheimer's / dementia (population), cellular / molecular (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions
Keywords: Genetics, Molecular biology, Neuroscience
Topic: Epigenetics and DNA Methylation (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: University of Miami; University of Pennsylvania; Johns Hopkins University; Rush University (P30AG10161, R01AG36042, R01 AG15819); National Institutes of Health (U24‐AG041689, U01AG016976, P30‐AG072946, AG072946, R01 AG15819, AG016976, R01HG010485, Z01‐AG000949‐02, Z01-ES101986, R01AG42210, T32HG012344, P30AG10161, Z01 AG000949, R01NS78009, 1ZIANS003154, R01AG36042, AG10161, R01‐AG34374, U01AG46152, R01AG39478, R01 AG17917, U18NS82140, R01AG36836); National Institute on Aging (U24 AG041689, AG072946, Z01‐AG000949, AG 016976, R01AG36042, Z01-AG000949–02, P30-AG 10161, Z01‐ES101986, R01-AG34374, R01‐AG17917, R01 AG15819, R01-AG36836, P30‐AG072946, 1ZIANS003154, U01-AG016976, R01NS78009, R01AG42210, U01‐AG46152)
Citations: cited by 1 paper (Europe PMC); 67 references in the paper

Abstract

The APOE gene encodes a lipid transport protein central to Alzheimer’s disease (AD) pathogenesis. Three common alleles—ε2 (rs7412(C > T)), ε3 (reference), and ε4 (rs429358(T > C))—arise from two coding variants in exon 4 and confer distinct AD risk profiles, with ε4 increasing risk and ε2 being protective. The ε3-linked APOE variant rs769455[T] has also been associated with increased AD risk among individuals of African ancestry who also carry the APOE ε4 allele. Determining how genetic variation influences CpG methylation requires methQTL-type analyses, but conventional bisulfite and array-based approaches offer limited resolution for distinguishing allele-specific effects. Here, we use high-accuracy long-read sequencing to generate haplotype-resolved methylation profiles across the APOE locus in 332 postmortem brain tissue samples from ancestrally diverse cohorts, including 201 samples from individuals of European ancestry and 131 samples from individuals of African and African admixed ancestry. Treating each haplotype as an independent observation, OLS regression identified 18 novel differentially methylated CpG sites associated with ε2, ε4, and rs769455[T] across the APOE locus (TOMM40, APOE, APOC1, and APOC4-APOC2 genes). These findings reveal distinct allele-specific methylation signatures and demonstrate the utility of long-read sequencing for resolving epigenetic variation relevant to AD risk.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.

NIH-CARD/CARDlongread_data_standardization

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 2b782f21911348ba91127a515c62ac3f367d7fa3, 17 April 2026
Languages: Python (15), Shell (2)
Size: 26 files, 17 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (10 files), pandas (10 files), Matplotlib (3 files), pysam (3 files), seaborn (3 files), BCFtools (2 files), statsmodels (2 files), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
18 files

nih-card/scmaverics

License: other
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: dce3d79def4e08190aca964183e009b1e90a9041, 22 September 2026
Languages: Python (64), Shell (11)
Size: 98 files, 75 scripts
Software Heritage: not archived
Found in: the text, “Cell-type proportions from single nuclei express”
Holds: README, license file, documentation
Not found: CITATION.cff, environment file, tests, continuous integration
Tools: Scanpy (54 files), pandas (53 files), NumPy (44 files), SciPy (18 files), anndata (16 files), Matplotlib (10 files), seaborn (9 files), PyTorch (5 files), statsmodels (5 files), BEDTools (1 file), pysam (1 file), Snakemake (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
77 files

NIH-CARD/APOE_CpG_methylation

License: none: the authors keep all their rights
State: the link is dead, verified on 27 September 2026
Evidence: found in the paper
Software Heritage: not archived
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link is dead
  • 27 September 2026: the link is dead

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 92 scripts, each with its path and the digest of its content;
  • 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

Access requests for these controlled datasets are available through dbGaP under accessions phs001300.v5.p1 and phs000979.v4.p2 and the data is available at https://explore.anvilproject.org/datasets or https://duos.org/datalibrary/anvil. The code used to process and analyze the data for this study is publicly available at https://github.com/NIH-CARD/APOE_CpG_methylation. This repository includes scripts designed to conduct an allele-specific, CpG-site specific beta-binomial regression analysis given a modkit bed file for a specific region and sample haplotype designations. The scripts used to generate all main and supplementary figures are also included. Additional details about the files and file formats and the scripts used to generate the PC values are located on the NIH-CARD Github page (https://github.com/NIH-CARD/CARDlongread_data_standardization).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Springer Science+Business Media
  • Funding: added University of Miami; University of Pennsylvania; Johns Hopkins University; Rush University: P30AG10161, R01AG36042, R01 AG15819; National Institutes of Health: U24‐AG041689, U01AG016976, P30‐AG072946, AG072946, R01 AG15819, AG016976, R01HG010485, Z01‐AG000949‐02, Z01-ES101986, R01AG42210, T32HG012344, P30AG10161, Z01 AG000949, R01NS78009, 1ZIANS003154, R01AG36042, AG10161, R01‐AG34374, U01AG46152, R01AG39478, R01 AG17917, U18NS82140, R01AG36836; National Institute on Aging: U24 AG041689, AG072946, Z01‐AG000949, AG 016976, R01AG36042, Z01-AG000949–02, P30-AG 10161, Z01‐ES101986, R01-AG34374, R01‐AG17917, R01 AG15819, R01-AG36836, P30‐AG072946, 1ZIANS003154, U01-AG016976, R01NS78009, R01AG42210, U01‐AG46152

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 27 authors, 3 keywords, 66 references.

Cite

This paper

Genner, R. M., Meredith, M., Daida, K., Moller, A., Weller, C., Ayuketah, A., Jerez, P. A., Akeson, S., Malik, L., Baker, B., Kouam, C., Paquette, K., Catching, A., Bromberek, S., Hu, F., Reed, X., Marenco, S., Auluck, P., Mandal, A., . . . Billingsley, K. J. (2026). Haplotype-resolved DNA methylation at the <i>APOE</i> locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk. NPJ dementia, 2(1), 45. https://doi.org/10.1038/s44400-026-00094-8

BibTeX

@article{genner2026haplotype,
author = {Genner, Rylee M and Meredith, Melissa and Daida, Kensuke and Moller, Abraham and Weller, Cory and Ayuketah, Alexis and Jerez, Pilar Alvarez and Akeson, Stuart and Malik, Laksh and Baker, Breeana and Kouam, Cedric and Paquette, Kimberly and Catching, Adam and Bromberek, Sarah and Hu, Fangle and Reed, Xylena and Marenco, Stefano and Auluck, Pavan and Mandal, Ajeet and Paten, Benedict and Gibbs, J Raphael and Jain, Miten and Cookson, Mark R and Singleton, Andrew B and Nalls, Mike and Blauwendraat, Cornelis and Billingsley, Kimberley J},
title = {{Haplotype-resolved DNA methylation at the \<i\>APOE\</i\> locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk}},
journal = {NPJ dementia},
year = {2026},
month = jun,
volume = {2},
number = {1},
pages = {45},
publisher = {Springer Science+Business Media},
issn = {3005-1940},
doi = {10.1038/s44400-026-00094-8},
url = {https://doi.org/10.1038/s44400-026-00094-8},
pmid = {42327426},
pmcid = {PMC13282175}
}

RIS

TY - JOUR
AU - Genner, Rylee M
AU - Meredith, Melissa
AU - Daida, Kensuke
AU - Moller, Abraham
AU - Weller, Cory
AU - Ayuketah, Alexis
AU - Jerez, Pilar Alvarez
AU - Akeson, Stuart
AU - Malik, Laksh
AU - Baker, Breeana
AU - Kouam, Cedric
AU - Paquette, Kimberly
AU - Catching, Adam
AU - Bromberek, Sarah
AU - Hu, Fangle
AU - Reed, Xylena
AU - Marenco, Stefano
AU - Auluck, Pavan
AU - Mandal, Ajeet
AU - Paten, Benedict
AU - Gibbs, J Raphael
AU - Jain, Miten
AU - Cookson, Mark R
AU - Singleton, Andrew B
AU - Nalls, Mike
AU - Blauwendraat, Cornelis
AU - Billingsley, Kimberley J
TI - Haplotype-resolved DNA methylation at the <i>APOE</i> locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk
T2 - NPJ dementia
J2 - NPJ Dement
PY - 2026
DA - 2026/06/19
VL - 2
IS - 1
SP - 45
SN - 3005-1940
PB - Springer Science+Business Media
DO - 10.1038/s44400-026-00094-8
UR - https://doi.org/10.1038/s44400-026-00094-8
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s44400-026-00094-8",
"type": "article-journal",
"title": "Haplotype-resolved DNA methylation at the <i>APOE</i> locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk",
"container-title": "NPJ dementia",
"author": [
{
"family": "Genner",
"given": "Rylee M"
},
{
"family": "Meredith",
"given": "Melissa"
},
{
"family": "Daida",
"given": "Kensuke"
},
{
"family": "Moller",
"given": "Abraham"
},
{
"family": "Weller",
"given": "Cory"
},
{
"family": "Ayuketah",
"given": "Alexis"
},
{
"family": "Jerez",
"given": "Pilar Alvarez"
},
{
"family": "Akeson",
"given": "Stuart"
},
{
"family": "Malik",
"given": "Laksh"
},
{
"family": "Baker",
"given": "Breeana"
},
{
"family": "Kouam",
"given": "Cedric"
},
{
"family": "Paquette",
"given": "Kimberly"
},
{
"family": "Catching",
"given": "Adam"
},
{
"family": "Bromberek",
"given": "Sarah"
},
{
"family": "Hu",
"given": "Fangle"
},
{
"family": "Reed",
"given": "Xylena"
},
{
"family": "Marenco",
"given": "Stefano"
},
{
"family": "Auluck",
"given": "Pavan"
},
{
"family": "Mandal",
"given": "Ajeet"
},
{
"family": "Paten",
"given": "Benedict"
},
{
"family": "Gibbs",
"given": "J Raphael"
},
{
"family": "Jain",
"given": "Miten"
},
{
"family": "Cookson",
"given": "Mark R"
},
{
"family": "Singleton",
"given": "Andrew B"
},
{
"family": "Nalls",
"given": "Mike"
},
{
"family": "Blauwendraat",
"given": "Cornelis"
},
{
"family": "Billingsley",
"given": "Kimberley J"
}
],
"container-title-short": "NPJ Dement",
"volume": "2",
"issue": "1",
"page": "45",
"DOI": "10.1038/s44400-026-00094-8",
"PMID": "42327426",
"PMCID": "PMC13282175",
"ISSN": "3005-1940",
"publisher": "Springer Science+Business Media",
"URL": "https://doi.org/10.1038/s44400-026-00094-8",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
19
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.celrep.2026.117110 [code]
Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.
Journal: Cell reports
In common: Snakemake, BCFtools, pysam, 10 other tools, genetics / omics, cellular / molecular, 3 references, author Adam Catching
[2] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: Snakemake, pysam, BEDTools, 8 other tools, genetics / omics, cellular / molecular, 2 references
[3] doi:10.1093/bioinformatics/btag652 [code]
mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.
Journal: Bioinformatics (Oxford, England)
In common: pysam, BEDTools, anndata, 9 other tools, genetics / omics
[4] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: pysam, BEDTools, anndata, 9 other tools, genetics / omics
[5] doi:10.1016/j.celrep.2026.117073 [code]
Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.
Journal: Cell reports
In common: Snakemake, pysam, BEDTools, 7 other tools, genetics / omics, cellular / molecular
[6] doi:10.3389/fnmol.2026.1844705 [code]
Risperidone regulates the expression of schizophrenia-related genes in the forebrain of adult male mice.
Journal: Frontiers in molecular neuroscience
In common: pysam, BEDTools, anndata, 8 other tools, genetics / omics, cellular / molecular
[7] doi:10.1038/s41467-026-73171-4 [code]
Dissecting epigenetic heterogeneity in single-cell DNA methylomes with a unified framework.
Journal: Nature communications
In common: pysam, BEDTools, anndata, 8 other tools, genetics / omics, cellular / molecular
[8] doi:10.1038/s42003-026-10462-y [code]
SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.
Journal: Communications biology
In common: pysam, BEDTools, anndata, 7 other tools, genetics / omics, 1 reference
[9] doi:10.1038/s41467-026-71790-5 [code]
Recurrent DNA break clusters drive replication-stress-induced copy number variants and genome diversification.
Journal: Nature communications
In common: Snakemake, BCFtools, pysam, 6 other tools, genetics / omics, cellular / molecular
[10] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: pysam, BEDTools, statsmodels, 6 other tools, genetics / omics, cellular / molecular, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.