OSCR

Local graph estimation with pathwise false discovery control.

Code ↔ Paper

19 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 19 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Results › Brain networks and cognition ↔ applications/hcp/data/clean_data.py, lines 1–19 · score 0.91 · HCP Young Adult, Fluid Cognition Composite, Human Connectome Project, NIH Toolbox, age adjusted, phenotypes
  2. [2] § Results › Brain networks and cognition ↔ applications/hcp/run_pfs.py, lines 1–19 · score 0.79 · HCP Young Adult, Fluid Cognition Composite, NIH Toolbox, neuroimaging, phenotypes, PFS
  3. [3] § Results › Cross-modal pathways in breast cancer ↔ applications/breast_cancer/run_methods.py, lines 1–19 · score 0.72 · microRNAs, breast cancer, pathologic stage, TCGA, RPPA, histological
  4. [4] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/figures/heatmaps.py, lines 37–48 · score 0.71 · sulfur dioxide, cancer mortality, pm2.5, Cancer incidence, mercury, TCE
  5. [5] § Results › Cross-modal pathways in breast cancer ↔ applications/breast_cancer/run_pfs.py, lines 1–18 · score 0.71 · miRNAs, breast cancer, pathologic stage, TCGA, RPPA, histological
  6. [6] § Results › Brain networks and cognition ↔ applications/hcp/enrichment/enrichment_analysis.py, lines 147–174 · score 0.67 · ventral attention, frontoparietal control, fluid cognition, FPCN, VAN, enrichment
  7. [7] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/figures/heatmaps.py, lines 37–48 · score 0.66 · sulfur dioxide, pm2.5, cancer incidence, SO2, mortality
  8. [8] § Methods › Simulation design for Fig. 1 ↔ simulations/simulate_block.py, lines 80–93 · score 0.66 · precision matrix, positive definiteness, eigenvalues, blocks, Simulation
  9. [9] § Results › Cell-type-specific gene networks in Alzheimer’s disease ↔ applications/alzheimers/figures/plot_results.py, lines 49–54 · score 0.59 · oligodendrocyte progenitor cells, astrocytes, microglia, OPCs, Alzheimer
  10. [10] § Results › Cross-modal pathways in breast cancer ↔ applications/breast_cancer/enrichment/gene_enrichment.py, lines 29–51 · score 0.58 · gene targets, ISCU, NDRG1, module, breast cancer, CDH1
  11. [11] § Results › Pathwise feature selection ↔ localgraph/pfs/main.py, the whole file · a weak match · score 0.56 · pathwise threshold, maximum radius, iterative, sum, local graph, selection
  12. [12] § Results › Brain networks and cognition ↔ applications/hcp/enrichment/enrichment_analysis.py, lines 147–174 · score 0.56 · fluid cognition, VIS, somatomotor, SM, FPCN, VAN
  13. [13] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/run_methods.py, lines 1–19 · score 0.55 · mortality rates, age adjusted, demographic, socioeconomic, Cancer, county
  14. [14] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/run_pfs.py, lines 1–19 · score 0.55 · mortality rates, age adjusted, demographic, socioeconomic, Cancer, county
  15. [15] § Results › Cross-modal pathways in breast cancer ↔ applications/breast_cancer/run_pfs.py, lines 1–18 · score 0.52 · miRNAs, breast cancer, protein, gene
  16. [16] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/run_methods.py, lines 1–19 · score 0.52 · environmental exposures, cancer incidence, socioeconomic, counties, mortality
  17. [17] § Results › Environmental and social drivers of cancer ↔ applications/env_cancer_study/run_pfs.py, lines 1–19 · score 0.52 · environmental exposures, cancer incidence, socioeconomic, counties, mortality
  18. [18] § Methods › Integrated path stability selection ↔ simulations/qvalue_comparison/analyze_qval_results.py, lines 109–132 · score 0.51 · random forests, stability selection, nonparametric, modeling, IPSS, discovery
  19. [19] § Results › Simulation studies ↔ simulations/qvalue_comparison/analyze_qval_results.py, lines 109–132 · score 0.51 · Local graph recovery, discovery rate, simulation, sparsity, TPR, FDR

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 151 lines · 4.2 KB · MIT · 2 matches

  1. # Plot heatmaps for target variables and select environmental exposure and social variables (Figure 3 in the paper)
  2. import os
  3. import geopandas as gpd
  4. import matplotlib.colors as mcolors
  5. import matplotlib.pyplot as plt
  6. import pandas as pd
  7. import numpy as np
  8. #--------------------------------
  9. # Configuration
  10. #--------------------------------
  11. save_fig = False
  12. dpi = 300
  13. # specify feature types ('targets', 'exposures', or 'social')
  14. feature_type = 'targets'
  15. if feature_type == 'targets':
  16. features_to_plot = ['Incidence', 'Mortality']
  17. nrows = 1
  18. elif feature_type == 'exposures':
  19. features_to_plot = ['pm2.5(A)', 'Hg(W)', 'SO2(A)', 'TCE(A)']
  20. nrows = 2
  21. elif feature_type == 'social':
  22. features_to_plot = ['Poverty', 'Education', 'Hispanic', 'Smoking']
  23. nrows = 2
  24. fig_name = f'heatmaps_{feature_type}'
  25. ncols = 2
  26. figsize = (16,9)
  27. # Color map
  28. cmap_base = plt.colormaps['Spectral'].reversed()
  29. cmap = mcolors.ListedColormap(cmap_base(np.linspace(1/4, 1, 256)))
  30. # format feature names
  31. feature_renames = {
  32. 'Incidence': 'Cancer incidence',
  33. 'Mortality': 'Cancer mortality',
  34. 'pm2.5(A)': 'PM$\\bf{_{2.5}}$',
  35. 'Hg(W)': 'Mercury (Hg)',
  36. 'SO2(A)': 'Sulfur dioxide (SO$\\bf{_{2}}$)',
  37. 'SO4(W)': 'Sulfate (SO$\\bf{_{4}}$)',
  38. 'TCE(A)': 'TCE (C$\\bf{_{2}}$HCl$\\bf{_{3}}$)',
  39. 'PSATest': 'PSA test',
  40. 'PapSmear': 'Pap smear'
  41. }
  42. #--------------------------------
  43. # Plotting function
  44. #--------------------------------
  45. def plot_heatmaps(
  46. feature_names,
  47. file,
  48. base_path="../data/raw_data",
  49. shapefile_path="./utils/tl_2022_us_county.shp",
  50. nrows=2,
  51. ncols=2,
  52. figsize=(16,9),
  53. cmap=cmap,
  54. fig_name=None,
  55. save_fig=False,
  56. dpi=300
  57. ):
  58. # Load feature name mapping
  59. mapping_path = os.path.join(base_path, file + "_feature_names.xlsx")
  60. mapping_df = pd.read_excel(mapping_path, engine="openpyxl", dtype=str)
  61. reverse_map = dict(zip(
  62. mapping_df["Updated Variable Name"].str.strip(),
  63. mapping_df["Variable Name"].str.strip()
  64. ))
  65. # Load county shapefile
  66. gdf = gpd.read_file(shapefile_path)
  67. gdf["FIPS"] = (gdf["STATEFP"] + gdf["COUNTYFP"]).astype(str).str.zfill(5)
  68. # Fix outdated Connecticut FIPS codes
  69. ct_fips_fix = {
  70. "09110": "09001", "09120": "09003", "09130": "09005", "09140": "09007",
  71. "09150": "09009", "09160": "09011", "09170": "09013", "09180": "09015"
  72. }
  73. gdf["FIPS"] = gdf["FIPS"].replace(ct_fips_fix)
  74. gdf = gdf[~gdf["STATEFP"].isin(["02", "15"])] # remove Alaska and Hawaii
  75. # Load environmental data
  76. data_path = os.path.join(base_path, file + ".csv")
  77. df = pd.read_csv(data_path, dtype={'FIPS': str})
  78. df['FIPS'] = df['FIPS'].str.zfill(5)
  79. # Setup figure
  80. fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=figsize)
  81. axes = axes.flatten()
  82. # Plot each feature
  83. for i, feature_name in enumerate(feature_names):
  84. feature_col = reverse_map.get(feature_name, feature_name)
  85. plot_col = feature_col
  86. df_feature = df[['FIPS', plot_col]].dropna()
  87. merged = gdf.merge(df_feature, on="FIPS", how="left").dropna(subset=[plot_col])
  88. merged[plot_col] = pd.to_numeric(merged[plot_col], errors='coerce')
  89. # Scale color range
  90. vmin = merged[plot_col].quantile(0.05)
  91. vmax = merged[plot_col].quantile(0.95)
  92. norm = mcolors.Normalize(vmin=vmin, vmax=vmax)
  93. # Use reversed colormap for Education
  94. current_cmap = cmap.reversed() if feature_name == 'Education' else cmap
  95. ax = axes[i]
  96. merged.plot(column=plot_col, cmap=current_cmap, linewidth=0.3,
  97. edgecolor="gray", ax=ax, legend=False, norm=norm)
  98. # Set title
  99. title = feature_renames.get(feature_name, feature_name)
  100. ax.set_title(title, fontsize=26, fontweight='bold', pad=1)
  101. # Standardize layout
  102. ax.set_xlim(merged.total_bounds[0], merged.total_bounds[2])
  103. ax.set_ylim(23, 50)
  104. ax.axis("off")
  105. ax.set_aspect(1.25)
  106. # Turn off any unused subplots
  107. for j in range(i + 1, len(axes)):
  108. axes[j].axis("off")
  109. wspace = 0.05 if nrows == 1 else -0.03
  110. plt.subplots_adjust(left=0, right=1, top=0.95, bottom=0, wspace=wspace, hspace=0.075)
  111. if save_fig:
  112. plt.savefig(f'{fig_name}_dpi{dpi}.png', dpi=dpi)
  113. plt.show()
  114. #--------------------------------
  115. # Call the function
  116. #--------------------------------
  117. plot_heatmaps(
  118. feature_names=features_to_plot,
  119. file="eqi2000",
  120. base_path="../data/raw_data",
  121. shapefile_path="./utils/tl_2022_us_county.shp",
  122. ncols=ncols,
  123. nrows=nrows,
  124. figsize=figsize,
  125. fig_name=fig_name,
  126. save_fig=save_fig,
  127. dpi=dpi
  128. )

heatmaps.py at commit 71c1f5c, under MIT · at the source

Overview

Authors: Omar Melikechi1, David B Dunson1, Noureddine Melikechi2, Jeffrey W Miller3
ORCID iDs: Omar Melikechi
  1. Department of Statistical Science, Duke University, Durham, NC USA
  2. Kennedy College of Sciences, University of Massachusetts Lowell, Lowell, MA USA
  3. Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA USA
Institutions: Duke University (United States); University of Massachusetts Lowell (United States); Harvard University (United States)
Journal: Nature communications, volume 17, issue 1, article 6353
Dates: received 22 August 2025; accepted 24 April 2026; published online 12 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-72796-9 · PMID 42120384 · PMCID PMC13376902 · OpenAlex W4414878489
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), other condition (population), methods / tools (subfield)
Methods: Statistics, Machine learning
Keywords: Statistical methods, Probabilistic data networks, Cancer genomics, Machine learning, Cancer epidemiology
MeSH: Algorithms*, Animals, Brain, Connectome, Humans, Multiomics (* major topic)
Topic: Advanced Graph Neural Networks (Artificial Intelligence, Computer Science), according to OpenAlex
Funding: NIEHS NIH HHS (R01 ES035625); Foundation for the National Institutes of Health (Foundation for the National Institutes of Health, Inc.) (R01ES035625, R01CA240299); NCI NIH HHS (R01 CA240299)
Citations: cited by 1 paper (Europe PMC); 48 references in the paper

Abstract

Many datasets include a small set of variables, such as biomarkers or clinical outcomes, whose relationships to the broader system are of primary scientific interest. Estimating the full network of inter-variable relationships in such settings often obscures local structures around these targets, limiting interpretability. To address this fundamental problem, we introduce local graph estimation, a statistical framework for inferring substructures around target variables. We show that traditional graph estimation methods often fail to recover local structure, and present pathwise feature selection (PFS) as an effective alternative. PFS estimates local subgraphs by iteratively applying feature selection and propagating uncertainty along network paths, providing rigorous finite-sample false discovery control even in settings with mixed variable types and nonlinear dependencies. In four distinct applications spanning environmental and public health, multiomics, brain connectomics, and single-nucleus RNA sequencing, PFS recovers interpretable networks consistent with domain knowledge, highlighting its ability to uncover established mechanisms and generate novel hypotheses.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.

omelikechi/localgraph-paper

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 71c1f5c992d4bec7a521db0ffc8c490b9dd791e2, 19 April 2026
Languages: Python (51)
Size: 125 files, 51 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (32 files), pandas (25 files), Matplotlib (18 files), NetworkX (9 files), scikit-learn (8 files), rpy2 (5 files), SciPy (2 files), statsmodels (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
53 files

omelikechi/localgraph

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 9e30b66b0f1702b363c63a35d2edee4b50468941, 28 May 2026
Languages: Python (12)
Size: 18 files, 12 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (pyproject.toml)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (6 files), Matplotlib (3 files), NetworkX (2 files)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
14 files

Zenodo 19655606

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (32 files), pandas (25 files), Matplotlib (18 files), NetworkX (9 files), scikit-learn (8 files), rpy2 (5 files), SciPy (2 files), statsmodels (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
  • 28 September 2026: the link answers (HTTP 200)
53 files
At the source:

Code availability

Code and processed data files required to reproduce all results in this paper are available at https://github.com/omelikechi/localgraph-paperand archived on Zenodo47. A Python package implementing local graph estimation and PFS is available at https://github.com/omelikechi/localgraphand can be installed via PyPI at https://pypi.org/project/localgraph/.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 114 scripts, each with its path and the digest of its content;
  • 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

All data used in this work are publicly available. For the environmental and sociodemographic cancer study, county-level data on cancer incidence, mortality, screening, and smoking prevalence are available from the State Cancer Profiles project at https://statecancerprofiles.cancer.gov/. Environmental and socioeconomic variables are available from the EPA Environmental Quality Index (EQI) at https://cfpub.epa.gov/ncea/risk/recordisplay.cfm?deid=316550, and demographic data are available from the U.S. Census Bureau at https://data.census.gov/. Data from the multiomic breast cancer study can be downloaded from LinkedOmics at https://www.linkedomics.org/data_download/TCGA-BRCA/. Data from the brain network and cognition analysis were obtained from the Human Connectome Project (HCP Young Adult cohort) via ConnectomeDB (https://www.humanconnectome.org), accessed through the BALSA data portal at https://balsa.wustl.edu. Data from the Alzheimer’s disease study can be downloaded from https://adsn.ddnetbio.comand are also available from the Gene Expression Omnibus (GEO) under the accession number GSE138852 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE138852).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 5 keywords, 6 MeSH terms, 3 funders, 34 references.

Cite

This paper

Melikechi, O., Dunson, D. B., Melikechi, N., & Miller, J. W. (2026). Local graph estimation with pathwise false discovery control. Nature communications, 17(1), 6353. https://doi.org/10.1038/s41467-026-72796-9

BibTeX

@article{melikechi2026local,
author = {Melikechi, Omar and Dunson, David B and Melikechi, Noureddine and Miller, Jeffrey W},
title = {{Local graph estimation with pathwise false discovery control}},
journal = {Nature communications},
year = {2026},
month = may,
volume = {17},
number = {1},
pages = {6353},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-72796-9},
url = {https://doi.org/10.1038/s41467-026-72796-9},
pmid = {42120384},
pmcid = {PMC13376902}
}

RIS

TY - JOUR
AU - Melikechi, Omar
AU - Dunson, David B
AU - Melikechi, Noureddine
AU - Miller, Jeffrey W
TI - Local graph estimation with pathwise false discovery control
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/05/12
VL - 17
IS - 1
SP - 6353
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-72796-9
UR - https://doi.org/10.1038/s41467-026-72796-9
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-72796-9",
"type": "article-journal",
"title": "Local graph estimation with pathwise false discovery control",
"container-title": "Nature communications",
"author": [
{
"family": "Melikechi",
"given": "Omar"
},
{
"family": "Dunson",
"given": "David B"
},
{
"family": "Melikechi",
"given": "Noureddine"
},
{
"family": "Miller",
"given": "Jeffrey W"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "6353",
"DOI": "10.1038/s41467-026-72796-9",
"PMID": "42120384",
"PMCID": "PMC13376902",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-72796-9",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pcbi.1014346 [code]
StPedf: Cell trajectory inference of spatial transcriptomics via spatial proximity embedding and spatial density-adaptive fusion.
Journal: PLoS computational biology
In common: rpy2, NetworkX, statsmodels, 5 other tools, genetics / omics
[2] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: rpy2, NetworkX, statsmodels, 5 other tools, genetics / omics
[3] doi:10.1038/s41593-026-02267-3 [code]
Spatial proteomic analysis in human Alzheimer's disease brains enables identification of microenvironment-dependent microglial cell states.
Journal: Nature neuroscience
In common: rpy2, NetworkX, statsmodels, 5 other tools, genetics / omics
[4] doi:10.1038/s41467-026-74694-6 [code]
Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Journal: Nature communications
In common: statsmodels, scikit-learn, pandas, 3 other tools, methods / tools, genetics / omics, 2 references
[5] doi:10.1093/bib/bbag485 [code]
Single-cell-level perturbation-induced and condition-related signal estimation with batch effect removal using NDreamer.
Journal: Briefings in bioinformatics
In common: rpy2, statsmodels, scikit-learn, 4 other tools, genetics / omics, 1 reference
[6] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: NetworkX, statsmodels, scikit-learn, 4 other tools, other condition, 2 references
[7] doi:10.64898/2026.03.04.709586 [code]
Mapping Higher-Order Topology in OCD Brain Networks with Hodge Laplacian
Journal: bioRxiv (preprint)
In common: NetworkX, statsmodels, scikit-learn, 4 other tools, other condition, 1 reference
[8] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: rpy2, NetworkX, scikit-learn, 4 other tools, genetics / omics, other condition
[9] doi:10.1038/s41467-026-75959-w [code]
Charting higher-order models of brain function beyond pairwise interactions.
Journal: Nature communications
In common: NetworkX, statsmodels, scikit-learn, 4 other tools, 2 references
[10] doi:10.1016/j.ebiom.2026.106312 [code]
Subgingival microbiota composition is associated with brain health in the general population-the PAROMIND study.
Journal: EBioMedicine
In common: rpy2, NetworkX, statsmodels, 4 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.