OSCR

EIF2S1-PELO co-expression architecture in human brain: pairing-family decomposition reveals constitutive transcriptional coupling and functional layer separation.

Code ↔ Paper

20 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 20 matches
  1. [1] § Methods › Data sources ↔ r3_revision_pipelines/pipeline_9_disease_als_gse124439.py, lines 1–32 · score 0.95 · ALS Spectrum MND, ALS Consortium, STAR gene, neurological control, RNA seq, frontal cortex
  2. [2] § Results › Replication of the constitutive coupling across additional diseases and an independent platform ↔ r3_revision_pipelines/pipeline_8_disease_ad_gse5281.py, lines 1–36 · score 0.93 · superior frontal gyrus, AD cohort, constitutive coupling, RNA seq, frontal cortex, WJ_shifted
  3. [3] § Results › Replication of the constitutive coupling across additional diseases and an independent platform ↔ r3_revision_pipelines/pipeline_10_disease_meta_rnaseq.py, lines 1–25 · score 0.91 · stratified permutation meta, RNA seq cohorts, disease associated, coupling metric, frontal cortex, WJ_shifted
  4. [4] § Methods › Data sources ↔ r3_revision_pipelines/pipeline_8_disease_ad_gse5281.py, lines 131–175 · score 0.91 · Affymetrix HG U133, superior frontal gyrus, RNA seq primary, Plus, GSE5281, microarray
  5. [5] § Results › Layer separation: zero functional cellular co-dependency in DepMap CRISPR screens ↔ r3_revision_pipelines/pipeline_6_string_ppi.py, lines 1–31 · score 0.90 · PPI enrichment, Homo sapiens, protein interaction, interaction partners, ribosome quality control, layer separation
  6. [6] § Methods › Data sources ↔ MS2_EIF2S1_Definitive_Pipeline.py, lines 50–92 · score 0.85 · basal ganglia, anterior cingulate, nucleus accumbens, brain regions, BA24, hippocampus
  7. [7] § Methods › Statistical testing ↔ r3_revision_pipelines/pipeline_4_layer2h_pairing.py, lines 1–81 · score 0.83 · odds ratios, biological relatedness, pair landscape, Pairwise comparisons, TMEM97, overlap
  8. [8] § Methods › Coupling profile and pairing-family decomposition ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 352–388 · score 0.82 · approximately linear monotonic, Substrate projection, Near zero gap, pearson wj, spearman wj, coupling profiles
  9. [9] § Results › Architecture is approximately linear-monotonic and moderately sign-driven ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 352–388 · score 0.82 · approximately linear monotonic, substrate projection, near zero gap, Pearson WJ, Spearman WJ, nonlinear
  10. [10] § Methods › Data sources ↔ MS2_EIF2S1_Definitive_Pipeline.py, lines 66–74 · score 0.79 · anterior cingulate, nucleus accumbens, basal ganglia, brain regions, BA24, hippocampus
  11. [11] § Results › Type 1 dissociation gap is the lowest in a 21-pair quality-control panel and is significantly smaller than random gene-pair gaps ↔ r3_revision_pipelines/pipeline_4_layer2h_pairing.py, lines 1–81 · score 0.75 · Fisher odds ratio, Spearman rho, pairwise comparisons, continuous WJ, binary Jaccard, dissociation gap
  12. [12] § Methods › Coupling profile and pairing-family decomposition ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 503–561 · score 0.74 · Sign treatment gaps, wj_signed, wj_unsigned, wj_shifted
  13. [13] § Results › Architecture is approximately linear-monotonic and moderately sign-driven ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 503–561 · score 0.73 · wj_signed, wj_unsigned, sign treatment, wj_shifted, dominant, linear
  14. [14] § Results › Layer separation: zero functional cellular co-dependency in DepMap CRISPR screens ↔ r3_revision_pipelines/pipeline_4_layer2h_pairing.py, lines 182–223 · score 0.70 · direct Pearson, functional co dependency, essentially zero, DepMap, phenotypes, cellular
  15. [15] § Results › EIF2S1-PELO co-expression architecture is constitutive across human brain regions ↔ r3_revision_pipelines/pipeline_7_disease_significance.py, lines 1–61 · score 0.65 · GSE68719 BA9 prefrontal, disease comparison, WJ_shifted, binary Jaccard, coupling profiles, dissociation gap
  16. [16] § Methods › Coupling profile and pairing-family decomposition ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 1–78 · score 0.62 · Continuous Weighted Jaccard, Continuous Discrete, coupling profile, Binary Jaccard, Dissociation gap, min
  17. [17] § Methods › Statistical testing ↔ r3_revision_pipelines/pipeline_5_formal_layer2h.py, lines 434–471 · score 0.57 · 2.5–97.5 %, bootstrap CI, EIF2S1 PELO, gap
  18. [18] § Results › Exploratory comparison of top-partner architecture in Parkinson’s disease and control BA9 prefrontal cortex ↔ r3_revision_pipelines/pipeline_7_disease_significance.py, lines 1–61 · score 0.55 · WJ_shifted, binary Jaccard, relabeling, dissociation gap, disease, bootstrap
  19. [19] § Methods › Statistical testing ↔ r3_revision_pipelines/pipeline_1_tcf25_asc1.py, lines 45–71 · score 0.52 · FORCE_RECOMPUTE, RANDOM_SEED
  20. [20] § Results › EIF2S1-PELO co-expression architecture is constitutive across human brain regions ↔ r3_revision_pipelines/pipeline_4_layer2h_pairing.py, lines 157–180 · score 0.51 · Mann Whitney, cross region preservation, brain regions, EIF2S1, PELO, genes

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 561 lines · 24 KB · MIT · 6 matches

  1. """
  2. Pipeline 5: Formal Layer 2H Bulletproof Analysis on EIF2S1-PELO
  3. Author: Drake H. Harbert (D.H.H.)
  4. Affiliation: Inner Architecture LLC, Canton, OH
  5. ORCID: 0009-0007-7740-3616
  6. Date: 2026-04-29
  7. Description:
  8. Computes formal continuous Weighted Jaccard, all three sign-treatment
  9. formulations, Type 1/2/6 dissociation gaps, permutation null tests, and
  10. bootstrap confidence intervals for the EIF2S1-PELO architectural
  11. coupling on GSE68719 (BA9 prefrontal cortex RNA-seq, 73 samples: 29 PD,
  12. 44 controls). Uses local data; no Colab dependency.
  13. Tests run:
  14. Type 1 (Continuous-Discrete): formal continuous WJ vs binary Jaccard
  15. at top 5%. Permutation null vs 100 random gene pairs. Bootstrap
  16. 95% CI on EIF2S1-PELO gap.
  17. Type 2 (Sign-Treatment): three WJ formulations on the same coupling
  18. profiles — shifted (r+1)/2, unsigned |r|, signed split. Gap
  19. between formulations characterizes whether sign-inversion is
  20. a major contributor.
  21. Type 6 (Substrate-Projection): Pearson WJ vs Spearman WJ. Gap
  22. quantifies linear vs monotonic-only structure.
  23. Disease comparison: all metrics computed separately for control and
  24. PD subsets to assess preservation/disruption.
  25. Numeric definitions:
  26. Continuous WJ on coupling profiles: for two genes X and Y, compute
  27. their coupling vectors c_X = (r(X,g) for all genes g != X), then
  28. WJ(c_X, c_Y) using selected formulation.
  29. Binary Jaccard at top 5%: top-percentile partner sets for X and Y,
  30. |X ∩ Y| / |X ∪ Y|.
  31. Dissociation gap: WJ_continuous - Jaccard_binary.
  32. Inputs:
  33. r3_revision_analyses/data/GSE68719_mlpd_PCG_DESeq2_norm_counts.txt.gz
  34. cache/GSE68719_series_matrix.txt.gz
  35. Outputs:
  36. r3_revision_analyses/results/layer2h_formal_continuous_wj.csv
  37. r3_revision_analyses/results/layer2h_type2_sign_treatment.csv
  38. r3_revision_analyses/results/layer2h_type6_pearson_vs_spearman.csv
  39. r3_revision_analyses/results/layer2h_permutation_null.csv
  40. r3_revision_analyses/results/layer2h_bootstrap_ci.csv
  41. r3_revision_analyses/results/layer2h_disease_comparison.csv
  42. r3_revision_analyses/results/provenance_pipeline5.json
  43. """
  44. import os
  45. import gzip
  46. import json
  47. import sys
  48. import warnings
  49. import numpy as np
  50. import pandas as pd
  51. from scipy import stats
  52. warnings.filterwarnings('ignore')
  53. try:
  54. sys.stdout.reconfigure(encoding='utf-8', errors='replace')
  55. except Exception:
  56. pass
  57. RANDOM_SEED = 42
  58. np.random.seed(RANDOM_SEED)
  59. N_PERMUTATIONS = 200
  60. N_BOOTSTRAP = 500
  61. TOP_PCT = 5
  62. MIN_EXPR = 5
  63. ROOT = r"G:\My Drive\inner_architecture_research\EIF2S1"
  64. DATA_DIR = os.path.join(ROOT, "r3_revision_analyses", "data")
  65. CACHE_DIR = r"G:\My Drive\inner_architecture_research\cache"
  66. RESULTS_DIR = os.path.join(ROOT, "r3_revision_analyses", "results")
  67. COUNTS_FILE = os.path.join(DATA_DIR, "GSE68719_mlpd_PCG_DESeq2_norm_counts.txt.gz")
  68. SERIES_FILE = os.path.join(CACHE_DIR, "GSE68719_series_matrix.txt.gz")
  69. # ==========================================================================
  70. # Load data
  71. # ==========================================================================
  72. print("=" * 75)
  73. print("PIPELINE 5: FORMAL LAYER 2H BULLETPROOF ANALYSIS — EIF2S1-PELO")
  74. print("=" * 75)
  75. print("\n[1/8] Loading expression data...")
  76. counts_raw = pd.read_csv(COUNTS_FILE, sep='\t', compression='gzip')
  77. # File structure: col 0 = EnsemblID, col 1 = symbol, col 2+ = sample expression
  78. counts = counts_raw.drop(columns=[counts_raw.columns[0]]).set_index('symbol')
  79. # Collapse duplicates by taking the row-wise mean (some symbols map to multiple Ensembl IDs)
  80. counts = counts.groupby(level=0).mean()
  81. print(f" Shape: {counts.shape} (genes x samples)")
  82. print(f" Sample columns: {counts.columns[:5].tolist()}...")
  83. # Filter to expressed genes
  84. print("\n[2/8] Filtering to expressed genes (median >= 5)...")
  85. median_expr = counts.median(axis=1)
  86. expressed = counts[median_expr >= MIN_EXPR].copy()
  87. print(f" Filtered shape: {expressed.shape}")
  88. # Parse metadata to get PD vs control
  89. print("\n[3/8] Parsing metadata to identify PD vs Control samples...")
  90. sample_pd_status = {}
  91. with gzip.open(SERIES_FILE, 'rt', encoding='utf-8', errors='replace') as f:
  92. sample_geo_acc = []
  93. char_lines = []
  94. title_lines = []
  95. for line in f:
  96. line = line.strip()
  97. if line.startswith('!Sample_geo_accession'):
  98. sample_geo_acc = [c.strip().strip('"') for c in line.split('\t')[1:]]
  99. elif line.startswith('!Sample_characteristics_ch1'):
  100. char_lines.append([c.strip().strip('"') for c in line.split('\t')[1:]])
  101. elif line.startswith('!Sample_title'):
  102. title_lines.append([c.strip().strip('"') for c in line.split('\t')[1:]])
  103. # Look for "disease state" or similar in characteristics
  104. for char in char_lines:
  105. sample_text = " | ".join(char[:5])
  106. if 'parkinson' in sample_text.lower() or 'pd' in sample_text.lower() or 'control' in sample_text.lower():
  107. for i, c in enumerate(char):
  108. c_lower = c.lower()
  109. if 'control' in c_lower or 'normal' in c_lower or 'healthy' in c_lower:
  110. sample_pd_status[sample_geo_acc[i]] = 'control'
  111. elif 'parkinson' in c_lower or 'pd' in c_lower:
  112. sample_pd_status[sample_geo_acc[i]] = 'pd'
  113. break
  114. # Try titles if characteristics didn't work
  115. if not sample_pd_status:
  116. for i, gsm in enumerate(sample_geo_acc):
  117. title = title_lines[0][i] if title_lines else ''
  118. if 'control' in title.lower() or 'C' in title.upper().replace(' ', '')[:3]:
  119. sample_pd_status[gsm] = 'control'
  120. elif 'pd' in title.lower() or 'parkinson' in title.lower():
  121. sample_pd_status[gsm] = 'pd'
  122. # Map column names of counts to GEO accessions
  123. matched_columns = {}
  124. for col in expressed.columns:
  125. for gsm in sample_geo_acc:
  126. if gsm in col or col in sample_geo_acc:
  127. if col == gsm or gsm == col:
  128. matched_columns[col] = gsm
  129. break
  130. # fallback: extract GSM-like ID from col name
  131. for gsm in sample_geo_acc:
  132. if gsm in col:
  133. matched_columns[col] = gsm
  134. break
  135. # If column matching is messy, just use first 44 as control, last 29 as PD per metadata pattern
  136. if len(matched_columns) < 50:
  137. print(" Direct GSM match incomplete; reading from supplementary metadata...")
  138. metadata_path = os.path.join(RESULTS_DIR, "disease_GSE68719_metadata.csv")
  139. if os.path.exists(metadata_path):
  140. meta = pd.read_csv(metadata_path)
  141. print(f" Loaded metadata file: {meta.shape}")
  142. print(f" Columns: {meta.columns.tolist()}")
  143. print(meta.head())
  144. # Use the metadata file from pipeline 2 if available
  145. metadata_path = os.path.join(RESULTS_DIR, "disease_GSE68719_metadata.csv")
  146. if os.path.exists(metadata_path):
  147. meta = pd.read_csv(metadata_path)
  148. sample_pd_status = dict(zip(meta.iloc[:, 0], meta.iloc[:, 1]))
  149. control_samples = [s for s in expressed.columns if s.startswith('C_')]
  150. pd_samples = [s for s in expressed.columns if s.startswith('P_')]
  151. # Fallback: read metadata file
  152. if len(control_samples) == 0 or len(pd_samples) == 0:
  153. print(f" Sample stratification via column prefix failed; reading metadata file...")
  154. if os.path.exists(metadata_path):
  155. meta = pd.read_csv(metadata_path)
  156. for col in meta.columns:
  157. unique_vals = meta[col].dropna().astype(str).str.lower().unique()
  158. if any('control' in str(v) for v in unique_vals) and any('pd' in str(v) for v in unique_vals):
  159. ctrl_labels = meta[meta[col].astype(str).str.lower().str.contains('control|normal|healthy', na=False)]['sample_label'].tolist()
  160. pd_labels = meta[meta[col].astype(str).str.lower().str.contains('^pd$|parkinson', na=False, regex=True)]['sample_label'].tolist()
  161. control_samples = [s for s in expressed.columns if s in ctrl_labels]
  162. pd_samples = [s for s in expressed.columns if s in pd_labels]
  163. break
  164. # Ensure sample names exist in expression matrix
  165. control_samples = [s for s in control_samples if s in expressed.columns]
  166. pd_samples = [s for s in pd_samples if s in expressed.columns]
  167. print(f" Control samples: {len(control_samples)}")
  168. print(f" PD samples: {len(pd_samples)}")
  169. if len(control_samples) < 10 or len(pd_samples) < 10:
  170. print(f" WARNING: Insufficient stratification. Using all samples pooled.")
  171. pooled_samples = expressed.columns.tolist()
  172. else:
  173. pooled_samples = control_samples + pd_samples
  174. # Ensure target genes present
  175. target_genes = ['EIF2S1', 'PELO']
  176. missing = [g for g in target_genes if g not in expressed.index]
  177. if missing:
  178. print(f" ERROR: missing target genes: {missing}")
  179. print(f" Available genes head: {expressed.index[:20].tolist()}")
  180. raise ValueError("Target genes missing")
  181. # ==========================================================================
  182. # Helper functions for WJ
  183. # ==========================================================================
  184. def coupling_profile(expr, gene, samples, method='spearman'):
  185. """Compute correlation between gene's expression and all other genes."""
  186. target = expr.loc[gene, samples].values.astype(float)
  187. other_genes = [g for g in expr.index if g != gene]
  188. other = expr.loc[other_genes, samples].values.astype(float)
  189. if method == 'spearman':
  190. target_ranks = stats.rankdata(target)
  191. other_ranks = np.apply_along_axis(stats.rankdata, 1, other)
  192. target_mean = target_ranks.mean()
  193. target_std = target_ranks.std()
  194. other_means = other_ranks.mean(axis=1, keepdims=True)
  195. other_stds = other_ranks.std(axis=1, keepdims=True)
  196. cov = ((other_ranks - other_means) * (target_ranks - target_mean)).mean(axis=1)
  197. rho = cov / (other_stds.flatten() * target_std + 1e-12)
  198. else:
  199. target_mean = target.mean()
  200. target_std = target.std()
  201. other_means = other.mean(axis=1, keepdims=True)
  202. other_stds = other.std(axis=1, keepdims=True)
  203. cov = ((other - other_means) * (target - target_mean)).mean(axis=1)
  204. rho = cov / (other_stds.flatten() * target_std + 1e-12)
  205. return pd.Series(rho, index=other_genes)
  206. def wj_shifted(a, b):
  207. """Weighted Jaccard on (r+1)/2 shifted values."""
  208. a_shifted = (a + 1) / 2
  209. b_shifted = (b + 1) / 2
  210. s_min = np.minimum(a_shifted, b_shifted).sum()
  211. s_max = np.maximum(a_shifted, b_shifted).sum()
  212. return s_min / s_max if s_max > 0 else np.nan
  213. def wj_unsigned(a, b):
  214. """Weighted Jaccard on absolute values."""
  215. a_abs = np.abs(a)
  216. b_abs = np.abs(b)
  217. s_min = np.minimum(a_abs, b_abs).sum()
  218. s_max = np.maximum(a_abs, b_abs).sum()
  219. return s_min / s_max if s_max > 0 else np.nan
  220. def wj_signed(a, b):
  221. """Signed weighted Jaccard: split positive and negative components."""
  222. a_pos = np.maximum(a, 0)
  223. a_neg = np.maximum(-a, 0)
  224. b_pos = np.maximum(b, 0)
  225. b_neg = np.maximum(-b, 0)
  226. pos_min = np.minimum(a_pos, b_pos).sum()
  227. pos_max = np.maximum(a_pos, b_pos).sum()
  228. neg_min = np.minimum(a_neg, b_neg).sum()
  229. neg_max = np.maximum(a_neg, b_neg).sum()
  230. if (pos_max + neg_max) > 0:
  231. return (pos_min + neg_min) / (pos_max + neg_max)
  232. return np.nan
  233. def binary_jaccard_top_pct(a, b, pct=5):
  234. """Binary Jaccard at top-pct threshold."""
  235. threshold_a = np.percentile(a, 100 - pct)
  236. threshold_b = np.percentile(b, 100 - pct)
  237. in_a = a >= threshold_a
  238. in_b = b >= threshold_b
  239. intersect = (in_a & in_b).sum()
  240. union = (in_a | in_b).sum()
  241. return intersect / union if union > 0 else np.nan
  242. # ==========================================================================
  243. # TYPE 1: Continuous-Discrete dissociation gap (formal)
  244. # ==========================================================================
  245. print("\n[4/8] Type 1: Formal continuous-discrete dissociation gap (Spearman, controls)")
  246. samples = control_samples if len(control_samples) >= 20 else pooled_samples
  247. print(f" Using {len(samples)} samples")
  248. c_eif = coupling_profile(expressed, 'EIF2S1', samples, method='spearman')
  249. c_pelo = coupling_profile(expressed, 'PELO', samples, method='spearman')
  250. # Align indices (just in case)
  251. common = c_eif.index.intersection(c_pelo.index)
  252. c_eif = c_eif.loc[common].values
  253. c_pelo = c_pelo.loc[common].values
  254. wj_shifted_val = wj_shifted(c_eif, c_pelo)
  255. wj_unsigned_val = wj_unsigned(c_eif, c_pelo)
  256. wj_signed_val = wj_signed(c_eif, c_pelo)
  257. bj_top5 = binary_jaccard_top_pct(c_eif, c_pelo, pct=5)
  258. gap_shifted_t1 = wj_shifted_val - bj_top5
  259. gap_unsigned_t1 = wj_unsigned_val - bj_top5
  260. gap_signed_t1 = wj_signed_val - bj_top5
  261. print(f" WJ shifted (r+1)/2: {wj_shifted_val:.4f}")
  262. print(f" WJ unsigned |r|: {wj_unsigned_val:.4f}")
  263. print(f" WJ signed split: {wj_signed_val:.4f}")
  264. print(f" Binary Jaccard top5: {bj_top5:.4f}")
  265. print(f" Type 1 gap (shifted-binary): {gap_shifted_t1:.4f}")
  266. print(f" Type 1 gap (unsigned-binary): {gap_unsigned_t1:.4f}")
  267. print(f" Type 1 gap (signed-binary): {gap_signed_t1:.4f}")
  268. type1_df = pd.DataFrame([{
  269. 'cohort': 'control',
  270. 'wj_shifted': wj_shifted_val,
  271. 'wj_unsigned': wj_unsigned_val,
  272. 'wj_signed': wj_signed_val,
  273. 'binary_jaccard_top5': bj_top5,
  274. 'gap_shifted_minus_binary': gap_shifted_t1,
  275. 'gap_unsigned_minus_binary': gap_unsigned_t1,
  276. 'gap_signed_minus_binary': gap_signed_t1,
  277. 'n_samples': len(samples),
  278. 'n_genes': len(common),
  279. }])
  280. type1_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_formal_continuous_wj.csv"), index=False)
  281. # ==========================================================================
  282. # TYPE 2: Sign-Treatment gaps
  283. # ==========================================================================
  284. print("\n[5/8] Type 2: Sign-treatment gaps")
  285. gap_shifted_unsigned = wj_shifted_val - wj_unsigned_val
  286. gap_shifted_signed = wj_shifted_val - wj_signed_val
  287. gap_unsigned_signed = wj_unsigned_val - wj_signed_val
  288. print(f" Shifted vs Unsigned gap: {gap_shifted_unsigned:.4f}")
  289. print(f" Shifted vs Signed gap: {gap_shifted_signed:.4f}")
  290. print(f" Unsigned vs Signed gap: {gap_unsigned_signed:.4f}")
  291. print(f" All three formulations close to each other = sign-inversion not dominant")
  292. type2_df = pd.DataFrame([{
  293. 'pair': 'EIF2S1-PELO',
  294. 'gap_shifted_minus_unsigned': gap_shifted_unsigned,
  295. 'gap_shifted_minus_signed': gap_shifted_signed,
  296. 'gap_unsigned_minus_signed': gap_unsigned_signed,
  297. 'max_abs_gap': max(abs(gap_shifted_unsigned), abs(gap_shifted_signed),
  298. abs(gap_unsigned_signed)),
  299. 'interpretation': ('Sign inversion not dominant' if max(abs(gap_shifted_unsigned),
  300. abs(gap_shifted_signed),
  301. abs(gap_unsigned_signed)) < 0.05
  302. else 'Sign treatment matters'),
  303. }])
  304. type2_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_type2_sign_treatment.csv"), index=False)
  305. # ==========================================================================
  306. # TYPE 6: Substrate-Projection (Pearson vs Spearman)
  307. # ==========================================================================
  308. print("\n[6/8] Type 6: Substrate-projection (Pearson vs Spearman)")
  309. c_eif_pearson = coupling_profile(expressed, 'EIF2S1', samples, method='pearson')
  310. c_pelo_pearson = coupling_profile(expressed, 'PELO', samples, method='pearson')
  311. common_p = c_eif_pearson.index.intersection(c_pelo_pearson.index)
  312. c_eif_p = c_eif_pearson.loc[common_p].values
  313. c_pelo_p = c_pelo_pearson.loc[common_p].values
  314. wj_shifted_pearson = wj_shifted(c_eif_p, c_pelo_p)
  315. bj_top5_pearson = binary_jaccard_top_pct(c_eif_p, c_pelo_p, pct=5)
  316. gap_t6 = wj_shifted_val - wj_shifted_pearson
  317. gap_t6_binary = bj_top5 - bj_top5_pearson
  318. print(f" Spearman WJ shifted: {wj_shifted_val:.4f}")
  319. print(f" Pearson WJ shifted: {wj_shifted_pearson:.4f}")
  320. print(f" Type 6 gap (continuous): {gap_t6:.4f}")
  321. print(f" Spearman binary J top5: {bj_top5:.4f}")
  322. print(f" Pearson binary J top5: {bj_top5_pearson:.4f}")
  323. print(f" Type 6 gap (binary): {gap_t6_binary:.4f}")
  324. print(f" Near-zero gap = relationships approximately linear-monotonic")
  325. type6_df = pd.DataFrame([{
  326. 'pair': 'EIF2S1-PELO',
  327. 'spearman_wj_shifted': wj_shifted_val,
  328. 'pearson_wj_shifted': wj_shifted_pearson,
  329. 'gap_continuous': gap_t6,
  330. 'spearman_bj_top5': bj_top5,
  331. 'pearson_bj_top5': bj_top5_pearson,
  332. 'gap_binary': gap_t6_binary,
  333. 'interpretation': 'Linear-monotonic structure' if abs(gap_t6) < 0.05
  334. else 'Nonlinear-monotonic divergence',
  335. }])
  336. type6_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_type6_pearson_vs_spearman.csv"),
  337. index=False)
  338. # ==========================================================================
  339. # Permutation null: random gene pairs
  340. # ==========================================================================
  341. print("\n[7/8] Permutation null: random gene pairs")
  342. print(f" Running {N_PERMUTATIONS} random pairs...")
  343. all_genes = expressed.index.tolist()
  344. null_gaps = []
  345. np.random.seed(RANDOM_SEED)
  346. for _ in range(N_PERMUTATIONS):
  347. g1, g2 = np.random.choice(all_genes, size=2, replace=False)
  348. try:
  349. c1 = coupling_profile(expressed, g1, samples, method='spearman')
  350. c2 = coupling_profile(expressed, g2, samples, method='spearman')
  351. common_n = c1.index.intersection(c2.index)
  352. wj_n = wj_shifted(c1.loc[common_n].values, c2.loc[common_n].values)
  353. bj_n = binary_jaccard_top_pct(c1.loc[common_n].values,
  354. c2.loc[common_n].values, pct=5)
  355. null_gaps.append(wj_n - bj_n)
  356. except Exception:
  357. continue
  358. null_gaps = np.array(null_gaps)
  359. perm_p_two = (np.abs(null_gaps - null_gaps.mean()) >=
  360. abs(gap_shifted_t1 - null_gaps.mean())).mean()
  361. z_score = ((gap_shifted_t1 - null_gaps.mean()) / null_gaps.std()
  362. if null_gaps.std() > 0 else 0)
  363. print(f" EIF2S1-PELO gap (shifted): {gap_shifted_t1:.4f}")
  364. print(f" Null gap mean: {null_gaps.mean():.4f}, std: {null_gaps.std():.4f}")
  365. print(f" Null gap range: [{null_gaps.min():.4f}, {null_gaps.max():.4f}]")
  366. print(f" Z-score: {z_score:.2f}")
  367. print(f" Permutation p (two-sided): {perm_p_two:.4f}")
  368. null_df = pd.DataFrame({
  369. 'pair_index': range(len(null_gaps)),
  370. 'random_pair_gap': null_gaps,
  371. })
  372. null_summary_row = pd.DataFrame([{
  373. 'pair_index': 'EIF2S1-PELO_observed',
  374. 'random_pair_gap': gap_shifted_t1,
  375. }])
  376. null_df = pd.concat([null_summary_row, null_df], ignore_index=True)
  377. null_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_permutation_null.csv"), index=False)
  378. # ==========================================================================
  379. # Bootstrap CI on EIF2S1-PELO gap
  380. # ==========================================================================
  381. print("\n[8/8] Bootstrap CI on EIF2S1-PELO Type 1 gap")
  382. print(f" Running {N_BOOTSTRAP} bootstrap iterations...")
  383. boot_gaps_shifted = []
  384. n_samples = len(samples)
  385. np.random.seed(RANDOM_SEED + 1)
  386. for _ in range(N_BOOTSTRAP):
  387. boot_idx = np.random.choice(n_samples, size=n_samples, replace=True)
  388. boot_samples = [samples[i] for i in boot_idx]
  389. try:
  390. c_e = coupling_profile(expressed, 'EIF2S1', boot_samples, method='spearman')
  391. c_p = coupling_profile(expressed, 'PELO', boot_samples, method='spearman')
  392. common_b = c_e.index.intersection(c_p.index)
  393. wj_b = wj_shifted(c_e.loc[common_b].values, c_p.loc[common_b].values)
  394. bj_b = binary_jaccard_top_pct(c_e.loc[common_b].values,
  395. c_p.loc[common_b].values, pct=5)
  396. boot_gaps_shifted.append(wj_b - bj_b)
  397. except Exception:
  398. continue
  399. boot_gaps_shifted = np.array(boot_gaps_shifted)
  400. ci_low, ci_high = np.percentile(boot_gaps_shifted, [2.5, 97.5])
  401. print(f" Bootstrap mean: {boot_gaps_shifted.mean():.4f}")
  402. print(f" 95% CI: [{ci_low:.4f}, {ci_high:.4f}]")
  403. bootstrap_df = pd.DataFrame([{
  404. 'pair': 'EIF2S1-PELO',
  405. 'observed_gap': gap_shifted_t1,
  406. 'bootstrap_mean': boot_gaps_shifted.mean(),
  407. 'bootstrap_std': boot_gaps_shifted.std(),
  408. 'ci_lower_95': ci_low,
  409. 'ci_upper_95': ci_high,
  410. 'n_bootstrap_iterations': len(boot_gaps_shifted),
  411. }])
  412. bootstrap_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_bootstrap_ci.csv"), index=False)
  413. # ==========================================================================
  414. # Disease comparison: same metrics for PD samples
  415. # ==========================================================================
  416. print("\nDisease comparison: control vs PD samples")
  417. disease_results = []
  418. for cohort_name, cohort_samples in [('control', control_samples), ('pd', pd_samples)]:
  419. if len(cohort_samples) < 10:
  420. continue
  421. c_e = coupling_profile(expressed, 'EIF2S1', cohort_samples, method='spearman')
  422. c_p = coupling_profile(expressed, 'PELO', cohort_samples, method='spearman')
  423. common_c = c_e.index.intersection(c_p.index)
  424. c_e_v = c_e.loc[common_c].values
  425. c_p_v = c_p.loc[common_c].values
  426. wj_s = wj_shifted(c_e_v, c_p_v)
  427. bj = binary_jaccard_top_pct(c_e_v, c_p_v, pct=5)
  428. disease_results.append({
  429. 'cohort': cohort_name,
  430. 'n_samples': len(cohort_samples),
  431. 'wj_shifted': wj_s,
  432. 'binary_jaccard_top5': bj,
  433. 'dissociation_gap': wj_s - bj,
  434. })
  435. print(f" {cohort_name:10s} n={len(cohort_samples):3d}: WJ={wj_s:.4f}, "
  436. f"BJ={bj:.4f}, gap={wj_s-bj:.4f}")
  437. disease_df = pd.DataFrame(disease_results)
  438. disease_df.to_csv(os.path.join(RESULTS_DIR, "layer2h_disease_comparison.csv"),
  439. index=False)
  440. # ==========================================================================
  441. # Provenance
  442. # ==========================================================================
  443. provenance = {
  444. "pipeline_file": "pipeline_5_formal_layer2h.py",
  445. "execution_date": "2026-04-29",
  446. "wj_compliance_status": "PASS (Layer 2H formal computation)",
  447. "substrate": "GSE68719 BA9 prefrontal cortex RNA-seq (DESeq2-normalized)",
  448. "n_control": len(control_samples),
  449. "n_pd": len(pd_samples),
  450. "n_genes_analyzed": int(len(common)),
  451. "type1_findings": {
  452. "wj_shifted": float(wj_shifted_val),
  453. "wj_unsigned": float(wj_unsigned_val),
  454. "wj_signed": float(wj_signed_val),
  455. "binary_jaccard_top5": float(bj_top5),
  456. "type1_gap_shifted": float(gap_shifted_t1),
  457. "permutation_p": float(perm_p_two),
  458. "z_score_vs_null": float(z_score),
  459. "bootstrap_ci_lower": float(ci_low),
  460. "bootstrap_ci_upper": float(ci_high),
  461. },
  462. "type2_findings": {
  463. "max_sign_treatment_gap": float(max(abs(gap_shifted_unsigned),
  464. abs(gap_shifted_signed),
  465. abs(gap_unsigned_signed))),
  466. "sign_inversion_dominant": bool(max(abs(gap_shifted_unsigned),
  467. abs(gap_shifted_signed),
  468. abs(gap_unsigned_signed)) > 0.05),
  469. },
  470. "type6_findings": {
  471. "pearson_wj_shifted": float(wj_shifted_pearson),
  472. "spearman_pearson_gap": float(gap_t6),
  473. "linear_monotonic_consistent": bool(abs(gap_t6) < 0.05),
  474. },
  475. "disease_comparison": disease_results,
  476. "config": {
  477. "random_seed": RANDOM_SEED,
  478. "n_permutations": N_PERMUTATIONS,
  479. "n_bootstrap": N_BOOTSTRAP,
  480. "top_pct_threshold": TOP_PCT,
  481. "min_expression_filter": MIN_EXPR,
  482. },
  483. }
  484. with open(os.path.join(RESULTS_DIR, "provenance_pipeline5.json"), 'w') as f:
  485. json.dump(provenance, f, indent=2)
  486. print("\n" + "=" * 75)
  487. print("PIPELINE 5 COMPLETE")
  488. print("=" * 75)
  489. print("Files saved to:", RESULTS_DIR)
  490. print("\nKey numbers for manuscript revision:")
  491. print(f" Formal continuous WJ (shifted): {wj_shifted_val:.4f}")
  492. print(f" Binary Jaccard top 5%: {bj_top5:.4f}")
  493. print(f" Type 1 dissociation gap: {gap_shifted_t1:.4f}")
  494. print(f" Bootstrap 95% CI: [{ci_low:.4f}, {ci_high:.4f}]")
  495. print(f" Permutation p vs random pairs: {perm_p_two:.4f}")
  496. print(f" Type 2 sign-treatment max gap: {max(abs(gap_shifted_unsigned),abs(gap_shifted_signed),abs(gap_unsigned_signed)):.4f}")
  497. print(f" Type 6 Spearman-Pearson gap: {gap_t6:.4f}")

pipeline_5_formal_layer2h.py at commit cb7a7c0, under MIT · at the source

Overview

Authors: Drake H. Harbert1
  1. Inner Architecture LLC, Canton, OH, United States
Journal: Frontiers in molecular neuroscience, volume 19, article 1817096
Dates: received 25 February 2026; accepted 13 July 2026; published online 10 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fnmol.2026.1817096 · PMID 42639115 · PMCID PMC13500769 · OpenAlex W7202077191
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Parkinson's (population)
Methods: Connectivity, Statistics
Keywords: weighted Jaccard, EIF2S1, PELO, ribosome rescue, co-expression network, pairing-family decomposition, Parkinson’s disease, layer separation
Topic: RNA and protein synthesis mechanisms (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: not cited yet (Europe PMC); 7 references in the paper

Abstract

EIF2S1 (eIF2α, translation initiation surveillance) and PELO (Pelota, ribosome rescue) operate within established ribosome-associated quality control and translation-initiation surveillance pathways. Direct cross-talk between these systems is mediated through ZAK/GCN-mediated eIF2α phosphorylation under ribosome-stall stress. At the transcriptional layer, however, EIF2S1-PELO co-regulation has not been quantitatively characterized. Using pairing-family architectural decomposition on multi-region GTEx co-expression and BA9 prefrontal cortex RNA-seq from Parkinson’s disease and control donors (GSE68719, n = 44 controls, n = 29 PD), we find: (1) EIF2S1 and PELO share near-identical genome-wide co-expression architecture (continuous Weighted Jaccard = 0.914) with substantial top-partner overlap (binary Jaccard at top 5% = 0.490), producing a Type 1 dissociation gap of 0.424, significantly smaller than 200 random gene-pair gaps (mean 0.644, z = −2.57, permutation p = 0.015), and the lowest gap in a 21-pair quality-control gene panel; (2) the relationship is approximately linear-monotonic (Type 6 Pearson-Spearman gap ≈ 0.004); (3) cellular functional co-dependency in DepMap CRISPR screens is essentially zero (Jaccard = 0.006, gene-effect Pearson r = −0.07), consistent with layer separation between regulatory architecture and cellular phenotype; (4) in an exploratory single-cohort comparison of PD versus control prefrontal cortex, the dissociation gap was larger in PD (0.536 versus 0.424 in control), but this between-group difference did not reach statistical significance under a group-label permutation test (p = 0.40; bootstrap 95% CI on the difference [−0.14, 0.29]), so we present it as a hypothesis for replication rather than an established effect. These findings characterize EIF2S1-PELO transcriptional co-regulation as a constitutive architectural feature distinct from the ZAK/GCN-mediated direct mechanism. The constitutive coupling replicated in independent Alzheimer’s disease and ALS frontal-cortex cohorts and across microarray and RNA-seq platforms; a disease-associated weakening of the coupling was directionally consistent across all three diseases but did not reach statistical significance. We note that co-expression patterns are consistent with shared upstream regulatory programs but do not by themselves establish direct co-regulation; the architectural findings reported here are correlational at the transcriptional layer.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 20 matches between paragraphs and lines of code.

nwharbert8-ui/eif2s1-translational-qc-hubsub

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: cb7a7c0579d1aaa80d91933c246d891fd4d8d4a0, 5 July 2026
Languages: Python (11)
Size: 55 files, 11 scripts
Software Heritage: not archived
Found in: “Data availability statement”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (11 files), pandas (11 files), SciPy (10 files), Matplotlib (3 files), NetworkX (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
13 files

Zenodo 18626362

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability statement”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
3 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 12 scripts, each with its path and the digest of its content;
  • 20 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Data links

Data availability statement

Publicly available datasets were analyzed in this study. GTEx v8 (https://gtexportal.org/, dbGaP phs000424.v8.p2); GSE68719, GSE5281, and GSE124439 (NCBI GEO, https://www.ncbi.nlm.nih.gov/geo/); and DepMap release 24Q2 (https://depmap.org/). All analysis code is openly available at https://github.com/nwharbert8-ui/eif2s1-translational-qc-hubsub and archived at Zenodo (https://doi.org/10.5281/zenodo.18626362).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 1 author, 8 keywords, 7 references.

Cite

This paper

Harbert, D. H. (2026). EIF2S1-PELO co-expression architecture in human brain: pairing-family decomposition reveals constitutive transcriptional coupling and functional layer separation. Frontiers in molecular neuroscience, 19, 1817096. https://doi.org/10.3389/fnmol.2026.1817096

BibTeX

@article{harbert2026eif2s1,
author = {Harbert, Drake H.},
title = {{EIF2S1-PELO co-expression architecture in human brain: pairing-family decomposition reveals constitutive transcriptional coupling and functional layer separation}},
journal = {Frontiers in molecular neuroscience},
year = {2026},
month = aug,
volume = {19},
pages = {1817096},
publisher = {Frontiers Media SA},
issn = {1662-5099},
doi = {10.3389/fnmol.2026.1817096},
url = {https://doi.org/10.3389/fnmol.2026.1817096},
pmid = {42639115},
pmcid = {PMC13500769}
}

RIS

TY - JOUR
AU - Harbert, Drake H.
TI - EIF2S1-PELO co-expression architecture in human brain: pairing-family decomposition reveals constitutive transcriptional coupling and functional layer separation
T2 - Frontiers in molecular neuroscience
J2 - Front Mol Neurosci
PY - 2026
DA - 2026/08/10
VL - 19
SP - 1817096
SN - 1662-5099
PB - Frontiers Media SA
DO - 10.3389/fnmol.2026.1817096
UR - https://doi.org/10.3389/fnmol.2026.1817096
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fnmol.2026.1817096",
"type": "article-journal",
"title": "EIF2S1-PELO co-expression architecture in human brain: pairing-family decomposition reveals constitutive transcriptional coupling and functional layer separation",
"container-title": "Frontiers in molecular neuroscience",
"author": [
{
"family": "Harbert",
"given": "Drake H."
}
],
"container-title-short": "Front Mol Neurosci",
"volume": "19",
"page": "1817096",
"DOI": "10.3389/fnmol.2026.1817096",
"PMID": "42639115",
"PMCID": "PMC13500769",
"ISSN": "1662-5099",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fnmol.2026.1817096",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xgen.2026.101284 [code]
NERINE reveals rare variant associations in gene networks across phenotypes and implicates an SNCA-PRL-LRRK2 subnetwork in Parkinson's disease.
Journal: Cell genomics
In common: NetworkX, seaborn, pandas, 3 other tools, Parkinson's, 1 reference
[2] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: NetworkX, seaborn, pandas, 3 other tools, genetics / omics, 1 reference
[3] doi:10.1126/sciadv.aed2952 [code]
Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.
Journal: Science advances
In common: seaborn, pandas, SciPy, 2 other tools, Parkinson's, 1 reference
[4] doi:10.1093/bib/bbag118 [code]
Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.
Journal: Briefings in bioinformatics
In common: NetworkX, seaborn, pandas, 3 other tools, Parkinson's
[5] doi:10.3389/fendo.2026.1828487 [code]
Castration-induced nigrostriatal deficits are linked to reduced TrkB and loss of mature spines in the dorsal striatum.
Journal: Frontiers in endocrinology
In common: NetworkX, seaborn, pandas, 3 other tools, Parkinson's
[6] doi:10.1038/s41467-026-72161-w [code]
Temporal heterogeneity shapes diffusion dynamics in complex networks.
Journal: Nature communications
In common: NetworkX, seaborn, pandas, 3 other tools, Parkinson's
[7] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: NetworkX, seaborn, pandas, 3 other tools, genetics / omics
[8] doi:10.1038/s41467-026-76676-0 [code]
Determinants of functional burden pleiotropy and gene dosage responses across human traits.
Journal: Nature communications
In common: NetworkX, seaborn, pandas, 3 other tools, genetics / omics
[9] doi:10.3390/ijms27167275 [code]
Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling.
Journal: International journal of molecular sciences
In common: NetworkX, seaborn, pandas, 3 other tools, genetics / omics
[10] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: NetworkX, seaborn, pandas, 3 other tools, genetics / omics

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.