OSCR

Sigma-1 and Sigma-2 receptors exhibit divergent genome-wide Co-expression architectures in human brain despite shared subcellular localization.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Materials and methods › Data source and processing ↔ MS3_Sigma_Divergence_Pipeline.py, lines 85–126 · score 0.85 · Anterior cingulate cortex, median TPM, Brain Frontal Cortex, GTEx v8, Nucleus accumbens, BA24
  2. [2] § Materials and methods › Data source and processing ↔ MS3_sigma_divergence_wj_pipeline.py, lines 91–101 · score 0.80 · Anterior cingulate cortex, median TPM, Brain Frontal Cortex, Nucleus accumbens, BA24, Putamen
  3. [3] § Results › Divergent tails encode distinct functional programs ↔ MS3_sigma_divergence_wj_pipeline.py, lines 916–975 · score 0.73 · TMEM97 unique GO, SIGMAR1 unique GO, TMEM97 Reactome, SIGMAR1 Reactome, SIGMAR1 KEGG, log10
  4. [4] § Results › SIGMAR1 and TMEM97 share global architecture but diverge at functional tails ↔ MS3_sigma_divergence_wj_pipeline.py, lines 1–67 · score 0.69 · genome wide Spearman, Genome wide co, co expression network, expression architecture, binary Jaccard, Weighted Jaccard
  5. [5] § Materials and methods › Genome-wide Co-expression analysis ↔ MS3_Sigma_Divergence_Pipeline.py, lines 85–126 · score 0.66 · ER stress, EIF2S1, target gene, NEMF, HSPA5, PELO
  6. [6] § Materials and methods › Functional enrichment ↔ MS3_Sigma_Divergence_Pipeline.py, lines 1–40 · score 0.65 · gProfiler, TMEM97 top, Custom gene, CC, MF, Reactome
  7. [7] § Materials and methods › Functional enrichment ↔ MS3_sigma_divergence_wj_pipeline.py, lines 653–725 · score 0.62 · gProfiler, Custom gene, TMEM97 unique, SIGMAR1 unique, CC, MF
  8. [8] § Results › Multi-region replication ↔ MS3_sigma_divergence_wj_pipeline.py, lines 91–101 · score 0.61 · NAcc, anterior cingulate cortex, nucleus accumbens, brain regions, BA24, putamen
  9. [9] § Materials and methods › Primary analysis: Weighted Jaccard on continuous vectors ↔ MS3_sigma_divergence_wj_pipeline.py, lines 154–173 · score 0.60 · r_shifted, denominator, numerator, mapped, max, WJ
  10. [10] § Materials and methods › Sensitivity analyses ↔ MS3_sigma_divergence_wj_pipeline.py, lines 1–67 · score 0.58 · covariate adjustment, sensitivity, brain regions, sex, age, deconvolution
  11. [11] § Results › Custom gene set enrichment ↔ MS3_sigma_divergence_wj_pipeline.py, lines 653–725 · score 0.57 · ribosome quality control, MAM mitochondrial, Vascular, pathway, enrichment, TMEM97
  12. [12] § Results › Custom gene set enrichment ↔ MS3_Sigma_Divergence_Pipeline.py, lines 637–678 · score 0.56 · ribosome quality control, MAM mitochondrial, Vascular, pathway, enrichment, TMEM97
  13. [13] § Materials and methods › Sensitivity analyses ↔ MS3_Sigma_Divergence_Pipeline.py, lines 1–40 · score 0.55 · covariate adjustment, brain regions, sex, age, deconvolution, validation
  14. [14] § Results › Divergent tails encode distinct functional programs ↔ MS3_Sigma_Divergence_Pipeline.py, lines 523–592 · score 0.52 · RAB1B, AAMP, PSMD3, YIPF3, partners, overlap
  15. [15] § Results › Sensitivity analyses ↔ MS3_Sigma_Divergence_Pipeline.py, lines 842–909 · score 0.51 · Covariate adjustment, rank preservation, VCP, thresholds, Cell, Spearman

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,084 lines · 47 KB · MIT · 8 matches

  1. #!/usr/bin/env python3
  2. """
  3. ===============================================================================
  4. MS3 SIGMA DIVERGENCE WJ-NATIVE PIPELINE — Local Execution (16GB RAM)
  5. ===============================================================================
  6. MANUSCRIPT: "Sigma-1 and Sigma-2 Receptors Exhibit Divergent Genome-Wide
  7. Co-Expression Networks in Human Brain Despite Shared Subcellular
  8. Localization"
  9. TARGET: Frontiers in Molecular Neuroscience
  10. AUTHOR: Drake H. Harbert (ORCID: 0009-0007-7740-3616)
  11. Inner Architecture LLC, Canton, OH 44721, USA
  12. METHODOLOGY:
  13. PRIMARY: Weighted Jaccard (WJ) on continuous genome-wide Spearman
  14. correlation vectors. Threshold-free comparison of full
  15. co-expression architectures.
  16. SUPPLEMENTARY: Binary Jaccard on top 5% network membership (consistency
  17. with other genomics papers in portfolio).
  18. COMPUTES:
  19. - Genome-wide Spearman correlations for SIGMAR1, TMEM97, and 5 additional
  20. targets across 5 brain regions
  21. - WJ on continuous correlation vectors (primary analysis)
  22. - WJ permutation testing (1000 permutations, seed=42)
  23. - Binary Jaccard on top 5% sets (supplementary)
  24. - FDR correction across all 21 pairwise comparisons
  25. - gProfiler GO/KEGG/Reactome enrichment
  26. - 6 custom gene set enrichments
  27. - Cell-type deconvolution sensitivity
  28. - Covariate adjustment (age, sex)
  29. - Multi-region replication (5 brain regions)
  30. - All figures (300 DPI, colorblind-safe)
  31. - Supplementary Tables S1-S4
  32. - provenance.json
  33. USAGE: py -3 MS3_sigma_divergence_wj_pipeline.py
  34. Dependencies: scipy, statsmodels, pandas, numpy, matplotlib, matplotlib-venn,
  35. gprofiler-official, python-docx, openpyxl, requests
  36. ===============================================================================
  37. """
  38. import os
  39. import gc
  40. import json
  41. import time
  42. import warnings
  43. import requests
  44. import gzip
  45. import numpy as np
  46. import pandas as pd
  47. from scipy import stats
  48. from scipy.stats import spearmanr, fisher_exact
  49. from statsmodels.stats.multitest import multipletests
  50. from collections import OrderedDict
  51. from itertools import combinations
  52. from numpy.linalg import lstsq
  53. import matplotlib
  54. matplotlib.use('Agg')
  55. matplotlib.rcParams['pdf.fonttype'] = 42
  56. matplotlib.rcParams['ps.fonttype'] = 42
  57. import matplotlib.pyplot as plt
  58. from matplotlib_venn import venn2
  59. warnings.filterwarnings('ignore')
  60. # ============================================================================
  61. # CONFIG
  62. # ============================================================================
  63. FORCE_RECOMPUTE = True
  64. RANDOM_SEED = 42
  65. np.random.seed(RANDOM_SEED)
  66. N_PERMUTATIONS = 1000
  67. # Paths
  68. DRIVE_BASE = r"G:\My Drive\inner_architecture_research\MS3_JNC_Submission"
  69. RESULTS_DIR = os.path.join(DRIVE_BASE, "MS3_Results")
  70. FIGURES_DIR = os.path.join(DRIVE_BASE, "MS3_Figures")
  71. SUPPL_DIR = os.path.join(DRIVE_BASE, "MS3_Supplementary_Tables")
  72. DATA_DIR = os.path.join(DRIVE_BASE, "data")
  73. for d in [RESULTS_DIR, FIGURES_DIR, SUPPL_DIR, DATA_DIR]:
  74. os.makedirs(d, exist_ok=True)
  75. # Targets
  76. PRIMARY_TARGETS = ['SIGMAR1', 'TMEM97']
  77. TARGET_GENES = ['EIF2S1', 'PELO', 'LTN1', 'NEMF', 'TMEM97', 'HSPA5', 'SIGMAR1']
  78. # Brain regions
  79. BRAIN_REGIONS = OrderedDict([
  80. ('BA9', 'Brain - Frontal Cortex (BA9)'),
  81. ('Putamen', 'Brain - Putamen (basal ganglia)'),
  82. ('Hippocampus', 'Brain - Hippocampus'),
  83. ('NAcc', 'Brain - Nucleus accumbens (basal ganglia)'),
  84. ('BA24', 'Brain - Anterior cingulate cortex (BA24)'),
  85. ])
  86. PRIMARY_REGION = 'BA9'
  87. MIN_MEDIAN_TPM = 1.0
  88. TOP_PERCENT = 5
  89. # Custom gene sets
  90. MAM_MITO_GENES = [
  91. 'VDAC1', 'VDAC2', 'VDAC3', 'MFN1', 'MFN2', 'RHOT1', 'RHOT2',
  92. 'ITPR1', 'ITPR2', 'ITPR3', 'VAPB', 'RMDN3', 'PACS2', 'FATE1',
  93. ]
  94. SIGMA_NETWORK_GENES = ['SIGMAR1', 'TMEM97', 'PGRMC1', 'NPC1']
  95. ER_STRESS_UPR_GENES = [
  96. 'HSPA5', 'HSP90B1', 'DDIT3', 'ATF4', 'ATF6', 'ERN1', 'EIF2AK3',
  97. 'XBP1', 'DNAJB9', 'HERPUD1', 'EDEM1', 'CALR', 'CANX', 'P4HB',
  98. ]
  99. METHYLATION_GENES = [
  100. 'MAT1A', 'MAT2A', 'MAT2B', 'AHCY', 'AHCYL1', 'AHCYL2',
  101. 'MTR', 'MTRR', 'MTHFR', 'BHMT', 'BHMT2', 'CBS', 'CTH',
  102. 'GNMT', 'PEMT', 'NNMT', 'INMT', 'DNMT1',
  103. ]
  104. VASCULAR_GENES = [
  105. 'PECAM1', 'CDH5', 'VWF', 'FLT1', 'KDR', 'ENG', 'CLDN5', 'ESAM',
  106. 'ERG', 'TIE1', 'TEK', 'ANGPT1', 'ANGPT2', 'NOS3', 'MCAM',
  107. 'PODXL', 'EMCN', 'ROBO4',
  108. ]
  109. RQC_GENES = [
  110. 'PELO', 'HBS1L', 'LTN1', 'NEMF', 'ANKZF1',
  111. 'VCP', 'UFD1', 'NPLOC4', 'ZNF598', 'RACK1', 'ABCE1', 'TCF25',
  112. ]
  113. # Cell-type deconvolution markers
  114. CELLTYPE_MARKERS = {
  115. 'Neurons': ['SNAP25', 'SYT1', 'GAD1', 'GAD2', 'SLC17A7', 'RBFOX3',
  116. 'STMN2', 'SYN1', 'NRGN'],
  117. 'Astrocytes': ['AQP4', 'GFAP', 'SLC1A2', 'SLC1A3', 'ALDH1L1', 'GJA1',
  118. 'S100B', 'SOX9', 'GLUL'],
  119. 'Oligodendrocytes': ['MBP', 'MOG', 'PLP1', 'MAG', 'MOBP', 'CLDN11', 'CNP',
  120. 'OPALIN', 'TF'],
  121. 'Microglia': ['CX3CR1', 'P2RY12', 'CSF1R', 'TMEM119', 'AIF1', 'ITGAM',
  122. 'CD68', 'HEXB', 'TREM2'],
  123. 'Endothelial': ['CLDN5', 'FLT1', 'PECAM1', 'VWF', 'CDH5', 'ERG',
  124. 'ESAM', 'TIE1'],
  125. 'OPCs': ['PDGFRA', 'CSPG4', 'OLIG1', 'OLIG2', 'SOX10', 'NKX2-2',
  126. 'GPR17', 'PCDH15', 'NEU4'],
  127. }
  128. # GTEx URLs
  129. GTEX_TPM_URL = "https://storage.googleapis.com/adult-gtex/bulk-gex/v8/rna-seq/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_tpm.gct.gz"
  130. GTEX_SAMPLE_URL = "https://storage.googleapis.com/adult-gtex/annotations/v8/metadata-files/GTEx_Analysis_v8_Annotations_SampleAttributesDS.txt"
  131. GTEX_SUBJECT_URL = "https://storage.googleapis.com/adult-gtex/annotations/v8/metadata-files/GTEx_Analysis_v8_Annotations_SubjectPhenotypesDS.txt"
  132. # ============================================================================
  133. # CORE WJ FUNCTIONS
  134. # ============================================================================
  135. def weighted_jaccard(vec_a, vec_b):
  136. """
  137. Compute Weighted Jaccard similarity on two continuous vectors.
  138. WJ = sum(min(a_i, b_i)) / sum(max(a_i, b_i))
  139. For correlation vectors that can be negative, shift to [0, 1] range first:
  140. r_shifted = (r + 1) / 2
  141. This maps r=-1 -> 0, r=0 -> 0.5, r=1 -> 1.
  142. """
  143. # Shift correlations from [-1, 1] to [0, 1]
  144. a = (np.array(vec_a) + 1) / 2
  145. b = (np.array(vec_b) + 1) / 2
  146. numerator = np.sum(np.minimum(a, b))
  147. denominator = np.sum(np.maximum(a, b))
  148. if denominator == 0:
  149. return 0.0
  150. return numerator / denominator
  151. def weighted_jaccard_permutation_test(vec_a, vec_b, n_perm=1000, seed=42):
  152. """
  153. Permutation test for WJ significance.
  154. Null hypothesis: the two correlation vectors are drawn from the same
  155. underlying distribution (i.e., the two genes see the same transcriptional
  156. architecture).
  157. Permutation strategy: shuffle the gene labels in one vector, breaking
  158. the gene-to-gene correspondence while preserving the marginal distribution.
  159. """
  160. rng = np.random.RandomState(seed)
  161. observed_wj = weighted_jaccard(vec_a, vec_b)
  162. null_distribution = np.zeros(n_perm)
  163. for i in range(n_perm):
  164. perm_b = rng.permutation(vec_b)
  165. null_distribution[i] = weighted_jaccard(vec_a, perm_b)
  166. # Two-sided: how extreme is the observed WJ?
  167. # For divergence testing: low WJ means more divergent
  168. # p-value = proportion of null WJ values <= observed WJ
  169. p_value = np.mean(null_distribution <= observed_wj)
  170. return observed_wj, p_value, null_distribution
  171. def vectorized_spearman(target_expr, other_expr):
  172. """
  173. Compute Spearman correlation between one target and many genes.
  174. Rank-transforms both, then computes Pearson on ranks.
  175. """
  176. from scipy.stats import rankdata
  177. x_rank = rankdata(target_expr)
  178. n = len(target_expr)
  179. # Rank each row of other_expr
  180. Y_rank = np.apply_along_axis(rankdata, 1, other_expr)
  181. # Pearson on ranks = Spearman
  182. x = x_rank - x_rank.mean()
  183. Y = Y_rank - Y_rank.mean(axis=1, keepdims=True)
  184. x_std = np.sqrt(np.sum(x**2))
  185. Y_std = np.sqrt(np.sum(Y**2, axis=1))
  186. valid = Y_std > 0
  187. r_values = np.full(len(other_expr), np.nan)
  188. r_values[valid] = np.dot(Y[valid], x) / (Y_std[valid] * x_std)
  189. return r_values
  190. def compute_custom_enrichment(network_genes, gene_set_list, set_name,
  191. all_genes_list, verbose=True):
  192. """Fisher's exact test for custom gene set enrichment."""
  193. universe = set(all_genes_list)
  194. expressed = [g for g in gene_set_list if g in universe]
  195. in_network = [g for g in expressed if g in network_genes]
  196. n_expressed = len(expressed)
  197. n_in_network = len(in_network)
  198. n_network = len(network_genes)
  199. n_universe = len(universe)
  200. if n_expressed == 0:
  201. return None
  202. expected = n_expressed * n_network / n_universe
  203. fold = n_in_network / expected if expected > 0 else 0
  204. a = n_in_network
  205. b = n_network - n_in_network
  206. c = n_expressed - n_in_network
  207. d = n_universe - n_network - c
  208. table = np.array([[a, b], [c, d]])
  209. _, p_val = fisher_exact(table, alternative='greater')
  210. result = {
  211. 'gene_set': set_name, 'expressed': n_expressed,
  212. 'in_network': n_in_network, 'fold_enrichment': fold,
  213. 'p_value': p_val, 'genes_found': ', '.join(sorted(in_network)),
  214. }
  215. if verbose:
  216. sig = "+" if p_val < 0.05 else "-"
  217. print(f" {sig} {set_name}: {n_in_network}/{n_expressed}, "
  218. f"{fold:.1f}x, p = {p_val:.2e}")
  219. if in_network:
  220. print(f" Genes: {', '.join(sorted(in_network))}")
  221. return result
  222. def partial_correlation_network(target, expr_df, covar_df, top_pct=5):
  223. """Genome-wide Spearman correlations after regressing covariates."""
  224. valid = covar_df.dropna().index
  225. valid = [s for s in valid if s in expr_df.columns]
  226. n_valid = len(valid)
  227. if n_valid < 30:
  228. return None, None, None, n_valid
  229. expr_sub = expr_df[valid]
  230. covar_matrix = covar_df.loc[valid].values
  231. X = np.column_stack([np.ones(n_valid), covar_matrix])
  232. target_vals = expr_sub.loc[target].values.astype(np.float64)
  233. beta, _, _, _ = lstsq(X, target_vals, rcond=None)
  234. target_resid = target_vals - X @ beta
  235. other_genes = [g for g in expr_sub.index if g != target]
  236. other_vals = expr_sub.loc[other_genes].values.astype(np.float64)
  237. betas = lstsq(X, other_vals.T, rcond=None)[0]
  238. resid = other_vals - (X @ betas).T
  239. r_vals = vectorized_spearman(target_resid, resid)
  240. corr_series = pd.Series(r_vals, index=other_genes).dropna().sort_values(ascending=False)
  241. n_top = int(np.ceil(len(corr_series) * top_pct / 100))
  242. top5_set = set(corr_series.head(n_top).index)
  243. threshold = corr_series.iloc[n_top - 1] if n_top <= len(corr_series) else np.nan
  244. return corr_series, threshold, top5_set, n_valid
  245. def sample_to_subject(sample_id):
  246. parts = sample_id.split('-')
  247. return '-'.join(parts[:2]) if len(parts) >= 2 else sample_id
  248. def download_file(url, dest_path):
  249. """Download a file, reuse if already present."""
  250. if os.path.exists(dest_path):
  251. size_mb = os.path.getsize(dest_path) / 1e6
  252. if size_mb > 1:
  253. print(f" Using cached: {os.path.basename(dest_path)} ({size_mb:.0f} MB)")
  254. return
  255. print(f" Downloading {os.path.basename(dest_path)}...")
  256. resp = requests.get(url, stream=True, timeout=600)
  257. resp.raise_for_status()
  258. total = int(resp.headers.get('content-length', 0))
  259. downloaded = 0
  260. with open(dest_path, 'wb') as f:
  261. for chunk in resp.iter_content(chunk_size=8192 * 16):
  262. f.write(chunk)
  263. downloaded += len(chunk)
  264. if total > 0 and downloaded % (50 * 1024 * 1024) < 8192 * 16:
  265. print(f" {downloaded / 1e6:.0f} / {total / 1e6:.0f} MB")
  266. print(f" Done: {os.path.basename(dest_path)} ({downloaded / 1e6:.0f} MB)")
  267. # ============================================================================
  268. # MAIN PIPELINE
  269. # ============================================================================
  270. def main():
  271. t_start = time.time()
  272. print("=" * 70)
  273. print("MS3 SIGMA DIVERGENCE WJ-NATIVE PIPELINE")
  274. print("Primary: Weighted Jaccard on continuous Spearman vectors")
  275. print("Supplementary: Binary Jaccard on top 5% sets")
  276. print("Target: Frontiers in Molecular Neuroscience")
  277. print("Drake H. Harbert -- Inner Architecture LLC")
  278. print("=" * 70)
  279. # ==================================================================
  280. # STEP 1: DOWNLOAD GTEx v8 DATA
  281. # ==================================================================
  282. print(f"\n{'='*70}\nSTEP 1: DOWNLOAD GTEx v8 DATA\n{'='*70}\n")
  283. tpm_file = os.path.join(DATA_DIR, "GTEx_v8_tpm.gct.gz")
  284. sample_file = os.path.join(DATA_DIR, "GTEx_v8_sample_attributes.txt")
  285. subject_file = os.path.join(DATA_DIR, "GTEx_v8_subject_phenotypes.txt")
  286. download_file(GTEX_TPM_URL, tpm_file)
  287. download_file(GTEX_SAMPLE_URL, sample_file)
  288. download_file(GTEX_SUBJECT_URL, subject_file)
  289. # ==================================================================
  290. # STEP 2: LOAD METADATA
  291. # ==================================================================
  292. print(f"\n{'='*70}\nSTEP 2: LOAD METADATA\n{'='*70}\n")
  293. sample_attr = pd.read_csv(sample_file, sep='\t', low_memory=False)
  294. brain_labels = sample_attr[
  295. sample_attr['SMTSD'].str.contains('Brain', na=False)
  296. ]['SMTSD'].unique()
  297. for key, smtsd in BRAIN_REGIONS.items():
  298. if smtsd not in brain_labels:
  299. raise ValueError(f"Region mismatch: {key}: '{smtsd}'")
  300. region_samples = {}
  301. for key, smtsd in BRAIN_REGIONS.items():
  302. region_samples[key] = list(sample_attr.loc[sample_attr['SMTSD'] == smtsd, 'SAMPID'])
  303. print(f" {key}: {len(region_samples[key])} samples")
  304. subject_pheno = pd.read_csv(subject_file, sep='\t')
  305. ba9_samples_list = region_samples[PRIMARY_REGION]
  306. ba9_subjects = [sample_to_subject(s) for s in ba9_samples_list]
  307. ba9_subject_df = pd.DataFrame({'SAMPID': ba9_samples_list, 'SUBJID': ba9_subjects})
  308. ba9_subject_df = ba9_subject_df.merge(subject_pheno, on='SUBJID', how='left')
  309. # ==================================================================
  310. # STEP 3: LOAD GTEx TPM
  311. # ==================================================================
  312. print(f"\n{'='*70}\nSTEP 3: LOAD GTEx TPM\n{'='*70}\n")
  313. print(" Reading header...")
  314. with gzip.open(tpm_file, 'rt') as f:
  315. f.readline(); f.readline()
  316. header = f.readline().strip().split('\t')
  317. all_brain_samples = set()
  318. for key in BRAIN_REGIONS:
  319. all_brain_samples.update(region_samples[key])
  320. sample_cols = {i: col for i, col in enumerate(header) if col in all_brain_samples}
  321. keep_cols = [0, 1] + sorted(sample_cols.keys())
  322. print(f" Brain samples: {len(sample_cols)}, columns to load: {len(keep_cols)}")
  323. print(" Loading expression data...")
  324. chunks = []
  325. with gzip.open(tpm_file, 'rt') as f:
  326. f.readline(); f.readline(); f.readline()
  327. row_buffer = []
  328. for line_num, line in enumerate(f):
  329. parts = line.strip().split('\t')
  330. row_buffer.append([parts[i] if i < len(parts) else '' for i in keep_cols])
  331. if len(row_buffer) >= 5000:
  332. chunks.append(pd.DataFrame(row_buffer))
  333. row_buffer = []
  334. if line_num % 10000 == 0:
  335. print(f" {line_num:,} genes...")
  336. if row_buffer:
  337. chunks.append(pd.DataFrame(row_buffer))
  338. tpm_raw = pd.concat(chunks, ignore_index=True)
  339. del chunks, row_buffer; gc.collect()
  340. col_names = ['Name', 'Description'] + [sample_cols[i] for i in sorted(sample_cols.keys())]
  341. tpm_raw.columns = col_names
  342. tpm_raw = tpm_raw.set_index('Name')
  343. gene_descriptions = tpm_raw['Description'].copy()
  344. tpm_raw = tpm_raw.drop('Description', axis=1)
  345. tpm_raw = tpm_raw.apply(pd.to_numeric, errors='coerce').astype(np.float32)
  346. print(f" Loaded: {tpm_raw.shape[0]} genes x {tpm_raw.shape[1]} samples")
  347. # ==================================================================
  348. # STEP 4: PREPARE REGION MATRICES
  349. # ==================================================================
  350. print(f"\n{'='*70}\nSTEP 4: PREPARE REGION MATRICES\n{'='*70}\n")
  351. def prepare_region(region_key):
  352. samples = [s for s in region_samples[region_key] if s in tpm_raw.columns]
  353. expr = tpm_raw[samples].copy()
  354. expr['gene_symbol'] = gene_descriptions.reindex(expr.index)
  355. expr = expr.dropna(subset=['gene_symbol'])
  356. expr = expr[expr['gene_symbol'] != '']
  357. expr['median_tpm'] = expr[samples].median(axis=1)
  358. expr = expr.sort_values('median_tpm', ascending=False)
  359. expr = expr.drop_duplicates(subset='gene_symbol', keep='first')
  360. expr = expr[expr['median_tpm'] >= MIN_MEDIAN_TPM]
  361. expr = expr.set_index('gene_symbol').drop('median_tpm', axis=1)
  362. return np.log2(expr + 1), len(samples), len(expr), list(expr.index)
  363. region_data = {}
  364. for key in BRAIN_REGIONS:
  365. log2_expr, n_samp, n_genes, genes = prepare_region(key)
  366. region_data[key] = {'expr': log2_expr, 'n_samples': n_samp,
  367. 'n_genes': n_genes, 'genes': genes}
  368. present = [g for g in TARGET_GENES if g in genes]
  369. print(f" {key}: n={n_samp}, genes={n_genes}, targets={len(present)}/7")
  370. del tpm_raw; gc.collect()
  371. # ==================================================================
  372. # STEP 5: GENOME-WIDE SPEARMAN CO-EXPRESSION
  373. # ==================================================================
  374. print(f"\n{'='*70}\nSTEP 5: GENOME-WIDE SPEARMAN CO-EXPRESSION -- {PRIMARY_REGION}\n{'='*70}\n")
  375. ba9 = region_data[PRIMARY_REGION]['expr']
  376. n_genes_ba9 = region_data[PRIMARY_REGION]['n_genes']
  377. n_samples_ba9 = region_data[PRIMARY_REGION]['n_samples']
  378. print(f"Matrix: {n_genes_ba9} genes x {n_samples_ba9} samples")
  379. print(f"Correlation method: Spearman (rank-based)\n")
  380. correlations = {}
  381. for target in TARGET_GENES:
  382. if target not in ba9.index:
  383. print(f" WARNING: {target} not found")
  384. continue
  385. target_expr = ba9.loc[target].values.astype(np.float64)
  386. other_genes = [g for g in ba9.index if g != target]
  387. other_expr = ba9.loc[other_genes].values.astype(np.float64)
  388. r_values = vectorized_spearman(target_expr, other_expr)
  389. corr_series = pd.Series(r_values, index=other_genes, name=target)
  390. correlations[target] = corr_series.dropna().sort_values(ascending=False)
  391. print(f" {target}: {len(corr_series.dropna())} genes, "
  392. f"top = {corr_series.dropna().idxmax()} "
  393. f"(rho = {corr_series.dropna().max():.3f})")
  394. # ==================================================================
  395. # STEP 6: PRIMARY ANALYSIS — WEIGHTED JACCARD ON CONTINUOUS VECTORS
  396. # ==================================================================
  397. print(f"\n{'='*70}")
  398. print("STEP 6: PRIMARY ANALYSIS -- WEIGHTED JACCARD (CONTINUOUS)")
  399. print(f"{'='*70}\n")
  400. print(f"Permutations: {N_PERMUTATIONS}, seed: {RANDOM_SEED}\n")
  401. # Align all correlation vectors to common gene set
  402. common_genes = sorted(set.intersection(*[set(correlations[g].index)
  403. for g in TARGET_GENES
  404. if g in correlations]))
  405. n_common = len(common_genes)
  406. print(f"Common genes across all 7 targets: {n_common}\n")
  407. # Build aligned matrix
  408. corr_matrix = pd.DataFrame({g: correlations[g].reindex(common_genes)
  409. for g in TARGET_GENES if g in correlations})
  410. # Compute WJ for all 21 pairwise comparisons
  411. wj_results = []
  412. for g1, g2 in combinations(TARGET_GENES, 2):
  413. if g1 not in corr_matrix.columns or g2 not in corr_matrix.columns:
  414. continue
  415. vec_a = corr_matrix[g1].values
  416. vec_b = corr_matrix[g2].values
  417. wj_obs, wj_pval, null_dist = weighted_jaccard_permutation_test(
  418. vec_a, vec_b, n_perm=N_PERMUTATIONS, seed=RANDOM_SEED)
  419. # Also compute Spearman rank correlation between vectors
  420. rho, rho_p = spearmanr(vec_a, vec_b)
  421. wj_results.append({
  422. 'gene1': g1, 'gene2': g2,
  423. 'wj': wj_obs,
  424. 'wj_perm_p': wj_pval,
  425. 'null_mean': np.mean(null_dist),
  426. 'null_std': np.std(null_dist),
  427. 'z_score': (wj_obs - np.mean(null_dist)) / np.std(null_dist) if np.std(null_dist) > 0 else 0,
  428. 'spearman_rho': rho,
  429. 'spearman_p': rho_p,
  430. 'n_genes': n_common,
  431. })
  432. wj_df = pd.DataFrame(wj_results).sort_values('wj', ascending=True)
  433. # FDR correction on WJ permutation p-values
  434. reject_wj, fdr_wj, _, _ = multipletests(wj_df['wj_perm_p'].values, method='fdr_bh')
  435. wj_df['wj_perm_p_fdr'] = fdr_wj
  436. wj_df['fdr_significant'] = reject_wj
  437. wj_df.to_csv(os.path.join(RESULTS_DIR, "wj_continuous_all_21_pairs.csv"), index=False)
  438. # Print all results
  439. print(f"{'Gene1':>10s} {'Gene2':>10s} {'WJ':>8s} {'z':>8s} {'p_perm':>10s} "
  440. f"{'FDR':>10s} {'rho':>8s}")
  441. print("-" * 70)
  442. for _, row in wj_df.iterrows():
  443. print(f"{row['gene1']:>10s} {row['gene2']:>10s} {row['wj']:>8.4f} "
  444. f"{row['z_score']:>8.2f} {row['wj_perm_p']:>10.4f} "
  445. f"{row['wj_perm_p_fdr']:>10.4f} {row['spearman_rho']:>8.3f}")
  446. # Key pair: SIGMAR1-TMEM97
  447. st_wj = wj_df[
  448. ((wj_df['gene1'] == 'SIGMAR1') & (wj_df['gene2'] == 'TMEM97')) |
  449. ((wj_df['gene1'] == 'TMEM97') & (wj_df['gene2'] == 'SIGMAR1'))
  450. ].iloc[0]
  451. print(f"\n{'='*60}")
  452. print("PRIMARY RESULT: SIGMAR1-TMEM97 WEIGHTED JACCARD")
  453. print(f" WJ = {st_wj['wj']:.6f}")
  454. print(f" z-score = {st_wj['z_score']:.2f}")
  455. print(f" Permutation p = {st_wj['wj_perm_p']:.4f}")
  456. print(f" FDR p = {st_wj['wj_perm_p_fdr']:.4f}")
  457. print(f" Null mean = {st_wj['null_mean']:.6f}, std = {st_wj['null_std']:.6f}")
  458. print(f" Spearman rho = {st_wj['spearman_rho']:.4f}")
  459. print(f"{'='*60}")
  460. # SIGMAR1-LTN1
  461. sl_wj = wj_df[
  462. ((wj_df['gene1'] == 'SIGMAR1') & (wj_df['gene2'] == 'LTN1')) |
  463. ((wj_df['gene1'] == 'LTN1') & (wj_df['gene2'] == 'SIGMAR1'))
  464. ].iloc[0]
  465. print(f"\nSIGMAR1-LTN1: WJ={sl_wj['wj']:.4f}, z={sl_wj['z_score']:.2f}, "
  466. f"p={sl_wj['wj_perm_p']:.4f}")
  467. # ==================================================================
  468. # STEP 7: SUPPLEMENTARY — BINARY JACCARD ON TOP 5% SETS
  469. # ==================================================================
  470. print(f"\n{'='*70}")
  471. print("STEP 7: SUPPLEMENTARY -- BINARY JACCARD (TOP 5% SETS)")
  472. print(f"{'='*70}\n")
  473. networks = {}
  474. thresholds = {}
  475. for target, corr in correlations.items():
  476. n_total = len(corr)
  477. n_top = int(np.ceil(n_total * TOP_PERCENT / 100))
  478. networks[target] = set(corr.head(n_top).index)
  479. thresholds[target] = corr.iloc[n_top - 1]
  480. print(f" {target}: top 5% = {n_top} genes, rho >= {thresholds[target]:.3f}")
  481. gene_universe = len(correlations[TARGET_GENES[0]])
  482. binary_results = []
  483. for g1, g2 in combinations(TARGET_GENES, 2):
  484. if g1 not in networks or g2 not in networks:
  485. continue
  486. set1, set2 = networks[g1], networks[g2]
  487. shared = set1 & set2
  488. union_set = set1 | set2
  489. jaccard = len(shared) / len(union_set) if len(union_set) > 0 else 0
  490. a, b, c = len(shared), len(set1 - set2), len(set2 - set1)
  491. d = gene_universe - len(union_set)
  492. fisher_or, fisher_p = fisher_exact(np.array([[a, b], [c, d]]),
  493. alternative='greater')
  494. binary_results.append({
  495. 'gene1': g1, 'gene2': g2,
  496. 'shared': len(shared), 'binary_jaccard': jaccard,
  497. 'fisher_or': fisher_or, 'fisher_p': fisher_p,
  498. 'set1_size': len(set1), 'set2_size': len(set2),
  499. })
  500. binary_df = pd.DataFrame(binary_results).sort_values('binary_jaccard', ascending=False)
  501. reject_b, fdr_b, _, _ = multipletests(binary_df['fisher_p'].values, method='fdr_bh')
  502. binary_df['fisher_p_fdr'] = fdr_b
  503. binary_df.to_csv(os.path.join(RESULTS_DIR, "binary_jaccard_top5pct.csv"), index=False)
  504. # Merge WJ and binary results for comparison
  505. comparison = wj_df[['gene1', 'gene2', 'wj', 'z_score', 'wj_perm_p']].merge(
  506. binary_df[['gene1', 'gene2', 'binary_jaccard', 'shared', 'fisher_p']],
  507. on=['gene1', 'gene2'], how='outer'
  508. )
  509. comparison.to_csv(os.path.join(RESULTS_DIR, "wj_vs_binary_comparison.csv"), index=False)
  510. print("\n--- WJ vs Binary Jaccard comparison ---")
  511. print(f"{'Pair':>25s} {'WJ':>8s} {'Binary J':>10s} {'Shared':>7s}")
  512. print("-" * 55)
  513. for _, row in comparison.sort_values('wj').iterrows():
  514. pair = f"{row['gene1']}-{row['gene2']}"
  515. bj = f"{row['binary_jaccard']:.3f}" if pd.notna(row.get('binary_jaccard')) else "N/A"
  516. sh = f"{int(row['shared'])}" if pd.notna(row.get('shared')) else "N/A"
  517. print(f"{pair:>25s} {row['wj']:>8.4f} {bj:>10s} {sh:>7s}")
  518. # Key validation
  519. st_binary = binary_df[
  520. ((binary_df['gene1'] == 'SIGMAR1') & (binary_df['gene2'] == 'TMEM97')) |
  521. ((binary_df['gene1'] == 'TMEM97') & (binary_df['gene2'] == 'SIGMAR1'))
  522. ].iloc[0]
  523. print(f"\n SIGMAR1-TMEM97 binary J = {st_binary['binary_jaccard']:.3f}, "
  524. f"shared = {int(st_binary['shared'])}")
  525. # Gene sets
  526. sigmar1_unique = sorted(networks['SIGMAR1'] - networks['TMEM97'])
  527. tmem97_unique = sorted(networks['TMEM97'] - networks['SIGMAR1'])
  528. shared_st = sorted(networks['SIGMAR1'] & networks['TMEM97'])
  529. # ==================================================================
  530. # STEP 8: TOP CO-EXPRESSION PARTNERS
  531. # ==================================================================
  532. print(f"\n{'='*70}\nSTEP 8: TOP CO-EXPRESSION PARTNERS\n{'='*70}\n")
  533. for target in PRIMARY_TARGETS:
  534. print(f" {target} top 10:")
  535. for rank, (gene, r_val) in enumerate(correlations[target].head(10).items(), 1):
  536. print(f" {rank:2d}. {gene:12s} rho = {r_val:.4f}")
  537. print()
  538. # ==================================================================
  539. # STEP 9: CUSTOM GENE SET ENRICHMENT
  540. # ==================================================================
  541. print(f"\n{'='*70}\nSTEP 9: CUSTOM GENE SET ENRICHMENT\n{'='*70}\n")
  542. custom_sets = OrderedDict([
  543. ('MAM-mitochondrial', MAM_MITO_GENES),
  544. ('Sigma receptor network', SIGMA_NETWORK_GENES),
  545. ('ER stress/UPR', ER_STRESS_UPR_GENES),
  546. ('Methylation pathway', METHYLATION_GENES),
  547. ('Vascular markers (neg ctrl)', VASCULAR_GENES),
  548. ('Ribosome quality control', RQC_GENES),
  549. ])
  550. all_custom_results = {}
  551. for target in PRIMARY_TARGETS:
  552. print(f"\n--- {target} top 5% custom enrichment ---\n")
  553. all_genes_t = list(correlations[target].index) + [target]
  554. target_results = []
  555. for set_name, gene_list in custom_sets.items():
  556. result = compute_custom_enrichment(
  557. networks[target], gene_list, set_name, all_genes_t)
  558. if result:
  559. result['target'] = target
  560. target_results.append(result)
  561. all_custom_results[target] = target_results
  562. pd.DataFrame(target_results).to_csv(
  563. os.path.join(RESULTS_DIR, f"{target}_custom_enrichment.csv"), index=False)
  564. # ==================================================================
  565. # STEP 10: gProfiler ENRICHMENT
  566. # ==================================================================
  567. print(f"\n{'='*70}\nSTEP 10: gProfiler ENRICHMENT\n{'='*70}\n")
  568. try:
  569. from gprofiler import GProfiler
  570. gp = GProfiler(return_dataframe=True)
  571. except ImportError:
  572. print("WARNING: gprofiler not installed")
  573. gp = None
  574. def run_gprofiler(gene_list, bg, label=""):
  575. if gp is None:
  576. return pd.DataFrame()
  577. try:
  578. df = gp.profile(organism='hsapiens', query=list(gene_list),
  579. background=list(bg),
  580. sources=['GO:BP', 'GO:MF', 'GO:CC', 'KEGG', 'REAC'],
  581. significance_threshold_method='g_SCS',
  582. user_threshold=0.05, no_evidences=False)
  583. if df is not None and len(df) > 0:
  584. print(f" {label}: {len(df)} terms")
  585. return df
  586. print(f" {label}: 0 terms")
  587. return pd.DataFrame()
  588. except Exception as e:
  589. print(f" {label}: error -- {e}")
  590. return pd.DataFrame()
  591. background = list(correlations['SIGMAR1'].index) + ['SIGMAR1']
  592. all_gprofiler = {}
  593. for name, genes in [
  594. ('SIGMAR1_full', list(networks['SIGMAR1'])),
  595. ('TMEM97_full', list(networks['TMEM97'])),
  596. ('SIGMAR1_unique', sigmar1_unique),
  597. ('TMEM97_unique', tmem97_unique),
  598. ('shared_SIGMAR1_TMEM97', shared_st),
  599. ]:
  600. df = run_gprofiler(genes, background, name)
  601. all_gprofiler[name] = df
  602. if len(df) > 0:
  603. df.to_csv(os.path.join(RESULTS_DIR, f"gProfiler_{name}.csv"), index=False)
  604. # ==================================================================
  605. # STEP 11: CELL-TYPE DECONVOLUTION
  606. # ==================================================================
  607. print(f"\n{'='*70}\nSTEP 11: CELL-TYPE DECONVOLUTION SENSITIVITY\n{'='*70}\n")
  608. celltype_proportions = pd.DataFrame(index=ba9.columns)
  609. for ct_name, markers in CELLTYPE_MARKERS.items():
  610. present = [g for g in markers if g in ba9.index]
  611. if present:
  612. celltype_proportions[ct_name] = ba9.loc[present].mean(axis=0)
  613. print(f" {ct_name}: {len(present)}/{len(markers)} markers")
  614. for target in PRIMARY_TARGETS:
  615. corr_ct, thr_ct, net_ct, n_valid = partial_correlation_network(
  616. target, ba9, celltype_proportions)
  617. if corr_ct is None:
  618. continue
  619. common = sorted(set(corr_ct.index) & set(correlations[target].index))
  620. rho_preserve, _ = spearmanr(corr_ct.reindex(common).values,
  621. correlations[target].reindex(common).values)
  622. print(f" {target}: rank preservation rho = {rho_preserve:.4f}")
  623. if target == 'SIGMAR1' and net_ct:
  624. print(f" VCP in adjusted top 5%: {'VCP' in net_ct}")
  625. # ==================================================================
  626. # STEP 12: COVARIATE ADJUSTMENT
  627. # ==================================================================
  628. print(f"\n{'='*70}\nSTEP 12: COVARIATE ADJUSTMENT (AGE + SEX)\n{'='*70}\n")
  629. ba9_covar = ba9_subject_df.set_index('SAMPID').reindex(ba9.columns)
  630. age_map = {'20-29': 25, '30-39': 35, '40-49': 45, '50-59': 55,
  631. '60-69': 65, '70-79': 75}
  632. ba9_covar['AGE_MID'] = ba9_covar['AGE'].map(age_map) if 'AGE' in ba9_covar.columns else np.nan
  633. ba9_covar['SEX_NUM'] = ba9_covar['SEX'].astype(float) if 'SEX' in ba9_covar.columns else np.nan
  634. covar_df = ba9_covar[['AGE_MID', 'SEX_NUM']]
  635. for target in PRIMARY_TARGETS:
  636. corr_as, thr_as, _, n_valid = partial_correlation_network(target, ba9, covar_df)
  637. if corr_as is None:
  638. continue
  639. common = sorted(set(corr_as.index) & set(correlations[target].index))
  640. rho_p, _ = spearmanr(corr_as.reindex(common).values,
  641. correlations[target].reindex(common).values)
  642. print(f" {target}: rank preservation rho = {rho_p:.4f}")
  643. # ==================================================================
  644. # STEP 13: MULTI-REGION REPLICATION (WJ + binary)
  645. # ==================================================================
  646. print(f"\n{'='*70}\nSTEP 13: MULTI-REGION REPLICATION\n{'='*70}\n")
  647. region_correlations = {}
  648. for region_key in BRAIN_REGIONS:
  649. expr = region_data[region_key]['expr']
  650. region_correlations[region_key] = {}
  651. for target in PRIMARY_TARGETS:
  652. if target not in expr.index:
  653. continue
  654. target_vals = expr.loc[target].values.astype(np.float64)
  655. other_genes = [g for g in expr.index if g != target]
  656. other_vals = expr.loc[other_genes].values.astype(np.float64)
  657. r_vals = vectorized_spearman(target_vals, other_vals)
  658. region_correlations[region_key][target] = pd.Series(
  659. r_vals, index=other_genes).dropna().sort_values(ascending=False)
  660. print(f" {region_key} (n={region_data[region_key]['n_samples']}): done")
  661. # Cross-region rank correlations
  662. region_keys = list(BRAIN_REGIONS.keys())
  663. cross_region_matrix = {}
  664. for target in PRIMARY_TARGETS:
  665. mat = np.zeros((5, 5))
  666. for i, r1 in enumerate(region_keys):
  667. for j, r2 in enumerate(region_keys):
  668. if i == j:
  669. mat[i, j] = 1.0
  670. continue
  671. if target in region_correlations.get(r1, {}) and target in region_correlations.get(r2, {}):
  672. c1, c2 = region_correlations[r1][target], region_correlations[r2][target]
  673. common = sorted(set(c1.index) & set(c2.index))
  674. if len(common) > 100:
  675. mat[i, j], _ = spearmanr(c1.reindex(common).values, c2.reindex(common).values)
  676. cross_region_matrix[target] = mat
  677. vals = mat[np.triu_indices(5, k=1)]
  678. print(f"\n {target} cross-region rho: {vals.min():.3f}-{vals.max():.3f}")
  679. # WJ across regions
  680. print("\n--- SIGMAR1-TMEM97 WJ across regions ---")
  681. region_wj_results = {}
  682. for region_key in BRAIN_REGIONS:
  683. if 'SIGMAR1' not in region_correlations.get(region_key, {}) or \
  684. 'TMEM97' not in region_correlations.get(region_key, {}):
  685. continue
  686. s_corr = region_correlations[region_key]['SIGMAR1']
  687. t_corr = region_correlations[region_key]['TMEM97']
  688. common_r = sorted(set(s_corr.index) & set(t_corr.index))
  689. wj_r = weighted_jaccard(s_corr.reindex(common_r).values,
  690. t_corr.reindex(common_r).values)
  691. # Binary Jaccard too
  692. n_s = int(np.ceil(len(s_corr) * TOP_PERCENT / 100))
  693. n_t = int(np.ceil(len(t_corr) * TOP_PERCENT / 100))
  694. s_set, t_set = set(s_corr.head(n_s).index), set(t_corr.head(n_t).index)
  695. shared_r = s_set & t_set
  696. bj_r = len(shared_r) / len(s_set | t_set) if len(s_set | t_set) > 0 else 0
  697. region_wj_results[region_key] = {'wj': wj_r, 'binary_j': bj_r,
  698. 'shared': len(shared_r)}
  699. print(f" {region_key}: WJ={wj_r:.4f}, binary J={bj_r:.3f}, shared={len(shared_r)}")
  700. # ==================================================================
  701. # STEP 14: EXPORT GENE LISTS
  702. # ==================================================================
  703. print(f"\n{'='*70}\nSTEP 14: EXPORT GENE LISTS\n{'='*70}\n")
  704. for target in PRIMARY_TARGETS:
  705. top5 = correlations[target].head(len(networks[target]))
  706. pd.DataFrame({'gene': top5.index, f'rho_with_{target}': top5.values}).to_csv(
  707. os.path.join(RESULTS_DIR, f"{target}_top5pct.csv"), index=False)
  708. pd.DataFrame({'gene': shared_st}).to_csv(
  709. os.path.join(RESULTS_DIR, "SIGMAR1_TMEM97_shared.csv"), index=False)
  710. pd.DataFrame({'gene': sigmar1_unique}).to_csv(
  711. os.path.join(RESULTS_DIR, "SIGMAR1_unique.csv"), index=False)
  712. pd.DataFrame({'gene': tmem97_unique}).to_csv(
  713. os.path.join(RESULTS_DIR, "TMEM97_unique.csv"), index=False)
  714. for target in TARGET_GENES:
  715. if target in correlations:
  716. corr = correlations[target]
  717. pd.DataFrame({'gene': corr.index, f'rho_with_{target}': corr.values,
  718. 'rank': range(1, len(corr) + 1)}).to_csv(
  719. os.path.join(RESULTS_DIR, f"{target}_genome_wide_rankings.csv"), index=False)
  720. print(f" Exported: {len(TARGET_GENES)} ranking files, 3 gene set files")
  721. # ==================================================================
  722. # STEP 15: FIGURES
  723. # ==================================================================
  724. print(f"\n{'='*70}\nSTEP 15: FIGURES\n{'='*70}\n")
  725. # Figure 1: Divergent Networks (Venn + scatter + top partners)
  726. fig = plt.figure(figsize=(18, 6))
  727. ax_a = fig.add_axes([0.02, 0.12, 0.28, 0.80])
  728. ax_a.text(-0.05, 1.08, 'A', fontsize=20, fontweight='bold', va='top',
  729. transform=ax_a.transAxes)
  730. v = venn2(subsets=(len(sigmar1_unique), len(tmem97_unique), len(shared_st)),
  731. set_labels=('SIGMAR1', 'TMEM97'), ax=ax_a)
  732. for pid, color in [('10', '#1565C0'), ('01', '#E65100'), ('11', '#7B1FA2')]:
  733. v.get_patch_by_id(pid).set_color(color)
  734. v.get_patch_by_id(pid).set_alpha(0.7)
  735. j_val = len(shared_st) / (len(sigmar1_unique) + len(tmem97_unique) + len(shared_st))
  736. ax_a.set_title(f'Binary Jaccard = {j_val:.3f}\nWJ = {st_wj["wj"]:.4f}',
  737. fontsize=12, fontweight='bold')
  738. ax_b = fig.add_axes([0.36, 0.12, 0.28, 0.80])
  739. ax_b.text(-0.05, 1.08, 'B', fontsize=20, fontweight='bold', va='top',
  740. transform=ax_b.transAxes)
  741. s_vals = corr_matrix['SIGMAR1'].values
  742. t_vals = corr_matrix['TMEM97'].values
  743. ax_b.scatter(s_vals, t_vals, s=1, alpha=0.15, c='#555', rasterized=True)
  744. ax_b.set_xlabel('SIGMAR1 Spearman rho', fontsize=12)
  745. ax_b.set_ylabel('TMEM97 Spearman rho', fontsize=12)
  746. ax_b.set_title(f'Genome-wide correlation vectors\nWJ = {st_wj["wj"]:.4f}',
  747. fontsize=12, fontweight='bold')
  748. ax_b.plot([-0.5, 1], [-0.5, 1], 'k--', alpha=0.3, lw=1)
  749. ax_c = fig.add_axes([0.70, 0.12, 0.28, 0.80])
  750. ax_c.text(-0.05, 1.08, 'C', fontsize=20, fontweight='bold', va='top',
  751. transform=ax_c.transAxes)
  752. top10_s = correlations['SIGMAR1'].head(10)
  753. top10_t = correlations['TMEM97'].head(10)
  754. y_pos = np.arange(10)
  755. bh = 0.35
  756. ax_c.barh(y_pos + bh/2, top10_s.values, bh, label='SIGMAR1', color='#1565C0', alpha=0.8)
  757. ax_c.barh(y_pos - bh/2, top10_t.values, bh, label='TMEM97', color='#E65100', alpha=0.8)
  758. ax_c.set_yticks(y_pos)
  759. ax_c.set_yticklabels([f"{top10_s.index[i]} | {top10_t.index[i]}" for i in range(10)], fontsize=8)
  760. ax_c.set_xlabel('Spearman rho', fontsize=12)
  761. ax_c.set_title('Top 10 co-expression partners', fontsize=12, fontweight='bold')
  762. ax_c.legend(fontsize=10)
  763. ax_c.invert_yaxis()
  764. plt.savefig(os.path.join(FIGURES_DIR, "Figure1_Divergent_Networks.png"),
  765. dpi=300, bbox_inches='tight')
  766. plt.savefig(os.path.join(FIGURES_DIR, "Figure1_Divergent_Networks.pdf"),
  767. bbox_inches='tight')
  768. plt.close()
  769. print(" Figure 1 saved")
  770. # Figure 2: GO Enrichment
  771. fig, axes = plt.subplots(2, 3, figsize=(20, 12))
  772. panel_data = [
  773. (all_gprofiler.get('SIGMAR1_unique', pd.DataFrame()), 'GO:BP', 'SIGMAR1-unique GO:BP'),
  774. (all_gprofiler.get('TMEM97_unique', pd.DataFrame()), 'GO:BP', 'TMEM97-unique GO:BP'),
  775. (all_gprofiler.get('shared_SIGMAR1_TMEM97', pd.DataFrame()), 'GO:BP', 'Shared GO:BP'),
  776. (all_gprofiler.get('SIGMAR1_full', pd.DataFrame()), 'REAC', 'SIGMAR1 Reactome'),
  777. (all_gprofiler.get('TMEM97_full', pd.DataFrame()), 'REAC', 'TMEM97 Reactome'),
  778. (all_gprofiler.get('SIGMAR1_full', pd.DataFrame()), 'KEGG', 'SIGMAR1 KEGG'),
  779. ]
  780. for idx, (go_df, source, title) in enumerate(panel_data):
  781. ax = axes.flat[idx]
  782. ax.text(-0.08, 1.05, chr(65+idx), fontsize=16, fontweight='bold',
  783. va='top', transform=ax.transAxes)
  784. if go_df is not None and len(go_df) > 0:
  785. subset = go_df[go_df['source'] == source].sort_values('p_value').head(8)
  786. if len(subset) > 0:
  787. names = [n[:45] for n in subset['name'].values]
  788. pvals = [-np.log10(p) for p in subset['p_value'].values]
  789. color = '#1565C0' if 'SIGMAR1' in title else '#E65100' if 'TMEM97' in title else '#7B1FA2'
  790. ax.barh(range(len(names)), pvals, color=color, alpha=0.8)
  791. ax.set_yticks(range(len(names)))
  792. ax.set_yticklabels(names, fontsize=8)
  793. ax.set_xlabel('-log10(p)', fontsize=10)
  794. ax.invert_yaxis()
  795. else:
  796. ax.text(0.5, 0.5, 'No terms', ha='center', va='center',
  797. transform=ax.transAxes, color='gray')
  798. else:
  799. ax.text(0.5, 0.5, 'No data', ha='center', va='center',
  800. transform=ax.transAxes, color='gray')
  801. ax.set_title(title, fontsize=11, fontweight='bold')
  802. plt.tight_layout()
  803. plt.savefig(os.path.join(FIGURES_DIR, "Figure2_GO_Enrichment.png"), dpi=300, bbox_inches='tight')
  804. plt.savefig(os.path.join(FIGURES_DIR, "Figure2_GO_Enrichment.pdf"), bbox_inches='tight')
  805. plt.close()
  806. print(" Figure 2 saved")
  807. # Figure 3: Multi-region replication
  808. fig, axes = plt.subplots(1, 2, figsize=(14, 6))
  809. for idx, target in enumerate(PRIMARY_TARGETS):
  810. ax = axes[idx]
  811. ax.text(-0.08, 1.05, chr(65+idx), fontsize=16, fontweight='bold',
  812. va='top', transform=ax.transAxes)
  813. mat = cross_region_matrix[target]
  814. im = ax.imshow(mat, cmap='viridis', vmin=0.8, vmax=1.0)
  815. for i in range(5):
  816. for j in range(5):
  817. ax.text(j, i, f'{mat[i,j]:.3f}', ha='center', va='center',
  818. fontsize=9, color='white' if mat[i,j] < 0.92 else 'black')
  819. ax.set_xticks(range(5)); ax.set_xticklabels(region_keys, fontsize=9, rotation=45, ha='right')
  820. ax.set_yticks(range(5)); ax.set_yticklabels(region_keys, fontsize=9)
  821. ax.set_title(f'{target} cross-region', fontsize=12, fontweight='bold')
  822. plt.colorbar(im, ax=ax, shrink=0.8, label='Spearman rho')
  823. plt.tight_layout()
  824. plt.savefig(os.path.join(FIGURES_DIR, "Figure3_MultiRegion_Replication.png"), dpi=300, bbox_inches='tight')
  825. plt.savefig(os.path.join(FIGURES_DIR, "Figure3_MultiRegion_Replication.pdf"), bbox_inches='tight')
  826. plt.close()
  827. print(" Figure 3 saved")
  828. # ==================================================================
  829. # STEP 16: SUPPLEMENTARY TABLES
  830. # ==================================================================
  831. print(f"\n{'='*70}\nSTEP 16: SUPPLEMENTARY TABLES\n{'='*70}\n")
  832. import openpyxl
  833. wb = openpyxl.Workbook()
  834. for i, target in enumerate(TARGET_GENES):
  835. if target in correlations:
  836. ws = wb.active if i == 0 else wb.create_sheet()
  837. ws.title = target
  838. ws.append(['Gene', f'rho_with_{target}', 'Rank'])
  839. for rank, (gene, r_val) in enumerate(correlations[target].items(), 1):
  840. ws.append([gene, round(r_val, 6), rank])
  841. wb.save(os.path.join(SUPPL_DIR, "Table_S1_Genome_Wide_Rankings.xlsx"))
  842. wb2 = openpyxl.Workbook()
  843. first = True
  844. for name, df in all_gprofiler.items():
  845. if len(df) > 0:
  846. ws = wb2.active if first else wb2.create_sheet()
  847. ws.title = name[:31]
  848. cols = [c for c in df.columns if c not in ['query', 'parents']]
  849. ws.append(cols)
  850. for _, row in df[cols].iterrows():
  851. ws.append([str(v) for v in row.values])
  852. first = False
  853. wb2.save(os.path.join(SUPPL_DIR, "Table_S2_gProfiler_Enrichment.xlsx"))
  854. all_custom_rows = []
  855. for target, results in all_custom_results.items():
  856. all_custom_rows.extend(results)
  857. pd.DataFrame(all_custom_rows).to_excel(
  858. os.path.join(SUPPL_DIR, "Table_S3_Custom_Gene_Set_Enrichment.xlsx"), index=False)
  859. shared_detail = [{'gene': g,
  860. 'rho_with_SIGMAR1': round(correlations['SIGMAR1'].get(g, np.nan), 6),
  861. 'rho_with_TMEM97': round(correlations['TMEM97'].get(g, np.nan), 6)}
  862. for g in shared_st]
  863. pd.DataFrame(shared_detail).to_excel(
  864. os.path.join(SUPPL_DIR, "Table_S4_Shared_Genes.xlsx"), index=False)
  865. print(" Tables S1-S4 saved")
  866. # ==================================================================
  867. # STEP 17: PROVENANCE
  868. # ==================================================================
  869. print(f"\n{'='*70}\nSTEP 17: PROVENANCE\n{'='*70}\n")
  870. from datetime import datetime
  871. provenance = {
  872. "methodology": "WJ-native",
  873. "fundamental_unit": f"individual gene (GTEx v8 RNA-seq, {n_genes_ba9} expressed genes in BA9)",
  874. "pairwise_matrix": "genome-wide Spearman correlation, each target vs all genes",
  875. "correlation_method": "Spearman",
  876. "primary_analysis": "Weighted Jaccard on continuous correlation vectors",
  877. "supplementary_analysis": "Binary Jaccard on top 5% network membership",
  878. "fdr_scope": f"all 21 pairwise WJ permutation p-values (Benjamini-Hochberg)",
  879. "permutations": N_PERMUTATIONS,
  880. "domain_conventional_methods": "gProfiler GO enrichment (comparison), Fisher exact (custom sets)",
  881. "random_seed": RANDOM_SEED,
  882. "pipeline_file": "MS3_sigma_divergence_wj_pipeline.py",
  883. "execution_date": datetime.now().strftime("%Y-%m-%d"),
  884. "execution_time_seconds": round(time.time() - t_start, 1),
  885. "wj_compliance_status": "PASS",
  886. "brain_regions": list(BRAIN_REGIONS.keys()),
  887. "n_samples_primary": n_samples_ba9,
  888. "n_genes_primary": n_genes_ba9,
  889. "key_results": {
  890. "SIGMAR1_TMEM97_wj": float(st_wj['wj']),
  891. "SIGMAR1_TMEM97_wj_z": float(st_wj['z_score']),
  892. "SIGMAR1_TMEM97_wj_perm_p": float(st_wj['wj_perm_p']),
  893. "SIGMAR1_TMEM97_binary_jaccard": float(st_binary['binary_jaccard']),
  894. "SIGMAR1_TMEM97_shared_genes": int(st_binary['shared']),
  895. "SIGMAR1_LTN1_wj": float(sl_wj['wj']),
  896. "cross_region_rho_SIGMAR1": f"{cross_region_matrix['SIGMAR1'][np.triu_indices(5, k=1)].min():.3f}-{cross_region_matrix['SIGMAR1'][np.triu_indices(5, k=1)].max():.3f}",
  897. "region_wj_values": {k: round(v['wj'], 4) for k, v in region_wj_results.items()},
  898. "region_binary_j_values": {k: round(v['binary_j'], 3) for k, v in region_wj_results.items()},
  899. },
  900. }
  901. provenance_path = os.path.join(RESULTS_DIR, "provenance.json")
  902. with open(provenance_path, 'w') as f:
  903. json.dump(provenance, f, indent=2)
  904. print(f" provenance.json saved")
  905. # ==================================================================
  906. # SUMMARY
  907. # ==================================================================
  908. elapsed = time.time() - t_start
  909. print(f"\n{'='*70}")
  910. print("PIPELINE COMPLETE")
  911. print(f"{'='*70}")
  912. print(f"\n Time: {elapsed:.0f}s ({elapsed/60:.1f} min)")
  913. print(f"\n--- PRIMARY: WEIGHTED JACCARD (continuous) ---")
  914. print(f" SIGMAR1-TMEM97 WJ = {st_wj['wj']:.6f} (z={st_wj['z_score']:.2f}, p={st_wj['wj_perm_p']:.4f})")
  915. print(f" SIGMAR1-LTN1 WJ = {sl_wj['wj']:.6f} (z={sl_wj['z_score']:.2f}, p={sl_wj['wj_perm_p']:.4f})")
  916. print(f"\n--- SUPPLEMENTARY: BINARY JACCARD (top 5%) ---")
  917. print(f" SIGMAR1-TMEM97 J = {st_binary['binary_jaccard']:.3f} (shared={int(st_binary['shared'])})")
  918. print(f"\n--- MULTI-REGION WJ ---")
  919. for rk, rv in region_wj_results.items():
  920. print(f" {rk}: WJ={rv['wj']:.4f}, binary J={rv['binary_j']:.3f}")
  921. return provenance
  922. if __name__ == '__main__':
  923. main()

MS3_sigma_divergence_wj_pipeline.py at commit bc9b836, under MIT · at the source

Overview

Authors: Drake H Harbert1
  1. Inner Architecture LLC, Canton, OH, United States
Journal: Frontiers in pharmacology, volume 17, article 1830847
Dates: received 14 March 2026; accepted 29 April 2026; published online 4 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fphar.2026.1830847 · PMID 42328650 · PMCID PMC13275387 · OpenAlex W7163580477
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), cellular / molecular (subfield)
Methods: Statistics, Connectivity
Keywords: co-expression architecture, mitochondria-associated membrane, neurodegeneration, sigma-1 receptor, sigma-2 receptor, subtype-selective pharmacology, TMEM97, weighted jaccard
Topic: Pharmacological Receptor Mechanisms and Effects (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: not cited yet (Europe PMC); 38 references in the paper

Abstract

The sigma-1 receptor (SIGMAR1) and sigma-2 receptor (TMEM97) are both enriched at the mitochondria-associated membrane (MAM) and have been pharmacologically co-classified for decades, yet their functional relationship at the transcriptomic level remains uncharacterized. We performed genome-wide co-expression analysis for both receptors across five brain regions from the GTEx v8 dataset (n = 209 in the primary region, 16,225 expressed genes) using Spearman correlations. Three Weighted Jaccard (WJ) formulations on continuous correlation vectors — (r+1)/2 shifted, unsigned |r|, and signed—all revealed that SIGMAR1 and TMEM97 share the majority of their global transcriptional architecture (WJ shifted = 0.964, unsigned = 0.907, signed = 0.906; all three rank-identical across 21 pairwise comparisons, ρ = 1.000), yet their top 5% co-expression networks overlap by only 10.0% (binary Jaccard = 0.100). Cosine similarity on raw vectors confirmed metric robustness (ρ = 0.856 with WJ, p = 7.5 × 10−7). Dissociation gap analysis across all 21 gene pairs showed the WJ-binary gap varies 2.1-fold (0.431–0.927), tracking known biological relatedness rather than reflecting a fixed methodological property. Gene Ontology analysis identified SIGMAR1-specific enrichment for mitochondrial translation and TCA cycle, and TMEM97-specific enrichment for ubiquitin-mediated proteolysis and neurodegeneration pathways. Multi-region replication confirmed the pattern across five brain regions, with the hippocampus showing tail-specific convergence. These findings are consistent with the hypothesis that dual sigma-1/sigma-2 ligands engage two functionally distinct co-expression programs within a shared cellular context, providing a transcriptomic rationale for subtype-selective pharmacological strategies.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

nwharbert8-ui/sigma-receptor-divergence

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 4acadaaec414c5191a251cd54b0ef52cfcacaf8d, 14 February 2026
Languages: Python (1)
Size: 5 files, 1 script
Software Heritage: not archived
Found in: “Software and reproducibility”
Holds: README, license file, CITATION.cff, environment (requirements.txt)
Not found: tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
3 files

Zenodo 19024710

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability statement”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), SciPy (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
3 files

nwharbert8-ui/sigma-receptor-divergence-wj

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bc9b836aae80e19aac860871508b205bb0957657, 14 March 2026
Languages: Python (1)
Size: 6 files, 1 script
Software Heritage: not archived
Found in: “Data availability statement”
Holds: README, license file, CITATION.cff, environment (requirements.txt)
Not found: tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), SciPy (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
3 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 3 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

Publicly available datasets were analyzed in this study. This data can be found here: GTEx v8 gene expression data (TPM) are available through the GTEx Portal (https://gtexportal.org/) under dbGaP accession phs000424.v8.p2. All analysis code is available at https://github.com/nwharbert8-ui/sigma-receptor-divergence-wj and archived at https://doi.org/10.5281/zenodo.19024710. Derived results (genome-wide correlation rankings, enrichment results, shared gene lists) are provided as Supplementary Data Sheets 1–4.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 1 author, 8 keywords, 38 references.

Cite

This paper

Harbert, D. H. (2026). Sigma-1 and Sigma-2 receptors exhibit divergent genome-wide Co-expression architectures in human brain despite shared subcellular localization. Frontiers in pharmacology, 17, 1830847. https://doi.org/10.3389/fphar.2026.1830847

BibTeX

@article{harbert2026sigma,
author = {Harbert, Drake H},
title = {{Sigma-1 and Sigma-2 receptors exhibit divergent genome-wide Co-expression architectures in human brain despite shared subcellular localization}},
journal = {Frontiers in pharmacology},
year = {2026},
month = jun,
volume = {17},
pages = {1830847},
publisher = {Frontiers Media SA},
issn = {1663-9812},
doi = {10.3389/fphar.2026.1830847},
url = {https://doi.org/10.3389/fphar.2026.1830847},
pmid = {42328650},
pmcid = {PMC13275387}
}

RIS

TY - JOUR
AU - Harbert, Drake H
TI - Sigma-1 and Sigma-2 receptors exhibit divergent genome-wide Co-expression architectures in human brain despite shared subcellular localization
T2 - Frontiers in pharmacology
J2 - Front Pharmacol
PY - 2026
DA - 2026/06/04
VL - 17
SP - 1830847
SN - 1663-9812
PB - Frontiers Media SA
DO - 10.3389/fphar.2026.1830847
UR - https://doi.org/10.3389/fphar.2026.1830847
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fphar.2026.1830847",
"type": "article-journal",
"title": "Sigma-1 and Sigma-2 receptors exhibit divergent genome-wide Co-expression architectures in human brain despite shared subcellular localization",
"container-title": "Frontiers in pharmacology",
"author": [
{
"family": "Harbert",
"given": "Drake H"
}
],
"container-title-short": "Front Pharmacol",
"volume": "17",
"page": "1830847",
"DOI": "10.3389/fphar.2026.1830847",
"PMID": "42328650",
"PMCID": "PMC13275387",
"ISSN": "1663-9812",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fphar.2026.1830847",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
4
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1007/s12035-026-06155-6
Sigma-1 Expression in Chronic Mouse Models of Temporal Lobe Epilepsy.
Journal: Molecular neurobiology
In common: cellular / molecular, 5 references
[2] doi:10.1016/j.isci.2026.116439 [code]
Decoding the role of transcriptomic clocks in the human prefrontal cortex.
Journal: iScience
In common: statsmodels, pandas, SciPy, 2 other tools, genetics / omics, 2 references
[3] doi:10.1038/s41380-026-03497-4 [code]
Transcriptome-informed brain cartography of polygenic risk and association with brain structure in major psychiatric disorders.
Journal: Molecular psychiatry
In common: statsmodels, pandas, SciPy, 2 other tools, genetics / omics, cellular / molecular, 2 references
[4] doi:10.3389/fgene.2026.1807347 [code]
Epigenetic regulators are preferentially coordinated with protocadherin gene expression across the human brain: a genome-wide co-expression analysis.
Journal: Frontiers in genetics
In common: statsmodels, pandas, SciPy, 2 other tools, genetics / omics, cellular / molecular, 2 references
[5] doi:10.1002/advs.202521254 [code]
Persistently Increased Expression of PKMzeta and Unbiased Gene Expression Profiles Identify Hippocampal Molecular Traces of a Long-Term Active Place Avoidance Memory and "Shadow" Proteins.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: statsmodels, pandas, SciPy, 2 other tools, genetics / omics, cellular / molecular, 2 references
[6] doi:10.1038/s41467-026-73428-y [code]
Regional heterogeneity in phenotypic and genetic associations between bone and brain in humans.
Journal: Nature communications
In common: pandas, SciPy, NumPy, genetics / omics, cellular / molecular, 3 references
[7] doi:10.1016/j.xgen.2026.101280 [code]
BMI-genome interactions regulate global gene expression with emphasis in brain and gut.
Journal: Cell genomics
In common: statsmodels, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
[8] doi:10.1038/s41467-026-75193-4 [code]
Multi-ancestry gene expression models amplify transcriptome-wide association study discovery and validation.
Journal: Nature communications
In common: statsmodels, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
[9] doi:10.1038/s41588-026-02646-3 [code]
Co-expression-based models improve eQTL predictions for transcriptome-wide association studies and highlight new schizophrenia-associated genes.
Journal: Nature genetics
In common: statsmodels, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
[10] doi:10.1038/s41380-026-03571-x [code]
Convergent coexpression reveals shared biological mechanisms underlying common and rare variant risk in six neuropsychiatric disorders.
Journal: Molecular psychiatry
In common: statsmodels, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.