OSCR

Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling.

Code ↔ Paper

40 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 40 matches
  1. [1] § 4. Materials and Methods › 4.8. Network and Pathway Analysis ↔ scripts/05_integration/disease_catalogue_ora.py, lines 1–37 · score 0.99 · GWAS_Catalog_2023, Jensen_DISEASES, background passed explicitly, Disease catalogue overlap, stratum RNA gene, DisGeNET
  2. [2] § 4. Materials and Methods › 4.6. Proteomic Processing ↔ scripts/03_proteomics/01cc_csf_astral_completecase.R, lines 110–191 · score 0.99 · left censored imputation, DEP behaviours, dep_bh_equivalence_check.R, test_diff, apparent abundance, dependent detection
  3. [3] § 4. Materials and Methods › 4.6. Proteomic Processing ↔ scripts/03_proteomics/02cc_csf_timstof_completecase.R, lines 109–177 · score 0.99 · left censored imputation, DEP behaviours, dep_bh_equivalence_check.R, test_diff, apparent abundance, dependent detection
  4. [4] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure4_proteomics.py, lines 1–42 · score 0.99 · Wang Julien brain, UK Biobank plasma, UK Biobank PPP, Bader Mann, CSF instrument, brain contrasts
  5. [5] § 2. Results › 2.4. Inverse-Concordant Discovery ↔ scripts/02_methylation/rebuild_S2_94genes.py, lines 1–43 · score 0.97 · methylation statistic aggregates, promoter region enrichment, correlates positively, gene body methylation, gene body probes, promoter silencing
  6. [6] § 4. Materials and Methods › 4.1. Data Sources and Cohort Assembly ↔ scripts/06_figures/figure_constants.py, lines 57–75 · score 0.96 · Astral DIA MS, MassIVE, raw spectra, Wang Julien, Bader Mann, region resolved brain
  7. [7] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure3_methylation.py, lines 1–49 · score 0.96 · uniform gene, cell baseline, limma DMP, fold change, stratum methylation, concordant proteomic anchors
  8. [8] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/02_methylation/14_wgbs_GSE173787_promoter.py, lines 1–38 · score 0.94 · genome bisulfite sequencing, CD14 monocytes, promoter reached, sorted populations, genes RUNX3, sorted immune
  9. [9] § 4. Materials and Methods › 4.1. Data Sources and Cohort Assembly ↔ scripts/05_integration/build_data_sources_table.py, lines 120–144 · score 0.94 · Mann CSF Orbitrap, Wang Julien brain, MassIVE, Astral DIA, CSF Astral, timsTOF
  10. [10] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure6_intersection_heatmap.py, lines 28–76 · score 0.92 · UK Biobank PPP, CSF instrument, brain region contrasts, proteomic compartments, single cell anchor, CSF Astral
  11. [11] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure6_intersection_heatmap.py, lines 28–76 · score 0.90 · fold change, pan tissue, proteomic compartments, single cell anchor, RNA FDR, auxiliary gene
  12. [12] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure4_proteomics.py, lines 1–42 · score 0.90 · Wang Julien brain, UK Biobank PPP, Bader Mann, Proteomic validation, concordant proteomic anchors, timsTOF
  13. [13] § 2. Results › 2.5. Proteomic Validation ↔ scripts/03_proteomics/11_itgb2_csf_pleocytosis.R, lines 1–55 · score 0.89 · Bader cohort records, leukocyte surface, detection rates, CSF protein, CSF leukocyte, CD45
  14. [14] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure7_string_network.py, lines 1–26 · score 0.89 · FOXP3 STAT1 STAT3, IKZF1 RUNX3, gene tiered, context genes, tiered candidates, CD79B
  15. [15] § 4. Materials and Methods › 4.4. DNA Methylation Processing ↔ scripts/02_methylation/14_wgbs_GSE173787_promoter.py, lines 1–38 · score 0.89 · GENCODE v49, genome bisulfite, BH correction, sorted cell, TSS, windows
  16. [16] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure7_string_network.py, lines 1–26 · score 0.89 · physical interaction network, canonical MS immune, pathway enrichment, IKZF1 RUNX3, gene tiered, context genes
  17. [17] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/09_inverse_concordance_scan.R, lines 1–60 · score 0.86 · treatment response contrast, disease candidate, genes entered, treatment context, tiered candidate, scan
  18. [18] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 167–250 · score 0.86 · 8–34 %, variance component, matched MS HC, composition shifts, donor limited, cell sampling
  19. [19] § 4. Materials and Methods › 4.4. DNA Methylation Processing ↔ scripts/02_methylation/normalize_beta_only.R, lines 1–52 · score 0.86 · sex chromosome probes, wateRmelon, normalize beta, quantile normalised, BMIQ, SNP
  20. [20] § 4. Materials and Methods › 4.9. Quality Control and Statistical Considerations ↔ scripts/04_singlecell/cell_level_wilcoxon_scan.py, lines 1–53 · score 0.86 · tie corrected Wilcoxon, Benjamini Hochberg correction, rank sum, rank genes, Scanpy, single cell
  21. [21] § 2. Results › 2.4. Inverse-Concordant Discovery ↔ scripts/02_methylation/update_S2_promoter_flag.py, lines 1–24 · score 0.84 · promoter region enrichment, gene body probes, promoter anchored, methylation call, mCSEA, attribution
  22. [22] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 167–250 · score 0.82 · deposited log normalised, CP10K, edgeR, limma voom, matched MS, limma trend
  23. [23] § 4. Materials and Methods › 4.3. Bulk Transcriptomic Processing ↔ scripts/03_proteomics/05_t_lineage_meta.R, lines 1–39 · score 0.80 · lmFit, gene symbols, eBayes, ComBat, sva, microarray
  24. [24] § 4. Materials and Methods › 4.3. Bulk Transcriptomic Processing ↔ scripts/01_transcriptome/harmonize_microarray_v2.py, lines 1–10 · score 0.80 · HGNC symbols, platform probe, microarray, GSE38010, GSE43591, GSE103005
  25. [25] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure2_rna_volcanoes.py, lines 53–63 · score 0.79 · FOXP3 STAT1 STAT3, IKZF1 RUNX3, gene tiered, context genes, CD79B, SH3BP4
  26. [26] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/04_singlecell/pseudobulk_brain_csf.R, lines 1–45 · score 0.78 · edgeR, limma voom, Kaufmann cohort, limma trend, weighting, Summed
  27. [27] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/02_methylation/09_inverse_concordance_scan.R, lines 1–60 · score 0.77 · inverse concordance scan, mCSEA prom, treatment response, treatment context, tiered candidates, baseline
  28. [28] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/05_integration/disease_catalogue_ora.py, lines 1–37 · score 0.77 · disease catalogues, DisGeNET, GWAS Catalog, Jensen DISEASES, background, inverse concordant
  29. [29] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/04_singlecell/cell_level_wilcoxon_scan.py, lines 1–53 · score 0.76 · IFI44L, LYN, CD79B, SH3BP4, MOSPD3, macrophages
  30. [30] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/helpers.R, lines 161–204 · score 0.75 · unweighted signed Stouffer, inverse variance, logFC, arithmetic, strata, BH
  31. [31] § 2. Results › 2.2. Per Cell-Type/Tissue Differential Expression ↔ scripts/06_figures/figure2_rna_volcanoes.py, lines 1–37 · score 0.75 · IFN PBMC, stratum RNA, concordant proteomic anchors, brain WM, auxiliary inverse concordant, inverse concordant Tier
  32. [32] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/06_figures/figure5_singlecell.py, lines 27–44 · score 0.74 · strong proteomic candidates, discovery arms, concordant proteomic anchor, UK Biobank, CD79B, SH3BP4
  33. [33] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure_constants.py, lines 57–75 · score 0.73 · Wang Julien, Bader Mann, DIA MS, UK Biobank, timsTOF, WML
  34. [34] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/04_singlecell/process_GSE118257_jakel_brain.py, lines 146–290 · score 0.73 · highly variable genes, Scanpy, Candidate gene, neighbour, Wilcoxon, clustering
  35. [35] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 1–25 · score 0.73 · IFI44L, immune context, CD79B, SH3BP4, MOSPD3, TYK2
  36. [36] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/rebuild_S2_94genes.py, lines 1–43 · score 0.69 · gene body probes, promoter silencing, mCSEA, S2, classified, composite
  37. [37] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/update_S2_promoter_flag.py, lines 1–24 · score 0.69 · gene body probes, promoter anchored, mCSEA, S2, silencing, classified
  38. [38] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure3_methylation.py, lines 1–49 · score 0.68 · cell baseline, concordant proteomic anchors, auxiliary inverse concordant, logFC, hypermethylation, DMF
  39. [39] § 2. Results › 2.2. Per Cell-Type/Tissue Differential Expression ↔ scripts/04_singlecell/pseudobulk_muscat_style.R, lines 1–17 · score 0.66 · IFI44L, expressed genes, MOSPD3, RPAP2, THRB, covariate
  40. [40] § 4. Materials and Methods › 4.9. Quality Control and Statistical Considerations ↔ scripts/06_figures/figure5_singlecell.py, lines 84–113 · score 0.62 · tie corrected Wilcoxon, Benjamini Hochberg, rank, Figure 5, log, candidate

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 250 lines · 13 KB · MIT · 3 matches

  1. #!/usr/bin/env python3
  2. """Assemble all donor-level pseudobulk results into one reviewer-ready Excel workbook."""
  3. import pandas as pd, numpy as np, os
  4. from statsmodels.stats.multitest import multipletests
  5. from openpyxl import Workbook
  6. from openpyxl.styles import Font, PatternFill, Alignment, Border, Side
  7. from openpyxl.utils import get_column_letter
  8. from openpyxl.utils.dataframe import dataframe_to_rows
  9. D="__MS_GEO_ROOT__/Poster_v2/figures/pseudobulk_proper"
  10. F="__MS_GEO_ROOT__/Poster_v2/figures"
  11. OUT="__MS_GEO_ROOT__/Poster_v2/Pseudobulk_DonorLevel_Results.xlsx"
  12. BH=lambda p: multipletests(p,method='fdr_bh')[1]
  13. T1=['ITGB2','IKZF1']
  14. SUG=['HLA-E']
  15. AUX=['CD79B','LXN','SH3BP4','CASP6','CASP8','DGKQ','MX1','IFIT1','NUP210','RUNX3']
  16. # FOXP3 was removed from this group on revision: it is not quantified in ANY of the seven
  17. # proteomic compartments (both CSF instruments, all four brain-region contrasts, UK Biobank-PPP),
  18. # so it could not be a "strong proteomic candidate", which is what defines this group. It is
  19. # retained in the STRING display as a canonical MS immune context gene, which is the role it
  20. # actually plays (the IKZF1-RUNX3/FOXP3-STAT1-STAT3 axis).
  21. PROT=['CTSZ','CHL1','ICAM1','ITGAL']
  22. CONTEXT=['THRB','IFI44L','RPAP2','SLAMF1','PCNP','STAT3','TYK2','MOSPD3']
  23. PANEL=T1+SUG+AUX+PROT+CONTEXT
  24. def tier(g):
  25. if g in T1: return 'Tier-1'
  26. if g in SUG: return 'suggestive (non-tier-1)'
  27. if g in AUX: return 'Tier-2 auxiliary inverse-concordant'
  28. if g in PROT: return 'Tier-2 non-concordant proteomic anchor'
  29. return 'context panel'
  30. # ---------------- load ----------------
  31. print("loading S1 (summed raw counts)...",flush=True)
  32. s1=pd.read_csv(f"{D}/pseudobulk_muscat_style.csv")
  33. bio=pd.read_csv(f"{D}/gene_biotypes.csv",index_col=0).biotype
  34. s1['biotype']=s1.gene.map(bio).fillna('unknown')
  35. print("loading S2/S3 (normalised aggregation)...",flush=True)
  36. s23=pd.read_csv(f"{D}/pseudobulk_norm_compare.csv")
  37. s23['biotype']=s23.gene.map(bio).fillna('unknown')
  38. # unify
  39. a=s1.rename(columns={'analysis':'design'})[['design','engine','cell_type','gene','logFC','PValue','FDR_local','FDR_global','biotype','n_ms','n_hc','n_pairs']]
  40. a['aggregation']='S1_sum_raw_counts'
  41. b=s23.rename(columns={'scheme':'design'})[['design','cell_type','gene','logFC','PValue','FDR_local','FDR_global','biotype','n_ms','n_hc','n_pairs']]
  42. b['engine']='limma_trend'
  43. b['aggregation']=np.where(b.design.str.startswith('S2'),'S2_mean_CP10K','S3_sum_CP10K_INVALID')
  44. allres=pd.concat([a,b],ignore_index=True)
  45. allres['key']=allres.aggregation+" | "+allres.design+" | "+allres.engine
  46. # panel-scope FDR within each design x engine
  47. allres['FDR_panel']=np.nan
  48. for k,idx in allres.groupby('key').groups.items():
  49. sub=allres.loc[idx]
  50. p=sub[sub.gene.isin(PANEL)]
  51. if len(p): allres.loc[p.index,'FDR_panel']=BH(p.PValue.values)
  52. # ---------------- sheets ----------------
  53. sheets={}
  54. # 1) design summary
  55. rows=[]
  56. for k,s in allres.groupby('key'):
  57. p=s[s.gene.isin(PANEL)]
  58. rows.append(dict(Aggregation=s.aggregation.iloc[0], Design=s.design.iloc[0], Engine=s.engine.iloc[0],
  59. Cell_types=s.cell_type.nunique(), Tests=len(s),
  60. Genes_FDRglobal_lt005=int((s.FDR_global<0.05).sum()),
  61. Genes_FDRlocal_lt005=int((s.FDR_local<0.05).sum()),
  62. Panel_tests=len(p), Panel_FDR_lt005=int((p.FDR_panel<0.05).sum()),
  63. MS_donors=int(s.n_ms.max()), HC_donors=int(s.n_hc.max())))
  64. sheets['01_Design_summary']=pd.DataFrame(rows).sort_values(['Aggregation','Design','Engine'])
  65. # 2) tier-1 headline: best result per gene per aggregation
  66. rows=[]
  67. for g in T1+SUG+['MOSPD3','IFI44L','MX1','DGKQ']:
  68. for agg,s in allres[allres.gene==g].groupby('aggregation'):
  69. if not len(s): continue
  70. # select the row that determines significance: smallest panel-scope FDR (tie-break on p)
  71. s=s.sort_values(['FDR_panel','PValue'], na_position='last')
  72. r=s.iloc[0]
  73. invalid = agg.endswith('INVALID')
  74. rows.append(dict(Tier=tier(g), Gene=g, Aggregation=agg,
  75. Scheme_valid="NO - do not cite" if invalid else "yes",
  76. Best_cell_type=r.cell_type,
  77. Design=r.design, Engine=r.engine, logFC=r.logFC, P_value=r.PValue,
  78. FDR_local=r.FDR_local, FDR_global=r.FDR_global, FDR_panel=r.FDR_panel,
  79. MS_donors=r.n_ms, HC_donors=r.n_hc,
  80. Significant_panel=("n/a (invalid scheme)" if invalid
  81. else ("YES" if r.FDR_panel<0.05 else "no"))))
  82. sheets['02_Candidates_best']=pd.DataFrame(rows).sort_values(['Scheme_valid','Tier','Gene'],
  83. ascending=[True,True,True])
  84. # 3) aggregation comparison (panel FDR side by side)
  85. piv=allres[allres.gene.isin(PANEL)].groupby(['gene','aggregation']).FDR_panel.min().unstack()
  86. pv =allres[allres.gene.isin(PANEL)].groupby(['gene','aggregation']).PValue.min().unstack()
  87. cmp=pd.DataFrame({'Gene':piv.index,'Tier':[tier(g) for g in piv.index]})
  88. for c in piv.columns:
  89. cmp[f'minP_{c}']=pv[c].values; cmp[f'panelFDR_{c}']=piv[c].values
  90. sheets['03_Aggregation_comparison']=cmp.sort_values(['Tier','Gene'])
  91. # 4) full candidate results (valid schemes only)
  92. full=allres[allres.gene.isin(PANEL) & (allres.aggregation!='S3_sum_CP10K_INVALID')].copy()
  93. full['Tier']=full.gene.map(tier)
  94. sheets['04_Candidates_full_valid']=full[['Tier','gene','cell_type','aggregation','design','engine',
  95. 'logFC','PValue','FDR_local','FDR_global','FDR_panel','n_ms','n_hc','n_pairs']]\
  96. .sort_values(['Tier','gene','PValue']).rename(columns={'gene':'Gene','cell_type':'Cell_type'})
  97. # 5) S3 artifact demonstration
  98. s3=allres[(allres.aggregation=='S3_sum_CP10K_INVALID')&(allres.gene.isin(PANEL))]
  99. art=s3[s3.design=='S3_sumCP10K_coarse8'].sort_values('PValue').head(40)
  100. sheets['05_S3_artifact_demo']=art[['gene','cell_type','logFC','PValue','FDR_local','FDR_global','FDR_panel']]\
  101. .rename(columns={'gene':'Gene','cell_type':'Cell_type'})
  102. # 6) differential abundance
  103. das=[]
  104. for tag in ['coarse','fine']:
  105. p=f"{D}/DA_{tag}.csv"
  106. if os.path.exists(p):
  107. d=pd.read_csv(p); d.insert(0,'Resolution',tag); das.append(d)
  108. if das:
  109. da=pd.concat(das,ignore_index=True)
  110. keep=[c for c in ['Resolution','cell_type','logFC','AveExpr','t','P.Value','adj.P.Val'] if c in da.columns]
  111. sheets['06_Differential_abundance']=da[keep]
  112. # 7) variance components
  113. M=pd.read_csv(f"{D}/PB_lognorm_matrix.csv",index_col=0)
  114. C=pd.read_csv(f"{D}/PB_lognorm_coldata.csv"); C=C[C['sample'].isin(M.columns)]
  115. rows=[]
  116. for g in T1+SUG:
  117. if g not in M.index: continue
  118. for ct in ['t_cells','monocytes','b_cells','nk_cells']:
  119. s=C[C.cell_type==ct]
  120. if not len(s): continue
  121. v=M.loc[g,s['sample'].values].astype(float).values; nc=s.n_cells.values
  122. dsd=float(np.std(v,ddof=1)); cse=float(np.mean(np.sqrt(max(v.mean(),1e-9)/nc)))
  123. rows.append(dict(Gene=g,Cell_type=ct,Donors=len(v),Mean_cells_per_donor=int(nc.mean()),
  124. Between_donor_SD=dsd,Cell_sampling_SE=cse,Ratio_SE_over_SD=cse/max(dsd,1e-9)))
  125. sheets['07_Variance_components']=pd.DataFrame(rows)
  126. # 8) cell-level vs donor-level
  127. p=f"{F}/scrna_PSEUDOBULK_COMPARISON.tsv"
  128. if os.path.exists(p):
  129. cd=pd.read_csv(p,sep='\t')
  130. cd=cd[cd.gene.isin(PANEL)].copy(); cd.insert(0,'Tier',cd.gene.map(tier))
  131. sheets['08_CellLevel_vs_Donor']=cd.sort_values(['Tier','cell_wilcox_FDR_OLD'])
  132. # 9) top genome-wide hits (pipeline sensitivity evidence)
  133. tops=[]
  134. for k,s in allres[allres.aggregation=='S1_sum_raw_counts'].groupby('key'):
  135. t=s[s.FDR_global<0.05].nsmallest(15,'FDR_global')
  136. if len(t): tops.append(t.assign(Source=k))
  137. if tops:
  138. tp=pd.concat(tops,ignore_index=True)
  139. sheets['09_TopGenomewide_S1']=tp[['Source','cell_type','gene','biotype','logFC','PValue','FDR_global']]\
  140. .rename(columns={'gene':'Gene','cell_type':'Cell_type'})
  141. # 10) biotype summary
  142. sheets['10_Gene_biotypes']=bio.value_counts().rename_axis('Biotype').reset_index(name='N_genes')
  143. # ---------------- write ----------------
  144. wb=Workbook(); wb.remove(wb.active)
  145. HDR=PatternFill('solid',start_color='1F4E78'); HF=Font(name='Arial',bold=True,color='FFFFFF',size=10)
  146. BF=Font(name='Arial',size=10); TF=Font(name='Arial',bold=True,size=12)
  147. SIG=PatternFill('solid',start_color='C6EFCE'); WARN=PatternFill('solid',start_color='FFC7CE')
  148. thin=Side(style='thin',color='BFBFBF'); BRD=Border(bottom=thin)
  149. # README
  150. ws=wb.create_sheet('00_README')
  151. readme=[
  152. ("Donor-level pseudobulk differential-state analysis - Kaufmann PBMC (GSE144744)",'title'),
  153. ("",''),
  154. ("Cohort: 31 MS / 31 HC donors, 497,705 cells, matched MS-HC pairs multiplexed per 10x run (31 pairs).",''),
  155. ("Design: paired ~batch_pair + condition. Pairs are sex-matched 31/31; mean age 38.48 (HC) vs 38.52 (MS),",''),
  156. ("so the pair term absorbs sex and age - no extra covariates needed.",''),
  157. ("",''),
  158. ("Raw counts were recovered exactly from the deposited log-normalised matrix:",''),
  159. (" count = expm1(lognorm) * nCount_RNA / 1e4 -- verified 100% integer; per-cell expm1 sums = 10,000;",''),
  160. (" reconstructed per-cell totals equal the recorded nCount_RNA exactly.",''),
  161. ("",''),
  162. ("AGGREGATION SCHEMES",'h'),
  163. ("S1_sum_raw_counts SUM of raw integer UMIs per donor x cell type -> edgeR-QL and limma-voom (muscat standard)",''),
  164. ("S2_mean_CP10K MEAN of per-cell CP10K, log2 -> limma-trend (equal weight per cell)",''),
  165. ("S3_sum_CP10K_INVALID SUM of per-cell CP10K - REJECTED, see sheet 05: cell number re-enters as an",''),
  166. (" uncontrolled covariate (21% of genes reach FDR<0.05; effects track cell proportion)",''),
  167. ("",''),
  168. ("RESOLUTIONS: 8 coarse lineages / 25 published subsets / whole-PBMC (one sample per donor)",''),
  169. ("FILTERS: min 10 cells per donor x cell type; filterByExpr or >=50% non-zero; protein-coding only (mygene.info)",''),
  170. ("FDR SCOPES reported separately: local (within cell type), global (all gene x cell type), panel (26 candidates)",''),
  171. ("",''),
  172. ("KEY RESULTS",'h'),
  173. ("IKZF1 panel FDR 0.047 (whole-PBMC, S2 mean-CP10K), logFC -0.078; direction concordant with the bulk RNA",''),
  174. (" down-call. Not significant under S1 (0.405). Most defensible candidate signal.",''),
  175. ("HLA-E panel FDR 0.047 (cDC, S1 voom), logFC +0.212 - direction DISCORDANT with the bulk down-call.",''),
  176. ("MOSPD3 panel FDR 0.043 (S1) / 0.047 (S2), whole-PBMC - but not one of the 17 named candidates.",''),
  177. ("ITGB2 null under every scheme (panel FDR 0.42-0.44). Its tier-1 anchoring rests on bulk RNA (q=3.4e-4),",''),
  178. (" promoter methylation (mCSEA q=0.016) and UK Biobank plasma proteomics (q=1.5e-17), not single-cell.",''),
  179. ("No candidate is robust across both valid aggregation schemes.",''),
  180. ("",''),
  181. ("WHY MORE POWER IS NOT AVAILABLE (sheet 07)",'h'),
  182. ("Cell-sampling SE is only 8-34% of between-donor SD, so the contrast is donor-limited: effective n = 62",''),
  183. ("donors, not 497,705 cells. Mixed models / kNN / cell-selection act on the cell dimension and cannot help.",''),
  184. ("Cell-type composition does not differ (sheet 06, all FDR > 0.19), so composition shift is not an explanation.",''),
  185. ("",''),
  186. ("SHEETS",'h'),
  187. ("01 Design summary | 02 Candidate best results | 03 Aggregation comparison | 04 Full candidate results",''),
  188. ("05 S3 artifact demo | 06 Differential abundance | 07 Variance components | 08 Cell- vs donor-level",''),
  189. ("09 Top genome-wide hits | 10 Gene biotypes",''),
  190. ]
  191. for i,(txt,kind) in enumerate(readme,1):
  192. c=ws.cell(row=i,column=1,value=txt)
  193. c.font=TF if kind=='title' else (Font(name='Arial',bold=True,size=10) if kind=='h' else BF)
  194. ws.column_dimensions['A'].width=120
  195. for name,df in sheets.items():
  196. ws=wb.create_sheet(name[:31])
  197. for r in dataframe_to_rows(df,index=False,header=True): ws.append(r)
  198. for c in ws[1]: c.fill=HDR; c.font=HF; c.alignment=Alignment(horizontal='center',vertical='center',wrap_text=True)
  199. ws.freeze_panes='A2'
  200. hdr=[c.value for c in ws[1]]
  201. for j,h in enumerate(hdr,1):
  202. L=get_column_letter(j)
  203. ws.column_dimensions[L].width=min(max(12,len(str(h))+3),30)
  204. for i in range(2,ws.max_row+1):
  205. cell=ws.cell(row=i,column=j); cell.font=BF; cell.border=BRD
  206. v=cell.value
  207. if isinstance(v,float):
  208. if h and ('P_value' in str(h) or 'PValue' in str(h) or str(h).startswith('minP') or 'P.Value' in str(h)):
  209. cell.number_format='0.00E+00'
  210. elif h and ('FDR' in str(h) or 'adj.P' in str(h)):
  211. cell.number_format='0.0000'
  212. if v<0.05: cell.fill=SIG
  213. elif h and ('logFC' in str(h) or 'SD' in str(h) or 'SE' in str(h) or 'Ratio' in str(h)):
  214. cell.number_format='0.0000'
  215. else: cell.number_format='0.000'
  216. if h=='Significant_panel' and v=='YES': cell.fill=SIG
  217. if h=='Scheme_valid' and isinstance(v,str) and v.startswith('NO'): cell.fill=WARN
  218. # grey out / flag rows from the invalid S3 scheme so they cannot be misread as findings
  219. if 'Aggregation' in hdr:
  220. ja=hdr.index('Aggregation')+1
  221. for i in range(2,ws.max_row+1):
  222. if str(ws.cell(row=i,column=ja).value).endswith('INVALID'):
  223. for j in range(1,len(hdr)+1):
  224. ws.cell(row=i,column=j).fill=WARN
  225. ws.cell(row=i,column=j).font=Font(name='Arial',size=10,italic=True,color='9C0006')
  226. if name=='05_S3_artifact_demo':
  227. for i in range(2,ws.max_row+1):
  228. for j in range(1,len(hdr)+1): ws.cell(row=i,column=j).fill=WARN
  229. wb.save(OUT)
  230. print(f"\nwrote {OUT}")
  231. for n,d in sheets.items(): print(f" {n:32s} {len(d):6d} rows x {len(d.columns)} cols")

build_pseudobulk_workbook.py at commit 056de0d, under MIT · at the source

Overview

  1. Department of Biostatistics and Bioinformatics, Graduate School of Health Sciences, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
  2. Department of Molecular Biology and Genetics, Faculty of Engineering and Natural Sciences, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
  3. Department of Molecular Biology and Genetics, Graduate School of Natural and Applied Science, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
Institutions: Acıbadem University (Türkiye)
Journal: International journal of molecular sciences, volume 27, issue 16, article 7275
Dates: received 29 June 2026; accepted 12 August 2026; published online 14 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/ijms27167275 · PMID 42653279 · PMCID PMC13513955 · OpenAlex W7203481781
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), multiple sclerosis (population), cellular / molecular (subfield)
Methods: Spectral & time-frequency, Statistics, Connectivity
Keywords: multiple sclerosis, multi-omics integration, DNA methylation, single-cell RNA sequencing, proteomics, network biology
MeSH: DNA Methylation*, Multiple Sclerosis*, Gene Expression Regulation, Humans, Multiomics, Proteomics, Transcriptome (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: not cited yet (Europe PMC); 110 references in the paper

Abstract

Multiple sclerosis (MS) is an immune-mediated central nervous system disease with mapped genetic risk loci but unresolved cross-layer regulatory mechanisms. We harmonised 30 primary public MS datasets spanning bulk transcriptomics (552 samples, 367 MS/185 healthy control [HC]), DNA methylation arrays (475 samples, 244 MS/231 HC) with sorted-cell whole-genome bisulfite sequencing (WGBS) assessment, single-cell RNA sequencing (RNA-seq; 517,533 cells from 81 donors), and cerebrospinal fluid (CSF) and brain white-matter proteomes, and UK Biobank Pharma Proteomics Project (UK Biobank-PPP) plasma proteomic statistics (407 MS/39,979 HC) for external validation. We prioritised MS-associated genes showing inverse-concordant, cohort-level methylation–expression associations and orthogonal support across layers. After harmonised reprocessing and per-stratum differential testing, a bulk-RNA × methylation inverse-concordance filter followed by an independent proteomic and/or donor-level single-cell anchor identified two Tier-1 candidates (IKZF1’s anchor is a candidate-panel-adjusted donor-level pseudobulk association, not transcriptome-wide significance): ITGB2 and IKZF1. ITGB2 showed the strongest cross-layer support: RNA down-regulation, promoter hypermethylation and reduced plasma abundance in UK Biobank MS cases, a pattern consistent with altered leukocyte-integrin adhesion. Pathway and network analyses linked the candidates predominantly to leukocyte and lymphocyte activation and differentiation, leukocyte-integrin adhesion, and interferon and cytokine-signalling immune modules. These findings nominate MS-relevant regulatory candidates and provide a reusable multi-omics framework for hypothesis generation and validation.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.

alperbulbul1/MS_Omics_Data_Analysis

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 056de0d6f94fa3cc3044ca5b0f33db3a8349c3c5, 6 August 2026
Languages: R (51), Python (50), Shell (1)
Size: 110 files, 102 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, license file, environment (env/requirements.txt), documentation
Not found: CITATION.cff, tests, continuous integration
Tools: pandas (46 files), NumPy (37 files), data.table (36 files), limma (36 files), ggplot2 (31 files), tidyverse (31 files), Matplotlib (17 files), SciPy (11 files), Scanpy (9 files), statsmodels (8 files), seaborn (7 files), anndata (6 files), pheatmap (6 files), Pillow (6 files), reshape2 (3 files), edgeR (2 files), NetworkX (2 files), neuroCombat (2 files), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
104 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 102 scripts, each with its path and the digest of its content;
  • 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

All primary public accessions used in this study are catalogued per-cohort in Supplementary Table S1. NCBI GEO: Bulk transcriptomics (15 series): GSE21942 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE21942), GSE38010 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE38010), GSE43591 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE43591), GSE103005 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE103005), GSE137143 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE137143), GSE138064 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE138064), GSE172009 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE172009), GSE173789 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE173789), GSE190847 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE190847), GSE207680 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE207680), GSE209596 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE209596), GSE211358 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE211358), GSE211739 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE211739), GSE214334 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE214334) and GSE288904 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE288904); GSE66573 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE66573) contributes to the harmonised expression matrix but not to the 552-sample differential analysis set; DNA methylation (eight Illumina 450K/EPIC array series (GSE40360 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE40360), GSE88824 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE88824), GSE106648 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE106648), GSE130029 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE130029), GSE130030 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE130030), GSE189255 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE189255), GSE189256 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE189256) and GSE219293 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE219293)) plus the GSE173787 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE173787) whole-genome bisulfite-sequencing series); single-cell RNA-sequencing (3 series): GSE118257 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE118257) (Jäkel et al., 2019 [11] brain snRNA-seq), GSE127969 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE127969) (Beltrán et al., 2019 [14] CSF/PBMC twin pairs) and GSE144744 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE144744) (Kaufmann et al., 2021 [13] PBMC 10× Chromium 3′ scRNA-seq/CITE-seq). PRIDE, proteomics: PXD064570 (Bader & Mann 2026 [8] CSF Orbitrap Astral DIA-MS) and PXD045058 (the same study’s timsTOF CSF dataset); the Wang & Julien 2026 [9] region-resolved brain white-matter proteome is deposited on MassIVE (MSV000096790, raw spectra) and Figshare (processed data). UK Biobank-PPP plasma proteomic associations were taken from the published summary statistics of Jacobs et al., 2024 [25] (Olink Explore; PMID 38282238. doi:10.1002/acn3.51990); individual-level UK Biobank data are available to approved researchers via the UK Biobank Access Management System. All analysis code supporting this study, covering the harmonisation, per-stratum differential testing, inverse-concordance scan, proteomic, single-cell, network and figure-generation steps, is available at https://github.com/alperbulbul1/MS_Omics_Data_Analysis (accessed on 5 August 2026).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 6 keywords, 7 MeSH terms, 110 references.

Cite

This paper

Bülbül, A., Yılmaz-İşgördü, Ö. H., Reda, M. D., & Turanlı, E. T. (2026). Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling. International journal of molecular sciences, 27(16), 7275. https://doi.org/10.3390/ijms27167275

BibTeX

@article{bulbul2026integrative,
author = {Bülbül, Alper and Yılmaz-İşgördü, Özdeyiş Hülya and Reda, Meziyet Dilara and Turanlı, Eda Tahir},
title = {{Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling}},
journal = {International journal of molecular sciences},
year = {2026},
month = aug,
volume = {27},
number = {16},
pages = {7275},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1422-0067},
doi = {10.3390/ijms27167275},
url = {https://doi.org/10.3390/ijms27167275},
pmid = {42653279},
pmcid = {PMC13513955}
}

RIS

TY - JOUR
AU - Bülbül, Alper
AU - Yılmaz-İşgördü, Özdeyiş Hülya
AU - Reda, Meziyet Dilara
AU - Turanlı, Eda Tahir
TI - Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling
T2 - International journal of molecular sciences
J2 - Int J Mol Sci
PY - 2026
DA - 2026/08/14
VL - 27
IS - 16
SP - 7275
SN - 1422-0067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/ijms27167275
UR - https://doi.org/10.3390/ijms27167275
LA - en
ER -

CSL-JSON

{
"id": "10.3390/ijms27167275",
"type": "article-journal",
"title": "Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling",
"container-title": "International journal of molecular sciences",
"author": [
{
"family": "Bülbül",
"given": "Alper"
},
{
"family": "Yılmaz-İşgördü",
"given": "Özdeyiş Hülya"
},
{
"family": "Reda",
"given": "Meziyet Dilara"
},
{
"family": "Turanlı",
"given": "Eda Tahir"
}
],
"container-title-short": "Int J Mol Sci",
"volume": "27",
"issue": "16",
"page": "7275",
"DOI": "10.3390/ijms27167275",
"PMID": "42653279",
"PMCID": "PMC13513955",
"ISSN": "1422-0067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/ijms27167275",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: edgeR, limma, anndata, 14 other tools, genetics / omics, cellular / molecular, 1 reference
[2] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: edgeR, limma, anndata, 13 other tools, genetics / omics, cellular / molecular
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: edgeR, limma, pheatmap, 13 other tools, cellular / molecular
[4] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: edgeR, limma, pheatmap, 11 other tools, genetics / omics, 2 references
[5] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: limma, anndata, Scanpy, 12 other tools, genetics / omics, 1 reference
[6] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: edgeR, anndata, Scanpy, 11 other tools, genetics / omics, cellular / molecular, 2 references
[7] doi:10.1016/j.celrep.2026.117073 [code]
Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.
Journal: Cell reports
In common: edgeR, anndata, Scanpy, 11 other tools, genetics / omics, cellular / molecular, 1 reference
[8] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: edgeR, limma, anndata, 12 other tools, genetics / omics
[9] doi:10.1038/s41467-026-71803-3 [code]
Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.
Journal: Nature communications
In common: anndata, Scanpy, pheatmap, 9 other tools, multiple sclerosis, genetics / omics, cellular / molecular, 3 references
[10] doi:10.1016/j.cpblue.2026.100007 [code]
An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.
Journal: Cell press blue
In common: edgeR, limma, anndata, 12 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.