Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling.
The 40 matches
- [1] § 4. Materials and Methods › 4.8. Network and Pathway Analysis ↔ scripts/05_integration/disease_catalogue_ora.py, lines 1–37 · score 0.99 · GWAS_Catalog_2023, Jensen_DISEASES, background passed explicitly, Disease catalogue overlap, stratum RNA gene, DisGeNET
- [2] § 4. Materials and Methods › 4.6. Proteomic Processing ↔ scripts/03_proteomics/01cc_csf_astral_completecase.R, lines 110–191 · score 0.99 · left censored imputation, DEP behaviours, dep_bh_equivalence_check.R, test_diff, apparent abundance, dependent detection
- [3] § 4. Materials and Methods › 4.6. Proteomic Processing ↔ scripts/03_proteomics/02cc_csf_timstof_completecase.R, lines 109–177 · score 0.99 · left censored imputation, DEP behaviours, dep_bh_equivalence_check.R, test_diff, apparent abundance, dependent detection
- [4] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure4_proteomics.py, lines 1–42 · score 0.99 · Wang Julien brain, UK Biobank plasma, UK Biobank PPP, Bader Mann, CSF instrument, brain contrasts
- [5] § 2. Results › 2.4. Inverse-Concordant Discovery ↔ scripts/02_methylation/rebuild_S2_94genes.py, lines 1–43 · score 0.97 · methylation statistic aggregates, promoter region enrichment, correlates positively, gene body methylation, gene body probes, promoter silencing
- [6] § 4. Materials and Methods › 4.1. Data Sources and Cohort Assembly ↔ scripts/06_figures/figure_constants.py, lines 57–75 · score 0.96 · Astral DIA MS, MassIVE, raw spectra, Wang Julien, Bader Mann, region resolved brain
- [7] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure3_methylation.py, lines 1–49 · score 0.96 · uniform gene, cell baseline, limma DMP, fold change, stratum methylation, concordant proteomic anchors
- [8] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/02_methylation/14_wgbs_GSE173787_promoter.py, lines 1–38 · score 0.94 · genome bisulfite sequencing, CD14 monocytes, promoter reached, sorted populations, genes RUNX3, sorted immune
- [9] § 4. Materials and Methods › 4.1. Data Sources and Cohort Assembly ↔ scripts/05_integration/build_data_sources_table.py, lines 120–144 · score 0.94 · Mann CSF Orbitrap, Wang Julien brain, MassIVE, Astral DIA, CSF Astral, timsTOF
- [10] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure6_intersection_heatmap.py, lines 28–76 · score 0.92 · UK Biobank PPP, CSF instrument, brain region contrasts, proteomic compartments, single cell anchor, CSF Astral
- [11] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure6_intersection_heatmap.py, lines 28–76 · score 0.90 · fold change, pan tissue, proteomic compartments, single cell anchor, RNA FDR, auxiliary gene
- [12] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure4_proteomics.py, lines 1–42 · score 0.90 · Wang Julien brain, UK Biobank PPP, Bader Mann, Proteomic validation, concordant proteomic anchors, timsTOF
- [13] § 2. Results › 2.5. Proteomic Validation ↔ scripts/03_proteomics/11_itgb2_csf_pleocytosis.R, lines 1–55 · score 0.89 · Bader cohort records, leukocyte surface, detection rates, CSF protein, CSF leukocyte, CD45
- [14] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure7_string_network.py, lines 1–26 · score 0.89 · FOXP3 STAT1 STAT3, IKZF1 RUNX3, gene tiered, context genes, tiered candidates, CD79B
- [15] § 4. Materials and Methods › 4.4. DNA Methylation Processing ↔ scripts/02_methylation/14_wgbs_GSE173787_promoter.py, lines 1–38 · score 0.89 · GENCODE v49, genome bisulfite, BH correction, sorted cell, TSS, windows
- [16] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure7_string_network.py, lines 1–26 · score 0.89 · physical interaction network, canonical MS immune, pathway enrichment, IKZF1 RUNX3, gene tiered, context genes
- [17] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/09_inverse_concordance_scan.R, lines 1–60 · score 0.86 · treatment response contrast, disease candidate, genes entered, treatment context, tiered candidate, scan
- [18] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 167–250 · score 0.86 · 8–34 %, variance component, matched MS HC, composition shifts, donor limited, cell sampling
- [19] § 4. Materials and Methods › 4.4. DNA Methylation Processing ↔ scripts/02_methylation/normalize_beta_only.R, lines 1–52 · score 0.86 · sex chromosome probes, wateRmelon, normalize beta, quantile normalised, BMIQ, SNP
- [20] § 4. Materials and Methods › 4.9. Quality Control and Statistical Considerations ↔ scripts/04_singlecell/cell_level_wilcoxon_scan.py, lines 1–53 · score 0.86 · tie corrected Wilcoxon, Benjamini Hochberg correction, rank sum, rank genes, Scanpy, single cell
- [21] § 2. Results › 2.4. Inverse-Concordant Discovery ↔ scripts/02_methylation/update_S2_promoter_flag.py, lines 1–24 · score 0.84 · promoter region enrichment, gene body probes, promoter anchored, methylation call, mCSEA, attribution
- [22] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 167–250 · score 0.82 · deposited log normalised, CP10K, edgeR, limma voom, matched MS, limma trend
- [23] § 4. Materials and Methods › 4.3. Bulk Transcriptomic Processing ↔ scripts/03_proteomics/05_t_lineage_meta.R, lines 1–39 · score 0.80 · lmFit, gene symbols, eBayes, ComBat, sva, microarray
- [24] § 4. Materials and Methods › 4.3. Bulk Transcriptomic Processing ↔ scripts/01_transcriptome/harmonize_microarray_v2.py, lines 1–10 · score 0.80 · HGNC symbols, platform probe, microarray, GSE38010, GSE43591, GSE103005
- [25] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/06_figures/figure2_rna_volcanoes.py, lines 53–63 · score 0.79 · FOXP3 STAT1 STAT3, IKZF1 RUNX3, gene tiered, context genes, CD79B, SH3BP4
- [26] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/04_singlecell/pseudobulk_brain_csf.R, lines 1–45 · score 0.78 · edgeR, limma voom, Kaufmann cohort, limma trend, weighting, Summed
- [27] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/02_methylation/09_inverse_concordance_scan.R, lines 1–60 · score 0.77 · inverse concordance scan, mCSEA prom, treatment response, treatment context, tiered candidates, baseline
- [28] § 2. Results › 2.7. STRING Physical-Interaction Network and Pathway Enrichment ↔ scripts/05_integration/disease_catalogue_ora.py, lines 1–37 · score 0.77 · disease catalogues, DisGeNET, GWAS Catalog, Jensen DISEASES, background, inverse concordant
- [29] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/04_singlecell/cell_level_wilcoxon_scan.py, lines 1–53 · score 0.76 · IFI44L, LYN, CD79B, SH3BP4, MOSPD3, macrophages
- [30] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/helpers.R, lines 161–204 · score 0.75 · unweighted signed Stouffer, inverse variance, logFC, arithmetic, strata, BH
- [31] § 2. Results › 2.2. Per Cell-Type/Tissue Differential Expression ↔ scripts/06_figures/figure2_rna_volcanoes.py, lines 1–37 · score 0.75 · IFN PBMC, stratum RNA, concordant proteomic anchors, brain WM, auxiliary inverse concordant, inverse concordant Tier
- [32] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/06_figures/figure5_singlecell.py, lines 27–44 · score 0.74 · strong proteomic candidates, discovery arms, concordant proteomic anchor, UK Biobank, CD79B, SH3BP4
- [33] § 2. Results › 2.5. Proteomic Validation ↔ scripts/06_figures/figure_constants.py, lines 57–75 · score 0.73 · Wang Julien, Bader Mann, DIA MS, UK Biobank, timsTOF, WML
- [34] § 4. Materials and Methods › 4.7. Single-Cell Processing ↔ scripts/04_singlecell/process_GSE118257_jakel_brain.py, lines 146–290 · score 0.73 · highly variable genes, Scanpy, Candidate gene, neighbour, Wilcoxon, clustering
- [35] § 2. Results › 2.6. Single-Cell Assessment in Brain, Blood and CSF ↔ scripts/06_figures/build_pseudobulk_workbook.py, lines 1–25 · score 0.73 · IFI44L, immune context, CD79B, SH3BP4, MOSPD3, TYK2
- [36] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/rebuild_S2_94genes.py, lines 1–43 · score 0.69 · gene body probes, promoter silencing, mCSEA, S2, classified, composite
- [37] § 4. Materials and Methods › 4.5. Inverse-Concordant Cross-Omics Scan ↔ scripts/02_methylation/update_S2_promoter_flag.py, lines 1–24 · score 0.69 · gene body probes, promoter anchored, mCSEA, S2, silencing, classified
- [38] § 2. Results › 2.3. Methylation Regulatory Landscape ↔ scripts/06_figures/figure3_methylation.py, lines 1–49 · score 0.68 · cell baseline, concordant proteomic anchors, auxiliary inverse concordant, logFC, hypermethylation, DMF
- [39] § 2. Results › 2.2. Per Cell-Type/Tissue Differential Expression ↔ scripts/04_singlecell/pseudobulk_muscat_style.R, lines 1–17 · score 0.66 · IFI44L, expressed genes, MOSPD3, RPAP2, THRB, covariate
- [40] § 4. Materials and Methods › 4.9. Quality Control and Statistical Considerations ↔ scripts/06_figures/figure5_singlecell.py, lines 84–113 · score 0.62 · tie corrected Wilcoxon, Benjamini Hochberg, rank, Figure 5, log, candidate
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 250 lines · 13 KB · MIT · 3 matches
- #!/usr/bin/env python3
- """Assemble all donor-level pseudobulk results into one reviewer-ready Excel workbook."""
- import pandas as pd, numpy as np, os
- from statsmodels.stats.multitest import multipletests
- from openpyxl import Workbook
- from openpyxl.styles import Font, PatternFill, Alignment, Border, Side
- from openpyxl.utils import get_column_letter
- from openpyxl.utils.dataframe import dataframe_to_rows
- D="__MS_GEO_ROOT__/Poster_v2/figures/pseudobulk_proper"
- F="__MS_GEO_ROOT__/Poster_v2/figures"
- OUT="__MS_GEO_ROOT__/Poster_v2/Pseudobulk_DonorLevel_Results.xlsx"
- BH=lambda p: multipletests(p,method='fdr_bh')[1]
- T1=['ITGB2','IKZF1']
- SUG=['HLA-E']
- AUX=['CD79B','LXN','SH3BP4','CASP6','CASP8','DGKQ','MX1','IFIT1','NUP210','RUNX3']
- # FOXP3 was removed from this group on revision: it is not quantified in ANY of the seven
- # proteomic compartments (both CSF instruments, all four brain-region contrasts, UK Biobank-PPP),
- # so it could not be a "strong proteomic candidate", which is what defines this group. It is
- # retained in the STRING display as a canonical MS immune context gene, which is the role it
- # actually plays (the IKZF1-RUNX3/FOXP3-STAT1-STAT3 axis).
- PROT=['CTSZ','CHL1','ICAM1','ITGAL']
- CONTEXT=['THRB','IFI44L','RPAP2','SLAMF1','PCNP','STAT3','TYK2','MOSPD3']
- PANEL=T1+SUG+AUX+PROT+CONTEXT
- def tier(g):
- if g in T1: return 'Tier-1'
- if g in SUG: return 'suggestive (non-tier-1)'
- if g in AUX: return 'Tier-2 auxiliary inverse-concordant'
- if g in PROT: return 'Tier-2 non-concordant proteomic anchor'
- return 'context panel'
- # ---------------- load ----------------
- print("loading S1 (summed raw counts)...",flush=True)
- s1=pd.read_csv(f"{D}/pseudobulk_muscat_style.csv")
- bio=pd.read_csv(f"{D}/gene_biotypes.csv",index_col=0).biotype
- s1['biotype']=s1.gene.map(bio).fillna('unknown')
- print("loading S2/S3 (normalised aggregation)...",flush=True)
- s23=pd.read_csv(f"{D}/pseudobulk_norm_compare.csv")
- s23['biotype']=s23.gene.map(bio).fillna('unknown')
- # unify
- a=s1.rename(columns={'analysis':'design'})[['design','engine','cell_type','gene','logFC','PValue','FDR_local','FDR_global','biotype','n_ms','n_hc','n_pairs']]
- a['aggregation']='S1_sum_raw_counts'
- b=s23.rename(columns={'scheme':'design'})[['design','cell_type','gene','logFC','PValue','FDR_local','FDR_global','biotype','n_ms','n_hc','n_pairs']]
- b['engine']='limma_trend'
- b['aggregation']=np.where(b.design.str.startswith('S2'),'S2_mean_CP10K','S3_sum_CP10K_INVALID')
- allres=pd.concat([a,b],ignore_index=True)
- allres['key']=allres.aggregation+" | "+allres.design+" | "+allres.engine
- # panel-scope FDR within each design x engine
- allres['FDR_panel']=np.nan
- for k,idx in allres.groupby('key').groups.items():
- sub=allres.loc[idx]
- p=sub[sub.gene.isin(PANEL)]
- if len(p): allres.loc[p.index,'FDR_panel']=BH(p.PValue.values)
- # ---------------- sheets ----------------
- sheets={}
- # 1) design summary
- rows=[]
- for k,s in allres.groupby('key'):
- p=s[s.gene.isin(PANEL)]
- rows.append(dict(Aggregation=s.aggregation.iloc[0], Design=s.design.iloc[0], Engine=s.engine.iloc[0],
- Cell_types=s.cell_type.nunique(), Tests=len(s),
- Genes_FDRglobal_lt005=int((s.FDR_global<0.05).sum()),
- Genes_FDRlocal_lt005=int((s.FDR_local<0.05).sum()),
- Panel_tests=len(p), Panel_FDR_lt005=int((p.FDR_panel<0.05).sum()),
- MS_donors=int(s.n_ms.max()), HC_donors=int(s.n_hc.max())))
- sheets['01_Design_summary']=pd.DataFrame(rows).sort_values(['Aggregation','Design','Engine'])
- # 2) tier-1 headline: best result per gene per aggregation
- rows=[]
- for g in T1+SUG+['MOSPD3','IFI44L','MX1','DGKQ']:
- for agg,s in allres[allres.gene==g].groupby('aggregation'):
- if not len(s): continue
- # select the row that determines significance: smallest panel-scope FDR (tie-break on p)
- s=s.sort_values(['FDR_panel','PValue'], na_position='last')
- r=s.iloc[0]
- invalid = agg.endswith('INVALID')
- rows.append(dict(Tier=tier(g), Gene=g, Aggregation=agg,
- Scheme_valid="NO - do not cite" if invalid else "yes",
- Best_cell_type=r.cell_type,
- Design=r.design, Engine=r.engine, logFC=r.logFC, P_value=r.PValue,
- FDR_local=r.FDR_local, FDR_global=r.FDR_global, FDR_panel=r.FDR_panel,
- MS_donors=r.n_ms, HC_donors=r.n_hc,
- Significant_panel=("n/a (invalid scheme)" if invalid
- else ("YES" if r.FDR_panel<0.05 else "no"))))
- sheets['02_Candidates_best']=pd.DataFrame(rows).sort_values(['Scheme_valid','Tier','Gene'],
- ascending=[True,True,True])
- # 3) aggregation comparison (panel FDR side by side)
- piv=allres[allres.gene.isin(PANEL)].groupby(['gene','aggregation']).FDR_panel.min().unstack()
- pv =allres[allres.gene.isin(PANEL)].groupby(['gene','aggregation']).PValue.min().unstack()
- cmp=pd.DataFrame({'Gene':piv.index,'Tier':[tier(g) for g in piv.index]})
- for c in piv.columns:
- cmp[f'minP_{c}']=pv[c].values; cmp[f'panelFDR_{c}']=piv[c].values
- sheets['03_Aggregation_comparison']=cmp.sort_values(['Tier','Gene'])
- # 4) full candidate results (valid schemes only)
- full=allres[allres.gene.isin(PANEL) & (allres.aggregation!='S3_sum_CP10K_INVALID')].copy()
- full['Tier']=full.gene.map(tier)
- sheets['04_Candidates_full_valid']=full[['Tier','gene','cell_type','aggregation','design','engine',
- 'logFC','PValue','FDR_local','FDR_global','FDR_panel','n_ms','n_hc','n_pairs']]\
- .sort_values(['Tier','gene','PValue']).rename(columns={'gene':'Gene','cell_type':'Cell_type'})
- # 5) S3 artifact demonstration
- s3=allres[(allres.aggregation=='S3_sum_CP10K_INVALID')&(allres.gene.isin(PANEL))]
- art=s3[s3.design=='S3_sumCP10K_coarse8'].sort_values('PValue').head(40)
- sheets['05_S3_artifact_demo']=art[['gene','cell_type','logFC','PValue','FDR_local','FDR_global','FDR_panel']]\
- .rename(columns={'gene':'Gene','cell_type':'Cell_type'})
- # 6) differential abundance
- das=[]
- for tag in ['coarse','fine']:
- p=f"{D}/DA_{tag}.csv"
- if os.path.exists(p):
- d=pd.read_csv(p); d.insert(0,'Resolution',tag); das.append(d)
- if das:
- da=pd.concat(das,ignore_index=True)
- keep=[c for c in ['Resolution','cell_type','logFC','AveExpr','t','P.Value','adj.P.Val'] if c in da.columns]
- sheets['06_Differential_abundance']=da[keep]
- # 7) variance components
- M=pd.read_csv(f"{D}/PB_lognorm_matrix.csv",index_col=0)
- C=pd.read_csv(f"{D}/PB_lognorm_coldata.csv"); C=C[C['sample'].isin(M.columns)]
- rows=[]
- for g in T1+SUG:
- if g not in M.index: continue
- for ct in ['t_cells','monocytes','b_cells','nk_cells']:
- s=C[C.cell_type==ct]
- if not len(s): continue
- v=M.loc[g,s['sample'].values].astype(float).values; nc=s.n_cells.values
- dsd=float(np.std(v,ddof=1)); cse=float(np.mean(np.sqrt(max(v.mean(),1e-9)/nc)))
- rows.append(dict(Gene=g,Cell_type=ct,Donors=len(v),Mean_cells_per_donor=int(nc.mean()),
- Between_donor_SD=dsd,Cell_sampling_SE=cse,Ratio_SE_over_SD=cse/max(dsd,1e-9)))
- sheets['07_Variance_components']=pd.DataFrame(rows)
- # 8) cell-level vs donor-level
- p=f"{F}/scrna_PSEUDOBULK_COMPARISON.tsv"
- if os.path.exists(p):
- cd=pd.read_csv(p,sep='\t')
- cd=cd[cd.gene.isin(PANEL)].copy(); cd.insert(0,'Tier',cd.gene.map(tier))
- sheets['08_CellLevel_vs_Donor']=cd.sort_values(['Tier','cell_wilcox_FDR_OLD'])
- # 9) top genome-wide hits (pipeline sensitivity evidence)
- tops=[]
- for k,s in allres[allres.aggregation=='S1_sum_raw_counts'].groupby('key'):
- t=s[s.FDR_global<0.05].nsmallest(15,'FDR_global')
- if len(t): tops.append(t.assign(Source=k))
- if tops:
- tp=pd.concat(tops,ignore_index=True)
- sheets['09_TopGenomewide_S1']=tp[['Source','cell_type','gene','biotype','logFC','PValue','FDR_global']]\
- .rename(columns={'gene':'Gene','cell_type':'Cell_type'})
- # 10) biotype summary
- sheets['10_Gene_biotypes']=bio.value_counts().rename_axis('Biotype').reset_index(name='N_genes')
- # ---------------- write ----------------
- wb=Workbook(); wb.remove(wb.active)
- HDR=PatternFill('solid',start_color='1F4E78'); HF=Font(name='Arial',bold=True,color='FFFFFF',size=10)
- BF=Font(name='Arial',size=10); TF=Font(name='Arial',bold=True,size=12)
- SIG=PatternFill('solid',start_color='C6EFCE'); WARN=PatternFill('solid',start_color='FFC7CE')
- thin=Side(style='thin',color='BFBFBF'); BRD=Border(bottom=thin)
- # README
- ws=wb.create_sheet('00_README')
- readme=[
- ("Donor-level pseudobulk differential-state analysis - Kaufmann PBMC (GSE144744)",'title'),
- ("",''),
- ("Cohort: 31 MS / 31 HC donors, 497,705 cells, matched MS-HC pairs multiplexed per 10x run (31 pairs).",''),
- ("Design: paired ~batch_pair + condition. Pairs are sex-matched 31/31; mean age 38.48 (HC) vs 38.52 (MS),",''),
- ("so the pair term absorbs sex and age - no extra covariates needed.",''),
- ("",''),
- ("Raw counts were recovered exactly from the deposited log-normalised matrix:",''),
- (" count = expm1(lognorm) * nCount_RNA / 1e4 -- verified 100% integer; per-cell expm1 sums = 10,000;",''),
- (" reconstructed per-cell totals equal the recorded nCount_RNA exactly.",''),
- ("",''),
- ("AGGREGATION SCHEMES",'h'),
- ("S1_sum_raw_counts SUM of raw integer UMIs per donor x cell type -> edgeR-QL and limma-voom (muscat standard)",''),
- ("S2_mean_CP10K MEAN of per-cell CP10K, log2 -> limma-trend (equal weight per cell)",''),
- ("S3_sum_CP10K_INVALID SUM of per-cell CP10K - REJECTED, see sheet 05: cell number re-enters as an",''),
- (" uncontrolled covariate (21% of genes reach FDR<0.05; effects track cell proportion)",''),
- ("",''),
- ("RESOLUTIONS: 8 coarse lineages / 25 published subsets / whole-PBMC (one sample per donor)",''),
- ("FILTERS: min 10 cells per donor x cell type; filterByExpr or >=50% non-zero; protein-coding only (mygene.info)",''),
- ("FDR SCOPES reported separately: local (within cell type), global (all gene x cell type), panel (26 candidates)",''),
- ("",''),
- ("KEY RESULTS",'h'),
- ("IKZF1 panel FDR 0.047 (whole-PBMC, S2 mean-CP10K), logFC -0.078; direction concordant with the bulk RNA",''),
- (" down-call. Not significant under S1 (0.405). Most defensible candidate signal.",''),
- ("HLA-E panel FDR 0.047 (cDC, S1 voom), logFC +0.212 - direction DISCORDANT with the bulk down-call.",''),
- ("MOSPD3 panel FDR 0.043 (S1) / 0.047 (S2), whole-PBMC - but not one of the 17 named candidates.",''),
- ("ITGB2 null under every scheme (panel FDR 0.42-0.44). Its tier-1 anchoring rests on bulk RNA (q=3.4e-4),",''),
- (" promoter methylation (mCSEA q=0.016) and UK Biobank plasma proteomics (q=1.5e-17), not single-cell.",''),
- ("No candidate is robust across both valid aggregation schemes.",''),
- ("",''),
- ("WHY MORE POWER IS NOT AVAILABLE (sheet 07)",'h'),
- ("Cell-sampling SE is only 8-34% of between-donor SD, so the contrast is donor-limited: effective n = 62",''),
- ("donors, not 497,705 cells. Mixed models / kNN / cell-selection act on the cell dimension and cannot help.",''),
- ("Cell-type composition does not differ (sheet 06, all FDR > 0.19), so composition shift is not an explanation.",''),
- ("",''),
- ("SHEETS",'h'),
- ("01 Design summary | 02 Candidate best results | 03 Aggregation comparison | 04 Full candidate results",''),
- ("05 S3 artifact demo | 06 Differential abundance | 07 Variance components | 08 Cell- vs donor-level",''),
- ("09 Top genome-wide hits | 10 Gene biotypes",''),
- ]
- for i,(txt,kind) in enumerate(readme,1):
- c=ws.cell(row=i,column=1,value=txt)
- c.font=TF if kind=='title' else (Font(name='Arial',bold=True,size=10) if kind=='h' else BF)
- ws.column_dimensions['A'].width=120
- for name,df in sheets.items():
- ws=wb.create_sheet(name[:31])
- for r in dataframe_to_rows(df,index=False,header=True): ws.append(r)
- for c in ws[1]: c.fill=HDR; c.font=HF; c.alignment=Alignment(horizontal='center',vertical='center',wrap_text=True)
- ws.freeze_panes='A2'
- hdr=[c.value for c in ws[1]]
- for j,h in enumerate(hdr,1):
- L=get_column_letter(j)
- ws.column_dimensions[L].width=min(max(12,len(str(h))+3),30)
- for i in range(2,ws.max_row+1):
- cell=ws.cell(row=i,column=j); cell.font=BF; cell.border=BRD
- v=cell.value
- if isinstance(v,float):
- if h and ('P_value' in str(h) or 'PValue' in str(h) or str(h).startswith('minP') or 'P.Value' in str(h)):
- cell.number_format='0.00E+00'
- elif h and ('FDR' in str(h) or 'adj.P' in str(h)):
- cell.number_format='0.0000'
- if v<0.05: cell.fill=SIG
- elif h and ('logFC' in str(h) or 'SD' in str(h) or 'SE' in str(h) or 'Ratio' in str(h)):
- cell.number_format='0.0000'
- else: cell.number_format='0.000'
- if h=='Significant_panel' and v=='YES': cell.fill=SIG
- if h=='Scheme_valid' and isinstance(v,str) and v.startswith('NO'): cell.fill=WARN
- # grey out / flag rows from the invalid S3 scheme so they cannot be misread as findings
- if 'Aggregation' in hdr:
- ja=hdr.index('Aggregation')+1
- for i in range(2,ws.max_row+1):
- if str(ws.cell(row=i,column=ja).value).endswith('INVALID'):
- for j in range(1,len(hdr)+1):
- ws.cell(row=i,column=j).fill=WARN
- ws.cell(row=i,column=j).font=Font(name='Arial',size=10,italic=True,color='9C0006')
- if name=='05_S3_artifact_demo':
- for i in range(2,ws.max_row+1):
- for j in range(1,len(hdr)+1): ws.cell(row=i,column=j).fill=WARN
- wb.save(OUT)
- print(f"\nwrote {OUT}")
- for n,d in sheets.items(): print(f" {n:32s} {len(d):6d} rows x {len(d.columns)} cols")
build_pseudobulk_workbook.py at commit 056de0d, under MIT · at the source
Overview
- Department of Biostatistics and Bioinformatics, Graduate School of Health Sciences, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
- Department of Molecular Biology and Genetics, Faculty of Engineering and Natural Sciences, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
- Department of Molecular Biology and Genetics, Graduate School of Natural and Applied Science, Acibadem Mehmet Ali Aydinlar University, 34638 Istanbul, Turkey
Abstract
Multiple sclerosis (MS) is an immune-mediated central nervous system disease with mapped genetic risk loci but unresolved cross-layer regulatory mechanisms. We harmonised 30 primary public MS datasets spanning bulk transcriptomics (552 samples, 367 MS/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.
alperbulbul1/MS_Omics_Data_Analysis
056de0d6f94fa3cc3044ca5b0f33db3a8349c3c5, 6 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
104 files
- configure.sh, Shell, 81 lines
- scripts/
00_data/ , Python, 87 linesdownload_bulk_rnaseq.py - scripts/
00_data/ , Python, 172 linesdownload_methylation.py - scripts/
00_data/ , Python, 99 linesselect_methylation_case_ control.py - scripts/
01_transcriptome/ , R, 48 lines00_run_all.R - scripts/
01_transcriptome/ , R, 57 lines01_pbmc_de.R - scripts/
01_transcriptome/ , R, 57 lines02_tcells_de.R - scripts/
01_transcriptome/ , R, 56 lines03_bcells_de.R - scripts/
01_transcriptome/ , R, 57 lines04_brainwm_de.R - scripts/
01_transcriptome/ , R, 56 lines05_whole_blood_de.R - scripts/
01_transcriptome/ , R, 57 lines06_pbmc_ifnb_de.R - scripts/
01_transcriptome/ , R, 86 lines07_total_combined_de.R - scripts/
01_transcriptome/ , R, 96 lines08_cross_stratum_master. R - scripts/
01_transcriptome/ , Python, 205 linesbuild_global_matrices.py - scripts/
01_transcriptome/ , Python, 241 linesbuild_matrices.py - scripts/
01_transcriptome/ , Python, 229 linesbuild_rna_sex_persample. py - scripts/
01_transcriptome/ , Python, 97 linesbuild_supplementary_tabl e_S3.py - scripts/
01_transcriptome/ , Python, 374 linescorrect_and_normalize.py - scripts/
01_transcriptome/ , Python, 102 lines, 1 matchharmonize_microarray_v2. py - scripts/
01_transcriptome/ , Python, 145 linesharmonize_rnaseq_v3.py - scripts/
01_transcriptome/ , R, 161 lineshelpers.R - scripts/
01_transcriptome/ , Python, 98 linesinfer_sex_from_expressio n.py - scripts/
01_transcriptome/ , Python, 77 linesmerge_and_correct.py - scripts/
01_transcriptome/ , Python, 40 linesmerge_harmonized_v2.py - scripts/
01_transcriptome/ , Python, 100 linesrerun_verified_case_cont rol.py - scripts/
01_transcriptome/ , R, 107 linesrun_expression_subgroup_ limma.R - scripts/
01_transcriptome/ , Python, 411 linesrun_stratified_omics.py - scripts/
01_transcriptome/ , R, 32 linessex_adjusted_sensitivity _rna.R - scripts/
02_methylation/ , R, 43 lines00_run_all.R - scripts/
02_methylation/ , R, 105 lines01_tcells_meth_dmp.R - scripts/
02_methylation/ , R, 105 lines02_wb_dmf_meth_dmp.R - scripts/
02_methylation/ , R, 104 lines03_wb_ocrelizumab_meth_d mp.R - scripts/
02_methylation/ , R, 105 lines04_tcells_remission_meth _dmp.R - scripts/
02_methylation/ , R, 94 lines05_combined_meth_dmp.R - scripts/
02_methylation/ , R, 91 lines06_mcsea_promoter_analys is.R - scripts/
02_methylation/ , R, 83 lines07_brainwm_rna_vs_meth.R - scripts/
02_methylation/ , R, 77 lines08_cross_stratum_meth_ma ster.R - scripts/
02_methylation/ , R, 257 lines, 2 matches09_inverse_concordance_s can.R - scripts/
02_methylation/ , R, 157 lines10_inverse_proteomics_va lidation.R - scripts/
02_methylation/ , Python, 239 lines11_inverse_scRNA_validat ion.py - scripts/
02_methylation/ , Python, 326 lines12_celltype_4layer_maste r.py - scripts/
02_methylation/ , Python, 342 lines13_perstudy_scRNA_valida tion.py - scripts/
02_methylation/ , Python, 107 lines, 2 matches14_wgbs_GSE173787_promot er.py - scripts/
02_methylation/ , R, 214 lines15_genelevel_weighting_c orrected.R - scripts/
02_methylation/ , Python, 107 linesbuild_meth_sex.py - scripts/
02_methylation/ , Python, 536 linesbuild_methylation_matrix .py - scripts/
02_methylation/ , R, 259 lines, 1 matchhelpers.R - scripts/
02_methylation/ , R, 16 linesinfer_sex_GSE88824_chrXY .R - scripts/
02_methylation/ , R, 368 lines, 1 matchnormalize_beta_only.R - scripts/
02_methylation/ , R, 436 linespreprocess_methylation_a rrays.R - scripts/
02_methylation/ , R, 101 linespromoter_vs_body_test.R - scripts/
02_methylation/ , Python, 189 lines, 2 matchesrebuild_S2_94genes.py - scripts/
02_methylation/ , R, 84 linesrun_all_methylation_comb at.R - scripts/
02_methylation/ , R, 45 linesrun_mcsea_combat.R - scripts/
02_methylation/ , R, 535 linesrun_methylation_subgroup _limma.R - scripts/
02_methylation/ , R, 94 linessex_adjusted_sensitivity .R - scripts/
02_methylation/ , Python, 100 lines, 2 matchesupdate_S2_promoter_flag. py - scripts/
02_methylation/ , Python, 92 linesupdate_S2_weighting.py - scripts/
03_proteomics/ , R, 70 lines00_run_all.R - scripts/
03_proteomics/ , R, 192 lines, 1 match01cc_csf_astral_complete case.R - scripts/
03_proteomics/ , R, 178 lines, 1 match02cc_csf_timstof_complet ecase.R - scripts/
03_proteomics/ , R, 176 lines03_csf_cross_platform_me ta.R - scripts/
03_proteomics/ , R, 117 lines04cc_magliozzi_brain_com pletecase.R - scripts/
03_proteomics/ , R, 187 lines, 1 match05_t_lineage_meta.R - scripts/
03_proteomics/ , R, 83 lines06_pegram_gse32915_de.R - scripts/
03_proteomics/ , R, 141 lines07_brainwm_rna_meth_reru n.R - scripts/
03_proteomics/ , R, 107 lines08_per_group_consistency .R - scripts/
03_proteomics/ , R, 105 lines09_cross_assay_lxn.R - scripts/
03_proteomics/ , R, 86 lines10_master_validation.R - scripts/
03_proteomics/ , R, 174 lines, 1 match11_itgb2_csf_pleocytosis .R - scripts/
03_proteomics/ , Python, 54 linesbuild_RDEP_CC_adapters.p y - scripts/
03_proteomics/ , R, 133 linesdep_bh_equivalence_check .R - scripts/
03_proteomics/ , R, 226 lineshelpers.R - scripts/
04_singlecell/ , Python, 79 linesbuild_pb_brain_csf.py - scripts/
04_singlecell/ , Python, 93 linesbuild_pb_counts.py - scripts/
04_singlecell/ , Python, 79 linesbuild_pb_lognorm.py - scripts/
04_singlecell/ , Python, 81 linesbuild_pb_norm.py - scripts/
04_singlecell/ , Python, 176 lines, 2 matchescell_level_wilcoxon_scan .py - scripts/
04_singlecell/ , Python, 232 linesdownload_singlecell_data sets.py - scripts/
04_singlecell/ , Python, 133 linesplot_GSE144744_kaufmann_ celltypes.py - scripts/
04_singlecell/ , Python, 294 lines, 1 matchprocess_GSE118257_jakel_ brain.py - scripts/
04_singlecell/ , Python, 277 linesprocess_GSE127969_beltra n_csf.py - scripts/
04_singlecell/ , Python, 181 linesprocess_GSE144744_kaufma nn_genes.py - scripts/
04_singlecell/ , Python, 203 linesprocess_GSE144744_kaufma nn_pbmc.py - scripts/
04_singlecell/ , R, 40 linespseudobulk_DA.R - scripts/
04_singlecell/ , R, 97 lines, 1 matchpseudobulk_brain_csf.R - scripts/
04_singlecell/ , R, 100 lines, 1 matchpseudobulk_muscat_style. R - scripts/
04_singlecell/ , R, 84 linespseudobulk_norm_compare. R - scripts/
04_singlecell/ , Python, 220 linesrun_pseudobulk_reanalysi s.py - scripts/
05_integration/ , Python, 192 lines, 1 matchbuild_data_sources_table .py - scripts/
05_integration/ , Python, 98 lines, 2 matchesdisease_catalogue_ora.py - scripts/
05_integration/ , Python, 654 linesexport_used_omics_invent ory.py - scripts/
05_integration/ , Python, 317 linesppi_analysis.py - scripts/
06_figures/ , Python, 250 lines, 3 matchesbuild_pseudobulk_workboo k.py - scripts/
06_figures/ , Python, 70 linesfigure1_workflow.py - scripts/
06_figures/ , Python, 206 lines, 2 matchesfigure2_rna_volcanoes.py - scripts/
06_figures/ , Python, 250 lines, 2 matchesfigure3_methylation.py - scripts/
06_figures/ , Python, 244 lines, 2 matchesfigure4_proteomics.py - scripts/
06_figures/ , Python, 310 lines, 2 matchesfigure5_singlecell.py - scripts/
06_figures/ , Python, 284 lines, 2 matchesfigure6_intersection_hea tmap.py - scripts/
06_figures/ , Python, 131 lines, 2 matchesfigure7_string_network.p y - scripts/
06_figures/ , Python, 108 lines, 2 matchesfigure_constants.py - LICENSE, License, 18 lines
- README.md, Text, 140 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 102 scripts, each with its path and the digest of its content;
- 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE173787, at NCBI GEO; found in the text, “4.1. Data Sources and Cohort Assembly”
Data Availability Statement
All primary public accessions used in this study are catalogued per-cohort in Supplementary Table S1. NCBI GEO: Bulk transcriptomics (15 series): GSE21942 (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 6 keywords, 7 MeSH terms, 110 references.
Cite
This paper
Bülbül, A., Yılmaz-İşgördü, Ö. H., Reda, M. D., & Turanlı, E. T. (2026). Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling. International journal of molecular sciences, 27(16), 7275. https://
BibTeX
@article{bulbul2026integ
author = {Bülbül, Alper and Yılmaz-İşgördü, Özdeyiş Hülya and Reda, Meziyet Dilara and Turanlı, Eda Tahir},
title = {{Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling}},
journal = {International journal of molecular sciences},
year = {2026},
month = aug,
volume = {27},
number = {16},
pages = {7275},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1422-0067},
doi = {10.3390/
url = {https://
pmid = {42653279},
pmcid = {PMC13513955}
}
RIS
TY - JOUR
AU - Bülbül, Alper
AU - Yılmaz-İşgördü, Özdeyiş Hülya
AU - Reda, Meziyet Dilara
AU - Turanlı, Eda Tahir
TI - Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling
T2 - International journal of molecular sciences
J2 - Int J Mol Sci
PY - 2026
DA - 2026/
VL - 27
IS - 16
SP - 7275
SN - 1422-0067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "Integrative Multi-Omics Analysis of Multiple Sclerosis Reveals Cell-Type-Specific Regulatory Landscapes and Discordant Methylation-Expression Coupling",
"container-title": "International journal of molecular sciences",
"author": [
{
"family": "Bülbül",
"given": "Alper"
},
{
"family": "Yılmaz-İşgördü",
"given": "Özdeyiş Hülya"
},
{
"family": "Reda",
"given": "Meziyet Dilara"
},
{
"family": "Turanlı",
"given": "Eda Tahir"
}
],
"container-title-short":
"volume": "27",
"issue": "16",
"page": "7275",
"DOI": "10.3390/
"PMID": "42653279",
"PMCID": "PMC13513955",
"ISSN": "1422-0067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
14
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: edgeR, limma, anndata, 14 other tools, genetics / omics, cellular / molecular, 1 reference
- [2] doi:10.1002/imt2.70163 [code]
- Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.Journal: iMetaIn common: edgeR, limma, anndata, 13 other tools, genetics / omics, cellular / molecular
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: edgeR, limma, pheatmap, 13 other tools, cellular / molecular
- [4] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: edgeR, limma, pheatmap, 11 other tools, genetics / omics, 2 references
- [5] doi:10.1038/s41592-026-03194-8 [code]
- Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.Journal: Nature methodsIn common: limma, anndata, Scanpy, 12 other tools, genetics / omics, 1 reference
- [6] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: edgeR, anndata, Scanpy, 11 other tools, genetics / omics, cellular / molecular, 2 references
- [7] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: edgeR, anndata, Scanpy, 11 other tools, genetics / omics, cellular / molecular, 1 reference
- [8] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: edgeR, limma, anndata, 12 other tools, genetics / omics
- [9] doi:10.1038/s41467-026-71803-3 [code]
- Charting the transition from in vitro gliogenesis to the in vivo maturation of human glial progenitor cells transplanted into the hypomyelinated mouse brain.Journal: Nature communicationsIn common: anndata, Scanpy, pheatmap, 9 other tools, multiple sclerosis, genetics / omics, cellular / molecular, 3 references
- [10] doi:10.1016/j.cpblue.2026.100007 [code]
- An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.Journal: Cell press blueIn common: edgeR, limma, anndata, 12 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 102 scripts, and 40 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:aa2cfe1176603242…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
