Machine-learning assisted subclassification of glioblastoma by developing an endoplasmic reticulum stress-related methylation signature.
The 4 matches
- [1] § Material and methods › GBM methylation data collection ↔ Preprocessing/config.R, lines 68–119 · score 0.86 · Raw IDAT, batch corrected, EPIC, FFPE, ambiguous, frozen
- [2] § Material and methods › GBM methylation data collection ↔ Preprocessing/main.R, lines 145–184 · score 0.75 · removeBatchEffect, batch corrected, limma, log2, unmethylated, preprocessing
- [3] § Material and methods › Feature selection ↔ RFECV/main.py, lines 94–174 · score 0.71 · Random Forest, XGBoost, feature selection, RFECV, voting, classification
- [4] § Material and methods › Feature selection ↔ RFECV/main.py, lines 94–174 · score 0.68 · multiple iterations, feature selection, RFECV, fold, training, stratified
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 373 lines · 14 KB · no license · 2 matches
- #!/usr/bin/env python3
- # -*- coding: utf-8 -*-
- """
- ===============================================================================
- Project: GBM Gene Feature Selection via Ensemble RFE Voting
- ===============================================================================
- """
- import logging
- import json
- import pickle
- from pathlib import Path
- from typing import Dict, List, Optional
- from datetime import datetime
- import pandas as pd
- import numpy as np
- from sklearn.ensemble import RandomForestClassifier
- from sklearn.svm import SVC
- from sklearn.feature_selection import RFECV
- from sklearn.model_selection import StratifiedKFold
- from sklearn.base import clone
- import xgboost as xgb
- # :: Configure Logging
- log_timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
- logging.basicConfig(
- level=logging.INFO,
- format='%(asctime)s - %(levelname)s - %(message)s',
- handlers=[
- logging.FileHandler(f"pipeline_{log_timestamp}.log"),
- logging.StreamHandler()
- ]
- )
- logger = logging.getLogger(__name__)
- class ConfigManager:
- """Handles loading and validating configuration settings."""
- def __init__(self, config_path: str = "config.json"):
- self.script_dir = Path(__file__).parent.resolve()
- self.config_path = self.script_dir / config_path
- self.config = self._load_config()
- self._resolve_paths()
- def _load_config(self) -> Dict:
- if not self.config_path.exists():
- logger.warning(f"Config file not found at {self.config_path}. Creating default...")
- self._create_default_config()
- with open(self.config_path, 'r', encoding='utf-8') as f:
- return json.load(f)
- def _create_default_config(self):
- default = {
- "paths": {
- "input_data": "data/GBM_NMF_group4.csv",
- "cache_dir": "cache",
- "output_dir": "results",
- "output_prefix": "GBM_group4"
- },
- "params": {
- "iteration_list": [10, 20, 30, 40, 50],
- "cv_splits": 5,
- "scoring": "accuracy",
- "vote_threshold": 1.0
- },
- "models": {
- "rf": {"type": "RandomForest", "random_state": 42},
- "svm": {"type": "SVC", "kernel": "linear", "random_state": 42},
- "xgb": {"type": "XGB", "random_state": 42, "use_label_encoder": False}
- }
- }
- with open(self.config_path, 'w', encoding='utf-8') as f:
- json.dump(default, f, indent=4)
- logger.info(f"Default config created at {self.config_path}")
- def _resolve_paths(self):
- """Converts relative paths in config to absolute paths based on script location."""
- for key, value in self.config['paths'].items():
- if isinstance(value, str):
- abs_path = self.script_dir / value
- self.config['paths'][key] = abs_path
- def get(self, *keys):
- """Nested access to config values."""
- val = self.config
- for k in keys:
- val = val[k]
- return val
- class EnsembleFeatureSelector:
- """
- Performs ensemble feature selection using RFE voting across multiple models.
- Supports multiple iteration counts for stability analysis.
- """
- def __init__(self, pd.DataFrame, config: ConfigManager, n_iterations: int):
- self.data = data
- self.config = config
- self.n_iterations = n_iterations
- self.X = None
- self.y = None
- self.voting_matrix = None
- self.feature_names = None
- def prepare_data(self):
- """Separate features and target labels."""
- if 'cluster' not in self.data.columns:
- raise ValueError("Target column 'cluster' not found in data.")
- self.X = self.data.drop(columns=['cluster'])
- self.y = self.data['cluster']
- self.feature_names = self.X.columns.tolist()
- logger.info(f"Data prepared. Features: {self.X.shape[1]}, Samples: {self.X.shape[0]}")
- def _get_estimator(self, model_name: str):
- """Factory method to create model instances based on config."""
- model_cfg = self.config.get('models', model_name)
- model_type = model_cfg.get('type')
- params = {k: v for k, v in model_cfg.items() if k != 'type'}
- if model_type == 'RandomForest':
- return RandomForestClassifier(**params)
- elif model_type == 'SVC':
- return SVC(**params)
- elif model_type == 'XGB':
- return xgb.XGBClassifier(**params)
- else:
- raise ValueError(f"Unknown model type: {model_type}")
- def _run_single_model_rfe(self, model_name: str) -> np.ndarray:
- """
- Run iterative RFECV for a single model type.
- Returns a matrix of shape (n_iterations, n_features) with rankings.
- """
- logger.info(f"Starting {self.n_iterations} RFE iterations for model: {model_name}")
- rankings = []
- cv_splits = self.config.get('params', 'cv_splits')
- scoring = self.config.get('params', 'scoring')
- cv = StratifiedKFold(n_splits=cv_splits, shuffle=True)
- for i in range(1, self.n_iterations + 1):
- estimator = self._get_estimator(model_name)
- if hasattr(estimator, 'random_state'):
- estimator.set_params(random_state=i)
- rfecv = RFECV(
- estimator=estimator,
- cv=cv,
- scoring=scoring,
- n_jobs=-1
- )
- y_train = self.y
- if model_name == 'xgb':
- y_train = self.y - 1
- try:
- rfecv.fit(self.X, y_train)
- rankings.append(rfecv.ranking_)
- except Exception as e:
- logger.warning(f"Iteration {i} failed for {model_name}: {e}")
- rankings.append(np.ones(len(self.feature_names)) * 2)
- return np.array(rankings)
- def run_ensemble_selection(self) -> pd.DataFrame:
- """Execute feature selection across all configured models."""
- cache_dir = Path(self.config.get('paths', 'cache_dir'))
- cache_dir.mkdir(parents=True, exist_ok=True)
- # Cache file name includes iteration count for differentiation
- cache_file = cache_dir / f"rfe_cache_iter{self.n_iterations}.pkl"
- # Try loading from cache
- if cache_file.exists():
- try:
- logger.info(f"Loading cached results from {cache_file}")
- with open(cache_file, 'rb') as f:
- cache_data = pickle.load(f)
- self.voting_matrix = cache_data['voting_matrix']
- self.feature_names = cache_data['feature_names']
- return self._process_votes()
- except Exception as e:
- logger.warning(f"Cache loading failed: {e}. Recomputing...")
- # Compute fresh results
- if self.X is None:
- self.prepare_data()
- all_rankings = []
- model_names = list(self.config.config['models'].keys())
- for name in model_names:
- rankings = self._run_single_model_rfe(name)
- all_rankings.append(rankings)
- self.voting_matrix = np.vstack(all_rankings)
- logger.info(f"Total voting matrix shape: {self.voting_matrix.shape}")
- # Save to cache
- try:
- with open(cache_file, 'wb') as f:
- pickle.dump({
- 'voting_matrix': self.voting_matrix,
- 'feature_names': self.feature_names
- }, f)
- logger.info(f"Results cached to {cache_file}")
- except Exception as e:
- logger.error(f"Failed to save cache: {e}")
- return self._process_votes()
- def _process_votes(self) -> pd.DataFrame:
- """Calculate vote counts and filter features based on threshold."""
- vote_counts = np.sum(self.voting_matrix == 1, axis=0)
- max_possible_votes = self.voting_matrix.shape[0]
- threshold_ratio = self.config.get('params', 'vote_threshold')
- threshold_count = int(max_possible_votes * threshold_ratio)
- logger.info(f"Total votes possible: {max_possible_votes}, Threshold: {threshold_count}")
- selected_indices = np.where(vote_counts >= threshold_count)[0]
- selected_features = [self.feature_names[i] for i in selected_indices]
- logger.info(f"Selected {len(selected_features)} features out of {len(self.feature_names)}")
- return self.data[selected_features + ['cluster']]
- def save_results(self, result_df: pd.DataFrame):
- """Save the final selected features and the list of names."""
- output_dir = Path(self.config.get('paths', 'output_dir'))
- output_dir.mkdir(parents=True, exist_ok=True)
- prefix = self.config.get('paths', 'output_prefix')
- threshold = int(self.config.get('params', 'vote_threshold') * 100)
- # File names include iteration count for differentiation
- file_name = f"{prefix}_iter{self.n_iterations}_selected_{threshold}pct.csv"
- result_path = output_dir / file_name
- result_df.to_csv(result_path, index=True)
- logger.info(f"Dataset saved to {result_path}")
- probe_list = [col for col in result_df.columns if col != 'cluster']
- probe_path = output_dir / f"{prefix}_iter{self.n_iterations}_probes_{threshold}pct.csv"
- pd.Series(probe_list).to_csv(probe_path, index=False, header=False)
- logger.info(f"Probe list saved to {probe_path}")
- return {
- 'n_iterations': self.n_iterations,
- 'n_selected': len(probe_list),
- 'n_total_features': len(self.feature_names),
- 'file_path': str(result_path)
- }
- class MultiIterationRunner:
- """
- Orchestrates multiple feature selection runs with different iteration counts.
- Generates summary statistics across all iterations.
- """
- def __init__(self, pd.DataFrame, config: ConfigManager):
- self.data = data
- self.config = config
- self.results_summary = []
- def run_all_iterations(self) -> pd.DataFrame:
- """Run feature selection for all configured iteration counts."""
- iteration_list = self.config.get('params', 'iteration_list')
- logger.info(f"Starting batch run for iterations: {iteration_list}")
- for n_iter in iteration_list:
- logger.info(f"{'='*60}")
- logger.info(f"Running iteration count: {n_iter}")
- logger.info(f"{'='*60}")
- selector = EnsembleFeatureSelector(self.data, self.config, n_iter)
- selected_data = selector.run_ensemble_selection()
- result_info = selector.save_results(selected_data)
- self.results_summary.append(result_info)
- # Optional: Add delay between runs to prevent resource exhaustion
- # time.sleep(1)
- # Generate summary report
- self._generate_summary_report()
- return pd.DataFrame(self.results_summary)
- def _generate_summary_report(self):
- """Generate and save a summary report of all iteration runs."""
- output_dir = Path(self.config.get('paths', 'output_dir'))
- output_dir.mkdir(parents=True, exist_ok=True)
- summary_df = pd.DataFrame(self.results_summary)
- summary_path = output_dir / "iteration_summary.csv"
- summary_df.to_csv(summary_path, index=False)
- logger.info(f"Summary report saved to {summary_path}")
- # Log summary statistics
- logger.info("\n" + "="*60)
- logger.info("ITERATION SUMMARY")
- logger.info("="*60)
- logger.info(summary_df.to_string(index=False))
- logger.info("="*60)
- # Generate stability analysis (optional)
- self._analyze_feature_stability()
- def _analyze_feature_stability(self):
- """
- Analyze how feature selection stabilizes across different iteration counts.
- """
- output_dir = Path(self.config.get('paths', 'output_dir'))
- output_dir.mkdir(parents=True, exist_ok=True)
- stability_data = []
- iteration_list = self.config.get('params', 'iteration_list')
- for result in self.results_summary:
- stability_data.append({
- 'iterations': result['n_iterations'],
- 'selected_features': result['n_selected'],
- 'selection_ratio': result['n_selected'] / result['n_total_features']
- })
- stability_df = pd.DataFrame(stability_data)
- stability_path = output_dir / "feature_stability_analysis.csv"
- stability_df.to_csv(stability_path, index=False)
- logger.info(f"Stability analysis saved to {stability_path}")
- def main():
- """Main execution pipeline."""
- logger.info("=== Starting Feature Selection Pipeline ===")
- logger.info(f"Script running from: {Path(__file__).parent.resolve()}")
- # 1. Load Configuration
- config = ConfigManager()
- # 2. Load Data
- input_path = Path(config.get('paths', 'input_data'))
- if not input_path.exists():
- logger.error(f"Input file not found: {input_path}")
- logger.error("Please ensure your data is placed correctly or update config.json")
- return
- try:
- data = pd.read_csv(input_path, sep=',', index_col=0, header=0)
- logger.info(f"Data loaded: {data.shape}")
- except Exception as e:
- logger.error(f"Failed to load {e}")
- return
- # 3. Initialize Multi-Iteration Runner
- runner = MultiIterationRunner(data, config)
- # 4. Run All Iterations
- summary = runner.run_all_iterations()
- # 5. Final Report
- logger.info("\n=== Pipeline Completed Successfully ===")
- logger.info(f"Total runs completed: {len(summary)}")
- logger.info(f"Results saved to: {config.get('paths', 'output_dir')}")
- if __name__ == "__main__":
- main()
main.py at commit 9e08935, no license · at the source
Overview
Abstract
Introduction: Glioblastoma (GBM) is a highly aggressive brain tumor with significant heterogeneity, leading to poor prognosis and limited treatment options. Developing innovative molecular subtyping approaches is important for gaining deeper insights into disease pathogenesis and optimizing treatment strategies. DNA methylation has been implicated in the regulation of endoplasmic reticulum stress (ERS), which disrupts protein folding and activates the unfolded protein response (UPR), ultimately determining cellular survival or apoptotic outcomes.
Methods: ERS-related DNA methylation profiles were integrated with non-negative matrix factorization (NMF) to establish a molecular classification framework for GBM. An ERS-based signature was further developed using recursive feature elimination with cross-validation (RFECV), and a random forest (RF) model was constructed for subtype prediction. The model was then applied to an external TCGA cohort for validation and downstream characterization.
Results: The NMF-based framework stratified GBM patients into four distinct subtypes. The RF model achieved an accuracy of 92.4% in the independent test set. Application of the model to the TCGA cohort revealed distinct molecular and clinical characteristics across subtypes. In particular, Subtype 2 was associated with an immune-inflamed phenotype, lower tumor purity, and poorer prognosis. Connectivity Map (CMap) analysis further identified MEK inhibitors as preliminary candidate compounds for specific subtypes.
Discussion: These findings support an association between ERS-related epigenetic modifications and GBM heterogeneity, and provide an epigenetic framework for refined molecular stratification and further exploration of subtype-related therapeutic strategies.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.
jyz30102/ERSR-GBM-Subtypes
9e0893566929d5aa6cac48ed19d8a018c8be40bf, 19 March 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
10 files
- NMF/
main.R , R, 103 lines - NMF/
nmf_analysis.R , R, 274 lines - NMF/
utils.R , R, 60 lines - Preprocessing/
config.R , R, 150 lines, 1 match - Preprocessing/
functions.R , R, 126 lines - Preprocessing/
main.R , R, 258 lines, 1 match - RFECV/
main.py , Python, 373 lines, 2 matches - mapping/
gene_to_probe.py , Python, 150 lines - mapping/
main.py , Python, 95 lines - README.md, Text, 20 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 9 scripts, each with its path and the digest of its content;
- 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE90496, at NCBI GEO; found in the text, “GBM methylation data collection”
Data availability statement
The original contributions presented in the study are included in the article/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 4 authors, 6 keywords, 62 references.
Cite
This paper
Zhao, J., Li, M., Pu, X., & Guo, Y. (2026). Machine-learning assisted subclassification of glioblastoma by developing an endoplasmic reticulum stress-related methylation signature. Frontiers in oncology, 16, 1750334. https://
BibTeX
@article{zhao2026machine
author = {Zhao, Jinyi and Li, Menglong and Pu, Xuemei and Guo, Yanzhi},
title = {{Machine-learning assisted subclassification of glioblastoma by developing an endoplasmic reticulum stress-related methylation signature}},
journal = {Frontiers in oncology},
year = {2026},
month = apr,
volume = {16},
pages = {1750334},
publisher = {Frontiers Media SA},
issn = {2234-943X},
doi = {10.3389/
url = {https://
pmid = {42100389},
pmcid = {PMC13143586}
}
RIS
TY - JOUR
AU - Zhao, Jinyi
AU - Li, Menglong
AU - Pu, Xuemei
AU - Guo, Yanzhi
TI - Machine-learning assisted subclassification of glioblastoma by developing an endoplasmic reticulum stress-related methylation signature
T2 - Frontiers in oncology
J2 - Front Oncol
PY - 2026
DA - 2026/
VL - 16
SP - 1750334
SN - 2234-943X
PB - Frontiers Media SA
DO - 10.3389/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3389/
"type": "article-journal",
"title": "Machine-learning assisted subclassification of glioblastoma by developing an endoplasmic reticulum stress-related methylation signature",
"container-title": "Frontiers in oncology",
"author": [
{
"family": "Zhao",
"given": "Jinyi"
},
{
"family": "Li",
"given": "Menglong"
},
{
"family": "Pu",
"given": "Xuemei"
},
{
"family": "Guo",
"given": "Yanzhi"
}
],
"container-title-short":
"volume": "16",
"page": "1750334",
"DOI": "10.3389/
"PMID": "42100389",
"PMCID": "PMC13143586",
"ISSN": "2234-943X",
"publisher": "Frontiers Media SA",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
22
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: limma, ggplot2, tidyverse, 3 other tools, genetics / omics, other condition, 1 reference
- [2] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: XGBoost, limma, ggplot2, 4 other tools
- [3] doi:10.1016/j.stemcr.2026.102977 [code]
- Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics.Journal: Stem cell reportsIn common: limma, scikit-learn, pandas, 1 other tool, genetics / omics, other condition, 2 references
- [4] doi:10.1038/s41398-026-04200-5 [code]
- Postmortem brain single-nucleus and bulk gene expression analyses identify shared and distinct abnormalities in bipolar disorder and major depressive disorder.Journal: Translational psychiatryIn common: XGBoost, limma, ggplot2, 3 other tools, genetics / omics
- [5] doi:10.1016/j.isci.2026.115982 [code]
- Multi-omics profiling-derived signature links cellular ecosystem to glioblastoma prognosis.Journal: iScienceIn common: limma, ggplot2, tidyverse, 2 other tools, genetics / omics, other condition, 1 reference
- [6] doi:10.1186/s13059-026-04125-8 [code]
- MLMarker: a machine learning framework for tissue inference and biomarker discovery.Journal: Genome biologyIn common: XGBoost, limma, tidyverse, 3 other tools, genetics / omics
- [7] doi:10.1038/s41746-026-02735-x [code]
- Multimodal interpretable deep learning for transcriptome-informed precision oncology and drug mechanism analysis.Journal: NPJ digital medicineIn common: scikit-learn, pandas, NumPy, genetics / omics, other condition, 3 references
- [8] doi:10.3390/ijms27136068 [code]
- Loss of Neuropeptide Y Signaling Accompanies the Neural-to-Mesenchymal Transcriptional Transition in Glioblastoma: A Multi-Scale Transcriptomic Analysis.Journal: International journal of molecular sciencesIn common: limma, ggplot2, tidyverse, genetics / omics, other condition, 2 references
- [9] doi:10.1186/s13293-026-00927-4 [code]
- Gene regulatory network analysis identifies dysregulation of hypoxia pathways as contributing to glioblastoma treatment resistance in females.Journal: Biology of sex differencesIn common: limma, ggplot2, tidyverse, 2 other tools, other condition, 1 reference
- [10] doi:10.1007/s00262-026-04390-3 [code]
- Identification and prioritisation of tumour antigen candidates from 79 glioblastoma transcriptomes.Journal: Cancer immunology, immunotherapy : CIIIn common: XGBoost, ggplot2, tidyverse, 3 other tools, genetics / omics, other condition
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 9 scripts, and 4 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:1257a23112becb21…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
