Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection.
The 5 matches
- [1] § 4. Materials and Methods › 4.4. Feature Selection: Pareto-Guided Recursive Neural Network (PRNN) ↔ code.py, lines 74–120 · score 0.79 · hidden layers, feature selection, Adam, ReLU, activation, PNN
- [2] § 4. Materials and Methods › 4.7. Discriminative Ability: ROC and Precision–Recall Analyses ↔ code.py, lines 250–384 · score 0.67 · PR curves, LightGBM, XGBoost, confidence, metrics, recall
- [3] § 4. Materials and Methods › 4.5. Model Development: Ensemble Learning Classifiers ↔ code.py, lines 250–384 · score 0.56 · LightGBM, XGBoost, ablation, preprocessing, boosting, pipeline
- [4] § 2. Results › 2.2. Model Performance ↔ code.py, lines 412–436 · score 0.54 · F1 score, LightGBM, XGBoost, sensitivity, metrics, accuracy
- [5] § 2. Results › 2.2. Model Performance ↔ code.py, lines 412–436 · score 0.53 · Training AUC, LightGBM, XGBoost, gap, overfitting, models
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 716 lines · 26 KB · no license · 5 matches
- # ============================================================
- # BÖLÜM 1: IMPORTS
- # ============================================================
- import numpy as np
- import pandas as pd
- import matplotlib
- matplotlib.use('Agg')
- import matplotlib.pyplot as plt
- import warnings, pickle, math
- from pathlib import Path
- from scipy.stats import mannwhitneyu, fisher_exact
- from sklearn.model_selection import (StratifiedShuffleSplit,
- GridSearchCV, StratifiedKFold)
- from sklearn.preprocessing import StandardScaler, MinMaxScaler, label_binarize
- from sklearn.neural_network import MLPClassifier
- from sklearn.svm import SVC
- from sklearn.metrics import (
- roc_auc_score, accuracy_score, f1_score, recall_score,
- roc_curve, auc, precision_recall_curve, average_precision_score
- )
- from interpret.glassbox import ExplainableBoostingClassifier
- from xgboost import XGBClassifier
- from lightgbm import LGBMClassifier
- warnings.filterwarnings('ignore')
- # ============================================================
- # BÖLÜM 2: VERİ YÜKLEME VE ÖN İŞLEME
- # ============================================================
- def load_data(filepath):
- df = pd.read_excel(filepath)
- y = df['factor'].values
- X = df.drop('factor', axis=1)
- feature_names = X.columns.tolist()
- print(f"Dataset: {X.shape[0]} samples, {X.shape[1]} features")
- print(f"ME/CFS(1): {y.sum()}, Control(0): {(y==0).sum()}")
- return X.values, y, feature_names
- def preprocess(X_train_raw, X_test_raw):
- X_tr = np.log2(X_train_raw + 1)
- X_te = np.log2(X_test_raw + 1)
- scaler = StandardScaler()
- return scaler.fit_transform(X_tr), scaler.transform(X_te)
- # ============================================================
- # BÖLÜM 3: PRNN FEATURE SELECTION
- # ============================================================
- def is_pareto_efficient(costs, return_mask=True):
- is_efficient = np.arange(costs.shape[0])
- n_points = costs.shape[0]
- next_point_index = 0
- while next_point_index < len(costs):
- nondominated = np.any(costs > costs[next_point_index], axis=1)
- nondominated[next_point_index] = True
- is_efficient = is_efficient[nondominated]
- costs = costs[nondominated]
- next_point_index = np.sum(nondominated[:next_point_index]) + 1
- if return_mask:
- mask = np.zeros(n_points, dtype=bool)
- mask[is_efficient] = True
- return mask
- return is_efficient
- def pnn_feature_selection(X_train, y_train, p, feature_names):
- mlp = MLPClassifier(
- hidden_layer_sizes=(64, 32),
- activation='relu',
- solver='adam',
- max_iter=100,
- early_stopping=True,
- n_iter_no_change=10,
- random_state=42
- )
- mlp.fit(X_train, y_train)
- # Ağırlık matrisi çarpımı → (n_features, n_classes)
- W = mlp.coefs_[0]
- for w in mlp.coefs_[1:]:
- W = W @ w
- P_W = np.abs(W)
- P_W_norm = MinMaxScaler().fit_transform(P_W)
- # Objective 1: max normalize skor
- Obj_max = P_W_norm.max(axis=1)
- # Objective 2: 1/rank
- P_W_rank = P_W_norm.copy()
- for j in range(P_W_norm.shape[1]):
- idx = np.argsort(-P_W_norm[:, j])
- P_W_rank[:, j] = np.argsort(idx) + 1
- Obj_rank_max = (1.0 / P_W_rank).max(axis=1)
- # Pareto frontlarına böl
- Obj = np.column_stack([Obj_max, Obj_rank_max])
- genes = list(feature_names)
- remaining = Obj.copy()
- rem_genes = genes.copy()
- fronts = []
- while len(remaining) > 0:
- mask = is_pareto_efficient(remaining)
- fronts.append([g for g, m in zip(rem_genes, mask) if m])
- rem_genes = [g for g, m in zip(rem_genes, mask) if not m]
- remaining = remaining[~mask]
- n_take = max(1, int(len(fronts) * p))
- selected = []
- for i in range(n_take):
- selected.extend(fronts[i])
- return selected
- def prnn_recursive(data_x, data_y, DX, DY, p=0.5,
- best_acc=0, best_gene=None,
- depth=0, max_depth=20):
- if depth >= max_depth:
- return best_gene if best_gene else data_x.columns.tolist()
- n_orig = DX.shape[1]
- n_curr = data_x.shape[1]
- # Dinamik p — Li et al. 2025 formülü (birebir)
- if p < 0.9:
- p = 0.5 + (math.tan((9 * math.pi / 40) * (1 - (n_curr / n_orig)))) / 2
- gene = pnn_feature_selection(
- data_x.values, data_y.values, p, data_x.columns.tolist()
- )
- if not gene:
- return best_gene if best_gene else data_x.columns.tolist()
- # SVM 5-fold CV — iç değerlendirme kriteri
- skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=111)
- accs = []
- for tr, te in skf.split(DX, DY):
- svm = SVC(random_state=42)
- svm.fit(DX.iloc[tr][gene], DY.iloc[tr])
- accs.append(accuracy_score(DY.iloc[te], svm.predict(DX.iloc[te][gene])))
- mean_acc = np.mean(accs)
- # En iyi subset güncelle (200'den az feature koşulu)
- if mean_acc > best_acc and len(gene) < 200:
- best_acc = mean_acc
- best_gene = gene
- # Durma koşulu
- if len(gene) <= 10:
- return best_gene if best_gene else gene
- # Recursive: seçilen gene'lerle devam
- return prnn_recursive(
- data_x[gene], data_y, DX, DY, p,
- best_acc, best_gene, depth + 1, max_depth
- )
- def run_prnn(X_scaled, y, feature_names, p=0.5):
- df_X = pd.DataFrame(X_scaled, columns=feature_names)
- df_y = pd.Series(y)
- selected = prnn_recursive(df_X, df_y, df_X, df_y, p=0.5)
- if not selected:
- print(" [WARN] PRNN seçim yapamadı, tüm feature'lar kullanılıyor")
- return np.arange(len(feature_names))
- name_to_idx = {n: i for i, n in enumerate(feature_names)}
- return np.array([name_to_idx[g] for g in selected if g in name_to_idx])
- # ============================================================
- # BÖLÜM 4: HİPERPARAMETRE GRİDLERİ
- # ============================================================
- EBM_GRID = {
- 'max_rounds': [100, 200, 500],
- 'learning_rate': [0.01, 0.05, 0.10],
- }
- XGB_GRID = {
- 'n_estimators': [100, 200, 500],
- 'max_depth': [3, 5, 7],
- 'learning_rate': [0.01, 0.05, 0.10],
- 'subsample': [0.6, 0.8, 1.0],
- 'reg_alpha': [0, 0.1, 1.0],
- 'reg_lambda': [0, 0.1, 1.0],
- }
- LGB_GRID = {
- 'n_estimators': [100, 200, 500],
- 'max_depth': [3, 5, 7],
- 'learning_rate': [0.01, 0.05, 0.10],
- 'subsample': [0.6, 0.8, 1.0],
- 'reg_alpha': [0, 0.1, 1.0],
- 'reg_lambda': [0, 0.1, 1.0],
- }
- def grid_search(model, param_grid, X_train, y_train):
- """5-fold CV, sadece training partition — leakage yok"""
- gs = GridSearchCV(model, param_grid, cv=5,
- scoring='roc_auc', n_jobs=-1, refit=True)
- gs.fit(X_train, y_train)
- return gs.best_estimator_
- # ============================================================
- # BÖLÜM 5: ANA 50-REPEAT DÖNGÜSÜ
- # ============================================================
- def _store_metrics(store, y_te, y_tr, model, X_te, X_tr,
- mean_fpr, mean_recall):
- """Her iterasyon için tüm metrikleri topla"""
- y_prob = model.predict_proba(X_te)[:, 1]
- y_pred = model.predict(X_te)
- y_prob_tr = model.predict_proba(X_tr)[:, 1]
- store['aucs'].append(roc_auc_score(y_te, y_prob))
- store['accs'].append(accuracy_score(y_te, y_pred))
- store['f1s'].append(f1_score(y_te, y_pred))
- store['sens'].append(recall_score(y_te, y_pred))
- store['aps'].append(average_precision_score(y_te, y_prob))
- store['train_aucs'].append(roc_auc_score(y_tr, y_prob_tr))
- tn = np.sum((y_te == 0) & (y_pred == 0))
- fp = np.sum((y_te == 0) & (y_pred == 1))
- store['specs'].append(tn / (tn + fp) if (tn + fp) > 0 else 0)
- fpr, tpr, _ = roc_curve(y_te, y_prob)
- tpr_i = np.interp(mean_fpr, fpr, tpr); tpr_i[0] = 0.0
- store['tprs'].append(tpr_i)
- prec, rec, _ = precision_recall_curve(y_te, y_prob)
- store['precs'].append(np.interp(mean_recall, rec[::-1], prec[::-1]))
- def run_50repeat(X, y, feature_names, n_iter=50, output_dir='results'):
- Path(output_dir).mkdir(exist_ok=True)
- mean_fpr = np.linspace(0, 1, 300)
- mean_recall = np.linspace(0, 1, 300)
- store = {m: {
- 'aucs':[], 'accs':[], 'f1s':[], 'sens':[], 'specs':[],
- 'aps':[], 'tprs':[], 'precs':[], 'train_aucs':[]
- } for m in ['EBM', 'XGBoost', 'LightGBM', 'EBM_NoInt']}
- # PR curve
- pr_pooled = {'y_true': [], 'y_prob_pos': [], 'y_prob_neg': []}
- # Local explanation için en yüksek confidence TP/FP/FN
- local_cands = {
- 'tp': None, 'fp': None, 'fn': None,
- 'tp_conf': -1, 'fp_conf': -1, 'fn_conf': -1
- }
- ebm_imp_dict = {} # term_name -> [importance per iter]
- noint_imp_dict = {} # term_name -> [importance per iter]
- sel_counts = []
- ebm_term_names = None
- sss = StratifiedShuffleSplit(
- n_splits=n_iter, test_size=0.20, random_state=42
- )
- for fold, (tr_idx, te_idx) in enumerate(sss.split(X, y)):
- print(f"\n[Iter {fold+1}/{n_iter}]", end=" ", flush=True)
- X_tr_raw, X_te_raw = X[tr_idx], X[te_idx]
- y_tr, y_te = y[tr_idx], y[te_idx]
- # Adım 1: Preprocess (train-only fit)
- X_tr, X_te = preprocess(X_tr_raw, X_te_raw)
- # Adım 2: PRNN — train-only, SVM internal wrapper
- sel_idx = run_prnn(X_tr, y_tr, feature_names, p=0.5)
- sel_counts.append(len(sel_idx))
- X_tr_s = X_tr[:, sel_idx]
- X_te_s = X_te[:, sel_idx]
- sel_names = [feature_names[i] for i in sel_idx]
- print(f"PRNN={len(sel_idx)}", end="", flush=True)
- # Adım 3: EBM (interactions='auto')
- ebm = grid_search(
- ExplainableBoostingClassifier(interactions='auto', random_state=42),
- EBM_GRID, X_tr_s, y_tr
- )
- ebm.fit(X_tr_s, y_tr)
- y_prob_ebm = ebm.predict_proba(X_te_s)[:, 1]
- y_pred_ebm = ebm.predict(X_te_s)
- _store_metrics(store['EBM'], y_te, y_tr, ebm,
- X_te_s, X_tr_s, mean_fpr, mean_recall)
- # Term adına göre importance topla (split bazında farklı term sırası)
- for tname, timp in zip(ebm.term_names_, ebm.term_importances()):
- ebm_imp_dict.setdefault(tname, []).append(timp)
- print(" EBM✓", end="", flush=True)
- # Pooled predictions
- pr_pooled['y_true'].extend(y_te.tolist())
- pr_pooled['y_prob_pos'].extend(y_prob_ebm.tolist())
- pr_pooled['y_prob_neg'].extend((1 - y_prob_ebm).tolist())
- # Local explanation candidates — en yüksek confidence TP/FP/FN
- for i, (yt, yp, ypr) in enumerate(zip(y_te, y_pred_ebm, y_prob_ebm)):
- conf = abs(ypr - 0.5)
- if yt == 1 and yp == 1 and conf > local_cands['tp_conf']:
- local_cands['tp'] = {
- 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
- 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
- }
- local_cands['tp_conf'] = conf
- elif yt == 0 and yp == 1 and conf > local_cands['fp_conf']:
- local_cands['fp'] = {
- 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
- 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
- }
- local_cands['fp_conf'] = conf
- elif yt == 1 and yp == 0 and conf > local_cands['fn_conf']:
- local_cands['fn'] = {
- 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
- 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
- }
- local_cands['fn_conf'] = conf
- # Adım 4: EBM Ablation (interactions=0)
- ebm_ni = ExplainableBoostingClassifier(
- interactions=0, random_state=42,
- max_rounds=ebm.max_rounds,
- learning_rate=ebm.learning_rate,
- )
- ebm_ni.fit(X_tr_s, y_tr)
- _store_metrics(store['EBM_NoInt'], y_te, y_tr, ebm_ni,
- X_te_s, X_tr_s, mean_fpr, mean_recall)
- # EBM_NoInt: term adına göre importance topla
- for tname, timp in zip(ebm_ni.term_names_, ebm_ni.term_importances()):
- noint_imp_dict.setdefault(tname, []).append(timp)
- print(" NI✓", end="", flush=True)
- # Adım 5: XGBoost
- xgb = grid_search(
- XGBClassifier(random_state=42, eval_metric='logloss', verbosity=0),
- XGB_GRID, X_tr_s, y_tr
- )
- xgb.fit(X_tr_s, y_tr)
- _store_metrics(store['XGBoost'], y_te, y_tr, xgb,
- X_te_s, X_tr_s, mean_fpr, mean_recall)
- print(" XGB✓", end="", flush=True)
- # Adım 6: LightGBM
- lgb = grid_search(
- LGBMClassifier(random_state=42, verbose=-1),
- LGB_GRID, X_tr_s, y_tr
- )
- lgb.fit(X_tr_s, y_tr)
- _store_metrics(store['LightGBM'], y_te, y_tr, lgb,
- X_te_s, X_tr_s, mean_fpr, mean_recall)
- print(" LGB✓", flush=True)
- # Kaydet
- with open(f'{output_dir}/pipeline_results.pkl', 'wb') as f:
- pickle.dump({
- 'store': store, 'ebm_imp_dict': ebm_imp_dict,
- 'noint_imp_dict': noint_imp_dict, 'sel_counts': sel_counts,
- 'mean_fpr': mean_fpr,
- 'mean_recall': mean_recall, 'feature_names': feature_names,
- 'pr_pooled': pr_pooled, 'local_cands': local_cands,
- }, f)
- print(f"\nResults saved → {output_dir}/pipeline_results.pkl")
- return (store, ebm_imp_dict, noint_imp_dict, sel_counts,
- mean_fpr, mean_recall, pr_pooled, local_cands)
- # ============================================================
- # BÖLÜM 6: SONUÇ ÖZETİ
- # ============================================================
- def summarise(store):
- keys = {
- 'aucs':'AUC', 'accs':'Accuracy', 'f1s':'F1-Score',
- 'sens':'Sensitivity', 'specs':'Specificity',
- 'aps':'AUPRC', 'train_aucs':'Train AUC'
- }
- summary = {}
- for model, data in store.items():
- summary[model] = {}
- for k, label in keys.items():
- if k not in data or not data[k]: continue
- arr = np.array(data[k])
- m, lo, hi = arr.mean(), np.percentile(arr,2.5), np.percentile(arr,97.5)
- summary[model][label] = {
- 'mean': round(m,3), 'ci_lo': round(lo,3), 'ci_hi': round(hi,3),
- 'fmt': f"{m:.3f} ({lo:.3f}–{hi:.3f})"
- }
- return summary
- def print_table2(summary):
- order = ['Accuracy','F1-Score','Sensitivity','Specificity','AUC']
- print("\n" + "="*75)
- print("TABLE 2: Classification Performance (mean; 95% CI)")
- print("="*75)
- print(f"{'Metric':<15}{'EBM':>20}{'XGBoost':>20}{'LightGBM':>20}")
- print("-"*75)
- for m in order:
- row = f"{m:<15}"
- for mdl in ['EBM','XGBoost','LightGBM']:
- row += f"{summary[mdl].get(m,{}).get('fmt','N/A'):>20}"
- print(row)
- print("\nOverfitting check (Train vs Test AUC):")
- for mdl in ['EBM','XGBoost','LightGBM']:
- tr = summary[mdl].get('Train AUC',{}).get('mean','?')
- te = summary[mdl].get('AUC',{}).get('mean','?')
- if isinstance(tr,float) and isinstance(te,float):
- print(f" {mdl}: Train={tr:.3f}, Test={te:.3f}, Gap={tr-te:.3f}")
- print("\nAblation (EBM with vs without interactions):")
- ai = summary.get('EBM',{}).get('AUC',{})
- an = summary.get('EBM_NoInt',{}).get('AUC',{})
- if ai and an:
- print(f" With interactions: {ai['fmt']}")
- print(f" Without interactions: {an['fmt']}")
- print(f" ΔAUC = {ai['mean']-an['mean']:.3f}")
- # ============================================================
- # BÖLÜM 7: FEATURE IMPORTANCE
- # ============================================================
- def get_top_importances(imp_dict, top_n=15):
- term_means = {name: np.mean(vals) for name, vals in imp_dict.items()}
- sorted_terms = sorted(term_means.items(), key=lambda x: x[1], reverse=True)
- top = sorted_terms[:top_n]
- names = [t[0] for t in top]
- scores = np.array([t[1] for t in top])
- return names, scores
- # ============================================================
- # BÖLÜM 8: GÖRSELLEŞTİRME
- # ============================================================
- def figure1_roc(store, mean_fpr, output_dir):
- cfg = {
- 'EBM': {'color':'#1f77b4','ls':'-', 'lw':2.2},
- 'XGBoost': {'color':'#d62728','ls':'--','lw':2.0},
- 'LightGBM': {'color':'#2ca02c','ls':'-.','lw':2.0},
- }
- fig, ax = plt.subplots(figsize=(6.5, 6.5), dpi=150)
- fig.patch.set_facecolor('white'); ax.set_facecolor('white')
- ax.plot([0,1],[0,1], color='gray', ls='--', lw=1.0, alpha=0.5)
- for name, c in cfg.items():
- tprs = np.array(store[name]['tprs'])
- aucs = np.array(store[name]['aucs'])
- mean_tpr = tprs.mean(axis=0); mean_tpr[-1] = 1.0
- lo_tpr = np.percentile(tprs, 2.5, axis=0)
- hi_tpr = np.percentile(tprs, 97.5, axis=0)
- m_auc = aucs.mean()
- lo_auc = np.percentile(aucs, 2.5)
- hi_auc = np.percentile(aucs, 97.5)
- ax.fill_between(mean_fpr, lo_tpr, hi_tpr,
- color=c['color'], alpha=0.10, linewidth=0)
- ax.plot(mean_fpr, mean_tpr, color=c['color'], lw=c['lw'], ls=c['ls'],
- label=(f"{name} (AUC = {m_auc:.3f}; "
- f"95% CI: {lo_auc:.3f}–{hi_auc:.3f})"))
- ax.set_xlim([0,1]); ax.set_ylim([0,1])
- ax.set_xlabel('False Positive Rate (1 − Specificity)', fontsize=11)
- ax.set_ylabel('True Positive Rate (Sensitivity)', fontsize=11)
- ax.legend(loc='lower right', fontsize=8.5, framealpha=0.95, edgecolor='#ccc')
- ax.grid(True, linestyle=':', alpha=0.3)
- for sp in ax.spines.values(): sp.set_linewidth(0.8); sp.set_color('#aaa')
- plt.tight_layout()
- plt.savefig(f'{output_dir}/figure1_roc.png', dpi=300,
- bbox_inches='tight', facecolor='white')
- plt.close()
- print("Figure 1 (ROC) saved.")
- def figure2_pr(pr_pooled, output_dir):
- y_true = np.array(pr_pooled['y_true'])
- y_prob_pos = np.array(pr_pooled['y_prob_pos'])
- y_prob_neg = np.array(pr_pooled['y_prob_neg'])
- fig, ax = plt.subplots(figsize=(7, 6), dpi=150)
- fig.patch.set_facecolor('white'); ax.set_facecolor('white')
- # Class 1 (ME/CFS)
- prec1, rec1, _ = precision_recall_curve(y_true, y_prob_pos, pos_label=1)
- ap1 = average_precision_score(y_true, y_prob_pos)
- ax.plot(rec1, prec1, color='#2ca02c', lw=2.0, ls='-',
- label=f'Precision-recall curve of class 1 (area = {ap1:.3f})')
- # Class 0 (Control)
- prec0, rec0, _ = precision_recall_curve(1-y_true, y_prob_neg, pos_label=1)
- ap0 = average_precision_score(1-y_true, y_prob_neg)
- ax.plot(rec0, prec0, color='black', lw=2.0, ls='-',
- label=f'Precision-recall curve of class 0 (area = {ap0:.3f})')
- # Micro-average (pooled)
- y_bin = label_binarize(y_true, classes=[0, 1])
- y_score = np.column_stack([y_prob_neg, y_prob_pos])
- prec_m, rec_m, _ = precision_recall_curve(y_bin.ravel(), y_score.ravel())
- ap_micro = average_precision_score(y_bin, y_score, average='micro')
- ax.plot(rec_m, prec_m, color='#1f77b4', lw=2.0, ls=':',
- label=f'micro-average Precision-recall curve (area = {ap_micro:.3f})')
- ax.set_xlim([0,1]); ax.set_ylim([0,1.02])
- ax.set_xlabel('Recall', fontsize=11)
- ax.set_ylabel('Precision', fontsize=11)
- ax.set_title('Precision-Recall Curve', fontsize=12)
- ax.legend(loc='lower left', fontsize=9, framealpha=0.95, edgecolor='#ccc')
- ax.grid(True, linestyle=':', alpha=0.3)
- for sp in ax.spines.values(): sp.set_linewidth(0.8); sp.set_color('#aaa')
- plt.tight_layout()
- plt.savefig(f'{output_dir}/figure2_pr.png', dpi=300,
- bbox_inches='tight', facecolor='white')
- plt.close()
- print(f"Figure 2 (PR) saved. Class1={ap1:.3f}, Class0={ap0:.3f}, Micro={ap_micro:.3f}")
- def figure3_importance(names_int, scores_int, names_ni, scores_ni,
- auc_int, auc_ni, output_dir):
- fig, axes = plt.subplots(1, 2, figsize=(22, 9))
- fig.patch.set_facecolor('white')
- BG='#EEF2F7'; BAR='#FF8C00'
- for ax, names, scores, label, auc_t in [
- (axes[0], names_int, scores_int, '(A)',
- f"AUC={auc_int[0]:.3f} (95% CI: {auc_int[1]:.3f}–{auc_int[2]:.3f})"),
- (axes[1], names_ni, scores_ni, '(B)',
- f"AUC={auc_ni[0]:.3f} (95% CI: {auc_ni[1]:.3f}–{auc_ni[2]:.3f})"),
- ]:
- ax.set_facecolor(BG)
- y_pos = np.arange(len(names))
- ax.barh(y_pos, scores[::-1], color=BAR, height=0.65,
- edgecolor='white', linewidth=0.4)
- ax.set_yticks(y_pos); ax.set_yticklabels(names[::-1], fontsize=12)
- ax.set_xlabel('Mean Absolute Score (Weighted)', fontsize=12)
- ax.text(-0.02, 1.03, label, transform=ax.transAxes,
- fontsize=15, fontweight='bold', va='bottom')
- ax.text(0.5, -0.07, auc_t, transform=ax.transAxes,
- ha='center', fontsize=10, style='italic', color='#444')
- ax.grid(axis='x', linestyle='--', alpha=0.35, color='white')
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.spines['left'].set_visible(False)
- ax.tick_params(left=False)
- ax.set_xlim(0, max(scores) * 1.08)
- plt.tight_layout(pad=3.0); plt.subplots_adjust(bottom=0.12)
- plt.savefig(f'{output_dir}/figure3_importance.png', dpi=300,
- bbox_inches='tight', facecolor='white')
- plt.close()
- print("Figure 3 (Importance) saved.")
- def figure4_local(local_cands, output_dir, top_n=15):
- cases = [
- ('tp', 'A', 'True Positive'),
- ('fp', 'B', 'False Positive'),
- ('fn', 'C', 'False Negative'),
- ]
- fig, axes = plt.subplots(1, 3, figsize=(21, 7))
- fig.patch.set_facecolor('white')
- for ax, (key, panel, label) in zip(axes, cases):
- case = local_cands.get(key)
- if case is None:
- ax.text(0.5, 0.5, f'No {label} found',
- ha='center', va='center', transform=ax.transAxes)
- continue
- ebm = case['ebm']
- x = case['x']
- prob = case['prob']
- y_true = case['y_true']
- y_pred = case['y_pred']
- # EBM local explanation
- expl = ebm.explain_local(x.reshape(1,-1), np.array([y_true]))
- scores_raw = expl.data(0)['scores']
- names_raw = expl.data(0)['names']
- abs_scores = np.abs(scores_raw)
- top_idx = np.argsort(abs_scores)[-top_n:]
- top_scores = np.array(scores_raw)[top_idx]
- top_names = np.array(names_raw)[top_idx]
- colors = ['#FF8C00' if s > 0 else '#1f77b4' for s in top_scores]
- ax.set_facecolor('#EEF2F7')
- y_pos = np.arange(len(top_names))
- ax.barh(y_pos, top_scores, color=colors, height=0.65,
- edgecolor='white', linewidth=0.3)
- ax.axvline(x=0, color='gray', lw=0.8, alpha=0.7)
- ax.set_yticks(y_pos); ax.set_yticklabels(top_names, fontsize=8.5)
- ax.set_xlabel('Contribution to Prediction', fontsize=10)
- pred_str = 'ME/CFS' if y_pred == 1 else 'Control'
- true_str = 'ME/CFS' if y_true == 1 else 'Control'
- prob_str = (f"Pr(y=1) = {prob:.3f}" if y_pred == 1
- else f"Pr(y=0) = {1-prob:.3f}")
- ax.set_title(f"Local Explanation (Actual: {true_str} | "
- f"Predicted: {pred_str}\n{prob_str})", fontsize=9, pad=6)
- ax.text(-0.02, 1.05, f'({panel})', transform=ax.transAxes,
- fontsize=13, fontweight='bold', va='bottom')
- ax.spines['top'].set_visible(False)
- ax.spines['right'].set_visible(False)
- ax.tick_params(left=False)
- ax.grid(axis='x', linestyle='--', alpha=0.3, color='white')
- plt.tight_layout(pad=2.5)
- plt.savefig(f'{output_dir}/figure4_local.png', dpi=300,
- bbox_inches='tight', facecolor='white')
- plt.close()
- print("Figure 4 (Local) saved.")
- # ============================================================
- # BÖLÜM 9: TANIMLAYICI İSTATİSTİK
- # ============================================================
- def descriptive_stats(df):
- mecfs = df[df['factor'] == 1]
- control = df[df['factor'] == 0]
- print("\n" + "="*60 + "\nTABLE 1: Descriptive Statistics\n" + "="*60)
- for label, col in [('Age','age'), ('BMI','bmi')]:
- if col not in df.columns: continue
- _, p = mannwhitneyu(mecfs[col].dropna(), control[col].dropna(),
- alternative='two-sided')
- print(f"{label}: ME/CFS={mecfs[col].mean():.1f}±{mecfs[col].std():.1f}, "
- f"Control={control[col].mean():.1f}±{control[col].std():.1f}, p={p:.3f} *")
- for col in ['ibs','probiotic_use','antidepressant_use','narcotic_use']:
- if col not in df.columns: continue
- ct = pd.crosstab(df['factor'], df[col])
- if ct.shape == (2,2):
- _, p = fisher_exact(ct.values)
- n1,n2 = mecfs[col].sum(), control[col].sum()
- print(f"{col}: ME/CFS={n1}({100*n1/len(mecfs):.1f}%), "
- f"Control={n2}({100*n2/len(control):.1f}%), p={p:.3f} **")
- print("* Mann-Whitney U ** Fisher's exact")
- # ============================================================
- # BÖLÜM 10: MAIN
- # ============================================================
- if __name__ == '__main__':
- DATA_PATH = 'df.xlsx' # <-- veri dosyasının yolu
- OUTPUT_DIR = 'results'
- N_ITER = 50
- # 1. Veri yükle
- X, y, feature_names = load_data(DATA_PATH)
- # 2. Ana pipeline
- (store, ebm_imp_dict, noint_imp_dict, sel_counts,
- mean_fpr, mean_recall, pr_pooled, local_cands) = run_50repeat(
- X, y, feature_names, n_iter=N_ITER, output_dir=OUTPUT_DIR
- )
- # 3. Performans özeti
- summary = summarise(store)
- print_table2(summary)
- # 4. PRNN seçim istatistikleri
- n_feat = len(feature_names)
- print(f"\nPRNN selected: {min(sel_counts)}–{max(sel_counts)} features "
- f"(mean={np.mean(sel_counts):.1f}, "
- f"~{np.mean(sel_counts)/n_feat*100:.1f}% of {n_feat})")
- n_mean = int(np.mean(sel_counts))
- print(f"Mean possible pairs: C({n_mean},2) = {n_mean*(n_mean-1)//2}")
- print(f"Unique EBM terms across 50 iters: {len(ebm_imp_dict)}")
- # 5. Feature importance
- names_int, scores_int = get_top_importances(ebm_imp_dict)
- names_ni, scores_ni = get_top_importances(noint_imp_dict)
- # 6. Görseller
- figure1_roc(store, mean_fpr, OUTPUT_DIR)
- figure2_pr(pr_pooled, OUTPUT_DIR)
- ai = summary['EBM']['AUC']
- an = summary['EBM_NoInt']['AUC']
- figure3_importance(
- names_int, scores_int, names_ni, scores_ni,
- (ai['mean'], ai['ci_lo'], ai['ci_hi']),
- (an['mean'], an['ci_lo'], an['ci_hi']),
- OUTPUT_DIR
- )
- figure4_local(local_cands, OUTPUT_DIR)
- print(f"\nTüm çıktılar → {OUTPUT_DIR}/")
code.py at commit 0ffe3c3, no license · at the source
Overview
- Department of Biostatistics, Faculty of Medicine, Malatya Turgut Ozal University, Malatya 44210, Türkiye
- Department of Family Medicine, Faculty of Medicine, Malatya Turgut Ozal University, Malatya 44210, Türkiye
- Department of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, Malatya 44280, Türkiye
- Department of Computer Sciences, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, P.O. Box 84428, Riyadh 11671, Saudi Arabia
- Department of Physiology, College of Medicine, King Khalid University, Abha 61421, Saudi Arabia
- Department of Ocean Operations and Civil Engineering, Norwegian University of Science and Technology (NTNU), 6009 Ålesund, Norway
- Department of Sustainable Systems Engineering (INATECH), Albert Ludwigs University of Freiburg, 79110 Freiburg, Germany
Abstract
Myalgic encephalomyelitis/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.
drhilal/ME-CFS-study
0ffe3c328255b3b48fc3076619a581fddebf7e17, 11 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
1 file
- code.py, Python, 716 lines, 5 matches
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 1 script, each with its path and the digest of its content;
- 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors on request.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 6 keywords, 9 MeSH terms, 1 funder, 34 references.
Cite
This paper
Yagin, F. H., Korkmaz, Y., Colak, C., Alzakari, S. A., Alkhalifa, A. K., Al-Hashem, F., & Aghaei, M. (2026). Metabolomic Classification of Myalgic Encephalomyelitis/
BibTeX
@article{yagin2026metabo
author = {Yagin, Fatma Hilal and Korkmaz, Yavuz and Colak, Cemil and Alzakari, Sarah A and Alkhalifa, Amal K and Al-Hashem, Fahaid and Aghaei, Mohammadreza},
title = {{Metabolomic Classification of Myalgic Encephalomyelitis/
journal = {International journal of molecular sciences},
year = {2026},
month = jun,
volume = {27},
number = {13},
pages = {5920},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1422-0067},
doi = {10.3390/
url = {https://
pmid = {42450188},
pmcid = {PMC13362375}
}
RIS
TY - JOUR
AU - Yagin, Fatma Hilal
AU - Korkmaz, Yavuz
AU - Colak, Cemil
AU - Alzakari, Sarah A
AU - Alkhalifa, Amal K
AU - Al-Hashem, Fahaid
AU - Aghaei, Mohammadreza
TI - Metabolomic Classification of Myalgic Encephalomyelitis/
T2 - International journal of molecular sciences
J2 - Int J Mol Sci
PY - 2026
DA - 2026/
VL - 27
IS - 13
SP - 5920
SN - 1422-0067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "Metabolomic Classification of Myalgic Encephalomyelitis/
"container-title": "International journal of molecular sciences",
"author": [
{
"family": "Yagin",
"given": "Fatma Hilal"
},
{
"family": "Korkmaz",
"given": "Yavuz"
},
{
"family": "Colak",
"given": "Cemil"
},
{
"family": "Alzakari",
"given": "Sarah A"
},
{
"family": "Alkhalifa",
"given": "Amal K"
},
{
"family": "Al-Hashem",
"given": "Fahaid"
},
{
"family": "Aghaei",
"given": "Mohammadreza"
}
],
"container-title-short":
"volume": "27",
"issue": "13",
"page": "5920",
"DOI": "10.3390/
"PMID": "42450188",
"PMCID": "PMC13362375",
"ISSN": "1422-0067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
30
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.3389/fnhum.2026.1869918 [code]
- Single-subject auditory ERP-BCI performance enhancement in ALS via an AI coding assistant prompt.Journal: Frontiers in human neuroscienceIn common: LightGBM, XGBoost, scikit-learn, 4 other tools, clinical / translational
- [2] doi:10.1073/pnas.2516601123 [code]
- Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.Journal: Proceedings of the National Academy of Sciences of the United States of AmericaIn common: LightGBM, XGBoost, scikit-learn, 4 other tools, clinical / translational
- [3] doi:10.7554/elife.110588 [code]
- Opening the black box toward a modular approach to spike sorting.Journal: eLifeIn common: LightGBM, XGBoost, scikit-learn, 4 other tools
- [4] doi:10.1186/s13244-026-02365-7 [code]
- Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral
radiomics in preoperative prediction of rectal cancer perineural invasion. Journal: Insights into imagingIn common: LightGBM, XGBoost, scikit-learn, 4 other tools - [5] doi:10.1371/journal.pone.0337726 [code]
- Classification of chronic pain and spinal cord stimulation response using machine learning in magnetoencephalography dataJournal: n/aIn common: LightGBM, XGBoost, scikit-learn, 4 other tools
- [6] doi:10.1371/journal.pgen.1012242 [code]
- Wiz regulates clustered protocadherin genes by restricting CTCF/
cohesin loop extrusion in a genomic-distance biased manner. Journal: PLoS geneticsIn common: LightGBM, XGBoost, scikit-learn, 4 other tools - [7] doi:10.3389/fpsyg.2026.1774068 [code]
- Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.Journal: Frontiers in psychologyIn common: LightGBM, XGBoost, scikit-learn, 4 other tools
- [8] doi:10.1038/s41591-026-04432-4 [code]
- Activity-dependent adaptive deep brain stimulation improves gait in Parkinson's disease.Journal: Nature medicineIn common: LightGBM, XGBoost, scikit-learn, 3 other tools, clinical / translational
- [9] doi:10.1371/journal.pone.0354510 [code]
- Personalized adaptive virtual reality experience driven by electroencephalography-b
ased pain recognition. Journal: PloS oneIn common: LightGBM, XGBoost, scikit-learn, 3 other tools - [10] doi:10.1212/wnl.0000000000218472 [code]
- Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.Journal: NeurologyIn common: LightGBM, scikit-learn, pandas, 3 other tools, clinical / translational
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 1 script, and 5 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:4e8d8c29a0ac4acd…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
