OSCR

Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches
  1. [1] § 4. Materials and Methods › 4.4. Feature Selection: Pareto-Guided Recursive Neural Network (PRNN) ↔ code.py, lines 74–120 · score 0.79 · hidden layers, feature selection, Adam, ReLU, activation, PNN
  2. [2] § 4. Materials and Methods › 4.7. Discriminative Ability: ROC and Precision–Recall Analyses ↔ code.py, lines 250–384 · score 0.67 · PR curves, LightGBM, XGBoost, confidence, metrics, recall
  3. [3] § 4. Materials and Methods › 4.5. Model Development: Ensemble Learning Classifiers ↔ code.py, lines 250–384 · score 0.56 · LightGBM, XGBoost, ablation, preprocessing, boosting, pipeline
  4. [4] § 2. Results › 2.2. Model Performance ↔ code.py, lines 412–436 · score 0.54 · F1 score, LightGBM, XGBoost, sensitivity, metrics, accuracy
  5. [5] § 2. Results › 2.2. Model Performance ↔ code.py, lines 412–436 · score 0.53 · Training AUC, LightGBM, XGBoost, gap, overfitting, models

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 716 lines · 26 KB · no license · 5 matches

  1. # ============================================================
  2. # BÖLÜM 1: IMPORTS
  3. # ============================================================
  4. import numpy as np
  5. import pandas as pd
  6. import matplotlib
  7. matplotlib.use('Agg')
  8. import matplotlib.pyplot as plt
  9. import warnings, pickle, math
  10. from pathlib import Path
  11. from scipy.stats import mannwhitneyu, fisher_exact
  12. from sklearn.model_selection import (StratifiedShuffleSplit,
  13. GridSearchCV, StratifiedKFold)
  14. from sklearn.preprocessing import StandardScaler, MinMaxScaler, label_binarize
  15. from sklearn.neural_network import MLPClassifier
  16. from sklearn.svm import SVC
  17. from sklearn.metrics import (
  18. roc_auc_score, accuracy_score, f1_score, recall_score,
  19. roc_curve, auc, precision_recall_curve, average_precision_score
  20. )
  21. from interpret.glassbox import ExplainableBoostingClassifier
  22. from xgboost import XGBClassifier
  23. from lightgbm import LGBMClassifier
  24. warnings.filterwarnings('ignore')
  25. # ============================================================
  26. # BÖLÜM 2: VERİ YÜKLEME VE ÖN İŞLEME
  27. # ============================================================
  28. def load_data(filepath):
  29. df = pd.read_excel(filepath)
  30. y = df['factor'].values
  31. X = df.drop('factor', axis=1)
  32. feature_names = X.columns.tolist()
  33. print(f"Dataset: {X.shape[0]} samples, {X.shape[1]} features")
  34. print(f"ME/CFS(1): {y.sum()}, Control(0): {(y==0).sum()}")
  35. return X.values, y, feature_names
  36. def preprocess(X_train_raw, X_test_raw):
  37. X_tr = np.log2(X_train_raw + 1)
  38. X_te = np.log2(X_test_raw + 1)
  39. scaler = StandardScaler()
  40. return scaler.fit_transform(X_tr), scaler.transform(X_te)
  41. # ============================================================
  42. # BÖLÜM 3: PRNN FEATURE SELECTION
  43. # ============================================================
  44. def is_pareto_efficient(costs, return_mask=True):
  45. is_efficient = np.arange(costs.shape[0])
  46. n_points = costs.shape[0]
  47. next_point_index = 0
  48. while next_point_index < len(costs):
  49. nondominated = np.any(costs > costs[next_point_index], axis=1)
  50. nondominated[next_point_index] = True
  51. is_efficient = is_efficient[nondominated]
  52. costs = costs[nondominated]
  53. next_point_index = np.sum(nondominated[:next_point_index]) + 1
  54. if return_mask:
  55. mask = np.zeros(n_points, dtype=bool)
  56. mask[is_efficient] = True
  57. return mask
  58. return is_efficient
  59. def pnn_feature_selection(X_train, y_train, p, feature_names):
  60. mlp = MLPClassifier(
  61. hidden_layer_sizes=(64, 32),
  62. activation='relu',
  63. solver='adam',
  64. max_iter=100,
  65. early_stopping=True,
  66. n_iter_no_change=10,
  67. random_state=42
  68. )
  69. mlp.fit(X_train, y_train)
  70. # Ağırlık matrisi çarpımı → (n_features, n_classes)
  71. W = mlp.coefs_[0]
  72. for w in mlp.coefs_[1:]:
  73. W = W @ w
  74. P_W = np.abs(W)
  75. P_W_norm = MinMaxScaler().fit_transform(P_W)
  76. # Objective 1: max normalize skor
  77. Obj_max = P_W_norm.max(axis=1)
  78. # Objective 2: 1/rank
  79. P_W_rank = P_W_norm.copy()
  80. for j in range(P_W_norm.shape[1]):
  81. idx = np.argsort(-P_W_norm[:, j])
  82. P_W_rank[:, j] = np.argsort(idx) + 1
  83. Obj_rank_max = (1.0 / P_W_rank).max(axis=1)
  84. # Pareto frontlarına böl
  85. Obj = np.column_stack([Obj_max, Obj_rank_max])
  86. genes = list(feature_names)
  87. remaining = Obj.copy()
  88. rem_genes = genes.copy()
  89. fronts = []
  90. while len(remaining) > 0:
  91. mask = is_pareto_efficient(remaining)
  92. fronts.append([g for g, m in zip(rem_genes, mask) if m])
  93. rem_genes = [g for g, m in zip(rem_genes, mask) if not m]
  94. remaining = remaining[~mask]
  95. n_take = max(1, int(len(fronts) * p))
  96. selected = []
  97. for i in range(n_take):
  98. selected.extend(fronts[i])
  99. return selected
  100. def prnn_recursive(data_x, data_y, DX, DY, p=0.5,
  101. best_acc=0, best_gene=None,
  102. depth=0, max_depth=20):
  103. if depth >= max_depth:
  104. return best_gene if best_gene else data_x.columns.tolist()
  105. n_orig = DX.shape[1]
  106. n_curr = data_x.shape[1]
  107. # Dinamik p — Li et al. 2025 formülü (birebir)
  108. if p < 0.9:
  109. p = 0.5 + (math.tan((9 * math.pi / 40) * (1 - (n_curr / n_orig)))) / 2
  110. gene = pnn_feature_selection(
  111. data_x.values, data_y.values, p, data_x.columns.tolist()
  112. )
  113. if not gene:
  114. return best_gene if best_gene else data_x.columns.tolist()
  115. # SVM 5-fold CV — iç değerlendirme kriteri
  116. skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=111)
  117. accs = []
  118. for tr, te in skf.split(DX, DY):
  119. svm = SVC(random_state=42)
  120. svm.fit(DX.iloc[tr][gene], DY.iloc[tr])
  121. accs.append(accuracy_score(DY.iloc[te], svm.predict(DX.iloc[te][gene])))
  122. mean_acc = np.mean(accs)
  123. # En iyi subset güncelle (200'den az feature koşulu)
  124. if mean_acc > best_acc and len(gene) < 200:
  125. best_acc = mean_acc
  126. best_gene = gene
  127. # Durma koşulu
  128. if len(gene) <= 10:
  129. return best_gene if best_gene else gene
  130. # Recursive: seçilen gene'lerle devam
  131. return prnn_recursive(
  132. data_x[gene], data_y, DX, DY, p,
  133. best_acc, best_gene, depth + 1, max_depth
  134. )
  135. def run_prnn(X_scaled, y, feature_names, p=0.5):
  136. df_X = pd.DataFrame(X_scaled, columns=feature_names)
  137. df_y = pd.Series(y)
  138. selected = prnn_recursive(df_X, df_y, df_X, df_y, p=0.5)
  139. if not selected:
  140. print(" [WARN] PRNN seçim yapamadı, tüm feature'lar kullanılıyor")
  141. return np.arange(len(feature_names))
  142. name_to_idx = {n: i for i, n in enumerate(feature_names)}
  143. return np.array([name_to_idx[g] for g in selected if g in name_to_idx])
  144. # ============================================================
  145. # BÖLÜM 4: HİPERPARAMETRE GRİDLERİ
  146. # ============================================================
  147. EBM_GRID = {
  148. 'max_rounds': [100, 200, 500],
  149. 'learning_rate': [0.01, 0.05, 0.10],
  150. }
  151. XGB_GRID = {
  152. 'n_estimators': [100, 200, 500],
  153. 'max_depth': [3, 5, 7],
  154. 'learning_rate': [0.01, 0.05, 0.10],
  155. 'subsample': [0.6, 0.8, 1.0],
  156. 'reg_alpha': [0, 0.1, 1.0],
  157. 'reg_lambda': [0, 0.1, 1.0],
  158. }
  159. LGB_GRID = {
  160. 'n_estimators': [100, 200, 500],
  161. 'max_depth': [3, 5, 7],
  162. 'learning_rate': [0.01, 0.05, 0.10],
  163. 'subsample': [0.6, 0.8, 1.0],
  164. 'reg_alpha': [0, 0.1, 1.0],
  165. 'reg_lambda': [0, 0.1, 1.0],
  166. }
  167. def grid_search(model, param_grid, X_train, y_train):
  168. """5-fold CV, sadece training partition — leakage yok"""
  169. gs = GridSearchCV(model, param_grid, cv=5,
  170. scoring='roc_auc', n_jobs=-1, refit=True)
  171. gs.fit(X_train, y_train)
  172. return gs.best_estimator_
  173. # ============================================================
  174. # BÖLÜM 5: ANA 50-REPEAT DÖNGÜSÜ
  175. # ============================================================
  176. def _store_metrics(store, y_te, y_tr, model, X_te, X_tr,
  177. mean_fpr, mean_recall):
  178. """Her iterasyon için tüm metrikleri topla"""
  179. y_prob = model.predict_proba(X_te)[:, 1]
  180. y_pred = model.predict(X_te)
  181. y_prob_tr = model.predict_proba(X_tr)[:, 1]
  182. store['aucs'].append(roc_auc_score(y_te, y_prob))
  183. store['accs'].append(accuracy_score(y_te, y_pred))
  184. store['f1s'].append(f1_score(y_te, y_pred))
  185. store['sens'].append(recall_score(y_te, y_pred))
  186. store['aps'].append(average_precision_score(y_te, y_prob))
  187. store['train_aucs'].append(roc_auc_score(y_tr, y_prob_tr))
  188. tn = np.sum((y_te == 0) & (y_pred == 0))
  189. fp = np.sum((y_te == 0) & (y_pred == 1))
  190. store['specs'].append(tn / (tn + fp) if (tn + fp) > 0 else 0)
  191. fpr, tpr, _ = roc_curve(y_te, y_prob)
  192. tpr_i = np.interp(mean_fpr, fpr, tpr); tpr_i[0] = 0.0
  193. store['tprs'].append(tpr_i)
  194. prec, rec, _ = precision_recall_curve(y_te, y_prob)
  195. store['precs'].append(np.interp(mean_recall, rec[::-1], prec[::-1]))
  196. def run_50repeat(X, y, feature_names, n_iter=50, output_dir='results'):
  197. Path(output_dir).mkdir(exist_ok=True)
  198. mean_fpr = np.linspace(0, 1, 300)
  199. mean_recall = np.linspace(0, 1, 300)
  200. store = {m: {
  201. 'aucs':[], 'accs':[], 'f1s':[], 'sens':[], 'specs':[],
  202. 'aps':[], 'tprs':[], 'precs':[], 'train_aucs':[]
  203. } for m in ['EBM', 'XGBoost', 'LightGBM', 'EBM_NoInt']}
  204. # PR curve
  205. pr_pooled = {'y_true': [], 'y_prob_pos': [], 'y_prob_neg': []}
  206. # Local explanation için en yüksek confidence TP/FP/FN
  207. local_cands = {
  208. 'tp': None, 'fp': None, 'fn': None,
  209. 'tp_conf': -1, 'fp_conf': -1, 'fn_conf': -1
  210. }
  211. ebm_imp_dict = {} # term_name -> [importance per iter]
  212. noint_imp_dict = {} # term_name -> [importance per iter]
  213. sel_counts = []
  214. ebm_term_names = None
  215. sss = StratifiedShuffleSplit(
  216. n_splits=n_iter, test_size=0.20, random_state=42
  217. )
  218. for fold, (tr_idx, te_idx) in enumerate(sss.split(X, y)):
  219. print(f"\n[Iter {fold+1}/{n_iter}]", end=" ", flush=True)
  220. X_tr_raw, X_te_raw = X[tr_idx], X[te_idx]
  221. y_tr, y_te = y[tr_idx], y[te_idx]
  222. # Adım 1: Preprocess (train-only fit)
  223. X_tr, X_te = preprocess(X_tr_raw, X_te_raw)
  224. # Adım 2: PRNN — train-only, SVM internal wrapper
  225. sel_idx = run_prnn(X_tr, y_tr, feature_names, p=0.5)
  226. sel_counts.append(len(sel_idx))
  227. X_tr_s = X_tr[:, sel_idx]
  228. X_te_s = X_te[:, sel_idx]
  229. sel_names = [feature_names[i] for i in sel_idx]
  230. print(f"PRNN={len(sel_idx)}", end="", flush=True)
  231. # Adım 3: EBM (interactions='auto')
  232. ebm = grid_search(
  233. ExplainableBoostingClassifier(interactions='auto', random_state=42),
  234. EBM_GRID, X_tr_s, y_tr
  235. )
  236. ebm.fit(X_tr_s, y_tr)
  237. y_prob_ebm = ebm.predict_proba(X_te_s)[:, 1]
  238. y_pred_ebm = ebm.predict(X_te_s)
  239. _store_metrics(store['EBM'], y_te, y_tr, ebm,
  240. X_te_s, X_tr_s, mean_fpr, mean_recall)
  241. # Term adına göre importance topla (split bazında farklı term sırası)
  242. for tname, timp in zip(ebm.term_names_, ebm.term_importances()):
  243. ebm_imp_dict.setdefault(tname, []).append(timp)
  244. print(" EBM✓", end="", flush=True)
  245. # Pooled predictions
  246. pr_pooled['y_true'].extend(y_te.tolist())
  247. pr_pooled['y_prob_pos'].extend(y_prob_ebm.tolist())
  248. pr_pooled['y_prob_neg'].extend((1 - y_prob_ebm).tolist())
  249. # Local explanation candidates — en yüksek confidence TP/FP/FN
  250. for i, (yt, yp, ypr) in enumerate(zip(y_te, y_pred_ebm, y_prob_ebm)):
  251. conf = abs(ypr - 0.5)
  252. if yt == 1 and yp == 1 and conf > local_cands['tp_conf']:
  253. local_cands['tp'] = {
  254. 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
  255. 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
  256. }
  257. local_cands['tp_conf'] = conf
  258. elif yt == 0 and yp == 1 and conf > local_cands['fp_conf']:
  259. local_cands['fp'] = {
  260. 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
  261. 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
  262. }
  263. local_cands['fp_conf'] = conf
  264. elif yt == 1 and yp == 0 and conf > local_cands['fn_conf']:
  265. local_cands['fn'] = {
  266. 'x': X_te_s[i], 'y_true': int(yt), 'y_pred': int(yp),
  267. 'prob': float(ypr), 'sel_names': sel_names, 'ebm': ebm
  268. }
  269. local_cands['fn_conf'] = conf
  270. # Adım 4: EBM Ablation (interactions=0)
  271. ebm_ni = ExplainableBoostingClassifier(
  272. interactions=0, random_state=42,
  273. max_rounds=ebm.max_rounds,
  274. learning_rate=ebm.learning_rate,
  275. )
  276. ebm_ni.fit(X_tr_s, y_tr)
  277. _store_metrics(store['EBM_NoInt'], y_te, y_tr, ebm_ni,
  278. X_te_s, X_tr_s, mean_fpr, mean_recall)
  279. # EBM_NoInt: term adına göre importance topla
  280. for tname, timp in zip(ebm_ni.term_names_, ebm_ni.term_importances()):
  281. noint_imp_dict.setdefault(tname, []).append(timp)
  282. print(" NI✓", end="", flush=True)
  283. # Adım 5: XGBoost
  284. xgb = grid_search(
  285. XGBClassifier(random_state=42, eval_metric='logloss', verbosity=0),
  286. XGB_GRID, X_tr_s, y_tr
  287. )
  288. xgb.fit(X_tr_s, y_tr)
  289. _store_metrics(store['XGBoost'], y_te, y_tr, xgb,
  290. X_te_s, X_tr_s, mean_fpr, mean_recall)
  291. print(" XGB✓", end="", flush=True)
  292. # Adım 6: LightGBM
  293. lgb = grid_search(
  294. LGBMClassifier(random_state=42, verbose=-1),
  295. LGB_GRID, X_tr_s, y_tr
  296. )
  297. lgb.fit(X_tr_s, y_tr)
  298. _store_metrics(store['LightGBM'], y_te, y_tr, lgb,
  299. X_te_s, X_tr_s, mean_fpr, mean_recall)
  300. print(" LGB✓", flush=True)
  301. # Kaydet
  302. with open(f'{output_dir}/pipeline_results.pkl', 'wb') as f:
  303. pickle.dump({
  304. 'store': store, 'ebm_imp_dict': ebm_imp_dict,
  305. 'noint_imp_dict': noint_imp_dict, 'sel_counts': sel_counts,
  306. 'mean_fpr': mean_fpr,
  307. 'mean_recall': mean_recall, 'feature_names': feature_names,
  308. 'pr_pooled': pr_pooled, 'local_cands': local_cands,
  309. }, f)
  310. print(f"\nResults saved → {output_dir}/pipeline_results.pkl")
  311. return (store, ebm_imp_dict, noint_imp_dict, sel_counts,
  312. mean_fpr, mean_recall, pr_pooled, local_cands)
  313. # ============================================================
  314. # BÖLÜM 6: SONUÇ ÖZETİ
  315. # ============================================================
  316. def summarise(store):
  317. keys = {
  318. 'aucs':'AUC', 'accs':'Accuracy', 'f1s':'F1-Score',
  319. 'sens':'Sensitivity', 'specs':'Specificity',
  320. 'aps':'AUPRC', 'train_aucs':'Train AUC'
  321. }
  322. summary = {}
  323. for model, data in store.items():
  324. summary[model] = {}
  325. for k, label in keys.items():
  326. if k not in data or not data[k]: continue
  327. arr = np.array(data[k])
  328. m, lo, hi = arr.mean(), np.percentile(arr,2.5), np.percentile(arr,97.5)
  329. summary[model][label] = {
  330. 'mean': round(m,3), 'ci_lo': round(lo,3), 'ci_hi': round(hi,3),
  331. 'fmt': f"{m:.3f} ({lo:.3f}–{hi:.3f})"
  332. }
  333. return summary
  334. def print_table2(summary):
  335. order = ['Accuracy','F1-Score','Sensitivity','Specificity','AUC']
  336. print("\n" + "="*75)
  337. print("TABLE 2: Classification Performance (mean; 95% CI)")
  338. print("="*75)
  339. print(f"{'Metric':<15}{'EBM':>20}{'XGBoost':>20}{'LightGBM':>20}")
  340. print("-"*75)
  341. for m in order:
  342. row = f"{m:<15}"
  343. for mdl in ['EBM','XGBoost','LightGBM']:
  344. row += f"{summary[mdl].get(m,{}).get('fmt','N/A'):>20}"
  345. print(row)
  346. print("\nOverfitting check (Train vs Test AUC):")
  347. for mdl in ['EBM','XGBoost','LightGBM']:
  348. tr = summary[mdl].get('Train AUC',{}).get('mean','?')
  349. te = summary[mdl].get('AUC',{}).get('mean','?')
  350. if isinstance(tr,float) and isinstance(te,float):
  351. print(f" {mdl}: Train={tr:.3f}, Test={te:.3f}, Gap={tr-te:.3f}")
  352. print("\nAblation (EBM with vs without interactions):")
  353. ai = summary.get('EBM',{}).get('AUC',{})
  354. an = summary.get('EBM_NoInt',{}).get('AUC',{})
  355. if ai and an:
  356. print(f" With interactions: {ai['fmt']}")
  357. print(f" Without interactions: {an['fmt']}")
  358. print(f" ΔAUC = {ai['mean']-an['mean']:.3f}")
  359. # ============================================================
  360. # BÖLÜM 7: FEATURE IMPORTANCE
  361. # ============================================================
  362. def get_top_importances(imp_dict, top_n=15):
  363. term_means = {name: np.mean(vals) for name, vals in imp_dict.items()}
  364. sorted_terms = sorted(term_means.items(), key=lambda x: x[1], reverse=True)
  365. top = sorted_terms[:top_n]
  366. names = [t[0] for t in top]
  367. scores = np.array([t[1] for t in top])
  368. return names, scores
  369. # ============================================================
  370. # BÖLÜM 8: GÖRSELLEŞTİRME
  371. # ============================================================
  372. def figure1_roc(store, mean_fpr, output_dir):
  373. cfg = {
  374. 'EBM': {'color':'#1f77b4','ls':'-', 'lw':2.2},
  375. 'XGBoost': {'color':'#d62728','ls':'--','lw':2.0},
  376. 'LightGBM': {'color':'#2ca02c','ls':'-.','lw':2.0},
  377. }
  378. fig, ax = plt.subplots(figsize=(6.5, 6.5), dpi=150)
  379. fig.patch.set_facecolor('white'); ax.set_facecolor('white')
  380. ax.plot([0,1],[0,1], color='gray', ls='--', lw=1.0, alpha=0.5)
  381. for name, c in cfg.items():
  382. tprs = np.array(store[name]['tprs'])
  383. aucs = np.array(store[name]['aucs'])
  384. mean_tpr = tprs.mean(axis=0); mean_tpr[-1] = 1.0
  385. lo_tpr = np.percentile(tprs, 2.5, axis=0)
  386. hi_tpr = np.percentile(tprs, 97.5, axis=0)
  387. m_auc = aucs.mean()
  388. lo_auc = np.percentile(aucs, 2.5)
  389. hi_auc = np.percentile(aucs, 97.5)
  390. ax.fill_between(mean_fpr, lo_tpr, hi_tpr,
  391. color=c['color'], alpha=0.10, linewidth=0)
  392. ax.plot(mean_fpr, mean_tpr, color=c['color'], lw=c['lw'], ls=c['ls'],
  393. label=(f"{name} (AUC = {m_auc:.3f}; "
  394. f"95% CI: {lo_auc:.3f}–{hi_auc:.3f})"))
  395. ax.set_xlim([0,1]); ax.set_ylim([0,1])
  396. ax.set_xlabel('False Positive Rate (1 − Specificity)', fontsize=11)
  397. ax.set_ylabel('True Positive Rate (Sensitivity)', fontsize=11)
  398. ax.legend(loc='lower right', fontsize=8.5, framealpha=0.95, edgecolor='#ccc')
  399. ax.grid(True, linestyle=':', alpha=0.3)
  400. for sp in ax.spines.values(): sp.set_linewidth(0.8); sp.set_color('#aaa')
  401. plt.tight_layout()
  402. plt.savefig(f'{output_dir}/figure1_roc.png', dpi=300,
  403. bbox_inches='tight', facecolor='white')
  404. plt.close()
  405. print("Figure 1 (ROC) saved.")
  406. def figure2_pr(pr_pooled, output_dir):
  407. y_true = np.array(pr_pooled['y_true'])
  408. y_prob_pos = np.array(pr_pooled['y_prob_pos'])
  409. y_prob_neg = np.array(pr_pooled['y_prob_neg'])
  410. fig, ax = plt.subplots(figsize=(7, 6), dpi=150)
  411. fig.patch.set_facecolor('white'); ax.set_facecolor('white')
  412. # Class 1 (ME/CFS)
  413. prec1, rec1, _ = precision_recall_curve(y_true, y_prob_pos, pos_label=1)
  414. ap1 = average_precision_score(y_true, y_prob_pos)
  415. ax.plot(rec1, prec1, color='#2ca02c', lw=2.0, ls='-',
  416. label=f'Precision-recall curve of class 1 (area = {ap1:.3f})')
  417. # Class 0 (Control)
  418. prec0, rec0, _ = precision_recall_curve(1-y_true, y_prob_neg, pos_label=1)
  419. ap0 = average_precision_score(1-y_true, y_prob_neg)
  420. ax.plot(rec0, prec0, color='black', lw=2.0, ls='-',
  421. label=f'Precision-recall curve of class 0 (area = {ap0:.3f})')
  422. # Micro-average (pooled)
  423. y_bin = label_binarize(y_true, classes=[0, 1])
  424. y_score = np.column_stack([y_prob_neg, y_prob_pos])
  425. prec_m, rec_m, _ = precision_recall_curve(y_bin.ravel(), y_score.ravel())
  426. ap_micro = average_precision_score(y_bin, y_score, average='micro')
  427. ax.plot(rec_m, prec_m, color='#1f77b4', lw=2.0, ls=':',
  428. label=f'micro-average Precision-recall curve (area = {ap_micro:.3f})')
  429. ax.set_xlim([0,1]); ax.set_ylim([0,1.02])
  430. ax.set_xlabel('Recall', fontsize=11)
  431. ax.set_ylabel('Precision', fontsize=11)
  432. ax.set_title('Precision-Recall Curve', fontsize=12)
  433. ax.legend(loc='lower left', fontsize=9, framealpha=0.95, edgecolor='#ccc')
  434. ax.grid(True, linestyle=':', alpha=0.3)
  435. for sp in ax.spines.values(): sp.set_linewidth(0.8); sp.set_color('#aaa')
  436. plt.tight_layout()
  437. plt.savefig(f'{output_dir}/figure2_pr.png', dpi=300,
  438. bbox_inches='tight', facecolor='white')
  439. plt.close()
  440. print(f"Figure 2 (PR) saved. Class1={ap1:.3f}, Class0={ap0:.3f}, Micro={ap_micro:.3f}")
  441. def figure3_importance(names_int, scores_int, names_ni, scores_ni,
  442. auc_int, auc_ni, output_dir):
  443. fig, axes = plt.subplots(1, 2, figsize=(22, 9))
  444. fig.patch.set_facecolor('white')
  445. BG='#EEF2F7'; BAR='#FF8C00'
  446. for ax, names, scores, label, auc_t in [
  447. (axes[0], names_int, scores_int, '(A)',
  448. f"AUC={auc_int[0]:.3f} (95% CI: {auc_int[1]:.3f}–{auc_int[2]:.3f})"),
  449. (axes[1], names_ni, scores_ni, '(B)',
  450. f"AUC={auc_ni[0]:.3f} (95% CI: {auc_ni[1]:.3f}–{auc_ni[2]:.3f})"),
  451. ]:
  452. ax.set_facecolor(BG)
  453. y_pos = np.arange(len(names))
  454. ax.barh(y_pos, scores[::-1], color=BAR, height=0.65,
  455. edgecolor='white', linewidth=0.4)
  456. ax.set_yticks(y_pos); ax.set_yticklabels(names[::-1], fontsize=12)
  457. ax.set_xlabel('Mean Absolute Score (Weighted)', fontsize=12)
  458. ax.text(-0.02, 1.03, label, transform=ax.transAxes,
  459. fontsize=15, fontweight='bold', va='bottom')
  460. ax.text(0.5, -0.07, auc_t, transform=ax.transAxes,
  461. ha='center', fontsize=10, style='italic', color='#444')
  462. ax.grid(axis='x', linestyle='--', alpha=0.35, color='white')
  463. ax.spines['top'].set_visible(False)
  464. ax.spines['right'].set_visible(False)
  465. ax.spines['left'].set_visible(False)
  466. ax.tick_params(left=False)
  467. ax.set_xlim(0, max(scores) * 1.08)
  468. plt.tight_layout(pad=3.0); plt.subplots_adjust(bottom=0.12)
  469. plt.savefig(f'{output_dir}/figure3_importance.png', dpi=300,
  470. bbox_inches='tight', facecolor='white')
  471. plt.close()
  472. print("Figure 3 (Importance) saved.")
  473. def figure4_local(local_cands, output_dir, top_n=15):
  474. cases = [
  475. ('tp', 'A', 'True Positive'),
  476. ('fp', 'B', 'False Positive'),
  477. ('fn', 'C', 'False Negative'),
  478. ]
  479. fig, axes = plt.subplots(1, 3, figsize=(21, 7))
  480. fig.patch.set_facecolor('white')
  481. for ax, (key, panel, label) in zip(axes, cases):
  482. case = local_cands.get(key)
  483. if case is None:
  484. ax.text(0.5, 0.5, f'No {label} found',
  485. ha='center', va='center', transform=ax.transAxes)
  486. continue
  487. ebm = case['ebm']
  488. x = case['x']
  489. prob = case['prob']
  490. y_true = case['y_true']
  491. y_pred = case['y_pred']
  492. # EBM local explanation
  493. expl = ebm.explain_local(x.reshape(1,-1), np.array([y_true]))
  494. scores_raw = expl.data(0)['scores']
  495. names_raw = expl.data(0)['names']
  496. abs_scores = np.abs(scores_raw)
  497. top_idx = np.argsort(abs_scores)[-top_n:]
  498. top_scores = np.array(scores_raw)[top_idx]
  499. top_names = np.array(names_raw)[top_idx]
  500. colors = ['#FF8C00' if s > 0 else '#1f77b4' for s in top_scores]
  501. ax.set_facecolor('#EEF2F7')
  502. y_pos = np.arange(len(top_names))
  503. ax.barh(y_pos, top_scores, color=colors, height=0.65,
  504. edgecolor='white', linewidth=0.3)
  505. ax.axvline(x=0, color='gray', lw=0.8, alpha=0.7)
  506. ax.set_yticks(y_pos); ax.set_yticklabels(top_names, fontsize=8.5)
  507. ax.set_xlabel('Contribution to Prediction', fontsize=10)
  508. pred_str = 'ME/CFS' if y_pred == 1 else 'Control'
  509. true_str = 'ME/CFS' if y_true == 1 else 'Control'
  510. prob_str = (f"Pr(y=1) = {prob:.3f}" if y_pred == 1
  511. else f"Pr(y=0) = {1-prob:.3f}")
  512. ax.set_title(f"Local Explanation (Actual: {true_str} | "
  513. f"Predicted: {pred_str}\n{prob_str})", fontsize=9, pad=6)
  514. ax.text(-0.02, 1.05, f'({panel})', transform=ax.transAxes,
  515. fontsize=13, fontweight='bold', va='bottom')
  516. ax.spines['top'].set_visible(False)
  517. ax.spines['right'].set_visible(False)
  518. ax.tick_params(left=False)
  519. ax.grid(axis='x', linestyle='--', alpha=0.3, color='white')
  520. plt.tight_layout(pad=2.5)
  521. plt.savefig(f'{output_dir}/figure4_local.png', dpi=300,
  522. bbox_inches='tight', facecolor='white')
  523. plt.close()
  524. print("Figure 4 (Local) saved.")
  525. # ============================================================
  526. # BÖLÜM 9: TANIMLAYICI İSTATİSTİK
  527. # ============================================================
  528. def descriptive_stats(df):
  529. mecfs = df[df['factor'] == 1]
  530. control = df[df['factor'] == 0]
  531. print("\n" + "="*60 + "\nTABLE 1: Descriptive Statistics\n" + "="*60)
  532. for label, col in [('Age','age'), ('BMI','bmi')]:
  533. if col not in df.columns: continue
  534. _, p = mannwhitneyu(mecfs[col].dropna(), control[col].dropna(),
  535. alternative='two-sided')
  536. print(f"{label}: ME/CFS={mecfs[col].mean():.1f}±{mecfs[col].std():.1f}, "
  537. f"Control={control[col].mean():.1f}±{control[col].std():.1f}, p={p:.3f} *")
  538. for col in ['ibs','probiotic_use','antidepressant_use','narcotic_use']:
  539. if col not in df.columns: continue
  540. ct = pd.crosstab(df['factor'], df[col])
  541. if ct.shape == (2,2):
  542. _, p = fisher_exact(ct.values)
  543. n1,n2 = mecfs[col].sum(), control[col].sum()
  544. print(f"{col}: ME/CFS={n1}({100*n1/len(mecfs):.1f}%), "
  545. f"Control={n2}({100*n2/len(control):.1f}%), p={p:.3f} **")
  546. print("* Mann-Whitney U ** Fisher's exact")
  547. # ============================================================
  548. # BÖLÜM 10: MAIN
  549. # ============================================================
  550. if __name__ == '__main__':
  551. DATA_PATH = 'df.xlsx' # <-- veri dosyasının yolu
  552. OUTPUT_DIR = 'results'
  553. N_ITER = 50
  554. # 1. Veri yükle
  555. X, y, feature_names = load_data(DATA_PATH)
  556. # 2. Ana pipeline
  557. (store, ebm_imp_dict, noint_imp_dict, sel_counts,
  558. mean_fpr, mean_recall, pr_pooled, local_cands) = run_50repeat(
  559. X, y, feature_names, n_iter=N_ITER, output_dir=OUTPUT_DIR
  560. )
  561. # 3. Performans özeti
  562. summary = summarise(store)
  563. print_table2(summary)
  564. # 4. PRNN seçim istatistikleri
  565. n_feat = len(feature_names)
  566. print(f"\nPRNN selected: {min(sel_counts)}–{max(sel_counts)} features "
  567. f"(mean={np.mean(sel_counts):.1f}, "
  568. f"~{np.mean(sel_counts)/n_feat*100:.1f}% of {n_feat})")
  569. n_mean = int(np.mean(sel_counts))
  570. print(f"Mean possible pairs: C({n_mean},2) = {n_mean*(n_mean-1)//2}")
  571. print(f"Unique EBM terms across 50 iters: {len(ebm_imp_dict)}")
  572. # 5. Feature importance
  573. names_int, scores_int = get_top_importances(ebm_imp_dict)
  574. names_ni, scores_ni = get_top_importances(noint_imp_dict)
  575. # 6. Görseller
  576. figure1_roc(store, mean_fpr, OUTPUT_DIR)
  577. figure2_pr(pr_pooled, OUTPUT_DIR)
  578. ai = summary['EBM']['AUC']
  579. an = summary['EBM_NoInt']['AUC']
  580. figure3_importance(
  581. names_int, scores_int, names_ni, scores_ni,
  582. (ai['mean'], ai['ci_lo'], ai['ci_hi']),
  583. (an['mean'], an['ci_lo'], an['ci_hi']),
  584. OUTPUT_DIR
  585. )
  586. figure4_local(local_cands, OUTPUT_DIR)
  587. print(f"\nTüm çıktılar → {OUTPUT_DIR}/")

code.py at commit 0ffe3c3, no license · at the source

Overview

  1. Department of Biostatistics, Faculty of Medicine, Malatya Turgut Ozal University, Malatya 44210, Türkiye
  2. Department of Family Medicine, Faculty of Medicine, Malatya Turgut Ozal University, Malatya 44210, Türkiye
  3. Department of Biostatistics and Medical Informatics, Faculty of Medicine, Inonu University, Malatya 44280, Türkiye
  4. Department of Computer Sciences, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, P.O. Box 84428, Riyadh 11671, Saudi Arabia
  5. Department of Physiology, College of Medicine, King Khalid University, Abha 61421, Saudi Arabia
  6. Department of Ocean Operations and Civil Engineering, Norwegian University of Science and Technology (NTNU), 6009 Ålesund, Norway
  7. Department of Sustainable Systems Engineering (INATECH), Albert Ludwigs University of Freiburg, 79110 Freiburg, Germany
Journal: International journal of molecular sciences, volume 27, issue 13, article 5920
Dates: received 10 May 2026; accepted 27 June 2026; published online 30 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/ijms27135920 · PMID 42450188 · PMCID PMC13362375 · OpenAlex W7166637608
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), clinical / translational (subfield)
Methods: Statistics, Machine learning, Preprocessing, Physiology & signal measures
Keywords: ME/CFS, omics, explainable boosting machine, ensemble learning, feature selection, PRNN
MeSH: Fatigue Syndrome, Chronic*, Metabolome*, Metabolomics*, Biomarkers, Boosting Machine Learning Algorithms, Classification Algorithms, Female, Humans, Male (* major topic)
Topic: Fibromyalgia and Chronic Fatigue Syndrome Research (Psychiatry and Mental health, Medicine), according to OpenAlex
Funding: Princess Nourah Bint Abdulrahman University (PNURSP2026R716)
Citations: not cited yet (Europe PMC); 38 references in the paper

Abstract

Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a debilitating multisystem illness characterised by post-exertional malaise, non-restorative sleep, and cognitive impairment, yet no objective diagnostic biomarkers have been established. Untargeted plasma metabolomics provides a broad view of the biochemical disturbances underlying ME/CFS; however, the high dimensionality of omics datasets and the limited interpretability of conventional classifiers nevertheless hinder translation into clinical practice. This study evaluates three ensemble classifiers—Explainable Boosting Machine (EBM), XGBoost, and LightGBM—for binary ME/CFS classification using plasma metabolomic and lipidomic profiles from 197 participants (106 ME/CFS; 91 healthy controls; 888 features). Feature dimensionality was reduced using a Pareto-Guided Recursive Neural Network (PRNN) pipeline. Model performance was assessed via 50-repeat stratified hold-out validation. EBM achieved the highest accuracy (0.909; 95% CI: 0.868–0.949) and area under the receiver operating characteristic curve (AUC: 0.940; 95% CI: 0.909–0.983), with XGBoost and LightGBM performing comparably. Interpretability analyses revealed that pairwise metabolite interaction terms—particularly proline & indole-3-lactate, tyrosine & N-acetylornithine, and maleic acid & arachidic acid—contributed the greatest discriminative signal. An ablation analysis comparing the full interaction-augmented EBM (AUC = 0.940) with a main-effects-only EBM (AUC = 0.882) confirmed that pairwise metabolite co-variation contributes additional discriminative value beyond individual metabolite levels, implicating amino acid catabolism, tryptophan–kynurenine pathway dysregulation, mitochondrial energy impairment, and lipid remodelling as central pathophysiological features. Global and instance-level explanations jointly demonstrated population-level metabolic signatures alongside individual heterogeneity, highlighting the added clinical value of explainable artificial intelligence (XAI) in metabolomics. These findings support EBM-based metabolomic profiling as an internally validated approach for ME/CFS classification, subject to external validation, calibration assessment, and prospective testing.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

drhilal/ME-CFS-study

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 0ffe3c328255b3b48fc3076619a581fddebf7e17, 11 June 2026
Languages: Python (1)
Size: 1 file, 1 script
Software Heritage: not archived
Found in: the text, “4.10. Software and Computational Environment”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: LightGBM (1 file), Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
1 file
  • code.py, Python, 716 lines, 5 matches

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 6 keywords, 9 MeSH terms, 1 funder, 34 references.

Cite

This paper

Yagin, F. H., Korkmaz, Y., Colak, C., Alzakari, S. A., Alkhalifa, A. K., Al-Hashem, F., & Aghaei, M. (2026). Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection. International journal of molecular sciences, 27(13), 5920. https://doi.org/10.3390/ijms27135920

BibTeX

@article{yagin2026metabolomic,
author = {Yagin, Fatma Hilal and Korkmaz, Yavuz and Colak, Cemil and Alzakari, Sarah A and Alkhalifa, Amal K and Al-Hashem, Fahaid and Aghaei, Mohammadreza},
title = {{Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection}},
journal = {International journal of molecular sciences},
year = {2026},
month = jun,
volume = {27},
number = {13},
pages = {5920},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1422-0067},
doi = {10.3390/ijms27135920},
url = {https://doi.org/10.3390/ijms27135920},
pmid = {42450188},
pmcid = {PMC13362375}
}

RIS

TY - JOUR
AU - Yagin, Fatma Hilal
AU - Korkmaz, Yavuz
AU - Colak, Cemil
AU - Alzakari, Sarah A
AU - Alkhalifa, Amal K
AU - Al-Hashem, Fahaid
AU - Aghaei, Mohammadreza
TI - Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection
T2 - International journal of molecular sciences
J2 - Int J Mol Sci
PY - 2026
DA - 2026/06/30
VL - 27
IS - 13
SP - 5920
SN - 1422-0067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/ijms27135920
UR - https://doi.org/10.3390/ijms27135920
LA - en
ER -

CSL-JSON

{
"id": "10.3390/ijms27135920",
"type": "article-journal",
"title": "Metabolomic Classification of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome via Explainable Ensemble Learning and Pareto-Guided Feature Selection",
"container-title": "International journal of molecular sciences",
"author": [
{
"family": "Yagin",
"given": "Fatma Hilal"
},
{
"family": "Korkmaz",
"given": "Yavuz"
},
{
"family": "Colak",
"given": "Cemil"
},
{
"family": "Alzakari",
"given": "Sarah A"
},
{
"family": "Alkhalifa",
"given": "Amal K"
},
{
"family": "Al-Hashem",
"given": "Fahaid"
},
{
"family": "Aghaei",
"given": "Mohammadreza"
}
],
"container-title-short": "Int J Mol Sci",
"volume": "27",
"issue": "13",
"page": "5920",
"DOI": "10.3390/ijms27135920",
"PMID": "42450188",
"PMCID": "PMC13362375",
"ISSN": "1422-0067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/ijms27135920",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
30
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3389/fnhum.2026.1869918 [code]
Single-subject auditory ERP-BCI performance enhancement in ALS via an AI coding assistant prompt.
Journal: Frontiers in human neuroscience
In common: LightGBM, XGBoost, scikit-learn, 4 other tools, clinical / translational
[2] doi:10.1073/pnas.2516601123 [code]
Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: LightGBM, XGBoost, scikit-learn, 4 other tools, clinical / translational
[3] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: LightGBM, XGBoost, scikit-learn, 4 other tools
[4] doi:10.1186/s13244-026-02365-7 [code]
Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion.
Journal: Insights into imaging
In common: LightGBM, XGBoost, scikit-learn, 4 other tools
[5] doi:10.1371/journal.pone.0337726 [code]
Classification of chronic pain and spinal cord stimulation response using machine learning in magnetoencephalography data
Journal: n/a
In common: LightGBM, XGBoost, scikit-learn, 4 other tools
[6] doi:10.1371/journal.pgen.1012242 [code]
Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.
Journal: PLoS genetics
In common: LightGBM, XGBoost, scikit-learn, 4 other tools
[7] doi:10.3389/fpsyg.2026.1774068 [code]
Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.
Journal: Frontiers in psychology
In common: LightGBM, XGBoost, scikit-learn, 4 other tools
[8] doi:10.1038/s41591-026-04432-4 [code]
Activity-dependent adaptive deep brain stimulation improves gait in Parkinson's disease.
Journal: Nature medicine
In common: LightGBM, XGBoost, scikit-learn, 3 other tools, clinical / translational
[9] doi:10.1371/journal.pone.0354510 [code]
Personalized adaptive virtual reality experience driven by electroencephalography-based pain recognition.
Journal: PloS one
In common: LightGBM, XGBoost, scikit-learn, 3 other tools
[10] doi:10.1212/wnl.0000000000218472 [code]
Lesion-Level Subtypes of White Matter Hyperintensity Evolution Beyond Spatial Location.
Journal: Neurology
In common: LightGBM, scikit-learn, pandas, 3 other tools, clinical / translational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.