OSCR

CT-based radiomic markers to predict late-onset seizures after traumatic brain injury.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § Methods › Model training ↔ NIFTItraining.py, lines 1–81 · score 0.94 · nested cross validation, training folds, variance thresholding, SelectKBest, feature selection, Logistic Regression
  2. [2] § Methods › NCCT processing ↔ NIFTIprocessing.py, lines 22–33 · score 0.68 · dcm2niix, DICOM images, NIfTI
  3. [3] § Methods › Feature extraction ↔ NIFTItraining.py, lines 1113–1159 · score 0.52 · social determinant, integer, sex, encoded

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,552 lines · 64 KB · no license · 2 matches

  1. """
  2. NIFTItrainingNestedCV.py
  3. -----------------------
  4. This script implements a robust, modular machine learning pipeline for tabular data (e.g., radiomics features) with support for nested cross-validation (nested CV) for unbiased model selection and evaluation.
  5. Key Features:
  6. - Modular pipeline classes for Random Forest and Logistic Regression, easily extensible to other models.
  7. - Automatic feature selection (SelectKBest) and preprocessing (imputation, scaling, variance thresholding).
  8. - Hyperparameter tuning using GridSearchCV or RandomizedSearchCV.
  9. - Nested cross-validation: outer loop for unbiased test evaluation, inner loop for hyperparameter search.
  10. - Centralized saving and reporting of all metrics, including nested CV results.
  11. Usage:
  12. - Place your patient feature CSV in the expected location (see FEATURES_CSV).
  13. - Customize the models list in main() to add/remove models.
  14. - Run the script. Results and metrics will be saved in the trainingML/ subdirectory.
  15. Class/Function Documentation:
  16. -----------------------------
  17. BaseMLPipeline:
  18. - Base class for ML pipelines. Handles pipeline construction, feature selection, and search_fit (hyperparameter tuning).
  19. - Methods:
  20. - __init__: Sets up the pipeline steps and estimator.
  21. - search_fit: Runs GridSearchCV or RandomizedSearchCV for hyperparameter tuning. Stores best estimator and metrics.
  22. - save: Saves the trained pipeline and all metrics, including nested CV results if provided.
  23. RandomForestMLPipeline, LogisticRegressionMLPipeline:
  24. - Subclasses of BaseMLPipeline with default estimators and parameter grids for each model type.
  25. nested_cv_evaluate:
  26. - Runs nested cross-validation for a given pipeline and dataset.
  27. - Outer loop: splits data into train/test folds for unbiased evaluation.
  28. - Inner loop: runs hyperparameter search (search_fit) on the training fold.
  29. - Returns test accuracies and best parameters for each outer fold.
  30. main:
  31. - Loads data and defines models.
  32. - Runs nested CV for each model, prints and saves results using the pipeline's save method.
  33. Pipeline Steps:
  34. - Imputation (median), scaling (StandardScaler), variance thresholding, feature selection (SelectKBest), classifier (clf).
  35. - 'clf' is the final classifier step (e.g., RandomForestClassifier or LogisticRegression).
  36. Outputs:
  37. - For each model, saves a metrics file with:
  38. - Best parameters (if available)
  39. - Accuracy, confusion matrix, classification report, top features (if available)
  40. - Nested CV results: mean and per-fold test accuracy, best params per fold
  41. Best Practices:
  42. - Nested CV provides an unbiased estimate of model performance, especially important for small datasets.
  43. - All reporting and saving is centralized for reproducibility and clarity.
  44. """
  45. import os
  46. import numpy as np
  47. import pandas as pd
  48. import warnings
  49. import NIFTIpatient
  50. from sklearn.model_selection import (
  51. GridSearchCV,
  52. RandomizedSearchCV,
  53. StratifiedKFold,
  54. )
  55. from sklearn.metrics import classification_report, confusion_matrix, accuracy_score
  56. from sklearn.pipeline import Pipeline
  57. from sklearn.impute import SimpleImputer
  58. from sklearn.preprocessing import StandardScaler
  59. from sklearn.compose import ColumnTransformer
  60. from sklearn.feature_selection import VarianceThreshold, SelectKBest, f_classif
  61. from sklearn.linear_model import LogisticRegression
  62. from sklearn.ensemble import RandomForestClassifier
  63. import joblib
  64. import random
  65. from collections import Counter, defaultdict
  66. import xgboost as xgb # For future use
  67. from xgboost import XGBClassifier, XGBRFClassifier
  68. # For ROC curve plotting
  69. from sklearn.metrics import roc_curve, auc
  70. from sklearn.model_selection import train_test_split
  71. import matplotlib.pyplot as plt
  72. from sklearn.metrics import (
  73. precision_score,
  74. recall_score,
  75. f1_score,
  76. roc_auc_score,
  77. )
  78. # Filter out expected warnings from scikit-learn
  79. # This occurs during feature selection when features have zero variance or cause numerical issues
  80. warnings.filterwarnings(
  81. "ignore", category=RuntimeWarning, message="invalid value encountered in divide"
  82. )
  83. # This was about the data format, which we fixed but the warning filters provide an extra safety net
  84. warnings.filterwarnings(
  85. "ignore", category=UserWarning, message=".*column-vector y was passed.*"
  86. )
  87. # This occurs when a class has no predicted samples, which is common in small datasets
  88. warnings.filterwarnings(
  89. "ignore", category=UserWarning, message=".*Precision is ill-defined.*"
  90. )
  91. # Standalone class for plotting ROC curve for binary classifiers
  92. class ROCCurvePlotter:
  93. """
  94. Utility class to plot ROC curve for binary classifiers.
  95. Usage:
  96. plotter = ROCCurvePlotter(y_true, y_score, title="ROC Curve")
  97. plotter.plot()
  98. Args:
  99. y_true: array-like of shape (n_samples,) - True binary labels (0 or 1)
  100. y_score: array-like of shape (n_samples,) - Target scores, can be probability estimates of the positive class
  101. title: str - Title for the plot
  102. """
  103. def __init__(self, y_true, y_score, title="ROC Curve"):
  104. """
  105. Initialize with true labels and predicted scores.
  106. Args:
  107. y_true (array-like of shape (n_samples,): True binary labels (0 or 1)
  108. y_score (array-like of shape (n_samples,): Target scores, can be probability estimates of the positive class
  109. title (str): Title for the plot
  110. Returns:
  111. None
  112. """
  113. self.y_true = y_true
  114. self.y_score = y_score
  115. self.title = title
  116. self.fpr = None
  117. self.tpr = None
  118. self.thresholds = None
  119. self.roc_auc = None
  120. def compute(self):
  121. """
  122. Computes FPR, TPR, thresholds, and AUC.
  123. Args:
  124. None
  125. Returns:
  126. None
  127. """
  128. self.fpr, self.tpr, self.thresholds = roc_curve(self.y_true, self.y_score)
  129. self.roc_auc = auc(self.fpr, self.tpr)
  130. def plot(self, show=True, save_path=None):
  131. """
  132. Plots the ROC curve.
  133. Args:
  134. show (bool): If True, display the plot interactively.
  135. save_path (str or None): If provided, save the plot to this path.
  136. Returns:
  137. None
  138. """
  139. self.compute()
  140. plt.figure()
  141. plt.plot(
  142. self.fpr,
  143. self.tpr,
  144. color="darkorange",
  145. lw=2,
  146. label=f"ROC curve (AUC = {self.roc_auc:.2f})",
  147. )
  148. plt.plot([0, 1], [0, 1], color="navy", lw=2, linestyle="--")
  149. plt.xlim([0.0, 1.0])
  150. plt.ylim([0.0, 1.05])
  151. plt.xlabel("False Positive Rate")
  152. plt.ylabel("True Positive Rate")
  153. plt.title(self.title)
  154. plt.legend(loc="lower right")
  155. if save_path:
  156. plt.savefig(save_path)
  157. if show:
  158. plt.show()
  159. plt.close()
  160. @staticmethod
  161. def plot_nested_cv_roc_curves(
  162. roc_data, save_path=None, title="Nested CV ROC Curves", csv_path=None
  163. ):
  164. """
  165. Plot all ROC curves from nested CV runs on the same graph.
  166. Args:
  167. roc_data (list of tuples): list of (fpr, tpr, auc, fold) tuples
  168. save_path (str or None): if provided, save the plot to this path
  169. title (str): plot title
  170. csv_path (str or None): if provided, save all ROC data to this CSV for reproducibility
  171. Returns:
  172. None
  173. """
  174. plt.figure(figsize=(8, 6))
  175. aucs = []
  176. # Optionally save all ROC data to CSV for reproducibility
  177. if csv_path is not None:
  178. # Flatten all ROC data into a DataFrame
  179. rows = []
  180. for fpr, tpr, roc_auc, fold in roc_data:
  181. for f, t in zip(fpr, tpr):
  182. rows.append({"fold": fold, "fpr": f, "tpr": t, "auc": roc_auc})
  183. pd.DataFrame(rows).to_csv(csv_path, index=False)
  184. for fpr, tpr, roc_auc, fold in roc_data:
  185. plt.plot(fpr, tpr, lw=2, label=f"Fold {fold} (AUC = {roc_auc:.2f})")
  186. aucs.append(roc_auc)
  187. mean_auc = np.mean(aucs)
  188. std_auc = np.std(aucs, ddof=1)
  189. n = len(aucs)
  190. # 95% CI for mean: mean ± 1.96 * (std / sqrt(n))
  191. if n > 1:
  192. ci95 = 1.96 * (std_auc / np.sqrt(n))
  193. ci_low = mean_auc - ci95
  194. ci_high = mean_auc + ci95
  195. ci_str = f"95% CI: [{ci_low:.3f}, {ci_high:.3f}]"
  196. else:
  197. ci_str = ""
  198. plt.plot([0, 1], [0, 1], color="navy", lw=2, linestyle="--", label="Chance")
  199. plt.xlim([0.0, 1.0])
  200. plt.ylim([0.0, 1.05])
  201. plt.xlabel("False Positive Rate")
  202. plt.ylabel("True Positive Rate")
  203. title_str = f"{title}\nMean AUC = {mean_auc:.3f} ± {std_auc:.3f}"
  204. if ci_str:
  205. title_str += f"\n{ci_str}"
  206. plt.title(title_str)
  207. plt.legend(loc="lower right")
  208. if save_path:
  209. plt.savefig(save_path)
  210. # Do not show the plot interactively
  211. plt.close()
  212. @staticmethod
  213. def plot_holdout_roc_curves_across_seeds(
  214. roc_data, save_path=None, title="Holdout ROC Curves Across Seeds", csv_path=None
  215. ):
  216. """
  217. Plot all holdout ROC curves from different seeds for a single model.
  218. Args:
  219. roc_data (list of tuples): list of (fpr, tpr, auc, seed) tuples
  220. save_path (str or None): if provided, save the plot to this path
  221. title (str): plot title
  222. csv_path (str or None): if provided, save all ROC data to this CSV for reproducibility
  223. Returns:
  224. None
  225. """
  226. plt.figure(figsize=(8, 6))
  227. aucs = []
  228. # Optionally save all ROC data to CSV for reproducibility
  229. if csv_path is not None:
  230. rows = []
  231. for fpr, tpr, roc_auc, seed in roc_data:
  232. for f, t in zip(fpr, tpr):
  233. rows.append({"seed": seed, "fpr": f, "tpr": t, "auc": roc_auc})
  234. pd.DataFrame(rows).to_csv(csv_path, index=False)
  235. for fpr, tpr, roc_auc, seed in roc_data:
  236. plt.plot(fpr, tpr, lw=2, label=f"Seed {seed} (AUC = {roc_auc:.2f})")
  237. aucs.append(roc_auc)
  238. mean_auc = np.mean(aucs)
  239. std_auc = np.std(aucs, ddof=1)
  240. n = len(aucs)
  241. # 95% CI for mean: mean ± 1.96 * (std / sqrt(n))
  242. if n > 1:
  243. ci95 = 1.96 * (std_auc / np.sqrt(n))
  244. ci_low = mean_auc - ci95
  245. ci_high = mean_auc + ci95
  246. ci_str = f"95% CI: [{ci_low:.3f}, {ci_high:.3f}]"
  247. else:
  248. ci_str = ""
  249. plt.plot([0, 1], [0, 1], color="navy", lw=2, linestyle="--", label="Chance")
  250. plt.xlim([0.0, 1.0])
  251. plt.ylim([0.0, 1.05])
  252. plt.xlabel("False Positive Rate")
  253. plt.ylabel("True Positive Rate")
  254. title_str = f"{title}\nMean AUC = {mean_auc:.3f} ± {std_auc:.3f}"
  255. if ci_str:
  256. title_str += f"\n{ci_str}"
  257. plt.title(title_str)
  258. plt.legend(loc="lower right")
  259. if save_path:
  260. plt.savefig(save_path)
  261. plt.close()
  262. class BaseMLPipeline:
  263. def get_preprocessed_feature_names(self, pipeline, X):
  264. """
  265. Retrieve feature names after preprocessing, with fallback and prefix stripping.
  266. Args:
  267. pipeline (skLearn Pipeline object): The fitted pipeline.
  268. X (pandas DataFrame): Original input features.
  269. Returns:
  270. feature_names (list of str): List of feature names after preprocessing.
  271. """
  272. if "preprocess" in pipeline.named_steps:
  273. try:
  274. X_preprocessed = pipeline.named_steps["preprocess"].transform(X)
  275. if hasattr(pipeline.named_steps["preprocess"], "get_feature_names_out"):
  276. feature_names = pipeline.named_steps[
  277. "preprocess"
  278. ].get_feature_names_out()
  279. else:
  280. feature_names = [f"f{i}" for i in range(X_preprocessed.shape[1])]
  281. except Exception:
  282. feature_names = X.columns
  283. else:
  284. feature_names = X.columns
  285. # Strip 'remainder__' prefix
  286. feature_names = [f.replace("remainder__", "") for f in feature_names]
  287. return feature_names
  288. """
  289. BaseMLPipeline: A flexible base class for building and running ML pipelines on tabular data.
  290. Customize the pipeline steps, estimator, and feature selection as needed for your use case.
  291. """
  292. def __init__(
  293. self,
  294. n_features=None,
  295. estimator=None,
  296. random_state=None,
  297. ord_categories=None,
  298. ):
  299. """
  300. Initialize the ML pipeline with preprocessing, feature selection, and estimator.
  301. Args:
  302. n_features (int or None): Number of top features to select (SelectKBest). If None, 'all' is used.
  303. estimator (sklearn estimator): The final estimator (classifier/regressor) to use.
  304. random_state (int or None): Random state for reproducibility.
  305. ord_categories (dict or None): Categorical feature categories for ordinal encoding (if needed).
  306. Returns:
  307. None
  308. """
  309. # If n_features is None, we'll let GridSearch tune select__k and default SelectKBest to 'all'
  310. self.n_features = n_features
  311. self.random_state = random_state
  312. # Define categorical columns for mhhcohort data (excluding target variable 'pts')
  313. self.ord_categories = ord_categories
  314. # Treat all features as numeric (already encoded). Single median imputation.
  315. # IF CATEGORICAL FEATURES ARE PRESENT, modify this step to include OrdinalEncoder
  316. preprocessor = SimpleImputer(strategy="median")
  317. pipeline_steps = [
  318. ("preprocess", preprocessor),
  319. ("scale", StandardScaler()),
  320. ("var_thresh", VarianceThreshold(threshold=0.01)),
  321. (
  322. "select",
  323. SelectKBest(
  324. score_func=f_classif,
  325. k=(self.n_features if self.n_features is not None else "all"),
  326. ),
  327. ),
  328. ]
  329. self.estimator = estimator
  330. pipeline_steps.append(("clf", estimator))
  331. self.pipeline = Pipeline(pipeline_steps)
  332. self.selected_features_ = None
  333. self.importances_ = None
  334. self.metrics_ = None
  335. def search_fit(
  336. self,
  337. X,
  338. y,
  339. random_state,
  340. search_type="grid",
  341. param_grid=None,
  342. n_iter=10,
  343. cv=5,
  344. scoring="accuracy",
  345. ):
  346. """
  347. Perform hyperparameter search (GridSearchCV or RandomizedSearchCV) and fit the pipeline.
  348. Args:
  349. X (array-like of shape (n_samples, n_features)): Input features.
  350. y (array-like of shape (n_samples,)): Target labels.
  351. random_state (int): Random state for reproducibility.
  352. search_type (str): "grid" for GridSearchCV, "random" for RandomizedSearchCV.
  353. param_grid (dict): Parameter grid or distributions for search.
  354. n_iter (int): Number of iterations for RandomizedSearchCV (ignored for GridSearchCV).
  355. cv (int): Number of cross-validation folds.
  356. scoring (str): Scoring metric for evaluation.
  357. Returns:
  358. search (sklearn GridSearchCV or RandomizedSearchCV object): Fitted search object.
  359. """
  360. if search_type == "grid":
  361. search = GridSearchCV(
  362. self.pipeline, param_grid, cv=cv, scoring=scoring, n_jobs=-1
  363. )
  364. elif search_type == "random":
  365. search = RandomizedSearchCV(
  366. self.pipeline,
  367. param_distributions=param_grid,
  368. n_iter=n_iter,
  369. cv=cv,
  370. scoring=scoring,
  371. random_state=random_state,
  372. n_jobs=-1,
  373. )
  374. else:
  375. raise ValueError("search_type must be 'grid' or 'random'")
  376. search.fit(X, y)
  377. best_pipeline = search.best_estimator_
  378. y_pred = best_pipeline.predict(X)
  379. cm = confusion_matrix(y, y_pred)
  380. # Specificity: TN / (TN + FP)
  381. if cm.shape == (2, 2):
  382. tn, fp, fn, tp = cm.ravel()
  383. specificity = tn / (tn + fp) if (tn + fp) > 0 else np.nan
  384. else:
  385. specificity = np.nan
  386. # AUC: only if y has both classes and proba/decision available
  387. try:
  388. if hasattr(best_pipeline.named_steps["clf"], "predict_proba"):
  389. y_score = best_pipeline.named_steps["clf"].predict_proba(X)[:, 1]
  390. else:
  391. y_score = best_pipeline.named_steps["clf"].decision_function(X)
  392. auc_val = roc_auc_score(y, y_score)
  393. except Exception:
  394. auc_val = np.nan
  395. metrics = {
  396. "best_params": search.best_params_,
  397. "accuracy": accuracy_score(y, y_pred),
  398. "precision": precision_score(y, y_pred, zero_division=0),
  399. "recall": recall_score(y, y_pred, zero_division=0),
  400. "specificity": specificity,
  401. "f1": f1_score(y, y_pred, zero_division=0),
  402. "auc": auc_val,
  403. "confusion_matrix": cm,
  404. "classification_report": classification_report(y, y_pred),
  405. }
  406. if (
  407. hasattr(best_pipeline.named_steps["clf"], "feature_importances_")
  408. and "select" in best_pipeline.named_steps
  409. ):
  410. # Get feature names after preprocessing
  411. preprocessed_feature_names = self.get_preprocessed_feature_names(
  412. best_pipeline, X
  413. )
  414. # Get the feature names after variance threshold but before selection
  415. if "var_thresh" in best_pipeline.named_steps:
  416. var_thresh_mask = best_pipeline.named_steps["var_thresh"].get_support()
  417. features_after_var_thresh = np.array(preprocessed_feature_names)[
  418. var_thresh_mask
  419. ]
  420. else:
  421. features_after_var_thresh = preprocessed_feature_names
  422. # Get the selection mask and apply it to the reduced feature set
  423. selected_mask = best_pipeline.named_steps["select"].get_support()
  424. selected_features = features_after_var_thresh[selected_mask]
  425. importances = best_pipeline.named_steps["clf"].feature_importances_
  426. importances_series = pd.Series(importances, index=selected_features)
  427. metrics["top_features"] = importances_series.sort_values(ascending=False)
  428. elif (
  429. hasattr(best_pipeline.named_steps["clf"], "coef_")
  430. and "select" in best_pipeline.named_steps
  431. ):
  432. # Get feature names after preprocessing
  433. preprocessed_feature_names = self.get_preprocessed_feature_names(
  434. best_pipeline, X
  435. )
  436. # Get the feature names after variance threshold but before selection
  437. if "var_thresh" in best_pipeline.named_steps:
  438. var_thresh_mask = best_pipeline.named_steps["var_thresh"].get_support()
  439. features_after_var_thresh = np.array(preprocessed_feature_names)[
  440. var_thresh_mask
  441. ]
  442. else:
  443. features_after_var_thresh = preprocessed_feature_names
  444. # Get the selection mask and apply it to the reduced feature set
  445. selected_mask = best_pipeline.named_steps["select"].get_support()
  446. selected_features = features_after_var_thresh[selected_mask]
  447. # Use absolute values of coefficients for importance
  448. coefs = np.abs(best_pipeline.named_steps["clf"].coef_[0])
  449. importances_series = pd.Series(coefs, index=selected_features)
  450. metrics["top_features"] = importances_series.sort_values(ascending=False)
  451. # Store signed coefficients for direction interpretation (positive = increases PTS odds, negative = decreases)
  452. signed_coefs = best_pipeline.named_steps["clf"].coef_[0]
  453. metrics["signed_coefficients"] = pd.Series(signed_coefs, index=selected_features)
  454. self.metrics_ = metrics
  455. self.pipeline = best_pipeline
  456. return search
  457. def save(self, model_path=None, metrics_path=None, nested_cv_results=None):
  458. """
  459. Save the trained pipeline and comprehensive metrics report.
  460. Args:
  461. model_path (str or None): If provided, save the trained pipeline to this path.
  462. metrics_path (str or None): If provided, save the comprehensive metrics report to this path.
  463. nested_cv_results (tuple or None): If provided, should be (test_scores, best_params_list, model_name)
  464. Returns:
  465. None
  466. """
  467. if model_path is not None:
  468. joblib.dump(self.pipeline, model_path)
  469. if metrics_path is not None:
  470. with open(metrics_path, "w") as f:
  471. f.write("ML Pipeline Comprehensive Report\n")
  472. f.write("=" * 50 + "\n\n")
  473. # Nested CV Results (Unbiased Performance Estimate)
  474. if nested_cv_results is not None:
  475. test_scores, best_params_list, model_name = nested_cv_results
  476. f.write("NESTED CROSS-VALIDATION RESULTS (Unbiased Performance)\n")
  477. f.write("-" * 55 + "\n")
  478. f.write(
  479. f"Mean test accuracy: {np.mean(test_scores):.3f} ± {np.std(test_scores):.3f}\n"
  480. )
  481. f.write(f"All test accuracies: {test_scores}\n")
  482. f.write(f"Best hyperparameters per fold:\n")
  483. for i, params in enumerate(best_params_list):
  484. f.write(f" Fold {i+1}: {params}\n")
  485. f.write("\n")
  486. # Final / Hold-out Results (Feature Interpretation)
  487. if self.metrics_ is not None:
  488. f.write("FINAL MODEL RESULTS\n")
  489. f.write("-" * 45 + "\n")
  490. if "best_params" in self.metrics_:
  491. f.write(
  492. f"Best hyperparameters: {self.metrics_['best_params']}\n"
  493. )
  494. # Training metrics (if provided)
  495. if any(k.startswith("train_") for k in self.metrics_.keys()):
  496. f.write("\nTRAINING SET PERFORMANCE\n")
  497. f.write("~" * 30 + "\n")
  498. if self.metrics_.get("train_accuracy") is not None:
  499. f.write(
  500. f"Accuracy: {self.metrics_['train_accuracy']:.3f}\n"
  501. )
  502. if self.metrics_.get("train_confusion_matrix") is not None:
  503. f.write("Confusion Matrix:\n")
  504. f.write(str(self.metrics_["train_confusion_matrix"]) + "\n")
  505. if self.metrics_.get("train_classification_report") is not None:
  506. f.write("Classification Report:\n")
  507. f.write(self.metrics_["train_classification_report"] + "\n")
  508. # Add extra metrics if available
  509. for k in [
  510. "train_precision",
  511. "train_recall",
  512. "train_specificity",
  513. "train_f1",
  514. "train_auc",
  515. ]:
  516. if self.metrics_.get(k) is not None:
  517. f.write(
  518. f"{k.replace('train_','').capitalize()}: {self.metrics_[k]:.3f}\n"
  519. )
  520. else:
  521. # Backward compatibility: older single-set metrics
  522. if self.metrics_.get("accuracy") is not None:
  523. f.write(
  524. f"Training accuracy: {self.metrics_.get('accuracy'):.3f}\n"
  525. )
  526. if self.metrics_.get("precision") is not None:
  527. f.write(
  528. f"Precision: {self.metrics_.get('precision'):.3f}\n"
  529. )
  530. if self.metrics_.get("recall") is not None:
  531. f.write(f"Recall: {self.metrics_.get('recall'):.3f}\n")
  532. if self.metrics_.get("specificity") is not None:
  533. f.write(
  534. f"Specificity: {self.metrics_.get('specificity'):.3f}\n"
  535. )
  536. if self.metrics_.get("f1") is not None:
  537. f.write(f"F1 score: {self.metrics_.get('f1'):.3f}\n")
  538. if self.metrics_.get("auc") is not None:
  539. f.write(f"AUC: {self.metrics_.get('auc'):.3f}\n")
  540. if self.metrics_.get("confusion_matrix") is not None:
  541. f.write("Confusion Matrix (Training Data):\n")
  542. f.write(str(self.metrics_.get("confusion_matrix")) + "\n")
  543. if self.metrics_.get("classification_report") is not None:
  544. f.write("Classification Report (Training Data):\n")
  545. f.write(self.metrics_.get("classification_report") + "\n")
  546. # Test metrics (hold-out set)
  547. if any(k.startswith("test_") for k in self.metrics_.keys()):
  548. f.write("\nHOLD-OUT TEST SET PERFORMANCE (20%)\n")
  549. f.write("~" * 45 + "\n")
  550. if self.metrics_.get("test_accuracy") is not None:
  551. f.write(f"Accuracy: {self.metrics_['test_accuracy']:.3f}\n")
  552. if self.metrics_.get("test_precision") is not None:
  553. f.write(
  554. f"Precision: {self.metrics_['test_precision']:.3f}\n"
  555. )
  556. if self.metrics_.get("test_recall") is not None:
  557. f.write(f"Recall: {self.metrics_['test_recall']:.3f}\n")
  558. if self.metrics_.get("test_specificity") is not None:
  559. f.write(
  560. f"Specificity: {self.metrics_['test_specificity']:.3f}\n"
  561. )
  562. if self.metrics_.get("test_f1") is not None:
  563. f.write(f"F1 score: {self.metrics_['test_f1']:.3f}\n")
  564. if self.metrics_.get("test_auc") is not None:
  565. f.write(f"ROC AUC: {self.metrics_['test_auc']:.3f}\n")
  566. if self.metrics_.get("test_confusion_matrix") is not None:
  567. f.write("Confusion Matrix:\n")
  568. f.write(str(self.metrics_["test_confusion_matrix"]) + "\n")
  569. if self.metrics_.get("test_classification_report") is not None:
  570. f.write("Classification Report:\n")
  571. f.write(self.metrics_["test_classification_report"] + "\n")
  572. # Feature importance / coefficients
  573. if self.metrics_.get("top_features", None) is not None:
  574. all_feats_series = self.metrics_["top_features"]
  575. signed_coefs = self.metrics_.get("signed_coefficients", None)
  576. # Preview
  577. f.write("\nTOP FEATURES PREVIEW (first 20 shown):\n")
  578. f.write("-" * 45 + "\n")
  579. preview = all_feats_series.head(20)
  580. f.write(preview.to_string() + "\n\n")
  581. # Full list with direction for signed coefficients
  582. f.write("FULL RANKED FEATURE LIST:\n")
  583. f.write("-" * 45 + "\n")
  584. if signed_coefs is not None:
  585. f.write("Feature, Importance (abs), Coefficient (signed), Direction\n")
  586. f.write("-" * 45 + "\n")
  587. for feature, importance in all_feats_series.items():
  588. coef = signed_coefs[feature]
  589. direction = "INCREASES PTS" if coef > 0 else "DECREASES PTS"
  590. f.write(f"{feature:60s} | {importance:10.6f} | {coef:10.6f} | {direction}\n")
  591. f.write("\n")
  592. else:
  593. f.write(all_feats_series.to_string() + "\n\n")
  594. selected_features = all_feats_series.index.tolist()
  595. f.write("ALL SELECTED FEATURES (ordered):\n")
  596. f.write("-" * 45 + "\n")
  597. for i, feature in enumerate(selected_features, 1):
  598. if signed_coefs is not None:
  599. coef = signed_coefs[feature]
  600. direction = "[+]" if coef > 0 else "[-]"
  601. f.write(f"{i:3d}. {feature:55s} {direction}\n")
  602. else:
  603. f.write(f"{i:3d}. {feature}\n")
  604. if model_path is not None:
  605. print(f"Final model saved to: {model_path}")
  606. if metrics_path is not None:
  607. print(f"Comprehensive metrics saved to: {metrics_path}")
  608. class RandomForestMLPipeline(BaseMLPipeline):
  609. def __init__(
  610. self, n_features=None, random_state=None, ord_categories=None, param_grid=None
  611. ):
  612. """
  613. Initialize RandomForestMLPipeline with default RandomForestClassifier and parameter grid.
  614. Args:
  615. n_features (int or None): Number of top features to select (SelectKBest). If None, 'all' is used.
  616. random_state (int or None): Random state for reproducibility.
  617. ord_categories (dict or None): Categorical feature categories for ordinal encoding (if needed).
  618. param_grid (dict or None): Parameter grid for hyperparameter search. If None, default grid is used.
  619. Returns:
  620. None
  621. """
  622. super().__init__(
  623. n_features=n_features,
  624. estimator=RandomForestClassifier(
  625. n_estimators=100, random_state=random_state, n_jobs=-1
  626. ),
  627. random_state=random_state,
  628. ord_categories=ord_categories,
  629. )
  630. # Default or user-supplied param grid for search
  631. self.param_grid = param_grid or {
  632. "select__k": [10, 20, 30, 40, 50, 60, 70, 80, 100],
  633. "clf__n_estimators": [100, 200],
  634. "clf__max_depth": [None, 5, 10],
  635. }
  636. class LogisticRegressionMLPipeline(BaseMLPipeline):
  637. def __init__(
  638. self, n_features=None, random_state=None, ord_categories=None, param_grid=None
  639. ):
  640. """
  641. Initialize LogisticRegressionMLPipeline with default LogisticRegression and parameter grid.
  642. Args:
  643. n_features (int or None): Number of top features to select (SelectKBest). If None, 'all' is used.
  644. random_state (int or None): Random state for reproducibility.
  645. ord_categories (dict or None): Categorical feature categories for ordinal encoding (if needed).
  646. param_grid (dict or None): Parameter grid for hyperparameter search. If None, default grid is used.
  647. Returns:
  648. None
  649. """
  650. super().__init__(
  651. n_features=n_features,
  652. estimator=LogisticRegression(max_iter=1000, random_state=random_state),
  653. random_state=random_state,
  654. ord_categories=ord_categories,
  655. )
  656. # Default or user-supplied param grid for search
  657. self.param_grid = param_grid or {
  658. "select__k": [10, 20, 30, 40],
  659. "clf__C": [0.1, 1, 10],
  660. "clf__penalty": ["l2"],
  661. }
  662. class XGBoostMLPipeline(BaseMLPipeline):
  663. def __init__(
  664. self, n_features=None, random_state=None, ord_categories=None, param_grid=None
  665. ):
  666. """
  667. Initialize XGBoostMLPipeline with default XGBClassifier and parameter grid.
  668. Args:
  669. n_features (int or None): Number of top features to select (SelectKBest). If None, 'all' is used.
  670. random_state (int or None): Random state for reproducibility.
  671. ord_categories (dict or None): Categorical feature categories for ordinal encoding (if needed).
  672. param_grid (dict or None): Parameter grid for hyperparameter search. If None, default grid is used.
  673. Returns:
  674. None
  675. """
  676. super().__init__(
  677. n_features=n_features,
  678. estimator=XGBClassifier(
  679. use_label_encoder=False,
  680. eval_metric="logloss",
  681. random_state=random_state,
  682. n_jobs=-1,
  683. ),
  684. random_state=random_state,
  685. ord_categories=ord_categories,
  686. )
  687. # Default or user-supplied param grid for search
  688. self.param_grid = param_grid or {
  689. "select__k": [10, 20, 30, 40, 50],
  690. "clf__max_depth": [3, 5, 7],
  691. "clf__n_estimators": [100, 200],
  692. "clf__learning_rate": [0.01, 0.1, 0.2],
  693. }
  694. class XGBoostRFMLPipeline(BaseMLPipeline):
  695. def __init__(
  696. self, n_features=None, random_state=None, ord_categories=None, param_grid=None
  697. ):
  698. """
  699. Initialize XGBoostRFMLPipeline with default XGBRFClassifier and parameter grid.
  700. Args:
  701. n_features (int or None): Number of top features to select (SelectKBest). If None, 'all' is used.
  702. random_state (int or None): Random state for reproducibility.
  703. ord_categories (dict or None): Categorical feature categories for ordinal encoding (if needed).
  704. param_grid (dict or None): Parameter grid for hyperparameter search. If None, default grid is used.
  705. Returns:
  706. None
  707. """
  708. super().__init__(
  709. n_features=n_features,
  710. estimator=XGBRFClassifier(
  711. use_label_encoder=False,
  712. eval_metric="logloss",
  713. random_state=random_state,
  714. n_jobs=-1,
  715. ),
  716. random_state=random_state,
  717. ord_categories=ord_categories,
  718. )
  719. # Default or user-supplied param grid for search
  720. self.param_grid = param_grid or {
  721. "select__k": [10, 20, 30, 40, 50],
  722. "clf__max_depth": [3, 5, 7],
  723. "clf__n_estimators": [100, 200],
  724. "clf__learning_rate": [0.01, 0.1, 0.2],
  725. "clf__subsample": [0.8, 1.0],
  726. }
  727. def nested_cv_evaluate(pipeline, X, y, outer_cv=3, inner_cv=3, search_type="random"):
  728. """
  729. Perform nested cross-validation for unbiased model selection and evaluation.
  730. Returns a list of test scores and best params for each outer fold.
  731. Args:
  732. pipeline (BaseMLPipeline instance): The ML pipeline to evaluate.
  733. X (pandas DataFrame): Input features.
  734. y (array-like): Target labels.
  735. outer_cv (int): Number of outer CV folds.
  736. inner_cv (int): Number of inner CV folds.
  737. search_type (str): "grid" or "random" for hyperparameter search.
  738. Returns:
  739. test_scores (list of floats): Test accuracies for each outer fold.
  740. best_params_list (list of dicts): Best hyperparameters for each outer fold.
  741. roc_data (list of tuples [fpr,tpr,auc,fold]): ROC curve data for each outer fold.
  742. """
  743. outer = StratifiedKFold(
  744. random_state=pipeline.random_state, n_splits=outer_cv, shuffle=True
  745. )
  746. test_scores = []
  747. best_params_list = []
  748. roc_data = [] # List of (fpr, tpr, auc, fold) for each fold
  749. fold = 1
  750. for train_idx, test_idx in outer.split(X, y):
  751. print(f"Indices are {train_idx}, {test_idx}")
  752. print(f"\n=== Outer Fold {fold} ===")
  753. X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
  754. y_train, y_test = y[train_idx], y[test_idx]
  755. # Recreate a fresh model of the same type; do not pass n_features so select__k is tuned via param_grid
  756. model = type(pipeline)(
  757. random_state=pipeline.random_state, ord_categories=pipeline.ord_categories
  758. )
  759. param_grid = getattr(model, "param_grid", None)
  760. model.search_fit(
  761. X_train,
  762. y_train,
  763. random_state=pipeline.random_state,
  764. search_type=search_type,
  765. param_grid=param_grid,
  766. cv=inner_cv,
  767. )
  768. y_pred = model.pipeline.predict(X_test)
  769. acc = accuracy_score(y_test, y_pred)
  770. print(f"Test accuracy (outer fold): {acc:.3f}")
  771. print(f"Best params (inner CV): {model.metrics_['best_params']}")
  772. print(f"Classification report:\n{classification_report(y_test, y_pred)}")
  773. test_scores.append(acc)
  774. best_params_list.append(model.metrics_["best_params"])
  775. # ROC curve for this fold
  776. if hasattr(model.pipeline.named_steps["clf"], "predict_proba"):
  777. y_score = model.pipeline.predict_proba(X_test)[:, 1]
  778. else:
  779. y_score = model.pipeline.decision_function(X_test)
  780. fpr, tpr, _ = roc_curve(y_test, y_score)
  781. roc_auc = auc(fpr, tpr)
  782. roc_data.append((fpr, tpr, roc_auc, fold))
  783. fold += 1
  784. return test_scores, best_params_list, roc_data
  785. def aggregate_seed_metrics(aggregate_records, ml_dir):
  786. """
  787. Aggregate per-seed performance across models and write raw + summary outputs.
  788. Artifacts written (if records exist):
  789. - aggregate_seed_metrics_raw.csv
  790. - aggregate_seed_metrics_summary.csv
  791. - aggregate_seed_metrics_summary.txt
  792. Uses sample standard deviation (ddof=1) for across-seed variability.
  793. Args:
  794. aggregate_records (list of dict):
  795. Each dict has keys:
  796. {'model': str, 'seed': int,
  797. 'nested_mean_acc': float,
  798. 'holdout_accuracy': float,
  799. 'holdout_auc': float}
  800. ml_dir (str): Base training output directory.
  801. Returns:
  802. summary_df (pandas.DataFrame or None): Summary dataframe or None if no records.
  803. """
  804. if not aggregate_records:
  805. print("No aggregate records to summarize.")
  806. return None
  807. agg_df = pd.DataFrame(aggregate_records)
  808. agg_csv = os.path.join(ml_dir, "aggregate_seed_metrics_raw.csv")
  809. agg_df.to_csv(agg_csv, index=False)
  810. # Compute summary stats for all metrics if present
  811. summary_rows = []
  812. for model_name, grp in agg_df.groupby("model"):
  813. n = len(grp)
  814. def ci95(mean, std, n):
  815. if n > 1:
  816. ci = 1.96 * (std / np.sqrt(n))
  817. return ci
  818. else:
  819. return np.nan
  820. # Helper to get mean, std, ci for a column if present
  821. def metric_stats(col):
  822. if col in grp:
  823. mean = grp[col].mean()
  824. std = grp[col].std(ddof=1)
  825. ci = ci95(mean, std, n)
  826. return mean, ci
  827. else:
  828. return np.nan, np.nan
  829. nested_mean_acc_mean, nested_mean_acc_ci = metric_stats("nested_mean_acc")
  830. holdout_acc_mean, holdout_acc_ci = metric_stats("holdout_accuracy")
  831. holdout_auc_mean, holdout_auc_ci = metric_stats("holdout_auc")
  832. holdout_precision_mean, holdout_precision_ci = metric_stats("holdout_precision")
  833. holdout_recall_mean, holdout_recall_ci = metric_stats("holdout_recall")
  834. holdout_specificity_mean, holdout_specificity_ci = metric_stats(
  835. "holdout_specificity"
  836. )
  837. holdout_f1_mean, holdout_f1_ci = metric_stats("holdout_f1")
  838. summary_rows.append(
  839. {
  840. "model": model_name,
  841. "seeds": grp.seed.tolist(),
  842. "n_seeds": n,
  843. "nested_mean_acc_mean": nested_mean_acc_mean,
  844. "nested_mean_acc_ci": nested_mean_acc_ci,
  845. "holdout_acc_mean": holdout_acc_mean,
  846. "holdout_acc_ci": holdout_acc_ci,
  847. "holdout_auc_mean": holdout_auc_mean,
  848. "holdout_auc_ci": holdout_auc_ci,
  849. "holdout_precision_mean": holdout_precision_mean,
  850. "holdout_precision_ci": holdout_precision_ci,
  851. "holdout_recall_mean": holdout_recall_mean,
  852. "holdout_recall_ci": holdout_recall_ci,
  853. "holdout_specificity_mean": holdout_specificity_mean,
  854. "holdout_specificity_ci": holdout_specificity_ci,
  855. "holdout_f1_mean": holdout_f1_mean,
  856. "holdout_f1_ci": holdout_f1_ci,
  857. }
  858. )
  859. summary_df = pd.DataFrame(summary_rows)
  860. summary_csv = os.path.join(ml_dir, "aggregate_seed_metrics_summary.csv")
  861. summary_df.to_csv(summary_csv, index=False)
  862. summary_txt = os.path.join(ml_dir, "aggregate_seed_metrics_summary.txt")
  863. with open(summary_txt, "w") as f:
  864. f.write("Aggregated Seed Performance Summary\n")
  865. f.write("=" * 40 + "\n\n")
  866. for _, row in summary_df.iterrows():
  867. f.write(f"Model: {row['model']}\n")
  868. f.write(f"Seeds: {row['seeds']}\n")
  869. # Nested CV accuracy
  870. if not np.isnan(row["nested_mean_acc_ci"]):
  871. f.write(
  872. f"Nested Mean Acc (mean, 95% CI): {row['nested_mean_acc_mean']:.3f} ± {row['nested_mean_acc_ci']:.3f} [{row['nested_mean_acc_mean']-row['nested_mean_acc_ci']:.3f}, {row['nested_mean_acc_mean']+row['nested_mean_acc_ci']:.3f}]\n"
  873. )
  874. else:
  875. f.write(f"Nested Mean Acc (mean): {row['nested_mean_acc_mean']:.3f}\n")
  876. # Holdout metrics
  877. def write_metric(label, mean, ci):
  878. if not np.isnan(ci):
  879. f.write(
  880. f"{label} (mean, 95% CI): {mean:.3f} ± {ci:.3f} [{mean-ci:.3f}, {mean+ci:.3f}]\n"
  881. )
  882. else:
  883. f.write(f"{label} (mean): {mean:.3f}\n")
  884. write_metric(
  885. "Hold-out Accuracy", row["holdout_acc_mean"], row["holdout_acc_ci"]
  886. )
  887. write_metric(
  888. "Hold-out Precision",
  889. row["holdout_precision_mean"],
  890. row["holdout_precision_ci"],
  891. )
  892. write_metric(
  893. "Hold-out Recall", row["holdout_recall_mean"], row["holdout_recall_ci"]
  894. )
  895. write_metric(
  896. "Hold-out Specificity",
  897. row["holdout_specificity_mean"],
  898. row["holdout_specificity_ci"],
  899. )
  900. write_metric(
  901. "Hold-out F1 Score", row["holdout_f1_mean"], row["holdout_f1_ci"]
  902. )
  903. write_metric("Hold-out AUC", row["holdout_auc_mean"], row["holdout_auc_ci"])
  904. f.write("-" * 40 + "\n")
  905. print(f"Aggregation raw CSV written to: {agg_csv}")
  906. print(f"Aggregation summary CSV written to: {summary_csv}")
  907. print(f"Aggregation summary TXT written to: {summary_txt}")
  908. return summary_df
  909. def aggregate_feature_usage(feature_usage, ml_dir, top_n=None):
  910. """
  911. Aggregate feature usage across seeds for each model.
  912. Args:
  913. feature_usage (dict):
  914. Keys are model names (str).
  915. Values are lists of dicts with keys:
  916. {'seed': int, 'features': list of str (ordered by importance), 'signed_coefficients': dict (optional, for LR)}
  917. ml_dir (str): Base training output directory.
  918. top_n (int or None): If provided, only consider top N features from each seed
  919. Returns
  920. bool or None: True if aggregation done, None if no records.
  921. """
  922. if not feature_usage:
  923. print("No feature usage records to aggregate.")
  924. return None
  925. report_lines = ["Aggregated Feature Usage Across Models", "=" * 50, ""]
  926. for model_name, records in feature_usage.items():
  927. if not records:
  928. continue
  929. model_dir = os.path.join(ml_dir, model_name)
  930. os.makedirs(model_dir, exist_ok=True)
  931. # Collect all features and ranks
  932. rank_maps = [] # list of dict(feature -> rank index starting at 1)
  933. coef_maps = [] # list of dict(feature -> signed coefficient) for LR models
  934. has_coefs = False
  935. for rec in records:
  936. feats = rec["features"]
  937. if top_n is not None:
  938. feats = feats[:top_n]
  939. rank_map = {f: i + 1 for i, f in enumerate(feats)}
  940. rank_maps.append(rank_map)
  941. if "signed_coefficients" in rec:
  942. coef_maps.append(rec["signed_coefficients"])
  943. has_coefs = True
  944. # Count occurrences
  945. all_features = [f for rm in rank_maps for f in rm.keys()]
  946. counter = Counter(all_features)
  947. n_seeds = len(rank_maps)
  948. rows = []
  949. for feat, cnt in counter.items():
  950. ranks = [rm[feat] for rm in rank_maps if feat in rm]
  951. row_dict = {
  952. "feature": feat,
  953. "count": cnt,
  954. "proportion": cnt / n_seeds,
  955. "in_all_seeds": cnt == n_seeds,
  956. "mean_rank": float(np.mean(ranks)),
  957. "median_rank": float(np.median(ranks)),
  958. "best_rank": int(np.min(ranks)),
  959. "worst_rank": int(np.max(ranks)),
  960. }
  961. # Add mean signed coefficient and direction if available (for LR)
  962. if has_coefs:
  963. coefs = [cm[feat] for cm in coef_maps if feat in cm]
  964. if coefs:
  965. row_dict["mean_signed_coef"] = float(np.mean(coefs))
  966. row_dict["direction"] = "INCREASES_PTS" if np.mean(coefs) > 0 else "DECREASES_PTS"
  967. rows.append(row_dict)
  968. df = pd.DataFrame(rows).sort_values(
  969. ["in_all_seeds", "count", "mean_rank"], ascending=[False, False, True]
  970. )
  971. out_csv = os.path.join(model_dir, "feature_usage_summary.csv")
  972. df.to_csv(out_csv, index=False)
  973. print(f"Feature usage summary written for model '{model_name}' -> {out_csv}")
  974. # Append to report
  975. report_lines.append(f"Model: {model_name}")
  976. report_lines.append(f"Seeds analyzed: {n_seeds}")
  977. report_lines.append("Top recurring features (first 15 shown):")
  978. head_df = df.head(15)
  979. # Format output with direction indicator if available
  980. if has_coefs and "direction" in df.columns:
  981. report_lines.append(head_df.to_string(index=False))
  982. else:
  983. report_lines.append(head_df.to_string(index=False))
  984. report_lines.append("-" * 50)
  985. # Write consolidated report
  986. consolidated_path = os.path.join(ml_dir, "feature_usage_across_models.txt")
  987. with open(consolidated_path, "w") as f:
  988. f.write("\n".join(report_lines))
  989. print(f"Consolidated feature usage report written to: {consolidated_path}")
  990. return True
  991. def main():
  992. """
  993. Main function to run and save multiple ML pipelines on patient features.
  994. Args:
  995. None
  996. Returns:
  997. None
  998. """
  999. BASE_FOLDER = os.path.dirname(os.path.abspath(__file__))
  1000. FEATURES_CSV = os.path.join(BASE_FOLDER, "trainingML", "patient_features.csv")
  1001. ML_DIR = os.path.join(BASE_FOLDER, "trainingML")
  1002. os.makedirs(ML_DIR, exist_ok=True)
  1003. exclude_col = ["subj_id", "mrn"]
  1004. ground_truth = ["pts"]
  1005. # Parameter Setup
  1006. # Option A: All social determinant columns (and sex) are assumed already numerically encoded.
  1007. # We therefore disable special categorical processing and treat them as numeric features.
  1008. ORD_CATEGORIES = None
  1009. # ------- CHANGE THIS FOR RUNS -------
  1010. EXCLUDE_COL_MATCH = []
  1011. # Can exclude columns that match any of these substrings
  1012. # e.g. EXCLUDE_COL_MATCH = ["RL_diff","RL_ratio", "RL_absdiff"]
  1013. # N_FEATURES removed: k will be tuned via param_grid select__k
  1014. # def generate_seeds(num_seeds=10, seed_min=1, seed_max=10000):
  1015. # """Generate a list of random integers for use as SEEDS."""
  1016. # return random.sample(range(seed_min, seed_max + 1), num_seeds)
  1017. # SEEDS = generate_seeds(num_seeds=25, seed_min=1, seed_max=10000)
  1018. SEEDS = [
  1019. # Seeds generated using Google RNG 1000-9999
  1020. 1292,
  1021. 4337,
  1022. 5590,
  1023. 3348,
  1024. 3930,
  1025. 2253,
  1026. 2400,
  1027. 7705,
  1028. 6993,
  1029. 7322,
  1030. ]
  1031. # param_grids for models
  1032. random_forest_params = {
  1033. "select__k": [10, 20, 30, 40, None],
  1034. "clf__n_estimators": [100, 200],
  1035. "clf__max_depth": [None, 5, 10],
  1036. }
  1037. logistic_regression_params = {
  1038. "select__k": [10, 20, 30, 40, None],
  1039. "clf__C": [0.1, 1, 10],
  1040. "clf__penalty": ["l2"],
  1041. }
  1042. xgboost_params = {
  1043. "select__k": [10, 20, 30, 40, 50],
  1044. "clf__max_depth": [3, 5, 7],
  1045. "clf__n_estimators": [100, 200],
  1046. "clf__learning_rate": [0.01, 0.1, 0.2],
  1047. }
  1048. xgboost_rf_params = {
  1049. "select__k": [10, 20, 30, 40, 50],
  1050. "clf__max_depth": [3, 5, 7],
  1051. "clf__n_estimators": [100, 200],
  1052. "clf__learning_rate": [0.01, 0.1, 0.2],
  1053. "clf__subsample": [0.8, 1.0],
  1054. }
  1055. OUTER_CV = 5
  1056. INNER_CV = 5
  1057. # ------------------------------------
  1058. X, y = NIFTIpatient.load_all_patients(
  1059. FEATURES_CSV,
  1060. exclude_col=exclude_col,
  1061. ground_truth=ground_truth,
  1062. exclude_col_match=EXCLUDE_COL_MATCH,
  1063. )
  1064. # Convert y to 1D array to avoid warnings
  1065. y = y.values.ravel()
  1066. # Replace sentinel -1 with NaN so median imputation can appropriately handle missingness.
  1067. # (Assumes -1 represents missing / not on file; adjust if you want to keep it as a separate value.)
  1068. X = X.replace(-1, np.nan)
  1069. # Aggregation container: list of dict records
  1070. aggregate_records = []
  1071. # Feature usage across seeds per model: model -> list of {seed, features}
  1072. feature_usage = defaultdict(list)
  1073. # Holdout ROC data for summary plot: model -> list of (fpr, tpr, auc, seed)
  1074. holdout_roc_data = defaultdict(list)
  1075. for seed in SEEDS:
  1076. # Define as many models as you want here
  1077. models = [
  1078. (
  1079. "rf",
  1080. RandomForestMLPipeline(
  1081. param_grid=random_forest_params,
  1082. random_state=seed,
  1083. ord_categories=ORD_CATEGORIES,
  1084. ),
  1085. ),
  1086. (
  1087. "lr",
  1088. LogisticRegressionMLPipeline(
  1089. param_grid=logistic_regression_params,
  1090. random_state=seed,
  1091. ord_categories=ORD_CATEGORIES,
  1092. ),
  1093. ),
  1094. (
  1095. "xgb",
  1096. XGBoostMLPipeline(
  1097. param_grid=xgboost_params,
  1098. random_state=seed,
  1099. ord_categories=ORD_CATEGORIES,
  1100. ),
  1101. ),
  1102. (
  1103. "xgbrf",
  1104. XGBoostRFMLPipeline(
  1105. param_grid=xgboost_rf_params,
  1106. random_state=seed,
  1107. ord_categories=ORD_CATEGORIES,
  1108. ),
  1109. ),
  1110. ]
  1111. # Save initial predictors (before pipeline steps) to each model folder
  1112. initial_predictors = list(X.columns)
  1113. exclude_str = "-".join(EXCLUDE_COL_MATCH)
  1114. for name, pipeline in models:
  1115. # New directory hierarchy: trainingML/{model}/seed{seed}_excl{exclude}/
  1116. model_base_dir = os.path.join(ML_DIR, name)
  1117. os.makedirs(model_base_dir, exist_ok=True)
  1118. exclude_segment = f"_excl{exclude_str}" if exclude_str else ""
  1119. run_dir = os.path.join(model_base_dir, f"seed{seed}{exclude_segment}")
  1120. nested_dir = os.path.join(run_dir, f"nested_outer{OUTER_CV}inner{INNER_CV}")
  1121. final_dir = os.path.join(run_dir, "final")
  1122. os.makedirs(nested_dir, exist_ok=True)
  1123. os.makedirs(final_dir, exist_ok=True)
  1124. # --- Nested CV artifacts ---
  1125. initial_predictors_csv = os.path.join(
  1126. nested_dir, f"{name}_initial_predictors.csv"
  1127. )
  1128. pd.Series(initial_predictors, name="predictor").to_csv(
  1129. initial_predictors_csv, index=False
  1130. )
  1131. print(f"Initial predictors saved to: {initial_predictors_csv}")
  1132. print(f"\nRunning nested CV for pipeline: {name}")
  1133. test_scores, best_params_list, cv_roc_data = nested_cv_evaluate(
  1134. pipeline, X, y, outer_cv=OUTER_CV, inner_cv=INNER_CV, search_type="grid"
  1135. )
  1136. print(
  1137. f"\n{name} mean test accuracy (nested CV): {np.mean(test_scores):.3f} ± {np.std(test_scores):.3f}"
  1138. )
  1139. nested_roc_plot_path = os.path.join(
  1140. nested_dir, f"{name}_nestedcv_roc_curves.png"
  1141. )
  1142. nested_roc_csv_path = os.path.join(
  1143. nested_dir, f"{name}_nestedcv_roc_data.csv"
  1144. )
  1145. ROCCurvePlotter.plot_nested_cv_roc_curves(
  1146. cv_roc_data,
  1147. save_path=nested_roc_plot_path,
  1148. title=f"Nested CV ROC Curves: {name}",
  1149. csv_path=nested_roc_csv_path,
  1150. )
  1151. print(f"Nested CV ROC curves saved to: {nested_roc_plot_path}")
  1152. print(f"Nested CV ROC data saved to: {nested_roc_csv_path}")
  1153. # --- Final model training with 80/20 hold-out using best hyperparameters from nested CV ---
  1154. print(
  1155. f"\nDeriving final hyperparameters for {name} from nested CV folds..."
  1156. )
  1157. # Strategy: choose the params from the fold with highest test score
  1158. best_fold_idx = int(np.argmax(test_scores))
  1159. chosen_params = best_params_list[best_fold_idx]
  1160. print(f"Selected params from fold {best_fold_idx+1}: {chosen_params}")
  1161. # 80/20 stratified split
  1162. X_train, X_test, y_train, y_test = train_test_split(
  1163. X,
  1164. y,
  1165. test_size=0.20,
  1166. random_state=seed,
  1167. stratify=y,
  1168. )
  1169. print(
  1170. f"Training final {name} model on 80% of data (hold-out 20% for unbiased evaluation)..."
  1171. )
  1172. # Instantiate clean pipeline of same type
  1173. final_model = type(pipeline)(
  1174. random_state=pipeline.random_state,
  1175. ord_categories=pipeline.ord_categories,
  1176. )
  1177. # Apply chosen hyperparameters directly (except select__k if None -> 'all')
  1178. final_model.pipeline.set_params(**chosen_params)
  1179. # Fit on training split only
  1180. final_model.pipeline.fit(X_train, y_train)
  1181. from sklearn.metrics import (
  1182. precision_score,
  1183. recall_score,
  1184. f1_score,
  1185. roc_auc_score,
  1186. )
  1187. # Compute train metrics
  1188. y_train_pred = final_model.pipeline.predict(X_train)
  1189. train_acc = accuracy_score(y_train, y_train_pred)
  1190. train_cm = confusion_matrix(y_train, y_train_pred)
  1191. train_cr = classification_report(y_train, y_train_pred)
  1192. # Train specificity
  1193. if train_cm.shape == (2, 2):
  1194. tn, fp, fn, tp = train_cm.ravel()
  1195. train_specificity = tn / (tn + fp) if (tn + fp) > 0 else np.nan
  1196. else:
  1197. train_specificity = np.nan
  1198. # Train AUC
  1199. try:
  1200. if hasattr(final_model.pipeline.named_steps["clf"], "predict_proba"):
  1201. y_train_score = final_model.pipeline.predict_proba(X_train)[:, 1]
  1202. else:
  1203. y_train_score = final_model.pipeline.decision_function(X_train)
  1204. train_auc_val = roc_auc_score(y_train, y_train_score)
  1205. except Exception:
  1206. train_auc_val = np.nan
  1207. # Compute test metrics
  1208. y_test_pred = final_model.pipeline.predict(X_test)
  1209. test_acc = accuracy_score(y_test, y_test_pred)
  1210. test_cm = confusion_matrix(y_test, y_test_pred)
  1211. test_cr = classification_report(y_test, y_test_pred)
  1212. # Test specificity
  1213. if test_cm.shape == (2, 2):
  1214. tn, fp, fn, tp = test_cm.ravel()
  1215. test_specificity = tn / (tn + fp) if (tn + fp) > 0 else np.nan
  1216. else:
  1217. test_specificity = np.nan
  1218. # Test AUC
  1219. if hasattr(final_model.pipeline.named_steps["clf"], "predict_proba"):
  1220. y_test_score = final_model.pipeline.predict_proba(X_test)[:, 1]
  1221. else:
  1222. y_test_score = final_model.pipeline.decision_function(X_test)
  1223. test_fpr, test_tpr, _ = roc_curve(y_test, y_test_score)
  1224. test_auc_val = auc(test_fpr, test_tpr)
  1225. # Store holdout ROC data for summary plot
  1226. holdout_roc_data[name].append((test_fpr, test_tpr, test_auc_val, seed))
  1227. # Feature importances/coefficients after fit
  1228. top_series = None
  1229. signed_coefs_series = None
  1230. if (
  1231. hasattr(final_model.pipeline.named_steps["clf"], "feature_importances_")
  1232. and "select" in final_model.pipeline.named_steps
  1233. ):
  1234. preprocessed_feature_names = final_model.get_preprocessed_feature_names(
  1235. final_model.pipeline, X_train
  1236. )
  1237. if "var_thresh" in final_model.pipeline.named_steps:
  1238. var_thresh_mask = final_model.pipeline.named_steps[
  1239. "var_thresh"
  1240. ].get_support()
  1241. features_after_var = np.array(preprocessed_feature_names)[
  1242. var_thresh_mask
  1243. ]
  1244. else:
  1245. features_after_var = preprocessed_feature_names
  1246. selected_mask = final_model.pipeline.named_steps["select"].get_support()
  1247. selected_features = features_after_var[selected_mask]
  1248. importances = final_model.pipeline.named_steps[
  1249. "clf"
  1250. ].feature_importances_
  1251. top_series = pd.Series(
  1252. importances, index=selected_features
  1253. ).sort_values(ascending=False)
  1254. elif (
  1255. hasattr(final_model.pipeline.named_steps["clf"], "coef_")
  1256. and "select" in final_model.pipeline.named_steps
  1257. ):
  1258. preprocessed_feature_names = final_model.get_preprocessed_feature_names(
  1259. final_model.pipeline, X_train
  1260. )
  1261. if "var_thresh" in final_model.pipeline.named_steps:
  1262. var_thresh_mask = final_model.pipeline.named_steps[
  1263. "var_thresh"
  1264. ].get_support()
  1265. features_after_var = np.array(preprocessed_feature_names)[
  1266. var_thresh_mask
  1267. ]
  1268. else:
  1269. features_after_var = preprocessed_feature_names
  1270. selected_mask = final_model.pipeline.named_steps["select"].get_support()
  1271. selected_features = features_after_var[selected_mask]
  1272. coefs = np.abs(final_model.pipeline.named_steps["clf"].coef_[0])
  1273. top_series = pd.Series(coefs, index=selected_features).sort_values(
  1274. ascending=False
  1275. )
  1276. # Store signed coefficients for direction interpretation
  1277. signed_coefs = final_model.pipeline.named_steps["clf"].coef_[0]
  1278. signed_coefs_series = pd.Series(signed_coefs, index=selected_features)
  1279. else:
  1280. top_series = None
  1281. signed_coefs_series = None
  1282. final_model.metrics_ = {
  1283. "best_params": chosen_params,
  1284. # Train
  1285. "train_accuracy": train_acc,
  1286. "train_precision": precision_score(
  1287. y_train, y_train_pred, zero_division=0
  1288. ),
  1289. "train_recall": recall_score(y_train, y_train_pred, zero_division=0),
  1290. "train_specificity": train_specificity,
  1291. "train_f1": f1_score(y_train, y_train_pred, zero_division=0),
  1292. "train_auc": train_auc_val,
  1293. "train_confusion_matrix": train_cm,
  1294. "train_classification_report": train_cr,
  1295. # Test
  1296. "test_accuracy": test_acc,
  1297. "test_precision": precision_score(y_test, y_test_pred, zero_division=0),
  1298. "test_recall": recall_score(y_test, y_test_pred, zero_division=0),
  1299. "test_specificity": test_specificity,
  1300. "test_f1": f1_score(y_test, y_test_pred, zero_division=0),
  1301. "test_auc": test_auc_val,
  1302. "test_confusion_matrix": test_cm,
  1303. "test_classification_report": test_cr,
  1304. }
  1305. if top_series is not None:
  1306. final_model.metrics_["top_features"] = top_series
  1307. # Record ordered list for feature usage aggregation (includes signed coefficients if LR)
  1308. feature_record = {"seed": seed, "features": top_series.index.tolist()}
  1309. if signed_coefs_series is not None:
  1310. feature_record["signed_coefficients"] = signed_coefs_series.to_dict()
  1311. feature_usage[name].append(feature_record)
  1312. else:
  1313. # Still record placeholder to keep seed count alignment (optional)
  1314. feature_usage[name].append({"seed": seed, "features": []})
  1315. print(
  1316. f"Hold-out test accuracy (20%): {test_acc:.3f} | AUC: {test_auc_val:.3f}"
  1317. )
  1318. # Save final model + metrics incorporating nested CV context (still reported in header)
  1319. model_path = os.path.join(final_dir, f"{name}_final_model.pkl")
  1320. metrics_path = os.path.join(final_dir, f"{name}_comprehensive_metrics.txt")
  1321. final_model.save(
  1322. model_path=model_path,
  1323. metrics_path=metrics_path,
  1324. nested_cv_results=(test_scores, best_params_list, name),
  1325. )
  1326. # Selected predictors after final training (top features if available else all)
  1327. predictors_csv = os.path.join(final_dir, f"{name}_predictors.csv")
  1328. if top_series is not None:
  1329. selected_features = top_series.index.tolist()
  1330. else:
  1331. selected_features = list(X.columns)
  1332. pd.Series(selected_features, name="predictor").to_csv(
  1333. predictors_csv, index=False
  1334. )
  1335. print(f"Predictors saved to: {predictors_csv}")
  1336. # Final ROC curve (hold-out test curve)
  1337. holdout_roc_plot_path = os.path.join(
  1338. final_dir, f"{name}_holdout_roc_curve.png"
  1339. )
  1340. plotter = ROCCurvePlotter(
  1341. y_true=y_test,
  1342. y_score=y_test_score,
  1343. title=f"Hold-out ROC Curve: {name} (seed={seed})",
  1344. )
  1345. plotter.plot(show=False, save_path=holdout_roc_plot_path)
  1346. print(f"Hold-out ROC curve saved to: {holdout_roc_plot_path}")
  1347. # Record aggregation metrics
  1348. aggregate_records.append(
  1349. {
  1350. "model": name,
  1351. "seed": seed,
  1352. "nested_mean_acc": float(np.mean(test_scores)),
  1353. "nested_std_acc": float(np.std(test_scores)),
  1354. "holdout_accuracy": float(test_acc),
  1355. "holdout_precision": float(final_model.metrics_.get("test_precision", np.nan)),
  1356. "holdout_recall": float(final_model.metrics_.get("test_recall", np.nan)),
  1357. "holdout_specificity": float(final_model.metrics_.get("test_specificity", np.nan)),
  1358. "holdout_f1": float(final_model.metrics_.get("test_f1", np.nan)),
  1359. "holdout_auc": float(test_auc_val),
  1360. "best_params": chosen_params,
  1361. "num_features_top": (
  1362. len(top_series) if top_series is not None else 0
  1363. ),
  1364. }
  1365. )
  1366. # After all seeds processed, plot summary ROC curves for each model
  1367. for name, roc_data in holdout_roc_data.items():
  1368. summary_roc_path = os.path.join(
  1369. ML_DIR, name, f"{name}_holdout_roc_curves_across_seeds.png"
  1370. )
  1371. summary_roc_csv = os.path.join(
  1372. ML_DIR, name, f"{name}_holdout_roc_data_across_seeds.csv"
  1373. )
  1374. ROCCurvePlotter.plot_holdout_roc_curves_across_seeds(
  1375. roc_data,
  1376. save_path=summary_roc_path,
  1377. title=f"Holdout ROC Curves Across Seeds: {name}",
  1378. csv_path=summary_roc_csv,
  1379. )
  1380. print(f"Summary holdout ROC curves saved to: {summary_roc_path}")
  1381. print(f"Summary holdout ROC data saved to: {summary_roc_csv}")
  1382. # Aggregate & write summaries
  1383. aggregate_seed_metrics(aggregate_records, ML_DIR)
  1384. # Aggregate feature usage across seeds per model
  1385. aggregate_feature_usage(feature_usage, ML_DIR, top_n=None)
  1386. if __name__ == "__main__":
  1387. main()

NIFTItraining.py at commit 550fe2e, no license · at the source

Overview

Authors: Mark Chao1, Monica Mallavarapu1, Matthew De Guzman1, Lauren Nguyen1, Nathan Nguyen1, Jerome Jeevarajan1
ORCID iDs: Mark Chao
  1. Department of Neurology, McGovern Medical School, The University of Texas Health Houston, 6431 Fannin St, MSB 7.005E, Houston, TX 77030 USA
Journal: Scientific reports, volume 16, issue 1, article 21655
Dates: received 31 December 2025; accepted 3 April 2026; published online 12 May 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41598-026-47942-4 · PMID 42120454 · PMCID PMC13357838 · OpenAlex W7160954675
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: other (modality), human (organism), other condition (population), epilepsy (population), traumatic brain injury (population), clinical / translational (subfield)
Methods: Connectivity, Statistics, Machine learning
Keywords: Traumatic brain injury, Post-traumatic epilepsy, Computed tomography, Radiomics, Machine learning, Biomarkers, Medical research, Neurology, Neuroscience
MeSH: Brain Injuries, Traumatic*, Epilepsy, Post-Traumatic*, Seizures*, Tomography, X-Ray Computed*, Adult, Biomarkers, Female, Glasgow Coma Scale, Humans, Machine Learning, Male, Middle Aged, Pilot Projects, Radiomics (* major topic)
Topic: Traumatic Brain Injury and Neurovascular Disturbances (Neurology, Medicine), according to OpenAlex
Citations: not cited yet (Europe PMC); 19 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

makchow/NIFTIProcessing-Published

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 550fe2ee3a3c4da28b6727907de305a178306fdf, 23 February 2026
Languages: Python (3)
Size: 305 files, 3 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (3 files), pandas (3 files), NiBabel (2 files), dcm2niix (1 file), FSL (1 file), Matplotlib (1 file), PyRadiomics (1 file), scikit-learn (1 file), SciPy (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
4 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 3 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41598-026-47942-4.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 9 keywords, 14 MeSH terms, 17 references.

Cite

This paper

Chao, M., Mallavarapu, M., De Guzman, M., Nguyen, L., Nguyen, N., & Jeevarajan, J. (2026). CT-based radiomic markers to predict late-onset seizures after traumatic brain injury. Scientific reports, 16(1), 21655. https://doi.org/10.1038/s41598-026-47942-4

BibTeX

@article{chao2026ct,
author = {Chao, Mark and Mallavarapu, Monica and De Guzman, Matthew and Nguyen, Lauren and Nguyen, Nathan and Jeevarajan, Jerome},
title = {{CT-based radiomic markers to predict late-onset seizures after traumatic brain injury}},
journal = {Scientific reports},
year = {2026},
month = may,
volume = {16},
number = {1},
pages = {21655},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-47942-4},
url = {https://doi.org/10.1038/s41598-026-47942-4},
pmid = {42120454},
pmcid = {PMC13357838}
}

RIS

TY - JOUR
AU - Chao, Mark
AU - Mallavarapu, Monica
AU - De Guzman, Matthew
AU - Nguyen, Lauren
AU - Nguyen, Nathan
AU - Jeevarajan, Jerome
TI - CT-based radiomic markers to predict late-onset seizures after traumatic brain injury
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/05/12
VL - 16
IS - 1
SP - 21655
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-47942-4
UR - https://doi.org/10.1038/s41598-026-47942-4
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-47942-4",
"type": "article-journal",
"title": "CT-based radiomic markers to predict late-onset seizures after traumatic brain injury",
"container-title": "Scientific reports",
"author": [
{
"family": "Chao",
"given": "Mark"
},
{
"family": "Mallavarapu",
"given": "Monica"
},
{
"family": "De Guzman",
"given": "Matthew"
},
{
"family": "Nguyen",
"given": "Lauren"
},
{
"family": "Nguyen",
"given": "Nathan"
},
{
"family": "Jeevarajan",
"given": "Jerome"
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "21655",
"DOI": "10.1038/s41598-026-47942-4",
"PMID": "42120454",
"PMCID": "PMC13357838",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-47942-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41598-026-56688-y [code]
On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.
Journal: Scientific reports
In common: PyRadiomics, XGBoost, FSL, 6 other tools, 1 reference
[2] doi:10.1162/netn.a.547 [code]
An evaluation of the efficacy of single-echo and multi-echo fMRI denoising strategies.
Journal: Network neuroscience (Cambridge, Mass.)
In common: dcm2niix, FSL, NiBabel, 5 other tools, 1 reference
[3] doi:10.1371/journal.pcbi.1014555 [code]
Body surface potential driven personalisation of electrophysiological digital twins in hypertrophic cardiomyopathy.
Journal: PLoS computational biology
In common: PyRadiomics, XGBoost, NiBabel, 5 other tools, other
[4] doi:10.1186/s13244-026-02365-7 [code]
Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion.
Journal: Insights into imaging
In common: PyRadiomics, XGBoost, NiBabel, 5 other tools, other condition
[5] doi:10.3390/cancers18162636 [code]
Exploring the Impact of T2-Weighted MRI Fat Saturation on Radiomics Stability for Brain Radionecrosis Prediction After Skull-Base Proton Therapy: A Pilot Study.
Journal: Cancers
In common: PyRadiomics, FSL, scikit-learn, 4 other tools, clinical / translational, 1 reference
[6] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: XGBoost, FSL, NiBabel, 5 other tools, epilepsy, clinical / translational
[7] doi:10.1371/journal.pone.0356382 [code]
Evaluating the impact of segmentation strategies on radiomic feature stability for paediatric brain tumour diagnosis using diffusion weighted imaging.
Journal: PloS one
In common: PyRadiomics, NiBabel, scikit-learn, 3 other tools, clinical / translational, other condition, 1 reference
[8] doi:10.1038/s41467-026-71918-7 [code]
Developmental disinhibition gates language lateralization in childhood.
Journal: Nature communications
In common: dcm2niix, FSL, NiBabel, 5 other tools
[9] doi:10.1038/s41597-026-07248-6 [code]
A large-scale fMRI dataset for vision-language semantic association.
Journal: Scientific data
In common: dcm2niix, FSL, NiBabel, 5 other tools
[10] doi:10.1038/s41467-026-71719-y [code]
Brain functional-structural gradient coupling reflects development, behavior and genetic influences.
Journal: Nature communications
In common: dcm2niix, FSL, NiBabel, 5 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.