OSCR

6PPDQ Exposure Exacerbates Seizure-Induced Neuronal Damage via the TP53/Nrf2 Axis: An Integrated Strategy Combining Network Toxicology and Experimental Validation.

Code ↔ Paper

1 match between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 1 match
  1. [1] § 2. Materials and Methods › 2.4. Construction and Validation of a Diagnostic Machine Learning Signature ↔ 113-Machine Learning Modeling(Two GEO dataset).py, lines 191–209 · score 0.80 · NaiveBayes, XGBoost, machine learning, Random Forest, feature selection, Elastic

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 510 lines · 22 KB · no license · 1 match

  1. # ==================== Import Libraries ====================
  2. import pandas as pd
  3. import numpy as np
  4. import warnings
  5. import time
  6. import matplotlib.pyplot as plt
  7. import seaborn as sns
  8. from sklearn.model_selection import train_test_split
  9. from sklearn.preprocessing import StandardScaler
  10. from sklearn.base import clone
  11. from sklearn.metrics import roc_auc_score
  12. # Feature Selection & Modeling Algorithm Libraries
  13. from sklearn.linear_model import (
  14. Lasso, Ridge, LogisticRegression, ElasticNet
  15. )
  16. from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
  17. from sklearn.ensemble import (
  18. RandomForestClassifier,
  19. GradientBoostingClassifier,
  20. AdaBoostClassifier
  21. )
  22. from sklearn.svm import SVC
  23. from sklearn.naive_bayes import GaussianNB
  24. from sklearn.neighbors import KNeighborsClassifier
  25. from sklearn.tree import DecisionTreeClassifier
  26. from xgboost import XGBClassifier
  27. # Filter warnings to keep output clean
  28. warnings.filterwarnings('ignore')
  29. print("=" * 80)
  30. print("Machine Learning Modeling Pipeline - 113 Combinations (Based on SCI papers)")
  31. print("Supports external validation set")
  32. print("=" * 80)
  33. # ==================== Configuration Parameters ====================
  34. RANDOM_SEED = 123
  35. TRAIN_SIZE = 0.7 # Training set ratio 70%
  36. TEST_SIZE = 0.3 # Test set ratio 30%
  37. MAX_ITER = 5000 # Maximum iterations for linear models
  38. # Dataset file paths
  39. TRAIN_DATASET_PATH = 'ML_dataset.csv' # Main dataset (for training + testing)
  40. VALIDATION_DATASET_PATH = 'ML_dataset_validation.csv' # External validation set
  41. # ==================== 1. Data Loading & Preprocessing ====================
  42. def load_raw_data(filepath, dataset_name="Main Dataset"):
  43. """
  44. Read raw dataset (without preprocessing)
  45. """
  46. print(f"\n -> Loading {dataset_name}: {filepath}...")
  47. try:
  48. ml_data = pd.read_csv(filepath, encoding='utf-8-sig')
  49. except FileNotFoundError:
  50. print(f" ⚠️ Warning: File not found {filepath}")
  51. return None, None
  52. # Handle index
  53. first_col = ml_data.columns[0]
  54. if first_col == '' or 'Unnamed' in first_col:
  55. ml_data = ml_data.rename(columns={first_col: 'Sample_ID'})
  56. if 'Sample_ID' in ml_data.columns:
  57. ml_data = ml_data.set_index('Sample_ID')
  58. # Check grouping column
  59. if 'Group' not in ml_data.columns:
  60. raise ValueError(f"{dataset_name} must contain a 'Group' column.")
  61. print(f" Original dimensions: {ml_data.shape}")
  62. print(f" Group distribution:")
  63. print(ml_data['Group'].value_counts())
  64. # Separate features (X) and labels (y)
  65. y = ml_data['Group']
  66. X = ml_data.drop(columns=['Group'])
  67. return X, y
  68. def prepare_datasets(train_path, validation_path):
  69. """
  70. Prepare training, test, and external validation sets
  71. Critical fix: Align features first, then standardize
  72. """
  73. print("\n" + "="*60)
  74. print("[Step 1] Data Loading & Preprocessing")
  75. print("="*60)
  76. # Load main dataset
  77. X_main, y_main = load_raw_data(train_path, "Main Dataset")
  78. if X_main is None:
  79. return None, None, None, None, None, None, None, None, None, False, None
  80. # Label encoding (auto-detect)
  81. unique_classes = y_main.unique()
  82. pos_class = 'Seizures' if 'Seizures' in unique_classes else unique_classes[0]
  83. y_main_encoded = (y_main == pos_class).astype(int)
  84. print(f"\n Label encoding: {pos_class}=1, Others=0")
  85. # Split train/test sets (70%/30%)
  86. print(f"\n -> Splitting train/test sets ({int(TRAIN_SIZE*100)}%/{int(TEST_SIZE*100)}%)...")
  87. X_train, X_test, y_train, y_test = train_test_split(
  88. X_main, y_main_encoded,
  89. test_size=TEST_SIZE,
  90. random_state=RANDOM_SEED,
  91. stratify=y_main_encoded
  92. )
  93. print(f" Training set: {X_train.shape[0]} samples, {X_train.shape[1]} genes")
  94. print(f" Test set: {X_test.shape[0]} samples, {X_test.shape[1]} genes")
  95. # Load external validation set (if exists)
  96. X_val, y_val_encoded, has_validation = None, None, False
  97. common_features = X_train.columns.tolist() # Default to using all training features
  98. try:
  99. X_val_raw, y_val = load_raw_data(validation_path, "External Validation Set")
  100. if X_val_raw is not None:
  101. has_validation = True
  102. print(f"\n External validation set original dimensions: {X_val_raw.shape}")
  103. print(f" External validation set group distribution:")
  104. print(y_val.value_counts())
  105. # 🔥 Critical fix: Find common features first
  106. train_features = set(X_train.columns)
  107. val_features = set(X_val_raw.columns)
  108. common_features = sorted(list(train_features & val_features))
  109. print(f"\n Training set gene count: {len(train_features)}")
  110. print(f" Validation set gene count: {len(val_features)}")
  111. print(f" Common gene count: {len(common_features)}")
  112. if len(common_features) < len(train_features):
  113. missing_in_val = len(train_features) - len(common_features)
  114. print(f" ⚠️ Validation set is missing {missing_in_val} genes")
  115. if len(common_features) < 2:
  116. print(f" ❌ Too few common genes, cannot proceed with analysis")
  117. has_validation = False
  118. else:
  119. # 🔥 Critical fix: Align features across all datasets before standardization
  120. X_train = X_train[common_features]
  121. X_test = X_test[common_features]
  122. X_val = X_val_raw[common_features]
  123. # Fill missing values in validation set (using training set mean)
  124. if X_val.isnull().sum().sum() > 0:
  125. print(f" ⚠️ Missing values found in validation set, filling with training set mean")
  126. X_val = X_val.fillna(X_train.mean())
  127. # Validation set label encoding (using the same positive class as main dataset)
  128. y_val_encoded = (y_val == pos_class).astype(int)
  129. print(f" ✓ Feature alignment complete, external validation set: {X_val.shape[0]} samples × {X_val.shape[1]} genes")
  130. except FileNotFoundError:
  131. print(f"\n ⚠️ External validation set file not found, will use only main dataset for evaluation")
  132. # Handle missing values in training and test sets
  133. if X_train.isnull().sum().sum() > 0:
  134. print(f"\n ⚠️ Missing values found in training set, filling with mean")
  135. X_train = X_train.fillna(X_train.mean())
  136. if X_test.isnull().sum().sum() > 0:
  137. X_test = X_test.fillna(X_train.mean())
  138. # 🔥 Critical fix: Standardize AFTER feature alignment
  139. print(f"\n -> Standardizing data (Z-score)...")
  140. scaler = StandardScaler()
  141. X_train_scaled_array = scaler.fit_transform(X_train)
  142. X_test_scaled_array = scaler.transform(X_test)
  143. # Convert back to DataFrame (keep column names consistent)
  144. X_train_scaled = pd.DataFrame(X_train_scaled_array, columns=common_features, index=X_train.index)
  145. X_test_scaled = pd.DataFrame(X_test_scaled_array, columns=common_features, index=X_test.index)
  146. # Standardize validation set (using training set scaler)
  147. X_val_scaled = None
  148. if has_validation and X_val is not None:
  149. X_val_scaled_array = scaler.transform(X_val)
  150. X_val_scaled = pd.DataFrame(X_val_scaled_array, columns=common_features, index=X_val.index)
  151. print(f" ✓ Data standardization complete")
  152. else:
  153. print(f" ✓ Data standardization complete (no external validation set)")
  154. return (X_train, X_test, X_val,
  155. X_train_scaled, X_test_scaled, X_val_scaled,
  156. y_train, y_test, y_val_encoded, has_validation, common_features)
  157. # ==================== 2. Define Feature Selection Algorithms (12 types) ====================
  158. def get_feature_selectors():
  159. """
  160. Return a dictionary of feature selection algorithms. 12 types in total.
  161. """
  162. return {
  163. 'Lasso': Lasso(alpha=0.01, random_state=RANDOM_SEED, max_iter=MAX_ITER),
  164. 'Ridge': Ridge(alpha=1.0, random_state=RANDOM_SEED),
  165. 'Stepglm': LogisticRegression(penalty='l1', solver='liblinear', C=1.0, random_state=RANDOM_SEED),
  166. 'XGBoost': XGBClassifier(n_estimators=100, random_state=RANDOM_SEED, use_label_encoder=False, eval_metric='logloss', verbosity=0),
  167. 'RF': RandomForestClassifier(n_estimators=100, random_state=RANDOM_SEED, n_jobs=-1),
  168. 'Enet': ElasticNet(alpha=0.01, l1_ratio=0.5, random_state=RANDOM_SEED, max_iter=MAX_ITER),
  169. 'plsRglm': Ridge(alpha=0.5, random_state=RANDOM_SEED),
  170. 'GBM': GradientBoostingClassifier(n_estimators=100, random_state=RANDOM_SEED),
  171. 'NaiveBayes': GaussianNB(),
  172. 'LDA': LinearDiscriminantAnalysis(),
  173. 'glmBoost': GradientBoostingClassifier(n_estimators=50, learning_rate=0.1, random_state=RANDOM_SEED),
  174. 'SVM': SVC(kernel='linear', random_state=RANDOM_SEED)
  175. }
  176. # ==================== 3. Define Classification Modeling Algorithms (11 types) ====================
  177. def get_classifiers():
  178. """
  179. Return a dictionary of classification modeling algorithms. 11 types in total.
  180. """
  181. return {
  182. 'LDA': LinearDiscriminantAnalysis(),
  183. 'Ridge': LogisticRegression(penalty='l2', C=1.0, max_iter=1000, random_state=RANDOM_SEED),
  184. 'Lasso': LogisticRegression(penalty='l1', solver='liblinear', C=1.0, random_state=RANDOM_SEED),
  185. 'glmnet': LogisticRegression(penalty='elasticnet', solver='saga', l1_ratio=0.5, C=1.0, max_iter=1000, random_state=RANDOM_SEED),
  186. 'RF': RandomForestClassifier(n_estimators=100, random_state=RANDOM_SEED, n_jobs=-1),
  187. 'SVM': SVC(kernel='rbf', probability=True, random_state=RANDOM_SEED),
  188. 'GBM': GradientBoostingClassifier(n_estimators=100, random_state=RANDOM_SEED),
  189. 'XGBoost': XGBClassifier(n_estimators=100, random_state=RANDOM_SEED, use_label_encoder=False, eval_metric='logloss', verbosity=0),
  190. 'NaiveBayes': GaussianNB(),
  191. 'AdaBoost': AdaBoostClassifier(n_estimators=50, random_state=RANDOM_SEED),
  192. 'DT': DecisionTreeClassifier(random_state=RANDOM_SEED)
  193. }
  194. # ==================== 4. Helper Function: Execute Feature Selection ====================
  195. def run_feature_selection(X_train, X_train_scaled, y_train):
  196. """
  197. Execute feature selection on the training set
  198. """
  199. print("\n[Step 2] Executing Feature Selection (12 methods)...")
  200. selectors = get_feature_selectors()
  201. fs_results = {}
  202. for name, model in selectors.items():
  203. start_time = time.time()
  204. selected_feats = []
  205. try:
  206. # A. Linear models (Based on Coefficients)
  207. if name in ['Lasso', 'Ridge', 'Enet', 'plsRglm']:
  208. model.fit(X_train_scaled, y_train)
  209. coef = np.abs(model.coef_).flatten()
  210. top_idx = np.argsort(coef)[-50:]
  211. if name in ['Lasso', 'Enet']:
  212. real_top = [i for i in top_idx if coef[i] > 1e-5]
  213. if len(real_top) < 2:
  214. real_top = top_idx
  215. top_idx = real_top
  216. selected_feats = X_train.columns[top_idx].tolist()
  217. # B. Tree model ensembles (Based on Feature Importance)
  218. elif name in ['XGBoost', 'RF', 'GBM', 'glmBoost']:
  219. X_curr = X_train
  220. if X_train.shape[1] > 2000:
  221. vars_idx = np.argsort(X_train.var())[-2000:]
  222. X_curr = X_train.iloc[:, vars_idx]
  223. model.fit(X_curr, y_train)
  224. importances = model.feature_importances_
  225. top_idx = np.argsort(importances)[-50:]
  226. selected_feats = X_curr.columns[top_idx].tolist()
  227. # C. Stepglm (L1 Logistic Regression simulation)
  228. elif name == 'Stepglm':
  229. model.fit(X_train_scaled, y_train)
  230. coef = np.abs(model.coef_).flatten()
  231. top_idx = np.argsort(coef)[-50:]
  232. selected_feats = X_train.columns[top_idx].tolist()
  233. # D. SVM / LDA (Based on Coefficients)
  234. elif name in ['SVM', 'LDA']:
  235. model.fit(X_train_scaled, y_train)
  236. coef = np.abs(model.coef_).flatten()
  237. top_idx = np.argsort(coef)[-50:]
  238. selected_feats = X_train.columns[top_idx].tolist()
  239. # E. NaiveBayes (Based on Variance)
  240. elif name == 'NaiveBayes':
  241. top_idx = np.argsort(X_train.var())[-50:]
  242. selected_feats = X_train.columns[top_idx].tolist()
  243. fs_results[name] = selected_feats
  244. print(f" -> {name:12s}: {len(selected_feats):3d} features ({time.time()-start_time:.1f}s)")
  245. except Exception as e:
  246. print(f" -> {name:12s}: ✗ Failed ({str(e)[:30]}). Falling back to Top 50 by variance.")
  247. fs_results[name] = X_train.columns[np.argsort(X_train.var())[-50:]].tolist()
  248. return fs_results
  249. # ==================== 5. Main Execution Flow ====================
  250. # 1. Load data
  251. (X_train, X_test, X_val,
  252. X_train_scaled, X_test_scaled, X_val_scaled,
  253. y_train, y_test, y_val,
  254. has_validation, feature_columns) = prepare_datasets(TRAIN_DATASET_PATH, VALIDATION_DATASET_PATH)
  255. if X_train is not None:
  256. # 2. Execute feature selection (only on training set)
  257. feature_selection_results = run_feature_selection(X_train, X_train_scaled, y_train)
  258. # 3. Generate combinations
  259. print("\n[Step 3] Generating model combinations...")
  260. classifiers = get_classifiers()
  261. all_combinations = []
  262. fs_methods = list(feature_selection_results.keys())
  263. clf_methods = list(classifiers.keys())
  264. print(f" Feature Selection methods (FS): {len(fs_methods)}")
  265. print(f" Classification Modeling methods (CLF): {len(clf_methods)}")
  266. print(f" Theoretical combinations: {len(fs_methods)} * {len(clf_methods)} = {len(fs_methods)*len(clf_methods)}")
  267. # Exclude duplicate combinations
  268. duplicates_to_skip = {
  269. ('RF', 'RF'),
  270. ('GBM', 'GBM'),
  271. ('XGBoost', 'XGBoost'),
  272. ('NaiveBayes', 'NaiveBayes'),
  273. ('LDA', 'LDA'),
  274. ('SVM', 'SVM'),
  275. ('Lasso', 'Lasso'),
  276. ('Ridge', 'Ridge'),
  277. ('Enet', 'glmnet'),
  278. ('Stepglm', 'Lasso'),
  279. ('Stepglm', 'Ridge'),
  280. ('Stepglm', 'glmnet'),
  281. ('Stepglm', 'LDA'),
  282. ('plsRglm', 'Lasso'),
  283. ('plsRglm', 'Ridge'),
  284. ('plsRglm', 'glmnet'),
  285. ('plsRglm', 'LDA'),
  286. ('glmBoost', 'GBM'),
  287. ('glmBoost', 'AdaBoost')
  288. }
  289. count_generated = 0
  290. count_skipped = 0
  291. for fs_name in fs_methods:
  292. for clf_name in clf_methods:
  293. if (fs_name, clf_name) in duplicates_to_skip:
  294. count_skipped += 1
  295. continue
  296. all_combinations.append({
  297. 'name': f"{fs_name} + {clf_name}",
  298. 'fs': fs_name,
  299. 'clf': clf_name,
  300. })
  301. count_generated += 1
  302. print(f" Excluded duplicate combinations: {count_skipped}")
  303. print(f" Final valid combinations: {len(all_combinations)} (Target: 113)")
  304. # 4. Train models and evaluate
  305. print("\n[Step 4] Training models and calculating AUC...")
  306. results = []
  307. total_models = len(all_combinations)
  308. for i, combo in enumerate(all_combinations, 1):
  309. try:
  310. feats = feature_selection_results[combo['fs']]
  311. if len(feats) < 2:
  312. continue
  313. # Select data based on model type
  314. if combo['clf'] in ['Lasso', 'Ridge', 'glmnet', 'LDA', 'SVM']:
  315. train_data = X_train_scaled[feats]
  316. test_data = X_test_scaled[feats]
  317. val_data = X_val_scaled[feats] if (has_validation and X_val_scaled is not None) else None
  318. else:
  319. train_data = X_train[feats]
  320. test_data = X_test[feats]
  321. val_data = X_val[feats] if (has_validation and X_val is not None) else None
  322. model = clone(classifiers[combo['clf']])
  323. model.fit(train_data, y_train)
  324. # Predict
  325. if hasattr(model, "predict_proba"):
  326. y_train_pred = model.predict_proba(train_data)[:, 1]
  327. y_test_pred = model.predict_proba(test_data)[:, 1]
  328. y_val_pred = model.predict_proba(val_data)[:, 1] if (val_data is not None) else None
  329. else:
  330. y_train_pred = model.decision_function(train_data)
  331. y_test_pred = model.decision_function(test_data)
  332. y_val_pred = model.decision_function(val_data) if (val_data is not None) else None
  333. # Calculate metrics
  334. train_auc = roc_auc_score(y_train, y_train_pred)
  335. test_auc = roc_auc_score(y_test, y_test_pred)
  336. val_auc = roc_auc_score(y_val, y_val_pred) if (y_val_pred is not None and has_validation) else np.nan
  337. # Print progress
  338. if i <= 5 or i % 10 == 0 or i == total_models:
  339. if has_validation and not np.isnan(val_auc):
  340. print(f" [{i:3d}/{total_models}] {combo['name']:35s} | Train: {train_auc:.3f} | Test: {test_auc:.3f} | Val: {val_auc:.3f}")
  341. else:
  342. print(f" [{i:3d}/{total_models}] {combo['name']:35s} | Train: {train_auc:.3f} | Test: {test_auc:.3f}")
  343. results.append({
  344. 'Method': combo['name'],
  345. 'Feature_Selection': combo['fs'],
  346. 'Classifier': combo['clf'],
  347. 'N_Features': len(feats),
  348. 'Train_AUC': train_auc,
  349. 'Test_AUC': test_auc,
  350. 'Validation_AUC': val_auc if has_validation else np.nan,
  351. 'Mean_AUC': (train_auc + test_auc) / 2
  352. })
  353. except Exception as e:
  354. print(f" [{i:3d}/{total_models}] {combo['name']:35s} | ✗ Failed: {str(e)[:30]}")
  355. # 5. Save results
  356. print("\n[Step 5] Saving results...")
  357. if results:
  358. results_df = pd.DataFrame(results).sort_values('Mean_AUC', ascending=False)
  359. # Save main CSV
  360. results_df.to_csv('03_all_models_results.csv', index=False)
  361. # Save AUC matrix for R (including validation set)
  362. if has_validation:
  363. auc_matrix = results_df[['Method', 'Train_AUC', 'Test_AUC', 'Validation_AUC']].set_index('Method')
  364. auc_matrix.columns = ['Train', 'Test', 'Validation']
  365. else:
  366. auc_matrix = results_df[['Method', 'Train_AUC', 'Test_AUC']].set_index('Method')
  367. auc_matrix.columns = ['Train', 'Test']
  368. auc_matrix.to_csv('04_AUC_matrix_for_R.txt', sep='\t')
  369. # Save Top 10
  370. results_df.head(10).to_csv('05_top10_models.csv', index=False)
  371. # Save feature selection list
  372. with open('01_feature_selection_results.txt', 'w') as f:
  373. for m, feats in feature_selection_results.items():
  374. f.write(f"{m}\t{','.join(feats)}\n")
  375. print(" ✓ All result files saved.")
  376. print(f" ✓ Successfully trained models: {len(results)}")
  377. # 6. Plotting preview
  378. fig, axes = plt.subplots(1, 2, figsize=(20, 10))
  379. top_plot = results_df.head(30)
  380. # Left plot: Train & Test AUC
  381. plot_data_left = top_plot[['Train_AUC', 'Test_AUC']]
  382. plot_data_left.index = top_plot['Method']
  383. sns.heatmap(plot_data_left, annot=True, fmt='.3f', cmap='RdYlGn', vmin=0.5, vmax=1.0, ax=axes[0])
  384. axes[0].set_title('Top 30 Models Performance (Train & Test AUC)')
  385. # Right plot: If validation set exists, show Validation AUC
  386. if has_validation and 'Validation_AUC' in results_df.columns:
  387. plot_data_right = top_plot[['Test_AUC', 'Validation_AUC']]
  388. plot_data_right.index = top_plot['Method']
  389. sns.heatmap(plot_data_right, annot=True, fmt='.3f', cmap='RdYlGn', vmin=0.5, vmax=1.0, ax=axes[1])
  390. axes[1].set_title('Top 30 Models Performance (Test & Validation AUC)')
  391. else:
  392. plot_data_right = top_plot[['Train_AUC', 'Mean_AUC']]
  393. plot_data_right.index = top_plot['Method']
  394. sns.heatmap(plot_data_right, annot=True, fmt='.3f', cmap='RdYlGn', vmin=0.5, vmax=1.0, ax=axes[1])
  395. axes[1].set_title('Top 30 Models Performance (Train & Mean AUC)')
  396. plt.tight_layout()
  397. plt.savefig('06_python_preview_heatmap.png', dpi=150)
  398. print(" ✓ Preview heatmap saved: 06_python_preview_heatmap.png")
  399. # 7. Print best model summary
  400. print("\n" + "="*60)
  401. print("[Best Model Summary]")
  402. print("="*60)
  403. if has_validation and 'Validation_AUC' in results_df.columns:
  404. best_by_val = results_df.loc[results_df['Validation_AUC'].idxmax()]
  405. print(f"\n Best model ranked by Validation AUC:")
  406. print(f" Method: {best_by_val['Method']}")
  407. print(f" Train AUC: {best_by_val['Train_AUC']:.3f}")
  408. print(f" Test AUC: {best_by_val['Test_AUC']:.3f}")
  409. print(f" Validation AUC: {best_by_val['Validation_AUC']:.3f}")
  410. best_by_test = results_df.loc[results_df['Test_AUC'].idxmax()]
  411. print(f"\n Best model ranked by Test AUC:")
  412. print(f" Method: {best_by_test['Method']}")
  413. print(f" Train AUC: {best_by_test['Train_AUC']:.3f}")
  414. print(f" Test AUC: {best_by_test['Test_AUC']:.3f}")
  415. if has_validation and not np.isnan(best_by_test['Validation_AUC']):
  416. print(f" Validation AUC: {best_by_test['Validation_AUC']:.3f}")
  417. else:
  418. print("❌ No models trained successfully.")
  419. print("\n" + "=" * 80)
  420. print("Pipeline execution completed")
  421. print("=" * 80)

113-Machine Learning Modeling(Two GEO dataset).py at commit ce65602, no license · at the source

Overview

Authors: Ruijin Xie1,2, Wei Xiao2, Hua Xu2, Yufan Luo2, Xue Xiao2, Qiyang Pan3,4, Shengjie Xu2, Li Liu5, Chenyu Sun4,6,7, Yueying Liu2
  1. School of Normal Education, Yangzhou Polytechnic College, Yangzhou 225009, China
  2. School of Medicine, Jiangnan University, Wuxi 214122, China
  3. Department of Psychiatry and Psychology, Mayo Clinic, Rochester, MN 55905, USA
  4. Mayo Clinic School of Graduate Medical Education, Mayo Clinic College of Medicine and Science, Rochester, MN 55905, USA
  5. Department of Internal Medicine, The Second People’s Hospital of Hefei, Guangde Road, Hefei 230061, China
  6. Division of Public Health, Infectious Diseases, and Occupational Medicine, Mayo Clinic, Rochester, MN 55905, USA
  7. School of Public Health, University of Minnesota-Twin Cities, Minneapolis, MN 55455, USA
Institutions: Jiangnan University (China); Yangzhou Polytechnic Institute (China); Mayo Clinic (United States); University of Minnesota (United States)
Journal: Toxics, volume 14, issue 5, article 443
Dates: received 18 April 2026; accepted 17 May 2026; published online 19 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/toxics14050443 · PMID 42198569 · PMCID PMC13211372 · OpenAlex W7161625102
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: epilepsy (population), cellular / molecular (subfield)
Methods: Statistics, Machine learning
Keywords: seizure, neuronal damage, inflammation, 6PPDQ, Nrf2, TP53
Topic: Neuroinflammation and Neurodegeneration Mechanisms (Neurology, Neuroscience), according to OpenAlex
Funding: National Natural Science Foundation of China (82371462); Provincial Subsidy Funds for Pairing Assistance (DHBFZ202505); Jiangsu Traditional Chinese Medicine Research Project (PDJH2024027); Qing Lan Project of Jiangsu (JS2023-27); Traditional Chinese Medicine Science and Technology Development Plan Project of Zhejiang Province (2024ZL140); Hangzhou Medical and Health Science and Technology Project (B20252321); Yangzhou Polytechnic College Research Project (2025XJ17); Social Science Research Project on Yangzhou’s Child-Friendly Initiative (2025YFE-057); Yangzhou Basic Research Program (2025YBR-225)
Citations: not cited yet (Europe PMC); 47 references in the paper

Abstract

As an emerging tire wear-derived environmental contaminant, 6PPD-quinone (6PPDQ) has raised significant concerns regarding its neurotoxic potential, particularly for children exposed to recycled tire crumb rubber in playgrounds. However, the molecular mechanisms by which 6PPDQ influences neurological disorders such as epilepsy remain poorly understood. In this study, we employed an integrative framework combining network toxicology, bulk analysis of human epileptic brain tissues, Mendelian randomization, and molecular dynamics simulations to elucidate these mechanisms. Our findings, validated through CETSA-WB and SPR, identify 6PPDQ as a direct ligand that binds to and stabilizes neuronal TP53. Through a synergistic double-hit mechanism, 6PPDQ directly engages the TP53 pathway while simultaneously triggering microglial interleukin-6 secretion. These converging pathways lead to the suppression of the master antioxidant regulator Nrf2, resulting in glutathione depletion, excessive reactive oxygen species accumulation, and exacerbated neuronal damage under excitotoxic stress. Experimental validation using glutamate-induced HT22 cell models and microglia–neuron crosstalk systems confirmed that targeting the TP53/Nrf2 axis or scavenging ROS significantly attenuates 6PPDQ-induced neurotoxicity. Our findings highlight critical risks to pediatric neurological health posed by tire-derived contaminants and identify the TP53/Nrf2 axis as a promising therapeutic target. Furthermore, this work provides a robust scientific basis for refining risk assessment frameworks and developing regulatory strategies to mitigate environmental exposure to 6PPDQ.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.

PediatricLab-Jiangnan/6PPDQ-network

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: ce656024aff015e37670569edfedd57d8153c9d1, 18 April 2026
Languages: Python (1), R (1)
Size: 4 files, 2 scripts
Software Heritage: not archived
Found in: the text, “2.16. Statistical Analysis”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: ggplot2 (1 file), Matplotlib (1 file), NumPy (1 file), pandas (1 file), pROC (1 file), randomForest (1 file), scikit-learn (1 file), seaborn (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
3 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 2 scripts, each with its path and the digest of its content;
  • 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 6 keywords, 9 funders, 47 references.

Cite

This paper

Xie, R., Xiao, W., Xu, H., Luo, Y., Xiao, X., Pan, Q., Xu, S., Liu, L., Sun, C., & Liu, Y. (2026). 6PPDQ Exposure Exacerbates Seizure-Induced Neuronal Damage via the TP53/Nrf2 Axis: An Integrated Strategy Combining Network Toxicology and Experimental Validation. Toxics, 14(5), 443. https://doi.org/10.3390/toxics14050443

BibTeX

@article{xie20266ppdq,
author = {Xie, Ruijin and Xiao, Wei and Xu, Hua and Luo, Yufan and Xiao, Xue and Pan, Qiyang and Xu, Shengjie and Liu, Li and Sun, Chenyu and Liu, Yueying},
title = {{6PPDQ Exposure Exacerbates Seizure-Induced Neuronal Damage via the TP53/Nrf2 Axis: An Integrated Strategy Combining Network Toxicology and Experimental Validation}},
journal = {Toxics},
year = {2026},
month = may,
volume = {14},
number = {5},
pages = {443},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {2305-6304},
doi = {10.3390/toxics14050443},
url = {https://doi.org/10.3390/toxics14050443},
pmid = {42198569},
pmcid = {PMC13211372}
}

RIS

TY - JOUR
AU - Xie, Ruijin
AU - Xiao, Wei
AU - Xu, Hua
AU - Luo, Yufan
AU - Xiao, Xue
AU - Pan, Qiyang
AU - Xu, Shengjie
AU - Liu, Li
AU - Sun, Chenyu
AU - Liu, Yueying
TI - 6PPDQ Exposure Exacerbates Seizure-Induced Neuronal Damage via the TP53/Nrf2 Axis: An Integrated Strategy Combining Network Toxicology and Experimental Validation
T2 - Toxics
J2 - Toxics
PY - 2026
DA - 2026/05/19
VL - 14
IS - 5
SP - 443
SN - 2305-6304
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/toxics14050443
UR - https://doi.org/10.3390/toxics14050443
LA - en
ER -

CSL-JSON

{
"id": "10.3390/toxics14050443",
"type": "article-journal",
"title": "6PPDQ Exposure Exacerbates Seizure-Induced Neuronal Damage via the TP53/Nrf2 Axis: An Integrated Strategy Combining Network Toxicology and Experimental Validation",
"container-title": "Toxics",
"author": [
{
"family": "Xie",
"given": "Ruijin"
},
{
"family": "Xiao",
"given": "Wei"
},
{
"family": "Xu",
"given": "Hua"
},
{
"family": "Luo",
"given": "Yufan"
},
{
"family": "Xiao",
"given": "Xue"
},
{
"family": "Pan",
"given": "Qiyang"
},
{
"family": "Xu",
"given": "Shengjie"
},
{
"family": "Liu",
"given": "Li"
},
{
"family": "Sun",
"given": "Chenyu"
},
{
"family": "Liu",
"given": "Yueying"
}
],
"container-title-short": "Toxics",
"volume": "14",
"issue": "5",
"page": "443",
"DOI": "10.3390/toxics14050443",
"PMID": "42198569",
"PMCID": "PMC13211372",
"ISSN": "2305-6304",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/toxics14050443",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
19
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s12951-026-04599-5 [code]
Long-term exposure to polystyrene microplastics exacerbates seizure symptoms via lipid metabolic disruption and ferroptosis: insights from multi-omics analyses.
Journal: Journal of nanobiotechnology
In common: XGBoost, ggplot2, scikit-learn, 3 other tools, epilepsy, cellular / molecular, 5 references
[2] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: randomForest, XGBoost, ggplot2, 5 other tools, cellular / molecular
[3] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: randomForest, pROC, ggplot2, 5 other tools
[4] doi:10.1073/pnas.2516601123 [code]
Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: pROC, XGBoost, ggplot2, 5 other tools
[5] doi:10.1016/j.phro.2026.101056 [code]
Toward uncertainty-aware manual delineation of brain tumours using eye-tracking and image-derived features.
Journal: Physics and imaging in radiation oncology
In common: randomForest, XGBoost, seaborn, 4 other tools
[6] doi:10.3390/ijms27156925 [code]
XGBoost-SHAP Interpretable Modeling Identifies and Validates an Eight-Gene Biomarker for Hepatic Encephalopathy Risk Prediction in Cirrhosis.
Journal: International journal of molecular sciences
In common: randomForest, pROC, XGBoost, 1 other tool
[7] doi:10.21037/jtd-2026-0997 [code]
Machine learning models based on XGBoost algorithm to predict prognosis of lung cancer brain metastases.
Journal: Journal of thoracic disease
In common: randomForest, pROC, XGBoost, 1 other tool
[8] doi:10.1371/journal.pcbi.1014327 [code]
Supervised deep learning with gene functional annotation for cell classification.
Journal: PLoS computational biology
In common: pROC, ggplot2, seaborn, 4 other tools, cellular / molecular
[9] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: pROC, ggplot2, seaborn, 4 other tools, cellular / molecular
[10] doi:10.3390/toxics14060504
Molecular Mechanisms of 6PPD and 6PPD-Q Toxicity in Neurodegenerative Diseases: A Network Toxicology and Experimental Validation Study.
Journal: Toxics
In common: cellular / molecular, 4 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.