OSCR

Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § Materials and methods › Analytical methodology ↔ onekey_algo/custom/components/comp1.py, lines 334–389 · score 0.62 · logistic regression, calibration curves, Python, predictors, Model
  2. [2] § Materials and methods › Development of the 2.5D deep learning framework ↔ onekey_core/models/densenet.py, lines 304–315 · score 0.61 · densenet201, ImageNet, pre trained, efficiency, model
  3. [3] § Materials and methods › Development of the 2.5D deep learning framework ↔ onekey_core/models/resnet.py, lines 310–319 · score 0.59 · resnet101, ImageNet, pre trained, deep, model

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,335 lines · 55 KB · no license · 1 match

  1. # -*- coding: UTF-8 -*-
  2. # Authorized by Vlon Jang
  3. # Created on 2021/12/22
  4. # Blog: www.wangqingbaidu.cn
  5. # Email: [email hidden]
  6. # Copyright 2015-2021 All Rights Reserved.
  7. import copy
  8. import math
  9. import os
  10. import platform
  11. import sys
  12. import warnings
  13. from typing import List, Union, Iterable
  14. import matplotlib.pyplot as plt
  15. import numpy as np
  16. import pandas as pd
  17. import scipy
  18. import seaborn as sns
  19. import statsmodels.api as sm
  20. from IPython.core.display import display
  21. from lightgbm import LGBMClassifier, LGBMRegressor
  22. # import matplotlib
  23. # matplotlib.use('Agg')
  24. from matplotlib.pyplot import MultipleLocator
  25. from pandas import DataFrame
  26. from scipy.stats import ttest_1samp
  27. from sklearn import manifold
  28. from sklearn.cluster import KMeans
  29. from sklearn.decomposition import PCA
  30. from sklearn.decomposition import TruncatedSVD
  31. from sklearn.ensemble import ExtraTreesClassifier, ExtraTreesRegressor, GradientBoostingClassifier, AdaBoostClassifier, \
  32. AdaBoostRegressor
  33. from sklearn.ensemble import RandomForestClassifier, RandomForestRegressor, GradientBoostingRegressor
  34. from sklearn.ensemble import RandomTreesEmbedding
  35. from sklearn.linear_model import LassoCV, LogisticRegression, ElasticNetCV, MultiTaskElasticNetCV, MultiTaskLassoCV
  36. from sklearn.metrics import accuracy_score
  37. from sklearn.metrics import roc_auc_score, confusion_matrix
  38. from sklearn.metrics import roc_curve, auc, mean_squared_error
  39. from sklearn.model_selection import train_test_split, StratifiedKFold
  40. from sklearn.naive_bayes import GaussianNB
  41. from sklearn.neighbors import KNeighborsClassifier, KNeighborsRegressor
  42. from sklearn.neighbors import NeighborhoodComponentsAnalysis
  43. from sklearn.neural_network import MLPClassifier, MLPRegressor
  44. from sklearn.pipeline import make_pipeline
  45. from sklearn.preprocessing import MinMaxScaler
  46. from sklearn.random_projection import SparseRandomProjection
  47. from sklearn.svm import SVC, SVR
  48. from sklearn.tree import DecisionTreeClassifier, DecisionTreeRegressor
  49. from xgboost import XGBClassifier, XGBRegressor
  50. from onekey_algo import get_param_in_cwd
  51. from onekey_algo.custom.components.delong import calc_95_CI
  52. from onekey_algo.custom.components.metrics import analysis_pred_binary
  53. from onekey_algo.utils import create_dir_if_not_exists
  54. from onekey_algo.utils.about_log import logger
  55. def init_CN():
  56. warnings.filterwarnings("ignore")
  57. plt.rcParams['figure.dpi'] = get_param_in_cwd('figure.dpi', 300)
  58. plt.rcParams['figure.figsize'] = get_param_in_cwd('figure.figsize', (10.0, 8.0))
  59. plt.rcParams['font.size'] = get_param_in_cwd('font.size', 15)
  60. plt.rcParams[
  61. 'font.sans-serif'] = 'Microsoft YaHei' if 'windows' in platform.platform().lower() else 'Arial Unicode MS'
  62. plt.rcParams['axes.unicode_minus'] = False
  63. pd.set_option('display.precision', get_param_in_cwd('display.precision', 3))
  64. def read_data(filename: str):
  65. assert os.path.exists(filename), f'{filename} 不存在!'
  66. if filename.endswith('.csv'):
  67. return pd.read_csv(filename, header=0)
  68. elif filename.endswith('.xlsx'):
  69. pd.read_excel(filename, header=0)
  70. else:
  71. raise ValueError(f'文件类型未定义!')
  72. def compress_feature(features: np.array, dim: int) -> np.array:
  73. assert len(features.shape) == 2, '特征的维度必须为2维'
  74. if dim > features.shape[1]:
  75. logger.warning(f"降维的维度({dim})不能多于样本数({features.shape[0]}), 使用{features.shape[1]}作为降维维度!")
  76. dim = features.shape[1]
  77. pca = PCA(dim)
  78. features = pca.fit_transform(features)
  79. return features
  80. def compress_df_feature(features: pd.DataFrame, dim: int, not_compress: Union[str, List[str]] = None,
  81. prefix='') -> pd.DataFrame:
  82. """
  83. 压缩深度学习特征
  84. Args:
  85. features: 特征DataFrame
  86. dim: 需要压缩到的维度,此值需要小于样本数
  87. not_compress: 不进行压缩的列。
  88. prefix: 所有特征的前缀。
  89. Returns:
  90. """
  91. if not_compress is not None:
  92. if isinstance(not_compress, str):
  93. not_compress = [not_compress]
  94. elif not isinstance(not_compress, Iterable):
  95. raise ValueError(f"not_compress设置出错!")
  96. not_compress_data = features[not_compress]
  97. features = features[[c for c in features.columns if c not in not_compress]]
  98. features = compress_feature(features, dim=dim)
  99. features = pd.DataFrame(features, columns=[f"{prefix}{i}" for i in range(features.shape[1])])
  100. return pd.concat([not_compress_data, features], axis=1)
  101. else:
  102. features = compress_feature(features, dim=dim)
  103. features = pd.DataFrame(features, columns=[f"{prefix}{i}" for i in range(features.shape[1])])
  104. return features
  105. def split_dataset(X_data, y_data, test_size=0.2, random_state=0, reset_index: bool = True):
  106. def __reset_index(d_):
  107. d_ = d_.reset_index()
  108. d_ = d_.drop(['index'], axis=1)
  109. return d_
  110. X_train, X_test, y_train, y_test = train_test_split(X_data, y_data, random_state=random_state, test_size=test_size)
  111. if reset_index:
  112. X_train = __reset_index(X_train)
  113. X_test = __reset_index(X_test)
  114. y_train = __reset_index(y_train)
  115. y_test = __reset_index(y_test)
  116. return X_train, X_test, y_train, y_test
  117. def cluster_analysis(features, n_clusters=2):
  118. cluster = KMeans(n_clusters=n_clusters, random_state=0)
  119. cluster.fit(features)
  120. return cluster
  121. def normalize_df(data: pd.DataFrame, not_norm: Union[str, List[str]] = None,
  122. method='gaussian', group: str = None, sub_group: Union[List, List[str]] = None) -> pd.DataFrame:
  123. """
  124. Normalize data frame
  125. Args:
  126. data: DataFrame to be Normalized.
  127. not_norm: Columns not be normalized.
  128. method: method to be used, choice: gaussian, min_max, default gaussian.
  129. group: Normalize by group separately, default None treated as a whole.
  130. sub_group: sub group to normalize.
  131. Returns: Normalized Dataframe
  132. """
  133. if not_norm is None:
  134. not_norm = []
  135. elif isinstance(not_norm, str):
  136. not_norm = [not_norm]
  137. if group is not None:
  138. not_norm = not_norm + [group]
  139. assert group in data.columns, f"group: {group} 分组没有在{data.columns}!"
  140. columns = [c for c in data.columns if c not in not_norm]
  141. new_data = data.copy(deep=True)
  142. def _norm_data(data_):
  143. desc = data_.describe()
  144. for column in columns:
  145. try:
  146. if method == 'gaussian':
  147. data_[column] = (data_[column] - desc[column]['mean']) / desc[column]['std']
  148. else:
  149. data_[column] = (data_[column] - desc[column]['min']) / (desc[column]['max'] - desc[column]['min'])
  150. except Exception as e:
  151. logger.warning(f"特征:{column}存在问题:{e},造成z-score失败!")
  152. return data_
  153. if group is None:
  154. _norm_data(new_data)
  155. else:
  156. ug_list = []
  157. ugroups = np.unique(data[group])
  158. for ug in ugroups:
  159. ug_data = new_data[new_data[group] == ug]
  160. if sub_group is None or ug in sub_group:
  161. _norm_data(ug_data)
  162. ug_list.append(ug_data)
  163. new_data = pd.concat(ug_list, axis=0)
  164. return new_data
  165. def convert2onehot(data, n_classes):
  166. data = np.reshape(data, -1)
  167. onehot_encoder = []
  168. for d in data:
  169. onehot = [0] * n_classes
  170. onehot[d] = 1
  171. onehot_encoder.append(onehot)
  172. return np.array(onehot_encoder)
  173. def draw_roc_per_class(y_test, y_score, n_classes, title='ROC per Class', include_spec_class: bool = True,
  174. mapping=None):
  175. """
  176. Args:
  177. mapping: label的映射
  178. y_test: 真实标签
  179. y_score: 预测标签
  180. n_classes: 类别数
  181. title: 标题
  182. include_spec_class: 是否包括每个细分标签的ROC曲线。
  183. Returns:
  184. """
  185. if mapping is None:
  186. mapping = {}
  187. y_test_binary = convert2onehot(y_test, n_classes=n_classes)
  188. # Compute ROC curve and ROC area for each class
  189. fpr = dict()
  190. tpr = dict()
  191. roc_auc = dict()
  192. # Compute micro-average ROC curve and ROC area
  193. # print(y_test_binary.ravel().shape, y_score.ravel().shape)
  194. fpr["micro"], tpr["micro"], _ = roc_curve(y_test_binary.ravel(), y_score.ravel())
  195. roc_auc["micro"] = auc(fpr["micro"], tpr["micro"])
  196. lw = 2
  197. plt.plot(
  198. fpr["micro"],
  199. tpr["micro"],
  200. label="micro-average ROC (area = {0:0.2f})".format(roc_auc["micro"]),
  201. color="deeppink",
  202. linestyle=":",
  203. linewidth=4,
  204. )
  205. try:
  206. for i in range(n_classes):
  207. try:
  208. fpr[i], tpr[i], _ = roc_curve(y_test_binary[:, i], y_score[:, i])
  209. roc_auc[i] = auc(fpr[i], tpr[i])
  210. except Exception as e:
  211. logger.error(f'解析{i}类别出错, {e}')
  212. # First aggregate all false positive rates
  213. all_fpr = np.unique(np.concatenate([fpr[i] for i in range(n_classes)]))
  214. # Then interpolate all ROC curves at this points
  215. mean_tpr = np.zeros_like(all_fpr)
  216. for i in range(n_classes):
  217. mean_tpr += np.interp(all_fpr, fpr[i], tpr[i])
  218. # Finally average it and compute AUC
  219. mean_tpr /= n_classes
  220. fpr["macro"] = all_fpr
  221. tpr["macro"] = mean_tpr
  222. roc_auc["macro"] = auc(fpr["macro"], tpr["macro"])
  223. plt.plot(
  224. fpr["macro"],
  225. tpr["macro"],
  226. label="macro-average ROC (area = {0:0.2f})".format(roc_auc["macro"]),
  227. color="navy",
  228. linestyle=":",
  229. linewidth=4,
  230. )
  231. if include_spec_class:
  232. for i in range(n_classes):
  233. plt.plot(
  234. fpr[i],
  235. tpr[i],
  236. lw=lw,
  237. label="class {0} ROC (area = {1:0.2f})".format(mapping[i] if i in mapping else i, roc_auc[i]),
  238. )
  239. except:
  240. logger.error(f'解析每个类别的ROC出错,大概率是应为数据没有指定类别的样本!')
  241. plt.plot([0, 1], [0, 1], "k--", lw=lw)
  242. plt.xlim([0.0, 1.0])
  243. plt.ylim([0.0, 1.05])
  244. plt.xlabel("1 - Specificity")
  245. plt.ylabel("Sensitivity")
  246. plt.title(title)
  247. plt.legend(loc="lower right")
  248. def draw_roc(y_test, y_score, title='ROC', labels=None):
  249. """
  250. 绘制ROC曲线
  251. Args:
  252. y_test: list或者array,为真实结果。
  253. y_score: list orray或者array,为模型预测结果。
  254. title: 图标题
  255. labels: 图例名称
  256. Returns:
  257. """
  258. precision = get_param_in_cwd('display.precision', 3)
  259. if not isinstance(y_test, (list, tuple)):
  260. y_test = [y_test]
  261. if not isinstance(y_score, (list, tuple)):
  262. y_score = [y_score]
  263. if labels is None:
  264. labels = [''] * len(y_score)
  265. assert len(y_test) == len(y_score) == len(labels)
  266. colors = ["deeppink", "navy", "aqua", "darkorange", "cornflowerblue"]
  267. ls = ['-', ':', '--', ':']
  268. for idx, (y_test_, y_score_, label) in enumerate(zip(y_test, y_score, labels)):
  269. y_score_ = np.array(y_score_)
  270. # enc = OneHotEncoder(handle_unknown='ignore')
  271. # y_test_binary = enc.fit_transform(y_test_.reshape(-1, 1)).toarray()
  272. if len(y_score_.shape) == 1:
  273. y_score_1 = y_score_
  274. else:
  275. y_score_1 = y_score_[:, 1]
  276. fpr, tpr, _ = roc_curve(y_test_, y_score_1)
  277. # print(y_test_, y_score_1)
  278. auc_, ci = calc_95_CI(np.squeeze(y_test_), y_score_1)
  279. roc_auc = auc(fpr, tpr)
  280. label_info = f"{label} AUC: {{roc_auc:.{precision}f}} (95%CI {{lo:.{precision}f}}-{{up:.{precision}f}})"
  281. label_info = label_info.format(roc_auc=roc_auc, lo=ci[0], up=ci[1])
  282. plt.plot(fpr, tpr, label=label_info,
  283. color=colors[idx % len(colors)], linestyle=ls[idx % len(ls)], linewidth=4)
  284. plt.plot([0, 1], [0, 1], "k--", lw=2)
  285. plt.xlim([0.0, 1.0])
  286. plt.ylim([0.0, 1.05])
  287. plt.xlabel("1 - Specificity")
  288. plt.ylabel("Sensitivity")
  289. plt.title(title)
  290. plt.legend(loc="lower right")
  291. def draw_calibration(y_test, pred_scores, model_names, remap: bool = False, **kwargs):
  292. """
  293. Args:
  294. y_test:
  295. pred_scores:
  296. model_names:
  297. remap: 是否使用逻辑回归重新映射。
  298. **kwargs:
  299. Returns:
  300. """
  301. version_info = sys.version_info
  302. assert version_info.major >= 3 and version_info.minor > 6, "Python版本必须3.7及以上。"
  303. if not isinstance(y_test, (tuple, list)):
  304. y_test = [y_test] * len(model_names)
  305. assert len(pred_scores) == len(model_names)
  306. from sklearn.calibration import CalibrationDisplay
  307. from matplotlib.gridspec import GridSpec
  308. gs = GridSpec(1, 1)
  309. fig = plt.figure(figsize=(10, 10))
  310. ax_calibration_curve = fig.add_subplot(gs[0, 0])
  311. def moving_average(interval, windowsize):
  312. window = np.ones(int(windowsize)) / float(windowsize)
  313. return np.convolve(interval, window, 'full')
  314. for model_name, scores, gt in zip(model_names, pred_scores, y_test):
  315. if remap:
  316. if kwargs.get('EX', None) is None:
  317. model = LogisticRegression()
  318. else:
  319. model = RandomForestClassifier(**kwargs.get('EX'))
  320. if len(scores.shape) == 1:
  321. scores = np.reshape(np.array(scores), [-1, 1])
  322. model.fit(scores, gt)
  323. else:
  324. model.fit(scores, gt)
  325. scores = model.predict_proba(scores)
  326. # display(pd.DataFrame(scores))
  327. scores = np.array(normalize_df(pd.DataFrame(scores), method='minmax'))
  328. if kwargs.get('smooth', False):
  329. if len(scores.shape) == 1:
  330. scores = scipy.signal.savgol_filter(scores, kwargs.get('window_length', 9), kwargs.get('k', 3))
  331. else:
  332. scores = scipy.signal.savgol_filter(scores[:, 1], kwargs.get('window_length', 9), kwargs.get('k', 3))
  333. scores = np.clip(scores, 0, 1)
  334. if len(scores.shape) == 1:
  335. pred = scores
  336. else:
  337. pred = scores[:, 1]
  338. disp = CalibrationDisplay.from_predictions(gt, pred,
  339. n_bins=kwargs.get('n_bins', 5),
  340. ax=ax_calibration_curve, name=model_name, strategy='quantile')
  341. def get_bst_split(X_data: pd.DataFrame, y_data: pd.DataFrame,
  342. models: dict, test_size=0.2, metric_fn=accuracy_score, n_trails=10,
  343. cv: bool = False, shuffle: bool = False, metric_cut_off: float = None, random_state=None,
  344. use_smote: bool = False, **kwargs):
  345. """
  346. 寻找数据集中最好的数据划分。
  347. Args:
  348. X_data: 训练数据
  349. y_data: 监督数据
  350. models: 模型名称,Dict类型、
  351. test_size: 测试集比例
  352. metric_fn: 评价模型好坏的函数,默认准确率,可选roc_auc_score。
  353. n_trails: 尝试多少次寻找最佳数据集划分。
  354. cv: 是否是交叉验证,默认是False,当为True时,n_trails为交叉验证的n_fold
  355. shuffle: 是否进行随机打乱
  356. metric_cut_off: 当metric_fn的值达到多少时进行截断。
  357. random_state: 随机种子
  358. use_smote: bool, 是否使用SMOTE技术,进行重采样。
  359. kwargs: 其他模型训练的参数。
  360. Returns: {'max_idx': max_idx, "max_model": max_model, "max_metric": max_metric, "results": results}
  361. """
  362. assert metric_fn in (roc_auc_score, accuracy_score, mean_squared_error)
  363. results = []
  364. max_model = None
  365. max_model_name = None
  366. max_idx = 0
  367. max_metric = None
  368. metrics = {}
  369. dataset = []
  370. if not isinstance(X_data, pd.DataFrame) or not isinstance(y_data, pd.DataFrame):
  371. X_data = pd.DataFrame(X_data)
  372. y_data = pd.DataFrame(y_data)
  373. logger.warning('你的数据不是DataFrame类型,可能遇到未知错误!')
  374. if cv:
  375. skf = StratifiedKFold(n_splits=n_trails, shuffle=shuffle or random_state is not None, random_state=random_state)
  376. for train_index, test_index in skf.split(X_data, y_data):
  377. X_train, X_test = X_data.loc[train_index], X_data.loc[test_index]
  378. y_train, y_test = y_data.loc[train_index], y_data.loc[test_index]
  379. dataset.append([X_train, X_test, y_train, y_test])
  380. for idx in range(n_trails):
  381. trail = []
  382. if cv:
  383. X_train, X_test, y_train, y_test = dataset[idx]
  384. else:
  385. rs = None if random_state is None else (idx + random_state)
  386. X_train, X_test, y_train, y_test = train_test_split(X_data, y_data, test_size=test_size, random_state=rs)
  387. X_train_smote, y_train_smote = X_train, y_train
  388. if use_smote:
  389. X_train_smote, y_train_smote = smote_resample(X_train, y_train)
  390. copy_models = copy.deepcopy(models)
  391. for model_name, model in copy_models.items():
  392. # model.fit(X_train, y_train)
  393. # sample_weight = [1 if i == 0 else 0.5 for i in list(np.array(y_train))]
  394. if kwargs:
  395. try:
  396. model.fit(X_train_smote, y_train_smote, **kwargs)
  397. logger.info(f'正在训练{model_name}, 使用{kwargs}。')
  398. except Exception as e:
  399. model.fit(X_train_smote, y_train_smote)
  400. logger.warning(f'因为:{e},训练{model_name}使用{kwargs}失败。')
  401. else:
  402. model.fit(X_train_smote, y_train_smote)
  403. y_pred = model.predict(X_test)
  404. if metric_fn == roc_auc_score:
  405. y_proba = model.predict_proba(X_test)[:, 1]
  406. metric = metric_fn(y_test, y_proba)
  407. else:
  408. metric = metric_fn(y_test, y_pred)
  409. if model_name not in metrics:
  410. metrics[model_name] = []
  411. metrics[model_name].append((idx, metric))
  412. if max_metric is None or metric > max_metric:
  413. max_metric = metric
  414. max_idx = idx
  415. max_model = model
  416. max_model_name = model_name
  417. trail.append(metric)
  418. results.append((trail, (X_train, X_test, y_train, y_test)))
  419. # 当满足用户需求的时候,可以停止。
  420. if metric_cut_off is not None and max_metric is not None and max_metric > metric_cut_off:
  421. logger.info(f'Get best split cut off on {idx + 1} trails!')
  422. break
  423. return {'max_idx': max_idx, "max_model": max_model, "max_metric": max_metric, 'max_model_name': max_model_name,
  424. "results": results, 'metrics': metrics}
  425. def get_cv_metric_binary_task(models, data_splits, labels, model_names=None):
  426. """
  427. 生成CV结果,只对二分类模型有效。
  428. Args:
  429. models: 需要使用的模型
  430. data_splits: 数据划分
  431. labels: 数据对应任务的label
  432. model_names: 模型名称
  433. Returns: 所有交叉验证的结果。
  434. """
  435. metric = []
  436. for cv_index, (X_train_sel, X_test_sel, y_train_sel, y_test_sel) in enumerate(data_splits):
  437. predictions = [[(model.predict(X_train_sel), model.predict(X_test_sel))
  438. for model in target] for label, target in zip(labels, models)]
  439. pred_scores = [[(model.predict_proba(X_train_sel), model.predict_proba(X_test_sel))
  440. for model in target] for label, target in zip(labels, models)]
  441. pred_sel_idx = []
  442. for model, label, prediction, scores in zip(models, labels, predictions, pred_scores):
  443. pred_sel_idx_label = []
  444. if model_names is None:
  445. model_names = [str(m.__class__) for m in model]
  446. assert len(prediction) == len(model_names), "模型名称必须与模型长度相同"
  447. for mname, (train_pred, test_pred), (train_score, test_score) in zip(model_names, prediction, scores):
  448. # 计算训练集指数
  449. acc, auc, ci, tpr, tnr, ppv, npv, precision, recall, f1, thres = analysis_pred_binary(
  450. y_train_sel[label],
  451. train_score[:, 1])
  452. ci = f"{ci[0]:.4f} - {ci[1]:.4f}"
  453. metric.append((mname, acc, auc, ci, tpr, tnr, ppv, npv, precision, recall, f1, thres,
  454. f"CV-{cv_index}-{label}-train"))
  455. # 计算验证集指标
  456. acc, auc, ci, tpr, tnr, ppv, npv, precision, recall, f1, thres = analysis_pred_binary(y_test_sel[label],
  457. test_score[:, 1])
  458. ci = f"{ci[0]:.4f} - {ci[1]:.4f}"
  459. metric.append((mname, acc, auc, ci, tpr, tnr, ppv, npv, precision, recall, f1, thres,
  460. f"CV-{cv_index}-{label}-test"))
  461. # 计算thres对应的sel idx
  462. pred_sel_idx_label.append(np.logical_or(test_score[:, 0] >= thres, test_score[:, 1] >= thres))
  463. pred_sel_idx.append(pred_sel_idx_label)
  464. metric = pd.DataFrame(metric, index=None, columns=['model_name', 'Accuracy', 'AUC', '95% CI',
  465. 'Sensitivity', 'Specificity',
  466. 'PPV', 'NPV', 'Precision', 'Recall', 'F1',
  467. 'Threshold', 'Task'])
  468. return metric
  469. def create_clf_model(model_names):
  470. models = {}
  471. # 判断是纯字符串,使用默认参数进行配置
  472. if isinstance(model_names, (list, tuple)):
  473. if 'lr' in model_names or 'LR' in model_names:
  474. models['LR'] = LogisticRegression(random_state=0)
  475. # SVM
  476. if 'svm' in model_names or 'SVM' in model_names:
  477. models['SVM'] = SVC(probability=True, random_state=0)
  478. # KNN
  479. if 'knn' in model_names or 'KNN' in model_names:
  480. models['KNN'] = KNeighborsClassifier(algorithm='kd_tree')
  481. # RandomForest
  482. if 'rf' in model_names or 'RandomForest' in model_names:
  483. models['RandomForest'] = RandomForestClassifier(n_estimators=10, max_depth=None,
  484. min_samples_split=2, random_state=0)
  485. elif isinstance(model_names, dict):
  486. for model_name, params in model_names.items():
  487. if 'svm' == model_name or 'SVM' == model_name:
  488. models['SVM'] = SVC(**params)
  489. # KNN
  490. if 'knn' == model_name or 'KNN' == model_name:
  491. models['KNN'] = KNeighborsClassifier(**params)
  492. # RandomForest
  493. if 'rf' == model_name or 'RandomForest' == model_name:
  494. models['RandomForest'] = RandomForestClassifier(**params)
  495. return models
  496. def create_reg_model(model_names):
  497. models = {}
  498. # 判断是纯字符串,使用默认参数进行配置
  499. if isinstance(model_names, (list, tuple)):
  500. if 'svm' in model_names or 'SVM' in model_names:
  501. models['SVM'] = SVR()
  502. # KNN
  503. if 'knn' in model_names or 'KNN' in model_names:
  504. models['KNN'] = KNeighborsRegressor(algorithm='kd_tree')
  505. # DecisionTree
  506. if 'dt' in model_names or 'DecisionTree' in model_names:
  507. models['DecisionTree'] = DecisionTreeRegressor(max_depth=None,
  508. min_samples_split=2, random_state=0)
  509. # RandomForest
  510. if 'rf' in model_names or 'RandomForest' in model_names:
  511. models['RandomForest'] = RandomForestRegressor(n_estimators=10, max_depth=None,
  512. min_samples_split=2, random_state=0)
  513. elif isinstance(model_names, dict):
  514. for model_name, params in model_names.items():
  515. if 'svm' == model_name or 'SVM' == model_name:
  516. models['SVM'] = SVR(**params)
  517. # KNN
  518. if 'knn' == model_name or 'KNN' == model_name:
  519. models['KNN'] = KNeighborsRegressor(**params)
  520. # DecisionTree
  521. if 'dt' == model_name or 'DecisionTree' == model_name:
  522. models['DecisionTree'] = DecisionTreeRegressor(**params)
  523. # RandomForest
  524. if 'rf' == model_name or 'RandomForest' == model_name:
  525. models['RandomForest'] = RandomForestRegressor(**params)
  526. return models
  527. def calc_confusion_matrix(prediction: List[int], gt: List[int], sel_idx: Union[List[bool], np.ndarray] = None,
  528. class_mapping: Union[str, dict, list, tuple] = None, num_classes: int = None):
  529. """
  530. Args:
  531. prediction: Prediction of each results.
  532. gt: Ground truth of each results.
  533. sel_idx: Use which index of data to calculate cm. default None for all.
  534. class_mapping: mapping class index to readable classes.
  535. num_classes: Number of classes.
  536. Returns:
  537. """
  538. num_classes = num_classes or len(set(gt))
  539. if num_classes != len(set(gt)):
  540. logger.warning(f'num_classes({num_classes}) is not equal to labels in gt({len(set(gt))}).')
  541. cm = np.zeros((num_classes, num_classes))
  542. if sel_idx is not None:
  543. logger.info(f"使用筛选阈值的数据绘制混淆矩阵!样本量从{len(gt)}变到{sum(sel_idx)}.")
  544. prediction = prediction[sel_idx]
  545. gt = gt[sel_idx]
  546. for pred, y in zip(prediction, gt):
  547. # cm[int(pred), int(y)] += 1
  548. cm[int(y), int(pred)] += 1
  549. mapping = {}
  550. if isinstance(class_mapping, dict):
  551. mapping = class_mapping
  552. elif isinstance(class_mapping, (list, tuple)):
  553. mapping = dict(enumerate(class_mapping))
  554. elif class_mapping and os.path.exists(class_mapping):
  555. label_names = [l_.strip() for l_ in open(class_mapping).readlines()]
  556. mapping = {i: l for i, l in enumerate(label_names)}
  557. labels = [mapping[i] if i in mapping else f"label_{i}" for i in range(num_classes)]
  558. return pd.DataFrame(cm, index=labels, columns=labels)
  559. def draw_matrix(data: pd.DataFrame, norm: bool = False, **kwargs):
  560. if norm:
  561. data = data.div(np.sum(data, axis=1), axis=0)
  562. if 'fmt' not in kwargs:
  563. kwargs['fmt'] = ".2g" if norm else ".3f"
  564. sns.heatmap(data, **kwargs)
  565. def draw_predict_score(pred_score, y_test, threshold=0.5):
  566. threshold = float(threshold)
  567. if np.array(pred_score).shape[1] == 1:
  568. ps = pred_score
  569. else:
  570. ps = pred_score[:, 1]
  571. d = pd.concat([pd.DataFrame(ps), y_test.reset_index()], axis=1)
  572. # plt.axis('off')
  573. d.columns = ['predict_score', 'index', 'label']
  574. d['predict_score'] = (d['predict_score'] - threshold)
  575. d['predict_score'] = ((d['predict_score'] - d['predict_score'].min()) / (
  576. d['predict_score'].max() - d['predict_score'].min()) - 0.5) * 2
  577. d = d.sort_values('predict_score')
  578. d['index'] = range(d.shape[0])
  579. plt.bar(d[d['label'] == 0]['index'], d[d['label'] == 0]["predict_score"])
  580. plt.bar(d[d['label'] == 1]['index'], d[d['label'] == 1]["predict_score"])
  581. def draw_cv_box(data, x, y, hue=None, **kwargs):
  582. if hue:
  583. sns.boxplot(data=data, x=x, y=y, hue=hue, **kwargs)
  584. else:
  585. sns.boxplot(data=data, x=x, y=y, **kwargs)
  586. def analysis_features(rad_features, labels, not_use: Union[List[str], str] = ['ID', 'label'],
  587. methods: Union[List[str], str] = 't-SNE', n_neighbors=30, save_dir=None, prefix=''):
  588. """
  589. Args:
  590. rad_features: 待分析的特征
  591. labels: 每个样本对饮的标签
  592. not_use: 那些列不使用
  593. methods: 选择可视化的方法, Random projection, Truncated SVD, Isomap, Standard LLE, Modified LLE,
  594. MDS, Random Trees, Spectral, t-SNE, NCA
  595. n_neighbors:
  596. save_dir: 保存目录
  597. prefix: 前缀
  598. Returns:
  599. """
  600. if not isinstance(not_use, (list, tuple)):
  601. not_use = [not_use]
  602. if methods is not None and not isinstance(methods, (list, tuple)):
  603. methods = [methods]
  604. data = np.array(rad_features[[c for c in rad_features.columns if c not in not_use]])
  605. embeddings = {
  606. "Random projection": SparseRandomProjection(n_components=2, random_state=42),
  607. "Truncated SVD": TruncatedSVD(n_components=2),
  608. "Isomap": manifold.Isomap(n_neighbors=n_neighbors, n_components=2),
  609. "Standard LLE": manifold.LocallyLinearEmbedding(n_neighbors=n_neighbors, n_components=2, method="standard"),
  610. "Modified LLE": manifold.LocallyLinearEmbedding(n_neighbors=n_neighbors, n_components=2, method="modified"),
  611. # "Hessian LLE": manifold.LocallyLinearEmbedding(n_neighbors=n_neighbors, n_components=2, method="hessian"),
  612. # "LTSA LLE": manifold.LocallyLinearEmbedding(n_neighbors=n_neighbors, n_components=2, method="ltsa"),
  613. "MDS": manifold.MDS(n_components=2, n_init=1, max_iter=120, n_jobs=2),
  614. "Random Trees": make_pipeline(RandomTreesEmbedding(n_estimators=200, max_depth=5, random_state=0),
  615. TruncatedSVD(n_components=2), ),
  616. "Spectral": manifold.SpectralEmbedding(n_components=2, random_state=0, eigen_solver="arpack"),
  617. "t-SNE": manifold.TSNE(n_components=2, init="pca", learning_rate="auto", n_iter=500,
  618. n_iter_without_progress=150, n_jobs=2, random_state=0, ),
  619. "NCA": NeighborhoodComponentsAnalysis(n_components=2, init="pca", random_state=0),
  620. }
  621. projections = {}
  622. for name, transformer in embeddings.items():
  623. if methods is None or name in methods:
  624. try:
  625. projections[name] = transformer.fit_transform(data, labels)
  626. except Exception as e:
  627. logger.error(f"使用{name}进行降维可视化过程中出错。{e}")
  628. for name in projections:
  629. X = MinMaxScaler().fit_transform(projections[name])
  630. color_list = sns.color_palette("hls", len(np.unique(labels)))
  631. for idx, label in enumerate(np.unique(labels)):
  632. plt.scatter(*X[labels == label].T, s=60, alpha=0.425, zorder=2, color=color_list[idx])
  633. plt.legend(labels=np.unique(labels), loc="lower right")
  634. plt.title(f'Method: {name}')
  635. if save_dir is not None:
  636. try:
  637. os.makedirs(save_dir, exist_ok=True)
  638. plt.savefig(os.path.join(save_dir, f'{prefix}feature_{name}_viz.svg'), bbox_inches='tight')
  639. except Exception as e:
  640. logger.error(f'保存{name}可视化功能失效,因为:「{e}」')
  641. plt.show()
  642. def select_feature(corr, threshold: float = 0.9, keep: int = 1, verbose: bool = False, topn: int = 1):
  643. feature_names = corr.columns
  644. drop_feature_names = [x for x, y in np.array(corr.isna().all().reset_index()) if y]
  645. has_corr_features = True
  646. while has_corr_features:
  647. has_corr_features = False
  648. corr_fname = {}
  649. feature2drop = [fname for fname in feature_names if fname not in drop_feature_names]
  650. for i, fi in enumerate(feature2drop):
  651. corr_num = 0
  652. for j in range(i + 1, len(feature2drop)):
  653. if abs(corr[fi][feature2drop[j]]) > threshold:
  654. corr_num += 1
  655. corr_fname[fi] = corr_num
  656. corr_fname = sorted(corr_fname.items(), key=lambda x: x[1], reverse=True)
  657. for fname, corr_num in corr_fname[:topn]:
  658. if corr_num >= keep:
  659. has_corr_features = True
  660. drop_feature_names.append(fname)
  661. if verbose:
  662. logger.info(f'len {len(feature2drop)}, {fname} has {corr_num} features')
  663. return [fname for fname in feature_names if fname not in drop_feature_names]
  664. def select_feature_ttest(data: pd.DataFrame, popmean: Union[float, np.ndarray], threshold: float = 0.05):
  665. feature_names = data.columns
  666. if not isinstance(popmean, (list, tuple)):
  667. popmean = [popmean] * len(feature_names)
  668. elif len(popmean) != len(feature_names):
  669. raise ValueError('mean is not equal to feature length!')
  670. res = ttest_1samp(data, popmean)
  671. return [fname for fname, flag in zip(feature_names, res.pvalue < threshold) if flag]
  672. def lasso_cv_coefs(X_data, y_data, alpha_logmin=-3, points=50, column_names: List[str] = None, cv: int = 10,
  673. ensure_lastn: int = None, model_name: str = 'lasso', force_alpha=None,
  674. random_state=0, **kwargs):
  675. """
  676. Args:
  677. X_data: 训练数据
  678. y_data: 监督数据
  679. alpha_logmin: alpha的log最小值
  680. points: 打印多少个点。默认50
  681. column_names: 列名,默认为None,当选择的数据很多的时候,建议不要添加此参数
  682. cv: 交叉验证次数。
  683. ensure_lastn: bool, 确保是最后一个MSE不降的alpha
  684. model_name: 可选使用Lasso还是ElasticNet。
  685. force_alpha: deprecated!
  686. random_state: 0
  687. **kwargs: 其他用于打印控制的参数。
  688. """
  689. # 每个特征值随lambda的变化
  690. alphas = np.logspace(alpha_logmin, 0, points)
  691. if points != 50 or cv != 10:
  692. print(f"Points: {points}, CV: {cv}")
  693. is_multi_task = len(np.array(y_data).shape) > 1 and np.array(y_data).shape[1] > 1
  694. if model_name.lower() == 'lasso' and is_multi_task:
  695. lasso_cv = MultiTaskLassoCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  696. elif model_name.lower() == 'lasso' and not is_multi_task:
  697. lasso_cv = LassoCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  698. elif is_multi_task:
  699. lasso_cv = MultiTaskElasticNetCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  700. else:
  701. lasso_cv = ElasticNetCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  702. _, coefs, _ = lasso_cv.path(X_data, y_data, alphas=alphas)
  703. coefs = np.squeeze(coefs).T
  704. MSEs = lasso_cv.mse_path_
  705. MSEs_mean = np.mean(MSEs, axis=1)
  706. # plt.rcParams['font.sans-serif'] = 'stixgeneral'
  707. def draw_cv(coefs_):
  708. plt.semilogx(lasso_cv.alphas_, coefs_, '-', **kwargs)
  709. if column_names is not None:
  710. # print(column_names, coefs.shape[1])
  711. assert len(column_names) == coefs.shape[1]
  712. plt.legend(labels=column_names, loc='best')
  713. if ensure_lastn is not None and lasso_cv.alpha_ == 1.0:
  714. lasso_cv.alpha_ = alphas[-ensure_lastn]
  715. logger.warning(f'你正在使用ensure_lastn={ensure_lastn}...')
  716. lambda_info = ''
  717. if force_alpha is not None:
  718. lasso_cv.alpha_ = force_alpha
  719. if lasso_cv.alpha_ != 1.0:
  720. plt.axvline(lasso_cv.alpha_, color='black', ls="--", **kwargs)
  721. precision = get_param_in_cwd('display.precision', 4)
  722. lambda_info = f"(λ={{l:.{precision}f}})"
  723. lambda_info = lambda_info.format(l=lasso_cv.alpha_)
  724. plt.xlabel(f'Lambda{lambda_info}')
  725. plt.ylabel('Coefficients')
  726. if is_multi_task:
  727. num_tasks = np.array(y_data).shape[1]
  728. plt.figure(figsize=(12 * num_tasks, 10))
  729. for task_idx in range(num_tasks):
  730. plt.subplot(1, num_tasks, task_idx + 1)
  731. draw_cv(coefs[:, :, task_idx])
  732. # plt.show()
  733. else:
  734. draw_cv(coefs)
  735. return lasso_cv.alpha_
  736. def lasso_cv_efficiency(X_data, y_data, alpha_logmin=-3, points=50, cv: int = 10, ensure_lastn: int = None,
  737. model_name: str = 'lasso', force_alpha=None, random_state=0, **kwargs):
  738. """
  739. Args:
  740. Xdata: 训练数据
  741. ydata: 测试数据
  742. alpha_logmin: alpha的log最小值
  743. points: 打印的数据密度
  744. cv: 交叉验证次数
  745. ensure_lastn: bool, 确保是最后一个MSE不降的alpha
  746. model_name: 可选使用Lasso还是ElasticNet。
  747. force_alpha: deprecated!
  748. random_state: random state
  749. **kwargs: 其他的图像样式
  750. # 数据点标记, fmt="o"
  751. # 数据点大小, ms=3
  752. # 数据点颜色, mfc="r"
  753. # 数据点边缘颜色, mec="r"
  754. # 误差棒颜色, ecolor="b"
  755. # 误差棒线宽, elinewidth=2
  756. # 误差棒边界线长度, capsize=2
  757. # 误差棒边界厚度, capthick=1
  758. Returns:
  759. """
  760. precision = get_param_in_cwd('display.precision', 4)
  761. alphas = np.logspace(alpha_logmin, 0, points)
  762. is_multi_task = np.array(y_data).shape[1] > 1
  763. if model_name.lower() == 'lasso' and is_multi_task:
  764. lasso_cv = MultiTaskLassoCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  765. elif model_name.lower() == 'lasso' and not is_multi_task:
  766. lasso_cv = LassoCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  767. elif is_multi_task:
  768. lasso_cv = MultiTaskElasticNetCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  769. else:
  770. lasso_cv = ElasticNetCV(alphas=alphas, cv=cv, n_jobs=-1, random_state=random_state).fit(X_data, y_data)
  771. MSEs = lasso_cv.mse_path_
  772. MSEs_mean = np.mean(MSEs, axis=1)
  773. MSEs_std = np.std(MSEs, axis=1)
  774. # plt.rcParams['figure.figsize'] = (10.0, 8.0)
  775. default_params = {'fmt': "o", 'ms': 3, 'mfc': 'r', 'mec': 'r', 'ecolor': 'b', 'elinewidth': 2, 'capsize': 2,
  776. 'capthick': 1}
  777. default_params.update(kwargs)
  778. plt.errorbar(lasso_cv.alphas_, MSEs_mean, yerr=MSEs_std, **default_params)
  779. plt.semilogx()
  780. if ensure_lastn is not None and lasso_cv.alpha_ == 1.0:
  781. lasso_cv.alpha_ = alphas[-ensure_lastn]
  782. logger.warning(f'你正在使用ensure_lastn={ensure_lastn}...')
  783. lambda_info = ''
  784. if force_alpha is not None:
  785. lasso_cv.alpha_ = force_alpha
  786. if lasso_cv.alpha_ != 1.0:
  787. plt.axvline(lasso_cv.alpha_, color='black', ls="--", **kwargs)
  788. lambda_info = f"(λ={{l:.{precision}f}})"
  789. lambda_info = lambda_info.format(l=lasso_cv.alpha_)
  790. # lambda_info = f"(λ={:.6f})"
  791. plt.xlabel(f'Lambda{lambda_info}')
  792. plt.ylabel('MSE')
  793. ax = plt.gca()
  794. y_major_locator = MultipleLocator(0.1)
  795. ax.yaxis.set_major_locator(y_major_locator)
  796. # plt.show()
  797. def calc_icc(Y, icc_type="icc(3,1)"):
  798. """
  799. Args:
  800. Y: 待计算的数据
  801. icc_type: 共支持 icc(2,1), icc(2,k), icc(3,1), icc(3,k)四种
  802. """
  803. [n, k] = Y.shape
  804. # Degrees of Freedom
  805. dfc = k - 1
  806. dfe = (n - 1) * (k - 1)
  807. dfr = n - 1
  808. # Sum Square Total
  809. mean_Y = np.mean(Y)
  810. SST = ((Y - mean_Y) ** 2).sum()
  811. # create the design matrix for the different levels
  812. x = np.kron(np.eye(k), np.ones((n, 1))) # sessions
  813. x0 = np.tile(np.eye(n), (k, 1)) # subjects
  814. X = np.hstack([x, x0])
  815. # Sum Square Error
  816. predicted_Y = np.dot(
  817. np.dot(np.dot(X, np.linalg.pinv(np.dot(X.T, X))), X.T), Y.flatten("F")
  818. )
  819. residuals = Y.flatten("F") - predicted_Y
  820. SSE = (residuals ** 2).sum()
  821. MSE = SSE / dfe
  822. # Sum square column effect - between colums
  823. SSC = ((np.mean(Y, 0) - mean_Y) ** 2).sum() * n
  824. # MSC = SSC / dfc / n
  825. MSC = SSC / dfc
  826. # Sum Square subject effect - between rows/subjects
  827. SSR = SST - SSC - SSE
  828. MSR = SSR / dfr
  829. if icc_type == "icc(2,1)" or icc_type == 'icc(2,k)':
  830. if icc_type == 'icc(2,k)':
  831. k = 1
  832. ICC = (MSR - MSE) / (MSR + (k - 1) * MSE + k * (MSC - MSE) / n + 1e-6)
  833. elif icc_type == "icc(3,1)" or icc_type == 'icc(3,k)':
  834. if icc_type == 'icc(3,k)':
  835. k = 1
  836. # print(SSR, dfr, SSE, dfe, MSR, MSE, MSR - MSE)
  837. ICC = (MSR - MSE) / (MSR + (k - 1) * MSE + 1e-6)
  838. else:
  839. raise ValueError(f"icc_type{icc_type} 没有找到!")
  840. return max(0, ICC)
  841. def icc_filter_features(dfs: List[DataFrame], threshold: float = 0.9, with_icc: bool = False,
  842. columns_start: int = 0, not_icc: List[str] = None) -> Union[list, dict]:
  843. """
  844. ICC校验,取大于threshold阈值的特征
  845. Args:
  846. dfs: DataFrame,一般对应于不同人标注的结果。
  847. threshold: ICC阈值,大于此阈值返回。默认0.9,此值为None,返回每个特征的ICC dict.
  848. with_icc: Boolean, 是否返回每个特征具体的ICC值。
  849. columns_start: int,筛选的特征起始列。
  850. not_icc: 不进行ICC的列名。
  851. Returns:
  852. """
  853. assert isinstance(dfs, (list, tuple)) and len(dfs) > 1, "做ICC校验的数据至少需要2组数据。"
  854. assert all(isinstance(df, DataFrame) for df in dfs), "所有的数据必须是DataFrame。"
  855. assert all(dfs[0].shape == df.shape and len(df.shape) == 2 for df in dfs), '所有的数据维度必须相同'
  856. columns = dfs[0].columns
  857. data = np.array(dfs)
  858. selected_features = {}
  859. # if isinstance(not_icc, (list, tuple)):
  860. # not_icc = [not_icc]
  861. for idx, column in enumerate(columns):
  862. if not_icc is not None and column in not_icc:
  863. logger.info(f'{column}不进行ICC。')
  864. continue
  865. if idx >= columns_start:
  866. try:
  867. icc = calc_icc(data[:, :, idx].T)
  868. selected_features[column] = icc
  869. except Exception as e:
  870. logger.error(f"{column}计算失败,因为:{e}")
  871. if threshold is not None:
  872. sel_features = [k for k, v in selected_features.items() if v > threshold]
  873. else:
  874. sel_features = selected_features
  875. if with_icc:
  876. return sel_features, selected_features
  877. else:
  878. return sel_features
  879. def calculate_net_benefit_model(thresh_group, y_pred_score, y_label):
  880. net_benefit_model = np.array([])
  881. for thresh in thresh_group:
  882. y_pred_label = y_pred_score > thresh
  883. tn, fp, fn, tp = confusion_matrix(y_label, y_pred_label).ravel()
  884. n = len(y_label)
  885. net_benefit = (tp / n) - (fp / n) * (thresh / (1 - thresh))
  886. net_benefit_model = np.append(net_benefit_model, net_benefit)
  887. return net_benefit_model
  888. def calculate_net_benefit_all(thresh_group, y_label):
  889. net_benefit_all = np.array([])
  890. tn, fp, fn, tp = confusion_matrix(y_label, y_label).ravel()
  891. total = tp + tn
  892. for thresh in thresh_group:
  893. net_benefit = (tp / total) - (tn / total) * (thresh / (1 - thresh))
  894. net_benefit_all = np.append(net_benefit_all, net_benefit)
  895. return net_benefit_all
  896. def plot_DCA(y_pred_score, y_label, title='Model DCA', labels=None, y_min=None):
  897. """
  898. plot DCA曲线
  899. Args:
  900. y_pred_score: list或者array,为模型预测结果。
  901. y_label: list或者array,为真实结果。
  902. title: 图标题
  903. labels: 图例名称
  904. y_min: 最小值
  905. Returns:
  906. """
  907. thresh_group = np.arange(0, 1, 0.01)
  908. fig, ax = plt.subplots()
  909. net_benefit_all = calculate_net_benefit_all(thresh_group, y_label)
  910. net_benefit_model = None
  911. if not isinstance(y_pred_score, (tuple, list)):
  912. y_pred_score = [y_pred_score]
  913. if not isinstance(labels, (tuple, list)):
  914. labels = [labels]
  915. assert len(y_pred_score) == len(labels)
  916. for y_pred, label in zip(y_pred_score, labels):
  917. net_benefit_model = calculate_net_benefit_model(thresh_group, y_pred, y_label)
  918. ax.plot(thresh_group, net_benefit_model, label=label if label is not None else 'Model')
  919. ax.plot(thresh_group, net_benefit_all, color='black', label='Treat all')
  920. ax.plot((0, 1), (0, 0), color='black', linestyle=':', label='Treat none')
  921. if len(y_pred_score) == 1:
  922. # Fill,显示出模型较于treat all和treat none好的部分
  923. y2 = np.maximum(net_benefit_all, 0)
  924. y1 = np.maximum(net_benefit_model, y2)
  925. ax.fill_between(thresh_group, y1, y2, color='crimson', alpha=0.2)
  926. # Figure Configuration, 美化一下细节
  927. ax.set_xlim(0, 1)
  928. if y_min is not None:
  929. y_min = max(y_min, net_benefit_model.min() - 0.15)
  930. else:
  931. y_min = net_benefit_model.min() - 0.15
  932. ax.set_ylim(y_min, net_benefit_model.max() + 0.15) # justify the y axis limitation
  933. ax.set_xlabel(
  934. xlabel='Threshold Probability',
  935. fontdict={'family': 'Times New Roman', 'fontsize': 15}
  936. )
  937. ax.set_ylabel(
  938. ylabel='Net Benefit',
  939. fontdict={'family': 'Times New Roman', 'fontsize': 15}
  940. )
  941. ax.grid('major')
  942. ax.spines['right'].set_color((0.8, 0.8, 0.8))
  943. ax.spines['top'].set_color((0.8, 0.8, 0.8))
  944. ax.legend(loc='upper right')
  945. plt.title(title)
  946. def fillna(x: pd.DataFrame, inplace: bool = False, fill_mod: str = 'best', verobse: bool = False):
  947. """
  948. 填充样本空值
  949. Args:
  950. x: 待填充的数据DataFrame
  951. inplace: 是否直接改变原始x的结果
  952. fill_mod: 填充模式,目前支持制定数值填充,以及
  953. * best: 自动寻找最优方式,连续值填充均值,离散值填充中位数。
  954. * 50%: 中位数填充。
  955. * mean: 均值填充。
  956. verobse: 是否输出日志
  957. Returns:
  958. """
  959. assert isinstance(fill_mod, (int, float)) or fill_mod in ['best', '50%', 'mean']
  960. if inplace:
  961. data = x
  962. else:
  963. data = copy.deepcopy(x)
  964. desc = x.describe()
  965. for c in desc.columns:
  966. if fill_mod == 'best':
  967. if data[c].dtype == float:
  968. if verobse:
  969. logger.info(f"{c}使用【平均值】进行填充")
  970. fill_value = desc[c]['mean']
  971. else:
  972. if verobse:
  973. logger.info(f"{c}使用【中位数】进行填充")
  974. fill_value = desc[c]['50%']
  975. elif isinstance(fill_mod, (float, int)):
  976. fill_value = fill_mod
  977. else:
  978. fill_value = desc[c][fill_mod]
  979. data[c] = data[c].fillna(fill_value)
  980. return data
  981. def plot_feature_importance(model, feature_names=None, save_dir=None, prefix='', trace: bool = False, topk=None):
  982. """
  983. Args:
  984. model: 模型类
  985. feature_names: 特征可读名称。
  986. save_dir: 保存图像目录
  987. prefix: prefix, 默认为空。
  988. topk: int, 最重要的k个特征
  989. trace: 是否跟踪错误信息。
  990. Returns:
  991. """
  992. try:
  993. imp = model.feature_importances_
  994. if feature_names is None:
  995. feature_names = [f"f{i}" for i in range(len(imp))]
  996. imp_ = sorted(list(np.array([feature_names, imp]).T), key=lambda x: x[1], reverse=True)[:topk]
  997. importance = pd.DataFrame(imp_, columns=['feature_name', 'importance'])
  998. importance['importance'] = importance['importance'].astype(float)
  999. sns.barplot(y="feature_name", x="importance", data=importance)
  1000. plt.title(model.__class__.__name__)
  1001. if save_dir is not None:
  1002. os.makedirs(save_dir, exist_ok=True)
  1003. plt.savefig(os.path.join(save_dir, f'{prefix}model_{model.__class__.__name__}_feature_importance.svg'),
  1004. bbox_inches='tight')
  1005. importance.to_csv(os.path.join(save_dir, f'{prefix}model_{model.__class__.__name__}_importance_weight.csv'),
  1006. index=False)
  1007. plt.show()
  1008. return imp_
  1009. except:
  1010. if trace:
  1011. import traceback
  1012. traceback.print_exc()
  1013. pass
  1014. def variable_analysis(data: pd.DataFrame, features: Union[str, List[str]] = None, label_column: str = 'label',
  1015. need_norm: Union[bool, List[bool]] = False, alpha=0.1, method='uni'):
  1016. """
  1017. 进行单因素或者多因素分析,请确保所有的分析特征都是数值类型的。
  1018. Args:
  1019. data: 数据
  1020. features: 需要分析的特征,默认除了ID和label_column列,其他的特征都进行分析。
  1021. label_column: 目标列
  1022. need_norm: 是否标准化所有分析的数据, 默认为False
  1023. alpha: CI alpha, alpha/2 %
  1024. method: 单因素回归分析还是多因素回归分析,uni 单因素回归,multi 多因素回归,默认单因素
  1025. Returns:
  1026. """
  1027. if features is None:
  1028. features = [c for c in data.columns if c not in ['ID', label_column, 'group']]
  1029. if isinstance(features, str):
  1030. features = [features]
  1031. else:
  1032. features = list(features)
  1033. features = [c for c in features if c != label_column]
  1034. assert all(c in data.columns for c in list(features) + [label_column]), \
  1035. "请检查analysis_features或者label_column参数配置,并不是所有的特征都在数据中。"
  1036. if isinstance(need_norm, bool):
  1037. need_norm = [need_norm] * len(features)
  1038. else:
  1039. assert len(need_norm) == len(features), "标准化参数(need_norm)与特征参数(features)长度必须相同!"
  1040. if need_norm:
  1041. not_norm = [label_column] + [feat for nn, feat in zip(need_norm, features) if not nn]
  1042. data = normalize_df(data.copy()[features + [label_column]], not_norm=not_norm)
  1043. else:
  1044. data = data.copy()[features + [label_column]]
  1045. Y = data[label_column]
  1046. features_un = []
  1047. if method == 'uni':
  1048. for c in features:
  1049. model = sm.formula.ols(f'Y~{c}', data=data).fit()
  1050. print_model = model.summary(alpha=alpha)
  1051. column_names = print_model.tables[1].data[0]
  1052. column_names[0] = 'feature_name'
  1053. features_un.append(pd.DataFrame(print_model.tables[1].data[2:], columns=column_names))
  1054. elif method == 'multi':
  1055. model = sm.formula.ols(f'Y~{"+".join(features)}', data=data).fit()
  1056. print_model = model.summary(alpha=alpha)
  1057. column_names = print_model.tables[1].data[0]
  1058. column_names[0] = 'feature_name'
  1059. features_un.append(pd.DataFrame(print_model.tables[1].data[2:], columns=column_names))
  1060. features_un = pd.concat(features_un, axis=0)
  1061. features_un.index = features_un['feature_name']
  1062. for c in features_un.columns[1:]:
  1063. features_un[c] = features_un[c].astype(np.float)
  1064. features_un = features_un.drop(['feature_name', 'std err', 't'], axis=1)
  1065. features_un.columns = ['Log(HR)', 'p_value', f'lower {int(100 * (1 - alpha / 2))}%CI',
  1066. f'upper {int(100 * (1 - alpha / 2))}%CI']
  1067. features_un['HR'] = features_un['Log(HR)'].map(lambda x: math.exp(x))
  1068. return features_un[['HR', 'Log(HR)', 'p_value', f'lower {int(100 * (1 - alpha / 2))}%CI',
  1069. f'upper {int(100 * (1 - alpha / 2))}%CI']]
  1070. def plot_HR(data, columns=None, hazard_ratios=False, ax=None, coef='Log(HR)', alpha=0.1, figsize=None,
  1071. **errorbar_kwargs):
  1072. """
  1073. Produces a visual representation of the coefficients (i.e. log hazard ratios), including
  1074. their standard errors and magnitudes.
  1075. Args:
  1076. data:
  1077. columns: list, optional, specify a subset of the columns to plot
  1078. hazard_ratios: bool, optional. ill present the log-hazard ratios (the coefficients).
  1079. However, by turning this flag to True, the hazard ratios are presented instead.
  1080. ax: matplotlib axis, the matplotlib axis that be edited.
  1081. coef: 系数的列明
  1082. alpha: CI alpha, alpha/2 %
  1083. figsize: 图的尺寸
  1084. **errorbar_kwargs:
  1085. Returns:
  1086. ax: matplotlib axis
  1087. the matplotlib axis that be edited.
  1088. """
  1089. alpha = alpha / 2
  1090. errorbar_kwargs.setdefault("c", "k")
  1091. errorbar_kwargs.setdefault("fmt", "s")
  1092. errorbar_kwargs.setdefault("markerfacecolor", "white")
  1093. errorbar_kwargs.setdefault("markeredgewidth", 1.25)
  1094. errorbar_kwargs.setdefault("elinewidth", 1.25)
  1095. errorbar_kwargs.setdefault("capsize", 3)
  1096. if columns is None:
  1097. columns = data.index
  1098. else:
  1099. data = data.loc[columns]
  1100. data.sort_values(by=coef, inplace=True)
  1101. yaxis_locations = list(range(len(columns)))
  1102. log_hazards = data['Log(HR)'].values.copy()
  1103. if ax is None:
  1104. plt.figure(figsize=figsize or (10, 0.5 * len(columns) + 1))
  1105. ax = plt.gca()
  1106. if hazard_ratios:
  1107. exp_log_hazards = np.exp(log_hazards)
  1108. upper = np.exp(data[f"upper {int(100 * (1 - alpha))}%CI"].values)
  1109. lower = np.exp(data[f"lower {int(100 * (1 - alpha))}%CI"].values)
  1110. ax.errorbar(
  1111. exp_log_hazards,
  1112. yaxis_locations,
  1113. xerr=np.vstack([exp_log_hazards - lower, upper - exp_log_hazards]),
  1114. **errorbar_kwargs,
  1115. )
  1116. ax.set_xlabel("HR (%g%% CI)" % ((1 - alpha) * 100))
  1117. else:
  1118. exp_log_hazards = log_hazards
  1119. xerr = (data[f"upper {int(100 * (1 - alpha))}%CI"].values - data[
  1120. f"lower {int(100 * (1 - alpha))}%CI"].values) / 2
  1121. ax.errorbar(
  1122. exp_log_hazards,
  1123. yaxis_locations,
  1124. xerr=xerr,
  1125. **errorbar_kwargs,
  1126. )
  1127. ax.set_xlabel("log(HR) (%g%% CI)" % ((1 - alpha) * 100))
  1128. best_ylim = ax.get_ylim()
  1129. ax.vlines(1 if hazard_ratios else 0, -2, len(columns) + 1, linestyles="dashed", linewidths=1, alpha=0.65, color="k")
  1130. ax.set_ylim(best_ylim)
  1131. tick_labels = data.index
  1132. ax.set_yticks(yaxis_locations)
  1133. ax.set_yticklabels(tick_labels)
  1134. return ax
  1135. def uni_multi_variable_analysis(data: pd.DataFrame, features: Union[str, List[str]] = None, label_column: str = 'label',
  1136. need_norm: Union[bool, List[bool]] = False, alpha=0.1,
  1137. p_value4multi: float = 0.05, save_dir: Union[str] = None, prefix: str = '',
  1138. **kwargs):
  1139. """
  1140. 单因素,步进多因素分析,使用p_value4multi参数指定多因素分析的阈值
  1141. Args:
  1142. data: 数据
  1143. features: 需要分析的特征,默认除了ID和label_column列,其他的特征都进行分析。
  1144. label_column: 目标列
  1145. need_norm: 是否标准化所有分析的数据, 默认为False
  1146. alpha: CI alpha, alpha/2 %;默认为0.1即95% CI
  1147. p_value4multi: 参数指定多因素分析的阈值,默认为0.05
  1148. save_dir: 保存位置
  1149. prefix: 前缀
  1150. **kwargs:
  1151. Returns:
  1152. """
  1153. if features is None:
  1154. features = [c for c in data.columns if c not in ['ID', label_column, 'group']]
  1155. if isinstance(features, str):
  1156. features = [features]
  1157. assert all(c in data.columns for c in features + [label_column]), \
  1158. "请检查analysis_features或者label_column参数配置,并不是所有的特征都在数据中。"
  1159. if isinstance(need_norm, bool):
  1160. need_norm = [need_norm] * len(features)
  1161. else:
  1162. assert len(need_norm) == len(features), "标准化参数(need_norm)与特征参数(features)长度必须相同!"
  1163. uni = variable_analysis(data, features, label_column, need_norm, alpha)
  1164. uni = uni.round(decimals=get_param_in_cwd('display.precision', 3))
  1165. display(uni)
  1166. plot_HR(uni, **kwargs)
  1167. if save_dir is not None:
  1168. save_dir = create_dir_if_not_exists(save_dir)
  1169. plt.savefig(os.path.join(save_dir, 'univariable_reg.svg'), bbox_inches='tight')
  1170. uni.to_csv(os.path.join(save_dir, f'{prefix}univariable_reg.csv'))
  1171. plt.show()
  1172. multi_features = list(uni[uni['p_value'] < p_value4multi].index)
  1173. if multi_features:
  1174. need_norm = [nn for nn, feat in zip(need_norm, features) if feat in multi_features]
  1175. multi = variable_analysis(data, multi_features, label_column, need_norm, alpha, method='multi')
  1176. multi = multi.round(decimals=get_param_in_cwd('display.precision', 3))
  1177. display(multi)
  1178. plot_HR(multi, **kwargs)
  1179. if save_dir is not None:
  1180. save_dir = create_dir_if_not_exists(save_dir)
  1181. plt.savefig(os.path.join(save_dir, 'multivariable_reg.svg'), bbox_inches='tight')
  1182. multi.to_csv(os.path.join(save_dir, f'{prefix}multivariable_reg.csv'))
  1183. plt.show()
  1184. else:
  1185. logger.warning(f'没有通过单因素分析,找到任何pvalue<{p_value4multi}的特征!')
  1186. init_CN()

comp1.py at commit e0fb8bf, no license · at the source

Overview

Authors: Yuhang Wang1, Dandan Dong2, Shengming Shi1, Yupeng Wu1, Apekshya Singh1, Jiayi Xie1, Qiuyang Chen1, Jianwei Zhu1, Xiaofu Li1
ORCID iDs: Xiaofu Li
  1. Department of Magnetic Resonance Imaging Diagnostic, The Second Affiliated Hospital of Harbin Medical University, Harbin, China
  2. Medical Imaging Center, Zhuhai People’s Hospital (The Affiliated Hospital of Beijing Institute of Technology, Zhuhai Clinical Medical College of Jinan University), Zhuhai, China
Journal: Insights into imaging, volume 17, issue 1, article 205
Dates: received 10 November 2025; accepted 5 July 2026; published online 22 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s13244-026-02365-7 · PMID 42632855 · PMCID PMC13499786 · OpenAlex W7203986908
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: structural MRI / diffusion (modality), human (organism), other condition (population)
Methods: Connectivity, Statistics, Machine learning, fMRI & imaging
Keywords: Rectal neoplasms, Perineural invasion, Magnetic resonance imaging, Deep learning, Radiomics
Topic: Colorectal Cancer Surgical Treatments (Oncology, Medicine), according to OpenAlex
Citations: not cited yet (Europe PMC); 39 references in the paper

Abstract

Objectives: Preoperative prediction of perineural invasion (PNI) in rectal cancer (RC) is challenging due to limited MRI resolution and the neglect of extramural microenvironmental features. We developed a noninvasive framework integrating super-resolution MRI, 2.5D deep learning (DL), and intratumoral-peritumoral radiomics for preoperative PNI prediction.

Materials and methods: A dual-center cohort of 312 RC patients with pathologically confirmed PNI status was analyzed. Preoperative MRI underwent 4× super-resolution reconstruction using a generative adversarial network (GAN) to enhance tissue definition. Radiomic features were extracted from the tumor and peritumoral regions (1–5 mm). For the 2.5D DL model, the largest tumor cross-section and adjacent axial slices served as multichannel inputs to a ResNet101 backbone using transfer learning. Slice-level features were aggregated to patient-level predictions via multi-instance learning (MIL). Feature selection employed univariate analysis, Pearson correlation, mRMR, and LASSO regression. Model performance was validated via 5-fold cross-validation and an external testing cohort.

Results: The combined model, integrating MIL features, intratumoral radiomics, the optimal (2-mm) peritumoral radiomics, and clinical variables, achieved an area under the curve (AUC) of 0.912 (95% CI: 0.850–0.974) on internal validation and 0.868 (95% CI: 0.783–0.952) on external testing. Gradient-weighted class activation mapping (Grad-CAM) highlighted the tumor-neural interface, and tumor length was an independent predictor of PNI (OR = 1.062, p = 0.032).

Conclusion: We propose a hybrid radiomic-DL framework leveraging super-resolution MRI and 2.5D spatial context to enhance preoperative prediction of PNI in RC. The model shows potential for risk stratification to support personalized neoadjuvant therapy decisions, with external testing suggesting promising generalizability.

Key Points: Question How can super-resolution MRI and 2.5D deep learning address the limitations of conventional imaging in predicting preoperative rectal cancer perineural invasion?

Findings The integrated framework achieved an external testing AUC of 0.868, demonstrating superior performance compared to standalone radiomics and deep learning models.

Critical relevance statement This study demonstrates a noninvasive framework integrating super-resolution MRI and 2.5D deep learning with intratumoral-peritumoral radiomics, offering a novel approach for preoperative rectal cancer perineural invasion prediction and potential assistance in personalized neoadjuvant therapy decisions.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

OnekeyAI-Platform/onekey

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: e0fb8bf8153df4951b6dc2a144df00371f00bcbe, 15 September 2025
Languages: Python (60)
Size: 152 files, 60 scripts
Software Heritage: not archived
Found in: the text, “Image collection and pretreatment”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: PyTorch (34 files), NumPy (14 files), pandas (8 files), SciPy (6 files), Pillow (5 files), NiBabel (3 files), scikit-learn (2 files), statsmodels (2 files), LightGBM (1 file), Matplotlib (1 file), MONAI (1 file), PyRadiomics (1 file), rpy2 (1 file), seaborn (1 file), survival (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
62 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 60 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The datasets generated and analyzed during the current study are not publicly available due to the PACS system regulations by the Second Affiliated Hospital of Harbin Medical University and Zhuhai People’s Hospital but are available from the corresponding author upon reasonable request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 5 keywords, 39 references.

Cite

This paper

Wang, Y., Dong, D., Shi, S., Wu, Y., Singh, A., Xie, J., Chen, Q., Zhu, J., & Li, X. (2026). Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion. Insights into imaging, 17(1), 205. https://doi.org/10.1186/s13244-026-02365-7

BibTeX

@article{wang2026super,
author = {Wang, Yuhang and Dong, Dandan and Shi, Shengming and Wu, Yupeng and Singh, Apekshya and Xie, Jiayi and Chen, Qiuyang and Zhu, Jianwei and Li, Xiaofu},
title = {{Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion}},
journal = {Insights into imaging},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {205},
publisher = {Springer},
issn = {1869-4101},
doi = {10.1186/s13244-026-02365-7},
url = {https://doi.org/10.1186/s13244-026-02365-7},
pmid = {42632855},
pmcid = {PMC13499786}
}

RIS

TY - JOUR
AU - Wang, Yuhang
AU - Dong, Dandan
AU - Shi, Shengming
AU - Wu, Yupeng
AU - Singh, Apekshya
AU - Xie, Jiayi
AU - Chen, Qiuyang
AU - Zhu, Jianwei
AU - Li, Xiaofu
TI - Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion
T2 - Insights into imaging
J2 - Insights Imaging
PY - 2026
DA - 2026/08/22
VL - 17
IS - 1
SP - 205
SN - 1869-4101
PB - Springer
DO - 10.1186/s13244-026-02365-7
UR - https://doi.org/10.1186/s13244-026-02365-7
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s13244-026-02365-7",
"type": "article-journal",
"title": "Super-resolution MRI and 2.5D deep learning for intratumoral-peritumoral radiomics in preoperative prediction of rectal cancer perineural invasion",
"container-title": "Insights into imaging",
"author": [
{
"family": "Wang",
"given": "Yuhang"
},
{
"family": "Dong",
"given": "Dandan"
},
{
"family": "Shi",
"given": "Shengming"
},
{
"family": "Wu",
"given": "Yupeng"
},
{
"family": "Singh",
"given": "Apekshya"
},
{
"family": "Xie",
"given": "Jiayi"
},
{
"family": "Chen",
"given": "Qiuyang"
},
{
"family": "Zhu",
"given": "Jianwei"
},
{
"family": "Li",
"given": "Xiaofu"
}
],
"container-title-short": "Insights Imaging",
"volume": "17",
"issue": "1",
"page": "205",
"DOI": "10.1186/s13244-026-02365-7",
"PMID": "42632855",
"PMCID": "PMC13499786",
"ISSN": "1869-4101",
"publisher": "Springer",
"URL": "https://doi.org/10.1186/s13244-026-02365-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
22
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pcbi.1014555 [code]
Body surface potential driven personalisation of electrophysiological digital twins in hypertrophic cardiomyopathy.
Journal: PLoS computational biology
In common: PyRadiomics, MONAI, XGBoost, 9 other tools, structural MRI / diffusion
[2] doi:10.1038/s41467-026-73996-z [code]
Genetic architecture of white matter microstructure captured by unsupervised deep representation learning of fractional anisotropy maps.
Journal: Nature communications
In common: LightGBM, MONAI, NiBabel, 8 other tools, structural MRI / diffusion
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: survival, XGBoost, Pillow, 8 other tools
[4] doi:10.1162/imag.a.1164 [code]
Bias and generalizability of brain age prediction models: A multi-cohort evaluation with anatomical and interpretability insights.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: MONAI, Pillow, NiBabel, 8 other tools, structural MRI / diffusion
[5] doi:10.1016/j.nicl.2026.104012 [code]
Structural-functional multilayer brain network properties and outcome of combined repetitive transcranial magnetic stimulation and psychotherapy for obsessive-compulsive disorder.
Journal: NeuroImage. Clinical
In common: survival, XGBoost, NiBabel, 7 other tools, structural MRI / diffusion, other condition
[6] doi:10.1038/s41598-026-56688-y [code]
On the value of radiomics in addition to clinical measures in emotional conflict fMRI for predicting sertraline response in major depressive disorder.
Journal: Scientific reports
In common: PyRadiomics, XGBoost, NiBabel, 7 other tools
[7] doi:10.3389/fonc.2026.1816015 [code]
Modality-level attribution and redundancy-aware radiomics for MRI-based differentiation of melanoma and NSCLC brain metastases.
Journal: Frontiers in oncology
In common: PyRadiomics, LightGBM, NiBabel, 6 other tools, structural MRI / diffusion, other condition
[8] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: MONAI, XGBoost, NiBabel, 7 other tools, structural MRI / diffusion
[9] doi:10.1073/pnas.2516601123 [code]
Unveiling the glymphatic system's role in brain aging: A comprehensive biomarker and modifiable intervention target.
Journal: Proceedings of the National Academy of Sciences of the United States of America
In common: LightGBM, survival, XGBoost, 6 other tools, structural MRI / diffusion
[10] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: LightGBM, rpy2, statsmodels, 7 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.