OSCR

Developing and Validating a Machine Learning Model to Predict Brain Injury in Preterm Infants Using Multisource Data from the Early Postnatal Period

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § 2. Materials and Methods › 2.3. Candidate Features and Data Collection ↔ app2.py, lines 385–401 · score 0.89 · MgSO4, birth weight, gestational age, excess, hemoglobin, invasive
  2. [2] § 2. Materials and Methods › 2.10. Web Tool Construction ↔ app2.py, lines 27–82 · score 0.51 · clinical diagnosis, predicted probability, force, waterfall, prototype, thresholds

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 858 lines · 36 KB · no license · 2 matches

  1. # app2.py — PBI 风险预测(与旧版交互一致;对接新模型 14 特征 & 二元口径)
  2. # ------------------------------------------------------------------
  3. # 依赖(requirements.txt 建议):
  4. # streamlit==1.38.0
  5. # pandas==2.2.3
  6. # numpy==1.26.4
  7. # scikit-learn==1.3.2
  8. # joblib==1.3.2
  9. # skops==0.13.0
  10. import json
  11. from pathlib import Path
  12. import numpy as np
  13. import pandas as pd
  14. import json, io, os, zipfile
  15. import matplotlib.pyplot as plt
  16. import streamlit as st # ← 新增:修复 NameError
  17. import joblib # ← 新增:回退用 joblib.load 时需要
  18. from sklearn.pipeline import Pipeline
  19. from sklearn.calibration import CalibratedClassifierCV
  20. BINARY_FEATURES = {"antenatal_mgso4","surgery","NRDS","AOP","sex_male"}
  21. # =========================
  22. # 1) 多语言文本
  23. # =========================
  24. TEXT = {
  25. "lang_label": {"zh": "界面语言", "en": "Interface language"},
  26. "title": {"zh": "PBI 风险预测 · 原型(研究性)", "en": "PBI Risk Prediction · Prototype"},
  27. "caption": {
  28. "zh": "说明:本工具仅用于科研/质控,不构成临床诊断依据。",
  29. "en": "Note: Research/QC only. Not for clinical diagnosis.",
  30. },
  31. "band_title": {"zh": "阈值与分级(四档)", "en": "Thresholds & four-level risk bands"},
  32. "view_imgs": {"zh": "查看单独图片", "en": "View individual images"},
  33. "download_zip_all": {
  34. "zh": "下载导出(包含分级、waterfall/force、贡献CSV)",
  35. "en": "Download (bands + waterfall/force + contributions CSV)"
  36. },
  37. "meta": {"zh": "版本/来源信息", "en": "Release / Meta"},
  38. "input_section": {"zh": "输入特征", "en": "Input features"},
  39. "thresh_mode": {"zh": "阈值策略", "en": "Threshold mode"},
  40. "predict": {"zh": "预测", "en": "Predict"},
  41. "youd": {"zh": "Youden", "en": "Youden"},
  42. "hsens": {"zh": "高敏感", "en": "High sensitivity"},
  43. "probability": {"zh": "预测概率", "en": "Predicted probability"},
  44. "two_rules": {"zh": "两套阈值判定:", "en": "Decisions under two thresholds:"},
  45. "pos": {"zh": "阳性", "en": "Positive"},
  46. "neg": {"zh": "阴性", "en": "Negative"},
  47. "current_mode": {"zh": "当前策略", "en": "Current mode"},
  48. "result": {"zh": "结果", "en": "Result"},
  49. "download_html": {"zh": "下载 HTML 报告", "en": "Download HTML report"},
  50. "batch_title": {"zh": "批量 CSV 预测", "en": "Batch CSV prediction"},
  51. "batch_caption": {
  52. "zh": "上传包含同名列的 CSV(列顺序将自动对齐;缺失值由模型内的 SimpleImputer 处理)",
  53. "en": "Upload CSV with matching columns (reordered automatically; missing values imputed by model).",
  54. },
  55. "upload_csv": {"zh": "选择 CSV 文件", "en": "Select CSV file"},
  56. "download_csv": {"zh": "下载结果 CSV", "en": "Download result CSV"},
  57. "range_help": {"zh": "范围", "en": "Range"},
  58. "step_help": {"zh": "步进", "en": "Step"},
  59. "binary_help": {"zh": "二元变量:1=发生/是,0=未发生/否;其中 sex_male:1=男性,0=女性", "en": "Binary: 1=present/yes, 0=absent/no; sex_male: 1=male, 0=female"},
  60. "risk_low": {"zh": "低", "en": "low"},
  61. "risk_mid": {"zh": "中", "en": "medium"},
  62. "risk_high": {"zh": "高", "en": "high"},
  63. "risk_sentence": {
  64. "zh": "该患儿脑损伤的风险较{level}(概率 {p:.1%})。",
  65. "en": "The infant's risk of brain injury is {level} (probability {p:.1%}).",
  66. },
  67. "skops_warn": {
  68. "zh": "安全提示:已根据 skops>=0.10 的要求使用 get_untrusted_types(file=...) 获取受信类型列表进行加载。",
  69. "en": "Security note: Using get_untrusted_types(file=...) per skops>=0.10 to build trusted list.",
  70. },
  71. "no_model": {
  72. "zh": "未找到模型:请将 final_pipeline.skops 放在仓库根目录(或提供 final_pipeline.joblib 作为回退)。",
  73. "en": "Model not found: place final_pipeline.skops at repo root (or provide final_pipeline.joblib as fallback).",
  74. },
  75. "done_n": {"zh": "预测完成:{n} 条", "en": "Done: {n} rows"},
  76. }
  77. # =============== 2) 语言选择 ===============
  78. st.set_page_config(page_title="PBI Risk Prediction · Prototype", layout="centered")
  79. lang_choice = st.sidebar.radio(
  80. f'{TEXT["lang_label"]["zh"]} / {TEXT["lang_label"]["en"]}',
  81. ["中文", "English"],
  82. index=0,
  83. horizontal=True,
  84. )
  85. LANG = "zh" if lang_choice == "中文" else "en"
  86. st.title(TEXT["title"][LANG])
  87. st.caption(TEXT["caption"][LANG])
  88. # =============== 3) 常量:新模型 14 特征与二元口径 ===============
  89. # 顺序与 step5 的 final_features.json 一致
  90. DEFAULT_FEATURE_ORDER = [
  91. "PLT","LAC","GA_weeks_decimal","inv_vent_days","ALB","birth_weight_g",
  92. "WBC","Hb","BE","antenatal_mgso4","surgery","NRDS","AOP","sex_male"
  93. ]
  94. # —— 风险分级:用两条阈值派生四档(极低/低/中/高)——
  95. def risk_band_from_prob(p: float, hs_thr: float, y_thr: float) -> str:
  96. """基于高敏阈值与 Youden 阈值给出四分级。
  97. 规则:p >= y_thr → 高;hs_thr ≤ p < y_thr → 中;0.5*hs_thr ≤ p < hs_thr → 低;否则 极低
  98. """
  99. if p >= y_thr:
  100. return "高"
  101. elif p >= hs_thr:
  102. return "中"
  103. elif p >= 0.5 * hs_thr:
  104. return "低"
  105. else:
  106. return "极低"
  107. # —— 计算 LightGBM 贡献:使用 pipeline 的“前处理 + LGBMClassifier”——
  108. def compute_lgbm_contrib(pipe, x_df):
  109. """
  110. 返回 base(log-odds) 与逐特征贡献(list[dict])。
  111. 兼容:
  112. - 任意多层 Pipeline
  113. - CalibratedClassifierCV(不同版本字段名:estimator / classifier / base_estimator)
  114. 只使用 getattr(..., None) 安全取属性,避免 AttributeError。
  115. """
  116. def _find_inner_estimator(m):
  117. # Pipeline → 递归到最后一步
  118. if isinstance(m, Pipeline):
  119. return _find_inner_estimator(m.steps[-1][1])
  120. # Calibrated → 遍历已拟合的 calibrated_classifiers_ 列表
  121. if isinstance(m, CalibratedClassifierCV):
  122. ccs = getattr(m, "calibrated_classifiers_", None)
  123. if ccs:
  124. # 逐个尝试拿内部估计器
  125. for cc in ccs:
  126. for key in ("estimator", "classifier", "base_estimator"):
  127. est = getattr(cc, key, None)
  128. if est is not None:
  129. return _find_inner_estimator(est)
  130. # 有些版本还会在 CalibratedClassifierCV 自身挂一个 estimator/base_estimator
  131. for key in ("estimator", "base_estimator", "classifier"):
  132. est = getattr(m, key, None)
  133. if est is not None:
  134. return _find_inner_estimator(est)
  135. raise RuntimeError("CalibratedClassifierCV 未暴露内部估计器(estimator/classifier/base_estimator 均不存在)。")
  136. # 若对象本身就是模型(如 LGBMClassifier),直接返回
  137. return m
  138. try:
  139. # 拆前处理与最后一步
  140. if hasattr(pipe, "steps") and len(pipe.steps) > 1:
  141. pre = Pipeline(steps=pipe.steps[:-1])
  142. last = pipe.steps[-1][1]
  143. else:
  144. pre, last = None, pipe
  145. Xp = pre.transform(x_df) if pre is not None else x_df
  146. # 剥离到底层 LGBMClassifier
  147. raw_est = _find_inner_estimator(last)
  148. # 先尝试 raw_est.predict(..., pred_contrib=True)
  149. contrib = None
  150. if hasattr(raw_est, "predict"):
  151. try:
  152. contrib = raw_est.predict(Xp, pred_contrib=True)
  153. except TypeError:
  154. contrib = None
  155. # 再尝试 booster_.predict
  156. if contrib is None and hasattr(raw_est, "booster_"):
  157. try:
  158. contrib = raw_est.booster_.predict(Xp, pred_contrib=True)
  159. except Exception:
  160. contrib = None
  161. if contrib is None:
  162. raise RuntimeError(
  163. "底层估计器不支持 LightGBM 的 pred_contrib;"
  164. f"实际类型:{type(raw_est)},可用属性:{dir(raw_est)}"
  165. )
  166. contrib = np.asarray(contrib)
  167. vec = contrib[0] if contrib.ndim == 2 else contrib
  168. base = float(vec[-1]) # 最后一列是 bias/base value
  169. vals = vec[:-1]
  170. items = []
  171. for name, val in zip(x_df.columns.tolist(), vals):
  172. items.append({
  173. "feature": name,
  174. "value": float(x_df.iloc[0][name]),
  175. "contribution": float(val)
  176. })
  177. return base, items
  178. except Exception as e:
  179. raise RuntimeError(f"无法计算特征贡献(pred_contrib):{e}")
  180. # —— 升级版:更贴近论文中 SHAP 报告风格(顶部一句话 + Waterfall + Force)——
  181. import matplotlib.pyplot as plt
  182. from matplotlib.patches import Rectangle
  183. def _format_risk_text(prob, band, lang="zh"):
  184. band_map_en = {"极低":"very low", "低":"low", "中":"medium", "高":"high"}
  185. if lang == "en":
  186. return f"This patient has {band_map_en.get(band, band)} risk of PBI with the probability of {prob:.4f}."
  187. else:
  188. return f"该患儿脑损伤风险为「{band}」;预测概率 {prob:.4f}"
  189. def save_waterfall_tif(base, items, out_path: str, top_k: int = 14, lang="zh"):
  190. """
  191. 论文风格:横向逐步累计 Waterfall
  192. - y 轴是按 |贡献| 排序后的特征(从上到下)
  193. - 每一行是一个“水平浮动条”,从累计起点 left 到 left+width
  194. - 左侧标 E[f(X)],右侧标 f(x)
  195. """
  196. # 1) 取前 top_k 项(按贡献绝对值大到小)
  197. items_sorted = sorted(items, key=lambda d: abs(d["contribution"]), reverse=True)[:top_k]
  198. feats = [f'{it["feature"]}={it["value"]:.2f}' for it in items_sorted]
  199. deltas = [float(it["contribution"]) for it in items_sorted]
  200. # 2) 逐步累计(横向)
  201. starts = []
  202. cur = float(base)
  203. for d in deltas:
  204. starts.append(cur if d >= 0 else cur + d)
  205. cur += d
  206. fx = float(cur)
  207. # 3) 画图
  208. n = len(items_sorted)
  209. y = np.arange(n)[::-1] # 让“最大贡献”在最上
  210. colors = ["#d62728" if d >= 0 else "#1f77b4" for d in deltas]
  211. fig, ax = plt.subplots(figsize=(10, 5.4), dpi=300)
  212. ax.barh(y, width=[abs(d) for d in deltas], left=starts, height=0.62, color=colors, edgecolor="none")
  213. # 4) 每段标注 +0.75/-0.22
  214. for yi, left, d in zip(y, starts, deltas):
  215. x_pos = left + (abs(d) * (1 if d >= 0 else -1)) # 段末尾位置
  216. ax.text(x_pos, yi, f"{d:+.2f}", va="center",
  217. ha="left" if d >= 0 else "right", fontsize=8, color="black")
  218. # 5) 坐标/标签/标题
  219. ax.set_yticks(y)
  220. ax.set_yticklabels([feats[i] for i in range(n)][::-1], fontsize=8)
  221. ax.set_xlabel("f(x) (log-odds)")
  222. ax.set_title("Waterfall plot of this patient")
  223. # 6) 范围与基线/终点标注
  224. xmin = min(base, fx, *(s if d >= 0 else s for s, d in zip(starts, deltas))) - 0.5
  225. xmax = max(base, fx, *(s + abs(d) for s, d in zip(starts, deltas))) + 0.5
  226. ax.set_xlim(xmin, xmax)
  227. ax.axvline(0, color="k", lw=0.6)
  228. ax.text(base, n + 0.2, f"E[f(X)] = {base:.3f}", ha="center", va="bottom", fontsize=9)
  229. ax.text(fx, n + 0.2, f"f(x) = {fx:.3f}", ha="center", va="bottom", fontsize=9)
  230. fig.tight_layout()
  231. fig.savefig(out_path, dpi=300, bbox_inches="tight")
  232. plt.close(fig)
  233. def save_force_tif(base, items, out_path: str, top_k: int = 14, lang="zh"):
  234. """
  235. 论文风格:单轴 Force-like 图
  236. - 在一条水平轴上,从 base 向右(红)/向左(蓝)逐段累计到 f(x)
  237. - 顶部显示 higher / lower 标签;底部显示 base value 与 f(x)
  238. """
  239. items_sorted = sorted(items, key=lambda d: abs(d["contribution"]), reverse=True)[:top_k]
  240. # 累计区段
  241. segs = []
  242. cur = float(base)
  243. for it in items_sorted:
  244. c = float(it["contribution"])
  245. left = cur if c >= 0 else cur + c
  246. right = cur + c if c >= 0 else cur
  247. segs.append((left, right, it))
  248. cur += c
  249. fx = float(cur)
  250. # 计算显示范围
  251. xs = [base, fx] + [p for lr in segs for p in lr[:2]]
  252. xmin, xmax = min(xs) - 0.5, max(xs) + 0.5
  253. # 作图
  254. fig, ax = plt.subplots(figsize=(10, 3.6), dpi=300)
  255. ax.set_ylim(0, 1)
  256. ax.set_yticks([])
  257. # 主轴
  258. ax.hlines(0.5, xmin, xmax, colors="#888", lw=1)
  259. ax.axvline(0, color="k", lw=0.6)
  260. # 区段 + 标签
  261. for left, right, it in segs:
  262. color = "#d62728" if right >= left else "#1f77b4"
  263. ax.add_patch(Rectangle((min(left, right), 0.36), abs(right-left), 0.28,
  264. facecolor=color, edgecolor="none", alpha=0.95))
  265. ax.text((left+right)/2, 0.50,
  266. f'{it["feature"]}={it["value"]:.2f}',
  267. ha="center", va="center", fontsize=8, color="white")
  268. # 标注 base 与 f(x)
  269. ax.text(base, 0.86, f"base value = {base:.3f}", ha="center", va="bottom", fontsize=9)
  270. ax.text(fx, 0.86, f"f(x) = {fx:.3f}", ha="center", va="bottom", fontsize=9)
  271. # higher / lower
  272. ax.text(xmax, 0.95, "higher", color="#d62728", ha="right", va="top", fontsize=9)
  273. ax.text(xmin, 0.95, "lower", color="#1f77b4", ha="left", va="top", fontsize=9)
  274. ax.set_xlim(xmin, xmax)
  275. ax.set_xlabel("contribution to f(x) (log-odds)")
  276. ax.set_title("The force plot of this patient")
  277. fig.tight_layout()
  278. fig.savefig(out_path, dpi=300, bbox_inches="tight")
  279. plt.close(fig)
  280. def save_shap_report_tif(base, items, prob, band_text, out_path: str, lang="zh", top_k: int = 14):
  281. """
  282. 一页式报告:顶部“这名患儿…概率 0.xxxx”,下方 Waterfall + Force
  283. - 将“高/中/低/极低”与概率高亮红色,模仿论文排版
  284. """
  285. tmp_wf = "_tmp_wf.tif"
  286. tmp_fc = "_tmp_fc.tif"
  287. save_waterfall_tif(base, items, tmp_wf, top_k=top_k, lang=lang)
  288. save_force_tif(base, items, tmp_fc, top_k=top_k, lang=lang)
  289. wf = plt.imread(tmp_wf)
  290. fc = plt.imread(tmp_fc)
  291. fig, ax = plt.subplots(figsize=(10, 11.2), dpi=300)
  292. ax.axis("off")
  293. # —— 顶部大标题(分两次写字以高亮“风险等级”和“概率”)——
  294. if lang == "en":
  295. ax.text(0.02, 0.975, "This patient has ", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  296. ax.text(0.31, 0.975, f"{band_text} risk", color="crimson", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  297. ax.text(0.47, 0.975, " of PBI with the probability of ", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  298. ax.text(0.86, 0.975, f"{prob:.4f}", color="crimson", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  299. ax.text(0.02, 0.94,
  300. "Waterfall plots show how each feature moves the model output from the expected value (E[f(X)]) "
  301. "to the prediction f(x). Red increases risk; blue decreases risk.", fontsize=10, color="#444",
  302. ha="left", va="top", transform=ax.transAxes)
  303. else:
  304. ax.text(0.02, 0.975, "该患儿脑损伤风险为 ", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  305. ax.text(0.30, 0.975, f"「{band_text}」", color="crimson", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  306. ax.text(0.40, 0.975, ";预测概率 ", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  307. ax.text(0.52, 0.975, f"{prob:.4f}", color="crimson", ha="left", va="top", fontsize=18, transform=ax.transAxes)
  308. ax.text(0.02, 0.94,
  309. "瀑布图展示从期望输出 E[f(X)] 到个体输出 f(x) 的逐步贡献;红色↑风险,蓝色↓风险。",
  310. fontsize=10, color="#444", ha="left", va="top", transform=ax.transAxes)
  311. # —— 拼图:上半瀑布、下半 force ——
  312. ax.imshow(wf, extent=[0, 1, 0.44, 0.92])
  313. ax.imshow(fc, extent=[0, 1, 0.03, 0.41])
  314. fig.savefig(out_path, dpi=300, bbox_inches="tight")
  315. plt.close(fig)
  316. # 清理临时图
  317. try:
  318. import os
  319. os.remove(tmp_wf); os.remove(tmp_fc)
  320. except Exception:
  321. pass
  322. # =============== 4) 路径与资产 ===============
  323. ROOT = Path(__file__).parent
  324. SKOPS_PATH = ROOT / "final_pipeline.skops" # 推荐
  325. PIPE_PATH = ROOT / "final_pipeline.joblib" # 回退
  326. SCHEMA_PATH = ROOT / "feature_schema.json" # 若无则使用默认顺序
  327. THR_PATH = ROOT / "thresholds.json"
  328. RECAL_PATH = ROOT / "external_recal.json"
  329. META_PATHS = [ROOT / "version.json", ROOT / "release_meta.json"]
  330. # 显示映射(中英)
  331. DISPLAY = {
  332. "PLT": {"zh": "血小板计数 (PLT, 10^9/L)", "en": "Platelets (PLT, 10^9/L)"},
  333. "LAC": {"zh": "乳酸 (LAC, mmol/L)", "en": "Lactate (LAC, mmol/L)"},
  334. "GA_weeks_decimal": {"zh": "胎龄(周)", "en": "Gestational age (weeks, decimal)"},
  335. "inv_vent_days": {"zh": "有创通气天数 (d)", "en": "Invasive ventilation (days)"},
  336. "ALB": {"zh": "白蛋白 (ALB, g/L)", "en": "Albumin (ALB, g/L)"},
  337. "birth_weight_g": {"zh": "出生体重 (g)", "en": "Birth weight (g)"},
  338. "WBC": {"zh": "白细胞计数 (WBC, 10^9/L)", "en": "WBC (10^9/L)"},
  339. "Hb": {"zh": "血红蛋白 (Hb, g/L)", "en": "Hemoglobin (g/L)"},
  340. "BE": {"zh": "碱剩余 (BE, mmol/L)", "en": "Base excess (mmol/L)"},
  341. "antenatal_mgso4": {"zh": "产前硫酸镁治疗 (0/1)", "en": "Antenatal MgSO4 (0/1)"},
  342. "surgery": {"zh": "手术 (0/1)", "en": "Surgery (0/1)"},
  343. "NRDS": {"zh": "新生儿呼吸窘迫综合征 (0/1)", "en": "NRDS (0/1)"},
  344. "AOP": {"zh": "早产儿贫血 (0/1)", "en": "AOP (0/1)"},
  345. "sex_male": {"zh": "性别:1=男性,0=女性", "en": "Sex: 1=male, 0=female"},
  346. }
  347. # =============== 5) 加载模型与配置(缓存) ===============
  348. @st.cache_resource
  349. def load_assets():
  350. # 模型
  351. pipe = None
  352. if SKOPS_PATH.exists():
  353. try:
  354. import skops.io as sio, inspect
  355. # 兼容 >=0.10 的关键字参数名
  356. sig = None
  357. try:
  358. sig = inspect.signature(sio.get_untrusted_types)
  359. except Exception:
  360. sig = None
  361. trusted = True
  362. if sig:
  363. params = sig.parameters
  364. if "file" in params:
  365. trusted = sio.get_untrusted_types(file=str(SKOPS_PATH))
  366. elif "path" in params:
  367. trusted = sio.get_untrusted_types(path=str(SKOPS_PATH))
  368. else:
  369. trusted = sio.get_untrusted_types(str(SKOPS_PATH))
  370. pipe = sio.load(str(SKOPS_PATH), trusted=trusted)
  371. st.info(TEXT["skops_warn"][LANG])
  372. except Exception as e:
  373. st.warning(f"读取 SKOPS 失败,将回退到 joblib:{e}")
  374. if pipe is None:
  375. if PIPE_PATH.exists():
  376. pipe = joblib.load(PIPE_PATH)
  377. else:
  378. st.error(TEXT["no_model"][LANG]); st.stop()
  379. # schema / 阈值
  380. # 若没有 schema,则使用 DEFAULT_FEATURE_ORDER 并构造简单定义
  381. order = DEFAULT_FEATURE_ORDER.copy()
  382. feat_defs = {n: {"name": n, "dtype": ("binary" if n in BINARY_FEATURES else "float"),
  383. "allowed_range": [None, None], "step": (1 if n in BINARY_FEATURES else 0.1)}
  384. for n in order}
  385. if SCHEMA_PATH.exists():
  386. try:
  387. schema = json.loads(SCHEMA_PATH.read_text(encoding="utf-8"))
  388. order = schema.get("order") or [d["name"] for d in schema.get("features", [])] or order
  389. feats = schema.get("features", [])
  390. if feats:
  391. feat_defs = {d["name"]: d for d in feats}
  392. except Exception as e:
  393. st.warning(f"读取 feature_schema.json 失败,将使用默认顺序:{e}")
  394. if not THR_PATH.exists():
  395. st.error(f"缺少 {THR_PATH.name}"); st.stop()
  396. thr = json.loads(THR_PATH.read_text(encoding="utf-8"))
  397. def pick(obj, *keys):
  398. for k in keys:
  399. if k in obj:
  400. v = obj[k]
  401. return float(v["thr"]) if isinstance(v, dict) else float(v)
  402. return None
  403. youden = pick(thr, "youden", "Youden")
  404. highs = pick(thr, "high_sensitivity", "HighSens", "highsens")
  405. # 元信息(有则显示)
  406. meta = {}
  407. for m in META_PATHS:
  408. if m.exists():
  409. try:
  410. meta = json.loads(m.read_text(encoding="utf-8"))
  411. break
  412. except Exception:
  413. pass
  414. recal = None
  415. if RECAL_PATH.exists():
  416. try:
  417. recal = json.loads(RECAL_PATH.read_text(encoding="utf-8"))
  418. except Exception as e:
  419. st.warning(f"读取 {RECAL_PATH.name} 失败:{e}")
  420. return pipe, order, feat_defs, {"youden": youden, "highs": highs}, meta, recal
  421. pipe, order, featdefs, thrs, meta, recal = load_assets()
  422. with st.expander(TEXT["meta"][LANG], expanded=False):
  423. st.json(meta or {"note": "N/A"})
  424. # =============== 6) 工具函数 ===============
  425. def is_binary(name: str, defn: dict) -> bool:
  426. if name in BINARY_FEATURES:
  427. return True
  428. dtype = str(defn.get("dtype","")).lower()
  429. if dtype == "binary":
  430. return True
  431. rng = defn.get("allowed_range") or [None, None]
  432. if rng[0] == 0 and rng[1] == 1:
  433. return True
  434. return False
  435. def label_for(name: str) -> str:
  436. d = DISPLAY.get(name, {})
  437. return d.get(LANG, name)
  438. def help_for(defn: dict, is_bin: bool) -> str:
  439. lo, hi = (defn.get("allowed_range") or [None, None])
  440. step = defn.get("step", 1 if is_bin else 0.1)
  441. parts = []
  442. if isinstance(lo, (int, float)) and isinstance(hi, (int, float)):
  443. parts.append(f'{TEXT["range_help"][LANG]}: {lo}–{hi}')
  444. parts.append(f'{TEXT["step_help"][LANG]}: {step}')
  445. if is_bin:
  446. parts.append(TEXT["binary_help"][LANG])
  447. return " · ".join(parts)
  448. # =============== 7) 表单输入 ===============
  449. st.markdown(f"### {TEXT['input_section'][LANG]}")
  450. cols = st.columns(2)
  451. values = {}
  452. for i, name in enumerate(order):
  453. d = featdefs.get(name, {})
  454. bin_flag = is_binary(name, d)
  455. lbl = label_for(name)
  456. help_text = help_for(d, bin_flag)
  457. lo, hi = (d.get("allowed_range") or [None, None])
  458. with cols[i % 2]:
  459. if bin_flag:
  460. # 0/1 选择;sex_male: 1=男性, 0=女性
  461. values[name] = st.selectbox(lbl, options=[0, 1], index=0, help=help_text, key=f"bin_{name}")
  462. else:
  463. if isinstance(lo, (int, float)) and isinstance(hi, (int, float)):
  464. default = float((lo + hi) / 2.0)
  465. step = float(d.get("step", 0.1))
  466. values[name] = st.number_input(lbl, value=default,
  467. min_value=float(lo), max_value=float(hi),
  468. step=step, format="%.3f", help=help_text)
  469. else:
  470. values[name] = st.number_input(lbl, value=0.0,
  471. step=float(d.get("step", 0.1)),
  472. format="%.3f", help=help_text)
  473. st.caption(TEXT["binary_help"][LANG])
  474. st.divider()
  475. mode = st.radio(TEXT["thresh_mode"][LANG], [TEXT["youd"][LANG], TEXT["hsens"][LANG]], horizontal=True)
  476. btn_predict = st.button(TEXT["predict"][LANG], type="primary")
  477. # =============== 8) 阈值处理 ===============
  478. def pick_thresholds(thrs_dict):
  479. t_youden = thrs_dict.get("youden")
  480. t_hsens = thrs_dict.get("highs")
  481. if t_youden is None: t_youden = 0.5
  482. if t_hsens is None: t_hsens = 0.2
  483. lo, hi = sorted([float(t_youden), float(t_hsens)])
  484. return lo, hi, float(t_youden), float(t_hsens)
  485. low_thr, high_thr, t_youden, t_hsens = pick_thresholds(thrs)
  486. def risk_bucket(p: float, lang="zh") -> str:
  487. if p < low_thr: return TEXT["risk_low"][lang]
  488. elif p < high_thr: return TEXT["risk_mid"][lang]
  489. else: return TEXT["risk_high"][lang]
  490. # =============== 9) 单例预测 + 报告下载 ===============
  491. if btn_predict:
  492. x = pd.DataFrame([values])[order] # 严格列顺序
  493. try:
  494. p = float(pipe.predict_proba(x)[:, 1][0])
  495. # 部署端再校准(可选)
  496. if recal:
  497. a = float(recal.get("intercept", 0.0))
  498. b = float(recal.get("slope", 1.0))
  499. eps = 1e-12
  500. z = np.log((p + eps) / (1 - p + eps))
  501. p = 1.0 / (1.0 + np.exp(-(a + b * z)))
  502. except Exception as e:
  503. st.error(f"预测失败:{e}")
  504. st.stop()
  505. youden_thr = float(thrs["youden"])
  506. hs_thr = float(thrs["highs"])
  507. st.subheader(TEXT["probability"][LANG])
  508. st.metric(TEXT["probability"][LANG], f"{p:.4f}")
  509. # —— 两种策略下的四档风险分级 ——
  510. band_youden = risk_band_from_prob(p, hs_thr, youden_thr)
  511. band_highs = risk_band_from_prob(p, hs_thr, youden_thr) # 分级规则一致,只是口径不同时你也可单独给阈值
  512. st.write(f"**{TEXT['band_title'][LANG]}**")
  513. c1, c2 = st.columns(2)
  514. # 小工具:把“极低/低/中/高”翻成英文
  515. def _band_disp(b):
  516. return {"极低": "very low", "低": "low", "中": "medium", "高": "high"}.get(b, b)
  517. with c1:
  518. if LANG == "en":
  519. dec_y = "Positive" if p >= youden_thr else "Negative"
  520. st.info(
  521. f"**Youden**: threshold={youden_thr:.6f} → decision: **{dec_y}**; "
  522. f"risk band: **{_band_disp(band_youden)}**"
  523. )
  524. else:
  525. dec_y = "阳性" if p >= youden_thr else "阴性"
  526. st.info(
  527. f"**Youden**:阈值={youden_thr:.6f} → 判定:**{dec_y}**;风险分级:**{band_youden}**"
  528. )
  529. with c2:
  530. if LANG == "en":
  531. dec_h = "Positive" if p >= hs_thr else "Negative"
  532. st.info(
  533. f"**{TEXT['hsens'][LANG]}**: threshold={hs_thr:.6f} → decision: **{dec_h}**; "
  534. f"risk band: **{_band_disp(band_highs)}**"
  535. )
  536. else:
  537. dec_h = "阳性" if p >= hs_thr else "阴性"
  538. st.info(
  539. f"**高敏**:阈值={hs_thr:.6f} → 判定:**{dec_h}**;风险分级:**{band_highs}**"
  540. )
  541. # —— 计算 LightGBM 贡献并绘图/导出 ——
  542. import pandas as pd, numpy as np, io, zipfile, os
  543. out_files = {}
  544. try:
  545. base, items = compute_lgbm_contrib(pipe, x)
  546. # 保存 CSV(贡献明细)
  547. df_contrib = pd.DataFrame(items) # feature, value, contribution
  548. out_files["shap_contrib.csv"] = df_contrib.to_csv(index=False).encode("utf-8")
  549. # 保存 summary.csv(概率、阈值、分级)
  550. df_summary = pd.DataFrame([{
  551. "p_hat": p,
  552. "youden_thr": youden_thr, "decision_youden": ("POS" if p >= youden_thr else "NEG"),
  553. "highs_thr": hs_thr, "decision_highs": ("POS" if p >= hs_thr else "NEG"),
  554. "band_youden": band_youden, "band_highs": band_highs,
  555. "base_logit": base
  556. }])
  557. out_files["summary.csv"] = df_summary.to_csv(index=False).encode("utf-8")
  558. # 仅生成与展示两张单图(不再生成“顶部结论”的整页报告)
  559. import matplotlib
  560. import matplotlib.pyplot as plt
  561. def _set_step5_style():
  562. try:
  563. matplotlib.rcParams["font.family"] = "Times New Roman"
  564. except Exception:
  565. pass
  566. matplotlib.rcParams["axes.unicode_minus"] = False
  567. def save_waterfall_tif(base, items, out_path: str, top_k: int = 14, lang="zh", dpi: int = 600):
  568. _set_step5_style()
  569. # 取贡献绝对值最大的 top_k 个
  570. items_sorted = sorted(items, key=lambda d: abs(d["contribution"]), reverse=True)[:top_k]
  571. deltas, labels = [], []
  572. for it in items_sorted:
  573. deltas.append(float(it["contribution"]))
  574. v = it["value"]
  575. v_txt = f"{int(v)}" if isinstance(v, (int, float)) and float(v).is_integer() else f"{float(v):.2f}"
  576. labels.append(f'{it["feature"]}={v_txt}')
  577. # 逐段累计(竖向 waterfall:每个柱子的 bottom=累计到上一段的值)
  578. starts, cur = [], float(base)
  579. for d in deltas:
  580. starts.append(cur)
  581. cur += d
  582. colors = ["#d62728" if d >= 0 else "#1f77b4" for d in deltas]
  583. fig, ax = plt.subplots(figsize=(10, 5.2), dpi=dpi)
  584. bars = ax.bar(range(len(deltas)), deltas, bottom=starts, color=colors, width=0.6, edgecolor="none")
  585. ax.axhline(0, color="#333", lw=0.6)
  586. ax.set_xticks(range(len(labels)))
  587. ax.set_xticklabels(labels, rotation=55, ha="right", fontsize=9)
  588. ax.set_ylabel("f(x) (log-odds)", fontsize=11)
  589. ax.set_title("Waterfall plot of this patient", fontsize=12)
  590. # —— 在每个条上标出具体的正负数值 —— #
  591. # 计算 y 方向边距,用于文字轻微偏移,避免与柱边缘重合
  592. tops = [s + d for s, d in zip(starts, deltas)]
  593. ymin = min([base] + starts + tops)
  594. ymax = max([base] + starts + tops)
  595. yspan = max(1e-6, ymax - ymin)
  596. pad = 0.015 * yspan
  597. for i, (s, d) in enumerate(zip(starts, deltas)):
  598. y_text = s + d + (pad if d >= 0 else -pad)
  599. ax.text(i, y_text, f"{d:+.2f}", ha="center",
  600. va=("bottom" if d >= 0 else "top"),
  601. fontsize=9, color="black")
  602. # 标注 E[f(X)]
  603. ax.text(-0.5, base, f"E[f(X)] = {base:.3f}", va="center", fontsize=9, color="#444")
  604. # y 轴留一点空隙,避免最上/最下文字被裁切
  605. ax.set_ylim(ymin - 0.05 * yspan, ymax + 0.08 * yspan)
  606. fig.tight_layout()
  607. fig.savefig(out_path, dpi=dpi, bbox_inches="tight")
  608. plt.close(fig)
  609. def save_force_tif(base, items, out_path: str, top_k: int = 14, lang="zh", dpi: int = 600):
  610. _set_step5_style()
  611. # 先按贡献从小到大排,随后做“负向前半 + 正向后半”的截取,保证两边信息对称
  612. items_sorted = sorted(items, key=lambda d: d["contribution"])
  613. if top_k and top_k < len(items_sorted):
  614. neg = [it for it in items_sorted if it["contribution"] < 0]
  615. pos = [it for it in items_sorted if it["contribution"] >= 0]
  616. k2 = max(1, top_k // 2)
  617. items_sorted = neg[:k2] + pos[-k2:]
  618. vals, labels = [], []
  619. for it in items_sorted:
  620. c = float(it["contribution"])
  621. vals.append(c)
  622. v = it["value"]
  623. v_txt = f"{int(v)}" if isinstance(v, (int, float)) and float(v).is_integer() else f"{float(v):.2f}"
  624. labels.append(f'{it["feature"]}={v_txt}')
  625. colors = ["#1f77b4" if v < 0 else "#d62728" for v in vals]
  626. h = max(2.8, 0.55 * len(vals) + 1.8) # 动态高度,避免“堆在一起”
  627. fig, ax = plt.subplots(figsize=(10, h), dpi=dpi)
  628. y = list(range(len(vals)))
  629. ax.barh(y, vals, color=colors, edgecolor="none")
  630. ax.axvline(0, color="#333", lw=0.6)
  631. ax.set_yticks(y)
  632. ax.set_yticklabels(labels, fontsize=9)
  633. ax.set_xlabel("contribution to f(x) (log-odds)", fontsize=11)
  634. ax.set_title("The force plot of this patient", fontsize=12)
  635. # —— 为每个横条标出数值;并自动留左右边距 —— #
  636. xmin = min(0.0, min(vals)) # 基于数据估一个范围
  637. xmax = max(0.0, max(vals))
  638. xspan = max(1e-6, xmax - xmin)
  639. pad = 0.02 * xspan
  640. for yi, v in enumerate(vals):
  641. # 数值贴在条形末端外侧:正数靠右、负数靠左
  642. x_text = v + (pad if v >= 0 else -pad)
  643. ax.text(x_text, yi, f"{v:+.2f}",
  644. va="center",
  645. ha=("left" if v >= 0 else "right"),
  646. fontsize=9, color="black")
  647. # 给坐标两端各加 8% 空间,确保文字不被裁切
  648. ax.set_xlim(xmin - 0.08 * xspan, xmax + 0.08 * xspan)
  649. fig.tight_layout()
  650. fig.savefig(out_path, dpi=dpi, bbox_inches="tight")
  651. plt.close(fig)
  652. # === 实际调用:生成图片 + 页面展示 + 加入导出 ===
  653. wf_path = "waterfall_patient.tif"
  654. fc_path = "force_patient.tif"
  655. save_waterfall_tif(base, items, wf_path, top_k=14, dpi=600)
  656. save_force_tif(base, items, fc_path, top_k=14, dpi=600)
  657. with st.expander(TEXT["view_imgs"][LANG], expanded=True):
  658. st.image(wf_path, caption="Waterfall", use_column_width=True)
  659. st.image(fc_path, caption="Force", use_column_width=True)
  660. from pathlib import Path
  661. out_files["waterfall.tif"] = Path(wf_path).read_bytes()
  662. out_files["force.tif"] = Path(fc_path).read_bytes()
  663. except Exception as e:
  664. st.warning(f"未能生成贡献图:{e}")
  665. # —— 打包导出 ——
  666. if out_files:
  667. zip_buf = io.BytesIO()
  668. with zipfile.ZipFile(zip_buf, "w", compression=zipfile.ZIP_DEFLATED) as zf:
  669. for nm, data in out_files.items():
  670. zf.writestr(nm, data)
  671. st.download_button(TEXT["download_zip_all"][LANG],
  672. data=zip_buf.getvalue(),
  673. file_name="pbi_single_prediction_exports.zip",
  674. mime="application/zip")
  675. y1 = int(p >= t_youden)
  676. y2 = int(p >= t_hsens)
  677. st.metric(TEXT["probability"][LANG], f"{p:.4f}")
  678. st.write(f"**{TEXT['two_rules'][LANG]}**")
  679. st.write(f"- Youden @ {t_youden:.6f} → **{TEXT['pos'][LANG] if y1 else TEXT['neg'][LANG]}**")
  680. st.write(f"- {TEXT['hsens'][LANG]} @ {t_hsens:.6f} → **{TEXT['pos'][LANG] if y2 else TEXT['neg'][LANG]}**")
  681. final_label = y1 if mode == TEXT["youd"][LANG] else y2
  682. st.success(f"{TEXT['current_mode'][LANG]}:**{mode}** → {TEXT['result'][LANG]}:**{TEXT['pos'][LANG] if final_label else TEXT['neg'][LANG]}**")
  683. risk_sent = TEXT["risk_sentence"][LANG].format(level=risk_bucket(p, LANG), p=p)
  684. st.write(risk_sent)
  685. # 简易 HTML 报告
  686. rows_html = "".join([f"<tr><td>{n}</td><td>{values[n]}</td></tr>" for n in order])
  687. html = f"""<!doctype html>
  688. <html><head><meta charset="utf-8"><title>PBI Report</title>
  689. <style>body{{font-family:Arial,Helvetica,sans-serif;max-width:900px;margin:24px auto;}}
  690. h1{{font-size:22px;}} table{{border-collapse:collapse;width:100%;}}
  691. td,th{{border:1px solid #ddd;padding:6px 8px;}} .small{{color:#666;font-size:12px;}}
  692. </style></head>
  693. <body>
  694. <h1>{TEXT['title'][LANG]}</h1>
  695. <div class="small">version: {meta.get('model_name','N/A')} · build: {meta.get('timestamp','N/A')}</div>
  696. <h2>{TEXT['input_section'][LANG]}</h2>
  697. <table><tr><th>Feature</th><th>Value</th></tr>{rows_html}</table>
  698. <h2>{TEXT['result'][LANG]}</h2>
  699. <p>{TEXT['probability'][LANG]}:<b>{p:.4f}</b></p>
  700. <ul>
  701. <li>Youden@{t_youden:.6f} → {'POS' if y1 else 'NEG'}</li>
  702. <li>{TEXT['hsens'][LANG]}@{t_hsens:.6f} → {'POS' if y2 else 'NEG'}</li>
  703. </ul>
  704. <p>{risk_sent}</p>
  705. <p class="small">{TEXT['caption'][LANG]}</p>
  706. </body></html>"""
  707. st.download_button(TEXT["download_html"][LANG], data=html.encode("utf-8"), file_name="pbi_report.html",
  708. mime="text/html")
  709. # =============== 10) 批量 CSV 预测 ===============
  710. st.divider()
  711. st.markdown(f"### {TEXT['batch_title'][LANG]}")
  712. st.caption(TEXT["batch_caption"][LANG])
  713. up = st.file_uploader(TEXT["upload_csv"][LANG], type=["csv"])
  714. if up is not None:
  715. try:
  716. df = pd.read_csv(up)
  717. # 丢失列补 NaN;多余列忽略;严格顺序
  718. for col in order:
  719. if col not in df.columns:
  720. df[col] = np.nan
  721. X = df[order]
  722. proba = pipe.predict_proba(X)[:, 1]
  723. df_out = df.copy()
  724. df_out["p_hat"] = proba
  725. df_out["label_youden"] = (df_out["p_hat"] >= t_youden).astype(int)
  726. df_out["label_highsens"] = (df_out["p_hat"] >= t_hsens).astype(int)
  727. csv_bytes = df_out.to_csv(index=False, encoding="utf-8-sig").encode("utf-8-sig")
  728. st.success(TEXT["done_n"][LANG].format(n=len(df_out)))
  729. st.download_button(TEXT["download_csv"][LANG], data=csv_bytes, file_name="pbi_pred_results.csv",
  730. mime="text/csv")
  731. except Exception as e:
  732. st.error(f"读取或预测失败:{e}")

app2.py at commit 6e3d3b3, no license · at the source

Overview

Authors: Pu Xu1,2, Ying Li2, Ying Chen2, Tongying Han2, Peicen Zou1,2, Qinglin Lu1,2, Dongmiao Zhang2, Jie Chen2, Yajuan Wang1,2
  1. Capital Institute of Pediatrics, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100020, China; (P.X.); (P.Z.); (Q.L.)
  2. Department of Neonatology, Capital Center for Children’s Health, Capital Medical University, Capital Institute of Pediatrics, Beijing 100020, China; (Y.L.); (Y.C.); (T.H.); (D.Z.); (J.C.)
Journal: Children (Basel, Switzerland), volume 13, issue 6, article 796
Dates: received 25 May 2026; accepted 8 June 2026; published online 9 June 2026; in print June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI · PMCID PMC13297451
Status: code verified
Categories: human (organism), other condition (population), developmental (subfield)
Methods: Statistics, Machine learning, Preprocessing, Connectivity
Keywords: preterm, brain injuries, neurodevelopmental impairment, machine learning, predictive model
Funding: Capital’s Funds for Health Improvement and Research (2026-2-2102, 2024-2-2102); Beijing Municipal Health Commission (Academic leader-03-02)
Citations: not cited yet (Europe PMC); 31 references in the paper

Abstract

Background: Moderate-to-severe preterm brain injury (PBI), including intraventricular hemorrhage (IVH) and periventricular leukomalacia (PVL), remains an important cause of adverse neurodevelopmental outcomes in preterm infants. Early risk stratification using routinely collected clinical data may help prioritize surveillance in vulnerable infants. Methods: We retrospectively included 318 preterm infants admitted between 2015 and 2024 as the development cohort. Thirty-three candidate predictors derived from perinatal factors, first laboratory tests within 24 h of admission, and selected early hospitalization variables were evaluated. Seven machine-learning algorithms were developed using stratified 10 × 5 nested cross-validation with prespecified preprocessing, class-balancing, and feature-selection procedures. Candidate models were compared primarily using the mean fold-level area under the receiver operating characteristic curve (AUROC). After model selection, the finalized LightGBM model was calibrated using Platt scaling, and its pooled out-of-fold (OOF) performance was summarized. Two prespecified thresholds (Youden and high-sensitivity) were used for risk stratification. A small independent temporal cohort of 35 infants was used for preliminary external validation. Results: PBI occurred in 62/318 infants (19.5%) in the development cohort and 6/35 infants (17.1%) in the temporal external cohort. During candidate-model comparison, LightGBM achieved the highest mean fold-level AUROC (0.768, 95% CI 0.708–0.825). The finalized 14-feature LightGBM model, evaluated using pooled OOF predictions after Platt calibration, yielded an AUROC of 0.747 (95% CI 0.679–0.811), a PR-AUC of 0.392, and a Brier score of 0.136. At the Youden threshold (0.18), sensitivity was approximately 0.70 and specificity approximately 0.85; at the high-sensitivity threshold (0.10), sensitivity was approximately 0.95 and specificity approximately 0.50. Key predictors included ventilation status and early physiologic and laboratory indicators. In the small temporal external cohort (n = 35), the AUROC was 0.897 (95% CI 0.672–1.000); however, this high point estimate should not be overinterpreted because of the limited sample size, wide confidence interval, and suboptimal calibration, and should therefore be considered preliminary. Conclusions: We developed an interpretable LightGBM model using routinely available early postnatal and early hospitalization data to support risk stratification for PBI in preterm infants. The model showed moderate internal discrimination and a positive net benefit across clinically relevant thresholds. Preliminary temporal external validation in a small cohort yielded highly uncertain estimates; larger multicenter studies are needed to confirm generalizability, refine calibration, and determine the most appropriate implementation strategy before routine clinical use.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

lyrperciver/PBI-APP

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 6e3d3b3ee120795a0d7db1d21d8c0d44ac7ecf21, 20 November 2025
Languages: Python (1)
Size: 8 files, 1 script
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, environment (requirements.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
2 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The clinical data are not publicly available because of privacy and institutional restrictions. Researchers interested in the data for academic purposes may contact the corresponding author. The code for the research-use web application is available at https://github.com/lyrperciver/PBI-APP (20 November 2025).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 5 keywords, 2 funders, 31 references.

Cite

This paper

Xu, P., Li, Y., Chen, Y., Han, T., Zou, P., Lu, Q., Zhang, D., Chen, J., & Wang, Y. (2026). Developing and Validating a Machine Learning Model to Predict Brain Injury in Preterm Infants Using Multisource Data from the Early Postnatal Period. Children (Basel, Switzerland), 13(6), 796.

BibTeX

@article{xu2026developing,
author = {Xu, Pu and Li, Ying and Chen, Ying and Han, Tongying and Zou, Peicen and Lu, Qinglin and Zhang, Dongmiao and Chen, Jie and Wang, Yajuan},
title = {{Developing and Validating a Machine Learning Model to Predict Brain Injury in Preterm Infants Using Multisource Data from the Early Postnatal Period}},
journal = {Children (Basel, Switzerland)},
year = {2026},
month = jun,
volume = {13},
number = {6},
pages = {796},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {2227-9067},
pmcid = {PMC13297451}
}

RIS

TY - JOUR
AU - Xu, Pu
AU - Li, Ying
AU - Chen, Ying
AU - Han, Tongying
AU - Zou, Peicen
AU - Lu, Qinglin
AU - Zhang, Dongmiao
AU - Chen, Jie
AU - Wang, Yajuan
TI - Developing and Validating a Machine Learning Model to Predict Brain Injury in Preterm Infants Using Multisource Data from the Early Postnatal Period
T2 - Children (Basel, Switzerland)
J2 - Children (Basel)
PY - 2026
DA - 2026/06/01
VL - 13
IS - 6
SP - 796
SN - 2227-9067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
LA - en
ER -

CSL-JSON

{
"id": "pmcid:PMC13297451",
"type": "article-journal",
"title": "Developing and Validating a Machine Learning Model to Predict Brain Injury in Preterm Infants Using Multisource Data from the Early Postnatal Period",
"container-title": "Children (Basel, Switzerland)",
"author": [
{
"family": "Xu",
"given": "Pu"
},
{
"family": "Li",
"given": "Ying"
},
{
"family": "Chen",
"given": "Ying"
},
{
"family": "Han",
"given": "Tongying"
},
{
"family": "Zou",
"given": "Peicen"
},
{
"family": "Lu",
"given": "Qinglin"
},
{
"family": "Zhang",
"given": "Dongmiao"
},
{
"family": "Chen",
"given": "Jie"
},
{
"family": "Wang",
"given": "Yajuan"
}
],
"container-title-short": "Children (Basel)",
"volume": "13",
"issue": "6",
"page": "796",
"PMCID": "PMC13297451",
"ISSN": "2227-9067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.isci.2026.117012 [code]
Multi-level static and dynamic graph-theoretical analyses of resting-state functional networks in a Chinese cohort of preterm neonates.
Journal: iScience
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental, other condition, 1 reference
[2] doi:10.1162/imag.a.1347 [code]
Neural and behavioural correlates of theory of mind reasoning in five-year-old children born preterm.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: scikit-learn, pandas, Matplotlib, 1 other tool, other condition, 1 reference
[3] doi:10.1016/j.dcn.2026.101775 [code]
Neonatal brain-age models in full- and preterm infants.
Journal: Developmental cognitive neuroscience
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental, other condition
[4] doi:10.1038/s41467-026-73770-1 [code]
Non-coding structural variants disrupt FOXG1 transcriptional regulation in early neurodevelopment.
Journal: Nature communications
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental, other condition
[5] doi:10.1186/s40708-026-00312-2 [code]
Synergistic and redundant information dynamics exhibit dissociable alterations across schizophrenia and neurodevelopmental conditions.
Journal: Brain informatics
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental, other condition
[6] doi:10.1038/s42003-026-09908-0 [code]
Offspring genetic diversity regulates rearing experiences that predict differential susceptibility to Chd8 haploinsufficiency.
Journal: Communications biology
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental, other condition
[7] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental
[8] doi:10.21203/rs.3.rs-10335768/v1 [code]
Interneurons in the subplate are associated with cortical maturation across the foetal human cortex
Journal: Research Square (preprint)
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental
[9] doi:10.1038/s41467-026-75722-1 [code]
Single-nucleus analysis of the adult human olfactory epithelium uncovers shared neurogenesis programs with the brain.
Journal: Nature communications
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental
[10] doi:10.7554/elife.107088 [code]
Development of auditory and spontaneous movement responses to music over the first postnatal year.
Journal: eLife
In common: scikit-learn, pandas, Matplotlib, 1 other tool, developmental

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.