OSCR

Detection of early-stage Parkinson's disease using wearable sensors at multiple body locations and convolutional neural networks.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Methods › CNN model training ↔ LOSO_Bootstrap CI_Sensitivity Analysis/Main_v4_Train_TSImgs_v3.py, lines 1425–1474 · score 0.80 · SqueezeNet, DenseNet, ResNet, TS images, CNN models, training
  2. [2] § Methods › Statistical analysis ↔ Machine Learning_Lasso/LASSO_Logistic_BinaryClassification.py, lines 395–455 · score 0.72 · LASSO logistic, LASSO selected feature, F1 score, zero, coefficients, ROC
  3. [3] § Methods › CNN model training ↔ Machine Learning_Lasso/LASSO_Logistic_BinaryClassification.py, lines 203–224 · score 0.69 · confusion matrix, binary classification, F1 score, recall, precision, metrics
  4. [4] § Results › CNN classification performance across gait phases ↔ LOSO_Bootstrap CI_Sensitivity Analysis/Main_v4_Train_TSImgs_v3.py, lines 1425–1474 · score 0.69 · SqueezeNet, DenseNet, ResNet, TS images, CNN, model
  5. [5] § Methods › CNN model training ↔ LOSO_Bootstrap CI_Sensitivity Analysis/Main_v4_Train_TSImgs_v3.py, lines 713–768 · score 0.67 · confusion matrix, binary classification, recall, sensitivity, precision, metrics
  6. [6] § Methods › Data generation ↔ LOSO_Bootstrap CI_Sensitivity Analysis/Main_v4_Train_TSImgs_v3.py, lines 436–536 · score 0.63 · minority class, randomly oversampled, training fold, segment, PD
  7. [7] § Methods › Statistical analysis ↔ LOSO_Bootstrap CI_Sensitivity Analysis/Main_v4_Train_TSImgs_v3.py, lines 436–536 · score 0.54 · minority class, random oversampling, sensitivity, training, fold

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,722 lines · 62 KB · no license · 5 matches

  1. from pathlib import Path
  2. import os
  3. import re
  4. import json
  5. import warnings
  6. from collections import Counter
  7. from itertools import product
  8. import matplotlib.pyplot as plt
  9. import numpy as np
  10. import pandas as pd
  11. import torch
  12. import torchvision
  13. from sklearn.model_selection import train_test_split
  14. from sklearn.metrics import accuracy_score, confusion_matrix, classification_report, roc_auc_score
  15. # =========================================
  16. # Analysis Options
  17. # =========================================
  18. RUN_LOSO = True
  19. RUN_BOOTSTRAP_CI = True
  20. RUN_LEARNING_CURVE = True # learning curve는 추후 가장 의미있는 센서가 나오면 따로 돌리는걸로 변경(시간절약 차원)
  21. # Learning curve settings
  22. # LEARNING_CURVE_TRAIN_RATIOS = [0.50, 1.00]
  23. # LEARNING_CURVE_REPEATS = 2
  24. LEARNING_CURVE_TRAIN_RATIOS = [0.25, 0.50, 0.75, 1.00]
  25. LEARNING_CURVE_REPEATS = 5
  26. # LOSO validation settings
  27. VAL_RATIO = 0.20
  28. RANDOM_STATE = 42
  29. # In this LOSO sensitivity analysis, oversampling is intentionally not applied.
  30. # This avoids ambiguity about whether the held-out subject/test fold was affected
  31. # by RandomOverSampler/SMOTE.
  32. USE_OVERSAMPLING_IN_LOSO = True
  33. # Output tag. This prevents results from training-fold oversampling from being mixed
  34. # with previous no-oversampling LOSO outputs.
  35. OVERSAMPLING_TAG = "TrainFoldOS" if USE_OVERSAMPLING_IN_LOSO else "NoOS"
  36. # Expected number of straight-walking segments/images per participant
  37. # Set to None if the number of segments differs legitimately by participant.
  38. EXPECTED_SEGMENTS_PER_PARTICIPANT = 3
  39. # Since each analysis folder contains one predefined condition
  40. # e.g., Larm, Rarm, MoreAffectedArm, LessAffectedArm, DominantArm, NonDominantArm,
  41. # the LOSO grouping should be based on participant identity, not sensor name.
  42. APPEND_SENSOR_TO_LOSO_SUBJECT = False
  43. from LibKIME.LibGeneral import (
  44. MakeFolder,
  45. StartTimer,
  46. StopTimer,
  47. Save_dict2json,
  48. LOGGER,
  49. GetNowString,
  50. )
  51. from LibKIME.LibML_Exp import Exp_Detail, CV_Loop
  52. # =========================================
  53. # Working Folder
  54. # =========================================
  55. code_root_folder = r'C:\Users\User\Desktop\CHJ_v3\TSC_TS2ImgCNN_240924'
  56. csv_data_root_folder = r'C:\Users\User\Desktop\CHJ_v3\Raw_linear_EarlyPDs\MASarm_1min'
  57. tsimg_root_path = r'C:\Users\User\Desktop\CHJ_v3\Imaging_linear_EarlyPDs\MASarm_1min'
  58. dl_result_root_path = r'C:\Users\User\Desktop\CHJ_v3\Results_linear_EarlyPDs\MASarm_1min\CNN_LOSO'
  59. import data_specific.Walk_6m_Choi.Info as Info
  60. import data_specific.PD_Park.Task as Task
  61. def safe_save_json(path, obj):
  62. """Save JSON with UTF-8 encoding. Falls back to LibKIME Save_dict2json if possible."""
  63. try:
  64. with open(path, "w", encoding="utf-8") as f:
  65. json.dump(obj, f, indent=4, ensure_ascii=False)
  66. except TypeError:
  67. # If some object is not serializable, convert to string
  68. def default(o):
  69. try:
  70. return o.__dict__
  71. except Exception:
  72. return str(o)
  73. with open(path, "w", encoding="utf-8") as f:
  74. json.dump(obj, f, indent=4, ensure_ascii=False, default=default)
  75. def Add_SensorName_From_Dataset(df, ds_pt):
  76. """
  77. Add SensorName to df using file names in ds_pt.samples or ds_pt.imgs.
  78. Note:
  79. - This is used for traceability and metadata only.
  80. - In the final LOSO split, sensor name is NOT appended to LOSO_SUBJECT by default
  81. because each folder is assumed to contain one predefined sensor/alignment condition.
  82. """
  83. df = df.copy()
  84. if hasattr(ds_pt, "samples"):
  85. sample_paths = [x[0] for x in ds_pt.samples]
  86. elif hasattr(ds_pt, "imgs"):
  87. sample_paths = [x[0] for x in ds_pt.imgs]
  88. else:
  89. raise AttributeError("ds_pt에서 samples 또는 imgs 속성을 찾을 수 없습니다.")
  90. sensor_list = []
  91. for idx in df["SampleIdx"]:
  92. path = sample_paths[int(idx)]
  93. fname = os.path.basename(path)
  94. m = re.search(r"(Larm|Rarm|Lthi|Rthi|Lumba|Thora)", fname)
  95. if m:
  96. sensor_list.append(m.group(1))
  97. else:
  98. sensor_list.append("Unknown")
  99. df["SensorName"] = sensor_list
  100. print(df[["UNIQ_CLASS_SID", "SensorName", "SampleIdx"]].head(20))
  101. return df
  102. def make_loso_subject_id(x):
  103. """
  104. Convert sample/segment-level ID to participant-level ID.
  105. Expected examples:
  106. - Cons_AGR1, Cons_AGR2, Cons_AGR3 -> Cons_AGR
  107. - EarlyPDs_BMY1, EarlyPDs_BMY2, EarlyPDs_BMY3 -> EarlyPDs_BMY
  108. - If file naming includes sensor information, it is ignored here because
  109. LOSO should be participant-level within a single predefined analysis folder.
  110. IMPORTANT:
  111. If the original ID system contains real participant numbers ending in 1/2/3,
  112. replace this function with a metadata-based subject_id column.
  113. """
  114. x = str(x)
  115. # Remove optional 6MWT/sensor/trial suffix if present
  116. # Example: EarlyPDs_CGS1_6MWT_Larm(1) -> EarlyPDs_CGS1
  117. x0 = re.sub(r"_6MWT.*$", "", x)
  118. x0 = re.sub(r"_(Larm|Rarm|Lthi|Rthi|Lumba|Thora).*$", "", x0)
  119. # Remove final segment marker 1, 2, or 3
  120. # Example: Cons_AGR1 -> Cons_AGR
  121. base = re.sub(r"[123]$", "", x0)
  122. return base
  123. def create_subject_table(df, subject_col="LOSO_SUBJECT", label_col="Label"):
  124. """
  125. Create subject-level table and check whether each subject has a single label.
  126. """
  127. label_nunique = df.groupby(subject_col)[label_col].nunique()
  128. bad_subjects = label_nunique[label_nunique > 1]
  129. if len(bad_subjects) > 0:
  130. raise ValueError(
  131. f"다음 LOSO subject에 두 개 이상의 label이 포함되어 있습니다: {bad_subjects.index.tolist()}"
  132. )
  133. subj_df = (
  134. df.groupby(subject_col)
  135. .agg(
  136. Label=(label_col, "first"),
  137. n_samples=("SampleIdx", "count"),
  138. ClassName=("ClassName", "first") if "ClassName" in df.columns else (label_col, "first"),
  139. )
  140. .reset_index()
  141. )
  142. return subj_df
  143. def stratified_subject_train_val_split(
  144. train_val_df,
  145. subject_col="LOSO_SUBJECT",
  146. label_col="Label",
  147. val_ratio=0.2,
  148. random_state=42,
  149. ):
  150. """
  151. Split remaining subjects into training and validation sets at the subject level.
  152. Stratification is used when possible. If stratification is impossible because
  153. class counts are too small after leaving one subject out, this function falls
  154. back to unstratified subject-level splitting and issues a warning.
  155. Returns
  156. -------
  157. train_subjects : np.ndarray
  158. val_subjects : np.ndarray
  159. split_info : dict
  160. """
  161. subj_df = create_subject_table(train_val_df, subject_col=subject_col, label_col=label_col)
  162. remain_subjects = subj_df[subject_col].to_numpy()
  163. remain_labels = subj_df[label_col].to_numpy()
  164. n_subjects = len(remain_subjects)
  165. if n_subjects < 2:
  166. raise ValueError("Train/validation split을 수행하기에 남은 subject 수가 너무 적습니다.")
  167. # Convert ratio to an integer test_size to avoid zero validation subject
  168. n_val = int(np.ceil(n_subjects * val_ratio))
  169. n_val = max(1, n_val)
  170. n_val = min(n_val, n_subjects - 1)
  171. unique_labels, counts = np.unique(remain_labels, return_counts=True)
  172. n_classes = len(unique_labels)
  173. # Stratification is possible only when:
  174. # 1) every class has at least 2 subjects
  175. # 2) validation split can contain all classes
  176. # 3) training split can contain all classes
  177. can_stratify = (
  178. n_classes >= 2 and
  179. np.min(counts) >= 2 and
  180. n_val >= n_classes and
  181. (n_subjects - n_val) >= n_classes
  182. )
  183. split_info = {
  184. "n_remaining_subjects": int(n_subjects),
  185. "n_validation_subjects": int(n_val),
  186. "class_counts_remaining": {str(k): int(v) for k, v in zip(unique_labels, counts)},
  187. "stratified": bool(can_stratify),
  188. "fallback_reason": None,
  189. }
  190. try:
  191. if can_stratify:
  192. train_subjects, val_subjects = train_test_split(
  193. remain_subjects,
  194. test_size=n_val,
  195. random_state=random_state,
  196. stratify=remain_labels,
  197. )
  198. else:
  199. split_info["fallback_reason"] = (
  200. "Stratified split not feasible due to small class counts or insufficient validation subjects."
  201. )
  202. warnings.warn(split_info["fallback_reason"])
  203. train_subjects, val_subjects = train_test_split(
  204. remain_subjects,
  205. test_size=n_val,
  206. random_state=random_state,
  207. stratify=None,
  208. )
  209. except ValueError as e:
  210. split_info["stratified"] = False
  211. split_info["fallback_reason"] = f"Stratified split failed: {str(e)}"
  212. warnings.warn(split_info["fallback_reason"])
  213. train_subjects, val_subjects = train_test_split(
  214. remain_subjects,
  215. test_size=n_val,
  216. random_state=random_state,
  217. stratify=None,
  218. )
  219. return np.array(train_subjects), np.array(val_subjects), split_info
  220. def Make_LOSO_Index_bySubject(
  221. df,
  222. sensor_name=None,
  223. val_ratio=0.2,
  224. random_state=42,
  225. append_sensor_to_subject=False,
  226. ):
  227. """
  228. Generate LOSO Train/Val/Test indices at the participant level.
  229. Important:
  230. - This assumes that each analysis folder contains one predefined condition
  231. such as one sensor, more-affected side, less-affected side, dominant side, etc.
  232. - By default, sensor name is NOT appended to LOSO_SUBJECT.
  233. """
  234. print("df columns:", df.columns)
  235. print(df.head())
  236. subject_col = "UNIQ_CLASS_SID"
  237. index_col = "SampleIdx"
  238. label_col = "Label"
  239. df = df.copy()
  240. df["BASE_SUBJECT"] = df[subject_col].apply(make_loso_subject_id)
  241. if append_sensor_to_subject:
  242. if "SensorName" in df.columns:
  243. df["LOSO_SUBJECT"] = df["BASE_SUBJECT"] + "_" + df["SensorName"].astype(str)
  244. elif sensor_name is not None:
  245. df["LOSO_SUBJECT"] = df["BASE_SUBJECT"] + "_" + str(sensor_name)
  246. else:
  247. df["LOSO_SUBJECT"] = df["BASE_SUBJECT"]
  248. else:
  249. df["LOSO_SUBJECT"] = df["BASE_SUBJECT"]
  250. loso_subject_col = "LOSO_SUBJECT"
  251. subjects = sorted(df[loso_subject_col].unique())
  252. Train_Index = []
  253. Val_Index = []
  254. Test_Index = []
  255. split_info_rows = []
  256. for fold_i, test_subject in enumerate(subjects):
  257. test_df = df[df[loso_subject_col] == test_subject]
  258. train_val_df = df[df[loso_subject_col] != test_subject]
  259. train_subjects, val_subjects, split_info = stratified_subject_train_val_split(
  260. train_val_df,
  261. subject_col=loso_subject_col,
  262. label_col=label_col,
  263. val_ratio=val_ratio,
  264. random_state=random_state,
  265. )
  266. train_df = train_val_df[train_val_df[loso_subject_col].isin(train_subjects)]
  267. val_df = train_val_df[train_val_df[loso_subject_col].isin(val_subjects)]
  268. Train_Index.append(train_df[index_col].tolist())
  269. Val_Index.append(val_df[index_col].tolist())
  270. Test_Index.append(test_df[index_col].tolist())
  271. split_info_row = {
  272. "fold": fold_i,
  273. "test_subject": test_subject,
  274. "test_label": int(test_df[label_col].iloc[0]),
  275. "n_train_subjects": int(train_df[loso_subject_col].nunique()),
  276. "n_val_subjects": int(val_df[loso_subject_col].nunique()),
  277. "n_test_subjects": int(test_df[loso_subject_col].nunique()),
  278. "n_train_samples": int(len(train_df)),
  279. "n_val_samples": int(len(val_df)),
  280. "n_test_samples": int(len(test_df)),
  281. **split_info,
  282. }
  283. split_info_rows.append(split_info_row)
  284. print(
  285. f"Fold {fold_i:03d} | Test subject: {test_subject} | "
  286. f"Train subjects: {train_df[loso_subject_col].nunique()} | "
  287. f"Val subjects: {val_df[loso_subject_col].nunique()} | "
  288. f"Test samples: {len(test_df)} | Stratified: {split_info['stratified']}"
  289. )
  290. out_cv_ind_sample = {
  291. "Train_Index": Train_Index,
  292. "Val_Index": Val_Index,
  293. "Test_Index": Test_Index,
  294. }
  295. split_info_df = pd.DataFrame(split_info_rows)
  296. return out_cv_ind_sample, len(subjects), df, split_info_df
  297. def check_no_subject_overlap(df, cv_index, subject_col="LOSO_SUBJECT"):
  298. """
  299. Verify that Train/Val/Test subject sets do not overlap for any fold.
  300. """
  301. rows = []
  302. for i, (tr, va, te) in enumerate(zip(
  303. cv_index["Train_Index"],
  304. cv_index["Val_Index"],
  305. cv_index["Test_Index"],
  306. )):
  307. tr_sub = set(df.loc[df["SampleIdx"].isin(tr), subject_col])
  308. va_sub = set(df.loc[df["SampleIdx"].isin(va), subject_col])
  309. te_sub = set(df.loc[df["SampleIdx"].isin(te), subject_col])
  310. train_test_overlap = sorted(tr_sub.intersection(te_sub))
  311. val_test_overlap = sorted(va_sub.intersection(te_sub))
  312. train_val_overlap = sorted(tr_sub.intersection(va_sub))
  313. rows.append({
  314. "fold": i,
  315. "n_train_subjects": len(tr_sub),
  316. "n_val_subjects": len(va_sub),
  317. "n_test_subjects": len(te_sub),
  318. "train_test_overlap": ",".join(train_test_overlap),
  319. "val_test_overlap": ",".join(val_test_overlap),
  320. "train_val_overlap": ",".join(train_val_overlap),
  321. "no_overlap": (
  322. len(train_test_overlap) == 0 and
  323. len(val_test_overlap) == 0 and
  324. len(train_val_overlap) == 0
  325. )
  326. })
  327. assert len(train_test_overlap) == 0, f"Train-Test leakage at fold {i}: {train_test_overlap}"
  328. assert len(val_test_overlap) == 0, f"Val-Test leakage at fold {i}: {val_test_overlap}"
  329. assert len(train_val_overlap) == 0, f"Train-Val overlap at fold {i}: {train_val_overlap}"
  330. print("No subject-level overlap detected.")
  331. return pd.DataFrame(rows)
  332. def summarize_fold_distribution(df, cv_index, subject_col="LOSO_SUBJECT", label_col="Label"):
  333. """
  334. Summarize fold-wise class distributions at both sample and subject levels.
  335. """
  336. rows = []
  337. for i, (tr, va, te) in enumerate(zip(
  338. cv_index["Train_Index"],
  339. cv_index["Val_Index"],
  340. cv_index["Test_Index"],
  341. )):
  342. for split_name, idx in [("train", tr), ("val", va), ("test", te)]:
  343. temp = df[df["SampleIdx"].isin(idx)]
  344. subj_temp = create_subject_table(
  345. temp,
  346. subject_col=subject_col,
  347. label_col=label_col
  348. ) if len(temp) > 0 else pd.DataFrame(columns=[subject_col, label_col])
  349. rows.append({
  350. "fold": i,
  351. "split": split_name,
  352. "n_samples": int(len(temp)),
  353. "n_subjects": int(temp[subject_col].nunique()) if len(temp) > 0 else 0,
  354. "n_control_samples": int((temp[label_col] == 0).sum()) if len(temp) > 0 else 0,
  355. "n_pd_samples": int((temp[label_col] == 1).sum()) if len(temp) > 0 else 0,
  356. "n_control_subjects": int((subj_temp[label_col] == 0).sum()) if len(subj_temp) > 0 else 0,
  357. "n_pd_subjects": int((subj_temp[label_col] == 1).sum()) if len(subj_temp) > 0 else 0,
  358. })
  359. return pd.DataFrame(rows)
  360. def oversample_train_indices_by_subject(
  361. df,
  362. train_indices,
  363. subject_col="LOSO_SUBJECT",
  364. label_col="Label",
  365. index_col="SampleIdx",
  366. random_state=42,
  367. ):
  368. """
  369. Apply subject-level random oversampling strictly within one training fold.
  370. This function balances the PD/HC class counts at the subject level by duplicating
  371. minority-class training participants. When a participant is duplicated, all of
  372. that participant's segment/image indices are duplicated together.
  373. Important:
  374. - Only Train_Index is modified.
  375. - Val_Index and Test_Index must remain untouched.
  376. - This does not create new independent participants; it only duplicates
  377. training participants within the current fold.
  378. """
  379. rng = np.random.default_rng(random_state)
  380. train_indices = list(train_indices)
  381. train_df = df[df[index_col].isin(train_indices)].copy()
  382. subj_df = (
  383. train_df.groupby(subject_col)
  384. .agg(
  385. label=(label_col, "first"),
  386. sample_indices=(index_col, lambda x: list(x)),
  387. n_samples=(index_col, "count"),
  388. )
  389. .reset_index()
  390. )
  391. before_subject_counts = subj_df["label"].value_counts().to_dict()
  392. before_sample_counts = train_df[label_col].value_counts().to_dict()
  393. info = {
  394. "oversampling_applied": False,
  395. "oversampling_unit": "participant",
  396. "before_subject_counts": json.dumps({str(k): int(v) for k, v in before_subject_counts.items()}),
  397. "before_sample_counts": json.dumps({str(k): int(v) for k, v in before_sample_counts.items()}),
  398. "after_subject_counts_including_duplicates": None,
  399. "after_sample_counts_including_duplicates": None,
  400. "n_original_train_indices": int(len(train_indices)),
  401. "n_oversampled_train_indices": int(len(train_indices)),
  402. "n_added_subject_duplicates": 0,
  403. "n_added_sample_indices": 0,
  404. "warning": None,
  405. }
  406. if len(before_subject_counts) < 2:
  407. info["warning"] = "Only one class exists in this training fold. Oversampling skipped."
  408. warnings.warn(info["warning"])
  409. return train_indices, info
  410. max_count = max(before_subject_counts.values())
  411. oversampled_indices = list(train_indices)
  412. duplicated_subject_labels = []
  413. for cls, count in before_subject_counts.items():
  414. if count < max_count:
  415. need = max_count - count
  416. minority_subjects = subj_df[subj_df["label"] == cls]
  417. # subject-level random oversampling with replacement
  418. sampled_subject_rows = minority_subjects.sample(
  419. n=need,
  420. replace=True,
  421. random_state=random_state,
  422. )
  423. for _, row in sampled_subject_rows.iterrows():
  424. duplicate_indices = list(row["sample_indices"])
  425. oversampled_indices.extend(duplicate_indices)
  426. duplicated_subject_labels.append(int(cls))
  427. # Summarize after oversampling based on duplicated subject labels and duplicated sample indices
  428. after_subject_counts = dict(before_subject_counts)
  429. for cls in duplicated_subject_labels:
  430. after_subject_counts[cls] = after_subject_counts.get(cls, 0) + 1
  431. # sample-level counts including duplicated indices
  432. label_map = df.set_index(index_col)[label_col].to_dict()
  433. after_sample_labels = [label_map[i] for i in oversampled_indices]
  434. after_sample_counts = Counter(after_sample_labels)
  435. info["oversampling_applied"] = len(duplicated_subject_labels) > 0
  436. info["after_subject_counts_including_duplicates"] = json.dumps(
  437. {str(k): int(v) for k, v in after_subject_counts.items()}
  438. )
  439. info["after_sample_counts_including_duplicates"] = json.dumps(
  440. {str(k): int(v) for k, v in after_sample_counts.items()}
  441. )
  442. info["n_oversampled_train_indices"] = int(len(oversampled_indices))
  443. info["n_added_subject_duplicates"] = int(len(duplicated_subject_labels))
  444. info["n_added_sample_indices"] = int(len(oversampled_indices) - len(train_indices))
  445. return oversampled_indices, info
  446. def apply_training_fold_oversampling(
  447. df,
  448. cv_index,
  449. subject_col="LOSO_SUBJECT",
  450. label_col="Label",
  451. index_col="SampleIdx",
  452. random_state=42,
  453. ):
  454. """
  455. Apply subject-level oversampling only to the training indices of each LOSO fold.
  456. Validation and test indices are copied unchanged.
  457. """
  458. oversampled_train_indices = []
  459. oversampling_rows = []
  460. for fold_i, train_idx in enumerate(cv_index["Train_Index"]):
  461. os_train_idx, info = oversample_train_indices_by_subject(
  462. df=df,
  463. train_indices=train_idx,
  464. subject_col=subject_col,
  465. label_col=label_col,
  466. index_col=index_col,
  467. random_state=random_state + fold_i,
  468. )
  469. oversampled_train_indices.append(os_train_idx)
  470. row = {
  471. "fold": fold_i,
  472. **info,
  473. }
  474. oversampling_rows.append(row)
  475. out_cv_ind_os = {
  476. "Train_Index": oversampled_train_indices,
  477. "Val_Index": cv_index["Val_Index"],
  478. "Test_Index": cv_index["Test_Index"],
  479. }
  480. oversampling_info_df = pd.DataFrame(oversampling_rows)
  481. return out_cv_ind_os, oversampling_info_df
  482. def summarize_fold_distribution_from_indices_with_duplicates(
  483. df,
  484. cv_index,
  485. subject_col="LOSO_SUBJECT",
  486. label_col="Label",
  487. index_col="SampleIdx",
  488. ):
  489. """
  490. Summarize fold-wise class distributions while preserving duplicated training indices.
  491. This is useful after oversampling because df[df.SampleIdx.isin(indices)] removes
  492. duplicate indices and therefore cannot show oversampled counts.
  493. """
  494. rows = []
  495. label_map = df.set_index(index_col)[label_col].to_dict()
  496. subject_map = df.set_index(index_col)[subject_col].to_dict()
  497. for fold_i, (tr, va, te) in enumerate(zip(
  498. cv_index["Train_Index"],
  499. cv_index["Val_Index"],
  500. cv_index["Test_Index"],
  501. )):
  502. for split_name, idx in [("train", tr), ("val", va), ("test", te)]:
  503. idx = list(idx)
  504. labels = [label_map[i] for i in idx]
  505. subjects = [subject_map[i] for i in idx]
  506. label_counts = Counter(labels)
  507. unique_subjects = sorted(set(subjects))
  508. # Unique subject label counts
  509. subj_labels = []
  510. for sub in unique_subjects:
  511. sub_label_values = df.loc[df[subject_col] == sub, label_col].unique()
  512. if len(sub_label_values) == 1:
  513. subj_labels.append(int(sub_label_values[0]))
  514. subject_label_counts = Counter(subj_labels)
  515. rows.append({
  516. "fold": fold_i,
  517. "split": split_name,
  518. "n_indices_including_duplicates": int(len(idx)),
  519. "n_unique_subjects": int(len(unique_subjects)),
  520. "n_control_indices": int(label_counts.get(0, 0)),
  521. "n_pd_indices": int(label_counts.get(1, 0)),
  522. "n_control_unique_subjects": int(subject_label_counts.get(0, 0)),
  523. "n_pd_unique_subjects": int(subject_label_counts.get(1, 0)),
  524. })
  525. return pd.DataFrame(rows)
  526. def save_subject_segment_count(df, result_path, title, expected_segments=3):
  527. """
  528. Save subject-level segment counts and warn if a participant does not have the expected number of samples.
  529. """
  530. seg_count = (
  531. df.groupby("LOSO_SUBJECT")
  532. .agg(
  533. n_samples=("SampleIdx", "count"),
  534. label=("Label", "first"),
  535. class_name=("ClassName", "first") if "ClassName" in df.columns else ("Label", "first"),
  536. )
  537. .reset_index()
  538. )
  539. seg_count.to_csv(
  540. os.path.join(result_path, f"Subject_Segment_Count_{title}.csv"),
  541. index=False,
  542. encoding="utf-8-sig",
  543. )
  544. if expected_segments is not None:
  545. abnormal = seg_count[seg_count["n_samples"] != expected_segments]
  546. if len(abnormal) > 0:
  547. msg = (
  548. f"Warning: Some subjects do not have exactly {expected_segments} segments. "
  549. "Check whether this is expected."
  550. )
  551. print(msg)
  552. print(abnormal)
  553. abnormal.to_csv(
  554. os.path.join(result_path, f"Subject_Segment_Count_Abnormal_{title}.csv"),
  555. index=False,
  556. encoding="utf-8-sig",
  557. )
  558. return seg_count
  559. def majority_vote(preds):
  560. """
  561. Majority vote for subject-level label.
  562. If tie occurs, choose the smallest label and return tie=True.
  563. For three segments, tie is unlikely.
  564. """
  565. preds = list(map(int, preds))
  566. counts = Counter(preds)
  567. max_count = max(counts.values())
  568. winners = sorted([label for label, count in counts.items() if count == max_count])
  569. tie = len(winners) > 1
  570. return winners[0], tie
  571. def extract_test_probs_if_available(cv_obj):
  572. """
  573. Try to extract test probabilities from a CV object if the LibKIME CV_Loop stores them.
  574. Returns None if no probability field exists.
  575. Because existing result objects only appear to include r_test_labels and r_test_preds,
  576. subject-level analysis defaults to majority voting.
  577. """
  578. prob_attr_candidates = [
  579. "r_test_probs",
  580. "r_test_prob",
  581. "r_test_pred_probs",
  582. "r_test_pred_prob",
  583. "r_test_outputs",
  584. "r_test_logits",
  585. "r_test_scores",
  586. ]
  587. for attr in prob_attr_candidates:
  588. if hasattr(cv_obj, attr):
  589. val = getattr(cv_obj, attr)
  590. if val is not None:
  591. arr = np.asarray(val)
  592. if arr.ndim == 2:
  593. return arr
  594. return None
  595. def compute_binary_metrics(y_true, y_pred, y_score=None, positive_label=1):
  596. """
  597. Compute binary classification metrics explicitly.
  598. Label convention:
  599. - 0 = control / Cons
  600. - 1 = PD / EarlyPDs
  601. sensitivity = recall for positive_label=1
  602. specificity = recall for negative class=0
  603. AUC is calculated only when y_score is available.
  604. """
  605. y_true = np.asarray(y_true).astype(int)
  606. y_pred = np.asarray(y_pred).astype(int)
  607. metrics = {
  608. "accuracy": np.nan,
  609. "sensitivity": np.nan,
  610. "specificity": np.nan,
  611. "precision": np.nan,
  612. "f1": np.nan,
  613. "balanced_accuracy": np.nan,
  614. "auc": np.nan,
  615. }
  616. if len(y_true) == 0:
  617. return metrics
  618. metrics["accuracy"] = float(accuracy_score(y_true, y_pred))
  619. # Force binary confusion-matrix order [0, 1]
  620. cm = confusion_matrix(y_true, y_pred, labels=[0, 1])
  621. tn, fp, fn, tp = cm.ravel()
  622. metrics["sensitivity"] = float(tp / (tp + fn)) if (tp + fn) > 0 else np.nan
  623. metrics["specificity"] = float(tn / (tn + fp)) if (tn + fp) > 0 else np.nan
  624. metrics["precision"] = float(tp / (tp + fp)) if (tp + fp) > 0 else np.nan
  625. metrics["f1"] = (
  626. float(2 * tp / (2 * tp + fp + fn))
  627. if (2 * tp + fp + fn) > 0
  628. else np.nan
  629. )
  630. if not np.isnan(metrics["sensitivity"]) and not np.isnan(metrics["specificity"]):
  631. metrics["balanced_accuracy"] = float((metrics["sensitivity"] + metrics["specificity"]) / 2)
  632. if y_score is not None:
  633. y_score = np.asarray(y_score, dtype=float)
  634. valid = ~np.isnan(y_score)
  635. if valid.sum() > 0 and len(np.unique(y_true[valid])) == 2:
  636. try:
  637. metrics["auc"] = float(roc_auc_score(y_true[valid], y_score[valid]))
  638. except Exception:
  639. metrics["auc"] = np.nan
  640. return metrics
  641. def collect_loso_subject_level_results(exp, loso_cv, df, result_path, title):
  642. """
  643. Save both segment-level and subject-level LOSO results for each model.
  644. Subject-level prediction:
  645. - If test probabilities are available: average probabilities across segments.
  646. - Otherwise: majority vote across segment-level predictions.
  647. """
  648. all_model_summary = []
  649. all_fold_rows = []
  650. for mdl in exp.models:
  651. model_name = mdl.MdlInfo.model_name
  652. subject_correct_list = []
  653. segment_acc_list = []
  654. subject_true_list = []
  655. subject_pred_list = []
  656. subject_score_list = [] # score/probability for class 1, if available
  657. all_segment_labels = []
  658. all_segment_preds = []
  659. all_segment_scores = [] # score/probability for class 1, if available
  660. for cv_i in range(loso_cv):
  661. cv_obj = mdl.CVs[cv_i]
  662. labels = np.asarray(cv_obj.r_test_labels).astype(int)
  663. preds = np.asarray(cv_obj.r_test_preds).astype(int)
  664. if len(labels) == 0:
  665. warnings.warn(f"{model_name} fold {cv_i}: no test labels.")
  666. continue
  667. # Get test subject ID from fold indices
  668. if hasattr(cv_obj, "ind_test"):
  669. test_indices = list(cv_obj.ind_test)
  670. else:
  671. test_indices = []
  672. test_subjects = sorted(
  673. set(df.loc[df["SampleIdx"].isin(test_indices), "LOSO_SUBJECT"])
  674. ) if len(test_indices) > 0 else []
  675. if len(test_subjects) != 1:
  676. warnings.warn(
  677. f"{model_name} fold {cv_i}: expected one test subject, found {test_subjects}"
  678. )
  679. test_subject = test_subjects[0] if len(test_subjects) > 0 else f"fold_{cv_i}"
  680. # Segment-level accuracy within the held-out subject
  681. seg_acc = accuracy_score(labels, preds)
  682. segment_acc_list.append(seg_acc)
  683. # Subject true label
  684. if len(np.unique(labels)) > 1:
  685. warnings.warn(
  686. f"{model_name} fold {cv_i}: test labels are not identical within subject. Using majority true label."
  687. )
  688. subject_true, true_tie = majority_vote(labels)
  689. else:
  690. subject_true = int(labels[0])
  691. true_tie = False
  692. probs = extract_test_probs_if_available(cv_obj)
  693. if probs is not None and probs.shape[0] == len(labels):
  694. mean_prob = probs.mean(axis=0)
  695. subject_pred = int(np.argmax(mean_prob))
  696. pred_method = "mean_probability"
  697. pred_tie = False
  698. else:
  699. subject_pred, pred_tie = majority_vote(preds)
  700. pred_method = "majority_vote"
  701. subject_correct = int(subject_pred == subject_true)
  702. subject_correct_list.append(subject_correct)
  703. subject_true_list.append(subject_true)
  704. subject_pred_list.append(subject_pred)
  705. all_segment_labels.extend(labels.tolist())
  706. all_segment_preds.extend(preds.tolist())
  707. all_fold_rows.append({
  708. "model": model_name,
  709. "fold": cv_i,
  710. "test_subject": test_subject,
  711. "n_test_segments": int(len(labels)),
  712. "segment_accuracy_within_subject": float(seg_acc),
  713. "subject_true": int(subject_true),
  714. "subject_pred": int(subject_pred),
  715. "subject_correct": int(subject_correct),
  716. "prediction_method": pred_method,
  717. "tie_in_prediction": bool(pred_tie),
  718. "tie_in_true_label": bool(true_tie),
  719. })
  720. # Explicit subject-level metrics
  721. subject_metrics = compute_binary_metrics(
  722. subject_true_list,
  723. subject_pred_list,
  724. y_score=subject_score_list,
  725. positive_label=1,
  726. )
  727. # Explicit segment-level metrics
  728. segment_metrics = compute_binary_metrics(
  729. all_segment_labels,
  730. all_segment_preds,
  731. y_score=all_segment_scores,
  732. positive_label=1,
  733. )
  734. subject_accuracy = subject_metrics["accuracy"]
  735. segment_accuracy_global = segment_metrics["accuracy"]
  736. segment_accuracy_mean_by_fold = float(np.mean(segment_acc_list)) if len(segment_acc_list) else np.nan
  737. cm_subject = confusion_matrix(subject_true_list, subject_pred_list, labels=[0, 1]).tolist() if len(subject_true_list) else []
  738. report_subject = classification_report(
  739. subject_true_list,
  740. subject_pred_list,
  741. output_dict=True,
  742. zero_division=0
  743. ) if len(subject_true_list) else {}
  744. all_model_summary.append({
  745. "model": model_name,
  746. "n_loso_subjects": int(len(subject_correct_list)),
  747. "subject_level_accuracy": subject_accuracy,
  748. "subject_level_sensitivity": subject_metrics["sensitivity"],
  749. "subject_level_specificity": subject_metrics["specificity"],
  750. "subject_level_precision": subject_metrics["precision"],
  751. "subject_level_f1": subject_metrics["f1"],
  752. "subject_level_balanced_accuracy": subject_metrics["balanced_accuracy"],
  753. "subject_level_auc": subject_metrics["auc"],
  754. # Segment-level metrics, secondary
  755. "segment_level_accuracy_global": segment_accuracy_global,
  756. "segment_level_sensitivity": segment_metrics["sensitivity"],
  757. "segment_level_specificity": segment_metrics["specificity"],
  758. "segment_level_precision": segment_metrics["precision"],
  759. "segment_level_f1": segment_metrics["f1"],
  760. "segment_level_balanced_accuracy": segment_metrics["balanced_accuracy"],
  761. "segment_level_auc": segment_metrics["auc"],
  762. "segment_level_accuracy_mean_by_fold": segment_accuracy_mean_by_fold,
  763. # Raw diagnostic objects
  764. "subject_confusion_matrix": json.dumps(cm_subject),
  765. "subject_classification_report": json.dumps(report_subject),
  766. "auc_note": (
  767. "AUC is calculated only if probability/scores are available from CV_Loop; "
  768. "otherwise it is saved as NaN."
  769. ),
  770. })
  771. fold_df = pd.DataFrame(all_fold_rows)
  772. summary_df = pd.DataFrame(all_model_summary)
  773. fold_df.to_csv(
  774. os.path.join(result_path, f"LOSO_SubjectLevel_FoldResults_{title}.csv"),
  775. index=False,
  776. encoding="utf-8-sig",
  777. )
  778. summary_df.to_csv(
  779. os.path.join(result_path, f"LOSO_SubjectLevel_Summary_{title}.csv"),
  780. index=False,
  781. encoding="utf-8-sig",
  782. )
  783. return fold_df, summary_df
  784. # =========================================
  785. # Bootstrap CI: subject-level
  786. # =========================================
  787. # 기존 방식은 fold별 accuracy list를 resampling함
  788. # 그러나 현재 segment-level accuracy이기 때문에 subject-level accuracy로 변경해서 subject 단위 bootstrap CI를 산출해야 함
  789. # 결과는 이 모델이 이 데이터 샘플수에서 불확실성 크다/작다를 해석하기 위해 사용
  790. def bootstrap_ci_from_binary_correct(correct_list, n_bootstrap=2000, ci=95, random_state=42):
  791. """
  792. Participant-level bootstrap CI using subject-level correct/incorrect values.
  793. """
  794. rng = np.random.default_rng(random_state)
  795. correct_list = np.asarray(correct_list, dtype=float)
  796. if len(correct_list) == 0:
  797. return np.nan, np.nan, np.nan
  798. boot_means = []
  799. for _ in range(n_bootstrap):
  800. sample = rng.choice(correct_list, size=len(correct_list), replace=True)
  801. boot_means.append(np.mean(sample))
  802. alpha = (100 - ci) / 2
  803. lower = np.percentile(boot_means, alpha)
  804. upper = np.percentile(boot_means, 100 - alpha)
  805. mean_acc = np.mean(correct_list)
  806. return float(mean_acc), float(lower), float(upper)
  807. def save_bootstrap_ci_from_subject_fold_results(subject_fold_df, result_path, title):
  808. """
  809. Save subject-level bootstrap CI for each model.
  810. """
  811. bootstrap_results = []
  812. for model_name, temp in subject_fold_df.groupby("model"):
  813. correct_list = temp["subject_correct"].astype(int).tolist()
  814. mean_acc, ci_lower, ci_upper = bootstrap_ci_from_binary_correct(
  815. correct_list,
  816. n_bootstrap=2000,
  817. ci=95,
  818. random_state=RANDOM_STATE,
  819. )
  820. bootstrap_results.append({
  821. "model": model_name,
  822. "subject_level_mean_accuracy": mean_acc,
  823. "ci_lower_95": ci_lower,
  824. "ci_upper_95": ci_upper,
  825. "n_loso_subjects": int(len(correct_list)),
  826. "bootstrap_unit": "held-out participant",
  827. })
  828. bootstrap_df = pd.DataFrame(bootstrap_results)
  829. bootstrap_df.to_csv(
  830. os.path.join(result_path, f"Bootstrap_CI_LOSO_SubjectLevel_{title}.csv"),
  831. index=False,
  832. encoding="utf-8-sig",
  833. )
  834. return bootstrap_df
  835. # ================================================================
  836. # Learning Curve Analysis: repeated subject-level subsampling
  837. # ================================================================
  838. # 기존 방식에서 그래프가 error bar 너무 크게 나타남 -> 불확실성 매우 큼
  839. # 결과는 현재 표본 수가 충분하다기 보다 샘플 사이즈 제한점을 뒷받침 하는 보완 분석으로 사용
  840. def subsample_train_subjects_stratified(
  841. train_subjects,
  842. df,
  843. train_ratio,
  844. subject_col="LOSO_SUBJECT",
  845. label_col="Label",
  846. random_state=42,
  847. ):
  848. """
  849. Subsample training subjects for learning curve.
  850. - If train_ratio >= 1.0, use all training subjects.
  851. - If the calculated number of selected subjects is equal to or greater than
  852. the available training subjects, use all subjects without calling train_test_split.
  853. - Stratified subsampling is used only when feasible.
  854. - If stratification is not feasible due to small class counts, random
  855. subject-level subsampling is used.
  856. """
  857. train_subjects = np.asarray(train_subjects)
  858. subj_df = create_subject_table(
  859. df[df[subject_col].isin(train_subjects)],
  860. subject_col=subject_col,
  861. label_col=label_col
  862. )
  863. subjects = subj_df[subject_col].to_numpy()
  864. labels = subj_df[label_col].to_numpy()
  865. n_total = len(subjects)
  866. info = {
  867. "subsampled": False,
  868. "stratified": False,
  869. "fallback_reason": None,
  870. "n_original_train_subjects": int(n_total),
  871. "n_subsampled_train_subjects": int(n_total),
  872. }
  873. # If too few subjects, do not subsample
  874. if n_total <= 2:
  875. info["fallback_reason"] = "Too few training subjects; all subjects were used."
  876. return subjects, info
  877. # Full training set
  878. if train_ratio >= 1.0:
  879. return subjects, info
  880. # Use floor instead of ceil to avoid n_train == n_total when train_ratio < 1.0
  881. n_train = int(np.floor(n_total * train_ratio))
  882. # At least 1 subject, but not all subjects
  883. n_train = max(1, n_train)
  884. # If n_train becomes equal to or larger than n_total, use all subjects
  885. if n_train >= n_total:
  886. info["fallback_reason"] = (
  887. "Calculated subsample size was equal to the total number of training subjects; "
  888. "all subjects were used."
  889. )
  890. return subjects, info
  891. unique_labels, counts = np.unique(labels, return_counts=True)
  892. n_classes = len(unique_labels)
  893. # Stratification requires:
  894. # 1) at least two classes
  895. # 2) each class has at least two subjects
  896. # 3) selected training subset can contain all classes
  897. # 4) remaining subset can also contain all classes
  898. can_stratify = (
  899. n_classes >= 2 and
  900. np.min(counts) >= 2 and
  901. n_train >= n_classes and
  902. (n_total - n_train) >= n_classes
  903. )
  904. info["subsampled"] = True
  905. info["n_subsampled_train_subjects"] = int(n_train)
  906. rng = np.random.default_rng(random_state)
  907. try:
  908. if can_stratify:
  909. selected_subjects, _ = train_test_split(
  910. subjects,
  911. train_size=n_train,
  912. random_state=random_state,
  913. stratify=labels
  914. )
  915. info["stratified"] = True
  916. else:
  917. info["fallback_reason"] = (
  918. "Stratified training subsampling not feasible due to small class counts; "
  919. "random subject-level subsampling was used."
  920. )
  921. warnings.warn(info["fallback_reason"])
  922. selected_subjects = rng.choice(
  923. subjects,
  924. size=n_train,
  925. replace=False
  926. )
  927. except ValueError as e:
  928. info["stratified"] = False
  929. info["fallback_reason"] = (
  930. f"Stratified training subsampling failed: {str(e)}; "
  931. "random subject-level subsampling was used."
  932. )
  933. warnings.warn(info["fallback_reason"])
  934. selected_subjects = rng.choice(
  935. subjects,
  936. size=n_train,
  937. replace=False
  938. )
  939. return np.asarray(selected_subjects), info
  940. def Make_LOSO_LearningCurve_Index_bySubject(
  941. df,
  942. train_ratio=1.0,
  943. val_ratio=0.2,
  944. random_state=42,
  945. ):
  946. """
  947. Generate LOSO indices for learning curve with:
  948. - subject-level held-out testing
  949. - stratified subject-level train/validation split when feasible
  950. - repeated stratified subject-level subsampling for training set when feasible
  951. """
  952. index_col = "SampleIdx"
  953. subject_col = "LOSO_SUBJECT"
  954. subjects = sorted(df[subject_col].unique())
  955. Train_Index = []
  956. Val_Index = []
  957. Test_Index = []
  958. split_info_rows = []
  959. for fold_i, test_subject in enumerate(subjects):
  960. test_df = df[df[subject_col] == test_subject]
  961. train_val_df = df[df[subject_col] != test_subject]
  962. train_subjects, val_subjects, split_info = stratified_subject_train_val_split(
  963. train_val_df,
  964. subject_col=subject_col,
  965. label_col="Label",
  966. val_ratio=val_ratio,
  967. random_state=random_state,
  968. )
  969. selected_train_subjects, subsample_info = subsample_train_subjects_stratified(
  970. train_subjects,
  971. df=train_val_df,
  972. train_ratio=train_ratio,
  973. subject_col=subject_col,
  974. label_col="Label",
  975. random_state=random_state,
  976. )
  977. train_df = train_val_df[train_val_df[subject_col].isin(selected_train_subjects)]
  978. val_df = train_val_df[train_val_df[subject_col].isin(val_subjects)]
  979. Train_Index.append(train_df[index_col].tolist())
  980. Val_Index.append(val_df[index_col].tolist())
  981. Test_Index.append(test_df[index_col].tolist())
  982. split_info_rows.append({
  983. "fold": fold_i,
  984. "test_subject": test_subject,
  985. "train_ratio": float(train_ratio),
  986. "n_train_subjects": int(train_df[subject_col].nunique()),
  987. "n_val_subjects": int(val_df[subject_col].nunique()),
  988. "n_test_subjects": int(test_df[subject_col].nunique()),
  989. "val_split_stratified": bool(split_info.get("stratified", False)),
  990. "train_subsample_stratified": bool(subsample_info.get("stratified", False)),
  991. "val_fallback_reason": split_info.get("fallback_reason", None),
  992. "train_subsample_fallback_reason": subsample_info.get("fallback_reason", None),
  993. })
  994. return {
  995. "Train_Index": Train_Index,
  996. "Val_Index": Val_Index,
  997. "Test_Index": Test_Index,
  998. }, len(subjects), pd.DataFrame(split_info_rows)
  999. def run_learning_curve(
  1000. df,
  1001. ds_pt,
  1002. mdlinfos,
  1003. tsimg,
  1004. fsname,
  1005. col,
  1006. result_path,
  1007. title,
  1008. train_ratios=None,
  1009. n_repeats=5,
  1010. ):
  1011. if train_ratios is None:
  1012. train_ratios = LEARNING_CURVE_TRAIN_RATIOS
  1013. learning_curve_results = []
  1014. learning_curve_fold_results = []
  1015. split_info_all = []
  1016. for train_ratio in train_ratios:
  1017. for repeat_i in range(n_repeats):
  1018. seed = RANDOM_STATE + repeat_i
  1019. print(f"\n===== Learning Curve | Train Ratio: {train_ratio} | Repeat: {repeat_i+1}/{n_repeats} =====")
  1020. out_cv_ind_sample, loso_cv, split_info_df = Make_LOSO_LearningCurve_Index_bySubject(
  1021. df,
  1022. train_ratio=train_ratio,
  1023. val_ratio=VAL_RATIO,
  1024. random_state=seed,
  1025. )
  1026. split_info_df["repeat"] = repeat_i
  1027. split_info_df["random_state"] = seed
  1028. split_info_all.append(split_info_df)
  1029. # No oversampling in LOSO learning curve sensitivity analysis
  1030. if USE_OVERSAMPLING_IN_LOSO:
  1031. out_cv_ind_imb, lc_os_info_df = apply_training_fold_oversampling(
  1032. df,
  1033. out_cv_ind_sample,
  1034. subject_col="LOSO_SUBJECT",
  1035. label_col="Label",
  1036. index_col="SampleIdx",
  1037. random_state=seed,
  1038. )
  1039. else:
  1040. out_cv_ind_imb = out_cv_ind_sample
  1041. exp_lc = Exp_Detail(
  1042. (tsimg, fsname, col),
  1043. mdlinfos,
  1044. ds_pt,
  1045. out_cv_ind_imb["Train_Index"],
  1046. out_cv_ind_imb["Val_Index"],
  1047. out_cv_ind_imb["Test_Index"],
  1048. )
  1049. for mdl in exp_lc.models:
  1050. for cv_i in range(loso_cv):
  1051. mdl.CVs[cv_i] = CV_Loop(mdl.MdlInfo, mdl.CVs[cv_i])
  1052. fold_df, summary_df = collect_loso_subject_level_results(
  1053. exp_lc,
  1054. loso_cv,
  1055. df,
  1056. result_path,
  1057. title=f"LC_tmp_{title}_ratio{train_ratio}_rep{repeat_i}"
  1058. )
  1059. # Remove temporary per-repeat files generated by collect_loso_subject_level_results if desired
  1060. # For transparency, keeping them is acceptable but may produce many files.
  1061. for _, row in summary_df.iterrows():
  1062. learning_curve_results.append({
  1063. "train_ratio": float(train_ratio),
  1064. "repeat": int(repeat_i),
  1065. "random_state": int(seed),
  1066. "model": row["model"],
  1067. "subject_level_accuracy": float(row["subject_level_accuracy"]),
  1068. "segment_level_accuracy_global": float(row["segment_level_accuracy_global"]),
  1069. "segment_level_accuracy_mean_by_fold": float(row["segment_level_accuracy_mean_by_fold"]),
  1070. "n_loso_subjects": int(row["n_loso_subjects"]),
  1071. })
  1072. fold_df["train_ratio"] = float(train_ratio)
  1073. fold_df["repeat"] = int(repeat_i)
  1074. fold_df["random_state"] = int(seed)
  1075. learning_curve_fold_results.append(fold_df)
  1076. learning_curve_df = pd.DataFrame(learning_curve_results)
  1077. learning_curve_df.to_csv(
  1078. os.path.join(result_path, f"LearningCurve_LOSO_SubjectLevel_Repeated_{title}.csv"),
  1079. index=False,
  1080. encoding="utf-8-sig",
  1081. )
  1082. if len(learning_curve_fold_results) > 0:
  1083. pd.concat(learning_curve_fold_results, ignore_index=True).to_csv(
  1084. os.path.join(result_path, f"LearningCurve_LOSO_SubjectLevel_FoldResults_{title}.csv"),
  1085. index=False,
  1086. encoding="utf-8-sig",
  1087. )
  1088. if len(split_info_all) > 0:
  1089. pd.concat(split_info_all, ignore_index=True).to_csv(
  1090. os.path.join(result_path, f"LearningCurve_LOSO_SplitInfo_{title}.csv"),
  1091. index=False,
  1092. encoding="utf-8-sig",
  1093. )
  1094. # Aggregate across repeats for plotting
  1095. agg = (
  1096. learning_curve_df.groupby(["model", "train_ratio"])
  1097. .agg(
  1098. mean_accuracy=("subject_level_accuracy", "mean"),
  1099. std_accuracy=("subject_level_accuracy", "std"),
  1100. n_repeats=("subject_level_accuracy", "count"),
  1101. )
  1102. .reset_index()
  1103. )
  1104. agg["std_accuracy"] = agg["std_accuracy"].fillna(0)
  1105. agg.to_csv(
  1106. os.path.join(result_path, f"LearningCurve_LOSO_SubjectLevel_Aggregated_{title}.csv"),
  1107. index=False,
  1108. encoding="utf-8-sig",
  1109. )
  1110. # Plot with clipped asymmetric error bars to keep within [0, 1]
  1111. for model_name in agg["model"].unique():
  1112. temp = agg[agg["model"] == model_name].sort_values("train_ratio")
  1113. y = temp["mean_accuracy"].to_numpy()
  1114. sd = temp["std_accuracy"].to_numpy()
  1115. lower_err = np.minimum(sd, y - 0.0)
  1116. upper_err = np.minimum(sd, 1.0 - y)
  1117. yerr = np.vstack([lower_err, upper_err])
  1118. plt.figure(figsize=(8, 6))
  1119. plt.errorbar(
  1120. temp["train_ratio"],
  1121. y,
  1122. yerr=yerr,
  1123. marker="o",
  1124. capsize=5
  1125. )
  1126. plt.xlabel("Training data proportion")
  1127. plt.ylabel("Subject-level LOSO accuracy")
  1128. plt.ylim(-0.02, 1.02)
  1129. plt.title(f"Learning Curve - {model_name}")
  1130. plt.grid(True)
  1131. plt.savefig(
  1132. os.path.join(result_path, f"LearningCurve_LOSO_SubjectLevel_{model_name}_{title}.png"),
  1133. dpi=300,
  1134. bbox_inches="tight"
  1135. )
  1136. plt.close()
  1137. return learning_curve_df, agg
  1138. def save_loso_analysis_info(result_path, title, loso_cv, df):
  1139. """
  1140. Save metadata that clearly distinguishes LOSO sensitivity analysis from original 5-CV settings.
  1141. """
  1142. analysis_info = {
  1143. "analysis_name": "LOSO CNN sensitivity analysis",
  1144. "validation_method": "LOSO",
  1145. "n_loso_folds": int(loso_cv),
  1146. "split_level": "participant",
  1147. "test_set": "held-out participant",
  1148. "validation_split": "stratified subject-level split when feasible; unstratified fallback for very small class counts",
  1149. "analysis_folder_csv": csv_data_root_folder,
  1150. "analysis_folder_tsimg": tsimg_root_path,
  1151. "result_folder": result_path,
  1152. "sensor_or_alignment_condition": os.path.basename(os.path.normpath(csv_data_root_folder)),
  1153. "note_on_alignment": (
  1154. "Each analysis folder is assumed to contain a single predefined sensor or side-alignment condition "
  1155. "such as anatomical side, more-affected side, less-affected side, dominant side, or non-dominant side."
  1156. ),
  1157. "expected_segments_per_participant": EXPECTED_SEGMENTS_PER_PARTICIPANT,
  1158. "actual_n_subjects": int(df["LOSO_SUBJECT"].nunique()),
  1159. "actual_n_samples": int(len(df)),
  1160. "oversampling": (
  1161. "Subject-level random oversampling applied strictly within each training fold after LOSO split. "
  1162. "Validation and test folds were not oversampled."
  1163. if USE_OVERSAMPLING_IN_LOSO
  1164. else "Not applied in LOSO sensitivity analysis"
  1165. ),
  1166. "use_oversampling_in_loso": bool(USE_OVERSAMPLING_IN_LOSO),
  1167. "bootstrap_ci": bool(RUN_BOOTSTRAP_CI),
  1168. "bootstrap_unit": "held-out participant",
  1169. "learning_curve": bool(RUN_LEARNING_CURVE),
  1170. "learning_curve_train_ratios": LEARNING_CURVE_TRAIN_RATIOS,
  1171. "learning_curve_repeats": LEARNING_CURVE_REPEATS,
  1172. "subject_id_generation": "UNIQ_CLASS_SID converted to BASE_SUBJECT by removing final segment marker 1/2/3",
  1173. "append_sensor_to_loso_subject": bool(APPEND_SENSOR_TO_LOSO_SUBJECT),
  1174. "software": {
  1175. "python": f"{os.sys.version}",
  1176. "torch": torch.__version__,
  1177. "torchvision": torchvision.__version__,
  1178. "device": getattr(Info.TSImg_Exp_Info, "exp_device", "unknown"),
  1179. },
  1180. "original_info_warning": (
  1181. "Original TSImg_Exp_Info may contain exp_cv=5 or exp_imbalanced=RandomOverSampler. "
  1182. "This JSON records the actual LOSO sensitivity analysis settings."
  1183. ),
  1184. }
  1185. safe_save_json(
  1186. os.path.join(result_path, f"Analysis_Info_LOSO_{title}.json"),
  1187. analysis_info
  1188. )
  1189. return analysis_info
  1190. # =========================================
  1191. # Main
  1192. # =========================================
  1193. def main():
  1194. tsimg_methods = Info.TSImg_Exp_Info.exp_tsimg
  1195. info_fs = Info.TSImg_Exp_Info.exp_fsname
  1196. include_classes = Info.TSImg_Exp_Info.exp_include_classes
  1197. split_ratio = Info.TSImg_Exp_Info.exp_split_ratio
  1198. cv = Info.TSImg_Exp_Info.exp_cv
  1199. num_classes = Info.TSImg_Exp_Info.exp_num_classes
  1200. mdlinfos = Info.Models
  1201. # Select final CNN model for LOSO sensitivity analysis
  1202. # 기존 사용한 모델들 중 하나를 선정해서 LOSO 검증
  1203. TARGET_MODEL_NAMES = ["ResNet"] # "ResNet", "DenseNet", "SqueezeNet" 중 선택
  1204. TARGET_OPTIM_NAMES = ["Adam"] # 기존 최종 모델 optimizer 기준
  1205. def get_model_attr(model_info, attr_name):
  1206. if hasattr(model_info, attr_name):
  1207. return getattr(model_info, attr_name)
  1208. if hasattr(model_info, "MdlInfo") and hasattr(model_info.MdlInfo, attr_name):
  1209. return getattr(model_info.MdlInfo, attr_name)
  1210. return None
  1211. mdlinfos = [
  1212. m for m in mdlinfos
  1213. if get_model_attr(m, "model_name") in TARGET_MODEL_NAMES
  1214. and get_model_attr(m, "optim_name") in TARGET_OPTIM_NAMES
  1215. ]
  1216. print("Selected models for LOSO sensitivity analysis:")
  1217. for m in mdlinfos:
  1218. print(
  1219. get_model_attr(m, "model_name"),
  1220. get_model_attr(m, "optim_name"),
  1221. "batch:", get_model_attr(m, "batch_size"),
  1222. "lr:", get_model_attr(m, "lr")
  1223. )
  1224. # Save original experiment info for traceability only
  1225. original_info_path = os.path.join(
  1226. dl_result_root_path,
  1227. f"TSImg_Exp_Info_original_{GetNowString(bFileFormat=True)}.json"
  1228. )
  1229. MakeFolder(dl_result_root_path)
  1230. Save_dict2json(original_info_path, dict(Info.TSImg_Exp_Info.__dict__))
  1231. start = StartTimer()
  1232. Results = {}
  1233. Loop_for_ExcludeCalculating = []
  1234. # Make result folders and detect previously saved results
  1235. for tsimg, fsname in product(tsimg_methods, info_fs):
  1236. ColNames = info_fs[fsname][1].get_ColNames() # 0: All, 1: Select
  1237. for col in ColNames:
  1238. result_path = os.path.join(dl_result_root_path, str(tsimg), fsname, col)
  1239. MakeFolder(result_path)
  1240. exp_key = (tsimg, fsname, col)
  1241. Inter_Result_fn = os.path.join(result_path, f"Intermediate_Result_LOS_TrainFoldOS.pt")
  1242. if os.path.isfile(Inter_Result_fn):
  1243. try:
  1244. Results[exp_key] = torch.load(Inter_Result_fn)
  1245. Loop_for_ExcludeCalculating.append(exp_key)
  1246. LOGGER.info(f"Loaded existing LOSO result: {Inter_Result_fn}")
  1247. except FileNotFoundError:
  1248. LOGGER.error("There is no saved file. Continue training.")
  1249. except Exception as err:
  1250. LOGGER.error(f"Loading Intermediate Result - {err}")
  1251. # Do LOSO experiment
  1252. for tsimg, fsname in product(tsimg_methods, info_fs):
  1253. ColNames = info_fs[fsname][1].get_ColNames()
  1254. for col in ColNames:
  1255. result_path = os.path.join(dl_result_root_path, str(tsimg), fsname, col)
  1256. exp_key = (tsimg, fsname, col)
  1257. if exp_key in Loop_for_ExcludeCalculating:
  1258. LOGGER.info(f"Skip existing experiment: {exp_key}")
  1259. continue
  1260. LOGGER.info(f"{tsimg}-{fsname}-{col}")
  1261. # Load Pytorch Dataset
  1262. ds_pt = Task.Load_Dataset(tsimg_root_path, tsimg, fsname, col, include_classes)
  1263. assert num_classes == len(ds_pt.classes)
  1264. # Get dataframe information.
  1265. # This function is used only to obtain sample metadata from the existing pipeline.
  1266. # The original 5-CV indices returned by this function are not used for LOSO.
  1267. df, _ = Task.Split_Dataset_bySubject(
  1268. ds_pt,
  1269. split_ratio,
  1270. cv
  1271. )
  1272. title = f"{str(tsimg)}_{fsname}_{col}"
  1273. run_title = f"{title}_{OVERSAMPLING_TAG}"
  1274. Save_dict2json(
  1275. os.path.join(result_path, f"df_raw_{run_title}.json"),
  1276. df.to_dict()
  1277. )
  1278. # Add sensor name for traceability
  1279. df = Add_SensorName_From_Dataset(df, ds_pt)
  1280. # LOSO index generation at participant level
  1281. out_cv_ind_sample, loso_cv, df, split_info_df = Make_LOSO_Index_bySubject(
  1282. df,
  1283. sensor_name=None,
  1284. val_ratio=VAL_RATIO,
  1285. random_state=RANDOM_STATE,
  1286. append_sensor_to_subject=APPEND_SENSOR_TO_LOSO_SUBJECT,
  1287. )
  1288. Save_dict2json(
  1289. os.path.join(result_path, f"df_LOSO_{run_title}.json"),
  1290. df.to_dict()
  1291. )
  1292. # Save analysis metadata to avoid confusion with original 5-CV/RandomOverSampler settings
  1293. save_loso_analysis_info(result_path, title, loso_cv, df)
  1294. # Save subject segment count
  1295. save_subject_segment_count(
  1296. df,
  1297. result_path,
  1298. title,
  1299. expected_segments=EXPECTED_SEGMENTS_PER_PARTICIPANT,
  1300. )
  1301. # Check train/val/test subject-level overlap
  1302. overlap_df = check_no_subject_overlap(
  1303. df,
  1304. out_cv_ind_sample,
  1305. subject_col="LOSO_SUBJECT"
  1306. )
  1307. overlap_df.to_csv(
  1308. os.path.join(result_path, f"LOSO_NoSubjectOverlap_Check_{title}.csv"),
  1309. index=False,
  1310. encoding="utf-8-sig",
  1311. )
  1312. # Save fold-wise class distribution
  1313. fold_dist_df = summarize_fold_distribution(
  1314. df,
  1315. out_cv_ind_sample,
  1316. subject_col="LOSO_SUBJECT",
  1317. label_col="Label"
  1318. )
  1319. fold_dist_df.to_csv(
  1320. os.path.join(result_path, f"LOSO_Fold_ClassDistribution_{title}.csv"),
  1321. index=False,
  1322. encoding="utf-8-sig",
  1323. )
  1324. split_info_df.to_csv(
  1325. os.path.join(result_path, f"LOSO_ValSplit_Info_{title}.csv"),
  1326. index=False,
  1327. encoding="utf-8-sig",
  1328. )
  1329. Save_dict2json(
  1330. os.path.join(result_path, f"LOSO_SamplesbySub_{run_title}.json"),
  1331. out_cv_ind_sample
  1332. )
  1333. # LOSO sensitivity analysis: Train oversampling
  1334. if USE_OVERSAMPLING_IN_LOSO:
  1335. out_cv_ind_imb, os_info_df = apply_training_fold_oversampling(
  1336. df,
  1337. out_cv_ind_sample,
  1338. subject_col="LOSO_SUBJECT",
  1339. label_col="Label",
  1340. index_col="SampleIdx",
  1341. random_state=RANDOM_STATE,
  1342. )
  1343. os_info_df.to_csv(
  1344. os.path.join(result_path, f"LOSO_TrainFoldOversampling_Info_{run_title}.csv"),
  1345. index=False,
  1346. encoding="utf-8-sig",
  1347. )
  1348. os_fold_dist_df = summarize_fold_distribution_from_indices_with_duplicates(
  1349. df,
  1350. out_cv_ind_imb,
  1351. subject_col="LOSO_SUBJECT",
  1352. label_col="Label",
  1353. index_col="SampleIdx",
  1354. )
  1355. os_fold_dist_df.to_csv(
  1356. os.path.join(result_path, f"LOSO_Fold_ClassDistribution_AfterOversampling_{run_title}.csv"),
  1357. index=False,
  1358. encoding="utf-8-sig",
  1359. )
  1360. Save_dict2json(
  1361. os.path.join(result_path, f"LOSO_Samples_TrainFoldOversampled_{run_title}.json"),
  1362. out_cv_ind_imb
  1363. )
  1364. else:
  1365. out_cv_ind_imb = out_cv_ind_sample
  1366. # Create experiment
  1367. exp = Exp_Detail(
  1368. (tsimg, fsname, col),
  1369. mdlinfos,
  1370. ds_pt,
  1371. out_cv_ind_imb["Train_Index"],
  1372. out_cv_ind_imb["Val_Index"],
  1373. out_cv_ind_imb["Test_Index"],
  1374. )
  1375. # Train/evaluate LOSO folds
  1376. for mdl in exp.models:
  1377. for cv_i in range(loso_cv):
  1378. mdl.CVs[cv_i] = CV_Loop(mdl.MdlInfo, mdl.CVs[cv_i])
  1379. mdl.CVs[cv_i].SavePlot(
  1380. result_path,
  1381. title=f"{mdl.MdlInfo.model_name}-LOSO{cv_i}"
  1382. )
  1383. mdl.SavePlot(result_path, title=f"{mdl.MdlInfo.model_name}_LOSO")
  1384. # Save intermediate results
  1385. Results[exp_key] = exp
  1386. torch.save(
  1387. Results[exp_key],
  1388. os.path.join(result_path, f"Intermediate_Result_LOSO_{OVERSAMPLING_TAG}.pt")
  1389. )
  1390. # Existing LibKIME result summaries/plots
  1391. Results[exp_key].SaveResults(
  1392. result_path,
  1393. title=f"LOSO_{str(tsimg)}_{fsname}_{col}"
  1394. )
  1395. Results[exp_key].SavePlot(
  1396. result_path,
  1397. title=f"LOSO_{str(tsimg)}_{fsname}_{col}"
  1398. )
  1399. # New subject-level LOSO results
  1400. subject_fold_df, subject_summary_df = collect_loso_subject_level_results(
  1401. exp,
  1402. loso_cv,
  1403. df,
  1404. result_path,
  1405. title=run_title
  1406. )
  1407. if RUN_BOOTSTRAP_CI:
  1408. save_bootstrap_ci_from_subject_fold_results(
  1409. subject_fold_df,
  1410. result_path,
  1411. run_title
  1412. )
  1413. if RUN_LEARNING_CURVE:
  1414. run_learning_curve(
  1415. df,
  1416. ds_pt,
  1417. mdlinfos,
  1418. tsimg,
  1419. fsname,
  1420. col,
  1421. result_path,
  1422. title,
  1423. train_ratios=LEARNING_CURVE_TRAIN_RATIOS,
  1424. n_repeats=LEARNING_CURVE_REPEATS
  1425. )
  1426. # 추후 대표변수에서만 learning curve를 그리려면 아래 코드 사용
  1427. # if RUN_LEARNING_CURVE and fsname == "Acc" and col == "Gyr_X":
  1428. # run_learning_curve(
  1429. # df,
  1430. # ds_pt,
  1431. # mdlinfos,
  1432. # tsimg,
  1433. # fsname,
  1434. # col,
  1435. # result_path,
  1436. # run_title,
  1437. # train_ratios=LEARNING_CURVE_TRAIN_RATIOS,
  1438. # n_repeats=LEARNING_CURVE_REPEATS
  1439. # )
  1440. StopTimer(start)
  1441. if __name__ == "__main__":
  1442. torch.multiprocessing.freeze_support()
  1443. print("PyTorch Version:", torch.__version__)
  1444. print("Torchvision Version:", torchvision.__version__)
  1445. plt.close("all")
  1446. main()

Main_v4_Train_TSImgs_v3.py at commit 123c6a5, no license · at the source

Overview

Authors: Hyejin Choi1,2, Changhong Youm1,3, Hwayoung Park1, Bohyun Kim1, Juseon Hwang1,3, Sang-Myung Cheon4
  1. Biomechanics Laboratory, Dong-A University,37 Nakdong-Daero 550 Beon-gil, Saha-gu, Busan, 49315 Republic of Korea
  2. DAU G-LAMP Project Group, Innovation Center for Atomic Science Dong-A University,37 Nakdong-Daero 550 Beon-gil, Saha-gu, Busan, 49315 Republic of Korea
  3. Department of Health Sciences, The Graduate School of Dong-A University,37 Nakdong-Daero 550 Beon-gil, Saha-gu, Busan, 49315 Republic of Korea
  4. Department of Neurology, School of Medicine, Dong-A University,32 Daesingongwon-ro, Seo-gu, Busan, 49201 Republic of Korea
Institutions: Dong-A University (South Korea)
Journal: Scientific reports, volume 16, issue 1, article 26521
Dates: received 19 February 2026; accepted 7 July 2026; published online 10 July 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41598-026-61801-2 · PMID 42432239 · PMCID PMC13503714 · OpenAlex W7167918293
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: other (modality), human (organism), Parkinson's (population), clinical / translational (subfield)
Methods: Spectral & time-frequency, Connectivity, Statistics, Machine learning, Complexity
Keywords: Parkinson’s disease, Gait, Wearable sensors, Artificial intelligence, Deep learning, Neurodegeneration, Biomarkers, Computational biology and bioinformatics, Engineering, Health care, Neurology, Neuroscience
MeSH: Parkinson Disease*, Wearable Electronic Devices*, Aged, Convolutional Neural Networks, Early Diagnosis, Female, Gait, Humans, Machine Learning, Male, Middle Aged, Neural Networks, Computer, Walking (* major topic)
Topic: Balance, Gait, and Falls Prevention (Physical Therapy, Sports Therapy and Rehabilitation, Health Professions), according to OpenAlex
Funding: National Research Foundation of Korea (2022R1A2C100933711); Basic Science Research Program through the NRF (2022R1A6A3A0108756411); Ministry of Education of the Republic of Korea and the NRF (2024S1A5B5A16021673)
Citations: not cited yet (Europe PMC); 73 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

Zenodo 20839352

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (3 files), NumPy (3 files), pandas (2 files), PyTorch (2 files), scikit-learn (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
5 files
At the source:

hyejin-choi1/early-pd-wearable-sensor-cnn

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 123c6a541999a115cf5c2c275df38dbca3f1f75f, 25 June 2026
Languages: Python (4)
Size: 11 files, 4 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, environment (Deep Learning_CNN/requirements_conda.txt)
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: Matplotlib (3 files), NumPy (3 files), pandas (2 files), PyTorch (2 files), scikit-learn (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
5 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it points to the authors' code: Zenodo 20839352

Read it in the paper: doi.org/10.1038/s41598-026-61801-2.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 8 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it says that the data are available on request

Read it in the paper: doi.org/10.1038/s41598-026-61801-2.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 12 keywords, 13 MeSH terms, 3 funders, 66 references.

Cite

This paper

Choi, H., Youm, C., Park, H., Kim, B., Hwang, J., & Cheon, S.-M. (2026). Detection of early-stage Parkinson's disease using wearable sensors at multiple body locations and convolutional neural networks. Scientific reports, 16(1), 26521. https://doi.org/10.1038/s41598-026-61801-2

BibTeX

@article{choi2026detection,
author = {Choi, Hyejin and Youm, Changhong and Park, Hwayoung and Kim, Bohyun and Hwang, Juseon and Cheon, Sang-Myung},
title = {{Detection of early-stage Parkinson's disease using wearable sensors at multiple body locations and convolutional neural networks}},
journal = {Scientific reports},
year = {2026},
month = jul,
volume = {16},
number = {1},
pages = {26521},
publisher = {Nature Publishing Group},
issn = {2045-2322},
doi = {10.1038/s41598-026-61801-2},
url = {https://doi.org/10.1038/s41598-026-61801-2},
pmid = {42432239},
pmcid = {PMC13503714}
}

RIS

TY - JOUR
AU - Choi, Hyejin
AU - Youm, Changhong
AU - Park, Hwayoung
AU - Kim, Bohyun
AU - Hwang, Juseon
AU - Cheon, Sang-Myung
TI - Detection of early-stage Parkinson's disease using wearable sensors at multiple body locations and convolutional neural networks
T2 - Scientific reports
J2 - Sci Rep
PY - 2026
DA - 2026/07/10
VL - 16
IS - 1
SP - 26521
SN - 2045-2322
PB - Nature Publishing Group
DO - 10.1038/s41598-026-61801-2
UR - https://doi.org/10.1038/s41598-026-61801-2
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41598-026-61801-2",
"type": "article-journal",
"title": "Detection of early-stage Parkinson's disease using wearable sensors at multiple body locations and convolutional neural networks",
"container-title": "Scientific reports",
"author": [
{
"family": "Choi",
"given": "Hyejin"
},
{
"family": "Youm",
"given": "Changhong"
},
{
"family": "Park",
"given": "Hwayoung"
},
{
"family": "Kim",
"given": "Bohyun"
},
{
"family": "Hwang",
"given": "Juseon"
},
{
"family": "Cheon",
"given": "Sang-Myung"
}
],
"container-title-short": "Sci Rep",
"volume": "16",
"issue": "1",
"page": "26521",
"DOI": "10.1038/s41598-026-61801-2",
"PMID": "42432239",
"PMCID": "PMC13503714",
"ISSN": "2045-2322",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41598-026-61801-2",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s43856-026-01606-6 [code]
Validation of remote multimodal AI screening for Parkinson disease across diverse settings.
Journal: Communications medicine
In common: PyTorch, scikit-learn, pandas, 2 other tools, Parkinson's, clinical / translational, 1 reference
[2] doi:10.1162/imag.a.1241 [code]
Multispectral 7 Tesla MRI as a potential predictor of dopamine transporter deficiency in Parkinson's disease.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: PyTorch, scikit-learn, pandas, 2 other tools, Parkinson's, 1 reference
[3] doi:10.1038/s41598-026-55584-9 [code]
Reproducible benchmark of wavelet-enhanced intrabody communication biometric identification.
Journal: Scientific reports
In common: PyTorch, scikit-learn, pandas, 2 other tools, other, 1 reference
[4] doi:10.1002/mco2.70980 [code]
An Intraoperative EEG Biomarker for Postoperative Delirium Predicting Based on Interpretable Deep Learning Framework.
Journal: MedComm
In common: PyTorch, scikit-learn, pandas, 2 other tools, clinical / translational, 1 reference
[5] doi:10.1038/s41598-026-47814-x [code]
Predicting post-stroke functional outcome using explainable machine learning and integrated data.
Journal: Scientific reports
In common: PyTorch, scikit-learn, pandas, 2 other tools, clinical / translational, 1 reference
[6] doi:10.1038/s41746-026-02651-0 [code]
AI Augmented Confocal Laser Endomicroscopy for Rapid Intraoperative Diagnosis of Brain Tumors.
Journal: NPJ digital medicine
In common: PyTorch, pandas, Matplotlib, 1 other tool, clinical / translational, 1 reference
[7] doi:10.1002/advs.202600020 [code]
Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: scikit-learn, pandas, Matplotlib, 1 other tool, 2 references
[8] doi:10.1371/journal.pone.0351872 [code]
Decoding visual object recognition from EEG signals.
Journal: PloS one
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[9] doi:10.1038/s41467-026-71831-z [code]
Dysfunction of the episodic memory network in the Alzheimer's disease cascade.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[10] doi:10.1371/journal.pone.0346575 [code]
Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.
Journal: PloS one
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.