OSCR

Validation of remote multimodal AI screening for Parkinson disease across diverse settings.

Code ↔ Paper

18 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 18 matches
  1. [1] § Methods › Statistics and reproducibility ↔ code/analyses/statistical_analysis.ipynb, lines 471–519 · score 0.73 · logistic regression, Benjamini Hochberg, variables, FDR, race, age
  2. [2] § Methods › Dataset splits ↔ code/analyses/figure_1_dataset_details.ipynb, lines 394–516 · score 0.71 · ROOSTER PD, ROUTE PD, Cluster PD, Parktest, training, model
  3. [3] § Methods › Model selection and performance reporting ↔ code/unimodal_models/quick_brown_fox/unimodal_fox_wavlm_baal.py, lines 428–496 · score 0.69 · min max, PD samples, minority oversampling, configurations, SMOTE, shallow
  4. [4] § Methods › Model selection and performance reporting ↔ code/unimodal_models/facial_expression_smile/unimodal_smile_baal.py, lines 388–455 · score 0.69 · min max, PD samples, minority oversampling, configurations, SMOTE, shallow
  5. [5] § Methods › Model training ↔ code/fusion_models/ufnet/UFNet_withhold_predictions.py, lines 447–495 · score 0.68 · withhold predictions, ReLU, layer normalization, linear, networks, dropout
  6. [6] § Methods › Model training ↔ code/fusion_models/ufnet/YoutubePD/uncertainty_aware_fusion_no_drop_youtubePD.py, lines 438–486 · score 0.65 · ReLU, layer normalization, uncertainty aware, linear, networks, dropout
  7. [7] § Methods › Dataset splits ↔ code/performance_analysis/bias_analysis.py, lines 199–258 · score 0.64 · SuperPD, InMotion, Cluster PD
  8. [8] § Methods › Model selection and performance reporting ↔ code/fusion_models/ufnet/analysis/draw_final_plots.py, lines 125–176 · score 0.63 · Expected Calibration Error, Brier Score, Negative Predictive, Positive Predictive, F1 score, Curve
  9. [9] § Methods › Statistics and reproducibility ↔ code/analyses/statistical_analysis.ipynb, lines 471–519 · score 0.62 · Benjamini Hochberg, FDR correction, coefficient, predicted, model
  10. [10] § Methods › Explainability analysis with SHAP ↔ code/analyses/figure_10_interpretability_shap.ipynb, lines 192–236 · score 0.62 · feature block, features predictions, flat, modality, unimodal, matrix
  11. [11] § Results › Study participants ↔ code/analyses/figure_1_dataset_details.ipynb, lines 394–516 · score 0.61 · ROOSTER PD, ROUTE PD, Cluster PD, train, models
  12. [12] § Methods › Extracting computational features ↔ code/feature_extraction_pipeline/facial_expression_smile/smile_feature_extraction.py, lines 146–230 · score 0.57 · MediaPipe, facial features, facial expressions, eye, Smile, videos
  13. [13] § Methods › Video quality analysis ↔ code/analyses/figure_5_9_clinical_comparison_error_analysis.ipynb, lines 590–695 · score 0.57 · finger tapping videos, poor quality, visibility, positioning, smile
  14. [14] § Results › Consistency with clinician evaluation ↔ code/analyses/statistical_analysis.ipynb, lines 549–619 · score 0.55 · Mann Whitney, FDR adjusted, clinicians, bootstrapped, PPV, sensitivity
  15. [15] § Methods › Extracting computational features ↔ code/feature_extraction_pipeline/finger_tapping/feature_extraction.py, lines 611–723 · score 0.52 · MediaPipe, interruptions, Finger tapping, speed, amplitude
  16. [16] § Results › Predictive performance ↔ code/analyses/figures_2_3_4_table_1_predictive_performance.ipynb, lines 1406–1482 · score 0.52 · predictive uncertainty, Fair, bins, 1.5 %, NPV, PPV
  17. [17] § Results › Investigation of model errors ↔ code/analyses/figure_5_9_clinical_comparison_error_analysis.ipynb, lines 418–526 · score 0.51 · finger tapping model, speech model, incorrect, errors, smile, accuracy
  18. [18] § Results › Investigation of model errors ↔ code/analyses/figure_9_c_d_error_analyses.ipynb, lines 82–108 · score 0.50 · unexplained error, PD symptoms, video quality, instructions

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 634 lines · 18 KB · MIT · 3 matches

  1. # %%
  2. import pandas as pd
  3. import numpy as np
  4. from sklearn.metrics import roc_auc_score, confusion_matrix, accuracy_score
  5. from sklearn.utils import resample
  6. # %%
  7. df_test_results = pd.read_csv('../../data/test_data_big.csv')
  8. df_test_results['pred_park'] = (df_test_results['pred_score_fusion'] >= 0.5).astype(int)
  9. df_test_results.head()
  10. # %%
  11. import pandas as pd
  12. import numpy as np
  13. from sklearn.metrics import roc_auc_score, accuracy_score, confusion_matrix
  14. from sklearn.utils import resample
  15. def calculate_metrics_bootstrapped(true_labels, pred_labels, n_bootstraps=100):
  16. rng = np.random.RandomState(seed=42)
  17. # Convert to numpy arrays
  18. true_labels = np.array(true_labels)
  19. pred_labels = np.array(pred_labels)
  20. # Lists to store bootstrap results
  21. boot_auroc, boot_acc, boot_ppv, boot_npv, boot_sens, boot_spec, boot_f1 = [], [], [], [], [], [], []
  22. boot_count = 0
  23. while boot_count < n_bootstraps:
  24. # Sample indices with replacement
  25. indices = rng.choice(len(true_labels), size=len(true_labels), replace=True)
  26. # Sample labels
  27. y_true = true_labels[indices]
  28. y_pred = pred_labels[indices]
  29. try:
  30. auroc = roc_auc_score(y_true, y_pred)
  31. except ValueError:
  32. continue # Retry this bootstrap
  33. acc = accuracy_score(y_true, y_pred)
  34. try:
  35. tn, fp, fn, tp = confusion_matrix(y_true, y_pred, labels=[0, 1]).ravel()
  36. except ValueError:
  37. continue # Retry this bootstrap
  38. # Compute metrics safely
  39. ppv = tp / (tp + fp) if (tp + fp) > 0 else np.nan
  40. npv = tn / (tn + fn) if (tn + fn) > 0 else np.nan
  41. sensitivity = tp / (tp + fn) if (tp + fn) > 0 else np.nan
  42. specificity = tn / (tn + fp) if (tn + fp) > 0 else np.nan
  43. f1_score = 2 * (ppv * sensitivity) / (ppv + sensitivity) if (ppv + sensitivity) > 0 else np.nan
  44. if np.isnan([ppv, npv, sensitivity, specificity, f1_score]).any():
  45. continue # Retry this bootstrap
  46. # Append to result lists
  47. boot_auroc.append(auroc)
  48. boot_acc.append(acc)
  49. boot_ppv.append(ppv)
  50. boot_npv.append(npv)
  51. boot_sens.append(sensitivity)
  52. boot_spec.append(specificity)
  53. boot_f1.append(f1_score)
  54. boot_count += 1 # only increment if the sample is valid
  55. # Aggregate results
  56. bootstrap_results = {
  57. 'AUROC': boot_auroc,
  58. 'Accuracy': boot_acc,
  59. 'PPV': boot_ppv,
  60. 'NPV': boot_npv,
  61. 'Sensitivity': boot_sens,
  62. 'Specificity': boot_spec,
  63. 'F1 Score': boot_f1
  64. }
  65. # Summary with mean ± 95% CI
  66. summary = {}
  67. for metric, values in bootstrap_results.items():
  68. mean_val = np.nanmean(values)
  69. lower_ci = np.nanpercentile(values, 2.5)
  70. upper_ci = np.nanpercentile(values, 97.5)
  71. margin = (upper_ci - lower_ci) / 2
  72. summary[metric] = f"{round(mean_val * 100, 1)} ± {round(margin * 100, 1)}"
  73. summary_df = pd.DataFrame.from_dict(summary, orient='index', columns=['Mean ± 95% CI'])
  74. return summary_df, bootstrap_results
  75. # %%
  76. df_test_results['test_split'].value_counts()
  77. # %%
  78. df_global = df_test_results[df_test_results['test_split'] == 'global']
  79. true_labels_global = np.asarray(df_global['true_label'])
  80. pred_labels_global = np.asarray(df_global['pred_park'])
  81. summary_df_global, bootstrap_results_global = calculate_metrics_bootstrapped(true_labels_global, pred_labels_global)
  82. summary_df_global
  83. # %%
  84. df_val_1 = df_test_results[df_test_results['test_split'] == 'validation_1']
  85. true_labels_val_1 = np.asarray(df_val_1['true_label'])
  86. pred_labels_val_1 = np.asarray(df_val_1['pred_park'])
  87. summary_df_val_1, bootstrap_results_val_1 = calculate_metrics_bootstrapped(true_labels_val_1, pred_labels_val_1)
  88. summary_df_val_1
  89. # %%
  90. df_val_2 = df_test_results[df_test_results['test_split'] == 'validation_2']
  91. true_labels_val_2 = np.asarray(df_val_2['true_label'])
  92. pred_labels_val_2 = np.asarray(df_val_2['pred_park'])
  93. summary_df_val_2, bootstrap_results_val_2 = calculate_metrics_bootstrapped(true_labels_val_2, pred_labels_val_2)
  94. summary_df_val_2
  95. # %%
  96. from scipy.stats import mannwhitneyu
  97. # from statsmodels.stats.multitest import multipletests
  98. import itertools
  99. # Define the distributions (re-using group_a, group_b, group_c as stand-ins)
  100. distribution_sets = [
  101. bootstrap_results_global,
  102. bootstrap_results_val_1,
  103. bootstrap_results_val_2
  104. ]
  105. labels = ['Balanced Test Set', 'Validation Study 1', 'Validation Study 2']
  106. metric_name = 'PPV'
  107. # Perform pairwise Mann-Whitney U tests
  108. results = []
  109. for (i, j) in itertools.combinations(range(len(distribution_sets)), 2):
  110. group1 = distribution_sets[i][metric_name]
  111. group2 = distribution_sets[j][metric_name]
  112. label1 = labels[i]
  113. label2 = labels[j]
  114. stat, p = mannwhitneyu(group1, group2, alternative='two-sided')
  115. results.append({
  116. 'Comparison': f'{label1} vs {label2}',
  117. 'U Statistic': stat,
  118. 'p-value': p
  119. })
  120. # Convert results to a DataFrame
  121. results_df = pd.DataFrame(results)
  122. results_df
  123. # %%
  124. from scipy.stats import mannwhitneyu
  125. # Define the distributions (re-using group_a, group_b, group_c as stand-ins)
  126. distribution_sets = [
  127. bootstrap_results_global,
  128. bootstrap_results_val_1,
  129. bootstrap_results_val_2
  130. ]
  131. labels = ['Balanced Test Set', 'Validation Study 1', 'Validation Study 2']
  132. metric_name = 'Sensitivity'
  133. # Perform pairwise Mann-Whitney U tests
  134. results = []
  135. for (i, j) in itertools.combinations(range(len(distribution_sets)), 2):
  136. group1 = distribution_sets[i][metric_name]
  137. group2 = distribution_sets[j][metric_name]
  138. label1 = labels[i]
  139. label2 = labels[j]
  140. stat, p = mannwhitneyu(group1, group2, alternative='two-sided')
  141. results.append({
  142. 'Comparison': f'{label1} vs {label2}',
  143. 'U Statistic': stat,
  144. 'p-value': p
  145. })
  146. # Convert results to a DataFrame
  147. results_df = pd.DataFrame(results)
  148. results_df
  149. # %%
  150. df = df_test_results.copy()
  151. # %%
  152. import numpy as np
  153. import pandas as pd
  154. from scipy.stats import chi2_contingency
  155. np.random.seed(42)
  156. # Observed counts (correct, incorrect) per PD stage
  157. data = np.array([
  158. [11, 2], # Stage 1
  159. [43, 4], # Stage 2
  160. [8, 3] # Stage 3
  161. ])
  162. # Compute observed chi-square statistic
  163. observed_stat, _, _, _ = chi2_contingency(data, correction=False)
  164. print(observed_stat)
  165. # Flatten into array of 1s and 0s
  166. flat = np.concatenate([[1]*c + [0]*i for c, i in data])
  167. group_sizes = data.sum(axis=1)
  168. # Monte Carlo simulation
  169. n_sim = 10000
  170. simulated_stats = []
  171. for _ in range(n_sim):
  172. shuffled = np.random.permutation(flat)
  173. reshaped = []
  174. start = 0
  175. for size in group_sizes:
  176. group = shuffled[start:start+size]
  177. correct = (group == 1).sum()
  178. incorrect = size - correct
  179. reshaped.append([correct, incorrect])
  180. start += size
  181. reshaped = np.array(reshaped)
  182. stat, _, _, _ = chi2_contingency(reshaped, correction=False)
  183. simulated_stats.append(stat)
  184. simulated_stats = np.array(simulated_stats)
  185. p_mc = (simulated_stats >= observed_stat).mean()
  186. observed_stat, p_mc
  187. # %%
  188. summary = {'gender': [{'group': 'Female',
  189. 'mean': 0.22,
  190. 'se': 0.03,
  191. 'ci_lower': 0.16,
  192. 'ci_upper': 0.29,
  193. 'n_total': 171,
  194. 'n_wrong': 38},
  195. {'group': 'Male',
  196. 'mean': 0.17,
  197. 'se': 0.03,
  198. 'ci_lower': 0.11,
  199. 'ci_upper': 0.23,
  200. 'n_total': 147,
  201. 'n_wrong': 25},
  202. {'group': 'Unknown',
  203. 'mean': 0.0,
  204. 'se': 0.0,
  205. 'ci_lower': 0.0,
  206. 'ci_upper': 0.0,
  207. 'n_total': 2,
  208. 'n_wrong': 0},
  209. {'group': 'all',
  210. 'mean': 0.2,
  211. 'se': 0.02,
  212. 'ci_lower': 0.15,
  213. 'ci_upper': 0.24,
  214. 'n_total': 320,
  215. 'n_wrong': 63}],
  216. 'age_group': [{'group': '60 - 79',
  217. 'mean': 0.2,
  218. 'se': 0.03,
  219. 'ci_lower': 0.14,
  220. 'ci_upper': 0.25,
  221. 'n_total': 194,
  222. 'n_wrong': 39},
  223. {'group': '40 - 59',
  224. 'mean': 0.18,
  225. 'se': 0.04,
  226. 'ci_lower': 0.1,
  227. 'ci_upper': 0.26,
  228. 'n_total': 85,
  229. 'n_wrong': 15},
  230. {'group': 'Not Mentioned',
  231. 'mean': 0.0,
  232. 'se': 0.0,
  233. 'ci_lower': 0.0,
  234. 'ci_upper': 0.0,
  235. 'n_total': 3,
  236. 'n_wrong': 0},
  237. {'group': '>= 80',
  238. 'mean': 0.0,
  239. 'se': 0.0,
  240. 'ci_lower': 0.0,
  241. 'ci_upper': 0.0,
  242. 'n_total': 7,
  243. 'n_wrong': 0},
  244. {'group': '20 - 39',
  245. 'mean': 0.33,
  246. 'se': 0.09,
  247. 'ci_lower': 0.16,
  248. 'ci_upper': 0.5,
  249. 'n_total': 27,
  250. 'n_wrong': 9},
  251. {'group': '< 20',
  252. 'mean': 0.0,
  253. 'se': 0.0,
  254. 'ci_lower': 0.0,
  255. 'ci_upper': 0.0,
  256. 'n_total': 4,
  257. 'n_wrong': 0},
  258. {'group': 'all',
  259. 'mean': 0.2,
  260. 'se': 0.02,
  261. 'ci_lower': 0.15,
  262. 'ci_upper': 0.24,
  263. 'n_total': 320,
  264. 'n_wrong': 63}],
  265. 'race': [{'group': 'white',
  266. 'mean': 0.18,
  267. 'se': 0.03,
  268. 'ci_lower': 0.13,
  269. 'ci_upper': 0.23,
  270. 'n_total': 226,
  271. 'n_wrong': 40},
  272. {'group': 'Non-white',
  273. 'mean': 0.24,
  274. 'se': 0.06,
  275. 'ci_lower': 0.12,
  276. 'ci_upper': 0.35,
  277. 'n_total': 59,
  278. 'n_wrong': 14},
  279. {'group': 'Unknown',
  280. 'mean': 0.26,
  281. 'se': 0.07,
  282. 'ci_lower': 0.12,
  283. 'ci_upper': 0.4,
  284. 'n_total': 35,
  285. 'n_wrong': 9},
  286. {'group': 'all',
  287. 'mean': 0.2,
  288. 'se': 0.02,
  289. 'ci_lower': 0.15,
  290. 'ci_upper': 0.24,
  291. 'n_total': 320,
  292. 'n_wrong': 63}]}
  293. # %%
  294. from statsmodels.stats.proportion import proportions_ztest
  295. # Extract gender data for male and female only
  296. gender_data = summary['gender']
  297. female = next(g for g in gender_data if g['group'] == 'Female')
  298. male = next(g for g in gender_data if g['group'] == 'Male')
  299. # Gather counts
  300. counts = [female['n_wrong'], male['n_wrong']]
  301. totals = [female['n_total'], male['n_total']]
  302. props = [counts[i] / totals[i] for i in range(2)]
  303. # Check if sample size conditions are met for z-test
  304. conditions_met = all([
  305. totals[i] * props[i] >= 5 and totals[i] * (1 - props[i]) >= 5
  306. for i in range(2)
  307. ])
  308. # Perform test if valid
  309. if conditions_met:
  310. stat, pval = proportions_ztest(counts, totals)
  311. gender_result = {
  312. "test": "z-test for two proportions",
  313. "z_statistic": round(stat, 4),
  314. "p_value": round(pval, 4),
  315. "female_error_rate": round(props[0], 3),
  316. "male_error_rate": round(props[1], 3),
  317. "sample_size_conditions_met": True
  318. }
  319. else:
  320. result = {
  321. "error": "Sample size requirements not met for z-test",
  322. "sample_size_conditions_met": False
  323. }
  324. gender_result
  325. # %%
  326. # Extract race data for White and Non-White
  327. race_data = summary['race']
  328. white = next(r for r in race_data if r['group'].lower() == 'white')
  329. non_white = next(r for r in race_data if r['group'].lower() == 'non-white')
  330. # Gather counts
  331. counts = [white['n_wrong'], non_white['n_wrong']]
  332. totals = [white['n_total'], non_white['n_total']]
  333. props = [counts[i] / totals[i] for i in range(2)]
  334. # Check sample size conditions
  335. conditions_met_race = all([
  336. totals[i] * props[i] >= 5 and totals[i] * (1 - props[i]) >= 5
  337. for i in range(2)
  338. ])
  339. # Perform test if valid
  340. if conditions_met_race:
  341. stat, pval = proportions_ztest(counts, totals)
  342. race_result = {
  343. "test": "z-test for two proportions",
  344. "z_statistic": round(stat, 4),
  345. "p_value": round(pval, 4),
  346. "white_error_rate": round(props[0], 3),
  347. "non_white_error_rate": round(props[1], 3),
  348. "sample_size_conditions_met": True
  349. }
  350. else:
  351. race_result = {
  352. "error": "Sample size requirements not met for z-test",
  353. "sample_size_conditions_met": False
  354. }
  355. race_result
  356. # %%
  357. # Extract age group data and filter valid ones
  358. age_data = summary['age_group']
  359. valid_groups = [g for g in age_data if g['group'] in ['20 - 39', '40 - 59', '60 - 79']]
  360. # Create contingency table: [correct, incorrect] for each group
  361. age_contingency = []
  362. age_labels = []
  363. for group in valid_groups:
  364. correct = group['n_total'] - group['n_wrong']
  365. incorrect = group['n_wrong']
  366. age_contingency.append([correct, incorrect])
  367. age_labels.append(group['group'])
  368. # Run chi-square test of independence
  369. chi2_stat, p_val, dof, expected = chi2_contingency(age_contingency)
  370. # Create expected frequencies DataFrame and check for < 5
  371. expected_df = pd.DataFrame(expected, columns=['Correct_exp', 'Incorrect_exp'], index=age_labels)
  372. expected_check = (expected_df < 5).any(axis=1)
  373. any_cell_under_5 = expected_check.any()
  374. # Format result
  375. age_result = {
  376. "test": "Chi-square test of independence",
  377. "chi2_statistic": round(chi2_stat, 4),
  378. "p_value": round(p_val, 4),
  379. "degrees_of_freedom": dof,
  380. "age_groups_compared": age_labels,
  381. "all_expected_freqs_>=5": not any_cell_under_5
  382. }
  383. age_result
  384. # %%
  385. from statsmodels.stats.multitest import multipletests
  386. # Raw p-values from your tests
  387. pvals = [gender_result['p_value'], race_result['p_value'], age_result['p_value']]
  388. labels = ['Gender', 'Race', 'Age Group']
  389. # Apply FDR correction (Benjamini-Hochberg)
  390. reject, pvals_corrected, _, _ = multipletests(pvals, alpha=0.05, method='fdr_bh')
  391. # Display results
  392. for i in range(len(pvals)):
  393. print(f"{labels[i]}: raw p = {pvals[i]:.4f}, FDR-corrected p = {pvals_corrected[i]:.4f}, significant = {reject[i]}")
  394. # %%
  395. import pandas as pd
  396. import statsmodels.api as sm
  397. from statsmodels.stats.multitest import multipletests
  398. df['race_'] = df['race'].replace({
  399. 'White': 'white',
  400. 'Black or African American': 'Non-white',
  401. 'American Indian or Alaska Native': 'Non-white',
  402. 'Asian': 'Non-white',
  403. 'Others': 'Non-white',
  404. 'Not Mentioned': 'Unknown'
  405. })
  406. # Convert misclassified column to binary (1 = misclassified, 0 = correct)
  407. df['error'] = df['misclassified_fusion'].astype(int)
  408. # Filter rows to include only the desired levels
  409. df_filtered = df[
  410. df['gender'].isin(['Female', 'Male']) &
  411. df['race_'].isin(['white', 'Non-white']) &
  412. df['age_group'].isin(['20 - 39', '40 - 59', '60 - 79'])
  413. ].copy()
  414. # One-hot encode categorical predictors (drop_first avoids multicollinearity)
  415. X = pd.get_dummies(df_filtered[['gender', 'race_', 'age_group']], drop_first=True)
  416. X = sm.add_constant(X)
  417. # Ensure all X columns are numeric
  418. X = X.astype(float)
  419. y = df_filtered['error']
  420. # Fit logistic regression model
  421. model = sm.Logit(y, X).fit()
  422. pvals = model.pvalues
  423. # Apply Benjamini-Hochberg FDR correction
  424. reject, pvals_corrected, _, _ = multipletests(pvals, method='fdr_bh')
  425. # Display results in a table
  426. results = pd.DataFrame({
  427. 'Variable': X.columns,
  428. 'Coefficient': model.params.round(4),
  429. 'Raw p-value': pvals.round(4),
  430. 'FDR-adjusted p': pvals_corrected.round(4),
  431. 'Significant (FDR < 0.05)': reject
  432. })
  433. results
  434. # %%
  435. from statsmodels.stats.contingency_tables import cochrans_q
  436. df_clinician_ratings = df[~df['neurologist_label_ray'].isna()]
  437. # %%
  438. ray_preds = np.asarray(df_clinician_ratings['neurologist_label_ray'])
  439. ruth_preds = np.asarray(df_clinician_ratings['neurologist_label_ruth'])
  440. jamie_preds = np.asarray(df_clinician_ratings['neurologist_label_jamie'])
  441. group_prediction = (ray_preds + ruth_preds + jamie_preds >= 2).astype(int)
  442. # Then compare boots_park['PPV'] to boots_group['PPV']
  443. df_clinician_ratings['pred_park'] = (df_clinician_ratings['pred_score_fusion'] >= 0.5).astype(int)
  444. park_preds = np.asarray(df_clinician_ratings['pred_park'])
  445. true_labels = np.asarray(df_clinician_ratings['true_label'])
  446. summary_ray, boots_ray = calculate_metrics_bootstrapped(true_labels, ray_preds)
  447. summary_ruth, boots_ruth = calculate_metrics_bootstrapped(true_labels, ruth_preds)
  448. summary_jamie, boots_jamie = calculate_metrics_bootstrapped(true_labels, jamie_preds)
  449. summary_park, boots_park = calculate_metrics_bootstrapped(true_labels, park_preds)
  450. summary_group, boots_group = calculate_metrics_bootstrapped(true_labels, group_prediction)
  451. # %%
  452. from scipy.stats import mannwhitneyu
  453. from statsmodels.stats.multitest import multipletests
  454. def compare_park_to_clinicians(boots_park, boots_clinicians: dict, metrics: list):
  455. """
  456. Compare PARK's metrics to clinicians using one-sided Mann-Whitney U test (PARK < Clinician).
  457. Args:
  458. boots_park: dict of bootstrapped metrics for PARK
  459. boots_clinicians: dict of dicts (e.g., {'Ray': boots_ray, ...})
  460. metrics: list of metric names (e.g., ['PPV', 'Sensitivity'])
  461. Returns:
  462. A dictionary of test results and a printed summary message.
  463. """
  464. results = {}
  465. significantly_lower_metrics = []
  466. for metric in metrics:
  467. pvals = []
  468. pairs = []
  469. # inside the for metric loop
  470. for name, boots in boots_clinicians.items():
  471. dist_park = np.array(boots_park[metric])
  472. dist_clinician = np.array(boots[metric])
  473. # Check for valid data
  474. if np.all(np.isnan(dist_park)) or np.all(np.isnan(dist_clinician)):
  475. pvals.append(np.nan)
  476. else:
  477. try:
  478. stat = mannwhitneyu(dist_park, dist_clinician, alternative='less')
  479. pvals.append(stat.pvalue)
  480. except ValueError:
  481. pvals.append(np.nan)
  482. pairs.append(f"PARK vs {name}")
  483. # FDR correction
  484. if all(np.isfinite(pvals)):
  485. reject, pvals_fdr, _, _ = multipletests(pvals, method='fdr_bh')
  486. else:
  487. reject = [False] * len(pvals)
  488. pvals_fdr = [np.nan] * len(pvals)
  489. # Store result
  490. results[metric] = {
  491. 'pairs': pairs,
  492. 'raw_pvals': pvals,
  493. 'fdr_pvals': pvals_fdr,
  494. 'reject': reject
  495. }
  496. # Track significance
  497. if any(reject):
  498. significantly_lower_metrics.append(metric)
  499. # Print detailed result per metric
  500. print(f"\nMetric: {metric}")
  501. for i, pair in enumerate(pairs):
  502. print(f"{pair}: raw p = {pvals[i]:.4f}, FDR-adjusted p = {pvals_fdr[i]:.4f}, significant = {reject[i]}")
  503. # Summary message
  504. if significantly_lower_metrics:
  505. print(f"\n⚠️ PARK showed significantly lower performance (FDR < 0.05) than at least one clinician for: {', '.join(significantly_lower_metrics)}")
  506. else:
  507. print("\n✅ PARK did not show significantly lower performance than any clinician for any metric (FDR-adjusted p > 0.05).")
  508. return results
  509. # %%
  510. # Define the clinicians and the metrics to compare
  511. clinicians_boots = {
  512. 'Ray': boots_ray,
  513. 'Ruth': boots_ruth,
  514. 'Jamie': boots_jamie,
  515. 'Group': boots_group
  516. }
  517. metrics_to_compare = ['Accuracy', 'Specificity', 'Sensitivity', 'PPV', 'NPV', 'F1 Score']
  518. # Run the function
  519. results = compare_park_to_clinicians(boots_park, clinicians_boots, metrics_to_compare)

statistical_analysis.ipynb at commit 187b906, under MIT · at the source

Overview

Authors: Md Saiful Islam1,2, Tariq Adnan1,2, Abdelrahman Abdelkader1, Zipei Liu1, Evelyn Ma1, Sooyong Park1, Asif Azad2,3, Pai Liu1, Meghan Pawlik4, Emily Hartman4, Erin Shelton5, Kristina B Larson6, M Saifur Rahman2, Cathe Schwartz5, Karen Jaffe5, Jamie L Adams4, Ruth B Schneider4, Jan Freyberg7, E Ray Dorsey3,4, Ehsan Hoque1,8
  1. University of Rochester, Rochester, NY USA
  2. Bangladesh University of Engineering & Technology, Dhaka, Bangladesh
  3. Atria Health and Research Institute, New York, NY USA
  4. University of Rochester Medical Center, Rochester, NY USA
  5. InMotion, Beachwood, OH USA
  6. Harvard Medical School, Boston, MA USA
  7. Google Research, London, UK
  8. Ministry of Defense, Riyadh, Saudi Arabia
Journal: Communications medicine, volume 6, issue 1, article 488
Dates: received 15 October 2025; accepted 10 April 2026; published online 6 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s43856-026-01606-6 · PMID 42092025 · PMCID PMC13586151 · OpenAlex W7160408048
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), Parkinson's (population), clinical / translational (subfield)
Methods: Connectivity, Statistics, Machine learning, Physiology & signal measures
Keywords: Parkinson's disease
Topic: Voice and Speech Disorders (Physiology, Medicine), according to OpenAlex
Funding: NINDS NIH HHS (P50 NS108676)
Citations: not cited yet (Europe PMC); 63 references in the paper

Abstract

Background: Timely detection of Parkinson’s disease (PD) remains limited by reliance on in-person neurological evaluations that are often costly and geographically inaccessible. To address these barriers, we develop PARK (Parkinson’s Analysis with Remote Kinetic-tasks) – a web-based artificial intelligence (AI) tool that screens for PD using short webcam recordings of facial expression, motor, and speech tasks.

Methods: Across eight independent studies (n = 1,865 participants; 670 with PD), participants completed three standardized tasks (smile mimicry, finger tapping, and pangram utterance) via webcam. Task-specific neural networks estimate PD risk and uncertainty, which are integrated through an uncertainty-calibrated fusion model (UFNet). Model performance is evaluated on one internal and two external test sets representing supervised and unsupervised real-world environments. Three movement disorder specialists also reviewed videos from 30 participants to benchmark clinical agreement of the PARK tool. User experience is assessed through structured surveys containing open-ended or multiple-choice questions.

Results: PARK achieves accuracies of 80.2–80.6% and AUROC of 0.85-0.87 across all evaluation cohorts, with 83.3–86.5% sensitivity and 71.2–78.4% specificity. Predictive performance remains stable across sex, age, and ethnicity. Agreement with clinician judgments reaches Cohen’s κ = 0.59. Uncertainty estimates reflect diagnostic confidence, and performance declines at high-uncertainty levels. Usability is rated highly (System Usability Scale > 70) in both supervised and unsupervised settings, with low perceived risk and strong user preference for remote screening.

Conclusions: PARK demonstrates promising accuracy and favorable user acceptance for remote PD screening, highlighting its potential as an accessible, equitable, and uncertainty-aware tool for neurological assessment when traditional care is challenging to obtain.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 18 matches between paragraphs and lines of code.

ROC-HCI/UFNet

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 5ece2c65ba184faccf6c8cdccdc03132427c464b, 30 December 2024
Languages: Python (50)
Size: 188 files, 50 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (requirements.txt)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: pandas (42 files), NumPy (41 files), PyTorch (38 files), imbalanced-learn (36 files), Matplotlib (36 files), scikit-learn (36 files), SciPy (36 files), seaborn (33 files), Plotly (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
52 files

Zenodo 18940836

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: pandas (42 files), NumPy (41 files), PyTorch (38 files), imbalanced-learn (36 files), Matplotlib (36 files), scikit-learn (36 files), SciPy (36 files), seaborn (33 files), Plotly (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
  • 28 September 2026: the link answers (HTTP 200)
52 files
At the source:

saiful1105020/park_communications_medicine

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 187b90651e80d313169690f18d716d7466b15b0c, 11 March 2026
Languages: Python (16), Jupyter (9)
Size: 145 files, 25 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (environment.yml), 9 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (20 files), pandas (19 files), scikit-learn (12 files), Matplotlib (11 files), SciPy (8 files), PyTorch (7 files), seaborn (7 files), imbalanced-learn (4 files), OpenCV (2 files), statsmodels (2 files), Numba (1 file), Plotly (1 file), SHAP (1 file), SymPy (1 file), Hugging Face Transformers (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
27 files

Zenodo 18940702

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
  • 28 September 2026: the link answers (HTTP 200)
At the source:

Code availability

All custom code used in this study, including scripts for video processing, feature extraction, model training and evaluation, statistical analyses, and figure generation, is publicly available via the GitHub repository: https://github.com/saiful1105020/park_communications_medicine(https://doi.org/10.5281/zenodo.18940702)63. All data processing and analyses were performed using Python. The repository includes instructions for setting up the Python environment and the exact package versions required to reproduce the experiments and analyses reported in this study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 125 scripts, each with its path and the digest of its content;
  • 18 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

The recorded videos are collected using a web-based tool. The tool is publicly accessible at https://parktest.net26. The datasets used in this study consist of participant videos and associated clinical information collected under institutional review board (IRB) approval across multiple studies. Because these data contain sensitive health information and identifiable video recordings, they cannot be made publicly available due to IRB restrictions and compliance with the Health Insurance Portability and Accountability Act (HIPAA). Access to the data may be considered for collaborative research purposes and will be subject to approval under the relevant IRB protocols and a data use agreement. Researchers interested in accessing the data for collaboration may contact the principal investigator, Ehsan Hoque, to discuss potential data-sharing arrangements that are consistent with ethical and regulatory guidelines. Such correspondence will be evaluated on a case-by-case basis, and we will aim to respond within 5 working days. All source data underlying the figures and tables in this study are publicly available through a repository on Figshare at https://doi.org/10.60593/ur.d.2941070361. The repository includes the numerical data used to generate the plots and summary statistics reported in the manuscript, along with documentation describing each file and its associated data columns. The UFNet dataset and preprocessing code used in this study are available at the GitHub repository https://github.com/ROC-HCI/UFNet (https://doi.org/10.5281/zenodo.18940836)62.

All custom code used in this study, including scripts for video processing, feature extraction, model training and evaluation, statistical analyses, and figure generation, is publicly available via the GitHub repository: https://github.com/saiful1105020/park_communications_medicine(https://doi.org/10.5281/zenodo.18940702)63. All data processing and analyses were performed using Python. The repository includes instructions for setting up the Python environment and the exact package versions required to reproduce the experiments and analyses reported in this study.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 20 authors, 1 keyword, 1 funder, 42 references.

Cite

This paper

Islam, M. S., Adnan, T., Abdelkader, A., Liu, Z., Ma, E., Park, S., Azad, A., Liu, P., Pawlik, M., Hartman, E., Shelton, E., Larson, K. B., Rahman, M. S., Schwartz, C., Jaffe, K., Adams, J. L., Schneider, R. B., Freyberg, J., Dorsey, E. R., & Hoque, E. (2026). Validation of remote multimodal AI screening for Parkinson disease across diverse settings. Communications medicine, 6(1), 488. https://doi.org/10.1038/s43856-026-01606-6

BibTeX

@article{islam2026validation,
author = {Islam, Md Saiful and Adnan, Tariq and Abdelkader, Abdelrahman and Liu, Zipei and Ma, Evelyn and Park, Sooyong and Azad, Asif and Liu, Pai and Pawlik, Meghan and Hartman, Emily and Shelton, Erin and Larson, Kristina B and Rahman, M Saifur and Schwartz, Cathe and Jaffe, Karen and Adams, Jamie L and Schneider, Ruth B and Freyberg, Jan and Dorsey, E Ray and Hoque, Ehsan},
title = {{Validation of remote multimodal AI screening for Parkinson disease across diverse settings}},
journal = {Communications medicine},
year = {2026},
month = may,
volume = {6},
number = {1},
pages = {488},
publisher = {Nature Publishing Group},
issn = {2730-664X},
doi = {10.1038/s43856-026-01606-6},
url = {https://doi.org/10.1038/s43856-026-01606-6},
pmid = {42092025},
pmcid = {PMC13586151}
}

RIS

TY - JOUR
AU - Islam, Md Saiful
AU - Adnan, Tariq
AU - Abdelkader, Abdelrahman
AU - Liu, Zipei
AU - Ma, Evelyn
AU - Park, Sooyong
AU - Azad, Asif
AU - Liu, Pai
AU - Pawlik, Meghan
AU - Hartman, Emily
AU - Shelton, Erin
AU - Larson, Kristina B
AU - Rahman, M Saifur
AU - Schwartz, Cathe
AU - Jaffe, Karen
AU - Adams, Jamie L
AU - Schneider, Ruth B
AU - Freyberg, Jan
AU - Dorsey, E Ray
AU - Hoque, Ehsan
TI - Validation of remote multimodal AI screening for Parkinson disease across diverse settings
T2 - Communications medicine
J2 - Commun Med (Lond)
PY - 2026
DA - 2026/05/06
VL - 6
IS - 1
SP - 488
SN - 2730-664X
PB - Nature Publishing Group
DO - 10.1038/s43856-026-01606-6
UR - https://doi.org/10.1038/s43856-026-01606-6
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s43856-026-01606-6",
"type": "article-journal",
"title": "Validation of remote multimodal AI screening for Parkinson disease across diverse settings",
"container-title": "Communications medicine",
"author": [
{
"family": "Islam",
"given": "Md Saiful"
},
{
"family": "Adnan",
"given": "Tariq"
},
{
"family": "Abdelkader",
"given": "Abdelrahman"
},
{
"family": "Liu",
"given": "Zipei"
},
{
"family": "Ma",
"given": "Evelyn"
},
{
"family": "Park",
"given": "Sooyong"
},
{
"family": "Azad",
"given": "Asif"
},
{
"family": "Liu",
"given": "Pai"
},
{
"family": "Pawlik",
"given": "Meghan"
},
{
"family": "Hartman",
"given": "Emily"
},
{
"family": "Shelton",
"given": "Erin"
},
{
"family": "Larson",
"given": "Kristina B"
},
{
"family": "Rahman",
"given": "M Saifur"
},
{
"family": "Schwartz",
"given": "Cathe"
},
{
"family": "Jaffe",
"given": "Karen"
},
{
"family": "Adams",
"given": "Jamie L"
},
{
"family": "Schneider",
"given": "Ruth B"
},
{
"family": "Freyberg",
"given": "Jan"
},
{
"family": "Dorsey",
"given": "E Ray"
},
{
"family": "Hoque",
"given": "Ehsan"
}
],
"container-title-short": "Commun Med (Lond)",
"volume": "6",
"issue": "1",
"page": "488",
"DOI": "10.1038/s43856-026-01606-6",
"PMID": "42092025",
"PMCID": "PMC13586151",
"ISSN": "2730-664X",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s43856-026-01606-6",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
6
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: imbalanced-learn, SHAP, Hugging Face Transformers, 11 other tools
[2] doi:10.1016/j.isci.2026.116825 [code]
Social hierarchy shapes behavioral and transcriptional responses to chronic stress and ketamine in male mice.
Journal: iScience
In common: imbalanced-learn, SHAP, Numba, 10 other tools
[3] doi:10.1186/s13059-026-04125-8 [code]
MLMarker: a machine learning framework for tissue inference and biomarker discovery.
Journal: Genome biology
In common: imbalanced-learn, SHAP, Plotly, 8 other tools
[4] doi:10.1016/j.isci.2026.115329 [code]
Brain metastases converge on shared geometric architecture and transcriptomic landscape yet remain distinct from gliomas.
Journal: iScience
In common: imbalanced-learn, SHAP, OpenCV, 7 other tools, clinical / translational
[5] doi:10.1371/journal.pone.0348866 [code]
Using deep learning to identify inherited retinal diseases based on wide-field retinal imaging data.
Journal: PloS one
In common: imbalanced-learn, Plotly, OpenCV, 7 other tools, clinical / translational
[6] doi:10.1038/s41467-026-76837-1 [code]
Drug screen and machine learning predict neuroprotective agents in a preclinical human model of childhood dementia.
Journal: Nature communications
In common: SHAP, Hugging Face Transformers, OpenCV, 7 other tools
[7] doi:10.1093/nargab/lqag050 [code]
TSProm: deep learning framework to predict tissue-specific regulatory logic.
Journal: NAR genomics and bioinformatics
In common: SHAP, Hugging Face Transformers, statsmodels, 7 other tools
[8] doi:10.1371/journal.pone.0346575 [code]
Statistically valid explainable black-box machine learning: applications in sex classification across species using brain imaging.
Journal: PloS one
In common: Hugging Face Transformers, Numba, Plotly, 7 other tools
[9] doi:10.7554/elife.109717 [code]
Retrosplenial cortex enables context-dependent goal-directed sensorimotor transformation.
Journal: eLife
In common: SHAP, Numba, OpenCV, 7 other tools
[10] doi:10.1038/s41467-026-74358-5 [code]
Brain-inspired spatial intelligence for embodied agents.
Journal: Nature communications
In common: Hugging Face Transformers, Numba, OpenCV, 7 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.