OSCR

Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity.

Code ↔ Paper

1 match between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 1 match
  1. [1] § Experimental Section › Machine Learning Prediction ↔ uchl1_analysis.ipynb, lines 221–359 · score 0.66 · random forest, SHAP, instantiated, fraction, iteration, absolute

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 556 lines · 19 KB · no license · 1 match

  1. # %% [markdown]
  2. # # Retinal Ubiquitin C-Terminal Hydrolase L1 Dysregulation Converges with Synaptic Vulnerability and Predicts Alzheimer’s Disease Severity
  3. # %% [markdown]
  4. # This notebook contains the ML components for the manuscript. The data loading portion is for a specific dataset; the details can be swapped out to load your own dataset.
  5. # %%
  6. data_file = "retinal_alzheimers.xlsx" # path to the data file
  7. required_features = ["Age (yrs)"] # features that must be defined to retain a subject
  8. features_to_encode = ["Sex", "Ethnicity"] # categorical features that we want to use for prediction
  9. features_to_drop = ["Column1", "Column2", "Column3", "APOE genotype", "APOE4 presence"]
  10. stratification_features = ["Clinical Diagnosis", "Sex"]
  11. savepath = "results/" # path where to save figures/data
  12. num_bars = 7 # number of features to plot in the feature identification
  13. # %%
  14. # %%
  15. import pandas as pd
  16. import numpy as np
  17. import sklearn as sk
  18. from sklearn.preprocessing import OneHotEncoder
  19. import anndata as ad
  20. from matplotlib import pyplot as plt
  21. from ast import literal_eval
  22. import scipy
  23. from sklearn.linear_model import Lasso
  24. from sklearn.ensemble import RandomForestRegressor
  25. import seaborn, shap
  26. import os
  27. import sklearn
  28. from sklearn.model_selection import cross_val_score
  29. from sklearn.model_selection import train_test_split
  30. from copy import deepcopy
  31. from pathlib import Path
  32. from sklearn.utils import shuffle
  33. from pathlib import Path
  34. import pickle as pkl # for saving
  35. from typing import Optional, Sequence, Tuple
  36. if not os.path.exists(savepath):
  37. p = Path(savepath)
  38. p.mkdir(parents=True, exist_ok=True)
  39. # %% [markdown]
  40. # # Data Loading
  41. # %% [markdown]
  42. # Note that the data file is a .xlsx file and is irregularly-shaped. To simplify loading your own data, you can have a unique index for each row and have column values in the first row.
  43. # %%
  44. def encode_apoe(df: pd.DataFrame,
  45. apoe_colname: str = "APOE genotype"):
  46. """
  47. Encodes APOE into a binary value depending on the presence of e4 in the string value.
  48. """
  49. if apoe_colname in df.columns:
  50. df["has_e4"] = df.apply(lambda x: type(x["APOE genotype"]) is str and "e4" in x["APOE genotype"], axis=1).astype(float)
  51. return
  52. # %%
  53. data = pd.read_excel(data_file, index_col="Pt", skiprows=2)
  54. encode_apoe(data)
  55. # Remove extraneous rows
  56. s = [bool(a) for a in data.index.isna()] # Some rows don't have an ID; data file is malformed.
  57. s = [not(a) and b != "average" for (a,b) in zip(s, data.index)] # Some rows have "average" in the ID column; data file is malformed.
  58. data = data.loc[s] # exclude rows
  59. # Some entries have things that they shouldn't; fix
  60. columns = data.columns
  61. cols_to_convert = set()
  62. # Gather the columns that are typed as strings, remove non-numeric data and attempt to convert.
  63. for i in range(data.shape[0]):
  64. for j in range(data.shape[1]):
  65. if isinstance(data.iloc[i,j], str):
  66. if data.iloc[i,j].endswith("*"): # some entries end in *
  67. data.iloc[i,j] = float(data.iloc[i,j][:-1])
  68. cols_to_convert.add(columns[j])
  69. elif data.iloc[i,j] == "na": # some entries are "na" instead of blank
  70. data.iloc[i,j] = 0
  71. cols_to_convert.add(columns[j])
  72. elif "," in data.iloc[i,j]: # some entries use "," as the floating point
  73. data.iloc[i,j] = literal_eval(data.iloc[i,j].replace(",","."))
  74. cols_to_convert.add(columns[j])
  75. for col in cols_to_convert:
  76. try:
  77. data[col] = data[col].astype(float)
  78. except Exception:
  79. continue
  80. # %%
  81. # Drop uninformative columns (blank or previously encoded)
  82. to_drop = [False for _ in range(data.shape[0])]
  83. for f in required_features:
  84. to_drop |= data[f].isna().values
  85. data = data.loc[~to_drop, :]
  86. # %%
  87. # Encode columns for one-hot encoding.
  88. for feat in features_to_encode:
  89. ohe = OneHotEncoder(sparse_output=False)
  90. recoded = ohe.fit_transform(data[[feat]])
  91. data.loc[:, ohe.get_feature_names_out()] = recoded
  92. if feat in stratification_features:
  93. stratification_features.pop(stratification_features.index(feat))
  94. stratification_features += list(ohe.get_feature_names_out())
  95. data.drop(feat, axis=1, inplace=True)
  96. # %%
  97. for feat in features_to_drop:
  98. data.drop(feat, axis=1, inplace=True)
  99. # %%
  100. cp_features = ["Retinal Cp % area", "Brain Cp % area"]
  101. retinal_biomarkers = list(data.columns[39:-7])
  102. for c in cp_features:
  103. retinal_biomarkers.pop(retinal_biomarkers.index(c)) # Retinal Cp and Brain Cp are stored in the middle of the retinal features, but we don't want to use them
  104. predicting_features = retinal_biomarkers
  105. target_features = ["Braak Stage"]
  106. # target_features = ["MMSE [score]"] # alternative target
  107. # %%
  108. data_train, data_test, label_train, label_test = train_test_split(data[predicting_features],
  109. data[target_features[0]],
  110. train_size=0.8,
  111. stratify=data[stratification_features].fillna(0),
  112. random_state=0)
  113. train_subjects = set(data_train.index)
  114. # %%
  115. # %%
  116. model_list = [RandomForestRegressor]
  117. model_names = {RandomForestRegressor: "RandomForest"}
  118. # %%
  119. def make_learning_curve(model, data_train, label_train, train_sizes=np.linspace(0.1,1,10), savepath=None):
  120. ret = sklearn.model_selection.learning_curve(model, data_train, label_train, train_sizes=train_sizes)
  121. plt.plot(ret[0], ret[2], ".-")
  122. plt.xlabel("Data samples")
  123. plt.ylabel("Score")
  124. # plt.title(f"Learning curve: {label_train.columns[0]}")
  125. plt.title(f"Learning curve: {label_train.name}")
  126. plt.savefig(savepath)
  127. plt.close()
  128. return
  129. def get_permutation_scores(model, data_train, label_train, savepath, n_permutations=200) -> (float, float):
  130. score, perm_scores, pvalue = sklearn.model_selection.permutation_test_score(model, data_train, label_train, n_permutations=200)
  131. f = open(savepath, "w")
  132. f.write(f"Score: {score}\n")
  133. f.write(f"p value: {pvalue}")
  134. f.close()
  135. return score, pvalue
  136. def plot_shap(model, data_train, target_name, savepath=None, num_kmeans=10, title=""):
  137. data_train_summary = shap.kmeans(data_train, num_kmeans)
  138. explainer = shap.KernelExplainer(model.predict, data_train_summary)
  139. shap_values = explainer.shap_values(data_train)
  140. shap.summary_plot(shap_values, data_train, title=f"Features predictive of {target_name}", show=False)
  141. plt.title(f"Features predictive of {target_name}" + f"{title}")
  142. if savepath is not None:
  143. plt.savefig(savepath, bbox_inches="tight")
  144. plt.close()
  145. return shap_values
  146. # %%
  147. model_parameters = {RandomForestRegressor: {"n_estimators":80, "criterion":"absolute_error"}}
  148. # %%
  149. for t in target_features:
  150. print("---------------")
  151. print(f"Target: {t}")
  152. data2 = deepcopy(data)
  153. to_drop = data2[t].isna().values
  154. data2 = data2.loc[~to_drop, :]
  155. data2.fillna(0, inplace=True)
  156. ssubj = set(data2.index)
  157. train_subj = list(train_subjects.intersection(ssubj))
  158. # val_subj = list(val_subjects.intersection(ssubj))
  159. pred_features2 = deepcopy(predicting_features)
  160. if t in pred_features2:
  161. pred_features2.pop(pred_features2.index(t))
  162. # data_train = data2.loc[train_subj, predicting_features]
  163. data_train = data2.loc[train_subj, pred_features2]
  164. label_train = data2.loc[train_subj, t]
  165. data_train, label_train = shuffle(data_train, label_train, random_state=0)
  166. if data_train.shape[0] < 20:
  167. print(f"too few subjects for measure {t}")
  168. continue
  169. for m in model_list:
  170. model = m(**model_parameters[m], random_state=0)
  171. model.fit(data_train, label_train)
  172. outpath = f"{savepath}/{model_names[m]}/{t.replace('/','div')}/"
  173. print(f"Model: {model_names[m]}")
  174. Path(outpath).mkdir(exist_ok=True, parents=True)
  175. make_learning_curve(model, data_train, label_train, savepath=f"{outpath}/learning_curve.png")
  176. score = cross_val_score(model, data_train, label_train)
  177. with open(f"{savepath}/permutation_score.txt", "w") as f:
  178. f.write(f"Score: {np.mean(score)}\n")
  179. print(f"Cross val. scores: {score}")
  180. print(f"Mean score: {np.mean(score)}")
  181. shap_values = plot_shap(model, data_train, t, savepath=f"{outpath}/shaps_{model.n_estimators}_{model.criterion}.png")
  182. print("Done")
  183. # %%
  184. import numpy as np
  185. import pandas as pd
  186. from typing import Callable, Dict, Any, Optional, Union
  187. import shap
  188. from collections import defaultdict as dd
  189. def feature_importance_stability(
  190. model_callable: Callable,
  191. X_train: pd.DataFrame,
  192. y_train: pd.Series,
  193. X_shap: Optional[pd.DataFrame] = None,
  194. params: Optional[Dict[str, Any]] = None,
  195. n_iterations: int = 10,
  196. n_top: Union[int, list] = 10,
  197. shap_kwargs: Optional[Dict[str, Any]] = None
  198. ) -> Dict[str, float]:
  199. """
  200. Train a model multiple times and compute the stability of top features based on SHAP values.
  201. Parameters
  202. ----------
  203. model_callable : Callable
  204. A callable that returns a new model instance when called with **params.
  205. Example: lambda **kwargs: RandomForestClassifier(**kwargs)
  206. X_train : pd.DataFrame
  207. Training features (n_samples, n_features)
  208. y_train : pd.Series
  209. Training labels (n_samples,)
  210. X_shap : pd.DataFrame, optional
  211. Data to use for SHAP value computation. If None, uses X_train.
  212. params : dict, optional
  213. Dictionary of parameters to pass to model_callable
  214. n_iterations : int, default=10
  215. Number of times to instantiate, train, and evaluate the model
  216. n_top : int, default=10
  217. Number of top features to track
  218. shap_kwargs : dict, optional
  219. Additional arguments to pass to shap.Explainer
  220. Returns
  221. -------
  222. dict
  223. Dictionary with feature names as keys and fraction of runs (0-1)
  224. where the feature appeared in top n as values.
  225. shap_value_dict
  226. SHAP values for each feature.
  227. """
  228. if params is None:
  229. params = {}
  230. if shap_kwargs is None:
  231. shap_kwargs = {}
  232. if X_shap is None:
  233. X_shap = X_train
  234. if isinstance(n_top, int):
  235. n_top = [n_top]
  236. # Extract feature names from DataFrame
  237. feature_names = list(X_train.columns)
  238. n_features = len(feature_names)
  239. # Track which features appear in top n for each iteration
  240. top_n_feature_dict = {}
  241. shap_value_dict = dd(list)
  242. for n in n_top:
  243. top_n_feature_dict[n] = dd(int)
  244. r_state = 0 # control initial random state
  245. for iteration in range(n_iterations):
  246. # Instantiate and train model
  247. model = model_callable(**params, random_state=r_state)
  248. r_state += 1 # ensure that we have a different state for the next iteration
  249. model.fit(X_train, y_train)
  250. # Compute SHAP values
  251. explainer = shap.Explainer(model, **shap_kwargs)
  252. shap_values = explainer(X_shap)
  253. # Get mean absolute SHAP values for each feature
  254. if hasattr(shap_values, 'values'):
  255. # For newer SHAP versions that return Explanation objects
  256. shap_array = shap_values.values
  257. else:
  258. shap_array = shap_values
  259. # Handle multi-class case (take mean across classes if needed)
  260. if len(shap_array.shape) == 3:
  261. mean_abs_shap = np.mean(np.abs(shap_array), axis=(0, 2))
  262. else:
  263. mean_abs_shap = np.mean(np.abs(shap_array), axis=0)
  264. # Get top n feature indices
  265. for idx in range(mean_abs_shap.shape[0]):
  266. shap_value_dict[feature_names[idx]].append(mean_abs_shap[idx]) # get SHAP values for each feature
  267. for n in n_top:
  268. top_n_indices = _select_top_n(mean_abs_shap, min(n, n_features))
  269. # Update counts for top features
  270. for idx in top_n_indices:
  271. top_n_feature_dict[n][feature_names[idx]] += 1
  272. shap_value_dict[feature_names[idx]].append(mean_abs_shap[idx])
  273. if iteration % 100 == 0:
  274. print(f"Iteration {iteration + 1}/{n_iterations} completed")
  275. print(f"Complete")
  276. # Convert counts to fractions
  277. for n in n_top:
  278. feature_fractions = {
  279. name: count / n_iterations
  280. for name, count in top_n_feature_dict[n].items()
  281. }
  282. # Sort by fraction (descending)
  283. top_n_feature_dict[n] = dict(
  284. sorted(feature_fractions.items(), key=lambda x: x[1], reverse=True)
  285. )
  286. return top_n_feature_dict, shap_value_dict
  287. def _select_top_n(scores: np.ndarray, n_top: int):
  288. """Select indices of top n scores."""
  289. n_from = scores.shape[0]
  290. reference_indices = np.arange(n_from, dtype=int)
  291. partition = np.argpartition(scores, -n_top)[-n_top:]
  292. partial_indices = np.argsort(scores[partition])[::-1]
  293. global_indices = reference_indices[partition][partial_indices]
  294. return global_indices
  295. def extract_nonzero(d):
  296. nd = {}
  297. for k, v in d.items():
  298. if v > 0:
  299. nd[k] = v
  300. return nd
  301. # %%
  302. # Get feature importance stability for a model.
  303. model = model_list[0]
  304. rfr_params = {"n_estimators": 80, "criterion": "absolute_error"}
  305. top_feats, shap_value_dict = feature_importance_stability(model,
  306. X_train=data_train,
  307. y_train=label_train,
  308. params=rfr_params,
  309. n_iterations=1000,
  310. n_top=[5,10])
  311. # %%
  312. top_5_feats = extract_nonzero(top_feats[5])
  313. top_10_feats = extract_nonzero(top_feats[10])
  314. # %%
  315. top_5_feats
  316. # %%
  317. top_10_feats
  318. # %%
  319. idx = 0
  320. top_feat_path = f"{savepath}/{target_features[0].replace(' ', '_')}_top_feats_{idx}.pkl"
  321. to_save = [top_feats, shap_value_dict]
  322. while os.path.exists(top_feat_path):
  323. idx += 1
  324. top_feat_path = f"{savepath}/{target_features[0].replace(' ', '_')}_top_feats_{idx}.pkl"
  325. print(f"Output file: {top_feat_path}")
  326. with open(top_feat_path, "wb") as f:
  327. pkl.dump(to_save, f)
  328. # %%
  329. plot_dict = dict()
  330. feats_list = []
  331. shap_frac = []
  332. shap_val = []
  333. for k in top_feats[5].keys():
  334. plot_dict[k] = np.zeros(2)
  335. plot_dict[k][0] = top_feats[5][k]
  336. plot_dict[k][1] = np.mean(shap_value_dict[k])
  337. feats_list.append(k)
  338. shap_frac.append(plot_dict[k][0])
  339. shap_val.append(plot_dict[k][1])
  340. # %%
  341. def plot_split_bars(
  342. features: Sequence[str],
  343. left_values: Sequence[float],
  344. right_values: Sequence[float],
  345. left_label: str = "Left",
  346. right_label: str = "Right",
  347. left_color: str = "#1f77b4",
  348. right_color: str = "#ff7f0e",
  349. figsize: Tuple[float, float] = (10, 6),
  350. ax: Optional[plt.Axes] = None
  351. ) -> plt.Axes:
  352. """
  353. Create a split horizontal bar plot with two values for each feature.
  354. Both sides are scaled from 0 (at center) to their maximum values,
  355. with each half occupying exactly 50% of the plot width.
  356. Parameters
  357. ----------
  358. features : sequence of str
  359. Names of the features (y-axis labels)
  360. left_values : sequence of float
  361. Values for the left side bars (assumed to start from 0)
  362. right_values : sequence of float
  363. Values for the right side bars (assumed to start from 0)
  364. left_label : str, default="Left"
  365. Label for the left bars in the legend
  366. right_label : str, default="Right"
  367. Label for the right bars in the legend
  368. left_color : str, default="#1f77b4"
  369. Color for the left bars
  370. right_color : str, default="#ff7f0e"
  371. Color for the right bars
  372. figsize : tuple of float, default=(10, 6)
  373. Figure size (width, height) in inches
  374. ax : matplotlib.axes.Axes, optional
  375. Axes to plot on. If None, creates a new figure.
  376. Returns
  377. -------
  378. ax : matplotlib.axes.Axes
  379. The axes object with the plot
  380. Examples
  381. --------
  382. >>> features = ['Feature A', 'Feature B', 'Feature C', 'Feature D']
  383. >>> left_vals = [0.8, 0.6, 0.9, 0.4]
  384. >>> right_vals = [120, 85, 150, 95]
  385. >>> ax = plot_split_bars(features, left_vals, right_vals,
  386. ... left_label="Score", right_label="Count")
  387. >>> plt.show()
  388. """
  389. # Convert to numpy arrays
  390. left_values = np.array(left_values)
  391. right_values = np.array(right_values)
  392. if len(features) != len(left_values) or len(features) != len(right_values):
  393. raise ValueError("features, left_values, and right_values must have the same length")
  394. # Create figure if no axes provided
  395. if ax is None:
  396. fig, ax = plt.subplots(figsize=figsize)
  397. # Get maximum for each side (assuming min is 0)
  398. left_max = np.max(left_values)
  399. right_max = np.max(right_values)
  400. # Avoid division by zero
  401. left_max = left_max if left_max != 0 else 1
  402. right_max = right_max if right_max != 0 else 1
  403. # Normalize each side to [0, 1] based on its own max
  404. left_normalized = left_values / left_max
  405. right_normalized = right_values / right_max
  406. # Y positions for bars
  407. y_pos = np.arange(len(features))
  408. # Plot left bars (extend from 0 to -1, scaled by normalized values)
  409. ax.barh(y_pos, -left_normalized, height=0.8,
  410. color=left_color, label=left_label, alpha=0.8)
  411. # Plot right bars (extend from 0 to 1, scaled by normalized values)
  412. ax.barh(y_pos, right_normalized, height=0.8,
  413. color=right_color, label=right_label, alpha=0.8)
  414. # Set y-axis labels
  415. ax.set_yticks(y_pos)
  416. ax.set_yticklabels(features)
  417. # Add vertical line at center
  418. ax.axvline(0, color='black', linewidth=1.5, linestyle='-', alpha=0.3)
  419. # Set x-axis limits to ensure each half occupies exactly 50%
  420. ax.set_xlim(-1, 1)
  421. # Create custom x-tick labels showing actual data values
  422. # Place ticks at key positions: left edge, left middle, center, right middle, right edge
  423. left_ticks = [-1, -0.5, 0]
  424. right_ticks = [0.5, 1]
  425. # Calculate corresponding actual values
  426. left_labels = [
  427. f'{left_max:.2g}',
  428. f'{0.5 * left_max:.2g}',
  429. '0'
  430. ]
  431. right_labels = [
  432. f'{0.5 * right_max:.2g}',
  433. f'{right_max:.2g}'
  434. ]
  435. # Set ticks and labels
  436. ax.set_xticks(left_ticks + right_ticks)
  437. ax.set_xticklabels(left_labels + right_labels)
  438. ax.set_xlabel('Value')
  439. # Add legend
  440. ax.legend(loc='best')
  441. # Add grid for easier reading
  442. ax.grid(axis='x', alpha=0.3, linestyle='--')
  443. ax.set_axisbelow(True)
  444. # Invert y-axis so first feature is at top
  445. ax.invert_yaxis()
  446. plt.tight_layout()
  447. return ax
  448. # %%
  449. if "MMSE" in target_features[0]:
  450. target_title = "MMSE"
  451. elif "Braak" in target_features[0]:
  452. target_title = "Braak Stage"
  453. else:
  454. target_title = target_features[0]
  455. ax = plot_split_bars(feats_list[:num_bars], shap_frac[:num_bars], shap_val[:num_bars], left_label="Iteration Fraction", right_label="Mean Abs. SHAP")
  456. ax.set_title(f"Predictive Features of {target_title}")
  457. fig = ax.get_figure()
  458. fig.savefig(f"{savepath}/{target_features[0].replace(' ', '_')}_featimpact_{idx}.png", bbox_inches='tight')
  459. fig.savefig(f"{savepath}/{target_features[0].replace(' ', '_')}_featimpact_{idx}.svg", bbox_inches='tight')
  460. # %%

uchl1_analysis.ipynb at commit 23d5d9f, no license · at the source

Overview

Authors: Altan Rentsendorj1, Jean‐Philippe Vit1, Alexandre Hutton2, Yosef Koronyo1, Saba Shahin1, Edward Robinson1, Ernesto Barron3, Anthony Rodriguez3, Ousman Jallow1, Julia Sheyn4, Bhakta Prasad Gaire1, Lalita Subedi5,6, Miyah R Davis1, Barathan Gnanabharathi4,7,8,9, Alexander V Ljubimov4,8,10, Stuart L Graham11, Vivek K Gupta11, Mehdi Mirzaei11,12, Debra Hawes13, Katlin Silm4,7,8,9
and 6 other authorsKeith L Black1, Lon S Schneider14, Alfredo A Sadun3, Jesse G Meyer2,15, Dieu‐Trang Fuchs1, Maya Koronyo‐Hamaoui1,8,9
15 affiliations
  1. Department of Neurosurgery, Maxine Dunitz Neurosurgical Institute, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  2. Department of Computational Biomedicine, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  3. Doheny Eye Institute, University of California, Los Angeles, California, USA
  4. Board of Governors Regenerative Medicine Institute, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  5. Division of Infectious Diseases and Immunology, Department of Pediatrics, Guerin Children's at Cedars‐Sinai Medical Center, Los Angeles, California, USA
  6. Department of Biomedical Sciences, Infectious and Immunologic Diseases Research Center, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  7. Center For Neural Science and Medicine, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  8. Department of Biomedical Sciences, Division of Applied Cell Biology and Physiology, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  9. Department of Neurology, Cedars‐Sinai Medical Center, Los Angeles, California, USA
  10. David Geffen School of Medicine, University of California, Los Angeles, California, USA
  11. Macquarie Medical School, Faculty of Medicine, Health and Human Sciences, Macquarie University, Sydney, NSW, Australia
  12. ProHeme Diagnostics Pty Ltd, Sydney, NSW, Australia
  13. Department of Pathology, Keck School of Medicine, University of Southern California, Los Angeles, California, USA
  14. Department of Psychiatry and The Behavioral Sciences, Department of Neurology, Keck School of Medicine, University of Southern California, Los Angeles, California, USA
  15. Smidt Heart Institute, Cedars‐Sinai Medical Center, Los Angeles, California, USA
Institutions: Cedars-Sinai Medical Center (United States); University of California, Los Angeles (United States); Doheny Eye Institute (United States); Macquarie University (Australia); University of Southern California (United States); Cedars-Sinai Smidt Heart Institute (United States)
Dates: received 9 January 2026; accepted 12 August 2026; published online 29 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/advs.202600020 · PMID 42666080 · PMCID PMC13525425 · OpenAlex W7204637716
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), Alzheimer's / dementia (population), cellular / molecular (subfield)
Methods: Statistics, Machine learning, Connectivity
Keywords: dementia, excitatory synapse, MCI, N‐methyl‐D‐aspartate receptor subunit 2A, p75NTR, postsynaptic density protein 95, ubiquitin–proteasome system, vesicular glutamate transporter 1
Topic: Alzheimer's disease research and treatments (Physiology, Medicine), according to OpenAlex
Funding: the Hertz Innovation Funds; the Wilstein Foundation; NIA NIH HHS (R01 AG055865, R01 AG056478, R01AG055865, AG056478-04S1, R01AG056478); NIGMS NIH HHS (R35GM142502, R35 GM142502); the Ray Charles Foundation; the Saban Foundation; NIH HHS; the Gordon Foundation
Citations: not cited yet (Europe PMC); 188 references in the paper
Research resources: NMDAR2A pAb RRID:AB_10001400, GAPDH (D16H11) mAb RRID:AB_10622025, GAPDH mAb RRID:AB_1078991, UCHL‐1 (PGP9.5) mAb RRID:AB_10891773, IRDye 680RD pAb RRID:AB_10956166, IRDye 680RD pAb RRID:AB_10956588, PSD95 mAb RRID:AB_1310620, Bassoon mAb (SAP7F407) RRID:AB_2038857, Parvalbumin pAb RRID:AB_2173902, AT8 (Ser202/Thr205) phospho‐tau) mAb RRID:AB_223647, βIII‐tubulin (TUBB3) RRID:AB_2256751, VGLUT1 pAb RRID:AB_2301751, GFAP mAb (2.2B10) RRID:AB_2532994, 12F4 (Aβ42) mAb RRID:AB_2564682, Synaptophysin mAb RRID:AB_2565370, Calbindin D‐28K Clone CB‐955 mAb RRID:AB_476894, IBA1 pAb RRID:AB_521594, IRDye 800CW pAb RRID:AB_621842, IRDye 800CW pAb RRID:AB_621843, PKC alpha mAb RRID:AB_777294, GFAP pAb RRID:AB_829022, IBA1 pAb RRID:AB_839504, p75NGFR (p75NTR) mAb RRID:AB_881682, C57BL/6J RRID:IMSR_JAX:000664, GraphPad Prism (Version 10.0.3) RRID:SCR_002798

Abstract

Synaptic dysfunction is a major driver of cognitive decline in Alzheimer's disease (AD), yet its extent and molecular basis in the retina remain poorly defined. We integrated postmortem retinal and matched brain histopathology with ultrastructural, proteomic, biochemical, and machine‐learning analyses across cognitively normal, mild cognitive impairment, and AD cohorts. Retinal glutamatergic synapses exhibited early, progressive degeneration, marked by loss of presynaptic vesicular glutamate transporter 1 (VGLUT1) and synaptophysin and postsynaptic density protein 95 (PSD95) and N‐methyl‐D‐aspartate receptor subunit 2A (NMDAR2A), along with ribbon synapse ultrastructural disruption. Synaptic deficits correlated with amyloid‐β42 (Aβ42), pathogenic tau, oxidative stress, the Aβ‐binding p75 neurotrophin receptor, and glial activation that paralleled disease progression. Proteomics revealed widespread synaptic remodeling accompanied by disease‐associated microglia, astrocyte‐mediated excitotoxicity, and pyroptotic pathways. The synapse‐enriched deubiquitinase ubiquitin C‐terminal hydrolase L1 (UCHL1) was dysregulated early, particularly in horizontal and bipolar interneurons, and strongly associated with synaptic loss and neuroinflammation. Mechanistically, fibrillar Aβ42 induced rapid UCHL1 and synaptic depletion in human and murine neurons before overt neurodegeneration. Machine‐learning models identified retinal UCHL1 as the strongest predictor of Braak stage and cognitive impairment. These findings establish the retina as an early site of AD synaptopathy and position UCHL1 as a candidate biomarker and mechanistic mediator linking amyloid pathology, neuroinflammation, and synaptic vulnerability.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.

Xomicsdatascience/Ubiquitin_Ad

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 23d5d9fd2b84d6c372a82e9fc83e7ed48c597f20, 2 December 2025
Languages: Jupyter (1)
Size: 3 files, 1 script
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README, environment (requirements.txt), 1 notebook
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: anndata (1 file), Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), seaborn (1 file), SHAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
2 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

Mass spectrometry data are publicly available via the PRIDE Database under the ProteomeXchange accession number PXD040225. All data are available in the main text, tables, and figures or in the supplementary materials (Figures S1–S12 and Tables S1–S7). All other material or data requests should be addressed to the corresponding author. The Jupyter Notebook containing the analysis with random forests has been uploaded to https://Github.com/Xomicsdatascience/Ubiquitin_Ad. Data will become publicly available as of the date of publication.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, pages, dates, 26 authors, 8 keywords, 8 funders, 188 references, 25 RRIDs.

Cite

This paper

Rentsendorj, A., Vit, J., Hutton, A., Koronyo, Y., Shahin, S., Robinson, E., Barron, E., Rodriguez, A., Jallow, O., Sheyn, J., Gaire, B. P., Subedi, L., Davis, M. R., Gnanabharathi, B., Ljubimov, A. V., Graham, S. L., Gupta, V. K., Mirzaei, M., Hawes, D., . . . Koronyo‐Hamaoui, M. (2026). Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity. Advanced science (Weinheim, Baden-Wurttemberg, Germany), e00020. https://doi.org/10.1002/advs.202600020

BibTeX

@article{rentsendorj2026early,
author = {Rentsendorj, Altan and Vit, Jean‐Philippe and Hutton, Alexandre and Koronyo, Yosef and Shahin, Saba and Robinson, Edward and Barron, Ernesto and Rodriguez, Anthony and Jallow, Ousman and Sheyn, Julia and Gaire, Bhakta Prasad and Subedi, Lalita and Davis, Miyah R and Gnanabharathi, Barathan and Ljubimov, Alexander V and Graham, Stuart L and Gupta, Vivek K and Mirzaei, Mehdi and Hawes, Debra and Silm, Katlin and Black, Keith L and Schneider, Lon S and Sadun, Alfredo A and Meyer, Jesse G and Fuchs, Dieu‐Trang and Koronyo‐Hamaoui, Maya},
title = {{Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity}},
journal = {Advanced science (Weinheim, Baden-Wurttemberg, Germany)},
year = {2026},
month = aug,
pages = {e00020},
publisher = {Wiley},
issn = {2198-3844},
doi = {10.1002/advs.202600020},
url = {https://doi.org/10.1002/advs.202600020},
pmid = {42666080},
pmcid = {PMC13525425}
}

RIS

TY - JOUR
AU - Rentsendorj, Altan
AU - Vit, Jean‐Philippe
AU - Hutton, Alexandre
AU - Koronyo, Yosef
AU - Shahin, Saba
AU - Robinson, Edward
AU - Barron, Ernesto
AU - Rodriguez, Anthony
AU - Jallow, Ousman
AU - Sheyn, Julia
AU - Gaire, Bhakta Prasad
AU - Subedi, Lalita
AU - Davis, Miyah R
AU - Gnanabharathi, Barathan
AU - Ljubimov, Alexander V
AU - Graham, Stuart L
AU - Gupta, Vivek K
AU - Mirzaei, Mehdi
AU - Hawes, Debra
AU - Silm, Katlin
AU - Black, Keith L
AU - Schneider, Lon S
AU - Sadun, Alfredo A
AU - Meyer, Jesse G
AU - Fuchs, Dieu‐Trang
AU - Koronyo‐Hamaoui, Maya
TI - Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity
T2 - Advanced science (Weinheim, Baden-Wurttemberg, Germany)
J2 - Adv Sci (Weinh)
PY - 2026
DA - 2026/08/29
SP - e00020
SN - 2198-3844
PB - Wiley
DO - 10.1002/advs.202600020
UR - https://doi.org/10.1002/advs.202600020
LA - en
ER -

CSL-JSON

{
"id": "10.1002/advs.202600020",
"type": "article-journal",
"title": "Early Retinal UCHL1 Dysregulation Coupled With Synaptic Loss Reflects Alzheimer's Disease Severity",
"container-title": "Advanced science (Weinheim, Baden-Wurttemberg, Germany)",
"author": [
{
"family": "Rentsendorj",
"given": "Altan"
},
{
"family": "Vit",
"given": "Jean‐Philippe"
},
{
"family": "Hutton",
"given": "Alexandre"
},
{
"family": "Koronyo",
"given": "Yosef"
},
{
"family": "Shahin",
"given": "Saba"
},
{
"family": "Robinson",
"given": "Edward"
},
{
"family": "Barron",
"given": "Ernesto"
},
{
"family": "Rodriguez",
"given": "Anthony"
},
{
"family": "Jallow",
"given": "Ousman"
},
{
"family": "Sheyn",
"given": "Julia"
},
{
"family": "Gaire",
"given": "Bhakta Prasad"
},
{
"family": "Subedi",
"given": "Lalita"
},
{
"family": "Davis",
"given": "Miyah R"
},
{
"family": "Gnanabharathi",
"given": "Barathan"
},
{
"family": "Ljubimov",
"given": "Alexander V"
},
{
"family": "Graham",
"given": "Stuart L"
},
{
"family": "Gupta",
"given": "Vivek K"
},
{
"family": "Mirzaei",
"given": "Mehdi"
},
{
"family": "Hawes",
"given": "Debra"
},
{
"family": "Silm",
"given": "Katlin"
},
{
"family": "Black",
"given": "Keith L"
},
{
"family": "Schneider",
"given": "Lon S"
},
{
"family": "Sadun",
"given": "Alfredo A"
},
{
"family": "Meyer",
"given": "Jesse G"
},
{
"family": "Fuchs",
"given": "Dieu‐Trang"
},
{
"family": "Koronyo‐Hamaoui",
"given": "Maya"
}
],
"container-title-short": "Adv Sci (Weinh)",
"page": "e00020",
"DOI": "10.1002/advs.202600020",
"PMID": "42666080",
"PMCID": "PMC13525425",
"ISSN": "2198-3844",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/advs.202600020",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
29
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1111/ejn.70480 [code]
Astrocyte Proximity Protects Synapses From Human Amyloid-Beta Induced Degeneration in a Mouse Ex Vivo Model of Early Alzheimer's Disease.
Journal: The European journal of neuroscience
In common: seaborn, scikit-learn, pandas, 3 other tools, Alzheimer's / dementia, cellular / molecular, 7 references
[2] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: anndata, seaborn, scikit-learn, 4 other tools, Alzheimer's / dementia, cellular / molecular, 2 references
[3] doi:10.1162/imag.a.1174 [code]
Retinal nerve fibre layer thickness reflects characteristics of brain grey and white matter.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Matplotlib, NumPy, 5 references
[4] doi:10.1038/s41514-026-00443-0 [code]
Network-based discovery of regulatory drivers of cognitive decline in alzheimer's disease.
Journal: npj aging
In common: seaborn, scikit-learn, pandas, 3 other tools, Alzheimer's / dementia, 3 references
[5] doi:10.1038/s41593-026-02267-3 [code]
Spatial proteomic analysis in human Alzheimer's disease brains enables identification of microenvironment-dependent microglial cell states.
Journal: Nature neuroscience
In common: anndata, seaborn, scikit-learn, 4 other tools, Alzheimer's / dementia, cellular / molecular, 2 references
[6] doi:10.1002/ana.78300
Distribution of Big Tau Isoforms in the Human Central and Peripheral Nervous System.
Journal: Annals of neurology
In common: Alzheimer's / dementia, 6 references
[7] doi:10.1038/s42003-026-09923-1 [code]
Scalable discovery of spatial multicellular patterns via neighborhood-to-sequence transformation.
Journal: Communications biology
In common: anndata, scikit-learn, pandas, 2 other tools, Alzheimer's / dementia, 3 references
[8] doi:10.1038/s41467-026-76837-1 [code]
Drug screen and machine learning predict neuroprotective agents in a preclinical human model of childhood dementia.
Journal: Nature communications
In common: SHAP, seaborn, scikit-learn, 4 other tools, Alzheimer's / dementia, 1 reference
[9] doi:10.1038/s41467-026-71831-z [code]
Dysfunction of the episodic memory network in the Alzheimer's disease cascade.
Journal: Nature communications
In common: seaborn, scikit-learn, pandas, 2 other tools, Alzheimer's / dementia, 3 references
[10] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: anndata, seaborn, scikit-learn, 4 other tools, cellular / molecular, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.