OSCR

Pursuit of biomarkers of brain diseases: beyond cohort comparisons.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § The case of brain activity-based biomarkers › Brain Swap ↔ AI_model.ipynb, lines 1–114 · score 0.56 · readout layer, RNN2, trained, RNN1, loss, RNNs
  2. [2] § The case of brain activity-based biomarkers › Brain Swap ↔ AI_model.ipynb, lines 1–114 · score 0.52 · hidden state, Adam, RNN2, layer, RNN1, loss

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 299 lines · 9.8 KB · GPL-3.0 · 2 matches

  1. # %%
  2. import torch
  3. import torch.nn as nn
  4. import torch.optim as optim
  5. import numpy as np
  6. import matplotlib.pyplot as plt
  7. # Set random seeds for reproducibility
  8. seed_num = 42
  9. torch.manual_seed(seed_num)
  10. np.random.seed(seed_num)
  11. # Parameters
  12. input_size = 10
  13. hidden_size = 50
  14. output_size = 1
  15. seq_length = 20
  16. batch_size = 32
  17. n_epochs = 100
  18. learning_rate = 0.01
  19. # Generate synthetic data
  20. def generate_data(batch_size, seq_length, input_size):
  21. X = torch.randn(batch_size, seq_length, input_size)
  22. # Simple task: predict the sum of the last 5 inputs
  23. y = torch.sum(X[:, -5:, :], dim=(1, 2)).unsqueeze(1)
  24. return X, y
  25. # Define the RNN model
  26. class RNNModel(nn.Module):
  27. def __init__(self, input_size, hidden_size, output_size):
  28. super(RNNModel, self).__init__()
  29. self.rnn = nn.RNN(input_size, hidden_size, batch_first=True)
  30. self.readout = nn.Linear(hidden_size, output_size)
  31. def forward(self, x):
  32. rnn_out, _ = self.rnn(x)
  33. # Take the last hidden state
  34. last_hidden = rnn_out[:, -1, :]
  35. output = self.readout(last_hidden)
  36. return output
  37. # Create two identical RNN models
  38. rnn1 = RNNModel(input_size, hidden_size, output_size)
  39. rnn2 = RNNModel(input_size, hidden_size, output_size)
  40. # Make rnn2's RNN identical to rnn1's RNN initially
  41. rnn2.rnn.load_state_dict(rnn1.rnn.state_dict())
  42. #rnn2.bias.load_state_dict(rnn1.bias.state_dict())
  43. # Verify they have the same parameters initially
  44. print("Initial parameter comparison:")
  45. for (name1, param1), (name2, param2) in zip(rnn1.named_parameters(), rnn2.named_parameters()):
  46. print(f"{name1} and {name2} equal: {torch.allclose(param1, param2)}")
  47. # Training setup
  48. criterion = nn.MSELoss()
  49. optimizer1 = optim.Adam(rnn1.parameters(), lr=learning_rate)
  50. optimizer2 = optim.Adam(rnn2.parameters(), lr=learning_rate)
  51. # Training loop
  52. for epoch in range(n_epochs):
  53. X, y = generate_data(batch_size, seq_length, input_size)
  54. # Train RNN1
  55. optimizer1.zero_grad()
  56. output1 = rnn1(X)
  57. loss1 = criterion(output1, y)
  58. loss1.backward()
  59. optimizer1.step()
  60. # Train RNN2
  61. optimizer2.zero_grad()
  62. output2 = rnn2(X)
  63. loss2 = criterion(output2, y)
  64. loss2.backward()
  65. optimizer2.step()
  66. if epoch % 10 == 0:
  67. print(f"Epoch {epoch}, Loss1: {loss1.item():.4f}, Loss2: {loss2.item():.4f}")
  68. # After training, compare the readout layers
  69. print("\nAfter training parameter comparison:")
  70. for (name1, param1), (name2, param2) in zip(rnn1.named_parameters(), rnn2.named_parameters()):
  71. print(f"{name1} and {name2} equal: {torch.allclose(param1, param2, atol=1e-4)}")
  72. # Test the models
  73. X_test, y_test = generate_data(batch_size, seq_length, input_size)
  74. with torch.no_grad():
  75. output1 = rnn1(X_test)
  76. output2 = rnn2(X_test)
  77. print(f"\nTest MSE RNN1: {criterion(output1, y_test).item():.4f}")
  78. print(f"Test MSE RNN2: {criterion(output2, y_test).item():.4f}")
  79. # Now drive the second readout with the first RNN's activity
  80. with torch.no_grad():
  81. rnn1_activity, _ = rnn1.rnn(X_test)
  82. rnn1_last_hidden = rnn1_activity[:, -1, :]
  83. output2_driven = rnn2.readout(rnn1_last_hidden)
  84. print(f"\nTest MSE when driving rnn2's readout with rnn1's activity: {criterion(output2_driven, y_test).item():.4f}")
  85. # Plot results
  86. plt.figure(figsize=(6, 3))
  87. plt.plot(y_test.numpy(), label='True')
  88. plt.plot(output1.numpy(), 'o', label='Model 1')
  89. plt.plot(output2.numpy(), 'o', label='Model 2')
  90. plt.plot(output2_driven.numpy(), 'x', label='Model 3')
  91. plt.legend(ncol=4,loc='upper left')
  92. plt.ylim((-15,24))
  93. #plt.title("Comparison of Outputs")
  94. plt.xlabel("Input ID")
  95. plt.ylabel("Output")
  96. plt.show()
  97. # %%
  98. import torch
  99. import torch.nn as nn
  100. import torch.optim as optim
  101. import numpy as np
  102. import matplotlib.pyplot as plt
  103. # Set random seeds for reproducibility
  104. seed_num = 0
  105. torch.manual_seed(seed_num)
  106. np.random.seed(seed_num)
  107. # Parameters
  108. input_size = 3 # Reduced for easier visualization
  109. seq_length = 20 # Reduced for easier visualization
  110. batch_size = 32
  111. n_epochs = 100
  112. learning_rate = 0.01
  113. # Generate synthetic data
  114. def generate_data(batch_size, seq_length, input_size):
  115. X = torch.randn(batch_size, seq_length, input_size)
  116. # Simple task: predict the sum of the last 5 inputs
  117. y = torch.sum(X[:, -5:, :], dim=(1, 2)).unsqueeze(1)
  118. return X, y
  119. # 1. Let's generate some sample data just for visualization
  120. X_vis, y_vis = generate_data(batch_size=3, seq_length=seq_length, input_size=input_size)
  121. # 2. Create a plot
  122. fig, axes = plt.subplots(3, 1, figsize=(3, 6))
  123. # Color map for the input features
  124. colors = ['black', 'black', 'black']
  125. matplot_colors = ["C{}".format(i) for i in range(20)]
  126. for i in range(3):
  127. ax = axes[i]
  128. # Plot each feature of the i-th sequence in the batch
  129. for feat in range(input_size):
  130. ax.plot(X_vis[i, :, feat].numpy(),
  131. color=colors[i],
  132. linestyle='-',
  133. marker='o',
  134. markersize=4,
  135. label=f'Feature {feat+1}' if i == 0 else ""
  136. )
  137. # Highlight the last 5 timesteps, which are used for the target sum
  138. ax.axvspan(seq_length-5, seq_length-1, alpha=0.3, label='Used for Target' if i == 0 else "")
  139. # Add the target value as text on the plot
  140. ax.text(0.40, 0.85, f'Output {i*32} = {y_vis[i].item():.2f}', transform=ax.transAxes,
  141. bbox=dict(boxstyle='round', facecolor=matplot_colors[0], alpha=0.3))
  142. ax.set_xlabel('Timestep')
  143. ax.set_title(f'Input {i*32}')
  144. ax.grid(True, linestyle='--', alpha=0.6)
  145. ax.set_xticks(np.arange(0, seq_length, 2))
  146. ax.set_xlim(0, 19)
  147. plt.tight_layout(rect=[0, 0.05, 1, 0.95]) # Adjust layout to make room for suptitle and legend
  148. plt.show()
  149. # %%
  150. import numpy as np
  151. import matplotlib.pyplot as plt
  152. from scipy.stats import multivariate_normal
  153. from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
  154. def create_rotated_covariance(major_var, minor_var, angle_deg):
  155. """
  156. Create a covariance matrix with specified rotation and axis variances.
  157. Parameters:
  158. major_var: variance along the major axis (long direction)
  159. minor_var: variance along the minor axis (short direction)
  160. angle_deg: rotation angle in degrees (0° = aligned with x-axis)
  161. """
  162. # Convert angle to radians
  163. angle_rad = np.deg2rad(angle_deg)
  164. # Create rotation matrix
  165. cos_angle = np.cos(angle_rad)
  166. sin_angle = np.sin(angle_rad)
  167. R = np.array([[cos_angle, -sin_angle],
  168. [sin_angle, cos_angle]])
  169. # Create eigenvalue matrix (diagonal)
  170. Lambda = np.array([[major_var, 0],
  171. [0, minor_var]])
  172. # Compute covariance matrix: Σ = R Λ Rᵀ
  173. covariance = R @ Lambda @ R.T
  174. return covariance
  175. # Set parameters
  176. angle = -45 # Keep rotation constant at 45°
  177. major_var_1, minor_var_1 = 3.0, 0.1 # First Gaussian: long and thin
  178. major_var_2, minor_var_2 = 2., 0.2 # Second Gaussian: also long and thin, but different
  179. # Create covariance matrices with same rotation but different shapes
  180. cov1 = create_rotated_covariance(major_var_1, minor_var_1, angle)
  181. cov2 = create_rotated_covariance(major_var_2, minor_var_2, angle)
  182. # Set random seed for reproducibility
  183. np.random.seed(42)
  184. # Create the figure with subplots
  185. plt.figure(figsize=(6, 5))
  186. # Generate data
  187. n_samples = 1000
  188. # Create two correlated 2D distributions
  189. mean1 = [1.2, 1]
  190. #cov1 = [[2, -1.6], [-0.4, .5]] # Positive correlation
  191. mean2 = [2, 2]
  192. #cov2 = [[1, -0.8], [-0.8, 1]] # Negative correlation
  193. # Generate samples
  194. samples1 = np.random.multivariate_normal(mean1, cov1, n_samples)
  195. samples2 = np.random.multivariate_normal(mean2, cov2, n_samples)
  196. # Prepare data for LDA
  197. X = np.vstack((samples1, samples2)) # Combine all features
  198. y = np.hstack((np.zeros(n_samples), np.ones(n_samples))) # Create labels: Class1=0, Class2=1
  199. # Fit Linear Discriminant Analysis (LDA)
  200. lda = LinearDiscriminantAnalysis()
  201. lda.fit(X, y)
  202. # Main 2D scatter plot
  203. ax_scatter = plt.subplot2grid((3, 3), (1, 0), colspan=2, rowspan=2)
  204. ax_scatter.scatter(samples1[:, 0], samples1[:, 1], alpha=0.6, label='Group 1', s=10)
  205. ax_scatter.scatter(samples2[:, 0], samples2[:, 1], alpha=0.6, label='Group 2', s=10)
  206. ax_scatter.set_xlabel('Feature 1')
  207. ax_scatter.set_ylabel('Feature 2')
  208. #ax_scatter.set_title('2D View: Easy Separation')
  209. ax_scatter.legend()
  210. ax_scatter.grid(True, alpha=0.1)
  211. # --- Plot the LDA separatrix ---
  212. # Create a mesh grid covering the plot area
  213. x_min, x_max = ax_scatter.get_xlim()
  214. y_min, y_max = ax_scatter.get_ylim()
  215. xx, yy = np.meshgrid(np.linspace(x_min, x_max, 100),
  216. np.linspace(y_min, y_max, 100))
  217. # Predict class probabilities for each grid point
  218. Z = lda.predict_proba(np.c_[xx.ravel(), yy.ravel()])[:, 1]
  219. Z = Z.reshape(xx.shape)
  220. # Plot the decision boundary (separatrix) at P(class=1) = 0.5
  221. contour = ax_scatter.contour(xx, yy, Z, levels=[0.5], colors='red', linewidths=3, linestyles='dashed')
  222. ax_scatter.clabel(contour, inline=True, fontsize=0)
  223. #, fmt='P(Class2)=0.5'
  224. # X marginal (top)
  225. ax_x_marginal = plt.subplot2grid((3, 3), (0, 0), colspan=2)
  226. ax_x_marginal.hist(samples1[:, 0], bins=30, alpha=0.5, density=True, label='Class 1')
  227. ax_x_marginal.hist(samples2[:, 0], bins=30, alpha=0.5, density=True, label='Class 2')
  228. ax_x_marginal.set_title('Feature 1')
  229. ax_x_marginal.set_ylabel('Density')
  230. #ax_x_marginal.legend()
  231. #ax_x_marginal.grid(True, alpha=0.3)
  232. # Y marginal (right)
  233. ax_y_marginal = plt.subplot2grid((3, 3), (1, 2), rowspan=2)
  234. ax_y_marginal.hist(samples1[:, 1], bins=30, orientation='horizontal', alpha=0.5, density=True, label='Class 1')
  235. ax_y_marginal.hist(samples2[:, 1], bins=30, orientation='horizontal', alpha=0.5, density=True, label='Class 2')
  236. ax_y_marginal.set_title('Feature 2')
  237. ax_y_marginal.set_xlabel('Density')
  238. #ax_y_marginal.legend()
  239. #ax_y_marginal.grid(True, alpha=0.3)
  240. # Add some empty space for better layout
  241. #ax_empty = plt.subplot2grid((4, 4), (0, 2), rowspan=1)
  242. #ax_empty.axis('off')
  243. plt.tight_layout()
  244. plt.show()

AI_model.ipynb at commit d159a15, under GPL-3.0 · at the source

Overview

Authors: Pascal Helson1,2,3, Arvind Kumar1,2
  1. School of Electrical Engineering and Computer Science and Digital Futures, KTH Royal Institute of Technology Stockholm,Stockholm, Sweden
  2. Science For Life Laboratory,Solna, Sweden
  3. UMR 5293, IMN, University of Bordeaux, CNRS,Bordeaux, France
Institutions: Digital Futures (Sweden); Science for Life Laboratory (Sweden); Université de Bordeaux (France)
Journal: NPJ digital medicine, volume 9, issue 1, article 361
Dates: received 11 September 2025; accepted 29 March 2026; published online 10 April 2026
Type: Review · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41746-026-02614-5 · PMID 41963496 · PMCID PMC13156286 · OpenAlex W4414597599
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), clinical / translational (subfield)
Methods: Single-unit activity, calcium imaging, Machine learning
Keywords: Biomarkers, Computational biology and bioinformatics, Neurology, Neuroscience
Topic: S100 Proteins and Annexins (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: not cited yet (Europe PMC); 51 references in the paper

Abstract

Despite the diversity and volume of brain data acquired and advanced AI-based algorithms to analyze them, brain features are rarely used in clinics for diagnosis and prognosis. Here we argue that the field continues to rely on cohort comparisons to seek biomarkers, despite the well-established degeneracy of brain features. Using a thought experiment (Brain Swap), we show that more data and more powerful algorithms will not be sufficient to identify biomarkers of brain diseases. We argue that instead of comparing patient versus healthy controls using single data type, we should use multimodal (e.g. brain activity, neurotransmitters, neuromodulators, brain imaging) and longitudinal brain data to guide the grouping before defining multidimensional biomarkers for brain diseases.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

paschels/brain_disease_biomarker

License: GPL-3.0
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: d159a15239d4e885edfc89c3a8e6783416a42411, 24 November 2025
Languages: Jupyter (1)
Size: 4 files, 1 script
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, 1 notebook
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), PyTorch (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
3 files

Code availability

We used Pytorch for our code, which you can find at https://github.com/paschels/brain_disease_biomarker.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 4 keywords, 1 funder, 46 references.

Cite

This paper

Helson, P., & Kumar, A. (2026). Pursuit of biomarkers of brain diseases: beyond cohort comparisons. NPJ digital medicine, 9(1), 361. https://doi.org/10.1038/s41746-026-02614-5

BibTeX

@article{helson2026pursuit,
author = {Helson, Pascal and Kumar, Arvind},
title = {{Pursuit of biomarkers of brain diseases: beyond cohort comparisons}},
journal = {NPJ digital medicine},
year = {2026},
month = apr,
volume = {9},
number = {1},
pages = {361},
publisher = {Nature Publishing Group},
issn = {2398-6352},
doi = {10.1038/s41746-026-02614-5},
url = {https://doi.org/10.1038/s41746-026-02614-5},
pmid = {41963496},
pmcid = {PMC13156286}
}

RIS

TY - JOUR
AU - Helson, Pascal
AU - Kumar, Arvind
TI - Pursuit of biomarkers of brain diseases: beyond cohort comparisons
T2 - NPJ digital medicine
J2 - NPJ Digit Med
PY - 2026
DA - 2026/04/10
VL - 9
IS - 1
SP - 361
SN - 2398-6352
PB - Nature Publishing Group
DO - 10.1038/s41746-026-02614-5
UR - https://doi.org/10.1038/s41746-026-02614-5
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41746-026-02614-5",
"type": "article-journal",
"title": "Pursuit of biomarkers of brain diseases: beyond cohort comparisons",
"container-title": "NPJ digital medicine",
"author": [
{
"family": "Helson",
"given": "Pascal"
},
{
"family": "Kumar",
"given": "Arvind"
}
],
"container-title-short": "NPJ Digit Med",
"volume": "9",
"issue": "1",
"page": "361",
"DOI": "10.1038/s41746-026-02614-5",
"PMID": "41963496",
"PMCID": "PMC13156286",
"ISSN": "2398-6352",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41746-026-02614-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
10
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-71918-7 [code]
Developmental disinhibition gates language lateralization in childhood.
Journal: Nature communications
In common: PyTorch, scikit-learn, SciPy, 2 other tools, 3 references
[2] doi:10.1186/s12916-026-04903-y [code]
Structural connectome architecture and biological vulnerability shape cortical atrophy in cocaine use disorder.
Journal: BMC medicine
In common: scikit-learn, SciPy, Matplotlib, 1 other tool, 3 references
[3] doi:10.1038/s41593-026-02205-3 [code]
Competitive interactions shape mammalian brain network dynamics and computation.
Journal: Nature neuroscience
In common: PyTorch, scikit-learn, SciPy, 2 other tools, 2 references
[4] doi:10.1371/journal.pone.0344600 [code]
Robust disease prognosis via diagnostic knowledge preservation: A sequential learning approach.
Journal: PloS one
In common: PyTorch, scikit-learn, SciPy, 2 other tools, clinical / translational, 1 reference
[5] doi:10.1038/s41467-026-71555-0 [code]
A deep representation learning model to predict response to vagus nerve stimulation.
Journal: Nature communications
In common: PyTorch, scikit-learn, SciPy, 2 other tools, clinical / translational, 1 reference
[6] doi:10.1093/braincomms/fcag253 [code]
Disease detection and classification in temporal lobe epilepsy: step-wise versus simultaneous AI decision models in a multisite neuroimaging study.
Journal: Brain communications
In common: PyTorch, scikit-learn, SciPy, 2 other tools, clinical / translational, 1 reference
[7] doi:10.1038/s41467-026-71961-4 [code]
Spatiotemporal asymmetries on brain energy landscape uncover system entrapment related to depression severity.
Journal: Nature communications
In common: scikit-learn, SciPy, Matplotlib, 1 other tool, 2 references
[8] doi:10.1038/s43856-026-01606-6 [code]
Validation of remote multimodal AI screening for Parkinson disease across diverse settings.
Journal: Communications medicine
In common: PyTorch, scikit-learn, SciPy, 2 other tools, clinical / translational, 1 reference
[9] doi:10.1038/s41398-026-04101-7 [code]
Using deep learning to identify brain networks mediating cognitive and motor impairments in alcohol use disorder.
Journal: Translational psychiatry
In common: PyTorch, scikit-learn, SciPy, 2 other tools, 1 reference
[10] doi:10.1162/imag.a.1269 [code]
From early to contemporary normative modeling: Mapping individual differences in neurophysiological signals.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: scikit-learn, SciPy, Matplotlib, 1 other tool, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.