OSCR

Label-free biochemical imaging and time point analysis of neural organoids via deep learning-enhanced Raman microspectroscopy.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § MATERIALS AND METHODS › Spectral preprocessing of Raman imaging data ↔ colab/Deep_learning_enhanced_Raman_microspectroscopy.ipynb, lines 244–291 · score 0.91 · global vector normalization, cosmic rays, baseline correction, Hayes, Whitaker, Whittaker
  2. [2] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ colab/Deep_learning_enhanced_Raman_microspectroscopy.ipynb, lines 513–585 · score 0.77 · Batch normalization, ReLU, dropout, rectified, filters, blocks
  3. [3] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ src/autoencoder.py, lines 34–111 · score 0.77 · Batch normalization, ReLU, dropout, rectified, filters, blocks
  4. [4] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ analyse.py, lines 11–48 · score 0.73 · training epochs, training loss, learning rate decay, Adam, MSE, SAD
  5. [5] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ colab/Deep_learning_enhanced_Raman_microspectroscopy.ipynb, lines 513–585 · score 0.72 · linear decoder layer, kernel constraint, activation, encoded, bias, weight
  6. [6] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ src/autoencoder.py, lines 34–111 · score 0.71 · linear decoder layer, kernel constraint, activation, encoded, bias, weight
  7. [7] § MATERIALS AND METHODS › Hyperspectral unmixing analysis ↔ src/autoencoder.py, lines 114–177 · score 0.60 · learning rate decay, exponential, Adam, epochs, SAD, optimizer

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 768 lines · 27 KB · Apache-2.0 · 3 matches

  1. # %% [markdown]
  2. # # 🔬 Deep Learning-Enhanced Raman Microspectroscopy
  3. #
  4. # ---
  5. #
  6. # ## ✨ Welcome
  7. #
  8. # This notebook provides an interactive version of the Raman imaging pipeline described in [Georgiev et al., Science Advances, 2026](https://doi.org/10.1126/sciadv.aec5080).
  9. #
  10. # ⚡ **Purpose:** Extract chemical components and their spatial maps from hyperspectral Raman imaging data.
  11. #
  12. # 💡 **How:** Using blind, unsupervised hyperspectral unmixing via physics-constrained autoencoder neural networks.
  13. #
  14. # ⚠️ **Note:** Uploading large files in Google Colab can be slow and unreliable. For large datasets (e.g. >500 MB-1 GB), we recommend running the pipeline locally using the provided GUI/CLI tools. See our [GitHub page](https://github.com/barahona-research-group/dl-raman) for more information.
  15. #
  16. # ---
  17. # ## ⚙️ Workflow
  18. #
  19. # This notebook guides you through the full analysis workflow:
  20. #
  21. # 1. Upload Raman imaging data
  22. # 2. Preprocess the spectra
  23. # 3. Run PCA to guide the selection of the number of endmembers
  24. # 4. Unmix the data into chemical components and abundance maps
  25. # 5. Save figures, results and metadata
  26. #
  27. # ---
  28. #
  29. # ## 📚 Related Publications
  30. #
  31. # This notebook provides an implementation of the analysis pipeline described in:
  32. #
  33. # > 📌 **Label-free biochemical imaging and timepoint analysis of neural organoids via deep learning-enhanced Raman microspectroscopy.**
  34. # > Dimitar Georgiev, Ruoxiao Xie, Daniel Reumann, Xiaoyu Zhao, Álvaro Fernández-Galiana, Mauricio Barahona, Molly M. Stevens.
  35. # > Science Advances, 12(33), eaec5080, 2026.
  36. # > https://doi.org/10.1126/sciadv.aec5080
  37. #
  38. # This work builds upon our previous research:
  39. #
  40. # > 📌 **Hyperspectral unmixing for Raman spectroscopy via physics-constrained autoencoders.**
  41. # > Dimitar Georgiev, Álvaro Fernández-Galiana, Simon Vilms Pedersen, Georgios Papadopoulos, Ruoxiao Xie, Molly M. Stevens, Mauricio Barahona.
  42. # > PNAS, 121(45), p.e2407439121, 2024.
  43. # > https://doi.org/10.1073/pnas.2407439121
  44. #
  45. # > 📌 **RamanSPy: An open-source Python package for integrative Raman spectroscopy data analysis.**
  46. # > Dimitar Georgiev, Simon Vilms Pedersen, Ruoxiao Xie, Álvaro Fernández-Galiana, Molly M. Stevens, Mauricio Barahona.
  47. # > Analytical Chemistry, 96(21), pp.8492-8500, 2024.
  48. # > https://doi.org/10.1021/acs.analchem.4c00383
  49. #
  50. # ⚠️ **If you use this notebook in your research, please cite our work.**
  51. # %%
  52. #@title **Install & Import Dependencies**
  53. #@markdown ---
  54. #@markdown Click the play button ▶️ to install and import relevant packages.
  55. import os
  56. import sys
  57. import glob
  58. import pickle
  59. import tifffile
  60. from google.colab import files
  61. import io
  62. from datetime import datetime
  63. import random
  64. import matplotlib.pyplot as plt
  65. import numpy as np
  66. from sklearn.decomposition import PCA
  67. import tensorflow as tf
  68. from IPython.display import clear_output
  69. !pip install ramanspy
  70. import ramanspy as rp
  71. os.environ['TF_CUDNN_DETERMINISTIC'] = '1'
  72. os.environ['TF_DETERMINISTIC_OPS'] = '1'
  73. os.environ['CUBLAS_WORKSPACE_CONFIG'] = ':4096:8'
  74. tf.config.experimental.enable_op_determinism()
  75. def set_seed(seed=None):
  76. # Set seeds for reproducibility
  77. os.environ['PYTHONHASHSEED'] = str(seed)
  78. random.seed(seed)
  79. np.random.seed(seed)
  80. tf.random.set_seed(seed)
  81. tf.keras.utils.set_random_seed(seed)
  82. def save_rp_object(spectral_container, caption, *, folder=None):
  83. if folder is not None:
  84. os.makedirs(folder, exist_ok=True)
  85. caption = os.path.join(folder, caption)
  86. spectral_container.save(f'{caption}_rp_spectral_container.pkl')
  87. np.save(f'{caption}_spectral_data.npy', spectral_container.spectral_data)
  88. if spectral_container.spectral_data.ndim >= 2:
  89. spectral_data = np.moveaxis(np.copy(spectral_container.spectral_data), -1, 0)
  90. spectral_data = np.moveaxis(spectral_data, -1, 0) if len(spectral_data.shape) > 3 else spectral_data # TZCYX order
  91. tifffile.imwrite(f'{caption}_spectral_data.tiff', spectral_data)
  92. np.savetxt(f'{caption}_spectral_axis.txt', spectral_container.spectral_axis)
  93. clear_output()
  94. print("✅ Packages installed successfully.")
  95. # %%
  96. #@title **Experiment setup**
  97. #@markdown General settings
  98. #@markdown ---
  99. experiment_name = "experiment" #@param {type:"string"}
  100. seed = 42 #@param {type:"number"}
  101. # Generate timestamp
  102. timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
  103. # Build folder name
  104. save_folder = f"{experiment_name}_{timestamp}"
  105. # Create directory
  106. os.makedirs(save_folder, exist_ok=True)
  107. print(f"✅ Results will be saved in: `{save_folder}/`")
  108. # %%
  109. #@title **Upload Raman imaging data and optional reference spectra** (.pkl, .mat, .wdf)
  110. #@markdown ---
  111. #@markdown Click the play button ▶️ to open the upload dialog.
  112. #@markdown
  113. #@markdown **Supported formats:**
  114. #@markdown - `.pkl`: `Spectrum`, `SpectralContainer`, `SpectralImage`, or `SpectralVolume` exported from RamanSPy
  115. #@markdown - `.mat`: WITec MATLAB export
  116. #@markdown - `.wdf`: Renishaw WiRE export
  117. #@markdown
  118. #@markdown At least one uploaded file must be Raman imaging data. Optional reference spectra are included in preprocessing and unmixing, but spatial abundance-map outputs are skipped for non-imaging inputs.
  119. #@markdown
  120. #@markdown ✅ Multiple files can be analyzed together in batch mode.
  121. #@markdown
  122. #@markdown ---
  123. #@markdown ⚠️ **Note:** Uploading large files in Google Colab can be slow and unreliable.
  124. #@markdown For large datasets (e.g. >500 MB-1 GB), we recommend running the pipeline locally using the local GUI or command-line interface provided on our [GitHub page](https://github.com/barahona-research-group/dl-raman).
  125. def _is_imaging_object(spectral_obj):
  126. return len(spectral_obj.spectral_data.shape) in [3, 4]
  127. def _as_2d_spectral_data(data):
  128. data = np.asarray(data)
  129. if data.ndim == 1:
  130. return data.reshape(1, -1)
  131. return data
  132. def _validate_raman_object(spectral_obj, filename):
  133. data_shape = spectral_obj.spectral_data.shape
  134. if len(data_shape) not in [1, 2, 3, 4]:
  135. raise ValueError(
  136. f"{filename} is not a supported Raman dataset. Expected shape "
  137. "(wavenumbers), (spectra, wavenumbers), "
  138. "(x, y, wavenumbers), or (x, y, z, wavenumbers), "
  139. f"but found {data_shape}."
  140. )
  141. uploaded = files.upload()
  142. assert uploaded, "No file uploaded."
  143. uploaded_filenames = list(uploaded.keys())
  144. print(f"✅ {len(uploaded_filenames)} file(s) uploaded.")
  145. spectral_objects = []
  146. spectral_object_names = []
  147. raw_data_folder = os.path.join(save_folder, 'Raw data')
  148. os.makedirs(raw_data_folder, exist_ok=True)
  149. for filename in uploaded_filenames:
  150. print(f"\n📁 Parsing: {filename}")
  151. ext = os.path.splitext(filename)[1].lower()
  152. # Save file to disk
  153. scan_folder = os.path.join(raw_data_folder, os.path.splitext(filename)[0])
  154. os.makedirs(scan_folder, exist_ok=True)
  155. with open(os.path.join(scan_folder, filename), 'wb') as f:
  156. f.write(uploaded[filename])
  157. # Load depending on extension
  158. if ext == ".pkl":
  159. spectral_obj = rp.SpectralContainer.load(filename)
  160. elif ext == ".mat":
  161. spectral_obj = rp.load.witec(filename)
  162. elif ext == ".wdf":
  163. if hasattr(rp.load, 'renishaw'):
  164. spectral_obj = rp.load.renishaw(filename)
  165. elif hasattr(rp.load, 'wdf'):
  166. spectral_obj = rp.load.wdf(filename)
  167. else:
  168. raise RuntimeError("Loading .wdf files requires a RamanSPy version with a Renishaw/WDF loader.")
  169. else:
  170. print(f"⚠️ Skipping unsupported file: {filename}")
  171. continue
  172. _validate_raman_object(spectral_obj, filename)
  173. # Store object
  174. spectral_objects.append(spectral_obj)
  175. spectral_object_names.append(filename)
  176. input_type = "imaging" if _is_imaging_object(spectral_obj) else "reference/non-imaging"
  177. # Display info
  178. print(f" ✅ Shape: {spectral_obj.shape}")
  179. print(f" Type: {input_type}")
  180. print(f" Wavenumber axis length: {len(spectral_obj.spectral_axis)}")
  181. # Save parsed object
  182. save_rp_object(
  183. spectral_obj,
  184. caption='raw',
  185. folder=scan_folder
  186. )
  187. assert spectral_objects, "No supported Raman datasets were loaded."
  188. assert any(_is_imaging_object(obj) for obj in spectral_objects), "At least one Raman image or volume is required. Reference spectra can be included alongside imaging data, but cannot be the only inputs."
  189. print(f"\n💾 All parsed spectral objects saved in: `{raw_data_folder}/`")
  190. print(f"📦 Total spectral datasets loaded: {len(spectral_objects)}")
  191. print(f"🖼️ Imaging datasets: {sum(_is_imaging_object(obj) for obj in spectral_objects)}")
  192. print(f"📈 Reference/non-imaging datasets: {sum(not _is_imaging_object(obj) for obj in spectral_objects)}")
  193. print("ℹ️ Run the preprocessing cell before PCA and unmixing.")
  194. # %%
  195. #@title **Spectral preprocessing**
  196. #@title 🧹 **Spectral preprocessing**
  197. #@markdown ---
  198. #@markdown This cell performs a standard Raman spectral preprocessing pipeline consisting of the following steps:
  199. #@markdown
  200. #@markdown 1. **Cropping** – trims the spectral axis to remove low-wavenumber noise (starting from 300 cm⁻¹).
  201. #@markdown 2. **Despiking** – removes cosmic ray spikes using the Whitaker-Hayes filter.
  202. #@markdown 3. **Denoising** – smooths the spectra using Whittaker smoothing (λ=1e3, d=3).
  203. #@markdown 4. **Baseline Correction** – removes fluorescence/background using ASLS (λ=1e5).
  204. #@markdown 5. **Normalization** – normalizes spectra using vector normalization (global, not pixelwise).
  205. #@markdown
  206. #@markdown ✏️ *To customize preprocessing, click on `Show code` and modify the `PREPROCESSING_PIPELINE` object below.*
  207. #@markdown
  208. #@markdown ℹ️ The processed data will be saved in the `Preprocessed data/` subfolder of your run directory.
  209. PREPROCESSING_PIPELINE = rp.preprocessing.Pipeline([
  210. rp.preprocessing.misc.Cropper(region=(300, None)), # 1. Crop low-wavenumber region
  211. rp.preprocessing.despike.WhitakerHayes(kernel_size=3, threshold=6), # 2. Despike
  212. rp.preprocessing.denoise.Whittaker(lam=1e3, d=3), # 3. Denoise
  213. rp.preprocessing.baseline.ASLS(lam=1e5), # 4. Baseline correction
  214. rp.preprocessing.normalise.Vector(pixelwise=False) # 5. Global vector normalization
  215. ])
  216. print("🧹 Preprocessing starting...")
  217. set_seed(seed) # Ensures reproducibility
  218. preprocessed_spectral_objs = PREPROCESSING_PIPELINE.apply(spectral_objects)
  219. preprocessed_spectral_axis = preprocessed_spectral_objs[0].spectral_axis
  220. print("✅ Preprocessing complete.")
  221. # Save results
  222. preprocessed_data_folder = os.path.join(save_folder, 'Preprocessed data')
  223. os.makedirs(preprocessed_data_folder, exist_ok=True)
  224. # Save updated objects
  225. for filename, spectral_obj in zip(spectral_object_names, preprocessed_spectral_objs):
  226. save_rp_object(
  227. spectral_obj,
  228. caption=f'preprocessed',
  229. folder=os.path.join(preprocessed_data_folder, os.path.splitext(filename)[0])
  230. )
  231. print(f"💾 Preprocessed data saved in: `{preprocessed_data_folder}/`")
  232. # %%
  233. #@title **PCA plot**
  234. def pca_explained_variance_plot(data, n_components=None, save_path=None, show_plot=True):
  235. """
  236. Performs PCA on a 2D data matrix and produces an explained variance plot.
  237. Parameters
  238. ----------
  239. data : array-like, shape (n_samples, n_features)
  240. Input data matrix where rows correspond to samples and columns to features.
  241. n_components : int or None, optional (default=None)
  242. Number of principal components to compute. If None or larger than the number of features,
  243. all components are computed.
  244. save_path : str or None, optional (default=None)
  245. File path where the plot image will be saved (e.g., 'fig1.png'). If None, the plot is not saved.
  246. show_plot : bool, optional (default=True)
  247. If True, displays the plot. If False, the figure is closed after saving.
  248. Returns
  249. -------
  250. pca : sklearn.decomposition.PCA object
  251. The fitted PCA object.
  252. explained_variance_ratio : ndarray, shape (n_components,)
  253. Percentage of variance explained by each selected principal component.
  254. """
  255. # Determine the number of components to use
  256. n_samples, n_features = data.shape
  257. max_components = min(n_samples, n_features)
  258. if n_components is None or n_components > max_components:
  259. n_components = max_components
  260. # Perform PCA (assuming the data is already centered/scaled as needed)
  261. pca = PCA(n_components=n_components)
  262. pca.fit(data)
  263. explained_variance_ratio = pca.explained_variance_ratio_
  264. # Create the plot
  265. fig, ax = plt.subplots(figsize=(12, 6))
  266. components = np.arange(1, n_components + 1)
  267. variance_percentage = explained_variance_ratio * 100 # Convert to percentage
  268. # Create a bar plot for the explained variance
  269. bars = ax.bar(components, variance_percentage, color='skyblue', edgecolor='black')
  270. # Set axes labels and title with appropriate font sizes and padding
  271. ax.set_xlabel('Principal Component', fontsize=16)
  272. ax.set_ylabel('Explained Variance (%)', fontsize=16)
  273. ax.set_title('Explained Variance by Principal Component', fontsize=18, pad=15)
  274. ax.set_xticks(components)
  275. ax.set_ylim(0, variance_percentage.max() * 1.1) # Add some headroom above the highest bar
  276. # Optionally annotate each bar with its percentage value
  277. for comp, var in zip(components, variance_percentage):
  278. ax.text(comp, var + variance_percentage.max() * 0.01, f"{var:.2f}%",
  279. ha='center', va='bottom', fontsize=8)
  280. plt.tight_layout()
  281. # Save the plot if a save path is provided
  282. if save_path:
  283. plt.savefig(save_path, dpi=300, bbox_inches='tight')
  284. # Show or close the plot based on the user's preference
  285. if show_plot:
  286. plt.show()
  287. else:
  288. plt.close(fig)
  289. return pca, explained_variance_ratio
  290. analysis_data_folder = os.path.join(save_folder, 'Analysis data')
  291. os.makedirs(analysis_data_folder, exist_ok=True)
  292. assert 'preprocessed_spectral_objs' in globals(), "Run the spectral preprocessing cell before PCA."
  293. spectral_data = np.vstack([_as_2d_spectral_data(s.flat.spectral_data) for s in preprocessed_spectral_objs])
  294. _ = pca_explained_variance_plot(spectral_data, n_components=20, save_path=os.path.join(analysis_data_folder, 'pca_plot.png'))
  295. # %%
  296. #@title **Hyperspectral unmixing analysis**
  297. #@markdown ---
  298. #@markdown Unmixing settings
  299. #@markdown ---
  300. num_endmembers = 10 #@param {type:"number"}
  301. vca_init = True #@param {type:"boolean"}
  302. #@markdown ---
  303. #@markdown Model training settings
  304. #@markdown ---
  305. loss = "MSE_SAD" # @param ["MSE", "SAD", "MSE_SAD"]
  306. epochs = 5 #@param {type:"number"}
  307. learning_rate = 0.001 #@param {type:"number"}
  308. batch_size = 32 #@param {type:"number"}
  309. #@markdown ---
  310. #@markdown Advanced settings
  311. #@markdown ---
  312. verbose = 1 #@param {type:"number"}
  313. use_bias = False #@param {type:"boolean"}
  314. learning_rate_decay = True #@param {type:"boolean"}
  315. assert 'preprocessed_spectral_objs' in globals(), "Run the spectral preprocessing cell before unmixing."
  316. analysis_data_folder = os.path.join(save_folder, 'Analysis data')
  317. os.makedirs(analysis_data_folder, exist_ok=True)
  318. # Pack into a dictionary
  319. experiment_params = {
  320. "num_endmembers": num_endmembers,
  321. "vca_init": vca_init,
  322. "loss": loss,
  323. "epochs": epochs,
  324. "learning_rate": learning_rate,
  325. "batch_size": batch_size,
  326. "verbose": verbose,
  327. "use_bias": use_bias,
  328. "learning_rate_decay": learning_rate_decay,
  329. "seed": seed
  330. }
  331. # Save to JSON inside the timestamped save folder
  332. import json
  333. params_path = os.path.join(analysis_data_folder, "params.json")
  334. with open(params_path, 'w') as f:
  335. json.dump(experiment_params, f, indent=2)
  336. print(f"✅ Parameters saved to: `{params_path}`")
  337. def save_results(abundance_maps, endmembers, x_axis, *, save_to):
  338. os.makedirs(save_to, exist_ok=True)
  339. # save x-axis
  340. np.savetxt(os.path.join(save_to, 'wavenumber_axis.txt'), x_axis)
  341. # save endmembers
  342. np.savetxt(os.path.join(save_to, 'endmembers.txt'), endmembers)
  343. # save abundances
  344. abundance_array = np.array(abundance_maps)
  345. abundance_array = np.moveaxis(abundance_array, -1, 0) if len(abundance_array.shape) > 3 else abundance_array # TZCYX order
  346. tifffile.imwrite(os.path.join(save_to, 'abundances.tiff'), abundance_array)
  347. def save_plots(abundance_maps, endmembers, spectral_axis, *, cmap='gray', save_to=None, plot=True):
  348. endmember_array = [rp.Spectrum(endmember, spectral_axis) for endmember in endmembers]
  349. # Plot the endmembers
  350. plt.figure(figsize=(8, 4))
  351. ax = rp.plot.spectra(endmember_array, title='Endmembers', color='gray', plot_type='single stacked')
  352. if save_to is not None:
  353. ax.get_figure().savefig(os.path.join(save_to, 'endmembers.png'), dpi=600)
  354. if plot is True:
  355. plt.show()
  356. else:
  357. plt.close()
  358. abundance_array = np.array(abundance_maps)
  359. if len(abundance_array.shape) == 3:
  360. abundance_array = abundance_array[..., np.newaxis]
  361. # Plot the abundance maps
  362. fig, axs = plt.subplots(abundance_array.shape[-1], len(endmember_array), figsize=(1.75 * len(endmember_array), 2*abundance_array.shape[-1]), sharey=True)
  363. for j in range(abundance_array.shape[-1]):
  364. if abundance_array.shape[-1] == 1:
  365. axs = [axs]
  366. for i, ax in enumerate(axs[j]):
  367. ax.imshow(abundance_array[i, :, :, j], cmap=cmap)
  368. ax.set_axis_off()
  369. if save_to is not None:
  370. os.makedirs(save_to, exist_ok=True)
  371. fig.savefig(os.path.join(save_to, 'abundances.png'), dpi=600)
  372. plt.show()
  373. def _soft_rectified_tanh(x, gamma=10.0):
  374. return (1 / gamma) * tf.math.log(1 + tf.math.exp(gamma * tf.keras.activations.tanh(x)))
  375. cos = tf.losses.CosineSimilarity()
  376. mse = tf.losses.MeanSquaredError()
  377. def SAD(y_true, y_pred):
  378. sad_loss = tf.math.acos(-1*cos(y_true, y_pred))
  379. return sad_loss
  380. def MSE_SAD(beta1=0.5, beta2=1):
  381. def mse_sad(y_true, y_pred):
  382. mse_loss = mse(y_true, y_pred)
  383. sad_loss = tf.math.acos(-1*cos(y_true, y_pred))
  384. return beta1*mse_loss + beta2*sad_loss
  385. return mse_sad
  386. loss_dict = {
  387. 'MSE': mse,
  388. 'SAD': SAD,
  389. 'MSE_SAD': MSE_SAD(beta1=1e4, beta2=1)
  390. }
  391. loss = loss_dict[loss]
  392. class ConvUnmixingAEModel(tf.keras.Model):
  393. def __init__(self, input_dim, bottleneck_dim, use_bias):
  394. super(ConvUnmixingAEModel, self).__init__()
  395. kernel_sizes=[5, 10, 15, 20]
  396. num_filters=[32, 32, 32, 32]
  397. encoder_hidden_dims=[256, 128, 64, 32]
  398. # ENCODER
  399. self.conv_layers = []
  400. for kernel_size, num_filter in zip(kernel_sizes, num_filters):
  401. conv_block = tf.keras.models.Sequential([
  402. tf.keras.layers.Conv1D(num_filter, kernel_size=kernel_size, padding='same', activation='relu', kernel_initializer=tf.keras.initializers.HeNormal()),
  403. tf.keras.layers.BatchNormalization(),
  404. tf.keras.layers.Dropout(0.2)
  405. ])
  406. self.conv_layers.append(conv_block)
  407. self.concat = tf.keras.layers.Concatenate(axis=-1)
  408. self.linear_dim_reduce = tf.keras.layers.Dense(1)
  409. # ENCODER
  410. self.ffn_layers = tf.keras.models.Sequential()
  411. for dim in encoder_hidden_dims:
  412. self.ffn_layers.add(tf.keras.layers.Dense(dim, activation='relu', kernel_initializer=tf.keras.initializers.HeNormal()))
  413. self.ffn_layers.add(tf.keras.layers.BatchNormalization())
  414. self.ffn_layers.add(tf.keras.layers.Dropout(0.5))
  415. self.bottleneck_layer = tf.keras.layers.Dense(bottleneck_dim, activation=_soft_rectified_tanh)
  416. # DECODER
  417. self.linear_decoder_layer = tf.keras.layers.Dense(input_dim, activation='linear', kernel_constraint=tf.keras.constraints.NonNeg(), use_bias=use_bias)
  418. def _encode(self, input_x):
  419. x = tf.expand_dims(input_x, axis=-1)
  420. conv_outputs = []
  421. for conv_layer in self.conv_layers:
  422. conv_outputs.append(conv_layer(x))
  423. x = self.concat(conv_outputs)
  424. x = self.linear_dim_reduce(x)
  425. x = tf.squeeze(x, axis=-1)
  426. x = self.ffn_layers(x)
  427. x = self.bottleneck_layer(x)
  428. return x
  429. def predict_abundances(self, data):
  430. return self._encode(data)
  431. def _initialize_endmembers(self, endmembers):
  432. endmembers = np.array(endmembers)
  433. self.linear_decoder_layer.build(input_shape=(None, endmembers.shape[0]))
  434. weigths = [endmembers, self.linear_decoder_layer.get_weights()[1]] if len(self.linear_decoder_layer.get_weights())>1 else [endmembers]
  435. self.linear_decoder_layer.set_weights(weigths)
  436. def predict_endmembers(self):
  437. return self.linear_decoder_layer.trainable_variables[0].numpy()
  438. def _decode(self, x):
  439. return self.linear_decoder_layer(x)
  440. def call(self, x):
  441. z = self._encode(x)
  442. x_hat = self._decode(z)
  443. return x_hat
  444. class ConvUnmixingAE:
  445. def apply(self, raman_objects, *, num_endmembers, use_bias=False, loss=SAD, batch_size=32, epochs=5, vca_init=False, seed=None, learning_rate=0.001, learning_rate_decay=False, verbose=0, return_endmember_init=False):
  446. if not isinstance(raman_objects, list):
  447. raman_objects = [raman_objects]
  448. flat_spectral_data = [_as_2d_spectral_data(v.flat.spectral_data) for v in raman_objects]
  449. spectral_data_concat = np.concatenate(flat_spectral_data)
  450. input_dim = len(raman_objects[0].spectral_axis)
  451. set_seed(seed)
  452. # create the model
  453. model = ConvUnmixingAEModel(input_dim, num_endmembers, use_bias=use_bias)
  454. if learning_rate_decay:
  455. steps_per_epoch = len(spectral_data_concat) / batch_size
  456. decay_epochs = 1
  457. learning_rate = tf.keras.optimizers.schedules.ExponentialDecay(
  458. initial_learning_rate=learning_rate,
  459. decay_steps=steps_per_epoch * decay_epochs,
  460. decay_rate=0.9,
  461. staircase=True
  462. )
  463. model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate), loss=loss)
  464. if vca_init:
  465. endmember_init = _vca(spectral_data_concat, num_endmembers)
  466. model._initialize_endmembers(endmember_init)
  467. # fit the model
  468. model.fit(spectral_data_concat, spectral_data_concat, epochs=epochs, verbose=verbose, batch_size=batch_size, shuffle=True)
  469. # model.summary()
  470. # get endmembers
  471. endmembers = model.predict_endmembers()
  472. # get abundances
  473. abundance_maps = []
  474. for scan, flat_data in zip(raman_objects, flat_spectral_data):
  475. spectral_dataset = tf.data.Dataset.from_tensor_slices(flat_data).batch(batch_size)
  476. predictions = []
  477. for batch in spectral_dataset:
  478. batch_predictions = model.predict_abundances(batch)
  479. predictions.append(batch_predictions)
  480. # If needed, concatenate the predictions into one array
  481. predictions = np.concatenate(predictions, axis=0)
  482. abundances = predictions.reshape(list(scan.shape) + [-1])
  483. # move channels to the last axis to comply with ramanspy unmixing
  484. abundances = np.moveaxis(abundances, -1, 0)
  485. abundance_maps.append(abundances)
  486. if return_endmember_init:
  487. return abundance_maps, endmembers, endmember_init
  488. return abundance_maps, endmembers
  489. def _vca(data, n_endmembers, *, snr_input=0):
  490. """
  491. Copyright 2018 Adrien Lagrange
  492. Licensed under the Apache License, Version 2.0.
  493. Source code available at: https://github.com/Laadr/VCA
  494. Changes:
  495. - sp -> np;
  496. - removed comments;
  497. - removed semicolons at the end of lines;
  498. - removed verbose option and print statements;
  499. - renamed some variables;
  500. - only return endmembers;
  501. """
  502. import scipy.linalg as splin
  503. # Transpose data to comply with code
  504. data = data.T
  505. N = data.shape[1]
  506. # Estimate SNR
  507. # ------------
  508. if snr_input == 0:
  509. data_mean = np.mean(data, axis=1, keepdims=True)
  510. data_centered = data - data_mean # data with zero-mean
  511. Ud = splin.svd(np.dot(data_centered, data_centered.T) / float(N))[0][:, :n_endmembers]
  512. x_p = np.dot(Ud.T, data_centered)
  513. P_y = np.sum(data ** 2) / float(N)
  514. P_x = np.sum(x_p ** 2) / float(N) + np.sum(data_mean ** 2)
  515. SNR = 10 * np.log10((P_x - n_endmembers / data.shape[0] * P_y) / (P_y - P_x))
  516. else:
  517. SNR = snr_input
  518. SNR_th = 15 + 10 * np.log10(n_endmembers)
  519. # Projection to n_endmembers-1 subspace if SNR < SNR_th; else, no projective projection
  520. # -------------------------------------------------------------------------------------
  521. if SNR < SNR_th:
  522. d = n_endmembers - 1
  523. if snr_input == 0:
  524. Ud = Ud[:, :d]
  525. else:
  526. data_mean = np.mean(data, axis=1, keepdims=True)
  527. data_centered = data - data_mean
  528. Ud = splin.svd(np.dot(data_centered, data_centered.T) / float(N))[0][:, :d] # computes the p-projection matrix
  529. x_p = np.dot(Ud.T, data_centered)
  530. Yp = np.dot(Ud, x_p[:d, :]) + data_mean
  531. x = x_p[:d, :]
  532. c = np.amax(np.sum(x ** 2, axis=0)) ** 0.5
  533. y = np.vstack((x, c * np.ones((1, N))))
  534. else:
  535. d = n_endmembers
  536. Ud = splin.svd(np.dot(data, data.T) / float(N))[0][:, :d]
  537. x_p = np.dot(Ud.T, data)
  538. Yp = np.dot(Ud, x_p[:d, :])
  539. x = np.dot(Ud.T, data)
  540. u = np.mean(x, axis=1, keepdims=True)
  541. y = x / np.dot(u.T, x)
  542. # VCA algorithm
  543. # -------------
  544. indice = np.zeros((n_endmembers), dtype=int)
  545. A = np.zeros((n_endmembers, n_endmembers))
  546. A[-1, 0] = 1
  547. for i in range(n_endmembers):
  548. w = np.random.rand(n_endmembers, 1)
  549. f = w - np.dot(A, np.dot(splin.pinv(A), w))
  550. f = f / splin.norm(f)
  551. v = np.dot(f.T, y)
  552. indice[i] = np.argmax(np.absolute(v))
  553. A[:, i] = y[:, indice[i]]
  554. Ae = Yp[:, indice]
  555. return Ae.T
  556. print("🔧 Initializing model with selected parameters...")
  557. conv_ae_unmixer = ConvUnmixingAE()
  558. print("📊 Starting the unmixing process. This may take a few moments...")
  559. abundances, endmembers = conv_ae_unmixer.apply(preprocessed_spectral_objs, num_endmembers=num_endmembers, epochs=epochs, seed=seed, use_bias=use_bias, loss=loss, batch_size=batch_size, verbose=verbose, learning_rate=learning_rate, learning_rate_decay=learning_rate_decay, vca_init=vca_init)
  560. print("✅ Unmixing complete. Extracted endmembers and abundances are ready.")
  561. # Save results
  562. for i, (filename, spectral_obj, amaps) in enumerate(zip(spectral_object_names, preprocessed_spectral_objs, abundances)):
  563. save_to = os.path.join(analysis_data_folder, os.path.splitext(filename)[0])
  564. if _is_imaging_object(spectral_obj):
  565. save_results(amaps, endmembers, preprocessed_spectral_axis, save_to=save_to)
  566. to_plot = True if i < 1 else False
  567. save_plots(amaps, endmembers, preprocessed_spectral_axis, save_to=save_to, plot=to_plot)
  568. else:
  569. os.makedirs(save_to, exist_ok=True)
  570. np.savetxt(os.path.join(save_to, 'wavenumber_axis.txt'), preprocessed_spectral_axis)
  571. np.savetxt(os.path.join(save_to, 'endmembers.txt'), endmembers)
  572. np.savetxt(os.path.join(save_to, 'reference_abundances.txt'), np.asarray(amaps))
  573. print(f"ℹ️ Skipped spatial abundance map plotting for non-imaging input: {filename}")
  574. print(f"✅ Unmixing data saved in {analysis_data_folder}/")

Deep_learning_enhanced_Raman_microspectroscopy.ipynb at commit 5eaa744, under Apache-2.0 · at the source

Overview

  1. Department of Computing, and UKRI Centre for Doctoral Training in AI for Healthcare, Imperial College London, London, UK SW7 2AZ
  2. Department of Materials, Department of Bioengineering, and Institute of Biomedical Engineering, Imperial College London, London, UK SW7 2AZ
  3. Department of Physiology, Anatomy and Genetics, Department of Engineering Science, and Kavli Institute for Nanoscience Discovery, University of Oxford, Oxford, UK OX1 3QU
  4. Department of Engineering Science, University of Oxford, Oxford, UK OX1 3PJ
  5. Department of Mathematics, Imperial College London, London, UK SW7 2AZ
Institutions: University of Oxford (United Kingdom); Imperial College London (United Kingdom)
Journal: Science advances, volume 12, issue 33, article eaec5080
Dates: received 26 September 2025; accepted 6 July 2026; published online 12 August 2026; in print August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1126/sciadv.aec5080 · PMID 42585326 · PMCID PMC13464642 · OpenAlex W4414720004
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), developmental (subfield)
Methods: Smoothing, state filtering, decompositions, Preprocessing, Machine learning
MeSH: Deep Learning*, Neurons*, Organoids*, Spectrum Analysis, Raman*, Animals, Humans, Imaging, Three-Dimensional (* major topic)
Topic: Spectroscopy Techniques in Biomedical and Chemical Research (Biophysics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Advanced Research and Invention Agency (SCNI-PR01-P11); UK Research and Innovation (10145749); Royal Academy of Engineering (CiET2021∖94); Norges Forskningsråd (#262613); HORIZON EUROPE Framework Programme (101071203); Engineering and Physical Sciences Research Council (EP/T027258/1, EP/N014529/1, EP/P00114/1, EP/T020792/1, EP/Z534870/1)
Citations: not cited yet (Europe PMC); 98 references in the paper
Research resources: Human embryonic stem cells RRID:CVCL_9773

Abstract

Three-dimensional (3D) organoids have emerged as powerful models for studying human development, disease, and drug response in vitro. Yet, their analysis remains constrained by standard imaging and characterization techniques, which are invasive, require exogenous labeling, and offer limited multiplexing. Here, we present a noninvasive, label-free imaging platform that integrates Raman microspectroscopy with deep learning–based hyperspectral unmixing for unsupervised, spatially resolved biochemical analysis of neural organoids. Our approach enables 2D and 3D mapping of cellular and subcellular structures in both cryosectioned and intact organoids, achieving improved imaging accuracy and robustness compared to conventional methods for hyperspectral analysis. Using our platform, we demonstrate volumetric imaging of a neural rosette within a neural organoid and interrogate changes in biochemical composition during early developmental stages in intact neural organoids, revealing spatiotemporal variations in lipids, proteins, and nucleic acids. This work establishes a versatile framework for high-content, label-free (bio)chemical phenotyping with broad applications in organoid research and beyond.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

Zenodo 18661874

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 7 files
Software Heritage: not checked
Found in: “Data, code, and materials availability:”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)

barahona-research-group/dl-raman

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 5eaa7445ad4243646de9c31c1912c662fcbdb499, 13 August 2026
Languages: Python (6), Jupyter (1)
Size: 13 files, 7 scripts
Software Heritage: not archived
Found in: “Data, code, and materials availability:”
Holds: README, license file, environment (requirements.txt), 1 notebook
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (4 files), TensorFlow (4 files), Matplotlib (3 files), tifffile (3 files), scikit-learn (2 files), SciPy (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
9 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 7 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data, code, and materials availability

The datasets and source code generated in this study are archived on Zenodo at https://doi.org/10.5281/zenodo.18661874. The source code is also available on GitHub at https://github.com/barahona-research-group/dl-raman and includes graphical user interface (GUI) and command-line interface (CLI) tools for running the pipeline. This study did not generate new materials. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Funding: added Advanced Research and Invention Agency: SCNI-PR01-P11; UK Research and Innovation: 10145749; Royal Academy of Engineering: CiET2021∖94; Norges Forskningsråd: #262613; HORIZON EUROPE Framework Programme: 101071203; Engineering and Physical Sciences Research Council: EP/T027258/1, EP/N014529/1, EP/P00114/1, EP/T020792/1, EP/Z534870/1

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 7 MeSH terms, 68 references, 1 RRID.

Cite

This paper

Georgiev, D., Xie, R., Reumann, D., Zhao, X., Fernández-Galiana, Á., Barahona, M., & Stevens, M. M. (2026). Label-free biochemical imaging and time point analysis of neural organoids via deep learning-enhanced Raman microspectroscopy. Science advances, 12(33), eaec5080. https://doi.org/10.1126/sciadv.aec5080

BibTeX

@article{georgiev2026label,
author = {Georgiev, Dimitar and Xie, Ruoxiao and Reumann, Daniel and Zhao, Xiaoyu and Fernández-Galiana, Álvaro and Barahona, Mauricio and Stevens, Molly M},
title = {{Label-free biochemical imaging and time point analysis of neural organoids via deep learning-enhanced Raman microspectroscopy}},
journal = {Science advances},
year = {2026},
month = aug,
volume = {12},
number = {33},
pages = {eaec5080},
publisher = {American Association for the Advancement of Science},
issn = {2375-2548},
doi = {10.1126/sciadv.aec5080},
url = {https://doi.org/10.1126/sciadv.aec5080},
pmid = {42585326},
pmcid = {PMC13464642}
}

RIS

TY - JOUR
AU - Georgiev, Dimitar
AU - Xie, Ruoxiao
AU - Reumann, Daniel
AU - Zhao, Xiaoyu
AU - Fernández-Galiana, Álvaro
AU - Barahona, Mauricio
AU - Stevens, Molly M
TI - Label-free biochemical imaging and time point analysis of neural organoids via deep learning-enhanced Raman microspectroscopy
T2 - Science advances
J2 - Sci Adv
PY - 2026
DA - 2026/08/12
VL - 12
IS - 33
SP - eaec5080
SN - 2375-2548
PB - American Association for the Advancement of Science
DO - 10.1126/sciadv.aec5080
UR - https://doi.org/10.1126/sciadv.aec5080
LA - en
ER -

CSL-JSON

{
"id": "10.1126/sciadv.aec5080",
"type": "article-journal",
"title": "Label-free biochemical imaging and time point analysis of neural organoids via deep learning-enhanced Raman microspectroscopy",
"container-title": "Science advances",
"author": [
{
"family": "Georgiev",
"given": "Dimitar"
},
{
"family": "Xie",
"given": "Ruoxiao"
},
{
"family": "Reumann",
"given": "Daniel"
},
{
"family": "Zhao",
"given": "Xiaoyu"
},
{
"family": "Fernández-Galiana",
"given": "Álvaro"
},
{
"family": "Barahona",
"given": "Mauricio"
},
{
"family": "Stevens",
"given": "Molly M"
}
],
"container-title-short": "Sci Adv",
"volume": "12",
"issue": "33",
"page": "eaec5080",
"DOI": "10.1126/sciadv.aec5080",
"PMID": "42585326",
"PMCID": "PMC13464642",
"ISSN": "2375-2548",
"publisher": "American Association for the Advancement of Science",
"URL": "https://doi.org/10.1126/sciadv.aec5080",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-74320-5 [code]
Spatial architecture of autism pathogenesis reveals mosaic structural disarray during early development.
Journal: Nature communications
In common: tifffile, SciPy, Matplotlib, 1 other tool, developmental, 3 references
[2] doi:10.7554/elife.98340 [code]
Human adherent cortical organoids in a multi-well format.
Journal: eLife
In common: Matplotlib, NumPy, 5 references
[3] doi:10.64898/2026.03.30.715222 [code]
Synthetic lumen rounding directs neural progenitor division mode
Journal: bioRxiv (preprint)
In common: tifffile, scikit-learn, SciPy, 2 other tools, developmental, 2 references
[4] doi:10.1002/advs.202519893 [code]
NeuroSuite for Long-Term Functional and Structural Studies of Air-Liquid Interface Cerebral Organoids.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: SciPy, Matplotlib, NumPy, developmental, 4 references
[5] doi:10.1016/j.cell.2026.07.023
Metabolic atlas of early human cortex reveals glycolytic remodeling and pentose phosphate pathway control of cell fate transitions.
Journal: Cell
In common: 5 references
[6] doi:10.1038/s41467-026-71458-0 [code]
Early differential impact of MeCP2 mutations on functional networks in Rett syndrome patient-derived human cortical organoids.
Journal: Nature communications
In common: tifffile, scikit-learn, SciPy, 2 other tools, 2 references
[7] doi:10.1371/journal.pbio.3003757 [code]
Cell type-agnostic transcriptomic signatures enable uniform comparisons of neural maturation.
Journal: PLoS biology
In common: SciPy, Matplotlib, NumPy, 4 references
[8] doi:10.1093/stcltm/szag071
Temporal mapping of radiation-induced neural injury and mitigation in human cortical organoids.
Journal: Stem cells translational medicine
In common: 5 references
[9] doi:10.1038/s41467-026-76956-9 [code]
Innervated human cardiac muscle model reveals sympathetic drivers of KCNH2-associated arrhythmias.
Journal: Nature communications
In common: tifffile, SciPy, Matplotlib, 1 other tool, 2 references
[10] doi:10.1371/journal.pcbi.1014571 [code]
SynAPSeg: A novel dataset and image analysis framework for deep learning-based synapse detection and quantification.
Journal: PLoS computational biology
In common: tifffile, TensorFlow, scikit-learn, 3 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.