OSCR

Human neuron activity during an 83-minute movie from 2,286 neurons and 29 patients.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches
  1. [1] § Data Records › Single-neuron data ↔ nwb_generation/write_nwb.py, lines 553–566 · score 0.92 · isi_violations, peak_SNR, waveform_sem, putative pyramidal cell, csc_nr, brain_region
  2. [2] § Methods › Spike sorting and data processing ↔ nwb_generation/stats/collect_mad_values.m, lines 1–47 · score 0.88 · median absolute deviation, 300–3000 Hz, peak signal, noise ratio, Spike sorting, bandpassed
  3. [3] § Technical Validation › Single-neuron data and spike-sorting quality metrics ↔ nwb_generation/stats/collect_mad_values.m, lines 1–47 · score 0.86 · median absolute deviation, 300–3000 Hz, peak signal, noise ratio, spike sorted, bandpassed
  4. [4] § Data Records › Movie stimulus ↔ ML_framework/nwb_loading/nwb_loading.py, lines 161–223 · score 0.84 · machine_learning, corresponding frame, movie bin, binned spike, bin length, movie frame
  5. [5] § Data Records › Movie stimulus ↔ nwb_generation/write_nwb.py, lines 635–682 · score 0.81 · movie_binning_info, machine_learning, movie bin edges, frame onsets, bin length, corresponding frame
  6. [6] § Methods › Spike sorting and data processing ↔ visualization/plot_code/spike_sorting/plot_spike_sorting.ipynb, lines 433–478 · score 0.81 · isolation distances, negative peaks, firing rates, spike sorting, MUs, PCs
  7. [7] § Background & Summary ↔ nwb_generation/session_info.py, lines 3–64 · score 0.77 · University Hospital Bonn, entorhinal cortex, parahippocampal cortex, microwires implanted, hippocampus, amygdala
  8. [8] § Methods › Spike sorting and data processing ↔ nwb_generation/stats/metrics.py, lines 108–131 · score 0.77 · median absolute deviation, noise ratio, Spike sorting, SNR, peak, amplitude
  9. [9] § Methods › Spike sorting and data processing ↔ nwb_generation/stats/metrics.py, lines 54–85 · score 0.76 · inter spike interval, ISI violations, spike events, spike sorted, ratio
  10. [10] § Technical Validation › Single-neuron data and spike-sorting quality metrics ↔ nwb_generation/stats/metrics.py, lines 108–131 · score 0.74 · median absolute deviation, noise ratio, spike sorted, SNR, metrics, peak
  11. [11] § Data Records › Single-neuron data ↔ nwb_generation/write_nwb.py, lines 761–799 · score 0.70 · bundle_index, csc_nr, brain_region, MNI, Hemisphere, localizations
  12. [12] § Background & Summary ↔ nwb_generation/session_info.py, lines 3–64 · score 0.67 · Behnke Fried microwire, entorhinal cortex, parahippocampal cortex, hippocampus, amygdala, EC
  13. [13] § Data Records › Movie annotations ↔ nwb_generation/write_nwb.py, lines 593–615 · score 0.66 · annotations_base, entry_index, label_name, absent, offset, intervals
  14. [14] § Technical Validation › Responses of individual neurons to movie features ↔ visualization/plot_code/decoding_results/decoding_results/notebooks/decoding_performance.ipynb, lines 26–70 · score 0.66 · McKenzie, Scene Cuts, Camera Cut, Tom, Indoor, Summer
  15. [15] § Data Records › Movie annotations ↔ ML_framework/nwb_loading/nwb_loading.py, lines 161–223 · score 0.55 · movie_annotations_indicator_functions, machine_learning, raw, NWB, Summer, frames
  16. [16] § Methods › Task and stimulus ↔ src/epiphyte/preprocessing/data_preprocessing/data_utils.py, lines 277–374 · score 0.55 · neural recording system, movie frame, DAQ, linear, log, timestamp

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,053 lines · 46 KB · no license · 4 matches

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. """
  4. Code for structuring and saving the file-system-based movies data into the NWB format.
  5. Included as an example -- will not run without access to original server with clinical data.
  6. Author: Alana Darcher, Uniklinikum Bonn
  7. Date: 8 April 2025
  8. License: MIT License
  9. Dependencies: numpy, pandas, pynwb, mat73, locals
  10. """
  11. import logging
  12. import shutil
  13. from collections import Counter
  14. from pathlib import Path
  15. from datetime import datetime
  16. from dateutil import tz
  17. from uuid import uuid4
  18. import re
  19. import numpy as np
  20. import pandas as pd
  21. from scipy.stats import sem
  22. from pynwb import NWBHDF5IO, NWBFile
  23. from pynwb.file import Subject, ProcessingModule
  24. from pynwb.misc import Units, TimeSeries
  25. from pynwb.core import DynamicTable, VectorData
  26. from pynwb.epoch import TimeIntervals
  27. import mat73
  28. from session_info import *
  29. from stats.cell_type import SpikeWidth
  30. from stats.metrics import calc_cv2, calculate_snr, count_isi_violations
  31. from nwb_generation.utils.data_io import *
  32. from utils.process_labels import *
  33. class SessionInfo:
  34. """
  35. Manages session info for a given patient.
  36. """
  37. def __init__(self, subject_dir):
  38. self.subject_dir = subject_dir
  39. self.session_info = self._load_session_info()
  40. self.session_start_time = self._get_session_start_time()
  41. def _load_session_info(self):
  42. session_info_path = self.subject_dir / "session_info.npy"
  43. return np.load(session_info_path, allow_pickle=True).item()
  44. def _get_session_start_time(self):
  45. session_time = f"{self.session_info['date']}"
  46. session_start_time = datetime.strptime(session_time, '%Y-%m-%d',)
  47. # For privacy reasons, we only include the year. All other time values are set to a common value.
  48. session_start_time = datetime.strptime(f"{session_start_time.year}-1-1_12:00:00", '%Y-%m-%d_%H:%M:%S')
  49. return session_start_time.replace(tzinfo=tz.gettz('Europe/Berlin'))
  50. class SpikeInfo:
  51. def __init__(self, subject_dir):
  52. self.subject_dir = subject_dir
  53. self.names_cscs = self._load_cscs()
  54. self.cscs_with_units = self._filter_cscs()
  55. def _load_cscs(self):
  56. """
  57. Loads and sorts the names of the unit data files from the refractored spiking_data directory.
  58. Uses natural sorting/human readable sorting.
  59. Filenames have the following pattern: CSC6_SUA1.npy
  60. """
  61. filepaths_cscs = (self.subject_dir / "spiking_data").glob(f"CSC*")
  62. names_cscs = [p.name for p in filepaths_cscs]
  63. return sorted(names_cscs, key=natural_keys)
  64. def _filter_cscs(self):
  65. """
  66. Grab the channel numbers from the filenames and return all unique numbers.
  67. Because the CSC ids are pulled from the spike files, by definition this list will only
  68. include channels with valid units.
  69. """
  70. return np.unique([int(re.findall(r'-?\d+', f)[0]) for f in self.names_cscs])
  71. class WatchlogInfo:
  72. def __init__(self, subject_dir, pat_id):
  73. self.subject_dir = subject_dir
  74. self.pat_id = pat_id
  75. self.clean_wl = self._load_clean_wl()
  76. self.clean_rec = self._load_clean_rec()
  77. self.raw_wl = self._load_raw_wl()
  78. self.raw_rec = self._load_raw_rec()
  79. self.run_check()
  80. def _load_clean_wl(self):
  81. path = Path(self.subject_dir, "cleaned_watchlog", f"{self.pat_id}_pts.npy")
  82. return np.load(path, allow_pickle=True)
  83. def _load_clean_rec(self):
  84. path = Path(self.subject_dir, "cleaned_watchlog", f"{self.pat_id}_rec.npy")
  85. return np.load(path, allow_pickle=True)
  86. def _load_raw_wl(self):
  87. path = Path(self.subject_dir, "watchlogs", f"raw_{self.pat_id}_pts.npy")
  88. return np.load(path, allow_pickle=True)
  89. def _load_raw_rec(self):
  90. path = Path(self.subject_dir, "watchlogs", f"raw_{self.pat_id}_rec.npy")
  91. return np.load(path, allow_pickle=True) / 1000 # convert to milliseconds
  92. def run_check(self):
  93. if self.clean_rec[0] != self.raw_rec[0]:
  94. print(f"cleaned onset {self.clean_rec[0]}")
  95. print(f"raw onset {self.raw_rec[0]}")
  96. assert self.clean_rec[0] >= self.raw_rec[0]
  97. class MovieBinningData:
  98. def __init__(self, data_dir):
  99. self.data_dir = data_dir
  100. self.bin_lengths = [40, 100, 200, 500, 1000] # multiples of the frame rate (0.04s)
  101. self.actual_bin_length = {"40": 40, "100": 80, "200": 200, "500": 480, "1000": 1000} # hack to get around filenames using other numbers
  102. self.df = self.build_df()
  103. def build_df(self):
  104. bin_info_dicts = []
  105. for bl in self.bin_lengths:
  106. edges = np.load(self.data_dir.parent / "movie_edges" / f"edges_movie_{bl}.npy", allow_pickle=True)
  107. frames = np.load(self.data_dir.parent / "movie_edges" / f"relevant_frames_{bl}.npy", allow_pickle=True)
  108. assert len(edges) == len(frames) + 1
  109. bin_info_dicts.append(
  110. {"bin_length": self.actual_bin_length[str(bl)],
  111. "edges": edges,
  112. "frames": frames}
  113. )
  114. return pd.DataFrame(bin_info_dicts)
  115. class MovieData:
  116. def __init__(self, data_dir):
  117. self.data_dir = data_dir
  118. self.movie_pts = self._load_movie_pts()
  119. def _load_movie_pts(self):
  120. """PTS used for analysis, excludes production credits and end credits."""
  121. return np.load(self.data_dir.parent / "movie_edges" / f"pts_movie.npy", allow_pickle=True)
  122. class LocalizationData:
  123. def __init__(self, subject_dir):
  124. self.subject_dir = subject_dir
  125. self.patient_id = int(self.subject_dir.parent.name)
  126. self.processed_localizations = self._load_processed_localizations()
  127. self.processed_localizations = self._rename_hippocampal_electrodes(self.processed_localizations)
  128. self.df = self.build_df()
  129. self.df = self._rename_hippocampal_electrodes(self.df)
  130. self.df_mtl_only = self.build_df(restrict=True)
  131. self.df_mtl_only = self._rename_hippocampal_electrodes(self.df_mtl_only)
  132. self.df_electrode_groups = self.build_df_electrode_groups(manual_localizations=True)
  133. self.df_electrode_groups = self._rename_hippocampal_electrodes(self.df_electrode_groups)
  134. def _load_legui_csv(self):
  135. localizations = pd.read_csv(list((self.subject_dir / "localizations").glob("*MNI*"))[0])
  136. return localizations
  137. def _load_processed_localizations(self):
  138. loaded_df = pd.read_csv(self.subject_dir / "localizations" / f"{self.patient_id}_finalized_localizations.csv")
  139. bundle_ids = [entry[-1] for entry in loaded_df["channel_name_postop"]]
  140. loaded_df["bundle_index"] = bundle_ids
  141. return loaded_df
  142. def _rename_hippocampal_electrodes(self, df):
  143. """Correct for the outdated/original naming scheme of middle hippocampal electrodes (MH).
  144. Replaces MH with PH (posterior hippocampus) for implantation schemes with only two hippocampal electrodes.
  145. Ignores schemes with three hippocampal electrodes. Uses the region_pre_review (original clincial assignment)
  146. to decide if there were two or three hippocampal electrodes.
  147. Args:
  148. df (pd.DataFrame): dataframe containing electrode information
  149. Returns:
  150. df: dataframe with corrected hippocampal region abbreviation
  151. """
  152. regions_in_dataframe = np.unique(self.processed_localizations["region_pre_review"])
  153. if "AH" and "MH" in regions_in_dataframe:
  154. if "PH" not in regions_in_dataframe:
  155. df.replace({"brain_region":"MH"}, {"brain_region":"PH"}, inplace=True)
  156. return df
  157. def build_df(self, restrict=False):
  158. """
  159. Creates the DataFrame used for populating the NWB electrodes and electrode_groups.
  160. Averages across the MNI coordinates for the micros (within coordinate) and sets the average for the micro locations.
  161. Args:
  162. restrict (bool, optional): Restricts to just originally-labeled MTL neurons when true. Defaults to False.
  163. Returns:
  164. pandas dataframe: DataFrame containing localization information
  165. """
  166. if restrict:
  167. region_set = set(zip(self.processed_localizations["region_pre_review"], self.processed_localizations["hemisphere"]))
  168. region_set = {item for item in region_set if any(item[0].startswith(prefix) for prefix in region_restriction.keys())}
  169. else:
  170. region_set = set(zip(self.processed_localizations["region_pre_review"], self.processed_localizations["hemisphere"]))
  171. for p, pair in enumerate(region_set):
  172. # The matching and averaging is done for using the pre-review locations as these are unique within each patient.
  173. # After localization, it's possible to have "duplicated" regions.
  174. region_matches = self.processed_localizations["region_pre_review"] == pair[0]
  175. hemisphere_matches = self.processed_localizations["hemisphere"] == pair[1]
  176. matching_rows = self.processed_localizations[region_matches & hemisphere_matches].copy()
  177. assert len(matching_rows) == 8, f"{len(matching_rows)}"
  178. matching_rows["mni_x_micro"] = [np.mean(matching_rows["mni_x_micro"])] * 8
  179. matching_rows["mni_y_micro"] = [np.mean(matching_rows["mni_y_micro"])] * 8
  180. matching_rows["mni_z_micro"] = [np.mean(matching_rows["mni_z_micro"])] * 8
  181. if p == 0:
  182. df = matching_rows.copy()
  183. else:
  184. df = pd.concat([df, matching_rows], ignore_index=True)
  185. df = df.sort_values(by='csc_nr')
  186. return df
  187. def handle_duplicate_electrode_groups(self, df):
  188. region_hemisphere_combinations = list((zip(df["brain_region"], df["hemisphere"])))
  189. region_hemisphere_set = set(region_hemisphere_combinations)
  190. if len(region_hemisphere_combinations) == len(region_hemisphere_set):
  191. dummy_column = [""] * len(df)
  192. df["unique_identifier_tags_for_duplicated_locations"] = dummy_column
  193. return df
  194. else:
  195. print("Handling duplicated localization groups.")
  196. duplicate_region_dict = Counter(region_hemisphere_combinations)
  197. unique_region_identifier_tags = []
  198. temp_duplicate_list = []
  199. temp_duplicate_list_2 = []
  200. for entry in region_hemisphere_combinations:
  201. if duplicate_region_dict[entry] == 2:
  202. if entry not in temp_duplicate_list:
  203. unique_region_identifier_tags.append("a")
  204. temp_duplicate_list.append(entry)
  205. else:
  206. unique_region_identifier_tags.append("b")
  207. elif duplicate_region_dict[entry] == 3:
  208. if entry not in temp_duplicate_list:
  209. unique_region_identifier_tags.append("a")
  210. temp_duplicate_list.append(entry)
  211. elif entry not in temp_duplicate_list_2:
  212. unique_region_identifier_tags.append("b")
  213. temp_duplicate_list_2.append(entry)
  214. else:
  215. unique_region_identifier_tags.append("c")
  216. elif duplicate_region_dict[entry] == 1:
  217. unique_region_identifier_tags.append('')
  218. assert len(unique_region_identifier_tags) == len(df["brain_region"])
  219. df["unique_identifier_tags_for_duplicated_locations"] = unique_region_identifier_tags
  220. return df
  221. def build_df_electrode_groups(self, manual_localizations=True):
  222. """
  223. Create the DataFrame used for populating the NWB electrode groups entry.
  224. """
  225. if manual_localizations:
  226. tmp = []
  227. for row in self.processed_localizations.itertuples():
  228. name = row.channel_name_postop
  229. region = row.region_pre_review
  230. hemisphere = row.hemisphere
  231. if int(name[-1]) == 1:
  232. print(f" including {hemisphere}{region} in electrode_groups.")
  233. tmp.append(row)
  234. df = pd.DataFrame(tmp)
  235. df = self.handle_duplicate_electrode_groups(df)
  236. else:
  237. tmp = []
  238. for row in self.processed_localizations.itertuples():
  239. name = row.channel_name_postop
  240. region = row.region_pre_review
  241. if region not in region_restriction.keys():
  242. print(f" excluding {region}")
  243. continue
  244. if int(name[-1]) == 1:
  245. tmp.append(row)
  246. df = pd.DataFrame(tmp)
  247. return df
  248. class ChannelData:
  249. def __init__(self, subject_dir, patient_id):
  250. self.subject_dir = subject_dir
  251. self.patient_id = patient_id
  252. self.micro_channels = self._load_channel_info()
  253. def _load_channel_info(self):
  254. micro_channels = pd.read_csv(self.subject_dir / "ChannelNames.txt", delimiter="\t", header=None, names=['Filename',])
  255. micro_channels["csc_nr"] = micro_channels.index + 1
  256. return micro_channels
  257. class MetricsData:
  258. def __init__(self, subject_dir):
  259. self.subject_dir = subject_dir
  260. self.median_abs_deviation_file = self._load_mad_file()
  261. self.mad = self.median_abs_deviation_file["MAD_channels"]
  262. self.iso_distances = self._load_isolation_distances()
  263. self.iso_channels = self._grab_cscs_with_units()
  264. def _load_mad_file(self):
  265. mad_all_channels = mat73.loadmat(self.subject_dir / 'metrics' / 'median_abs_deviations.mat')
  266. if sum(np.isnan(mad_all_channels["MAD_channels"])) != 0:
  267. print(f"Unassigned value for MAD for a channel. {np.where(np.isnan(mad_all_channels['MAD_channels']))}")
  268. return mad_all_channels
  269. def _load_isolation_distances(self):
  270. return np.load(self.subject_dir / "metrics" / "isolation_distances.npy", allow_pickle=True).item()
  271. def _grab_cscs_with_units(self):
  272. return np.unique([int(name.split("_")[0][3:]) for name in self.iso_distances["filenames"]])
  273. class BaseLabelsData:
  274. def __init__(self, subject_dir):
  275. self.subject_dir = subject_dir
  276. self.base_labels = self._load_base_labels()
  277. self.df = self.build_df()
  278. def _load_base_labels(self):
  279. base_labels = Path(self.subject_dir, "base_labels.npy")
  280. base_labels = np.load(base_labels, allow_pickle=True).item()
  281. df_base_labels = pd.DataFrame(base_labels)
  282. return df_base_labels
  283. def build_df(self):
  284. for i, label in enumerate(self.base_labels.itertuples()):
  285. name = label.names
  286. values = label.values
  287. starts = label.starts
  288. stops = label.stops
  289. if name in dataset_labels:
  290. pass
  291. else:
  292. print(f" {name} not included in base labels.")
  293. continue
  294. assert len(values) == len(starts) == len(stops), "Number of label entries is inconsistent."
  295. entry_index = np.arange(1, len(values)+1) # using 1-indexing for compatibility with matlab
  296. for e, entry in enumerate(entry_index):
  297. d_ = {
  298. "names": [name.lower()],
  299. "entry_index": [entry],
  300. "values": [values[e]],
  301. "starts": [starts[e]],
  302. "stops": [stops[e]]
  303. }
  304. df = pd.DataFrame(d_,)
  305. if i == 0 and e == 0:
  306. df_base_labels_db_format = df.copy()
  307. else:
  308. df_base_labels_db_format = pd.concat([df_base_labels_db_format, df])
  309. return df_base_labels_db_format
  310. class AlignedLabelsData:
  311. def __init__(self, subject_dir, rec):
  312. self.subject_dir = subject_dir
  313. self.rec = rec
  314. self.aligned_labels = self._load_aligned_labels()
  315. self.df = self.build_df()
  316. def _load_aligned_labels(self):
  317. aligned_labels = Path(self.subject_dir, "aligned_labels.npy")
  318. aligned_labels = np.load(aligned_labels, allow_pickle=True).item()
  319. df_aligned_labels = pd.DataFrame(aligned_labels)
  320. return df_aligned_labels
  321. def build_df(self):
  322. for i, label in enumerate(self.aligned_labels.itertuples()):
  323. name = label.names
  324. values = label.values
  325. starts = label.starts / 1000 - self.rec[0]
  326. stops = label.stops / 1000 - self.rec[0]
  327. if name in dataset_labels:
  328. pass
  329. else:
  330. print(f" {name} not included in aligned labels.")
  331. continue
  332. assert len(values) == len(starts) == len(stops), "Number of label entries is inconsistent."
  333. assert len(values) > 1, f"Only one entry for label {label}."
  334. entry_index = np.arange(1, len(values)+1) # using 1-indexing for compatibility with matlab
  335. for e, entry in enumerate(entry_index):
  336. d_ = {
  337. "names": [name.lower()],
  338. "entry_index": [entry],
  339. "values": [values[e]],
  340. "starts": [starts[e]],
  341. "stops": [stops[e]]
  342. }
  343. df = pd.DataFrame(d_,)
  344. if i == 0 and e == 0:
  345. df_aligned_labels_db_format = df.copy()
  346. else:
  347. df_aligned_labels_db_format = pd.concat([df_aligned_labels_db_format, df])
  348. return df_aligned_labels_db_format
  349. ##
  350. # Workhorse Class
  351. ##
  352. class PatientNWB:
  353. def __init__(self, patient_id, sub_id, data_dir, save_dir,):
  354. self.pat_id = patient_id
  355. self.data_dir = data_dir
  356. self.save_dir = save_dir
  357. self.sub_id = sub_id
  358. self.subject_dir = Path(self.data_dir, str(self.pat_id), "session_1")
  359. self.logger = self._init_logger()
  360. self.sw = SpikeWidth()
  361. self.session_info = SessionInfo(self.subject_dir)
  362. self.spike_info = SpikeInfo(self.subject_dir)
  363. self.watchlog_info = WatchlogInfo(self.subject_dir, self.pat_id)
  364. self.movie_binning_data = MovieBinningData(self.data_dir)
  365. self.movie_pts = MovieData(self.data_dir).movie_pts
  366. self.localizations_data = LocalizationData(self.subject_dir)
  367. self.channel_data = ChannelData(self.subject_dir, self.pat_id)
  368. self.metrics = MetricsData(self.subject_dir)
  369. self.median_abs_deviations = self.metrics.mad
  370. assert len(self.median_abs_deviations) == len(self.localizations_data.processed_localizations["csc_nr"]), "Discrepancy in number of channels in MAD results."
  371. self.isolation_distances = self.metrics.iso_distances
  372. self.df_isolation_distances = pd.DataFrame(self.isolation_distances)
  373. assert np.all(self.metrics.iso_channels == self.spike_info.cscs_with_units), f"Mismatch in the number of channels with units. Isodist: {np.unique(self.df_isolation_distances['filenames'])}, SpikeInfos: {self.spike_info.cscs_with_units}"
  374. self.base_labels_data = BaseLabelsData(self.subject_dir)
  375. self.aligned_labels_data = AlignedLabelsData(self.subject_dir, self.watchlog_info.raw_rec)
  376. self.sr = 32768. # sampling rate from the NLX amplifier
  377. self.frame_rate = 0.04 # frame rate of the movie file
  378. self.len_movie = 5029.68
  379. self.fps = 25
  380. self.num_frames = int(self.len_movie / self.frame_rate)+1
  381. self.pts_full_movie = [round((x * self.frame_rate), 2) for x in range(0, self.num_frames)] # includes production and end credits.
  382. self.t_r = 3 # refractory period, ms
  383. self.t_c = (1 / self.sr) * 1000 # censored time period
  384. # init nwbfile components
  385. self.nwbfile = self._init_nwbfile()
  386. self.device = self._define_device()
  387. self._init_electrodes()
  388. self._init_units()
  389. self._init_unit_columns()
  390. self.annotations_patient_aligned = self._init_aligned_annotations()
  391. self.annotations_base = self._init_base_annotations()
  392. self._init_processing_module()
  393. #### initialization functions ####
  394. def _init_logger(self,):
  395. timestamp = datetime.now().strftime("%d-%m-%Y-%H-%M")
  396. logging.basicConfig(
  397. format="{asctime} - {levelname} - {message}",
  398. style="{",
  399. datefmt="%Y-%m-%d %H:%M",
  400. level=logging.DEBUG,
  401. filename=f"logs/write_nwb_{timestamp}.log")
  402. logger = logging.getLogger(__name__)
  403. logger.info(f"patient {self.pat_id}, subject id: {self.sub_id}")
  404. logger.info(f"Data location: {self.data_dir}")
  405. logger.info(f"Save location: {self.save_dir}")
  406. logger.info(f"Subject Dir: {self.subject_dir}")
  407. return logger
  408. def _init_nwbfile(self, ):
  409. # note: session_description includes patient-identifying information
  410. # and is excluded from the codebase.
  411. nwbfile = NWBFile(
  412. session_description=f"{session_description_base_text}{self.sub_id}",
  413. identifier=str(uuid4()),
  414. session_start_time=self.session_info.session_start_time,
  415. session_id=f"sub{self.sub_id}",
  416. lab=lab,
  417. institution=institution,
  418. experiment_description=experiment_description,
  419. keywords=keywords,
  420. related_publications=related_publications
  421. )
  422. self.logger.info("NWBfile created.")
  423. return nwbfile
  424. def _define_device(self, ):
  425. device = self.nwbfile.create_device(
  426. name="NeuraLynx ATLAS",
  427. description="Acquisition amplifier",
  428. manufacturer="NeuraLynx"
  429. )
  430. self.logger.info("Device defined.")
  431. return device
  432. def _init_electrodes(self,):
  433. self.nwbfile.add_electrode_column(name="hemisphere", description="Hemisphere in which electrode was implanted")
  434. self.nwbfile.add_electrode_column(name="brain_region", description="Brain region in which electrode was implanted")
  435. self.nwbfile.add_electrode_column(name="bundle_index", description="Microwire channel number within the implanted bundle (unordered)")
  436. self.nwbfile.add_electrode_column(name="csc_nr", description="Continuously sampling channel (CSC) id assigned to the microwire channel")
  437. self.logger.info("Electrode table created.")
  438. def _init_units(self):
  439. self.nwbfile.units = Units(
  440. name="units",
  441. waveform_rate=self.sr,
  442. waveform_unit="microvolts",
  443. description=f"Spike times (milliseconds), waveforms (uV), and associated spike sorting metrics for {self.sub_id}'s movie session. Spike times are reference to the start of the movie, such that a spike time == 0 would indicate a spike at the exact onset of the movie."
  444. )
  445. self.logger.info("Unit table created.")
  446. def _init_unit_columns(self):
  447. self.nwbfile.add_unit_column(name="unit_id", description="unique integer id for each unit within a session")
  448. self.nwbfile.add_unit_column(name="csc_nr", description="id for the electrode channel from which the unit was sorted, links to electrodes table")
  449. self.nwbfile.add_unit_column(name="brain_region", description="brain region from which unit was recorded.")
  450. self.nwbfile.add_unit_column(name="hemisphere", description="hemisphere from which unit was recorded")
  451. self.nwbfile.add_unit_column(name="is_single_unit", description="indicates if a given unit is a putative single neuron or putative multi-unit activity")
  452. self.nwbfile.add_unit_column(name="peak_SNR", description="signal-to-noise ratio calculated via max. amp of the mean spike waveform")
  453. self.nwbfile.add_unit_column(name="isi_violations", description="percentage of refractory-period violations in the spike train")
  454. self.nwbfile.add_unit_column(name="cv2", description="degree of variation in the inter-spike intervals")
  455. self.nwbfile.add_unit_column(name="iso_dist", description="isolation distance of cluster from other spike events on the same channel")
  456. self.nwbfile.add_unit_column(name="cell_type", description="indicates if a given unit is a putative pyramidal cell or interneuron")
  457. self.nwbfile.add_unit_column(name="waveform_sem", description="standard error of the mean waveform across all spikes for the unit")
  458. self.logger.info("Custom unit columns added.")
  459. def _init_aligned_annotations(self):
  460. annotations_patient_aligned = TimeIntervals(
  461. name="annotations_patient_aligned",
  462. description=("Annotation data for the labeled movie features, aligned to the watchlog of a given participant (handles pauses in movie playback). Can be non-monotonic due to patient-led skips in playbak. Time given in milliseconds. "
  463. " Contents: starts: onset of a labeled segment, stops: offset of the labeled segment, values: whether the feature was present in the labeled segment (1) or not (0)")
  464. )
  465. annotations_patient_aligned.add_column(
  466. name="label_name",
  467. description="label name"
  468. )
  469. annotations_patient_aligned.add_column(
  470. name="entry_index",
  471. description="index for the occurrence of the label within the movie"
  472. )
  473. annotations_patient_aligned.add_column(
  474. name="value",
  475. description="indicates if time span corresponds to labeled entity being present (=1) or absent (=0)"
  476. )
  477. self.logger.info("Patient aligned annotations table created.")
  478. return annotations_patient_aligned
  479. def _init_base_annotations(self):
  480. annotations_base = TimeIntervals(
  481. name="annotations_base",
  482. description=("Annotation data for the labeled movie features, aligned to directly to the movie, given in presentation time stamps (seconds)."
  483. " Contents: starts: onset of a labeled segment, stops: offset of the labeled segment, values: whether the feature was present in the labeled segment (1) or not (0)")
  484. )
  485. annotations_base.add_column(
  486. name="label_name",
  487. description="label name"
  488. )
  489. annotations_base.add_column(
  490. name="entry_index",
  491. description="index for the occurrence of the label within the movie"
  492. )
  493. annotations_base.add_column(
  494. name="value",
  495. description="indicates if time span corresponds to labeled entity being present (=1) or absent (=0)"
  496. )
  497. self.logger.info("Base annotations table created.")
  498. return annotations_base
  499. def _init_processing_module(self):
  500. mod = ProcessingModule(name="machine_learning", description="movie annotation and frame information processed for machine learning applications")
  501. self.nwbfile.add_processing_module(mod)
  502. self.logger.info("Processing table added.")
  503. ##### population functions #####
  504. def populate_subject(self, ):
  505. subject = Subject(
  506. subject_id=str(self.sub_id),
  507. age=f"P{patient_ages[self.pat_id]}Y",
  508. sex=patient_sex[self.pat_id],
  509. species="Homo sapiens",
  510. description="BF-implanted epilepsy patient"
  511. )
  512. self.nwbfile.subject = subject
  513. self.logger.info(f"Subject info populated: {subject}")
  514. def populate_movie_binning_data(self):
  515. """
  516. Adds movie binning data to the NWB file using an explicit DynamicTable with ragged columns for
  517. edges and frames. Populates from MovieBinningData.
  518. Note: differs from other stimulus info and tables as this is not pre-defined.
  519. """
  520. df = self.movie_binning_data.df
  521. n = len(df)
  522. # create the table
  523. table = DynamicTable(
  524. name="movie_binning_info",
  525. description=(
  526. "Edges and corresponding frames indices (given as filenames) for binning the data using the frame onsets as edges."
  527. "Each row corresponds to one bin_length."
  528. ),
  529. id=np.arange(n, dtype=np.int64),
  530. columns=[]
  531. )
  532. # add column: bin_length
  533. table.add_column(
  534. name="bin_length",
  535. description="Length of the bin in milliseconds.",
  536. data=np.array(df["bin_length"].values, dtype=np.int64),
  537. )
  538. edges = [np.asarray(x, dtype=np.float64) for x in df["edges"].tolist()]
  539. table.add_column(
  540. name="edges",
  541. description="Bin edges aligned to frame onsets.",
  542. data=edges,
  543. index=True
  544. )
  545. frames = [list(map(str, x)) for x in df["frames"].tolist()]
  546. table.add_column(
  547. name="frames",
  548. description="Frames indices corresponding to the bin edges.",
  549. data=frames,
  550. index=True
  551. )
  552. mod = self.nwbfile.processing["machine_learning"]
  553. mod.add(table)
  554. self.logger.info("Movie binning info populated.")
  555. def populate_movie_annotations(self):
  556. """
  557. Creates and adds a table to the NWB file containing the indicator functions for
  558. each labeled feature.
  559. """
  560. df = self.base_labels_data.df
  561. table = DynamicTable(
  562. name="movie_annotations_indicator_functions",
  563. description=(
  564. "Indicator functions (0=feature absent, 1=feature present) for all labels, spanning the entire film."
  565. "Each value in the function corresponds to a movie frame, monotonically increasing."
  566. )
  567. )
  568. table.add_column(
  569. name="label_name",
  570. description="Label name (string)."
  571. )
  572. table.add_column(
  573. name="indicator_function",
  574. description=f"Indicates presence/absence of labeled feature for all {self.num_frames} frames in the movie.",
  575. index=True
  576. )
  577. for label, sub in df.groupby("names"):
  578. print(f" {label}")
  579. vec = make_label_from_start_stop_times(sub["values"].to_numpy(), sub["starts"].to_numpy(), sub["stops"].to_numpy(), self.pts_full_movie, )
  580. table.add_row(label_name=str(label).lower(), indicator_function=vec)
  581. print(sum(vec))
  582. mod = self.nwbfile.processing["machine_learning"]
  583. mod.add(table)
  584. self.logger.info("Movie anntation table populated.")
  585. def populate_movie_pts(self, ):
  586. """
  587. Adds movie PTS (presentation timestamps) data to the NWB file from dataframe built by MovieData.
  588. """
  589. movie_pts_nwb = TimeSeries(name="movie_frame_times_analysis",
  590. description="onset times for each frame in the movie for the portion used in analysis (excludes production credits and the end credits), given in the presentation time stamps (seconds, referenced to movie start)",
  591. unit="seconds",
  592. rate=self.frame_rate,
  593. data=self.movie_pts)
  594. self.nwbfile.add_stimulus(stimulus=movie_pts_nwb)
  595. movie_pts_nwb = TimeSeries(name="movie_frame_times_base",
  596. description="onset times for each frame in the movie for the entire movie, given in the presentation time stamps (seconds, referenced to movie start)",
  597. unit="seconds",
  598. rate=self.frame_rate,
  599. data=self.pts_full_movie)
  600. self.nwbfile.add_stimulus(stimulus=movie_pts_nwb)
  601. self.logger.info("Movie PTS info populated (frames analyzed and complete set).")
  602. def populate_electrode_groups(self):
  603. """
  604. Populates the NWB file electrode groups table by adding all macros in the recording.
  605. Uses the dataframe built by LocalizationData.df_electrode_groups.
  606. Name includes post-localization brain region and a character tag in case of duplicated regions.
  607. """
  608. self.logger.info("Populating Electrode Groups table:")
  609. for macro in self.localizations_data.df_electrode_groups.itertuples():
  610. self.nwbfile.create_electrode_group(
  611. name=f"{macro.hemisphere}{macro.brain_region}{macro.unique_identifier_tags_for_duplicated_locations}",
  612. description=f"depth electrode {macro.region_pre_review}{macro.hemisphere} (macro contact), {macro.brain_region}{macro.hemisphere} manual review",
  613. device=self.device,
  614. location=f"{hemisphere_full_names[macro.hemisphere]} {region_full_names[macro.brain_region]}"
  615. )
  616. self.logger.info(f" manual location:{macro.hemisphere}{macro.brain_region} / pre-review: {macro.region_pre_review} added.")
  617. def populate_micro_channels(self):
  618. """
  619. Populates the NWB file electrode table by iterating through all micro channels with valid units.
  620. Only includes channels with valid units AND from regions of interest (e.g. excluding temporal FCD electrodes).
  621. Process:
  622. - processed_localizations (LocalizationData) is filtered to include only channels with valid units
  623. - for each channel, the averaged micro location is taken from LocalizationData.df
  624. """
  625. self.logger.info("Populating Electrode table: ")
  626. for row in self.localizations_data.processed_localizations[self.localizations_data.processed_localizations["csc_nr"].isin(self.spike_info.cscs_with_units)].itertuples():
  627. if row.region_pre_review in region_exclusion:
  628. self.logger.info(f" excluding {row.region_pre_review}")
  629. continue
  630. # manual localizations version
  631. original_channel_name = row.channel_name_postop
  632. original_channel_name_first_microwire = f"{original_channel_name[:-1]}1" # using this as a unique identifier
  633. unique_identifier_tags_for_duplicated_locations = self.localizations_data.df_electrode_groups[self.localizations_data.df_electrode_groups["channel_name_postop"] == original_channel_name_first_microwire]["unique_identifier_tags_for_duplicated_locations"].iloc[0]
  634. # take averaged micro locations
  635. x = self.localizations_data.df[self.localizations_data.df["channel_name_postop"] == original_channel_name]["mni_x_micro"].iloc[0]
  636. y = self.localizations_data.df[self.localizations_data.df["channel_name_postop"] == original_channel_name]["mni_y_micro"].iloc[0]
  637. z = self.localizations_data.df[self.localizations_data.df["channel_name_postop"] == original_channel_name]["mni_z_micro"].iloc[0]
  638. electrode_group = self.nwbfile.electrode_groups[f"{row.hemisphere}{row.brain_region}{unique_identifier_tags_for_duplicated_locations}"]
  639. self.nwbfile.add_electrode(
  640. group=electrode_group,
  641. location=f"{hemisphere_full_names[row.hemisphere]} {region_full_names[row.brain_region]}",
  642. hemisphere=row.hemisphere,
  643. brain_region=row.brain_region,
  644. bundle_index=row.bundle_index,
  645. csc_nr=row.csc_nr,
  646. x=x,
  647. y=y,
  648. z=z
  649. )
  650. self.logger.info(f" {row.hemisphere}{row.brain_region}{row.bundle_index} added.")
  651. def restrict_spiketimes_to_movie(self, unit_path):
  652. unit = np.load(unit_path, allow_pickle=True).item()
  653. times_raw = unit["times"] # in milliseconds
  654. amps_raw = unit["amps"]
  655. # restrict to just the movie-related activity
  656. left_idx = np.searchsorted(times_raw, self.watchlog_info.raw_rec[0], side='left')
  657. right_idx = np.searchsorted(times_raw, self.watchlog_info.raw_rec[-1], side='right')
  658. times_filtered = times_raw[left_idx:right_idx] - self.watchlog_info.raw_rec[0] ## reindex times to the start of the movie, instead of amplifier time
  659. amps_filtered = amps_raw[left_idx:right_idx]
  660. # spikes, but not during the movie period.
  661. if times_raw.any() == 1 and times_filtered.any() == 0:
  662. pass
  663. return times_filtered, amps_filtered
  664. def restrict_spiketimes_to_analyzed_movie(self, times_filtered, amps_filtered):
  665. """
  666. Optionally restrict the spiketimes to just those
  667. included in the portion of the movie analyzed.
  668. Args:
  669. times_filtered (np.ndarray): spike times, standardized to the start of the movie
  670. amps_filtered (np.ndarray): amplitudes of each spike in times_filtered
  671. """
  672. pts_analysis_start, pts_analysis_stop = movie_analysis_pts
  673. left_idx = np.searchsorted(times_filtered, pts_analysis_start, side='left')
  674. right_idx = np.searchsorted(times_filtered, pts_analysis_stop, side='right')
  675. times_filtered_analysis = times_filtered[left_idx:right_idx]
  676. amps_filtered_analysis = amps_filtered[left_idx:right_idx]
  677. assert np.min(times_filtered_analysis) >= pts_analysis_start, f"Something went wrong with the spike time filtering. Min spike time is f{np.min(times_filtered_analysis)}, shouldn't be less than f{pts_analysis_start}."
  678. assert np.max(times_filtered_analysis) <= pts_analysis_stop, f"Something went wrong with the spike time filtering. Max spike time is f{np.max(times_filtered_analysis)}, shouldn't be more than f{pts_analysis_stop}."
  679. return times_filtered_analysis, amps_filtered_analysis
  680. def get_unit_type(self, name, waveform):
  681. unit_class_info = name.split("_")[1]
  682. if unit_class_info[0] == "M":
  683. single_unit = 0
  684. cell_type = "multi-unit"
  685. else:
  686. single_unit = 1
  687. if waveform[19] > 0:
  688. spike_width, cell_type = self.sw.calculate_spike_width(waveform)
  689. else:
  690. cell_type = "negative-peak"
  691. unit_index = int(unit_class_info[-1])
  692. return single_unit, cell_type, unit_index
  693. def isi_contamination_wrapper(self, times):
  694. """
  695. Wraps the ISI contamination functions.
  696. Args:
  697. times (np.array): list of spike times for a unit in milliseconds
  698. Returns:
  699. float: percentage of violations
  700. """
  701. r = count_isi_violations(times, self.t_r)
  702. if len(times) <= 1:
  703. p = 0
  704. else:
  705. p = r / (len(times) - 1) # denomniator is number of ISIs
  706. return p * 100
  707. def grab_isolation_distance(self, filename_csc):
  708. print(filename_csc)
  709. fname = filename_csc.split(".")[0]
  710. matches = self.df_isolation_distances.loc[self.df_isolation_distances["filenames"] == fname, "isolation_distances"]
  711. isolation_distance = float(matches.iloc[0]) if not matches.empty else np.nan
  712. return isolation_distance
  713. def populate_units(self):
  714. self.logger.info(f"Populating unit table:")
  715. nm_units_included = 0
  716. nm_units_excluded = 0
  717. all_regions = []
  718. unit_id = 0
  719. for filename_csc in self.spike_info.names_cscs:
  720. unit_path = self.subject_dir / "spiking_data" / filename_csc
  721. csc_nr = int(re.findall(r'-?\d+', filename_csc)[0])
  722. name = filename_csc.split(".")[0]
  723. brain_region = self.localizations_data.processed_localizations[self.localizations_data.processed_localizations["csc_nr"] == csc_nr]["brain_region"]
  724. if not brain_region.empty:
  725. brain_region = brain_region.iloc[0]
  726. else:
  727. AssertionError
  728. hemisphere = self.localizations_data.processed_localizations[self.localizations_data.processed_localizations["csc_nr"] == csc_nr]["hemisphere"]
  729. if not hemisphere.empty:
  730. hemisphere = hemisphere.iloc[0]
  731. else:
  732. AssertionError
  733. # ############# REMOVE BEFORE RELEASE
  734. region_pre_review = self.localizations_data.processed_localizations[self.localizations_data.processed_localizations["csc_nr"] == csc_nr]["region_pre_review"].iloc[0]
  735. # #####################################
  736. if region_pre_review in region_exclusion:
  737. self.logger.info(f" excluding {name}, {hemisphere}{brain_region}, original region {region_pre_review}")
  738. nm_units_excluded += 1
  739. continue
  740. else:
  741. self.logger.info(f" including {name}, {hemisphere}{brain_region}. original region {region_pre_review}")
  742. nm_units_included += 1
  743. all_regions.append(brain_region)
  744. times, amps = self.restrict_spiketimes_to_movie(unit_path)
  745. waveform_mean = np.mean(amps, axis=0)
  746. waveform_sem = sem(amps, axis=0)
  747. assert len(waveform_mean) == len(waveform_sem), "Waveform mean and sem have different lengths."
  748. single_unit, cell_type, _ = self.get_unit_type(name, waveform_mean)
  749. # metrics
  750. channel_mad = self.median_abs_deviations[csc_nr-1]
  751. peak_snr = calculate_snr(waveform_mean, channel_mad)
  752. isi_below = self.isi_contamination_wrapper(times)
  753. cv2 = calc_cv2(times)
  754. iso_dist = self.grab_isolation_distance(filename_csc)
  755. self.nwbfile.add_unit(spike_times=times, unit_id=unit_id, csc_nr=csc_nr, brain_region=brain_region, hemisphere=hemisphere, is_single_unit=bool(single_unit),
  756. peak_SNR=peak_snr, isi_violations=isi_below, cv2=cv2, iso_dist=iso_dist, cell_type=cell_type,
  757. waveform_mean=waveform_mean, waveform_sem=waveform_sem,
  758. )
  759. unit_id += 1
  760. self.logger.info(f"Number of units included for patient {self.pat_id}: {nm_units_included}")
  761. self.logger.info(f"Number of units excluded for patient {self.pat_id}: {nm_units_excluded}")
  762. self.logger.info(f"Regions kept: {Counter(all_regions)}")
  763. def populate_aligned_annotations(self):
  764. for row in self.aligned_labels_data.df.itertuples():
  765. self.annotations_patient_aligned.add_row(start_time=row.starts, stop_time=row.stops,
  766. label_name=row.names, entry_index=row.entry_index, value=row.values)
  767. self.nwbfile.add_stimulus(stimulus=self.annotations_patient_aligned)
  768. self.logger.info(f"Populated patient aligned annotations.")
  769. def populate_base_annotations(self):
  770. for row in self.base_labels_data.df.itertuples():
  771. self.annotations_base.add_row(start_time=row.starts, stop_time=row.stops,
  772. label_name=row.names, entry_index=row.entry_index, value=int(row.values))
  773. self.nwbfile.add_stimulus(stimulus=self.annotations_base)
  774. self.logger.info(f"Populated base annotations.")
  775. def populate_clean_watchlogs(self):
  776. pts_column = VectorData(
  777. name="pts",
  778. data=self.watchlog_info.clean_wl,
  779. description="presentation timestamps (pts, secoonds) for each frame in the movie shown to the patient during the experiment",
  780. )
  781. neural_rectime_column = VectorData(
  782. name="neural_recording_time",
  783. data=self.watchlog_info.clean_rec - self.watchlog_info.raw_rec[0],
  784. description="neural recording system timestamp (milliseconds) for each frame in the movie shown to the patient during the experiment, corresponds to the pts",
  785. )
  786. watchlogs = DynamicTable(
  787. name="cleaned_watchlogs",
  788. description=("Log of movie presentation to participant. Processed: paused portions unified, inconsistencies in frame rate corrected. "
  789. "Lookup table containing: frame times (presentation time stamps, seconds) and the corresponding timestamp from the neural recording system (milliseconds)."),
  790. colnames=[
  791. "pts",
  792. "neural_recording_time",
  793. ],
  794. columns=[
  795. pts_column,
  796. neural_rectime_column,
  797. ],
  798. )
  799. self.nwbfile.add_stimulus(stimulus=watchlogs)
  800. self.logger.info(f"Populated cleaned watchlogs.")
  801. def populate_raw_watchlogs(self):
  802. pts_column = VectorData(
  803. name="pts",
  804. data=self.watchlog_info.raw_wl,
  805. description="presentation timestamps (pts, seconds) for each frame in the movie shown to the patient during the experiment (in seconds, movie play time)",
  806. )
  807. neural_rectime_column = VectorData(
  808. name="neural_recording_time",
  809. data=self.watchlog_info.raw_rec - self.watchlog_info.raw_rec[0],
  810. description="neural recording system timestamp (milliseconds) for each frame in the movie shown to the patient during the experiment, corresponds to the pts",
  811. )
  812. watchlogs = DynamicTable(
  813. name="raw_watchlogs",
  814. description=("Log of movie presentation to participant. Raw watchlog file, for data record purposes. "
  815. "Lookup table containing: frame times (presentation time stamps, seconds) and the corresponding timestamp from the neural recording system (milliseconds)."),
  816. colnames=[
  817. "pts",
  818. "neural_recording_time",
  819. ],
  820. columns=[
  821. pts_column,
  822. neural_rectime_column,
  823. ],
  824. )
  825. self.nwbfile.add_stimulus(stimulus=watchlogs)
  826. self.logger.info(f"Populated raw watchlogs")
  827. def save_nwbfile(self, save_path=None):
  828. if save_path is None:
  829. save_path = Path(self.subject_dir, f"sub{self.sub_id}.nwb")
  830. io = NWBHDF5IO(save_path, mode="w")
  831. io.write(self.nwbfile)
  832. io.close()
  833. complete_dataset_path = Path(self.save_dir, f"sub{self.sub_id}.nwb")
  834. shutil.copy(save_path, complete_dataset_path)
  835. self.logger.info(f"NWB file saved! \n")
  836. def generate_nwb(patient_id, sub_id, data_dir, save_dir):
  837. patNWB = PatientNWB(patient_id=patient_id, sub_id=sub_id, data_dir=data_dir, save_dir=save_dir)
  838. patNWB.populate_subject()
  839. patNWB.populate_movie_binning_data()
  840. patNWB.populate_movie_annotations()
  841. patNWB.populate_movie_pts()
  842. patNWB.populate_electrode_groups()
  843. patNWB.populate_micro_channels()
  844. patNWB.populate_units()
  845. patNWB.populate_aligned_annotations()
  846. patNWB.populate_base_annotations()
  847. patNWB.populate_clean_watchlogs()
  848. patNWB.populate_raw_watchlogs()
  849. patNWB.save_nwbfile()
  850. if __name__ == "__main__":
  851. data_dir = Path("/media/al/refractored_data/patient_data")
  852. save_dir = Path("/media/al/movies_dataset_nwb")
  853. patient_id = 1 #
  854. sub_id = 1
  855. generate_nwb(patient_id=patient_id, sub_id=sub_id, data_dir=data_dir, save_dir=save_dir)

write_nwb.py at commit 59fc242, no license · at the source

Overview

Authors: Alana Darcher1, Franziska Gerken2, Johannes Niediek1,3, Marcel S. Kehl1,4, Thomas P. Reber1,5, Stefanie Liebe1,6,7,8, Laura Nett1, Attila Racz1, Lukas Kunz1, Bernhard Staresina4,9, Rachel Rapp10, Pedro J. Gonçalves10,11,12,13, Ismail Elezi2,14, Valeri Borger15, Rainer Surges1, Laura Leal-Taixé2,16, Jakob H. Macke10,17,18, Florian Mormann1
18 affiliations
  1. Department of Epileptology, University Hospital Bonn,Bonn, Germany
  2. Dynamic Vision and Learning Group, Technical University of Munich,Munich, Germany
  3. Machine Learning Group, Technische Universität Berlin,Berlin, Germany
  4. Department of Experimental Psychology, University of Oxford,Oxford, UK
  5. Faculty of Psychology, UniDistance Suisse,Brig, Switzerland
  6. Department of Epileptology and Neurology, University Hospital Tübingen,Tübingen, Germany
  7. Hertie Institute for Clinical Brain Science, University of Tübingen,Tübingen, Germany
  8. Hertie Institute for Artificial Intelligence for Brain Health, University of Tübingen,Tübingen, Germany
  9. Oxford Centre for Human Brain Activity, Centre for Integrative Neuroimaging, Department of Psychiatry, University of Oxford,Oxford, UK
  10. Machine Learning in Science, Excellence Cluster Machine Learning and Tübingen AI Center, University of Tübingen,Tübingen, Germany
  11. VIB Center for AI and Computational Biology, VIB,Leuven, Belgium
  12. VIB Center for Neuroscience Leuven, VIB,Leuven, Belgium
  13. Departments of Computer Science and Electrical Engineering, KU Leuven,Leuven, Belgium
  14. Present Address: Huawei, London, United Kingdom
  15. Department of Neurosurgery, University Hospital Bonn,Bonn, Germany
  16. Present Address: NVIDIA Srl Italy, Milan, Italy
  17. Empirical Inference, Max Planck Institute for Intelligent Systems,Tübingen, Germany
  18. Present Address: ELLIS Institute Tübingen,Tübingen, Germany
Journal: Scientific data, volume 13, issue 1, article 1105
Dates: received 27 February 2026; accepted 20 July 2026; published online 30 July 2026
Type: Data paper · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41597-026-07955-0 · PMID 42533002 · PMCID PMC13424620 · OpenAlex W7171799126
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), methods / tools (subfield)
Methods: Spectral & time-frequency, Smoothing, state filtering, decompositions, Machine learning, Single-unit activity, calcium imaging
MeSH: Motion Pictures*, Neurons*, Hippocampus, Humans, Machine Learning, Temporal Lobe (* major topic)
Journal subjects: Data Descriptor
Funding: Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) (DeepHumanVision, FKZ 031L0197A-C); Deutsche Forschungsgemeinschaft (DFG) (EXC 2064/1, PN 391 390727645), SFB 1233 (PN 276693517), SFB 1089 (PN 227953431), SPP 2411 (PN 520287829), MO 930/4-2, M 930/15-1)
Citations: cited by 1 paper (Europe PMC); 60 references in the paper

Abstract

Despite a growing trend towards more naturalistic experiments, few single-unit datasets collected during dynamic and naturalistic stimuli have been released, and none using a full-length movie. Here, we present SUMMER (Single Unit activity during a Movie in the human Medial temporal lobe via Electrophysiological Recordings), a dataset containing recordings from 2,286 neurons from the human amygdala, hippocampus, entorhinal cortex, parahippocampal cortex, and neighboring structures during the complete presentation of the commercial film 500 Days of Summer to 29 intracranially implanted patients. We provide a rich set of frame-wise annotations spanning the entirety of the movie’s 83-minute runtime, which systematically label the most salient narrative and visual elements, from characters and locations to camera cuts and main character speech. This Neurodata Without Borders-formatted dataset contains the spike times and corresponding mean waveforms from all recorded neurons, along with demographic information and electrode localizations. For technical validation, we provide spike-sorting metrics and demonstrate the tuning of individual neurons to movie features. We additionally provide decoding results from a machine learning-based pipeline for predicting movie features from neuronal population activity. To facilitate immediate use, we offer the dataset in a machine learning-ready format alongside a codebase demonstrating both neural feature and movie feature prediction tasks. This dataset offers a valuable foundation for exploring how the human brain processes semantic content, particularly in real-world contexts involving dynamic and naturalistic stimuli.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

mormannlab/SUMMER

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 59fc2424615150481494edf3e31d707729d395ea, 6 July 2026
Languages: Python (43), Jupyter (14), MATLAB (3)
Size: 93 files, 60 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, environment (environment_no_builds.yml, ML_framework/requirements-train.txt), 14 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (27 files), Matplotlib (15 files), pandas (12 files), seaborn (11 files), Neurodata Without Borders (PyNWB, MatNWB) (5 files), Nilearn (4 files), SciPy (4 files), PyTorch Lightning (3 files), PyTorch (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
61 files

Zenodo 21165561

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (27 files), Matplotlib (15 files), pandas (12 files), seaborn (11 files), Neurodata Without Borders (PyNWB, MatNWB) (5 files), Nilearn (4 files), SciPy (4 files), PyTorch Lightning (3 files), PyTorch (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
61 files
At the source:

mackelab/epiphyte

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 48632c8cc60d2fa1faa707229fb61e75d3174a17, 12 June 2026
Languages: JavaScript (36), Python (26), Jupyter (9), Shell (1)
Size: 133 files, 72 scripts
Software Heritage: archived
Found in: the references
Holds: README, license file, CITATION.cff, environment (setup.py, docs/tutorials/docker files/docker-compose.yaml), continuous integration, documentation, 9 notebooks
Not found: tests
Tools: NumPy (15 files), pandas (5 files), Matplotlib (2 files), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
74 files

Code availability

All code used to generate NWB files, perform technical validation, synchronize movie stimulus versions, and visualize results is provided in a GitHub repository, https://github.com/mormannlab/SUMMER36. The calculation of the median absolute deviation (normalization factor for obtaining the peak signal-to-noise ratio of a channel43) is implemented in MATLAB. All other code is implemented in Python. We provide a conda environment file containing all used dependencies, as well as instructions for creating and using an environment. The previously published work on this data34 was performed using a relational database built for this dataset59.

Beyond traditional analytical workflows, we prioritize accessibility for the machine learning community and offer the dataset in an ML-ready format. To further lower the barrier to entry, our repository includes several modules for training a simple neural network on two demonstrative tasks. Detailed setup instructions are provided to ensure ease of use and facilitate rapid implementation.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 192 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The complete dataset, encompassing both neural data and stimulus data, is formatted as Neurodata Without Borders (NWB) files. The NWB-formatted dataset is available on the DANDI Archive: 10.48324/dandi.001616/0.260702.082426.

We are not able to directly release the movie stimulus used in this experiment. To facilitate access, we have created a set of Python-based modules (see following Code Availability section) that convert a specific DVD version of the movie, 500 Days of Summer, to the version used in our experiment–DVD: Cine Project (2010 Release), EAN: 4010232049162, ASIN: B0030FXXLK.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 18 authors, 6 MeSH terms, 2 funders, 54 references.

Cite

This paper

Darcher, A., Gerken, F., Niediek, J., Kehl, M. S., Reber, T. P., Liebe, S., Nett, L., Racz, A., Kunz, L., Staresina, B., Rapp, R., Gonçalves, P. J., Elezi, I., Borger, V., Surges, R., Leal-Taixé, L., Macke, J. H., & Mormann, F. (2026). Human neuron activity during an 83-minute movie from 2,286 neurons and 29 patients. Scientific data, 13(1), 1105. https://doi.org/10.1038/s41597-026-07955-0

BibTeX

@article{darcher2026human,
author = {Darcher, Alana and Gerken, Franziska and Niediek, Johannes and Kehl, Marcel S. and Reber, Thomas P. and Liebe, Stefanie and Nett, Laura and Racz, Attila and Kunz, Lukas and Staresina, Bernhard and Rapp, Rachel and Gonçalves, Pedro J. and Elezi, Ismail and Borger, Valeri and Surges, Rainer and Leal-Taixé, Laura and Macke, Jakob H. and Mormann, Florian},
title = {{Human neuron activity during an 83-minute movie from 2,286 neurons and 29 patients}},
journal = {Scientific data},
year = {2026},
month = jul,
volume = {13},
number = {1},
pages = {1105},
publisher = {Nature Publishing Group},
issn = {2052-4463},
doi = {10.1038/s41597-026-07955-0},
url = {https://doi.org/10.1038/s41597-026-07955-0},
pmid = {42533002},
pmcid = {PMC13424620}
}

RIS

TY - JOUR
AU - Darcher, Alana
AU - Gerken, Franziska
AU - Niediek, Johannes
AU - Kehl, Marcel S.
AU - Reber, Thomas P.
AU - Liebe, Stefanie
AU - Nett, Laura
AU - Racz, Attila
AU - Kunz, Lukas
AU - Staresina, Bernhard
AU - Rapp, Rachel
AU - Gonçalves, Pedro J.
AU - Elezi, Ismail
AU - Borger, Valeri
AU - Surges, Rainer
AU - Leal-Taixé, Laura
AU - Macke, Jakob H.
AU - Mormann, Florian
TI - Human neuron activity during an 83-minute movie from 2,286 neurons and 29 patients
T2 - Scientific data
J2 - Sci Data
PY - 2026
DA - 2026/07/30
VL - 13
IS - 1
SP - 1105
SN - 2052-4463
PB - Nature Publishing Group
DO - 10.1038/s41597-026-07955-0
UR - https://doi.org/10.1038/s41597-026-07955-0
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41597-026-07955-0",
"type": "article-journal",
"title": "Human neuron activity during an 83-minute movie from 2,286 neurons and 29 patients",
"container-title": "Scientific data",
"author": [
{
"family": "Darcher",
"given": "Alana"
},
{
"family": "Gerken",
"given": "Franziska"
},
{
"family": "Niediek",
"given": "Johannes"
},
{
"family": "Kehl",
"given": "Marcel S."
},
{
"family": "Reber",
"given": "Thomas P."
},
{
"family": "Liebe",
"given": "Stefanie"
},
{
"family": "Nett",
"given": "Laura"
},
{
"family": "Racz",
"given": "Attila"
},
{
"family": "Kunz",
"given": "Lukas"
},
{
"family": "Staresina",
"given": "Bernhard"
},
{
"family": "Rapp",
"given": "Rachel"
},
{
"family": "Gonçalves",
"given": "Pedro J."
},
{
"family": "Elezi",
"given": "Ismail"
},
{
"family": "Borger",
"given": "Valeri"
},
{
"family": "Surges",
"given": "Rainer"
},
{
"family": "Leal-Taixé",
"given": "Laura"
},
{
"family": "Macke",
"given": "Jakob H."
},
{
"family": "Mormann",
"given": "Florian"
}
],
"container-title-short": "Sci Data",
"volume": "13",
"issue": "1",
"page": "1105",
"DOI": "10.1038/s41597-026-07955-0",
"PMID": "42533002",
"PMCID": "PMC13424620",
"ISSN": "2052-4463",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41597-026-07955-0",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
30
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-76939-w [code]
HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities.
Journal: Nature communications
In common: Neurodata Without Borders (PyNWB, MatNWB), PyTorch Lightning, PyTorch, 5 other tools, methods / tools, 1 reference
[2] doi:10.1016/j.crmeth.2026.101421 [code]
EthoPy provides an accessible platform for reproducible behavioral neuroscience.
Journal: Cell reports methods
In common: Neurodata Without Borders (PyNWB, MatNWB), seaborn, pandas, 3 other tools, methods / tools, 3 references
[3] doi:10.7554/elife.109717 [code]
Retrosplenial cortex enables context-dependent goal-directed sensorimotor transformation.
Journal: eLife
In common: Neurodata Without Borders (PyNWB, MatNWB), seaborn, pandas, 3 other tools, 3 references
[4] doi:10.1016/j.celrep.2026.117420 [code]
Neural population dynamics of direct electrical stimulation of neocortex.
Journal: Cell reports
In common: Neurodata Without Borders (PyNWB, MatNWB), seaborn, pandas, 3 other tools, 3 references
[5] doi:10.1038/s41467-026-71443-7 [code]
Large-scale single-neuron recording in the human cortex using an ultra-flexible electrode array.
Journal: Nature communications
In common: Neurodata Without Borders (PyNWB, MatNWB), 5 references
[6] doi:10.1523/jneurosci.0038-26.2026 [code]
Multidimensional Feature Tuning in Category Selective Areas of Human Visual Cortex.
Journal: The Journal of neuroscience : the official journal of the Society for Neuroscience
In common: PyTorch, seaborn, pandas, 3 other tools, 3 references
[7] doi:10.7554/elife.107933 [code]
Modality-agnostic decoding of vision and language from fMRI.
Journal: eLife
In common: Nilearn, PyTorch, seaborn, 4 other tools, 2 references
[8] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: Neurodata Without Borders (PyNWB, MatNWB), PyTorch, seaborn, 4 other tools, methods / tools, 1 reference
[9] doi:10.1038/s41592-026-03076-z [code]
Neuropixels Opto: combining high-resolution electrophysiology and optogenetics.
Journal: Nature methods
In common: Neurodata Without Borders (PyNWB, MatNWB), PyTorch, pandas, 3 other tools, 2 references
[10] doi:10.1038/s41467-026-73996-z [code]
Genetic architecture of white matter microstructure captured by unsupervised deep representation learning of fractional anisotropy maps.
Journal: Nature communications
In common: PyTorch Lightning, Nilearn, PyTorch, 5 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.