OSCR

HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities.

Code ↔ Paper

40 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 40 matches
  1. [1] § Methods › Benchmarking methods ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 1–65 · score 1.00 · raw waveform variance, repolarization slope, Allen Institute ecephys, waveform metrics convention, handcrafted waveform features, peak channel waveform
  2. [2] § Methods › Classification algorithms ↔ comparison_methods/nemo/scripts/nemo_benchmark_evaluation.py, lines 410–532 · score 0.91 · MLPClassifier, StandardScaler, validation fraction, hidden layer, evaluates embedding, neighbor classification
  3. [3] § Methods › Classification algorithms ↔ comparison_methods/physmap/physmap_script.r, lines 570–657 · score 0.90 · cross validation iteration, maximizes balanced accuracy, selection rules identical, fold local best, L2 normalized, caret
  4. [4] § Methods › Waveform polarity effect analysis ↔ analysis/figure_polarity_supplement.py, lines 357–450 · score 0.89 · Pearson correlations, polarity composition, polarity stratified, waveform polarity, polarity class, CellExplorer
  5. [5] § Methods › Code library details › Data augmentation ↔ hippie/augmentations.py, lines 116–214 · score 0.87 · warp strength, Gaussian smoothing, Gaussian noise, baseline shifts, amplitude scaling, augmentations
  6. [6] § Methods › Classification algorithms ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 244–295 · score 0.87 · MLPClassifier, StandardScaler, validation fraction, hidden layer, training split, MLP probing
  7. [7] § Methods › Code library details › Generative evaluation setup and decoding procedures ↔ analysis/CVAE_only_experiment_latent_flow_crosstech.py, lines 1–51 · score 0.86 · SOMnan, CellExplorer, source_id, cross technology, Pyra, FS
  8. [8] § Methods › Code library details › Generative evaluation setup and decoding procedures ↔ analysis/CVAE_only_experiment_train_hippie.py, lines 53–99 · score 0.82 · prevent posterior collapse, free bits, source_id, weight decay, nats, cVAE
  9. [9] § Methods › Benchmarking methods › Hyperparameter parity across methods ↔ analysis/figure_2_supp_nemo.py, lines 1–59 · score 0.80 · stability tiebreaker rule, hyperparameter sweeps, weight decay, temperature, selection, parity
  10. [10] § Methods › Benchmarking methods › Hyperparameter parity across methods ↔ analysis/make_source_data.py, lines 49–195 · score 0.77 · dimV, HIPPIE hyperparameter, weight decay, cross validation fold, hyperparameter sweeps, PhysMAP
  11. [11] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ analysis/CVAE_only_experiment_plot.py, lines 1–83 · score 0.75 · Cross modal imputation, waveform waterfall, Cross species, PkC_ss, latent space, UMAP
  12. [12] § Results › Brain region classification from single-unit features remains challenging across methods ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 1–65 · score 0.74 · WF RF handcrafted, random forest, MLP probes, Allen Institute, waveform features, component
  13. [13] § Results › Cell-type classification from waveform and spike-timing features ↔ analysis/figure_3_panel_ab_profile.py, lines 1–67 · score 0.72 · auditory cortex, Juxtacellular Mouse S1, somatosensory cortex, silicon probe, brain regions, waveform
  14. [14] § Methods › Code library details › Datasets ↔ data_wrangling_scripts/allen_nwb_to_csv_converter.ipynb, lines 200–299 · score 0.72 · amplitude cutoff, ISI violations, presence ratio, Allen Institute
  15. [15] § Methods › Benchmarking methods ↔ hippie_wf3dacg/acg.py, lines 36–114 · score 0.72 · instantaneous firing rate, lag bins, deciles, window, autocorrelogram, NEMO
  16. [16] § Methods › Benchmarking methods ↔ hippie_wf3dacg/model.py, lines 148–219 · score 0.71 · cross modal contrastive, weight decay, deciles, resolution, window, KL
  17. [17] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ scripts/run_all_figures.sh, lines 43–99 · score 0.71 · Cross modal imputation, Cross species counterfactual, latent interpolation, Linear probe, bar, decoding
  18. [18] § Methods › Benchmarking methods ↔ comparison_methods/physmap/physmap_script.r, lines 144–214 · score 0.71 · distance metric, dimV, Seurat, UMAP, euclidean, nfeatures
  19. [19] § Results › Cell-type classification from waveform and spike-timing features ↔ analysis/figure_3_panel_ab_profile.py, lines 1–67 · score 0.71 · bar charts, auditory cortex, Juxtacellular Mouse S1, somatosensory cortex, row, Figure 3
  20. [20] § Methods › Benchmarking methods ↔ hippie/multimodal_model.py, lines 15–62 · score 0.69 · hyperparameter sweeps, amplitude scaling, cross species, brain region, std, temperature
  21. [21] § Results › Ablation and hyperparameter analysis ↔ analysis/make_source_data.py, lines 49–195 · score 0.69 · Ablation ladder, MLP classification head, architectural variants, error bar, hyperparameter sweep, latent dimensionality
  22. [22] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ examples/cross_dataset_tutorial.ipynb, lines 194–259 · score 0.66 · predicted mouse, shared latent space, PkC_ss, mouse cerebellar, cross species, geometric
  23. [23] § Methods › Code library details › HIPPIE training details ↔ examples/cross_dataset_tutorial.ipynb, lines 194–259 · score 0.66 · cross species transfer, PkC_cs, PkC_ss, MFB, macaque, MLI
  24. [24] § Methods › Statistics and reproducibility ↔ data_wrangling_scripts/allen_nwb_to_csv_converter.ipynb, lines 301–353 · score 0.65 · amplitude cutoff, ISI violation, presence ratio, Allen, spikes
  25. [25] § Methods › Classification algorithms ↔ hippie/inference.py, lines 16–89 · score 0.64 · training fold, L2 normalized, maximizes, iteration, selection, inference
  26. [26] § Results › HIPPIE generalizes to multiple electrophysiological modalities ↔ analysis/figure_supp_umap_clustering.py, lines 49–78 · score 0.64 · macaque cerebellar, mouse cerebellar, PkC_cs, PkC_ss, expert, MFB
  27. [27] § Methods › Benchmarking methods ↔ hippie_wf3dacg/model.py, lines 148–219 · score 0.63 · amplitude scaling, weight decay, std, temperature, noise, probability
  28. [28] § Methods › Data processing pipeline ↔ hippie_nwb_classify.py, lines 196–290 · score 0.62 · quality filtered, spike trains, spike sorting, NWB, Raw, waveforms
  29. [29] § Methods › Code library details › Datasets ↔ analysis/figure_supp_umap_clustering.py, lines 49–78 · score 0.61 · C4 database, PkC_cs, PkC_ss, MFBs, macaques, Cerebellar
  30. [30] § Methods › Code library details › HIPPIE training details ↔ hippie/multimodal_model.py, lines 15–62 · score 0.61 · swept, forces, frozen, internalized, hidden, gradients
  31. [31] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ scripts/run_all_figures.sh, lines 43–99 · score 0.61 · HIPPIE cVAE, cross species, brain regions, imputation, Watson, counterfactual
  32. [32] § Methods › Statistics and reproducibility ↔ analysis/figure_polarity_supplement.py, lines 357–450 · score 0.60 · waveform polarity composition, Pearson correlations, deviation, Balanced accuracy, neurons, folds
  33. [33] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ analysis/CVAE_only_experiment_latent_flow_crosstech.py, lines 1–51 · score 0.60 · HIPPIE latent space, CellExplorer, Cross technology, S1, encodes, ISI
  34. [34] § Methods › Code library details ↔ comparison_methods/nemo/scripts/nemo_benchmark_evaluation.py, lines 410–532 · score 0.59 · KNN evaluation, F1 score, PhysMAP, zero, masks, NEMO
  35. [35] § Methods › Code library details › HIPPIE training details ↔ scripts/train.py, lines 101–240 · score 0.58 · Recording technology covariates, phases, architecture, decoder, error, KL
  36. [36] § Methods › Web application and software stack ↔ examples/extract_embeddings.py, lines 168–218 · score 0.57 · hippie_techcond_v1.ckpt, Hugging Face Hub, downloaded, checkpoint, pretrained, embeddings
  37. [37] § Methods › Web application and software stack ↔ extract_embeddings.py, lines 146–196 · score 0.57 · hippie_techcond_v1.ckpt, Hugging Face Hub, downloaded, checkpoint, pretrained, embeddings
  38. [38] § Methods › Benchmarking methods ↔ scripts/pick_locked_configs.py, lines 1–61 · score 0.57 · locked hyperparameter, L2 normalized, PhysMAP, sweep, bins, scores
  39. [39] § Methods › Code library details › Datasets ↔ data_wrangling_scripts/ibl_one_to_csv_converter.ipynb, lines 232–285 · score 0.56 · IBL quality, presence ratio, brain regions, Map, spikes, probe
  40. [40] § Results › Ablation and hyperparameter analysis ↔ analysis/figure_2_supp_nemo.py, lines 1–59 · score 0.55 · stability tiebreaker selected, lowest, SEM, locked, sweep, hyperparameter

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 638 lines · 27 KB · BSD-3-Clause · 3 matches

  1. #!/usr/bin/env python3
  2. """
  3. WF-RF — handcrafted waveform features + Random Forest baseline.
  4. Referred to as "WF-RF" in the paper (Methods § Baselines). The four core
  5. shape descriptors follow Jia et al. (2019, Journal of Neurophysiology);
  6. the five additional features follow the Allen Institute ecephys waveform
  7. metrics convention.
  8. Implementation details for the single-channel variant: since the benchmark
  9. datasets in this repository expose only the peak-channel waveform, we
  10. extract nine scalar descriptors from that waveform plus PCA components
  11. that retain ≥ 90 % of the raw waveform variance.
  12. Features extracted on the fly from the 90-sample peak-channel waveform:
  13. 1. PT ratio — positive peak / abs(negative trough)
  14. 2. Duration — (peak_idx - trough_idx) / T [signed, normalised]
  15. 3. Halfwidth — fraction of samples above half-max abs amplitude
  16. 4. Repolarisation slope — (peak_val - trough_val) / |peak_idx - trough_idx|
  17. 5. Recovery slope — slope from later extremum to waveform end
  18. 6. Peak amplitude — raw positive peak value
  19. 7. Trough amplitude — raw negative trough value
  20. 8. Total amplitude — peak_val - trough_val
  21. 9. Peak location — index of abs-maximum / T
  22. plus PCA components that retain ≥ 90 % of variance (matching the paper).
  23. All of the above are concatenated into a single feature vector and fed to a
  24. Random Forest (primary), MLP, and Logistic Regression classifier, matching
  25. the multi-probe evaluation protocol used by NEMO.
  26. No neural network training required — pure sklearn, CPU-only.
  27. Output format matches the benchmark standard:
  28. transductive_predictions.csv — (pred, true) — Random Forest
  29. mlp_predictions.csv — (pred, true) — MLP probe
  30. linear_probe_predictions.csv — (pred, true) — Logistic Regression
  31. Usage
  32. -----
  33. python comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py \\
  34. --dataset hull_cell_type \\
  35. --cv-fold 0 \\
  36. --data-root /data/datasets \\
  37. --output-dir /data/results/wf-rf/hull_cell_type/fold_0
  38. """
  39. from __future__ import annotations
  40. import argparse
  41. import glob
  42. import json
  43. import os
  44. import sys
  45. import time
  46. from pathlib import Path
  47. import numpy as np
  48. import pandas as pd
  49. from sklearn.decomposition import PCA
  50. from sklearn.ensemble import RandomForestClassifier
  51. from sklearn.linear_model import LogisticRegression
  52. from sklearn.metrics import balanced_accuracy_score
  53. from sklearn.model_selection import StratifiedKFold
  54. from sklearn.neural_network import MLPClassifier
  55. from sklearn.preprocessing import LabelEncoder, StandardScaler
  56. # ---------------------------------------------------------------------------
  57. # Repo-relative import of the bundled hippie_wf3dacg H5 loader. This is a
  58. # drop-in replacement for the upstream C4 H5 loader — the dict
  59. # returned by load_c4_dataset has identical keys and semantics.
  60. # ---------------------------------------------------------------------------
  61. _REPO = Path(__file__).resolve().parent.parent.parent
  62. if str(_REPO) not in sys.path:
  63. sys.path.insert(0, str(_REPO))
  64. from hippie_wf3dacg.load_c4 import load_c4_dataset # noqa: E402
  65. # ---------------------------------------------------------------------------
  66. # Waveform feature extraction (Buccino 2018 single-channel variant)
  67. # ---------------------------------------------------------------------------
  68. def extract_waveform_features(waveforms: np.ndarray) -> np.ndarray:
  69. """Compute 9 shape descriptors from a batch of peak-channel waveforms.
  70. Args:
  71. waveforms: (N, T) float32 array — peak-channel waveform for each unit.
  72. Returns:
  73. feats: (N, 9) float32 feature array.
  74. Feature index legend
  75. --------------------
  76. 0 PT ratio — peak_val / (|trough_val| + ε)
  77. 1 duration — (peak_idx – trough_idx) / T [signed, normalised]
  78. 2 halfwidth — fraction of samples ≥ half max abs amplitude
  79. 3 repol_slope — (peak_val – trough_val) / |peak_idx – trough_idx|
  80. 4 recovery_slope — slope from later extremum to last sample
  81. 5 peak_val — positive peak amplitude
  82. 6 trough_val — negative trough amplitude
  83. 7 amplitude — peak_val – trough_val
  84. 8 peak_loc — index_of_abs_max / T
  85. """
  86. N, T = waveforms.shape
  87. feats = np.zeros((N, 9), dtype=np.float32)
  88. for i, wvf in enumerate(waveforms):
  89. trough_idx = int(np.argmin(wvf))
  90. peak_idx = int(np.argmax(wvf))
  91. trough_val = float(wvf[trough_idx])
  92. peak_val = float(wvf[peak_idx])
  93. # 0. PT ratio
  94. feats[i, 0] = peak_val / (abs(trough_val) + 1e-10)
  95. # 1. Duration (signed: positive if peak after trough → typical RS cell)
  96. feats[i, 1] = (peak_idx - trough_idx) / T
  97. # 2. Halfwidth — fraction of time spent ≥ half the max absolute amplitude
  98. abs_wvf = np.abs(wvf)
  99. half_max = abs_wvf.max() / 2.0
  100. feats[i, 2] = float(np.sum(abs_wvf >= half_max)) / T
  101. # 3. Repolarisation slope (trough → peak, regardless of order)
  102. dt = abs(peak_idx - trough_idx)
  103. feats[i, 3] = (peak_val - trough_val) / (dt + 1e-10)
  104. # 4. Recovery slope (later of the two extrema → end of waveform)
  105. later = max(trough_idx, peak_idx)
  106. remaining = T - 1 - later
  107. if remaining > 0:
  108. feats[i, 4] = (float(wvf[-1]) - float(wvf[later])) / remaining
  109. # 5–8. Amplitude stats
  110. feats[i, 5] = peak_val
  111. feats[i, 6] = trough_val
  112. feats[i, 7] = peak_val - trough_val
  113. feats[i, 8] = float(np.argmax(abs_wvf)) / T
  114. # Clip any NaN / Inf that could arise from degenerate (zero) waveforms
  115. np.nan_to_num(feats, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
  116. return feats
  117. def build_features(
  118. waveforms: np.ndarray,
  119. pca_variance: float = 0.90,
  120. pca: PCA | None = None,
  121. fit_pca: bool = True,
  122. ) -> tuple[np.ndarray, PCA]:
  123. """Build combined feature matrix: shape descriptors + PCA components.
  124. Replicates Buccino 2018 which combines:
  125. - Extracted waveform features (PT ratio, duration, …)
  126. - PCA on the raw waveform retaining 90 % of explained variance.
  127. Args:
  128. waveforms: (N, T) raw waveform array.
  129. pca_variance: Fraction of variance to retain (default 0.90).
  130. pca: Pre-fitted PCA object. Used when fit_pca=False.
  131. fit_pca: Fit PCA on this batch (True for training data).
  132. Returns:
  133. features: (N, n_shape + n_pca) combined feature matrix.
  134. pca: Fitted PCA object (to apply identically to test data).
  135. """
  136. shape_feats = extract_waveform_features(waveforms)
  137. N, T = waveforms.shape
  138. if fit_pca:
  139. # Upper bound: never more components than min(N-1, T)
  140. n_max = min(N - 1, T)
  141. pca = PCA(n_components=n_max, random_state=42)
  142. pca_all = pca.fit_transform(waveforms)
  143. # Trim to the fewest components that capture ≥ pca_variance
  144. cum_var = np.cumsum(pca.explained_variance_ratio_)
  145. n_keep = int(np.searchsorted(cum_var, pca_variance)) + 1
  146. pca_feats = pca_all[:, :n_keep]
  147. else:
  148. pca_all = pca.transform(waveforms)
  149. cum_var = np.cumsum(pca.explained_variance_ratio_)
  150. n_keep = int(np.searchsorted(cum_var, pca_variance)) + 1
  151. pca_feats = pca_all[:, :n_keep]
  152. return np.concatenate([shape_feats, pca_feats], axis=1), pca
  153. # ---------------------------------------------------------------------------
  154. # Holdout (inductive) split helpers
  155. # ---------------------------------------------------------------------------
  156. def _build_animal_splits(
  157. animal_ids: np.ndarray,
  158. holdout_fold: int,
  159. n_holdout_folds: int,
  160. min_cells: int,
  161. ) -> tuple[set, set]:
  162. """Partition animals into n_holdout_folds groups; return (train_set, test_set).
  163. Replicates the logic in the holdout training script::_build_animal_splits.
  164. Animals with < min_cells neurons are excluded entirely.
  165. """
  166. from collections import Counter
  167. counts = Counter(animal_ids.tolist())
  168. valid_animals = sorted(a for a, c in counts.items() if c >= min_cells and a)
  169. if not valid_animals:
  170. raise ValueError(
  171. f"No animals with ≥ {min_cells} cells. Counts: {dict(counts)}"
  172. )
  173. rng = np.random.default_rng(42)
  174. shuffled = list(rng.permutation(valid_animals))
  175. # Interleave assignment to balance group sizes
  176. fold_groups = [shuffled[i::n_holdout_folds] for i in range(n_holdout_folds)]
  177. test_set = set(fold_groups[holdout_fold])
  178. train_set = set(a for a in valid_animals if a not in test_set)
  179. return train_set, test_set
  180. # ---------------------------------------------------------------------------
  181. # I/O helpers
  182. # ---------------------------------------------------------------------------
  183. def _find_h5(data_root: str, dataset: str) -> str:
  184. """Return the single .h5 file under <data_root>/<dataset>/."""
  185. pattern = os.path.join(data_root, dataset, "*.h5")
  186. hits = sorted(glob.glob(pattern))
  187. if not hits:
  188. raise FileNotFoundError(
  189. f"No .h5 file found at {pattern}. "
  190. "Download the C4 H5 file via `python scripts/download_c4_h5.py "
  191. "--dataset <key> --out-dir <data_root>/<dataset>/` and re-run. "
  192. "See docs/DATASETS.md for the available keys and the C4 access "
  193. "procedure."
  194. )
  195. if len(hits) > 1:
  196. print(f"WARNING: multiple H5 files found; using {hits[0]}", file=sys.stderr)
  197. return hits[0]
  198. # ---------------------------------------------------------------------------
  199. # Probe suite (matches the HIPPIE training script / nemo_benchmark_evaluation.py)
  200. # ---------------------------------------------------------------------------
  201. def _run_probes(
  202. train_feats: np.ndarray,
  203. train_labels: np.ndarray,
  204. test_feats: np.ndarray,
  205. test_labels: np.ndarray,
  206. n_trees: int = 200,
  207. ) -> tuple[np.ndarray, np.ndarray, np.ndarray]:
  208. """Random Forest (primary) + MLP + Logistic Regression probes.
  209. Features are z-scored with a scaler fitted on the training split.
  210. Returns:
  211. rf_preds, mlp_preds, lr_preds — integer arrays aligned with test set.
  212. """
  213. sc = StandardScaler()
  214. Xtr = sc.fit_transform(train_feats)
  215. Xte = sc.transform(test_feats)
  216. # Guard against NaN / Inf after scaling
  217. np.nan_to_num(Xtr, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
  218. np.nan_to_num(Xte, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
  219. # 1. Random Forest — primary classifier (Buccino 2018)
  220. rf = RandomForestClassifier(
  221. n_estimators=n_trees,
  222. max_features="sqrt",
  223. random_state=42,
  224. n_jobs=-1,
  225. )
  226. rf.fit(Xtr, train_labels)
  227. rf_preds = rf.predict(Xte)
  228. print(f" RF BA={balanced_accuracy_score(test_labels, rf_preds):.4f}")
  229. # 2. MLP probe (256→128, same as NEMO)
  230. mlp = MLPClassifier(
  231. hidden_layer_sizes=(256, 128),
  232. max_iter=500,
  233. random_state=42,
  234. early_stopping=True,
  235. validation_fraction=0.1,
  236. )
  237. mlp.fit(Xtr, train_labels)
  238. mlp_preds = mlp.predict(Xte)
  239. print(f" MLP BA={balanced_accuracy_score(test_labels, mlp_preds):.4f}")
  240. # 3. Logistic Regression probe
  241. lr = LogisticRegression(max_iter=1000, random_state=42)
  242. lr.fit(Xtr, train_labels)
  243. lr_preds = lr.predict(Xte)
  244. print(f" LR BA={balanced_accuracy_score(test_labels, lr_preds):.4f}")
  245. return rf_preds, mlp_preds, lr_preds
  246. # ---------------------------------------------------------------------------
  247. # Holdout (inductive) evaluation
  248. # ---------------------------------------------------------------------------
  249. def _run_holdout(data: dict, args: argparse.Namespace) -> int:
  250. """Run Buccino in inductive holdout mode.
  251. Splits neurons by animal_id: trains PCA + classifiers on (N-1)/N animal
  252. groups and evaluates on the held-out group. Mirrors the protocol in
  253. the holdout training script for cross-method consistency.
  254. Outputs holdout_predictions.csv, mlp_predictions.csv,
  255. linear_probe_predictions.csv (all with an animal_id column).
  256. """
  257. N = len(data["labels"])
  258. all_animal_ids = np.array(data.get("animal_ids", [""] * N), dtype=object)
  259. n_unique = len(set(a for a in all_animal_ids if a))
  260. if n_unique == 0:
  261. print(
  262. "ERROR: C4 H5 has no animal_id metadata.\n"
  263. "Rebuild the H5 with the C4-format preprocessing script after the ANIMAL_ID_FIELD update.",
  264. file=sys.stderr,
  265. )
  266. return 1
  267. print(f"Animal IDs: {n_unique} unique animals across {N} neurons")
  268. # ── Split animals into train / test ──────────────────────────────────────
  269. train_animals, test_animals = _build_animal_splits(
  270. all_animal_ids,
  271. holdout_fold=args.holdout_fold,
  272. n_holdout_folds=args.n_holdout_folds,
  273. min_cells=args.min_cells_per_animal,
  274. )
  275. print(f"Holdout fold {args.holdout_fold}: "
  276. f"train={len(train_animals)} animals test={len(test_animals)} animals")
  277. train_mask = np.array([a in train_animals for a in all_animal_ids])
  278. test_mask = np.array([a in test_animals for a in all_animal_ids])
  279. # ── Filter to labeled neurons within each split ──────────────────────────
  280. unlabeled = set(args.unlabeled_strings)
  281. def _labeled_in_mask(mask):
  282. return np.where(
  283. mask & np.array([s not in unlabeled for s in data["str_labels"]])
  284. )[0]
  285. train_idx = _labeled_in_mask(train_mask)
  286. test_idx = _labeled_in_mask(test_mask)
  287. print(f" Labeled train : {len(train_idx)} | Labeled test : {len(test_idx)}")
  288. if len(train_idx) == 0 or len(test_idx) == 0:
  289. print("ERROR: no labeled neurons in train or test split.", file=sys.stderr)
  290. return 1
  291. # ── Label encoding (fit on union so vocabulary is shared) ────────────────
  292. all_str = ([data["str_labels"][i] for i in train_idx] +
  293. [data["str_labels"][i] for i in test_idx])
  294. le = LabelEncoder()
  295. le.fit(all_str)
  296. train_labels_int = le.transform([data["str_labels"][i] for i in train_idx])
  297. test_labels_int = le.transform([data["str_labels"][i] for i in test_idx])
  298. print(f" Classes: {list(le.classes_)}")
  299. # ── Waveform arrays ──────────────────────────────────────────────────────
  300. waveforms_all = data["waveforms"]
  301. if waveforms_all.dtype == object:
  302. lengths = [len(waveforms_all[i]) for i in np.concatenate([train_idx, test_idx])]
  303. T = int(np.median(lengths))
  304. def _to_fixed(idx_arr):
  305. out = np.zeros((len(idx_arr), T), dtype=np.float32)
  306. for j, i in enumerate(idx_arr):
  307. w = waveforms_all[i]
  308. l = min(len(w), T)
  309. out[j, :l] = w[:l]
  310. return out
  311. train_wvf = _to_fixed(train_idx)
  312. test_wvf = _to_fixed(test_idx)
  313. else:
  314. train_wvf = waveforms_all[train_idx].astype(np.float32)
  315. test_wvf = waveforms_all[test_idx].astype(np.float32)
  316. print(f"Waveform shape: train={train_wvf.shape} test={test_wvf.shape}")
  317. # ── Feature extraction (fit PCA on training neurons only) ────────────────
  318. print("\nExtracting waveform features ...")
  319. t_feat = time.time()
  320. train_feats, pca_obj = build_features(
  321. train_wvf, pca_variance=args.pca_variance, fit_pca=True)
  322. test_feats, _ = build_features(
  323. test_wvf, pca_variance=args.pca_variance, pca=pca_obj, fit_pca=False)
  324. feat_time = time.time() - t_feat
  325. print(f"Feature shape: train={train_feats.shape} test={test_feats.shape} "
  326. f"({feat_time:.2f}s)")
  327. # ── Probes ───────────────────────────────────────────────────────────────
  328. print("\nRunning probes ...")
  329. t_probe = time.time()
  330. rf_preds, mlp_preds, lr_preds = _run_probes(
  331. train_feats, train_labels_int,
  332. test_feats, test_labels_int,
  333. n_trees=args.n_trees,
  334. )
  335. probe_time = time.time() - t_probe
  336. print(f"Probes done in {probe_time:.1f}s")
  337. # ── Write prediction CSVs ────────────────────────────────────────────────
  338. test_true_str = le.inverse_transform(test_labels_int)
  339. test_animal_ids_ = all_animal_ids[test_idx]
  340. def _write(preds, fname):
  341. df = pd.DataFrame({
  342. "pred": le.inverse_transform(preds.astype(int)),
  343. "true": test_true_str,
  344. "animal_id": test_animal_ids_,
  345. })
  346. df.to_csv(os.path.join(args.output_dir, fname), index=False)
  347. print(f" → {os.path.join(args.output_dir, fname)}")
  348. _write(rf_preds, "holdout_predictions.csv")
  349. _write(mlp_preds, "mlp_predictions.csv")
  350. _write(lr_preds, "linear_probe_predictions.csv")
  351. with open(os.path.join(args.output_dir, "animal_split.json"), "w") as f:
  352. json.dump({
  353. "holdout_fold": args.holdout_fold,
  354. "n_holdout_folds": args.n_holdout_folds,
  355. "train_animals": sorted(train_animals),
  356. "test_animals": sorted(test_animals),
  357. "min_cells_per_animal": args.min_cells_per_animal,
  358. "pca_variance": args.pca_variance,
  359. "n_pca_components": int(
  360. np.searchsorted(
  361. np.cumsum(pca_obj.explained_variance_ratio_),
  362. args.pca_variance,
  363. ) + 1
  364. ),
  365. "n_trees": args.n_trees,
  366. "n_features": int(train_feats.shape[1]),
  367. }, f, indent=2)
  368. total = feat_time + probe_time
  369. print(f"\nDone in {total:.1f}s (feat={feat_time:.1f}s probe={probe_time:.1f}s)")
  370. print(f"Results written to {args.output_dir}")
  371. return 0
  372. # ---------------------------------------------------------------------------
  373. # Main
  374. # ---------------------------------------------------------------------------
  375. def main() -> int:
  376. parser = argparse.ArgumentParser(
  377. description="Buccino 2018 single-channel waveform feature baseline.",
  378. formatter_class=argparse.ArgumentDefaultsHelpFormatter,
  379. )
  380. parser.add_argument("--dataset", required=True,
  381. help="Dataset name, e.g. hull_cell_type")
  382. parser.add_argument("--cv-fold", type=int, default=0,
  383. help="CV fold index (0 to n-cv-folds-1)")
  384. parser.add_argument("--n-cv-folds", type=int, default=5)
  385. parser.add_argument("--data-root", default="./datasets",
  386. help="Root directory containing downloaded H5 files")
  387. parser.add_argument("--output-dir", required=True,
  388. help="Write prediction CSVs and cv_split.json here")
  389. parser.add_argument("--pca-variance", type=float, default=0.90,
  390. help="Fraction of variance retained by PCA (Buccino 2018: 0.90)")
  391. parser.add_argument("--n-trees", type=int, default=200,
  392. help="Number of trees in the Random Forest")
  393. parser.add_argument("--unlabeled-strings", nargs="+",
  394. default=["", "unlabeled"],
  395. help="str_label values treated as unlabeled/excluded from probes")
  396. # ── Holdout (inductive) mode ──────────────────────────────────────────────
  397. parser.add_argument("--holdout-fold", type=int, default=None,
  398. help="Enable holdout mode: index of animal fold to hold out "
  399. "(0 to n-holdout-folds-1). When set, --cv-fold is ignored.")
  400. parser.add_argument("--n-holdout-folds", type=int, default=5,
  401. help="Number of animal holdout folds (default: 5)")
  402. parser.add_argument("--min-cells-per-animal", type=int, default=50,
  403. help="Minimum cells per animal to include in holdout splits "
  404. "(default: 50, matching HIPPIE)")
  405. args = parser.parse_args()
  406. os.makedirs(args.output_dir, exist_ok=True)
  407. mode = "holdout" if args.holdout_fold is not None else "transductive"
  408. print("=" * 60)
  409. print("Buccino 2018 Waveform Feature Baseline")
  410. print(f" dataset : {args.dataset}")
  411. print(f" mode : {mode}")
  412. if mode == "holdout":
  413. print(f" holdout_fold : {args.holdout_fold}/{args.n_holdout_folds}")
  414. print(f" min_cells : {args.min_cells_per_animal}")
  415. else:
  416. print(f" cv_fold : {args.cv_fold}/{args.n_cv_folds}")
  417. print(f" pca_variance : {args.pca_variance}")
  418. print(f" n_trees : {args.n_trees}")
  419. print("=" * 60)
  420. # ── 1. Load data from C4 H5 format ─────────────────────────────────
  421. t_load = time.time()
  422. h5_path = _find_h5(args.data_root, args.dataset)
  423. print(f"\nLoading {h5_path} ...")
  424. data = load_c4_dataset(h5_path, args.dataset, source_id=0)
  425. N = len(data["labels"])
  426. load_time = time.time() - t_load
  427. print(f"Loaded {N} neurons in {load_time:.1f}s")
  428. label_counts: dict[str, int] = {}
  429. for s in data["str_labels"]:
  430. label_counts[s] = label_counts.get(s, 0) + 1
  431. print(f"Label distribution: {label_counts}")
  432. # ── 2. Dispatch to holdout mode if requested ─────────────────────────────
  433. if args.holdout_fold is not None:
  434. return _run_holdout(data, args)
  435. # ── 3. (Transductive) Identify labeled neurons ───────────────────────────
  436. unlabeled = set(args.unlabeled_strings)
  437. is_labeled = np.array([s not in unlabeled for s in data["str_labels"]])
  438. labeled_idx = np.where(is_labeled)[0]
  439. print(f"\nLabeled neurons: {len(labeled_idx)} / {N}")
  440. if len(labeled_idx) < 2 * args.n_cv_folds:
  441. print(
  442. f"ERROR: only {len(labeled_idx)} labeled neurons — too few for "
  443. f"{args.n_cv_folds}-fold CV",
  444. file=sys.stderr,
  445. )
  446. return 1
  447. # ── 3. Label encoder + CV split (same seed as NEMO) ──────────
  448. labeled_str = [data["str_labels"][i] for i in labeled_idx]
  449. le = LabelEncoder()
  450. le.fit(labeled_str)
  451. labeled_int = le.transform(labeled_str)
  452. skf = StratifiedKFold(n_splits=args.n_cv_folds, shuffle=True, random_state=42)
  453. splits = list(skf.split(labeled_idx, labeled_int))
  454. train_pos, test_pos = splits[args.cv_fold]
  455. print(f"CV fold {args.cv_fold}: train={len(train_pos)} test={len(test_pos)}")
  456. print(f"Classes: {list(le.classes_)}")
  457. # ── 4. Extract waveforms for labeled neurons ──────────────────────────────
  458. waveforms_all = data["waveforms"]
  459. # Ensure uniform shape — load_c4_dataset may return object array if ragged
  460. waveforms_labeled = waveforms_all[labeled_idx]
  461. if waveforms_labeled.dtype == object:
  462. # Pad / truncate to the most common length
  463. lengths = [len(w) for w in waveforms_labeled]
  464. T = int(np.median(lengths))
  465. wvf_fixed = np.zeros((len(labeled_idx), T), dtype=np.float32)
  466. for j, w in enumerate(waveforms_labeled):
  467. l = min(len(w), T)
  468. wvf_fixed[j, :l] = w[:l]
  469. waveforms_labeled = wvf_fixed
  470. else:
  471. waveforms_labeled = waveforms_labeled.astype(np.float32)
  472. train_wvf = waveforms_labeled[train_pos]
  473. test_wvf = waveforms_labeled[test_pos]
  474. train_labels_int = labeled_int[train_pos]
  475. test_labels_int = labeled_int[test_pos]
  476. print(f"\nWaveform shape: {waveforms_labeled.shape} (dtype {waveforms_labeled.dtype})")
  477. # ── 5. Feature extraction + PCA ──────────────────────────────────────────
  478. print("\nExtracting waveform features ...")
  479. t_feat = time.time()
  480. train_feats, pca_obj = build_features(
  481. train_wvf,
  482. pca_variance=args.pca_variance,
  483. fit_pca=True,
  484. )
  485. test_feats, _ = build_features(
  486. test_wvf,
  487. pca_variance=args.pca_variance,
  488. pca=pca_obj,
  489. fit_pca=False,
  490. )
  491. feat_time = time.time() - t_feat
  492. print(f"Feature shape: train={train_feats.shape} test={test_feats.shape} "
  493. f"({feat_time:.2f}s)")
  494. # ── 6. Probes ─────────────────────────────────────────────────────────────
  495. print("\nRunning probes ...")
  496. t_probe = time.time()
  497. rf_preds, mlp_preds, lr_preds = _run_probes(
  498. train_feats, train_labels_int,
  499. test_feats, test_labels_int,
  500. n_trees=args.n_trees,
  501. )
  502. probe_time = time.time() - t_probe
  503. print(f"Probes done in {probe_time:.1f}s")
  504. # ── 7. Write prediction CSVs ─────────────────────────────────────────────
  505. test_true_str = le.inverse_transform(test_labels_int)
  506. def _write_preds(preds: np.ndarray, fname: str):
  507. df = pd.DataFrame({
  508. "pred": le.inverse_transform(preds.astype(int)),
  509. "true": test_true_str,
  510. })
  511. df.to_csv(os.path.join(args.output_dir, fname), index=False)
  512. print(f" → {os.path.join(args.output_dir, fname)}")
  513. _write_preds(rf_preds, "transductive_predictions.csv")
  514. _write_preds(mlp_preds, "mlp_predictions.csv")
  515. _write_preds(lr_preds, "linear_probe_predictions.csv")
  516. # ── 8. CV split metadata ──────────────────────────────────────────────────
  517. with open(os.path.join(args.output_dir, "cv_split.json"), "w") as f:
  518. json.dump(
  519. {
  520. "train_global_indices": labeled_idx[train_pos].tolist(),
  521. "test_global_indices": labeled_idx[test_pos].tolist(),
  522. "fold": args.cv_fold,
  523. "n_folds": args.n_cv_folds,
  524. "pca_variance": args.pca_variance,
  525. "n_pca_components": int(
  526. np.searchsorted(
  527. np.cumsum(pca_obj.explained_variance_ratio_),
  528. args.pca_variance,
  529. ) + 1
  530. ),
  531. "n_trees": args.n_trees,
  532. "n_features": int(train_feats.shape[1]),
  533. },
  534. f,
  535. indent=2,
  536. )
  537. total_time = load_time + feat_time + probe_time
  538. print(f"\nDone in {total_time:.1f}s total")
  539. print(f"Results written to {args.output_dir}")
  540. return 0
  541. if __name__ == "__main__":
  542. sys.exit(main())

wf_rf_benchmark_evaluation.py at commit 6165dac, under BSD-3-Clause · at the source

Overview

Authors: Jesus Gonzalez-Ferrer1,2,3, Julian Lehrer1,2, Bruno Alvarez-Esteban1,2, Avelina Moreno-Ochando1,2, Hunter E Schweiger1,2,4, Jinghui Geng1,2,5, Luiz F S Eugenio dos Santos1,2,5, Sebastian Hernandez1,2,5, Francisco Reyes1,2,6, Jess L Sevetson1,4, Aidan Schneider7,8, Sofie R Salama1,4, Mircea Teodorescu1,2,3,5, David Haussler1,2,3, Mohammed A Mostajo-Radji1,2
  1. Genomics Institute, University of California Santa Cruz, Santa Cruz, CA USA
  2. Live Cell Biotechnology Discovery Lab, University of California Santa Cruz, Santa Cruz, CA USA
  3. Department of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA USA
  4. Department of Molecular, Cellular and Developmental Biology, University of California Santa Cruz, Santa Cruz, CA USA
  5. Department of Electrical and Computer Engineering, University of California Santa Cruz, Santa Cruz, CA USA
  6. Biotechnology Program, Berkeley City College, Berkeley, CA USA
  7. Department of Statistics & Data Science, Yale University, New Haven, CT USA
  8. Wu Tsai Institute, Yale University, New Haven, CT USA
Institutions: University of California, Santa Cruz (United States); Berkeley City College (United States); Yale University (United States)
Journal: Nature communications, volume 17, issue 1, article 10007
Dates: received 27 February 2026; accepted 11 August 2026; published online 21 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-76939-w · PMID 42764354 · PMCID PMC13590638 · OpenAlex W7171767096
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism), rat (organism), non-human primate (organism), methods / tools (subfield)
Methods: Connectivity, Smoothing, state filtering, decompositions, Machine learning, Evoked potentials, Graphs, Statistics, Single-unit activity, calcium imaging
Keywords: Neurophysiology, Data processing
MeSH: Deep Learning*, Electrophysiological Phenomena*, Models, Neurological*, Neurons*, Action Potentials, Animals, Autoencoder, Electrophysiology, Generative Artificial Intelligence, Macaca, Mice, Rats, Species Specificity (* major topic)
Topic: Neural dynamics and brain function (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: U.S. Department of Health & Human Services | NIH | National Institute of Neurological Disorders and Stroke (NINDS) (U24NS146314); Brain and Behavior Research Foundation (Brain & Behavior Research Foundation) (33184); California Institute for Regenerative Medicine (CIRM) (DISC4-16285, DISC4-16337, DISC4-19334); U.S. Department of Health & Human Services | NIH | National Institute of Mental Health (NIMH) (U24MH132628); NHGRI NIH HHS (RM1 HG011543); NIMH NIH HHS (U24 MH132628); NINDS NIH HHS (U24 NS146314); U.S. Department of Health & Human Services | NIH | National Human Genome Research Institute (NHGRI) (RM1HG011543)
Citations: not cited yet (Europe PMC); 50 references in the paper

Abstract

Neuronal classification from extracellular electrophysiological recordings is challenging due to intrinsic waveform variability, noise, and technical differences across experiments, technologies, and species. We introduce HIPPIE (High-dimensional Interpretation of Physiological Patterns In Intercellular Electrophysiology), a deep learning framework that combines self-supervised pretraining on unlabeled datasets with supervised fine-tuning to classify neurons from extracellular recordings. Using conditional convolutional joint autoencoders, HIPPIE learns technology-adjusted representations of waveforms and spiking dynamics. Here we show, across mouse, rat, and macaque recordings, that HIPPIE classifies cell types competitively with existing methods while additionally supporting generative analyses that discriminative models cannot perform, including counterfactual decoding of electrophysiological signals under changed experimental conditioning, cross-species latent interpolation, and a cross-modal analysis revealing that spike-timing modalities and waveform morphology encode largely independent dimensions of neuronal identity. HIPPIE is available as both a Python package and a coding-free web application, providing a unified framework for multimodal neuronal classification across technologies, experimental conditions, and species.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.

huggingface.co/jesusgf23/hippie

License: apache
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 18ae97027016c2a401c386999ce213176d00b20e, 2 August 2026
Languages: Python (1)
Size: 8 files, 1 script
Software Heritage: not archived
Found in: “Code availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (1 file), pandas (1 file), PyTorch (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
2 files

braingeneers/HIPPIE

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 455598c1d893bf5c7f28c1adf214cd27c4065878, 25 July 2026
Languages: Python (27), Jupyter (5)
Size: 81 files, 32 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (Dockerfile, pyproject.toml), tests, 5 notebooks
Not found: CITATION.cff, continuous integration, documentation
Tools: NumPy (22 files), PyTorch (19 files), pandas (15 files), scikit-learn (6 files), PyTorch Lightning (5 files), Matplotlib (5 files), SciPy (4 files), Neurodata Without Borders (PyNWB, MatNWB) (3 files), AllenSDK (2 files), UMAP (2 files), SpikeInterface (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
34 files

JesusGF1/hippie_benchmarking_release

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 6165daca6da283e5309f1c7840cf39bcb919dcf5, 31 July 2026
Languages: Python (116), Shell (12), R (6), Jupyter (2)
Size: 828 files, 136 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, CITATION.cff, environment (environment.yml, pyproject.toml, requirements.txt, comparison_methods/nemo/environment.yml, comparison_methods/nemo/pyproject.toml, comparison_methods/nemo/setup.py, comparison_methods/physmap/renv.lock), tests, documentation, 2 notebooks
Not found: continuous integration
Tools: NumPy (95 files), pandas (62 files), PyTorch (57 files), scikit-learn (46 files), Matplotlib (42 files), PyTorch Lightning (13 files), h5py (7 files), SciPy (7 files), caret (6 files), seaborn (6 files), tidyverse (6 files), ggplot2 (5 files), imbalanced-learn (5 files), patchwork (4 files), reshape2 (4 files), UMAP (4 files), Seurat (3 files), statsmodels (3 files), Neurodata Without Borders (PyNWB, MatNWB) (2 files), Pillow (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
138 files

EricKenjiLee/PhysMAP_Manuscript

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 08819f45c7742001a760096c64bb00109bd8062e, 11 February 2026
Languages: R (17), Python (3), Jupyter (2), MATLAB (1)
Size: 77 files, 23 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (poetry.lock, pyproject.toml, juxtacellular/requirements.txt), 2 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: caret (9 files), reshape2 (9 files), tidyverse (8 files), NumPy (5 files), pandas (5 files), Matplotlib (4 files), seaborn (3 files), anndata (2 files), ggplot2 (2 files), ggpubr (2 files), reticulate (2 files), Scanpy (2 files), scikit-learn (2 files), SciPy (2 files), Seurat (2 files), cowplot (1 file), PyTorch (1 file), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
24 files

Haansololfp/NEMO_ICLR

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bc173508ad7cb8b664a8673412c6437ba4416ac9, 22 May 2025
Languages: Python (46), Shell (6), Jupyter (1)
Size: 59 files, 53 scripts
Software Heritage: not archived
Found in: the text, “Benchmarking methods”
Holds: README, license file, environment (environment.yml, pyproject.toml, setup.py), 1 notebook
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (38 files), PyTorch (24 files), pandas (16 files), scikit-learn (16 files), Matplotlib (14 files), imbalanced-learn (5 files), seaborn (4 files), SciPy (3 files), statsmodels (3 files), Pillow (1 file), UMAP (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
55 files

Zenodo 21707907

License: BSD-3-Clause
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

Code availability

The user facing code for HIPPIE is available on GitHub: https://github.com/braingeneers/HIPPIE, and the version used in this study has been archived on Zenodo (https://doi.org/10.5281/zenodo.21707907). We have added an extensive tutorial in the shape of a commented Jupyter Notebook. The benchmarking code for the paper release is available in https://github.com/JesusGF1/hippie_benchmarking_release. The pretrained checkpoint (hippie_techcond_v1.ckpt) is publicly available on Hugging Face Hub at https://huggingface.co/Jesusgf23/hippie under the Apache 2.0 license.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 6 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 245 scripts, each with its path and the digest of its content;
  • 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

All data analyzed in this study are previously published and publicly available; no new data or materials were generated.

The Hausser, Hull, and Lisberger cerebellar data used in this study are available in the C4 database (https://www.c4-database.com)23. The Juxtacellular Mouse S132, Extracellular Mouse A131, and CellExplorer46–48 data used in this study are available in the PhysMAP repository (https://github.com/EricKenjiLee/PhysMAP_Manuscript)21. The Watson rat neocortex data used in this study are available in the DANDI Archive under accession code 000041 (https://dandiarchive.org/dandiset/000041)38. The Calvigioni mouse prefrontal cortex data used in this study are available in the DANDI Archive under accession code 000473 (https://dandiarchive.org/dandiset/000473)49. The Ramachandran rat somatosensory cortex data used in this study are available in the DANDI Archive under accession code 000955 (https://dandiarchive.org/dandiset/000955)33. The Allen Institute Visual Coding Neuropixels data used in this study are available in the Allen Brain Observatory via the AllenSDK (https://brain-map.org/our-research/circuits-behavior/visual-coding)41. The International Brain Laboratory Brainwide Map data used in this study are available via the ONE API (https://int-brain-lab.github.io/iblenv/notebooks_external/data_release_brainwidemap.html). Source data are provided with this paper.

The user facing code for HIPPIE is available on GitHub: https://github.com/braingeneers/HIPPIE, and the version used in this study has been archived on Zenodo (https://doi.org/10.5281/zenodo.21707907). We have added an extensive tutorial in the shape of a commented Jupyter Notebook. The benchmarking code for the paper release is available in https://github.com/JesusGF1/hippie_benchmarking_release. The pretrained checkpoint (hippie_techcond_v1.ckpt) is publicly available on Hugging Face Hub at https://huggingface.co/Jesusgf23/hippie under the Apache 2.0 license.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 15 authors, 2 keywords, 13 MeSH terms, 8 funders, 44 references.

Cite

This paper

Gonzalez-Ferrer, J., Lehrer, J., Alvarez-Esteban, B., Moreno-Ochando, A., Schweiger, H. E., Geng, J., Eugenio dos Santos, L. F. S., Hernandez, S., Reyes, F., Sevetson, J. L., Schneider, A., Salama, S. R., Teodorescu, M., Haussler, D., & Mostajo-Radji, M. A. (2026). HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities. Nature communications, 17(1), 10007. https://doi.org/10.1038/s41467-026-76939-w

BibTeX

@article{gonzalezferrer2026hippie,
author = {Gonzalez-Ferrer, Jesus and Lehrer, Julian and Alvarez-Esteban, Bruno and Moreno-Ochando, Avelina and Schweiger, Hunter E and Geng, Jinghui and Eugenio dos Santos, Luiz F S and Hernandez, Sebastian and Reyes, Francisco and Sevetson, Jess L and Schneider, Aidan and Salama, Sofie R and Teodorescu, Mircea and Haussler, David and Mostajo-Radji, Mohammed A},
title = {{HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities}},
journal = {Nature communications},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {10007},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-76939-w},
url = {https://doi.org/10.1038/s41467-026-76939-w},
pmid = {42764354},
pmcid = {PMC13590638}
}

RIS

TY - JOUR
AU - Gonzalez-Ferrer, Jesus
AU - Lehrer, Julian
AU - Alvarez-Esteban, Bruno
AU - Moreno-Ochando, Avelina
AU - Schweiger, Hunter E
AU - Geng, Jinghui
AU - Eugenio dos Santos, Luiz F S
AU - Hernandez, Sebastian
AU - Reyes, Francisco
AU - Sevetson, Jess L
AU - Schneider, Aidan
AU - Salama, Sofie R
AU - Teodorescu, Mircea
AU - Haussler, David
AU - Mostajo-Radji, Mohammed A
TI - HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/08/21
VL - 17
IS - 1
SP - 10007
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-76939-w
UR - https://doi.org/10.1038/s41467-026-76939-w
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-76939-w",
"type": "article-journal",
"title": "HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities",
"container-title": "Nature communications",
"author": [
{
"family": "Gonzalez-Ferrer",
"given": "Jesus"
},
{
"family": "Lehrer",
"given": "Julian"
},
{
"family": "Alvarez-Esteban",
"given": "Bruno"
},
{
"family": "Moreno-Ochando",
"given": "Avelina"
},
{
"family": "Schweiger",
"given": "Hunter E"
},
{
"family": "Geng",
"given": "Jinghui"
},
{
"family": "Eugenio dos Santos",
"given": "Luiz F S"
},
{
"family": "Hernandez",
"given": "Sebastian"
},
{
"family": "Reyes",
"given": "Francisco"
},
{
"family": "Sevetson",
"given": "Jess L"
},
{
"family": "Schneider",
"given": "Aidan"
},
{
"family": "Salama",
"given": "Sofie R"
},
{
"family": "Teodorescu",
"given": "Mircea"
},
{
"family": "Haussler",
"given": "David"
},
{
"family": "Mostajo-Radji",
"given": "Mohammed A"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "10007",
"DOI": "10.1038/s41467-026-76939-w",
"PMID": "42764354",
"PMCID": "PMC13590638",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-76939-w",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
21
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-71331-0 [code]
A multimodal approach for visualizing and identifying electrophysiological cell types in vivo.
Journal: Nature communications
In common: caret, reticulate, UMAP, 15 other tools, mouse, 10 references
[2] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: imbalanced-learn, PyTorch Lightning, caret, 17 other tools, mouse
[3] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: AllenSDK, reticulate, UMAP, 14 other tools, non-human primate, mouse, 1 reference
[4] doi:10.1016/j.xcrm.2026.102651 [code]
Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.
Journal: Cell reports. Medicine
In common: PyTorch Lightning, reticulate, UMAP, 15 other tools
[5] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: UMAP, anndata, Scanpy, 15 other tools, 1 reference
[6] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: reticulate, UMAP, anndata, 15 other tools, mouse
[7] doi:10.7554/elife.108224 [code]
Arrayed single-gene perturbations identify drivers of human anterior neural tube closure.
Journal: eLife
In common: reticulate, anndata, Scanpy, 15 other tools
[8] doi:10.1038/s41586-026-10214-2 [code]
Multidimensional profiling of heterogeneity in supratentorial ependymomas.
Journal: Nature
In common: PyTorch Lightning, reticulate, anndata, 14 other tools, mouse
[9] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: reticulate, UMAP, anndata, 14 other tools, mouse
[10] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: reticulate, UMAP, anndata, 14 other tools, methods / tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.