HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities.
The 40 matches
- [1] § Methods › Benchmarking methods ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 1–65 · score 1.00 · raw waveform variance, repolarization slope, Allen Institute ecephys, waveform metrics convention, handcrafted waveform features, peak channel waveform
- [2] § Methods › Classification algorithms ↔ comparison_methods/nemo/scripts/nemo_benchmark_evaluation.py, lines 410–532 · score 0.91 · MLPClassifier, StandardScaler, validation fraction, hidden layer, evaluates embedding, neighbor classification
- [3] § Methods › Classification algorithms ↔ comparison_methods/physmap/physmap_script.r, lines 570–657 · score 0.90 · cross validation iteration, maximizes balanced accuracy, selection rules identical, fold local best, L2 normalized, caret
- [4] § Methods › Waveform polarity effect analysis ↔ analysis/figure_polarity_supplement.py, lines 357–450 · score 0.89 · Pearson correlations, polarity composition, polarity stratified, waveform polarity, polarity class, CellExplorer
- [5] § Methods › Code library details › Data augmentation ↔ hippie/augmentations.py, lines 116–214 · score 0.87 · warp strength, Gaussian smoothing, Gaussian noise, baseline shifts, amplitude scaling, augmentations
- [6] § Methods › Classification algorithms ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 244–295 · score 0.87 · MLPClassifier, StandardScaler, validation fraction, hidden layer, training split, MLP probing
- [7] § Methods › Code library details › Generative evaluation setup and decoding procedures ↔ analysis/CVAE_only_experiment_latent_flow_crosstech.py, lines 1–51 · score 0.86 · SOMnan, CellExplorer, source_id, cross technology, Pyra, FS
- [8] § Methods › Code library details › Generative evaluation setup and decoding procedures ↔ analysis/CVAE_only_experiment_train_hippie.py, lines 53–99 · score 0.82 · prevent posterior collapse, free bits, source_id, weight decay, nats, cVAE
- [9] § Methods › Benchmarking methods › Hyperparameter parity across methods ↔ analysis/figure_2_supp_nemo.py, lines 1–59 · score 0.80 · stability tiebreaker rule, hyperparameter sweeps, weight decay, temperature, selection, parity
- [10] § Methods › Benchmarking methods › Hyperparameter parity across methods ↔ analysis/make_source_data.py, lines 49–195 · score 0.77 · dimV, HIPPIE hyperparameter, weight decay, cross validation fold, hyperparameter sweeps, PhysMAP
- [11] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ analysis/CVAE_only_experiment_plot.py, lines 1–83 · score 0.75 · Cross modal imputation, waveform waterfall, Cross species, PkC_ss, latent space, UMAP
- [12] § Results › Brain region classification from single-unit features remains challenging across methods ↔ comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py, lines 1–65 · score 0.74 · WF RF handcrafted, random forest, MLP probes, Allen Institute, waveform features, component
- [13] § Results › Cell-type classification from waveform and spike-timing features ↔ analysis/figure_3_panel_ab_profile.py, lines 1–67 · score 0.72 · auditory cortex, Juxtacellular Mouse S1, somatosensory cortex, silicon probe, brain regions, waveform
- [14] § Methods › Code library details › Datasets ↔ data_wrangling_scripts/allen_nwb_to_csv_converter.ipynb, lines 200–299 · score 0.72 · amplitude cutoff, ISI violations, presence ratio, Allen Institute
- [15] § Methods › Benchmarking methods ↔ hippie_wf3dacg/acg.py, lines 36–114 · score 0.72 · instantaneous firing rate, lag bins, deciles, window, autocorrelogram, NEMO
- [16] § Methods › Benchmarking methods ↔ hippie_wf3dacg/model.py, lines 148–219 · score 0.71 · cross modal contrastive, weight decay, deciles, resolution, window, KL
- [17] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ scripts/run_all_figures.sh, lines 43–99 · score 0.71 · Cross modal imputation, Cross species counterfactual, latent interpolation, Linear probe, bar, decoding
- [18] § Methods › Benchmarking methods ↔ comparison_methods/physmap/physmap_script.r, lines 144–214 · score 0.71 · distance metric, dimV, Seurat, UMAP, euclidean, nfeatures
- [19] § Results › Cell-type classification from waveform and spike-timing features ↔ analysis/figure_3_panel_ab_profile.py, lines 1–67 · score 0.71 · bar charts, auditory cortex, Juxtacellular Mouse S1, somatosensory cortex, row, Figure 3
- [20] § Methods › Benchmarking methods ↔ hippie/multimodal_model.py, lines 15–62 · score 0.69 · hyperparameter sweeps, amplitude scaling, cross species, brain region, std, temperature
- [21] § Results › Ablation and hyperparameter analysis ↔ analysis/make_source_data.py, lines 49–195 · score 0.69 · Ablation ladder, MLP classification head, architectural variants, error bar, hyperparameter sweep, latent dimensionality
- [22] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ examples/cross_dataset_tutorial.ipynb, lines 194–259 · score 0.66 · predicted mouse, shared latent space, PkC_ss, mouse cerebellar, cross species, geometric
- [23] § Methods › Code library details › HIPPIE training details ↔ examples/cross_dataset_tutorial.ipynb, lines 194–259 · score 0.66 · cross species transfer, PkC_cs, PkC_ss, MFB, macaque, MLI
- [24] § Methods › Statistics and reproducibility ↔ data_wrangling_scripts/allen_nwb_to_csv_converter.ipynb, lines 301–353 · score 0.65 · amplitude cutoff, ISI violation, presence ratio, Allen, spikes
- [25] § Methods › Classification algorithms ↔ hippie/inference.py, lines 16–89 · score 0.64 · training fold, L2 normalized, maximizes, iteration, selection, inference
- [26] § Results › HIPPIE generalizes to multiple electrophysiological modalities ↔ analysis/figure_supp_umap_clustering.py, lines 49–78 · score 0.64 · macaque cerebellar, mouse cerebellar, PkC_cs, PkC_ss, expert, MFB
- [27] § Methods › Benchmarking methods ↔ hippie_wf3dacg/model.py, lines 148–219 · score 0.63 · amplitude scaling, weight decay, std, temperature, noise, probability
- [28] § Methods › Data processing pipeline ↔ hippie_nwb_classify.py, lines 196–290 · score 0.62 · quality filtered, spike trains, spike sorting, NWB, Raw, waveforms
- [29] § Methods › Code library details › Datasets ↔ analysis/figure_supp_umap_clustering.py, lines 49–78 · score 0.61 · C4 database, PkC_cs, PkC_ss, MFBs, macaques, Cerebellar
- [30] § Methods › Code library details › HIPPIE training details ↔ hippie/multimodal_model.py, lines 15–62 · score 0.61 · swept, forces, frozen, internalized, hidden, gradients
- [31] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ scripts/run_all_figures.sh, lines 43–99 · score 0.61 · HIPPIE cVAE, cross species, brain regions, imputation, Watson, counterfactual
- [32] § Methods › Statistics and reproducibility ↔ analysis/figure_polarity_supplement.py, lines 357–450 · score 0.60 · waveform polarity composition, Pearson correlations, deviation, Balanced accuracy, neurons, folds
- [33] § Results › HIPPIE’s generative architecture enables cross-species and cross-technology analysis ↔ analysis/CVAE_only_experiment_latent_flow_crosstech.py, lines 1–51 · score 0.60 · HIPPIE latent space, CellExplorer, Cross technology, S1, encodes, ISI
- [34] § Methods › Code library details ↔ comparison_methods/nemo/scripts/nemo_benchmark_evaluation.py, lines 410–532 · score 0.59 · KNN evaluation, F1 score, PhysMAP, zero, masks, NEMO
- [35] § Methods › Code library details › HIPPIE training details ↔ scripts/train.py, lines 101–240 · score 0.58 · Recording technology covariates, phases, architecture, decoder, error, KL
- [36] § Methods › Web application and software stack ↔ examples/extract_embeddings.py, lines 168–218 · score 0.57 · hippie_techcond_v1.ckpt, Hugging Face Hub, downloaded, checkpoint, pretrained, embeddings
- [37] § Methods › Web application and software stack ↔ extract_embeddings.py, lines 146–196 · score 0.57 · hippie_techcond_v1.ckpt, Hugging Face Hub, downloaded, checkpoint, pretrained, embeddings
- [38] § Methods › Benchmarking methods ↔ scripts/pick_locked_configs.py, lines 1–61 · score 0.57 · locked hyperparameter, L2 normalized, PhysMAP, sweep, bins, scores
- [39] § Methods › Code library details › Datasets ↔ data_wrangling_scripts/ibl_one_to_csv_converter.ipynb, lines 232–285 · score 0.56 · IBL quality, presence ratio, brain regions, Map, spikes, probe
- [40] § Results › Ablation and hyperparameter analysis ↔ analysis/figure_2_supp_nemo.py, lines 1–59 · score 0.55 · stability tiebreaker selected, lowest, SEM, locked, sweep, hyperparameter
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 638 lines · 27 KB · BSD-3-Clause · 3 matches
- #!/usr/bin/env python3
- """
- WF-RF — handcrafted waveform features + Random Forest baseline.
- Referred to as "WF-RF" in the paper (Methods § Baselines). The four core
- shape descriptors follow Jia et al. (2019, Journal of Neurophysiology);
- the five additional features follow the Allen Institute ecephys waveform
- metrics convention.
- Implementation details for the single-channel variant: since the benchmark
- datasets in this repository expose only the peak-channel waveform, we
- extract nine scalar descriptors from that waveform plus PCA components
- that retain ≥ 90 % of the raw waveform variance.
- Features extracted on the fly from the 90-sample peak-channel waveform:
- 1. PT ratio — positive peak / abs(negative trough)
- 2. Duration — (peak_idx - trough_idx) / T [signed, normalised]
- 3. Halfwidth — fraction of samples above half-max abs amplitude
- 4. Repolarisation slope — (peak_val - trough_val) / |peak_idx - trough_idx|
- 5. Recovery slope — slope from later extremum to waveform end
- 6. Peak amplitude — raw positive peak value
- 7. Trough amplitude — raw negative trough value
- 8. Total amplitude — peak_val - trough_val
- 9. Peak location — index of abs-maximum / T
- plus PCA components that retain ≥ 90 % of variance (matching the paper).
- All of the above are concatenated into a single feature vector and fed to a
- Random Forest (primary), MLP, and Logistic Regression classifier, matching
- the multi-probe evaluation protocol used by NEMO.
- No neural network training required — pure sklearn, CPU-only.
- Output format matches the benchmark standard:
- transductive_predictions.csv — (pred, true) — Random Forest
- mlp_predictions.csv — (pred, true) — MLP probe
- linear_probe_predictions.csv — (pred, true) — Logistic Regression
- Usage
- -----
- python comparison_methods/wf-rf/wf_rf_benchmark_evaluation.py \\
- --dataset hull_cell_type \\
- --cv-fold 0 \\
- --data-root /data/datasets \\
- --output-dir /data/results/wf-rf/hull_cell_type/fold_0
- """
- from __future__ import annotations
- import argparse
- import glob
- import json
- import os
- import sys
- import time
- from pathlib import Path
- import numpy as np
- import pandas as pd
- from sklearn.decomposition import PCA
- from sklearn.ensemble import RandomForestClassifier
- from sklearn.linear_model import LogisticRegression
- from sklearn.metrics import balanced_accuracy_score
- from sklearn.model_selection import StratifiedKFold
- from sklearn.neural_network import MLPClassifier
- from sklearn.preprocessing import LabelEncoder, StandardScaler
- # ---------------------------------------------------------------------------
- # Repo-relative import of the bundled hippie_wf3dacg H5 loader. This is a
- # drop-in replacement for the upstream C4 H5 loader — the dict
- # returned by load_c4_dataset has identical keys and semantics.
- # ---------------------------------------------------------------------------
- _REPO = Path(__file__).resolve().parent.parent.parent
- if str(_REPO) not in sys.path:
- sys.path.insert(0, str(_REPO))
- from hippie_wf3dacg.load_c4 import load_c4_dataset # noqa: E402
- # ---------------------------------------------------------------------------
- # Waveform feature extraction (Buccino 2018 single-channel variant)
- # ---------------------------------------------------------------------------
- def extract_waveform_features(waveforms: np.ndarray) -> np.ndarray:
- """Compute 9 shape descriptors from a batch of peak-channel waveforms.
- Args:
- waveforms: (N, T) float32 array — peak-channel waveform for each unit.
- Returns:
- feats: (N, 9) float32 feature array.
- Feature index legend
- --------------------
- 0 PT ratio — peak_val / (|trough_val| + ε)
- 1 duration — (peak_idx – trough_idx) / T [signed, normalised]
- 2 halfwidth — fraction of samples ≥ half max abs amplitude
- 3 repol_slope — (peak_val – trough_val) / |peak_idx – trough_idx|
- 4 recovery_slope — slope from later extremum to last sample
- 5 peak_val — positive peak amplitude
- 6 trough_val — negative trough amplitude
- 7 amplitude — peak_val – trough_val
- 8 peak_loc — index_of_abs_max / T
- """
- N, T = waveforms.shape
- feats = np.zeros((N, 9), dtype=np.float32)
- for i, wvf in enumerate(waveforms):
- trough_idx = int(np.argmin(wvf))
- peak_idx = int(np.argmax(wvf))
- trough_val = float(wvf[trough_idx])
- peak_val = float(wvf[peak_idx])
- # 0. PT ratio
- feats[i, 0] = peak_val / (abs(trough_val) + 1e-10)
- # 1. Duration (signed: positive if peak after trough → typical RS cell)
- feats[i, 1] = (peak_idx - trough_idx) / T
- # 2. Halfwidth — fraction of time spent ≥ half the max absolute amplitude
- abs_wvf = np.abs(wvf)
- half_max = abs_wvf.max() / 2.0
- feats[i, 2] = float(np.sum(abs_wvf >= half_max)) / T
- # 3. Repolarisation slope (trough → peak, regardless of order)
- dt = abs(peak_idx - trough_idx)
- feats[i, 3] = (peak_val - trough_val) / (dt + 1e-10)
- # 4. Recovery slope (later of the two extrema → end of waveform)
- later = max(trough_idx, peak_idx)
- remaining = T - 1 - later
- if remaining > 0:
- feats[i, 4] = (float(wvf[-1]) - float(wvf[later])) / remaining
- # 5–8. Amplitude stats
- feats[i, 5] = peak_val
- feats[i, 6] = trough_val
- feats[i, 7] = peak_val - trough_val
- feats[i, 8] = float(np.argmax(abs_wvf)) / T
- # Clip any NaN / Inf that could arise from degenerate (zero) waveforms
- np.nan_to_num(feats, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
- return feats
- def build_features(
- waveforms: np.ndarray,
- pca_variance: float = 0.90,
- pca: PCA | None = None,
- fit_pca: bool = True,
- ) -> tuple[np.ndarray, PCA]:
- """Build combined feature matrix: shape descriptors + PCA components.
- Replicates Buccino 2018 which combines:
- - Extracted waveform features (PT ratio, duration, …)
- - PCA on the raw waveform retaining 90 % of explained variance.
- Args:
- waveforms: (N, T) raw waveform array.
- pca_variance: Fraction of variance to retain (default 0.90).
- pca: Pre-fitted PCA object. Used when fit_pca=False.
- fit_pca: Fit PCA on this batch (True for training data).
- Returns:
- features: (N, n_shape + n_pca) combined feature matrix.
- pca: Fitted PCA object (to apply identically to test data).
- """
- shape_feats = extract_waveform_features(waveforms)
- N, T = waveforms.shape
- if fit_pca:
- # Upper bound: never more components than min(N-1, T)
- n_max = min(N - 1, T)
- pca = PCA(n_components=n_max, random_state=42)
- pca_all = pca.fit_transform(waveforms)
- # Trim to the fewest components that capture ≥ pca_variance
- cum_var = np.cumsum(pca.explained_variance_ratio_)
- n_keep = int(np.searchsorted(cum_var, pca_variance)) + 1
- pca_feats = pca_all[:, :n_keep]
- else:
- pca_all = pca.transform(waveforms)
- cum_var = np.cumsum(pca.explained_variance_ratio_)
- n_keep = int(np.searchsorted(cum_var, pca_variance)) + 1
- pca_feats = pca_all[:, :n_keep]
- return np.concatenate([shape_feats, pca_feats], axis=1), pca
- # ---------------------------------------------------------------------------
- # Holdout (inductive) split helpers
- # ---------------------------------------------------------------------------
- def _build_animal_splits(
- animal_ids: np.ndarray,
- holdout_fold: int,
- n_holdout_folds: int,
- min_cells: int,
- ) -> tuple[set, set]:
- """Partition animals into n_holdout_folds groups; return (train_set, test_set).
- Replicates the logic in the holdout training script::_build_animal_splits.
- Animals with < min_cells neurons are excluded entirely.
- """
- from collections import Counter
- counts = Counter(animal_ids.tolist())
- valid_animals = sorted(a for a, c in counts.items() if c >= min_cells and a)
- if not valid_animals:
- raise ValueError(
- f"No animals with ≥ {min_cells} cells. Counts: {dict(counts)}"
- )
- rng = np.random.default_rng(42)
- shuffled = list(rng.permutation(valid_animals))
- # Interleave assignment to balance group sizes
- fold_groups = [shuffled[i::n_holdout_folds] for i in range(n_holdout_folds)]
- test_set = set(fold_groups[holdout_fold])
- train_set = set(a for a in valid_animals if a not in test_set)
- return train_set, test_set
- # ---------------------------------------------------------------------------
- # I/O helpers
- # ---------------------------------------------------------------------------
- def _find_h5(data_root: str, dataset: str) -> str:
- """Return the single .h5 file under <data_root>/<dataset>/."""
- pattern = os.path.join(data_root, dataset, "*.h5")
- hits = sorted(glob.glob(pattern))
- if not hits:
- raise FileNotFoundError(
- f"No .h5 file found at {pattern}. "
- "Download the C4 H5 file via `python scripts/download_c4_h5.py "
- "--dataset <key> --out-dir <data_root>/<dataset>/` and re-run. "
- "See docs/DATASETS.md for the available keys and the C4 access "
- "procedure."
- )
- if len(hits) > 1:
- print(f"WARNING: multiple H5 files found; using {hits[0]}", file=sys.stderr)
- return hits[0]
- # ---------------------------------------------------------------------------
- # Probe suite (matches the HIPPIE training script / nemo_benchmark_evaluation.py)
- # ---------------------------------------------------------------------------
- def _run_probes(
- train_feats: np.ndarray,
- train_labels: np.ndarray,
- test_feats: np.ndarray,
- test_labels: np.ndarray,
- n_trees: int = 200,
- ) -> tuple[np.ndarray, np.ndarray, np.ndarray]:
- """Random Forest (primary) + MLP + Logistic Regression probes.
- Features are z-scored with a scaler fitted on the training split.
- Returns:
- rf_preds, mlp_preds, lr_preds — integer arrays aligned with test set.
- """
- sc = StandardScaler()
- Xtr = sc.fit_transform(train_feats)
- Xte = sc.transform(test_feats)
- # Guard against NaN / Inf after scaling
- np.nan_to_num(Xtr, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
- np.nan_to_num(Xte, nan=0.0, posinf=0.0, neginf=0.0, copy=False)
- # 1. Random Forest — primary classifier (Buccino 2018)
- rf = RandomForestClassifier(
- n_estimators=n_trees,
- max_features="sqrt",
- random_state=42,
- n_jobs=-1,
- )
- rf.fit(Xtr, train_labels)
- rf_preds = rf.predict(Xte)
- print(f" RF BA={balanced_accuracy_score(test_labels, rf_preds):.4f}")
- # 2. MLP probe (256→128, same as NEMO)
- mlp = MLPClassifier(
- hidden_layer_sizes=(256, 128),
- max_iter=500,
- random_state=42,
- early_stopping=True,
- validation_fraction=0.1,
- )
- mlp.fit(Xtr, train_labels)
- mlp_preds = mlp.predict(Xte)
- print(f" MLP BA={balanced_accuracy_score(test_labels, mlp_preds):.4f}")
- # 3. Logistic Regression probe
- lr = LogisticRegression(max_iter=1000, random_state=42)
- lr.fit(Xtr, train_labels)
- lr_preds = lr.predict(Xte)
- print(f" LR BA={balanced_accuracy_score(test_labels, lr_preds):.4f}")
- return rf_preds, mlp_preds, lr_preds
- # ---------------------------------------------------------------------------
- # Holdout (inductive) evaluation
- # ---------------------------------------------------------------------------
- def _run_holdout(data: dict, args: argparse.Namespace) -> int:
- """Run Buccino in inductive holdout mode.
- Splits neurons by animal_id: trains PCA + classifiers on (N-1)/N animal
- groups and evaluates on the held-out group. Mirrors the protocol in
- the holdout training script for cross-method consistency.
- Outputs holdout_predictions.csv, mlp_predictions.csv,
- linear_probe_predictions.csv (all with an animal_id column).
- """
- N = len(data["labels"])
- all_animal_ids = np.array(data.get("animal_ids", [""] * N), dtype=object)
- n_unique = len(set(a for a in all_animal_ids if a))
- if n_unique == 0:
- print(
- "ERROR: C4 H5 has no animal_id metadata.\n"
- "Rebuild the H5 with the C4-format preprocessing script after the ANIMAL_ID_FIELD update.",
- file=sys.stderr,
- )
- return 1
- print(f"Animal IDs: {n_unique} unique animals across {N} neurons")
- # ── Split animals into train / test ──────────────────────────────────────
- train_animals, test_animals = _build_animal_splits(
- all_animal_ids,
- holdout_fold=args.holdout_fold,
- n_holdout_folds=args.n_holdout_folds,
- min_cells=args.min_cells_per_animal,
- )
- print(f"Holdout fold {args.holdout_fold}: "
- f"train={len(train_animals)} animals test={len(test_animals)} animals")
- train_mask = np.array([a in train_animals for a in all_animal_ids])
- test_mask = np.array([a in test_animals for a in all_animal_ids])
- # ── Filter to labeled neurons within each split ──────────────────────────
- unlabeled = set(args.unlabeled_strings)
- def _labeled_in_mask(mask):
- return np.where(
- mask & np.array([s not in unlabeled for s in data["str_labels"]])
- )[0]
- train_idx = _labeled_in_mask(train_mask)
- test_idx = _labeled_in_mask(test_mask)
- print(f" Labeled train : {len(train_idx)} | Labeled test : {len(test_idx)}")
- if len(train_idx) == 0 or len(test_idx) == 0:
- print("ERROR: no labeled neurons in train or test split.", file=sys.stderr)
- return 1
- # ── Label encoding (fit on union so vocabulary is shared) ────────────────
- all_str = ([data["str_labels"][i] for i in train_idx] +
- [data["str_labels"][i] for i in test_idx])
- le = LabelEncoder()
- le.fit(all_str)
- train_labels_int = le.transform([data["str_labels"][i] for i in train_idx])
- test_labels_int = le.transform([data["str_labels"][i] for i in test_idx])
- print(f" Classes: {list(le.classes_)}")
- # ── Waveform arrays ──────────────────────────────────────────────────────
- waveforms_all = data["waveforms"]
- if waveforms_all.dtype == object:
- lengths = [len(waveforms_all[i]) for i in np.concatenate([train_idx, test_idx])]
- T = int(np.median(lengths))
- def _to_fixed(idx_arr):
- out = np.zeros((len(idx_arr), T), dtype=np.float32)
- for j, i in enumerate(idx_arr):
- w = waveforms_all[i]
- l = min(len(w), T)
- out[j, :l] = w[:l]
- return out
- train_wvf = _to_fixed(train_idx)
- test_wvf = _to_fixed(test_idx)
- else:
- train_wvf = waveforms_all[train_idx].astype(np.float32)
- test_wvf = waveforms_all[test_idx].astype(np.float32)
- print(f"Waveform shape: train={train_wvf.shape} test={test_wvf.shape}")
- # ── Feature extraction (fit PCA on training neurons only) ────────────────
- print("\nExtracting waveform features ...")
- t_feat = time.time()
- train_feats, pca_obj = build_features(
- train_wvf, pca_variance=args.pca_variance, fit_pca=True)
- test_feats, _ = build_features(
- test_wvf, pca_variance=args.pca_variance, pca=pca_obj, fit_pca=False)
- feat_time = time.time() - t_feat
- print(f"Feature shape: train={train_feats.shape} test={test_feats.shape} "
- f"({feat_time:.2f}s)")
- # ── Probes ───────────────────────────────────────────────────────────────
- print("\nRunning probes ...")
- t_probe = time.time()
- rf_preds, mlp_preds, lr_preds = _run_probes(
- train_feats, train_labels_int,
- test_feats, test_labels_int,
- n_trees=args.n_trees,
- )
- probe_time = time.time() - t_probe
- print(f"Probes done in {probe_time:.1f}s")
- # ── Write prediction CSVs ────────────────────────────────────────────────
- test_true_str = le.inverse_transform(test_labels_int)
- test_animal_ids_ = all_animal_ids[test_idx]
- def _write(preds, fname):
- df = pd.DataFrame({
- "pred": le.inverse_transform(preds.astype(int)),
- "true": test_true_str,
- "animal_id": test_animal_ids_,
- })
- df.to_csv(os.path.join(args.output_dir, fname), index=False)
- print(f" → {os.path.join(args.output_dir, fname)}")
- _write(rf_preds, "holdout_predictions.csv")
- _write(mlp_preds, "mlp_predictions.csv")
- _write(lr_preds, "linear_probe_predictions.csv")
- with open(os.path.join(args.output_dir, "animal_split.json"), "w") as f:
- json.dump({
- "holdout_fold": args.holdout_fold,
- "n_holdout_folds": args.n_holdout_folds,
- "train_animals": sorted(train_animals),
- "test_animals": sorted(test_animals),
- "min_cells_per_animal": args.min_cells_per_animal,
- "pca_variance": args.pca_variance,
- "n_pca_components": int(
- np.searchsorted(
- np.cumsum(pca_obj.explained_variance_ratio_),
- args.pca_variance,
- ) + 1
- ),
- "n_trees": args.n_trees,
- "n_features": int(train_feats.shape[1]),
- }, f, indent=2)
- total = feat_time + probe_time
- print(f"\nDone in {total:.1f}s (feat={feat_time:.1f}s probe={probe_time:.1f}s)")
- print(f"Results written to {args.output_dir}")
- return 0
- # ---------------------------------------------------------------------------
- # Main
- # ---------------------------------------------------------------------------
- def main() -> int:
- parser = argparse.ArgumentParser(
- description="Buccino 2018 single-channel waveform feature baseline.",
- formatter_class=argparse.ArgumentDefaultsHelpFormatter,
- )
- parser.add_argument("--dataset", required=True,
- help="Dataset name, e.g. hull_cell_type")
- parser.add_argument("--cv-fold", type=int, default=0,
- help="CV fold index (0 to n-cv-folds-1)")
- parser.add_argument("--n-cv-folds", type=int, default=5)
- parser.add_argument("--data-root", default="./datasets",
- help="Root directory containing downloaded H5 files")
- parser.add_argument("--output-dir", required=True,
- help="Write prediction CSVs and cv_split.json here")
- parser.add_argument("--pca-variance", type=float, default=0.90,
- help="Fraction of variance retained by PCA (Buccino 2018: 0.90)")
- parser.add_argument("--n-trees", type=int, default=200,
- help="Number of trees in the Random Forest")
- parser.add_argument("--unlabeled-strings", nargs="+",
- default=["", "unlabeled"],
- help="str_label values treated as unlabeled/excluded from probes")
- # ── Holdout (inductive) mode ──────────────────────────────────────────────
- parser.add_argument("--holdout-fold", type=int, default=None,
- help="Enable holdout mode: index of animal fold to hold out "
- "(0 to n-holdout-folds-1). When set, --cv-fold is ignored.")
- parser.add_argument("--n-holdout-folds", type=int, default=5,
- help="Number of animal holdout folds (default: 5)")
- parser.add_argument("--min-cells-per-animal", type=int, default=50,
- help="Minimum cells per animal to include in holdout splits "
- "(default: 50, matching HIPPIE)")
- args = parser.parse_args()
- os.makedirs(args.output_dir, exist_ok=True)
- mode = "holdout" if args.holdout_fold is not None else "transductive"
- print("=" * 60)
- print("Buccino 2018 Waveform Feature Baseline")
- print(f" dataset : {args.dataset}")
- print(f" mode : {mode}")
- if mode == "holdout":
- print(f" holdout_fold : {args.holdout_fold}/{args.n_holdout_folds}")
- print(f" min_cells : {args.min_cells_per_animal}")
- else:
- print(f" cv_fold : {args.cv_fold}/{args.n_cv_folds}")
- print(f" pca_variance : {args.pca_variance}")
- print(f" n_trees : {args.n_trees}")
- print("=" * 60)
- # ── 1. Load data from C4 H5 format ─────────────────────────────────
- t_load = time.time()
- h5_path = _find_h5(args.data_root, args.dataset)
- print(f"\nLoading {h5_path} ...")
- data = load_c4_dataset(h5_path, args.dataset, source_id=0)
- N = len(data["labels"])
- load_time = time.time() - t_load
- print(f"Loaded {N} neurons in {load_time:.1f}s")
- label_counts: dict[str, int] = {}
- for s in data["str_labels"]:
- label_counts[s] = label_counts.get(s, 0) + 1
- print(f"Label distribution: {label_counts}")
- # ── 2. Dispatch to holdout mode if requested ─────────────────────────────
- if args.holdout_fold is not None:
- return _run_holdout(data, args)
- # ── 3. (Transductive) Identify labeled neurons ───────────────────────────
- unlabeled = set(args.unlabeled_strings)
- is_labeled = np.array([s not in unlabeled for s in data["str_labels"]])
- labeled_idx = np.where(is_labeled)[0]
- print(f"\nLabeled neurons: {len(labeled_idx)} / {N}")
- if len(labeled_idx) < 2 * args.n_cv_folds:
- print(
- f"ERROR: only {len(labeled_idx)} labeled neurons — too few for "
- f"{args.n_cv_folds}-fold CV",
- file=sys.stderr,
- )
- return 1
- # ── 3. Label encoder + CV split (same seed as NEMO) ──────────
- labeled_str = [data["str_labels"][i] for i in labeled_idx]
- le = LabelEncoder()
- le.fit(labeled_str)
- labeled_int = le.transform(labeled_str)
- skf = StratifiedKFold(n_splits=args.n_cv_folds, shuffle=True, random_state=42)
- splits = list(skf.split(labeled_idx, labeled_int))
- train_pos, test_pos = splits[args.cv_fold]
- print(f"CV fold {args.cv_fold}: train={len(train_pos)} test={len(test_pos)}")
- print(f"Classes: {list(le.classes_)}")
- # ── 4. Extract waveforms for labeled neurons ──────────────────────────────
- waveforms_all = data["waveforms"]
- # Ensure uniform shape — load_c4_dataset may return object array if ragged
- waveforms_labeled = waveforms_all[labeled_idx]
- if waveforms_labeled.dtype == object:
- # Pad / truncate to the most common length
- lengths = [len(w) for w in waveforms_labeled]
- T = int(np.median(lengths))
- wvf_fixed = np.zeros((len(labeled_idx), T), dtype=np.float32)
- for j, w in enumerate(waveforms_labeled):
- l = min(len(w), T)
- wvf_fixed[j, :l] = w[:l]
- waveforms_labeled = wvf_fixed
- else:
- waveforms_labeled = waveforms_labeled.astype(np.float32)
- train_wvf = waveforms_labeled[train_pos]
- test_wvf = waveforms_labeled[test_pos]
- train_labels_int = labeled_int[train_pos]
- test_labels_int = labeled_int[test_pos]
- print(f"\nWaveform shape: {waveforms_labeled.shape} (dtype {waveforms_labeled.dtype})")
- # ── 5. Feature extraction + PCA ──────────────────────────────────────────
- print("\nExtracting waveform features ...")
- t_feat = time.time()
- train_feats, pca_obj = build_features(
- train_wvf,
- pca_variance=args.pca_variance,
- fit_pca=True,
- )
- test_feats, _ = build_features(
- test_wvf,
- pca_variance=args.pca_variance,
- pca=pca_obj,
- fit_pca=False,
- )
- feat_time = time.time() - t_feat
- print(f"Feature shape: train={train_feats.shape} test={test_feats.shape} "
- f"({feat_time:.2f}s)")
- # ── 6. Probes ─────────────────────────────────────────────────────────────
- print("\nRunning probes ...")
- t_probe = time.time()
- rf_preds, mlp_preds, lr_preds = _run_probes(
- train_feats, train_labels_int,
- test_feats, test_labels_int,
- n_trees=args.n_trees,
- )
- probe_time = time.time() - t_probe
- print(f"Probes done in {probe_time:.1f}s")
- # ── 7. Write prediction CSVs ─────────────────────────────────────────────
- test_true_str = le.inverse_transform(test_labels_int)
- def _write_preds(preds: np.ndarray, fname: str):
- df = pd.DataFrame({
- "pred": le.inverse_transform(preds.astype(int)),
- "true": test_true_str,
- })
- df.to_csv(os.path.join(args.output_dir, fname), index=False)
- print(f" → {os.path.join(args.output_dir, fname)}")
- _write_preds(rf_preds, "transductive_predictions.csv")
- _write_preds(mlp_preds, "mlp_predictions.csv")
- _write_preds(lr_preds, "linear_probe_predictions.csv")
- # ── 8. CV split metadata ──────────────────────────────────────────────────
- with open(os.path.join(args.output_dir, "cv_split.json"), "w") as f:
- json.dump(
- {
- "train_global_indices": labeled_idx[train_pos].tolist(),
- "test_global_indices": labeled_idx[test_pos].tolist(),
- "fold": args.cv_fold,
- "n_folds": args.n_cv_folds,
- "pca_variance": args.pca_variance,
- "n_pca_components": int(
- np.searchsorted(
- np.cumsum(pca_obj.explained_variance_ratio_),
- args.pca_variance,
- ) + 1
- ),
- "n_trees": args.n_trees,
- "n_features": int(train_feats.shape[1]),
- },
- f,
- indent=2,
- )
- total_time = load_time + feat_time + probe_time
- print(f"\nDone in {total_time:.1f}s total")
- print(f"Results written to {args.output_dir}")
- return 0
- if __name__ == "__main__":
- sys.exit(main())
wf_rf_benchmark_evaluation.py at commit 6165dac, under BSD-3-Clause · at the source
Overview
- Genomics Institute, University of California Santa Cruz, Santa Cruz, CA USA
- Live Cell Biotechnology Discovery Lab, University of California Santa Cruz, Santa Cruz, CA USA
- Department of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA USA
- Department of Molecular, Cellular and Developmental Biology, University of California Santa Cruz, Santa Cruz, CA USA
- Department of Electrical and Computer Engineering, University of California Santa Cruz, Santa Cruz, CA USA
- Biotechnology Program, Berkeley City College, Berkeley, CA USA
- Department of Statistics & Data Science, Yale University, New Haven, CT USA
- Wu Tsai Institute, Yale University, New Haven, CT USA
Abstract
Neuronal classification from extracellular electrophysiological recordings is challenging due to intrinsic waveform variability, noise, and technical differences across experiments, technologies, and species. We introduce HIPPIE (High-dimensional Interpretation of Physiological Patterns In Intercellular Electrophysiology), a deep learning framework that combines self-supervised pretraining on unlabeled datasets with supervised fine-tuning to classify neurons from extracellular recordings. Using conditional convolutional joint autoencoders, HIPPIE learns technology-adjusted representations of waveforms and spiking dynamics. Here we show, across mouse, rat, and macaque recordings, that HIPPIE classifies cell types competitively with existing methods while additionally supporting generative analyses that discriminative models cannot perform, including counterfactual decoding of electrophysiological signals under changed experimental conditioning, cross-species latent interpolation, and a cross-modal analysis revealing that spike-timing modalities and waveform morphology encode largely independent dimensions of neuronal identity. HIPPIE is available as both a Python package and a coding-free web application, providing a unified framework for multimodal neuronal classification across technologies, experimental conditions, and species.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 40 matches between paragraphs and lines of code.
huggingface.co/jesusgf23/hippie
18ae97027016c2a401c386999ce213176d00b20e, 2 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
2 files
- extract_embeddings.py, Python, 200 lines, 1 match
- README.md, Text, 358 lines
braingeneers/HIPPIE
455598c1d893bf5c7f28c1adf214cd27c4065878, 25 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
34 files
- data_wrangling_scripts/
_internal/ , Python, 165 linesdownload_sessions_to_jso n.py - data_wrangling_scripts/
acqm_to_csv_converter.ip , Jupyter, 435 linesynb - data_wrangling_scripts/
allen_nwb_to_csv_convert , Jupyter, 579 lines, 2 matcheser.ipynb - data_wrangling_scripts/
ibl_one_to_csv_converter , Jupyter, 353 lines, 1 match.ipynb - data_wrangling_scripts/
neurocurator.py , Python, 294 lines - examples/
cross_dataset_tutorial.i , Jupyter, 268 lines, 2 matchespynb - examples/
extract_embeddings.py , Python, 222 lines, 1 match - examples/
train_on_your_own_data.i , Jupyter, 196 linespynb - hippie/
__init__.py , Python, 6 lines - hippie/
augmentations.py , Python, 214 lines, 1 match - hippie/
backbones.py , Python, 183 lines - hippie/
checkpoint.py , Python, 93 lines - hippie/
cli.py , Python, 336 lines - hippie/
dataloading.py , Python, 178 lines - hippie/
inference.py , Python, 297 lines, 1 match - hippie/
multimodal_model.py , Python, 1,061 lines, 2 matches - hippie/
optimizers.py , Python, 209 lines - hippie/
unimodal_model.py , Python, 385 lines - hippie/
utils.py , Python, 40 lines - hippie/
vae.py , Python, 281 lines - hippie_nwb_classify.py, Python, 774 lines, 1 match
- scripts/
generate_toy_dataset.py , Python, 108 lines - scripts/
train.py , Python, 244 lines, 1 match - tests/
__init__.py , Python, 1 line - tests/
conftest.py , Python, 35 lines - tests/
test_augmentations.py , Python, 211 lines - tests/
test_cli.py , Python, 81 lines - tests/
test_configs.py , Python, 132 lines - tests/
test_dataset_schema.py , Python, 93 lines - tests/
test_free_bits.py , Python, 72 lines - tests/
test_inference.py , Python, 75 lines - tests/
test_k_selection.py , Python, 205 lines - LICENSE, License, 28 lines
- README.md, Text, 543 lines
JesusGF1/hippie_benchmarking_release
6165daca6da283e5309f1c7840cf39bcb919dcf5, 31 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
138 files
- analysis/
CVAE_only_experiment_G0_ , Python, 198 linesclassification_accuracy. py - analysis/
CVAE_only_experiment_G1_ , Python, 110 linescross_modal_imputation.p y - analysis/
CVAE_only_experiment_G2_ , Python, 252 linescounterfactual_swap.py - analysis/
CVAE_only_experiment_G2_ , Python, 223 linescross_species.py - analysis/
CVAE_only_experiment_G3_ , Python, 152 lineslatent_interpolation.py - analysis/
CVAE_only_experiment_bui , Python, 105 linesld_datasets.py - analysis/
CVAE_only_experiment_lat , Python, 187 lines, 2 matchesent_flow_crosstech.py - analysis/
CVAE_only_experiment_plo , Python, 490 lines, 1 matcht.py - analysis/
CVAE_only_experiment_tra , Python, 258 lines, 1 matchin_hippie.py - analysis/
CVAE_only_experiment_tra , Python, 215 linesin_hippie_wf3dacg.py - analysis/
CVAE_only_experiment_uti , Python, 381 linesls.py - analysis/
CVAE_only_latent_flow.py , Python, 108 lines - analysis/
figure_2_panel_a_pipelin , Python, 118 linese.py - analysis/
figure_2_panel_b_hull_pr , Python, 177 linesofile.py - analysis/
figure_2_panel_c_hausser , Python, 190 lines_ablation.py - analysis/
figure_2_panel_d_hausser , Python, 226 lines_hyperparam.py - analysis/
figure_2_supp_nemo.py , Python, 243 lines, 2 matches - analysis/
figure_2_supp_physmap.py , Python, 238 lines - analysis/
figure_2_supp_pipelines. , Python, 100 linespy - analysis/
figure_3_panel_ab_profil , Python, 239 lines, 2 matchese.py - analysis/
figure_3_panel_bimodal_b , Python, 218 linesenchmark.py - analysis/
figure_3_panel_c_l4l5.py , Python, 226 lines - analysis/
figure_3_panel_confusion , Python, 150 lines_matrices.py - analysis/
figure_3_supp_isiacg_ben , Python, 201 lineschmark.py - analysis/
figure_3_supp_isiacg_pro , Python, 232 linesfile.py - analysis/
figure_4_panel_profiles. , Python, 268 linespy - analysis/
figure_4_supp_trimodal_b , Python, 288 linesenchmark.py - analysis/
figure_4_trimodal_benchm , Python, 490 linesark.py - analysis/
figure_5_panel_a_method_ , Python, 117 linesbars.py - analysis/
figure_5_panel_c_confusi , Python, 119 lineson.py - analysis/
figure_5_panel_d_umap.py , Python, 182 lines - analysis/
figure_5_supp_analysis.p , Python, 335 linesy - analysis/
figure_5_supp_nopretrain , Python, 129 lines.py - analysis/
figure_7_brain_region_be , Python, 343 linesnchmark.py - analysis/
figure_polarity_suppleme , Python, 454 lines, 2 matchesnt.py - analysis/
figure_supp_umap_cluster , Python, 608 lines, 2 matchesing.py - analysis/
make_source_data.py , Python, 225 lines, 2 matches - comparison_methods/
nemo/ , Python, 136 linesiclr/ eval/ c4_nested_cross_validati on.py - comparison_methods/
nemo/ , Python, 34 linesiclr/ eval/ c4_supervise_evaluation. py - comparison_methods/
nemo/ , Python, 118 linesiclr/ eval/ c4_vae_evaluation.py - comparison_methods/
nemo/ , Python, 426 linesiclr/ eval/ ibl_label_ratio_MLP.py - comparison_methods/
nemo/ , Python, 163 linesiclr/ eval/ ibl_label_ratio_linear.p y - comparison_methods/
nemo/ , Python, 162 linesiclr/ eval/ label_ratio_sweep_per_cl ass.py - comparison_methods/
nemo/ , Python, 516 linesiclr/ eval/ nearest_neighbor_region_ decoder.py - comparison_methods/
nemo/ , Jupyter, 105 linesiclr/ predict_region.ipynb - comparison_methods/
nemo/ , Python, 116 linesiclr/ predict_region.py - comparison_methods/
nemo/ , Python, 267 linesiclr/ train/ acg_VAE_training_seed_sw eep.py - comparison_methods/
nemo/ , Python, 286 linesiclr/ train/ wvf_VAE_training_seed_sw eep.py - comparison_methods/
nemo/ , Jupyter, 215 linesreplicate_c4_results.ipy nb - comparison_methods/
nemo/ , Python, 497 linesscripts/ allen_nwb_loader.py - comparison_methods/
nemo/ , Python, 170 linesscripts/ compute_3dACG_IBL.py - comparison_methods/
nemo/ , Python, 444 linesscripts/ compute_3dacg_from_spike s.py - comparison_methods/
nemo/ , Python, 149 linesscripts/ compute_NP_Ultra_dataset .py - comparison_methods/
nemo/ , Python, 260 linesscripts/ curate_and_process_ibl_d ataset.py - comparison_methods/
nemo/ , Python, 257 linesscripts/ gather_ibl_wvf_acg_pair. py - comparison_methods/
nemo/ , Python, 91 linesscripts/ get_c4_wvf_acg3d_pairs.p y - comparison_methods/
nemo/ , Shell, 1 linescripts/ ibl_label_sweep.sh - comparison_methods/
nemo/ , Python, 1,240 lines, 2 matchesscripts/ nemo_benchmark_evaluatio n.py - comparison_methods/
nemo/ , Python, 540 linesscripts/ nemo_transductive_test.p y - comparison_methods/
nemo/ , Shell, 1 linescripts/ run_c4_acg_vaes.sh - comparison_methods/
nemo/ , Shell, 13 linesscripts/ run_c4_wvf_vaes.sh - comparison_methods/
nemo/ , Python, 73 linesscripts/ run_mlp_probe_ultra.py - comparison_methods/
nemo/ , Shell, 186 linesscripts/ run_nemo_benchmark.sh - comparison_methods/
nemo/ , Shell, 74 linesscripts/ run_transductive_nemo.sh - comparison_methods/
nemo/ , Shell, 19 linesscripts/ train_c4_nemo.sh - comparison_methods/
nemo/ , Shell, 6 linesscripts/ train_ibl_nemo.sh - comparison_methods/
nemo/ , Shell, 11 linesscripts/ train_ultra_nemo.sh - comparison_methods/
nemo/ , Python, 6 linessetup.py - comparison_methods/
nemo/ , Python, 144 linessrc/ celltype_ibl/ models/ ACG_augmentation_dataloa der.py - comparison_methods/
nemo/ , Python, 173 linessrc/ celltype_ibl/ models/ AcgAug.py - comparison_methods/
nemo/ , Python, 364 linessrc/ celltype_ibl/ models/ BiModalEmbedding.py - comparison_methods/
nemo/ , Python, 88 linessrc/ celltype_ibl/ models/ WvfAug.py - comparison_methods/
nemo/ , Python, 1 linesrc/ celltype_ibl/ models/ __init__.py - comparison_methods/
nemo/ , Python, 1,052 linessrc/ celltype_ibl/ models/ bimodal_embedding_main_k nn.py - comparison_methods/
nemo/ , Python, 255 linessrc/ celltype_ibl/ models/ encoders.py - comparison_methods/
nemo/ , Python, 682 linessrc/ celltype_ibl/ models/ linear_classifier.py - comparison_methods/
nemo/ , Python, 681 linessrc/ celltype_ibl/ models/ linear_classifier_multi_ folds.py - comparison_methods/
nemo/ , Python, 140 linessrc/ celltype_ibl/ models/ linear_mlp_test.py - comparison_methods/
nemo/ , Python, 121 linessrc/ celltype_ibl/ models/ linear_probe.py - comparison_methods/
nemo/ , Python, 40 linessrc/ celltype_ibl/ models/ metrics.py - comparison_methods/
nemo/ , Python, 971 linessrc/ celltype_ibl/ models/ mlp_classifier.py - comparison_methods/
nemo/ , Python, 1,019 linessrc/ celltype_ibl/ models/ mlp_classifier_multi_fol ds.py - comparison_methods/
nemo/ , Python, 79 linessrc/ celltype_ibl/ params/ config.py - comparison_methods/
nemo/ , Python, 161 linessrc/ celltype_ibl/ params/ set_params.py - comparison_methods/
nemo/ , Python, 1,381 linessrc/ celltype_ibl/ utils/ c4_data_utils.py - comparison_methods/
nemo/ , Python, 259 linessrc/ celltype_ibl/ utils/ c4_vae_util.py - comparison_methods/
nemo/ , Python, 92 linessrc/ celltype_ibl/ utils/ cell_explorer_data_util. py - comparison_methods/
nemo/ , Python, 20 linessrc/ celltype_ibl/ utils/ cross_session_validation .py - comparison_methods/
nemo/ , Python, 274 linessrc/ celltype_ibl/ utils/ ibl_data_util.py - comparison_methods/
nemo/ , Python, 319 linessrc/ celltype_ibl/ utils/ ibl_data_visualization.p y - comparison_methods/
nemo/ , Python, 93 linessrc/ celltype_ibl/ utils/ ibl_eval_utils.py - comparison_methods/
nemo/ , Python, 73 linessrc/ celltype_ibl/ utils/ kenji_allen_data_utils.p y - comparison_methods/
nemo/ , Python, 58 linessrc/ celltype_ibl/ utils/ nested_cv_util.py - comparison_methods/
nemo/ , Python, 102 linessrc/ celltype_ibl/ utils/ preprocess.py - comparison_methods/
nemo/ , Python, 139 linessrc/ celltype_ibl/ utils/ ultra_data_util.py - comparison_methods/
nemo/ , Python, 522 linessrc/ celltype_ibl/ utils/ visualize_util.py - comparison_methods/
physmap/ , R, 460 linesphysmap_knn_crossdataset .r - comparison_methods/
physmap/ , R, 785 lines, 2 matchesphysmap_script.r - comparison_methods/
physmap/ , R, 585 linesphysmap_script_cross_dat aset.r - comparison_methods/
physmap/ , R, 657 linesphysmap_script_holdout.r - comparison_methods/
physmap/ , R, 586 linesphysmap_script_pca.r - comparison_methods/
physmap/ , Shell, 79 linesrun_physmap.sh - comparison_methods/
physmap/ , Shell, 86 linesrun_physmap_crossdataset .sh - comparison_methods/
physmap/ , Shell, 142 linesrun_physmap_holdout.sh - comparison_methods/
physmap/ , R, 51 linessetup_script.R - comparison_methods/
wf-rf/ , Python, 638 lines, 3 matcheswf_rf_benchmark_evaluati on.py - examples/
smoke_test_hf_checkpoint , Python, 160 lines.py - hippie/
__init__.py , Python, 1 line - hippie/
augmentations.py , Python, 214 lines - hippie/
backbones.py , Python, 183 lines - hippie/
cli.py , Python, 161 lines - hippie/
compute_parity.py , Python, 297 lines - hippie/
dataloading.py , Python, 162 lines - hippie/
multimodal_model.py , Python, 866 lines - hippie/
optimizers.py , Python, 209 lines - hippie/
utils.py , Python, 191 lines - hippie_wf3dacg/
__init__.py , Python, 56 lines - hippie_wf3dacg/
acg.py , Python, 186 lines, 1 match - hippie_wf3dacg/
dataloading.py , Python, 297 lines - hippie_wf3dacg/
load_c4.py , Python, 239 lines - hippie_wf3dacg/
model.py , Python, 695 lines, 2 matches - scripts/
check_dataset_cache_cons , Python, 140 linesistency.py - scripts/
cross_dataset_script.py , Python, 1,785 lines - scripts/
download_c4_h5.py , Python, 94 lines - scripts/
pick_locked_configs.py , Python, 262 lines, 1 match - scripts/
run_all_figures.sh , Shell, 99 lines, 2 matches - scripts/
train_hippie_wf3dacg.py , Python, 487 lines - scripts/
train_hippie_wf3dacg_cro , Python, 560 linesss_dataset.py - scripts/
train_multimodal_holdout , Python, 1,314 lines.py - scripts/
train_multimodal_holdout , Python, 1,113 lines_waveisi.py - scripts/
train_multimodal_transdu , Python, 2,182 linesctive.py - scripts/
train_multimodal_transdu , Python, 1,471 linesctive_isiacg.py - scripts/
train_multimodal_transdu , Python, 1,465 linesctive_waveisi.py - scripts/
train_multimodal_transdu , Python, 1,401 linesctive_waveonly.py - tests/
test_augmentations.py , Python, 264 lines - tests/
test_configs.py , Python, 165 lines - LICENSE, License, 28 lines
- README.md, Text, 189 lines
EricKenjiLee/PhysMAP_Manuscript
08819f45c7742001a760096c64bb00109bd8062e, 11 February 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
24 files
- CellExplorerv2/
AllenOnly.r , R, 245 lines - CellExplorerv2/
classifyCellExp.r , R, 118 lines - CellExplorerv2/
processCellExp.m , MATLAB, 110 lines - CellExplorerv2/
processCellExp.r , R, 248 lines - ConfMatPlotting.R, R, 39 lines
- InvivoA1/
processSantiago.r , R, 331 lines - PairOfVAEs_cuda.ipynb, Jupyter, 1,278 lines
- constants.R, R, 17 lines
- juxtacellular/
Generate_Traces.R , R, 47 lines - juxtacellular/
classifyData.r , R, 264 lines - juxtacellular/
classifyDataTwoModality. , R, 235 linesr - juxtacellular/
classifyRaw.R , R, 125 lines - juxtacellular/
generate_example_data.py , Python, 130 lines - juxtacellular/
helperFunctions.r , R, 156 lines - juxtacellular/
physmap_app.py , Python, 471 lines - juxtacellular/
physmap_utils.py , Python, 490 lines - juxtacellular/
processJianing.ipynb , Jupyter, 1,152 lines - juxtacellular/
processJianing.r , R, 199 lines - juxtacellular/
processJianingTwoModalit , R, 189 linesy.R - lookupTable/
lookupTable.r , R, 38 lines - lookupTable/
mapReference.r , R, 231 lines - lookupTable/
mapReferenceAllen.r , R, 229 lines - lookupTable/
mapReferenceAllen_Kenji. , R, 242 linesr - README.md, Text, 70 lines
Haansololfp/NEMO_ICLR
bc173508ad7cb8b664a8673412c6437ba4416ac9, 22 May 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
55 files
- iclr/
eval/ , Python, 136 linesc4_nested_cross_validati on.py - iclr/
eval/ , Python, 34 linesc4_supervise_evaluation. py - iclr/
eval/ , Python, 118 linesc4_vae_evaluation.py - iclr/
eval/ , Python, 426 linesibl_label_ratio_MLP.py - iclr/
eval/ , Python, 163 linesibl_label_ratio_linear.p y - iclr/
eval/ , Python, 162 lineslabel_ratio_sweep_per_cl ass.py - iclr/
eval/ , Python, 516 linesnearest_neighbor_region_ decoder.py - iclr/
predict_region.ipynb , Jupyter, 105 lines - iclr/
predict_region.py , Python, 116 lines - iclr/
train/ , Python, 267 linesacg_VAE_training_seed_sw eep.py - iclr/
train/ , Python, 286 lineswvf_VAE_training_seed_sw eep.py - scripts/
compute_3dACG_IBL.py , Python, 170 lines - scripts/
compute_NP_Ultra_dataset , Python, 149 lines.py - scripts/
curate_and_process_ibl_d , Python, 260 linesataset.py - scripts/
gather_ibl_wvf_acg_pair. , Python, 257 linespy - scripts/
get_c4_wvf_acg3d_pairs.p , Python, 91 linesy - scripts/
ibl_label_sweep.sh , Shell, 1 line - scripts/
run_c4_acg_vaes.sh , Shell, 1 line - scripts/
run_c4_wvf_vaes.sh , Shell, 13 lines - scripts/
run_mlp_probe_ultra.py , Python, 73 lines - scripts/
train_c4_nemo.sh , Shell, 9 lines - scripts/
train_ibl_nemo.sh , Shell, 6 lines - scripts/
train_ultra_nemo.sh , Shell, 11 lines - setup.py, Python, 6 lines
- src/
celltype_ibl/ , Python, 144 linesmodels/ ACG_augmentation_dataloa der.py - src/
celltype_ibl/ , Python, 173 linesmodels/ AcgAug.py - src/
celltype_ibl/ , Python, 364 linesmodels/ BiModalEmbedding.py - src/
celltype_ibl/ , Python, 88 linesmodels/ WvfAug.py - src/
celltype_ibl/ , Python, 1 linemodels/ __init__.py - src/
celltype_ibl/ , Python, 818 linesmodels/ bimodal_embedding_main.p y - src/
celltype_ibl/ , Python, 255 linesmodels/ encoders.py - src/
celltype_ibl/ , Python, 675 linesmodels/ linear_classifier.py - src/
celltype_ibl/ , Python, 681 linesmodels/ linear_classifier_multi_ folds.py - src/
celltype_ibl/ , Python, 140 linesmodels/ linear_mlp_test.py - src/
celltype_ibl/ , Python, 121 linesmodels/ linear_probe.py - src/
celltype_ibl/ , Python, 40 linesmodels/ metrics.py - src/
celltype_ibl/ , Python, 971 linesmodels/ mlp_classifier.py - src/
celltype_ibl/ , Python, 1,019 linesmodels/ mlp_classifier_multi_fol ds.py - src/
celltype_ibl/ , Python, 85 linesparams/ config.py - src/
celltype_ibl/ , Python, 151 linesparams/ set_params.py - src/
celltype_ibl/ , Python, 149 linesutils/ c4_data_utils.py - src/
celltype_ibl/ , Python, 259 linesutils/ c4_vae_util.py - src/
celltype_ibl/ , Python, 92 linesutils/ cell_explorer_data_util. py - src/
celltype_ibl/ , Python, 20 linesutils/ cross_session_validation .py - src/
celltype_ibl/ , Python, 274 linesutils/ ibl_data_util.py - src/
celltype_ibl/ , Python, 319 linesutils/ ibl_data_visualization.p y - src/
celltype_ibl/ , Python, 93 linesutils/ ibl_eval_utils.py - src/
celltype_ibl/ , Python, 73 linesutils/ kenji_allen_data_utils.p y - src/
celltype_ibl/ , Python, 58 linesutils/ nested_cv_util.py - src/
celltype_ibl/ , Python, 102 linesutils/ preprocess.py - src/
celltype_ibl/ , Python, 139 linesutils/ ultra_data_util.py - src/
celltype_ibl/ , Python, 511 linesutils/ visualize_util.py - src/
celltype_ibl/ , Python, 51 linesutils/ wandb.py - LICENSE, License, 21 lines
- README.md, Text, 68 lines
Zenodo 21707907
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
Code availability
The user facing code for HIPPIE is available on GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 6 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 245 scripts, each with its path and the digest of its content;
- 40 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- dandi:000041, at DANDI; found in “Data availability”
- dandi:000473, at DANDI; found in “Data availability”
- dandi:000955, at DANDI; found in “Data availability”
Data Availability Statement
All data analyzed in this study are previously published and publicly available; no new data or materials were generated.
The Hausser, Hull, and Lisberger cerebellar data used in this study are available in the C4 database (https://
The user facing code for HIPPIE is available on GitHub: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 15 authors, 2 keywords, 13 MeSH terms, 8 funders, 44 references.
Cite
This paper
Gonzalez-Ferrer, J., Lehrer, J., Alvarez-Esteban, B., Moreno-Ochando, A., Schweiger, H. E., Geng, J., Eugenio dos Santos, L. F. S., Hernandez, S., Reyes, F., Sevetson, J. L., Schneider, A., Salama, S. R., Teodorescu, M., Haussler, D., & Mostajo-Radji, M. A. (2026). HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities. Nature communications, 17(1), 10007. https://
BibTeX
@article{gonzalezferrer2
author = {Gonzalez-Ferrer, Jesus and Lehrer, Julian and Alvarez-Esteban, Bruno and Moreno-Ochando, Avelina and Schweiger, Hunter E and Geng, Jinghui and Eugenio dos Santos, Luiz F S and Hernandez, Sebastian and Reyes, Francisco and Sevetson, Jess L and Schneider, Aidan and Salama, Sofie R and Teodorescu, Mircea and Haussler, David and Mostajo-Radji, Mohammed A},
title = {{HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities}},
journal = {Nature communications},
year = {2026},
month = aug,
volume = {17},
number = {1},
pages = {10007},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42764354},
pmcid = {PMC13590638}
}
RIS
TY - JOUR
AU - Gonzalez-Ferrer, Jesus
AU - Lehrer, Julian
AU - Alvarez-Esteban, Bruno
AU - Moreno-Ochando, Avelina
AU - Schweiger, Hunter E
AU - Geng, Jinghui
AU - Eugenio dos Santos, Luiz F S
AU - Hernandez, Sebastian
AU - Reyes, Francisco
AU - Sevetson, Jess L
AU - Schneider, Aidan
AU - Salama, Sofie R
AU - Teodorescu, Mircea
AU - Haussler, David
AU - Mostajo-Radji, Mohammed A
TI - HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 10007
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "HIPPIE: a generative model for electrophysiological analysis across species, technologies, and modalities",
"container-title": "Nature communications",
"author": [
{
"family": "Gonzalez-Ferrer",
"given": "Jesus"
},
{
"family": "Lehrer",
"given": "Julian"
},
{
"family": "Alvarez-Esteban",
"given": "Bruno"
},
{
"family": "Moreno-Ochando",
"given": "Avelina"
},
{
"family": "Schweiger",
"given": "Hunter E"
},
{
"family": "Geng",
"given": "Jinghui"
},
{
"family": "Eugenio dos Santos",
"given": "Luiz F S"
},
{
"family": "Hernandez",
"given": "Sebastian"
},
{
"family": "Reyes",
"given": "Francisco"
},
{
"family": "Sevetson",
"given": "Jess L"
},
{
"family": "Schneider",
"given": "Aidan"
},
{
"family": "Salama",
"given": "Sofie R"
},
{
"family": "Teodorescu",
"given": "Mircea"
},
{
"family": "Haussler",
"given": "David"
},
{
"family": "Mostajo-Radji",
"given": "Mohammed A"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "10007",
"DOI": "10.1038/
"PMID": "42764354",
"PMCID": "PMC13590638",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
21
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41467-026-71331-0 [code]
- A multimodal approach for visualizing and identifying electrophysiological cell types in vivo.Journal: Nature communicationsIn common: caret, reticulate, UMAP, 15 other tools, mouse, 10 references
- [2] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: imbalanced-learn, PyTorch Lightning, caret, 17 other tools, mouse
- [3] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: AllenSDK, reticulate, UMAP, 14 other tools, non-human primate, mouse, 1 reference
- [4] doi:10.1016/j.xcrm.2026.102651 [code]
- Integrative CSF profiling identifies disease-specific immune responses in leptomeningeal disease.Journal: Cell reports. MedicineIn common: PyTorch Lightning, reticulate, UMAP, 15 other tools
- [5] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: UMAP, anndata, Scanpy, 15 other tools, 1 reference
- [6] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: reticulate, UMAP, anndata, 15 other tools, mouse
- [7] doi:10.7554/elife.108224 [code]
- Arrayed single-gene perturbations identify drivers of human anterior neural tube closure.Journal: eLifeIn common: reticulate, anndata, Scanpy, 15 other tools
- [8] doi:10.1038/s41586-026-10214-2 [code]
- Multidimensional profiling of heterogeneity in supratentorial ependymomas.Journal: NatureIn common: PyTorch Lightning, reticulate, anndata, 14 other tools, mouse
- [9] doi:10.1038/s44318-026-00818-9 [code]
- FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.Journal: The EMBO journalIn common: reticulate, UMAP, anndata, 14 other tools, mouse
- [10] doi:10.21203/rs.3.rs-9676637/v1 [code]
- A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic TechnologiesJournal: Research Square (preprint)In common: reticulate, UMAP, anndata, 14 other tools, methods / tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 6 repositories of the authors' code, each at its verified commit and with its license, 245 scripts, and 40 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:aa7358efacd92b70…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
