OSCR

Tissueformer: extending single-cell foundation models to predict population-level phenotypes.

Code ↔ Paper

13 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 13 matches
  1. [1] § Methods › Pseudobulk and type composition benchmarks ↔ applications/brain_annotation/benchmarks.py, lines 1–52 · score 0.88 · random forest classifier, max depth, logistic regression, scikit-learn, lbfgs, leaf
  2. [2] § Methods › Deep learning benchmarks ↔ applications/brain_annotation/benchmarks.py, lines 740–824 · score 0.84 · weight decay, ScRAT, cosine, schedule, warmup, dropout
  3. [3] § Methods › Data and tokenization › Data description: mouse brains ↔ colormycells/colormap.py, lines 48–141 · score 0.80 · perceptual distance, perceptually uniform, pseudobulk expression, ColorMyCells, LUV, MDS
  4. [4] § Methods › Brain map visualization ↔ applications/brain_annotation/paper_figures/fig3/compute_visp_area.py, lines 443–531 · score 0.74 · flatmap pixel, cuML, smooth, GPU, kernel, SVC
  5. [5] § Methods › Data and tokenization › Data description: mouse brains ↔ applications/brain_annotation/figures/utils.py, lines 104–159 · score 0.69 · perceptual distance, perceptually uniform, LUV, MDS, matrix, colormaps
  6. [6] § Methods › TissueFormer architecture: hyperparameters of this study ↔ applications/brain_annotation/benchmarks.py, lines 740–824 · score 0.67 · learning rate warmup, PyTorch, decay, hidden, optimizer, Adam
  7. [7] § Results › Annotating cell groups with TissueFormer ↔ applications/brain_annotation/benchmarks.py, lines 827–889 · score 0.64 · Logistic regression, deep learning, random forest, bulk, CellCnn, ScRAT
  8. [8] § Methods › Deep learning benchmarks ↔ applications/brain_annotation/paper_figures/fig2/hyperparameters.ipynb, lines 139–170 · score 0.63 · CellCNN, ScAGG, ScRAT, Hyperparameters, benchmarked
  9. [9] § Methods › Data and tokenization › Data description: mouse brains ↔ applications/brain_annotation/paper_figures/fig3/cell_type_viz.ipynb, lines 518–657 · score 0.61 · brain slice, AnnData, ccf streamlines, axis
  10. [10] § Results › Single-cell profiles are not informative of cortical area ↔ applications/brain_annotation/paper_figures/fig2/single_cell_accuracy.py, lines 6–25 · score 0.61 · Murine Geneformer, logistic regression, random forest, neighbors, single cell, classifier
  11. [11] § Results › Predicting COVID-19 severity from blood transcriptomics ↔ applications/brain_annotation/paper_figures/fig2/hyperparameters.ipynb, lines 139–170 · score 0.56 · CellCNN, ScAGG, ScRAT, accuracy
  12. [12] § Results › Single-cell profiles are not informative of cortical area ↔ applications/brain_annotation/paper_figures/fig2/hyperparameters.ipynb, lines 172–212 · score 0.53 · logistic regression, brain accuracy, Validation accuracy, LR, bars, Figure 2
  13. [13] § Results › Predicted cortical maps ↔ applications/brain_annotation/paper_figures/fig3/gene_viz.ipynb, lines 489–498 · score 0.53 · Brinp3, Coro6, Nell1, Rcan2, Tshz2, Zfpm2

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 892 lines · 32 KB · MIT · 4 matches

  1. """
  2. Benchmarking Configuration
  3. =========================
  4. This configuration file controls the benchmarking pipeline for brain cell type classification models.
  5. It defines parameters for Random Forest and Logistic Regression classifiers, along with data handling
  6. and experiment tracking settings.
  7. Dependencies
  8. -----------
  9. - hydra-core
  10. - wandb (Weights & Biases)
  11. - scikit-learn
  12. - numpy
  13. - anndata (when using AnnData features)
  14. Configuration Sections
  15. --------------------
  16. random_forest:
  17. Parameters for sklearn RandomForestClassifier
  18. - n_estimators: Number of trees (default: 200)
  19. - max_depth: Maximum tree depth (default: 15)
  20. - min_samples_split: Minimum samples for split (default: 5)
  21. - min_samples_leaf: Minimum samples per leaf (default: 2)
  22. - max_features: Features per tree (default: 0.33)
  23. - bootstrap: Bootstrap samples (default: true)
  24. - class_weight: Class weighting strategy (default: null)
  25. logistic_regression:
  26. Parameters for sklearn LogisticRegression
  27. - max_iter: Maximum iterations (default: 1000)
  28. - multi_class: Classification strategy (default: 'multinomial')
  29. - solver: Optimization algorithm (default: 'lbfgs')
  30. Debug Options
  31. ------------
  32. debug: false # Enable debug mode with reduced dataset
  33. debug_args:
  34. on_adata: false # Use AnnData features
  35. resample_adata: true # Resample AnnData to match dataset size
  36. Experiment Types
  37. --------------
  38. run_bulk_expression_rf: false # Run Random Forest on bulk expression
  39. run_bulk_expression_lr: true # Run Logistic Regression on bulk expression
  40. run_h3type_rf: false # Run Random Forest on H3 types
  41. run_h3type_lr: false # Run Logistic Regression on H3 types
  42. Usage Examples
  43. ------------
  44. # Run default benchmarks
  45. python benchmarks.py
  46. # Enable debug mode
  47. python benchmarks.py debug=true
  48. # Run specific experiment
  49. python benchmarks.py run_bulk_expression_rf=true
  50. # Override Random Forest parameters
  51. python benchmarks.py random_forest.n_estimators=300 random_forest.max_depth=20
  52. Output
  53. ------
  54. Results are saved to ${output_dir} (default: 'benchmarks/') including:
  55. - Model performance metrics
  56. - Feature importance plots
  57. - Confusion matrices
  58. - Weights & Biases logs (if enabled)
  59. Notes
  60. -----
  61. - Set debug=true for quick iteration with reduced dataset
  62. - Enable wandb tracking by configuring wandb/default.yaml
  63. - Use debug_args.on_adata=true when working with AnnData features
  64. """
  65. import os
  66. import hydra
  67. import wandb
  68. import numpy as np
  69. import anndata as ad
  70. from omegaconf import DictConfig, OmegaConf
  71. from sklearn.ensemble import RandomForestClassifier
  72. from sklearn.linear_model import LogisticRegression, LogisticRegressionCV
  73. from sklearn.preprocessing import StandardScaler
  74. from collections import Counter
  75. from typing import Dict, List, Tuple, Optional
  76. from datasets import load_from_disk, DatasetDict, Dataset
  77. import torch
  78. from torch.utils.data import DataLoader
  79. from sklearn.model_selection import train_test_split
  80. from transformers import PreTrainedModel, TrainingArguments
  81. import json
  82. import scipy
  83. from sklearn.utils import check_random_state
  84. from sklearn.preprocessing import OneHotEncoder
  85. # Import necessary components from train.py
  86. from tissueformer.samplers import (
  87. GroupedSpatialTrainer
  88. )
  89. from transformers import PretrainedConfig
  90. class DummyConfig(PretrainedConfig):
  91. def __init__(self, **kwargs):
  92. super().__init__(**kwargs)
  93. class DummyModel(PreTrainedModel):
  94. config_class = DummyConfig
  95. def __init__(self, config):
  96. super().__init__(config)
  97. def forward(self, *args, **kwargs):
  98. return None
  99. def get_dataloaders(datasets: DatasetDict, cfg: DictConfig) -> Dict[str, DataLoader]:
  100. """
  101. Create dataloaders using GroupedSpatialTrainer infrastructure.
  102. """
  103. # Create minimal training arguments
  104. training_args = TrainingArguments(
  105. output_dir=cfg.output_dir,
  106. per_device_train_batch_size=cfg.data.group_size*32,
  107. per_device_eval_batch_size=cfg.data.group_size*32,
  108. remove_unused_columns=False, # Important for GroupedSpatialTrainer
  109. )
  110. dummy_config = DummyConfig()
  111. # Initialize trainer with dummy model
  112. trainer = GroupedSpatialTrainer(
  113. model=DummyModel(dummy_config),
  114. args=training_args,
  115. train_dataset=datasets["train"],
  116. eval_dataset=datasets["validation"],
  117. spatial_group_size=cfg.data.group_size,
  118. spatial_label_key="labels",
  119. coordinate_key='CCF_streamlines',
  120. additional_feature_keys=['raw_counts'],
  121. sampling_strategy=cfg.data.sampling.strategy,
  122. hex_scaling=cfg.data.sampling.hex_scaling,
  123. reflect_points=cfg.data.sampling.reflect_points,
  124. use_train_hex_grid_on_eval=cfg.data.sampling.use_train_hex_grid_on_eval,
  125. max_radius_expansions=cfg.data.sampling.max_radius_expansions,
  126. group_within_keys=cfg.data.sampling.group_within_keys
  127. )
  128. # Get dataloaders
  129. dataloaders = {
  130. "train": trainer.get_train_dataloader(),
  131. "validation": trainer.get_eval_dataloader(datasets["validation"]),
  132. "test": trainer.get_test_dataloader(datasets["test"])
  133. }
  134. return dataloaders
  135. def load_and_align_anndata(
  136. train_filenames: List[str],
  137. test_filenames: List[str],
  138. data_dir: str,
  139. dataset: DatasetDict,
  140. coordinate_key = "CCF_streamlines",
  141. ) -> Tuple[ad.AnnData, ad.AnnData]:
  142. """
  143. Load and concatenate AnnData files, ensuring alignment with dataset indices.
  144. Handles train and test files separately to match the original tokenization.
  145. Args:
  146. train_filenames: List of h5ad filenames for training data
  147. test_filenames: List of h5ad filenames for test data
  148. data_dir: Directory containing h5ad files
  149. dataset: HuggingFace dataset to align with
  150. Returns:
  151. Tuple of (train_adata, test_adata)
  152. """
  153. print("Loading and processing AnnData files...")
  154. def process_files(filenames: List[str]) -> ad.AnnData:
  155. """Helper function to process a list of files"""
  156. adatas = []
  157. for filename in filenames:
  158. filepath = os.path.join(data_dir, filename)
  159. print(f"Loading {filepath}")
  160. adata = ad.read_h5ad(filepath)
  161. # Filter cells with invalid CCF coordinates
  162. valid_mask = ~np.isnan(adata.obsm[coordinate_key]).any(axis=1)
  163. adata = adata[valid_mask]
  164. adatas.append(adata)
  165. return ad.concat(adatas, join='outer', fill_value=0)
  166. # Process train and test files separately
  167. train_adata = process_files(train_filenames)
  168. test_adata = process_files(test_filenames)
  169. # Verify alignment with dataset
  170. print("Verifying alignment with dataset...")
  171. # Check first 10000 cells in train dataset
  172. test_dataset = dataset['test']
  173. dataset_h3types = np.array(test_dataset[:10000]['H3_type'])
  174. adata_h3types = test_adata.obs['H3_type'].values[:10000]
  175. if not np.array_equal(dataset_h3types, adata_h3types):
  176. mismatches = np.where(dataset_h3types != adata_h3types)[0]
  177. mismatch_info = [
  178. f"Index {i}: Dataset H3_type: {dataset_h3types[i]}, AnnData H3_type: {adata_h3types[i]}"
  179. for i in mismatches
  180. ]
  181. raise ValueError(
  182. "Mismatch found in H3 types:\n" + "\n".join(mismatch_info)
  183. )
  184. print("Alignment verification passed!")
  185. return train_adata, test_adata
  186. def prepare_datasets(dataset_dict: DatasetDict, cfg: DictConfig) -> DatasetDict:
  187. """
  188. Prepare train/validation split from dataset.
  189. """
  190. train_dataset = dataset_dict["train"]
  191. # Create train/validation split
  192. train_idx, val_idx = train_test_split(
  193. np.arange(len(train_dataset)),
  194. test_size=cfg.data.validation_split,
  195. random_state=cfg.seed
  196. )
  197. # Add unique ids if not present
  198. if 'uuid' not in train_dataset.features:
  199. dataset_dict["test"] = dataset_dict["test"].add_column(
  200. "uuid",
  201. np.arange(len(dataset_dict["test"]))
  202. )
  203. train_dataset = train_dataset.add_column(
  204. "uuid",
  205. np.arange(len(dataset_dict["train"]))
  206. )
  207. # Select train and validation datasets
  208. val_dataset = train_dataset.select(val_idx)
  209. train_dataset = train_dataset.select(train_idx)
  210. # Limit dataset size if in debug mode
  211. if hasattr(cfg.data, 'max_train_samples') and cfg.data.max_train_samples is not None:
  212. train_dataset = train_dataset.select(range(min(len(train_dataset), cfg.data.max_train_samples)))
  213. # Limit validation/test size if specified
  214. if hasattr(cfg.data, 'max_eval_samples') and cfg.data.max_eval_samples is not None:
  215. val_dataset = val_dataset.select(range(min(len(val_dataset), cfg.data.max_eval_samples)))
  216. dataset_dict["test"] = dataset_dict["test"].select(range(min(len(dataset_dict["test"]), cfg.data.max_eval_samples)))
  217. # rename the labels to match the model's expected input
  218. if hasattr(cfg.data, 'label_key'):
  219. train_dataset = train_dataset.rename_column(cfg.data.label_key, "labels")
  220. val_dataset = val_dataset.rename_column(cfg.data.label_key, "labels")
  221. dataset_dict["test"] = dataset_dict["test"].rename_column(cfg.data.label_key, "labels")
  222. return DatasetDict({
  223. "train": train_dataset,
  224. "validation": val_dataset,
  225. "test": dataset_dict["test"]
  226. })
  227. def setup_wandb(cfg: DictConfig):
  228. """Initialize W&B logging"""
  229. wandb.init(
  230. project=cfg.wandb.project,
  231. entity=cfg.wandb.entity,
  232. name=cfg.wandb.name,
  233. group=cfg.wandb.group,
  234. tags=cfg.wandb.tags,
  235. notes=cfg.wandb.notes,
  236. config=OmegaConf.to_container(cfg, resolve=True),
  237. )
  238. def evaluate_method(
  239. predictions: np.ndarray,
  240. labels: np.ndarray,
  241. indices: np.ndarray,
  242. prefix: str,
  243. label_names: Dict,
  244. output_dir: str
  245. ) -> Dict:
  246. """Evaluate predictions and save results."""
  247. # Extract metrics
  248. from sklearn.metrics import balanced_accuracy_score
  249. metrics = {
  250. "accuracy": (predictions == labels).mean(),
  251. "balanced_accuracy": balanced_accuracy_score(labels, predictions),
  252. }
  253. # Save predictions
  254. output_dict = {
  255. 'predictions': predictions,
  256. 'labels': labels,
  257. 'indices': indices,
  258. 'label_names': label_names
  259. }
  260. np.save(os.path.join(output_dir, f"{prefix}_predictions.npy"), output_dict)
  261. # Log to wandb
  262. wandb.log({f"{prefix}_{k}": v for k, v in metrics.items()})
  263. return metrics
  264. def prepare_h3type_data(dataset: DatasetDict) -> Tuple[Dict[str, np.ndarray], Dict[str, int], Dict[str, Dict[int, int]]]:
  265. """
  266. Prepare H3 type data for fast access during training.
  267. Creates both type mapping and index mapping.
  268. """
  269. # Create type mapping
  270. all_h3_types = set()
  271. for split in dataset.values():
  272. all_h3_types.update(split['H3_type'])
  273. type_to_idx = {h3_type: idx for idx, h3_type in enumerate(sorted(list(all_h3_types)))}
  274. # Create numpy arrays and index mappings for fast access
  275. h3_arrays = {}
  276. index_maps = {}
  277. for split, dset in dataset.items():
  278. # Create mapping from dataset index to array index
  279. index_maps[split] = {
  280. idx: i for i, idx in enumerate(dset['uuid'])
  281. }
  282. # Create array with H3 type indices
  283. h3_arrays[split] = np.array([
  284. type_to_idx[t] for t in dset['H3_type']
  285. ])
  286. return h3_arrays, type_to_idx, index_maps
  287. def get_h3type_histogram(
  288. indices: np.ndarray,
  289. h3_array: np.ndarray,
  290. index_map: Dict[int, int],
  291. n_types: int
  292. ) -> np.ndarray:
  293. """
  294. Create histogram of H3 types for a group of cells using vectorized operations.
  295. Args:
  296. indices: Original dataset indices
  297. h3_array: Pre-computed array of H3 type indices
  298. index_map: Mapping from dataset indices to array indices
  299. n_types: Total number of H3 types
  300. """
  301. # Ensure indices is 2D: (n_groups, group_size)
  302. indices = np.asarray(indices)
  303. if indices.ndim == 1:
  304. indices = indices.reshape(1, -1) # One group with multiple cells
  305. batch_size = indices.shape[0]
  306. histogram = np.zeros((batch_size, n_types))
  307. for i, batch_indices in enumerate(indices):
  308. # Map dataset indices to array indices
  309. array_indices = [index_map[idx] for idx in batch_indices]
  310. # Get H3 types using mapped indices
  311. type_indices = h3_array[array_indices]
  312. histogram[i] = np.bincount(type_indices, minlength=n_types)
  313. total = histogram[i].sum()
  314. if total > 0:
  315. histogram[i] /= total
  316. return histogram
  317. def create_dataset_from_anndata(adata: ad.AnnData, cfg: DictConfig) -> Dataset:
  318. """
  319. Create a HuggingFace Dataset directly from AnnData object.
  320. """
  321. # Extract features (gene expression)
  322. features = np.array(adata.X.todense() if scipy.sparse.issparse(adata.X) else adata.X)
  323. print(f"Features shape: {features.shape}")
  324. # Extract coordinates
  325. coordinates = adata.obsm["CCF_streamlines"]
  326. # Extract area labels using the same logic as in tokenize_cells.py
  327. with open('data/files/area_ancestor_id_map.json', 'r') as f:
  328. area_ancestor_id_map = json.load(f)
  329. with open('data/files/area_name_map.json', 'r') as f:
  330. area_name_map = json.load(f)
  331. area_name_map['0'] = 'outside_brain'
  332. annotation2area_int = {0.0:0}
  333. for a in area_ancestor_id_map.keys():
  334. higher_area_id = area_ancestor_id_map[str(int(a))][1] if len(area_ancestor_id_map[str(int(a))])>1 else a
  335. annotation2area_int[float(a)] = int(higher_area_id)
  336. unique_areas = np.unique(list(annotation2area_int.values()))
  337. area_classes = np.arange(len(unique_areas))
  338. id2id = {float(k):v for (k,v) in zip(unique_areas, area_classes)}
  339. annotation2area_class = {k: id2id[int(v)] for k,v in annotation2area_int.items()}
  340. # Convert CCF annotations to area labels
  341. labels = np.array([annotation2area_class[x] for x in adata.obs['CCFano']])
  342. # Extract H3 types
  343. h3types = adata.obs['H3_type'].values
  344. # # Filter dataset to only include cells for which the CCF_streamlines is not nans
  345. # same as tokenized_dataset = tokenized_dataset.filter(lambda x: not np.isnan(np.sum(x['CCF_streamlines'])))
  346. # Filter out indices where CCF_streamlines contains NaN values
  347. valid_mask = ~np.isnan(coordinates).any(axis=1)
  348. features = features[valid_mask]
  349. coordinates = coordinates[valid_mask]
  350. labels = labels[valid_mask]
  351. h3types = h3types[valid_mask]
  352. indices = np.arange(len(adata))[valid_mask]
  353. # Create dataset
  354. return Dataset.from_dict({
  355. 'expression': features,
  356. 'CCF_streamlines': coordinates,
  357. 'labels': labels,
  358. 'H3_type': h3types,
  359. 'uuid': indices
  360. })
  361. def prepare_features_from_anndata(
  362. train_adata: ad.AnnData,
  363. test_adata: ad.AnnData,
  364. cfg: DictConfig,
  365. scaler: StandardScaler,
  366. feature_type: str
  367. ) -> Dict[str, Tuple[np.ndarray, np.ndarray, np.ndarray]]:
  368. """
  369. Prepare features directly from AnnData objects.
  370. Returns dict with keys 'train', 'validation', 'test', each containing (features, labels, indices)
  371. """
  372. # Create datasets from AnnData
  373. train_valid_dataset = create_dataset_from_anndata(train_adata, cfg)
  374. test_dataset = create_dataset_from_anndata(test_adata, cfg)
  375. rng = check_random_state(cfg.seed)
  376. # Split train/validation
  377. train_idx, val_idx = train_test_split(
  378. np.arange(len(train_valid_dataset)),
  379. test_size=cfg.data.validation_split,
  380. random_state=rng
  381. )
  382. # Select splits using Dataset.select()
  383. train_dataset = train_valid_dataset.select(train_idx)
  384. val_dataset = train_valid_dataset.select(val_idx)
  385. if feature_type == "h3type":
  386. # One-hot encode H3 types
  387. # Collect all unique H3 types from all splits
  388. all_h3_types = np.concatenate([
  389. train_dataset['H3_type'],
  390. val_dataset['H3_type'],
  391. test_dataset['H3_type']
  392. ])
  393. encoder = OneHotEncoder(sparse_output=False)
  394. # Fit on all types
  395. encoder.fit(all_h3_types.reshape(-1, 1))
  396. # Transform each split
  397. train_features = encoder.transform(np.array(train_dataset['H3_type']).reshape(-1, 1))
  398. val_features = encoder.transform(np.array(val_dataset['H3_type']).reshape(-1, 1))
  399. test_features = encoder.transform(np.array(test_dataset['H3_type']).reshape(-1, 1))
  400. else:
  401. # Extract and scale continuous features
  402. train_features = np.array(train_dataset['expression'])
  403. val_features = np.array(val_dataset['expression'])
  404. test_features = np.array(test_dataset['expression'])
  405. # Only apply scaling for non-categorical features
  406. train_features = scaler.fit_transform(train_features)
  407. val_features = scaler.transform(val_features)
  408. test_features = scaler.transform(test_features)
  409. # Extract features and labels as numpy arrays
  410. train_features = scaler.fit_transform(train_features)
  411. val_features = scaler.transform(val_features)
  412. test_features = scaler.transform(test_features)
  413. # Verify no NaN or infinite values
  414. if np.any(np.isnan(train_features)) or np.any(np.isinf(train_features)):
  415. raise ValueError("Training features contain NaN or infinite values")
  416. print(f"Train features final shape: {train_features.shape}")
  417. print(f"Test features final shape: {test_features.shape}")
  418. output = {
  419. 'train': (
  420. train_features,
  421. np.array(train_dataset['labels']),
  422. train_idx
  423. ),
  424. 'validation': (
  425. val_features,
  426. np.array(val_dataset['labels']),
  427. val_idx
  428. ),
  429. 'test': (
  430. test_features,
  431. np.array(test_dataset['labels']),
  432. np.arange(len(test_dataset))
  433. )
  434. }
  435. if cfg.debug_args.resample_adata:
  436. # Resample the data
  437. print("Resampling the adata")
  438. for name, (features, labels, indices) in output.items():
  439. resampled_indices = np.random.choice(range(len(indices)), len(indices), replace=True)
  440. output[name] = (features[resampled_indices], labels[resampled_indices], indices[resampled_indices])
  441. return output
  442. def run_classifier(
  443. datasets: DatasetDict,
  444. adata: Optional[Tuple[ad.AnnData, ad.AnnData]],
  445. cfg: DictConfig,
  446. classifier_type: str,
  447. feature_type: str
  448. ) -> None:
  449. """Generic function to run different classifiers on different feature types."""
  450. # Initialize classifier and scaler
  451. if classifier_type == "random_forest":
  452. clf = RandomForestClassifier(**cfg.random_forest)
  453. elif classifier_type == "logistic_regression":
  454. clf = LogisticRegression(**cfg.logistic_regression)
  455. else:
  456. raise ValueError(f"Unknown classifier type: {classifier_type}")
  457. scaler = StandardScaler()
  458. # If using direct AnnData path
  459. if cfg.data.group_size == 1 and cfg.debug_args.on_adata:
  460. train_adata, test_adata = adata
  461. splits = prepare_features_from_anndata(train_adata, test_adata, cfg, scaler, feature_type)
  462. # Train classifier
  463. train_features, train_labels, train_indices = splits['train']
  464. clf.fit(train_features, train_labels)
  465. # Evaluate
  466. for name, (features, labels, indices) in splits.items():
  467. predictions = clf.predict(features)
  468. evaluate_method(
  469. predictions,
  470. labels,
  471. indices,
  472. f"{classifier_type}_{feature_type}_{name}",
  473. cfg.data.label_names,
  474. cfg.output_dir
  475. )
  476. return
  477. # Original dataloader-based path
  478. dataloaders = get_dataloaders(datasets, cfg)
  479. # Prepare H3 type data if needed
  480. if feature_type == "h3type":
  481. h3_arrays, type_to_idx, index_maps = prepare_h3type_data(datasets)
  482. n_types = len(type_to_idx)
  483. # Training
  484. print(f"Collecting training features for {classifier_type} on {feature_type}...")
  485. train_features = []
  486. train_labels = []
  487. train_indices = []
  488. for batch in dataloaders["train"]:
  489. indices = batch['indices'].cpu().numpy() # these are uuid = indices in the original dataset['train'] before splitting
  490. train_indices.extend(indices)
  491. if feature_type == "bulk_expression":
  492. features = batch['raw_counts'].mean(1).cpu().numpy()
  493. else: # h3type
  494. features = get_h3type_histogram(indices, h3_arrays['train'], index_maps['train'], n_types)
  495. # get
  496. train_features.append(features)
  497. train_labels.append(batch['labels'].cpu().numpy())
  498. train_features = np.vstack(train_features)
  499. train_features = scaler.fit_transform(train_features)
  500. train_labels = np.concatenate(train_labels)
  501. print("Train features:", train_features.shape)
  502. print("Train labels:", train_labels.shape)
  503. print(f"Training {classifier_type}...")
  504. # Check for NaN/Inf values
  505. assert not np.any(np.isnan(train_features)), "Features contain NaN values"
  506. assert not np.any(np.isinf(train_features)), "Features contain infinite values"
  507. # Verify feature array is contiguous
  508. if not train_features.flags['C_CONTIGUOUS']:
  509. train_features = np.ascontiguousarray(train_features)
  510. # Fit with verbose logging
  511. clf.fit(train_features, train_labels)
  512. # Evaluate on all sets
  513. for name in ['train', 'validation', 'test']:
  514. loader = dataloaders[name]
  515. print(f"Evaluating on {name} set...")
  516. predictions = []
  517. labels = []
  518. indices = []
  519. for batch in loader:
  520. batch_indices = batch['indices'].cpu().numpy()
  521. indices.extend(batch_indices)
  522. if feature_type == "bulk_expression":
  523. features = batch['raw_counts'].mean(1).cpu().numpy()
  524. else: # h3type
  525. features = get_h3type_histogram(
  526. batch_indices,
  527. h3_arrays[name],
  528. index_maps[name],
  529. n_types
  530. )
  531. features = scaler.transform(features)
  532. pred = clf.predict(features)
  533. predictions.extend(pred)
  534. labels.extend(batch['labels'].cpu().numpy())
  535. evaluate_method(
  536. np.array(predictions),
  537. np.array(labels),
  538. np.array(indices),
  539. f"{classifier_type}_{feature_type}_{name}",
  540. cfg.data.label_names,
  541. cfg.output_dir
  542. )
  543. def run_dl_benchmark_brain(
  544. datasets: DatasetDict,
  545. cfg: DictConfig,
  546. model_name: str,
  547. ) -> None:
  548. """Run a DL benchmark on brain annotation data.
  549. Brain annotation differs from COVID: data is spatial (not donor-level),
  550. each cell has a label, and we group spatially adjacent cells into bags.
  551. """
  552. from torch.utils.data import DataLoader
  553. from tissueformer.benchmark_models.cellcnn import CellCnn
  554. from tissueformer.benchmark_models.scagg import ScAGG
  555. from tissueformer.benchmark_models.scrat import ScRAT
  556. from tissueformer.benchmark_models.data import MILDataset, CroppedMILDataset, mil_collate_fn
  557. from tissueformer.benchmark_models.trainer import BenchmarkTrainer
  558. import transformers
  559. print(f"\n--- DL benchmark: {model_name} ---")
  560. device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
  561. torch.manual_seed(cfg.seed)
  562. np.random.seed(cfg.seed)
  563. # Load model config
  564. model_cfg = OmegaConf.load(
  565. os.path.join(
  566. os.path.dirname(__file__),
  567. "config", "benchmark_models", f"{model_name}.yaml",
  568. )
  569. )
  570. if hasattr(cfg, "benchmark_models") and hasattr(cfg.benchmark_models, model_name):
  571. model_cfg = OmegaConf.merge(model_cfg, cfg.benchmark_models[model_name])
  572. # Use the spatial group dataloaders to get grouped bags
  573. dataloaders = get_dataloaders(datasets, cfg)
  574. # Collect all data through dataloaders into bags
  575. def collect_bags(loader, split_name):
  576. all_cells = []
  577. all_labels = []
  578. for batch in loader:
  579. raw = batch['raw_counts'].cpu().numpy() # (batch, group_size, n_genes)
  580. labs = batch['labels'].cpu().numpy() # (batch,)
  581. for i in range(raw.shape[0]):
  582. all_cells.append(raw[i])
  583. all_labels.append(labs[i])
  584. return all_cells, np.array(all_labels)
  585. train_cells, train_labels = collect_bags(dataloaders["train"], "train")
  586. val_cells, val_labels = collect_bags(dataloaders["validation"], "validation")
  587. test_cells, test_labels = collect_bags(dataloaders["test"], "test")
  588. # Each item in train_cells is already a (group_size, n_genes) array = one bag
  589. # Stack into flat expression matrix with sample indices
  590. def bags_to_mil_format(bags_list, labels_arr):
  591. expression_parts = []
  592. sample_indices = []
  593. cursor = 0
  594. for bag in bags_list:
  595. n = bag.shape[0]
  596. expression_parts.append(bag)
  597. sample_indices.append(np.arange(cursor, cursor + n))
  598. cursor += n
  599. expression = np.vstack(expression_parts).astype(np.float32)
  600. return expression, sample_indices, labels_arr
  601. train_X, train_si, train_y = bags_to_mil_format(train_cells, train_labels)
  602. val_X, val_si, val_y = bags_to_mil_format(val_cells, val_labels)
  603. test_X, test_si, test_y = bags_to_mil_format(test_cells, test_labels)
  604. # Apply preprocessing
  605. from tissueformer.benchmark_models.data import preprocess_zscore, preprocess_cp10k_log1p
  606. preprocess = model_cfg.get("preprocess", "raw")
  607. if preprocess == "zscore":
  608. scaler = StandardScaler()
  609. scaler.fit(train_X)
  610. train_X, _ = preprocess_zscore(train_X, scaler)
  611. val_X, _ = preprocess_zscore(val_X, scaler)
  612. test_X, _ = preprocess_zscore(test_X, scaler)
  613. elif preprocess == "cp10k_log1p":
  614. train_X = preprocess_cp10k_log1p(train_X)
  615. val_X = preprocess_cp10k_log1p(val_X)
  616. test_X = preprocess_cp10k_log1p(test_X)
  617. n_genes = train_X.shape[1]
  618. n_classes = int(max(train_y.max(), val_y.max(), test_y.max())) + 1
  619. # Compute class weights for imbalanced labels
  620. from tissueformer.class_weights import calculate_class_weights
  621. class_weights = calculate_class_weights(train_y, method="balanced", n_classes=n_classes)
  622. class_weights_tensor = torch.tensor(class_weights, dtype=torch.float32).to(device)
  623. # Build datasets (bags already formed, just wrap)
  624. train_ds = MILDataset(train_X, train_si, train_y)
  625. val_ds = MILDataset(val_X, val_si, val_y)
  626. test_ds = MILDataset(test_X, test_si, test_y)
  627. batch_size = model_cfg.batch_size
  628. train_loader = DataLoader(train_ds, batch_size=batch_size, shuffle=True, collate_fn=mil_collate_fn)
  629. val_loader = DataLoader(val_ds, batch_size=batch_size, shuffle=False, collate_fn=mil_collate_fn)
  630. test_loader = DataLoader(test_ds, batch_size=batch_size, shuffle=False, collate_fn=mil_collate_fn)
  631. # Build model
  632. if model_name == "cellcnn":
  633. cells_per_input = model_cfg.get("cells_per_input", 200)
  634. model = CellCnn(
  635. n_genes=n_genes, n_classes=n_classes,
  636. n_filters=model_cfg.n_filters,
  637. maxpool_percentage=model_cfg.maxpool_percentage,
  638. dropout=model_cfg.dropout,
  639. cells_per_input=cells_per_input,
  640. )
  641. optimizer = torch.optim.Adam(model.parameters(), lr=model_cfg.lr, weight_decay=model_cfg.l2_reg)
  642. l1_fn = lambda: model.l1_loss(model_cfg.l1_reg)
  643. loss_fn = torch.nn.CrossEntropyLoss(weight=class_weights_tensor)
  644. trainer = BenchmarkTrainer(
  645. model=model, optimizer=optimizer,
  646. train_loader=train_loader, val_loader=val_loader,
  647. device=device, n_epochs=model_cfg.n_epochs,
  648. early_stopping_patience=model_cfg.early_stopping_patience,
  649. l1_loss_fn=l1_fn, loss_fn=loss_fn,
  650. model_name=model_name, n_classes=n_classes,
  651. )
  652. elif model_name == "scagg":
  653. model = ScAGG(
  654. n_genes=n_genes, n_classes=n_classes,
  655. hidden_dim=model_cfg.hidden_dim,
  656. n_heads=model_cfg.n_heads, n_heads2=model_cfg.n_heads2,
  657. dropout=model_cfg.dropout,
  658. )
  659. optimizer = torch.optim.Adam(model.parameters(), lr=model_cfg.lr, weight_decay=model_cfg.weight_decay)
  660. loss_fn = torch.nn.CrossEntropyLoss(weight=class_weights_tensor, reduction=model_cfg.loss_reduction)
  661. trainer = BenchmarkTrainer(
  662. model=model, optimizer=optimizer,
  663. train_loader=train_loader, val_loader=val_loader,
  664. device=device, n_epochs=model_cfg.n_epochs,
  665. loss_fn=loss_fn, early_stopping_patience=None,
  666. model_name=model_name, n_classes=n_classes,
  667. )
  668. elif model_name == "scrat":
  669. model = ScRAT(
  670. n_genes=n_genes, n_classes=n_classes,
  671. hidden_dim=model_cfg.hidden_dim,
  672. n_heads=model_cfg.n_heads, n_layers=model_cfg.n_layers,
  673. dropout=model_cfg.dropout,
  674. )
  675. optimizer = torch.optim.Adam(model.parameters(), lr=model_cfg.lr, weight_decay=model_cfg.weight_decay)
  676. scheduler = transformers.get_cosine_schedule_with_warmup(
  677. optimizer, num_warmup_steps=model_cfg.lr_warmup_epochs,
  678. num_training_steps=model_cfg.n_epochs,
  679. )
  680. loss_fn = torch.nn.CrossEntropyLoss(weight=class_weights_tensor)
  681. trainer = BenchmarkTrainer(
  682. model=model, optimizer=optimizer,
  683. train_loader=train_loader, val_loader=val_loader,
  684. device=device, n_epochs=model_cfg.n_epochs,
  685. early_stopping_patience=model_cfg.early_stopping_patience,
  686. early_stopping_start_epoch=model_cfg.early_stopping_start_epoch,
  687. scheduler=scheduler, loss_fn=loss_fn,
  688. model_name=model_name, n_classes=n_classes,
  689. )
  690. else:
  691. raise ValueError(f"Unknown DL benchmark: {model_name}")
  692. # Train
  693. trainer.train()
  694. # Evaluate on test
  695. test_loss, test_preds, test_labels_out, test_probs = trainer.eval_epoch(test_loader)
  696. evaluate_method(
  697. test_preds, test_labels_out,
  698. np.arange(len(test_preds)),
  699. f"{model_name}_test",
  700. cfg.data.label_names if hasattr(cfg.data, 'label_names') else {},
  701. cfg.output_dir,
  702. )
  703. @hydra.main(version_base=None, config_path="config", config_name="benchmarks")
  704. def main(cfg: DictConfig) -> None:
  705. print("Running benchmarks...")
  706. # Print config
  707. print(OmegaConf.to_yaml(cfg))
  708. if cfg.debug:
  709. # Limit dataset size for faster iteration
  710. cfg.data.max_train_samples = 10000
  711. cfg.data.max_eval_samples = 10000
  712. # Setup wandb
  713. setup_wandb(cfg)
  714. # Load dataset
  715. dataset_dict = load_from_disk(cfg.data.dataset_path)
  716. datasets = prepare_datasets(dataset_dict, cfg)
  717. print(f"Loaded datasets: {datasets}")
  718. # Load AnnData if needed
  719. if cfg.debug_args.on_adata:
  720. adata = load_and_align_anndata(
  721. cfg.data.train_h5ad_files,
  722. cfg.data.test_h5ad_files,
  723. cfg.data.h5ad_directory,
  724. datasets
  725. )
  726. else:
  727. adata = None
  728. # ensure that the lengths are the same
  729. if cfg.debug_args.on_adata and cfg.debug:
  730. subsampled_idx_train = np.random.choice(len(adata[0]), len(datasets['train']), replace=False)
  731. subsampled_idx_test = np.random.choice(len(adata[1]), len(datasets['test']), replace=False)
  732. adata = (adata[0][subsampled_idx_train], adata[1][subsampled_idx_test])
  733. # Run benchmarks
  734. if cfg.run_bulk_expression_rf:
  735. run_classifier(datasets, adata, cfg, "random_forest", "bulk_expression")
  736. if cfg.run_bulk_expression_lr:
  737. run_classifier(datasets, adata, cfg, "logistic_regression", "bulk_expression")
  738. if cfg.run_h3type_rf:
  739. run_classifier(datasets, adata, cfg, "random_forest", "h3type")
  740. if cfg.run_h3type_lr:
  741. run_classifier(datasets, adata, cfg, "logistic_regression", "h3type")
  742. # Deep learning benchmarks
  743. dl_models = {
  744. "cellcnn": cfg.get("run_cellcnn", False),
  745. "scagg": cfg.get("run_scagg", False),
  746. "scrat": cfg.get("run_scrat", False),
  747. }
  748. for model_name, should_run in dl_models.items():
  749. if not should_run:
  750. continue
  751. run_dl_benchmark_brain(datasets, cfg, model_name)
  752. # Close wandb run
  753. wandb.finish()
  754. if __name__ == "__main__":
  755. main()

benchmarks.py at commit ff2b881, under MIT · at the source

Overview

Authors: Ari S. Benjamin1, Anthony Zador1
ORCID iDs: Anthony Zador
  1. Cold Spring Harbor Laboratory,Cold Spring Harbor, NY 11724 USA
Institutions: Cold Spring Harbor Laboratory (United States)
Journal: BMC bioinformatics, volume 27, issue 1, article 172
Dates: received 10 December 2025; accepted 13 May 2026; published online 4 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s12859-026-06490-4 · PMID 42243667 · PMCID PMC13466306 · OpenAlex W7163530187
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), mouse (organism), other condition (population), cellular / molecular (subfield)
Methods: Machine learning, Statistics, Preprocessing
Keywords: Spatial transcriptomics, Single-cell foundation models, Sample-level diagnostics, Brain annotation, Tissue labeling, AI for biology
MeSH: COVID-19*, Neural Networks, Computer*, Single-Cell Analysis*, Animals, Brain, Mice, Phenotype, Sequence Analysis, RNA, Single-Cell Gene Expression Analysis, Spatial Transcriptomics, Transcriptome (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NIH (RF1 MH126883-01A1)
Citations: not cited yet (Europe PMC); 69 references in the paper

Abstract

Background: Single-cell RNA sequencing technologies have enabled unprecedented insights into gene expression and opened new pathways for diagnostics and tissue annotation. At present, most computational approaches for interpreting single-cell data predict labels or properties based on isolated single-cell transcriptomic profiles. This approach overlooks the cellular composition within a sample, which is often critical for inferring tissue identity or other sample-level phenotypes.

Results: To address this limitation, we introduce TissueFormer, a Transformer-based neural network that infers population-level labels from groups of single-cell RNA profiles while retaining single-cell resolution. We applied TissueFormer to two tasks: predicting COVID-19 severity from single-cell RNA sequencing of blood samples, and predicting cortical area identity from spatial transcriptomic data in mouse brains. TissueFormer outperformed single-cell foundation models and machine learning methods applied to pseudobulk and cell type composition.

Conclusions: TissueFormer’s higher performance promises more accurate diagnostics and enables the automated construction of high-resolution brain region maps in individual mice directly from spatial transcriptomic data. Applied to mice with developmental perturbations to visual input, these maps revealed a significant reduction in predicted visual cortex area, illustrating how individual differences in neuroanatomy can be quantified. More broadly, TissueFormer provides a framework for predicting any population-level phenotypes which are influenced by cellular diversity and tissue-level organization.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.

ZadorLaboratory/TissueFormer

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: ff2b881ab40e09757facc4838eea7fce536327dd, 3 April 2026
Languages: Python (44), Jupyter (12), Shell (8)
Size: 257 files, 64 scripts
Software Heritage: archived
Found in: “Data availability”
Holds: README, environment (pyproject.toml), tests, 12 notebooks
Not found: license file, CITATION.cff, continuous integration, documentation
Tools: NumPy (19 files), Matplotlib (13 files), anndata (12 files), pandas (12 files), seaborn (12 files), SciPy (7 files), PyTorch (4 files), scikit-learn (3 files), Hugging Face Transformers (3 files), CuPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
21 files

Zenodo 15595324

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (2 files), NumPy (2 files), pandas (2 files), scikit-learn (2 files), SciPy (2 files), seaborn (2 files), anndata (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
5 files
At the source:

zadorlaboratory/colormycells

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 3edd4a073e9c5a052b536ea729cec5ee2cce5305, 17 July 2025
Languages: Python (2)
Size: 10 files, 2 scripts
Software Heritage: not archived
Found in: the Zenodo archive record
Holds: README, license file, CITATION.cff, environment (pyproject.toml)
Not found: tests, continuous integration, documentation
Tools: anndata (1 file), Matplotlib (1 file), NumPy (1 file), pandas (1 file), scikit-learn (1 file), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
4 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 25 scripts, each with its path and the digest of its content;
  • 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

All code is available at https://github.com/ZadorLaboratory/TissueFormer. Brain spatial transcriptomic data are hosted on Mendeley Data (https://doi.org/10.17632/5xfzcb4kn8.1 and https://doi.org/10.17632/8bhhk7c5n9.1). COVID-19 scRNA-seq datasets are publicly available from their original publications. All datasets used in this study are listed in Table 1.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 6 keywords, 11 MeSH terms, 1 funder, 58 references.

Cite

This paper

Benjamin, A. S., & Zador, A. (2026). Tissueformer: extending single-cell foundation models to predict population-level phenotypes. BMC bioinformatics, 27(1), 172. https://doi.org/10.1186/s12859-026-06490-4

BibTeX

@article{benjamin2026tissueformer,
author = {Benjamin, Ari S. and Zador, Anthony},
title = {{Tissueformer: extending single-cell foundation models to predict population-level phenotypes}},
journal = {BMC bioinformatics},
year = {2026},
month = jun,
volume = {27},
number = {1},
pages = {172},
publisher = {BMC},
issn = {1471-2105},
doi = {10.1186/s12859-026-06490-4},
url = {https://doi.org/10.1186/s12859-026-06490-4},
pmid = {42243667},
pmcid = {PMC13466306}
}

RIS

TY - JOUR
AU - Benjamin, Ari S.
AU - Zador, Anthony
TI - Tissueformer: extending single-cell foundation models to predict population-level phenotypes
T2 - BMC bioinformatics
J2 - BMC Bioinformatics
PY - 2026
DA - 2026/06/04
VL - 27
IS - 1
SP - 172
SN - 1471-2105
PB - BMC
DO - 10.1186/s12859-026-06490-4
UR - https://doi.org/10.1186/s12859-026-06490-4
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s12859-026-06490-4",
"type": "article-journal",
"title": "Tissueformer: extending single-cell foundation models to predict population-level phenotypes",
"container-title": "BMC bioinformatics",
"author": [
{
"family": "Benjamin",
"given": "Ari S."
},
{
"family": "Zador",
"given": "Anthony"
}
],
"container-title-short": "BMC Bioinformatics",
"volume": "27",
"issue": "1",
"page": "172",
"DOI": "10.1186/s12859-026-06490-4",
"PMID": "42243667",
"PMCID": "PMC13466306",
"ISSN": "1471-2105",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s12859-026-06490-4",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
4
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: anndata, PyTorch, seaborn, 5 other tools, 8 references
[2] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, 8 references
[3] doi:10.1038/s41467-026-71759-4 [code]
CellNiche represents cellular microenvironments in atlas-scale spatial omics data with contrastive learning.
Journal: Nature communications
In common: anndata, PyTorch, seaborn, 5 other tools, other condition, mouse, cellular / molecular, 7 references
[4] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, 5 references
[5] doi:10.1002/advs.77003 [code]
SemanticST: A Scalable Multi-Contextual Graph Learning Framework for Uncovering Spatial Niches and Robust Multi-Sample Integration in Spatial Transcriptomics.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, 5 references
[6] doi:10.1093/bib/bbag298 [code]
Empowering multifaceted analysis of spatial transcriptomics data with RGAST.
Journal: Briefings in bioinformatics
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, mouse, 4 references
[7] doi:10.1186/s13073-026-01704-z [code]
Gene expression profiling enables refined parcellation of cortical layers in the heterogeneous human cerebral cortex.
Journal: Genome medicine
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, mouse, cellular / molecular, 4 references
[8] doi:10.1016/j.isci.2026.117206 [code]
ReliST: A model-agnostic risk layer for spatial transcriptomics deconvolution.
Journal: iScience
In common: anndata, PyTorch, seaborn, 5 other tools, genetics / omics, mouse, 4 references
[9] doi:10.64898/2026.03.30.714220 [code]
An integrated single cell and spatial omics atlas of human prenatal development
Journal: bioRxiv (preprint)
In common: CuPy, Hugging Face Transformers, anndata, 7 other tools, 1 reference
[10] doi:10.1038/s41467-026-74171-0 [code]
Cluster replicability in single-cell and single-nucleus atlases of the mouse brain.
Journal: Nature communications
In common: anndata, pandas, SciPy, 1 other tool, genetics / omics, mouse, cellular / molecular, 6 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.