OSCR

RECOMBINE identifies recurrent composite markers of cell types and states.

Code ↔ Paper

17 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 17 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Results › RECOMBINE identifies composite marker sets for hierarchical cell identities ↔ R/recombine-package.R, the whole file · a weak match · score 0.81 · slab LASSO, sparse hierarchical clustering, recurrent composite, nearest neighbors, cell subpopulation, SHC SSL
  2. [2] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Identifying markers that characterize a rare subpopulation of mouse intestine ↔ R/recombine-package.R, the whole file · a weak match · score 0.81 · intestinal organoid, rare cell subpopulations, scRNA, single cell, transformation, mouse
  3. [3] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Selecting targeted panel from scRNA data for spatial molecular profiling of mouse visual cortex ↔ inst/scripts/recombinePipeline_MVCscRNA.R, lines 1–45 · score 0.80 · Allen Brain Atlas, mouse visual cortex, scRNA, downloaded, seq, genes
  4. [4] § Results › RECOMBINE identifies composite marker sets for hierarchical cell identities ↔ R/fl.R, lines 273–327 · score 0.75 · fused LASSO, L0 norm, sparse hierarchical clustering, SHC FL, feature selection, penalty
  5. [5] § Results › RECOBMINE optimizes marker panels for targeted spatial transcriptomics ↔ README.Rmd, lines 114–159 · score 0.74 · Rab3b, recurrent composite, Cck, Nrgn, Synpr, Gad1
  6. [6] § Results › RECOMBINE identifies composite marker sets for hierarchical cell identities ↔ README.Rmd, lines 17–30 · score 0.74 · unbiased selection, high dimensional, hierarchical cell subpopulations, slab LASSO, computational framework, sparse hierarchical clustering
  7. [7] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Selecting discriminant markers of transcriptional variation shaped by spatial gradients in the mouse cerebellum ↔ R/fl.R, lines 366–429 · score 0.73 · gap statistic profile, maximum gap statistic, standard error, nearest neighbors, dimensionality, optimal
  8. [8] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies ↔ R/ssl.R, lines 361–416 · score 0.72 · gap statistic profile, dissimilarity metric, squared distance, nearest neighbors, SHC SSL, matrix
  9. [9] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Selecting discriminant markers of transcriptional variation shaped by spatial gradients in the mouse cerebellum ↔ R/lasso.R, lines 286–334 · score 0.71 · gap statistic profile, maximum gap statistic, standard error, nearest neighbors, seq, genes
  10. [10] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies ↔ R/fl.R, lines 366–429 · score 0.71 · gap statistic profile, dissimilarity metric, squared distance, nearest neighbors, optimal, SHC
  11. [11] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Identifying markers that discriminate heterogeneous cells of human tissues in Tabula Sapiens ↔ R/pipeline.R, lines 190–226 · score 0.65 · log normalized, neighborhood recurrence, batch, Seurat, Harmony, resolution
  12. [12] § Results › RECOMBINE is robust to hyperparameter variation and data sparsity, and outperforms other feature selection methods ↔ R/pipeline.R, lines 190–226 · score 0.64 · fRECOMBINE, nonzero weights, fixed hyperparameter, graph, HVGs, match
  13. [13] § Methods › Benchmarking of RECOMBINE and other feature selection methods using biological data sets › Data preprocessing ↔ inst/scripts/recombinePipeline_MVCscRNA.R, lines 47–95 · score 0.60 · Low quality cells, log normalized, filtered, clustering
  14. [14] § Methods › Overview of RECOMBINE ↔ R/ssl.R, lines 361–416 · score 0.60 · slab LASSO penalty, gap statistics, SHC SSL, spike, metrics, hyperparameters
  15. [15] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Selecting targeted panel from scRNA data for spatial molecular profiling of mouse visual cortex ↔ inst/scripts/recombinePipeline_MVCscRNA.R, lines 1–45 · score 0.57 · mouse visual cortex, scRNA, downloaded, genes, clustering, cell
  16. [16] § Methods › Applying RECOMBINE to biological data sets for in-depth case studies › Selecting targeted panel from scRNA data for spatial molecular profiling of mouse visual cortex ↔ README.Rmd, lines 114–159 · score 0.54 · cell subpopulations, Gad1, Pcp4, Sst, Vip, filtered
  17. [17] § Methods › Benchmarking of RECOMBINE and other feature selection methods using biological data sets › Evaluation metrics ↔ R/lasso.R, lines 76–119 · score 0.54 · sparse hierarchical clustering, feature selection, objective, uniformly, norm, weights

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 95 lines · 2.5 KB · Apache-2.0 · 3 matches

  1. rm(list = ls())
  2. library(tidyverse)
  3. library(Seurat)
  4. library(recombine)
  5. # Here we use a scRNA-seq data of the mouse visual cortex from Allen Brain Atlas, which can be downloaded from:
  6. # http://celltypes.brain-map.org/api/v2/well_known_file_download/694413985
  7. # The following files are used after uncompressing the downloaded file:
  8. # mouse_VISp_2018-06-14_exon-matrix.csv
  9. # mouse_VISp_2018-06-14_genes-rows.csv
  10. # mouse_VISp_2018-06-14_samples-columns.csv
  11. # get expression matrix and cell annotation ------
  12. # expression
  13. tb <- read_csv("mouse_VISp_2018-06-14_exon-matrix.csv")
  14. colnames(tb)[1] <- "gene_entrez_id"
  15. tb_data <- tb
  16. # gene names
  17. tb <- read_csv("mouse_VISp_2018-06-14_genes-rows.csv")
  18. stopifnot(all.equal(tb_data$gene_entrez_id, tb$gene_entrez_id))
  19. tb_data$gene_symbol <- tb$gene_symbol
  20. tb_data <- tb_data %>%
  21. select(-gene_entrez_id) %>%
  22. select(gene_symbol, everything())
  23. # transpose
  24. mt <- as.matrix(tb_data[, -1])
  25. rownames(mt) <- tb_data$gene_symbol
  26. mt <- t(mt)
  27. tb_data <- tibble(cell = rownames(mt)) %>%
  28. bind_cols(as_tibble(mt))
  29. # anno
  30. tb <- read_csv("mouse_VISp_2018-06-14_samples-columns.csv")
  31. stopifnot(all.equal(tb_data$cell, tb$sample_name))
  32. tb <- tb %>%
  33. rename(cell = sample_name) %>%
  34. select(cell, class, subclass, cluster)
  35. tb_anno <- tb
  36. # filter low quality cells
  37. low_quality = c('No Class', 'Low Quality')
  38. tb_anno <- tb_anno %>%
  39. filter(!(class %in% low_quality) &
  40. !(subclass %in% low_quality) &
  41. !(cluster %in% low_quality))
  42. tb_data <- tb_data %>%
  43. filter(cell %in% tb_anno$cell)
  44. # generate log-normalized data --------
  45. # Initialize the Seurat object
  46. mt <- as.matrix(tb_data[, -1])
  47. rownames(mt) <- tb_data$cell
  48. mt <- t(mt)
  49. sobj <- CreateSeuratObject(counts = mt,
  50. project = "mvc",
  51. min.features = 0,
  52. min.cells = 0)
  53. # add cell annotation
  54. stopifnot(all.equal(rownames([email hidden]), tb_anno$cell))
  55. [email hidden] <- [email hidden] %>%
  56. cbind(tb_anno[, -1] %>% as.data.frame())
  57. # Normalizing the data
  58. sobj <- NormalizeData(sobj, normalization.method = "LogNormalize", scale.factor = 10000)
  59. # run RECOMBINE pipeline --------
  60. recombine.out <- recombine_pipeline(sobj, subpop_name = "subclass")
  61. # discriminant markers and weights
  62. recombine.out$df_w
  63. # recurrent composite markers of cell subpopulations
  64. recombine.out$df_rcm_subpop %>%
  65. filter(fract_signif_cells > 0.5 & avg_nhood_zscore > 2 & PR_AUC > 0.5) %>%
  66. group_by(subclass) %>%
  67. slice_max(avg_nhood_zscore, n = 3) %>%
  68. ungroup()
  69. recombine.out %>%
  70. saveRDS("recombine.out.rds")
  71. sessionInfo()

recombinePipeline_MVCscRNA.R at commit 8e88ef7, under Apache-2.0 · at the source

Overview

Authors: Xubin Li1, Justin Nguyen1, Anil Korkut1
  1. Department of Bioinformatics and Computational Biology, The University of Texas MD Anderson Cancer Center, Houston, Texas 77030, USA
Journal: Genome research, volume 36, issue 6, pages 1221-1237
Dates: received 21 April 2025; accepted 15 April 2026; published online June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1101/gr.280817.125 · PMID 42161585 · PMCID PMC13262949 · OpenAlex W7161790956
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), mouse (organism), zebrafish (organism)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing
MeSH: Algorithms*, Biomarkers*, Animals, Cerebellum, Humans, Mice, Visual Cortex, Zebrafish (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: the Bioinformatics Shared Resource (U01CA253472); Andrew Sabin Family Foundation; MD Anderson Cancer Center; Cancer Prevention and Research Institute of Texas (RP240293); NCI NIH HHS (U01 CA253472, P30 CA016672); National Cancer Institute (P30 CA016672)
Citations: not cited yet (Europe PMC); 61 references in the paper

Abstract

Biological function is mediated by the hierarchical organization of cell types and states within tissue ecosystems. Identifying interpretable composite marker sets that both define and distinguish hierarchical cell identities is essential for decoding biological complexity yet remains a major challenge. Here, we present RECOMBINE, an algorithm that identifies recurrent composite marker sets to define hierarchical cell identities. Validation using both simulated and biological data sets demonstrates that RECOMBINE is robust to hyperparameter variation and data sparsity, and achieves higher accuracy in identifying discriminant markers compared with existing approaches. As a partition-free framework, RECOMBINE is particularly powerful for data sets characterized by continuous cell-state transitions, in which defining discrete boundaries is inappropriate. This capability is demonstrated by its application to zebrafish development, revealing gradual transcriptional transitions across embryonic stages, and to the mouse cerebellum, in which it uncovers transcriptional variation shaped by spatial gradients. When applied to single-cell data and validated with spatial transcriptomic data from the mouse visual cortex, RECOMBINE identifies key cell-type markers and generates a robust gene panel for targeted spatial profiling. It also uncovers markers of CD8+ T cell states, including GZMK+HAVCR2− effector memory cells associated with anti-PD-1 therapy response. Finally, using data from the Tabula Sapiens project, RECOMBINE identifies composite marker sets across a broad range of human tissues. Together, these results highlight RECOMBINE as a robust, data-driven framework for optimized marker selection, enabling the discovery and validation of hierarchical cell identities across diverse tissue contexts.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 17 matches between paragraphs and lines of code.

korkutlab/recombine

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 8e88ef7252920745cdfd8f1fc42c3cc8d1ee18df, 20 April 2026
Languages: R (12), C++ (6), C/C++ (2)
Size: 82 files, 20 scripts
Software Heritage: archived
Found in: “Code availability”
Holds: README, license file, environment (DESCRIPTION), documentation, 1 notebook
Not found: CITATION.cff, tests, continuous integration
Tools: tidyverse (5 files), Seurat (3 files), Harmony (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
22 files

Code availability

The RECOMBINE R package and simulation data are available at GitHub (https://github.com/korkutlab/recombine) and as Supplemental Code.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 20 scripts, each with its path and the digest of its content;
  • 17 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Other data links

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 8 MeSH terms, 6 funders, 60 references.

Cite

This paper

Li, X., Nguyen, J., & Korkut, A. (2026). RECOMBINE identifies recurrent composite markers of cell types and states. Genome research, 36(6), 1221-1237. https://doi.org/10.1101/gr.280817.125

BibTeX

@article{li2026recombine,
author = {Li, Xubin and Nguyen, Justin and Korkut, Anil},
title = {{RECOMBINE identifies recurrent composite markers of cell types and states}},
journal = {Genome research},
year = {2026},
month = jun,
volume = {36},
number = {6},
pages = {1221--1237},
publisher = {Cold Spring Harbor Laboratory Press},
issn = {1088-9051},
doi = {10.1101/gr.280817.125},
url = {https://doi.org/10.1101/gr.280817.125},
pmid = {42161585},
pmcid = {PMC13262949}
}

RIS

TY - JOUR
AU - Li, Xubin
AU - Nguyen, Justin
AU - Korkut, Anil
TI - RECOMBINE identifies recurrent composite markers of cell types and states
T2 - Genome research
J2 - Genome Res
PY - 2026
DA - 2026/06/01
VL - 36
IS - 6
SP - 1221
EP - 1237
SN - 1088-9051
PB - Cold Spring Harbor Laboratory Press
DO - 10.1101/gr.280817.125
UR - https://doi.org/10.1101/gr.280817.125
LA - en
ER -

CSL-JSON

{
"id": "10.1101/gr.280817.125",
"type": "article-journal",
"title": "RECOMBINE identifies recurrent composite markers of cell types and states",
"container-title": "Genome research",
"author": [
{
"family": "Li",
"given": "Xubin"
},
{
"family": "Nguyen",
"given": "Justin"
},
{
"family": "Korkut",
"given": "Anil"
}
],
"container-title-short": "Genome Res",
"volume": "36",
"issue": "6",
"page": "1221-1237",
"DOI": "10.1101/gr.280817.125",
"PMID": "42161585",
"PMCID": "PMC13262949",
"ISSN": "1088-9051",
"publisher": "Cold Spring Harbor Laboratory Press",
"URL": "https://doi.org/10.1101/gr.280817.125",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.21203/rs.3.rs-9676637/v1 [code]
A Comprehensive Benchmarking of Spatial Deconvolution and Domain Detection Methods across Diverse Tissues and Spatial Transcriptomic Technologies
Journal: Research Square (preprint)
In common: Seurat, tidyverse, 5 references
[2] doi:10.7554/elife.107531 [code]
CellCover defines marker gene panels capturing developmental progression in neocortical neural stem cell identity.
Journal: eLife
In common: mouse, 6 references
[3] doi:10.1038/s41593-026-02293-1 [code]
Optics-free spatial genomics for mapping mammalian brain aging by IRISeq.
Journal: Nature neuroscience
In common: Seurat, tidyverse, mouse, 5 references
[4] doi:10.1002/imt2.70163 [code]
Spatial multi-omics unveils sphingolipid metabolic reprogramming within the retinal pathological niche.
Journal: iMeta
In common: Harmony, Seurat, tidyverse, mouse, 2 references
[5] doi:10.1038/s41467-026-76232-w [code]
Th17 effector cytokines induce shared and distinct microglial and endothelial cell responses in a mouse model for post-streptococcal encephalitis.
Journal: Nature communications
In common: Harmony, Seurat, tidyverse, mouse, 2 references
[6] doi:10.1093/nar/gkag621 [code]
Optimal gene panel selection for targeted spatial transcriptomics experiments.
Journal: Nucleic acids research
In common: mouse, 5 references
[7] doi:10.1038/s41467-026-71595-6 [code]
A single-cell and spatial atlas of early human olfactory development.
Journal: Nature communications
In common: Harmony, Seurat, tidyverse, 2 references
[8] doi:10.1038/s44320-026-00208-7 [code]
Interpretable deep generative ensemble learning for single-cell omics with Hydra.
Journal: Molecular systems biology
In common: Seurat, tidyverse, 3 references
[9] doi:10.1093/bib/bbag404 [code]
Navigating cell maps by deep learning integration of single-cell and spatially resolved transcriptomics.
Journal: Briefings in bioinformatics
In common: Seurat, tidyverse, mouse, 3 references
[10] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: Seurat, tidyverse, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.