A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence.
The 16 matches
- [1] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kge_training/train_pykeen_model.py, lines 113–244 · score 0.96 · adversarial temperature, NSSALoss, negative sampler, random seed, Validation MRR, RotatE
- [2] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kge_training/evaluate_regulates_lp.py, lines 1–50 · score 0.80 · full entity vocabulary, filtered evaluation, tail prediction, head prediction, metrics, Hits
- [3] § 4. Materials and Methods › 4.6. In Silico Knockout Validation ↔ scripts/virtual_knockout/run_virtual_knockout.py, lines 1–21 · score 0.78 · virtual knockout, silico knockout, microglia regulator candidates, KGE, Validation
- [4] § 4. Materials and Methods › 4.6. In Silico Knockout Validation ↔ scripts/virtual_knockout/run_virtual_knockout.py, lines 88–175 · score 0.76 · virtual knockout, scTenifoldKnk, expression matrix, distance, networks, gene
- [5] § 4. Materials and Methods › 4.2. Single-Cell Transcriptomic Marker Gene Identification and Validation ↔ scripts/kg_construction/import_marker_genes.py, lines 37–111 · score 0.75 · gene biotype, gene symbols, gene records, Ensembl, pct, protein
- [6] § 2. Results › 2.1. A Hierarchical Neural Cell Type Taxonomy for Knowledge Graph Construction ↔ scripts/virtual_knockout/run_virtual_knockout.py, lines 23–40 · score 0.67 · excitatory neurons, inhibitory neurons, hippocampal CA1, Glial cell, astrocytes
- [7] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kg_construction/export_triples.py, lines 1–15 · score 0.65 · Neo4j, Paper nodes, KGE training, exporting, tail, head
- [8] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kge_training/evaluate_marker_retrieval.py, lines 1–69 · score 0.63 · reciprocal rank, filtered evaluation, metrics, vocabulary, candidate, scoring
- [9] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kge_training/train_pykeen_model.py, lines 1–48 · score 0.60 · link prediction evaluation, REGULATES triples, absent, splitting, validation, Embedding
- [10] § 4. Materials and Methods › 4.2. Single-Cell Transcriptomic Marker Gene Identification and Validation ↔ scripts/kg_construction/import_marker_genes.py, lines 30–34 · score 0.58 · CA3 pyramidal neurons, hippocampal CA1, node, Gene
- [11] § 4. Materials and Methods › 4.7. Statistical Analysis and Reproducibility ↔ scripts/kge_training/evaluate_marker_retrieval.py, lines 1–69 · score 0.57 · model training, KGE train, reproducing, held, metrics, splits
- [12] § 4. Materials and Methods › 4.6. In Silico Knockout Validation ↔ scripts/virtual_knockout/run_virtual_knockout.py, lines 1–21 · score 0.56 · virtual knockout, prioritization, Silico, validation, regulated
- [13] § 2. Results › 2.4. Mapping Molecular Fingerprints and Literature Evidence onto the Hierarchical Backbone ↔ scripts/kg_construction/export_triples.py, lines 1–15 · score 0.56 · Neo4j, Paper nodes, KGE training, exported, triple, entities
- [14] § 2. Results › 2.1. A Hierarchical Neural Cell Type Taxonomy for Knowledge Graph Construction ↔ scripts/kg_construction/import_marker_genes.py, lines 30–34 · score 0.55 · CA3 pyramidal neurons, hippocampal CA1, CA2, nodes
- [15] § 4. Materials and Methods › 4.4. Multi-Source Knowledge Graph Integration and Neo4j Implementation ↔ scripts/kg_construction/import_hierarchy.py, lines 1–16 · score 0.55 · Level3 nodes, Neo4j, Level1, Level2
- [16] § 4. Materials and Methods › 4.5. Knowledge Graph Embedding and Link Prediction ↔ scripts/kge_training/split_triples.py, lines 1–68 · score 0.52 · transductive safe, ratio, splitting, validation, interacts, training
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 179 lines · 6.7 KB · MIT · 4 matches
- #!/usr/bin/env python3
- """run_virtual_knockout.py
- In silico knockout validation for KGE-prioritized Microglia regulator candidates
- using the scTenifoldKnk framework (Cabezas-Bratesco et al., Patterns, 2022).
- Usage:
- python run_virtual_knockout.py
- """
- import io
- import time
- import warnings
- import numpy as np
- import pandas as pd
- from scipy.io import mmread
- from pathlib import Path
- from scTenifold import scTenifoldKnk
- warnings.filterwarnings("ignore")
- # ──────────────────────────────────────────────
- # Configuration
- # ──────────────────────────────────────────────
- BASE_DIR = Path(__file__).parent
- GENES_TO_KO = ["FMR1", "PTEN", "FKBP5"]
- N_SAMP_CELLS = 600 # cells sampled per network
- MAX_GENES = 3000 # gene cap after preprocessing
- KGE_INFO = {
- "FMR1": {"novel_rank": 6, "global_rank": 763, "score": -5.755,
- "cells": "Astrocyte; Bergmann glial cell; Excitatory neuron; "
- "Hippocampal CA1 PN; LAMP5+ interneuron; NSC; Radial glial cell"},
- "PTEN": {"novel_rank": 26, "global_rank": 787, "score": -6.113,
- "cells": "Cerebellar inhibitory neuron; Hippocampal DG granule cell; "
- "LAMP5+ interneuron; NSC; Pericyte"},
- "FKBP5": {"novel_rank": 38, "global_rank": 801, "score": -6.197,
- "cells": "Astrocyte; Fibroblast"},
- }
- # ──────────────────────────────────────────────
- # Data loading and preprocessing
- # ──────────────────────────────────────────────
- def load_expression_matrix():
- """Load Microglia expression matrix and downsample to 500 cells."""
- mtx_file = BASE_DIR / "data" / "microglia_matrix.mtx"
- gene_file = BASE_DIR / "data" / "microglia_gene_names.txt"
- print("\nLoading expression matrix...")
- with open(mtx_file, "rb") as f:
- mat = mmread(io.BytesIO(f.read())).tocsr()
- gene_names = open(gene_file, encoding="utf-8").read().strip().split("\n")
- if mat.shape[0] != len(gene_names):
- mat = mat.T.tocsr()
- # Downsample to 500 cells for input loading
- n_total = mat.shape[1]
- n_qc = min(500, n_total)
- np.random.seed(42)
- sel_cols = np.random.choice(n_total, n_qc, replace=False)
- mat_sub = mat[:, sel_cols]
- df = pd.DataFrame(mat_sub.toarray(), index=gene_names)
- print(f" Loaded: {df.shape[0]} genes x {df.shape[1]} cells (downsampled from {n_total})")
- return df
- def preprocess_genes(df, max_genes=3000):
- """Filter low-expression genes and cap gene count for memory."""
- gene_mask = (df.mean(axis=1) >= 0.05) & (df.sum(axis=1) >= 25)
- df = df.loc[gene_mask]
- if df.shape[0] > max_genes:
- top_genes = df.sum(axis=1).nlargest(max_genes).index
- ko_genes = [g for g in GENES_TO_KO if g in df.index]
- keep = set(top_genes) | set(ko_genes)
- df = df.loc[sorted(keep, key=lambda x: df.index.get_loc(x))]
- print(f" After gene filtering: {df.shape[0]} genes")
- return df
- # ──────────────────────────────────────────────
- # Main pipeline using scTenifoldKnk
- # ──────────────────────────────────────────────
- def main():
- print("=" * 60)
- print("Virtual Knockout - scTenifoldKnk")
- print(f"Genes: {', '.join(GENES_TO_KO)}")
- print(f"Sampled cells per network: {N_SAMP_CELLS}")
- print("=" * 60)
- df = load_expression_matrix()
- df = preprocess_genes(df, max_genes=MAX_GENES)
- # Check genes
- for gene in GENES_TO_KO:
- if gene in df.index:
- pct = (df.loc[gene] > 0).sum() / df.shape[1] * 100
- print(f" {gene}: OK ({pct:.0f}% cells)")
- else:
- print(f" {gene}: MISSING!")
- output_dir = BASE_DIR / "results"
- output_dir.mkdir(exist_ok=True)
- summary = []
- total_t0 = time.time()
- for gene in GENES_TO_KO:
- print(f"\n{'='*60}")
- print(f"[KO] {gene} (KGE Novel Rank: {KGE_INFO[gene]['novel_rank']})")
- print(f"{'='*60}")
- t0 = time.time()
- try:
- knk = scTenifoldKnk(
- data=df,
- ko_genes=gene,
- ko_method="default",
- qc_kws={"min_exp_avg": 0.05, "min_exp_sum": 25},
- nc_kws={"n_nets": 10, "n_samp_cells": N_SAMP_CELLS,
- "n_comp": 3, "q": 0.95},
- ma_kws={"d": 30},
- dr_kws={"n_ko_genes": 1},
- )
- dr_df = knk.build()
- elapsed = time.time() - t0
- # Save results
- dr_df.to_csv(output_dir / f"{gene}_vk_results.csv", index=False)
- sig = dr_df[dr_df["adjusted p-value"] < 0.05]
- sig.to_csv(output_dir / f"{gene}_dr_genes.csv", index=False)
- n_sig = len(sig)
- summary.append({
- "Gene": gene,
- "KGE_Novel_Rank": KGE_INFO[gene]["novel_rank"],
- "KGE_Global_Rank": KGE_INFO[gene]["global_rank"],
- "KGE_Score": KGE_INFO[gene]["score"],
- "Known_Regulated_Cells": KGE_INFO[gene]["cells"],
- "Total_DR_genes": len(dr_df),
- "Significant_DR_genes": n_sig,
- "Top_DR_gene": dr_df.iloc[0]["Gene"] if len(dr_df) > 0 else "N/A",
- "Status": "OK",
- })
- print(f" Done in {elapsed:.1f}s | Total: {len(dr_df)} genes, "
- f"Significant (adj.p<0.05): {n_sig}")
- if len(dr_df) > 0:
- print(" Top 5 DR genes:")
- for _, row in dr_df.head(5).iterrows():
- print(f" {row['Gene']}: dist={row['Distance']:.3f}, "
- f"adj.p={row['adjusted p-value']:.4e}")
- except Exception as e:
- print(f" FAILED: {e}")
- summary.append({
- "Gene": gene, "Status": "FAILED",
- "KGE_Novel_Rank": KGE_INFO[gene]["novel_rank"],
- "KGE_Global_Rank": KGE_INFO[gene]["global_rank"],
- "Total_DR_genes": 0, "Significant_DR_genes": 0,
- })
- # Summary
- total_elapsed = time.time() - total_t0
- summary_df = pd.DataFrame(summary)
- summary_df.to_csv(output_dir / "knockout_summary.csv", index=False)
- print(f"\n{'='*60}")
- print(f"ANALYSIS COMPLETE ({total_elapsed:.1f}s total)")
- print(f"{'='*60}")
- print(summary_df.to_string(index=False))
- print(f"\nResults saved to: {output_dir}")
- if __name__ == "__main__":
- main()
run_virtual_knockout.py at commit ba99138, under MIT · at the source
Overview
- Institute of Biomedical and Health Engineering, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China; (C.C.); (X.N.); (Y.M.); (Z.W.); (Z.X.)
- School of Life Sciences, University of Chinese Academy of Sciences, Beijing 100049, China
- Institute of Molecular Physiology, Shenzhen Bay Laboratory, Shenzhen 518132, China
- The Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences, Shenzhen 518055, China
Abstract
The nervous system comprises highly diverse cell types governed by cell-type-specific molecular regulatory programs. However, regulatory evidence is scattered across unstructured literature and described using inconsistent cell-type nomenclature and granularity, hindering systematic integration and cross-study comparison. Here, we construct a neural-cell-centric multimodal knowledge graph that transforms fragmented regulatory evidence into a standardized, computable substrate. We establish a three-level hierarchical cell-type taxonomy anchored to the Cell Ontology (79 nodes), integrate two large-scale human brain single-cell transcriptomic datasets (over 4 million cells) to derive molecular fingerprints, and use a large language model to retain 25,812 curated regulatory evidence records from PubMed abstracts. The resulting Neo4j graph contains 41,532 directed relationships. For knowledge graph embedding, we export a deduplicated non-paper training subgraph containing 19,819 triples over 10,660 entities, supporting cell-type-specific link prediction that prioritizes candidate regulators and markers, illustrated here for microglia. This framework provides a structured basis for cross-study comparison, hypothesis generation and knowledge-guided reasoning in neural cell-type-specific regulation.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.
SiatBioInf/NeuroCellKG
ba99138a38aed3d52a8b97826460ff5430ba467e, 29 May 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
12 files
- scripts/
kg_construction/ , Python, 76 lines, 2 matchesexport_triples.py - scripts/
kg_construction/ , Python, 103 lines, 1 matchimport_hierarchy.py - scripts/
kg_construction/ , Python, 129 lines, 3 matchesimport_marker_genes.py - scripts/
kge_training/ , Python, 190 linescompare_models.py - scripts/
kge_training/ , Python, 303 lines, 2 matchesevaluate_marker_retrieva l.py - scripts/
kge_training/ , Python, 270 linesevaluate_ppi_lp.py - scripts/
kge_training/ , Python, 286 lines, 1 matchevaluate_regulates_lp.py - scripts/
kge_training/ , Python, 517 lines, 1 matchsplit_triples.py - scripts/
kge_training/ , Python, 248 lines, 2 matchestrain_pykeen_model.py - scripts/
virtual_knockout/ , Python, 179 lines, 4 matchesrun_virtual_knockout.py - LICENSE, License, 21 lines
- README.md, Text, 274 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 10 scripts, each with its path and the digest of its content;
- 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
All source code, data files, and Supplementary Materials are publicly available at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 6 keywords, 8 MeSH terms, 4 funders, 33 references.
Cite
This paper
Chen, C., Ni, X., Min, Y., Wang, Z., Xu, Z., Zhang, Y., & Yu, H. (2026). A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence. International journal of molecular sciences, 27(15), 6842. https://
BibTeX
@article{chen2026hierarc
author = {Chen, Chuangyu and Ni, Xiaomin and Min, Yang and Wang, Zhen and Xu, Zhilan and Zhang, Yang and Yu, Hao},
title = {{A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence}},
journal = {International journal of molecular sciences},
year = {2026},
month = jul,
volume = {27},
number = {15},
pages = {6842},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1422-0067},
doi = {10.3390/
url = {https://
pmid = {42589508},
pmcid = {PMC13466513}
}
RIS
TY - JOUR
AU - Chen, Chuangyu
AU - Ni, Xiaomin
AU - Min, Yang
AU - Wang, Zhen
AU - Xu, Zhilan
AU - Zhang, Yang
AU - Yu, Hao
TI - A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence
T2 - International journal of molecular sciences
J2 - Int J Mol Sci
PY - 2026
DA - 2026/
VL - 27
IS - 15
SP - 6842
SN - 1422-0067
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "A Hierarchical Multimodal Knowledge Graph for Neural Cell-Type-Specific Regulation Integrating Single-Cell Transcriptomics and Literature Evidence",
"container-title": "International journal of molecular sciences",
"author": [
{
"family": "Chen",
"given": "Chuangyu"
},
{
"family": "Ni",
"given": "Xiaomin"
},
{
"family": "Min",
"given": "Yang"
},
{
"family": "Wang",
"given": "Zhen"
},
{
"family": "Xu",
"given": "Zhilan"
},
{
"family": "Zhang",
"given": "Yang"
},
{
"family": "Yu",
"given": "Hao"
}
],
"container-title-short":
"volume": "27",
"issue": "15",
"page": "6842",
"DOI": "10.3390/
"PMID": "42589508",
"PMCID": "PMC13466513",
"ISSN": "1422-0067",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
30
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41597-026-07173-8 [code]
- The Cell Ontology in the age of single-cell omics.Journal: Scientific dataIn common: scikit-learn, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
- [2] doi:10.1038/s41514-026-00391-9 [code]
- Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.Journal: npj agingIn common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, cellular / molecular, 2 references
- [3] doi:10.1038/s41467-026-73171-4 [code]
- Dissecting epigenetic heterogeneity in single-cell DNA methylomes with a unified framework.Journal: Nature communicationsIn common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, cellular / molecular, 1 reference
- [4] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: PyTorch, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
- [5] doi:10.1101/gr.281350.125 [code]
- High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND.Journal: Genome researchIn common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, cellular / molecular, 1 reference
- [6] doi:10.3389/fnmol.2026.1844705 [code]
- Risperidone regulates the expression of schizophrenia-related genes in the forebrain of adult male mice.Journal: Frontiers in molecular neuroscienceIn common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, cellular / molecular, 1 reference
- [7] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: scikit-learn, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 2 references
- [8] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, 1 reference
- [9] doi:10.21203/rs.3.rs-9060414/v1 [code]
- Cross-Species Aging Knowledge Integration into Agentic AI Platform Uncovers Conserved MechanismsJournal: Research Square (preprint)In common: PyTorch, scikit-learn, pandas, 2 other tools, cellular / molecular, 1 reference
- [10] doi:10.1038/s41467-026-75723-0 [code]
- Spatial transcriptomics reveals distinct cell type dynamics following opioid dependence in female mice with the common human μ-opioid receptor variant Oprm1 A118G.Journal: Nature communicationsIn common: scikit-learn, pandas, SciPy, 1 other tool, genetics / omics, cellular / molecular, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 10 scripts, and 16 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:732f9d1c63ff381b…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
