OSCR

DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation.

Code ↔ Paper

1 match between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 1 match
  1. [1] § 4. Materials and Methods › 4.3. Data Curation and Presumed-Negative Augmentation ↔ DeepBBB_training_code/BBB_binary_pred_large_ratio/read_smi_n.py, lines 39–43 · score 0.64 · Enamine Hit Locator, library 200k, plated, DrugBank

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 99 lines · 2.5 KB · no license · 1 match

  1. from rdkit import Chem
  2. import glob
  3. import pandas as pd
  4. import sys
  5. input_f=sys.argv[1]
  6. #Test=pd.read_csv("Test_"+num+".csv", header=0, sep=",")
  7. Test=pd.read_csv(input_f+".tsv", header=0, sep="\t")
  8. #TAID,Name,IUPAC Name,PubChem CID,Canonical SMILES,InChIKey,Toxicity Value
  9. #dict_smi={}
  10. #dict_name={}
  11. #dict_t_valu={}
  12. dict_smi_n_v={}
  13. for index, row in Test.iterrows():
  14. smi=row['SMILES']
  15. print (smi)
  16. InChI=row['Inchi']
  17. if str(smi) != 'nan' and str(InChI) != 'nan':
  18. mol = Chem.MolFromSmiles(smi)
  19. #InChI=row['Inchi']
  20. Name=row['compound_name']
  21. print (Name)
  22. T_value=row['BBB+/BBB-']
  23. if T_value=='BBB+':
  24. T_value_n=1
  25. elif T_value=='BBB-':
  26. T_value_n=0
  27. if mol is not None:
  28. dict_smi_n_v[InChI] = [smi,Name,T_value_n]
  29. #dict_name[InChIKey]= Name
  30. # 读取 CSV 文件
  31. approved_df = pd.read_csv('drugbank_approved_structure_links.csv')
  32. experimental_df = pd.read_csv('drugbank_experimental_structure_links.csv')
  33. investigational_df = pd.read_csv('drugbank_investigational_structure_links.csv')
  34. Enamine_df = pd.read_csv('Enamine_Hit_Locator_Library_200K_Set_plated_200000cmpds_20250427.smiles',sep='\t')
  35. # 合并所有 DataFrame
  36. combined_df = pd.concat([approved_df, experimental_df, investigational_df], ignore_index=True)
  37. for index, row in combined_df.iterrows():
  38. smi=row['SMILES']
  39. print (smi)
  40. InChI=row['InChI']
  41. if str(smi) != 'nan' and str(InChI) != 'nan':
  42. mol = Chem.MolFromSmiles(smi)
  43. #InChI=row['Inchi']
  44. Name=row['Name']
  45. print (Name)
  46. T_value_n=0
  47. if mol is not None and InChI not in dict_smi_n_v.keys():
  48. dict_smi_n_v[InChI] = [smi,Name,T_value_n]
  49. Enamine_df = pd.read_csv('Enamine_Hit_Locator_Library_200K_Set_plated_200000cmpds_20250427.smiles',sep='\t')
  50. for index, row in Enamine_df.iterrows():
  51. smi=row['SMILES']
  52. print (smi)
  53. InChI=row['Catalog ID']
  54. if str(smi) != 'nan' and str(InChI) != 'nan':
  55. mol = Chem.MolFromSmiles(smi)
  56. #InChI=row['Inchi']
  57. Name=row['Catalog ID']
  58. print (Name)
  59. T_value_n=0
  60. if mol is not None and InChI not in dict_smi_n_v.keys():
  61. dict_smi_n_v[InChI] = [smi,Name,T_value_n]
  62. import numpy as np
  63. np.save(input_f+'_dict_comb.npy',dict_smi_n_v)
  64. new_dic=np.load(input_f+'_dict_comb.npy', allow_pickle='TRUE').item()
  65. print (new_dic)
  66. import json
  67. with open(input_f+'_dict_comb.json', 'w') as f:
  68. json.dump(dict_smi_n_v, f)
  69. '''
  70. import json
  71. with open('my_dict.json', 'r') as f:
  72. my_dict = json.load(f)
  73. print(my_dict)
  74. '''

read_smi_n.py at commit bd2b9a1, no license · at the source

Overview

Authors: Ziying Xu1, Wei Xia2, Haiqiang Wu1, Haiping Zhang3
  1. School of Pharmacy, Shenzhen University Medical School, Shenzhen University, Shenzhen 518055, China
  2. Department of Chemistry, New York University, New York, NY 10003, USA
  3. Faculty of Pharmaceutical Sciences, Shenzhen University of Advanced Technology, Shenzhen 518055, China
Journal: Pharmaceuticals (Basel, Switzerland), volume 19, issue 8, article 1319
Dates: received 1 July 2026; accepted 11 August 2026; published online 21 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/ph19081319 · PMID 42653814 · PMCID PMC13516556 · OpenAlex W7203887190
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: clinical / translational (subfield)
Methods: Connectivity, Machine learning
Keywords: blood–brain barrier, PAMPA-BBB, graph neural networks, CNS drug discovery, class imbalance, virtual screening, BBB-focused library
Topic: Computational Drug Discovery Methods (Computational Theory and Mathematics, Computer Science), according to OpenAlex
Funding: Shenzhen Science and Technology Program (KCXFZ20230731092802004, JCYJ20220818100804009); National Natural Science Foundation of China (52573336, 52273299); Shenzhen Medical Research Fund (B2404003, D2403008)
Citations: not cited yet (Europe PMC); 34 references in the paper

Abstract

Background/Objectives: Blood–brain barrier (BBB) permeability is a major practical obstacle in central nervous system (CNS) drug discovery, because only a small fraction of drug-like molecules achieve sufficient brain exposure. Methods: We present DeepBBB, a graph-based, data-composition-aware screening workflow for predicting BBB permeability and for constructing BBB-focused screening libraries from commercial chemical space. Rather than introducing a new graph-learning architecture, the workflow combines standard graph convolutional and graph-transformer models with deliberate control of training-set composition, commercial-library filtering, chemical-space profiling, and prospective experimental evaluation. Three model variants were trained on the Blood–Brain Barrier Database (B3DB): a baseline classifier/regressor pair (DeepBBB_V1_BC/RG), a variant trained with a more strongly negative-enriched configuration (DeepBBB_V2_BC), and a graph-transformer counterpart (DeepBBB_trans_BC/RG). Because the sample-level split assignments and per-compound predictions from the original runs were not recoverable, the archived summary metrics are reported descriptively in the main text and are not used to support calibration, scaffold-level validity, or generalization. Applying the workflow to the ChemDiv collection (~1.5 million compounds) and the Enamine REAL lead-like space (~1.7 billion compounds) produced three progressively more stringently filtered BBB-focused libraries (21,991; 4,808,885; and 151,790 compounds). Results: Analysis of available processed data indicated that the predicted BBB-permeable set occupies a compact, BBB-compatible property region. Physicochemical, fragment, and scaffold summaries were interpreted descriptively at the constructed-library level. In a first prospective campaign, one of 12 tested candidate compounds was PAMPA-BBB-positive (all-tested molecular-level positive fraction 8.3%; exact 95% CI 0.2–38.5%). In a second campaign, five of 35 tested candidate compounds were PAMPA-BBB-positive (14.3%; exact 95% CI 4.8–30.3%); 13 compounds were not quantifiable and were not treated as ordinary CNS-negative measurements, and the two campaigns differed in compound source, selection strategy, and assay setting, so the numerical difference is reported descriptively rather than causally. Conclusions: Together, these results support the feasibility of BBB-focused computational filtering and a PAMPA-BBB evaluation workflow for CNS-oriented discovery.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 1 match between paragraphs and lines of code.

haiping1010/DeepBBB

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: bd2b9a115348161bf46dcfc0bfdf69273c5e9131, 1 May 2026
Languages: Python (59), Shell (28)
Size: 135 files, 87 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (34 files), pandas (27 files), RDKit (26 files), PyTorch Geometric (24 files), PyTorch (24 files), NetworkX (8 files), scikit-learn (8 files), SciPy (8 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
88 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 87 scripts, each with its path and the digest of its content;
  • 1 match between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

4.1. Evidence Boundary and Data Availability

This manuscript reports a workflow assembled from models, processed datasets, summary tables, and prospective assay results. To keep claims aligned with the supporting evidence, model-performance values are reported as archived retrospective metrics from the original training/test split; per-sample predictions, exact split assignments, training checkpoints, and random seeds were not re-derived here, so the reported metrics are not described as scaffold-split or external-validation performance. PAMPA-BBB outcomes are reported at the level of per-compound mean permeability and categorical assignment; raw plate-level replicate measurements were not re-analyzed, so no new standard deviations, replicate confidence intervals, or reproducibility statistics are introduced. Physicochemical, fragment, and scaffold analyses are based on available processed descriptor and summary tables. The original ultra-large screening databases (Enamine REAL, ChemDiv) and licensed third-party datasets (DrugBank) are not redistributed; reported library sizes are retained as summary counts. A complete account of data provenance and access routes is given in the Data Availability Statement and Supplementary Information.

Reproduced under the paper's license (CC BY), from the paper cited above.

Data Availability Statement

Model implementations, associated scripts, and usage examples are available in the public GitHub repository (https://github.com/haiping1010/DeepBBB, accessed on 10 August 2026); trained model checkpoints are not included in the repository. The processed summary tables, extracted PAMPA-BBB tables, redrawn figure plot-data, figure scripts, and model-metric summary tables generated for this revision are provided as Supplementary Data accompanying this manuscript. The Enamine REAL lead-like collection, ChemDiv collection, Enamine Hit Locator Library, and DrugBank are third-party resources and cannot be redistributed by the authors. The original sample-level split assignments, per-compound predictions, model checkpoints, random seeds, raw plate-level PAMPA-BBB records, and verified machine-readable structures matching the assay identifiers were not found in the local revision archive and therefore cannot be provided with this revision. A GitHub commit hash corresponding to the original training runs could not be verified from the local archive and is therefore not stated. The original contributions presented in this study are included in the article/Supplementary Materials. Further inquiries can be directed to the corresponding authors.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 7 keywords, 3 funders, 34 references.

Cite

This paper

Xu, Z., Xia, W., Wu, H., & Zhang, H. (2026). DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation. Pharmaceuticals (Basel, Switzerland), 19(8), 1319. https://doi.org/10.3390/ph19081319

BibTeX

@article{xu2026deepbbb,
author = {Xu, Ziying and Xia, Wei and Wu, Haiqiang and Zhang, Haiping},
title = {{DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation}},
journal = {Pharmaceuticals (Basel, Switzerland)},
year = {2026},
month = aug,
volume = {19},
number = {8},
pages = {1319},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {1424-8247},
doi = {10.3390/ph19081319},
url = {https://doi.org/10.3390/ph19081319},
pmid = {42653814},
pmcid = {PMC13516556}
}

RIS

TY - JOUR
AU - Xu, Ziying
AU - Xia, Wei
AU - Wu, Haiqiang
AU - Zhang, Haiping
TI - DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation
T2 - Pharmaceuticals (Basel, Switzerland)
J2 - Pharmaceuticals (Basel)
PY - 2026
DA - 2026/08/21
VL - 19
IS - 8
SP - 1319
SN - 1424-8247
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/ph19081319
UR - https://doi.org/10.3390/ph19081319
LA - en
ER -

CSL-JSON

{
"id": "10.3390/ph19081319",
"type": "article-journal",
"title": "DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation",
"container-title": "Pharmaceuticals (Basel, Switzerland)",
"author": [
{
"family": "Xu",
"given": "Ziying"
},
{
"family": "Xia",
"given": "Wei"
},
{
"family": "Wu",
"given": "Haiqiang"
},
{
"family": "Zhang",
"given": "Haiping"
}
],
"container-title-short": "Pharmaceuticals (Basel)",
"volume": "19",
"issue": "8",
"page": "1319",
"DOI": "10.3390/ph19081319",
"PMID": "42653814",
"PMCID": "PMC13516556",
"ISSN": "1424-8247",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/ph19081319",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
21
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pone.0345854 [code]
Shedding light on neural learning to rank models for anticancer drug prioritization.
Journal: PloS one
In common: RDKit, PyTorch Geometric, NetworkX, 5 other tools
[2] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: RDKit, PyTorch Geometric, NetworkX, 5 other tools
[3] doi:10.1016/j.apsb.2026.05.021 [code]
Design, preclinical evaluation, and multicenter phase 1 clinical study of HZ-A-018 for relapsed or refractory central nervous system lymphoma.
Journal: Acta pharmaceutica Sinica. B
In common: RDKit, PyTorch Geometric, NetworkX, 4 other tools, clinical / translational
[4] doi:10.1093/bioinformatics/btag153 [code]
MAISNet: a multi-species integrated graph neural network for acetylcholinesterase inhibitor screening.
Journal: Bioinformatics (Oxford, England)
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, clinical / translational
[5] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits <i>Porphyromonas gingivalis</i> Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools
[6] doi:10.1038/s41586-026-10670-w [code]
Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
Journal: Nature
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools
[7] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools
[8] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools
[9] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: RDKit, NetworkX, PyTorch, 4 other tools
[10] doi:10.1186/s13321-026-01177-7 [code]
A pipeline for developing AI-driven models to predict molecular initiating events: a case study on neural tube defects.
Journal: Journal of cheminformatics
In common: RDKit, NetworkX, PyTorch, 4 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.