A benchmark dataset and interpretable deep learning framework for drug-induced developmental neurotoxicity prediction.
The 13 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Materials and methods › Model construction and training strategy › Model training and evaluation ↔ 3-DNN_MACCS_modeling/DNN-test_selection.py, lines 33–67 · score 0.70 · monitoring validation, model selection, cross validation, weights, optimal, epochs
- [2] § Materials and methods › Performance metrics and statistical analysis ↔ 3-DNN_MACCS_modeling/DNN-independent validation.py, lines 208–293 · score 0.69 · F1 score, Matthew, TN, TP, FN, precision
- [3] § Materials and methods › Performance metrics and statistical analysis ↔ 3-DNN_MACCS_modeling/DNN-test_selection.py, lines 560–599 · score 0.67 · predicted toxic, evaluation metrics, F1 score, precision, recall, accuracy
- [4] § Materials and methods › Model construction and training strategy › Model training and evaluation ↔ 3-DNN_MACCS_modeling/DNN-training-5-fold-CV.py, lines 34–65 · score 0.65 · fold cross validation, network architectures, overfitting, epochs, training, Model
- [5] § Results and discussion › Analysis of the training stability and generalization ability of the DNN_MACCS model ↔ 3-DNN_MACCS_modeling/DNN-test_selection.py, lines 889–968 · score 0.64 · developmental neurotoxicity prediction, validation AUC, F1 score, dnn maccs, overfitting, configuration
- [6] § Materials and methods › Model interpretability analysis ↔ 4-applicability_domain_SHAP_analysis/SHAP-analysisl.py, lines 1056–1121 · score 0.62 · SHapley, exPlanations, Model interpretability, Additive, global, transparency
- [7] § Materials and methods › Benchmark dataset preparation › Data partitioning of the benchmark dataset ↔ 1-Benchmark_Dataset_Preparation/plot_umap_chemical_space.py, the whole file · a weak match · score 0.59 · chemical space, DNT positive, UMAP, benchmark, SMILES, fingerprints
- [8] § Results and discussion › Strengths and limitations ↔ 4-applicability_domain_SHAP_analysis/Application_Domain_UMAP_tSNE.py, lines 429–462 · score 0.58 · broader chemical, applicability domain, experimentally validated, model generalizability, chemical space, expand
- [9] § Results and discussion › Analysis of the training stability and generalization ability of the DNN_MACCS model ↔ 3-DNN_MACCS_modeling/DNN-test_selection.py, lines 889–968 · score 0.56 · potential overfitting, best validation, dnn maccs, loss, epochs, accuracy
- [10] § Materials and methods › Benchmark dataset preparation › Preprocessing of positive samples ↔ 1-Benchmark_Dataset_Preparation/DNTREF_curation.py, lines 25–46 · score 0.54 · heavy atoms, mixtures, parent, curation, SMILES
- [11] § Materials and methods › Chemical space visualization and applicability domain analysis ↔ 4-applicability_domain_SHAP_analysis/Application_Domain_UMAP_tSNE.py, lines 198–201 · score 0.52 · Dimensionality reduction, UMAP, Neighbor, SNE, domain
- [12] § Materials and methods › Benchmark dataset preparation ↔ 3-DNN_MACCS_modeling/DNN-test_selection.py, lines 1014–1134 · score 0.52 · model training, developmental neurotoxicity prediction, pipeline, selection, preprocessing
- [13] § Materials and methods › Benchmark dataset preparation › Preprocessing of presumed negative samples and construction of a balanced benchmark dataset ↔ 1-Benchmark_Dataset_Preparation/hard_negative_matching.py, lines 54–90 · score 0.50 · hard negative matching, greedy, MW, log, benchmark
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 1,147 lines · 51 KB · no license · 5 matches
DNN-test_selection.py at commit 514073b, no license · at the source
Overview
- Shandong Key Laboratory of Digital Diagnosis and Treatment of Thoracic Oncology, Shandong Engineering Research Center of Precision Diagnosis and Treatment Technology for Neuro-Oncology, Department of Clinical Pharmacy, The First Affiliated Hospital of Shandong First Medical University, Shandong Provincial Qianfoshan Hospital Jinan 250014 China
- Department of Pharmacy, Beijing Hospital, National Center of Gerontology, Institute of Geriatric Medicine, Chinese Academy of Medical Sciences Beijing 100730 China
- State Key Laboratory of Neurology and Oncology Drug Development, The First Affiliated Hospital of Shandong First Medical University, Shandong Provincial Qianfoshan Hospital Jinan 250014 China
Abstract
Developmental neurotoxicity (DNT) represents a critical yet underevaluated toxicity endpoint within current chemical safety assessment frameworks, particularly in the context of drug development. In the present study, we constructed a standardized benchmark dataset comprising 2724 structurally curated compounds, including 1362 DNT-positive chemicals derived from the Developmental Neurotox Reference List (DNTREF) of the U.S. Environmental Protection Agency (EPA) and 1362 property-matched presumed negative compounds (non-neuroactive drug-like compounds) selected from the ChEMBL database. Using this benchmark dataset, we systematically evaluated seven representative quantitative structure activity relationship (QSAR) modeling paradigms, covering traditional machine learning, deep neural networks, molecular language models, and graph neural networks. Across all models, the area under the receiver operating characteristic curve (AUC) values ranged from 0.85 to 0.90 on an independent test set. Among these, a deep neural network based on MACCS fingerprints achieved the highest sensitivity in identifying DNT-positive compounds, supporting its utility in early-stage drug safety screening where minimizing false negatives is critical. Furthermore, SHAP-based interpretability analysis revealed key structural motifs associated with DNT, highlighting the roles of hydrophobic aromatic frameworks, electrophilic reactivity, and polar functional groups in modulating predicted neurodevelopmental risk. Collectively, this study provides a reproducible benchmark dataset and an interpretable QSAR framework that supports early DNT risk prioritization and offers actionable structural insights for medicinal chemistry optimization in drug discovery.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.
lixiao1688/DNT_Benchmark_Dataset_Model
514073b77308152898322ed1e5380f12c0f7b426, 27 January 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
9 files, not copied: shown from their source
OSCR keeps no copy of these files: this repository has no license that allows it. The reader above shows each one from its source, fetched by your browser at commit 514073b, when its fingerprint is the one OSCR verified. How this works.
- 1-Benchmark_Dataset_Prep
aration/ — Python, 136 lines, 1 match, shown from its sourceDNTREF_curation.py - 1-Benchmark_Dataset_Prep
aration/ — Python, 105 lines, 1 match, shown from its sourcehard_negative_matching.p y - 1-Benchmark_Dataset_Prep
aration/ — Python, 59 lines, 1 match, shown from its sourceplot_umap_chemical_space .py - 3-DNN_MACCS_modeling/
DNN-independent validation.py — Python, 778 lines, 1 match, shown from its source - 3-DNN_MACCS_modeling/
DNN-test_selection.py — Python, 1,147 lines, 5 matches, shown from its source - 3-DNN_MACCS_modeling/
DNN-training-5-fold-CV.p — Python, 628 lines, 1 match, shown from its sourcey - 4-applicability_domain_S
HAP_analysis/ — Python, 462 lines, 2 matches, shown from its sourceApplication_Domain_UMAP_ tSNE.py - 4-applicability_domain_S
HAP_analysis/ — Python, 1,220 lines, 1 match, shown from its sourceSHAP-analysisl.py - README.md — Text, 46 lines, shown from its source
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 8 scripts, each with its path and the digest of its content;
- 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data and model availability
The developmental neurotoxicity benchmark dataset constructed in this study includes standardized SMILES representations of all compounds, corresponding binary toxicity labels, and the fixed training, validation, and independent test set partitions used for model development and evaluation. These data are deposited in a public repository to ensure full reproducibility of the reported results. The dataset is expected to serve as a valuable resource for the computational toxicology and drug discovery communities, enabling further methodological development and benchmarking.
In addition, the scripts for MACCS fingerprint generation, the trained weights of the optimal DNN_MACCS model, and all scripts used for model training, inference, and SHAP-based interpretability analysis will be made publicly available.
All scripts for AD analysis and chemical space visualization using UMAP and t-SNE will also be released to support transparent assessment of model scope and result interpretation.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Data availability
The data utilized in this study were derived from the DNTREF from U.S. EPA, and ChEMBL database. Complete data (including processed datasets, and all scripts necessary to reproduce the study) is available on GitHub at https://
Supplementary information (SI): detailed data processing for benchmark dataset; detailed network architectures, hyperparameter settings, and training procedures; a summary of predictive performance for all models on validation set; unified chemical standardization process for data collection, structure preprocessing, and sample selection; training history and model comparison of the DNN_MACCS model. See DOI: https://
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 8 authors, 4 funders, 46 references.
Cite
This paper
Ma, H., Zhang, W., Liu, F., Ni, R., Wang, X., Sun, Y., Sun, X., & Li, X. (2026). A benchmark dataset and interpretable deep learning framework for drug-induced developmental neurotoxicity prediction. RSC advances, 16(36), 37450-37464. https://
BibTeX
@article{ma2026benchmark
author = {Ma, Hongting and Zhang, Wenhui and Liu, Fengxi and Ni, Rong and Wang, Xue and Sun, Yanying and Sun, Xuelin and Li, Xiao},
title = {{A benchmark dataset and interpretable deep learning framework for drug-induced developmental neurotoxicity prediction}},
journal = {RSC advances},
year = {2026},
month = jul,
volume = {16},
number = {36},
pages = {37450--37464},
publisher = {Royal Society of Chemistry},
issn = {2046-2069},
doi = {10.1039/
url = {https://
pmid = {42440932},
pmcid = {PMC13334446}
}
RIS
TY - JOUR
AU - Ma, Hongting
AU - Zhang, Wenhui
AU - Liu, Fengxi
AU - Ni, Rong
AU - Wang, Xue
AU - Sun, Yanying
AU - Sun, Xuelin
AU - Li, Xiao
TI - A benchmark dataset and interpretable deep learning framework for drug-induced developmental neurotoxicity prediction
T2 - RSC advances
J2 - RSC Adv
PY - 2026
DA - 2026/
VL - 16
IS - 36
SP - 37450
EP - 37464
SN - 2046-2069
PB - Royal Society of Chemistry
DO - 10.1039/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1039/
"type": "article-journal",
"title": "A benchmark dataset and interpretable deep learning framework for drug-induced developmental neurotoxicity prediction",
"container-title": "RSC advances",
"author": [
{
"family": "Ma",
"given": "Hongting"
},
{
"family": "Zhang",
"given": "Wenhui"
},
{
"family": "Liu",
"given": "Fengxi"
},
{
"family": "Ni",
"given": "Rong"
},
{
"family": "Wang",
"given": "Xue"
},
{
"family": "Sun",
"given": "Yanying"
},
{
"family": "Sun",
"given": "Xuelin"
},
{
"family": "Li",
"given": "Xiao"
}
],
"container-title-short":
"volume": "16",
"issue": "36",
"page": "37450-37464",
"DOI": "10.1039/
"PMID": "42440932",
"PMCID": "PMC13334446",
"ISSN": "2046-2069",
"publisher": "Royal Society of Chemistry",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
6
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: RDKit, SHAP, Keras, 8 other tools
- [2] doi:10.1093/nar/gkag706 [code]
- scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.Journal: Nucleic acids researchIn common: RDKit, Keras, UMAP, 7 other tools, 1 reference
- [3] doi:10.1038/s41598-026-48613-0 [code]
- An snRNA-seq aging clock for the fruit fly head sheds light on sex-biased aging.Journal: Scientific reportsIn common: SHAP, Keras, TensorFlow, 6 other tools, 2 references
- [4] doi:10.1126/sciadv.aed3650 [code]
- Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &
lt;i& gt;m/ z& lt;/ i& gt; mapping and exploration. Journal: Science advancesIn common: SHAP, Keras, UMAP, 6 other tools, methods / tools - [5] doi:10.1177/13872877261453512 [code]
- Fluorescence spectroscopy and machine learning methods for detection of Alzheimer's disease from circulating white blood cells.Journal: Journal of Alzheimer's disease : JADIn common: Keras, UMAP, TensorFlow, 6 other tools, methods / tools
- [6] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: Keras, UMAP, TensorFlow, 6 other tools, methods / tools
- [7] doi:10.3390/s26175327 [code]
- Subject Identity Confounds qEEG Emotion Recognition on DEAP and DREAMER.Journal: Sensors (Basel, Switzerland)In common: SHAP, Keras, TensorFlow, 6 other tools
- [8] doi:10.1371/journal.pcbi.1014615 [code]
- Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.Journal: PLoS computational biologyIn common: SHAP, Keras, TensorFlow, 6 other tools
- [9] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: SHAP, Keras, TensorFlow, 6 other tools
- [10] doi:10.1038/s41598-026-47047-y [code]
- Imaging-genetics-based dementia risk prediction using deep survival neural networks in the Rotterdam Study.Journal: Scientific reportsIn common: SHAP, Keras, TensorFlow, 6 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 8 scripts, and 13 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:5e9110b98771798d…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
