OSCR

TSProm: deep learning framework to predict tissue-specific regulatory logic.

Code ↔ Paper

16 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 16 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Results › Deciphering brain regulatory language and its’ disease relevance ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 18–31 · score 0.83 · C2H2 zinc finger, basic helix loop, bHLH, ZNFs, motifs
  2. [2] § Materials and methods › Fine tuning of DFMs ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 14–54 · score 0.82 · gena lm bigbird, AIRI Institute, fine tuned, t2t, classify, tokenizes
  3. [3] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 18–31 · score 0.74 · basic helix loop, zinc finger, bHLH, ZnF, Overlap, motifs
  4. [4] § Results › TSp promoter prediction models ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 142–219 · score 0.66 · weight decay, F1 scores, fine tuning, epochs, accuracy, batch
  5. [5] § Materials and methods › Attention-based motif discovery › Motif positional analysis ↔ TSProm.zip/src/3_attention/2A_save_meme.py, lines 394–447 · score 0.65 · gap free, attention score, heuristic, histograms, positions, tokens
  6. [6] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/notebooks/runs/3_TransSHAP.ipynb, lines 94–121 · score 0.65 · KernelExplainer, mask token, SHAP, logits, background, tokenized
  7. [7] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/src/3_attention/3_SHAP.py, lines 85–111 · score 0.65 · KernelExplainer, mask token, SHAP, logits, background, tokenized
  8. [8] § DNA foundation models ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 14–54 · score 0.64 · Gena LM, fine tuned, sequence length, batch, tokenization, training
  9. [9] § Materials and methods › Model evaluation metrics ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 142–219 · score 0.63 · F1 Score, fine tuned, correlation, Recall, Precision, metrics
  10. [10] § Materials and methods › Fine tuning of DFMs ↔ TSProm.zip/run_scripts/1_dnabert2_finetune.sh, lines 3–51 · score 0.62 · 2–117, fine tuned, zhihan1996, activated, DNABERT2, tokenizes
  11. [11] § Results › TSp promoter prediction models ↔ TSProm.zip/run_scripts/0_generate_data_run.sh, the whole file · a weak match · score 0.61 · MouseTranstex, TSProm, trained, tissues, promoter, model
  12. [12] § Results › TSp promoter prediction models ↔ TSProm.zip/run_scripts/1_dnabert2_finetune.sh, lines 3–51 · score 0.58 · weight decay, Hyperparameter, fine tuning, epochs, batch, DNABERT2
  13. [13] § Datasets ↔ TSProm.zip/run_scripts/0_generate_data_run.sh, the whole file · a weak match · score 0.58 · Mouse TransTEx, fine tuning, liver, testis, species, brain
  14. [14] § Results › Language of the global tissue specificity ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 39–157 · score 0.57 · C2H2 ZNFs, bHLH, hits, global, TF, gene
  15. [15] § Datasets › Varying sequence lengths tested ↔ TSProm.zip/src/0_generate_data/make_modelA_negSet.sh, lines 1–52 · score 0.53 · promoter window, configurations, mm39, hg38, genomes, Mouse
  16. [16] § Materials and methods › Model evaluation metrics ↔ TSProm.zip/src/1_finetune/DNABERT2_AttentionExtracted.py, lines 232–250 · score 0.50 · F1 Score, correlation, Recall, Precision, metrics, Accuracy

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 286 lines · 11 KB · CC-BY-4.0 · 4 matches

  1. # Usage instructions:
  2. # conda activate dnabert2.0
  3. # # cd /data/private/psurana/ramanaServer_code/scripts/OtherModel_FineTune/final
  4. # # CUDA_VISIBLE_DEVICES=0,1,2 TOKENIZERS_PARALLELISM=false PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True nohup python aug_2025_genalm_1.py > finetune_all.log 2>&1 &
  5. # cd /data/private/psurana/TSProm/src/1_finetune/
  6. # CUDA_VISIBLE_DEVICES=5,6,7 TOKENIZERS_PARALLELISM=false PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python FineTune_GENALM.py
  7. # export NEPTUNE_PROJECT="lavi.ana/GenaLM-human-mouse-Finetune"
  8. # export TOKENIZERS_PARALLELISM=false
  9. def run_genalm_finetune(data_base_path, output_base_path, tissue_folder, length,
  10. learning_rate=2e-5, batch_size=12, max_length=2000, num_train_epochs=10):
  11. """
  12. Fine-tune GENA-LM on tissue-specific promoter data.
  13. Parameters:
  14. - data_base_path (str): Root directory containing all data folders
  15. - output_base_path (str): Where to save model and evaluation results
  16. - tissue_folder (str): e.g., "tsp_brain_low"
  17. - length (str): e.g., "2000", "3000", "4000"
  18. - learning_rate (float): Training learning rate
  19. - batch_size (int): Batch size per device
  20. - max_length (int): Max token length for tokenizer
  21. - num_train_epochs (int): Number of training epochs
  22. """
  23. import torch
  24. import os
  25. import pandas as pd
  26. import json
  27. from transformers import AutoTokenizer, AutoModel, TrainingArguments, Trainer
  28. from transformers import DataCollatorWithPadding
  29. from datasets import Dataset, DatasetDict
  30. import numpy as np
  31. from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, matthews_corrcoef
  32. import importlib
  33. # Debug print
  34. print(f"🔧 Running with: LR={learning_rate}, Batch={batch_size}, Epochs={num_train_epochs}, Max_len={max_length}")
  35. print(f"📁 Tissue: {tissue_folder} | Length: {length}")
  36. print(f"🔢 CUDA devices available: {torch.cuda.device_count()}")
  37. # Model selection based on sequence length
  38. if length == "2000":
  39. tokenizer = AutoTokenizer.from_pretrained('AIRI-Institute/gena-lm-bert-large-t2t')
  40. model_name = 'AIRI-Institute/gena-lm-bert-large-t2t'
  41. model_class_name = "BertForSequenceClassification"
  42. else: # For 3000 and 4000
  43. tokenizer = AutoTokenizer.from_pretrained('AIRI-Institute/gena-lm-bigbird-base-t2t')
  44. model_name = 'AIRI-Institute/gena-lm-bigbird-base-t2t'
  45. model_class_name = "BigBirdForSequenceClassification"
  46. # Fix padding token if needed
  47. if tokenizer.pad_token is None:
  48. tokenizer.pad_token = tokenizer.eos_token
  49. # Load base model to get module info
  50. model = AutoModel.from_pretrained(model_name, trust_remote_code=True)
  51. gena_module_name = model.__class__.__module__
  52. print(f"📦 Using module: {gena_module_name}")
  53. # Get classification model class
  54. cls = getattr(importlib.import_module(gena_module_name), model_class_name)
  55. # Create classification model
  56. model = cls.from_pretrained(model_name, num_labels=2)
  57. print(f'🧠 Classification head: {model.classifier}')
  58. # Setup paths
  59. trial_data = os.path.join(data_base_path, length, tissue_folder)
  60. output_finish_path = os.path.join(output_base_path, tissue_folder, length, f"lr_{learning_rate:.0e}")
  61. # Create output directory
  62. os.makedirs(output_finish_path, exist_ok=True)
  63. print(f"📂 Output directory: {output_finish_path}")
  64. # Load data with debugging
  65. try:
  66. dev = pd.read_csv(os.path.join(trial_data, 'dev.csv'))
  67. train = pd.read_csv(os.path.join(trial_data, 'train.csv'))
  68. test = pd.read_csv(os.path.join(trial_data, 'test.csv'))
  69. # Debug: Check column names and data types
  70. print(f"📊 Train columns: {list(train.columns)}")
  71. print(f"🔍 Sample train data:\n{train.head(2)}")
  72. print(f"🔍 Sequence column type: {type(train['Sequence'].iloc[0])}")
  73. print(f"📊 Data loaded - Train: {len(train)}, Dev: {len(dev)}, Test: {len(test)}")
  74. except Exception as e:
  75. print(f"❌ Error loading data: {e}")
  76. raise
  77. # Clean and prepare data - remove pandas index column that might cause issues
  78. train = train.reset_index(drop=True)
  79. dev = dev.reset_index(drop=True)
  80. test = test.reset_index(drop=True)
  81. # Convert to the format expected by the model
  82. def prepare_data(df):
  83. return {
  84. 'sequence': df['Sequence'].astype(str).tolist(),
  85. 'label': df['Label'].astype(int).tolist()
  86. }
  87. train_data = prepare_data(train)
  88. dev_data = prepare_data(dev)
  89. test_data = prepare_data(test)
  90. # Create datasets from dictionaries
  91. train_dataset = Dataset.from_dict(train_data)
  92. dev_dataset = Dataset.from_dict(dev_data)
  93. test_dataset = Dataset.from_dict(test_data)
  94. dataset = DatasetDict({
  95. "train": train_dataset,
  96. "validation": dev_dataset,
  97. "test": test_dataset,
  98. })
  99. # Debug: Check first example
  100. print("🔍 First train example:", dataset["train"][0])
  101. # Tokenization - simpler approach
  102. def preprocess_function(examples):
  103. # Tokenize the sequences
  104. tokenized = tokenizer(
  105. examples["sequence"],
  106. truncation=True,
  107. max_length=max_length,
  108. padding=False, # Let DataCollator handle padding
  109. return_tensors=None
  110. )
  111. # Keep the labels
  112. tokenized["labels"] = examples["label"] # Use "labels" not "label"
  113. return tokenized
  114. tokenized_dataset = dataset.map(preprocess_function, batched=True)
  115. # Remove original columns to avoid conflicts
  116. tokenized_dataset = tokenized_dataset.remove_columns(["sequence", "label"])
  117. # Debug: Check tokenized data
  118. print("🔍 Tokenized sample:", {k: (type(v), len(v) if isinstance(v, list) else v) for k, v in tokenized_dataset["train"][0].items()})
  119. # Metrics computation
  120. def compute_metrics(eval_pred):
  121. predictions, labels = eval_pred
  122. predictions = np.argmax(predictions, axis=1)
  123. return {
  124. "accuracy": accuracy_score(labels, predictions),
  125. "precision": precision_score(labels, predictions, zero_division=0),
  126. "recall": recall_score(labels, predictions, zero_division=0),
  127. "f1": f1_score(labels, predictions, zero_division=0),
  128. "matthews_correlation": matthews_corrcoef(labels, predictions)
  129. }
  130. # Create data collator
  131. data_collator = DataCollatorWithPadding(tokenizer=tokenizer, return_tensors="pt")
  132. # Training arguments
  133. training_args = TrainingArguments(
  134. output_dir=output_finish_path,
  135. learning_rate=learning_rate,
  136. lr_scheduler_type="constant_with_warmup",
  137. warmup_ratio=0.1,
  138. optim='adamw_torch',
  139. weight_decay=0.0,
  140. per_device_train_batch_size=batch_size,
  141. per_device_eval_batch_size=batch_size,
  142. num_train_epochs=num_train_epochs,
  143. evaluation_strategy="epoch",
  144. save_strategy="epoch",
  145. logging_strategy="epoch",
  146. load_best_model_at_end=True,
  147. dataloader_pin_memory=False,
  148. remove_unused_columns=False,
  149. report_to=None,
  150. # Added for better multi-GPU support
  151. dataloader_num_workers=4,
  152. fp16=True, # Enable mixed precision for better memory usage
  153. )
  154. # Create trainer
  155. trainer = Trainer(
  156. model=model,
  157. args=training_args,
  158. train_dataset=tokenized_dataset["train"],
  159. eval_dataset=tokenized_dataset["validation"],
  160. tokenizer=tokenizer,
  161. data_collator=data_collator,
  162. compute_metrics=compute_metrics,
  163. )
  164. print("🚀 Starting training...")
  165. trainer.train()
  166. # Evaluate on test set
  167. print("📈 Evaluating on test set...")
  168. results = trainer.evaluate(eval_dataset=tokenized_dataset["test"])
  169. print("📊 Test Results:")
  170. for key, value in results.items():
  171. if isinstance(value, float):
  172. print(f" {key}: {value:.4f}")
  173. else:
  174. print(f" {key}: {value}")
  175. # Save results
  176. output_json_path = os.path.join(output_finish_path, "eval_results.json")
  177. with open(output_json_path, "w") as f:
  178. json.dump(results, f, indent=4)
  179. print(f"💾 Results saved to: {output_json_path}")
  180. # Cleanup memory
  181. del model, trainer
  182. torch.cuda.empty_cache()
  183. print("Memory cleaned up")
  184. # ============================================================================
  185. # Main execution loop with error handling
  186. # ============================================================================
  187. # tissue_dirs = [
  188. # "tsp_brain_low", "tsp_brain_tenh", "tsp_liver_low", "tsp_liver_tenh",
  189. # "tsp_testis_low", "tsp_brain_null", "tsp_brain_wide",
  190. # "tsp_liver_null", "tsp_liver_wide", "tsp_testis_null", "tsp_testis_wide", "tsp_testis_tenh"
  191. # ]
  192. # tissue_dirs =["tsp_spleen_null", "tsp_spleen_wide", "tsp_spleen_tenh", "tsp_spleen_low",
  193. # "tsp_muscle_null", "tsp_muscle_wide"]
  194. tissue_dirs = ["tsp_muscle_tenh", "tsp_muscle_low"]
  195. lengths = ["3000", "4000", "2000"]
  196. length_map = {"2000": 2000, "3000": 3000, "4000": 4000}
  197. learning_rates = [2e-5, 3e-4, 5e-6]
  198. import os
  199. print("🎯 Starting GENA-LM fine-tuning experiments...")
  200. print(f"📋 Total experiments: {len(tissue_dirs)} × {len(lengths)} × {len(learning_rates)} = {len(tissue_dirs) * len(lengths) * len(learning_rates)}")
  201. experiment_count = 0
  202. total_experiments = len(tissue_dirs) * len(lengths) * len(learning_rates)
  203. for tissue in tissue_dirs:
  204. for length in lengths:
  205. for lr in learning_rates:
  206. experiment_count += 1
  207. batch_size = 8 if "testis" in tissue else 12
  208. lr_tag = f"lr_{lr:.0e}"
  209. output_dir = os.path.join(
  210. "/data/projects/dna/pallavi/DNABERT_runs/DATA_RUN/gena-lm_finetune/unique_TSS/human_mouse",
  211. tissue, length, lr_tag
  212. )
  213. print(f"\n{'='*80}")
  214. print(f"🧪 Experiment {experiment_count}/{total_experiments}")
  215. print(f"🧬 Tissue: {tissue} | Length: {length} | LR: {lr} | Batch: {batch_size}")
  216. print(f"{'='*80}")
  217. if os.path.exists(os.path.join(output_dir, "eval_results.json")):
  218. print(f"✅ Skipping: {tissue} | Length: {length} | LR: {lr} (Already completed)")
  219. continue
  220. try:
  221. run_genalm_finetune(
  222. data_base_path="/data/projects/dna/pallavi/data_TSp_Vs_Rest",
  223. output_base_path="/data/projects/dna/pallavi/DNABERT_runs/DATA_RUN/gena-lm_finetune/unique_TSS/human_mouse",
  224. tissue_folder=tissue,
  225. length=length,
  226. learning_rate=lr,
  227. batch_size=batch_size,
  228. max_length=length_map[length],
  229. num_train_epochs=10
  230. )
  231. print(f"✅ Completed: {tissue} | Length: {length} | LR: {lr}")
  232. except Exception as e:
  233. print(f"❌ Error for {tissue} | Length: {length} | LR: {lr}: {str(e)}")
  234. import traceback
  235. print(f"📝 Full traceback:\n{traceback.format_exc()}")
  236. continue
  237. print(f"\n🎉 All experiments completed! Check logs for any errors.")

FineTune_GENALM.py, under CC-BY-4.0 · at the source

Overview

Authors: Pallavi Surana1, Pratik Dutta1, Nimisha Papineni1, Rekha Sathian1, Zhihan Zhou2, Han Liu2, Ramana V Davuluri1
  1. Department of Biomedical Informatics, Stony Brook University, NY 11794, United States
  2. Department of Computer Science, Northwestern University, IL 60208, United States
Institutions: Stony Brook University (United States); Northwestern University (United States)
Journal: NAR genomics and bioinformatics, volume 8, issue 2, article lqag050
Dates: received 30 November 2025; accepted 10 March 2026; published online 3 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1093/nargab/lqag050 · PMID 42244857 · PMCID PMC13233143 · OpenAlex W7163593589
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism)
Methods: Statistics, Preprocessing, Machine learning
MeSH: Deep Learning*, Gene Expression Regulation*, Promoter Regions, Genetic*, Brain, Humans, Liver, Organ Specificity, Testis, Transcription Factors (* major topic)
Topic: Genomics and Chromatin Dynamics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Library of Medicine (R01LM013722); National Institutes of Health
Citations: cited by 1 paper (Europe PMC); 56 references in the paper

Abstract

Characterizing tissue-specific (TSp) gene expression is crucial for understanding development and disease; however, traditional expression-based methods often overlook the latent “regulatory grammar” embedded in non-coding DNA, particularly across distal promoter regions. Here, we introduce TSProm, a framework that adapts a DNA foundation model (DNABERT2) to decode the regulatory logic of TSp promoters at the isoform level. Our contributions are two-fold: (i) a comparative design that trains two specialized models: Model A for general promoter biology and Model B for tissue-specific regulation enabling precise isolation of sequence motifs surrounding the transcription start site that uniquely define tissue identity; (ii) we develop an explainable AI module that integrates attention-based motif discovery with model-agnostic SHAP analysis to yield cross-validated interpretations of learned features. Applying TSProm to human brain, liver, and testis promoters, we identified clinically relevant transcription factors (TFs) in the brain, including SP1, MYC, and HES6, whose associations with gliomas and neuroblastomas highlight clinical relevance. Moreover, our results highlight C2H2 zinc finger proteins as a dominant family shaping the global landscape of TSp gene regulation. TSProm provides an interpretable and generalizable framework for identifying tissue-specific regulatory elements, offering powerful computational tools to investigate gene regulation in both normal and disease contexts.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.

figshare 30747257

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 2 files
Software Heritage: not checked
Found in: “Data availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: pandas (23 files), NumPy (22 files), PyTorch (20 files), Hugging Face Transformers (20 files), Matplotlib (18 files), SciPy (12 files), statsmodels (10 files), Biopython (8 files), seaborn (8 files), SHAP (8 files), tidyverse (8 files), data.table (7 files), caret (5 files), scikit-learn (4 files), ggplot2 (3 files), BEDTools (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
44 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 43 scripts, each with its path and the digest of its content;
  • 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

TSProm code and Model weights: https://doi.org/10.6084/m9.figshare.30747257

Isoform expression data: GTEx (see GTEx Consortium 2015, 2020).

Human TransTEx groupings and code: Described in Surana et al. (2024) and available at TransTExdb.

Mouse

BodyMap dataset: Available from NCBI BioProject PRJNA375882 (Li et al., 2017).

Hugging Face DNA language models (Fishman et al., 2025; Zhou et al., 2021; Dalla Torre et al., 2025):

AIRI-Institute/gena-lm-bert-large-t2t

AIRI-Institute/gena-lm-bigbird-base-t2t

InstaDeepAI/nucleotide-transformer-500m-1000g

jaandoui/DNABERT2-AttentionExtracted

zhihan1996/DNABERT-2–117M

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 9 MeSH terms, 2 funders, 50 references.

Cite

This paper

Surana, P., Dutta, P., Papineni, N., Sathian, R., Zhou, Z., Liu, H., & Davuluri, R. V. (2026). TSProm: deep learning framework to predict tissue-specific regulatory logic. NAR genomics and bioinformatics, 8(2), lqag050. https://doi.org/10.1093/nargab/lqag050

BibTeX

@article{surana2026tsprom,
author = {Surana, Pallavi and Dutta, Pratik and Papineni, Nimisha and Sathian, Rekha and Zhou, Zhihan and Liu, Han and Davuluri, Ramana V},
title = {{TSProm: deep learning framework to predict tissue-specific regulatory logic}},
journal = {NAR genomics and bioinformatics},
year = {2026},
month = jun,
volume = {8},
number = {2},
pages = {lqag050},
publisher = {Oxford University Press},
issn = {2631-9268},
doi = {10.1093/nargab/lqag050},
url = {https://doi.org/10.1093/nargab/lqag050},
pmid = {42244857},
pmcid = {PMC13233143}
}

RIS

TY - JOUR
AU - Surana, Pallavi
AU - Dutta, Pratik
AU - Papineni, Nimisha
AU - Sathian, Rekha
AU - Zhou, Zhihan
AU - Liu, Han
AU - Davuluri, Ramana V
TI - TSProm: deep learning framework to predict tissue-specific regulatory logic
T2 - NAR genomics and bioinformatics
J2 - NAR Genom Bioinform
PY - 2026
DA - 2026/06/03
VL - 8
IS - 2
SP - lqag050
SN - 2631-9268
PB - Oxford University Press
DO - 10.1093/nargab/lqag050
UR - https://doi.org/10.1093/nargab/lqag050
LA - en
ER -

CSL-JSON

{
"id": "10.1093/nargab/lqag050",
"type": "article-journal",
"title": "TSProm: deep learning framework to predict tissue-specific regulatory logic",
"container-title": "NAR genomics and bioinformatics",
"author": [
{
"family": "Surana",
"given": "Pallavi"
},
{
"family": "Dutta",
"given": "Pratik"
},
{
"family": "Papineni",
"given": "Nimisha"
},
{
"family": "Sathian",
"given": "Rekha"
},
{
"family": "Zhou",
"given": "Zhihan"
},
{
"family": "Liu",
"given": "Han"
},
{
"family": "Davuluri",
"given": "Ramana V"
}
],
"container-title-short": "NAR Genom Bioinform",
"volume": "8",
"issue": "2",
"page": "lqag050",
"DOI": "10.1093/nargab/lqag050",
"PMID": "42244857",
"PMCID": "PMC13233143",
"ISSN": "2631-9268",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/nargab/lqag050",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
3
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Biopython, SHAP, caret, 12 other tools
[2] doi:10.1038/s41592-026-03211-w [code]
Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.
Journal: Nature methods
In common: Biopython, BEDTools, caret, 10 other tools
[3] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: Biopython, BEDTools, statsmodels, 7 other tools, 2 references
[4] doi:10.1038/s42003-026-10462-y [code]
SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.
Journal: Communications biology
In common: Biopython, SHAP, BEDTools, 7 other tools, 1 reference
[5] doi:10.1038/s41467-026-76837-1 [code]
Drug screen and machine learning predict neuroprotective agents in a preclinical human model of childhood dementia.
Journal: Nature communications
In common: SHAP, Hugging Face Transformers, data.table, 9 other tools
[6] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: Biopython, SHAP, BEDTools, 7 other tools
[7] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: Biopython, BEDTools, data.table, 8 other tools
[8] doi:10.1038/s44318-026-00818-9 [code]
FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.
Journal: The EMBO journal
In common: Biopython, BEDTools, data.table, 8 other tools
[9] doi:10.1093/bib/bbag339 [code]
scDeepAPA: a deep learning framework for single-cell alternative polyadenylation identification.
Journal: Briefings in bioinformatics
In common: Biopython, BEDTools, Hugging Face Transformers, 7 other tools
[10] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: BEDTools, data.table, ggplot2, 7 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.