TSProm: deep learning framework to predict tissue-specific regulatory logic.
The 16 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
- [1] § Results › Deciphering brain regulatory language and its’ disease relevance ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 18–31 · score 0.83 · C2H2 zinc finger, basic helix loop, bHLH, ZNFs, motifs
- [2] § Materials and methods › Fine tuning of DFMs ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 14–54 · score 0.82 · gena lm bigbird, AIRI Institute, fine tuned, t2t, classify, tokenizes
- [3] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 18–31 · score 0.74 · basic helix loop, zinc finger, bHLH, ZnF, Overlap, motifs
- [4] § Results › TSp promoter prediction models ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 142–219 · score 0.66 · weight decay, F1 scores, fine tuning, epochs, accuracy, batch
- [5] § Materials and methods › Attention-based motif discovery › Motif positional analysis ↔ TSProm.zip/src/3_attention/2A_save_meme.py, lines 394–447 · score 0.65 · gap free, attention score, heuristic, histograms, positions, tokens
- [6] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/notebooks/runs/3_TransSHAP.ipynb, lines 94–121 · score 0.65 · KernelExplainer, mask token, SHAP, logits, background, tokenized
- [7] § Materials and methods › Attention-based motif discovery › Validation of motifs: SHAP analysis ↔ TSProm.zip/src/3_attention/3_SHAP.py, lines 85–111 · score 0.65 · KernelExplainer, mask token, SHAP, logits, background, tokenized
- [8] § DNA foundation models ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 14–54 · score 0.64 · Gena LM, fine tuned, sequence length, batch, tokenization, training
- [9] § Materials and methods › Model evaluation metrics ↔ TSProm.zip/src/1_finetune/FineTune_GENALM.py, lines 142–219 · score 0.63 · F1 Score, fine tuned, correlation, Recall, Precision, metrics
- [10] § Materials and methods › Fine tuning of DFMs ↔ TSProm.zip/run_scripts/1_dnabert2_finetune.sh, lines 3–51 · score 0.62 · 2–117, fine tuned, zhihan1996, activated, DNABERT2, tokenizes
- [11] § Results › TSp promoter prediction models ↔ TSProm.zip/run_scripts/0_generate_data_run.sh, the whole file · a weak match · score 0.61 · MouseTranstex, TSProm, trained, tissues, promoter, model
- [12] § Results › TSp promoter prediction models ↔ TSProm.zip/run_scripts/1_dnabert2_finetune.sh, lines 3–51 · score 0.58 · weight decay, Hyperparameter, fine tuning, epochs, batch, DNABERT2
- [13] § Datasets ↔ TSProm.zip/run_scripts/0_generate_data_run.sh, the whole file · a weak match · score 0.58 · Mouse TransTEx, fine tuning, liver, testis, species, brain
- [14] § Results › Language of the global tissue specificity ↔ TSProm.zip/notebooks/Fig4.ipynb, lines 39–157 · score 0.57 · C2H2 ZNFs, bHLH, hits, global, TF, gene
- [15] § Datasets › Varying sequence lengths tested ↔ TSProm.zip/src/0_generate_data/make_modelA_negSet.sh, lines 1–52 · score 0.53 · promoter window, configurations, mm39, hg38, genomes, Mouse
- [16] § Materials and methods › Model evaluation metrics ↔ TSProm.zip/src/1_finetune/DNABERT2_AttentionExtracted.py, lines 232–250 · score 0.50 · F1 Score, correlation, Recall, Precision, metrics, Accuracy
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 286 lines · 11 KB · CC-BY-4.0 · 4 matches
- # Usage instructions:
- # conda activate dnabert2.0
- # # cd /data/private/psurana/ramanaServer_code/scripts/OtherModel_FineTune/final
- # # CUDA_VISIBLE_DEVICES=0,1,2 TOKENIZERS_PARALLELISM=false PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True nohup python aug_2025_genalm_1.py > finetune_all.log 2>&1 &
- # cd /data/private/psurana/TSProm/src/1_finetune/
- # CUDA_VISIBLE_DEVICES=5,6,7 TOKENIZERS_PARALLELISM=false PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python FineTune_GENALM.py
- # export NEPTUNE_PROJECT="lavi.ana/GenaLM-human-mouse-Finetune"
- # export TOKENIZERS_PARALLELISM=false
- def run_genalm_finetune(data_base_path, output_base_path, tissue_folder, length,
- learning_rate=2e-5, batch_size=12, max_length=2000, num_train_epochs=10):
- """
- Fine-tune GENA-LM on tissue-specific promoter data.
- Parameters:
- - data_base_path (str): Root directory containing all data folders
- - output_base_path (str): Where to save model and evaluation results
- - tissue_folder (str): e.g., "tsp_brain_low"
- - length (str): e.g., "2000", "3000", "4000"
- - learning_rate (float): Training learning rate
- - batch_size (int): Batch size per device
- - max_length (int): Max token length for tokenizer
- - num_train_epochs (int): Number of training epochs
- """
- import torch
- import os
- import pandas as pd
- import json
- from transformers import AutoTokenizer, AutoModel, TrainingArguments, Trainer
- from transformers import DataCollatorWithPadding
- from datasets import Dataset, DatasetDict
- import numpy as np
- from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, matthews_corrcoef
- import importlib
- # Debug print
- print(f"🔧 Running with: LR={learning_rate}, Batch={batch_size}, Epochs={num_train_epochs}, Max_len={max_length}")
- print(f"📁 Tissue: {tissue_folder} | Length: {length}")
- print(f"🔢 CUDA devices available: {torch.cuda.device_count()}")
- # Model selection based on sequence length
- if length == "2000":
- tokenizer = AutoTokenizer.from_pretrained('AIRI-Institute/gena-lm-bert-large-t2t')
- model_name = 'AIRI-Institute/gena-lm-bert-large-t2t'
- model_class_name = "BertForSequenceClassification"
- else: # For 3000 and 4000
- tokenizer = AutoTokenizer.from_pretrained('AIRI-Institute/gena-lm-bigbird-base-t2t')
- model_name = 'AIRI-Institute/gena-lm-bigbird-base-t2t'
- model_class_name = "BigBirdForSequenceClassification"
- # Fix padding token if needed
- if tokenizer.pad_token is None:
- tokenizer.pad_token = tokenizer.eos_token
- # Load base model to get module info
- model = AutoModel.from_pretrained(model_name, trust_remote_code=True)
- gena_module_name = model.__class__.__module__
- print(f"📦 Using module: {gena_module_name}")
- # Get classification model class
- cls = getattr(importlib.import_module(gena_module_name), model_class_name)
- # Create classification model
- model = cls.from_pretrained(model_name, num_labels=2)
- print(f'🧠 Classification head: {model.classifier}')
- # Setup paths
- trial_data = os.path.join(data_base_path, length, tissue_folder)
- output_finish_path = os.path.join(output_base_path, tissue_folder, length, f"lr_{learning_rate:.0e}")
- # Create output directory
- os.makedirs(output_finish_path, exist_ok=True)
- print(f"📂 Output directory: {output_finish_path}")
- # Load data with debugging
- try:
- dev = pd.read_csv(os.path.join(trial_data, 'dev.csv'))
- train = pd.read_csv(os.path.join(trial_data, 'train.csv'))
- test = pd.read_csv(os.path.join(trial_data, 'test.csv'))
- # Debug: Check column names and data types
- print(f"📊 Train columns: {list(train.columns)}")
- print(f"🔍 Sample train data:\n{train.head(2)}")
- print(f"🔍 Sequence column type: {type(train['Sequence'].iloc[0])}")
- print(f"📊 Data loaded - Train: {len(train)}, Dev: {len(dev)}, Test: {len(test)}")
- except Exception as e:
- print(f"❌ Error loading data: {e}")
- raise
- # Clean and prepare data - remove pandas index column that might cause issues
- train = train.reset_index(drop=True)
- dev = dev.reset_index(drop=True)
- test = test.reset_index(drop=True)
- # Convert to the format expected by the model
- def prepare_data(df):
- return {
- 'sequence': df['Sequence'].astype(str).tolist(),
- 'label': df['Label'].astype(int).tolist()
- }
- train_data = prepare_data(train)
- dev_data = prepare_data(dev)
- test_data = prepare_data(test)
- # Create datasets from dictionaries
- train_dataset = Dataset.from_dict(train_data)
- dev_dataset = Dataset.from_dict(dev_data)
- test_dataset = Dataset.from_dict(test_data)
- dataset = DatasetDict({
- "train": train_dataset,
- "validation": dev_dataset,
- "test": test_dataset,
- })
- # Debug: Check first example
- print("🔍 First train example:", dataset["train"][0])
- # Tokenization - simpler approach
- def preprocess_function(examples):
- # Tokenize the sequences
- tokenized = tokenizer(
- examples["sequence"],
- truncation=True,
- max_length=max_length,
- padding=False, # Let DataCollator handle padding
- return_tensors=None
- )
- # Keep the labels
- tokenized["labels"] = examples["label"] # Use "labels" not "label"
- return tokenized
- tokenized_dataset = dataset.map(preprocess_function, batched=True)
- # Remove original columns to avoid conflicts
- tokenized_dataset = tokenized_dataset.remove_columns(["sequence", "label"])
- # Debug: Check tokenized data
- print("🔍 Tokenized sample:", {k: (type(v), len(v) if isinstance(v, list) else v) for k, v in tokenized_dataset["train"][0].items()})
- # Metrics computation
- def compute_metrics(eval_pred):
- predictions, labels = eval_pred
- predictions = np.argmax(predictions, axis=1)
- return {
- "accuracy": accuracy_score(labels, predictions),
- "precision": precision_score(labels, predictions, zero_division=0),
- "recall": recall_score(labels, predictions, zero_division=0),
- "f1": f1_score(labels, predictions, zero_division=0),
- "matthews_correlation": matthews_corrcoef(labels, predictions)
- }
- # Create data collator
- data_collator = DataCollatorWithPadding(tokenizer=tokenizer, return_tensors="pt")
- # Training arguments
- training_args = TrainingArguments(
- output_dir=output_finish_path,
- learning_rate=learning_rate,
- lr_scheduler_type="constant_with_warmup",
- warmup_ratio=0.1,
- optim='adamw_torch',
- weight_decay=0.0,
- per_device_train_batch_size=batch_size,
- per_device_eval_batch_size=batch_size,
- num_train_epochs=num_train_epochs,
- evaluation_strategy="epoch",
- save_strategy="epoch",
- logging_strategy="epoch",
- load_best_model_at_end=True,
- dataloader_pin_memory=False,
- remove_unused_columns=False,
- report_to=None,
- # Added for better multi-GPU support
- dataloader_num_workers=4,
- fp16=True, # Enable mixed precision for better memory usage
- )
- # Create trainer
- trainer = Trainer(
- model=model,
- args=training_args,
- train_dataset=tokenized_dataset["train"],
- eval_dataset=tokenized_dataset["validation"],
- tokenizer=tokenizer,
- data_collator=data_collator,
- compute_metrics=compute_metrics,
- )
- print("🚀 Starting training...")
- trainer.train()
- # Evaluate on test set
- print("📈 Evaluating on test set...")
- results = trainer.evaluate(eval_dataset=tokenized_dataset["test"])
- print("📊 Test Results:")
- for key, value in results.items():
- if isinstance(value, float):
- print(f" {key}: {value:.4f}")
- else:
- print(f" {key}: {value}")
- # Save results
- output_json_path = os.path.join(output_finish_path, "eval_results.json")
- with open(output_json_path, "w") as f:
- json.dump(results, f, indent=4)
- print(f"💾 Results saved to: {output_json_path}")
- # Cleanup memory
- del model, trainer
- torch.cuda.empty_cache()
- print("Memory cleaned up")
- # ============================================================================
- # Main execution loop with error handling
- # ============================================================================
- # tissue_dirs = [
- # "tsp_brain_low", "tsp_brain_tenh", "tsp_liver_low", "tsp_liver_tenh",
- # "tsp_testis_low", "tsp_brain_null", "tsp_brain_wide",
- # "tsp_liver_null", "tsp_liver_wide", "tsp_testis_null", "tsp_testis_wide", "tsp_testis_tenh"
- # ]
- # tissue_dirs =["tsp_spleen_null", "tsp_spleen_wide", "tsp_spleen_tenh", "tsp_spleen_low",
- # "tsp_muscle_null", "tsp_muscle_wide"]
- tissue_dirs = ["tsp_muscle_tenh", "tsp_muscle_low"]
- lengths = ["3000", "4000", "2000"]
- length_map = {"2000": 2000, "3000": 3000, "4000": 4000}
- learning_rates = [2e-5, 3e-4, 5e-6]
- import os
- print("🎯 Starting GENA-LM fine-tuning experiments...")
- print(f"📋 Total experiments: {len(tissue_dirs)} × {len(lengths)} × {len(learning_rates)} = {len(tissue_dirs) * len(lengths) * len(learning_rates)}")
- experiment_count = 0
- total_experiments = len(tissue_dirs) * len(lengths) * len(learning_rates)
- for tissue in tissue_dirs:
- for length in lengths:
- for lr in learning_rates:
- experiment_count += 1
- batch_size = 8 if "testis" in tissue else 12
- lr_tag = f"lr_{lr:.0e}"
- output_dir = os.path.join(
- "/data/projects/dna/pallavi/DNABERT_runs/DATA_RUN/gena-lm_finetune/unique_TSS/human_mouse",
- tissue, length, lr_tag
- )
- print(f"\n{'='*80}")
- print(f"🧪 Experiment {experiment_count}/{total_experiments}")
- print(f"🧬 Tissue: {tissue} | Length: {length} | LR: {lr} | Batch: {batch_size}")
- print(f"{'='*80}")
- if os.path.exists(os.path.join(output_dir, "eval_results.json")):
- print(f"✅ Skipping: {tissue} | Length: {length} | LR: {lr} (Already completed)")
- continue
- try:
- run_genalm_finetune(
- data_base_path="/data/projects/dna/pallavi/data_TSp_Vs_Rest",
- output_base_path="/data/projects/dna/pallavi/DNABERT_runs/DATA_RUN/gena-lm_finetune/unique_TSS/human_mouse",
- tissue_folder=tissue,
- length=length,
- learning_rate=lr,
- batch_size=batch_size,
- max_length=length_map[length],
- num_train_epochs=10
- )
- print(f"✅ Completed: {tissue} | Length: {length} | LR: {lr}")
- except Exception as e:
- print(f"❌ Error for {tissue} | Length: {length} | LR: {lr}: {str(e)}")
- import traceback
- print(f"📝 Full traceback:\n{traceback.format_exc()}")
- continue
- print(f"\n🎉 All experiments completed! Check logs for any errors.")
FineTune_GENALM.py, under CC-BY-4.0 · at the source
Overview
- Department of Biomedical Informatics, Stony Brook University, NY 11794, United States
- Department of Computer Science, Northwestern University, IL 60208, United States
Abstract
Characterizing tissue-specific (TSp) gene expression is crucial for understanding development and disease; however, traditional expression-based methods often overlook the latent “regulatory grammar” embedded in non-coding DNA, particularly across distal promoter regions. Here, we introduce TSProm, a framework that adapts a DNA foundation model (DNABERT2) to decode the regulatory logic of TSp promoters at the isoform level. Our contributions are two-fold: (i) a comparative design that trains two specialized models: Model A for general promoter biology and Model B for tissue-specific regulation enabling precise isolation of sequence motifs surrounding the transcription start site that uniquely define tissue identity; (ii) we develop an explainable AI module that integrates attention-based motif discovery with model-agnostic SHAP analysis to yield cross-validated interpretations of learned features. Applying TSProm to human brain, liver, and testis promoters, we identified clinically relevant transcription factors (TFs) in the brain, including SP1, MYC, and HES6, whose associations with gliomas and neuroblastomas highlight clinical relevance. Moreover, our results highlight C2H2 zinc finger proteins as a dominant family shaping the global landscape of TSp gene regulation. TSProm provides an interpretable and generalizable framework for identifying tissue-specific regulatory elements, offering powerful computational tools to investigate gene regulation in both normal and disease contexts.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 16 matches between paragraphs and lines of code.
figshare 30747257
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
44 files
- TSProm.zip/
install_R_requirements.R — R, 28 lines - TSProm.zip/
notebooks/ — Jupyter, 381 lines.ipynb_checkpoints/ Fig2-checkpoint.ipynb - TSProm.zip/
notebooks/ — Jupyter, 69 lines.ipynb_checkpoints/ Fig4-checkpoint.ipynb - TSProm.zip/
notebooks/ — Jupyter, 174 lines.ipynb_checkpoints/ Fig5-checkpoint.ipynb - TSProm.zip/
notebooks/ — Jupyter, 381 linesFig2.ipynb - TSProm.zip/
notebooks/ — Jupyter, 69 linesFig3.ipynb - TSProm.zip/
notebooks/ — Jupyter, 174 lines, 3 matchesFig4.ipynb - TSProm.zip/
notebooks/ — Jupyter, 660 linesruns/ 1_Attention_motif_TEST.i pynb - TSProm.zip/
notebooks/ — Jupyter, 901 linesruns/ 2A_TEST.ipynb - TSProm.zip/
notebooks/ — Jupyter, 541 linesruns/ 2C_Biclustering.ipynb - TSProm.zip/
notebooks/ — Jupyter, 363 lines, 1 matchruns/ 3_TransSHAP.ipynb - TSProm.zip/
run_scripts/ — Shell, 42 lines, 2 matches0_generate_data_run.sh - TSProm.zip/
run_scripts/ — Shell, 37 lines1_GENA-LM_finetune.sh - TSProm.zip/
run_scripts/ — Shell, 130 lines, 2 matches1_dnabert2_finetune.sh - TSProm.zip/
run_scripts/ — Shell, 24 lines2_run_predict.sh - TSProm.zip/
run_scripts/ — Shell, 92 lines3_run_attention.sh - TSProm.zip/
src/ — R, 106 lines0_generate_data/ .ipynb_checkpoints/ 1_make_fine_tune-checkpo int.R - TSProm.zip/
src/ — R, 681 lines0_generate_data/ .ipynb_checkpoints/ 2A_Final_make_data_diff_ length-checkpoint.R - TSProm.zip/
src/ — R, 474 lines0_generate_data/ .ipynb_checkpoints/ 2_Final_bed_fa_dnabert2- checkpoint.R - TSProm.zip/
src/ — R, 450 lines0_generate_data/ .ipynb_checkpoints/ 3B_cpg_ncpg-checkpoint.R - TSProm.zip/
src/ — R, 371 lines0_generate_data/ .ipynb_checkpoints/ 3_CPG_Non_CPG-checkpoint .R - TSProm.zip/
src/ — R, 353 lines0_generate_data/ make_data.R - TSProm.zip/
src/ — Shell, 132 lines, 1 match0_generate_data/ make_modelA_negSet.sh - TSProm.zip/
src/ — R, 230 lines0_generate_data/ make_modelB_nullSeqs.R - TSProm.zip/
src/ — Python, 287 lines1_finetune/ .ipynb_checkpoints/ FineTune_GENALM-checkpoi nt.py - TSProm.zip/
src/ — Python, 458 lines, 1 match1_finetune/ DNABERT2_AttentionExtrac ted.py - TSProm.zip/
src/ — Python, 286 lines, 4 matches1_finetune/ FineTune_GENALM.py - TSProm.zip/
src/ — Python, 30 lines2_predict/ .ipynb_checkpoints/ 1_runAll_predict-checkpo int.py - TSProm.zip/
src/ — Jupyter, 170 lines2_predict/ .ipynb_checkpoints/ 2_merge_data-checkpoint. ipynb - TSProm.zip/
src/ — Python, 171 lines2_predict/ 1_predict.py - TSProm.zip/
src/ — Python, 30 lines2_predict/ 1_runAll_predict.py - TSProm.zip/
src/ — Jupyter, 660 lines3_attention/ .ipynb_checkpoints/ 1_Attention_motif_TEST-c heckpoint.ipynb - TSProm.zip/
src/ — Python, 485 lines3_attention/ .ipynb_checkpoints/ 1_raw_attention_extract- gpu-checkpoint.py - TSProm.zip/
src/ — Python, 484 lines3_attention/ .ipynb_checkpoints/ 1_raw_attention_extract- gpu-oneJob-checkpoint.py - TSProm.zip/
src/ — Jupyter, 898 lines3_attention/ .ipynb_checkpoints/ 2A_TEST-checkpoint.ipynb - TSProm.zip/
src/ — Python, 534 lines3_attention/ .ipynb_checkpoints/ 2A_save_meme-checkpoint. py - TSProm.zip/
src/ — Shell, 45 lines3_attention/ .ipynb_checkpoints/ 2B_meme-checkpoint.sh - TSProm.zip/
src/ — Shell, 56 lines3_attention/ .ipynb_checkpoints/ 2C_meme_testis_subsample -checkpoint.sh - TSProm.zip/
src/ — Python, 148 lines3_attention/ .ipynb_checkpoints/ 3_SHAP-checkpoint.py - TSProm.zip/
src/ — Jupyter, 360 lines3_attention/ .ipynb_checkpoints/ 3_TransSHAP-checkpoint.i pynb - TSProm.zip/
src/ — Python, 485 lines3_attention/ 1_raw_attention_extract- gpu.py - TSProm.zip/
src/ — Python, 534 lines, 1 match3_attention/ 2A_save_meme.py - TSProm.zip/
src/ — Python, 148 lines, 1 match3_attention/ 3_SHAP.py - TSProm.zip/
README.md — Text, 156 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 43 scripts, each with its path and the digest of its content;
- 16 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data availability
TSProm code and Model weights: https://
Isoform expression data: GTEx (see GTEx Consortium 2015, 2020).
Human TransTEx groupings and code: Described in Surana et al. (2024) and available at TransTExdb.
Mouse
BodyMap dataset: Available from NCBI BioProject PRJNA375882 (Li et al., 2017).
Hugging Face DNA language models (Fishman et al., 2025; Zhou et al., 2021; Dalla Torre et al., 2025):
AIRI-Institute/
AIRI-Institute/
InstaDeepAI/
jaandoui/
zhihan1996/
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 9 MeSH terms, 2 funders, 50 references.
Cite
This paper
Surana, P., Dutta, P., Papineni, N., Sathian, R., Zhou, Z., Liu, H., & Davuluri, R. V. (2026). TSProm: deep learning framework to predict tissue-specific regulatory logic. NAR genomics and bioinformatics, 8(2), lqag050. https://
BibTeX
@article{surana2026tspro
author = {Surana, Pallavi and Dutta, Pratik and Papineni, Nimisha and Sathian, Rekha and Zhou, Zhihan and Liu, Han and Davuluri, Ramana V},
title = {{TSProm: deep learning framework to predict tissue-specific regulatory logic}},
journal = {NAR genomics and bioinformatics},
year = {2026},
month = jun,
volume = {8},
number = {2},
pages = {lqag050},
publisher = {Oxford University Press},
issn = {2631-9268},
doi = {10.1093/
url = {https://
pmid = {42244857},
pmcid = {PMC13233143}
}
RIS
TY - JOUR
AU - Surana, Pallavi
AU - Dutta, Pratik
AU - Papineni, Nimisha
AU - Sathian, Rekha
AU - Zhou, Zhihan
AU - Liu, Han
AU - Davuluri, Ramana V
TI - TSProm: deep learning framework to predict tissue-specific regulatory logic
T2 - NAR genomics and bioinformatics
J2 - NAR Genom Bioinform
PY - 2026
DA - 2026/
VL - 8
IS - 2
SP - lqag050
SN - 2631-9268
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "TSProm: deep learning framework to predict tissue-specific regulatory logic",
"container-title": "NAR genomics and bioinformatics",
"author": [
{
"family": "Surana",
"given": "Pallavi"
},
{
"family": "Dutta",
"given": "Pratik"
},
{
"family": "Papineni",
"given": "Nimisha"
},
{
"family": "Sathian",
"given": "Rekha"
},
{
"family": "Zhou",
"given": "Zhihan"
},
{
"family": "Liu",
"given": "Han"
},
{
"family": "Davuluri",
"given": "Ramana V"
}
],
"container-title-short":
"volume": "8",
"issue": "2",
"page": "lqag050",
"DOI": "10.1093/
"PMID": "42244857",
"PMCID": "PMC13233143",
"ISSN": "2631-9268",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
3
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: Biopython, SHAP, caret, 12 other tools
- [2] doi:10.1038/s41592-026-03211-w [code]
- Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.Journal: Nature methodsIn common: Biopython, BEDTools, caret, 10 other tools
- [3] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: Biopython, BEDTools, statsmodels, 7 other tools, 2 references
- [4] doi:10.1038/s42003-026-10462-y [code]
- SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.Journal: Communications biologyIn common: Biopython, SHAP, BEDTools, 7 other tools, 1 reference
- [5] doi:10.1038/s41467-026-76837-1 [code]
- Drug screen and machine learning predict neuroprotective agents in a preclinical human model of childhood dementia.Journal: Nature communicationsIn common: SHAP, Hugging Face Transformers, data.table, 9 other tools
- [6] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: Biopython, SHAP, BEDTools, 7 other tools
- [7] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: Biopython, BEDTools, data.table, 8 other tools
- [8] doi:10.1038/s44318-026-00818-9 [code]
- FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.Journal: The EMBO journalIn common: Biopython, BEDTools, data.table, 8 other tools
- [9] doi:10.1093/bib/bbag339 [code]
- scDeepAPA: a deep learning framework for single-cell alternative polyadenylation identification.Journal: Briefings in bioinformaticsIn common: Biopython, BEDTools, Hugging Face Transformers, 7 other tools
- [10] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: BEDTools, data.table, ggplot2, 7 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 43 scripts, and 16 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:b159c5cc4237f21c…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
