OSCR

INB<sup>3</sup>P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Materials and Methods › Data Scarcity and Augmentation in Peptide Prediction ↔ BBBPPINFP/constants.py, the whole file · a weak match · score 0.83 · molar refractivity, interfacial hydrophobicity, molecular weight, TPSA, hydropathy, SASA
  2. [2] § Results and Discussion › Conformity of the Augmented Data to Biochemical Regularities ↔ BBBPPINFP/constants.py, the whole file · a weak match · score 0.82 · Wimley White interfacial, isoelectric point, molecular weight, TPSA, hydropathy, SASA
  3. [3] § Results and Discussion › Contribution of Model Components and Loss Functions › Loss‐Function Ablation and Contrastive Objectives ↔ model/threshold_model/model.py, lines 65–118 · score 0.63 · supervised contrastive losses, loss weight, Stable MCC, InfoNCE, SCL, threshold
  4. [4] § Materials and Methods › Loss Functions › Stage 1: Representation Learning (Alignment First, Supervision Later) ↔ model/original_model/model.py, lines 92–204 · score 0.59 · curriculum weights, PDB graph, con, supervision, trains, sequence
  5. [5] § Materials and Methods › Stratified Mini‐Batch Sampling Under Class Imbalance ↔ model/threshold_model/model.py, lines 167–261 · score 0.58 · STRATIFIED BATCH SAMPLER, Algorithm, seeds, ratio, focal, augmented
  6. [6] § Materials and Methods › Stratified Mini‐Batch Sampling Under Class Imbalance ↔ BBBPPINFP/__init__.py, lines 56–99 · score 0.57 · STRATIFIED BATCH SAMPLER, focal loss, gradients, seeds, MCC
  7. [7] § Results and Discussion › Evaluation of the Proposed Model and Its Predictive Performance Against Prior Baselines ↔ model/threshold_model/model.py, lines 167–261 · score 0.56 · DataLoader, scheduler, freezing, split, AP, CPU
  8. [8] § Materials and Methods › Loss Functions ↔ model/original_model/model.py, lines 92–204 · score 0.53 · supervised contrastive, wn, wp, logit, InfoNCE, smoothing
  9. [9] § Materials and Methods › Loss Functions ↔ model/threshold_model/model.py, lines 65–118 · score 0.53 · supervised contrastive, wn, wp, logit, InfoNCE, smoothing
  10. [10] § Materials and Methods › Evaluation Metrics and Threshold Selection › Threshold‐Free Curves and Areas ↔ BBBPPINFP/evaluation.py, lines 14–33 · score 0.52 · ROC AUC, ACC, Precision, AP, Sn, Sp
  11. [11] § Materials and Methods › Evaluation Metrics and Threshold Selection › Threshold‐Free Curves and Areas ↔ BBBPPINFP/evaluation.py, lines 36–64 · score 0.52 · ROC curve, FPR, TPR
  12. [12] § Results and Discussion › Interpretability Study and Biological Insights ↔ BBBPPINFP/__init__.py, lines 56–99 · score 0.52 · contact maps, supervised contrastive, heatmap, enrichment, alignment, properties
  13. [13] § Materials and Methods › Property Embedding and Normalization › Mutation Policy ↔ BBBPPINFP/augmentation.py, lines 100–188 · score 0.51 · fused substitution, identity, optional, BLOSUM, temperature, max
  14. [14] § Materials and Methods › Loss Functions › Stage 1: Representation Learning (Alignment First, Supervision Later) ↔ BBBPPINFP/bp_infp.py, lines 228–255 · score 0.50 · fusion blocks, PDB graph, linear, encoder

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 264 lines · 18 KB · no license · 4 matches

  1. #!/usr/bin/env python3
  2. # -*- coding: utf-8 -*-
  3. import os
  4. os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":16:8"
  5. import warnings
  6. warnings.filterwarnings("ignore", message=".*scatter_reduce_cuda does not have a deterministic implementation.*")
  7. import random
  8. import argparse
  9. import numpy as np
  10. import pandas as pd
  11. import torch
  12. import torch.nn.functional as F
  13. from torch.utils.data import DataLoader
  14. from sklearn.model_selection import StratifiedShuffleSplit
  15. from tqdm import tqdm
  16. import esm
  17. import copy
  18. import io
  19. import sys
  20. import logging
  21. import time
  22. import hashlib
  23. import torch_geometric
  24. here = os.path.abspath(os.path.dirname(__file__))
  25. probe = here
  26. while True:
  27. if os.path.exists(os.path.join(probe, 'BBBPPINFP', '__init__.py')):
  28. if probe not in sys.path: sys.path.insert(0, probe)
  29. break
  30. parent = os.path.dirname(probe)
  31. if parent == probe: raise RuntimeError("Could not locate project root containing BBBPPINFP/")
  32. probe = parent
  33. project_root = os.path.abspath(os.path.join(os.path.dirname(__file__), '..', '..'))
  34. if project_root not in sys.path: sys.path.insert(0, project_root)
  35. from BBBPPINFP import (
  36. BP_INFP, SubstitutionMatrix, generate_biologically_plausible_mutation,
  37. SeqPDBDataset, collate_seq_pdb, focal_loss, stable_mcc_loss,
  38. info_nce_loss, supervised_contrastive_loss, StratifiedBatchSampler,
  39. compute_metrics, plot_metrics_curves, plot_score_kde_seaborn,
  40. dump_hard_cases, seed_worker, infer_pdb_cached, PDB_NODE_FEATURE_DIM
  41. )
  42. ESMFOLD_MODEL = None
  43. ESMFOLD_CACHE_DIR = None
  44. def _make_rng(seed: int, epoch: int, sample_id: str):
  45. h = hashlib.blake2b(digest_size=8)
  46. h.update(str(seed).encode()); h.update(str(epoch).encode()); h.update(str(sample_id).encode())
  47. s = int.from_bytes(h.digest(), 'little') % (2 ** 32)
  48. return np.random.default_rng(s), random.Random(s)
  49. def compute_class_weights(lbls_all):
  50. arr = np.array(lbls_all)
  51. n_pos, n_neg = np.sum(arr == 1), np.sum(arr == 0)
  52. total = len(arr)
  53. wp = min(20.0, total / (2.0 * n_pos) if n_pos > 0 else 1.0)
  54. wn = min(20.0, total / (2.0 * n_neg) if n_neg > 0 else 1.0)
  55. return wp, wn
  56. def train_one_epoch(model, ldr, optim, dev, wp, wn, args, substitution_matrix, current_stage=1, current_epoch=0, total_epochs_in_stage=1):
  57. model.train()
  58. total_loss_val, all_lbls, all_probs = 0.0, [], []
  59. use_contrastive_loss = (current_stage == 1)
  60. use_supervised_loss = (current_stage == 2) or (current_stage == 1 and args.stage1_cls_loss_weight > 0)
  61. pdb_parser_instance = ldr.dataset
  62. mutated_count_epoch, folded_missing_count_epoch = 0, 0
  63. w_contrastive, w_supervised = 1.0, 1.0
  64. if current_stage == 1 and total_epochs_in_stage > 1:
  65. progress = current_epoch / (total_epochs_in_stage - 1)
  66. w_contrastive, w_supervised = 1.0 - progress, progress
  67. is_augmentation_active = (current_stage == 1 and args.aug_prob > 0) or (current_stage == 2 and args.enable_stage2_augmentation)
  68. pbar = tqdm(ldr, desc=f"S{current_stage} Train", leave=False, ncols=120)
  69. for grph_pdb_orig, seq_esm_orig, lbl_cpu, is_pos_tensor in pbar:
  70. if (not hasattr(grph_pdb_orig, 'x')) or grph_pdb_orig.num_graphs == 0 or len(seq_esm_orig) == 0 or lbl_cpu.numel() == 0: continue
  71. batch_size = grph_pdb_orig.num_graphs
  72. final_graphs_list, final_seqs_list = [], []
  73. original_graph_list = grph_pdb_orig.to_data_list()
  74. for i in range(batch_size):
  75. current_graph, current_seq_tuple = original_graph_list[i], seq_esm_orig[i]
  76. np_rng, py_rng = _make_rng(args.seed, current_epoch, current_seq_tuple[0])
  77. a_p, m_p, m_f = (args.stage2_aug_prob, args.stage2_mutation_prob, args.stage2_max_mutation_fraction) if (current_stage == 2 and args.enable_stage2_augmentation) else (args.aug_prob, args.mutation_prob, args.max_mutation_fraction)
  78. if args.aug_warmup_epochs > 0: a_p *= min(1.0, (current_epoch + 1) / float(args.aug_warmup_epochs))
  79. if is_augmentation_active and is_pos_tensor[i] and (py_rng.random() < a_p):
  80. mutant = generate_biologically_plausible_mutation(current_seq_tuple[1], substitution_matrix, m_p, m_f, np_rng, py_rng, current_epoch, f"S{current_stage}", str(current_seq_tuple[0]))
  81. if ESMFOLD_MODEL is not None and mutant != current_seq_tuple[1]:
  82. pdb_txt = infer_pdb_cached(mutant, ESMFOLD_MODEL, ESMFOLD_CACHE_DIR)
  83. if pdb_txt:
  84. folded, _ = pdb_parser_instance._pdb_to_graph_and_seq(io.StringIO(pdb_txt))
  85. if folded and not folded.is_placeholder:
  86. current_graph, current_seq_tuple = copy.deepcopy(folded), (current_seq_tuple[0], mutant)
  87. mutated_count_epoch += 1
  88. if current_graph.is_placeholder and ESMFOLD_MODEL is not None:
  89. pdb_txt = infer_pdb_cached(current_seq_tuple[1], ESMFOLD_MODEL, ESMFOLD_CACHE_DIR)
  90. if pdb_txt:
  91. folded, _ = pdb_parser_instance._pdb_to_graph_and_seq(io.StringIO(pdb_txt))
  92. if folded and not folded.is_placeholder: current_graph = copy.deepcopy(folded); folded_missing_count_epoch += 1
  93. final_graphs_list.append(current_graph); final_seqs_list.append(current_seq_tuple)
  94. grph_pdb = torch_geometric.data.Batch.from_data_list(final_graphs_list).to(dev)
  95. optim.zero_grad(set_to_none=True)
  96. z_s, z_p, z_s_scl, z_p_scl, temp, logits = model(grph_pdb, final_seqs_list, mode="train")
  97. eff_bs, eff_lbls = z_s.size(0), lbl_cpu.to(dev)[:z_s.size(0)]
  98. loss = torch.tensor(0.0, device=dev)
  99. if use_contrastive_loss:
  100. loss += w_contrastive * (args.lambda_contrastive * info_nce_loss(z_s, z_p, temp, model.pdb_graph_bank, model.seq_bank, args.k_neg_infonce) + args.lambda_scl_seq * supervised_contrastive_loss(z_s_scl, eff_lbls, args.scl_temperature) + args.lambda_scl_pdb * supervised_contrastive_loss(z_p_scl, eff_lbls, args.scl_temperature))
  101. if use_supervised_loss and logits.numel() > 0:
  102. prbs = torch.sigmoid(logits)
  103. loss += (w_supervised * args.stage1_cls_loss_weight if current_stage == 1 else 1.0) * (args.w_focal * focal_loss(logits, eff_lbls, args.gamma_focal, wp, wn, args.label_smoothing) + args.w_mcc * stable_mcc_loss(prbs, eff_lbls, wp, wn, args.label_smoothing))
  104. all_probs.extend(prbs.detach().cpu().numpy().tolist())
  105. all_lbls.extend(eff_lbls.cpu().numpy().tolist())
  106. loss.backward(); optim.step(); total_loss_val += loss.item() * eff_bs
  107. return (total_loss_val / len(all_lbls) if all_lbls else 0), compute_metrics(all_lbls, all_probs)
  108. def validate(model, ldr, dev, ep, odir, mode_desc="Valid", threshold=None, args=None, plot=False):
  109. model.eval()
  110. total_loss, all_lbls, all_probs, all_ids = 0.0, [], [], []
  111. with torch.no_grad():
  112. for grph_pdb, seq_esm, lbl_cpu, _ in tqdm(ldr, desc=mode_desc, leave=False):
  113. if (not hasattr(grph_pdb, 'x')) or grph_pdb.num_graphs == 0 or len(seq_esm) == 0: continue
  114. grph_pdb = grph_pdb.to(dev)
  115. _, _, _, _, _, logits = model(grph_pdb, seq_esm, mode="classify")
  116. if logits is not None and logits.numel() > 0:
  117. eff_bs, eff_lbls = logits.size(0), lbl_cpu.to(dev)[:logits.size(0)].float()
  118. total_loss += F.binary_cross_entropy_with_logits(logits, eff_lbls).item() * eff_bs
  119. all_lbls.extend(eff_lbls.cpu().numpy().tolist())
  120. all_probs.extend(torch.sigmoid(logits).cpu().numpy().tolist())
  121. all_ids.extend([int(t[0]) for t in seq_esm][:logits.size(0)])
  122. best_thr = 0.5
  123. if all_lbls:
  124. if threshold is not None:
  125. best_thr = threshold
  126. else:
  127. thrs = np.linspace(0.05, 0.95, 91)
  128. candidate_metrics = [compute_metrics(all_lbls, all_probs, thr=t) for t in thrs]
  129. if args.monitor_metric == 'mcc':
  130. best_thr = thrs[np.argmax([m['mcc'] for m in candidate_metrics])]
  131. elif args.monitor_metric == 'f1':
  132. best_thr = thrs[np.argmax([m['f1'] for m in candidate_metrics])]
  133. elif args.monitor_metric == 'sn': # Target Sensitivity
  134. target = args.target_value
  135. diffs = [abs(m['sn'] - target) for m in candidate_metrics]
  136. best_thr = thrs[np.argmin(diffs)]
  137. elif args.monitor_metric == 'sp': # Target Specificity
  138. target = args.target_value
  139. diffs = [abs(m['sp'] - target) for m in candidate_metrics]
  140. best_thr = thrs[np.argmin(diffs)]
  141. elif args.monitor_metric == 'min_fnr': # Minimum FNR (Max SN)
  142. best_thr = thrs[np.argmax([m['sn'] for m in candidate_metrics])]
  143. else:
  144. best_thr = thrs[np.argmax([m['ap'] for m in candidate_metrics])]
  145. metrics = compute_metrics(all_lbls, all_probs, thr=best_thr)
  146. metrics['best_threshold'] = best_thr
  147. if plot and odir:
  148. plot_metrics_curves(all_lbls, all_probs, ep, odir, mode_desc)
  149. plot_score_kde_seaborn(all_lbls, all_probs, ep, odir, mode_desc, thr_lines=(best_thr,))
  150. return (total_loss / len(all_lbls) if all_lbls else 0), metrics
  151. def run():
  152. parser = argparse.ArgumentParser()
  153. parser.add_argument("--hid", type=int, default=1024); parser.add_argument("--batch", type=int, default=64); parser.add_argument("--seed", type=int, default=45)
  154. parser.add_argument("--train_csv", type=str, required=True); parser.add_argument("--test_csv", type=str, required=True); parser.add_argument("--output_dir", type=str, default="./outputs")
  155. parser.add_argument("--val_split_ratio", type=float, default=0.0); parser.add_argument("--deterministic", action="store_true")
  156. parser.add_argument("--monitor_metric", type=str, default="ap", choices=["ap", "mcc", "f1", "sn", "sp", "min_fnr"])
  157. parser.add_argument("--target_value", type=float, default=0.9, help="Used for sn/sp target mode")
  158. parser.add_argument("--stage1_epochs", type=int, default=35); parser.add_argument("--lr_stage1", type=float, default=2e-5)
  159. parser.add_argument("--stage1_cls_loss_weight", type=float, default=0.5); parser.add_argument("--freeze_ratio_esm", type=float, default=0.7)
  160. parser.add_argument("--stage2_epochs", type=int, default=15); parser.add_argument("--lr_stage2_head", type=float, default=2e-4); parser.add_argument("--lr_stage2_encoder", type=float, default=1e-6)
  161. parser.add_argument("--enable_esmfold", action="store_true"); parser.add_argument("--esmfold_cache_dir", type=str, default=None)
  162. parser.add_argument("--aug_prob", type=float, default=0.5); parser.add_argument("--mutation_prob", type=float, default=0.3); parser.add_argument("--max_mutation_fraction", type=float, default=0.2)
  163. parser.add_argument("--enable_stage2_augmentation", action='store_true'); parser.add_argument("--stage2_aug_prob", type=float, default=1.0)
  164. parser.add_argument("--stage2_mutation_prob", type=float, default=0.25); parser.add_argument("--stage2_max_mutation_fraction", type=float, default=0.25); parser.add_argument("--aug_warmup_epochs", type=int, default=0)
  165. parser.add_argument("--bank_size", type=int, default=2048); parser.add_argument("--k_neg_infonce", type=int, default=64); parser.add_argument("--lambda_contrastive", type=float, default=1.0)
  166. parser.add_argument("--gamma_focal", type=float, default=2.0); parser.add_argument("--w_focal", type=float, default=0.2); parser.add_argument("--w_mcc", type=float, default=0.8)
  167. parser.add_argument("--lambda_scl_seq", type=float, default=2.0); parser.add_argument("--lambda_scl_pdb", type=float, default=2.0); parser.add_argument("--normalize_esm_channels", action='store_true')
  168. parser.add_argument("--esm_channel_stats_path", type=str, default=None); parser.add_argument("--scl_embedding_dim", type=int, default=128); parser.add_argument("--scl_temperature", type=float, default=0.1)
  169. parser.add_argument("--w_blosum", type=float, default=0.5); parser.add_argument("--w_biochem", type=float, default=0.5); parser.add_argument("--softmax_temp", type=float, default=0.3)
  170. parser.add_argument("--label_smoothing", type=float, default=0.1); parser.add_argument("--num_fusion_layers", type=int, default=1); parser.add_argument("--esm_grad_ckpt", action="store_true")
  171. parser.add_argument("--stage2_unfreeze_ratio", type=float, default=None); parser.add_argument("--dropout", type=float, default=0.5); parser.add_argument("--stage2_dropout", type=float, default=0.5)
  172. args = parser.parse_args(); global ESMFOLD_MODEL, ESMFOLD_CACHE_DIR
  173. ESMFOLD_CACHE_DIR = args.esmfold_cache_dir; os.makedirs(args.output_dir, exist_ok=True)
  174. logging.basicConfig(level=logging.INFO, format='%(asctime)s [%(levelname)s] - %(message)s', handlers=[logging.FileHandler(os.path.join(args.output_dir, 'train.log')), logging.StreamHandler(sys.stdout)])
  175. if args.deterministic: torch.use_deterministic_algorithms(True, warn_only=True)
  176. random.seed(args.seed); np.random.seed(args.seed); torch.manual_seed(args.seed)
  177. device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
  178. substitution_matrix = SubstitutionMatrix(args.w_blosum, args.w_biochem, args.softmax_temp)
  179. if args.enable_esmfold:
  180. try: ESMFOLD_MODEL = esm.pretrained.esmfold_v1().to(device).eval()
  181. except: ESMFOLD_MODEL = None
  182. df_tv = pd.read_csv(args.train_csv).dropna(subset=['sequence', 'label'])
  183. s, p, l = df_tv["sequence"].tolist(), df_tv.get("pdb_dir", pd.Series([None]*len(df_tv))).tolist(), df_tv["label"].astype(int).tolist()
  184. val_ldr = None
  185. if args.val_split_ratio > 0:
  186. sss = StratifiedShuffleSplit(n_splits=1, test_size=args.val_split_ratio, random_state=args.seed)
  187. tr_idx, v_idx = next(sss.split(np.zeros(len(l)), l))
  188. train_ds = SeqPDBDataset([s[i] for i in tr_idx], [p[i] for i in tr_idx], [l[i] for i in tr_idx])
  189. val_ldr = DataLoader(SeqPDBDataset([s[i] for i in v_idx], [p[i] for i in v_idx], [l[i] for i in v_idx]), batch_size=args.batch, shuffle=False, num_workers=4, collate_fn=collate_seq_pdb)
  190. else: train_ds = SeqPDBDataset(s, p, l)
  191. train_sampler = StratifiedBatchSampler(train_ds.valid_labels, args.batch, True, seed=args.seed)
  192. train_ldr = DataLoader(train_ds, batch_sampler=train_sampler, num_workers=4, collate_fn=collate_seq_pdb, worker_init_fn=seed_worker)
  193. model = BP_INFP(hid=args.hid, pdb_node_feature_dim=PDB_NODE_FEATURE_DIM, freeze_ratio_esm=args.freeze_ratio_esm, normalize_esm_channels=args.normalize_esm_channels, esm_channel_stats_path=args.esm_channel_stats_path, bank_size=args.bank_size, scl_embedding_dim=args.scl_embedding_dim, dropout=args.dropout, pdb_edge_dim=16, num_fusion_layers=args.num_fusion_layers).to(device)
  194. if args.esm_grad_ckpt:
  195. try: model.seq_enc.esm_model.set_grad_checkpointing(True)
  196. except: pass
  197. wp, wn = compute_class_weights(l)
  198. logging.info("--- STAGE 1 ---")
  199. optim_s1 = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=args.lr_stage1, weight_decay=5e-3)
  200. for ep in range(args.stage1_epochs):
  201. train_sampler.set_epoch(ep)
  202. tr_loss, tr_met = train_one_epoch(model, train_ldr, optim_s1, device, wp, wn, args, substitution_matrix, 1, ep, args.stage1_epochs)
  203. logging.info(f"S1 E{ep+1} Loss: {tr_loss:.4f}, AUC: {tr_met['auc']:.4f}")
  204. stage1_path = os.path.join(args.output_dir, "stage1.pt")
  205. torch.save(model.state_dict(), stage1_path)
  206. logging.info("--- STAGE 2 ---")
  207. if args.stage2_unfreeze_ratio is not None:
  208. start = int(model.seq_enc.num_layers_esm * args.stage2_unfreeze_ratio)
  209. for i in range(start, model.seq_enc.num_layers_esm):
  210. if i < len(model.seq_enc.esm_model.layers):
  211. for p in model.seq_enc.esm_model.layers[i].parameters(): p.requires_grad = True
  212. enc_p, head_p = model.get_parameter_groups()
  213. optim_s2 = torch.optim.AdamW([{'params': enc_p, 'lr': args.lr_stage2_encoder}, {'params': head_p, 'lr': args.lr_stage2_head}], weight_decay=5e-3)
  214. sched_s2 = torch.optim.lr_scheduler.ReduceLROnPlateau(optim_s2, mode='max', factor=0.2, patience=3)
  215. best_key, best_thr, es_cnt, stage2_path = -1.0, 0.5, 0, os.path.join(args.output_dir, "stage2_best.pt")
  216. for ep in range(args.stage2_epochs):
  217. train_sampler.set_epoch(args.stage1_epochs + ep)
  218. tr_loss, tr_met = train_one_epoch(model, train_ldr, optim_s2, device, wp, wn, args, substitution_matrix, 2, ep, args.stage2_epochs)
  219. if val_ldr:
  220. v_loss, v_met = validate(model, val_ldr, device, ep, args.output_dir, "Valid", args=args, plot=True)
  221. score = v_met.get(args.monitor_metric if args.monitor_metric in ['ap','mcc','f1'] else 'mcc', 0.0)
  222. logging.info(f"S2 E{ep+1} Val {args.monitor_metric.upper()}: {v_met.get(args.monitor_metric,0):.4f}, Thr: {v_met['best_threshold']:.2f}")
  223. sched_s2.step(score)
  224. if score > best_key:
  225. best_key, best_thr, es_cnt = score, v_met['best_threshold'], 0
  226. torch.save({'state_dict': model.state_dict(), 'thr': best_thr}, stage2_path)
  227. else: es_cnt += 1
  228. if es_cnt >= 10: break
  229. else: torch.save({'state_dict': model.state_dict(), 'thr': 0.5}, stage2_path)
  230. logging.info("--- FINAL TEST ---")
  231. ckpt = torch.load(stage2_path)
  232. model.load_state_dict(ckpt['state_dict'])
  233. df_test = pd.read_csv(args.test_csv).dropna(subset=['sequence', 'label'])
  234. test_ldr = DataLoader(SeqPDBDataset(df_test["sequence"].tolist(), df_test.get("pdb_dir", pd.Series([None]*len(df_test))).tolist(), df_test["label"].astype(int).tolist()), batch_size=args.batch, shuffle=False, num_workers=4, collate_fn=collate_seq_pdb)
  235. t_loss, t_met = validate(model, test_ldr, device, None, args.output_dir, "Test", threshold=ckpt['thr'], args=args, plot=True)
  236. logging.info(f"[FINAL] Loss: {t_loss:.4f}, MCC: {t_met['mcc']:.4f}, SN: {t_met['sn']:.4f}, SP: {t_met['sp']:.4f}, Thr: {ckpt['thr']:.2f}")
  237. if __name__ == "__main__":
  238. run()

model.py at commit 2d1fd6e, no license · at the source

Overview

Authors: Jingwei Lv1, Qianyang Wu1, Jian Liu1, Binlu Yang1, Yuanhao Li1, Junlin Xu2, Yajie Meng3, Leyi Wei4, Zilong Zhang1, Quan Zou5, Xiulai Li1, Feifei Cui1
  1. School of Computer Science and Technology Hainan University Haikou China
  2. School of Computer Science and Technology Wuhan University of Science and Technology Wuhan Hubei China
  3. School of Computer Science and Artificial Intelligence Wuhan Textile University Wuhan Hubei China
  4. Centre For Artificial Intelligence‐Driven Drug Discovery Faculty of Applied Science Macao Polytechnic University Macao SAR China
  5. Institute of Fundamental and Frontier Sciences University of Electronic Science and Technology of China Chengdu China
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany), volume 13, issue 34, article e23984
Dates: received 25 November 2025; accepted 25 March 2026; published online 3 April 2026; in print June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1002/advs.202523984 · PMID 41933929 · PMCID PMC13285127 · OpenAlex W7149395500
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), methods / tools (subfield)
Methods: Connectivity, Statistics
Keywords: contrastive learning, cns drug delivery, interpretability, multi‐modal deep learning, physicochemical‐guided augmentation
MeSH: Blood-Brain Barrier*, Cell-Penetrating Peptides*, Deep Learning*, Humans (* major topic)
Topic: Biochemical and Structural Characterization (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Natural Science Foundation of China (National Science Foundation of China) (62450002 62572156); Science and Technology Development Fund of Macau (0177/2023/RIA3)
Citations: not cited yet (Europe PMC); 56 references in the paper

Abstract

Functional peptide discovery, particularly for blood–brain barrier‐penetrating peptides (BBBPPs), is strictly limited by extreme data scarcity and the “black‐box” nature of deep learning. Here, INB3P is presented as a physics‐informed, multi‐modal framework designed to address these challenges. Physicochemical‐guided mutagenesis (PCGM), a novel augmentation strategy that enforces biochemical constraints to expand training diversity without violating the biological manifold. INB3P integrates PCGM with a bi‐directional co‐attention mechanism fusing sequence and structure, optimized via contrastive learning and a Stable‐MCC loss. INB3P significantly outperforms state‐of‐the‐art baselines on the same independent test set used in a prior study. Crucially, the model autonomously rediscovers known biophysical mechanisms—including amphipathic motifs and long‐range contact stabilization—providing strong in silico validation of its learned representations. This work establishes a generalizable paradigm for learning from small, imbalanced biological datasets. To facilitate community adoption, a web server is provided at http://www.bioai‐lab.com/INBP, featuring a standalone PCGM module, empowering researchers to apply physics‐guided augmentation strategy to their own sparse datasets.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

EuclidLv/INB-P

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 2d1fd6ead71294da8846cbe653a07ed40f8179dd, 24 March 2026
Languages: Python (14)
Size: 24 files, 14 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (11 files), PyTorch (10 files), pandas (4 files), PyTorch Geometric (4 files), scikit-learn (3 files), Biopython (2 files), Matplotlib (2 files), SciPy (1 file), seaborn (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
15 files

bioai-lab.com/inbp

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: the link answers
Software Heritage: not checked
Found in: “Data Availability Statement”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 28 September 2026: the link answers (HTTP 200)
  • 28 September 2026: the link answers (HTTP 200)
At the source: bioai-lab.com/INBP

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 14 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

The INB3P web server, comprising the predictor, interpretability dashboard, and the standalone PCGM augmentation tool, is freely available at http://www.bioai‐lab.com/INBP (http://www.bioai-lab.com/INBP). The source code and training logs are available to researchers and developers at https://github.com/EuclidLv/INB‐P (https://github.com/EuclidLv/INB-P). The hyperparameters and training commands used in this study are detailed in the Supporting Information. The Zenodo archive containing the materials provided to the reviewers during peer review is now publicly available at:https://zenodo.org/records/17667996

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 12 authors, 5 keywords, 4 MeSH terms, 2 funders, 49 references.

Cite

This paper

Lv, J., Wu, Q., Liu, J., Yang, B., Li, Y., Xu, J., Meng, Y., Wei, L., Zhang, Z., Zou, Q., Li, X., & Cui, F. (2026). INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery. Advanced science (Weinheim, Baden-Wurttemberg, Germany), 13(34), e23984. https://doi.org/10.1002/advs.202523984

BibTeX

@article{lv2026inb,
author = {Lv, Jingwei and Wu, Qianyang and Liu, Jian and Yang, Binlu and Li, Yuanhao and Xu, Junlin and Meng, Yajie and Wei, Leyi and Zhang, Zilong and Zou, Quan and Li, Xiulai and Cui, Feifei},
title = {{INB\&lt;sup\&gt;3\&lt;/sup\&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery}},
journal = {Advanced science (Weinheim, Baden-Wurttemberg, Germany)},
year = {2026},
month = apr,
volume = {13},
number = {34},
pages = {e23984},
publisher = {Wiley},
issn = {2198-3844},
doi = {10.1002/advs.202523984},
url = {https://doi.org/10.1002/advs.202523984},
pmid = {41933929},
pmcid = {PMC13285127}
}

RIS

TY - JOUR
AU - Lv, Jingwei
AU - Wu, Qianyang
AU - Liu, Jian
AU - Yang, Binlu
AU - Li, Yuanhao
AU - Xu, Junlin
AU - Meng, Yajie
AU - Wei, Leyi
AU - Zhang, Zilong
AU - Zou, Quan
AU - Li, Xiulai
AU - Cui, Feifei
TI - INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery
T2 - Advanced science (Weinheim, Baden-Wurttemberg, Germany)
J2 - Adv Sci (Weinh)
PY - 2026
DA - 2026/04/03
VL - 13
IS - 34
SP - e23984
SN - 2198-3844
PB - Wiley
DO - 10.1002/advs.202523984
UR - https://doi.org/10.1002/advs.202523984
LA - en
ER -

CSL-JSON

{
"id": "10.1002/advs.202523984",
"type": "article-journal",
"title": "INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery",
"container-title": "Advanced science (Weinheim, Baden-Wurttemberg, Germany)",
"author": [
{
"family": "Lv",
"given": "Jingwei"
},
{
"family": "Wu",
"given": "Qianyang"
},
{
"family": "Liu",
"given": "Jian"
},
{
"family": "Yang",
"given": "Binlu"
},
{
"family": "Li",
"given": "Yuanhao"
},
{
"family": "Xu",
"given": "Junlin"
},
{
"family": "Meng",
"given": "Yajie"
},
{
"family": "Wei",
"given": "Leyi"
},
{
"family": "Zhang",
"given": "Zilong"
},
{
"family": "Zou",
"given": "Quan"
},
{
"family": "Li",
"given": "Xiulai"
},
{
"family": "Cui",
"given": "Feifei"
}
],
"container-title-short": "Adv Sci (Weinh)",
"volume": "13",
"issue": "34",
"page": "e23984",
"DOI": "10.1002/advs.202523984",
"PMID": "41933929",
"PMCID": "PMC13285127",
"ISSN": "2198-3844",
"publisher": "Wiley",
"URL": "https://doi.org/10.1002/advs.202523984",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
3
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.molcel.2026.07.006 [code]
DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.
Journal: Molecular cell
In common: Biopython, PyTorch Geometric, PyTorch, 6 other tools, 1 reference
[2] doi:10.1523/eneuro.0362-25.2026 [code]
Similarities between &lt;i&gt;Ciona&lt;/i&gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
Journal: eNeuro
In common: Biopython, PyTorch Geometric, PyTorch, 6 other tools
[3] doi:10.1021/acsomega.5c09368 [code]
Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.
Journal: ACS omega
In common: Biopython, PyTorch Geometric, PyTorch, 6 other tools
[4] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: Biopython, PyTorch Geometric, PyTorch, 5 other tools
[5] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: Biopython, PyTorch Geometric, PyTorch, 5 other tools
[6] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: Biopython, PyTorch Geometric, PyTorch, 5 other tools
[7] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits &lt;i&gt;Porphyromonas gingivalis&lt;/i&gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: Biopython, PyTorch Geometric, PyTorch, 5 other tools
[8] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: Biopython, PyTorch, seaborn, 5 other tools, methods / tools
[9] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: PyTorch Geometric, PyTorch, seaborn, 5 other tools, methods / tools
[10] doi:10.1093/bioinformatics/btag540 [code]
Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.
Journal: Bioinformatics (Oxford, England)
In common: PyTorch Geometric, PyTorch, seaborn, 5 other tools, methods / tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.