OSCR

SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches
  1. [1] § Methods › α-Helix classification use case ↔ examples/3DCNN-SpheronizaTor.py, lines 791–913 · score 0.90 · weight decay, AdamW, Dice loss, BCE, optimizer, IoU
  2. [2] § Methods › α-Helix classification use case ↔ examples/3DCNN-SpheronizaTor.py, lines 188–237 · score 0.81 · linear layers, LeakyReLU, dilation, head, stem, padding
  3. [3] § Methods › Construction of residue-centered spatial representations ↔ src/spheronizator/voxelBuilder.py, lines 143–154 · score 0.57 · buildSphere, get_boxProjection, position, spheres, voxel, atoms
  4. [4] § Methods › Protein structure acquisition and preprocessing ↔ src/spheronizator/mol2parser.py, lines 13–139 · score 0.56 · mol2parser, Tripos, Biopython, parsed, matching, atomic
  5. [5] § Methods › Construction of residue-centered spatial representations ↔ src/spheronizator/functions.py, lines 168–221 · score 0.50 · get_boxProjection, vector, distance, position, atoms, residue

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,095 lines · 35 KB · BSD-3-Clause · 2 matches

  1. #!/usr/bin/env python
  2. # coding: utf-8
  3. # In[1]:
  4. #!/usr/bin/env python3
  5. # -*- coding: utf-8 -*-
  6. from __future__ import annotations
  7. import os
  8. import json
  9. import warnings
  10. import argparse
  11. from dataclasses import dataclass, asdict
  12. from pathlib import Path
  13. import numpy as np
  14. import pandas as pd
  15. import matplotlib.pyplot as plt
  16. import torch
  17. import torch.nn as nn
  18. import torch.optim as optim
  19. import torch.nn.functional as F
  20. from torch.utils.data import Dataset, DataLoader
  21. import MDAnalysis as mda
  22. from MDAnalysis.analysis.dssp import DSSP
  23. from pyuul import utils, VolumeMaker
  24. from sklearn.metrics import (
  25. roc_curve,
  26. auc,
  27. confusion_matrix,
  28. ConfusionMatrixDisplay,
  29. roc_auc_score,
  30. precision_score,
  31. recall_score,
  32. f1_score,
  33. matthews_corrcoef,
  34. accuracy_score,
  35. )
  36. from sklearn.model_selection import train_test_split
  37. # =========================================================
  38. # CONFIG
  39. # =========================================================
  40. @dataclass
  41. class Config:
  42. backend: str = "spheronizator" # pyuul | spheronizator | compare
  43. pdb_dir: Path = Path("data/classification_data")
  44. pdb_pattern: str = "*fixed.pdb"
  45. sphero_dataset_dir: Path = Path(".")
  46. sphero_atoms_dir: Path = Path("output_vox_atoms")
  47. sphero_bonds_dir: Path = Path("output_vox_bonds")
  48. resolution: float = 1.0
  49. epochs: int = 10
  50. learning_rate: float = 1e-3
  51. batch_size: int = 1
  52. num_workers: int = 0
  53. pin_memory: bool = False
  54. use_amp: bool = True
  55. plot_first_sample: bool = False
  56. save_checkpoint_every: int = 1
  57. sphero_aggregate: str = "sum" # sum | mean | max
  58. device: str = "cuda" if torch.cuda.is_available() else "cpu"
  59. target_build_device: str = "cuda" if torch.cuda.is_available() else "cpu"
  60. debug: bool = False
  61. test_size: float = 0.2
  62. random_state: int = 42
  63. cls_threshold: float = 0.5
  64. output_dir: Path = Path("outputs_classifier")
  65. def parse_args():
  66. parser = argparse.ArgumentParser(description="3D CNN protein classifier with pyuul vs spheronizator comparison")
  67. parser.add_argument("--backend", default="spheronizator", choices=["spheronizator", "pyuul", "compare"])
  68. parser.add_argument("--epochs", type=int, default=10)
  69. parser.add_argument("--lr", type=float, default=1e-2)
  70. parser.add_argument("--resolution", type=float, default=1)
  71. parser.add_argument("--aggregate", default="sum", choices=["sum", "mean", "max"])
  72. parser.add_argument("--batch_size", type=int, default=1)
  73. parser.add_argument("--num_workers", type=int, default=0)
  74. parser.add_argument("--plot_sample", action="store_true")
  75. parser.add_argument("--debug", action="store_true")
  76. parser.add_argument("--test_size", type=float, default=0.2)
  77. parser.add_argument("--cls_threshold", type=float, default=0.25)
  78. parser.add_argument("--output_dir", type=str, default="outputs_classifier")
  79. return parser.parse_args()
  80. def build_config():
  81. args = parse_args()
  82. cfg = Config()
  83. cfg.backend = args.backend
  84. cfg.epochs = args.epochs
  85. cfg.learning_rate = args.lr
  86. cfg.resolution = args.resolution
  87. cfg.sphero_aggregate = args.aggregate
  88. cfg.batch_size = args.batch_size
  89. cfg.num_workers = args.num_workers
  90. cfg.plot_first_sample = args.plot_sample
  91. cfg.debug = args.debug
  92. cfg.test_size = args.test_size
  93. cfg.cls_threshold = args.cls_threshold
  94. cfg.output_dir = Path(args.output_dir)
  95. cfg.sphero_atoms_dir = cfg.sphero_dataset_dir / "output_vox_atoms"
  96. cfg.sphero_bonds_dir = cfg.sphero_dataset_dir / "output_vox_bonds"
  97. os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
  98. cfg.output_dir.mkdir(parents=True, exist_ok=True)
  99. return cfg
  100. def backend_output_dir(cfg: Config, backend: str) -> Path:
  101. out = cfg.output_dir / backend
  102. out.mkdir(parents=True, exist_ok=True)
  103. return out
  104. def debug_print(cfg: Config, *args):
  105. if cfg.debug:
  106. print(*args)
  107. # =========================================================
  108. # MODEL
  109. # =========================================================
  110. class ResidualBlock3D(nn.Module):
  111. def __init__(self, in_channels: int, out_channels: int, groups: int = 4, dropout: float = 0.0):
  112. super().__init__()
  113. g = min(groups, out_channels)
  114. while out_channels % g != 0:
  115. g -= 1
  116. self.conv1 = nn.Conv3d(in_channels, out_channels, kernel_size=3, padding=1, bias=False)
  117. self.norm1 = nn.GroupNorm(g, out_channels)
  118. self.act1 = nn.LeakyReLU(0.1, inplace=True)
  119. self.conv2 = nn.Conv3d(out_channels, out_channels, kernel_size=3, padding=1, bias=False)
  120. self.norm2 = nn.GroupNorm(g, out_channels)
  121. self.dropout = nn.Dropout3d(dropout) if dropout > 0 else nn.Identity()
  122. if in_channels != out_channels:
  123. self.skip = nn.Conv3d(in_channels, out_channels, kernel_size=1, bias=False)
  124. else:
  125. self.skip = nn.Identity()
  126. self.act2 = nn.LeakyReLU(0.1, inplace=True)
  127. def forward(self, x: torch.Tensor) -> torch.Tensor:
  128. identity = self.skip(x)
  129. out = self.conv1(x)
  130. out = self.norm1(out)
  131. out = self.act1(out)
  132. out = self.dropout(out)
  133. out = self.conv2(out)
  134. out = self.norm2(out)
  135. out = out + identity
  136. out = self.act2(out)
  137. return out
  138. class Conv3dVoxelClassifier(nn.Module):
  139. def __init__(self, ch_in: int):
  140. super().__init__()
  141. self.stem = nn.Sequential(
  142. nn.Conv3d(ch_in, 16, kernel_size=3, padding=1, bias=False),
  143. nn.GroupNorm(4, 16),
  144. nn.LeakyReLU(0.1, inplace=True),
  145. )
  146. self.features = nn.Sequential(
  147. ResidualBlock3D(16, 32, groups=4, dropout=0.05),
  148. ResidualBlock3D(32, 32, groups=4, dropout=0.10),
  149. ResidualBlock3D(32, 32, groups=4, dropout=0.10),
  150. ResidualBlock3D(32, 16, groups=4, dropout=0.05),
  151. )
  152. self.context = nn.Sequential(
  153. nn.Conv3d(16, 16, kernel_size=3, padding=2, dilation=2, bias=False),
  154. nn.GroupNorm(4, 16),
  155. nn.LeakyReLU(0.1, inplace=True),
  156. )
  157. self.head = nn.Sequential(
  158. nn.Conv3d(32, 16, kernel_size=3, padding=1, bias=False),
  159. nn.GroupNorm(4, 16),
  160. nn.LeakyReLU(0.1, inplace=True),
  161. nn.Conv3d(16, 1, kernel_size=1),
  162. )
  163. #linear layer
  164. self.classifier = nn.Sequential(
  165. nn.Linear(1, 8),
  166. nn.LeakyReLU(0.1, inplace=True),
  167. nn.Dropout(0.1),
  168. nn.Linear(8, 1)
  169. )
  170. def forward(self, x: torch.Tensor) -> torch.Tensor:
  171. x = self.stem(x)
  172. feat = self.features(x)
  173. ctx = self.context(feat)
  174. x = torch.cat([feat, ctx], dim=1)
  175. x = self.head(x)
  176. # NOVO: global pooling
  177. x = x.mean(dim=(2, 3, 4)) # (N, 1)
  178. x = self.classifier(x) # (N, 1)
  179. return x
  180. # =========================================================
  181. # HELPERS
  182. # =========================================================
  183. def ensure_coords_shape(coords: torch.Tensor) -> torch.Tensor:
  184. if coords.ndim == 2:
  185. if coords.shape[1] != 3:
  186. raise ValueError(f"coords expected [N,3], got {tuple(coords.shape)}")
  187. coords = coords.unsqueeze(0)
  188. elif coords.ndim == 3:
  189. if coords.shape[-1] != 3:
  190. raise ValueError(f"coords expected [B,N,3], got {tuple(coords.shape)}")
  191. else:
  192. raise ValueError(f"coords must have 2 or 3 dims, got {coords.ndim}")
  193. return coords.float()
  194. def ensure_radii_shape(radii: torch.Tensor) -> torch.Tensor:
  195. if radii.ndim == 1:
  196. radii = radii.unsqueeze(0)
  197. elif radii.ndim == 2:
  198. if radii.shape[0] == 1:
  199. pass
  200. elif radii.shape[1] == 1:
  201. radii = radii.squeeze(1).unsqueeze(0)
  202. else:
  203. raise ValueError(f"Unexpected radii shape: {tuple(radii.shape)}")
  204. elif radii.ndim == 3:
  205. if radii.shape[0] == 1 and radii.shape[2] == 1:
  206. radii = radii.squeeze(-1)
  207. else:
  208. raise ValueError(f"Unexpected radii shape: {tuple(radii.shape)}")
  209. else:
  210. raise ValueError(f"radii must have 1, 2 or 3 dims, got {radii.ndim}")
  211. return radii.float()
  212. def ensure_atomwise_2d(x: torch.Tensor, n_atoms_hint: int | None = None, name: str = "tensor") -> torch.Tensor:
  213. if x.ndim == 1:
  214. x = x.unsqueeze(0)
  215. elif x.ndim == 2:
  216. if x.shape[0] == 1:
  217. pass
  218. elif x.shape[1] == 1:
  219. x = x.squeeze(1).unsqueeze(0)
  220. else:
  221. if n_atoms_hint is not None:
  222. if x.shape[0] == n_atoms_hint:
  223. x = x.unsqueeze(0)
  224. elif x.shape[1] == n_atoms_hint:
  225. pass
  226. else:
  227. raise ValueError(f"Unexpected {name} shape: {tuple(x.shape)}")
  228. else:
  229. raise ValueError(f"Unexpected {name} shape: {tuple(x.shape)}")
  230. elif x.ndim == 3:
  231. if x.shape[0] != 1:
  232. raise ValueError(f"Unexpected {name} batch shape: {tuple(x.shape)}")
  233. if x.shape[2] == 1:
  234. x = x.squeeze(-1)
  235. elif x.shape[1] == 1:
  236. x = x.squeeze(1)
  237. else:
  238. if n_atoms_hint is not None:
  239. if x.shape[1] == n_atoms_hint:
  240. x = x[:, :, 0]
  241. elif x.shape[2] == n_atoms_hint:
  242. x = x[:, 0, :]
  243. else:
  244. raise ValueError(f"Unexpected {name} shape: {tuple(x.shape)}")
  245. else:
  246. raise ValueError(f"Unexpected {name} shape: {tuple(x.shape)}")
  247. else:
  248. raise ValueError(f"{name} must have 1, 2, or 3 dims, got {x.ndim}")
  249. if x.ndim != 2:
  250. raise ValueError(f"{name} could not be normalized to [B,N], got {tuple(x.shape)}")
  251. return x.float()
  252. def align_atomwise_tensors(coords, atom_channels, radii, atom_labels):
  253. n = min(coords.shape[1], atom_channels.shape[1], radii.shape[1], atom_labels.shape[1])
  254. coords = coords[:, :n, :]
  255. atom_channels = atom_channels[:, :n]
  256. radii = radii[:, :n]
  257. atom_labels = atom_labels[:, :n]
  258. return coords, atom_channels, radii, atom_labels
  259. def ensure_volume_shape_for_cnn(volume: torch.Tensor) -> torch.Tensor:
  260. if volume.ndim == 4:
  261. volume = volume.unsqueeze(1)
  262. elif volume.ndim == 5:
  263. if volume.shape[-1] <= 64 and volume.shape[1] > 64:
  264. volume = volume.permute(0, 4, 1, 2, 3).contiguous()
  265. else:
  266. raise ValueError(f"Unexpected volume shape: {tuple(volume.shape)}")
  267. return volume.float()
  268. # =========================================================
  269. # LABELS
  270. # =========================================================
  271. """
  272. def extract_atom_labels_from_dssp(pdb_file: Path) -> np.ndarray:
  273. with warnings.catch_warnings():
  274. warnings.simplefilter("ignore")
  275. u = mda.Universe(str(pdb_file))
  276. protein = u.select_atoms("protein")
  277. residues = protein.residues
  278. dssp = DSSP(u).run()
  279. res_labels = dssp.results.dssp[0]
  280. n = min(len(residues), len(res_labels))
  281. atom_labels = []
  282. for i in range(n):
  283. residue = residues[i]
  284. ss = res_labels[i]
  285. label = 1.0 if ss == "H" else 0.0
  286. for _atom in residue.atoms:
  287. atom_labels.append(label)
  288. return np.asarray(atom_labels, dtype=np.float32)
  289. """
  290. def extract_atom_labels_from_dssp(pdb_file: Path) -> np.ndarray:
  291. with warnings.catch_warnings():
  292. warnings.simplefilter("ignore")
  293. u = mda.Universe(str(pdb_file))
  294. protein = u.select_atoms("protein")
  295. residues = protein.residues
  296. dssp = DSSP(u).run()
  297. res_labels = dssp.results.dssp[0]
  298. n = min(len(residues), len(res_labels))
  299. atom_labels = []
  300. for i in range(n):
  301. residue = residues[i]
  302. ss = res_labels[i]
  303. label = 1.0 if ss == "H" else 0.0
  304. #for _atom in residue.atoms:
  305. atom_labels.append(label)
  306. return np.asarray(atom_labels, dtype=np.float32)
  307. def get_pdb_files(pdb_dir: Path, pattern: str) -> list[Path]:
  308. files = sorted(pdb_dir.glob(pattern))
  309. if not files:
  310. raise FileNotFoundError(f"No PDB files found in: {pdb_dir.resolve()}")
  311. return files
  312. # =========================================================
  313. # BUILDERS
  314. # =========================================================
  315. """
  316. def build_pyuul_input_and_target(pdb_file: Path, device: str = "cpu", resolution: float = 0.5):
  317. coords, atname = utils.parsePDB(str(pdb_file))
  318. atom_channels = utils.atomlistToChannels(atname)
  319. radii = utils.atomlistToRadius(atname)
  320. coords = ensure_coords_shape(coords.to(device))
  321. radii = ensure_radii_shape(radii.to(device))
  322. atom_labels_np = extract_atom_labels_from_dssp(pdb_file)
  323. atom_labels = torch.tensor(atom_labels_np, dtype=torch.float32, device=device)
  324. n_atoms_hint = coords.shape[1]
  325. atom_channels = ensure_atomwise_2d(atom_channels.to(device), n_atoms_hint=n_atoms_hint, name="atom_channels")
  326. #atom_labels = ensure_atomwise_2d(atom_labels, n_atoms_hint=n_atoms_hint, name="atom_labels")
  327. residue_labels_np = extract_atom_labels_from_dssp(pdb_file) # [N]
  328. target_volume = torch.tensor(residue_labels_np, dtype=torch.float32).unsqueeze(1) # [N, 1]
  329. coords, atom_channels, radii, atom_labels = align_atomwise_tensors(coords, atom_channels, radii, atom_labels)
  330. volmaker = VolumeMaker.Voxels(device=device)
  331. input_volume = volmaker(coords, radii, atom_channels, resolution=resolution).to_dense().float()
  332. input_volume = ensure_volume_shape_for_cnn(input_volume)
  333. if input_volume.shape[0] != target_volume.shape[0]:
  334. n = min(input_volume.shape[0], target_volume.shape[0])
  335. input_volume = input_volume[:n]
  336. target_volume = target_volume[:n]
  337. target_volume = (target_volume > 0).float()
  338. # NOVO: reduzir voxel → label global
  339. target_volume = target_volume.amax(dim=(2, 3, 4)) # (N, 1)
  340. return input_volume, target_volume
  341. """
  342. def build_pyuul_input_and_target(
  343. pdb_file: Path,
  344. device: str = "cpu",
  345. resolution: float = 1.0,
  346. box_size: int = 41,
  347. ):
  348. """
  349. Build per-residue local PyUUL volumes for residue-level classification.
  350. Returns
  351. -------
  352. input_volume : torch.Tensor
  353. Shape [N_res, C, box_size, box_size, box_size]
  354. target : torch.Tensor
  355. Shape [N_res, 1]
  356. """
  357. # --------------------------------------------------
  358. # Parse structure with PyUUL
  359. # --------------------------------------------------
  360. coords, atname = utils.parsePDB(str(pdb_file))
  361. atom_channels = utils.atomlistToChannels(atname)
  362. radii = utils.atomlistToRadius(atname)
  363. coords = ensure_coords_shape(coords.to(device)) # [1, N_atoms, 3]
  364. radii = ensure_radii_shape(radii.to(device)) # [1, N_atoms]
  365. n_atoms_hint = coords.shape[1]
  366. atom_channels = ensure_atomwise_2d(
  367. atom_channels.to(device),
  368. n_atoms_hint=n_atoms_hint,
  369. name="atom_channels"
  370. ) # [1, N_atoms]
  371. # number of atom-type channels expected by the model
  372. n_input_channels = int(atom_channels.max().item()) + 1
  373. # --------------------------------------------------
  374. # Residue-level labels from DSSP
  375. # IMPORTANT: extract_atom_labels_from_dssp must return
  376. # one label per residue for this pipeline
  377. # --------------------------------------------------
  378. residue_labels_np = extract_atom_labels_from_dssp(pdb_file) # [N_res]
  379. target = torch.tensor(
  380. residue_labels_np,
  381. dtype=torch.float32,
  382. device=device
  383. ).unsqueeze(1) # [N_res, 1]
  384. # --------------------------------------------------
  385. # Load residues with MDAnalysis
  386. # --------------------------------------------------
  387. with warnings.catch_warnings():
  388. warnings.simplefilter("ignore")
  389. u = mda.Universe(str(pdb_file))
  390. protein = u.select_atoms("protein")
  391. residues = protein.residues
  392. n_res = min(len(residues), target.shape[0])
  393. if n_res == 0:
  394. raise ValueError(f"No valid residues found for {pdb_file}")
  395. # trim target if needed
  396. target = target[:n_res]
  397. volmaker = VolumeMaker.Voxels(device=device)
  398. half_box_ang = (box_size * resolution) / 2.0
  399. local_volumes = []
  400. # global coords tensor without batch dim -> [N_atoms, 3]
  401. global_coords = coords[0]
  402. # --------------------------------------------------
  403. # Build one local volume per residue
  404. # --------------------------------------------------
  405. for i in range(n_res):
  406. res = residues[i]
  407. # Prefer backbone centroid (N, CA, C); fallback to residue center
  408. backbone = res.atoms.select_atoms("name N CA C")
  409. if len(backbone) > 0:
  410. center_np = backbone.positions.mean(axis=0)
  411. else:
  412. center_np = res.atoms.center_of_geometry()
  413. center = torch.tensor(center_np, dtype=torch.float32, device=device)
  414. # Local coordinates centered on the residue
  415. local_coords = global_coords - center # [N_atoms, 3]
  416. # Cubic local crop around the residue center
  417. mask = (
  418. (local_coords[:, 0].abs() <= half_box_ang) &
  419. (local_coords[:, 1].abs() <= half_box_ang) &
  420. (local_coords[:, 2].abs() <= half_box_ang)
  421. )
  422. # Fallback: empty local box with correct channel count
  423. if mask.sum().item() == 0:
  424. local_volume = torch.zeros(
  425. (1, n_input_channels, box_size, box_size, box_size),
  426. dtype=torch.float32,
  427. device=device
  428. )
  429. local_volumes.append(local_volume)
  430. continue
  431. sel_coords = local_coords[mask].unsqueeze(0) # [1, n_sel, 3]
  432. sel_channels = atom_channels[:, mask] # [1, n_sel]
  433. sel_radii = radii[:, mask] # [1, n_sel]
  434. # Voxelize local environment
  435. local_volume = volmaker(
  436. sel_coords,
  437. sel_radii,
  438. sel_channels,
  439. resolution=resolution
  440. ).to_dense().float()
  441. # Normalize tensor layout to [1, C, D, H, W]
  442. if local_volume.ndim == 5:
  443. # Sometimes dense output may come as [1, D, H, W, C]
  444. if local_volume.shape[1] != n_input_channels and local_volume.shape[-1] == n_input_channels:
  445. local_volume = local_volume.permute(0, 4, 1, 2, 3).contiguous()
  446. elif local_volume.ndim == 4:
  447. local_volume = local_volume.unsqueeze(1)
  448. else:
  449. raise ValueError(
  450. f"Unexpected local_volume shape for residue {i} in {pdb_file}: "
  451. f"{tuple(local_volume.shape)}"
  452. )
  453. # If channel dimension is still wrong, force a safe fix when possible
  454. if local_volume.shape[1] != n_input_channels:
  455. if local_volume.shape[1] == 1 and local_volume.shape[-1] == n_input_channels:
  456. local_volume = local_volume.permute(0, 4, 1, 2, 3).contiguous()
  457. elif local_volume.shape[1] != n_input_channels:
  458. raise ValueError(
  459. f"Channel mismatch for residue {i} in {pdb_file}: "
  460. f"expected {n_input_channels}, got {local_volume.shape[1]}"
  461. )
  462. # Force fixed spatial size
  463. if local_volume.shape[-3:] != (box_size, box_size, box_size):
  464. local_volume = F.interpolate(
  465. local_volume,
  466. size=(box_size, box_size, box_size),
  467. mode="nearest"
  468. )
  469. local_volumes.append(local_volume)
  470. input_volume = torch.cat(local_volumes, dim=0) # [N_res, C, box, box, box]
  471. # Final safety check
  472. if input_volume.shape[0] != target.shape[0]:
  473. n = min(input_volume.shape[0], target.shape[0])
  474. input_volume = input_volume[:n]
  475. target = target[:n]
  476. return input_volume.float(), target.float()
  477. def build_spheronizator_input_and_target(
  478. pdb_file: Path,
  479. atoms_dir: Path,
  480. bonds_dir: Path,
  481. target_device: str = "cpu",
  482. resolution: float = 0.5,
  483. aggregate: str = "sum",
  484. ):
  485. input_volume = load_spheronizator_volume(pdb_file, atoms_dir, bonds_dir) # [N, C, D, H, W]
  486. residue_labels_np = extract_atom_labels_from_dssp(pdb_file) # [N]
  487. target_volume = torch.tensor(residue_labels_np, dtype=torch.float32, device=target_device).unsqueeze(1) # [N, 1]
  488. if input_volume.shape[0] != target_volume.shape[0]:
  489. n = min(input_volume.shape[0], target_volume.shape[0])
  490. input_volume = input_volume[:n]
  491. target_volume = target_volume[:n]
  492. return input_volume, target_volume
  493. # =========================================================
  494. # DATASETS
  495. # =========================================================
  496. class PyuulDataset(Dataset):
  497. def __init__(self, pdb_files: list[Path], resolution: float = 0.5, build_device: str = "cpu"):
  498. self.pdb_files = list(pdb_files)
  499. self.resolution = resolution
  500. self.build_device = build_device
  501. def __len__(self):
  502. return len(self.pdb_files)
  503. def __getitem__(self, idx):
  504. pdb_file = self.pdb_files[idx]
  505. x, y = build_pyuul_input_and_target(pdb_file, device=self.build_device, resolution=self.resolution)
  506. return x, y, pdb_file.name
  507. class SpheronizatorDataset(Dataset):
  508. def __init__(
  509. self,
  510. pdb_files: list[Path],
  511. atoms_dir: Path,
  512. bonds_dir: Path,
  513. resolution: float = 0.5,
  514. target_build_device: str = "cuda",
  515. aggregate: str = "sum",
  516. ):
  517. self.pdb_files = list(pdb_files)
  518. self.atoms_dir = Path(atoms_dir)
  519. self.bonds_dir = Path(bonds_dir)
  520. self.resolution = resolution
  521. self.target_build_device = target_build_device
  522. self.aggregate = aggregate
  523. def __len__(self):
  524. return len(self.pdb_files)
  525. def __getitem__(self, idx):
  526. pdb_file = self.pdb_files[idx]
  527. x, y = build_spheronizator_input_and_target(
  528. pdb_file=pdb_file,
  529. atoms_dir=self.atoms_dir,
  530. bonds_dir=self.bonds_dir,
  531. target_device=self.target_build_device,
  532. resolution=self.resolution,
  533. aggregate=self.aggregate,
  534. )
  535. return x, y, pdb_file.name
  536. def single_item_collate(batch):
  537. return batch[0]
  538. def build_loader(backend: str, pdb_files: list[Path], cfg: Config) -> DataLoader:
  539. if backend == "pyuul":
  540. dataset = PyuulDataset(
  541. pdb_files=pdb_files,
  542. resolution=cfg.resolution,
  543. build_device=cfg.target_build_device,
  544. )
  545. elif backend == "spheronizator":
  546. dataset = SpheronizatorDataset(
  547. pdb_files=pdb_files,
  548. atoms_dir=cfg.sphero_atoms_dir,
  549. bonds_dir=cfg.sphero_bonds_dir,
  550. resolution=cfg.resolution,
  551. target_build_device=cfg.target_build_device,
  552. aggregate=cfg.sphero_aggregate,
  553. )
  554. else:
  555. raise ValueError(f"Unknown backend: {backend}")
  556. return DataLoader(
  557. dataset,
  558. batch_size=cfg.batch_size,
  559. shuffle=True,
  560. num_workers=cfg.num_workers,
  561. pin_memory=cfg.pin_memory,
  562. collate_fn=single_item_collate,
  563. )
  564. # =========================================================
  565. # METRICS
  566. # =========================================================
  567. def compute_voxel_metrics(logits, y, threshold=0.5, eps=1e-8):
  568. probs = torch.sigmoid(logits)
  569. preds = (probs > threshold).float()
  570. y = y.float()
  571. tp = ((preds == 1) & (y == 1)).sum().float()
  572. tn = ((preds == 0) & (y == 0)).sum().float()
  573. fp = ((preds == 1) & (y == 0)).sum().float()
  574. fn = ((preds == 0) & (y == 1)).sum().float()
  575. acc = (tp + tn) / (tp + tn + fp + fn + eps)
  576. precision = tp / (tp + fp + eps)
  577. recall = tp / (tp + fn + eps)
  578. f1 = 2 * precision * recall / (precision + recall + eps)
  579. dice = (2 * tp) / (2 * tp + fp + fn + eps)
  580. iou = tp / (tp + fp + fn + eps)
  581. return {
  582. "acc": acc.item(),
  583. "precision": precision.item(),
  584. "recall": recall.item(),
  585. "f1": f1.item(),
  586. "dice": dice.item(),
  587. "iou": iou.item(),
  588. }
  589. def dice_loss_from_logits(logits, targets, eps=1e-6):
  590. probs = torch.sigmoid(logits)
  591. probs = probs.reshape(probs.shape[0], -1)
  592. targets = targets.reshape(targets.shape[0], -1).float()
  593. intersection = (probs * targets).sum(dim=1)
  594. union = probs.sum(dim=1) + targets.sum(dim=1)
  595. dice = (2 * intersection + eps) / (union + eps)
  596. return 1 - dice.mean()
  597. def plot_roc_curve(y_true, y_score, output_path="roc_curve.pdf", title="ROC Curve"):
  598. y_true = np.asarray(y_true).astype(int)
  599. y_score = np.asarray(y_score).astype(float)
  600. if len(np.unique(y_true)) < 2:
  601. print(f"[WARNING] ROC curve skipped for {title}: only one class present.")
  602. return float("nan")
  603. fpr, tpr, _ = roc_curve(y_true, y_score)
  604. roc_auc = auc(fpr, tpr)
  605. plt.figure(figsize=(6, 6))
  606. plt.plot(fpr, tpr, label=f"AUC = {roc_auc:.4f}")
  607. plt.plot([0, 1], [0, 1], linestyle="--")
  608. plt.xlabel("False Positive Rate")
  609. plt.ylabel("True Positive Rate")
  610. plt.title(title)
  611. plt.legend(loc="lower right")
  612. plt.tight_layout()
  613. plt.savefig(output_path)
  614. plt.close()
  615. return roc_auc
  616. def plot_confusion_matrix(y_true, y_pred, output_path="confusion_matrix.pdf", title="Confusion Matrix"):
  617. y_true = np.asarray(y_true).astype(int)
  618. y_pred = np.asarray(y_pred).astype(int)
  619. #cm = confusion_matrix(y_true, y_pred)
  620. cm = confusion_matrix(y_true, y_pred, labels=[0, 1])
  621. fig, ax = plt.subplots(figsize=(5, 5))
  622. disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=[0, 1])
  623. disp.plot(ax=ax, colorbar=False)
  624. ax.set_title(title)
  625. plt.tight_layout()
  626. plt.savefig(output_path)
  627. plt.close()
  628. return cm
  629. # =========================================================
  630. # TRAIN / EVAL
  631. # =========================================================
  632. def train_voxel_classification(
  633. model,
  634. train_loader,
  635. test_loader,
  636. epochs,
  637. lr,
  638. device,
  639. use_amp,
  640. save_checkpoint_every,
  641. backend,
  642. output_dir,
  643. ):
  644. bce = nn.BCEWithLogitsLoss()
  645. optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-4)
  646. use_amp = use_amp and device.startswith("cuda")
  647. scaler = torch.amp.GradScaler("cuda", enabled=use_amp)
  648. history = {
  649. "train_loss": [],
  650. "test_loss": [],
  651. "accuracy": [],
  652. "precision": [],
  653. "recall": [],
  654. "f1": [],
  655. "dice": [],
  656. "iou": [],
  657. }
  658. for epoch in range(1, epochs + 1):
  659. model.train()
  660. train_loss = 0.0
  661. for x, y, _ in train_loader:
  662. x = x.to(device, non_blocking=True)
  663. y = y.to(device, non_blocking=True).float()
  664. optimizer.zero_grad(set_to_none=True)
  665. with torch.amp.autocast(device_type="cuda", enabled=use_amp):
  666. logits = model(x)
  667. if logits.shape != y.shape:
  668. raise RuntimeError(
  669. f"Shape mismatch: output={tuple(logits.shape)} target={tuple(y.shape)}"
  670. )
  671. loss = bce(logits, y) + dice_loss_from_logits(logits, y)
  672. scaler.scale(loss).backward()
  673. scaler.step(optimizer)
  674. scaler.update()
  675. train_loss += loss.item()
  676. train_loss /= max(len(train_loader), 1)
  677. model.eval()
  678. test_loss = 0.0
  679. metric_sums = {
  680. "acc": 0.0,
  681. "precision": 0.0,
  682. "recall": 0.0,
  683. "f1": 0.0,
  684. "dice": 0.0,
  685. "iou": 0.0,
  686. }
  687. n_test = 0
  688. with torch.no_grad():
  689. for x, y, _ in test_loader:
  690. x = x.to(device, non_blocking=True)
  691. y = y.to(device, non_blocking=True).float()
  692. with torch.amp.autocast(device_type="cuda", enabled=use_amp):
  693. logits = model(x)
  694. loss = bce(logits, y) + dice_loss_from_logits(logits, y)
  695. test_loss += loss.item()
  696. metrics = compute_voxel_metrics(logits, y)
  697. for k, v in metrics.items():
  698. metric_sums[k] += v
  699. n_test += 1
  700. test_loss /= max(n_test, 1)
  701. epoch_metrics = {k: v / max(n_test, 1) for k, v in metric_sums.items()}
  702. history["train_loss"].append(train_loss)
  703. history["test_loss"].append(test_loss)
  704. history["accuracy"].append(epoch_metrics["acc"])
  705. history["precision"].append(epoch_metrics["precision"])
  706. history["recall"].append(epoch_metrics["recall"])
  707. history["f1"].append(epoch_metrics["f1"])
  708. history["dice"].append(epoch_metrics["dice"])
  709. history["iou"].append(epoch_metrics["iou"])
  710. print(
  711. f"[{backend}] Epoch {epoch:02d}/{epochs:02d} | "
  712. f"train_loss={train_loss:.4f} | "
  713. f"test_loss={test_loss:.4f} | "
  714. f"acc={epoch_metrics['acc']:.4f} | "
  715. f"prec={epoch_metrics['precision']:.4f} | "
  716. f"rec={epoch_metrics['recall']:.4f} | "
  717. f"f1={epoch_metrics['f1']:.4f} | "
  718. f"dice={epoch_metrics['dice']:.4f} | "
  719. f"iou={epoch_metrics['iou']:.4f}"
  720. )
  721. if save_checkpoint_every > 0 and epoch % save_checkpoint_every == 0:
  722. torch.save(
  723. {
  724. "epoch": epoch,
  725. "model_state_dict": model.state_dict(),
  726. "optimizer_state_dict": optimizer.state_dict(),
  727. "train_loss": train_loss,
  728. "test_loss": test_loss,
  729. "metrics": epoch_metrics,
  730. "backend": backend,
  731. },
  732. output_dir / f"checkpoint_voxel_{backend}_epoch_{epoch:02d}.pt",
  733. )
  734. return history
  735. # =========================================================
  736. # PLOTS
  737. # =========================================================
  738. def plot_sample_voxels(input_volume: torch.Tensor, target_volume: torch.Tensor, title: str = "", output_dir: Path | None = None):
  739. x = input_volume.detach().cpu()
  740. y = target_volume.detach().cpu()
  741. input_mask = (x[0].sum(dim=0) > 0)
  742. target_mask = (y[0, 0] > 0)
  743. fig = plt.figure(figsize=(12, 5))
  744. ax1 = fig.add_subplot(1, 2, 1, projection="3d")
  745. ax1.voxels(input_mask.numpy())
  746. ax1.set_title(f"{title} - Input")
  747. ax2 = fig.add_subplot(1, 2, 2, projection="3d")
  748. ax2.voxels(target_mask.numpy())
  749. ax2.set_title(f"{title} - Target")
  750. plt.tight_layout()
  751. out = (output_dir / f"{title}_sample_voxels.png") if output_dir else Path(f"{title}_sample_voxels.png")
  752. plt.savefig(out, dpi=200)
  753. plt.close(fig)
  754. def plot_classification_history(history: dict, prefix: str, output_dir: Path):
  755. plt.figure(figsize=(7, 4))
  756. plt.plot(history["train_loss"], marker="o", label="train_loss")
  757. plt.plot(history["test_loss"], marker="o", label="test_loss")
  758. plt.xlabel("Epoch")
  759. plt.ylabel("Loss")
  760. plt.title(f"Classification Loss - {prefix}")
  761. plt.legend()
  762. plt.tight_layout()
  763. plt.savefig(output_dir / f"{prefix}_cls_loss.pdf")
  764. plt.close()
  765. metrics_to_plot = ["accuracy", "precision", "recall", "f1"]
  766. plt.figure(figsize=(10, 5))
  767. for metric in metrics_to_plot:
  768. plt.plot(history[metric], marker="o", label=metric)
  769. plt.xlabel("Epoch")
  770. plt.ylabel("Score")
  771. plt.title(f"Classification Metrics - {prefix}")
  772. plt.legend()
  773. plt.tight_layout()
  774. plt.savefig(output_dir / f"{prefix}_cls_metrics.pdf")
  775. plt.close()
  776. def plot_backend_comparison(results_df: pd.DataFrame, output_path: Path):
  777. metrics = ["accuracy", "precision", "recall", "f1"]
  778. x = np.arange(len(matrics))
  779. width = 0.12
  780. offsets = np.linspace(-2.5 * width, 2.5 * width, len(metrics))
  781. plt.figure(figsize=(12, 6))
  782. for metric, offset in zip(metrics, offsets):
  783. plt.bar(x + offset, results_df[metric].values, width=width, label=metric)
  784. plt.xticks(x, backends)
  785. plt.ylim(0, 1)
  786. plt.ylabel("Score")
  787. plt.title("Backend Comparison - Final Test Metrics")
  788. plt.legend()
  789. plt.tight_layout()
  790. plt.savefig(output_path)
  791. plt.close()
  792. # =========================================================
  793. # EXPERIMENT RUNNER
  794. # =========================================================
  795. def run_backend_classification(cfg: Config, backend: str, train_files: list[Path], test_files: list[Path]):
  796. out_dir = backend_output_dir(cfg, backend)
  797. train_loader = build_loader(backend, train_files, cfg)
  798. test_loader = build_loader(backend, test_files, cfg)
  799. x0, y0, name0 = next(iter(train_loader))
  800. print(f"\n=== Backend: {backend} ===")
  801. print("First sample:")
  802. print(" name :", name0)
  803. print(" x :", tuple(x0.shape))
  804. print(" y :", tuple(y0.shape))
  805. if cfg.plot_first_sample:
  806. plot_sample_voxels(x0, y0, title=name0, output_dir=out_dir)
  807. ch_in = x0.shape[1]
  808. model = Conv3dVoxelClassifier(ch_in=ch_in).to(cfg.device)
  809. history = train_voxel_classification(
  810. model=model,
  811. train_loader=train_loader,
  812. test_loader=test_loader,
  813. epochs=cfg.epochs,
  814. lr=cfg.learning_rate,
  815. device=cfg.device,
  816. use_amp=cfg.use_amp,
  817. save_checkpoint_every=cfg.save_checkpoint_every,
  818. backend=backend,
  819. output_dir=out_dir)
  820. plot_classification_history(history, prefix=backend, output_dir=out_dir)
  821. model_path = out_dir / f"conv3d_classifier_{backend}.pt"
  822. torch.save(model.state_dict(), model_path)
  823. print(f"Model saved to: {model_path}")
  824. with open(out_dir / "config.json", "w") as f:
  825. json.dump(
  826. {k: str(v) if isinstance(v, Path) else v for k, v in asdict(cfg).items()},
  827. f,
  828. indent=2,
  829. )
  830. return history
  831. # =========================================================
  832. # MAIN
  833. # =========================================================
  834. def main():
  835. cfg = build_config()
  836. print("Device:", cfg.device)
  837. print("Backend mode:", cfg.backend)
  838. print("Resolution:", cfg.resolution)
  839. print("Output dir:", cfg.output_dir)
  840. pdb_files = get_pdb_files(cfg.pdb_dir, cfg.pdb_pattern)
  841. print(f"Found {len(pdb_files)} PDB files")
  842. train_files, test_files = train_test_split(
  843. pdb_files,
  844. test_size=cfg.test_size,
  845. random_state=cfg.random_state,
  846. shuffle=True,
  847. )
  848. if cfg.backend == "compare":
  849. summaries = []
  850. for backend in ["pyuul","spheronizator"]:
  851. summary = run_backend_classification(cfg, backend, train_files, test_files)
  852. summaries.append(summary)
  853. results_df = pd.DataFrame(summaries)
  854. results_df.to_csv(cfg.output_dir / "backend_comparison.csv", index=False)
  855. plot_backend_comparison(
  856. results_df,
  857. output_path=cfg.output_dir / "backend_comparison.pdf",
  858. )
  859. print("\n=== Final Backend Comparison ===")
  860. print(results_df.to_string(index=False))
  861. else:
  862. summary = run_backend_classification(cfg, cfg.backend, train_files, test_files)
  863. summary_df = pd.DataFrame([summary])
  864. summary_df.to_csv(cfg.output_dir / f"{cfg.backend}_summary.csv", index=False)
  865. print("\n=== Final Summary ===")
  866. print(summary_df.to_string(index=False))
  867. if __name__ == "__main__":
  868. main()
  869. # In[ ]:
  870. # python train_classifier.py --backend compare
  871. # python train_classifier.py --backend spheronizator
  872. # python train_classifier.py --backend compare

3DCNN-SpheronizaTor.py at commit 36b80d6, under BSD-3-Clause · at the source

Overview

  1. Department of Microbiology and Cell Science, Institute of Food and Agricultural Sciences, University of Florida, Gainesville, FL, USA
  2. Department of Chemical Engineering, Herbert Wertheim College of Engineering, University of Florida, Gainesville, FL, USA
Institutions: University of Florida (United States); Institute of Food and Agricultural Sciences (United States)
Journal: Computational and structural biotechnology journal, volume 35, issue 1, article 0076
Dates: received 17 December 2025; accepted 9 April 2026; published online 7 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.34133/csbj.0076 · PMID 42110217 · PMCID PMC13150068 · OpenAlex W7153140616
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: cellular / molecular (subfield)
Methods: Machine learning
Journal subjects: Software/Web Server Article, General
Topic: Cell Image Analysis Techniques (Biophysics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: not cited yet (Europe PMC); 32 references in the paper

Abstract

Artificial intelligence (AI) has expanded the reach of structural biology by enabling models to extract biochemical and geometric features directly from 3-dimensional (3D) protein structures. Yet, the effectiveness of these models depends critically on how protein environments are encoded. Most existing volumetric representations rely on Cartesian voxel grids derived from smoothed atomic densities, an approach that offers broad applicability but struggles to reconcile rotational invariance, residue-level specificity, and explicit biochemical detail. We present SpheronizaTor, a residue-centered voxelization framework for protein structures that builds local spherical voxel maps centered on each residue. Each spherical map encodes atom types, covalent bonding information, and whether atoms belong to the central residue or neighboring residues. By producing one voxel representation per residue, SpheronizaTor emphasizes the structural and functional granularity through which proteins organize catalysis, recognition, and stability. The combination of spherical alignment with chemically explicit feature channels enables richer interpretability and enhances compatibility with 3D convolutional and hybrid neural architectures. Designed specifically for proteins and engineered for extensibility, SpheronizaTor provides a voxelization strategy that is both chemically realistic and computationally efficient. The residue-centric approach bridges the gap between global volumetric encoders and graph-based models, offering a versatile foundation for downstream tasks such as mutation effect prediction, binding site analysis, and structural comparison across protein families.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

Dias-Lab/spheronizator

License: BSD-3-Clause
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: 36b80d66e065474d30036a931dff43bc4ad83291, 8 April 2026
Languages: Python (9), Jupyter (1)
Size: 45 files, 10 scripts
Software Heritage: not archived
Found in: “Data Availability”
Holds: README, license file, environment (pyproject.toml, setup.py), tests, continuous integration, documentation, 1 notebook
Not found: CITATION.cff
Tools: NumPy (7 files), Biopython (5 files), pandas (2 files), Matplotlib (1 file), MDAnalysis (1 file), PyTorch (1 file), scikit-learn (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
12 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 10 scripts, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability

The SpheronizaTor software is freely available on GitHub at https://github.com/Dias-Lab/spheronizator and can also be installed via pip (https://pypi.org/project/spheronizator/). Detailed documentation is available at https://dias-lab.github.io/spheronizator/.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 1 funder, 26 references.

Cite

This paper

Silva, J. C. F., Richardson, M., Cediel-Becerra, J. D. D., Schuster, L., & Dias, R. (2026). SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling. Computational and structural biotechnology journal, 35(1), 0076. https://doi.org/10.34133/csbj.0076

BibTeX

@article{silva2026spheronizator,
author = {Silva, Jose Cleydson Ferreira and Richardson, Matthew and Cediel-Becerra, José D. D. and Schuster, Layla and Dias, Raquel},
title = {{SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling}},
journal = {Computational and structural biotechnology journal},
year = {2026},
month = may,
volume = {35},
number = {1},
pages = {0076},
publisher = {AAAS Science Partner Journal Program},
issn = {2001-0370},
doi = {10.34133/csbj.0076},
url = {https://doi.org/10.34133/csbj.0076},
pmid = {42110217},
pmcid = {PMC13150068}
}

RIS

TY - JOUR
AU - Silva, Jose Cleydson Ferreira
AU - Richardson, Matthew
AU - Cediel-Becerra, José D. D.
AU - Schuster, Layla
AU - Dias, Raquel
TI - SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling
T2 - Computational and structural biotechnology journal
J2 - Comput Struct Biotechnol J
PY - 2026
DA - 2026/05/07
VL - 35
IS - 1
SP - 0076
SN - 2001-0370
PB - AAAS Science Partner Journal Program
DO - 10.34133/csbj.0076
UR - https://doi.org/10.34133/csbj.0076
LA - en
ER -

CSL-JSON

{
"id": "10.34133/csbj.0076",
"type": "article-journal",
"title": "SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling",
"container-title": "Computational and structural biotechnology journal",
"author": [
{
"family": "Silva",
"given": "Jose Cleydson Ferreira"
},
{
"family": "Richardson",
"given": "Matthew"
},
{
"family": "Cediel-Becerra",
"given": "José D. D."
},
{
"family": "Schuster",
"given": "Layla"
},
{
"family": "Dias",
"given": "Raquel"
}
],
"container-title-short": "Comput Struct Biotechnol J",
"volume": "35",
"issue": "1",
"page": "0076",
"DOI": "10.34133/csbj.0076",
"PMID": "42110217",
"PMCID": "PMC13150068",
"ISSN": "2001-0370",
"publisher": "AAAS Science Partner Journal Program",
"URL": "https://doi.org/10.34133/csbj.0076",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
7
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.molcel.2026.07.006 [code]
DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.
Journal: Molecular cell
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular, 2 references
[2] doi:10.1038/s41467-026-75444-4 [code]
Structural insights enable drug discovery for the neuronal NBCn2 carbonate transporter.
Journal: Nature communications
In common: MDAnalysis, Biopython, pandas, 2 other tools, cellular / molecular
[3] doi:10.1002/advs.202523984 [code]
INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: Biopython, PyTorch, scikit-learn, 3 other tools, 1 reference
[4] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular, 1 reference
[5] doi:10.1038/s41586-026-10658-6 [code]
An AI system to help scientists write expert-level empirical software.
Journal: Nature
In common: PyTorch, scikit-learn, pandas, 2 other tools, 2 references
[6] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular
[7] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits &lt;i&gt;Porphyromonas gingivalis&lt;/i&gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular
[8] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular
[9] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular
[10] doi:10.1093/bib/bbag182 [code]
NyxBind: enhancing deep neural representations for transcription factor binding site prediction via contrastive learning.
Journal: Briefings in bioinformatics
In common: Biopython, PyTorch, scikit-learn, 3 other tools, cellular / molecular

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.