OSCR

Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.

Code ↔ Paper

19 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 19 matches
  1. [1] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Preprocessing and Pharmacophore Feature Extraction ↔ pharmacophore_modeling.py, lines 223–255 · score 0.95 · aliphatic rings, aromatic rings, rdMolDescriptors, MolLogP, MolWt, NumRotatableBonds
  2. [2] § Methods › AI-Assisted Literature Search for Target Identification ↔ language model.py, lines 18–67 · score 0.91 · FLAN T5 XXL, L6 v2, MiniLM, prompts, FAISS, Sentence
  3. [3] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 223–255 · score 0.86 · MolLogP, MolWt, NumRotatableBonds, partial charges, Gasteiger, HBA
  4. [4] § Methods › Hybrid Model Architecture › 3D CNN for Voxelized Spatial Encoding ↔ Binding free energy model.py, lines 300–341 · score 0.81 · BatchNorm3d, Conv3d, ReLU, kernel, CNN, flattened
  5. [5] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 177–221 · score 0.76 · AllChem.EmbedMolecule, Chem.AddHs, random seed, graph, optimization, pharmacophoric
  6. [6] § Methods › Model Training and Evaluation ↔ Binding free energy model.py, lines 462–483 · score 0.76 · AdamW, hybrid model, binding free energies, training, batch, Optimization
  7. [7] § Methods › AI-Assisted Literature Search for Target Identification ↔ scoring assistant.py, lines 203–219 · score 0.76 · L6 v2, MiniLM, Sentence, paraphrase, semantic, embedder
  8. [8] § Methods › Deep Neural Network for Binding Affinity Prediction › Deep Graph Neural Network Model ↔ pharmacophore_modeling.py, lines 270–308 · score 0.74 · LeakyReLU, multihead attention, layer normalization, PyTorch, modules, GNN
  9. [9] § Methods › Deep Neural Network for Binding Affinity Prediction › Protein Structure Featurization ↔ pharmacophore_modeling.py, lines 154–175 · score 0.73 · ANM modes, surface points, active site, volume, dummy, hydrophobicity
  10. [10] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ prediction.py, lines 33–100 · score 0.72 · standard deviation, Boltzmann weighted, binding free energy, MD trajectory, Delta, kcal
  11. [11] § Methods › Deep Neural Network for Binding Affinity Prediction ↔ pharmacophore_modeling.py, lines 348–451 · score 0.67 · rotatable bonds, entropy loss, SDF, Batch, embedding, centroid
  12. [12] § Results › Binding Free Energy and Dissociation Constant ↔ prediction.py, lines 33–100 · score 0.66 · Boltzmann weighted binding, standard deviation, binding free energies, kcal, predicted, frame
  13. [13] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 177–221 · score 0.65 · edge_index, surface point, PMI, ESP, graph, bonds
  14. [14] § Results › Model Evaluation ↔ Binding free energy model.py, lines 343–405 · score 0.63 · MD trajectories processed, binding free energies, protein ligand, voxelized, solvation, fingerprints
  15. [15] § Results › Generative AI and PPI Framework for Glioma Target Identification ↔ language model.py, lines 18–67 · score 0.60 · flan t5 xxl, pretrained, google, database, transformer, glioma
  16. [16] § Methods › Deep Neural Network for Binding Affinity Prediction › Deep Graph Neural Network Model ↔ pharmacophore_modeling.py, lines 270–308 · score 0.60 · cross attention, multihead attention, layer, heads, ligand, protein
  17. [17] § Methods › Deep Neural Network for Binding Affinity Prediction › Protein Structure Featurization ↔ pharmacophore_modeling.py, lines 506–600 · score 0.55 · distance matrix, edge_index, graph, protein
  18. [18] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ Binding free energy model.py, lines 343–405 · score 0.54 · binding free energies, protein ligand, slicing, voxelization, featurization, topology
  19. [19] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ Binding free energy model.py, lines 232–264 · score 0.53 · Lennard Jones, binding free energies, ProLIF, RDKit, electrostatic, Interaction

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 694 lines · 27 KB · no license · 9 matches

  1. import os
  2. import sys
  3. import gc
  4. import numpy as np
  5. import pandas as pd
  6. import torch
  7. import torch.nn as nn
  8. from rdkit import Chem
  9. from rdkit.Chem import AllChem, Descriptors, rdMolDescriptors
  10. from prody import *
  11. from torch_geometric.data import Data, Batch
  12. from torch_geometric.nn import GATv2Conv, global_mean_pool
  13. from sklearn.model_selection import KFold
  14. from sklearn.metrics import r2_score
  15. from sklearn.preprocessing import StandardScaler
  16. from scipy.spatial.distance import cdist
  17. from tqdm import tqdm
  18. import warnings
  19. import datetime
  20. import psutil
  21. import traceback
  22. import pickle
  23. import json
  24. from contextlib import contextmanager
  25. from rdkit import RDLogger
  26. RDLogger.DisableLog('rdApp.*')
  27. warnings.filterwarnings("ignore")
  28. # Configuration
  29. PDB_FILE = os.path.normpath(os.path.join(os.getcwd(), "model_01_5AE1.pdb"))
  30. LIGAND_FOLDER = os.path.normpath("Ligands")
  31. ACTIVE_SITE_RESIDUES = [616,617]
  32. CHAIN_ID = "A"
  33. R = 1.987 # Gas constant in cal/(mol?K)
  34. # Checkpointing configuration
  35. CHECKPOINT_DIR = "checkpoints"
  36. DESCRIPTOR_CHECKPOINT_SIZE = 5000 # Save descriptor checkpoint every N molecules
  37. AFFINITY_CHECKPOINT_SIZE = 100 # Save affinity checkpoint every N molecules
  38. # GPU Configuration
  39. def setup_device():
  40. if torch.cuda.is_available():
  41. device = torch.device('cuda')
  42. print(f"Using GPU: {torch.cuda.get_device_name(0)}")
  43. print(f"GPU Memory: {torch.cuda.get_device_properties(0).total_memory / 1024**3:.2f} GB")
  44. # Set more conservative memory fraction
  45. torch.cuda.set_per_process_memory_fraction(0.6)
  46. else:
  47. device = torch.device('cpu')
  48. print("Using CPU - GPU not available")
  49. return device
  50. DEVICE = setup_device()
  51. # Create checkpoint directory
  52. os.makedirs(CHECKPOINT_DIR, exist_ok=True)
  53. @contextmanager
  54. def memory_cleanup():
  55. """Context manager for aggressive memory cleanup"""
  56. try:
  57. yield
  58. finally:
  59. gc.collect()
  60. if torch.cuda.is_available():
  61. torch.cuda.empty_cache()
  62. torch.cuda.synchronize()
  63. def log_memory_usage(stage=""):
  64. """Log current memory usage"""
  65. process = psutil.Process(os.getpid())
  66. cpu_mem = process.memory_info().rss / 1024 / 1024
  67. if torch.cuda.is_available():
  68. gpu_mem = torch.cuda.memory_allocated() / 1024**3
  69. gpu_cached = torch.cuda.memory_reserved() / 1024**3
  70. print(f"{stage} - CPU: {cpu_mem:.2f} MB, GPU: {gpu_mem:.2f} GB (cached: {gpu_cached:.2f} GB)")
  71. else:
  72. print(f"{stage} - CPU: {cpu_mem:.2f} MB")
  73. def save_checkpoint(data, filename):
  74. """Save checkpoint data"""
  75. try:
  76. filepath = os.path.join(CHECKPOINT_DIR, filename)
  77. if isinstance(data, pd.DataFrame):
  78. # Clean dataframe before saving
  79. data_clean = data.copy()
  80. if 'MolObj' in data_clean.columns:
  81. data_clean = data_clean.drop('MolObj', axis=1)
  82. data_clean = data_clean.replace([np.inf, -np.inf], np.nan)
  83. data_clean.to_excel(filepath, index=False)
  84. else:
  85. with open(filepath, 'wb') as f:
  86. pickle.dump(data, f)
  87. print(f"💾 Checkpoint saved: {filename}")
  88. return True
  89. except Exception as e:
  90. print(f"❌ Failed to save checkpoint {filename}: {e}")
  91. return False
  92. def load_checkpoint(filename):
  93. """Load checkpoint data"""
  94. try:
  95. filepath = os.path.join(CHECKPOINT_DIR, filename)
  96. if not os.path.exists(filepath):
  97. return None
  98. if filename.endswith('.xlsx'):
  99. return pd.read_excel(filepath)
  100. else:
  101. with open(filepath, 'rb') as f:
  102. return pickle.load(f)
  103. except Exception as e:
  104. print(f"❌ Failed to load checkpoint {filename}: {e}")
  105. return None
  106. def get_processed_files():
  107. """Get list of already processed files"""
  108. progress_file = os.path.join(CHECKPOINT_DIR, "processed_files.json")
  109. if os.path.exists(progress_file):
  110. try:
  111. with open(progress_file, 'r') as f:
  112. return json.load(f)
  113. except:
  114. return {}
  115. return {}
  116. def save_processed_files(processed_files):
  117. """Save list of processed files"""
  118. progress_file = os.path.join(CHECKPOINT_DIR, "processed_files.json")
  119. try:
  120. with open(progress_file, 'w') as f:
  121. json.dump(processed_files, f, indent=2)
  122. except Exception as e:
  123. print(f"❌ Failed to save progress: {e}")
  124. def calcAtomicFeatures(active_site):
  125. return np.zeros((active_site.numAtoms(), 10))
  126. def calcHydrophobicity(active_site):
  127. return np.zeros(active_site.numAtoms())
  128. def calcVolume(active_site):
  129. return 0.0
  130. class Electrostatics:
  131. def __init__(self, active_site):
  132. self.potentials = np.zeros(active_site.numAtoms())
  133. def getPotentials(self):
  134. return self.potentials
  135. def process_protein(pdb_file, active_site_residues):
  136. try:
  137. protein = parsePDB(pdb_file, fetch=False)
  138. sel_str = f'chain {CHAIN_ID} and resnum {" ".join(map(str, active_site_residues))}'
  139. active_site = protein.select(sel_str)
  140. if active_site is None or active_site.numAtoms() == 0:
  141. raise ValueError(f"Invalid active site selection: {sel_str}")
  142. coords = active_site.getCoords()
  143. centroid = np.mean(coords, axis=0)
  144. dummy_modes = np.zeros((active_site.numAtoms() * 3, 3))
  145. es = Electrostatics(active_site)
  146. return {
  147. 'residue_features': calcAtomicFeatures(active_site),
  148. 'anm_modes': dummy_modes,
  149. 'electrostatic': es.getPotentials(),
  150. 'hydrophobicity': calcHydrophobicity(active_site),
  151. 'volume': calcVolume(active_site),
  152. 'surface_points': coords,
  153. 'centroid': centroid
  154. }
  155. except Exception as e:
  156. raise RuntimeError(f"Failed to process protein: {e}")
  157. class LigandProcessor:
  158. def __init__(self, mol):
  159. self.mol = Chem.AddHs(mol)
  160. AllChem.EmbedMolecule(self.mol, randomSeed=42)
  161. AllChem.MMFFOptimizeMolecule(self.mol)
  162. def get_3dpharma_features(self):
  163. try:
  164. return {
  165. 'esp': 0.0,
  166. 'pmi': Descriptors.PMI1(self.mol) if hasattr(Descriptors, 'PMI1') else 0.0,
  167. 'steric': Descriptors.NPR1(self.mol) if hasattr(Descriptors, 'NPR1') else 0.0,
  168. 'affinity': 0.0
  169. }
  170. except:
  171. return {'esp': 0.0, 'pmi': 0.0, 'steric': 0.0, 'affinity': 0.0}
  172. def get_graph_representation(self):
  173. pos = torch.tensor([list(self.mol.GetConformer().GetAtomPosition(i))
  174. for i in range(self.mol.GetNumAtoms())], dtype=torch.float)
  175. features = []
  176. for atom in self.mol.GetAtoms():
  177. atom_features = [
  178. atom.GetAtomicNum(),
  179. atom.GetDegree(),
  180. atom.GetFormalCharge(),
  181. float(atom.GetHybridization().real),
  182. float(atom.GetIsAromatic())
  183. ]
  184. features.append(atom_features)
  185. x = torch.tensor(features, dtype=torch.float)
  186. bonds = list(self.mol.GetBonds())
  187. if len(bonds) > 0:
  188. edge_index = torch.tensor(
  189. [[b.GetBeginAtomIdx(), b.GetEndAtomIdx()]
  190. for b in bonds] +
  191. [[b.GetEndAtomIdx(), b.GetBeginAtomIdx()]
  192. for b in bonds], dtype=torch.long).t().contiguous()
  193. else:
  194. edge_index = torch.empty((2,0), dtype=torch.long)
  195. return Data(x=x, edge_index=edge_index, pos=pos, batch=torch.zeros(x.shape[0], dtype=torch.long))
  196. def get_surface_points(self):
  197. conf = self.mol.GetConformer()
  198. return np.array([conf.GetAtomPosition(i) for i in range(self.mol.GetNumAtoms())])
  199. def compute_descriptors(mol):
  200. try:
  201. AllChem.ComputeGasteigerCharges(mol)
  202. charges = []
  203. for atom in mol.GetAtoms():
  204. if atom.HasProp('_GasteigerCharge'):
  205. charges.append(float(atom.GetDoubleProp('_GasteigerCharge')))
  206. return {
  207. 'MolWt': Descriptors.MolWt(mol),
  208. 'MolLogP': Descriptors.MolLogP(mol),
  209. 'NumRotatableBonds': Descriptors.NumRotatableBonds(mol),
  210. 'TPSA': Descriptors.TPSA(mol),
  211. 'HBA': rdMolDescriptors.CalcNumHBA(mol),
  212. 'HBD': rdMolDescriptors.CalcNumHBD(mol),
  213. 'AromaticRings': rdMolDescriptors.CalcNumAromaticRings(mol),
  214. 'FractionCSP3': rdMolDescriptors.CalcFractionCSP3(mol),
  215. 'HeavyAtomCount': Descriptors.HeavyAtomCount(mol),
  216. 'AliphaticRings': rdMolDescriptors.CalcNumAliphaticRings(mol),
  217. 'MaxPartialCharge': max(charges) if charges else 0.0,
  218. 'MinPartialCharge': min(charges) if charges else 0.0,
  219. 'LabuteASA': rdMolDescriptors.CalcLabuteASA(mol),
  220. 'MolMR': Descriptors.MolMR(mol),
  221. 'BalabanJ': Descriptors.BalabanJ(mol),
  222. 'BertzCT': Descriptors.BertzCT(mol)
  223. }
  224. except:
  225. return {key: 0.0 for key in [
  226. 'MolWt', 'MolLogP', 'NumRotatableBonds', 'TPSA', 'HBA', 'HBD',
  227. 'AromaticRings', 'FractionCSP3', 'HeavyAtomCount', 'AliphaticRings',
  228. 'MaxPartialCharge', 'MinPartialCharge', 'LabuteASA', 'MolMR',
  229. 'BalabanJ', 'BertzCT'
  230. ]}
  231. def compute_entropy(rot_bonds):
  232. delta_s_torsion = -R * np.log(rot_bonds + 1)
  233. return delta_s_torsion
  234. def compute_distance_to_site(mol, centroid):
  235. conf = mol.GetConformer()
  236. distances = []
  237. for atom in mol.GetAtoms():
  238. pos = np.array(conf.GetAtomPosition(atom.GetIdx()))
  239. dist = np.linalg.norm(pos - centroid)
  240. distances.append(dist)
  241. return np.mean(distances), np.min(distances), np.max(distances)
  242. class PharmaGNN(nn.Module):
  243. def __init__(self, protein_feature_size=10, ligand_feature_size=5):
  244. super().__init__()
  245. self.device = DEVICE
  246. self.protein_conv1 = GATv2Conv(protein_feature_size, 128, heads=2) # Reduced complexity
  247. self.protein_conv2 = GATv2Conv(128*2, 256)
  248. self.ligand_conv1 = GATv2Conv(ligand_feature_size, 128, heads=2)
  249. self.ligand_conv2 = GATv2Conv(128*2, 256)
  250. self.cross_attention = nn.MultiheadAttention(256, 4, batch_first=True) # Reduced heads
  251. self.fc = nn.Sequential(
  252. nn.Linear(512, 256),
  253. nn.LayerNorm(256),
  254. nn.LeakyReLU(),
  255. nn.Dropout(0.2),
  256. nn.Linear(256, 1)
  257. )
  258. self.to(self.device)
  259. def forward(self, protein_data, ligand_data):
  260. # Move data to device
  261. protein_data = protein_data.to(self.device)
  262. ligand_data = ligand_data.to(self.device)
  263. # Process protein
  264. p_x = self.protein_conv1(protein_data.x, protein_data.edge_index)
  265. p_x = self.protein_conv2(p_x, protein_data.edge_index)
  266. p_x = global_mean_pool(p_x, protein_data.batch)
  267. # Process ligand
  268. l_x = self.ligand_conv1(ligand_data.x, ligand_data.edge_index)
  269. l_x = self.ligand_conv2(l_x, ligand_data.edge_index)
  270. l_x = global_mean_pool(l_x, ligand_data.batch)
  271. # Cross attention
  272. attn_out, _ = self.cross_attention(p_x.unsqueeze(1), l_x.unsqueeze(1), l_x.unsqueeze(1))
  273. # Combine features
  274. combined = torch.cat([p_x, attn_out.squeeze(1)], dim=1)
  275. return self.fc(combined)
  276. def calculate_binding_score(protein_features, ligand_features, descriptors, distances, entropy):
  277. try:
  278. sc = 1 - cdist(protein_features['surface_points'], ligand_features['surface_points'], 'cosine').mean()
  279. ec = np.corrcoef(protein_features['electrostatic'], [ligand_features['esp']])[0,1]
  280. if np.isnan(ec):
  281. ec = 0.0
  282. dist_score = 1 / (1 + distances['AvgDistToSite'])
  283. entropy_score = -entropy / R
  284. energy = (0.3 * sc) + (0.2 * ec) + (0.2 * ligand_features['affinity']) + \
  285. (0.2 * dist_score) + (0.1 * entropy_score)
  286. return energy
  287. except:
  288. return 0.0
  289. def safe_save_dataframe(df, filename, max_retries=3):
  290. """Safely save DataFrame with retries"""
  291. for attempt in range(max_retries):
  292. try:
  293. df_clean = df.copy()
  294. if 'MolObj' in df_clean.columns:
  295. df_clean = df_clean.drop('MolObj', axis=1)
  296. df_clean = df_clean.replace([np.inf, -np.inf], np.nan)
  297. with pd.ExcelWriter(filename, engine='openpyxl') as writer:
  298. df_clean.to_excel(writer, index=False, sheet_name='Results')
  299. print(f"✅ Successfully saved {filename}")
  300. return True
  301. except Exception as e:
  302. print(f"❌ Attempt {attempt + 1} failed to save {filename}: {e}")
  303. if attempt < max_retries - 1:
  304. import time
  305. time.sleep(1)
  306. else:
  307. print(f"❌ Failed to save {filename} after {max_retries} attempts")
  308. return False
  309. def process_single_sdf_file(sdf_path, sdf_file, centroid, processed_files):
  310. """Process a single SDF file with continuous checkpointing"""
  311. results = []
  312. skipped = []
  313. # Check if file already processed
  314. if sdf_file in processed_files:
  315. print(f"📁 Loading cached results for {sdf_file}...")
  316. cached_file = f"descriptors_{sdf_file.replace('.sdf', '')}.xlsx"
  317. cached_df = load_checkpoint(cached_file)
  318. if cached_df is not None:
  319. print(f"✅ Loaded {len(cached_df)} cached results for {sdf_file}")
  320. return cached_df, []
  321. print(f"🔄 Processing {sdf_file}...")
  322. try:
  323. suppl = Chem.SDMolSupplier(sdf_path, sanitize=False)
  324. batch_results = []
  325. processed_count = 0
  326. total_processed = 0
  327. for idx, mol in enumerate(suppl):
  328. if idx % 1000 == 0:
  329. print(f"\r{sdf_file}: Processing molecule {idx + 1}...", end="", flush=True)
  330. if mol is None:
  331. skipped.append((sdf_file, idx, "Unreadable molecule"))
  332. continue
  333. try:
  334. # Sanitize and prepare molecule
  335. Chem.SanitizeMol(mol)
  336. mol = Chem.AddHs(mol)
  337. res = AllChem.EmbedMolecule(mol, randomSeed=42)
  338. if res != 0:
  339. skipped.append((sdf_file, idx, "3D embedding failed"))
  340. continue
  341. AllChem.MMFFOptimizeMolecule(mol)
  342. # Compute descriptors
  343. desc = compute_descriptors(mol)
  344. avg_d, min_d, max_d = compute_distance_to_site(mol, centroid)
  345. entropy = compute_entropy(desc['NumRotatableBonds'])
  346. desc.update({
  347. 'Ligand': f"{sdf_file}_mol_{idx}",
  348. 'AvgDistToSite': avg_d,
  349. 'MinDistToSite': min_d,
  350. 'MaxDistToSite': max_d,
  351. 'EntropyLoss': entropy,
  352. 'SourceFile': sdf_file,
  353. 'MolIndex': idx
  354. })
  355. batch_results.append(desc)
  356. processed_count += 1
  357. total_processed += 1
  358. # Save checkpoint every N molecules
  359. if processed_count >= DESCRIPTOR_CHECKPOINT_SIZE:
  360. results.extend(batch_results)
  361. # Save intermediate checkpoint
  362. checkpoint_df = pd.DataFrame(results)
  363. checkpoint_file = f"descriptors_{sdf_file.replace('.sdf', '')}_temp.xlsx"
  364. save_checkpoint(checkpoint_df, checkpoint_file)
  365. batch_results = []
  366. processed_count = 0
  367. # Memory cleanup
  368. with memory_cleanup():
  369. pass
  370. except Exception as e:
  371. skipped.append((sdf_file, idx, f"Processing error: {str(e)[:100]}"))
  372. continue
  373. # Add remaining results
  374. if batch_results:
  375. results.extend(batch_results)
  376. # Save final results for this file
  377. if results:
  378. final_df = pd.DataFrame(results)
  379. final_file = f"descriptors_{sdf_file.replace('.sdf', '')}.xlsx"
  380. save_checkpoint(final_df, final_file)
  381. # Mark file as processed
  382. processed_files[sdf_file] = {
  383. 'processed_count': total_processed,
  384. 'timestamp': datetime.datetime.now().isoformat()
  385. }
  386. save_processed_files(processed_files)
  387. print(f"\r{sdf_file}: Completed - {total_processed} molecules processed")
  388. except Exception as e:
  389. print(f"\n❌ Error processing {sdf_file}: {e}")
  390. return pd.DataFrame(), skipped
  391. return pd.DataFrame(results), skipped
  392. def process_ligands_from_multiple_sdfs(folder, centroid):
  393. """Process ligands with file-level checkpointing"""
  394. all_results = []
  395. all_skipped = []
  396. sdf_files = [f for f in os.listdir(folder) if f.endswith(".sdf")]
  397. processed_files = get_processed_files()
  398. print(f"📂 Found {len(sdf_files)} SDF files")
  399. print(f"📋 {len(processed_files)} files already processed")
  400. for file_idx, sdf_file in enumerate(sdf_files):
  401. print(f"\n🔄 File {file_idx + 1}/{len(sdf_files)}: {sdf_file}")
  402. sdf_path = os.path.join(folder, sdf_file)
  403. with memory_cleanup():
  404. results_df, skipped = process_single_sdf_file(sdf_path, sdf_file, centroid, processed_files)
  405. if not results_df.empty:
  406. all_results.append(results_df)
  407. all_skipped.extend(skipped)
  408. log_memory_usage(f"After processing {sdf_file}")
  409. # Combine all results
  410. if all_results:
  411. final_df = pd.concat(all_results, ignore_index=True)
  412. print(f"\n✅ Total molecules processed: {len(final_df)}")
  413. else:
  414. final_df = pd.DataFrame()
  415. print(f"\n❌ No molecules processed successfully")
  416. if all_skipped:
  417. print(f"⚠️ Skipped {len(all_skipped)} molecules. First 10 errors:")
  418. for skip_info in all_skipped[:10]:
  419. print(f" {skip_info[0]}_mol_{skip_info[1]}: {skip_info[2]}")
  420. return final_df
  421. class PharmaPipeline:
  422. def __init__(self):
  423. print("🔬 Initializing PharmaPipeline...")
  424. log_memory_usage("Initial")
  425. self.protein_features = process_protein(PDB_FILE, ACTIVE_SITE_RESIDUES)
  426. self.model = PharmaGNN()
  427. self.model.eval()
  428. self.scaler = StandardScaler()
  429. log_memory_usage("After initialization")
  430. def process_affinity_predictions(self, ligand_df):
  431. """Process affinity predictions with continuous checkpointing"""
  432. results = []
  433. # Prepare protein data once
  434. coords = self.protein_features['surface_points']
  435. dist_matrix = cdist(coords, coords)
  436. edge_list = np.where(dist_matrix < 5.0)
  437. edge_index = torch.tensor([edge_list[0], edge_list[1]], dtype=torch.long).contiguous()
  438. protein_data = Data(
  439. x=torch.tensor(self.protein_features['residue_features'], dtype=torch.float),
  440. edge_index=edge_index,
  441. batch=torch.zeros(len(self.protein_features['residue_features']), dtype=torch.long)
  442. )
  443. print(f"🧬 Predicting affinities for {len(ligand_df)} ligands...")
  444. # Process in smaller batches
  445. batch_size = AFFINITY_CHECKPOINT_SIZE
  446. num_batches = (len(ligand_df) - 1) // batch_size + 1
  447. for batch_idx in range(num_batches):
  448. batch_start = batch_idx * batch_size
  449. batch_end = min(batch_start + batch_size, len(ligand_df))
  450. print(f"\n🔄 Affinity batch {batch_idx + 1}/{num_batches} ({batch_start + 1}-{batch_end})")
  451. batch_results = []
  452. with memory_cleanup():
  453. for idx in range(batch_start, batch_end):
  454. row = ligand_df.iloc[idx]
  455. if (idx - batch_start) % 20 == 0:
  456. progress = f"Progress: {idx - batch_start + 1}/{batch_end - batch_start}"
  457. print(f"\r{progress}", end="", flush=True)
  458. try:
  459. # Create molecule from SMILES or use stored mol object
  460. mol_data = row.to_dict()
  461. # Create a dummy molecule for testing (replace with actual mol reconstruction)
  462. mol = Chem.MolFromSmiles('CCO') # Placeholder
  463. if mol is None:
  464. continue
  465. with torch.no_grad():
  466. lp = LigandProcessor(mol)
  467. pharma_features = lp.get_3dpharma_features()
  468. graph = lp.get_graph_representation()
  469. # Predict affinity
  470. affinity = self.model(protein_data, graph)
  471. pharma_features['affinity'] = affinity.item()
  472. # Calculate physical score
  473. phys_score = calculate_binding_score(
  474. self.protein_features,
  475. {'surface_points': lp.get_surface_points(), **pharma_features},
  476. mol_data,
  477. {'AvgDistToSite': row['AvgDistToSite']},
  478. row['EntropyLoss']
  479. )
  480. final_score = 0.7 * affinity.item() + 0.3 * phys_score
  481. result = {
  482. 'Drug': row['Ligand'],
  483. 'Affinity': affinity.item(),
  484. 'PhysScore': phys_score,
  485. 'FinalScore': final_score,
  486. **pharma_features,
  487. **mol_data
  488. }
  489. batch_results.append(result)
  490. except Exception as e:
  491. print(f"\n❌ Error predicting for {row['Ligand']}: {e}")
  492. continue
  493. # Save batch results
  494. results.extend(batch_results)
  495. if results:
  496. checkpoint_df = pd.DataFrame(results)
  497. timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
  498. checkpoint_file = f"affinity_results_batch_{batch_idx + 1}_{timestamp}.xlsx"
  499. save_checkpoint(checkpoint_df, checkpoint_file)
  500. print(f"\n💾 Saved {len(results)} affinity predictions")
  501. log_memory_usage(f"After affinity batch {batch_idx + 1}")
  502. return pd.DataFrame(results)
  503. def run(self):
  504. """Main pipeline execution with comprehensive checkpointing"""
  505. try:
  506. print("🚀 Starting PharmaPipeline execution...")
  507. # Validate inputs
  508. pdb_path = os.path.normpath(PDB_FILE)
  509. ligand_folder = os.path.normpath(LIGAND_FOLDER)
  510. if not os.path.exists(pdb_path):
  511. raise FileNotFoundError(f"PDB file {pdb_path} not found")
  512. if not os.path.exists(ligand_folder):
  513. raise FileNotFoundError(f"Ligand folder {ligand_folder} not found")
  514. # Step 1: Process ligand descriptors
  515. print("\n📊 Step 1: Processing ligand descriptors...")
  516. ligand_df = process_ligands_from_multiple_sdfs(ligand_folder, self.protein_features['centroid'])
  517. if ligand_df.empty:
  518. print("❌ No ligands processed successfully")
  519. return None
  520. # Step 2: Process affinity predictions
  521. print("\n🧬 Step 2: Processing affinity predictions...")
  522. results_df = self.process_affinity_predictions(ligand_df)
  523. if results_df.empty:
  524. print("❌ No affinity predictions completed")
  525. return None
  526. # Step 3: Final processing and saving
  527. print("\n📈 Step 3: Final processing...")
  528. results_df = results_df.sort_values('FinalScore', ascending=False)
  529. # Save final results
  530. timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
  531. output_file = f"Nature_Methods_Pharma_Ranking_{timestamp}.xlsx"
  532. success = safe_save_dataframe(results_df, output_file)
  533. if success:
  534. print(f"\n📄 Final output saved to: {output_file}")
  535. print(f"🏆 Top 5 compounds by FinalScore:")
  536. display_cols = ['Drug', 'FinalScore', 'Affinity', 'PhysScore']
  537. available_cols = [col for col in display_cols if col in results_df.columns]
  538. print(results_df[available_cols].head().to_string())
  539. return results_df
  540. except Exception as e:
  541. print(f"❌ Pipeline failed: {e}")
  542. print("🔍 Full traceback:")
  543. traceback.print_exc()
  544. return None
  545. finally:
  546. print("\n🧹 Performing final cleanup...")
  547. with memory_cleanup():
  548. pass
  549. log_memory_usage("Final cleanup")
  550. if __name__ == "__main__":
  551. try:
  552. log_memory_usage("Program start")
  553. print("🔬 Nature Methods Pharma Analysis Pipeline")
  554. print("=" * 50)
  555. pipeline = PharmaPipeline()
  556. result_df = pipeline.run()
  557. if result_df is not None:
  558. print("\n🎉 Nature Methods-ready analysis complete!")
  559. print(f"✅ Successfully processed {len(result_df)} compounds")
  560. else:
  561. print("\n❌ Analysis failed - check error messages above")
  562. log_memory_usage("Program end")
  563. except KeyboardInterrupt:
  564. print("\n⚠️ Process interrupted by user")
  565. print("💾 Checkpoint files have been saved - you can resume processing")
  566. sys.exit(1)
  567. except Exception as e:
  568. print(f"\n💥 Unexpected error: {e}")
  569. print("🔍 Full traceback:")
  570. traceback.print_exc()
  571. sys.exit(1)
  572. finally:
  573. print("\n🧹 Final cleanup...")
  574. with memory_cleanup():
  575. pass
  576. print("🏁 Program finished.")

pharmacophore_modeling.py at commit 9e48132, no license · at the source

Overview

Authors: Amritha Thaikkad1, Sonet Daniel Thomas1,2, Leona Dcunha1, Levin John1, Jishna Vijayan1, Naveen Joy1, Radul R Dev1, Rajesh Raju1, Abhithaj Jayanandan1
  1. Centre for Integrative Omics Data Science (CIODS), Yenepoya (Deemed to be University), Mangalore 575018, Karnataka, India
  2. Center for Systems Biology and Molecular Medicine (CSBMM), Yenepoya (Deemed to be University), Mangalore 575018, Karnataka, India
Institutions: Yenepoya University (India)
Journal: ACS omega, volume 11, issue 11, pages 17249-17265
Dates: received 9 September 2025; accepted 29 January 2026; published online 11 March 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1021/acsomega.5c09368 · PMID 41908407 · PMCID PMC13019191 · OpenAlex W7134966412
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: other condition (population)
Methods: Smoothing, state filtering, decompositions, Machine learning, Connectivity
Topic: Biological Research and Disease Studies (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: cited by 1 paper (Europe PMC); 46 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.

naveen-joy-18/Feature-driven-AI-Guided-Drug-Repurposing-Pipeline-for-AGPS-Targeted-Glioma-Therapy-

License: none: the authors keep all their rights
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 9e481328d5f3663cf67336f5a1a036a791a8def3, 18 December 2025
Languages: Python (5)
Size: 8 files, 5 scripts
Software Heritage: not archived
Found in: “Data Availability Statement”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (5 files), PyTorch (5 files), pandas (4 files), Matplotlib (2 files), RDKit (2 files), seaborn (2 files), Hugging Face Transformers (2 files), Biopython (1 file), Plotly (1 file), PyTorch Geometric (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
6 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1021/acsomega.5c09368.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 44 references.

Cite

This paper

Thaikkad, A., Thomas, S. D., Dcunha, L., John, L., Vijayan, J., Joy, N., R Dev, R., Raju, R., & Jayanandan, A. (2026). Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening. ACS omega, 11(11), 17249-17265. https://doi.org/10.1021/acsomega.5c09368

BibTeX

@article{thaikkad2026structure,
author = {Thaikkad, Amritha and Thomas, Sonet Daniel and Dcunha, Leona and John, Levin and Vijayan, Jishna and Joy, Naveen and R Dev, Radul and Raju, Rajesh and Jayanandan, Abhithaj},
title = {{Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening}},
journal = {ACS omega},
year = {2026},
month = mar,
volume = {11},
number = {11},
pages = {17249--17265},
publisher = {American Chemical Society},
issn = {2470-1343},
doi = {10.1021/acsomega.5c09368},
url = {https://doi.org/10.1021/acsomega.5c09368},
pmid = {41908407},
pmcid = {PMC13019191}
}

RIS

TY - JOUR
AU - Thaikkad, Amritha
AU - Thomas, Sonet Daniel
AU - Dcunha, Leona
AU - John, Levin
AU - Vijayan, Jishna
AU - Joy, Naveen
AU - R Dev, Radul
AU - Raju, Rajesh
AU - Jayanandan, Abhithaj
TI - Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening
T2 - ACS omega
J2 - ACS Omega
PY - 2026
DA - 2026/03/11
VL - 11
IS - 11
SP - 17249
EP - 17265
SN - 2470-1343
PB - American Chemical Society
DO - 10.1021/acsomega.5c09368
UR - https://doi.org/10.1021/acsomega.5c09368
LA - en
ER -

CSL-JSON

{
"id": "10.1021/acsomega.5c09368",
"type": "article-journal",
"title": "Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening",
"container-title": "ACS omega",
"author": [
{
"family": "Thaikkad",
"given": "Amritha"
},
{
"family": "Thomas",
"given": "Sonet Daniel"
},
{
"family": "Dcunha",
"given": "Leona"
},
{
"family": "John",
"given": "Levin"
},
{
"family": "Vijayan",
"given": "Jishna"
},
{
"family": "Joy",
"given": "Naveen"
},
{
"family": "R Dev",
"given": "Radul"
},
{
"family": "Raju",
"given": "Rajesh"
},
{
"family": "Jayanandan",
"given": "Abhithaj"
}
],
"container-title-short": "ACS Omega",
"volume": "11",
"issue": "11",
"page": "17249-17265",
"DOI": "10.1021/acsomega.5c09368",
"PMID": "41908407",
"PMCID": "PMC13019191",
"ISSN": "2470-1343",
"publisher": "American Chemical Society",
"URL": "https://doi.org/10.1021/acsomega.5c09368",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
11
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: RDKit, Biopython, Hugging Face Transformers, 8 other tools
[2] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: RDKit, Biopython, PyTorch Geometric, 6 other tools, other condition
[3] doi:10.1038/s41586-026-10670-w [code]
Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
Journal: Nature
In common: RDKit, PyTorch Geometric, Plotly, 7 other tools
[4] doi:10.1002/hbm.70469 [code]
VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.
Journal: Human brain mapping
In common: PyTorch Geometric, Hugging Face Transformers, Plotly, 7 other tools
[5] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits &lt;i&gt;Porphyromonas gingivalis&lt;/i&gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: RDKit, Biopython, PyTorch Geometric, 6 other tools
[6] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: RDKit, Biopython, PyTorch Geometric, 6 other tools
[7] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: RDKit, Biopython, PyTorch Geometric, 6 other tools
[8] doi:10.1371/journal.pone.0345854 [code]
Shedding light on neural learning to rank models for anticancer drug prioritization.
Journal: PloS one
In common: RDKit, PyTorch Geometric, PyTorch, 6 other tools, other condition
[9] doi:10.34133/csbj.0036 [code]
HYG-mol: An Interpretable Multimodal Hypergraph Framework for Molecular Property Prediction.
Journal: Computational and structural biotechnology journal
In common: RDKit, PyTorch Geometric, Hugging Face Transformers, 5 other tools
[10] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: RDKit, PyTorch Geometric, PyTorch, 6 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.