Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.
The 19 matches
- [1] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Preprocessing and Pharmacophore Feature Extraction ↔ pharmacophore_modeling.py, lines 223–255 · score 0.95 · aliphatic rings, aromatic rings, rdMolDescriptors, MolLogP, MolWt, NumRotatableBonds
- [2] § Methods › AI-Assisted Literature Search for Target Identification ↔ language model.py, lines 18–67 · score 0.91 · FLAN T5 XXL, L6 v2, MiniLM, prompts, FAISS, Sentence
- [3] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 223–255 · score 0.86 · MolLogP, MolWt, NumRotatableBonds, partial charges, Gasteiger, HBA
- [4] § Methods › Hybrid Model Architecture › 3D CNN for Voxelized Spatial Encoding ↔ Binding free energy model.py, lines 300–341 · score 0.81 · BatchNorm3d, Conv3d, ReLU, kernel, CNN, flattened
- [5] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 177–221 · score 0.76 · AllChem.EmbedMolecule, Chem.AddHs, random seed, graph, optimization, pharmacophoric
- [6] § Methods › Model Training and Evaluation ↔ Binding free energy model.py, lines 462–483 · score 0.76 · AdamW, hybrid model, binding free energies, training, batch, Optimization
- [7] § Methods › AI-Assisted Literature Search for Target Identification ↔ scoring assistant.py, lines 203–219 · score 0.76 · L6 v2, MiniLM, Sentence, paraphrase, semantic, embedder
- [8] § Methods › Deep Neural Network for Binding Affinity Prediction › Deep Graph Neural Network Model ↔ pharmacophore_modeling.py, lines 270–308 · score 0.74 · LeakyReLU, multihead attention, layer normalization, PyTorch, modules, GNN
- [9] § Methods › Deep Neural Network for Binding Affinity Prediction › Protein Structure Featurization ↔ pharmacophore_modeling.py, lines 154–175 · score 0.73 · ANM modes, surface points, active site, volume, dummy, hydrophobicity
- [10] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ prediction.py, lines 33–100 · score 0.72 · standard deviation, Boltzmann weighted, binding free energy, MD trajectory, Delta, kcal
- [11] § Methods › Deep Neural Network for Binding Affinity Prediction ↔ pharmacophore_modeling.py, lines 348–451 · score 0.67 · rotatable bonds, entropy loss, SDF, Batch, embedding, centroid
- [12] § Results › Binding Free Energy and Dissociation Constant ↔ prediction.py, lines 33–100 · score 0.66 · Boltzmann weighted binding, standard deviation, binding free energies, kcal, predicted, frame
- [13] § Methods › Deep Neural Network for Binding Affinity Prediction › Ligand Quantum and Graph Representation ↔ pharmacophore_modeling.py, lines 177–221 · score 0.65 · edge_index, surface point, PMI, ESP, graph, bonds
- [14] § Results › Model Evaluation ↔ Binding free energy model.py, lines 343–405 · score 0.63 · MD trajectories processed, binding free energies, protein ligand, voxelized, solvation, fingerprints
- [15] § Results › Generative AI and PPI Framework for Glioma Target Identification ↔ language model.py, lines 18–67 · score 0.60 · flan t5 xxl, pretrained, google, database, transformer, glioma
- [16] § Methods › Deep Neural Network for Binding Affinity Prediction › Deep Graph Neural Network Model ↔ pharmacophore_modeling.py, lines 270–308 · score 0.60 · cross attention, multihead attention, layer, heads, ligand, protein
- [17] § Methods › Deep Neural Network for Binding Affinity Prediction › Protein Structure Featurization ↔ pharmacophore_modeling.py, lines 506–600 · score 0.55 · distance matrix, edge_index, graph, protein
- [18] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ Binding free energy model.py, lines 343–405 · score 0.54 · binding free energies, protein ligand, slicing, voxelization, featurization, topology
- [19] § Results › Deep Interaction Featurization Using ProLIF and RDKit ↔ Binding free energy model.py, lines 232–264 · score 0.53 · Lennard Jones, binding free energies, ProLIF, RDKit, electrostatic, Interaction
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 694 lines · 27 KB · no license · 9 matches
- import os
- import sys
- import gc
- import numpy as np
- import pandas as pd
- import torch
- import torch.nn as nn
- from rdkit import Chem
- from rdkit.Chem import AllChem, Descriptors, rdMolDescriptors
- from prody import *
- from torch_geometric.data import Data, Batch
- from torch_geometric.nn import GATv2Conv, global_mean_pool
- from sklearn.model_selection import KFold
- from sklearn.metrics import r2_score
- from sklearn.preprocessing import StandardScaler
- from scipy.spatial.distance import cdist
- from tqdm import tqdm
- import warnings
- import datetime
- import psutil
- import traceback
- import pickle
- import json
- from contextlib import contextmanager
- from rdkit import RDLogger
- RDLogger.DisableLog('rdApp.*')
- warnings.filterwarnings("ignore")
- # Configuration
- PDB_FILE = os.path.normpath(os.path.join(os.getcwd(), "model_01_5AE1.pdb"))
- LIGAND_FOLDER = os.path.normpath("Ligands")
- ACTIVE_SITE_RESIDUES = [616,617]
- CHAIN_ID = "A"
- R = 1.987 # Gas constant in cal/(mol?K)
- # Checkpointing configuration
- CHECKPOINT_DIR = "checkpoints"
- DESCRIPTOR_CHECKPOINT_SIZE = 5000 # Save descriptor checkpoint every N molecules
- AFFINITY_CHECKPOINT_SIZE = 100 # Save affinity checkpoint every N molecules
- # GPU Configuration
- def setup_device():
- if torch.cuda.is_available():
- device = torch.device('cuda')
- print(f"Using GPU: {torch.cuda.get_device_name(0)}")
- print(f"GPU Memory: {torch.cuda.get_device_properties(0).total_memory / 1024**3:.2f} GB")
- # Set more conservative memory fraction
- torch.cuda.set_per_process_memory_fraction(0.6)
- else:
- device = torch.device('cpu')
- print("Using CPU - GPU not available")
- return device
- DEVICE = setup_device()
- # Create checkpoint directory
- os.makedirs(CHECKPOINT_DIR, exist_ok=True)
- @contextmanager
- def memory_cleanup():
- """Context manager for aggressive memory cleanup"""
- try:
- yield
- finally:
- gc.collect()
- if torch.cuda.is_available():
- torch.cuda.empty_cache()
- torch.cuda.synchronize()
- def log_memory_usage(stage=""):
- """Log current memory usage"""
- process = psutil.Process(os.getpid())
- cpu_mem = process.memory_info().rss / 1024 / 1024
- if torch.cuda.is_available():
- gpu_mem = torch.cuda.memory_allocated() / 1024**3
- gpu_cached = torch.cuda.memory_reserved() / 1024**3
- print(f"{stage} - CPU: {cpu_mem:.2f} MB, GPU: {gpu_mem:.2f} GB (cached: {gpu_cached:.2f} GB)")
- else:
- print(f"{stage} - CPU: {cpu_mem:.2f} MB")
- def save_checkpoint(data, filename):
- """Save checkpoint data"""
- try:
- filepath = os.path.join(CHECKPOINT_DIR, filename)
- if isinstance(data, pd.DataFrame):
- # Clean dataframe before saving
- data_clean = data.copy()
- if 'MolObj' in data_clean.columns:
- data_clean = data_clean.drop('MolObj', axis=1)
- data_clean = data_clean.replace([np.inf, -np.inf], np.nan)
- data_clean.to_excel(filepath, index=False)
- else:
- with open(filepath, 'wb') as f:
- pickle.dump(data, f)
- print(f"💾 Checkpoint saved: {filename}")
- return True
- except Exception as e:
- print(f"❌ Failed to save checkpoint {filename}: {e}")
- return False
- def load_checkpoint(filename):
- """Load checkpoint data"""
- try:
- filepath = os.path.join(CHECKPOINT_DIR, filename)
- if not os.path.exists(filepath):
- return None
- if filename.endswith('.xlsx'):
- return pd.read_excel(filepath)
- else:
- with open(filepath, 'rb') as f:
- return pickle.load(f)
- except Exception as e:
- print(f"❌ Failed to load checkpoint {filename}: {e}")
- return None
- def get_processed_files():
- """Get list of already processed files"""
- progress_file = os.path.join(CHECKPOINT_DIR, "processed_files.json")
- if os.path.exists(progress_file):
- try:
- with open(progress_file, 'r') as f:
- return json.load(f)
- except:
- return {}
- return {}
- def save_processed_files(processed_files):
- """Save list of processed files"""
- progress_file = os.path.join(CHECKPOINT_DIR, "processed_files.json")
- try:
- with open(progress_file, 'w') as f:
- json.dump(processed_files, f, indent=2)
- except Exception as e:
- print(f"❌ Failed to save progress: {e}")
- def calcAtomicFeatures(active_site):
- return np.zeros((active_site.numAtoms(), 10))
- def calcHydrophobicity(active_site):
- return np.zeros(active_site.numAtoms())
- def calcVolume(active_site):
- return 0.0
- class Electrostatics:
- def __init__(self, active_site):
- self.potentials = np.zeros(active_site.numAtoms())
- def getPotentials(self):
- return self.potentials
- def process_protein(pdb_file, active_site_residues):
- try:
- protein = parsePDB(pdb_file, fetch=False)
- sel_str = f'chain {CHAIN_ID} and resnum {" ".join(map(str, active_site_residues))}'
- active_site = protein.select(sel_str)
- if active_site is None or active_site.numAtoms() == 0:
- raise ValueError(f"Invalid active site selection: {sel_str}")
- coords = active_site.getCoords()
- centroid = np.mean(coords, axis=0)
- dummy_modes = np.zeros((active_site.numAtoms() * 3, 3))
- es = Electrostatics(active_site)
- return {
- 'residue_features': calcAtomicFeatures(active_site),
- 'anm_modes': dummy_modes,
- 'electrostatic': es.getPotentials(),
- 'hydrophobicity': calcHydrophobicity(active_site),
- 'volume': calcVolume(active_site),
- 'surface_points': coords,
- 'centroid': centroid
- }
- except Exception as e:
- raise RuntimeError(f"Failed to process protein: {e}")
- class LigandProcessor:
- def __init__(self, mol):
- self.mol = Chem.AddHs(mol)
- AllChem.EmbedMolecule(self.mol, randomSeed=42)
- AllChem.MMFFOptimizeMolecule(self.mol)
- def get_3dpharma_features(self):
- try:
- return {
- 'esp': 0.0,
- 'pmi': Descriptors.PMI1(self.mol) if hasattr(Descriptors, 'PMI1') else 0.0,
- 'steric': Descriptors.NPR1(self.mol) if hasattr(Descriptors, 'NPR1') else 0.0,
- 'affinity': 0.0
- }
- except:
- return {'esp': 0.0, 'pmi': 0.0, 'steric': 0.0, 'affinity': 0.0}
- def get_graph_representation(self):
- pos = torch.tensor([list(self.mol.GetConformer().GetAtomPosition(i))
- for i in range(self.mol.GetNumAtoms())], dtype=torch.float)
- features = []
- for atom in self.mol.GetAtoms():
- atom_features = [
- atom.GetAtomicNum(),
- atom.GetDegree(),
- atom.GetFormalCharge(),
- float(atom.GetHybridization().real),
- float(atom.GetIsAromatic())
- ]
- features.append(atom_features)
- x = torch.tensor(features, dtype=torch.float)
- bonds = list(self.mol.GetBonds())
- if len(bonds) > 0:
- edge_index = torch.tensor(
- [[b.GetBeginAtomIdx(), b.GetEndAtomIdx()]
- for b in bonds] +
- [[b.GetEndAtomIdx(), b.GetBeginAtomIdx()]
- for b in bonds], dtype=torch.long).t().contiguous()
- else:
- edge_index = torch.empty((2,0), dtype=torch.long)
- return Data(x=x, edge_index=edge_index, pos=pos, batch=torch.zeros(x.shape[0], dtype=torch.long))
- def get_surface_points(self):
- conf = self.mol.GetConformer()
- return np.array([conf.GetAtomPosition(i) for i in range(self.mol.GetNumAtoms())])
- def compute_descriptors(mol):
- try:
- AllChem.ComputeGasteigerCharges(mol)
- charges = []
- for atom in mol.GetAtoms():
- if atom.HasProp('_GasteigerCharge'):
- charges.append(float(atom.GetDoubleProp('_GasteigerCharge')))
- return {
- 'MolWt': Descriptors.MolWt(mol),
- 'MolLogP': Descriptors.MolLogP(mol),
- 'NumRotatableBonds': Descriptors.NumRotatableBonds(mol),
- 'TPSA': Descriptors.TPSA(mol),
- 'HBA': rdMolDescriptors.CalcNumHBA(mol),
- 'HBD': rdMolDescriptors.CalcNumHBD(mol),
- 'AromaticRings': rdMolDescriptors.CalcNumAromaticRings(mol),
- 'FractionCSP3': rdMolDescriptors.CalcFractionCSP3(mol),
- 'HeavyAtomCount': Descriptors.HeavyAtomCount(mol),
- 'AliphaticRings': rdMolDescriptors.CalcNumAliphaticRings(mol),
- 'MaxPartialCharge': max(charges) if charges else 0.0,
- 'MinPartialCharge': min(charges) if charges else 0.0,
- 'LabuteASA': rdMolDescriptors.CalcLabuteASA(mol),
- 'MolMR': Descriptors.MolMR(mol),
- 'BalabanJ': Descriptors.BalabanJ(mol),
- 'BertzCT': Descriptors.BertzCT(mol)
- }
- except:
- return {key: 0.0 for key in [
- 'MolWt', 'MolLogP', 'NumRotatableBonds', 'TPSA', 'HBA', 'HBD',
- 'AromaticRings', 'FractionCSP3', 'HeavyAtomCount', 'AliphaticRings',
- 'MaxPartialCharge', 'MinPartialCharge', 'LabuteASA', 'MolMR',
- 'BalabanJ', 'BertzCT'
- ]}
- def compute_entropy(rot_bonds):
- delta_s_torsion = -R * np.log(rot_bonds + 1)
- return delta_s_torsion
- def compute_distance_to_site(mol, centroid):
- conf = mol.GetConformer()
- distances = []
- for atom in mol.GetAtoms():
- pos = np.array(conf.GetAtomPosition(atom.GetIdx()))
- dist = np.linalg.norm(pos - centroid)
- distances.append(dist)
- return np.mean(distances), np.min(distances), np.max(distances)
- class PharmaGNN(nn.Module):
- def __init__(self, protein_feature_size=10, ligand_feature_size=5):
- super().__init__()
- self.device = DEVICE
- self.protein_conv1 = GATv2Conv(protein_feature_size, 128, heads=2) # Reduced complexity
- self.protein_conv2 = GATv2Conv(128*2, 256)
- self.ligand_conv1 = GATv2Conv(ligand_feature_size, 128, heads=2)
- self.ligand_conv2 = GATv2Conv(128*2, 256)
- self.cross_attention = nn.MultiheadAttention(256, 4, batch_first=True) # Reduced heads
- self.fc = nn.Sequential(
- nn.Linear(512, 256),
- nn.LayerNorm(256),
- nn.LeakyReLU(),
- nn.Dropout(0.2),
- nn.Linear(256, 1)
- )
- self.to(self.device)
- def forward(self, protein_data, ligand_data):
- # Move data to device
- protein_data = protein_data.to(self.device)
- ligand_data = ligand_data.to(self.device)
- # Process protein
- p_x = self.protein_conv1(protein_data.x, protein_data.edge_index)
- p_x = self.protein_conv2(p_x, protein_data.edge_index)
- p_x = global_mean_pool(p_x, protein_data.batch)
- # Process ligand
- l_x = self.ligand_conv1(ligand_data.x, ligand_data.edge_index)
- l_x = self.ligand_conv2(l_x, ligand_data.edge_index)
- l_x = global_mean_pool(l_x, ligand_data.batch)
- # Cross attention
- attn_out, _ = self.cross_attention(p_x.unsqueeze(1), l_x.unsqueeze(1), l_x.unsqueeze(1))
- # Combine features
- combined = torch.cat([p_x, attn_out.squeeze(1)], dim=1)
- return self.fc(combined)
- def calculate_binding_score(protein_features, ligand_features, descriptors, distances, entropy):
- try:
- sc = 1 - cdist(protein_features['surface_points'], ligand_features['surface_points'], 'cosine').mean()
- ec = np.corrcoef(protein_features['electrostatic'], [ligand_features['esp']])[0,1]
- if np.isnan(ec):
- ec = 0.0
- dist_score = 1 / (1 + distances['AvgDistToSite'])
- entropy_score = -entropy / R
- energy = (0.3 * sc) + (0.2 * ec) + (0.2 * ligand_features['affinity']) + \
- (0.2 * dist_score) + (0.1 * entropy_score)
- return energy
- except:
- return 0.0
- def safe_save_dataframe(df, filename, max_retries=3):
- """Safely save DataFrame with retries"""
- for attempt in range(max_retries):
- try:
- df_clean = df.copy()
- if 'MolObj' in df_clean.columns:
- df_clean = df_clean.drop('MolObj', axis=1)
- df_clean = df_clean.replace([np.inf, -np.inf], np.nan)
- with pd.ExcelWriter(filename, engine='openpyxl') as writer:
- df_clean.to_excel(writer, index=False, sheet_name='Results')
- print(f"✅ Successfully saved {filename}")
- return True
- except Exception as e:
- print(f"❌ Attempt {attempt + 1} failed to save {filename}: {e}")
- if attempt < max_retries - 1:
- import time
- time.sleep(1)
- else:
- print(f"❌ Failed to save {filename} after {max_retries} attempts")
- return False
- def process_single_sdf_file(sdf_path, sdf_file, centroid, processed_files):
- """Process a single SDF file with continuous checkpointing"""
- results = []
- skipped = []
- # Check if file already processed
- if sdf_file in processed_files:
- print(f"📁 Loading cached results for {sdf_file}...")
- cached_file = f"descriptors_{sdf_file.replace('.sdf', '')}.xlsx"
- cached_df = load_checkpoint(cached_file)
- if cached_df is not None:
- print(f"✅ Loaded {len(cached_df)} cached results for {sdf_file}")
- return cached_df, []
- print(f"🔄 Processing {sdf_file}...")
- try:
- suppl = Chem.SDMolSupplier(sdf_path, sanitize=False)
- batch_results = []
- processed_count = 0
- total_processed = 0
- for idx, mol in enumerate(suppl):
- if idx % 1000 == 0:
- print(f"\r{sdf_file}: Processing molecule {idx + 1}...", end="", flush=True)
- if mol is None:
- skipped.append((sdf_file, idx, "Unreadable molecule"))
- continue
- try:
- # Sanitize and prepare molecule
- Chem.SanitizeMol(mol)
- mol = Chem.AddHs(mol)
- res = AllChem.EmbedMolecule(mol, randomSeed=42)
- if res != 0:
- skipped.append((sdf_file, idx, "3D embedding failed"))
- continue
- AllChem.MMFFOptimizeMolecule(mol)
- # Compute descriptors
- desc = compute_descriptors(mol)
- avg_d, min_d, max_d = compute_distance_to_site(mol, centroid)
- entropy = compute_entropy(desc['NumRotatableBonds'])
- desc.update({
- 'Ligand': f"{sdf_file}_mol_{idx}",
- 'AvgDistToSite': avg_d,
- 'MinDistToSite': min_d,
- 'MaxDistToSite': max_d,
- 'EntropyLoss': entropy,
- 'SourceFile': sdf_file,
- 'MolIndex': idx
- })
- batch_results.append(desc)
- processed_count += 1
- total_processed += 1
- # Save checkpoint every N molecules
- if processed_count >= DESCRIPTOR_CHECKPOINT_SIZE:
- results.extend(batch_results)
- # Save intermediate checkpoint
- checkpoint_df = pd.DataFrame(results)
- checkpoint_file = f"descriptors_{sdf_file.replace('.sdf', '')}_temp.xlsx"
- save_checkpoint(checkpoint_df, checkpoint_file)
- batch_results = []
- processed_count = 0
- # Memory cleanup
- with memory_cleanup():
- pass
- except Exception as e:
- skipped.append((sdf_file, idx, f"Processing error: {str(e)[:100]}"))
- continue
- # Add remaining results
- if batch_results:
- results.extend(batch_results)
- # Save final results for this file
- if results:
- final_df = pd.DataFrame(results)
- final_file = f"descriptors_{sdf_file.replace('.sdf', '')}.xlsx"
- save_checkpoint(final_df, final_file)
- # Mark file as processed
- processed_files[sdf_file] = {
- 'processed_count': total_processed,
- 'timestamp': datetime.datetime.now().isoformat()
- }
- save_processed_files(processed_files)
- print(f"\r{sdf_file}: Completed - {total_processed} molecules processed")
- except Exception as e:
- print(f"\n❌ Error processing {sdf_file}: {e}")
- return pd.DataFrame(), skipped
- return pd.DataFrame(results), skipped
- def process_ligands_from_multiple_sdfs(folder, centroid):
- """Process ligands with file-level checkpointing"""
- all_results = []
- all_skipped = []
- sdf_files = [f for f in os.listdir(folder) if f.endswith(".sdf")]
- processed_files = get_processed_files()
- print(f"📂 Found {len(sdf_files)} SDF files")
- print(f"📋 {len(processed_files)} files already processed")
- for file_idx, sdf_file in enumerate(sdf_files):
- print(f"\n🔄 File {file_idx + 1}/{len(sdf_files)}: {sdf_file}")
- sdf_path = os.path.join(folder, sdf_file)
- with memory_cleanup():
- results_df, skipped = process_single_sdf_file(sdf_path, sdf_file, centroid, processed_files)
- if not results_df.empty:
- all_results.append(results_df)
- all_skipped.extend(skipped)
- log_memory_usage(f"After processing {sdf_file}")
- # Combine all results
- if all_results:
- final_df = pd.concat(all_results, ignore_index=True)
- print(f"\n✅ Total molecules processed: {len(final_df)}")
- else:
- final_df = pd.DataFrame()
- print(f"\n❌ No molecules processed successfully")
- if all_skipped:
- print(f"⚠️ Skipped {len(all_skipped)} molecules. First 10 errors:")
- for skip_info in all_skipped[:10]:
- print(f" {skip_info[0]}_mol_{skip_info[1]}: {skip_info[2]}")
- return final_df
- class PharmaPipeline:
- def __init__(self):
- print("🔬 Initializing PharmaPipeline...")
- log_memory_usage("Initial")
- self.protein_features = process_protein(PDB_FILE, ACTIVE_SITE_RESIDUES)
- self.model = PharmaGNN()
- self.model.eval()
- self.scaler = StandardScaler()
- log_memory_usage("After initialization")
- def process_affinity_predictions(self, ligand_df):
- """Process affinity predictions with continuous checkpointing"""
- results = []
- # Prepare protein data once
- coords = self.protein_features['surface_points']
- dist_matrix = cdist(coords, coords)
- edge_list = np.where(dist_matrix < 5.0)
- edge_index = torch.tensor([edge_list[0], edge_list[1]], dtype=torch.long).contiguous()
- protein_data = Data(
- x=torch.tensor(self.protein_features['residue_features'], dtype=torch.float),
- edge_index=edge_index,
- batch=torch.zeros(len(self.protein_features['residue_features']), dtype=torch.long)
- )
- print(f"🧬 Predicting affinities for {len(ligand_df)} ligands...")
- # Process in smaller batches
- batch_size = AFFINITY_CHECKPOINT_SIZE
- num_batches = (len(ligand_df) - 1) // batch_size + 1
- for batch_idx in range(num_batches):
- batch_start = batch_idx * batch_size
- batch_end = min(batch_start + batch_size, len(ligand_df))
- print(f"\n🔄 Affinity batch {batch_idx + 1}/{num_batches} ({batch_start + 1}-{batch_end})")
- batch_results = []
- with memory_cleanup():
- for idx in range(batch_start, batch_end):
- row = ligand_df.iloc[idx]
- if (idx - batch_start) % 20 == 0:
- progress = f"Progress: {idx - batch_start + 1}/{batch_end - batch_start}"
- print(f"\r{progress}", end="", flush=True)
- try:
- # Create molecule from SMILES or use stored mol object
- mol_data = row.to_dict()
- # Create a dummy molecule for testing (replace with actual mol reconstruction)
- mol = Chem.MolFromSmiles('CCO') # Placeholder
- if mol is None:
- continue
- with torch.no_grad():
- lp = LigandProcessor(mol)
- pharma_features = lp.get_3dpharma_features()
- graph = lp.get_graph_representation()
- # Predict affinity
- affinity = self.model(protein_data, graph)
- pharma_features['affinity'] = affinity.item()
- # Calculate physical score
- phys_score = calculate_binding_score(
- self.protein_features,
- {'surface_points': lp.get_surface_points(), **pharma_features},
- mol_data,
- {'AvgDistToSite': row['AvgDistToSite']},
- row['EntropyLoss']
- )
- final_score = 0.7 * affinity.item() + 0.3 * phys_score
- result = {
- 'Drug': row['Ligand'],
- 'Affinity': affinity.item(),
- 'PhysScore': phys_score,
- 'FinalScore': final_score,
- **pharma_features,
- **mol_data
- }
- batch_results.append(result)
- except Exception as e:
- print(f"\n❌ Error predicting for {row['Ligand']}: {e}")
- continue
- # Save batch results
- results.extend(batch_results)
- if results:
- checkpoint_df = pd.DataFrame(results)
- timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
- checkpoint_file = f"affinity_results_batch_{batch_idx + 1}_{timestamp}.xlsx"
- save_checkpoint(checkpoint_df, checkpoint_file)
- print(f"\n💾 Saved {len(results)} affinity predictions")
- log_memory_usage(f"After affinity batch {batch_idx + 1}")
- return pd.DataFrame(results)
- def run(self):
- """Main pipeline execution with comprehensive checkpointing"""
- try:
- print("🚀 Starting PharmaPipeline execution...")
- # Validate inputs
- pdb_path = os.path.normpath(PDB_FILE)
- ligand_folder = os.path.normpath(LIGAND_FOLDER)
- if not os.path.exists(pdb_path):
- raise FileNotFoundError(f"PDB file {pdb_path} not found")
- if not os.path.exists(ligand_folder):
- raise FileNotFoundError(f"Ligand folder {ligand_folder} not found")
- # Step 1: Process ligand descriptors
- print("\n📊 Step 1: Processing ligand descriptors...")
- ligand_df = process_ligands_from_multiple_sdfs(ligand_folder, self.protein_features['centroid'])
- if ligand_df.empty:
- print("❌ No ligands processed successfully")
- return None
- # Step 2: Process affinity predictions
- print("\n🧬 Step 2: Processing affinity predictions...")
- results_df = self.process_affinity_predictions(ligand_df)
- if results_df.empty:
- print("❌ No affinity predictions completed")
- return None
- # Step 3: Final processing and saving
- print("\n📈 Step 3: Final processing...")
- results_df = results_df.sort_values('FinalScore', ascending=False)
- # Save final results
- timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
- output_file = f"Nature_Methods_Pharma_Ranking_{timestamp}.xlsx"
- success = safe_save_dataframe(results_df, output_file)
- if success:
- print(f"\n📄 Final output saved to: {output_file}")
- print(f"🏆 Top 5 compounds by FinalScore:")
- display_cols = ['Drug', 'FinalScore', 'Affinity', 'PhysScore']
- available_cols = [col for col in display_cols if col in results_df.columns]
- print(results_df[available_cols].head().to_string())
- return results_df
- except Exception as e:
- print(f"❌ Pipeline failed: {e}")
- print("🔍 Full traceback:")
- traceback.print_exc()
- return None
- finally:
- print("\n🧹 Performing final cleanup...")
- with memory_cleanup():
- pass
- log_memory_usage("Final cleanup")
- if __name__ == "__main__":
- try:
- log_memory_usage("Program start")
- print("🔬 Nature Methods Pharma Analysis Pipeline")
- print("=" * 50)
- pipeline = PharmaPipeline()
- result_df = pipeline.run()
- if result_df is not None:
- print("\n🎉 Nature Methods-ready analysis complete!")
- print(f"✅ Successfully processed {len(result_df)} compounds")
- else:
- print("\n❌ Analysis failed - check error messages above")
- log_memory_usage("Program end")
- except KeyboardInterrupt:
- print("\n⚠️ Process interrupted by user")
- print("💾 Checkpoint files have been saved - you can resume processing")
- sys.exit(1)
- except Exception as e:
- print(f"\n💥 Unexpected error: {e}")
- print("🔍 Full traceback:")
- traceback.print_exc()
- sys.exit(1)
- finally:
- print("\n🧹 Final cleanup...")
- with memory_cleanup():
- pass
- print("🏁 Program finished.")
pharmacophore_modeling.py at commit 9e48132, no license · at the source
Overview
- Centre for Integrative Omics Data Science (CIODS), Yenepoya (Deemed to be University), Mangalore 575018, Karnataka, India
- Center for Systems Biology and Molecular Medicine (CSBMM), Yenepoya (Deemed to be University), Mangalore 575018, Karnataka, India
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.
naveen-joy-18/Feature-driven-AI-Guided-Drug-Repurposing-Pipeline-for-AGPS-Targeted-Glioma-Therapy-
9e481328d5f3663cf67336f5a1a036a791a8def3, 18 December 2025Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
6 files
- Binding free energy model.py, Python, 506 lines, 5 matches
- language model.py, Python, 89 lines, 2 matches
- pharmacophore_modeling.p
y , Python, 694 lines, 9 matches - prediction.py, Python, 112 lines, 2 matches
- scoring assistant.py, Python, 551 lines, 1 match
- README.md, Text, 74 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 5 scripts, each with its path and the digest of its content;
- 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- rcsb.org/
structure/ , at PDB; found in the text, “Homology Modeling of AGPS”4bca - rcsb.org/
structure/ , at PDB; found in the text, “Homology Modeling and Protein Preparation”5ae1
Code and data availability statement
The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: naveen-joy-18/
Feature-driven-AI-Guided -Drug-Repurposing-Pipeli ne-for-AGPS-Targeted-Gli oma-Therapy-
Read it in the paper: doi.org/10.1021/acsomega.5c09368.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 44 references.
Cite
This paper
Thaikkad, A., Thomas, S. D., Dcunha, L., John, L., Vijayan, J., Joy, N., R Dev, R., Raju, R., & Jayanandan, A. (2026). Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening. ACS omega, 11(11), 17249-17265. https://
BibTeX
@article{thaikkad2026str
author = {Thaikkad, Amritha and Thomas, Sonet Daniel and Dcunha, Leona and John, Levin and Vijayan, Jishna and Joy, Naveen and R Dev, Radul and Raju, Rajesh and Jayanandan, Abhithaj},
title = {{Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening}},
journal = {ACS omega},
year = {2026},
month = mar,
volume = {11},
number = {11},
pages = {17249--17265},
publisher = {American Chemical Society},
issn = {2470-1343},
doi = {10.1021/
url = {https://
pmid = {41908407},
pmcid = {PMC13019191}
}
RIS
TY - JOUR
AU - Thaikkad, Amritha
AU - Thomas, Sonet Daniel
AU - Dcunha, Leona
AU - John, Levin
AU - Vijayan, Jishna
AU - Joy, Naveen
AU - R Dev, Radul
AU - Raju, Rajesh
AU - Jayanandan, Abhithaj
TI - Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening
T2 - ACS omega
J2 - ACS Omega
PY - 2026
DA - 2026/
VL - 11
IS - 11
SP - 17249
EP - 17265
SN - 2470-1343
PB - American Chemical Society
DO - 10.1021/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1021/
"type": "article-journal",
"title": "Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening",
"container-title": "ACS omega",
"author": [
{
"family": "Thaikkad",
"given": "Amritha"
},
{
"family": "Thomas",
"given": "Sonet Daniel"
},
{
"family": "Dcunha",
"given": "Leona"
},
{
"family": "John",
"given": "Levin"
},
{
"family": "Vijayan",
"given": "Jishna"
},
{
"family": "Joy",
"given": "Naveen"
},
{
"family": "R Dev",
"given": "Radul"
},
{
"family": "Raju",
"given": "Rajesh"
},
{
"family": "Jayanandan",
"given": "Abhithaj"
}
],
"container-title-short":
"volume": "11",
"issue": "11",
"page": "17249-17265",
"DOI": "10.1021/
"PMID": "41908407",
"PMCID": "PMC13019191",
"ISSN": "2470-1343",
"publisher": "American Chemical Society",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
11
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: RDKit, Biopython, Hugging Face Transformers, 8 other tools
- [2] doi:10.1038/s41598-026-53415-5 [code]
- Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.Journal: Scientific reportsIn common: RDKit, Biopython, PyTorch Geometric, 6 other tools, other condition
- [3] doi:10.1038/s41586-026-10670-w [code]
- Zero-shot design of drug-binding proteins via neural iterative selection-expansion.Journal: NatureIn common: RDKit, PyTorch Geometric, Plotly, 7 other tools
- [4] doi:10.1002/hbm.70469 [code]
- VarCoNet: A Variability-Aware Self-Supervised Framework for Functional Connectome Extraction From Resting-State fMRI.Journal: Human brain mappingIn common: PyTorch Geometric, Hugging Face Transformers, Plotly, 7 other tools
- [5] doi:10.3390/ijms27156614 [code]
- Candidalysin Inhibits &
lt;i& gt;Porphyromonas gingivalis& lt;/ i& gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions. Journal: International journal of molecular sciencesIn common: RDKit, Biopython, PyTorch Geometric, 6 other tools - [6] doi:10.1038/s41586-026-10391-0 [code]
- Cell-type-targeted mitochondrial transplantation rescues cell degeneration.Journal: NatureIn common: RDKit, Biopython, PyTorch Geometric, 6 other tools
- [7] doi:10.1021/acs.biochem.5c00596 [code]
- Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.Journal: BiochemistryIn common: RDKit, Biopython, PyTorch Geometric, 6 other tools
- [8] doi:10.1371/journal.pone.0345854 [code]
- Shedding light on neural learning to rank models for anticancer drug prioritization.Journal: PloS oneIn common: RDKit, PyTorch Geometric, PyTorch, 6 other tools, other condition
- [9] doi:10.34133/csbj.0036 [code]
- HYG-mol: An Interpretable Multimodal Hypergraph Framework for Molecular Property Prediction.Journal: Computational and structural biotechnology journalIn common: RDKit, PyTorch Geometric, Hugging Face Transformers, 5 other tools
- [10] doi:10.1093/nar/gkag706 [code]
- scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.Journal: Nucleic acids researchIn common: RDKit, PyTorch Geometric, PyTorch, 6 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 5 scripts, and 19 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:35dd99783c75c177…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
