OSCR

Zero-shot design of drug-binding proteins via neural iterative selection-expansion.

Code ↔ Paper

13 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 13 matches
  1. [1] § LASErMPNN neural network ↔ train_lasermpnn_streptavidin_heldout_split.py, lines 695–738 · score 0.77 · pretrainable ligand encoder, dihedral angles, backbone noise, backbone atoms, idealized, head
  2. [2] § LASErMPNN neural network ↔ train_lasermpnn.py, lines 696–729 · score 0.75 · dihedral angles, backbone noise, ligand encoder, backbone atoms, idealized, head
  3. [3] § LASErMPNN neural network ↔ train_lasermpnn_streptavidin_heldout_split.py, lines 695–738 · score 0.75 · pretrained ligand encoder, dihedral angles, node embeddings, ligand information, backbone atoms, streptavidin
  4. [4] § NISE sampling algorithm ↔ carp_dock-pykeops.py, lines 3–18 · score 0.72 · brute force, ligand orientation, rigid body, docking, position, pose
  5. [5] § NISE sampling algorithm ↔ carp_dock.py, lines 3–30 · score 0.72 · brute force, ligand orientation, rigid body, docking, position, pose
  6. [6] § LASErMPNN neural network ↔ train_ligandmpnn_streptavidin_heldout_split.py, lines 575–661 · score 0.66 · dihedral angles, node embeddings, streptavidin, backbone atoms, pretrained, autoregressively
  7. [7] § LASErMPNN neural network ↔ run_nise_boltz2x.py, lines 555–630 · score 0.58 · helix bundle, structures predicted, binding site, objective, polar, LASErMPNN
  8. [8] § Characterization of exatecan binders ↔ run_nise_boltz2x.py, lines 280–396 · score 0.57 · structure ligand, designed proteins, ligand pLDDT, stem, pose, affinities
  9. [9] § Characterization of exatecan binders ↔ run_nise_boltz2x_ligandmpnn.py, lines 259–353 · score 0.57 · structure ligand, designed proteins, ligand pLDDT, stem, pose, affinities
  10. [10] § Design of apixaban binders using NISE ↔ run_nise_boltz2x.py, lines 280–396 · score 0.52 · affinity probability, binary, ligand pLDDT, score, poses, Boltz
  11. [11] § Design of apixaban binders using NISE ↔ run_nise_boltz2x_ligandmpnn.py, lines 259–353 · score 0.52 · affinity probability, binary, ligand pLDDT, score, poses, Boltz
  12. [12] § Design of apixaban binders using NISE ↔ carp_dock-pykeops.py, lines 3–18 · score 0.50 · ligand positions, orientations, clashing, docked, translations, rotations
  13. [13] § Design of apixaban binders using NISE ↔ carp_dock.py, lines 3–30 · score 0.50 · ligand positions, orientations, clashing, docked, translations, rotations

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 630 lines · 35 KB · MIT · 3 matches

  1. import io
  2. import os
  3. import sys
  4. import time
  5. import json
  6. import shutil
  7. import warnings
  8. import subprocess
  9. from typing import *
  10. from pathlib import Path
  11. from collections import defaultdict
  12. warnings.filterwarnings("ignore")
  13. # NOTE: If you move the run_nise*.py script, adjust the NISE_DIRECTORY_PATH
  14. NISE_DIRECTORY_PATH = str(Path(os.path.abspath(__file__)).parent)
  15. LASER_PATH = str(Path(NISE_DIRECTORY_PATH) / 'LASErMPNN')
  16. sys.path.append(NISE_DIRECTORY_PATH)
  17. import wandb
  18. import torch
  19. import plotly
  20. import numpy as np
  21. import prody as pr
  22. import pandas as pd
  23. from rdkit import Chem
  24. import plotly.express as px
  25. from rdkit.Chem import AllChem
  26. from rdkit_to_params import Params
  27. from LASErMPNN.run_inference import load_model_from_parameter_dict # type: ignore
  28. from LASErMPNN.run_batch_inference import _run_inference, output_protein_structure, output_ligand_structure # type: ignore
  29. from utility_scripts.burial_calc import compute_fast_ligand_burial_mask
  30. from utility_scripts.calc_symmetry_aware_rmsd import _main as calc_rmsd
  31. def compute_objective_function(confidence_metrics_dict: dict, objective_function: str) -> float:
  32. """
  33. Compute the objective function which will be MAXIMIZED by NISE.
  34. The input to this function is a dictionary of metrics extracted from Boltz-2x
  35. including design_ligand_plddt, iptm, and affinity metrics if running with
  36. boltz2_predict_affinity = True.
  37. The default behavior is to return design_ligand_plddt, but other metrics
  38. and combinations of metrics could be used here.
  39. """
  40. # NOTE: add your own objective function in another elif block here!
  41. if objective_function == 'ligand_plddt':
  42. return confidence_metrics_dict['design_ligand_plddt']
  43. elif objective_function == 'ligand_plddt_and_iptm':
  44. return confidence_metrics_dict['design_ligand_plddt'] + confidence_metrics_dict['iptm']
  45. elif objective_function == 'iptm':
  46. return confidence_metrics_dict['iptm']
  47. elif objective_function == 'pbind':
  48. return confidence_metrics_dict['affinity_probability_binary']
  49. elif objective_function == 'ligand_plddt_and_pbind':
  50. return confidence_metrics_dict['design_ligand_plddt'] + confidence_metrics_dict['affinity_probability_binary']
  51. elif objective_function == 'iptm_and_pbind':
  52. return confidence_metrics_dict['iptm'] + confidence_metrics_dict['affinity_probability_binary']
  53. else:
  54. raise ValueError(f'Objective_function strategy {objective_function} is not implemented.')
  55. def get_boltz_yaml_boilerplate(sequence: str, smiles: str, predict_affinity: bool):
  56. boltz_yaml = f"version: 1\nsequences:\n - protein:\n id: A\n sequence: {sequence}\n msa: empty\n - ligand:\n id: B\n smiles: '{smiles}'\n"
  57. if predict_affinity:
  58. boltz_yaml += "properties:\n - affinity:\n binder: B\n"
  59. return boltz_yaml
  60. def check_input(dir_path: Path, model_weights_path: Path):
  61. if not dir_path.exists():
  62. raise FileNotFoundError(f"Path {dir_path} does not exist.")
  63. if not dir_path.is_dir():
  64. raise NotADirectoryError(f"Path {dir_path} is not a directory.")
  65. num_files = len([x for x in dir_path.iterdir() if x.is_file() and x.suffix == '.pdb'])
  66. if num_files == 0:
  67. raise FileNotFoundError(f"No PDB files found in {dir_path}")
  68. if not model_weights_path.exists():
  69. raise FileNotFoundError(f"Model weights file {model_weights_path} does not exist.")
  70. def handle_directory_creation(input_dir, model_checkpoint):
  71. input_backbones_path = input_dir / 'input_backbones'
  72. check_input(input_backbones_path, model_checkpoint)
  73. sampling_dataframe_path = input_dir / 'sampling_dataframes'
  74. sampled_backbones_path = input_dir / 'sampled_backbones'
  75. sampling_dataframe_path.mkdir(exist_ok=True)
  76. sampled_backbones_path.mkdir(exist_ok=True)
  77. sdf_path = input_dir / 'lig_from_input_pdb.sdf'
  78. params_path = input_dir / 'lig_from_input_pdb.params'
  79. return input_backbones_path, sampling_dataframe_path, sampled_backbones_path, sdf_path, params_path
  80. def construct_helper_files(sdf_path, params_path, backbone_path, ligand_smiles):
  81. input_protein = pr.parsePDB(str(backbone_path))
  82. assert isinstance(input_protein, pr.AtomGroup), f"Error loading backbone {backbone_path}"
  83. ligand = input_protein.select('not protein').copy()
  84. assert ligand.numResidues() == 1, f"Error selecting ligand from backbone {backbone_path}"
  85. # Create a pdb string for the ligand.
  86. pdb_stream = io.StringIO()
  87. pr.writePDBStream(pdb_stream, ligand)
  88. ligand_string = pdb_stream.getvalue()
  89. lignames_set = set(ligand.getResnames())
  90. # Write ligand to PDB file.
  91. pdb_path = Path(str(sdf_path.resolve()).rsplit('.')[0] + '.pdb')
  92. with pdb_path.open('w') as f:
  93. f.write(ligand_string)
  94. pdb_mol = Chem.MolFromPDBBlock(ligand_string)
  95. smi_mol = Chem.MolFromSmiles(ligand_smiles)
  96. pdb_mol = AllChem.AssignBondOrdersFromTemplate(smi_mol, pdb_mol)
  97. pdb_mol = AllChem.AddHs(pdb_mol, addCoords=True)
  98. AllChem.ComputeGasteigerCharges(pdb_mol)
  99. Chem.MolToMolFile(pdb_mol, str(sdf_path.resolve()))
  100. ligname = lignames_set.pop()
  101. p = Params.from_mol(pdb_mol, name=ligname)
  102. p.dump(params_path) # type: ignore
  103. with open(Path(params_path).parent / '.gitignore', 'w') as f:
  104. f.write('*')
  105. class DesignCampaign:
  106. def __init__(self,
  107. model_checkpoint, input_dir, ligand_rmsd_mask_atoms, ligand_atoms_enforce_buried, ligand_atoms_enforce_exposed, laser_inference_device, debug, ligand_3lc,
  108. rmsd_use_chirality, self_consistency_ligand_rmsd_threshold, self_consistency_protein_rmsd_threshold,
  109. laser_inference_dropout, num_iterations, num_top_backbones_per_round, laser_sampling_params, sequences_sampled_per_backbone,
  110. sequences_sampled_at_once, boltz_inference_devices, ligand_smiles, boltz2x_executable_path,
  111. use_reduce_protonation, keep_input_backbone_in_queue, keep_best_generator_backbone, use_boltz_conformer_potentials,
  112. boltz2_predict_affinity, drop_rmsd_mask_atoms_from_ligand_plddt_calc, use_boltz_1x,
  113. boltz2_disable_kernels, boltz2_disable_nccl_p2p, objective_function, fixed_identity_residue_indices,
  114. align_on_binding_site, burial_mask_alpha_hull_alpha, boltz2_cache_directory, boltz2_sampling_steps, **kwargs
  115. ):
  116. self.debug = debug
  117. self.ligand_3lc = ligand_3lc
  118. self.boltz_inference_devices = boltz_inference_devices
  119. self.use_boltz_conformer_potentials = use_boltz_conformer_potentials
  120. self.use_boltz_1x = use_boltz_1x
  121. self.boltz2x_executable_path = boltz2x_executable_path
  122. self.predict_affinity = boltz2_predict_affinity
  123. self.use_reduce_protonation = use_reduce_protonation
  124. self.boltz2_cache_directory = boltz2_cache_directory
  125. self.boltz2_sampling_steps = boltz2_sampling_steps
  126. self.boltz2_disable_kernels = boltz2_disable_kernels
  127. self.boltz2_disable_nccl_p2p = boltz2_disable_nccl_p2p
  128. self.objective_function = objective_function
  129. self.align_on_binding_site = align_on_binding_site
  130. self.fixed_identity_residue_indices = fixed_identity_residue_indices
  131. if self.fixed_identity_residue_indices is not None:
  132. print(f'Fixing residues: {self.fixed_identity_residue_indices}')
  133. if self.use_boltz_1x and self.predict_affinity:
  134. raise ValueError('Cannot use boltz1x with affinity prediction.')
  135. if 'pbind' in self.objective_function and (not self.predict_affinity):
  136. raise ValueError(f"predict_affinity must be True to use objective function {self.objective_function}")
  137. self.rmsd_use_chirality = rmsd_use_chirality
  138. self.self_consistency_ligand_rmsd_threshold = self_consistency_ligand_rmsd_threshold
  139. self.self_consistency_protein_rmsd_threshold = self_consistency_protein_rmsd_threshold
  140. self.keep_input_backbone_in_queue = keep_input_backbone_in_queue
  141. self.drop_rmsd_mask_atoms_from_ligand_plddt_calc = drop_rmsd_mask_atoms_from_ligand_plddt_calc
  142. self.num_iterations = num_iterations
  143. self.sequences_sampled_per_backbone = sequences_sampled_per_backbone
  144. self.sequences_sampled_at_once = sequences_sampled_at_once
  145. self.top_k = num_top_backbones_per_round
  146. self.ligand_rmsd_mask_atoms = ligand_rmsd_mask_atoms
  147. self.ligand_atoms_enforce_buried = ligand_atoms_enforce_buried
  148. self.ligand_atoms_enforce_exposed = ligand_atoms_enforce_exposed
  149. self.burial_mask_alpha_hull_alpha = burial_mask_alpha_hull_alpha
  150. self.ligand_smiles = ligand_smiles
  151. self.laser_sampling_params = laser_sampling_params
  152. self.laser_inference_dropout = laser_inference_dropout
  153. self.sampling_metadata = defaultdict(float)
  154. self.sampling_metadata['min_protein_rmsd'] = float('inf')
  155. self.sampling_metadata['min_ligand_rmsd'] = float('inf')
  156. self.input_backbones_path, self.sampling_dataframe_path, self.sampled_backbones_path, self.sdf_path, self.params_path = handle_directory_creation(input_dir, model_checkpoint)
  157. self.model, self.model_params = load_model_from_parameter_dict(model_checkpoint, torch.device(laser_inference_device))
  158. # Set model to eval mode, enable inference dropout if specified.
  159. self.model.eval()
  160. if self.laser_inference_dropout:
  161. for module in self.model.modules():
  162. if isinstance(module, torch.nn.Dropout):
  163. module.train()
  164. # Set path and priority for input backbones.
  165. if self.keep_input_backbone_in_queue:
  166. self.backbone_queue = [(x, torch.inf) for x in self.input_backbones_path.iterdir() if x.is_file() and x.suffix == '.pdb']
  167. else:
  168. self.backbone_queue = [(x, 0) for x in self.input_backbones_path.iterdir() if x.is_file() and x.suffix == '.pdb']
  169. # Whether to keep the pose that generates the highest confidence sequences.
  170. self.keep_best_generator_backbone = keep_best_generator_backbone
  171. self.backbone_to_best_generation = defaultdict(float)
  172. if len(self.backbone_queue) > 1:
  173. raise NotImplementedError(f"More than one input backbone not currently supported.")
  174. construct_helper_files(self.sdf_path, self.params_path, self.backbone_queue[0][0], ligand_smiles)
  175. self.ligand_rmsd_mask_atoms = ligand_rmsd_mask_atoms
  176. assert self.sdf_path.exists(), f"Error creating SDF file {self.sdf_path}"
  177. assert self.params_path.exists(), f"Error creating params file {self.params_path}"
  178. def sample_sequences(self, backbone_path: str) -> Tuple[List[pr.AtomGroup], List[str], List[float], List[float]]:
  179. laser_nll, laser_bs_nll = [], []
  180. sampled_proteins, sampled_sequences = [], []
  181. num_sampled = 0
  182. while num_sampled < self.sequences_sampled_per_backbone:
  183. remaining_samples = self.sequences_sampled_per_backbone - num_sampled
  184. num_seq_to_sample = min(remaining_samples, self.sequences_sampled_at_once)
  185. backbone_path_ = backbone_path
  186. if self.fixed_identity_residue_indices is not None:
  187. # NOTE: there might be a better way to do this so that we don't lose information stored in the b-factor columns of backbones that we sample,
  188. # but the important pLDDT information should be recorded in the log steps anyways.
  189. protein = pr.parsePDB(str(backbone_path))
  190. protein.setBetas(0.0)
  191. fixed_selection = protein.select(self.fixed_identity_residue_indices)
  192. if fixed_selection is None:
  193. raise ValueError(f'ProDy encounted an error selecting residues with selection string: {self.fixed_identity_residue_indices}')
  194. fixed_selection.setBetas(1.0)
  195. backbone_path_ = Path(backbone_path).parent / f'{Path(backbone_path).stem}_fixbeta.pdb'
  196. pr.writePDB(str(backbone_path_), protein)
  197. self.laser_sampling_params.update({'fix_beta': True})
  198. sampling_output, full_atom_coords, nh_coords, sampled_probs, batch_data, protein_complex_data = _run_inference(self.model, self.model_params, Path(backbone_path_), num_seq_to_sample, **self.laser_sampling_params)
  199. protein_complex_data = protein_complex_data[0] # type: ignore
  200. for idx in range(num_seq_to_sample):
  201. curr_batch_mask = batch_data.batch_indices == idx
  202. curr_probs = sampled_probs[curr_batch_mask]
  203. nll = (-1 * torch.log10(curr_probs)).cpu().numpy().mean()
  204. bs_nll = (-1 * torch.log10(sampled_probs[curr_batch_mask][batch_data.first_shell_ligand_contact_mask[curr_batch_mask]])).cpu().numpy().mean()
  205. out_prot = output_protein_structure(full_atom_coords[curr_batch_mask], sampling_output.sampled_sequence_indices[curr_batch_mask], protein_complex_data.residue_identifiers, nh_coords[curr_batch_mask], curr_probs)
  206. out_lig = output_ligand_structure(protein_complex_data.ligand_info)
  207. out_complex = out_prot + out_lig
  208. out_complex.setTitle('LASErMPNN/NISE Generated Protein')
  209. sampled_proteins.append(out_complex)
  210. sampled_sequences.append(out_prot.ca.getSequence())
  211. laser_nll.append(nll)
  212. laser_bs_nll.append(bs_nll)
  213. num_sampled += num_seq_to_sample
  214. return sampled_proteins, sampled_sequences, laser_nll, laser_bs_nll
  215. def identify_backbone_candidates(self, sorted_designs_boltz: Sequence[Path], sorted_designs_laser: Sequence[Path], sampled_backbone_paths: Sequence[Path], reduce_executable_path, reduce_hetdict_path):
  216. assert len(sorted_designs_laser) == len(sorted_designs_boltz), f"Error: different number of designs in" # {boltz_output_subdir} and {laser_output_subdir}."
  217. smi_mol = Chem.MolFromSmiles(self.ligand_smiles)
  218. log_data = defaultdict(list)
  219. for laser, boltz, bb_path in zip(sorted_designs_laser, sorted_designs_boltz, sampled_backbone_paths):
  220. laser_prot = pr.parsePDB(str(laser))
  221. boltz_string_str = open(boltz, 'r').read()
  222. boltz_string_io = io.StringIO(boltz_string_str)
  223. boltz_prot = pr.parsePDBStream(boltz_string_io)
  224. confidence_data = {'laser_output_pdb_path': laser, 'boltz_output_pdb_path': boltz}
  225. with (boltz.parent / f'confidence_{boltz.stem}.json').open('r') as f:
  226. confidence_data.update(json.load(f))
  227. design_iptm = confidence_data['iptm']
  228. design_bind_probability, design_predicted_affinity = torch.nan, torch.nan
  229. if self.predict_affinity:
  230. with (boltz.parent / f'affinity_{boltz.stem.replace("_model_0", "")}.json').open('r') as f:
  231. confidence_data.update(json.load(f))
  232. design_bind_probability = confidence_data['affinity_probability_binary']
  233. design_predicted_affinity = confidence_data['affinity_pred_value']
  234. log_data['affinity_probability_binary'].append(design_bind_probability)
  235. log_data['affinity_pred_value'].append(design_predicted_affinity)
  236. log_data['iptms'].append(design_iptm)
  237. try:
  238. protein_rmsd, ligand_rmsd, laser_to_boltz_name_mapping = calc_rmsd(laser_prot, boltz_prot, self.ligand_smiles, self.rmsd_use_chirality, self.ligand_rmsd_mask_atoms, align_on_binding_site=self.align_on_binding_site)
  239. if len(laser_to_boltz_name_mapping) == 0:
  240. raise ValueError('No atoms in common between laser and boltz structures.')
  241. boltz_to_laser_name_mapping = {v: k for k, v in laser_to_boltz_name_mapping.items()}
  242. except:
  243. protein_rmsd = np.nan
  244. ligand_rmsd = np.nan
  245. laser_to_boltz_name_mapping = {}
  246. log_data['ligand_rmsds'].append(ligand_rmsd)
  247. log_data['protein_rmsds'].append(protein_rmsd)
  248. if len(laser_to_boltz_name_mapping) == 0:
  249. print(f'{boltz}: Failed to map names between laser and boltz structures.')
  250. log_data['ligand_is_buried'].append(False)
  251. log_data['ligand_plddts'].append(torch.nan)
  252. log_data['protein_plddts'].append(torch.nan)
  253. continue
  254. # Remap the boltz structure ligand atoms with the name mapping.
  255. boltz_prot_only = boltz_prot.select('chid A')
  256. boltz_lig_only = boltz_prot.select('chid B and not element H')
  257. boltz_lig_only.setNames([boltz_to_laser_name_mapping[x] for x in boltz_lig_only.getNames()])
  258. boltz_lig_only.setResnames([self.ligand_3lc for _ in range(len(boltz_lig_only.getResnames()))])
  259. boltz_coords = boltz_lig_only.getCoords()
  260. atoms_enforced_buried_mask = np.array([x in self.ligand_atoms_enforce_buried for x in boltz_lig_only.getNames()])
  261. atoms_enforced_exposed_mask = np.array([x in self.ligand_atoms_enforce_exposed for x in boltz_lig_only.getNames()])
  262. pdb_output_path = str(boltz)
  263. pr.writePDB(pdb_output_path, boltz_prot_only + boltz_lig_only)
  264. # Compute ligand pLDDT over relevant atoms.
  265. rmsd_mask = np.array([x not in self.ligand_rmsd_mask_atoms for x in boltz_lig_only.getNames()])
  266. if not self.drop_rmsd_mask_atoms_from_ligand_plddt_calc:
  267. rmsd_mask = np.ones_like(rmsd_mask)
  268. design_ligand_plddt = boltz_lig_only.getBetas()[rmsd_mask].mean() / 100
  269. design_protein_plddt = boltz_prot_only.copy().ca.getBetas().mean() / 100
  270. log_data['ligand_plddts'].append(design_ligand_plddt)
  271. log_data['protein_plddts'].append(design_protein_plddt)
  272. confidence_data['design_ligand_plddt'] = design_ligand_plddt
  273. confidence_data['design_protein_plddt'] = design_protein_plddt
  274. # Check ligand burial constraints are obeyed in predicted structure.
  275. all_buried_mask = compute_fast_ligand_burial_mask(boltz_prot.ca.getCoords(), boltz_coords[atoms_enforced_buried_mask], num_rays=5, alpha=self.burial_mask_alpha_hull_alpha)
  276. none_buried_mask = compute_fast_ligand_burial_mask(boltz_prot.ca.getCoords(), boltz_coords[atoms_enforced_exposed_mask], num_rays=5, alpha=self.burial_mask_alpha_hull_alpha)
  277. if (all_buried_mask.all().item() and (not none_buried_mask.any().item())) or self.debug:
  278. log_data['ligand_is_buried'].append(True)
  279. if (ligand_rmsd < self.self_consistency_ligand_rmsd_threshold and protein_rmsd < self.self_consistency_protein_rmsd_threshold) or self.debug:
  280. if self.use_reduce_protonation:
  281. subprocess.run(
  282. f'{reduce_executable_path} -DB {reduce_hetdict_path} -DROP_HYDROGENS_ON_ATOM_RECORDS -BUILD {pdb_output_path} > {pdb_output_path}_',
  283. shell=True, check=False,
  284. stdout=subprocess.DEVNULL if not self.debug else subprocess.PIPE,
  285. stderr=subprocess.DEVNULL if not self.debug else subprocess.PIPE
  286. )
  287. shutil.move(f'{pdb_output_path}_', pdb_output_path)
  288. else:
  289. # Protonate using RDKit
  290. ligand_string = io.StringIO()
  291. pr.writePDBStream(ligand_string, boltz_lig_only.copy())
  292. pdb_mol = Chem.MolFromPDBBlock(ligand_string.getvalue())
  293. pdb_mol = AllChem.AssignBondOrdersFromTemplate(smi_mol, pdb_mol)
  294. pdb_mol = AllChem.AddHs(pdb_mol, addCoords=True)
  295. ligand_prody = pr.parsePDBStream(io.StringIO(Chem.MolToPDBBlock(pdb_mol)))
  296. ligand_prody.setResnames(self.ligand_3lc)
  297. ligand_prody.setChids('B')
  298. pr.writePDB(pdb_output_path, boltz_prot_only.copy() + ligand_prody)
  299. score = compute_objective_function(confidence_data, self.objective_function)
  300. self.backbone_to_best_generation[bb_path] = max(score, self.backbone_to_best_generation[bb_path])
  301. self.backbone_queue.append((pdb_output_path, score))
  302. else:
  303. log_data['ligand_is_buried'].append(False)
  304. self.backbone_queue = sorted(self.backbone_queue, key=lambda x: float(x[1]), reverse=True)[:self.top_k]
  305. # Adds the backbone which has generated the best scoring pose if it's not in the queue already.
  306. if self.keep_best_generator_backbone and len(self.backbone_to_best_generation) > 0:
  307. top_backbone = max(self.backbone_to_best_generation.items(), key=lambda x: x[1])
  308. print(top_backbone)
  309. if not (top_backbone[0] in [x[0] for x in self.backbone_queue]):
  310. self.backbone_queue = self.backbone_queue[:-1] + [top_backbone]
  311. return sorted_designs_laser, sorted_designs_boltz, log_data
  312. def log(self, use_wandb: bool, dataframe: pd.DataFrame):
  313. self.sampling_metadata['min_protein_rmsd'] = min(self.sampling_metadata['min_protein_rmsd'], dataframe['protein_rmsds'].dropna().min())
  314. self.sampling_metadata['min_ligand_rmsd'] = min(self.sampling_metadata['min_ligand_rmsd'], dataframe['ligand_rmsds'].dropna().min())
  315. logs = dict(self.sampling_metadata)
  316. logs['mean_sampled_ligand_plddt'] = dataframe['ligand_plddts'].dropna().mean()
  317. logs['mean_sampled_protein_plddt'] = dataframe['protein_plddts'].dropna().mean()
  318. logs['mean_sampled_protein_rmsd'] = dataframe['protein_rmsds'].dropna().mean()
  319. logs['mean_sampled_ligand_rmsd'] = dataframe['ligand_rmsds'].dropna().mean()
  320. logs['num_sequences_sampled'] = len(dataframe)
  321. logs['mean_affinity_probability'] = dataframe['affinity_probability_binary'].mean()
  322. logs['mean_affinity_pred_value'] = dataframe['affinity_pred_value'].mean()
  323. logs['mean_iptm'] = dataframe['iptms'].mean()
  324. logs['mean_laser_score'] = dataframe['laser_nll'].mean()
  325. logs['mean_laser_bs_score'] = dataframe['laser_bs_nll'].mean()
  326. try:
  327. if self.keep_input_backbone_in_queue:
  328. logs['max_sampled_backbone_priority'] = max(self.backbone_queue[1:], key=lambda x: x[1])[1]
  329. else:
  330. logs['max_sampled_backbone_priority'] = max(self.backbone_queue, key=lambda x: x[1])[1]
  331. except:
  332. logs['max_sampled_backbone_priority'] = 0
  333. print(logs)
  334. if use_wandb:
  335. scatter_fig = px.scatter(dataframe, x='ligand_rmsds', y='ligand_plddts', color='protein_rmsds', hover_data=['protein_rmsds', 'sequences'], range_color=[0.0, 2.5], range_y=[0.0, 1.0], range_x=[0.0, 5.0])
  336. logs['ligand_RMSD_vs_pLDDT_scatter'] = wandb.Html(plotly.io.to_html(scatter_fig)) # type: ignore
  337. wandb.log(logs)
  338. def predict_complex_structures(
  339. boltz_inputs_dir, boltz2x_executable_path, boltz_inference_devices,
  340. boltz_output_dir, use_potentials, use_boltz_1x, disable_kernels, disable_nccl_p2p, boltz2_cache_directory, boltz2_sampling_steps, debug
  341. ):
  342. device_ints = [x.split(':')[-1] for x in boltz_inference_devices]
  343. command = f'{boltz2x_executable_path} predict {boltz_inputs_dir} --devices {len(device_ints)} --out_dir {boltz_output_dir} --output_format pdb --override --sampling_steps {int(boltz2_sampling_steps)} --sampling_steps_affinity {int(boltz2_sampling_steps)}'
  344. command = f'CUDA_VISIBLE_DEVICES={",".join(device_ints)} {command}'
  345. if disable_nccl_p2p:
  346. command = f'NCCL_P2P_DISABLE=1 {command}'
  347. if use_potentials:
  348. command += f' --use_potentials'
  349. if use_boltz_1x:
  350. command += f' --model boltz1'
  351. if disable_kernels:
  352. command += f' --no_kernels'
  353. if boltz2_cache_directory is not None:
  354. command += f' --cache {boltz2_cache_directory}'
  355. print(command)
  356. try:
  357. # Boltz sometimes completes with a nonzero exit code despite completing successfully.
  358. # If not all expected files were generated NISE will crash at the log step.
  359. subprocess.run(command, shell=True, check=False, stdout=subprocess.DEVNULL if not debug else None, stderr=subprocess.DEVNULL if not debug else None)
  360. except:
  361. print('Boltz crashed! This might be fine, trying to recover...')
  362. pass
  363. def main(use_wandb, reduce_executable_path, reduce_hetdict_path, **kwargs):
  364. design_campaign = DesignCampaign(**kwargs)
  365. for iidx in range(design_campaign.num_iterations):
  366. # Run laser on all backbone queue inputs.
  367. all_sampled_proteins, backbone_sample_indices, sampled_backbone_path, all_sampled_sequences = [], [], [], []
  368. laser_nll, laser_bs_nll = [], []
  369. for bidx, (backbone_path, score) in enumerate(design_campaign.backbone_queue):
  370. sampled_proteins, sampled_sequences, nlls, bs_nlls = design_campaign.sample_sequences(backbone_path)
  371. all_sampled_proteins.extend(sampled_proteins)
  372. backbone_sample_indices.extend([bidx] * len(sampled_proteins))
  373. sampled_backbone_path.extend([design_campaign.backbone_queue[bidx][0]] * len(sampled_proteins))
  374. all_sampled_sequences.extend(sampled_sequences)
  375. laser_nll.extend(nlls)
  376. laser_bs_nll.extend(bs_nlls)
  377. # Sanity check output shapes before trying to fold or log anything.
  378. assert len(all_sampled_proteins) == len(backbone_sample_indices), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(backbone_sample_indices)}'
  379. assert len(all_sampled_proteins) == len(laser_nll), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(laser_nll)}'
  380. assert len(all_sampled_proteins) == len(laser_bs_nll), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(laser_bs_nll)}'
  381. # Make subdirectories to write outputs to disk.
  382. sampling_subdir = design_campaign.sampled_backbones_path / f'iter_{iidx}'
  383. laser_output_subdir = sampling_subdir / 'laser_outputs'
  384. boltz_input_dir = sampling_subdir / 'boltz_inputs'
  385. laser_output_subdir.mkdir(exist_ok=True, parents=True)
  386. boltz_input_dir.mkdir(exist_ok=True)
  387. # Write all the boltz input directories.
  388. all_boltz_input_path_names = []
  389. sampled_sequence_chunks = np.array_split(all_sampled_sequences, len(design_campaign.boltz_inference_devices))
  390. for idx, chunk_sequences in enumerate(sampled_sequence_chunks):
  391. for seq_idx, seq in enumerate(chunk_sequences):
  392. boltz_input_output_path = boltz_input_dir / f'chunk_{idx}_seq_{seq_idx}.yaml'
  393. all_boltz_input_path_names.append(boltz_input_output_path.stem)
  394. with boltz_input_output_path.open('w') as f:
  395. f.write(get_boltz_yaml_boilerplate(seq, design_campaign.ligand_smiles, design_campaign.predict_affinity))
  396. all_boltz_model_paths = [(sampling_subdir / 'boltz_results_boltz_inputs' / 'predictions' / x / f'{x}_model_0.pdb') for x in all_boltz_input_path_names]
  397. all_laser_output_paths = []
  398. for laser_output_structure, boltz_output_path in zip(all_sampled_proteins, all_boltz_model_paths):
  399. laser_output_path = laser_output_subdir / f'laser_{boltz_output_path.parent.stem}.pdb'
  400. pr.writePDB(str(laser_output_path), laser_output_structure)
  401. all_laser_output_paths.append(laser_output_path)
  402. curr_tries = 0
  403. max_tries = 10
  404. while curr_tries < max_tries and not all([x.exists() for x in all_boltz_model_paths]):
  405. if curr_tries != 0:
  406. print('Not all boltz predictions were completed or file system not updated.. retrying...')
  407. time.sleep(30)
  408. predict_complex_structures(
  409. boltz_input_dir, design_campaign.boltz2x_executable_path, design_campaign.boltz_inference_devices,
  410. sampling_subdir, design_campaign.use_boltz_conformer_potentials, design_campaign.use_boltz_1x,
  411. design_campaign.boltz2_disable_kernels, design_campaign.boltz2_disable_nccl_p2p, design_campaign.boltz2_cache_directory, design_campaign.boltz2_sampling_steps, design_campaign.debug
  412. )
  413. curr_tries += 1
  414. assert all([x.exists() for x in all_boltz_model_paths]), f"Error: not all boltz predictions were written to disk."
  415. # Identify any new backbone candidates.
  416. sorted_designs_laser, sorted_designs_rosetta, log_data = design_campaign.identify_backbone_candidates(all_boltz_model_paths, all_laser_output_paths, sampled_backbone_path, reduce_executable_path, reduce_hetdict_path)
  417. with open(sampling_subdir / 'backbone_queue.txt', 'w') as f:
  418. f.write('\n'.join([f"{x[1]}\t{x[0]}" for x in design_campaign.backbone_queue]))
  419. iidx_data = {
  420. 'sampled_backbone_path': sampled_backbone_path,
  421. 'laser_nll': laser_nll,
  422. 'laser_bs_nll': laser_bs_nll,
  423. 'laser_paths': sorted_designs_laser,
  424. 'rosetta_paths': sorted_designs_rosetta,
  425. 'sequences': all_sampled_sequences,
  426. 'curr_idx': [iidx] * len(laser_nll),
  427. }
  428. iidx_data.update(dict(log_data))
  429. # Write data to disk.
  430. iidx_dataframe = pd.DataFrame(iidx_data)
  431. iidx_dataframe.to_pickle(design_campaign.sampling_dataframe_path / f'iter_{iidx}_data.pkl')
  432. design_campaign.log(use_wandb, iidx_dataframe)
  433. if __name__ == "__main__":
  434. laser_sampling_params = {
  435. 'sequence_temp': 0.5, 'first_shell_sequence_temp': 0.7,
  436. 'chi_temp': 1e-6, 'seq_min_p': 0.0, 'chi_min_p': 0.0,
  437. 'disable_pbar': True, 'disabled_residues_list': ['X', 'C'], # Disables cysteine sampling by default.
  438. # ====================================================================================================
  439. # If constrain_ala_gly_sampling_to_exposed_non_secondary_structure is True (recommended),
  440. # the ala_budget and gly_budget parameters are
  441. # used to constrain the sampling of ALA and GLY residues to exposed non-secondary structured residues.
  442. # You can override the specific residues with the budget_residue_sele_string parameter (not recommended).
  443. # If constrain_ala_gly_sampling_to_exposed_non_secondary_structure is False and budget_residue_sele_string is None,
  444. # no constraints are applied to the sampling of ALA and GLY residues.
  445. # The reason we suggest doing this is to constrain the generated sequences to the manifold of sequences that are likely to exist in nature
  446. # and not exploiting a propensity of structure prediction networks to predict structure in alanine rich sequences
  447. # ====================================================================================================
  448. 'constrain_ala_gly_sampling_to_exposed_non_secondary_structure': True,
  449. 'budget_residue_sele_string': None,
  450. 'ala_budget': 4, 'gly_budget': 0, # May sample up to 4 Ala and 0 Gly over the selected region if not None.
  451. 'disable_charged_fs': True, # Disables sampling D,E,K,R residues for buried residues around the ligand.
  452. }
  453. params = dict(
  454. debug = (debug := False),
  455. use_wandb = (use_wandb := False),
  456. input_dir = Path('./debug/').resolve(),
  457. ligand_3lc = 'GG2', # Should match CCD code if using reduce.
  458. ligand_rmsd_mask_atoms = set(), # Atoms to IGNORE in RMSD calculation.
  459. ligand_atoms_enforce_buried = set(), # Atoms to enforce remain buried inside convex hull when selecting new backbones.
  460. ligand_atoms_enforce_exposed = set(), # Atoms to enforce remain exposed relative to the convex hull when selecting new backbones. I would suggest only using this for linker regions attached to your ligand or clearly exposed charged polar groups.
  461. laser_sampling_params = laser_sampling_params,
  462. ligand_smiles = 'COC1=CC=C(C=C1)N2C3=C(CCN(C3=O)C4=CC=C(C=C4)N5CCCCC5=O)C(=N2)C(=O)N',
  463. objective_function = (objective_function := 'ligand_plddt'), # Current options: {'ligand_plddt', 'iptm', 'ligand_plddt_and_iptm', 'pbind', 'ligand_plddt_and_pbind', 'iptm_and_pbind'}, Check the top of the file for implemented strategies, if you find an alternative strategy to work well please make a git commit so others can test it out as well!
  464. drop_rmsd_mask_atoms_from_ligand_plddt_calc = True,
  465. keep_input_backbone_in_queue = False,
  466. keep_best_generator_backbone = True, # The highest scoring pose may not necessarily generate higher scoring poses, keeps the pose that has generated the best poses after the first iteration in the queue if not already the best scoring pose.
  467. rmsd_use_chirality = False, # Will fail to compute RMSD on mismatched chirality ligands, might be bugged...
  468. self_consistency_ligand_rmsd_threshold = 2.5,
  469. self_consistency_protein_rmsd_threshold = 2.5,
  470. align_on_binding_site = False, # If align on binding site is True, protein RMSD (and self_consistency_protein_rmsd_threshold above) becomes binding site RMSD. Useful if you have a large protein with floppy regions away from the ligand. Binding site is computed as residues with sidechain atoms within 5.0A of the ligand in the lasermpnn output structure.
  471. fixed_identity_residue_indices = None, # An optional prody selection string of the form "resindex 0 2 3 4 5" or "resnum 1 3 5"...
  472. use_reduce_protonation = False, # If False, will use RDKit to protonate, these hydrogens will not preserve the input names and aren't placed conditioned on the sidechains but REDUCE sometimes drops hydrogens if geometry changes outside of expected bounds.
  473. reduce_hetdict_path = Path('./modified_hetdict.txt').absolute(), # Can set to None if use_reduce_protonation False
  474. reduce_executable_path = None, # Can set to None if use_reduce_protonation False
  475. model_checkpoint = Path(LASER_PATH) / 'model_weights/laser_weights_0p1A_nothing_heldout.pt',
  476. num_iterations = 35,
  477. num_top_backbones_per_round = 3,
  478. sequences_sampled_at_once = 30,
  479. boltz2x_executable_path = str((Path(NISE_DIRECTORY_PATH) / '.venv/bin/boltz').absolute()),
  480. boltz2_cache_directory = None, # Optional path to the boltz weights, can be used to avoid redownloading weights that have already been cached on your machine not in the default location.
  481. boltz2_sampling_steps = 200,
  482. boltz_inference_devices = (boltz_inference_devices := ['cuda:0',]), # a list of multiple torch-style device strings
  483. use_boltz_conformer_potentials = True, # Use Boltz-x mode, this is almost always better.
  484. boltz2_predict_affinity = True if ('pbind' in objective_function) else False,
  485. use_boltz_1x = False, # Run the same script using --model boltz-1, multi-device inference with this seems bugged with boltz v2.1.1
  486. boltz2_disable_kernels = False, # Disables cuEquivariance kernels, this is likely not necessary.
  487. boltz2_disable_nccl_p2p = False, # On some systems with certain graphics cards, NCCL can hang indefinitely. This flag fixes this issue allowing running boltz / NISE with multiple GPUs. https://github.com/NVIDIA/nccl/issues/631
  488. sequences_sampled_per_backbone = 64 if not debug else 1 * len(boltz_inference_devices),
  489. burial_mask_alpha_hull_alpha = 9.0, # Set to a larger number for folds with wider pockets (ex: 7-helix bundle) (Ex: 100.0), see https://github.com/benf549/CARPdock/blob/main/visualize_hull.ipynb
  490. laser_inference_device = boltz_inference_devices[0],
  491. laser_inference_dropout = True,
  492. )
  493. if use_wandb:
  494. wandb.init(project='design-campaigns', entity='benf549', config=params)
  495. main(**params)

run_nise_boltz2x.py at commit 61b7500, under MIT · at the source

Overview

Authors: Benjamin Fry1,2,3, Kaia Slaw2,3, Nicholas F. Polizzi2,3
  1. Harvard Graduate Program in Biophysics, Harvard University,Boston, MA USA
  2. Department of Cancer Biology, Dana-Farber Cancer Institute,Boston, MA USA
  3. Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School,Boston, MA USA
Institutions: Harvard University (United States); Dana-Farber Cancer Institute (United States)
Journal: Nature, volume 656, issue 8126, pages 237-249
Dates: received 22 April 2025; accepted 15 May 2026; published online 24 June 2026; in print 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41586-026-10670-w · PMID 42343133 · PMCID PMC13441969 · OpenAlex W7165762908
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: computational modeling (no new data) (modality), none (in silico) (organism), cellular / molecular (subfield)
Keywords: Protein design, Computational biophysics
MeSH: Drug Design*, Protein Engineering*, Proteins*, Algorithms, Graph Neural Networks, Ligands, Models, Molecular, Molecular Docking Simulation, Protein Binding, Protein Conformation, Pyrazoles (* major topic)
Topic: Protein Structure and Dynamics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: cited by 4 papers (Europe PMC); 58 references in the paper

Abstract

The design of proteins that bind to small molecules has been challenging because it requires simultaneous optimization of the protein sequence, protein structure and ligand conformation1–7. Current deep-learning algorithms have struggled to navigate this landscape, precluding the zero-shot design of binders. Here we show that by combining two neural networks in an iterative design algorithm, small-molecule binding proteins can be created from scratch with high accuracy. We trained a graph neural network—ligand-aware sequence engineering message-passing neural network (LASErMPNN)—to design compatible protein sequences for an input protein backbone and docked ligand. We paired LASErMPNN with a structure predictor that models a three-dimensional protein–ligand complex for an input protein sequence and ligand identity. The closed-loop iteration of these reciprocal networks optimized sequence–structure–ligand compatibility, and outperformed a comparable design loop using a physics-based energy function. We used our strategy, termed neural iterative selection–expansion (NISE), to design proteins that, using different folds, specifically bind to two chemically distinct small-molecule drugs, exatecan and apixaban, with success rates of 100% and 83%, respectively. The tightest NISE binders had nanomolar-to-picomolar affinities, surpassing those of the next-leading method by 70-fold for exatecan and nearly 10,000-fold for apixaban. LASErMPNN then suggested two amino-acid substitutions that improved the affinity of the tightest exatecan binder by 100-fold without any experimental input. The optimized binder protected the labile lactone ring of exatecan from hydrolysis for days. Our work describes a general recipe for using neural networks to automate the design of small-molecule binding proteins for applications in drug delivery, sensing and catalysis.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.

polizzilab/LASErMPNN

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: e70f2c6d765416f7e29d51bfd6d4e08496438878, 4 September 2026
Languages: Python (31), Shell (2), Jupyter (1)
Size: 60 files, 34 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (conda_env.yml, conda_env_cuda_12p4.yml), tests, 1 notebook
Not found: CITATION.cff, continuous integration, documentation
Tools: PyTorch (30 files), NumPy (16 files), PyTorch Geometric (14 files), pandas (8 files), Matplotlib (6 files), SciPy (3 files), scikit-learn (2 files), h5py (1 file), Plotly (1 file), RDKit (1 file), seaborn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
36 files

polizzilab/nise

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 61b7500e99ab37295ae9510b3db05aefe6a5fdfb, 11 August 2026
Languages: Python (9), Jupyter (2), Shell (2)
Size: 26 files, 13 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (NISE.def, setup.py), 2 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (5 files), RDKit (5 files), Plotly (4 files), PyTorch (4 files), pandas (3 files), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
15 files

benf549/CARPdock

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 9e11474052a0934444acfb5392e56b67f1a12c58, 18 February 2026
Languages: Python (4), Jupyter (2)
Size: 101 files, 6 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, 2 notebooks
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: PyTorch (4 files), NumPy (3 files), SciPy (3 files), scikit-learn (2 files), Plotly (1 file), RDKit (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
8 files

Zenodo 18308430

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 6 files
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

Code availability

All model inference code, training code, processed datasets and dataset split information can be found at GitHub (https://github.com/polizzilab/LASErMPNN). A Google Colab notebook implementing LASErMPNN sequence design and proofreading can be found at https://colab.research.google.com/github/polizzilab/LASErMPNN/blob/main/run_lasermpnn.ipynb. An implementation of the NISE protocol using Boltz-1/2 can be found at GitHub (https://github.com/polizzilab/NISE). A minimal implementation of NISE using LASErMPNN and Boltz-2 can be found at https://colab.research.google.com/github/polizzilab/NISE/blob/main/NISE_LASErMPNN.ipynb. An implementation of the rigid ligand–protein docking protocol used to generate NISE inputs for apixaban binders can be found at GitHub (https://github.com/benf549/CARPdock). Analysis code and PDB files of designs can be found at Zenodo58.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 53 scripts, each with its path and the digest of its content;
  • 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

Coordinates and data files for the X-ray crystal structures of exatecan-bound EPIC and exatecan-bound EPIC(Q51N) have been deposited in the PDB with accession codes 9NZE (http://doi.org/10.2210/pdb9nze/pdb) (EPIC) and 9NZG (http://doi.org/10.2210/pdb9NZG/pdb) (EPIC(Q51N)). Other referenced PDB IDs include 6W70 (http://doi.org/10.2210/pdb6W70/pdb), 4JNJ (http://doi.org/10.2210/pdb4JNJ/pdb), 4L9K (http://doi.org/10.2210/pdb4L9K/pdb) and 8TN6 (http://doi.org/10.2210/pdb8TN6/pdb). The X-ray crystal structure of the camptothecin derivative used to build exatecan conformers was retrieved from Cambridge Structure Database (CSD) deposition number 1504071. All data used to train LASErMPNN are publicly available from the PDB or the SPICE v.2.0.1 dataset56. The processed datasets used to train the LASErMPNN model are available at Zenodo57.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 2 keywords, 11 MeSH terms, 58 references.

Cite

This paper

Fry, B., Slaw, K., & Polizzi, N. F. (2026). Zero-shot design of drug-binding proteins via neural iterative selection-expansion. Nature, 656(8126), 237-249. https://doi.org/10.1038/s41586-026-10670-w

BibTeX

@article{fry2026zero,
author = {Fry, Benjamin and Slaw, Kaia and Polizzi, Nicholas F.},
title = {{Zero-shot design of drug-binding proteins via neural iterative selection-expansion}},
journal = {Nature},
year = {2026},
month = jun,
volume = {656},
number = {8126},
pages = {237--249},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/s41586-026-10670-w},
url = {https://doi.org/10.1038/s41586-026-10670-w},
pmid = {42343133},
pmcid = {PMC13441969}
}

RIS

TY - JOUR
AU - Fry, Benjamin
AU - Slaw, Kaia
AU - Polizzi, Nicholas F.
TI - Zero-shot design of drug-binding proteins via neural iterative selection-expansion
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/06/24
VL - 656
IS - 8126
SP - 237
EP - 249
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/s41586-026-10670-w
UR - https://doi.org/10.1038/s41586-026-10670-w
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41586-026-10670-w",
"type": "article-journal",
"title": "Zero-shot design of drug-binding proteins via neural iterative selection-expansion",
"container-title": "Nature",
"author": [
{
"family": "Fry",
"given": "Benjamin"
},
{
"family": "Slaw",
"given": "Kaia"
},
{
"family": "Polizzi",
"given": "Nicholas F."
}
],
"container-title-short": "Nature",
"volume": "656",
"issue": "8126",
"page": "237-249",
"DOI": "10.1038/s41586-026-10670-w",
"PMID": "42343133",
"PMCID": "PMC13441969",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41586-026-10670-w",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.molcel.2026.07.006 [code]
DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.
Journal: Molecular cell
In common: PyTorch Geometric, h5py, PyTorch, 6 other tools, cellular / molecular, 4 references
[2] doi:10.1093/nar/gkag706 [code]
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.
Journal: Nucleic acids research
In common: RDKit, PyTorch Geometric, h5py, 7 other tools, 1 reference
[3] doi:10.1021/acsomega.5c09368 [code]
Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.
Journal: ACS omega
In common: RDKit, PyTorch Geometric, Plotly, 7 other tools
[4] doi:10.1093/bib/bbag118 [code]
Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.
Journal: Briefings in bioinformatics
In common: RDKit, PyTorch Geometric, PyTorch, 6 other tools, computational modeling (no new data), cellular / molecular
[5] doi:10.1016/j.isci.2026.115554
Computationally guided discovery of Ly6e/LY6E-dependent AAV capsid variants.
Journal: iScience
In common: cellular / molecular, 7 references
[6] doi:10.1016/j.isci.2026.116055 [code]
Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.
Journal: iScience
In common: PyTorch Geometric, Plotly, h5py, 7 other tools
[7] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, cellular / molecular, 1 reference
[8] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, cellular / molecular, 1 reference
[9] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, computational modeling (no new data), cellular / molecular
[10] doi:10.1093/bioinformatics/btag153 [code]
MAISNet: a multi-species integrated graph neural network for acetylcholinesterase inhibitor screening.
Journal: Bioinformatics (Oxford, England)
In common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.