Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
The 13 matches
- [1] § LASErMPNN neural network ↔ train_lasermpnn_streptavidin_heldout_split.py, lines 695–738 · score 0.77 · pretrainable ligand encoder, dihedral angles, backbone noise, backbone atoms, idealized, head
- [2] § LASErMPNN neural network ↔ train_lasermpnn.py, lines 696–729 · score 0.75 · dihedral angles, backbone noise, ligand encoder, backbone atoms, idealized, head
- [3] § LASErMPNN neural network ↔ train_lasermpnn_streptavidin_heldout_split.py, lines 695–738 · score 0.75 · pretrained ligand encoder, dihedral angles, node embeddings, ligand information, backbone atoms, streptavidin
- [4] § NISE sampling algorithm ↔ carp_dock-pykeops.py, lines 3–18 · score 0.72 · brute force, ligand orientation, rigid body, docking, position, pose
- [5] § NISE sampling algorithm ↔ carp_dock.py, lines 3–30 · score 0.72 · brute force, ligand orientation, rigid body, docking, position, pose
- [6] § LASErMPNN neural network ↔ train_ligandmpnn_streptavidin_heldout_split.py, lines 575–661 · score 0.66 · dihedral angles, node embeddings, streptavidin, backbone atoms, pretrained, autoregressively
- [7] § LASErMPNN neural network ↔ run_nise_boltz2x.py, lines 555–630 · score 0.58 · helix bundle, structures predicted, binding site, objective, polar, LASErMPNN
- [8] § Characterization of exatecan binders ↔ run_nise_boltz2x.py, lines 280–396 · score 0.57 · structure ligand, designed proteins, ligand pLDDT, stem, pose, affinities
- [9] § Characterization of exatecan binders ↔ run_nise_boltz2x_ligandmpnn.py, lines 259–353 · score 0.57 · structure ligand, designed proteins, ligand pLDDT, stem, pose, affinities
- [10] § Design of apixaban binders using NISE ↔ run_nise_boltz2x.py, lines 280–396 · score 0.52 · affinity probability, binary, ligand pLDDT, score, poses, Boltz
- [11] § Design of apixaban binders using NISE ↔ run_nise_boltz2x_ligandmpnn.py, lines 259–353 · score 0.52 · affinity probability, binary, ligand pLDDT, score, poses, Boltz
- [12] § Design of apixaban binders using NISE ↔ carp_dock-pykeops.py, lines 3–18 · score 0.50 · ligand positions, orientations, clashing, docked, translations, rotations
- [13] § Design of apixaban binders using NISE ↔ carp_dock.py, lines 3–30 · score 0.50 · ligand positions, orientations, clashing, docked, translations, rotations
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 630 lines · 35 KB · MIT · 3 matches
- import io
- import os
- import sys
- import time
- import json
- import shutil
- import warnings
- import subprocess
- from typing import *
- from pathlib import Path
- from collections import defaultdict
- warnings.filterwarnings("ignore")
- # NOTE: If you move the run_nise*.py script, adjust the NISE_DIRECTORY_PATH
- NISE_DIRECTORY_PATH = str(Path(os.path.abspath(__file__)).parent)
- LASER_PATH = str(Path(NISE_DIRECTORY_PATH) / 'LASErMPNN')
- sys.path.append(NISE_DIRECTORY_PATH)
- import wandb
- import torch
- import plotly
- import numpy as np
- import prody as pr
- import pandas as pd
- from rdkit import Chem
- import plotly.express as px
- from rdkit.Chem import AllChem
- from rdkit_to_params import Params
- from LASErMPNN.run_inference import load_model_from_parameter_dict # type: ignore
- from LASErMPNN.run_batch_inference import _run_inference, output_protein_structure, output_ligand_structure # type: ignore
- from utility_scripts.burial_calc import compute_fast_ligand_burial_mask
- from utility_scripts.calc_symmetry_aware_rmsd import _main as calc_rmsd
- def compute_objective_function(confidence_metrics_dict: dict, objective_function: str) -> float:
- """
- Compute the objective function which will be MAXIMIZED by NISE.
- The input to this function is a dictionary of metrics extracted from Boltz-2x
- including design_ligand_plddt, iptm, and affinity metrics if running with
- boltz2_predict_affinity = True.
- The default behavior is to return design_ligand_plddt, but other metrics
- and combinations of metrics could be used here.
- """
- # NOTE: add your own objective function in another elif block here!
- if objective_function == 'ligand_plddt':
- return confidence_metrics_dict['design_ligand_plddt']
- elif objective_function == 'ligand_plddt_and_iptm':
- return confidence_metrics_dict['design_ligand_plddt'] + confidence_metrics_dict['iptm']
- elif objective_function == 'iptm':
- return confidence_metrics_dict['iptm']
- elif objective_function == 'pbind':
- return confidence_metrics_dict['affinity_probability_binary']
- elif objective_function == 'ligand_plddt_and_pbind':
- return confidence_metrics_dict['design_ligand_plddt'] + confidence_metrics_dict['affinity_probability_binary']
- elif objective_function == 'iptm_and_pbind':
- return confidence_metrics_dict['iptm'] + confidence_metrics_dict['affinity_probability_binary']
- else:
- raise ValueError(f'Objective_function strategy {objective_function} is not implemented.')
- def get_boltz_yaml_boilerplate(sequence: str, smiles: str, predict_affinity: bool):
- boltz_yaml = f"version: 1\nsequences:\n - protein:\n id: A\n sequence: {sequence}\n msa: empty\n - ligand:\n id: B\n smiles: '{smiles}'\n"
- if predict_affinity:
- boltz_yaml += "properties:\n - affinity:\n binder: B\n"
- return boltz_yaml
- def check_input(dir_path: Path, model_weights_path: Path):
- if not dir_path.exists():
- raise FileNotFoundError(f"Path {dir_path} does not exist.")
- if not dir_path.is_dir():
- raise NotADirectoryError(f"Path {dir_path} is not a directory.")
- num_files = len([x for x in dir_path.iterdir() if x.is_file() and x.suffix == '.pdb'])
- if num_files == 0:
- raise FileNotFoundError(f"No PDB files found in {dir_path}")
- if not model_weights_path.exists():
- raise FileNotFoundError(f"Model weights file {model_weights_path} does not exist.")
- def handle_directory_creation(input_dir, model_checkpoint):
- input_backbones_path = input_dir / 'input_backbones'
- check_input(input_backbones_path, model_checkpoint)
- sampling_dataframe_path = input_dir / 'sampling_dataframes'
- sampled_backbones_path = input_dir / 'sampled_backbones'
- sampling_dataframe_path.mkdir(exist_ok=True)
- sampled_backbones_path.mkdir(exist_ok=True)
- sdf_path = input_dir / 'lig_from_input_pdb.sdf'
- params_path = input_dir / 'lig_from_input_pdb.params'
- return input_backbones_path, sampling_dataframe_path, sampled_backbones_path, sdf_path, params_path
- def construct_helper_files(sdf_path, params_path, backbone_path, ligand_smiles):
- input_protein = pr.parsePDB(str(backbone_path))
- assert isinstance(input_protein, pr.AtomGroup), f"Error loading backbone {backbone_path}"
- ligand = input_protein.select('not protein').copy()
- assert ligand.numResidues() == 1, f"Error selecting ligand from backbone {backbone_path}"
- # Create a pdb string for the ligand.
- pdb_stream = io.StringIO()
- pr.writePDBStream(pdb_stream, ligand)
- ligand_string = pdb_stream.getvalue()
- lignames_set = set(ligand.getResnames())
- # Write ligand to PDB file.
- pdb_path = Path(str(sdf_path.resolve()).rsplit('.')[0] + '.pdb')
- with pdb_path.open('w') as f:
- f.write(ligand_string)
- pdb_mol = Chem.MolFromPDBBlock(ligand_string)
- smi_mol = Chem.MolFromSmiles(ligand_smiles)
- pdb_mol = AllChem.AssignBondOrdersFromTemplate(smi_mol, pdb_mol)
- pdb_mol = AllChem.AddHs(pdb_mol, addCoords=True)
- AllChem.ComputeGasteigerCharges(pdb_mol)
- Chem.MolToMolFile(pdb_mol, str(sdf_path.resolve()))
- ligname = lignames_set.pop()
- p = Params.from_mol(pdb_mol, name=ligname)
- p.dump(params_path) # type: ignore
- with open(Path(params_path).parent / '.gitignore', 'w') as f:
- f.write('*')
- class DesignCampaign:
- def __init__(self,
- model_checkpoint, input_dir, ligand_rmsd_mask_atoms, ligand_atoms_enforce_buried, ligand_atoms_enforce_exposed, laser_inference_device, debug, ligand_3lc,
- rmsd_use_chirality, self_consistency_ligand_rmsd_threshold, self_consistency_protein_rmsd_threshold,
- laser_inference_dropout, num_iterations, num_top_backbones_per_round, laser_sampling_params, sequences_sampled_per_backbone,
- sequences_sampled_at_once, boltz_inference_devices, ligand_smiles, boltz2x_executable_path,
- use_reduce_protonation, keep_input_backbone_in_queue, keep_best_generator_backbone, use_boltz_conformer_potentials,
- boltz2_predict_affinity, drop_rmsd_mask_atoms_from_ligand_plddt_calc, use_boltz_1x,
- boltz2_disable_kernels, boltz2_disable_nccl_p2p, objective_function, fixed_identity_residue_indices,
- align_on_binding_site, burial_mask_alpha_hull_alpha, boltz2_cache_directory, boltz2_sampling_steps, **kwargs
- ):
- self.debug = debug
- self.ligand_3lc = ligand_3lc
- self.boltz_inference_devices = boltz_inference_devices
- self.use_boltz_conformer_potentials = use_boltz_conformer_potentials
- self.use_boltz_1x = use_boltz_1x
- self.boltz2x_executable_path = boltz2x_executable_path
- self.predict_affinity = boltz2_predict_affinity
- self.use_reduce_protonation = use_reduce_protonation
- self.boltz2_cache_directory = boltz2_cache_directory
- self.boltz2_sampling_steps = boltz2_sampling_steps
- self.boltz2_disable_kernels = boltz2_disable_kernels
- self.boltz2_disable_nccl_p2p = boltz2_disable_nccl_p2p
- self.objective_function = objective_function
- self.align_on_binding_site = align_on_binding_site
- self.fixed_identity_residue_indices = fixed_identity_residue_indices
- if self.fixed_identity_residue_indices is not None:
- print(f'Fixing residues: {self.fixed_identity_residue_indices}')
- if self.use_boltz_1x and self.predict_affinity:
- raise ValueError('Cannot use boltz1x with affinity prediction.')
- if 'pbind' in self.objective_function and (not self.predict_affinity):
- raise ValueError(f"predict_affinity must be True to use objective function {self.objective_function}")
- self.rmsd_use_chirality = rmsd_use_chirality
- self.self_consistency_ligand_rmsd_threshold = self_consistency_ligand_rmsd_threshold
- self.self_consistency_protein_rmsd_threshold = self_consistency_protein_rmsd_threshold
- self.keep_input_backbone_in_queue = keep_input_backbone_in_queue
- self.drop_rmsd_mask_atoms_from_ligand_plddt_calc = drop_rmsd_mask_atoms_from_ligand_plddt_calc
- self.num_iterations = num_iterations
- self.sequences_sampled_per_backbone = sequences_sampled_per_backbone
- self.sequences_sampled_at_once = sequences_sampled_at_once
- self.top_k = num_top_backbones_per_round
- self.ligand_rmsd_mask_atoms = ligand_rmsd_mask_atoms
- self.ligand_atoms_enforce_buried = ligand_atoms_enforce_buried
- self.ligand_atoms_enforce_exposed = ligand_atoms_enforce_exposed
- self.burial_mask_alpha_hull_alpha = burial_mask_alpha_hull_alpha
- self.ligand_smiles = ligand_smiles
- self.laser_sampling_params = laser_sampling_params
- self.laser_inference_dropout = laser_inference_dropout
- self.sampling_metadata = defaultdict(float)
- self.sampling_metadata['min_protein_rmsd'] = float('inf')
- self.sampling_metadata['min_ligand_rmsd'] = float('inf')
- self.input_backbones_path, self.sampling_dataframe_path, self.sampled_backbones_path, self.sdf_path, self.params_path = handle_directory_creation(input_dir, model_checkpoint)
- self.model, self.model_params = load_model_from_parameter_dict(model_checkpoint, torch.device(laser_inference_device))
- # Set model to eval mode, enable inference dropout if specified.
- self.model.eval()
- if self.laser_inference_dropout:
- for module in self.model.modules():
- if isinstance(module, torch.nn.Dropout):
- module.train()
- # Set path and priority for input backbones.
- if self.keep_input_backbone_in_queue:
- self.backbone_queue = [(x, torch.inf) for x in self.input_backbones_path.iterdir() if x.is_file() and x.suffix == '.pdb']
- else:
- self.backbone_queue = [(x, 0) for x in self.input_backbones_path.iterdir() if x.is_file() and x.suffix == '.pdb']
- # Whether to keep the pose that generates the highest confidence sequences.
- self.keep_best_generator_backbone = keep_best_generator_backbone
- self.backbone_to_best_generation = defaultdict(float)
- if len(self.backbone_queue) > 1:
- raise NotImplementedError(f"More than one input backbone not currently supported.")
- construct_helper_files(self.sdf_path, self.params_path, self.backbone_queue[0][0], ligand_smiles)
- self.ligand_rmsd_mask_atoms = ligand_rmsd_mask_atoms
- assert self.sdf_path.exists(), f"Error creating SDF file {self.sdf_path}"
- assert self.params_path.exists(), f"Error creating params file {self.params_path}"
- def sample_sequences(self, backbone_path: str) -> Tuple[List[pr.AtomGroup], List[str], List[float], List[float]]:
- laser_nll, laser_bs_nll = [], []
- sampled_proteins, sampled_sequences = [], []
- num_sampled = 0
- while num_sampled < self.sequences_sampled_per_backbone:
- remaining_samples = self.sequences_sampled_per_backbone - num_sampled
- num_seq_to_sample = min(remaining_samples, self.sequences_sampled_at_once)
- backbone_path_ = backbone_path
- if self.fixed_identity_residue_indices is not None:
- # NOTE: there might be a better way to do this so that we don't lose information stored in the b-factor columns of backbones that we sample,
- # but the important pLDDT information should be recorded in the log steps anyways.
- protein = pr.parsePDB(str(backbone_path))
- protein.setBetas(0.0)
- fixed_selection = protein.select(self.fixed_identity_residue_indices)
- if fixed_selection is None:
- raise ValueError(f'ProDy encounted an error selecting residues with selection string: {self.fixed_identity_residue_indices}')
- fixed_selection.setBetas(1.0)
- backbone_path_ = Path(backbone_path).parent / f'{Path(backbone_path).stem}_fixbeta.pdb'
- pr.writePDB(str(backbone_path_), protein)
- self.laser_sampling_params.update({'fix_beta': True})
- sampling_output, full_atom_coords, nh_coords, sampled_probs, batch_data, protein_complex_data = _run_inference(self.model, self.model_params, Path(backbone_path_), num_seq_to_sample, **self.laser_sampling_params)
- protein_complex_data = protein_complex_data[0] # type: ignore
- for idx in range(num_seq_to_sample):
- curr_batch_mask = batch_data.batch_indices == idx
- curr_probs = sampled_probs[curr_batch_mask]
- nll = (-1 * torch.log10(curr_probs)).cpu().numpy().mean()
- bs_nll = (-1 * torch.log10(sampled_probs[curr_batch_mask][batch_data.first_shell_ligand_contact_mask[curr_batch_mask]])).cpu().numpy().mean()
- out_prot = output_protein_structure(full_atom_coords[curr_batch_mask], sampling_output.sampled_sequence_indices[curr_batch_mask], protein_complex_data.residue_identifiers, nh_coords[curr_batch_mask], curr_probs)
- out_lig = output_ligand_structure(protein_complex_data.ligand_info)
- out_complex = out_prot + out_lig
- out_complex.setTitle('LASErMPNN/NISE Generated Protein')
- sampled_proteins.append(out_complex)
- sampled_sequences.append(out_prot.ca.getSequence())
- laser_nll.append(nll)
- laser_bs_nll.append(bs_nll)
- num_sampled += num_seq_to_sample
- return sampled_proteins, sampled_sequences, laser_nll, laser_bs_nll
- def identify_backbone_candidates(self, sorted_designs_boltz: Sequence[Path], sorted_designs_laser: Sequence[Path], sampled_backbone_paths: Sequence[Path], reduce_executable_path, reduce_hetdict_path):
- assert len(sorted_designs_laser) == len(sorted_designs_boltz), f"Error: different number of designs in" # {boltz_output_subdir} and {laser_output_subdir}."
- smi_mol = Chem.MolFromSmiles(self.ligand_smiles)
- log_data = defaultdict(list)
- for laser, boltz, bb_path in zip(sorted_designs_laser, sorted_designs_boltz, sampled_backbone_paths):
- laser_prot = pr.parsePDB(str(laser))
- boltz_string_str = open(boltz, 'r').read()
- boltz_string_io = io.StringIO(boltz_string_str)
- boltz_prot = pr.parsePDBStream(boltz_string_io)
- confidence_data = {'laser_output_pdb_path': laser, 'boltz_output_pdb_path': boltz}
- with (boltz.parent / f'confidence_{boltz.stem}.json').open('r') as f:
- confidence_data.update(json.load(f))
- design_iptm = confidence_data['iptm']
- design_bind_probability, design_predicted_affinity = torch.nan, torch.nan
- if self.predict_affinity:
- with (boltz.parent / f'affinity_{boltz.stem.replace("_model_0", "")}.json').open('r') as f:
- confidence_data.update(json.load(f))
- design_bind_probability = confidence_data['affinity_probability_binary']
- design_predicted_affinity = confidence_data['affinity_pred_value']
- log_data['affinity_probability_binary'].append(design_bind_probability)
- log_data['affinity_pred_value'].append(design_predicted_affinity)
- log_data['iptms'].append(design_iptm)
- try:
- protein_rmsd, ligand_rmsd, laser_to_boltz_name_mapping = calc_rmsd(laser_prot, boltz_prot, self.ligand_smiles, self.rmsd_use_chirality, self.ligand_rmsd_mask_atoms, align_on_binding_site=self.align_on_binding_site)
- if len(laser_to_boltz_name_mapping) == 0:
- raise ValueError('No atoms in common between laser and boltz structures.')
- boltz_to_laser_name_mapping = {v: k for k, v in laser_to_boltz_name_mapping.items()}
- except:
- protein_rmsd = np.nan
- ligand_rmsd = np.nan
- laser_to_boltz_name_mapping = {}
- log_data['ligand_rmsds'].append(ligand_rmsd)
- log_data['protein_rmsds'].append(protein_rmsd)
- if len(laser_to_boltz_name_mapping) == 0:
- print(f'{boltz}: Failed to map names between laser and boltz structures.')
- log_data['ligand_is_buried'].append(False)
- log_data['ligand_plddts'].append(torch.nan)
- log_data['protein_plddts'].append(torch.nan)
- continue
- # Remap the boltz structure ligand atoms with the name mapping.
- boltz_prot_only = boltz_prot.select('chid A')
- boltz_lig_only = boltz_prot.select('chid B and not element H')
- boltz_lig_only.setNames([boltz_to_laser_name_mapping[x] for x in boltz_lig_only.getNames()])
- boltz_lig_only.setResnames([self.ligand_3lc for _ in range(len(boltz_lig_only.getResnames()))])
- boltz_coords = boltz_lig_only.getCoords()
- atoms_enforced_buried_mask = np.array([x in self.ligand_atoms_enforce_buried for x in boltz_lig_only.getNames()])
- atoms_enforced_exposed_mask = np.array([x in self.ligand_atoms_enforce_exposed for x in boltz_lig_only.getNames()])
- pdb_output_path = str(boltz)
- pr.writePDB(pdb_output_path, boltz_prot_only + boltz_lig_only)
- # Compute ligand pLDDT over relevant atoms.
- rmsd_mask = np.array([x not in self.ligand_rmsd_mask_atoms for x in boltz_lig_only.getNames()])
- if not self.drop_rmsd_mask_atoms_from_ligand_plddt_calc:
- rmsd_mask = np.ones_like(rmsd_mask)
- design_ligand_plddt = boltz_lig_only.getBetas()[rmsd_mask].mean() / 100
- design_protein_plddt = boltz_prot_only.copy().ca.getBetas().mean() / 100
- log_data['ligand_plddts'].append(design_ligand_plddt)
- log_data['protein_plddts'].append(design_protein_plddt)
- confidence_data['design_ligand_plddt'] = design_ligand_plddt
- confidence_data['design_protein_plddt'] = design_protein_plddt
- # Check ligand burial constraints are obeyed in predicted structure.
- all_buried_mask = compute_fast_ligand_burial_mask(boltz_prot.ca.getCoords(), boltz_coords[atoms_enforced_buried_mask], num_rays=5, alpha=self.burial_mask_alpha_hull_alpha)
- none_buried_mask = compute_fast_ligand_burial_mask(boltz_prot.ca.getCoords(), boltz_coords[atoms_enforced_exposed_mask], num_rays=5, alpha=self.burial_mask_alpha_hull_alpha)
- if (all_buried_mask.all().item() and (not none_buried_mask.any().item())) or self.debug:
- log_data['ligand_is_buried'].append(True)
- if (ligand_rmsd < self.self_consistency_ligand_rmsd_threshold and protein_rmsd < self.self_consistency_protein_rmsd_threshold) or self.debug:
- if self.use_reduce_protonation:
- subprocess.run(
- f'{reduce_executable_path} -DB {reduce_hetdict_path} -DROP_HYDROGENS_ON_ATOM_RECORDS -BUILD {pdb_output_path} > {pdb_output_path}_',
- shell=True, check=False,
- stdout=subprocess.DEVNULL if not self.debug else subprocess.PIPE,
- stderr=subprocess.DEVNULL if not self.debug else subprocess.PIPE
- )
- shutil.move(f'{pdb_output_path}_', pdb_output_path)
- else:
- # Protonate using RDKit
- ligand_string = io.StringIO()
- pr.writePDBStream(ligand_string, boltz_lig_only.copy())
- pdb_mol = Chem.MolFromPDBBlock(ligand_string.getvalue())
- pdb_mol = AllChem.AssignBondOrdersFromTemplate(smi_mol, pdb_mol)
- pdb_mol = AllChem.AddHs(pdb_mol, addCoords=True)
- ligand_prody = pr.parsePDBStream(io.StringIO(Chem.MolToPDBBlock(pdb_mol)))
- ligand_prody.setResnames(self.ligand_3lc)
- ligand_prody.setChids('B')
- pr.writePDB(pdb_output_path, boltz_prot_only.copy() + ligand_prody)
- score = compute_objective_function(confidence_data, self.objective_function)
- self.backbone_to_best_generation[bb_path] = max(score, self.backbone_to_best_generation[bb_path])
- self.backbone_queue.append((pdb_output_path, score))
- else:
- log_data['ligand_is_buried'].append(False)
- self.backbone_queue = sorted(self.backbone_queue, key=lambda x: float(x[1]), reverse=True)[:self.top_k]
- # Adds the backbone which has generated the best scoring pose if it's not in the queue already.
- if self.keep_best_generator_backbone and len(self.backbone_to_best_generation) > 0:
- top_backbone = max(self.backbone_to_best_generation.items(), key=lambda x: x[1])
- print(top_backbone)
- if not (top_backbone[0] in [x[0] for x in self.backbone_queue]):
- self.backbone_queue = self.backbone_queue[:-1] + [top_backbone]
- return sorted_designs_laser, sorted_designs_boltz, log_data
- def log(self, use_wandb: bool, dataframe: pd.DataFrame):
- self.sampling_metadata['min_protein_rmsd'] = min(self.sampling_metadata['min_protein_rmsd'], dataframe['protein_rmsds'].dropna().min())
- self.sampling_metadata['min_ligand_rmsd'] = min(self.sampling_metadata['min_ligand_rmsd'], dataframe['ligand_rmsds'].dropna().min())
- logs = dict(self.sampling_metadata)
- logs['mean_sampled_ligand_plddt'] = dataframe['ligand_plddts'].dropna().mean()
- logs['mean_sampled_protein_plddt'] = dataframe['protein_plddts'].dropna().mean()
- logs['mean_sampled_protein_rmsd'] = dataframe['protein_rmsds'].dropna().mean()
- logs['mean_sampled_ligand_rmsd'] = dataframe['ligand_rmsds'].dropna().mean()
- logs['num_sequences_sampled'] = len(dataframe)
- logs['mean_affinity_probability'] = dataframe['affinity_probability_binary'].mean()
- logs['mean_affinity_pred_value'] = dataframe['affinity_pred_value'].mean()
- logs['mean_iptm'] = dataframe['iptms'].mean()
- logs['mean_laser_score'] = dataframe['laser_nll'].mean()
- logs['mean_laser_bs_score'] = dataframe['laser_bs_nll'].mean()
- try:
- if self.keep_input_backbone_in_queue:
- logs['max_sampled_backbone_priority'] = max(self.backbone_queue[1:], key=lambda x: x[1])[1]
- else:
- logs['max_sampled_backbone_priority'] = max(self.backbone_queue, key=lambda x: x[1])[1]
- except:
- logs['max_sampled_backbone_priority'] = 0
- print(logs)
- if use_wandb:
- scatter_fig = px.scatter(dataframe, x='ligand_rmsds', y='ligand_plddts', color='protein_rmsds', hover_data=['protein_rmsds', 'sequences'], range_color=[0.0, 2.5], range_y=[0.0, 1.0], range_x=[0.0, 5.0])
- logs['ligand_RMSD_vs_pLDDT_scatter'] = wandb.Html(plotly.io.to_html(scatter_fig)) # type: ignore
- wandb.log(logs)
- def predict_complex_structures(
- boltz_inputs_dir, boltz2x_executable_path, boltz_inference_devices,
- boltz_output_dir, use_potentials, use_boltz_1x, disable_kernels, disable_nccl_p2p, boltz2_cache_directory, boltz2_sampling_steps, debug
- ):
- device_ints = [x.split(':')[-1] for x in boltz_inference_devices]
- command = f'{boltz2x_executable_path} predict {boltz_inputs_dir} --devices {len(device_ints)} --out_dir {boltz_output_dir} --output_format pdb --override --sampling_steps {int(boltz2_sampling_steps)} --sampling_steps_affinity {int(boltz2_sampling_steps)}'
- command = f'CUDA_VISIBLE_DEVICES={",".join(device_ints)} {command}'
- if disable_nccl_p2p:
- command = f'NCCL_P2P_DISABLE=1 {command}'
- if use_potentials:
- command += f' --use_potentials'
- if use_boltz_1x:
- command += f' --model boltz1'
- if disable_kernels:
- command += f' --no_kernels'
- if boltz2_cache_directory is not None:
- command += f' --cache {boltz2_cache_directory}'
- print(command)
- try:
- # Boltz sometimes completes with a nonzero exit code despite completing successfully.
- # If not all expected files were generated NISE will crash at the log step.
- subprocess.run(command, shell=True, check=False, stdout=subprocess.DEVNULL if not debug else None, stderr=subprocess.DEVNULL if not debug else None)
- except:
- print('Boltz crashed! This might be fine, trying to recover...')
- pass
- def main(use_wandb, reduce_executable_path, reduce_hetdict_path, **kwargs):
- design_campaign = DesignCampaign(**kwargs)
- for iidx in range(design_campaign.num_iterations):
- # Run laser on all backbone queue inputs.
- all_sampled_proteins, backbone_sample_indices, sampled_backbone_path, all_sampled_sequences = [], [], [], []
- laser_nll, laser_bs_nll = [], []
- for bidx, (backbone_path, score) in enumerate(design_campaign.backbone_queue):
- sampled_proteins, sampled_sequences, nlls, bs_nlls = design_campaign.sample_sequences(backbone_path)
- all_sampled_proteins.extend(sampled_proteins)
- backbone_sample_indices.extend([bidx] * len(sampled_proteins))
- sampled_backbone_path.extend([design_campaign.backbone_queue[bidx][0]] * len(sampled_proteins))
- all_sampled_sequences.extend(sampled_sequences)
- laser_nll.extend(nlls)
- laser_bs_nll.extend(bs_nlls)
- # Sanity check output shapes before trying to fold or log anything.
- assert len(all_sampled_proteins) == len(backbone_sample_indices), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(backbone_sample_indices)}'
- assert len(all_sampled_proteins) == len(laser_nll), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(laser_nll)}'
- assert len(all_sampled_proteins) == len(laser_bs_nll), f'Error in laser sampling: {len(all_sampled_proteins)} != {len(laser_bs_nll)}'
- # Make subdirectories to write outputs to disk.
- sampling_subdir = design_campaign.sampled_backbones_path / f'iter_{iidx}'
- laser_output_subdir = sampling_subdir / 'laser_outputs'
- boltz_input_dir = sampling_subdir / 'boltz_inputs'
- laser_output_subdir.mkdir(exist_ok=True, parents=True)
- boltz_input_dir.mkdir(exist_ok=True)
- # Write all the boltz input directories.
- all_boltz_input_path_names = []
- sampled_sequence_chunks = np.array_split(all_sampled_sequences, len(design_campaign.boltz_inference_devices))
- for idx, chunk_sequences in enumerate(sampled_sequence_chunks):
- for seq_idx, seq in enumerate(chunk_sequences):
- boltz_input_output_path = boltz_input_dir / f'chunk_{idx}_seq_{seq_idx}.yaml'
- all_boltz_input_path_names.append(boltz_input_output_path.stem)
- with boltz_input_output_path.open('w') as f:
- f.write(get_boltz_yaml_boilerplate(seq, design_campaign.ligand_smiles, design_campaign.predict_affinity))
- all_boltz_model_paths = [(sampling_subdir / 'boltz_results_boltz_inputs' / 'predictions' / x / f'{x}_model_0.pdb') for x in all_boltz_input_path_names]
- all_laser_output_paths = []
- for laser_output_structure, boltz_output_path in zip(all_sampled_proteins, all_boltz_model_paths):
- laser_output_path = laser_output_subdir / f'laser_{boltz_output_path.parent.stem}.pdb'
- pr.writePDB(str(laser_output_path), laser_output_structure)
- all_laser_output_paths.append(laser_output_path)
- curr_tries = 0
- max_tries = 10
- while curr_tries < max_tries and not all([x.exists() for x in all_boltz_model_paths]):
- if curr_tries != 0:
- print('Not all boltz predictions were completed or file system not updated.. retrying...')
- time.sleep(30)
- predict_complex_structures(
- boltz_input_dir, design_campaign.boltz2x_executable_path, design_campaign.boltz_inference_devices,
- sampling_subdir, design_campaign.use_boltz_conformer_potentials, design_campaign.use_boltz_1x,
- design_campaign.boltz2_disable_kernels, design_campaign.boltz2_disable_nccl_p2p, design_campaign.boltz2_cache_directory, design_campaign.boltz2_sampling_steps, design_campaign.debug
- )
- curr_tries += 1
- assert all([x.exists() for x in all_boltz_model_paths]), f"Error: not all boltz predictions were written to disk."
- # Identify any new backbone candidates.
- sorted_designs_laser, sorted_designs_rosetta, log_data = design_campaign.identify_backbone_candidates(all_boltz_model_paths, all_laser_output_paths, sampled_backbone_path, reduce_executable_path, reduce_hetdict_path)
- with open(sampling_subdir / 'backbone_queue.txt', 'w') as f:
- f.write('\n'.join([f"{x[1]}\t{x[0]}" for x in design_campaign.backbone_queue]))
- iidx_data = {
- 'sampled_backbone_path': sampled_backbone_path,
- 'laser_nll': laser_nll,
- 'laser_bs_nll': laser_bs_nll,
- 'laser_paths': sorted_designs_laser,
- 'rosetta_paths': sorted_designs_rosetta,
- 'sequences': all_sampled_sequences,
- 'curr_idx': [iidx] * len(laser_nll),
- }
- iidx_data.update(dict(log_data))
- # Write data to disk.
- iidx_dataframe = pd.DataFrame(iidx_data)
- iidx_dataframe.to_pickle(design_campaign.sampling_dataframe_path / f'iter_{iidx}_data.pkl')
- design_campaign.log(use_wandb, iidx_dataframe)
- if __name__ == "__main__":
- laser_sampling_params = {
- 'sequence_temp': 0.5, 'first_shell_sequence_temp': 0.7,
- 'chi_temp': 1e-6, 'seq_min_p': 0.0, 'chi_min_p': 0.0,
- 'disable_pbar': True, 'disabled_residues_list': ['X', 'C'], # Disables cysteine sampling by default.
- # ====================================================================================================
- # If constrain_ala_gly_sampling_to_exposed_non_secondary_structure is True (recommended),
- # the ala_budget and gly_budget parameters are
- # used to constrain the sampling of ALA and GLY residues to exposed non-secondary structured residues.
- # You can override the specific residues with the budget_residue_sele_string parameter (not recommended).
- # If constrain_ala_gly_sampling_to_exposed_non_secondary_structure is False and budget_residue_sele_string is None,
- # no constraints are applied to the sampling of ALA and GLY residues.
- # The reason we suggest doing this is to constrain the generated sequences to the manifold of sequences that are likely to exist in nature
- # and not exploiting a propensity of structure prediction networks to predict structure in alanine rich sequences
- # ====================================================================================================
- 'constrain_ala_gly_sampling_to_exposed_non_secondary_structure': True,
- 'budget_residue_sele_string': None,
- 'ala_budget': 4, 'gly_budget': 0, # May sample up to 4 Ala and 0 Gly over the selected region if not None.
- 'disable_charged_fs': True, # Disables sampling D,E,K,R residues for buried residues around the ligand.
- }
- params = dict(
- debug = (debug := False),
- use_wandb = (use_wandb := False),
- input_dir = Path('./debug/').resolve(),
- ligand_3lc = 'GG2', # Should match CCD code if using reduce.
- ligand_rmsd_mask_atoms = set(), # Atoms to IGNORE in RMSD calculation.
- ligand_atoms_enforce_buried = set(), # Atoms to enforce remain buried inside convex hull when selecting new backbones.
- ligand_atoms_enforce_exposed = set(), # Atoms to enforce remain exposed relative to the convex hull when selecting new backbones. I would suggest only using this for linker regions attached to your ligand or clearly exposed charged polar groups.
- laser_sampling_params = laser_sampling_params,
- ligand_smiles = 'COC1=CC=C(C=C1)N2C3=C(CCN(C3=O)C4=CC=C(C=C4)N5CCCCC5=O)C(=N2)C(=O)N',
- objective_function = (objective_function := 'ligand_plddt'), # Current options: {'ligand_plddt', 'iptm', 'ligand_plddt_and_iptm', 'pbind', 'ligand_plddt_and_pbind', 'iptm_and_pbind'}, Check the top of the file for implemented strategies, if you find an alternative strategy to work well please make a git commit so others can test it out as well!
- drop_rmsd_mask_atoms_from_ligand_plddt_calc = True,
- keep_input_backbone_in_queue = False,
- keep_best_generator_backbone = True, # The highest scoring pose may not necessarily generate higher scoring poses, keeps the pose that has generated the best poses after the first iteration in the queue if not already the best scoring pose.
- rmsd_use_chirality = False, # Will fail to compute RMSD on mismatched chirality ligands, might be bugged...
- self_consistency_ligand_rmsd_threshold = 2.5,
- self_consistency_protein_rmsd_threshold = 2.5,
- align_on_binding_site = False, # If align on binding site is True, protein RMSD (and self_consistency_protein_rmsd_threshold above) becomes binding site RMSD. Useful if you have a large protein with floppy regions away from the ligand. Binding site is computed as residues with sidechain atoms within 5.0A of the ligand in the lasermpnn output structure.
- fixed_identity_residue_indices = None, # An optional prody selection string of the form "resindex 0 2 3 4 5" or "resnum 1 3 5"...
- use_reduce_protonation = False, # If False, will use RDKit to protonate, these hydrogens will not preserve the input names and aren't placed conditioned on the sidechains but REDUCE sometimes drops hydrogens if geometry changes outside of expected bounds.
- reduce_hetdict_path = Path('./modified_hetdict.txt').absolute(), # Can set to None if use_reduce_protonation False
- reduce_executable_path = None, # Can set to None if use_reduce_protonation False
- model_checkpoint = Path(LASER_PATH) / 'model_weights/laser_weights_0p1A_nothing_heldout.pt',
- num_iterations = 35,
- num_top_backbones_per_round = 3,
- sequences_sampled_at_once = 30,
- boltz2x_executable_path = str((Path(NISE_DIRECTORY_PATH) / '.venv/bin/boltz').absolute()),
- boltz2_cache_directory = None, # Optional path to the boltz weights, can be used to avoid redownloading weights that have already been cached on your machine not in the default location.
- boltz2_sampling_steps = 200,
- boltz_inference_devices = (boltz_inference_devices := ['cuda:0',]), # a list of multiple torch-style device strings
- use_boltz_conformer_potentials = True, # Use Boltz-x mode, this is almost always better.
- boltz2_predict_affinity = True if ('pbind' in objective_function) else False,
- use_boltz_1x = False, # Run the same script using --model boltz-1, multi-device inference with this seems bugged with boltz v2.1.1
- boltz2_disable_kernels = False, # Disables cuEquivariance kernels, this is likely not necessary.
- boltz2_disable_nccl_p2p = False, # On some systems with certain graphics cards, NCCL can hang indefinitely. This flag fixes this issue allowing running boltz / NISE with multiple GPUs. https://github.com/NVIDIA/nccl/issues/631
- sequences_sampled_per_backbone = 64 if not debug else 1 * len(boltz_inference_devices),
- burial_mask_alpha_hull_alpha = 9.0, # Set to a larger number for folds with wider pockets (ex: 7-helix bundle) (Ex: 100.0), see https://github.com/benf549/CARPdock/blob/main/visualize_hull.ipynb
- laser_inference_device = boltz_inference_devices[0],
- laser_inference_dropout = True,
- )
- if use_wandb:
- wandb.init(project='design-campaigns', entity='benf549', config=params)
- main(**params)
run_nise_boltz2x.py at commit 61b7500, under MIT · at the source
Overview
- Harvard Graduate Program in Biophysics, Harvard University,Boston, MA USA
- Department of Cancer Biology, Dana-Farber Cancer Institute,Boston, MA USA
- Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School,Boston, MA USA
Abstract
The design of proteins that bind to small molecules has been challenging because it requires simultaneous optimization of the protein sequence, protein structure and ligand conformation1–7. Current deep-learning algorithms have struggled to navigate this landscape, precluding the zero-shot design of binders. Here we show that by combining two neural networks in an iterative design algorithm, small-molecule binding proteins can be created from scratch with high accuracy. We trained a graph neural network—ligand-aware sequence engineering message-passing neural network (LASErMPNN)—to design compatible protein sequences for an input protein backbone and docked ligand. We paired LASErMPNN with a structure predictor that models a three-dimensional protein–ligand complex for an input protein sequence and ligand identity. The closed-loop iteration of these reciprocal networks optimized sequence–structure–ligan
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.
polizzilab/LASErMPNN
e70f2c6d765416f7e29d51bfd6d4e08496438878, 4 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
36 files
- download_ligand_encoder_
training_dataset.sh , Shell, 9 lines - download_protonated_pdb_
training_dataset.sh , Shell, 9 lines - pretrain_ligand_encoder.
py , Python, 239 lines - run_batch_inference.py, Python, 411 lines
- run_batch_inference_liga
ndmpnn.py , Python, 355 lines - run_inference.py, Python, 814 lines
- run_inference_ligandmpnn
.py , Python, 812 lines - run_inference_tied.py, Python, 926 lines
- run_lasermpnn.ipynb, Jupyter, 400 lines
- run_predict_partial_char
ges.py , Python, 84 lines - run_proofreading.py, Python, 237 lines
- scripts/
thread_and_score_sequenc , Python, 182 lineses.py - tests/
test_equivariance.py , Python, 239 lines - tests/
test_inference.py , Python, 26 lines - tests/
test_model_generics.py , Python, 41 lines - train_lasermpnn.py, Python, 812 lines, 1 match
- train_lasermpnn_streptav
idin_heldout_split.py , Python, 811 lines, 2 matches - train_ligandmpnn.py, Python, 668 lines
- train_ligandmpnn_strepta
vidin_heldout_split.py , Python, 662 lines, 1 match - utils/
__init__.py , Python, 1 line - utils/
build_rotamers.py , Python, 603 lines - utils/
burial_calc.py , Python, 381 lines - utils/
constants.py , Python, 283 lines - utils/
hbond_network.py , Python, 460 lines - utils/
helper_functions.py , Python, 199 lines - utils/
ligand_featurization.py , Python, 115 lines - utils/
ligand_featurization_lig , Python, 134 linesandmpnn.py - utils/
model.py , Python, 1,361 lines - utils/
model_generics.py , Python, 656 lines - utils/
model_ligandmpnn.py , Python, 539 lines - utils/
optimizer.py , Python, 50 lines - utils/
pdb_dataset.py , Python, 1,935 lines - utils/
pdb_dataset_ligandmpnn.p , Python, 1,617 linesy - utils/
spice_dataset.py , Python, 372 lines - LICENSE, License, 13 lines
- README.md, Text, 208 lines
polizzilab/nise
61b7500e99ab37295ae9510b3db05aefe6a5fdfb, 11 August 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
15 files
- NISE_LASErMPNN.ipynb, Jupyter, 30 lines
- build_container.sh, Shell, 39 lines
- identify_surface_residue
s.ipynb , Jupyter, 102 lines - inject_ligand_into_hetdi
ct.py , Python, 121 lines - protonate_and_add_conect
_records.py , Python, 88 lines - run_nise_boltz1x.py, Python, 513 lines
- run_nise_boltz2x.py, Python, 630 lines, 3 matches
- run_nise_boltz2x_cli.py, Python, 625 lines
- run_nise_boltz2x_ligandm
pnn.py , Python, 560 lines, 2 matches - setup.py, Python, 106 lines
- setup.sh, Shell, 16 lines
- utility_scripts/
burial_calc.py , Python, 317 lines - utility_scripts/
calc_symmetry_aware_rmsd , Python, 164 lines.py - LICENSE, License, 13 lines
- README.md, Text, 215 lines
benf549/CARPdock
9e11474052a0934444acfb5392e56b67f1a12c58, 18 February 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
8 files
- carp_dock-pykeops.py, Python, 470 lines, 2 matches
- carp_dock.py, Python, 371 lines, 2 matches
- generate_initial_ligand_
conformer.py , Python, 42 lines - run_CARPdock.ipynb, Jupyter, 44 lines
- utils/
burial_calc.py , Python, 445 lines - visualize_hull.ipynb, Jupyter, 20 lines
- LICENSE, License, 21 lines
- README.md, Text, 106 lines
Zenodo 18308430
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
Code availability
All model inference code, training code, processed datasets and dataset split information can be found at GitHub (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 53 scripts, each with its path and the digest of its content;
- 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- zenodo:10975225, at Zenodo; found in the references
- zenodo:17990180, at Zenodo; found in the references
Data availability
Coordinates and data files for the X-ray crystal structures of exatecan-bound EPIC and exatecan-bound EPIC(Q51N) have been deposited in the PDB with accession codes 9NZE (http://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Nature Portfolio
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 2 keywords, 11 MeSH terms, 58 references.
Cite
This paper
Fry, B., Slaw, K., & Polizzi, N. F. (2026). Zero-shot design of drug-binding proteins via neural iterative selection-expansion. Nature, 656(8126), 237-249. https://
BibTeX
@article{fry2026zero,
author = {Fry, Benjamin and Slaw, Kaia and Polizzi, Nicholas F.},
title = {{Zero-shot design of drug-binding proteins via neural iterative selection-expansion}},
journal = {Nature},
year = {2026},
month = jun,
volume = {656},
number = {8126},
pages = {237--249},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/
url = {https://
pmid = {42343133},
pmcid = {PMC13441969}
}
RIS
TY - JOUR
AU - Fry, Benjamin
AU - Slaw, Kaia
AU - Polizzi, Nicholas F.
TI - Zero-shot design of drug-binding proteins via neural iterative selection-expansion
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/
VL - 656
IS - 8126
SP - 237
EP - 249
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Zero-shot design of drug-binding proteins via neural iterative selection-expansion",
"container-title": "Nature",
"author": [
{
"family": "Fry",
"given": "Benjamin"
},
{
"family": "Slaw",
"given": "Kaia"
},
{
"family": "Polizzi",
"given": "Nicholas F."
}
],
"container-title-short":
"volume": "656",
"issue": "8126",
"page": "237-249",
"DOI": "10.1038/
"PMID": "42343133",
"PMCID": "PMC13441969",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.molcel.2026.07.006 [code]
- DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.Journal: Molecular cellIn common: PyTorch Geometric, h5py, PyTorch, 6 other tools, cellular / molecular, 4 references
- [2] doi:10.1093/nar/gkag706 [code]
- scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data.Journal: Nucleic acids researchIn common: RDKit, PyTorch Geometric, h5py, 7 other tools, 1 reference
- [3] doi:10.1021/acsomega.5c09368 [code]
- Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.Journal: ACS omegaIn common: RDKit, PyTorch Geometric, Plotly, 7 other tools
- [4] doi:10.1093/bib/bbag118 [code]
- Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.Journal: Briefings in bioinformaticsIn common: RDKit, PyTorch Geometric, PyTorch, 6 other tools, computational modeling (no new data), cellular / molecular
- [5] doi:10.1016/j.isci.2026.115554
- Computationally guided discovery of Ly6e/
LY6E-dependent AAV capsid variants. Journal: iScienceIn common: cellular / molecular, 7 references - [6] doi:10.1016/j.isci.2026.116055 [code]
- Mapping the transcriptional diversity of calcium signaling in the mouse and human brain.Journal: iScienceIn common: PyTorch Geometric, Plotly, h5py, 7 other tools
- [7] doi:10.1038/s41586-026-10391-0 [code]
- Cell-type-targeted mitochondrial transplantation rescues cell degeneration.Journal: NatureIn common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, cellular / molecular, 1 reference
- [8] doi:10.1021/acs.biochem.5c00596 [code]
- Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.Journal: BiochemistryIn common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, cellular / molecular, 1 reference
- [9] doi:10.1038/s41598-026-53415-5 [code]
- Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.Journal: Scientific reportsIn common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, computational modeling (no new data), cellular / molecular
- [10] doi:10.1093/bioinformatics/btag153 [code]
- MAISNet: a multi-species integrated graph neural network for acetylcholinesterase inhibitor screening.Journal: Bioinformatics (Oxford, England)In common: RDKit, PyTorch Geometric, PyTorch, 5 other tools, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 4 repositories of the authors' code, each at its verified commit and with its license, 53 scripts, and 13 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:ce409203059de803…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
