OSCR

DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.

Code ↔ Paper

15 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 15 matches
  1. [1] § Star★Methods › Method Details › Model architecture and training ↔ DeorphaNN_training.ipynb, lines 568–692 · score 0.81 · AdamW, CrossEntropyLoss, DataLoader, smoothing, optimizer, PyTorch
  2. [2] § Star★Methods › Method Details › Model architecture and training ↔ DeorphaNN_training.ipynb, lines 568–692 · score 0.74 · precision score, Optuna, GPCR families, subsampled, shuffle, seed
  3. [3] § Results › AF-multimer confidence metrics partially discriminate peptide agonists from non-agonists ↔ src/analysis/vis_analyze.py, lines 194–288 · score 0.70 · interface contacts, pDockQ, plDDT, metrics, AF, chains
  4. [4] § Star★Methods › Method Details › AlphaFold2 ↔ structure_prediction/colabfold_runner.py, lines 226–287 · score 0.69 · max_msa, multimer v3, auto, v1, ColabFold, AlphaFold2
  5. [5] § Star★Methods › Method Details › AlphaFold2 ↔ preprocessing/minimum_distance.ipynb, lines 35–157 · score 0.68 · minimum distance, DeepTMHMM, binding pocket, atoms, error, positions
  6. [6] § Star★Methods › Method Details › AlphaFold2 ↔ src/alphafold/notebooks/AlphaFold.ipynb, lines 334–397 · score 0.67 · predicted aligned error, residue pLDDT, confidence metric, PAE, atoms, model
  7. [7] § Star★Methods › Method Details › AlphaFold2 ↔ alphafold/common/mmcif_metadata.py, lines 72–213 · score 0.58 · top ranked, pLDDT, coevolutionary, protocol, databases, conformations
  8. [8] § Star★Methods › Method Details › Model architecture and training ↔ esm/inverse_folding/gvp_modules.py, lines 331–475 · score 0.57 · node embeddings, convolutional, PyTorch, aggregates, ReLU, Geometric
  9. [9] § Results › AF-multimer confidence metrics partially discriminate peptide agonists from non-agonists ↔ structure_prediction/alphafold/common/confidence.py, lines 111–168 · score 0.56 · ipTM, predicted aligned error, confidence, interface, alignment, chains
  10. [10] § Star★Methods › Method Details › Phylogenetic analysis of GPCRs ↔ alphafold/data/tools/hmmbuild.py, lines 27–144 · score 0.54 · scoring matrix, aligned sequences, substitution, amino, model
  11. [11] § Star★Methods › Method Details › Phylogenetic analysis of GPCRs ↔ structure_prediction/alphafold/data/tools/hmmbuild.py, lines 26–138 · score 0.54 · scoring matrix, aligned sequences, substitution, amino, model
  12. [12] § Star★Methods › Method Details › Arpeggio ↔ structure_prediction/alphafold/model/all_atom.py, lines 744–850 · score 0.53 · van der Waals, bonds, atom, distance, position, residue
  13. [13] § Star★Methods › Method Details › Arpeggio ↔ alphafold/model/all_atom.py, lines 847–971 · score 0.52 · van der Waals, bonds, atom, distance, position, residue
  14. [14] § Star★Methods › Method Details › Model architecture and training ↔ esm/inverse_folding/gvp_modules.py, lines 331–475 · score 0.52 · attention heads, PyTorch, aggregating, Network, Geometric, dropout
  15. [15] § Results › AF-multimer confidence metrics partially discriminate peptide agonists from non-agonists ↔ src/analysis/vis_analyze.py, lines 292–375 · score 0.52 · pDockQ, top scoring, curves, ROC, plDDT, metric

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 700 lines · 25 KB · MIT · 2 matches

  1. # %% [markdown]
  2. # <a href="https://colab.research.google.com/github/Zebreu/DeorphaNN/blob/main/DeorphaNN_training.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>
  3. # %%
  4. #location to save results
  5. save_to = '/content/'
  6. # %%
  7. #@title Install Dependencies
  8. %%capture
  9. !pip uninstall torch -y
  10. !pip install torch==2.4.0
  11. !pip install pyg_lib torch_scatter torch_sparse torch_cluster torch_spline_conv -f https://data.pyg.org/whl/torch-2.4.0+cu121.html
  12. !pip install torch-geometric
  13. !pip install optuna
  14. !pip install numpy-indexed
  15. import glob
  16. import warnings
  17. import numpy as np
  18. import numpy_indexed as npi
  19. import pandas as pd
  20. import scipy
  21. from collections import defaultdict
  22. from sklearn.metrics import roc_auc_score, confusion_matrix, average_precision_score
  23. import torch
  24. from torch.nn import Linear
  25. import torch.nn.functional as F
  26. import torch_geometric
  27. from torch_geometric.data import Data
  28. from torch_geometric.loader import DataLoader
  29. from torch_geometric.nn import GCNConv, GATv2Conv
  30. from torch_geometric.nn import global_mean_pool, global_add_pool, global_max_pool
  31. from torch_geometric.nn import aggr
  32. from torch_geometric.nn.norm import GraphNorm
  33. from sklearn import preprocessing
  34. import optuna
  35. torch.manual_seed(111)
  36. from huggingface_hub import hf_hub_download, list_repo_files
  37. import h5py
  38. import random
  39. # %%
  40. #@title Import files from HuggingFace repo
  41. %%capture
  42. repo_id = "lariferg/DeorphaNN"
  43. all_files = list_repo_files(repo_id, repo_type="dataset")
  44. pdbs_paths = sorted(
  45. hf_hub_download(repo_id, f, repo_type="dataset")
  46. for f in all_files
  47. if f.startswith("DeorphaNN_training/nov7relaxed") and f.endswith(".parquet")
  48. )
  49. labels = hf_hub_download(
  50. repo_id,
  51. next(f for f in all_files if f.startswith("DeorphaNN_training/") and f.endswith("Dataset_Labels_full - Sheet1.csv")),
  52. repo_type="dataset"
  53. )
  54. beetsdata = pd.read_csv(labels)
  55. min_dis = hf_hub_download(
  56. repo_id,
  57. next(f for f in all_files if f.startswith("DeorphaNN_training/") and f.endswith("mindistance_active_bias.csv")),
  58. repo_type="dataset"
  59. )
  60. outsidepocket = pd.read_csv(min_dis)
  61. new_contacts_paths = sorted(
  62. hf_hub_download(repo_id, f, repo_type="dataset")
  63. for f in all_files
  64. if f.startswith("DeorphaNN_training/nov7relaxed") and f.endswith("_arpeggio_contacts3.parquet")
  65. )
  66. hdfs_int = sorted(
  67. hf_hub_download(repo_id, f, repo_type="dataset")
  68. for f in all_files
  69. if f.startswith("pair_representations/") and f.endswith("_interaction.h5")
  70. )
  71. hdfs_t = sorted(
  72. hf_hub_download(repo_id, f, repo_type="dataset")
  73. for f in all_files
  74. if f.startswith("pair_representations/") and f.endswith("T.h5")
  75. )
  76. # %% [markdown]
  77. # ###Preparing data
  78. # %%
  79. gpcrs = []
  80. peptides = []
  81. plddts = []
  82. paths = []
  83. plddt_peptides = []
  84. plddt_gpcrs = []
  85. plddt_atoms = []
  86. pdb_frames = dict()
  87. for pdbs in pdbs_paths:
  88. print(pdbs)
  89. pdbs = pd.read_parquet(pdbs)
  90. for key, st in pdbs.groupby('path'):
  91. if 'amber_r_' in key:
  92. original_key = key
  93. key = key.replace('amber_r_', '')
  94. gpcrs.append(key.split('/')[-1].split('_')[0])
  95. peptides.append(key.split('/')[-1].split('_')[1])
  96. plddt_peptides.append(st[st['chain_id'] == 'B'].groupby('residue_seq_id')['b_factor'].first().mean())
  97. plddt_gpcrs.append(st[st['chain_id'] == 'A'].groupby('residue_seq_id')['b_factor'].first().mean())
  98. plddt_atoms.append(st[st['chain_id'] == 'B']['b_factor'].mean())
  99. paths.append(original_key)
  100. pdb_frames[original_key] = st
  101. del pdbs
  102. # %%
  103. st = pd.DataFrame({'path': paths, 'gpcr': gpcrs, 'peptide': peptides, 'plddt_peptides': plddt_peptides, 'plddt_gpcrs': plddt_gpcrs, 'plddt_atoms': plddt_atoms})
  104. merged = pd.merge(beetsdata, st, how='left', left_on=['GPCR name', 'Peptide'], right_on=['gpcr', 'peptide'])
  105. merged['gpcr_family'] = merged['GPCR name'].str[:-2]
  106. merged['y'] = merged['binds'].apply(lambda x: 1 if x == True else 0)
  107. # %%
  108. gpcr_hits = merged[merged['gpcr'].isna() == False]
  109. gpcrs_lens = []
  110. peps_lens = []
  111. for index, st in gpcr_hits.iterrows():
  112. pdb = pdb_frames[st['path']]
  113. rec = pdb[pdb['chain_id'] == 'A']
  114. pep = pdb[pdb['chain_id'] == 'B']
  115. gpcrs_lens.append(rec['residue_seq_id'].max())
  116. peps_lens.append(pep['residue_seq_id'].max())
  117. gpcr_hits['gpcr_len'] = gpcrs_lens
  118. gpcr_hits['pep_len'] = peps_lens
  119. # %%
  120. outsidepocket['pair'] = outsidepocket['GPCR name']+'_'+outsidepocket['Peptide']
  121. # %%
  122. gpcr_hits = gpcr_hits[-gpcr_hits['pair'].isin(set(outsidepocket['pair'].values))]
  123. # %%
  124. allcontacts_new = []
  125. for cpath_new in new_contacts_paths:
  126. allcontacts_new.append(pd.read_parquet(cpath_new))
  127. allcontacts_new = pd.concat(allcontacts_new)
  128. # %%
  129. len(allcontacts_new)
  130. # %%
  131. allcontacts_new
  132. # %%
  133. total_interactions = allcontacts_new.groupby('gpcr_peptide')['contacts'].apply(lambda x: sum(len(c) for c in x)).reset_index()
  134. total_interactions.columns = ['gpcr_peptide', 'total_interactions']
  135. # %%
  136. total_interactions
  137. # %%
  138. gpcr_hits
  139. # %%
  140. gpcr_hits_bonds = pd.merge(
  141. gpcr_hits,
  142. total_interactions,
  143. how='left',
  144. left_on='pair',
  145. right_on='gpcr_peptide'
  146. )
  147. # Optionally drop the redundant 'gpcr_peptide' column
  148. gpcr_hits_bonds = gpcr_hits_bonds.drop(columns=['gpcr_peptide'])
  149. # %%
  150. gpcr_hits_bonds
  151. # %%
  152. interactions = dict()
  153. for key, st in allcontacts_new.groupby('gpcr_peptide'):
  154. interactions[key] = st
  155. # %%
  156. gpcr_hits_interaction_edges = dict()
  157. for index, g in gpcr_hits.iterrows():
  158. pdb = pdb_frames[g['path']].copy()
  159. if g['path'] not in interactions:
  160. continue
  161. bonds = interactions[g['path']]
  162. gpcr_len = g['gpcr_len']
  163. # they're 1-indexed so -1
  164. bonds['source'] = bonds['bgn'].apply(lambda x: x['auth_seq_id'] if x['auth_asym_id'] == "A" else x['auth_seq_id'] + gpcr_len) - 1
  165. bonds['target'] = bonds['end'].apply(lambda x: x['auth_seq_id'] if x['auth_asym_id'] == "A" else x['auth_seq_id'] + gpcr_len) - 1
  166. bonds = bonds.groupby(['source', 'target'])['contact'].agg(lambda x: {bondtype for array in x for bondtype in array}).reset_index()
  167. sources = bonds['source'].values
  168. targets = bonds['target'].values
  169. h_edge_index = np.vstack([sources,targets])
  170. key = g['gpcr']+'_'+g['peptide']
  171. gpcr_hits_interaction_edges[key] = h_edge_index
  172. # %%
  173. gpcr_hits_interaction_edges_new = dict()
  174. for _, row in allcontacts_new.iterrows():
  175. key = row['gpcr_peptide']
  176. contacts = row['contacts']
  177. # Check if contacts is None, NaN, or empty
  178. if contacts is None or len(contacts) == 0:
  179. # create empty 2x0 array
  180. #gpcr_hits_interaction_edges_new[key] = np.empty((2,0), dtype=int)
  181. continue
  182. # Stack the pairs vertically and transpose
  183. arr = np.vstack(contacts).T # shape: 2 x N_pairs
  184. arr -= 1 #convert 1-indexed to 0-indexed
  185. gpcr_hits_interaction_edges_new[key] = arr
  186. # %%
  187. len(gpcr_hits_interaction_edges_new)
  188. # %%
  189. len(interactions)
  190. # %%
  191. gpcr_hits_interaction_edges_new['DMSR-5-1_FLP-1-7']
  192. # %%
  193. lens = gpcr_hits.groupby(['GPCR name'])['gpcr_len'].first()
  194. # %%
  195. len(hdfs_int)
  196. # %%
  197. emb_map_interaction = dict()
  198. emb_map_interaction_gpcrindex = dict()
  199. pairmissed = []
  200. for hdf in hdfs_int:
  201. with h5py.File(hdf, "r") as f:
  202. keys = list(f.keys())
  203. print(keys)
  204. gpcr = keys[0].split('_')[0]
  205. for k in keys:
  206. try:
  207. array = np.nan_to_num(f[k][()],0)
  208. peptide = k.split('_')[1]
  209. mapkey = gpcr+'_'+peptide
  210. indices_to_keep = set()
  211. maximum = lens[gpcr]
  212. indices_to_keep.update(set(gpcr_hits_interaction_edges_new[mapkey][0]))
  213. indices_to_keep.update(set(gpcr_hits_interaction_edges_new[mapkey][1]))
  214. indices_to_keep = sorted([i for i in indices_to_keep if i < maximum])
  215. emb_map_interaction[mapkey] = array[:,indices_to_keep,:]
  216. emb_map_interaction_gpcrindex[mapkey] = np.array(indices_to_keep)
  217. print("success "+k)
  218. except:
  219. pairmissed.append(k)
  220. print("missed "+k)
  221. continue
  222. # %%
  223. len(emb_map_interaction)
  224. # %%
  225. len(pairmissed)
  226. # %%
  227. all_peptide_arrays = []
  228. peptide_keys = []
  229. all_gpcr_arrays = []
  230. gpcr_keys = []
  231. for hdf in hdfs_t:
  232. with h5py.File(hdf, "r") as f:
  233. arrays = []
  234. keys = list(f.keys())
  235. for k in keys:
  236. arrays.append(f[k][()])
  237. if "_pep_T" in hdf:
  238. all_peptide_arrays.append(arrays)
  239. peptide_keys.append(keys)
  240. if "_gpcr_T" in hdf:
  241. all_gpcr_arrays.append(arrays)
  242. gpcr_keys.append(keys)
  243. # %%
  244. emb_map_gpcr = dict()
  245. for keys, arrays in zip(gpcr_keys, all_gpcr_arrays):
  246. try:
  247. gpcr = keys[0].split('_')[0]
  248. for i,array in enumerate(arrays):
  249. peptide = keys[i].split('_')[1]
  250. emb_map_gpcr[gpcr+'_'+peptide] = array
  251. except:
  252. print('oops')
  253. continue
  254. # %%
  255. emb_map_peptide = dict()
  256. for keys, arrays in zip(peptide_keys, all_peptide_arrays):
  257. try:
  258. gpcr = keys[0].split('_')[0]
  259. for i,array in enumerate(arrays):
  260. peptide = keys[i].split('_')[1]
  261. emb_map_peptide[gpcr+'_'+peptide] = array
  262. except:
  263. print('oops')
  264. continue
  265. # %%
  266. embst = pd.DataFrame({'gpcr_keys': [kk for k in gpcr_keys for kk in k ], 'gpcr_embedding': [aa.mean(axis=0) for a in all_gpcr_arrays for aa in a ], 'peptide_keys': [kk for k in peptide_keys for kk in k], 'peptide_embedding': [aa.mean(axis=0) for a in all_peptide_arrays for aa in a]})
  267. # %%
  268. embst['peptide'] = embst['peptide_keys'].apply(lambda x: x.split('_')[1])
  269. embst['gpcr'] = embst['gpcr_keys'].apply(lambda x: x.split('_')[0])
  270. # %%
  271. gpcrweight = 1/gpcr_hits.groupby(['gpcr']).agg({'y': 'sum'}).sort_values(by='y')
  272. # %%
  273. gpcrweight
  274. # %%
  275. gpcr_hits
  276. # %%
  277. subgraphing = True
  278. subgraph_hops = 1
  279. with_edge_weights = True
  280. missed = []
  281. all_graphs = []
  282. for index, g in gpcr_hits.iterrows():
  283. mapkey = g['gpcr']+'_'+g['peptide']
  284. if mapkey not in gpcr_hits_interaction_edges_new:
  285. missed.append((mapkey, g['y']))
  286. continue
  287. gpcr_len = g['gpcr_len']
  288. h_edge_index = gpcr_hits_interaction_edges_new[g['gpcr']+'_'+g['peptide']]
  289. xg = emb_map_gpcr[g['gpcr']+'_'+g['peptide']]
  290. xp = emb_map_peptide[g['gpcr']+'_'+g['peptide']]
  291. x = np.concatenate([xg, xp])
  292. x = torch.from_numpy(x).type(torch.float32)
  293. pep_edge_index = np.vstack([np.array(range(g['gpcr_len'], len(x)-1)), np.array(range(g['gpcr_len']+1, len(x)))])
  294. edge_index = torch.cat([torch.from_numpy(h_edge_index), torch.from_numpy(pep_edge_index)], dim=1)
  295. if with_edge_weights:
  296. if mapkey not in emb_map_interaction:
  297. missed.append((mapkey, g['y']))
  298. continue
  299. edgefeatures = emb_map_interaction[mapkey]
  300. edgeindices = emb_map_interaction_gpcrindex[mapkey]
  301. sources = npi.remap(h_edge_index[0], edgeindices, np.arange(len(edgeindices)))
  302. targets = npi.remap(h_edge_index[1], edgeindices, np.arange(len(edgeindices)))
  303. sourcewherever = np.where(sources >= gpcr_len)[0]
  304. targetwherever = np.where(targets < gpcr_len)[0]
  305. newsources = np.array(sources)
  306. newtargets = np.array(targets)
  307. newsources[sourcewherever] = targets[sourcewherever]
  308. newtargets[targetwherever] = sources[targetwherever]
  309. newtargets -= gpcr_len
  310. edge_attrs = edgefeatures[newtargets, newsources, :]
  311. pep_edge_attrs = np.ones(shape=(len(pep_edge_index[0]),128))*edge_attrs.mean(axis=0)
  312. edge_attrs = torch.from_numpy(edge_attrs).type(torch.float32)
  313. edge_attrs = torch.cat([edge_attrs, torch.from_numpy(pep_edge_attrs)], dim=0)
  314. if with_edge_weights:
  315. # convert to undirected first, so hops are symmetric
  316. edge_index, edge_attrs = torch_geometric.utils.to_undirected(edge_index, edge_attrs, reduce='mean')
  317. if subgraphing:
  318. to_keep = torch.tensor([i for i in range(gpcr_len, len(x))]) #hopping from peptide nodes
  319. # to_keep = torch.unique(torch.from_numpy(h_edge_index[0])) #hopping from gpcr nodes
  320. nodes, edges, _, _ = torch_geometric.utils.k_hop_subgraph(to_keep, subgraph_hops, edge_index, relabel_nodes=True, num_nodes=len(x))
  321. # mask = (nodes >= gpcr_len) | (torch.isin(nodes, to_keep))
  322. # nodes = nodes[mask]
  323. if with_edge_weights:
  324. edges, new_edge_attrs = torch_geometric.utils.subgraph(nodes, edge_index, edge_attrs, relabel_nodes=True)
  325. # edges, new_edge_attrs = torch_geometric.utils.to_undirected(edges, new_edge_attrs, reduce='mean')
  326. graph = Data(x=x[nodes], edge_index=edges, edge_attr=new_edge_attrs, y=torch.tensor(g['y']))
  327. else:
  328. graph = Data(x=x[nodes], edge_index=edges, y=torch.tensor(g['y']))
  329. else:
  330. graph = Data(x=x, edge_index=edge_index, y=torch.tensor(g['y']))
  331. graph.peptide = g['peptide']
  332. graph.gpcr = g['gpcr']
  333. graph.gpcr_family = g['gpcr_family']
  334. #graph.zscore = g['modifiedzscore']
  335. gpcrw = gpcrweight.loc[g['gpcr']].iloc[0]
  336. all_graphs.append({'graph': graph, 'peptide':g['peptide'], 'gpcr':g['gpcr'], 'gpcr_family': g['gpcr_family'], 'y': g['y'], 'gpcrweight': gpcrw})
  337. # %%
  338. len(all_graphs)
  339. # %%
  340. len(missed)
  341. # %%
  342. print(missed)
  343. # %% [markdown]
  344. # #Train
  345. # %%
  346. from torch_geometric.loader import DataLoader
  347. from sklearn.metrics import roc_auc_score
  348. def train(model, criterion, optimizer, train_loader):
  349. model.train()
  350. total_loss = 0
  351. for data in train_loader:
  352. optimizer.zero_grad()
  353. logits = model(data.x, data.edge_index, data.edge_attr, data.batch) # edge_attr
  354. loss = criterion(logits, data.y)
  355. loss.backward()
  356. optimizer.step()
  357. total_loss += float(loss) * data.num_graphs
  358. return total_loss / len(train_loader.dataset)
  359. def train_weighted(model, criterion, optimizer, train_loader):
  360. model.train()
  361. total_loss = 0
  362. for data in train_loader:
  363. optimizer.zero_grad()
  364. logits = model(data.x, data.edge_index, data.edge_attr, data.batch) # edge_attr
  365. loss = criterion(logits, data.y)
  366. loss = (loss*data.gpcrweight).mean()
  367. loss.backward()
  368. optimizer.step()
  369. total_loss += float(loss) * data.num_graphs
  370. return total_loss / len(train_loader.dataset)
  371. @torch.no_grad()
  372. def test_roc(model, criterion, loader):
  373. model.eval()
  374. aucs = 0
  375. total = len(loader.dataset)
  376. correct = 0
  377. for data in loader: # Iterate in batches over the training/test dataset.
  378. out = model(data.x, data.edge_index, data.edge_attr, data.batch) # edge_attr
  379. aucs += roc_auc_score(data.y.detach().cpu(), torch.softmax(out.detach(),dim=1).cpu()[:, 1])*(len(out)/total)
  380. return aucs
  381. @torch.no_grad()
  382. def test_without_crash(model, criterion, loader):
  383. model.eval()
  384. all_logits = []
  385. atrues = []
  386. for data in loader:
  387. logits = model(data.x, data.edge_index, data.edge_attr, data.batch) # data.edge_attr
  388. all_logits.append(logits.cpu().detach()[:,1])
  389. atrues.append(data.y.cpu())
  390. return roc_auc_score(np.concatenate(atrues), np.concatenate(all_logits))
  391. @torch.no_grad()
  392. def nope_test_without_crash(model, criterion, loader):
  393. model.eval()
  394. all_logits = []
  395. atrues = []
  396. for data in loader:
  397. logits = model(data.x, data.edge_index, data.edge_attr, data.batch) # data.edge_attr
  398. all_logits.append(torch.sigmoid(logits.squeeze()).cpu().detach())
  399. atrues.append(data.y.cpu())
  400. return roc_auc_score(np.concatenate(atrues), np.concatenate(all_logits))
  401. @torch.no_grad()
  402. def test(model, criterion, loader):
  403. model.eval()
  404. total_correct = 0
  405. for data in loader:
  406. logits = model(data.x, data.edge_index, data.edge_attr, data.batch) # data.edge_attr
  407. pred = logits.argmax(dim=-1)
  408. total_correct += int((pred == data.y).sum())
  409. return total_correct / len(test_loader.dataset)
  410. # %%
  411. def move_to_cuda(g):
  412. g.x = g.x.cuda()
  413. g.edge_index = g.edge_index.cuda()
  414. g.edge_attr = g.edge_attr.cuda().type(torch.float32)
  415. g.y = g.y.cuda()
  416. return g
  417. # %%
  418. from torch_geometric.nn.norm import LayerNorm, BatchNorm
  419. from torch_geometric.nn import global_add_pool
  420. from torch_geometric.nn import aggr
  421. import sklearn
  422. class PeptideGNN(torch.nn.Module):
  423. def __init__(self, hidden_channels, input_channels=4, gatheads=10, gatdropout=0.5, finaldropout=0.5):
  424. super(PeptideGNN, self).__init__()
  425. self.finaldropout = finaldropout
  426. torch.manual_seed(111)
  427. self.norm = BatchNorm(input_channels)
  428. self.conv1 = GATv2Conv(input_channels, hidden_channels, dropout=gatdropout, heads=gatheads, concat=False, edge_dim=128)
  429. self.pooling = global_mean_pool
  430. self.lin = Linear(hidden_channels, 2)
  431. def forward(self, x, edge_index, edge_attr, batch, hidden=False):
  432. x = self.norm(x)
  433. x = self.conv1(x, edge_index, edge_attr)
  434. x = x.relu()
  435. if hidden:
  436. return x
  437. x = self.pooling(x, batch)
  438. x = F.dropout(x, p=self.finaldropout, training=self.training)
  439. x = self.lin(x)
  440. return x
  441. # %%
  442. gpcr_hits.groupby(['gpcr']).agg({'y': 'sum'}).sort_values(by='y')
  443. # %%
  444. for g in all_graphs:
  445. g['graph'].gpcrweight = torch.tensor(g['gpcrweight']).cuda()
  446. # %%
  447. # original splits (gpcr)
  448. # validation_peptides = [['NPR-43', 'CKR-1', 'NPR-39', 'AEX-2', 'DMSR-2', 'NPR-41'],
  449. # ['NPR-11', 'SPRR-2', 'SPRR-1', 'NPR-10', 'DMSR-3', 'GNRR-6'],
  450. # ['NPR-5', 'DMSR-8', 'NPR-2', 'FRPR-9', 'NPR-42', 'NPR-32'],
  451. # ['FRPR-8', 'NPR-40', 'FRPR-16', 'NPR-1', 'FRPR-6', 'FRPR-4'],
  452. # ['NPR-6', 'NMUR-2', 'FRPR-7', 'NPR-13', 'FRPR-19', 'TRHR-1'],
  453. # ['GNRR-1', 'FRPR-18', 'NPR-37', 'PDFR-1', 'FRPR-3'],
  454. # ['NPR-22', 'EGL-6', 'CKR-2', 'NMUR-1', 'NPR-4', 'FRPR-15'],
  455. # ['NPR-24', 'SEB-3', 'DMSR-6', 'NPR-12', 'DMSR-7'],
  456. # ['GNRR-3', 'NPR-35', 'TKR-2', 'NTR-1', 'DMSR-5'],
  457. # ['NPR-8', 'DMSR-1', 'NPR-3', 'TKR-1']]
  458. #phylogenetic splits (gpcr)
  459. validation_peptides = [['FRPR-16', 'FRPR-18', 'FRPR-4', 'FRPR-6', 'NPR-22', 'NMUR-2'],
  460. ['AEX-2', 'DMSR-5', 'DMSR-6', 'DMSR-7', 'DMSR-8','NPR-32'],
  461. ['FRPR-7', 'FRPR-7', 'NPR-6', 'GNRR-3', 'EGL-6'],
  462. ['TKR-1', 'TKR-2', 'DMSR-1', 'DMSR-2', 'NPR-40', 'GNRR-6'],
  463. ['NPR-42', 'FRPR-9', 'FRPR-15', 'FRPR-19', 'NMUR-1'],
  464. ['NPR-8', 'NPR-24', 'NPR-37', 'NPR-43', 'FRPR-3'],
  465. ['FRPR-8', 'GNRR-1', 'SPRR-1', 'SPRR-2', 'NPR-11', 'NPR-12'],
  466. ['NPR-41', 'NPR-1', 'NPR-2', 'NPR-3', 'PDFR-1', 'NTR-1'],
  467. ['TRHR-1', 'NPR-35', 'NPR-13', 'NPR-5', 'SEB-3'],
  468. ['DMSR-3', 'NPR-10', 'NPR-4', 'NPR-39', 'CKR-1', 'CKR-2']]
  469. # %%
  470. hpt_peptides = validation_peptides[-2:]+validation_peptides[0:-2]
  471. # %%
  472. all_graphs
  473. # %%
  474. test_logits = []
  475. test_labels = []
  476. candidatesgnn = []
  477. hitmapgnn = []
  478. subsampling_factor = 4
  479. average_precisions = dict()
  480. hpt_results = []
  481. for validation_peptide,hpt_peptide in zip(validation_peptides, hpt_peptides):
  482. training_graphs = [g['graph'] for g in all_graphs if g['gpcr_family'] not in validation_peptide and g['y'] == 1]
  483. hpt_training_graphs = [g['graph'] for g in all_graphs if g['gpcr_family'] not in validation_peptide and g['gpcr_family'] not in hpt_peptide and g['y'] == 1]
  484. hpt_to_shuffle = [g['graph'] for g in all_graphs if g['gpcr_family'] not in validation_peptide and g['gpcr_family'] not in hpt_peptide and g['y'] == 0]
  485. to_shuffle = [g['graph'] for g in all_graphs if g['gpcr_family'] not in validation_peptide and g['y'] == 0]
  486. random.Random(111).shuffle(to_shuffle)
  487. random.Random(111).shuffle(hpt_to_shuffle)
  488. trainings = []
  489. hpt_trainings = []
  490. for i in range(20):
  491. random.Random(i).shuffle(to_shuffle)
  492. trainings.append(training_graphs + to_shuffle[0:len(training_graphs)*subsampling_factor])
  493. trainings = [list(map(move_to_cuda, h)) for h in trainings]
  494. random.Random(i).shuffle(hpt_to_shuffle)
  495. hpt_trainings.append(hpt_training_graphs + hpt_to_shuffle[0:len(hpt_training_graphs)*subsampling_factor])
  496. hpt_trainings = [list(map(move_to_cuda, h)) for h in hpt_trainings]
  497. validation_graphs = [g['graph'] for g in all_graphs if g['gpcr_family'] in validation_peptide]
  498. hpt_validation_graphs = [g['graph'] for g in all_graphs if g['gpcr_family'] in hpt_peptide]
  499. validation_graphs = list(map(move_to_cuda, validation_graphs))
  500. hpt_validation_graphs = list(map(move_to_cuda, hpt_validation_graphs))
  501. test_loader = DataLoader(validation_graphs, batch_size=256, shuffle=False)
  502. hpt_test_loader = DataLoader(hpt_validation_graphs, batch_size=256, shuffle=False)
  503. hpt_maps = []
  504. hpt_logits = []
  505. hpt_labels = []
  506. def objective(trial):
  507. hidden_channels = trial.suggest_int('hidden_units', 50, 100)
  508. batch_size = trial.suggest_int('batch_size', 50, 200)
  509. lr = 0.0005
  510. model = PeptideGNN(hidden_channels, input_channels=128).cuda()
  511. optimizer = torch.optim.AdamW(model.parameters(), lr=lr)
  512. criterion = torch.nn.CrossEntropyLoss(label_smoothing=0.5, reduction='none')
  513. for epoch in range(1, 30):
  514. training_graphs = random.Random(epoch+111).choice(hpt_trainings)
  515. train_loader = DataLoader(training_graphs, batch_size=batch_size, shuffle=True)
  516. loss = train_weighted(model, criterion, optimizer, train_loader)
  517. model.eval()
  518. all_logits = []
  519. atrues = []
  520. with torch.no_grad():
  521. for data in hpt_test_loader:
  522. logits = model(data.x, data.edge_index, data.edge_attr, data.batch)
  523. if torch.isnan(data.x).any():
  524. print(f"NaN in node features for {hpt_peptide}")
  525. if torch.isnan(data.edge_attr).any():
  526. print(f"NaN in edge attributes for {hpt_peptide}")
  527. if torch.isinf(data.x).any() or torch.isinf(data.edge_attr).any():
  528. print(f"Infinite values in data for {hpt_peptide}")
  529. # 🧩 Check for NaNs in model output
  530. if torch.isnan(logits).any():
  531. print(f"NaN detected in model output during Optuna eval for {hpt_peptide}")
  532. continue # skip this batch safely
  533. all_logits.append(logits.cpu().detach()[:,1])
  534. atrues.append(data.y.cpu())
  535. hpt_logits.append(np.concatenate(all_logits))
  536. hpt_labels.append(np.concatenate(atrues))
  537. hpt_average_precisions = []
  538. val_gpcr = [g.gpcr for g in hpt_validation_graphs]
  539. val_peptide = [g.peptide for g in hpt_validation_graphs]
  540. result = pd.DataFrame(zip(hpt_labels[-1], hpt_logits[-1], val_gpcr, val_peptide))
  541. for gpcr, r in result.groupby(2):
  542. if r[0].sum() > 0:
  543. if np.isnan(r[0]).any():
  544. print(f"NaNs in TRUE labels for group: {gpcr}")
  545. if np.isnan(r[1]).any():
  546. print(f"NaNs in PREDICTIONS for group: {gpcr}")
  547. hpt_average_precisions.append(sklearn.metrics.average_precision_score(r[0], r[1]))
  548. else:
  549. print('what')
  550. hpt_maps.append((np.mean(hpt_average_precisions), (hidden_channels, batch_size, lr)))
  551. return np.mean(hpt_average_precisions)
  552. sampler = optuna.samplers.RandomSampler(seed=111)
  553. study = optuna.create_study(direction='maximize', sampler=sampler)
  554. study.optimize(objective, n_trials=40)
  555. hpt_results.append(hpt_maps)
  556. _, params = sorted(hpt_maps)[-1]
  557. model = PeptideGNN(params[0], input_channels=128).cuda()
  558. optimizer = torch.optim.AdamW(model.parameters(), lr=params[2])
  559. criterion = torch.nn.CrossEntropyLoss(label_smoothing=0.5, reduction='none')
  560. print(validation_peptide)
  561. for epoch in range(1, 30):
  562. training_graphs = random.Random(epoch+111).choice(trainings)
  563. train_loader = DataLoader(training_graphs, batch_size=params[1], shuffle=True)
  564. loss = train_weighted(model, criterion, optimizer, train_loader)
  565. if epoch % 14 == 0:
  566. test_acc = test_without_crash(model, criterion, test_loader)
  567. print(f'Epoch: {epoch:02d}, Train Acc: {test_without_crash(model, criterion, train_loader):.4f}, Test AUC: {test_acc:.4f}')
  568. torch.save(model.state_dict(), f'{save_to}pretrained_{validation_peptide[0]}.pth')
  569. model.eval()
  570. all_logits = []
  571. atrues = []
  572. with torch.no_grad():
  573. for data in test_loader:
  574. logits = model(data.x, data.edge_index, data.edge_attr, data.batch)
  575. all_logits.append(logits.cpu().detach()[:,1])
  576. atrues.append(data.y.cpu())
  577. test_logits.append(np.concatenate(all_logits))
  578. test_labels.append(np.concatenate(atrues))
  579. val_gpcr = [g.gpcr for g in validation_graphs]
  580. val_peptide = [g.peptide for g in validation_graphs]
  581. result = pd.DataFrame(zip(test_labels[-1], test_logits[-1], val_gpcr, val_peptide))
  582. for gpcr, r in result.groupby(2):
  583. hitmapgnn.append((gpcr, r.sort_values(by=1).iloc[-17:][0].sum()))
  584. candidatesgnn.append((gpcr,r.sort_values(by=1)))
  585. if r[0].sum() > 0:
  586. average_precisions[gpcr] = sklearn.metrics.average_precision_score(r[0], r[1])
  587. print(gpcr, average_precisions[gpcr])
  588. print(roc_auc_score(np.concatenate(test_labels), np.concatenate(test_logits)))
  589. # %%
  590. np.mean(list(average_precisions.values()))
  591. # %%
  592. map_value = int(round(np.mean(list(average_precisions.values())),3)*1000)
  593. st = pd.concat([st for g,st in candidatesgnn])
  594. st.to_csv(f'/content/average_precision_values.csv', index=False)

DeorphaNN_training.ipynb at commit 0ddab30, under MIT · at the source

Overview

Authors: Larissa Ferguson1, Sébastien Ouellet2, Elke Vandewyer3, Christopher Wang2, Zaw Wunna1, Tony KY Lim4, William R Schafer1,3, Isabel Beets3
  1. Neurobiology Division, MRC Laboratory of Molecular Biology, Cambridge, UK
  2. Independent Researcher, Ottawa, ON, Canada
  3. Department of Biology, KU Leuven, Leuven, Belgium
  4. Department of Pharmacology, University of Cambridge, Cambridge, UK
Institutions: MRC Laboratory of Molecular Biology (United Kingdom); KU Leuven (Belgium); University of Cambridge (United Kingdom)
Journal: Molecular cell, volume 86, issue 15, pages 3102-3117.e5
Dates: published online 23 July 2026; in print 6 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.molcel.2026.07.006 · PMID 42492505 · PMCID PMC7619462 · OpenAlex W4408742047
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), C. elegans (organism), cellular / molecular (subfield)
Methods: Smoothing, state filtering, decompositions, Machine learning, Statistics
Keywords: Neuropeptides, G protein-coupled receptors, Peptide hormone, Structural Bioinformatics, Deorphanization, Protein Representations, Alphafold, Peptide Agonists, Active-state Structures, Protein Embeddings
MeSH: Caenorhabditis elegans Proteins*, Deep Learning*, Drug Discovery*, Peptides*, Receptors, G-Protein-Coupled*, Animals, Caenorhabditis elegans, Graph Neural Networks, Humans, Neuropeptides, Protein Binding (* major topic)
Topic: Receptor Mechanisms and Signaling (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Medical Research Council (MC_U105185857); KU Leuven; UK Research and Innovation Medical Research Council; European Biodiversity Partnership; Research Foundation Flanders
Citations: not cited yet (Europe PMC); 100 references in the paper

Abstract

Peptide-activated G protein-coupled receptors (GPCRs) regulate physiological processes through interaction with neuropeptides and peptide hormones. Identifying endogenous peptide agonists remains challenging, as peptide-GPCR pairings often follow gene-family relationships that offer limited predictive insight for orphan GPCRs without characterized homologs. Using a dataset of experimentally validated peptide-GPCR interactions from Caenorhabditis elegans, we demonstrate that AF-multimer confidence metrics partially discriminate agonist from non-agonist complexes, with improved discrimination using AF-Multistate-derived active-state templates. Feature analysis revealed that AF-multimer’s pair representations outperform single representations, with distinct subregions providing complementary signals. Leveraging these insights, we developed DeorphaNN, a graph neural network integrating active-state GPCR-peptide structural predictions, interatomic interactions, and deep learning embeddings to prioritize putative peptide agonists for experimental screening. DeorphaNN generalized across diverse species, as shown by performance on annelid and human retrospective benchmarks. Experimental validation confirmed predicted agonists for two orphan GPCRs, demonstrating its utility for accelerating peptide-GPCR deorphanization.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 15 matches between paragraphs and lines of code.

Zebreu/DeorphaNN

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 0ddab304da08d48540ce09a3072a5c7d1e5a9b98, 25 June 2026
Languages: Python (7), Jupyter (3)
Size: 13 files, 10 scripts
Software Heritage: not archived
Found in: “Data and code availability”
Holds: README, license file, environment (requirements.txt), 3 notebooks
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (7 files), PyTorch (7 files), pandas (5 files), PyTorch Geometric (3 files), Biopython (2 files), h5py (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
12 files

Zenodo 20862123

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 70 files
Software Heritage: not checked
Found in: “Data and code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

Zenodo 20865768

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data and code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (7 files), PyTorch (7 files), pandas (5 files), PyTorch Geometric (3 files), Biopython (2 files), h5py (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
12 files
At the source:

Maryam-Haghani/NEFFy

License: GPL-3.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c6bc480369201d030b7a18243e71fc261c2d1406, 6 May 2026
Languages: JavaScript (93), Python (12), C++ (7), C/C++ (5)
Size: 392 files, 117 scripts
Software Heritage: not archived
Found in: the resources table
Holds: README, license file, CITATION.cff, environment (pyproject.toml, setup.py), documentation
Not found: tests, continuous integration
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
119 files

facebookresearch/esm

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 2b369911bb5b4b0dda914521b9475cad1656b2ac, 27 June 2023
Languages: Python (79), Jupyter (8), Shell (1)
Size: 476 files, 88 scripts
Software Heritage: archived
Found in: the resources table
Holds: README, license file, environment (environment.yml, pyproject.toml, setup.py), tests, 8 notebooks
Not found: CITATION.cff, continuous integration, documentation
Tools: PyTorch (52 files), NumPy (22 files), SciPy (6 files), Matplotlib (4 files), pandas (3 files), PyTorch Geometric (3 files), Biopython (2 files), seaborn (2 files), scikit-learn (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
90 files

ElofssonLab/FoldDock

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 9a1a26ced4f6b8b9bc65a7ac76999118c292b80d, 18 April 2023
Languages: Python (99), Shell (36), Jupyter (2)
Size: 19,182 files, 137 scripts
Software Heritage: not archived
Found in: the resources table
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (64 files), pandas (26 files), JAX (17 files), TensorFlow (11 files), Biopython (10 files), SciPy (9 files), scikit-learn (4 files), Matplotlib (3 files), seaborn (2 files)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
139 files

huhlim/alphafold-multistate

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 7014212f2111d82ec0bf19c7cd70441e7c8ef7f2, 12 April 2024
Languages: Python (90), Shell (3), Jupyter (2)
Size: 132 files, 95 scripts
Software Heritage: archived
Found in: the resources table
Holds: README, environment (structure_prediction/requirements.txt), tests, 2 notebooks
Not found: license file, CITATION.cff, continuous integration, documentation
Tools: NumPy (43 files), JAX (29 files), TensorFlow (11 files), Biopython (5 files), SciPy (3 files), Matplotlib (2 files), pandas (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
96 files

google-deepmind/alphafold

License: Apache-2.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c77e5d2a8961d1a353632c462914ff0a32a950f6, 22 April 2026
Languages: Python (87), Shell (11), Jupyter (1)
Size: 121 files, 99 scripts
Software Heritage: archived
Found in: the resources table
Holds: README, license file, environment (pyproject.toml, requirements.txt, docker/Dockerfile, docker/requirements.txt), tests, documentation, 1 notebook
Not found: CITATION.cff, continuous integration
Tools: NumPy (43 files), JAX (31 files), TensorFlow (8 files), Biopython (3 files), Matplotlib (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
101 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 8 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 556 scripts, each with its path and the digest of its content;
  • 15 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data and code availability

All datasets utilized in this study are available as supplemental materials and in the associated Zenodo repository (DOI: 10.5281/zenodo.20862123 (https://www.doi.org/10.5281/zenodo.20862123)). The DeorphaNN code is available at https://github.com/Zebreu/DeorphaNN (DOI: 10.5281/zenodo.20865768 (https://www.doi.org/10.5281/zenodo.20865768)). Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Elsevier BV
  • Authors: added Sébastien Ouellet (0000-0002-3098-4907); removed Sébastien Ouellet

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 8 authors, 10 keywords, 11 MeSH terms, 5 funders, 99 references.

Cite

This paper

Ferguson, L., Ouellet, S., Vandewyer, E., Wang, C., Wunna, Z., Lim, T. K., Schafer, W. R., & Beets, I. (2026). DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings. Molecular cell, 86(15), 3102-3117.e5. https://doi.org/10.1016/j.molcel.2026.07.006

BibTeX

@article{ferguson2026deorphann,
author = {Ferguson, Larissa and Ouellet, Sébastien and Vandewyer, Elke and Wang, Christopher and Wunna, Zaw and Lim, Tony KY and Schafer, William R and Beets, Isabel},
title = {{DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings}},
journal = {Molecular cell},
year = {2026},
month = jul,
volume = {86},
number = {15},
pages = {3102--3117.e5},
publisher = {Elsevier BV},
issn = {1097-2765},
doi = {10.1016/j.molcel.2026.07.006},
url = {https://doi.org/10.1016/j.molcel.2026.07.006},
pmid = {42492505},
pmcid = {PMC7619462}
}

RIS

TY - JOUR
AU - Ferguson, Larissa
AU - Ouellet, Sébastien
AU - Vandewyer, Elke
AU - Wang, Christopher
AU - Wunna, Zaw
AU - Lim, Tony KY
AU - Schafer, William R
AU - Beets, Isabel
TI - DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings
T2 - Molecular cell
J2 - Mol Cell
PY - 2026
DA - 2026/07/23
VL - 86
IS - 15
SP - 3102
EP - 3117.e5
SN - 1097-2765
PB - Elsevier BV
DO - 10.1016/j.molcel.2026.07.006
UR - https://doi.org/10.1016/j.molcel.2026.07.006
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.molcel.2026.07.006",
"type": "article-journal",
"title": "DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings",
"container-title": "Molecular cell",
"author": [
{
"family": "Ferguson",
"given": "Larissa"
},
{
"family": "Ouellet",
"given": "Sébastien"
},
{
"family": "Vandewyer",
"given": "Elke"
},
{
"family": "Wang",
"given": "Christopher"
},
{
"family": "Wunna",
"given": "Zaw"
},
{
"family": "Lim",
"given": "Tony KY"
},
{
"family": "Schafer",
"given": "William R"
},
{
"family": "Beets",
"given": "Isabel"
}
],
"container-title-short": "Mol Cell",
"volume": "86",
"issue": "15",
"page": "3102-3117.e5",
"DOI": "10.1016/j.molcel.2026.07.006",
"PMID": "42492505",
"PMCID": "PMC7619462",
"ISSN": "1097-2765",
"publisher": "Elsevier BV",
"URL": "https://doi.org/10.1016/j.molcel.2026.07.006",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
23
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: JAX, Biopython, PyTorch Geometric, 7 other tools, cellular / molecular, 3 references
[2] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: JAX, Biopython, PyTorch Geometric, 7 other tools, cellular / molecular, 2 references
[3] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: JAX, Biopython, PyTorch Geometric, 7 other tools, cellular / molecular, 2 references
[4] doi:10.1038/s41586-026-10670-w [code]
Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
Journal: Nature
In common: PyTorch Geometric, h5py, PyTorch, 6 other tools, cellular / molecular, 4 references
[5] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits &lt;i&gt;Porphyromonas gingivalis&lt;/i&gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: JAX, Biopython, PyTorch Geometric, 7 other tools, cellular / molecular, 1 reference
[6] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: JAX, Biopython, TensorFlow, 7 other tools, cellular / molecular
[7] doi:10.1002/advs.202523984 [code]
INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: Biopython, PyTorch Geometric, PyTorch, 6 other tools, 1 reference
[8] doi:10.1038/s41586-026-10658-6 [code]
An AI system to help scientists write expert-level empirical software.
Journal: Nature
In common: JAX, TensorFlow, h5py, 6 other tools, 1 reference
[9] doi:10.1523/eneuro.0362-25.2026 [code]
Similarities between &lt;i&gt;Ciona&lt;/i&gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
Journal: eNeuro
In common: Biopython, PyTorch Geometric, h5py, 7 other tools
[10] doi:10.1016/j.xops.2026.101249 [code]
A Disorder-Aware Computational Framework to Identify Structurally Tractable Targets in Proliferative Vitreoretinopathy.
Journal: Ophthalmology science
In common: JAX, Biopython, TensorFlow, 5 other tools, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.