OSCR

Cross-Species Multitask Learning with Molecular and ADME Descriptors for Liver Microsomal Metabolic Stability.

Code ↔ Paper

19 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 19 matches
  1. [1] § Materials and Methods › Molecular representation and embedding › Molecular ADME and physicochemical descriptors ↔ Analysis/Edgeshaper/Dataset.py, lines 50–92 · score 0.99 · CYP2D6, synthetic accessibility, CYP1A2, iLOGP, bond acceptor, bioavailability scores
  2. [2] § Materials and Methods › Molecular representation and embedding › Molecular ADME and physicochemical descriptors ↔ Notebooks/dataset_scaffold_modelready.py, lines 75–109 · score 0.99 · CYP2D6, synthetic accessibility, CYP1A2, iLOGP, bond acceptor, bioavailability scores
  3. [3] § Materials and Methods › Molecular representation and embedding › Molecular graph ↔ Notebooks/model.py, lines 212–358 · score 0.95 · global max pooling, residual connections, Jumping Knowledge, LeakyReLU, attention heads, edge feature
  4. [4] § Materials and Methods › Molecular representation and embedding › Molecular graph ↔ Analysis/Edgeshaper/Model.py, lines 179–291 · score 0.82 · global max pooling, Jumping Knowledge, LeakyReLU, hidden, Node, residual
  5. [5] § Results › SHAP-based interpretation of ADME descriptors for liver microsomal stability across species ↔ Analysis/Edgeshaper/Dataset.py, lines 50–92 · score 0.78 · synthetic accessibility, bond donors, CYP2C9 inhibition, CYP2C19, WLOGP, ADME
  6. [6] § Results › SHAP-based interpretation of ADME descriptors for liver microsomal stability across species ↔ Notebooks/dataset_scaffold_modelready.py, lines 75–109 · score 0.78 · synthetic accessibility, bond donors, CYP2C9 inhibition, CYP2C19, WLOGP
  7. [7] § Materials and Methods › Loss function ↔ Analysis/SHAP/Train.py, lines 256–392 · score 0.73 · margin ranking loss, uncertainty weighting, focal loss, learnable, encourages, logits
  8. [8] § Materials and Methods › Loss function ↔ Analysis/SHAP/SHAP.py, lines 300–414 · score 0.69 · margin ranking loss, uncertainty weighting, focal loss, encourages, logits
  9. [9] § Materials and Methods › Training and evaluation ↔ Analysis/Edgeshaper/Utile.py, lines 402–492 · score 0.68 · precision recall curve, F1 score, Matthews, accuracy, MCC, AUPR
  10. [10] § Materials and Methods › Loss function ↔ Analysis/SHAP/SHAP.py, lines 1–56 · score 0.67 · margin ranking loss, class imbalance, focal loss
  11. [11] § Materials and Methods › Molecular representation and embedding › Molecular fingerprint ↔ Notebooks/dataset_scaffold_modelready.py, lines 426–515 · score 0.61 · hashed fingerprints, RDKit, bit, MACCS, Morgan, atom
  12. [12] § Materials and Methods › Training and evaluation ↔ Analysis/Edgeshaper/Edgeshaper.py, lines 732–815 · score 0.59 · AdamW, cosine, warmup, scheduling, optimized, metric
  13. [13] § Materials and Methods › Prediction ↔ Analysis/Edgeshaper/Model.py, lines 418–549 · score 0.57 · LayerNorm, GELU, fused, encoders, dropout, head
  14. [14] § Materials and Methods › Prediction ↔ Notebooks/model.py, lines 492–624 · score 0.57 · LayerNorm, GELU, fused, encoders, dropout, head
  15. [15] § Materials and Methods › Molecular representation and embedding › Molecular graph ↔ Notebooks/model.py, lines 204–210 · score 0.57 · formal charge, hybridization, node, aromaticity, atoms, encoded
  16. [16] § Materials and Methods › Loss function ↔ Analysis/SHAP/Train.py, lines 256–392 · score 0.54 · margin ranking loss, focal loss
  17. [17] § Materials and Methods › External evaluation dataset ↔ Notebooks/dataset_scaffold_modelready.py, lines 426–515 · score 0.53 · isomeric SMILES, RDKit, canonical, scaffold
  18. [18] § Materials and Methods › Molecular representation and embedding › Molecular fingerprint ↔ Notebooks/model.py, lines 18–94 · score 0.51 · RDKit, hashing, MACCS, Morgan, linear, fingerprints
  19. [19] § Materials and Methods › Training dataset ↔ Notebooks/predict_colab.ipynb, lines 1–31 · score 0.50 · RDKit, sparse, liver microsomal stability, Cross species, cutoff, CYP

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,215 lines · 34 KB · no license · 4 matches

  1. # dataset_scaffold.py
  2. # ============================================================
  3. # Dataset preprocessing for multi-task molecular classification
  4. # - SMILES -> graph
  5. # - dense fingerprints: Morgan / MACCS / RDKit
  6. # - descriptor preprocessing with train-only scaler
  7. # - scaffold-disjoint K-fold with fold-level descriptor scaler
  8. # ============================================================
  9. import os
  10. import random
  11. import time
  12. import numpy as np
  13. import pandas as pd
  14. from typing import Optional, List, Dict, Tuple
  15. from collections import defaultdict, Counter
  16. import torch
  17. from torch_geometric.data import InMemoryDataset, Data
  18. from torch_geometric.loader import DataLoader as GeometricDataLoader
  19. from tqdm import tqdm
  20. from sklearn.preprocessing import StandardScaler
  21. from rdkit import Chem
  22. from rdkit.Chem import MACCSkeys
  23. from rdkit.Chem.Scaffolds import MurckoScaffold
  24. from rdkit import DataStructs
  25. from rdkit.Chem.rdFingerprintGenerator import (
  26. GetMorganGenerator,
  27. GetRDKitFPGenerator,
  28. )
  29. # ============================================================
  30. # 1. SMILES sequence encoding
  31. # ============================================================
  32. smi_to_seq = "(.02468@BDFHLNPRTVZ/bdfhlnprt#*%)+-/13579=ACEGIKMOSUWY[]acegimosuy\\"
  33. seq_dict_smi = {ch: (i + 1) for i, ch in enumerate(smi_to_seq)}
  34. MAX_SEQ_SMI_LEN = 100
  35. def seq_smi(smile: str, max_seq_smi_len: int = MAX_SEQ_SMI_LEN) -> np.ndarray:
  36. idx = np.array(
  37. [seq_dict_smi.get(ch, 0) for ch in str(smile)[:max_seq_smi_len]],
  38. dtype=int,
  39. )
  40. if len(idx) < max_seq_smi_len:
  41. idx = np.pad(
  42. idx,
  43. (0, max_seq_smi_len - len(idx)),
  44. "constant",
  45. constant_values=0,
  46. )
  47. return idx
  48. # ============================================================
  49. # 2. Fingerprint generators
  50. # ============================================================
  51. MORGAN_GEN = GetMorganGenerator(radius=2, fpSize=2048)
  52. RDK_GEN = GetRDKitFPGenerator(fpSize=2048)
  53. # ============================================================
  54. # 3. Descriptor settings
  55. # ============================================================
  56. USE_EXPLICIT_COLS = True
  57. PHYS_COLS = [
  58. "MW", "TPSA", "iLOGP", "XLOGP3", "WLOGP", "MLOGP",
  59. "Silicos-IT Log P", "Consensus Log P",
  60. "ESOL Log S", "Ali Log S", "Silicos-IT LogSw",
  61. "#Heavy atoms", "#Aromatic heavy atoms", "Fraction Csp3",
  62. "#Rotatable bonds", "#H-bond acceptors", "#H-bond donors",
  63. "MR", "log Kp (cm/s)",
  64. ]
  65. CONT_COLS = [
  66. "Lipinski #violations", "Ghose #violations", "Veber #violations",
  67. "Egan #violations", "Muegge #violations", "Bioavailability Score",
  68. "PAINS #alerts", "Brenk #alerts", "Leadlikeness #violations",
  69. "Synthetic Accessibility",
  70. ]
  71. CAT_COLS = [
  72. "ESOL Class", "Ali Class", "Silicos-IT class",
  73. "GI absorption", "BBB permeant", "Pgp substrate",
  74. "CYP1A2 inhibitor", "CYP2C19 inhibitor",
  75. "CYP2C9 inhibitor", "CYP2D6 inhibitor", "CYP3A4 inhibitor",
  76. ]
  77. BINARY_POSITIVE_MAP = {
  78. "GI absorption": "High",
  79. "BBB permeant": "Yes",
  80. "Pgp substrate": "Yes",
  81. "CYP1A2 inhibitor": "Yes",
  82. "CYP2C19 inhibitor": "Yes",
  83. "CYP2C9 inhibitor": "Yes",
  84. "CYP2D6 inhibitor": "Yes",
  85. "CYP3A4 inhibitor": "Yes",
  86. }
  87. def _coerce_numeric_series(s: pd.Series) -> pd.Series:
  88. return pd.to_numeric(
  89. s.astype(str).str.replace(",", "", regex=False).str.strip(),
  90. errors="coerce",
  91. )
  92. def _make_desc_df_raw(df_: pd.DataFrame, tasks: List[str]) -> pd.DataFrame:
  93. """
  94. Build the descriptor subset in raw form.
  95. No NaN/Inf imputation or scaling is done here.
  96. """
  97. exclude = set(
  98. [
  99. "Cano_Smile",
  100. "SMILES",
  101. "PUBCHEM_EXT_DATASOURCE_SMILES",
  102. "_canonical_smiles",
  103. ]
  104. + list(tasks)
  105. )
  106. if USE_EXPLICIT_COLS:
  107. phys_in = [c for c in PHYS_COLS if c in df_.columns and c not in exclude]
  108. cont_in = [c for c in CONT_COLS if c in df_.columns and c not in exclude]
  109. cat_in = [c for c in CAT_COLS if c in df_.columns and c not in exclude]
  110. # numeric descriptors
  111. if phys_in or cont_in:
  112. num_df = df_[phys_in + cont_in].copy()
  113. for c in num_df.columns:
  114. num_df[c] = _coerce_numeric_series(num_df[c])
  115. else:
  116. num_df = pd.DataFrame(index=df_.index)
  117. # categorical descriptors
  118. if cat_in:
  119. cat_cols = []
  120. for c in cat_in:
  121. col = df_[c].astype(str).str.strip()
  122. col = col.replace({"nan": np.nan, "None": np.nan, "": np.nan})
  123. uniq = sorted([u for u in col.dropna().unique()])
  124. if 0 < len(uniq) <= 2:
  125. positive = BINARY_POSITIVE_MAP.get(c, uniq[-1])
  126. bin_series = (col == positive).astype(float)
  127. bin_series = bin_series.fillna(0.0)
  128. bin_series.name = f"{c}__{positive}"
  129. cat_cols.append(bin_series)
  130. else:
  131. dummies = pd.get_dummies(col, prefix=c, dummy_na=False)
  132. cat_cols.append(dummies)
  133. cat_df = pd.concat(cat_cols, axis=1)
  134. else:
  135. cat_df = pd.DataFrame(index=df_.index)
  136. desc_df = pd.concat([num_df, cat_df], axis=1)
  137. else:
  138. candidate_numeric_cols = []
  139. for c in df_.columns:
  140. if c in exclude:
  141. continue
  142. coerced = _coerce_numeric_series(df_[c])
  143. if coerced.notna().mean() > 0.9:
  144. candidate_numeric_cols.append(c)
  145. if candidate_numeric_cols:
  146. desc_df = df_[candidate_numeric_cols].copy()
  147. for c in desc_df.columns:
  148. desc_df[c] = _coerce_numeric_series(desc_df[c])
  149. else:
  150. desc_df = pd.DataFrame(index=df_.index)
  151. if desc_df.shape[1] == 0:
  152. desc_df = pd.DataFrame(
  153. {"__desc_dummy__": np.zeros(len(df_), dtype=float)},
  154. index=df_.index,
  155. )
  156. return desc_df
  157. def build_desc_df_scaled(
  158. df: pd.DataFrame,
  159. tasks: List[str],
  160. logger=None,
  161. fixed_cols: Optional[List[str]] = None,
  162. scaler: Optional[StandardScaler] = None,
  163. impute_values: Optional[Dict[str, float]] = None,
  164. add_missing_indicators: bool = True,
  165. ):
  166. """
  167. Descriptor preprocessing.
  168. Train:
  169. fixed_cols=None, scaler=None, impute_values=None
  170. -> build columns, compute medians, fit StandardScaler
  171. Val/Test/External:
  172. fixed_cols=train_cols, scaler=train_scaler, impute_values=train_impute
  173. -> reindex to train schema, impute with train medians, transform with train scaler
  174. """
  175. desc_df_raw = _make_desc_df_raw(df, tasks)
  176. numeric_base_cols = [
  177. c for c in desc_df_raw.columns
  178. if c in set(PHYS_COLS + CONT_COLS)
  179. ]
  180. desc_df_proc = desc_df_raw.replace([np.inf, -np.inf], np.nan)
  181. if add_missing_indicators:
  182. for c in numeric_base_cols:
  183. miss_col = f"{c}__missing"
  184. desc_df_proc[miss_col] = desc_df_proc[c].isna().astype(float)
  185. if fixed_cols is None:
  186. cols = list(desc_df_proc.columns)
  187. else:
  188. for c in fixed_cols:
  189. if c not in desc_df_proc.columns:
  190. desc_df_proc[c] = 0.0
  191. desc_df_proc = desc_df_proc.reindex(columns=fixed_cols)
  192. cols = list(fixed_cols)
  193. if impute_values is None:
  194. impute_values = {}
  195. for c in cols:
  196. med = desc_df_proc[c].median()
  197. if pd.isna(med):
  198. med = 0.0
  199. impute_values[c] = float(med)
  200. else:
  201. impute_values = dict(impute_values)
  202. for c in cols:
  203. desc_df_proc[c] = desc_df_proc[c].fillna(impute_values.get(c, 0.0))
  204. X = desc_df_proc.values.astype("float32")
  205. if scaler is None:
  206. scaler = StandardScaler()
  207. X_scaled = scaler.fit_transform(X)
  208. else:
  209. X_scaled = scaler.transform(X)
  210. desc_df_scaled = pd.DataFrame(
  211. X_scaled,
  212. columns=cols,
  213. index=df.index,
  214. )
  215. if logger:
  216. n_missing_indicators = sum(c.endswith("__missing") for c in cols)
  217. logger.info(
  218. f"[Descriptor] cols={desc_df_scaled.shape[1]} | "
  219. f"missing_indicators={n_missing_indicators} | "
  220. f"fit_scaler={scaler is not None and fixed_cols is None}"
  221. )
  222. return desc_df_scaled, cols, scaler, impute_values
  223. # ============================================================
  224. # 4. Graph feature encoding
  225. # ============================================================
  226. def one_of_k_encoding_unk(x, allowable_set):
  227. if x not in allowable_set:
  228. x = allowable_set[-1]
  229. return [x == s for s in allowable_set]
  230. def _explicit_valence(atom):
  231. if hasattr(Chem, "ValenceType"):
  232. return atom.GetValence(Chem.ValenceType.EXPLICIT)
  233. return atom.GetExplicitValence()
  234. # NOTE:
  235. # To stay compatible with DEFAULT_NODE_IN_DIM=89 in model.py,
  236. # the atom list dim is kept; if you shrink the element list here, update model.py too.
  237. ATOM_LIST = [
  238. "C", "N", "O", "S", "F", "Si", "P", "Cl", "Br", "Mg",
  239. "Na", "Ca", "Fe", "As", "Al", "I", "B", "V", "K", "Tl",
  240. "Yb", "Sb", "Sn", "Ag", "Pd", "Co", "Se", "Ti", "Zn", "H",
  241. "Li", "Ge", "Cu", "Au", "Ni", "Cd", "In", "Mn", "Zr", "Cr",
  242. "Pt", "Hg", "Pb", "Nd", "Ru", "W", "Unknown", "Mo", "Sr",
  243. "Bi", "S", "Ba", "Be", "Dy",
  244. ]
  245. DEGREE_LIST = [0, 1, 2, 3, 4, 5, 6]
  246. FORMAL_CHARGE_LIST = [-1, 0, 1]
  247. EXPLICIT_VALENCE_LIST = [0, 1, 2, 3, 4, 5, 6]
  248. NUM_H_LIST = [0, 1, 2, 3, 4, 5]
  249. HYBRIDIZATION_LIST = [
  250. Chem.rdchem.HybridizationType.SP,
  251. Chem.rdchem.HybridizationType.SP2,
  252. Chem.rdchem.HybridizationType.SP3,
  253. Chem.rdchem.HybridizationType.SP3D,
  254. Chem.rdchem.HybridizationType.SP3D2,
  255. "UNSPECIFIED",
  256. "S",
  257. ]
  258. CHIRALITY_LIST = [
  259. Chem.rdchem.ChiralType.CHI_UNSPECIFIED,
  260. Chem.rdchem.ChiralType.CHI_TETRAHEDRAL_CW,
  261. Chem.rdchem.ChiralType.CHI_TETRAHEDRAL_CCW,
  262. Chem.rdchem.ChiralType.CHI_OTHER,
  263. ]
  264. def atom_features(atom):
  265. features = (
  266. one_of_k_encoding_unk(atom.GetSymbol(), ATOM_LIST)
  267. + one_of_k_encoding_unk(atom.GetDegree(), DEGREE_LIST)
  268. + one_of_k_encoding_unk(atom.GetFormalCharge(), FORMAL_CHARGE_LIST)
  269. + one_of_k_encoding_unk(_explicit_valence(atom), EXPLICIT_VALENCE_LIST)
  270. + one_of_k_encoding_unk(atom.GetTotalNumHs(), NUM_H_LIST)
  271. + one_of_k_encoding_unk(atom.GetHybridization(), HYBRIDIZATION_LIST)
  272. + [atom.GetIsAromatic()]
  273. + one_of_k_encoding_unk(atom.GetChiralTag(), CHIRALITY_LIST)
  274. )
  275. return np.array(features, dtype=float)
  276. def bond_features(bond):
  277. bt = bond.GetBondType()
  278. return np.array(
  279. [
  280. bt == Chem.rdchem.BondType.SINGLE,
  281. bt == Chem.rdchem.BondType.DOUBLE,
  282. bt == Chem.rdchem.BondType.TRIPLE,
  283. bt == Chem.rdchem.BondType.AROMATIC,
  284. bond.GetIsConjugated(),
  285. bond.IsInRing(),
  286. ],
  287. dtype=float,
  288. )
  289. # ============================================================
  290. # 5. CSV and Data conversion helpers
  291. # ============================================================
  292. SMI_CANDIDATES = ["Cano_Smile", "SMILES", "PUBCHEM_EXT_DATASOURCE_SMILES"]
  293. def find_smiles_column(df: pd.DataFrame) -> str:
  294. smi_col = next((c for c in SMI_CANDIDATES if c in df.columns), None)
  295. if smi_col is None:
  296. raise ValueError(f"Input CSV must contain one of {SMI_CANDIDATES}")
  297. return smi_col
  298. def prepare_dataframe(
  299. dataset_path: str,
  300. tasks: List[str],
  301. logger=None,
  302. ) -> Tuple[pd.DataFrame, str]:
  303. df = pd.read_csv(dataset_path)
  304. smi_col = find_smiles_column(df)
  305. for t in tasks:
  306. if t not in df.columns:
  307. df[t] = -1
  308. df[t] = _coerce_numeric_series(df[t]).fillna(-1).astype(np.float32)
  309. label_mat = df[tasks].values
  310. valid_label_mask = (label_mat != -1).any(axis=1)
  311. df = df[valid_label_mask].reset_index(drop=True)
  312. canonical_smiles = []
  313. keep_idx = []
  314. invalid_smiles = 0
  315. for idx, smi in enumerate(df[smi_col].astype(str)):
  316. mol = Chem.MolFromSmiles(smi.strip())
  317. if mol is None:
  318. invalid_smiles += 1
  319. continue
  320. canonical_smiles.append(Chem.MolToSmiles(mol, isomericSmiles=True))
  321. keep_idx.append(idx)
  322. df = df.iloc[keep_idx].reset_index(drop=True)
  323. df["_canonical_smiles"] = canonical_smiles
  324. if logger:
  325. logger.info(
  326. f"[prepare_dataframe] {os.path.basename(dataset_path)} | "
  327. f"kept={len(df)} | invalid_smiles={invalid_smiles}"
  328. )
  329. return df, "_canonical_smiles"
  330. def dataframe_to_data_list(
  331. df: pd.DataFrame,
  332. tasks: List[str],
  333. desc_df: pd.DataFrame,
  334. smi_col: str,
  335. logger=None,
  336. ):
  337. data_list = []
  338. smiles_attr = []
  339. invalid_smiles = 0
  340. for i, row in tqdm(
  341. df.iterrows(),
  342. total=len(df),
  343. desc="Graph Conversion",
  344. ):
  345. smi = str(row[smi_col]).strip()
  346. mol = Chem.MolFromSmiles(smi)
  347. if mol is None:
  348. invalid_smiles += 1
  349. continue
  350. canonical_smi = Chem.MolToSmiles(mol, isomericSmiles=True)
  351. atom_feat = [atom_features(atom) for atom in mol.GetAtoms()]
  352. x = torch.tensor(np.array(atom_feat), dtype=torch.float)
  353. edge_indices = []
  354. edge_attrs = []
  355. for bond in mol.GetBonds():
  356. start = bond.GetBeginAtomIdx()
  357. end = bond.GetEndAtomIdx()
  358. b_feat = bond_features(bond)
  359. edge_indices.append([start, end])
  360. edge_attrs.append(b_feat)
  361. edge_indices.append([end, start])
  362. edge_attrs.append(b_feat)
  363. if len(edge_indices) == 0:
  364. edge_index = torch.empty((2, 0), dtype=torch.long)
  365. edge_attr = torch.empty((0, 6), dtype=torch.float)
  366. else:
  367. edge_index = torch.tensor(edge_indices, dtype=torch.long).t().contiguous()
  368. edge_attr = torch.tensor(np.array(edge_attrs), dtype=torch.float)
  369. data = Data(x=x, edge_index=edge_index, edge_attr=edge_attr)
  370. data.smiles = canonical_smi
  371. # Morgan / RDKit hashed fingerprints
  372. for gen, name in [(MORGAN_GEN, "morgan_fp"), (RDK_GEN, "rdit_fp")]:
  373. bv = gen.GetFingerprint(mol)
  374. arr = np.zeros((bv.GetNumBits(),), dtype=np.int8)
  375. DataStructs.ConvertToNumpyArray(bv, arr)
  376. setattr(
  377. data,
  378. name,
  379. torch.from_numpy(arr.astype(np.float32)).unsqueeze(0),
  380. )
  381. # MACCS
  382. bv_maccs = MACCSkeys.GenMACCSKeys(mol)
  383. arr_maccs = np.zeros((bv_maccs.GetNumBits(),), dtype=np.int8)
  384. DataStructs.ConvertToNumpyArray(bv_maccs, arr_maccs)
  385. data.maccs_fp = torch.from_numpy(arr_maccs.astype(np.float32)).unsqueeze(0)
  386. data.y = torch.tensor(
  387. row[tasks].values.astype(np.float32),
  388. dtype=torch.float,
  389. ).unsqueeze(0)
  390. data.desc = torch.tensor(
  391. desc_df.iloc[i].values.astype(np.float32),
  392. dtype=torch.float,
  393. ).unsqueeze(0)
  394. data.smil2vec = torch.LongTensor(seq_smi(canonical_smi)).unsqueeze(0)
  395. data_list.append(data)
  396. smiles_attr.append(canonical_smi)
  397. if logger:
  398. logger.info(
  399. f"[dataframe_to_data_list] valid={len(data_list)} | "
  400. f"invalid_smiles={invalid_smiles}"
  401. )
  402. return data_list, smiles_attr
  403. # ============================================================
  404. # 6. Safe torch load
  405. # ============================================================
  406. def _safe_torch_load(path: str):
  407. try:
  408. return torch.load(path, map_location="cpu", weights_only=False)
  409. except TypeError:
  410. return torch.load(path, map_location="cpu")
  411. # ============================================================
  412. # 7. MolDataset
  413. # ============================================================
  414. class MolDataset(InMemoryDataset):
  415. def __init__(
  416. self,
  417. root,
  418. dataset,
  419. task_type,
  420. tasks,
  421. logger=None,
  422. transform=None,
  423. pre_transform=None,
  424. pre_filter=None,
  425. desc_cols=None,
  426. desc_scaler=None,
  427. desc_impute_values=None,
  428. processed_suffix="default",
  429. ):
  430. self.tasks = tasks
  431. self.dataset = dataset
  432. self.task_type = task_type
  433. self.logger = logger
  434. self.fixed_desc_cols = desc_cols
  435. self.fixed_desc_scaler = desc_scaler
  436. self.fixed_desc_impute_values = desc_impute_values
  437. self.processed_suffix = processed_suffix
  438. super().__init__(root, transform, pre_transform, pre_filter)
  439. loaded = _safe_torch_load(self.processed_paths[0])
  440. if isinstance(loaded, tuple) and len(loaded) == 6:
  441. (
  442. self.data,
  443. self.slices,
  444. self.smiles_list,
  445. self.desc_cols_,
  446. self.desc_scaler_,
  447. self.desc_impute_values_,
  448. ) = loaded
  449. elif isinstance(loaded, tuple) and len(loaded) == 5:
  450. (
  451. self.data,
  452. self.slices,
  453. self.smiles_list,
  454. self.desc_cols_,
  455. self.desc_scaler_,
  456. ) = loaded
  457. self.desc_impute_values_ = None
  458. else:
  459. self.data, self.slices, self.smiles_list = loaded[:3]
  460. self.desc_cols_ = None
  461. self.desc_scaler_ = None
  462. self.desc_impute_values_ = None
  463. @property
  464. def raw_file_names(self):
  465. return [self.dataset]
  466. @property
  467. def processed_file_names(self):
  468. base = os.path.splitext(os.path.basename(self.dataset))[0]
  469. return [f"{base}_{self.processed_suffix}.pt"]
  470. def process(self):
  471. dataset_path = os.path.join(self.root, self.dataset)
  472. df, smi_col = prepare_dataframe(
  473. dataset_path=dataset_path,
  474. tasks=self.tasks,
  475. logger=self.logger,
  476. )
  477. desc_df, used_cols, scaler, impute_values = build_desc_df_scaled(
  478. df,
  479. tasks=self.tasks,
  480. logger=self.logger,
  481. fixed_cols=self.fixed_desc_cols,
  482. scaler=self.fixed_desc_scaler,
  483. impute_values=self.fixed_desc_impute_values,
  484. )
  485. self.desc_cols_ = list(used_cols)
  486. self.desc_scaler_ = scaler
  487. self.desc_impute_values_ = impute_values
  488. data_list, smiles_attr = dataframe_to_data_list(
  489. df=df,
  490. tasks=self.tasks,
  491. desc_df=desc_df,
  492. smi_col=smi_col,
  493. logger=self.logger,
  494. )
  495. if len(data_list) == 0:
  496. raise RuntimeError(f"[MolDataset] No valid molecules found in {self.dataset}")
  497. data, slices = self.collate(data_list)
  498. torch.save(
  499. (
  500. data,
  501. slices,
  502. smiles_attr,
  503. self.desc_cols_,
  504. self.desc_scaler_,
  505. self.desc_impute_values_,
  506. ),
  507. self.processed_paths[0],
  508. )
  509. # ============================================================
  510. # 8. Scaffold split utilities
  511. # ============================================================
  512. def _murcko_scaffold(smi: str, include_chirality: bool = False) -> str:
  513. try:
  514. mol = Chem.MolFromSmiles(smi)
  515. if mol is None:
  516. return ""
  517. return MurckoScaffold.MurckoScaffoldSmiles(
  518. mol=mol,
  519. includeChirality=include_chirality,
  520. ) or ""
  521. except Exception:
  522. return ""
  523. def _group_indices_by_scaffold(
  524. smiles_list: List[str],
  525. include_chirality: bool = False,
  526. ):
  527. buckets = defaultdict(list)
  528. for i, smi in enumerate(smiles_list):
  529. scaf = _murcko_scaffold(smi, include_chirality=include_chirality)
  530. key = scaf if len(scaf) > 0 else smi
  531. buckets[key].append(i)
  532. return buckets
  533. def _greedy_pack_scaffolds_to_folds(
  534. scaffold_buckets: Dict[str, List[int]],
  535. n_splits: int,
  536. seed: int = 2026,
  537. ):
  538. groups = list(scaffold_buckets.items())
  539. groups.sort(key=lambda kv: (-len(kv[1]), kv[0]))
  540. fold_bins = [[] for _ in range(n_splits)]
  541. fold_sizes = [0] * n_splits
  542. for _, idx_list in groups:
  543. k = min(range(n_splits), key=lambda f: fold_sizes[f])
  544. fold_bins[k].extend(idx_list)
  545. fold_sizes[k] += len(idx_list)
  546. return fold_bins
  547. def _log_fold_stats(data_list, tasks, logger, tag=""):
  548. if not logger:
  549. return
  550. cnt = Counter()
  551. pos_cnt = Counter()
  552. neg_cnt = Counter()
  553. for d in data_list:
  554. yrow = d.y.view(-1)
  555. for i, t in enumerate(tasks):
  556. val = yrow[i].item()
  557. if val != -1:
  558. cnt[t] += 1
  559. if val == 1:
  560. pos_cnt[t] += 1
  561. elif val == 0:
  562. neg_cnt[t] += 1
  563. logger.info(
  564. f"{tag} size={len(data_list)} | "
  565. f"available={dict(cnt)} | "
  566. f"pos={dict(pos_cnt)} | "
  567. f"neg={dict(neg_cnt)}"
  568. )
  569. # ============================================================
  570. # 9. Scaffold K-fold loader with fold-level descriptor scaler
  571. # ============================================================
  572. import hashlib
  573. import json
  574. def _safe_torch_load_cache(path: str):
  575. try:
  576. return torch.load(path, map_location="cpu", weights_only=False)
  577. except TypeError:
  578. return torch.load(path, map_location="cpu")
  579. def _make_fold_cache_path(
  580. data_path,
  581. dataset_name,
  582. tasks,
  583. n_splits,
  584. include_chirality,
  585. seed,
  586. cache_dir=None,
  587. cache_version="v1",
  588. ):
  589. """
  590. Build the per-fold graph/data_list cache path.
  591. Includes dataset file size/mtime, so a changed CSV automatically uses a different cache.
  592. """
  593. dataset_path = os.path.join(data_path, dataset_name)
  594. if not os.path.exists(dataset_path):
  595. raise FileNotFoundError(dataset_path)
  596. stat = os.stat(dataset_path)
  597. meta = {
  598. "dataset_name": dataset_name,
  599. "tasks": list(tasks),
  600. "n_splits": int(n_splits),
  601. "include_chirality": bool(include_chirality),
  602. "seed": int(seed),
  603. "file_size": int(stat.st_size),
  604. "file_mtime": int(stat.st_mtime),
  605. "cache_version": str(cache_version),
  606. }
  607. key = hashlib.md5(
  608. json.dumps(meta, sort_keys=True).encode("utf-8")
  609. ).hexdigest()[:12]
  610. if cache_dir is None:
  611. cache_dir = os.path.join(data_path, "fold_graph_cache")
  612. os.makedirs(cache_dir, exist_ok=True)
  613. base = os.path.splitext(os.path.basename(dataset_name))[0]
  614. cache_path = os.path.join(
  615. cache_dir,
  616. f"{base}_scaffold{n_splits}_seed{seed}_{key}.pt",
  617. )
  618. return cache_path, meta
  619. def _build_loaders_from_cached_fold_lists(
  620. cached_train_lists,
  621. cached_val_lists,
  622. batch_size,
  623. ):
  624. train_loaders = []
  625. val_loaders = []
  626. for train_data_list, val_data_list in zip(cached_train_lists, cached_val_lists):
  627. train_loader = GeometricDataLoader(
  628. train_data_list,
  629. batch_size=batch_size,
  630. shuffle=True,
  631. num_workers=0,
  632. drop_last=False,
  633. )
  634. val_loader = GeometricDataLoader(
  635. val_data_list,
  636. batch_size=batch_size,
  637. shuffle=False,
  638. num_workers=0,
  639. drop_last=False,
  640. )
  641. train_loaders.append(train_loader)
  642. val_loaders.append(val_loader)
  643. return train_loaders, val_loaders
  644. def build_scaffold_kfold_loader(
  645. data_path: str,
  646. dataset_name: str,
  647. task_type: str,
  648. batch_size: int,
  649. tasks: List[str],
  650. logger=None,
  651. n_splits: int = 10,
  652. include_chirality: bool = False,
  653. seed: int = 2026,
  654. use_cache: bool = True,
  655. force_rebuild: bool = False,
  656. cache_dir: str = None,
  657. cache_version: str = "v1",
  658. ):
  659. """
  660. Scaffold K-fold loader with fold-level descriptor scaler + fold graph cache.
  661. First run:
  662. raw CSV -> scaffold split -> per-fold descriptor scaler fit/transform
  663. -> graph conversion -> save fold cache
  664. Subsequent runs:
  665. load fold cache -> build DataLoader
  666. Returns
  667. -------
  668. train_loaders : list[DataLoader]
  669. val_loaders : list[DataLoader]
  670. fold_desc_info: list[dict]
  671. """
  672. random.seed(seed)
  673. np.random.seed(seed)
  674. torch.manual_seed(seed)
  675. # --------------------------------------------------------
  676. # 0. build cache path
  677. # --------------------------------------------------------
  678. cache_path, cache_meta = _make_fold_cache_path(
  679. data_path=data_path,
  680. dataset_name=dataset_name,
  681. tasks=tasks,
  682. n_splits=n_splits,
  683. include_chirality=include_chirality,
  684. seed=seed,
  685. cache_dir=cache_dir,
  686. cache_version=cache_version,
  687. )
  688. # --------------------------------------------------------
  689. # 1. Cache load
  690. # --------------------------------------------------------
  691. if use_cache and (not force_rebuild) and os.path.exists(cache_path):
  692. if logger:
  693. logger.info("=" * 70)
  694. logger.info(f"[Scaffold Cache] Loading cached folds:")
  695. logger.info(f" {cache_path}")
  696. logger.info("=" * 70)
  697. payload = _safe_torch_load_cache(cache_path)
  698. cached_train_lists = payload["train_data_lists"]
  699. cached_val_lists = payload["val_data_lists"]
  700. fold_desc_info = payload["fold_desc_info"]
  701. train_loaders, val_loaders = _build_loaders_from_cached_fold_lists(
  702. cached_train_lists=cached_train_lists,
  703. cached_val_lists=cached_val_lists,
  704. batch_size=batch_size,
  705. )
  706. if logger:
  707. logger.info(
  708. f"[Scaffold Cache] Loaded {len(train_loaders)} folds from cache."
  709. )
  710. for fold_idx, (tr_list, va_list) in enumerate(
  711. zip(cached_train_lists, cached_val_lists)
  712. ):
  713. logger.info(
  714. f"[Cached Fold {fold_idx + 1}/{len(train_loaders)}] "
  715. f"Train={len(tr_list)}, Val={len(va_list)}"
  716. )
  717. _log_fold_stats(tr_list, tasks, logger, tag=" Train")
  718. _log_fold_stats(va_list, tasks, logger, tag=" Val ")
  719. return train_loaders, val_loaders, fold_desc_info
  720. # --------------------------------------------------------
  721. # 2. build fresh if no cache
  722. # --------------------------------------------------------
  723. if logger:
  724. logger.info("=" * 70)
  725. logger.info("[Scaffold Cache] Cache not found or force_rebuild=True.")
  726. logger.info("[Scaffold Cache] Building folds from raw CSV.")
  727. logger.info(f"[Scaffold Cache] Cache will be saved to: {cache_path}")
  728. logger.info("=" * 70)
  729. dataset_path = os.path.join(data_path, dataset_name)
  730. df, smi_col = prepare_dataframe(
  731. dataset_path=dataset_path,
  732. tasks=tasks,
  733. logger=logger,
  734. )
  735. smiles_list = df[smi_col].tolist()
  736. buckets = _group_indices_by_scaffold(
  737. smiles_list,
  738. include_chirality=include_chirality,
  739. )
  740. fold_bins = _greedy_pack_scaffolds_to_folds(
  741. buckets,
  742. n_splits=n_splits,
  743. seed=seed,
  744. )
  745. train_loaders = []
  746. val_loaders = []
  747. fold_desc_info = []
  748. cached_train_lists = []
  749. cached_val_lists = []
  750. # --------------------------------------------------------
  751. # 3. build per-fold data_list
  752. # --------------------------------------------------------
  753. for fold_idx in range(n_splits):
  754. val_index = sorted(fold_bins[fold_idx])
  755. train_index = sorted(
  756. [
  757. i
  758. for k in range(n_splits)
  759. if k != fold_idx
  760. for i in fold_bins[k]
  761. ]
  762. )
  763. train_df = df.iloc[train_index].reset_index(drop=True)
  764. val_df = df.iloc[val_index].reset_index(drop=True)
  765. # fit descriptor scaler on fold-train only
  766. train_desc_df, desc_cols, desc_scaler, desc_impute_values = build_desc_df_scaled(
  767. train_df,
  768. tasks=tasks,
  769. logger=logger,
  770. fixed_cols=None,
  771. scaler=None,
  772. impute_values=None,
  773. )
  774. # transform fold-val with the train scaler
  775. val_desc_df, _, _, _ = build_desc_df_scaled(
  776. val_df,
  777. tasks=tasks,
  778. logger=logger,
  779. fixed_cols=desc_cols,
  780. scaler=desc_scaler,
  781. impute_values=desc_impute_values,
  782. )
  783. train_data_list, _ = dataframe_to_data_list(
  784. df=train_df,
  785. tasks=tasks,
  786. desc_df=train_desc_df,
  787. smi_col=smi_col,
  788. logger=logger,
  789. )
  790. val_data_list, _ = dataframe_to_data_list(
  791. df=val_df,
  792. tasks=tasks,
  793. desc_df=val_desc_df,
  794. smi_col=smi_col,
  795. logger=logger,
  796. )
  797. train_loader = GeometricDataLoader(
  798. train_data_list,
  799. batch_size=batch_size,
  800. shuffle=True,
  801. num_workers=0,
  802. drop_last=False,
  803. )
  804. val_loader = GeometricDataLoader(
  805. val_data_list,
  806. batch_size=batch_size,
  807. shuffle=False,
  808. num_workers=0,
  809. drop_last=False,
  810. )
  811. if logger:
  812. logger.info(
  813. f"[Scaffold Fold {fold_idx + 1}/{n_splits}] "
  814. f"Train={len(train_data_list)}, Val={len(val_data_list)}"
  815. )
  816. _log_fold_stats(train_data_list, tasks, logger, tag=" Train")
  817. _log_fold_stats(val_data_list, tasks, logger, tag=" Val ")
  818. train_loaders.append(train_loader)
  819. val_loaders.append(val_loader)
  820. cached_train_lists.append(train_data_list)
  821. cached_val_lists.append(val_data_list)
  822. fold_desc_info.append(
  823. {
  824. "desc_cols": desc_cols,
  825. "desc_scaler": desc_scaler,
  826. "desc_impute_values": desc_impute_values,
  827. "train_index": train_index,
  828. "val_index": val_index,
  829. }
  830. )
  831. # --------------------------------------------------------
  832. # 4. Save cache
  833. # --------------------------------------------------------
  834. if use_cache:
  835. payload = {
  836. "meta": cache_meta,
  837. "train_data_lists": cached_train_lists,
  838. "val_data_lists": cached_val_lists,
  839. "fold_desc_info": fold_desc_info,
  840. }
  841. torch.save(payload, cache_path)
  842. if logger:
  843. logger.info("=" * 70)
  844. logger.info(f"[Scaffold Cache] Saved fold cache:")
  845. logger.info(f" {cache_path}")
  846. logger.info("=" * 70)
  847. return train_loaders, val_loaders, fold_desc_info
  848. # ============================================================
  849. # 10. Standard train/val/test loader
  850. # ============================================================
  851. def build_loader(
  852. data_path,
  853. dataset_names,
  854. task_type,
  855. batch_size,
  856. tasks,
  857. logger=None,
  858. ):
  859. train_loader = None
  860. val_loader = None
  861. test_loader = None
  862. train_desc_cols = None
  863. train_desc_scaler = None
  864. train_desc_impute_values = None
  865. if "train" in dataset_names:
  866. train_dataset = MolDataset(
  867. root=data_path,
  868. dataset=dataset_names["train"],
  869. task_type=task_type,
  870. tasks=tasks,
  871. logger=logger,
  872. desc_cols=None,
  873. desc_scaler=None,
  874. desc_impute_values=None,
  875. processed_suffix="train_fit",
  876. )
  877. train_desc_cols = getattr(train_dataset, "desc_cols_", None)
  878. train_desc_scaler = getattr(train_dataset, "desc_scaler_", None)
  879. train_desc_impute_values = getattr(train_dataset, "desc_impute_values_", None)
  880. train_loader = GeometricDataLoader(
  881. train_dataset,
  882. batch_size=batch_size,
  883. shuffle=True,
  884. num_workers=0,
  885. drop_last=False,
  886. )
  887. if "val" in dataset_names:
  888. val_dataset = MolDataset(
  889. root=data_path,
  890. dataset=dataset_names["val"],
  891. task_type=task_type,
  892. tasks=tasks,
  893. logger=logger,
  894. desc_cols=train_desc_cols,
  895. desc_scaler=train_desc_scaler,
  896. desc_impute_values=train_desc_impute_values,
  897. processed_suffix="val_using_train_scaler",
  898. )
  899. val_loader = GeometricDataLoader(
  900. val_dataset,
  901. batch_size=batch_size,
  902. shuffle=False,
  903. num_workers=0,
  904. drop_last=False,
  905. )
  906. if "test" in dataset_names:
  907. test_dataset = MolDataset(
  908. root=data_path,
  909. dataset=dataset_names["test"],
  910. task_type=task_type,
  911. tasks=tasks,
  912. logger=logger,
  913. desc_cols=train_desc_cols,
  914. desc_scaler=train_desc_scaler,
  915. desc_impute_values=train_desc_impute_values,
  916. processed_suffix="test_using_train_scaler",
  917. )
  918. test_loader = GeometricDataLoader(
  919. test_dataset,
  920. batch_size=batch_size,
  921. shuffle=False,
  922. num_workers=0,
  923. drop_last=False,
  924. )
  925. return (
  926. train_loader,
  927. val_loader,
  928. test_loader,
  929. train_desc_cols,
  930. train_desc_scaler,
  931. train_desc_impute_values,
  932. )
  933. # ============================================================
  934. # 11. External loader
  935. # ============================================================
  936. def build_external_loader(
  937. data_path,
  938. dataset_name,
  939. task_type,
  940. batch_size,
  941. tasks,
  942. logger,
  943. train_desc_cols,
  944. train_desc_scaler,
  945. train_desc_impute_values,
  946. processed_suffix="external_using_train_scaler",
  947. ):
  948. """
  949. External HLM/RLM test loader.
  950. You must pass the train descriptor schema/scaler/impute_values.
  951. """
  952. external_dataset = MolDataset(
  953. root=data_path,
  954. dataset=dataset_name,
  955. task_type=task_type,
  956. tasks=tasks,
  957. logger=logger,
  958. desc_cols=train_desc_cols,
  959. desc_scaler=train_desc_scaler,
  960. desc_impute_values=train_desc_impute_values,
  961. processed_suffix=processed_suffix,
  962. )
  963. external_loader = GeometricDataLoader(
  964. external_dataset,
  965. batch_size=batch_size,
  966. shuffle=False,
  967. num_workers=0,
  968. drop_last=False,
  969. )
  970. return external_loader
  971. # ============================================================
  972. # 12. Quick feature dimension check
  973. # ============================================================
  974. def check_atom_feature_dim(smiles: str = "CCO") -> int:
  975. mol = Chem.MolFromSmiles(smiles)
  976. if mol is None:
  977. raise ValueError(f"Invalid SMILES: {smiles}")
  978. return len(atom_features(mol.GetAtomWithIdx(0)))

dataset_scaffold_modelready.py at commit 2e085d8, no license · at the source

Overview

Authors: Subhin Seomun1, Sunyong Yoo1,2
ORCID iDs: Sunyong Yoo
  1. Department of Intelligent Electronics and Computer Engineering, Chonnam National University, Gwangju, Republic of Korea
  2. R&D Center, MATILO AI Inc., Gwangju, Republic of Korea
Institutions: Chonnam National University (South Korea)
Journal: Computational and structural biotechnology journal, volume 35, issue 1, article 0184
Dates: received 10 May 2026; accepted 13 July 2026; published online 11 August 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.34133/csbj.0184 · PMID 42582689 · PMCID PMC13457899 · OpenAlex W7168294012
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: cellular / molecular (subfield)
Methods: Statistics, Machine learning, Graphs
Topic: Machine Learning in Bioinformatics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Ministry of Food and Drug Safety (RS-2025-02215961, RS-2024-00332003); Ministry of Science and ICT, South Korea (RS-2025-16063391)
Citations: not cited yet (Europe PMC); 70 references in the paper

Abstract

Liver microsomal metabolic stability is a key determinant of in vivo exposure and an essential filter in lead optimization, yet cross-species prediction remains difficult because of heterogeneous metabolic pathways and limited model interpretability. We propose a cross-species multitask learning framework that integrates complementary molecular modalities—SMILES-derived fingerprints (Morgan and MACCS/RDKit), molecular graphs, and in silico absorption, distribution, metabolism, and excretion (ADME)/physicochemical descriptors—to predict binary microsomal stability (unstable: t1/2 ≤ 30 min; stable: t1/2 > 30 min) in human (HLM), rat (RLM), and mouse (MLM) liver microsomes. We curated 18,921 PubChem BioAssay measurements (6,685 HLM; 5,753 RLM; 6,483 MLM). Under stratified 10-fold Bemis–Murcko scaffold cross-validation with ensemble prediction and species-specific thresholds, the model achieved AUROC values of 0.811, 0.806, and 0.794 and AUPR values of 0.854, 0.860, and 0.862 for HLM, RLM, and MLM, respectively, consistently outperforming conventional machine-learning and single-task deep-learning baselines. SHapley Additive exPlanations (SHAP) identified transport/permeability indicators, CYP interaction flags, and the lipophilicity–polarity axis as the features most strongly associated with predicted stability. EdgeSHAPer, a graph neural network explanation method based on SHAP, highlighted stabilizing and destabilizing substructures. Recurrent destabilizing attributions were observed in alkene and allylic/benzylic contexts, whereas amide/carbamate motifs exhibited stabilizing attributions, with nitriles and halogens showing context-dependent effects. Fragment–ADME enrichment analysis characterized associations between local structural motifs and whole-molecule properties including lipophilicity, solubility, and blood–brain barrier permeability. This multi-modal, cross-species framework demonstrates that integrating structural encodings with ADME descriptors enhances both predictive performance and interpretability, yielding hypothesis-generating attributions for structural optimization that warrant prospective experimental validation.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 19 matches between paragraphs and lines of code.

bmil-jnu/cross-species_metabolic_stability_prediction

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 2e085d8bbc018abac4877a2686f481388ad47d8f, 20 August 2026
Languages: Python (23), Jupyter (2)
Size: 42 files, 25 scripts
Software Heritage: not archived
Found in: “Data Availability”
Holds: README, environment (Main/requirements.txt, Notebooks/pyproject.toml, Notebooks/requirements.txt), 2 notebooks
Not found: license file, CITATION.cff, tests, continuous integration, documentation
Tools: PyTorch (21 files), PyTorch Geometric (15 files), NumPy (14 files), pandas (13 files), RDKit (7 files), scikit-learn (7 files), Matplotlib (6 files), SHAP (4 files), Pillow (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
26 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 25 scripts, each with its path and the digest of its content;
  • 19 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability

The model implementation, the trained model, the inference and explanation software, and the preprocessed training and external-validation datasets are available at https://github.com/bmil-jnu/cross-species_metabolic_stability_prediction. A Google Colab notebook for running predictions without local installation is also provided in the repository.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Funding: added Ministry of Food and Drug Safety: RS-2025-02215961, RS-2024-00332003; Ministry of Science and ICT, South Korea: RS-2025-16063391

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 62 references.

Cite

This paper

Seomun, S., & Yoo, S. (2026). Cross-Species Multitask Learning with Molecular and ADME Descriptors for Liver Microsomal Metabolic Stability. Computational and structural biotechnology journal, 35(1), 0184. https://doi.org/10.34133/csbj.0184

BibTeX

@article{seomun2026cross,
author = {Seomun, Subhin and Yoo, Sunyong},
title = {{Cross-Species Multitask Learning with Molecular and ADME Descriptors for Liver Microsomal Metabolic Stability}},
journal = {Computational and structural biotechnology journal},
year = {2026},
month = aug,
volume = {35},
number = {1},
pages = {0184},
publisher = {AAAS Science Partner Journal Program},
issn = {2001-0370},
doi = {10.34133/csbj.0184},
url = {https://doi.org/10.34133/csbj.0184},
pmid = {42582689},
pmcid = {PMC13457899}
}

RIS

TY - JOUR
AU - Seomun, Subhin
AU - Yoo, Sunyong
TI - Cross-Species Multitask Learning with Molecular and ADME Descriptors for Liver Microsomal Metabolic Stability
T2 - Computational and structural biotechnology journal
J2 - Comput Struct Biotechnol J
PY - 2026
DA - 2026/08/11
VL - 35
IS - 1
SP - 0184
SN - 2001-0370
PB - AAAS Science Partner Journal Program
DO - 10.34133/csbj.0184
UR - https://doi.org/10.34133/csbj.0184
LA - en
ER -

CSL-JSON

{
"id": "10.34133/csbj.0184",
"type": "article-journal",
"title": "Cross-Species Multitask Learning with Molecular and ADME Descriptors for Liver Microsomal Metabolic Stability",
"container-title": "Computational and structural biotechnology journal",
"author": [
{
"family": "Seomun",
"given": "Subhin"
},
{
"family": "Yoo",
"given": "Sunyong"
}
],
"container-title-short": "Comput Struct Biotechnol J",
"volume": "35",
"issue": "1",
"page": "0184",
"DOI": "10.34133/csbj.0184",
"PMID": "42582689",
"PMCID": "PMC13457899",
"ISSN": "2001-0370",
"publisher": "AAAS Science Partner Journal Program",
"URL": "https://doi.org/10.34133/csbj.0184",
"language": "en",
"issued": {
"date-parts": [
[
2026,
8,
11
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.34133/csbj.0036 [code]
HYG-mol: An Interpretable Multimodal Hypergraph Framework for Molecular Property Prediction.
Journal: Computational and structural biotechnology journal
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, cellular / molecular, 3 references
[2] doi:10.1371/journal.pone.0345854 [code]
Shedding light on neural learning to rank models for anticancer drug prioritization.
Journal: PloS one
In common: RDKit, SHAP, PyTorch Geometric, 5 other tools, 1 reference
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: RDKit, SHAP, Pillow, 5 other tools, cellular / molecular
[4] doi:10.1093/bib/bbag118 [code]
Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.
Journal: Briefings in bioinformatics
In common: RDKit, PyTorch Geometric, Pillow, 5 other tools, cellular / molecular
[5] doi:10.1093/bioinformatics/btag153 [code]
MAISNet: a multi-species integrated graph neural network for acetylcholinesterase inhibitor screening.
Journal: Bioinformatics (Oxford, England)
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, 1 reference
[6] doi:10.1002/pro.70695 [code]
MIF-MAPMS: Enhancing identification of myelin autoantigenic peptides in multiple sclerosis through multimodal information fusion.
Journal: Protein science : a publication of the Protein Society
In common: RDKit, PyTorch, scikit-learn, 3 other tools, cellular / molecular, 2 references
[7] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits &lt;i&gt;Porphyromonas gingivalis&lt;/i&gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, cellular / molecular
[8] doi:10.1038/s41586-026-10670-w [code]
Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
Journal: Nature
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, cellular / molecular
[9] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, cellular / molecular
[10] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: RDKit, PyTorch Geometric, PyTorch, 4 other tools, cellular / molecular

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.