OSCR

A non-linear game for two: genetic parameters and prediction of fertilization success using Bayesian and machine learning frameworks.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § Methods › Two-tower MLP model ↔ 2tower_MLP/AC_fert_2tower.ipynb, lines 349–428 · score 0.60 · class weighted, PyTorch, epoch, binary, accuracy, Training
  2. [2] § Methods › Estimation of genetic parameters for latent fertility traits ↔ 2tower_MLP/AC_fert_2tower.ipynb, lines 349–428 · score 0.52 · latent liability, binary fertilization, matrix, sires, model

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 492 lines · 17 KB · no license · 2 matches

  1. # %%
  2. import numpy as np
  3. import pandas as pd
  4. import torch
  5. import math
  6. import torch.nn as nn
  7. import torch.nn.functional as F
  8. from sklearn.model_selection import StratifiedGroupKFold
  9. from sklearn.metrics import accuracy_score, roc_auc_score, average_precision_score, confusion_matrix, roc_curve
  10. import matplotlib.pyplot as plt
  11. from torch.utils.data import Dataset, DataLoader
  12. import warnings
  13. warnings.filterwarnings("ignore")
  14. SEED = 42
  15. np.random.seed(SEED)
  16. torch.manual_seed(SEED)
  17. # =============================
  18. # Data loading and preprocessing
  19. # =============================
  20. A_df = pd.read_csv("ApIcl.csv", index_col=0)
  21. A_df.index = (
  22. A_df.index.astype(str)
  23. .str.strip()
  24. .str.upper()
  25. .str.replace(r"\.0$", "", regex=True)
  26. )
  27. A_df.columns = A_df.columns.astype(str).str.strip().str.upper()
  28. A_df = A_df.replace([np.inf, -np.inf], np.nan).fillna(0.0).astype(np.float32)
  29. # Founders
  30. A_features_founder = {idx: row.values.astype(np.float32) for idx, row in A_df.iterrows()}
  31. def assign_founder_features(df, lookup):
  32. df = df.copy()
  33. df["Sire"] = df["Sire"].astype(str)
  34. df["Dam"] = df["Dam"].astype(str)
  35. df["Sire_features"] = df["Sire"].map(lookup)
  36. df["Dam_features"] = df["Dam"].map(lookup)
  37. return df
  38. crosses_df = pd.read_csv("EggsIcl.csv", index_col=0)
  39. crosses_df = assign_founder_features(crosses_df, A_features_founder)
  40. # Feature vector
  41. EXPECTED_DIM = A_df.shape[1]
  42. DEFAULT_VEC = np.zeros(EXPECTED_DIM, dtype=np.float32)
  43. def clean_features(x):
  44. if isinstance(x, (list, np.ndarray)):
  45. v = np.asarray(x, dtype=np.float32)
  46. if v.shape[0] == EXPECTED_DIM:
  47. return np.nan_to_num(v, nan=0.0, posinf=0.0, neginf=0.0)
  48. return DEFAULT_VEC
  49. for col in ["Sire_features", "Dam_features"]:
  50. crosses_df[col] = crosses_df[col].apply(clean_features)
  51. crosses_df = crosses_df.reset_index(drop=True)
  52. # =========
  53. # Dataloader
  54. # ============
  55. def build_inputs_matrix(df):
  56. sire = np.stack(df["Sire_features"]).astype(np.float32)
  57. dam = np.stack(df["Dam_features"]).astype(np.float32)
  58. return sire, dam
  59. class FertilityDataset(Dataset):
  60. def __init__(self, df):
  61. sire_feats, dam_feats = build_inputs_matrix(df)
  62. sire_feats = np.nan_to_num(sire_feats, nan=0.0, posinf=0.0, neginf=0.0).astype(np.float32)
  63. dam_feats = np.nan_to_num(dam_feats, nan=0.0, posinf=0.0, neginf=0.0).astype(np.float32)
  64. self.sire_input = torch.tensor(sire_feats, dtype=torch.float32)
  65. self.dam_input = torch.tensor(dam_feats, dtype=torch.float32)
  66. self.y = torch.tensor(df["BinPheno"].values, dtype=torch.float32)
  67. # IDs
  68. sire_col, dam_col, cross_col = ["Sire", "Dam", "Cohort.Sib"]
  69. self.sire_ids = (df[sire_col].astype(str).tolist() if sire_col else df.index.astype(str).tolist())
  70. self.dam_ids = (df[dam_col].astype(str).tolist() if dam_col else df.index.astype(str).tolist())
  71. self.cross_ids = (df[cross_col].astype(str).tolist() if cross_col else ["" for _ in range(len(df))])
  72. self._sire_col_name = sire_col or "SireID_fallback"
  73. self._dam_col_name = dam_col or "DamID_fallback"
  74. self._cross_col_name = cross_col or "CrossID"
  75. def __len__(self):
  76. return len(self.y)
  77. def __getitem__(self, idx):
  78. return (self.sire_input[idx],
  79. self.dam_input[idx],
  80. self.y[idx],
  81. self.sire_ids[idx],
  82. self.dam_ids[idx],
  83. self.cross_ids[idx])
  84. # =====
  85. # Model
  86. # ======
  87. class BinaryFertilityModel(nn.Module):
  88. def __init__(self, input_dim, dropout1=0.2, dropout2=0.2):
  89. super().__init__()
  90. self.sire_net = self._make_tower(input_dim, dropout1, dropout2)
  91. self.dam_net = self._make_tower(input_dim, dropout1, dropout2)
  92. @staticmethod
  93. def _make_tower(input_dim, d1, d2):
  94. return nn.Sequential(
  95. nn.Linear(input_dim, 128), nn.ReLU(), nn.Dropout(d1),
  96. nn.Linear(128, 64), nn.ReLU(), nn.Dropout(d2),
  97. nn.Linear(64, 32), nn.ReLU(),
  98. nn.Linear(32, 1)
  99. )
  100. def forward(self, sire_input, dam_input):
  101. a = self.sire_net(sire_input).squeeze(-1)
  102. b = self.dam_net(dam_input).squeeze(-1)
  103. return a, b
  104. # =========================
  105. # Helpers: probs, loss, eval
  106. # =========================
  107. def _phi(x: torch.Tensor) -> torch.Tensor:
  108. # Standard normal CDF
  109. if hasattr(torch.special, "ndtr"):
  110. return torch.special.ndtr(x)
  111. return 0.5 * (1.0 + torch.erf(x / math.sqrt(2.0)))
  112. def probs_from_scores(a, b, link="probit", stable=True, eps=1e-7):
  113. """
  114. Turn tower outputs a, b into a joint (fused) probability.
  115. """
  116. if link == "probit":
  117. if stable and hasattr(torch.special, "log_ndtr"):
  118. # log version
  119. log_p = torch.special.log_ndtr(a) + torch.special.log_ndtr(b)
  120. p = torch.exp(log_p)
  121. else:
  122. p = _phi(a) * _phi(b)
  123. elif link == "logit":
  124. if stable:
  125. # log
  126. log_p = -F.softplus(-a) + -F.softplus(-b)
  127. p = torch.exp(log_p)
  128. else:
  129. p = torch.sigmoid(a) * torch.sigmoid(b)
  130. else:
  131. raise ValueError(f"Unknown link: {link}")
  132. return torch.clamp(p, min=eps, max=1.0 - eps)
  133. def weighted_bce_on_prob(p, y, pos_weight_value, device):
  134. per_sample_nll = -(y * torch.log(p) + (1 - y) * torch.log(1 - p))
  135. pos_w = torch.as_tensor(pos_weight_value, dtype=torch.float32, device=device)
  136. w = torch.where(y > 0.5, pos_w, torch.tensor(1.0, device=device))
  137. return (per_sample_nll * w).sum() / w.sum()
  138. @torch.no_grad()
  139. def evaluate_split(model, loader, device="cpu", threshold=0.5,
  140. return_ids=False, return_scores=False):
  141. model.eval()
  142. all_probs, all_labels = [], []
  143. all_sire_ids, all_dam_ids, all_cross_ids = [], [], []
  144. all_a, all_b = [], []
  145. for batch in loader:
  146. if len(batch) == 3:
  147. sire_x, dam_x, y = batch
  148. sire_ids = dam_ids = cross_ids = None
  149. else:
  150. sire_x, dam_x, y, sire_ids, dam_ids, cross_ids = batch
  151. sire_x, dam_x = sire_x.to(device), dam_x.to(device)
  152. a, b = model(sire_x, dam_x)
  153. p_joint = probs_from_scores(a, b, link="probit", stable=True)
  154. if return_scores:
  155. all_a.extend(a.detach().cpu().numpy().ravel())
  156. all_b.extend(b.detach().cpu().numpy().ravel())
  157. all_probs.extend(p_joint.cpu().numpy().ravel())
  158. all_labels.extend(y.numpy().ravel())
  159. if return_ids and sire_ids is not None:
  160. all_sire_ids.extend(sire_ids)
  161. all_dam_ids.extend(dam_ids)
  162. all_cross_ids.extend(cross_ids)
  163. y_true = np.asarray(all_labels)
  164. y_prob = np.asarray(all_probs)
  165. y_pred = (y_prob >= threshold).astype(int)
  166. acc = accuracy_score(y_true, y_pred)
  167. if len(np.unique(y_true)) > 1:
  168. auc = roc_auc_score(y_true, y_prob)
  169. prau = average_precision_score(y_true, y_prob)
  170. else:
  171. auc, prau = np.nan, np.nan
  172. cm = confusion_matrix(y_true, y_pred)
  173. out = (acc, auc, prau, cm, y_true, y_prob)
  174. if return_ids and len(all_sire_ids) > 0:
  175. out += (np.array(all_sire_ids), np.array(all_dam_ids), np.array(all_cross_ids))
  176. if return_scores:
  177. out += (np.array(all_a), np.array(all_b)) # latent liabilities
  178. return out
  179. # =========================
  180. # Training (early stopping)
  181. # =========================
  182. def train_model(model, train_loader, val_loader, pos_weight_value,
  183. epochs=200, lr=1e-3, weight_decay=1e-4, device="cpu",
  184. patience=15, max_grad_norm=5.0, return_best_epoch=False):
  185. model.to(device)
  186. optim = torch.optim.Adam(model.parameters(), lr=lr, weight_decay=weight_decay)
  187. best_auc = -float("inf")
  188. best_state = None
  189. best_epoch = 0
  190. patience_ctr = 0
  191. for epoch in range(1, epochs + 1):
  192. model.train()
  193. # training loop
  194. for sire_x, dam_x, y, _sire_ids, _dam_ids, _cross_ids in train_loader:
  195. sire_x, dam_x, y = sire_x.to(device), dam_x.to(device), y.to(device)
  196. a, b = model(sire_x, dam_x)
  197. p = probs_from_scores(a, b, link="probit", stable=False)
  198. loss = weighted_bce_on_prob(p, y, pos_weight_value, device)
  199. if not torch.isfinite(loss):
  200. continue
  201. optim.zero_grad()
  202. loss.backward()
  203. if max_grad_norm is not None:
  204. nn.utils.clip_grad_norm_(model.parameters(), max_grad_norm)
  205. optim.step()
  206. acc_v, auc_v, pr_v, _, _, _ = evaluate_split(model, val_loader, device=device, threshold=0.5)
  207. print(f"Epoch {epoch:03d} | Val Acc {acc_v:.3f} | Val AUC {auc_v:.3f} | Val PR-AUC {pr_v:.3f}")
  208. if auc_v > best_auc:
  209. best_auc = auc_v
  210. best_epoch = epoch
  211. best_state = {k: v.detach().cpu().clone() for k, v in model.state_dict().items()}
  212. patience_ctr = 0
  213. else:
  214. patience_ctr += 1
  215. if patience_ctr >= patience:
  216. print(f"Early stopping at epoch {epoch} — best Val AUC {best_auc:.3f} (epoch {best_epoch})")
  217. break
  218. if best_state is not None:
  219. model.load_state_dict(best_state)
  220. print(f"Restored best model (Val AUC {best_auc:.3f} at epoch {best_epoch})")
  221. return (model, best_epoch) if return_best_epoch else model
  222. # ====================================
  223. # 80/20 split + inner 4-fold grid search
  224. # =======================================
  225. device = "cuda" if torch.cuda.is_available() else "cpu"
  226. input_dim = EXPECTED_DIM
  227. # Grid for first/second layer dropout
  228. dropout1_grid = [0.0, 0.1, 0.2, 0.3, 0.4, 0.5]
  229. dropout2_grid = [0.0, 0.1, 0.2, 0.3, 0.4, 0.5]
  230. # 80/20 stratified split
  231. groups_all = crosses_df["Cohort"].astype(str).values
  232. y_all = crosses_df["BinPheno"].values.astype(int)
  233. sgkf_80_20 = StratifiedGroupKFold(n_splits=5, shuffle=True, random_state=SEED)
  234. train_idx_all, test_idx = next(sgkf_80_20.split(np.zeros(len(y_all)), y_all, groups_all))
  235. train_df = crosses_df.iloc[train_idx_all].reset_index(drop=True).copy()
  236. test_df = crosses_df.iloc[test_idx].reset_index(drop=True).copy()
  237. print(f"Train size: {len(train_df)} | Test size: {len(test_df)}")
  238. assert set(train_df["Cohort"]).isdisjoint(set(test_df["Cohort"])), "Group leakage"
  239. # Inner 4-fold CV on 80% for dropout search
  240. inner_groups = train_df["Cohort"].astype(str).values
  241. inner_y = train_df["BinPheno"].values.astype(int)
  242. inner_cv = StratifiedGroupKFold(n_splits=4, shuffle=True, random_state=SEED)
  243. grid_scores = [] # (d1, d2, mean_auc, mean_prauc)
  244. epoch_book = {} # key: (d1, d2) -> list of best_epoch
  245. for d1 in dropout1_grid:
  246. for d2 in dropout2_grid:
  247. fold_metrics = []
  248. epoch_book[(d1, d2)] = []
  249. print(f"\nGrid candidate: dropout1={d1}, dropout2={d2}")
  250. for tr_idx, val_idx in inner_cv.split(np.zeros(len(inner_y)), inner_y, inner_groups):
  251. tr_df = train_df.iloc[tr_idx].reset_index(drop=True).copy()
  252. va_df = train_df.iloc[val_idx].reset_index(drop=True).copy()
  253. tr_loader = DataLoader(FertilityDataset(tr_df), batch_size=32, shuffle=True)
  254. va_loader = DataLoader(FertilityDataset(va_df), batch_size=64, shuffle=False)
  255. n_pos = int((tr_df["BinPheno"] == 1).sum())
  256. n_neg = int((tr_df["BinPheno"] == 0).sum())
  257. pos_weight_value = float(n_neg) / max(1.0, float(n_pos))
  258. model = BinaryFertilityModel(input_dim=input_dim, dropout1=d1, dropout2=d2)
  259. model, best_ep = train_model(
  260. model, tr_loader, va_loader, pos_weight_value,
  261. epochs=200, lr=1e-3, weight_decay=1e-4,
  262. device=device, patience=15, max_grad_norm=5.0,
  263. return_best_epoch=True
  264. )
  265. acc, auc, prauc, _, _, _ = evaluate_split(model, va_loader, device=device, threshold=0.5)
  266. fold_metrics.append((acc, auc, prauc))
  267. epoch_book[(d1, d2)].append(best_ep)
  268. fm = np.array(fold_metrics, dtype=float)
  269. mean_auc = np.nanmean(fm[:, 1])
  270. mean_prauc = np.nanmean(fm[:, 2])
  271. grid_scores.append((d1, d2, mean_auc, mean_prauc))
  272. print(f" 4-fold CV -> Mean AUC={mean_auc:.3f} | Mean PR-AUC={mean_prauc:.3f}")
  273. # Pick best by mean ROC AUC
  274. grid_scores.sort(key=lambda t: (t[2], t[3]), reverse=True)
  275. best_d1, best_d2, best_auc_inner, best_pra_inner = grid_scores[0]
  276. best_epochs = epoch_book[(best_d1, best_d2)]
  277. fixed_epochs = int(np.median(best_epochs)) if len(best_epochs) else 100
  278. print("\n=== GRID SEARCH SUMMARY ===")
  279. print(f"Best (dropout1, dropout2) = ({best_d1}, {best_d2})")
  280. print(f"4-fold mean AUC = {best_auc_inner:.3f} | mean PR-AUC = {best_pra_inner:.3f}")
  281. print("Grid (top 5):")
  282. for row in grid_scores[:5]:
  283. print(f" d1={row[0]} d2={row[1]} | mean AUC={row[2]:.3f} | mean PR-AUC={row[3]:.3f}")
  284. # %%
  285. # =========================
  286. # Final training on 80% with val, test on 20%
  287. # ============================================
  288. device = "cuda" if torch.cuda.is_available() else "cpu"
  289. input_dim = EXPECTED_DIM
  290. #best_d1 = 0.0
  291. #best_d2 = 0.5
  292. # Split 80% train_df into train_sub and val_sub
  293. inner_sgkf = StratifiedGroupKFold(n_splits=5, shuffle=True, random_state=SEED)
  294. tr_idx, va_idx = next(inner_sgkf.split(
  295. np.zeros(len(train_df)),
  296. train_df["BinPheno"].values,
  297. train_df["Cohort"].astype(str).values
  298. ))
  299. tr_df = train_df.iloc[tr_idx].reset_index(drop=True).copy()
  300. va_df = train_df.iloc[va_idx].reset_index(drop=True).copy()
  301. # DataLoaders
  302. tr_loader = DataLoader(FertilityDataset(tr_df), batch_size=32, shuffle=True)
  303. va_loader = DataLoader(FertilityDataset(va_df), batch_size=64, shuffle=False)
  304. te_loader = DataLoader(FertilityDataset(test_df), batch_size=64, shuffle=False)
  305. # Class weight from train_sub
  306. n_pos = int((tr_df["BinPheno"] == 1).sum())
  307. n_neg = int((tr_df["BinPheno"] == 0).sum())
  308. pos_weight_value = float(n_neg) / max(1.0, float(n_pos))
  309. # Train final model with best dropout and early stopping
  310. final_model = BinaryFertilityModel(input_dim=input_dim, dropout1=best_d1, dropout2=best_d2)
  311. final_model = train_model(
  312. final_model, tr_loader, va_loader, pos_weight_value,
  313. epochs=300, lr=1e-3, weight_decay=1e-4,
  314. device=device, patience=15, max_grad_norm=5.0)
  315. acc_t, auc_t, pr_t, cm_t, y_true_t, y_prob_t, sire_ids_t, dam_ids_t, cross_ids_t = evaluate_split(
  316. final_model, te_loader, device=device, threshold=0.5, return_ids=True)
  317. print("\nFINAL TEST RESULTS")
  318. print(f"Acc={acc_t:.3f} | AUC={auc_t:.3f} | PR-AUC={pr_t:.3f}")
  319. print("Confusion matrix:\n", cm_t)
  320. # IDs + latent scores
  321. res = evaluate_split(
  322. final_model, tr_loader, device=device, threshold=0.5,
  323. return_ids=True, return_scores=True
  324. )
  325. # Unpack
  326. (acc_t, auc_t, pr_t, cm_t,
  327. y_true_t, _y_prob_joint,
  328. sire_ids_t, dam_ids_t, cross_ids_t,
  329. a_sire_t, b_dam_t) = res
  330. # Output dataframe: IDs + latent liabilities
  331. out = {
  332. "SireID": sire_ids_t,
  333. "DamID": dam_ids_t,
  334. "a_sire": a_sire_t,
  335. "b_dam": b_dam_t,
  336. }
  337. if any(cross_ids_t):
  338. out["CrossID"] = cross_ids_t
  339. if isinstance(y_true_t, np.ndarray) and y_true_t.size == a_sire_t.size:
  340. out["y_true"] = y_true_t
  341. cols = [c for c in ["CrossID","SireID","DamID","y_true","a_sire","b_dam"] if c in out]
  342. latent_df = pd.DataFrame(out, columns=cols)
  343. latent_df.to_csv("final_test_latent_liabilities_with_ids.csv", index=False)
  344. print(latent_df.head())
  345. np.save("latents_sire_a.npy", a_sire_t)
  346. np.save("latents_dam_b.npy", b_dam_t)
  347. # %%
  348. full_loader = DataLoader(
  349. FertilityDataset(crosses_df),
  350. batch_size=64,
  351. shuffle=False
  352. )
  353. res_full = evaluate_split(
  354. final_model,
  355. full_loader,
  356. device=device,
  357. threshold=0.5,
  358. return_ids=True,
  359. return_scores=True
  360. )
  361. (acc_f, auc_f, pr_f, cm_f,
  362. y_true_f, _yprob_f,
  363. sire_ids_f, dam_ids_f, cross_ids_f,
  364. a_sire_f, b_dam_f) = res_full
  365. np.save("latents_full_sire_a.npy", a_sire_f)
  366. np.save("latents_full_dam_b.npy", b_dam_f)
  367. # %%
  368. final_model.eval()
  369. all_labels, all_probs = [], []
  370. with torch.no_grad():
  371. for sire_x, dam_x, y, _, _, _ in te_loader:
  372. sire_x, dam_x = sire_x.to(device), dam_x.to(device)
  373. a, b = final_model(sire_x, dam_x)
  374. probs = probs_from_scores(a, b, link="probit", stable=True)
  375. all_probs.extend(probs)
  376. all_labels.extend(y.numpy())
  377. all_labels = np.array(all_labels)
  378. all_probs = np.array(all_probs)
  379. pred_labels = (all_probs >= 0.5).astype(int)
  380. # Confusion matrix
  381. cm = confusion_matrix(all_labels, pred_labels)
  382. print("\nConfusion Matrix:\n", cm)
  383. # ROC + AUC
  384. auc_val = roc_auc_score(all_labels, all_probs)
  385. fpr, tpr, _ = roc_curve(all_labels, all_probs)
  386. plt.figure(figsize=(6,5))
  387. plt.plot(fpr, tpr, label=f"ROC Curve (AUC = {auc_val:.3f})")
  388. plt.plot([0, 1], [0, 1], linestyle="--")
  389. plt.xlabel("False Positive Rate")
  390. plt.ylabel("True Positive Rate")
  391. plt.title("ROC Curve — Final Test")
  392. plt.legend()
  393. plt.grid(True)
  394. plt.show()

AC_fert_2tower.ipynb at commit 9a0d895, no license · at the source

Overview

  1. Department of Animal Biosciences, Swedish University of Agricultural Sciences, Uppsala, Sweden
  2. Department of Aquaculture and Fish Biology, Hólar University, Sauðárkrókur, Iceland
Journal: Genetics, selection, evolution : GSE, volume 58, issue 1, article 34
Dates: received 10 November 2025; accepted 6 July 2026; published online 16 July 2026
Type: Brief report · Language: English
License: CC BY
Identifiers: DOI 10.1186/s12711-026-01070-9 · PMID 42464095 · PMCID PMC13377850 · OpenAlex W4416394785
Open access: gold, a free copy (OpenAlex)
Status: code verified
Methods: Machine learning, Statistics
MeSH: Fertility*, Fertilization*, Machine Learning*, Animals, Bayes Theorem, Breeding, Female, Male, Models, Genetic, Prediction Algorithms, Predictive Learning Models (* major topic)
Topic: Genetic and phenotypic traits in livestock (Genetics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Rannís (2410430)
Citations: not cited yet (Europe PMC); 30 references in the paper

Abstract

Fertility is an important but often cryptic and intrinsic characteristic of domesticated animals. Predicting reproductive potential is of great importance for the industry but assessment through indirect proxies is laborious and often impractical. Among other biological factors, genetic effects are expected to play a crucial role in shaping male and female fertility. In cases where heritable components are strong, polygenic merit could be a valuable tool for decision-making in breeding schemes. Here we estimate sex-specific variance components affecting fertilization success by analyzing outcomes of over 3000 controlled mating events in an Arctic charr breeding nucleus from Iceland. Furthermore, a machine learning framework using relationships-to-founders vectors as input and a two-tower neural network architecture is proposed and tested for prediction of fertilization success. Both approaches seem to capture a meaningful biological signal and offer alternative tools for ranking, selecting or even allocating matings between breeding candidates.

Supplementary Information: The online version contains supplementary material available at 10.1186/s12711-026-01070-9.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

pappasfotios/AC_Iceland_fertility

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 9a0d8952581e1e1458798aeee1d158cf78fdadf9, 25 September 2026
Languages: Jupyter (1), Stan (1)
Size: 6 files, 2 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, 1 notebook
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), PyTorch (1 file), scikit-learn (1 file), Stan (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
3 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 2 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

Data and code available at: https://github.com/pappasfotios/AC_Iceland_fertility.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 11 MeSH terms, 1 funder, 14 references.

Cite

This paper

Pappas, F., Debes, P. V., Johnsson, M., & Palaiokostas, C. (2026). A non-linear game for two: genetic parameters and prediction of fertilization success using Bayesian and machine learning frameworks. Genetics, selection, evolution : GSE, 58(1), 34. https://doi.org/10.1186/s12711-026-01070-9

BibTeX

@article{pappas2026non,
author = {Pappas, Fotis and Debes, Paul Vincent and Johnsson, Martin and Palaiokostas, Christos},
title = {{A non-linear game for two: genetic parameters and prediction of fertilization success using Bayesian and machine learning frameworks}},
journal = {Genetics, selection, evolution : GSE},
year = {2026},
month = jul,
volume = {58},
number = {1},
pages = {34},
publisher = {BMC},
issn = {0999-193X},
doi = {10.1186/s12711-026-01070-9},
url = {https://doi.org/10.1186/s12711-026-01070-9},
pmid = {42464095},
pmcid = {PMC13377850}
}

RIS

TY - JOUR
AU - Pappas, Fotis
AU - Debes, Paul Vincent
AU - Johnsson, Martin
AU - Palaiokostas, Christos
TI - A non-linear game for two: genetic parameters and prediction of fertilization success using Bayesian and machine learning frameworks
T2 - Genetics, selection, evolution : GSE
J2 - Genet Sel Evol
PY - 2026
DA - 2026/07/16
VL - 58
IS - 1
SP - 34
SN - 0999-193X
PB - BMC
DO - 10.1186/s12711-026-01070-9
UR - https://doi.org/10.1186/s12711-026-01070-9
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s12711-026-01070-9",
"type": "article-journal",
"title": "A non-linear game for two: genetic parameters and prediction of fertilization success using Bayesian and machine learning frameworks",
"container-title": "Genetics, selection, evolution : GSE",
"author": [
{
"family": "Pappas",
"given": "Fotis"
},
{
"family": "Debes",
"given": "Paul Vincent"
},
{
"family": "Johnsson",
"given": "Martin"
},
{
"family": "Palaiokostas",
"given": "Christos"
}
],
"container-title-short": "Genet Sel Evol",
"volume": "58",
"issue": "1",
"page": "34",
"DOI": "10.1186/s12711-026-01070-9",
"PMID": "42464095",
"PMCID": "PMC13377850",
"ISSN": "0999-193X",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s12711-026-01070-9",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
16
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-75653-x [code]
Concept2Brain: an AI model for predicting neurophysiological responses to text and pictures.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 2 other tools, 2 references
[2] doi:10.7554/elife.107423 [code]
A context-free model of savings in motor learning.
Journal: eLife
In common: PyTorch, scikit-learn, pandas, 2 other tools, 2 references
[3] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Stan, PyTorch, scikit-learn, 3 other tools
[4] doi:10.7554/elife.108223 [code]
Two time scales of adaptation in human learning rates.
Journal: eLife
In common: Stan, pandas, Matplotlib, 1 other tool, 1 reference
[5] doi:10.3390/jcm15176501 [code]
Saliency-Curated Deep Learning for Predicting Receptor Status in Breast Cancer Brain Metastases.
Journal: Journal of clinical medicine
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[6] doi:10.34133/csbj.0076 [code]
SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling.
Journal: Computational and structural biotechnology journal
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[7] doi:10.1038/s41598-026-47814-x [code]
Predicting post-stroke functional outcome using explainable machine learning and integrated data.
Journal: Scientific reports
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[8] doi:10.64898/2026.03.30.715222 [code]
Synthetic lumen rounding directs neural progenitor division mode
Journal: bioRxiv (preprint)
In common: PyTorch, scikit-learn, pandas, 2 other tools, 1 reference
[9] doi:10.1186/s13293-026-00927-4 [code]
Gene regulatory network analysis identifies dysregulation of hypoxia pathways as contributing to glioblastoma treatment resistance in females.
Journal: Biology of sex differences
In common: Stan, PyTorch, pandas, 2 other tools
[10] doi:10.1016/j.stemcr.2026.102977 [code]
Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics.
Journal: Stem cell reports
In common: Stan, scikit-learn, pandas, 2 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.