OSCR

AI solutions for evolutionary genomics of nonmodel species.

Code ↔ Paper

The paper beside its authors' code: matches between them have not been computed for this paper yet.

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 253 lines · 9 KB · MIT

  1. import numpy as np
  2. import argparse
  3. import tensorflow as tf
  4. from tensorflow import keras
  5. from tensorflow.keras.layers import (
  6. Input, Conv2D, MaxPooling2D, BatchNormalization, Activation,
  7. Dropout, Flatten, Dense, concatenate
  8. )
  9. from tensorflow.keras.models import Model
  10. from tensorflow.keras.optimizers import Adam
  11. # ============================================================
  12. # Argparse
  13. # ============================================================
  14. ### >>> CHANGE: user CLI options
  15. parser = argparse.ArgumentParser(description="Train DANN model")
  16. parser.add_argument("--src_sweep", required=True, help="Path to source sweep .npy")
  17. parser.add_argument("--src_neut", required=True, help="Path to source neutral .npy")
  18. parser.add_argument("--tgt_sweep", required=True, help="Path to target sweep .npy")
  19. parser.add_argument("--tgt_neut", required=True, help="Path to target neutral .npy")
  20. parser.add_argument("--src_train", type=int, default=5000, help="Source train samples per class")
  21. parser.add_argument("--src_val", type=int, default=1000, help="Source val samples per class")
  22. parser.add_argument("--tgt_train", type=int, default=5000, help="Target train samples per class")
  23. parser.add_argument("--tgt_val", type=int, default=1000, help="Target val samples per class")
  24. parser.add_argument("--epochs", type=int, default=50, help="Training epochs")
  25. parser.add_argument("--batch", type=int, default=32, help="Batch size")
  26. args = parser.parse_args()
  27. print("\n=== Loaded Arguments ===")
  28. print(args)
  29. # ============================================================
  30. # Gradient Reversal Layer
  31. # ============================================================
  32. @tf.custom_gradient
  33. def gradient_reversal(x, alpha):
  34. alpha = tf.cast(alpha, x.dtype)
  35. def grad(dy):
  36. return -alpha * dy, None
  37. return x, grad
  38. class GradientReversalLayer(keras.layers.Layer):
  39. def __init__(self, alpha=1.0, **kwargs):
  40. super().__init__(**kwargs)
  41. self.alpha = tf.cast(alpha, tf.float32)
  42. def call(self, x):
  43. return gradient_reversal(x, self.alpha)
  44. def set_alpha(self, alpha):
  45. self.alpha = tf.cast(alpha, tf.float32)
  46. # ============================================================
  47. # Feature extractor (YOUR CNN)
  48. # ============================================================
  49. ### >>> CHANGE: replacing previous 3-branch CNN
  50. def build_feature_extractor(inputs):
  51. m = Conv2D(32, (3,3), padding='same')(inputs)
  52. m = BatchNormalization()(m)
  53. m = Activation('relu')(m)
  54. m = MaxPooling2D((2,2), strides=2)(m)
  55. m = Conv2D(32, (3,3), padding='same')(m)
  56. m = BatchNormalization()(m)
  57. m = Activation('relu')(m)
  58. m = Flatten()(m)
  59. features = Dense(128, activation='relu')(m)
  60. return features
  61. # ============================================================
  62. # Full DANN model
  63. # ============================================================
  64. def build_dann(input_shape, alpha=1.0):
  65. inp = Input(input_shape)
  66. features = build_feature_extractor(inp)
  67. # Label predictor
  68. label_out = Dense(1, activation="sigmoid", name="label")(features)
  69. # Domain classifier
  70. grl_layer = GradientReversalLayer(alpha=alpha)
  71. grl = grl_layer(features)
  72. domain_out = Dense(2, activation="softmax", name="domain")(grl)
  73. model = Model(inputs=inp, outputs=[label_out, domain_out])
  74. return model, grl_layer
  75. # ============================================================
  76. # Load data
  77. # ============================================================
  78. sweep_src = np.load(args.src_sweep)
  79. neutral_src = np.load(args.src_neut)
  80. sweep_tgt = np.load(args.tgt_sweep)
  81. neutral_tgt = np.load(args.tgt_neut)
  82. print("\n=== Loaded Datasets ===")
  83. print("Source sweep:", sweep_src.shape)
  84. print("Source neutral:", neutral_src.shape)
  85. print("Target sweep:", sweep_tgt.shape)
  86. print("Target neutral:", neutral_tgt.shape)
  87. # ============================================================
  88. # Auto-detect image shape
  89. # ============================================================
  90. ### >>> CHANGE: automatically infer shape from .npy
  91. img_h, img_w = sweep_src.shape[1], sweep_src.shape[2]
  92. input_shape = (img_h, img_w, 1)
  93. # Add channel dimension
  94. sweep_src = sweep_src.reshape(-1, img_h, img_w, 1)
  95. neutral_src = neutral_src.reshape(-1, img_h, img_w, 1)
  96. sweep_tgt = sweep_tgt.reshape(-1, img_h, img_w, 1)
  97. neutral_tgt = neutral_tgt.reshape(-1, img_h, img_w, 1)
  98. # ============================================================
  99. # Extract counts
  100. # ============================================================
  101. Ns_tr = args.src_train
  102. Ns_val = args.src_val
  103. Nt_tr = args.tgt_train
  104. Nt_val = args.tgt_val
  105. # ============================================================
  106. # Build train/val/test splits
  107. # ============================================================
  108. # Source splits
  109. sweep_src_train = sweep_src[:Ns_tr]
  110. sweep_src_val = sweep_src[Ns_tr:Ns_tr+Ns_val]
  111. neutral_src_train = neutral_src[:Ns_tr]
  112. neutral_src_val = neutral_src[Ns_tr:Ns_tr+Ns_val]
  113. # Target splits
  114. sweep_tgt_train = sweep_tgt[:Nt_tr]
  115. sweep_tgt_val = sweep_tgt[Nt_tr:Nt_tr+Nt_val]
  116. sweep_tgt_test = sweep_tgt[Nt_tr+Nt_val:]
  117. neutral_tgt_train = neutral_tgt[:Nt_tr]
  118. neutral_tgt_val = neutral_tgt[Nt_tr:Nt_tr+Nt_val]
  119. neutral_tgt_test = neutral_tgt[Nt_tr+Nt_val:]
  120. print("\n=== Using image size:", input_shape, "===")
  121. # ============================================================
  122. # Assemble datasets
  123. # ============================================================
  124. X_train_src = np.concatenate([sweep_src_train, neutral_src_train])
  125. y_train_src = np.array([1]*Ns_tr + [0]*Ns_tr)
  126. d_train_src = np.zeros(len(X_train_src), dtype="int")
  127. X_train_tgt = np.concatenate([sweep_tgt_train, neutral_tgt_train])
  128. y_train_tgt = np.array([-1]*len(X_train_tgt))
  129. d_train_tgt = np.ones(len(X_train_tgt), dtype="int")
  130. X_train = np.concatenate([X_train_src, X_train_tgt])
  131. y_train = np.concatenate([y_train_src, y_train_tgt])
  132. d_train = np.concatenate([d_train_src, d_train_tgt])
  133. X_val_src = np.concatenate([sweep_src_val, neutral_src_val])
  134. y_val_src = np.array([1]*Ns_val + [0]*Ns_val)
  135. d_val_src = np.zeros(len(X_val_src), dtype="int")
  136. X_val_tgt = np.concatenate([sweep_tgt_val, neutral_tgt_val])
  137. y_val_tgt = np.array([-1]*len(X_val_tgt))
  138. d_val_tgt = np.ones(len(X_val_tgt), dtype="int")
  139. X_val = np.concatenate([X_val_src, X_val_tgt])
  140. y_val = np.concatenate([y_val_src, y_val_tgt])
  141. d_val = np.concatenate([d_val_src, d_val_tgt])
  142. X_test = np.concatenate([sweep_tgt_test, neutral_tgt_test])
  143. y_test = np.array([1]*len(sweep_tgt_test) + [0]*len(neutral_tgt_test))
  144. d_test = np.ones(len(X_test), dtype="int")
  145. # ============================================================
  146. # Prepare model inputs
  147. # ============================================================
  148. label_mask_train = (y_train != -1).astype("float32")
  149. label_mask_val = (y_val != -1).astype("float32")
  150. y_train_use = np.where(y_train == -1, 0, y_train).astype("float32").reshape(-1, 1)
  151. y_val_use = np.where(y_val == -1, 0, y_val).astype("float32").reshape(-1, 1)
  152. y_test_use = y_test.astype("float32").reshape(-1, 1)
  153. d_train_onehot = keras.utils.to_categorical(d_train, 2)
  154. d_val_onehot = keras.utils.to_categorical(d_val, 2)
  155. d_test_onehot = keras.utils.to_categorical(d_test, 2)
  156. X_train = X_train.astype("float32")
  157. X_val = X_val.astype("float32")
  158. X_test = X_test.astype("float32")
  159. # ============================================================
  160. # Build + compile model
  161. # ============================================================
  162. model, grl_layer = build_dann(input_shape=input_shape, alpha=1.0)
  163. model.compile(
  164. optimizer=Adam(learning_rate=0.001),
  165. loss={"label": "binary_crossentropy", "domain": "categorical_crossentropy"},
  166. loss_weights={"label": 1.0, "domain": 1.0},
  167. metrics={"label": "accuracy", "domain": "accuracy"}
  168. )
  169. # ============================================================
  170. # Training
  171. # ============================================================
  172. epochs = args.epochs
  173. batch = args.batch
  174. for epoch in range(epochs):
  175. p = epoch / epochs
  176. alpha = 2.0 / (1.0 + np.exp(-10*p)) - 1.0
  177. grl_layer.set_alpha(alpha)
  178. model.fit(
  179. X_train,
  180. [y_train_use, d_train_onehot],
  181. sample_weight=[label_mask_train, np.ones(len(X_train))],
  182. batch_size=batch,
  183. epochs=1,
  184. verbose=1
  185. )
  186. # Validation stats
  187. label_pred_src, _ = model.predict(X_val_src, verbose=0)
  188. acc_src = np.mean((label_pred_src.flatten() > 0.5) == y_val_src)
  189. _, dom_pred = model.predict(X_val, verbose=0)
  190. acc_dom = np.mean(np.argmax(dom_pred, axis=1) == d_val)
  191. print(f"EPOCH {epoch+1}/{epochs} α={alpha:.3f} SRC_ACC={acc_src:.4f} DOM_ACC={acc_dom:.4f}")
  192. # ============================================================
  193. # Test Evaluation
  194. # ============================================================
  195. test_results = model.evaluate(X_test, [y_test_use, d_test_onehot], verbose=0)
  196. print("\nFINAL TEST:")
  197. print(f"Label Loss: {test_results[1]:.4f}")
  198. print(f"Label Acc: {test_results[3]:.4f}")
  199. print(f"Domain Acc: {test_results[4]:.4f}")
  200. label_pred, domain_pred = model.predict(X_test)
  201. np.savetxt("class_pred.txt", label_pred)
  202. np.savetxt("domain_pred.txt", domain_pred)
  203. print("\nSaved predictions to class_pred.txt & domain_pred.txt")

DANN.py at commit 5462f65, under MIT · at the source

Overview

Authors: Michael DeGiorgio1,2,3, Sandipan Paul Arnab1,3, Matteo Fumagalli4,5
  1. Department of Electrical Engineering and Computer Science, Florida Atlantic University, Boca Raton, FL, United States
  2. Department of Biomedical Engineering, Florida Atlantic University, Boca Raton, FL, United States
  3. Center for Omics technologies and Data Engineering, Florida Atlantic University, Boca Raton, FL, United States
  4. School of Biological and Behavioural Sciences, Queen Mary University of London, London, United Kingdom
  5. The Alan Turing Institute, London, United Kingdom
Institutions: Florida Atlantic University (United States); Queen Mary University of London (United Kingdom); The Alan Turing Institute (United Kingdom)
Journal: Evolution letters, volume 10, issue 2, pages 135-146
Dates: received 29 October 2025; accepted 4 January 2026; published online 3 March 2026
Type: Review · Language: English
License: CC BY
Identifiers: DOI 10.1093/evlett/qrag004 · PMID 41938207 · PMCID PMC13043892 · OpenAlex W7133340451
Open access: gold, a free copy (OpenAlex)
Status: code verified
Keywords: evolutionary genomics, adaptation, population genetics, positive selection
Topic: Evolution and Genetic Dynamics (Genetics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Science Foundation (1636933, 1920920); UKRI (NE/Y003519/1); National Institutes of Health (2R25GM128590); NIGMS NIH HHS (R35 GM128590)
Citations: cited by 1 paper (Europe PMC); 90 references in the paper

Abstract

As large-scale genomic datasets are becoming abundant, new questions can now be posed in evolutionary biology. Although innovative methodological approaches are constantly developed to test these new hypotheses, their application to the study of nonmodel species is hampered by technical challenges associated with such systems. In recent years, artificial intelligence (AI) solutions, mostly in the form of deep neural networks, have been successfully introduced to analyse genomic data from nonmodel species. Here, we highlight the latest trends in deep learning to infer demographic history and signals of natural selection, and offer novel research directions to develop AI algorithms for the study of nonmodel organisms. Specifically, we identify strategies to process data missingness and uncertainty, to infer selective events in the face of unknown genomic and demographic parameters, and to generate interpretable and explainable predictions. We demonstrate our arguments by showcasing an original implementation to detect selective sweeps from an experimental setting with low sample size, uncertain sequencing data, and unknown demographic model, as typical in studies of nonmodel species. We argue that the study of nonmodel organisms is an opportunity to develop general-purpose data-driven methodologies for evolutionary inferences. Fair sharing of resources and inclusive frameworks are key to enabling researchers to benefit the most from this new wave of technologies.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above.

sandipanpaul06/DANN

License: MIT
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 5462f6532b71710a90be8d7e1d7f4907f44f90f8, 4 December 2025
Languages: Python (1)
Size: 3 files, 1 script
Software Heritage: not archived
Found in: “Data and code availability”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Keras (1 file), NumPy (1 file), TensorFlow (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
3 files

Zenodo 18166412

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data and code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • no match between paragraphs and code yet;
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data and code availability

The open-source implementation is available at https://github.com/sandipanpaul06/DANN, and the simulated replicates through Zenodo at https://doi.org/10.5281/zenodo.18166412.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 3 authors, 4 keywords, 4 funders, 73 references.

Cite

This paper

DeGiorgio, M., Arnab, S. P., & Fumagalli, M. (2026). AI solutions for evolutionary genomics of nonmodel species. Evolution letters, 10(2), 135-146. https://doi.org/10.1093/evlett/qrag004

BibTeX

@article{degiorgio2026ai,
author = {DeGiorgio, Michael and Arnab, Sandipan Paul and Fumagalli, Matteo},
title = {{AI solutions for evolutionary genomics of nonmodel species}},
journal = {Evolution letters},
year = {2026},
month = mar,
volume = {10},
number = {2},
pages = {135--146},
publisher = {Oxford University Press},
issn = {2056-3744},
doi = {10.1093/evlett/qrag004},
url = {https://doi.org/10.1093/evlett/qrag004},
pmid = {41938207},
pmcid = {PMC13043892}
}

RIS

TY - JOUR
AU - DeGiorgio, Michael
AU - Arnab, Sandipan Paul
AU - Fumagalli, Matteo
TI - AI solutions for evolutionary genomics of nonmodel species
T2 - Evolution letters
J2 - Evol Lett
PY - 2026
DA - 2026/03/03
VL - 10
IS - 2
SP - 135
EP - 146
SN - 2056-3744
PB - Oxford University Press
DO - 10.1093/evlett/qrag004
UR - https://doi.org/10.1093/evlett/qrag004
LA - en
ER -

CSL-JSON

{
"id": "10.1093/evlett/qrag004",
"type": "article-journal",
"title": "AI solutions for evolutionary genomics of nonmodel species",
"container-title": "Evolution letters",
"author": [
{
"family": "DeGiorgio",
"given": "Michael"
},
{
"family": "Arnab",
"given": "Sandipan Paul"
},
{
"family": "Fumagalli",
"given": "Matteo"
}
],
"container-title-short": "Evol Lett",
"volume": "10",
"issue": "2",
"page": "135-146",
"DOI": "10.1093/evlett/qrag004",
"PMID": "41938207",
"PMCID": "PMC13043892",
"ISSN": "2056-3744",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/evlett/qrag004",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
3
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/genetics/iyag107 [code]
Neural posterior estimation for population genetics.
Journal: Genetics
In common: NumPy, 6 references
[2] doi: [code]
Going deeper with morphologically detailed neural networks by simulation-based gradient propagation
Journal: Frontiers in computational neuroscience
In common: Keras, TensorFlow, NumPy
[3] doi: [code]
Real-time closed-loop feedback system for mouse mesoscale cortical signal and movement control
Journal: eLife
In common: Keras, TensorFlow, NumPy
[4] doi:10.1371/journal.pone.0356243 [code]
Functional organization and natural scene responses across mouse visual cortical areas revealed with encoding manifolds.
Journal: PloS one
In common: Keras, TensorFlow, NumPy
[5] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: Keras, TensorFlow, NumPy
[6] doi:10.1162/imag.a.1366 [code]
MICAFlow: Fast and robust MRI preprocessing bridging research neuroimaging and clinical practice.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: Keras, TensorFlow, NumPy
[7] doi:10.7554/elife.104053 [code]
Brain cognition gaps reveal associations with dopamine and factors related to brain health through artificial intelligence prediction of functional connectome.
Journal: eLife
In common: Keras, TensorFlow, NumPy
[8] doi:10.1093/neuonc/noag128 [code]
Spatially-resolved single-cell imaging of melanoma brain metastases identifies localized immune patterns predictive of immune checkpoint blockade response.
Journal: Neuro-oncology
In common: Keras, TensorFlow, NumPy
[9] doi:10.1126/sciadv.aee6952 [code]
Wafer-scale SOT-MRAM for analog crossbar array applications.
Journal: Science advances
In common: Keras, TensorFlow, NumPy
[10] doi:10.1126/sciadv.aeh7220 [code]
Central complex representations of self-movement are sufficient to compute wind direction in flight.
Journal: Science advances
In common: Keras, TensorFlow, NumPy

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.