Neural posterior estimation for population genetics.
The 21 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § Methods › Population genetic tasks › Comparisons to MSMC2 ↔ experiments/variable-popn-size/baselines/msmc2_utils.py, lines 279–351 · score 0.91 · EM iterations, segment pattern, diploid individuals, fixed recombination, mutation rate, tree sequences
- [2] § Methods › Population genetic tasks › Embedding networks and summary statistics ↔ workflow/scripts/ts_processors.py, lines 325–410 · score 0.89 · Rogers Huff, distance bin, SNP pairs, gap, logarithmically, r2
- [3] § Methods › Population genetic tasks › Embedding networks and summary statistics ↔ workflow/scripts/embedding_networks.py, lines 328–397 · score 0.75 · bi directional GRU, dropout, PyTorch, vector, concatenated, layer
- [4] § Methods › Population genetic tasks › Embedding networks and summary statistics ↔ workflow/scripts/embedding_networks.py, lines 71–179 · score 0.71 · permutation invariance, exchangeable CNN, haplotype, LD, SNPs, genotypes
- [5] § Results › Inference of population bottleneck parameters ↔ workflow/scripts/ts_simulators.py, lines 163–212 · score 0.70 · South Middle Atlas, Arabidopsis thaliana, demographic model, stdpopsim, event, population
- [6] § Methods › Population genetic tasks › Comparisons to ABC ↔ experiments/aratha-2epoch-extra/abc_utils.py, lines 274–334 · score 0.66 · closest simulations, ABC rejection, training simulations, posterior sample, quantile, NPE
- [7] § Results › Inference of population bottleneck parameters ↔ experiments/aratha-2epoch-extra/plot-posterior-surfaces.py, lines 20–73 · score 0.63 · npe cnn, npe rnn, npe sfs, curvature, surfaces, match
- [8] § Results › Inference of population bottleneck parameters ↔ experiments/aratha-2epoch-extra/plot-posterior-surfaces.py, lines 20–73 · score 0.63 · npe cnn, npe rnn, npe sfs, surface, rows, likelihood
- [9] § Methods › Population genetic tasks › Application to Drosophila melanogaster ↔ workflow/scripts/ts_simulators.py, lines 95–160 · score 0.62 · ancestral population, isolation, haploid, CO, FR, migration
- [10] § Results › D. melanogaster out-of-Africa model ↔ experiments/variable-popn-size/compare-posteriors.py, lines 325–370 · score 0.59 · marginal distributions, generations ago, credible intervals, posterior samples, model, population
- [11] § Results › D. melanogaster out-of-Africa model ↔ experiments/dependent-variable-popn-size/compare-posteriors.py, lines 333–377 · score 0.59 · marginal distributions, generations ago, credible intervals, posterior samples, model, population
- [12] § Results › Inference of historical population sizes in a one-population-model ↔ experiments/dependent-variable-popn-size/compare-posteriors.py, lines 488–543 · score 0.57 · credible interval width, CI width, geometric, recombination rate, log, models
- [13] § Results › Inference of historical population sizes in a one-population-model ↔ experiments/variable-popn-size/compare-posteriors.py, lines 467–522 · score 0.57 · credible interval width, CI width, geometric, recombination rate, log, models
- [14] § Results › D. melanogaster out-of-Africa model ↔ depr/amortized_dadi_workflow/scripts/dadi_simulators.py, lines 15–84 · score 0.56 · migration rates, ancestral population, demographic model, split, growth
- [15] § Results › Inference of historical population sizes in a one-population-model ↔ workflow/scripts/ts_simulators.py, lines 380–519 · score 0.56 · 100–100000, uniform distribution, tree sequences, log10, msprime, spaced
- [16] § Results › Inference of population bottleneck parameters ↔ depr/amortized_dadi_workflow/scripts/dadi_simulators.py, lines 86–125 · score 0.56 · Arabidopsis thaliana, catalog, demographic model, stdpopsim, event, population
- [17] § Methods › Population genetic tasks › Application to Drosophila melanogaster ↔ depr/amortized_dadi_workflow/scripts/dadi_simulators.py, lines 15–84 · score 0.54 · migration rates, ancestral population, split, growth, uniform, Demographic
- [18] § Methods › NPE workflow ↔ experiments/dependent-variable-popn-size/simulate-for-posterior.py, lines 98–220 · score 0.52 · simulating tree sequences, model parameters, Python, recombination rate, raw, flow
- [19] § Methods › NPE workflow ↔ experiments/variable-popn-size/simulate-for-posterior.py, lines 224–366 · score 0.52 · simulating tree sequences, model parameters, Python, recombination rate, raw, flow
- [20] § Results › Inference of historical population sizes in a one-population-model ↔ experiments/variable-popn-size/run-ne-comparison-three-scenarios.sh, the whole file · a weak match · score 0.51 · linear decline, linear growth, CNN, RNN, population, model
- [21] § Methods › Population genetic tasks › Application to Drosophila melanogaster ↔ workflow/scripts/ts_simulators.py, lines 95–160 · score 0.50 · mutation rate, recombination rate, Drosophila, melanogaster, uniform, demographic
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 519 lines · 20 KB · MIT · 4 matches
- ## ts_simulators outputs tree sequence. This cannot be used as a simulator for simulate_for_sbi!
- import tskit
- import msprime
- import demes
- import torch
- import numpy as np
- import stdpopsim
- from sbi.utils import BoxUniform
- class BaseSimulator:
- def __init__(self, config: dict, default: dict):
- for key in config:
- if key == "class_name": continue
- assert key in default, f"Option {key} not available for simulator"
- for key, default in default.items():
- setattr(self, key, config.get(key, default))
- class YRI_CEU(BaseSimulator):
- """
- Simulate a model defined in dadi's manual
- (https://dadi.readthedocs.io/en/latest/examples/YRI_CEU/YRI_CEU/).
- Ancestral population changes in size; splits into YRI/CEU; CEU population
- undergoes a bottleneck upon splitting and subsequently grows exponentially.
- There is continuous symmetric migration between YRI and CEU.
- """
- default_config = {
- # FIXED PARAMETERS
- "samples": {"YRI": 10, "CEU": 10},
- "sequence_length": 10e6,
- "recombination_rate": 1.5e-8,
- "mutation_rate": 1.5e-8,
- # RANDOM PARAMETERS (UNIFORM)
- "N_A": [1e2, 1e5],
- "N_YRI": [1e2, 1e5],
- "N_CEU_initial": [1e2, 1e5],
- "N_CEU_final": [1e2, 1e5],
- "M": [0, 5e-4],
- "Tp": [0, 6e4],
- "T": [0, 6e4],
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- self.parameters = ["N_A", "N_YRI", "N_CEU_initial", "N_CEU_final", "M", "Tp", "T"]
- self.prior = BoxUniform(
- low=torch.tensor([getattr(self, p)[0] for p in self.parameters]),
- high=torch.tensor([getattr(self, p)[1] for p in self.parameters]),
- )
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- torch.manual_seed(seed)
- theta = self.prior.sample().numpy()
- N_A, N_YRI, N_CEU_initial, N_CEU_final, M, Tp, T = theta
- graph = demes.Builder()
- graph.add_deme(
- "ancestral",
- epochs=[dict(start_size=N_A, end_time=Tp + T)]
- )
- graph.add_deme(
- "AMH",
- ancestors=["ancestral"],
- epochs=[dict(start_size=N_YRI, end_time=T)],
- )
- graph.add_deme(
- "CEU",
- ancestors=["AMH"],
- epochs=[dict(start_size=N_CEU_initial, end_size=N_CEU_final)],
- )
- graph.add_deme(
- "YRI",
- ancestors=["AMH"],
- epochs=[dict(start_size=N_YRI)],
- )
- graph.add_migration(demes=["CEU", "YRI"], rate=M)
- demog = msprime.Demography.from_demes(graph.resolve())
- ts = msprime.sim_ancestry(
- self.samples,
- demography=demog,
- sequence_length=self.sequence_length,
- recombination_rate=self.recombination_rate,
- random_seed=seed,
- )
- ts = msprime.sim_mutations(ts, rate=self.mutation_rate, random_seed=seed)
- return ts, theta
- class DroMel_CO_FR(BaseSimulator):
- """
- Simulate a two-population isolation-with-migration model for
- Drosophila melanogaster, Congolese and French populations,
- with growth in the French population
- """
- default_config = {
- # FIXED PARAMETERS
- "samples": {"CO": 10, "FR": 10},
- "sequence_length": 1e6,
- "recombination_rate": [0.5e-8, 3.0e-8],
- "mutation_rate": 5.49e-9,
- "haploid": True,
- # RANDOM PARAMETERS (UNIFORM IN LOGSPACE)
- "N_ANC": [3.5, 6.5], # log10 ancestral population size
- "N_CO": [3.5, 6.5], # log10 Congolese population size
- "N_FR0": [3.5, 6.5], # log10 French population size after split
- "N_FR1": [3.5, 6.5], # log10 French population size after growth
- "T": [3, 6], # log10 split time
- "m_CO_FR": [-8, -3], # log10 migration Congolese to French
- "m_FR_CO": [-8, -3], # log10 migration French to Congolese
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- self.parameters = ["N_ANC", "N_CO", "N_FR0", "N_FR1", "T", "m_CO_FR", "m_FR_CO"]
- self.prior = BoxUniform(
- low=torch.tensor([getattr(self, p)[0] for p in self.parameters]),
- high=torch.tensor([getattr(self, p)[1] for p in self.parameters]),
- )
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- torch.manual_seed(seed)
- theta = self.prior.sample().numpy()
- seeds = torch.randint(2 ** 32 - 1, (2, )).numpy()
- r = self.recombination_rate[0] + \
- (self.recombination_rate[1] - self.recombination_rate[0]) * \
- torch.rand(1).item()
- N_ANC, N_CO, N_FR0, N_FR1, T, m_CO_FR, m_FR_CO = (10 ** theta)
- G_FR = np.log(N_FR1 / N_FR0) / T
- demogr = msprime.Demography()
- demogr.add_population(name="CO", initial_size=N_CO)
- demogr.add_population(name="FR", initial_size=N_FR1, growth_rate=G_FR)
- demogr.add_population(name="ANC", initial_size=N_ANC)
- demogr.migration_matrix = np.array([[0, m_CO_FR, 0], [m_FR_CO, 0, 0], [0, 0, 0]])
- demogr.add_population_split(time=T, derived=["CO", "FR"], ancestral="ANC")
- samples = [
- msprime.SampleSet(n, population=p, ploidy=1 if self.haploid else 2)
- for p, n in self.samples.items()
- ]
- ts = msprime.sim_ancestry(
- samples,
- demography=demogr,
- sequence_length=self.sequence_length,
- recombination_rate=r,
- random_seed=seeds[0],
- )
- ts = msprime.sim_mutations(
- ts,
- rate=self.mutation_rate,
- random_seed=seeds[1],
- )
- return ts, theta
- class AraTha_2epoch(BaseSimulator):
- """
- Simulate the African2Epoch_1H18 model from stdpopsim for Arabidopsis thaliana.
- The model consists of a single population that undergoes a size change.
- """
- species = stdpopsim.get_species("AraTha")
- model = species.get_demographic_model("African2Epoch_1H18")
- default_config = {
- # FIXED PARAMETERS
- "samples": {"SouthMiddleAtlas": 10},
- "sequence_length": 10e6,
- # RANDOM PARAMETERS (UNIFORM)
- "nu": [0.01, 1], # Ratio of current to ancestral population size
- "T": [0.01, 1.5], # Time of size change (scaled)
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- self.parameters = ["nu", "T"]
- self.prior = BoxUniform(
- low=torch.tensor([getattr(self, p)[0] for p in self.parameters]),
- high=torch.tensor([getattr(self, p)[1] for p in self.parameters]),
- )
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- torch.manual_seed(seed)
- theta = self.prior.sample().numpy()
- nu, T = theta
- species = self.species
- contig = species.get_contig(
- length=self.sequence_length,
- )
- model = self.model
- # Scale the population size and time parameters
- N_A = model.model.events[0].initial_size # ancestral population size
- model.populations[0].initial_size = nu * N_A # current population size
- model.model.events[0].time = T * 2 * N_A # time of size change
- engine = stdpopsim.get_engine("msprime")
- ts = engine.simulate(
- model,
- contig,
- samples=self.samples,
- random_seed=seed
- )
- return ts, theta
- class VariablePopulationSize(BaseSimulator):
- """
- Simulate a model with varying population size across multiple time windows.
- The model consists of a single population that undergoes multiple size changes.
- """
- default_config = {
- # FIXED PARAMETERS
- "samples": {"pop": 10},
- "sequence_length": 10e6,
- "mutation_rate": 1.5e-8,
- "num_time_windows": 3,
- # RANDOM PARAMETERS (UNIFORM)
- "pop_sizes": [1e2, 1e5], # Range for population sizes (log10 space)
- "recomb_rate": [1e-9, 2e-8], # Range for recombination rate
- # TIME PARAMETERS
- "max_time": 100000, # Maximum time for population events
- "time_rate": 0.1, # Rate at which time changes across windows
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- # Set up parameters list
- self.parameters = [f"N_{i}" for i in range(self.num_time_windows)] + ["recomb_rate"]
- # Create parameter ranges in the same format as AraTha_2epoch
- # Population sizes (in log10 space)
- pop_size_ranges = [[np.log10(self.pop_sizes[0]), np.log10(self.pop_sizes[1])]
- for _ in range(self.num_time_windows)]
- # Add recombination rate range
- param_ranges = pop_size_ranges + [self.recomb_rate]
- # Set up prior using BoxUniform
- self.prior = BoxUniform(
- low=torch.tensor([r[0] for r in param_ranges]),
- high=torch.tensor([r[1] for r in param_ranges])
- )
- # Calculate fixed time points for population size changes
- self.change_times = self._calculate_change_times()
- def _calculate_change_times(self) -> np.ndarray:
- """Calculate the times at which population size changes occur using an exponential spacing."""
- #times = [(np.exp(np.log(1 + self.time_rate * self.max_time) * i /
- # (self.num_time_windows - 1)) - 1) / self.time_rate
- # for i in range(self.num_time_windows)]
- win = np.logspace(2, np.log10(self.max_time), self.num_time_windows)
- win[0] = 0
- times = win
- return np.around(times).astype(int)
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- if seed is not None:
- torch.manual_seed(seed)
- min_snps = 400
- max_attempts = 100
- attempt = 0
- while attempt < max_attempts:
- # Sample parameters directly from prior (like AraTha_2epoch)
- theta = self.prior.sample().numpy()
- # Convert population sizes from log10 space
- pop_sizes = 10 ** theta[:-1] # All but last element are population sizes
- recomb_rate = theta[-1] # Last element is recombination rate
- # Create demography
- demography = msprime.Demography()
- demography.add_population(name="pop0", initial_size=float(pop_sizes[0]))
- # Add population size changes at calculated time intervals
- for i in range(1, len(pop_sizes)):
- demography.add_population_parameters_change(
- time=self.change_times[i],
- initial_size=float(pop_sizes[i]),
- growth_rate=0,
- population="pop0"
- )
- # Simulate ancestry
- ts = msprime.sim_ancestry(
- samples={"pop0": self.samples["pop"]},
- demography=demography,
- sequence_length=self.sequence_length,
- recombination_rate=recomb_rate,
- random_seed=seed
- )
- # Add mutations
- ts = msprime.sim_mutations(ts, rate=self.mutation_rate, random_seed=seed)
- # Check if we have enough SNPs after MAF filtering
- geno = ts.genotype_matrix().T
- num_sample = geno.shape[0]
- if (geno==2).any():
- num_sample *= 2
- row_sum = np.sum(geno, axis=0)
- keep = np.logical_and.reduce([
- row_sum != 0,
- row_sum != num_sample,
- row_sum > num_sample * 0.05,
- num_sample - row_sum > num_sample * 0.05
- ])
- if np.sum(keep) >= min_snps:
- return ts, theta
- attempt += 1
- if seed is not None:
- seed += 1
- raise RuntimeError(f"Failed to generate tree sequence with at least {min_snps} SNPs after {max_attempts} attempts")
- class recombination_rate(BaseSimulator):
- """
- Simulate a one population model where recombination rate varies
- among replicates. The prior is a beta distribution shifted/scaled
- to a given interval (by default, a noninformative beta).
- """
- default_config = {
- # FIXED PARAMETERS
- "samples": {0: 10},
- "sequence_length": 1e6,
- "mutation_rate": 1.5e-8,
- "pop_size": 1e4,
- "mean_and_dispersion": [0.5, 0.5], # beta mean and dispersion
- # RANDOM PARAMETERS (UNIFORM)
- "recombination_rate": [0, 1e-8], # bounds on recombination rate
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- self.parameters = ["recombination_rate"]
- self.prior = BoxUniform(
- low=torch.tensor([getattr(self, p)[0] for p in self.parameters]),
- high=torch.tensor([getattr(self, p)[1] for p in self.parameters]),
- )
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- torch.manual_seed(seed)
- mean, dispersion = self.mean_and_dispersion
- alpha, beta = mean / dispersion, (1 - mean) / dispersion
- prior = torch.distributions.Beta(torch.FloatTensor([alpha]), torch.FloatTensor([beta]))
- low, high = self.recombination_rate
- theta = low + (high - low) * prior.sample().numpy()
- recombination_rate = theta.item()
- ts = msprime.sim_ancestry(
- self.samples,
- population_size=self.pop_size,
- sequence_length=self.sequence_length,
- recombination_rate=recombination_rate,
- random_seed=seed,
- discrete_genome=False, # don't want overlapping mutations
- )
- ts = msprime.sim_mutations(ts, rate=self.mutation_rate, random_seed=seed)
- return ts, theta
- class DependentVariablePopulationSize(BaseSimulator):
- """
- Simulate a population with variable population size across multiple time windows, with each
- population size dependent on the previous one.
- The model consists of a single population that undergoes multiple size changes.
- """
- default_config = {
- # FIXED PARAMETERS
- "samples": {"pop0": 25},
- "sequence_length": 2e6,
- "mutation_rate": 1e-8,
- "num_time_windows": 21,
- "maf": 0.05,
- # RANDOM PARAMETERS (UNIFORM)
- "pop_sizes": [1e2, 1e5], # Range for population sizes (log10 space)
- "pop_changes": [-1, 1], # Range for population size changes (* 10 ** beta)
- "recomb_rate": [1e-9, 1e-8], # Range for recombination rate
- # TIME PARAMETERS
- "max_time": 130000, # Maximum time for population events
- "time_rate": 0.1, # Rate at which time changes across windows
- }
- def __init__(self, config: dict):
- super().__init__(config, self.default_config)
- # Set up parameters list
- self.parameters = [f"N_{i}" for i in range(self.num_time_windows)] + ["recomb_rate"]
- # Create parameter ranges in the same format as AraTha_2epoch
- # Population sizes (in log10 space)
- pop_size_range = [[np.log10(self.pop_sizes[0]), np.log10(self.pop_sizes[1])]
- for _ in range(self.num_time_windows)]
- pop_change_ranges = [[self.pop_changes[0], self.pop_changes[1]]
- for _ in range(1,self.num_time_windows)]
- # Add recombination rate range
- param_ranges = pop_size_range + [self.recomb_rate]
- # Set up prior using BoxUniform
- self.prior = BoxUniform(
- low=torch.tensor([r[0] for r in param_ranges]),
- high=torch.tensor([r[1] for r in param_ranges])
- )
- self.beta_prior = BoxUniform(
- low=torch.tensor([b[0] for b in pop_change_ranges]),
- high=torch.tensor([b[1] for b in pop_change_ranges])
- )
- # Calculate fixed time points for population size changes
- self.change_times = self._calculate_change_times()
- def _calculate_change_times(self) -> np.ndarray:
- """Calculate the times at which population size changes occur using an exponential spacing."""
- times = [(np.exp(np.log(1 + self.time_rate * self.max_time) * i /
- (self.num_time_windows - 1)) - 1) / self.time_rate
- for i in range(self.num_time_windows)]
- return np.around(times).astype(int)
- def _generate_dependent_pop_sizes(self) -> np.ndarray:
- """
- Generate a sequence of population sizes where the first population size N_1
- is sampled from a uniform distribution corresponding to the most recent
- time window. The following population sizes are generated following
- N_i = N_{i-1} * 10 ^ β for i in [2,...,num_time_windows], unless N_i
- is outside of pop_ranges. If so N_i is set to the max/min population size
- """
- prior_sample = self.prior.sample().numpy()
- # only the first sample from prior will be used
- beta_sample = self.beta_prior.sample().numpy()
- modified_prior = prior_sample.copy()
- # The first value is uniformly sampled within the log10 bounds
- for i in range(self.num_time_windows-1):
- # For subsequent time windows, calculate the new value based on the previous one and beta
- new_value = modified_prior[i] + beta_sample[i]
- # If the new value is outside the bounds, set it to the max/min
- if new_value > np.log10(self.pop_sizes[1]):
- new_value = np.log10(self.pop_sizes[1])
- if new_value < np.log10(self.pop_sizes[0]):
- new_value = np.log10(self.pop_sizes[0])
- modified_prior[i+1] = new_value
- # Return the sampled prior and recombination rate
- return modified_prior
- def __call__(self, seed: int = None) -> (tskit.TreeSequence, np.ndarray):
- if seed is not None:
- torch.manual_seed(seed)
- min_snps = 400
- max_attempts = 100
- attempt = 0
- while attempt < max_attempts:
- # Sample parameters directly from prior (like AraTha_2epoch)
- theta = self._generate_dependent_pop_sizes()
- pop_sizes = 10**theta[:-1] # All but last element are for the population sizes
- recomb_rate = theta[-1] # Last element is recombination rate
- # Create demography
- demography = msprime.Demography()
- demography.add_population(name="pop0", initial_size=float(pop_sizes[0]))
- # Add population size changes at calculated time intervals
- for i in range(1, len(pop_sizes)):
- demography.add_population_parameters_change(
- time=self.change_times[i],
- initial_size=float(pop_sizes[i]),
- growth_rate=0,
- population="pop0"
- )
- # Simulate ancestry
- ts = msprime.sim_ancestry(
- samples=self.samples,
- demography=demography,
- sequence_length=self.sequence_length,
- recombination_rate=recomb_rate,
- random_seed=seed
- )
- # Add mutations
- ts = msprime.sim_mutations(ts, rate=self.mutation_rate, random_seed=seed)
- # Check if we have enough SNPs after MAF filtering
- geno = ts.genotype_matrix().T
- num_sample = ts.num_samples
- row_sum = np.sum(geno, axis=0)
- keep = np.logical_and.reduce([
- row_sum > num_sample * self.maf,
- num_sample - row_sum > num_sample * self.maf
- ])
- if np.sum(keep) >= min_snps:
- return ts, theta
- attempt += 1
- if seed is not None:
- seed += 1
- raise RuntimeError(f"Failed to generate tree sequence with at least {min_snps} SNPs after {max_attempts} attempts")
ts_simulators.py at commit e81dfe6, under MIT · at the source
Overview
- Institute of Ecology and Evolution, University of Oregon, Eugene, OR 97403, United States
- Institute for Bioinformatics and Medical Informatics (IBMI), University of Tübingen, Tübingen 72076, Germany
- Big Data Analytics in Bioinformatics, Justus-Liebig University, Gießen 35390, Germany
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repository
Its files are read in the Code ↔ Paper reader above, with 21 matches between paragraphs and lines of code.
kr-colab/popgen-npe
e81dfe66d32ff194e42c3c68941e647f824e0a32, 22 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
64 files
- depr/
amortized_dadi_workflow/ , Python, 184 lines, 3 matchesscripts/ dadi_simulators.py - depr/
amortized_dadi_workflow/ , Python, 25 linesscripts/ plot_2d_comp_multinom.py - depr/
amortized_dadi_workflow/ , Python, 36 linesscripts/ plot_confidence_interval s.py - depr/
amortized_dadi_workflow/ , Python, 49 linesscripts/ plotting.py - depr/
amortized_dadi_workflow/ , Python, 22 linesscripts/ posterior_ensemble.py - depr/
amortized_dadi_workflow/ , Python, 17 linesscripts/ simulate_default.py - depr/
amortized_dadi_workflow/ , Python, 22 linesscripts/ simulate_from_posterior. py - depr/
amortized_dadi_workflow/ , Python, 28 linesscripts/ simulate_fs.py - depr/
amortized_dadi_workflow/ , Python, 88 linesscripts/ train_npe.py - depr/
plots/ , Jupyter, 238 linesAraTha/ compare_to_dadi_2epoch_m odel.ipynb - docs/
conf.py , Python, 55 lines - experiments/
aratha-2epoch-extra/ , Python, 334 lines, 1 matchabc_utils.py - experiments/
aratha-2epoch-extra/ , Python, 738 linescompare-posteriors.py - experiments/
aratha-2epoch-extra/ , Shell, 13 linesexperiment.sh - experiments/
aratha-2epoch-extra/ , Python, 73 lines, 2 matchesplot-posterior-surfaces. py - experiments/
aratha-2epoch/ , Python, 334 linesabc_utils.py - experiments/
aratha-2epoch/ , Python, 727 linescompare-posteriors.py - experiments/
aratha-2epoch/ , Shell, 13 linesexperiment.sh - experiments/
aratha-2epoch/ , Python, 73 linesplot-posterior-surfaces. py - experiments/
dependent-variable-popn- , Python, 264 linessize/ compare-analysis.py - experiments/
dependent-variable-popn- , Python, 638 lines, 2 matchessize/ compare-posteriors.py - experiments/
dependent-variable-popn- , Python, 189 linessize/ plot-6-scenarios.py - experiments/
dependent-variable-popn- , Shell, 130 linessize/ run-ne-comparison-six-sc enarios.sh - experiments/
dependent-variable-popn- , Shell, 33 linessize/ simulate-examples.sh - experiments/
dependent-variable-popn- , Python, 224 lines, 1 matchsize/ simulate-for-posterior.p y - experiments/
dmel-isolation/ , Shell, 17 linesexperiment.sh - experiments/
dmel-isolation/ , Shell, 114 linespackage-data.sh - experiments/
dmel-isolation/ , Python, 342 linesplot-joint-posteriors.py - experiments/
dmel-isolation/ , Python, 224 linesplot-rescaled-posteriors .py - experiments/
dmel-isolation/ , Python, 230 linesplot-summary-stats.py - experiments/
relernn-prior/ , Shell, 17 linesexperiment.sh - experiments/
relernn-prior/ , Python, 111 linesplot_coverage_comparison .py - experiments/
relernn-prior/ , Python, 204 linespredict_on_true.py - experiments/
variable-popn-size/ , Python, 351 lines, 1 matchbaselines/ msmc2_utils.py - experiments/
variable-popn-size/ , Python, 112 linesbaselines/ run_baselines.py - experiments/
variable-popn-size/ , Python, 253 linescompare-analysis.py - experiments/
variable-popn-size/ , Python, 617 lines, 2 matchescompare-posteriors.py - experiments/
variable-popn-size/ , Python, 200 linescreate-3x3-panel-with-ba selines.py - experiments/
variable-popn-size/ , Python, 189 linescreate-3x3-panel.py - experiments/
variable-popn-size/ , Shell, 86 lines, 1 matchrun-ne-comparison-three- scenarios.sh - experiments/
variable-popn-size/ , Shell, 61 linessimulate-examples.sh - experiments/
variable-popn-size/ , Python, 370 lines, 1 matchsimulate-for-posterior.p y - resources/
util/ , Python, 141 linessimulate-vcf.py - tests/
test_polarise.py , Python, 110 lines - tests/
test_workflow_script.py , Python, 381 lines - workflow/
scripts/ , Python, 156 linesdata_handlers.py - workflow/
scripts/ , Python, 397 lines, 2 matchesembedding_networks.py - workflow/
scripts/ , Python, 60 linesinfer_batch.py - workflow/
scripts/ , Python, 230 linesplot_diagnostics.py - workflow/
scripts/ , Python, 105 linesplot_simulation_stats.py - workflow/
scripts/ , Python, 84 linesplot_tree_stats.py - workflow/
scripts/ , Python, 141 linespredict_windows.py - workflow/
scripts/ , Python, 27 linesprocess_batch.py - workflow/
scripts/ , Python, 156 linessetup_prediction.py - workflow/
scripts/ , Python, 54 linessetup_training.py - workflow/
scripts/ , Python, 32 linessimulate_batch.py - workflow/
scripts/ , Python, 190 linestrain_embedding_network. py - workflow/
scripts/ , Python, 143 linestrain_npe_on_embeddings. py - workflow/
scripts/ , Python, 142 linestrain_npe_on_features.py - workflow/
scripts/ , Python, 462 lines, 1 matchts_processors.py - workflow/
scripts/ , Python, 519 lines, 4 matchests_simulators.py - workflow/
scripts/ , Python, 53 linesutils.py - LICENSE, License, 21 lines
- README.md, Text, 110 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 62 scripts, each with its path and the digest of its content;
- 21 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Code and data availability statement
The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: kr-colab/
popgen-npe
Read it in the paper: doi.org/10.1093/genetics/iyag107.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Publisher: n/a → Oxford University Press
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 5 keywords, 8 MeSH terms, 7 funders, 61 references.
Cite
This paper
Min, J., Ning, Y., Pope, N. S., Baumdicker, F., & Kern, A. D. (2026). Neural posterior estimation for population genetics. Genetics, 233(3), iyag107. https://
BibTeX
@article{min2026neural,
author = {Min, Jiseon and Ning, Yuxin and Pope, Nathaniel S and Baumdicker, Franz and Kern, Andrew D},
title = {{Neural posterior estimation for population genetics}},
journal = {Genetics},
year = {2026},
month = jul,
volume = {233},
number = {3},
pages = {iyag107},
publisher = {Oxford University Press},
issn = {0016-6731},
doi = {10.1093/
url = {https://
pmid = {42032815},
pmcid = {PMC13334119}
}
RIS
TY - JOUR
AU - Min, Jiseon
AU - Ning, Yuxin
AU - Pope, Nathaniel S
AU - Baumdicker, Franz
AU - Kern, Andrew D
TI - Neural posterior estimation for population genetics
T2 - Genetics
J2 - Genetics
PY - 2026
DA - 2026/
VL - 233
IS - 3
SP - iyag107
SN - 0016-6731
PB - Oxford University Press
DO - 10.1093/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1093/
"type": "article-journal",
"title": "Neural posterior estimation for population genetics",
"container-title": "Genetics",
"author": [
{
"family": "Min",
"given": "Jiseon"
},
{
"family": "Ning",
"given": "Yuxin"
},
{
"family": "Pope",
"given": "Nathaniel S"
},
{
"family": "Baumdicker",
"given": "Franz"
},
{
"family": "Kern",
"given": "Andrew D"
}
],
"container-title-short":
"volume": "233",
"issue": "3",
"page": "iyag107",
"DOI": "10.1093/
"PMID": "42032815",
"PMCID": "PMC13334119",
"ISSN": "0016-6731",
"publisher": "Oxford University Press",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s44400-026-00094-8 [code]
- Haplotype-resolved DNA methylation at the &
lt;i& gt;APOE& lt;/ i& gt; locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk. Journal: NPJ dementiaIn common: Snakemake, BCFtools, pysam, 6 other tools - [2] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: Snakemake, BCFtools, pysam, 6 other tools
- [3] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: PyMC, BCFtools, PyTorch Lightning, 6 other tools
- [4] doi:10.1093/evlett/qrag004 [code]
- AI solutions for evolutionary genomics of nonmodel species.Journal: Evolution lettersIn common: NumPy, 6 references
- [5] doi:10.1038/s41467-026-71790-5 [code]
- Recurrent DNA break clusters drive replication-stress-induc
ed copy number variants and genome diversification. Journal: Nature communicationsIn common: Snakemake, BCFtools, pysam, 5 other tools - [6] doi:10.1016/j.xcrm.2026.102904 [code]
- High-dose furmonertinib as first-line treatment for untreated EGFR-mutated advanced NSCLC with central nervous system metastases: A phase 2 trial.Journal: Cell reports. MedicineIn common: PyMC, pysam, PyTorch Lightning, 4 other tools
- [7] doi:10.1371/journal.pcbi.1014364 [code]
- A comparative study of simulation-based inference methods for epidemic models with identifiability considerations.Journal: PLoS computational biologyIn common: PyTorch, pandas, SciPy, 2 other tools, computational, 3 references
- [8] doi:10.1093/bioinformatics/btag652 [code]
- mmVelo: a deep generative model for estimating cell state-dependent dynamics across multiple modalities.Journal: Bioinformatics (Oxford, England)In common: pysam, PyTorch Lightning, PyTorch, 5 other tools
- [9] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: Snakemake, pysam, seaborn, 4 other tools
- [10] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: Snakemake, pysam, seaborn, 4 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 62 scripts, and 21 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:e64a158ae0679a2e…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
