An integrative mendelian randomisation and drug mechanism framework for target prioritisation and therapeutic repurposing in major depression.
The 5 matches
- [1] § Materials and methods › Instrument selection ↔ gwas_norm/normalise.py, lines 503–962 · score 0.60 · chi square, minor allele frequencies, minimal, GWAS
- [2] § Materials and methods › Instrument selection ↔ merit/coloc/coloc_core.py, lines 1367–1445 · score 0.60 · minor allele frequencies, removing variants, squares, selection, trait, correlations
- [3] § Materials and methods › Sensitivity analyses and target annotation › Colocalisation ↔ merit/coloc/coloc_core.py, lines 334–437 · score 0.58 · posterior probability, PPH3, abf, PPH4, colocalisation, traits
- [4] § Materials and methods › Primary Mendelian randomisation (MR) analysis ↔ merit/mr/ivw.py, lines 21–164 · score 0.57 · inverse variance weighted, IVW, instrumented, correlated, MR, variant
- [5] § Materials and methods › Sensitivity analyses and target annotation › Estimated therapeutic relevance ↔ bio_misc/drug_lookups/therapeutic_direction.py, lines 102–230 · score 0.54 · therapeutic directional, ChEMBL, activator, inhibitor, unknown, risk
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 2,566 lines · 104 KB · GPL-3.0 · 2 matches
- """
- Python implementation of the coloc R package.
- See the original R code here (06-04-2021 last update):
- `Here <https://github.com/chr1swallace/coloc/blob/main/R/split.R>`__.
- Relevant documentation:
- `Here <https://chr1swallace.github.io/coloc/articles/a02_data.html>`__.
- And relevant papers:
- `Here <https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1004383>`__.
- `Here <https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1008720>`__.
- `Here <https://www.biorxiv.org/content/10.1101/2021.02.23.432421v1>`__.
- Note, the `Sum of Single Effects` SuSIE version is not yet implementied
- Attribution
- -----------
- All the credit goes towards the original authors.
- All mistakes are our own.
- Note:
- -----
- Some function were omitted, because there were not used in R coloc:
- -. est.cond.nometa
- -. bin2lin.nometa
- -. estgeno.1.ctl
- -. estgeno.1.cse
- """
- import numpy as np
- import pandas as pd
- import statsmodels.formula.api as smf
- import warnings
- import itertools
- import scipy.stats
- # from pandas.core.frame import DataFrame
- from collections import OrderedDict
- from merit.errors import(
- _process_corr_mat,
- are_columns_in_df,
- are_Series_equal,
- error_on_non_finite_pd_corr_matrix,
- is_type,
- )
- from merit import constants as c
- from typing import Any, List, Type, Union, Tuple
- SEP = ''
- """
- TODO:
- Nothing ...
- """
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- # Constants
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- class Coloc_Expected(object):
- """
- Collecting some expected constants, getting around all the `SEP.join` that
- would otherwise clutter the code.
- Might move this to merit.constants, depending on the development of
- the coloc implementation into separate scripts.
- Notice:
- -------
- In R coloc `standard_deviation`, `sample_size` and `event_rate` are floats,
- not columns. Here we have included them as columns and addapted to code to
- take the average.
- """
- # defining names
- exp_name = 'exposure'
- out_name = 'outcome'
- VARIANT_ID=c.UNIVERSAL_ID.name
- EXPOSURE_EFFECT_SIZE= SEP.join(
- [c.EFFECT_SIZE.name, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_EFFECT_SIZE= SEP.join(
- [c.EFFECT_SIZE.name, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # simply se^2
- EXPOSURE_VARBETA= SEP.join(
- [c.COLOC_VARBETA, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_VARBETA= SEP.join(
- [c.COLOC_VARBETA, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # standard deviation of the gwas trait
- EXPOSURE_STANDARD_DEVIATION = SEP.join(
- [c.COLOC_STANDARD_DEVIATION, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_STANDARD_DEVIATION = SEP.join(
- [c.COLOC_STANDARD_DEVIATION, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # GWAS sample size
- EXPOSURE_SAMPLE_SIZE = SEP.join(
- [c.COLOC_SAMPLE_SIZE, c.MERGE_SUFFIX_DELIMITER_UP, exp_name]
- )
- OUTCOME_SAMPLE_SIZE = SEP.join(
- [c.COLOC_SAMPLE_SIZE, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # PROPORTION OF CASES
- EXPOSURE_EVENT_RATE = SEP.join(
- [c.COLOC_EVENT_RATE, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_EVENT_RATE = SEP.join(
- [c.COLOC_EVENT_RATE, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # z-statistics
- EXPOSURE_Z_STATISTICS = SEP.join(
- [c.COLOC_NORMAL_DEVIATE, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_Z_STATISTICS = SEP.join(
- [c.COLOC_NORMAL_DEVIATE, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # Shrinkage factor: ratio of the prior variance to the total variance
- EXPOSURE_SHRINKAGE_FACTOR = SEP.join(
- [c.COLOC_SHRINKAGE_FACTOR, c.MERGE_SUFFIX_DELIMITER_UP, exp_name])
- OUTCOME_SHRINKAGE_FACTOR = SEP.join(
- [c.COLOC_SHRINKAGE_FACTOR, c.MERGE_SUFFIX_DELIMITER_UP, out_name])
- # The log (approximate) Bayes Factor for colocalization
- EXPOSURE_ABF = SEP.join(
- [c.COLOC_LOG_APPROX_BAYES_FACTOR, c.MERGE_SUFFIX_DELIMITER_UP,
- exp_name])
- OUTCOME_ABF = SEP.join(
- [c.COLOC_LOG_APPROX_BAYES_FACTOR, c.MERGE_SUFFIX_DELIMITER_UP,
- out_name])
- # minor allele frequence
- # NOTE does this need to be exp/outcome stratified?
- MINOR_ALLELE_FREQ = c.MINOR_ALLELE_FREQ.name
- # types of gwas traits
- TRAIT_TYPES = ['quant', 'bin']
- # methods
- METHODS=["single", "cond", "mask"]
- # modes
- MODES=["iterative", "allbutone"]
- # colnames for dataset 1
- exposure_columns = [
- VARIANT_ID,
- EXPOSURE_EFFECT_SIZE,
- EXPOSURE_VARBETA,
- EXPOSURE_STANDARD_DEVIATION,
- EXPOSURE_SAMPLE_SIZE,
- EXPOSURE_EVENT_RATE,
- MINOR_ALLELE_FREQ
- ]
- # names for dataset 2
- outcome_columns = [
- VARIANT_ID,
- OUTCOME_EFFECT_SIZE,
- OUTCOME_VARBETA,
- OUTCOME_STANDARD_DEVIATION,
- OUTCOME_SAMPLE_SIZE,
- OUTCOME_EVENT_RATE,
- MINOR_ALLELE_FREQ
- ]
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- # Example datasets
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def _coloc_example_data():
- """
- Returns a coloc example dataset, and corr_matrix, for five variants
- """
- nvariants = 5
- idx = ['snp_' + str(l) for l in range(1, 6)]
- # dataset
- dataset = pd.DataFrame(
- {
- Coloc_Expected.VARIANT_ID: idx,
- Coloc_Expected.EXPOSURE_EFFECT_SIZE : [ 1.14517929e+00,
- 7.90320821e-01,
- -1.33679243e-02,
- 1.78273663e-01,
- -3.77606312e-01],
- Coloc_Expected.EXPOSURE_VARBETA : [0.03924328, 0.03880072,
- 0.04455547, 0.04602953,
- 0.04237982],
- Coloc_Expected.EXPOSURE_SAMPLE_SIZE : [1000] * nvariants,
- Coloc_Expected.EXPOSURE_STANDARD_DEVIATION : [2.32924174]*nvariants,
- Coloc_Expected.EXPOSURE_EVENT_RATE : [0.10] * nvariants,
- Coloc_Expected.OUTCOME_EFFECT_SIZE : [ 1.07128929, 0.92979073,
- -0.14149679, -0.02395321,
- -0.03075124],
- Coloc_Expected.OUTCOME_VARBETA : [0.04028734, 0.04478292,
- 0.04728015, 0.04748579,
- 0.04415969],
- Coloc_Expected.OUTCOME_SAMPLE_SIZE : [1000] * nvariants,
- Coloc_Expected.OUTCOME_STANDARD_DEVIATION : [2.33203611]*nvariants,
- Coloc_Expected.OUTCOME_EVENT_RATE : [0.10, 0.11, 0.09, 0.12, 0.08],
- Coloc_Expected.MINOR_ALLELE_FREQ : [0.1349454, 0.11783439,
- 0.13162119, 0.12019231, 0.125]
- }, index = idx
- )
- dataset.index_name=Coloc_Expected.VARIANT_ID
- # correlation matrix
- corr_matrix = pd.DataFrame(
- {
- idx[0] : [1.000000, 0.738238, 0.634454, 0.555827, 0.478299],
- idx[1] : [0.738238, 1.000000, 0.730387, 0.636565, 0.550489],
- idx[2] : [0.634454, 0.730387, 1.000000, 0.735947, 0.618572],
- idx[3] : [0.555827, 0.636565, 0.735947, 1.000000, 0.746561],
- idx[4] : [0.478299, 0.550489, 0.618572, 0.746561, 1.000000]
- }, index = idx)
- # return
- return dataset, corr_matrix
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # empty results dataframe
- def _empty_coloc_results():
- '''
- Returns an empty coloc results pd.DataFrame
- '''
- results = pd.DataFrame(columns = [c.COLOC_HIT1, c.COLOC_HIT2, c.MR_NSNPS,
- c.COLOC_PPH0, c.COLOC_PPH1, c.COLOC_PPH2,
- c.COLOC_PPH3, c.COLOC_PPH4, c.COLOC_BEST1,
- c.COLOC_BEST2, c.COLOC_BEST4,
- c.COLOC_ZSTAT_HIT1, c.COLOC_ZSTAT_HIT2
- ], index=[0])
- return results
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- # Helper functions
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def _add_missing_columns(check_data: pd.DataFrame, complete_data: pd.DataFrame,
- columns : list, index_col: Coloc_Expected.VARIANT_ID
- ):
- '''
- Add missing columns
- Arguments
- ---------
- data : pd.DataFrame
- columns : list of strings
- index_col : str
- '''
- # Shared index
- check_data.set_index(check_data[index_col], drop=False, inplace=True)
- complete_data.set_index(complete_data[index_col], drop=False, inplace=True)
- check_data.index.name = None
- complete_data.index.name = None
- # mapp columns
- for col in columns:
- if not col in check_data.columns:
- check_data[col] = complete_data[col]
- # returns stuff
- return check_data
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def _logsum(x):
- """
- Calculate the sum of terms.
- Used by the approximate Bayes factor function to calculate the posterior
- probabilities. Depends on numpy (np).
- Arguments
- ---------
- x : pd.Series, numpy.array or some other kind of numeric data type
- """
- max_res = np.max(x)
- results = max_res + np.log(np.sum(np.exp(x - max_res)))
- return results
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def _logdiff(x, y):
- """
- Calculate the difference of between two vectors.
- Used by the approximate Bayes factor function to calculate the posterior
- probablities. Depends on numpy (np), used in combination with _logsum.
- In the case of coloc its output should be a constant.
- Arguments
- ---------
- x, y : pd.Series, numpy.array or some other kind of numeric data type
- Returns
- -------
- A pd.Series with the element wise difference of x-y after relevant
- transformation. Results will be NA if x-y < 0.
- """
- max_res = np.max(x + y)
- results = max_res + np.log(np.exp(x - max_res) - np.exp(y - max_res))
- return results
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- # Functions originally in coloc/split.R
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def sd_pheno_approx(variancebeta: pd.Series, maf: pd.Series,
- n: Union[pd.Series, float]):
- """
- Get an estimate of the standard deviation of the phenotype distribution.
- Note
- ----
- Implements the `sdY.est` function of R coloc.
- Arguments
- ---------
- variancebeta : pd.Series
- Squared standard errors of the relevant point estimates.
- maf : pd.Series
- The minor allele frequencies of the supplied variants.
- n : pd.Series
- The GWAS sample size.
- Returns
- -------
- np.float64, reflecting an estimate of the phenotypic standard deviation
- (of 'Y').
- """
- # @@@ check input
- is_type(variancebeta, pd.Series)
- is_type(maf, pd.Series)
- is_type(n,(pd.Series, float))
- if not variancebeta.shape[0] == maf.shape[0]:
- raise IndexError('Input does not have the same shape')
- # @@@ Do the actual calculations
- oneover = np.divide(1, variancebeta)
- # remove accidental inf
- oneover = oneover[~np.isinf(oneover)]
- if isinstance(n, pd.Series):
- n = n.mean()
- nvx = np.multiply(2 * n, maf * np.subtract(1, maf))
- # OLS to approximate the SD of the phenotype
- ols_res = smf.ols(
- formula="nvx ~ oneover -1",
- data=pd.DataFrame({"nvx": nvx, "oneover": oneover}),
- missing='drop'
- )
- cf = ols_res.fit().params["oneover"]
- if cf < 0:
- raise ValueError("The standard deviation is negative. \
- Please provide your own estimate of the phenotype SD")
- else:
- return np.sqrt(cf)
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def approx_bf(ln_abf1: pd.Series, ln_abf2: pd.Series,
- p1: float=10**-4,
- p2: float=10**-4,
- p12: float=5*10**-6,
- extra_info: bool=True,
- masking_ln_abf3: pd.Series=pd.Series(dtype=float),
- masking_ln_abf4: pd.Series=pd.Series(dtype=float)):
- """
- Calculate the posterior probabilities of colocalized signals.
- Depends on numpy (np). Its arguments include the natural logarithm of the
- base factor for GWAS data set 1 and GWAS data set 2. P1-2 are constants (
- or column vectors of length equal to the number of variants) representing
- the probability of a random variant being associated only with GWAS
- phenotype 1 or 2, but not both. P12 on the other hand is the prior
- probability of colocalization.
- Note
- ----
- Implements `.f`, from `coloc.process` from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L419>`__.
- Implements `combine.abf` coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/claudia.R#L115>`__.
- Arguments
- ---------
- ln_abf1, ln_abf2 : pd.Series
- The log approximate base factor for phenotypes 1 or 2.
- p1, p2, p12 : floats
- The prior probabilities that a random variant is associated with either
- trait 1 or 2, or both, respectively.
- extra_info : boolean, default True
- If we want addiontal detail on the best variants per hypothesis.
- Similar to the information provided by R coloc.detail and the `.f`
- function in R coloc.process.
- masking_ln_abf3, masking_ln_abf4 : pd.Series
- the log approximate Bayes factors when using `method=mask`
- Returns
- -------
- Returns a dictionary of posterior probabilities
- """
- # @@@ CHECKS
- # TODO
- # @@@ actual calculations
- lsum = ln_abf1 + ln_abf2
- # the log posterior probability
- lh0_abf = 0
- lh1_abf = np.add(np.log(p1), _logsum(ln_abf1))
- lh2_abf = np.add(np.log(p2), _logsum(ln_abf2))
- # NOTE the _logdiff here is specific to `combine.abf` and not repeated
- # below when extra_info == True
- lh3_abf = np.add(
- np.add(np.log(p1), np.log(p2)),
- _logdiff(_logsum(ln_abf1) + _logsum(ln_abf2), _logsum(lsum))
- )
- lh4_abf = np.add(np.log(p12), _logsum(lsum))
- # returning a summary
- temp_tuple = (lh0_abf, lh1_abf, lh2_abf, lh3_abf, lh4_abf)
- denom = _logsum(temp_tuple)
- # if extra_info is True
- if extra_info == True:
- best1 = ln_abf1.idxmax()
- best2 = ln_abf2.idxmax()
- best4 = lsum.idxmax()
- # NOTE _logsum(ln_abf1) + _logsum(ln_abf2)) == lbf3
- # NOTE this does NOT include an _logdiff following `coloc.process`
- # see https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L423
- lh3_abf = np.add(np.add(np.log(p1), np.log(p2)),
- _logsum(ln_abf1) + _logsum(ln_abf2))
- if masking_ln_abf3.size != 0:
- lh3_abf = np.add(np.add(np.log(p1), np.log(p2)),
- _logsum(masking_ln_abf3))
- if masking_ln_abf4.size != 0:
- lh4_abf = np.add(np.log(p12), _logsum(masking_ln_abf4))
- best4 = masking_ln_abf4.idxmax()
- temp_tuple = (lh0_abf, lh1_abf, lh2_abf, lh3_abf, lh4_abf)
- denom = _logsum(temp_tuple)
- pp_abf = np.exp(temp_tuple - denom)
- # NOTE best? are the variants that match to their respective hypotheses.
- results = {c.COLOC_PPH0: pp_abf[0],
- c.COLOC_PPH1: pp_abf[1],
- c.COLOC_PPH2: pp_abf[2],
- c.COLOC_PPH3: pp_abf[3],
- c.COLOC_PPH4: pp_abf[4],
- c.COLOC_BEST1: best1,
- c.COLOC_BEST2: best2,
- c.COLOC_BEST4: best4
- }
- else:
- # if extra_info is False
- posterior = np.exp(np.subtract(temp_tuple, denom))
- results = {c.COLOC_PPH0: posterior[0],
- c.COLOC_PPH1: posterior[1],
- c.COLOC_PPH2: posterior[2],
- c.COLOC_PPH3: posterior[3],
- c.COLOC_PPH4: posterior[4]
- }
- return results
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def vestgeno_1_ctl(f: pd.Series):
- '''
- Helper function to estimate the relative allele frequency of genotypes
- 0, 1, 2 in control subjects.
- Note
- ----
- Implements `vestgeno.1.ctl`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L543>`__.
- Arguments
- ---------
- f : pd.Series
- A vector of minor allele frequencies.
- Returns
- -------
- A pd.DF with the genotype frequencies.
- '''
- # @@@ check input
- is_type(f, pd.Series)
- if (0 > f.min()) or (f.max() > 0.50):
- raise ValueError('Supplied values should range between {} and {}'.
- format(0, 0.50))
- # @@@ calulcating the MAF in controls
- dt = pd.DataFrame(np.array([(1-f)**2,
- 2*f*(1-f),
- f**2 ])).T
- return dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def vestgeno_1_cse(G_0: pd.DataFrame, b: np.ndarray):
- '''
- Helper function to estimate the relative allele frequency of genotypes
- 0, 1, 2 in case subjects.
- Implements `vestgeno.1.cse`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L559>`__.
- Arguments
- ---------
- G_0 : pd.DF
- The control group genotype frequencies;
- Results from `vestgeno_1_ctl`. With:
- -. Column 0 recording the frequency of genotype 0,
- -. Column 1 the frequency of genotype 1,
- -. Column 2 the frequency of genotype 2.
- b : pd.Series
- A column vector of effect estimates.
- Returns
- -------
- A pd.DF
- '''
- # @@@ check input
- is_type(G_0, pd.DataFrame)
- is_type(b, pd.Series)
- b = b.to_numpy()
- # @@@ Do the actual calculations
- g_0 = 1
- g_1 = np.exp(b - np.log(G_0.iloc[:,0]/G_0.iloc[:,1]))
- g_2 = np.exp(2*b - np.log(G_0.iloc[:,0]/G_0.iloc[:,2]))
- dt = pd.DataFrame({"g_0" : g_0,
- "g_1" : g_1,
- "g_2" : g_2})
- dt1 = dt.div(g_0+g_1+g_2, axis='rows')
- return dt1
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def VMAF_cc(f_0: pd.Series, b: pd.Series, N_0: float, N_1: float) -> np.ndarray:
- '''
- Estimate the MAF for a case-control analysis, based on MAF from controls,
- the log OR (`b`) and the number of controls `N_0` and cases `N_1`.
- Implements `VMAF.cc`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L566>`__.
- Arguments
- ---------
- f_O : pd.Series
- The minor allele frequency estimated from control subjects.
- b : pd.Series
- A vector of log odds ratio.
- N_0 : float
- The number of control subjects.
- N_1 : float
- The number of subjects with an event (the cases).
- Returns
- -------
- An np.ndarray
- '''
- # @@@ check input
- is_type(f_0, pd.Series)
- is_type(b, pd.Series)
- is_type(N_0, float)
- is_type(N_1, float)
- # @@@ the calculations
- G_0 = vestgeno_1_ctl(f_0)*N_0
- G_1 = vestgeno_1_cse(G_0, b)*N_1
- # set same names
- G_0.columns = G_1.columns
- G = G_0.add(G_1)
- E_2 = np.divide((G @ np.array([0,1,4]).reshape(-1,1)), (N_0+N_1))
- E_1 = np.divide((G @ np.array([0,1,2]).reshape(-1,1)), (N_0+N_1))
- cc = np.divide((E_2 - E_1**2) * (N_0+N_1), (N_0+N_1-1))
- # to numpy
- # NOTE this may break stuff - was pd.DF before
- if cc.shape[1] == 1:
- cc = cc.iloc[:, 0].to_numpy()
- else:
- raise IndexError('`cc` has more than one column')
- return cc
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def map_cond(dt,
- corr_matrix,
- YY,
- sigsnps=[],
- r2_threshold = 0.80,
- effect_type='quant',
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- event_rate=c.COLOC_EVENT_RATE,
- sample_size=c.COLOC_SAMPLE_SIZE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ,
- label=""):
- """
- Internal helper function for finemap.signals. Finds the next most significant
- variant, after conditioning on content of `sigsnps`.
- Implements `map_cond`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L256>`__.
- Arguments
- ---------
- dt : pd.DF,
- A variant dataframe with columns ...
- corr_matrix : pd.DF,
- The correlation matrix of the varaints in `dt`.
- YY : float,
- The residual sum of squares.
- sigsnps : list,
- A list of row indices of the varaint(s) we want to condition on.
- effect_size : str,
- The name of the effect size column.
- r2_threshold : float, default 0.80,
- Pruning variants with a squared pairwise correlation above
- `r2_threshold`. This is done to improve model stability, nothing else.
- effect_type : str, default `quant`
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- varbeta : str,
- The name of the variance of the effect size column
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent). Only relevant for quantitative trait GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- The name of the effect size column.
- label : str
- ...
- Returns
- -------
- A pd.DF with the conditionally significant variants.
- """
- # @@@ check input
- is_type(dt, pd.DataFrame)
- core_cols = [variant_id, effect_size, varbeta,
- sample_size, minor_allele_freq]
- are_columns_in_df(dt, core_cols)
- is_type(effect_type, str)
- # check effect_type
- if not effect_type in Coloc_Expected.TRAIT_TYPES:
- raise AttributeError("`effect_type` should be either {} or {}.".\
- format(*Coloc_Expected.TRAIT_TYPES))
- # Check if we have a correlation matrix
- if corr_matrix is None:
- raise AttributeError(" The `corr_matrix` is missing")
- if corr_matrix.empty == True:
- raise AttributeError(" The `corr_matrix` is missing")
- is_type(corr_matrix, pd.DataFrame)
- # check if corr_matrix and dt have the same entries
- # and align dt and corr_matrix entries
- _process_corr_mat(data=dt, corr_matrix=corr_matrix,
- index_col=Coloc_Expected.VARIANT_ID
- )
- # @@@ The calculations
- # if there are no pre-defined varaints
- # take max unconditional abs(z)
- if len(sigsnps) == 0:
- Z = np.divide(dt[effect_size], np.sqrt(dt[varbeta])).to_numpy()
- wh = np.argmax(np.abs(Z))
- sig_dt = pd.DataFrame([Z[wh]], columns=[dt[variant_id].iloc[wh]])
- elif len(sigsnps) > 0:
- # place holder
- if effect_size == 'quant':
- event_rate = c.COLOC_EVENT_RATE
- est = est_cond(dt=dt, corr_matrix=corr_matrix, YY=YY,
- sigsnps=sigsnps, effect_type=effect_type,
- r2_threshold=r2_threshold,
- variant_id=variant_id,
- effect_size=effect_size,
- varbeta=varbeta,
- event_rate=event_rate,
- sample_size=sample_size,
- minor_allele_freq=minor_allele_freq,
- label=label
- )
- Z = np.divide(est[c.EFFECT_SIZE.name+"_"+label],
- np.sqrt(est[c.COLOC_VARBETA+"_"+label]))
- wh = np.argmax(np.abs(Z))
- # NOTE the iloc is a dirty fix - should go back to this and actually
- # address the index
- sig_dt = pd.DataFrame([Z.iloc[wh]],
- columns=[est.iloc[wh][c.UNIVERSAL_ID.name]])
- # return
- return sig_dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def est_cond(dt,
- corr_matrix,
- YY,
- sigsnps,
- effect_type="quant",
- xtx=None,
- r2_threshold = 0.80,
- effect_size=c.EFFECT_SIZE.name,
- variant_id=Coloc_Expected.VARIANT_ID,
- varbeta=c.COLOC_VARBETA,
- event_rate=c.COLOC_EVENT_RATE,
- sample_size=c.COLOC_SAMPLE_SIZE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ,
- label=""):
- """
- Internal helper function for est_all_cond. Estimates conditional GWAS summary
- statistics.
- Implements `est_cond`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L115>`__.
- Arguments
- ---------
- dt : pd.DF,
- A variant dataframe with columns ...
- corr_matrix : pd.DF,
- The correlation matrix of the variants in `dt`.
- YY : float,
- The residual sum of squares.
- sigsnps : list,
- A list of row indices of the varaint(s) we want to condition on.
- effect_type : str, default `quant`
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- xtx : pd.DF?, default None,
- The "Gram matrix" matrix -- essentially the variance-covariance matrix
- obtained by X.transpose(). Will be internally calculated when set to
- `None`, based on the corr_matrix, MAF, beta and sample size (the last
- three are supplied through `dt`).
- r2_threshold : float, default 0.90,
- Pruning variants with a squared pairwise correlation above
- `r2_threshold`.
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent). Only relevant for quantitative trait GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- label : str
- ...
- Returns
- -------
- pd.DF with conditional GWAS estimates
- """
- # @@@ check input
- is_type(dt, pd.DataFrame)
- core_cols = [variant_id, effect_size, varbeta,
- sample_size, minor_allele_freq]
- are_columns_in_df(dt, core_cols)
- # Check if we have a correlation matrix
- if corr_matrix is None:
- raise AttributeError(" The `corr_matrix` is missing")
- if corr_matrix.empty == True:
- raise AttributeError(" The `corr_matrix` is missing")
- is_type(corr_matrix, pd.DataFrame)
- # check if corr_matrix and dt have the same entries
- # and align dt and corr_matrix entries
- _process_corr_mat(data=dt, corr_matrix=corr_matrix,
- index_col=Coloc_Expected.VARIANT_ID
- )
- # Guard the matrix-inverting path (np.linalg.solve, below) against
- # undefined-LD NaN rows and non-positive-definite matrices. This runs
- # AFTER _process_corr_mat has aligned dt/corr_matrix and BEFORE the XX /
- # solve computation. It is matrix-only (no instrument matching), so it is
- # safe for coloc's dt whose variant IDs live in a column with a RangeIndex.
- error_on_non_finite_pd_corr_matrix(corr_matrix)
- # further input
- is_type(YY, float)
- is_type(sigsnps, list)
- if not effect_type in Coloc_Expected.TRAIT_TYPES:
- raise AttributeError("`effect_type` should be either {} or {}.".\
- format(*Coloc_Expected.TRAIT_TYPES))
- # @@@ The actual calculations
- # Extracting dt.variants we want to condition one
- nuse = []
- for snp in sigsnps:
- nuse.append(dt[variant_id].tolist().index(snp))
- # to find new beta for, conditional on nuse, excluding variants in r-squared
- # > r2_threshold. With nuse because small inaccuracies can blow up, and
- # we generally don't expect a second detectable signal in
- # such high LD with a first signal
- ld_mask = np.max(np.power(corr_matrix.iloc[nuse], 2), axis=0) > r2_threshold
- nuse_ld = np.flatnonzero(np.asarray(ld_mask)).tolist()
- # adding high_ld variants
- nuse_comb = list(set(nuse + nuse_ld))
- nuse_comb_nam = dt[variant_id].iloc[nuse_comb].to_list()
- # position
- use_names = [x for x in dt[variant_id].tolist() if x not in nuse_comb_nam]
- # use_names = [x for x in dt[variant_id].tolist() if x not in sigsnps]
- use = []
- # get position
- for snp in use_names:
- use.append(dt[variant_id].tolist().index(snp))
- # getting the variance-covariance matrix
- if xtx is None:
- # depending on the trait, do
- if effect_type == "quant":
- # if quant, assume MAF is correct estimate for sample
- # expected variance of X, h_buf in GCTA, Dj/N in Yang et al
- VX = 2 * dt[minor_allele_freq] * (1 - dt[minor_allele_freq])
- A1 = np.array(VX).reshape(1,-1)
- A2 = np.array(VX).reshape(-1,1)
- # E(XtX/Ni)
- P = np.sqrt(A2 @ A1)
- elif effect_type == "bin":
- # Correcting the MAF
- # NOTE Taking the average event_rate and sample_size
- VW = VMAF_cc(f_0 = dt[minor_allele_freq],
- b = dt[effect_size],
- N_0 = (1-dt[event_rate].mean()) * dt[sample_size].mean(),
- N_1 = dt[event_rate].mean() * dt[sample_size].mean(),
- )
- A1 = np.array(VW).reshape(-1, 1)
- A2 = np.array(VW).reshape(1,-1)
- # np.sqrt(A1 @ A2)
- P = np.sqrt(A1 @ A2)
- else:
- raise AttributeError("effect_type does not exist, either supply \
- effect type as `quant` or `bin`, or supply a variance-covariance matrix through \
- xtx")
- else:
- XX = xtx
- # calculate XX
- XX = dt[sample_size].mean() * corr_matrix * P
- # The variance vectors
- D = np.diag(XX)
- D_1 = D[nuse]
- D_2 = D[use]
- # the covariance matrices (_1, _2, will be vectors)
- XX_1 = XX.iloc[nuse, nuse]
- XX_2 = XX.iloc[use, use]
- XX_12 = XX.iloc[nuse, use]
- XX_21 = XX.iloc[use, nuse]
- # accommodate nuse/use of length 1 or mor
- D_1 = np.array(D_1).reshape(1,1) if len(nuse) == 1 else np.diag(D_1)
- D_2 = np.array(D_2).reshape(1,1) if sum(use) == 1 else np.diag(D_2)
- # joint effects at nuse, if needed
- # (X'X)^-1 @ var @ beta
- b_1 = np.linalg.solve(XX_1, np.identity(len(XX_1))) @ D_1 @\
- dt[effect_size].iloc[nuse].to_numpy().reshape(-1,1) # single column vector
- ## conditional effects 2 | 1
- b_2 = np.subtract(
- # term 1
- dt[effect_size].iloc[use].to_numpy(),
- # subtract this from term 1
- np.divide(
- # numerator
- XX_21.to_numpy() @ np.linalg.solve(XX_1, np.identity(len(XX_1))) @ D_1 @\
- np.array(dt[effect_size].iloc[nuse]),
- # denominator
- np.diag(D_2)
- )
- ).reshape(1,-1)
- ## residual and conditional b2 variance
- # row vector
- # b_1 = b_1.reshape(1,-1)
- # column vector
- beta_mat = dt[effect_size].iloc[nuse].to_numpy().reshape(-1,1)
- # Sc
- # NOTE taking the mean -- this will always be a float
- Sc = YY - b_1.T @ D_1 @ beta_mat
- Sc = (
- # numerator
- (Sc - b_2 * np.diag(D_2) *
- dt[effect_size].iloc[use].to_numpy().reshape(1,-1)) /
- # denominator
- (dt[sample_size].mean() - len(sigsnps) - 1)
- )
- # variance b2
- vb_2 = np.divide(
- # numerator
- np.array(Sc) * (
- np.diag(D_2) -\
- np.diag(XX_21 @ np.linalg.solve(XX_1, np.identity(len(XX_1))) @\
- np.array(XX_12)
- )
- ),
- # denominator
- np.diag(D_2)**2)
- # abs here because very occasionally can get a small negative vb2 due to
- # approximation. In this case, replace by a positive vb2 of same small
- # magnitude
- vb_2 = np.abs(vb_2)
- # results table
- cond_dt = pd.concat(
- [
- # table 1
- pd.DataFrame({
- c.UNIVERSAL_ID.name: dt[variant_id].iloc[use].to_list(),
- c.EFFECT_SIZE.name+"_"+label: list(itertools.chain.from_iterable(b_2)),
- c.COLOC_VARBETA+"_"+label: list(itertools.chain.from_iterable(vb_2))
- }) ,
- # table 2
- # NOTE USING nuse_comb, here to ensure multicolinear variants
- # get returned as well
- pd.DataFrame({
- c.UNIVERSAL_ID.name: dt[variant_id].iloc[nuse_comb].to_list(),
- c.EFFECT_SIZE.name+"_"+label: 0*len(nuse_comb),
- c.COLOC_VARBETA+"_"+label: dt[varbeta].iloc[nuse_comb].to_list()
- })
- ]
- )
- # reset index
- cond_dt.reset_index(drop=True, inplace=True)
- # returns stuff
- return cond_dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # Done
- def find_best_signal(dt: pd.DataFrame,
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- ):
- """
- Internal helper function to select the most significant variant.
- Implements `find.best.signal`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L285>`__.
- Arguments
- ---------
- dt : pd.DF,
- A variant dataframe with columns ...
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column
- Returns
- -------
- pd.DF with most significant variant
- """
- # @@@ Cehck input
- is_type(dt, pd.DataFrame)
- are_columns_in_df(dt, [effect_size, varbeta])
- # @@@ caclculate z-statistics
- # NOTE HAVE removed the YY part, as far as I can tell it does nothing.
- Z = dt[effect_size]/dt[varbeta].pow(1/2)
- # largest Z
- wh = np.argmax(np.abs(np.array(Z)))
- # return
- # NOTE perhaps this works better as a dictionary?
- # NOTE 2 - what if there are ties?
- Zs = Z.iloc[wh]
- nam = [dt[variant_id].iloc[wh]]
- return_dt = pd.DataFrame([Zs], columns=nam)
- return return_dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def est_all_cond(dt: pd.DataFrame,
- corr_matrix: Union[pd.DataFrame, type(None)],
- fm: pd.DataFrame,
- mode="iterative",
- effect_type="quant",
- r2_threshold = 0.80,
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- event_rate=c.COLOC_EVENT_RATE,
- sample_size=c.COLOC_SAMPLE_SIZE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ,
- standard_deviation=c.COLOC_STANDARD_DEVIATION,
- label="") -> dict:
- """
- Implements `est_all_cond`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L671>`__.
- Arguments
- ---------
- dt : pd.DF,
- A variant dataframe with columns ...
- corr_matrix : pd.DF,
- The correlation matrix of the variants in `dt`.
- fm : pd.DF,
- The best variants selected through `find_best_signal`.
- mode : float, default `iterative`,
- `iterative` or `allbutone`; to either successively condition on signals
- find all putative signals and condition on all but one of them in each
- analysis.
- effect_type : str, default `quant`
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- r2_threshold : float, default 0.80,
- Pruning variants with a squared pairwise correlation above
- `r2_threshold`.
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column
- event_rate : str,
- The column name of the proportion of cases in the source GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent). Only relevant for quantitative trait GWAS.
- label : str
- ...
- Returns
- -------
- A dictionary with one unconditional and x many conditional GWAS estimates
- """
- # @@@ Check input
- # Check if core data is present
- is_type(dt, pd.DataFrame)
- is_type(fm, pd.DataFrame)
- is_type(corr_matrix, (pd.DataFrame, type(None)))
- is_type(effect_type, str)
- is_type(mode, str)
- is_type(r2_threshold, float)
- # Check if effect types are correct
- if not effect_type in Coloc_Expected.TRAIT_TYPES:
- raise AttributeError("`effect_type` should be either {} or {}.".\
- format(*Coloc_Expected.TRAIT_TYPES))
- # check modes
- if not mode in Coloc_Expected.MODES:
- raise AttributeError(
- "`modes` should be a str: {}".format(Coloc_Expected.MODES)
- )
- # Checking columns content
- cor_cols = [variant_id, effect_size, varbeta]
- are_columns_in_df(dt, cor_cols)
- # @@@ the actual calculations
- if effect_type == 'bin':
- dt, _ = bin2lin(dt,
- variant_id=variant_id,
- effect_size=effect_size,
- varbeta=varbeta,
- event_rate=event_rate,
- sample_size=sample_size,
- minor_allele_freq=minor_allele_freq
- )
- # get z-statistics
- dt[c.COLOC_NORMAL_DEVIATE + label] = dt[effect_size]/dt[varbeta].pow(1/2)
- # calculate YY
- YY = est_YY(dt=dt, effect_type=effect_type,
- standard_deviation=standard_deviation,
- sample_size=sample_size,
- event_rate=event_rate)
- # replaced the `sigs` function with the following
- snps = fm.columns.tolist()
- # snps = ['snp_1', 'snp_2', 'snp_3']
- if mode == "allbutone":
- sigs = list(itertools.combinations(snps, len(snps)-1))
- sigs.reverse()
- elif mode == "iterative":
- sigs = [tuple(snps[:i]) for i in range(len(snps))]
- # getting the conditional dataframe
- # NOTE USING ORDERED_DICT
- cond_dt = OrderedDict()
- for sigsnps in sigs:
- # This will be True if sigsnps are not in sigs
- if not sigsnps:
- # making a tuple with the same structure as sigsnps
- temp = dt.copy()
- # making sure the index is correct
- temp.set_index(temp[variant_id], drop=False, inplace=True)
- temp.index.name = None
- cond_dt[(None, None)] = temp
- del temp
- else:
- # NOTE mapping sigsnps to a list in `est_cond`
- temp = est_cond(dt=dt, corr_matrix=corr_matrix, YY=YY,
- sigsnps=list(sigsnps),
- effect_type=effect_type,
- r2_threshold=r2_threshold,
- variant_id=variant_id,
- effect_size=effect_size,
- varbeta=varbeta,
- event_rate=event_rate,
- sample_size=sample_size,
- minor_allele_freq=minor_allele_freq,
- label=label
- )
- # ensuring the index is correct
- temp.set_index(temp[variant_id], drop=False, inplace=True)
- temp.index.name = None
- cond_dt[sigsnps] = temp
- del temp
- # return
- return cond_dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def map_mask(dt: pd.DataFrame,
- corr_matrix: pd.DataFrame,
- r2thr=0.01,
- sigsnps=[],
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- zstatistic=c.COLOC_NORMAL_DEVIATE
- ):
- """
- Internal helper function for finemap.signals. Finds the next most
- significant variants, masking/removing variants in `sigsnps`.
- Implements `map_mask`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L72>`__.
- Arguments
- ---------
- dt : pd.DF,
- A coloc dataframe with either the exposure or outcome data.
- corr_matrix : pd.DF,
- The correlation matrix of the variants in `dt`.
- r2thr : float, default 0.01
- Mask variants with a squared correlation with any of the `sigsnps`
- larger than `r2thr`.
- sigsnps : list,
- A list of row indices of the varaint(s) we want to condition on.
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column.
- zstatistic : str,
- The name of the z-statistic column. Will be estimates if effect_size
- and varbeta are supplied and the column is not present in `dt`.
- Returns
- -------
- A pd.DF with the significant variants
- """
- # @@@ check input
- is_type(dt, pd.DataFrame)
- core_cols = [variant_id, varbeta ]
- are_columns_in_df(dt, core_cols)
- # Check if we have a correlation matrix
- if corr_matrix is None:
- raise AttributeError(" The `corr_matrix` is missing")
- if corr_matrix.empty == True:
- raise AttributeError(" The `corr_matrix` is missing")
- is_type(corr_matrix, pd.DataFrame)
- # check if corr_matrix and dt have the same entries
- are_Series_equal(dt[variant_id], corr_matrix.index.to_series(),
- objects_names=['dt', 'corr_matrix']
- )
- # Files have the same entries, now make sure they are in the same order
- dt.sort_values(by=[variant_id], ascending=True, inplace=True)
- corr_matrix.sort_index(axis=0, ascending=True, inplace=True)
- corr_matrix.sort_index(axis=1, ascending=True, inplace=True)
- # @@@ actual calculations
- # calculate z-statistic
- if not zstatistic in dt.columns:
- Z = np.divide(dt[effect_size], np.sqrt(dt[varbeta])).to_numpy()
- else:
- Z = dt[zstatistic].to_numpy()
- uni_id = np.array(dt[variant_id])
- use = np.repeat(True, dt.shape[0])
- # if there are sigsnps, do
- if len(sigsnps) > 0:
- expectedz = np.repeat(0, dt.shape[0])
- # select just the column for the sigsnps
- a = np.array(corr_matrix[sigsnps]).T
- friends = np.any(np.abs(a) > np.sqrt(r2thr), axis=0)
- use = np.logical_xor(use, friends)
- # NOTE check if this needs to be an empty df with the same rows and columns
- # It seems to only fill the column name by uni_id, so presumably you
- # can keep it empty.
- if np.any(use) == False:
- sig_dt = pd.DataFrame()
- else:
- # otherwise go on
- idx = np.flatnonzero(
- dt[variant_id].isin(sigsnps).to_numpy()
- )
- expectedz = np.array(corr_matrix[sigsnps]) @ Z[idx]
- zdiff = np.abs(Z.T[use]) - np.abs(expectedz.T[use])
- wh = np.argmax(zdiff)
- sig_dt = pd.DataFrame([Z.T[use][wh]],
- columns=[uni_id[use][wh]])
- elif len(sigsnps) == 0:
- zdiff = np.abs(Z.T)
- wh = np.argmax(zdiff)
- sig_dt = pd.DataFrame([Z[wh]],
- columns=[dt[variant_id].iloc[wh]])
- # return
- return sig_dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def wmean(estimates: Union[pd.DataFrame, pd.Series, np.ndarray],
- weights: Union[pd.DataFrame, pd.Series, np.ndarray, float]
- ) -> np.ndarray:
- '''
- Weighted mean -- fixed effect meta-analysis.
- Arguments
- ---------
- estimates, weights : pd.DataFrame or pd.Series
- Returns
- -------
- A np.array
- '''
- # @@@ check
- is_type(estimates, (pd.DataFrame, pd.Series, np.ndarray))
- is_type(weights, (pd.DataFrame, pd.Series, np.ndarray, float))
- # @@@ estimates
- # remove indices if needed
- if isinstance(estimates, ( pd.DataFrame, pd.Series )):
- estimates = estimates.to_numpy()
- if isinstance(weights, ( pd.DataFrame, pd.Series )):
- weights = weights.to_numpy()
- # make 2d
- if len(weights.shape) == 1:
- weights = weights.reshape(-1,1)
- try:
- # colsums
- res = np.sum(estimates * weights, axis=0)/np.sum(weights, axis=0)
- except ValueError:
- # print('testing here')
- res = np.sum(estimates * weights, axis=0)/np.sum(weights)
- # return
- return res
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def bin2lin(dt,
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- event_rate=c.COLOC_EVENT_RATE,
- sample_size=c.COLOC_SAMPLE_SIZE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ):
- """
- Estimate beta and varbeta if a linear regression had been run on a binary
- outcome, given log odds ratio and their variance + MAF in controls
- Implements `bin2lin` and `bin2lin.nometa`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L633>`__.
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L574>`__.
- Arguments
- ---------
- dt : pd.DF,
- A variant dataframe with columns ...
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column
- event_rate : str,
- The column name of the proportion of cases in the source GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- Returns
- Unpacks an updated `dt` and `linear_approx`. Here `dt` replaces the orignal
- logOR and logOR variance with the estimated mean difference and its variance.
- The `linear_approx` compares the z-statistics from both the mean difference
- and the logOR, should be close to 1.
- """
- # @@@ check input
- is_type(dt, pd.DataFrame)
- core_cols = [variant_id, effect_size, varbeta,
- minor_allele_freq, event_rate, sample_size]
- are_columns_in_df(dt, core_cols)
- # @@@ estimate the genotype frequencies
- # control frequency
- g0 = vestgeno_1_ctl(dt[minor_allele_freq])
- # case frequency
- g1 = vestgeno_1_cse(g0, dt[effect_size])
- # expected genotypes
- geno_mat = np.array([[0,1,2]])
- # expected proportion of case and control subjects
- ex1 = geno_mat @ g1.transpose()
- ex0 = geno_mat @ g0.transpose()
- # squared expression
- vx1 = geno_mat**2 @ g1.transpose()
- vx0 = geno_mat**2 @ g0.transpose()
- # combined
- ex = (np.array([1 - dt[event_rate].mean()]).T @ ex0 +
- np.array([dt[event_rate].mean()]).T @ ex1)
- vx = (np.array([1 - dt[event_rate].mean()]).T @ vx0 +
- np.array([dt[event_rate].mean()]).T @ vx1 - ex**2)
- # binomial variance
- vy = np.array([dt[event_rate].mean() * (1 - dt[event_rate].mean())])
- # centred
- vxy = vy.transpose() @ ( ex1 - ex0 )
- # @@@ estimating varbeta and beta on the linear scale
- # essentially getting mean differences
- varbeta_mat = (vy/vx - vxy**2/vx**2) / (dt[sample_size].mean() - 1)
- # row vector
- # varbeta_mat = varbeta_mat.to_numpy().reshape(-1, 1)
- # storing original values
- beta_org = dt[effect_size].copy()
- varbeta_org = dt[varbeta].copy()
- if len(vxy.shape) == 1:
- dt[varbeta] = varbeta_mat.to_list()
- dt[effect_size] = (vxy/vx).to_list()
- else:
- warnings.warn("Performing meta-analysis")
- # meta-analysis, axis=1 is colsum
- dt[varbeta] = 1/np.sum(1/varbeta_mat, axis=1).tolist()
- dt[effect_size] = wmean(vxy/vx, 1/varbeta_mat).tolist()
- # @@@ diagnostic results
- zstat = dt[effect_size]/np.sqrt(dt[varbeta])
- zstat_org = beta_org/np.sqrt(varbeta_org)
- fit = smf.ols(formula='zstat ~ zstat_org - 1',
- data = pd.DataFrame(
- {'zstat': zstat, 'zstat_org': zstat_org}
- ), missing = 'drop').fit()
- linear_approx = fit.params['zstat_org']
- if not (0.90 <= linear_approx <= 1.1):
- warnings.warn("Poor linear approximation: {}".format(linear_approx))
- # returns
- return dt, linear_approx
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def est_YY(dt : pd.DataFrame,
- effect_type: list,
- variant_id=Coloc_Expected.VARIANT_ID,
- standard_deviation=c.COLOC_STANDARD_DEVIATION,
- sample_size=c.COLOC_SAMPLE_SIZE,
- event_rate=c.COLOC_EVENT_RATE
- ) -> float:
- """
- Calculates the squared residuals.
- Implements code chunck from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L293>`__.
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L355>`__.
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L680>`__.
- Arguments
- ---------
- dt : pd.DF,
- A coloc dataframe.
- effect_type : str, default `quant`,
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent). Only relevant for quantitative trait GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- event_rate : str,
- The column name of the proportion of cases in the source GWAS. Only
- applicable for a binary trait GWAS.
- Returns
- -------
- A float
- """
- # @@@ Check if core data is present
- is_type(dt, pd.DataFrame)
- core_cols = [variant_id, sample_size ]
- are_columns_in_df(dt, core_cols)
- if not effect_type in Coloc_Expected.TRAIT_TYPES:
- raise AttributeError("`effect_type` should be either {} or {}.".\
- format(*Coloc_Expected.TRAIT_TYPES))
- # @@@ performing the actual calulcations
- if effect_type == "quant":
- # check the correct columns are present
- are_columns_in_df(dt, [sample_size, standard_deviation])
- # do the actual calculations
- # NOTE taking the mean
- YY = dt[sample_size].mean() * dt[standard_deviation].mean()**2
- elif effect_type == "bin":
- # check input
- are_columns_in_df(dt, [sample_size, event_rate])
- # NOTE taking the mean here, reflecting R coloc
- # NOTE 2, in R `find.best.signal` divides YY by N, however
- # the value is never returned - dead code, not used here.
- YY = (dt[sample_size].mean() * dt[event_rate].mean() *
- (1 - dt[event_rate].mean()))
- # return
- return YY
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def finemap_signals(dt: pd.DataFrame,
- corr_matrix=None,
- method="single",
- r2thr=0.01,
- sigsnps=[],
- pthr=1e-6,
- maxhits=3,
- effect_type="quant",
- variant_id=Coloc_Expected.VARIANT_ID,
- effect_size=c.EFFECT_SIZE.name,
- varbeta=c.COLOC_VARBETA,
- zstatistic=c.COLOC_NORMAL_DEVIATE,
- standard_deviation=c.COLOC_STANDARD_DEVIATION,
- sample_size=c.COLOC_SAMPLE_SIZE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ,
- event_rate=c.COLOC_EVENT_RATE,
- verbose=True
- ):
- """
- This is an analogue to finemap.abf, adapted to find multiple signals where
- they exist, via conditioning or masking - ie a stepwise procedure
- Finemap multiple signals in a single dataset
- Internal helper function for finemap.signals. Finds the next most
- significant variants, masking/removing variants in `sigsnps`.
- Implements `finemap.signals`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L328>`__.
- Arguments
- ---------
- dt : pd.DF,
- A coloc dataframe.
- corr_matrix : pd.DF,
- The correlation matrix of the variants in `dt`. Only relevant when
- method == 'cond'.
- r2thr : float, default 0.01
- Mask variants with a squared correlation with any of the `sigsnps`
- larger than `r2thr`.
- method : str, default `single`,
- Use `single` to perform coloc assuming there is a single causal variant,
- Use `cond` to perform a stepwise conditional analysis,
- Use `mask` to perform the stepwise selection algorithm using masking
- instead.
- sigsnps : list,
- A list of row indices of the varaint(s) we want to condition on.
- NOTE:-> this does not seem to do anything in the R-version.
- pthr : float, default 1 \times 10^{-6},
- Stops when p-values are larger than `pthr`, where input should be
- defined between 0 and 1.
- maxhits : int, default 3,
- Stops when the procedures finds `maxhits`.
- effect_type : str, default `quant`,
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- effect_size : str,
- The name of the effect size column.
- varbeta : str,
- The name of the variance of the effect size column
- zstatistic : str,
- The name of the z-statistic column. Will be estimates if effect_size
- and varbeta are supplied and the column is not present in `dt`.
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent). Only relevant for quantitative trait GWAS.
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- event_rate : str,
- The column name of the proportion of cases in the source GWAS. Only
- applicable for a binary trait GWAS.
- Returns
- -------
- a pd.df with the remaining significant variants after masking.
- """
- # @@@ Check if core data is present
- is_type(dt, pd.DataFrame)
- is_type(corr_matrix, (pd.DataFrame, type(None)))
- core_cols = [variant_id, sample_size ]
- are_columns_in_df(dt, core_cols)
- # Checking methods
- if not isinstance(method, str):
- raise AttributeError( "The method type should be a string")
- if not method in Coloc_Expected.METHODS:
- raise AttributeError(
- "`method` should be a string. Pick from: {}".\
- format(Coloc_Expected.METHODS)
- )
- # Check input for `cond` or 'mask'
- if method == "cond" or method =='mask':
- # Check if we have an correlation matrix
- if corr_matrix is None:
- raise AttributeError(" The `corr_matrix` is missing")
- # check if corr_matrix and dt have the same entries
- # and align dt and corr_matrix entries
- _process_corr_mat(data=dt, corr_matrix=corr_matrix,
- index_col=Coloc_Expected.VARIANT_ID
- )
- # Check input for `cond`
- if method == "cond":
- # Check if MAF is present
- are_columns_in_df(dt, minor_allele_freq)
- # Is the trait standard deviation present or do we need to calculate it
- if not standard_deviation in dt.columns:
- dt[standard_deviation] = sd_pheno_approx(variancebeta=dt[varbeta],
- maf=dt[minor_allele_freq],
- n=dt[sample_size]
- )
- # Check if the sigsnps and dt have the same entries
- # if len(sigsnps) > 0:
- # are_Series_equal(dt[variant_id], pd.Series(sigsnps),
- # objects_names=['dt', 'sigsnps']
- # )
- # map p-values to z-values
- if 0 <= pthr <= 1:
- zthr = scipy.stats.norm.ppf(1- pthr/2)
- else:
- raise ValueError("Please supply `pthr` as a value between 0 and 1")
- # Calculate mean differences
- if method == "cond" and effect_type == "bin":
- if verbose is True:
- warnings.warn("[info] Approximating linear analysis of binary trait")
- # bin2lin will also check dt columns are correct.
- dt, _ = bin2lin(dt,
- variant_id=variant_id,
- effect_size=effect_size,
- varbeta=varbeta,
- event_rate=event_rate,
- sample_size=sample_size,
- minor_allele_freq=minor_allele_freq
- )
- # get YY
- YY = est_YY(dt=dt, effect_type=effect_type,
- standard_deviation=standard_deviation,
- sample_size=sample_size,
- event_rate=event_rate)
- # getting variants hits
- hits = pd.DataFrame()
- while hits.shape[1] < maxhits:
- if method == "mask":
- # Masking
- newhit = map_mask(dt=dt, corr_matrix=corr_matrix, r2thr=r2thr,
- sigsnps=hits.columns.unique().tolist(),
- variant_id=variant_id, effect_size=effect_size,
- zstatistic=zstatistic,
- varbeta=varbeta
- )
- else:
- # getting conditional hits
- newhit = map_cond(dt=dt, corr_matrix=corr_matrix, YY=YY,
- sigsnps=hits.columns.tolist(),
- effect_type=effect_type,
- variant_id=variant_id,
- effect_size=effect_size,
- varbeta=varbeta,
- event_rate=event_rate,
- sample_size=sample_size,
- minor_allele_freq=minor_allele_freq
- )
- # DO we need to exit?
- if newhit is None or newhit.empty:
- break
- # otherwise get value
- # NOTE might change newhit to a dictionary - this is just painfull
- value = np.array(newhit)[0][0]
- # if not sufficiently significant stop
- if np.abs(value) < zthr:
- break
- hits = pd.concat([hits, newhit], axis=1)
- # stop after one hit
- if method=="single":
- break
- return hits
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def gethits(hits, i, mode="iterative"):
- """
- Implements `gethits`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L378>`__.
- Arguments
- ---------
- hits : list of strings
- A list of strings of variant ids
- i : float
- mode : str, default iterative,
- Mode is either `iterative` or `allbutone`.
- Returns
- -------
- A numpy.array
- """
- # @@@ checking input
- is_type(hits, list)
- is_type(i, (np.int32,np.int16, np.int32, np.int64, int,float))
- is_type(mode, str)
- # check modes
- if not mode in Coloc_Expected.MODES:
- raise AttributeError(
- "`modes` should be a str: {}".format(Coloc_Expected.MODES)
- )
- # @@@ algorithm
- # copy, so we save the original
- hits_c = hits.copy()
- if mode == "allbutone":# mask everything else
- # remove the ith element
- del hits_c[i]
- ret = hits_c
- elif mode == "iterative": # mask everything earlier in list
- if i == 0:
- return np.empty(0, dtype='U5')
- else:
- # everything until i (excluding i)
- ret = hits_c[:i]
- # ret = hits[0:(i)]
- # return
- return np.setdiff1d(ret, "")
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def coloc_process(coloc_res,
- hits1: list=None,
- hits2: list=None,
- corr_exp: pd.DataFrame=None,
- corr_out: pd.DataFrame=None,
- r2thr=0.01,
- p1=10**-4,
- p2=10**-4,
- p12=5*10**-6,
- mode="iterative",
- verbose=True):
- """Generates a coloc results table, used by coloc_signal.
- Parameters
- ----------
- coloc_res : ColocResults,
- ColocResults object output by `coloc_single`.
- hits1, hits2 : list of strings,
- strings from `gethits` representing the hits in dataset1 and
- dataset2. If the array contain more than one lead variant -> masking.
- corr_(exp|out) : pd.DF,
- The correlation matrix of the variants in `dt`, for the exposure and
- outcome data respectively.
- r2thr : float, default 0.01
- Mask variants with a squared correlation with any of the `sigsnps`
- larger than `r2thr`.
- p1 : `float`, optional, default: `1E-4`
- Prior probability for H1 (variant only associated with trait 1)
- p2 : `float`, optional, default: `1E-4`
- Prior probability for H2 (variant only associated with trait 2)
- p12 : `float`, optional, default: `5E-6`,
- Prior probability for H4 (associated with both traits).
- mode : float, default: `iterative`,
- `iterative` or `allbutone`; to either successively condition on signals
- find all putative signals and condition on all but one of them in each
- analysis.
- Returns
- -------
- results : `pandas.DataFrame`
- The coloc results.
- Notes
- -----
- Implements `finemap.signals`, from coloc R, see the `chris wallace
- <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df8
- 4d7eba39d92a8c1/R/split.R#L408>`_ implementation.
- The H3 probability (causal variants for distinct traits) is calculated
- internally from the product of p1 and p2. See supplemental of the
- `original publication` <https://journals.plos.org/plosgenetics/article
- ?id=10.1371/journal.pgen.1004383#s5>`_.
- """
- # @@@ Check
- is_type(coloc_res, ColocResults)
- is_type(corr_exp, (pd.DataFrame, type(None)))
- is_type(corr_out, (pd.DataFrame, type(None)))
- is_type(hits1, (list, type(None)))
- is_type(hits2, (list, type(None)))
- is_type(r2thr, float)
- is_type(p1, float)
- is_type(p2, float)
- is_type(p12, float)
- is_type(mode, str)
- # check modes
- if not mode in Coloc_Expected.MODES:
- raise AttributeError(
- "`modes` should be a str: {}".format(Coloc_Expected.MODES)
- )
- # check rsquared
- if not (0 <= r2thr <= 1):
- raise ValueError('Please define r2thr within the range [0, 1]')
- # @@@ Actual calculations
- # NOTE `.f` is moved to `approx_bf`
- # see R coloc `.f`: https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L455
- if hits1 is None:
- hits1 = ['']
- #
- if hits2 is None:
- hits2 = ['']
- # one signal per trait
- if len(hits1) <= 1 and len(hits2) <= 1:
- res_dict = {
- c.COLOC_SUMMARY: coloc_res.results_dataframe(hit1=hits1,
- hit2=hits2),
- c.COLOC_RESULTS: coloc_res.variant_dataframe(),
- c.COLOC_PRIORS: pd.DataFrame(
- {
- c.COLOC_PRIOR1 : p1,
- c.COLOC_PRIOR2 : p2,
- c.COLOC_PRIOR12 : p12
- }, index=[0]
- )
- }
- # break function
- return res_dict
- # otherwise mask
- # check corr_matrix
- if (len(hits1) > 1 or len(hits2) > 1) and (corr_exp is None or corr_out is None):
- raise AttributeError('Please supply `corr_matrix` as a pd.DataFrame')
- # - allbutone : all but one signal from each trait with >1 signal
- # - iterative : successively one more signal from each trait with >1 signal
- ldfriends1 = np.power(corr_exp.loc[hits1, :], 2)
- ldfriends2 = np.power(corr_out.loc[hits2, :], 2)
- i = [idx for idx,value in enumerate(hits1)]
- j = [idx for idx,value in enumerate(hits2)]
- # procudes an i*j by 2 matrix
- todo=pd.DataFrame(itertools.product(i,j))
- # interchanges columns, followed by a transpose
- todo=todo.iloc[:, [1,0]].T
- # getting pd.Series with variant names
- newresult = coloc_res.variant_dataframe()[[c.UNIVERSAL_ID.name]].copy()
- res = pd.DataFrame()
- for r in list(range(0,todo.shape[1])):
- i = todo.iloc[1, r]
- j = todo.iloc[0, r]
- if verbose == True:
- print("[info] {0}".format(r+1))
- # collecting hits
- drop1, drop2 = ([] for i in range(2))
- # hit 1
- if len(hits1) > 1:
- # int(i) to get non-numpy ints
- ihits1 = gethits(hits1, int(i), mode)
- if len(ihits1) >= 1:
- ld_out1 = ldfriends1.loc[ihits1, ]
- drop1 = ldfriends1[ld_out1 > r2thr].\
- dropna(thresh=1, axis=1).columns.tolist()
- # hit 2
- if len(hits2) > 1:
- ihits2 = gethits(hits2, int(j), mode)
- if len(ihits2) >= 1:
- # if ihits2 != None:
- ld_out2 = ldfriends2.loc[ihits2, ]
- drop2 = ldfriends2[ld_out2 > r2thr].\
- dropna(thresh=1, axis=1).columns.tolist()
- # dropping variants
- # finding the union
- dropsnps = list(set(drop1+drop2))
- if verbose == True:
- # COLOR R uses message -> warnings
- warnings.warn("[info] dropping {0}/{1}: {2} (hits1) + {3} (hits2)"
- "".format(
- len(dropsnps),
- coloc_res.variant_dataframe().shape[0],
- len(drop1),
- len(drop2)
- )
- )
- # NOTE returns an empty dict
- if len(np.setdiff1d(coloc_res.variant_dataframe()[c.UNIVERSAL_ID.name],
- dropsnps)) <=1:
- res_dict = {
- c.COLOC_SUMMARY: pd.DataFrame(),
- c.COLOC_RESULTS: pd.DataFrame(),
- c.COLOC_PRIORS: pd.DataFrame(
- {
- c.COLOC_PRIOR1 : p1,
- c.COLOC_PRIOR2 : p2,
- c.COLOC_PRIOR12 : p12
- }, index=[0]
- )
- }
- return res_dict
- ### getting results object
- # NOTE ADD some of these strings to merit.constant
- df = coloc_res.variant_dataframe().copy()
- # get ABF
- df[Coloc_Expected.EXPOSURE_ABF] = np.where(
- df[c.UNIVERSAL_ID.name].isin(drop1), -1.1,
- df[Coloc_Expected.EXPOSURE_ABF])
- df[Coloc_Expected.OUTCOME_ABF] = np.where(
- df[c.UNIVERSAL_ID.name].isin(drop2), -1.1,
- df[Coloc_Expected.OUTCOME_ABF])
- # isum
- df.loc[:,c.COLOC_INTERNAL_SUM] =\
- np.where(df[c.UNIVERSAL_ID.name].isin(dropsnps), -1.1,
- df[c.COLOC_INTERNAL_SUM].copy()).copy()
- my_denom_log_abf = _logsum(df[c.COLOC_INTERNAL_SUM].values)
- # new results
- sum_temp1 = np.exp(
- df[c.COLOC_INTERNAL_SUM].copy().values - my_denom_log_abf
- ).tolist()
- newresult["SNP.PP.H4.row"+str(r)] = sum_temp1
- newresult.loc[:,"z.df1.row"+str(r)] = np.where(
- df[c.UNIVERSAL_ID.name].isin(drop1), 0,
- df[Coloc_Expected.EXPOSURE_Z_STATISTICS].copy()
- ).copy()
- newresult.loc[:, "z.df2.row"+str(r)] = np.where(
- df[c.UNIVERSAL_ID.name].isin(drop2), 0,
- df[Coloc_Expected.OUTCOME_ABF].copy()
- ).copy()
- # subsetting df3
- df3 = pd.DataFrame(itertools.product(
- coloc_res.variant_dataframe()[c.UNIVERSAL_ID.name],
- coloc_res.variant_dataframe()[c.UNIVERSAL_ID.name]),
- columns=["uni_id_exposure", "uni_id_outcome"]
- )
- df3 = df3.merge(
- df[[Coloc_Expected.EXPOSURE_ABF, c.UNIVERSAL_ID.name]],
- left_on = "uni_id_exposure", right_on = c.UNIVERSAL_ID.name
- )
- df3 = df3.merge(
- df[[Coloc_Expected.OUTCOME_ABF, c.UNIVERSAL_ID.name]],
- left_on = "uni_id_outcome", right_on = c.UNIVERSAL_ID.name
- )
- df3 = df3[["uni_id_exposure",
- "uni_id_outcome",
- Coloc_Expected.EXPOSURE_ABF,
- Coloc_Expected.OUTCOME_ABF
- ]]
- # subsetting ln_abs_h3
- df3["ln_abf_h3"] = df3[Coloc_Expected.EXPOSURE_ABF] +\
- df3[Coloc_Expected.OUTCOME_ABF].copy()
- df3["ln_abf_h3"] = np.where(
- df3["uni_id_exposure"].isin(drop1), 0, df3["ln_abf_h3"])
- df3["ln_abf_h3"] = np.where(
- df3["uni_id_outcome"].isin(drop2), 0, df3["ln_abf_h3"])
- # res object
- # df3.ln_abf_h3 and df.isum are empty unless `method == mask`
- ret = pd.DataFrame(approx_bf(df.ln_abf_exposure,
- df.ln_abf_outcome,
- masking_ln_abf3=df3.ln_abf_h3,
- masking_ln_abf4=df.isum,
- p1=p1,
- p2=p2,
- p12=p12,
- extra_info=True),
- index=[0])
- # get best varaints
- # ret[c.COLOC_BEST1] = hits1[todo.iloc[1,r]]
- # ret[c.COLOC_BEST2] = hits2[todo.iloc[0,r]]
- # hits
- ret[c.COLOC_HIT1] = hits1[todo.iloc[1,r]]
- ret[c.COLOC_HIT2] = hits2[todo.iloc[0,r]]
- # number of variants
- ret['nsnps'] = df.shape[0]
- res = pd.concat([res, ret])
- # adding hits
- res_dict = {
- c.COLOC_SUMMARY: res,
- c.COLOC_RESULTS: newresult,
- c.COLOC_PRIORS: pd.DataFrame(
- {
- c.COLOC_PRIOR1 : p1,
- c.COLOC_PRIOR2 : p2,
- c.COLOC_PRIOR12 : p12
- }, index=[0]
- )
- }
- return res_dict
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def coloc_single(inst: pd.DataFrame,
- p1: float=10**-4,
- p2: float=10**-4,
- p12: float=5*10**-6,
- effect_type: list=["quant","quant"],
- variant_id: str=Coloc_Expected.VARIANT_ID,
- exposure_effect_size: str=Coloc_Expected.EXPOSURE_EFFECT_SIZE,
- outcome_effect_size: str=Coloc_Expected.OUTCOME_EFFECT_SIZE,
- exposure_varbeta: str=Coloc_Expected.EXPOSURE_VARBETA,
- outcome_varbeta: str=Coloc_Expected.OUTCOME_VARBETA,
- exposure_zstatistic: str=Coloc_Expected.EXPOSURE_Z_STATISTICS,
- outcome_zstatistic: str=Coloc_Expected.OUTCOME_Z_STATISTICS,
- exposure_standard_deviation: str=Coloc_Expected.EXPOSURE_STANDARD_DEVIATION,
- outcome_standard_deviation: str=Coloc_Expected.OUTCOME_STANDARD_DEVIATION,
- exposure_sample_size: str=Coloc_Expected.EXPOSURE_SAMPLE_SIZE,
- outcome_sample_size: str=Coloc_Expected.OUTCOME_SAMPLE_SIZE,
- minor_allele_freq: str=Coloc_Expected.MINOR_ALLELE_FREQ,
- extra_info: bool=False):
- """
- Coloc main function, is similar to coloc.details in R coloc. It started
- out as a port of `coloc.abf` (assuming a single causal variant).
- Will calculate `sdY`, the phenotypic standard deviation if this is not
- provided by `inst` and type == `quant`.
- Will compute the z-statistic if not provided from the effect_size and
- varbeta.
- Similar to `coloc.detail`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L21>`__.
- And to `coloc.abf`:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L21>`__.
- Arguments
- ---------
- inst : pd.DF
- A coloc.DataFrame.
- p1, p2, p12 : float,
- prior probabilities for H1 (variant only associated with trait 1),
- for H2 (only with trait 2), or H4 (associated with both traits).
- See supplemental of:
- `Here <https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1004383#s5>`__.
- effect_type : list of str, default `quant`,
- The effect_types or the `_exposure` and `_outcome` data.
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- effect_size : str,
- The effect size column names for the `exposure` and `outcome` data.
- varbeta : str,
- The variance of the effect size column names for the
- `exposure` and `outcome` data.
- zstatistics : str,
- The z-statistics column names for the `exposure` and `outcome` data.
- Note these can be readily obtained from the GWAS p-values. Can be
- estimated from the effect_size and varbeta columns, if absent.
- standard_deviation : str,
- The column name of the gwas trait variance (will be estimated if
- absent).
- sample_size : str,
- The column name of the GWAS sample size.
- minor_allele_freq : str,
- The column name of the MAF.
- extra_info : boolean, default False
- If additional information should be returned (mimicking coloc.details),
- otherwise (defaults) output will be similar to coloc.abf.
- Returns
- -------
- A `ColocResults` object.
- """
- # @@@ checks
- # core data needed
- core_cols = [variant_id, exposure_varbeta, outcome_varbeta,
- exposure_zstatistic, outcome_zstatistic, minor_allele_freq ]
- # are_columns_in_df(inst, core_cols)
- # unpacking effect type
- if np.all([ t in Coloc_Expected.TRAIT_TYPES for t in effect_type ]):
- type_exposure, type_outcome = effect_type
- # if needed estimate sdY
- # check if standard deviation is provided
- if (not exposure_standard_deviation in inst.columns) and\
- (type_exposure == "quant"):
- inst[exposure_standard_deviation] =\
- sd_pheno_approx(
- variancebeta=inst[exposure_varbeta],
- maf=inst[minor_allele_freq],
- n=inst[exposure_sample_size]
- )
- if (not outcome_standard_deviation in inst.columns) and\
- (type_outcome == "quant"):
- inst[outcome_standard_deviation] =\
- sd_pheno_approx(
- variancebeta=inst[outcome_varbeta],
- maf=inst[minor_allele_freq],
- n=inst[outcome_sample_size]
- )
- # set the variance prior
- if type_exposure == "quant":
- sd_prior_exp = 0.15 * inst[exposure_standard_deviation]
- elif type_exposure == "bin":
- # for case control (binary data)
- sd_prior_exp = 0.20
- else:
- raise ValueError("Please specify `effect_type` two element list with {0} or {1}".\
- format(*Coloc_Expected.TRAIT_TYPES))
- if type_outcome == "quant":
- sd_prior_out = 0.15 * inst[outcome_standard_deviation]
- elif type_outcome == "bin":
- # for case-control (binary data)
- sd_prior_out = 0.20
- else:
- raise ValueError("Please specify `effect_type` two element list with {0} or {1}".\
- format(*Coloc_Expected.TRAIT_TYPES))
- # if needed calculate z normal deviate
- if not exposure_zstatistic in inst.columns:
- inst[exposure_zstatistic] =\
- np.divide( inst[exposure_effect_size], np.sqrt(inst[exposure_varbeta])
- )
- if not outcome_zstatistic in inst.columns:
- inst[outcome_zstatistic] =\
- np.divide( inst[outcome_effect_size], np.sqrt(inst[outcome_varbeta])
- )
- # extract core columns
- dt = inst[core_cols].copy()
- # Shrinkage factor: ratio of the prior variance to the total variance
- dt[Coloc_Expected.EXPOSURE_SHRINKAGE_FACTOR] = \
- np.divide(
- np.power(sd_prior_exp, 2),
- (np.power(sd_prior_exp, 2) + dt[exposure_varbeta])
- )
- dt[Coloc_Expected.OUTCOME_SHRINKAGE_FACTOR] = \
- np.divide(
- np.power(sd_prior_out, 2),
- (np.power(sd_prior_out, 2) + dt[outcome_varbeta])
- )
- # appending the log (approximate) Bayes Factor for colocalization
- dt[Coloc_Expected.EXPOSURE_ABF] =\
- (0.5 *
- (
- np.log(1 - dt[Coloc_Expected.EXPOSURE_SHRINKAGE_FACTOR]) +
- dt[Coloc_Expected.EXPOSURE_SHRINKAGE_FACTOR] *
- np.power( dt[Coloc_Expected.EXPOSURE_Z_STATISTICS], 2)
- )
- )
- dt[Coloc_Expected.OUTCOME_ABF] =\
- (0.5 *
- (
- np.log(1 - dt[Coloc_Expected.OUTCOME_SHRINKAGE_FACTOR]) +
- dt[Coloc_Expected.OUTCOME_SHRINKAGE_FACTOR] *
- np.power( dt[Coloc_Expected.OUTCOME_Z_STATISTICS], 2)
- )
- )
- # INTERNAL SUM Just wrapping in tuple to get better code-alignment
- dt[c.COLOC_INTERNAL_SUM] = (dt[Coloc_Expected.EXPOSURE_ABF] +
- dt[Coloc_Expected.OUTCOME_ABF]
- )
- # posterior prob that each SNP is THE causal variant for a shared signal
- dt[c.COLOC_PPH4] = np.exp(dt[c.COLOC_INTERNAL_SUM] -
- _logsum(dt[c.COLOC_INTERNAL_SUM]))
- # coloc summary following classic coloc
- if extra_info is False:
- summary = approx_bf(ln_abf1=dt[Coloc_Expected.EXPOSURE_ABF],
- ln_abf2=dt[Coloc_Expected.OUTCOME_ABF],
- p1=p1, p2=p2, p12=p12, extra_info=extra_info
- )
- # results
- res = {c.COLOC_SUMMARY: summary, c.COLOC_VARIANT_RESULTS: dt}
- elif extra_info is True:
- summary = approx_bf(ln_abf1=dt[Coloc_Expected.EXPOSURE_ABF],
- ln_abf2=dt[Coloc_Expected.OUTCOME_ABF],
- p1=p1, p2=p2, p12=p12, extra_info=extra_info
- )
- summary[c.COLOC_BEST1] = dt.loc[ summary[c.COLOC_BEST1], variant_id]
- summary[c.COLOC_BEST2] = dt.loc[ summary[c.COLOC_BEST2], variant_id]
- summary[c.COLOC_BEST4] = dt.loc[ summary[c.COLOC_BEST4], variant_id]
- # results
- res = {c.COLOC_SUMMARY: summary, c.COLOC_VARIANT_RESULTS: dt}
- else:
- raise ValueError("`extra_info` is a boolean")
- # returning
- return ColocResults(**res)
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- class ColocResults(object):
- """Return Results object."""
- ALL_ARGS = [c.COLOC_SUMMARY, c.COLOC_VARIANT_RESULTS]
- # SET_ARGS = ['uni_id', 'b', 'se_fixed', 'pvalue_fixed', 'instrument']
- IGNORE_ARGS = [c.COLOC_VARIANT_RESULTS]
- # OUT_ARGS = [el for el in ALL_ARGS if el not in IGNORE_ARGS]
- OUT_ARGS = [c.COLOC_SUMMARY]
- def __init__(self, **kwargs):
- """
- Initialise
- """
- # Check that we have not supplied any arguments that can't be set
- for k in kwargs.keys():
- if k not in self.__class__.ALL_ARGS:
- raise AttributeError("unrecognised argument '{0}'".format(k))
- for s in self.__class__.ALL_ARGS:
- try:
- setattr(self, s, kwargs[s])
- except KeyError:
- warnings.warn("argument '{0}' is set to 'None'".format(s))
- setattr(self, s, None)
- def __repr__(self):
- values = ",".join(["{0}={1}".format(s, getattr(self, s))
- for s in self.__class__.ALL_ARGS if s not in
- self.__class__.IGNORE_ARGS])
- return "{0}({1})".format(self.__class__.__name__, values)
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- @property
- def nsnps(self):
- """Return the number of variants."""
- return len(self.variant_results[c.UNIVERSAL_ID.name])
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def results_dataframe(self, **kwargs):
- """Return dataframe with summary results."""
- pp_dict = [getattr(self, s) for s in self.__class__.OUT_ARGS][0]
- # check if is a dict or a pd.DF
- # if a pd.DF - assume it is already formatted.
- if isinstance(pp_dict, dict):
- dt = pd.DataFrame([[float(self.nsnps)] + list(pp_dict.values())],
- columns=[c.MR_NSNPS] + list(pp_dict.keys()))
- for k, v in kwargs.items():
- dt.loc[:, k] = v
- elif isinstance(pp_dict, pd.DataFrame):
- dt = pp_dict
- # return stuff
- return dt
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- def variant_dataframe(self, **kwargs):
- """Return dataframe with individual variant results."""
- return self.variant_results
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # def results_H3_dataframe(self, **kwargs):
- # """Return dataframe with individual variant results."""
- # return self.results_H3
- # @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
- # MAIN ENTRY POINT
- # ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- # DONE
- def coloc_signals(input_data,
- corr_matrix,
- method=["single", "single"],
- mode="iterative",
- p1=1e-4,
- p2=1e-4,
- p12=5*10**-6,
- maxhits=3,
- r2thr=0.01,
- pthr = 1e-06,
- verbose=True,
- effect_type=["quant","quant"],
- variant_id=Coloc_Expected.VARIANT_ID,
- exposure_effect_size=Coloc_Expected.EXPOSURE_EFFECT_SIZE,
- outcome_effect_size=Coloc_Expected.OUTCOME_EFFECT_SIZE,
- exposure_varbeta=Coloc_Expected.EXPOSURE_VARBETA,
- outcome_varbeta=Coloc_Expected.OUTCOME_VARBETA,
- exposure_zstatistic: str=Coloc_Expected.EXPOSURE_Z_STATISTICS,
- outcome_zstatistic: str=Coloc_Expected.OUTCOME_Z_STATISTICS,
- exposure_standard_deviation=Coloc_Expected.EXPOSURE_STANDARD_DEVIATION,
- outcome_standard_deviation=Coloc_Expected.OUTCOME_STANDARD_DEVIATION,
- exposure_sample_size=Coloc_Expected.EXPOSURE_SAMPLE_SIZE,
- outcome_sample_size=Coloc_Expected.OUTCOME_SAMPLE_SIZE,
- exposure_event_rate=Coloc_Expected.EXPOSURE_EVENT_RATE,
- outcome_event_rate=Coloc_Expected.OUTCOME_EVENT_RATE,
- minor_allele_freq=Coloc_Expected.MINOR_ALLELE_FREQ
- ):
- """
- The main entry point for coloc to consider multiple causal signals.
- Implements `coloc.signals`, from coloc R:
- `Here <https://github.com/chr1swallace/coloc/blob/eca81cdd0e5cc9fc0c5dd4df84d7eba39d92a8c1/R/split.R#L760>`__.
- Arguments
- ---------
- input_data : pd.DF,
- A coloc pd.dataframe, with columns ... . The original coloc `data1` and
- `data2` are combined in this single dataframe, with column suffixes
- `_exposure` and `_outcome` respectivly.
- corr_matrix : pd.DF,
- The correlation matrix of the variants in `input_data`. If the expousre
- and outcome GWAS are from distinct populations please supply a list of
- two matrices [exposure_cor, outcome_cor]. Note is only required when
- `method` is `cond` or 'mask'.
- method : list of two strings, default `single`,
- The method to apply to the `_exposure` and `_outcome` data.
- Use `single` to perform coloc assuming there is a single causal
- variant - equivalent to the original coloc implementation,
- Use `cond` to perform a stepwise conditional analysis,
- Use `mask` to perform the stepwise selection algorithm using masking
- instead.
- mode : float, default `iterative`,
- `iterative` or `allbutone`; to either successively condition on signals
- find all putative signals and condition on all but one of them in each
- analysis.
- p1, p2, p12 : float,
- prior probabilities for H1 (any random variant only associated with
- trait 1), for H2 (only with trait 2), or H4 (associated with both traits).
- H3 (causal variants for distinct traits) is simply the product of
- p1 and p2. See supplemental of:
- `Here <https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1004383#s5>`__.
- Note that when the data is sufficiently large, p1 and p2 can often be
- estimated from the emperical data under consideration.
- pthr : float, default 1 \times 10^{-6},
- Stops when p-values are larger than `pthr`, where input should be
- defined between 0 and 1.
- maxhits : int, default 3,
- Stops when the procedures finds `maxhits`.
- r2thr : float, default 0.01
- Mask variants with a squared correlation with any of the `sigsnps`
- larger than `r2thr`.
- effect_type : list of str, default `quant`,
- The effect_types or the `_exposure` and `_outcome` data.
- If the GWAS trait was `quant`: continuous/quantative, or `bin`:
- case-control/binary.
- Returns
- -------
- Resutns a pd.DF of coloc results.
- """
- # @@@ CHECKING INPUT
- # Check if core data is present
- is_type(input_data, pd.DataFrame)
- is_type(corr_matrix, (pd.DataFrame, type(None), list))
- is_type(effect_type, list)
- is_type(method, list)
- is_type(mode, str)
- # expected columns
- Exposure_col = [variant_id, exposure_effect_size, exposure_varbeta,
- exposure_zstatistic, exposure_standard_deviation,
- exposure_sample_size, exposure_event_rate,
- minor_allele_freq]
- Outcome_col = [variant_id, outcome_effect_size, outcome_varbeta,
- outcome_zstatistic, outcome_standard_deviation,
- outcome_sample_size, outcome_event_rate,
- minor_allele_freq]
- # Check if effect types are correct
- if np.all([ t in Coloc_Expected.TRAIT_TYPES for t in effect_type ]):
- effect_type_exp = effect_type[0]
- effect_type_out = effect_type[1]
- else:
- raise ValueError(
- "`effect_type` should be a lis with two entries. Pick from: {}".\
- format(Coloc_Expected.TRAIT_TYPES)
- )
- # Checking methods
- if len(method) != 2:
- raise AttributeError( "`method` should be a list with two entries")
- if not np.all([ t in Coloc_Expected.METHODS for t in method ]):
- raise ValueError(
- "`method` should be a list with two entries. Pick from: {}".\
- format(Coloc_Expected.METHODS)
- )
- else:
- method_exp=method[0]
- method_out=method[1]
- # check modes
- if not mode in Coloc_Expected.MODES:
- raise ValueError(
- "`modes` should be a str: {}".format(Coloc_Expected.MODES)
- )
- # Checking columns content
- # are_columns_in_df(input_data, list(set(Coloc_Expected.exposure_columns +
- # Coloc_Expected.outcome_columns)))
- # Check if all SNPs are in the corr_matrix
- if np.any(method_exp in ['cond', 'mask']) or np.any(method_out in ['cond', 'mask']):
- if isinstance(corr_matrix, list):
- if len(corr_matrix) == 2:
- corr_exp = corr_matrix[0]
- corr_out = corr_matrix[1]
- else:
- raise IndexError("`corr_matrix`, please supply a single pd.DataFrame, or a list of two pd.DataFrame's")
- else:
- corr_exp = corr_matrix.copy()
- corr_out = corr_matrix.copy()
- print('test1')
- # align with dt
- _process_corr_mat(data=input_data, corr_matrix=corr_exp,
- index_col=Coloc_Expected.VARIANT_ID
- )
- _process_corr_mat(data=input_data, corr_matrix=corr_out,
- index_col=Coloc_Expected.VARIANT_ID
- )
- else:
- # if single just assing whatever is corr_matrix
- corr_exp = corr_matrix
- corr_out = corr_matrix
- # add standard deviation of Y
- if not Coloc_Expected.EXPOSURE_STANDARD_DEVIATION in input_data.columns:
- input_data[exposure_standard_deviation] =\
- sd_pheno_approx(
- variancebeta=input_data[exposure_varbeta],
- maf=input_data[minor_allele_freq],
- n=input_data[exposure_sample_size]
- )
- if not Coloc_Expected.OUTCOME_STANDARD_DEVIATION in input_data.columns:
- input_data[outcome_standard_deviation] =\
- sd_pheno_approx(
- variancebeta=input_data[outcome_varbeta],
- maf=input_data[minor_allele_freq],
- n=input_data[outcome_sample_size]
- )
- # map p-value to z-statistics
- if 0 <= pthr <= 1:
- zthr = scipy.stats.norm.ppf(1- pthr/2)
- else:
- raise ValueError("Please supply `pthr` as a value between 0 and 1")
- # @@@ DOING ACTUAL ANALYSES
- # Finemaping
- if verbose is True:
- print("[info] Fine-mapping exposure")
- # NOTE the .copy() is essential here because it overwrites some columns
- # when set to effect_type = 'bin'
- fm1 = finemap_signals(
- input_data.loc[:,input_data.columns.isin(Exposure_col)].copy(),
- corr_matrix=corr_exp, method=method_exp, maxhits=maxhits,
- r2thr=r2thr, pthr=pthr, effect_type=effect_type_exp,
- variant_id=variant_id,
- effect_size=exposure_effect_size,
- varbeta=exposure_varbeta,
- zstatistic=exposure_zstatistic,
- standard_deviation=exposure_standard_deviation,
- sample_size=exposure_sample_size,
- event_rate=exposure_event_rate,
- minor_allele_freq=minor_allele_freq
- )
- if verbose is True:
- print("[info] Variants independently associated in the exposure dataset")
- print(fm1)
- print('')
- print("[info] Fine-mapping outcome")
- fm2 = finemap_signals(
- input_data.loc[:,input_data.columns.isin(Outcome_col)].copy(),
- corr_matrix=corr_out, method=method_out, maxhits=maxhits,
- r2thr=r2thr, pthr=pthr, effect_type=effect_type_out,
- variant_id=variant_id,
- effect_size=outcome_effect_size,
- varbeta=outcome_varbeta,
- zstatistic=outcome_zstatistic,
- standard_deviation=outcome_standard_deviation,
- sample_size=outcome_sample_size,
- event_rate=outcome_event_rate,
- minor_allele_freq=minor_allele_freq
- )
- if verbose is True:
- print("[info] Variants independently associated in the outcome dataset")
- print(fm2)
- if fm1.empty:
- fm1 = find_best_signal(input_data,
- variant_id=variant_id,
- effect_size=exposure_effect_size,
- varbeta=exposure_varbeta,
- )
- if fm2.empty:
- fm2 = find_best_signal(input_data,
- variant_id=variant_id,
- effect_size=outcome_effect_size,
- varbeta=outcome_varbeta,
- )
- # conditionals if needed
- if (not fm1.empty) and (method_exp == "cond"):
- # NOTE the copy() is intended, the function overwrites columns
- cond1 = est_all_cond(
- dt=input_data.loc[:,input_data.columns.isin(Exposure_col)].copy(),
- corr_matrix=corr_exp,
- fm=fm1, mode=mode, effect_type=effect_type_exp,
- variant_id=variant_id,
- effect_size=exposure_effect_size,
- varbeta=exposure_varbeta,
- event_rate=exposure_event_rate,
- sample_size=exposure_sample_size,
- standard_deviation=exposure_standard_deviation,
- minor_allele_freq=minor_allele_freq,
- label=Coloc_Expected.exp_name
- )
- if (not fm2.empty) and (method_out == "cond"):
- cond2 = est_all_cond(
- dt=input_data.loc[:,input_data.columns.isin(Outcome_col)].copy(),
- corr_matrix=corr_out,
- fm=fm2, mode=mode, effect_type=effect_type_out,
- variant_id=variant_id,
- effect_size=outcome_effect_size,
- varbeta=outcome_varbeta,
- event_rate=outcome_event_rate,
- sample_size=outcome_sample_size,
- standard_deviation=outcome_standard_deviation,
- minor_allele_freq=minor_allele_freq,
- label=Coloc_Expected.out_name
- )
- # doublemask
- doublemask_method = ['single', 'mask']
- if not np.all([ t in Coloc_Expected.METHODS for t in doublemask_method ]):
- raise AttributeError(
- "`doublemask_method` should be a list with two entries. Pick from: {}".\
- format(Coloc_Expected.METHODS)
- )
- if (method_exp in doublemask_method) and (method_out in doublemask_method):
- # coloc_single == coloc.detail
- coloc_res = coloc_single(inst =input_data.copy(),
- p1=p1, p2=p2, p12=p12,
- effect_type=effect_type,
- variant_id=variant_id,
- exposure_effect_size=exposure_effect_size,
- outcome_effect_size=outcome_effect_size,
- exposure_varbeta=exposure_varbeta,
- outcome_varbeta=outcome_varbeta,
- exposure_zstatistic=exposure_zstatistic,
- outcome_zstatistic=outcome_zstatistic,
- exposure_standard_deviation=exposure_standard_deviation,
- outcome_standard_deviation=outcome_standard_deviation,
- exposure_sample_size=exposure_sample_size,
- outcome_sample_size=outcome_sample_size,
- minor_allele_freq=minor_allele_freq,
- extra_info=True)
- # get results
- proc_res_1 = coloc_process(coloc_res,
- hits1=fm1.columns.tolist(),
- hits2=fm2.columns.tolist(),
- corr_exp = corr_exp,
- corr_out = corr_out,
- r2thr=r2thr, mode=mode,
- p1=p1, p2=p2, p12=p12,
- verbose=verbose
- )
- # get summary
- proc_res = proc_res_1[c.COLOC_SUMMARY]
- # double cond
- if (method_exp in ["cond"]) and (method_out in ["cond"]):
- i = [idx for idx, value in enumerate(cond1)]
- j = [idx for idx, value in enumerate(cond2)]
- todo = pd.DataFrame(itertools.product(i,j), columns=["i", "j"])
- # resuts pd.DF
- proc_res = pd.DataFrame()
- proc_res_variants = pd.DataFrame(
- {c.UNIVERSAL_ID.name: input_data[variant_id]})
- # looping over rows
- for k in list(range(todo.shape[0])):
- cond_dt = pd.merge( cond1[list(cond1.keys())[todo.loc[k,"i"]]],
- cond2[list(cond2.keys())[todo.loc[k,"j"]]])
- ### Check if all the columns are there
- column_check = [exposure_sample_size, outcome_sample_size,
- exposure_standard_deviation,
- outcome_standard_deviation,
- minor_allele_freq]
- cond_dt = _add_missing_columns(cond_dt, input_data.copy(),
- columns=column_check,
- index_col=variant_id)
- # Getting coloc results
- coloc_res = coloc_single(inst=cond_dt,
- p1=p1, p2=p2, p12=p12,
- effect_type=effect_type,
- variant_id=variant_id,
- exposure_effect_size=exposure_effect_size,
- outcome_effect_size=outcome_effect_size,
- exposure_varbeta=exposure_varbeta,
- outcome_varbeta=outcome_varbeta,
- exposure_zstatistic=exposure_zstatistic,
- outcome_zstatistic=outcome_zstatistic,
- exposure_standard_deviation=exposure_standard_deviation,
- outcome_standard_deviation=outcome_standard_deviation,
- exposure_sample_size=exposure_sample_size,
- outcome_sample_size=outcome_sample_size,
- minor_allele_freq=minor_allele_freq,
- extra_info=True)
- # Processing
- proc_res_1 = coloc_process(coloc_res,
- hits1=[fm1.columns.tolist()[todo.loc[k,"i"]]],
- hits2=[fm2.columns.tolist()[todo.loc[k,"j"]]],
- corr_exp = corr_exp,
- corr_out = corr_out,
- r2thr=r2thr, mode=mode, p1=p1, p2=p2,
- p12=p12,
- verbose=verbose
- )
- # Getting Summary reults
- proc_res= pd.concat([proc_res, proc_res_1[c.COLOC_SUMMARY]])
- # merging — use iteration-specific suffixes to avoid
- # duplicate column names across signal-pair iterations
- proc_res_variants = pd.merge(
- proc_res_variants, coloc_res.variant_dataframe(),
- left_on=c.UNIVERSAL_ID.name, right_on=c.UNIVERSAL_ID.name,
- suffixes=("", "_{0}".format(k))
- )
- # cond mask/-
- if method_exp in ["cond"] and method_out in ["mask", "single"]:
- proc_res = pd.DataFrame()
- # Looping over the keys
- for k in list(range(len(cond1))):
- # this merges cond1 = exposure with outcome data
- cond_dt = pd.merge(
- cond1[list(cond1.keys())[k]],
- input_data.loc[:,input_data.columns.isin(Outcome_col)].copy(),
- )
- ### adding back potentially missing data
- column_check = [exposure_sample_size, outcome_sample_size,
- exposure_standard_deviation,
- outcome_standard_deviation,
- minor_allele_freq]
- cond_dt = _add_missing_columns(cond_dt, input_data,
- columns=column_check,
- index_col=variant_id)
- # run coloc
- coloc_res = coloc_single(inst=cond_dt,
- p1=p1, p2=p2, p12=p12,
- effect_type=effect_type,
- variant_id=variant_id,
- exposure_effect_size=exposure_effect_size,
- outcome_effect_size=outcome_effect_size,
- exposure_varbeta=exposure_varbeta,
- outcome_varbeta=outcome_varbeta,
- exposure_zstatistic=exposure_zstatistic,
- outcome_zstatistic=outcome_zstatistic,
- exposure_standard_deviation=exposure_standard_deviation,
- outcome_standard_deviation=outcome_standard_deviation,
- exposure_sample_size=exposure_sample_size,
- outcome_sample_size=outcome_sample_size,
- minor_allele_freq=minor_allele_freq,
- extra_info=True)
- # get summary
- proc_res_1 = coloc_process(coloc_res,
- hits1=[fm1.columns.tolist()[k]],
- hits2=fm2.columns.tolist(),
- corr_exp = corr_exp,
- corr_out = corr_out,
- r2thr=r2thr,
- mode=mode, p1=p1, p2=p2, p12=p12,
- verbose=verbose
- )
- proc_res = pd.concat([proc_res, proc_res_1[c.COLOC_SUMMARY]])
- ## mask/- cond
- if method_exp in ["mask", "single"] and method_out in ["cond"]:
- proc_res = pd.DataFrame()
- # NOTE change with `enumerate`?
- for k in list(range(0,len(cond2))):
- # this merges cond2 = outcomes with exposure data
- cond_dt = pd.merge(
- cond2[list(cond2.keys())[k]],
- input_data.loc[:,input_data.columns.isin(Exposure_col)].copy()
- )
- ### adding back potentially missing data
- column_check = [exposure_sample_size, outcome_sample_size,
- exposure_standard_deviation,
- outcome_standard_deviation,
- minor_allele_freq]
- cond_dt = _add_missing_columns(cond_dt, input_data,
- columns=column_check,
- index_col=variant_id)
- # COLOC results
- coloc_res = coloc_single(inst=cond_dt,
- p1=p1, p2=p2, p12=p12,
- effect_type=effect_type,
- variant_id=variant_id,
- exposure_effect_size=exposure_effect_size,
- outcome_effect_size=outcome_effect_size,
- exposure_varbeta=exposure_varbeta,
- outcome_varbeta=outcome_varbeta,
- exposure_zstatistic=exposure_zstatistic,
- outcome_zstatistic=outcome_zstatistic,
- exposure_standard_deviation=exposure_standard_deviation,
- outcome_standard_deviation=outcome_standard_deviation,
- exposure_sample_size=exposure_sample_size,
- outcome_sample_size=outcome_sample_size,
- minor_allele_freq=minor_allele_freq,
- extra_info=True)
- # get summary
- proc_res_1 = coloc_process(coloc_res,
- hits1=fm1.columns.tolist(),
- hits2=[fm2.columns.tolist()[k]],
- corr_exp = corr_exp,
- corr_out = corr_out,
- r2thr=r2thr,
- mode=mode, p12=p12, p1=p1, p2=p2,
- verbose=verbose
- )
- proc_res = pd.concat([proc_res,
- proc_res_1[c.COLOC_SUMMARY]])
- # returning results objects
- proc_res.reset_index(inplace=True)
- proc_res[c.MR_NSNPS] = proc_res[c.MR_NSNPS].astype(str)
- results = proc_res[
- [c.COLOC_HIT1, c.COLOC_HIT2,
- c.MR_NSNPS, c.COLOC_PPH0, c.COLOC_PPH1,
- c.COLOC_PPH2, c.COLOC_PPH3, c.COLOC_PPH4,
- c.COLOC_BEST1, c.COLOC_BEST2, c.COLOC_BEST4]
- ]
- # add zstats
- try:
- results[c.COLOC_ZSTAT_HIT1] = fm1[results[c.COLOC_HIT1]].apply(pd.Series)
- results[c.COLOC_ZSTAT_HIT2] = fm2[results[c.COLOC_HIT2]].apply(pd.Series)
- except:
- results[c.COLOC_ZSTAT_HIT1] = fm1[results[c.COLOC_HIT1]].iloc[0].to_list()
- results[c.COLOC_ZSTAT_HIT2] = fm2[results[c.COLOC_HIT2]].iloc[0].to_list()
- # returning ColocResults objects
- res_obj = {c.COLOC_SUMMARY: results,
- c.COLOC_VARIANT_RESULTS: proc_res_1[c.COLOC_RESULTS]}
- # return stuff
- # return results
- return ColocResults(**res_obj)
coloc_core.py at commit 0e05762, under GPL-3.0 · at the source
Overview
- Department of Clinical, Educational and Health Psychology, University College London,London, UK
- Institute of Cardiovascular Science, University College London,London, UK
- The National Institute for Health Research University College London Hospitals Biomedical Research Centre, University College London,London, UK
- Department of Cardiology, Division of Heart and Lungs, University Medical Centre Utrecht, Utrecht University,Utrecht, the Netherlands
- Department of Cardiology, Amsterdam Cardiovascular Sciences, Amsterdam University Medical Centre, University of Amsterdam,Amsterdam, the Netherlands
- Institute of Health Informatics, University College London,London, UK
- Developmental Neurosciences, Zayed Centre for Research into Rare Disease in Children, UCL GOS Institute of Child Health, University College London,London, UK
- Division of Psychiatry, University College London,London, UK
- Social Genetic and Developmental Psychiatry, King’s College London,London, UK
Abstract
Major depression (MD) treatments have limited efficacy and target few mechanisms, highlighting the need for innovative drug discovery. Drugs targeting genetically supported proteins are 2.6 times more likely to succeed in drug development. Here, we use genetic methods to identify and prioritise MD drug targets, leveraging genome-wide association study (GWAS) summary statistics from >525,000 MD cases. We derived exposure data from 10 datasets measuring protein quantitative trait loci (pQTLs) and gene expression levels (eQTLs) in blood, cerebrospinal fluid, and brain tissues. We performed cis-Mendelian randomisation (MR) on 3469 druggable targets (genes encoding proteins targeted by existing compounds or experimentally predicted to be druggable). To strengthen causal inference, we implemented robust MR estimators, colocalisation, external replication, and assessed directional consistency across tissues. We integrated cis-MR effect directions with drug mechanisms and clinical annotations to infer potential therapeutic effects. Validation analyses showed that 82% of drugs approved for depression/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.
cfinan/merit
0e05762f86400756a0152cb5aa645edb81d8d831, 4 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
476 files
- docs/
source/ , Python, 98 linesconf.py - merit/
__init__.py , Python, 16 lines - merit/
coloc/ , Python, 17 lines__init__.py - merit/
coloc/ , Python, 2,566 lines, 2 matchescoloc_core.py - merit/
coloc/ , Python, 359 linescoloc_sensitivity.py - merit/
common.py , Python, 24 lines - merit/
constants.py , Python, 794 lines - merit/
errors.py , Python, 1,702 lines - merit/
example_data/ , Python, 1 line__init__.py - merit/
example_data/ , Python, 79 linesdownloads.py - merit/
example_data/ , Python, 312 linesdummy_ipd_data.py - merit/
example_data/ , Python, 1,126 linesdummy_summary_data.py - merit/
example_data/ , Shell, 1 lineexample_datasets/ test_file_pools/ cohort_c/ chr_hint/ create_bgen.sh - merit/
example_data/ , Python, 1,213 linesexamples.py - merit/
gwas/ , Python, 9 lines__init__.py - merit/
gwas/ , Python, 2,897 linesdataset.py - merit/
gwas/ , Python, 3,684 linesutils.py - merit/
ipd/ , Python, 80 lines__init__.py - merit/
ipd/ , Python, 753 linesadmission.py - merit/
ipd/ , Python, 1,766 linescommon.py - merit/
ipd/ , Python, 2,127 linescorr_factory.py - merit/
ipd/ , Python, 4,026 linescorr_matrix.py - merit/
ipd/ , Python, 1,313 linesdosage_cache.py - merit/
ipd/ , Python, 2,856 linesfiles.py - merit/
ipd/ , Python, 346 linesmanifest.py - merit/
ipd/ , Python, 3,606 linesregions.py - merit/
ld/ , Python, 31 lines__init__.py - merit/
ld/ , Python, 940 linesclumping.py - merit/
ld/ , Python, 1,022 linesfactory_clumping.py - merit/
ld/ , Python, 781 lineshandlers.py - merit/
ld/ , Python, 81 linesvariants.py - merit/
misc/ , Python, 9 lines__init__.py - merit/
misc/ , Python, 550 linessojo.py - merit/
misc/ , Python, 1,247 linessteiger_filtering.py - merit/
mr/ , Python, 47 lines__init__.py - merit/
mr/ , Python, 282 linesbase.py - merit/
mr/ , Python, 229 linesconmix.py - merit/
mr/ , Python, 226 linesegger.py - merit/
mr/ , Python, 207 lines, 1 matchivw.py - merit/
mr/ , Python, 321 linesmedian.py - merit/
mr/ , Python, 47 linesmendelian_randomization. py - merit/
mr/ , Python, 549 linesmode.py - merit/
mr/ , Python, 432 linesmvmr.py - merit/
mr/ , Python, 127 linesselection.py - merit/
mr/ , Python, 71 linesutils.py - merit/
mr/ , Python, 118 lineswald.py - merit/
pipelines/ , Python, 306 lines__init__.py - merit/
pipelines/ , Python, 482 lines_outcome_loader.py - merit/
pipelines/ , Python, 1,066 linesassembly.py - merit/
pipelines/ , Python, 10 linescli/ __init__.py - merit/
pipelines/ , Python, 3,107 linescli/ _args.py - merit/
pipelines/ , Python, 353 linescli/ _log.py - merit/
pipelines/ , Python, 331 linescli/ run_coloc.py - merit/
pipelines/ , Python, 562 linescli/ run_extract.py - merit/
pipelines/ , Python, 741 linescli/ run_mr.py - merit/
pipelines/ , Python, 754 linescli/ run_mr_hpc.py - merit/
pipelines/ , Python, 779 linescli/ run_mr_parallel.py - merit/
pipelines/ , Python, 1,383 linescoloc.py - merit/
pipelines/ , Python, 61 linescolumns.py - merit/
pipelines/ , Python, 370 linesconstants.py - merit/
pipelines/ , Python, 1 linedata/ __init__.py - merit/
pipelines/ , Python, 3,557 linesextract.py - merit/
pipelines/ , Python, 1,043 linesfeatures.py - merit/
pipelines/ , Python, 1,690 linesfiltering.py - merit/
pipelines/ , Python, 320 linesformatting.py - merit/
pipelines/ , Python, 27 lineshpc/ __init__.py - merit/
pipelines/ , Python, 401 lineshpc/ jaf.py - merit/
pipelines/ , Python, 1,171 lineshpc/ planner.py - merit/
pipelines/ , Python, 512 lineshpc/ task_script.py - merit/
pipelines/ , Python, 766 linesio.py - merit/
pipelines/ , Python, 587 linesld_reference.py - merit/
pipelines/ , Python, 1,586 linesmaf.py - merit/
pipelines/ , Python, 25 linesmethods/ __init__.py - merit/
pipelines/ , Python, 267 linesmethods/ _config.py - merit/
pipelines/ , Python, 190 linesmethods/ _protocol.py - merit/
pipelines/ , Python, 272 linesmethods/ _registry.py - merit/
pipelines/ , Python, 170 linesmethods/ _support.py - merit/
pipelines/ , Python, 198 linesmethods/ egger.py - merit/
pipelines/ , Python, 180 linesmethods/ ivw.py - merit/
pipelines/ , Python, 189 linesmethods/ median.py - merit/
pipelines/ , Python, 223 linesmethods/ mode.py - merit/
pipelines/ , Python, 148 linesmethods/ mvmr.py - merit/
pipelines/ , Python, 111 linesmethods/ wald.py - merit/
pipelines/ , Python, 4,608 linesmr.py - merit/
pipelines/ , Python, 40 linesparallel/ __init__.py - merit/
pipelines/ , Python, 581 linesparallel/ executor.py - merit/
pipelines/ , Python, 164 linesparallel/ options.py - merit/
pipelines/ , Python, 58 linesparallel/ publish.py - merit/
pipelines/ , Python, 566 linesparallel/ runtime.py - merit/
pipelines/ , Python, 331 linesparallel/ worker.py - merit/
pipelines/ , Python, 613 linesparameters.py - merit/
pipelines/ , Python, 167 linesregion_plan.py - merit/
pipelines/ , Python, 57 linesresult.py - merit/
pipelines/ , Python, 81 linessample_size.py - merit/
pipelines/ , Python, 283 linesvalidation.py - merit/
stats.py , Python, 49 lines - resources/
analysis/ , Python, 14 linesfn_12_missingness_failur e_prediction/ __init__.py - resources/
analysis/ , Python, 52 linesfn_12_missingness_failur e_prediction/ _paths.py - resources/
analysis/ , Python, 154 linesfn_12_missingness_failur e_prediction/ definedness.py - resources/
analysis/ , Python, 243 linesfn_12_missingness_failur e_prediction/ features.py - resources/
analysis/ , Python, 369 linesfn_12_missingness_failur e_prediction/ index_build.py - resources/
analysis/ , Python, 102 linesfn_12_missingness_failur e_prediction/ index_parse.py - resources/
analysis/ , Python, 126 linesfn_12_missingness_failur e_prediction/ io_util.py - resources/
analysis/ , Python, 71 linesfn_12_missingness_failur e_prediction/ metrics.py - resources/
analysis/ , Python, 388 linesfn_12_missingness_failur e_prediction/ model.py - resources/
analysis/ , Python, 433 linesfn_12_missingness_failur e_prediction/ pipeline.py - resources/
analysis/ , Python, 294 linesfn_12_missingness_failur e_prediction/ plot_roc.py - resources/
analysis/ , Python, 201 linesfn_12_missingness_failur e_prediction/ predictors.py - resources/
analysis/ , Python, 212 linesfn_12_missingness_failur e_prediction/ project.py - resources/
analysis/ , Python, 532 linesfn_12_missingness_failur e_prediction/ report.py - resources/
analysis/ , Python, 171 linesfn_12_missingness_failur e_prediction/ roc.py - resources/
analysis/ , Python, 714 linesfn_12_missingness_failur e_prediction/ run.py - resources/
analysis/ , Python, 55 linesfn_12_missingness_failur e_prediction/ sampling.py - resources/
analysis/ , Python, 639 linesfn_12_missingness_failur e_prediction/ simulate_production.py - resources/
analysis/ , Python, 617 linesfn_12_missingness_failur e_prediction/ subsets.py - resources/
analysis/ , Python, 126 linesfn_12_missingness_failur e_prediction/ windows_io.py - resources/
analysis/ , Python, 14 linesfn_51_missingness_vs_d/ __init__.py - resources/
analysis/ , Python, 53 linesfn_51_missingness_vs_d/ _paths.py - resources/
analysis/ , Python, 339 linesfn_51_missingness_vs_d/ calibrate.py - resources/
analysis/ , Python, 137 linesfn_51_missingness_vs_d/ io_util.py - resources/
analysis/ , Python, 233 linesfn_51_missingness_vs_d/ metrics.py - resources/
analysis/ , Python, 121 linesfn_51_missingness_vs_d/ pairwise.py - resources/
analysis/ , Python, 149 linesfn_51_missingness_vs_d/ production_parity.py - resources/
analysis/ , Python, 55 linesfn_51_missingness_vs_d/ sampling.py - resources/
analysis/ , Python, 292 linesfn_51_missingness_vs_d/ select_windows.py - resources/
analysis/ , Python, 367 linesfn_51_missingness_vs_d/ summarize.py - resources/
analysis/ , Python, 108 linesfn_51_missingness_vs_d/ thresholds.py - resources/
analysis/ , Python, 389 linesfn_51_missingness_vs_d/ vcf_io.py - resources/
analysis/ , Python, 383 linesfn_51_missingness_vs_d/ verify_pairwise.py - resources/
analysis/ , Python, 453 linesfn_51_missingness_vs_d/ windows.py - resources/
bin/ , Python, 652 linesdemo_corr_factory.py - resources/
bin/ , Python, 462 linesdemo_dosage_cache.py - resources/
bin/ , Shell, 470 linesdownload_1kgp.sh - resources/
bin/ , Python, 114 linesgen_corr_golden.py - resources/
bin/ , Python, 215 linesprobe_query_edge_truncat ion.py - resources/
ci_cd/ , Shell, 30 linesbefore_script.sh - resources/
ci_cd/ , Shell, 11 linesbuild.sh - resources/
ci_cd/ , Shell, 27 linespages.sh - resources/
ci_cd/ , Shell, 68 linesrun_docker.sh - resources/
ci_cd/ , Shell, 32 linesrun_pages.sh - resources/
conda/ , Shell, 4 linesconda-build.sh - resources/
examples/ , Jupyter, 178 lines1kgp_01_setup.ipynb - resources/
examples/ , Jupyter, 220 lines1kgp_02_variant_queries. ipynb - resources/
examples/ , Jupyter, 185 lines1kgp_03_dosage_caches.ip ynb - resources/
examples/ , Jupyter, 351 lines1kgp_04_ld_analysis.ipyn b - resources/
examples/ , Jupyter, 238 linescoloc.ipynb - resources/
examples/ , Jupyter, 316 linesdrug_target_mr.ipynb - resources/
examples/ , Jupyter, 346 linesld_clumping.ipynb - resources/
examples/ , Jupyter, 237 linesmendelian_randomization. ipynb - resources/
misc/ , Python, 159 lines2026-sort/ merit/ ipd/ cohort.py - resources/
misc/ , Python, 1,889 lines2026-sort/ merit/ ipd/ common.py - resources/
misc/ , Python, 184 lines2026-sort/ merit/ ipd/ engine.py - resources/
misc/ , Python, 386 lines2026-sort/ merit/ ipd/ file_container.py - resources/
misc/ , Python, 276 lines2026-sort/ merit/ ipd/ file_pool.py - resources/
misc/ , Python, 1,613 lines2026-sort/ merit/ ipd/ readers.py - resources/
misc/ , Python, 70 lines2026-sort/ merit/ ipd/ workers.py - resources/
misc/ , R, 181 linescode_other_packages/ MendelianRandomization/ mr_conmix-methods.R - resources/
misc/ , R, 57 linescode_other_packages/ MendelianRandomization/ mr_mode.R - resources/
misc/ , R, 38 linescode_other_packages/ TwoSampleMR/ mr/ mr_wald_ratio.R - resources/
misc/ , R, 391 linescode_other_packages/ coloc/ coloc_baseline.R - resources/
misc/ , R, 114 linescode_other_packages/ coloc/ coloc_sensitivity.R - resources/
misc/ , R, 66 linescode_other_packages/ mvmr_wes_spiller/ code/ phenocov_mvmr.R - resources/
misc/ , R, 52 linescode_other_packages/ mvmr_wes_spiller/ code/ pl2_mvmr_one.R - resources/
misc/ , R, 39 linescode_other_packages/ mvmr_wes_spiller/ code/ pl2_mvmr_two.R - resources/
misc/ , R, 73 linescode_other_packages/ mvmr_wes_spiller/ code/ pl_mvmr.R - resources/
misc/ , R, 218 linescode_other_packages/ mvmr_wes_spiller/ code/ qhet_mvmr.R - resources/
misc/ , R, 112 linescode_other_packages/ mvmr_wes_spiller/ code/ qmin.R - resources/
misc/ , R, 91 linescode_other_packages/ mvmr_wes_spiller/ code/ snpcov_mvmr.R - resources/
misc/ , R, 213 linescode_other_packages/ mvmr_wes_spiller/ code/ strength_mvmr.R - resources/
misc/ , R, 220 linescode_other_packages/ mvmr_wes_spiller/ code/ strhet_mvmr.R - resources/
misc/ , R, 174 linescode_other_packages/ mvmr_wes_spiller/ docs/ MVMR.rmd - resources/
misc/ , Python, 68 linescode_other_packages/ sojo/ load_data_and_arguments. py - resources/
misc/ , R, 311 linescode_other_packages/ sojo/ sojo_newest.R - resources/
misc/ , R, 288 linescode_other_packages/ sojo/ sojo_pytest.r - resources/
misc/ , Python, 180 linesdead_code/ analysis_tools_dead_code .py - resources/
misc/ , Python, 2,042 linesdead_code/ cohorts.py - resources/
misc/ , Python, 884 linesdead_code/ config.py - resources/
misc/ , Python, 154 linesdead_code/ config2.py - resources/
misc/ , Python, 774 linesdead_code/ dataset_dead_code.py - resources/
misc/ , Python, 208 linesdead_code/ emerald.py - resources/
misc/ , Python, 1 linedead_code/ ensembl/ __init__.py - resources/
misc/ , Python, 500 linesdead_code/ ensembl/ constants.py - resources/
misc/ , Python, 167 linesdead_code/ ensembl/ coords.py - resources/
misc/ , Python, 129 linesdead_code/ ensembl/ ensembl_errors.py - resources/
misc/ , Python, 1,183 linesdead_code/ ensembl/ rest.py - resources/
misc/ , Python, 2,577 linesdead_code/ ensembl/ variants.py - resources/
misc/ , Jupyter, 101 linesdead_code/ examples/ allele_norm.ipynb - resources/
misc/ , Jupyter, 150 linesdead_code/ examples/ chrpos_parsing.ipynb - resources/
misc/ , Jupyter, 60 linesdead_code/ examples/ cohort_init_from_genomic _config.ipynb - resources/
misc/ , Jupyter, 374 linesdead_code/ examples/ drug_target_mr_example.i pynb - resources/
misc/ , Jupyter, 125 linesdead_code/ examples/ liftover.ipynb - resources/
misc/ , Jupyter, 454 linesdead_code/ examples/ run_simple_prs.ipynb - resources/
misc/ , Jupyter, 303 linesdead_code/ examples/ run_simple_prs_example.i pynb - resources/
misc/ , Jupyter, 278 linesdead_code/ examples/ sojo.ipynb - resources/
misc/ , Python, 317 linesdead_code/ interface.py - resources/
misc/ , Python, 1 linedead_code/ ipd/ __init__.py - resources/
misc/ , Python, 1 linedead_code/ ipd/ agnostic_handler.py - resources/
misc/ , Python, 1 linedead_code/ ipd/ bed_handler.py - resources/
misc/ , Python, 2,624 linesdead_code/ ipd/ bgen_handler.py - resources/
misc/ , Python, 982 linesdead_code/ ipd/ ipd_tools.py - resources/
misc/ , Python, 1 linedead_code/ ipd/ vcf_handler.py - resources/
misc/ , Python, 1 linedead_code/ mrbase/ __init__.py - resources/
misc/ , Python, 1,377 linesdead_code/ mrbase/ mrbase_api.py - resources/
misc/ , Python, 1 linedead_code/ pools/ __init__.py - resources/
misc/ , Python, 43 linesdead_code/ pools/ bgen_pool.py - resources/
misc/ , Python, 1 linedead_code/ pools/ cohort_pool.py - resources/
misc/ , Python, 415 linesdead_code/ pools/ file_pool.py - resources/
misc/ , Python, 88 linesdead_code/ pools/ vcf_pool.py - resources/
misc/ , Python, 1 linedead_code/ prs/ __init__.py - resources/
misc/ , Python, 1,257 linesdead_code/ prs/ scores.py - resources/
misc/ , Python, 1,263 linesdead_code/ prs/ simple_prs.py - resources/
misc/ , Python, 465 linesdead_code/ prs/ validation.py - resources/
misc/ , Python, 509 linesdead_code/ prs/ validation_calibration.p y - resources/
misc/ , Python, 439 linesdead_code/ qctool_dead_code.py - resources/
misc/ , Python, 1,531 linesdead_code/ refpops.py - resources/
misc/ , Python, 787 linesdead_code/ utils/ liftover.py - resources/
misc/ , Python, 140 linesdead_code/ utils/ plink.py - resources/
misc/ , Python, 72 linesdead_code/ utils/ vcf.py - resources/
misc/ , Python, 178 linesdead_code/ visual/ visualisation.py - resources/
misc/ , Python, 1 linedead_code/ wrappers/ __init__.py - resources/
misc/ , Python, 534 linesdead_code/ wrappers/ bgenix.py - resources/
misc/ , Python, 523 linesdead_code/ wrappers/ plink.py - resources/
misc/ , Python, 745 linesdead_code/ wrappers/ qctool.py - resources/
misc/ , Python, 30 linesdead_code/ wrappers/ wrapper_errors.py - resources/
misc/ , Jupyter, 382 linesexamples/ COLOC_example.ipynb - resources/
misc/ , Jupyter, 347 linesexamples/ LD_tests.ipynb - resources/
misc/ , Jupyter, 227 linesexamples/ MR_example.ipynb - resources/
misc/ , Jupyter, 202 linesexamples/ allele_conversion_testin g.ipynb - resources/
misc/ , Jupyter, 153 linesexamples/ coloc.ipynb - resources/
misc/ , Python, 7 linesexamples/ example_path.py - resources/
misc/ , Jupyter, 121 linesexamples/ mrbase_query.ipynb - resources/
misc/ , Python, 113 linesexamples/ mrbase_query.py - resources/
misc/ , Jupyter, 211 lineshelper_scripts/ LD_clumping_tests.ipynb - resources/
misc/ , Shell, 71 lineshelper_scripts/ build_druggable_list.sh - resources/
misc/ , Jupyter, 18 lineshelper_scripts/ dup_test.ipynb - resources/
misc/ , Jupyter, 187 lineshelper_scripts/ emerald.ipynb - resources/
misc/ , Python, 817 lineshelper_scripts/ ensembl.py - resources/
misc/ , Shell, 52 lineshelper_scripts/ ensembl_regulatory_activ ity.sh - resources/
misc/ , Shell, 56 lineshelper_scripts/ ensembl_regulatory_peaks .sh - resources/
misc/ , Shell, 108 lineshelper_scripts/ extract_eqtl.sh - resources/
misc/ , Shell, 83 lineshelper_scripts/ extract_somalogic.sh - resources/
misc/ , Shell, 37 lineshelper_scripts/ fetch_gene_coords.sh - resources/
misc/ , Shell, 35 lineshelper_scripts/ fetch_on_asid.sh - resources/
misc/ , Shell, 22 lineshelper_scripts/ fetch_targets.sh - resources/
misc/ , Shell, 69 lineshelper_scripts/ generate_plink_testset.s h - resources/
misc/ , Shell, 57 lineshelper_scripts/ get_chembl_indication.sh - resources/
misc/ , Shell, 10 lineshelper_scripts/ make_physical_regions.sh - resources/
misc/ , Shell, 42 lineshelper_scripts/ run_simple_prs.sh - resources/
misc/ , Python, 18 lineshelper_scripts/ test_bgenix.py - resources/
misc/ , Python, 66 lineshelper_scripts/ test_corr_plink.py - resources/
misc/ , Python, 28 lineshelper_scripts/ test_emerald.py - resources/
misc/ , Python, 99 lineshelper_scripts/ test_filter.py - resources/
misc/ , Python, 34 lineshelper_scripts/ test_ld_matrix.py - resources/
misc/ , Python, 66 lineshelper_scripts/ test_park_amd.py - resources/
misc/ , Python, 49 lineshelper_scripts/ test_plink.py - resources/
misc/ , Python, 96 lineshelper_scripts/ test_simple_prs.py - resources/
misc/ , Python, 18 lineshelper_scripts/ test_vcf_write.py - resources/
misc/ , Python, 49 linesissue_debug/ 51_issue_debug.py - resources/
misc/ , Python, 90 linesissue_debug/ 53_issue_debug.py - resources/
misc/ , Python, 62 linesissue_debug/ 57_issue_debug.py - resources/
misc/ , Jupyter, 120 linesissue_debug/ 63_issue_debug.ipynb - resources/
misc/ , Python, 45 linesissue_debug/ 63_issue_debug.py - resources/
misc/ , Python, 1 linemeta_analysis/ __init__.py - resources/
misc/ , Python, 159 linesmeta_analysis/ meta_analysis.py - resources/
misc/ , R, 35 linesmv_mr_examples/ multivariatemr_test.R - resources/
misc/ , Python, 1,131 linesold_tests/ test_bgen_handler.py - resources/
misc/ , Python, 234 linesold_tests/ test_ipd_tools.py - resources/
misc/ , Python, 129 linesold_tests/ tests/ example_data/ test_examples.py - resources/
misc/ , Python, 720 linesold_tests/ tests/ mrbase/ test_mrbase_api.py - resources/
misc/ , Python, 580 linesold_tests/ tests/ prs/ conftest.py - resources/
misc/ , Python, 303 linesold_tests/ tests/ prs/ test_prs.py - resources/
misc/ , Python, 429 linesold_tests/ tests/ prs/ test_scores.py - resources/
misc/ , Python, 335 linesold_tests/ tests/ prs/ test_simple_prs.py - resources/
misc/ , Python, 118 linesold_tests/ tests/ prs/ test_validation.py - resources/
misc/ , Python, 229 linesold_tests/ tests/ prs/ test_validation_calibrat ion.py - resources/
misc/ , Python, 554 linesold_tests/ tests/ test_errors.py - resources/
misc/ , Python, 1,051 linesold_tests/ tests/ test_refpops.py - resources/
misc/ , Python, 96 linesold_tests/ tests/ utils/ test_pyplink.py - resources/
misc/ , Python, 375 linesold_tests/ tests/ utils/ test_steiger_filtering.p y - resources/
misc/ , Python, 19 linesold_tests/ tests/ utils/ test_vcf.py - resources/
misc/ , Python, 101 linesold_tests/ tests/ wrappers/ test_bgenix.py - resources/
misc/ , Python, 150 linesold_tests/ tests/ wrappers/ test_plink.py - resources/
misc/ , Python, 485 linesold_tests/ tests/ wrappers/ test_qctool.py - resources/
misc/ , Python, 32 linestests_to_check/ genotypes/ conftest.py - resources/
misc/ , Python, 32 linestests_to_check/ genotypes/ test_bgen_pool.py - resources/
misc/ , Python, 151 linestests_to_check/ genotypes/ test_file_pool.py - resources/
misc/ , Python, 22 linestests_to_check/ genotypes/ test_vcf_pool.py - resources/
misc/ , Python, 1,032 linestests_to_check/ test_checks.py - resources/
misc/ , Python, 180 linestests_to_check/ test_config.py - resources/
misc/ , Python, 99 linestests_to_check/ test_data/ test_mr_data.py - resources/
misc/ , Python, 298 linestests_to_check/ test_dummy_data.py - resources/
misc/ , Python, 161 linestests_to_check/ test_ld_calcs.py - resources/
misc/ , Python, 111 linestests_to_check/ test_ld_clumping.py - resources/
misc/ , Python, 55 linestests_to_check/ test_mrbase.py - resources/
misc/ , Python, 169 linestests_to_check/ test_plink_utils.py - resources/
misc/ , Python, 1,535 linestests_to_check/ test_read_gwas.py - resources/
misc/ , Python, 997 linestests_to_check/ test_refpops.py - resources/
misc/ , Python, 175 linestests_to_check/ test_vcf_utils.py - resources/
misc/ , Python, 124 linestests_to_check/ utils/ test_liftover.py - resources/
misc/ , R, 451 lineswork_in_progress/ bma_mr/ summary_mvMR_BF.R - resources/
misc/ , R, 503 lineswork_in_progress/ bma_mr/ summary_mvMR_SSS.R - resources/
misc/ , Python, 134 lineswork_in_progress/ mr_sim/ mr_twosample_sim.py - resources/
misc/ , R, 1,586 lineswork_in_progress/ mvmr_lasso/ mvMR_glmnet.R - resources/
misc/ , R, 97 lineswork_in_progress/ pleio_lasso/ mr_lasso-methods.R - resources/
misc/ , R, 106 lineswork_in_progress/ pleio_lasso/ mr_mvlasso-methods.R - resources/
misc/ , Python, 111 lineswork_in_progress/ selection/ selection_pointestimate. py - resources/
misc/ , Python, 132 lineswork_in_progress/ selection/ selection_precision.py - resources/
qa-tests/ , Shell, 148 linesfn-51-missingness-vs-d/ run.sh - resources/
qa-tests/ , Shell, 297 linesfn-57-external-ld-corpus / run.sh - resources/
test_data/ , Python, 5,003 linesld_clumping/ build_clumping_corpus.py - resources/
test_data/ , R, 78 linesld_clumping/ fixtures/ naive_clump_r_input.R - resources/
test_data/ , R, 517 linesld_clumping/ naive_clump.R - resources/
test_data/ , Python, 440 linesld_clumping/ naive_clump.py - resources/
test_data/ , Python, 1,850 linesld_clumping/ preflight.py - resources/
test_data/ , R, 29 linesld_clumping/ roundtrip_numeric.R - resources/
test_data/ , Python, 149 linesld_clumping/ run_f1_parity.py - resources/
test_data/ , Python, 533 linesld_clumping/ test_agreement_gate.py - resources/
test_data/ , Python, 178 linesld_clumping/ test_artifact_emission.p y - resources/
test_data/ , Python, 85 linesld_clumping/ test_auxiliary_fixtures. py - resources/
test_data/ , Python, 146 linesld_clumping/ test_build_clumping_corp us.py - resources/
test_data/ , R, 264 linesld_clumping/ test_naive_clump.R - resources/
test_data/ , Python, 143 linesld_clumping/ test_naive_clump.py - resources/
test_data/ , Python, 514 linesld_clumping/ test_parameter_sweeps.py - resources/
test_data/ , Python, 170 linesld_clumping/ test_sumstats_generation .py - resources/
test_data/ , Python, 1,448 linesld_clumping/ validate_clumping_corpus .py - resources/
test_data/ , Python, 3,927 linesld_reference/ build_reference_corpus.p y - resources/
test_data/ , R, 307 linesld_reference/ reference_corr.R - resources/
test_data/ , Python, 1,319 linesld_reference/ validate_reference_corpu s.py - tests/
benchmarks/ , Python, 570 linesr_code/ test_mendelian_randomiza tion.py - tests/
benchmarks/ , Python, 322 linesr_code/ test_merit_vs_r.py - tests/
benchmarks/ , Python, 275 linesr_code/ test_steiger_filtering.p y - tests/
coloc/ , Python, 1,461 linestest_coloc_core.py - tests/
coloc/ , Python, 102 linestest_coloc_sensitivity.p y - tests/
conftest.py , Python, 60 lines - tests/
example_data/ , Python, 127 linestest_dummy_ipd_data.py - tests/
example_data/ , Python, 553 linestest_dummy_summary_data. py - tests/
example_data/ , Python, 115 linestest_plink_write.py - tests/
gwas/ , Python, 3,377 linestest_dataset.py - tests/
gwas/ , Python, 2,534 linestest_utils.py - tests/
ipd/ , Python, 188 linesadmission_fixtures.py - tests/
ipd/ , Python, 788 linesconftest.py - tests/
ipd/ , Python, 671 linescorr_test_oracle.py - tests/
ipd/ , Python, 216 lineshelpers.py - tests/
ipd/ , Python, 340 linesproperty_oracle.py - tests/
ipd/ , Python, 1,093 linestest_admission.py - tests/
ipd/ , Python, 205 linestest_benchmark_corrcoef. py - tests/
ipd/ , Python, 259 linestest_cohort_init.py - tests/
ipd/ , Python, 124 linestest_cohort_reader.py - tests/
ipd/ , Python, 511 linestest_cohort_sample_subse tting.py - tests/
ipd/ , Python, 475 linestest_common.py - tests/
ipd/ , Python, 668 linestest_common_sidecar.py - tests/
ipd/ , Python, 3,036 linestest_corr_factory.py - tests/
ipd/ , Python, 5,486 linestest_corr_matrix.py - tests/
ipd/ , Python, 1,301 linestest_corr_matrix_cleanup _primitives.py - tests/
ipd/ , Python, 149 linestest_corr_matrix_default _goldens.py - tests/
ipd/ , Python, 255 linestest_corr_matrix_e2e.py - tests/
ipd/ , Python, 478 linestest_corr_test_oracle.py - tests/
ipd/ , Python, 674 linestest_corrcoef.py - tests/
ipd/ , Python, 1,223 linestest_dosage_cache.py - tests/
ipd/ , Python, 932 linestest_external_ld_referen ce_e2e.py - tests/
ipd/ , Python, 1,456 linestest_external_ld_referen ce_generation.py - tests/
ipd/ , Python, 2,525 linestest_files.py - tests/
ipd/ , Python, 676 linestest_manifest.py - tests/
ipd/ , Python, 852 linestest_metamorphic_admissi on.py - tests/
ipd/ , Python, 462 linestest_metamorphic_common. py - tests/
ipd/ , Python, 417 linestest_missingness_calibra tion.py - tests/
ipd/ , Python, 1,190 linestest_missingness_failure _prediction.py - tests/
ipd/ , Python, 103 linestest_package_exports.py - tests/
ipd/ , Python, 1,064 linestest_property_admission. py - tests/
ipd/ , Python, 794 linestest_property_common.py - tests/
ipd/ , Python, 497 linestest_property_manifest.p y - tests/
ipd/ , Python, 542 linestest_property_regions.py - tests/
ipd/ , Python, 180 linestest_random_samples.py - tests/
ipd/ , Python, 3,843 linestest_regions.py - tests/
ipd/ , Python, 266 linestest_tools.py - tests/
ld/ , Python, 1 line__init__.py - tests/
ld/ , Python, 552 linesconftest.py - tests/
ld/ , Python, 295 linesexternal_clumping_suppor t.py - tests/
ld/ , Python, 832 linestest_clumping.py - tests/
ld/ , Python, 414 linestest_cross_validation.py - tests/
ld/ , Python, 262 linestest_external_clumping_a greement_e2e.py - tests/
ld/ , Python, 314 linestest_external_clumping_a uxiliary_e2e.py - tests/
ld/ , Python, 1,137 linestest_external_clumping_e 2e.py - tests/
ld/ , Python, 417 linestest_external_clumping_f actories.py - tests/
ld/ , Python, 373 linestest_external_clumping_f ile_routes.py - tests/
ld/ , Python, 180 linestest_external_clumping_i dentity.py - tests/
ld/ , Python, 194 linestest_external_clumping_s election_e2e.py - tests/
ld/ , Python, 1,152 linestest_external_clumping_s emantics.py - tests/
ld/ , Python, 302 linestest_external_clumping_s umstats_e2e.py - tests/
ld/ , Python, 1,264 linestest_external_clumping_t hresholds.py - tests/
ld/ , Python, 623 linestest_external_clumping_u nresolved.py - tests/
ld/ , Python, 408 linestest_factory_clumping.py - tests/
ld/ , Python, 323 linestest_file_clumping.py - tests/
ld/ , Python, 216 linestest_handlers.py - tests/
ld/ , Python, 85 linestest_helpers.py - tests/
ld/ , Python, 337 linestest_multiprocess.py - tests/
ld/ , Python, 84 linestest_variants.py - tests/
misc/ , Python, 83 linestest_sojo.py - tests/
misc/ , Python, 70 linestest_steiger_filtering.p y - tests/
mr/ , Python, 1 line__init__.py - tests/
mr/ , Python, 24 linestest_api_exports.py - tests/
mr/ , Python, 126 linestest_egger.py - tests/
mr/ , Python, 166 linestest_ivw.py - tests/
mr/ , Python, 1,066 linestest_mendelian_randomiza tion.py - tests/
mr/ , Python, 143 linestest_mvmr.py - tests/
pipelines/ , Python, 1 line__init__.py - tests/
pipelines/ , Python, 1 linecli/ __init__.py - tests/
pipelines/ , Python, 49 linescli/ conftest.py - tests/
pipelines/ , Python, 921 linescli/ test_args.py - tests/
pipelines/ , Python, 106 linescli/ test_command_skeletons.p y - tests/
pipelines/ , Python, 178 linescli/ test_log.py - tests/
pipelines/ , Python, 102 linescli/ test_output_hash_reprodu cibility.py - tests/
pipelines/ , Python, 675 linescli/ test_run_coloc.py - tests/
pipelines/ , Python, 921 linescli/ test_run_extract.py - tests/
pipelines/ , Python, 2,011 linescli/ test_run_mr.py - tests/
pipelines/ , Python, 680 linescli/ test_run_mr_hpc.py - tests/
pipelines/ , Python, 809 linescli/ test_run_mr_parallel.py - tests/
pipelines/ , Python, 1,781 linescli/ test_shared_cli.py - tests/
pipelines/ , Python, 1,594 linescli/ test_subprocess_contract s.py - tests/
pipelines/ , Python, 676 linescli/ test_typed_ld_reference_ e2e.py - tests/
pipelines/ , Python, 639 linesconftest.py - tests/
pipelines/ , Python, 1 linehpc/ __init__.py - tests/
pipelines/ , Python, 412 lineshpc/ test_jaf.py - tests/
pipelines/ , Python, 589 lineshpc/ test_planner.py - tests/
pipelines/ , Python, 426 lineshpc/ test_task_script.py - tests/
pipelines/ , Python, 1 linemethods/ __init__.py - tests/
pipelines/ , Python, 343 linesmethods/ test_config.py - tests/
pipelines/ , Python, 370 linesmethods/ test_egger.py - tests/
pipelines/ , Python, 331 linesmethods/ test_ivw.py - tests/
pipelines/ , Python, 391 linesmethods/ test_median.py - tests/
pipelines/ , Python, 61 linesmethods/ test_min_snps.py - tests/
pipelines/ , Python, 541 linesmethods/ test_mode.py - tests/
pipelines/ , Python, 613 linesmethods/ test_registry.py - tests/
pipelines/ , Python, 280 linesmethods/ test_support.py - tests/
pipelines/ , Python, 206 linesmethods/ test_wald.py - tests/
pipelines/ , Python, 177 linesmr_results_report_contra ct.py - tests/
pipelines/ , Python, 1 lineparallel/ __init__.py - tests/
pipelines/ , Python, 964 linesparallel/ test_executor.py - tests/
pipelines/ , Python, 132 linesparallel/ test_publish.py - tests/
pipelines/ , Python, 477 linesparallel/ test_runtime.py - tests/
pipelines/ , Python, 551 linesparallel/ test_worker.py - tests/
pipelines/ , Python, 816 linestest_apply_maf.py - tests/
pipelines/ , Python, 569 linestest_assembly.py - tests/
pipelines/ , Python, 1,579 linestest_coloc.py - tests/
pipelines/ , Python, 1,281 linestest_coloc_e2e.py - tests/
pipelines/ , Python, 58 linestest_columns.py - tests/
pipelines/ , Python, 197 linestest_constants.py - tests/
pipelines/ , Python, 1,785 linestest_extract.py - tests/
pipelines/ , Python, 649 linestest_features.py - tests/
pipelines/ , Python, 2,344 linestest_filtering.py - tests/
pipelines/ , Python, 228 linestest_fn17_e2e.py - tests/
pipelines/ , Python, 278 linestest_formatting.py - tests/
pipelines/ , Python, 68 linestest_human_chromosomes.p y - tests/
pipelines/ , Python, 685 linestest_io.py - tests/
pipelines/ , Python, 149 linestest_io_integration.py - tests/
pipelines/ , Python, 1,238 linestest_ld_reference.py - tests/
pipelines/ , Python, 289 linestest_maf_crosscheck.py - tests/
pipelines/ , Python, 525 linestest_maf_source.py - tests/
pipelines/ , Python, 5,018 linestest_mr.py - tests/
pipelines/ , Python, 2,111 linestest_mr_e2e.py - tests/
pipelines/ , Python, 89 linestest_mr_results_contract .py - tests/
pipelines/ , Python, 1,106 linestest_mr_results_report.p y - tests/
pipelines/ , Python, 543 linestest_outcome_loader.py - tests/
pipelines/ , Python, 263 linestest_outcome_loader_isol ation.py - tests/
pipelines/ , Python, 226 linestest_package_exports.py - tests/
pipelines/ , Python, 599 linestest_parameters.py - tests/
pipelines/ , Python, 62 linestest_result.py - tests/
test_ci_image_config.py , Python, 89 lines - tests/
test_constants.py , Python, 52 lines - tests/
test_errors_formatting.p , Python, 80 linesy - tests/
test_imports.py , Python, 62 lines - tests/
test_pyproject_metadata. , Python, 127 linespy - tests/
test_stats.py , Python, 36 lines - tests/
test_version_config.py , Python, 35 lines - LICENCE.txt, License, 674 lines
- README.md, Text, 54 lines
cfinan/bio-misc
4060630a617911862a403c065d2cd0f5bebe0cc3, 15 February 2024Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
104 files
- _version.py, Python, 1 line
- bio_misc/
__init__.py , Python, 4 lines - bio_misc/
annotation/ , Python, 1 line__init__.py - bio_misc/
annotation/ , Python, 428 linesboolean_matrix.py - bio_misc/
annotation/ , Python, 540 linesgenecards.py - bio_misc/
annotation/ , Python, 239 linesvcf_id_mapper.py - bio_misc/
common/ , Python, 1 line__init__.py - bio_misc/
common/ , Python, 394 linesarguments.py - bio_misc/
common/ , Python, 35 linesconstants.py - bio_misc/
common/ , Python, 248 lineslog.py - bio_misc/
common/ , Python, 717 linesutils.py - bio_misc/
deprecated.py , Python, 8 lines - bio_misc/
drug_lookups/ , Python, 1 line__init__.py - bio_misc/
drug_lookups/ , Python, 67 linesbnf_queries.py - bio_misc/
drug_lookups/ , Python, 403 lineschembl_queries.py - bio_misc/
drug_lookups/ , Python, 2 linescommon.py - bio_misc/
drug_lookups/ , Python, 595 linessearch_effects.py - bio_misc/
drug_lookups/ , Python, 489 linestarget_effects.py - bio_misc/
drug_lookups/ , Python, 766 linestarget_summary.py - bio_misc/
drug_lookups/ , Python, 441 lines, 1 matchtherapeutic_direction.py - bio_misc/
example_data/ , Python, 1 line__init__.py - bio_misc/
example_data/ , Python, 1 lineexample_datasets/ __init__.py - bio_misc/
example_data/ , Python, 1 lineexample_datasets/ clump_data/ __init__.py - bio_misc/
example_data/ , Python, 1 lineexample_datasets/ ipd_data/ __init__.py - bio_misc/
example_data/ , Python, 1 lineexample_datasets/ overlaps/ __init__.py - bio_misc/
example_data/ , Python, 1 lineexample_datasets/ zarr_test/ __init__.py - bio_misc/
example_data/ , Python, 1,296 linesexamples.py - bio_misc/
format/ , Python, 1 line__init__.py - bio_misc/
format/ , Python, 231 lineschr_pos.py - bio_misc/
format/ , Python, 268 linescommon.py - bio_misc/
format/ , Python, 246 linesuni_id.py - bio_misc/
id_map/ , Python, 1 line__init__.py - bio_misc/
id_map/ , Python, 287 linesensembl_gene.py - bio_misc/
id_map/ , Python, 261 linesensembl_human_gene_unipr ot.py - bio_misc/
id_map/ , Python, 237 linesensembl_human_uniprot_ge ne.py - bio_misc/
id_map/ , Python, 334 linesensembl_uniprot.py - bio_misc/
id_map/ , Python, 232 linessomalogic_seq_ids.py - bio_misc/
ipd/ , Python, 1 line__init__.py - bio_misc/
ipd/ , Python, 1,171 linescommon.py - bio_misc/
ipd/ , Python, 600 linescorr_matrix.py - bio_misc/
ipd/ , Python, 345 linescorr_matrix_single.py - bio_misc/
ipd/ , Python, 294 linescorr_matrix_zarr.py - bio_misc/
ipd/ , Python, 341 linesdosage_zarr.py - bio_misc/
ipd/ , Python, 311 linesextract_allele_counts.py - bio_misc/
ipd/ , Python, 108 linesindex_bgen.py - bio_misc/
ipd/ , Python, 1,478 linesld_clump.py - bio_misc/
ipd/ , Python, 159 linesld_factory.py - bio_misc/
ipd/ , Python, 1,078 linesld_handlers.py - bio_misc/
ipd/ , Python, 1,613 linesreaders.py - bio_misc/
liftover/ , Python, 1 line__init__.py - bio_misc/
liftover/ , Python, 270 linescrossmap.py - bio_misc/
liftover/ , Python, 357 linesquick_lift.py - bio_misc/
overlaps/ , Python, 1 line__init__.py - bio_misc/
overlaps/ , Python, 293 linesgenes_within.py - bio_misc/
overlaps/ , Python, 240 linesnearest_genes.py - bio_misc/
overlaps/ , Python, 693 linessite_overlap.py - bio_misc/
overlaps/ , Python, 432 linestrans_qtl_link.py - bio_misc/
overlaps/ , Python, 307 linestrans_within.py - bio_misc/
pubmed/ , Python, 1 line__init__.py - bio_misc/
pubmed/ , Python, 1,258 linespubmed_util.py - bio_misc/
pubmed/ , Python, 694 linesquery_genes.py - bio_misc/
reactome/ , Python, 1 line__init__.py - bio_misc/
reactome/ , Python, 247 linescommon.py - bio_misc/
reactome/ , Python, 226 linesdrug_path_dist.py - bio_misc/
reactome/ , Python, 132 linespairwise_path_dist.py - bio_misc/
reactome/ , Python, 174 linesset_path_dist.py - bio_misc/
transforms/ , Python, 1 line__init__.py - bio_misc/
transforms/ , Python, 544 lineseffect_flipper.py - bio_misc/
transforms/ , Python, 272 lineslog_pvalue.py - bio_misc/
umls/ , Python, 1 line__init__.py - bio_misc/
umls/ , Python, 318 linesmesh_mapping.py - docs/
source/ , Python, 103 linesconf.py - resources/
bin/ , Shell, 22 linescat-bgen-wrap.sh - resources/
bin/ , Shell, 280 linesliftover-bgen.sh - resources/
ci_cd/ , Shell, 26 linesbefore_script.sh - resources/
ci_cd/ , Shell, 32 linespages.sh - resources/
ci_cd/ , Shell, 57 linesrun_docker.sh - resources/
ci_cd/ , Shell, 29 linesrun_pages.sh - resources/
conda/ , Shell, 3 linesbuild/ py310/ build.sh - resources/
conda/ , Shell, 5 linesbuild/ py310/ run_test.sh - resources/
conda/ , Shell, 3 linesbuild/ py38/ build.sh - resources/
conda/ , Shell, 5 linesbuild/ py38/ run_test.sh - resources/
conda/ , Shell, 3 linesbuild/ py39/ build.sh - resources/
conda/ , Shell, 5 linesbuild/ py39/ run_test.sh - resources/
examples/ , Jupyter, 196 linesbgen_reader_test.ipynb - resources/
tasks/ , Shell, 43 linescat-bgen.task.sh - resources/
tasks/ , Shell, 76 linesliftover-bgen.task.sh - resources/
test_scripts/ , Python, 202 linesmake_clump_test.py - resources/
test_scripts/ , Python, 110 linestest_bgen_read.py - resources/
test_scripts/ , Python, 504 linestest_corr_dp.py - setup.py, Python, 220 lines
- tests/
test_imports.py , Python, 65 lines - tests/
test_ipd/ , Python, 246 linesconftest.py - tests/
test_ipd/ , Python, 82 linestest_common.py - tests/
test_ipd/ , Python, 374 linestest_corr_matrix.py - tests/
test_ipd/ , Python, 113 linestest_corrcoef.py - tests/
test_ipd/ , Python, 235 linestest_ld_clump.py - tests/
test_ipd/ , Python, 210 linestest_np_commands.py - tests/
test_overlaps/ , Python, 65 linesconftest.py - tests/
test_overlaps/ , Python, 67 linestest_genes_within.py - tests/
test_overlaps/ , Python, 90 linestest_site_overlaps.py - tests/
test_overlaps/ , Python, 68 linestest_trans_within.py - LICENSE.txt, License, 674 lines
- README.md, Text, 79 lines
cfinan/gwas-norm
04b05b9eeb43a67266e4350a0e27f75ec0f2e0a3, 26 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
134 files
- docs/
source/ , Python, 86 linesconf.py - gwas_norm/
__init__.py , Python, 1 line - gwas_norm/
_version.py , Python, 1 line - gwas_norm/
columns.py , Python, 1,083 lines - gwas_norm/
common.py , Python, 1,167 lines - gwas_norm/
constants.py , Python, 318 lines - gwas_norm/
controllers.py , Python, 2,368 lines - gwas_norm/
crossmap.py , Python, 647 lines - gwas_norm/
errors.py , Python, 28 lines - gwas_norm/
example_data/ , Python, 1 line__init__.py - gwas_norm/
example_data/ , Python, 817 linesexamples.py - gwas_norm/
files.py , Python, 1,314 lines - gwas_norm/
genome_assemblies.py , Python, 196 lines - gwas_norm/
gwas_norm_run.py , Python, 1,093 lines - gwas_norm/
handlers.py , Python, 1,619 lines - gwas_norm/
hpc/ , Python, 1 line__init__.py - gwas_norm/
hpc/ , Python, 712 lineshpc_gwas_norm.py - gwas_norm/
metadata/ , Python, 1 line__init__.py - gwas_norm/
metadata/ , Python, 1,086 linesanalysis.py - gwas_norm/
metadata/ , Python, 550 linesbase.py - gwas_norm/
metadata/ , Python, 1,397 linescohort.py - gwas_norm/
metadata/ , Python, 251 linescolumn.py - gwas_norm/
metadata/ , Python, 964 linesconvert.py - gwas_norm/
metadata/ , Python, 1,633 linesfile.py - gwas_norm/
metadata/ , Python, 503 linesgwas_data.py - gwas_norm/
metadata/ , Python, 247 linesinfo.py - gwas_norm/
metadata/ , Python, 955 linesphenotype.py - gwas_norm/
metadata/ , Python, 1,505 linesstudy.py - gwas_norm/
metadata/ , Python, 623 linestest.py - gwas_norm/
normalise.py , Python, 3,150 lines, 1 match - gwas_norm/
parsers.py , Python, 890 lines - gwas_norm/
stats.py , Python, 182 lines - gwas_norm/
tests.py , Python, 4,684 lines - resources/
bin/ , Shell, 99 lineslist-norm-files.sh - resources/
ci_cd/ , Shell, 23 linesbefore_script.sh - resources/
ci_cd/ , Shell, 32 linespages.sh - resources/
ci_cd/ , Shell, 8 linespypi-upload.sh - resources/
ci_cd/ , Shell, 57 linesrun_docker.sh - resources/
ci_cd/ , Shell, 32 linesrun_pages.sh - resources/
examples/ , Jupyter, 291 linesmetadata_api.ipynb - resources/
misc/ , Python, 108 linescohort.py - resources/
misc/ , Shell, 115 linesconvert_xmls.sh - resources/
misc/ , Python, 372 linesextract-xml-info.py - resources/
misc/ , Python, 729 linesmake-test-data.py - resources/
misc/ , Python, 485 linesold/ old_code/ 1kg.py - resources/
misc/ , Python, 274 linesold/ old_code/ alfa.py - resources/
misc/ , Python, 847 linesold/ old_code/ arrays.py - resources/
misc/ , Python, 342 linesold/ old_code/ batch_proc_old.py - resources/
misc/ , Python, 119 linesold/ old_code/ define_match.py - resources/
misc/ , Jupyter, 69 linesold/ old_code/ liftover_test.ipynb - resources/
misc/ , Python, 283 linesold/ old_code/ mapper_file_code.py - resources/
misc/ , Python, 3,020 linesold/ old_code/ mappers.py - resources/
misc/ , Python, 1,248 linesold/ old_code/ mapping_files.py - resources/
misc/ , Python, 388 linesold/ old_code/ mapping_utils.py - resources/
misc/ , Shell, 10 linesold/ old_code/ merge_1kg_freq.sh - resources/
misc/ , Shell, 29 linesold/ old_code/ process_alfa.sh - resources/
misc/ , Shell, 33 linesold/ old_code/ process_ensembl_dbsnp.sh - resources/
misc/ , Shell, 17 linesold/ old_code/ process_ukbb.sh - resources/
misc/ , Python, 388 linesold/ old_code/ test_config_obj.py - resources/
misc/ , Python, 250 linesold/ old_code/ test_config_parser.py - resources/
misc/ , Python, 12 linesold/ old_code/ vcf.py - resources/
misc/ , Python, 49 linesold/ tests/ utils/ conftest.py - resources/
misc/ , Python, 221 linesold/ tests/ utils/ test_upgrade_norm.py - resources/
misc/ , Python, 604 linesold/ v0.2-code/ config.py - resources/
misc/ , Python, 1 lineold/ v0.2-code/ database/ __init__.py - resources/
misc/ , Python, 594 linesold/ v0.2-code/ database/ add_to_db.py - resources/
misc/ , Python, 56 linesold/ v0.2-code/ external.py - resources/
misc/ , Python, 935 linesold/ v0.2-code/ gwas_norm_premap.py - resources/
misc/ , Python, 248 linesold/ v0.2-code/ log.py - resources/
misc/ , Python, 1 lineold/ v0.2-code/ map_file/ __init__.py - resources/
misc/ , Python, 1 lineold/ v0.2-code/ map_file/ dbsnp/ __init__.py - resources/
misc/ , Python, 391 linesold/ v0.2-code/ map_file/ dbsnp/ clinvar.py - resources/
misc/ , Python, 813 linesold/ v0.2-code/ map_file/ dbsnp/ download.py - resources/
misc/ , Python, 2,450 linesold/ v0.2-code/ map_file/ dbsnp/ parse.py - resources/
misc/ , Python, 1,316 linesold/ v0.2-code/ old/ config_obj.py - resources/
misc/ , Python, 2,105 linesold/ v0.2-code/ old/ config_parser.py - resources/
misc/ , Python, 975 linesold/ v0.2-code/ old/ crossmap.py - resources/
misc/ , Python, 273 linesold/ v0.2-code/ old/ crossmap_api.py - resources/
misc/ , Python, 281 linesold/ v0.2-code/ old/ gtex_xml_writer.py - resources/
misc/ , Python, 746 linesold/ v0.2-code/ old/ readers.py - resources/
misc/ , Python, 1 lineold/ v0.2-code/ old/ writers.py - resources/
misc/ , Python, 1,436 linesold/ v0.2-code/ processors.py - resources/
misc/ , Python, 676 linesold/ v0.2-code/ sort.py - resources/
misc/ , Python, 1 lineold/ v0.2-code/ utils/ __init__.py - resources/
misc/ , Python, 334 linesold/ v0.2-code/ utils/ extract_pvalue.py - resources/
misc/ , Python, 748 linesold/ v0.2-code/ utils/ genome_chunks.py - resources/
misc/ , Python, 133 linesold/ v0.2-code/ utils/ make_extract_pvalue_job_ array.py - resources/
misc/ , Python, 343 linesold/ v0.2-code/ utils/ quick_lift.py - resources/
misc/ , Python, 1,658 linesold/ v0.2-code/ utils/ upgrade_norm.py - resources/
misc/ , Python, 1,745 linesold/ v0.2-code/ utils/ upgrade_top_hits.py - resources/
misc/ , Python, 32 linestest-norm-structure.py - resources/
misc/ , Python, 16 linestest_decompress.py - resources/
misc/ , Python, 44 linestest_variants/ create_test_mapper.py - resources/
misc/ , Shell, 24 linestest_variants/ ensembl-cmd-line.sh - resources/
misc/ , Shell, 41 linestest_variants/ other-impute.sh - resources/
misc/ , Shell, 40 linestest_variants/ scan-cmd-line.sh - resources/
misc/ , Shell, 41 linestest_variants/ scan-suhre.sh - resources/
misc/ , Shell, 42 linestest_variants/ tabix-cmd-line.sh - resources/
misc/ , Python, 33 linestest_variants/ test_mapper.py - resources/
misc/ , Python, 528 linesupdate-xml-cohorts.py - resources/
misc/ , Python, 789 linesupdate-xmls.py - resources/
tasks/ , Shell, 196 linesgwas-norm.task.sh - tests/
__init__.py , Python, 1 line - tests/
conftest.py , Python, 491 lines - tests/
data/ , Shell, 5 linesdummy_master/ rename.sh - tests/
data/ , Shell, 102 linesref_genome_tests/ old/ ref_vcf_setup.sh - tests/
data/ , Shell, 177 linesref_genome_tests/ ref_vcf_setup.sh - tests/
metadata/ , Python, 193 linesconftest.py - tests/
metadata/ , Python, 417 linestest_analysis.py - tests/
metadata/ , Python, 53 linestest_base.py - tests/
metadata/ , Python, 539 linestest_cohort.py - tests/
metadata/ , Python, 112 linestest_column.py - tests/
metadata/ , Python, 111 linestest_convert.py - tests/
metadata/ , Python, 453 linestest_file.py - tests/
metadata/ , Python, 97 linestest_gwas_data.py - tests/
metadata/ , Python, 86 linestest_metadata_info.py - tests/
metadata/ , Python, 159 linestest_phenotypes.py - tests/
metadata/ , Python, 826 linestest_study.py - tests/
metadata/ , Python, 562 linestest_test.py - tests/
test_columns.py , Python, 342 lines - tests/
test_controllers.py , Python, 600 lines - tests/
test_crossmap.py , Python, 140 lines - tests/
test_gwas_norm.py , Python, 504 lines - tests/
test_info_columns_withou , Python, 185 linest_info_element.py - tests/
test_info_handlers.py , Python, 220 lines - tests/
test_merge_log_index_err , Python, 176 linesor.py - tests/
test_normalise.py , Python, 1,190 lines - tests/
test_parsers.py , Python, 655 lines - tests/
test_safe_writerow.py , Python, 160 lines - tests/
test_stats.py , Python, 122 lines - tests/
test_target_genome_assem , Python, 110 linesbly_normalise.py - tests/
test_test_data.py , Python, 431 lines - LICENCE.txt, License, 21 lines
- README.md, Text, 35 lines
chembl/tractability_pipeline_v2
84f53f49664455f89a802f0e2caf7927f7ff06c9, 9 February 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
16 files
- ot_tractability_pipeline
_v2/ , Python, 12 linesSpaCy_NER_PROTAC_model_p ackaged/ en_NER_PROTAC-0.2.5/ en_NER_PROTAC/ __init__.py - ot_tractability_pipeline
_v2/ , Python, 64 linesSpaCy_NER_PROTAC_model_p ackaged/ en_NER_PROTAC-0.2.5/ setup.py - ot_tractability_pipeline
_v2/ , Python, 1 line__init__.py - ot_tractability_pipeline
_v2/ , Python, 1 linebin/ __init__.py - ot_tractability_pipeline
_v2/ , Python, 825 linesbin/ run_pipeline.py - ot_tractability_pipeline
_v2/ , Python, 759 linesbuckets_ab.py - ot_tractability_pipeline
_v2/ , Python, 346 linesbuckets_othercl.py - ot_tractability_pipeline
_v2/ , Python, 1,318 linesbuckets_protac.py - ot_tractability_pipeline
_v2/ , Python, 753 linesbuckets_sm.py - ot_tractability_pipeline
_v2/ , Python, 49 linesqueries_ab.py - ot_tractability_pipeline
_v2/ , Python, 52 linesqueries_othercl.py - ot_tractability_pipeline
_v2/ , Python, 106 linesqueries_protac.py - ot_tractability_pipeline
_v2/ , Python, 373 linesqueries_sm.py - setup.py, Python, 26 lines
- LICENSE, License, 21 lines
- README.md, Text, 256 lines
Code availability
Cis-Mendelian randomisation and colocalisation analyses were performed using purpose-built Python (v3.9) packages (Mendelian Randomisation for Identifying Targets [MeRIT], v0.3.3a0; https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 722 scripts, each with its path and the digest of its content;
- 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- ukbiobank.ac.uk/
enable-your-research/ , at UK Biobank; found in “Data availability”apply-for-access - zenodo:19135140, at Zenodo; found in DataCite
- zenodo:19135141, at Zenodo; found in DataCite
Data availability
GWAS summary statistics for MD are available from the Psychiatric Genomics Consortium at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 3 keywords, 7 MeSH terms, 4 funders, 101 references.
Cite
This paper
ter Kuile, A. R., Finan, C., Chopade, S., van Vugt, M., Hukerikar, N., Barral, S., Stringaris, A., Schmidt, A. F., Kuchenbaecker, K., & Pingault, J.-B. (2026). An integrative mendelian randomisation and drug mechanism framework for target prioritisation and therapeutic repurposing in major depression. Translational psychiatry, 16(1), 392. https://
BibTeX
@article{terkuile2026int
author = {ter Kuile, Abigail R. and Finan, Chris and Chopade, Sandesh and van Vugt, Marion and Hukerikar, Nikita and Barral, Serena and Stringaris, Argyris and Schmidt, Amand F. and Kuchenbaecker, Karoline and Pingault, Jean-Baptiste},
title = {{An integrative mendelian randomisation and drug mechanism framework for target prioritisation and therapeutic repurposing in major depression}},
journal = {Translational psychiatry},
year = {2026},
month = jun,
volume = {16},
number = {1},
pages = {392},
publisher = {Nature Publishing Group},
issn = {2158-3188},
doi = {10.1038/
url = {https://
pmid = {42230543},
pmcid = {PMC13443190}
}
RIS
TY - JOUR
AU - ter Kuile, Abigail R.
AU - Finan, Chris
AU - Chopade, Sandesh
AU - van Vugt, Marion
AU - Hukerikar, Nikita
AU - Barral, Serena
AU - Stringaris, Argyris
AU - Schmidt, Amand F.
AU - Kuchenbaecker, Karoline
AU - Pingault, Jean-Baptiste
TI - An integrative mendelian randomisation and drug mechanism framework for target prioritisation and therapeutic repurposing in major depression
T2 - Translational psychiatry
J2 - Transl Psychiatry
PY - 2026
DA - 2026/
VL - 16
IS - 1
SP - 392
SN - 2158-3188
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "An integrative mendelian randomisation and drug mechanism framework for target prioritisation and therapeutic repurposing in major depression",
"container-title": "Translational psychiatry",
"author": [
{
"family": "ter Kuile",
"given": "Abigail R."
},
{
"family": "Finan",
"given": "Chris"
},
{
"family": "Chopade",
"given": "Sandesh"
},
{
"family": "van Vugt",
"given": "Marion"
},
{
"family": "Hukerikar",
"given": "Nikita"
},
{
"family": "Barral",
"given": "Serena"
},
{
"family": "Stringaris",
"given": "Argyris"
},
{
"family": "Schmidt",
"given": "Amand F."
},
{
"family": "Kuchenbaecker",
"given": "Karoline"
},
{
"family": "Pingault",
"given": "Jean-Baptiste"
}
],
"container-title-short":
"volume": "16",
"issue": "1",
"page": "392",
"DOI": "10.1038/
"PMID": "42230543",
"PMCID": "PMC13443190",
"ISSN": "2158-3188",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
2
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41562-026-02476-7 [code]
- Genome-wide meta-analysis of quantitatively measured generalized anxiety symptoms in individuals of European ancestry.Journal: Nature human behaviourIn common: BCFtools, psych, data.table, clinical / translational, genetics / omics, 5 references, author Abigail R. ter Kuile
- [2] doi:10.1038/s41467-026-76676-0 [code]
- Determinants of functional burden pleiotropy and gene dosage responses across human traits.Journal: Nature communicationsIn common: Numba, data.table, statsmodels, 5 other tools, ukbiobank.ac.uk/enable-your-research/apply-for-access, genetics / omics, 1 reference
- [3] doi:10.1038/s41592-026-03211-w [code]
- Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.Journal: Nature methodsIn common: pysam, Biopython, Numba, 7 other tools, genetics / omics
- [4] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: BCFtools, Biopython, Numba, 7 other tools
- [5] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: pysam, rpy2, Numba, 6 other tools, genetics / omics
- [6] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: pysam, Biopython, Numba, 6 other tools, genetics / omics
- [7] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: BCFtools, pysam, statsmodels, 5 other tools, genetics / omics, 1 reference
- [8] doi:10.1038/s43587-026-01106-1 [code]
- Repurposing drugs for the prevention of vascular dementia using evidence from drug target Mendelian randomization.Journal: Nature agingIn common: genetics / omics, 7 references
- [9] doi:10.1038/s41562-026-02486-5 [code]
- Genome-wide association studies of infant and toddler temperament in European and multi-ancestry populations.Journal: Nature human behaviourIn common: Biopython, psych, data.table, 5 other tools, genetics / omics, 1 reference
- [10] doi:10.7554/elife.108109 [code]
- Multimodal MRI marker of cognition explains the association between cognition and mental health in the UK Biobank.Journal: eLifeIn common: psych, data.table, seaborn, 4 other tools, ukbiobank.ac.uk/enable-your-research/apply-for-access
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 4 repositories of the authors' code, each at its verified commit and with its license, 722 scripts, and 5 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:34b11ae760e8d0c3…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
