OSCR

FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Results › FM-GPT accurately detects true causal genes across multiple traits while controlling false positives ↔ R/gen_gene_by_snp.R, lines 2–61 · score 0.66 · cis SNPs, phenotypic heritability, gene expression, variance, LD, binary
  2. [2] § Results › FM-GPT accurately detects true causal genes across multiple traits while controlling false positives ↔ example/fmgpt_simulation_code.R, the whole file · a weak match · score 0.53 · cis SNPs, ROC, AUC, FDR, power, variable

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 179 lines · 6.1 KB · no license · 1 match

  1. #'@param n_qtl The sample size of the reference qtl dataset
  2. #'@param n_gwas The sample size of the GWAS dataset
  3. #'@param snps_per_gene The total number of cis SNPs for each gene
  4. #'@param block_sizes How many genes are in each block
  5. #'@param rho_between The correlation between cis SNPs of different genes
  6. #'@param rho_within The correlation of cis SNPs for the same gene
  7. #'@param response_types A vector containing the number of continuous, binary and count variables (in that order)
  8. #'@param r The parameter controlling the negative binomial distribution for count variables
  9. #'@param k The total number of factors
  10. #'@param maf The minor allele frequency
  11. #'@param sigma2 (what is this again?)
  12. #'@param choice (what is this again?)
  13. #'@param sig_rho (what is this again?)
  14. #'@param binflate Inflation factor for the beta coefficients for the covariates. Used to control SNP heritability on gene expression
  15. #'@param vinfale Inflation factor for outcome variance. Used to control heritability
  16. #'@param pi0 The underlying probability of selecting any given group of genes
  17. #'@param pi1 The underlying probability of selecting any gene in a group
  18. #'@param pi2 The underlying probability of selecting any factor for a gene
  19. #'@param expression_heritability Used to control how much of an effect gene expression has on phenotype heritability
  20. gen_twas_sets <- function(n_qtl, n_gwas, n_genes, snps_per_gene, block_sizes, rho_between, rho_within, response_types = c(1,1,1), r = NULL, k, maf, sigma2 = NULL, choice = "Z", sig_rho = 0.66, binflate = 1, vinflate = 1, pi0, pi1, pi2, expression_heritability = 0.1) {
  21. response_vec <- rep(c("continuous", "binary", "count"), response_types)
  22. p <- sum(response_types)
  23. r2 <- numeric(p)
  24. if(is.null(r)) {
  25. r2[which(response_vec == "count")] <- 50
  26. } else {
  27. r2[which(response_vec == "count")] <- r
  28. }
  29. r <- r2
  30. n <- n_qtl + n_gwas
  31. xp <- sum(block_sizes) #total number of covariates
  32. LD_corr_matrix <- matrix(rho_between, xp, xp)
  33. row_start <- 1
  34. col_start <- 1
  35. for (size in block_sizes) {
  36. block_matrix <- matrix(rho_within, nrow = size, ncol = size)
  37. LD_corr_matrix[row_start:(row_start + size - 1), col_start:(col_start + size - 1)] <- block_matrix
  38. row_start <- row_start + size
  39. col_start <- col_start + size
  40. }
  41. diag(LD_corr_matrix) <-1
  42. mu <- numeric(xp)
  43. Z <- mvrnorm(n = n, mu, Sigma = LD_corr_matrix)
  44. prob_g0 <- dbinom(0, 2, maf)
  45. prob_g1 <- dbinom(1, 2, maf)
  46. prob_g2 <- dbinom(2, 2, maf)
  47. c1 <- numeric(xp)
  48. c2 <- numeric(xp)
  49. for (j in 1:xp) {
  50. c1[j] <- qnorm(prob_g0, mean = 0, sd = LD_corr_matrix[j,j])
  51. c2[j] <- qnorm(prob_g2, mean = 0, sd = LD_corr_matrix[j,j], lower.tail = FALSE)
  52. }
  53. G <- Z
  54. intervals <- c(-Inf, c1[1], c2[1], Inf)
  55. replacement_values <- c(0, 1, 2)
  56. for (i in seq_along(intervals)) {
  57. lower_bound <- intervals[i]
  58. upper_bound <- intervals[i + 1]
  59. G[(Z > lower_bound) & (Z <= upper_bound)] <- replacement_values[i]
  60. }
  61. G_qtl <- G[1:n_qtl,]
  62. G_gwas <- G[seq(n_qtl+1,n,1),]
  63. gene_mat <- matrix(nrow = n_qtl, ncol = n_genes)
  64. snp_sets <- seq(1, n_genes*snps_per_gene, by = snps_per_gene)
  65. snp_coef_mat <- matrix(0, nrow = snps_per_gene * n_genes, ncol = n_genes)
  66. counter <- 1
  67. for(i in snp_sets) {
  68. snp_coefs <- snp_coef_mat[seq(i,i+(snps_per_gene-1),1),counter] <- rnorm(n = snps_per_gene, sd = 1)
  69. eta <- as.numeric(G_qtl[,seq(i,i+(snps_per_gene-1),1)] %*% snp_coefs)
  70. gene_mat[,counter] <- rnorm(n_qtl, mean = eta, sd = sqrt(1-expression_heritability))
  71. counter <- counter + 1
  72. }
  73. group <- rep(1:length(block_sizes), times = block_sizes)
  74. Sig <- diag(k)
  75. Sig[Sig == 0] <- sig_rho
  76. if(choice == "G") {
  77. alpha <- rbinom(n = length(block_sizes), size = 1, prob = pi0)
  78. gamma <- rbinom(n = xp, size = 1, prob = pi1)
  79. omega <- rbinom(n = xp * k, size = 1, prob = pi2)
  80. amat <- matrix(rep(rep(alpha, times = table(group)), k), nrow = xp, ncol = k, byrow = FALSE)
  81. gmat <- matrix(rep(gamma, k), nrow = xp, k, byrow = FALSE)
  82. omat <- matrix(omega, ncol = k, byrow = FALSE)
  83. b <- MASS::mvrnorm(n = xp, mu = rep(0, k), Sigma = binflate * vinflate * Sig)
  84. beta_true <- amat * gmat * omat * b
  85. } else {
  86. gamma <- rbinom(n = n_genes, size = 1, prob = pi1)
  87. omega <- rbinom(n = n_genes * k, size = 1, prob = pi2)
  88. gmat <- matrix(rep(gamma, k), nrow = n_genes, k, byrow = FALSE)
  89. omat <- matrix(omega, ncol = k, byrow = FALSE)
  90. b <- MASS::mvrnorm(n = n_genes, mu = rep(0, k), Sigma = binflate * vinflate * Sig)
  91. beta_true <- gmat * omat * b
  92. }
  93. E_gwas <- G_gwas %*% snp_coef_mat
  94. A <- as.numeric(rowSums(abs(beta_true)) != 0)
  95. if(choice == "G") {
  96. XB <- G_gwas %*% beta_true
  97. } else {
  98. XB <- E_gwas %*% beta_true
  99. }
  100. Y_latent <- XB + MASS::mvrnorm(n = n_gwas, mu = rep(0,k), Sigma = vinflate * Sig)
  101. Lambda.true<- matrix(0,p,k)
  102. if(k == p) {
  103. for(h in 1:k){
  104. Lambda.true[sample(1:p,p),h] <- rnorm(k,0,1)
  105. }
  106. } else {
  107. for(h in 1:k){
  108. Lambda.true[sample(1:p,p)[1:(2*k-(h-1))],h] <- rnorm(2*k-(h-1),0,1)
  109. }
  110. }
  111. if(is.null(sigma2)) {
  112. Ucov <- diag(1/rgamma(p,shape=1,scale=0.25))
  113. } else {
  114. Ucov <- sigma2 * diag(p)
  115. }
  116. U <- MASS::mvrnorm(n = n_gwas, rep(0, p), Ucov)
  117. theta <- Y_latent %*% t(Lambda.true) + U
  118. Y_observed <- matrix(0, n_gwas, p)
  119. sigmoid <- function(x) {
  120. 1/(1 + exp(-x))
  121. }
  122. for(q in 1:length(response_vec)) {
  123. if(response_vec[q] == "continuous") {
  124. Y_observed[,q] <- matrix(Y_latent %*% Lambda.true[q,], n_gwas, 1) + U[,q] + rnorm(n_gwas, 0, 1)
  125. } else if(response_vec[q] == "binary") {
  126. prob <- sigmoid(matrix(Y_latent %*% Lambda.true[q, ], n_gwas, 1) + U[, q])
  127. Y_observed[,q] <- rbinom(n_gwas, 1, prob)
  128. } else {
  129. prob <- sigmoid(matrix(Y_latent %*% Lambda.true[q, ], n_gwas, 1) + U[, q])
  130. prob[prob > 0.9999] <- 0.9999
  131. Y_observed[,q] <- stats::rnbinom(n_gwas, size = r[q], prob = 1-prob)
  132. }
  133. }
  134. return(list(G_qtl = G_qtl, G_gwas = G_gwas, E_qtl = gene_mat, SNP_coef = snp_coef_mat, Yl = Y_latent, Yo = Y_observed, Lambda = Lambda.true, Sigma = vinflate * Sig,
  135. group = group, A = A, beta = beta_true, r = r, responses = response_vec, U = U, Ucov = diag(Ucov), LD = LD_corr_matrix))
  136. }

gen_gene_by_snp.R at commit 97b0994, no license · at the source

Overview

Authors: Travis Canida1,2, Zhenyao Ye3,4, Shao-Hsuan Wang5, Hsin-Hsiung Huang6, Yezhi Pan2,3, Menglu Liang1, Shuo Chen3,4, Tianzhou Ma1,3
  1. Department of Epidemiology and Biostatistics, School of Public Health, University of Maryland, College Park, Maryland, United States of America
  2. Department of Mathematics, University of Maryland, College Park, Maryland, United States of America
  3. Maryland Psychiatric Research Center, Department of Psychiatry, School of Medicine, University of Maryland, Baltimore, Maryland, United States of America
  4. Department of Epidemiology and Public Health, School of Medicine, University of Maryland, Baltimore, Maryland, United States of America
  5. Graduate Institute of Statistics, National Central University, Taoyuan City, Taiwan
  6. Department of Statistics and Data Science, University of Central Florida, Orlando, Florida, United States of America
Journal: PLoS genetics, volume 22, issue 9, article e1012126
Dates: received 6 April 2026; accepted 24 August 2026; published online 8 September 2026
Type: Research article · Language: English
License: none stated
Identifiers: DOI 10.1371/journal.pgen.1012126 · PMID 42709879 · PMCID PMC13581217 · OpenAlex W7153858399
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), structural MRI / diffusion (modality), human (organism), cellular / molecular (subfield)
Methods: Statistics, Machine learning, fMRI & imaging, Preprocessing
MeSH: Chromosome Mapping*, Genome-Wide Association Study*, Phenomics*, Transcriptome*, Bayes Theorem, Gene Expression Profiling, Humans, Linkage Disequilibrium, Phenotype, Quantitative Trait Loci (* major topic)
Topic: Genetic Associations and Epidemiology (Genetics, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NIDA NIH HHS (K01 DA059603)
Citations: not cited yet (Europe PMC); 64 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (none stated) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

tacanida/fm-gpt

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 97b099405913e854c86dd403df7b0dd0c8c5442e, 10 July 2026
Languages: R (5), C++ (4)
Size: 23 files, 9 scripts
Software Heritage: not archived
Found in: “Data Availability”
Holds: README, environment (DESCRIPTION), documentation
Not found: license file, CITATION.cff, tests, continuous integration
Tools: pROC (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
10 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 9 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Code and data availability statement

The paper has a code and data availability statement. Its license (none stated) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1371/journal.pgen.1012126.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 8 authors, 10 MeSH terms, 1 funder, 62 references.

Cite

This paper

Canida, T., Ye, Z., Wang, S.-H., Huang, H.-H., Pan, Y., Liang, M., Chen, S., & Ma, T. (2026). FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies. PLoS genetics, 22(9), e1012126. https://doi.org/10.1371/journal.pgen.1012126

BibTeX

@article{canida2026fm,
author = {Canida, Travis and Ye, Zhenyao and Wang, Shao-Hsuan and Huang, Hsin-Hsiung and Pan, Yezhi and Liang, Menglu and Chen, Shuo and Ma, Tianzhou},
title = {{FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies}},
journal = {PLoS genetics},
year = {2026},
month = sep,
volume = {22},
number = {9},
pages = {e1012126},
publisher = {PLOS},
issn = {1553-7390},
doi = {10.1371/journal.pgen.1012126},
url = {https://doi.org/10.1371/journal.pgen.1012126},
pmid = {42709879},
pmcid = {PMC13581217}
}

RIS

TY - JOUR
AU - Canida, Travis
AU - Ye, Zhenyao
AU - Wang, Shao-Hsuan
AU - Huang, Hsin-Hsiung
AU - Pan, Yezhi
AU - Liang, Menglu
AU - Chen, Shuo
AU - Ma, Tianzhou
TI - FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies
T2 - PLoS genetics
J2 - PLoS Genet
PY - 2026
DA - 2026/09/08
VL - 22
IS - 9
SP - e1012126
SN - 1553-7390
PB - PLOS
DO - 10.1371/journal.pgen.1012126
UR - https://doi.org/10.1371/journal.pgen.1012126
LA - en
ER -

CSL-JSON

{
"id": "10.1371/journal.pgen.1012126",
"type": "article-journal",
"title": "FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies",
"container-title": "PLoS genetics",
"author": [
{
"family": "Canida",
"given": "Travis"
},
{
"family": "Ye",
"given": "Zhenyao"
},
{
"family": "Wang",
"given": "Shao-Hsuan"
},
{
"family": "Huang",
"given": "Hsin-Hsiung"
},
{
"family": "Pan",
"given": "Yezhi"
},
{
"family": "Liang",
"given": "Menglu"
},
{
"family": "Chen",
"given": "Shuo"
},
{
"family": "Ma",
"given": "Tianzhou"
}
],
"container-title-short": "PLoS Genet",
"volume": "22",
"issue": "9",
"page": "e1012126",
"DOI": "10.1371/journal.pgen.1012126",
"PMID": "42709879",
"PMCID": "PMC13581217",
"ISSN": "1553-7390",
"publisher": "PLOS",
"URL": "https://doi.org/10.1371/journal.pgen.1012126",
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
8
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10030-4 [code]
Cell-type-aware transcriptome-wide association studies identify 91 independent risk genes for Alzheimer's disease dementia.
Journal: Communications biology
In common: genetics / omics, cellular / molecular, 6 references
[2] doi:10.1038/s41467-026-75193-4 [code]
Multi-ancestry gene expression models amplify transcriptome-wide association study discovery and validation.
Journal: Nature communications
In common: genetics / omics, cellular / molecular, 6 references
[3] doi:10.3390/biomedicines14081677 [code]
Exploratory Genome and Transcriptome-Wide Association Analyses of Addiction-Related Phenotypes in a Twin Cohort.
Journal: Biomedicines
In common: genetics / omics, cellular / molecular, 5 references
[4] doi:10.1186/s12967-026-08266-z [code]
Single-cell multi-omic integration analysis prioritizes druggable genes and reveals cell-type-specific causal effects in glioblastomagenesis.
Journal: Journal of translational medicine
In common: genetics / omics, cellular / molecular, 5 references
[5] doi:10.1371/journal.pcbi.1014422 [code]
Deciphering cell type-specific causal genetic effects on brain imaging-derived phenotypes and disorders with single-cell Mendelian randomization.
Journal: PLoS computational biology
In common: genetics / omics, cellular / molecular, 6 references
[6] doi:10.21203/rs.3.rs-10380518/v1 [code]
Shared Genetic Architecture of Premenstrual Disorder and Postpartum Depression: Registry-Based and Genetic Evidence
Journal: Research Square (preprint)
In common: genetics / omics, cellular / molecular, 4 references
[7] doi:10.3390/genes17070813
Multilayer Genomic Characterization of a Shared Genetic Factor Linking Depression-Related Liability and Reduced Physical Function.
Journal: Genes
In common: genetics / omics, cellular / molecular, 4 references
[8] doi:10.1038/s41588-026-02646-3 [code]
Co-expression-based models improve eQTL predictions for transcriptome-wide association studies and highlight new schizophrenia-associated genes.
Journal: Nature genetics
In common: genetics / omics, cellular / molecular, 4 references
[9] doi:10.34133/csbj.0108 [code]
Latent Factor Modeling Reveals Unexpected Spatial Heterogeneity in Human Alzheimer's Disease Brain Transcriptomes.
Journal: Computational and structural biotechnology journal
In common: pROC, genetics / omics, cellular / molecular, 2 references
[10] doi:10.1038/s41467-026-72139-8 [code]
Early and late RNA eQTL are driven by different genetic mechanisms.
Journal: Nature communications
In common: genetics / omics, cellular / molecular, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.