OSCR

scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches · 2 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Methods › Real data applications › Prediction model training and evaluation ↔ realdata_code/ROSMAP/GE_generate_CTS.R, lines 1–61 · score 0.68 · excitatory neuron, inhibitory neuron, Onek1k, OPC, astrocyte, oligodendrocyte
  2. [2] § Methods › Data sets › Genotype and scRNA-seq data › The OneK1K study ↔ realdata_code/Onek1k/CTS_GE_generate.R, the whole file · a weak match · score 0.66 · MonoNC, MonoC, BMem, DC, donors, CD4
  3. [3] § Methods › Data sets › Genotype and scRNA-seq data › The ROSMAP study ↔ realdata_code/ROSMAP/GE_generate_CTS.R, lines 1–61 · score 0.64 · excitatory neurons, inhibitory neurons, cell subtypes, OPCs, astrocytes, oligodendrocytes
  4. [4] § Methods › Data sets › Genotype and scRNA-seq data › The OneK1K study ↔ realdata_code/Onek1k/summary_result.R, the whole file · a weak match · score 0.60 · MonoNC, MonoC, BMem, DC, CD4, filtered
  5. [5] § Methods › Other TWAS methods under comparison ↔ realdata_code/ROSMAP/GE_generate_CTS.R, lines 178–241 · score 0.56 · log transformation, pseudo bulk, TMM, aggregation, matrix, TWAS
  6. [6] § Methods › Real data applications › Prediction model training and evaluation ↔ realdata_code/ROSMAP/GE_generate_BULK.R, lines 31–65 · score 0.54 · inhibitory neuron, external, log, OPC, astrocyte, oligodendrocyte

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 320 lines · 14 KB · no license · 3 matches

  1. library(Seurat)
  2. library(SeuratDisk) # new, for h5Seurat objects
  3. library(dplyr)
  4. library(data.table)
  5. library(Matrix)
  6. require(SeqArray)
  7. source('../../R/scTWAS_IRLS.R')
  8. gene_info <- fread('../Onek1k/1k1k_gene_GRCh37.txt')
  9. suppressMessages(library("optparse"))
  10. option_list = list(
  11. make_option("--run_subtype", action="store", default="FALSE", type="character",
  12. help="whether to analyze cell subtypes"),
  13. make_option("--ct", action="store", default="Microglia", type="character",
  14. help="which cell type"),
  15. make_option("--match_with_ct", action="store", default="FALSE", type="character",
  16. help="whether to match the gene evauated in subtypes with cell type's"))
  17. opt = parse_args(OptionParser(option_list=option_list))
  18. run_subtype <- (opt$run_subtype == 'TRUE')
  19. cts <- opt$ct
  20. match_with_ct <- (opt$match_with_ct == 'TRUE')
  21. file_suffix <- ifelse(match_with_ct, '_matched', '')
  22. sprintf('%s, run subtype: %s, %s', ct, run_subtype, file_suffix) %>% print
  23. data_dir <- './data/ROSMAP/Gene Expression (snRNAseq - DLPFC, Experiment 2)/processed (March 2024 update)'
  24. if(run_subtype){
  25. cell_annotation <- read.csv(sprintf('%s/cell-annotation.n424.csv', data_dir))
  26. # handle Ex and cux2+, cux2-
  27. annotation_ct <- ifelse(grepl('cux2', ct), 'Excitatory Neurons', ct)
  28. # get gene expression data for all subtypes
  29. sub_cts <- table(cell_annotation$state[cell_annotation$cell.type == annotation_ct]) %>% sort(decreasing=T) %>% names
  30. print(sprintf('%s subtypes:', ct))
  31. print(sub_cts)
  32. }
  33. ct_seurat_fn <- c('microglia', 'astrocytes',
  34. 'cux2+', 'cux2-',
  35. 'oligodendroglia', 'oligodendroglia',
  36. 'inhibitory',
  37. 'vascular.niche', 'vascular.niche', 'vascular.niche', 'vascular.niche', 'excitatory')
  38. names(ct_seurat_fn) <- c('Microglia', 'Astrocyte',
  39. 'cux2+', 'cux2-',
  40. 'Oligodendrocytes', 'OPCs',
  41. 'Inhibitory Neurons',
  42. 'Endothelial', 'Fibroblast', 'Pericytes', 'SMC', 'excitatory')
  43. subj_var <- 'individualID'
  44. batch_var <- 'batch_combined'
  45. if(ct == 'excitatory'){
  46. obj_list <- list()
  47. obj_list[['cux2-']] <- LoadH5Seurat(sprintf("%s/%s.h5Seurat", data_dir, 'cux2-'), assays='RNA')
  48. obj_list[['cux2+']] <- LoadH5Seurat(sprintf("%s/%s.h5Seurat", data_dir, 'cux2+'), assays='RNA')
  49. obj <- merge(obj_list[['cux2-']], obj_list[['cux2+']], add.cell.ids = c("cux2-", "cux2+"))
  50. print('cux2- and cux2+ combined')
  51. print(dim(obj))
  52. }else{
  53. obj <- LoadH5Seurat(sprintf("%s/%s.h5Seurat", data_dir, ct_seurat_fn[ct]), assays='RNA')
  54. }
  55. # handle Ex:
  56. # only part of Ex subtypes are saved in either cux2+/cux2-
  57. # focus on those subtypes
  58. if(grepl('cux2', ct) & run_subtype){
  59. print('handle Excitatory neurons data')
  60. # subset cell subtypes for Ex
  61. covered_state <- unique(cell_annotation$state[cell_annotation$barcode %in% colnames(obj)]) %>% sort()
  62. sub_cts <- sub_cts[sub_cts %in% covered_state]
  63. print('overlapped cell states:')
  64. print(sub_cts)
  65. }
  66. # handle cell types underlying vascular niche:
  67. # the seurat object contain cell types other than those annotated in cell_annotation
  68. # remove those cells when performing cell-type level aggregation
  69. if(ct_seurat_fn[ct] %in% c('vascular.niche', 'oligodendroglia')){
  70. print('handle vascular niche / oligodendroglia data')
  71. cell_annotation <- read.csv(sprintf('%s/cell-annotation.n424.csv', data_dir))
  72. obj$barcode <- rownames([email hidden])
  73. obj <- subset(obj, subset = barcode %in% cell_annotation$barcode[cell_annotation$cell.type == ct])
  74. }
  75. # -
  76. # remove cells that failed to be mapped to an individual (individualID='NA') or sequenced in a duplicate batch
  77. # following Fujita et al., 2024
  78. # -
  79. print(sprintf('total #cells: %i', ncol(obj)))
  80. # remove cells that were not matched to the 424 subjects
  81. obj <- subset(obj, subset = individualID != 'NA')
  82. cell_annotation <- read.csv(sprintf('%s/cell-annotation.n424.csv', data_dir))
  83. obj$barcode <- rownames([email hidden])
  84. obj <- subset(obj, subset = barcode %in% cell_annotation$barcode)
  85. print(sprintf('#cells matched to 424 subjects: %i', ncol(obj)))
  86. # Combine ...-A and ...-B, e.g. 190403-B4-A and 190403-B4-B as one batch
  87. # "Libraries from four channels were pooled and sequenced on one lane of the Illu- mina HiSeq X "
  88. # As there are a total of eight channels on 10x GEM machine, we assume that -A represents the first four libraries, and -B the second set of four libraries.
  89. [email hidden]$batch_combined <- sapply(1:nrow([email hidden]),
  90. function(i){
  91. x = [email hidden]$batch[i]
  92. substr(x, 1, nchar(x)-2)})
  93. # Identify individualIDs sampled in multiple batches
  94. meta_data1 <- [email hidden] %>% select('individualID', 'batch_combined') %>% unique()
  95. inds_with_dup_batches <- names(which(table(meta_data1$individualID)>1))
  96. keep_batch_list <- character(length(inds_with_dup_batches))
  97. names(keep_batch_list) <- inds_with_dup_batches
  98. for(ind in inds_with_dup_batches){
  99. keep_batch_list[ind] <- names(which.max(table(obj$batch_combined[obj$individualID == ind])))
  100. }
  101. # Remove the batch with less samples
  102. # following the pre-processing in the original paper
  103. # "Among the remaining 436 specimens, 12 individuals were sequenced twice in distinct batches. After comparing sequencing metrics, one of these duplicates was excluded from further analyses."
  104. keep_cell_inds <- rep(T, ncol(obj))
  105. for(ind in names(keep_batch_list)){
  106. subset_inds <- (obj$individualID == ind & obj$batch_combined != keep_batch_list[ind])
  107. keep_cell_inds[subset_inds] <- F
  108. }
  109. obj$keep_cell_inds <- keep_cell_inds
  110. obj <- subset(obj, subset = keep_cell_inds)
  111. print(sprintf('#cells with de-duplicated batch: %i', ncol(obj)))
  112. # -
  113. # match ID between snRNA-seq and WGS data
  114. # -
  115. # load WGS IDs
  116. wgs_samples <- seqVCF_SampID(sprintf('%s/ROSMAP/WGS/DEJ_11898_B01_GRM_WGS_2017-05-15_15.recalibrated_variants.vcf.gz', data_dir))
  117. clinical_covar <- fread(sprintf('%s/ROSMAP_clinical.csv', data_dir))
  118. specimen_covar <- fread(sprintf('%s/ROSMAP_biospecimen_metadata.csv', data_dir))
  119. merged_covar <- merge(clinical_covar, specimen_covar, by = 'individualID', all = TRUE)
  120. merged_covar$newID <- paste0(merged_covar$Study, merged_covar$projid)
  121. # load gene IDs
  122. gene_samples <- unique(obj$individualID)
  123. #colnames(Agg_obj[['RNA']]$counts)
  124. # extract covariate data that covers the gene sample
  125. gene_covar <- merged_covar[merged_covar$individualID %in% gene_samples,]
  126. gene_covar$geno_ID <- rep(NA, nrow(gene_covar))
  127. # some WGS genotype ID are based on Study+projid
  128. gene_covar$geno_ID[gene_covar$newID %in% wgs_samples] <- gene_covar$newID[gene_covar$newID %in% wgs_samples]
  129. # others are based on specimenID
  130. gene_covar$geno_ID[gene_covar$specimenID %in% wgs_samples] <- gene_covar$specimenID[gene_covar$specimenID %in% wgs_samples]
  131. geno_covar_uni <- unique(gene_covar[!is.na(gene_covar$geno_ID), c('individualID', 'geno_ID')])
  132. print(dim(geno_covar_uni)) # some individuals with multiple WGS samples
  133. geno_covar_matched <- geno_covar_uni[match(gene_samples, geno_covar_uni$individualID),] # remove one of the WGS samples for those who have more than one
  134. print(dim(geno_covar_matched))
  135. obj$geno_ID <- geno_covar_matched$geno_ID[match(obj$individualID, geno_covar_matched$individualID)] # match the new geno_IDs to cells
  136. subj_var <- 'geno_ID'
  137. if(ct == 'microglia'){
  138. # save the matching between WGS ID and snRNAseq ID
  139. write.table(geno_covar_matched, sprintf('%s/WGS_snRNAseq_sample_matching.txt', data_dir))
  140. # create a fam file for WGS samples that are matched
  141. wgs_samples_matched <- wgs_samples[match(geno_covar_matched$geno_ID, wgs_samples)]
  142. print(length(wgs_samples_matched))
  143. write.table(data.frame(FID=0, IID=wgs_samples_matched),
  144. sprintf('%s/ROSMAP/WGS/plink_files/match_WGS_samples_subset.txt', data_dir), sep = '\t', quote = F, col.names = F, row.names = F)
  145. }
  146. if(ct %in% c('Endothelial', 'Fibroblast', 'Pericytes', 'SMC')){
  147. print(dim(obj))
  148. saveRDS(dim(obj), sprintf('%s/scTransform_by_celltype_afterAgg/dimension_%s.rds', data_dir, ct))
  149. q()
  150. }
  151. if(!run_subtype){
  152. sub_cts <- cts
  153. }else{
  154. sub_ct_sum <- matrix(nrow=length(sub_cts), ncol=3)
  155. rownames(sub_ct_sum) <- sub_cts
  156. colnames(sub_ct_sum) <- c('ncells', 'ngenes', 'nsubj')
  157. }
  158. # -
  159. # save pseudo-bulk data by cel ltype
  160. # -
  161. for(sub_ct in sub_cts){
  162. print('---------------------------')
  163. print(sprintf('-----------%s----------', sub_ct))
  164. print('---------------------------')
  165. # -
  166. # save scTWAS object
  167. # -
  168. if(run_subtype){
  169. obj_sub <- subset(obj, barcode %in% cell_annotation$barcode[cell_annotation$state == sub_ct])
  170. }else{
  171. obj_sub <- obj
  172. }
  173. sprintf('#cells in %s: %i', sub_ct, ncol(obj_sub))
  174. if(run_subtype) sub_ct_sum[sub_ct, 1] <- ncol(obj_sub)
  175. Agg_obj <- AggregateExpression(object=obj_sub, assay = 'RNA', group.by=subj_var, return.seurat = TRUE)
  176. meta_data1 <- [email hidden] %>% select(all_of(subj_var), batch_combined) %>% unique()
  177. meta_data1 <- meta_data1[match(colnames(Agg_obj),meta_data1[[subj_var]]),]
  178. rownames(meta_data1) <- meta_data1[[subj_var]]
  179. [email hidden] <- meta_data1
  180. ## scTransform normalization
  181. data <- SCTransform(object = Agg_obj, return.only.var.genes = FALSE)
  182. sprintf('#individual samples with %s cells after aggregation: %i', sub_ct, ncol(data)) %>% print
  183. sub_ct_save <- ifelse(grepl('cux2', ct),
  184. sprintf('%s_%s', ct,sub_ct),
  185. sub_ct)
  186. # scTWAS object
  187. saveRDS(data, file = sprintf('%s/scTransform_by_celltype_afterAgg/%s.rds',data_dir,sub_ct_save))
  188. # -
  189. # AN-TWAS
  190. # -
  191. ## TMM+log+batch correct
  192. library(edgeR) # for TMM
  193. library(sva)
  194. pb_mat <- GetAssayData(Agg_obj, assay = 'RNA', layer = 'counts')
  195. # select genes
  196. if(match_with_ct){
  197. print(sprintf('Match with %s', ct))
  198. # use the same genes for subtypes that were selected at the cell type level
  199. # this is used for microglia subtypes to enable evaluating the same genes as in microglia
  200. ct_PB_df <- read.table(sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s_AN.tsv',data_dir,ct), header = T)
  201. min_count <- which(!rownames(pb_mat) %in% ct_PB_df$gene_name)
  202. }else{
  203. # otherwise, select genes based on total gene counts
  204. min_count <- which(rowSums(pb_mat)<1000)
  205. }
  206. pb_mat <- pb_mat[-min_count,]
  207. pb_mat <- pb_mat[order(rowSums(pb_mat),decreasing=T),]
  208. print(dim(pb_mat))
  209. dgList <- DGEList(pb_mat, genes=rownames(pb_mat))
  210. dgList <- calcNormFactors(dgList, method = "TMM")
  211. expr_norm = log_transform(dgList) # same as voom(dgList)$E
  212. donor_pool <- ([email hidden])[[batch_var]]
  213. # SKIP quantile normalization
  214. # https://support.bioconductor.org/p/77664/#77665
  215. expr_norm_inrt <- matrix(NA, nrow = nrow(expr_norm), ncol = ncol(expr_norm))
  216. for(i in 1:nrow(expr_norm)){
  217. expr_norm_inrt[i,] <- INT(expr_norm[i,])
  218. }
  219. rownames(expr_norm_inrt) = rownames(expr_norm)
  220. colnames(expr_norm_inrt) = colnames(expr_norm)
  221. PB_df = data.frame(gene_name = rownames(expr_norm_inrt))
  222. PB_df = cbind(PB_df,expr_norm_inrt)
  223. PB_df = merge(PB_df,gene_info %>% dplyr::select(-ensembl_gene_id),by.x='gene_name',by.y='external_gene_name',sort=FALSE)
  224. PB_df = PB_df %>% relocate(gene_name, chromosome_name, start_position, end_position)
  225. fwrite(PB_df, sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s_AN%s.tsv',data_dir,sub_ct_save, file_suffix),sep='\t')
  226. write.table(colnames(PB_df),sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s%s.header',
  227. data_dir,sub_ct_save,file_suffix),quote=F,row.names=F,col.names=F)
  228. if(run_subtype){
  229. sub_ct_sum[sub_ct,2] <- nrow(expr_norm)
  230. sub_ct_sum[sub_ct,3] <- ncol(expr_norm)
  231. }
  232. # -
  233. # NA-TWAS
  234. # -
  235. print('scTWAS')
  236. print(dim(obj_sub))
  237. sct_obj <- NormalizeData(obj_sub,normalization.method = "LogNormalize", scale.factor = 1e6)
  238. unique_donor <- sort(unique([email hidden][[subj_var]]))
  239. # match with PBINT
  240. print(all(unique_donor == colnames(expr_norm_inrt)))
  241. unique_donor <- unique_donor[match(colnames(expr_norm_inrt), unique_donor)]
  242. print(all(unique_donor == colnames(expr_norm_inrt)))
  243. pr = matrix(NA, nrow = nrow(sct_obj), ncol = length(unique_donor))
  244. print(length(unique_donor))
  245. colnames(pr) = unique_donor; rownames(pr) = rownames(sct_obj)
  246. n_cells <- numeric(length(unique_donor))
  247. names(n_cells) <- unique_donor
  248. for(i in unique_donor){
  249. ind = which([email hidden][[subj_var]] == i)
  250. pr[,i] = rowMeans((sct_obj[['RNA']]$data[,ind,drop=FALSE]))
  251. n_cells[i] = length(ind)
  252. }
  253. if(grepl('cux', sub_ct_save)){
  254. saveRDS(pr, sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s_pr_mean%s.rds',data_dir,sub_ct_save, file_suffix))
  255. saveRDS(n_cells, sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s_pr_n_count%s.rds', data_dir,sub_ct_save,file_suffix))
  256. print('save PR mean before INT normalization')
  257. }
  258. # # SKIP quantile normalization
  259. # # https://support.bioconductor.org/p/77664/#77665
  260. pr_inrt <- matrix(NA, nrow = nrow(pr), ncol = ncol(pr))
  261. print(anyNA(pr))
  262. for(i in 1:nrow(pr)){
  263. pr_inrt[i,] <- INT(pr[i,])
  264. }
  265. rownames(pr_inrt) = rownames(pr)
  266. colnames(pr_inrt) = colnames(pr)
  267. PR_df = data.frame(gene_name = rownames(pr_inrt))
  268. PR_df = cbind(PR_df,pr_inrt)
  269. PR_df = merge(PR_df,gene_info %>% dplyr::select(-ensembl_gene_id),by.x='gene_name',by.y='external_gene_name',sort=FALSE)
  270. PR_df = PR_df %>% relocate(gene_name, chromosome_name, start_position, end_position)
  271. PR_df = PR_df[match(PB_df$gene_name,PR_df$gene_name),]
  272. fwrite(PR_df, sprintf('%s/scTransform_by_celltype_afterAgg/Expression_matrices/%s_NA%s.tsv',
  273. data_dir,sub_ct_save,file_suffix),sep='\t')
  274. }
  275. print('Data saved')
  276. if(run_subtype){
  277. write.table(sub_ct_sum, sprintf('%s/scTransform_by_celltype_afterAgg/%s_subtypes_summary%s.txt',
  278. data_dir,ct,file_suffix))
  279. print(sprintf('%s/scTransform_by_celltype_afterAgg/%s_subtypes_summary%s.txt',data_dir,ct,file_suffix))
  280. }

GE_generate_CTS.R at commit f4120fa, no license · at the source

Overview

Authors: Zhaotong Lin1, Chang Su2,3
  1. Department of Statistics, Florida State University,Tallahassee, FL USA
  2. Department of Biostatistics and Bioinformatics, Emory University,Atlanta, GA USA
  3. Department of Human Genetics, Emory University,Atlanta, GA USA
Institutions: Florida State University (United States); Emory University (United States)
Journal: Nature communications, volume 17, issue 1, article 3853
Dates: received 1 May 2025; accepted 25 February 2026; published online 12 March 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-70374-7 · PMID 41820391 · PMCID PMC13121454 · OpenAlex W7135087009
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Alzheimer's / dementia (population), cellular / molecular (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing
Keywords: Genome-wide association studies, Gene expression profiling, Gene expression, Statistical methods, Computational models
MeSH: Blood*, Brain*, Single-Cell Gene Expression Analysis*, Aging, Datasets as Topic, Dementia, Gene Expression Regulation, Genetic Association Studies, Humans, Leukocytes, Mononuclear (* major topic)
Topic: Single-cell and spatial transcriptomics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: URC Research Award at Emory University
Citations: cited by 3 papers (Europe PMC); 114 references in the paper

Abstract

Transcriptome-wide association studies (TWAS) have successfully identified genes associated with complex traits and diseases, but most have been performed using bulk gene expression data, which aggregate signals across heterogeneous cell types. Population-scale single-cell RNA sequencing data now make it possible to perform TWAS at the cell-type resolution, but present unique challenges due to strong noises, technical variations, and high sparsity. Here, we propose scTWAS, a statistical method to conduct cell-type-specific TWAS using single-cell data. Leveraging a latent-variable model and moment-based estimation to address the challenges of single-cell data, scTWAS consistently improves the prediction of genetically regulated gene expression across cell types in both blood and brain tissues. Compared to existing methods, scTWAS identifies substantially more gene-trait associations across 29 hematological traits and three immune-related diseases in immune cell types. An application to Alzheimer’s disease also reveals cell-subtype-specific associations, including MS4A6A in the disease-associated microglial subtype and PPP1R37 in the inflammatory microglial subtype.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

ZhaotongL/scTWAS

License: other
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: bbd6bf099325721fb1de2b66948445662336d64f, 24 December 2025
Languages: R (6)
Size: 34 files, 6 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (DESCRIPTION), documentation
Not found: CITATION.cff, tests, continuous integration
Tools: glmnet (1 file), Seurat (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
9 files

ZhaotongL/scTWAS_paper

License: none: the authors keep all their rights
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: f4120fac41ee6a87caedfed976dc644d05ac303a, 31 December 2025
Languages: R (12), Shell (4)
Size: 32 files, 16 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: data.table (8 files), tidyverse (8 files), Seurat (7 files), edgeR (4 files), caret (2 files), glmnet (2 files), ggplot2 (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
17 files

Zenodo 18434920

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: glmnet (1 file), Seurat (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
9 files
At the source:

Zenodo 18434979

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: data.table (8 files), tidyverse (8 files), Seurat (7 files), edgeR (4 files), caret (2 files), glmnet (2 files), ggplot2 (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
17 files
At the source:

Code availability

R package for scTWAS is available at https://github.com/ZhaotongL/scTWAS113. Code for real data analyses is available at https://github.com/ZhaotongL/scTWAS_paper114.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 44 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The GReX prediction models for immune and brain cell types generated in this study have been deposited in Figshare at [10.6084/m9.figshare.31329184] and [10.6084/m9.figshare.31329139]111,112. The availability for datasets used to train or validate GReX models is as follows: OneK1K genotype and scRNA-seq data are publicly available from the OneK1K study17. Genotype and gene expression data of immune cell types from the DICE project16 are available under controlled access in dbGap under accession code phs001703.v5.p1 [https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs001703.v5.p1], access can be obtained by requesting Authorized Access on dbGap. WGS and snRNA-seq data from the ROSMAP study are available on Synapse under controlled access with accession codes syn11724057 [https://www.synapse.org/Synapse:syn11724057], syn53366818 [https://www.synapse.org/Synapse:syn53366818], and syn52293417 [https://www.synapse.org/Synapse:syn52293417], access can be obtained by submitting a Data Use Certificate to Synapse.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 5 keywords, 10 MeSH terms, 1 funder, 112 references.

Cite

This paper

Lin, Z., & Su, C. (2026). scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies. Nature communications, 17(1), 3853. https://doi.org/10.1038/s41467-026-70374-7

BibTeX

@article{lin2026sctwas,
author = {Lin, Zhaotong and Su, Chang},
title = {{scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies}},
journal = {Nature communications},
year = {2026},
month = mar,
volume = {17},
number = {1},
pages = {3853},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-70374-7},
url = {https://doi.org/10.1038/s41467-026-70374-7},
pmid = {41820391},
pmcid = {PMC13121454}
}

RIS

TY - JOUR
AU - Lin, Zhaotong
AU - Su, Chang
TI - scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/03/12
VL - 17
IS - 1
SP - 3853
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-70374-7
UR - https://doi.org/10.1038/s41467-026-70374-7
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-70374-7",
"type": "article-journal",
"title": "scTWAS: a powerful statistical framework for single-cell transcriptome-wide association studies",
"container-title": "Nature communications",
"author": [
{
"family": "Lin",
"given": "Zhaotong"
},
{
"family": "Su",
"given": "Chang"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "3853",
"DOI": "10.1038/s41467-026-70374-7",
"PMID": "41820391",
"PMCID": "PMC13121454",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-70374-7",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s42003-026-10030-4 [code]
Cell-type-aware transcriptome-wide association studies identify 91 independent risk genes for Alzheimer's disease dementia.
Journal: Communications biology
In common: ggplot2, tidyverse, Alzheimer's / dementia, genetics / omics, cellular / molecular, 12 references
[2] doi:10.1038/s41467-026-73007-1 [code]
Single-nucleus epigenomic dysregulation unmasks genetic risk-associated neurodegenerative glia states.
Journal: Nature communications
In common: Seurat, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics, cellular / molecular, 10 references
[3] doi:10.1038/s41514-026-00391-9 [code]
Region-specific transcriptional signatures of brain aging in the absence of neuropathology at the single-cell level.
Journal: npj aging
In common: edgeR, Seurat, data.table, 2 other tools, genetics / omics, cellular / molecular, 7 references
[4] doi:10.1038/s41588-026-02646-3 [code]
Co-expression-based models improve eQTL predictions for transcriptome-wide association studies and highlight new schizophrenia-associated genes.
Journal: Nature genetics
In common: glmnet, caret, data.table, 1 other tool, genetics / omics, cellular / molecular, 5 references
[5] doi:10.1016/j.isci.2026.116412 [code]
KOLF2.1J iTF-Microglia: A standardized platform to study microglial transcriptional regulatory networks in CNS disease.
Journal: iScience
In common: edgeR, data.table, ggplot2, 1 other tool, genetics / omics, 6 references
[6] doi:10.1371/journal.pgen.1012126 [code]
FM-GPT: Bayesian fine mapping for phenome-wide transcriptome-wide association studies.
Journal: PLoS genetics
In common: genetics / omics, cellular / molecular, 8 references
[7] doi:10.1038/s41467-026-75193-4 [code]
Multi-ancestry gene expression models amplify transcriptome-wide association study discovery and validation.
Journal: Nature communications
In common: Seurat, data.table, ggplot2, 1 other tool, genetics / omics, cellular / molecular, 5 references
[8] doi:10.1186/s12967-026-08266-z [code]
Single-cell multi-omic integration analysis prioritizes druggable genes and reveals cell-type-specific causal effects in glioblastomagenesis.
Journal: Journal of translational medicine
In common: edgeR, Seurat, data.table, 2 other tools, genetics / omics, cellular / molecular, 3 references
[9] doi:10.1002/alz.71823 [code]
Cellular transcriptomic signatures underpinning the heterogeneity of depression in Alzheimer's disease.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: Seurat, data.table, ggplot2, 1 other tool, Alzheimer's / dementia, genetics / omics, cellular / molecular, 4 references
[10] doi:10.1002/alz.71558 [code]
Allele specific expression in Alzheimer's disease.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: Seurat, ggplot2, synapse.org/synapse:syn52293417, Alzheimer's / dementia, genetics / omics, cellular / molecular, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.