OSCR

Comprehensive analysis of interactions between brain aging and late-onset psychoses using heuristic mapping models.

Code ↔ Paper

The paper beside its authors' code: matches between them have not been computed for this paper yet.

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 693 lines · 35 KB · no license

  1. #' This function performs variable selection on gene expression matrices.
  2. #' It can, for instance, remove genes with low variation.
  3. #' @param exprMat A matrix of gene expression levels. rownames() are genes, and colnames() are samples.
  4. #' @param removeLowVaryingGenes The proportion of low varying genes to be removed.The default is .2
  5. #' @return A vector of row/genes to keep.
  6. #' @export
  7. doVariableSelection <- function(exprMat, removeLowVaryingGenes=.2)
  8. {
  9. vars <- apply(exprMat, 1, var)
  10. return(order(vars, decreasing=TRUE)[seq(1:as.integer(nrow(exprMat)*(1-removeLowVaryingGenes)))])
  11. }
  12. #'This function takes two gene expression matrices (like trainExprMat and testExprMat) and returns homogenized versions of the matrices by employing the homogenization method specified.
  13. #'By default, the Combat method from the sva library is used.
  14. #'In both matrices, genes are row names and samples are column names.
  15. #'It will deal with duplicated gene names, as it subsets and orders the matrices correctly.
  16. #'@param testExprMat A gene expression matrix for samples on which we wish to predict a phenotype.Genes are rows, samples are columns.
  17. #'@param trainExprMat A gene expression matrix for samples for which the phenotype is already known.Genes are rows, samples are columns.
  18. #'@param batchCorrect The type of batch correction to be used. Options are 'eb' for Combat, 'none', or 'qn' for quantile normalization.
  19. #'#The default is 'eb'.
  20. #'@param selection This parameter can be used to specify how duplicates are handled. The default value of -1 means to ask the user.
  21. #'#Other options include '1' to summarize duplicates by their mean, and '2'to discard all duplicated genes.
  22. #'@param printOutput To suppress output, set to false. Default is TRUE.
  23. #'@import sva
  24. #'@import preprocessCore
  25. #'@keywords Homogenize gene expression data.
  26. #'@return A list containing two entries $train and $test, which are the homogenized input matrices.
  27. #'@export
  28. homogenizeData<-function (testExprMat, trainExprMat, batchCorrect = "eb", selection = -1, printOutput = TRUE)
  29. {
  30. #Check the batchCorrect parameter
  31. if (!(batchCorrect %in% c("eb", "qn", "none",
  32. "rank", "rank_then_eb", "standardize")))
  33. stop("\"batchCorrect\" must be one of \"eb\", \"qn\", \"rank\", \"rank_then_eb\", \"standardize\" or \"none\"")
  34. #Check if both row and column names have been specified.
  35. if (is.null(rownames(trainExprMat)) || is.null(rownames(testExprMat))) {
  36. stop("ERROR: Gene identifiers must be specified as \"rownames()\" on both training and test expression matrices. Both matices must have the same type of gene identifiers.")
  37. }
  38. #Check that some of the row names overlap between both datasets (print an error if none overlap)
  39. if (sum(rownames(trainExprMat) %in% rownames(testExprMat)) == 0) {
  40. stop("ERROR: The rownames() of the supplied expression matrices do not match. Note that these are case-sensitive.")
  41. }
  42. else {
  43. if (printOutput)
  44. message(paste("\n", sum(rownames(trainExprMat) %in% rownames(testExprMat)), " gene identifiers overlap between the supplied expression matrices... \n", paste = ""))
  45. }
  46. #If there are duplicate gene names, give the option of removing them or summarizing them by their mean.
  47. if ((sum(duplicated(rownames(trainExprMat))) > 0) || sum(sum(duplicated(rownames(testExprMat))) > 0)) {
  48. if (selection == -1) {
  49. message("\nExpression matrix contain duplicated gene identifiers (i.e. duplicate rownames()), how would you like to proceed:")
  50. message("\n1. Summarize duplicated gene ids by their mean value (acceptable in most cases)")
  51. message("\n2. Disguard all duplicated genes (recommended if unsure)")
  52. message("\n3. Abort (if you want to deal with duplicate genes ids manually)\n")
  53. }
  54. while (is.na(selection) | selection <= 0 | selection > 3) {
  55. selection <- readline("Selection: ")
  56. selection <- ifelse(grepl("[^1-3.]", selection), -1, as.numeric(selection))
  57. }
  58. message("\n")
  59. if (selection == 1) #Summarize duplicates by their mean.
  60. {
  61. if ((sum(duplicated(rownames(trainExprMat))) > 0)) {
  62. trainExprMat <- summarizeGenesByMean(trainExprMat)
  63. }
  64. if ((sum(duplicated(rownames(testExprMat))) > 0)) {
  65. testExprMat <- summarizeGenesByMean(testExprMat)
  66. }
  67. }
  68. else if (selection == 2) #Disguard all duplicated genes.
  69. {
  70. if ((sum(duplicated(rownames(trainExprMat))) > 0)) {
  71. keepGenes <- names(which(table(rownames(trainExprMat)) == 1))
  72. trainExprMat <- trainExprMat[keepGenes, ]
  73. }
  74. if ((sum(duplicated(rownames(testExprMat))) > 0)) {
  75. keepGenes <- names(which(table(rownames(testExprMat)) == 1))
  76. testExprMat <- testExprMat[keepGenes, ]
  77. }
  78. }
  79. else {
  80. stop("Exectution Aborted!")
  81. }
  82. }
  83. #Subset and order gene ids on the expression matrices.
  84. commonGenesIds <- rownames(trainExprMat)[rownames(trainExprMat) %in%
  85. rownames(testExprMat)]
  86. trainExprMat <- trainExprMat[commonGenesIds, ]
  87. testExprMat <- testExprMat[commonGenesIds, ]
  88. #Subset and order the 2 expression matrices.
  89. if (batchCorrect == "eb") {
  90. #Subset to common genes and batch correct using ComBat.
  91. dataMat <- cbind(trainExprMat, testExprMat)
  92. mod <- data.frame(`(Intercept)` = rep(1, ncol(dataMat)))
  93. rownames(mod) <- colnames(dataMat)
  94. whichbatch <- as.factor(c(rep("train", ncol(trainExprMat)),
  95. rep("test", ncol(testExprMat))))
  96. # Added
  97. # Filter out genes with low variances to make sure comBat run correctly
  98. dataMat <- cbind(trainExprMat, testExprMat)
  99. gene_vars = apply(dataMat, 1, var)
  100. genes<-as.vector(gene_vars)
  101. if (length(which(genes <= 1e-3) != 0)){ #If some genes have low variances (if the variance is not 0), remove them.
  102. dataMat = dataMat[-(which(genes <= 1e-3)),]
  103. }
  104. # End added
  105. combatout <- ComBat(dataMat, whichbatch, mod = mod)
  106. return(list(train = combatout[, whichbatch == "train"],
  107. test = combatout[, whichbatch == "test"], selection = selection))
  108. }
  109. else if (batchCorrect == "standardize") #Standardize to mean 0 and variance 1 in each dataset using a non EB based approach.
  110. {
  111. for (i in 1:nrow(trainExprMat)) {
  112. row <- trainExprMat[i, ]
  113. trainExprMat[i, ] <- ((row - mean(row))/sd(row))
  114. }
  115. for (i in 1:nrow(testExprMat)) {
  116. row <- testExprMat[i, ]
  117. testExprMat[i, ] <- ((row - mean(row))/sd(row))
  118. }
  119. return(list(train = trainExprMat, test = testExprMat,
  120. selection = selection))
  121. }
  122. else if (batchCorrect == "rank") #The random-rank transform approach, that may be better when applying models to RNA-seq data.
  123. {
  124. for (i in 1:nrow(trainExprMat)) {
  125. trainExprMat[i, ] <- rank(trainExprMat[i, ], ties.method = "random")
  126. }
  127. for (i in 1:nrow(testExprMat)) {
  128. testExprMat[i, ] <- rank(testExprMat[i, ], ties.method = "random")
  129. }
  130. return(list(train = trainExprMat, test = testExprMat,
  131. selection = selection))
  132. }
  133. else if (batchCorrect == "rank_then_eb") #Rank-transform the RNAseq data, then apply ComBat
  134. {
  135. #First, rank transform the RNA-seq data.
  136. for (i in 1:nrow(testExprMat)) {
  137. testExprMat[i, ] <- rank(testExprMat[i, ], ties.method = "random")
  138. }
  139. #Subset to common genes and batch correct using ComBat.
  140. dataMat <- cbind(trainExprMat, testExprMat)
  141. mod <- data.frame(`(Intercept)` = rep(1, ncol(dataMat)))
  142. rownames(mod) <- colnames(dataMat)
  143. whichbatch <- as.factor(c(rep("train", ncol(trainExprMat)),
  144. rep("test", ncol(testExprMat))))
  145. combatout <- ComBat(dataMat, whichbatch, mod = mod)
  146. return(list(train = combatout[, whichbatch == "train"],
  147. test = combatout[, whichbatch == "test"], selection = selection))
  148. }
  149. else if (batchCorrect == "qn")
  150. {
  151. dataMat <- cbind(trainExprMat, testExprMat)
  152. dataMatNorm <- normalize.quantiles(dataMat)
  153. whichbatch <- as.factor(c(rep("train", ncol(trainExprMat)),
  154. rep("test", ncol(testExprMat))))
  155. return(list(train = dataMatNorm[, whichbatch == "train"],
  156. test = dataMatNorm[, whichbatch == "test"],
  157. selection = selection))
  158. }
  159. else {
  160. return(list(train = trainExprMat, test = testExprMat,
  161. selection = selection))
  162. }
  163. }
  164. #'This function takes a gene expression matrix and if duplicate genes are measured, summarizes them by their means.
  165. #'@param exprMat A gene expression matrix with genes as rownames() and samples as colnames().
  166. #'@return A gene expression matrix that does not contain duplicate genes.
  167. #'@keywords Summarize duplicate genes by their mean.
  168. #'@export
  169. summarizeGenesByMean <- function(exprMat)
  170. {
  171. geneIds <- rownames(exprMat)
  172. t <- table(geneIds) #How many times is each gene name duplicated.
  173. allNumDups <- unique(t)
  174. allNumDups <- allNumDups[-which(allNumDups == 1)]
  175. #Create a *new* gene expression matrix with everything in the correct order....
  176. #Starrt by just adding stuff that isn't duplicated
  177. exprMatUnique <- exprMat[which(geneIds %in% names(t[t == 1])), ]
  178. gnamesUnique <- geneIds[which(geneIds %in% names(t[t == 1]))]
  179. #Add all the duplicated genes to the bottom of "exprMatUniqueHuman", summarizing as you go
  180. for(numDups in allNumDups)
  181. {
  182. geneList <- names(which(t == numDups))
  183. for(i in 1:length(geneList))
  184. {
  185. exprMatUnique <- rbind(exprMatUnique, colMeans(exprMat[which(geneIds == geneList[i]), ]))
  186. gnamesUnique <- c(gnamesUnique, geneList[i])
  187. # print(i)
  188. }
  189. }
  190. if(class(exprMatUnique) == "numeric")
  191. {
  192. exprMatUnique <- matrix(exprMatUnique, ncol=1)
  193. }
  194. rownames(exprMatUnique) <- gnamesUnique
  195. return(exprMatUnique)
  196. }
  197. #'This function predicts a phenotype (drug sensitivity score) when provided with microarray or bulk RNAseq gene expression data of different platforms.
  198. #'The imputations are performed using ridge regression, training on a gene expression matrix where phenotype is already known.
  199. #'This function integrates training and testing datasets via a user-defined procedure, and power transforming the known phenotype.
  200. #'@param trainingExprData The training data. A matrix of expression levels. rownames() are genes, colnames() are samples (cell line names or cosmic ides, etc.). rownames() must be specified and must contain the same type of gene ids as "testExprData"
  201. #'@param trainingPtype The known phenotype for "trainingExprData". This data must be a matrix of drugs/rows x cell lines/columns or cosmic ids/columns. This matrix can contain NA values, that is ok (they are removed in the calcPhenotype() function).
  202. #'@param testExprData The test data where the phenotype will be estimated. It is a matrix of expression levels, rows contain genes and columns contain samples, "rownames()" must be specified and must contain the same type of gene ids as "trainingExprData".
  203. #'@param batchCorrect How should training and test data matrices be homogenized. Choices are "eb" (default) for ComBat, "qn" for quantiles normalization or "none" for no homogenization.
  204. #'@param powerTransformPhenotype Should the phenotype be power transformed before we fit the regression model? Default to TRUE, set to FALSE if the phenotype is already known to be highly normal.
  205. #'@param removeLowVaryingGenes What proportion of low varying genes should be removed? 20 percent be default
  206. #'@param minNumSamples How many training and test samples are required. Print an error if below this threshold
  207. #'@param selection How should duplicate gene ids be handled. Default is -1 which asks the user. 1 to summarize by their or 2 to disguard all duplicates.
  208. #'@param printOutput Set to FALSE to supress output.
  209. #'@param pcr Indicates whether or not you'd like to use pcr for feature (gene) reduction. Options are 'TRUE' and 'FALSE'. If you indicate 'report_pc=TRUE' you need to also indicate 'pcr=TRUE'
  210. #'@param removeLowVaringGenesFrom Determine method to remove low varying genes. Options are 'homogenizeData' and 'rawData'.
  211. #'@param report_pc Indicates whether you want to output the training principal components. Options are 'TRUE' and 'FALSE'.
  212. #'@param cc Indicate if you want correlation coefficients for biomarker discovery.
  213. #'@param percent Indicate percent variability (of the training data) you'd like principal components to reflect if pcr=TRUE. Default is 80 for 80%
  214. #'These are the correlations between a given gene of interest across all samples vs. a given drug response across samples.
  215. #'These correlations can be ranked to obtain a ranked correlation to determine highly correlated drug-gene associations.
  216. #'@param rsq Indicate whether or not you want to output the R^2 values for the data you train on from true and predicted values.
  217. #'These values represent the percentage in which the optimal model accounts for the variance in the training data.
  218. #'Options are 'TRUE' and 'FALSE'.
  219. #'@return .txt files will be saved into your working directory. Depending on the parameter specified, the .txt file outputs of this function can include the estimated phenotype/sensitivity predictions,
  220. #'the R^2 data, and the correlation coefficients. Principal components are stored as .RData files for each drug in your drug dataset.
  221. #'@import sva
  222. #'@import ridge
  223. #'@import car
  224. #'@import ridge
  225. #'@import glmnet
  226. #'@import tidyverse
  227. #'@import utils
  228. #'@import stats
  229. #'@importFrom pls explvar explvar
  230. #'@keywords predict drug sensitivity and phenotype
  231. #'@export
  232. calcPhenotype<-function (trainingExprData,
  233. trainingPtype,
  234. testExprData,
  235. batchCorrect,
  236. powerTransformPhenotype=TRUE,
  237. removeLowVaryingGenes=0.2,
  238. minNumSamples,
  239. selection=1,
  240. printOutput,
  241. pcr=FALSE,
  242. removeLowVaringGenesFrom,
  243. report_pc=FALSE,
  244. cc=FALSE,
  245. percent=80,
  246. rsq=FALSE)
  247. {
  248. #Initiate empty lists for each data type you'd like to collect.
  249. #_______________________________________________________________
  250. DrugPredictions<-list() #Collects drug predictions.
  251. rsqs<-list() #Collects R^2 values.
  252. cors<-list() #Collects correlation coefficient for each gene across all samples vs. each drug across all samples.
  253. pvalues<-list() #Collects p-values for the correlation coefficients.
  254. #vs=c()
  255. drugs<-colnames(trainingPtype) #Store all the possible drugs in a vector.
  256. #Check the supplied data and parameters.
  257. #_______________________________________________________________
  258. if (class(testExprData)[1] != "matrix")
  259. stop("\nERROR: \"testExprData\" must be a matrix.")
  260. if (class(trainingExprData)[1] != "matrix")
  261. stop("\nERROR: \"trainingExprData\" must be a matrix.")
  262. if (class(trainingPtype)[1] != "matrix")
  263. stop("\nERROR: \"trainingPtype\" must be a matrix.")
  264. if (report_pc)
  265. if (pcr == FALSE)
  266. stop("\nERROR: pcr must be TRUE if report_pc is TRUE")
  267. if (pcr)
  268. if (cc)
  269. stop("\nERROR: pcr must be FALSE if cc is TRUE")
  270. #Make sure training samples are equivalent in both matrices.
  271. if (!any(colnames(trainingExprData) %in% rownames(trainingPtype)))
  272. stop("\nERROR: No Cell Lines Found in Common: Sample names must be consistent in training matrices")
  273. #Subset and order the training Expr and trainingPtype to the cell lines in common (and order them)
  274. commonCellLines<-colnames(trainingExprData)[colnames(trainingExprData) %in% rownames(trainingPtype)]
  275. trainingExprData <- trainingExprData[,commonCellLines]
  276. trainingPtype <- trainingPtype[commonCellLines,]
  277. #Check if an adequate number of training and test samples have been supplied.
  278. #_______________________________________________________________
  279. if ((nrow(trainingExprData) < minNumSamples) || (nrow(testExprData) < minNumSamples)) {
  280. stop(paste("\nThere are less than", minNumSamples, "samples in your test or training set. It is strongly recommended that you use larger numbers of samples in order to (a) correct for batch effects and (b) fit a reliable model. To supress this message, change the \"minNumSamples\" parameter to this function."))
  281. }
  282. #Get the homogenized data.
  283. #_______________________________________________________________
  284. homData <- homogenizeData(testExprMat=testExprData, trainExprMat=trainingExprData, batchCorrect, selection, printOutput)
  285. #Remove low varying genes.
  286. #_______________________________________________________________
  287. #Do variable selection if specified. By default, we remove 20% of least varying genes from the homogenized dataset.
  288. #We can also remove the intersection of the lowest 20% from both training and test sets (for the gene ids remaining in the homogenized data).
  289. #Otherwise, keep all genes.
  290. #Check batchCorrect parameter.
  291. if (!(removeLowVaringGenesFrom %in% c("homogenizeData", "rawData"))) {
  292. stop("\nremoveLowVaringGenesFrom\" must be one of \"homogenizeData\", \"rawData\"")
  293. }
  294. keepRows <- seq(1:nrow(homData$train)) #By default we will keep all the genes.
  295. if (removeLowVaryingGenes > 0 && removeLowVaryingGenes < 1) { #If the proportion of variability to keep is between 0 and 1.
  296. if (removeLowVaringGenesFrom == "homogenizeData") { #If you're filtering based on homogenized data.
  297. keepRows <- doVariableSelection(cbind(homData$test, homData$train), removeLowVaryingGenes = removeLowVaryingGenes)
  298. numberGenesRemoved <- nrow(homData$test) - length(keepRows)
  299. if (printOutput) message(paste("\n", numberGenesRemoved, "low variabilty genes filtered."));
  300. }
  301. else if (removeLowVaringGenesFrom == "rawData") { #If we are filtering based on the raw data i.e. the intersection of the things filtered from both datasets.
  302. evaluabeGenes <- rownames(homData$test)
  303. keepRowsTrain <- doVariableSelection(trainingExprData[evaluabeGenes,], removeLowVaryingGenes = removeLowVaryingGenes)
  304. keepRowsTest <- doVariableSelection(testExprData[evaluabeGenes,], removeLowVaryingGenes = removeLowVaryingGenes)
  305. keepRows <- intersect(keepRowsTrain, keepRowsTest)
  306. numberGenesRemoved <- nrow(homData$test) - length(keepRows)
  307. if (printOutput)
  308. message(paste("\n", numberGenesRemoved, "low variabilty genes filtered."));
  309. }
  310. }
  311. #Predict for each drug.
  312. #_______________________________________________________________
  313. for(a in 1:length(drugs)){ #For each drug...
  314. #Modify trainingPtype and trainingExprData so that you only use cell lines for which you have expression and response data for.
  315. #_______________________________________________________________
  316. trainingPtype2<-trainingPtype[,a] #Obtain the response data for the compound of interest.
  317. NonNAindex <- which(!is.na(trainingPtype2)) #Get the indices of the non NAs. You only want the cell lines/cosmic ids that you have drug info for.
  318. samps<-rownames(trainingPtype)[NonNAindex] #Obtain cell lines you have expression and response data for.
  319. if (length(samps) == 1){ #Make sure training data has more than just 1 sample. If the drug has one sample, it will be skipped.
  320. drugs = drugs[-a]
  321. message(paste("\n", drugs[a], "is skipped due to insufficient cell lines to fit the model."))
  322. next
  323. } else {
  324. trainingPtype4<-as.numeric(trainingPtype2[NonNAindex]) #This makes sure you use all the response data for the drug without NA values.
  325. #PowerTransform the phenotype if specified.
  326. #_______________________________________________________________
  327. offset = 0
  328. if (powerTransformPhenotype){
  329. if (min(trainingPtype4) < 0){ #All numbers must be positive for a powertransform to work, so make them positive.
  330. offset <- -min(trainingPtype4) + 1
  331. trainingPtype4 <- trainingPtype4 + offset
  332. }
  333. transForm <- powerTransform(trainingPtype4)[[6]]
  334. trainingPtype4 <- trainingPtype4^transForm
  335. }
  336. #Create the ridge regression model on the training data using pcr (keeping the number of components required for 80% variance).
  337. #_______________________________________________________________
  338. if (pcr){
  339. #There are many ways for pcr in R...here I will use pcr() (principal component regression).
  340. #Use pcr to predict for testing data.
  341. #_______________________________________________________________
  342. train_x<-(t(homData$train)[samps,keepRows]) #samps represent the cell lines that have been filtered, keepRows represents the genes.
  343. train_y<-trainingPtype4
  344. test_x<-(t(homData$test)[,keepRows])
  345. #Remove genes that you have no transcriptome data for (aka columns are filled with 0's).
  346. #If you don't, it causes problems in linearRidge().
  347. #This is weird, but there were 2 genes in CTRPv2 data that had 0's for one drug (I suppose this occurs depending on sample filtration).
  348. #Make sure you remove those same genes from the test data...want to make sure the genes are the same in both train_x and test_x
  349. x<-as.vector(colSums(train_x)) #Sum each column.
  350. bad<-which(x == 0) #Column/genes indices that are filled with only 0's.
  351. if(length(bad) != 0){
  352. train_x<-train_x[,-bad]
  353. test_x<-data.frame(test_x[,-bad])
  354. }
  355. #Remove genes with 0 variance. If you don't, this will also cause problems in linearRidge().
  356. variance<-c()
  357. for(i in 1:ncol(train_x)){
  358. variance[i]<-var(as.vector(train_x[,i]))
  359. }
  360. bi<-which(variance %in% 0) #Bad index...gene has 0 variance.
  361. if(length(bi) != 0){
  362. train_x<-train_x[,-bi]
  363. test_x<-data.frame(test_x[,-bi])
  364. }
  365. #Check to make sure you have more than 1 training sample for the drug's model.
  366. #_______________________________________________________________
  367. trainFrame<-try(data.frame(Resp=train_y, train_x), silent = TRUE)
  368. if (dim(trainFrame)[1] == 1){ #Make sure you have more than 1 sample.
  369. drugs = drugs[-a]
  370. message(paste("\n", drugs[a], "is skipped due to insufficient cell lines to fit the model."))
  371. next
  372. } else {
  373. pcr_model<-pcr(Resp~., data=trainFrame, validation='CV')
  374. v=cumsum(explvar(pcr_model)) #A vector of all the pcs and their percent of variance.
  375. ncomp=min(which(v > percent)) #Identify which pcs will represent 80% variance.
  376. #vs[a]<-ncomp
  377. if(printOutput) message("\nCalculating predicted phenotype using pcr...")
  378. preds<-predict(pcr_model, newdata=test_x, ncomp=ncomp)
  379. #You can compute an R^2 value for the data you train on from true and predicted values.
  380. #The rsq value represents the percentage in which the optimal model accounts for the variance in the training data.
  381. #_______________________________________________________________
  382. if (rsq){
  383. if (dim(train_x)[1] < 4){ #The code will result in an error if you have 3 samples (which is enough for the model fitting but not when you do a 70/30% split)...
  384. message(paste("\n", drugs[a], 'is skipped for R^2 analysis'))
  385. }else{
  386. #trainFrame<-data.frame(Resp=train_y, train_x)
  387. data<-(cbind(train_x, train_y)) #Rows are samples, columns are genes.
  388. dt<-sort(sample(nrow(data), nrow(data)*.7)) #sample() randomly picks 70% of rows/samples from the dataset. It samples without replacement.
  389. #Prepare the training data (70% of original training data)
  390. train_x<-data[dt,]
  391. ncol<-dim(train_x)[2]
  392. train_y<-train_x[,ncol]
  393. train_x<-train_x[,-ncol]
  394. #Prepare the testing data (30% of original training data)
  395. test_x<-data[-dt,]
  396. ncol<-dim(test_x)[2]
  397. test_y<-test_x[,ncol]
  398. test_x<-test_x[,-ncol]
  399. #Remove genes that you have no transcriptome data for aka columns are filled with 0's.
  400. #If you don't, it causes problems in linearRidge()
  401. #Make sure you remove these same genes from the test data...want to make sure the genes are the same in both.
  402. x<-as.vector(colSums(train_x))
  403. bad<-which(x == 0)
  404. if (length(bad) != 0){
  405. train_x<-train_x[,-bad]
  406. test_x<-data.frame(test_x[,-bad])
  407. }
  408. #Remove genes with 0 variance. If you don't, this will also cause problems in linearRidge()
  409. variance<-c()
  410. for(i in 1:ncol(train_x)){
  411. variance[i]<-var(as.vector(train_x[,i]))
  412. }
  413. bi<-which(variance %in% 0) #Bad index...gene has 0 variance.
  414. if (length(bi) != 0){
  415. train_x<-train_x[,-bi]
  416. }
  417. data<-data.frame(Resp=train_y, train_x)
  418. pcr_model<-pcr(Resp~., data=data, validation='CV') #Set validation argument to CV.
  419. v=cumsum(explvar(pcr_model)) #A vector of all the pcs and their percent of variance.
  420. ncomp=min(which(v > percent)) #Identify which pcs will represent 80% variance.
  421. pcr_pred<-predict(pcr_model, test_x, ncomp=ncomp)
  422. if (printOutput) message("\nCalculating R^2...")
  423. sst<-sum((test_y - mean(test_y))^2) #Compute the sum of squares total.
  424. sse<-sum((pcr_pred - test_y)^2) #Compute the sum of squares error.
  425. rsq_value<-1 - sse/sst #Compute the rsq value.
  426. }
  427. }
  428. }
  429. if (report_pc){
  430. if (printOutput) message("\nObtaining principal components...")
  431. pcs<-coef(pcr_model, comps = ncomp) #comps: numeric, which components to return.
  432. dir.create("./calcPhenotype_Output")
  433. path<-paste('./calcPhenotype_Output/', drugs[a], '.RData', sep="")
  434. save(pcs, file=path)
  435. }
  436. } else {
  437. #Create the ridge regression model on our training data to predict for our actual testing data without pcr.
  438. #_______________________________________________________________
  439. if(printOutput) message("\nFitting Ridge Regression model...");
  440. expression<-(t(homData$train)[samps,keepRows]) #samps represent the cell lines that have been filtered, keepRows represents the genes.
  441. test_x<-(t(homData$test)[,keepRows])
  442. #Remove genes that you have no transcriptome data for (aka columns are filled with 0's).
  443. #If you don't, it causes problems in linearRidge().
  444. #This is weird, but there were 2 genes in CTRPv2 data that had 0's for one drug (I suppose this occurs depending on sample filtration).
  445. #Make sure you remove those same genes from the test data...want to make sure the genes are the same in both train_x and test_x
  446. x<-as.vector(colSums(expression))
  447. bad<-which(x == 0) #Column/genes indices that are filled with only 0's.
  448. if (length(bad) != 0){
  449. expression<-expression[,-bad]
  450. test_x<-data.frame(test_x[,-bad])
  451. }
  452. #Remove genes with 0 variance. If you don't, this will also cause problems in linearRidge().
  453. variance<-c()
  454. for (i in 1:ncol(expression)){
  455. variance[i]<-var(as.vector(expression[,i]))
  456. }
  457. bi<-which(variance %in% 0) #Bad index...gene has 0 variance.
  458. if (length(bi) != 0){ #If there are actually bad indices/genes that had 0 variance across all samples, remove them.
  459. expression<-expression[,-bi]
  460. test_x<-data.frame(test_x[,-bi])
  461. }
  462. #Check to make sure you have more than 1 training sample for the drug's model.
  463. #_______________________________________________________________
  464. trainFrame<-try(data.frame(Resp=trainingPtype4, expression), silent = TRUE)
  465. if (dim(trainFrame)[1] == 1){ #Make sure you have more than 1 sample.
  466. drugs = drugs[-a]
  467. message(paste("\n", drugs[a], "is skipped due to insufficient cell lines to fit the model."))
  468. next
  469. } else {
  470. if(printOutput) message("\nCalculating predicted phenotype...")
  471. rrModel<-linearRidge(Resp ~., data=trainFrame)
  472. preds<-predict(rrModel, newdata=data.frame(test_x))
  473. #You can compute an R^2 value for the data you train on from true and predicted values.
  474. #The R^2 value represents the percentage in which the optimal model accounts for the variance in the training data.
  475. #_______________________________________________________________
  476. if(rsq){
  477. if (dim(expression)[1] < 4){ #The code will result in an error if you have 3 samples...this makes sure you have more than 3 samples.
  478. #It results in an error because there is a 70/30% split of training data.
  479. message(paste("\n", drugs[a], 'is skipped for R^2 analysis'))
  480. } else {
  481. expression<-(cbind(expression, trainingPtype4))
  482. dt<-sort(sample(nrow(expression), nrow(expression)*.7)) #sample() randomly picks 70% of rows/samples from the dataset. It samples without replacement.
  483. #Prepare the training data (70% of original training data)
  484. train_x<-expression[dt,]
  485. ncol<-dim(train_x)[2]
  486. train_y<-train_x[,ncol]
  487. train_x<-train_x[,-ncol]
  488. #Prepare the testing data (30% of original training data)
  489. test_x<-expression[-dt,]
  490. ncol<-dim(test_x)[2]
  491. test_y<-test_x[,ncol]
  492. test_x<-test_x[,-ncol]
  493. #Remove genes that you have no transcriptome data for (aka columns are filled with 0's).
  494. #If you don't, it causes problems in linearRidge().
  495. #Make sure you remove those same genes from the test data...want to make sure the genes are the same in both train_x and test_x
  496. x<-as.vector(colSums(train_x))
  497. bad<-which(x == 0) #Column/genes indices that are filled with only 0's.
  498. if (length(bad) != 0){
  499. train_x<-train_x[,-bad]
  500. test_x<-data.frame(test_x[,-bad])
  501. }
  502. #Remove genes with 0 variance. If you don't, this will also cause problems in linearRidge().
  503. variance<-c()
  504. for (i in 1:ncol(train_x)){
  505. variance[i]<-var(as.vector(train_x[,i]))
  506. }
  507. bi<-which(variance %in% 0) #Bad index...gene has 0 variance.
  508. if (length(bi) != 0){ #If there are actually bad indices/genes that had 0 variance across all samples, remove them.
  509. train_x<-train_x[,-bi]
  510. test_x<-data.frame(test_x[,-bi])
  511. }
  512. trainFrame<-data.frame(Resp=train_y, train_x)
  513. rrModel<-linearRidge(Resp ~., data=trainFrame)
  514. testFrame<-data.frame(test_x)
  515. pred<-predict(rrModel, newdata=testFrame)
  516. if(printOutput) message("\nCalculating R^2...")
  517. sst<-sum((test_y - mean(test_y))^2) #Compute the sum of squares total.
  518. sse<-sum((pred - test_y)^2) #Compute the sum of squares error.
  519. rsq_value<-1 - sse/sst #Compute the rsq value.
  520. }
  521. }
  522. }
  523. }
  524. #If the response variable was transformed (aka powerTransformPhenotype=TRUE), untransform it here.
  525. #_______________________________________________________________
  526. if(powerTransformPhenotype) {
  527. preds <- preds^(1/transForm)
  528. preds <- preds - offset
  529. }
  530. #Find correlation between imputed response and expression of a gene.
  531. #Each gene is corrected differently; therefore, it may not be ideal to determine weight or percentage or weight in which each gene contributes to the prediction.
  532. #_______________________________________________________________
  533. if(cc){ #You can only collect correlations if pcr=FALSE!
  534. if(pcr){
  535. stop('ERROR: pcr must equal FALSE in order to compute correlations') #It doesn't make sense to compute correlations when the features have changed from genes to pcs.
  536. }
  537. if(printOutput) message("\nCalculating correlation coefficients...") #This is only relevant if you aren't using pcr.
  538. cors_vec<-c() #This vector will store the correlation coefficients.
  539. cors_vec2<-c() #This vector will store the p values.
  540. matrix<-homData$test[keepRows,] #Matrix of genes x cell lines/cosmic ids.
  541. for(d in 1:nrow(matrix)){ #For each gene...
  542. cors_vec[d]<-cor.test(as.vector(matrix[d,]), as.vector(preds))$estimate #Compute correlation coefficient for expression of a given gene across patients vs.imputed values for a given drug across patients.
  543. cors_vec2[d]<-cor.test(as.vector(matrix[d,]), as.vector(preds))$p.value
  544. }
  545. #indices<-order(cor, decreasing=TRUE)
  546. #ordered_cor<-cor[indices] #Order the correlation coefficients from big to small.
  547. #genes<-rownames(matrix)
  548. #ordered_genes<-genes[indices]
  549. #names(ordered_cor)<-ordered_genes
  550. }
  551. if(printOutput) message(paste("\nDone making prediction for drug", a, "of", ncol(trainingPtype)))
  552. #Store the data in your lists.
  553. #_______________________________________________________________
  554. DrugPredictions[[a]]<-preds
  555. if(rsq){
  556. rsqs[[a]]<-rsq_value
  557. }
  558. if(cc){
  559. cors[[a]]<-cors_vec
  560. pvalues[[a]]<-cors_vec2
  561. }
  562. }
  563. }
  564. #Time to save the data!
  565. #_______________________________________________________________
  566. #Save drug prediction data to your home directory as a .txt file.
  567. names(DrugPredictions)<-drugs
  568. DrugPredictions_mat<-do.call(cbind, DrugPredictions)
  569. colnames(DrugPredictions_mat)<-drugs
  570. rownames(DrugPredictions_mat)<-colnames(testExprData)
  571. dir.create("./calcPhenotype_Output")
  572. write.csv(DrugPredictions_mat, file="./calcPhenotype_Output/DrugPredictions.csv", row.names = TRUE, col.names = TRUE)
  573. #If rsq=TRUE, save R^2 data.
  574. if(rsq){
  575. names(rsqs)<-drugs
  576. rsqs_mat<-do.call(cbind, rsqs)
  577. dir.create("./calcPhenotype_Output")
  578. write.table(rsqs_mat, file="./calcPhenotype_Output/R^2.txt")
  579. }
  580. #If CC=TRUE, save correlation coefficient data.
  581. if(cc){
  582. names(cors)<-drugs
  583. cor_mat<-do.call(cbind, cors)
  584. rownames(cor_mat)<-rownames(homData$train[keepRows,NonNAindex])
  585. colnames(cor_mat)<-drugs
  586. dir.create("./calcPhenotype_Output")
  587. write.table(cor_mat, file="./calcPhenotype_Output/cors.txt")
  588. names(pvalues)<-drugs
  589. p_mat<-do.call(cbind, pvalues)
  590. rownames(p_mat)<-rownames(homData$train[keepRows, NonNAindex])
  591. colnames(p_mat)<-drugs
  592. dir.create("./calcPhenotype_Output")
  593. write.table(p_mat, file="./calcPhenotype_Output/pvalues.txt")
  594. }
  595. #print(vs)
  596. }

CALCPHENOTYPE.R at commit f7c1944, no license · at the source

Overview

Authors: Siwei Ren1,2,3, Mengchu Xu1, Wenyan Yang1, Bing Wang1, Yixue Li4,5,6,7,8,9, Yin Wang1,10
ORCID iDs: Yixue Li, Yin Wang
  1. Department of Biomedical Engineering, School of Intelligent Medicine, China Medical University, Shenyang, Liaoning Province China
  2. National Genomics Data Center, China National Center for Bioinformation, Beijing, China
  3. Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China
  4. Guangzhou National Laboratory, Guangzhou International Bio Island, Guangzhou, Guangdong Province China
  5. GZMU-GIBH Joint School of Life Sciences, The Guangdong-Hong Kong-Macau Joint Laboratory for Cell Fate Regulation and Diseases, Guangzhou Medical University, Guangzhou, China
  6. Key Laboratory of Systems Health Science of Zhejiang Province, School of Life Science, Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China
  7. School of Life Sciences and Biotechnology, Shanghai Jiao Tong University, Shanghai, China
  8. Shanghai Institute of Nutrition and Health, Chinese Academy of Sciences, Shanghai, China
  9. Bioland Laboratory, Guangzhou, China
  10. Tumor Etiology and Screening Department of Cancer Institute and General Surgery, The First Affiliated Hospital of China Medical University, Shenyang, Liaoning Province China
Journal: Communications medicine, volume 6, issue 1, article 367
Dates: received 12 December 2024; accepted 18 May 2026; published online 24 June 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s43856-026-01689-1 · PMID 42342841 · PMCID PMC13320173 · OpenAlex W7165782009
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: schizophrenia / psychosis (population), clinical / translational (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing, Connectivity, Spectral & time-frequency
Keywords: Computational biology and bioinformatics, Systems biology
Topic: Tryptophan and brain disorders (Biological Psychiatry, Neuroscience), according to OpenAlex
Funding: National Natural Science Foundation of China (32000478, 2023JH2, XDB38050200); Natural Science Foundation of Liaoning Province
Citations: not cited yet (Europe PMC); 193 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above.

maese005/oncoPredict

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: f7c19444bf98f112a4a6ed3c1a4da5788eb4f535, 8 September 2023
Languages: R (7)
Size: 28 files, 7 scripts
Software Heritage: not archived
Found in: the text, “Drug-gene interaction analysis”
Holds: README, environment (DESCRIPTION), documentation, 4 notebooks
Not found: license file, CITATION.cff, tests, continuous integration
Tools: car (2 files), glmnet (2 files), tidyverse (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
8 files

figshare 32204805

License: CC-BY-4.0
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Size: 1 file, 0 scripts
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
  • 27 September 2026: the link answers (HTTP 200)
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s43856-026-01689-1.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 7 scripts, each with its path and the digest of its content;
  • no match between paragraphs and code yet;
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Other data links

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • it points to a dataset: NCBI

Read it in the paper: doi.org/10.1038/s43856-026-01689-1.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Funding: added National Natural Science Foundation of China: 32000478, 2023JH2, XDB38050200; Natural Science Foundation of Liaoning Province

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 2 keywords, 192 references.

Cite

This paper

Ren, S., Xu, M., Yang, W., Wang, B., Li, Y., & Wang, Y. (2026). Comprehensive analysis of interactions between brain aging and late-onset psychoses using heuristic mapping models. Communications medicine, 6(1), 367. https://doi.org/10.1038/s43856-026-01689-1

BibTeX

@article{ren2026comprehensive,
author = {Ren, Siwei and Xu, Mengchu and Yang, Wenyan and Wang, Bing and Li, Yixue and Wang, Yin},
title = {{Comprehensive analysis of interactions between brain aging and late-onset psychoses using heuristic mapping models}},
journal = {Communications medicine},
year = {2026},
month = jun,
volume = {6},
number = {1},
pages = {367},
publisher = {Nature Publishing Group},
issn = {2730-664X},
doi = {10.1038/s43856-026-01689-1},
url = {https://doi.org/10.1038/s43856-026-01689-1},
pmid = {42342841},
pmcid = {PMC13320173}
}

RIS

TY - JOUR
AU - Ren, Siwei
AU - Xu, Mengchu
AU - Yang, Wenyan
AU - Wang, Bing
AU - Li, Yixue
AU - Wang, Yin
TI - Comprehensive analysis of interactions between brain aging and late-onset psychoses using heuristic mapping models
T2 - Communications medicine
J2 - Commun Med (Lond)
PY - 2026
DA - 2026/06/24
VL - 6
IS - 1
SP - 367
SN - 2730-664X
PB - Nature Publishing Group
DO - 10.1038/s43856-026-01689-1
UR - https://doi.org/10.1038/s43856-026-01689-1
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s43856-026-01689-1",
"type": "article-journal",
"title": "Comprehensive analysis of interactions between brain aging and late-onset psychoses using heuristic mapping models",
"container-title": "Communications medicine",
"author": [
{
"family": "Ren",
"given": "Siwei"
},
{
"family": "Xu",
"given": "Mengchu"
},
{
"family": "Yang",
"given": "Wenyan"
},
{
"family": "Wang",
"given": "Bing"
},
{
"family": "Li",
"given": "Yixue"
},
{
"family": "Wang",
"given": "Yin"
}
],
"container-title-short": "Commun Med (Lond)",
"volume": "6",
"issue": "1",
"page": "367",
"DOI": "10.1038/s43856-026-01689-1",
"PMID": "42342841",
"PMCID": "PMC13320173",
"ISSN": "2730-664X",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s43856-026-01689-1",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
24
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.bbi.2026.106598
Choroid plexus inflammation in bipolar disorder.
Journal: Brain, behavior, and immunity
In common: 4 references
[2] doi:10.1016/j.isci.2026.115657 [code]
Integration of machine learning to develop a disulfidptosis model for predicting glioma prognosis, immunotherapy response, and drug.
Journal: iScience
In common: glmnet, car, tidyverse, clinical / translational
[3] doi:10.1002/hbm.70496 [code]
Transdiagnostic Profiles of BOLD Signal Variability in Autism and Schizophrenia Spectrum Disorders: Associations With Cognition and Functioning.
Journal: Human brain mapping
In common: car, tidyverse, schizophrenia / psychosis, 1 reference
[4] doi:10.64898/2026.03.09.26347914 [code]
Multimodal Ageing Biomarkers and Plasma Proteomic Signatures Associated with All-Cause Mortality
Journal: medRxiv (preprint)
In common: glmnet, tidyverse, clinical / translational, 1 reference
[5] doi:10.1093/schbul/sbaf091 [code]
Setd1a Loss-of-function Disrupts Epigenetic Regulation of Ribosomal Genes via Altered DNA Methylation.
Journal: Schizophrenia bulletin
In common: car, tidyverse, schizophrenia / psychosis, 1 reference
[6] doi:10.1126/sciadv.aec9291 [code]
Computational mechanisms of perception in autism revealed using games inspired by rodent operant tasks.
Journal: Science advances
In common: glmnet, car, tidyverse
[7] doi:10.1016/j.isci.2026.116601 [code]
Random auditory stimulation during sleep disturbs traveling slow waves and declarative memory.
Journal: iScience
In common: glmnet, car, tidyverse
[8] doi:10.1038/s41398-026-04142-y [code]
Phocaeicola vulgatus improves anxiety-like behavior by ameliorating amygdala neuroinflammation and the neurite impairment in IBS.
Journal: Translational psychiatry
In common: glmnet, car, tidyverse
[9] doi:10.1186/s12916-026-04869-x [code]
The tissue-specific effects of glucose-lowering drug targets on aging mediated through DNA methylation: a multi-omics genetic study.
Journal: BMC medicine
In common: glmnet, car, tidyverse
[10] doi:10.1016/j.isci.2026.115982 [code]
Multi-omics profiling-derived signature links cellular ecosystem to glioblastoma prognosis.
Journal: iScience
In common: glmnet, car, tidyverse

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.