OSCR

Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § Materials and methods › Identifying and analyzing transcriptional regulatory modules ↔ eGRAMv2R1.py, lines 1–49 · score 0.98 · clustering algorithms, regulatory modules, eGRAM, transcriptionally regulate, independently regulated, co regulated
  2. [2] § Appendix 8 › The analysis of HS TFs and their DBSs ↔ eGRAMv2R1.py, lines 1–49 · score 0.79 · CellOracle, Pearson correlation, TF target, database, gene expression, intersection

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 776 lines · 41 KB · MIT · 2 matches

  1. '''
  2. Several notes:
  3. 1. There are two kinds of input files - (a) user data and (b) system data. The former has three files - (a) expression profile, (b) lncRNA target
  4. prediction, (c) TF target prediction. LncRNAs' targets are predicted using LongTarget (or the more rapid version Farsim); TFs' targets are
  5. predicted using any method, such as CellOracle. Since the primary goal of eGRAM is analyzing lncRNA function in transcriptional regulation, the
  6. TF target prediction file is optional. System data include human/mouse KEGG and wikipathway pathways downloaded from the KEGG and Wikipathways
  7. websites.
  8. 2. All variables begin with their data type: df_, list_, dict_, dict2D_, set_.
  9. 3. It is the user's responsibility to determine correct gene symbols and gene types, and keep gene symbols and types consistent in the input files.
  10. In the given examples, targets of lncRNAs and TFs are only lncRNA genes, protein-coding genes, and TF genes. The user can explore other gene
  11. types (but they are barely annotated in pathway databases).
  12. 4. eGRAM performs transcription analysis based on both correlated gene expression and lncRNA/TF target prediction. For small sample size, Pearson
  13. correlation is recommended, for large sample size, Spearman correlation can also be chosen.
  14. 5. eGRAM identifies "all-mutually-correlated" lncRNAs and "all-mutually-correlated" TFs from lncRNAs and TFs in the gene expression file. We call
  15. this method "COCO", which is akin to clustering genes into clusters allowing overlap. Clustering with overlapping is reasonable because
  16. a lncRNA may work with different partners in different situations to regulate transcription, so do TFs. Two command-line arguments,
  17. --lncCorr and --tfCorr, work with the option. Since there are many clustering algorithm, future versions will allow the use of other algorithms
  18. (e.g. FLAME).
  19. 6. To make eGRAM able to handle lncRNAs and TFs of any names in any numbers, we organize lncRNAs and TFs into dictionaries, each lncRNA/TF is the
  20. key and its COCO set (a list) is the value.
  21. 7. Several pending issues
  22. (1) The relations between regulators and targets may be (a) co-upregulation, (b) co-downregulation, (c) differential expression - regulator high
  23. and target low, (d) differential expression - regulator low and target high. --- we have not explore this???
  24. (2) Whether a cutoff is needed for controlling the smallest COCO set size. For example, a set of <=3 lncRNAs or a set of <=2 TFs may not indicate
  25. a reliable regulator set.
  26. (3) There are three kinds of modules - lncRNAs' target sets, TFs' target sets, and merged target sets (intersection and union). Pathway enrichment
  27. analysis can be applied to each of the sets to reveal whether lncRNAs' target sets and TFs' target sets are enriched in different pathways.
  28. (4) Examining the intersection or union of merged sets is also biologically different, indicating co-regulation coordinated regulation by lncRNAs
  29. and TFs, respectively.
  30. 8. A gene with all 0s is removed, but otherwise kept. 0s in RNA-seq datasets are assumed true value, so no imputation is performed. If scRNA-seq data
  31. is analyzed, it is the user's responsibility to perform imputation using any method. We assume that imputation is finished before running eGRAM.
  32. Main running steps:
  33. step 1 - read in user files and command line arguments
  34. step 2 - read in kegg files and wikipathway files
  35. step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
  36. step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
  37. step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
  38. step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
  39. step 7 - identify typically co-regulated modules and independently-regulated modules
  40. step 8 - perform pathway enrichment analysis for module genes in df_path_module
  41. step 9 - write results to files
  42. step 10 - generate files for cytoscape
  43. An explanatory command line:
  44. eGRAMv2R1.py --exp 2024May-DEG-exp-A549-2513WT.csv --lncDBS 2024May-lncRNA-DBS-A549-2513.csv --tfDBS 2024May-TF-DBS-A549-2513.csv --species 1
  45. --lncCutoff 100 --tfCutoff 8 --lncCorr 0.6 --tfCorr 0.6 --moduleSize 50 --corr Pearson --cluster COCO --fdr 0.01 --out outA5492513WT
  46. The user can output intermediate results to a specific file by adding "> file_name" at the end of the command line in ther Terminal mode.
  47. '''
  48. #!/usr/bin/python
  49. import numpy as np
  50. import pandas as pd
  51. import re
  52. import os
  53. import csv
  54. import sys, getopt
  55. from scipy.stats import hypergeom
  56. #from sklearn.impute import SimpleImputer
  57. ########################################################################################################################
  58. # step 2 - read in kegg files and wikipathway files
  59. def read_kegg(geneFile, pathFile, linkFile):
  60. '''
  61. Read kegg genes, pathways and pathway-gene links.
  62. '''
  63. gene_symbols = {} # debug info: gene_symbols={dict:0} {}, 0 indicating no element.
  64. kegg_pathway = {} # debug info: gene_symbols={dict:0} {}.
  65. kegg_link = {}
  66. gene_id_set = set() # debug info: gene_id_set={set:0} {}.
  67. with open(geneFile) as f:
  68. for line in f: # debug: first = 'hsa:102466751\tmiRNA\t1:complement(17369..17436)\tMIR6859-1, hsa-mir-6859-1; microRNA 6859-1\n'
  69. cols = line.strip('\r\n').split('\t') # debug: ['hsa:102466751', 'miRNA', '1:complement(17369..17436)', 'MIR6859-1, hsa-mir-6859-1; microRNA 6859-1']
  70. gene_id = cols[0] # debug: gene_id = 'hsa:102466751', the 1st element in cols, i.e., cols[0].
  71. descri = cols[-1] # descri is the last element of cols, e.g., 'MIR6859-1, hsa-mir-6859-1; microRNA 6859-1'.
  72. if (',' in descri) or (';' in descri):
  73. symbol_descri = descri.split(';')[0]
  74. symbols = symbol_descri.split(", ")
  75. if gene_id not in gene_symbols: #
  76. gene_symbols[gene_id] = set()
  77. for symbol in symbols:
  78. gene_symbols[gene_id].add(symbol)
  79. gene_id_set.add(gene_id)
  80. with open(pathFile) as f:
  81. for line in f:
  82. pathID, pathwayName = line.strip('\r\n').split('\t')
  83. pathID = pathID.replace("path:", '')
  84. pathwayName = pathwayName[:pathwayName.rfind(' -')]
  85. kegg_pathway[pathID] = pathwayName
  86. with open(linkFile) as f:
  87. for line in f:
  88. pathID, gene_id = line.strip('\r\n').split('\t')
  89. pathID = pathID.replace("path:", '')
  90. if pathID in kegg_pathway:
  91. if gene_id in gene_id_set:
  92. if pathID not in kegg_link:
  93. kegg_link[pathID] = set()
  94. kegg_link[pathID].add(gene_id)
  95. totalGeneNum = len(gene_id_set) # totalGeneNum = 22101
  96. kegg_gene = {}
  97. for gene_id in gene_symbols:
  98. for gene_symbol in gene_symbols[gene_id]:
  99. kegg_gene[gene_symbol] = gene_id
  100. return kegg_gene, kegg_pathway, kegg_link, totalGeneNum
  101. def read_wikip(wikiDir, kegg_gene):
  102. '''
  103. Read genes and pathways in Wikipathways.
  104. '''
  105. wiki_pathway = {}
  106. wiki_link = {}
  107. pathway_files = os.listdir(wikiDir)
  108. pathway_comp = re.compile(' Name="(.+?)" ')
  109. gene_comp = re.compile(' TextLabel="(.+?)" ')
  110. for pathway_file in pathway_files:
  111. pathway_id = pathway_file.split('_')[-2]
  112. with open(wikiDir + pathway_file, encoding='utf-8') as f:
  113. for line in f:
  114. pathway_re = re.search(pathway_comp, line)
  115. if pathway_re:
  116. pathway = pathway_re.group(1)
  117. wiki_pathway[pathway_id] = pathway_re.group(1)
  118. gene_re = re.search(gene_comp, line)
  119. if gene_re:
  120. symbol = gene_re.group(1)
  121. if symbol in kegg_gene:
  122. gene_id = kegg_gene[symbol]
  123. if pathway_id in wiki_pathway:
  124. if pathway_id not in wiki_link:
  125. wiki_link[pathway_id] = set()
  126. wiki_link[pathway_id].add(gene_id)
  127. return wiki_pathway, wiki_link
  128. ########################################################################################################################
  129. # step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
  130. def generate_lncRNA_files(df_gene_sample_exp):
  131. df_lncRNA_sample_exp = df_gene_sample_exp.copy(deep=True)
  132. # When deep=True (default), a new object is created.
  133. df_lncRNA_sample_exp = df_lncRNA_sample_exp.drop(df_lncRNA_sample_exp[df_lncRNA_sample_exp['genetype'] != 'lncRNA'].index)
  134. # dropping non-lncRNA rows.
  135. df_lncRNA_sample_exp = df_lncRNA_sample_exp.reset_index()
  136. # "drop" operation does change the original index.
  137. # This .reset_index() re-sets index (also keeps the original index in df_gene_sample_exp).
  138. df_lncRNA_sample_exp.drop(["index"], axis=1, inplace=True)
  139. # drop the old index.
  140. list_lncRNA = df_lncRNA_sample_exp['genesymbol_GRCh38']
  141. df_lncRNA_sample_exp.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
  142. df_sample_lncRNA_exp = df_lncRNA_sample_exp.T
  143. # correlation is computed between columns.
  144. df_lncRNA_lncRNA_corr = df_sample_lncRNA_exp.corr(method='pearson')
  145. df_lncRNA_lncRNA_corr.columns = list_lncRNA
  146. # This replaces column names (but not column content) with lncRNA names in list_lncRNA (i.e., re-index columns).
  147. df_gene_sample_exp2 = df_gene_sample_exp.copy(deep=True)
  148. # df_gene_sample_exp2 is a temporary dataframe
  149. list_gene = df_gene_sample_exp2['genesymbol_GRCh38']
  150. list_type = df_gene_sample_exp2['genetype']
  151. df_gene_sample_exp2.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
  152. df_sample_gene_exp = df_gene_sample_exp2.T
  153. print("\nafter transferred\n", df_sample_gene_exp)
  154. if arg_Pearson == 'Spearman':
  155. df_gene_gene_corr = df_sample_gene_exp.corr(method='spearman')
  156. else:
  157. df_gene_gene_corr = df_sample_gene_exp.corr(method='pearson')
  158. print(df_gene_gene_corr)
  159. df_gene_gene_corr.columns = list_gene
  160. # Using gene names in list_gene to replace the 0-N column names.
  161. print("\nafter putting column names\n", df_gene_gene_corr)
  162. for j in range(len(list_gene)):
  163. if list_type[j] != 'lncRNA':
  164. df_gene_gene_corr.drop(columns=list_gene[j], inplace=True)
  165. print("\nafter deleting non-lncRNA columns\n", df_gene_gene_corr)
  166. df_lncRNA_gene_corr = df_gene_gene_corr.copy(deep=True)
  167. df_lncRNA_gene_corr.index = list_gene
  168. return df_lncRNA_lncRNA_corr, df_lncRNA_gene_corr, list_lncRNA
  169. def generate_lncRNA_sets(list_lncRNA, df_lncRNA_lncRNA_corr):
  170. dict_lncRNAsets = {}
  171. # For each lncRNA in list_lncRNA, a dict is generated with the lncRNA as the key and a list as the value.
  172. # The list contains all-mutually-correlated lncRNAs of this lncRNA.
  173. # Process lncRNAs in df_lncRNA_lncRNA_corr row by row (lncRNAs in list_lncRNA and df_lncRNA_lncRNA_corr have the same order).
  174. for row in df_lncRNA_lncRNA_corr.index:
  175. deleted = []
  176. for col in range(0, len(list_lncRNA)):
  177. if df_lncRNA_lncRNA_corr.iat[row, col] < arg_LncRNACorrCutoff:
  178. df_lncRNA_lncRNA_corr.iat[row, col] = -100 # -100 is a number for marking the current column
  179. df_lncRNA_lncRNA_corr.iat[col, row] = -100 # -100 is a number for marking the corresponding row(diagonal elements are equal)
  180. deleted.append(list_lncRNA[col]) # Put this lncRNA to deleted[]
  181. # For lncRNAs with high pairwise corr with the current lncRNA (row), the two-for-loop check/ensure all-mutually-correlated lncRNAs.
  182. for i in range(0, len(list_lncRNA)):
  183. if i != row and list_lncRNA[i] not in deleted:
  184. for j in range(0, len(list_lncRNA)):
  185. if list_lncRNA[j] not in deleted:
  186. if df_lncRNA_lncRNA_corr.iat[i, j] < arg_LncRNACorrCutoff:
  187. df_lncRNA_lncRNA_corr.iat[i, j] = -100 # -100 for marking the newly identified invalid lncRNA
  188. deleted.append(list_lncRNA[j]) # Put the newly identified invalid lncRNA into deleted[]
  189. else:
  190. df_lncRNA_lncRNA_corr.iat[i, j] = -100 # If the col name is already in deleted[], we still set the value to -100.
  191. lncRNAset = set(list_lncRNA) - set(deleted)
  192. if list_lncRNA[row] not in deleted:
  193. lncRNAlist = sorted(list(lncRNAset))
  194. else:
  195. lncRNAlist = []
  196. dict_lncRNAsets.update({list_lncRNA[row]: lncRNAlist})
  197. print(row, {list_lncRNA[row]: lncRNAlist})
  198. return dict_lncRNAsets
  199. ########################################################################################################################
  200. # step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
  201. def generate_TF_files(df_gene_sample_exp):
  202. df_TF_sample_exp = df_gene_sample_exp.copy(deep=True)
  203. # When deep=True (default), a new object is created with a copy of the calling object’s data and indices.
  204. df_TF_sample_exp = df_TF_sample_exp.drop(df_TF_sample_exp[df_TF_sample_exp['genetype'] != 'TF'].index)
  205. # drop non-TF rows.
  206. df_TF_sample_exp = df_TF_sample_exp.reset_index()
  207. # "drop" operation does change the original index.
  208. # This .reset_index() re-sets index and also keeps the original index in df_gene_sample_exp.
  209. df_TF_sample_exp.drop(["index"], axis=1, inplace=True)
  210. # drop the old index.
  211. list_TF = df_TF_sample_exp['genesymbol_GRCh38']
  212. # get all TFs in the first column that are gene names (this list also obtains the column name "genesymbol_GRCh38")
  213. df_TF_sample_exp.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
  214. # drop "genesymbol_GRCh38" and "genetype", because all fileds should be numeric to generate the corr dataframe.
  215. df_sample_TF_exp = df_TF_sample_exp.T
  216. # correlation is computed between columns.
  217. df_TF_TF_corr = df_sample_TF_exp.corr(method='pearson')
  218. df_TF_TF_corr.columns = list_TF
  219. # This replaces column names (but not column content) with TF names in list_TF (i.e., re-index columns).
  220. df_gene_sample_exp2 = df_gene_sample_exp.copy(deep=True)
  221. # df_gene_sample_exp2 is a temporary dataframe
  222. list_gene = df_gene_sample_exp2['genesymbol_GRCh38']
  223. list_type = df_gene_sample_exp2['genetype']
  224. df_gene_sample_exp2.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
  225. df_sample_gene_exp = df_gene_sample_exp2.T
  226. print("\nafter transferred\n", df_sample_gene_exp)
  227. if arg_Pearson == 'Spearman':
  228. df_gene_gene_corr = df_sample_gene_exp.corr(method='spearman')
  229. else:
  230. df_gene_gene_corr = df_sample_gene_exp.corr(method='pearson')
  231. print(df_gene_gene_corr)
  232. df_gene_gene_corr.columns = list_gene
  233. # Using gene names in list_gene to replace the 0-N column names.
  234. print("\nafter putting column names\n", df_gene_gene_corr)
  235. for j in range(len(list_gene)):
  236. if list_type[j] != 'TF':
  237. df_gene_gene_corr.drop(columns=list_gene[j], inplace=True)
  238. print("\nafter deleting non-TF columns\n", df_gene_gene_corr)
  239. df_TF_gene_corr = df_gene_gene_corr.copy(deep=True)
  240. df_TF_gene_corr.index = list_gene
  241. return df_TF_gene_corr, df_TF_TF_corr, list_TF
  242. def generate_TF_sets(list_TF, df_TF_TF_corr):
  243. dict_TFsets = {}
  244. # For each TF in list_TF, a dict is generated with the TF as key and a list as the value.
  245. # The list contains all-mutually-correlated TFs of this TF.
  246. # Process TFs in df_TF_TF_corr row by row. Note TFs in list_TF and df_TF_TF_corr have the same order.
  247. for row in df_TF_TF_corr.index: # This row-for-loop check all TFs (rows)
  248. deleted = []
  249. for col in range(0, len(list_TF)):
  250. if df_TF_TF_corr.iat[row, col] < arg_TFCorrCutoff:
  251. df_TF_TF_corr.iat[row, col] = -100
  252. df_TF_TF_corr.iat[col, row] = -100
  253. deleted.append(list_TF[col])
  254. # For TFs with high pairwise corr with the current TF (row), the two-for-loop check/ensure all-mutually-correlated TFs.
  255. for i in range(0, len(list_TF)):
  256. if i != row and list_TF[i] not in deleted:
  257. for j in range(0, len(list_TF)):
  258. if list_TF[j] not in deleted:
  259. if df_TF_TF_corr.iat[i, j] < arg_TFCorrCutoff:
  260. df_TF_TF_corr.iat[i, j] = -100
  261. deleted.append(list_TF[j])
  262. else:
  263. df_TF_TF_corr.iat[i, j] = -100
  264. TFset = set(list_TF) - set(deleted)
  265. if list_TF[row] not in deleted:
  266. TFlist = sorted(list(TFset))
  267. else:
  268. TFlist = []
  269. dict_TFsets.update({list_TF[row]: TFlist})
  270. print(row, {list_TF[row]: TFlist})
  271. return dict_TFsets
  272. ########################################################################################################################
  273. # step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
  274. def generate_lncRNA_target(df_lncRNA_gene_DBS, df_lncRNA_gene_corr, dict_lncRNAsets):
  275. '''
  276. Upon df_lncRNA_gene_corr and df_lncRNA_gene_DBS, we get targets of every lncRNA set, which are the
  277. intersection of set_lncRNAtarget_DBS (DBS>cutoff1) and set_lncRNAtarget_CORR (|corr|>cutoff2).
  278. cutoff2 (lncRNA-gene-corr) = lncRNA-lncRNA-corr (arg_LncRNACorrCutoff). A separable cutoff can be considered.
  279. :
  280. traverse key in dict_lncRNAsets
  281. traverse lncRNA in lncRNAlist
  282. get targets of lncRNA from df_lncRNA_gene_DBS, append targets to dict_lncRNAtarget_DBS
  283. get targets of lncRNA from df_lncRNA_gene_CORR, append targets to dict_lncRNAtarget_CORR
  284. get overlap of the two target sets, append the overlap to dict_lncRNAtarget_shared
  285. '''
  286. list_lncRNACORR = df_lncRNA_gene_corr.columns
  287. list_lncRNADBS = df_lncRNA_gene_DBS.columns
  288. dict_lncRNAtarget_DBS = {}
  289. dict_lncRNAtarget_CORR = {}
  290. dict_lncRNAtarget_shared = {}
  291. rows1 = df_lncRNA_gene_corr.shape[0]
  292. cols1 = df_lncRNA_gene_corr.shape[1]
  293. rows2 = df_lncRNA_gene_DBS.shape[0]
  294. cols2 = df_lncRNA_gene_DBS.shape[1]
  295. for key, value in dict_lncRNAsets.items():
  296. set_tpmCORR = set()
  297. set_tmpDBS = set()
  298. if key in list_lncRNACORR and value != []: # if the COCO set of this lncRNA is non-empty
  299. for lncRNA in value: # then process all lncRNAs in the lncRNAlist
  300. if lncRNA in df_lncRNA_gene_corr.columns:
  301. for target in df_lncRNA_gene_corr.index:
  302. if abs(df_lncRNA_gene_corr[lncRNA][target]) > arg_LncRNACorrCutoff:
  303. set_tpmCORR.add(target)
  304. dict_lncRNAtarget_CORR.update({key: set_tpmCORR})
  305. if key in list_lncRNADBS and value != []:
  306. for lncRNA in value:
  307. if lncRNA in df_lncRNA_gene_DBS.columns:
  308. for target in df_lncRNA_gene_DBS.index:
  309. if df_lncRNA_gene_DBS[lncRNA][target] > arg_LncRNADBSCutoff:
  310. set_tmpDBS.add(target)
  311. dict_lncRNAtarget_DBS.update({key: set_tmpDBS})
  312. # dict_lncRNAtarget_CORR and dict_lncRNAtarget_DBS have the same keys and key order.
  313. for key, value1 in dict_lncRNAtarget_DBS.items():
  314. if key in dict_lncRNAtarget_CORR:
  315. value2 = dict_lncRNAtarget_CORR[key]
  316. shared = value1 & value2
  317. dict_lncRNAtarget_shared.update({key: shared})
  318. return dict_lncRNAtarget_shared
  319. ########################################################################################################################
  320. # step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
  321. def generate_TF_target(df_TF_gene_DBS, df_TF_gene_corr, dict_TFsets):
  322. list_TFCORR = df_TF_gene_corr.columns
  323. list_TFDBS = df_TF_gene_DBS.columns
  324. dict_TFtarget_DBS = {}
  325. dict_TFtarget_CORR = {}
  326. dict_TFtarget_shared = {}
  327. rows1 = df_TF_gene_corr.shape[0]
  328. cols1 = df_TF_gene_corr.shape[1]
  329. rows2 = df_TF_gene_DBS.shape[0]
  330. cols2 = df_TF_gene_DBS.shape[1]
  331. for key, value in dict_TFsets.items():
  332. set_tmpCORR = set()
  333. set_tmpDBS = set()
  334. if key in list_TFCORR and value != []:
  335. for TF in value:
  336. if TF in df_TF_gene_corr.columns:
  337. for target in df_TF_gene_corr.index:
  338. if abs(df_TF_gene_corr[TF][target]) > arg_TFCorrCutoff:
  339. set_tmpCORR.add(target)
  340. dict_TFtarget_CORR.update({key: set_tmpCORR})
  341. if key in list_TFDBS and value != []:
  342. for TF in value:
  343. if TF in df_TF_gene_DBS.columns:
  344. for target in df_TF_gene_DBS.index:
  345. if df_TF_gene_DBS[TF][target] > arg_TFDBSCutoff:
  346. set_tmpDBS.add(target)
  347. dict_TFtarget_DBS.update({key: set_tmpDBS})
  348. # dict_TFtarget_CORR and dict_TFtarget_DBS have the same keys and key order.
  349. for key, value1 in dict_TFtarget_DBS.items():
  350. if key in dict_TFtarget_CORR:
  351. value2 = dict_TFtarget_CORR[key]
  352. shared = value1 & value2
  353. dict_TFtarget_shared.update({key: shared})
  354. return dict_TFtarget_shared
  355. ########################################################################################################################
  356. # step 7 - identify modules and merge lncRNAs/TFs that regulate the same modules
  357. def collect_modules(dict_lncRNAtarget_shared, dict_TFtarget_shared, dict_lncRNAsets, dict_TFsets):
  358. '''
  359. This function collects all lncRNAs' modules (lncRNA target sets) and TFs' modules (TF target sets) into a dict,
  360. with target genes in the module (a tuple) as key and lncRNAs/TFs that co-regulate the module as values.
  361. '''
  362. dict_module_regulators = {}
  363. dict_module_regulator_sets = {}
  364. # zhu
  365. # dict_lncRNAtarget_shared = {'AC090152.1': {'LRMDA', 'FGFR2', 'ABLIM1', 'TNFSF9'}, 'LINC02783': set(), 'Z99572.1': set(), 'AC010889.1': {'GABRA3', 'ABLIM1'}, 'ZNF790-AS1': {'COL18A1-AS1', 'MED14'}, 'AC125807.2': set()}
  366. # dict_TFtarget_shared = {'ZNF790': {'GABRA3', 'FGFR2', 'CEACAM6', 'TNFSF9'}, 'SOX21': {'LINC02577', 'COL18A1-AS1', 'MED14'}, 'ALX1': {'DENND2A', 'TSPAN7'}, 'HNF4A': {'DENND2A', 'TSPAN7'}, 'NR5A2': set()}
  367. #dict_xx_shared = {'1': {'1', '2', '3', '4'}, '5': set(), '6': set(), '7': {'8', '9'}, '10': {'11', '12'}, '13': set()}
  368. #dict_zz_shared = {'a': {'b', 'd', 'e', 'f'}, 'g': {'h', 'i', 'j'}, 'k': {'l', 'm'}, 'n': {'o', 'p'}, 'q': set()}
  369. #allModules = [tuple(sorted(module)) for module in list(dict_xx_shared.values()) + list(dict_zz_shared.values()) if len(module) > 1]
  370. # We sort the order of target genes in modules so that we can conveniently compare and determine whether two modules are the same
  371. allModules = [tuple(sorted(module)) for module in list(dict_lncRNAtarget_shared.values()) + list(dict_TFtarget_shared.values()) if len(module) > arg_ModuleSize]
  372. for module in allModules:
  373. dict_module_regulators[module] = set()
  374. dict_module_regulator_sets[module] = []
  375. for lncRNA in dict_lncRNAtarget_shared:
  376. module = tuple(sorted(dict_lncRNAtarget_shared[lncRNA]))
  377. # dict_lncRNAtarget_shared has 102 elements, each has a key (a lncRNA) and a value (the set of lncRNAs corresponding to the key)
  378. # module is the key, now the key of lncRNA set is the value, better to use the value of lncRNA set as the value
  379. if module in dict_module_regulators:
  380. dict_module_regulators[module].add(lncRNA + '(*)') # we use * to indicate lncRNA
  381. regulator_set = dict_lncRNAsets[lncRNA]
  382. if regulator_set and regulator_set not in dict_module_regulator_sets[module]:
  383. dict_module_regulator_sets[module].append(regulator_set)
  384. for TF in dict_TFtarget_shared:
  385. module = tuple(sorted(dict_TFtarget_shared[TF]))
  386. if module in dict_module_regulators:
  387. dict_module_regulators[module].add(TF + '(#)') # we use # to indicate TF
  388. regulator_set = dict_TFsets[TF]
  389. if regulator_set and regulator_set not in dict_module_regulator_sets[module]:
  390. dict_module_regulator_sets[module].append(regulator_set)
  391. print("dict_module_regulators=", dict_module_regulators)
  392. return dict_module_regulators, dict_module_regulator_sets
  393. ########################################################################################################################
  394. # step 8 - perform pathway enrichment analysis for module genes ### in df_path_module
  395. def pathway_analysis(dict_module_regulators, kegg_gene, kegg_pathway, kegg_link, totalGeneNum, arg_fdrCutoff):
  396. dict_module_kegg = {}
  397. results = []
  398. pvalues = []
  399. for module in dict_module_regulators:
  400. module_geneIDs = {}
  401. #a gene may have serveral symbols
  402. #we convert its symbol which user assign to KEGG geneID to obtain the interset between the module and a pathway
  403. for gene in module:
  404. if gene in kegg_gene:
  405. geneID = kegg_gene[gene]
  406. module_geneIDs[geneID] = gene
  407. moduleSize = len(module_geneIDs)
  408. for pathID in kegg_link:
  409. interSet = set(module_geneIDs.keys()) & kegg_link[pathID]
  410. if interSet:
  411. hitNum = len(interSet)
  412. thisPathGeneNum = len(kegg_link[pathID])
  413. # Indentify significantly enriched pathways using hypergeometric distribution test
  414. pVal = hypergeom.sf(hitNum-1, totalGeneNum, thisPathGeneNum, moduleSize)
  415. if pVal<0.05:
  416. pathName = kegg_pathway[pathID]
  417. hitGenes = set([module_geneIDs[geneID] for geneID in interSet])
  418. results.append([module, pathID, pathName, hitGenes, pVal])
  419. pvalues.append(pVal)
  420. #Benjamini-Hochberg method for p-value correction
  421. pValues = np.asfarray(pvalues)
  422. by_descend = pValues.argsort()[::-1]
  423. by_orig = by_descend.argsort()
  424. steps = float(len(pValues)) / np.arange(len(pValues), 0, -1)
  425. fdr_values = np.minimum(1, np.minimum.accumulate(steps * pValues[by_descend]))
  426. fdrs = []
  427. for i in range(len(fdr_values[by_orig])):
  428. fdrs.append(fdr_values[by_orig][i])
  429. for i in range(len(fdrs)):
  430. results[i].append(fdrs[i])
  431. for module, pathID, pathName, hitGenes, pVal, fdr in results:
  432. if fdr < arg_fdrCutoff:
  433. if module not in dict_module_kegg:
  434. dict_module_kegg[module] = [(fdr, pathID, pathName, hitGenes)]
  435. else:
  436. dict_module_kegg[module].append((fdr, pathID, pathName, hitGenes))
  437. return dict_module_kegg
  438. ########################################################################################################################
  439. # step 9 - write eGRAM results to files in the directory "eGRAM_results/"
  440. def write_modules(dict_module_regulators, dict_module_regulator_sets, dict_module_kegg, dict_module_wiki, outFile):
  441. with open('eGRAM_results/' + outFile + '.csv', 'w', newline='') as f:
  442. writer = csv.writer(f)
  443. writer.writerow(['ModuleID',
  444. 'LncRNA(*)/TF(#)',
  445. 'Regulator set',
  446. 'Target gene',
  447. 'KEGG pathway, pathwayID, and FDR',
  448. 'Hit genes of enriched KEGG pathways',
  449. 'WikiPathway, pathwayID, and FDR',
  450. 'Hit genes of enriched Wikipathways'])
  451. moduleID = 0
  452. for module in dict_module_regulators:
  453. moduleID += 1
  454. if module in dict_module_kegg:
  455. this_module_kegg = []
  456. this_module_kegg_hitGenes = set()
  457. for fdr, pathID, pathName, hitGenes in sorted(dict_module_kegg[module]):
  458. this_module_kegg.append(pathName + ' (' + pathID + ', ' + str('{0:3f}'.format(fdr)) + ')')
  459. this_module_kegg_hitGenes = this_module_kegg_hitGenes | hitGenes
  460. pathway_kegg = '; '.join(this_module_kegg)
  461. hitGenes_kegg = ', '.join(this_module_kegg_hitGenes)
  462. else:
  463. pathway_kegg = hitGenes_kegg = ''
  464. if module in dict_module_wiki:
  465. this_module_wiki = []
  466. this_module_wiki_hitGenes = set()
  467. for fdr, pathID, pathName, hitGenes in sorted(dict_module_wiki[module]):
  468. #this_module_wiki.append(pathName + ' (' + pathID + ', ' + str(fdr) + ')')
  469. this_module_wiki.append(pathName + ' (' + pathID + ', ' + str('{0:3f}'.format(fdr)) + ')')
  470. this_module_wiki_hitGenes = this_module_wiki_hitGenes | hitGenes
  471. pathway_wiki = '; '.join(this_module_wiki)
  472. hitGenes_wiki = ', '.join(this_module_wiki_hitGenes)
  473. else:
  474. pathway_wiki = hitGenes_wiki = ''
  475. regulators = ' | '.join(dict_module_regulators[module])
  476. regulatory_sets = ' | '.join([str(regulatory_set).replace("'",'') for regulatory_set in dict_module_regulator_sets[module]])
  477. targetGenes = ' | '.join(module)
  478. writer.writerow(['ME'+str(moduleID),
  479. regulators,
  480. regulatory_sets,
  481. targetGenes,
  482. pathway_kegg,
  483. hitGenes_kegg,
  484. pathway_wiki,
  485. hitGenes_wiki])
  486. # Generate and write files for module visulization using cytoscape in the 'eGRAM_results/' directory.
  487. def generate_cytosc_file(moduleFile):
  488. if '/' not in moduleFile:
  489. moduleFile = 'eGRAM_results/' + moduleFile
  490. fileName = moduleFile.split('/')[-1]
  491. f_re = open('eGRAM_results/' + fileName + '_edge', 'w')
  492. f_re.write('\t'.join(['lncRNA', 'module', 'weight'])+'\n')
  493. allLncs = set()
  494. allTFs = set()
  495. with open(moduleFile + '.csv') as f_mod:
  496. reader = csv.reader(f_mod)
  497. first_row = next(reader)
  498. for row in reader:
  499. moduleID = row[0]
  500. regulators = row[1].split(' | ')
  501. for regulator in regulators:
  502. regulator_symbol = regulator.split('(')[0]
  503. f_re.write('\t'.join([regulator_symbol, moduleID, 'NA']) + '\n')
  504. f_re.close()
  505. # Merge modules with the same gene sets to avoid redundancy
  506. def remove_redundancy(dict_module_regulators, dict_module_regulator_sets):
  507. dict_module_regulators_noRedundancy = {}
  508. dict_module_regulator_sets_noRedundancy = {}
  509. for module1 in dict_module_regulators:
  510. merge_flag = 0
  511. for module2 in dict_module_regulators_noRedundancy:
  512. set1 = set(module1)
  513. set2 = set(module2)
  514. larger_set = max(set1, set2, key=len)
  515. intersection = set1.intersection(set2)
  516. ratio = len(intersection) / len(larger_set)
  517. if ratio > 0.9:
  518. intersection_module = tuple(sorted(list(intersection)))
  519. dict_module_regulators_noRedundancy[intersection_module] = dict_module_regulators[module1] | dict_module_regulators_noRedundancy[module2]
  520. for regulator_set in dict_module_regulator_sets[module1]:
  521. if regulator_set not in dict_module_regulator_sets_noRedundancy[module2]:
  522. dict_module_regulator_sets_noRedundancy[module2].append(regulator_set)
  523. dict_module_regulator_sets_noRedundancy[intersection_module] = dict_module_regulator_sets_noRedundancy[module2]
  524. del dict_module_regulators_noRedundancy[module2]
  525. del dict_module_regulator_sets_noRedundancy[module2]
  526. merge_flag = 1
  527. break
  528. if merge_flag == 0:
  529. dict_module_regulators_noRedundancy[module1] = dict_module_regulators[module1]
  530. dict_module_regulator_sets_noRedundancy[module1] = dict_module_regulator_sets[module1]
  531. return dict_module_regulators_noRedundancy, dict_module_regulator_sets_noRedundancy
  532. # Modify the resulting file name to prevent overwriting existing files
  533. def check_output(arg_OutFile, resultDir):
  534. maxFileNum = -1
  535. for file in os.listdir(resultDir):
  536. if file==arg_OutFile + '.csv' and maxFileNum==-1:
  537. maxFileNum = 0
  538. elif arg_OutFile+'_' in file:
  539. this_fileNum = file.replace(arg_OutFile+'_', '').replace('.csv', '')
  540. if this_fileNum.isdigit():
  541. this_fileNum = int(this_fileNum)
  542. if this_fileNum > maxFileNum:
  543. maxFileNum = this_fileNum
  544. if maxFileNum == -1:
  545. return
  546. else:
  547. newFileNum = maxFileNum + 1
  548. newOutFile = arg_OutFile + '_' + str(newFileNum) + '.csv'
  549. return newOutFile
  550. ########################################################################################################################
  551. def usage():
  552. print('Usage: python eGRAMv2R1.py --exp gene_exp_file --lncDBS lncRNA_DBS_file --tfDBS TF_DBS_file --species 1 --lncCutoff lncRNA-DBS-cutoff --tfCutoff TF-DBS-cutoff --lncCorr lncRNA-corr-cutoff --tfCorr TF-corr-cutoff --moduleSize modulesize --corr Pearson/Spearman --remove_redundancy y --fdr 0.01 --out output')
  553. print()
  554. print('Example: python eGRAMv2R1.py --exp 2024May-DEG-exp-A549-2513WT.csv --lncDBS 2024May-lncRNA-DBS-A549-2513.csv --tfDBS 2024May-TF-DBS-A549-2513.csv --species 1 --lncCutoff 100 --tfCutoff 8 --lncCorr 0.5 --tfCorr 0.5 --moduleSize 50 --corr Pearson --remove_redundancy y --fdr 0.01 --out myout')
  555. print()
  556. print("""Parameters:
  557. --exp gene_exp_file is required; if genetype does not have TF, all TF-related functions are not performed.
  558. --lncDBS lncRNA_DBS_file is required; integer or 0.
  559. --tfDBS TF_DBS_file is not a must; required if gene_exp_file contains TF.
  560. --species 1=human, 2=mouse.
  561. --lncCutoff the DBS threshold for determining lncRNAs' targets (defult=100).
  562. --tfCutoff the DBS threshold for determining TFs's targets (default=10).
  563. --lncCorr the correlation threshold for lncRNA/target and lncRNA/lncRNA (default=0.5).
  564. --tfCorr the correlation threshold for TF/target and TF/TF (default=0.5).
  565. --moduleSize the minimum gene number of a module (default=50).
  566. --corr Pearson/Spearman (default=pearson).
  567. --remove_redundancy y/n, whether or not merge modules with the same gene sets (default=y)
  568. --fdr the FDR value to determine significance of pathway enrichment (default 0.01).
  569. --out output file name.
  570. """)
  571. ########################################################################################################################
  572. ########################################################################################################################
  573. if __name__ == '__main__':
  574. param = sys.argv[1:]
  575. try:
  576. opts, args = getopt.getopt(param, '-h', ['exp=', 'lncDBS=', 'tfDBS=', 'species=', 'lncCutoff=', 'tfCutoff=', 'lncCorr=', 'tfCorr=', 'moduleSize=', 'corr=', 'remove_redundancy=', 'fdr=', 'out='])
  577. except getopt.GetoptError:
  578. usage()
  579. sys.exit(2)
  580. # step 1 - read in user files and command line arguments
  581. for opt, arg in opts:
  582. if opt == '-h':
  583. usage()
  584. sys.exit(2)
  585. elif opt == '--exp':
  586. df_gene_sample_exp = pd.read_csv(arg)
  587. elif opt == '--lncDBS':
  588. df_lncRNA_gene_DBS = pd.read_csv(arg, index_col = [0])
  589. elif opt == '--tfDBS':
  590. df_TF_gene_DBS = pd.read_csv(arg, index_col = [0])
  591. elif opt == '--species':
  592. arg_SpeciesPath = int(arg)
  593. elif opt == '--lncCutoff':
  594. arg_LncRNADBSCutoff = int(arg)
  595. elif opt == '--tfCutoff':
  596. arg_TFDBSCutoff = int(arg)
  597. elif opt == '--lncCorr':
  598. arg_LncRNACorrCutoff = float(arg)
  599. elif opt == '--tfCorr':
  600. arg_TFCorrCutoff = float(arg)
  601. elif opt == '--moduleSize':
  602. arg_ModuleSize = int(arg)
  603. elif opt == '--corr':
  604. arg_Pearson = arg
  605. elif opt == '--remove_redundancy':
  606. arg_RemoveRedundancy = arg
  607. elif opt == '--fdr':
  608. arg_fdrCutoff = float(arg)
  609. elif opt == '--out':
  610. arg_OutFile = arg
  611. # default arguments
  612. if 'arg_SpeciesPath' not in dir():
  613. arg_SpeciesPath = 1
  614. if 'arg_LncRNADBSCutoff' not in dir():
  615. arg_LncRNADBSCutoff = 100
  616. if 'arg_TFDBSCutoff' not in dir():
  617. arg_TFDBSCutoff = 10
  618. if 'arg_LncRNACorrCutoff' not in dir():
  619. arg_LncRNACorrCutoff = 0.5
  620. if 'arg_TFCorrCutoff' not in dir():
  621. arg_TFCorrCutoff = 0.5
  622. if 'arg_ModuleSize' not in dir():
  623. arg_ModuleSize = 50
  624. if 'arg_Pearson' not in dir():
  625. arg_Pearson = 'Pearson'
  626. if 'remove_redundancy' not in dir():
  627. arg_RemoveRedundancy = 'y'
  628. elif arg_RemoveRedundancy != 'n' and arg_RemoveRedundancy != 'y':
  629. arg_RemoveRedundancy = 'y'
  630. if 'arg_fdrCutoff' not in dir():
  631. arg_fdrCutoff = 0.01
  632. if arg_SpeciesPath == 1:
  633. keggGeneFile = "KEGG/keggGene-hsa"
  634. keggPathFile = "KEGG/keggPathway-hsa"
  635. keggLinkFile = "KEGG/keggLink-hsa"
  636. wikiPath = 'WikiPathways/wikipathways-20240310-gpml-Homo_sapiens/'
  637. elif arg_SpeciesPath == 2:
  638. keggGeneFile = "KEGG/keggGene-mmu"
  639. keggPathFile = "KEGG/keggPathway-mmu"
  640. keggLinkFile = "KEGG/keggLink-mmu"
  641. wikiPath = 'WikiPathways/wikipathways-20240310-gpml-Mus_musculus/'
  642. else:
  643. sys.exit(2)
  644. ########################################################################################################################
  645. # step 2 - read in kegg files and wikipathway files
  646. kegg_gene, kegg_pathway, kegg_link, totalGeneNum = read_kegg(keggGeneFile, keggPathFile, keggLinkFile)
  647. wiki_pathway, wiki_link = read_wikip(wikiPath, kegg_gene)
  648. resultDir = 'eGRAM_results/'
  649. if not os.path.exists(resultDir):
  650. os.mkdir(resultDir)
  651. # step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
  652. # i.e., clustering lncRNAs into 'COCO' sets (this could be done by introduing an outside program such as FLAME)
  653. df_lncRNA_lncRNA_corr, df_lncRNA_gene_corr, list_lncRNA = generate_lncRNA_files(df_gene_sample_exp)
  654. dict_lncRNAsets = generate_lncRNA_sets(list_lncRNA, df_lncRNA_lncRNA_corr)
  655. # step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
  656. # This is the counterpart of the step 3 for TFs.
  657. df_TF_gene_corr, df_TF_TF_corr, list_TF = generate_TF_files(df_gene_sample_exp)
  658. dict_TFsets = generate_TF_sets(list_TF, df_TF_TF_corr)
  659. # step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
  660. dict_lncRNAtarget_shared = generate_lncRNA_target(df_lncRNA_gene_DBS, df_lncRNA_gene_corr, dict_lncRNAsets)
  661. # step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
  662. if 'df_TF_gene_DBS' in dir():
  663. dict_TFtarget_shared = generate_TF_target(df_TF_gene_DBS, df_TF_gene_corr, dict_TFsets)
  664. else:
  665. dict_TFtarget_shared = {}
  666. # step 7 - identify modules and merge lncRNAs/TFs that regulate the same modules
  667. dict_module_regulators, dict_module_regulator_sets = collect_modules(dict_lncRNAtarget_shared, dict_TFtarget_shared, dict_lncRNAsets, dict_TFsets)
  668. if arg_RemoveRedundancy == 'y':
  669. dict_module_regulators, dict_module_regulator_sets = remove_redundancy(dict_module_regulators, dict_module_regulator_sets)
  670. # step 8 - perform pathway enrichment analysis for module genes in df_path_module. KEGG gene annotation is used.
  671. dict_module_kegg = pathway_analysis(dict_module_regulators, kegg_gene, kegg_pathway, kegg_link, totalGeneNum, arg_fdrCutoff)
  672. dict_module_wiki = pathway_analysis(dict_module_regulators, kegg_gene, wiki_pathway, wiki_link, totalGeneNum, arg_fdrCutoff)
  673. # step 9 - write results to files
  674. newOutFile = check_output(arg_OutFile, resultDir)
  675. if newOutFile:
  676. arg_OutFile = newOutFile
  677. print('A result file with the same name already exists, the newly generated file is named with '+ newOutFile + '.')
  678. write_modules(dict_module_regulators, dict_module_regulator_sets, dict_module_kegg, dict_module_wiki, arg_OutFile)
  679. generate_cytosc_file(arg_OutFile)

eGRAMv2R1.py at commit 5af5db9, under MIT · at the source

Overview

Authors: Jie Lin1,2, Yujian Wen1, Ji Tang1, Xuecong Zhang1, Huanlin Zhang1, Hao Zhu1,3,4
ORCID iDs: Jie Lin, Hao Zhu
  1. Bioinformatics Section, School of Basic Medical Sciences, Southern Medical University, Guangzhou, China
  2. College of Biological and Food Engineering, Guangdong University of Petrochemical Technology, Maoming, China
  3. Guangdong-Hong Kong-Macao Greater Bay Area Center for Brain Science and Brain-Inspired Intelligence, Southern Medical University, Guangzhou, China
  4. Guangdong Provincial Key Lab of Single Cell Technology and Application, Southern Medical University, Guangzhou, China
Journal: eLife, volume 12, article RP89001
Dates: published online 13 March 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.7554/elife.89001 · PMID 41823704 · PMCID PMC12987650 · OpenAlex W4385275559
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), cellular / molecular (subfield)
Methods: Statistics, Connectivity
Keywords: Human
MeSH: Evolution, Molecular*, Gene Expression Regulation*, RNA, Long Noncoding*, Animals, Binding Sites, Humans, Neanderthals, Species Specificity (* major topic)
Topic: Cancer-related molecular mechanisms research (Cancer Research, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: China Postdoctoral Science Foundation (2020M682788); National Natural Science Foundation of China (31771456)
Citations: cited by 4 papers (Europe PMC); 100 references in the paper

Abstract

What genes and regulatory sequences critically differentiate modern humans from apes and archaic humans, which share highly similar genomes but show distinct phenotypes, has puzzled researchers for decades. Previous studies examined species-specific protein-coding genes and related regulatory sequences, revealing that birth, loss, and changes in these genes and sequences drive speciation and evolution. However, investigations of species-specific lncRNA genes and related regulatory sequences, which regulate substantial genes, remain limited. We identified human-specific (HS) lncRNAs from GENCODE-annotated human lncRNAs, predicted their DNA-binding domains (DBDs) and DNA-binding sites (DBSs), analyzed DBS sequences in modern humans (CEU, CHB, and YRI), archaic humans (Altai Neanderthals, Denisovans, and Vindija Neanderthals), and chimpanzees, and investigated how HS lncRNAs and their DBSs have influenced gene expression in archaic and modern humans. Our results suggest that these lncRNAs and DBSs have substantially reshaped gene expression, and this reshaping has evolved continuously from archaic to modern humans, enabling humans to adapt to new environments and lifestyles, promoting brain evolution, and resulting in cross-population differences. The parallel analysis of gene expression in GTEx tissues by HS transcription factors (TFs) and their DBSs indicates that HS lncRNAs have reshaped gene expression in the brain more significantly than HS TFs.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

linjiecodes/egramv2r1

License: MIT
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 5af5db9572d694eab0a0ef28431e0e23daa105c2, 25 February 2026
Languages: Python (1)
Size: 17 files, 1 script
Software Heritage: not archived
Found in: the references
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: NumPy (1 file), pandas (1 file), SciPy (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
3 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1 script, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Other data links

Data availability

Data is provided in the Supplementary file 1 and at the following locations: GENCODE human lncRNAs are freely available at the website (https://www.gencodegenes.org/human/). The LongTarget program and human lncRNAs' orthologs in 16 mammals are freely available at the LongTarget website (http://www.gaemons.net/), which consists of the program, the lncRNA database LongMan, the full sequences of many mammamlain genomes, and a cluster of servers (also available at Zenodo (https://doi.org/10.5281/zenodo.18919871)). Fasim-LongTarget is a fast, simple, and standalone version of LongTarget and is freely available at Zenodo (https://doi.org/10.5281/zenodo.18919871). The eGRAM program with examples is freely available the at Zenodo (https://doi.org/10.5281/zenodo.18919871) and on the GitHub website (https://github.com/LinjieCodes/eGRAMv2R1 copy archived at Lin, 2026). The data on favored and hitchhiking mutations are available at Zenodo (https://doi.org/10.5281/zenodo.18919871). RNA-seq data of cell lines before and after DBD KO of lncRNAs were deposited to the NCBI GEO database (https://www.ncbi.nlm.nih.gov/geo), with two accession numbers (accession number GSE213231 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE213231); accession number is GSE229846 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE229846)).

The following datasets were generated:

Zhu H. 2022. Knockout of the DNA binding domains in human-specific lncRNAs causes changed expression of target genes. NCBI Gene Expression Omnibus. GSE213231

Zhu H. 2024. A multi-cancer CRISPR/Cas9 reveals distinct gene dysregulation by noncoding MSTRG transcripts. NCBI Gene Expression Omnibus. GSE229846

Lin J. 2026. Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression. Zenodo.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 1 keyword, 8 MeSH terms, 2 funders, 98 references.

Cite

This paper

Lin, J., Wen, Y., Tang, J., Zhang, X., Zhang, H., & Zhu, H. (2026). Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression. eLife, 12, RP89001. https://doi.org/10.7554/elife.89001

BibTeX

@article{lin2026human,
author = {Lin, Jie and Wen, Yujian and Tang, Ji and Zhang, Xuecong and Zhang, Huanlin and Zhu, Hao},
title = {{Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression}},
journal = {eLife},
year = {2026},
month = mar,
volume = {12},
pages = {RP89001},
publisher = {eLife Sciences Publications, Ltd},
issn = {2050-084X},
doi = {10.7554/elife.89001},
url = {https://doi.org/10.7554/elife.89001},
pmid = {41823704},
pmcid = {PMC12987650}
}

RIS

TY - JOUR
AU - Lin, Jie
AU - Wen, Yujian
AU - Tang, Ji
AU - Zhang, Xuecong
AU - Zhang, Huanlin
AU - Zhu, Hao
TI - Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression
T2 - eLife
J2 - Elife
PY - 2026
DA - 2026/03/13
VL - 12
SP - RP89001
SN - 2050-084X
PB - eLife Sciences Publications, Ltd
DO - 10.7554/elife.89001
UR - https://doi.org/10.7554/elife.89001
LA - en
ER -

CSL-JSON

{
"id": "10.7554/elife.89001",
"type": "article-journal",
"title": "Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression",
"container-title": "eLife",
"author": [
{
"family": "Lin",
"given": "Jie"
},
{
"family": "Wen",
"given": "Yujian"
},
{
"family": "Tang",
"given": "Ji"
},
{
"family": "Zhang",
"given": "Xuecong"
},
{
"family": "Zhang",
"given": "Huanlin"
},
{
"family": "Zhu",
"given": "Hao"
}
],
"container-title-short": "Elife",
"volume": "12",
"page": "RP89001",
"DOI": "10.7554/elife.89001",
"PMID": "41823704",
"PMCID": "PMC12987650",
"ISSN": "2050-084X",
"publisher": "eLife Sciences Publications, Ltd",
"URL": "https://doi.org/10.7554/elife.89001",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
13
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s13059-026-04098-8 [code]
A LINE-1 insertion upstream of FOXP2 promotes neuronal differentiation during primate evolution.
Journal: Genome biology
In common: pandas, SciPy, NumPy, 7 references
[2] doi:10.1016/j.xhgg.2026.100629 [code]
Positive selection on brain cis-regulatory elements in the human lineage drives changes in gene expression and susceptibility to neuropsychiatric disorders.
Journal: HGG advances
In common: pandas, SciPy, NumPy, cellular / molecular, 5 references
[3] doi:10.1093/molbev/msag140 [code]
Segmentally Duplicated Regulatory Elements Undergo Human-Specific Rewiring.
Journal: Molecular biology and evolution
In common: cellular / molecular, 5 references
[4] doi:10.1016/j.xgen.2026.101278 [code]
Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures.
Journal: Cell genomics
In common: pandas, NumPy, 4 references
[5] doi:10.1038/s42003-026-10802-y
Transcriptome of fetal cortex of tree shrew underlying the emergence of outer subventricular zone.
Journal: Communications biology
In common: 4 references
[6] doi:10.1038/s41467-026-74153-2 [code]
Regional, functional and transcriptomic decoding of multidimensional brain structure alterations in obsessive-compulsive disorder.
Journal: Nature communications
In common: pandas, SciPy, NumPy, cellular / molecular, 3 references
[7] doi:10.1093/gbe/evag086 [code]
Genome Scans Reveal Species-Specific Selection in the Genus Lynx.
Journal: Genome biology and evolution
In common: pandas, cellular / molecular, 3 references
[8] doi:10.1016/j.mocell.2026.100363 [code]
Tandem repeats in human brain evolution and disease susceptibility.
Journal: Molecules and cells
In common: NumPy, cellular / molecular, 3 references
[9] doi:10.1038/s41586-026-10699-x [code]
Competing programs shape cortical sensorimotor-association axis development.
Journal: Nature
In common: NumPy, 4 references
[10] doi:10.1038/s41588-026-02646-3 [code]
Co-expression-based models improve eQTL predictions for transcriptome-wide association studies and highlight new schizophrenia-associated genes.
Journal: Nature genetics
In common: pandas, SciPy, NumPy, cellular / molecular, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.