Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression.
The 2 matches
- [1] § Materials and methods › Identifying and analyzing transcriptional regulatory modules ↔ eGRAMv2R1.py, lines 1–49 · score 0.98 · clustering algorithms, regulatory modules, eGRAM, transcriptionally regulate, independently regulated, co regulated
- [2] § Appendix 8 › The analysis of HS TFs and their DBSs ↔ eGRAMv2R1.py, lines 1–49 · score 0.79 · CellOracle, Pearson correlation, TF target, database, gene expression, intersection
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 776 lines · 41 KB · MIT · 2 matches
- '''
- Several notes:
- 1. There are two kinds of input files - (a) user data and (b) system data. The former has three files - (a) expression profile, (b) lncRNA target
- prediction, (c) TF target prediction. LncRNAs' targets are predicted using LongTarget (or the more rapid version Farsim); TFs' targets are
- predicted using any method, such as CellOracle. Since the primary goal of eGRAM is analyzing lncRNA function in transcriptional regulation, the
- TF target prediction file is optional. System data include human/mouse KEGG and wikipathway pathways downloaded from the KEGG and Wikipathways
- websites.
- 2. All variables begin with their data type: df_, list_, dict_, dict2D_, set_.
- 3. It is the user's responsibility to determine correct gene symbols and gene types, and keep gene symbols and types consistent in the input files.
- In the given examples, targets of lncRNAs and TFs are only lncRNA genes, protein-coding genes, and TF genes. The user can explore other gene
- types (but they are barely annotated in pathway databases).
- 4. eGRAM performs transcription analysis based on both correlated gene expression and lncRNA/TF target prediction. For small sample size, Pearson
- correlation is recommended, for large sample size, Spearman correlation can also be chosen.
- 5. eGRAM identifies "all-mutually-correlated" lncRNAs and "all-mutually-correlated" TFs from lncRNAs and TFs in the gene expression file. We call
- this method "COCO", which is akin to clustering genes into clusters allowing overlap. Clustering with overlapping is reasonable because
- a lncRNA may work with different partners in different situations to regulate transcription, so do TFs. Two command-line arguments,
- --lncCorr and --tfCorr, work with the option. Since there are many clustering algorithm, future versions will allow the use of other algorithms
- (e.g. FLAME).
- 6. To make eGRAM able to handle lncRNAs and TFs of any names in any numbers, we organize lncRNAs and TFs into dictionaries, each lncRNA/TF is the
- key and its COCO set (a list) is the value.
- 7. Several pending issues
- (1) The relations between regulators and targets may be (a) co-upregulation, (b) co-downregulation, (c) differential expression - regulator high
- and target low, (d) differential expression - regulator low and target high. --- we have not explore this???
- (2) Whether a cutoff is needed for controlling the smallest COCO set size. For example, a set of <=3 lncRNAs or a set of <=2 TFs may not indicate
- a reliable regulator set.
- (3) There are three kinds of modules - lncRNAs' target sets, TFs' target sets, and merged target sets (intersection and union). Pathway enrichment
- analysis can be applied to each of the sets to reveal whether lncRNAs' target sets and TFs' target sets are enriched in different pathways.
- (4) Examining the intersection or union of merged sets is also biologically different, indicating co-regulation coordinated regulation by lncRNAs
- and TFs, respectively.
- 8. A gene with all 0s is removed, but otherwise kept. 0s in RNA-seq datasets are assumed true value, so no imputation is performed. If scRNA-seq data
- is analyzed, it is the user's responsibility to perform imputation using any method. We assume that imputation is finished before running eGRAM.
- Main running steps:
- step 1 - read in user files and command line arguments
- step 2 - read in kegg files and wikipathway files
- step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
- step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
- step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
- step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
- step 7 - identify typically co-regulated modules and independently-regulated modules
- step 8 - perform pathway enrichment analysis for module genes in df_path_module
- step 9 - write results to files
- step 10 - generate files for cytoscape
- An explanatory command line:
- eGRAMv2R1.py --exp 2024May-DEG-exp-A549-2513WT.csv --lncDBS 2024May-lncRNA-DBS-A549-2513.csv --tfDBS 2024May-TF-DBS-A549-2513.csv --species 1
- --lncCutoff 100 --tfCutoff 8 --lncCorr 0.6 --tfCorr 0.6 --moduleSize 50 --corr Pearson --cluster COCO --fdr 0.01 --out outA5492513WT
- The user can output intermediate results to a specific file by adding "> file_name" at the end of the command line in ther Terminal mode.
- '''
- #!/usr/bin/python
- import numpy as np
- import pandas as pd
- import re
- import os
- import csv
- import sys, getopt
- from scipy.stats import hypergeom
- #from sklearn.impute import SimpleImputer
- ########################################################################################################################
- # step 2 - read in kegg files and wikipathway files
- def read_kegg(geneFile, pathFile, linkFile):
- '''
- Read kegg genes, pathways and pathway-gene links.
- '''
- gene_symbols = {} # debug info: gene_symbols={dict:0} {}, 0 indicating no element.
- kegg_pathway = {} # debug info: gene_symbols={dict:0} {}.
- kegg_link = {}
- gene_id_set = set() # debug info: gene_id_set={set:0} {}.
- with open(geneFile) as f:
- for line in f: # debug: first = 'hsa:102466751\tmiRNA\t1:complement(17369..17436)\tMIR6859-1, hsa-mir-6859-1; microRNA 6859-1\n'
- cols = line.strip('\r\n').split('\t') # debug: ['hsa:102466751', 'miRNA', '1:complement(17369..17436)', 'MIR6859-1, hsa-mir-6859-1; microRNA 6859-1']
- gene_id = cols[0] # debug: gene_id = 'hsa:102466751', the 1st element in cols, i.e., cols[0].
- descri = cols[-1] # descri is the last element of cols, e.g., 'MIR6859-1, hsa-mir-6859-1; microRNA 6859-1'.
- if (',' in descri) or (';' in descri):
- symbol_descri = descri.split(';')[0]
- symbols = symbol_descri.split(", ")
- if gene_id not in gene_symbols: #
- gene_symbols[gene_id] = set()
- for symbol in symbols:
- gene_symbols[gene_id].add(symbol)
- gene_id_set.add(gene_id)
- with open(pathFile) as f:
- for line in f:
- pathID, pathwayName = line.strip('\r\n').split('\t')
- pathID = pathID.replace("path:", '')
- pathwayName = pathwayName[:pathwayName.rfind(' -')]
- kegg_pathway[pathID] = pathwayName
- with open(linkFile) as f:
- for line in f:
- pathID, gene_id = line.strip('\r\n').split('\t')
- pathID = pathID.replace("path:", '')
- if pathID in kegg_pathway:
- if gene_id in gene_id_set:
- if pathID not in kegg_link:
- kegg_link[pathID] = set()
- kegg_link[pathID].add(gene_id)
- totalGeneNum = len(gene_id_set) # totalGeneNum = 22101
- kegg_gene = {}
- for gene_id in gene_symbols:
- for gene_symbol in gene_symbols[gene_id]:
- kegg_gene[gene_symbol] = gene_id
- return kegg_gene, kegg_pathway, kegg_link, totalGeneNum
- def read_wikip(wikiDir, kegg_gene):
- '''
- Read genes and pathways in Wikipathways.
- '''
- wiki_pathway = {}
- wiki_link = {}
- pathway_files = os.listdir(wikiDir)
- pathway_comp = re.compile(' Name="(.+?)" ')
- gene_comp = re.compile(' TextLabel="(.+?)" ')
- for pathway_file in pathway_files:
- pathway_id = pathway_file.split('_')[-2]
- with open(wikiDir + pathway_file, encoding='utf-8') as f:
- for line in f:
- pathway_re = re.search(pathway_comp, line)
- if pathway_re:
- pathway = pathway_re.group(1)
- wiki_pathway[pathway_id] = pathway_re.group(1)
- gene_re = re.search(gene_comp, line)
- if gene_re:
- symbol = gene_re.group(1)
- if symbol in kegg_gene:
- gene_id = kegg_gene[symbol]
- if pathway_id in wiki_pathway:
- if pathway_id not in wiki_link:
- wiki_link[pathway_id] = set()
- wiki_link[pathway_id].add(gene_id)
- return wiki_pathway, wiki_link
- ########################################################################################################################
- # step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
- def generate_lncRNA_files(df_gene_sample_exp):
- df_lncRNA_sample_exp = df_gene_sample_exp.copy(deep=True)
- # When deep=True (default), a new object is created.
- df_lncRNA_sample_exp = df_lncRNA_sample_exp.drop(df_lncRNA_sample_exp[df_lncRNA_sample_exp['genetype'] != 'lncRNA'].index)
- # dropping non-lncRNA rows.
- df_lncRNA_sample_exp = df_lncRNA_sample_exp.reset_index()
- # "drop" operation does change the original index.
- # This .reset_index() re-sets index (also keeps the original index in df_gene_sample_exp).
- df_lncRNA_sample_exp.drop(["index"], axis=1, inplace=True)
- # drop the old index.
- list_lncRNA = df_lncRNA_sample_exp['genesymbol_GRCh38']
- df_lncRNA_sample_exp.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
- df_sample_lncRNA_exp = df_lncRNA_sample_exp.T
- # correlation is computed between columns.
- df_lncRNA_lncRNA_corr = df_sample_lncRNA_exp.corr(method='pearson')
- df_lncRNA_lncRNA_corr.columns = list_lncRNA
- # This replaces column names (but not column content) with lncRNA names in list_lncRNA (i.e., re-index columns).
- df_gene_sample_exp2 = df_gene_sample_exp.copy(deep=True)
- # df_gene_sample_exp2 is a temporary dataframe
- list_gene = df_gene_sample_exp2['genesymbol_GRCh38']
- list_type = df_gene_sample_exp2['genetype']
- df_gene_sample_exp2.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
- df_sample_gene_exp = df_gene_sample_exp2.T
- print("\nafter transferred\n", df_sample_gene_exp)
- if arg_Pearson == 'Spearman':
- df_gene_gene_corr = df_sample_gene_exp.corr(method='spearman')
- else:
- df_gene_gene_corr = df_sample_gene_exp.corr(method='pearson')
- print(df_gene_gene_corr)
- df_gene_gene_corr.columns = list_gene
- # Using gene names in list_gene to replace the 0-N column names.
- print("\nafter putting column names\n", df_gene_gene_corr)
- for j in range(len(list_gene)):
- if list_type[j] != 'lncRNA':
- df_gene_gene_corr.drop(columns=list_gene[j], inplace=True)
- print("\nafter deleting non-lncRNA columns\n", df_gene_gene_corr)
- df_lncRNA_gene_corr = df_gene_gene_corr.copy(deep=True)
- df_lncRNA_gene_corr.index = list_gene
- return df_lncRNA_lncRNA_corr, df_lncRNA_gene_corr, list_lncRNA
- def generate_lncRNA_sets(list_lncRNA, df_lncRNA_lncRNA_corr):
- dict_lncRNAsets = {}
- # For each lncRNA in list_lncRNA, a dict is generated with the lncRNA as the key and a list as the value.
- # The list contains all-mutually-correlated lncRNAs of this lncRNA.
- # Process lncRNAs in df_lncRNA_lncRNA_corr row by row (lncRNAs in list_lncRNA and df_lncRNA_lncRNA_corr have the same order).
- for row in df_lncRNA_lncRNA_corr.index:
- deleted = []
- for col in range(0, len(list_lncRNA)):
- if df_lncRNA_lncRNA_corr.iat[row, col] < arg_LncRNACorrCutoff:
- df_lncRNA_lncRNA_corr.iat[row, col] = -100 # -100 is a number for marking the current column
- df_lncRNA_lncRNA_corr.iat[col, row] = -100 # -100 is a number for marking the corresponding row(diagonal elements are equal)
- deleted.append(list_lncRNA[col]) # Put this lncRNA to deleted[]
- # For lncRNAs with high pairwise corr with the current lncRNA (row), the two-for-loop check/ensure all-mutually-correlated lncRNAs.
- for i in range(0, len(list_lncRNA)):
- if i != row and list_lncRNA[i] not in deleted:
- for j in range(0, len(list_lncRNA)):
- if list_lncRNA[j] not in deleted:
- if df_lncRNA_lncRNA_corr.iat[i, j] < arg_LncRNACorrCutoff:
- df_lncRNA_lncRNA_corr.iat[i, j] = -100 # -100 for marking the newly identified invalid lncRNA
- deleted.append(list_lncRNA[j]) # Put the newly identified invalid lncRNA into deleted[]
- else:
- df_lncRNA_lncRNA_corr.iat[i, j] = -100 # If the col name is already in deleted[], we still set the value to -100.
- lncRNAset = set(list_lncRNA) - set(deleted)
- if list_lncRNA[row] not in deleted:
- lncRNAlist = sorted(list(lncRNAset))
- else:
- lncRNAlist = []
- dict_lncRNAsets.update({list_lncRNA[row]: lncRNAlist})
- print(row, {list_lncRNA[row]: lncRNAlist})
- return dict_lncRNAsets
- ########################################################################################################################
- # step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
- def generate_TF_files(df_gene_sample_exp):
- df_TF_sample_exp = df_gene_sample_exp.copy(deep=True)
- # When deep=True (default), a new object is created with a copy of the calling object’s data and indices.
- df_TF_sample_exp = df_TF_sample_exp.drop(df_TF_sample_exp[df_TF_sample_exp['genetype'] != 'TF'].index)
- # drop non-TF rows.
- df_TF_sample_exp = df_TF_sample_exp.reset_index()
- # "drop" operation does change the original index.
- # This .reset_index() re-sets index and also keeps the original index in df_gene_sample_exp.
- df_TF_sample_exp.drop(["index"], axis=1, inplace=True)
- # drop the old index.
- list_TF = df_TF_sample_exp['genesymbol_GRCh38']
- # get all TFs in the first column that are gene names (this list also obtains the column name "genesymbol_GRCh38")
- df_TF_sample_exp.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
- # drop "genesymbol_GRCh38" and "genetype", because all fileds should be numeric to generate the corr dataframe.
- df_sample_TF_exp = df_TF_sample_exp.T
- # correlation is computed between columns.
- df_TF_TF_corr = df_sample_TF_exp.corr(method='pearson')
- df_TF_TF_corr.columns = list_TF
- # This replaces column names (but not column content) with TF names in list_TF (i.e., re-index columns).
- df_gene_sample_exp2 = df_gene_sample_exp.copy(deep=True)
- # df_gene_sample_exp2 is a temporary dataframe
- list_gene = df_gene_sample_exp2['genesymbol_GRCh38']
- list_type = df_gene_sample_exp2['genetype']
- df_gene_sample_exp2.drop(["genesymbol_GRCh38", "genetype"], axis=1, inplace=True)
- df_sample_gene_exp = df_gene_sample_exp2.T
- print("\nafter transferred\n", df_sample_gene_exp)
- if arg_Pearson == 'Spearman':
- df_gene_gene_corr = df_sample_gene_exp.corr(method='spearman')
- else:
- df_gene_gene_corr = df_sample_gene_exp.corr(method='pearson')
- print(df_gene_gene_corr)
- df_gene_gene_corr.columns = list_gene
- # Using gene names in list_gene to replace the 0-N column names.
- print("\nafter putting column names\n", df_gene_gene_corr)
- for j in range(len(list_gene)):
- if list_type[j] != 'TF':
- df_gene_gene_corr.drop(columns=list_gene[j], inplace=True)
- print("\nafter deleting non-TF columns\n", df_gene_gene_corr)
- df_TF_gene_corr = df_gene_gene_corr.copy(deep=True)
- df_TF_gene_corr.index = list_gene
- return df_TF_gene_corr, df_TF_TF_corr, list_TF
- def generate_TF_sets(list_TF, df_TF_TF_corr):
- dict_TFsets = {}
- # For each TF in list_TF, a dict is generated with the TF as key and a list as the value.
- # The list contains all-mutually-correlated TFs of this TF.
- # Process TFs in df_TF_TF_corr row by row. Note TFs in list_TF and df_TF_TF_corr have the same order.
- for row in df_TF_TF_corr.index: # This row-for-loop check all TFs (rows)
- deleted = []
- for col in range(0, len(list_TF)):
- if df_TF_TF_corr.iat[row, col] < arg_TFCorrCutoff:
- df_TF_TF_corr.iat[row, col] = -100
- df_TF_TF_corr.iat[col, row] = -100
- deleted.append(list_TF[col])
- # For TFs with high pairwise corr with the current TF (row), the two-for-loop check/ensure all-mutually-correlated TFs.
- for i in range(0, len(list_TF)):
- if i != row and list_TF[i] not in deleted:
- for j in range(0, len(list_TF)):
- if list_TF[j] not in deleted:
- if df_TF_TF_corr.iat[i, j] < arg_TFCorrCutoff:
- df_TF_TF_corr.iat[i, j] = -100
- deleted.append(list_TF[j])
- else:
- df_TF_TF_corr.iat[i, j] = -100
- TFset = set(list_TF) - set(deleted)
- if list_TF[row] not in deleted:
- TFlist = sorted(list(TFset))
- else:
- TFlist = []
- dict_TFsets.update({list_TF[row]: TFlist})
- print(row, {list_TF[row]: TFlist})
- return dict_TFsets
- ########################################################################################################################
- # step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
- def generate_lncRNA_target(df_lncRNA_gene_DBS, df_lncRNA_gene_corr, dict_lncRNAsets):
- '''
- Upon df_lncRNA_gene_corr and df_lncRNA_gene_DBS, we get targets of every lncRNA set, which are the
- intersection of set_lncRNAtarget_DBS (DBS>cutoff1) and set_lncRNAtarget_CORR (|corr|>cutoff2).
- cutoff2 (lncRNA-gene-corr) = lncRNA-lncRNA-corr (arg_LncRNACorrCutoff). A separable cutoff can be considered.
- :
- traverse key in dict_lncRNAsets
- traverse lncRNA in lncRNAlist
- get targets of lncRNA from df_lncRNA_gene_DBS, append targets to dict_lncRNAtarget_DBS
- get targets of lncRNA from df_lncRNA_gene_CORR, append targets to dict_lncRNAtarget_CORR
- get overlap of the two target sets, append the overlap to dict_lncRNAtarget_shared
- '''
- list_lncRNACORR = df_lncRNA_gene_corr.columns
- list_lncRNADBS = df_lncRNA_gene_DBS.columns
- dict_lncRNAtarget_DBS = {}
- dict_lncRNAtarget_CORR = {}
- dict_lncRNAtarget_shared = {}
- rows1 = df_lncRNA_gene_corr.shape[0]
- cols1 = df_lncRNA_gene_corr.shape[1]
- rows2 = df_lncRNA_gene_DBS.shape[0]
- cols2 = df_lncRNA_gene_DBS.shape[1]
- for key, value in dict_lncRNAsets.items():
- set_tpmCORR = set()
- set_tmpDBS = set()
- if key in list_lncRNACORR and value != []: # if the COCO set of this lncRNA is non-empty
- for lncRNA in value: # then process all lncRNAs in the lncRNAlist
- if lncRNA in df_lncRNA_gene_corr.columns:
- for target in df_lncRNA_gene_corr.index:
- if abs(df_lncRNA_gene_corr[lncRNA][target]) > arg_LncRNACorrCutoff:
- set_tpmCORR.add(target)
- dict_lncRNAtarget_CORR.update({key: set_tpmCORR})
- if key in list_lncRNADBS and value != []:
- for lncRNA in value:
- if lncRNA in df_lncRNA_gene_DBS.columns:
- for target in df_lncRNA_gene_DBS.index:
- if df_lncRNA_gene_DBS[lncRNA][target] > arg_LncRNADBSCutoff:
- set_tmpDBS.add(target)
- dict_lncRNAtarget_DBS.update({key: set_tmpDBS})
- # dict_lncRNAtarget_CORR and dict_lncRNAtarget_DBS have the same keys and key order.
- for key, value1 in dict_lncRNAtarget_DBS.items():
- if key in dict_lncRNAtarget_CORR:
- value2 = dict_lncRNAtarget_CORR[key]
- shared = value1 & value2
- dict_lncRNAtarget_shared.update({key: shared})
- return dict_lncRNAtarget_shared
- ########################################################################################################################
- # step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
- def generate_TF_target(df_TF_gene_DBS, df_TF_gene_corr, dict_TFsets):
- list_TFCORR = df_TF_gene_corr.columns
- list_TFDBS = df_TF_gene_DBS.columns
- dict_TFtarget_DBS = {}
- dict_TFtarget_CORR = {}
- dict_TFtarget_shared = {}
- rows1 = df_TF_gene_corr.shape[0]
- cols1 = df_TF_gene_corr.shape[1]
- rows2 = df_TF_gene_DBS.shape[0]
- cols2 = df_TF_gene_DBS.shape[1]
- for key, value in dict_TFsets.items():
- set_tmpCORR = set()
- set_tmpDBS = set()
- if key in list_TFCORR and value != []:
- for TF in value:
- if TF in df_TF_gene_corr.columns:
- for target in df_TF_gene_corr.index:
- if abs(df_TF_gene_corr[TF][target]) > arg_TFCorrCutoff:
- set_tmpCORR.add(target)
- dict_TFtarget_CORR.update({key: set_tmpCORR})
- if key in list_TFDBS and value != []:
- for TF in value:
- if TF in df_TF_gene_DBS.columns:
- for target in df_TF_gene_DBS.index:
- if df_TF_gene_DBS[TF][target] > arg_TFDBSCutoff:
- set_tmpDBS.add(target)
- dict_TFtarget_DBS.update({key: set_tmpDBS})
- # dict_TFtarget_CORR and dict_TFtarget_DBS have the same keys and key order.
- for key, value1 in dict_TFtarget_DBS.items():
- if key in dict_TFtarget_CORR:
- value2 = dict_TFtarget_CORR[key]
- shared = value1 & value2
- dict_TFtarget_shared.update({key: shared})
- return dict_TFtarget_shared
- ########################################################################################################################
- # step 7 - identify modules and merge lncRNAs/TFs that regulate the same modules
- def collect_modules(dict_lncRNAtarget_shared, dict_TFtarget_shared, dict_lncRNAsets, dict_TFsets):
- '''
- This function collects all lncRNAs' modules (lncRNA target sets) and TFs' modules (TF target sets) into a dict,
- with target genes in the module (a tuple) as key and lncRNAs/TFs that co-regulate the module as values.
- '''
- dict_module_regulators = {}
- dict_module_regulator_sets = {}
- # zhu
- # dict_lncRNAtarget_shared = {'AC090152.1': {'LRMDA', 'FGFR2', 'ABLIM1', 'TNFSF9'}, 'LINC02783': set(), 'Z99572.1': set(), 'AC010889.1': {'GABRA3', 'ABLIM1'}, 'ZNF790-AS1': {'COL18A1-AS1', 'MED14'}, 'AC125807.2': set()}
- # dict_TFtarget_shared = {'ZNF790': {'GABRA3', 'FGFR2', 'CEACAM6', 'TNFSF9'}, 'SOX21': {'LINC02577', 'COL18A1-AS1', 'MED14'}, 'ALX1': {'DENND2A', 'TSPAN7'}, 'HNF4A': {'DENND2A', 'TSPAN7'}, 'NR5A2': set()}
- #dict_xx_shared = {'1': {'1', '2', '3', '4'}, '5': set(), '6': set(), '7': {'8', '9'}, '10': {'11', '12'}, '13': set()}
- #dict_zz_shared = {'a': {'b', 'd', 'e', 'f'}, 'g': {'h', 'i', 'j'}, 'k': {'l', 'm'}, 'n': {'o', 'p'}, 'q': set()}
- #allModules = [tuple(sorted(module)) for module in list(dict_xx_shared.values()) + list(dict_zz_shared.values()) if len(module) > 1]
- # We sort the order of target genes in modules so that we can conveniently compare and determine whether two modules are the same
- allModules = [tuple(sorted(module)) for module in list(dict_lncRNAtarget_shared.values()) + list(dict_TFtarget_shared.values()) if len(module) > arg_ModuleSize]
- for module in allModules:
- dict_module_regulators[module] = set()
- dict_module_regulator_sets[module] = []
- for lncRNA in dict_lncRNAtarget_shared:
- module = tuple(sorted(dict_lncRNAtarget_shared[lncRNA]))
- # dict_lncRNAtarget_shared has 102 elements, each has a key (a lncRNA) and a value (the set of lncRNAs corresponding to the key)
- # module is the key, now the key of lncRNA set is the value, better to use the value of lncRNA set as the value
- if module in dict_module_regulators:
- dict_module_regulators[module].add(lncRNA + '(*)') # we use * to indicate lncRNA
- regulator_set = dict_lncRNAsets[lncRNA]
- if regulator_set and regulator_set not in dict_module_regulator_sets[module]:
- dict_module_regulator_sets[module].append(regulator_set)
- for TF in dict_TFtarget_shared:
- module = tuple(sorted(dict_TFtarget_shared[TF]))
- if module in dict_module_regulators:
- dict_module_regulators[module].add(TF + '(#)') # we use # to indicate TF
- regulator_set = dict_TFsets[TF]
- if regulator_set and regulator_set not in dict_module_regulator_sets[module]:
- dict_module_regulator_sets[module].append(regulator_set)
- print("dict_module_regulators=", dict_module_regulators)
- return dict_module_regulators, dict_module_regulator_sets
- ########################################################################################################################
- # step 8 - perform pathway enrichment analysis for module genes ### in df_path_module
- def pathway_analysis(dict_module_regulators, kegg_gene, kegg_pathway, kegg_link, totalGeneNum, arg_fdrCutoff):
- dict_module_kegg = {}
- results = []
- pvalues = []
- for module in dict_module_regulators:
- module_geneIDs = {}
- #a gene may have serveral symbols
- #we convert its symbol which user assign to KEGG geneID to obtain the interset between the module and a pathway
- for gene in module:
- if gene in kegg_gene:
- geneID = kegg_gene[gene]
- module_geneIDs[geneID] = gene
- moduleSize = len(module_geneIDs)
- for pathID in kegg_link:
- interSet = set(module_geneIDs.keys()) & kegg_link[pathID]
- if interSet:
- hitNum = len(interSet)
- thisPathGeneNum = len(kegg_link[pathID])
- # Indentify significantly enriched pathways using hypergeometric distribution test
- pVal = hypergeom.sf(hitNum-1, totalGeneNum, thisPathGeneNum, moduleSize)
- if pVal<0.05:
- pathName = kegg_pathway[pathID]
- hitGenes = set([module_geneIDs[geneID] for geneID in interSet])
- results.append([module, pathID, pathName, hitGenes, pVal])
- pvalues.append(pVal)
- #Benjamini-Hochberg method for p-value correction
- pValues = np.asfarray(pvalues)
- by_descend = pValues.argsort()[::-1]
- by_orig = by_descend.argsort()
- steps = float(len(pValues)) / np.arange(len(pValues), 0, -1)
- fdr_values = np.minimum(1, np.minimum.accumulate(steps * pValues[by_descend]))
- fdrs = []
- for i in range(len(fdr_values[by_orig])):
- fdrs.append(fdr_values[by_orig][i])
- for i in range(len(fdrs)):
- results[i].append(fdrs[i])
- for module, pathID, pathName, hitGenes, pVal, fdr in results:
- if fdr < arg_fdrCutoff:
- if module not in dict_module_kegg:
- dict_module_kegg[module] = [(fdr, pathID, pathName, hitGenes)]
- else:
- dict_module_kegg[module].append((fdr, pathID, pathName, hitGenes))
- return dict_module_kegg
- ########################################################################################################################
- # step 9 - write eGRAM results to files in the directory "eGRAM_results/"
- def write_modules(dict_module_regulators, dict_module_regulator_sets, dict_module_kegg, dict_module_wiki, outFile):
- with open('eGRAM_results/' + outFile + '.csv', 'w', newline='') as f:
- writer = csv.writer(f)
- writer.writerow(['ModuleID',
- 'LncRNA(*)/TF(#)',
- 'Regulator set',
- 'Target gene',
- 'KEGG pathway, pathwayID, and FDR',
- 'Hit genes of enriched KEGG pathways',
- 'WikiPathway, pathwayID, and FDR',
- 'Hit genes of enriched Wikipathways'])
- moduleID = 0
- for module in dict_module_regulators:
- moduleID += 1
- if module in dict_module_kegg:
- this_module_kegg = []
- this_module_kegg_hitGenes = set()
- for fdr, pathID, pathName, hitGenes in sorted(dict_module_kegg[module]):
- this_module_kegg.append(pathName + ' (' + pathID + ', ' + str('{0:3f}'.format(fdr)) + ')')
- this_module_kegg_hitGenes = this_module_kegg_hitGenes | hitGenes
- pathway_kegg = '; '.join(this_module_kegg)
- hitGenes_kegg = ', '.join(this_module_kegg_hitGenes)
- else:
- pathway_kegg = hitGenes_kegg = ''
- if module in dict_module_wiki:
- this_module_wiki = []
- this_module_wiki_hitGenes = set()
- for fdr, pathID, pathName, hitGenes in sorted(dict_module_wiki[module]):
- #this_module_wiki.append(pathName + ' (' + pathID + ', ' + str(fdr) + ')')
- this_module_wiki.append(pathName + ' (' + pathID + ', ' + str('{0:3f}'.format(fdr)) + ')')
- this_module_wiki_hitGenes = this_module_wiki_hitGenes | hitGenes
- pathway_wiki = '; '.join(this_module_wiki)
- hitGenes_wiki = ', '.join(this_module_wiki_hitGenes)
- else:
- pathway_wiki = hitGenes_wiki = ''
- regulators = ' | '.join(dict_module_regulators[module])
- regulatory_sets = ' | '.join([str(regulatory_set).replace("'",'') for regulatory_set in dict_module_regulator_sets[module]])
- targetGenes = ' | '.join(module)
- writer.writerow(['ME'+str(moduleID),
- regulators,
- regulatory_sets,
- targetGenes,
- pathway_kegg,
- hitGenes_kegg,
- pathway_wiki,
- hitGenes_wiki])
- # Generate and write files for module visulization using cytoscape in the 'eGRAM_results/' directory.
- def generate_cytosc_file(moduleFile):
- if '/' not in moduleFile:
- moduleFile = 'eGRAM_results/' + moduleFile
- fileName = moduleFile.split('/')[-1]
- f_re = open('eGRAM_results/' + fileName + '_edge', 'w')
- f_re.write('\t'.join(['lncRNA', 'module', 'weight'])+'\n')
- allLncs = set()
- allTFs = set()
- with open(moduleFile + '.csv') as f_mod:
- reader = csv.reader(f_mod)
- first_row = next(reader)
- for row in reader:
- moduleID = row[0]
- regulators = row[1].split(' | ')
- for regulator in regulators:
- regulator_symbol = regulator.split('(')[0]
- f_re.write('\t'.join([regulator_symbol, moduleID, 'NA']) + '\n')
- f_re.close()
- # Merge modules with the same gene sets to avoid redundancy
- def remove_redundancy(dict_module_regulators, dict_module_regulator_sets):
- dict_module_regulators_noRedundancy = {}
- dict_module_regulator_sets_noRedundancy = {}
- for module1 in dict_module_regulators:
- merge_flag = 0
- for module2 in dict_module_regulators_noRedundancy:
- set1 = set(module1)
- set2 = set(module2)
- larger_set = max(set1, set2, key=len)
- intersection = set1.intersection(set2)
- ratio = len(intersection) / len(larger_set)
- if ratio > 0.9:
- intersection_module = tuple(sorted(list(intersection)))
- dict_module_regulators_noRedundancy[intersection_module] = dict_module_regulators[module1] | dict_module_regulators_noRedundancy[module2]
- for regulator_set in dict_module_regulator_sets[module1]:
- if regulator_set not in dict_module_regulator_sets_noRedundancy[module2]:
- dict_module_regulator_sets_noRedundancy[module2].append(regulator_set)
- dict_module_regulator_sets_noRedundancy[intersection_module] = dict_module_regulator_sets_noRedundancy[module2]
- del dict_module_regulators_noRedundancy[module2]
- del dict_module_regulator_sets_noRedundancy[module2]
- merge_flag = 1
- break
- if merge_flag == 0:
- dict_module_regulators_noRedundancy[module1] = dict_module_regulators[module1]
- dict_module_regulator_sets_noRedundancy[module1] = dict_module_regulator_sets[module1]
- return dict_module_regulators_noRedundancy, dict_module_regulator_sets_noRedundancy
- # Modify the resulting file name to prevent overwriting existing files
- def check_output(arg_OutFile, resultDir):
- maxFileNum = -1
- for file in os.listdir(resultDir):
- if file==arg_OutFile + '.csv' and maxFileNum==-1:
- maxFileNum = 0
- elif arg_OutFile+'_' in file:
- this_fileNum = file.replace(arg_OutFile+'_', '').replace('.csv', '')
- if this_fileNum.isdigit():
- this_fileNum = int(this_fileNum)
- if this_fileNum > maxFileNum:
- maxFileNum = this_fileNum
- if maxFileNum == -1:
- return
- else:
- newFileNum = maxFileNum + 1
- newOutFile = arg_OutFile + '_' + str(newFileNum) + '.csv'
- return newOutFile
- ########################################################################################################################
- def usage():
- print('Usage: python eGRAMv2R1.py --exp gene_exp_file --lncDBS lncRNA_DBS_file --tfDBS TF_DBS_file --species 1 --lncCutoff lncRNA-DBS-cutoff --tfCutoff TF-DBS-cutoff --lncCorr lncRNA-corr-cutoff --tfCorr TF-corr-cutoff --moduleSize modulesize --corr Pearson/Spearman --remove_redundancy y --fdr 0.01 --out output')
- print()
- print('Example: python eGRAMv2R1.py --exp 2024May-DEG-exp-A549-2513WT.csv --lncDBS 2024May-lncRNA-DBS-A549-2513.csv --tfDBS 2024May-TF-DBS-A549-2513.csv --species 1 --lncCutoff 100 --tfCutoff 8 --lncCorr 0.5 --tfCorr 0.5 --moduleSize 50 --corr Pearson --remove_redundancy y --fdr 0.01 --out myout')
- print()
- print("""Parameters:
- --exp gene_exp_file is required; if genetype does not have TF, all TF-related functions are not performed.
- --lncDBS lncRNA_DBS_file is required; integer or 0.
- --tfDBS TF_DBS_file is not a must; required if gene_exp_file contains TF.
- --species 1=human, 2=mouse.
- --lncCutoff the DBS threshold for determining lncRNAs' targets (defult=100).
- --tfCutoff the DBS threshold for determining TFs's targets (default=10).
- --lncCorr the correlation threshold for lncRNA/target and lncRNA/lncRNA (default=0.5).
- --tfCorr the correlation threshold for TF/target and TF/TF (default=0.5).
- --moduleSize the minimum gene number of a module (default=50).
- --corr Pearson/Spearman (default=pearson).
- --remove_redundancy y/n, whether or not merge modules with the same gene sets (default=y)
- --fdr the FDR value to determine significance of pathway enrichment (default 0.01).
- --out output file name.
- """)
- ########################################################################################################################
- ########################################################################################################################
- if __name__ == '__main__':
- param = sys.argv[1:]
- try:
- opts, args = getopt.getopt(param, '-h', ['exp=', 'lncDBS=', 'tfDBS=', 'species=', 'lncCutoff=', 'tfCutoff=', 'lncCorr=', 'tfCorr=', 'moduleSize=', 'corr=', 'remove_redundancy=', 'fdr=', 'out='])
- except getopt.GetoptError:
- usage()
- sys.exit(2)
- # step 1 - read in user files and command line arguments
- for opt, arg in opts:
- if opt == '-h':
- usage()
- sys.exit(2)
- elif opt == '--exp':
- df_gene_sample_exp = pd.read_csv(arg)
- elif opt == '--lncDBS':
- df_lncRNA_gene_DBS = pd.read_csv(arg, index_col = [0])
- elif opt == '--tfDBS':
- df_TF_gene_DBS = pd.read_csv(arg, index_col = [0])
- elif opt == '--species':
- arg_SpeciesPath = int(arg)
- elif opt == '--lncCutoff':
- arg_LncRNADBSCutoff = int(arg)
- elif opt == '--tfCutoff':
- arg_TFDBSCutoff = int(arg)
- elif opt == '--lncCorr':
- arg_LncRNACorrCutoff = float(arg)
- elif opt == '--tfCorr':
- arg_TFCorrCutoff = float(arg)
- elif opt == '--moduleSize':
- arg_ModuleSize = int(arg)
- elif opt == '--corr':
- arg_Pearson = arg
- elif opt == '--remove_redundancy':
- arg_RemoveRedundancy = arg
- elif opt == '--fdr':
- arg_fdrCutoff = float(arg)
- elif opt == '--out':
- arg_OutFile = arg
- # default arguments
- if 'arg_SpeciesPath' not in dir():
- arg_SpeciesPath = 1
- if 'arg_LncRNADBSCutoff' not in dir():
- arg_LncRNADBSCutoff = 100
- if 'arg_TFDBSCutoff' not in dir():
- arg_TFDBSCutoff = 10
- if 'arg_LncRNACorrCutoff' not in dir():
- arg_LncRNACorrCutoff = 0.5
- if 'arg_TFCorrCutoff' not in dir():
- arg_TFCorrCutoff = 0.5
- if 'arg_ModuleSize' not in dir():
- arg_ModuleSize = 50
- if 'arg_Pearson' not in dir():
- arg_Pearson = 'Pearson'
- if 'remove_redundancy' not in dir():
- arg_RemoveRedundancy = 'y'
- elif arg_RemoveRedundancy != 'n' and arg_RemoveRedundancy != 'y':
- arg_RemoveRedundancy = 'y'
- if 'arg_fdrCutoff' not in dir():
- arg_fdrCutoff = 0.01
- if arg_SpeciesPath == 1:
- keggGeneFile = "KEGG/keggGene-hsa"
- keggPathFile = "KEGG/keggPathway-hsa"
- keggLinkFile = "KEGG/keggLink-hsa"
- wikiPath = 'WikiPathways/wikipathways-20240310-gpml-Homo_sapiens/'
- elif arg_SpeciesPath == 2:
- keggGeneFile = "KEGG/keggGene-mmu"
- keggPathFile = "KEGG/keggPathway-mmu"
- keggLinkFile = "KEGG/keggLink-mmu"
- wikiPath = 'WikiPathways/wikipathways-20240310-gpml-Mus_musculus/'
- else:
- sys.exit(2)
- ########################################################################################################################
- # step 2 - read in kegg files and wikipathway files
- kegg_gene, kegg_pathway, kegg_link, totalGeneNum = read_kegg(keggGeneFile, keggPathFile, keggLinkFile)
- wiki_pathway, wiki_link = read_wikip(wikiPath, kegg_gene)
- resultDir = 'eGRAM_results/'
- if not os.path.exists(resultDir):
- os.mkdir(resultDir)
- # step 3 - computing lncRNA-gene correlations, lncRNA-lncRNA correlations, and upon which, lncRNAsets
- # i.e., clustering lncRNAs into 'COCO' sets (this could be done by introduing an outside program such as FLAME)
- df_lncRNA_lncRNA_corr, df_lncRNA_gene_corr, list_lncRNA = generate_lncRNA_files(df_gene_sample_exp)
- dict_lncRNAsets = generate_lncRNA_sets(list_lncRNA, df_lncRNA_lncRNA_corr)
- # step 4 - computing TF-gene correlations, TF-TF correlations, and upon which, TFsets.
- # This is the counterpart of the step 3 for TFs.
- df_TF_gene_corr, df_TF_TF_corr, list_TF = generate_TF_files(df_gene_sample_exp)
- dict_TFsets = generate_TF_sets(list_TF, df_TF_TF_corr)
- # step 5 - generate shared lncRNA target sets upon df_lncRNA_gene_DBS and df_lncRNA_gene_corr
- dict_lncRNAtarget_shared = generate_lncRNA_target(df_lncRNA_gene_DBS, df_lncRNA_gene_corr, dict_lncRNAsets)
- # step 6 - generate shared TF target sets upon df_TF_gene_DBS and df_TF_gene_corr
- if 'df_TF_gene_DBS' in dir():
- dict_TFtarget_shared = generate_TF_target(df_TF_gene_DBS, df_TF_gene_corr, dict_TFsets)
- else:
- dict_TFtarget_shared = {}
- # step 7 - identify modules and merge lncRNAs/TFs that regulate the same modules
- dict_module_regulators, dict_module_regulator_sets = collect_modules(dict_lncRNAtarget_shared, dict_TFtarget_shared, dict_lncRNAsets, dict_TFsets)
- if arg_RemoveRedundancy == 'y':
- dict_module_regulators, dict_module_regulator_sets = remove_redundancy(dict_module_regulators, dict_module_regulator_sets)
- # step 8 - perform pathway enrichment analysis for module genes in df_path_module. KEGG gene annotation is used.
- dict_module_kegg = pathway_analysis(dict_module_regulators, kegg_gene, kegg_pathway, kegg_link, totalGeneNum, arg_fdrCutoff)
- dict_module_wiki = pathway_analysis(dict_module_regulators, kegg_gene, wiki_pathway, wiki_link, totalGeneNum, arg_fdrCutoff)
- # step 9 - write results to files
- newOutFile = check_output(arg_OutFile, resultDir)
- if newOutFile:
- arg_OutFile = newOutFile
- print('A result file with the same name already exists, the newly generated file is named with '+ newOutFile + '.')
- write_modules(dict_module_regulators, dict_module_regulator_sets, dict_module_kegg, dict_module_wiki, arg_OutFile)
- generate_cytosc_file(arg_OutFile)
eGRAMv2R1.py at commit 5af5db9, under MIT · at the source
Overview
- Bioinformatics Section, School of Basic Medical Sciences, Southern Medical University, Guangzhou, China
- College of Biological and Food Engineering, Guangdong University of Petrochemical Technology, Maoming, China
- Guangdong-Hong Kong-Macao Greater Bay Area Center for Brain Science and Brain-Inspired Intelligence, Southern Medical University, Guangzhou, China
- Guangdong Provincial Key Lab of Single Cell Technology and Application, Southern Medical University, Guangzhou, China
Abstract
What genes and regulatory sequences critically differentiate modern humans from apes and archaic humans, which share highly similar genomes but show distinct phenotypes, has puzzled researchers for decades. Previous studies examined species-specific protein-coding genes and related regulatory sequences, revealing that birth, loss, and changes in these genes and sequences drive speciation and evolution. However, investigations of species-specific lncRNA genes and related regulatory sequences, which regulate substantial genes, remain limited. We identified human-specific (HS) lncRNAs from GENCODE-annotated human lncRNAs, predicted their DNA-binding domains (DBDs) and DNA-binding sites (DBSs), analyzed DBS sequences in modern humans (CEU, CHB, and YRI), archaic humans (Altai Neanderthals, Denisovans, and Vindija Neanderthals), and chimpanzees, and investigated how HS lncRNAs and their DBSs have influenced gene expression in archaic and modern humans. Our results suggest that these lncRNAs and DBSs have substantially reshaped gene expression, and this reshaping has evolved continuously from archaic to modern humans, enabling humans to adapt to new environments and lifestyles, promoting brain evolution, and resulting in cross-population differences. The parallel analysis of gene expression in GTEx tissues by HS transcription factors (TFs) and their DBSs indicates that HS lncRNAs have reshaped gene expression in the brain more significantly than HS TFs.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.
linjiecodes/egramv2r1
5af5db9572d694eab0a0ef28431e0e23daa105c2, 25 February 2026Availability: 1 check, the latest on 30 September 2026: the link answers
- 30 September 2026: the link answers
3 files
- eGRAMv2R1.py, Python, 776 lines, 2 matches
- LICENSE, License, 21 lines
- README.md, Text, 68 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 1 script, each with its path and the digest of its content;
- 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE213231, at NCBI GEO; found in “Data availability”
Other data links
- ncbi.nlm.nih.gov/
geo , NCBI; found in “Data availability”
Data availability
Data is provided in the Supplementary file 1 and at the following locations: GENCODE human lncRNAs are freely available at the website (https://
The following datasets were generated:
Zhu H. 2022. Knockout of the DNA binding domains in human-specific lncRNAs causes changed expression of target genes. NCBI Gene Expression Omnibus. GSE213231
Zhu H. 2024. A multi-cancer CRISPR/
Lin J. 2026. Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression. Zenodo.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 30 September 2026: the first record
Recorded: type, language, journal, volume, pages, dates, 6 authors, 1 keyword, 8 MeSH terms, 2 funders, 98 references.
Cite
This paper
Lin, J., Wen, Y., Tang, J., Zhang, X., Zhang, H., & Zhu, H. (2026). Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression. eLife, 12, RP89001. https://
BibTeX
@article{lin2026human,
author = {Lin, Jie and Wen, Yujian and Tang, Ji and Zhang, Xuecong and Zhang, Huanlin and Zhu, Hao},
title = {{Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression}},
journal = {eLife},
year = {2026},
month = mar,
volume = {12},
pages = {RP89001},
publisher = {eLife Sciences Publications, Ltd},
issn = {2050-084X},
doi = {10.7554/
url = {https://
pmid = {41823704},
pmcid = {PMC12987650}
}
RIS
TY - JOUR
AU - Lin, Jie
AU - Wen, Yujian
AU - Tang, Ji
AU - Zhang, Xuecong
AU - Zhang, Huanlin
AU - Zhu, Hao
TI - Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression
T2 - eLife
J2 - Elife
PY - 2026
DA - 2026/
VL - 12
SP - RP89001
SN - 2050-084X
PB - eLife Sciences Publications, Ltd
DO - 10.7554/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.7554/
"type": "article-journal",
"title": "Human-specific lncRNAs contributed critically to human evolution by distinctly regulating gene expression",
"container-title": "eLife",
"author": [
{
"family": "Lin",
"given": "Jie"
},
{
"family": "Wen",
"given": "Yujian"
},
{
"family": "Tang",
"given": "Ji"
},
{
"family": "Zhang",
"given": "Xuecong"
},
{
"family": "Zhang",
"given": "Huanlin"
},
{
"family": "Zhu",
"given": "Hao"
}
],
"container-title-short":
"volume": "12",
"page": "RP89001",
"DOI": "10.7554/
"PMID": "41823704",
"PMCID": "PMC12987650",
"ISSN": "2050-084X",
"publisher": "eLife Sciences Publications, Ltd",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
13
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s13059-026-04098-8 [code]
- A LINE-1 insertion upstream of FOXP2 promotes neuronal differentiation during primate evolution.Journal: Genome biologyIn common: pandas, SciPy, NumPy, 7 references
- [2] doi:10.1016/j.xhgg.2026.100629 [code]
- Positive selection on brain cis-regulatory elements in the human lineage drives changes in gene expression and susceptibility to neuropsychiatric disorders.Journal: HGG advancesIn common: pandas, SciPy, NumPy, cellular / molecular, 5 references
- [3] doi:10.1093/molbev/msag140 [code]
- Segmentally Duplicated Regulatory Elements Undergo Human-Specific Rewiring.Journal: Molecular biology and evolutionIn common: cellular / molecular, 5 references
- [4] doi:10.1016/j.xgen.2026.101278 [code]
- Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures.Journal: Cell genomicsIn common: pandas, NumPy, 4 references
- [5] doi:10.1038/s42003-026-10802-y
- Transcriptome of fetal cortex of tree shrew underlying the emergence of outer subventricular zone.Journal: Communications biologyIn common: 4 references
- [6] doi:10.1038/s41467-026-74153-2 [code]
- Regional, functional and transcriptomic decoding of multidimensional brain structure alterations in obsessive-compulsive disorder.Journal: Nature communicationsIn common: pandas, SciPy, NumPy, cellular / molecular, 3 references
- [7] doi:10.1093/gbe/evag086 [code]
- Genome Scans Reveal Species-Specific Selection in the Genus Lynx.Journal: Genome biology and evolutionIn common: pandas, cellular / molecular, 3 references
- [8] doi:10.1016/j.mocell.2026.100363 [code]
- Tandem repeats in human brain evolution and disease susceptibility.Journal: Molecules and cellsIn common: NumPy, cellular / molecular, 3 references
- [9] doi:10.1038/s41586-026-10699-x [code]
- Competing programs shape cortical sensorimotor-association
axis development. Journal: NatureIn common: NumPy, 4 references - [10] doi:10.1038/s41588-026-02646-3 [code]
- Co-expression-based models improve eQTL predictions for transcriptome-wide association studies and highlight new schizophrenia-associated
genes. Journal: Nature geneticsIn common: pandas, SciPy, NumPy, cellular / molecular, 2 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 1 script, and 2 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:611e7cefec2694d4…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
