Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics.
The 2 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
- [1] § STAR★Methods › Method details › Generation of heterozygous and homozygous SQSTM1 knockout PSC lines ↔ crispor.py, lines 2721–2779 · score 0.70 · spCas9, length polymorphism, mutation induced, enzyme, RFLP, CRISPOR
- [2] § STAR★Methods › Method details › Prazosin treatment of zebrafish with sqstm1 knockdown ↔ js/jquery-ui.min.js, the whole file · a weak match · score 0.52 · duration, Touch, escape, submitted, LI, fast
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 4,761 lines · 195 KB · other · 1 match
- #!/data/www/crispor/venv/bin/python3
- # if you do not want the hardcoded PATH above, delete this line and the one above to use the default Python3 interpreter
- #!/usr/bin/env python3
- # I know that this line looks unprofessional to you, but modifying the PATH on a shared Apache webserver is not obvious.
- # the tefor crispr tool
- # can be run as a CGI or from the command line
- # python std library
- import subprocess, tempfile, optparse, logging, atexit, glob, shutil, signal, pdb
- import http.cookies, time, sys, cgi, re, random, platform, os, pipes
- import hashlib, base64, string, logging, operator, urllib.request, urllib.parse, urllib.error, time
- import traceback, json, pwd, gzip, zlib
- from io import StringIO
- from collections import defaultdict, namedtuple
- from datetime import datetime
- from itertools import product
- from os.path import abspath, basename, dirname, isdir, isfile, join, relpath
- try:
- # prefer the pip package, it's more up-to-date than the native package
- import pysqlite3 as sqlite3
- SQLITEERROR=pysqlite3.dbapi2.OperationalError
- except:
- import sqlite3
- SQLITEERROR=sqlite3.OperationalError
- try:
- from collections import OrderedDict
- except ImportError:
- from ordereddict import OrderedDict # python2.6 users: run 'sudo pip install ordereddict'
- # for matplotlib, improves "import" performance
- os.environ["MPLCONFIGDIR"] = "/tmp/matplotlib-cache"
- # try to load external dependencies
- # we're going into great lengths to create a readable error message
- needModules = set(["pytabix", "twobitreader", "pandas", "matplotlib", "scipy"])
- try:
- import tabix # if not found, install with 'pip install pytabix'
- needModules.remove("pytabix")
- except:
- pass
- try:
- import twobitreader # if not found, install with 'pip install twobitreader'
- needModules.remove("twobitreader")
- except:
- pass
- try:
- import pandas # required by doench2016 score. install with 'pip install pandas'
- needModules.remove("pandas")
- import scipy # required by doench2016 score. install with 'pip install scipy'
- needModules.remove("scipy")
- import matplotlib # required by doench2016 score. install with 'pip install matplotlib'
- needModules.remove("matplotlib")
- import numpy # required by doench2016 score. install with 'pip install numpy'
- needModules.remove("numpy")
- except:
- pass
- if len(needModules)!=0:
- print("Content-type: text/html\n")
- print(("Python interpreter path: %s<p>" % sys.executable))
- print(("These python modules were not found: %s<p>" % ",".join(needModules)))
- print("To install all requirements in one line, run: sudo pip install biopython numpy scikit-learn==0.16.1 pandas twobitreader<p>")
- sys.exit(0)
- # our own eff scoring library
- import crisporEffScores
- # don't report print as an error
- # pylint: disable=E1601
- # optional module for Excel export as native .xls files
- # install with 'apt-get install python-xlwt' or 'pip install xlwt'
- xlwtLoaded = True
- try:
- import xlwt
- except:
- sys.stderr.write("crispor.py - warning - the python xlwt module is not available\n")
- xlwtLoaded = False
- # optional module for mysql support
- #try:
- #import MySQLdb
- #mysqldbLoaded = True
- #except:
- #mysqldbLoaded = False
- # version of crispor
- versionStr = "5.2"
- # contact email
- contactEmail='[email hidden]'
- # url to this server
- ctBaseUrl = "http://crispor-max.tefor.net/temp/customTracks"
- # write debug output to stdout
- DEBUG = False
- #DEBUG = True
- # use bowtie for off-target search?
- useBowtie = False
- # calculate the efficienc scores?
- doEffScoring = True
- # system-wide temporary directory
- #TEMPDIR = os.environ.get("TMPDIR", "/var/tmp")
- TEMPDIR = "/var/tmp"
- # a hack for cluster jobs at UCSC:
- # - default to ramdisk
- if isdir("/scratch/tmp"):
- TEMPDIR = "/dev/shm/"
- # skipAlign is useful if your input sequence is not in the genome at all
- # - don't do bwasw
- # - this will trigger auto-ontarget: any perfect match is the on-target
- # - do not calculate efficiency scores
- skipAlign = False
- # prefix in html statements before the directories "image/", "style/" and "js/"
- HTMLPREFIX = ""
- # alternative directory on local disk where image/, style/ and js/ are located
- HTMLDIR = "/usr/local/apache/htdocs/crispor/"
- # directory of crispor.py
- baseDir = dirname(__file__)
- # filename of this script, usually crispor.py
- myName = basename(__file__)
- # the segments.bed files use abbreviated genomic region names
- segTypeConv = {"ex":"exon", "in":"intron", "ig":"intergenic"}
- # directory for processed batches of offtargets ("cache" of bwa results)
- batchDir = join(baseDir,"temp")
- # sqlite3 db with gzipped old json batch files, to avoid hitting the ext4 inode limits
- batchArchive = "/data/crisporJobArchive.db"
- # the file where the sqlite job queue is stored
- #JOBQUEUEDB = join(TEMPDIR, "crisporJobs.db") # TEMPDIR is mapped away for security reasons under Redhat/Centos for CGIs
- JOBQUEUEDB = "/data/www/temp/crisporJobs.db"
- # alternatively: connection info for mysql
- jobQueueMysqlConn = {"socket":None, "host":None, "user": None, "password" : None}
- # directory for platform-independent scripts (e.g. Heng Li's perl SAM parser)
- scriptDir = join(baseDir, "scripts")
- # directory for helper binaries (e.g. BWA)
- # system() is one of 'Linux', 'Darwin', 'Windows', machine() is one of 'x86_64', 'arm64', 'aarch64'
- os_name = platform.system() # 'Linux', 'Darwin', 'Windows'
- arch = platform.machine() # 'x86_64', 'arm64', 'aarch64'
- binDir = abspath(join(baseDir, "bin", platform.system()+"-"+platform.machine()))
- # directory for genomes
- genomesDir = join(baseDir, "genomes")
- DEFAULTORG = 'hg19'
- DEFAULTSEQ = 'cttcctttgtccccaatctgggcgcgcgccggcgccccctggcggcctaaggactcggcgcgccggaagtggccagggcgggggcgacctcggctcacagcgcgcccggctattctcgcagctcaccatgGATGATGATATCGCCGCGCTCGTCGTCGACAACGGCTCCGGCATGTGCAAGGCCGGCTTCGCGGGCGACGATGCCCCCCGGGCCGTCTTCCCCTCCATCGTGGGGCGCC'
- # used if hg19 is not available
- ALTORG = 'sacCer3'
- ALTSEQ = 'ATTCTACTTTTCAACAATAATACATAAACatattggcttgtggtagCAACACTATCATGGTATCACTAACGTAAAAGTTCCTCAATATTGCAATTTGCTTGAACGGATGCTATTTCAGAATATTTCGTACTTACACAGGCCATACATTAGAATAATATGTCACATCACTGTCGTAACACTCT'
- pamDesc = [ ('NGG','20bp-NGG - Sp Cas9, SpCas9-HF1, eSpCas9 1.1'),
- ('NNG','20bp-NNG - Cas9 S. canis'),
- ('NGN','20bp-NGN - SpG'),
- ('NNGT','20bp-NNGT - Cas9 S. canis - high efficiency PAM, recommended'),
- ('NAA','20bp-NAA - iSpyMacCas9'),
- ('TTN', 'TTN-23bp - hfCas12Max - as recommended by Synthego'), # Casey Jowdy by email
- ('TNN','TNN-23bp - hfCas12Max, broader PAM, as recommended by Synthego'), #
- ('NGG-22', 'NGG-22bp - eSpOT-ON (ePsCas9), as recommended by Synthego'),
- ('NNGRRT','21bp-NNG(A/G)(A/G)T - Cas9 S. Aureus'),
- ('NNGRRT-20','20bp-NNG(A/G)(A/G)T - Cas9 S. Aureus with 20bp-guides'),
- ('NGK','20bp-NG(G/T) - xCas9, recommended PAM, see notes'),
- #('NGN','20bp-NGN or GA(A/T) - xCas9 (low efficiency, not recommended)'),
- #('NGG-BE1','20bp-NGG - BaseEditor1, modifies C->T'),
- ('NNNRRT','21bp-NNN(A/G)(A/G)T - KKH SaCas9'),
- ('NNNRRT-20','20bp-NNN(A/G)(A/G)T - KKH SaCas9 with 20bp-guides'),
- ('NGA','20bp-NGA - Cas9 S. Pyogenes mutant VQR'),
- ('NNNNCC','24bp-NNNNCC - Nme2Cas9'),
- ('NGCG','20bp-NGCG - Cas9 S. Pyogenes mutant VRER'),
- ('NNAGAA','20bp-NNAGAA - Cas9 S. Thermophilus'),
- ('NGGNG','20bp-NGGNG - Cas9 S. Thermophilus'),
- ('NNNNGMTT','20bp-NNNNG(A/C)TT - Cas9 N. Meningitidis'),
- ('NNNNACA','20bp-NNNNACA - Cas9 Campylobacter jejuni, original PAM'),
- ('NNNNRYAC','22bp-NNNNRYAC - Cas9 Campylobacter jejuni, revised PAM'),
- ('NNNVRYAC','22bp-NNNVRYAC - Cas9 Campylobacter jejuni, opt. efficiency'),
- ('TTCN','TTCN-20bp - CasX'),
- ('TTTV','TTT(A/C/G)-23bp - Cas12a (Cpf1) - recommended, 23bp guides'),
- ('TTTV-21','TTT(A/C/G)-21bp - Cas12a (Cpf1) - 21bp guides recommended by IDT'),
- ('TTTN','TTTN-23bp - Cas12a (Cpf1) - low efficiency'),
- ('ATTN','ATTN-23bp - BhCas12b v4'),
- ('NGTN','NGTN-23bp - ShCAST/AcCAST, Strecker et al, Science 2019'),
- ('TYCV','T(C/T)C(A/C/G)-23bp - TYCV As-Cpf1 K607R'),
- ('TATV','TAT(A/C/G)-23bp - TATV As-Cpf1 K548V'),
- ('TTTA','TTTA-23bp - TTTA LbCpf1'),
- ('TCTA','TCTA-23bp - TCTA LbCpf1'),
- ('TCCA','TCCA-23bp - TCCA LbCpf1'),
- ('CCCA','CCCA-23bp - CCCA LbCpf1'),
- ('GGTT','GGTT-23bp - CCCA LbCpf1'),
- ('YTTV','YTTV-20bp - MAD7 Nuclease, Lui, Schiel, Maksimova et al, CRISPR J 2020'),
- ('TTYN','TTYN- or VTTV- or TRTV-23bp - enCas12a E174R/S542R/K548R - Kleinstiver et al Nat Biot 2019'),
- ('NNNNCNAA','20bp-NNNNCNAA - Thermo Cas9 - Walker et al, Metab Eng Comm 2020'),
- ('NNN','20bp-NNN - SpRY, Walton et al Science 2020'), # https://science.sciencemag.org/content/368/6488/290.abstract
- ('NRN','20bp-NRN - SpRY (high efficiency PAM)'),
- ('NYN','20bp-NYN - SpRY (low efficiency PAM)'),
- #('VTTV','(A/C)TT(A/C)-23bp - enCas12a S542R - Kleinstiver et al Nat Biot 2019'),
- #('TRTV','T(A/G)T(A/C)-23bp - enCas12a K548R - Kleinstiver et al Nat Biot 2019'),
- ]
- DEFAULTPAM = 'NGG'
- # the default base editor modification window
- DEFAULTBEWIN = "1-7"
- # for some PAMs, there are alternative main PAMs. These are also shown on the main sequence panel
- multiPams = {
- #"NGN" : ["GAW"],
- "TTYN" : ["VTTV", "TRTV"]
- }
- # these PAMs are not specific. Allow only short sequences for them.
- slowPams = ["TTYN", "NNG"]
- # allow only very short sequences for these
- verySlowPams = ["NNN", "NRN", "NYN"]
- # for some PAMs, we allow other alternative motifs when searching for offtargets
- # MIT and eCrisp do that, they use the motif NGG + NAG, we add one more, based on the
- # on the guideSeq results in Tsai et al, Nat Biot 2014
- # The NGA -> NGG rule was described by Kleinstiver...Young 2015 "Improved Cas9 Specificity..."
- # NNGTRRT rule for S. aureus is in the new protocol "SaCas9 User manual"
- # ! the length of the alternate PAM has to be the same as the original PAM!
- offtargetPams = {
- "NGG" : ["NAG","NGA"],
- #"NGN" : ["GAW"],
- "NGK" : ["GAW"],
- "NGA" : ["NGG"],
- "NNGRRT" : ["NNGRRN"],
- "TTTV" : ["TTTN"],
- 'ATTN' : ["TTTN", "GTTN"],
- "TTYN" : ["VTTV", "TRTV"]
- }
- # maximum size of an input sequence
- MAXSEQLEN = 2300
- # maximum input size when specifying "no genome"
- MAXSEQLEN_NOGENOME = 25000
- # maximum input size when using xCas9 or sCanis
- MAXSEQLEN2 = 600
- # maximum input size for NNN SpRY or similar PAMs
- MAXSEQLEN3 = 150
- # BWA: allow up to X mismatches
- maxMMs=4
- # maximum number of occurences in the genome to get flagged as repeats.
- # This is used in bwa samse, when converting the same file
- # and for warnings in the table output.
- MAXOCC = 60000
- # the BWA queue size is 2M by default. We derive the queue size from MAXOCC
- MFAC = 2000000/MAXOCC
- # the length of the guide sequence, set by setupPamInfo
- GUIDELEN=None
- # length of the PAM sequence
- PAMLEN=None
- # the name of the base editor, if any. This is the flag to activate
- # baseEditor mode in the UI
- baseEditor = None
- # input sequences are extended by X basepairs so we can calculate the efficiency scores
- # and can better design primers
- FLANKLEN=100
- # the name of the currently processed batch, assigned only once
- # in readBatchParams and only for json-type batches
- batchName = ""
- # are we doing a Cpf1 run?
- # this variable changes almost all processing and
- # has to be set on program start, as soon as we know
- # the PAM we're running on
- pamIsFirst=None
- saCas9Mode=False
- # Highly-sensitive mode (not for CLI mode):
- # MAXOCC is increased in processSubmission() and in the html UI if only one
- # guide seq is run
- # Also, the number of allowed mismatches is increased to 5 instead of 4
- #HIGH_MAXOCC=600000
- #HIGH_maxMMs=5
- # minimum off-target score of standard off-targets (those that end with NGG)
- # This should probably be based on the CFD score these days
- # But for now, I'll let the user do the filtering
- MINSCORE = 0.0
- # minimum off-target score for alternative PAM off-targets
- # There is not a lot of data to support this cutoff, but it seems
- # reasonable to have at least some cutoff, as otherwise we would show
- # NAG and NGA like NGG and the data shows clearly that the alternative
- # PAMs are not recognized as well as the main NGG PAM.
- # so for now, I just filter out very degenerative ones. the best solution
- # would be to have a special penalty on the CFD score, but CFS does not
- # support non-NGG PAMs (is this actually true?)
- ALTPAMMINSCORE = 1.0
- # how much shall we extend the guide after the PAM to match restriction enzymes?
- pamPlusLen = 5
- # global flag to indicate if we're run from command line or as a CGI
- commandLineMode = False
- # names/order of efficiency scores to show in UI
- cas9ScoreNames = ["fusi", "crisprScan", "rs3"]
- allScoreNames = ["fusi", "chariRank", "ssc", "wuCrispr", "doench", "wang", "crisprScan", "ccTop", "rs3"]
- mutScoreNames = []
- spCas9MutScoreNames = ["oof", 'lindel'] # lindel is only added for spCas9
- otherMutScoreNames = ["oof"] # lindel is only added for spCas9
- cpf1ScoreNames = ["seqDeepCpf1"]
- saCas9ScoreNames = ["najm"]
- # to make the CFD more comparable to the MIT score, Nicholas Parkinson suggests to multiply it with 100.
- # can be switched on with the URL argument fixCfd=1
- doCfdFix=False
- # how many digits shall we show for each score? default is 0
- scoreDigits = {
- "ssc" : 1,
- }
- # List of AddGene plasmids, their long and short names:
- addGenePlasmids = [
- ("43860", ("MLM3636 (Joung lab)", "MLM3636")),
- ("49330", ("pAc-sgRNA-Cas9 (Liu lab)", "pAcsgRnaCas9")),
- ("42230", ("pX330-U6-Chimeric_BB-CBh-hSpCas9 (Zhang lab) + derivatives", "pX330")),
- ("52961", ("lentiCRISPR v2 (Zhang lab)", "lentiCrispr")),
- ("52963", ("lentiGuide-Puro (Zhang lab)", "lentiGuide-Puro")),
- ]
- addGenePlasmidsAureus = [
- ("61591", ("pX601-AAV-CMV::NLS-SaCas9-NLS-3xHA-bGHpA;U6::BsaI-sgRNA (Zhang lab)", "pX601")),
- ("61592", ("pX600-AAV-CMV::NLS-SaCas9-NLS-3xHA-bGHpA (Zhang lab)", "pX600")),
- ("61593", ("pX602-AAV-TBG::NLS-SaCas9-NLS-HA-OLLAS-bGHpA;U6::BsaI-sgRNA (Zhang lab)", "pX602")),
- ("65779", ("VVT1 (Joung lab)", "VVT1"))
- ]
- # list of AddGene primer 5' and 3' extensions, one for each AddGene plasmid
- # format: prefixFw, prefixRw, u6-G-suffix, restriction enzyme, link to protocol
- addGenePlasmidInfo = {
- "43860" : ("ACACC", "AAAAC", "G", "BsmBI", "https://www.addgene.org/static/data/plasmids/43/43860/43860-attachment_T35tt6ebKxov.pdf"),
- "49330" : ("TTC", "AAC", "", "Bsp QI", "http://bio.biologists.org/content/3/1/42#sec-9"),
- "42230" : ("CACC", "AAAC", "", "Bbs1", "https://www.addgene.org/static/data/plasmids/52/52961/52961-attachment_B3xTwla0bkYD.pdf"),
- "52961" : ("CACC", "AAAC", "", "BsmBI", "https://www.addgene.org/static/data/plasmids/52/52961/52961-attachment_B3xTwla0bkYD.pdf"),
- "61591" : ("CACC", "AAAC", "", "BsaI", "https://www.addgene.org/static/data/plasmids/61/61591/61591-attachment_it03kn5x5O6E.pdf"),
- "61592" : ("CACC", "AAAC", "", "BsaI", "https://www.addgene.org/static/data/plasmids/61/61592/61592-attachment_iAbvIKnbqNRO.pdf"),
- "61593" : ("CACC", "AAAC", "", "BsaI", "https://www.addgene.org/static/data/plasmids/61/61592/61592-attachment_iAbvIKnbqNRO.pdf"),
- "65779": ("CACC", "AAAC", "", "BsmBI (aka Esp3l)", "https://www.addgene.org/static/data/plasmids/65/65779/65779-attachment_G8oNyvV6pA78.pdf"),
- "52963": ("CACC", "AAAC", "", "BsmBI (aka Esp3l)", "https://www.addgene.org/static/data/plasmids/52/52963/52963-attachment_IPB7ZL_hJcbm.pdf")
- }
- # the barcodes for subpool tagging for oligo pool tables
- satMutBarcodes = [
- (0, "No Subpool barcode"),
- (1, "Subpool 1: CGGGTTCCGT/GCTTAGAATAGAA"),
- (2, "Subpool 2: GTTTATCGGGC/ACTTACTGTACC"),
- (3, "Subpool 3: ACCGATGTTGAC/CTCGTAATAGC"),
- (4, "Subpool 4: GAGGTCTTTCATGC/CACAACATA"),
- (5, "Subpool 5: TATCCCGTGAAGCT/TTCGGTTAA"),
- (6, "Subpool 6: TAGTAGTTCAGACGC/ATGTACCC"),
- (7, "Subpool 7: GGATGCATGATCTAG/CATCAAGC"),
- (8, "Subpool 8: ATGAGGACGAATCT/CACCTAAAG"),
- (9, "Subpool 9: GGTAGGCACG/TAAACTTAGAACC"),
- (10, "Subpool 10: AGTCATGATTCAG/GTTGCAAGTCTAG"),
- ]
- # Restriction enzyme supplier codes
- rebaseSuppliers = {
- "B":"Life Technologies",
- "C":"Minotech",
- "E":"Agilent",
- "I":"SibEnzyme",
- "J":"Nippon Gene",
- "K":"Takara",
- "M":"Roche",
- "N":"NEB",
- "O":"Toyobo",
- "Q":"Molecular Biology Resources",
- "R":"Promega",
- "S":"Sigma",
- "V":"Vivantis",
- "X":"EURx",
- "Y":"SinaClon BioScience"
- }
- # labels and descriptions of eff. scores
- scoreDescs = {
- "doench" : ("Doench '14", "Range: 0-100. Linear regression model trained on 880 guides transfected into human MOLM13/NB4/TF1 cells (three genes) and mouse cells (six genes). Delivery: lentivirus. The Fusi score can be considered an updated version this score, as their training data overlaps a lot. See <a target='_blank' href='http://www.nature.com/nbt/journal/v32/n12/full/nbt.3026.html'>Doench et al.</a>"),
- "wuCrispr" : ("Wu-Crispr", "Range 0-100. Aka 'Wong score'. SVM model trained on previously published data. The aim is to identify only a subset of efficient guides, many guides will have a score of 0. Takes into account RNA structure. See <a target='_blank' href='https://genomebiology.biomedcentral.com/articles/10.1186/s13059-015-0784-0'>Wong et al., Gen Biol 2015</a>"),
- "ssc" : ("Xu", "Range ~ -2 - +2. Aka 'SSC score'. Linear regression model trained on data from >1000 genes in human KBM7/HL60 cells (Wang et al) and mouse (Koike-Yusa et al.). Delivery: lentivirus. Ranges mostly -2 to +2. See <a target='_blank' href='http://genome.cshlp.org/content/early/2015/06/10/gr.191452.115'>Xu et al.</a>"),
- "crisprScan" : ["Moreno-Mateos", "Also called 'CrisprScan'. Range: mostly 0-100. Linear regression model, trained on data from 1000 guides on >100 genes, from zebrafish 1-cell stage embryos injected with mRNA. See <a target=_blank href='http://www.nature.com/nmeth/journal/v12/n10/full/nmeth.3543.html'>Moreno-Mateos et al.</a>. Recommended for guides transcribed <i>in-vitro</i> (T7 promoter). Click to sort by this score. Note that under 'Show all scores', you can find a Doench2016 model trained on Zebrafish scores, Azimuth in-vitro, which should be slightly better than this model for zebrafish."],
- "wang" : ("Wang", "Range: 0-100. SVM model trained on human cell culture data on guides from >1000 genes. The Xu score can be considered an updated version of this score, as the training data overlaps a lot. Delivery: lentivirus. See <a target='_blank' href='http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3972032/'>Wang et al.</a>"),
- "chariRank" : ("Chari", "Range: 0-100. Support Vector Machine, converted to rank-percent, trained on data from 1235 guides targeting sequences that were also transfected with a lentivirus into human 293T cells. See <a target='_blank' href='http://www.nature.com/nmeth/journal/v12/n9/abs/nmeth.3473.html'>Chari et al.</a>"),
- "fusi" : ("Doench '16", "Aka the 'Fusi-Score', since V4.4 using the version 'Azimuth', scores are slightly different than before April 2018 but very similar (click 'show all' to see the old scores). Range: 0-100. Boosted Regression Tree model, trained on data produced by Doench et al (881 guides, MOLM13/NB4/TF1 cells + unpublished additional data). Delivery: lentivirus. See <a target='_blank' href='http://biorxiv.org/content/early/2015/06/26/021568'>Fusi et al. 2015</a> and <a target='_blank' href='http://www.nature.com/nbt/journal/v34/n2/full/nbt.3437.html'>Doench et al. 2016</a> and <a target=_blank href='https://crispr.ml/'>crispr.ml</a>. Recommended for guides expressed in cells (U6 promoter). Click to sort the table by this score."),
- "fusiOld" : ("OldDoench '16", "The original implementation of the Doench 2016 score, as received from John Doench. The scores are similar, but not exactly identical to the 'Azimuth' version of the Doench 2016 model that is currently the default on this site, since Apr 2018."),
- "rs3" : ("Doench-RuleSet3", "The Doench Rule Set 3 (RS3) score (-200-+200). Similar to the Doench 2014 and Doench 2016/Fusi/Azimuth score, but updated and more accurate. See <a href='https://www.nature.com/articles/s41467-022-33024-2' target=_blank>. Scores shown are multiplied with 100 for easier display. RS3 is configured here to use the Hsu-TRACR sequence."),
- "najm" : ("Najm 2018", "A modified version of the Doench 2016 score ('Azimuth'), by Mudra Hegde for S. aureus Cas9. Range 0-100. See <a target=_blank href='https://www.nature.com/articles/nbt.4048'>Najm et al 2018</a>."),
- "ccTop" : ("CCTop", "The efficiency score used by CCTop, called 'crisprRank'."),
- "aziInVitro" : ("Azimuth in-vitro", "The Doench 2016 model trained on the Moreno-Mateos zebrafish data. Unpublished model, gratefully provided by J. Listgarden. This should be better than Moreno-Mateos, but we have not found the time to evaluate it yet."),
- "housden" : ("Housden", "Range: ~ 1-10. Weight matrix model trained on data from Drosophila mRNA injections. See <a target='_blank' href='http://stke.sciencemag.org/content/8/393/rs9.long'>Housden et al.</a>"),
- "proxGc" : ("ProxGCCount", "Number of GCs in the last 4pb before the PAM"),
- "seqDeepCpf1" : ("DeepCpf1", "Range: ~ 0-100. Convolutional Neural Network trained on ~20k Cpf1 lentiviral guide results. This is the score without DNAse information, 'Seq-DeepCpf1' in the paper. See <a target='_blank' href='https://www.nature.com/articles/nbt.4061'>Kim et al. 2018</a>"),
- "oof" : ("Out-of-Frame", "Range: 0-100. Out-of-Frame score, only for deletions. Predicts the percentage of clones that will carry out-of-frame deletions, based on the micro-homology in the sequence flanking the target site. See <a target='_blank' href='http://www.nature.com/nmeth/journal/v11/n7/full/nmeth.3015.html'>Bae et al. 2014</a>. Click the score to show the predicted deletions."),
- "lindel": ("Lindel", "Wei Chen Frameshift ratio (0-100). Predicts probability of a frameshift caused by any type of insertion or deletion. See <a href='https://academic.oup.com/nar/article/47/15/7989/5511473'>Wei Chen et al, Bioinf 2018</a>. Click the score to see the most likely deletions and insertions."),
- }
- # the headers for the guide and offtarget output files
- guideHeaders = ["guideId", "targetSeq", "mitSpecScore", "cfdSpecScore", "offtargetCount", "targetGenomeGeneLocus"]
- offtargetHeaders = ["guideId", "guideSeq", "offtargetSeq", "mismatchPos", "mismatchCount", "mitOfftargetScore", "cfdOfftargetScore", "chrom", "start", "end", "strand", "locusDesc"]
- # library descriptions
- libLabels = [
- # https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4486245/
- ("human_brunello" , "Human, Brunello, Doench Nat Bio 2016 (recommended)"),
- ("human_avana" , "Human, Avana, Doench Nat Bio 2016"),
- ("human_geckov2" , "Human, GeCKO V2, Sanjana Nat Meth 2014"),
- ("mouse_brie" , "Mouse, Brie, Doench Nat Bio 2016 (recommended)"),
- ("mouse_geckov2" , "Mouse, GeCKO V2, Sanjana Nat Meth 2014"),
- ("mouse_asiago" , "Mouse, Asiago, Doench Nat Bio 2016"),
- ]
- # a file crispor.conf in the directory of the script allows to override any global variable
- myDir = dirname(__file__)
- confPath =join(myDir, "crispor.conf")
- if isfile(confPath):
- exec(open(confPath).read())
- #execfile(confPath)
- cgiParams = None
- # ====== END GLOBALS ============
- def setupPamInfo(pam):
- " modify a few globals based on the current pam "
- global GUIDELEN
- global pamIsFirst
- global addGenePlasmids
- global PAMLEN
- global scoreNames
- global baseEditor
- global saCas9Mode
- global mutScoreNames
- global isSpg
- PAMLEN = len(pam)
- pamIsFirst = False
- scoreNames = cas9ScoreNames
- pamOpt = None
- if "-" in pam:
- pam, pamOpt = pam.split("-")
- if pamOpt=="BE1":
- baseEditor = "BE1"
- elif pamOpt=="spg":
- isSpg = True
- if pamIsCasX(pam):
- logging.debug("switching on CasX mode, guide length is 20bp")
- GUIDELEN = 20
- pamIsFirst = True
- scoreNames = cpf1ScoreNames
- if pamIsCas12max(pam):
- logging.debug("switching on hfCas12max mode, guide length is 20bp")
- GUIDELEN = 20
- pamIsFirst = True
- scoreNames = cpf1ScoreNames
- elif pamIsCpf1(pam):
- logging.debug("switching on Cpf1 mode, guide length is 23bp")
- GUIDELEN = 23
- pamIsFirst = True
- scoreNames = cpf1ScoreNames
- #if pamOpt:
- #GUIDELEN=int(pamOpt)
- elif pam=="NGTN":
- logging.debug("switching on Cpf1 mode for ShCAST, guide length is 23bp")
- GUIDELEN = 23
- pamIsFirst = True
- elif pam=="NNNNRYAC" or pam=="NNNVRYAC":
- GUIDELEN = 22
- elif pam=="NNGRRT" or pam=="NNNRRT":
- logging.debug("switching on S. aureus mode, guide length is 21bp")
- addGenePlasmids = addGenePlasmidsAureus
- GUIDELEN = 21
- #if pamOpt=="20":
- #GUIDELEN=20
- saCas9Mode = True
- scoreNames = saCas9ScoreNames
- elif pam=="NNNNCC":
- GUIDELEN = 24
- else:
- GUIDELEN = 20
- if pamOpt and pamOpt.isnumeric():
- GUIDELEN = int(pamOpt)
- if (GUIDELEN==20 or GUIDELEN==22) and pam=="NGG":
- mutScoreNames = spCas9MutScoreNames
- else:
- mutScoreNames = otherMutScoreNames
- logging.debug("Enzyme info: pam=%s, guideLen=%d, pamIsFirst=%s, saCas9Mode=%s" %
- (pam, GUIDELEN, pamIsFirst, saCas9Mode))
- return pam
- # ==== CLASSES =====
- class JobQueue:
- """
- simple job queue, using a db table as a backend
- jobs have different types and status. status can be updated while they run
- job running times are kept and old job info is kept in a separate table
- >>> q = JobQueue()
- >>> q.openSqlite()
- >>> q.clearJobs()
- >>> q.waitCount()
- 0
- >>> q.addJob("search", "abc123", "myParams")
- True
- only one job per jobId
- >>> q.addJob("search", "abc123", "myParams")
- False
- >>> q.waitCount()
- 1
- >>> q.getStatus("abc123")
- 'Waiting'
- >>> q.startStep("abc123", "bwa", "Alignment with BWA")
- >>> q.getStatus("abc123")
- 'Alignment with BWA'
- >>> jobType, jobId, paramStr = q.popJob()
- >>> q.waitCount()
- 0
- >>> q.jobDone("abc123")
- >>> q.waitCount()
- 0
- can't pop from an empty queue
- #>>> q.popJob()
- #(None, None, None)
- #>>> os.system("rm /tmp/tempCrisporTest.db")
- #0
- """
- _queueDef = (
- 'CREATE TABLE IF NOT EXISTS %s '
- '('
- ' jobType text,' # either "index" or "search"
- ' jobId text %s,' # unique identifier
- ' paramStr text,' # parameters for jobs, like db, options, etc.
- ' isRunning int DEFAULT 0,' # indicates steps have started, done jobs are moved to doneJobs table
- ' stepName text,' # currently step, internal step name for timings
- ' stepLabel text,' # current step, human-readable status of job, for UI
- ' lastUpdate float,' # time of last update
- ' stepTimes text,' # comma-sep list of whole msecs, one per step
- ' startTime text ' # date+time when job was put into queue
- ')')
- #def __init__(self):
- #" no inheritance needed here "
- #self.openSqlite(JOBQUEUEDB)
- def openSqlite(self, dbName=JOBQUEUEDB):
- self.dbName = dbName
- # isolation_level=None = autocommit mode: we manage transactions explicitly
- # timeout=30: wait up to 30s for locks (multiple daemons + CGI share this DB)
- self.conn = sqlite3.connect(dbName, timeout=30, isolation_level=None)
- # WAL mode allows concurrent readers + writer, essential for CGI + daemon(s)
- result = self.conn.execute("PRAGMA journal_mode=WAL;").fetchone()
- if result[0] != "wal":
- logging.warn("Could not set WAL mode, journal_mode is: %s (is another connection open with the old mode?)" % result[0])
- #self.conn.set_trace_callback(print) # for debugging: print all sql statements
- self._chmodJobDb()
- try:
- self.conn.execute(self._queueDef % ("queue", "PRIMARY KEY"))
- except SQLITEERROR as ex:
- errAbort("cannot open the sqlite jobs file %s: %s" % (JOBQUEUEDB, ex))
- def _chmodJobDb(self):
- # umask is not respected by sqlite, bug http://www.mail-archive.com/[email hidden]/msg59080.html
- try:
- os.chmod(JOBQUEUEDB, 0o666)
- except OSError:
- # if the file was created by other job, we can't chmod, as we're the CGI. Just silently ignore this
- pass
- def addJob(self, jobType, jobId, paramStr):
- " create a new job, returns False if not successful "
- self._chmodJobDb()
- sql = 'INSERT INTO queue (jobType, jobId, isRunning, lastUpdate, ' \
- 'stepTimes, paramStr, stepName, stepLabel, startTime) VALUES (:jobType, :jobId, :isRunning, :lastUpdate, ' \
- ':stepTimes, :paramStr, :stepName, :stepLabel, :startTime)'
- now = "%.3f" % time.time()
- values = {'jobType' : jobType, 'jobId' : jobId, 'isRunning' : 0, 'lastUpdate' : now, 'stepTimes':"", 'paramStr':paramStr, 'stepName':"wait",
- "stepLabel":"Waiting", "startTime":now}
- try:
- # in autocommit mode, this INSERT commits immediately
- self.conn.execute(sql, values)
- return True
- except sqlite3.IntegrityError:
- # job already in queue (e.g. resubmit, or daemon restart) - that's fine
- return True
- except SQLITEERROR:
- errAbort("Cannot open DB file %s. Please contact %s" % (self.dbName, contactEmail))
- def getStatus(self, jobId):
- " return current job status label or None if job is not in queue"
- sql = 'SELECT stepLabel FROM queue WHERE jobId=?'
- try:
- rows = self.conn.execute(sql, (jobId,)).fetchmany(1)
- if len(rows) == 0:
- status = None
- else:
- status = rows[0][0]
- except (StopIteration, IndexError):
- logging.debug("getStatus: job %s not found" % jobId)
- status = None
- return status
- def dump(self):
- " for debugging, write the whole queue table to stdout "
- sql = 'SELECT * FROM queue'
- for row in self.conn.execute(sql):
- print("\t".join([str(x) for x in row]))
- def jobInfo(self, jobId, isDone=False):
- " for debugging, return all job info as a tuple "
- print("job info<br>")
- if isDone:
- sql = 'SELECT * FROM doneJobs WHERE jobId=?'
- else:
- sql = 'SELECT * FROM queue WHERE jobId=?'
- try:
- row = next(self.conn.execute(sql, (jobId,)))
- except StopIteration:
- return []
- return row
- def startStep(self, jobId, newName, newLabel):
- " start a new step. Update lastUpdate, status and stepTime "
- try:
- self.conn.execute('BEGIN IMMEDIATE')
- sql = 'SELECT lastUpdate, stepTimes, stepName FROM queue WHERE jobId=?'
- logging.debug(sql)
- rows = self.conn.execute(sql, (jobId,)).fetchmany(1)
- if len(rows)==0:
- logging.error("startStep: no row for jobId %s" % jobId)
- self.conn.commit()
- return
- lastTime, timeStr, lastStep = rows[0]
- lastTime = float(lastTime)
- # append a string in format "stepName:milliSecs" to the timeStr
- now = time.time()
- timeDiff = "%d" % int((1000.0*(now - lastTime)))
- newTimeStr = timeStr+"%s=%s" % (lastStep, timeDiff)+","
- sql = 'UPDATE queue SET lastUpdate=?, stepName=?, stepLabel=?, stepTimes=?, isRunning=? WHERE jobId=?'
- self.conn.execute(sql, (now, newName, newLabel, newTimeStr, 1, jobId))
- self.conn.commit()
- except:
- self.conn.rollback()
- raise
- def jobDone(self, jobId):
- " remove the job from the queue and add it to the queue log"
- print("job done<br>")
- try:
- self.conn.execute('BEGIN IMMEDIATE')
- sql = 'SELECT * FROM queue WHERE jobId=?'
- row = self.conn.execute(sql, (jobId,)).fetchone()
- if row is None:
- logging.warn("jobDone - job %s has been removed already" % jobId)
- self.conn.commit()
- return
- sql = 'DELETE FROM queue WHERE jobId=?'
- self.conn.execute(sql, (jobId,))
- self.conn.commit()
- except:
- self.conn.rollback()
- raise
- # good to have a log file of the old jobs
- with open("doneJobs.tsv", "a") as ofh: # if this triggers an error: run 'touch doneJobs.tsv && chmod a+rw doneJobs.tsv' in the crispor dir.
- row = [str(x) for x in row]
- line = "\t".join(row)
- ofh.write(line)
- ofh.write("\n")
- def waitCount(self):
- " return number of waiting jobs "
- sql = 'SELECT count(*) FROM queue WHERE isRunning=0'
- return self.conn.execute(sql).fetchone()[0]
- def popJob(self):
- " return (jobType, jobId, params) of first waiting job and set it to running state "
- print('pop job<br>')
- try:
- self.conn.execute('BEGIN IMMEDIATE')
- sql = 'SELECT jobType, jobId, paramStr FROM queue WHERE isRunning=0 ORDER BY lastUpdate LIMIT 1'
- row = self.conn.execute(sql).fetchone()
- if row is None:
- logging.debug("popJob: no waiting jobs")
- self.conn.commit()
- return None, None, None
- jobType, jobId, paramStr = row
- sql = 'UPDATE queue SET isRunning=1 where jobId=?'
- self.conn.execute(sql, (jobId,))
- self.conn.commit()
- except:
- self.conn.rollback()
- raise
- return jobType, jobId, paramStr
- def clearJobs(self):
- " clear the job table, removing running jobs, too "
- self.conn.execute("DELETE from queue")
- def close(self):
- " "
- self.conn.close()
- # ====== FUNCTIONS =====
- contentLineDone = False
- # the queue workers should be able to never abort
- doAbort = True
- def getTwoBitFname(db):
- " return the name of the twoBit file for a genome "
- # at UCSC, try to use local disk, if possible
- locPath = join("/scratch", "data", db, db+".2bit")
- if isfile(locPath):
- return locPath
- path = join(genomesDir, db, db+".2bit")
- return path
- def errAbort(msg, isWarn=False):
- " print err msg and exit "
- if commandLineMode:
- raise Exception(msg)
- if not contentLineDone:
- print("Content-type: text/html\n")
- print('<div style="position: absolute; padding: 10px; left: 100; top: 100; border: 10px solid black; background-color: white; text-align:left; width: 800px; font-size: 18px">')
- if isWarn:
- print("<strong>Warning:</strong><p> ")
- else:
- print("<strong>Error:</strong><p> ")
- print((msg+"<p>"))
- print(("If you think this is a bug or you have any other suggestions, please do not hesitate to contact us %s<p>" % contactEmail))
- if isWarn:
- print("In the email, please also send us the full URL of the page.")
- else:
- print("Please also send us the full URL of the page where you see the error. Thanks!")
- print('</div>')
- if doAbort:
- sys.exit(0) # cgi must not exit with 1
- # allow only dashes, digits, characters, underscores and colons in the CGI parameters
- # and +
- notOkChars = re.compile(r'[^+a-zA-Z0-9/:\n\r_. -]')
- def checkVal(key, inStr):
- """ remove special characters from input string, to protect against injection attacks """
- if key!="geneIds":
- if len(inStr) > 10000:
- errAbort("input parameter %s is too long" % key)
- else:
- if len(inStr) > 100000:
- errAbort("Pasting more than tens of thousands of gene IDs makes little sense. Copy/paste error?")
- matchObj =notOkChars.search(inStr)
- if matchObj!=None:
- errAbort("input parameter %s contains an invalid character %s (ASCII %d)" % (key, repr(matchObj.group()), ord(matchObj.group())))
- return inStr
- def cgiGetParams():
- " get CGI parameters and return as dict "
- form = cgi.FieldStorage()
- global cgiParams
- cgiParams = {}
- # parameters are:
- #"pamId", "batchId", "pam", "seq", "org", "download", "sortBy", "format", "ajax
- for key in list(form.keys()):
- val = form.getfirst(key)
- if val!=None:
- # "seq" is cleaned by cleanSeq later
- val = urllib.parse.unquote(val)
- if key not in ["seq", "name"]:
- checkVal(key, val)
- cgiParams[key] = val
- if "pam" in cgiParams:
- legalChars = set("ACTGNMKRYVBE120345/-")
- illegalChars = set(cgiParams["pam"])-legalChars
- if len(illegalChars)!=0:
- errAbort("Illegal character in PAM-sequence. Only %s are allowed."+"".join(legalChars))
- if "batchId" in cgiParams:
- batchId = cgiParams["batchId"]
- if not batchId.isalnum() or len(batchId) > 30:
- errAbort("Invalid batchId")
- return cgiParams
- def cgiGetStr(params, argName, default=None):
- val = params.get(argName, None)
- if val==None and default==None:
- errAbort("'%s' parameter must be specified" % argName)
- if val==None:
- return default
- return val
- def cgiGetNum(params, argName, default):
- " get CGI parameter which must be a number "
- val = params.get(argName, None)
- if val==None:
- return default
- if not val.isdigit():
- errAbort("'%s' parameter must be a number" % argName)
- val = int(val)
- return val
- transTab = str.maketrans("-=/+_", "abcde")
- def makeTempBase(seq, org, pam, batchName):
- "create the base name of temp files using a hash function and some prettyfication "
- hasher = hashlib.sha1(seq.encode("latin1")+org.encode("latin1")+pam.encode("latin1")+batchName.encode("latin1"))
- shortHash = hasher.digest()[0:20]
- batchId = base64.urlsafe_b64encode(shortHash).decode('latin1').translate(transTab)[:20]
- return batchId
- def makeTempFile(prefix, suffix):
- " return a temporary file that is deleted upon exit, unless DEBUG is set "
- if DEBUG:
- fname = join("/tmp", prefix+suffix)
- fh = open(fname, "wt")
- else:
- fh = tempfile.NamedTemporaryFile(mode="wt", dir=TEMPDIR, prefix="primer3In", suffix=".txt")
- return fh
- def pamIsCpf1(pam):
- " if you change this, also change bin/filterFaToBed and bin/samToBed!!! "
- return (pam in ["TNN", "TTN", "TTTN", "TYCV", "TATV", "TTTV", "TTTR", "ATTN", "TTTA", "TCTA", "TCCA", "CCCA", "YTTV", "TTYN"])
- # test : modify cleavage sites for hfCas12max according to Synthego specifications
- def pamIsCas12max(pam):
- return (pam in ["TNN", "TTN"])
- def pamIsCasX(pam):
- " if you change this, also change bin/filterFaToBed and bin/samToBed!!! "
- return (pam in ["TTCN"])
- def pamIsSaCas9(pam):
- " only used for notes and efficiency scores, unlike its Cpf1 cousin function "
- return (pam.split("-")[0] in ["NNGRRT", "NNNRRT"])
- def isSlowPam(pam):
- " do not allow input sequences > 500 bp "
- if pamIsXCas9(pam) or pam=="TTYN" or pam=="NNG" or pam=="TNN":
- return True
- else:
- return False
- def pamIsXCas9(pam):
- " "
- return (pam in ["NGK", "NGN"])
- def pamIsSpCas9(pam):
- " only used for notes and efficiency scores, unlike its Cpf1 cousin function "
- return (pam in ["NGG", "NGA", "NGCG"])
- def saveSeqOrgPamToCookies(seq, org, pam):
- " create a cookie with seq, org and pam and print it"
- cookies=http.cookies.SimpleCookie()
- expires = 365 * 24 * 60 * 60
- if len(seq)<3000:
- cookies['lastseq'] = seq
- else:
- cookies['lastseq'] = "(last sequence was too long, could not be saved in Internet Browser cookie)"
- cookies['lastseq']['expires'] = expires
- cookies['lastorg'] = org
- cookies['lastorg']['expires'] = expires
- cookies['lastpam'] = pam
- cookies['lastpam']['expires'] = expires
- print(cookies)
- def debug(msg):
- if commandLineMode:
- logging.debug(msg)
- elif DEBUG:
- print(msg)
- print("<br>")
- def gcContent(seq):
- " return GC content as a float "
- c = 0
- for x in seq:
- if x in ["G","C"]:
- c+= 1
- return (float(c)/len(seq))
- def findPat(seq, pat):
- """ yield positions where pat matches seq, stupid brute force search
- """
- seq = seq.upper()
- pat = pat.upper()
- patLen = len(pat)
- for i in range(0, len(seq)-patLen+1):
- subseq = seq[i:i+patLen]
- if patMatch(subseq, pat):
- yield i
- def rndSeq(seqLen):
- " return random seq "
- seq = []
- alf = "ACTG"
- for i in range(0, seqLen):
- seq.append(alf[random.randint(0,3)])
- return "".join(seq)
- def cleanSeq(seq, db):
- """ remove fasta header, check seq for illegal chars and return (filtered
- seq, user message) special value "random" returns a random sequence.
- """
- #print repr(seq)
- if seq.startswith("random"):
- seq = rndSeq(800)
- lines = seq.strip().splitlines()
- #print "<br>"
- #print "before fasta cleaning", "|".join(lines)
- if len(lines)>0 and lines[0].startswith(">"):
- line1 = lines.pop(0)
- #print "<br>"
- #print "after fasta cleaning", "|".join(lines)
- #print "<br>"
- newSeq = []
- nCount = 0
- for l in lines:
- if len(l)==0:
- continue
- for c in l:
- if c not in "actgACTGNn":
- nCount +=1
- else:
- newSeq.append(c)
- seq = "".join(newSeq)
- msgs = []
- tooLongHint = """
- Please split your input sequence into shorter sequences or use
- the <a href='downloads/'>stand-alone version</a> on your own Linux or Mac server to process longer sequences in batch.<br>
- """
- if len(seq)>MAXSEQLEN and db!="noGenome":
- errMsg = "<strong>Sorry, this tool cannot handle sequences longer than %d bp</strong><br>" % (MAXSEQLEN)
- errAbort(errMsg+tooLongHint)
- if len(seq)>MAXSEQLEN_NOGENOME and db=="noGenome":
- errMsg = "<strong>Sorry, this tool cannot handle sequences longer than %d bp when using the 'No Genome' option.</strong><br>" % (MAXSEQLEN_NOGENOME)
- errAbort(errMsg+tooLongHint)
- if nCount!=0:
- msgs.append("Sequence contained %d non-ACTGN letters. They were removed." % nCount)
- return seq, "<br>".join(msgs)
- revTbl = {'A' : 'T', 'C' : 'G', 'G' : 'C', 'T' : 'A', 'N' : 'N' , 'M' : 'K', 'K' : 'M',
- "R" : "Y" , "Y":"R" , "g":"c", "a":"t", "c":"g","t":"a", "n":"n", "V" : "B", "v":"b",
- "B" : "V", "b": "v", "W" : "W", "w" : "w"}
- def revComp(seq):
- " rev-comp a dna sequence with UIPAC characters "
- newSeq = []
- for c in reversed(seq):
- newSeq.append(revTbl[c])
- return "".join(newSeq)
- def docTestInit(isCpf1, guideLen):
- global pamIsFirst
- global GUIDELEN
- pamIsFirst=isCpf1
- GUIDELEN=guideLen
- def findPams (seq, pam, strand, startDict, endSet):
- """ return two values: dict with pos -> strand of PAM and set of end positions of PAMs
- Makes sure to return only values with at least GUIDELEN bp left (if strand "+") or to the
- right of the match (if strand "-")
- If the PAM is cpf1, then this is inversed: pos-strand matches must have at least GUIDELEN
- basepairs to the right, neg-strand matches must have at least GUIDELEN bp on their left
- >>> docTestInit(False, 20)
- >>> findPams("GGGGGGGGGGGGGGGGGGGGGGG", "NGG", "+", {}, set())
- ({20: '+'}, {23})
- >>> findPams("CCAGCCCCCCCCCCCCCCCCCCC", "CCA", "-", {}, set())
- ({0: '-'}, {3})
- >>> docTestInit(True, 20)
- >>> findPams("TTTNCCCCCCCCCCCCCCCCCTTTN", "TTTN", "+", {}, set())
- ({0: '+'}, {4})
- >>> docTestInit(False, 20)
- >>> findPams("CCCCCCCCCCCCCCCCCCCCCAAAA", "NAA", "-", {}, set())
- ({}, set())
- >>> findPams("AAACCCCCCCCCCCCCCCCCCCCC", "NAA", "-", {}, set())
- ({0: '-'}, {3})
- >>> findPams("CCCCCCCCCCCCCCCCCCCCCCCCCAA", "NAA", "-", {}, set())
- ({}, set())
- >>> findPams("GTTGTGTTTTACAATGCAGAGAGTGGAGGATGCTTTTTATACATTGGTGAGAGAGATCCGACAGTACAGATTGAAAAAAATCAGCAAAGAAGAAAAGACTCCTGGCTGTGTGAAAATTAAAAAATGCGTTATAATGTAATCTGGTAAGTTGAGCATATTCATTCTGGTACAAAGCAGATGTCTTCAGAGGTAACA", "TATV", "-", {}, set())
- ({37: '-', 129: '-'}, {41, 133})
- >>> findPams("GTTGTGTTTTACAATGCAGAGAGTGGAGGATGCTTTTTATACATTGGTGAGAGAGATCCGACAGTACAGATTGAAAAAAATCAGCAAAGAAGAAAAGACTCCTGGCTGTGTGAAAATTAAAAAATGCGTTATAATGTAATCTGGTAAGTTGAGCATATTCATTCTGGTACAAAGCAGATGTCTTCAGAGGTAACA", "TATV", "+", {}, set())
- ({37: '+', 129: '+'}, {41, 133})
- """
- assert(pamIsFirst is not None)
- if pamIsFirst:
- maxPosPlus = len(seq)-(GUIDELEN+len(pam))
- minPosMinus = GUIDELEN
- else:
- # -------------------
- # OKOKOKOKOK
- minPosPlus = GUIDELEN
- # -------------------
- # OKOKOKOKOK
- maxPosMinus = len(seq)-(GUIDELEN+len(pam))
- #print "new search", seq, pam, "minPosPlus=",minPosPlus, "guideLen=", GUIDELEN, "<br>"
- for start in findPat(seq, pam):
- if pamIsFirst:
- # need enough flanking seq on one side
- #return("Cpf1 mode found", start,"<br>")
- if strand == "+" and start > maxPosPlus:
- continue
- if strand == "-" and start < minPosMinus:
- continue
- else:
- # return("non-Cpf1 mode found", start,"<br>")
- if strand=="+" and start < minPosPlus:
- continue
- if strand=="-" and start > maxPosMinus:
- continue
- #print "match", strand, start, end, "<br>"
- startDict[start] = strand
- end = start+len(pam)
- endSet.add(end)
- return startDict, endSet
- def rulerString(maxLen):
- " return line with positions every 10 chars "
- texts = []
- for i in range(0, maxLen, 10):
- numStr = str(i)
- texts.append(numStr)
- spacer = "".join([" "]*(10-len(numStr)))
- texts.append(spacer)
- return "".join(texts)
- def varDictToHtml(varDict, seq, varShortLabel):
- " make a list of one html string per position in the sequence "
- if varDict is None:
- return None
- varHtmls = []
- for i in range(0, len(seq)):
- if not i in varDict:
- varHtmls.append(".")
- else:
- varHooverLines = []
- showStar = False # show a star if change is non-simple SNP
- varInfos = varDict[i]
- for chrom, pos, refAll, altAll, infoDict in varInfos:
- varHooverLines.append("%s: %s → %s<br>" % (varShortLabel, refAll, altAll))
- if "freq" in infoDict:
- varHooverLines.append(" <b>Freq:</b> %s<br>" % infoDict["freq"])
- #if "dbg" in infoDict:
- #varHooverLines.append("%s<br>" % infoDict["dbg"])
- if "varId" in infoDict:
- varHooverLines.append(" <b>ID:</b> %s<br>" % infoDict["varId"])
- if "ExAC" in varDict["label"]:
- endPos = int(pos)+len(refAll)
- varHooverLines.append(' <a target=_blank href="http://exac.broadinstitute.org/region/%s-%s-%d">ExAC Browser</a><br>' % (chrom, pos, endPos))
- if len(refAll)!=1 or len(altAll)!=1:
- showStar = True
- if len(varInfos)!=1:
- showStar = True
- varDesc = "".join(varHooverLines)
- if showStar:
- dispChar = "*"
- else:
- dispChar = altAll
- varHtmls.append("<u class='tooltipsterInteract' title='%s'>%s</u>" % (varDesc, dispChar))
- return varHtmls
- def cssClassesFromSeq(guideSeq, suffix=""):
- " The CSS class of guide row and links in seq viewer depend on the first nucl of guide "
- classNames = ["guideRow"]
- if guideSeq[0].upper()!="G":
- classNames.append("guideRowNoPrefixG"+suffix)
- if not guideSeq.startswith("GG"):
- classNames.append("guideRowNoPrefixGG"+suffix)
- if guideSeq[0].upper()!="A":
- classNames.append("guideRowNoPrefixA"+suffix)
- classStr = " ".join(classNames)
- return classStr
- def buildCodonTable():
- " from http://www.petercollingridge.co.uk/tutorials/bioinformatics/codon-table/ "
- bases = "TCAG"
- codons = [a + b + c for a in bases for b in bases for c in bases]
- amino_acids = 'FFLLSSSSYY**CC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG'
- codon_table = dict(list(zip(codons, amino_acids)))
- return codon_table
- def buildOneToThree():
- " return one-letter -> three-letter conversion table for amino acids "
- oneToThree = \
- {'C':'Cys', 'D':'Asp', 'S':'Ser', 'Q':'Gln', 'K':'Lys',
- 'I':'Ile', 'P':'Pro', 'T':'Thr', 'F':'Phe', 'N':'Asn',
- 'G':'Gly', 'H':'His', 'L':'Leu', 'R':'Arg', 'W':'Trp',
- 'A':'Ala', 'V':'Val', 'E':'Glu', 'Y':'Tyr', 'M':'Met',
- 'U':'Sec', '*':'Stop',
- 'X':'Stop', # is this really used like that?
- 'Z':'Glx', # special case: asparagine or aspartic acid
- 'B':'Asx' # special case: glutamine or glutamic acid
- }
- return oneToThree
- def makeExonLines(exonInfo, seq, selTransId):
- """ create text that draws exons, input is transId -> (exonNumber, exStart, exEnd, exFrame).
- returns a list of (transId (=label), symbol (=mouseover), ASCII-line) """
- lines = []
- #maxLabelLen = 0
- codonTable = buildCodonTable()
- #oneToThree = buildOneToThree()
- seqLen = len(seq)
- seq = seq.upper()
- for (transId, symbol), exRows in exonInfo.items():
- if selTransId!="allTrans" and transId!=selTransId:
- continue
- line = [" "]*seqLen
- mouseOvers = {} # position -> mouseOver-text or null, for end of mouse over
- for exIdx, (exNum, exStart, exEnd, exFrame, nextFrame, exStrand) in enumerate(exRows):
- if exFrame==-1:
- for i in range(exStart, exEnd):
- line[i]="="
- exonLabel = "noncoding"
- if (exEnd-exStart)>len(exonLabel)+4:
- # center the exon label on the exon
- mid = exStart+int((exEnd-exStart)*0.5)
- halfLen = int(len(exonLabel)*0.5)
- labStart = mid-halfLen
- for i in range(0, len(exonLabel)):
- line[labStart+i] = exonLabel[i]
- line[labStart-1] = " "
- line[labStart+len(exonLabel)] = " "
- else:
- exonDesc = "gene %s<br>transcript %s<br>exon %d<br>start phase %s" % (symbol, transId, exNum+1, exFrame)
- if nextFrame is not None:
- exonDesc += "<br>end phase %s" % nextFrame
- if (exFrame+nextFrame) % 3 == 0:
- exonDesc += "<br>Removing the exon retains the reading frame"
- else:
- exonDesc += "<br>Removing the exon will destroy the reading frame"
- mouseOvers[exStart] = exonDesc
- mouseOvers[exEnd] = None
- for i in range(exStart, exStart+exFrame):
- line[i] = "-"
- for i in range(exStart+exFrame, exEnd, 3):
- codon = seq[i:i+3]
- if len(codon)==3:
- shortAa = codonTable[codon]
- if exStrand=="+":
- longAa = shortAa+"]]"
- else:
- longAa = "[["+shortAa # highlighting rev. dir. more
- else:
- # codon is split by splice site
- longAa = "-"
- for j in range(0, len(longAa)):
- line[i+j] = longAa[j]
- # now merge the mouse overs as span tags into the ASCII line
- newLine = []
- for pos, char in enumerate(line):
- if pos in mouseOvers:
- overString = mouseOvers[pos]
- if overString is not None:
- newLine.append("<span class='tooltipsterInteract' title='%s'>" % overString)
- if char=="-":
- newLine.append("<span class='tooltipsterInteract' title='This codon goes over a splice site. The nucleotides of the split codon are not translated to amino acids but shown as dashes.'>-</span>")
- else:
- newLine.append(char)
- if overString is None:
- newLine.append("</span>")
- else:
- newLine.append(char)
- # and fix up the < signs
- newLineStr = "".join(newLine)
- newLineStr = newLineStr.replace("[[", "<lc><<</lc>")
- newLineStr = newLineStr.replace("]]", "<lc>>></lc>")
- lines.append((symbol, transId, newLineStr))
- #maxLabelLen = max(maxLabelLen, len(transId))
- return lines
- def getGeneModels(org):
- " read possible gene models for org and return as list (name, desc) or None if no gene models "
- mask = join(genomesDir, org, "*.bb")
- fnames = glob.glob(mask)
- descFname = join(genomesDir, org, "genes.tsv")
- if not isfile(descFname):
- return None
- geneDescs = {}
- for line in open(descFname):
- fname, desc = line.split(maxsplit=1)
- geneDescs[fname] = desc
- ret = []
- for fname in fnames:
- baseName = basename(fname)
- name = baseName.split('.')[0]
- desc = geneDescs.get(baseName, name)
- ret.append((name, desc))
- return ret
- def getSelGeneModel(org):
- " return (list of (name, desc) of models, selected gene model name) "
- geneModels = getGeneModels(org)
- selGeneModel = None
- selTransId = None
- if geneModels:
- #selGeneModel = cgiParams.get("geneModel", geneModels[0][0])
- selGeneModel = cgiParams.get("geneModel", "noGenes")
- geneModels.insert(0, ("noGenes", "Do not show"))
- possNames = [x for x,y in geneModels]
- if not selGeneModel in possNames:
- errAbort("The gene model name specified with the argument geneModel is invalid")
- selTransId = cgiParams.get("transId", "allTrans")
- return geneModels, selGeneModel, selTransId
- def printSeqForCopy(seq):
- " print a hidden text area so we can copy the sequence to the clipboard "
- print('<input id="seqAsText" type="text" style="display:none">')
- print(seq)
- print("</input>")
- def calcKomorScore(guideSeq, pos):
- " return base editing score given the guide sequence and the position "
- return pos/7.0 # temporary hack
- def makeEditLines(seq, pamSeqs, winStart, winEnd, guideScores):
- " create the lines that show the possible baseEditor edits "
- editInfos = []
- for i in range(0, len(seq)):
- editInfos.append(defaultdict(list))
- upSeq = seq.upper()
- for pamId, pamStart, guideStart, strand, guideSeq, pamSeq, pamPlusSeq in pamSeqs:
- specScore = guideScores[pamId]
- if strand=="+":
- fromPos = guideStart+winStart
- toPos = guideStart+winEnd
- fromNucl = "C"
- toNucl = "T"
- else:
- guideEnd = guideStart+GUIDELEN
- fromPos = guideEnd-winEnd
- toPos = guideEnd-winStart
- fromNucl = "G"
- toNucl = "A"
- for pos in range(fromPos, toPos):
- # position of mutated nucl on forw strand guide
- if strand=="+":
- mutPos = pos-guideStart
- else:
- mutPos = GUIDELEN - (pos - guideStart) - 1
- if upSeq[pos]==fromNucl:
- beScore = calcKomorScore(guideSeq, mutPos)
- editInfos[pos][toNucl].append((pamId, guideSeq, pamSeq, mutPos, beScore, specScore))
- altNucls = ["A", "T"]
- editLabels = []
- for an in altNucls:
- editLabels.append("Edits to "+an)
- editLines = []
- for i in range(0, len(altNucls)):
- editLines.append([" "]*len(seq))
- # rearrange into lines of text + JSON
- jsonData = defaultdict(list)
- for pos, eiDict in enumerate(editInfos):
- if not eiDict:
- continue
- jsonData[pos] = eiDict
- for nucl, guideData in eiDict.items():
- yPos = altNucls.index(nucl)
- editLines[yPos][pos] = "<d pos=%d>%s</d>" % (pos, nucl)
- ret = []
- for label, lineChars in zip(editLabels, editLines):
- ret.append( (label, None, "".join(lineChars)) )
- return ret, jsonData
- def makePamLines(lines, maxY, pamIdToSeq, guideScores):
- for y in range(0, maxY+1):
- texts = []
- lastEnd = 0
- for start, end, name, strand, pamId in lines[y]:
- guideSeq = pamIdToSeq.get(pamId)
- if guideSeq==None:
- # when there is an N in the guide, the PAM is valid, but the guide is not
- continue
- classStr = cssClassesFromSeq(guideSeq, suffix="Seq")
- spacer = "".join([" "]*((start-lastEnd)))
- lastEnd = end
- texts.append(spacer)
- score = guideScores[pamId]
- # XX How can this happen for non-Cpf1 enzymes? Can this ever happen?
- if score is None and not pamIsFirst:
- continue
- color = scoreToColor(score)[0]
- texts.append('''<a class='%s' style="text-shadow: 1px 1px 1px #bbb; color: %s" id="list%s" href="#%s">''' % (classStr, color, pamId,pamId))
- texts.append(name)
- texts.append("</a>")
- yield ("", None, ''.join(texts))
- def getBeWin(winVal):
- " return (start, end) of base editor window given CGI variable "
- fs = winVal.split("-")
- if len(fs)!=2:
- errAbort("parameter beWin must contain only one dash")
- start = fs[0].strip()
- end = fs[1].strip()
- if not start.isdigit() or not end.isdigit():
- errAbort("parameter beWin must be two dash-separated numbers")
- start = int(start)
- end = int(end)
- return start, end
- def printLines(lines, labelLen):
- " print list of (label, string) such that label is at least labelLen characters long "
- for label, mouseOver, line in lines:
- if mouseOver is not None:
- labelStr =('<span class="tooltipsterInteract" title="{:s}">{:'+str(labelLen)+'s} </span>').format(label, mouseOver)
- else:
- labelStr = ('{:'+str(labelLen)+'s} ').format(label)
- print((labelStr), end=' ')
- print(line)
- def getMaxLen(lines):
- " given a list of tuples where first element is the label, return the longest label len "
- maxLen = 0
- for l in lines:
- label = l[0]
- maxLen = max(maxLen, len(label))
- return maxLen
- def printJson(name, obj):
- print("<script>")
- print((name), end=' ')
- print(("="), end=' ')
- print((json.dumps(obj)))
- print("</script>")
- def showSeqAndPams(org, seq, startDict, pam, guideScores, varHtmls, varDbs, varDb, minFreq, position, pamIdToSeq):
- " show the sequence and the PAM sites underneath in a sequence viewer "
- pamSeqs = list(flankSeqIter(seq, startDict, len(pam), True))
- lines, maxY = distrOnLines(seq.upper(), startDict, len(pam), pam)
- posLabel = "Position"
- varLabel = "Variants"
- seqLabel = "Sequence"
- exonLabelLen = 0
- editLines = []
- exonLines = []
- geneModels, selGeneModel, selTransId = getSelGeneModel(org)
- #selGeneModel = None
- #geneModels = None
- if baseEditor:
- beWinStart, beWinEnd = getBeWin(cgiParams.get("beWin", DEFAULTBEWIN))
- editLines, jsonData = makeEditLines(seq, pamSeqs, beWinStart, beWinEnd, guideScores)
- printJson("editData", jsonData)
- pamLines = list(makePamLines(lines, maxY, pamIdToSeq, guideScores))
- labelLen = max(len(varLabel), len(seqLabel), len(posLabel), getMaxLen(pamLines))
- if selGeneModel!=None:
- exonInfo, maxTransIdLen = getExonInfo(org, selGeneModel, position)
- labelLen = max(labelLen, maxTransIdLen)
- if baseEditor:
- labelLen = max(labelLen, getMaxLen(editLines))
- if selGeneModel:
- labelLen = max(labelLen, exonLabelLen)
- print("<div class='substep'>")
- print('<a id="seqStart"></a>')
- print("Your input sequence is %d bp long. It contains %d possible guide sequences.<br>" % (len(seq), len(guideScores)))
- if not pamIsFirst:
- print("Shown below are their PAM sites and the expected cleavage position located -3bp 5' of the PAM site.<br>")
- print("Click on a match for the PAM %s below to show its %d bp-long guide sequence. " % (pam, GUIDELEN))
- print("(Need help? Look at the <a target=_blank href='manual/#annotseq'>CRISPOR manual</a>)<br>")
- print('''Colors <span style="color:#32cd32; text-shadow: 1px 1px 1px #bbb">green</span>, <span style="color:#ffff00; text-shadow: 1px 1px 1px #888">yellow</span> and <span style="text-shadow: 1px 1px 1px #f01; color:#aa0014">red</span> indicate high, medium and low specificity of the PAM's guide sequence in the genome.<p>''')
- else:
- print("Click on a match for the PAM %s below to show its %d bp-long guide sequence.<br>" % (pam, GUIDELEN))
- if baseEditor or varDb or selGeneModel:
- print(("""<form style="display:inline" id="paramForm" action="%s" method="GET">""" % basename(__file__)))
- if geneModels:
- print ("Gene Models:")
- printDropDown("geneModel", geneModels, selGeneModel, style="width:20em")
- if selGeneModel!="noGenes":
- print ("Transcript:")
- transIdInfo = [("allTrans", "All Transcripts")]
- for transId, sym in list(exonInfo.keys()):
- transIdInfo.append( (transId, sym+" / "+transId) )
- printDropDown("transId", transIdInfo, selTransId, style="width:20em")
- # XX XXXXXX
- exonLines = makeExonLines(exonInfo, seq, selTransId)
- #exonLines = []
- print("""<input style="height:18px;margin:0px;font-size:10px;line-height:normal" type="submit" name="submit" value="Update">""")
- print("""<br>""")
- if baseEditor:
- print ("Base Editor modification window:")
- print(("""<input type="text" name="beWin" size="10" value="%s">""" % DEFAULTBEWIN))
- print("""<input style="height:18px;margin:0px;font-size:10px;line-height:normal" type="submit" name="submit" value="Update">""")
- print("<br>")
- if varDb is not None:
- print ("Variant database:")
- varDbList = [(b,c) for a,b,c,d in varDbs] # only keep fname+label
- printDropDown("varDb", varDbList, varDb)
- if minFreq==0.0:
- minFreq="0.0"
- else:
- minFreq = str(minFreq)
- # pull out the hasAF field for this varDb
- varDbHasAF = False
- for shortLabel, fname, desc, hasAF in varDbs:
- if fname==varDb:
- varDbHasAF = hasAF
- break
- if varDbHasAF:
- print(""" Min. frequency: """)
- print(("""<input type="text" name="minFreq" size="8" value="%s">""" % minFreq))
- print("""<input style="height:18px;margin:0px;font-size:10px;line-height:normal" type="submit" name="submit" value="Update">""")
- print(("<small style='margin-left:30px'><a href='mailto:%s'>Missing a variant database? We can add it.</a></small>" % contactEmail))
- if position=="?":
- print("<small style=''>Input sequence not in genome, cannot show genome variants.</small>")
- elif varDb is None:
- print(("<small style=''><a href='mailto:%s'>Suggest a genome variants database to show on this page</a></small>" % contactEmail))
- print("</div>")
- print('''<div class="blueHighlight" style="text-align: left; overflow-x:scroll; width:98vw; background:#DDDDDD; border-style: solid; border-width: 1px">''')
- print('''<pre style="font-family: Source Code Pro; font-size: 80%; display:inline; line-height: 0.95em; text-align:left">''')
- print(('{:'+str(labelLen)+'s} ').format(posLabel), end=' ')
- print(rulerString(len(seq)))
- if varHtmls is not None:
- print(('{:'+str(labelLen)+'s} ').format(varLabel), end=' ')
- print("".join(varHtmls))
- print(('{:'+str(labelLen)+'s} ').format(seqLabel), end=' ')
- print (seq)
- printLines(exonLines, labelLen)
- if baseEditor:
- printLines(editLines, labelLen)
- printLines(pamLines, labelLen)
- print("</pre><br>")
- print('''</div>''')
- #printSeqForCopy(seq)
- if pamIsCas12max(pam):
- print('<div style="line-height: 1.0; padding-top: 5px; font-size: 15px">Cpf1 has a staggered site: cleavage occurs between the 14th and 16th base on the non-targeted strand (indicated by "\\" in the schema above). Cleavage mostly occurs after the 24rd base on the targeted strand (indicated by "/" in the schema above). See on <a target=_blank href="https://www.synthego.com/products/nuclease/hfcas12max-hifi">Synthego</a></div>')
- elif pamIsCpf1(pam):
- print('<div style="line-height: 1.0; padding-top: 5px; font-size: 15px">Cpf1 has a staggered site: cleavage occurs usually - but not always - after the 18th base on the non-targeted strand which has the TTTV PAM motif (indicated by "\\" in the schema above). Cleavage mostly occurs after the 23rd base on the targeted strand which has the AAAN motif (indicated by "/" in the schema above). See <a target=_blank href="http://www.sciencedirect.com/science/article/pii/S0092867415012003">Zetsche et al 2015</a>, in particular <a target=_blank href="http://www.sciencedirect.com/science?_ob=MiamiCaptionURL&_method=retrieve&_eid=1-s2.0-S0092867415012003&_image=1-s2.0-S0092867415012003-gr3.jpg&_cid=272196&_explode=defaultEXP_LIST&_idxType=defaultREF_WORK_INDEX_TYPE&_alpha=defaultALPHA&_ba=&_rdoc=1&_fmt=FULL&_issn=00928674&_pii=S0092867415012003&md5=11771263f3e390e444320cacbcfae323">Fig 3</a>.</div>')
- elif pamIsCasX(pam):
- print('<div style="line-height: 1.0; padding-top: 5px; font-size: 15px">We have no description yet on how exactly the CasX cleavage looks like. Please contact [email hidden] if you have an idea how to describe the cleavage site.</div>')
- def iterOneDelSeqs(seq):
- """ given a seq, create versions with each bp removed. Avoid duplicates
- yields (delPos, seq)
- >>> list(iterOneDelSeqs("AATGG"))
- [(0, 'ATGG'), (2, 'AAGG'), (3, 'AATG')]
- """
- doneSeqs = set()
- for i in range(0, len(seq)):
- delSeq = seq[:i]+seq[i+1:]
- if delSeq not in doneSeqs:
- yield i, delSeq
- doneSeqs.add(delSeq)
- def flankSeqIter(seq, startDict, pamLen, doFilterNs):
- """ given a seq and dictionary of pamPos -> strand and the length of the pamSite
- yield tuples of (name, pamStart, guideStart, strand, flankSeq, pamSeq)
- flankSeq is the guide sequence (=flanking the PAM).
- if doFilterNs is set, will not return any sequences that contain an N character
- pamPlusSeq are the 5bp after the PAM. If not enough space, pamPlusSeq is None
- """
- startList = sorted(startDict.keys())
- for pamStart in startList:
- strand = startDict[pamStart]
- pamPlusSeq = None
- if pamIsFirst: # Cpf1: get the sequence to the right of the PAM
- if strand=="+":
- guideStart = pamStart+pamLen
- flankSeq = seq[guideStart:guideStart+GUIDELEN]
- pamSeq = seq[pamStart:pamStart+pamLen]
- if pamStart-pamPlusLen >= 0:
- pamPlusSeq = seq[pamStart-pamPlusLen:pamStart]
- else: # strand is minus
- guideStart = pamStart-GUIDELEN
- flankSeq = revComp(seq[guideStart:pamStart])
- pamSeq = revComp(seq[pamStart:pamStart+pamLen])
- if pamStart+pamLen+pamPlusLen < len(seq):
- pamPlusSeq = revComp(seq[pamStart+pamLen:pamStart+pamLen+pamPlusLen])
- else: # common case: get the sequence on the left side of the PAM
- if strand=="+":
- guideStart = pamStart-GUIDELEN
- flankSeq = seq[guideStart:pamStart]
- pamSeq = seq[pamStart:pamStart+pamLen]
- if pamStart+pamLen+pamPlusLen < len(seq):
- pamPlusSeq = seq[pamStart+pamLen:pamStart+pamLen+pamPlusLen]
- else: # strand is minus
- guideStart = pamStart+pamLen
- flankSeq = revComp(seq[guideStart:guideStart+GUIDELEN])
- pamSeq = revComp(seq[pamStart:pamStart+pamLen])
- if pamStart-pamPlusLen >= 0:
- pamPlusSeq = revComp(seq[pamStart-pamPlusLen:pamStart])
- if "N" in flankSeq and doFilterNs:
- continue
- yield "s%d%s" % (pamStart, strand), pamStart, guideStart, strand, flankSeq, pamSeq, pamPlusSeq
- def makeBrowserLink(dbInfo, pos, text, title, cssClasses, ctUrl=None):
- " return link to genome browser (ucsc or ensembl) at pos, with given text "
- if dbInfo is None:
- errAbort("Your batchID relates to a genome that is not present anymore. You will have to change the version of the site. Or contact us and send us the full URL of this page.")
- if dbInfo.server.startswith("Ensembl"):
- baseUrl = "www.ensembl.org"
- urlLabel = "Ensembl"
- # link back to archive, if possible
- if dbInfo.description.startswith("Ensembl "):
- ensVersion = dbInfo.description.split()[1]
- if ensVersion.isdigit():
- baseUrl = "e%s.ensembl.org" % ensVersion
- elif dbInfo.server=="EnsemblPlants":
- baseUrl = "plants.ensembl.org"
- elif dbInfo.server=="EnsemblMetazoa":
- baseUrl = "metazoa.ensembl.org"
- elif dbInfo.server=="EnsemblProtists":
- baseUrl = "protists.ensembl.org"
- org = dbInfo.scientificName.replace(" ", "_")
- pos = pos.replace(":+","").replace(":-","") # remove the strand
- url = "http://%s/%s/Location/View?r=%s" % (baseUrl, org, pos)
- elif dbInfo.server=="ucsc" or dbInfo.name.startswith("GCA_") or dbInfo.name.startswith("GCF_"):
- urlLabel = "UCSC"
- if len(pos)>0 and pos[0].isdigit():
- pos = "chr"+pos
- # remove the strand
- pos = pos.replace(":+","").replace(":-","")
- url = "http://genome.ucsc.edu/cgi-bin/hgTracks?db=%s&position=%s" % (dbInfo.name, pos)
- if ctUrl is not None:
- url+= "&hgt.customText=%s" % ctUrl
- # some limited support for gbrowse
- elif dbInfo.server.startswith("http://"):
- urlLabel = "GBrowse"
- chrom, start, end, strand = parsePos(pos)
- start = start+1
- url = "%s/?name=%s:%d..%d" % (dbInfo.server, chrom, start, end)
- else:
- chrom, start, end, strand = parsePos(pos)
- if chrom is not None and chrom.startswith("NC_"):
- start = start+1
- url = "https://www.ncbi.nlm.nih.gov/nuccore/%s?report=graph&log$=seqview&v=%d-%d" % \
- (chrom, start, end)
- urlLabel = "NCBI "
- else:
- #return "unknown genome browser server %s, please email [email hidden]" % dbInfo.server
- urlLabel = None
- url = "javascript:void(0)"
- classStr = ""
- if len(cssClasses)!=0:
- classStr = ' class="%s"' % (" ".join(cssClasses))
- if title is None:
- if urlLabel != None:
- title = "Link to %s Genome Browser" % urlLabel
- else:
- title = "No Genome Browser link available yet for this organism"
- return '''<a title="%s"%s target="_blank" href="%s">%s</a>''' % (title, classStr, url, text)
- def highlightMismatches(guide, offTarget, pamLen):
- " return a string that marks mismatches between guide and offtarget with * "
- if pamLen!=0:
- if pamIsFirst:
- offTarget = offTarget[pamLen:]
- else:
- offTarget = offTarget[:-pamLen]
- assert(len(guide)==len(offTarget))
- s = []
- for x, y in zip(guide, offTarget):
- if x==y:
- s.append(".")
- else:
- s.append("*")
- return "".join(s)
- def parseNewAlias(ifh):
- " part of parseAlias(): IGV-compatible format: first is UCSC, all other columns are aliases "
- toUcsc = {}
- for line in ifh:
- if line.startswith("#"):
- continue
- row = line.rstrip("\n").split("\t")
- for i in range(1, len(row)):
- toUcsc[row[i]] = row[0]
- return toUcsc
- def parseAlias(fname):
- " parse tsv file with at least two columns, orig chrom name and new chrom name. copied from chromToUcsc script from the UCSC tools. "
- logging.debug("alias file is in IGV-format")
- toUcsc = {}
- if fname.startswith("http://") or fname.startswith("https://"):
- ifh = urlopen(fname)
- if fname.endswith(".gz"):
- data = gzip.GzipFile(fileobj=ifh).read().decode()
- ifh = data.splitlines()
- elif fname.endswith(".gz"):
- ifh = gzip.open(fname, "rt")
- else:
- ifh = open(fname)
- firstLine = True
- for line in ifh:
- if line.startswith("#") and firstLine:
- return parseNewAlias(ifh)
- if line.startswith("alias"):
- continue
- row = line.rstrip("\n").split("\t")
- toUcsc[row[0]] = row[1]
- firstLine = False
- return toUcsc
- chromAlias = None
- def applyChromAlias(db, chrom):
- " if chrom is in chromAlias, return the human-readable name "
- global chromAlias
- if chromAlias==-1: # == chromAlias file not present
- return chrom
- elif chromAlias is None:
- chromAliasFname = join("genomes", db, db+".chromAlias.txt")
- if not isfile(chromAliasFname):
- chromAlias = -1
- return chrom
- else:
- chromAlias = parseAlias(chromAliasFname)
- return chromAlias.get(chrom, chrom)
- def makeAlnStr(org, seq1, seq2, pam, mitScore, cfdScore, posStr, chromDist):
- """ given two strings of equal length, return a html-formatted string of several lines
- that show the two sequences and a line that highlights where they differ
- """
- lines = [ [], [], [] ]
- last12MmCount = 0
- inLinkage = False
- hlSeed = False
- if pamIsSpCas9(pam):
- hlSeed = True
- if pamIsFirst:
- lines[0].append("<i>"+seq1[:len(pam)]+"</i> ")
- lines[1].append("<i>"+seq2[:len(pam)]+"</i> ")
- lines[2].append("".join([" "]*(len(pam)+1)))
- if pamIsFirst:
- guideStart = len(pam)
- guideEnd = len(seq1)
- else:
- guideStart = 0
- guideEnd = len(seq1)-len(pam)
- for i in range(guideStart, guideEnd):
- if hlSeed and i==10:
- lines[1].append("<u>")
- if seq1[i]==seq2[i]:
- lines[0].append(seq1[i])
- lines[1].append(seq2[i])
- lines[2].append(" ")
- else:
- lines[0].append("<b>%s</b>" % seq1[i])
- lines[1].append("<b>%s</b>" % seq2[i])
- lines[2].append("*")
- if i>7:
- last12MmCount += 1
- if hlSeed and i==guideEnd-1:
- lines[1].append("</u>")
- if not pamIsFirst:
- lines[0].append(" <i>"+seq1[-len(pam):]+"</i>")
- lines[1].append(" <i>"+seq2[-len(pam):]+"</i>")
- lines = ["".join(l) for l in lines]
- chrom, chromPos, strand = posStr.split(":")
- chrom = applyChromAlias(org, chrom)
- posStr = ":".join((chrom, chromPos, strand))
- if len(posStr)>1 and posStr[0].isdigit():
- posStr = "chr"+posStr
- htmlText1 = "<small><pre>guide: %s<br>off-target: %s<br> %s</pre>" \
- % (lines[0], lines[1], lines[2])
- if pamIsCpf1(pam) or pamIsCasX(pam):
- htmlText2 = "Cpf1/CasX: No off-target scores available</small>"
- elif saCas9Mode:
- htmlText2 = "SaCas9 Tycko Score: %s" % mitScore
- else:
- if cfdScore==None:
- cfdStr = "Cannot calculate CFD score on non-ACTG characters"
- else:
- cfdStr = "%f" % cfdScore
- htmlText2 = "CFD Off-target score: %s<br>MIT Off-target score: %.2f<br>Position: %s</small>" % (cfdStr, mitScore, posStr)
- if chromDist!=None and org!=None:
- htmlText2 += "<br><small>Distance from target: %.3f Mbp</small>" % (float(chromDist)/1000000.0)
- if org.startswith("mm") or org.startswith("hg") or org.startswith("rn"):
- if chromDist > 20000000:
- htmlText2 += "<br><small>>20Mbp = unlikely to be in linkage with target</small>"
- else:
- htmlText2 += "<br><small><20Mbp= likely to be in linkage with "
- "target! Even if no linkage: beware of chromosomal rearrangements "
- "when using this guide!</small>"
- inLinkage = True
- hasLast12Mm = last12MmCount>0
- return htmlText1+htmlText2, hasLast12Mm, inLinkage
- def parsePos(text):
- """ parse a string of format chr:start-end:strand and return a 4-tuple
- Strand defaults to + and end defaults to start+23
- """
- if text!=None and len(text)!=0 and text!="?":
- fields = text.split(":")
- if len(fields)==2:
- chrom, posRange = fields
- strand = "+"
- else:
- chrom, posRange, strand = fields
- posRange = posRange.replace(",","")
- if "-" in posRange:
- start, end = posRange.split("-")
- start, end = int(start), int(end)
- if start > end:
- start, end = end, start
- strand = "-"
- else:
- # if the end position is not specified (as by default done by UCSC outlinks), use start+23
- start = int(posRange)
- end = start+23
- else:
- chrom, start, end, strand = "", 0, 0, "+"
- return chrom, start, end, strand
- def annotateOfftargets(org, countDict, guideSeq, pam, inputPos):
- """ for a given guide sequence, return a list of tuples that
- describes the offtargets sorted by score and a string to describe the offtargets in the
- format x/y/z/w of mismatch counts
- inputPos has format "chrom:start-end:strand". All 0MM matches in this range
- are ignored from scoring ("ontargets")
- Also return the same description for just the last 12 bp and the score
- of the guide sequence (calculated using all offtargets).
- """
- inChrom, inStart, inEnd, inStrand = parsePos(inputPos)
- count = 0
- otCounts = []
- posList = []
- mitOtScores = []
- cfdScores = []
- last12MmCounts = []
- ontargetDesc = ""
- repCount = 0 # if repCount for a guide is !=0, then the guide should not be used. repCount is then the number
- # of matches for the guide in the genome (not looking at the PAM)
- # for each edit distance, get the off targets and iterate over them
- foundOneOntarget = False
- isSaCas9 = pamIsSaCas9(pam)
- isCpf1 = pamIsCpf1(pam)
- for editDist in range(0, maxMMs+1):
- #print countDict,"<p>"
- matches = countDict.get(editDist, [])
- #print otCounts,"<p>"
- last12MmOtCount = 0
- # create html and score for every offtarget
- otCount = 0
- for chrom, start, end, otSeq, strand, segType, geneNameStr, totalAlnCount, isRep in matches:
- # if repCount is > 0, then this means that the guide should not be used and we cannot
- # even get any off-targets
- if (totalAlnCount > MAXOCC) or (totalAlnCount > 1 and isRep):
- repCount = totalAlnCount # any off-target with this condition will trigger the whole guide to be suppressed
- # skip on-targets
- if segType!="":
- segTypeDesc = segTypeConv[segType]
- geneDesc = segTypeDesc+":"+geneNameStr
- geneDesc = geneDesc.replace("|", "-")
- else:
- geneDesc = geneNameStr
- # is this not an off-target but the on-target?
- # if we got a genome position, use it. Otherwise use a random off-target with 0MMs
- # as the on-target ("auto-ontarget" mode)
- if editDist==0 and \
- repCount==0 and \
- ((chrom==inChrom and start >= inStart and end <= inEnd) \
- or (inChrom=='' and foundOneOntarget==False)):
- foundOneOntarget = True
- ontargetDesc = geneDesc
- continue
- otCount += 1
- guideNoPam = guideSeq[:len(guideSeq)-len(pam)]
- otSeqNoPam = otSeq[:len(otSeq)-len(pam)]
- if len(otSeqNoPam)==19:
- otSeqNoPam = "A"+otSeqNoPam # should not change the score a lot, weight0 is very low
- guideNoPam = "A"+guideNoPam
- if isCpf1:
- # Cpf1 has no off-target scores yet
- mitScore=0.0
- cfdScore=0.0
- elif isSaCas9:
- mitScore = calcSaHitScore(guideNoPam, otSeqNoPam)
- cfdScore = -1
- else:
- # MIT score must not include the PAM
- mitScore = calcHitScore(guideNoPam, otSeqNoPam)
- # this is a heuristic based on the guideSeq data where alternative
- # PAMs represent only ~10% of all cleaveage events.
- # We divide the MIT score by 5 to make sure that these off-targets
- # are not ranked among the top but still appear in the list somewhat
- if pam=="NGG" and otSeq[-2:]!="GG":
- mitScore = mitScore * 0.2
- # CFD score must include the PAM
- cfdScore = calcCfdScore(guideSeq, otSeq)
- mitOtScores.append(mitScore)
- if cfdScore != -1:
- cfdScores.append(cfdScore)
- posStr = "%s:%d-%s:%s" % (chrom, start+1,end, strand)
- if (chrom==inChrom):
- dist = abs(start-inStart)
- else:
- dist = None
- parNum = isInPar(org, chrom, start, end)
- if parNum is not None:
- posStr += " PAR%s" % parNum
- alnHtml, hasLast12Mm, inLinkage = makeAlnStr(org, guideSeq, otSeq,
- pam, mitScore, cfdScore, posStr, dist)
- if not hasLast12Mm:
- last12MmOtCount+=1
- posList.append( (otSeq, mitScore, cfdScore, editDist, posStr, geneDesc,
- alnHtml, inLinkage) )
- last12MmCounts.append(str(last12MmOtCount))
- # create a list of number of offtargets for this edit dist
- otCounts.append( str(otCount) )
- # calculate the guide scores
- if pamIsCpf1(pam):
- guideScore = -1
- guideCfdScore = -1
- else:
- if repCount>0:
- guideScore = 0
- guideCfdScore = 0
- else:
- guideScore = calcMitGuideScore(sum(mitOtScores))
- if doCfdFix:
- guideCfdScore = calcCfdGuideScore(sum(cfdScores))
- else:
- guideCfdScore = calcMitGuideScore(sum(cfdScores))
- # obtain the off-target info: coordinates, descriptions and off-target counts
- if repCount>0:
- posList = []
- ontargetDesc = ""
- last12DescStr = ""
- otDescStr = ""
- else:
- otDescStr = " - ".join(otCounts)
- last12DescStr = " - ".join(last12MmCounts)
- if pamIsCpf1(pam):
- # sort by edit dist if using Cfp1
- posList.sort(key=operator.itemgetter(3))
- else:
- # sort by CFD score if we have it
- posList.sort(reverse=True, key=operator.itemgetter(2))
- return posList, otDescStr, guideScore, guideCfdScore, last12DescStr, \
- ontargetDesc, repCount
- # --- START OF SCORING ROUTINES
- saGuide = None
- saScorer = None
- def calcSaHitScore(guideSeq, otSeq):
- """
- saCas9 offtarget scoring from Tycko et al, https://www.nature.com/articles/s41467-018-05391-2
- see bin/src/pairwise-library-screen/
- """
- global saScorer
- global saGuide
- if guideSeq!=saGuide:
- sys.path.append("bin/src/pairwise-library-screen")
- import predictSingle
- saGuide = guideSeq
- saScorer = predictSingle.SaCas9Scorer(len(guideSeq))
- # to be compatible with the MIT score, has to be in the range 0-100
- # for the MIT aggregate guide specificity score
- return 100.0*saScorer.calcScore(guideSeq, otSeq)
- # MIT offtarget scoring, "Hsu score"
- # aka Matrix "M"
- hitScoreM = [0,0,0.014,0,0,0.395,0.317,0,0.389,0.079,0.445,0.508,0.613,0.851,0.732,0.828,0.615,0.804,0.685,0.583]
- def calcHitScore(string1,string2):
- " see 'Scores of single hits' on http://crispr.mit.edu/about "
- # The Patrick Hsu weighting scheme
- # S. aureus requires 21bp long guides. We fudge by using only last 20bp
- matrixStart = 0
- maxDist = 19
- assert(string1[0].isupper())
- assert(len(string1)==len(string2))
- #for nmCas9 and a few others with longer guides, we limit ourselves to 20bp
- if len(string1)>20:
- string1 = string1[-20:]
- string2 = string2[-20:]
- # for 19bp guides, we fudge a little, but first pos has no weight anyways
- elif len(string1)==19:
- string1 = "A"+string1
- string2 = "A"+string2
- # for shorter guides, I'm not sure if this score makes sense anymore, we force things
- elif len(string1)<19:
- matrixStart = 20-len(string1)
- maxDist = len(string1)-1
- assert(len(string1)==len(string2))
- dists = [] # distances between mismatches, for part 2
- mmCount = 0 # number of mismatches, for part 3
- lastMmPos = None # position of last mismatch, used to calculate distance
- score1 = 1.0
- for pos in range(matrixStart, len(string1)):
- if string1[pos]!=string2[pos]:
- mmCount+=1
- if lastMmPos!=None:
- dists.append(pos-lastMmPos)
- score1 *= 1-hitScoreM[pos]
- lastMmPos = pos
- # 2nd part of the score
- if mmCount<2: # special case, not shown in the paper
- score2 = 1.0
- else:
- avgDist = sum(dists)/len(dists)
- score2 = 1.0 / (((maxDist-avgDist)/float(maxDist)) * 4 + 1)
- # 3rd part of the score
- if mmCount==0: # special case, not shown in the paper
- score3 = 1.0
- else:
- score3 = 1.0 / (mmCount**2)
- score = score1 * score2 * score3 * 100
- return score
- def calcMitGuideScore(hitSum):
- """ Sguide defined on http://crispr.mit.edu/about
- Input is the sum of all off-target hit scores. Returns the specificity of the guide.
- """
- score = 100 / (100+hitSum)
- score = int(round(score*100))
- return score
- def calcCfdGuideScore(hitSum):
- " suggested by Nicholas Parkinson "
- norm_score = 100.* 100.0 / (hitSum)
- return norm_score
- # === SOURCE CODE cfd-score-calculator.py provided by John Doench =====
- # The CFD score is an improved specificity score
- def get_mm_pam_scores():
- """
- """
- import pickle
- dataDir = join(dirname(__file__), 'CFD_Scoring')
- mm_scores = pickle.load(open(join(dataDir, 'mismatch_score.pkl'),'rb'))
- pam_scores = pickle.load(open(join(dataDir, 'pam_scores.pkl'),'rb'))
- return (mm_scores,pam_scores)
- #Reverse complements a given string
- def revcom(s):
- basecomp = {'A': 'T', 'C': 'G', 'G': 'C', 'T': 'A','U':'A'}
- letters = list(s[::-1])
- letters = [basecomp[base] for base in letters]
- return ''.join(letters)
- #Calculates CFD score
- def calc_cfd(wt,sg,pam):
- #mm_scores,pam_scores = get_mm_pam_scores()
- score = 1
- sg = sg.replace('T','U')
- wt = wt.replace('T','U')
- s_list = list(sg)
- wt_list = list(wt)
- for i,sl in enumerate(s_list):
- if wt_list[i] == sl:
- score*=1
- else:
- key = 'r'+wt_list[i]+':d'+revcom(sl)+','+str(i+1)
- score*= mm_scores[key]
- score*=pam_scores[pam]
- return (score)
- mm_scores, pam_scores = None, None
- def calcCfdScore(guideSeq, otSeq):
- """ based on source code provided by John Doench
- >>> calcCfdScore("GGGGGGGGGGGGGGGGGGGGGGG", "GGGGGGGGGGGGGGGGGAAAGGG")
- 0.4635989007074176
- >>> calcCfdScore("GGGGGGGGGGGGGGGGGGGGGGG", "GGGGGGGGGGGGGGGGGGGGGGG")
- 1.0
- >>> calcCfdScore("GGGGGGGGGGGGGGGGGGGGGGG", "aaaaGaGaGGGGGGGGGGGGGGG")
- 0.5140384614450001
- # mismatches: * !!
- >>> calcCfdScore("ATGGTCGGACTCCCTGCCAGAGG", "ATGGTGGGACTCCCTGCCAGAGG")
- 0.5
- # mismatches: * ** *
- >>> calcCfdScore("ATGGTCGGACTCCCTGCCAGAGG", "ATGATCCAAATCCCTGCCAGAGG")
- 0.53625000020625
- >>> calcCfdScore("ATGTGGAGATTGCCACCTACCGG", "ATCTGGAGATTGCCACCTACAGG")
- 0.384615385
- """
- global mm_scores, pam_scores
- if mm_scores is None:
- mm_scores,pam_scores = get_mm_pam_scores()
- wt = guideSeq.upper()
- off = otSeq.upper()
- m_wt = re.search('[^ATCG]',wt)
- m_off = re.search('[^ATCG]',off)
- if (m_wt is None) and (m_off is None):
- pam = off[-2:]
- sg = off[:20]
- cfd_score = calc_cfd(wt,sg,pam)
- if doCfdFix:
- cfd_score = cfd_score*100.0
- return cfd_score
- return -1
- # ==== END CFD score source provided by John Doench
- # --- END OF SCORING ROUTINES
- def getSizeFname(genome):
- " return name of chrom.sizes file "
- genomeDir = genomesDir # make local
- sizeFname = "%(genomeDir)s/%(genome)s/%(genome)s.sizes" % locals()
- return sizeFname
- def parseChromSizes(genome):
- " return chrom sizes as dict chrom -> size "
- sizeFname = getSizeFname(genome)
- ret = {}
- for line in open(sizeFname).read().splitlines():
- fields = line.split()
- chrom, size = fields[:2]
- ret[chrom] = int(size)
- return ret
- def extendAndGetSeq(db, chrom, start, end, strand, oldSeq, flank=FLANKLEN):
- """ extend (start, end) by flank and get sequence for it using twoBitTwoFa.
- Return None if not possible to extend.
- #>>> extendAndGetSeq("hg19", "chr21", 10000000, 10000005, "+", flank=3)
- #'AAGGAATGTAG'
- """
- assert("|" not in chrom) # we are using | to split info in BED files. | is not allowed in the fasta
- chromSizes = parseChromSizes(db)
- maxEnd = chromSizes[chrom]+1
- start -= flank
- end += flank
- if start < 0 or end > maxEnd:
- return None
- genomeDir = genomesDir
- twoBitFname = "%(genomeDir)s/%(db)s/%(db)s.2bit" % locals()
- progDir = binDir
- genome = db
- cmd = "%(progDir)s/twoBitToFa %(genomeDir)s/%(genome)s/%(genome)s.2bit stdout -seq='%(chrom)s' -start=%(start)s -end=%(end)s" % locals()
- proc = subprocess.Popen(cmd, shell=True, stdout=subprocess.PIPE, encoding="utf8")
- seqStr = proc.stdout.read()
- proc.wait()
- if proc.returncode!=0:
- errAbort("Could not run '%s'. Return code %s" % (cmd, str(proc.returncode)))
- faFile = StringIO(seqStr)
- seqs = parseFasta(faFile)
- assert(len(seqs)==1)
- seq = list(seqs.values())[0].upper()
- if strand=="-":
- seq = revComp(seq)
- genomeSeq = seq[FLANKLEN:(FLANKLEN+len(oldSeq))].upper()
- if oldSeq.upper() not in genomeSeq:
- logging.warn("Input sequence has SNPs compared to genome, not returning extended seq:")
- logging.warn("- Input sequence: %s" % oldSeq)
- logging.warn("- Genome sequence: %s" % genomeSeq)
- logging.warn("- Diff String : %s" % highlightMismatches(oldSeq, genomeSeq, 0))
- return None
- # ? make sure that user annotations, like added Ns, are retained in the long sequence
- #fixedSeq = seq[:100]+oldSeq+seq[-100:]
- #assert(len(fixedSeq)==len(seq))
- return seq
- def getExtSeq(seq, start, end, strand, extUpstream, extDownstream, extSeq=None, extFlank=FLANKLEN):
- """ extend (start,end) by extUpstream and extDownstream and return the subsequence
- at this position in seq.
- Return None if there is not enough space to extend (start, end).
- extSeq is a sequence with extFlank additional flanking bases on each side. It can be provided
- optionally and is used if needed to return a subseq.
- Careful: returned sequence might contain lowercase letters.
- >>> getExtSeq("AACCTTGG", 2, 4, "+", 2, 4)
- 'AACCTTGG'
- >>> getExtSeq("CCAACCTTGGCC", 4, 6, "-", 2, 3)
- 'AAGGTTG'
- >>> getExtSeq("AA", 0, 2, "+", 2, 3)
- >>> getExtSeq("AA", 0, 2, "+", 2, 3, extSeq="CAGAATGA", extFlank=3)
- 'AGAATGA'
- >>> getExtSeq("AA", 0, 2, "-", 2, 3, extSeq="CAGAATGA", extFlank=3)
- 'CATTCTG'
- """
- assert(start>=0)
- assert(end<=len(seq))
- # extend
- if strand=="+":
- extStart, extEnd = start-extUpstream, end+extDownstream
- else:
- extStart, extEnd = start-extDownstream, end+extUpstream
- # check for out of bounds and get seq
- if extStart >= 0 and extEnd <= len(seq):
- logging.debug("using input seq, pos %d-%d" % (extStart, extEnd))
- subSeq = seq[extStart:extEnd]
- else:
- if extSeq==None:
- return None
- # lift to extSeq coords and get seq
- extStart += extFlank
- extEnd += extFlank
- assert(extStart >= 0)
- assert(extEnd <= len(extSeq))
- subSeq = extSeq[extStart:extEnd]
- logging.debug("using extended seq, pos %d-%d" % (extStart, extEnd))
- if strand=="-":
- logging.debug("revcomp'ing result")
- subSeq = revComp(subSeq)
- # check that the extended sequence really contains the whole input seq
- # e.g. when user has added nucleotides to a otherwise matching sequence
- #if seq.upper() not in subSeq.upper():
- #debug("seq is not in extSeq")
- #subSeq = None
- logging.debug("Got -%d/+%d-extended seq for (%d, %d, %s) = %s. Result: %s." %
- (extUpstream, extDownstream, start, end, strand, seq[start:end], subSeq))
- return subSeq
- def pamStartToGuideRange(startPos, strand, pamLen):
- """ given a PAM start position and its strand, return the (start,end) of the guide.
- Coords can be negative or exceed the length of the input sequence.
- """
- if not pamIsFirst:
- if strand=="+":
- return (startPos-GUIDELEN, startPos)
- else: # strand is minus
- return (startPos+pamLen, startPos+pamLen+GUIDELEN)
- else:
- if strand=="+":
- return (startPos+pamLen, startPos+pamLen+GUIDELEN)
- else: # strand is minus
- return (startPos-GUIDELEN, startPos)
- def htmlHelp(text):
- " show help text with tooltip or modal dialog "
- className = "tooltipster"
- if "href" in text:
- className = "tooltipsterInteract"
- print('''<img style="padding-bottom: 3px; height:1.1em; width:1.0em" src="%simage/info-small.png" class="help %s" title="%s" />''' % (HTMLPREFIX, className, text))
- def htmlWarn(text):
- " show help text with tooltip "
- print('''<img style="height:0.9em; width:0.8em; padding-bottom: 2px" src="%simage/warning-32.png" class="help tooltipster" title="%s" />''' % (HTMLPREFIX, text))
- def readRestrEnzymes():
- """ parse restrSites.txt and
- return as dict length -> list of (name, suppliers, seq) """
- fname = "restrSites.txt"
- enzList = {}
- for line in open(join(baseDir, fname)):
- if line.startswith("#"):
- continue
- seq, name, suppliers = line.rstrip("\n").rstrip("\r").split("\t")
- suppliers = tuple(suppliers.split(","))
- enzList.setdefault(len(seq), []).append( (name, suppliers, seq) )
- return enzList
- def patMatch(seq, pat, notDegPos=None):
- """ return true if pat matches seq, both have to be same length
- do not match degenerate codes at position notDegPos (0-based)
- """
- assert(len(seq)==len(pat))
- for x in range(0, len(pat)):
- patChar = pat[x]
- nuc = seq[x]
- assert(patChar in "MKYRACTGNWSDVBH")
- assert(nuc in "MKYRACTGNWSDX")
- if notDegPos!=None and x==notDegPos and patChar!=nuc:
- return False
- if nuc=="X":
- return False
- if patChar=="N":
- continue
- if patChar=="H" and nuc in "ACT":
- continue
- if patChar=="D" and nuc in "AGT":
- continue
- if patChar=="B" and nuc in "CGT":
- continue
- if patChar=="V" and nuc in "ACG":
- continue
- if patChar=="W" and nuc in "AT":
- continue
- if patChar=="S" and nuc in "GC":
- continue
- if patChar=="M" and nuc in "AC":
- continue
- if patChar=="K" and nuc in "TG":
- continue
- if patChar=="R" and nuc in "AG":
- continue
- if patChar=="Y" and nuc in "CT":
- continue
- if patChar!=nuc:
- return False
- return True
- def findSite(seq, restrSite):
- """ return the positions where restrSite matches seq
- seq can be longer than restrSite
- Do not allow degenerate characters to match at position len(restrSite) in seq
- """
- posList = []
- for i in range(0, len(seq)-len(restrSite)+1):
- subseq = seq[i:i+len(restrSite)]
- # JP does not want any potential site to be suppressed
- #if i<len(restrSite):
- #isMatch = patMatch(subseq, restrSite, len(restrSite)-i-1)
- #else:
- #isMatch = patMatch(subseq, restrSite)
- isMatch = patMatch(subseq, restrSite)
- if isMatch:
- posList.append( (i, i+len(restrSite)) )
- return posList
- def matchRestrEnz(allEnzymes, guideSeq, pamSeq, pamPlusSeq, pamPat):
- """ return list of enzymes that overlap the -3 position in guideSeq
- returns dict (name, pattern, suppliers) -> list of matching positions
- """
- matches = defaultdict(set)
- if pamPlusSeq is None:
- pamPlusSeq = "XXXXX" # make sure that we never match a restriction site outside the seq boundaries
- fullSeq = concatGuideAndPam(guideSeq, pamSeq, pamPlusSeq)
- #print guideSeq, pamSeq, pamPlusSeq, fullSeq, "<br>"
- for siteLen, sites in allEnzymes.items():
- if pamIsCpf1(pamPat):
- # most modified position: 4nt from the end
- # see http://www.nature.com/nbt/journal/v34/n8/full/nbt.3620.html
- # Figure 1
- startSeq = len(fullSeq)-4-pamPlusLen-(siteLen)+1
- else:
- # most modified position for Cas9: 3bp from the end
- startSeq = len(fullSeq)-len(pamSeq)-3-pamPlusLen-(siteLen)+1
- seq = fullSeq[startSeq:].upper()
- for name, suppliers, restrSite in sites:
- posList = findSite(seq, restrSite)
- if len(posList)!=0:
- liftOffset = startSeq
- posList = [(liftOffset+x, liftOffset+y) for x,y in posList]
- matches.setdefault((name, restrSite, suppliers), set()).update(posList)
- return matches
- def mergeGuideInfo(seq, startDict, pamPat, otMatches, inputPos, effScores, sortBy=None, org=None):
- """
- merges guide information from the sequence, the efficiency scores and the off-targets.
- creates rows with too many fields. needs refactoring.
- for each pam in startDict, retrieve the guide sequence next to it and score it
- sortBy can be "effScore", "mhScore", "oofScore" or "pos"
- """
- allEnzymes = readRestrEnzymes()
- guideData = []
- guideScores = {}
- hasNotFound = False
- pamIdToSeq = {}
- pamSeqs = list(flankSeqIter(seq.upper(), startDict, len(pamPat), True))
- for pamId, pamStart, guideStart, strand, guideSeq, pamSeq, pamPlusSeq in pamSeqs:
- # matches in genome
- # one desc in last column per OT seq
- if pamId in otMatches:
- pamMatches = otMatches[pamId]
- guideSeqFull = concatGuideAndPam(guideSeq, pamSeq)
- mutEnzymes = matchRestrEnz(allEnzymes, guideSeq, pamSeq, pamPlusSeq, pamPat)
- posList, otDesc, guideScore, guideCfdScore, last12Desc, ontargetDesc, \
- repCount = \
- annotateOfftargets(org, pamMatches, guideSeqFull, pamPat, inputPos)
- if repCount!=0:
- guideScore = 0
- guideCfdScore = 0
- # no off-targets found?
- else:
- posList, otDesc, guideScore = None, "Not found", -1
- guideCfdScore = -1
- last12Desc = ""
- hasNotFound = True
- mutEnzymes = []
- ontargetDesc = ""
- repCount = 0
- seq34Mer = None
- guideRow = [guideScore, guideCfdScore, effScores.get(pamId, {}), pamStart, guideStart, strand, pamId, guideSeq, pamSeq, posList, otDesc, last12Desc, mutEnzymes, ontargetDesc, repCount]
- guideData.append( guideRow )
- guideScores[pamId] = guideScore
- pamIdToSeq[pamId] = guideSeq
- if sortBy == "pos":
- sortFunc = (lambda row: row[3])
- reverse = False
- elif sortBy == "offCount":
- sortFunc = (lambda row: len(row[9]))
- reverse = False
- elif sortBy == "cfdSpec":
- sortFunc = operator.itemgetter(1)
- reverse = True
- elif sortBy == "spec" or sortBy is None:
- sortFunc = (lambda row: row[0])
- reverse = True
- elif sortBy is not None and not sortBy.endswith("pec"):
- sortFunc = (lambda row: row[2].get(sortBy, 0))
- reverse = True
- else:
- errAbort("Unknown sortBy value. This is a bug. Please contact us.")
- guideData.sort(reverse=reverse, key=sortFunc)
- return guideData, guideScores, hasNotFound, pamIdToSeq
- def printDownloadTableLinks(batchId, addTsv=False):
- print('<div id="downloads" style="text-align:left">')
- print("Download as Excel tables: ", end=' ')
- print('<a href="crispor.py?batchId=%s&download=guides&format=xls">Guides</a> / ' % batchId, end=' ')
- if not pamIsFirst and not saCas9Mode:
- print('<a href="crispor.py?batchId=%s&showAllScores=1&download=guides&format=xls">Guides, all scores</a> / ' % batchId, end=' ')
- print('<a href="crispor.py?batchId=%s&download=offtargets&format=xls">Off-targets</a> / ' % batchId, end=' ')
- print(('<a href="crispor.py?batchId=%s&satMut=1">Saturating mutagenesis assistant</a><br>' % batchId))
- #print "<small>Plasmid Editor: ",
- #print '<a href="crispor.py?batchId=%s&download=genbank">Guides</a></small>' % batchId,
- if addTsv:
- print("<small>Tab-sep format: ", end=' ')
- print('<a href="crispor.py?batchId=%s&download=guides&format=tsv">Guides</a> / ' % batchId, end=' ')
- print('<a href="crispor.py?batchId=%s&download=offtargets&format=tsv">Off-targets</a></small>' % batchId, end=' ')
- print('</div>')
- def hasGeneModels(org):
- " return true if this organism has gene model information "
- geneFname = join(genomesDir, org, org+".segments.bed")
- return isfile(geneFname)
- def printTableHead(pam, batchId, chrom, org, varHtmls, showColumns):
- " print guide score table description and columns "
- # one row per guide sequence
- if not pamIsCpf1(pam):
- print('''<div class='substep'>Ranked by default from highest to lowest specificity score (<a target='_blank' href='http://dx.doi.org/10.1038/nbt.2647'>Hsu et al., Nat Biot 2013</a>). Click on a column title to rank by a score.<br>''')
- #print("""<b>Our recommendation:</b> Use Fusi for in-vivo (U6) transcribed guides, Moreno-Mateos for in-vitro (T7) guides injected into Zebrafish/Mouse oocytes.<br>""")
- print('''If you use this website, please cite our <a href="https://academic.oup.com/nar/article/46/W1/W242/4995687">paper in NAR 2018</a>.''')
- print("Too much information? Look at the <a target=_blank href='manual/'>CRISPOR manual</a>.<p>")
- print('</div>')
- printDownloadTableLinks(batchId)
- print("""
- <script type="text/javascript">
- function allRows() {
- $("guideRow").show();
- }
- //function copySeq() {
- //var c = new ClipboardJS('#seqAsText');
- //var copyText = document.getElementById("seqAsText");
- //var selRes = copyText.select();
- //var val = copyText.value;
- //var res = document.execCommand("copy");
- //alert("The input sequence is now in your clipboard. You can paste it into other programs.");
- //}
- $(document).ready( function() {
- //#$('#copyLink').click( copySeq );
- var clipboard = new ClipboardJS('#copyLink');
- clipboard.on('success', function(e) {
- alert("The input sequence is now in your clipboard. You can paste it into other programs.");
- console.info('Action:', e.action);
- console.info('Text:', e.text);
- console.info('Trigger:', e.trigger);
- e.clearSelection();
- });
- $('d').mouseenter( onEditHover );
- $('d').mouseleave ( onEditOut );
- });
- function onEditOut() {
- $('#editHover').hide();
- }
- function colorChar(str, pos) {
- /* put a span-color tag around the char at pos in str and return result */
- var prefix = str.substring(0, pos);
- var hlChar = str[pos];
- var suffix = str.substring(pos+1);
- return prefix+"<mut>"+hlChar+"</mut>"+suffix;
- }
- function onEditHover(ev) {
- /* user hovers over an edit letter */
- ev.preventDefault();
- var oldEl = document.getElementById("editHover");
- if (oldEl)
- oldEl.remove();
- console.log(ev.target);
- const boundBox = ev.target.getBoundingClientRect();
- var x = boundBox.left;
- var y = boundBox.top;
- y += 14;
- var div = document.createElement('div');
- div.id = "editHover";
- div.style.width="400px";
- div.style.height="200px";
- div.style.border="1px solid black";
- div.style.padding="10px";
- div.style.position="fixed";
- div.style.backgroundColor="white";
- div.style.left=x+"px";
- div.style.top=y+"px";
- var pos = parseInt(this.getAttribute("pos"));
- var nucl = this.textContent;
- if (nucl.toUpperCase()==="T")
- origNucl = "C";
- else
- origNucl = "G";
- var htmls=[];
- htmls.push("The following guides can mutate "+origNucl+" to "+nucl+" at position "+pos+":<br>");
- htmls.push("<table class='editTable'>");
- htmls.push("<tr><th>Guide ID</th><th>Guide Sequence</th><th>Komor score</th><th>Spec. Score</th></tr>");
- var guides = editData[pos][nucl];
- guides.sort( function (a, b) { a[4] - b[4] } ); // sort by komor score
- for (var i=0; i<guides.length; i++) {
- guide = guides[i];
- pamId = guide[0];
- guideSeq = guide[1];
- pam = guide[2];
- mutPos = guide[3];
- beScore = guide[4];
- specScore = guide[5];
- htmls.push("<tr>");
- htmls.push("<td>"+pamId+"</td>");
- htmls.push("<td><tt>"+colorChar(guideSeq, mutPos)+" "+pam+"</tt></td>");
- htmls.push("<td>"+beScore.toFixed(2)+"</td>");
- htmls.push("<td>"+specScore+"</td>");
- htmls.push("</tr>");
- }
- htmls.push("</table>");
- $(div).append(htmls.join(""));
- document.body.appendChild(div);
- }
- function onlyWith(doPrefix) {
- /* show only guide rows and guide sequence viewer features that start with a prefix */
- if ($("#onlyWith"+doPrefix+"Box").prop("checked"))
- {
- $(".prefixBox").prop("checked", false);
- $("#onlyWith"+doPrefix+"Box").prop("checked", true);
- //$(".guideRow").show();
- $(".guideRow").css("visibility", "visible");
- $(".guideRowNoPrefix"+doPrefix).hide();
- // special handling for sequence viewer: hide() would destroy the layout there
- $(".guideRowNoPrefix"+doPrefix+"Seq").css("visibility", "hidden");
- }
- else
- {
- $(".prefixBox").prop("checked", false);
- $(".guideRow").show();
- $(".guideRowNoPrefix"+doPrefix+"Seq").css("visibility", "visible");
- }
- }
- function displayClass(className, dispVal) {
- /* hide in a loop, works around Safari stack size limits that crash jquery functions */
- var els = document.getElementsByClassName(className);
- for (var el of els) {
- el.style.display = dispVal;
- }
- }
- function onlyExons() {
- /* show only off-targets in exons */
- if ($("#onlyExonBox").prop("checked")) {
- $(".otMore").show();
- $(".otMoreLink").hide();
- $(".otLessLink").hide();
- displayClass("notExon", "none");
- }
- else {
- if ($("#onlySameChromBox").prop("checked")) {
- $(".notExon:not(.diffChrom)").show();
- }
- else {
- displayClass("notExon", "block");
- $(".otMoreLink").show();
- $(".otMore").hide();
- }
- }
- }
- function onlySameChrom() {
- if ($("#onlySameChromBox").prop("checked"))
- {
- $(".otMore").show();
- $(".otMoreLink").hide();
- $(".otLessLink").hide();
- $(".diffChrom").hide();
- }
- else {
- if ($("#onlyExonBox").prop("checked")) {
- $(".diffChrom:not(.notExon)").show();
- }
- else {
- $(".diffChrom").show();
- $(".otMoreLink").show();
- $(".otMore").hide();
- }
- }
- }
- function showAllOts(classId) {
- $("#"+classId).show();
- $("#"+classId+"MoreLink").hide();
- $("#"+classId+"LessLink").show();
- }
- function showLessOts(classId) {
- $("#"+classId).hide();
- $("#"+classId+"MoreLink").show();
- $("#"+classId+"LessLink").hide();
- }
- </script>
- """)
- print('<table id="otTable" style="background:white;table-layout:fixed; overflow:scroll; width:100%">')
- print('<thead>')
- print('<tr style="border-bottom:none; border-left:5px solid black; background-color:#F0F0F0">')
- print('<th style="width:80px; border-bottom:none"><a href="crispor.py?batchId=%s&sortBy=pos" class="tooltipster" title="Click to sort the table by the position of the PAM site">Position/<br>Strand</a>' % batchId)
- htmlHelp("You can click on the links in this column to highlight the <br>PAM site in the sequence viewer at the top of the page.")
- print('</th>')
- print('<th style="width:235px; border-bottom:none">Guide Sequence + <i>PAM</i><br>')
- print ('+ Restriction Enzymes')
- htmlHelp("Restriction enzymes can be very useful for screening mutations induced by the guide RNA using PCR and Restrictrion frament length polymorphism (RFLP).<br>Enzyme sites shown here overlap the main cleavage site 3bp 5' to the PAM.<br>Digestion of the PCR product with these enzymes will not cut the product if the genome was mutated by Cas9. This is a lot easier than screening with the T7 assay, Surveyor or sequencing.")
- print('<br>')
- if varHtmls is not None:
- print(' + Variants')
- htmlHelp("Variants that overlap the guide sequence are shown. You can change the variant database with the drop-down box above the sequence viewer at the top of the page.")
- print('<br>')
- print('''<small>''')
- print('''<input type="checkbox" class="prefixBox" id="onlyWithGBox" onchange="onlyWith('G')">Only G-''')
- print('''<input type="checkbox" class="prefixBox" id="onlyWithGGBox" onchange="onlyWith('GG')">Only GG-''')
- print('''<input type="checkbox" class="prefixBox" id="onlyWithABox" onchange="onlyWith('A')">Only A-''')
- htmlHelp("The three checkboxes allow you to show only guides that start with GG-, G- or A-. While we recommend prefixing a 20bp guide with G for U6 expression with spCas9, some protocols recommend using only guides with a G- prefix for U6 and A- for U3.")
- print('''</small>''')
- if not pamIsCpf1(pam):
- print('<th style="width:80px; border-bottom:none"><a href="crispor.py?batchId=%s&sortBy=spec" class="tooltipster" title="Click to sort the table by specificity score. Hover over the (i) bubble on the right to get more information about the specificity score.">MIT Specificity Score</a>' % batchId)
- if pamIsSaCas9(pam):
- htmlHelp("The higher the specificity score, the lower are off-target effects in the genome.<br>This specificity score has been adapted for SaCas9 and based on the off-target scores shown on mouse-over. The algorithm was provided by Josh Tycko. Like the MIT score for spCas9, it is aggregated from all off-target scores and ranges 0-100. See <a href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6063963/'>Tycko et al. Nat Comm 2018</a> for details.")
- else:
- htmlHelp("The higher the specificity score, the lower are off-target effects in the genome.<br>The specificity score ranges from 0-100 and measures the uniqueness of a guide in the genome. See <a href='http://dx.doi.org/10.1038/nbt.2647'>Hsu et al. Nat Biotech 2013</a>. We recommend values >50, where possible. See <a target=_blank href='manual/#offs'>the CRISPOR manual</a>")
- print("</th>")
- if "cfdGuideScore" in showColumns:
- print('<th style="width:60px; border-bottom:none"><a href="crispor.py?batchId=%s&sortBy=cfdSpec" class="tooltipster" title="Click to sort the table by CFD specificity score">CFD Spec. score</a>' % batchId)
- htmlHelp("The CFD specificity score, inspired like guidescan.com, behaves like the MIT specificity score, but it is based on the more accurate CFD off-target model, from <a href='http://www.nature.com/nbt/journal/v34/n2/full/nbt.3437.html'>Doench 2016</a>, which is also used by Crispor to rank the off-targets. The CFD specificity score correlates better than the MIT score with the total off-target cleavage fraction of a guide, see <a target=_blank href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6731277/'>Tycko et al, Nat Comm 2019</a> and also the <a target=_blank href='/manual/#faq'>CRISPOR manual</a>.")
- print("</th>")
- if len(scoreNames)==2 or pamIsCpf1(pam) or pamIsSaCas9(pam):
- print('<th style="width:150px; height:100px; border-bottom:none" colspan="%d">Predicted Efficiency' % (len(scoreNames)))
- else:
- print('<th style="width:270px; border-bottom:none" colspan="%d">Predicted Efficiency' % (len(scoreNames))) # -1 because proxGc is in scoreNames but has no column
- htmlHelp("The higher the efficiency score, the more likely is cleavage at this position. For details on the scores, mouseover their titles below.<br>Note that these predictions are not very accurate, they merely enrich for more efficient guides by a factor of 2-3 so you have to test a few guides to see the effect. <a target=_blank href='manual/#onEff'>Read the CRISPOR manual</a>")
- if not pamIsCpf1(pam) and not pamIsSaCas9(pam):
- if cgiParams.get("showAllScores", "0")=="0":
- print(("""<br><a style="font-size:12px" href="%s" class="tooltipsterInteract" title="By default, only the two most relevant scores are shown, based on our study <a href='http://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-1012-2'>Haeussler et al. 2016</a>. Click this link to show all efficiency scores.">Show all scores</a>""" % cgiGetSelfUrl({"showAllScores":"1"}, anchor="otTable")))
- scoreDescs["crisprScan"][0] = "Mor.-Mateos"
- else:
- print(("""<br><a style="font-size:12px" href="%s" class="tooltipsterInteract" title="Show only the two main scores">Show main scores</a>""" % cgiGetSelfUrl({"showAllScores":None}, anchor="otTable")))
- print('</th>')
- mhColName="Outcome"
- if not baseEditor:
- if len(mutScoreNames)<=1:
- mhColName = ""
- oofWidth=45
- #oofDesc = "Click on score to show micro-homology"
- #oofDesc = ""
- else:
- oofWidth=67
- #oofDesc = "Click score for details"
- colSpan = len(mutScoreNames)
- print('<th colspan=%d style="width:%dpx; border-bottom:none"><a href="crispor.py?batchId=%s&sortBy=oof" class="tooltipster" title="Prediction of the DNA sequence after strand break repair. Click to sort the table by frameshift/out-of-frame scores. Hover over the score names to show information about a particular score. Click a score number to see the predicted indel pattern around the guide.">%s</a>' % (colSpan, oofWidth, batchId, mhColName))
- #htmlHelp(scoreDescs["oof"][1])
- #print "<small>%s</small>" % oofDesc
- print('</th>')
- print('<th style="width:117px; border-bottom:none"><a href="crispor.py?batchId=%s&sortBy=offCount" class="tooltipster" title="Click to sort the table by number of off-targets">Off-targets for <br>0-1-2-3-4 mismatches<br></a><span style="color:grey">+ next to PAM </span>' % (batchId))
- altPamsHelp = [pam]
- if pam in offtargetPams:
- altPamsHelp.extend(offtargetPams[pam])
- htmlHelp("For each number of mismatches, the number of off-targets is indicated.<br>Example: 1-3-20-50-60 means 1 off-target with 0 mismatches, 3 off-targets with 1 mismatch, <br>20 off-targets with 2 mismatches, etc.<br>The CRISPOR website only searches up to four mismatches (use the command line version for 5 or 6). Off-targets are considered if they are flanked by one of these motifs: %s .<br>Shown in grey are the off-targets that have no mismatches in the 12 bp adjacent to the PAM. These are the most likely off-targets." % (", ".join(altPamsHelp)))
- print("</th>")
- print('<th style="width:*; border-bottom:none">Genome Browser links to matches sorted by CFD off-target score')
- htmlHelp("For each off-target the number of mismatches is indicated and linked to a genome browser. <br>Matches are ranked by CFD off-target score (see Doench 2016 et al) from most to least likely.<br>Matches can be filtered to show only off-targets in exons or on the same chromosome as the input sequence.<br>On most organisms, you can click the links below to open a window with a genome browser at this position.")
- print('<br><small>')
- print('<input type="hidden" name="batchId" value="%s">' % batchId)
- if hasGeneModels(org):
- print('''<input type="checkbox" id="onlyExonBox" onchange="onlyExons()">exons only''')
- else:
- print('<small title="When this genome was loaded into CRISPOR, gene models were not available. Contact us if you want to filter for off-targets in exons and think that a gene models are now available for this genome." style="color:grey">No exons.</small>')
- if chrom!="":
- if chrom[0].isdigit():
- chrom = "chrom "+chrom
- print('''<input type="checkbox" id="onlySameChromBox" onchange="onlySameChrom()">%s only''' % chrom)
- else:
- print('<small style="color:grey"> No match, no chrom filter</small>')
- print("</small>")
- print("</th>")
- print("</tr>")
- # subheaders
- print('<tr style="border-top:none; border-left: solid black 5px; background-color:#F0F0F0">')
- print('<th style="border-top:none"></th>')
- print('<th style="border-top:none"></th>')
- if "cfdGuideScore" in showColumns:
- print('<th style="border-top:none"></th>')
- if not pamIsCpf1(pam):
- print('<th style="border-top:none"></th>')
- for scoreName in scoreNames:
- if scoreName in ["oof", "proxGc"] or "oof" in scoreName:
- continue
- scoreLabel, scoreDesc = scoreDescs[scoreName]
- print('<th style="width: 10px; border: none; border-top:none; border-right: none" class="rotate"><div><span><a title="%s" class="tooltipsterInteract" href="crispor.py?batchId=%s&sortBy=%s">%s</a></span></div></th>' % (scoreDesc, batchId, scoreName, scoreLabel))
- if "proxGc" in scoreNames:
- # the ProxGC score comes next
- print('''<th style="border: none; border-top:none; border-right: none; border-left:none" class="rotate">''')
- print('''<div><span style="border-bottom:none">''')
- print('''<a title="This column shows two heuristics based on observations rather than computational models: <a href='http://www.cell.com/cell-reports/abstract/S2211-1247%2814%2900827-4'>Ren et al</a> 2014 obtained the highest cleavage in Drosophila when the final 6bp contained >= 4 GCs, based on data from 39 guides. <a href='http://www.genetics.org/content/early/2015/02/18/genetics.115.175166.abstract'>Farboud et al.</a> obtained the highest cleavage in C. elegans for the 10 guides that ended with -GG, out of the 50 guides they tested.<br>The column contains + if the final GC count is >= 4 and GG if the guide ends with GG." href="crispor.py?batchId=%s&sortBy=finalGc6" class="tooltipsterInteract">Prox GC</span></div></th>''' % (batchId))
- # these are empty cells to fill up the row and avoid white space
- for scoreName in mutScoreNames:
- scoreLabel, scoreDesc = scoreDescs[scoreName]
- print('<th style="width: 10px; border-top:none; border-right: none" class="rotate"><div><span><a title="%s" class="tooltipsterInteract" href="crispor.py?batchId=%s&sortBy=%s">%s</a></span></div></th>' % (scoreDesc, batchId, scoreName, scoreLabel))
- print('<th style="border-top:none"></th>')
- print('<th style="border-top:none"></th>')
- print("</tr>")
- print('</thead>')
- def scoreToColor(guideScore):
- if guideScore is None:
- color = ("#000000", "black")
- elif guideScore > 50:
- color = ("#32cd32", "green")
- elif guideScore > 20:
- color = ("#ffff00", "yellow")
- elif guideScore==-1:
- color = ("#000000", "black")
- else:
- color = ("#aa0114", "red")
- return color
- def hexToRgb(hexCode):
- " convert hex color to RGB in UCSC format, https://stackoverflow.com/questions/29643352/converting-hex-to-rgb-value-in-python "
- hexCode = hexCode.lstrip("#")
- return ",".join(tuple(str(int(hexCode[i:i+2], 16)) for i in (0, 2 ,4)))
- def makeOtBrowserLinks(otData, chrom, dbInfo, pamId):
- " return a list with the html texts of the offtarget links "
- links = []
- i = 0
- for otSeq, score, cfdScore, editDist, pos, gene, alnHtml, inLinkage in otData:
- cssClasses = ["tooltipster"]
- if not gene.startswith("exon:"):
- cssClasses.append("notExon")
- if pos.split(":")[0]!=chrom:
- cssClasses.append("diffChrom")
- if inLinkage:
- cssClasses.append("inLinkage")
- classStr = ""
- if len(cssClasses)!=0:
- classStr = ' class="%s"' % " ".join(cssClasses)
- link = makeBrowserLink(dbInfo, pos, gene, alnHtml, ["tooltipster"])
- editDist = str(editDist)
- links.append( '''<div%(classStr)s>%(editDist)s:%(link)s</div>''' % locals() )
- return links
- def filterOts(otDatas, minScore):
- " remove all offtargets with score < minScore "
- newList = []
- for otData in otDatas:
- score = otData[1]
- if score > minScore:
- newList.append(otData)
- return newList
- def findOtCutoff(otData):
- " try cutoffs 0.5, 1.0, 2.0, 3.0 until not more than 20 offtargets left "
- for cutoff in [0.3, 0.5, 1.0, 2.0, 3.0, 10.0, 99.9]:
- otData = filterOts(otData, cutoff)
- if len(otData)<=30:
- return otData, cutoff
- if len(otData)>30:
- return otData[:30], None
- return otData, 1000
- def printNote(s):
- print('<div style="text-align:left; background-color: aliceblue; padding:5px; border: 1px solid black"><strong>Note:</strong>')
- print(s)
- print("</div>")
- def printWarning(s):
- print('<div style="text-align:left; background-color: #FFDDDD; padding:5px; border: 1px solid black"><strong>Warning:</strong>')
- print(s)
- print('</div>')
- def printNoEffScoreFoundWarn(effScoresCount, pam):
- if effScoresCount==0 and not pamIsCpf1(pam):
- note = "No guide could be scored for efficiency. This happens when the input sequence is shorter than 100bp and there is no genome available to extend it or if there is simply not guide socring method. In the first case, please add flanking 50bp on both sides of the input sequence and submit this new, longer sequence. For the second case, you can contact me and suggest an efficiency scoring method, send me the published paper in this case."
- printNote(note)
- def showGuideTable(guideData, pam, otMatches, dbInfo, batchId, org, chrom, varHtmls):
- " shows table of all PAM motif matches "
- print("<br><div class='title'>Predicted guide sequences for PAMs</div>")
- global scoreNames
- if (cgiParams.get("showAllScores", "0")=="1"):
- scoreNames = allScoreNames
- showColumns = set()
- # show the CFD guide score?
- if pamIsSpCas9(pam):
- showColumns.add("cfdGuideScore")
- showPamWarning(pam)
- showNoGenomeWarning(dbInfo)
- printTableHead(pam, batchId, chrom, org, varHtmls, showColumns)
- count = 0
- effScoresCount = 0
- showProxGcCol = ("proxGc" in scoreNames)
- for guideRow in guideData:
- guideScore, guideCfdScore, effScores, pamStart, guideStart, strand, pamId, guideSeq, \
- pamSeq, otData, otDesc, last12Desc, mutEnzymes, ontargetDesc, repCount = guideRow
- color = scoreToColor(guideScore)[0]
- classStr = cssClassesFromSeq(guideSeq)
- print('<tr id="%s" class="%s" style="border-left: 5px solid %s">' % (pamId, classStr, color))
- # position and strand
- #print '<td id="%s">' % pamId
- print('<td>')
- print('<a href="#list%s">' % (pamId))
- print(str(pamStart+1)+" /")
- if strand=="+":
- print('fw')
- else:
- print('rev')
- print('</a>')
- print("</td>")
- # sequence with variants and PCR primer link
- print("<td>")
- print("<small>")
- # guide sequence + PAM sequence
- if pamIsFirst:
- fullGuideHtml = "<tt><i>"+pamSeq+"</i> " + guideSeq+"</tt>"
- spacePos = len(pamSeq)
- else:
- fullGuideHtml = "<tt>"+guideSeq + " <i>" + pamSeq+"</i></tt>"
- spacePos = len(guideSeq)
- print(fullGuideHtml)
- print("<br>")
- # variant-string
- if varHtmls is not None:
- varFound = False
- varStrs = []
- guideHtmlStart = min(guideStart, pamStart)
- guideHtmls = varHtmls[guideHtmlStart:guideHtmlStart+len(guideSeq)+len(pamSeq)]
- if strand=="-":
- guideHtmls = list(reversed(guideHtmls))
- for i in range(len(guideHtmls)):
- html = guideHtmls[i]
- if html!=".":
- varFound = True
- if i==spacePos:
- varStrs.append(" ")
- varStrs.append(html)
- print(("<tt style='color:#888888'>%s</tt><br>" % ("".join(varStrs))))
- if "TTTT" in guideSeq.upper():
- text = "This guide contains the sequence TTTT. It cannot be transcribed with a U6 or U3 promoter, as TTTT terminates the transcription."
- htmlWarn(text)
- print(' Not with U6/U3')
- print("<br>")
- if pam=="NGG":
- grafType = crisporEffScores.getGrafType(guideSeq)
- if grafType:
- if grafType=="tt":
- grafText = "The guide ends with TTC or TTT or contains only T and C in the last four nucleotides and more than 2 Ts or at least one TT and one T or C ('TT-motif'). These guides should be avoided in polymerase III (Pol III)-based gene editing experiments requiring high sgRNA expression levels."
- elif grafType=="gcc":
- grafText = "The guide ends with [AGT]GCC or GCCT ('GCC motif'). These sgRNAs appear to be inefficient irrespective of the delivery method and should thus be generally avoided."
- text = "This guide contains one of the motifs described by <a target=_blank href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6352712/'>Graf et al, Cell Reports 2019</a>. %s " % grafText
- htmlWarn(text)
- print(' Inefficient')
- print("<br>")
- if gcContent(guideSeq)>0.75:
- text = "This sequence has a GC content higher than 75%.<br>In the data of Tsai et al Nat Biotech 2015, the two guide sequences with a high GC content had almost as many off-targets as all other sequences combined. We do not recommend using guide sequences with such a high GC content."
- htmlWarn(text)
- print(' High GC content')
- print("<br>")
- if gcContent(guideSeq)<0.25:
- text = "This sequence has a GC content lower than 25%.<br>In the data of Wang/Sabatini/Lander Science 2014, guides with a very low GC content had low cleavage efficiency."
- htmlWarn(text)
- print(' Low GC content<br>')
- print("<br>")
- if len(mutEnzymes)!=0:
- print("<div style='margin-top: 3px'>Enzymes: <i>", end=' ')
- print(", ".join([x.split("/")[0] for x,y,z in list(mutEnzymes.keys())]))
- print("</i></div>")
- scriptName = basename(__file__)
- if otData!=None and repCount == 0:
- print((' <a href="%s?batchId=%s&pamId=%s&pam=%s" target="_blank"><strong>Cloning / PCR primers</strong></a>' % (scriptName, batchId, urllib.parse.quote(str(pamId)), pam) ))
- print("</small>")
- print("</td>")
- # off-target score, aka specificity score aka MIT score
- if not pamIsCpf1(pam):
- print("<td>")
- if guideScore==None:
- print("No matches")
- else:
- print("%d" % guideScore)
- print("</td>")
- # guide score based on CFD scores, aka guidescan score
- if "cfdGuideScore" in showColumns:
- print("<td>")
- if guideCfdScore==None:
- print("No matches")
- else:
- print("%d" % guideCfdScore)
- print("</td>")
- # eff scores
- if effScores==None:
- print('<td colspan="%d">Too close to end</td>' % len(scoreNames))
- htmlHelp("The efficiency scores require some flanking sequence<br>This guide does not have enough flanking sequence in your input sequence and could not be extended as it was not found in the genome.<br>")
- else:
- for scoreName in scoreNames:
- # out-of-frame and prox. gc need special treatment
- if scoreName in ["oof", "proxGc"]:
- continue
- score = effScores.get(scoreName, None)
- if score!=None:
- effScoresCount += 1
- if score==None:
- print('''<td>--</td>''')
- elif scoreName=="ssc":
- # save some space
- numStr = '%.1f' % (float(score))
- print('''<td style="font-size:small">%s</td>''' % numStr)
- elif scoreDigits.get(scoreName, 0)==0:
- print('''<td>%d</td>''' % int(score))
- else:
- print('''<td>%0.1f</td>''' % (float(score)))
- #print "<!-- %s -->" % seq30Mer
- if showProxGcCol:
- print("<td>")
- # close GC > 4
- finalGc = int(effScores.get("finalGc6", -1))
- if finalGc==1:
- print("+")
- elif finalGc==0:
- print("-")
- else:
- print("--")
- # main motif is "NGG" and last nucleotides are GGNGG
- if int(effScores.get("finalGg", 0))==1:
- print("<br>")
- print("<small>-GG</small>")
- print("</td>")
- if not baseEditor:
- for mutScoreName in mutScoreNames:
- print("<td>")
- oofScore = str(effScores.get(mutScoreName, None))
- if mutScoreName=="oof":
- scoreDesc = "out-of-frame deletions"
- else:
- scoreDesc = "frameshift mutations"
- if oofScore==None or oofScore=="None":
- print("--")
- else:
- print("""<a href="%s?batchId=%s&pamId=%s&showMh=%s" target=_blank class="tooltipster" title="This score indicates how likely %s are. Click to show the induced deletions based on the micro-homology around the cleavage site.">%s</a>""" % (myName, batchId, urllib.parse.quote(pamId), mutScoreName, scoreDesc, oofScore))
- #print """<br><br><small><a href="%s?batchId=%s&pamId=%s&showMh=1" target=_blank class="tooltipster">Micro-homology</a></small>""" % (myName, batchId, pamId)
- print("</td>")
- # mismatch description
- print("<td>")
- #otCount = sum([int(x) for x in otDesc.split("/")])
- if otData==None:
- # no genome match
- print(otDesc)
- htmlHelp("This exact sequence was not found in the genome.<br>If you have pasted a cDNA multi-exon sequence, note that sequences that overlap a splice site cannot be used as guide sequences. If you only have a cDNA sequence, please BLAST or BLAT your sequence first against the genome, then use the resulting exon from the genome for CRISPOR.<br>This warning also appears if you have selected the wrong or no genome.")
- elif repCount > 0:
- print ("Repeat")
- htmlHelp("At <= 4 mismatches, %d alignments were found in the genome for this sequence, without looking at the PAM sequence around these alignments.<br>This guide is a repeated region, it is too unspecific.<br>Usually, CRISPR cannot be used to target repeats. Also, note that sequences that include long repeats will make the CRISPOR website slow. You can mask repeats with Ns to speed up the search." % repCount)
- else:
- print(otDesc)
- print("<br>")
- # mismatch description, last 12 bp
- print('<small style="color:grey">'+last12Desc+"</small><br>")
- otCount = len(otData)
- print("<br><small>%d off-targets</small>" % otCount)
- print("</td>")
- # links to offtargets
- print("<td><small>")
- if otData!=None:
- if len(otData)>500 and len(guideData)>1:
- otData, cutoff = findOtCutoff(otData)
- if cutoff==None:
- print("More than 1000 off-targets, showing only top "+str(len(otData)))
- else:
- print("More than 500 off-targets, showing %d with score >%0.1f " % (len(otData), cutoff))
- htmlHelp("This guide sequence has a high number of off-targets, its use is discouraged.<br>To show all off-targets, paste only the guide sequence into the input sequence box.")
- otLinks = makeOtBrowserLinks(otData, chrom, dbInfo, pamId)
- print("\n".join(otLinks[:3]))
- if len(otLinks)>3:
- cssPamId = pamId.replace("-","minus").replace("+","plus") # +/-: not valid in css
- cssPamId = cssPamId+"More"
- print('<div id="%s" class="otMore" style="display:none; width:100%%">' % cssPamId)
- print("\n".join(otLinks[3:]))
- print('''<a style="float:right;text-decoration:underline" href="%s?batchId=%s&pamId=%s&otPrimers=1" id="%s">''' % (myName, batchId, urllib.parse.quote(pamId), cssPamId))
- print('<strong>Off-target primers</strong></a>')
- print('</div>')
- print('''<a id="%sMoreLink" class="otMoreLink" onclick="showAllOts('%s')">''' % (cssPamId, cssPamId))
- print('show all...</a>')
- print('''<a id="%sLessLink" class="otLessLink" style="display:none" onclick="showLessOts('%s')">''' % (cssPamId, cssPamId))
- print('show less...</a>')
- print("</small></td>")
- print("</tr>")
- count = count+1
- print("</table>")
- printDownloadTableLinks(batchId, addTsv=True)
- printNoEffScoreFoundWarn(effScoresCount, pam)
- def linkLocalFiles(listFname):
- """ write a <link> statement for each filename in listFname. Version them via mtime
- (-> browser cache)
- """
- for fname in open(listFname).read().splitlines():
- fname = fname.strip()
- if not isfile(fname):
- fname = join(HTMLDIR, fname)
- if not isfile(fname):
- print("missing: %s<br>" % fname)
- continue
- mTime = str(os.path.getmtime(fname)).split(".")[0] # seconds is enough
- if fname.endswith(".css"):
- #url = fname.replace("/var/www/", "http://tefor.net/")
- print("<link rel='stylesheet' media='screen' type='text/css' href='%s%s?%s'/>" % (HTMLPREFIX, fname, mTime))
- def printHeader(batchId, title):
- " print the html header "
- print('''<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">''')
- print("<html><head>")
- if title==None:
- if batchName!="":
- print("""<title>CRISPOR - %s</title>""" % batchName)
- else:
- print("""<title>CRISPOR</title>""")
- else:
- print("""<title>%s</title>""" % title)
- print("""
- <meta name='description' content='Design CRISPR guides with off-target and efficiency predictions, for more than 100 genomes.'/>
- <meta http-equiv='Content-Type' content='text/html; charset=utf-8' />
- <meta property='fb:admins' content='692090743' />
- <meta name="google-site-verification" content="OV5GRHyp-xVaCc76rbCuFj-CIizy2Es0K3nN9FbIBig" />
- <meta property='og:type' content='website' />
- <meta property='og:url' content='http://crispor.gi.ucsc.edu/' />
- <meta property='og:image' content='http://crispor.gi.ucsc.edu/image/CRISPOR.png' />
- <script src="https://cdn.jsdelivr.net/npm/clipboard@2/dist/clipboard.min.js"></script>
- """)
- # load jquery from local copy, not from CDN, for offline use
- print(("""<script src='%sjs/jquery.min.js'></script>
- <script src='%sjs/jquery-ui.min.js'></script>
- """ % (HTMLPREFIX, HTMLPREFIX)))
- #print('<link rel="stylesheet" href="//fonts.googleapis.com/css?family=Roboto:300,300italic,700,700italic" />')
- #print('<link rel="stylesheet" type="text/css" href="https://cdnjs.cloudflare.com/ajax/libs/normalize/5.0.0/normalize.min.css" />')
- #print('<link rel="stylesheet" href="//cdn.rawgit.com/milligram/milligram/master/dist/milligram.min.css">')
- linkLocalFiles("includes.txt")
- print('<link rel="stylesheet" type="text/css" href="%sstyle/tooltipster.css" />' % HTMLPREFIX)
- print('<link rel="stylesheet" type="text/css" href="%sstyle/tooltipster-shadow.css" />' % HTMLPREFIX)
- print('<link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/chosen/1.6.2/chosen.css" />')
- print('<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/source-code-pro@2.38.0/source-code-pro.css" />')
- # the UFD combobox, https://code.google.com/p/ufd/wiki/Usage
- # patched to allow mouse wheel
- # https://code.google.com/p/ufd/issues/detail?id=86&q=mouse%20wheel
- print('<script type="text/javascript" src="%sjs/jquery.ui.ufd.js"></script>' % HTMLPREFIX)
- #print '<link rel="stylesheet" type="text/css" href="%sstyle/ufd-base.css" />' % HTMLPREFIX
- print('<link rel="stylesheet" type="text/css" href="%sstyle/plain.css" />' % HTMLPREFIX)
- print('<link rel="stylesheet" type="text/css" href="%sstyle/jquery-ui.css" />' % HTMLPREFIX)
- print('<script type="text/javascript" src="js/jquery.tooltipster.min.js"></script>')
- print('<script src="https://cdnjs.cloudflare.com/ajax/libs/chosen/1.6.2/chosen.jquery.min.js"></script>')
- # override the main TEFOR css
- print("""
- <style>
- select { font-size: 80%; }
- body {
- text-align: left;
- /* float: left; */
- }
- p {
- }
- ul {
- -webkit-margin-before: 0;
- -webkit-margin-after: 0;
- }
- .editTable {
- border: 1px solid black;
- background-color: white;
- }
- mut {
- color: blue;
- background-color: yellow;
- }
- .editTable th {
- background-color: #F0F0F0;
- }
- tt { font-size: 90% }
- div.contentcentral { text-align: left; float: left}
- /* for chosen.js */
- .chosen-container { width: 600px }
- .chosen-container .chosen-results li.active-result { float: left}
- </style>
- """)
- # activate tooltipster
- #theme: 'tooltipster-shadow',
- # activate jqueryUI tooltips
- print ("""
- <script>
- $(function () {
- $(".tooltip").tooltip({
- relative : true,
- tooltipClass : "alignStyle",
- content: function () {
- return '<div style="width:300px">'+$(this).prop('title')+"</div>";
- }
- });
- });
- $(function () {
- $(".tooltipAuto").tooltip({
- contentAsHtml : true
- });
- });
- </script>""")
- # style of Jquery UI tooltips, default style is div.ui-tooltip
- print("""<style>
- .alignStyle {
- background-color: #FFFFFF;
- width: 350px;
- max-width: 400px;
- height: 110px;
- position : absolute;
- text-align: left;
- border:1px solid #cccccc;
- }
- </style>""")
- # style from https://css-tricks.com/rotated-table-column-headers/ to rotate table headers
- print("""<style>
- th.rotate {
- /* Something you can count on */
- /* height: 10px; */
- white-space: nowrap;
- }
- th.rotate > div {
- float:left;
- white-space: nowrap;
- position: relative;
- border-style: none;
- """)
- # if we're showing all scores, we have very little space, so turn by
- # 90degrees. otherwise, we can afford 45 degrees, which is easier to read.
- #if cgiParams.get("showAllScores", "0")=="1":
- print("""
- -webkit-transform: rotate(-90);
- -moz-transform: rotate(270deg);
- -ms-transform: rotate(270deg);
- -o-transform: rotate(270deg);
- transform: rotate(270deg);
- width: 25px;""")
- #else:
- #print("""
- #-webkit-transform: rotate(-45deg);
- #-moz-transform: rotate(315deg);
- #-ms-transform: rotate(315deg);
- #-o-transform: rotate(315deg);
- #transform: rotate(315deg);
- #width: 25px;""")
- print("""
- }
- th.rotate > div > span {
- /* border-bottom: 1px solid #ccc; */
- padding: 0px 3px;
- white-space: nowrap;
- }
- </style>""")
- print("</head>")
- print('<body id="wrapper">')
- def firstFreeLine(lineMasks, y, start, end):
- " recursively search for first free line to place a feature (start, end) "
- #print "first free line called with y", y, "<br>"
- if y>=len(lineMasks):
- return None
- lineMask = lineMasks[y]
- for x in range(start, end):
- #print "checking pos", x, "<br>"
- if lineMask[x]!=0:
- return firstFreeLine(lineMasks, y+1, start, end)
- return y
- #return None
- def distrOnLines(seq, startDict, featLen, pam):
- """ given a dict with start -> (start,end,name,strand) and a motif len, create lines of annotations such that
- the motifs don't overlap on the lines
- """
- # max number of lines in y direction to draw
- MAXLINES = 18
- # amount of free space around each feature
- SLOP = 2
- # bitmask, one per line, 1 = we have a feature here, 0 = no feature here
- lineMasks = []
- for i in range(0, MAXLINES):
- lineMasks.append( [0]* (len(seq)+10) )
- # dict with lineCount (0...MAXLINES) -> list of (start, strand) tuples
- ftsByLine = defaultdict(list)
- maxY = 0
- for start in sorted(startDict):
- end = start+featLen
- strand = startDict[start]
- # Cannot use Unicode here: these symbols are not part of the
- # monospace font on some platforms and therefore their width
- # is not the same as the other characters
- #arrNE = u'\u2197'
- #arrSE = u'\u2198'
- arrNE = '/'
- arrSE = '\\'
- #arrNE = u'\u2a3c' # hebrew
- #arrSE = u'\ufb27' # math
- ftSeq = seq[start:end]
- if strand=="+":
- if pamIsFirst:
- if pamIsCas12max(pam): #modify cleavage sites to 14-16 (target strand) and 24nt (non-target strand) / hfCas12Max
- label = '%s'%(ftSeq)+'.............%s%s%s.......%s' % (arrNE, arrNE, arrNE, arrSE)
- else:
- label = '%s'%(ftSeq)+'.................%s....%s' % (arrNE, arrSE)
- startFt = start
- endFt = start+len(label)
- else:
- #label = '%s..%s'%(seq[start-3].lower(), ftSeq)
- #label = '---%s'%(ftSeq)
- #label = '---%s'%(ftSeq)
- label = '−−−%s'%(ftSeq)
- startFt = start - 3
- endFt = end
- else:
- if pamIsFirst:
- if pamIsCas12max(pam): #modify cleavage sites to 14-16 (target strand) and 24nt (non-target strand) / hfCas12Max
- spc1 = "......."
- spc2 = "............."
- labelPrefix = '%s%s%s%s%s%s' % (arrSE, spc1, arrNE, arrNE, arrNE, spc2)
- label = labelPrefix + ftSeq
- else:
- spc1 = "...."
- spc2 = "................."
- labelPrefix = '%s%s%s%s' % (arrSE, spc1, arrNE, spc2)
- label = labelPrefix + ftSeq
- startFt = start - len(labelPrefix)
- endFt = startFt+len(label)
- else:
- #label = '%s..%s'%(ftSeq, seq[end+2].lower())
- label = '%s---'%(ftSeq)
- startFt = start
- endFt = end + 3
- #print "feature", strand, start, startFt, endFt, SLOP,"<br>"
- #print "mask", lineMasks[0][startFt:endFt], "<br>"
- y = firstFreeLine(lineMasks, 0, startFt, endFt)
- #print "free line: %s<br>" % y
- if y==None:
- errAbort("not enough space to plot features")
- # fill the current mask
- mask = lineMasks[y]
- maskStart = max(startFt-SLOP, 0)
- maskEnd = min(endFt+SLOP, len(seq))
- #print "mask:", maskStart, maskEnd
- for i in range(maskStart, maskEnd):
- mask[i]=1
- maxY = max(y, maxY)
- pamId = "s%d%s" % (start, strand)
- ft = (startFt, endFt, label, strand, pamId)
- #print "labelLen: %d<br>" % len(label)
- #print "ft: %s<br>" % repr(ft)
- ftsByLine[y].append(ft )
- return ftsByLine, maxY
- def writePamFlank(seq, startDict, pam, faFname):
- " write pam flanking sequences to fasta file, optionally with versions where each nucl is removed "
- #print "writing pams to %s<br>" % faFname
- faFh = open(faFname, "w")
- for pamId, pamStart, guideStart, strand, flankSeq, pamSeq, pamPlusSeq in flankSeqIter(seq, startDict, len(pam), True):
- faFh.write(">%s\n%s\n" % (pamId, flankSeq))
- faFh.close()
- def runCmd(cmd, ignoreExitCode=False, useShell=True):
- " run shell command, check ret code, replaces BIN and SCRIPTS special variables "
- if useShell:
- cmd = cmd.replace("$BIN", binDir)
- cmd = cmd.replace("$PYTHON", sys.executable)
- cmd = cmd.replace("$SCRIPT", scriptDir)
- cmd = "set -o pipefail; " + cmd
- executable = "/bin/bash"
- else:
- cmd = [x.replace("$BIN", binDir).replace("$PYTHON", sys.executable).replace("$SCRIPT", scriptDir) for x in cmd]
- executable=None
- debug("Running %s" % cmd)
- ret = subprocess.call(cmd, shell=useShell, executable=executable)
- if ret!=0 and not ignoreExitCode:
- if not useShell:
- cmd = " ".join(cmd)
- if commandLineMode:
- logging.error("Error: could not run command %s." % cmd)
- sys.exit(1)
- else:
- print("Server error: could not run command %s, error %d.<p>" % (cmd, ret))
- print("please send us an email, we will fix this error as quickly as possible. %s " % contactEmail)
- raise
- sys.exit(0)
- def isAltChrom(chrom):
- """ return true is chrom name looks like it's not on the primary assembly. This is mostly relevant for hg38.
- examples: chr6_*_alt (hg38)
- """
- return chrom.endswith("_alt")
- def parseOfftargets(db, batchId, onTargetChrom=""):
- """ parse a bed file with annotataed off target matches from overlapSelect,
- has two name fields, one with the pam position/strand and one with the
- overlapped segment
- return as dict pamId -> editDist -> (chrom, start, end, seq, strand, segType, segName, totalAlnCount, isRep)
- segType is "ex" "int" or "ig" (=intergenic)
- if intergenic, geneNameStr is two genes, split by |
- The isRep flag is true if BWA reported more than one alignment with X0+X1 but didn't report these with the
- XA tag. It means that we can't get the alignments for this sequence from BWA (=repeats).
- """
- # edge case: target is on chr6_alt -> we remove a single off-target with 0 mismatches on chr6
- targetIsAlt = isAltChrom(onTargetChrom)
- # keep track of pamIds already handled for this edge case
- skippedPams = set()
- # ideally we would check if the offtarget falls into the chrom area that gave rise to the alt
- # but that would mean parsing yet another non-small file and this case should be sufficiently rare
- # to not bother 99% of users.
- batchBase = join(batchDir, batchId)
- bedFname = batchBase+".bed.gz"
- # example input:
- # chrIV 9864393 9864410 s41-|-|5|ACTTGACTG|0 chrIV 9864303 9864408 ex:K07F5.16
- # chrIV 9864393 9864410 s41-|-|5|ACTGTAGCTAGCT|9999 chrIV 9864408 9864470 in:K07F5.16
- debug("reading offtargets from %s" % bedFname)
- # first sort into dict (pamId,chrom,start,end,editDist,strand)
- # -> (segType, segName)
- pamData = {}
- #ifh = open(bedFname) # switched to gzip compression in Dec 2018, converted old files with bash script
- try:
- ifh = gzip.open(bedFname, "rt")
- except FileNotFoundError:
- print("Off-target results from this link were temporarily removed to save space. ")
- linkUrl = "crispor.py?batchId="+batchId
- print("Please go back to <a href='%s'>your job page</a>, which will rerun the job, then try the link again." % linkUrl)
- exit(0)
- maxOtLines = 500000
- count = 0
- for line in ifh:
- fields = line.rstrip("\n").split("\t")
- count +=1
- if count > maxOtLines:
- print("Error: More than %d off-targets. CRISPR has trouble with handling extremely unspecific inputs. Please email us at "
- "%s and discuss. Your input sequence most likely includes a recent L1HS, SVA or similar repetitive elements. "
- "It is hard to design guides for these, you can try to re-run your input sequence, but without the repetitive element. "%
- (maxOtLines, contactEmail))
- sys.exit(1)
- chrom, start, end, name, segment = fields
- logging.debug("off-target: %s" % name)
- # hg38: ignore alternate chromosomes otherwise the
- # regions on the main chroms look as if they could not be
- # targeted at all with Cas9
- if isAltChrom(chrom):
- logging.debug("skipping off-target: on alt-chromosome")
- continue
- nameFields = name.split("|")
- pamId, strand, editDist, seq = nameFields[:4]
- #print pamId, strand, editDist, seq, chrom, start, end, name, segment, "<br>"
- if targetIsAlt:
- if editDist=='0' and onTargetChrom.split("_")[0]==chrom.split("_")[0] and not pamId in skippedPams:
- logging.debug("altChrom edge case: target is on alt-chrom, skipping a single 0-mismatch off-target on primary chrom")
- skippedPams.add(pamId)
- continue
- isRep = 0
- totalAlnCount = 0
- # for compatibility with old bed files, only parse these fields if they are present
- # note: for some reason, the MIT hitScore was always written to these files
- # However, it's not parsed here and never was. In order to not break the old files
- # I kept it in the files, but am not reading it here
- # these are the different formats until now:
- # seqId+"|"+strand+"|"+editDist+"|"+seq+"|"+str(hitScore) # five fields
- # seqId+"|"+strand+"|"+editDist+"|"+seq+"|"+x1Score+"|"+str(hitScore) # six fields
- # (x1Score was roughly the alnCount, similar enough for practical purposes, fixed in 2019)
- # seqId+"|"+strand+"|"+editDist+"|"+seq+"|"+alnCount+"|"+str(hitScore)+"|"+isRep # seven fields
- if len(nameFields)>5:
- totalAlnCount = int(nameFields[4])
- if len(nameFields)>6:
- isRep = bool(int(nameFields[6]))
- editDist = int(editDist)
- # some gene models include colons
- if ":" in segment:
- segType, segName = segment.split(":", maxsplit=1)
- else:
- segType, segName = "", segment
- start, end = int(start), int(end)
- otKey = (pamId, chrom, start, end, editDist, seq, strand, totalAlnCount, isRep)
- # if an offtarget is in the PAR region, we keep only the chrY off-target
- parNum = isInPar(db, chrom, start, end)
- # keep only matches on chrX
- if parNum is not None and chrom=="chrX":
- logging.debug("off-target on PAR region, skipping")
- continue
- # if a offtarget overlaps an intron/exon or ig/exon boundary it will
- # appear twice; in this case, we only keep the exon offtarget
- if otKey in pamData and segType!="ex":
- logging.debug("skipping off-target: ex/ig boundary edge case")
- continue
- pamData[otKey] = (segType, segName)
- # index by pamId and edit distance
- indexedOts = defaultdict(dict)
- for otKey, otVal in pamData.items():
- pamId, chrom, start, end, editDist, seq, strand, totalAlnCount, isRep = otKey
- segType, segName = otVal
- otTuple = (chrom, start, end, seq, strand, segType, segName, totalAlnCount, isRep)
- indexedOts[pamId].setdefault(editDist, []).append( otTuple )
- return indexedOts
- class ConsQueue:
- """ a pseudo job queue that does nothing but report progress to the console """
- def startStep(self, batchId, desc, label):
- logging.info("Progress %s - %s - %s" % (batchId, desc, label))
- def annotateBedWithPos(inBed, outBed, genome):
- """
- given an input bed4 and an output bed filename, add an additional column 5 to the bed file
- that is a descriptive text of the chromosome pos (e.g. chr1:1.23 Mbp).
- """
- ofh = gzip.open(outBed, "wt")
- for line in open(inBed):
- chrom, start = line.split("\t")[:2]
- chrom = applyChromAlias(genome, chrom)
- start = int(start)
- if start>1000000:
- startStr = "%.2f Mbp" % (float(start)/1000000)
- else:
- startStr = "%.2f Kbp" % (float(start)/1000)
- desc = "%s %s" % (chrom, startStr)
- ofh.write(line.rstrip("\n"))
- ofh.write("\t")
- ofh.write(desc)
- ofh.write("\n")
- ofh.close()
- def findAllGuides(seq, pam):
- startDict, endSet = findAllPams(seq, pam)
- pamInfo = list(flankSeqIter(seq, startDict, len(pam), False))
- return pamInfo
- def extractMutScores(scoreDict, pamIds):
- " make a list of the guide-related outcome scores in the order of pamIds "
- res = []
- for pamId in pamIds:
- res.append(scoreDict[pamId][0])
- return res
- def calcSaveEffScores(batchId, seq, extSeq, pam, queue):
- """ given a sequence and an extended sequence, get all potential guides
- with pam, extend them to 100mers and score them with various eff. scores.
- Return a
- list of rows [headers, (guideSeq, 100mer, score1, score2, score3,...), ... ]
- Also write the results to a database so they can be retrieved later.
- extSeq can be None, if we were unable to extend the sequence
- """
- seq = seq.upper()
- if extSeq:
- extSeq = extSeq.upper()
- pamInfo = findAllGuides(seq, pam)
- pamIds = []
- guides = []
- longSeqs = []
- for pamId, startPos, guideStart, strand, guideSeq, pamSeq, pamPlusSeq in pamInfo:
- logging.debug("PAM ID: %s - guideSeq %s" % (pamId, guideSeq))
- gStart, gEnd = pamStartToGuideRange(startPos, strand, len(pam))
- longSeq = getExtSeq(seq, gStart, gEnd, strand, 50-GUIDELEN, 50, extSeq) # +-50 bp from the end of the guide
- if longSeq!=None:
- longSeqs.append(longSeq)
- pamIds.append(pamId)
- guides.append(guideSeq+pamSeq)
- if len(longSeqs)>0 and doEffScoring:
- enz = None
- if pamIsCpf1(pam) and not pam=="NGTN":
- enz = "cpf1"
- elif pamIsSaCas9(pam):
- enz = "sacas9"
- # for spcas9, we use the extended list for the calculation
- global scoreNames
- if enz is None:
- scoreNames = allScoreNames
- effScores = crisporEffScores.calcAllScores(longSeqs, enzyme=enz, scoreNames=scoreNames)
- # these are slow algorithms, so store the results for later
- queue.startStep(batchId, "outcome", "Calculating editing outcomes")
- mutScores = crisporEffScores.calcMutSeqs(pamIds, longSeqs, enz, scoreNames=mutScoreNames)
- saveOutcomeData(batchId, mutScores)
- # for output and sorting, it's easier to treat the outcome-derived scores like an efficiency score
- for mutScoreName in mutScoreNames:
- if mutScoreName in mutScores:
- effScores[mutScoreName] = extractMutScores(mutScores[mutScoreName], pamIds)
- # make sure the "N bug" reported by Alberto does never happen again:
- # we must get back as many scores as we have sequences
- for scoreName, scores in effScores.items():
- if len(scores)!=len(longSeqs):
- print("Internal error when calculating score %s" % scoreName)
- assert(False)
- else:
- effScores = {}
- activeScoreNames = list(effScores.keys())
- # reformat to rows, write all scores to file
- assert(len(pamIds)==len(guides)==len(longSeqs))
- rows = []
- for i, (guideId, guide, longSeq) in enumerate(zip(pamIds, guides, longSeqs)):
- row = [guideId, guide, longSeq]
- for scoreName in activeScoreNames:
- scoreList = effScores[scoreName]
- if len(scoreList) > 0:
- row.append(scoreList[i])
- else:
- row.append("noScore?")
- rows.append(row)
- headerRow = ["guideId", "guide", "longSeq"]
- headerRow.extend(activeScoreNames)
- rows.insert(0, headerRow)
- return rows
- def writeRow(ofh, row):
- " write list to file as tab-sep row "
- row = [str(x) for x in row]
- ofh.write("\t".join(row))
- ofh.write("\n")
- def createBatchEffScoreTable(batchId, queue):
- """ annotate all potential guides with efficiency scores and write to file.
- tab-sep file for easier debugging, no pickling
- """
- outFname = join(batchDir, batchId+".effScores.tab")
- # Todo: why don't we get these from the caller as arguments instead of reading the batch?
- batchInfo = readBatchAsDict(batchId)
- seq = batchInfo["seq"]
- extSeq = batchInfo.get("extSeq") # cannot always extend a sequence, e.g. when no perfect match
- pam = batchInfo["pam"]
- pam = setupPamInfo(pam)
- seq = seq.upper()
- if extSeq:
- extSeq = extSeq.upper()
- guideRows = calcSaveEffScores(batchId, seq, extSeq, pam, queue)
- guideFh = open(outFname, "w")
- for row in guideRows:
- writeRow(guideFh, row)
- guideFh.close()
- logging.info("Wrote eff scores to %s" % guideFh.name)
- def readEffScores(batchId):
- " parse eff scores from tab sep file and return as dict pamId -> dict of scoreName -> value "
- effScoreFname = join(batchDir, batchId)+".effScores.tab"
- seqToScores = {}
- if isfile(effScoreFname):
- for row in lineFileNext(open(effScoreFname)):
- scoreDict = {}
- rowDict = row._asdict()
- # the first three fields are the pamId, shortSeq, longSeq, they are not scores
- allScoreNames = row._fields[3:]
- for scoreName in allScoreNames:
- score = rowDict[scoreName]
- if score=="None":
- score = "NA"
- elif "." in score or "e" in score:
- score = float(score)
- else:
- score = int(score)
- scoreDict[scoreName] = score
- seqToScores[row.guideId] = scoreDict
- return seqToScores
- def findOfftargetsBwa(queue, batchId, batchBase, faFname, genome, pamDesc, bedFname):
- " align faFname to genome and create matchedBedFname "
- matchesBedFname = batchBase+".matches.bed"
- saFname = batchBase+".sa"
- pam = setupPamInfo(pamDesc)
- pamLen = len(pam)
- genomeDir = genomesDir # make var local, see below
- open(matchesBedFname, "w") # truncate to 0 size
- # increase MAXOCC if there is only a single query, but only in CGI mode
- #if len(parseFasta(open(faFname)))==1 and not commandLineMode:
- #global MAXOCC
- #global maxMMs
- #MAXOCC=max(HIGH_MAXOCC, MAXOCC)
- #maxMMs=HIGH_maxMMs
- maxDiff = maxMMs
- queue.startStep(batchId, "bwa", "Alignment of potential guides, mismatches <= %d" % maxDiff)
- convertMsg = "Converting alignments"
- seqLen = GUIDELEN
- bwaM = MFAC*MAXOCC # -m is queue size in bwa
- cmd = "$BIN/bwa aln -o 0 -m %(bwaM)s -n %(maxDiff)d -k %(maxDiff)d -N -l %(seqLen)d %(genomeDir)s/%(genome)s/%(genome)s.fa %(faFname)s > %(saFname)s" % locals()
- runCmd(cmd)
- queue.startStep(batchId, "saiToBed", convertMsg)
- maxOcc = MAXOCC # create local var from global
- # EXTRACTION OF POSITIONS + CONVERSION + SORT/CLIP
- # the sorting should improve the twoBitToFa runtime
- python = sys.executable
- cmd = "$BIN/bwa samse -n %(maxOcc)d %(genomeDir)s/%(genome)s/%(genome)s.fa %(saFname)s %(faFname)s | $SCRIPT/xa2multi.pl | %(python)s $SCRIPT/samToBed %(pam)s %(seqLen)d | sort -k1,1 -k2,2n | $BIN/bedClip stdin %(genomeDir)s/%(genome)s/%(genome)s.sizes stdout >> %(matchesBedFname)s " % locals()
- runCmd(cmd)
- filtMatchesBedFname = batchBase+".filtMatches.bed"
- queue.startStep(batchId, "filter", "Removing matches without a PAM motif")
- altPats = ",".join(offtargetPams.get(pam, ["na"]))
- bedFnameTmp = bedFname+".tmp"
- altPamMinScore = str(ALTPAMMINSCORE)
- shmFaFname = join("/dev/shm", genome+".fa")
- # EXTRACTION OF SEQUENCES + ANNOTATION - big headache!!
- # twoBitToFa was 15x slower than python's twobitreader, after markd's fix it is better
- # but bedtools uses an fa.idx file and also mmap, so is a LOT faster
- # arguments: guideSeq, mainPat, altPats, altScore, passTotalAlnCount
- if isfile(shmFaFname):
- logging.info("Using bedtools and genome fasta on ramdisk, %s" % shmFaFname)
- cmd = "time bedtools getfasta -s -name -fi %(shmFaFname)s -bed %(matchesBedFname)s -fo /dev/stdout | $SCRIPT/filterFaToBed %(faFname)s %(pam)s %(altPats)s %(altPamMinScore)s > %(filtMatchesBedFname)s" % locals()
- else:
- cmd = "time $BIN/twoBitToFa %(genomeDir)s/%(genome)s/%(genome)s.2bit stdout -bed=%(matchesBedFname)s | %(python)s $SCRIPT/filterFaToBed %(faFname)s %(pam)s %(altPats)s %(altPamMinScore)s > %(filtMatchesBedFname)s" % locals()
- #cmd = "$SCRIPT/twoBitToFaPython %(genomeDir)s/%(genome)s/%(genome)s.2bit %(matchesBedFname)s | $SCRIPT/filterFaToBed %(faFname)s %(pam)s %(altPats)s %(altPamMinScore)s %(maxOcc)d > %(filtMatchesBedFname)s" % locals()
- runCmd(cmd)
- segFname = "%(genomeDir)s/%(genome)s/%(genome)s.segments.bed" % locals()
- # if we have gene model segments, annotate them, otherwise just use the chrom position
- if isfile(segFname):
- queue.startStep(batchId, "genes", "Annotating matches with genes")
- cmd = "cat %(filtMatchesBedFname)s | $BIN/overlapSelect %(segFname)s stdin stdout -mergeOutput -selectFmt=bed -inFmt=bed | cut -f1,2,3,4,8 | gzip > %(bedFnameTmp)s " % locals()
- runCmd(cmd)
- else:
- queue.startStep(batchId, "chromPos", "Annotating matches with chromosome position")
- annotateBedWithPos(filtMatchesBedFname, bedFnameTmp, genome)
- # make sure the final bed file is never in a half-written state,
- # as it is our signal that the job is complete
- shutil.move(bedFnameTmp, bedFname)
- queue.startStep(batchId, "done", "Job completed")
- # remove the temporary files
- tempFnames = [saFname, matchesBedFname, filtMatchesBedFname]
- if not DEBUG:
- for tfn in tempFnames:
- if isfile(tfn):
- os.remove(tfn)
- return bedFname
- def makeVariants(seq):
- " generate all possible variants of sequence at 1bp-distance"
- seqs = []
- for i in range(0, len(seq)):
- for l in "ACTG":
- if l==seq[i]:
- continue
- newSeq = seq[:i]+l+seq[i+1:]
- seqs.append((i, seq[i], l, newSeq))
- return seqs
- def expandIupac(seq):
- """ expand all IUPAC characters to nucleotides, returns list.
- >>> expandIupac("NY")
- ['GC', 'GT', 'AC', 'AT', 'TC', 'TT', 'CC', 'CT']
- """
- # http://stackoverflow.com/questions/27551921/how-to-extend-ambiguous-dna-sequence
- d = {'A': 'A', 'C': 'C', 'B': 'CGT', 'D': 'AGT', 'G': 'G', \
- 'H': 'ACT', 'K': 'GT', 'M': 'AC', 'N': 'GATC', 'S': 'CG', \
- 'R': 'AG', 'T': 'T', 'W': 'AT', 'V': 'ACG', 'Y': 'CT', 'X': 'GATC'}
- seqs = []
- for i in product(*[d[j] for j in seq]):
- seqs.append("".join(i))
- return seqs
- def writeBowtieSequences(inFaFname, outFname, pamPat):
- """ write the sequence and one-bp-distant-sequences + all possible PAM sequences to outFname
- Return dict querySeqId -> querySeq and a list of all
- possible PAMs, as nucleotide sequences (not IUPAC-patterns)
- """
- ofh = open(outFname, "w")
- outCount = 0
- inCount = 0
- guideSeqs = {} # 20mer guide sequences
- qSeqs = {} # 23mer query sequences for bowtie, produced by expanding guide sequences
- allPamSeqs = expandIupac(pamPat)
- for seqId, seq in parseFastaAsList(open(inFaFname)):
- inCount += 1
- guideSeqs[seqId] = seq
- for pamSeq in allPamSeqs:
- # the input sequence + the PAM
- newSeqId = "%s.%s" % (seqId, pamSeq)
- newFullSeq = seq+pamSeq
- ofh.write(">%s\n%s\n" % (newSeqId, newFullSeq))
- qSeqs[newSeqId] = newFullSeq
- # all one-bp mutations of the input sequence + the PAM
- for nPos, fromNucl, toNucl, newSeq in makeVariants(seq):
- newSeqId = "%s.%s.%d:%s>%s" % (seqId, pamSeq, nPos, fromNucl, toNucl)
- newFullSeq = newSeq+pamSeq
- ofh.write(">%s\n%s\n" % (newSeqId, newFullSeq))
- qSeqs[newSeqId] = newFullSeq
- outCount += 1
- ofh.close()
- logging.debug("Wrote %d variants+expandedPam of %d sequences to %s" % (outCount, inCount, outFname))
- return guideSeqs, qSeqs, allPamSeqs
- def applyModifStr(seq, modifStrs, strand):
- """ bowtie: given a list of pos:toNucl>fromNucl and a seq, return the original seq.
- position is 0-based
- >>> applyModifStr("ACAATAAGACATAAACATATCGG", "14:T>A,21:A>G,22:C>G".split(","), "+")
- 'ACAATAAGACATAATCATATCAC'
- """
- seq = list(seq)
- for modifStr in modifStrs:
- #logging.debug( modifStr)
- pos, toFromNucl = modifStr.split(":")
- fromNucl, toNucl = toFromNucl.split(">")
- pos = int(pos)
- if strand=="-":
- fromNucl = revComp(fromNucl)
- seq[pos] = fromNucl
- return "".join(seq)
- def parseRefout(tmpDir, guideSeqs, pamLen):
- """ parse all .map file in tmpDir and return as list of chrom,start,end,strand,guideSeq,tSeq
- """
- fnames = glob.glob(join(tmpDir, "*.map"))
- # while parsing, make sure we keep only the hit with the lowest number of mismatches
- # to the guide. Saves time when parsing.
- posToHit = {}
- hitBestMismCount = {}
- for fname in fnames:
- for line in open(fname):
- # s20+.17:A>G - chr8 26869044 CCAGCACGTGCAAGGCCGGCTTC IIIIIIIIIIIIIIIIIIIIIII 7 4:C>G,13:T>G,15:C>G
- guideIdWithMod, strand, chrom, start, tSeq, weird, someScore, alnModifStr = \
- line.rstrip("\n").split("\t")
- guideId = guideIdWithMod.split(".")[0]
- modifParts = alnModifStr.split(",")
- if modifParts==['']:
- modifParts = []
- mismCount = len(modifParts)
- hitId = (guideId, chrom, start, strand)
- oldMismCount = hitBestMismCount.get(hitId, 9999)
- if mismCount < oldMismCount:
- hit = (mismCount, guideIdWithMod, strand, chrom, start, tSeq, modifParts)
- posToHit[hitId] = hit
- hitBestMismCount[hitId] = mismCount # thanks to github user mbsimonovic
- ret = []
- for guideId, hit in posToHit.items():
- mismCount, guideIdWithMod, strand, chrom, start, tSeq, modifParts = hit
- if strand=="-":
- tSeq = revComp(tSeq)
- guideId = guideIdWithMod.split(".")[0]
- guideSeq = guideSeqs[guideId]
- genomeSeq = applyModifStr(tSeq, modifParts, strand)
- start = int(start)
- bedRow = (guideId, chrom, start, start+GUIDELEN+pamLen, strand, guideSeq, genomeSeq)
- ret.append( bedRow )
- return ret
- def getEditDist(str1, str2):
- """ return edit distance between two strings of equal length
- >>> getEditDist("HIHI", "HAHA")
- 2
- """
- assert(len(str1)==len(str2))
- str1 = str1.upper()
- str2 = str2.upper()
- editDist = 0
- for c1, c2 in zip(str1, str2):
- if c1!=c2:
- editDist +=1
- return editDist
- def findOfftargetsBowtie(queue, batchId, batchBase, faFname, genome, pamPat, bedFname):
- " align guides with pam in faFname to genome and write off-targets to bedFname "
- tmpDir = batchBase+".bowtie.tmp"
- os.mkdir(tmpDir)
- # make sure this directory gets removed, no matter what
- global tmpDirsDelExit
- tmpDirsDelExit.append(tmpDir)
- if not DEBUG:
- atexit.register(delTmpDirs)
- # write out the sequences for bowtie
- queue.startStep(batchId, "seqPrep", "preparing sequences")
- bwFaFname = abspath(join(tmpDir, "bowtieIn.fa"))
- guideSeqs, qSeqs, allPamSeqs = writeBowtieSequences(faFname, bwFaFname, pamPat)
- genomePath = abspath(join(genomesDir, genome, genome))
- oldCwd = os.getcwd()
- # run bowtie
- queue.startStep(batchId, "bowtie", "aligning with bowtie")
- os.chdir(tmpDir) # bowtie writes to hardcoded output filenames with --refout
- # -v 3 = up to three mismatches
- # -y = try hard
- # -t = print time it took
- # -k = output up to X alignments
- # -m = do not output any hit if a read has more than X hits
- # --max = write all reads that exceed -m to this file
- # --refout = output in bowtie format, not SAM
- # --maxbts=2000 maximum number of backtracks
- # -p 4 = use four threads
- # --mm = use mmap
- maxOcc = MAXOCC # meaning in BWA: includes any PAM, in bowtie we have the PAM in the input sequence
- cmd = "$BIN/bowtie -e 1000 %(genomePath)s -f %(bwFaFname)s -v 3 -y -t -k %(maxOcc)d -m %(maxOcc)d dummy --max tooManyHits.txt --mm --refout --maxbts=2000 -p 4" % locals()
- runCmd(cmd)
- os.chdir(oldCwd)
- queue.startStep(batchId, "parse", "parsing alignments")
- pamLen = len(pamPat)
- hits = parseRefout(tmpDir, guideSeqs, pamLen)
- queue.startStep(batchId, "scoreOts", "scoring off-targets")
- # make the list of alternative PAM sequences
- altPats = offtargetPams.get(pamPat, [])
- altPamSeqs = []
- for altPat in altPats:
- altPamSeqs.extend(expandIupac(altPat))
- # iterate over bowtie hits and write to a BED file with scores
- # if the hit looks OK (right PAM + score is high enough)
- tempBedPath = join(tmpDir, "bowtieHits.bed")
- tempFh = open(tempBedPath, "w")
- offTargets = {}
- isSaCas9 = pamIsSaCas9(pamPat)
- for guideIdWithMod, chrom, start, end, strand, _, tSeq in hits:
- guideId = guideIdWithMod.split(".")[0]
- guideSeq = guideSeqs[guideId]
- genomePamSeq = tSeq[-pamLen:]
- logging.debug( "PAM seq: %s of %s" % (genomePamSeq, tSeq))
- if genomePamSeq in altPamSeqs:
- minScore = ALTPAMMINSCORE
- elif genomePamSeq in allPamSeqs:
- minScore = MINSCORE
- else:
- logging.debug("Skipping off-target for %s: %s:%d-%d" % (guideId, chrom, start, end))
- continue
- logging.debug("off-target minScore = %f" % minScore )
- # check if this match passes the off-target score limit
- if pamIsCpf1(pamPat):
- otScore = 0.0
- else:
- tSeqNoPam = tSeq[:-pamLen]
- if isSaCas9:
- otScore = calcSaHitScore(guideSeq, tSeqNoPam)
- else:
- otScore = calcHitScore(guideSeq, tSeqNoPam)
- if otScore < minScore:
- logging.debug("off-target not accepted")
- continue
- editDist = getEditDist(guideSeq, tSeqNoPam)
- guideHitCount = 0
- guideId = guideId.split(".")[0] # full guide ID looks like s33+.0:A>T
- name = guideId+"|"+strand+"|"+str(editDist)+"|"+tSeq+"|"+str(guideHitCount)+"|"+str(otScore)
- row = [chrom, str(start), str(end), name]
- # this way of collecting the features will remove the duplicates
- otKey = (chrom, start, end, strand, guideId)
- logging.debug("off-target key is %s" % str(otKey))
- offTargets[ otKey ] = row
- for rowKey, row in offTargets.items():
- tempFh.write("\t".join(row))
- tempFh.write("\n")
- tempFh.flush()
- # create a tempfile which is moved over upon success
- # makes sure we do not leave behind a half-written file if
- # we crash later
- tmpFd, tmpAnnotOffsPath = tempfile.mkstemp(dir=tmpDir, prefix="annotOfftargets")
- tmpFh = open(tmpAnnotOffsPath, "w")
- # get name of file with genome locus names
- genomeDir = genomesDir # make local var
- segFname = "%(genomeDir)s/%(genome)s/%(genome)s.segments.bed" % locals()
- # annotate with genome locus names
- cmd = "$BIN/overlapSelect %(segFname)s %(tempBedPath)s stdout -mergeOutput -selectFmt=bed -inFmt=bed | cut -f1,2,3,4,8 > %(tmpAnnotOffsPath)s" % locals()
- runCmd(cmd)
- shutil.move(tmpAnnotOffsPath, bedFname)
- queue.startStep(batchId, "done", "Job completed")
- if DEBUG:
- logging.info("debug mode: Not deleting %s" % tmpDir)
- else:
- shutil.rmtree(tmpDir)
- def processSubmission(faFname, genome, pamDesc, bedFname, batchBase, batchId, queue):
- """ search fasta file against genome, filter for pam matches and write to bedFName
- optionally write status updates to work queue. Remove faFname.
- """
- batchInfo = readBatchAsDict(batchId)
- if genome=="noGenome":
- posStr = "?"
- elif "batchName" in batchInfo and batchInfo["batchName"].count(":")==2: # chrom:start-end:strand
- posStr = batchInfo["batchName"]
- else:
- queue.startStep(batchId, "bwasw", "Searching genome for one 100% identical match to input sequence")
- posStr = findPerfectMatch(batchId)
- batchInfo["posStr"] = posStr
- if posStr!="?":
- # get a 100bp-extended version of the input seq
- chrom, start, end, strand = parsePos(posStr)
- extSeq = extendAndGetSeq(genome, chrom, start, end, strand, batchInfo["seq"])
- if extSeq is None:
- # this can only happen if there is a 100%-M match but small SNPs in it compared to the input sequence
- # so the extension of the input fails.
- # in this case, we also invalidate the position, as there was no perfect match and the user
- # has to do something to fix it
- batchInfo["posStr"] = "?"
- else:
- logging.debug("100pb-extended seq (len: %d) is: %s" % (len(extSeq), extSeq))
- batchInfo["extSeq"] = extSeq
- # must save the batch again, as otherwise display won't work, we need the position saved
- writeBatchAsDict(batchInfo, batchId)
- if doEffScoring:
- queue.startStep(batchId, "effScores", "Calculating guide efficiency scores")
- createBatchEffScoreTable(batchId, queue)
- if genome=="noGenome":
- # skip the off-target search entirely
- open(bedFname, "w") # create a 0-byte file to signal job completion
- queue.startStep(batchId, "done", "Job completed")
- return
- if useBowtie:
- findOfftargetsBowtie(queue, batchId, batchBase, faFname, genome, pamDesc, bedFname)
- else:
- findOfftargetsBwa(queue, batchId, batchBase, faFname, genome, pamDesc, bedFname)
- if not DEBUG:
- os.remove(faFname)
- return bedFname
- def lineFileNext(fh):
- """
- parses tab-sep file with headers as field names
- yields collection.namedtuples
- strips "#"-prefix from header line
- """
- line1 = fh.readline()
- while line1.startswith("##"):
- line1 = fh.readline()
- line1 = line1.strip("\n").strip("#")
- headers = line1.split("\t")
- Record = namedtuple('tsvRec', headers)
- for line in fh:
- line = line.rstrip("\n")
- fields = line.split("\t")
- try:
- rec = Record(*fields)
- except Exception as msg:
- logging.error("Exception occured while parsing line, %s" % msg)
- logging.error("Filename %s" % fh.name)
- logging.error("Line was: %s" % repr(line))
- logging.error("Does number of fields match headers?")
- logging.error("Headers are: %s" % headers)
- #raise Exception("wrong field count in line %s" % line)
- continue
- # convert fields to correct data type
- yield rec
- allGenomes = None
- def readGenomes():
- " return list of all genomes supported "
- global allGenomes
- if allGenomes:
- return allGenomes
- genomes = {}
- myDir = dirname(__file__)
- genomesDir = join(myDir, "genomes")
- inFnames = []
- globalFname = join(genomesDir, "genomeInfo.all.tab")
- if isfile(globalFname):
- inFnames = [globalFname]
- else:
- for subDir in os.listdir(genomesDir):
- infoFname = join(genomesDir, subDir, "genomeInfo.tab")
- if isfile(infoFname):
- inFnames.append(infoFname)
- for infoFname in inFnames:
- for row in lineFileNext(open(infoFname)):
- # add a note to identify UCSC genomes
- if row.server.startswith("ucsc"):
- addStr="UCSC "
- else:
- addStr = ""
- genomes[row.name] = row.scientificName+" - "+row.genome+" - "+addStr+row.description
- genomes = list(genomes.items())
- genomes.sort(key=operator.itemgetter(1))
- allGenomes = genomes
- return allGenomes
- def printOrgDropDown(lastorg, genomes):
- " print the organism drop down box. "
- print('<select id="genomeDropDown" class style="max-width:600px" name="org" tabindex="2">')
- print('<option ')
- if lastorg == "noGenome":
- print('selected ')
- print('value="noGenome">-- No Genome: no specificity, only cleavage efficiency scores (max. len 25kbp)</option>')
- for db, desc in genomes:
- print('<option ')
- if db == lastorg :
- print('selected ')
- print('value="%s">%s</option>' % (db, desc))
- print("</select>")
- #print ('''
- #<script type="text/javascript">
- #$("#genomeDropDown").ufd({maxWidth:350, listWidthFixed:false});
- #</script>''')
- print ('''<br>''')
- print ("""<script>
- $("#genomeDropDown").chosen();
- $(".chosen-choices li").css("background","red");
- </script>
- """)
- def printPamDropDown(lastpam):
- print('<select style="float:left" name="pam" tabindex="3">')
- for key,value in pamDesc:
- print('<option ')
- if key == lastpam :
- print('selected ')
- print('value="%s">%s</option>' % (key, value))
- print("</select>")
- def printForm(params):
- " print html input form "
- scriptName = basename(__file__)
- genomes = readGenomes()
- haveHuman = False
- for g in genomes:
- if g[0]=="hg19":
- haveHuman = True
- # The returned cookie is available in the os.environ dictionary
- cookies=http.cookies.SimpleCookie(os.environ.get('HTTP_COOKIE'))
- if "lastorg" in cookies and "lastseq" in cookies and "lastpam" in cookies:
- lastorg = cookies['lastorg'].value
- lastseq = cookies['lastseq'].value
- lastpam = cookies['lastpam'].value
- else:
- if not haveHuman:
- global DEFAULTSEQ
- global DEFAULTORG
- DEFAULTSEQ = ALTSEQ
- DEFAULTORG = ALTORG
- lastorg = DEFAULTORG
- lastseq = DEFAULTSEQ
- lastpam = DEFAULTPAM
- # SerialCloner is sending us the sequence via a HTTP get parameter
- if "seq" in params:
- lastseq = params["seq"]
- if "org" in params:
- lastorg = params["org"]
- seqName = ""
- if "seqName" in params:
- seqName = params["seqName"]
- printTeforBodyStart()
- #print('''March 6 2023: Sorry, no CRISPOR on the new UCSC-based-server (with RS3 scores) today. Too many performance problems on the new server. We were able to renew the old server. Please use the <a href="http://37.187.154.234">old server</a> temporarily.''')
- #sys.exit(0)
- print("""
- <form id="main-form" method="post" action="%s">
- <div style="text-align:left; margin-left: 10px">
- CRISPOR (<a href="https://academic.oup.com/nar/article/46/W1/W242/4995687">citation</a>) is a program that helps design, evaluate and clone guide sequences for the CRISPR/Cas9 system. <a target=_blank href="/manual/">CRISPOR Manual</a>
- <br><i>July 2025: Added hasCas12Max and e-SpotOn. Also allowing old primer links to Crispor to work again. See <a href="doc/changes.html">Full list of changes</a></i><br>
- </div>
- <div class="windowstep subpanel" style="width:40%%;">
- <div class="substep">
- <div class="title">
- Step 1
- </div>
- Planning a lentiviral gene knockout screen? Use <a href="crispor.py?libDesign=1">CRISPOR Batch</a><br>
- Sequence name (optional): <input type="text" name="name" size="20" value="%s"><br>
- Enter a single genomic sequence, < %d bp, typically an exon
- <img src="%simage/info-small.png" title="CRISPOR conserves the lowercase and uppercase format of your sequence, allowing to highlight sequence features of interest such as ATG or STOP codons.<br>Avoid using cDNA sequences as input, CRISPR guides that straddle splice sites are unlikely to work.<br>You can paste a single >23bp sequence and even multiple sequences, separated by N characters." class="tooltipster">
- <br>
- <small><a href="javascript:clearInput()">Clear Box</a> - </small>
- <small><a href="javascript:resetToExample()">Reset to default</a></small>
- </div>
- <textarea tabindex="1" style="width:100%%" name="seq" rows="12"
- placeholder="Paste here the genomic - not a cDNA - sequence of the exon you want to target. The sequence has to include the PAM site for your enzyme of interest, e.g. NGG. Maximum size %d bp. If you only have a cDNA, please BLAST or BLAT the cDNA first to find the right exon sequence for CRISPOR.">%s</textarea>
- <small>Text case is preserved, e.g. you can mark ATGs with lowercase.<br>Instead of a sequence, you can paste a chromosome range, e.g. chr1:11,130,540-11,130,751</small>
- </div>
- <div class="windowstep subpanel" style="width:50%%">
- <div class="substep" style="margin-bottom: 1px">
- <div class="title" style="cursor:pointer;" onclick="$('#helpstep2').toggle('fast')">
- Step 2
- </div>
- Select a genome
- </div>
- """% (scriptName, seqName, MAXSEQLEN, HTMLPREFIX, MAXSEQLEN, lastseq))
- printOrgDropDown(lastorg, genomes)
- print("""
- <div id="trackHubNote" style="margin-bottom:5px">
- <small>Note: pre-calculated exonic guides for this species are on the <a id='hgTracksLink' target=_blank href="">UCSC Genome Browser</a>.</small>
- </div>
- """)
- print('<small style="float:left">We have %d genomes, but not yours? Search <a href="https://www.ncbi.nlm.nih.gov/assembly">NCBI assembly</a> and send a GCF_/GCA_ ID to <a href="mailto:%s">CRISPOR support</a>.</small>' % (len(genomes), contactEmail))
- print("""
- </div>
- <div class="windowstep subpanel" style="width:50%%; height:158px">
- <div class="substep">
- <div class="title" style="cursor:pointer;" onclick="$('#helpstep3').toggle('fast')">
- Step 3
- <img src="%simage/info-small.png" title="The most common system uses the NGG PAM recognized by Cas9 from S. <i>pyogenes</i>. The VRER and VQR mutants were described by <a href='http://www.nature.com/nature/journal/vaop/ncurrent/abs/nature14592.html' target='_blank'>Kleinstiver et al</a>, Cas9-HF1 by <a href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4851738/'>Kleinstiver 2016</a>, eSpCas1.1 by <a href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4714946/'>Slaymaker 2016</a>, Cpf1 by <a href='http://www.cell.com/abstract/S0092-8674(15)01200-3'>Zetsche 2015</a>, SaCas9 by <a href='https://www.ncbi.nlm.nih.gov/pmc/articles/pmid/25830891/'>Ran 2015</a> and KKH-SaCas9 by <a href='https://www.ncbi.nlm.nih.gov/pmc/articles/pmid/26524662/'>Kleinstiver 2015</a>, modified As-Cpf1s by <a href='http://biorxiv.org/content/early/2016/12/04/091611'>Gao et al. 2017</a>." class="tooltipsterInteract">
- </div>
- Select a Protospacer Adjacent Motif (PAM)
- </div>
- """ % HTMLPREFIX)
- printPamDropDown(lastpam)
- print("""<br>See <a target=_blank href="manual/manual.html#enzymes">notes on enzymes</a> in the manual.<br>""")
- print("""
- <div style="width:40%; margin-top: 10px; margin-left:50px; text-align:center; display:block">
- <input type="submit" name="submit" value="SUBMIT" tabindex="4"/>
- </div>
- </div>
- """)
- print("""
- <script>
- /* set the dropbox to hg19 and paste the example sequence into the input box. */
- function resetToExample() {
- $("textarea[name='seq']").val("%s");
- $("#genomeDropDown").val("%s");
- $("select[name='pam']").val("NGG");
- }
- /* clear the sequence input box */
- function clearInput() {
- $("textarea[name='seq']").val("");
- }
- </script>
- <script>
- /* hide the track hub note if genome is not hg19 */
- ucscTrackDbs=['hg19', 'hg38', 'rn5', 'mm10', 'mm9', 'ci2', 'danRer7', 'sacCer3', 'dm6'];
- function showHideHubNote() {
- var valSel = $("#genomeDropDown").val();
- if (jQuery.inArray(valSel, ucscTrackDbs)!=-1)
- {
- $("#trackHubNote").css('visibility', 'visible');
- $("#hgTracksLink").attr("href", "http://genome.ucsc.edu/cgi-bin/hgTracks?db="+valSel+"&crispr=show");
- }
- else
- $("#trackHubNote").css('visibility', 'hidden');
- }
- $("#genomeDropDown").on('change', showHideHubNote);
- showHideHubNote();
- </script>
- </form>
- """ % (DEFAULTSEQ, DEFAULTORG))
- def readBatchAsDict(batchId):
- " return contents of batch as a dictionary or None "
- batchBase = join(batchDir, batchId)
- jsonFname = batchBase+".json"
- if isfile(jsonFname):
- params = json.load(open(jsonFname))
- else:
- db = sqlite3.connect(batchArchive)
- c = db.cursor()
- c.execute("select data from jobArchive where id=?", (batchId,))
- data = None
- for row in c.fetchall():
- data = row[0]
- db.close()
- if data is None:
- return None
- jsonStr = gzip.decompress(data).decode("utf8")
- params = json.loads(jsonStr)
- if "batchName" in params:
- global batchName
- batchName = params["batchName"]
- return params
- def writeBatchAsDict(batchInfo, batchId):
- batchBase = join(batchDir, batchId)
- tmpFname = batchBase+".json.tmp"
- ofh = open(tmpFname, "w")
- json.dump(batchInfo, ofh)
- ofh.close()
- jsonFname = batchBase+".json"
- os.rename(tmpFname, jsonFname)
- logging.debug("Wrote batch info to %s: %s" % (jsonFname, batchInfo))
- def readBatchParams(batchId):
- """ given a batchId, return the genome, the pam, the input sequence and the
- chrom pos and extSeq, a 100bp-extended version of the input sequence.
- Returns None for pos if not found. """
- params = readBatchAsDict(batchId)
- if params != None:
- return params["seq"], params["org"], params["pam"], params.get("posStr"), params.get("extSeq")
- # FROM HERE UP TO END OF FUNCTION: legacy cold for old batches pre-end-2016 (no json files back then)
- # remove in 2017
- batchBase = join(batchDir, batchId)
- inputFaFname = batchBase+".input.fa"
- if not isfile(inputFaFname):
- errAbort('Could not find the batch %s. We cannot keep Crispor runs for more than '
- 'a few months. Please resubmit your input sequence via'
- ' <a href="crispor.py">the query input form</a>' % batchId)
- ifh = open(inputFaFname, encoding="utf8")
- ifhFields = ifh.readline().replace(">","").strip().split()
- if len(ifhFields)==2:
- genome, pamSeq = ifhFields
- position = None
- else:
- genome, pamSeq, position = ifhFields
- inSeq = ifh.readline().strip()
- ifh.seek(0)
- seqs = parseFasta(ifh)
- ifh.close()
- extSeq = None
- if "extSeq" in seqs:
- extSeq = seqs["extSeq"]
- return inSeq, genome, pamSeq, position, extSeq
- def gzipStr(s):
- " compress a string with gzip and return "
- out = StringIO()
- with gzip.GzipFile(fileobj=out, mode="w") as f:
- f.write(s)
- return out.getvalue()
- def gunzipStr(s):
- " uncompress a string with gzip and return "
- print(len(s), type(s), dir(s))
- f = gzip.GzipFile(StringIO(s))
- result = f.read()
- f.close()
- return result
- def openDbm(dbFname, mode):
- " some distributions don't include the dbm module anymore "
- #import dbm.ndbm
- #dbMod = dbm
- #import dbm.gnu
- #dbMod = gdbm
- #import semidbm
- # lmdbm is faster than everything else: https://pypi.org/project/lmdbm/
- # though semidbm is not bad either
- # Also see leveldb Wiki page
- #import lmdb
- #import dbm.dumb
- from lmdbm import Lmdb
- db = Lmdb.open(dbFname+".lmdb", mode)
- if mode=="c":
- # hack, I don't know how to set the permissions on the open call
- cmd = ["chmod", "-R", "a+rw",dbFname+".lmdb"]
- runCmd(cmd, useShell=False)
- return db
- def saveOutcomeData(batchId, data):
- """ save outcome data of batch. data is a dictionary with key = score name """
- batchBase = join(batchDir, batchId)
- dbFname = batchBase
- db = openDbm(dbFname, "c")
- #conn = sqlite3.connect(dbFname, "w")
- #c = conn.cursor()
- #c.execute('''CREATE TABLE outcomes (id text PRIMARY KEY, data blob))''' % scoreName)
- #c.commit()
- for scoreName, data in data.items():
- #c.execute("INSERT INTO outcomes values (?, ?)", (scoreName, gzipStr(json.dumps(data))))
- db[scoreName] = zlib.compress(json.dumps(data).encode("utf8"))
- db.close()
- #c.commit()
- def readOutcomeData(batchId, scoreName):
- """ open outcome data of batch, key is score name """
- batchBase = join(batchDir, batchId)
- #conn = sqlite3.connect(dbFname, "r")
- #c = conn.cursor()
- #binData = c.execute("SELECT data FROM outcomes where id=?", scoreName)
- #try:
- # import dbm.ndbm
- # db = dbm.ndbm.open(batchBase, "r") # dbm always adds .db to the file name
- #except:
- # # old batches on crispor.org are still using gdbm
- #dbFname = batchBase+".dbm"
- # import dbm.gnu
- # db = dbm.gnu.open(dbFname, "r")
- #import dbm.dumb
- #db = dbm.dumb.open(dbFname, "r") # dbm always adds .db to the file name
- db = openDbm(batchBase, "r")
- dbObj = db[scoreName]
- jsonStr = zlib.decompress(dbObj)
- data = json.loads(jsonStr)
- db.close()
- return data
- def findAllPams(seq, pam):
- """ find all matches for PAM and return as dict startPos -> strand and a set
- of end positions. The start positions for the negative strand are for the
- rev-complemented PAM
- """
- seq = seq.upper()
- startDict, endSet = findPams(seq, pam, "+", {}, set())
- startDict, endSet = findPams(seq, revComp(pam), "-", startDict, endSet)
- if pam in multiPams:
- for pam2 in multiPams[pam]:
- startDict, endSet = findPams(seq, pam2, "+", startDict, endSet)
- startDict, endSet = findPams(seq, revComp(pam2), "-", startDict, endSet)
- return startDict, endSet
- def newBatch(batchName, seq, org, pam):
- """ obtain a batch ID and write seq/org/pam to their files.
- Return batchId.
- """
- batchId = makeTempBase(seq, org, pam, batchName)
- batchData = {}
- batchData["org"] = org
- batchData["pam"] = pam
- batchData["batchName"] = batchName
- batchData["seq"] = seq
- batchData["posStr"] = ""
- writeBatchAsDict(batchData, batchId)
- return batchId
- def readDbInfo(org):
- " return a dbInfo object with the columsn in the genomeInfo.tab file "
- myDir = dirname(__file__)
- genomesDir = join(myDir, "genomes")
- infoFname = join(genomesDir, org, "genomeInfo.tab")
- if not isfile(infoFname):
- return None
- dbInfo = next(lineFileNext(open(infoFname)))
- return dbInfo
- def printQueryNotFoundNote(dbInfo):
- print("<div class='title'>Query sequence, not found in the selected genome, %s (%s)</div>" % (dbInfo.scientificName, dbInfo.name))
- print("<div class='substep' style='border: 1px black solid; padding:5px; background-color: aliceblue'>")
- print("<strong>Warning:</strong> The query sequence was not found in the selected genome.")
- print("This can be a valid query, e.g. a GFP sequence.<br>")
- print("If not, you might want to check if you selected the right genome for your query sequence.<br>")
- print("Use a tool like <a target=_blank href='http://genome.ucsc.edu/cgi-bin/hgBlat'>BLAT</a> to check if the " \
- "sequence really has a 100% identical match in the target genome.<p>")
- print("When reading the list of guide sequences and off-targets below, bear in mind that in case that the input sequence is really in the genome and just has a few differences, the software will use the first found match as the on-target as it cannot distinguish 0-mismatch off-targets from 0-mismatch on-targets. In this case, the specificity scores of guide sequences are too low. In other words, some guides may be fine, the problem may just be that the on-target is shown as an off-target. <br>")
- print("Because there is no flanking sequence available, the guides in your sequence that are within 50bp of the ends will have no efficiency scores. The efficiency scores will instead be shown as '--'. Include more flanking sequence > 50bp to obtain these scores.")
- print("</div>")
- def getOfftargets(seq, org, pamDesc, batchId, startDict, queue):
- """ write guides to fasta and run bwa or use cached results.
- Return name of the BED file with the matches or None if not yet available.
- Write progress status updates to queue object.
- """
- pam = setupPamInfo(pamDesc)
- assert('-' not in pam)
- batchBase = join(batchDir, batchId)
- otBedFname = batchBase+".bed.gz"
- batchInfo = readBatchAsDict(batchId)
- flagFile = batchBase+".running"
- if isfile(flagFile):
- errAbort("This sequence is still being processed. Please wait for ~20 seconds "
- "and try again, e.g. by reloading this page. If you see this message for "
- "more than 2-3 minutes, please send an email to %s. Thanks!" % contactEmail)
- if not batchInfo or not isfile(otBedFname) or commandLineMode or not "posStr" in batchInfo or \
- (batchInfo["posStr"]=="" and not batchInfo["org"]=="noGenome"): # pre-4.8 batches don't have a posStr at all
- # write potential PAM sites to file
- faFname = batchBase+".fa"
- writePamFlank(seq, startDict, pam, faFname)
- if commandLineMode:
- processSubmission(faFname, org, pamDesc, otBedFname, batchBase, batchId, queue)
- else:
- q = JobQueue()
- q.openSqlite()
- ip = os.environ.get("REMOTE_ADDR", "noIp")
- if ip=="195.176.112.240":
- errAbort("IP address blocked.")
- wasOk = q.addJob("search", batchId, "ip=%s,org=%s,pam=%s" % (ip, org, pamDesc))
- if not wasOk:
- print("CRISPOR job %s failed-running..." % batchId)
- pass
- q.close()
- return None
- return otBedFname
- def showPamWarning(pam):
- if pamIsCpf1(pam):
- print('<div style="text-align:left; border: 1px solid; background-color: aliceblue; padding: 3px">')
- print("<strong>Note:</strong> You are using the Cpf1 enzyme or related enzyme.")
- print("While there is an efficiency score specificially for Cpf1, there is no off-target ranking algorithm available in the literature, to our knowledge. We use Hsu and CFD scores below for off-target ranking, but they were developed for spCas9. There is not enough data yet to support their usefulness for Cpf1. Contact us for more info if you need to rank Cpf1 off-targets for validation or if you have a dataset that could elucidate this question. We are showing out-of-frame scores, but they are based on micro-homology that assumes a spCas9 cut site, so most likely the out-of-frame scores are not accurate for the staggered cut of Cpf1 either.")
- print('</div>')
- #elif pamIsSaCas9(pam):
- #print '<div style="text-align:left; border: 1px solid; background-color: aliceblue; padding: 3px">'
- #print "<strong>Note:</strong> Your query is using a Cas9 from S. aureus.<br>"
- #print "Please note that while the efficiency scoring was built for saCas9, the off-target ranking below and specificity scores are based on CFD/Hsu models, which were developed for spCas9. The ranking of off-targets could be very inaccurate. If you have a saCas9 off-target dataset, you can contact us for further info, we are only aware of the BLESS dataset by <a href='https://www.nature.com/articles/nature14299' target=_blank>Ran et al. 2015</a>.<br>As for out-of-frame and micro-homology, this model is also based on spCas9, but <a target=_blank href='https://www.nature.com/articles/nature14299'>Ran et al 2015</a> showed that the saCas9 cleavage pattern looks identical to spCas9's, so the OOF micro-homology model should work with saCas9."
- #print '</div>'
- elif not pamIsSpCas9(pam) and not pamIsSaCas9(pam):
- print('<div style="text-align:left; border: 1px solid; background-color: aliceblue; padding: 3px">')
- print("<strong>Warning:</strong> Your query involves a Cas9 that is not from S. Pyogenes and is also not Cpf1 nor saCas9.")
- print("Please bear in mind that specificity and efficiency scores were designed using data with S. Pyogenes Cas9 and will very likely not be applicable to this particular Cas9. There is nothing we can do about this, we are unaware of a published dataset for this enzyme. If you know one, please contact us. Also contact us if you think another one of the existing scoring model would be more appropriate for this enzyme.<br>")
- print('</div>')
- if pam=="NNNNACA":
- printNote("You selected the old version of the CjCas9 PAM. You may want to select the more recent "+
- "PAMs from the menu on the first page, based on the study by "+
- "<a target=_blank href='https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5473640/'>Kim et al 2017</a>.")
- if pam=="NGN":
- printNote("You have selected the NGN pam for xCas9. While this PAM is documented to work, "+
- "if you read the paper in detail, you will notice that the editing efficiency is much lower. "+
- "For optimal efficiency, consider going back and switching to the 'high-efficiency' xCas9 PAM.")
- if pam=="NGK":
- printNote("You have selected the most efficient PAM for xCas9. You can also select the more general/flexible"+
- " NGN PAM from the menu when you submit your job. If you read the xCas9 paper in detail, you will find "+
- "that NGN is not as efficient though.")
- if pam=="TTTN":
- printWarning("You selected TTTN as the PAM for Cpf1. " +
- "This is not the best PAM. The actual PAM is TTTV, as shown in Fig. 2a of " +
- "<a href='https://www.ncbi.nlm.nih.gov/pubmed/27992409'>Kim HK et al. Nat Meth 2017</a>.<br>")
- def showNoGe
crispor.py at commit 3aa198f, under other · at the source
Overview
- Université Paris Saclay, Université d’Evry, Inserm, IStem, UMR861, 91100 Corbeil-Essonnes, France
- IStem, CECS, 91100 Corbeil-Essonnes, France
- IStem, CECS, the research and innovation team, 91100 Corbeil-Essonnes, France
- Université Paris-Saclay, CEA, Centre National de Recherche en Génomique Humaine (CNRGH), 91057 Evry, France
- Université Grenoble Alpes, Inserm, CEA, UA13, BGE, 38000 Grenoble, France
- Ecole Normale Supérieure de Lyon, Inserm, U1293, CNRS, UMR 5239, Université Claude Bernard Lyon 1, Laboratory of Biology and Modelling of the Cell, 46 allée d'Italie 69364 Lyon, France
- Department of Pediatrics and Adolescent Medicine, University Medical Center, Göttingen, Germany
- Laboratory of Translational Research for Neurological Disorders, Imagine Institute, Université de Paris, INSERM, UMR 1163, 75015 Paris, France
- Sorbonne Université, Institut du Cerveau - Paris Brain Institute - ICM, APHP, Inserm, CNRS, Département de Neurologie, Centre SLA de Paris, Hôpital Pitié-Salpêtrière, 75013 Paris, France
- Alliance on Clinical Trials for ALS-MND (ACT4ALS-MND), Neuroscience Clinical Investigation Center, Paris Brain Institute, 75013 Paris, France
Abstract
The classical paradigm of drug screening often faces significant limitations due to the challenges associated with identifying molecular or cellular read-outs that are relevant to specific genetic diseases. To remedy this, an alternative approach of reverse phenotypic mapping was tested: Compounds were evaluated for their effects on gene expression and alternative splicing in a healthy cell model, and the resulting data were matched to molecular signatures of diseases. A subset of 50 drugs was tested on mesenchymal stem cells derived from a human pluripotent stem cell line. Over half of the compounds altered gene expression, many affecting pathways linked to monogenic diseases. One hit, increased SQSTM1 expression induced by prazosin, was further validated in FTD/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.
maximilianh/crisporWebsite
3aa198f1371bb15733acd98aaa097aad052cdbe3, 26 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
860 files
- CFD_Scoring/
cfd-score-calculator.py , Python, 69 lines - bin/
Azimuth-2.0/ , Python, 1 lineazimuth/ __init__.py - bin/
Azimuth-2.0/ , Python, 54 linesazimuth/ cli_run_model.py - bin/
Azimuth-2.0/ , Python, 118 linesazimuth/ cluster_job.py - bin/
Azimuth-2.0/ , Python, 120 linesazimuth/ corrstats.py - bin/
Azimuth-2.0/ , Python, 1 lineazimuth/ features/ __init__.py - bin/
Azimuth-2.0/ , Python, 552 linesazimuth/ features/ featurization.py - bin/
Azimuth-2.0/ , Python, 102 linesazimuth/ features/ microhomology.py - bin/
Azimuth-2.0/ , Python, 481 linesazimuth/ load_data.py - bin/
Azimuth-2.0/ , Python, 30 linesazimuth/ local_multiprocessing.py - bin/
Azimuth-2.0/ , Python, 655 linesazimuth/ metrics.py - bin/
Azimuth-2.0/ , Python, 649 linesazimuth/ model_comparison.py - bin/
Azimuth-2.0/ , Python, 74 linesazimuth/ models/ DNN.py - bin/
Azimuth-2.0/ , Python, 114 linesazimuth/ models/ GP.py - bin/
Azimuth-2.0/ , Python, 1 lineazimuth/ models/ __init__.py - bin/
Azimuth-2.0/ , Python, 97 linesazimuth/ models/ baselines.py - bin/
Azimuth-2.0/ , Python, 212 linesazimuth/ models/ ensembles.py - bin/
Azimuth-2.0/ , Python, 56 linesazimuth/ models/ gpy_ssk.py - bin/
Azimuth-2.0/ , Python, 287 linesazimuth/ models/ regression.py - bin/
Azimuth-2.0/ , Python, 38 linesazimuth/ models/ ssk.py - bin/
Azimuth-2.0/ , Python, 370 linesazimuth/ predict.py - bin/
Azimuth-2.0/ , Python, 22 linesazimuth/ tests/ test_saved_models.py - bin/
Azimuth-2.0/ , Python, 1,331 linesazimuth/ util.py - bin/
Azimuth-2.0/ , Python, 649 linesmodel_comparison.py - bin/
Azimuth-2.0/ , Python, 15 linessetup.py - bin/
Darwin/ , C, 177 linesWU-CRISPR/ libsvm-2.82/ svm-predict.c - bin/
Darwin/ , C, 308 linesWU-CRISPR/ libsvm-2.82/ svm-scale.c - bin/
Darwin/ , C++, 422 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ callbacks.cpp - bin/
Darwin/ , C/C++, 54 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ callbacks.h - bin/
Darwin/ , C, 164 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ interface.c - bin/
Darwin/ , C/C++, 14 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ interface.h - bin/
Darwin/ , C, 23 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ main.c - bin/
Darwin/ , C++, 433 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ qt/ svm-toy.cpp - bin/
Darwin/ , C++, 456 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ windows/ svm-toy.cpp - bin/
Darwin/ , C, 318 linesWU-CRISPR/ libsvm-2.82/ svm-train.c - bin/
Darwin/ , C++, 3,075 linesWU-CRISPR/ libsvm-2.82/ svm.cpp - bin/
Darwin/ , C/C++, 70 linesWU-CRISPR/ libsvm-2.82/ svm.h - bin/
Darwin/ , Perl, 688 linesWU-CRISPR/ wu-crispr.pl - bin/
Linux-aarch64/ , Perl, 688 linesWU-CRISPR/ wu-crispr.pl - bin/
Linux-x86_64/ , Java, 2,811 linesWU-CRISPR/ libsvm-2.82/ java/ libsvm/ svm.java - bin/
Linux-x86_64/ , Java, 21 linesWU-CRISPR/ libsvm-2.82/ java/ libsvm/ svm_model.java - bin/
Linux-x86_64/ , Java, 6 linesWU-CRISPR/ libsvm-2.82/ java/ libsvm/ svm_node.java - bin/
Linux-x86_64/ , Java, 47 linesWU-CRISPR/ libsvm-2.82/ java/ libsvm/ svm_parameter.java - bin/
Linux-x86_64/ , Java, 7 linesWU-CRISPR/ libsvm-2.82/ java/ libsvm/ svm_problem.java - bin/
Linux-x86_64/ , Java, 149 linesWU-CRISPR/ libsvm-2.82/ java/ svm_predict.java - bin/
Linux-x86_64/ , Java, 471 linesWU-CRISPR/ libsvm-2.82/ java/ svm_toy.java - bin/
Linux-x86_64/ , Java, 299 linesWU-CRISPR/ libsvm-2.82/ java/ svm_train.java - bin/
Linux-x86_64/ , Python, 31 linesWU-CRISPR/ libsvm-2.82/ python/ cross_validation.py - bin/
Linux-x86_64/ , Python, 283 linesWU-CRISPR/ libsvm-2.82/ python/ svm.py - bin/
Linux-x86_64/ , Python, 57 linesWU-CRISPR/ libsvm-2.82/ python/ svm_test.py - bin/
Linux-x86_64/ , C, 3,569 linesWU-CRISPR/ libsvm-2.82/ python/ svmc_wrap.c - bin/
Linux-x86_64/ , Python, 27 linesWU-CRISPR/ libsvm-2.82/ python/ test_cross_validation.py - bin/
Linux-x86_64/ , C, 177 linesWU-CRISPR/ libsvm-2.82/ svm-predict.c - bin/
Linux-x86_64/ , C, 308 linesWU-CRISPR/ libsvm-2.82/ svm-scale.c - bin/
Linux-x86_64/ , C++, 422 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ callbacks.cpp - bin/
Linux-x86_64/ , C/C++, 54 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ callbacks.h - bin/
Linux-x86_64/ , C, 164 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ interface.c - bin/
Linux-x86_64/ , C/C++, 14 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ interface.h - bin/
Linux-x86_64/ , C, 23 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ gtk/ main.c - bin/
Linux-x86_64/ , C++, 433 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ qt/ svm-toy.cpp - bin/
Linux-x86_64/ , C++, 456 linesWU-CRISPR/ libsvm-2.82/ svm-toy/ windows/ svm-toy.cpp - bin/
Linux-x86_64/ , C, 318 linesWU-CRISPR/ libsvm-2.82/ svm-train.c - bin/
Linux-x86_64/ , C++, 3,075 linesWU-CRISPR/ libsvm-2.82/ svm.cpp - bin/
Linux-x86_64/ , C/C++, 70 linesWU-CRISPR/ libsvm-2.82/ svm.h - bin/
Linux-x86_64/ , Python, 78 linesWU-CRISPR/ libsvm-2.82/ tools/ easy.py - bin/
Linux-x86_64/ , Python, 350 linesWU-CRISPR/ libsvm-2.82/ tools/ grid.py - bin/
Linux-x86_64/ , Python, 349 linesWU-CRISPR/ libsvm-2.82/ tools/ grid_bk.py - bin/
Linux-x86_64/ , Python, 144 linesWU-CRISPR/ libsvm-2.82/ tools/ subset.py - bin/
Linux-x86_64/ , Perl, 688 linesWU-CRISPR/ wu-crispr.pl - bin/
azimuthMoreno/ , Python, 56 linespredict_moreno.py - bin/
deepCpf1/ , Python, 106 linesDeepCpf1-orig.py - bin/
deepCpf1/ , Python, 140 linesDeepCpf1.py - bin/
fusiDoench/ , Python, 1 lineanalysis/ features/ __init__.py - bin/
fusiDoench/ , Python, 465 linesanalysis/ features/ featurization.py - bin/
fusiDoench/ , Python, 102 linesanalysis/ features/ microhomology.py - bin/
fusiDoench/ , Python, 365 linesanalysis/ load_data.py - bin/
fusiDoench/ , Python, 30 linesanalysis/ local_multiprocessing.py - bin/
fusiDoench/ , Python, 376 linesanalysis/ metrics.py - bin/
fusiDoench/ , Python, 609 linesanalysis/ model_comparison.py - bin/
fusiDoench/ , Python, 74 linesanalysis/ models/ DNN.py - bin/
fusiDoench/ , Python, 114 linesanalysis/ models/ GP.py - bin/
fusiDoench/ , Python, 1 lineanalysis/ models/ __init__.py - bin/
fusiDoench/ , Python, 76 linesanalysis/ models/ baselines.py - bin/
fusiDoench/ , Python, 157 linesanalysis/ models/ ensembles.py - bin/
fusiDoench/ , Python, 56 linesanalysis/ models/ gpy_ssk.py - bin/
fusiDoench/ , Python, 257 linesanalysis/ models/ regression.py - bin/
fusiDoench/ , Python, 38 linesanalysis/ models/ ssk.py - bin/
fusiDoench/ , Python, 337 linesanalysis/ predict.py - bin/
fusiDoench/ , Python, 50 linesanalysis/ rs2_score_calculator.py - bin/
fusiDoench/ , Python, 786 linesanalysis/ util.py - bin/
fusiDoench/ , Python, 162 linesanalysis/ util_new.py - bin/
najm2018/ , Python, 39 linessaureus_scoring_v1.py - bin/
src/ , C/C++, 17 linesSSC0.1/ include/ performance.h - bin/
src/ , C/C++, 22 linesSSC0.1/ include/ words.h - bin/
src/ , C, 286 linesSSC0.1/ src/ Fasta2Spacer.c - bin/
src/ , C, 274 linesSSC0.1/ src/ SSC.c - bin/
src/ , C, 32 linesSSC0.1/ src/ performance.c - bin/
src/ , C, 133 linesSSC0.1/ src/ words.c - bin/
src/ , C, 85 linesViennaRNA-2.1.9/ Cluster/ AD_main.c - bin/
src/ , C, 294 linesViennaRNA-2.1.9/ Cluster/ AS_main.c - bin/
src/ , C, 143 linesViennaRNA-2.1.9/ Cluster/ PS3D.c - bin/
src/ , C/C++, 2 linesViennaRNA-2.1.9/ Cluster/ PS3D.h - bin/
src/ , C/C++, 107 linesViennaRNA-2.1.9/ Cluster/ StrEdit_CostMatrix.h - bin/
src/ , C, 281 linesViennaRNA-2.1.9/ Cluster/ cluster.c - bin/
src/ , C/C++, 16 linesViennaRNA-2.1.9/ Cluster/ cluster.h - bin/
src/ , C, 568 linesViennaRNA-2.1.9/ Cluster/ distance_matrix.c - bin/
src/ , C/C++, 14 linesViennaRNA-2.1.9/ Cluster/ distance_matrix.h - bin/
src/ , C, 332 linesViennaRNA-2.1.9/ Cluster/ split.c - bin/
src/ , C/C++, 14 linesViennaRNA-2.1.9/ Cluster/ split.h - bin/
src/ , C, 515 linesViennaRNA-2.1.9/ Cluster/ statgeom.c - bin/
src/ , C/C++, 5 linesViennaRNA-2.1.9/ Cluster/ statgeom.h - bin/
src/ , C, 358 linesViennaRNA-2.1.9/ Cluster/ treeplot.c - bin/
src/ , C/C++, 1 lineViennaRNA-2.1.9/ Cluster/ treeplot.h - bin/
src/ , C/C++, 134 linesViennaRNA-2.1.9/ H/ 2Dfold.h - bin/
src/ , C/C++, 173 linesViennaRNA-2.1.9/ H/ 2Dpfold.h - bin/
src/ , C/C++, 164 linesViennaRNA-2.1.9/ H/ LPfold.h - bin/
src/ , C/C++, 81 linesViennaRNA-2.1.9/ H/ Lfold.h - bin/
src/ , C/C++, 32 linesViennaRNA-2.1.9/ H/ MEA.h - bin/
src/ , C/C++, 25 linesViennaRNA-2.1.9/ H/ PKplex.h - bin/
src/ , C/C++, 197 linesViennaRNA-2.1.9/ H/ PS_dot.h - bin/
src/ , C/C++, 14 linesViennaRNA-2.1.9/ H/ ProfileAln.h - bin/
src/ , C/C++, 152 linesViennaRNA-2.1.9/ H/ RNAstruct.h - bin/
src/ , C/C++, 40 linesViennaRNA-2.1.9/ H/ ali_plex.h - bin/
src/ , C/C++, 378 linesViennaRNA-2.1.9/ H/ alifold.h - bin/
src/ , C/C++, 10 linesViennaRNA-2.1.9/ H/ aln_util.h - bin/
src/ , C/C++, 182 linesViennaRNA-2.1.9/ H/ cofold.h - bin/
src/ , C/C++, 91 linesViennaRNA-2.1.9/ H/ convert_epars.h - bin/
src/ , C/C++, 775 linesViennaRNA-2.1.9/ H/ data_structures.h - bin/
src/ , C/C++, 51 linesViennaRNA-2.1.9/ H/ dist_vars.h - bin/
src/ , C/C++, 28 linesViennaRNA-2.1.9/ H/ duplex.h - bin/
src/ , C/C++, 53 linesViennaRNA-2.1.9/ H/ edit_cost.h - bin/
src/ , C/C++, 33 linesViennaRNA-2.1.9/ H/ energy_const.h - bin/
src/ , C/C++, 100 linesViennaRNA-2.1.9/ H/ energy_par.h - bin/
src/ , C/C++, 53 linesViennaRNA-2.1.9/ H/ findpath.h - bin/
src/ , C/C++, 604 linesViennaRNA-2.1.9/ H/ fold.h - bin/
src/ , C/C++, 221 linesViennaRNA-2.1.9/ H/ fold_vars.h - bin/
src/ , C/C++, 726 linesViennaRNA-2.1.9/ H/ gquad.h - bin/
src/ , C/C++, 67 linesViennaRNA-2.1.9/ H/ inverse.h - bin/
src/ , C/C++, 661 linesViennaRNA-2.1.9/ H/ loop_energies.h - bin/
src/ , C/C++, 21 linesViennaRNA-2.1.9/ H/ mm.h - bin/
src/ , C/C++, 89 linesViennaRNA-2.1.9/ H/ move_set.h - bin/
src/ , C/C++, 16 linesViennaRNA-2.1.9/ H/ naview.h - bin/
src/ , C/C++, 148 linesViennaRNA-2.1.9/ H/ pair_mat.h - bin/
src/ , C/C++, 134 linesViennaRNA-2.1.9/ H/ params.h - bin/
src/ , C/C++, 444 linesViennaRNA-2.1.9/ H/ part_func.h - bin/
src/ , C/C++, 235 linesViennaRNA-2.1.9/ H/ part_func_co.h - bin/
src/ , C/C++, 141 linesViennaRNA-2.1.9/ H/ part_func_up.h - bin/
src/ , C/C++, 80 linesViennaRNA-2.1.9/ H/ plex.h - bin/
src/ , C/C++, 106 linesViennaRNA-2.1.9/ H/ plot_layouts.h - bin/
src/ , C/C++, 58 linesViennaRNA-2.1.9/ H/ profiledist.h - bin/
src/ , C/C++, 57 linesViennaRNA-2.1.9/ H/ read_epars.h - bin/
src/ , C/C++, 8 linesViennaRNA-2.1.9/ H/ ribo.h - bin/
src/ , C/C++, 58 linesViennaRNA-2.1.9/ H/ snofold.h - bin/
src/ , C/C++, 284 linesViennaRNA-2.1.9/ H/ snoop.h - bin/
src/ , C/C++, 30 linesViennaRNA-2.1.9/ H/ stringdist.h - bin/
src/ , C/C++, 117 linesViennaRNA-2.1.9/ H/ subopt.h - bin/
src/ , C/C++, 46 linesViennaRNA-2.1.9/ H/ svm_utils.h - bin/
src/ , C/C++, 42 linesViennaRNA-2.1.9/ H/ treedist.h - bin/
src/ , C/C++, 615 linesViennaRNA-2.1.9/ H/ utils.h - bin/
src/ , Perl, 16 linesViennaRNA-2.1.9/ Kinfold/ Laplace/ extract_data.pl - bin/
src/ , Shell, 26 linesViennaRNA-2.1.9/ Kinfold/ Laplace/ laplace.sh - bin/
src/ , R, 6 linesViennaRNA-2.1.9/ Kinfold/ Laplace/ to_boxplot.R - bin/
src/ , C, 764 linesViennaRNA-2.1.9/ Kinfold/ baum.c - bin/
src/ , C/C++, 20 linesViennaRNA-2.1.9/ Kinfold/ baum.h - bin/
src/ , C, 133 linesViennaRNA-2.1.9/ Kinfold/ cache.c - bin/
src/ , C/C++, 33 linesViennaRNA-2.1.9/ Kinfold/ cache_util.h - bin/
src/ , C, 1,689 linesViennaRNA-2.1.9/ Kinfold/ cmdline.c - bin/
src/ , C/C++, 259 linesViennaRNA-2.1.9/ Kinfold/ cmdline.h - bin/
src/ , C/C++, 91 linesViennaRNA-2.1.9/ Kinfold/ config.h - bin/
src/ , C, 584 linesViennaRNA-2.1.9/ Kinfold/ globals.c - bin/
src/ , C/C++, 79 linesViennaRNA-2.1.9/ Kinfold/ globals.h - bin/
src/ , C, 224 linesViennaRNA-2.1.9/ Kinfold/ main.c - bin/
src/ , C, 440 linesViennaRNA-2.1.9/ Kinfold/ nachbar.c - bin/
src/ , C/C++, 21 linesViennaRNA-2.1.9/ Kinfold/ nachbar.h - bin/
src/ , C, 344 linesViennaRNA-2.1.9/ Progs/ RNA2Dfold.c - bin/
src/ , C, 1,939 linesViennaRNA-2.1.9/ Progs/ RNA2Dfold_cmdl.c - bin/
src/ , C/C++, 308 linesViennaRNA-2.1.9/ Progs/ RNA2Dfold_cmdl.h - bin/
src/ , C, 393 linesViennaRNA-2.1.9/ Progs/ RNALalifold.c - bin/
src/ , C, 1,681 linesViennaRNA-2.1.9/ Progs/ RNALalifold_cmdl.c - bin/
src/ , C/C++, 338 linesViennaRNA-2.1.9/ Progs/ RNALalifold_cmdl.h - bin/
src/ , C, 186 linesViennaRNA-2.1.9/ Progs/ RNALfold.c - bin/
src/ , C, 1,502 linesViennaRNA-2.1.9/ Progs/ RNALfold_cmdl.c - bin/
src/ , C/C++, 289 linesViennaRNA-2.1.9/ Progs/ RNALfold_cmdl.h - bin/
src/ , C, 520 linesViennaRNA-2.1.9/ Progs/ RNAPKplex.c - bin/
src/ , C, 1,417 linesViennaRNA-2.1.9/ Progs/ RNAPKplex_cmdl.c - bin/
src/ , C/C++, 268 linesViennaRNA-2.1.9/ Progs/ RNAPKplex_cmdl.h - bin/
src/ , C, 172 linesViennaRNA-2.1.9/ Progs/ RNAaliduplex.c - bin/
src/ , C, 1,428 linesViennaRNA-2.1.9/ Progs/ RNAaliduplex_cmdl.c - bin/
src/ , C/C++, 262 linesViennaRNA-2.1.9/ Progs/ RNAaliduplex_cmdl.h - bin/
src/ , C, 700 linesViennaRNA-2.1.9/ Progs/ RNAalifold.c - bin/
src/ , C, 1,995 linesViennaRNA-2.1.9/ Progs/ RNAalifold_cmdl.c - bin/
src/ , C/C++, 428 linesViennaRNA-2.1.9/ Progs/ RNAalifold_cmdl.h - bin/
src/ , C, 687 linesViennaRNA-2.1.9/ Progs/ RNAcofold.c - bin/
src/ , C, 1,718 linesViennaRNA-2.1.9/ Progs/ RNAcofold_cmdl.c - bin/
src/ , C/C++, 327 linesViennaRNA-2.1.9/ Progs/ RNAcofold_cmdl.h - bin/
src/ , C, 548 linesViennaRNA-2.1.9/ Progs/ RNAdistance.c - bin/
src/ , C, 1,216 linesViennaRNA-2.1.9/ Progs/ RNAdistance_cmdl.c - bin/
src/ , C/C++, 207 linesViennaRNA-2.1.9/ Progs/ RNAdistance_cmdl.h - bin/
src/ , C, 204 linesViennaRNA-2.1.9/ Progs/ RNAduplex.c - bin/
src/ , C, 1,458 linesViennaRNA-2.1.9/ Progs/ RNAduplex_cmdl.c - bin/
src/ , C/C++, 268 linesViennaRNA-2.1.9/ Progs/ RNAduplex_cmdl.h - bin/
src/ , C, 244 linesViennaRNA-2.1.9/ Progs/ RNAeval.c - bin/
src/ , C, 1,367 linesViennaRNA-2.1.9/ Progs/ RNAeval_cmdl.c - bin/
src/ , C/C++, 251 linesViennaRNA-2.1.9/ Progs/ RNAeval_cmdl.h - bin/
src/ , C, 415 linesViennaRNA-2.1.9/ Progs/ RNAfold.c - bin/
src/ , C, 1,754 linesViennaRNA-2.1.9/ Progs/ RNAfold_cmdl.c - bin/
src/ , C/C++, 340 linesViennaRNA-2.1.9/ Progs/ RNAfold_cmdl.h - bin/
src/ , C, 253 linesViennaRNA-2.1.9/ Progs/ RNAheat.c - bin/
src/ , C, 1,506 linesViennaRNA-2.1.9/ Progs/ RNAheat_cmdl.c - bin/
src/ , C/C++, 289 linesViennaRNA-2.1.9/ Progs/ RNAheat_cmdl.h - bin/
src/ , C, 249 linesViennaRNA-2.1.9/ Progs/ RNAinverse.c - bin/
src/ , C, 1,513 linesViennaRNA-2.1.9/ Progs/ RNAinverse_cmdl.c - bin/
src/ , C/C++, 293 linesViennaRNA-2.1.9/ Progs/ RNAinverse_cmdl.h - bin/
src/ , C, 314 linesViennaRNA-2.1.9/ Progs/ RNApaln.c - bin/
src/ , C, 1,591 linesViennaRNA-2.1.9/ Progs/ RNApaln_cmdl.c - bin/
src/ , C/C++, 318 linesViennaRNA-2.1.9/ Progs/ RNApaln_cmdl.h - bin/
src/ , C, 111 linesViennaRNA-2.1.9/ Progs/ RNAparconv.c - bin/
src/ , C, 1,517 linesViennaRNA-2.1.9/ Progs/ RNAparconv_cmdl.c - bin/
src/ , C/C++, 310 linesViennaRNA-2.1.9/ Progs/ RNAparconv_cmdl.h - bin/
src/ , C, 326 linesViennaRNA-2.1.9/ Progs/ RNApdist.c - bin/
src/ , C, 1,531 linesViennaRNA-2.1.9/ Progs/ RNApdist_cmdl.c - bin/
src/ , C/C++, 280 linesViennaRNA-2.1.9/ Progs/ RNApdist_cmdl.h - bin/
src/ , C, 2,390 linesViennaRNA-2.1.9/ Progs/ RNAplex.c - bin/
src/ , C, 1,689 linesViennaRNA-2.1.9/ Progs/ RNAplex_cmdl.c - bin/
src/ , C/C++, 346 linesViennaRNA-2.1.9/ Progs/ RNAplex_cmdl.h - bin/
src/ , C, 434 linesViennaRNA-2.1.9/ Progs/ RNAplfold.c - bin/
src/ , C, 1,642 linesViennaRNA-2.1.9/ Progs/ RNAplfold_cmdl.c - bin/
src/ , C/C++, 338 linesViennaRNA-2.1.9/ Progs/ RNAplfold_cmdl.h - bin/
src/ , C, 144 linesViennaRNA-2.1.9/ Progs/ RNAplot.c - bin/
src/ , C, 1,198 linesViennaRNA-2.1.9/ Progs/ RNAplot_cmdl.c - bin/
src/ , C/C++, 211 linesViennaRNA-2.1.9/ Progs/ RNAplot_cmdl.h - bin/
src/ , C, 1,083 linesViennaRNA-2.1.9/ Progs/ RNAsnoop.c - bin/
src/ , C, 1,886 linesViennaRNA-2.1.9/ Progs/ RNAsnoop_cmdl.c - bin/
src/ , C/C++, 446 linesViennaRNA-2.1.9/ Progs/ RNAsnoop_cmdl.h - bin/
src/ , C, 383 linesViennaRNA-2.1.9/ Progs/ RNAsubopt.c - bin/
src/ , C, 1,745 linesViennaRNA-2.1.9/ Progs/ RNAsubopt_cmdl.c - bin/
src/ , C/C++, 338 linesViennaRNA-2.1.9/ Progs/ RNAsubopt_cmdl.h - bin/
src/ , C, 1,142 linesViennaRNA-2.1.9/ Progs/ RNAup.c - bin/
src/ , C, 2,082 linesViennaRNA-2.1.9/ Progs/ RNAup_cmdl.c - bin/
src/ , C/C++, 354 linesViennaRNA-2.1.9/ Progs/ RNAup_cmdl.h - bin/
src/ , C/C++, 80 linesViennaRNA-2.1.9/ RNAforester/ config.h - bin/
src/ , C++, 40 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ demo_cpp.cpp - bin/
src/ , C, 60 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ g2_anim.c - bin/
src/ , C, 148 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ g2_arc.c - bin/
src/ , C, 143 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ g2_splines_demo.c - bin/
src/ , C, 254 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ g2_test.c - bin/
src/ , C, 58 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ handles.c - bin/
src/ , C, 63 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ penguin.c - bin/
src/ , C, 46 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ pointer.c - bin/
src/ , C, 30 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ simple_FIG.c - bin/
src/ , C, 30 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ simple_PS.c - bin/
src/ , C, 31 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ simple_X11.c - bin/
src/ , C, 31 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ simple_gd.c - bin/
src/ , C, 31 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ demo/ simple_win32.c - bin/
src/ , Perl, 10 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ g2_perl/ Makefile.PL - bin/
src/ , Perl, 213 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ g2_perl/ test.pl - bin/
src/ , C/C++, 200 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ include/ g2.h - bin/
src/ , C/C++, 54 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ include/ g2_FIG.h - bin/
src/ , C/C++, 118 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ include/ g2_PS.h - bin/
src/ , C/C++, 45 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ include/ g2_X11.h - bin/
src/ , C/C++, 67 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ include/ g2_gd.h - bin/
src/ , Perl, 58 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ perl/ Makefile.PL - bin/
src/ , C, 2,172 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ perl/ g2_wrap.c - bin/
src/ , C, 486 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ FIG/ g2_FIG.c - bin/
src/ , C/C++, 54 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ FIG/ g2_FIG.h - bin/
src/ , C/C++, 82 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ FIG/ g2_FIG_P.h - bin/
src/ , C/C++, 56 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ FIG/ g2_FIG_funix.h - bin/
src/ , C, 348 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ GD/ g2_gd.c - bin/
src/ , C/C++, 67 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ GD/ g2_gd.h - bin/
src/ , C/C++, 118 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ GD/ g2_gd_P.h - bin/
src/ , C/C++, 54 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ GD/ g2_gd_funix.h - bin/
src/ , C, 624 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ PS/ g2_PS.c - bin/
src/ , C/C++, 118 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ PS/ g2_PS.h - bin/
src/ , C/C++, 99 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ PS/ g2_PS_P.h - bin/
src/ , C/C++, 104 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ PS/ g2_PS_definitions.h - bin/
src/ , C/C++, 56 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ PS/ g2_PS_funix.h - bin/
src/ , C, 671 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2_win32.c - bin/
src/ , C/C++, 68 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2_win32.h - bin/
src/ , C/C++, 119 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2_win32_P.h - bin/
src/ , C/C++, 56 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2_win32_funix.h - bin/
src/ , C, 213 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2_win32_thread.c - bin/
src/ , C, 7 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ g2res.c - bin/
src/ , C/C++, 21 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ Win32/ resource.h - bin/
src/ , C, 797 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ X11/ g2_X11.c - bin/
src/ , C/C++, 45 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ X11/ g2_X11.h - bin/
src/ , C/C++, 94 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ X11/ g2_X11_P.h - bin/
src/ , C/C++, 60 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ X11/ g2_X11_funix.h - bin/
src/ , C/C++, 200 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2.h - bin/
src/ , C/C++, 42 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_config.h - bin/
src/ , C, 314 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_control_pd.c - bin/
src/ , C/C++, 39 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_control_pd.h - bin/
src/ , C, 206 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_device.c - bin/
src/ , C/C++, 64 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_device.h - bin/
src/ , C, 499 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_fif.c - bin/
src/ , C/C++, 160 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_funix.h - bin/
src/ , C, 684 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_graphic_pd.c - bin/
src/ , C/C++, 66 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_graphic_pd.h - bin/
src/ , C, 78 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_physical_device.c - bin/
src/ , C/C++, 61 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_physical_device.h - bin/
src/ , C, 880 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_splines.c - bin/
src/ , C, 612 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_ui_control.c - bin/
src/ , C, 196 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_ui_device.c - bin/
src/ , C, 1,045 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_ui_graphic.c - bin/
src/ , C, 149 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_ui_virtual_device.c - bin/
src/ , C, 224 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_util.c - bin/
src/ , C/C++, 48 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_util.h - bin/
src/ , C, 82 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_virtual_device.c - bin/
src/ , C/C++, 32 linesViennaRNA-2.1.9/ RNAforester/ g2-0.70/ src/ g2_virtual_device.h - bin/
src/ , C++, 220 linesViennaRNA-2.1.9/ RNAforester/ src/ Arguments.cpp - bin/
src/ , C/C++, 110 linesViennaRNA-2.1.9/ RNAforester/ src/ Arguments.h - bin/
src/ , C/C++, 66 linesViennaRNA-2.1.9/ RNAforester/ src/ algebra.h - bin/
src/ , C/C++, 94 linesViennaRNA-2.1.9/ RNAforester/ src/ alignment.h - bin/
src/ , C++, 1,045 linesViennaRNA-2.1.9/ RNAforester/ src/ alignment.t.cpp - bin/
src/ , C/C++, 30 linesViennaRNA-2.1.9/ RNAforester/ src/ debug.h - bin/
src/ , C/C++, 8 linesViennaRNA-2.1.9/ RNAforester/ src/ fold.h - bin/
src/ , C/C++, 37 linesViennaRNA-2.1.9/ RNAforester/ src/ fold_vars.h - bin/
src/ , C, 355 linesViennaRNA-2.1.9/ RNAforester/ src/ glib.c - bin/
src/ , C/C++, 89 linesViennaRNA-2.1.9/ RNAforester/ src/ graphtypes.h - bin/
src/ , C++, 1,127 linesViennaRNA-2.1.9/ RNAforester/ src/ main.cpp - bin/
src/ , C/C++, 61 linesViennaRNA-2.1.9/ RNAforester/ src/ matrix.h - bin/
src/ , C/C++, 21 linesViennaRNA-2.1.9/ RNAforester/ src/ misc.h - bin/
src/ , C++, 37 linesViennaRNA-2.1.9/ RNAforester/ src/ misc.t.cpp - bin/
src/ , C, 140 linesViennaRNA-2.1.9/ RNAforester/ src/ pairs.c - bin/
src/ , C/C++, 13 linesViennaRNA-2.1.9/ RNAforester/ src/ part_func.h - bin/
src/ , C, 80 linesViennaRNA-2.1.9/ RNAforester/ src/ pointer.c - bin/
src/ , C/C++, 124 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforest.h - bin/
src/ , C++, 226 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforest.t.cpp - bin/
src/ , C/C++, 52 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforestali.h - bin/
src/ , C++, 203 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforestbase.cpp - bin/
src/ , C/C++, 225 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforestbase.h - bin/
src/ , C/C++, 69 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforestsz.h - bin/
src/ , C++, 77 linesViennaRNA-2.1.9/ RNAforester/ src/ ppforestsz.t.cpp - bin/
src/ , C++, 343 linesViennaRNA-2.1.9/ RNAforester/ src/ progressive_align.cpp - bin/
src/ , C/C++, 32 linesViennaRNA-2.1.9/ RNAforester/ src/ progressive_align.h - bin/
src/ , C, 129 linesViennaRNA-2.1.9/ RNAforester/ src/ readgraph.c - bin/
src/ , C++, 89 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_algebra.cpp - bin/
src/ , C/C++, 638 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_algebra.h - bin/
src/ , C++, 191 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_alignment.cpp - bin/
src/ , C/C++, 87 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_alignment.h - bin/
src/ , C++, 33 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_alphabet.cpp - bin/
src/ , C/C++, 99 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_alphabet.h - bin/
src/ , C++, 1,294 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_profile_alignment.cp p - bin/
src/ , C/C++, 147 linesViennaRNA-2.1.9/ RNAforester/ src/ rna_profile_alignment.h - bin/
src/ , C++, 183 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforest.cpp - bin/
src/ , C/C++, 61 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforest.h - bin/
src/ , C++, 244 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforester_options.cpp - bin/
src/ , C/C++, 140 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforester_options.h - bin/
src/ , C++, 114 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforestsz.cpp - bin/
src/ , C/C++, 49 linesViennaRNA-2.1.9/ RNAforester/ src/ rnaforestsz.h - bin/
src/ , C++, 1,127 linesViennaRNA-2.1.9/ RNAforester/ src/ rnafuncs.cpp - bin/
src/ , C/C++, 68 linesViennaRNA-2.1.9/ RNAforester/ src/ rnafuncs.h - bin/
src/ , C, 66 linesViennaRNA-2.1.9/ RNAforester/ src/ term.c - bin/
src/ , C/C++, 72 linesViennaRNA-2.1.9/ RNAforester/ src/ treeedit.h - bin/
src/ , C++, 174 linesViennaRNA-2.1.9/ RNAforester/ src/ treeedit.t.cpp - bin/
src/ , C/C++, 7 linesViennaRNA-2.1.9/ RNAforester/ src/ types.h - bin/
src/ , C, 112 linesViennaRNA-2.1.9/ RNAforester/ src/ unpairs.c - bin/
src/ , C/C++, 52 linesViennaRNA-2.1.9/ RNAforester/ src/ utils.h - bin/
src/ , C, 147 linesViennaRNA-2.1.9/ RNAforester/ src/ wmatch.c - bin/
src/ , C/C++, 40 linesViennaRNA-2.1.9/ RNAforester/ src/ wmatch.h - bin/
src/ , C, 302 linesViennaRNA-2.1.9/ Readseq/ macinit.c - bin/
src/ , R, 412 linesViennaRNA-2.1.9/ Readseq/ macinit.r - bin/
src/ , C, 1,139 linesViennaRNA-2.1.9/ Readseq/ readseq.c - bin/
src/ , C, 305 linesViennaRNA-2.1.9/ Readseq/ ureadasn.c - bin/
src/ , C, 1,876 linesViennaRNA-2.1.9/ Readseq/ ureadseq.c - bin/
src/ , C/C++, 169 linesViennaRNA-2.1.9/ Readseq/ ureadseq.h - bin/
src/ , C, 152 linesViennaRNA-2.1.9/ Utils/ b2ct.c - bin/
src/ , Perl, 42 linesViennaRNA-2.1.9/ Utils/ b2mt.pl - bin/
src/ , Perl, 152 linesViennaRNA-2.1.9/ Utils/ cmount.pl - bin/
src/ , Perl, 564 linesViennaRNA-2.1.9/ Utils/ coloraln.pl - bin/
src/ , Perl, 111 linesViennaRNA-2.1.9/ Utils/ colorrna.pl - bin/
src/ , Perl, 27 linesViennaRNA-2.1.9/ Utils/ ct2b.pl - bin/
src/ , C, 198 linesViennaRNA-2.1.9/ Utils/ ct2db.c - bin/
src/ , C, 1,127 linesViennaRNA-2.1.9/ Utils/ ct2db_cmdl.c - bin/
src/ , C/C++, 183 linesViennaRNA-2.1.9/ Utils/ ct2db_cmdl.h - bin/
src/ , Perl, 46 linesViennaRNA-2.1.9/ Utils/ dpzoom.pl - bin/
src/ , Perl, 135 linesViennaRNA-2.1.9/ Utils/ mountain.pl - bin/
src/ , C, 88 linesViennaRNA-2.1.9/ Utils/ popt.c - bin/
src/ , Perl, 327 linesViennaRNA-2.1.9/ Utils/ refold.pl - bin/
src/ , Perl, 200 linesViennaRNA-2.1.9/ Utils/ relplot.pl - bin/
src/ , Perl, 138 linesViennaRNA-2.1.9/ Utils/ rotate_ss.pl - bin/
src/ , Perl, 608 linesViennaRNA-2.1.9/ Utils/ switch.pl - bin/
src/ , C/C++, 163 linesViennaRNA-2.1.9/ config.h - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ 2Dfold_8h.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ 2Dpfold_8h.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ LPfold_8h.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ Lfold_8h.js - bin/
src/ , JavaScript, 4 linesViennaRNA-2.1.9/ doc/ html/ MEA_8h.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ PS__dot_8h.js - bin/
src/ , JavaScript, 19 linesViennaRNA-2.1.9/ doc/ html/ RNAstruct_8h.js - bin/
src/ , JavaScript, 21 linesViennaRNA-2.1.9/ doc/ html/ alifold_8h.js - bin/
src/ , JavaScript, 41 linesViennaRNA-2.1.9/ doc/ html/ annotated.js - bin/
src/ , JavaScript, 13 linesViennaRNA-2.1.9/ doc/ html/ cofold_8h.js - bin/
src/ , JavaScript, 26 linesViennaRNA-2.1.9/ doc/ html/ convert__epars_8h.js - bin/
src/ , JavaScript, 36 linesViennaRNA-2.1.9/ doc/ html/ data__structures_8h.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ dir_97aefd0d527b934f1d99 a682da8fe6a9.js - bin/
src/ , JavaScript, 49 linesViennaRNA-2.1.9/ doc/ html/ dir_d72344b28b4f2089ce25 682c4e6eba22.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ dist__vars_8h.js - bin/
src/ , JavaScript, 97 linesViennaRNA-2.1.9/ doc/ html/ dynsections.js - bin/
src/ , JavaScript, 11 linesViennaRNA-2.1.9/ doc/ html/ energy__const_8h.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ files.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ findpath_8h.js - bin/
src/ , JavaScript, 30 linesViennaRNA-2.1.9/ doc/ html/ fold_8h.js - bin/
src/ , JavaScript, 29 linesViennaRNA-2.1.9/ doc/ html/ fold__vars_8h.js - bin/
src/ , JavaScript, 27 linesViennaRNA-2.1.9/ doc/ html/ globals_dup.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ gquad_8h.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ group__centroid__fold.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ group__class__fold.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ group__cofold.js - bin/
src/ , JavaScript, 17 linesViennaRNA-2.1.9/ doc/ html/ group__consensus__fold.j s - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ group__consensus__mfe__f old.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ group__consensus__pf__fo ld.js - bin/
src/ , JavaScript, 4 linesViennaRNA-2.1.9/ doc/ html/ group__consensus__stochb t.js - bin/
src/ , JavaScript, 4 linesViennaRNA-2.1.9/ doc/ html/ group__dos.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ group__energy__parameter s.js - bin/
src/ , JavaScript, 27 linesViennaRNA-2.1.9/ doc/ html/ group__energy__parameter s__convert.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ group__energy__parameter s__rw.js - bin/
src/ , JavaScript, 10 linesViennaRNA-2.1.9/ doc/ html/ group__eval.js - bin/
src/ , JavaScript, 13 linesViennaRNA-2.1.9/ doc/ html/ group__folding__routines .js - bin/
src/ , JavaScript, 10 linesViennaRNA-2.1.9/ doc/ html/ group__inverse__fold.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ group__kl__neighborhood. js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ group__kl__neighborhood_ _mfe.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ group__kl__neighborhood_ _pf.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ group__kl__neighborhood_ _stochbt.js - bin/
src/ , JavaScript, 4 linesViennaRNA-2.1.9/ doc/ html/ group__local__consensus_ _fold.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ group__local__fold.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ group__local__mfe__fold. js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ group__local__pf__fold.j s - bin/
src/ , JavaScript, 10 linesViennaRNA-2.1.9/ doc/ html/ group__mfe__cofold.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ group__mfe__fold.js - bin/
src/ , JavaScript, 14 linesViennaRNA-2.1.9/ doc/ html/ group__pf__cofold.js - bin/
src/ , JavaScript, 22 linesViennaRNA-2.1.9/ doc/ html/ group__pf__fold.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ group__subopt__fold.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ group__subopt__stochbt.j s - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ group__subopt__wuchty.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ group__subopt__zuker.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ group__up__cofold.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ inverse_8h.js - bin/
src/ , JavaScript, 77 linesViennaRNA-2.1.9/ doc/ html/ jquery.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ loop__energies_8h.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ modules.js - bin/
src/ , JavaScript, 542 linesViennaRNA-2.1.9/ doc/ html/ navtree.js - bin/
src/ , JavaScript, 253 linesViennaRNA-2.1.9/ doc/ html/ navtreeindex0.js - bin/
src/ , JavaScript, 253 linesViennaRNA-2.1.9/ doc/ html/ navtreeindex1.js - bin/
src/ , JavaScript, 253 linesViennaRNA-2.1.9/ doc/ html/ navtreeindex2.js - bin/
src/ , JavaScript, 56 linesViennaRNA-2.1.9/ doc/ html/ navtreeindex3.js - bin/
src/ , JavaScript, 10 linesViennaRNA-2.1.9/ doc/ html/ params_8h.js - bin/
src/ , JavaScript, 27 linesViennaRNA-2.1.9/ doc/ html/ part__func_8h.js - bin/
src/ , JavaScript, 15 linesViennaRNA-2.1.9/ doc/ html/ part__func__co_8h.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ part__func__up_8h.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ plot__layouts_8h.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ profiledist_8h.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ read__epars_8h.js - bin/
src/ , JavaScript, 93 linesViennaRNA-2.1.9/ doc/ html/ resize.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ stringdist_8h.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ structConcEnt.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ structSOLUTION.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ structTwoDfold__solution .js - bin/
src/ , JavaScript, 16 linesViennaRNA-2.1.9/ doc/ html/ structTwoDfold__vars.js - bin/
src/ , JavaScript, 6 linesViennaRNA-2.1.9/ doc/ html/ structTwoDpfold__solutio n.js - bin/
src/ , JavaScript, 15 linesViennaRNA-2.1.9/ doc/ html/ structTwoDpfold__vars.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ structcofoldF.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ structinteract.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ structintermediate__t.js - bin/
src/ , JavaScript, 12 linesViennaRNA-2.1.9/ doc/ html/ structmodel__detailsT.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ structpair__info.js - bin/
src/ , JavaScript, 5 linesViennaRNA-2.1.9/ doc/ html/ structparamT.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ structpf__paramT.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ structpu__contrib.js - bin/
src/ , JavaScript, 8 linesViennaRNA-2.1.9/ doc/ html/ structpu__out.js - bin/
src/ , JavaScript, 9 linesViennaRNA-2.1.9/ doc/ html/ subopt_8h.js - bin/
src/ , JavaScript, 319 linesViennaRNA-2.1.9/ doc/ html/ svgpan.js - bin/
src/ , JavaScript, 7 linesViennaRNA-2.1.9/ doc/ html/ treedist_8h.js - bin/
src/ , JavaScript, 65 linesViennaRNA-2.1.9/ doc/ html/ utils_8h.js - bin/
src/ , C/C++, 848 linesViennaRNA-2.1.9/ doc/ mainpage.h - bin/
src/ , Perl, 51 linesViennaRNA-2.1.9/ interfaces/ Perl/ Makefile.PL - bin/
src/ , C, 7,231 linesViennaRNA-2.1.9/ interfaces/ Perl/ RNA_wrap.c - bin/
src/ , Perl, 144 linesViennaRNA-2.1.9/ interfaces/ Perl/ RNAfold.pl - bin/
src/ , Perl, 166 linesViennaRNA-2.1.9/ interfaces/ Perl/ test.pl - bin/
src/ , C, 6,691 linesViennaRNA-2.1.9/ interfaces/ Python/ RNA_wrap.c - bin/
src/ , Python, 1,124 linesViennaRNA-2.1.9/ interfaces/ Python/ __init__.py - bin/
src/ , Python, 35 linesViennaRNA-2.1.9/ interfaces/ Python/ setup.py - bin/
src/ , Python, 10 linesViennaRNA-2.1.9/ interfaces/ Python/ version_test.py - bin/
src/ , C/C++, 366 linesViennaRNA-2.1.9/ lib/ 1.8.4_epars.h - bin/
src/ , C/C++, 2,992 linesViennaRNA-2.1.9/ lib/ 1.8.4_intloops.h - bin/
src/ , C, 4,024 linesViennaRNA-2.1.9/ lib/ 2Dfold.c - bin/
src/ , C, 4,489 linesViennaRNA-2.1.9/ lib/ 2Dpfold.c - bin/
src/ , C, 1,379 linesViennaRNA-2.1.9/ lib/ LPfold.c - bin/
src/ , C, 1,284 linesViennaRNA-2.1.9/ lib/ Lfold.c - bin/
src/ , C, 278 linesViennaRNA-2.1.9/ lib/ MEA.c - bin/
src/ , C, 2,277 linesViennaRNA-2.1.9/ lib/ PS_dot.c - bin/
src/ , C, 272 linesViennaRNA-2.1.9/ lib/ ProfileAln.c - bin/
src/ , C, 247 linesViennaRNA-2.1.9/ lib/ ProfileDist.c - bin/
src/ , C, 579 linesViennaRNA-2.1.9/ lib/ RNAstruct.c - bin/
src/ , C, 972 linesViennaRNA-2.1.9/ lib/ aliLfold.c - bin/
src/ , C, 1,294 linesViennaRNA-2.1.9/ lib/ ali_plex.c - bin/
src/ , C, 1,726 linesViennaRNA-2.1.9/ lib/ alifold.c - bin/
src/ , C, 1,290 linesViennaRNA-2.1.9/ lib/ alipfold.c - bin/
src/ , C, 171 linesViennaRNA-2.1.9/ lib/ aln_util.c - bin/
src/ , C, 1,172 linesViennaRNA-2.1.9/ lib/ c_plex.c - bin/
src/ , C, 1,510 linesViennaRNA-2.1.9/ lib/ cofold.c - bin/
src/ , C, 950 linesViennaRNA-2.1.9/ lib/ convert_epars.c - bin/
src/ , C, 9 linesViennaRNA-2.1.9/ lib/ dist_vars.c - bin/
src/ , C, 565 linesViennaRNA-2.1.9/ lib/ duplex.c - bin/
src/ , C, 843 linesViennaRNA-2.1.9/ lib/ energy_par.c - bin/
src/ , C, 543 linesViennaRNA-2.1.9/ lib/ findpath.c - bin/
src/ , C, 2,792 linesViennaRNA-2.1.9/ lib/ fold.c - bin/
src/ , C, 91 linesViennaRNA-2.1.9/ lib/ fold_vars.c - bin/
src/ , C, 1,051 linesViennaRNA-2.1.9/ lib/ gquad.c - bin/
src/ , C/C++, 393 linesViennaRNA-2.1.9/ lib/ intl11.h - bin/
src/ , C/C++, 393 linesViennaRNA-2.1.9/ lib/ intl11dH.h - bin/
src/ , C/C++, 1,993 linesViennaRNA-2.1.9/ lib/ intl21.h - bin/
src/ , C/C++, 1,993 linesViennaRNA-2.1.9/ lib/ intl21dH.h - bin/
src/ , C/C++, 5,749 linesViennaRNA-2.1.9/ lib/ intl22.h - bin/
src/ , C/C++, 5,749 linesViennaRNA-2.1.9/ lib/ intl22dH.h - bin/
src/ , C, 533 linesViennaRNA-2.1.9/ lib/ inverse.c - bin/
src/ , C, 411 linesViennaRNA-2.1.9/ lib/ list.c - bin/
src/ , C/C++, 65 linesViennaRNA-2.1.9/ lib/ list.h - bin/
src/ , C, 99 linesViennaRNA-2.1.9/ lib/ mm.c - bin/
src/ , C, 1,130 linesViennaRNA-2.1.9/ lib/ move_set.c - bin/
src/ , C, 1,167 linesViennaRNA-2.1.9/ lib/ naview.c - bin/
src/ , C, 595 linesViennaRNA-2.1.9/ lib/ params.c - bin/
src/ , C, 1,724 linesViennaRNA-2.1.9/ lib/ part_func.c - bin/
src/ , C, 1,053 linesViennaRNA-2.1.9/ lib/ part_func_co.c - bin/
src/ , C, 1,458 linesViennaRNA-2.1.9/ lib/ part_func_up.c - bin/
src/ , C, 2,853 linesViennaRNA-2.1.9/ lib/ plex.c - bin/
src/ , C, 301 linesViennaRNA-2.1.9/ lib/ plex_functions.c - bin/
src/ , C, 157 linesViennaRNA-2.1.9/ lib/ plot_layouts.c - bin/
src/ , C, 1,075 linesViennaRNA-2.1.9/ lib/ read_epars.c - bin/
src/ , C, 1,104 linesViennaRNA-2.1.9/ lib/ ribo.c - bin/
src/ , C, 1,287 linesViennaRNA-2.1.9/ lib/ snofold.c - bin/
src/ , C, 2,586 linesViennaRNA-2.1.9/ lib/ snoop.c - bin/
src/ , C, 433 linesViennaRNA-2.1.9/ lib/ stringdist.c - bin/
src/ , C, 1,510 linesViennaRNA-2.1.9/ lib/ subopt.c - bin/
src/ , C, 474 linesViennaRNA-2.1.9/ lib/ svm_utils.c - bin/
src/ , C, 657 linesViennaRNA-2.1.9/ lib/ treedist.c - bin/
src/ , C, 1,154 linesViennaRNA-2.1.9/ lib/ utils.c - bin/
src/ , C++, 3,069 linesViennaRNA-2.1.9/ libsvm-2.91/ svm.cpp - bin/
src/ , C/C++, 76 linesViennaRNA-2.1.9/ libsvm-2.91/ svm.h - bin/
src/ , Shell, 15 linesViennaRNA-2.1.9/ man/ cmdlopt.sh - bin/
src/ , C, 402 linesbwa-0.7.15/ QSufSort.c - bin/
src/ , C/C++, 45 linesbwa-0.7.15/ QSufSort.h - bin/
src/ , C, 210 linesbwa-0.7.15/ bamlite.c - bin/
src/ , C/C++, 114 linesbwa-0.7.15/ bamlite.h - bin/
src/ , C, 446 linesbwa-0.7.15/ bntseq.c - bin/
src/ , C/C++, 92 linesbwa-0.7.15/ bntseq.h - bin/
src/ , C, 450 linesbwa-0.7.15/ bwa.c - bin/
src/ , C/C++, 69 linesbwa-0.7.15/ bwa.h - bin/
src/ , JavaScript, 524 linesbwa-0.7.15/ bwakit/ bwa-postalt.js - bin/
src/ , JavaScript, 62 linesbwa-0.7.15/ bwakit/ typeHLA-selctg.js - bin/
src/ , JavaScript, 496 linesbwa-0.7.15/ bwakit/ typeHLA.js - bin/
src/ , Shell, 49 linesbwa-0.7.15/ bwakit/ typeHLA.sh - bin/
src/ , C, 1,201 linesbwa-0.7.15/ bwamem.c - bin/
src/ , C/C++, 184 linesbwa-0.7.15/ bwamem.h - bin/
src/ , C, 140 linesbwa-0.7.15/ bwamem_extra.c - bin/
src/ , C, 388 linesbwa-0.7.15/ bwamem_pair.c - bin/
src/ , C, 783 linesbwa-0.7.15/ bwape.c - bin/
src/ , C, 603 linesbwa-0.7.15/ bwase.c - bin/
src/ , C/C++, 29 linesbwa-0.7.15/ bwase.h - bin/
src/ , C, 235 linesbwa-0.7.15/ bwaseqio.c - bin/
src/ , C, 213 linesbwa-0.7.15/ bwashm.c - bin/
src/ , C, 469 linesbwa-0.7.15/ bwt.c - bin/
src/ , C/C++, 130 linesbwa-0.7.15/ bwt.h - bin/
src/ , C, 1,632 linesbwa-0.7.15/ bwt_gen.c - bin/
src/ , C, 98 linesbwa-0.7.15/ bwt_lite.c - bin/
src/ , C/C++, 29 linesbwa-0.7.15/ bwt_lite.h - bin/
src/ , C, 320 linesbwa-0.7.15/ bwtaln.c - bin/
src/ , C/C++, 153 linesbwa-0.7.15/ bwtaln.h - bin/
src/ , C, 264 linesbwa-0.7.15/ bwtgap.c - bin/
src/ , C/C++, 40 linesbwa-0.7.15/ bwtgap.h - bin/
src/ , C, 324 linesbwa-0.7.15/ bwtindex.c - bin/
src/ , C/C++, 69 linesbwa-0.7.15/ bwtsw2.h - bin/
src/ , C, 776 linesbwa-0.7.15/ bwtsw2_aux.c - bin/
src/ , C, 112 linesbwa-0.7.15/ bwtsw2_chain.c - bin/
src/ , C, 619 linesbwa-0.7.15/ bwtsw2_core.c - bin/
src/ , C, 89 linesbwa-0.7.15/ bwtsw2_main.c - bin/
src/ , C, 268 linesbwa-0.7.15/ bwtsw2_pair.c - bin/
src/ , C, 60 linesbwa-0.7.15/ example.c - bin/
src/ , C, 441 linesbwa-0.7.15/ fastmap.c - bin/
src/ , C, 223 linesbwa-0.7.15/ is.c - bin/
src/ , C/C++, 388 linesbwa-0.7.15/ kbtree.h - bin/
src/ , C/C++, 614 linesbwa-0.7.15/ khash.h - bin/
src/ , C, 374 linesbwa-0.7.15/ kopen.c - bin/
src/ , C/C++, 239 linesbwa-0.7.15/ kseq.h - bin/
src/ , C/C++, 273 linesbwa-0.7.15/ ksort.h - bin/
src/ , C, 39 linesbwa-0.7.15/ kstring.c - bin/
src/ , C/C++, 115 linesbwa-0.7.15/ kstring.h - bin/
src/ , C, 713 linesbwa-0.7.15/ ksw.c - bin/
src/ , C/C++, 114 linesbwa-0.7.15/ ksw.h - bin/
src/ , C, 146 linesbwa-0.7.15/ kthread.c - bin/
src/ , C/C++, 94 linesbwa-0.7.15/ kvec.h - bin/
src/ , C, 104 linesbwa-0.7.15/ main.c - bin/
src/ , C, 57 linesbwa-0.7.15/ malloc_wrap.c - bin/
src/ , C/C++, 47 linesbwa-0.7.15/ malloc_wrap.h - bin/
src/ , C, 67 linesbwa-0.7.15/ maxk.c - bin/
src/ , C, 291 linesbwa-0.7.15/ pemerge.c - bin/
src/ , Perl, 27 linesbwa-0.7.15/ qualfa2fq.pl - bin/
src/ , C, 191 linesbwa-0.7.15/ rle.c - bin/
src/ , C/C++, 77 linesbwa-0.7.15/ rle.h - bin/
src/ , C, 318 linesbwa-0.7.15/ rope.c - bin/
src/ , C/C++, 58 linesbwa-0.7.15/ rope.h - bin/
src/ , C, 295 linesbwa-0.7.15/ utils.c - bin/
src/ , C/C++, 111 linesbwa-0.7.15/ utils.h - bin/
src/ , Perl, 25 linesbwa-0.7.15/ xa2multi.pl - bin/
src/ , Python, 712 linescctop_standalone/ CCTop.py - bin/
src/ , Python, 86 linescctop_standalone/ bedInterval.py - bin/
src/ , Shell, 54 linescufflinks-2.2.1/ make_bin.sh - bin/
src/ , C++, 376 linescufflinks-2.2.1/ src/ GArgs.cpp - bin/
src/ , C/C++, 98 linescufflinks-2.2.1/ src/ GArgs.h - bin/
src/ , C++, 780 linescufflinks-2.2.1/ src/ GBase.cpp - bin/
src/ , C/C++, 458 linescufflinks-2.2.1/ src/ GBase.h - bin/
src/ , C++, 319 linescufflinks-2.2.1/ src/ GFaSeqGet.cpp - bin/
src/ , C/C++, 136 linescufflinks-2.2.1/ src/ GFaSeqGet.h - bin/
src/ , C++, 170 linescufflinks-2.2.1/ src/ GFastaIndex.cpp - bin/
src/ , C/C++, 79 linescufflinks-2.2.1/ src/ GFastaIndex.h - bin/
src/ , C++, 1,354 linescufflinks-2.2.1/ src/ GStr.cpp - bin/
src/ , C/C++, 223 linescufflinks-2.2.1/ src/ GStr.h - bin/
src/ , C++, 5,165 linescufflinks-2.2.1/ src/ abundances.cpp - bin/
src/ , C/C++, 806 linescufflinks-2.2.1/ src/ abundances.h - bin/
src/ , C++, 568 linescufflinks-2.2.1/ src/ assemble.cpp - bin/
src/ , C/C++, 44 linescufflinks-2.2.1/ src/ assemble.h - bin/
src/ , C++, 835 linescufflinks-2.2.1/ src/ biascorrection.cpp - bin/
src/ , C/C++, 118 linescufflinks-2.2.1/ src/ biascorrection.h - bin/
src/ , C++, 2,104 linescufflinks-2.2.1/ src/ bundles.cpp - bin/
src/ , C/C++, 405 linescufflinks-2.2.1/ src/ bundles.h - bin/
src/ , C++, 130 linescufflinks-2.2.1/ src/ clustering.cpp - bin/
src/ , C/C++, 168 linescufflinks-2.2.1/ src/ clustering.h - bin/
src/ , C++, 90 linescufflinks-2.2.1/ src/ codons.cpp - bin/
src/ , C/C++, 54 linescufflinks-2.2.1/ src/ codons.h - bin/
src/ , C++, 507 linescufflinks-2.2.1/ src/ common.cpp - bin/
src/ , C/C++, 817 linescufflinks-2.2.1/ src/ common.h - bin/
src/ , C++, 428 linescufflinks-2.2.1/ src/ compress_gtf.cpp - bin/
src/ , C++, 2,684 linescufflinks-2.2.1/ src/ cuffcompare.cpp - bin/
src/ , C++, 2,721 linescufflinks-2.2.1/ src/ cuffdiff.cpp - bin/
src/ , C++, 1,802 linescufflinks-2.2.1/ src/ cufflinks.cpp - bin/
src/ , C++, 1,995 linescufflinks-2.2.1/ src/ cuffnorm.cpp - bin/
src/ , C++, 1,623 linescufflinks-2.2.1/ src/ cuffquant.cpp - bin/
src/ , C++, 1,705 linescufflinks-2.2.1/ src/ differential.cpp - bin/
src/ , C/C++, 217 linescufflinks-2.2.1/ src/ differential.h - bin/
src/ , C++, 1,151 linescufflinks-2.2.1/ src/ filters.cpp - bin/
src/ , C/C++, 42 linescufflinks-2.2.1/ src/ filters.h - bin/
src/ , C++, 90 linescufflinks-2.2.1/ src/ gdna.cpp - bin/
src/ , C/C++, 15 linescufflinks-2.2.1/ src/ gdna.h - bin/
src/ , C++, 155 linescufflinks-2.2.1/ src/ genes.cpp - bin/
src/ , C/C++, 217 linescufflinks-2.2.1/ src/ genes.h - bin/
src/ , C++, 2,124 linescufflinks-2.2.1/ src/ gff.cpp - bin/
src/ , C/C++, 1,094 linescufflinks-2.2.1/ src/ gff.h - bin/
src/ , C++, 676 linescufflinks-2.2.1/ src/ gff_utils.cpp - bin/
src/ , C/C++, 610 linescufflinks-2.2.1/ src/ gff_utils.h - bin/
src/ , C++, 1,060 linescufflinks-2.2.1/ src/ gffread.cpp - bin/
src/ , C++, 751 linescufflinks-2.2.1/ src/ graph_optimize.cpp - bin/
src/ , C/C++, 94 linescufflinks-2.2.1/ src/ graph_optimize.h - bin/
src/ , C++, 349 linescufflinks-2.2.1/ src/ gtf_to_sam.cpp - bin/
src/ , C++, 778 linescufflinks-2.2.1/ src/ gtf_tracking.cpp - bin/
src/ , C/C++, 1,347 linescufflinks-2.2.1/ src/ gtf_tracking.h - bin/
src/ , C++, 1,276 linescufflinks-2.2.1/ src/ hits.cpp - bin/
src/ , C/C++, 1,188 linescufflinks-2.2.1/ src/ hits.h - bin/
src/ , C++, 171 linescufflinks-2.2.1/ src/ jensen_shannon.cpp - bin/
src/ , C/C++, 30 linescufflinks-2.2.1/ src/ jensen_shannon.h - bin/
src/ , C/C++, 1,597 linescufflinks-2.2.1/ src/ lemon/ bfs.h - bin/
src/ , C/C++, 346 linescufflinks-2.2.1/ src/ lemon/ bin_heap.h - bin/
src/ , C/C++, 1,732 linescufflinks-2.2.1/ src/ lemon/ bipartite_matching.h - bin/
src/ , C/C++, 485 linescufflinks-2.2.1/ src/ lemon/ bits/ alteration_notifier.h - bin/
src/ , C/C++, 346 linescufflinks-2.2.1/ src/ lemon/ bits/ array_map.h - bin/
src/ , C/C++, 495 linescufflinks-2.2.1/ src/ lemon/ bits/ base_extender.h - bin/
src/ , C/C++, 382 linescufflinks-2.2.1/ src/ lemon/ bits/ debug_map.h - bin/
src/ , C/C++, 181 linescufflinks-2.2.1/ src/ lemon/ bits/ default_map.h - bin/
src/ , C/C++, 742 linescufflinks-2.2.1/ src/ lemon/ bits/ graph_adaptor_extender.h - bin/
src/ , C/C++, 1,397 linescufflinks-2.2.1/ src/ lemon/ bits/ graph_extender.h - bin/
src/ , C/C++, 54 linescufflinks-2.2.1/ src/ lemon/ bits/ invalid.h - bin/
src/ , C/C++, 321 linescufflinks-2.2.1/ src/ lemon/ bits/ map_extender.h - bin/
src/ , C/C++, 174 linescufflinks-2.2.1/ src/ lemon/ bits/ path_dump.h - bin/
src/ , C/C++, 346 linescufflinks-2.2.1/ src/ lemon/ bits/ traits.h - bin/
src/ , C/C++, 140 linescufflinks-2.2.1/ src/ lemon/ bits/ utility.h - bin/
src/ , C/C++, 508 linescufflinks-2.2.1/ src/ lemon/ bits/ variant.h - bin/
src/ , C/C++, 243 linescufflinks-2.2.1/ src/ lemon/ bits/ vector_map.h - bin/
src/ , C/C++, 831 linescufflinks-2.2.1/ src/ lemon/ bucket_heap.h - bin/
src/ , C/C++, 105 linescufflinks-2.2.1/ src/ lemon/ concept_check.h - bin/
src/ , C/C++, 1,004 linescufflinks-2.2.1/ src/ lemon/ concepts/ bpugraph.h - bin/
src/ , C/C++, 453 linescufflinks-2.2.1/ src/ lemon/ concepts/ graph.h - bin/
src/ , C/C++, 2,093 linescufflinks-2.2.1/ src/ lemon/ concepts/ graph_components.h - bin/
src/ , C/C++, 226 linescufflinks-2.2.1/ src/ lemon/ concepts/ heap.h - bin/
src/ , C/C++, 208 linescufflinks-2.2.1/ src/ lemon/ concepts/ maps.h - bin/
src/ , C/C++, 207 linescufflinks-2.2.1/ src/ lemon/ concepts/ matrix_maps.h - bin/
src/ , C/C++, 307 linescufflinks-2.2.1/ src/ lemon/ concepts/ path.h - bin/
src/ , C/C++, 702 linescufflinks-2.2.1/ src/ lemon/ concepts/ ugraph.h - bin/
src/ , C/C++, 1,543 linescufflinks-2.2.1/ src/ lemon/ dfs.h - bin/
src/ , C/C++, 683 linescufflinks-2.2.1/ src/ lemon/ error.h - bin/
src/ , C/C++, 464 linescufflinks-2.2.1/ src/ lemon/ fib_heap.h - bin/
src/ , C/C++, 2,720 linescufflinks-2.2.1/ src/ lemon/ graph_adaptor.h - bin/
src/ , C/C++, 3,179 linescufflinks-2.2.1/ src/ lemon/ graph_utils.h - bin/
src/ , C/C++, 2,249 linescufflinks-2.2.1/ src/ lemon/ list_graph.h - bin/
src/ , C/C++, 1,633 linescufflinks-2.2.1/ src/ lemon/ maps.h - bin/
src/ , C/C++, 63 linescufflinks-2.2.1/ src/ lemon/ math.h - bin/
src/ , C/C++, 1,163 linescufflinks-2.2.1/ src/ lemon/ smart_graph.h - bin/
src/ , C/C++, 454 linescufflinks-2.2.1/ src/ lemon/ tolerance.h - bin/
src/ , C/C++, 1,590 linescufflinks-2.2.1/ src/ lemon/ topology.h - bin/
src/ , C, 195 linescufflinks-2.2.1/ src/ locfit/ adap.c - bin/
src/ , C, 105 linescufflinks-2.2.1/ src/ locfit/ ar_funs.c - bin/
src/ , C, 619 linescufflinks-2.2.1/ src/ locfit/ arith.c - bin/
src/ , C, 383 linescufflinks-2.2.1/ src/ locfit/ band.c - bin/
src/ , C, 89 linescufflinks-2.2.1/ src/ locfit/ c_args.c - bin/
src/ , C, 471 linescufflinks-2.2.1/ src/ locfit/ c_plot.c - bin/
src/ , C, 801 linescufflinks-2.2.1/ src/ locfit/ cmd.c - bin/
src/ , C, 190 linescufflinks-2.2.1/ src/ locfit/ dens_haz.c - bin/
src/ , C, 223 linescufflinks-2.2.1/ src/ locfit/ dens_int.c - bin/
src/ , C, 515 linescufflinks-2.2.1/ src/ locfit/ dens_odi.c - bin/
src/ , C, 509 linescufflinks-2.2.1/ src/ locfit/ density.c - bin/
src/ , C/C++, 28 linescufflinks-2.2.1/ src/ locfit/ design.h - bin/
src/ , C, 162 linescufflinks-2.2.1/ src/ locfit/ dist.c - bin/
src/ , C, 204 linescufflinks-2.2.1/ src/ locfit/ ev_atree.c - bin/
src/ , C, 273 linescufflinks-2.2.1/ src/ locfit/ ev_interp.c - bin/
src/ , C, 332 linescufflinks-2.2.1/ src/ locfit/ ev_kdtre.c - bin/
src/ , C, 287 linescufflinks-2.2.1/ src/ locfit/ ev_main.c - bin/
src/ , C, 457 linescufflinks-2.2.1/ src/ locfit/ ev_trian.c - bin/
src/ , C, 607 linescufflinks-2.2.1/ src/ locfit/ family.c - bin/
src/ , C, 154 linescufflinks-2.2.1/ src/ locfit/ fitted.c - bin/
src/ , C, 409 linescufflinks-2.2.1/ src/ locfit/ frend.c - bin/
src/ , C, 88 linescufflinks-2.2.1/ src/ locfit/ help.c - bin/
src/ , C/C++, 36 linescufflinks-2.2.1/ src/ locfit/ imatlb.h - bin/
src/ , C, 51 linescufflinks-2.2.1/ src/ locfit/ lf_dercor.c - bin/
src/ , C, 258 linescufflinks-2.2.1/ src/ locfit/ lf_fitfun.c - bin/
src/ , C, 127 linescufflinks-2.2.1/ src/ locfit/ lf_robust.c - bin/
src/ , C, 157 linescufflinks-2.2.1/ src/ locfit/ lf_vari.c - bin/
src/ , C/C++, 280 linescufflinks-2.2.1/ src/ locfit/ lfcons.h - bin/
src/ , C, 319 linescufflinks-2.2.1/ src/ locfit/ lfd.c - bin/
src/ , C/C++, 215 linescufflinks-2.2.1/ src/ locfit/ lffuns.h - bin/
src/ , C, 150 linescufflinks-2.2.1/ src/ locfit/ lfstr.c - bin/
src/ , C/C++, 102 linescufflinks-2.2.1/ src/ locfit/ lfstruc.h - bin/
src/ , C/C++, 117 linescufflinks-2.2.1/ src/ locfit/ lfwin.h - bin/
src/ , C, 337 linescufflinks-2.2.1/ src/ locfit/ linalg.c - bin/
src/ , C/C++, 148 linescufflinks-2.2.1/ src/ locfit/ local.h - bin/
src/ , C, 263 linescufflinks-2.2.1/ src/ locfit/ locfit.c - bin/
src/ , C, 72 linescufflinks-2.2.1/ src/ locfit/ m_chol.c - bin/
src/ , C, 141 linescufflinks-2.2.1/ src/ locfit/ m_eigen.c - bin/
src/ , C, 118 linescufflinks-2.2.1/ src/ locfit/ m_jacob.c - bin/
src/ , C, 215 linescufflinks-2.2.1/ src/ locfit/ m_max.c - bin/
src/ , C, 256 linescufflinks-2.2.1/ src/ locfit/ makecmd.c - bin/
src/ , C, 158 linescufflinks-2.2.1/ src/ locfit/ math.c - bin/
src/ , C, 302 linescufflinks-2.2.1/ src/ locfit/ minmax.c - bin/
src/ , C/C++, 56 linescufflinks-2.2.1/ src/ locfit/ mutil.h - bin/
src/ , C, 226 linescufflinks-2.2.1/ src/ locfit/ nbhd.c - bin/
src/ , C, 191 linescufflinks-2.2.1/ src/ locfit/ pcomp.c - bin/
src/ , C, 780 linescufflinks-2.2.1/ src/ locfit/ pout.c - bin/
src/ , C, 222 linescufflinks-2.2.1/ src/ locfit/ preplot.c - bin/
src/ , C, 147 linescufflinks-2.2.1/ src/ locfit/ random.c - bin/
src/ , C, 81 linescufflinks-2.2.1/ src/ locfit/ readfile.c - bin/
src/ , C, 326 linescufflinks-2.2.1/ src/ locfit/ scb.c - bin/
src/ , C, 342 linescufflinks-2.2.1/ src/ locfit/ scb_cons.c - bin/
src/ , C, 219 linescufflinks-2.2.1/ src/ locfit/ simul.c - bin/
src/ , C, 119 linescufflinks-2.2.1/ src/ locfit/ solve.c - bin/
src/ , C, 447 linescufflinks-2.2.1/ src/ locfit/ startlf.c - bin/
src/ , C, 140 linescufflinks-2.2.1/ src/ locfit/ strings.c - bin/
src/ , C++, 299 linescufflinks-2.2.1/ src/ locfit/ vari.cpp - bin/
src/ , C, 260 linescufflinks-2.2.1/ src/ locfit/ wdiag.c - bin/
src/ , C, 462 linescufflinks-2.2.1/ src/ locfit/ weight.c - bin/
src/ , C++, 103 linescufflinks-2.2.1/ src/ matching_merge.cpp - bin/
src/ , C/C++, 113 linescufflinks-2.2.1/ src/ matching_merge.h - bin/
src/ , C++, 132 linescufflinks-2.2.1/ src/ multireads.cpp - bin/
src/ , C/C++, 72 linescufflinks-2.2.1/ src/ multireads.h - bin/
src/ , C/C++, 223 linescufflinks-2.2.1/ src/ negative_binomial_distri bution.h - bin/
src/ , C/C++, 110 linescufflinks-2.2.1/ src/ progressbar.h - bin/
src/ , C++, 1,147 linescufflinks-2.2.1/ src/ replicates.cpp - bin/
src/ , C/C++, 508 linescufflinks-2.2.1/ src/ replicates.h - bin/
src/ , C/C++, 191 linescufflinks-2.2.1/ src/ rounding.h - bin/
src/ , C++, 69 linescufflinks-2.2.1/ src/ sampling.cpp - bin/
src/ , C/C++, 307 linescufflinks-2.2.1/ src/ sampling.h - bin/
src/ , C++, 304 linescufflinks-2.2.1/ src/ scaffold_graph.cpp - bin/
src/ , C/C++, 39 linescufflinks-2.2.1/ src/ scaffold_graph.h - bin/
src/ , C++, 1,837 linescufflinks-2.2.1/ src/ scaffolds.cpp - bin/
src/ , C/C++, 708 linescufflinks-2.2.1/ src/ scaffolds.h - bin/
src/ , C++, 39 linescufflinks-2.2.1/ src/ tokenize.cpp - bin/
src/ , C/C++, 18 linescufflinks-2.2.1/ src/ tokenize.h - bin/
src/ , C++, 99 linescufflinks-2.2.1/ src/ tracking.cpp - bin/
src/ , C/C++, 118 linescufflinks-2.2.1/ src/ tracking.h - bin/
src/ , C/C++, 391 linescufflinks-2.2.1/ src/ transitive_closure.h - bin/
src/ , C/C++, 129 linescufflinks-2.2.1/ src/ transitive_reduction.h - bin/
src/ , C/C++, 113 linescufflinks-2.2.1/ src/ update_check.h - bin/
src/ , C, 177 lineslibsvm-260/ svm-predict.c - bin/
src/ , C, 290 lineslibsvm-260/ svm-scale.c - bin/
src/ , C, 299 lineslibsvm-260/ svm-train.c - bin/
src/ , C++, 2,771 lineslibsvm-260/ svm.cpp - bin/
src/ , C/C++, 70 lineslibsvm-260/ svm.h - bin/
src/ , Python, 284 lineslindel/ Lindel/ Predictor.py - bin/
src/ , Python, 1 linelindel/ Lindel/ __init__.py - bin/
src/ , Python, 30 lineslindel/ Lindel_prediction.py - bin/
src/ , Shell, 1 linelindel/ sampleRun.sh - bin/
src/ , Python, 24 lineslindel/ setup.py - bin/
src/ , Python, 53 linespairwise-library-screen/ encode_full_matrix.py - bin/
src/ , Stan, 44 linespairwise-library-screen/ exp_model.stan - bin/
src/ , Python, 30 linespairwise-library-screen/ filter_dataset.py - bin/
src/ , Python, 120 linespairwise-library-screen/ pairwise-library-analysi s/ PWL_library_features_ann otate.py - bin/
src/ , Python, 78 linespairwise-library-screen/ pairwise-library-analysi s/ PWL_make_dataframe.py - bin/
src/ , Python, 460 linespairwise-library-screen/ pairwise-library-analysi s/ PWL_screen_analysis.py - bin/
src/ , Python, 74 linespairwise-library-screen/ predictSingle.py - bin/
src/ , Python, 92 linespairwise-library-screen/ predict_activity_single. py - bin/
src/ , R, 75 linespairwise-library-screen/ train_rstan_model.R - bin/
src/ , Perl, 113 linesprimer3-2.3.6/ cmp_settings.pl - bin/
src/ , C, 1,166 linesprimer3-2.3.6/ src/ dpal.c - bin/
src/ , C/C++, 167 linesprimer3-2.3.6/ src/ dpal.h - bin/
src/ , C, 950 linesprimer3-2.3.6/ src/ format_output.c - bin/
src/ , C/C++, 61 linesprimer3-2.3.6/ src/ format_output.h - bin/
src/ , C, 5,934 linesprimer3-2.3.6/ src/ libprimer3.c - bin/
src/ , C/C++, 1,275 linesprimer3-2.3.6/ src/ libprimer3.h - bin/
src/ , C, 88 linesprimer3-2.3.6/ src/ long_seq_tm_test_main.c - bin/
src/ , C, 200 linesprimer3-2.3.6/ src/ ntdpal_main.c - bin/
src/ , C, 708 linesprimer3-2.3.6/ src/ oligotm.c - bin/
src/ , C/C++, 203 linesprimer3-2.3.6/ src/ oligotm.h - bin/
src/ , C, 226 linesprimer3-2.3.6/ src/ oligotm_main.c - bin/
src/ , C, 473 linesprimer3-2.3.6/ src/ p3_seq_lib.c - bin/
src/ , C/C++, 93 linesprimer3-2.3.6/ src/ p3_seq_lib.h - bin/
src/ , Shell, 14 linesprimer3-2.3.6/ src/ paranoid_tests.sh - bin/
src/ , C, 565 linesprimer3-2.3.6/ src/ primer3_boulder_main.c - bin/
src/ , C, 536 linesprimer3-2.3.6/ src/ print_boulder.c - bin/
src/ , C/C++, 51 linesprimer3-2.3.6/ src/ print_boulder.h - bin/
src/ , C, 1,302 linesprimer3-2.3.6/ src/ read_boulder.c - bin/
src/ , C/C++, 89 linesprimer3-2.3.6/ src/ read_boulder.h - bin/
src/ , C, 2,829 linesprimer3-2.3.6/ src/ thal.c - bin/
src/ , C/C++, 150 linesprimer3-2.3.6/ src/ thal.h - bin/
src/ , C, 339 linesprimer3-2.3.6/ src/ thal_main.c - bin/
src/ , Perl, 357 linesprimer3-2.3.6/ test/ cmdline_test.pl - bin/
src/ , Perl, 278 linesprimer3-2.3.6/ test/ dpal_test.pl - bin/
src/ , Perl, 17 linesprimer3-2.3.6/ test/ long_seq_tm_test.pl - bin/
src/ , Perl, 133 linesprimer3-2.3.6/ test/ oligotm_test.pl - bin/
src/ , Perl, 718 linesprimer3-2.3.6/ test/ p3test.pl - bin/
src/ , Perl, 231 linesprimer3-2.3.6/ test/ primer3_io_translator.pl - bin/
src/ , Perl, 291 linesprimer3-2.3.6/ test/ thal_test.pl - bin/
src/ , Perl, 9 linesprimer3-2.3.6/ test/ vgrep.pl - bin/
src/ , Python, 48 linessgRNA.Scorer.1.0/ generateSVMFile.FASTA.py - bin/
src/ , Python, 114 linessgRNA.Scorer.1.0/ identifyPutativegRNASite s.py - bin/
src/ , Python, 64 linessgRNA.Scorer.1.0/ makeFinalTable.py - bin/
src/ , Python, 100 linessgRNA.Scorer.1.0/ scoreMySites.py - bin/
src/ , Python, 16 linessgRNA.Scorer.1.0/ simplifyRankPerc.py - bin/
src/ , C/C++, 40 linessvm_light/ kernel.h - bin/
src/ , C, 197 linessvm_light/ svm_classify.c - bin/
src/ , C, 985 linessvm_light/ svm_common.c - bin/
src/ , C/C++, 301 linessvm_light/ svm_common.h - bin/
src/ , C, 1,062 linessvm_light/ svm_hideo.c - bin/
src/ , C, 4,147 linessvm_light/ svm_learn.c - bin/
src/ , C/C++, 169 linessvm_light/ svm_learn.h - bin/
src/ , C, 397 linessvm_light/ svm_learn_main.c - bin/
src/ , C, 211 linessvm_light/ svm_loqo.c - bin/
src/ , R, 16 lineswangSabatiniSvm/ scorer.R - bin/
twobitreader/ , Python, 698 lines__init__.py - bin/
twobitreader/ , Python, 3 lines__main__.py - bin/
twobitreader/ , Python, 50 linesdownload.py - crispor.py, Python, 4,761 lines, 1 match
- crisporEffScores.py, Python, 1,359 lines
- doenchScore.py, Python, 54 lines
- js/
jquery-ui.min.js , JavaScript, 12 lines, 1 match - js/
jquery.min.js , JavaScript, 4 lines - js/
jquery.tooltipster.min.j , JavaScript, 1 lines - js/
jquery.ui.ufd.js , JavaScript, 1,320 lines - microHomScore.py, Python, 71 lines
- scripts/
xa2multi.pl , Perl, 25 lines - startWorkers.sh, Shell, 18 lines
- stopWorkers.sh, Shell, 6 lines
- tools/
usrLocalBin/ , Python, 474 linesbed.py - tools/
usrLocalBin/ , Python, 473 linestabfile.py - LICENSE.txt, License, 42 lines
- README.md, Text, 53 lines
usadellab/trimmomatic
ef98d6252abeae80cfee36acf9e0e1055097da0b, 3 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
152 files
- src/
main/ , Java, 181 linesjava/ org/ usadellab/ trimmomatic/ Pairomatic.java - src/
main/ , Java, 112 linesjava/ org/ usadellab/ trimmomatic/ TrimStats.java - src/
main/ , Java, 153 linesjava/ org/ usadellab/ trimmomatic/ Trimmomatic.java - src/
main/ , Java, 412 linesjava/ org/ usadellab/ trimmomatic/ TrimmomaticPE.java - src/
main/ , Java, 234 linesjava/ org/ usadellab/ trimmomatic/ TrimmomaticSE.java - src/
main/ , Java, 66 linesjava/ org/ usadellab/ trimmomatic/ fasta/ FastaParser.java - src/
main/ , Java, 95 linesjava/ org/ usadellab/ trimmomatic/ fasta/ FastaRecord.java - src/
main/ , Java, 46 linesjava/ org/ usadellab/ trimmomatic/ fasta/ FastaSerializer.java - src/
main/ , Java, 165 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParser.java - src/
main/ , Java, 230 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqRecord.java - src/
main/ , Java, 54 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqSerializer.java - src/
main/ , Java, 112 linesjava/ org/ usadellab/ trimmomatic/ fastq/ PairingValidator.java - src/
main/ , Java, 79 linesjava/ org/ usadellab/ trimmomatic/ fastq/ RecordNamePattern.java - src/
main/ , Java, 80 linesjava/ org/ usadellab/ trimmomatic/ threading/ BlockOfRecords.java - src/
main/ , Java, 229 linesjava/ org/ usadellab/ trimmomatic/ threading/ BlockOfWork.java - src/
main/ , Java, 23 linesjava/ org/ usadellab/ trimmomatic/ threading/ ExceptionHolder.java - src/
main/ , Java, 21 linesjava/ org/ usadellab/ trimmomatic/ threading/ parser/ ParasiteSerialParser.jav a - src/
main/ , Java, 59 linesjava/ org/ usadellab/ trimmomatic/ threading/ parser/ Parser.java - src/
main/ , Java, 69 linesjava/ org/ usadellab/ trimmomatic/ threading/ parser/ SelfThreadedParser.java - src/
main/ , Java, 20 linesjava/ org/ usadellab/ trimmomatic/ threading/ pipeline/ ParasiteSerialPipeline.j ava - src/
main/ , Java, 21 linesjava/ org/ usadellab/ trimmomatic/ threading/ pipeline/ Pipeline.java - src/
main/ , Java, 57 linesjava/ org/ usadellab/ trimmomatic/ threading/ pipeline/ ThreadedPipeline.java - src/
main/ , Java, 43 linesjava/ org/ usadellab/ trimmomatic/ threading/ serializer/ ParasiteSerializer.java - src/
main/ , Java, 55 linesjava/ org/ usadellab/ trimmomatic/ threading/ serializer/ SelfThreadedSerializer.j ava - src/
main/ , Java, 47 linesjava/ org/ usadellab/ trimmomatic/ threading/ serializer/ SerializedBlock.java - src/
main/ , Java, 117 linesjava/ org/ usadellab/ trimmomatic/ threading/ serializer/ SerializedBlockQueue.jav a - src/
main/ , Java, 86 linesjava/ org/ usadellab/ trimmomatic/ threading/ serializer/ Serializer.java - src/
main/ , Java, 27 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimlog/ ParasiteTrimLogCollector .java - src/
main/ , Java, 67 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimlog/ SelfThreadedTrimLogColle ctor.java - src/
main/ , Java, 54 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimlog/ TrimLogCollector.java - src/
main/ , Java, 3 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimlog/ TrimLogRecord.java - src/
main/ , Java, 29 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimstats/ ParasiteTrimStatsCollect or.java - src/
main/ , Java, 69 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimstats/ SelfThreadedTrimStatsCol lector.java - src/
main/ , Java, 39 linesjava/ org/ usadellab/ trimmomatic/ threading/ trimstats/ TrimStatsCollector.java - src/
main/ , Java, 24 linesjava/ org/ usadellab/ trimmomatic/ trim/ AbstractSingleRecordTrim mer.java - src/
main/ , Java, 32 linesjava/ org/ usadellab/ trimmomatic/ trim/ AvgQualTrimmer.java - src/
main/ , Java, 110 linesjava/ org/ usadellab/ trimmomatic/ trim/ BarcodeSplitter.java - src/
main/ , Java, 51 linesjava/ org/ usadellab/ trimmomatic/ trim/ BaseCountTrimmer.java - src/
main/ , Java, 35 linesjava/ org/ usadellab/ trimmomatic/ trim/ CropTrimmer.java - src/
main/ , Java, 95 linesjava/ org/ usadellab/ trimmomatic/ trim/ HeadCropTrimmer.java - src/
main/ , Java, 968 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmer. java - src/
main/ , Java, 29 linesjava/ org/ usadellab/ trimmomatic/ trim/ LeadingTrimmer.java - src/
main/ , Java, 24 linesjava/ org/ usadellab/ trimmomatic/ trim/ MaxLenTrimmer.java - src/
main/ , Java, 111 linesjava/ org/ usadellab/ trimmomatic/ trim/ MaximumInformationTrimme r.java - src/
main/ , Java, 24 linesjava/ org/ usadellab/ trimmomatic/ trim/ MinLenTrimmer.java - src/
main/ , Java, 69 linesjava/ org/ usadellab/ trimmomatic/ trim/ SlidingWindowTrimmer.jav a - src/
main/ , Java, 94 linesjava/ org/ usadellab/ trimmomatic/ trim/ TailCropTrimmer.java - src/
main/ , Java, 29 linesjava/ org/ usadellab/ trimmomatic/ trim/ ToPhred33Trimmer.java - src/
main/ , Java, 29 linesjava/ org/ usadellab/ trimmomatic/ trim/ ToPhred64Trimmer.java - src/
main/ , Java, 28 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrailingTrimmer.java - src/
main/ , Java, 7 linesjava/ org/ usadellab/ trimmomatic/ trim/ Trimmer.java - src/
main/ , Java, 70 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactory.java - src/
main/ , Java, 61 linesjava/ org/ usadellab/ trimmomatic/ util/ Logger.java - src/
main/ , Java, 100 linesjava/ org/ usadellab/ trimmomatic/ util/ PositionTrackingInputStr eam.java - src/
main/ , Java, 5 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ BlockData.java - src/
main/ , Java, 25 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ BlockOutputStream.java - src/
main/ , Java, 10 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ Bzip2BlockData.java - src/
main/ , Java, 48 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ Bzip2ParallelCompressor. java - src/
main/ , Java, 76 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ CompressionFormat.java - src/
main/ , Java, 96 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ ConcatGZIPInputStream.ja va - src/
main/ , Java, 8 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ GzipBlockData.java - src/
main/ , Java, 155 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ GzipParallelCompressor.j ava - src/
main/ , Java, 16 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ ParallelCompressor.java - src/
main/ , Java, 12 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ TunableGZIPOutputStream. java - src/
main/ , Java, 63 linesjava/ org/ usadellab/ trimmomatic/ util/ compression/ UncompressedBlockData.ja va - src/
test/ , Java, 88 linesjava/ org/ usadellab/ trimmomatic/ CompressionTest.java - src/
test/ , Java, 79 linesjava/ org/ usadellab/ trimmomatic/ MultiThreadedCompression Test.java - src/
test/ , Java, 22 linesjava/ org/ usadellab/ trimmomatic/ TrimmomaticTest.java - src/
test/ , Java, 108 linesjava/ org/ usadellab/ trimmomatic/ fasta/ FastaParserTest.java - src/
test/ , Java, 97 linesjava/ org/ usadellab/ trimmomatic/ fastq/ CompressionIntegrationTe st.java - src/
test/ , Java, 118 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserAutoDetectTes t.java - src/
test/ , Java, 40 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserEdgeCasesTest .java - src/
test/ , Java, 43 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserIteratorTest. java - src/
test/ , Java, 34 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserMultiLineTest .java - src/
test/ , Java, 35 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserNoNewlineTest .java - src/
test/ , Java, 54 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserProgressTest. java - src/
test/ , Java, 42 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserStrictnessTes t.java - src/
test/ , Java, 194 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserTest.java - src/
test/ , Java, 36 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserTrailingNewli neTest.java - src/
test/ , Java, 38 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserTruncatedTest .java - src/
test/ , Java, 38 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserValidVariatio nsTest.java - src/
test/ , Java, 47 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqParserWhitespaceTes t.java - src/
test/ , Java, 215 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqRecordGetLengthTest .java - src/
test/ , Java, 47 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqRecordHeadPosTest.j ava - src/
test/ , Java, 297 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqRecordTest.java - src/
test/ , Java, 17 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqRecordToStringTest. java - src/
test/ , Java, 69 linesjava/ org/ usadellab/ trimmomatic/ fastq/ FastqSerializerTest.java - src/
test/ , Java, 77 linesjava/ org/ usadellab/ trimmomatic/ fastq/ PairingValidatorTest.jav a - src/
test/ , Java, 31 linesjava/ org/ usadellab/ trimmomatic/ fastq/ RecordNamePatternTest.ja va - src/
test/ , Java, 66 linesjava/ org/ usadellab/ trimmomatic/ trim/ AvgQualTrimmerTest.java - src/
test/ , Java, 104 linesjava/ org/ usadellab/ trimmomatic/ trim/ BarcodeSplitterTest.java - src/
test/ , Java, 61 linesjava/ org/ usadellab/ trimmomatic/ trim/ BaseCountTrimmerAdvanced Test.java - src/
test/ , Java, 70 linesjava/ org/ usadellab/ trimmomatic/ trim/ BaseCountTrimmerTest.jav a - src/
test/ , Java, 98 linesjava/ org/ usadellab/ trimmomatic/ trim/ CropTrimmerTest.java - src/
test/ , Java, 151 linesjava/ org/ usadellab/ trimmomatic/ trim/ GeneralOvertrimTest.java - src/
test/ , Java, 158 linesjava/ org/ usadellab/ trimmomatic/ trim/ HeadCropTrimmerTest.java - src/
test/ , Java, 19 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaBitwiseExtendedT est.java - src/
test/ , Java, 81 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaBitwiseTest.java - src/
test/ , Java, 55 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerA dapterDimerTest.java - src/
test/ , Java, 62 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerC aseTest.java - src/
test/ , Java, 85 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerD irectionalTest.java - src/
test/ , Java, 47 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerD uplicateTest.java - src/
test/ , Java, 76 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerE dgeCasesTest.java - src/
test/ , Java, 67 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerI nterleaveTest.java - src/
test/ , Java, 54 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerI nvalidCharTest.java - src/
test/ , Java, 111 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerL oadTest.java - src/
test/ , Java, 55 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerM ediumPartialTest.java - src/
test/ , Java, 101 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerM ergeTest.java - src/
test/ , Java, 55 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerM inPrefixSimpleTest.java - src/
test/ , Java, 61 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerO verlapCapTest.java - src/
test/ , Java, 58 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerO verlapTest.java - src/
test/ , Java, 125 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerO vertrimTest.java - src/
test/ , Java, 83 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerP EvsSETest.java - src/
test/ , Java, 92 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerP riorityTest.java - src/
test/ , Java, 62 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerR obustnessExtendedTest.ja va - src/
test/ , Java, 91 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerS coreTest.java - src/
test/ , Java, 76 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerS eedMismatchTest.java - src/
test/ , Java, 45 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerS hortReadTest.java - src/
test/ , Java, 54 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerS impleShortOverlapTest.ja va - src/
test/ , Java, 284 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerT est.java - src/
test/ , Java, 49 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerT inyAdapterTest.java - src/
test/ , Java, 67 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerW ildcardTest.java - src/
test/ , Java, 46 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaClippingTrimmerZ eroThresholdTest.java - src/
test/ , Java, 86 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaPalindromeMinLen gthTest.java - src/
test/ , Java, 83 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaPalindromeOverla pTest.java - src/
test/ , Java, 83 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaPalindromeTest.j ava - src/
test/ , Java, 67 linesjava/ org/ usadellab/ trimmomatic/ trim/ IlluminaPrefixPairTest.j ava - src/
test/ , Java, 64 linesjava/ org/ usadellab/ trimmomatic/ trim/ LeadingTrailingTrimmerEd geTest.java - src/
test/ , Java, 65 linesjava/ org/ usadellab/ trimmomatic/ trim/ LeadingTrimmerTest.java - src/
test/ , Java, 58 linesjava/ org/ usadellab/ trimmomatic/ trim/ MaxLenTrimmerTest.java - src/
test/ , Java, 41 linesjava/ org/ usadellab/ trimmomatic/ trim/ MaximumInformationTrimme rLongReadTest.java - src/
test/ , Java, 114 linesjava/ org/ usadellab/ trimmomatic/ trim/ MaximumInformationTrimme rTest.java - src/
test/ , Java, 67 linesjava/ org/ usadellab/ trimmomatic/ trim/ MinLenTrimmerTest.java - src/
test/ , Java, 107 linesjava/ org/ usadellab/ trimmomatic/ trim/ SensitivitySpecificityTe st.java - src/
test/ , Java, 60 linesjava/ org/ usadellab/ trimmomatic/ trim/ SlidingWindowTrimmerEdge Test.java - src/
test/ , Java, 47 linesjava/ org/ usadellab/ trimmomatic/ trim/ SlidingWindowTrimmerInva lidArgsTest.java - src/
test/ , Java, 114 linesjava/ org/ usadellab/ trimmomatic/ trim/ SlidingWindowTrimmerTest .java - src/
test/ , Java, 95 linesjava/ org/ usadellab/ trimmomatic/ trim/ TailCropTrimmerTest.java - src/
test/ , Java, 53 linesjava/ org/ usadellab/ trimmomatic/ trim/ ToPhred33TrimmerTest.jav a - src/
test/ , Java, 53 linesjava/ org/ usadellab/ trimmomatic/ trim/ ToPhred64TrimmerTest.jav a - src/
test/ , Java, 90 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrailingTrimmerTest.java - src/
test/ , Java, 72 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerArrayProcessingTe st.java - src/
test/ , Java, 19 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactoryEmptyTest. java - src/
test/ , Java, 45 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactoryFailureTes t.java - src/
test/ , Java, 65 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactoryIlluminaCl ipFailureTest.java - src/
test/ , Java, 188 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactoryTest.java - src/
test/ , Java, 20 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerFactoryWhitespace Test.java - src/
test/ , Java, 107 linesjava/ org/ usadellab/ trimmomatic/ trim/ TrimmerPipelineTest.java - src/
test/ , Java, 44 linesjava/ org/ usadellab/ trimmomatic/ util/ LoggerTest.java - src/
test/ , Java, 59 linesjava/ org/ usadellab/ trimmomatic/ util/ PositionTrackingInputStr eamTest.java - LICENSE, License, 677 lines
- README.md, Text, 323 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 1,008 scripts, each with its path and the digest of its content;
- 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE261648, at NCBI GEO; found in “Data and code availability”
Data and code availability
Bulk RNA-seq data have been deposited at GEO: GSE261648 (https://
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 2, 28 September 2026
- Authors: added Cécile Martinat (0000-0002-5234-1064); Sandrine Baghdoyan (0000-0002-5904-7394); removed Cécile Martinat; Sandrine Baghdoyan
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 19 authors, 9 keywords, 12 MeSH terms, 3 funders, 54 references, 11 RRIDs.
Cite
This paper
Roussange, F., Gide, J., Tournois, J., Cailleret, M., Boland, A., Battail, C., Deleuze, J.-F., Polvèche, H., Auboeuf, D., Brockmann, K., Kabashi, E., Marian, A., El Kassar, L., Blondel, S., Salachas, F., Bruneteau, G., Peschanski, M., Martinat, C., & Baghdoyan, S. (2026). Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics. Stem cell reports, 21(7), 102977. https://
BibTeX
@article{roussange2026in
author = {Roussange, Florine and Gide, Jacqueline and Tournois, Johana and Cailleret, Michel and Boland, Anne and Battail, Christophe and Deleuze, Jean-François and Polvèche, Hélène and Auboeuf, Didier and Brockmann, Knut and Kabashi, Edor and Marian, Anca and El Kassar, Lina and Blondel, Sophie and Salachas, François and Bruneteau, Gaëlle and Peschanski, Marc and Martinat, Cécile and Baghdoyan, Sandrine},
title = {{Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics}},
journal = {Stem cell reports},
year = {2026},
month = jun,
volume = {21},
number = {7},
pages = {102977},
publisher = {Elsevier},
issn = {2213-6711},
doi = {10.1016/
url = {https://
pmid = {42349423},
pmcid = {PMC13385447}
}
RIS
TY - JOUR
AU - Roussange, Florine
AU - Gide, Jacqueline
AU - Tournois, Johana
AU - Cailleret, Michel
AU - Boland, Anne
AU - Battail, Christophe
AU - Deleuze, Jean-François
AU - Polvèche, Hélène
AU - Auboeuf, Didier
AU - Brockmann, Knut
AU - Kabashi, Edor
AU - Marian, Anca
AU - El Kassar, Lina
AU - Blondel, Sophie
AU - Salachas, François
AU - Bruneteau, Gaëlle
AU - Peschanski, Marc
AU - Martinat, Cécile
AU - Baghdoyan, Sandrine
TI - Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics
T2 - Stem cell reports
J2 - Stem Cell Reports
PY - 2026
DA - 2026/
VL - 21
IS - 7
SP - 102977
SN - 2213-6711
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "Integrative analysis of drug-gene signatures in human pluripotent stem cells reveals prazosin as a novel SQSTM1 regulator for ALS therapeutics",
"container-title": "Stem cell reports",
"author": [
{
"family": "Roussange",
"given": "Florine"
},
{
"family": "Gide",
"given": "Jacqueline"
},
{
"family": "Tournois",
"given": "Johana"
},
{
"family": "Cailleret",
"given": "Michel"
},
{
"family": "Boland",
"given": "Anne"
},
{
"family": "Battail",
"given": "Christophe"
},
{
"family": "Deleuze",
"given": "Jean-François"
},
{
"family": "Polvèche",
"given": "Hélène"
},
{
"family": "Auboeuf",
"given": "Didier"
},
{
"family": "Brockmann",
"given": "Knut"
},
{
"family": "Kabashi",
"given": "Edor"
},
{
"family": "Marian",
"given": "Anca"
},
{
"family": "El Kassar",
"given": "Lina"
},
{
"family": "Blondel",
"given": "Sophie"
},
{
"family": "Salachas",
"given": "François"
},
{
"family": "Bruneteau",
"given": "Gaëlle"
},
{
"family": "Peschanski",
"given": "Marc"
},
{
"family": "Martinat",
"given": "Cécile"
},
{
"family": "Baghdoyan",
"given": "Sandrine"
}
],
"container-title-short":
"volume": "21",
"issue": "7",
"page": "102977",
"DOI": "10.1016/
"PMID": "42349423",
"PMCID": "PMC13385447",
"ISSN": "2213-6711",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
25
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: Stan, Biopython, limma, 7 other tools, zebrafish, cellular / molecular
- [2] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: Biopython, limma, reshape2, 5 other tools, genetics / omics, 2 references
- [3] doi:10.1016/j.xcrm.2026.102766 [code]
- A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.Journal: Cell reports. MedicineIn common: limma, Keras, reshape2, 5 other tools, genetics / omics, other condition, cellular / molecular, 1 reference
- [4] doi:10.1038/s44318-026-00818-9 [code]
- FAM134B-mediated ER-phagy degrades APP and suppresses Alzheimer's disease pathology.Journal: The EMBO journalIn common: Biopython, limma, reshape2, 5 other tools, cellular / molecular
- [5] doi:10.1186/s13293-026-00927-4 [code]
- Gene regulatory network analysis identifies dysregulation of hypoxia pathways as contributing to glioblastoma treatment resistance in females.Journal: Biology of sex differencesIn common: Stan, limma, reshape2, 4 other tools, other condition, cellular / molecular
- [6] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: Biopython, Keras, scikit-learn, 4 other tools, zebrafish, genetics / omics
- [7] doi:10.1038/s43587-026-01207-x [code]
- A microprotein atlas of the human frontal cortex in Alzheimer's disease.Journal: Nature agingIn common: Biopython, limma, reshape2, 4 other tools, genetics / omics
- [8] doi:10.1126/sciadv.aed2952 [code]
- Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.Journal: Science advancesIn common: Biopython, limma, reshape2, 4 other tools, cellular / molecular
- [9] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: Biopython, Keras, scikit-learn, 4 other tools, genetics / omics, cellular / molecular
- [10] doi:10.1038/s41467-026-71391-2 [code]
- Accelerating Leigh syndrome drug discovery through deep learning screening in brain organoids.Journal: Nature communicationsIn common: scikit-learn, pandas, SciPy, 2 other tools, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 1008 scripts, and 2 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:4d22377aea51cb5c…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
