OSCR

Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform.

Code ↔ Paper

4 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 4 matches
  1. [1] § Methods › Scoring compound-protein interactions and generating interaction signatures ↔ CANDO_tutorial.ipynb, lines 478–510 · score 0.89 · Sorenson Dice coefficient, COACH algorithm, Binding site prediction, RDKit, binding strength, scores serve
  2. [2] § Methods › Benchmarking ↔ cando/cando.py, lines 2829–2916 · score 0.58 · normalized discounted cumulative, NDCG, score, metric, ranked, Benchmarking
  3. [3] § Methods › Generating drug predictions and corroborating them using literature searches ↔ CANDO_tutorial.ipynb, lines 86–117 · score 0.55 · drug indication mapping, DrugBank, MeSH, interaction signatures, protein, predict
  4. [4] § Methods › Benchmarking ↔ CANDO_tutorial.ipynb, lines 384–425 · score 0.54 · normalized discounted cumulative, nNDCG, prioritizes, cutoffs, score, CANDO

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Jupyter notebook · 719 lines · 56 KB · BSD-3-Clause · 3 matches

  1. # %% [markdown]
  2. # # CANDO tutorial
  3. # *Updated January 2025*
  4. #
  5. # This notebook will walk you through the basics of using the Computational Analysis of Novel Drug Opportunities platform, or CANDO. This includes preparing data, setting up a CANDO object, making therapeutic predictions, and benchmarking the platform.
  6. #
  7. # ## Table of contents
  8. #
  9. # - [Introduction](#Introduction)
  10. # - [Basic tutorial](#Basic-tutorial)
  11. # - [Getting started](#Getting-started)
  12. # - [Accessing data](#Accessing-data)
  13. # - [Drug prediction](#Drug-prediction)
  14. # - [Benchmarking](#Benchmarking)
  15. # - [Advanced topics](#Advanced-topics)
  16. # - [Customizing dataset](#Customizing-dataset)
  17. # - [Advanced drug prediction](#Advanced-drug-prediction)
  18. #
  19. # We recommend starting with the Getting started section, but if you wish to skip directly to a later section, or if you re-start the notebook and must re-create your CANDO object from scratch, please run the code below.
  20. # %%
  21. import cando as cnd # shortened to cnd for ease of typing/clarity
  22. import os
  23. while os.getcwd().split(os.sep)[-1] == 'tutorial':
  24. os.chdir('..')
  25. if not os.path.exists('tutorial'):
  26. cnd.get_tutorial()
  27. cnd.get_data(v='test.0',org='tutorial')
  28. os.chdir("tutorial")
  29. # Define variables with filepaths for our data
  30. cmpd_map='cmpds-v2.2.tsv' # Drug identifiers
  31. ind_map='cmpds2inds-v2.2.tsv' # Drug-indication associations
  32. matrix_file='tutorial_matrix-approved.tsv' # Drug-protein interaction matrix
  33. # Define additional variables used later in the tutorial; will be discussed as they appear
  34. dist_metric='cosine'
  35. ncpus=1
  36. # Create CANDO object
  37. cando = cnd.CANDO(cmpd_map, ind_map, matrix=matrix_file, compound_set='approved', compute_distance=True,
  38. dist_metric=dist_metric, ncpus=ncpus)
  39. # %% [markdown]
  40. # ---
  41. # %% [markdown]
  42. # # Introduction
  43. # %% [markdown]
  44. # This introduction will overview the fundamental ideas of the **C**omputational **A**nalysis of **N**ovel **D**rug **O**pportunities (CANDO) platform. If you want to continue directly to the practical portion of the tutorial, please skip to the [Getting started](#Getting-started) section.
  45. # %% [markdown]
  46. # CANDO is a similarity-based drug discovery platform, meaning it predicts new drugs for given indications based on their similarity to existing drugs within that indication. It calculates similarity by comparing the multitarget interactions of a given compound to those of the drugs of interest. Any quantifiable interaction between a compound and biological system can be used for this purpose, and multiple types of interactions can even be combined. For this tutorial, we will be using a series of compound-protein binding scores as our interaction signature.
  47. # %% [markdown]
  48. # Drug-protein binding scores can be generated through a variety of methods. CANDO has multiple in-house pipelines to create such scores based on small molecule drug chemical structure and protein binding pocket data (generated by [the Zhang group's protein binding site prediction algorithm, COACH](https://zhanggroup.org/COACH/)). These pipelines, CANDOCK and BANDOCK, can be used to create a protein library or add new proteins you think might be relevant to an existing protein library. Once these binding scores are generated, they are arranged into a vector or list corresponding to each compound: the *proteomic signature* of that compound. These proteomic signatures are arranged into a drug-protein interaction matrix. Once this matrix is created, alongside a corresponding lists of compounds and any known compound-indication associations, CANDO can be used to make novel drug predictions.
  49. # %% [markdown]
  50. # To prepare for novel drug or indication prediction, CANDO first creates a compound-compound similarity matrix. Similarity scores are calculated based on the compound interaction signatures, with more similar interaction signatures having a higher similarity score. Finally, for each drug with a known effect, a sorted list of the most similar compounds to that drug is created. To predict a new drug for a given indication, the most similar compounds to all drugs already in that indication are compared. If a non-indicated drug is very similar to multiple inidcated drugs, that is, it appears in the top ranks of multiple similar lists, it will be predicted as a possible new drug in that indication. Likewise, to predict a new indication for a given compound, the indications of the drugs most similar to it are found. If multiple drugs similar to the compound have the same inidcation, the indication is returned as a potential use for that compound.
  51. # %% [markdown]
  52. # Benchmarking is useful for determining how trustworthy these predictions are. In most cases, the overall performance of any major release of CANDO will already be benchmarked. However, if you customize the set of proteins/interaction scores, indications, and/or compounds available to CANDO, or if you want to know CANDO's performance on a subset of indications or compounds of interest, re-benchmarking will allow you to estimate the trustworthiness of your results more specifically. CANDO has multiple built-in benchmarking processes to allow rapid and easy benchmarking assessments to be carried out as necessary.
  53. #
  54. # ---
  55. # %% [markdown]
  56. # # Basic tutorial
  57. #
  58. # This section will go through the basics of using CANDO for simple drug discovery purposes.
  59. # %% [markdown]
  60. # ## Getting started
  61. #
  62. # The first step of using CANDO is importing the package. We will also be importing the os module for use in the next step.
  63. #
  64. # If this code throws an ImportError, check to see that you have cando.py downloaded and all relevant modules are installed. CANDO can be installed using `pip install cando-py` or using an Anaconda environment. See the [CANDO Github page](https://github.com/ram-compbio/CANDO) for more installation information.
  65. # %%
  66. import cando as cnd # shortened to cnd for ease of typing/clarity
  67. import os
  68. # %% [markdown]
  69. # Next, you must obtain all input data required to run CANDO. Most of the basic functions of CANDO require only three files:
  70. #
  71. # - **c_map** (compound map) - str, the filename of a file containing information on all drugs to be examined
  72. # - Each compound has two identifiers (a CANDO identifier and a DrugBank identifier), its name, and its approval status (e.g. approved or investigational) listed
  73. # - **i_map** (indication map) - str, the filename of a file containing information on all indications to be examined
  74. # - Each line maps one CANDO compound identifier to an indication, represented by its name and two identifiers (a MeSH identifier and a CANDO identifier)
  75. # - **matrix** - str, the filename of a file containing drug-protein binding scores (i.e. the protein interaction signature)
  76. # - Must contain the same number of compounds as the compound map file
  77. #
  78. # In this case, we will be using real compound and indication maps, but we will use a **tutorial matrix** with a limited number of drug-protein interactions so that we can generate predictions rapidly. Other datasets can be downloaded using the command `cnd.get_data()` or [via this website](http://protinfo.compbio.buffalo.edu/cando/data/v2.2+/).
  79. #
  80. # If you have previously downloaded an older version of this tutorial, run `cnd.clear_cache()` before the other commands below to clear out the old tutorial data.
  81. # %%
  82. # In case this box is run 2+ times, navigates up to the first non-tutorial folder
  83. while os.getcwd().split(os.sep)[-1] == 'tutorial':
  84. os.chdir('..')
  85. # Pull data for this tutorial
  86. cnd.get_tutorial()
  87. cnd.get_data(v='test.0',org='tutorial')
  88. os.chdir("tutorial") # navigates to the tutorial folder
  89. # Define variables with filepaths for our data
  90. cmpd_map='cmpds-v2.2.tsv' # Drug identifiers
  91. ind_map='cmpds2inds-v2.2.tsv' # Drug-indication associations
  92. matrix_file='tutorial_matrix-approved.tsv' # Drug-protein interaction matrix
  93. # Define additional variables used later in the tutorial; will be discussed as they appear
  94. dist_metric='cosine'
  95. ncpus=1
  96. # %% [markdown]
  97. # CANDO is an object-oriented program, and the majority of its methods require the creation of a CANDO object before they can be used. The CANDO object organizes and processes all drug, indication, and matrix data. To create a CANDO object, you will use the command `cnd.CANDO()` and pass the three files discussed above as arguments. We will be using a couple of optional arguments in addition to ensure our CANDO object is set up how we want it:
  98. #
  99. # 1. **compound_set** - string, indicates which compounds to include in the CANDO object. This is "all" by default, which allows for novel drug prediction. Using "approved" instead allows us to focus on drug repurposing (finding new uses for existing drugs) instsead.
  100. # 2. **compute_distance** - boolean, determines whether CANDO will compute the distances between the compounds based on the similarity of their interaction signatures. Unless you have a saved distance file (we do not), you have to use compute_distance For larger protein sets. Note, either `compute_distance` or `read_dists` must be used.
  101. # 3. **save_dists** - string, takes a filename to save the distances calculated (e.g. "big_protein_dists.tsv"). Must be used alongside `compute_distance`. This is useful for larger protein sets (e.g. >100,000 proteins) where it might be costly to repeatedly compute these distances. *Since we are using a small protein set, we will not use this.*
  102. # 4. **read_dists** - string, takes the filename of a saved distance file (e.g. "big_protein_dists.tsv"). This can only be used if you have previously computed and saved distances. *Since we have not done that, we will not use this.*
  103. # 5. **dist_metric** - string, distance metric used for compute_distance. This is "rmsd" (root mean squared distance) by default, but "cosine" can also be used. This is only used if compute_distance is True.
  104. # 6. **ncpus** - integer, parallelizes certain processes across the specified number of CPUs, making them faster. You can increase or decrease this number based on the number of CPUs available to you.
  105. #
  106. # There are many other optional argments available for the CANDO object instantiation. If you want to know more, please consult [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf).
  107. # %%
  108. cando = cnd.CANDO(cmpd_map, ind_map, matrix=matrix_file, compound_set='approved', compute_distance=True,
  109. dist_metric=dist_metric, ncpus=ncpus)
  110. # %% [markdown]
  111. # Creating or instantiating the CANDO object, as above, makes all of the functions of CANDO available. However, before continuing onto the actual usage of CANDO, we will first cover how to access information about the major objects that make up CANDO. This will allow you to ensure that all data was integrated into the CANDO object as expected and that you are creating the predictions you intend to in later steps.
  112. #
  113. # ---
  114. # %% [markdown]
  115. # ## Accessing data
  116. #
  117. # CANDO contains a large quantity of data, smooth access to which is essential to usage of its prediction functions.
  118. # %% [markdown]
  119. # CANDO consists of multiple objects, each working in tandem to create a drug prediction pipeline. The major objects are:
  120. #
  121. # - **CANDO object** - object, contains and organizes most variables and methods necessary to run CANDO
  122. # - **Compound** - object, represents a single small molecule compound/drug
  123. # - **Indication** - object, represents a single disease or other indication
  124. # - **Protein** - object, represents a protein in the proteomic signature
  125. #
  126. # Other objects also exist, but are less commonly used, including **Compound_pair**, **Pathway**, and **ADR**. You will not use these objects in most applications
  127. #
  128. # We will start by going over the CANDO object, which we just created/instantiated. The CANDO object contains lists of all objects in your model. The CANDO object also has many other properties that might be useful; consult [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf) for more information.
  129. #
  130. # - **cando.compounds** - list, all Compound objects (taken from the compound map, `cmpd_map`)
  131. # - Only compounds in this list can be associated with indications or predicted as novel therapies
  132. # - **cando.indications** - list, all Indication objects (taken from the indication map, `ind_map`)
  133. # - Only indications in this list can be associated with drugs or have novel therapies predicted
  134. # - **cando.proteins** - list, all Protein objects (taken from the drug-protein interaction matrix, `matrix_file`)
  135. #
  136. # To start with, we will check how many objects are in each group.
  137. # %%
  138. print(len(cando.compounds), 'compounds')
  139. print(len(cando.indications), 'indications')
  140. print(len(cando.proteins), 'proteins')
  141. # %% [markdown]
  142. # For this tutorial, our input files have 2449 compounds, 2214 indications, and 64 proteins. As a reminder, **most real matrices contain significantly more than 64 proteins**, but the protein count has been greatly reduced for the purposes of this tutorial.
  143. #
  144. # Let's look into the Compound objects in a little more depth. We can find a Compound if we know its CANDO ID, but we are unlikely to know that information offhand. Luckily, we can look up our compound by name using the `search_compound` function. This function will search for drugs with a similar name to your search term and return the top results.
  145. # %%
  146. print('Searching for "siprofloxasin"...')
  147. cando.search_compound('siprofloxasin')
  148. # %% [markdown]
  149. # Despite some mispellings, we found "ciprofloxacin," the antibiotic we were looking for! Once we have our compound ID, we can use the `get_compound` function to get the corresponding Compound object without going through the cando.compounds list.
  150. # %%
  151. cipro = cando.get_compound(421)
  152. print('The variable cipro contains the compound:', cipro.name)
  153. # %% [markdown]
  154. # A Compound contains the information CANDO "knows" about the drug, metabolite, or small molecule it represents. This includes the following:
  155. #
  156. # - **compound.id_** - integer, the CANDO ID of the compound (i.e. 421)
  157. # - **compound.name** - string, the name of the compound (i.e. ciprofloxacin)
  158. # - **compound.status** - the clinical trial status of the compound (from the compound map)
  159. # - All included compounds will be "approved" if `compound_set = approved`; otherwise, can be "approved" or "other"
  160. # - **compound.sig** - list, the (proteomic) signature of the drug, meaning all protein binding scores in order (from the matrix file)
  161. # - **compound.indications** - list of string, the MeSH IDs of the indications associated with the drug (from the indication map)
  162. # - The drug in question is considered effective against all indications in this list
  163. #
  164. # Other variables may also be associated with the Compound object. Consult [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf) for more information.
  165. #
  166. # To explore ciprofloxacin, we will look into some of the properties listed above.
  167. # %%
  168. print(cipro.name, cipro.id_)
  169. print(cipro.name, 'is', cipro.status)
  170. print('It is associated with', len(cipro.indications), 'indications')
  171. print('Its signature contains', len(cipro.sig), 'proteins')
  172. print('Signature:', cipro.sig)
  173. # %% [markdown]
  174. # Ciprofloxacin has existed since the 1980s, and it has accumulated a large number of indication associations in that time. Its signature represents its interactions with the 64 proteins we included in our matrix. These scores are from 0.0 to 1.0; we can see that ciprofloxacin does not have a high predicted binding affinity with any of these proteins.
  175. # %% [markdown]
  176. # Next, let's look at Indication objects. CANDO uses [**Me**dical **S**ubject **H**eadings (MeSH)](https://meshb.nlm.nih.gov/) IDs to internally catalogue indications. In order to predict new drugs for a given indication, you must first choose the most appropriate MeSH term. While you could search MeSH itself for the right term, not all terms have an Indication associated with them in CANDO. Searching CANDO for a MeSH ID ensures that your term is already present in your model.
  177. #
  178. # Finding an Indication is similar to finding a Compound. You can start by using `cando.search_indication` to find the closest text match to the disease you are interested in. If we wanted to find pneumonia, we could start with the following:
  179. # %%
  180. cando.search_indication('Pneumonia')
  181. # %% [markdown]
  182. # Many diseases have multiple comorbidities, etiologies, or otherwise related indications, so it is very common to get multiple results for an indication search. In this case, it seems like viral, aspiration, and bacterial pneumonia should have different causes and possibly different treatments, so it makes sense to focus on a specific type rather than the general term. If we want to study bacterial pneumonia, we would look at "Pneumonia, Bacterial" with a MeSH ID of [MESH:D018410](https://meshb.nlm.nih.gov/record/ui?ui=D018410).
  183. #
  184. # We can use this identifier to access the Indication itself using `get_indication`.
  185. # %%
  186. pneu_bact = cando.get_indication("MESH:D018410")
  187. print('The variable pneu_bact contains the indication:', pneu_bact.name)
  188. # %% [markdown]
  189. # Like a Compound, an Indication contains various information about the indication as it exists in your CANDO object. This includes:
  190. # - **indication.id_** - string, the MeSH identifier of the indication (i.e. MESH:D018410)
  191. # - **indication.name** - string, the English name of the indication (i.e. Pneumonia, Bacterial)
  192. # - **indication.compounds** - list of integer, the CANDO ID of all compounds associated with the indication
  193. # - If a compound is listed here, it is considered effective against our indication
  194. #
  195. # Other properties may also be associated with the Indication object. Consult [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf) for more information.
  196. #
  197. # We can find out more about how many compounds known to treat bacterial pneumonia using these properties.
  198. # %%
  199. print(pneu_bact.name, pneu_bact.id_)
  200. print(len(pneu_bact.compounds), 'compounds are associated with bacterial pneumonia')
  201. # %% [markdown]
  202. # By combining Compound and Indication properties, we can get even more information about what drugs are associated with bacterial pneumonia in our CANDO object.
  203. # %%
  204. print('The following compounds are associated with bacterial pneumonia:')
  205. print(*[cando.get_compound(cmpd).name for cmpd in pneu_bact.compounds], sep=', ')
  206. # %% [markdown]
  207. # If you check the list above, you will see a familiar compound. Bacterial pneumonia is one of the 77 indications associated with ciprofloxacin.
  208. #
  209. # Now that we can access drug and indication information, we can continue on to completing practical predictions and assessments in the [Drug prediction](#Drug-prediction) and [Benchmarking](#Benchmarking) sections.
  210. #
  211. # ---
  212. # %% [markdown]
  213. # ## Drug prediction
  214. #
  215. # One of the most common ways to use CANDO is to predict new drugs for an indication of interest. Let's say we want to predict new drugs for human immunodeficiency virus (HIV). We can start by searching for "HIV" to find the relevant MeSH term.
  216. # %%
  217. cando.search_indication('HIV')
  218. # %% [markdown]
  219. # HIV is associated with the development of multiple other conditions, so multiple results are returned. However, if we are specifically looking for drugs effective against HIV as a *virus*, not treatments for the resulting conditions, "HIV Infections" is the result we are looking for. This is associated with [MESH:D015658](https://meshb.nlm.nih.gov/record/ui?ui=D015658).
  220. #
  221. # Because CANDO is a similarity-based drug discovery method, it is generally more effective at predicting new drugs for indications that already have multiple drugs associated with them. In addition, if there are no drugs associated with the indication we want to predict effective drugs for, we would have to use a different approach to predict novel drugs (explored in the [Advanced topics](#Advanced-topics)).
  222. #
  223. # To ensure we have enough drugs to create effective predictions against HIV, we should find the Indication associated with HIV infections and check how many drugs are already associated with it.
  224. # %%
  225. hiv = cando.get_indication("MESH:D015658")
  226. print(len(hiv.compounds), 'compounds are associated with', hiv.name)
  227. print(*[cando.get_compound(cmpd).name for cmpd in hiv.compounds], sep=', ')
  228. # %% [markdown]
  229. # There are 32 drugs associated with HIV Infections in our data, which makes sense as it is a well-studied disease. Most seem to be protease inhibitors (saquinavir, fosamprenavir) or nucleoside analogs (zalcitabine, didanosine).
  230. #
  231. # We can now make predictions for other drugs/compounds that may be useful against HIV using the `canpredict_compounds` function. Given an indication, canpredict_compounds takes the lists of most similar compounds to each drug already associated with that indication. It then checks to see which compounds are similar to multiple indicated drugs; compounds that are similar to more indicated drugs are ranked higher, and ties are broken based on average rank within the similar lists. Possible parameters include:
  232. #
  233. # - **ind_id** - string, the indication for which new compounds should be predicted (i.e. MESH:D015658)
  234. # - **n** - integer, the number of similar compounds to be considered per indicated compoud (10 by default)
  235. # - **topX** - integer, the maximum number of predictions to be outputted (10 by default)
  236. # - **consensus** - boolean, outputs only compounds that appear in multiple similar lists if True (True by default)
  237. # - **keep_associated** - boolean, includes indicated drugs in the output list if True (False by default)
  238. # - **cmpd_set** - "all" "approved" or "other", determines which compound set to use ("all" by default)
  239. # - Note, if the CANDO object was created with only approved compounds, "all" will still only include aproved compounds
  240. # - **save** - string, name of an output file to save results (tsv format, nothing/not saved by default)
  241. #
  242. # Because we have 32 drugs associated with HIV, the full output of canpredict_compounds would be very long. Therefore, we will only look at the top 10 results (`topX = 10`). There are plenty of opportunities for a compound to appear in at least 2 of the 32 similar lists, so looking at only the top 10 most similar compounds (`n = 10`) and excluding any compounds that appear in only one similar list (`consensus = True`) are also appropriate.
  243. # %%
  244. cando.canpredict_compounds("MESH:D015658", n=10, topX=10)
  245. # %% [markdown]
  246. # These results are a little complex at first glance, so let's break down some key columns:
  247. #
  248. # - **score1** - integer, the number of similar lists in which the compound appears at or above rank n (here, rank 10)
  249. # - This is the primary ranking factor
  250. # - **score2** - float, the average rank of each compound in the "score1" similar lists it appears in
  251. # - When score1 is tied for two compounds, this breaks the tie
  252. # - **id** - integer, the CANDO ID of the compound; can be used to directly access the compound using `cando.get_compound()`
  253. # - **approved** - boolean, true if the compound is listed as approved in the input data; false otherwise
  254. # - **name** - string, the name of the compound
  255. #
  256. # The results can change greatly if we alter the input parameters. Instead of considering only the top 10 most similar compounds to each indicated compounds, let's change it to the top 25 (`n = 25`) and see what changes.
  257. # %%
  258. cando.canpredict_compounds("MESH:D015658", n=25, topX=10)
  259. # %% [markdown]
  260. # Although the first compound is the same (darunavir), the second and third, paclitaxel and decitabine, are new; previously, these two did not even appear in the top 10 predictions. Meanwhile, the previous third place compound, entecavir, is not in the top 10 anymore. To find out where it ended up, we can see more results for the same assessment using `topX = 25`.
  261. # %%
  262. cando.canpredict_compounds("MESH:D015658", n=25, topX=25)
  263. # %% [markdown]
  264. # It turns out entecavir only appeared in one more similar list when we considered the top 25 instead of the top 10 most similar compounds, pushing it down to rank 12.
  265. #
  266. # If we want to, we can also see where compounds already associated with HIV appear relative to newly predicted compounds using the same assessment using `keep_associated = True`. Technically, compounds associated with HIV are at a slight disadvantage, since they cannot appear in their own similar lists, but, because they are effective against HIV, we would still expect them to be pretty highly ranked overall if CANDO is functioning well. Let's run this assessment with `n = 10` and `topX = 10`.
  267. # %%
  268. cando.canpredict_compounds("MESH:D015658", n=10, topX=10,
  269. keep_associated=True)
  270. # %% [markdown]
  271. # If an asterisk * appears next to the label in the "approved" column, that means that drug is already associated with HIV. We can see there are three such compounds that appear in this list: amprenavir, saquinavir, and nelfinavir. As it turns out, CANDO ranks darunavir higher than every compound approved to treat HIV in this prediction.
  272. #
  273. # Before moving on, let's repeat this assessment on an indication with many fewer approved drugs. When we searched "HIV" earlier, one of our results was "HIV Wasting Syndrome," MESH:D019247. Let's look at that indication now.
  274. # %%
  275. hiv_ws = cando.get_indication('MESH:D019247')
  276. print(hiv_ws.name, hiv_ws.id_)
  277. print(len(hiv_ws.compounds))
  278. # %% [markdown]
  279. # There are only 4 drugs associated with HIV wasting syndrome in our dataset. This will likely reduce the number of results available through `canpredict_compounds`. Let's try running it with the same parameters as we did previously.
  280. # %%
  281. cando.canpredict_compounds('MESH:D019247', n=10, topX=10)
  282. # %% [markdown]
  283. # As you can see, though we still requested the top 10 predictions, only five appeared because only five compounds appear in multiple similar lists of the four drugs associated with HIV wasting syndrome. Stanozolol, the top ranked prediction, appears in only two similar lists (score1 is 2), but it appears in rank 2 on average (score2 is 1.0; rank 1 is 0.0).
  284. #
  285. # If we want to increase the number of predictions we receive, we can try a couple things. First, as we have done previously, we can increase `n` to 25, which is likely to increase the number of compounds appearing in multiple similar lists.
  286. # %%
  287. cando.canpredict_compounds('MESH:D019247', n=25, topX=10)
  288. # %% [markdown]
  289. # Now we have a full 10 results, but the ranks of most of our previous results (aside from testosterone propionate) have fallen. If we want to extend the previous list instead of changing the results, we can keep `n = 10` but set `consensus = False`. This will include in the ranks compounds that only appear in one similar list, but are highly ranked in that list.
  290. # %%
  291. cando.canpredict_compounds('MESH:D019247', n=10, topX=10, consensus=False)
  292. # %% [markdown]
  293. # As you can see, this output has the same top ranks as our first results, but the table has been extended to include additional compounds.
  294. #
  295. # While CANDO is more often used to predict novel drugs for an indication, it can also be used for the reverse, predicting new indications for an existing compound. This could be useful, for example, if one discovers a new small molecule in nature and wonders if it could have any pharmaceutical applications. `canpredict_indications` finds the most similar drugs to the compound in question, determines what those drugs are indicated to treat, and then returns any indications associated with multiple most similar drugs as potential uses.
  296. #
  297. # Key parameters of `canpredict_indications` are largely the same as those of `canpredict_compounds`, with three exceptions: the first argument is the Compound object, as opposed to the indication id; there is no `keep_associated` parameter; and there is an additional `sorting` variable that takes a string to determine whether to rank the outputtted indications by probability (`"prob"`) or score (`"score"`).
  298. #
  299. # Let's look at the antimicrobial paromomycin as an example. First, we have to find its id:
  300. # %%
  301. cando.search_compound('paromomycin')
  302. # %% [markdown]
  303. # Then, we get the Compound object and check which indications paromomycin is already associated with:
  304. # %%
  305. paro = cando.get_compound(1225)
  306. for indic in paro.indications:
  307. print(cando.get_indication(indic).name, cando.get_indication(indic).id_)
  308. # %% [markdown]
  309. # It looks like paromomycin is associated with six infection-related indications. We will use the top 10 most similar compounds to paromomycin (`n = 10`) and print the top 10 results (`topX = 10`).
  310. # %%
  311. cando.canpredict_indications(paro, n=10, topX=10)
  312. # %% [markdown]
  313. # All of paromomycin's existing indications are infections, so it is unsurprising that all of the predictions are also infections. As with the output of `canpredict_compounds`, the score indicates how many times an indication appears associated with the top 10 (or top n) most similar compounds to paromomycin.
  314. # %% [markdown]
  315. # This section covered the most basic ways to predict a new drug for an indication and an new indication for a drug. To learn about additional drug discovery methods within CANDO, you may continue to the [Advanced topics](#Advanced-topics) section or consult [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf).
  316. #
  317. # ---
  318. # %% [markdown]
  319. # ## Benchmarking
  320. #
  321. # Benchmarking is essential for assessing how well CANDO is working, particularly when using a new dataset or prediction function, but also when looking at predictions for a specific indication in greater depth. Benchmarking ensures that our platform has predictive value and that we know how much to trust the predictions we are working with.
  322. #
  323. # CANDO has a primary built-in benchmarking protocol, as well as multiple older benchmarking functions. We will focus on the primary benchmarking protocol, `canbenchmark_new`, here. This method uses a leave-one-out benchmarking protocol on every indication with two or more associated drugs. Every drug is withheld from its indication, and CANDO predicts novel drugs for that indication based on the remaining drugs. If the left out drug appears in these predictions, that is considered a success, as CANDO is able to predict the efficacy of that drug.
  324. #
  325. # The primary benchmarking function calculates four metrics:
  326. #
  327. # - **New average indication accuracy (nAIA)** - float, average % of indicated drugs that can be predicted within a given cutoff
  328. # - The "new" distinguishes these metrics from old average indication accuracy, as calculated by our older benchmarking functions
  329. # - A **Control nAIA** is also provided; this is the nAIA we would anticipate if CANDO worked as well as random chance
  330. # - **Pairwise accuracy (PA)** - float, % of drugs that appear within a given cutoff in the similarity lists of drugs with the same indication
  331. # - PA assesses similarity list quality, whereas nAIA assesses prediction quality
  332. # - **Indication coverage (IC)** - integer, the number of indications with a non-zero indication accuracy
  333. # - **Normalized discounted cumulative gain (nNDCG)** - float, scores # indicated drugs predicted within a given cutoff with prioritization of early recall
  334. #
  335. # Higher scores in these metrics should lead to higher confidence in our predictions. Additional metrics, such as precision at a given rank or area under the receiver operating characteristic (AUROC/AUC), may also be calculated. See the [the CANDO documentation](https://github.com/ram-compbio/CANDO/blob/master/docs/CANDO-v2.2.pdf) or [this paper on evaluating drug repurposing technologies](https://www.biorxiv.org/content/10.1101/2020.12.03.410274v1) for more information about metrics.
  336. # %% [markdown]
  337. # `canbenchmark_new` has a couple key arguments:
  338. # - **file_name** - string, a descriptive name to be assigned to all benchmarking files created
  339. # - **n** - integer, the number of similar compounds to be considered per indicated compoud (10 by default)
  340. # - Use the same value for `n` as you used in `canpredict_compounds` to ensure your benchmarking results are applicable
  341. # - **indications** - list of string, a list of the subset of indications to be examined
  342. # - If this argument is not used, all indications with at least two associated drugs will be assessed
  343. # - **associated** - boolean, if True (default), excludes all drugs that are not associated with an indication from consideration
  344. #
  345. # As with other CANDO functions, additional arguments are available. `canbenchmark_new` has not yet been added to the documentation, but some arguments are explained in the function header within cando.py.
  346. #
  347. # We can choose to assess CANDO in two ways: we can assess the overall performance of CANDO, or we can assess a subset of indications of interest. The former is especially useful when we are using an unusual dataset that might affect the performance of CANDO. In this tutorial, we are using a matrix with only 64 proteins, which might alter performance, so it makes sense to run an overall assessment. We will give this assessment the name "tutorial_assessment" and use n=10, as we did when running `canpredict_compounds`. **Note that this function will take a long time to run.**
  348. # %%
  349. results = cando.canbenchmark_new('tutorial_assessment', n=10)
  350. # %% [markdown]
  351. # The primary results of `canbenchmark_new` are printed out directly for you to read. Based on the nAIA results, across all indications, CANDO is able to recall 8.6% of indicated drugs in the top 10 compounds (out of 1464), 13.3% in the top 25, and 23.8% in the top 100. CANDO also outperforms random chance (control-nAIA) at every cutoff. In addition, the indication coverage (IC) tells us that 523 indications (out of 1595 assessed; see top100% column) have non-zero performance at the top 10 threshold, 721 at the top 25 threshold, and 982 at the top 100 threshold. Note that these results correspond to the 64-protein matrix used in this tutorial (plus the compound and indication mapping); as a general rule, performance is better when a larger protein set is considered.
  352. #
  353. # In addition, multiple additional files that may be of interest are generated:
  354. #
  355. # - **Summary file** - TSV, contains overall metrics; top-level folder with a name starting "summary"
  356. # - **Raw results** - CSV, contains the rank at which each indicated drug was predicted; raw_results folder
  357. # - **Pairwise results** - CSV, contains the ranks of each pair of drugs with the same indications in each others' similarity lists; pairwise_results folder
  358. # - **Results analysed** - TSV, contains indication accuracy or NDCG scores for each individual indication; results_analysed_named folder
  359. #
  360. # These files can be viewed in the tutorial folder, and they can be opened in a text editor like Notepad or in certain spreadsheet editors, like Excel. For now, we will just look at the results of the summary file, which should reflet the results printed out when running `canbenchmark_new` above:
  361. # %%
  362. with open('summary-tutorial_assessment-10-associated.tsv', 'r') as f:
  363. print(f.read()) # Note: do NOT do this with the longer, non-summary files; it may crash or hang
  364. # %% [markdown]
  365. # The other reason one might want to complete benchmarking is to determine how well CANDO performs on a given indication. For example, earlier we were looking into HIV infections, MESH:D015658. Let's replicate those results here.
  366. # %%
  367. cando.canpredict_compounds("MESH:D015658", n=10)
  368. # %% [markdown]
  369. # We have the same 10 results as before, but now we want to estimate how trustworthy these results are. We can use `canbenchmark_new` with the same `n` value as we used in our prediction to get such an estimate. Since we *only* want to assess how well CANDO performs on HIV here, we can pass a list with only the MeSH ID we are interested in to the `indications` parameter.
  370. # %%
  371. results = cando.canbenchmark_new("HIV_assessment", n=10, indications=["MESH:D015658"])
  372. # %% [markdown]
  373. # We found that 9.375% of drugs associated with HIV infections, or 3 of the 32 drugs, were ranked within the top 10 compounds predicted for this indication, as compared to 2449 compounds ranked overall. If you recall, this is also what we found when we ran `canpredict_compounds` with `keep_associated=True`. When we look at the top 25 compounds, this increases to 21.875%, or 7 out of 32 drugs.
  374. #
  375. # Based on these results, we can expect that not all compounds effective against HIV infections are going to be in our top 10 predictions, which is unsurprising. However, by comparing our nAIA to the control nAIA, we can see that CANDO is performing far better than random chance on predicting effective drugs for HIV. CANDO also performs better on HIV infections than on indications in general. Therefore, it is reasonable to use CANDO to predict novel drugs for HIV infections.
  376. # %% [markdown]
  377. # Historically, we have assessed the performance of CANDO based on calculating accuracies for the similar lists of every compound within an indication. This is the idea behind the original CANDO benchmarking algorithm, `canbenchmark`. Although the newer `canbenchmark_new` is more relevant to most applications, the original algorithm is still preserved for posterity and potential in the development of CANDO.
  378. #
  379. # As an example of the original `canbenchmark`, let's look at hyperparathyroidism, MESH:D006961. There are four drugs associated with this indication: cinacalcet (884), paricalcitol (786), dihydrotachysterol (938), and calcitriol (32). As it turns out, cinacalcet does not have any of the other three drugs in its top 10 most similar compounds. However, paricalcitol, dihydrotachysterol, and calcitriol are all in each other's top 10 most similar compounds lists. Since three out of the four compounds' similar lists have at least one other indicated compound in the top 10, we calculate an indication accuracy of 75% at the top 10 threshold for hyperparathyroidism. This can be repeated for multiple rank thresholds (top 25, 100, etc) and for every other indication with at least two associated compounds. From this, the unweighted AIA, weighted APA, and IC can be calculated for these thresholds.
  380. # %% [markdown]
  381. # Doing this work by hand is tedious and repetitive, which is why we created `canbenchmark` to rapidly calculate these metrics for every indication. `canbenchmark` only requires one argument, a string that determines what the output files it creates will be named. The protein (or other interaction) matrix, compound list, and compond-indication associations given to the CANDO object as input may all affect the bnechmarking performance.
  382. # %%
  383. cando.canbenchmark('tutorial')
  384. # %% [markdown]
  385. # Besides the abbreviated results it prints out, `canbenchmark` creates three additional files: a raw results file that contains information on every compound-indication pair (found at raw_results/raw_results-tutorial.csv), an analyzed results file that contains every individual indication accuracy (found at results_analysed_named/results_analysed_named-tutorial.csv), and a summary file containing overall metrics (found at summary-tutorial.tsv). You can open these results from your file navigator into spreadsheet software like Excel or text editors like Notepad, and these files can also be directly opened and examined via your code, as below:
  386. # %%
  387. with open('summary-tutorial.tsv', 'r') as f:
  388. print(f.read()) # Note: do NOT do this with the longer, non-summary files; it may crash or hang
  389. # %% [markdown]
  390. # As we can see, we have an AIA of 20% at the top 10 threshold, meaning that, across all indications, about 20% of compounds had at least one other compound with the same indication in their top 10 most similar lists. The APA is higher because the chance of recovering an indicated compound in the top ranks is higher when there are more indicated compounds, so more heavily weighting larger indications leads to a better result. Indication coverage is 819, meaning that 819 out of 1595 indications with at least two associated compounds had an indication accuracy greater than 0%.
  391. # %% [markdown]
  392. # ---
  393. # %% [markdown]
  394. # # Advanced topics
  395. #
  396. # This section will cover more advanced and customizable ways of working with CANDO. Note that this section will not go into the same depth as the basic tutorial, as it is assumed you already understand the basics of CANDO.
  397. # %% [markdown]
  398. # ## Customizing dataset
  399. # %% [markdown]
  400. # ### Interaction matrix
  401. #
  402. # The example interaction matrix, matrix_file, has already been downloaded via get_tutorial(). This function may take anywhere from ~1-10 mins on 3 cores, depending on your computer's processor.
  403. #
  404. # In this step we will generate a matrix of 2,449 approved compounds by 64 proteins, populated with the corresponding interaction score betwen each drug and protein. The final matrix will have drugs as the columns (indexed according to the compound mapping file), and proteins as the indices (indexed by PDB and chain ID).
  405. #
  406. # The function generate_matrix() first creates a dataframe for all chemical fingerprint similarity scores comparing each compound in the specified version library v (or only approved drugs with approved_only=True) to every potential binding site ligand from the PDB using RDKit to compute the chemical fingerprints in which fp denotes the type/radius of fingerprint (rd_ecfp4, rd_ecfp8, etc). The vector type, vect, can be binary ("1024_bit" or "2048_bit") or integer ("int") and denote the presence or absence or count of molecular substructures in the molecule, respectively. If using integer vectors, dist should be set to "dice", but binary vectors can be set to Tanimoto ("tani"). Next, it creates a dataframe of all potential binding sites for each protein in the specified protein library, protlib, with their corresponding binding site scores from the specified binding site prediction method, bs ("coach", "cof", "ssite", "tms"). The function then iterates over all drugs and protein binding sites to populate the matrix with the best score based on the following input parameters: 1) i_score - the scoring protocol of choice, which can be 'C', 'dC', 'P', 'CxP', 'dCxP', where C is the fingerprint similarity score (either Tanimoto or Sorenson-Dice coefficient), dC is the percentile of the C score for the compound compared to all ligands in the library, and P is binding site score associated with the ligand predicted by the COACH algorithm, 2) c_cutoff and p_cutoff - set cutoffs which ignore any C or P scores below each threshold, respectively, and 3) percentile_cutoff - similar to c_cutoff but with the dC score (overrides c_cutoff if not None). These scores serve as a proxy for binding strength/probability of the drug and protein target. We then output the matrix to a tsv file, out_file, which in this case is "tutorial_matrix-all.tsv" (out_path can be set to write the file to a specific directory).
  407. #
  408. # We will create a matrix using the 'CxP' protocol, which just chooses the top Dice score to any ligand predicted to bind to a given protein, without any cutoffs. This function is parallelized, so setting the variable ncpus will change the number of processors that are used for this function. NOTE: percentile cutoff protocols ('dC', 'dCxP') take much longer to compute than 'C' and 'P' protocols.
  409. # %%
  410. # generate example cando interaction matrix (2,449 compounds x 64 proteins)
  411. cnd.generate_matrix(v="test.0", fp="rd_ecfp4", vect="int",
  412. dist="dice", org="tutorial", bs="coach",
  413. c_cutoff=0.0, p_cutoff=0.0, percentile_cutoff=0.0,
  414. i_score="dCxP", out_file='', out_path=".",
  415. nr_ligs=True, approved_only=True,
  416. lib_path='', prot_path='', lig_name=False, ncpus=ncpus)
  417. # %% [markdown]
  418. # ### Novel compound
  419. #
  420. # The CANDO platform contains an extensive library of approved drugs and other compounds from DrugBank. However, if you wish to predict indications or similar drugs for a compound that is not present in our library, we make it possible with the `generate_signature()` function.
  421. #
  422. # First, you must have the compound properly formatted in mol file format. There are many programs that provide conversion between many chemical file formats, such as OpenBabel.
  423. #
  424. # Next, run `generate_signature()`. This will populate a tsv file with Tanimoto/Sorenson-Dice similarity scores of the provided compound to all binding site ligands in our database. These values will be used for the generation of the drug-proteome signature. The input parameters for this function are very similar to those for `generate_matrix()`, however the first argument is the path to the compound structure file in mol format. The output signature file will be saved in tsv format with the name you provide for "out_file" (with the appended path from "out_path").
  425. #
  426. # **NOTE: the interaction score protocol must match the input matrix protocol, which was 'CxP' as above. If we used a different protocol, these scores would not be directly comparable.**
  427. # %%
  428. cmpd_file = "lmk235.mol"
  429. signature_file = "lmk235_signature.tsv"
  430. cnd.generate_signature(cmpd_file, fp="rd_ecfp4", vect="int", dist="dice",
  431. org="tutorial", bs="coach", c_cutoff=0.0, p_cutoff=0.0,
  432. percentile_cutoff=0.0, i_score="CxP", out_file=signature_file,
  433. out_path=".", nr_ligs=True)
  434. # %% [markdown]
  435. # We then must add the compound to the existing `CANDO` object with the `add_cmpd()` function. The inputs are the newly generated signature file and the desired name (which below is "lmk-235").
  436. # %%
  437. cando.add_cmpd(signature_file, new_name='lmk-235')
  438. # %% [markdown]
  439. # Now that our new compound is added to the platform, we can see what other compounds to which it is similar. We can use the `similar_compounds()` function from before to print those compounds.
  440. # %%
  441. lmk235 = cando.get_compound(10674)
  442. cando.similar_compounds(lmk235, n=10)
  443. # %% [markdown]
  444. # Finally, we can predict potential indications for which our new compound may be useful. We can use the `canpredict_indications()` as before to print those results.
  445. # %%
  446. cando.canpredict_indications(lmk235, n=25, topX=15)
  447. # %% [markdown]
  448. # Interestingly, both Hypertension (MESH:D006973) and Pain (MESH:D010146) are top predictions - the analgesic and hypotensive properties of lmk-235 are both supported by in vivo studies in the literature.
  449. # %% [markdown]
  450. # ### Compound library
  451. #
  452. # In order to create a new compound set, you must first have a TSV (tab separated values) file that contains one of two chemical file types:
  453. #
  454. # 1. SMILES - `file_type='smi'`. The file must have the SMILES string as the first column and the corresponding compound name in the second column, e.g.
  455. # `C1CNCCN(C1)S(=O)(=O)C2=CC=CC3=C2C=CN=C3 fasudil`
  456. #
  457. # 2. Mol - `file_type='mol'`. The file must have the name of the file, without the file extension. In addition, the path to the files must be given in the argument `cmpd_dir`.
  458. #
  459. # In our example, we will create a new compound library of select tyrosine kinase inhibitors (TKIs) that we will then use to create a new matrix with just these drugs. First, we generate the library.
  460. # %%
  461. cnd.add_cmpds("tki_set-test.smi", file_type='smi', fp="rd_ecfp4", vect="int", cmpd_dir=".", v='tki')
  462. # %% [markdown]
  463. # We have now generated a new compound library, located at `./tki/cmpds`. This location is partially hardcoded, so it will also create a directory in your current working directory, `./`, named after the `v` you choose, `./tki`. This makes it easier for the end-user.
  464. #
  465. # Notice we chose to use the rdkit fingerprint ecfp4, `fp="rd_ecfp4"`, and the vector type int, `vect="int"`. We need to keep track of this for our matrix generation.
  466. # %% [markdown]
  467. # Next we will generate a matrix using this **tki** compound library. The generate matrix function can take in a customized matrix, as opposed to the pregenerated/downlaoded matrices based upon our predefined `v`, e.g. v2.2, v2.3, etc. By setting `v` to the same name you used in the previous `add_cmpds` function, you can load those cmpds to create a new matrix.
  468. #
  469. # We will set `v='tki'` and `lib_path='.'`. This means we will use the **tki** library, which is located in the current working dir. In addition, sicne we created our compound library using fingerprint type ecfp4, we need to make sure we use it here, as well. The remaining arguemnts has already been discussed in the previous [Interaction matrix](#Interaction-matrix) section.
  470. # %%
  471. cnd.generate_matrix(v="tki", lib_path='.',
  472. fp="rd_ecfp4", vect="int",
  473. dist="dice", org="tutorial", bs="coach",
  474. c_cutoff=0.0, p_cutoff=0.0, percentile_cutoff=0.0,
  475. i_score="CxP", out_file='', out_path="",
  476. nr_ligs=True, approved_only=False,
  477. lig_name=False, ncpus=ncpus)
  478. # %% [markdown]
  479. # We now have a brand new matrix with all of our new TKIs scored against our set of 64 test proteins. This file is located at `./tki/matrices/`. Again, these paths are mostly hardcoded. You can control where the matrix is written and what the name is using the `out_path` and `out_file` arguments. When these arugments are set to '' the hardcoded paths are used, as shown in this example.
  480. #
  481. # Now we will load the newly created matrix into a CANDO object.
  482. # %%
  483. # Set compound mapping variable to the new TKI compound mapping file
  484. tki_map = 'tki/mappings/cmpds-tki.tsv'
  485. tki_matrix = 'tki/matrices/rd_ecfp4-int-dice-tutorial-coach-c0.0-p0.0-CxP.tsv'
  486. # Create CANDO object using the new compound mapping and TKI matrix
  487. tki_cando = cnd.CANDO(tki_map, ind_map, matrix=tki_matrix, compound_set='all', compute_distance=True,
  488. dist_metric=dist_metric, ncpus=ncpus)
  489. # %% [markdown]
  490. # Check the new CANDO object data and the most similar compounds to the first TKI.
  491. # %%
  492. # print cando object stats
  493. print('compounds', len(tki_cando.compounds))
  494. print('indications', len(tki_cando.indications))
  495. print('proteins', len(tki_cando.proteins))
  496. print('')
  497. # print first TKI name and signature
  498. c = tki_cando.compounds[0]
  499. print(c.name, len(c.sig))
  500. print(c.sig)
  501. print('')
  502. # top5 most similar compounds to first TKI
  503. for s in c.similar[0:5]:
  504. print(s[0].name, round(s[1], 3))
  505. # %% [markdown]
  506. # ---
  507. # %% [markdown]
  508. # ## Advanced drug prediction
  509. # %% [markdown]
  510. # ### De novo prediction
  511. #
  512. # Sometimes there are no drugs associated with an indication (whether it be in reality or in your input indication mapping); in those cases we can use the `canpredict_denovo()` function to suggest putative candidates. Basically, this function counts/sums the number of protein interaction scores above a set threshold for each compound and ranks them according to frequency and strength. This is particularly useful for pathogen proteomes (like SARS-CoV-2 or bacterial proteomes) or for finding which compounds target a subset of proteins of interest (e.g. kinases). In our case, we have a diverse set of proteins, which would render the results meaningless. Instead, we can input a list of protein IDs in which we are interested. We have precompiled a list of bacterial proteins already present in the sample matrix for convenience. Note: if we had read in an indication-genes mapping file, we could use the `ind_id=` parameter to automatically select all proteins associated with said indication.
  513. # %%
  514. bacterial_proteins = ["3mk7C", "3eziA", "1u2mC", "3atsA", "2x4mA",
  515. "4wliA", "1t4aA", "4zxkA", "1zhhA", "1eb0A",
  516. "4nqwB", "2gqrA", "2qlcA", "3rf1A", "2xpwA"]
  517. cando.canpredict_denovo(method='sum', threshold=0.6, topX=30, proteins=bacterial_proteins)
  518. # %% [markdown]
  519. # ### Machine learning
  520. #
  521. # The "proteomic vectors" within CANDO lend themselves well to machine learning to perhaps learn more complex relationships between the proteins within the vector and their impacts on the treatment of diseases. CANDO has built-in ML algorithms that allow for two main functionalities:
  522. # 1. Benchmark the platform using a hold-one-out protocol very similar to canbenchmark
  523. # 2. Make predictions for novel or non-associated compounds that may be therapeutic for a given disease
  524. #
  525. # The ML module currently supports 2 algorithms: random forests and logistic regression. The models are trained on drugs approved for the disease (positive classes) and an equal number of randomly selected "neutral samples", which are drugs/compounds not approved for the disease (negative samples). Random seeds may be set to ensure the same compounds are used in training.
  526. #
  527. # We have the option to benchmark the platform with an ML algorithm - this module outputs files very similar to canbenchmark. For this tutorial, we will skip this function as it requires a great deal of time to complete (training a separate model for EVERY drug-disease association, basically). The command to do so with a logistic regression classifier is below, feel free to run it! The `'out='` flag defines the name of the output files. Again, only diseases with 2+ compounds associated are benchmarked.
  528. #
  529. # `cnd.ml(method='rf', benchmark=True, seed=50, out='test_rf')`
  530. #
  531. # We can also use this module to predict if a certain compound may be therapeutic for a given disease. We can use the
  532. # `'predict='` flag to specify a list of compounds that we wish to predict with the classifier. Let's use three drugs, imatinib, buprenorphine, and lisdexamfetamine, and see if they are predicted to have antibiotic activity using a random forest classifier.
  533. # %%
  534. baci = cando.get_indication('MESH:D001424')
  535. imat = cando.get_compound(503)
  536. bup = cando.get_compound(797)
  537. lamf = cando.get_compound(1121)
  538. cando.ml(method='rf', effect=baci, benchmark=False, seed=50, predict=[imat, bup, lamf])
  539. # %% [markdown]
  540. # The accuracy is quite poor, so fine-tuning and hyperparameter optimization is necessary to enhance performance.
  541. # %% [markdown]
  542. # ### Custom protein sets
  543. #
  544. # It may be useful for some users to probe compound-protein interaction similarity, but only in the context of a few particular proteins (e.g. set of kinase inhibitors). Instead of generating a new matrix with all of these proteins and their corresponding interaction values, which can begin to take up a lot of storage if done multiple times, the `'protein_set='` flag can be specified during the instantiation of the CANDO object. This flag contains the path to the protein subset the user wishes to use, which is simply a list of UniProt protein IDs. The CANDO object will automatically check for each ID if it either simply matches any UniProt IDs within the matrix or if that UniProt ID is associated with any PDB chains within the matrix (based on a mapping from the SIFTs project). If there are matches, the CANDO object will now contain Compound objects with only those protein interaction values in their signatures. Below is how to use this functionality - this time we can search for the bacterial proteins from the previous `canpredict_denovo()` section, but filter the proteins from the start using a list of UniProt IDs.
  545. # %%
  546. cando_subset = cnd.CANDO(cmpd_map, ind_map, matrix=matrix_file, compound_set='approved',
  547. compute_distance=True, protein_set=protein_set,
  548. dist_metric=dist_metric, ncpus=ncpus)
  549. print("Number of proteins in new signature =", len(cando_subset.proteins))
  550. # %% [markdown]
  551. # The signature was successfully edited to 15 proteins. Note: this does not nececessarily mean the each UniProt ID had a corresponding PDB match -- multiple PDB chains can be associated with a given UniProt ID. If the matrix contains UniProt IDs in the first column instead of PDB IDs, as in this case, the "Direct UniProt matches" value would be incremented if they match.
  552. #
  553. # We can also repeat all benchmarks and predictive algorithms with the new signatures. Below is the default benchmarking results with the new signatures.
  554. # %%
  555. cando_subset.canbenchmark('test_subset')
  556. # %% [markdown]
  557. # Let's repeat the ML code from above, but this time with the reduced protein subset.
  558. # %%
  559. baci = cando_subset.get_indication('MESH:D001424')
  560. imat = cando_subset.get_compound(503)
  561. bup = cando_subset.get_compound(797)
  562. lamf = cando_subset.get_compound(1121)
  563. cando_subset.ml(method='rf', effect=baci, benchmark=False,
  564. seed=50, predict=[imat, bup, lamf])
  565. # %% [markdown]
  566. # As you can see, the output probability significantly change for one of the compounds, lisdexamfetamine, from 0.340 to 0.600, illustrating the impact of protein subset composition on model behavior.
  567. # %% [markdown]
  568. # ### Virtual screening
  569. #
  570. # Though the CANDO platform is mainly intended for multitarget drug discovery and repurposing, there still exists the option to check the top compound hits for a given protein. This is accomplished via the `virtual_screen()` function. Let's test the top hits for two of the known bacterial proteins from above, namely "1u2mC", "3atsA".
  571. # %%
  572. cando.virtual_screen("1u2mC")
  573. cando.virtual_screen("3atsA")
  574. # %% [markdown]
  575. # Streptozocin is the top hit for 1u2mC, which suggests it shares significant structural similarity to a ligand known/predicted to bind to this protein. Similarly, kanamycin is the 5th hit for 3atsA. These scores/ranks can change significantly based on scoring protocol used for the matrix - changing i_score in `generate_matrix()` to "dC", for example, would greatly affect the output.
  576. # %% [markdown]
  577. # *Credits: CANDO tutorial created by Zackary Falls; revised by Melissa Van Norden*
  578. # %%

CANDO_tutorial.ipynb at commit 955eed2, under BSD-3-Clause · at the source

Overview

Authors: Sumei Xu1,2,3, Yakun Hu3, William Mangione3, Melissa Van Norden3, Katherine Elefteriou3, Zackary Falls3, Ram Samudrala3
ORCID iDs: Zackary Falls
  1. Phase I Clinical Trial Center, Xiangya Hospital, Central South University, 87 Xiangya Rd, Changsha, 410008 Hunan China
  2. National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University, 87 Xiangya Rd, Changsha, 410008 Hunan China
  3. Department of Biomedical Informatics, University at Buffalo, 77 Goodell Street, Buffalo, NY 14203 USA
Journal: Journal of cheminformatics, volume 18, issue 1, article 65
Dates: received 22 May 2025; accepted 20 March 2026; published online 12 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s13321-026-01191-9 · PMID 41968358 · PMCID PMC13185303 · OpenAlex W4410716747
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: other condition (population), clinical / translational (subfield)
Keywords: Glioma, Multiscale drug discovery, Computational drug repurposing, Translational bioinformatics, Deep learning, Systems biology
Topic: Computational Drug Discovery Methods (Computational Theory and Mathematics, Computer Science), according to OpenAlex
Funding: NIH NLM R25 Award (R25LM014213); startup funds from the Department of Biomedical Informatics at the University at Buffalo; NCATS NIH HHS (UL1 TR001412); NIH NCATS ASPIRE Design Challenge Award; NLM NIH HHS (R25 LM014213, T15 LM012495); National Institutes of Health (NIH) Director’s Pioneer Award (DP1OD006779); NIDA NIH HHS (K01 DA056690); NIH National Library of Medicine (NLM) T15 Award (T15LM012495); NIH Clinical and Translational Sciences (NCATS) Award (UL1TR001412); NIH NCATS ASPIRE Reduction-to-Practice Award; National Institute of Standards of Technology (NIST) Award (60NANB22D168); NIDA Mentored Research Scientist Development Award (K01DA056690); NIH HHS (DP1 OD006779)
Citations: cited by 1 paper (Europe PMC); 166 references in the paper

Abstract

Glioma is a highly malignant brain tumor with limited treatment options. We employed the Computational Analysis of Novel Drug Opportunities (CANDO) platform for multiscale therapeutic discovery to predict new glioma therapies. We began by computing interaction scores between extensive libraries of drugs/compounds and proteins to generate “interaction signatures” that model compound behavior on a proteomic scale. Compounds with signatures most similar to those of drugs approved for a given indication were considered potential treatments. These compounds were further ranked by degree of consensus in corresponding similarity lists. We benchmarked performance by measuring the recovery of approved drugs in these similarity and consensus lists at various cutoffs, using multiple metrics and comparing results to random controls and performance across all indications. Compounds ranked highly by consensus but not previously associated with the indication of interest were considered new predictions. Our benchmarking results showed that CANDO improved accuracy in identifying glioma-associated drugs across all cutoffs compared to random controls. Our predictions, supported by literature-based analysis, identified 24 potential glioma treatments, including approved drugs like vitamin D, taxanes, vinca alkaloids, topoisomerase inhibitors, and folic acid, as well as investigational compounds such as ginsenosides, chrysin, resiniferatoxin, and cryptotanshinone. Further functional annotation-based analysis of the top targets with the strongest interactions to these predictions identified Vitamin D3 receptor, thyroid hormone receptor, acetylcholinesterase, cyclin-dependent kinase 2, tubulin alpha chain, dihydrofolate reductase, and thymidylate synthase. These findings indicate that CANDO’s multitarget, multiscale framework is effective in identifying glioma drug candidates thereby informing new strategies for improving treatment.

Scientific contribution

(1) We present a robust, multiscale drug discovery framework that accurately recovers known glioma therapies and uncovers 24 novel candidates with strong literature and mechanistic support. (2) By modeling compound behavior across the proteome, our method pinpoints key targets–including VDR, CDK2, and DHFR–implicated in glioma biology. (3) This work positions CANDO as a powerful tool for rational repurposing and discovery of urgently needed treatments for aggressive brain tumors.

Supplementary Information: The online version contains supplementary material available at 10.1186/s13321-026-01191-9.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 4 matches between paragraphs and lines of code.

ram-compbio/CANDO

License: BSD-3-Clause
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 955eed289e444fa34dc02a916d812202e9abda8c, 23 September 2025
Languages: Python (4), Jupyter (1)
Size: 14 files, 5 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file, environment (setup.py), continuous integration, documentation, 1 notebook
Not found: CITATION.cff, tests
Tools: Matplotlib (1 file), NumPy (1 file), pandas (1 file), RDKit (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
7 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 4 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

CANDO is publicly available through Github at https://github.com/ram-compbio/CANDO. The code and data files are available https://doi.org/10.5061/dryad.g4f4qrg3j.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 7 authors, 6 keywords, 13 funders, 150 references.

Cite

This paper

Xu, S., Hu, Y., Mangione, W., Van Norden, M., Elefteriou, K., Falls, Z., & Samudrala, R. (2026). Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform. Journal of cheminformatics, 18(1), 65. https://doi.org/10.1186/s13321-026-01191-9

BibTeX

@article{xu2026multiscale,
author = {Xu, Sumei and Hu, Yakun and Mangione, William and Van Norden, Melissa and Elefteriou, Katherine and Falls, Zackary and Samudrala, Ram},
title = {{Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform}},
journal = {Journal of cheminformatics},
year = {2026},
month = apr,
volume = {18},
number = {1},
pages = {65},
publisher = {BMC},
issn = {1758-2946},
doi = {10.1186/s13321-026-01191-9},
url = {https://doi.org/10.1186/s13321-026-01191-9},
pmid = {41968358},
pmcid = {PMC13185303}
}

RIS

TY - JOUR
AU - Xu, Sumei
AU - Hu, Yakun
AU - Mangione, William
AU - Van Norden, Melissa
AU - Elefteriou, Katherine
AU - Falls, Zackary
AU - Samudrala, Ram
TI - Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform
T2 - Journal of cheminformatics
J2 - J Cheminform
PY - 2026
DA - 2026/04/12
VL - 18
IS - 1
SP - 65
SN - 1758-2946
PB - BMC
DO - 10.1186/s13321-026-01191-9
UR - https://doi.org/10.1186/s13321-026-01191-9
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s13321-026-01191-9",
"type": "article-journal",
"title": "Multiscale analysis and optimal glioma therapeutic candidate discovery using the CANDO platform",
"container-title": "Journal of cheminformatics",
"author": [
{
"family": "Xu",
"given": "Sumei"
},
{
"family": "Hu",
"given": "Yakun"
},
{
"family": "Mangione",
"given": "William"
},
{
"family": "Van Norden",
"given": "Melissa"
},
{
"family": "Elefteriou",
"given": "Katherine"
},
{
"family": "Falls",
"given": "Zackary"
},
{
"family": "Samudrala",
"given": "Ram"
}
],
"container-title-short": "J Cheminform",
"volume": "18",
"issue": "1",
"page": "65",
"DOI": "10.1186/s13321-026-01191-9",
"PMID": "41968358",
"PMCID": "PMC13185303",
"ISSN": "1758-2946",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s13321-026-01191-9",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/bioinformatics/btag153 [code]
MAISNet: a multi-species integrated graph neural network for acetylcholinesterase inhibitor screening.
Journal: Bioinformatics (Oxford, England)
In common: RDKit, scikit-learn, pandas, 3 other tools, clinical / translational, 1 reference
[2] doi:10.3390/ph19081319 [code]
DeepBBB: A Data-Composition-Aware Graph Screening Workflow for BBB-Focused CNS Library Construction and Prospective PAMPA-BBB Evaluation.
Journal: Pharmaceuticals (Basel, Switzerland)
In common: RDKit, scikit-learn, pandas, 2 other tools, clinical / translational, 1 reference
[3] doi:10.1038/s41586-026-10670-w [code]
Zero-shot design of drug-binding proteins via neural iterative selection-expansion.
Journal: Nature
In common: RDKit, scikit-learn, pandas, 3 other tools, 1 reference
[4] doi:10.1021/acs.biochem.5c00596 [code]
Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.
Journal: Biochemistry
In common: RDKit, scikit-learn, pandas, 3 other tools, 1 reference
[5] doi:10.1021/acs.jcim.6c01299 [code]
Targeting BCL-2 through Deep Learning-Based Drug Repurposing: A Multimodal Approach Combining Diffusion-Based Generative Modeling, Neural Relational Inference, and In Vitro Validation.
Journal: Journal of chemical information and modeling
In common: RDKit, pandas, SciPy, 2 other tools, other condition, 1 reference
[6] doi:10.1371/journal.pone.0345854 [code]
Shedding light on neural learning to rank models for anticancer drug prioritization.
Journal: PloS one
In common: RDKit, scikit-learn, pandas, 3 other tools, other condition
[7] doi:10.1038/s41598-026-53415-5 [code]
Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.
Journal: Scientific reports
In common: RDKit, scikit-learn, pandas, 3 other tools, other condition
[8] doi:10.1021/acsomega.5c09368 [code]
Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.
Journal: ACS omega
In common: RDKit, scikit-learn, pandas, 3 other tools, other condition
[9] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: RDKit, scikit-learn, pandas, 3 other tools
[10] doi:10.3390/ijms27156614 [code]
Candidalysin Inhibits <i>Porphyromonas gingivalis</i> Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions.
Journal: International journal of molecular sciences
In common: RDKit, scikit-learn, pandas, 3 other tools

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.