Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 498 lines · 23 KB · MIT
- # %% [markdown]
- # <a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/AlphaFold2.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>
- # %% [markdown]
- # <img src="https://raw.githubusercontent.com/sokrypton/ColabFold/main/.github/ColabFold_Marv_Logo_Small.png" height="200" align="right" style="height:240px">
- #
- # ##ColabFold v1.6.3: AlphaFold2 using MMseqs2
- #
- # Easy to use protein structure and complex prediction using [AlphaFold2](https://www.nature.com/articles/s41586-021-03819-2) and [Alphafold2-multimer](https://www.biorxiv.org/content/10.1101/2021.10.04.463034v1). Sequence alignments/templates are generated through [MMseqs2](mmseqs.com) and [HHsearch](https://github.com/soedinglab/hh-suite). For more details, see <a href="#Instructions">bottom</a> of the notebook, checkout the [ColabFold GitHub](https://github.com/sokrypton/ColabFold) and [Nature Protocols](https://www.nature.com/articles/s41596-024-01060-5).
- #
- # Old versions: [v1.4](https://colab.research.google.com/github/sokrypton/ColabFold/blob/v1.4.0/AlphaFold2.ipynb), [v1.5.1](https://colab.research.google.com/github/sokrypton/ColabFold/blob/v1.5.1/AlphaFold2.ipynb), [v1.5.2](https://colab.research.google.com/github/sokrypton/ColabFold/blob/v1.5.2/AlphaFold2.ipynb), [v1.5.3-patch](https://colab.research.google.com/github/sokrypton/ColabFold/blob/56c72044c7d51a311ca99b953a71e552fdc042e1/AlphaFold2.ipynb)
- #
- # [Mirdita M, Schütze K, Moriwaki Y, Heo L, Ovchinnikov S, Steinegger M. ColabFold: Making protein folding accessible to all.
- # *Nature Methods*, 2022](https://www.nature.com/articles/s41592-022-01488-1)
- # %%
- #@title Input protein sequence(s), then hit `Runtime` -> `Run all`
- from google.colab import files
- import os
- import re
- import hashlib
- import random
- def add_hash(x,y):
- return x+"_"+hashlib.sha1(y.encode()).hexdigest()[:5]
- query_sequence = 'PIAQIHILEGRSDEQKETLIREVSEAISRSLDAPLTSVRVIITEMAKGHFGIGGELASK' #@param {type:"string"}
- #@markdown - Use `:` to specify inter-protein chainbreaks for **modeling complexes** (supports homo- and hetro-oligomers). For example **PI...SK:PI...SK** for a homodimer
- jobname = 'test' #@param {type:"string"}
- # number of models to use
- num_relax = 0 #@param [0, 1, 5] {type:"raw"}
- #@markdown - specify how many of the top ranked structures to relax using amber
- template_mode = "none" #@param ["none", "pdb100","custom"]
- #@markdown - `none` = no template information is used. `pdb100` = detect templates in pdb100 (see [notes](#pdb100)). `custom` - upload and search own templates (PDB or mmCIF format, see [notes](#custom_templates))
- use_amber = num_relax > 0
- # remove whitespaces
- query_sequence = "".join(query_sequence.split())
- basejobname = "".join(jobname.split())
- basejobname = re.sub(r'\W+', '', basejobname)
- jobname = add_hash(basejobname, query_sequence)
- # check if directory with jobname exists
- def check(folder):
- if os.path.exists(folder):
- return False
- else:
- return True
- if not check(jobname):
- n = 0
- while not check(f"{jobname}_{n}"): n += 1
- jobname = f"{jobname}_{n}"
- # make directory to save results
- os.makedirs(jobname, exist_ok=True)
- # save queries
- queries_path = os.path.join(jobname, f"{jobname}.csv")
- with open(queries_path, "w") as text_file:
- text_file.write(f"id,sequence\n{jobname},{query_sequence}")
- if template_mode == "pdb100":
- use_templates = True
- custom_template_path = None
- elif template_mode == "custom":
- custom_template_path = os.path.join(jobname,f"template")
- os.makedirs(custom_template_path, exist_ok=True)
- uploaded = files.upload()
- use_templates = True
- for fn in uploaded.keys():
- os.rename(fn,os.path.join(custom_template_path,fn))
- else:
- custom_template_path = None
- use_templates = False
- print("jobname",jobname)
- print("sequence",query_sequence)
- print("length",len(query_sequence.replace(":","")))
- # %%
- #@title Install dependencies
- %%time
- import os
- USE_AMBER = use_amber
- EXTRAS = "alphafold-minus-jax,openmm" if USE_AMBER else "alphafold-minus-jax"
- OPENMM = ""
- if USE_AMBER:
- from importlib.metadata import distributions
- installed = {d.metadata["Name"] for d in distributions()}
- for cuda in ("cuda13", "cuda12"):
- if f"jax-{cuda}-plugin" in installed:
- OPENMM = f" 'openmm[{cuda}]'"
- break
- READY = "COLABFOLD_AMBER_READY" if USE_AMBER else "COLABFOLD_READY"
- if not os.path.isfile(READY):
- print("installing colabfold...")
- os.system(f"pip install -q --no-warn-conflicts 'colabfold[{EXTRAS}] @ git+https://github.com/sokrypton/ColabFold'{OPENMM}")
- if os.environ.get('TPU_NAME', False) != False:
- os.system("pip uninstall -y jax jaxlib")
- os.system("pip install --no-warn-conflicts --upgrade dm-haiku==0.0.10 'jax[cuda12_pip]'==0.3.25 -f https://storage.googleapis.com/jax-releases/jax_cuda_releases.html")
- os.system("ln -s /usr/local/lib/python3.*/dist-packages/colabfold colabfold")
- os.system("ln -s /usr/local/lib/python3.*/dist-packages/alphafold alphafold")
- os.system(f"touch {READY}")
- # %%
- #@markdown ### MSA options (custom MSA upload, single sequence, pairing mode)
- msa_mode = "mmseqs2_uniref_env" #@param ["mmseqs2_uniref_env", "mmseqs2_uniref","single_sequence","custom"]
- pair_mode = "unpaired_paired" #@param ["unpaired_paired","paired","unpaired"] {type:"string"}
- #@markdown - "unpaired_paired" = pair sequences from same species + unpaired MSA, "unpaired" = seperate MSA for each chain, "paired" - only use paired sequences.
- # decide which a3m to use
- if "mmseqs2" in msa_mode:
- a3m_file = os.path.join(jobname,f"{jobname}.a3m")
- elif msa_mode == "custom":
- a3m_file = os.path.join(jobname,f"{jobname}.custom.a3m")
- if not os.path.isfile(a3m_file):
- custom_msa_dict = files.upload()
- custom_msa = list(custom_msa_dict.keys())[0]
- header = 0
- import fileinput
- for line in fileinput.FileInput(custom_msa,inplace=1):
- if line.startswith(">"):
- header = header + 1
- if not line.rstrip():
- continue
- if line.startswith(">") == False and header == 1:
- query_sequence = line.rstrip()
- print(line, end='')
- os.rename(custom_msa, a3m_file)
- queries_path=a3m_file
- print(f"moving {custom_msa} to {a3m_file}")
- else:
- a3m_file = os.path.join(jobname,f"{jobname}.single_sequence.a3m")
- with open(a3m_file, "w") as text_file:
- text_file.write(">1\n%s" % query_sequence)
- # %%
- #@markdown ### Advanced settings
- model_type = "auto" #@param ["auto", "alphafold2_ptm", "alphafold2_multimer_v1", "alphafold2_multimer_v2", "alphafold2_multimer_v3", "deepfold_v1", "alphafold2"]
- #@markdown - if `auto` selected, will use `alphafold2_ptm` for monomer prediction and `alphafold2_multimer_v3` for complex prediction.
- #@markdown Any of the mode_types can be used (regardless if input is monomer or complex).
- num_recycles = "3" #@param ["auto", "0", "1", "3", "6", "12", "24", "48"]
- #@markdown - if `auto` selected, will use `num_recycles=20` if `model_type=alphafold2_multimer_v3`, else `num_recycles=3` .
- recycle_early_stop_tolerance = "auto" #@param ["auto", "0.0", "0.5", "1.0"]
- #@markdown - if `auto` selected, will use `tol=0.5` if `model_type=alphafold2_multimer_v3` else `tol=0.0`.
- relax_max_iterations = 200 #@param [0, 200, 2000] {type:"raw"}
- #@markdown - max amber relax iterations, `0` = unlimited (AlphaFold2 default, can take very long)
- pairing_strategy = "greedy" #@param ["greedy", "complete"] {type:"string"}
- #@markdown - `greedy` = pair any taxonomically matching subsets, `complete` = all sequences have to match in one line.
- calc_extra_ptm = False #@param {type:"boolean"}
- #@markdown - return pairwise chain iptm/actifptm
- use_fast_kernels = False #@param {type:"boolean"}
- #@markdown - use fused kernels for faster prediction (experimental)
- #@markdown #### Sample settings
- #@markdown - enable dropouts and increase number of seeds to sample predictions from uncertainty of the model.
- #@markdown - decrease `max_msa` to increase uncertainity
- max_msa = "auto" #@param ["auto", "512:1024", "256:512", "64:128", "32:64", "16:32"]
- num_seeds = 1 #@param [1,2,4,8,16] {type:"raw"}
- use_dropout = False #@param {type:"boolean"}
- num_recycles = None if num_recycles == "auto" else int(num_recycles)
- recycle_early_stop_tolerance = None if recycle_early_stop_tolerance == "auto" else float(recycle_early_stop_tolerance)
- if max_msa == "auto": max_msa = None
- #@markdown #### Save settings
- save_all = False #@param {type:"boolean"}
- save_recycles = False #@param {type:"boolean"}
- save_to_google_drive = False #@param {type:"boolean"}
- #@markdown - if the save_to_google_drive option was selected, the result zip will be uploaded to your Google Drive
- dpi = 200 #@param {type:"integer"}
- #@markdown - set dpi for image resolution
- if save_to_google_drive:
- from pydrive2.drive import GoogleDrive
- from pydrive2.auth import GoogleAuth
- from google.colab import auth
- from oauth2client.client import GoogleCredentials
- auth.authenticate_user()
- gauth = GoogleAuth()
- gauth.credentials = GoogleCredentials.get_application_default()
- drive = GoogleDrive(gauth)
- print("You are logged into Google Drive and are good to go!")
- #@markdown Don't forget to hit `Runtime` -> `Run all` after updating the form.
- # %%
- #@title Run Prediction
- display_images = True #@param {type:"boolean"}
- import sys
- import warnings
- warnings.simplefilter(action='ignore', category=FutureWarning)
- from Bio import BiopythonDeprecationWarning
- warnings.simplefilter(action='ignore', category=BiopythonDeprecationWarning)
- from pathlib import Path
- from colabfold.download import download_alphafold_params, default_data_dir
- from colabfold.utils import setup_logging
- from colabfold.batch import get_queries, run, set_model_type
- from colabfold.plot import plot_msa_v2
- import os
- import numpy as np
- try:
- K80_chk = os.popen('nvidia-smi | grep "Tesla K80" | wc -l').read()
- except:
- K80_chk = "0"
- pass
- if "1" in K80_chk:
- print("WARNING: found GPU Tesla K80: limited to total length < 1000")
- if "TF_FORCE_UNIFIED_MEMORY" in os.environ:
- del os.environ["TF_FORCE_UNIFIED_MEMORY"]
- if "XLA_PYTHON_CLIENT_MEM_FRACTION" in os.environ:
- del os.environ["XLA_PYTHON_CLIENT_MEM_FRACTION"]
- from colabfold.colabfold import plot_protein
- from pathlib import Path
- import matplotlib.pyplot as plt
- def input_features_callback(input_features):
- if display_images:
- plot_msa_v2(input_features)
- plt.show()
- plt.close()
- def prediction_callback(protein_obj, length,
- prediction_result, input_features, mode):
- model_name, relaxed = mode
- if not relaxed:
- if display_images:
- fig = plot_protein(protein_obj, Ls=length, dpi=150)
- plt.show()
- plt.close()
- result_dir = jobname
- log_filename = os.path.join(jobname,"log.txt")
- setup_logging(Path(log_filename))
- queries, is_complex = get_queries(queries_path)
- model_type = set_model_type(is_complex, model_type)
- if "multimer" in model_type and max_msa is not None:
- use_cluster_profile = False
- else:
- use_cluster_profile = True
- download_alphafold_params(model_type, Path("."))
- results = run(
- queries=queries,
- result_dir=result_dir,
- use_templates=use_templates,
- custom_template_path=custom_template_path,
- num_relax=num_relax,
- msa_mode=msa_mode,
- model_type=model_type,
- num_models=5,
- num_recycles=num_recycles,
- relax_max_iterations=relax_max_iterations,
- recycle_early_stop_tolerance=recycle_early_stop_tolerance,
- num_seeds=num_seeds,
- use_dropout=use_dropout,
- use_fast_kernels=use_fast_kernels,
- model_order=[1,2,3,4,5],
- is_complex=is_complex,
- data_dir=Path("."),
- keep_existing_results=False,
- rank_by="auto",
- pair_mode=pair_mode,
- pairing_strategy=pairing_strategy,
- stop_at_score=float(100),
- prediction_callback=prediction_callback,
- dpi=dpi,
- zip_results=False,
- save_all=save_all,
- max_msa=max_msa,
- use_cluster_profile=use_cluster_profile,
- input_features_callback=input_features_callback,
- save_recycles=save_recycles,
- user_agent="colabfold/google-colab-main",
- calc_extra_ptm=calc_extra_ptm,
- )
- results_zip = f"{jobname}.result.zip"
- os.system(f"zip -r {results_zip} {jobname}")
- # %%
- #@title Display 3D structure {run: "auto"}
- import py3Dmol
- import glob
- import matplotlib.pyplot as plt
- from colabfold.colabfold import plot_plddt_legend
- from colabfold.colabfold import pymol_color_list, alphabet_list
- rank_num = 1 #@param ["1", "2", "3", "4", "5"] {type:"raw"}
- color = "lDDT" #@param ["chain", "lDDT", "rainbow"]
- show_sidechains = False #@param {type:"boolean"}
- show_mainchains = False #@param {type:"boolean"}
- tag = results["rank"][0][rank_num - 1]
- jobname_prefix = ".custom" if msa_mode == "custom" else ""
- pdb_filename = f"{jobname}/{jobname}{jobname_prefix}_unrelaxed_{tag}.pdb"
- pdb_file = glob.glob(pdb_filename)
- def show_pdb(rank_num=1, show_sidechains=False, show_mainchains=False, color="lDDT"):
- model_name = f"rank_{rank_num}"
- view = py3Dmol.view(js='https://3dmol.org/build/3Dmol.js',)
- view.addModel(open(pdb_file[0],'r').read(),'pdb')
- if color == "lDDT":
- view.setStyle({'cartoon': {'colorscheme': {'prop':'b','gradient': 'roygb','min':50,'max':90}}})
- elif color == "rainbow":
- view.setStyle({'cartoon': {'color':'spectrum'}})
- elif color == "chain":
- chains = len(queries[0][1]) + 1 if is_complex else 1
- for n,chain,color in zip(range(chains),alphabet_list,pymol_color_list):
- view.setStyle({'chain':chain},{'cartoon': {'color':color}})
- if show_sidechains:
- BB = ['C','O','N']
- view.addStyle({'and':[{'resn':["GLY","PRO"],'invert':True},{'atom':BB,'invert':True}]},
- {'stick':{'colorscheme':f"WhiteCarbon",'radius':0.3}})
- view.addStyle({'and':[{'resn':"GLY"},{'atom':'CA'}]},
- {'sphere':{'colorscheme':f"WhiteCarbon",'radius':0.3}})
- view.addStyle({'and':[{'resn':"PRO"},{'atom':['C','O'],'invert':True}]},
- {'stick':{'colorscheme':f"WhiteCarbon",'radius':0.3}})
- if show_mainchains:
- BB = ['C','O','N','CA']
- view.addStyle({'atom':BB},{'stick':{'colorscheme':f"WhiteCarbon",'radius':0.3}})
- view.zoomTo()
- return view
- show_pdb(rank_num, show_sidechains, show_mainchains, color).show()
- if color == "lDDT":
- plot_plddt_legend().show()
- # %%
- #@title Plots {run: "auto"}
- from IPython.display import display, HTML
- import base64
- from html import escape
- # see: https://stackoverflow.com/a/53688522
- def image_to_data_url(filename):
- ext = filename.split('.')[-1]
- prefix = f'data:image/{ext};base64,'
- with open(filename, 'rb') as f:
- img = f.read()
- return prefix + base64.b64encode(img).decode('utf-8')
- pae = ""
- pae_file = os.path.join(jobname,f"{jobname}{jobname_prefix}_pae.png")
- if os.path.isfile(pae_file):
- pae = image_to_data_url(pae_file)
- cov = image_to_data_url(os.path.join(jobname,f"{jobname}{jobname_prefix}_coverage.png"))
- plddt = image_to_data_url(os.path.join(jobname,f"{jobname}{jobname_prefix}_plddt.png"))
- display(HTML(f"""
- <style>
- img {{
- float:left;
- }}
- .full {{
- max-width:100%;
- }}
- .half {{
- max-width:50%;
- }}
- @media (max-width:640px) {{
- .half {{
- max-width:100%;
- }}
- }}
- </style>
- <div style="max-width:90%; padding:2em;">
- <h1>Plots for {escape(jobname)}</h1>
- { '<!--' if pae == '' else '' }<img src="{pae}" class="full" />{ '-->' if pae == '' else '' }
- <img src="{cov}" class="half" />
- <img src="{plddt}" class="half" />
- </div>
- """))
- # %%
- #@title Package and download results
- #@markdown If you are having issues downloading the result archive, try disabling your adblocker and run this cell again. If that fails click on the little folder icon to the left, navigate to file: `jobname.result.zip`, right-click and select \"Download\" (see [screenshot](https://pbs.twimg.com/media/E6wRW2lWUAEOuoe?format=jpg&name=small)).
- if msa_mode == "custom":
- print("Don't forget to cite your custom MSA generation method.")
- files.download(f"{jobname}.result.zip")
- if save_to_google_drive == True and drive:
- uploaded = drive.CreateFile({'title': f"{jobname}.result.zip"})
- uploaded.SetContentFile(f"{jobname}.result.zip")
- uploaded.Upload()
- print(f"Uploaded {jobname}.result.zip to Google Drive with ID {uploaded.get('id')}")
- # %% [markdown]
- # # Instructions <a name="Instructions"></a>
- # For detailed instructions, tips and tricks, see recently published paper at [Nature Protocols](https://www.nature.com/articles/s41596-024-01060-5)
- #
- # **Quick start**
- # 1. Paste your protein sequence(s) in the input field.
- # 2. Press "Runtime" -> "Run all".
- # 3. The pipeline consists of 5 steps. The currently running step is indicated by a circle with a stop sign next to it.
- #
- # **Result zip file contents**
- #
- # 1. PDB formatted structures sorted by avg. pLDDT and complexes are sorted by pTMscore. (unrelaxed and relaxed if `use_amber` is enabled).
- # 2. Plots of the model quality.
- # 3. Plots of the MSA coverage.
- # 4. Parameter log file.
- # 5. A3M formatted input MSA.
- # 6. A `predicted_aligned_error_v1.json` using [AlphaFold-DB's format](https://alphafold.ebi.ac.uk/faq#faq-7) and a `scores.json` for each model which contains an array (list of lists) for PAE, a list with the average pLDDT and the pTMscore.
- # 7. BibTeX file with citations for all used tools and databases.
- #
- # At the end of the job a download modal box will pop up with a `jobname.result.zip` file. Additionally, if the `save_to_google_drive` option was selected, the `jobname.result.zip` will be uploaded to your Google Drive.
- #
- # **MSA generation for complexes**
- #
- # For the complex prediction we use unpaired and paired MSAs. Unpaired MSA is generated the same way as for the protein structures prediction by searching the UniRef100 and environmental sequences three iterations each.
- #
- # The paired MSA is generated by searching the UniRef100 database and pairing the best hits sharing the same NCBI taxonomic identifier (=species or sub-species). We only pair sequences if all of the query sequences are present for the respective taxonomic identifier.
- #
- # **Using a custom MSA as input**
- #
- # To predict the structure with a custom MSA (A3M formatted): (1) Change the `msa_mode`: to "custom", (2) Wait for an upload box to appear at the end of the "MSA options ..." box. Upload your A3M. The first fasta entry of the A3M must be the query sequence without gaps.
- #
- # It is also possilbe to provide custom MSAs for complex predictions. Read more about the format [here](https://github.com/sokrypton/ColabFold/issues/76).
- #
- # As an alternative for MSA generation the [HHblits Toolkit server](https://toolkit.tuebingen.mpg.de/tools/hhblits) can be used. After submitting your query, click "Query Template MSA" -> "Download Full A3M". Download the A3M file and upload it in this notebook.
- #
- # **PDB100** <a name="pdb100"></a>
- #
- # As of 23/06/08, we have transitioned from using the PDB70 to a 100% clustered PDB, the PDB100. The construction methodology of PDB100 differs from that of PDB70.
- #
- # The PDB70 was constructed by running each PDB70 representative sequence through [HHblits](https://github.com/soedinglab/hh-suite) against the [Uniclust30](https://uniclust.mmseqs.com/). On the other hand, the PDB100 is built by searching each PDB100 representative structure with [Foldseek](https://github.com/steineggerlab/foldseek) against the [AlphaFold Database](https://alphafold.ebi.ac.uk).
- #
- # To maintain compatibility with older Notebook versions and local installations, the generated files and API responses will continue to be named "PDB70", even though we're now using the PDB100.
- #
- # **Using custom templates** <a name="custom_templates"></a>
- #
- # To predict the structure with a custom template (PDB or mmCIF formatted): (1) change the `template_mode` to "custom" in the execute cell and (2) wait for an upload box to appear at the end of the "Input Protein" box. Select and upload your templates (multiple choices are possible).
- #
- # * Templates must follow the four letter PDB naming with lower case letters.
- #
- # * Templates in mmCIF format must contain `_entity_poly_seq`. An error is thrown if this field is not present. The field `_pdbx_audit_revision_history.revision_date` is automatically generated if it is not present.
- #
- # * Templates in PDB format are automatically converted to the mmCIF format. `_entity_poly_seq` and `_pdbx_audit_revision_history.revision_date` are automatically generated.
- #
- # If you encounter problems, please report them to this [issue](https://github.com/sokrypton/ColabFold/issues/177).
- #
- # **Comparison to the full AlphaFold2 and AlphaFold2 Colab**
- #
- # This notebook replaces the homology detection and MSA pairing of AlphaFold2 with MMseqs2. For a comparison against the [AlphaFold2 Colab](https://colab.research.google.com/github/deepmind/alphafold/blob/main/notebooks/AlphaFold.ipynb) and the full [AlphaFold2](https://github.com/deepmind/alphafold) system read our [paper](https://www.nature.com/articles/s41592-022-01488-1).
- #
- # **Troubleshooting**
- # * Check that the runtime type is set to GPU at "Runtime" -> "Change runtime type".
- # * Try to restart the session "Runtime" -> "Factory reset runtime".
- # * Check your input sequence.
- #
- # **Known issues**
- # * Google Colab assigns different types of GPUs with varying amount of memory. Some might not have enough memory to predict the structure for a long sequence.
- # * Your browser can block the pop-up for downloading the result file. You can choose the `save_to_google_drive` option to upload to Google Drive instead or manually download the result file: Click on the little folder icon to the left, navigate to file: `jobname.result.zip`, right-click and select \"Download\" (see [screenshot](https://pbs.twimg.com/media/E6wRW2lWUAEOuoe?format=jpg&name=small)).
- #
- # **Limitations**
- # * Computing resources: Our MMseqs2 API can handle ~20-50k requests per day.
- # * MSAs: MMseqs2 is very precise and sensitive but might find less hits compared to HHblits/HMMer searched against BFD or MGnify.
- # * We recommend to additionally use the full [AlphaFold2 pipeline](https://github.com/deepmind/alphafold).
- #
- # **Description of the plots**
- # * **Number of sequences per position** - We want to see at least 30 sequences per position, for best performance, ideally 100 sequences.
- # * **Predicted lDDT per position** - model confidence (out of 100) at each position. The higher the better.
- # * **Predicted Alignment Error** - For homooligomers, this could be a useful metric to assess how confident the model is about the interface. The lower the better.
- #
- # **Bugs**
- # - If you encounter any bugs, please report the issue to https://github.com/sokrypton/ColabFold/issues
- #
- # **License**
- #
- # The source code of ColabFold is licensed under [MIT](https://raw.githubusercontent.com/sokrypton/ColabFold/main/LICENSE). Additionally, this notebook uses the AlphaFold2 source code and its parameters licensed under [Apache 2.0](https://raw.githubusercontent.com/deepmind/alphafold/main/LICENSE) and [CC BY 4.0](https://creativecommons.org/licenses/by-sa/4.0/) respectively. Read more about the AlphaFold license [here](https://github.com/deepmind/alphafold).
- #
- # **Acknowledgments**
- # - We thank the AlphaFold team for developing an excellent model and open sourcing the software.
- #
- # - [KOBIC](https://kobic.re.kr) and [Söding Lab](https://www.mpinat.mpg.de/soeding) for providing the computational resources for the MMseqs2 MSA server.
- #
- # - Richard Evans for helping to benchmark the ColabFold's Alphafold-multimer support.
- #
- # - [David Koes](https://github.com/dkoes) for his awesome [py3Dmol](https://3dmol.csb.pitt.edu/) plugin, without whom these notebooks would be quite boring!
- #
- # - Do-Yoon Kim for creating the ColabFold logo.
- #
- # - A colab by Sergey Ovchinnikov ([@sokrypton](https://twitter.com/sokrypton)), Milot Mirdita ([@milot_mirdita](https://twitter.com/milot_mirdita)) and Martin Steinegger ([@thesteinegger](https://twitter.com/thesteinegger)).
AlphaFold2.ipynb at commit efbf31c, under MIT · at the source
Overview
and 9 other authors
Ji Hoon Han1,2, Annalisa Bucci1, Christoph Hess5,6, Simone Picelli1, Magdalena Renner1, Daniel J Müller3, Cameron S Cowan1, Simon Hansen1, Botond Roska1,2- Institute of Molecular and Clinical Ophthalmology Basel, Basel, Switzerland
- Department of Ophthalmology, University of Basel, Basel, Switzerland
- Department of Biosystems Science and Engineering, ETH Zürich, Basel, Switzerland
- Facility for Advanced Imaging and Microscopy, Friedrich Miescher Institute for Biomedical Research, Basel, Switzerland
- Immunobiology Laboratory, Department of Biomedicine, University of Basel and University Hospital of Basel, Basel, Switzerland
- Cambridge Institute of Therapeutic Immunology & Infectious Disease (CITIID), Department of Medicine, University of Cambridge, Cambridge, UK
Abstract
The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.
Repositories
Its files are read in the Code ↔ Paper reader above.
sokrypton/colabfold
efbf31c37cedb38cd09c69c1b991910a9866480e, 19 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
64 files
- AlphaFold2.ipynb, Jupyter, 498 lines
- AlphaFold3_of3.ipynb, Jupyter, 554 lines
- BioEmu.ipynb, Jupyter, 390 lines
- Boltz1.ipynb, Jupyter, 258 lines
- ColabFold2_preview.ipynb
, Jupyter, 666 lines - ESMFold.ipynb, Jupyter, 265 lines
- MsaServer/
restart-systemd.sh , Shell, 5 lines - MsaServer/
setup-and-start-local.sh , Shell, 111 lines - RoseTTAFold.ipynb, Jupyter, 213 lines
- RoseTTAFold2.ipynb, Jupyter, 411 lines
- batch/
AlphaFold2_batch.ipynb , Jupyter, 180 lines - beta/
AlphaFold2_advanced.ipyn , Jupyter, 615 linesb - beta/
AlphaFold2_advanced_beta , Jupyter, 7 lines.ipynb - beta/
AlphaFold2_advanced_old. , Jupyter, 1,021 linesipynb - beta/
AlphaFold2_complexes.ipy , Jupyter, 514 linesnb - beta/
AlphaFold_wJackhmmer.ipy , Jupyter, 697 linesnb - beta/
Alphafold_single.ipynb , Jupyter, 207 lines - beta/
ESMFold.ipynb , Jupyter, 385 lines - beta/
ESMFold_advanced.ipynb , Jupyter, 496 lines - beta/
ESMFold_api.ipynb , Jupyter, 104 lines - beta/
RoseTTAFold.ipynb , Jupyter, 300 lines - beta/
RoseTTAFold_install.sh , Shell, 5 lines - beta/
RoseTTAFold_run.sh , Shell, 26 lines - beta/
alphafold_output_at_each , Jupyter, 152 lines_recycle.ipynb - beta/
colabfold.py , Python, 711 lines - beta/
colabfold_alphafold.py , Python, 821 lines - beta/
convert_256_to_384_rep.i , Jupyter, 39 linespynb - beta/
omegafold.ipynb , Jupyter, 167 lines - beta/
omegafold_hacks.ipynb , Jupyter, 206 lines - beta/
pairmsa.py , Python, 237 lines - beta/
relax_amber.ipynb , Jupyter, 154 lines - colabfold/
__init__.py , Python, 1 line - colabfold/
alphafold/ , Python, 1 line__init__.py - colabfold/
alphafold/ , Python, 435 linesextra_ptm.py - colabfold/
alphafold/ , Python, 98 linesipsae.py - colabfold/
alphafold/ , Python, 277 linesmodels.py - colabfold/
alphafold/ , Python, 44 linesmsa.py - colabfold/
batch.py , Python, 2,364 lines - colabfold/
citations.py , Python, 160 lines - colabfold/
colabfold.py , Python, 831 lines - colabfold/
download.py , Python, 140 lines - colabfold/
input.py , Python, 414 lines - colabfold/
mmseqs/ , Python, 1 line__init__.py - colabfold/
mmseqs/ , Python, 63 linesmerge_and_split_msas.py - colabfold/
mmseqs/ , Python, 640 linessearch.py - colabfold/
mmseqs/ , Python, 55 linessplit_msas.py - colabfold/
pdb.py , Python, 69 lines - colabfold/
plot.py , Python, 158 lines - colabfold/
relax.py , Python, 108 lines - colabfold/
utils.py , Python, 406 lines - colabfold_search.sh, Shell, 68 lines
- setup_databases.sh, Shell, 218 lines
- tests/
__init__.py , Python, 1 line - tests/
mock.py , Python, 230 lines - tests/
reindent_ipynb.py , Python, 8 lines - tests/
test_colabfold.py , Python, 510 lines - tests/
test_msa.py , Python, 40 lines - tests/
test_utils.py , Python, 141 lines - utils/
convert_deepfold_weights , Python, 9 lines.py - utils/
plot_scores.ipynb , Jupyter, 28 lines - verbose/
alphafold_noTemplates_no , Jupyter, 267 linesMD.ipynb - verbose/
alphafold_noTemplates_ye , Jupyter, 323 linessMD.ipynb - LICENSE, License, 21 lines
- README.md, Text, 293 lines
alphafoldserver.com
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
- 29 September 2026: the link answers (HTTP 200)
Zenodo 17909630
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
- 29 September 2026: the link answers (HTTP 200)
Code availability statement
The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to the authors' code: Zenodo 17909630
- it says that the code is available on request
Read it in the paper: doi.org/10.1038/s41586-026-10391-0.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 62 scripts, each with its path and the digest of its content;
- no match between paragraphs and code yet;
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- bioproject:PRJNA1380367, at NCBI BioProject; found in “Data availability”
- uniprot.org/
uniprot/ , at UniProt; found in the text, “Expression and purification of TOMM20…”q15388
Data availability statement
The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:
- it points to a dataset: NCBI BioProject PRJNA1380367
- it says that the data are available on request
Read it in the paper: doi.org/10.1038/s41586-026-10391-0.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 29 authors, 2 keywords, 12 MeSH terms, 2 funders, 94 references.
Cite
This paper
Ayupov, T., Moreno-Juan, V., Curtoni, S., Fratzl, A., Sharma, U., Posada-Céspedes, S., Ratiu, R., Morikawa, R., Meyer, A. G., Pezzoli, M., Bantug, G., Chevalier, M., Hou, Y., Nadeau, S. A., Herrero-Navarro, Á., Ayinampudi, V., Kastanaki, E., Whitehead, N., Siwicki, R. A., . . . Roska, B. (2026). Cell-type-targeted mitochondrial transplantation rescues cell degeneration. Nature, 653(8113), 221-231. https://
BibTeX
@article{ayupov2026cell,
author = {Ayupov, Temurkhan and Moreno-Juan, Verónica and Curtoni, Serena and Fratzl, Alex and Sharma, Upnishad and Posada-Céspedes, Susana and Ratiu, Ramona and Morikawa, Rei and Meyer, Alexandra Graff and Pezzoli, Margherita and Bantug, Glenn and Chevalier, Morgan and Hou, Yanyan and Nadeau, Sarah A and Herrero-Navarro, Álvaro and Ayinampudi, Vikram and Kastanaki, Elizabeth and Whitehead, Natasha and Siwicki, Rebecca A and Ribeiro, Mariana M and Han, Ji Hoon and Bucci, Annalisa and Hess, Christoph and Picelli, Simone and Renner, Magdalena and Müller, Daniel J and Cowan, Cameron S and Hansen, Simon and Roska, Botond},
title = {{Cell-type-targeted mitochondrial transplantation rescues cell degeneration}},
journal = {Nature},
year = {2026},
month = apr,
volume = {653},
number = {8113},
pages = {221--231},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/
url = {https://
pmid = {41986718},
pmcid = {PMC13149334}
}
RIS
TY - JOUR
AU - Ayupov, Temurkhan
AU - Moreno-Juan, Verónica
AU - Curtoni, Serena
AU - Fratzl, Alex
AU - Sharma, Upnishad
AU - Posada-Céspedes, Susana
AU - Ratiu, Ramona
AU - Morikawa, Rei
AU - Meyer, Alexandra Graff
AU - Pezzoli, Margherita
AU - Bantug, Glenn
AU - Chevalier, Morgan
AU - Hou, Yanyan
AU - Nadeau, Sarah A
AU - Herrero-Navarro, Álvaro
AU - Ayinampudi, Vikram
AU - Kastanaki, Elizabeth
AU - Whitehead, Natasha
AU - Siwicki, Rebecca A
AU - Ribeiro, Mariana M
AU - Han, Ji Hoon
AU - Bucci, Annalisa
AU - Hess, Christoph
AU - Picelli, Simone
AU - Renner, Magdalena
AU - Müller, Daniel J
AU - Cowan, Cameron S
AU - Hansen, Simon
AU - Roska, Botond
TI - Cell-type-targeted mitochondrial transplantation rescues cell degeneration
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/
VL - 653
IS - 8113
SP - 221
EP - 231
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Cell-type-targeted mitochondrial transplantation rescues cell degeneration",
"container-title": "Nature",
"author": [
{
"family": "Ayupov",
"given": "Temurkhan"
},
{
"family": "Moreno-Juan",
"given": "Verónica"
},
{
"family": "Curtoni",
"given": "Serena"
},
{
"family": "Fratzl",
"given": "Alex"
},
{
"family": "Sharma",
"given": "Upnishad"
},
{
"family": "Posada-Céspedes",
"given": "Susana"
},
{
"family": "Ratiu",
"given": "Ramona"
},
{
"family": "Morikawa",
"given": "Rei"
},
{
"family": "Meyer",
"given": "Alexandra Graff"
},
{
"family": "Pezzoli",
"given": "Margherita"
},
{
"family": "Bantug",
"given": "Glenn"
},
{
"family": "Chevalier",
"given": "Morgan"
},
{
"family": "Hou",
"given": "Yanyan"
},
{
"family": "Nadeau",
"given": "Sarah A"
},
{
"family": "Herrero-Navarro",
"given": "Álvaro"
},
{
"family": "Ayinampudi",
"given": "Vikram"
},
{
"family": "Kastanaki",
"given": "Elizabeth"
},
{
"family": "Whitehead",
"given": "Natasha"
},
{
"family": "Siwicki",
"given": "Rebecca A"
},
{
"family": "Ribeiro",
"given": "Mariana M"
},
{
"family": "Han",
"given": "Ji Hoon"
},
{
"family": "Bucci",
"given": "Annalisa"
},
{
"family": "Hess",
"given": "Christoph"
},
{
"family": "Picelli",
"given": "Simone"
},
{
"family": "Renner",
"given": "Magdalena"
},
{
"family": "Müller",
"given": "Daniel J"
},
{
"family": "Cowan",
"given": "Cameron S"
},
{
"family": "Hansen",
"given": "Simon"
},
{
"family": "Roska",
"given": "Botond"
}
],
"container-title-short":
"volume": "653",
"issue": "8113",
"page": "221-231",
"DOI": "10.1038/
"PMID": "41986718",
"PMCID": "PMC13149334",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
15
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.3390/ijms27156614 [code]
- Candidalysin Inhibits &
lt;i& gt;Porphyromonas gingivalis& lt;/ i& gt; Lipoprotein-Induced IL-1β Production in BV-2 Microglia via Hydrophobic Microbial Interactions. Journal: International journal of molecular sciencesIn common: RDKit, JAX, PyTorch Lightning, 9 other tools, mouse, cellular / molecular, 1 reference - [2] doi:10.1038/s41598-026-53415-5 [code]
- Computational design and immunoinformatics validation of a T cell multi-epitope vaccine targeting glioblastoma stem cells.Journal: Scientific reportsIn common: RDKit, JAX, PyTorch Lightning, 9 other tools, cellular / molecular, 1 reference
- [3] doi:10.1021/acs.biochem.5c00596 [code]
- Cargo Recognition of Nesprin-2 by the Dynein Adapter Bicaudal D2 for a Nuclear Positioning Pathway That Is Important for Brain Development.Journal: BiochemistryIn common: RDKit, JAX, PyTorch Lightning, 9 other tools, cellular / molecular, 1 reference
- [4] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: RDKit, JAX, PyTorch Lightning, 8 other tools, mouse, cellular / molecular
- [5] doi:10.1016/j.molcel.2026.07.006 [code]
- DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.Journal: Molecular cellIn common: JAX, Biopython, PyTorch Geometric, 7 other tools, cellular / molecular, 2 references
- [6] doi:10.1016/j.xcrm.2026.102904 [code]
- High-dose furmonertinib as first-line treatment for untreated EGFR-mutated advanced NSCLC with central nervous system metastases: A phase 2 trial.Journal: Cell reports. MedicineIn common: PyTorch Lightning, Biopython, PyTorch, 4 other tools, 3 references
- [7] doi:10.1038/s41467-026-75444-4 [code]
- Structural insights enable drug discovery for the neuronal NBCn2 carbonate transporter.Journal: Nature communicationsIn common: RDKit, JAX, Biopython, 4 other tools, mouse, cellular / molecular, 1 reference
- [8] doi:10.1038/s41467-026-72130-3 [code]
- Retinoic acid drives cell fate specification, maturation and retinal regionality in human retinal organoids.Journal: Nature communicationsIn common: PyTorch, SciPy, Matplotlib, 1 other tool, 5 references
- [9] doi:10.1021/acsomega.5c09368 [code]
- Structure-Based and AI-Assisted Identification of AGPS Inhibitors for Glioma via Integrated Docking, Molecular Dynamics, and Binding Affinity Screening.Journal: ACS omegaIn common: RDKit, Biopython, PyTorch Geometric, 6 other tools
- [10] doi:10.1523/eneuro.0362-25.2026 [code]
- Similarities between &
lt;i& gt;Ciona& lt;/ i& gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/ V Subdivision within a Hindbrain Precursor? Journal: eNeuroIn common: PyTorch Lightning, Biopython, PyTorch Geometric, 6 other tools
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 62 scripts, and 0 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:eb45cf0d1c05d4d7…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
