Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder.
The 2 matches
- [1] § 2. Materials and Methods › 2.6. Bioinformatic and Statistical Analysis ↔ humann/humann.py, lines 3–36 · score 0.66 · Unified Metabolic, HUMAnN, pipeline, Network, diamond, Metagenomic
- [2] § 2. Materials and Methods › 2.6. Bioinformatic and Statistical Analysis ↔ humann/search/prescreen.py, lines 245–386 · score 0.65 · MetaPhlAn, taxonomic profiles, mpa, gene families, chocophlansgb, v4
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 1,151 lines · 50 KB · MIT · 1 match
- #!/usr/bin/env python
- """
- HUMAnN : HMP Unified Metabolic Analysis Network 3
- HUMAnN is a pipeline for efficiently and accurately determining
- the coverage and abundance of microbial pathways in a community
- from metagenomic data. Sequencing a metagenome typically produces millions
- of short DNA/RNA reads.
- Dependencies: MetaPhlAn2, ChocoPhlAn, Bowtie2, and ( Diamond or Rapsearch2 or Usearch )
- To Run: humann -i <input.fastq> -o <output_dir>
- Copyright (c) 2014 Harvard School of Public Health
- Permission is hereby granted, free of charge, to any person obtaining a copy
- of this software and associated documentation files (the "Software"), to deal
- in the Software without restriction, including without limitation the rights
- to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
- copies of the Software, and to permit persons to whom the Software is
- furnished to do so, subject to the following conditions:
- The above copyright notice and this permission notice shall be included in
- all copies or substantial portions of the Software.
- THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
- IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
- FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
- AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
- LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
- OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
- THE SOFTWARE.
- """
- import sys
- # Try to load one of the humann modules to check the installation
- try:
- from . import check
- except ImportError:
- sys.exit("CRITICAL ERROR: Unable to find the HUMAnN python package." +
- " Please check your install.")
- # Check the python version
- check.python_version()
- import argparse
- import subprocess
- import os
- import time
- import tempfile
- import re
- import logging
- from . import config
- from . import store
- from . import utilities
- from .search import prescreen
- from .search import nucleotide
- from .search import translated
- from .quantify import families
- from .quantify import modules
- # name global logging instance
- logger=logging.getLogger(__name__)
- VERSION="4.0.0.alpha.2"
- MAX_SIZE_DEMO_INPUT_FILE=35
- def parse_arguments(args):
- """
- Parse the arguments from the user
- """
- parser = argparse.ArgumentParser(
- description= "HUMAnN : HMP Unified Metabolic Analysis Network 3\n",
- formatter_class=argparse.RawTextHelpFormatter,
- prog="humann")
- common_settings=parser.add_argument_group("[0] Common settings")
- common_settings.add_argument(
- "-i", "--input",
- help="input file of type {" +",".join(config.input_format_choices)+ "} \n[REQUIRED]",
- metavar="<input.fastq>",
- required=True)
- common_settings.add_argument(
- "-o", "--output",
- help="directory to write output files\n[REQUIRED]",
- metavar="<output>",
- required=True)
- common_settings.add_argument(
- "--threads",
- help="number of threads/processes\n[DEFAULT: " + str(config.threads) + "]",
- metavar="<" + str(config.threads) + ">",
- type=int,
- default=config.threads)
- common_settings.add_argument(
- "--version",
- action="version",
- version="%(prog)s v"+VERSION)
- workflow_refinement=parser.add_argument_group("[1] Workflow refinement")
- workflow_refinement.add_argument(
- "-r","--resume",
- help="bypass commands if the output files exist\n",
- action="store_true",
- default=config.resume)
- workflow_refinement.add_argument(
- "--bypass-nucleotide-index",
- help="bypass the nucleotide index step and run on the indexed ChocoPhlAn database\n",
- action="store_true",
- default=config.bypass_nucleotide_index)
- workflow_refinement.add_argument(
- "--bypass-nucleotide-search",
- help="bypass the nucleotide search steps\n",
- action="store_true",
- default=config.bypass_nucleotide_search)
- workflow_refinement.add_argument(
- "--bypass-prescreen",
- help="bypass the prescreen step and run on the full ChocoPhlAn database\n",
- action="store_true",
- default=config.bypass_prescreen)
- workflow_refinement.add_argument(
- "--bypass-translated-search",
- help="bypass the translated search step\n",
- action="store_true",
- default=config.bypass_translated_search)
- workflow_refinement.add_argument(
- "--taxonomic-profile",
- help="a taxonomic profile (the output file created by metaphlan)\n[DEFAULT: file will be created]",
- metavar="<taxonomic_profile.tsv>")
- workflow_refinement.add_argument(
- "--memory-use",
- help="the amount of memory to use\n[DEFAULT: " +
- config.memory_use + "]",
- default=config.memory_use,
- choices=config.memory_use_options)
- workflow_refinement.add_argument(
- "--input-format",
- help="the format of the input file\n[DEFAULT: format identified by software]",
- choices=config.input_format_choices)
- workflow_refinement.add_argument(
- "-v","--verbose",
- help="additional output is printed\n",
- action="store_true",
- default=config.verbose)
- tier1_prescreen=parser.add_argument_group("[2] Configure tier 1: prescreen")
- tier1_prescreen.add_argument(
- "--metaphlan",
- help="directory containing the MetaPhlAn software\n[DEFAULT: $PATH]",
- metavar="<metaphlan>")
- tier1_prescreen.add_argument(
- "--metaphlan-options",
- help="options to be provided to the MetaPhlAn software\n[DEFAULT: \"" + str(" ".join(config.metaphlan_opts)) + "\"]",
- metavar="<metaphlan_options>",
- default=" ".join(config.metaphlan_opts))
- tier1_prescreen.add_argument(
- "--prescreen-threshold",
- help="minimum estimated genome coverage for inclusion in pangenome search\n[DEFAULT: "
- + str(config.prescreen_threshold) + "]",
- metavar="<" + str(config.prescreen_threshold) + ">",
- type=float,
- default=config.prescreen_threshold)
- tier1_prescreen.add_argument(
- "--average-read-length",
- help="average read length for input file\n[DEFAULT: "
- + str(config.average_read_length) + "]",
- metavar="<" + str(config.average_read_length) + ">",
- type=float,
- default=config.average_read_length)
- tier2_nucleotide_search=parser.add_argument_group("[3] Configure tier 2: nucleotide search")
- tier2_nucleotide_search.add_argument(
- "--bowtie2",
- help="directory containing the bowtie2 executable\n[DEFAULT: $PATH]",
- metavar="<bowtie2>")
- tier2_nucleotide_search.add_argument(
- "--bowtie-options",
- help="options to be provided to the bowtie software\n[DEFAULT: \"" + str(" ".join(config.bowtie2_align_opts)) + "\"]",
- metavar="<bowtie_options>",
- default=" ".join(config.bowtie2_align_opts))
- tier2_nucleotide_search.add_argument(
- "--nucleotide-database",
- help="directory containing the nucleotide database\n[DEFAULT: "
- + config.nucleotide_database + "]",
- metavar="<nucleotide_database>")
- tier2_nucleotide_search.add_argument(
- "--nucleotide-identity-threshold",
- help="identity threshold for nucleotide alignments\n[DEFAULT: " + str(config.nucleotide_identity_threshold) + "]",
- metavar="<" + str(config.nucleotide_identity_threshold) + ">",
- default=config.nucleotide_identity_threshold,
- type=float)
- tier2_nucleotide_search.add_argument(
- "--nucleotide-query-coverage-threshold",
- help="query coverage threshold for nucleotide alignments\n[DEFAULT: "
- + str(config.nucleotide_query_coverage_threshold) + "]",
- metavar="<" + str(config.nucleotide_query_coverage_threshold) + ">",
- type=float,
- default=config.nucleotide_query_coverage_threshold)
- tier2_nucleotide_search.add_argument(
- "--nucleotide-subject-coverage-threshold",
- help="subject coverage threshold for nucleotide alignments\n[DEFAULT: "
- + str(config.nucleotide_subject_coverage_threshold) + "]",
- metavar="<" + str(config.nucleotide_subject_coverage_threshold) + ">",
- type=float,
- default=config.nucleotide_subject_coverage_threshold)
- tier3_translated_search=parser.add_argument_group("[3] Configure tier 2: translated search")
- tier3_translated_search.add_argument(
- "--diamond",
- help="directory containing the diamond executable\n[DEFAULT: $PATH]",
- metavar="<diamond>")
- tier3_translated_search.add_argument(
- "--diamond-options",
- help="options to be provided to the diamond software\n[DEFAULT: \"" + str(" ".join(config.diamond_opts_uniref50)) + "\"]",
- metavar="<diamond_options>",
- default=config.diamond_opts_uniref50)
- tier3_translated_search.add_argument(
- "--evalue",
- help="the evalue threshold to use with the translated search\n[DEFAULT: " + str(config.evalue_threshold) + "]",
- metavar="<" + str(config.evalue_threshold) + ">",
- type=float,
- default=config.evalue_threshold)
- tier3_translated_search.add_argument(
- "--protein-database",
- help="directory containing the protein database\n[DEFAULT: "
- + config.protein_database + "]",
- metavar="<protein_database>")
- tier3_translated_search.add_argument(
- "--translated-alignment",
- help="software to use for translated alignment\n[DEFAULT: " +
- config.translated_alignment_selected + "]",
- default=config.translated_alignment_selected,
- choices=config.translated_alignment_choices)
- tier3_translated_search.add_argument(
- "--translated-identity-threshold",
- help="identity threshold for translated alignments\n[DEFAULT: " + str(config.identity_threshold_uniref50_mode) + "]",
- metavar="<" + str(config.identity_threshold_uniref50_mode) + ">",
- default=config.identity_threshold_uniref50_mode,
- type=float,
- dest="identity_threshold")
- tier3_translated_search.add_argument(
- "--translated-query-coverage-threshold",
- help="query coverage threshold for translated alignments\n[DEFAULT: "
- + str(config.translated_query_coverage_threshold) + "]",
- metavar="<" + str(config.translated_query_coverage_threshold) + ">",
- type=float,
- default=config.translated_query_coverage_threshold)
- tier3_translated_search.add_argument(
- "--translated-subject-coverage-threshold",
- help="subject coverage threshold for translated alignments\n[DEFAULT: "
- + str(config.translated_subject_coverage_threshold) + "]",
- metavar="<" + str(config.translated_subject_coverage_threshold) + ">",
- type=float,
- default=config.translated_subject_coverage_threshold)
- gene_and_pathway=parser.add_argument_group("[5] Gene and pathway quantification")
- gene_and_pathway.add_argument(
- "--count-normalization",
- help="normalization mode for results from both nucleotide and translated search\n[DEFAULT: " +
- config.count_normalization + "]",
- default=config.count_normalization,
- choices=config.count_normalization_choices)
- gene_and_pathway.add_argument(
- "--utility-database",
- help="directory containing the utility database\n[DEFAULT: "
- + config.utility_mapping_database + "]",
- metavar="<utility_mapping_database>")
- gene_and_pathway.add_argument(
- "--gap-fill",
- help="turn on/off the gap fill computation\n[DEFAULT: " +
- config.gap_fill_toggle + "]",
- default=config.gap_fill_toggle,
- choices=config.toggle_choices)
- gene_and_pathway.add_argument(
- "--minpath",
- help="turn on/off the minpath computation\n[DEFAULT: " +
- config.minpath_toggle + "]",
- default=config.minpath_toggle,
- choices=config.toggle_choices)
- gene_and_pathway.add_argument(
- "--pathways",
- help="the database to use for pathway computations\n[DEFAULT: " +
- config.pathways_database + "]",
- default=config.pathways_database,
- choices=config.pathways_database_choices)
- gene_and_pathway.add_argument(
- "--pathways-database",
- help="mapping file (or files, at most two in a comma-delimited list) to use for pathway computations\n[DEFAULT: " +
- config.pathways_database + " database ]",
- metavar=("<pathways_database.tsv>"),
- nargs=1)
- gene_and_pathway.add_argument(
- "--xipe",
- help="turn on/off the xipe computation\n[DEFAULT: " +
- config.xipe_toggle + "]",
- default=config.xipe_toggle,
- choices=config.toggle_choices)
- gene_and_pathway.add_argument(
- "--annotation-gene-index",
- help="the index of the gene in the sequence annotation\n[DEFAULT: "
- + ",".join(str(i) for i in config.chocophlan_gene_indexes) + "]",
- metavar="<"+",".join(str(i) for i in config.chocophlan_gene_indexes)+">",
- default=",".join(str(i) for i in config.chocophlan_gene_indexes))
- gene_and_pathway.add_argument(
- "--id-mapping",
- help="id mapping file for alignments\n[DEFAULT: alignment reference used]",
- metavar="<id_mapping.tsv>")
- more_output_config=parser.add_argument_group("[6] More output configuration")
- more_output_config.add_argument(
- "--remove-temp-output",
- help="remove temp output files\n" +
- "[DEFAULT: temp files are not removed]",
- action="store_true")
- more_output_config.add_argument(
- "--log-level",
- help="level of messages to display in log\n" +
- "[DEFAULT: " + config.log_level + "]",
- default=config.log_level,
- choices=config.log_level_choices)
- more_output_config.add_argument(
- "--o-log",
- help="log file\n" +
- "[DEFAULT: sample_0.log]",
- metavar="<sample_0.log>")
- more_output_config.add_argument(
- "--output-basename",
- help="the basename for the output files\n[DEFAULT: " +
- "input file basename]",
- default=config.file_basename,
- metavar="<sample_name>")
- more_output_config.add_argument(
- "--output-format",
- help="the format of the output files\n[DEFAULT: " +
- config.output_format + "]",
- default=config.output_format,
- choices=config.output_format_choices)
- more_output_config.add_argument(
- "--output-max-decimals",
- help="the number of decimals to output\n[DEFAULT: " +
- str(config.output_max_decimals) + "]",
- metavar="<" + str(config.output_max_decimals) + ">",
- type=int,
- default=config.output_max_decimals)
- more_output_config.add_argument(
- "--remove-column-description-output",
- help="remove the description in the output column\n" +
- "[DEFAULT: output column includes description]",
- action="store_true",
- default=config.remove_column_description_output)
- more_output_config.add_argument(
- "--remove-stratified-output",
- help="remove stratification from output\n" +
- "[DEFAULT: output is stratified]",
- action="store_true",
- default=config.remove_stratified_output)
- return parser.parse_args()
- def update_configuration(args):
- """
- Update the configuration settings based on the arguments
- """
- # Use the full path to the input file
- args.input=os.path.abspath(args.input)
- # set the version header
- config.version_header=config.version_header+VERSION
- # If set, append paths executable locations
- if args.metaphlan:
- utilities.add_exe_to_path(os.path.abspath(args.metaphlan))
- if args.bowtie2:
- utilities.add_exe_to_path(os.path.abspath(args.bowtie2))
- if args.diamond:
- utilities.add_exe_to_path(os.path.abspath(args.diamond))
- # Set the metaphlan options, removing any extra spaces
- if args.metaphlan_options:
- config.metaphlan_opts=list(filter(None,args.metaphlan_options.split(" ")))
- # check for custom diamond options
- config.diamond_opts=args.diamond_options
- if args.diamond_options and args.diamond_options != config.diamond_opts_uniref50:
- config.diamond_options_custom = True
- config.diamond_opts=list(filter(None,args.diamond_options.split(" ")))
- if args.bowtie_options:
- config.bowtie2_align_opts=list(filter(None,args.bowtie_options.split(" ")))
- # Set the pathways database selection
- if args.pathways == "metacyc":
- config.pathways_database_part1=config.metacyc_gene_to_reactions
- config.pathways_database_part2=config.metacyc_reactions_to_pathways
- config.pathways_ec_column=True
- elif args.pathways == "unipathway":
- config.pathways_database_part1=config.unipathway_database_part1
- config.pathways_database_part2=config.unipathway_database_part2
- config.pathways_ec_column=True
- # Set the locations of the pathways databases
- # If provided by the user, this will take precedence over the pathways database selection
- if args.pathways_database:
- custom_pathways_files=args.pathways_database[0].split(",")
- if len(custom_pathways_files)==2:
- config.pathways_database_part1=os.path.abspath(custom_pathways_files[0])
- config.pathways_database_part2=os.path.abspath(custom_pathways_files[1])
- elif len(custom_pathways_files)==1:
- config.pathways_database_part1=None
- config.pathways_database_part2=os.path.abspath(custom_pathways_files[0])
- else:
- sys.exit("ERROR: Please provide one or two pathways files.")
- config.pathways_ec_column=False
- # Set the locations of the other databases
- if args.nucleotide_database:
- config.nucleotide_database=os.path.abspath(args.nucleotide_database)
- if args.protein_database:
- config.protein_database=os.path.abspath(args.protein_database)
- # if set, update the config run mode to resume
- if args.resume:
- config.resume=True
- # if set, update the config run mode to verbose
- if args.verbose:
- config.verbose=True
- # if set, update the config run mode to bypass prescreen step
- if args.bypass_prescreen:
- config.bypass_prescreen=True
- # if set, update the config run mode to bypass nucleotide index step
- if args.bypass_nucleotide_index:
- config.bypass_nucleotide_index=True
- config.bypass_prescreen=True
- # if set, update the config run mode to bypass translated search step
- # set the pick_frames toggle based on the bypass
- config.pick_frames_toggle="off"
- if args.bypass_translated_search:
- config.bypass_translated_search=True
- # if set, update the config run mode to bypass nucleotide search steps
- if args.bypass_nucleotide_search:
- config.bypass_prescreen=True
- config.bypass_nucleotide_index=True
- config.bypass_nucleotide_search=True
- # update the average read length
- config.average_read_length=args.average_read_length
- # Update thresholds
- config.prescreen_threshold=args.prescreen_threshold
- config.translated_subject_coverage_threshold=args.translated_subject_coverage_threshold
- config.nucleotide_subject_coverage_threshold=args.nucleotide_subject_coverage_threshold
- config.translated_query_coverage_threshold=args.translated_query_coverage_threshold
- config.nucleotide_query_coverage_threshold=args.nucleotide_query_coverage_threshold
- # Update the max decimals output
- config.output_max_decimals=args.output_max_decimals
- # Update memory use
- config.memory_use=args.memory_use
- # Update threads
- config.threads=args.threads
- # Update the evalue threshold
- config.evalue_threshold=args.evalue
- # Update translated alignment software
- config.translated_alignment_selected=args.translated_alignment
- # Update the computation toggle choices
- config.xipe_toggle=args.xipe
- config.minpath_toggle=args.minpath
- config.gap_fill_toggle=args.gap_fill
- config.count_normalization=args.count_normalization
- if args.utility_database:
- config.utility_mapping_database=os.path.abspath(args.utility_database)
- # Check that the input file exists and is readable
- if not os.path.isfile(args.input):
- sys.exit("CRITICAL ERROR: Can not find input file selected: "+ args.input)
- if not os.access(args.input, os.R_OK):
- sys.exit("CRITICAL ERROR: Not able to read input file selected: " + args.input)
- # Update the output format
- config.remove_stratified_output=args.remove_stratified_output
- config.remove_column_description_output=args.remove_column_description_output
- # Check that the output directory is writeable
- output_dir = os.path.abspath(args.output)
- if not os.path.isdir(output_dir):
- try:
- print("Creating output directory: " + output_dir)
- os.mkdir(output_dir)
- except EnvironmentError:
- sys.exit("CRITICAL ERROR: Unable to create output directory.")
- if not os.access(output_dir, os.W_OK):
- sys.exit("CRITICAL ERROR: The output directory is not " +
- "writeable. This software needs to write files to this directory.\n" +
- "Please select another directory.")
- print("Output files will be written to: " + output_dir)
- # Set the basename of the output files if specified as an option
- if args.output_basename:
- config.file_basename=args.output_basename
- else:
- # Determine the basename of the input file to use as output file basename
- input_file_basename=os.path.basename(args.input)
- # Remove gzip extension if present
- if re.search('.gz$',input_file_basename):
- input_file_basename='.'.join(input_file_basename.split('.')[:-1])
- # Remove input file extension if present
- if '.' in input_file_basename:
- input_file_basename='.'.join(input_file_basename.split('.')[:-1])
- config.file_basename=input_file_basename
- # Set the output format
- config.output_format=args.output_format
- # Set final output file names and location
- config.pathabundance_file=os.path.join(output_dir,
- config.file_basename + config.pathabundance_file + "." +
- config.output_format)
- config.pathcoverage_file=os.path.join(output_dir,
- config.file_basename + config.pathcoverage_file + "." +
- config.output_format)
- config.genefamilies_file=os.path.join(output_dir,
- config.file_basename + config.genefamilies_file + "." +
- config.output_format)
- config.profile_file=os.path.join(output_dir,
- config.file_basename + config.profile_file + "." +
- config.output_format)
- config.reactions_file=os.path.join(output_dir,
- config.file_basename + config.reactions_file + "." +
- config.output_format)
- # set the location of the temp directory
- if not args.remove_temp_output:
- config.temp_dir=os.path.join(output_dir,config.file_basename+"_humann_temp")
- if not os.path.isdir(config.temp_dir):
- try:
- os.mkdir(config.temp_dir)
- except EnvironmentError:
- sys.exit("Unable to create temp directory: " + config.temp_dir)
- else:
- config.temp_dir=tempfile.mkdtemp(
- prefix=config.file_basename+'_humann_temp_',dir=output_dir)
- # create the unnamed temp directory
- config.unnamed_temp_dir=tempfile.mkdtemp(dir=config.temp_dir)
- # set the name of the log file
- log_file=os.path.join(output_dir,config.file_basename+"_0.log")
- # change file name if set
- if args.o_log:
- log_file=args.o_log
- # configure the logger
- logging.basicConfig(filename=log_file,format='%(asctime)s - %(name)s - %(levelname)s: %(message)s',
- level=getattr(logging,args.log_level), filemode='w', datefmt='%m/%d/%Y %I:%M:%S %p')
- # write the version of the software to the log
- logger.info("Running humann v"+VERSION)
- # write the location of the output files to the log
- logger.info("Output files will be written to: " + output_dir)
- # write the location of the temp file directory to the log
- message="Writing temp files to directory: " + config.temp_dir
- logger.info(message)
- if config.verbose:
- print("\n"+message+"\n")
- return log_file
- def parse_chocophlan_gene_indexes(annotation_gene_index):
- """ Parse the chocophlan gene index input """
- # Update the chocophlan gene indexes
- chocophlan_gene_indexes=[]
- for index in annotation_gene_index.split(","):
- # Look for array range
- if ":" in index:
- split_index=index.split(":")
- start=split_index[0]
- end=split_index[1]
- try:
- start=int(start)
- end=int(end)
- chocophlan_gene_indexes+=range(start,end)
- except ValueError:
- pass
- else:
- # Convert to int
- try:
- index=int(index)
- chocophlan_gene_indexes.append(index)
- except ValueError:
- pass
- return chocophlan_gene_indexes
- def check_requirements(args):
- """
- Check requirements (file format, dependencies, permissions)
- """
- # Check the pathways database files exist and are readable
- if config.pathways_database_part1:
- utilities.file_exists_readable(config.pathways_database_part1)
- utilities.file_exists_readable(config.pathways_database_part2)
- # Determine the input file format if not provided
- if not args.input_format:
- args.input_format=utilities.determine_file_format(args.input)
- if args.input_format == "unknown":
- sys.exit("CRITICAL ERROR: Unable to determine the input file format." +
- " Please provide the format with the --input_format argument.")
- config.input_format = args.input_format
- # If the input file is compressed, then decompress
- if args.input_format.endswith(".gz"):
- new_file=utilities.gunzip_file(args.input)
- if new_file:
- args.input=new_file
- args.input_format=args.input_format.split(".")[0]
- else:
- sys.exit("CRITICAL ERROR: Unable to use gzipped input file. " +
- " Please check the format of the input file.")
- # check if the input file has sequence identifiers of the new illumina casava v1.8+ format
- # these have spaces causing the paired end reads to have the same identifier after
- # delimiting by space (causing bowtie2/diamond to label the reads with the same identifier)
- if args.input_format in ["fasta","fastq"]:
- if utilities.space_in_identifier(args.input):
- message="Removing spaces from identifiers in input file"
- print(message+" ...\n")
- logger.info(message)
- new_file = utilities.remove_spaces_from_file(args.input)
- if new_file:
- args.input=new_file
- else:
- sys.exit("CRITICAL ERROR: Unable to remove spaces from identifiers in input file.")
- # If the input format is in binary then convert to sam (tab-delimited text)
- if args.input_format == "bam":
- # Check for the samtools software
- if not utilities.find_exe_in_path("samtools"):
- sys.exit("CRITICAL ERROR: The samtools executable can not be found. "
- "Please check the install or select another input format.")
- new_file=utilities.bam_to_sam(args.input)
- if new_file:
- args.input=new_file
- args.input_format="sam"
- else:
- sys.exit("CRITICAL ERROR: Unable to convert bam input file to sam.")
- # If the input format is in biom then convert to tsv
- if args.input_format == "biom":
- # Check for the biom software
- if not utilities.find_exe_in_path("biom"):
- sys.exit("CRITICAL ERROR: The biom executable can not be found. "
- "Please check the install or select another input format.")
- new_file=utilities.biom_to_tsv(args.input)
- if new_file:
- args.input=new_file
- # determine the format of the file
- args.input_format=utilities.determine_file_format(args.input)
- else:
- sys.exit("CRITICAL ERROR: Unable to convert biom input file to tsv.")
- # If the biom output format is selected, check for the biom package
- if config.output_format=="biom":
- if not utilities.find_exe_in_path("biom"):
- sys.exit("CRITICAL ERROR: The biom executable can not be found. "
- "Please check the install or select another output format.")
- try:
- import biom
- except ImportError:
- sys.exit("Could not find the biom software."+
- " This software is required since the output file is a biom file.")
- if os.path.basename(config.utility_mapping_database) == "utility_DEMO":
- # Check the input file is a demo input if running with demo database
- try:
- input_file_size=os.path.getsize(args.input)/1024**2
- except EnvironmentError:
- input_file_size=0
- if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
- sys.exit("ERROR: You are using the demo utility database with "
- + "a non-demo input file. If you have not already done so, please "
- + "run humann_databases to download the full utility database. "
- + "If you have downloaded the full database, use the option "
- + "--utility-database to provide the location. "
- + "You can also run humann_config to update the default "
- + "database location. For additional information, please "
- + "see the HUMAnN User Manual.")
- # If the file is fasta/fastq check for requirements
- if args.input_format in ["fasta","fastq"]:
- # Check that the chocophlan directory exists
- if not config.bypass_nucleotide_index:
- if not os.path.isdir(config.nucleotide_database):
- if args.nucleotide_database:
- sys.exit("CRITICAL ERROR: The directory provided for the ChocoPhlAn database at "
- + args.nucleotide_database + " does not exist. Please select another directory.")
- else:
- sys.exit("CRITICAL ERROR: The default ChocoPhlAn database directory of "
- + config.nucleotide_database + " does not exist. Please provide the location "
- + "of the ChocoPhlAn directory using the --nucleotide-database option.")
- # Check that the files in the chocophlan folder are of the right format
- if not config.bypass_nucleotide_index:
- valid_format_count=0
- for file in os.listdir(config.nucleotide_database):
- # expect most of the file names to be of the format g__*s__*
- if re.search("^SGB",file):
- valid_format_count+=1
- if not config.metaphlan_v4_db_matching_uniref in file:
- sys.exit("\n\nCRITICAL ERROR: The directory provided for ChocoPhlAn contains files ( "+file+" )"+\
- " that are not of the expected version. Please install the latest version"+\
- " of the database: "+config.metaphlan_v4_db_matching_uniref)
- if valid_format_count == 0:
- sys.exit("CRITICAL ERROR: The directory provided for ChocoPhlAn does not "
- + "contain files of the expected format (ie \'^SGB\').")
- # Check if running with the demo database
- if not config.bypass_nucleotide_index:
- if os.path.basename(config.nucleotide_database) == "chocophlan_DEMO":
- # Check the input file is a demo input if running with demo database
- try:
- input_file_size=os.path.getsize(args.input)/1024**2
- except EnvironmentError:
- input_file_size=0
- if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
- sys.exit("ERROR: You are using the demo ChocoPhlAn database with "
- + "a non-demo input file. If you have not already done so, please "
- + "run humann_databases to download the full ChocoPhlAn database. "
- + "If you have downloaded the full database, use the option "
- + "--nucleotide-database to provide the location. "
- + "You can also run humann_config to update the default "
- + "database location. For additional information, please "
- + "see the HUMAnN User Manual.")
- # Check that the metaphlan2 executable can be found
- if not config.bypass_prescreen and not config.bypass_nucleotide_index:
- if not utilities.find_exe_in_path("metaphlan"):
- sys.exit("CRITICAL ERROR: The metaphlan executable can not be found. "
- "Please check the install.")
- # Check the metaphlan2 version
- # utilities.check_software_version("metaphlan",config.metaphlan_version)
- # Check that the bowtie2 executable can be found
- if not config.bypass_nucleotide_search:
- if not utilities.find_exe_in_path("bowtie2"):
- sys.exit("CRITICAL ERROR: The bowtie2 executable can not be found. "
- "Please check the install.")
- # Check the bowtie2 version
- utilities.check_software_version("bowtie2", config.bowtie2_version, warning=True)
- if not config.bypass_translated_search:
- # Check that the protein database directory exists
- if not os.path.isdir(config.protein_database):
- if args.protein_database:
- sys.exit("CRITICAL ERROR: The directory provided for the protein database at "
- + args.protein_database + " does not exist. Please select another directory.")
- else:
- sys.exit("CRITICAL ERROR: The default protein database directory of "
- + config.protein_database + " does not exist. Please provide the location "
- + "of the directory using the --protein-database option.")
- # Check that some files in the protein database folder are of the expected extension
- expected_database_extension=""
- if config.translated_alignment_selected == "usearch":
- expected_database_extension=config.usearch_database_extension
- elif config.translated_alignment_selected == "rapsearch":
- expected_database_extension=config.rapsearch_database_extension
- elif config.translated_alignment_selected == "diamond":
- expected_database_extension=config.diamond_database_extension
- valid_format_count=0
- database_files=os.listdir(config.protein_database)
- valid_format_database_files=[]
- for file in database_files:
- if not config.matching_uniref in file:
- sys.exit("\n\nCRITICAL ERROR: The directory provided for the translated database contains files ( "+file+" )"+\
- " that are not of the expected version. Please install the latest version"+\
- " of the database: "+config.matching_uniref)
- if file.endswith(expected_database_extension):
- # if rapsearch check for the second database file
- if config.translated_alignment_selected == "rapsearch":
- database_file=re.sub(config.rapsearch_database_extension+"$","",file)
- if database_file in database_files:
- valid_format_count+=1
- valid_format_database_files.append(database_file)
- else:
- valid_format_count+=1
- valid_format_database_files.append(file)
- if valid_format_count == 0:
- sys.exit("CRITICAL ERROR: The protein database directory provided ( " + config.protein_database
- + " ) does not contain any files that have been formatted to run with"
- " the translated alignment software selected ( " +
- config.translated_alignment_selected + " ). Please format these files so"
- + " they are of the expected extension ( " + expected_database_extension +" ).")
- # Check if running with the demo database
- if not config.bypass_translated_search:
- if os.path.basename(config.protein_database) == "uniref_DEMO":
- # Check the input file is a demo input if running with demo database
- try:
- input_file_size=os.path.getsize(args.input)/1024**2
- except EnvironmentError:
- input_file_size=0
- if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
- sys.exit("ERROR: You are using the demo UniRef database with "
- + "a non-demo input file. If you have not already done so, please "
- + "run humann_databases to download the full UniRef database. "
- + "If you have downloaded the full database, use the option "
- + "--protein-database to provide the location. "
- + "You can also run humann_config to update the default "
- + "database location. For additional information, please "
- + "see the HUMAnN User Manual.")
- # Check that the translated alignment executable can be found
- if not utilities.find_exe_in_path(config.translated_alignment_selected):
- sys.exit("CRITICAL ERROR: The " + config.translated_alignment_selected +
- " executable can not be found. Please check the install.")
- # Check for correct usearch version
- if config.translated_alignment_selected == "usearch":
- utilities.check_software_version("usearch",config.usearch_version)
- # Check for correct rapsearch version
- if config.translated_alignment_selected == "rapsearch":
- utilities.check_software_version("rapsearch", config.rapsearch_version)
- # Check for the correct diamond version
- if config.translated_alignment_selected == "diamond":
- utilities.check_software_version("diamond", config.diamond_version)
- # parse the chocophlan gene index
- chocophlan_gene_indexes=parse_chocophlan_gene_indexes(args.annotation_gene_index)
- config.chocophlan_gene_indexes = chocophlan_gene_indexes
- # set the identity thresholds if provided by the user
- config.nucleotide_identity_threshold=args.nucleotide_identity_threshold
- config.identity_threshold=args.identity_threshold
- def timestamp_message(task, start_time):
- """
- Print and log a message about the task completed and the time
- Log messages are tab delimited for quick task/time access with awk
- Return the new start time
- """
- message="TIMESTAMP: Completed \t" + task + " \t:\t " + \
- str(int(round(time.time() - start_time))) + "\t seconds"
- logger.info(message)
- if config.verbose:
- print("\n"+message.replace("\t","")+"\n")
- return time.time()
- def main():
- # Parse arguments from command line
- args=parse_arguments(sys.argv)
- # Update the configuration settings based on the arguments
- log_file=update_configuration(args)
- # Check for required files, software, databases, and also permissions
- check_requirements(args)
- # Write the config settings to the log file
- config.log_settings()
- # Initialize alignments and gene scores
- minimize_memory_use=True
- if config.memory_use == "maximum":
- minimize_memory_use=False
- alignments=store.Alignments(minimize_memory_use=minimize_memory_use)
- unaligned_reads_store=store.Reads(minimize_memory_use=minimize_memory_use)
- gene_scores=store.GeneScores()
- # If id mapping is provided then process
- if args.id_mapping:
- alignments.process_id_mapping(args.id_mapping)
- # Load in the reactions database
- reactions_database=None
- if config.pathways_database_part1:
- reactions_database=store.ReactionsDatabase(config.pathways_database_part1)
- message="Load pathways database part 1: " + config.pathways_database_part1
- logger.info(message)
- # Load in the pathways database
- pathways_database=store.PathwaysDatabase(config.pathways_database_part2, reactions_database)
- if config.pathways_database_part1:
- message="Load pathways database part 2: " + config.pathways_database_part2
- else:
- message="Load pathways database: " + config.pathways_database_part2
- logger.info(message)
- # Start timer
- start_time=time.time()
- # Process fasta or fastq input files
- output_files=[]
- if args.input_format in ["fasta","fastq"]:
- # Run prescreen to identify bugs
- bug_file = "Empty"
- if args.taxonomic_profile:
- bug_file = os.path.abspath(args.taxonomic_profile)
- else:
- if not config.bypass_prescreen:
- bug_file = prescreen.alignment(args.input)
- output_files.append(bug_file)
- start_time=timestamp_message("prescreen",start_time)
- # Create the custom database from the bugs list
- custom_database = ""
- if not config.bypass_nucleotide_index:
- custom_database = prescreen.create_custom_database(config.nucleotide_database, bug_file)
- start_time=timestamp_message("custom database creation",start_time)
- else:
- custom_database = "Bypass"
- # Run nucleotide search on custom database
- if custom_database != "Empty" and not config.bypass_nucleotide_search:
- if not config.bypass_nucleotide_index:
- nucleotide_index_file = nucleotide.index(custom_database)
- start_time=timestamp_message("database index",start_time)
- else:
- nucleotide_index_file = nucleotide.find_index(config.nucleotide_database)
- nucleotide_alignment_file = nucleotide.alignment(args.input,
- nucleotide_index_file)
- start_time=timestamp_message("nucleotide alignment",start_time)
- # Determine which reads are unaligned and reduce aligned reads file
- # Remove the alignment_file as we only need the reduced aligned reads file
- [ unaligned_reads_file_fasta, reduced_aligned_reads_file ] = nucleotide.unaligned_reads(
- nucleotide_alignment_file, alignments, unaligned_reads_store, keep_sam=True)
- start_time=timestamp_message("nucleotide alignment post-processing",start_time)
- # Print out total alignments per bug
- message="Total bugs from nucleotide alignment: " + str(alignments.count_bugs())
- logger.info(message)
- print(message)
- message=alignments.counts_by_bug()
- logger.info("\n"+message)
- print(message)
- message="Total gene families from nucleotide alignment: " + str(alignments.count_genes())
- logger.info(message)
- print("\n"+message)
- # Report reads unaligned
- message="Unaligned reads after nucleotide alignment: " + utilities.estimate_unaligned_reads_stored(
- args.input, unaligned_reads_store) + " %"
- logger.info(message)
- print("\n"+message+"\n")
- else:
- logger.debug("Custom database is empty")
- reduced_aligned_reads_file = "Empty"
- unaligned_reads_file_fasta=args.input
- unaligned_reads_store=store.Reads(unaligned_reads_file_fasta, minimize_memory_use=minimize_memory_use)
- # Do not run if set to bypass translated search in config file
- if not config.bypass_translated_search:
- # Run translated search on UniRef database if unaligned reads exit
- if unaligned_reads_store.count_reads()>0:
- translated_alignment_file = translated.alignment(config.protein_database,
- unaligned_reads_file_fasta)
- start_time=timestamp_message("translated alignment",start_time)
- # Determine which reads are unaligned
- translated_unaligned_reads_file_fastq = translated.unaligned_reads(
- unaligned_reads_store, translated_alignment_file, alignments)
- start_time=timestamp_message("translated alignment post-processing",start_time)
- # Print out total alignments per bug
- message="Total bugs after translated alignment: " + str(alignments.count_bugs())
- logger.info(message)
- print(message)
- message=alignments.counts_by_bug()
- logger.info("\n"+message)
- print(message)
- message="Total gene families after translated alignment: " + str(alignments.count_genes())
- logger.info(message)
- print("\n"+message)
- # Report reads unaligned
- message="Unaligned reads after translated alignment: " + utilities.estimate_unaligned_reads_stored(
- args.input, unaligned_reads_store) + " %"
- logger.info(message)
- print("\n"+message+"\n")
- else:
- message="All reads are aligned so translated alignment will not be run"
- logger.info(message)
- print(message)
- else:
- message="Bypass translated search"
- logger.info(message)
- print(message)
- # Process input files of sam format
- elif args.input_format in ["sam"]:
- # Store the sam mapping results
- message="Process the sam mapping results ..."
- logger.info(message)
- print("\n"+message)
- [unaligned_reads_file_fasta, reduced_aligned_reads_file] = nucleotide.unaligned_reads(
- args.input, alignments, unaligned_reads_store, keep_sam=True)
- start_time=timestamp_message("alignment post-processing",start_time)
- # Process input files of tab-delimited blast format
- elif args.input_format in ["blastm8"]:
- # Store the blastm8 mapping results
- message="Process the blastm8 mapping results ..."
- logger.info(message)
- print("\n"+message)
- translated_unaligned_reads_file_fastq = translated.unaligned_reads(
- unaligned_reads_store, args.input, alignments)
- start_time=timestamp_message("alignment post-processing",start_time)
- # Get the number of remaining unaligned reads
- unaligned_reads_count=unaligned_reads_store.count_reads()
- # Clear all of the unaligned reads as they are no longer needed
- unaligned_reads_store.clear()
- # Compute or load in gene families
- if args.input_format in ["fasta","fastq","sam","blastm8"]:
- # Compute the gene families
- message="Computing gene families ..."
- logger.info(message)
- print("\n"+message)
- families_file=families.gene_families(alignments,gene_scores,unaligned_reads_count)
- output_files.append(families_file)
- start_time=timestamp_message("computing gene families",start_time)
- elif args.input_format in ["genetable"]:
- # Load the gene scores
- message="Process the gene table ..."
- logger.info(message)
- print("\n"+message)
- unaligned_reads_count=gene_scores.add_from_file(args.input,id_mapping_file=args.id_mapping)
- start_time=timestamp_message("processing gene table",start_time)
- # Handle input files of unknown formats
- else:
- sys.exit("CRITICAL ERROR: Input file of unknown format.")
- # Clear all of the alignments data as they are no longer needed
- alignments.clear()
- # Identify reactions and then pathways from the alignments
- message="Computing reaction and pathway abundance ..."
- logger.info(message)
- print("\n"+message)
- pathways_and_reactions_store=modules.identify_reactions_and_pathways(
- gene_scores, reactions_database, pathways_database, unaligned_reads_count)
- # Compute pathway abundance and coverage
- abundance_file, coverage_file, reaction_file=modules.compute_pathways_abundance_and_coverage(
- gene_scores, reactions_database, pathways_and_reactions_store, pathways_database, unaligned_reads_count)
- output_files.append(reaction_file)
- output_files.append(abundance_file)
- output_files.append(log_file)
- #output_files.append(coverage_file)
- start_time=timestamp_message("computing pathways",start_time)
- message="\nOutput files created: \n" + "\n".join(output_files) + "\n"
- logger.info(message)
- print(message)
- # Remove the unnamed temp files
- utilities.remove_directory(config.unnamed_temp_dir)
- # Remove named temp directory
- if args.remove_temp_output:
- utilities.remove_directory(config.temp_dir)
humann.py at commit e07b3a3, under MIT · at the source
Overview
- Doctorado en Ciencias Biológicas y de la Salud, Universidad Autónoma Metropolitana, Ciudad de Mexico 04960, Mexico
- Departamento de Salud Mental, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
- Departamento de Ciencias Ambientales, Universidad Autónoma Metropolitana-Lerma, Lerma 52006, Estado de Mexico, Mexico
- Grupo Neurológico, Neuroquirúrgico y de Columna, Hospital Ángeles Acoxpa, Ciudad de Mexico 14308, Mexico
- Laboratorio de Errores Innatos del Metabolismo y Tamiz, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
- Departamento de Higiene y Tecnología de los Alimentos, Facultad de Veterinaria, Universidad de León, 24071 Leon, Spain
- Laboratorio de Oncología Experimental, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
Abstract
Background/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.
biobakery/humann
e07b3a34d0b94c09a8ac5d28ff95009611178be2, 10 July 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
73 files
- humann/
__init__.py — Python, 1 line - humann/
check.py — Python, 54 lines - humann/
config.py — Python, 535 lines - humann/
humann.py — Python, 1,151 lines, 1 match - humann/
quantify/ — Python, 947 linesMinPath12hmp.py - humann/
quantify/ — Python, 1 line__init__.py - humann/
quantify/ — Python, 92 lineschi2cdf.py - humann/
quantify/ — Python, 90 linesfamilies.py - humann/
quantify/ — Python, 658 linesmodules.py - humann/
quantify/ — Python, 518 linesxipe.py - humann/
search/ — Python, 1 line__init__.py - humann/
search/ — Python, 128 linesblastx_coverage.py - humann/
search/ — Python, 404 linesnucleotide.py - humann/
search/ — Python, 134 linespick_frames.py - humann/
search/ — Python, 387 lines, 1 matchprescreen.py - humann/
search/ — Python, 323 linestranslated.py - humann/
store.py — Python, 1,507 lines - humann/
tests/ — Python, 1 line__init__.py - humann/
tests/ — Python, 206 linesadvanced_tests_blastx_co verage.py - humann/
tests/ — Python, 475 linesadvanced_tests_nucleotid e_search.py - humann/
tests/ — Python, 193 linesadvanced_tests_quantify_ families.py - humann/
tests/ — Python, 807 linesadvanced_tests_quantify_ modules.py - humann/
tests/ — Python, 562 linesadvanced_tests_store.py - humann/
tests/ — Python, 326 linesadvanced_tests_translate d_search.py - humann/
tests/ — Python, 44 linesadvanced_tests_utilities .py - humann/
tests/ — Python, 196 linesbasic_tests_nucleotide_s earch.py - humann/
tests/ — Python, 187 linesbasic_tests_quantify_mod ules.py - humann/
tests/ — Python, 1,143 linesbasic_tests_store.py - humann/
tests/ — Python, 464 linesbasic_tests_utilities.py - humann/
tests/ — Python, 204 linescfg.py - humann/
tests/ — Shell, 4 linesdata/ tooltest-infer_taxonomy/ infer_taxonomy-command.s h - humann/
tests/ — Shell, 5 linesdata/ tooltest-merge_abundance _tables/ merge_abundance-command. sh - humann/
tests/ — Shell, 4 linesdata/ tooltest-regroup_table/ regroup_table-builtin_co mmand.sh - humann/
tests/ — Shell, 4 linesdata/ tooltest-regroup_table/ regroup_table-custom_com mand.sh - humann/
tests/ — Shell, 4 linesdata/ tooltest-rename_table/ rename_table-builtin_com mand.sh - humann/
tests/ — Shell, 4 linesdata/ tooltest-rename_table/ rename_table-custom.sh - humann/
tests/ — Shell, 4 linesdata/ tooltest-rename_table/ rename_table-custom_comm and.sh - humann/
tests/ — Shell, 5 linesdata/ tooltest-renorm_table/ renorm_table-command.sh - humann/
tests/ — Shell, 6 linesdata/ tooltest-rna_dna_norm/ rna_dna_norm-command.sh - humann/
tests/ — Shell, 5 linesdata/ tooltest-strain_profiler / strain_profiler-command. sh - humann/
tests/ — Python, 79 linesfunctional_tests_biom_hu mann.py - humann/
tests/ — Python, 91 linesfunctional_tests_biom_to ols.py - humann/
tests/ — Python, 250 linesfunctional_tests_humann. py - humann/
tests/ — Python, 259 linesfunctional_tests_tools.p y - humann/
tests/ — Python, 135 lineshumann_test.py - humann/
tests/ — Python, 150 linesutils.py - humann/
tools/ — Python, 1 line__init__.py - humann/
tools/ — Python, 215 linesbuild_custom_database.py - humann/
tools/ — Python, 140 linesexpand_cluster.py - humann/
tools/ — Python, 92 linesgenefamilies_genus_level .py - humann/
tools/ — Python, 200 lineshumann_associate.py - humann/
tools/ — Python, 733 lineshumann_barplot.py - humann/
tools/ — Python, 101 lineshumann_benchmark.py - humann/
tools/ — Python, 84 lineshumann_config.py - humann/
tools/ — Python, 184 lineshumann_databases.py - humann/
tools/ — Python, 143 lineshumann_expand_results_ta xonomy.py - humann/
tools/ — Python, 262 lineshumann_humann1_kegg.py - humann/
tools/ — Python, 328 linesinfer_taxonomy.py - humann/
tools/ — Python, 251 linesjoin_tables.py - humann/
tools/ — Python, 319 linesmerge_abundance.py - humann/
tools/ — Python, 155 linesreduce_table.py - humann/
tools/ — Python, 238 linesregroup_table.py - humann/
tools/ — Python, 147 linesrename_table.py - humann/
tools/ — Python, 130 linesrenorm_table.py - humann/
tools/ — Python, 211 linesrna_dna_norm.py - humann/
tools/ — Python, 148 linessplit_stratified_table.p y - humann/
tools/ — Python, 349 linessplit_table.py - humann/
tools/ — Python, 156 linesstrain_profiler.py - humann/
tools/ — Python, 488 linesutil.py - humann/
utilities.py — Python, 1,444 lines - setup.py — Python, 727 lines
- LICENSE — License, 22 lines
- readme.md — Text, 1,516 lines
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 71 scripts, each with its path and the digest of its content;
- 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- bioproject:PRJNA1425957 — at NCBI BioProject; found in “Data Availability Statement”
Data Availability Statement
All 16S rDNA and Whole Metagenomic Sequence files and corresponding mapping files for the samples utilized in this study have been deposited in the NCBI BioSample repository. Interested parties may access these resources using the provided Accession Number PRJNA1425957, which can be found at the following link: https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 7 keywords, 16 MeSH terms, 1 funder, 90 references.
Cite
This paper
De Sales-Millan, A., Reyes-Ferreira, P., González-Cervantes, R. M., Luna-Álvarez, M., Guillén-López, S., Cobo-Díaz, J. F., Ramos, S., Aguirre-Garrido, J. F., & Velázquez-Aragón, J. A. (2026). Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder. Nutrients, 18(15), 2441. https://
BibTeX
@article{desalesmillan20
author = {De Sales-Millan, Amapola and Reyes-Ferreira, Paulina and González-Cervantes, Rina María and Luna-Álvarez, Mariana and Guillén-López, Sara and Cobo-Díaz, José F. and Ramos, Sandra and Aguirre-Garrido, José Félix and Velázquez-Aragón, José Antonio},
title = {{Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder}},
journal = {Nutrients},
year = {2026},
month = jul,
volume = {18},
number = {15},
pages = {2441},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {2072-6643},
doi = {10.3390/
url = {https://
pmid = {42588064},
pmcid = {PMC13468271}
}
RIS
TY - JOUR
AU - De Sales-Millan, Amapola
AU - Reyes-Ferreira, Paulina
AU - González-Cervantes, Rina María
AU - Luna-Álvarez, Mariana
AU - Guillén-López, Sara
AU - Cobo-Díaz, José F.
AU - Ramos, Sandra
AU - Aguirre-Garrido, José Félix
AU - Velázquez-Aragón, José Antonio
TI - Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder
T2 - Nutrients
J2 - Nutrients
PY - 2026
DA - 2026/
VL - 18
IS - 15
SP - 2441
SN - 2072-6643
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.3390/
"type": "article-journal",
"title": "Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder",
"container-title": "Nutrients",
"author": [
{
"family": "De Sales-Millan",
"given": "Amapola"
},
{
"family": "Reyes-Ferreira",
"given": "Paulina"
},
{
"family": "González-Cervantes",
"given": "Rina María"
},
{
"family": "Luna-Álvarez",
"given": "Mariana"
},
{
"family": "Guillén-López",
"given": "Sara"
},
{
"family": "Cobo-Díaz",
"given": "José F."
},
{
"family": "Ramos",
"given": "Sandra"
},
{
"family": "Aguirre-Garrido",
"given": "José Félix"
},
{
"family": "Velázquez-Aragón",
"given": "José Antonio"
}
],
"container-title-short":
"volume": "18",
"issue": "15",
"page": "2441",
"DOI": "10.3390/
"PMID": "42588064",
"PMCID": "PMC13468271",
"ISSN": "2072-6643",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
26
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1186/s12866-026-05296-x
- Gut bacterial characteristics in children with autism spectrum disorder according to symptom severity: a cross-sectional study.Journal: BMC microbiologyIn common: autism, 4 references
- [2] doi:10.1186/s12888-026-08178-8 [code]
- Machine learning model to identify gut microbiome-derived metabolites as potential biomarkers of autism spectrum disorder: a pilot study.Journal: BMC psychiatryIn common: Matplotlib, NumPy, autism, clinical / translational, 2 references
- [3] doi:10.1016/j.bbih.2026.101275 [code]
- A protocol for the Teen Bugs study: An integrative, multi-omics approach to understanding the role of the gut microbiome and mesocorticolimbic system in adolescent mental health following early adverse caregiving.Journal: Brain, behavior, & immunity - healthIn common: 4 references
- [4] doi:10.1016/j.gutmic.2026.100006 [code]
- From microbes to milestones: Gut bacterial abundances and functional pathways associate with neurodevelopment following preterm birth.Journal: Gut microbiologyIn common: autism, 3 references
- [5] doi:10.1002/hbm.70496 [code]
- Transdiagnostic Profiles of BOLD Signal Variability in Autism and Schizophrenia Spectrum Disorders: Associations With Cognition and Functioning.Journal: Human brain mappingIn common: SciPy, Matplotlib, NumPy, autism, 1 reference
- [6] doi:10.3389/fnins.2026.1858005 [code]
- Targeted stool metabolomics suggests exploratory catecholamine- and tryptophan-linked metabolic features in autism spectrum disorder.Journal: Frontiers in neuroscienceIn common: autism, clinical / translational, 2 references
- [7] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: h5py, SciPy, Matplotlib, 1 other tool, autism
- [8] doi:10.1038/s41467-026-74320-5 [code]
- Spatial architecture of autism pathogenesis reveals mosaic structural disarray during early development.Journal: Nature communicationsIn common: h5py, SciPy, Matplotlib, 1 other tool, autism
- [9] doi:10.1002/advs.202519479 [code]
- Diminished Signal-to-Noise Ratio Disrupts Somatosensory Population Encoding and Drives Tactile Hyposensitivity in the Fmr1&
lt;sup& gt;-/ y& lt;/ sup& gt; Autism Model. Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: h5py, SciPy, Matplotlib, 1 other tool, autism - [10] doi:10.1016/j.celrep.2026.117590 [code]
- Impaired behavioral inhibition in Fmr1 KO mice is linked to disrupted visual cortex theta oscillations.Journal: Cell reportsIn common: SciPy, Matplotlib, NumPy, autism, 1 reference
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 71 scripts, and 2 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:98005fc3634aa3df…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
