OSCR

Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § 2. Materials and Methods › 2.6. Bioinformatic and Statistical Analysis ↔ humann/humann.py, lines 3–36 · score 0.66 · Unified Metabolic, HUMAnN, pipeline, Network, diamond, Metagenomic
  2. [2] § 2. Materials and Methods › 2.6. Bioinformatic and Statistical Analysis ↔ humann/search/prescreen.py, lines 245–386 · score 0.65 · MetaPhlAn, taxonomic profiles, mpa, gene families, chocophlansgb, v4

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 1,151 lines · 50 KB · MIT · 1 match

  1. #!/usr/bin/env python
  2. """
  3. HUMAnN : HMP Unified Metabolic Analysis Network 3
  4. HUMAnN is a pipeline for efficiently and accurately determining
  5. the coverage and abundance of microbial pathways in a community
  6. from metagenomic data. Sequencing a metagenome typically produces millions
  7. of short DNA/RNA reads.
  8. Dependencies: MetaPhlAn2, ChocoPhlAn, Bowtie2, and ( Diamond or Rapsearch2 or Usearch )
  9. To Run: humann -i <input.fastq> -o <output_dir>
  10. Copyright (c) 2014 Harvard School of Public Health
  11. Permission is hereby granted, free of charge, to any person obtaining a copy
  12. of this software and associated documentation files (the "Software"), to deal
  13. in the Software without restriction, including without limitation the rights
  14. to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
  15. copies of the Software, and to permit persons to whom the Software is
  16. furnished to do so, subject to the following conditions:
  17. The above copyright notice and this permission notice shall be included in
  18. all copies or substantial portions of the Software.
  19. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
  20. IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
  21. FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
  22. AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
  23. LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
  24. OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
  25. THE SOFTWARE.
  26. """
  27. import sys
  28. # Try to load one of the humann modules to check the installation
  29. try:
  30. from . import check
  31. except ImportError:
  32. sys.exit("CRITICAL ERROR: Unable to find the HUMAnN python package." +
  33. " Please check your install.")
  34. # Check the python version
  35. check.python_version()
  36. import argparse
  37. import subprocess
  38. import os
  39. import time
  40. import tempfile
  41. import re
  42. import logging
  43. from . import config
  44. from . import store
  45. from . import utilities
  46. from .search import prescreen
  47. from .search import nucleotide
  48. from .search import translated
  49. from .quantify import families
  50. from .quantify import modules
  51. # name global logging instance
  52. logger=logging.getLogger(__name__)
  53. VERSION="4.0.0.alpha.2"
  54. MAX_SIZE_DEMO_INPUT_FILE=35
  55. def parse_arguments(args):
  56. """
  57. Parse the arguments from the user
  58. """
  59. parser = argparse.ArgumentParser(
  60. description= "HUMAnN : HMP Unified Metabolic Analysis Network 3\n",
  61. formatter_class=argparse.RawTextHelpFormatter,
  62. prog="humann")
  63. common_settings=parser.add_argument_group("[0] Common settings")
  64. common_settings.add_argument(
  65. "-i", "--input",
  66. help="input file of type {" +",".join(config.input_format_choices)+ "} \n[REQUIRED]",
  67. metavar="<input.fastq>",
  68. required=True)
  69. common_settings.add_argument(
  70. "-o", "--output",
  71. help="directory to write output files\n[REQUIRED]",
  72. metavar="<output>",
  73. required=True)
  74. common_settings.add_argument(
  75. "--threads",
  76. help="number of threads/processes\n[DEFAULT: " + str(config.threads) + "]",
  77. metavar="<" + str(config.threads) + ">",
  78. type=int,
  79. default=config.threads)
  80. common_settings.add_argument(
  81. "--version",
  82. action="version",
  83. version="%(prog)s v"+VERSION)
  84. workflow_refinement=parser.add_argument_group("[1] Workflow refinement")
  85. workflow_refinement.add_argument(
  86. "-r","--resume",
  87. help="bypass commands if the output files exist\n",
  88. action="store_true",
  89. default=config.resume)
  90. workflow_refinement.add_argument(
  91. "--bypass-nucleotide-index",
  92. help="bypass the nucleotide index step and run on the indexed ChocoPhlAn database\n",
  93. action="store_true",
  94. default=config.bypass_nucleotide_index)
  95. workflow_refinement.add_argument(
  96. "--bypass-nucleotide-search",
  97. help="bypass the nucleotide search steps\n",
  98. action="store_true",
  99. default=config.bypass_nucleotide_search)
  100. workflow_refinement.add_argument(
  101. "--bypass-prescreen",
  102. help="bypass the prescreen step and run on the full ChocoPhlAn database\n",
  103. action="store_true",
  104. default=config.bypass_prescreen)
  105. workflow_refinement.add_argument(
  106. "--bypass-translated-search",
  107. help="bypass the translated search step\n",
  108. action="store_true",
  109. default=config.bypass_translated_search)
  110. workflow_refinement.add_argument(
  111. "--taxonomic-profile",
  112. help="a taxonomic profile (the output file created by metaphlan)\n[DEFAULT: file will be created]",
  113. metavar="<taxonomic_profile.tsv>")
  114. workflow_refinement.add_argument(
  115. "--memory-use",
  116. help="the amount of memory to use\n[DEFAULT: " +
  117. config.memory_use + "]",
  118. default=config.memory_use,
  119. choices=config.memory_use_options)
  120. workflow_refinement.add_argument(
  121. "--input-format",
  122. help="the format of the input file\n[DEFAULT: format identified by software]",
  123. choices=config.input_format_choices)
  124. workflow_refinement.add_argument(
  125. "-v","--verbose",
  126. help="additional output is printed\n",
  127. action="store_true",
  128. default=config.verbose)
  129. tier1_prescreen=parser.add_argument_group("[2] Configure tier 1: prescreen")
  130. tier1_prescreen.add_argument(
  131. "--metaphlan",
  132. help="directory containing the MetaPhlAn software\n[DEFAULT: $PATH]",
  133. metavar="<metaphlan>")
  134. tier1_prescreen.add_argument(
  135. "--metaphlan-options",
  136. help="options to be provided to the MetaPhlAn software\n[DEFAULT: \"" + str(" ".join(config.metaphlan_opts)) + "\"]",
  137. metavar="<metaphlan_options>",
  138. default=" ".join(config.metaphlan_opts))
  139. tier1_prescreen.add_argument(
  140. "--prescreen-threshold",
  141. help="minimum estimated genome coverage for inclusion in pangenome search\n[DEFAULT: "
  142. + str(config.prescreen_threshold) + "]",
  143. metavar="<" + str(config.prescreen_threshold) + ">",
  144. type=float,
  145. default=config.prescreen_threshold)
  146. tier1_prescreen.add_argument(
  147. "--average-read-length",
  148. help="average read length for input file\n[DEFAULT: "
  149. + str(config.average_read_length) + "]",
  150. metavar="<" + str(config.average_read_length) + ">",
  151. type=float,
  152. default=config.average_read_length)
  153. tier2_nucleotide_search=parser.add_argument_group("[3] Configure tier 2: nucleotide search")
  154. tier2_nucleotide_search.add_argument(
  155. "--bowtie2",
  156. help="directory containing the bowtie2 executable\n[DEFAULT: $PATH]",
  157. metavar="<bowtie2>")
  158. tier2_nucleotide_search.add_argument(
  159. "--bowtie-options",
  160. help="options to be provided to the bowtie software\n[DEFAULT: \"" + str(" ".join(config.bowtie2_align_opts)) + "\"]",
  161. metavar="<bowtie_options>",
  162. default=" ".join(config.bowtie2_align_opts))
  163. tier2_nucleotide_search.add_argument(
  164. "--nucleotide-database",
  165. help="directory containing the nucleotide database\n[DEFAULT: "
  166. + config.nucleotide_database + "]",
  167. metavar="<nucleotide_database>")
  168. tier2_nucleotide_search.add_argument(
  169. "--nucleotide-identity-threshold",
  170. help="identity threshold for nucleotide alignments\n[DEFAULT: " + str(config.nucleotide_identity_threshold) + "]",
  171. metavar="<" + str(config.nucleotide_identity_threshold) + ">",
  172. default=config.nucleotide_identity_threshold,
  173. type=float)
  174. tier2_nucleotide_search.add_argument(
  175. "--nucleotide-query-coverage-threshold",
  176. help="query coverage threshold for nucleotide alignments\n[DEFAULT: "
  177. + str(config.nucleotide_query_coverage_threshold) + "]",
  178. metavar="<" + str(config.nucleotide_query_coverage_threshold) + ">",
  179. type=float,
  180. default=config.nucleotide_query_coverage_threshold)
  181. tier2_nucleotide_search.add_argument(
  182. "--nucleotide-subject-coverage-threshold",
  183. help="subject coverage threshold for nucleotide alignments\n[DEFAULT: "
  184. + str(config.nucleotide_subject_coverage_threshold) + "]",
  185. metavar="<" + str(config.nucleotide_subject_coverage_threshold) + ">",
  186. type=float,
  187. default=config.nucleotide_subject_coverage_threshold)
  188. tier3_translated_search=parser.add_argument_group("[3] Configure tier 2: translated search")
  189. tier3_translated_search.add_argument(
  190. "--diamond",
  191. help="directory containing the diamond executable\n[DEFAULT: $PATH]",
  192. metavar="<diamond>")
  193. tier3_translated_search.add_argument(
  194. "--diamond-options",
  195. help="options to be provided to the diamond software\n[DEFAULT: \"" + str(" ".join(config.diamond_opts_uniref50)) + "\"]",
  196. metavar="<diamond_options>",
  197. default=config.diamond_opts_uniref50)
  198. tier3_translated_search.add_argument(
  199. "--evalue",
  200. help="the evalue threshold to use with the translated search\n[DEFAULT: " + str(config.evalue_threshold) + "]",
  201. metavar="<" + str(config.evalue_threshold) + ">",
  202. type=float,
  203. default=config.evalue_threshold)
  204. tier3_translated_search.add_argument(
  205. "--protein-database",
  206. help="directory containing the protein database\n[DEFAULT: "
  207. + config.protein_database + "]",
  208. metavar="<protein_database>")
  209. tier3_translated_search.add_argument(
  210. "--translated-alignment",
  211. help="software to use for translated alignment\n[DEFAULT: " +
  212. config.translated_alignment_selected + "]",
  213. default=config.translated_alignment_selected,
  214. choices=config.translated_alignment_choices)
  215. tier3_translated_search.add_argument(
  216. "--translated-identity-threshold",
  217. help="identity threshold for translated alignments\n[DEFAULT: " + str(config.identity_threshold_uniref50_mode) + "]",
  218. metavar="<" + str(config.identity_threshold_uniref50_mode) + ">",
  219. default=config.identity_threshold_uniref50_mode,
  220. type=float,
  221. dest="identity_threshold")
  222. tier3_translated_search.add_argument(
  223. "--translated-query-coverage-threshold",
  224. help="query coverage threshold for translated alignments\n[DEFAULT: "
  225. + str(config.translated_query_coverage_threshold) + "]",
  226. metavar="<" + str(config.translated_query_coverage_threshold) + ">",
  227. type=float,
  228. default=config.translated_query_coverage_threshold)
  229. tier3_translated_search.add_argument(
  230. "--translated-subject-coverage-threshold",
  231. help="subject coverage threshold for translated alignments\n[DEFAULT: "
  232. + str(config.translated_subject_coverage_threshold) + "]",
  233. metavar="<" + str(config.translated_subject_coverage_threshold) + ">",
  234. type=float,
  235. default=config.translated_subject_coverage_threshold)
  236. gene_and_pathway=parser.add_argument_group("[5] Gene and pathway quantification")
  237. gene_and_pathway.add_argument(
  238. "--count-normalization",
  239. help="normalization mode for results from both nucleotide and translated search\n[DEFAULT: " +
  240. config.count_normalization + "]",
  241. default=config.count_normalization,
  242. choices=config.count_normalization_choices)
  243. gene_and_pathway.add_argument(
  244. "--utility-database",
  245. help="directory containing the utility database\n[DEFAULT: "
  246. + config.utility_mapping_database + "]",
  247. metavar="<utility_mapping_database>")
  248. gene_and_pathway.add_argument(
  249. "--gap-fill",
  250. help="turn on/off the gap fill computation\n[DEFAULT: " +
  251. config.gap_fill_toggle + "]",
  252. default=config.gap_fill_toggle,
  253. choices=config.toggle_choices)
  254. gene_and_pathway.add_argument(
  255. "--minpath",
  256. help="turn on/off the minpath computation\n[DEFAULT: " +
  257. config.minpath_toggle + "]",
  258. default=config.minpath_toggle,
  259. choices=config.toggle_choices)
  260. gene_and_pathway.add_argument(
  261. "--pathways",
  262. help="the database to use for pathway computations\n[DEFAULT: " +
  263. config.pathways_database + "]",
  264. default=config.pathways_database,
  265. choices=config.pathways_database_choices)
  266. gene_and_pathway.add_argument(
  267. "--pathways-database",
  268. help="mapping file (or files, at most two in a comma-delimited list) to use for pathway computations\n[DEFAULT: " +
  269. config.pathways_database + " database ]",
  270. metavar=("<pathways_database.tsv>"),
  271. nargs=1)
  272. gene_and_pathway.add_argument(
  273. "--xipe",
  274. help="turn on/off the xipe computation\n[DEFAULT: " +
  275. config.xipe_toggle + "]",
  276. default=config.xipe_toggle,
  277. choices=config.toggle_choices)
  278. gene_and_pathway.add_argument(
  279. "--annotation-gene-index",
  280. help="the index of the gene in the sequence annotation\n[DEFAULT: "
  281. + ",".join(str(i) for i in config.chocophlan_gene_indexes) + "]",
  282. metavar="<"+",".join(str(i) for i in config.chocophlan_gene_indexes)+">",
  283. default=",".join(str(i) for i in config.chocophlan_gene_indexes))
  284. gene_and_pathway.add_argument(
  285. "--id-mapping",
  286. help="id mapping file for alignments\n[DEFAULT: alignment reference used]",
  287. metavar="<id_mapping.tsv>")
  288. more_output_config=parser.add_argument_group("[6] More output configuration")
  289. more_output_config.add_argument(
  290. "--remove-temp-output",
  291. help="remove temp output files\n" +
  292. "[DEFAULT: temp files are not removed]",
  293. action="store_true")
  294. more_output_config.add_argument(
  295. "--log-level",
  296. help="level of messages to display in log\n" +
  297. "[DEFAULT: " + config.log_level + "]",
  298. default=config.log_level,
  299. choices=config.log_level_choices)
  300. more_output_config.add_argument(
  301. "--o-log",
  302. help="log file\n" +
  303. "[DEFAULT: sample_0.log]",
  304. metavar="<sample_0.log>")
  305. more_output_config.add_argument(
  306. "--output-basename",
  307. help="the basename for the output files\n[DEFAULT: " +
  308. "input file basename]",
  309. default=config.file_basename,
  310. metavar="<sample_name>")
  311. more_output_config.add_argument(
  312. "--output-format",
  313. help="the format of the output files\n[DEFAULT: " +
  314. config.output_format + "]",
  315. default=config.output_format,
  316. choices=config.output_format_choices)
  317. more_output_config.add_argument(
  318. "--output-max-decimals",
  319. help="the number of decimals to output\n[DEFAULT: " +
  320. str(config.output_max_decimals) + "]",
  321. metavar="<" + str(config.output_max_decimals) + ">",
  322. type=int,
  323. default=config.output_max_decimals)
  324. more_output_config.add_argument(
  325. "--remove-column-description-output",
  326. help="remove the description in the output column\n" +
  327. "[DEFAULT: output column includes description]",
  328. action="store_true",
  329. default=config.remove_column_description_output)
  330. more_output_config.add_argument(
  331. "--remove-stratified-output",
  332. help="remove stratification from output\n" +
  333. "[DEFAULT: output is stratified]",
  334. action="store_true",
  335. default=config.remove_stratified_output)
  336. return parser.parse_args()
  337. def update_configuration(args):
  338. """
  339. Update the configuration settings based on the arguments
  340. """
  341. # Use the full path to the input file
  342. args.input=os.path.abspath(args.input)
  343. # set the version header
  344. config.version_header=config.version_header+VERSION
  345. # If set, append paths executable locations
  346. if args.metaphlan:
  347. utilities.add_exe_to_path(os.path.abspath(args.metaphlan))
  348. if args.bowtie2:
  349. utilities.add_exe_to_path(os.path.abspath(args.bowtie2))
  350. if args.diamond:
  351. utilities.add_exe_to_path(os.path.abspath(args.diamond))
  352. # Set the metaphlan options, removing any extra spaces
  353. if args.metaphlan_options:
  354. config.metaphlan_opts=list(filter(None,args.metaphlan_options.split(" ")))
  355. # check for custom diamond options
  356. config.diamond_opts=args.diamond_options
  357. if args.diamond_options and args.diamond_options != config.diamond_opts_uniref50:
  358. config.diamond_options_custom = True
  359. config.diamond_opts=list(filter(None,args.diamond_options.split(" ")))
  360. if args.bowtie_options:
  361. config.bowtie2_align_opts=list(filter(None,args.bowtie_options.split(" ")))
  362. # Set the pathways database selection
  363. if args.pathways == "metacyc":
  364. config.pathways_database_part1=config.metacyc_gene_to_reactions
  365. config.pathways_database_part2=config.metacyc_reactions_to_pathways
  366. config.pathways_ec_column=True
  367. elif args.pathways == "unipathway":
  368. config.pathways_database_part1=config.unipathway_database_part1
  369. config.pathways_database_part2=config.unipathway_database_part2
  370. config.pathways_ec_column=True
  371. # Set the locations of the pathways databases
  372. # If provided by the user, this will take precedence over the pathways database selection
  373. if args.pathways_database:
  374. custom_pathways_files=args.pathways_database[0].split(",")
  375. if len(custom_pathways_files)==2:
  376. config.pathways_database_part1=os.path.abspath(custom_pathways_files[0])
  377. config.pathways_database_part2=os.path.abspath(custom_pathways_files[1])
  378. elif len(custom_pathways_files)==1:
  379. config.pathways_database_part1=None
  380. config.pathways_database_part2=os.path.abspath(custom_pathways_files[0])
  381. else:
  382. sys.exit("ERROR: Please provide one or two pathways files.")
  383. config.pathways_ec_column=False
  384. # Set the locations of the other databases
  385. if args.nucleotide_database:
  386. config.nucleotide_database=os.path.abspath(args.nucleotide_database)
  387. if args.protein_database:
  388. config.protein_database=os.path.abspath(args.protein_database)
  389. # if set, update the config run mode to resume
  390. if args.resume:
  391. config.resume=True
  392. # if set, update the config run mode to verbose
  393. if args.verbose:
  394. config.verbose=True
  395. # if set, update the config run mode to bypass prescreen step
  396. if args.bypass_prescreen:
  397. config.bypass_prescreen=True
  398. # if set, update the config run mode to bypass nucleotide index step
  399. if args.bypass_nucleotide_index:
  400. config.bypass_nucleotide_index=True
  401. config.bypass_prescreen=True
  402. # if set, update the config run mode to bypass translated search step
  403. # set the pick_frames toggle based on the bypass
  404. config.pick_frames_toggle="off"
  405. if args.bypass_translated_search:
  406. config.bypass_translated_search=True
  407. # if set, update the config run mode to bypass nucleotide search steps
  408. if args.bypass_nucleotide_search:
  409. config.bypass_prescreen=True
  410. config.bypass_nucleotide_index=True
  411. config.bypass_nucleotide_search=True
  412. # update the average read length
  413. config.average_read_length=args.average_read_length
  414. # Update thresholds
  415. config.prescreen_threshold=args.prescreen_threshold
  416. config.translated_subject_coverage_threshold=args.translated_subject_coverage_threshold
  417. config.nucleotide_subject_coverage_threshold=args.nucleotide_subject_coverage_threshold
  418. config.translated_query_coverage_threshold=args.translated_query_coverage_threshold
  419. config.nucleotide_query_coverage_threshold=args.nucleotide_query_coverage_threshold
  420. # Update the max decimals output
  421. config.output_max_decimals=args.output_max_decimals
  422. # Update memory use
  423. config.memory_use=args.memory_use
  424. # Update threads
  425. config.threads=args.threads
  426. # Update the evalue threshold
  427. config.evalue_threshold=args.evalue
  428. # Update translated alignment software
  429. config.translated_alignment_selected=args.translated_alignment
  430. # Update the computation toggle choices
  431. config.xipe_toggle=args.xipe
  432. config.minpath_toggle=args.minpath
  433. config.gap_fill_toggle=args.gap_fill
  434. config.count_normalization=args.count_normalization
  435. if args.utility_database:
  436. config.utility_mapping_database=os.path.abspath(args.utility_database)
  437. # Check that the input file exists and is readable
  438. if not os.path.isfile(args.input):
  439. sys.exit("CRITICAL ERROR: Can not find input file selected: "+ args.input)
  440. if not os.access(args.input, os.R_OK):
  441. sys.exit("CRITICAL ERROR: Not able to read input file selected: " + args.input)
  442. # Update the output format
  443. config.remove_stratified_output=args.remove_stratified_output
  444. config.remove_column_description_output=args.remove_column_description_output
  445. # Check that the output directory is writeable
  446. output_dir = os.path.abspath(args.output)
  447. if not os.path.isdir(output_dir):
  448. try:
  449. print("Creating output directory: " + output_dir)
  450. os.mkdir(output_dir)
  451. except EnvironmentError:
  452. sys.exit("CRITICAL ERROR: Unable to create output directory.")
  453. if not os.access(output_dir, os.W_OK):
  454. sys.exit("CRITICAL ERROR: The output directory is not " +
  455. "writeable. This software needs to write files to this directory.\n" +
  456. "Please select another directory.")
  457. print("Output files will be written to: " + output_dir)
  458. # Set the basename of the output files if specified as an option
  459. if args.output_basename:
  460. config.file_basename=args.output_basename
  461. else:
  462. # Determine the basename of the input file to use as output file basename
  463. input_file_basename=os.path.basename(args.input)
  464. # Remove gzip extension if present
  465. if re.search('.gz$',input_file_basename):
  466. input_file_basename='.'.join(input_file_basename.split('.')[:-1])
  467. # Remove input file extension if present
  468. if '.' in input_file_basename:
  469. input_file_basename='.'.join(input_file_basename.split('.')[:-1])
  470. config.file_basename=input_file_basename
  471. # Set the output format
  472. config.output_format=args.output_format
  473. # Set final output file names and location
  474. config.pathabundance_file=os.path.join(output_dir,
  475. config.file_basename + config.pathabundance_file + "." +
  476. config.output_format)
  477. config.pathcoverage_file=os.path.join(output_dir,
  478. config.file_basename + config.pathcoverage_file + "." +
  479. config.output_format)
  480. config.genefamilies_file=os.path.join(output_dir,
  481. config.file_basename + config.genefamilies_file + "." +
  482. config.output_format)
  483. config.profile_file=os.path.join(output_dir,
  484. config.file_basename + config.profile_file + "." +
  485. config.output_format)
  486. config.reactions_file=os.path.join(output_dir,
  487. config.file_basename + config.reactions_file + "." +
  488. config.output_format)
  489. # set the location of the temp directory
  490. if not args.remove_temp_output:
  491. config.temp_dir=os.path.join(output_dir,config.file_basename+"_humann_temp")
  492. if not os.path.isdir(config.temp_dir):
  493. try:
  494. os.mkdir(config.temp_dir)
  495. except EnvironmentError:
  496. sys.exit("Unable to create temp directory: " + config.temp_dir)
  497. else:
  498. config.temp_dir=tempfile.mkdtemp(
  499. prefix=config.file_basename+'_humann_temp_',dir=output_dir)
  500. # create the unnamed temp directory
  501. config.unnamed_temp_dir=tempfile.mkdtemp(dir=config.temp_dir)
  502. # set the name of the log file
  503. log_file=os.path.join(output_dir,config.file_basename+"_0.log")
  504. # change file name if set
  505. if args.o_log:
  506. log_file=args.o_log
  507. # configure the logger
  508. logging.basicConfig(filename=log_file,format='%(asctime)s - %(name)s - %(levelname)s: %(message)s',
  509. level=getattr(logging,args.log_level), filemode='w', datefmt='%m/%d/%Y %I:%M:%S %p')
  510. # write the version of the software to the log
  511. logger.info("Running humann v"+VERSION)
  512. # write the location of the output files to the log
  513. logger.info("Output files will be written to: " + output_dir)
  514. # write the location of the temp file directory to the log
  515. message="Writing temp files to directory: " + config.temp_dir
  516. logger.info(message)
  517. if config.verbose:
  518. print("\n"+message+"\n")
  519. return log_file
  520. def parse_chocophlan_gene_indexes(annotation_gene_index):
  521. """ Parse the chocophlan gene index input """
  522. # Update the chocophlan gene indexes
  523. chocophlan_gene_indexes=[]
  524. for index in annotation_gene_index.split(","):
  525. # Look for array range
  526. if ":" in index:
  527. split_index=index.split(":")
  528. start=split_index[0]
  529. end=split_index[1]
  530. try:
  531. start=int(start)
  532. end=int(end)
  533. chocophlan_gene_indexes+=range(start,end)
  534. except ValueError:
  535. pass
  536. else:
  537. # Convert to int
  538. try:
  539. index=int(index)
  540. chocophlan_gene_indexes.append(index)
  541. except ValueError:
  542. pass
  543. return chocophlan_gene_indexes
  544. def check_requirements(args):
  545. """
  546. Check requirements (file format, dependencies, permissions)
  547. """
  548. # Check the pathways database files exist and are readable
  549. if config.pathways_database_part1:
  550. utilities.file_exists_readable(config.pathways_database_part1)
  551. utilities.file_exists_readable(config.pathways_database_part2)
  552. # Determine the input file format if not provided
  553. if not args.input_format:
  554. args.input_format=utilities.determine_file_format(args.input)
  555. if args.input_format == "unknown":
  556. sys.exit("CRITICAL ERROR: Unable to determine the input file format." +
  557. " Please provide the format with the --input_format argument.")
  558. config.input_format = args.input_format
  559. # If the input file is compressed, then decompress
  560. if args.input_format.endswith(".gz"):
  561. new_file=utilities.gunzip_file(args.input)
  562. if new_file:
  563. args.input=new_file
  564. args.input_format=args.input_format.split(".")[0]
  565. else:
  566. sys.exit("CRITICAL ERROR: Unable to use gzipped input file. " +
  567. " Please check the format of the input file.")
  568. # check if the input file has sequence identifiers of the new illumina casava v1.8+ format
  569. # these have spaces causing the paired end reads to have the same identifier after
  570. # delimiting by space (causing bowtie2/diamond to label the reads with the same identifier)
  571. if args.input_format in ["fasta","fastq"]:
  572. if utilities.space_in_identifier(args.input):
  573. message="Removing spaces from identifiers in input file"
  574. print(message+" ...\n")
  575. logger.info(message)
  576. new_file = utilities.remove_spaces_from_file(args.input)
  577. if new_file:
  578. args.input=new_file
  579. else:
  580. sys.exit("CRITICAL ERROR: Unable to remove spaces from identifiers in input file.")
  581. # If the input format is in binary then convert to sam (tab-delimited text)
  582. if args.input_format == "bam":
  583. # Check for the samtools software
  584. if not utilities.find_exe_in_path("samtools"):
  585. sys.exit("CRITICAL ERROR: The samtools executable can not be found. "
  586. "Please check the install or select another input format.")
  587. new_file=utilities.bam_to_sam(args.input)
  588. if new_file:
  589. args.input=new_file
  590. args.input_format="sam"
  591. else:
  592. sys.exit("CRITICAL ERROR: Unable to convert bam input file to sam.")
  593. # If the input format is in biom then convert to tsv
  594. if args.input_format == "biom":
  595. # Check for the biom software
  596. if not utilities.find_exe_in_path("biom"):
  597. sys.exit("CRITICAL ERROR: The biom executable can not be found. "
  598. "Please check the install or select another input format.")
  599. new_file=utilities.biom_to_tsv(args.input)
  600. if new_file:
  601. args.input=new_file
  602. # determine the format of the file
  603. args.input_format=utilities.determine_file_format(args.input)
  604. else:
  605. sys.exit("CRITICAL ERROR: Unable to convert biom input file to tsv.")
  606. # If the biom output format is selected, check for the biom package
  607. if config.output_format=="biom":
  608. if not utilities.find_exe_in_path("biom"):
  609. sys.exit("CRITICAL ERROR: The biom executable can not be found. "
  610. "Please check the install or select another output format.")
  611. try:
  612. import biom
  613. except ImportError:
  614. sys.exit("Could not find the biom software."+
  615. " This software is required since the output file is a biom file.")
  616. if os.path.basename(config.utility_mapping_database) == "utility_DEMO":
  617. # Check the input file is a demo input if running with demo database
  618. try:
  619. input_file_size=os.path.getsize(args.input)/1024**2
  620. except EnvironmentError:
  621. input_file_size=0
  622. if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
  623. sys.exit("ERROR: You are using the demo utility database with "
  624. + "a non-demo input file. If you have not already done so, please "
  625. + "run humann_databases to download the full utility database. "
  626. + "If you have downloaded the full database, use the option "
  627. + "--utility-database to provide the location. "
  628. + "You can also run humann_config to update the default "
  629. + "database location. For additional information, please "
  630. + "see the HUMAnN User Manual.")
  631. # If the file is fasta/fastq check for requirements
  632. if args.input_format in ["fasta","fastq"]:
  633. # Check that the chocophlan directory exists
  634. if not config.bypass_nucleotide_index:
  635. if not os.path.isdir(config.nucleotide_database):
  636. if args.nucleotide_database:
  637. sys.exit("CRITICAL ERROR: The directory provided for the ChocoPhlAn database at "
  638. + args.nucleotide_database + " does not exist. Please select another directory.")
  639. else:
  640. sys.exit("CRITICAL ERROR: The default ChocoPhlAn database directory of "
  641. + config.nucleotide_database + " does not exist. Please provide the location "
  642. + "of the ChocoPhlAn directory using the --nucleotide-database option.")
  643. # Check that the files in the chocophlan folder are of the right format
  644. if not config.bypass_nucleotide_index:
  645. valid_format_count=0
  646. for file in os.listdir(config.nucleotide_database):
  647. # expect most of the file names to be of the format g__*s__*
  648. if re.search("^SGB",file):
  649. valid_format_count+=1
  650. if not config.metaphlan_v4_db_matching_uniref in file:
  651. sys.exit("\n\nCRITICAL ERROR: The directory provided for ChocoPhlAn contains files ( "+file+" )"+\
  652. " that are not of the expected version. Please install the latest version"+\
  653. " of the database: "+config.metaphlan_v4_db_matching_uniref)
  654. if valid_format_count == 0:
  655. sys.exit("CRITICAL ERROR: The directory provided for ChocoPhlAn does not "
  656. + "contain files of the expected format (ie \'^SGB\').")
  657. # Check if running with the demo database
  658. if not config.bypass_nucleotide_index:
  659. if os.path.basename(config.nucleotide_database) == "chocophlan_DEMO":
  660. # Check the input file is a demo input if running with demo database
  661. try:
  662. input_file_size=os.path.getsize(args.input)/1024**2
  663. except EnvironmentError:
  664. input_file_size=0
  665. if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
  666. sys.exit("ERROR: You are using the demo ChocoPhlAn database with "
  667. + "a non-demo input file. If you have not already done so, please "
  668. + "run humann_databases to download the full ChocoPhlAn database. "
  669. + "If you have downloaded the full database, use the option "
  670. + "--nucleotide-database to provide the location. "
  671. + "You can also run humann_config to update the default "
  672. + "database location. For additional information, please "
  673. + "see the HUMAnN User Manual.")
  674. # Check that the metaphlan2 executable can be found
  675. if not config.bypass_prescreen and not config.bypass_nucleotide_index:
  676. if not utilities.find_exe_in_path("metaphlan"):
  677. sys.exit("CRITICAL ERROR: The metaphlan executable can not be found. "
  678. "Please check the install.")
  679. # Check the metaphlan2 version
  680. # utilities.check_software_version("metaphlan",config.metaphlan_version)
  681. # Check that the bowtie2 executable can be found
  682. if not config.bypass_nucleotide_search:
  683. if not utilities.find_exe_in_path("bowtie2"):
  684. sys.exit("CRITICAL ERROR: The bowtie2 executable can not be found. "
  685. "Please check the install.")
  686. # Check the bowtie2 version
  687. utilities.check_software_version("bowtie2", config.bowtie2_version, warning=True)
  688. if not config.bypass_translated_search:
  689. # Check that the protein database directory exists
  690. if not os.path.isdir(config.protein_database):
  691. if args.protein_database:
  692. sys.exit("CRITICAL ERROR: The directory provided for the protein database at "
  693. + args.protein_database + " does not exist. Please select another directory.")
  694. else:
  695. sys.exit("CRITICAL ERROR: The default protein database directory of "
  696. + config.protein_database + " does not exist. Please provide the location "
  697. + "of the directory using the --protein-database option.")
  698. # Check that some files in the protein database folder are of the expected extension
  699. expected_database_extension=""
  700. if config.translated_alignment_selected == "usearch":
  701. expected_database_extension=config.usearch_database_extension
  702. elif config.translated_alignment_selected == "rapsearch":
  703. expected_database_extension=config.rapsearch_database_extension
  704. elif config.translated_alignment_selected == "diamond":
  705. expected_database_extension=config.diamond_database_extension
  706. valid_format_count=0
  707. database_files=os.listdir(config.protein_database)
  708. valid_format_database_files=[]
  709. for file in database_files:
  710. if not config.matching_uniref in file:
  711. sys.exit("\n\nCRITICAL ERROR: The directory provided for the translated database contains files ( "+file+" )"+\
  712. " that are not of the expected version. Please install the latest version"+\
  713. " of the database: "+config.matching_uniref)
  714. if file.endswith(expected_database_extension):
  715. # if rapsearch check for the second database file
  716. if config.translated_alignment_selected == "rapsearch":
  717. database_file=re.sub(config.rapsearch_database_extension+"$","",file)
  718. if database_file in database_files:
  719. valid_format_count+=1
  720. valid_format_database_files.append(database_file)
  721. else:
  722. valid_format_count+=1
  723. valid_format_database_files.append(file)
  724. if valid_format_count == 0:
  725. sys.exit("CRITICAL ERROR: The protein database directory provided ( " + config.protein_database
  726. + " ) does not contain any files that have been formatted to run with"
  727. " the translated alignment software selected ( " +
  728. config.translated_alignment_selected + " ). Please format these files so"
  729. + " they are of the expected extension ( " + expected_database_extension +" ).")
  730. # Check if running with the demo database
  731. if not config.bypass_translated_search:
  732. if os.path.basename(config.protein_database) == "uniref_DEMO":
  733. # Check the input file is a demo input if running with demo database
  734. try:
  735. input_file_size=os.path.getsize(args.input)/1024**2
  736. except EnvironmentError:
  737. input_file_size=0
  738. if input_file_size > MAX_SIZE_DEMO_INPUT_FILE:
  739. sys.exit("ERROR: You are using the demo UniRef database with "
  740. + "a non-demo input file. If you have not already done so, please "
  741. + "run humann_databases to download the full UniRef database. "
  742. + "If you have downloaded the full database, use the option "
  743. + "--protein-database to provide the location. "
  744. + "You can also run humann_config to update the default "
  745. + "database location. For additional information, please "
  746. + "see the HUMAnN User Manual.")
  747. # Check that the translated alignment executable can be found
  748. if not utilities.find_exe_in_path(config.translated_alignment_selected):
  749. sys.exit("CRITICAL ERROR: The " + config.translated_alignment_selected +
  750. " executable can not be found. Please check the install.")
  751. # Check for correct usearch version
  752. if config.translated_alignment_selected == "usearch":
  753. utilities.check_software_version("usearch",config.usearch_version)
  754. # Check for correct rapsearch version
  755. if config.translated_alignment_selected == "rapsearch":
  756. utilities.check_software_version("rapsearch", config.rapsearch_version)
  757. # Check for the correct diamond version
  758. if config.translated_alignment_selected == "diamond":
  759. utilities.check_software_version("diamond", config.diamond_version)
  760. # parse the chocophlan gene index
  761. chocophlan_gene_indexes=parse_chocophlan_gene_indexes(args.annotation_gene_index)
  762. config.chocophlan_gene_indexes = chocophlan_gene_indexes
  763. # set the identity thresholds if provided by the user
  764. config.nucleotide_identity_threshold=args.nucleotide_identity_threshold
  765. config.identity_threshold=args.identity_threshold
  766. def timestamp_message(task, start_time):
  767. """
  768. Print and log a message about the task completed and the time
  769. Log messages are tab delimited for quick task/time access with awk
  770. Return the new start time
  771. """
  772. message="TIMESTAMP: Completed \t" + task + " \t:\t " + \
  773. str(int(round(time.time() - start_time))) + "\t seconds"
  774. logger.info(message)
  775. if config.verbose:
  776. print("\n"+message.replace("\t","")+"\n")
  777. return time.time()
  778. def main():
  779. # Parse arguments from command line
  780. args=parse_arguments(sys.argv)
  781. # Update the configuration settings based on the arguments
  782. log_file=update_configuration(args)
  783. # Check for required files, software, databases, and also permissions
  784. check_requirements(args)
  785. # Write the config settings to the log file
  786. config.log_settings()
  787. # Initialize alignments and gene scores
  788. minimize_memory_use=True
  789. if config.memory_use == "maximum":
  790. minimize_memory_use=False
  791. alignments=store.Alignments(minimize_memory_use=minimize_memory_use)
  792. unaligned_reads_store=store.Reads(minimize_memory_use=minimize_memory_use)
  793. gene_scores=store.GeneScores()
  794. # If id mapping is provided then process
  795. if args.id_mapping:
  796. alignments.process_id_mapping(args.id_mapping)
  797. # Load in the reactions database
  798. reactions_database=None
  799. if config.pathways_database_part1:
  800. reactions_database=store.ReactionsDatabase(config.pathways_database_part1)
  801. message="Load pathways database part 1: " + config.pathways_database_part1
  802. logger.info(message)
  803. # Load in the pathways database
  804. pathways_database=store.PathwaysDatabase(config.pathways_database_part2, reactions_database)
  805. if config.pathways_database_part1:
  806. message="Load pathways database part 2: " + config.pathways_database_part2
  807. else:
  808. message="Load pathways database: " + config.pathways_database_part2
  809. logger.info(message)
  810. # Start timer
  811. start_time=time.time()
  812. # Process fasta or fastq input files
  813. output_files=[]
  814. if args.input_format in ["fasta","fastq"]:
  815. # Run prescreen to identify bugs
  816. bug_file = "Empty"
  817. if args.taxonomic_profile:
  818. bug_file = os.path.abspath(args.taxonomic_profile)
  819. else:
  820. if not config.bypass_prescreen:
  821. bug_file = prescreen.alignment(args.input)
  822. output_files.append(bug_file)
  823. start_time=timestamp_message("prescreen",start_time)
  824. # Create the custom database from the bugs list
  825. custom_database = ""
  826. if not config.bypass_nucleotide_index:
  827. custom_database = prescreen.create_custom_database(config.nucleotide_database, bug_file)
  828. start_time=timestamp_message("custom database creation",start_time)
  829. else:
  830. custom_database = "Bypass"
  831. # Run nucleotide search on custom database
  832. if custom_database != "Empty" and not config.bypass_nucleotide_search:
  833. if not config.bypass_nucleotide_index:
  834. nucleotide_index_file = nucleotide.index(custom_database)
  835. start_time=timestamp_message("database index",start_time)
  836. else:
  837. nucleotide_index_file = nucleotide.find_index(config.nucleotide_database)
  838. nucleotide_alignment_file = nucleotide.alignment(args.input,
  839. nucleotide_index_file)
  840. start_time=timestamp_message("nucleotide alignment",start_time)
  841. # Determine which reads are unaligned and reduce aligned reads file
  842. # Remove the alignment_file as we only need the reduced aligned reads file
  843. [ unaligned_reads_file_fasta, reduced_aligned_reads_file ] = nucleotide.unaligned_reads(
  844. nucleotide_alignment_file, alignments, unaligned_reads_store, keep_sam=True)
  845. start_time=timestamp_message("nucleotide alignment post-processing",start_time)
  846. # Print out total alignments per bug
  847. message="Total bugs from nucleotide alignment: " + str(alignments.count_bugs())
  848. logger.info(message)
  849. print(message)
  850. message=alignments.counts_by_bug()
  851. logger.info("\n"+message)
  852. print(message)
  853. message="Total gene families from nucleotide alignment: " + str(alignments.count_genes())
  854. logger.info(message)
  855. print("\n"+message)
  856. # Report reads unaligned
  857. message="Unaligned reads after nucleotide alignment: " + utilities.estimate_unaligned_reads_stored(
  858. args.input, unaligned_reads_store) + " %"
  859. logger.info(message)
  860. print("\n"+message+"\n")
  861. else:
  862. logger.debug("Custom database is empty")
  863. reduced_aligned_reads_file = "Empty"
  864. unaligned_reads_file_fasta=args.input
  865. unaligned_reads_store=store.Reads(unaligned_reads_file_fasta, minimize_memory_use=minimize_memory_use)
  866. # Do not run if set to bypass translated search in config file
  867. if not config.bypass_translated_search:
  868. # Run translated search on UniRef database if unaligned reads exit
  869. if unaligned_reads_store.count_reads()>0:
  870. translated_alignment_file = translated.alignment(config.protein_database,
  871. unaligned_reads_file_fasta)
  872. start_time=timestamp_message("translated alignment",start_time)
  873. # Determine which reads are unaligned
  874. translated_unaligned_reads_file_fastq = translated.unaligned_reads(
  875. unaligned_reads_store, translated_alignment_file, alignments)
  876. start_time=timestamp_message("translated alignment post-processing",start_time)
  877. # Print out total alignments per bug
  878. message="Total bugs after translated alignment: " + str(alignments.count_bugs())
  879. logger.info(message)
  880. print(message)
  881. message=alignments.counts_by_bug()
  882. logger.info("\n"+message)
  883. print(message)
  884. message="Total gene families after translated alignment: " + str(alignments.count_genes())
  885. logger.info(message)
  886. print("\n"+message)
  887. # Report reads unaligned
  888. message="Unaligned reads after translated alignment: " + utilities.estimate_unaligned_reads_stored(
  889. args.input, unaligned_reads_store) + " %"
  890. logger.info(message)
  891. print("\n"+message+"\n")
  892. else:
  893. message="All reads are aligned so translated alignment will not be run"
  894. logger.info(message)
  895. print(message)
  896. else:
  897. message="Bypass translated search"
  898. logger.info(message)
  899. print(message)
  900. # Process input files of sam format
  901. elif args.input_format in ["sam"]:
  902. # Store the sam mapping results
  903. message="Process the sam mapping results ..."
  904. logger.info(message)
  905. print("\n"+message)
  906. [unaligned_reads_file_fasta, reduced_aligned_reads_file] = nucleotide.unaligned_reads(
  907. args.input, alignments, unaligned_reads_store, keep_sam=True)
  908. start_time=timestamp_message("alignment post-processing",start_time)
  909. # Process input files of tab-delimited blast format
  910. elif args.input_format in ["blastm8"]:
  911. # Store the blastm8 mapping results
  912. message="Process the blastm8 mapping results ..."
  913. logger.info(message)
  914. print("\n"+message)
  915. translated_unaligned_reads_file_fastq = translated.unaligned_reads(
  916. unaligned_reads_store, args.input, alignments)
  917. start_time=timestamp_message("alignment post-processing",start_time)
  918. # Get the number of remaining unaligned reads
  919. unaligned_reads_count=unaligned_reads_store.count_reads()
  920. # Clear all of the unaligned reads as they are no longer needed
  921. unaligned_reads_store.clear()
  922. # Compute or load in gene families
  923. if args.input_format in ["fasta","fastq","sam","blastm8"]:
  924. # Compute the gene families
  925. message="Computing gene families ..."
  926. logger.info(message)
  927. print("\n"+message)
  928. families_file=families.gene_families(alignments,gene_scores,unaligned_reads_count)
  929. output_files.append(families_file)
  930. start_time=timestamp_message("computing gene families",start_time)
  931. elif args.input_format in ["genetable"]:
  932. # Load the gene scores
  933. message="Process the gene table ..."
  934. logger.info(message)
  935. print("\n"+message)
  936. unaligned_reads_count=gene_scores.add_from_file(args.input,id_mapping_file=args.id_mapping)
  937. start_time=timestamp_message("processing gene table",start_time)
  938. # Handle input files of unknown formats
  939. else:
  940. sys.exit("CRITICAL ERROR: Input file of unknown format.")
  941. # Clear all of the alignments data as they are no longer needed
  942. alignments.clear()
  943. # Identify reactions and then pathways from the alignments
  944. message="Computing reaction and pathway abundance ..."
  945. logger.info(message)
  946. print("\n"+message)
  947. pathways_and_reactions_store=modules.identify_reactions_and_pathways(
  948. gene_scores, reactions_database, pathways_database, unaligned_reads_count)
  949. # Compute pathway abundance and coverage
  950. abundance_file, coverage_file, reaction_file=modules.compute_pathways_abundance_and_coverage(
  951. gene_scores, reactions_database, pathways_and_reactions_store, pathways_database, unaligned_reads_count)
  952. output_files.append(reaction_file)
  953. output_files.append(abundance_file)
  954. output_files.append(log_file)
  955. #output_files.append(coverage_file)
  956. start_time=timestamp_message("computing pathways",start_time)
  957. message="\nOutput files created: \n" + "\n".join(output_files) + "\n"
  958. logger.info(message)
  959. print(message)
  960. # Remove the unnamed temp files
  961. utilities.remove_directory(config.unnamed_temp_dir)
  962. # Remove named temp directory
  963. if args.remove_temp_output:
  964. utilities.remove_directory(config.temp_dir)

humann.py at commit e07b3a3, under MIT · at the source

Overview

  1. Doctorado en Ciencias Biológicas y de la Salud, Universidad Autónoma Metropolitana, Ciudad de Mexico 04960, Mexico
  2. Departamento de Salud Mental, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
  3. Departamento de Ciencias Ambientales, Universidad Autónoma Metropolitana-Lerma, Lerma 52006, Estado de Mexico, Mexico
  4. Grupo Neurológico, Neuroquirúrgico y de Columna, Hospital Ángeles Acoxpa, Ciudad de Mexico 14308, Mexico
  5. Laboratorio de Errores Innatos del Metabolismo y Tamiz, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
  6. Departamento de Higiene y Tecnología de los Alimentos, Facultad de Veterinaria, Universidad de León, 24071 Leon, Spain
  7. Laboratorio de Oncología Experimental, Instituto Nacional de Pediatría, Ciudad de Mexico 04530, Mexico
Journal: Nutrients, volume 18, issue 15, article 2441
Dates: received 27 May 2026; accepted 22 July 2026; published online 26 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.3390/nu18152441 · PMID 42588064 · PMCID PMC13468271 · OpenAlex W7171466161
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), autism (population), clinical / translational (subfield)
Methods: Statistics, Preprocessing, Connectivity
Keywords: autism spectrum disorder, probiotics, synbiotics, gut microbiota, microbiome function, ASD severity, gastrointestinal symptoms
MeSH: Autism Spectrum Disorder*, Gastrointestinal Microbiome*, Synbiotics*, Child, Child, Preschool, Dietary Supplements, Dysbiosis, Feces, Female, Humans, Longitudinal Studies, Male, Mexico, Probiotics, RNA, Ribosomal, 16S, Treatment Outcome (* major topic)
Topic: Gut microbiota and health (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Institute of Pediatrics
Citations: not cited yet (Europe PMC); 90 references in the paper

Abstract

Background/Objectives: Gut dysbiosis in children with autism spectrum disorder (ASD) has been associated with alterations in microbial ecology and metabolic function that may contribute to gastrointestinal dysfunction and the severity of clinical manifestations. Synbiotic and probiotic supplementation has emerged as a promising microbiome-targeted strategy for ASD; however, its effects on gut microbiome composition, functional potential, and clinical outcomes remain incompletely understood. We conducted a longitudinal study of Mexican children diagnosed with ASD to analyze changes in the composition, diversity, and functional potential of the gut microbiome during six months of multi-strain synbiotic supplementation. Methods: Stool samples were collected from 25 children with ASD at baseline and after 3 and 6 months of multi-strain synbiotic supplementation. Gut microbiome composition and diversity were analyzed by 16S rRNA gene sequencing, whereas whole metagenome sequencing (WMS) was performed in a subset of samples to evaluate the functional potential of the fecal microbiome. Gastrointestinal symptoms were assessed using the Rome IV criteria, and ASD severity was evaluated with the Childhood Autism Rating Scale (CARS). Results: Twenty-five children with ASD completed the 6 months of synbiotic supplementation. Overall, ASD severity decreased, reflected by a reduction in total CARS score, and improvements in several CARS domains. Gastrointestinal symptoms also decreased significantly. Longitudinal microbiome profiling revealed significant taxonomic and diversity changes over the supplementation period, while WMS identified changes in microbial metabolic potential, including enrichment of tryptophan biosynthesis pathways and reduced L-rhamnose degradation. Conclusions: This exploratory research provides proof-of-concept evidence supporting multi-strain synbiotic supplementation in children with ASD. Larger controlled studies are needed to confirm these findings and clarify their relevance to microbiota–gut–brain axis interactions. The observed concordance between clinical improvements and microbiome remodeling supports further investigation of microbiome-targeted interventions according to ASD severity and duration of supplementation.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

biobakery/humann

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: e07b3a34d0b94c09a8ac5d28ff95009611178be2, 10 July 2026
Languages: Python (61), Shell (10)
Size: 206 files, 71 scripts
Software Heritage: archived
Found in: the text, “2.6. Bioinformatic and Statistical Analysis”
Holds: README, license file, environment (setup.py), tests
Not found: CITATION.cff, continuous integration, documentation
Tools: NumPy (4 files), h5py (3 files), SciPy (2 files), Matplotlib (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
73 files

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 71 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data Availability Statement

All 16S rDNA and Whole Metagenomic Sequence files and corresponding mapping files for the samples utilized in this study have been deposited in the NCBI BioSample repository. Interested parties may access these resources using the provided Accession Number PRJNA1425957, which can be found at the following link: https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA1425957 (accessed on 27 May, 2026).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 9 authors, 7 keywords, 16 MeSH terms, 1 funder, 90 references.

Cite

This paper

De Sales-Millan, A., Reyes-Ferreira, P., González-Cervantes, R. M., Luna-Álvarez, M., Guillén-López, S., Cobo-Díaz, J. F., Ramos, S., Aguirre-Garrido, J. F., & Velázquez-Aragón, J. A. (2026). Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder. Nutrients, 18(15), 2441. https://doi.org/10.3390/nu18152441

BibTeX

@article{desalesmillan2026clinical,
author = {De Sales-Millan, Amapola and Reyes-Ferreira, Paulina and González-Cervantes, Rina María and Luna-Álvarez, Mariana and Guillén-López, Sara and Cobo-Díaz, José F. and Ramos, Sandra and Aguirre-Garrido, José Félix and Velázquez-Aragón, José Antonio},
title = {{Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder}},
journal = {Nutrients},
year = {2026},
month = jul,
volume = {18},
number = {15},
pages = {2441},
publisher = {Multidisciplinary Digital Publishing Institute (MDPI)},
issn = {2072-6643},
doi = {10.3390/nu18152441},
url = {https://doi.org/10.3390/nu18152441},
pmid = {42588064},
pmcid = {PMC13468271}
}

RIS

TY - JOUR
AU - De Sales-Millan, Amapola
AU - Reyes-Ferreira, Paulina
AU - González-Cervantes, Rina María
AU - Luna-Álvarez, Mariana
AU - Guillén-López, Sara
AU - Cobo-Díaz, José F.
AU - Ramos, Sandra
AU - Aguirre-Garrido, José Félix
AU - Velázquez-Aragón, José Antonio
TI - Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder
T2 - Nutrients
J2 - Nutrients
PY - 2026
DA - 2026/07/26
VL - 18
IS - 15
SP - 2441
SN - 2072-6643
PB - Multidisciplinary Digital Publishing Institute (MDPI)
DO - 10.3390/nu18152441
UR - https://doi.org/10.3390/nu18152441
LA - en
ER -

CSL-JSON

{
"id": "10.3390/nu18152441",
"type": "article-journal",
"title": "Clinical Improvement and Taxonomic-Functional Gut Microbiome Remodeling After Six Months of Multi-Strain Synbiotic Supplementation in Mexican Children with Autism Spectrum Disorder",
"container-title": "Nutrients",
"author": [
{
"family": "De Sales-Millan",
"given": "Amapola"
},
{
"family": "Reyes-Ferreira",
"given": "Paulina"
},
{
"family": "González-Cervantes",
"given": "Rina María"
},
{
"family": "Luna-Álvarez",
"given": "Mariana"
},
{
"family": "Guillén-López",
"given": "Sara"
},
{
"family": "Cobo-Díaz",
"given": "José F."
},
{
"family": "Ramos",
"given": "Sandra"
},
{
"family": "Aguirre-Garrido",
"given": "José Félix"
},
{
"family": "Velázquez-Aragón",
"given": "José Antonio"
}
],
"container-title-short": "Nutrients",
"volume": "18",
"issue": "15",
"page": "2441",
"DOI": "10.3390/nu18152441",
"PMID": "42588064",
"PMCID": "PMC13468271",
"ISSN": "2072-6643",
"publisher": "Multidisciplinary Digital Publishing Institute (MDPI)",
"URL": "https://doi.org/10.3390/nu18152441",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
26
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s12866-026-05296-x
Gut bacterial characteristics in children with autism spectrum disorder according to symptom severity: a cross-sectional study.
Journal: BMC microbiology
In common: autism, 4 references
[2] doi:10.1186/s12888-026-08178-8 [code]
Machine learning model to identify gut microbiome-derived metabolites as potential biomarkers of autism spectrum disorder: a pilot study.
Journal: BMC psychiatry
In common: Matplotlib, NumPy, autism, clinical / translational, 2 references
[3] doi:10.1016/j.bbih.2026.101275 [code]
A protocol for the Teen Bugs study: An integrative, multi-omics approach to understanding the role of the gut microbiome and mesocorticolimbic system in adolescent mental health following early adverse caregiving.
Journal: Brain, behavior, & immunity - health
In common: 4 references
[4] doi:10.1016/j.gutmic.2026.100006 [code]
From microbes to milestones: Gut bacterial abundances and functional pathways associate with neurodevelopment following preterm birth.
Journal: Gut microbiology
In common: autism, 3 references
[5] doi:10.1002/hbm.70496 [code]
Transdiagnostic Profiles of BOLD Signal Variability in Autism and Schizophrenia Spectrum Disorders: Associations With Cognition and Functioning.
Journal: Human brain mapping
In common: SciPy, Matplotlib, NumPy, autism, 1 reference
[6] doi:10.3389/fnins.2026.1858005 [code]
Targeted stool metabolomics suggests exploratory catecholamine- and tryptophan-linked metabolic features in autism spectrum disorder.
Journal: Frontiers in neuroscience
In common: autism, clinical / translational, 2 references
[7] doi:10.1016/j.celrep.2026.117590 [code]
Impaired behavioral inhibition in Fmr1 KO mice is linked to disrupted visual cortex theta oscillations.
Journal: Cell reports
In common: SciPy, Matplotlib, NumPy, autism, 1 reference
[8] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: h5py, SciPy, Matplotlib, 1 other tool, autism
[9] doi:10.1038/s41467-026-74320-5 [code]
Spatial architecture of autism pathogenesis reveals mosaic structural disarray during early development.
Journal: Nature communications
In common: h5py, SciPy, Matplotlib, 1 other tool, autism
[10] doi:10.1002/advs.202519479 [code]
Diminished Signal-to-Noise Ratio Disrupts Somatosensory Population Encoding and Drives Tactile Hyposensitivity in the Fmr1&lt;sup&gt;-/y&lt;/sup&gt; Autism Model.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: h5py, SciPy, Matplotlib, 1 other tool, autism

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.