OSCR

Blood-based circular RNAs for early diagnosis of Alzheimer's disease.

Code ↔ Paper

2 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 2 matches
  1. [1] § Methods › Linear transcript identification ↔ linear_pipeline_circ_alignment/src/linear_RNA_pipeline.py, lines 318–325 · score 0.82 · Picard Collect RNA, Collect Alignment, Summary Metrics, linear RNA, FastQC, seq
  2. [2] § Methods › Linear transcript identification ↔ linear_pipeline_circ_alignment/src/run_Picard_QC.py, lines 24–36 · score 0.68 · Collect Alignment Summary, Picard Collect, Metrics, sequence, QC, linear

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 353 lines · 18 KB · no license · 1 match

  1. #!/usr/bin/env python3
  2. import argparse
  3. import logging
  4. import sys
  5. import shutil
  6. import re
  7. sys.path.append("/src")
  8. from fastq_to_ubam import *
  9. from run_STAR_js import align_STAR_normal, cram_samtools, index_samtools, sort_samtools, cram_samtools
  10. from run_Picard_QC import picard_collect_RNA_metrics, picard_collect_alignment_metrics, picard_mark_dups
  11. from run_Salmon import salmon_quant
  12. from run_tin import calc_tin
  13. from run_fastqc import fastqc_fastq_PE, fastqc_SE
  14. ################################################################################
  15. # Setup
  16. ################################################################################
  17. # Argparse to get all the input
  18. parser = argparse.ArgumentParser(description='Run linear RNAseq pipeline')
  19. # to do maybe make this one argument that sometimes needs two values
  20. parser.add_argument('-r1', '--raw_input', help='Path to raw data (Read 1 for PE data)', required = True)
  21. parser.add_argument('-r2', '--input_read_2', help = 'Path to raw data read 2 file for PE data')
  22. parser.add_argument('--sample', help = 'Sample name', required = True)
  23. # TODO maybe add a crammed option?
  24. parser.add_argument('--file_type', help = 'Input file type for -r1 and -r2', choices = ["ubam", "fastq", "bam"], required = True)
  25. parser.add_argument('--tmp_dir', help ='Path to directory for temporary files', required = True)
  26. parser.add_argument('--read_type', help = 'Paired end or Single end reads', choices = ['PE', 'SE'], required = True)
  27. parser.add_argument('--stranded', help = 'Include for strand specific libraries', default = "NONE", choices = ["NONE", "FIRST_READ_TRANSCRIPTION_STRAND", "SECOND_READ_TRANSCRIPTION_STRAND"])
  28. parser.add_argument('--cohort', help = 'Full RNAseq cohort name', required = True )
  29. parser.add_argument('--tissue', help = 'Sample tissue type for hydra dir path', choices = ['brain', 'blood_pax', 'ipsc', 'plasma', 'csf'])
  30. parser.add_argument('--STAR_index', help = "Path to the directory with the STAR genome index refernce", required = True)
  31. parser.add_argument('--ref_flat', help = "Path to reference annotation in flat format", required = True)
  32. parser.add_argument('--annotation', help = "Path to annotaiton file for Salmon", required = True)
  33. parser.add_argument('--annote_bed', help = "Path to annotation in bed format for TIN", required = True)
  34. parser.add_argument('--rib_int', help = "Ribosomal (rRNA) interval list", required = True)
  35. parser.add_argument('--transcripts', help = "Path to Salmon transcripts index", required = True)
  36. parser.add_argument('--adapter_seq', default = 'null', help ="Adapter sequence for picard")
  37. parser.add_argument('--out_dir', help = 'Directory for output', required = True)
  38. parser.add_argument('--out_struct', help = "Directory structure of output", default = 'none', choices = ['none', 'hydra'])
  39. parser.add_argument( '-d', '--delete',\
  40. nargs = '*',\
  41. choices = ["input", "aligned_bam", "aligned_sorted_bam", "bais", "STAR_extras", "ubam", "aligned_transcriptome_bam", "md_bam"],\
  42. help = 'Include to delete one or more following files: input files, .Aligned.out.bam, Aligned.sortedByCoord.out.bam, all .bai files,\
  43. ._STARgenome and ._STARpass1, _unmapped.bam, .Aligned.toTranscriptome.out.bam, and .Aligned.sortedByCoord.out.md.bam respectively')
  44. parser.add_argument('-c', '--cram', action = 'store_true', help = 'Cram the following bams: Aligned.sortedByCoord.out.md.bam')
  45. parser.add_argument('--ref', help = 'Path to reference genome fasta (for cramming)', required= True) # TODO change this to be required only when cramming
  46. parser.add_argument('--input_to_merge', help = 'Path to secondary input file to concatenate with primary input (ie MSBB which is separated into aligned and unaligend reads)')
  47. parser.add_argument('--merge_file_type', help = 'Input file type for --input_to_merge', choices = ["ubam", "fastq", "bam"], default = "fastq")
  48. args = parser.parse_args()
  49. if args.read_type == 'PE' and args.file_type == 'fastq' and args.input_read_2 is None:
  50. parser.error("--input_read_2 required if input is fastq and read type is PE")
  51. if args.out_struct == 'hydra' and args.tissue is None:
  52. parser.error("--tissue required if output structure is hydra")
  53. # Maybe make this an error later
  54. if args.raw_input == args.input_read_2:
  55. parser.error('--input_read_2 and --raw_input cannot be the same file. Please check input.')
  56. # set up logs
  57. log_file_name = args.cohort + '_linear_pipeline.log'
  58. log_path = os.path.join(args.out_dir, log_file_name)
  59. # Create logger
  60. linear_logs = logging.getLogger('linear_RNA_pipeline')
  61. linear_logs.setLevel(logging.DEBUG)
  62. # Handler 1: file for all info+ logs
  63. log_file = logging.FileHandler(log_path)
  64. file_format = logging.Formatter("%(asctime)s:%(levelname)s:%(message)s")
  65. log_file.setLevel(logging.INFO)
  66. log_file.setFormatter(file_format)
  67. # Handler 2: stream for all warning +
  68. stream = logging.StreamHandler()
  69. streamformat = logging.Formatter("%(levelname)s:%(module)s:%(message)s")
  70. stream.setLevel(logging.WARNING)
  71. stream.setFormatter(streamformat)
  72. # Add handlers to logs
  73. linear_logs.addHandler(log_file)
  74. linear_logs.addHandler(stream)
  75. ################################################################################
  76. # Pipeline Steps
  77. ################################################################################
  78. def check_refs_exist(star_ref, flat_ref, annot_ref, annot_bed_ref, rib_int_ref, transcript_ref):
  79. star_exists = os.path.exists(star_ref)
  80. flat_exists = os.path.exists(flat_ref)
  81. annot_exists = os.path.exists(annot_ref)
  82. annot_bed_exists = os.path.exists(annot_bed_ref)
  83. rib_int_exists = os.path.exists(rib_int_ref)
  84. transcript_exists = os.path.exists(transcript_ref)
  85. if not star_exists:
  86. sys.exit(f'STAR reference not found: {star_ref}, exiting')
  87. if not flat_exists:
  88. sys.exit(f'flat reference not found: {flat_ref}, exiting')
  89. if not annot_exists:
  90. sys.exit(f'Annotation not found: {annot_ref}, exiting')
  91. if not annot_bed_exists:
  92. sys.exit(f'Bed format annotation not found: {annot_bed_ref}, exiting')
  93. if not rib_int_exists:
  94. sys.exit(f'Ribosomal interval list not found: {rib_int_ref}, exiting')
  95. if not transcript_exists:
  96. sys.exit(f'Salmon transcripts index not found: {transcript_ref}, exiting')
  97. def create_out_dir(dir_to_create):
  98. already_exists = os.path.exists(dir_to_create)
  99. if not already_exists:
  100. os.makedirs(dir_to_create, exist_ok = True) #note this will make any dirs missing on the path
  101. def setup_output_dirs(output_struct, out_dir, cohort_name, tissue, sample_id):
  102. if output_struct == "hydra":
  103. if tissue == "brain":
  104. tissue_folder = "01-Brain"
  105. elif tissue == "blood_pax":
  106. tissue_folder = "02-Blood_PAXgene"
  107. elif tissue == "ipsc":
  108. tissue_folder = "03-iPSC"
  109. elif tissue == "plasma":
  110. tissue_folder = "04-Plasma"
  111. else:
  112. tissue_folder = "05-CSF"
  113. linear_logs.info(f'Setting up output structure to comply with Hydras dir structure and placing in {out_dir}')
  114. #If outstructure = hydra create paths for :
  115. # 02-Processed/<tissue>/<cohort>/01-FastQC/${SAMPLEID} -- fastqc.htlm, fastqc.zip, picard qc
  116. fastqc_out_dir = os.path.join(out_dir, '02-Processed/02-GRCh38/', tissue_folder, cohort_name, '01-FastQC', sample_id)
  117. create_out_dir(fastqc_out_dir)
  118. # 02-Processed/<tissue>/<cohort>/02-Linear_TIN/${SAMPLEID} -- .bam, .bai, .tin.csv, .tin.xls, salmon
  119. linear_tin_processed_dir = os.path.join(out_dir, '02-Processed/02-GRCh38/', tissue_folder, cohort_name, '02-Linear_TIN', sample_id)
  120. create_out_dir(linear_tin_processed_dir)
  121. output_dirs = {"fastqc": fastqc_out_dir, "linear_tin": linear_tin_processed_dir}
  122. else:
  123. print(f'Placing all output in dir: {out_dir}')
  124. output_dirs = {"fastqc": out_dir, "linear_tin": out_dir }
  125. return output_dirs
  126. def convert_to_ubam(out_dir, sample_name, file_type, read_type, raw_input, input_read_2, tmp_dir, rg_name = 'A'):
  127. ubam_out = os.path.join(out_dir, f"{sample_name}_unmapped.bam")
  128. if file_type == 'fastq':
  129. print('Converting fastq to ubam')
  130. if read_type == 'PE':
  131. print('Using ' + raw_input + ' and ' + input_read_2 +' as input for ummaped bam')
  132. fastq_to_ubam_PE(raw_input, input_read_2, ubam_out, sample_name, rg_name)
  133. star_input = ubam_out
  134. else:
  135. print('Converting single fq to unmapped bam')
  136. fastq_to_ubam_SE(raw_input, ubam_out, sample_name, rg_name)
  137. star_input = ubam_out
  138. elif args.file_type == 'bam':
  139. bam_to_ubam(raw_input, ubam_out, tmp_dir, by_readgroup = 'false'),
  140. star_input = ubam_out
  141. else:
  142. print('Input data already ubam, proceeding to alignment')
  143. star_input = raw_input
  144. return star_input
  145. def align_with_star(star_input, sample_name, read_type, STAR_index, out_dir, tmp_dir):
  146. out_prefix = os.path.join(out_dir, f'{sample_name}.')
  147. print('Aligning ubam with STAR')
  148. STAR_file_input = "SAM "+ read_type
  149. align_STAR_normal( star_input, out_prefix, STAR_file_input, STAR_index, tmp_dir)
  150. aligned_bam_out = f"{out_prefix}Aligned.out.bam"
  151. aligned_transcript_bam = f"{out_prefix}Aligned.toTranscriptome.out.bam"
  152. return (aligned_bam_out, aligned_transcript_bam)
  153. def sort(out_dir, sample_name, aligned_bam_in):
  154. print('Sorting with samtools')
  155. out_prefix = os.path.join(out_dir, sample_name)
  156. sort_samtools(aligned_bam_in, out_prefix)
  157. sorted_bam_out = f"{out_prefix}.Aligned.sortedByCoord.out.bam"
  158. return sorted_bam_out
  159. def index(sorted_bam_out):
  160. print('Indexing with samools')
  161. index_samtools(sorted_bam_out)
  162. indexed_bam_out = f"{sorted_bam_out}.bai"
  163. return indexed_bam_out
  164. def fastqc(out_dir, sample_name, input_file_type, input_read_type, raw_input, raw_input2):
  165. if input_file_type == 'fastq' and input_read_type == 'PE':
  166. fastqc_fastq_PE(raw_input, raw_input2, sample_name, out_dir)
  167. else:
  168. fastqc_SE(raw_input, sample_name, out_dir)
  169. def delete_extras(delete, sample_name, aligned_bam, sorted_bam, indexed_bam, STAR_dir, indexed_md_bam, merged_ubam, input_2, transcriptome_bam, md_bam):
  170. if delete is not None:
  171. print("Deleting files specified with --delete option")
  172. if 'input' in args.delete:
  173. print('Deleting input file(s)')
  174. os.remove(args.raw_input)
  175. linear_logs.info(f'Deleting file: {args.raw_input}')
  176. if args.input_read_2 is not None:
  177. os.remove(args.input_read_2)
  178. linear_logs.info(f'Deleting file: {input_2}')
  179. if 'aligned_bam' in args.delete:
  180. print('Deleting aligned bam')
  181. os.remove(aligned_bam)
  182. linear_logs.info(f'Deleting file: {aligned_bam}')
  183. if 'md_bam' in args.delete and args.cram is not True: # added cram qualification -- if cramming is turned on then this file won't exist
  184. print('Deleting md bam')
  185. os.remove(md_bam)
  186. linear_logs.info(f'Deleting file: {md_bam}')
  187. if 'aligned_sorted_bam' in args.delete:
  188. print('Deleting STAR aligned & sorted bam')
  189. os.remove(sorted_bam)
  190. linear_logs.info(f'Deleting file: {sorted_bam}')
  191. if 'bais' in args.delete:
  192. print('Deleting bai files')
  193. os.remove(indexed_bam)
  194. os.remove(indexed_md_bam)
  195. linear_logs.info(f'Deleting file: {indexed_bam} and {indexed_md_bam}')
  196. if 'ubam' in args.delete:
  197. print('Deleting unmapped bam file')
  198. unmapped_bam = os.path.join(STAR_dir, f"{sample_name}_unmapped.bam")
  199. os.remove(unmapped_bam)
  200. linear_logs.info(f'Deleting file: {unmapped_bam}')
  201. # remove the input to merge if it exists
  202. if input_2 is not None:
  203. os.remove(merged_ubam)
  204. unmapped_bam_input2 = os.path.join(STAR_dir, f"{sample_name}_input2_unmapped.bam")
  205. os.remove(unmapped_bam_input2)
  206. if 'STAR_extras' in args.delete:
  207. print('Deleting ._STARgenome and ._STARpass1')
  208. STAR_pass1 = os.path.join(STAR_dir, f"{sample_name}._STARpass1")
  209. STAR_genome = os.path.join(STAR_dir,f"{sample_name}._STARgenome" )
  210. shutil.rmtree(STAR_pass1)
  211. shutil.rmtree(STAR_genome)
  212. linear_logs.info(f'Deleting files: {STAR_pass1} and {STAR_genome}')
  213. if 'aligned_transcriptome_bam' in args.delete:
  214. print('Deleting STAR aligned to transcriptome bam file')
  215. os.remove(transcriptome_bam)
  216. linear_logs.info(f'Deleting file: {transcriptome_bam}')
  217. def cram_bams(cram, mark_dups_bam, ref, sample_name, out_dir ):
  218. if cram is True:
  219. print(f"Cramming {mark_dups_bam_out}")
  220. out_prefix = os.path.join(out_dir, sample_name)
  221. md_cram_out = f"{out_prefix}.Aligned.sortedByCoord.out.md.cram"
  222. cram_samtools(mark_dups_bam, md_cram_out, ref )
  223. linear_logs.info(f"Cramming: {mark_dups_bam}")
  224. #then delete the orginal bam
  225. linear_logs.info(f"Deleting {mark_dups_bam}")
  226. os.remove(mark_dups_bam)
  227. def get_MSBB_read_group(sample_name):
  228. #MSBB read group ID is just the sample name UNLESS it has a third _ in name -- then it is everything before this underscore
  229. regex = re.compile('^[^_]+_[^_]+_[^_]+')
  230. read_group_id = regex.findall(sample_name)[0]
  231. return read_group_id
  232. # Note: currently only setup to work with SE reads -- for MSBB
  233. def merge_files(out_dir, sample_name, input_1_ubam, input_2, input_2_file_type, read_type, read_2, tmp_dir):
  234. if input_2 is not None:
  235. # fix sample names so they are different
  236. sample_name_input2 = f"{sample_name}_input2"
  237. # convert second file to ubam (first file should already be)
  238. #get the read group for MSBB -- TODO probably should make this check if this is MSBB
  239. msbb_read_group = get_MSBB_read_group(sample_name)
  240. input_2_ubam = convert_to_ubam(out_dir, sample_name_input2, input_2_file_type, read_type, input_2, read_2, tmp_dir, msbb_read_group)
  241. # concatenate with samtools
  242. unmapped_cat_bam = os.path.join(out_dir, f"{sample_name}_input1_input2_cat.bam")
  243. concat_ubams(input_1_ubam, input_2_ubam, unmapped_cat_bam)
  244. return unmapped_cat_bam
  245. else:
  246. return input_1_ubam
  247. # 0.1 check all references exist (so that don't have to quit half way thorugh):
  248. check_refs_exist(args.STAR_index, args.ref_flat, args.annotation, args.annote_bed, args.rib_int, args.transcripts)
  249. # 0.2 Set up output paths -- have an option to have this automatically structure like hydra
  250. out_dirs = setup_output_dirs(args.out_struct, args.out_dir, args.cohort, args.tissue, args.sample)
  251. print(f'''Output locations: \n fastqc: {out_dirs["fastqc"]} \n linear and tin processed: {out_dirs["linear_tin"]}
  252. salmon quant: {out_dirs["linear_tin"]} \n tin_summary: {out_dirs["linear_tin"]}''')
  253. # 0.3 set up tmp dir -- include JOB ID in path to prevent conflicts
  254. tmp_dir_path = os.path.join(args.tmp_dir, os.getenv('LSB_JOBID'))
  255. create_out_dir(tmp_dir_path)
  256. print(f'''tmp location: {tmp_dir_path}''')
  257. # 1. Run fastqc (if we don't need to merge files first)
  258. if args.input_to_merge is None:
  259. fastqc(out_dirs["fastqc"], args.sample, args.file_type, args.read_type, args.raw_input, args.input_read_2)
  260. # 2. convert to Ubam (from fastq or aligned bam)
  261. ubam_out = convert_to_ubam(out_dirs["linear_tin"], args.sample, args.file_type, args.read_type, args.raw_input, args.input_read_2, tmp_dir_path)
  262. # 2.1 concatenate if needed
  263. merged_ubam_out = merge_files(out_dirs["linear_tin"], args.sample, ubam_out, args.input_to_merge, args.merge_file_type, args.read_type, args.input_read_2, tmp_dir_path)
  264. # 2.1.2 run fastqc on concatenated bams
  265. if args.input_to_merge is not None:
  266. fastqc(out_dirs["fastqc"], args.sample, "ubam", args.read_type, merged_ubam_out, args.input_read_2)
  267. # 3. Align with STAR
  268. aligned_bam_out, aligned_transcript_bam_out = align_with_star(merged_ubam_out, args.sample, args.read_type, args.STAR_index, out_dirs["linear_tin"], tmp_dir_path)
  269. # 4. samtools sort
  270. sorted_bam_out = sort(out_dirs["linear_tin"], args.sample, aligned_bam_out)
  271. # 5. samtools index
  272. indexed_bam_out = index(sorted_bam_out)
  273. # 6. Post Alignment Picard QC ( Collect RNAseq metrics, Collect Alignemt summary metrics, Mark dups)
  274. RNAseq_metrics_out = os.path.join(out_dirs["fastqc"], f"{args.sample}.RNA_Metrics.txt")
  275. picard_collect_RNA_metrics(sorted_bam_out, RNAseq_metrics_out, args.ref_flat, args.rib_int, tmp_dir_path, args.stranded)
  276. Alignment_metrics_out = os.path.join(out_dirs["fastqc"], f"{args.sample}.Summary_metrics.txt")
  277. picard_collect_alignment_metrics(sorted_bam_out, Alignment_metrics_out, tmp_dir_path, args.adapter_seq)
  278. mark_dups_bam_out = os.path.join(out_dirs["linear_tin"], f"{args.sample}.Aligned.sortedByCoord.out.md.bam")
  279. mark_dups_txt_out = os.path.join(out_dirs["fastqc"],f"{args.sample}.marked_dup_metrics.txt")
  280. picard_mark_dups(sorted_bam_out, mark_dups_bam_out, mark_dups_txt_out, tmp_dir_path )
  281. # 7. quantify with salmon
  282. salmon_out = os.path.join(out_dirs["linear_tin"],f'{args.sample}_salmon')
  283. salmon_quant(aligned_transcript_bam_out, salmon_out, args.annotation, args.transcripts)
  284. # 8. TIN
  285. # first index md bam
  286. indexed_md_bam_out = index(mark_dups_bam_out)
  287. #then run tin.py
  288. calc_tin(mark_dups_bam_out, args.annote_bed)
  289. # move summary tin ouput
  290. tin_summary_current = f"{args.sample}.Aligned.sortedByCoord.out.md.summary.txt"
  291. tin_summary_new_loc = os.path.join(out_dirs["linear_tin"], f"{args.sample}.Aligned.sortedByCoord.out.md.summary.txt")
  292. tin_xls_current = f"{args.sample}.Aligned.sortedByCoord.out.md.tin.xls"
  293. tin_xls_new_loc = os.path.join(out_dirs["linear_tin"], f"{args.sample}.Aligned.sortedByCoord.out.md.tin.xls")
  294. shutil.move(tin_summary_current, tin_summary_new_loc)
  295. shutil.move(tin_xls_current, tin_xls_new_loc)
  296. # cram bams
  297. cram_bams(args.cram, mark_dups_bam_out, args.ref, args.sample, out_dirs["linear_tin"])
  298. ## 9. Clean up
  299. # delete the intermediate files we don't normally keep -- as requested by user with --delete argument
  300. delete_extras(args.delete, args.sample, aligned_bam_out, sorted_bam_out, indexed_bam_out, out_dirs["linear_tin"], \
  301. indexed_md_bam_out, merged_ubam_out, args.input_to_merge, aligned_transcript_bam_out, mark_dups_bam_out)

linear_RNA_pipeline.py at commit c6fc549, no license · at the source

Overview

Authors: Bridget Phillips1,2, Jessie Sanford1,2, Vaibhav A Janve3,4, Menghan Liu1,2, Matt Johnson1,2, Katherine Gong1,2, Kristy Bergmann1,2, Joseph Lowery1,2, Allison Flynn1,2, William Brock1,2, Brenda Sanchez Montejo1,2, Nicholas Sykora1,2, John Budde1,2, Nikolaos Mellios5, Grigorios Papageorgiou5, Reisa A Sperling6, Rachel F Buckley6,7,8, Mabel Seto7, John C Morris9,10,11, Joel S Perlmutter9,12, Paul T Kotzbauer9, Richard J Perrin9,10,11, Timothy J Hohman3,4, Laura Ibanez1,2,9, Carlos Cruchaga1,2
  1. Department of Psychiatry, Washington University School of Medicine, St. Louis, MO USA
  2. NeuroGenomics and Informatics Center, Washington University School of Medicine, St. Louis, MO USA
  3. Vanderbilt Memory and Alzheimer’s Center, Vanderbilt University Medical Center, Nashville, TN USA
  4. Department of Neurology, Vanderbilt University Medical Center, Nashville, TN USA
  5. Circular Genomics, San Diego, CA USA
  6. Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA USA
  7. Center for Alzheimer’s Research and Treatment, Brigham and Women’s Hospital, Harvard Medical School, Boston, MA USA
  8. School of Psychological Sciences, University of Melbourne, Melbourne, Victoria Australia
  9. Department of Neurology, Washington University School of Medicine, St. Louis, MO USA
  10. Department of Pathology and Immunology, Washington University School of Medicine, St. Louis, MO USA
  11. Knight Alzheimer Disease Research Center, Washington University School of Medicine, St. Louis, MO USA
  12. Department of Radiology, Washington University School of Medicine, St. Louis, MO USA
Institutions: Washington University in St. Louis (United States); Vanderbilt University Medical Center (United States); Harvard University (United States); Massachusetts General Hospital (United States); Brigham and Women's Hospital (United States); The University of Melbourne (Australia)
Journal: Nature medicine, volume 32, issue 8, pages 2857-2864
Dates: received 3 November 2025; accepted 27 May 2026; published online 1 July 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41591-026-04485-5 · PMID 42387213 · PMCID PMC13472853 · OpenAlex W7166805726
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Alzheimer's / dementia (population), clinical / translational (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, fMRI & imaging
Keywords: Alzheimer's disease, RNA sequencing, Predictive markers, Machine learning
MeSH: Alzheimer Disease*, RNA, Circular*, Aged, Aged, 80 and over, Amyloid beta-Peptides, Biomarkers, Early Diagnosis, Female, Humans, Male, tau Proteins (* major topic)
Topic: Circular RNAs in diseases (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NIA NIH HHS (R01 AG044546); U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging) (R01AG044546); U.S. Department of Health &amp; Human Services | NIH | National Institute on Aging (R01AG044546)
Citations: not cited yet (Europe PMC); 41 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repository

Its files are read in the Code ↔ Paper reader above, with 2 matches between paragraphs and lines of code.

neurogenomicsandinformatics/rnaseq_pipeline

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: c6fc54934abcbc152f2df339a29fc07126cae036, 18 March 2026
Languages: Python (10)
Size: 16 files, 10 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: environment (circ_quant_DCC/Dockerfile, linear_pipeline_circ_alignment/Dockerfile.dockerfile, linear_pipeline_circ_alignment/src/gffread.dockerfile)
Not found: README, license file, CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (1 file), pysam (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
10 files

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41591-026-04485-5.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 10 scripts, each with its path and the digest of its content;
  • 2 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

  • no repository, dataset or request procedure was recognized in it

Read it in the paper: doi.org/10.1038/s41591-026-04485-5.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 2, 28 September 2026

  • Publisher: n/a → Nature Portfolio

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 25 authors, 4 keywords, 11 MeSH terms, 3 funders, 39 references.

Cite

This paper

Phillips, B., Sanford, J., Janve, V. A., Liu, M., Johnson, M., Gong, K., Bergmann, K., Lowery, J., Flynn, A., Brock, W., Montejo, B. S., Sykora, N., Budde, J., Mellios, N., Papageorgiou, G., Sperling, R. A., Buckley, R. F., Seto, M., Morris, J. C., . . . Cruchaga, C. (2026). Blood-based circular RNAs for early diagnosis of Alzheimer's disease. Nature medicine, 32(8), 2857-2864. https://doi.org/10.1038/s41591-026-04485-5

BibTeX

@article{phillips2026blood,
author = {Phillips, Bridget and Sanford, Jessie and Janve, Vaibhav A and Liu, Menghan and Johnson, Matt and Gong, Katherine and Bergmann, Kristy and Lowery, Joseph and Flynn, Allison and Brock, William and Montejo, Brenda Sanchez and Sykora, Nicholas and Budde, John and Mellios, Nikolaos and Papageorgiou, Grigorios and Sperling, Reisa A and Buckley, Rachel F and Seto, Mabel and Morris, John C and Perlmutter, Joel S and Kotzbauer, Paul T and Perrin, Richard J and Hohman, Timothy J and Ibanez, Laura and Cruchaga, Carlos},
title = {{Blood-based circular RNAs for early diagnosis of Alzheimer's disease}},
journal = {Nature medicine},
year = {2026},
month = jul,
volume = {32},
number = {8},
pages = {2857--2864},
publisher = {Nature Portfolio},
issn = {1078-8956},
doi = {10.1038/s41591-026-04485-5},
url = {https://doi.org/10.1038/s41591-026-04485-5},
pmid = {42387213},
pmcid = {PMC13472853}
}

RIS

TY - JOUR
AU - Phillips, Bridget
AU - Sanford, Jessie
AU - Janve, Vaibhav A
AU - Liu, Menghan
AU - Johnson, Matt
AU - Gong, Katherine
AU - Bergmann, Kristy
AU - Lowery, Joseph
AU - Flynn, Allison
AU - Brock, William
AU - Montejo, Brenda Sanchez
AU - Sykora, Nicholas
AU - Budde, John
AU - Mellios, Nikolaos
AU - Papageorgiou, Grigorios
AU - Sperling, Reisa A
AU - Buckley, Rachel F
AU - Seto, Mabel
AU - Morris, John C
AU - Perlmutter, Joel S
AU - Kotzbauer, Paul T
AU - Perrin, Richard J
AU - Hohman, Timothy J
AU - Ibanez, Laura
AU - Cruchaga, Carlos
TI - Blood-based circular RNAs for early diagnosis of Alzheimer's disease
T2 - Nature medicine
J2 - Nat Med
PY - 2026
DA - 2026/07/01
VL - 32
IS - 8
SP - 2857
EP - 2864
SN - 1078-8956
PB - Nature Portfolio
DO - 10.1038/s41591-026-04485-5
UR - https://doi.org/10.1038/s41591-026-04485-5
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41591-026-04485-5",
"type": "article-journal",
"title": "Blood-based circular RNAs for early diagnosis of Alzheimer's disease",
"container-title": "Nature medicine",
"author": [
{
"family": "Phillips",
"given": "Bridget"
},
{
"family": "Sanford",
"given": "Jessie"
},
{
"family": "Janve",
"given": "Vaibhav A"
},
{
"family": "Liu",
"given": "Menghan"
},
{
"family": "Johnson",
"given": "Matt"
},
{
"family": "Gong",
"given": "Katherine"
},
{
"family": "Bergmann",
"given": "Kristy"
},
{
"family": "Lowery",
"given": "Joseph"
},
{
"family": "Flynn",
"given": "Allison"
},
{
"family": "Brock",
"given": "William"
},
{
"family": "Montejo",
"given": "Brenda Sanchez"
},
{
"family": "Sykora",
"given": "Nicholas"
},
{
"family": "Budde",
"given": "John"
},
{
"family": "Mellios",
"given": "Nikolaos"
},
{
"family": "Papageorgiou",
"given": "Grigorios"
},
{
"family": "Sperling",
"given": "Reisa A"
},
{
"family": "Buckley",
"given": "Rachel F"
},
{
"family": "Seto",
"given": "Mabel"
},
{
"family": "Morris",
"given": "John C"
},
{
"family": "Perlmutter",
"given": "Joel S"
},
{
"family": "Kotzbauer",
"given": "Paul T"
},
{
"family": "Perrin",
"given": "Richard J"
},
{
"family": "Hohman",
"given": "Timothy J"
},
{
"family": "Ibanez",
"given": "Laura"
},
{
"family": "Cruchaga",
"given": "Carlos"
}
],
"container-title-short": "Nat Med",
"volume": "32",
"issue": "8",
"page": "2857-2864",
"DOI": "10.1038/s41591-026-04485-5",
"PMID": "42387213",
"PMCID": "PMC13472853",
"ISSN": "1078-8956",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41591-026-04485-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41467-026-71682-8 [code]
GWAS meta-analysis of cerebrospinal fluid Alzheimer's biomarkers reveals loci regulating lipids, brain volume and autophagy.
Journal: Nature communications
In common: NumPy, Alzheimer's / dementia, genetics / omics, 1 reference, 4 authors
[2] doi:10.1038/s41588-026-02722-8 [code]
A multiancestry polygenic risk score for Alzheimer's disease is associated with cognitive decline and neuropathological hallmarks in diverse populations.
Journal: Nature genetics
In common: Alzheimer's / dementia, genetics / omics, 3 references, 2 authors
[3] doi:10.64898/2026.05.06.26352540 [code]
Generating synthetic tau-PET scans in Alzheimer’s disease from MRI, blood biomarkers and demographics with deep learning
Journal: medRxiv (preprint)
In common: NumPy, Alzheimer's / dementia, clinical / translational, 4 references
[4] doi:10.1038/s41467-026-74515-w [code]
Advancing fair and explainable machine learning for neuroimaging dementia pattern classification in multi-racial and multi-ethnic populations.
Journal: Nature communications
In common: NumPy, Alzheimer's / dementia, 1 reference, author Timothy Hohman
[5] doi:10.1093/brain/awaf375 [code]
Reference proteins to improve Core 1 and Core 2 Alzheimer's disease CSF and plasma biomarkers.
Journal: Brain : a journal of neurology
In common: NumPy, Alzheimer's / dementia, clinical / translational, 4 references
[6] doi:10.1186/s13195-026-02119-z
Age-dependent diagnostic and correlational architecture of multiplex plasma biomarkers in Alzheimer's disease: a cross-ethnic, cross-platform validation study.
Journal: Alzheimer's research & therapy
In common: Alzheimer's / dementia, clinical / translational, 4 references
[7] doi:10.1093/braincomms/fcag074 [code]
Plasma p-tau217 and glucose metabolism correlate in neocortical association areas in Alzheimer's disease.
Journal: Brain communications
In common: Alzheimer's / dementia, 4 references
[8] doi:10.1038/s44400-026-00094-8 [code]
Haplotype-resolved DNA methylation at the &lt;i&gt;APOE&lt;/i&gt; locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk.
Journal: NPJ dementia
In common: pysam, NumPy, Alzheimer's / dementia, genetics / omics, 2 references
[9] doi:10.3389/fneur.2026.1822479 [code]
Circulating neuron-derived cfDNA for blood-based detection of Alzheimer's and other neurodegenerative conditions.
Journal: Frontiers in neurology
In common: pysam, NumPy, Alzheimer's / dementia, clinical / translational, genetics / omics, 1 reference
[10] doi:10.1038/s41586-026-10454-2 [code]
White matter micro- and macrostructure brain charts for the human lifespan.
Journal: Nature
In common: NumPy, author Timothy Hohman

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.