OSCR

Population-scale repeat expansions elucidate disease risk and brain atrophy.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches
  1. [1] § Methods › Repeat expansion genotyping and QC ↔ src/main_gangstr.cpp, lines 43–106 · score 0.68 · GangSTR, reference genome, nonuniform, enclose, probes, score
  2. [2] § Methods › Genetic relatedness analysis ↔ 1.9/plink_calc.h, lines 61–91 · score 0.60 · pi hat, allele frequency, unrelated, exp, cutoffs, PLINK
  3. [3] § Methods › Genetic relatedness analysis ↔ 2.0/plink2.cc, lines 3041–3100 · score 0.57 · minor allele frequency, IBD, unrelated, components, cutoffs, ancestral
  4. [4] § Disease risk increases with repeat length ↔ src/Step2_Models.cpp, lines 1160–1254 · score 0.57 · Firth logistic regression, carrier status, 0.01 %, models, covariates, thresholds
  5. [5] § Methods › Repeat expansion genotyping and QC ↔ src/likelihood_maximizer.h, lines 55–171 · score 0.55 · confidence interval, enclose, likelihood, score, locus, expansion
  6. [6] § Methods › Ancestry assignment ↔ 2.0/plink2_filter.cc, lines 4389–4475 · score 0.51 · Hardy Weinberg equilibrium, allele frequency, missingness

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

C++ · 463 lines · 18 KB · GPL-3.0 · 1 match

  1. /*
  2. Copyright (C) 2017 Melissa Gymrek <[email hidden]>
  3. and Nima Mousavi ([email hidden])
  4. This file is part of GangSTR.
  5. GangSTR is free software: you can redistribute it and/or modify
  6. it under the terms of the GNU General Public License as published by
  7. the Free Software Foundation, either version 3 of the License, or
  8. (at your option) any later version.
  9. GangSTR is distributed in the hope that it will be useful,
  10. but WITHOUT ANY WARRANTY; without even the implied warranty of
  11. MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
  12. GNU General Public License for more details.
  13. You should have received a copy of the GNU General Public License
  14. along with GangSTR. If not, see <http://www.gnu.org/licenses/>.
  15. */
  16. #include <getopt.h>
  17. #include <stdlib.h>
  18. #include <iostream>
  19. #include <set>
  20. #include <sstream>
  21. //#include "src/bam_reader.h"
  22. #include "src/bam_info_extract.h"
  23. #include "src/bam_io.h"
  24. #include "src/common.h"
  25. #include "src/genotyper.h"
  26. #include "src/options.h"
  27. #include "src/ref_genome.h"
  28. #include "src/region_reader.h"
  29. #include "src/sample_info.h"
  30. #include "src/str_info.h"
  31. #include "src/stringops.h"
  32. #include "src/vcf_writer.h"
  33. #include "GangSTRConfig.h"
  34. using namespace std;
  35. void show_help() {
  36. Options options;
  37. std::stringstream help_msg;
  38. help_msg << "\nUsage: GangSTR [OPTIONS] "
  39. << "--bam <file1[,file2,...]> "
  40. << "--ref <reference.fa> "
  41. << "--regions <regions.bed> "
  42. << "--out <outprefix> "
  43. << "\n\n Required options:\n"
  44. << "\t" << "--bam <file.bam,[file2.bam]>" << "\t" << "Comma separated list of input BAM files" << "\n"
  45. << "\t" << "--ref <genome.fa> " << "\t" << "FASTA file for the reference genome" << "\n"
  46. << "\t" << "--regions <regions.bed> " << "\t" << "BED file containing TR coordinates" << "\n"
  47. << "\t" << "--out <outprefix> " << "\t" << "Prefix to name output files" << "\n"
  48. << "\n Additional general options:\n"
  49. << "\t" << "--targeted " << "\t" << "Targeted mode" << "\n"
  50. << "\t" << "--chrom " << "\t" << "Only genotype regions on this chromosome" << "\n"
  51. << "\t" << "--bam-samps <string> " << "\t" << "Comma separated list of sample IDs for --bam" << "\n"
  52. << "\t" << "--samp-sex <string> " << "\t" << "Comma separated list of sample sex for each sample ID (--bam-samps must be provided)" << "\n"
  53. << "\t" << "--str-info <string> " << "\t" << "Tab file with additional per-STR info (see docs)" << "\n"
  54. << "\t" << "--period <string> " << "\t" << "Only genotype loci with periods (motif lengths) in this comma-separated list." << "\n"
  55. << "\t" << "--skip-qscore " << "\t" << "Skip calculation of Q-score" << "\n"
  56. << "\n Options for different sequencing settings\n"
  57. << "\t" << "--readlength <int> " << "\t" << "Read length. Default: " << options.read_len << "\n"
  58. << "\t" << "--coverage <float> " << "\t" << "Average coverage. must be set for exome/targeted data. Comma separated list to specify for each BAM" << "\n"
  59. << "\t" << "--model-gc-coverage " << "\t" << "Model coverage as a function of GC content. Requires genome-wide data" << "\n"
  60. << "\t" << "--insertmean <float> " << "\t" << "Fragment length mean. Comma separated list to specify for each BAM separately." << "\n"
  61. << "\t" << "--insertsdev <float> " << "\t" << "Fragment length standard deviation. Comma separated list to specify for each BAM separately. " << "\n"
  62. << "\t" << "--nonuniform " << "\t" << "Indicate whether data has non-uniform coverage (i.e., exome)" << "\n"
  63. << "\t" << "--min-sample-reads <int> " << "\t" << "Minimum number of reads per sample." << "\n"
  64. << "\n Advanced paramters for likelihood model:\n"
  65. << "\t" << "--frrweight <float> " << "\t" << "Weight for FRR reads. Default: " << options.frr_weight << "\n"
  66. << "\t" << "--enclweight <float> " << "\t" << "Weight for enclosing reads. Default: " << options.enclosing_weight << "\n"
  67. << "\t" << "--spanweight <float> " << "\t" << "Weight for spanning reads. Default: " << options.spanning_weight << "\n"
  68. << "\t" << "--flankweight <float> " << "\t" << "Weight for flanking reads. Default: " << options.flanking_weight << "\n"
  69. << "\t" << "--ploidy <int> " << "\t" << "Indicate whether data is haploid (1) or diploid (2). Default: " << options.ploidy << "\n"
  70. << "\t" << "--skipofftarget " << "\t" << "Skip off target regions included in the BED file." << "\n"
  71. << "\t" << "--read-prob-mode " << "\t" << "Use only read probability (ignore class probability)" << "\n"
  72. << "\t" << "--numbstrap <int> " << "\t" << "Number of bootstrap samples. Default: " << options.num_boot_samp << "\n"
  73. << "\t" << "--grid-threshold <int> " << "\t" << "Use optimization rather than grid search to find MLE if more than this many possible alleles. Default: " << options.grid_threshold << "\n"
  74. << "\t" << "--rescue-count <int> " << "\t" << "Number of regions that GangSTR attempts to rescue mates from (excluding off-target regions) Default: " << options.rescue_count << "\n"
  75. << "\t" << "--max-proc-read <int> " << "\t" << "Maximum number of processed reads per sample before a region is skipped. Default: " << options.max_processed_reads_per_sample << "\n"
  76. << "\n Parameters for local realignment:\n"
  77. << "\t" << "--minscore <int> " << "\t" << "Minimum alignment score (out of 100). Default: " << options.min_score << "\n"
  78. << "\t" << "--minmatch <int> " << "\t" << "Minimum number of matching basepairs on each end of enclosing reads. Default: " << options.min_match<< "\n"
  79. << "\n Default stutter model parameters:\n"
  80. << "\t" << "--stutterup <float> " << "\t" << "Stutter insertion probability. Default: " << options.stutter_up << "\n"
  81. << "\t" << "--stutterdown <float> " << "\t" << "Stutter deletion probability. Default: " << options.stutter_down << "\n"
  82. << "\t" << "--stutterprob <float> " << "\t" << "Stutter step size parameter. Default: " << options.stutter_p << "\n"
  83. << "\n Parameters for more detailed info about each locus:\n"
  84. << "\t" << "--output-bootstraps " << "\t" << "Output file with bootstrap samples" << "\n"
  85. << "\t" << "--output-readinfo " << "\t" << "Output read class info (for debugging)" << "\n"
  86. << "\t" << "--include-ggl " << "\t" << "Output GGL (special GL field) in VCF" << "\n"
  87. << "\n Additional optional paramters:\n"
  88. << "\t" << "-h,--help " << "\t" << "display this help screen" << "\n"
  89. << "\t" << "--seed " << "\t" << "Random number generator initial seed" << "\n"
  90. << "\t" << "-v,--verbose " << "\t" << "Print out useful progress messages" << "\n"
  91. << "\t" << "--very " << "\t" << "Print out more detailed progress messages for debugging" << "\n"
  92. << "\t" << "--quiet " << "\t" << "Don't print anything" << "\n"
  93. << "\t" << "--version " << "\t" << "Print out the version of this software.\n"
  94. << "\n\nThis program takes in aligned reads in BAM format\n"
  95. << "and outputs estimated genotypes at each TR in VCF format.\n\n";
  96. cerr << help_msg.str();
  97. exit(1);
  98. }
  99. void parse_commandline_options(int argc, char* argv[], Options* options) {
  100. enum LONG_OPTIONS {
  101. OPT_MAXPROCREAD,
  102. OPT_SKIPQ,
  103. OPT_MINREAD,
  104. OPT_PERIOD,
  105. OPT_GGL,
  106. OPT_GRIDTHRESH,
  107. OPT_RESCUE,
  108. OPT_BAMFILES,
  109. OPT_BAMSAMP,
  110. OPT_SAMPSEX,
  111. OPT_STRINFO,
  112. OPT_CHROM,
  113. OPT_REFFA,
  114. OPT_REGIONS,
  115. OPT_OUT,
  116. OPT_HELP,
  117. OPT_WFRR,
  118. OPT_WENCLOSE,
  119. OPT_WSPAN,
  120. OPT_WFLANK,
  121. OPT_TARGETED,
  122. OPT_PLOIDY,
  123. OPT_READLEN,
  124. OPT_COVERAGE,
  125. OPT_GCCOV,
  126. OPT_SKIPOFF,
  127. OPT_NONUNIF,
  128. OPT_INSMEAN,
  129. OPT_INSSDEV,
  130. OPT_MINSCORE,
  131. OPT_MINMATCH,
  132. OPT_STUTUP,
  133. OPT_STUTDW,
  134. OPT_STUTPR,
  135. OPT_NBSTRAP,
  136. OPT_RDPROB,
  137. OPT_OUTBS,
  138. OPT_OUTREADINFO,
  139. OPT_SEED,
  140. OPT_VERBOSE,
  141. OPT_VERYVERBOSE,
  142. OPT_QUIET,
  143. OPT_VERSION,
  144. };
  145. static struct option long_options[] = {
  146. {"max-proc-read", required_argument, NULL, OPT_MAXPROCREAD},
  147. {"skip-qscore", no_argument, NULL, OPT_SKIPQ},
  148. {"min-sample-reads", required_argument, NULL, OPT_MINREAD},
  149. {"period", required_argument, NULL, OPT_PERIOD},
  150. {"include-ggl", no_argument, NULL, OPT_GGL},
  151. {"grid-threshold", required_argument, NULL, OPT_GRIDTHRESH},
  152. {"rescue-count", required_argument, NULL, OPT_RESCUE},
  153. {"bam", required_argument, NULL, OPT_BAMFILES},
  154. {"bam-samps", required_argument, NULL, OPT_BAMSAMP},
  155. {"samp-sex", required_argument, NULL, OPT_SAMPSEX},
  156. {"str-info", required_argument, NULL, OPT_STRINFO},
  157. {"chrom", required_argument, NULL, OPT_CHROM},
  158. {"ref", required_argument, NULL, OPT_REFFA},
  159. {"regions", required_argument, NULL, OPT_REGIONS},
  160. {"out", required_argument, NULL, OPT_OUT},
  161. {"help", no_argument, NULL, OPT_HELP},
  162. {"frrweight", required_argument, NULL, OPT_WFRR},
  163. {"enclweight", required_argument, NULL, OPT_WENCLOSE},
  164. {"spanweight", required_argument, NULL, OPT_WSPAN},
  165. {"flankweight", required_argument, NULL, OPT_WFLANK},
  166. {"targeted", no_argument, NULL, OPT_TARGETED},
  167. {"ploidy", required_argument, NULL, OPT_PLOIDY},
  168. {"readlength", required_argument, NULL, OPT_READLEN},
  169. {"coverage", required_argument, NULL, OPT_COVERAGE},
  170. {"model-gc-coverage", no_argument, NULL, OPT_GCCOV},
  171. {"nonuniform", no_argument, NULL, OPT_NONUNIF},
  172. {"skipofftarget",no_argument, NULL, OPT_SKIPOFF},
  173. {"insertmean", required_argument, NULL, OPT_INSMEAN},
  174. {"insertsdev", required_argument, NULL, OPT_INSSDEV},
  175. {"minscore", required_argument, NULL, OPT_MINSCORE},
  176. {"minmatch", required_argument, NULL, OPT_MINMATCH},
  177. {"stutterup", required_argument, NULL, OPT_STUTUP},
  178. {"stutterdown", required_argument, NULL, OPT_STUTDW},
  179. {"stutterprob", required_argument, NULL, OPT_STUTPR},
  180. {"numbstrap", required_argument, NULL, OPT_NBSTRAP},
  181. {"read-prob-mode", no_argument, NULL, OPT_RDPROB},
  182. {"output-bootstraps", no_argument, NULL, OPT_OUTBS},
  183. {"output-readinfo", no_argument, NULL, OPT_OUTREADINFO},
  184. {"seed", required_argument, NULL, OPT_SEED},
  185. {"verbose", no_argument, NULL, OPT_VERBOSE},
  186. {"very", no_argument, NULL, OPT_VERYVERBOSE},
  187. {"quiet", no_argument, NULL, OPT_QUIET},
  188. {"version", no_argument, NULL, OPT_VERSION},
  189. {NULL, no_argument, NULL, 0},
  190. };
  191. std::vector<std::string> dist_means_str, dist_sdev_str, coverage_str, pers;
  192. int ch;
  193. int option_index = 0;
  194. ch = getopt_long(argc, argv, "hv?",
  195. long_options, &option_index);
  196. while (ch != -1) {
  197. switch (ch) {
  198. case OPT_MAXPROCREAD:
  199. options->max_processed_reads_per_sample = atoi(optarg);
  200. break;
  201. case OPT_SKIPQ:
  202. options->skip_qscore = true;
  203. break;
  204. case OPT_MINREAD:
  205. options->min_reads_per_sample = atoi(optarg);
  206. break;
  207. case OPT_PERIOD:
  208. split_by_delim(optarg, ',', pers);
  209. options->period.clear();
  210. for (size_t i=0; i<pers.size(); i++) {
  211. options->period.push_back(atoi(pers[i].c_str()));
  212. }
  213. break;
  214. case OPT_GGL:
  215. options->include_ggl = true;
  216. break;
  217. case OPT_GRIDTHRESH:
  218. options->grid_threshold = atoi(optarg);
  219. break;
  220. case OPT_RESCUE:
  221. options->rescue_count = atoi(optarg);
  222. break;
  223. case OPT_BAMFILES:
  224. options->bamfiles.clear();
  225. split_by_delim(optarg, ',', options->bamfiles);
  226. break;
  227. case OPT_BAMSAMP:
  228. options->rg_sample_string = optarg;
  229. break;
  230. case OPT_SAMPSEX:
  231. options->sample_sex_string = optarg;
  232. break;
  233. case OPT_STRINFO:
  234. options->str_info_file = optarg;
  235. break;
  236. case OPT_CHROM:
  237. options->chrom = optarg;
  238. break;
  239. case OPT_REFFA:
  240. options->reffa = optarg;
  241. break;
  242. case OPT_REGIONS:
  243. options->regionsfile = optarg;
  244. break;
  245. case OPT_OUT:
  246. options->outprefix = optarg;
  247. break;
  248. case OPT_HELP:
  249. case 'h':
  250. show_help();
  251. case OPT_WFRR:
  252. options->frr_weight = atof(optarg);
  253. break;
  254. case OPT_WENCLOSE:
  255. options->enclosing_weight = atof(optarg);
  256. break;
  257. case OPT_WSPAN:
  258. options->spanning_weight = atof(optarg);
  259. break;
  260. case OPT_WFLANK:
  261. options->flanking_weight = atof(optarg);
  262. break;
  263. case OPT_TARGETED:
  264. options->genome_wide = false;
  265. break;
  266. case OPT_PLOIDY:
  267. options->ploidy = atoi(optarg);
  268. break;
  269. case OPT_READLEN:
  270. options->read_len = atoi(optarg);
  271. break;
  272. case OPT_COVERAGE:
  273. split_by_delim(optarg, ',', coverage_str);
  274. options->coverage.clear();
  275. for (size_t i=0; i<coverage_str.size(); i++) {
  276. options->coverage.push_back(atoi(coverage_str[i].c_str()));
  277. }
  278. break;
  279. case OPT_GCCOV:
  280. options->model_gc_cov = true;
  281. break;
  282. case OPT_NONUNIF:
  283. options->use_cov = false;
  284. options->use_off = false;
  285. break;
  286. case OPT_SKIPOFF:
  287. options->use_off = false;
  288. break;
  289. case OPT_INSMEAN:
  290. split_by_delim(optarg, ',', dist_means_str);
  291. options->dist_mean.clear();
  292. for (size_t i=0; i<dist_means_str.size(); i++) {
  293. options->dist_mean.push_back(strtof(dist_means_str[i].c_str(), NULL));
  294. }
  295. options->dist_man_set = true;
  296. break;
  297. case OPT_INSSDEV:
  298. split_by_delim(optarg, ',', dist_sdev_str);
  299. options->dist_sdev.clear();
  300. for (size_t i=0; i<dist_sdev_str.size(); i++) {
  301. options->dist_sdev.push_back(strtof(dist_sdev_str[i].c_str(), NULL));
  302. }
  303. options->dist_man_set = true;
  304. break;
  305. case OPT_MINSCORE:
  306. options->min_score = atoi(optarg);
  307. break;
  308. case OPT_MINMATCH:
  309. options->min_match = atoi(optarg);
  310. break;
  311. case OPT_STUTUP:
  312. options->stutter_up = atof(optarg);
  313. break;
  314. case OPT_STUTDW:
  315. options->stutter_down = atof(optarg);
  316. break;
  317. case OPT_STUTPR:
  318. options->stutter_p = atof(optarg);
  319. break;
  320. case OPT_NBSTRAP:
  321. options->num_boot_samp = atoi(optarg);
  322. break;
  323. case OPT_OUTBS:
  324. options->output_bootstrap = true;
  325. break;
  326. case OPT_RDPROB:
  327. options->read_prob_mode = true;
  328. break;
  329. case OPT_OUTREADINFO:
  330. options->output_readinfo= true;
  331. break;
  332. case OPT_SEED:
  333. options->seed = atoi(optarg);
  334. break;
  335. case OPT_VERBOSE:
  336. case 'v':
  337. options->verbose = true;
  338. break;
  339. case OPT_VERYVERBOSE:
  340. options->very_verbose = true;
  341. break;
  342. case OPT_QUIET:
  343. options->quiet = true;
  344. break;
  345. case OPT_VERSION:
  346. cerr << GangSTR_VER << endl;
  347. exit(0);
  348. case '?':
  349. show_help();
  350. default:
  351. show_help();
  352. };
  353. ch = getopt_long(argc, argv, "hv?",
  354. long_options, &option_index);
  355. };
  356. // Leftover arguments are errors
  357. if (optind < argc) {
  358. PrintMessageDieOnError("Unnecessary leftover arguments", M_ERROR, false);
  359. }
  360. // Perform other error checking
  361. if (options->bamfiles.empty()) {
  362. PrintMessageDieOnError("No --bam files specified", M_ERROR, false);
  363. }
  364. if (options->regionsfile.empty()) {
  365. PrintMessageDieOnError("No --regions option specified", M_ERROR, false);
  366. }
  367. if (options->reffa.empty()) {
  368. PrintMessageDieOnError("No --ref option specified", M_ERROR, false);
  369. }
  370. if (options->outprefix.empty()) {
  371. PrintMessageDieOnError("No --out option specified", M_ERROR, false);
  372. }
  373. if (options->min_match < 0 or (options->read_len != -1 and options->min_match > options->read_len)){
  374. PrintMessageDieOnError("--minmatch parameter must be in (0, read_len) range", M_ERROR, false);
  375. }
  376. if (options->min_score < 0 and options->min_score > 100){
  377. PrintMessageDieOnError("--min_score parameter must be in (0, 100) range", M_ERROR, false);
  378. }
  379. // Check if bam-samps provided when user uses sam-sex
  380. if (options->rg_sample_string.empty() and !options->sample_sex_string.empty()){
  381. PrintMessageDieOnError("--sample-sex requires input argument --bam-samps to be set", M_ERROR, false);
  382. }
  383. }
  384. int main(int argc, char* argv[]) {
  385. // Set up
  386. Options options;
  387. parse_commandline_options(argc, argv, &options);
  388. stringstream full_command_ss;
  389. full_command_ss << "GangSTR-" << GangSTR_VER;
  390. for (int i = 1; i < argc; i++) {
  391. full_command_ss << " " << argv[i];
  392. }
  393. std::string full_command = full_command_ss.str();
  394. RegionReader region_reader(options.regionsfile);
  395. Locus locus;
  396. int merge_type = BamCramMultiReader::ORDER_ALNS_BY_FILE;
  397. BamCramMultiReader bamreader(options.bamfiles, options.reffa, merge_type);
  398. // Extract sample info
  399. SampleInfo sample_info;
  400. if (!options.rg_sample_string.empty()) {
  401. if (!sample_info.SetCustomReadGroups(options)) {
  402. PrintMessageDieOnError("Error setting custom read groups", M_ERROR, false);
  403. }
  404. } else {
  405. if (!sample_info.LoadReadGroups(options, bamreader)) {
  406. PrintMessageDieOnError("Error loading read groups", M_ERROR, false);
  407. }
  408. }
  409. // Extract information from bam file (read length, insert size distribution, ..)
  410. RefGenome refgenome(options.reffa);
  411. if (!sample_info.ExtractBamInfo(options, bamreader, region_reader, refgenome)) {
  412. PrintMessageDieOnError("Error extracting info from BAM file", M_ERROR, false);
  413. }
  414. options.realignment_flanklen = sample_info.GetReadLength();
  415. // Write out values found for each sample
  416. sample_info.PrintSampleInfo(options.outprefix + ".samplestats.tab");
  417. // Process each region
  418. region_reader.Reset();
  419. STRInfo str_info(options);
  420. VCFWriter vcfwriter(options.outprefix + ".vcf", full_command,
  421. refgenome,
  422. sample_info, options.include_ggl);
  423. Genotyper genotyper(refgenome, options, sample_info, str_info);
  424. stringstream ss;
  425. while (region_reader.GetNextRegion(&locus)) {
  426. if (!options.chrom.empty() && locus.chrom != options.chrom) {continue;}
  427. if (!options.period.empty() &&
  428. std::find(options.period.begin(), options.period.end(), locus.period) == options.period.end()) {
  429. continue;
  430. }
  431. ss.str("");
  432. ss.clear();
  433. ss << "Processing " << locus.chrom << ":" << locus.start;
  434. PrintMessageDieOnError(ss.str(), M_PROGRESS, options.quiet);
  435. if (options.use_off){
  436. locus.offtarget_share = 1.0;
  437. }
  438. else{
  439. locus.offtarget_share = 0.0;
  440. }
  441. if (genotyper.ProcessLocus(&bamreader, &locus)) {
  442. vcfwriter.WriteRecord(locus);
  443. }
  444. locus.Reset();
  445. };
  446. }

main_gangstr.cpp at commit 6ea9b2b, under GPL-3.0 · at the source

Overview

Authors: Vijay Kumar Pounraja1, Jae Hoon Sul1, Joseph Herman1, Sean O’Keeffe1, Veera Rajagopal1, Xiaodong Bai1, Michael D. Kessler1, Neelroop Parikshak1, Karl Landheer1, Xingmin Zhang1, Sean Yu1, Lance Zhang1, Michelle G. LeBlanc1, Jennifer Rico-Varela1, Frederic Grau2, Sarah Wolf1, Sriramkumar Sundaramoorthy2, Farshid Sepehrband2, Eli A. Stahl1, Yuda Huo2
and 15 other authorsMohsin Ahmed2, Susan Croll2, GHS-RGC DiscovEHR Collaboration, Mayo-RGC Project Generation, Penn Medicine BioBank, William Salerno1, John D. Overton1, Jonathan Marchini1, Jeffrey Reid1, Luca A. Lotta1, Aris Baras1, Regeneron Genetics Center, Goncalo R. Abecasis1, Giovanni Coppola1, Sahar Gelfman1
  1. Regeneron Genetics Center,Tarrytown, NY USA
  2. Regeneron Pharmaceuticals,Tarrytown, NY USA
Institutions: Regeneron (United States) (United States)
Journal: Nature, volume 653, issue 8115, pages 796-808
Dates: received 16 May 2025; accepted 2 March 2026; published online 8 April 2026; in print 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s41586-026-10345-6 · PMID 41951733 · PMCID PMC13190288 · OpenAlex W7151843168
Open access: hybrid, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), structural MRI / diffusion (modality), human (organism), other condition (population)
Methods: Statistics, Connectivity, Preprocessing
Keywords: DNA sequencing, Microsatellite instability, Neurodegeneration, Structural variation, Genetic predisposition to disease
MeSH: Brain*, DNA Repeat Expansion*, Genetic Predisposition to Disease*, Microsatellite Repeats*, Atrophy, C9orf72 Protein, Cerebellum, Female, Humans, Huntingtin Protein, Huntington Disease, Male, Middle Aged, Myotonin-Protein Kinase, Neurofilament Proteins, Organ Size (* major topic)
Topic: Genetic Neurodegenerative Diseases (Cellular and Molecular Neuroscience, Neuroscience), according to OpenAlex
Funding: NCI NIH HHS (P30 CA015083)
Citations: cited by 3 papers (Europe PMC); 76 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

rgcgithub/regenie

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 7c2b153456e1e4c921789f2c1e26cf3757e6fb4d, 8 September 2026
Languages: C/C++ (358), C++ (31), Shell (8), JavaScript (7), R (6), Python (1), C (1)
Size: 669 files, 412 scripts
Software Heritage: archived
Found in: “Code availability”
Holds: README, license file, environment (Dockerfile), tests, continuous integration, documentation
Not found: CITATION.cff
Tools: data.table (4 files), tidyverse (4 files), cowplot (2 files), ggplot2 (2 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
414 files

gymreklab/GangSTR

License: GPL-3.0
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 6ea9b2b8daca51dcab1f0e46210622b94b52ff17, 19 April 2021
Languages: C++ (31), C/C++ (31), Shell (14), Python (12), C (1)
Size: 184 files, 89 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, tests
Not found: CITATION.cff, environment file, continuous integration, documentation
Tools: BCFtools (1 file), Matplotlib (1 file), NumPy (1 file), pandas (1 file), SAMtools (1 file), SciPy (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
92 files

Illumina/ExpansionHunter

License: MIT
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: dc477e1379ed202bbfa103a75e023107a2b02c36, 30 January 2024
Languages: C++ (156), Shell (13), C/C++ (12), Python (1)
Size: 352 files, 182 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (tools/docker/debug/mallinfo2/Dockerfile, tools/docker/deploy/centos7gcc9src/Dockerfile, tools/docker/test/centos8Build/Dockerfile, tools/docker/test/ubuntu2004Build/Dockerfile), tests, continuous integration, documentation
Not found: CITATION.cff
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
184 files, not copied: shown from their source

OSCR keeps no copy of these files: the license of this repository (MIT) is not one it has verified to allow it. The reader above shows each one from its source, fetched by your browser at commit dc477e1, when its fingerprint is the one OSCR verified. How this works.

chrchang/plink-ng

License: GPL
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 04003c63e54ee1728bc93163e153200ba2902bef, 28 September 2026
Languages: C/C++ (577), C (108), Shell (87), C++ (81), Python (43), R (6), CUDA (1)
Size: 1,228 files, 903 scripts
Software Heritage: archived
Found in: “Code availability”
Holds: README, environment (docker/Dockerfile, 2.0/cindex/DESCRIPTION, 2.0/pgenlibr/DESCRIPTION, 2.0/Python/pyproject.toml, 2.0/Python/setup.py), tests, continuous integration, documentation
Not found: license file, CITATION.cff
Tools: NumPy (14 files), pandas (5 files), Matplotlib (2 files), SciPy (1 file), statsmodels (1 file)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
879 files, not copied: shown from their source

OSCR keeps no copy of these files: the license of this repository (GPL) is not one it has verified to allow it. The reader above shows each one from its source, fetched by your browser at commit 04003c6, when its fingerprint is the one OSCR verified. How this works.

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41586-026-10345-6.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 1,562 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability statement

The paper has a data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s41586-026-10345-6.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 35 authors, 5 keywords, 16 MeSH terms, 1 funder, 70 references.

Cite

This paper

Pounraja, V. K., Sul, J. H., Herman, J., O’Keeffe, S., Rajagopal, V., Bai, X., Kessler, M. D., Parikshak, N., Landheer, K., Zhang, X., Yu, S., Zhang, L., LeBlanc, M. G., Rico-Varela, J., Grau, F., Wolf, S., Sundaramoorthy, S., Sepehrband, F., Stahl, E. A., . . . Gelfman, S. (2026). Population-scale repeat expansions elucidate disease risk and brain atrophy. Nature, 653(8115), 796-808. https://doi.org/10.1038/s41586-026-10345-6

BibTeX

@article{pounraja2026population,
author = {Pounraja, Vijay Kumar and Sul, Jae Hoon and Herman, Joseph and O’Keeffe, Sean and Rajagopal, Veera and Bai, Xiaodong and Kessler, Michael D. and Parikshak, Neelroop and Landheer, Karl and Zhang, Xingmin and Yu, Sean and Zhang, Lance and LeBlanc, Michelle G. and Rico-Varela, Jennifer and Grau, Frederic and Wolf, Sarah and Sundaramoorthy, Sriramkumar and Sepehrband, Farshid and Stahl, Eli A. and Huo, Yuda and Ahmed, Mohsin and Croll, Susan and {GHS-RGC DiscovEHR Collaboration} and {Mayo-RGC Project Generation} and {Penn Medicine BioBank} and Salerno, William and Overton, John D. and Marchini, Jonathan and Reid, Jeffrey and Lotta, Luca A. and Baras, Aris and {Regeneron Genetics Center} and Abecasis, Goncalo R. and Coppola, Giovanni and Gelfman, Sahar},
title = {{Population-scale repeat expansions elucidate disease risk and brain atrophy}},
journal = {Nature},
year = {2026},
month = apr,
volume = {653},
number = {8115},
pages = {796--808},
publisher = {Nature Portfolio},
issn = {0028-0836},
doi = {10.1038/s41586-026-10345-6},
url = {https://doi.org/10.1038/s41586-026-10345-6},
pmid = {41951733},
pmcid = {PMC13190288}
}

RIS

TY - JOUR
AU - Pounraja, Vijay Kumar
AU - Sul, Jae Hoon
AU - Herman, Joseph
AU - O’Keeffe, Sean
AU - Rajagopal, Veera
AU - Bai, Xiaodong
AU - Kessler, Michael D.
AU - Parikshak, Neelroop
AU - Landheer, Karl
AU - Zhang, Xingmin
AU - Yu, Sean
AU - Zhang, Lance
AU - LeBlanc, Michelle G.
AU - Rico-Varela, Jennifer
AU - Grau, Frederic
AU - Wolf, Sarah
AU - Sundaramoorthy, Sriramkumar
AU - Sepehrband, Farshid
AU - Stahl, Eli A.
AU - Huo, Yuda
AU - Ahmed, Mohsin
AU - Croll, Susan
AU - GHS-RGC DiscovEHR Collaboration
AU - Mayo-RGC Project Generation
AU - Penn Medicine BioBank
AU - Salerno, William
AU - Overton, John D.
AU - Marchini, Jonathan
AU - Reid, Jeffrey
AU - Lotta, Luca A.
AU - Baras, Aris
AU - Regeneron Genetics Center
AU - Abecasis, Goncalo R.
AU - Coppola, Giovanni
AU - Gelfman, Sahar
TI - Population-scale repeat expansions elucidate disease risk and brain atrophy
T2 - Nature
J2 - Nature
PY - 2026
DA - 2026/04/08
VL - 653
IS - 8115
SP - 796
EP - 808
SN - 0028-0836
PB - Nature Portfolio
DO - 10.1038/s41586-026-10345-6
UR - https://doi.org/10.1038/s41586-026-10345-6
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41586-026-10345-6",
"type": "article-journal",
"title": "Population-scale repeat expansions elucidate disease risk and brain atrophy",
"container-title": "Nature",
"author": [
{
"family": "Pounraja",
"given": "Vijay Kumar"
},
{
"family": "Sul",
"given": "Jae Hoon"
},
{
"family": "Herman",
"given": "Joseph"
},
{
"family": "O’Keeffe",
"given": "Sean"
},
{
"family": "Rajagopal",
"given": "Veera"
},
{
"family": "Bai",
"given": "Xiaodong"
},
{
"family": "Kessler",
"given": "Michael D."
},
{
"family": "Parikshak",
"given": "Neelroop"
},
{
"family": "Landheer",
"given": "Karl"
},
{
"family": "Zhang",
"given": "Xingmin"
},
{
"family": "Yu",
"given": "Sean"
},
{
"family": "Zhang",
"given": "Lance"
},
{
"family": "LeBlanc",
"given": "Michelle G."
},
{
"family": "Rico-Varela",
"given": "Jennifer"
},
{
"family": "Grau",
"given": "Frederic"
},
{
"family": "Wolf",
"given": "Sarah"
},
{
"family": "Sundaramoorthy",
"given": "Sriramkumar"
},
{
"family": "Sepehrband",
"given": "Farshid"
},
{
"family": "Stahl",
"given": "Eli A."
},
{
"family": "Huo",
"given": "Yuda"
},
{
"family": "Ahmed",
"given": "Mohsin"
},
{
"family": "Croll",
"given": "Susan"
},
{
"literal": "GHS-RGC DiscovEHR Collaboration"
},
{
"literal": "Mayo-RGC Project Generation"
},
{
"literal": "Penn Medicine BioBank"
},
{
"family": "Salerno",
"given": "William"
},
{
"family": "Overton",
"given": "John D."
},
{
"family": "Marchini",
"given": "Jonathan"
},
{
"family": "Reid",
"given": "Jeffrey"
},
{
"family": "Lotta",
"given": "Luca A."
},
{
"family": "Baras",
"given": "Aris"
},
{
"literal": "Regeneron Genetics Center"
},
{
"family": "Abecasis",
"given": "Goncalo R."
},
{
"family": "Coppola",
"given": "Giovanni"
},
{
"family": "Gelfman",
"given": "Sahar"
}
],
"container-title-short": "Nature",
"volume": "653",
"issue": "8115",
"page": "796-808",
"DOI": "10.1038/s41586-026-10345-6",
"PMID": "41951733",
"PMCID": "PMC13190288",
"ISSN": "0028-0836",
"publisher": "Nature Portfolio",
"URL": "https://doi.org/10.1038/s41586-026-10345-6",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
8
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: BCFtools, SAMtools, cowplot, 7 other tools, genetics / omics, other condition, 2 references
[2] doi:10.1038/s41467-026-72598-z [code]
Functional impact of genetic background on variable expressivity in neurodevelopmental disorders.
Journal: Nature communications
In common: BCFtools, SAMtools, data.table, 7 other tools, other condition, 1 reference
[3] doi:10.1038/s41467-026-73902-7 [code]
GWAS on short tandem repeats identifies genetic mechanisms in Alzheimer's disease.
Journal: Nature communications
In common: BCFtools, SAMtools, data.table, 6 other tools, genetics / omics, 1 reference
[4] doi:10.21203/rs.3.rs-9927928/v1 [code]
Genome-wide and allele-resolved maps of the radial architecture of the mouse genome
Journal: Research Square (preprint)
In common: BCFtools, SAMtools, cowplot, 7 other tools, genetics / omics
[5] doi:10.1038/s41467-026-71790-5 [code]
Recurrent DNA break clusters drive replication-stress-induced copy number variants and genome diversification.
Journal: Nature communications
In common: BCFtools, SAMtools, cowplot, 7 other tools, genetics / omics
[6] doi:10.1186/s13148-026-02082-4 [code]
DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum.
Journal: Clinical epigenetics
In common: cowplot, ggplot2, tidyverse, genetics / omics, other condition, 5 references
[7] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: BCFtools, SAMtools, data.table, 7 other tools
[8] doi:10.1002/alz.71271 [code]
Intracellular protein GBF1 displays significant associations with amyloid pathology in Alzheimer's disease.
Journal: Alzheimer's & dementia : the journal of the Alzheimer's Association
In common: BCFtools, SAMtools, statsmodels, 6 other tools, 1 reference
[9] doi:10.1038/s41592-026-03211-w [code]
Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.
Journal: Nature methods
In common: SAMtools, cowplot, data.table, 7 other tools, genetics / omics
[10] doi:10.1038/s41586-026-10612-6 [code]
Acquired genetic and cell-state changes in IDH-mutant glioma progression.
Journal: Nature
In common: BCFtools, SAMtools, cowplot, 5 other tools, other condition

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.