OSCR

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches
  1. [1] § Methods › Variant scoring with SuPreMo-Akita ↔ scripts/SuPreMo.py, lines 312–387 · score 0.79 · reverse complement, map corresponds, contact frequency maps, alternate allele, Akita scores, augmented
  2. [2] § Methods › Variant exclusions ↔ scripts/SuPreMo.py, lines 134–188 · score 0.76 · generate mutated sequences, prediction models, genome folding, SuPreMo, pipeline, kb
  3. [3] § Results › ASD dnSVs disrupt neighboring neuronal regulatory element contacts ↔ walkthroughs/SuPreMo-Akita_weighted_scores.ipynb, lines 155–193 · score 0.63 · ROI bins, weight track, weighted scoring, disruption track, maps
  4. [4] § Results › ASD dnSVs disrupt neighboring neuronal regulatory element contacts ↔ scripts/get_Akita_scores_utils.py, lines 305–337 · score 0.63 · ROI bins, weight track, weighted scoring, disruption track, maps
  5. [5] § Methods › Akita model ↔ scripts/SuPreMo.py, lines 6–64 · score 0.60 · H1hESC, GM12878, HCT116, HFF, Spearman, MSE
  6. [6] § Methods › CREint weighted scoring ↔ walkthroughs/SuPreMo-Akita_weighted_scores.ipynb, lines 155–193 · score 0.60 · weight track, weighted score, disruption track, BED, ROI, bin
  7. [7] § Methods › Criteria for variant prioritization ↔ scripts/SuPreMo.py, lines 312–387 · score 0.53 · reverse complement, predicted contacts, augmenting, shifting, sequence, Akita

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 876 lines · 40 KB · no license · 4 matches

  1. #!/usr/bin/env python
  2. # SuPreMo: Pipeline for generating mutated sequences for input into predictive models and for scoring variants for disruption to genome folding.
  3. # Written in Python v 3.7.11
  4. '''
  5. usage: SuPreMo [-h] [--sequences SEQUENCES] [--fa FASTA] [--roi ROI] [--roi_scales ROI_SCALES [ROI_SCALES ...]] [--genome {hg19,hg38}]
  6. [--Akita_cell_types {HFF,H1hESC,GM12878,IMR90,HCT116} [{HFF,H1hESC,GM12878,IMR90,HCT116} ...]] [--scores {mse,corr,ssi,scc,ins,di,dec,tri,pca} [{mse,corr,ssi,scc,ins,di,dec,tri,pca} ...]]
  7. [--shift_by SHIFT_WINDOW [SHIFT_WINDOW ...]] [--shifts_file SHIFTS_FILE] [--file OUT_FILE] [--dir OUT_DIR] [--limit SVLEN_LIMIT] [--seq_len SEQ_LEN]
  8. [--revcomp {no_revcomp,add_revcomp,only_revcomp}] [--augment] [--get_seq] [--get_tracks] [--get_maps] [--get_Akita_scores] [--nrows NROWS]
  9. Input file
  10. Pipeline for generating mutated sequences for input into predictive models and for scoring variants for disruption to genome folding.
  11. positional arguments:
  12. Input file Input file with variants. Accepted formats are:
  13. VCF file,
  14. TSV file (SVannot output),
  15. BED file,
  16. TXT file.
  17. Can be gzipped. Coordinates should be 1-based and left open (,], except for SNPs (follow vcf 4.1/4.2 specifications). See custom_perturbations.ipynb for BED or TXT file input specifications.
  18. options:
  19. -h, --help show this help message and exit
  20. --sequences SEQUENCES
  21. Input fasta file with sequences. This file is one outputted by SuPreMo.
  22. (default: None)
  23. --fa FASTA Path to reference genome fasta file. Default: data/{genome}.fa, where genome is hg38 or hg19.
  24. (default: None)
  25. --roi ROI Path to text file with regions of intererst (roi) to use for scoring. Tab-delimited text file with required columns: chrom, start, end or chrom, start1, end1, start2, end2 for paired regions. Alternatively, use 'genes' to weight 2 kb window around protein-coding gene transcription start sites.
  26. (default: None)
  27. --roi_scales ROI_SCALES [ROI_SCALES ...]
  28. Scale by which regions of interest (roi) are upweighted compared to the rest of the map. Default is 10 meaning regions of interest have weights of 10 whereas the rest of the bins have weights of 1.
  29. (default: [10])
  30. --genome {hg19,hg38} Genome to be used: hg19 or hg38. (default: ['hg38'])
  31. --Akita_cell_types {HFF,H1hESC,GM12878,IMR90,HCT116} [{HFF,H1hESC,GM12878,IMR90,HCT116} ...]
  32. Cell types to make Akita predictions for: HFF, H1hESC, GM12878, IMR90, and/or HCT116.
  33. (default: ['HFF'])
  34. --scores {mse,corr,ssi,scc,ins,di,dec,tri,pca} [{mse,corr,ssi,scc,ins,di,dec,tri,pca} ...]
  35. Method(s) used to calculate disruption scores. Use abbreviations as follows:
  36. mse: Mean squared error
  37. corr: Spearman correlation
  38. ssi: Structural similarity index measure
  39. scc: Stratum adjusted correlation coefficient
  40. ins: Insulation
  41. di: Directionality index
  42. dec: Contact decay
  43. tri: Triangle method
  44. pca: Principal component method.
  45. (default: ['mse', 'corr'])
  46. --shift_by SHIFT_WINDOW [SHIFT_WINDOW ...]
  47. Values for shifting prediciton windows inputted as space-separated integers (e.g. -1 0 1). Values outside of range -450000 ≤ x ≤ 450000 will be ignored. Prediction windows at the edge of chromosome arms will only be shifted in the direction that is possible (ex. for window at chrom start, a -1 shift will be treated as a 1 shift since it is not possible to shift left).
  48. (default: [0])
  49. --shifts_file SHIFTS_FILE
  50. Path to file with values to shift each variant by. Required when not all variants are shifted by the same amount. File should be a text file with the same number of rows as the input file and a column which contains a shift value (negative or positive integer) by which to shift the variants in the corresponding row of the input file. (default: None)
  51. --file OUT_FILE Prefix for output files. Saved files will overwrite any existing files.
  52. (default: SuPreMo)
  53. --dir OUT_DIR Output directory. If directory already exists, files will be saved in existing directory. If the same files already exists in that directory, new files will overwrite them. (default: SuPreMo_output)
  54. --limit SVLEN_LIMIT Maximum length of variants to be scored. Filtering out variants that are too big can save time and memory. If not specified, will be set to 2/3 of seq_len.
  55. (default: None)
  56. --seq_len SEQ_LEN Length for sequences to generate. Default value is based on Akita requirement. If non-default value is set, get_Akita_scores must be false.
  57. (default: 1048576)
  58. --revcomp {no_revcomp,add_revcomp,only_revcomp}
  59. Option to use the reverse complement of the sequence:
  60. no_revcomp: no, only use the standard sequence;
  61. add_revcomp: yes, use both the standard sequence and its reverse complement;
  62. only_revcomp: yes, only use the reverse complement of the sequence.
  63. The reverse complement of the sequence is only taken with 0 shift.
  64. (default: ['no_revcomp'])
  65. --augment
  66. Only applicable if --get_Akita_scores is specified. Get the mean and median scores from sequences with specified shifts and reverse complement. If augment is used but shift and revcomp are not specified, the following four sequences will be used:
  67. 1) no augmentation: 0 shift and no reverse complement,
  68. 2) +1bp shift and no reverse complement,
  69. 3) -1bp shift and no reverse complement,
  70. 4) 0 shift and take reverse complement.
  71. (default: False)
  72. --get_seq Save sequences for the reference and alternate alleles in fa file format. If --get_seq is not specified, must specify --get_Akita_scores.
  73. Sequence name format: {var_index}_{shift}_{revcomp_annot}_{seq_index}_{var_rel_pos}
  74. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  75. shift: integer that window is shifted by;
  76. revcomp_annot: present only if reverse complement of sequence was taken;
  77. seq_index: index for sequences generated for that variant: 0-1 for non-BND reference and alternate sequences and 0-2 for BND left and right reference sequence and alternate sequence;
  78. var_rel_pos: relative position of variant in sequence: list of two for non-BND variant positions in reference and alternate sequence and an integer for BND breakend position in reference and alternate sequences.
  79. There are 2-3 entries per prediction (2 for non-BND variants and 3 for BND variants).
  80. To read fasta file: pysam.Fastafile(filename).fetch(seqname, start, end).upper().
  81. To get sequence names in fasta file: pysam.Fastafile(filename).references.
  82. (default: False)
  83. --get_tracks Save disruption score tracks (448 bins) in npy file format. Only possible for mse and corr scores.
  84. Dictionary item name format: {var_index}_{track}_{shift}_{revcomp_annot}
  85. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  86. track: disruption score track specified;
  87. shift: integer that window is shifted by;
  88. revcomp_annot: present only if reverse complement of sequence was taken.
  89. There is 1 entry per prediction: a 448x1 array.
  90. To read into a dictionary in python: np.load(filename, allow_pickle="TRUE").item()
  91. (default: False)
  92. --get_maps Save predicted contact frequency maps in npy file format.
  93. Dictionary item name format: {var_index}_{shift}_{revcomp_annot}
  94. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  95. shift: integer that window is shifted by;
  96. revcomp_annot: present only if reverse complement of sequence was taken.
  97. There is 1 entry per prediction. Each entry contains the following: 2 (3 for chromosomal rearrangements) arrays that correspond to the upper right triangle of the predicted contact frequency maps, the relative variant position in the map, and the first coordinate of the sequence that the map corresponds to.
  98. To read into a dictionary in python: np.load(filename, allow_pickle="TRUE").item()
  99. (default: False)
  100. --get_Akita_scores Get disruption scores. If --get_Akita_scores is not specified, must specify --get_seq. Scores saved in a dataframe with the same number of rows as the input. For multiple alternate alleles, the scores are separated by a comma. To convert the scores from strings to integers, use float(x), after separating rows with multiple alternate alleles. Scores go up to 20 decimal points.
  101. (default: False)
  102. --nrows NROWS Number of rows (perturbations) to read at a time from input. When dealing with large inputs, selecting a subset of rows to read at a time allows scores to be saved in increments and uses less memory. Files with scores and filtered out variants will be temporarily saved in output direcotry. The file names will have a suffix corresponding to the set of nrows (0-based), for example for an input with 2700 rows and with nrows = 1000, there will be 3 sets. At the end of the run, these files will be concatenated into a comprehensive file and the temporary files will be removed.
  103. (default: 1000)
  104. '''
  105. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  106. # Parse through arguments
  107. import argparse
  108. class CustomFormatter(argparse.RawTextHelpFormatter, argparse.ArgumentDefaultsHelpFormatter):
  109. pass
  110. parser = argparse.ArgumentParser(
  111. prog = 'SuPreMo',
  112. formatter_class = CustomFormatter,
  113. description='''Pipeline for generating mutated sequences for input into predictive models and for scoring variants for disruption to genome folding.''')
  114. parser.add_argument('in_file',
  115. metavar = 'Input file',
  116. help = ''' Input file with variants. Accepted formats are:
  117. VCF file,
  118. TSV file (SVannot output),
  119. BED file,
  120. TXT file.
  121. Can be gzipped. Coordinates should be 1-based and left open (,], except for SNPs (follow vcf 4.1/4.2 specifications). See custom_perturbations.ipynb for BED or TXT file input specifications.
  122. ''',
  123. type = str)
  124. parser.add_argument('--sequences',
  125. dest = 'sequences',
  126. help = ''' Input fasta file with sequences. This file is one outputted by SuPreMo.
  127. ''',
  128. type = str,
  129. required = False)
  130. parser.add_argument('--fa',
  131. dest = 'fasta',
  132. help = '''Path to reference genome fasta file. Default: data/{genome}.fa, where genome is hg38 or hg19.
  133. ''',
  134. type = str,
  135. required = False)
  136. parser.add_argument('--roi',
  137. dest = 'roi',
  138. help = '''Path to text file with regions of intererst (roi) to use for scoring. Tab-delimited text file with required columns: chrom, start, end or chrom, start1, end1, start2, end2 for paired regions. Alternatively, use 'genes' to weight 2 kb window around protein-coding gene transcription start sites.
  139. ''',
  140. type = str,
  141. required = False)
  142. parser.add_argument('--roi_scales',
  143. dest = 'roi_scales',
  144. nargs = '+',
  145. help = '''Scale by which regions of interest (roi) are upweighted compared to the rest of the map. Default is 10 meaning regions of interest have weights of 10 whereas the rest of the bins have weights of 1.
  146. ''',
  147. type = int,
  148. default = [10],
  149. required = False)
  150. parser.add_argument('--genome',
  151. dest = 'genome',
  152. nargs = 1,
  153. help = '''Genome to be used: hg19 or hg38.''',
  154. type = str,
  155. choices = ['hg19', 'hg38'],
  156. default = ['hg38'],
  157. required = False)
  158. parser.add_argument('--Akita_cell_types',
  159. dest = 'Akita_cell_types',
  160. nargs = '+',
  161. help = '''Cell types to make Akita predictions for: HFF, H1hESC, GM12878, IMR90, and/or HCT116.
  162. ''',
  163. type = str,
  164. choices = ['HFF', 'H1hESC', 'GM12878', 'IMR90', 'HCT116'],
  165. default = ['HFF'],
  166. required = False)
  167. parser.add_argument('--scores',
  168. dest = 'scores',
  169. nargs = '+',
  170. help = '''
  171. Method(s) used to calculate disruption scores. Use abbreviations as follows:
  172. mse: Mean squared error
  173. corr: Spearman correlation
  174. ssi: Structural similarity index measure
  175. scc: Stratum adjusted correlation coefficient
  176. ins: Insulation
  177. di: Directionality index
  178. dec: Contact decay
  179. tri: Triangle method
  180. pca: Principal component method.
  181. ''',
  182. type = str,
  183. choices = ['mse', 'corr', 'ssi', 'scc', 'ins', 'di', 'dec', 'tri', 'pca'],
  184. default = ['mse', 'corr'],
  185. required = False)
  186. parser.add_argument('--shift_by',
  187. dest = 'shift_window',
  188. nargs = '+',
  189. help = '''Values for shifting prediciton windows inputted as space-separated integers (e.g. -1 0 1). Values outside of range -450000 ≤ x ≤ 450000 will be ignored. Prediction windows at the edge of chromosome arms will only be shifted in the direction that is possible (ex. for window at chrom start, a -1 shift will be treated as a 1 shift since it is not possible to shift left).
  190. ''',
  191. type = int,
  192. default = [0],
  193. required = False)
  194. parser.add_argument('--shifts_file',
  195. dest = 'shifts_file',
  196. help = 'Path to file with values to shift each variant by. Required when not all variants are shifted by the same amount. File should be a text file with the same number of rows as the input file and a column which contains a shift value (negative or positive integer) by which to shift the variants in the corresponding row of the input file.',
  197. type = str,
  198. required = False)
  199. parser.add_argument('--file',
  200. dest = 'out_file',
  201. help = '''Prefix for output files. Saved files will overwrite any existing files.
  202. ''',
  203. type = str,
  204. default = 'SuPreMo',
  205. required = False)
  206. parser.add_argument('--dir',
  207. dest = 'out_dir',
  208. help = 'Output directory. If directory already exists, files will be saved in existing directory. If the same files already exists in that directory, new files will overwrite them.',
  209. type = str,
  210. default = 'SuPreMo_output',
  211. required = False)
  212. parser.add_argument('--limit',
  213. dest = 'svlen_limit',
  214. help = '''Maximum length of variants to be scored. Filtering out variants that are too big can save time and memory. If not specified, will be set to 2/3 of seq_len.
  215. ''',
  216. type = int,
  217. required = False)
  218. parser.add_argument('--seq_len',
  219. dest = 'seq_len',
  220. help = '''Length for sequences to generate. Default value is based on Akita requirement. If non-default value is set, get_Akita_scores must be false.
  221. ''',
  222. type = int,
  223. default = 1048576,
  224. required = False)
  225. parser.add_argument('--revcomp',
  226. dest = 'revcomp',
  227. nargs = 1,
  228. help = '''
  229. Option to use the reverse complement of the sequence:
  230. no_revcomp: no, only use the standard sequence;
  231. add_revcomp: yes, use both the standard sequence and its reverse complement;
  232. only_revcomp: yes, only use the reverse complement of the sequence.
  233. The reverse complement of the sequence is only taken with 0 shift.
  234. ''',
  235. type = str,
  236. choices = ['no_revcomp', 'add_revcomp', 'only_revcomp'],
  237. default = ['no_revcomp'],
  238. required = False)
  239. parser.add_argument('--augment',
  240. dest = 'augment',
  241. help = '''
  242. Only applicable if --get_Akita_scores is specified. Get the mean and median scores from sequences with specified shifts and reverse complement. If augment is used but shift and revcomp are not specified, the following four sequences will be used:
  243. 1) no augmentation: 0 shift and no reverse complement,
  244. 2) +1bp shift and no reverse complement,
  245. 3) -1bp shift and no reverse complement,
  246. 4) 0 shift and take reverse complement.
  247. ''',
  248. action='store_true',
  249. required = False)
  250. parser.add_argument('--get_seq',
  251. dest = 'get_seq',
  252. help = '''Save sequences for the reference and alternate alleles in fa file format. If --get_seq is not specified, must specify --get_Akita_scores.
  253. Sequence name format: {var_index}_{shift}_{revcomp_annot}_{seq_index}_{var_rel_pos}
  254. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  255. shift: integer that window is shifted by;
  256. revcomp_annot: present only if reverse complement of sequence was taken;
  257. seq_index: index for sequences generated for that variant: 0-1 for non-BND reference and alternate sequences and 0-2 for BND left and right reference sequence and alternate sequence;
  258. var_rel_pos: relative position of variant in sequence: list of two for non-BND variant positions in reference and alternate sequence and an integer for BND breakend position in reference and alternate sequences.
  259. There are 2-3 entries per prediction (2 for non-BND variants and 3 for BND variants).
  260. To read fasta file: pysam.Fastafile(filename).fetch(seqname, start, end).upper().
  261. To get sequence names in fasta file: pysam.Fastafile(filename).references.
  262. ''',
  263. action = 'store_true',
  264. required = False)
  265. parser.add_argument('--get_tracks',
  266. dest = 'get_tracks',
  267. help = '''Save disruption score tracks (448 bins) in npy file format. Only possible for mse and corr scores.
  268. Dictionary item name format: {var_index}_{track}_{shift}_{revcomp_annot}
  269. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  270. track: disruption score track specified;
  271. shift: integer that window is shifted by;
  272. revcomp_annot: present only if reverse complement of sequence was taken.
  273. There is 1 entry per prediction: a 448x1 array.
  274. To read into a dictionary in python: np.load(filename, allow_pickle="TRUE").item()
  275. ''',
  276. action='store_true',
  277. required = False)
  278. parser.add_argument('--get_maps',
  279. dest = 'get_maps',
  280. help = '''Save predicted contact frequency maps in npy file format.
  281. Dictionary item name format: {var_index}_{shift}_{revcomp_annot}
  282. var_index: input row number, followed by _0, _1, etc for each allele of variants with multiple alternate alleles;
  283. shift: integer that window is shifted by;
  284. revcomp_annot: present only if reverse complement of sequence was taken.
  285. There is 1 entry per prediction. Each entry contains the following: 2 (3 for chromosomal rearrangements) arrays that correspond to the upper right triangle of the predicted contact frequency maps, the relative variant position in the map, and the first coordinate of the sequence that the map corresponds to.
  286. To read into a dictionary in python: np.load(filename, allow_pickle="TRUE").item()
  287. ''',
  288. action='store_true',
  289. required = False)
  290. parser.add_argument('--get_Akita_scores',
  291. dest = 'get_Akita_scores',
  292. help = '''Get disruption scores. If --get_Akita_scores is not specified, must specify --get_seq. Scores saved in a dataframe with the same number of rows as the input. For multiple alternate alleles, the scores are separated by a comma. To convert the scores from strings to integers, use float(x), after separating rows with multiple alternate alleles. Scores go up to 20 decimal points.
  293. ''',
  294. action='store_true',
  295. required = False)
  296. parser.add_argument('--nrows',
  297. dest = 'nrows',
  298. help = '''Number of rows (perturbations) to read at a time from input. When dealing with large inputs, selecting a subset of rows to read at a time allows scores to be saved in increments and uses less memory. Files with scores and filtered out variants will be temporarily saved in output direcotry. The file names will have a suffix corresponding to the set of nrows (0-based), for example for an input with 2700 rows and with nrows = 1000, there will be 3 sets. At the end of the run, these files will be concatenated into a comprehensive file and the temporary files will be removed.
  299. ''',
  300. type = int,
  301. default = 1000,
  302. required = False)
  303. args = parser.parse_args()
  304. in_file = args.in_file
  305. input_sequences = args.sequences
  306. fasta_path = args.fasta
  307. roi = args.roi
  308. roi_scales = args.roi_scales
  309. genome = args.genome[0]
  310. Akita_cell_types = args.Akita_cell_types
  311. scores_to_use = args.scores
  312. shift_by = args.shift_window
  313. shifts_file = args.shifts_file
  314. out_file = args.out_file
  315. out_dir = args.out_dir
  316. svlen_limit = args.svlen_limit
  317. seq_len = args.seq_len
  318. revcomp = args.revcomp[0]
  319. augment = args.augment
  320. get_seq = args.get_seq
  321. get_tracks = args.get_tracks
  322. get_maps = args.get_maps
  323. get_Akita_scores = args.get_Akita_scores
  324. var_set_size = args.nrows
  325. __version__ = '1.0'
  326. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  327. # Adjust inputs from arguments
  328. # Handle paths
  329. import os
  330. from pathlib import Path
  331. # This file path and repo path
  332. repo_path = Path(__file__).parents[1]
  333. # Output directory and file
  334. if not os.path.exists(out_dir):
  335. os.mkdir(out_dir)
  336. out_file = os.path.join(out_dir, out_file)
  337. # Data path
  338. chrom_lengths_path = f'{repo_path}/data/chrom_lengths_{genome}'
  339. centromere_coords_path = f'{repo_path}/data/centromere_coords_{genome}'
  340. # Handle argument dependencies
  341. use_roi = False
  342. if fasta_path is None:
  343. fasta_path = f'{repo_path}/data/{genome}.fa'
  344. if seq_len != 1048576:
  345. get_Akita_scores = False
  346. if svlen_limit is None:
  347. svlen_limit = 2/3*seq_len
  348. elif svlen_limit > 2/3*seq_len:
  349. raise ValueError("Maximum SV length limit should not be >2/3 of sequence length.")
  350. if not get_seq and not get_Akita_scores:
  351. raise ValueError('Either get_seq and/or get_Akita_scores must be specified.')
  352. if roi is not None and not get_Akita_scores:
  353. raise ValueError('get_Akita_scores must be specified to use roi.')
  354. # Adjust shift input: Remove shifts that are outside of allowed range
  355. max_shift = 0.4*seq_len
  356. shift_by = [x for x in shift_by if x > -max_shift and x < max_shift]
  357. # Adjust input for taking the reverse complement
  358. if revcomp == 'no_revcomp':
  359. revcomp_decision = [False]
  360. elif revcomp == 'add_revcomp':
  361. revcomp_decision = [False, True]
  362. elif revcomp == 'only_revcomp':
  363. revcomp_decision = [True]
  364. if augment and shift_by == [0] and revcomp == 'no_revcomp':
  365. shift_by = [-1,0,1]
  366. revcomp_decision = [False, True]
  367. revcomp_decision_i = revcomp_decision
  368. import pysam
  369. if input_sequences is not None:
  370. seq_names = pysam.Fastafile(input_sequences).references
  371. # Create dictionaries to save sequences, maps, and disruption score tracks, if specified
  372. if get_seq:
  373. sequences = {}
  374. if get_maps:
  375. variant_maps = {}
  376. if get_Akita_scores == False:
  377. get_Akita_scores = True
  378. print('Must get scores to get maps. --get_Akita_scores was not specified but will be applied.')
  379. if get_tracks:
  380. variant_tracks = {}
  381. if get_Akita_scores == False:
  382. get_Akita_scores = True
  383. print('Must get scores to get tracks. --get_Akita_scores was not specified but will be applied.')
  384. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  385. # Read in (and adjust) data
  386. import pandas as pd
  387. chrom_lengths = pd.read_table(chrom_lengths_path, header = None, names = ['CHROM', 'chrom_max'])
  388. centromere_coords = pd.read_table(centromere_coords_path, sep = '\t')
  389. fasta_open = pysam.Fastafile(fasta_path)
  390. # Assign necessary values to variables across module
  391. # Reading utilities
  392. import sys
  393. sys.path.insert(0, 'scripts/')
  394. import reading_utils
  395. reading_utils.var_set_size = var_set_size
  396. # SuPreMo utilities
  397. import get_seq_utils
  398. get_seq_utils.fasta_open = fasta_open
  399. get_seq_utils.chrom_lengths = chrom_lengths
  400. get_seq_utils.centromere_coords = centromere_coords
  401. get_seq_utils.svlen_limit = svlen_limit
  402. get_seq_utils.seq_length = seq_len
  403. get_seq_utils.half_patch_size = round(seq_len/2)
  404. # SuPreMo-Akita utilities
  405. if get_Akita_scores:
  406. import get_Akita_scores_utils
  407. get_Akita_scores_utils.chrom_lengths = chrom_lengths
  408. get_Akita_scores_utils.centromere_coords = centromere_coords
  409. if roi is None or roi_scales == [0]:
  410. use_roi = False
  411. else:
  412. use_roi = True
  413. import get_roi_utils
  414. get_roi_utils.pred_len = get_Akita_scores_utils.target_length_cropped * get_Akita_scores_utils.bin_size
  415. get_roi_utils.d = get_Akita_scores_utils.bin_size * 5
  416. get_Akita_scores_utils.roi_coords_BED = get_roi_utils.get_roi(roi, genome)
  417. from pybedtools import BedTool
  418. get_Akita_scores_utils.BedTool = BedTool
  419. get_Akita_scores_utils.roi_scales = roi_scales
  420. import sys
  421. import numpy as np
  422. nt = ['A', 'T', 'C', 'G']
  423. var_set = 0
  424. var_set_list = []
  425. print(f'Log file being saved here: {out_file}_log')
  426. while True:
  427. # Read in variants
  428. variants = reading_utils.read_input(in_file, var_set)
  429. if len(variants) == 0:
  430. break
  431. if shifts_file is not None:
  432. variants = pd.concat([variants,
  433. pd.read_csv(shifts_file, sep = '\t', low_memory=False, names = ['shift_by'],
  434. skiprows = 1 + var_set*var_set_size, nrows = var_set_size)],
  435. axis = 1)
  436. # Index input based on row number and create output with same indexes
  437. variants['var_index'] = list(range(var_set*var_set_size, var_set*var_set_size + len(variants)))
  438. variants['var_index'] = variants['var_index'].astype(str)
  439. # If there are multiple alternate alleles, split those into new rows and indexes
  440. if any([',' in x for x in variants.ALT]):
  441. variants = (variants
  442. .set_index(['CHROM', 'POS', 'REF', 'var_index'])
  443. .apply(lambda x: x.str.split(',').explode())
  444. .reset_index())
  445. g = variants.groupby(['var_index'])
  446. variants.loc[g['var_index'].transform('size').gt(1),
  447. 'var_index'] += '-'+g.cumcount().astype(str)
  448. variant_scores = pd.DataFrame({'var_index':variants.var_index})
  449. if use_roi:
  450. variant_roi_ids = pd.DataFrame({'var_index':variants.var_index})
  451. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  452. # Filter out variants that cannot be scored
  453. # Get indexes for variants to exclude
  454. # Exclude mitochondrial variants
  455. chrM_var = pd.DataFrame({'var_index' : list(variants[variants.CHROM == 'chrM'].var_index),
  456. 'reason' : ' Mitochondrial chromosome.'})
  457. # Exclude variants larger than limit
  458. if 'SVLEN' in variants.columns:
  459. too_long_var = pd.DataFrame({'var_index' : [y for x,y in zip(variants.SVLEN, variants.var_index)
  460. if not pd.isnull(x) and abs(int(x)) > svlen_limit],
  461. 'reason' : f' SV longer than {svlen_limit}.'})
  462. unsuitable_var = pd.DataFrame({'var_index' : [y for x,y,z in zip(variants.SVTYPE, variants.var_index, variants.ALT)
  463. if not pd.isnull(x) and
  464. x not in ["DEL", "DUP", "INV", "BND"] and
  465. all([g not in nt for g in z])],
  466. 'reason' : ' SV type not compatible.'})
  467. else:
  468. too_long_var = pd.DataFrame()
  469. unsuitable_var = pd.DataFrame()
  470. filtered_out = pd.concat([chrM_var, too_long_var], axis = 0)
  471. filtered_out = pd.concat([filtered_out, unsuitable_var], axis = 0)
  472. filtered_out.var_index = filtered_out.var_index.astype('str')
  473. # Save filtered out variants into file
  474. filtered_out.to_csv(f'{out_file}_filtered_out_{var_set}', sep = ':', index = False, header = False)
  475. # Exclude
  476. variants = variants[[x not in filtered_out.var_index.values for x in variants.var_index]]
  477. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  478. # Run: Make Akita predictions and calculate disruption scores
  479. # Create log file to save standard output with error messages
  480. std_output = sys.stdout
  481. log_file = open(f'{out_file}_log_{var_set}','w')
  482. sys.stdout = log_file
  483. # Loop through each row (not index) and get disruption scores
  484. for i in range(len(variants)):
  485. variant = variants.iloc[i]
  486. var_index = variant.var_index
  487. CHR = variant.CHROM
  488. POS = variant.POS
  489. REF = variant.REF
  490. ALT = variant.ALT
  491. if 'SVTYPE' in variants.columns:
  492. END = variant.END
  493. SVTYPE = variant.SVTYPE
  494. SVLEN = variant.SVLEN
  495. else:
  496. END = np.nan
  497. SVTYPE = np.nan
  498. SVLEN = 0
  499. if shifts_file is not None:
  500. shift_by = [int(variant.shift_by)]
  501. for shift in shift_by:
  502. # Take reverse complement only with 0 shift
  503. if shift != 0 & True in revcomp_decision:
  504. revcomp_decision_i = [False]
  505. else:
  506. revcomp_decision_i = revcomp_decision
  507. for revcomp in revcomp_decision_i:
  508. try:
  509. if revcomp:
  510. revcomp_annot = '_revcomp'
  511. else:
  512. revcomp_annot = ''
  513. if input_sequences is not None:
  514. # Generate sequences_i from sequence input
  515. if revcomp_annot == '':
  516. sequence_names = [x for x in seq_names if x.startswith(f'{var_index}_{shift}') and
  517. 'revcomp' not in x]
  518. elif revcomp_annot == '_revcomp':
  519. sequence_names = [x for x in seq_names if x.startswith(f'{var_index}_{shift}{revcomp_annot}')]
  520. sequences_i = []
  521. for sequence_name in sequence_names:
  522. sequences_i.append(pysam.Fastafile(input_sequences).fetch(sequence_name, 0, seq_len).upper())
  523. sequences_i.append([int(x) for x in sequence_name.split('[')[1].split(']')[0].split('_')])
  524. else:
  525. # Create sequences_i from variant input
  526. sequences_i = get_seq_utils.get_sequences_SV(CHR, POS, REF, ALT, END, SVTYPE, shift, revcomp)
  527. if get_seq:
  528. # Get relative position of variant in sequence
  529. var_rel_pos = str(sequences_i[-1]).replace(', ', '_')
  530. for ii in range(len(sequences_i[:-1][:3])):
  531. sequences[f'{var_index}_{shift}{revcomp_annot}_{ii}_{var_rel_pos}'] = sequences_i[:-1][ii]
  532. if get_Akita_scores:
  533. scores = get_Akita_scores_utils.get_scores(CHR, POS, SVTYPE, SVLEN,
  534. sequences_i, scores_to_use,
  535. shift, revcomp,
  536. get_tracks, get_maps, use_roi, Akita_cell_types)
  537. if get_tracks:
  538. for track in [x for x in scores.keys() if 'track' in x]:
  539. variant_tracks[f'{var_index}_{track}_{shift}{revcomp_annot}'] = scores[track]
  540. del scores[track]
  541. if get_maps:
  542. for map in [x for x in scores.keys() if 'map' in x]:
  543. variant_maps[f'{var_index}_{map}_{shift}{revcomp_annot}'] = scores[map]
  544. del scores[map]
  545. if shifts_file is not None:
  546. shift = 'shifted'
  547. if use_roi:
  548. for roi_ids in [x for x in scores.keys() if 'roi_id' in x]:
  549. variant_roi_ids.loc[variant_scores.var_index == var_index, f'roi_ids_{shift}'] = scores[roi_ids]
  550. del scores[roi_ids]
  551. for score in scores:
  552. variant_scores.loc[variant_scores.var_index == var_index,
  553. f'{score}_{shift}{revcomp_annot}'] = scores[score]
  554. print(str(var_index) + ' (' + str(shift) + f' shift{revcomp_annot})')
  555. except Exception as e:
  556. print(str(var_index) + ' (' + str(shift) + f' shift{revcomp_annot})' + ': Error:', e)
  557. pass
  558. # Write standard output with error messages and warnings to log file
  559. sys.stdout = std_output
  560. log_file.close()
  561. # Combine results from all sets
  562. # Write sequences to fasta file
  563. if get_seq:
  564. if var_set == 0:
  565. sequences_all = sequences.copy()
  566. else:
  567. sequences_all.update(sequences)
  568. # Write scores to data frame
  569. if get_Akita_scores:
  570. # Take average of augmented sequences
  571. if augment:
  572. for score in scores:
  573. cols = [x for x in variant_scores.columns if score in x]
  574. variant_scores[f'{score}_mean'] = variant_scores[cols].mean(axis = 1)
  575. variant_scores[f'{score}_median'] = variant_scores[cols].median(axis = 1)
  576. variant_scores.drop(cols, axis = 1, inplace = True)
  577. # Convert scores from float to string so you can merge scores for variants with multiple alleles
  578. for col in variant_scores.iloc[:,1:].columns:
  579. variant_scores[col] = [format(x, '.20f') for x in variant_scores[col]]
  580. # Join scores for alternate alleles, separated by a comma
  581. if any(['-' in x for x in variant_scores.var_index]):
  582. variant_scores['var_index'] = variant_scores.var_index.str.split('-').str[0]
  583. variant_scores = (variant_scores
  584. .set_index(['var_index'], drop = False)
  585. .rename(columns = {'var_index':'var_index2'})
  586. .groupby('var_index2')
  587. .transform(','.join)
  588. .reset_index()
  589. .drop_duplicates())
  590. if var_set == 0:
  591. variant_scores.to_csv(f'{out_file}_scores_{var_set}', sep = '\t', index = False)
  592. if use_roi:
  593. variant_roi_ids.to_csv(f'{out_file}_roi_ids_{var_set}', sep = '\t', index = False)
  594. else:
  595. variant_scores.to_csv(f'{out_file}_scores_{var_set}', sep = '\t', index = False, header = False)
  596. if use_roi:
  597. variant_roi_ids.to_csv(f'{out_file}_roi_ids_{var_set}', sep = '\t', index = False, header = False)
  598. if get_tracks:
  599. if var_set == 0:
  600. variant_tracks_all = variant_tracks.copy()
  601. else:
  602. variant_tracks_all.update(variant_tracks)
  603. if get_maps:
  604. if var_set == 0:
  605. variant_maps_all = variant_maps.copy()
  606. else:
  607. variant_maps_all.update(variant_maps)
  608. var_set_list.append(var_set)
  609. var_set += 1
  610. # # # # # # # # # # # # # # # # # # # # # # # # # # # # #
  611. # Save results
  612. # Write sequences to fasta file
  613. if get_seq:
  614. sequences_fasta = open(f'{out_file}_sequences.fa','w')
  615. for seq_name, sequence in sequences_all.items():
  616. seq_name_line = ">" + seq_name + "\n"
  617. sequences_fasta.write(seq_name_line)
  618. sequence_line = sequence + "\n"
  619. sequences_fasta.write(sequence_line)
  620. sequences_fasta.close()
  621. # Write disruption tracks and/or predictions to array
  622. if get_tracks:
  623. np.save(f'{out_file}_tracks.npy', variant_tracks_all)
  624. if get_maps:
  625. np.save(f'{out_file}_maps.npy', variant_maps_all)
  626. # Combine subset files into one
  627. os.system(f'rm -f {out_file}_filtered_out; \
  628. for file in {out_file}_filtered_out_*; \
  629. do cat "$file" >> {out_file}_filtered_out && rm "$file"; \
  630. done')
  631. os.system(f'rm -f {out_file}_log; \
  632. for file in {out_file}_log_*; \
  633. do cat "$file" >> {out_file}_log && rm "$file"; \
  634. done')
  635. if get_Akita_scores:
  636. os.system(f'rm -f {out_file}_scores; \
  637. for file in {out_file}_scores_*; \
  638. do cat "$file" >> {out_file}_scores && rm "$file"; \
  639. done')
  640. if use_roi:
  641. os.system(f'rm -f {out_file}_roi_ids; \
  642. for file in {out_file}_roi_ids_*; \
  643. do cat "$file" >> {out_file}_roi_ids && rm "$file"; \
  644. done')
  645. # Adjust log file to only have 1 row per variant
  646. if os.path.exists(f'{out_file}_log'):
  647. log_file = pd.read_csv(f'{out_file}_log', names = ['output'], sep = '\t')
  648. # Move warnings (printed 1 line before variant) to variant line
  649. indexes = np.array([[index, index+1] for (index, item) in enumerate(log_file.output) if item.startswith('Warning')])
  650. if len(indexes) != 0:
  651. log_file.loc[indexes[:,1],'output'] = [x+': '+y for x,y in zip(list(log_file.loc[indexes[:,1],'output']),
  652. list(log_file.loc[indexes[:,0],'output']))]
  653. log_file.drop(indexes[:,0], axis = 0, inplace = True)
  654. log_file.to_csv(f'{out_file}_log', sep = '\t', header = None, index = False)

SuPreMo.py at commit 9799281, no license · at the source

Overview

  1. Gladstone Institute of Data Science and Biotechnology, San Francisco, California 94158, USA
  2. Department of Epidemiology and Biostatistics, University of California San Francisco, California 94158, USA
  3. Institute for Human Genetics, University of California San Francisco, San Francisco, California 94143, USA
  4. Department of Neurology, University of California San Francisco, San Francisco, California 94143, USA
  5. Weill Institute for Neurosciences, University of California San Francisco, San Francisco, California 94158, USA
  6. Bakar Computational Health Sciences Institute, University of California, San Francisco, California 94143, USA
  7. Chan Zuckerberg Biohub, San Francisco, California 94158, USA
Institutions: Gladstone Institutes (United States); University of California, San Francisco (United States); Chan Zuckerberg Biohub San Francisco (United States)
Journal: Genome research, volume 36, issue 9, pages 1825-1835
Dates: received 2 January 2025; accepted 2 July 2026; published online September 2026
Type: Research article · Language: English
License: CC BY-NC
Identifiers: DOI 10.1101/gr.280394.124 · PMID 42680562 · PMCID PMC13534236 · OpenAlex W4404144294
Open access: green, a free copy (OpenAlex)
Status: code verified
Categories: human (organism), autism (population), cellular / molecular (subfield)
Methods: Statistics, Connectivity, fMRI & imaging, Spectral & time-frequency
MeSH: Autism Spectrum Disorder*, Genomic Structural Variation*, Neurons*, Chromatin, Gene Expression Regulation, Humans, Promoter Regions, Genetic, Regulatory Sequences, Nucleic Acid, Software (* major topic)
Topic: Genomics and Chromatin Dynamics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: NIH HHS (R03 OD034499); NHLBI NIH HHS (U01 HL157989); National Institutes of Health (NIH) 4D Nucleome Project (U01HL157989); Biswas Family Foundation; Simons Foundation (SFI-AN-AR-Data Analysis-00019589); Achievement Rewards for College Scientists Scholarships; Gladstone Institutes; NIH Office of the Director (R03OD034499); Additional Ventures; W. M. Keck Foundation
Citations: cited by 2 papers (Europe PMC); 55 references in the paper

Abstract

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We develop a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implement it in the SuPreMo-Akita software. With this tool, we rank hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This reveals that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritize candidate variants that disrupt ASD CREints and validate our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study suggests that disrupted genome folding is a potential genetic mechanism in a subset of ASD cases and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

ketringjoni/SuPreMo

License: none: the authors keep all their rights
State: the link answers, verified on 26 September 2026
Evidence: files inventoried
Commit: 9799281ec9b4ea5ea03702931e57950903f75424, 14 November 2024
Languages: Python (10), Jupyter (5)
Size: 67 files, 15 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, 5 notebooks
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: pandas (14 files), NumPy (13 files), pysam (6 files), Matplotlib (5 files), SciPy (5 files), BEDTools (4 files), Biopython (3 files), scikit-learn (2 files), TensorFlow (2 files), scikit-image (1 file)
Availability: 1 check, the latest on 26 September 2026: the link answers
  • 26 September 2026: the link answers
16 files

Code availability

SuPreMo V2 code is available as Supplemental Code and at GitHub (https://github.com/ketringjoni/SuPreMo/).

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 15 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Other data links

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 3, 28 September 2026

  • Authors: added Xingjie Ren (0000-0002-2348-4162); Amanda Everitt (0000-0001-9720-1922); Yin Shen (0000-0001-9901-5613); removed Xingjie Ren; Amanda Everitt; Yin Shen

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 9 MeSH terms, 10 funders, 55 references.

Cite

This paper

Gjoni, K., Ren, X., Everitt, A., Shen, Y., & Pollard, K. S. (2026). De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes. Genome research, 36(9), 1825-1835. https://doi.org/10.1101/gr.280394.124

BibTeX

@article{gjoni2026de,
author = {Gjoni, Ketrin and Ren, Xingjie and Everitt, Amanda and Shen, Yin and Pollard, Katherine S},
title = {{De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes}},
journal = {Genome research},
year = {2026},
month = sep,
volume = {36},
number = {9},
pages = {1825--1835},
publisher = {Cold Spring Harbor Laboratory Press},
issn = {1088-9051},
doi = {10.1101/gr.280394.124},
url = {https://doi.org/10.1101/gr.280394.124},
pmid = {42680562},
pmcid = {PMC13534236}
}

RIS

TY - JOUR
AU - Gjoni, Ketrin
AU - Ren, Xingjie
AU - Everitt, Amanda
AU - Shen, Yin
AU - Pollard, Katherine S
TI - De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes
T2 - Genome research
J2 - Genome Res
PY - 2026
DA - 2026/09/01
VL - 36
IS - 9
SP - 1825
EP - 1835
SN - 1088-9051
PB - Cold Spring Harbor Laboratory Press
DO - 10.1101/gr.280394.124
UR - https://doi.org/10.1101/gr.280394.124
LA - en
ER -

CSL-JSON

{
"id": "10.1101/gr.280394.124",
"type": "article-journal",
"title": "De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes",
"container-title": "Genome research",
"author": [
{
"family": "Gjoni",
"given": "Ketrin"
},
{
"family": "Ren",
"given": "Xingjie"
},
{
"family": "Everitt",
"given": "Amanda"
},
{
"family": "Shen",
"given": "Yin"
},
{
"family": "Pollard",
"given": "Katherine S"
}
],
"container-title-short": "Genome Res",
"volume": "36",
"issue": "9",
"page": "1825-1835",
"DOI": "10.1101/gr.280394.124",
"PMID": "42680562",
"PMCID": "PMC13534236",
"ISSN": "1088-9051",
"publisher": "Cold Spring Harbor Laboratory Press",
"URL": "https://doi.org/10.1101/gr.280394.124",
"language": "en",
"issued": {
"date-parts": [
[
2026,
9,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.xhgg.2026.100652 [code]
CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling.
Journal: HGG advances
In common: pysam, BEDTools, pandas, 3 other tools, autism, cellular / molecular, 6 references
[2] doi:10.1038/s41467-026-71877-z [code]
Hi-Compass: a depth-aware deep learning framework for predicting cell-type-specific 3D genome organization from single-cell to spatial resolution.
Journal: Nature communications
In common: pysam, BEDTools, scikit-image, 3 other tools, 4 references
[3] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: pysam, Biopython, BEDTools, 5 other tools, autism, 3 references
[4] doi:10.21203/rs.3.rs-9927928/v1 [code]
Genome-wide and allele-resolved maps of the radial architecture of the mouse genome
Journal: Research Square (preprint)
In common: pysam, BEDTools, pandas, 3 other tools, 5 references
[5] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: pysam, Biopython, BEDTools, 6 other tools, cellular / molecular
[6] doi:10.1126/sciadv.aed2952 [code]
Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.
Journal: Science advances
In common: pysam, Biopython, BEDTools, 4 other tools, cellular / molecular, 3 references
[7] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: pysam, Biopython, BEDTools, 6 other tools
[8] doi:10.1186/s13059-026-04177-w [code]
Genomic sequence evolution underlying human neocortical interareal diversification.
Journal: Genome biology
In common: pysam, BEDTools, TensorFlow, 5 other tools, cellular / molecular, 1 reference
[9] doi:10.1038/s41586-026-10391-0 [code]
Cell-type-targeted mitochondrial transplantation rescues cell degeneration.
Journal: Nature
In common: Biopython, TensorFlow, scikit-learn, 4 other tools, cellular / molecular, 3 references
[10] doi:10.1038/s41586-026-10512-9 [code]
Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.
Journal: Nature
In common: pysam, Biopython, BEDTools, 4 other tools, cellular / molecular, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.