OSCR

Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction.

Code ↔ Paper

7 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 7 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Methods › SPLiT-seq single-nucleus read demultiplexing ↔ splitseqdemultiplex_0.2.1.sh, lines 271–324 · score 0.74 · featureCounts, Seq_demultiplexing, UMI tools, SPLiT, samtools, memory
  2. [2] § Methods › SPLiT-seq single-nucleus read demultiplexing ↔ splitseqdemultiplex_0.2.2.sh, lines 280–333 · score 0.74 · featureCounts, Seq_demultiplexing, UMI tools, SPLiT, samtools, memory
  3. [3] § Methods › SPLiT-seq single-nucleus read demultiplexing ↔ splitseqdemultiplex_0.2.1.sh, lines 1–41 · score 0.60 · SPLiT Seq demultiplexing, pipeline, python, tool, FASTQ, barcoded
  4. [4] § Methods › SPLiT-seq single-nucleus read demultiplexing ↔ Extract_BC_UMI.py, lines 35–37 · score 0.53 · Seq barcode sequence, SPLiT, bp
  5. [5] § Methods › SPLiT-seq single-nucleus read demultiplexing ↔ splitseqdemultiplex_0.2.1.sh, lines 368–422 · score 0.53 · Seq_demultiplexing, SPLiT, pipeline, threads, parallelization, FASTQ
  6. [6] § Methods › Jensen-Shannon Divergence (JSD) analysis ↔ prep_TCC_matrix.py, lines 133–136 · score 0.51 · Jensen Shannon, pairwise, distance
  7. [7] § Methods › SPLiT-seq sublibrary generation ↔ Collapse_RanHex_Odt.sh, the whole file · a weak match · score 0.50 · RT primer, plates, dT, seq, barcoding

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Shell · 532 lines · 21 KB · MIT · 3 matches

  1. #!/bin/bash
  2. #alias python='python'
  3. #####################################################################################
  4. # Example Use #
  5. # Script modified on 29the august, 2019
  6. # A python script Collapse_Ranhex_Odt.py ( Dipankar / Dumaatravaie ) had replaced the bash script Collapse_Ranhex_Odt.sh
  7. # Fixed the issue with 'Too many arguments error' in parallel ( Dipankar / Dumaatravaie ) in the case of Large Number of cells (> 50,000 or more )
  8. #####################################################################################
  9. #bash splitseqdemultiplex.sh \
  10. # -n 12 \
  11. # -v merged \
  12. # -e 1 \
  13. # -m 10 \
  14. # -1 Round1_barcodes_new4.txt \
  15. # -2 Round2_barcodes_new4.txt \
  16. # -3 Round3_barcodes_new4.txt \
  17. # -f SRR6750041_1_smalltest.fastq \
  18. # -r SRR6750041_2_smalltest.fastq \
  19. # -o results \
  20. # -t 8000 \
  21. # -g 100000 \
  22. # -a star \
  23. # -x /path/to/star/genome/index/folder/GRCm38/ \
  24. # -y /path/to/matching/genome/annotation/gtf/GRCm38.gtf \
  25. # -k /path/to/kallisto/index/.idx/ \
  26. # -i /path/to/kallisto/index/.fasta
  27. ################/media/bachar.d/ec530b5a-02c3-4ebe-8b79-8d8a7fc98220/MirCos_splitSeq/SPLiT-Seq_demultiplexing_annotation_pipeline/New_Script_Test/results_b
  28. # Dependencies #
  29. ################
  30. # Python3 must be installed and accessible as "python" from your system's path
  31. type python &>/dev/null || { echo "ERROR python is not installed or is not accessible from the PATH as python"; exit 1; }
  32. # UMI_Tools must be installed and accessible from the PATH as "umi_tools"
  33. type umi_tools &>/dev/null || { echo "ERROR umi_tools is not installed or is not accessible from the PATH as umi_tools"; exit 1; }
  34. # parallel must be installed and accessible from the path as "parallel"
  35. type parallel &>/dev/null || { echo "ERROR parallel is not installed or is not accessible from the PATH as parallel"; exit 1; }
  36. # Not to get the error "/bin/ls: Argument list too long"
  37. # So, a solution is to increase the amount of space available for the stack.
  38. # https://unix.stackexchange.com/questions/45583/argument-list-too-long-how-do-i-deal-with-it-without-changing-my-command
  39. ulimit -s 65536
  40. ###########################
  41. ### Manually Set Inputs ###
  42. ###########################
  43. export NUMCORES="4"
  44. export VERSION="fast"
  45. export ERRORS="1"
  46. export MINREADS="10"
  47. export ROUND1="Round1_barcodes_new5.txt"
  48. export ROUND2="Round2_barcodes_new4.txt"
  49. export ROUND3="Round3_barcodes_new4.txt"
  50. export FASTQ_F="SRR6750041_1_smalltest.fastq"
  51. export FASTQ_R="SRR6750041_2_smalltest.fastq"
  52. export OUTPUT_DIR="results_multiThread"
  53. export TARGET_MEMORY="8000"
  54. export GRANULARITY="100000"
  55. export COLLAPSE="true"
  56. export ALIGN="star"
  57. export STARGENOME="/mnt/isilon/davidson_lab/ranum/Tools/STAR_Genomes/mm10/"
  58. #STARGTF="GTF /mnt/isilon/davidson_lab/ranum/Tools/STAR_Genomes/mm10_Raw/Mus_musculus.GRCm38.96.chr.gtf"
  59. export SAF="SAF ../GRCm38_genes.saf"
  60. #KALLISTOINDEXIDX="/mnt/isilon/davidson_lab/ranum/Tools/Kallisto_Index/GRCm38.idx"
  61. #KALLISTOINDEXFASTA="/mnt/isilon/davidson_lab/ranum/Tools/Kallisto_Index/Mus_musculus.GRCm38.cdna.all.fa"
  62. ################################
  63. ### User Inputs Using Getopt ###
  64. ################################
  65. # NOTE on mac systems the options won't work because mac doesnt have the GNU version of getopt by default. GNU getopt can be installed on mac using homebrew. You can do this by running 'brew install gnu_getopt'
  66. # Once gnu_getopt is installed you can run it with using this '/usr/local/Cellar/gnu-getopt/1.1.6/bin/getopt' as the executable in the place of 'getopt' below.
  67. # read the options
  68. TEMP=`getopt -o n:v:e:m:1:2:3:f:r:o:t:g:c:a:x:y:s:k:i --long numcores:,errors:,minreads:,round1barcodes:,round2barcodes:,round3barcodes:,fastqF:,fastqR:,outputdir:,targetMemory:,granularity:,collapseRandomHexamers:,align:,starGenome:,starGTF:,geneAnnotationSAF:,kallistoIndexIDX:,kallistoIndexFASTA: -n 'test.sh' -- "$@"`
  69. eval set -- "$TEMP"
  70. # extract options and their arguments into variables.
  71. echo "Checking options..."
  72. while true ; do
  73. case "$1" in
  74. -n|--numcores)
  75. case "$2" in
  76. "") shift 2 ;;
  77. *) NUMCORES=$2 ; shift 2 ;;
  78. esac ;;
  79. -v|--version)
  80. case "$2" in
  81. "") shift 2 ;;
  82. *) VERSION=$2 ; shift 2 ;;
  83. esac ;;
  84. -e|--errors)
  85. case "$2" in
  86. "") shift 2 ;;
  87. *) ERRORS=$2 ; shift 2 ;;
  88. esac ;;
  89. -m|--minreads)
  90. case "$2" in
  91. "") shift 2 ;;
  92. *) MINREADS=$2 ; shift 2 ;;
  93. esac ;;
  94. -1|--round1barcodes)
  95. case "$2" in
  96. "") shift 2 ;;
  97. *) ROUND1=$2 ; shift 2 ;;
  98. esac ;;
  99. -2|--round2barcodes)
  100. case "$2" in
  101. "") shift 2 ;;
  102. *) ROUND2=$2 ; shift 2 ;;
  103. esac ;;
  104. -3|--round3barcodes)
  105. case "$2" in
  106. "") shift 2 ;;
  107. *) ROUND3=$2 ; shift 2 ;;
  108. esac ;;
  109. -f|--fastqF)
  110. case "$2" in
  111. "") shift 2 ;;
  112. *) FASTQ_F=$2 ; shift 2 ;;
  113. esac ;;
  114. -r|--fastqR)
  115. case "$2" in
  116. "") shift 2 ;;
  117. *) FASTQ_R=$2 ; shift 2 ;;
  118. esac ;;
  119. -o|--outputdir)
  120. case "$2" in
  121. "") shift 2 ;;
  122. *) OUTPUT_DIR=$2 ; shift 2 ;;
  123. esac ;;
  124. -t|--targetMemory)
  125. case "$2" in
  126. "") shift 2;;
  127. *) TARGET_MEMORY=$2 ; shift 2 ;;
  128. esac ;;
  129. -g|--granularity)
  130. case "$2" in
  131. "") shift 2;;
  132. *) GRANULARITY=$2 ; shift 2 ;;
  133. esac ;;
  134. -c|--collapseRandomHexamers)
  135. case "$2" in
  136. "") shift 2;;
  137. *) COLLAPSE=$2 ; shift 2 ;;
  138. esac ;;
  139. -a|--align)
  140. case "$2" in
  141. "") shift 2;;
  142. *) ALIGN=$2 ; shift 2 ;;
  143. esac ;;
  144. -x|--starGenome)
  145. case "$2" in
  146. "") shift 2;;
  147. *) STARGENOME=$2 ; shift 2 ;;
  148. esac ;;
  149. -y|--starGTF)
  150. case "$2" in
  151. "") shift 2;;
  152. *) STARGTF=$2 ; shift 2 ;;
  153. esac ;;
  154. -s|--geneAnnotationSAF)
  155. case "$2" in
  156. "") shift 2;;
  157. *) SAF=$2 ; shift 2 ;;
  158. esac ;;
  159. -k|--kallistoIndexIDX)
  160. case "$2" in
  161. "") shift 2;;
  162. *) KALLISTOINDEXIDX=$2 ; shift 2 ;;
  163. esac ;;
  164. -i|--kallistoIndexFASTA)
  165. case "$2" in
  166. "") shift 2;;
  167. *) KALLISTOINDEXFASTA=$2 ; shift 2 ;;
  168. esac ;;
  169. --) shift ; break ;;
  170. *) echo "Internal error!" ; exit 1 ;;
  171. esac
  172. done
  173. ###############################
  174. ### Write Input Args To Log ###
  175. ###############################
  176. # Print the arguments provided as input to splitseqdemultiplex.sh
  177. echo "splitseqdemultiplex.sh has been run with the following input arguments"
  178. echo "numcores = $NUMCORES"
  179. echo "errors = $ERRORS"
  180. echo "minreadspercell = $MINREADS"
  181. echo "round1_barcodes = $ROUND1"
  182. echo "round2_barcodes = $ROUND2"
  183. echo "round3_barcodes = $ROUND3"
  184. echo "fastq_f = $FASTQ_F"
  185. echo "fastq_r = $FASTQ_R"
  186. echo "targetMemory = $TARGET_MEMORY"
  187. echo "granularity = $GRANULARITY"
  188. echo "collapseRandomHexamers = $COLLAPSE"
  189. echo "align = $ALIGN"
  190. echo "starGenome = $STARGENOME"
  191. echo "starGTF = $STARGTF"
  192. echo "geneAnnotationSAF = $SAF"
  193. echo "kallistoIndexIDX = $KALLISTOINDEXIDX"
  194. echo "kallistoIndexFASTA = $KALLISTOINDEXFASTA"
  195. #if [ $COLLAPSE = true ]
  196. #then
  197. #ROUND1="Round1_barcodes_new5.txt"
  198. #fi
  199. if [ $VERSION = fast ]
  200. then
  201. #######################################
  202. # STEP 1: Demultiplex Using Barcodes #
  203. #######################################
  204. # Generate a progress message
  205. now=$(date '+%Y-%m-%d %H:%M:%S')
  206. echo "Beginning STEP1: Demultiplex using barcodes. Current time : $now"
  207. # Demultiplex the fastqr file using barcodes
  208. mkdir $OUTPUT_DIR
  209. # Set up a function to parallelize the Demultiplex Using Barcodes Step
  210. linesInInputFastq=$(wc -l < $FASTQ_R)
  211. num_linesPerSplitFastq=$(expr $linesInInputFastq / $NUMCORES)
  212. #split --lines=${num_linesPerSplitFastq} $FASTQ_R split_fastq_R_
  213. #split --lines=${num_linesPerSplitFastq} $FASTQ_F split_fastq_F_
  214. head -n 100 $FASTQ_R > position_learner_fastqr.fastq
  215. split --number="l/$NUMCORES" $FASTQ_R split_fastq_R_
  216. split --number="l/$NUMCORES" $FASTQ_F split_fastq_F_
  217. #my_func() {
  218. # python InDevOptimizations/DemultiplexUsingBarcodes_New_V1.py -f "split_fastq_F$1" -r "split_fastq_R$1" -b $GRANULARITY -o $OUTPUT_DIR -e $ERRORS -p -t $MINREADS
  219. # }
  220. #export -f my_func
  221. #ls split_fastq_F* | awk -F "split_fastq_F" '{print $1}' | parallel my_func {}
  222. ls split_fastq_F* | awk -F "split_fastq_F" '{print $2}' | parallel "python InDevOptimizations/DemultiplexUsingBarcodes_New_V1.py -f split_fastq_F{} -r split_fastq_R{} -b $GRANULARITY -o $OUTPUT_DIR -e $ERRORS -p -t $MINREADS"
  223. #python InDevOptimizations/DemultiplexUsingBarcodes_New_V1.py -f $FASTQ_F -r $FASTQ_R -b $GRANULARITY -o $OUTPUT_DIR -e $ERRORS -p -t $MINREADS
  224. #--minreads $MINREADS --round1barcodes $ROUND1 --round2barcodes $ROUND2 --round3barcodes $ROUND3 --fastqr $FASTQ_R --errors $ERRORS --outputdir $OUTPUT_DIR --targetMemory $TARGET_MEMORY --granularity $GRANULARITY
  225. #echo "$(ls $OUTPUT_DIR/*.fastq | wc -l) results files 'cells' were demultiplexed from the input .fastq file"
  226. rm position_learner_fastqr.fastq
  227. rm split_fastq_*
  228. ###########################
  229. # STEP 5: Perform Mapping #
  230. ###########################
  231. # generate batch file
  232. now=$(date '+%Y-%m-%d %H:%M:%S')
  233. echo "Beginning STEP5: Performing Mapping. Current time : $now"
  234. if [ $ALIGN = star ]
  235. then
  236. pushd $OUTPUT_DIR
  237. # Run alignment of merged .fastq file using STAR
  238. STAR --runThreadN $NUMCORES \
  239. --readFilesIn MergedCells_1.fastq \
  240. --outFilterMismatchNoverLmax 0.05 \
  241. --genomeDir $STARGENOME \
  242. --alignIntronMax 20000 \
  243. --outSAMtype BAM SortedByCoordinate
  244. #cp /media/bachar.d/ec530b5a-02c3-4ebe-8b79-8d8a7fc98220/MirCos_splitSeq/SPLiT-Seq_demultiplexing-master/results/*.bam /media/bachar.d/ec530b5a-02c3-4ebe-8b79-8d8a7fc98220/MirCos_splitSeq/SPLiT-Seq_demultiplexing-master/
  245. if [[ $(echo "$SAF" | awk '{print $1}') = SAF ]]
  246. then
  247. countsMode=$(echo "$SAF" | awk '{print $1}')
  248. countsFile=$(echo "$SAF" | awk '{print $2}')
  249. #echo $countsFile
  250. #echo $countsMode
  251. # Assign reads to genes
  252. # Removed -M parameters for excluding multimapping reads
  253. featureCounts -F $countsMode \
  254. -a $countsFile \
  255. -o gene_assigned \
  256. -R BAM Aligned.sortedByCoord.out.bam \
  257. -T $NUMCORES \
  258. -M
  259. else
  260. countsMode=$(echo "$STARGTF" | awk '{print $1}')
  261. countsFile=$(echo "$STARGTF" | awk '{print $2}')
  262. #echo "Hello world "
  263. #echo $countsFile
  264. #echo "Hello world "
  265. #echo $countsMode
  266. # Removed -M parameters for excluding multimapping reads
  267. featureCounts -F $countsMode \
  268. -a $countsFile \
  269. -o gene_assigned \
  270. -R BAM Aligned.sortedByCoord.out.bam \
  271. -T $NUMCORES \
  272. -M
  273. fi
  274. samtools sort Aligned.sortedByCoord.out.bam.featureCounts.bam -o assigned_sorted.bam
  275. samtools index assigned_sorted.bam
  276. # Count UMIs per gene per cell
  277. umi_tools count --wide-format-cell-counts --per-gene --gene-tag=XT --assigned-status-tag=XS --per-cell -I assigned_sorted.bam -S counts.tsv.gz
  278. popd
  279. fi
  280. fi
  281. if [ $VERSION = merged ]
  282. then
  283. #######################################
  284. # STEP 1: Demultiplex Using Barcodes #
  285. #######################################
  286. # Generate a progress message
  287. now=$(date '+%Y-%m-%d %H:%M:%S')
  288. echo "Beginning STEP1: Demultiplex using barcodes. Current time : $now"
  289. # Demultiplex the fastqr file using barcodes
  290. python demultiplex_using_barcodes.py --minreads $MINREADS --round1barcodes $ROUND1 --round2barcodes $ROUND2 --round3barcodes $ROUND3 --fastqr $FASTQ_R --errors $ERRORS --outputdir $OUTPUT_DIR --targetMemory $TARGET_MEMORY --granularity $GRANULARITY
  291. echo "$(ls $OUTPUT_DIR/*.fastq | wc -l) results files 'cells' were demultiplexed from the input .fastq file"
  292. ##########################################################################
  293. # STEP 2: Collapse OligoDT and RandomHexamer Barcodes from the same well #
  294. ##########################################################################
  295. # Generate a progress message
  296. now=$(date '+%Y-%m-%d %H:%M:%S')
  297. echo "Beginning STEP2: Collapse OligoDT and RandomHexamer Barcodes from the same well. Current time : $now"
  298. if [ $COLLAPSE = true ]
  299. then
  300. # Bash script is painfully slow
  301. # Replaced by python script, which is 100 times faster then the bash script
  302. #bash Collapse_RanHex_Odt.sh
  303. python Collapse_RanHex_Odt.py
  304. fi
  305. echo "after collapsing OligoDT and RandomHexamer Barcodes, $(ls $OUTPUT_DIR/*.fastq | wc -l) results files 'cells' remain."
  306. ##########################################################
  307. # STEP 3: For every cell find matching paired end reads #
  308. ##########################################################
  309. # Generate a progress message
  310. now=$(date '+%Y-%m-%d %H:%M:%S')
  311. echo "Beginning STEP3: Finding read mate pairs. Current time : $now"
  312. # Now we need to collect the other read pair. To do this we can collect read IDs from the $OUTPUT_DIR files we generated in step one.
  313. # Generate an array of cell filenames
  314. python matepair_finding.py --input $OUTPUT_DIR --fastqf $FASTQ_F --output $OUTPUT_DIR --targetMemory $TARGET_MEMORY --granularity $GRANULARITY
  315. ########################
  316. # STEP 4: Extract UMIs #
  317. ########################
  318. # Generate a progress message
  319. now=$(date '+%Y-%m-%d %H:%M:%S')
  320. echo "Beginning STEP4: Extracting UMIs. Current time : $now"
  321. # Implement new method for umi and cell barcode extraction
  322. pushd $OUTPUT_DIR
  323. #parallel python ../Extract_BC_UMI.py -R {} -F {}-MATEPAIR ::: $(ls *.fastq)
  324. # Modified the original parallel command to avoid TOO many arguments error
  325. ls | grep '\.fastq$' | parallel python ../Extract_BC_UMI.py -R {} -F {}-MATEPAIR
  326. cat *_1.fastq > MergedCells
  327. #parallel rm {} ::: $(ls *fastq*)
  328. # This command has been changed to avoid Arguments Lists Too Long error
  329. # Above command Gives error Argument List Too Long using parallel rm command, and the pipeline halts
  330. # So, we use either loop to delete the files one by one
  331. # for i in *.fastq*;do rm "$i";done
  332. # Or, we can increase stack limit by command ulimit -s 65536, done at the beginning of this script
  333. # https://unix.stackexchange.com/questions/45583/argument-list-too-long-how-do-i-deal-with-it-without-changing-my-command
  334. # and we modify the original parallel commands like below, to avoid too many parameters errors
  335. #parallel rm {} ::: $(ls *fastq*)
  336. ls | grep '\.fastq*' | parallel rm {}
  337. mv MergedCells MergedCells_1.fastq
  338. popd
  339. ###########################
  340. # STEP 5: Perform Mapping #
  341. ###########################
  342. # generate batch file
  343. now=$(date '+%Y-%m-%d %H:%M:%S')
  344. echo "Beginning STEP5: Performing Mapping. Current time : $now"
  345. if [ $ALIGN = star ]
  346. then
  347. pushd $OUTPUT_DIR
  348. # Run alignment of merged .fastq file using STAR
  349. STAR --runThreadN $NUMCORES \
  350. --readFilesIn MergedCells_1.fastq \
  351. --outFilterMismatchNoverLmax 0.05 \
  352. --genomeDir $STARGENOME \
  353. --alignIntronMax 20000 \
  354. --outSAMtype BAM SortedByCoordinate
  355. #cp /media/bachar.d/ec530b5a-02c3-4ebe-8b79-8d8a7fc98220/MirCos_splitSeq/SPLiT-Seq_demultiplexing-master/results/*.bam /media/bachar.d/ec530b5a-02c3-4ebe-8b79-8d8a7fc98220/MirCos_splitSeq/SPLiT-Seq_demultiplexing-master/
  356. if [ $(echo "$SAF" | awk '{print $1}') = SAF ]
  357. then
  358. countsMode=$(echo "$SAF" | awk '{print $1}')
  359. countsFile=$(echo "$SAF" | awk '{print $2}')
  360. #echo $countsFile
  361. #echo $countsMode
  362. # Assign reads to genes
  363. # Removed -M parameters for excluding multimapping reads
  364. featureCounts -F $countsMode \
  365. -a $countsFile \
  366. -o gene_assigned \
  367. -R BAM Aligned.sortedByCoord.out.bam \
  368. -T $NUMCORES \
  369. -M
  370. else
  371. countsMode=$(echo "$STARGTF" | awk '{print $1}')
  372. countsFile=$(echo "$STARGTF" | awk '{print $2}')
  373. #echo "Hello world "
  374. #echo $countsFile
  375. #echo "Hello world "
  376. #echo $countsMode
  377. # Removed -M parameters for excluding multimapping reads
  378. featureCounts -F $countsMode \
  379. -a $countsFile \
  380. -o gene_assigned \
  381. -R BAM Aligned.sortedByCoord.out.bam \
  382. -T $NUMCORES \
  383. -M
  384. fi
  385. samtools sort Aligned.sortedByCoord.out.bam.featureCounts.bam -o assigned_sorted.bam
  386. samtools index assigned_sorted.bam
  387. # Count UMIs per gene per cell
  388. umi_tools count --wide-format-cell-counts --per-gene --gene-tag=XT --assigned-status-tag=XS --per-cell -I assigned_sorted.bam -S counts.tsv.gz
  389. popd
  390. fi
  391. fi
  392. if [ $VERSION = split ]
  393. then
  394. #######################################
  395. # STEP 1: Demultiplex Using Barcodes #
  396. #######################################
  397. # Generate a progress message
  398. now=$(date '+%Y-%m-%d %H:%M:%S')
  399. echo "Beginning STEP1: Demultiplex using barcodes. Current time : $now"
  400. # Demultiplex the fastqr file using barcodes
  401. python demultiplex_using_barcodes.py --minreads $MINREADS --round1barcodes $ROUND1 --round2barcodes $ROUND2 --round3barcodes $ROUND3 --fastqr $FASTQ_R --errors $ERRORS --outputdir $OUTPUT_DIR --targetMemory $TARGET_MEMORY --granularity $GRANULARITY
  402. ##########################################################
  403. # STEP 2: For every cell find matching paired end reads #
  404. ##########################################################
  405. # Generate a progress message
  406. now=$(date '+%Y-%m-%d %H:%M:%S')
  407. echo "Beginning STEP2: Finding read mate pairs. Current time : $now"
  408. # Now we need to collect the other read pair. To do this we can collect read IDs from the $OUTPUT_DIR files we generated in step one.
  409. # Generate an array of cell filenames
  410. python matepair_finding.py --input $OUTPUT_DIR --fastqf $FASTQ_F --output $OUTPUT_DIR --targetMemory $TARGET_MEMORY --granularity $GRANULARITY
  411. ########################
  412. # STEP 3: Extract UMIs #
  413. ########################
  414. # Generate a progress message
  415. now=$(date '+%Y-%m-%d %H:%M:%S')
  416. echo "Beginning STEP3: Extracting UMIs. Current time : $now"
  417. rm -r $OUTPUT_DIR-UMI
  418. mkdir $OUTPUT_DIR-UMI
  419. # Parallelize UMI extraction
  420. {
  421. ls $OUTPUT_DIR | grep \.fastq$ | parallel -j $NUMCORES -k "umi_tools extract -I $OUTPUT_DIR/{} --read2-in=$OUTPUT_DIR/{}-MATEPAIR --bc-pattern=NNNNNNNNNN --log=processed.log --read2-out=$OUTPUT_DIR-UMI/{}"
  422. } &> /dev/null
  423. #################################
  424. # STEP 4: Collect Summary Stats #
  425. #################################
  426. # Print the number of lines and barcode ID for each cell to a file
  427. echo "$(wc -l $OUTPUT_DIR-UMI/*.fastq)" | sed '$d' | sed 's/$OUTPUT_DIR-UMI\///g' > linespercell.txt
  428. ###########################
  429. # STEP 5: Perform Mapping #
  430. ###########################
  431. # generate batch file
  432. if [ $ALIGN = kallisto ]
  433. then
  434. # generate batch file
  435. rm batch.txt
  436. for file in $(ls results-UMI/); do
  437. echo "$(echo $file | sed 's|.fastq||g')" "$(echo $file | sed 's|fastq|umi|g')" "$(echo $file)" >> batch.txt
  438. python align_kallisto.py -F results-UMI/$file
  439. done
  440. pushd results-UMI
  441. mkdir ../kallisto_output
  442. kallisto pseudo -i $KALLISTOINDEXIDX -o ../kallisto_output --single --umi -b ../batch.txt
  443. cp matrix.cells results4_colNames_cellIDs.txt
  444. popd
  445. pushd kallisto_output
  446. python ../prep_TCC_matrix.py -T matrix.tsv -E matrix.ec -O results -I $KALLISTOINDEXFASTA -G geneIDs
  447. popd
  448. fi
  449. number_of_cells=$(ls -1 "$OUTPUT_DIR-UMI" | wc -l)
  450. echo "a total of $number_of_cells cells were demultiplexed from the input .fastq"
  451. fi
  452. #All finished
  453. #number_of_cells=$(ls -1 "$OUTPUT_DIR-UMI" | wc -l)
  454. now=$(date '+%Y-%m-%d %H:%M:%S')
  455. #echo "a total of $number_of_cells cells were demultiplexed from the input .fastq"
  456. # Re Initialize the stack limit to default
  457. ulimit -s 8192
  458. echo "Current time : $now"
  459. echo "all finished goodbye"

splitseqdemultiplex_0.2.1.sh at commit a8015fb, under MIT · at the source

Overview

Authors: Ashley B Robbins1,2, Paul T Ranum3, Icnelia Huerta-Ocampo1, Michael Kuckyr1,4, Beverly L Davidson1,5
  1. Raymond G. Perelman Center for Cellular and Molecular Therapeutics, Children’s Hospital of Philadelphia, Philadelphia, PA 19104 USA
  2. Neuroscience Graduate Group, Biomedical Graduate Studies Program, University of Pennsylvania, Philadelphia, PA 19104 USA
  3. Latus Bio, N 30th Street, Philadelphia, PA 19104 USA
  4. Cell and Molecular Biology Graduate Group, Biomedical Graduate Studies Program, University of Pennsylvania, Philadelphia, PA 19104 USA
  5. Department of Pathology & Laboratory Medicine, University of Pennsylvania, Philadelphia, PA 19104 USA
Institutions: Children's Hospital of Philadelphia (United States); University of Pennsylvania (United States)
Journal: Molecular neurodegeneration, volume 21, issue 1, article 41
Dates: received 15 January 2026; accepted 23 May 2026; published online 28 May 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1186/s13024-026-00960-2 · PMID 42210302 · PMCID PMC13440178 · OpenAlex W7162674730
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), mouse (organism), other condition (population), systems (subfield)
Methods: Connectivity, Statistics, Smoothing, state filtering, decompositions, Machine learning, fMRI & imaging
Keywords: Huntington’s disease, Neurodegeneration, Neuronal vulnerability, Single-nucleus RNA sequencing, CAG repeat expansion, Striatum, Motor cortex, Omics
MeSH: Corpus Striatum*, Huntington Disease*, Motor Cortex*, Animals, Disease Models, Animal, Disease Progression, Humans, Huntingtin Protein, Mice, Mice, Transgenic, Neurons, Transcriptome (* major topic)
Topic: Genetic Neurodegenerative Diseases (Cellular and Molecular Neuroscience, Neuroscience), according to OpenAlex
Funding: National Human Genome Research Institute (T32 HG000046); CHOP Research Institute; National Institute of Neurological Disorders and Stroke (F31 NS122297); NHGRI NIH HHS (T32 HG000046); NINDS NIH HHS (F31 NS122297); Hereditary Disease Foundation
Citations: cited by 1 paper (Europe PMC); 110 references in the paper
Research resources: RRID:SCR_024672

Abstract

Background: Huntington’s disease (HD) involves progressive corticostriatal dysfunction, yet the temporal dynamics and cell type-specific vulnerability patterns remain incompletely understood. While recent single-cell studies in rapidly progressing models have revealed early developmental and regional changes, temporal profiling distinguishing pathogenic mechanisms from normal aging in full-length HTT models remains lacking. Resolving stage-specific temporal dynamics across interconnected striatal and cortical neuronal populations over protracted time is essential for identifying drivers of cellular dysfunction.

Methods: A temporal single-nucleus transcriptomic atlas was generated from striatum and motor cortex from heterozygous zQ175 knock-in mice at early symptomatic (6 months) and late symptomatic (18 months) stages. This full-length huntingtin model enables staging of progressive circuit dysfunction alongside physiological aging. The high inherited CAG repeat length of the zQ175 model places cells beyond the somatic expansion threshold associated with transcriptional dysregulation and identity erosion in vulnerable human neuronal populations, yet prior to the de-repression crisis and cell loss observed at the most extreme expansions in HD, providing a tractable window into the progressive molecular pathogenic cascade. Genotype-dependent effects were modeled to distinguish cell type-specific signatures of disease mechanisms from age-related and compensatory changes. Integration of weighted gene co-expression and transcription factor regulatory networks with protein-protein interaction databases predicted candidate regulators of stage-specific programs. Findings were validated across human HD datasets and the rapidly progressive R6/2 mouse model.

Results: Temporal gene and network analysis revealed diverging, converging and biphasic patterns of transcriptional changes, distinguishing progressive disease and neuronal identity loss from aging. 21 cell type-specific gene co-expression modules were validated in human HD and R6/2 mice datasets, revealing stage-specific shifts in cellular stress, proteostasis, and synaptic programs. Disease modules enriched for CAG repeat length-dependent genes resolved their temporal progression. Shared vulnerability across cortical and striatal projection neurons implicated epigenetic regulator Zswim6 and splicing factors Rbfox1 and Celf2 in corticostriatal dysfunction. Integrative network analysis identified Foxo1, Neurod2, and Npas2 as stage-specific transcriptional regulators. Cross-species validation established conserved gene regulatory modules in human HD, establishing generalizable cell type-specific gene modules of translational relevance.

Conclusions: This temporally resolved atlas reveals stage-specific transcriptional dynamics of disease-relevant gene expression programs and physiological trajectories in vulnerable neuronal populations. distinguished from aging alone. This work establishes an important framework for understanding the temporal and regional coordination of pathogenic mechanisms, providing molecular insights into stage-specific therapeutic intervention.

Supplementary Information: The online version contains supplementary material available at 10.1186/s13024-026-00960-2.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 7 matches between paragraphs and lines of code.

DavidsonLabCHOP/SPLiT-seq

License: MIT
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: a8015fb890e2129a19b0abfece3e1413a277a3ef, 18 December 2024
Languages: Python (17), Shell (6), C++ (1)
Size: 41 files, 24 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, license file
Not found: CITATION.cff, environment file, tests, continuous integration, documentation
Tools: SAMtools (2 files), STAR (2 files), Subread (featureCounts) (2 files), NumPy (1 file), scikit-learn (1 file), SciPy (1 file)
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
26 files

DavidsonLabCHOP/Robbins_HDsnRNAseq_2025

License: none: the authors keep all their rights
State: the link answers, verified on 28 September 2026
Evidence: files inventoried
Commit: c402d25cc96e1e42deb503a8a0b4832c20c0eaf4, 4 January 2026
Size: 1 file, 0 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 28 September 2026: the link answers
  • 28 September 2026: the link answers
1 file

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 24 scripts, each with its path and the digest of its content;
  • 7 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data availability

The datasets generated during the current study are available in theSequence Read Archive (SRA) under BioProject accession number: PRJNA1337572. All processed count matrices, supplementary datasets and expanded results tables are available on Synapse (Project SynID: syn69978129). Publicly available datasets analyzed in this study can be found on NCBI GEO query. The Human HD Caudate and Putamen snRNA-seq from Lee et al. (30) is available under GEO accession: GSE152058 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE152058), and the R6/2 snRNA-seq study from Lim et al. (51) in striatum is available under GEO accession: GSE180294 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE180294). All code associated with the SPLiT-seq demultiplexing package is available on GitHub at https://github.com/DavidsonLabCHOP/SPLiT-seq. All data processing and analysis code for this study is available on GitHub at https://github.com/DavidsonLabCHOP/Robbins_HDsnRNAseq_2025. An interactive web portal is available at https://davidsonlabchop.github.io/zQ175-Brain-Atlas-App/. Users can explore gene, module and regulon expression and download processed data.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 28 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 8 keywords, 12 MeSH terms, 6 funders, 110 references, 1 RRID.

Cite

This paper

Robbins, A. B., Ranum, P. T., Huerta-Ocampo, I., Kuckyr, M., & Davidson, B. L. (2026). Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction. Molecular neurodegeneration, 21(1), 41. https://doi.org/10.1186/s13024-026-00960-2

BibTeX

@article{robbins2026temporal,
author = {Robbins, Ashley B and Ranum, Paul T and Huerta-Ocampo, Icnelia and Kuckyr, Michael and Davidson, Beverly L},
title = {{Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction}},
journal = {Molecular neurodegeneration},
year = {2026},
month = may,
volume = {21},
number = {1},
pages = {41},
publisher = {BMC},
issn = {1750-1326},
doi = {10.1186/s13024-026-00960-2},
url = {https://doi.org/10.1186/s13024-026-00960-2},
pmid = {42210302},
pmcid = {PMC13440178}
}

RIS

TY - JOUR
AU - Robbins, Ashley B
AU - Ranum, Paul T
AU - Huerta-Ocampo, Icnelia
AU - Kuckyr, Michael
AU - Davidson, Beverly L
TI - Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction
T2 - Molecular neurodegeneration
J2 - Mol Neurodegener
PY - 2026
DA - 2026/05/28
VL - 21
IS - 1
SP - 41
SN - 1750-1326
PB - BMC
DO - 10.1186/s13024-026-00960-2
UR - https://doi.org/10.1186/s13024-026-00960-2
LA - en
ER -

CSL-JSON

{
"id": "10.1186/s13024-026-00960-2",
"type": "article-journal",
"title": "Temporal single-cell atlas of full-length Huntington's disease mouse model defines stage-specific signatures of corticostriatal dysfunction",
"container-title": "Molecular neurodegeneration",
"author": [
{
"family": "Robbins",
"given": "Ashley B"
},
{
"family": "Ranum",
"given": "Paul T"
},
{
"family": "Huerta-Ocampo",
"given": "Icnelia"
},
{
"family": "Kuckyr",
"given": "Michael"
},
{
"family": "Davidson",
"given": "Beverly L"
}
],
"container-title-short": "Mol Neurodegener",
"volume": "21",
"issue": "1",
"page": "41",
"DOI": "10.1186/s13024-026-00960-2",
"PMID": "42210302",
"PMCID": "PMC13440178",
"ISSN": "1750-1326",
"publisher": "BMC",
"URL": "https://doi.org/10.1186/s13024-026-00960-2",
"language": "en",
"issued": {
"date-parts": [
[
2026,
5,
28
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1186/s13148-026-02082-4 [code]
DNA methylation profiling in Huntington's disease reveals disease associated changes in the striatum.
Journal: Clinical epigenetics
In common: systems, genetics / omics, other condition, 8 references
[2] doi:10.1038/s41598-026-54412-4 [code]
Single-cell transcriptomics and mouse model phenotyping for biomarker screen of peripheral blood in Huntington's disease.
Journal: Scientific reports
In common: NCBI GEO GSE152058, genetics / omics, other condition, mouse, 4 references
[3] doi:10.1177/18796397261443137 [code]
Towards AI-driven prediction of &lt;i&gt;HTT&lt;/i&gt; CAG size in super-expanded human spiny projection neurons from Huntington disease donors.
Journal: Journal of Huntington's disease
In common: genetics / omics, other condition, 7 references
[4] doi:10.1038/s41467-026-73305-8 [code]
Comparative analysis of the cellular landscape in mammalian striatum.
Journal: Nature communications
In common: SAMtools, NumPy, genetics / omics, mouse, 5 references
[5] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: Subread (featureCounts), STAR, SAMtools, 3 other tools, genetics / omics, other condition, 1 reference
[6] doi:10.1038/s41586-026-10512-9 [code]
Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.
Journal: Nature
In common: Subread (featureCounts), STAR, SAMtools, 2 other tools, mouse, 2 references
[7] doi:10.1038/s44321-026-00459-9
Anle138b ameliorates pathological phenotypes in mouse and cellular models of Huntington's disease.
Journal: EMBO molecular medicine
In common: other condition, mouse, 6 references
[8] doi:10.1126/sciadv.aed2952 [code]
Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.
Journal: Science advances
In common: Subread (featureCounts), STAR, SAMtools, 2 other tools, 2 references
[9] doi:10.1016/j.xgen.2026.101278 [code]
Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures.
Journal: Cell genomics
In common: Subread (featureCounts), STAR, SAMtools, 1 other tool, genetics / omics, 2 references
[10] doi:10.1038/s41467-026-71432-w [code]
MeCP2 gene dosage-dependent neurodevelopmentally restricted defects arise by aberrant activation of cell fate-determining bivalent genes.
Journal: Nature communications
In common: Subread (featureCounts), STAR, SAMtools, 3 other tools, other condition, mouse

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.