Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures.
The 6 matches
- [1] § STAR★Methods › Quantification and statistical analysis › Bioinformatic alignment and feature quantification ↔ Notebooks/A06_mapping_star.ipynb, lines 16–63 · score 1.00 · BySJout, EndToEnd, alignEndsType, alignIntronMax, alignMatesGapMax, alignSJDBoverhangMin
- [2] § STAR★Methods › Quantification and statistical analysis › Bioinformatic alignment and feature quantification ↔ Notebooks/A06_mapping_star.ipynb, lines 545–630 · score 0.90 · pixel distance, optical duplicates, MarkDuplicates, samtools view, score, flags
- [3] § STAR★Methods › Quantification and statistical analysis › Bioinformatic alignment and feature quantification ↔ Notebooks/A04_mapping_bismark.ipynb, lines 425–501 · score 0.85 · pixel distance, optical duplicates, MarkDuplicates, conversion, samtools, mapped
- [4] § STAR★Methods › Quantification and statistical analysis › Bioinformatic alignment and feature quantification ↔ Notebooks/A03_trimming.ipynb, lines 57–196 · score 0.83 · adapter_fasta, cut_right, adapter sequences, snmCTseq, demultiplexing, trimmed
- [5] § STAR★Methods › Quantification and statistical analysis › Bioinformatic alignment and feature quantification ↔ Scripts/A02a_demultiplex_fastq.pl, lines 2–49 · score 0.53 · cell barcodes, demultiplexing, fasta, v0, nuclei, plates
- [6] § STAR★Methods › Quantification and statistical analysis › Enrichment analyses ↔ Notebooks/A00_environment_and_genome_setup.ipynb, lines 437–555 · score 0.51 · GENCODE v40, gtf, kb, downstream, exon, gene
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 1,120 lines · 33 KB · no license · 2 matches
- # %%
- # ## overall commands
- # ## might be cleaner to wait for STAR mapping to finish 100% (run A04a only --> check output)
- # ## in lieu of submitting all of the below, but technically could run all cmds at once
- # # * = job array based on "platenum"
- # # † = job array based on "batchnum" (two rows at a time)
- # qsub Scripts/A06a_star_mapping.sub # †
- # qsub Scripts/A06b_star_filtering.sub # †
- # qsub Scripts/A06c_check_star.sub
- # qsub Scripts/A06d_featurecounts.sub # *
- # qsub Scripts/A06e_star_bam_stats.sub # †
- # %% [markdown]
- # ## (A06a) star mapping
- # %%
- %%bash
- cat > ../Scripts/A06a_star_mapping.sub
- #!/bin/bash
- #$ -cwd
- #$ -o sublogs/A06a_star.$JOB_ID.$TASK_ID
- #$ -j y
- #$ -l h_rt=12:00:00,h_data=8G,exclusive
- #$ -pe shared 8
- #$ -N A06a_star
- #$ -t 1-512
- #$ -hold_jid_ad A03a_trim
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `date `
- echo " "
- # environment init -------------------------------------------------------------
- . /u/local/Modules/default/init/modules.sh # <--
- module load anaconda3 # <--
- conda activate snmCTseq # <--
- export $(cat snmCT_parameters.env | grep -v '^#' | xargs) # <--
- skip_complete=true # <-- for help with incomplete jobs
- overwrite_partial=true # <-- for help with incomplete jobs
- # (most of the) STAR settings
- # originally adapted from ENCODE guidelines
- star_params="--runThreadN 8 \
- --genomeDir ${ref_starfolder} --genomeLoad LoadAndKeep \
- –alignEndsType EndToEnd --outSAMtype BAM Unsorted \
- --outSAMattributes NH HI AS NM MD --outSAMstrandField intronMotif \
- --sjdbOverhang 149 --outFilterType BySJout \
- --outFilterMultimapNmax 20 --alignSJoverhangMin 8 \
- --alignSJDBoverhangMin 1 –outFilterMismatchNmax 999 \
- --outFilterMismatchNoverLmax 0.04 --alignIntronMin 20 \
- --alignIntronMax 1000000 --alignMatesGapMax 1000000 --readFilesCommand zcat "
- # extract target filepaths -----------------------------------------------------
- # helper functions
- query_metadat () {
- awk -F',' -v targetcol="$1" \
- 'NR==1 {
- for (i=1;i<=NF;i++) {
- if ($i==targetcol) {assayout=i; break} }
- print $assayout
- }
- NR>1 {
- print $assayout
- }' ${metadat_well}
- }
- # extract target wells, print values for log
- batchnum=($(query_metadat "batchnum"))
- nwells=${#batchnum[@]}
- target_well_rows=()
- for ((row=1; row<=nwells; row++))
- do
- if [[ "${batchnum[$row]}" == "${SGE_TASK_ID}" ]]
- then
- target_well_rows+=($row)
- fi
- done
- # filepaths associated with target rows in well-level metadata -----------------
- wellprefix=($(query_metadat "wellprefix"))
- dir_well=($(query_metadat "A06a_dir_star"))
- # .fastqs for input to PE mapping (properly paired read pairs)
- fastq_r1p=($(query_metadat "A03a_fqgz_paired_R1"))
- fastq_r2p=($(query_metadat "A03a_fqgz_paired_R2"))
- # .fastqs for input to SE mapping, including singletons from trimming & unaligned in PE-mapping
- fastq_r1singletrim=($(query_metadat "A03a_fqgz_singletrim_R1"))
- fastq_r2singletrim=($(query_metadat "A03a_fqgz_singletrim_R2"))
- # temporary/intermediate mapping files -----------------------------------------
- # (expected output generated by STAR, given outname prefixes PE., SE1., SE2.)
- fastq_pe_unmap1=PE.Unmapped.out.mate1
- fastq_pe_unmap2=PE.Unmapped.out.mate2
- bam_pe=PE.Aligned.out.bam
- bam_se1=SE1.Aligned.out.bam
- bam_se2=SE2.Aligned.out.bam
- # run STAR mapping -------------------------------------------------------------
- if [[ ! -s mapping_star ]]
- then
- mkdir mapping_star
- fi
- cd ${dir_proj}/mapping_star
- # load genome index [5~10 min]
- # creates some apparent .log, .sam out despite just loading genome
- # so putting in mapping_star to keep these files in one place
- STAR --runThreadN 8 --genomeDir ${ref_starfolder} --genomeLoad LoadAndExit
- # loop through each well to map ------------------------------------------------
- for row in ${target_well_rows[@]}
- do
- # check directory/prior mapping --------------------------------------------
- cd ${dir_proj}
- # check for existing mapping output
- # if final outputs exist, skip; else run mapping .bam
- if [[ -s ${dir_well[$row]}/${bam_pe} \
- && -s ${dir_well[$row]}/${bam_se1} \
- && -s ${dir_well[$row]}/${bam_se2} \
- && "${skip_complete}"=="true" ]]
- then
- echo -e "final aligned .bams for '${wellprefix[$row]}' already exist. skipping this well.'"
- else
- echo -e "\n\napplying STAR to '${wellprefix[$row]}'...\n\n"
- # remove old directory if one exists to deal with incomplete files
- # albeit the only issue i've seen are .bai and .tbi indices
- # (these often are not overwritten by software in the pipeline,
- # resulting in "index is older than file" errors later on)
- if [[ -e ${dir_well[$row]} && "${overwrite_partial}" == "true" ]]
- then
- echo -e "\n\nWARNING: folder for '${wellprefix[$row]}' exists, but not its final .bam alignments."
- echo "because overwrite_partial=true, deleting the directory and re-mapping."
- rm -rf ${dir_well[$row]}
- fi
- mkdir ${dir_well[$row]}
- cd ${dir_well[$row]}
- # run alignments -----------------------------------------------------------
- # in: .fastqs from trimming: four .fastqs,
- # properly paired ($fastq_r2p, $fastq_r1p) and
- # trimming singletons ($fastq_r1singletrim, $fastq_r2singletrim)
- # out: - paired-end, single-end .bam alignments out (${bam_pe}, ${bam_se1}, ${bam_se2})
- # - key log files (e.g., mapping rate)
- # --------------------------------------------------------------------------
- # (i) paired-end mapping [<1-3 min] ----------------------------------------
- # assumptions: pairs that map ambiguously in paired-end mode should be discarded
- STAR ${star_params} \
- --outFileNamePrefix PE. \
- --readFilesIn ${dir_proj}/${fastq_r1p[$row]} ${dir_proj}/${fastq_r2p[$row]} \
- --outReadsUnmapped Fastx
- # .fq --> .fq.gz for future storage/help match STAR's expected input type [<1-2 min]
- bgzip ${fastq_pe_unmap1}
- bgzip ${fastq_pe_unmap2}
- # (ii.) single-end, R1 [<1-3 min] ------------------------------------------
- # includes Read 1 singletons from trimming and STAR mapping in (i)
- STAR ${star_params} \
- --outFileNamePrefix SE1. \
- --readFilesIn ${dir_proj}/${fastq_r1singletrim[$row]},${fastq_pe_unmap1}.gz
- # (iii.) single-end, R2 [<1-3 min] -----------------------------------------
- # includes Read 2 singletons from trimming and STAR mapping in (i)
- STAR ${star_params} \
- --outFileNamePrefix SE2. \
- --readFilesIn ${dir_proj}/${fastq_r2singletrim[$row]},${fastq_pe_unmap2}.gz
- # (iv.) optional clean-up (comment out as desired)
- # *.Log.final.out contains mapping rate; other *.out
- # have STAR internals that can be discarded
- rm -rf *_STARtmp
- rm -rf *SJ.out.tab
- rm "${fastq_pe_unmap1}.gz" "${fastq_pe_unmap2}.gz"
- fi
- done
- # unload genome from mem at end ------------------------------------------------
- STAR --genomeDir ${ref_starfolder} --genomeLoad Remove
- echo -e "\n\n'A06a_star_mapping' completed.\n\n"
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `date `
- echo " "
- # %% [markdown]
- # ## (A06b) star filtering
- # %%
- %%bash
- cat > ../Scripts/A06b_classify_mCT_reads_STAR.pl
- #!/usr/bin/perl -w
- use strict;
- # A06b_classify_mCT_reads_STAR.pl, v0.2 =====================================================
- # based loosely on original perl script written by Dr. Chongyuan Luo (@luogenomics)
- # modifications by Choo Liu (@chooliu):
- # - readability/documentation
- # - changes in logic for paired-end mapping (fwd/rev strand) & more efficient parsing of MD:Z flag
- # incl looking at only C/G positions in read vs. ref changes, calculating # cytosines, etc
- # inputs: - .sam file from STAR (compatible with single-end or paired-end alignments)
- # outputs: - "_annotations" .tsv recording each alignment's:
- # number of cytosines, mC/C fraction, and call (DNA, RNA, ambiguous)
- # this annotation file is subsequently appended to the. bam to keep RNA reads
- # typical usage: perl A06b_classify_mCT_reads_STAR.pl alignments.sam
- # =====================================================================================
- # reading in .sam file
- my ($samfile)=($ARGV[0]);
- my @sample; my @samline;
- my $read; my @read; my @ref; my $ref; my @md; my $dir; my $unmch;
- # determining read classification
- my $filter_num_CH=3;
- my $filter_frac_mCHmin=0.5; # DNA (mCH/CH <0.5)
- my $filter_frac_mCHmax=0.9; # RNA (mCH/CH >0.9)
- my $mch_fraction; my $totalch; my $totcpg; my @call;
- # load .sam file, loop through lines in .sam
- # exports info on each read to "(input .sam file name)_annotations"
- # for each line in .sam file, --------------------------------------------------------
- # count unmethylated and methylated cytosines in CHG and CHH context based on XR-tag
- # notes:
- # - skip header header rows (starting with @)
- # - assumes the XR-tag is in column 10 (STAR output), $samline[9]
- # - currently ignores indels
- # - default settings: keep reads with >=3 CHNs and mCHN/CHN fraction >=0.9 [*]
- # at present, we only retain call="RNA" and discard all others
- open sam_in, "$samfile" or die $!;
- open sam_annotations, ">$samfile\_annotations" or die $!;
- while (<sam_in>)
- {
- chop $_;
- if (substr($_,0,1) eq '@') { print sam_annotations "\n"; }
- else
- {
- @samline = split(/\t/,$_);
- $read = $samline[9];
- @read = split(//,$read);
- @ref = @read;
- @md=split( /(\d+)/ , $samline[15]);
- # check mapping orientation relative to reference genome ----------------------------
- # dir=1 if fwd (99, 163, 0, 73), =0 if reverse (147, 83, 16, 89)
- # flaw in script is that these numbers manually specified--if see strange results,
- # check mapping output for other flags before filtering / add interpreter
- $dir = 0 + ( ($samline[1] eq 99) or ($samline[1] eq 163 ) or ($samline[1] eq 0) or ($samline[1] eq 73) );
- # compare read to reference genome --------------------------------------------------
- # split MD:Z: flag by numbers then examine resulting length
- # e.g., MD:Z:10A0B121 --> MD:Z: 10 A 0 B 121 has $num_md_features=5
- # assume fully methylated if perfect match to genomic ref
- # (if MD stores a single number indicating no changes, $#md = 1)
- my $unmch = 0;
- my $num_md_features=$#md;
- if ($num_md_features == 1) { }
- # <-- start ELSE for num_md_features != 1
- # otherwise, attempt to reconstruct reference sequence,
- # looping through positions in read where the read != ref nucleotide
- else {
- my $pos=0;
- for (my $i=1; $i<=$num_md_features; $i += 2) {
- # loop through bases where there are changes
- $pos = $pos + @md[$i];
- # do nothing if starts with "^" (deletion)
- # modify @ref sequence if differences
- if ( (rindex("@md[$i+1]", "^", 0) == 0) or ($i+1 > $num_md_features)) { }
- else {
- @ref[$pos] = @md[$i + 1];
- }
- # increment base by 1
- $pos += 1;
- }
- # tabulate # unmethylated cytosines ---------------------------------------------------
- # again looping through each position where read != ref
- # (although same loop as above, re-loop because potential cases where adjacent bases differ btwn ref & read)
- my $pos=0;
- for (my $i=1; $i<=$num_md_features; $i += 2) {
- $pos = $pos + @md[$i];
- # if read maps to forward strand of genome
- # check if at a CH-site, and unmethylated cytosine converted to "T"
- if ( ($dir eq 1) and
- (@ref[$pos] eq "C") and (@ref[$pos + 1] ne "G") and (@read[$pos] eq "T") ) {
- $unmch += 1;
- }
- # if read maps to reverse strand of genome
- # STAR rev compliments the $read sequence, so cytosines are represented by "G"
- # check if at a CH-site, and unmethylated cytosine converted to "A"
- # (note: mCH underestimated in niche case where dir=0 &
- # first base is cytosine, as @ref[$pos - 1] is undefined)
- if ( ($dir eq 0) and
- (@ref[$pos] eq "G") and (@ref[$pos - 1] ne "C") and (@read[$pos] eq "A") ) {
- $unmch += 1;
- }
- $pos += 1;
- }
- } # <-- end ELSE statement for num_md_features != 1
- $ref = join('', @ref);
- # count cytosines in CH-context --------------------------------------------------
- # note: since done via regex, faster to count [# C] and subtract [# CG] vs searching wild flag C[ACT]
- if ($dir eq 1) {
- my @totalch = $ref =~ /C/g;
- $totalch = scalar @totalch;
- my @totcpg = $ref =~ /CG/g;
- $totcpg = scalar @totcpg;
- }
- if ($dir eq 0) {
- my @totalch = $ref =~ /G/g;
- $totalch = scalar @totalch;
- my @totcpg = $ref =~ /CG/g;
- $totcpg = scalar @totcpg;
- }
- my $totalch = $totalch - $totcpg;
- # classify each read into modalities ----------------------------------------------
- if ($totalch==0) { # avoid div by zero error
- $mch_fraction=-999;
- @call="amb";
- } else {
- $mch_fraction = 1 - ($unmch/$totalch);
- if (($totalch>=$filter_num_CH) and ($mch_fraction < $filter_frac_mCHmin)) { @call="DNA"; }
- elsif (($totalch>=$filter_num_CH) and ($mch_fraction >= $filter_frac_mCHmax)) { @call="RNA"; }
- else { @call="amb"; } # exclude (low # CH or ambiguous mCH between 0.5-0.9)
- }
- print sam_annotations "${totalch}\t${mch_fraction}\t@call\n";
- }
- }
- close bam_in;
- # %%
- %%bash
- cat > ../Scripts/A06b_star_filtering.sub
- #!/bin/bash
- #$ -cwd
- #$ -o sublogs/A06b_starfilt.$JOB_ID.$TASK_ID
- #$ -j y
- #$ -l h_rt=6:00:00,h_data=16G
- #$ -N A06b_starfilt
- #$ -t 1-512
- #$ -hold_jid_ad A06a_star
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `date `
- echo " "
- # environment init -------------------------------------------------------------
- . /u/local/Modules/default/init/modules.sh # <--
- module load anaconda3 # <--
- conda activate snmCTseq # <--
- export $(cat snmCT_parameters.env | grep -v '^#' | xargs) # <--
- skip_complete=true # <-- for help with incomplete jobs
- # extract target filepaths -----------------------------------------------------
- # helper functions
- query_metadat () {
- awk -F',' -v targetcol="$1" \
- 'NR==1 {
- for (i=1;i<=NF;i++) {
- if ($i==targetcol) {assayout=i; break} }
- print $assayout
- }
- NR>1 {
- print $assayout
- }' ${metadat_well}
- }
- # extract target wells, print values for log
- batchnum=($(query_metadat "batchnum"))
- nwells=${#batchnum[@]}
- target_well_rows=()
- for ((row=1; row<=nwells; row++))
- do
- if [[ "${batchnum[$row]}" == "${SGE_TASK_ID}" ]]
- then
- target_well_rows+=($row)
- fi
- done
- # filepaths associated with target rows in well-level metadata -----------------
- wellprefix=($(query_metadat "wellprefix"))
- dir_well=($(query_metadat "A06a_dir_star"))
- # set naming convention within each well folder --------------------------------
- # .bam files in (from A06a)
- bam_in_pe=PE.Aligned.out.bam
- bam_in_se1=SE1.Aligned.out.bam
- bam_in_se2=SE2.Aligned.out.bam
- # de-duplicated .bam
- bam_dedupe_pe=pe_dedupe.bam
- bam_dedupe_se1=se1_dedupe.bam
- bam_dedupe_se2=se2_dedupe.bam
- # de-duplication logs
- picard_log_pe=star_dedupe_pe.log
- picard_log_se1=star_dedupe_se1.log
- picard_log_se2=star_dedupe_se2.log
- # final .sam
- sam_q10_pe=q10_pe.sam
- sam_q10_se1=q10_se1.sam
- sam_q10_se2=q10_se2.sam
- # final .bam
- bam_final_pe=PE.Final.bam
- bam_final_se1=SE1.Final.bam
- bam_final_se2=SE2.Final.bam
- # map each well
- for row in ${target_well_rows[@]}
- do
- # check directory/prior mapping -------------------------------------------
- # check for existing mapping output
- # if final outputs exist, skip; else run filtering of direct STAR align
- cd ${dir_proj}
- if [[ -s ${dir_well[$row]}/${bam_final_pe} \
- && -s ${dir_well[$row]}/${bam_final_se1} \
- && -s ${dir_well[$row]}/${bam_final_se2} \
- && "${skip_complete}"=="true" ]]
- then
- echo -e "final aligned .bams for '${wellprefix[$row]}' already exist. skipping this well.'"
- else
- echo -e "\n\napplying STAR to '${wellprefix[$row]}'...\n\n"
- cd ${dir_proj}/${dir_well[$row]}
- # proceed if all input files exist
- if [[ -s ${bam_in_pe} && -s ${bam_in_se1} && -s ${bam_in_se2} ]]
- then
- # clear intermediate files (if not all of them exist)
- rm ${bam_final_pe} ${bam_final_se1} ${bam_final_se2} 2>&1 >/dev/null
- # run RNA filtering -------------------------------------------------------
- # in: three .bams from STAR mapping: $bam_in_* (paired-end, single-end read 1, read 2)
- # out: - MAPQ and f ($bam_final_*)
- # - key log files (e.g., mapping rate)
- # -------------------------------------------------------------------------
- # STAR will output singletons in .bam file
- # (e.g., read 2 maps but read 1 doesn't; samtools view -f 8 ${bam_in_pe})
- # hence the "-f 0x0002" flag for "proper pairs" only, and SE alignments merged with...
- # in later sections, "-f 0x0048" (the read is R1; it mapped but R2 mate didn't map)
- # and "-f 0x0088" (the read is R2; it mapped but R1 mate didn't map)
- # if both R1 & R2 of pair need to pass filtering criteria (AND instead of OR),
- # needs an added annotations --> wide step like the below
- # sed '$!N;s/\n/ /' ${sam_q10_pe}_annotations \
- # | awk '{print ( ($3 == "RNA") && ($6 == "RNA") )}' \
- # | awk '{print $0}1' > ${sam_q10_pe}_annotations_bothpairs
- # then awk filter off of this 'bothpairs' file
- # each step usually ~3-5 min
- # (i) paired-end
- picard MarkDuplicates -I ${bam_in_pe} \
- --ASSUME_SORT_ORDER "queryname" --OPTICAL_DUPLICATE_PIXEL_DISTANCE 2500 \
- --ADD_PG_TAG_TO_READS false --REMOVE_DUPLICATES \
- -O ${bam_dedupe_pe} -M ${picard_log_pe}
- samtools view -h -q 10 -f 0x0002 ${bam_dedupe_pe} > ${sam_q10_pe}
- perl ${dir_proj}/Scripts/A06b_classify_mCT_reads_STAR.pl ${sam_q10_pe}
- awk 'NR == FNR { if ($0=="" || $3=="RNA") line[NR]; next } (FNR in line)' \
- ${sam_q10_pe}_annotations ${sam_q10_pe} |
- samtools view -b - | samtools sort - > ${bam_final_pe}
- samtools index ${bam_final_pe}
- # (ii) single-end, read 1
- picard MarkDuplicates -I ${bam_in_se1} \
- --ASSUME_SORT_ORDER "queryname" --OPTICAL_DUPLICATE_PIXEL_DISTANCE 2500 \
- --ADD_PG_TAG_TO_READS false --REMOVE_DUPLICATES \
- -O ${bam_dedupe_se1} -M ${picard_log_se1}
- samtools view -h -q 10 ${bam_dedupe_se1} > ${sam_q10_se1}
- perl ${dir_proj}/Scripts/A06b_classify_mCT_reads_STAR.pl ${sam_q10_se1}
- awk 'NR == FNR { if ($0=="" || $3=="RNA") line[NR]; next } (FNR in line)' \
- ${sam_q10_se1}_annotations ${sam_q10_se1} |
- samtools view -b - | samtools sort - > ${bam_final_se1}
- samtools index ${bam_final_se1}
- # (iii) single-end, read 2
- picard MarkDuplicates -I ${bam_in_se2} \
- --ASSUME_SORT_ORDER "queryname" --OPTICAL_DUPLICATE_PIXEL_DISTANCE 2500 \
- --ADD_PG_TAG_TO_READS false --DUPLICATE_SCORING_STRATEGY RANDOM --REMOVE_DUPLICATES \
- -O ${bam_dedupe_se2} -M ${picard_log_se2}
- samtools view -h -q 10 ${bam_dedupe_se2} > ${sam_q10_se2}
- perl ${dir_proj}/Scripts/A06b_classify_mCT_reads_STAR.pl ${sam_q10_se2}
- awk 'NR == FNR { if ($0=="" || $3=="RNA") line[NR]; next } (FNR in line)' \
- ${sam_q10_se2}_annotations ${sam_q10_se2} |
- samtools view -b - | samtools sort - > ${bam_final_se2}
- samtools index ${bam_final_se2}
- # (iv.) optional clean-up (comment out as desired)
- rm ${bam_dedupe_pe} ${bam_dedupe_se1} ${bam_dedupe_se2}
- rm ${sam_q10_pe} ${sam_q10_se1} ${sam_q10_se2}
- # rm *annotations
- fi
- fi
- done
- echo -e "\n\n'A06b_star_filtering' completed.\n\n"
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `date `
- echo " "
- # %% [markdown]
- # ## (A06b) check star output
- # %%
- %%bash
- cat > ../Scripts/A06c_check_star.sub
- #!/bin/bash
- #$ -cwd
- #$ -o sublogs/A06c_check_star.$JOB_ID
- #$ -j y
- #$ -l h_rt=1:00:00,h_data=8G
- #$ -N A06c_starcheck
- #$ -hold_jid A06a_star
- echo "Job $JOB_ID started on: " `hostname -s`
- echo "Job $JOB_ID started on: " `date `
- echo " "
- # environment init -------------------------------------------------------------
- export $(cat snmCT_parameters.env | grep -v '^#' | xargs) # <--
- # helper functions / extract target filepaths ----------------------------------
- query_metadat () {
- awk -F',' -v targetcol="$1" \
- 'NR==1 {
- for (i=1;i<=NF;i++) {
- if ($i==targetcol) {assayout=i; break} }
- print $assayout
- }
- NR>1 {
- print $assayout
- }' ${metadat_well}
- }
- check_filepaths_in_assay() {
- for file in $@
- do
- if [[ ! -s ${file} ]]
- then
- echo "missing '${file}'"
- fi
- done
- }
- check_filepath_by_batch() {
- target_array=($@)
- batches_to_rerun=()
- for ((target_batch=1; target_batch<=nbatches; target_batch++))
- do
- target_well_rows=()
- for ((row=1; row<=nwells; row++))
- do
- if [[ "${batchnum[$row]}" == "${target_batch}" ]]
- then
- target_well_rows+=($row)
- fi
- done
- batch_file_list=${target_array[@]: ${target_well_rows[0]}:${#target_well_rows[@]} }
- num_files_missing=$(check_filepaths_in_assay ${batch_file_list[@]} | wc -l)
- if [[ ${num_files_missing} > 0 ]]
- then
- batches_to_rerun+=(${target_batch})
- echo -e "${target_batch} \t ${num_files_missing}"
- fi
- done
- if [[ ${#batches_to_rerun[@]} > 0 ]]
- then
- echo "batches to re-run:"
- echo "${batches_to_rerun[*]}"
- fi
- }
- batchnum=($(query_metadat "batchnum"))
- nwells=${#batchnum[@]}
- nbatches=${batchnum[-1]}
- # apply checks for A06a output -------------------------------------------------
- echo "-----------------------------------------------------------------"
- echo "A. printing number of missing STAR mapping .bam files missing (by batch)... "
- echo "-----------------------------------------------------------------"
- bam_star_pe=($(query_metadat "A06a_bam_star_PE"))
- bam_star_se1=($(query_metadat "A06a_bam_star_SE1"))
- bam_star_se2=($(query_metadat "A06a_bam_star_SE2"))
- echo -e "\n\n\nchecking PE.Aligned.out.bam ----------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_star_pe[@]}
- echo -e "\n\n\nchecking SE1.Aligned.out.bam ---------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_star_se1[@]}
- echo -e "\n\n\nchecking SE2.Aligned.out.bam ---------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_star_se2[@]}
- echo "-----------------------------------------------------------------"
- echo "B. printing number of missing filtered .bam files missing (by batch)... "
- echo "-----------------------------------------------------------------"
- bam_starfilt_pe=($(query_metadat "A06b_bam_starfilt_PE"))
- bam_starfilt_se1=($(query_metadat "A06b_bam_starfilt_SE1"))
- bam_starfilt_se2=($(query_metadat "A06b_bam_starfilt_SE2"))
- echo -e "\n\n\nchecking PE.Final.out.bam ------------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_starfilt_pe[@]}
- echo -e "\n\n\nchecking SE1.Final.out.bam -----------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_starfilt_se1[@]}
- echo -e "\n\n\nchecking SE2.Final.out.bam -----------------------------------------------------"
- echo -e "batchnum\tnum_missing"
- check_filepath_by_batch ${bam_starfilt_se2[@]}
- echo -e "\n\n\nsuggest re-running and checking sublog output of above batches."
- echo -e "\n\n-----------------------------------------------------------------"
- echo "C. checking log files for issues."
- echo -e "-----------------------------------------------------------------\n"
- echo -e "\n\nchecking if 'completed' in sublogs/A06a_star* output."
- echo "if any filename is printed, the associated batch may have not completed mapping."
- grep -c "'A06a_star_mapping' completed" sublogs/A06a_star* | awk -F ":" '$2==0 {print $1}'
- echo -e "\n\nchecking if 'completed' is in sublogs/A06b_starfilt* output."
- echo "if any filename is printed, the associated batch may have not completed mapping."
- grep -c "A06b_star_filtering" sublogs/A06b_starfilt* | awk -F ":" '$2==0 {print $1}'
- echo -e "\n\nchecking if 'Exception' is in sublogs/A06b_starfilt* output."
- echo "if any filename is printed, may be issues with .bams from STAR mapping."
- echo "(e.g., stale file handle, premature end of file, unexpected compressed block length)"
- grep -c "Exception" sublogs/A06b_starfilt* | awk -F ":" '$2!=0 {print $1}'
- echo -e "\n\n'A06c_starcheck' completed.\n\n"
- echo "Job $JOB_ID ended on: " `hostname -s`
- echo "Job $JOB_ID ended on: " `date `
- echo " "
- # %% [markdown]
- # ## (A06d) feature counts
- # %%
- %%bash
- cat > ../Scripts/A06d_featurecounts.sub
- #!/bin/bash
- #$ -cwd
- #$ -o sublogs/A06d_featurecounts.$JOB_ID.$TASK_ID
- #$ -j y
- #$ -l h_rt=3:00:00,h_data=24G
- #$ -N A06d_featurecounts
- #$ -t 1-32
- #$ -hold_jid A06b_starfilt
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `date `
- echo " "
- # environment init -------------------------------------------------------------
- . /u/local/Modules/default/init/modules.sh # <--
- module load anaconda3 # <--
- conda activate snmCTseq # <--
- export $(cat snmCT_parameters.env | grep -v '^#' | xargs) # <--
- dir_out_gene=featurecounts_gene/ # <--
- dir_out_exon=featurecounts_exon/ # <--
- quantify_exons="true" # <-- (typically we analyze combined exon+intron)
- # extract target filepaths -----------------------------------------------------
- # helper functions
- query_metadat () {
- awk -F',' -v targetcol="$1" \
- 'NR==1 {
- for (i=1;i<=NF;i++) {
- if ($i==targetcol) {assayout=i; break} }
- print $assayout
- }
- NR>1 {f
- print $assayout
- }' ${metadat_well}
- }
- # extract target wells, print values for log
- platenum=($(query_metadat "platenum"))
- nwells=${#platenum[@]}
- target_well_rows=()
- for ((row=1; row<=nwells; row++))
- do
- if [[ "${platenum[$row]}" == "${SGE_TASK_ID}" ]]
- then
- target_well_rows+=($row)
- fi
- done
- # filepaths associated with target rows in well-level metadata -----------------
- wellprefix=($(query_metadat "wellprefix"))
- dir_well=($(query_metadat "A06a_dir_star"))
- bam_pe=($(query_metadat "A06b_bam_starfilt_PE"))
- bam_se1=($(query_metadat "A06b_bam_starfilt_SE1"))
- bam_se2=($(query_metadat "A06b_bam_starfilt_SE2"))
- # extract valid .bam filepaths -------------------------------------------------
- # checks only if the .bam exists (otherwise featureCounts may terminate with error)
- # could also wrap in check that number of alignments in file is > 0
- # for greater future compatibility with featureCounts
- pe_files=$(
- for file in ${bam_pe[@]: ${target_well_rows[0] }:${#target_well_rows[@] } }
- do
- if [[ -s ${file} ]]
- then
- echo ${file}
- fi
- done)
- se1_files=$(
- for file in ${bam_se1[@]: ${target_well_rows[0] }:${#target_well_rows[@] } }
- do
- if [[ -s ${file} ]]
- then
- echo ${file}
- fi
- done)
- se2_files=$(
- for file in ${bam_se2[@]: ${target_well_rows[0] }:${#target_well_rows[@] } }
- do
- if [[ -s ${file} ]]
- then
- echo ${file}
- fi
- done)
- # featurecounts on genes -------------------------------------------------------
- # usually <1 min/well --> <1 hr/plate
- if [[ ! -s ${dir_out_gene} ]]
- then
- mkdir ${dir_out_gene}
- fi
- echo "running gene featureCounts on paired-end alignments (mapping_star/*/PE.Final.bam)."
- featureCounts -p -T 4 -t gene -a ${ref_gtf} \
- -o ${dir_out_gene}/PE_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${pe_files}
- echo "running gene featureCounts on single-end R1 alignments (mapping_star/*/SE1.Final.bam)."
- featureCounts -T 4 -t gene -a ${ref_gtf} \
- -o ${dir_out_gene}/SE1_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${se1_files}
- echo "running gene featureCounts on single-end R2 alignments (mapping_star/*/SE2.Final.bam)."
- featureCounts -T 4 -t gene -a ${ref_gtf} \
- -o ${dir_out_gene}/SE2_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${se2_files}
- # featurecounts on exons -------------------------------------------------------
- # usually <1 min/well --> <1 hr/plate
- if [[ ${quantify_exons} == "true" ]]
- then
- if [[ ! -s ${dir_out_exon} ]]
- then
- mkdir ${dir_out_exon}
- fi
- echo "running exon featureCounts on paired-end alignments (mapping_star/*/PE.Final.bam)."
- featureCounts -p -T 4 -t exon -a ${ref_gtf} \
- -o ${dir_out_exon}/PE_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${pe_files}
- echo "running exon featureCounts on single-end R1 alignments (mapping_star/*/SE1.Final.bam)."
- featureCounts -T 4 -t exon -a ${ref_gtf} \
- -o ${dir_out_exon}/SE1_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${se1_files}
- echo "running exon featureCounts on single-end R2 alignments (mapping_star/*/SE2.Final.bam)."
- featureCounts -T 4 -t exon -a ${ref_gtf} \
- -o ${dir_out_exon}/SE2_${SGE_TASK_ID} --donotsort --tmpDir ${dir_scratch} ${se2_files}
- fi
- echo -e "\n\n'A06d_featurecounts' completed.\n\n"
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `date `
- echo " "
- # %% [markdown]
- # ## (A06e) final .bam stats like counts, % intergenic
- # %%
- %%bash
- cat > ../Scripts/A06e_star_bam_stats.sub
- #!/bin/bash
- #$ -cwd
- #$ -o sublogs/A06e_samstat_star.$JOB_ID.$TASK_ID
- #$ -j y
- #$ -l h_rt=2:00:00,h_data=24G
- #$ -N A06e_samstat_star
- #$ -t 1-512
- #$ -hold_jid_ad A06b_starfilt
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID started on: " `date `
- echo " "
- # environment init -------------------------------------------------------------
- . /u/local/Modules/default/init/modules.sh # <--
- module load anaconda3 # <--
- conda activate snmCTseq # <--
- export $(cat snmCT_parameters.env | grep -v '^#' | xargs) # <--
- skip_complete=true # <-- for help with incomplete jobs
- # extract target filepaths -----------------------------------------------------
- # helper functions
- query_metadat () {
- awk -F',' -v targetcol="$1" \
- 'NR==1 {
- for (i=1;i<=NF;i++) {
- if ($i==targetcol) {assayout=i; break} }
- print $assayout
- }
- NR>1 {
- print $assayout
- }' ${metadat_well}
- }
- # extract target wells, print values for log
- batchnum=($(query_metadat "batchnum"))
- nwells=${#batchnum[@]}
- target_well_rows=()
- for ((row=1; row<=$nwells; row++))
- do
- if [[ "${batchnum[$row]}" == "${SGE_TASK_ID}" ]]
- then
- target_well_rows+=($row)
- fi
- done
- # filepaths associated with target rows in well-level metadata -----------------
- wellprefix=($(query_metadat "wellprefix"))
- dir_well=($(query_metadat "A06a_dir_star"))
- # samtools stats on each well in the batch -------------------------------------
- # usually <5 seconds per .bam --> <12 min/24 wells
- for row in ${target_well_rows[@]}
- do
- cd ${dir_proj}
- if [[ -s ${dir_well[$row]}/samstats_PE \
- && -s ${dir_well[$row]}/samstats_SE1 \
- && -s ${dir_well[$row]}/samstats_SE2 \
- && -s ${dir_well[$row]}/picard_PE \
- && -s ${dir_well[$row]}/picard_SE1 \
- && -s ${dir_well[$row]}/picard_SE2 ]]
- then
- echo -e "final metrics for '${wellprefix[$row]}' already exist."
- if [[ ${skip_complete} = "true" ]]
- then
- echo "skip_complete = true. skipping this well."
- continue
- else
- echo "but skip_complete = false. re-running this well."
- fi
- fi
- echo -e "\n\ngetting metrics for '${wellprefix[$row]}'...\n\n"
- cd ${dir_well[$row]}
- # run samtools stats
- samtools stats PE.Final.bam | grep '^SN' | cut -f 2,3 > samstats_PE
- samtools stats SE1.Final.bam | grep '^SN' | cut -f 2,3 > samstats_SE1
- samtools stats SE2.Final.bam | grep '^SN' | cut -f 2,3 > samstats_SE2
- # run picard CollectRnaSeqMetrics
- samtools sort PE.Final.bam | \
- picard CollectRnaSeqMetrics -I /dev/stdin -O picard_PE \
- --REF_FLAT ${ref_flat} -STRAND "NONE" --RIBOSOMAL_INTERVALS ${ref_rrna}
- samtools sort SE1.Final.bam | \
- picard CollectRnaSeqMetrics -I /dev/stdin -O picard_SE1 \
- --REF_FLAT ${ref_flat} -STRAND "NONE" --RIBOSOMAL_INTERVALS ${ref_rrna}
- samtools sort SE2.Final.bam | \
- picard CollectRnaSeqMetrics -I /dev/stdin -O picard_SE2 \
- --REF_FLAT ${ref_flat} -STRAND "NONE" --RIBOSOMAL_INTERVALS ${ref_rrna}
- done
- echo -e "\n\n'A06e_samstat_star' completed.\n\n"
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `hostname -s`
- echo "Job $JOB_ID.$SGE_TASK_ID ended on: " `date `
- echo " "
A06_mapping_star.ipynb at commit 3a5f663, no license · at the source
Overview
- Department of Neurology, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Bioinformatics Interdepartmental Program, University of California, Los Angeles, Los Angeles, CA, USA
- Department of Human Genetics, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Department of Computer Science, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Department of Computational Medicine, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Center for Autism Research and Treatment, Semel Institute, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Department of Psychiatry, Semel Institute, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA 90095, USA
- Institute for Precision Health, University of California, Los Angeles, Los Angeles, CA 90095, USA
Abstract
Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition. Studies of postmortem ASD brain tissue have revealed convergent molecular changes across the cortex. Whether these features are reflected in cell-type-specific epigenetic signatures is unknown. Here, we present a single-cell analysis of DNA methylation (DNAm) coupled with transcriptomics in ASD. Using snmCT-seq, we profiled DNAm and transcript levels from over 60,000 nuclei derived from the prefrontal cortex of 49 donors. We identified over 30,000 differentially methylated regions (DMRs) in ASD that were enriched in promoters and cell-type-specific regulatory elements active across the lifespan. ASD-related methylation changes were uncorrelated with transcript levels and were small in magnitude compared with age-associated effects. Age-DMRs were concentrated in excitatory neurons and revealed distinct roles for CG and non-CG methylation. Age-varying methylation signatures of ASD identified neuron projection development as a key process perturbed in ASD, highlighting the heterogeneous impact of ASD across the lifespan.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.
chooliu/snmCTseq_Pipeline
3a5f663678849c9cdd1f666d79135b11e9c19b70, 17 March 2024Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
31 files
- Notebooks/
A00_environment_and_geno , Jupyter, 558 lines, 1 matchme_setup.ipynb - Notebooks/
A01_mergefastq_preptarge , Jupyter, 501 linests.ipynb - Notebooks/
A02_demultiplexing.ipynb , Jupyter, 1,337 lines - Notebooks/
A03_trimming.ipynb , Jupyter, 542 lines, 1 match - Notebooks/
A04_mapping_bismark.ipyn , Jupyter, 1,343 lines, 1 matchb - Notebooks/
A05_compile_metadata_DNA , Jupyter, 855 lines.ipynb - Notebooks/
A06_mapping_star.ipynb , Jupyter, 1,120 lines, 2 matches - Notebooks/
A07_compile_metadata_RNA , Jupyter, 933 lines.ipynb - Notebooks/
A08_compile_final_metada , Jupyter, 376 linesta.ipynb - Scripts/
A00d_gtf_annotations_bed , Python, 118 lines.py - Scripts/
A01b_plate_metadata.py , Python, 75 lines - Scripts/
A01c_well_filepaths.py , Python, 213 lines - Scripts/
A02a_demultiplex_fastq.p , Perl, 199 lines, 1 matchl - Scripts/
A04b_classify_mCT_reads_ , Perl, 65 linesbismark.pl - Scripts/
A05a_trimming.py , Python, 61 lines - Scripts/
A05b_DNA_maprate.py , Python, 242 lines - Scripts/
A05c_DNA_dedupe.py , Python, 99 lines - Scripts/
A05d_DNA_global_mCfracs. , Python, 44 linespy - Scripts/
A05e_DNA_samtools.py , Python, 106 lines - Scripts/
A05f_DNA_cov.py , Python, 133 lines - Scripts/
A06b_classify_mCT_reads_ , Perl, 164 linesSTAR.pl - Scripts/
A07a_trimming.py , Python, 61 lines - Scripts/
A07b_RNA_maprate.py , Python, 169 lines - Scripts/
A07c_RNA_dedupe.py , Python, 131 lines - Scripts/
A07d_RNA_featcounts.py , Python, 126 lines - Scripts/
A07e_RNA_samtools.py , Python, 145 lines - Scripts/
A07f_RNA_picard.py , Python, 132 lines - Scripts/
A08a_final_metadat_DNA.p , Python, 99 linesy - Scripts/
A08b_final_metadat_RNA.p , Python, 132 linesy - Scripts/
A08c_metadata_RNADNA.py , Python, 53 lines - README.md, Text, 73 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 30 scripts, each with its path and the digest of its content;
- 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- geo:GSE298243, at NCBI GEO; found in “Data and code availability”
Data and code availability
All original code has been deposited at https://
Processed data have been deposited at GEO as series no. GSE298243 (https://
Controlled-access data are available here: https://
Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Reproduced under the paper's license (CC BY-NC), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 10 authors, 7 keywords, 15 MeSH terms, 2 funders, 130 references, 3 RRIDs.
Cite
This paper
Eyring, K. W., Liu, C., Elhajjaoui, N., Abuhanna, K. D., Zhang, Y., von Behren, Z., Tsai, M. J., Eskin, E., Geschwind, D. H., & Luo, C. (2026). Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures. Cell genomics, 6(7), 101278. https://
BibTeX
@article{eyring2026singl
author = {Eyring, Katherine W and Liu, Cuining and Elhajjaoui, Nasser and Abuhanna, Kevin D and Zhang, Yi and von Behren, Zachary and Tsai, Min Jen and Eskin, Eleazar and Geschwind, Daniel H and Luo, Chongyuan},
title = {{Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures}},
journal = {Cell genomics},
year = {2026},
month = jun,
volume = {6},
number = {7},
pages = {101278},
publisher = {Elsevier},
issn = {2666-979X},
doi = {10.1016/
url = {https://
pmid = {42320469},
pmcid = {PMC13347942}
}
RIS
TY - JOUR
AU - Eyring, Katherine W
AU - Liu, Cuining
AU - Elhajjaoui, Nasser
AU - Abuhanna, Kevin D
AU - Zhang, Yi
AU - von Behren, Zachary
AU - Tsai, Min Jen
AU - Eskin, Eleazar
AU - Geschwind, Daniel H
AU - Luo, Chongyuan
TI - Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures
T2 - Cell genomics
J2 - Cell Genom
PY - 2026
DA - 2026/
VL - 6
IS - 7
SP - 101278
SN - 2666-979X
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures",
"container-title": "Cell genomics",
"author": [
{
"family": "Eyring",
"given": "Katherine W"
},
{
"family": "Liu",
"given": "Cuining"
},
{
"family": "Elhajjaoui",
"given": "Nasser"
},
{
"family": "Abuhanna",
"given": "Kevin D"
},
{
"family": "Zhang",
"given": "Yi"
},
{
"family": "von Behren",
"given": "Zachary"
},
{
"family": "Tsai",
"given": "Min Jen"
},
{
"family": "Eskin",
"given": "Eleazar"
},
{
"family": "Geschwind",
"given": "Daniel H"
},
{
"family": "Luo",
"given": "Chongyuan"
}
],
"container-title-short":
"volume": "6",
"issue": "7",
"page": "101278",
"DOI": "10.1016/
"PMID": "42320469",
"PMCID": "PMC13347942",
"ISSN": "2666-979X",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
19
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41593-026-02247-7 [code]
- Transcriptomic and phenotypic convergence of neurodevelopmental disorder risk genes in vitro and in vivo.Journal: Nature neuroscienceIn common: autism, genetics / omics, 12 references
- [2] doi:10.1038/s42003-026-10059-5 [code]
- Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways.Journal: Communications biologyIn common: FastQC, STAR, SAMtools, autism, genetics / omics, 6 references
- [3] doi:10.1016/j.celrep.2026.117073 [code]
- Single-cell epigenomics uncovers heterochromatin instability and transcription factor dysfunction during mouse brain aging.Journal: Cell reportsIn common: Subread (featureCounts), SAMtools, pandas, 1 other tool, genetics / omics, 7 references
- [4] doi:10.1016/j.xhgg.2026.100652 [code]
- CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling.Journal: HGG advancesIn common: STAR, SAMtools, pandas, 1 other tool, autism, 8 references
- [5] doi:10.1038/s41467-026-76675-1 [code]
- Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.Journal: Nature communicationsIn common: STAR, SAMtools, pandas, 1 other tool, autism, genetics / omics, 7 references
- [6] doi:10.21203/rs.3.rs-9927928/v1 [code]
- Genome-wide and allele-resolved maps of the radial architecture of the mouse genomeJournal: Research Square (preprint)In common: FastQC, Subread (featureCounts), STAR, 3 other tools, genetics / omics, 4 references
- [7] doi:10.1093/bioinformatics/btag592 [code]
- Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.Journal: Bioinformatics (Oxford, England)In common: FastQC, Subread (featureCounts), STAR, 3 other tools, genetics / omics, 4 references
- [8] doi:10.1038/s41467-026-74320-5 [code]
- Spatial architecture of autism pathogenesis reveals mosaic structural disarray during early development.Journal: Nature communicationsIn common: pandas, NumPy, autism, genetics / omics, 9 references
- [9] doi:10.1038/s41586-026-10512-9 [code]
- Astrocyte glucocorticoid receptor signalling restricts neuronal plasticity.Journal: NatureIn common: Subread (featureCounts), STAR, SAMtools, 2 other tools, 6 references
- [10] doi:10.1038/s41586-026-10679-1 [code]
- Cortical development dynamics across autism spectrum disorder mouse models.Journal: NatureIn common: pandas, NumPy, autism, genetics / omics, 10 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 30 scripts, and 6 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:4f5cf25b836b46a1…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
