Gene regulatory innovations from transposable elements in primate cerebellum development.
The 18 matches
- [1] § Methods › Regression analysis of sequence conservation at high attribution sites ↔ 03_HERVL/08_nucleotide_preservation.ipynb, lines 417–510 · score 0.97 · hot encoded nucleotide, absolute attribution score, low attribution positions, high attribution, outcome variable, linear regression
- [2] § Results › Lineage-specific co-option of TE subfamilies as regulatory elements ↔ 02_CRE_potential_screening/03_visualize_screening_candidates.ipynb, lines 159–172 · score 0.76 · LTR9B, MLT1M, ERVL B4, LTR1A2, MER72, MER130
- [3] § Results › Complex cis-regulatory sequences in ancestral TE sequences facilitate co-option of TEs ↔ 02_CRE_potential_screening/03_visualize_screening_candidates.ipynb, lines 159–172 · score 0.76 · LTR9B, MLT1M, ERVL B4, LTR1A2, MER72, MER130
- [4] § Methods › Logistic regression analysis of TE–cCRE overlap ↔ 01_CRE_TE_overlap/06a_regression_analysis_human.ipynb, lines 334–368 · score 0.74 · log_dist_TSS, GC_content, TE_overlap, regression model, mappability
- [5] § Methods › Logistic regression analysis of TE–cCRE overlap ↔ 01_CRE_TE_overlap/06b_regression_analysis_marmoset.ipynb, lines 334–368 · score 0.74 · log_dist_TSS, GC_content, TE_overlap, regression model, mappability
- [6] § Methods › Overlap between TEs and cCREs in mammalian cerebellum development ↔ 01_CRE_TE_overlap/01_make_repeat_beds.ipynb, lines 142–149 · score 0.70 · nf LO, liftOver, chain, transposable element, UCSC, Dfam
- [7] § Results › Complex cis-regulatory sequences in ancestral TE sequences facilitate co-option of TEs ↔ 03_HERVL/01_find_accessible_copies.py, lines 393–431 · score 0.67 · MamRTE1, MamTip2b, MLT1M, ancestral sequence, MER130, copies
- [8] § Results › Lineage-specific co-option of TE subfamilies as regulatory elements ↔ 03_HERVL/01_find_accessible_copies.py, lines 433–482 · score 0.65 · MamRTE1, MamTip2b, MLT1M, MER130, TEs, mice
- [9] § Methods › Logistic regression analysis of TE–cCRE overlap ↔ 01_CRE_TE_overlap/07_regression_analysis_catlas.ipynb, lines 233–254 · score 0.64 · accessibility_class, phyloP, GC_content, TE_overlap, regression, model
- [10] § Results › Determinants of copy-level TE co-option ↔ 03_HERVL/08_nucleotide_preservation.ipynb, lines 417–510 · score 0.58 · high attribution positions, linear regression, HERVL copies, ancestral sequences, motif, scores
- [11] § Methods › Logistic regression analysis of TE copy accessibility ↔ 01_CRE_TE_overlap/06a_regression_analysis_human.ipynb, lines 334–368 · score 0.58 · Odds ratios, GC content, confidence intervals, coefficients, TSS, fitted
- [12] § Methods › Logistic regression analysis of TE copy accessibility ↔ 01_CRE_TE_overlap/06b_regression_analysis_marmoset.ipynb, lines 334–368 · score 0.58 · Odds ratios, GC content, confidence intervals, coefficients, TSS, fitted
- [13] § Methods › Logistic regression analysis of TE–cCRE overlap ↔ 01_CRE_TE_overlap/03a_celltype_specific_peaks_human.ipynb, lines 113–138 · score 0.57 · peakType, n_NMF, intronic, exonic, distal, TE overlap
- [14] § Methods › Logistic regression analysis of TE–cCRE overlap ↔ 01_CRE_TE_overlap/06a_regression_analysis_human.ipynb, lines 523–558 · score 0.57 · n_NMF, postnatal, embryonic, adult, fetal, cons
- [15] § Methods › Evaluation of the contributions of species-specific TE copies to cross-species gene expression divergence ↔ 03_HERVL/14_HERVL_gene_expression.ipynb, lines 140–154 · score 0.55 · cross species comparisons, gene expression, macaque, marmoset, mouse, cells
- [16] § Methods › Mapping chromatin accessibility to the consensus sequences and identifying accessible copies of the screened TE subfamilies ↔ 03_HERVL/01_find_accessible_copies.py, lines 152–163 · score 0.54 · pyBigWig, tracks, copies, accessibility, cell
- [17] § Methods › Single-cell multi-omics atlas of cerebellum development ↔ 03_HERVL/11_distance_analysis.ipynb, lines 31–34 · score 0.52 · cross species comparison, highly variable, distance, NMF, peaks, human
- [18] § Results › Lineage-specific co-option of TE subfamilies as regulatory elements ↔ 03_HERVL/03_HERVL_accessibilitiy.ipynb, lines 424–438 · score 0.51 · UBC_diff, GC_diff_1, GCP, accessibility, HERVL
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Jupyter notebook · 582 lines · 18 KB · CC-BY-NC-SA-4.0 · 3 matches
- # %%
- import pandas as pd
- import numpy as np
- import pyranges as pr
- import pyBigWig
- from Bio import SeqIO
- from Bio.SeqUtils import gc_fraction as calc_GC
- from scipy import stats
- import statsmodels.api as sm
- import statsmodels.formula.api as smf
- from statsmodels.stats.outliers_influence import variance_inflation_factor
- # %%
- path_to_lsdf2 = '/work/tetsuya/sds/sd17d003/'
- # %% [markdown]
- # ## 1. Load data
- # %%
- # Load CREs
- cCREs = pr.read_bed(f"{path_to_lsdf2}/Ioannis/Comparative_Cereb/peak_annotations/Human/human_peaks.bed")
- print(f"Loaded {len(cCREs)} CREs")
- # Load TEs
- TEs = pr.read_bed(f"{path_to_lsdf2}/Tetsuya/data/transposable_elements/dfam/3.8/annotations/hg38/hg38.nrph.hits.te.bed")
- print(f"Loaded {len(TEs)} TEs")
- # %%
- TEs
- # %% [markdown]
- # ## 2. Create random background regions
- # %%
- def create_random_regions(n_regions, width, genome_file, exclude_regions=None):
- """
- Create random genomic regions
- Parameters:
- n_regions: number of regions to create
- width: width of each region (500 for your case)
- genome_file: path to chromosome sizes file
- exclude_regions: PyRanges object of regions to avoid
- """
- # Load chromosome sizes
- chrom_sizes = pd.read_csv(genome_file, sep='\t', header=None,
- names=['Chromosome', 'End'])
- chrom_sizes['Start'] = 0
- # Filter to main chromosomes
- chrom_sizes = chrom_sizes[chrom_sizes['Chromosome'].str.contains('^chr[0-9XY]+$')]
- regions = []
- attempts = 0
- max_attempts = n_regions * 10
- while len(regions) < n_regions and attempts < max_attempts:
- # Randomly select chromosome weighted by size
- weights = chrom_sizes['End'].values / chrom_sizes['End'].sum()
- chrom = np.random.choice(chrom_sizes['Chromosome'], p=weights)
- chrom_size = chrom_sizes[chrom_sizes['Chromosome'] == chrom]['End'].values[0]
- # Random position
- start = np.random.randint(0, max(1, chrom_size - width))
- end = start + width
- # Create region
- region = pr.PyRanges(chromosomes=[chrom], starts=[start], ends=[end])
- # Check if it overlaps excluded regions
- if exclude_regions is not None:
- if len(region.overlap(exclude_regions)) == 0:
- regions.append({'Chromosome': chrom, 'Start': start, 'End': end})
- else:
- regions.append({'Chromosome': chrom, 'Start': start, 'End': end})
- attempts += 1
- background = pr.PyRanges(pd.DataFrame(regions))
- return background
- # %%
- # Create background (5x CREs)
- np.random.seed(123)
- n_background = len(cCREs)
- background = create_random_regions(
- n_regions=n_background,
- width=500,
- genome_file=f'{path_to_lsdf2}/Tetsuya/data/genomes/hg38/hg38.chrom.sizes',
- exclude_regions=cCREs
- )
- print(f"Created {len(background)} background regions")
- # %% [markdown]
- # ## 3. Calculate TE overlap
- # %%
- def calculate_overlap(regions, features, min_overlap=250):
- """
- Calculate binary overlap: does region overlap any feature by at least min_overlap bp?
- Parameters:
- regions: PyRanges object of genomic regions
- features: PyRanges object of features (TEs)
- min_overlap: minimum overlap width in bp (default 250)
- Returns:
- PyRanges object with TE_overlap column added
- """
- # Get all overlaps
- overlaps = regions.join(features, how='left', suffix='_TE')
- if len(overlaps) == 0:
- # No overlaps at all
- regions_df = regions.df.copy()
- regions_df['TE_overlap'] = 0
- return pr.PyRanges(regions_df)
- # Calculate intersection width for each overlap
- overlaps_df = overlaps.df.copy()
- # Intersection start = max of the two starts
- # Intersection end = min of the two ends
- overlaps_df['intersection_start'] = overlaps_df[['Start', 'Start_TE']].max(axis=1)
- overlaps_df['intersection_end'] = overlaps_df[['End', 'End_TE']].min(axis=1)
- overlaps_df['overlap_width'] = overlaps_df['intersection_end'] - overlaps_df['intersection_start']
- # Filter to overlaps >= min_overlap
- significant_overlaps = overlaps_df[overlaps_df['overlap_width'] > min_overlap]
- # Get unique region indices with significant overlaps
- if len(significant_overlaps) > 0:
- # Need to match back to original regions
- # Create unique identifier for regions
- overlaps_df['region_id'] = (overlaps_df['Chromosome'].astype(str) + ':' +
- overlaps_df['Start'].astype(str) + '-' +
- overlaps_df['End'].astype(str))
- significant_overlaps['region_id'] = (significant_overlaps['Chromosome'].astype(str) + ':' +
- significant_overlaps['Start'].astype(str) + '-' +
- significant_overlaps['End'].astype(str))
- overlapped_ids = set(significant_overlaps['region_id'].unique())
- else:
- overlapped_ids = set()
- # Add binary overlap to original regions
- regions_df = regions.df.copy()
- regions_df['region_id'] = (regions_df['Chromosome'].astype(str) + ':' +
- regions_df['Start'].astype(str) + '-' +
- regions_df['End'].astype(str))
- regions_df['TE_overlap'] = regions_df['region_id'].isin(overlapped_ids).astype(int)
- regions_df = regions_df.drop('region_id', axis=1)
- return pr.PyRanges(regions_df)
- # %%
- # Calculate for both
- cCREs_with_TE = calculate_overlap(cCREs, TEs)
- background_with_TE = calculate_overlap(background, TEs)
- # Add group labels
- cCREs_df = cCREs_with_TE.df
- cCREs_df['is_CRE'] = 1
- background_df = background_with_TE.df
- background_df['is_CRE'] = 0
- # Combine
- all_regions_df = pd.concat([cCREs_df, background_df], ignore_index=True)
- print(f"Total regions: {len(all_regions_df)}")
- print(f"TE overlap in CREs: {all_regions_df[all_regions_df['is_CRE']==1]['TE_overlap'].mean():.3f}")
- print(f"TE overlap in background: {all_regions_df[all_regions_df['is_CRE']==0]['TE_overlap'].mean():.3f}")
- # %%
- len(cCREs_df[cCREs_df['TE_overlap']==1])
- # %% [markdown]
- # ## 4. Calculate GC content
- # %%
- from pyfaidx import Fasta
- def calculate_gc_content(regions_df, genome_fasta):
- """Calculate GC content for genomic regions"""
- genome = Fasta(genome_fasta)
- gc_values = []
- for idx, row in regions_df.iterrows():
- chrom = row['Chromosome']
- start = row['Start']
- end = row['End']
- try:
- seq = genome[chrom][start:end].seq.upper()
- gc = (seq.count('G') + seq.count('C')) / len(seq) if len(seq) > 0 else 0
- gc_values.append(gc)
- except:
- gc_values.append(np.nan)
- return gc_values
- # %%
- all_regions_df['GC_content'] = calculate_gc_content(all_regions_df, f'{path_to_lsdf2}/Tetsuya/data/genomes/hg38/hg38.fa')
- print(f"GC content calculated, mean: {all_regions_df['GC_content'].mean():.3f}")
- # %% [markdown]
- # ## 5. Calculate mappability
- # %%
- def calculate_mappability(regions_df, bigwig_file):
- """Calculate mean mappability score for each region"""
- bw = pyBigWig.open(bigwig_file)
- mappability_scores = []
- for idx, row in regions_df.iterrows():
- chrom = row['Chromosome']
- start = int(row['Start'])
- end = int(row['End'])
- try:
- # Get values for this region
- values = bw.values(chrom, start, end)
- # Filter out None values and calculate mean
- valid_values = [v for v in values if v is not None]
- if len(valid_values) > 0:
- mappability_scores.append(np.mean(valid_values))
- else:
- mappability_scores.append(0)
- except:
- mappability_scores.append(0)
- bw.close()
- return mappability_scores
- # %%
- all_regions_df['mappability'] = calculate_mappability(
- all_regions_df,
- f'{path_to_lsdf2}/Tetsuya/data/mappability/hg38.k50.Umap.MultiTrackMappability.bw'
- )
- print(f"Mappability calculated, mean: {all_regions_df['mappability'].mean():.3f}")
- # %% [markdown]
- # ## 6. Calculate distance to nearest TSS
- # %%
- def extract_tss_from_gtf(gtf_file, feature_type='gene'):
- """
- Extract TSS positions from GTF file
- Parameters:
- gtf_file: path to GTF file (e.g., gencode.v38.annotation.gtf.gz)
- feature_type: 'transcript' or 'gene' (transcript gives more TSSs)
- Returns:
- PyRanges object with TSS positions (single bp)
- """
- # Read GTF
- print(f"Reading {gtf_file}...")
- gtf = pr.read_gtf(gtf_file)
- # Filter to desired feature type
- features = gtf[gtf.Feature == feature_type]
- print(f"Found {len(features)} {feature_type} features")
- # Extract TSS based on strand
- # TSS = 5' end of transcript
- # For + strand: TSS = Start
- # For - strand: TSS = End
- df = features.df.copy()
- # Create TSS positions
- df['TSS'] = np.where(df['Strand'] == '+', df['Start'], df['End'])
- # Create single-bp regions for TSS
- # PyRanges uses 0-based half-open coordinates [start, end)
- # So a single position is represented as [pos, pos+1)
- tss_df = pd.DataFrame({
- 'Chromosome': df['Chromosome'],
- 'Start': df['TSS'],
- 'End': df['TSS'] + 1,
- 'Strand': df['Strand'],
- 'gene_id': df.get('gene_id', ''),
- 'gene_name': df.get('gene_name', ''),
- 'transcript_id': df.get('transcript_id', '')
- })
- # Remove duplicates (same TSS position)
- tss_df = tss_df.drop_duplicates(subset=['Chromosome', 'Start', 'End'])
- print(f"Extracted {len(tss_df)} unique TSS positions")
- return pr.PyRanges(tss_df)
- def calculate_distance_to_tss(regions, gtf_file):
- """
- Calculate distance from regions to nearest TSS
- Parameters:
- regions: PyRanges object of genomic regions
- gtf_file: path to GTF annotation file
- Returns:
- Array of distances to nearest TSS
- """
- # Extract TSS positions
- tss = extract_tss_from_gtf(gtf_file, feature_type='transcript')
- # Find nearest TSS for each region
- print("Calculating distances to nearest TSS...")
- nearest = regions.nearest(tss, overlap=True)
- # Extract distances
- # PyRanges nearest() returns Distance column
- # Negative distances mean overlap (upstream/downstream)
- # We want absolute distance
- distances = nearest.df['Distance'].abs().values
- return distances
- # %%
- # Usage
- gtf_file = f"{path_to_lsdf2}/Ioannis/Comparative_Cereb/genome_annotations/Human/hg38_ens92_filtered.gtf"
- all_regions_pr = pr.PyRanges(all_regions_df[['Chromosome', 'Start', 'End']])
- all_regions_df['dist_to_TSS'] = calculate_distance_to_tss(all_regions_pr, gtf_file)
- print(f"Distance to TSS - median: {np.median(all_regions_df['dist_to_TSS']):.0f} bp")
- print(f"Distance to TSS - mean: {np.mean(all_regions_df['dist_to_TSS']):.0f} bp")
- # %% [markdown]
- # ## 7. Main regression model
- # %%
- import statsmodels.api as sm
- # Remove any rows with missing data
- df_clean = all_regions_df.dropna(subset=['TE_overlap', 'is_CRE', 'GC_content', 'mappability', 'dist_to_TSS'])
- df_clean['log_dist_TSS'] = np.log(df_clean['dist_to_TSS'] + 1)
- # Main model: GC + mappability
- formula_main = 'TE_overlap ~ is_CRE + GC_content + mappability + log_dist_TSS'
- model_main = smf.logit(formula_main, data=df_clean).fit()
- print("\n" + "="*60)
- print("MAIN MODEL: TE_overlap ~ is_CRE + GC_content + mappability + log_dist_TSS")
- print("="*60)
- print(model_main.summary())
- # Get odds ratios
- odds_ratios = np.exp(model_main.params)
- conf_int = np.exp(model_main.conf_int())
- conf_int['OR'] = odds_ratios
- conf_int.columns = ['2.5%', '97.5%', 'OR']
- print("\nOdds Ratios and 95% Confidence Intervals:")
- print(conf_int)
- # Focus on is_CRE coefficient
- or_cre = odds_ratios['is_CRE']
- ci_low = conf_int.loc['is_CRE', '2.5%']
- ci_high = conf_int.loc['is_CRE', '97.5%']
- pval = model_main.pvalues['is_CRE']
- print(f"\nCRE effect: OR = {or_cre:.3f}, 95% CI: [{ci_low:.3f}, {ci_high:.3f}], P = {pval:.2e}")
- # %%
- print(model_main.pvalues)
- # %% [markdown]
- # ---
- # %% [markdown]
- # # Within cCRE comparisons:
- # %%
- human_peaks = pd.read_csv(f'{path_to_lsdf2}/Tetsuya/project/cerebellum/ioannis/Comparative_Cereb/peak_annotations/Human/human_peaks_repeat_overlap.tsv',
- sep='\t')
- human_peaks['Name'] = human_peaks['peak'].str.replace('hg38_', '')
- # %%
- human_peaks = human_peaks.merge(all_regions_df[all_regions_df['is_CRE']==1][['Name', 'TE_overlap', 'GC_content', 'mappability', 'dist_to_TSS']],
- how='left', on='Name')
- # %%
- mapping = {
- 1: "human_specific",
- 3: "primate_specific",
- 7: "conserved",
- }
- human_peaks["Cons_group_label"] = human_peaks["Cons_group"].map(mapping).fillna("others")
- # Set "Others" as reference
- human_peaks['Cons_group_label'] = pd.Categorical(
- human_peaks['Cons_group_label'],
- categories=['others', 'human_specific', 'primate_specific', 'conserved'],
- ordered=False
- )
- human_peaks.to_csv(f'{path_to_lsdf2}/Tetsuya/project/cerebellum/ioannis/Comparative_Cereb/peak_annotations/Human/human_peaks_regression.tsv',
- sep='\t', index=False)
- # %%
- human_peaks
- # %%
- print("\n" + "="*70)
- print("MAIN MODEL: Peak Type + Conservation + Covariates")
- print("="*70)
- formula = 'TE_overlap ~ C(peakType) + C(Cons_group_label) + GC_content'
- model_main = smf.logit(formula, data=human_peaks).fit()
- print(model_main.summary())
- # Get odds ratios
- odds_ratios = np.exp(model_main.params)
- conf_int = np.exp(model_main.conf_int())
- conf_int['OR'] = odds_ratios
- conf_int.columns = ['2.5%', '97.5%', 'OR']
- print("\n" + "="*70)
- print("ODDS RATIOS AND 95% CONFIDENCE INTERVALS")
- print("="*70)
- print(conf_int)
- model_main.pvalues
- # %% [markdown]
- # ### n_NMF
- # %%
- human_peaks = pd.read_csv(f'{path_to_lsdf2}/Tetsuya/project/cerebellum/ioannis/Comparative_Cereb/peak_annotations/Human/human_peaks_regression.tsv',
- sep='\t')
- # %%
- nmf_list = [f'NMF_{i}' for i in range(1, 19)]
- nmf_size_list = []
- for nmf in nmf_list:
- nmf_peak_df = pd.read_csv(f'{path_to_lsdf2}/Ioannis/Comparative_Cereb/crossSpecies_comparisons/atac_nmf/allFilteredPeaks_nnls/human/{nmf}.bed',
- sep='\t', names=['chrom', 'start', 'end', 'peak'])
- print(f'{nmf}: {len(nmf_peak_df)}')
- nmf_size_list.append(len(nmf_peak_df))
- nmf_peak_df[nmf] = 1
- human_peaks = human_peaks.merge(nmf_peak_df[['peak', nmf]], on='peak', how='left')
- human_peaks[nmf] = human_peaks[nmf].fillna(0).astype(int)
- human_peaks['n_NMF'] = human_peaks[nmf_list].sum(axis=1)
- human_peaks
- # %%
- # Add n_NMF_label column with classification
- human_peaks_subset = human_peaks[human_peaks['n_NMF'] > 0].copy()
- human_peaks_subset['n_NMF_label'] = human_peaks_subset['n_NMF'].apply(
- lambda x: str(int(x)) if x < 10 else '>9'
- )
- # Set "Others" as reference
- human_peaks_subset['Cons_group_label'] = pd.Categorical(
- human_peaks_subset['Cons_group_label'],
- categories=['others', 'human_specific', 'primate_specific', 'conserved'],
- ordered=False
- )
- # Convert to categorical with '0' as the reference level
- human_peaks_subset['n_NMF_label'] = pd.Categorical(
- human_peaks_subset['n_NMF_label'],
- categories=['1', '2', '3', '4', '5', '6', '7', '8', '9', '>9'],
- ordered=True
- )
- # %%
- print("\n" + "="*70)
- print("MAIN MODEL: Peak Type + Conservation + Covariates")
- print("="*70)
- formula = 'TE_overlap ~ C(peakType) + C(Cons_group_label) + C(n_NMF_label) + GC_content'
- model_main = smf.logit(formula, data=human_peaks_subset).fit()
- print(model_main.summary())
- # Get odds ratios
- odds_ratios = np.exp(model_main.params)
- conf_int = np.exp(model_main.conf_int())
- conf_int['OR'] = odds_ratios
- conf_int.columns = ['2.5%', '97.5%', 'OR']
- print("\n" + "="*70)
- print("ODDS RATIOS AND 95% CONFIDENCE INTERVALS")
- print("="*70)
- print(conf_int)
- model_main.pvalues
- # %% [markdown]
- # ### Highly variable NMF
- # %%
- nmf_list = [f'NMF_{i}' for i in range(1, 19)]
- nmf_size_list = []
- for nmf in nmf_list:
- nmf_peak_df = pd.read_csv(f'{path_to_lsdf2}/Ioannis/Comparative_Cereb/crossSpecies_comparisons/atac_nmf/highly_variable_peaks/human/{nmf}.bed',
- sep='\t', names=['chrom', 'start', 'end', 'peak'])
- print(f'{nmf}: {len(nmf_peak_df)}')
- nmf_size_list.append(len(nmf_peak_df))
- nmf_peak_df[nmf] = 1
- human_peaks = human_peaks.merge(nmf_peak_df[['peak', nmf]], on='peak', how='left')
- human_peaks[nmf] = human_peaks[nmf].fillna(0).astype(int)
- human_peaks['n_NMF'] = human_peaks[nmf_list].sum(axis=1)
- human_peaks
- # %%
- peak_info = pd.read_csv(f'{path_to_lsdf2}/Ioannis/Comparative_Cereb/peak_annotations/Human/human_peaks_info.txt',
- sep='\t')
- peak_info['maxAcc_devState'] = peak_info['maxAcc_Sample'].str.split(':', expand=True)[1]
- peak_info_subset = peak_info[['peak', 'maxAcc_devState']].copy()
- peak_info_subset
- # %%
- human_peaks_subset = human_peaks[human_peaks['n_NMF'] > 0]
- dev_state_mapping = {
- 'CS18-19': 'embryonic',
- 'CS20': 'embryonic',
- 'CS22-23': 'embryonic',
- '11wpc': 'fetal',
- '15-17wpc': 'fetal',
- 'newborn': 'postnatal',
- 'infant': 'postnatal',
- 'toddler': 'postnatal',
- 'adult': 'adult',
- 'adult_DN': 'adult'
- }
- # Map developmental states to numeric values
- peak_info_subset['dev_state_label'] = peak_info_subset['maxAcc_devState'].map(dev_state_mapping)
- human_peaks_subset = human_peaks_subset.merge(peak_info_subset[['peak', 'dev_state_label']], on='peak', how='left')
- # Set "Others" as reference
- human_peaks_subset['Cons_group_label'] = pd.Categorical(
- human_peaks_subset['Cons_group_label'],
- categories=['others', 'human_specific', 'primate_specific', 'conserved'],
- ordered=False
- )
- # Convert to categorical with '0' as the reference level
- human_peaks_subset['dev_state_label'] = pd.Categorical(
- human_peaks_subset['dev_state_label'],
- categories=['adult', 'postnatal', 'fetal', 'embryonic'],
- ordered=True
- )
- human_peaks_subset
- # %%
- print("\n" + "="*70)
- print("MAIN MODEL: Peak Type + Conservation + Covariates")
- print("="*70)
- formula = 'TE_overlap ~ C(peakType) + C(Cons_group_label) + C(dev_state_label) + GC_content'
- model_main = smf.logit(formula, data=human_peaks_subset).fit()
- print(model_main.summary())
- # Get odds ratios
- odds_ratios = np.exp(model_main.params)
- conf_int = np.exp(model_main.conf_int())
- conf_int['OR'] = odds_ratios
- conf_int.columns = ['2.5%', '97.5%', 'OR']
- print("\n" + "="*70)
- print("ODDS RATIOS AND 95% CONFIDENCE INTERVALS")
- print("="*70)
- print(conf_int)
- # %%
- model_main.pvalues
06a_regression_analysis_human.ipynb at commit d89c357, under CC-BY-NC-SA-4.0 · at the source
Overview
- Center for Molecular Biology of Heidelberg University (ZMBH), DKFZ-ZMBH Alliance, Heidelberg, Germany
- Present Address: Centre of Genomics, Evolution and Medicine (cGEM), Institute of Genomics, University of Tartu, Tartu, Estonia
- Present Address: Wellcome Sanger Institute, Cambridge, UK
- Present Address: Cambridge Stem Cell Institute and Department of Medicine, University of Cambridge, Cambridge, UK
Abstract
Transposable elements are hypothesized to have driven gene regulatory innovation, yet their contributions to primate brain development at the cell type level remain underexplored. Here, we use single-cell multiomics data from human, macaque, marmoset, and mouse cerebella to show that transposable element contributions to different cell types are shaped by varying degrees of constraints across cell types, as well as the preferential co-option of certain transposable elements in specific cell states. Using a sequence-based deep-learning model that predicts cell-type-specific chromatin accessibility, we systematically assess the co-option potential of transposable elements into cerebellar gene regulatory networks, identifying twelve transposable element subfamilies with complex regulatory sequences in their ancestral states that facilitate their co-option as cell-type-specific cis-regulatory elements. Preservation of these ancestral regulatory sequences, as well as the active chromatin environment surrounding the insertion site, is the major determinant of the accessibility of extant copies. Lineage-specific accessible copies contribute to human-specific gene expression. Broadly, we demonstrate how transposable elements can be flexibly co-opted into cell-type-specific gene regulatory networks, and introduce a generalizable analytical framework for dissecting their contribution to mammalian regulatory evolution.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 18 matches between paragraphs and lines of code.
kaessmannlab/cerebellum_te
d89c3571724b09827196d19ed0ca22d19a6cd6aa, 10 June 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
34 files
- 01_CRE_TE_overlap/
01_make_repeat_beds.ipyn , Jupyter, 299 lines, 1 matchb - 01_CRE_TE_overlap/
02a_species_specific_pea , Jupyter, 402 linesks_human.ipynb - 01_CRE_TE_overlap/
02b_species_specific_pea , Jupyter, 358 linesks_marmoset.ipynb - 01_CRE_TE_overlap/
02c_species_specific_pea , Jupyter, 355 linesks_mouse.ipynb - 01_CRE_TE_overlap/
03a_celltype_specific_pe , Jupyter, 141 lines, 1 matchaks_human.ipynb - 01_CRE_TE_overlap/
03b_celltype_specific_pe , Jupyter, 88 linesaks_marmoset.ipynb - 01_CRE_TE_overlap/
03c_celltype_specific_pe , Jupyter, 91 linesaks_mouse.ipynb - 01_CRE_TE_overlap/
04_catlas_analysis.ipynb , Jupyter, 220 lines - 01_CRE_TE_overlap/
05_te_subfamily_enrichme , Jupyter, 622 linesnt.ipynb - 01_CRE_TE_overlap/
06a_regression_analysis_ , Jupyter, 582 lines, 3 matcheshuman.ipynb - 01_CRE_TE_overlap/
06b_regression_analysis_ , Jupyter, 429 lines, 2 matchesmarmoset.ipynb - 01_CRE_TE_overlap/
06c_regression_analysis_ , Jupyter, 431 linesmouse.ipynb - 01_CRE_TE_overlap/
07_regression_analysis_c , Jupyter, 260 lines, 1 matchatlas.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 147 lines01_make_msa_human.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 114 lines01_make_msa_marmoset.ipy nb - 02_CRE_potential_screeni
ng/ , Jupyter, 111 lines01_make_msa_mouse.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 535 lines02_potential_enrichment_ screening.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 200 lines, 2 matches03_visualize_screening_c andidates.ipynb - 03_HERVL/
01_find_accessible_copie , Python, 594 lines, 3 matchess.py - 03_HERVL/
02_copy_analysis.py , Python, 572 lines - 03_HERVL/
03_HERVL_accessibilitiy. , Jupyter, 579 lines, 1 matchipynb - 03_HERVL/
04_HERVL_accessibility_o , Jupyter, 259 linesverlap.ipynb - 03_HERVL/
05_ERVL_family_tree.ipyn , Jupyter, 168 linesb - 03_HERVL/
06_ERVL_family_info.ipyn , Jupyter, 257 linesb - 03_HERVL/
07_ERVL_family_accessibi , Jupyter, 234 lineslity.ipynb - 03_HERVL/
08_nucleotide_preservati , Jupyter, 510 lines, 2 matcheson.ipynb - 03_HERVL/
09_KZFP.ipynb , Jupyter, 318 lines - 03_HERVL/
10_AB_compartment.ipynb , Jupyter, 204 lines - 03_HERVL/
11_distance_analysis.ipy , Jupyter, 348 lines, 1 matchnb - 03_HERVL/
12_HERVL_regression_anal , Jupyter, 294 linesysis.ipynb - 03_HERVL/
13_HERVL_orthology.ipynb , Jupyter, 210 lines - 03_HERVL/
14_HERVL_gene_expression , Jupyter, 300 lines, 1 match.ipynb - LICENSE, License, 11 lines
- README.md, Text, 76 lines
Zenodo 19354896
Availability: 1 check, the latest on 27 September 2026: the link answers (HTTP 200)
- 27 September 2026: the link answers (HTTP 200)
22 files
- 01_CRE_TE_overlap/
01_make_repeat_beds.ipyn , Jupyter, 299 linesb - 01_CRE_TE_overlap/
02a_species_specific_pea , Jupyter, 399 linesks_human.ipynb - 01_CRE_TE_overlap/
02b_species_specific_pea , Jupyter, 355 linesks_marmoset.ipynb - 01_CRE_TE_overlap/
02c_species_specific_pea , Jupyter, 355 linesks_mouse.ipynb - 01_CRE_TE_overlap/
03a_celltype_specific_pe , Jupyter, 141 linesaks_human.ipynb - 01_CRE_TE_overlap/
03b_celltype_specific_pe , Jupyter, 88 linesaks_marmoset.ipynb - 01_CRE_TE_overlap/
03c_celltype_specific_pe , Jupyter, 91 linesaks_mouse.ipynb - 01_CRE_TE_overlap/
04_catlas_analysis.ipynb , Jupyter, 220 lines - 01_CRE_TE_overlap/
05_te_subfamily_enrichme , Jupyter, 619 linesnt.ipynb - 01_CRE_TE_overlap/
06a_regression_analysis_ , Jupyter, 582 lineshuman.ipynb - 01_CRE_TE_overlap/
06b_regression_analysis_ , Jupyter, 429 linesmarmoset.ipynb - 01_CRE_TE_overlap/
06c_regression_analysis_ , Jupyter, 431 linesmouse.ipynb - 01_CRE_TE_overlap/
07_regression_analysis_c , Jupyter, 260 linesatlas.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 147 lines01_make_msa_human.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 114 lines01_make_msa_marmoset.ipy nb - 02_CRE_potential_screeni
ng/ , Jupyter, 111 lines01_make_msa_mouse.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 535 lines02_potential_enrichment_ screening.ipynb - 02_CRE_potential_screeni
ng/ , Jupyter, 200 lines03_visualize_screening_c andidates.ipynb - 03_HERVL/
01_HERVL_copy_analysis_h , Jupyter, 478 linesuman.ipynb - 03_HERVL/
01_HERVL_copy_analysis_m , Jupyter, 384 linesacaque.ipynb - 03_HERVL/
01_HERVL_copy_analysis_m , Jupyter, 406 linesarmoset.ipynb - 03_HERVL/
02_HERVL_info_human.ipyn , Jupyter, 501 linesb - repository limit reached (2,000 files or 30 MB): the rest is at the source
Code availability
All original code is available on GitLab [https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 2 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 54 scripts, each with its path and the digest of its content;
- 18 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- arrayexpress:E-MTAB-9765
, at ArrayExpress; found in “Data availability”
Data Availability Statement
Previously published datasets used in this study are available in the ArrayExpress database under accession codes E-MTAB-9765 (https://
All original code is available on GitLab [https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 5 keywords, 12 MeSH terms, 2 funders, 116 references.
Cite
This paper
Yamada, T., Sepp, M., Sarropoulos, I., & Kaessmann, H. (2026). Gene regulatory innovations from transposable elements in primate cerebellum development. Nature communications, 17(1), 7598. https://
BibTeX
@article{yamada2026gene,
author = {Yamada, Tetsuya and Sepp, Mari and Sarropoulos, Ioannis and Kaessmann, Henrik},
title = {{Gene regulatory innovations from transposable elements in primate cerebellum development}},
journal = {Nature communications},
year = {2026},
month = jul,
volume = {17},
number = {1},
pages = {7598},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42527403},
pmcid = {PMC13421694}
}
RIS
TY - JOUR
AU - Yamada, Tetsuya
AU - Sepp, Mari
AU - Sarropoulos, Ioannis
AU - Kaessmann, Henrik
TI - Gene regulatory innovations from transposable elements in primate cerebellum development
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 7598
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Gene regulatory innovations from transposable elements in primate cerebellum development",
"container-title": "Nature communications",
"author": [
{
"family": "Yamada",
"given": "Tetsuya"
},
{
"family": "Sepp",
"given": "Mari"
},
{
"family": "Sarropoulos",
"given": "Ioannis"
},
{
"family": "Kaessmann",
"given": "Henrik"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "7598",
"DOI": "10.1038/
"PMID": "42527403",
"PMCID": "PMC13421694",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
30
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: pysam, Biopython, BEDTools, 9 other tools, genetics / omics, mouse, 8 references
- [2] doi:10.1186/s13059-026-04177-w [code]
- Genomic sequence evolution underlying human neocortical interareal diversification.Journal: Genome biologyIn common: pysam, BEDTools, TensorFlow, 6 other tools, non-human primate, genetics / omics, mouse, 1 other category, 6 references
- [3] doi:10.1186/s13059-026-04050-w
- Transposable element-mediated evolutionary expansion of Sox2- and Brn2-binding regulatory modules for mammalian neural-cell differentiation.Journal: Genome biologyIn common: 12 references
- [4] doi:10.1038/s42003-026-10462-y [code]
- SpaDC enables sequence-based integrative analysis and regulatory inference of spatial chromatin accessibility data.Journal: Communications biologyIn common: pysam, Biopython, SHAP, 6 other tools, genetics / omics, mouse, 3 references
- [5] doi:10.1126/sciadv.aed2952 [code]
- Activation of transposable elements is linked to a region- and cell type-specific interferon response in Parkinson's disease.Journal: Science advancesIn common: pysam, Biopython, BEDTools, 5 other tools, cellular / molecular, 4 references
- [6] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: Biopython, SHAP, Keras, 8 other tools, mouse, cellular / molecular
- [7] doi:10.1016/j.celrep.2026.117110 [code]
- Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.Journal: Cell reportsIn common: pysam, BEDTools, statsmodels, 5 other tools, genetics / omics, cellular / molecular, 3 references
- [8] doi:10.1038/s44400-026-00094-8 [code]
- Haplotype-resolved DNA methylation at the &
lt;i& gt;APOE& lt;/ i& gt; locus identifies allele-specific epigenetic signatures relevant to Alzheimer's disease risk. Journal: NPJ dementiaIn common: pysam, BEDTools, statsmodels, 6 other tools, genetics / omics, cellular / molecular, 2 references - [9] doi:10.1016/j.xhgg.2026.100629 [code]
- Positive selection on brain cis-regulatory elements in the human lineage drives changes in gene expression and susceptibility to neuropsychiatric disorders.Journal: HGG advancesIn common: Biopython, statsmodels, seaborn, 5 other tools, cellular / molecular, 3 references
- [10] doi:10.1038/s41592-026-03211-w [code]
- Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.Journal: Nature methodsIn common: pysam, Biopython, BEDTools, 7 other tools, genetics / omics, mouse
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 2 repositories of the authors' code, each at its verified commit and with its license, 54 scripts, and 18 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:da1239a5770f3236…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[.
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
