Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.
The 14 matches · 3 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
- [1] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/benchmark.py, lines 88–125 · score 0.87 · passive aggressive classifier, support vector machine, decision tree, LightGBM, random forest, perceptron
- [2] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/draw_pCBS.py, the whole file · a weak match · score 0.86 · passive aggressive classifier, support vector machine, decision tree, LightGBM, random forest, perceptron
- [3] § Materials and methods › Artificial intelligence (AI) model › Tokenization of input sequences. ↔ AI/preprocess/data_collator.py, the whole file · a weak match · score 0.86 · undefined amino acids, padding token, 0–25, end token, start token, protein sequences
- [4] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/benchmark.py, lines 88–125 · score 0.84 · passive aggressive classifier, DeepZF, LightGBM, Random Forest, benchmarked, boosting
- [5] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ scripts/formal/test.sh, lines 13–57 · score 0.82 · passive aggressive classifier, DeepZF, LightGBM, Random Forest, boosting, perceptron
- [6] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 13–147 · score 0.78 · rotary positional embeddings, Multi head attention, RMS normalization, residual, dropout, module
- [7] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/model.py, lines 27–144 · score 0.72 · ProteinBert, attention layer, RMS normalized, C2H2, token, embedding
- [8] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 238–291 · score 0.71 · DNA protein cross, RMS normalized, cross attention, protein sequences, module, token
- [9] § Materials and methods › Artificial intelligence (AI) model › Tokenization of input sequences. ↔ AI/preprocess/data_collator.py, the whole file · a weak match · score 0.71 · padding token, 0–6, tokenized, C2H2, nucleotides, masking
- [10] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 293–367 · score 0.62 · RMSNorm, cross attention, GELU, FFN, dropout, head
- [11] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 369–413 · score 0.60 · feedforward network, cross attention, FFN, embeddings, COP, protein
- [12] § Materials and methods › Artificial intelligence (AI) model › Processing of mouse C2H2-ZFP data. ↔ AI/dataset/run.sh, lines 27–32 · score 0.59 · ft_zn_fing, organism, C2H2, UniProt, protein
- [13] § Materials and methods › Artificial intelligence (AI) model › Processing of mouse C2H2-ZFP data. ↔ AI/dataset/scripts/download_uniprot_C2H2_protein_table.py, lines 15–45 · score 0.59 · ft_zn_fing, UniProt, organism, id, C2H2, protein
- [14] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/model.py, lines 27–144 · score 0.59 · RMSNorm, logits, GELU, FFN, dropout, head
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 413 lines · 14 KB · MIT · 4 matches
- import torch
- from common_ai.non_causality_hyena import HyenaOperator
- from common_ai.utils import Residual
- from einops import einsum, repeat
- from einops.layers.torch import EinMix
- from torch import nn
- # torch does not import opt_einsum as backend by default. import opt_einsum manually will enable it.
- from torch.backends import opt_einsum
- from torchtune.modules import MultiHeadAttention, RotaryPositionalEmbeddings
- class SecondEncoder(nn.Module):
- def __init__(
- self,
- vocab: int,
- max_num_tokens: int,
- dim_emb: int,
- dim_head: int,
- heads: int,
- depth: int,
- dim_ffn: int,
- dropout: float,
- use_hyena: bool,
- hyena_order: int,
- hyena_filter_order: int,
- ) -> None:
- super().__init__()
- self.depth = depth
- self.use_hyena = use_hyena
- self.embed = nn.Sequential(
- nn.Embedding(
- vocab,
- dim_emb,
- ),
- nn.Dropout(dropout),
- )
- self.rms_norms = nn.ModuleList([nn.RMSNorm(dim_emb) for _ in range(depth)])
- self.self_attentions = nn.ModuleList(
- [
- (
- MultiHeadAttention(
- embed_dim=dim_emb,
- num_heads=heads,
- num_kv_heads=heads,
- head_dim=dim_head,
- q_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- k_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- v_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- output_proj=EinMix(
- "b s nhd -> b s d",
- weight_shape="nhd d",
- bias_shape="d",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- pos_embeddings=RotaryPositionalEmbeddings(
- dim=dim_head, max_seq_len=max_num_tokens
- ),
- max_seq_len=max_num_tokens,
- is_causal=False,
- attn_dropout=dropout,
- )
- if not use_hyena
- else HyenaOperator(
- d_model=dim_emb,
- l_max=max_num_tokens,
- order=hyena_order,
- filter_order=hyena_filter_order,
- dropout=dropout,
- filter_dropout=dropout,
- )
- )
- for _ in range(depth)
- ]
- )
- self.ffns = nn.ModuleList(
- [
- Residual(
- nn.Sequential(
- nn.RMSNorm(dim_emb),
- EinMix(
- "b s d -> b s d_f",
- weight_shape="d d_f",
- bias_shape="d_f",
- d=dim_emb,
- d_f=dim_ffn,
- ),
- nn.GELU(),
- EinMix(
- "b s d_f -> b s d",
- weight_shape="d_f d",
- bias_shape="d",
- d=dim_emb,
- d_f=dim_ffn,
- ),
- nn.Dropout(dropout),
- )
- )
- for _ in range(depth)
- ]
- )
- self.last_rms_norm = nn.RMSNorm(dim_emb)
- def forward(self, second_ids: torch.Tensor) -> torch.Tensor:
- mask = second_ids != 0
- embs = self.embed(second_ids)
- for i in range(self.depth):
- embs_rms = self.rms_norms[i](embs)
- if not self.use_hyena:
- embs = embs + self.self_attentions[i](
- x=embs_rms,
- y=embs_rms,
- mask=einsum(mask, mask, "b s1, b s2 -> b s1 s2"),
- )
- else:
- # self_attentions[i] is actually hyena
- embs = embs + self.self_attentions[i](embs_rms)
- embs = self.ffns[i](embs)
- embs = self.last_rms_norm(embs)
- return embs, mask
- class DNAEncoder(nn.Module):
- def __init__(
- self,
- vocab: int,
- max_num_tokens: int,
- dim_emb: int,
- dim_head: int,
- heads: int,
- depth: int,
- dim_ffn: int,
- dropout: float,
- use_hyena: bool,
- hyena_order: int,
- hyena_filter_order: int,
- ) -> None:
- super().__init__()
- self.depth = depth
- self.use_hyena = use_hyena
- self.embed = nn.Sequential(
- nn.Embedding(
- vocab,
- dim_emb,
- ),
- nn.Dropout(dropout),
- )
- self.rms_norms = nn.ModuleList([nn.RMSNorm(dim_emb) for _ in range(depth)])
- self.self_attentions = nn.ModuleList(
- [
- (
- MultiHeadAttention(
- embed_dim=dim_emb,
- num_heads=heads,
- num_kv_heads=heads,
- head_dim=dim_head,
- q_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- k_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- v_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- output_proj=EinMix(
- "b s nhd -> b s d",
- weight_shape="nhd d",
- bias_shape="d",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- # DNA starts with CLS token, so increase max_seq_len by 1
- pos_embeddings=RotaryPositionalEmbeddings(
- dim=dim_head, max_seq_len=max_num_tokens + 1
- ),
- max_seq_len=max_num_tokens + 1,
- is_causal=False,
- attn_dropout=dropout,
- )
- if not use_hyena
- else HyenaOperator(
- d_model=dim_emb,
- l_max=max_num_tokens + 1,
- order=hyena_order,
- filter_order=hyena_filter_order,
- dropout=dropout,
- filter_dropout=dropout,
- )
- )
- for _ in range(depth)
- ]
- )
- self.dna_protein_rms_norms = nn.ModuleList(
- [nn.RMSNorm(dim_emb) for _ in range(depth)]
- )
- self.dna_protein_cross_attentions = nn.ModuleList(
- [
- MultiHeadAttention(
- embed_dim=dim_emb,
- num_heads=heads,
- num_kv_heads=heads,
- head_dim=dim_head,
- q_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- k_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- v_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- output_proj=EinMix(
- "b s nhd -> b s d",
- weight_shape="nhd d",
- bias_shape="d",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- # protein bert clip the protein sequence by <start> and <end> tokens, so increase max_seq_len by 2
- pos_embeddings=RotaryPositionalEmbeddings(
- dim=dim_head, max_seq_len=max_num_tokens + 2
- ),
- max_seq_len=max_num_tokens + 2,
- is_causal=False,
- attn_dropout=dropout,
- )
- for _ in range(depth)
- ]
- )
- self.dna_second_rms_norms = nn.ModuleList(
- [nn.RMSNorm(dim_emb) for _ in range(depth)]
- )
- self.dna_second_cross_attentions = nn.ModuleList(
- [
- MultiHeadAttention(
- embed_dim=dim_emb,
- num_heads=heads,
- num_kv_heads=heads,
- head_dim=dim_head,
- q_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- k_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- v_proj=EinMix(
- "b s d -> b s nhd",
- weight_shape="d nhd",
- bias_shape="nhd",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- output_proj=EinMix(
- "b s nhd -> b s d",
- weight_shape="nhd d",
- bias_shape="d",
- d=dim_emb,
- nhd=heads * dim_head,
- ),
- # DNA starts with CLS token, so increase max_seq_len by 1
- pos_embeddings=RotaryPositionalEmbeddings(
- dim=dim_head, max_seq_len=max_num_tokens + 1
- ),
- max_seq_len=max_num_tokens + 1,
- is_causal=False,
- attn_dropout=dropout,
- )
- for _ in range(depth)
- ]
- )
- self.ffns = nn.ModuleList(
- [
- Residual(
- nn.Sequential(
- nn.RMSNorm(dim_emb),
- EinMix(
- "b s d -> b s d_f",
- weight_shape="d d_f",
- bias_shape="d_f",
- d=dim_emb,
- d_f=dim_ffn,
- ),
- nn.GELU(),
- EinMix(
- "b s d_f -> b s d",
- weight_shape="d_f d",
- bias_shape="d",
- d=dim_emb,
- d_f=dim_ffn,
- ),
- nn.Dropout(dropout),
- )
- )
- for _ in range(depth)
- ]
- )
- self.last_rms_norm = nn.RMSNorm(dim_emb)
- def forward(
- self,
- dna_ids: torch.Tensor,
- protein_embs: torch.Tensor,
- second_embs: torch.Tensor,
- second_mask: torch.Tensor,
- ) -> torch.Tensor:
- dna_mask = dna_ids != 0
- dna_embs = self.embed(dna_ids)
- for i in range(self.depth):
- dna_embs_rms = self.rms_norms[i](dna_embs)
- if not self.use_hyena:
- dna_embs = dna_embs + self.self_attentions[i](
- x=dna_embs_rms,
- y=dna_embs_rms,
- mask=einsum(dna_mask, dna_mask, "b s1, b s2 -> b s1 s2"),
- )
- else:
- # self_attentions[i] is actually hyena
- dna_embs = dna_embs + self.self_attentions[i](dna_embs_rms)
- # DNA and protein cross-attention
- dna_embs_rms = self.dna_protein_rms_norms[i](dna_embs)
- dna_embs = dna_embs + self.dna_protein_cross_attentions[i](
- x=dna_embs_rms,
- y=protein_embs,
- mask=repeat(dna_mask, "b s1 -> b s1 s2", s2=protein_embs.shape[1]),
- )
- # DNA and secondary structure cross-attention
- dna_embs_rms = self.dna_second_rms_norms[i](dna_embs)
- dna_embs = dna_embs + self.dna_second_cross_attentions[i](
- x=dna_embs_rms,
- y=second_embs,
- mask=einsum(dna_mask, second_mask, "b s1, b s2 -> b s1 s2"),
- )
- # DNA feedforward network
- dna_embs = self.ffns[i](dna_embs)
- # only use the embedding for CLS token
- dna_embs = dna_embs[:, 0, :]
- dna_embs = self.last_rms_norm(dna_embs)
- return dna_embs
encoder.py at commit 353997b, under MIT · at the source
Overview
- Center for Comparative Biomedicine, Key Laboratory of Systems Biomedicine (MOE), Institute of Systems Biomedicine, Shanghai Jiao Tong University, Shanghai, China
- Shanghai Key Laboratory of Gene Editing and Cell-based Immunotherapy for Hematological Diseases, State Key Laboratory of Medical Genomics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
Abstract
Zinc finger proteins (ZFPs or ZNFs) constitute the largest family of transcription factors in mammals; however, their regulatory mechanism remains largely elusive. Here we propose COP (C2H2-ZFP occupancy predictor), a deep learning-based heuristic screening tool that integrates DNA sequence with protein primary and secondary features to assess ZFP genomic enrichments. Applying COP to the mouse clustered protocadherin (cPcdh) gene locus, we identified dozens of C2H2-ZFPs potentially involved in CTCF-mediated gene regulation with Wiz (widely interspaced zinc finger-containing protein) having the highest number of 12 ZFs. We confirmed Wiz enrichments at all of the CTCF-binding site (CBS) elements across the three Pcdh clusters by Myc-tagging the endogenous Wiz gene. Genetic experiments revealed significant increases of expression levels of the cPcdh genes upon Wiz deletion in both neuronal cells in vitro and in mouse brain in vivo. Finally, integrated ChIP-seq, RNA-seq, and 4C-seq analyses demonstrated that Wiz regulates CTCF/
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.
ljw20180420/COP
353997b21fa1b80fc71654e48a3c406c70572702, 7 September 2026Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
80 files
- AI/
__init__.py , Python, 1 line - AI/
dataset.py , Python, 31 lines - AI/
dataset/ , Python, 77 linesextra/ scripts/ transform_protein_featur e.py - AI/
dataset/ , Shell, 382 lines, 1 matchrun.sh - AI/
dataset/ , Python, 34 linesscripts/ assess_SRR.py - AI/
dataset/ , Python, 19 linesscripts/ assess_peak_width.py - AI/
dataset/ , Python, 43 linesscripts/ balance_data.py - AI/
dataset/ , Python, 44 linesscripts/ download_alphafoldDB_mmc if.py - AI/
dataset/ , Python, 45 lines, 1 matchscripts/ download_uniprot_C2H2_pr otein_table.py - AI/
dataset/ , Python, 15 linesscripts/ generate_unittest_data.p y - AI/
dataset/ , Python, 63 linesscripts/ get_pCBS.py - AI/
dataset/ , Python, 60 linesscripts/ parse_protein_feature.py - AI/
dataset/ , Python, 23 linesscripts/ sample_data.py - AI/
dataset/ , Python, 22 linesscripts/ select_SRR.py - AI/
dataset/ , Python, 27 linesscripts/ select_seed.py - AI/
dataset/ , Python, 22 linesscripts/ split_data.py - AI/
gradio_fn.py , Python, 58 lines - AI/
inference.py , Python, 44 lines - AI/
metric.py , Python, 326 lines - AI/
metric/ , Python, 106 linesaccuracy.py - AI/
metric/ , Python, 134 linesbrier_score.py - AI/
metric/ , Python, 88 linesconfusion_matrix.py - AI/
metric/ , Python, 130 linesf1.py - AI/
metric/ , Python, 140 linesmatthews_correlation.py - AI/
metric/ , Python, 145 linesprecision.py - AI/
metric/ , Python, 135 linesrecall.py - AI/
metric/ , Python, 191 linesroc_auc.py - AI/
preprocess/ , Python, 1 lineCOP/ __init__.py - AI/
preprocess/ , Python, 413 lines, 4 matchesCOP/ encoder.py - AI/
preprocess/ , Python, 235 lines, 2 matchesCOP/ model.py - AI/
preprocess/ , Python, 1 lineDeepZF/ __init__.py - AI/
preprocess/ , Python, 74 linesDeepZF/ data_collator.py - AI/
preprocess/ , Python, 337 linesDeepZF/ model.py - AI/
preprocess/ , Python, 1 lineLightGBM/ __init__.py - AI/
preprocess/ , Python, 205 linesLightGBM/ model.py - AI/
preprocess/ , Python, 1 lineScikit/ __init__.py - AI/
preprocess/ , Python, 346 linesScikit/ model.py - AI/
preprocess/ , Python, 1 lineXGBoost/ __init__.py - AI/
preprocess/ , Python, 306 linesXGBoost/ model.py - AI/
preprocess/ , Python, 1 line__init__.py - AI/
preprocess/ , Python, 139 lines, 2 matchesdata_collator.py - AI/
preprocess/ , Python, 71 linesmodel.py - DeepZF/
BindZFPredictor.ipynb , Jupyter, 30 lines - DeepZF/
BindZF_predictor/ , Python, 54 linescode/ create_zf_pred_df_and_ca l_auc.py - DeepZF/
BindZF_predictor/ , Python, 249 linescode/ finetuning.py - DeepZF/
BindZF_predictor/ , Python, 140 linescode/ main_bindzfpredictor_pre dict.py - DeepZF/
BindZF_predictor/ , Python, 94 linescode/ main_loo_bindzfpredictor .py - DeepZF/
PWMpredictor.ipynb , Jupyter, 31 lines - DeepZF/
PWMpredictor/ , Python, 83 linescode/ create_b1h_input_and_lab el_data.py - DeepZF/
PWMpredictor/ , Python, 39 linescode/ create_c_rc_input_and_la bel.py - DeepZF/
PWMpredictor/ , Python, 321 linescode/ functions.py - DeepZF/
PWMpredictor/ , Python, 83 linescode/ main_PWMpredictor.py - DeepZF/
PWMpredictor/ , Python, 71 linescode/ main_loo_PWMpredictor.py - DeepZF/
PWMpredictor/ , Python, 16 linescode/ models_PWMpredictor.py - DeepZF/
PWMpredictor/ , Python, 12 linescode/ models_loo_PWMpredictor. py - DeepZF/
PWMpredictor/ , Python, 19 linescode/ utiles_loo_PWMpredictor. py - DeepZF/
PWMpredictor/ , Python, 152 linescode/ utiles_models_PWMpredict or.py - DeepZF/
PWMpredictor/ , Python, 97 linescode/ utiles_models_loo_PWMpre dictor.py - DeepZF/
PWMpredictor/ , Python, 114 linesresult_analysis/ create_mosbat_input_prot ein_bert.py - DeepZF/
PWMpredictor/ , Shell, 27 linesresult_analysis/ eval_DeepZF.sh - DeepZF/
PWMpredictor/ , Python, 51 linesresult_analysis/ eval_PWMpredictor.py - DeepZF/
PWMpredictor/ , Python, 49 linesresult_analysis/ eval_mosbat.py - DeepZF/
PWMpredictor/ , Python, 27 linesresult_analysis/ pearson_cor_per_pos.py - app_space.py, Python, 12 lines
- paper/
benchmark.py , Python, 125 lines, 2 matches - paper/
draw_pCBS.py , Python, 50 lines, 1 match - paper/
extract_ctcf_binding_sit , Python, 50 linese.py - paper/
pCBS.py , Python, 111 lines - run.py, Python, 72 lines
- scripts/
formal/ , Shell, 10 linesapp.sh - scripts/
formal/ , Shell, 43 lineshpo.sh - scripts/
formal/ , Shell, 45 linesinfer.sh - scripts/
formal/ , Python, 53 linesspace.py - scripts/
formal/ , Shell, 57 lines, 1 matchtest.sh - scripts/
formal/ , Shell, 103 linestrain.sh - scripts/
formal/ , Shell, 29 linesupload.sh - scripts/
formal/ , Shell, 19 linesupload_dataset.sh - scripts/
unittest/ , Shell, 41 lineshpo.sh - LICENSE.md, License, 9 lines
- README.md, Text, 72 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 78 scripts, each with its path and the digest of its content;
- 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability
The raw data of RNA-seq, ChIP-seq, and QHR-4C generated in this study were deposited in the Genome Sequence Archive (GSA) at the National Genomics Data Center (NGDC), China National Center for Bioinformation (CNCB), under accession number CRA037522. The source code of the COP has been uploaded to GitHub (https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 14 MeSH terms, 2 funders, 78 references.
Cite
This paper
Li, T., Li, J., Wang, L., Huang, H., & Wu, Q. (2026). Wiz regulates clustered protocadherin genes by restricting CTCF/
BibTeX
@article{li2026wiz,
author = {Li, Tianjie and Li, Jingwei and Wang, Leyang and Huang, Haiyan and Wu, Qiang},
title = {{Wiz regulates clustered protocadherin genes by restricting CTCF/
journal = {PLoS genetics},
year = {2026},
month = jul,
volume = {22},
number = {7},
pages = {e1012242},
publisher = {PLOS},
issn = {1553-7390},
doi = {10.1371/
url = {https://
pmid = {42461945},
pmcid = {PMC13395409}
}
RIS
TY - JOUR
AU - Li, Tianjie
AU - Li, Jingwei
AU - Wang, Leyang
AU - Huang, Haiyan
AU - Wu, Qiang
TI - Wiz regulates clustered protocadherin genes by restricting CTCF/
T2 - PLoS genetics
J2 - PLoS Genet
PY - 2026
DA - 2026/
VL - 22
IS - 7
SP - e1012242
SN - 1553-7390
PB - PLOS
DO - 10.1371/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1371/
"type": "article-journal",
"title": "Wiz regulates clustered protocadherin genes by restricting CTCF/
"container-title": "PLoS genetics",
"author": [
{
"family": "Li",
"given": "Tianjie"
},
{
"family": "Li",
"given": "Jingwei"
},
{
"family": "Wang",
"given": "Leyang"
},
{
"family": "Huang",
"given": "Haiyan"
},
{
"family": "Wu",
"given": "Qiang"
}
],
"container-title-short":
"volume": "22",
"issue": "7",
"page": "e1012242",
"DOI": "10.1371/
"PMID": "42461945",
"PMCID": "PMC13395409",
"ISSN": "1553-7390",
"publisher": "PLOS",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
16
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.3389/fgene.2026.1807347 [code]
- Epigenetic regulators are preferentially coordinated with protocadherin gene expression across the human brain: a genome-wide co-expression analysis.Journal: Frontiers in geneticsIn common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 9 references
- [2] doi:10.3389/fpsyg.2026.1774068 [code]
- Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.Journal: Frontiers in psychologyIn common: LightGBM, XGBoost, Keras, 8 other tools
- [3] doi:10.1038/s41592-026-03057-2 [code]
- CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.Journal: Nature methodsIn common: BEDTools, Keras, TensorFlow, 7 other tools, mouse, 1 reference
- [4] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: BEDTools, Keras, TensorFlow, 6 other tools, mouse, cellular / molecular, 1 reference
- [5] doi:10.7554/elife.110588 [code]
- Opening the black box toward a modular approach to spike sorting.Journal: eLifeIn common: LightGBM, XGBoost, TensorFlow, 7 other tools, mouse
- [6] doi:10.1038/s42003-026-10957-8 [code]
- Brain defence by the extracellular matrix protein Cochlin.Journal: Communications biologyIn common: XGBoost, Keras, TensorFlow, 7 other tools, mouse, cellular / molecular
- [7] doi:10.1038/s41598-026-48613-0 [code]
- An snRNA-seq aging clock for the fruit fly head sheds light on sex-biased aging.Journal: Scientific reportsIn common: XGBoost, Keras, TensorFlow, 7 other tools, cellular / molecular
- [8] doi:10.1371/journal.pcbi.1014615 [code]
- Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.Journal: PLoS computational biologyIn common: XGBoost, Keras, TensorFlow, 7 other tools
- [9] doi:10.1093/bib/bbag118 [code]
- Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.Journal: Briefings in bioinformaticsIn common: LightGBM, XGBoost, PyTorch, 6 other tools, cellular / molecular
- [10] doi:10.1126/sciadv.aed3650 [code]
- Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &
lt;i& gt;m/ z& lt;/ i& gt; mapping and exploration. Journal: Science advancesIn common: XGBoost, Keras, TensorFlow, 6 other tools, mouse
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 78 scripts, and 14 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:a6997ad74b48fa2a…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
