OSCR

Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.

Code ↔ Paper

14 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 14 matches · 3 of them tie a paragraph to a whole file, not to given lines: weak matches, whose lines are not tinted
  1. [1] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/benchmark.py, lines 88–125 · score 0.87 · passive aggressive classifier, support vector machine, decision tree, LightGBM, random forest, perceptron
  2. [2] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/draw_pCBS.py, the whole file · a weak match · score 0.86 · passive aggressive classifier, support vector machine, decision tree, LightGBM, random forest, perceptron
  3. [3] § Materials and methods › Artificial intelligence (AI) model › Tokenization of input sequences. ↔ AI/preprocess/data_collator.py, the whole file · a weak match · score 0.86 · undefined amino acids, padding token, 0–25, end token, start token, protein sequences
  4. [4] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ paper/benchmark.py, lines 88–125 · score 0.84 · passive aggressive classifier, DeepZF, LightGBM, Random Forest, benchmarked, boosting
  5. [5] § Materials and methods › Artificial intelligence (AI) model › Benchmarking. ↔ scripts/formal/test.sh, lines 13–57 · score 0.82 · passive aggressive classifier, DeepZF, LightGBM, Random Forest, boosting, perceptron
  6. [6] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 13–147 · score 0.78 · rotary positional embeddings, Multi head attention, RMS normalization, residual, dropout, module
  7. [7] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/model.py, lines 27–144 · score 0.72 · ProteinBert, attention layer, RMS normalized, C2H2, token, embedding
  8. [8] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 238–291 · score 0.71 · DNA protein cross, RMS normalized, cross attention, protein sequences, module, token
  9. [9] § Materials and methods › Artificial intelligence (AI) model › Tokenization of input sequences. ↔ AI/preprocess/data_collator.py, the whole file · a weak match · score 0.71 · padding token, 0–6, tokenized, C2H2, nucleotides, masking
  10. [10] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 293–367 · score 0.62 · RMSNorm, cross attention, GELU, FFN, dropout, head
  11. [11] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/encoder.py, lines 369–413 · score 0.60 · feedforward network, cross attention, FFN, embeddings, COP, protein
  12. [12] § Materials and methods › Artificial intelligence (AI) model › Processing of mouse C2H2-ZFP data. ↔ AI/dataset/run.sh, lines 27–32 · score 0.59 · ft_zn_fing, organism, C2H2, UniProt, protein
  13. [13] § Materials and methods › Artificial intelligence (AI) model › Processing of mouse C2H2-ZFP data. ↔ AI/dataset/scripts/download_uniprot_C2H2_protein_table.py, lines 15–45 · score 0.59 · ft_zn_fing, UniProt, organism, id, C2H2, protein
  14. [14] § Materials and methods › Artificial intelligence (AI) model › COP architecture. ↔ AI/preprocess/COP/model.py, lines 27–144 · score 0.59 · RMSNorm, logits, GELU, FFN, dropout, head

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 413 lines · 14 KB · MIT · 4 matches

  1. import torch
  2. from common_ai.non_causality_hyena import HyenaOperator
  3. from common_ai.utils import Residual
  4. from einops import einsum, repeat
  5. from einops.layers.torch import EinMix
  6. from torch import nn
  7. # torch does not import opt_einsum as backend by default. import opt_einsum manually will enable it.
  8. from torch.backends import opt_einsum
  9. from torchtune.modules import MultiHeadAttention, RotaryPositionalEmbeddings
  10. class SecondEncoder(nn.Module):
  11. def __init__(
  12. self,
  13. vocab: int,
  14. max_num_tokens: int,
  15. dim_emb: int,
  16. dim_head: int,
  17. heads: int,
  18. depth: int,
  19. dim_ffn: int,
  20. dropout: float,
  21. use_hyena: bool,
  22. hyena_order: int,
  23. hyena_filter_order: int,
  24. ) -> None:
  25. super().__init__()
  26. self.depth = depth
  27. self.use_hyena = use_hyena
  28. self.embed = nn.Sequential(
  29. nn.Embedding(
  30. vocab,
  31. dim_emb,
  32. ),
  33. nn.Dropout(dropout),
  34. )
  35. self.rms_norms = nn.ModuleList([nn.RMSNorm(dim_emb) for _ in range(depth)])
  36. self.self_attentions = nn.ModuleList(
  37. [
  38. (
  39. MultiHeadAttention(
  40. embed_dim=dim_emb,
  41. num_heads=heads,
  42. num_kv_heads=heads,
  43. head_dim=dim_head,
  44. q_proj=EinMix(
  45. "b s d -> b s nhd",
  46. weight_shape="d nhd",
  47. bias_shape="nhd",
  48. d=dim_emb,
  49. nhd=heads * dim_head,
  50. ),
  51. k_proj=EinMix(
  52. "b s d -> b s nhd",
  53. weight_shape="d nhd",
  54. bias_shape="nhd",
  55. d=dim_emb,
  56. nhd=heads * dim_head,
  57. ),
  58. v_proj=EinMix(
  59. "b s d -> b s nhd",
  60. weight_shape="d nhd",
  61. bias_shape="nhd",
  62. d=dim_emb,
  63. nhd=heads * dim_head,
  64. ),
  65. output_proj=EinMix(
  66. "b s nhd -> b s d",
  67. weight_shape="nhd d",
  68. bias_shape="d",
  69. d=dim_emb,
  70. nhd=heads * dim_head,
  71. ),
  72. pos_embeddings=RotaryPositionalEmbeddings(
  73. dim=dim_head, max_seq_len=max_num_tokens
  74. ),
  75. max_seq_len=max_num_tokens,
  76. is_causal=False,
  77. attn_dropout=dropout,
  78. )
  79. if not use_hyena
  80. else HyenaOperator(
  81. d_model=dim_emb,
  82. l_max=max_num_tokens,
  83. order=hyena_order,
  84. filter_order=hyena_filter_order,
  85. dropout=dropout,
  86. filter_dropout=dropout,
  87. )
  88. )
  89. for _ in range(depth)
  90. ]
  91. )
  92. self.ffns = nn.ModuleList(
  93. [
  94. Residual(
  95. nn.Sequential(
  96. nn.RMSNorm(dim_emb),
  97. EinMix(
  98. "b s d -> b s d_f",
  99. weight_shape="d d_f",
  100. bias_shape="d_f",
  101. d=dim_emb,
  102. d_f=dim_ffn,
  103. ),
  104. nn.GELU(),
  105. EinMix(
  106. "b s d_f -> b s d",
  107. weight_shape="d_f d",
  108. bias_shape="d",
  109. d=dim_emb,
  110. d_f=dim_ffn,
  111. ),
  112. nn.Dropout(dropout),
  113. )
  114. )
  115. for _ in range(depth)
  116. ]
  117. )
  118. self.last_rms_norm = nn.RMSNorm(dim_emb)
  119. def forward(self, second_ids: torch.Tensor) -> torch.Tensor:
  120. mask = second_ids != 0
  121. embs = self.embed(second_ids)
  122. for i in range(self.depth):
  123. embs_rms = self.rms_norms[i](embs)
  124. if not self.use_hyena:
  125. embs = embs + self.self_attentions[i](
  126. x=embs_rms,
  127. y=embs_rms,
  128. mask=einsum(mask, mask, "b s1, b s2 -> b s1 s2"),
  129. )
  130. else:
  131. # self_attentions[i] is actually hyena
  132. embs = embs + self.self_attentions[i](embs_rms)
  133. embs = self.ffns[i](embs)
  134. embs = self.last_rms_norm(embs)
  135. return embs, mask
  136. class DNAEncoder(nn.Module):
  137. def __init__(
  138. self,
  139. vocab: int,
  140. max_num_tokens: int,
  141. dim_emb: int,
  142. dim_head: int,
  143. heads: int,
  144. depth: int,
  145. dim_ffn: int,
  146. dropout: float,
  147. use_hyena: bool,
  148. hyena_order: int,
  149. hyena_filter_order: int,
  150. ) -> None:
  151. super().__init__()
  152. self.depth = depth
  153. self.use_hyena = use_hyena
  154. self.embed = nn.Sequential(
  155. nn.Embedding(
  156. vocab,
  157. dim_emb,
  158. ),
  159. nn.Dropout(dropout),
  160. )
  161. self.rms_norms = nn.ModuleList([nn.RMSNorm(dim_emb) for _ in range(depth)])
  162. self.self_attentions = nn.ModuleList(
  163. [
  164. (
  165. MultiHeadAttention(
  166. embed_dim=dim_emb,
  167. num_heads=heads,
  168. num_kv_heads=heads,
  169. head_dim=dim_head,
  170. q_proj=EinMix(
  171. "b s d -> b s nhd",
  172. weight_shape="d nhd",
  173. bias_shape="nhd",
  174. d=dim_emb,
  175. nhd=heads * dim_head,
  176. ),
  177. k_proj=EinMix(
  178. "b s d -> b s nhd",
  179. weight_shape="d nhd",
  180. bias_shape="nhd",
  181. d=dim_emb,
  182. nhd=heads * dim_head,
  183. ),
  184. v_proj=EinMix(
  185. "b s d -> b s nhd",
  186. weight_shape="d nhd",
  187. bias_shape="nhd",
  188. d=dim_emb,
  189. nhd=heads * dim_head,
  190. ),
  191. output_proj=EinMix(
  192. "b s nhd -> b s d",
  193. weight_shape="nhd d",
  194. bias_shape="d",
  195. d=dim_emb,
  196. nhd=heads * dim_head,
  197. ),
  198. # DNA starts with CLS token, so increase max_seq_len by 1
  199. pos_embeddings=RotaryPositionalEmbeddings(
  200. dim=dim_head, max_seq_len=max_num_tokens + 1
  201. ),
  202. max_seq_len=max_num_tokens + 1,
  203. is_causal=False,
  204. attn_dropout=dropout,
  205. )
  206. if not use_hyena
  207. else HyenaOperator(
  208. d_model=dim_emb,
  209. l_max=max_num_tokens + 1,
  210. order=hyena_order,
  211. filter_order=hyena_filter_order,
  212. dropout=dropout,
  213. filter_dropout=dropout,
  214. )
  215. )
  216. for _ in range(depth)
  217. ]
  218. )
  219. self.dna_protein_rms_norms = nn.ModuleList(
  220. [nn.RMSNorm(dim_emb) for _ in range(depth)]
  221. )
  222. self.dna_protein_cross_attentions = nn.ModuleList(
  223. [
  224. MultiHeadAttention(
  225. embed_dim=dim_emb,
  226. num_heads=heads,
  227. num_kv_heads=heads,
  228. head_dim=dim_head,
  229. q_proj=EinMix(
  230. "b s d -> b s nhd",
  231. weight_shape="d nhd",
  232. bias_shape="nhd",
  233. d=dim_emb,
  234. nhd=heads * dim_head,
  235. ),
  236. k_proj=EinMix(
  237. "b s d -> b s nhd",
  238. weight_shape="d nhd",
  239. bias_shape="nhd",
  240. d=dim_emb,
  241. nhd=heads * dim_head,
  242. ),
  243. v_proj=EinMix(
  244. "b s d -> b s nhd",
  245. weight_shape="d nhd",
  246. bias_shape="nhd",
  247. d=dim_emb,
  248. nhd=heads * dim_head,
  249. ),
  250. output_proj=EinMix(
  251. "b s nhd -> b s d",
  252. weight_shape="nhd d",
  253. bias_shape="d",
  254. d=dim_emb,
  255. nhd=heads * dim_head,
  256. ),
  257. # protein bert clip the protein sequence by <start> and <end> tokens, so increase max_seq_len by 2
  258. pos_embeddings=RotaryPositionalEmbeddings(
  259. dim=dim_head, max_seq_len=max_num_tokens + 2
  260. ),
  261. max_seq_len=max_num_tokens + 2,
  262. is_causal=False,
  263. attn_dropout=dropout,
  264. )
  265. for _ in range(depth)
  266. ]
  267. )
  268. self.dna_second_rms_norms = nn.ModuleList(
  269. [nn.RMSNorm(dim_emb) for _ in range(depth)]
  270. )
  271. self.dna_second_cross_attentions = nn.ModuleList(
  272. [
  273. MultiHeadAttention(
  274. embed_dim=dim_emb,
  275. num_heads=heads,
  276. num_kv_heads=heads,
  277. head_dim=dim_head,
  278. q_proj=EinMix(
  279. "b s d -> b s nhd",
  280. weight_shape="d nhd",
  281. bias_shape="nhd",
  282. d=dim_emb,
  283. nhd=heads * dim_head,
  284. ),
  285. k_proj=EinMix(
  286. "b s d -> b s nhd",
  287. weight_shape="d nhd",
  288. bias_shape="nhd",
  289. d=dim_emb,
  290. nhd=heads * dim_head,
  291. ),
  292. v_proj=EinMix(
  293. "b s d -> b s nhd",
  294. weight_shape="d nhd",
  295. bias_shape="nhd",
  296. d=dim_emb,
  297. nhd=heads * dim_head,
  298. ),
  299. output_proj=EinMix(
  300. "b s nhd -> b s d",
  301. weight_shape="nhd d",
  302. bias_shape="d",
  303. d=dim_emb,
  304. nhd=heads * dim_head,
  305. ),
  306. # DNA starts with CLS token, so increase max_seq_len by 1
  307. pos_embeddings=RotaryPositionalEmbeddings(
  308. dim=dim_head, max_seq_len=max_num_tokens + 1
  309. ),
  310. max_seq_len=max_num_tokens + 1,
  311. is_causal=False,
  312. attn_dropout=dropout,
  313. )
  314. for _ in range(depth)
  315. ]
  316. )
  317. self.ffns = nn.ModuleList(
  318. [
  319. Residual(
  320. nn.Sequential(
  321. nn.RMSNorm(dim_emb),
  322. EinMix(
  323. "b s d -> b s d_f",
  324. weight_shape="d d_f",
  325. bias_shape="d_f",
  326. d=dim_emb,
  327. d_f=dim_ffn,
  328. ),
  329. nn.GELU(),
  330. EinMix(
  331. "b s d_f -> b s d",
  332. weight_shape="d_f d",
  333. bias_shape="d",
  334. d=dim_emb,
  335. d_f=dim_ffn,
  336. ),
  337. nn.Dropout(dropout),
  338. )
  339. )
  340. for _ in range(depth)
  341. ]
  342. )
  343. self.last_rms_norm = nn.RMSNorm(dim_emb)
  344. def forward(
  345. self,
  346. dna_ids: torch.Tensor,
  347. protein_embs: torch.Tensor,
  348. second_embs: torch.Tensor,
  349. second_mask: torch.Tensor,
  350. ) -> torch.Tensor:
  351. dna_mask = dna_ids != 0
  352. dna_embs = self.embed(dna_ids)
  353. for i in range(self.depth):
  354. dna_embs_rms = self.rms_norms[i](dna_embs)
  355. if not self.use_hyena:
  356. dna_embs = dna_embs + self.self_attentions[i](
  357. x=dna_embs_rms,
  358. y=dna_embs_rms,
  359. mask=einsum(dna_mask, dna_mask, "b s1, b s2 -> b s1 s2"),
  360. )
  361. else:
  362. # self_attentions[i] is actually hyena
  363. dna_embs = dna_embs + self.self_attentions[i](dna_embs_rms)
  364. # DNA and protein cross-attention
  365. dna_embs_rms = self.dna_protein_rms_norms[i](dna_embs)
  366. dna_embs = dna_embs + self.dna_protein_cross_attentions[i](
  367. x=dna_embs_rms,
  368. y=protein_embs,
  369. mask=repeat(dna_mask, "b s1 -> b s1 s2", s2=protein_embs.shape[1]),
  370. )
  371. # DNA and secondary structure cross-attention
  372. dna_embs_rms = self.dna_second_rms_norms[i](dna_embs)
  373. dna_embs = dna_embs + self.dna_second_cross_attentions[i](
  374. x=dna_embs_rms,
  375. y=second_embs,
  376. mask=einsum(dna_mask, second_mask, "b s1, b s2 -> b s1 s2"),
  377. )
  378. # DNA feedforward network
  379. dna_embs = self.ffns[i](dna_embs)
  380. # only use the embedding for CLS token
  381. dna_embs = dna_embs[:, 0, :]
  382. dna_embs = self.last_rms_norm(dna_embs)
  383. return dna_embs

encoder.py at commit 353997b, under MIT · at the source

Overview

  1. Center for Comparative Biomedicine, Key Laboratory of Systems Biomedicine (MOE), Institute of Systems Biomedicine, Shanghai Jiao Tong University, Shanghai, China
  2. Shanghai Key Laboratory of Gene Editing and Cell-based Immunotherapy for Hematological Diseases, State Key Laboratory of Medical Genomics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China
Journal: PLoS genetics, volume 22, issue 7, article e1012242
Dates: received 12 February 2026; accepted 6 July 2026; published online 16 July 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1371/journal.pgen.1012242 · PMID 42461945 · PMCID PMC13395409 · OpenAlex W7168814918
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: mouse (organism), cellular / molecular (subfield)
Methods: Connectivity, Statistics, Machine learning
MeSH: Cadherins*, CCCTC-Binding Factor*, Cell Cycle Proteins*, Chromosomal Proteins, Non-Histone*, Animals, Binding Sites, Cohesins, Enhancer Elements, Genetic, Gene Expression Regulation, Mice, Promoter Regions, Genetic, Protocadherins, Transcription Factors, Zinc Fingers (* major topic)
Topic: Genomics and Chromatin Dynamics (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: National Key R&amp;D Program of China (2022YFC3400200); National Natural Science Foundation of China (32330016)
Citations: not cited yet (Europe PMC); 82 references in the paper

Abstract

Zinc finger proteins (ZFPs or ZNFs) constitute the largest family of transcription factors in mammals; however, their regulatory mechanism remains largely elusive. Here we propose COP (C2H2-ZFP occupancy predictor), a deep learning-based heuristic screening tool that integrates DNA sequence with protein primary and secondary features to assess ZFP genomic enrichments. Applying COP to the mouse clustered protocadherin (cPcdh) gene locus, we identified dozens of C2H2-ZFPs potentially involved in CTCF-mediated gene regulation with Wiz (widely interspaced zinc finger-containing protein) having the highest number of 12 ZFs. We confirmed Wiz enrichments at all of the CTCF-binding site (CBS) elements across the three Pcdh clusters by Myc-tagging the endogenous Wiz gene. Genetic experiments revealed significant increases of expression levels of the cPcdh genes upon Wiz deletion in both neuronal cells in vitro and in mouse brain in vivo. Finally, integrated ChIP-seq, RNA-seq, and 4C-seq analyses demonstrated that Wiz regulates CTCF/cohesin occupancy and long-range enhancer-promoter contacts in a genomic-distance biased manner. Together, these findings reveal a key role for Wiz in coupling cohesin occupancy to long-range cPcdh regulation and highlight important functions of C2H2-ZFPs in enhancer-promoter interactions.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 14 matches between paragraphs and lines of code.

ljw20180420/COP

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 353997b21fa1b80fc71654e48a3c406c70572702, 7 September 2026
Languages: Python (66), Shell (10), Jupyter (2)
Size: 157 files, 78 scripts
Software Heritage: not archived
Found in: “Data Availability”
Holds: README, license file, environment (environment.yml, requirements_space.txt, requirements_torch_pascal.txt, DeepZF/environment.yml), tests, 2 notebooks
Not found: CITATION.cff, continuous integration, documentation
Tools: pandas (35 files), NumPy (20 files), scikit-learn (17 files), PyTorch (11 files), TensorFlow (8 files), Keras (6 files), SciPy (4 files), Matplotlib (2 files), BEDTools (1 file), LightGBM (1 file), seaborn (1 file), XGBoost (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
80 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 78 scripts, each with its path and the digest of its content;
  • 14 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability

The raw data of RNA-seq, ChIP-seq, and QHR-4C generated in this study were deposited in the Genome Sequence Archive (GSA) at the National Genomics Data Center (NGDC), China National Center for Bioinformation (CNCB), under accession number CRA037522. The source code of the COP has been uploaded to GitHub (https://github.com/ljw20180420/COP). The underlying numerical data for all of graphs and summary statistics have been provided in spreadsheet form as Supporting information.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 5 authors, 14 MeSH terms, 2 funders, 78 references.

Cite

This paper

Li, T., Li, J., Wang, L., Huang, H., & Wu, Q. (2026). Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner. PLoS genetics, 22(7), e1012242. https://doi.org/10.1371/journal.pgen.1012242

BibTeX

@article{li2026wiz,
author = {Li, Tianjie and Li, Jingwei and Wang, Leyang and Huang, Haiyan and Wu, Qiang},
title = {{Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner}},
journal = {PLoS genetics},
year = {2026},
month = jul,
volume = {22},
number = {7},
pages = {e1012242},
publisher = {PLOS},
issn = {1553-7390},
doi = {10.1371/journal.pgen.1012242},
url = {https://doi.org/10.1371/journal.pgen.1012242},
pmid = {42461945},
pmcid = {PMC13395409}
}

RIS

TY - JOUR
AU - Li, Tianjie
AU - Li, Jingwei
AU - Wang, Leyang
AU - Huang, Haiyan
AU - Wu, Qiang
TI - Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner
T2 - PLoS genetics
J2 - PLoS Genet
PY - 2026
DA - 2026/07/16
VL - 22
IS - 7
SP - e1012242
SN - 1553-7390
PB - PLOS
DO - 10.1371/journal.pgen.1012242
UR - https://doi.org/10.1371/journal.pgen.1012242
LA - en
ER -

CSL-JSON

{
"id": "10.1371/journal.pgen.1012242",
"type": "article-journal",
"title": "Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner",
"container-title": "PLoS genetics",
"author": [
{
"family": "Li",
"given": "Tianjie"
},
{
"family": "Li",
"given": "Jingwei"
},
{
"family": "Wang",
"given": "Leyang"
},
{
"family": "Huang",
"given": "Haiyan"
},
{
"family": "Wu",
"given": "Qiang"
}
],
"container-title-short": "PLoS Genet",
"volume": "22",
"issue": "7",
"page": "e1012242",
"DOI": "10.1371/journal.pgen.1012242",
"PMID": "42461945",
"PMCID": "PMC13395409",
"ISSN": "1553-7390",
"publisher": "PLOS",
"URL": "https://doi.org/10.1371/journal.pgen.1012242",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
16
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3389/fgene.2026.1807347 [code]
Epigenetic regulators are preferentially coordinated with protocadherin gene expression across the human brain: a genome-wide co-expression analysis.
Journal: Frontiers in genetics
In common: pandas, SciPy, Matplotlib, 1 other tool, cellular / molecular, 9 references
[2] doi:10.3389/fpsyg.2026.1774068 [code]
Analysis of cognitive mechanisms in phoneme perception and pronunciation errors among Korean language learners.
Journal: Frontiers in psychology
In common: LightGBM, XGBoost, Keras, 8 other tools
[3] doi:10.1038/s41592-026-03057-2 [code]
CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species.
Journal: Nature methods
In common: BEDTools, Keras, TensorFlow, 7 other tools, mouse, 1 reference
[4] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: BEDTools, Keras, TensorFlow, 6 other tools, mouse, cellular / molecular, 1 reference
[5] doi:10.7554/elife.110588 [code]
Opening the black box toward a modular approach to spike sorting.
Journal: eLife
In common: LightGBM, XGBoost, TensorFlow, 7 other tools, mouse
[6] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: XGBoost, Keras, TensorFlow, 7 other tools, mouse, cellular / molecular
[7] doi:10.1038/s41598-026-48613-0 [code]
An snRNA-seq aging clock for the fruit fly head sheds light on sex-biased aging.
Journal: Scientific reports
In common: XGBoost, Keras, TensorFlow, 7 other tools, cellular / molecular
[8] doi:10.1371/journal.pcbi.1014615 [code]
Toward reliable machine learning models for neural circuit inference: A diagnostic study of CNNs on spike trains.
Journal: PLoS computational biology
In common: XGBoost, Keras, TensorFlow, 7 other tools
[9] doi:10.1093/bib/bbag118 [code]
Drug screening for α-synuclein aggregation inhibitors via multimodal graph neural network.
Journal: Briefings in bioinformatics
In common: LightGBM, XGBoost, PyTorch, 6 other tools, cellular / molecular
[10] doi:10.1126/sciadv.aed3650 [code]
Truthful visualizations for mass spectrometry imaging enable high-spatial-resolution interactive &lt;i&gt;m/z&lt;/i&gt; mapping and exploration.
Journal: Science advances
In common: XGBoost, Keras, TensorFlow, 6 other tools, mouse

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.