OSCR

RNA-binding protein family diversification correlates with neural complexity across metazoan evolution.

Code ↔ Paper

23 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 23 matches
  1. [1] § Results › Non-canonical RBPs and RBP-transcription factor overlap ↔ scripts/supplementary_analyses.py, lines 256–307 · score 0.94 · zf C2H2 multi, RBP TF overlap, 2.6–8.3 %, excluding zebrafish, AnimalTFDB, inflated
  2. [2] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 314–400 · score 0.87 · RAP domain, OAS domain, RNase, PARP domain, YTH domain, domain families
  3. [3] § STAR★Methods › Method details › Non-canonical RBP analysis ↔ scripts/supplementary_analyses.py, lines 1–50 · score 0.85 · RBP TF overlap, RBPWorld, AnimalTFDB, gene symbols, Canonical RBPs, EuRBPDB
  4. [4] § STAR★Methods › Method details › IDR analysis ↔ scripts/idr_analysis.py, lines 1–53 · score 0.83 · longest contiguous IDR, disorder predictions, disordered residues, MobiDB, database, RBP
  5. [5] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 314–400 · score 0.78 · RNase, PARP domain, YTH domain, domain expansion, absent, mammalian
  6. [6] § Results › RBP family diversity strongly correlates with organismal complexity ↔ scripts/statistical_validation.py, lines 1–50 · score 0.75 · bootstrap resampling, Pfam clan, control protein, validation, CI, Cohen
  7. [7] § STAR★Methods › Method details › Control protein classification ↔ scripts/analyze_correlations.py, lines 1–38 · score 0.74 · Pfam domain assignments, tf_family, AnimalTFDB, GPCR, Kinase, database
  8. [8] § Results › Intrinsically disordered regions expand with neural complexity ↔ scripts/fetch_idr_mobidb_v2.py, lines 44–115 · score 0.71 · disordered residues, longest IDR, disordered regions, MobiDB, consensus, predictions
  9. [9] § Results › Intrinsically disordered regions expand with neural complexity ↔ scripts/idr_analysis.py, lines 1–53 · score 0.70 · disordered residues, longest IDR, IDR content, MobiDB, aa, predictions
  10. [10] § STAR★Methods › Method details › RBP identification ↔ scripts/analyze_correlations.py, lines 1–38 · score 0.70 · AnimalTFDB, RBP gene, control protein, EuRBPDB, GPCR, Kinase
  11. [11] § Results › RBP family diversity strongly correlates with organismal complexity ↔ scripts/pgls_analysis.py, lines 1–51 · score 0.70 · phylogenetic signal, phylogenetic correction, correlation remained, PGLS, Pagel, optimal
  12. [12] § STAR★Methods › Quantification and statistical analysis ↔ scripts/pgls_analysis.py, lines 1–51 · score 0.69 · maximum likelihood, Spearman correlations, PGLS, Pagel, scipy, Phylogenetic
  13. [13] § STAR★Methods › Method details › IDR analysis ↔ scripts/fetch_idr_mobidb_v2.py, lines 44–115 · score 0.69 · disordered residues, MobiDB, disordered regions, IDR, consensus, longest
  14. [14] § STAR★Methods › Method details › RBP identification ↔ scripts/fetch_pfam_from_uniprot.py, lines 1–44 · score 0.69 · RBP.txt, UniProt, EuRBPDB, melanogaster, musculus, rerio
  15. [15] § STAR★Methods › Method details › RBP family classification ↔ scripts/fetch_pfam_from_uniprot.py, lines 1–44 · score 0.69 · Gene symbols, UniProt, EuRBPDB, Pfam domain, mapped, accessions
  16. [16] § STAR★Methods › Method details › LLPS analysis ↔ scripts/aggregate_llphyscore.py, lines 81–196 · score 0.67 · raw scores, LLPhyScore, aggregation, isoforms, LLPS, tropicalis
  17. [17] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 1–50 · score 0.65 · expansion ratios, canonical RBP, Pfam domains, emergence, fold, enrichment
  18. [18] § Results › LLPS propensity is evolutionarily conserved ↔ scripts/extract_phasepred_v2.py, lines 1–14 · score 0.61 · PScore, catGRANULE, PhaSePred
  19. [19] § STAR★Methods › Method details › Orthogroup analysis ↔ scripts_of/util.py, lines 327–387 · score 0.61 · MCL clustering, OrthoFinder, DIAMOND, alignment, Orthogroups, sequences
  20. [20] § STAR★Methods › Quantification and statistical analysis ↔ scripts/statistical_validation.py, lines 1–50 · score 0.57 · Bootstrap resampling, validation, Cohen, scipy, numpy, correlations
  21. [21] § Results › Non-canonical RBPs and RBP-transcription factor overlap ↔ scripts/supplementary_analyses.py, lines 171–249 · score 0.53 · RBPs lacking, RBPWorld, canonical RBP, RBDs, correlated, species
  22. [22] § STAR★Methods › Method details › RBP family classification ↔ scripts/statistical_validation.py, lines 324–393 · score 0.53 · Pfam clan, Robustness, assignments, mapped, EuRBPDB, matching
  23. [23] § STAR★Methods › Method details › LLPS analysis ↔ scripts/extract_llphyscore.py, lines 16–58 · score 0.53 · raw scores, LLPhyScore, protein, species

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 469 lines · 19 KB · no license · 6 matches

  1. #!/usr/bin/env python3
  2. """
  3. Supplementary Analyses Script
  4. ===============================
  5. Reproduces supplementary analyses in Yasuda et al. (iScience, 2026)
  6. Analyses performed:
  7. 1. Non-canonical RBP analysis
  8. - Count non-canonical RBPs per species (RBPWorld)
  9. - Spearman correlation vs neurons
  10. - Combined canonical + non-canonical correlation
  11. 2. RBP–TF overlap
  12. - Intersect EuRBPDB gene symbols with AnimalTFDB gene symbols
  13. - Report overlap % per species
  14. 3. Vertebrate domain expansion
  15. - Compute vertebrate/invertebrate gene count ratio per Pfam domain
  16. - Identify domains with ≥3-fold enrichment or vertebrate-specific emergence
  17. Usage:
  18. python3 supplementary_analyses.py <data_dir> <output_dir>
  19. <data_dir> : directory containing input files (default: ../data/processed)
  20. <output_dir> : directory for output files (default: ../results)
  21. Input files required:
  22. new_rbp_db.csv - RBP database (EuRBPDB, canonical + non-canonical flag)
  23. new_tf_db.csv - TF database (AnimalTFDB 4.0)
  24. eurbpdb_pfam_mapping.tsv - gene-level Pfam assignments
  25. Output files:
  26. noncanonical_correlations.csv - non-canonical RBP counts and correlations
  27. rbp_tf_overlap.csv - RBP-TF overlap per species
  28. domain_expansion.csv - vertebrate expansion ratios per Pfam domain
  29. Author: Kyota Yasuda
  30. Date: March 2026
  31. """
  32. import sys
  33. import os
  34. import csv
  35. import math
  36. from collections import defaultdict
  37. try:
  38. from scipy import stats
  39. HAS_SCIPY = True
  40. except ImportError:
  41. HAS_SCIPY = False
  42. print("[WARN] scipy not found – using built-in Spearman implementation")
  43. # ── Species metadata ─────────────────────────────────────────────────────────
  44. SPECIES_ORDER = [
  45. "Caenorhabditis_elegans",
  46. "Drosophila_melanogaster",
  47. "Danio_rerio",
  48. "Xenopus_tropicalis",
  49. "Mus_musculus",
  50. "Homo_sapiens",
  51. ]
  52. # RBPWorld available for 5 species (Xenopus excluded)
  53. SPECIES_NONCANONICAL = [
  54. "Caenorhabditis_elegans",
  55. "Drosophila_melanogaster",
  56. "Danio_rerio",
  57. "Mus_musculus",
  58. "Homo_sapiens",
  59. ]
  60. NEURON_COUNT = {
  61. "Caenorhabditis_elegans": 302,
  62. "Drosophila_melanogaster": 200_000,
  63. "Danio_rerio": 10_000_000,
  64. "Xenopus_tropicalis": 16_000_000,
  65. "Mus_musculus": 71_000_000,
  66. "Homo_sapiens": 86_000_000_000,
  67. }
  68. # Vertebrate vs invertebrate classification
  69. INVERTEBRATES = {"Caenorhabditis_elegans", "Drosophila_melanogaster"}
  70. VERTEBRATES = {"Danio_rerio", "Xenopus_tropicalis", "Mus_musculus", "Homo_sapiens"}
  71. # ════════════════════════════════════════════════════════════════════════════
  72. # Spearman utilities
  73. # ════════════════════════════════════════════════════════════════════════════
  74. def _rank(lst):
  75. sorted_idx = sorted(range(len(lst)), key=lambda i: lst[i])
  76. ranks = [0.0] * len(lst)
  77. i = 0
  78. while i < len(lst):
  79. j = i
  80. while j < len(lst) - 1 and lst[sorted_idx[j]] == lst[sorted_idx[j + 1]]:
  81. j += 1
  82. avg_rank = (i + j) / 2.0 + 1
  83. for k in range(i, j + 1):
  84. ranks[sorted_idx[k]] = avg_rank
  85. i = j + 1
  86. return ranks
  87. def _norm_cdf(x):
  88. return (1.0 + math.erf(x / math.sqrt(2))) / 2.0
  89. def spearman_builtin(x, y):
  90. n = len(x)
  91. rx, ry = _rank(x), _rank(y)
  92. mx = sum(rx) / n
  93. my = sum(ry) / n
  94. num = sum((rx[i] - mx) * (ry[i] - my) for i in range(n))
  95. dx = math.sqrt(sum((rx[i] - mx) ** 2 for i in range(n)))
  96. dy = math.sqrt(sum((ry[i] - my) ** 2 for i in range(n)))
  97. if dx == 0 or dy == 0:
  98. return 0.0, 1.0
  99. rho = num / (dx * dy)
  100. if abs(rho) >= 1.0:
  101. return rho, 0.0
  102. t_stat = rho * math.sqrt(n - 2) / math.sqrt(1 - rho ** 2)
  103. p = 2 * (1 - _norm_cdf(abs(t_stat)))
  104. return rho, p
  105. def spearman(x, y):
  106. if HAS_SCIPY:
  107. r = stats.spearmanr(x, y)
  108. return float(r.statistic), float(r.pvalue)
  109. return spearman_builtin(x, y)
  110. def sig_label(p):
  111. if p < 0.001: return "***"
  112. if p < 0.01: return "**"
  113. if p < 0.05: return "*"
  114. return "ns"
  115. # ════════════════════════════════════════════════════════════════════════════
  116. # File loaders
  117. # ════════════════════════════════════════════════════════════════════════════
  118. def load_csv(path, delimiter=","):
  119. if not os.path.isfile(path):
  120. raise FileNotFoundError(f"Required file not found: {path}")
  121. with open(path, encoding="utf-8") as fh:
  122. rows = list(csv.DictReader(fh, delimiter=delimiter))
  123. print(f" Loaded {len(rows):,} rows ← {os.path.basename(path)}")
  124. return rows
  125. def load_tsv(path):
  126. return load_csv(path, delimiter="\t")
  127. def write_csv(path, rows, fieldnames):
  128. with open(path, "w", newline="", encoding="utf-8") as fh:
  129. w = csv.DictWriter(fh, fieldnames=fieldnames)
  130. w.writeheader()
  131. w.writerows(rows)
  132. print(f" Saved → {path}")
  133. # ════════════════════════════════════════════════════════════════════════════
  134. # 1. Non-canonical RBP analysis
  135. # ════════════════════════════════════════════════════════════════════════════
  136. def run_noncanonical(rbp_rows):
  137. """
  138. Uses is_canonical column in new_rbp_db.csv to separate canonical / non-canonical.
  139. Non-canonical = experimentally identified RBPs lacking canonical RBDs (from RBPWorld).
  140. Expected result (from manuscript):
  141. Non-canonical count vs neurons: rho=0.900, p=0.037 (5 species, excl. Xenopus)
  142. Combined (canonical + non-canonical) vs neurons: rho=0.975, p=0.005
  143. """
  144. # Count by species and canonical status
  145. canonical_counts = defaultdict(int)
  146. noncanonical_counts = defaultdict(int)
  147. for r in rbp_rows:
  148. sp = r.get("species", "").strip()
  149. can = r.get("is_canonical", "").strip().lower()
  150. if sp not in SPECIES_ORDER:
  151. continue
  152. if can in ("true", "1", "yes"):
  153. canonical_counts[sp] += 1
  154. elif can in ("false", "0", "no"):
  155. noncanonical_counts[sp] += 1
  156. # Check non-canonical data exists
  157. total_noncan = sum(noncanonical_counts.values())
  158. if total_noncan == 0:
  159. print(" [WARN] No non-canonical RBPs found in new_rbp_db.csv.")
  160. print(" Ensure is_canonical=False entries from RBPWorld are present.")
  161. # Correlation: non-canonical vs neurons (5 species, excl. Xenopus)
  162. sp5 = SPECIES_NONCANONICAL
  163. neurons5 = [NEURON_COUNT[sp] for sp in sp5]
  164. noncan5 = [noncanonical_counts[sp] for sp in sp5]
  165. rho_nc, p_nc = spearman(neurons5, noncan5)
  166. # Correlation: combined vs neurons (5 species)
  167. combined5 = [canonical_counts[sp] + noncanonical_counts[sp] for sp in sp5]
  168. rho_cb, p_cb = spearman(neurons5, combined5)
  169. print(f"\n Non-canonical RBP counts (5 species):")
  170. for sp in sp5:
  171. print(f" {sp:35s} canonical={canonical_counts[sp]:5d} "
  172. f"non-canonical={noncanonical_counts[sp]:5d}")
  173. print(f"\n Non-canonical vs neurons: rho={rho_nc:.3f} p={p_nc:.4f} "
  174. f"{sig_label(p_nc)} (expected: rho=0.900, p=0.037)")
  175. print(f" Combined vs neurons: rho={rho_cb:.3f} p={p_cb:.4f} "
  176. f"{sig_label(p_cb)} (expected: rho=0.975, p=0.005)")
  177. results = []
  178. for sp in SPECIES_ORDER:
  179. results.append({
  180. "species": sp,
  181. "neurons": NEURON_COUNT[sp],
  182. "canonical_count": canonical_counts[sp],
  183. "noncanonical_count": noncanonical_counts[sp],
  184. "combined_count": canonical_counts[sp] + noncanonical_counts[sp],
  185. "xenopus_excluded": sp not in SPECIES_NONCANONICAL,
  186. })
  187. # Append correlation summary
  188. results.append({
  189. "species": "CORRELATION: non-canonical vs neurons (n=5)",
  190. "neurons": "",
  191. "canonical_count": "",
  192. "noncanonical_count": round(rho_nc, 3),
  193. "combined_count": round(p_nc, 4),
  194. "xenopus_excluded": sig_label(p_nc),
  195. })
  196. results.append({
  197. "species": "CORRELATION: combined vs neurons (n=5)",
  198. "neurons": "",
  199. "canonical_count": "",
  200. "noncanonical_count": round(rho_cb, 3),
  201. "combined_count": round(p_cb, 4),
  202. "xenopus_excluded": sig_label(p_cb),
  203. })
  204. return results
  205. # ════════════════════════════════════════════════════════════════════════════
  206. # 2. RBP–TF overlap
  207. # ════════════════════════════════════════════════════════════════════════════
  208. def run_rbp_tf_overlap(rbp_rows, tf_rows):
  209. """
  210. Intersect EuRBPDB gene symbols with AnimalTFDB gene symbols per species.
  211. Expected result (from manuscript):
  212. Overlap ranges 2.6–8.3% (excluding zebrafish where zf-C2H2 inflates to 16.0%)
  213. Human: 194/2961 = 6.6% overlap (86 zf-C2H2)
  214. """
  215. # Build gene symbol sets
  216. rbp_genes = defaultdict(set)
  217. tf_genes = defaultdict(set)
  218. for r in rbp_rows:
  219. sp = r.get("species", "").strip()
  220. sym = r.get("gene_symbol", "").strip().upper()
  221. if sp and sym:
  222. rbp_genes[sp].add(sym)
  223. for r in tf_rows:
  224. sp = r.get("species", "").strip()
  225. sym = r.get("gene_symbol", "").strip().upper()
  226. if sp and sym:
  227. tf_genes[sp].add(sym)
  228. results = []
  229. print(f"\n {'Species':35s} {'RBPs':>6} {'TFs':>6} {'Overlap':>8} {'%':>7}")
  230. for sp in SPECIES_ORDER:
  231. rbp_set = rbp_genes[sp]
  232. tf_set = tf_genes[sp]
  233. overlap = rbp_set & tf_set
  234. n_rbp = len(rbp_set)
  235. n_tf = len(tf_set)
  236. n_ov = len(overlap)
  237. pct = 100.0 * n_ov / n_rbp if n_rbp > 0 else 0.0
  238. note = ""
  239. if sp == "Danio_rerio":
  240. note = "inflated by zf-C2H2 multi-annotation"
  241. print(f" {sp:35s} {n_rbp:>6} {n_tf:>6} {n_ov:>8} {pct:>6.1f}%"
  242. + (f" [{note}]" if note else ""))
  243. results.append({
  244. "species": sp,
  245. "n_rbps": n_rbp,
  246. "n_tfs": n_tf,
  247. "n_overlap": n_ov,
  248. "overlap_pct": round(pct, 1),
  249. "note": note,
  250. })
  251. return results
  252. # ════════════════════════════════════════════════════════════════════════════
  253. # 3. Vertebrate domain expansion
  254. # ════════════════════════════════════════════════════════════════════════════
  255. def run_domain_expansion(pfam_rows):
  256. """
  257. Compute mean gene count per Pfam domain family in invertebrates vs vertebrates.
  258. Identify domains with ≥3-fold enrichment or vertebrate-specific emergence.
  259. Expected (from manuscript Figure 3A):
  260. PARP domain: ~8.9-fold
  261. YTH domain: ~8.0-fold
  262. RAP domain: ~11.0-fold
  263. RNase A domain: vertebrate-specific (mammalian)
  264. OAS domain: vertebrate-specific
  265. """
  266. # Count genes per (family, species)
  267. family_species_count = defaultdict(lambda: defaultdict(int))
  268. for row in pfam_rows:
  269. sp = row.get("species", "").strip()
  270. family = row.get("eurbpdb_family", "").strip()
  271. if sp not in SPECIES_ORDER:
  272. continue
  273. if not family or family == "Non-canonical":
  274. continue
  275. family_species_count[family][sp] += 1
  276. results = []
  277. for family, sp_counts in family_species_count.items():
  278. invert_counts = [sp_counts.get(sp, 0) for sp in INVERTEBRATES]
  279. vert_counts = [sp_counts.get(sp, 0) for sp in VERTEBRATES]
  280. mean_invert = sum(invert_counts) / len(invert_counts)
  281. mean_vert = sum(vert_counts) / len(vert_counts)
  282. total_genes = sum(sp_counts.values())
  283. n_species = len(sp_counts)
  284. # Expansion ratio
  285. if mean_invert == 0 and mean_vert > 0:
  286. ratio = float("inf")
  287. expansion_type = "vertebrate-specific"
  288. elif mean_invert == 0:
  289. ratio = 1.0
  290. expansion_type = "absent"
  291. else:
  292. ratio = mean_vert / mean_invert
  293. if ratio >= 3.0:
  294. expansion_type = "vertebrate-expanded (>=3x)"
  295. elif ratio >= 2.0:
  296. expansion_type = "vertebrate-enriched (2-3x)"
  297. else:
  298. expansion_type = "conserved"
  299. results.append({
  300. "pfam_family": family,
  301. "n_species": n_species,
  302. "total_genes": total_genes,
  303. "mean_invertebrate": round(mean_invert, 2),
  304. "mean_vertebrate": round(mean_vert, 2),
  305. "vert_invert_ratio": round(ratio, 2) if ratio != float("inf") else "inf",
  306. "expansion_type": expansion_type,
  307. "worm": sp_counts.get("Caenorhabditis_elegans", 0),
  308. "fly": sp_counts.get("Drosophila_melanogaster", 0),
  309. "zfish": sp_counts.get("Danio_rerio", 0),
  310. "frog": sp_counts.get("Xenopus_tropicalis", 0),
  311. "mouse": sp_counts.get("Mus_musculus", 0),
  312. "human": sp_counts.get("Homo_sapiens", 0),
  313. })
  314. # Sort by ratio descending (inf first)
  315. def sort_key(r):
  316. v = r["vert_invert_ratio"]
  317. return -float("inf") if v == "inf" else -float(v)
  318. results.sort(key=sort_key)
  319. # Print top expanded
  320. print(f"\n Top vertebrate-expanded Pfam domain families (≥3-fold or vertebrate-specific):")
  321. print(f" {'Family':30s} {'ratio':>8} {'type':35s} "
  322. f"{'worm':>5} {'fly':>5} {'zfish':>6} {'frog':>5} {'mouse':>6} {'human':>6}")
  323. for r in results:
  324. if r["expansion_type"] in ("vertebrate-specific", "vertebrate-expanded (>=3x)"):
  325. ratio_str = str(r["vert_invert_ratio"])
  326. print(f" {r['pfam_family']:30s} {ratio_str:>8} "
  327. f"{r['expansion_type']:35s} "
  328. f"{r['worm']:>5} {r['fly']:>5} {r['zfish']:>6} "
  329. f"{r['frog']:>5} {r['mouse']:>6} {r['human']:>6}")
  330. return results
  331. # ════════════════════════════════════════════════════════════════════════════
  332. # Main
  333. # ════════════════════════════════════════════════════════════════════════════
  334. def main():
  335. data_dir = sys.argv[1] if len(sys.argv) > 1 else os.path.join(
  336. os.path.dirname(__file__), "..", "data", "processed")
  337. output_dir = sys.argv[2] if len(sys.argv) > 2 else os.path.join(
  338. os.path.dirname(__file__), "..", "results")
  339. data_dir = os.path.abspath(data_dir)
  340. output_dir = os.path.abspath(output_dir)
  341. os.makedirs(output_dir, exist_ok=True)
  342. print("=" * 60)
  343. print("Supplementary Analyses")
  344. print(f" data_dir : {data_dir}")
  345. print(f" output_dir : {output_dir}")
  346. print("=" * 60)
  347. # ── Load data ─────────────────────────────────────────────────────────
  348. print("\n[0] Loading data …")
  349. rbp_rows = load_csv(os.path.join(data_dir, "new_rbp_db.csv"))
  350. tf_rows = load_csv(os.path.join(data_dir, "new_tf_db.csv"))
  351. pfam_rows = load_tsv(os.path.join(data_dir, "eurbpdb_pfam_mapping.tsv"))
  352. # ── 1. Non-canonical ──────────────────────────────────────────────────
  353. print("\n[1] Non-canonical RBP analysis …")
  354. nc_results = run_noncanonical(rbp_rows)
  355. write_csv(
  356. os.path.join(output_dir, "noncanonical_correlations.csv"),
  357. nc_results,
  358. ["species", "neurons", "canonical_count",
  359. "noncanonical_count", "combined_count", "xenopus_excluded"],
  360. )
  361. # ── 2. RBP-TF overlap ────────────────────────────────────────────────
  362. print("\n[2] RBP–TF overlap analysis …")
  363. overlap_results = run_rbp_tf_overlap(rbp_rows, tf_rows)
  364. write_csv(
  365. os.path.join(output_dir, "rbp_tf_overlap.csv"),
  366. overlap_results,
  367. ["species", "n_rbps", "n_tfs", "n_overlap", "overlap_pct", "note"],
  368. )
  369. # ── 3. Domain expansion ───────────────────────────────────────────────
  370. print("\n[3] Vertebrate domain expansion analysis …")
  371. expansion_results = run_domain_expansion(pfam_rows)
  372. write_csv(
  373. os.path.join(output_dir, "domain_expansion.csv"),
  374. expansion_results,
  375. ["pfam_family", "n_species", "total_genes",
  376. "mean_invertebrate", "mean_vertebrate", "vert_invert_ratio",
  377. "expansion_type", "worm", "fly", "zfish", "frog", "mouse", "human"],
  378. )
  379. print("\nDone.")
  380. print("\nExpected results (from manuscript):")
  381. print(" Non-canonical vs neurons (n=5): rho=0.900, p=0.037 *")
  382. print(" Combined vs neurons (n=5): rho=0.975, p=0.005 **")
  383. print(" RBP-TF overlap: 2.6–8.3% (excl. zebrafish 16.0%)")
  384. print(" Human overlap: 194/2961 = 6.6%")
  385. print(" Top expanded domains: RAP (11x), PARP (8.9x), YTH (8x)")
  386. if __name__ == "__main__":
  387. main()

supplementary_analyses.py at commit 66dafdc, no license · at the source

Overview

Authors: Kyota Yasuda1,2,3
ORCID iDs: Kyota Yasuda
  1. Graduate School of Integrated Sciences for Life, Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8526, Japan
  2. International Institute for Sustainability with Knotted Chiral Meta Matter (SKCM2), Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8526, Japan
  3. Research Center for the Mathematics on Chromatin Live Dynamics (RcMcD), Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8531, Japan
Institutions: Hiroshima University (Japan)
Journal: iScience, volume 29, issue 5, article 115766
Dates: received 22 December 2025; accepted 13 April 2026; published online 17 April 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1016/j.isci.2026.115766 · PMID 42111166 · PMCID PMC13157009 · OpenAlex W7154729910
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: cellular / molecular (subfield)
Methods: Connectivity, Machine learning, Statistics
Keywords: Evolutionary biology, Systems biology
Topic: RNA Research and Splicing (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Funding: Japan Society for the Promotion of Science; Japan Society for the Promotion of Science London (23K05147)
Citations: cited by 1 paper (Europe PMC); 56 references in the paper
Research resources: SciPy RRID:SCR_008058, Python 3.9 RRID:SCR_008394, Matplotlib RRID:SCR_008624, NumPy RRID:SCR_008633, Seaborn RRID:SCR_018132, Pandas RRID:SCR_018214

Abstract

RNA-binding proteins (RBPs) control post-transcriptional gene expression with critical roles in nervous system function. Here, we ask whether RBP family diversity tracks with neural complexity across animal evolution. Comparing six species ranging from a simple worm to humans, we find that the number of distinct RBP domain families increases progressively with neuron number—a relationship specific to RBPs and not seen for kinases or G-protein-coupled receptors. Transcription factors show a different pattern, reaching a ceiling in vertebrates while RBP families continue expanding. We further show that the disordered structural regions of RBPs grow longer in more complex organisms, whereas liquid-liquid phase separation propensity remains broadly conserved. Finally, the lengthening of messenger RNA regulatory tails parallels RBP diversification, suggesting co-evolution of post-transcriptional regulatory capacity. These findings establish RBP family diversification as a distinctive molecular signature of animal neural complexity.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repositories

Its files are read in the Code ↔ Paper reader above, with 23 matches between paragraphs and lines of code.

Zenodo 18002633

License: CC-BY-4.0
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Code”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: SciPy (6 files), pandas (4 files), NumPy (3 files)
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
  • 29 September 2026: the link answers (HTTP 200)
12 files

Kyotay-12/rbp-evolution-analysis

License: none: the authors keep all their rights
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 66dafdcaad56c2c9bb26baf69f007002759fa743, 9 March 2026
Languages: Python (11)
Size: 41 files, 11 scripts
Software Heritage: not archived
Found in: “Code”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: SciPy (6 files), pandas (4 files), NumPy (3 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
12 files

davidemms/OrthoFinder

License: GPL-3.0
State: the link answers, verified on 29 September 2026
Evidence: files inventoried
Commit: 43f9bc7273a33d2e7a4e8d1c55b5117ed7f08a42, 15 July 2025
Languages: Python (42), Perl (2)
Size: 4,207 files, 44 scripts
Software Heritage: not archived
Found in: the resources table
Holds: README, license file, environment (requirements.txt, setup.py)
Not found: CITATION.cff, tests, continuous integration, documentation
Tools: NumPy (17 files), SciPy (7 files), Biopython (2 files), scikit-learn (2 files)
Availability: 1 check, the latest on 29 September 2026: the link answers
  • 29 September 2026: the link answers
46 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 66 scripts, each with its path and the digest of its content;
  • 23 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Data and code availability

Data: All data reported in this paper are derived from publicly available databases (EuRBPDB, RBPWorld, AnimalTFDB 4.0, MobiDB, Ensembl BioMart, UniProt, PhaSePred, LLPhyScore). Processed datasets supporting this study are publicly available at Zenodo (DOI: https://doi.org/10.5281/zenodo.18002633). Data reported in this paper will be shared by the lead contact upon request.

Code: All analysis scripts used in this study are publicly available at GitHub (https://github.com/Kyotay-12/rbp-evolution-analysis) and archived at Zenodo (DOI: https://doi.org/10.5281/zenodo.18002633).

Other items: Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 29 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 1 author, 2 keywords, 2 funders, 53 references, 6 RRIDs.

Cite

This paper

Yasuda, K. (2026). RNA-binding protein family diversification correlates with neural complexity across metazoan evolution. iScience, 29(5), 115766. https://doi.org/10.1016/j.isci.2026.115766

BibTeX

@article{yasuda2026rna,
author = {Yasuda, Kyota},
title = {{RNA-binding protein family diversification correlates with neural complexity across metazoan evolution}},
journal = {iScience},
year = {2026},
month = apr,
volume = {29},
number = {5},
pages = {115766},
publisher = {Elsevier},
issn = {2589-0042},
doi = {10.1016/j.isci.2026.115766},
url = {https://doi.org/10.1016/j.isci.2026.115766},
pmid = {42111166},
pmcid = {PMC13157009}
}

RIS

TY - JOUR
AU - Yasuda, Kyota
TI - RNA-binding protein family diversification correlates with neural complexity across metazoan evolution
T2 - iScience
J2 - iScience
PY - 2026
DA - 2026/04/17
VL - 29
IS - 5
SP - 115766
SN - 2589-0042
PB - Elsevier
DO - 10.1016/j.isci.2026.115766
UR - https://doi.org/10.1016/j.isci.2026.115766
LA - en
ER -

CSL-JSON

{
"id": "10.1016/j.isci.2026.115766",
"type": "article-journal",
"title": "RNA-binding protein family diversification correlates with neural complexity across metazoan evolution",
"container-title": "iScience",
"author": [
{
"family": "Yasuda",
"given": "Kyota"
}
],
"container-title-short": "iScience",
"volume": "29",
"issue": "5",
"page": "115766",
"DOI": "10.1016/j.isci.2026.115766",
"PMID": "42111166",
"PMCID": "PMC13157009",
"ISSN": "2589-0042",
"publisher": "Elsevier",
"URL": "https://doi.org/10.1016/j.isci.2026.115766",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
17
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1038/s41586-026-10629-x [code]
Whole-genome duplication shaped cell-type evolution in the vertebrate brain.
Journal: Nature
In common: scikit-learn, pandas, SciPy, 1 other tool, cellular / molecular, 4 references
[2] doi:10.1523/eneuro.0362-25.2026 [code]
Similarities between &lt;i&gt;Ciona&lt;/i&gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/V Subdivision within a Hindbrain Precursor?
Journal: eNeuro
In common: Biopython, scikit-learn, pandas, 2 other tools, 2 references
[3] doi:10.3389/fgene.2026.1742595 [code]
An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis.
Journal: Frontiers in genetics
In common: scikit-learn, pandas, NumPy, 4 references
[4] doi:10.1038/s41467-026-75700-7 [code]
Gene regulatory innovations from transposable elements in primate cerebellum development.
Journal: Nature communications
In common: Biopython, scikit-learn, pandas, 2 other tools, cellular / molecular, 1 reference
[5] doi:10.1002/advs.202523984 [code]
INB&lt;sup&gt;3&lt;/sup&gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: Biopython, scikit-learn, pandas, 2 other tools, 1 reference
[6] doi:10.1038/s41467-026-76045-x [code]
A manufacturability-informed topology framework for AI-guided design of fibrous network materials.
Journal: Nature communications
In common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
[7] doi:10.1093/narmme/ugag028 [code]
MCVAE-based multi-omic anomaly detection in Fragile X Syndrome.
Journal: NAR molecular medicine
In common: scikit-learn, pandas, SciPy, 1 other tool, cellular / molecular, 2 references
[8] doi:10.1371/journal.pbio.3003856 [code]
Aging and metabolism contribute separately to brain-body health.
Journal: PLoS biology
In common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
[9] doi:10.1038/s41597-026-07077-7 [code]
Everyday Activity Science and Engineering Table Setting Dataset.
Journal: Scientific data
In common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
[10] doi:10.1016/j.crmeth.2026.101421 [code]
EthoPy provides an accessible platform for reproducible behavioral neuroscience.
Journal: Cell reports methods
In common: scikit-learn, pandas, SciPy, 1 other tool, 3 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.