RNA-binding protein family diversification correlates with neural complexity across metazoan evolution.
The 23 matches
- [1] § Results › Non-canonical RBPs and RBP-transcription factor overlap ↔ scripts/supplementary_analyses.py, lines 256–307 · score 0.94 · zf C2H2 multi, RBP TF overlap, 2.6–8.3 %, excluding zebrafish, AnimalTFDB, inflated
- [2] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 314–400 · score 0.87 · RAP domain, OAS domain, RNase, PARP domain, YTH domain, domain families
- [3] § STAR★Methods › Method details › Non-canonical RBP analysis ↔ scripts/supplementary_analyses.py, lines 1–50 · score 0.85 · RBP TF overlap, RBPWorld, AnimalTFDB, gene symbols, Canonical RBPs, EuRBPDB
- [4] § STAR★Methods › Method details › IDR analysis ↔ scripts/idr_analysis.py, lines 1–53 · score 0.83 · longest contiguous IDR, disorder predictions, disordered residues, MobiDB, database, RBP
- [5] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 314–400 · score 0.78 · RNase, PARP domain, YTH domain, domain expansion, absent, mammalian
- [6] § Results › RBP family diversity strongly correlates with organismal complexity ↔ scripts/statistical_validation.py, lines 1–50 · score 0.75 · bootstrap resampling, Pfam clan, control protein, validation, CI, Cohen
- [7] § STAR★Methods › Method details › Control protein classification ↔ scripts/analyze_correlations.py, lines 1–38 · score 0.74 · Pfam domain assignments, tf_family, AnimalTFDB, GPCR, Kinase, database
- [8] § Results › Intrinsically disordered regions expand with neural complexity ↔ scripts/fetch_idr_mobidb_v2.py, lines 44–115 · score 0.71 · disordered residues, longest IDR, disordered regions, MobiDB, consensus, predictions
- [9] § Results › Intrinsically disordered regions expand with neural complexity ↔ scripts/idr_analysis.py, lines 1–53 · score 0.70 · disordered residues, longest IDR, IDR content, MobiDB, aa, predictions
- [10] § STAR★Methods › Method details › RBP identification ↔ scripts/analyze_correlations.py, lines 1–38 · score 0.70 · AnimalTFDB, RBP gene, control protein, EuRBPDB, GPCR, Kinase
- [11] § Results › RBP family diversity strongly correlates with organismal complexity ↔ scripts/pgls_analysis.py, lines 1–51 · score 0.70 · phylogenetic signal, phylogenetic correction, correlation remained, PGLS, Pagel, optimal
- [12] § STAR★Methods › Quantification and statistical analysis ↔ scripts/pgls_analysis.py, lines 1–51 · score 0.69 · maximum likelihood, Spearman correlations, PGLS, Pagel, scipy, Phylogenetic
- [13] § STAR★Methods › Method details › IDR analysis ↔ scripts/fetch_idr_mobidb_v2.py, lines 44–115 · score 0.69 · disordered residues, MobiDB, disordered regions, IDR, consensus, longest
- [14] § STAR★Methods › Method details › RBP identification ↔ scripts/fetch_pfam_from_uniprot.py, lines 1–44 · score 0.69 · RBP.txt, UniProt, EuRBPDB, melanogaster, musculus, rerio
- [15] § STAR★Methods › Method details › RBP family classification ↔ scripts/fetch_pfam_from_uniprot.py, lines 1–44 · score 0.69 · Gene symbols, UniProt, EuRBPDB, Pfam domain, mapped, accessions
- [16] § STAR★Methods › Method details › LLPS analysis ↔ scripts/aggregate_llphyscore.py, lines 81–196 · score 0.67 · raw scores, LLPhyScore, aggregation, isoforms, LLPS, tropicalis
- [17] § Results › Vertebrate RBP diversification is driven by domain-level expansion beyond canonical neural RBPs ↔ scripts/supplementary_analyses.py, lines 1–50 · score 0.65 · expansion ratios, canonical RBP, Pfam domains, emergence, fold, enrichment
- [18] § Results › LLPS propensity is evolutionarily conserved ↔ scripts/extract_phasepred_v2.py, lines 1–14 · score 0.61 · PScore, catGRANULE, PhaSePred
- [19] § STAR★Methods › Method details › Orthogroup analysis ↔ scripts_of/util.py, lines 327–387 · score 0.61 · MCL clustering, OrthoFinder, DIAMOND, alignment, Orthogroups, sequences
- [20] § STAR★Methods › Quantification and statistical analysis ↔ scripts/statistical_validation.py, lines 1–50 · score 0.57 · Bootstrap resampling, validation, Cohen, scipy, numpy, correlations
- [21] § Results › Non-canonical RBPs and RBP-transcription factor overlap ↔ scripts/supplementary_analyses.py, lines 171–249 · score 0.53 · RBPs lacking, RBPWorld, canonical RBP, RBDs, correlated, species
- [22] § STAR★Methods › Method details › RBP family classification ↔ scripts/statistical_validation.py, lines 324–393 · score 0.53 · Pfam clan, Robustness, assignments, mapped, EuRBPDB, matching
- [23] § STAR★Methods › Method details › LLPS analysis ↔ scripts/extract_llphyscore.py, lines 16–58 · score 0.53 · raw scores, LLPhyScore, protein, species
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 469 lines · 19 KB · no license · 6 matches
- #!/usr/bin/env python3
- """
- Supplementary Analyses Script
- ===============================
- Reproduces supplementary analyses in Yasuda et al. (iScience, 2026)
- Analyses performed:
- 1. Non-canonical RBP analysis
- - Count non-canonical RBPs per species (RBPWorld)
- - Spearman correlation vs neurons
- - Combined canonical + non-canonical correlation
- 2. RBP–TF overlap
- - Intersect EuRBPDB gene symbols with AnimalTFDB gene symbols
- - Report overlap % per species
- 3. Vertebrate domain expansion
- - Compute vertebrate/invertebrate gene count ratio per Pfam domain
- - Identify domains with ≥3-fold enrichment or vertebrate-specific emergence
- Usage:
- python3 supplementary_analyses.py <data_dir> <output_dir>
- <data_dir> : directory containing input files (default: ../data/processed)
- <output_dir> : directory for output files (default: ../results)
- Input files required:
- new_rbp_db.csv - RBP database (EuRBPDB, canonical + non-canonical flag)
- new_tf_db.csv - TF database (AnimalTFDB 4.0)
- eurbpdb_pfam_mapping.tsv - gene-level Pfam assignments
- Output files:
- noncanonical_correlations.csv - non-canonical RBP counts and correlations
- rbp_tf_overlap.csv - RBP-TF overlap per species
- domain_expansion.csv - vertebrate expansion ratios per Pfam domain
- Author: Kyota Yasuda
- Date: March 2026
- """
- import sys
- import os
- import csv
- import math
- from collections import defaultdict
- try:
- from scipy import stats
- HAS_SCIPY = True
- except ImportError:
- HAS_SCIPY = False
- print("[WARN] scipy not found – using built-in Spearman implementation")
- # ── Species metadata ─────────────────────────────────────────────────────────
- SPECIES_ORDER = [
- "Caenorhabditis_elegans",
- "Drosophila_melanogaster",
- "Danio_rerio",
- "Xenopus_tropicalis",
- "Mus_musculus",
- "Homo_sapiens",
- ]
- # RBPWorld available for 5 species (Xenopus excluded)
- SPECIES_NONCANONICAL = [
- "Caenorhabditis_elegans",
- "Drosophila_melanogaster",
- "Danio_rerio",
- "Mus_musculus",
- "Homo_sapiens",
- ]
- NEURON_COUNT = {
- "Caenorhabditis_elegans": 302,
- "Drosophila_melanogaster": 200_000,
- "Danio_rerio": 10_000_000,
- "Xenopus_tropicalis": 16_000_000,
- "Mus_musculus": 71_000_000,
- "Homo_sapiens": 86_000_000_000,
- }
- # Vertebrate vs invertebrate classification
- INVERTEBRATES = {"Caenorhabditis_elegans", "Drosophila_melanogaster"}
- VERTEBRATES = {"Danio_rerio", "Xenopus_tropicalis", "Mus_musculus", "Homo_sapiens"}
- # ════════════════════════════════════════════════════════════════════════════
- # Spearman utilities
- # ════════════════════════════════════════════════════════════════════════════
- def _rank(lst):
- sorted_idx = sorted(range(len(lst)), key=lambda i: lst[i])
- ranks = [0.0] * len(lst)
- i = 0
- while i < len(lst):
- j = i
- while j < len(lst) - 1 and lst[sorted_idx[j]] == lst[sorted_idx[j + 1]]:
- j += 1
- avg_rank = (i + j) / 2.0 + 1
- for k in range(i, j + 1):
- ranks[sorted_idx[k]] = avg_rank
- i = j + 1
- return ranks
- def _norm_cdf(x):
- return (1.0 + math.erf(x / math.sqrt(2))) / 2.0
- def spearman_builtin(x, y):
- n = len(x)
- rx, ry = _rank(x), _rank(y)
- mx = sum(rx) / n
- my = sum(ry) / n
- num = sum((rx[i] - mx) * (ry[i] - my) for i in range(n))
- dx = math.sqrt(sum((rx[i] - mx) ** 2 for i in range(n)))
- dy = math.sqrt(sum((ry[i] - my) ** 2 for i in range(n)))
- if dx == 0 or dy == 0:
- return 0.0, 1.0
- rho = num / (dx * dy)
- if abs(rho) >= 1.0:
- return rho, 0.0
- t_stat = rho * math.sqrt(n - 2) / math.sqrt(1 - rho ** 2)
- p = 2 * (1 - _norm_cdf(abs(t_stat)))
- return rho, p
- def spearman(x, y):
- if HAS_SCIPY:
- r = stats.spearmanr(x, y)
- return float(r.statistic), float(r.pvalue)
- return spearman_builtin(x, y)
- def sig_label(p):
- if p < 0.001: return "***"
- if p < 0.01: return "**"
- if p < 0.05: return "*"
- return "ns"
- # ════════════════════════════════════════════════════════════════════════════
- # File loaders
- # ════════════════════════════════════════════════════════════════════════════
- def load_csv(path, delimiter=","):
- if not os.path.isfile(path):
- raise FileNotFoundError(f"Required file not found: {path}")
- with open(path, encoding="utf-8") as fh:
- rows = list(csv.DictReader(fh, delimiter=delimiter))
- print(f" Loaded {len(rows):,} rows ← {os.path.basename(path)}")
- return rows
- def load_tsv(path):
- return load_csv(path, delimiter="\t")
- def write_csv(path, rows, fieldnames):
- with open(path, "w", newline="", encoding="utf-8") as fh:
- w = csv.DictWriter(fh, fieldnames=fieldnames)
- w.writeheader()
- w.writerows(rows)
- print(f" Saved → {path}")
- # ════════════════════════════════════════════════════════════════════════════
- # 1. Non-canonical RBP analysis
- # ════════════════════════════════════════════════════════════════════════════
- def run_noncanonical(rbp_rows):
- """
- Uses is_canonical column in new_rbp_db.csv to separate canonical / non-canonical.
- Non-canonical = experimentally identified RBPs lacking canonical RBDs (from RBPWorld).
- Expected result (from manuscript):
- Non-canonical count vs neurons: rho=0.900, p=0.037 (5 species, excl. Xenopus)
- Combined (canonical + non-canonical) vs neurons: rho=0.975, p=0.005
- """
- # Count by species and canonical status
- canonical_counts = defaultdict(int)
- noncanonical_counts = defaultdict(int)
- for r in rbp_rows:
- sp = r.get("species", "").strip()
- can = r.get("is_canonical", "").strip().lower()
- if sp not in SPECIES_ORDER:
- continue
- if can in ("true", "1", "yes"):
- canonical_counts[sp] += 1
- elif can in ("false", "0", "no"):
- noncanonical_counts[sp] += 1
- # Check non-canonical data exists
- total_noncan = sum(noncanonical_counts.values())
- if total_noncan == 0:
- print(" [WARN] No non-canonical RBPs found in new_rbp_db.csv.")
- print(" Ensure is_canonical=False entries from RBPWorld are present.")
- # Correlation: non-canonical vs neurons (5 species, excl. Xenopus)
- sp5 = SPECIES_NONCANONICAL
- neurons5 = [NEURON_COUNT[sp] for sp in sp5]
- noncan5 = [noncanonical_counts[sp] for sp in sp5]
- rho_nc, p_nc = spearman(neurons5, noncan5)
- # Correlation: combined vs neurons (5 species)
- combined5 = [canonical_counts[sp] + noncanonical_counts[sp] for sp in sp5]
- rho_cb, p_cb = spearman(neurons5, combined5)
- print(f"\n Non-canonical RBP counts (5 species):")
- for sp in sp5:
- print(f" {sp:35s} canonical={canonical_counts[sp]:5d} "
- f"non-canonical={noncanonical_counts[sp]:5d}")
- print(f"\n Non-canonical vs neurons: rho={rho_nc:.3f} p={p_nc:.4f} "
- f"{sig_label(p_nc)} (expected: rho=0.900, p=0.037)")
- print(f" Combined vs neurons: rho={rho_cb:.3f} p={p_cb:.4f} "
- f"{sig_label(p_cb)} (expected: rho=0.975, p=0.005)")
- results = []
- for sp in SPECIES_ORDER:
- results.append({
- "species": sp,
- "neurons": NEURON_COUNT[sp],
- "canonical_count": canonical_counts[sp],
- "noncanonical_count": noncanonical_counts[sp],
- "combined_count": canonical_counts[sp] + noncanonical_counts[sp],
- "xenopus_excluded": sp not in SPECIES_NONCANONICAL,
- })
- # Append correlation summary
- results.append({
- "species": "CORRELATION: non-canonical vs neurons (n=5)",
- "neurons": "",
- "canonical_count": "",
- "noncanonical_count": round(rho_nc, 3),
- "combined_count": round(p_nc, 4),
- "xenopus_excluded": sig_label(p_nc),
- })
- results.append({
- "species": "CORRELATION: combined vs neurons (n=5)",
- "neurons": "",
- "canonical_count": "",
- "noncanonical_count": round(rho_cb, 3),
- "combined_count": round(p_cb, 4),
- "xenopus_excluded": sig_label(p_cb),
- })
- return results
- # ════════════════════════════════════════════════════════════════════════════
- # 2. RBP–TF overlap
- # ════════════════════════════════════════════════════════════════════════════
- def run_rbp_tf_overlap(rbp_rows, tf_rows):
- """
- Intersect EuRBPDB gene symbols with AnimalTFDB gene symbols per species.
- Expected result (from manuscript):
- Overlap ranges 2.6–8.3% (excluding zebrafish where zf-C2H2 inflates to 16.0%)
- Human: 194/2961 = 6.6% overlap (86 zf-C2H2)
- """
- # Build gene symbol sets
- rbp_genes = defaultdict(set)
- tf_genes = defaultdict(set)
- for r in rbp_rows:
- sp = r.get("species", "").strip()
- sym = r.get("gene_symbol", "").strip().upper()
- if sp and sym:
- rbp_genes[sp].add(sym)
- for r in tf_rows:
- sp = r.get("species", "").strip()
- sym = r.get("gene_symbol", "").strip().upper()
- if sp and sym:
- tf_genes[sp].add(sym)
- results = []
- print(f"\n {'Species':35s} {'RBPs':>6} {'TFs':>6} {'Overlap':>8} {'%':>7}")
- for sp in SPECIES_ORDER:
- rbp_set = rbp_genes[sp]
- tf_set = tf_genes[sp]
- overlap = rbp_set & tf_set
- n_rbp = len(rbp_set)
- n_tf = len(tf_set)
- n_ov = len(overlap)
- pct = 100.0 * n_ov / n_rbp if n_rbp > 0 else 0.0
- note = ""
- if sp == "Danio_rerio":
- note = "inflated by zf-C2H2 multi-annotation"
- print(f" {sp:35s} {n_rbp:>6} {n_tf:>6} {n_ov:>8} {pct:>6.1f}%"
- + (f" [{note}]" if note else ""))
- results.append({
- "species": sp,
- "n_rbps": n_rbp,
- "n_tfs": n_tf,
- "n_overlap": n_ov,
- "overlap_pct": round(pct, 1),
- "note": note,
- })
- return results
- # ════════════════════════════════════════════════════════════════════════════
- # 3. Vertebrate domain expansion
- # ════════════════════════════════════════════════════════════════════════════
- def run_domain_expansion(pfam_rows):
- """
- Compute mean gene count per Pfam domain family in invertebrates vs vertebrates.
- Identify domains with ≥3-fold enrichment or vertebrate-specific emergence.
- Expected (from manuscript Figure 3A):
- PARP domain: ~8.9-fold
- YTH domain: ~8.0-fold
- RAP domain: ~11.0-fold
- RNase A domain: vertebrate-specific (mammalian)
- OAS domain: vertebrate-specific
- """
- # Count genes per (family, species)
- family_species_count = defaultdict(lambda: defaultdict(int))
- for row in pfam_rows:
- sp = row.get("species", "").strip()
- family = row.get("eurbpdb_family", "").strip()
- if sp not in SPECIES_ORDER:
- continue
- if not family or family == "Non-canonical":
- continue
- family_species_count[family][sp] += 1
- results = []
- for family, sp_counts in family_species_count.items():
- invert_counts = [sp_counts.get(sp, 0) for sp in INVERTEBRATES]
- vert_counts = [sp_counts.get(sp, 0) for sp in VERTEBRATES]
- mean_invert = sum(invert_counts) / len(invert_counts)
- mean_vert = sum(vert_counts) / len(vert_counts)
- total_genes = sum(sp_counts.values())
- n_species = len(sp_counts)
- # Expansion ratio
- if mean_invert == 0 and mean_vert > 0:
- ratio = float("inf")
- expansion_type = "vertebrate-specific"
- elif mean_invert == 0:
- ratio = 1.0
- expansion_type = "absent"
- else:
- ratio = mean_vert / mean_invert
- if ratio >= 3.0:
- expansion_type = "vertebrate-expanded (>=3x)"
- elif ratio >= 2.0:
- expansion_type = "vertebrate-enriched (2-3x)"
- else:
- expansion_type = "conserved"
- results.append({
- "pfam_family": family,
- "n_species": n_species,
- "total_genes": total_genes,
- "mean_invertebrate": round(mean_invert, 2),
- "mean_vertebrate": round(mean_vert, 2),
- "vert_invert_ratio": round(ratio, 2) if ratio != float("inf") else "inf",
- "expansion_type": expansion_type,
- "worm": sp_counts.get("Caenorhabditis_elegans", 0),
- "fly": sp_counts.get("Drosophila_melanogaster", 0),
- "zfish": sp_counts.get("Danio_rerio", 0),
- "frog": sp_counts.get("Xenopus_tropicalis", 0),
- "mouse": sp_counts.get("Mus_musculus", 0),
- "human": sp_counts.get("Homo_sapiens", 0),
- })
- # Sort by ratio descending (inf first)
- def sort_key(r):
- v = r["vert_invert_ratio"]
- return -float("inf") if v == "inf" else -float(v)
- results.sort(key=sort_key)
- # Print top expanded
- print(f"\n Top vertebrate-expanded Pfam domain families (≥3-fold or vertebrate-specific):")
- print(f" {'Family':30s} {'ratio':>8} {'type':35s} "
- f"{'worm':>5} {'fly':>5} {'zfish':>6} {'frog':>5} {'mouse':>6} {'human':>6}")
- for r in results:
- if r["expansion_type"] in ("vertebrate-specific", "vertebrate-expanded (>=3x)"):
- ratio_str = str(r["vert_invert_ratio"])
- print(f" {r['pfam_family']:30s} {ratio_str:>8} "
- f"{r['expansion_type']:35s} "
- f"{r['worm']:>5} {r['fly']:>5} {r['zfish']:>6} "
- f"{r['frog']:>5} {r['mouse']:>6} {r['human']:>6}")
- return results
- # ════════════════════════════════════════════════════════════════════════════
- # Main
- # ════════════════════════════════════════════════════════════════════════════
- def main():
- data_dir = sys.argv[1] if len(sys.argv) > 1 else os.path.join(
- os.path.dirname(__file__), "..", "data", "processed")
- output_dir = sys.argv[2] if len(sys.argv) > 2 else os.path.join(
- os.path.dirname(__file__), "..", "results")
- data_dir = os.path.abspath(data_dir)
- output_dir = os.path.abspath(output_dir)
- os.makedirs(output_dir, exist_ok=True)
- print("=" * 60)
- print("Supplementary Analyses")
- print(f" data_dir : {data_dir}")
- print(f" output_dir : {output_dir}")
- print("=" * 60)
- # ── Load data ─────────────────────────────────────────────────────────
- print("\n[0] Loading data …")
- rbp_rows = load_csv(os.path.join(data_dir, "new_rbp_db.csv"))
- tf_rows = load_csv(os.path.join(data_dir, "new_tf_db.csv"))
- pfam_rows = load_tsv(os.path.join(data_dir, "eurbpdb_pfam_mapping.tsv"))
- # ── 1. Non-canonical ──────────────────────────────────────────────────
- print("\n[1] Non-canonical RBP analysis …")
- nc_results = run_noncanonical(rbp_rows)
- write_csv(
- os.path.join(output_dir, "noncanonical_correlations.csv"),
- nc_results,
- ["species", "neurons", "canonical_count",
- "noncanonical_count", "combined_count", "xenopus_excluded"],
- )
- # ── 2. RBP-TF overlap ────────────────────────────────────────────────
- print("\n[2] RBP–TF overlap analysis …")
- overlap_results = run_rbp_tf_overlap(rbp_rows, tf_rows)
- write_csv(
- os.path.join(output_dir, "rbp_tf_overlap.csv"),
- overlap_results,
- ["species", "n_rbps", "n_tfs", "n_overlap", "overlap_pct", "note"],
- )
- # ── 3. Domain expansion ───────────────────────────────────────────────
- print("\n[3] Vertebrate domain expansion analysis …")
- expansion_results = run_domain_expansion(pfam_rows)
- write_csv(
- os.path.join(output_dir, "domain_expansion.csv"),
- expansion_results,
- ["pfam_family", "n_species", "total_genes",
- "mean_invertebrate", "mean_vertebrate", "vert_invert_ratio",
- "expansion_type", "worm", "fly", "zfish", "frog", "mouse", "human"],
- )
- print("\nDone.")
- print("\nExpected results (from manuscript):")
- print(" Non-canonical vs neurons (n=5): rho=0.900, p=0.037 *")
- print(" Combined vs neurons (n=5): rho=0.975, p=0.005 **")
- print(" RBP-TF overlap: 2.6–8.3% (excl. zebrafish 16.0%)")
- print(" Human overlap: 194/2961 = 6.6%")
- print(" Top expanded domains: RAP (11x), PARP (8.9x), YTH (8x)")
- if __name__ == "__main__":
- main()
supplementary_analyses.py at commit 66dafdc, no license · at the source
Overview
- Graduate School of Integrated Sciences for Life, Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8526, Japan
- International Institute for Sustainability with Knotted Chiral Meta Matter (SKCM2), Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8526, Japan
- Research Center for the Mathematics on Chromatin Live Dynamics (RcMcD), Hiroshima University, 1-3-1 Kagamiyama, Higashi-Hiroshima, Hiroshima 739-8531, Japan
Abstract
RNA-binding proteins (RBPs) control post-transcriptional gene expression with critical roles in nervous system function. Here, we ask whether RBP family diversity tracks with neural complexity across animal evolution. Comparing six species ranging from a simple worm to humans, we find that the number of distinct RBP domain families increases progressively with neuron number—a relationship specific to RBPs and not seen for kinases or G-protein-coupled receptors. Transcription factors show a different pattern, reaching a ceiling in vertebrates while RBP families continue expanding. We further show that the disordered structural regions of RBPs grow longer in more complex organisms, whereas liquid-liquid phase separation propensity remains broadly conserved. Finally, the lengthening of messenger RNA regulatory tails parallels RBP diversification, suggesting co-evolution of post-transcriptional regulatory capacity. These findings establish RBP family diversification as a distinctive molecular signature of animal neural complexity.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repositories
Its files are read in the Code ↔ Paper reader above, with 23 matches between paragraphs and lines of code.
Zenodo 18002633
Availability: 1 check, the latest on 29 September 2026: the link answers (HTTP 200)
- 29 September 2026: the link answers (HTTP 200)
12 files
- scripts/
aggregate_llphyscore.py , Python, 200 lines - scripts/
analyze_correlations.py , Python, 377 lines - scripts/
extract_llphyscore.py , Python, 107 lines - scripts/
extract_phasepred_v2.py , Python, 124 lines - scripts/
fetch_controls_pfam.py , Python, 168 lines - scripts/
fetch_idr_mobidb_v2.py , Python, 197 lines - scripts/
fetch_pfam_from_uniprot. , Python, 225 linespy - scripts/
idr_analysis.py , Python, 384 lines - scripts/
pgls_analysis.py , Python, 378 lines - scripts/
statistical_validation.p , Python, 473 linesy - scripts/
supplementary_analyses.p , Python, 469 linesy - README.md, Text, 207 lines
Kyotay-12/rbp-evolution-analysis
66dafdcaad56c2c9bb26baf69f007002759fa743, 9 March 2026Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
12 files
- scripts/
aggregate_llphyscore.py , Python, 200 lines, 1 match - scripts/
analyze_correlations.py , Python, 377 lines, 2 matches - scripts/
extract_llphyscore.py , Python, 107 lines, 1 match - scripts/
extract_phasepred_v2.py , Python, 124 lines, 1 match - scripts/
fetch_controls_pfam.py , Python, 168 lines - scripts/
fetch_idr_mobidb_v2.py , Python, 197 lines, 2 matches - scripts/
fetch_pfam_from_uniprot. , Python, 225 lines, 2 matchespy - scripts/
idr_analysis.py , Python, 384 lines, 2 matches - scripts/
pgls_analysis.py , Python, 378 lines, 2 matches - scripts/
statistical_validation.p , Python, 473 lines, 3 matchesy - scripts/
supplementary_analyses.p , Python, 469 lines, 6 matchesy - README.md, Text, 207 lines
davidemms/OrthoFinder
43f9bc7273a33d2e7a4e8d1c55b5117ed7f08a42, 15 July 2025Availability: 1 check, the latest on 29 September 2026: the link answers
- 29 September 2026: the link answers
46 files
- orthofinder.py, Python, 7 lines
- scripts_of/
__init__.py , Python, 1 line - scripts_of/
__main__.py , Python, 1,430 lines - scripts_of/
accelerate.py , Python, 384 lines - scripts_of/
astral.py , Python, 27 lines - scripts_of/
bin/ , Perl, 464 linesmafft/ libexec/ mafftash_premafft.pl - scripts_of/
bin/ , Perl, 600 linesmafft/ libexec/ seekquencer_premafft.pl - scripts_of/
blast_file_processor.py , Python, 104 lines - scripts_of/
consensus_tree.py , Python, 292 lines - scripts_of/
fasta_writer.py , Python, 94 lines - scripts_of/
files.py , Python, 869 lines - scripts_of/
gathering.py , Python, 561 lines - scripts_of/
matrices.py , Python, 78 lines - scripts_of/
mcl.py , Python, 296 lines - scripts_of/
newick.py , Python, 431 lines - scripts_of/
orthologues.py , Python, 1,265 lines - scripts_of/
parallel_task_manager.py , Python, 368 lines - scripts_of/
probroot.py , Python, 461 lines - scripts_of/
program_caller.py , Python, 400 lines - scripts_of/
resolve.py , Python, 488 lines - scripts_of/
sample_genes.py , Python, 295 lines - scripts_of/
split_ortholog_files.py , Python, 51 lines - scripts_of/
stag.py , Python, 298 lines - scripts_of/
stats.py , Python, 235 lines - scripts_of/
stride.py , Python, 745 lines - scripts_of/
tree.py , Python, 1,858 lines - scripts_of/
trees2ologs_dlcpar.py , Python, 300 lines - scripts_of/
trees2ologs_of.py , Python, 1,591 lines - scripts_of/
trees_msa.py , Python, 426 lines - scripts_of/
trim.py , Python, 201 lines - scripts_of/
util.py , Python, 530 lines, 1 match - scripts_of/
wrapper_phyldog.py , Python, 223 lines - setup.py, Python, 101 lines
- tests/
TestArgumentCombinations , Python, 126 lines.py - tests/
test_consensus_tree.py , Python, 118 lines - tests/
test_hog_writer.py , Python, 209 lines - tests/
test_orthofinder.py , Python, 1,088 lines - tests/
test_program_caller.py , Python, 216 lines - tools/
__init__.py , Python, 1 line - tools/
convert_orthofinder_tree , Python, 79 lines_ids.py - tools/
create_files_for_hogs.py , Python, 241 lines - tools/
make_ultrametric.py , Python, 73 lines - tools/
orthogroup_gene_count.py , Python, 20 lines - tools/
primary_transcript.py , Python, 219 lines - License.md, License, 674 lines
- README.md, Text, 520 lines
The paper's code and data availability statement is in the Data section.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 3 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 66 scripts, each with its path and the digest of its content;
- 23 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
Datasets cited
- uniprot.org/
help/ , at UniProt; found in the resources tableapi
Data and code availability
Data: All data reported in this paper are derived from publicly available databases (EuRBPDB, RBPWorld, AnimalTFDB 4.0, MobiDB, Ensembl BioMart, UniProt, PhaSePred, LLPhyScore). Processed datasets supporting this study are publicly available at Zenodo (DOI: https://
Code: All analysis scripts used in this study are publicly available at GitHub (https://
Other items: Any additional information required to reanalyze the data reported in this paper is available from the lead contact upon request.
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 29 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 1 author, 2 keywords, 2 funders, 53 references, 6 RRIDs.
Cite
This paper
Yasuda, K. (2026). RNA-binding protein family diversification correlates with neural complexity across metazoan evolution. iScience, 29(5), 115766. https://
BibTeX
@article{yasuda2026rna,
author = {Yasuda, Kyota},
title = {{RNA-binding protein family diversification correlates with neural complexity across metazoan evolution}},
journal = {iScience},
year = {2026},
month = apr,
volume = {29},
number = {5},
pages = {115766},
publisher = {Elsevier},
issn = {2589-0042},
doi = {10.1016/
url = {https://
pmid = {42111166},
pmcid = {PMC13157009}
}
RIS
TY - JOUR
AU - Yasuda, Kyota
TI - RNA-binding protein family diversification correlates with neural complexity across metazoan evolution
T2 - iScience
J2 - iScience
PY - 2026
DA - 2026/
VL - 29
IS - 5
SP - 115766
SN - 2589-0042
PB - Elsevier
DO - 10.1016/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1016/
"type": "article-journal",
"title": "RNA-binding protein family diversification correlates with neural complexity across metazoan evolution",
"container-title": "iScience",
"author": [
{
"family": "Yasuda",
"given": "Kyota"
}
],
"container-title-short":
"volume": "29",
"issue": "5",
"page": "115766",
"DOI": "10.1016/
"PMID": "42111166",
"PMCID": "PMC13157009",
"ISSN": "2589-0042",
"publisher": "Elsevier",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
17
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1038/s41586-026-10629-x [code]
- Whole-genome duplication shaped cell-type evolution in the vertebrate brain.Journal: NatureIn common: scikit-learn, pandas, SciPy, 1 other tool, cellular / molecular, 4 references
- [2] doi:10.1523/eneuro.0362-25.2026 [code]
- Similarities between &
lt;i& gt;Ciona& lt;/ i& gt; Dorsal Motor Ganglion and Vertebrate Cerebellum: Did a Chordate Ancestor Already Show D/ V Subdivision within a Hindbrain Precursor? Journal: eNeuroIn common: Biopython, scikit-learn, pandas, 2 other tools, 2 references - [3] doi:10.3389/fgene.2026.1742595 [code]
- An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis.Journal: Frontiers in geneticsIn common: scikit-learn, pandas, NumPy, 4 references
- [4] doi:10.1038/s41467-026-75700-7 [code]
- Gene regulatory innovations from transposable elements in primate cerebellum development.Journal: Nature communicationsIn common: Biopython, scikit-learn, pandas, 2 other tools, cellular / molecular, 1 reference
- [5] doi:10.1002/advs.202523984 [code]
- INB&
lt;sup& gt;3& lt;/ sup& gt;P: A Multi-Modal and Interpretable Co-Attention Framework Integrating Property-Aware Explanations and Memory-Bank Contrastive Fusion for Blood-Brain Barrier Penetrating Peptide Discovery. Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)In common: Biopython, scikit-learn, pandas, 2 other tools, 1 reference - [6] doi:10.1038/s41467-026-76045-x [code]
- A manufacturability-inform
ed topology framework for AI-guided design of fibrous network materials. Journal: Nature communicationsIn common: scikit-learn, pandas, SciPy, 1 other tool, 3 references - [7] doi:10.1093/narmme/ugag028 [code]
- MCVAE-based multi-omic anomaly detection in Fragile X Syndrome.Journal: NAR molecular medicineIn common: scikit-learn, pandas, SciPy, 1 other tool, cellular / molecular, 2 references
- [8] doi:10.1371/journal.pbio.3003856 [code]
- Aging and metabolism contribute separately to brain-body health.Journal: PLoS biologyIn common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
- [9] doi:10.1038/s41597-026-07077-7 [code]
- Everyday Activity Science and Engineering Table Setting Dataset.Journal: Scientific dataIn common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
- [10] doi:10.1016/j.crmeth.2026.101421 [code]
- EthoPy provides an accessible platform for reproducible behavioral neuroscience.Journal: Cell reports methodsIn common: scikit-learn, pandas, SciPy, 1 other tool, 3 references
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 3 repositories of the authors' code, each at its verified commit and with its license, 66 scripts, and 23 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:076be9e3aa2807e0…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
